IncidentCrafterVhostDown

Severity: critical · Fires after: 5m · Rule: central/prometheus/rules/crafter.yml

What it means

A CrafterCMS-delivered site is failing its end-to-end HTTP probe. User-facing: jakefear.com or maiiavorobiova.com is down or erroring.

Expression

probe_success{job=~"integrations/blackbox/(jakefear|maiiavorobiova).com"} == 0

Probes run from the docker2 agent against the delivery nginx proxy (:8082), which fronts a single Engine serving both vhosts. Each vhost needs its own blackbox module carrying its Host header and a unique address, or the exporter dedups them.

This is the alert that matters

crafter-exporter is excluded from ServiceDown precisely because losing platform visibility is not the same as losing the site. This rule is the real "are the sites serving?" signal. If this is green and IncidentCrafterComponentDown is red, users are fine and you are chasing a monitoring problem.

First checks

# Through the proxy, from the host
ssh jakefear@docker2 'curl -s -o /dev/null -w "%{http_code}\n" -H "Host: jakefear.com" http://localhost:8082/'

ssh jakefear@docker2 'docker ps --filter name=personal-site-delivery'
ssh jakefear@docker2 'docker logs --tail 100 <delivery-engine-container>'

Then externally, to separate origin failure from edge failure:

curl -s -o /dev/null -w "%{http_code}\n" https://jakefear.com/

If the origin is healthy but the public URL is not, the problem is the tunnel — see IncidentCloudflaredConfigChanged.

How to clear

Restore the delivery stack. The probe recovers within one 60s scrape.

Notes

⚠️ The delivery Engine pins craftercms/delivery_tomcat:latest. A restart can therefore pull a new major version unannounced — that is exactly what happened on 2026-06-13, when Crafter 5.0 arrived under that tag and broke the Monitoring API routes. The site itself stayed up and this alert never fired, which was the tell that the failure was in monitoring, not delivery.