Severity: critical · Fires after: 5m · Rule: central/prometheus/rules/edge.yml
An edge vhost is failing its end-to-end HTTP probe against the origin nginx on
cloudflare. User-facing: the site is down or erroring.
probe_success{job=~"integrations/blackbox/(www\\.wikantik\\.com|wikantik\\.com)"} == 0
The probes are defined in hosts/cloudflare/apps.alloy as
prometheus.exporter.blackbox "edge_sites". Each vhost needs its own module
carrying its Host header and a unique address (via a harmless ?probe=
param) — the blackbox exporter keys targets by address and will silently dedup
two targets that share one.
This probes the origin, not Cloudflare. A green probe with users reporting errors means the problem is between Cloudflare and the origin — check IncidentCloudflaredConfigChanged and the tunnel journal instead.
The nginx stub_status exporter being down does not trigger this; that is a
separate non-critical poller.
# From the edge host itself, bypassing everything upstream
ssh jakefear@cloudflare 'curl -s -o /dev/null -w "%{http_code}\n" -H "Host: wikantik.com" http://localhost:8000/'
ssh jakefear@cloudflare 'sudo nginx -t && systemctl status nginx'
ssh jakefear@cloudflare 'tail -50 /var/log/nginx/error.log'
# What the probe sees
curl -sG 'http://192.168.0.5:9090/api/v1/query' --data-urlencode 'query=probe_success == 0'
Per-vhost traffic detail is in Loki, from the dedicated edgemon JSON access
log:
{host="cloudflare", job="edge-nginx"} | json | vhost="wikantik.com"
Restore nginx or the content it serves. The probe recovers within one scrape interval (60s).
The edgemon log format is not bind-mounted by jakemon — it is applied on
the host with hosts/cloudflare/nginx/reference/apply-edgemon.sh (run as root).
If the Loki stream is empty after an nginx reinstall, that script needs
re-running.
Adding a new site: traffic and referral data flow in automatically, but the
blackbox module, this rule's matcher, and the dashboard's $vhost value are
three manual edits.