Severity: critical · Fires after: 5m · Rule: central/prometheus/rules/service.yml
A scrape target that is expected to be up has been unreachable for 5 minutes. This is the fleet's broadest "something that should be running isn't" signal.
up{job!~"prometheus|ollama|cf-analytics-exporter|crafter-exporter|visibility-exporter|visibility-shipper|nginx|gpu|ollama-exporter"} == 0
Every name in that regex is exact, so e.g. cf-analytics-exporter does not also
exclude cloudflared.
| Excluded | Reason | Covered instead by |
|---|---|---|
prometheus | cannot alert on its own absence | external Watchdog |
ollama | no /metrics, so no up series | blackbox probe |
cf-analytics-exporter | non-critical analytics poller | — |
crafter-exporter | dashboards flatten, sites still serve | CrafterVhostDown |
visibility-exporter / visibility-shipper | nightly SEO pipeline | VisibilityShipper* |
nginx | edge stub_status internals only | EdgeVhostDown |
gpu | GPU visibility only, ollama keeps serving | GpuExporterDown |
ollama-exporter | model-state poller | blackbox probe |
cloudflared is covered here — a down edge tunnel pages.
The excluded exporters are not invisible: they are all on the reliability
dashboard's scrape/target-health table, which is the one place to see them.
# Exactly what is down
curl -sG 'http://192.168.0.5:9090/api/v1/query' --data-urlencode 'query=up == 0'
ssh jakefear@<host> 'docker ps -a'
ssh jakefear@<host> 'docker logs --tail 100 <container>'
ssh jakefear@<host> 'curl -s localhost:12345/-/ready' # is the agent itself ok?
If many jobs on one host are down at once, this is really a host problem — see IncidentHostNotReporting.
Restart the service, or redeploy the host's agent if the agent is the thing that died:
bin/deploy-agent.sh <host>
An empty or unroutable Postgres DSN will crash-loop the whole agent, taking
every job on that host down with it, not just postgres. deploy-agent.sh falls
back to a parseable placeholder DSN to prevent this. See CLAUDE.md gotcha #3.