IncidentServiceDown

Severity: critical · Fires after: 5m · Rule: central/prometheus/rules/service.yml

What it means

A scrape target that is expected to be up has been unreachable for 5 minutes. This is the fleet's broadest "something that should be running isn't" signal.

Expression

up{job!~"prometheus|ollama|cf-analytics-exporter|crafter-exporter|visibility-exporter|visibility-shipper|nginx|gpu|ollama-exporter"} == 0

What is deliberately excluded, and why

Every name in that regex is exact, so e.g. cf-analytics-exporter does not also exclude cloudflared.

ExcludedReasonCovered instead by
prometheuscannot alert on its own absenceexternal Watchdog
ollamano /metrics, so no up seriesblackbox probe
cf-analytics-exporternon-critical analytics poller—
crafter-exporterdashboards flatten, sites still serveCrafterVhostDown
visibility-exporter / visibility-shippernightly SEO pipelineVisibilityShipper*
nginxedge stub_status internals onlyEdgeVhostDown
gpuGPU visibility only, ollama keeps servingGpuExporterDown
ollama-exportermodel-state pollerblackbox probe

cloudflared is covered here — a down edge tunnel pages.

The excluded exporters are not invisible: they are all on the reliability dashboard's scrape/target-health table, which is the one place to see them.

First checks

# Exactly what is down
curl -sG 'http://192.168.0.5:9090/api/v1/query' --data-urlencode 'query=up == 0'

ssh jakefear@<host> 'docker ps -a'
ssh jakefear@<host> 'docker logs --tail 100 <container>'
ssh jakefear@<host> 'curl -s localhost:12345/-/ready'   # is the agent itself ok?

If many jobs on one host are down at once, this is really a host problem — see IncidentHostNotReporting.

How to clear

Restart the service, or redeploy the host's agent if the agent is the thing that died:

bin/deploy-agent.sh <host>

Notes

An empty or unroutable Postgres DSN will crash-loop the whole agent, taking every job on that host down with it, not just postgres. deploy-agent.sh falls back to a parseable placeholder DSN to prevent this. See CLAUDE.md gotcha #3.