Severity: none · Fires after: 0m · Rule: central/prometheus/rules/watchdog.yml
Nothing is wrong. This alert is always firing by design. It is the dead-man's switch for the alerting pipeline itself.
vector(1)
It routes to a dedicated watchdog receiver which POSTs to an external
healthchecks.io check every 5 minutes (repeat_interval: 5m, which must stay
below the external check's period plus grace).
You will never be paged by this rule. You get paged by healthchecks.io when it stops arriving. That means one of:
docker2) is off or off-networkIn other words: the entire local alerting path is dead, so no local alert could possibly tell you.
curl -s http://192.168.0.5:9090/-/healthy
ssh jakefear@docker2 'curl -s localhost:9093/-/healthy'
ssh jakefear@docker2 'docker compose -f /opt/jakemon/docker-compose.yml ps'
# Is the webhook actually being delivered?
curl -sG 'http://192.168.0.5:9090/api/v1/query' \
--data-urlencode 'query=alertmanager_notifications_failed_total'
The watchdog URL is a secret file at
central/alertmanager/secrets/watchdog_url (gitignored), referenced by
url_file: in alertmanager.yml.
Restore the stack. The heartbeat resumes by itself within one repeat_interval.
bin/deploy-central.sh
If you are deliberately taking the fleet dark (a move, an extended outage), pause the healthchecks.io check rather than letting it alarm for weeks. Un-pause it as the first step of bring-up, before you trust any other alert.