IncidentWatchdog

Severity: none · Fires after: 0m · Rule: central/prometheus/rules/watchdog.yml

What it means

Nothing is wrong. This alert is always firing by design. It is the dead-man's switch for the alerting pipeline itself.

vector(1)

It routes to a dedicated watchdog receiver which POSTs to an external healthchecks.io check every 5 minutes (repeat_interval: 5m, which must stay below the external check's period plus grace).

The alert you actually care about

You will never be paged by this rule. You get paged by healthchecks.io when it stops arriving. That means one of:

In other words: the entire local alerting path is dead, so no local alert could possibly tell you.

First checks

curl -s http://192.168.0.5:9090/-/healthy
ssh jakefear@docker2 'curl -s localhost:9093/-/healthy'
ssh jakefear@docker2 'docker compose -f /opt/jakemon/docker-compose.yml ps'

# Is the webhook actually being delivered?
curl -sG 'http://192.168.0.5:9090/api/v1/query' \
  --data-urlencode 'query=alertmanager_notifications_failed_total'

The watchdog URL is a secret file at central/alertmanager/secrets/watchdog_url (gitignored), referenced by url_file: in alertmanager.yml.

How to clear

Restore the stack. The heartbeat resumes by itself within one repeat_interval.

bin/deploy-central.sh

Notes

If you are deliberately taking the fleet dark (a move, an extended outage), pause the healthchecks.io check rather than letting it alarm for weeks. Un-pause it as the first step of bring-up, before you trust any other alert.