IncidentPostgresDown

Severity: critical · Fires after: 2m · Rule: central/prometheus/rules/service.yml

What it means

The Postgres exporter is reachable (up=1) but cannot connect to the database. This is strictly more informative than ServiceDown: the metrics path is healthy, so the database itself is the problem.

The DB backs Wikantik on docker1.

Expression

pg_up == 0

First checks

ssh jakefear@docker1 'docker ps --filter name=db'
ssh jakefear@docker1 'docker logs --tail 100 repo-db-1'

# Is it accepting connections on the published host port?
ssh jakefear@docker1 'ss -lntp | grep 5432'

The DSN lives in remote.env as PG_EXPORTER_DSN_<host> and is baked into the agent at deploy time. If the credentials or port changed, the exporter will sit at pg_up=0 indefinitely while the database is perfectly healthy — check that before assuming a DB outage.

How to clear

Restore the database, or correct PG_EXPORTER_DSN_docker1 in remote.env and redeploy:

bin/deploy-agent.sh docker1

Notes

The Wikantik DB must be published on a host port for the agent's embedded exporter to reach it — it connects through the host, not the compose network. A placeholder DSN (postgresql://disabled@disabled.invalid:...) is used when no real one is set; that produces pg_up=0 too, so on a host where DB metrics were never configured this alert is expected rather than an incident.