IncidentPiholeDnsDown

Severity: critical · Fires after: 3m · Rule: central/prometheus/rules/edge.yml

What it means

Pi-hole on cloudflare (192.168.0.3) has stopped answering DNS. It is the LAN's resolver, so name resolution is broken for every host — including this stack's own scrapes and remote-writes.

Expression

probe_success{job="integrations/blackbox/pihole-dns"} == 0

Probed from docker2, deliberately: that means a network path failure or a host failure trips it too, not just an FTL crash. 3 minutes avoids paging on a single dropped UDP packet.

Why critical

Nothing on the LAN works properly without DNS, and the failure presents as everything being slow or intermittently broken rather than as a DNS error. Expect a wave of other alerts to follow if this is not fixed quickly.

First checks

dig @192.168.0.3 wikantik.com +short
ssh jakefear@cloudflare 'systemctl status pihole-FTL'
ssh jakefear@cloudflare 'sudo journalctl -u pihole-FTL --since "15 min ago" --no-pager | tail -40'
ssh jakefear@cloudflare 'ss -lnup | grep :53'

How to clear

ssh jakefear@cloudflare 'sudo systemctl restart pihole-FTL'

If the host itself is unreachable, this is really IncidentHostNotReporting for cloudflare, and DNS is one of several things you have lost.

Fallback

While it is down, resolution can be pointed at a public resolver on the affected host to unblock work:

# temporary, on one host only
echo 'nameserver 1.1.1.1' | sudo tee /etc/resolv.conf

Remember to revert — Pi-hole's filtering is doing real work.

Notes

Added 2026-08 (feat(dns): probe Pi-hole DNS and alert when the LAN loses resolution). cloudflare is the LAN's single resolver, so this is a single point of failure by design; a second resolver would be the structural fix.