Severity: warning · Fires after: 5m · Rule: central/prometheus/rules/host.yml
A systemd unit is stuck in the failed state. {{ $labels.name }} names it.
node_systemd_unit_state{state="failed"} == 1
unless on(host, name) (
node_systemd_unit_state{host="docker1", name="fwupd-refresh.service"}
or
node_systemd_unit_state{host="nas", name=~"isnsdd.service|krb5-admin-server.service|krb5-kdc.service|rtslib-fb-targetctl.service|smbftpd.service|targetclid.service|targetclid.socket"}
)
Alerts per unit — off the per-unit metric rather than a count summary — so the unit name is visible in the alert instead of just "3 units failed".
The unless clause excludes units that are expected to be failed: a cosmetic
oneshot on docker1, and the NAS appliance's unconfigured vendor services.
It is keyed on(host, name), so the same unit name failing on a host not
listed still fires. That is deliberate — krb5-kdc.service being failed on the
NAS appliance is normal; it failing on docker1 would not be.
Regex dots are unescaped (. matches the literal dot too); no real unit collides
with these names.
ssh jakefear@<host> 'systemctl --failed'
ssh jakefear@<host> 'systemctl status <unit>'
ssh jakefear@<host> 'journalctl -u <unit> --since "1 hour ago" --no-pager | tail -50'
Fix the unit and systemctl reset-failed <unit>.
If the unit is genuinely expected to fail on that host forever, add it to the
allowlist above — with the host scoped — rather than muting the whole alert.
Changing the rule requires a bin/validate.sh run and a bin/deploy-central.sh.
Added 2026-07-02. Unit-tested in central/prometheus/rules/tests/host_test.yml,
including the allowlist behaviour — if you edit the matcher, that test is what
tells you whether you widened it too far.