IncidentHostSystemdUnitFailed

Severity: warning · Fires after: 5m · Rule: central/prometheus/rules/host.yml

What it means

A systemd unit is stuck in the failed state. {{ $labels.name }} names it.

Expression

node_systemd_unit_state{state="failed"} == 1
unless on(host, name) (
    node_systemd_unit_state{host="docker1", name="fwupd-refresh.service"}
  or
    node_systemd_unit_state{host="nas", name=~"isnsdd.service|krb5-admin-server.service|krb5-kdc.service|rtslib-fb-targetctl.service|smbftpd.service|targetclid.service|targetclid.socket"}
)

Alerts per unit — off the per-unit metric rather than a count summary — so the unit name is visible in the alert instead of just "3 units failed".

The allowlist

The unless clause excludes units that are expected to be failed: a cosmetic oneshot on docker1, and the NAS appliance's unconfigured vendor services.

It is keyed on(host, name), so the same unit name failing on a host not listed still fires. That is deliberate — krb5-kdc.service being failed on the NAS appliance is normal; it failing on docker1 would not be.

Regex dots are unescaped (. matches the literal dot too); no real unit collides with these names.

First checks

ssh jakefear@<host> 'systemctl --failed'
ssh jakefear@<host> 'systemctl status <unit>'
ssh jakefear@<host> 'journalctl -u <unit> --since "1 hour ago" --no-pager | tail -50'

How to clear

Fix the unit and systemctl reset-failed <unit>.

If the unit is genuinely expected to fail on that host forever, add it to the allowlist above — with the host scoped — rather than muting the whole alert. Changing the rule requires a bin/validate.sh run and a bin/deploy-central.sh.

Notes

Added 2026-07-02. Unit-tested in central/prometheus/rules/tests/host_test.yml, including the allowlist behaviour — if you edit the matcher, that test is what tells you whether you widened it too far.