IncidentHostNotReporting

Severity: critical · Fires after: 0m · Rule: central/prometheus/rules/host.yml

What it means

A 24/7 host has stopped shipping metrics entirely. Not "a service is down" — the host's whole integrations/unix series is gone.

Expression

One rule per covered host, because this is an absent() check and absence has no labels of its own to group by:

absent(up{job="integrations/unix", host="docker2"})
absent(up{job="integrations/unix", host="docker1"})
absent(up{job="integrations/unix", host="cloudflare"})

absent() rather than == 0 because the fleet is push-based: a dead host does not report up=0, its series simply stops arriving and ages out. There is nothing to compare against zero.

Who is covered, and who is not

HostCoveredWhy
docker2yescentral host; catches its agent dying while the stack lives
docker1yes24/7 server
cloudflareyes24/7 edge
minisforumnoworkstation, normal shutdowns must not page
inferencenoremoved 2026-07-21, boxed for the US move
nasnonever had an entry — see the note below

A full docker2 outage cannot fire this rule (the thing evaluating it is also gone). That case is caught by the external Watchdog.

Keep this list in sync with JAKEMON_HOSTS in remote.env.

First checks

ping -c1 <host>
ssh jakefear@<host> 'uptime; docker ps --filter name=alloy'
ssh jakefear@<host> 'curl -s localhost:12345/-/ready'
ssh jakefear@<host> 'docker logs --tail 100 agent-alloy-1'

If the agent container is up but nothing arrives centrally, suspect the remote-write path: the agent must reach JAKEMON_CENTRAL_ADDR (192.168.0.5) as a container-routable address. A bare hostname resolves on the LAN but not inside the Alloy container — CLAUDE.md gotcha #1.

How to clear

Bring the host back, then:

bin/deploy-agent.sh <host>

Notes

⚠️ nas has no entry here. It is push-based with no host-down coverage at all, which is why its silence went unnoticed for two months in 2026. Its only proxy signal is IncidentWikantikOffsiteBackupStale, whose absent() arm was added specifically to cover that hole. Consider adding a nas entry here once the hardware is back and expected to stay up.