Severity: warning · Fires after: 0m · Rule: central/prometheus/rules/backup.yml
There is no fresh off-box copy of the Wikantik backup. One of two things:
The alert text tells you which. The absent case carries no host label,
because there is no series to carry one.
(time() - wikantik_backup_offsite_last_success_timestamp_seconds > 36 * 3600)
or
absent(wikantik_backup_offsite_last_success_timestamp_seconds)
absent() arm existstime() - <gauge> can only detect staleness while the gauge still exists.
When the off-box host stops writing the series, the series vanishes, the
subtraction evaluates over an empty vector, and the rule goes silent — at
exactly the moment there is no off-box copy at all.
That is not hypothetical. Between the NAS shutdown on 2026-07-21 and 2026-09-20
the fleet had no off-box copy for roughly two months, and this alert was
structurally unable to say so. absent() closes that.
Only the offsite tier gets absence coverage. The docker1-sourced gauges
vanishing means docker1 itself is gone, which ServiceDown and
HostNotReporting already page for. The NAS is push-based with no
HostNotReporting entry, so nothing else notices its silence.
ping -c1 nas.lan
ssh jakefear@nas.lan 'docker ps --filter name=alloy'
curl -sG 'http://192.168.0.5:9090/api/v1/query' \
--data-urlencode 'query=wikantik_backup_offsite_last_success_timestamp_seconds'
Empty result = the absent arm. Anything returned = the stale arm.
If the NAS is simply not deployed yet:
bin/deploy-agent.sh nas
Then re-verify the off-box pull actually resumes — the agent reporting is not the same as the backup pull succeeding. Watch for the gauge to appear and then advance.
If the NAS is knowingly dark (hardware in transit), silence with an expiry rather than disabling the rule.
nas is in JAKEMON_TAR_HOSTS — its rsync rejects --server writes, so
deploys go over tar-over-ssh. See CLAUDE.md gotcha #7.
Related: IncidentWikantikBackupStale.