Severity: critical · Fires after: 0m · Rule: central/prometheus/rules/backup.yml
The Wikantik daily backup on docker1 — the source copy everything else is
derived from — has not succeeded in over 36 hours. The threshold is one missed
run plus grace, so this means at least one scheduled run did not complete.
This is the critical tier because it is the difference between "a restore would lose a day" and "there may be nothing fresh to restore from".
time() - wikantik_backup_last_success_timestamp_seconds{tier="daily"} > 36 * 3600
The gauge is written by the backup sidecar into the host textfile collector
(see agent/config.alloy), so a stale value means the backup job itself did not
report success — not that the metric pipeline broke.
ssh jakefear@docker1 'docker ps -a --filter name=backup'
ssh jakefear@docker1 'docker logs --tail 100 repo-backup-1'
ssh jakefear@docker1 'ls -la /var/lib/jakemon/textfile/' # is the gauge being written?
ssh jakefear@docker1 'df -h' # out of disk is the usual cause
Confirm the gauge actually is stale rather than absent:
curl -sG 'http://192.168.0.5:9090/api/v1/query' \
--data-urlencode 'query=wikantik_backup_last_success_timestamp_seconds'
Fix the underlying backup job and let one run complete; the gauge advances and the alert resolves on its own. Do not clear it by touching the textfile — it would report a success that never happened, which is exactly the failure this rule exists to catch.
Found 2026-06-08 already ~16 days stale with no alert at all. The off-box NAS pull kept "succeeding" by copying an already-stale backup, so the offsite signal looked healthy while the source was rotting. That masking is why the source tier is alerted separately and at a higher severity than the offsite pull.
Related: IncidentWikantikOffsiteBackupStale, IncidentWikantikArchiveBackupStale.