IncidentWikantikBackupStale

Severity: critical · Fires after: 0m · Rule: central/prometheus/rules/backup.yml

What it means

The Wikantik daily backup on docker1 — the source copy everything else is derived from — has not succeeded in over 36 hours. The threshold is one missed run plus grace, so this means at least one scheduled run did not complete.

This is the critical tier because it is the difference between "a restore would lose a day" and "there may be nothing fresh to restore from".

Expression

time() - wikantik_backup_last_success_timestamp_seconds{tier="daily"} > 36 * 3600

The gauge is written by the backup sidecar into the host textfile collector (see agent/config.alloy), so a stale value means the backup job itself did not report success — not that the metric pipeline broke.

First checks

ssh jakefear@docker1 'docker ps -a --filter name=backup'
ssh jakefear@docker1 'docker logs --tail 100 repo-backup-1'
ssh jakefear@docker1 'ls -la /var/lib/jakemon/textfile/'   # is the gauge being written?
ssh jakefear@docker1 'df -h'                                # out of disk is the usual cause

Confirm the gauge actually is stale rather than absent:

curl -sG 'http://192.168.0.5:9090/api/v1/query' \
  --data-urlencode 'query=wikantik_backup_last_success_timestamp_seconds'

How to clear

Fix the underlying backup job and let one run complete; the gauge advances and the alert resolves on its own. Do not clear it by touching the textfile — it would report a success that never happened, which is exactly the failure this rule exists to catch.

Notes

Found 2026-06-08 already ~16 days stale with no alert at all. The off-box NAS pull kept "succeeding" by copying an already-stale backup, so the offsite signal looked healthy while the source was rotting. That masking is why the source tier is alerted separately and at a higher severity than the offsite pull.

Related: IncidentWikantikOffsiteBackupStale, IncidentWikantikArchiveBackupStale.