Severity: warning · Fires after: 0m · Rule: central/prometheus/rules/backup.yml
A longer-cadence Wikantik archive tier has missed its window: weekly past 9 days, or monthly past 35 days. Each threshold is one full cadence plus grace.
Warning rather than critical on purpose. A missed archive costs retention depth — how far back you can restore to — not the ability to restore at all. The daily tier covers that, at critical.
(time() - wikantik_backup_last_success_timestamp_seconds{tier="weekly"} > 9 * 24 * 3600)
or
(time() - wikantik_backup_last_success_timestamp_seconds{tier="monthly"} > 35 * 24 * 3600)
{{ $labels.tier }} in the alert tells you which cadence missed.
# Which tiers exist and how old is each?
curl -sG 'http://192.168.0.5:9090/api/v1/query' \
--data-urlencode 'query=(time() - wikantik_backup_last_success_timestamp_seconds) / 86400'
ssh jakefear@docker1 'docker logs --tail 200 repo-backup-1 | grep -iE "weekly|monthly"'
ssh jakefear@docker1 'ls -la /var/backups/wikantik/ 2>/dev/null || true'
The most common cause is not a broken job but the host being off when the schedule fired — a monthly run that lands on the 1st simply never happens if the machine is powered down that day, and nothing retries it.
Trigger the missed tier manually, or wait for the next scheduled run if it is close. The gauge advances on success and the alert resolves.
Added 2026-09-20. The weekly and monthly tiers had existed for some time with no
rule watching them — the original WikantikBackupStale comment said other tiers
"would need their own threshold" and none was ever added. It was found with the
monthly tier 50 days stale: the September run never fired because the host
was in transit during the US move.
If a tier is expected to be missing for a known reason (hardware in transit), prefer a time-bounded Alertmanager silence over deleting the rule, so it re-surfaces on expiry:
ssh jakefear@docker2 'curl -s localhost:9093/api/v2/silences'
Related: IncidentWikantikBackupStale.