IncidentVisibilityShipperStale

Severity: warning · Fires after: 15m · Rule: central/prometheus/rules/visibility.yml

What it means

visibility-shipper is running and being scraped, but has not completed a clean cycle in over 26 hours. Usually a revoked credential or a moved endpoint.

Expression

visibility_shipper_configured == 1
  and (time() - visibility_shipper_last_success_timestamp_seconds) > 93600

The silent failure this was built for

A revoked credential makes every POST return 403. ship_all reports the failures rather than raising, the loop keeps running, and the next cycle is 24h away. No crash, no restart, no alert — broken ingest could sit unnoticed for a full day. The two visibility rules exist because the two failure modes are invisible to each other: a dead container publishes no timestamp, while a running container with a bad credential stays up=1 forever.

First checks

ssh jakefear@docker2 'docker compose -f /opt/jakemon/docker-compose.yml logs --tail 200 visibility-shipper | grep -iE "fail|403|401|error"'

curl -sG 'http://192.168.0.5:9090/api/v1/query' \
  --data-urlencode 'query=(time() - visibility_shipper_last_success_timestamp_seconds) / 3600'

HTTP 401/403 means the token; connection errors or 404 mean the URL moved.

How to clear

Fix WIKANTIK_INSIGHTS_URL / WIKANTIK_INSIGHTS_TOKEN in .env, redeploy, and let one clean cycle complete:

bin/deploy-central.sh

Notes

Ingest uses HTTP Basic against Wikantik's /admin endpoint. Related: IncidentVisibilityShipperDown.