IncidentWebmasterProviderStale

Severity: warning · Fires after: 15m · Rule: central/prometheus/rules/visibility.yml

What it means

An engine still reads up == 1, but its last successful poll is over 14h old — two polls have been missed. The poll loop has stopped or hung; up is simply the value left behind by the last good poll.

Expression

webmaster_exporter_up == 1
  and (time() - webmaster_exporter_last_success_timestamp_seconds) > 50400

Gated on up == 1 so a failing engine raises only IncidentWebmasterProviderDown.

First checks

ssh jakefear@docker2 'cd /opt/jakemon && docker compose logs --since 24h visibility-exporter | tail -50'

No recent log lines at all = the poll thread died (/healthz and /metrics keep answering from the main thread, so the container still looks healthy) or is blocked on a hung HTTP call.

How to clear

docker compose restart visibility-exporter. If it recurs, the traceback just before the silence names the unguarded raise.