Severity: warning · Fires after: 5m · Rule: central/prometheus/rules/gpu.yml
The DCGM GPU exporter on inference is not responding. GPU visibility is
lost — ollama keeps serving perfectly well, you just cannot see VRAM,
temperature or utilisation.
up{job="gpu"} == 0
Losing the exporter loses observation, not service. This is also why gpu is
excluded from ServiceDown — it would otherwise page at critical for a
monitoring-only failure. The trade-off is that
IncidentGpuVramSaturated and
IncidentGpuThermalThrottle are blind while this
is firing.
ssh jakefear@inference 'docker ps -a --filter name=dcgm'
ssh jakefear@inference 'docker logs --tail 100 <dcgm-container>'
ssh jakefear@inference 'nvidia-smi' # does the driver itself still work?
If nvidia-smi fails, the problem is the driver or the card, not the exporter —
and that is a much bigger deal than this warning suggests.
bin/deploy-gpu-exporter.sh inference
⚠️ inference is currently boxed for the US move, so this rule is inert —
there is no up{job="gpu"} series at all, and == 0 cannot match an absent
series. (Contrast
IncidentWikantikOffsiteBackupStale,
where an absent() arm was added precisely because absence was the failure that
mattered. Here absence is expected, so no absent() arm is wanted.)
The exporter pin is knowingly stale at 3.3.9-3.6.1-ubuntu22.04; upstream is
4.x. Revisit at the US rebuild.