IncidentGpuExporterDown

Severity: warning · Fires after: 5m · Rule: central/prometheus/rules/gpu.yml

What it means

The DCGM GPU exporter on inference is not responding. GPU visibility is lost — ollama keeps serving perfectly well, you just cannot see VRAM, temperature or utilisation.

Expression

up{job="gpu"} == 0

Why it is not critical

Losing the exporter loses observation, not service. This is also why gpu is excluded from ServiceDown — it would otherwise page at critical for a monitoring-only failure. The trade-off is that IncidentGpuVramSaturated and IncidentGpuThermalThrottle are blind while this is firing.

First checks

ssh jakefear@inference 'docker ps -a --filter name=dcgm'
ssh jakefear@inference 'docker logs --tail 100 <dcgm-container>'
ssh jakefear@inference 'nvidia-smi'          # does the driver itself still work?

If nvidia-smi fails, the problem is the driver or the card, not the exporter — and that is a much bigger deal than this warning suggests.

How to clear

bin/deploy-gpu-exporter.sh inference

Notes

⚠️ inference is currently boxed for the US move, so this rule is inert — there is no up{job="gpu"} series at all, and == 0 cannot match an absent series. (Contrast IncidentWikantikOffsiteBackupStale, where an absent() arm was added precisely because absence was the failure that mattered. Here absence is expected, so no absent() arm is wanted.)

The exporter pin is knowingly stale at 3.3.9-3.6.1-ubuntu22.04; upstream is 4.x. Revisit at the US rebuild.