Severity: warning · Fires after: 5m · Rule: central/prometheus/rules/gpu.yml
The GPU on inference has been above 85°C for 5 minutes. At this range the
card begins clocking itself down to protect itself, so sustained load gets
slower and slower.
DCGM_FI_DEV_GPU_TEMP > 85
ssh jakefear@inference 'nvidia-smi --query-gpu=temperature.gpu,clocks.sm,power.draw,utilization.gpu --format=csv'
ssh jakefear@inference 'nvidia-smi -q -d PERFORMANCE | grep -A10 "Clocks Throttle Reasons"'
Clocks Throttle Reasons confirms whether throttling is actually happening or
the card is merely warm.
Reduce sustained load, or fix cooling. Physical causes dominate: dust in the heatsink, a failed case fan, poor airflow, or high ambient temperature.
Ambient matters more than it sounds — check the host's other thermal sensors on
the host-health dashboard to see whether the whole box is hot or just the GPU.
⚠️ inference is currently boxed for the US move, so this rule is inert.
Worth re-checking on arrival: a machine that has been packed, shipped and unpacked is exactly when a heatsink comes loose or a fan cable gets knocked off. This alert firing shortly after the rebuild would be a real signal, not noise.