IncidentGpuThermalThrottle

Severity: warning · Fires after: 5m · Rule: central/prometheus/rules/gpu.yml

What it means

The GPU on inference has been above 85°C for 5 minutes. At this range the card begins clocking itself down to protect itself, so sustained load gets slower and slower.

Expression

DCGM_FI_DEV_GPU_TEMP > 85

First checks

ssh jakefear@inference 'nvidia-smi --query-gpu=temperature.gpu,clocks.sm,power.draw,utilization.gpu --format=csv'
ssh jakefear@inference 'nvidia-smi -q -d PERFORMANCE | grep -A10 "Clocks Throttle Reasons"'

Clocks Throttle Reasons confirms whether throttling is actually happening or the card is merely warm.

How to clear

Reduce sustained load, or fix cooling. Physical causes dominate: dust in the heatsink, a failed case fan, poor airflow, or high ambient temperature.

Ambient matters more than it sounds — check the host's other thermal sensors on the host-health dashboard to see whether the whole box is hot or just the GPU.

Notes

⚠️ inference is currently boxed for the US move, so this rule is inert.

Worth re-checking on arrival: a machine that has been packed, shipped and unpacked is exactly when a heatsink comes loose or a fan cable gets knocked off. This alert firing shortly after the rebuild would be a real signal, not noise.