Severity: warning · Fires after: 15m · Rule: central/prometheus/rules/gpu.yml
GPU VRAM on inference (RTX 4060 Ti) has been over 95% used for 15 minutes.
ollama may be spilling layers to CPU, which collapses generation throughput
without anything actually failing.
DCGM_FI_DEV_FB_USED / (DCGM_FI_DEV_FB_USED + DCGM_FI_DEV_FB_FREE) > 0.95
Nothing is broken — inference still works, just slowly. The user-visible symptom is latency, not errors, which is exactly why it needs a metric to surface it.
ssh jakefear@inference 'nvidia-smi'
ssh jakefear@inference 'curl -s localhost:11434/api/ps | python3 -m json.tool'
/api/ps shows which models are currently loaded and their sizes — that is
usually the whole story. Several models resident at once, or one model larger
than the card, will pin VRAM.
Cross-check throughput to confirm real impact: IncidentOllamaThroughputLow.
Unload models or reduce concurrency:
ssh jakefear@inference 'curl -s http://localhost:11434/api/generate -d "{\"model\":\"<model>\",\"keep_alive\":0}"'
Lowering OLLAMA_KEEP_ALIVE or OLLAMA_MAX_LOADED_MODELS is the durable fix if
this recurs.
⚠️ inference is currently boxed for the US move, so this alert cannot fire
and the GPU rules are inert. The nvidia/dcgm-exporter pin
(3.3.9-3.6.1-ubuntu22.04) is knowingly stale — upstream is 4.x, which also has
NVIDIA driver requirements that cannot be verified without the box. Revisit both
at the US rebuild.
Dashboard: inference-gpu ("Inference — RTX 4060 Ti + ollama").