IncidentGpuVramSaturated

Severity: warning · Fires after: 15m · Rule: central/prometheus/rules/gpu.yml

What it means

GPU VRAM on inference (RTX 4060 Ti) has been over 95% used for 15 minutes. ollama may be spilling layers to CPU, which collapses generation throughput without anything actually failing.

Expression

DCGM_FI_DEV_FB_USED / (DCGM_FI_DEV_FB_USED + DCGM_FI_DEV_FB_FREE) > 0.95

Why it is a warning

Nothing is broken — inference still works, just slowly. The user-visible symptom is latency, not errors, which is exactly why it needs a metric to surface it.

First checks

ssh jakefear@inference 'nvidia-smi'
ssh jakefear@inference 'curl -s localhost:11434/api/ps | python3 -m json.tool'

/api/ps shows which models are currently loaded and their sizes — that is usually the whole story. Several models resident at once, or one model larger than the card, will pin VRAM.

Cross-check throughput to confirm real impact: IncidentOllamaThroughputLow.

How to clear

Unload models or reduce concurrency:

ssh jakefear@inference 'curl -s http://localhost:11434/api/generate -d "{\"model\":\"<model>\",\"keep_alive\":0}"'

Lowering OLLAMA_KEEP_ALIVE or OLLAMA_MAX_LOADED_MODELS is the durable fix if this recurs.

Notes

⚠️ inference is currently boxed for the US move, so this alert cannot fire and the GPU rules are inert. The nvidia/dcgm-exporter pin (3.3.9-3.6.1-ubuntu22.04) is knowingly stale — upstream is 4.x, which also has NVIDIA driver requirements that cannot be verified without the box. Revisit both at the US rebuild.

Dashboard: inference-gpu ("Inference — RTX 4060 Ti + ollama").