Severity: warning · Fires after: 10m · Rule: central/prometheus/rules/ollama.yml
ollama's generation throughput on inference has averaged under 8 tokens/sec
while under load for 10 minutes. Inference still works; it is just slow enough
to be a bad experience.
rate(loki_process_custom_ollama_eval_tokens_per_second_sum[10m])
/ rate(loki_process_custom_ollama_eval_tokens_per_second_count[10m]) < 8
and rate(loki_process_custom_ollama_eval_tokens_per_second_count[10m]) > 0
The second clause is essential: without it, an idle GPU divides zero by zero
and the alert fires for no reason. > 0 means "only judge throughput when
requests are actually happening".
Not from an exporter. The agent tails ollama's journald print_timing lines
via loki.process stage.metrics, and the resulting metrics are surfaced by a
self-scrape of Alloy's :12345. The loki_process_custom_ prefix is Alloy's,
not something this repo chose — if you are searching for these names and finding
nothing, that prefix is why.
ssh jakefear@inference 'journalctl -u ollama --since "20 min ago" --no-pager | grep print_timing | tail -20'
ssh jakefear@inference 'nvidia-smi'
ssh jakefear@inference 'curl -s localhost:11434/api/ps | python3 -m json.tool'
The dominant cause is VRAM pressure forcing CPU offload — check IncidentGpuVramSaturated first. A large model on a 16GB card will simply be slow, which is a capacity fact rather than an incident.
Unload competing models, use a smaller quantisation, or accept the throughput for that model and raise the threshold.
⚠️ inference is currently boxed for the US move, so this rule is inert.
Known limitation: per-model throughput labels are not implemented, so this
is a single aggregate across whatever was running. A slow large model and a
degraded small one look identical here. That is the open follow-up recorded in
the 2026-07-11 addendum of docs/BRINGUP-STATUS.md.