Severity: warning · Fires after: 0m · Rule: central/prometheus/rules/host.yml
The kernel OOM killer terminated a process in the last 10 minutes. The host ran out of memory and something was sacrificed.
Fires immediately (for: 0m) because the kill already happened — waiting
would only delay the news.
increase(node_vmstat_oom_kill[10m]) > 0
The OOM killer picks by heuristic, not importance. It may have killed a database, a backup job mid-write, or the monitoring agent itself. The host looking healthy afterwards tells you nothing about what was lost.
ssh jakefear@<host> 'sudo dmesg -T | grep -i "killed process" | tail -20'
ssh jakefear@<host> 'sudo journalctl -k --since "30 min ago" | grep -i oom'
ssh jakefear@<host> 'docker ps -a --filter status=exited'
Then confirm what died actually came back:
curl -sG 'http://192.168.0.5:9090/api/v1/query' --data-urlencode 'query=up == 0'
The alert clears itself once 10 minutes pass without another kill. The real work
is preventing a repeat: add or lower a mem_limit on the offending container,
or reduce what the host is being asked to run.
Related: IncidentHostMemoryHigh is the warning before this; if you see that one and act, you never see this one. IncidentContainerRestartLooping often follows an OOM kill when the container restarts straight back into the same memory ceiling.