IncidentHostOOMKill

Severity: warning · Fires after: 0m · Rule: central/prometheus/rules/host.yml

What it means

The kernel OOM killer terminated a process in the last 10 minutes. The host ran out of memory and something was sacrificed.

Fires immediately (for: 0m) because the kill already happened — waiting would only delay the news.

Expression

increase(node_vmstat_oom_kill[10m]) > 0

Why it matters even though the host recovered

The OOM killer picks by heuristic, not importance. It may have killed a database, a backup job mid-write, or the monitoring agent itself. The host looking healthy afterwards tells you nothing about what was lost.

First checks

ssh jakefear@<host> 'sudo dmesg -T | grep -i "killed process" | tail -20'
ssh jakefear@<host> 'sudo journalctl -k --since "30 min ago" | grep -i oom'
ssh jakefear@<host> 'docker ps -a --filter status=exited'

Then confirm what died actually came back:

curl -sG 'http://192.168.0.5:9090/api/v1/query' --data-urlencode 'query=up == 0'

How to clear

The alert clears itself once 10 minutes pass without another kill. The real work is preventing a repeat: add or lower a mem_limit on the offending container, or reduce what the host is being asked to run.

Notes

Related: IncidentHostMemoryHigh is the warning before this; if you see that one and act, you never see this one. IncidentContainerRestartLooping often follows an OOM kill when the container restarts straight back into the same memory ceiling.