IncidentHostMemoryHigh

Severity: warning · Fires after: 10m · Rule: central/prometheus/rules/host.yml

What it means

Available memory has been under 10% for 10 minutes. Not yet fatal, but the host is one workload spike away from the OOM killer.

Expression

100 * (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) > 90

Uses MemAvailable, not MemFree, so page cache is correctly treated as reclaimable — this will not fire merely because Linux is caching aggressively.

First checks

ssh jakefear@<host> 'free -h'
ssh jakefear@<host> 'ps aux --sort=-%mem | head -15'
ssh jakefear@<host> 'docker stats --no-stream'

How to clear

Restart or constrain the offending workload. If a container has no memory limit and is growing unbounded, add a mem_limit in its compose file.

Notes

docker2 carries the whole central stack (Prometheus, Loki, Grafana, Alertmanager, three exporters) plus both CrafterCMS stacks and wikantik. It has 30GB, which is why it was chosen as the relocation target — but it is the host most likely to trip this.

If this fires and then the host goes quiet, check IncidentHostOOMKill — the kill may already have happened.