Severity: warning · Fires after: 10m · Rule: central/prometheus/rules/host.yml
Available memory has been under 10% for 10 minutes. Not yet fatal, but the host is one workload spike away from the OOM killer.
100 * (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) > 90
Uses MemAvailable, not MemFree, so page cache is correctly treated as
reclaimable — this will not fire merely because Linux is caching aggressively.
ssh jakefear@<host> 'free -h'
ssh jakefear@<host> 'ps aux --sort=-%mem | head -15'
ssh jakefear@<host> 'docker stats --no-stream'
Restart or constrain the offending workload. If a container has no memory limit
and is growing unbounded, add a mem_limit in its compose file.
docker2 carries the whole central stack (Prometheus, Loki, Grafana,
Alertmanager, three exporters) plus both CrafterCMS stacks and wikantik. It
has 30GB, which is why it was chosen as the relocation target — but it is the
host most likely to trip this.
If this fires and then the host goes quiet, check IncidentHostOOMKill — the kill may already have happened.