Severity: warning · Fires after: 5m · Rule: central/prometheus/rules/host.yml
The kernel connection-tracking table is over 80% of its limit. If it fills completely, the kernel silently drops new connections — no error, no log on the application side, just timeouts.
100 * node_nf_conntrack_entries / node_nf_conntrack_entries_limit > 80
ssh jakefear@<host> 'cat /proc/sys/net/netfilter/nf_conntrack_count /proc/sys/net/netfilter/nf_conntrack_max'
ssh jakefear@<host> 'sudo conntrack -L 2>/dev/null | awk "{print \$3}" | sort | uniq -c | sort -rn | head'
ssh jakefear@<host> 'ss -s'
Docker hosts track every container connection, so a host running many
containers with chatty outbound traffic fills this faster than its raw traffic
volume suggests. On this fleet docker2 (central stack + both Crafter stacks)
and cloudflare (edge proxy) are the likely candidates.
Short term, raise the limit:
ssh jakefear@<host> 'sudo sysctl -w net.netfilter.nf_conntrack_max=262144'
Persist it in /etc/sysctl.d/ if it should stick across reboots. Longer term,
find what is holding connections open — a leaking client or an absent keepalive
tends to be the real cause.
Added 2026-07-02 as a silent-failure alert: the failure mode is invisible from inside the application, which is exactly why it needs a host-level signal.