IncidentHostDiskFull

Severity: warning · Fires after: 5m · Rule: central/prometheus/rules/host.yml

What it means

A real, persistent filesystem is over 90% full.

Expression

100 * node_filesystem_avail_bytes{fstype=~"ext4|xfs|btrfs"}
    / node_filesystem_size_bytes{fstype=~"ext4|xfs|btrfs"} < 10

Scoped to ext4|xfs|btrfs on purpose — tmpfs, fuse and vfat mounts are excluded so they don't generate noise. {{ $labels.mountpoint }} names the filesystem.

First checks

ssh jakefear@<host> 'df -h'
ssh jakefear@<host> 'sudo du -xh --max-depth=1 /var | sort -h | tail -20'
ssh jakefear@<host> 'docker system df'

On this fleet the usual culprits, in order:

  1. Docker — dangling images and build cache after a deploy.
  2. Backups — docker1 holds the Wikantik dumps; retention may have grown.
  3. Logs — a container logging heavily to a json-file driver.
  4. Prometheus/Loki data on docker2 (90d retention, ~9GB+).

How to clear

ssh jakefear@<host> 'docker system prune -f'          # safe: dangling only
ssh jakefear@<host> 'docker image prune -a -f'        # check first, removes unused images
ssh jakefear@<host> 'sudo journalctl --vacuum-time=7d'

Be careful pruning on docker2 — images there are pinned and re-pulling the full stack takes a while.

Notes

The host-health dashboard has a disk-fill projection panel — use it to see whether this is a slow creep or something that will hit 100% in hours.

Related: IncidentHostInodesFull — a filesystem can refuse writes at 60% bytes used if inodes are exhausted.