IncidentHostFilesystemReadOnly

Severity: critical · Fires after: 1m · Rule: central/prometheus/rules/host.yml

What it means

A real filesystem has been remounted read-only by the kernel. Linux does this to protect data when it hits I/O errors — it usually signals failing storage, not a configuration mistake.

Treat as hardware failure until proven otherwise.

Expression

node_filesystem_readonly{fstype=~"ext4|xfs|btrfs"} == 1

Why it is critical

Everything on that mount silently stops being able to write. Services often keep running and appear healthy while losing every write — databases, backups and logs included. This is one of the quietest catastrophic failure modes there is, which is why it fires after only 1 minute.

First checks

ssh jakefear@<host> 'mount | grep " ro,"'
ssh jakefear@<host> 'sudo dmesg -T | grep -iE "I/O error|remount|ext4-fs|xfs|blk_update"'
ssh jakefear@<host> 'sudo smartctl -a /dev/<disk>'

The host-health dashboard has a Disk SMART row, but note SMART needs a manual per-host install (agent/textfile/smartmon.sh via root cron writing to /var/lib/jakemon/textfile/ — see agent/textfile/README.md). If the row is empty on this host, that step was never done.

How to clear

Do not simply remount read-write and carry on. Establish whether the disk is failing first:

  1. Capture dmesg and SMART output.
  2. Verify backups exist and are fresh — IncidentWikantikBackupStale.
  3. fsck the filesystem from a safe state.
  4. Replace the disk if SMART shows reallocated/pending sectors.

Notes

Added 2026-07-02 as one of five "silent host failure mode" alerts, on the grounds that it was being collected but had no panel or rule watching it.