IncidentCrafterOpenSearchDiskHigh

Severity: warning · Fires after: 10m · Rule: central/prometheus/rules/crafter.yml

What it means

An OpenSearch node's disk is over 85% used. OpenSearch's own low watermark is 90%, at which point it stops allocating shards to the node; at 95% it flips indices to read-only.

The 85% threshold is an early warning with room to act before OpenSearch starts refusing work.

Expression

crafter_opensearch_disk_used_percent > 85

Why it escalates on its own

Crossing the watermarks turns a disk-space warning into a cluster-health incident:

UsedOpenSearch behaviour
85%this alert
90%stops allocating shards → cluster goes yellow/red
95%indices flipped to read-only → writes fail

So this alert arriving and being ignored is how you get IncidentCrafterOpenSearchRed.

First checks

ssh jakefear@docker2 'df -h'
ssh jakefear@docker2 'docker system df'
ssh jakefear@docker2 'docker exec <opensearch-container> curl -s "localhost:9200/_cat/allocation?v"'
ssh jakefear@docker2 'docker exec <opensearch-container> curl -s "localhost:9200/_cat/indices?v&s=store.size:desc"'

How to clear

Free space on docker2. Remember this host carries the entire central stack plus both CrafterCMS stacks, so OpenSearch is competing with Prometheus and Loki data — the biggest win is often not in OpenSearch at all.

ssh jakefear@docker2 'docker system prune -f'

If an index is genuinely oversized, delete old indices or reduce retention.

Notes

If the read-only flag has already been set, clearing disk is not enough — the block must be released explicitly:

ssh jakefear@docker2 'docker exec <opensearch-container> curl -s -XPUT \
  -H "Content-Type: application/json" "localhost:9200/_all/_settings" \
  -d "{\"index.blocks.read_only_allow_delete\":null}"'

Related: IncidentHostDiskFull.