IncidentCrafterOpenSearchRed

Severity: critical · Fires after: 5m · Rule: central/prometheus/rules/crafter.yml

What it means

A CrafterCMS OpenSearch cluster is RED: at least one primary shard is unassigned. Some data is not just un-redundant, it is unavailable. Site search and any content query backed by that index will fail or return partial results.

Expression

crafter_opensearch_cluster_status == 2

Encoding: 0 = green, 1 = yellow, 2 = red. {{ $labels.cluster }} says which of the two (authoring / delivery).

First checks

ssh jakefear@docker2 'docker ps --filter name=opensearch'
ssh jakefear@docker2 'docker exec <opensearch-container> curl -s localhost:9200/_cluster/health?pretty'
ssh jakefear@docker2 'docker exec <opensearch-container> curl -s "localhost:9200/_cat/indices?v&health=red"'
ssh jakefear@docker2 'docker exec <opensearch-container> curl -s "localhost:9200/_cluster/allocation/explain?pretty"'

allocation/explain is the one that actually tells you why a shard will not assign. Check disk first — OpenSearch refuses allocation past its watermark, so IncidentCrafterOpenSearchDiskHigh and this alert often arrive together, and freeing disk resolves both.

How to clear

  1. If disk-driven, free space and shards reassign on their own.
  2. If a node is missing, bring it back.
  3. If the index is genuinely corrupt, rebuild it — CrafterCMS can reindex from the content store, so this is usually recoverable rather than data loss.

Notes

Related: IncidentCrafterOpenSearchYellow is the much less urgent replica-only case.