IncidentCrafterOpenSearchYellow

Severity: warning · Fires after: 1h · Rule: central/prometheus/rules/crafter.yml

What it means

A CrafterCMS OpenSearch cluster has been YELLOW for over an hour: replica shards are unassigned. All data is available and searchable — there is simply no redundancy.

The 1h delay is deliberate; yellow is normal and transient during restarts and reindexing.

Expression

crafter_opensearch_cluster_status == 1

The likely cause on this fleet

Both clusters are single-node. A single node cannot assign a replica to itself, so an index created with number_of_replicas: 1 will sit yellow forever. That is a configuration mismatch, not a fault.

The alert summary states the fix directly: set number_of_replicas=0 for a single node, or add a node.

First checks

ssh jakefear@docker2 'docker exec <opensearch-container> curl -s localhost:9200/_cluster/health?pretty'
ssh jakefear@docker2 'docker exec <opensearch-container> curl -s "localhost:9200/_cat/indices?v"'

Look at the rep column — any index with rep > 0 on a one-node cluster is the culprit.

How to clear

ssh jakefear@docker2 'docker exec <opensearch-container> curl -s -XPUT \
  -H "Content-Type: application/json" \
  "localhost:9200/_all/_settings" \
  -d "{\"index\":{\"number_of_replicas\":0}}"'

The cluster should go green within seconds.

Notes

Do not "fix" persistent yellow by lengthening the for: duration — that hides a real configuration mismatch rather than resolving it. Related: IncidentCrafterOpenSearchRed.