IncidentContainerRestartLooping

Severity: warning · Fires after: 0m · Rule: central/prometheus/rules/container.yml

What it means

A container has restarted more than 3 times in 15 minutes — it is crash-looping rather than running. The service may still look "up" intermittently between restarts, which is why this is worth its own signal.

Expression

changes(container_start_time_seconds{name!=""}[15m]) > 3

cAdvisor's containerd handler emits no restart counter, so the loop is detected by the container's start time changing repeatedly. {{ $labels.name }} is the container ID; {{ $labels.image }} is the human-readable part.

First checks

ssh jakefear@<host> 'docker ps -a --filter status=restarting'
ssh jakefear@<host> 'docker logs --tail 200 <container>'
ssh jakefear@<host> 'docker inspect <container> --format "{{.State.ExitCode}} {{.State.Error}}"'

Usual causes: bad config after a deploy, a missing secret or env var, a dependency that is not up yet, or OOM (check IncidentHostOOMKill).

How to clear

Fix the cause and let it start cleanly. The alert clears once restarts stop.

Notes

Containers Docker auto-named <adjective>_<surname> are collapsed into container="ephemeral" in the log pipeline so a docker run without --name cannot mint a permanent stream. That relabeling is in discovery.relabel "containers" in agent/config.alloy and does not affect this metric-side alert.