Alert info
- Severity
- Critical
- Detected by
- Grafana
- Alert
- Image Pull / CrashLoopBackOff on clickhouse/prod-clickhouse-keeper-0 (FOR=5m)
- Time
- 2026-10-05 02:49:30 UTC
- Service
- prod-clickhouse-keeper (clickhouse)
Kubernetes pod clickhouse/prod-clickhouse-keeper-0 (container keeper) entered CrashLoopBackOff, and the condition persisted for the configured FOR=5m window.
One ClickHouse Keeper pod crash-looped because its entrypoint could not be executed (exec format error, exitCode=255) on one node, while its two peers stayed Running. It recovered after kubelet pulled the image fresh on the same node, which is most consistent with a node-local cached image inconsistency. Confidence is medium.
- Exit Code
- 255 (exec format error)
- Affected Pods
- 1 of 3 keepers
- Confidence
- Medium
Entities identified
Infrastructure
Grafana
alerting
Image Pull / CrashLoopBackOff alert, FOR=5m.
ip-172-31-96-57
node
The node keeper-0 failed on, and started successfully on after the fresh pull. No reschedule.
Service
prod-clickhouse-keeper-0
container keeper
Entered CrashLoopBackOff. The only pod affected.
keeper-1 / keeper-2
clickhouse namespace
Stayed Running with restartCount=0.
Data
cached image
node-local
The node-local cached image the failed start ran from.
entrypoint.sh
/opt/bitnami/scripts/clickhouse-keeper
The resolved entrypoint that could not be executed.
External
keeper image tag
mutable
Rewritten by PutImage at 2026-10-05 02:26:42 UTC.
How they relate
The same relationships as text
- Grafana alerts on prod-clickhouse-keeper-0
- prod-clickhouse-keeper-0 runs on ip-172-31-96-57
- ip-172-31-96-57 starts from cached image
- cached image resolves entrypoint.sh. Resolved to a non-executable format (exec format error, exitCode=255).
- keeper image tag pulled ip-172-31-96-57. Kubelet pulled digest sha256:847fff... at 02:48:11 UTC and the container started.
Hypotheses tree
Ruled-out branches are collapsed. Open any one to read the check that eliminated it.
Cluster-wide misconfigurationRuled out
Compare keeper-0 with keeper-1 and keeper-2.
Blast radius was limited to keeper-0. keeper-1 and keeper-2 stayed Running with restartCount=0.
keeper-1 and keeper-2 Running, restartCount=0.
Only one pod diverged.
Application, probe or OOM failureRuled out
Read the previous container log.
The container start failed because the entrypoint could not be executed (exec format error, exitCode=255).
exec /opt/bitnami/scripts/clickhouse-keeper/entrypoint.sh: exec format error, exitCode=255.
The failure is at container start.
Container start failed at exec on node ip-172-31-96-57Confirmed
Follow the timeline from the failed start to recovery.
At 02:47:38 UTC the start on ip-172-31-96-57 failed with exec format error. At 02:48:11 UTC kubelet pulled image digest sha256:847fff... and the container started successfully on the same node, with no reschedule.
02:47:38 UTC, exec format error, exitCode=255, node ip-172-31-96-57.
The start failed on the cached image.
02:48:11 UTC, kubelet pulls digest sha256:847fff... and the container starts. No reschedule.
A fresh pull on the same node ended the incident.
Node-local cached image or runtime inconsistency (bad or partial cached layer)Root cause
Explain the fail on cached, succeed after repull pattern.
The pod recovered only after kubelet performed a fresh image pull and then started successfully on the same node. This is most consistent with a node-local cached image or runtime inconsistency (a bad or partial cached layer) rather than an application, probe or OOM failure or a cluster-wide misconfiguration.
Exec format error at 02:47:38 UTC, successful start after the 02:48:11 UTC pull, same node.
The fault was in the node's local image state.
Mutable image tag rewritten shortly before the incidentConfirmed
Check for changes to the keeper image tag before the incident.
The keeper image tag was mutable and was rewritten by PutImage at 2026-10-05 02:26:42 UTC, increasing the likelihood of per-node divergence. This explains why only one pod diverged, but does not by itself explain the fail on cached, succeed after repull pattern, so it is a contributing factor.
PutImage at 2026-10-05 02:26:42 UTC.
Nodes could hold different digests under the same tag.
Exact digest used by the failed cached startNeeds access
Identify the digest and content used by the failed start.
Kubernetes events did not expose it, so the root cause is stated as a node-local cached image or runtime inconsistency rather than a specific digest mismatch.
Final RCA
Node-local cached image or runtime inconsistency (bad or partial cached layer)
The pod recovered only after kubelet performed a fresh image pull and then started successfully on the same node. This is most consistent with a node-local cached image or runtime inconsistency (a bad or partial cached layer) rather than an application, probe or OOM failure or a cluster-wide misconfiguration.
Remediation
The incident ended without intervention: at 02:48:11 UTC kubelet pulled image digest sha256:847fff... and the container started on the same node. The RCA names the mutable keeper image tag, rewritten at 02:26:42 UTC, as a contributing factor.
Frequently asked questions
What caused the ClickHouse Keeper CrashLoopBackOff?
On node ip-172-31-96-57 the keeper container could not execute its entrypoint, /opt/bitnami/scripts/clickhouse-keeper/entrypoint.sh, and failed with exec format error and exitCode=255. The pod recovered only after kubelet pulled the image fresh and started it on the same node, which is most consistent with a bad or partial cached image layer on that node.
Why did only one ClickHouse Keeper pod fail?
keeper-1 and keeper-2 stayed Running with restartCount=0, so the blast radius was limited to keeper-0. The keeper image tag was mutable and had been rewritten by PutImage at 02:26:42 UTC, which increases the likelihood of nodes holding different digests under one tag and explains why a single pod diverged.
How did the incident end?
At 02:48:11 UTC kubelet pulled image digest sha256:847fff... and the container started successfully on the same node, with no reschedule. The Grafana Image Pull / CrashLoopBackOff alert fired at 02:49:30 UTC, after the condition had held for its configured FOR=5m window.
Why is the confidence for this root cause medium?
Kubernetes events did not expose the digest or content used by the failed cached start, so a specific digest mismatch cannot be proven. The root cause is therefore stated as a node-local cached image or runtime inconsistency, which fits the fail on cached, succeed after repull pattern without naming the exact content that failed.