ClickHouse Keeper CrashLoopBackOff from exec format error on one node

Alert info

Severity
Critical
Detected by
Grafana
Alert
Image Pull / CrashLoopBackOff on clickhouse/prod-clickhouse-keeper-0 (FOR=5m)
Time
2026-10-05 02:49:30 UTC
Service
prod-clickhouse-keeper (clickhouse)

Kubernetes pod clickhouse/prod-clickhouse-keeper-0 (container keeper) entered CrashLoopBackOff, and the condition persisted for the configured FOR=5m window.

One ClickHouse Keeper pod crash-looped because its entrypoint could not be executed (exec format error, exitCode=255) on one node, while its two peers stayed Running. It recovered after kubelet pulled the image fresh on the same node, which is most consistent with a node-local cached image inconsistency. Confidence is medium.

Exit Code
255 (exec format error)
Affected Pods
1 of 3 keepers
Confidence
Medium

Entities identified

Infrastructure

  • Grafana

    alerting

    Image Pull / CrashLoopBackOff alert, FOR=5m.

  • ip-172-31-96-57

    node

    The node keeper-0 failed on, and started successfully on after the fresh pull. No reschedule.

Service

  • prod-clickhouse-keeper-0

    container keeper

    Entered CrashLoopBackOff. The only pod affected.

  • keeper-1 / keeper-2

    clickhouse namespace

    Stayed Running with restartCount=0.

Data

  • cached image

    node-local

    The node-local cached image the failed start ran from.

  • entrypoint.sh

    /opt/bitnami/scripts/clickhouse-keeper

    The resolved entrypoint that could not be executed.

External

  • keeper image tag

    mutable

    Rewritten by PutImage at 2026-10-05 02:26:42 UTC.

How they relate

alerts onruns onstarts fromresolvesGrafanaalertingkeeper-1 / keeper-2clickhouse namespacekeeper image tagmutableprod-clickhouse-keeper…container keeperip-172-31-96-57nodecached imagenode-localentrypoint.sh/opt/bitnami/scripts/clic…
The same relationships as text
  • Grafana alerts on prod-clickhouse-keeper-0
  • prod-clickhouse-keeper-0 runs on ip-172-31-96-57
  • ip-172-31-96-57 starts from cached image
  • cached image resolves entrypoint.sh. Resolved to a non-executable format (exec format error, exitCode=255).
  • keeper image tag pulled ip-172-31-96-57. Kubelet pulled digest sha256:847fff... at 02:48:11 UTC and the container started.

Hypotheses tree

Ruled-out branches are collapsed. Open any one to read the check that eliminated it.

Cluster-wide misconfigurationRuled out

Compare keeper-0 with keeper-1 and keeper-2.

Blast radius was limited to keeper-0. keeper-1 and keeper-2 stayed Running with restartCount=0.

Peer pods, from Kubernetes pod status

keeper-1 and keeper-2 Running, restartCount=0.

Only one pod diverged.

Application, probe or OOM failureRuled out

Read the previous container log.

The container start failed because the entrypoint could not be executed (exec format error, exitCode=255).

Previous container log, from Previous container log

exec /opt/bitnami/scripts/clickhouse-keeper/entrypoint.sh: exec format error, exitCode=255.

The failure is at container start.

Container start failed at exec on node ip-172-31-96-57Confirmed

Follow the timeline from the failed start to recovery.

At 02:47:38 UTC the start on ip-172-31-96-57 failed with exec format error. At 02:48:11 UTC kubelet pulled image digest sha256:847fff... and the container started successfully on the same node, with no reschedule.

Failed start, from Previous container log

02:47:38 UTC, exec format error, exitCode=255, node ip-172-31-96-57.

The start failed on the cached image.

Recovery, from Kubernetes events

02:48:11 UTC, kubelet pulls digest sha256:847fff... and the container starts. No reschedule.

A fresh pull on the same node ended the incident.

Node-local cached image or runtime inconsistency (bad or partial cached layer)Root cause

Explain the fail on cached, succeed after repull pattern.

The pod recovered only after kubelet performed a fresh image pull and then started successfully on the same node. This is most consistent with a node-local cached image or runtime inconsistency (a bad or partial cached layer) rather than an application, probe or OOM failure or a cluster-wide misconfiguration.

Fail on cached, succeed after repull, from Previous container log and Kubernetes events

Exec format error at 02:47:38 UTC, successful start after the 02:48:11 UTC pull, same node.

The fault was in the node's local image state.

Mutable image tag rewritten shortly before the incidentConfirmed

Check for changes to the keeper image tag before the incident.

The keeper image tag was mutable and was rewritten by PutImage at 2026-10-05 02:26:42 UTC, increasing the likelihood of per-node divergence. This explains why only one pod diverged, but does not by itself explain the fail on cached, succeed after repull pattern, so it is a contributing factor.

Tag rewrite, from PutImage event

PutImage at 2026-10-05 02:26:42 UTC.

Nodes could hold different digests under the same tag.

Exact digest used by the failed cached startNeeds access

Identify the digest and content used by the failed start.

Kubernetes events did not expose it, so the root cause is stated as a node-local cached image or runtime inconsistency rather than a specific digest mismatch.

Final RCA

Node-local cached image or runtime inconsistency (bad or partial cached layer)

The pod recovered only after kubelet performed a fresh image pull and then started successfully on the same node. This is most consistent with a node-local cached image or runtime inconsistency (a bad or partial cached layer) rather than an application, probe or OOM failure or a cluster-wide misconfiguration.

Remediation

The incident ended without intervention: at 02:48:11 UTC kubelet pulled image digest sha256:847fff... and the container started on the same node. The RCA names the mutable keeper image tag, rewritten at 02:26:42 UTC, as a contributing factor.

Frequently asked questions

What caused the ClickHouse Keeper CrashLoopBackOff?

On node ip-172-31-96-57 the keeper container could not execute its entrypoint, /opt/bitnami/scripts/clickhouse-keeper/entrypoint.sh, and failed with exec format error and exitCode=255. The pod recovered only after kubelet pulled the image fresh and started it on the same node, which is most consistent with a bad or partial cached image layer on that node.

Why did only one ClickHouse Keeper pod fail?

keeper-1 and keeper-2 stayed Running with restartCount=0, so the blast radius was limited to keeper-0. The keeper image tag was mutable and had been rewritten by PutImage at 02:26:42 UTC, which increases the likelihood of nodes holding different digests under one tag and explains why a single pod diverged.

How did the incident end?

At 02:48:11 UTC kubelet pulled image digest sha256:847fff... and the container started successfully on the same node, with no reschedule. The Grafana Image Pull / CrashLoopBackOff alert fired at 02:49:30 UTC, after the condition had held for its configured FOR=5m window.

Why is the confidence for this root cause medium?

Kubernetes events did not expose the digest or content used by the failed cached start, so a specific digest mismatch cannot be proven. The root cause is therefore stated as a node-local cached image or runtime inconsistency, which fits the fail on cached, succeed after repull pattern without naming the exact content that failed.

Related investigations