CrashLoopBackOff · OOMKilled · Pending · Rollouts · Nodes

AI SRE for Kubernetes

Investigate Kubernetes incidents in minutes, not hours. The Sherlocks AI Watson data agent runs read-only inside your cluster. No standing production credentials. No logs leaving your infrastructure.

sherlocks / kubernetes production
sre@prod-bastion ~ $ sherlocks investigate --source kubernetesRESOURCE      SIGNAL              CORRELATEDpod           OOMKilled exit 137  22M-row querypod           CrashLoopBackOff    revision 1848node          DiskPressure        no storage requestpvc           volume 90% full     TSDB cardinality sre@prod-bastion ~ $ sherlocks explain oomkilled-kyc-query

What Sherlocks AI does on Kubernetes

Sherlocks AI runs 16+ specialized agents on top of your Kubernetes cluster to investigate incidents end to end. When a pod crashes, a deployment stalls, or a node runs out of resources, the agents correlate signals across logs, metrics, events, and recent deploys to return a draft root cause analysis in minutes, delivered directly in Slack, ready for the on-call engineer to confirm or edit.

Install via Helm. Read-only access, with Watson running as pods inside your cluster.

Why Kubernetes incidents take hours to explain

Kubernetes is good at telling you that something is wrong and bad at telling you why. kubectl describe pod gives you an exit code and a restart count. The cause is usually somewhere else: in the diff of the deploy that went out twenty minutes earlier, in a node that is short on disk, in a query that read more rows than anyone expected, or in a dependency the pod cannot reach.

So the on-call engineer becomes the correlation engine. Events from the scheduler and the kubelet, container logs, Prometheus or Datadog metrics, rollout history and CI all live in different tools, and each obvious first guess has to be checked and ruled out by hand before the real one is found.

That checking is the part Sherlocks AI takes over. The agents query each signal source in parallel, test the competing hypotheses against the evidence, and return the root cause along with the hypotheses they ruled out, so the engineer can verify the reasoning instead of redoing it.

Common Kubernetes incidents Sherlocks AI investigates

(05 Groups)

The agents are trained on real production patterns. These are the failure modes handled out of the box, each group shown next to real investigations from the Kubernetes examples, and the RCA guide and upstream documentation for each.

01

Pod state failures

  • CrashLoopBackOff
  • OOMKilled
  • ImagePullBackOff
  • ErrImagePull
  • Pod stuck in Pending
  • Pod stuck in ContainerCreating
  • Pod stuck in Terminating

The status says how a container died, never why. In the published cases behind this group, an OOMKilled pod traced to a single unbounded query, a CrashLoopBackOff to one release-candidate image, and a pod that never started to a Secret missing from the namespace.

02

Scheduling and resource failures

  • "0/N nodes are available"
  • Readiness probe failed
  • Liveness probe failed
  • Unschedulable workloads
  • Resource quota exceeded
  • Insufficient CPU or memory

Resource failures are rarely about one pod. The Trino pods here crash-looped from cluster-wide memory and CPU contention, not the bad image the alert suggested, and the Celery worker ran out of headroom while an Aurora load surge grew its task backlog.

03

Deployment and rollout failures

  • Failed rolling updates
  • ReplicaSet not scaling as expected
  • Deployment stuck mid-rollout
  • Rollback detection and root cause

A rollout changes many things at once, so the question is which change mattered. One outage here was a Deployment scaled to zero replicas. Another was rollout churn, CPU throttling and an HPA hard cap combining to stall a Kafka consumer.

04

Networking and storage failures

  • DNS resolution issues in-cluster
  • Service endpoints not updating
  • PVC stuck in Pending
  • Volume mount failures
  • Network policy misconfigurations

The failing pod is often only where the symptom shows up. A Prometheus PVC filled because metric cardinality grew under a 65-day retention window, and a Celery consumer crash-looped because it could not reach its AMQP broker.

05

Node-level issues

  • NodeNotReady
  • Kubelet failures
  • Node pressure (disk, memory, PID)

Node problems surface as pod problems on one machine. A pod was evicted because workloads without ephemeral-storage requests silently overcommitted the node, and a ClickHouse Keeper pod failed on one node while its two peers kept running.

How Sherlocks AI works on Kubernetes

(04 Steps)
  1. 1Step 1: Deploy the Watson data agent inside your cluster.

    Install via Helm. Watson runs as a pod in a dedicated namespace, with a read-only ClusterRole scoped to what the agent needs: pods, deployments, events, logs, metrics.

  2. 2Step 2: Connect your observability stack.

    Watson reads from your existing observability tools using service accounts. Native integrations for Datadog, New Relic, Prometheus, OpenTelemetry, and in-cluster Kubernetes events.

  3. 3Step 3: Install the Slack app.

    Agents post investigation summaries directly in your incident channel. Engineers ask follow-up questions in natural language.

  4. 4Step 4: Fire an alert.

    When Prometheus, Datadog, or any alerting tool fires, Watson spins up a parallel investigation across 16+ specialized agents, each querying a different signal source. A draft RCA lands in Slack within minutes.

Supported Kubernetes platforms

(02 Platforms)
Amazon EKS

Install with the universal installer, which handles the Helm release and the EKS-specific settings.

Google Kubernetes Engine (GKE)

Install with the universal installer, which handles the Helm release and the GKE-specific settings.

Running somewhere else? The cloud pages cover AWS, Google Cloud and Azure, and we will walk your platform team through any other cluster on a call.

What the agent reads in your cluster

(08 Sources)

Watson reads from inside the cluster and from your connected observability tools.

Pod data

States, exit codes, restart counts, resource usage.

Container logs

Streamed via the kubelet API or your log aggregator.

Kubernetes events

Scheduler, controller manager, kubelet.

Deployment history

Recent rollouts, ReplicaSet changes, config updates.

Metrics

CPU, memory, disk and network from Prometheus, Datadog or New Relic.

Traces

Distributed traces from OpenTelemetry, if available.

Service topology

Which services talk to which, based on NetworkPolicies and service meshes.

Recent changes

From kubectl apply history, GitOps commits, or CI/CD webhook events.

Kubernetes security and data handling

Read-only by default.

Watson has zero write permissions in your cluster. It cannot create, modify, or delete any resource. All actions are observation-only.

No logs leave your infrastructure.

Watson processes log data inside your cluster. Only the correlated findings (RCA summaries, timeline, deploy diffs) are sent to the Sherlocks AI control plane for display in Slack.

PII stripping at the edge.

Watson strips personally identifiable information from logs and traces before any data crosses the trust boundary. Configurable rules per team.

SOC 2 Type 2 certified.

Sherlocks AI cloud environment is annually audited against the SOC 2 Type 2 framework.

A real investigation, step by step

This is a published case from the examples section, not a mock-up. The pod was OOMKilled, and the obvious first guesses were wrong. Here is how the investigation got from the alert to the line of code.

  1. 1The alert
    CriticalKubernetes / prod-india-exchange2026-04-17 10:39:48 UTC

    Pod support-dashboard-staging-7cc9fd7874-gpf82 OOMKilled

    137 (OOMKilled)
    Exit Code
    2,500 MB
    Memory Limit
    22.36M rows
    KycStatusLog
  2. 25 hypotheses ruled out
    • Node-level memory pressure / noisy-neighbor eviction

      Node was at 3.12% of 30,808 MB allocated and the kill carried Reason=OOMKilled / Exit 137, which is a cgroup limit breach, not an eviction.

    • Deployment / config regression introduced the OOM

      No rollout at incident time; start command and memory settings unchanged (start_gunicorn.sh, 1500M/2500M). Favors a runtime workload trigger.

    • Gradual memory leak (recurring) rather than a one-shot spike

      RESTARTS=1 (a single event), sibling pods at 0 restarts, and the pod recovered and stayed Running. Event-driven, not a leak.

    • Memory leak in the show_manual_review_kyc endpoint

      Memory is garbage collected properly when small result sets are returned; the issue is the size of the single result set.

    • Background task consumed excessive memory

      No background tasks were running on the pod at the time of the crash.

  3. 3Root cause

    Unbounded query over 22.36M KycStatusLog rows breached 2.5 GB

    KycStatusLog holds 22,364,964 rows. At ~100 bytes/row that is ~2.1 GB in-memory; with ~400 MB runtime overhead it consumes the entire 2500M budget. The unbounded KYC query materialized the full rowset and tripped the OOM killer. Root cause; repeatable until the path is paginated/streamed.

  4. 4The fix
    • Paginate the show_manual_review_kyc path with an explicit page size (for example 500 rows) so no single request can materialize the full table.
    • Use QuerySet.iterator() with a bounded chunk_size for any code path that must stream large result sets, avoiding full-rowset materialization in the worker.
Open the full investigation, with the evidence graph

For the general method, read how to fix an OOMKilled container (exit code 137), or browse every Kubernetes investigation.

Why Sherlocks AI for Kubernetes

Specialist agents, not a generalist.

16+ specialized agents, each querying a different signal source in parallel. Faster and more accurate than a monolithic agent that tries to do everything.

Read-only by design.

Your cluster security team can approve Watson in one review. No standing production credentials to rotate.

Works with the stack you already have.

Native integrations with Datadog, New Relic, Prometheus and OpenTelemetry. Nothing new to deploy.

Slack-native.

No new dashboard to open at 3am. The agent lives where your incident response already happens.

Honest about what AI can and cannot do.

Every root cause comes with a confidence score. You decide when to trust and when to verify.

Frequently asked questions

(08 Questions)

What Kubernetes incidents can Sherlocks AI investigate?

Pod state failures such as CrashLoopBackOff, OOMKilled, ImagePullBackOff and pods stuck in Pending, ContainerCreating or Terminating; scheduling and resource failures such as failed readiness and liveness probes, exceeded resource quotas and "0/N nodes are available"; failed or stalled rollouts; in-cluster DNS, service endpoint, PVC, volume mount and network policy problems; and node-level issues such as NodeNotReady, kubelet failures and disk, memory or PID pressure.

How is an investigation different from running kubectl describe?

kubectl describe tells you what state a pod is in. Sherlocks AI works out why, by correlating Kubernetes events, container logs, metrics, rollout history and recent changes, and by testing competing hypotheses against that evidence. The result shows the root cause and the hypotheses that were ruled out, so the on-call engineer can check the reasoning rather than repeat it.

Does Sherlocks AI need write access to my cluster?

No. The Watson data agent uses a read-only ClusterRole scoped to what it needs: pods, deployments, events, logs and metrics. It cannot create, modify or delete any resource, cannot execute commands, and cannot deploy changes.

How is the Watson agent installed?

Watson installs with Helm, using the universal installer, which handles the Helm release and the cloud-specific settings. It runs as a pod in a dedicated namespace inside your cluster. The installer supports Amazon EKS and Google Kubernetes Engine today; for any other cluster, book a call and we will walk your platform team through it.

Do my logs leave the cluster?

No. Watson processes log data inside your cluster. Only the correlated findings, such as RCA summaries, the incident timeline and deploy diffs, are sent to the Sherlocks AI control plane for display in Slack. Personally identifiable information is stripped from logs and traces before anything crosses the trust boundary.

Which observability tools does Sherlocks AI work with on Kubernetes?

Prometheus, Datadog, New Relic and OpenTelemetry traces, alongside in-cluster Kubernetes events, container logs and rollout history. Watson reads from the tools you already run, so there is nothing new to deploy.

Where do investigation results show up?

In your incident channel in Slack. When an alert fires, a draft RCA lands in the channel within minutes, and engineers can ask follow-up questions in natural language in the same thread.

How do I know whether to trust a root cause?

Every root cause comes with a confidence score and the evidence behind it, including the hypotheses that were checked and ruled out. You decide when to act on it and when to verify further.

Get started

Start free with 30 investigations per month, or book a 20-minute demo and we will walk you through a live incident replay in your stack.

Questions? Get in touch.

Sherlocks AI is used by Fynd, Topmate, Wafeq, TradeIndia, and Lokal to investigate production incidents. SOC 2 Type 2 certified. Watson data agent runs read-only inside your VPC.