AI SRE for Kubernetes
Investigate Kubernetes incidents in minutes, not hours. The Sherlocks AI Watson data agent runs read-only inside your cluster. No standing production credentials. No logs leaving your infrastructure.
sre@prod-bastion ~ $ sherlocks investigate --source kubernetesRESOURCE SIGNAL CORRELATEDpod OOMKilled exit 137 22M-row querypod CrashLoopBackOff revision 1848node DiskPressure no storage requestpvc volume 90% full TSDB cardinality sre@prod-bastion ~ $ sherlocks explain oomkilled-kyc-query
What Sherlocks AI does on Kubernetes
Sherlocks AI runs 16+ specialized agents on top of your Kubernetes cluster to investigate incidents end to end. When a pod crashes, a deployment stalls, or a node runs out of resources, the agents correlate signals across logs, metrics, events, and recent deploys to return a draft root cause analysis in minutes, delivered directly in Slack, ready for the on-call engineer to confirm or edit.
Install via Helm. Read-only access, with Watson running as pods inside your cluster.
Why Kubernetes incidents take hours to explain
Kubernetes is good at telling you that something is wrong and bad at telling you why. kubectl describe pod gives you an exit code and a restart count. The cause is usually somewhere else: in the diff of the deploy that went out twenty minutes earlier, in a node that is short on disk, in a query that read more rows than anyone expected, or in a dependency the pod cannot reach.
So the on-call engineer becomes the correlation engine. Events from the scheduler and the kubelet, container logs, Prometheus or Datadog metrics, rollout history and CI all live in different tools, and each obvious first guess has to be checked and ruled out by hand before the real one is found.
That checking is the part Sherlocks AI takes over. The agents query each signal source in parallel, test the competing hypotheses against the evidence, and return the root cause along with the hypotheses they ruled out, so the engineer can verify the reasoning instead of redoing it.
Common Kubernetes incidents Sherlocks AI investigates
(05 Groups)The agents are trained on real production patterns. These are the failure modes handled out of the box, each group shown next to real investigations from the Kubernetes examples, and the RCA guide and upstream documentation for each.
Pod state failures
- CrashLoopBackOff
- OOMKilled
- ImagePullBackOff
- ErrImagePull
- Pod stuck in Pending
- Pod stuck in ContainerCreating
- Pod stuck in Terminating
The status says how a container died, never why. In the published cases behind this group, an OOMKilled pod traced to a single unbounded query, a CrashLoopBackOff to one release-candidate image, and a pod that never started to a Secret missing from the namespace.
- GuideHow to fix a CrashLoopBackOff
- GuideHow to fix OOMKilled (exit code 137)
- GuideHow to fix CreateContainerConfigError
- Kubernetes docsPod lifecycle
- Kubernetes docsDebugging pods
Worked examples
CrashLoopBackOff: a release candidate image crashed ad-service before it could bind its port
OOMKilled on Kubernetes: how an unbounded KYC query exhausted a 2.5 GB container limit
CreateContainerConfigError: a missing Kubernetes Secret blocked container startup
Scheduling and resource failures
- "0/N nodes are available"
- Readiness probe failed
- Liveness probe failed
- Unschedulable workloads
- Resource quota exceeded
- Insufficient CPU or memory
Resource failures are rarely about one pod. The Trino pods here crash-looped from cluster-wide memory and CPU contention, not the bad image the alert suggested, and the Celery worker ran out of headroom while an Aurora load surge grew its task backlog.
- Kubernetes docsResource management for pods
- Kubernetes docsLiveness, readiness and startup probes
Deployment and rollout failures
- Failed rolling updates
- ReplicaSet not scaling as expected
- Deployment stuck mid-rollout
- Rollback detection and root cause
A rollout changes many things at once, so the question is which change mattered. One outage here was a Deployment scaled to zero replicas. Another was rollout churn, CPU throttling and an HPA hard cap combining to stall a Kafka consumer.
- GuideHow to debug Kafka consumer lag in Kubernetes
- Kubernetes docsDeployments
Networking and storage failures
- DNS resolution issues in-cluster
- Service endpoints not updating
- PVC stuck in Pending
- Volume mount failures
- Network policy misconfigurations
The failing pod is often only where the symptom shows up. A Prometheus PVC filled because metric cardinality grew under a 65-day retention window, and a Celery consumer crash-looped because it could not reach its AMQP broker.
- Kubernetes docsDNS for Services and Pods
- Kubernetes docsPersistent volumes
Node-level issues
- NodeNotReady
- Kubelet failures
- Node pressure (disk, memory, PID)
Node problems surface as pod problems on one machine. A pod was evicted because workloads without ephemeral-storage requests silently overcommitted the node, and a ClickHouse Keeper pod failed on one node while its two peers kept running.
- GuideHow to fix a pod evicted for ephemeral storage
- Kubernetes docsNodes
- Kubernetes docsNode-pressure eviction
How Sherlocks AI works on Kubernetes
(04 Steps)- 1Step 1: Deploy the Watson data agent inside your cluster.
Install via Helm. Watson runs as a pod in a dedicated namespace, with a read-only ClusterRole scoped to what the agent needs: pods, deployments, events, logs, metrics.
- 2Step 2: Connect your observability stack.
Watson reads from your existing observability tools using service accounts. Native integrations for Datadog, New Relic, Prometheus, OpenTelemetry, and in-cluster Kubernetes events.
- 3Step 3: Install the Slack app.
Agents post investigation summaries directly in your incident channel. Engineers ask follow-up questions in natural language.
- 4Step 4: Fire an alert.
When Prometheus, Datadog, or any alerting tool fires, Watson spins up a parallel investigation across 16+ specialized agents, each querying a different signal source. A draft RCA lands in Slack within minutes.
Supported Kubernetes platforms
(02 Platforms)Install with the universal installer, which handles the Helm release and the EKS-specific settings.
Install with the universal installer, which handles the Helm release and the GKE-specific settings.
Running somewhere else? The cloud pages cover AWS, Google Cloud and Azure, and we will walk your platform team through any other cluster on a call.
What the agent reads in your cluster
(08 Sources)Watson reads from inside the cluster and from your connected observability tools.
States, exit codes, restart counts, resource usage.
Streamed via the kubelet API or your log aggregator.
Scheduler, controller manager, kubelet.
Recent rollouts, ReplicaSet changes, config updates.
CPU, memory, disk and network from Prometheus, Datadog or New Relic.
Distributed traces from OpenTelemetry, if available.
Which services talk to which, based on NetworkPolicies and service meshes.
From kubectl apply history, GitOps commits, or CI/CD webhook events.
Kubernetes security and data handling
Watson has zero write permissions in your cluster. It cannot create, modify, or delete any resource. All actions are observation-only.
Watson processes log data inside your cluster. Only the correlated findings (RCA summaries, timeline, deploy diffs) are sent to the Sherlocks AI control plane for display in Slack.
Watson strips personally identifiable information from logs and traces before any data crosses the trust boundary. Configurable rules per team.
Sherlocks AI cloud environment is annually audited against the SOC 2 Type 2 framework.
A real investigation, step by step
This is a published case from the examples section, not a mock-up. The pod was OOMKilled, and the obvious first guesses were wrong. Here is how the investigation got from the alert to the line of code.
- 1The alertCriticalKubernetes / prod-india-exchange2026-04-17 10:39:48 UTC
Pod support-dashboard-staging-7cc9fd7874-gpf82 OOMKilled
137 (OOMKilled)Exit Code2,500 MBMemory Limit22.36M rowsKycStatusLog - 25 hypotheses ruled out
Node-level memory pressure / noisy-neighbor eviction
Node was at 3.12% of 30,808 MB allocated and the kill carried Reason=OOMKilled / Exit 137, which is a cgroup limit breach, not an eviction.
Deployment / config regression introduced the OOM
No rollout at incident time; start command and memory settings unchanged (start_gunicorn.sh, 1500M/2500M). Favors a runtime workload trigger.
Gradual memory leak (recurring) rather than a one-shot spike
RESTARTS=1 (a single event), sibling pods at 0 restarts, and the pod recovered and stayed Running. Event-driven, not a leak.
Memory leak in the show_manual_review_kyc endpoint
Memory is garbage collected properly when small result sets are returned; the issue is the size of the single result set.
Background task consumed excessive memory
No background tasks were running on the pod at the time of the crash.
- 3Root cause
Unbounded query over 22.36M KycStatusLog rows breached 2.5 GB
KycStatusLog holds 22,364,964 rows. At ~100 bytes/row that is ~2.1 GB in-memory; with ~400 MB runtime overhead it consumes the entire 2500M budget. The unbounded KYC query materialized the full rowset and tripped the OOM killer. Root cause; repeatable until the path is paginated/streamed.
- 4The fix
- Paginate the show_manual_review_kyc path with an explicit page size (for example 500 rows) so no single request can materialize the full table.
- Use QuerySet.iterator() with a bounded chunk_size for any code path that must stream large result sets, avoiding full-rowset materialization in the worker.
For the general method, read how to fix an OOMKilled container (exit code 137), or browse every Kubernetes investigation.
Why Sherlocks AI for Kubernetes
16+ specialized agents, each querying a different signal source in parallel. Faster and more accurate than a monolithic agent that tries to do everything.
Your cluster security team can approve Watson in one review. No standing production credentials to rotate.
Native integrations with Datadog, New Relic, Prometheus and OpenTelemetry. Nothing new to deploy.
No new dashboard to open at 3am. The agent lives where your incident response already happens.
Every root cause comes with a confidence score. You decide when to trust and when to verify.
Frequently asked questions
(08 Questions)What Kubernetes incidents can Sherlocks AI investigate?
Pod state failures such as CrashLoopBackOff, OOMKilled, ImagePullBackOff and pods stuck in Pending, ContainerCreating or Terminating; scheduling and resource failures such as failed readiness and liveness probes, exceeded resource quotas and "0/N nodes are available"; failed or stalled rollouts; in-cluster DNS, service endpoint, PVC, volume mount and network policy problems; and node-level issues such as NodeNotReady, kubelet failures and disk, memory or PID pressure.
How is an investigation different from running kubectl describe?
kubectl describe tells you what state a pod is in. Sherlocks AI works out why, by correlating Kubernetes events, container logs, metrics, rollout history and recent changes, and by testing competing hypotheses against that evidence. The result shows the root cause and the hypotheses that were ruled out, so the on-call engineer can check the reasoning rather than repeat it.
Does Sherlocks AI need write access to my cluster?
No. The Watson data agent uses a read-only ClusterRole scoped to what it needs: pods, deployments, events, logs and metrics. It cannot create, modify or delete any resource, cannot execute commands, and cannot deploy changes.
How is the Watson agent installed?
Watson installs with Helm, using the universal installer, which handles the Helm release and the cloud-specific settings. It runs as a pod in a dedicated namespace inside your cluster. The installer supports Amazon EKS and Google Kubernetes Engine today; for any other cluster, book a call and we will walk your platform team through it.
Do my logs leave the cluster?
No. Watson processes log data inside your cluster. Only the correlated findings, such as RCA summaries, the incident timeline and deploy diffs, are sent to the Sherlocks AI control plane for display in Slack. Personally identifiable information is stripped from logs and traces before anything crosses the trust boundary.
Which observability tools does Sherlocks AI work with on Kubernetes?
Prometheus, Datadog, New Relic and OpenTelemetry traces, alongside in-cluster Kubernetes events, container logs and rollout history. Watson reads from the tools you already run, so there is nothing new to deploy.
Where do investigation results show up?
In your incident channel in Slack. When an alert fires, a draft RCA lands in the channel within minutes, and engineers can ask follow-up questions in natural language in the same thread.
How do I know whether to trust a root cause?
Every root cause comes with a confidence score and the evidence behind it, including the hypotheses that were checked and ruled out. You decide when to act on it and when to verify further.
Get started
Start free with 30 investigations per month, or book a 20-minute demo and we will walk you through a live incident replay in your stack.
Questions? Get in touch.
Sherlocks AI is used by Fynd, Topmate, Wafeq, TradeIndia, and Lokal to investigate production incidents. SOC 2 Type 2 certified. Watson data agent runs read-only inside your VPC.