Kubernetes · AWS · Databases · Linux VMs

Root cause analysis guides

Root cause analysis guides for the failures that page on-call engineers. Each covers one failure mode end to end: what the error means, the commands that confirm it, and the configuration change that stops it recurring.

oncall@prod — 03:14
sre@prod-bastion ~ $ column -t ~/oncall/last-four.txtSTACK         SYMPTOM          ROOT CAUSEkubernetes    pod Evicted      node ephemeral-storageaws           EBS latency 4x   gp3 throughput ceilingdatabases     pool exhausted   plan regression, migrationlinux-vm      / at 99%         logrotate rule dropped sre@prod-bastion ~ $ cat ~/oncall/kubernetes-eviction.md

Start with your stack

(04 Stacks, 05 Guides Published)

Looking for worked incidents instead?

A guide explains a class of failure. The examples section is the other half: real production incidents investigated end to end, including the hypotheses that were ruled out along the way.

See Sherlocks AI in action

Watch an AI SRE work a real incident from alert to root cause, on your stack, in 30 minutes.