AI Incident Investigation for AWS
Sherlocks automates AWS incident investigation from alert to root cause.
It maps affected services and dependencies, correlates operational signals, and tests likely causes using:
- CloudWatch metrics
- Application and infrastructure logs
- Distributed traces
- Deployments and code changes
- Databases and queues
- Kubernetes signals
Teams receive an evidence-backed AWS root cause analysis containing:
- Primary cause and confidence level
- Contributing factors
- Affected services and blast radius
- Reconstructed event timeline
- Supporting telemetry
- Recommended next steps
Findings are delivered directly in Slack.
Investigate an AWS incident with Sherlocks
- Typical alert analysis: 2–3 minutes
- Complex investigations: 5–6 minutes
- Potential MTTR reduction: Up to 70%
AWS Root Cause Analysis Across Telemetry, Dependencies, and Changes
An AWS alarm identifies a symptom. Sherlocks investigates the surrounding system to determine what happened, which services were affected, and why.
Starting from an alert or production question, Sherlocks:
- Identifies abnormal signals and impacted AWS resources.
- Follows application, infrastructure, database, and queue dependencies.
- Correlates logs, metrics, traces, deployments, and configuration changes.
- Generates possible root-cause hypotheses.
- Tests each hypothesis against available evidence.
- Ranks likely causes by confidence and impact.
- Reconstructs the incident timeline and blast radius.
- Recommends remediation or rollback steps.
Evidence, not unexplained conclusions
Sherlocks clearly separates hypotheses from supporting evidence.
Each finding links to relevant:
- Logs and metrics
- Traces and dashboards
- Deployments and pipeline runs
- Commits and code changes
- Infrastructure and application signals
Responders can verify the reasoning instead of relying on an unexplained AI answer.
Visit Sherlocks AI
Investigate common AWS incidents
Sherlocks helps teams investigate:
- Deployment regressions
- API latency and elevated error rates
- ECS service degradation
- Lambda failures
- RDS performance problems
- SQS backlogs and dead-letter queues
- Kubernetes crash loops
- Resource saturation
- Multi-service failures
Automated AWS Troubleshooting Across Services and Dependencies
AWS incidents rarely remain isolated to one resource.
A deployment might:
- Increase database queries
- Exhaust RDS connections
- Slow dependent APIs
- Trigger alarms across multiple services
Sherlocks uses its Awareness Graph to connect these signals and reveal the causal chain.
Continuously updated system context
The Awareness Graph maps:
- Applications and AWS infrastructure
- Service and resource dependencies
- Databases, queues, and external APIs
- Kubernetes workloads and communication patterns
- Normal system behavior
- Deployments and code changes
- Previous incidents and resolutions
- Runbooks and technical documentation
- Incident context from Slack
During an investigation, this topology helps Sherlocks move through:
Reported symptom → affected service → dependency → recent change → likely cause
Change correlation
Sherlocks correlates the incident window with:
- Commits and pull requests
- Builds and pipeline runs
- Releases and deployments
- Configuration updates
- Scaling events
This can reveal causes such as:
- A missing environment variable
- A partial deployment
- A failed pipeline step
- Replicas configured to zero
- An inefficient database query
- An application change behind an infrastructure symptom
Incident memory
Sherlocks preserves useful context from previous incidents, including:
- Past root causes
- Effective fixes
- Investigation trails
- Team decisions
- Impacted services
- Relevant runbooks
Previous incident knowledge can inform new investigations and reduce repeated troubleshooting.
AWS Telemetry and Integration Coverage
Sherlocks investigates AWS infrastructure together with the application and operational systems around it.
AWS services
- Amazon CloudWatch
- Amazon EC2
- Amazon ECS
- AWS Lambda
- Amazon RDS
- Amazon S3
- Amazon SQS
Telemetry
- CloudWatch metrics and resource metadata
- Application and infrastructure logs
- Load-balancer logs
- APM and distributed traces
- Database health signals
- Queue health signals
- Kubernetes state and events
- Deployment events
- Cloud resource data
Kubernetes workloads on AWS
- Pods
- Services
- Deployments
- Nodes
- Events
- Logs
- Resource metrics
Observability
- Prometheus
- Datadog
- New Relic
- Sentry
- Elastic
- Coralogix
- ELK
- Loki
- OpenTelemetry
Code and delivery
- GitHub
- GitHub Actions
- Jenkins
- Azure Pipelines
Incident response
Sherlocks queries relevant telemetry when an investigation requires it.
This reduces the need for responders to manually search across:
- AWS consoles
- Monitoring dashboards
- Log platforms
- Deployment systems
- Source control
- Incident channels
Sherlocks’ infrastructure graph supports:
- Multi-region environments
- Multi-Availability Zone environments
- Multi-cluster Kubernetes environments
Every AWS root cause analysis can include:
- Primary cause and confidence
- Contributing factors
- Affected services
- Blast radius
- Event timeline
- Evidence links
- Recommended actions
A Read-Only AI SRE for AWS
Sherlocks acts as an AI SRE for AWS incident investigation while engineers remain in control of remediation.
Least-privilege investigation
The Watson data agent uses read-only permissions.
It cannot:
- Modify AWS infrastructure
- Deploy code or configuration changes
- Retrieve secrets
- Access application data stored in databases
- Access application data stored in queues
It collects only the telemetry and metadata needed for investigation, such as:
- Resource metrics
- Error types
- Connection counts
- Query timing
- Replication lag
- Queue depth
- Message age
Flexible deployment
Sherlocks can run:
- As a cloud-hosted service
- With Watson inside your VPC
- Fully within your environment
- With AWS Bedrock in your AWS account
Data is protected using:
- TLS 1.3 in transit
- AES-256 at rest
Investigation and collaboration in Slack
Sherlocks delivers each investigation with:
- The most likely root cause
- Supporting evidence
- Impacted services
- Investigation timeline
- Recommended next steps
Responders can:
- Ask follow-up questions
- Review the investigation trail
- Tag service owners
- Share evidence
- Hand off complete incident context
Find the cause of your next AWS incident in minutes.
Book a Sherlocks demo