Book a Demo
AI SRE for Datadog

From Datadog alert
to root cause.

Stop piecing incidents together by hand. Sherlocks investigates your Datadog alerts across the stack and returns a root cause, supporting evidence, and next steps.

Your observability stays. The manual investigation shrinks.

Read-only access · No re-instrumentation

Datadog alert

Checkout latency is rising

Sherlocks finding

Connection pool reduced in the latest deployment.

Deployment changes and traces point to the same cause.

Recommended next step

Review the connection pool change

Find the cause beyond the alert Connect Datadog to the rest of your stack.
Get answers you can verify Review the evidence behind each finding.
Keep your team in control Your engineers choose the response.
An investigation, not another dashboard

See the path from alert to answer.

An AI SRE investigation that follows the evidence—and shows its work.

01 Start from the alert

Identify the affected service, incident window, and connected systems from the Datadog alert.

02 Generate and test hypotheses

Test possible causes against Datadog telemetry, infrastructure state, service dependencies, deployments, code changes, and similar incidents.

03 Rank the likely causes

Rank causes by likelihood and impact. Rule out explanations the evidence does not support.

04 Deliver an evidence-backed RCA

Deliver a reviewable RCA with the cause, confidence level, timeline, affected services, blast radius, evidence, and recommended next steps.

Engineers can inspect the investigation and ask follow-up questions in Slack before deciding how to respond.

Incident investigation
D
Datadog alert
checkout-api · elevated latency
A new deployment goes livecheckout-api configuration changes.
Database wait time increasesTraces show requests waiting for a connection.
Datadog latency alert firesSherlocks correlates the alert with the earlier changes.
AI incident investigation for Datadog alerts

Give on-call engineers
a head start.

Datadog tells your team when production behavior changes. Sherlocks investigates why.

Know what went wrong

Pinpoint what changed, what failed first, and how far the impact spreads.

See why it happened

Trace every finding back to metrics, logs, traces, and deployment context.

Decide what to do next

Review mitigation options and follow-up actions. Your engineers choose the response.

How Sherlocks investigates a Datadog alert

A Datadog alert starts an AI SRE investigation across your production stack.

Build context
Map service dependencies, deployment history, and previous incidents.
Test the evidence
Correlate Datadog metrics, logs, APM traces, dashboards, and events with Kubernetes, cloud infrastructure, databases, and code changes.
Explain the cause
Return the likely root cause, contributing factors, blast radius, and recommended next steps.
Follow every signal

Root cause analysis.
Built on your Datadog telemetry.

Turn the Datadog metrics, logs, and traces you already collect into root cause evidence.

Infrastructure metrics

Compare behavior across the incident window, identify abnormal resource patterns, and determine which signal changed first.

What’s in your root cause analysis?
The cause
Primary root cause, confidence level, and contributing factors.
The evidence
Incident timeline, affected services, blast radius, and links to metrics, logs, dashboards, or commits.
The response
Recommended mitigation, remediation steps, and follow-up actions for engineers to review.
Cross-stack AI troubleshooting

The alert is in Datadog.
The cause could be anywhere.

Connect the symptom in Datadog to the cause in your infrastructure, databases, code, or queues.

Kubernetes AWS Google Cloud PostgreSQL GitHub Kafka Redis Grafana
Explore connected systems

Investigate the cause beyond Datadog, without moving every signal into one platform.

Infrastructure
Kubernetes · AWS · Google Cloud · Microsoft Azure
Databases
PostgreSQL · MySQL · MongoDB · Redis
Queues
Kafka · RabbitMQ · Cloud queues
Code & CI/CD
GitHub · Deployment history · GitHub Actions · Jenkins · Other CI/CD systems
Observability
Prometheus · Grafana · Elasticsearch · Loki · Sentry
Team context
Previous incidents · Runbooks · Team knowledge
Your production. Your decisions.

An AI SRE that investigates.
Not one that changes production.

Autonomous investigation.
Human-controlled remediation.

  • Read-only by design. Watson gathers evidence with least-privilege access.
  • No production execution. It cannot deploy code, modify infrastructure, or change databases.
  • Your team makes the call. Engineers review findings and choose the response.
Read-only access. Engineer-controlled changes.
Access
Watson uses least-privilege, read-only permissions. It cannot execute commands, deploy code, modify infrastructure, change databases, or operate queues.
Recommendations
Mitigation options, configuration changes, runbook steps, rollback considerations, and longer-term fixes.
Control
Your engineers review the evidence, choose the response, and make production changes.
Watson access Read-only

Investigate autonomously.
Keep engineers in control.

Gather evidence Read-only
Test hypotheses Read-only
Production changes Human-controlled
Every conclusion has a trail

Evidence Engineers Can Inspect

Give every finding a paper trail.

Review the evidence, challenge a hypothesis in Slack, and save the context for the next incident.

What your team can review and reuse
Verify
Follow the investigation path through logs, metrics, deployment context, and evidence links.
Collaborate
Ask follow-up questions, test other hypotheses, and share findings with service owners in Slack.
Reuse
Keep the incident record and operational context for future investigations.
Investigation record
What happened Root cause · Confidence · Contributing factors
How we know Timeline · Logs · Metrics · Deployment context
What to do next Mitigation · Runbooks · Follow-up actions
Continue the investigation in Slack Ask a follow-up. Explore another hypothesis.
More context. Less manual work.

Why SRE Teams Use Sherlocks Alongside Datadog

01

Reduce manual correlation

Start with a structured investigation instead of assembling context from every tool by hand.

02

Investigate beyond the triggering service

Follow a symptom beyond the alerting service to the systems that caused it.

03

Preserve operational knowledge

Bring previous incidents, team knowledge, and runbooks into the next investigation.

04

Keep existing observability investments

Keep Datadog and your existing stack. Add AI investigation without an observability migration.

05

Keep production changes controlled

Let Sherlocks investigate. Keep production decisions with the engineers responsible for them.

Keep the stack you already trust

Connect Sherlocks AI to Datadog

01

Connect Datadog

Give Sherlocks read-only access to your existing Datadog telemetry.

02

Connect your stack

Add the infrastructure, databases, and deployment context around it.

03

Route your alerts

Let production alerts launch investigations automatically.

How the Datadog integration connects
01 · Datadog
Use read-only API access for metrics, events, and dashboards; include logs and APM traces as evidence.
02 · Your stack
Connect infrastructure, databases, cloud services, code repositories, CI/CD, and Slack.
03 · Alerts
Route Datadog alerts to Sherlocks to launch investigations automatically.
See Sherlocks work with Datadog. Walk through an evidence-backed incident investigation.
Book a Demo
A few more answers

Sherlocks AI and Datadog FAQ

Does Sherlocks AI integrate with Datadog?

Yes. Sherlocks uses read-only access to investigate Datadog infrastructure metrics, logs, APM traces, dashboards, and events.

Does Sherlocks replace Datadog?

No. Keep Datadog for observability. Sherlocks adds AI incident investigation and root cause analysis across your production stack.

Can a Datadog alert automatically trigger an investigation?

Yes. Datadog alerts can automatically launch investigations that gather evidence, test hypotheses, and return a reviewable RCA.

Can Sherlocks perform root cause analysis using Datadog data?

Yes. Sherlocks tests root-cause hypotheses against Datadog telemetry, then correlates findings with infrastructure, service dependencies, database health, deployments, code changes, and previous incidents.

Is Sherlocks an autonomous SRE tool?

Sherlocks investigates autonomously and recommends remediation. Watson has read-only permissions; production changes remain with your engineers.

Can Sherlocks investigate systems outside Datadog?

Yes. Investigations can span Kubernetes, cloud infrastructure, databases, queues, code repositories, CI/CD, other observability tools, and team knowledge.

Put your alerts to work

Your next Datadog alert.
A clearer path to resolution.

See how Sherlocks turns Datadog telemetry and cross-stack context into an investigation your engineers can review.