Sherlocks AI Changelog

Every AI SRE feature, integration, and platform improvement we ship. Updated with each release.

July 2026

Sherlocks AI 2.0 is here

Our biggest release yet, built entirely on your feedback. Sherlocks AI 2.0 is more transparent about what it knows and how it reasons, keeps every incident, signal, and action item in one place, and works as a real teammate you can DM directly and investigate alongside during an incident.

Watch the Sherlocks AI 2.0 demo

System Understanding

Infra Graph (/sher-infra-graph)·Your whole system in one view: connections, topology, recommendations from static analysis, and tribal knowledge and memories together. It surfaces an understanding % per system, so you can see how well Sherlocks knows your stack and review, fix, or improve investigations from there.

Collaborative Investigation

1:1 DMs·Sherlocks AI is no longer only alert-driven. DM it directly and ask anything about a live incident, a service, past incidents, or why something is behaving the way it is, and it answers right there using everything it knows from the infra graph and its investigations.
Collaborative RCA·Sherlocks works the incident with you: it lays out everything it has already deduced and the evidence it collected, lets on-call add new hypotheses to test or cancel ones that don't make sense, and takes in your own evidence so it can look at signals it isn't even connected to.

Remediation & Review

Remediations & Mitigations·Detailed remediation suggestions for what it finds, a complete step-by-step plan once you pick one, and draft PRs opened across multiple repos by reading your codebase. Now listens to Sentry alerts too.
Daily Incident Review (/sher-daily-incident-review)·One dashboard for everything: every incident triggered, MTTD, MTTR, evidence, action items, and more.
Feb – Mar 2026

Agent success rate: 35% to 74.8%

Agent success rate jumped from 35% to 74.8%. Eight new integrations, full ECS support, multi-region infra graph, and investigation time cut nearly in half.

Agent Success Rate

35.5%74.8%

Tool Call Success

57.3%72%

Investigation Time (p75)

15 min8 min

Alert Ingestion

43%65%

Classification Cost

↓ 70%

Conclusive RCAs

55%61%

AI SRE Agent Intelligence

Alert Context Agent·Pulls memory context across all connected platforms and past incidents before every investigation.
Alert Classification via Cypher Queries·Topology-aware classification over the graph. 30% faster, 70% cheaper.
Infra Q&A Agent v2·Rebuilt with structured formatting for the infra graph UI. More precise, more readable answers.
Agent Tool Call Tracking·Every tool call tracked for reliability metrics, debugging, and accountability.
OpenTelemetry Tracing·Full tracing and monitoring integrated into the agent runtime.

Infrastructure Topology Graph

ECS Full Support·ECS entities integrated with Redis and CubeAPM mapping, debugging skills, and hypothesis tree generation.
Multi-Region / Multi-Cluster / Multi-AZ·Cross-region, multi-cluster Kubernetes, multi-AZ support. Critical enterprise milestone.
Slack Memories in the Graph·Past incident conversations and runbook references embedded directly into the infra graph.
MongoDB Atlas + RabbitMQ + External API nodes·All three added with full edge support for richer topology context.
Full infrastructure topology graph showing service dependencies across clusters

Multi-region infrastructure topology view

Service detail view showing dependency connections for a single service

Service-level detail with dependency connections

Platform & Incident Management

Incidents with Context·Investigation trigger passes full alert context to the agent pipeline, not just the raw alert.
Impacted Entity Tracking·Every investigation stores impacted entity details for a richer audit trail.
Incident Conversation API·New API for listing incident conversations with namespace scoping.
GitHub + Slack as Data Providers·Both added as platform-level data sources for broader investigation context.

Performance & Reliability Fixes

Steampipe Query Optimization·Columns scoped to minimum required to eliminate permission blockers.
CubeAPM Service Normalization·Deduplicates and collapses similar services before graph enrichment.
ELK Latency Query Fix·Corrected metric source from transactions to spans for accurate APM data.
Throughput + MySQL + RDS fixes·Fixed interval conversions, host config key mapping, and RDS cluster resolution.

New AI SRE Integrations

IntegrationWhat it enables
Elastic APMFull dependency graph with K8s-to-Elastic service mapping; latency/throughput by transaction + percentile
MySQLSelf-hosted MySQL on VMs with discovery, index inspection, read-only query execution, K8s mapping
MongoDB AtlasManaged cloud MongoDB with projects, clusters, processes, time-series metrics
MongoDB (self-hosted)Raw query execution with safe formatting, full metrics provider
GrafanaAlert rules, firing alerts, alert-rule-by-UID for full alerting surface
GCPGCP integration support + kubectl tool + Steampipe query support
CubeAPMMetric provider, alert retrieval, service normalization, K8s mapping
PagerDutySlack message capturing and lifecycle management
Infra Q&A Agent answering questions about Kubernetes clusters and services with structured responses

Infra Q&A Agent v2 with structured responses and actionable recommendations

Get notified when we ship

Subscribe to get AI SRE platform updates delivered to your inbox.