AI SRE · Forecast

What Will AI SRE Look Like in the Next 5 Years?

A plain-language forecast of where AI SRE goes next, with the evidence behind each prediction.

By Sherlocks AIPublished on: Oct 8, 202614 min read

TL;DR

AI SRE is a young category, about two years old. Most of the first products appeared in 2024 and 2025. By late 2026 it has clear buyers and a clear set of competitors. The next 5 years will decide whether AI SRE stays its own category or becomes a feature inside something bigger. This piece lays out where we think it is heading.

Each section below explains one change in plain terms, with the evidence behind it. Each prediction is specific, so you can check it against what actually happens. Want to know where things stand today first? Start with our Top 13 AI SRE Tools in 2026 comparison.

The forecast at a glance

First: the category settles

These are the changes we expect to see in the next one to two years, as buyers and vendors start to agree on what an AI SRE actually is.

AI SRE tools get sorted into four levels of autonomy

Today, most vendors describe their tool in one of two ways: "read-only" (it only looks) or "autonomous" (it acts on its own). We think buyers will soon ask about four clear levels, by name:

  • Read-only. The AI looks at your systems and tells you what it found. You do everything else.
  • Suggest. The AI proposes a fix. A person applies it.
  • Approve-to-apply. The AI applies the fix, but only after someone clicks approve.
  • Autonomous. The AI applies fixes on its own, inside limits your team agreed on in advance.
Tier 1Read-only

Observe and recommend

Human does everything

Tier 2Suggest

Propose actions

Human applies them

Tier 3Approve-to-apply

Agent executes after a single click

Human approves each action

Tier 4Autonomous

Agent executes inside pre-approved guardrails

Human sets the guardrails

Less agent autonomyMore agent autonomy
The four autonomy tiers buyers will screen AI SRE vendors by.

Evidence: Vendors are already picking their spot on this scale. Rootly has openly asked whether teams can stop needing a human to approve fixes ("Could 2027 be the year you turn off human-in-the-loop remediation?", Sep 2026). Cleric starts as read-only and only makes changes through pull requests that a person approves. incident.io's Investigations product focuses on working out the root cause together with the team. Within two years, we expect these level names to be standard, and buying teams to filter vendors by them.

The AI thinks in the vendor's cloud but acts inside yours

Large companies are settling on a setup with two parts. The "brain" is the AI model, its reasoning and its memory. It runs in the vendor's cloud. The "harness" is everything that touches your systems: permissions, running commands and reading data. It runs inside your own network (your VPC, or virtual private cloud).

With this setup, your logs stay inside your infrastructure, and the vendor never holds permanent login access to your production systems.

Vendor cloud

The brain

  • Agent model
  • Reasoning
  • Memory
no standing prod credentials

Customer VPC

The harness

  • Permissions
  • Execution
  • Data access
  • Customer logs never leave
The harness and brain split: reasoning in the vendor's cloud, execution and data access inside the customer's VPC.

Evidence: Sherlocks' Watson data agent is built around this split. It runs read-only inside the customer's VPC. Cleric is read-only by default, and write access is something the customer has to turn on. More and more enterprise AI SRE products work this way. Security teams in regulated industries, like finance and healthcare, expect this setup, and they sign off on the largest deals.

Every AI SRE will write your postmortem for you

A postmortem is the report a team writes after an incident. It explains what broke, when, why, who was affected, and what the team will do so it does not happen again. Writing one usually takes a senior engineer two to four hours. It is also the step teams skip most often, because once the problem is fixed, nobody wants to spend an afternoon writing about it.

Within a couple of years, every AI SRE product will write the first draft for you. The AI reads the incident's Slack thread, builds the timeline from your logs and alerts, adds the key screenshots, and hands you an almost finished report to edit.

Once every tool can do this, buyers will compare the quality of the drafts. Three things decide quality:

  • Is the timeline right? AI often gets the order of events wrong, or blames the wrong deploy.
  • Are the action items useful? A good draft says "add a canary deploy for the payments service." A weak one says "improve monitoring."
  • How much does a person have to rewrite? If the engineer rewrites 80% of it, the feature saved no time at all.

Evidence: Rootly has published several pieces on AI postmortems in a few weeks, including "Can AI Actually Find the Root Cause? Automating Postmortems and RCA" (Sep 2026) and "Why our retrospective template doesn't ask for root cause" (Oct 2026). When a vendor writes about the same topic twice in a few weeks, it is usually because customers keep asking about it. The hardest case is a long, messy incident, such as a whole datacenter going down. A useful draft for an incident like that saves hours. A draft that fills gaps with made-up facts creates new problems.

AI agents make observability bills go up

An AI agent runs far more searches than a person does. When an engineer investigates an incident, they might run a handful of queries. One automated investigation (often called an RCA, or root cause analysis) can run dozens of log searches, metric queries and trace lookups.

So bills from tools like Datadog, New Relic, Honeycomb and Grafana Cloud start going up, even though no person changed how they work. Finance teams will start asking vendors to prove they are not inflating the bill, and someone will publish the first serious study of this "AI SRE observability tax".

Evidence: We are already hearing about this in early customer conversations. The big observability vendors still set prices based on how people use their tools. Teams are also adopting OpenTelemetry, an open-source standard for collecting this data, partly because it gives them more control over cost.

Then: the tools reorganize

Once the basics are settled, the tools around AI SRE start to change shape. We expect these shifts in the middle of the five-year window.

Many specialist AI agents replace one general agent

Teams will run several specialist agents: one for databases, one for Kubernetes, one for the network, one for unusual cloud costs. Each one knows its area deeply. The vendor's job becomes coordinating that team of specialists.

Evidence: PagerDuty launched a set of four AI agents in October 2025, each with its own job: SRE, Scribe (writing things down), Shift (on-call scheduling) and Insights. Komodor offers an AI SRE agent just for Kubernetes, built from several coordinated workflows. Research on multi-agent systems points the same way. We show how specialist agents work in practice in our guide to using AI for Kubernetes incidents.

"Agent Ops" becomes a real job title

Someone has to look after the AI. That means watching the decisions it makes, reviewing its postmortem drafts, changing what it is allowed to do, and noticing when its reasoning starts to slip. Today, senior SREs do this on top of their normal job. We think it becomes a proper role within three years, usually sitting next to the platform engineering team. By the end of the five years, a senior Agent Ops engineer will earn as much as an SRE, or more.

Evidence: This is how DevOps happened between 2010 and 2012. People were doing the work before it had a name. Once it had a title, it became a real career with its own budget. Job boards already show early "AI Operations Engineer" and "ML Reliability" roles appearing close to SRE.

AI SRE joins the deploy pipeline

When something breaks, the first question is usually "which deploy caused this?" Today, getting from the alert to the deploy takes a few steps. Soon it will take none. Deploy tools like Argo, Spinnaker and CircleCI will either buy AI SRE products or build the same ability themselves. The AI will see the code change that caused the incident before a person does.

Evidence: Rootly already writes about linking incidents to deploys. GitHub keeps adding more ways to plug into build pipelines. We walk through this kind of debugging in our CrashLoopBackOff root cause guide. Deploy platforms are well placed to add this.

Observability tools start charging differently for AI traffic

We expect Datadog, New Relic and Grafana Cloud to add separate pricing for queries made by AI agents, either as a new tier or as pay-per-use. The push for this will come from customers. The first vendor to publish clear, AI-aware pricing will win a large share of AI SRE customers.

Evidence: Pricing always follows how a product is used. AI agents use these tools very differently from people, and current pricing was designed around people.

Finally: the job changes

The biggest changes take the longest. By the end of the five years, we expect the people and the rules around reliability work to look different too.

The SRE job title splits into three

"Site Reliability Engineer" is heading the same way "Webmaster" went by 2008. The work splits into three clearer jobs:

  • Platform Engineer: owns infrastructure as code, networking and provisioning.
  • Reliability Engineer: owns SLOs (the reliability targets a service has to meet), on-call and how the team handles incidents.
  • Agent Ops Engineer: owns the AI SRE agents.

Teams that still use the "SRE" title will mostly be small teams where one person does all three jobs, or teams that keep the title because of how pay is set.

Site Reliability Engineer

Platform Engineer

Owns: Infrastructure as code, networking, provisioning

Reliability Engineer

Owns: SLOs, on-call, incident response practice

Agent Ops Engineer

Owns: The AI SRE fleet

The functions behind the SRE title split into three cleaner titles.

Evidence: Job boards in 2026 already show this split starting. Over the last three years, platform engineering has taken over about half of what used to be classic SRE work. The responsibilities described in the original Google SRE book are already being divided up this way in most mature engineering teams.

The AI joins the on-call rotation

In the long run, every page goes to a mixed team. The AI SRE looks first. It adds context to the alert and checks recent deploys. For small incidents, it fixes the problem itself. For everything else, it hands over to the human on-call, with a draft root cause analysis already waiting in the Slack thread.

The AI is treated as a member of the rotation. It has a clear set of things it is allowed to do. It can be taken off the rotation if it behaves badly. Its decisions are reviewed in retrospectives.

Page fires

→

AI SRE agent, first look

Enriches the alert. Correlates with recent deploys.

→

Low-tier incident

Resolved autonomously

Everything else

Escalated to the human on-call, draft RCA already in the Slack thread

The hybrid on-call page: the agent takes first look, then resolves or escalates.

Evidence: AI agents keep getting more capable, and on-call engineers are burning out. Both trends point toward this model. The open question is which vendor will lead it.

Compliance rules start treating AI agents as their own kind of identity

Security and compliance standards like SOC 2, ISO 27001, HIPAA and PCI DSS will treat AI agents as their own category of identity. That means separate rules from both human users and service accounts. Audit logs will need to record the agent's reasoning as well as its actions. Tools that manage these identities, like Oasis Security, will become standard. Rotating an agent's credentials will become as routine as rotating a person's password is today.

Evidence: There is already a group of vendors focused on non-human identities (NHI) in 2026. Oasis Security, which has agreed to be acquired by Cyera, focuses on credentials and time-limited access for AI agents. Compliance standards usually catch up three to five years after a change like this, which puts it inside our five-year window.

A new layer appears above AI SRE

Every few years, a new kind of reliability tool appears that looks at a bigger part of the problem than the last one:

  • Monitoring tells you what broke. "The payments service is down."
  • Observability tells you why it broke. "The payments service is down because the database connection pool is full."
  • AI SRE investigates the whole thing for you. "Payments is down because the connection pool is full. Last Tuesday's deploy caused it. Here is the fix."

The next layer stops focusing on single incidents. It looks for patterns across many incidents and tells you when the design of your system is causing the same incidents to repeat.

For example, imagine a tool that says: "Your region failover has triggered 14 times in six months. Your setup is the cause, and this is what to change." The tool is less concerned with tonight's incident and more with why the same incident keeps coming back.

Early 2000s

Monitoring

What broke?

"The payments service is down."

Late 2010s

Observability

Why did it break?

"The connection pool is exhausted."

2024 to 2025

AI SRE

What caused it, and what is the fix?

"Last Tuesday's deploy. Here is the fix."

Next

The next layer

Why does this keep happening?

"Your topology is wrong."

Each layer of SRE tooling looks at a wider slice of the problem than the one before.

Evidence: This pattern has repeated for more than twenty years. Monitoring tools like Nagios and Zabbix came out in the early 2000s. Observability became its own field in the late 2010s, with vendors like Honeycomb making the word popular. AI SRE took shape in 2024 and 2025. Each new layer took one more kind of work off people's plates. We expect the next one to do the same, though it does not have a name yet.

What stays the same

We expect three things to stay the same over the next five years.

Production stays messy. Automation reduces the mess in production, and some of it will always remain. The incident of the future will still have one engineer who did not read the Slack thread, one dashboard showing the wrong thing, and one deploy that went out before anyone noticed.

People still make the final call. Even when an AI can fix things on its own, a person decides whether it is allowed to. That decision stays with people. In five years, a senior reliability engineer will spend less time fixing incidents and more time deciding how much to trust the AI agents they manage.

Security review still decides every deal. The specific checks will change over time. Any AI SRE vendor that cannot clearly answer "what can this agent access, who approved that, and what gets logged?" will struggle to win enterprise customers, today and in five years. Our Azure integration page shows how we answer this for our own agent.

How to use this forecast

If you are buying an AI SRE tool, use this as a checklist. The questions that will matter in five years start to matter in two. Ask vendors now which autonomy level they sit at, how they keep AI query costs down, and whether they are moving toward specialist agents or sticking with one big agent.

AI SRE vendor scorecard

  1. 1Where on the autonomy spectrum does the agent sit?
  2. 2Where does customer data live during an investigation?
  3. 3How are agent-driven observability queries priced?
  4. 4What controls exist around non-human identity for the agent?
The four questions to ask any AI SRE vendor now, from the FAQ below.

If you are building one, use it as a map. What you build this quarter will ship into the market described in the later sections.

If you are an engineer, use it to plan your career. "Agent Ops" roles are rare on job boards today. We expect them to be common within two years, and people who build these skills early will be well placed.

Frequently asked questions

AI SRE (AI Site Reliability Engineering) means using AI agents to investigate production incidents. The agent looks across logs, metrics, traces and recent deploys, connects the dots, and finds the root cause faster than an on-call engineer could alone. Modern AI SRE platforms like Sherlocks run read-only inside the customer's own network (VPC), so the data stays on the customer's side.

No. The job changes shape. People still decide how much to trust the AI, make the big architecture decisions, and make the final call on when the AI is allowed to act. A senior reliability engineer's job becomes managing AI agents.

Two changes stand out. First, tools get sorted into four autonomy levels: read-only, suggest, approve-to-apply and autonomous. Second, the harness and brain split becomes the standard setup: the AI's reasoning runs with the vendor, while everything that touches your systems runs inside your network. Together, these will shape how companies buy AI SRE in the next two to three years.

Agent Ops is the work of looking after AI SRE agents in production, and an emerging job title for the people who do it. It means watching the agent's decisions, reviewing drafts like postmortems, changing what the agent is allowed to do, and noticing when its reasoning gets worse. Today the role sits next to platform engineering. We expect it to become a proper job title within three years.

Ask four questions. One: which autonomy level does the agent work at? Two: where does our data live during an investigation? Three: how are the AI's observability queries priced? Four: what controls are in place for the agent's own identity and access?

Yes. Within a couple of years, every AI SRE product will write the first draft. The AI reads the incident's Slack thread, builds the timeline from logs and alerts, and gives you an almost finished report to edit. Since every tool will do this, buyers will compare draft quality: is the timeline right, are the action items useful, and how much does a person have to rewrite?

Yes. One automated investigation can run dozens of log searches, metric queries and trace lookups, where a person would run a handful. So bills from tools like Datadog, New Relic, Honeycomb and Grafana Cloud go up even though nobody changed how they work. We expect observability vendors to add separate pricing for AI agent queries as a result.

The one line to end on

Five years from now, AI SRE will look less like a product you buy and more like a layer of how you run production. Buyers will care less about whether a product uses AI, and more about what the agent can do, what it can access, how much damage it could cause, and what it owns.

We plan to revisit these predictions and check how they held up.

Related reading

See an AI SRE work a real incident

Book 30 minutes with our team and watch an investigation run on your own stack.

Book a demo →