All case studies
Fynd logoCase Study

How Fynd cut incident resolution time by 70% with autonomous AI SRE

Reliance Retail's e-commerce platform replaced human-driven triage via AI tools with always-on Sherlocks AI agents, freeing up support and engineering teams while cutting time-to-resolution across the stack.

Most of our resolution time was spent extracting context and manually steering Claude through the investigation. Now the investigation is already running by the time we look at the ticket.
70%
reduction in MTTR
63%
support team hours saved
77%
infrastructure and engineering hours saved

Figures as reported by Fynd.

Now the investigation is already running by the time we look at the ticket. Sherlocks AI takes us straight to the root cause instead of handing us more dashboards, and we've been able to retire a few point tools along the way.
Amboj Goyal, Head of Infrastructure, Fynd

About Fynd

Fynd is an AI-native unified commerce platform built to help brands run their businesses more efficiently. With AI-first solutions across online commerce, in-store technology and supply chain operations, Fynd brings entire business functions together into one connected system. Designed for scale, the platform helps teams launch faster, operate smarter and deliver consistent customer experiences across every channel.

Headquartered in India and backed by Reliance, Fynd is expanding across the GCC, UK, Africa and Southeast Asia to support the next phase of modern commerce.

The engineering organisation is 800+ strong, runs primarily on GCP, and operates its own observability stack built on Grafana/Prometheus and Tempo.

Because Fynd's product powers commerce at scale, every failing transaction represents high-value business impact. Reliability isn't a nice-to-have. It maps directly to revenue.

The challenge: strong signals, slow answers

Fynd's incident response ran on two intake paths.

  1. 1Tickets came in directly from B2B customers whose end-users could not complete actions on their platforms.
  2. 2Alerts came from Fynd's own observability stack, either as tickets or as PagerDuty calls.

Both paths funneled into a tiered support chain. L1 triaged the initial signal, escalated to L2 for deeper investigation, and L2 passed on to infrastructure or engineering when needed. It worked. It was also slow. Each layer added minutes to resolution while the underlying incident kept costing money.

SignalL1 triageL2 investigationInfra / Engineering

Each layer added minutes while the incident kept costing money.

Before Sherlocks AI, Fynd's investigation flow leaned on Claude as a co-pilot. Human agents would spin up Claude and manually feed it context. Some of that context was organisation-level and available to any support agent. Most of it had to be extracted on the fly, and that surfaced two structural problems.

First, individual agents often did not have access to the observability data or telemetry they needed. They had to pull it from systems they were not authorised for, or route through people who were.

Second, Claude was never autonomous. Someone always had to steer it. It made root cause faster, but a human was still doing the work.

Their Grafana stack was strong at collecting signals but did not close the gap between something is wrong and here is why. That gap was where MTTR was quietly bleeding.

The evaluation

Fynd compared Sherlocks AI against their existing workflow: human triage assisted by Claude. Two criteria mattered most.

Accuracy

Could the tool actually identify the correct root cause, not just a plausible one?

Speed

Could it do so faster than a human agent driving Claude in real time?

Fynd's verdict: Sherlocks AI won on both.

The solution

Fynd deployed Sherlocks AI as a SaaS integration via the Watson agent, connecting it directly to their existing Grafana observability stack on GCP. No re-instrumentation. No changes to the telemetry pipeline.

Once connected, Sherlocks AI runs autonomously across Fynd's stack, correlating signals across services, deployments, and infrastructure, and surfacing root cause and remediation directly to the team.

The most important shift wasn't in what Sherlocks AI does. It was in when.

Because Sherlocks AI is always-on, every alert gets investigated the moment it fires, not after a human agent picks up the ticket, spins up a tool, and starts gathering context. By the time an engineer looks at the incident, the investigation is already running or done.

The impact

70%reduction in MTTR
63%support team hours saved
77%infrastructure and engineering hours saved
False alarms handled by Sherlocks AI rather than eating the support team's bandwidth

Beyond the numbers, the biggest qualitative shift was in what Sherlocks AI replaced. Instead of dashboards to interpret and tools to spin up, engineers get a resolution directly. Fynd reports the team is no longer drowning in false alarms, and that infrastructure work is far less often interrupted by triage tasks Sherlocks AI can handle first.

In their words

Sherlocks AI has fundamentally changed how we respond to incidents. Almost every problem on our platform eventually becomes an engineering problem, and before Sherlocks AI, most of our resolution time was spent extracting context and manually steering Claude through the investigation. Now the investigation is already running by the time we look at the ticket. Sherlocks AI takes us straight to the root cause instead of handing us more dashboards, and we've been able to retire a few point tools along the way.
Fynd logoAmboj GoyalHead of Infrastructure, Fynd

What's next

Fynd is expanding Sherlocks AI usage beyond the main e-commerce platform. Since Fynd operates a suite of SaaS products, the team is now deploying Sherlocks AI across additional offerings in the portfolio.

See Sherlocks AI in your stack