How Fynd cut incident resolution time by 70% with autonomous AI SRE
Reliance Retail's e-commerce platform replaced human-driven triage via AI tools with always-on Sherlocks AI agents, freeing up support and engineering teams while cutting time-to-resolution across the stack.
Most of our resolution time was spent extracting context and manually steering Claude through the investigation. Now the investigation is already running by the time we look at the ticket.
Figures as reported by Fynd.
“Now the investigation is already running by the time we look at the ticket. Sherlocks AI takes us straight to the root cause instead of handing us more dashboards, and we've been able to retire a few point tools along the way.”
About Fynd
Fynd is an AI-native unified commerce platform built to help brands run their businesses more efficiently. With AI-first solutions across online commerce, in-store technology and supply chain operations, Fynd brings entire business functions together into one connected system. Designed for scale, the platform helps teams launch faster, operate smarter and deliver consistent customer experiences across every channel.
Headquartered in India and backed by Reliance, Fynd is expanding across the GCC, UK, Africa and Southeast Asia to support the next phase of modern commerce.
The engineering organisation is 800+ strong, runs primarily on GCP, and operates its own observability stack built on Grafana/Prometheus and Tempo.
Because Fynd's product powers commerce at scale, every failing transaction represents high-value business impact. Reliability isn't a nice-to-have. It maps directly to revenue.
The challenge: strong signals, slow answers
Fynd's incident response ran on two intake paths.
- 1Tickets came in directly from B2B customers whose end-users could not complete actions on their platforms.
- 2Alerts came from Fynd's own observability stack, either as tickets or as PagerDuty calls.
Both paths funneled into a tiered support chain. L1 triaged the initial signal, escalated to L2 for deeper investigation, and L2 passed on to infrastructure or engineering when needed. It worked. It was also slow. Each layer added minutes to resolution while the underlying incident kept costing money.
Each layer added minutes while the incident kept costing money.
Before Sherlocks AI, Fynd's investigation flow leaned on Claude as a co-pilot. Human agents would spin up Claude and manually feed it context. Some of that context was organisation-level and available to any support agent. Most of it had to be extracted on the fly, and that surfaced two structural problems.
First, individual agents often did not have access to the observability data or telemetry they needed. They had to pull it from systems they were not authorised for, or route through people who were.
Second, Claude was never autonomous. Someone always had to steer it. It made root cause faster, but a human was still doing the work.
Their Grafana stack was strong at collecting signals but did not close the gap between something is wrong and here is why. That gap was where MTTR was quietly bleeding.
The evaluation
Fynd compared Sherlocks AI against their existing workflow: human triage assisted by Claude. Two criteria mattered most.
Could the tool actually identify the correct root cause, not just a plausible one?
Could it do so faster than a human agent driving Claude in real time?
Fynd's verdict: Sherlocks AI won on both.
The solution
Fynd deployed Sherlocks AI as a SaaS integration via the Watson agent, connecting it directly to their existing Grafana observability stack on GCP. No re-instrumentation. No changes to the telemetry pipeline.
Once connected, Sherlocks AI runs autonomously across Fynd's stack, correlating signals across services, deployments, and infrastructure, and surfacing root cause and remediation directly to the team.
The most important shift wasn't in what Sherlocks AI does. It was in when.
Because Sherlocks AI is always-on, every alert gets investigated the moment it fires, not after a human agent picks up the ticket, spins up a tool, and starts gathering context. By the time an engineer looks at the incident, the investigation is already running or done.
The impact
Beyond the numbers, the biggest qualitative shift was in what Sherlocks AI replaced. Instead of dashboards to interpret and tools to spin up, engineers get a resolution directly. Fynd reports the team is no longer drowning in false alarms, and that infrastructure work is far less often interrupted by triage tasks Sherlocks AI can handle first.
In their words
Sherlocks AI has fundamentally changed how we respond to incidents. Almost every problem on our platform eventually becomes an engineering problem, and before Sherlocks AI, most of our resolution time was spent extracting context and manually steering Claude through the investigation. Now the investigation is already running by the time we look at the ticket. Sherlocks AI takes us straight to the root cause instead of handing us more dashboards, and we've been able to retire a few point tools along the way.
What's next
Fynd is expanding Sherlocks AI usage beyond the main e-commerce platform. Since Fynd operates a suite of SaaS products, the team is now deploying Sherlocks AI across additional offerings in the portfolio.