Most on-call rotations fail for the same three reasons: too few engineers on the schedule, too many pages per shift, and no plan for recovery between shifts.
This guide covers the five rotation models that actually work in production, the math for how many engineers you really need, a 90-day playbook for fixing a broken rotation, and the honest cost of doing nothing. For the broader context on why rotations break and how to fix alert quality, see The On-Call Playbook for 2026.
8–9
Engineers for sustainable 24/7
single-site minimum
25%
Max on-call load per engineer
the sustainable ceiling
2–3
Actionable pages per shift
more means alert fatigue
40–60%
Alerts that are non-actionable
cut these first
Why Most On-Call Rotations Fail
On-call itself is not the problem. Most engineers signed up for it when they took the job. The problem is what happens when the rotation is designed poorly.
According to 2024 software engineering survey data, 65% of engineers reported experiencing burnout in the past year, and on-call stress is one of the leading contributors. Google's SRE book puts a specific number on the sustainable ceiling: no more than 30 to 40% of an engineer's bandwidth should be spent on incident work during their on-call period. Beyond that, quality drops, mistakes rise, and engineers start planning their exit.
The patterns behind broken rotations are almost always the same. Too few engineers carrying too many shifts. Alerts that are noisy without being actionable. No time to recover between shifts. No secondary to call when the primary is drowning. Managers who set the schedule but never take the pager themselves.
None of these problems are technical. All of them are design decisions.
The Math: How Many Engineers Do You Actually Need?
The math is unforgiving. If you want 24/7 single-site coverage without burning out your engineers, you need a minimum of eight to nine people on the rotation. This is not a soft recommendation. It is what falls out of the arithmetic once you cap on-call load at 25% of an engineer's time.
Here's the working number. A week of on-call is 168 hours. If you want each engineer to be on-call no more than 25% of the time, each engineer can take 42 hours per week. To cover the full 168 hours with a 25% cap, you need four engineers. Add a secondary layer, and you need eight. Add a buffer for vacations, sick days, and someone quitting, and you land at nine.
Below that number, rotations look sustainable on paper but drain the team in practice. Engineers end up on-call every three weeks instead of every six, cover for each other constantly, and burn out inside 12 months.
Here's the quick reference for how team size shapes what's realistic:
| Team Size | Realistic Rotation Model |
|---|---|
| 3-5 engineers | Primary plus backup, weekly swaps, business-hours coverage only |
| 6-8 engineers | Weekly primary with a secondary layer, still stretched |
| 9-15 engineers | Weekly primary plus secondary, sustainable single-site |
| 15-25 engineers | Follow-the-sun across two regions, no overnight shifts |
| 25+ engineers | Follow-the-sun across three regions, split shifts optional |
If your team is below the sustainable threshold, the answer is not a better schedule. It is more engineers, more automation, or a smaller scope of what you actually cover 24/7.
The Five Rotation Models: When to Use Each
Almost every functional on-call program uses one of five models. Picking the right one for your team size and geography matters more than any tool you buy.
One engineer holds the pager for seven days. Simple, predictable, best for teams of five to fifteen in a single time zone with moderate alert volume. It breaks the moment page volume gets heavy: five or more nightly interruptions on a weekly rotation is a burnout path, not a schedule.
The primary handles pages. The secondary backs up and takes over if an incident runs long or the primary misses the page. This is Google's recommended baseline for 24/7 coverage and works well from eight engineers up.
Each regional team covers their daylight hours and hands off to the next time zone. Nobody gets paged at 3 AM. Requires at least nine engineers spread across two or three regions and handoff discipline that most teams underestimate. Widely used at global companies. Reference: Atlassian's on-call guide.
Engineers cover fixed blocks of the day, usually 8 or 12 hours. Spreads night pages more evenly but adds handoff overhead. Best for larger teams running high-volume services where the pager fires every night.
A globally distributed team runs weekly primary within each region, with follow-the-sun handoffs between regions. This is the pattern most large tech companies converge on. Requires strong tooling and handoff culture, but eliminates overnight pages for individuals.
The rule of thumb: match rotation length to alert volume. Low volume, weekly works. High volume, shorter shifts. Global team, follow-the-sun.
The 90-Day Fix: From Broken to Healthy
Most guides tell you what a good rotation looks like. Almost none of them tell you how to get there when you have inherited a broken one. Here is a practical week-by-week plan for turning a bad rotation around without breaking the team in the process.
The 90-day turnaround at a glance
Weeks 1-2
Audit & stabilize
Weeks 3-6
Alert hygiene sprint
Weeks 7-10
Rotation redesign
Weeks 11-13
Tools & automation
Cut the load first, redesign the schedule second, add tools last. Broken to healthy in one quarter.
Start by measuring what you actually have. Pages per engineer per week. Alert-to-action ratio. Nightly page volume. How many engineers are effectively carrying the load versus how many are on the schedule. Ship no changes yet. You need a baseline before you can fix anything.
The single biggest lever in the first month is reducing alert noise. Most teams find that 40 to 60% of their alerts are non-actionable within one careful pass. Kill the alerts nobody acts on. Consolidate the noisy ones. Tighten thresholds on real signal. Do not touch the rotation yet. Cut the load first.
Now redesign the schedule. Use the team size math above to pick a realistic model. Add a secondary layer if you do not have one. If you are below the sustainable threshold, this is the moment to raise the headcount question with leadership. Bring the data from weeks 1 and 3 to make the case.
Only at the end do you invest in tooling. Better on-call management, better runbooks, automation for the routine restarts and known-issue responses. Tools amplify a good rotation. They do not fix a broken one.
Ship the changes gradually. Announce the new schedule with two weeks of notice. Review after 30 days. Adjust.
Recovery and Handoffs
The two most-skipped practices in on-call design are recovery time and structured handoffs. Both cost almost nothing to implement and both make an outsized difference to how sustainable the rotation feels.
Recovery time means giving engineers a day of reduced workload after a difficult shift. Multiple overnight pages, a long incident, a weekend outage. The team should know that after those shifts, the engineer gets time to catch up before diving back into feature work. It is not a reward. It is basic sustainability.
Structured handoffs are how one engineer transfers context to the next. A good handoff transfers operational context, not just a list of alerts. Five things, nothing more:
- •Active incidents and current mitigation state
- •Recent deployments in the last 24 hours
- •Anything unusual observed but not yet investigated
- •Dependencies that were flaky during the shift
- •Anything the incoming engineer specifically needs to watch for
The handoff should take five to ten minutes. Written. In a fixed template. Every time.
The 5 Most Common Mistakes
Ranked by how often they show up when a rotation is failing.
Too few engineers, faking sustainability. Running an eight-person rotation with six people. Everyone knows it is unsustainable. Nobody wants to raise it because the answer is unpopular. The rotation collapses when the first person quits.
No recovery time after difficult shifts. The engineer who spent Saturday night on an incident is expected to be at standup on Monday morning shipping features. Multiply this by a few months and you lose them.
Alert thresholds set too low. Pages that fire for conditions nobody acts on. Engineers start ignoring the pager. When a real issue hits, it gets missed in the noise. The Google SRE Workbook recommends targeting no more than two to three actionable pages per shift.
No secondary, no escalation. Only a primary. If the primary is asleep, in a meeting, or overwhelmed, there is no fallback. Single points of failure in on-call design cause the same problems as single points of failure in infrastructure.
Managers not participating in the rotation. When managers do not take the pager themselves, they underestimate how bad the rotation is. Engineers stop reporting the problems because they know it will not land.
The Real Cost of a Broken Rotation
Fixing a broken rotation is expensive. Not fixing one is more expensive. The math is worth doing out loud.
Replacing a senior SRE typically costs between $50,000 and $150,000 in the US, depending on seniority, region, and how long the role stays open. This includes recruiter fees, interview time, ramp-up time, and productivity lost during the vacancy. Even in lower-cost regions like India, the fully-loaded cost of losing and replacing an SRE lands in the $20,000 to $60,000 range.
Now apply attrition rates. Teams with broken on-call rotations lose engineers at 30 to 40% per year. A team of ten with a broken rotation loses three to four engineers annually. At $80,000 per replacement, that is $240,000 to $320,000 per year in retention cost alone. Not counting the incident quality that drops when experienced engineers leave, or the training burden on the ones who stay.
Compare that to what it costs to fix a rotation:
| Investment | Typical Cost |
|---|---|
| Hire 2 additional engineers to reach sustainable team size | $200,000 to $400,000 per year |
| Invest in alert hygiene and automation | $20,000 to $50,000 one-time |
| Roll out on-call management tooling | $10,000 to $30,000 per year |
| Total first-year cost of fixing | ~$230,000 to $480,000 |
| Annual retention cost of not fixing | ~$240,000 to $320,000 |
The math almost always works in favor of fixing the rotation. And that is before counting the second-order costs: slower incident response, worse postmortems, lower morale across engineering, and the reputational hit when your best engineers leave for competitors.
Every quarter you do not fix the rotation, you pay for it in attrition. The fix pays back within twelve months.
Metrics to Track (Health Dashboard)
A functional on-call program measures itself. Five metrics that predict burnout before your best engineers hand in notice:
- •Pages per engineer per week. Healthy: under 5. Danger: over 10.
- •Alert-to-action ratio. Healthy: 30 to 50%. Below 20% and engineers stop trusting the alerts.
- •Recovery-time compliance. Healthy: 100%. Anything less and you are burning your buffer.
- •Repeat incident rate. Healthy: zero within 30 days. Repeat incidents mean postmortems are not producing durable fixes.
- •Engineer satisfaction with on-call. Track quarterly. This is the earliest leading indicator of attrition.
Track these on a dashboard the whole team can see. Review monthly. When numbers slip, act before they become resignations.
Key Takeaways
- •Rotations fail from design, not from concept. Most engineers accept on-call. What they reject is a broken rotation.
- •The math is unforgiving. Below 8 to 9 engineers, single-site 24/7 coverage is not sustainable without automation.
- •Alert hygiene comes before rotation redesign. Cut noise first. Redesign the schedule second.
- •Recovery time is not a reward. It is a baseline sustainability practice.
- •Structured handoffs matter. Five items, written, every time.
- •The fix pays for itself in retention. Broken rotations cost more than fixing them, every year.
- •Track the leading indicators. Pages per week, alert-to-action ratio, and satisfaction predict attrition before it happens.
Frequently Asked Questions
Minimum eight to nine for single-site coverage with a 25% on-call cap per engineer, following Google's SRE recommendations.
Weekly for most teams, daily or shorter blocks only if alert volume is high and handoff discipline is strong.
A model where regional teams cover their local daylight hours and hand off to the next time zone, eliminating overnight pages for individuals.
Either a flat on-call stipend per shift or time-in-lieu after difficult shifts; industry norm is one of the two.
No more than 25% of an engineer's time on-call and no more than two to three actionable pages per shift.
Cap load, fix alert noise, add secondaries, enforce recovery time, and track engineer satisfaction quarterly.
Yes, at least occasionally, so they experience the rotation firsthand and stay honest about what needs fixing.
A two-layer model where a primary engineer takes the page and a secondary backs up if the primary is unavailable or overwhelmed.
Declining responsiveness during shifts, dread before the handoff, engineers pushing back on schedule changes, and rising attrition.
Three engineers is the absolute floor for business-hours coverage; eight to nine is the floor for sustainable 24/7.
Further Reading
The On-Call Playbook for 2026
Why rotations break and how to fix alert quality.
What Is SRE Toil? How to Measure and Reduce It
Measure and reduce the manual work that burns engineers out.
Why Incident Debugging Is Still Slow in 2026
Where incident response time actually goes.
Root Cause Analysis for Production Incidents
The methods that still work in production, and how AI changes them.
Never Miss What's Breaking in Prod
Breaking Prod is a weekly newsletter for SRE and DevOps engineers.
Subscribe on LinkedIn →