SDP Clouds
DevOps·4 min read

On-Call Without the Burnout: Designing a Rota People Survive

Page volume, escalation, and handoffs decide whether on-call is a job or a lifestyle — how to build a rotation the team will still volunteer for in a year.


Most on-call rotas fail for boring reasons. Not a lack of tooling — a schedule that pages one person nine times a night, an escalation path that resolves to "wake the person who wrote it," and a handoff that is a Slack message nobody reads.

On-call is a design problem. You can make it survivable with a small number of decisions, and the hard part is committing to them rather than optimising something else first.

Page volume is the only number that matters

You can argue about fairness, compensation, and tooling all day. If the average shift produces more than a couple of actionable pages, nothing else you do will make it pleasant.

Measure it honestly for one rotation: pages per shift, and of those, how many were actionable. Alerting on symptoms users feel is different from alerting on conditions you merely noticed. A CPU graph crossing a threshold is not a page; a checkout failing is.

Two habits move the number most:

  • Alert on SLO burn, not on metrics. A single rate-of-error window with a sensible threshold beats five rules on five graphs.
  • Delete rules nobody has acted on in ninety days. Most alert sets accumulate like unread email. Every alert that fires and gets silenced trains the team to ignore the next one.

Escalate twice, then wake someone

The worst rota is the one where a junior engineer hits an alert at 03:00, doesn't recognise it, and sits with it for forty minutes before deciding to ask. That's not a training problem — it's a missing path.

Write the escalation down, in the alert itself:

  1. Is it actionable now? (yes → step 2, no → acknowledge and sleep)
  2. Do you recognise it within five minutes? (yes → work it, no → page secondary)
  3. Secondary doesn't know either? (page the author or the team lead — by policy, not by judgement)

The point of step 2 is that it gives permission to stop. People who feel they must personally resolve everything will quietly grind through the night instead of escalating, which is how you get a hero and then a resignation.

Handoffs are a meeting, not a message

Changeover happens twice a week at 09:00, it takes ten minutes, and it has an agenda: what broke, what's half-finished, what's noisy right now, and which alerts are known-bad.

This is the highest-value ten minutes in the whole rotation and it is almost always the first thing cut. Without it, context lives in one person's head and the incoming engineer inherits an undiagnosed incident with no history.

Written form is fine if the rota is follow-the-sun. What doesn't work is a channel where the outgoing person types "fyi" and leaves.

Make the quiet hours quiet

Two structural choices matter more than any dashboard:

  • Split the week so nobody does back-to-back nights. Consecutive broken sleep is where the resentment comes from.
  • Give a comp day after a rough shift. Not as a perk — as a load balancer. If a night of pages is paid back in a day off, the rotation stays voluntary.

And if volume genuinely can't come down, split the rota by system so a small team isn't on the hook for everything. Two narrow on-calls beat one broad one.

The rota is a training programme

A rotation where the same senior engineer is always called teaches the rest of the team nothing, and it means the system is one holiday away from an outage.

Rotate deliberately. Pair the new person with the experienced one for the first two shifts. Require that every alert has a runbook link — see the restore-drill discipline for why a document nobody has ever executed is just a document. When the person on call can fix it without asking, you have built a team rather than a schedule.

Track it: time-to-acknowledge, time-to-resolve, pages per shift, and how many shifts per person per month. If those four are stable, the rota is working. If they aren't, you know where to look.

Summary

Get page volume under control, write an escalation path that gives permission to stop, hold a real handoff, and protect sleep with structure rather than goodwill. On-call that people survive is not a heroic culture — it's a rotation with few pages, clear escalation, and enough slack to be human the next morning.

#devops#on-call#sre#reliability#culture

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles