Clouds
Observability·4 min read

Error Budgets: The Contract That Ends the Uptime Argument

Ninety-nine point nine is a number nobody chose for a reason. An error budget turns availability from a preference into a shared spending limit.


"Can we just ship it? It's only a small change." I have had that conversation while a service was nominally at 99.97% and while it was at 99.4%, and the argument was identical both times. The difference was not the number — nobody in either room had picked it, and nobody could say what it entitled us to.

An error budget exists to replace that argument with arithmetic.

Ninety-nine point nine is a decision, not a fact

Take 99.9% monthly availability and you have bought yourself 43 minutes of downtime. 99.99% is four. 99.5% is three and a half hours.

Those numbers sound like properties of the system. They are not — they are the answer to "how much unreliability are we willing to trade for how much speed?", and that question is a product decision that engineering is usually left to guess at. When the target is guessed, every release conversation becomes a negotiation between two people's instincts.

The arithmetic, once

An error budget is simply the complement of the SLO over a window: if the objective is 99.9% of requests succeeding in 30 days, the budget is the 0.1% that may fail.

text
objective   99.9% successful requests / 30 days
budget      0.1%  =  43 minutes of full outage
            or     =  43,000 failed requests out of 43 million

The second line matters more than the first. Most services do not go down; they degrade. A checkout endpoint failing for 2% of requests over a week burns budget without a single page firing, which is why budget burn rate is the number to alert on rather than raw availability.

Two directions to spend it

The budget buys you two things, and they belong to different people.

Reliability work. If the budget is gone, the reliability backlog gets funded: the flaky dependency, the missing retry, the single point of failure everyone has been naming in design reviews.

Release velocity. If the budget is healthy and nearly untouched, the team should ship faster — that is what it was for. A budget that sits at 100% every month is usually a sign the SLO is set too loose to constrain anything, not a sign of excellence.

The rule that makes it a contract rather than a dashboard is symmetric: no budget means no non-essential releases, full budget means releasing should be the easy answer. Both halves have to be stated, or only the restrictive one gets applied and the budget becomes another reason to say no.

The meeting it replaces

Before: a disagreement about whether a change is risky, resolved by whoever is more senior or more tired.

After: a look at burn rate over the last seven days and a shared number. "We're at 40% of a 30-day budget in 12 days" ends the discussion without anyone needing to defend a position.

This is the actual product of the exercise — not the metric, but the removal of a recurring negotiation. It works precisely because the threshold was agreed when nobody was afraid of a specific release.

Where teams go wrong

Setting the objective before choosing the indicator (you cannot have an SLO for a number you do not collect), picking 99.99% because it sounds serious, and then having no instrumentation precise enough to observe the difference. Or measuring availability as uptime of the process, which stays green while every user request fails.

What it needs to be real

None of this works unless the indicator itself is trustworthy. It has to be collected continuously, at a granularity where a five-minute degradation is visible, and it has to describe what users experienced rather than whether a process stayed up — the same discipline as choosing metrics that answer a question instead of metrics that merely look healthy.

And a budget breach is a fact worth writing down. When you later ask why the quarter was quiet or why that release slipped, the burn record is the honest answer, and it belongs beside the postmortem rather than in a dashboard nobody reopens. Budget that went unspent in a quiet month is evidence the objective was set too loose to constrain anyone — worth raising, not celebrating.

Summary

An error budget converts availability from an aspiration into a spending limit that engineering and product share. Compute it from an objective you chose deliberately, measure it against user-visible success rather than process uptime, and agree in advance what happens when it runs out in both directions. The metric is easy; the value comes from the argument it makes unnecessary.

#slo#error-budgets#reliability#observability#sre

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles