SDP Clouds
Cloud·4 min read

Multi-Region Without the Fantasy: What Actually Fails Over

Active-active sounds like the answer until you price the second region and meet your first cross-region write. The three tiers worth considering, and the bill before the demo.


Every architecture diagram I have reviewed in the last five years had two regions and a cheerful arrow between them. Almost none of the systems behind those diagrams could have used one.

Multi-region is the most expensive thing you can casually agree to in a design review, and the cost is not primarily money. It is the cost of every subsequent decision getting harder: where does a write go, what happens when the two halves disagree, and who is paged at 2 AM when DNS itself is the thing that broke.

What are you actually protecting against?

Region-wide outage is the reason people give. It is also nearly the rarest one. The failures you will actually meet are:

  • A bad deploy — region-independent, and multi-region doubles your blast radius.
  • A dependency outage — your database vendor's global control plane. Two regions do not help.
  • A capacity shortfall — you cannot buy GPUs in us-east-1 today; eu-west-1 has the same problem on the same day.
  • A real region event — genuinely rare, genuinely catastrophic.

Only the last one is what multi-region buys, and it is worth knowing which of your requirements that is before you spend a year on it. If the actual requirement is "we cannot tolerate a bad deploy," you want rollback and restore drills, not a second region.

Three ways to do it

Active-passive. Everything runs in one region; a second sits warm with a copy of the data. Recovery is a DNS flip plus a promotion step, measured in tens of minutes. This is what most teams actually need, and it costs roughly 1.4× rather than 2×.

Active-standby with async replication. Same shape, but the standby is continuously caught up so the promotion does not lose the last few minutes. The interesting part stops being infrastructure and starts being: how much data are you willing to lose?

Active-active. Both regions serve writes. This is where diagrams get confident and engineers get quiet.

The write problem

Async replication means two copies can disagree, and a conflict has to resolve somewhere. Your choices are last-write-wins (silently loses data), a conflict-resolution function (now you have written a distributed system), or routing each entity to exactly one home region (which quietly turns active-active into active-parallel).

Almost every "we went active-active" story I have read resolved it by moving writes back to one region. That is not a failure of the idea — it is the idea, once it meets real data.

Nobody tells you this part because the benchmark did not have a customer record that two offices edited on the same afternoon.

The bill, before the demo

Double the compute is the obvious line item and usually the smaller one. The ones that surprise people:

  • Cross-region data transfer, charged per GB in both directions, on traffic you never see on a dashboard.
  • A second set of the boring things — NAT gateways, load balancers, backup storage, log retention. Each is cheap alone.
  • Twice the operational surface. Two of every runbook, every alert, every on-call skill. Cost optimisation gets harder when you cannot turn half of it off at night.

If you cannot articulate the region-failure requirement in one sentence with a number in it — RTO in minutes, RPO in seconds — you are buying insurance against a scenario you have not described.

When single-region is the right answer

When your data has a natural home, when your users are concentrated, when a thirty-minute recovery is survivable, or when the real threat is bad code rather than bad weather. Say so in the design doc. "We accepted single-region risk because X costs Y and our RTO target is Z" is a defensible engineering position; "we'll add the second region later" is how you end up with the arrow on the diagram and none of the behaviour.

Summary

Multi-region buys exactly one thing: survival of a region-level event. Price that specifically — transfer, duplicated fixed costs, doubled operations — before agreeing to it. Most teams want active-passive with a stated RTO, not active-active with an unstated conflict-resolution strategy. Write the requirement down as a number first; if you cannot, the answer is that you do not need the second region yet.

#multi-region#disaster-recovery#aws#architecture#cloud

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles