Disaster Recovery for Teams That Don't Have a DR Team
RPO and RTO in plain language, the four tiers that actually exist, and the one practice — the restore drill — that separates a plan from a hope.
Most small teams have a backup and no disaster recovery. The distinction matters: backups are a copy of data, DR is the ability to be running again. Having the first and being unable to demonstrate the second is how a two-hour incident becomes a two-week one.
Two numbers, decided in advance
Everything in DR reduces to two numbers you should write down before anything breaks:
RPO — Recovery Point Objective. How much data can you lose, measured in time? If your last successful backup was at 04:00 and you fail at 11:00, your RPO is seven hours. Anyone who finds that unacceptable has just declared that backups need to run more often.
RTO — Recovery Time Objective. How long until you are serving users again? Not until data is restored — until the application answers requests correctly.
Those two numbers are a budget, and the budget is set by the business, not by engineering. "How much downtime can we afford per year?" is a question about revenue and reputation. Answer it once, explicitly, and every technical decision after it becomes a comparison rather than a debate.
The four tiers, honestly described
There are four common patterns. Pick by loss tolerance, not by ambition.
| Tier | What it looks like | Realistic RTO |
|---|---|---|
| Backup & restore | Copies exist; you rebuild from them | Hours to days |
| Pilot light | Minimal core running; rest scaled up on demand | Tens of minutes to hours |
| Warm standby | Reduced-capacity clone, always running | Minutes |
| Multi-region active | Traffic shifted between live regions | Seconds to minutes |
The important part: tiers are priced by how much idle infrastructure you are willing to pay for. An active-active setup costs roughly a second full environment sitting there, permanently, in case something happens. For most teams that is the wrong trade, and saying so is a legitimate architectural decision rather than an admission of weakness.
Backup and restore is also a real tier, not a failure. If your RTO is "this evening," a solid backup with a tested restore path is the correct, proportionate answer. What is not acceptable is having chosen it by default rather than by decision.
The part everyone skips
A plan that has never been executed is a document. The restore drill is the only practice that converts one into the other, and it is cheap:
- Pick a resource. Start with something non-critical.
- Restore it into a separate account or VPC — not in place.
- Time how long it takes, including the steps you forgot to write down.
- Run the application against it and confirm it works.
- Write down what failed.
You will find problems in the first drill, every time. Credentials that only exist in one person's password manager. A restore that succeeds technically but produces a database missing its latest migrations. A runbook that says "provision the replacement" without saying from where.
Do it quarterly. Put the measured RTO in the document next to the target RTO — the gap between those two numbers is your actual risk, and it is far more useful than any compliance checkbox.
What I would set for a typical SaaS
For a product with a small team and a real but not catastrophic cost of downtime:
- Databases: automated backups with point-in-time recovery, retention long enough to catch a bad migration being discovered the next morning. Often 7–14 days.
- Everything reconstructible: infrastructure as code means compute, networking, and configuration are not backed up — they are redeployed. This is the single biggest advantage a modern stack has over a 2010 one, and it depends on keeping that code trustworthy.
- State you cannot rebuild: object storage with versioning and lifecycle rules, replicated at minimum across availability zones.
- Secrets and configuration: documented, not stored only in someone's head.
- One drill per quarter, on the least critical system first, escalating as confidence grows.
Write the runbook while it is boring
Incidents are the worst possible time to reconstruct a procedure. Capture the sequence in plain language: who declares the incident, where the runbook lives, the restore order, and how traffic is pointed at the new environment.
Keep it short enough that someone under pressure will actually read it. A twelve-page document nobody opens during an outage is decoration.
Summary
Decide RPO and RTO as a business question, pick a tier that fits the budget rather than the org chart, and treat the restore drill as the real deliverable. Backups without rehearsed recovery are storage spending; a tested, timed, documented restore path is disaster recovery — and it costs a few hours a quarter.
SDP Clouds Team
DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.
More about us →