Backup and disaster recovery get used interchangeably, and the confusion costs real money in both directions: businesses that buy expensive DR they don't need, and businesses that think nightly backups mean they'll be running again by lunch. The distinction is one sentence. A backup is a copy of your data. Disaster recovery is the plan and machinery for operating again.
Why the copy isn't the recovery
Your server dies. The backup is fine, sitting safely in the cloud. Now: restore onto what? The dead server needs parts or a replacement, which is days. The restore itself might be hours for terabytes pulled over your internet line. Reinstalling and reconfiguring the applications on a fresh box adds more. A perfectly good backup can still mean a week of downtime, because backup answers "is the data safe" and recovery answers "when are we working again."
The two numbers that define recovery
- RTO, recovery time objective: how long until you're operating again. "We can be down four hours" is an RTO.
- RPO, recovery point objective: how much recent work you can afford to lose. Nightly backups mean an RPO of up to 24 hours; whatever happened since last night's backup is gone.
Set both per system, not for the business as a whole. The invoicing system might justify an RTO of hours and an RPO of minutes. The archive of 2019 project files can have an RTO of "next week" and nobody suffers. Your priority-ordered systems list is where these numbers live, and your downtime cost is how you justify them.
The three recovery tiers
- Rebuild from backup. Cheapest. When disaster hits, you buy or provision hardware and restore. RTO: days. Right for businesses that can limp along on workarounds for a while.
- Standby recovery. A recovery environment exists ahead of time: a spare box, or more commonly now, backup software that can boot your server as a cloud VM within hours. Roughly $200 to $500 a month at small-business scale. RTO: hours. This tier is the sweet spot for most companies with a server they depend on.
- Live failover. A duplicate environment continuously replicated, taking over in minutes. Serious money and complexity. Right when downtime is measured in thousands per minute, which is a real number for some businesses and a vendor fantasy for most.
Cloud-first businesses get a quiet advantage here: with no server to resurrect, "recovery" collapses into data restore plus laptops plus Wi-Fi, one of several honest arguments in the cloud decision. More on the cloud version of this in cloud backups and disaster recovery.
The part everyone skips
A recovery plan you haven't executed is a hypothesis. The standby VM that's never been booted, the restore that's never been timed against the RTO you promised yourself, both have a way of underperforming on the day it counts. Timing a full recovery once a year is the difference between a plan and a hope, and it's the centerpiece of continuity testing.
Want this handled instead of homeworked? That's the job.
Email us →