Skip to content

Backups vs. DR: Why You Need Both

I’ve sat in more than one post-incident review where the room genuinely believed they were covered, right up until the moment they weren’t. Backups were running. Reports were green. And the business was still down for two days, or in a few cases weeks, because nobody had actually thought through what happens after the data comes back, or even what needs to come back with it, or beforehand.

That’s the gap. Backups and disaster recovery get talked about like they’re the same checkbox. They’re not. One gets your data back. The other gets your business back. You need both, and conflating them is one of the more expensive mistakes I see well-meaning IT teams make.

One more distinction worth naming, even though it’s not the focus here: business continuity planning (BCP) is a third, separate discipline from both of these. Where DR is about getting systems back, BCP is about keeping the business operating while that happens: alternate work locations, manual workarounds, which functions absolutely cannot stop even for an hour. It deserves its own full treatment rather than a paragraph tacked onto this one, but if your plan only covers backup and DR, you’re still missing a piece. I cover that gap in Business Continuity Planning: The Third Thing Backup and DR Don’t Cover.

What a backup actually promises you

A backup is a copy of your data, frozen at a point in time, sitting somewhere separate from the original. That’s it. That’s the whole promise.

It answers one question: if this file, this database, this server gets corrupted, deleted, or encrypted by ransomware, can I get it back?

Not all backups are built the same way, and it’s worth knowing the difference when you’re deciding what fits where:

  • Full: a complete snapshot of the entire system. Slowest to run, fastest to restore from.
  • Incremental / differential: only the changes since the last backup. Faster and lighter, but restoring means stitching several of these together in the right order.
  • Snapshot-based: common in virtualized environments, capturing the state of a VM or volume at a point in time.

Most real environments use some mix of all three, and that’s fine. What matters is that whatever mix you’re running is automated, tested, and retained long enough to actually reach back to the point you’d need. If your backups aren’t tested, they’re not backups. They’re a hope.

The classic guidance here is the 3-2-1 rule: three copies of your data, on two different types of media, with one copy stored offsite: cloud, a remote data center, or an air-gapped location, meaning physically or logically disconnected from your production network so nothing on that network can reach it. That’s still solid advice. I’d add a minimum threshold to it in 2026: at least one of those copies needs to be immutable. In plain terms, immutable means write-once: once that backup lands, nothing (not an admin, not a script, not the ransomware that just compromised your admin credentials) can alter or delete it until a retention timer you set in advance expires. You’ll sometimes see this called WORM storage (Write Once, Read Many), which is the same idea under an older name. Regular backup storage can be encrypted or wiped right along with everything else if an attacker gets the right credentials. Immutable storage can’t, by design, no matter who’s logged in. That single property is what turns “we have backups” into “we have backups an attacker can’t touch,” and combined with air-gapping to limit how far a compromised network can reach in the first place, it’s stopped being optional. Encryption at rest and in transit belongs in that same non-negotiable category: if your backup data is sitting or moving unencrypted, immutability alone isn’t covering you. And that protection is only as strong as your key management: if your encryption keys sit in the same environment as the backups they protect, or the same identity system an attacker just compromised, encryption isn’t actually buying you anything. You wouldn’t leave the garage door opener for your house in an unlocked car in your driveway would you? (Please tell me you don’t do that.)

You’ll sometimes see this whole extended version written as 3-2-1-1-0: three copies, two media types, one offsite, one immutable, and zero errors on your restore tests. That last number is the one people skip: a backup nobody’s verified restoring cleanly is a theory, not a plan.

But notice what a backup does not promise: it says nothing about how long it takes to restore, what order things come back in, whether your applications work once the data’s there, or whether anyone downstream even knows what’s happening while you’re doing it.

What disaster recovery actually promises you

Disaster recovery is the plan for keeping the business running, or getting it running again, when something takes out more than a file. A ransomware attack that hits your whole environment. A data center outage. A regional cloud provider issue: the kind of infrastructure dependency I walked through in Cloud Computing Explained. Fire, flood, the contractor who cuts through the wrong conduit.

DR is process, not just data. It’s the answer to: if our primary environment is gone, how do people keep working, in what order do systems come back, and how long is that actually going to take? That process is usually built from a handful of concrete pieces: failover infrastructure to run on while primary systems are down (on-prem, cloud, or hybrid), replication to keep a standby environment current, orchestration (runbooks and automation scripts, not tribal knowledge in someone’s head), and regular testing through tabletop exercises or full simulations.

And “in what order” is where I want to plant a flag early, because it’s the part people skip past fastest. Your data being restorable doesn’t mean much if the thing that controls who’s allowed to log in and touch it isn’t restorable too. Entra ID is the identity layer everything else depends on: file shares, applications, email, the works. And if you’re running hybrid identity, on-prem Active Directory is a second, equally critical system, not just an old name for the same thing: it can fail independently of Entra ID, and your plan needs to cover both, not one standing in for the other. If you’ve backed up your database perfectly but identity isn’t part of a tested, recoverable plan of its own, you don’t have a recovery plan. That plan also needs at least one breakglass account per identity system: credentials stored outside the identity system itself, tested periodically, for the scenario where the thing you’d normally log in with is exactly what’s down. You have a pile of data nobody can authenticate their way into. I’ve seen teams get their servers back online faster than they got their identity layer sorted out, which means “restored” systems just sat there, unusable, while everyone scrambled to rebuild trust relationships and permissions from scratch. Identity recovery isn’t a footnote to the plan. It’s step one, and it’s worth extending your everyday Zero Trust posture (verify explicitly, least privilege, assume breach) into your DR environment too, rather than treating the recovery environment as a trusted-by-default exception. A failover environment with looser access controls than production is just a second attack surface waiting for someone to notice it.

This is where two numbers matter more than almost anything else in your environment, and if you don’t know them off the top of your head for your critical systems, that’s the actual gap this article is trying to point at:

  • RTO (Recovery Time Objective): how long can this system be down before it seriously hurts the business?
  • RPO (Recovery Point Objective): how much data can you afford to lose, measured in time? An hour of transactions? A day?

A backup with a two-day restore time is fine for an archive server nobody touches. It is not fine for the system that runs your order processing. The backup being “good” doesn’t mean the RTO is acceptable: those are two completely different conversations, and I’ve watched them get treated as one far too often.

Two questions will get you most of the way to real numbers for both:

  • What’s actually recreatable after data loss? If your online ordering system gets hit, is there a way to manually rebuild the orders that happened between the last backup and the moment things went sideways? If not, you need real-time or near-real-time replication for RPO: not the nightly or weekly backup that’s been the old standard since “A Flock of Seagulls” had a #1 hit.
  • What does downtime actually cost, per week, per day, per hour, maybe even per minute? Rare is the company I’ve come across in thirty years that could readily answer that question, but it’s what actually drives the RTO number for each system. I’ve had a CEO tell me, without hesitation, “we can’t be down for even one minute,” right up until he saw the price tag on a fully redundant, zero-loss, automated failover system with hot standbys, at which point it became “well, maybe a few days is fine if client-facing systems are back in a few hours.” Another client could document that they’d lose $1.1 million per minute of downtime given the nature of their business, so a $5 million DR solution was cheap insurance against the economic disaster an extended outage would actually cause them.

Knowing your business and its real costs is what right-sizes the solution to your organization’s actual needs. You don’t need this 100% figured out before you put the basics of a DR plan in place, but it’s what pushes you from “the basics” toward something your organization can genuinely rely on long term.

Two scenarios that show the actual difference

Nothing makes the distinction click faster than putting it next to a real situation.

Scenario 1: someone deletes a file. A user accidentally (or not so accidentally) deletes a critical spreadsheet. You pull it from backup, restored in minutes. This is exactly what backups are for. DR never enters the picture, because the business never stopped functioning. One file was briefly missing.

Scenario 2: ransomware hits the environment. Everything’s encrypted, including whatever backups weren’t isolated from the blast radius. This is where DR takes over: failing over to a clean environment, restoring from the immutable copies that survived, and, per the point above, revalidating your identity systems before you trust anything running on top of them. This isn’t a faster version of scenario 1. It’s an entirely different operation, with an entirely different team, timeline, and set of decisions.

If your plan only accounts for scenario 1, you have a backup strategy wearing a DR strategy’s name tag.

Common misconceptions worth killing off

A few beliefs I run into constantly, usually right before they turn into the post-incident review nobody wanted to be in:

“We have backups, so we’re covered.” Backups don’t restore operations: they restore data. Covered means someone can answer, with a straight face and a tested number, how long it takes to get the business functioning again. See everything above.

“DR is only for large enterprises.” I’d have believed that a decade ago. Not anymore. Smaller organizations are targeted just as often now, sometimes more, precisely because attackers assume the DR maturity isn’t there. Company size was never the qualifying factor: what’s actually running the business is.

“Cloud providers handle DR for us.” Partially true, and the “partially” is where people get burned. Take Microsoft 365 as the example, since it’s the one I see trip people up most: Microsoft’s job is to keep the platform available: uptime, infrastructure, the service itself. Your data inside that platform is your responsibility, full stop, and that’s Microsoft’s own published position, not a technicality I’m reading into it. M365’s built-in retention will get you back most accidentally deleted items, but usually only within a default window measured in weeks, not indefinitely, and the exact window shifts depending on the service and your licensing. Outside that window, or outside what retention policy actually covers, it’s gone. A user who leaves and has their license reassigned before someone moves their data first, a retention policy that’s more forgiving in Exchange than it is in SharePoint, an internal actor deleting things on the way out, a rogue app with too much access: none of those are Microsoft’s problem to solve, because none of that data loss happens to the platform. It happens to your copy of it.

And this isn’t a hypothetical gap. HYCU’s 2025 State of SaaS Resilience Report found that 70% of businesses using SaaS applications have already lost data from one of those apps, and 74% still don’t have offsite protection for that data. Most of the causes are exactly what you’d expect: malicious deletion, accidental deletion, misconfiguration, ransomware, roughly in that order of frequency. If “the cloud has it” is your entire backup strategy for M365 or any other SaaS platform, you’re not in the minority, but you are in the group finding this out the hard way. The same due-diligence instinct applies before you’ve even picked a provider, not just after: I get into what actually differs between them for a small business in Azure vs. AWS vs. Google Cloud for Small Business.

So what do you actually do about it

You don’t need a 40-page DR plan nobody will ever read to close most of this gap. Start smaller than that:

  1. List your critical systems and assign an actual RTO/RPO to each. Not “as fast as possible”: an actual number. This forces the prioritization conversation you’re otherwise going to have during an outage, which is the worst possible time to have it.
  2. Test one real restore this quarter. Not a file. A full system, into an isolated environment, timed. And make it a real test: most “DR tests” are scheduled, announced, and run against a system everyone already knows is coming down, which tells you almost nothing. The tests that actually teach you something involve killing a system nobody planned around and seeing what breaks. You’ll learn more from one real test than from a stack of green backup reports.
  3. Write down the dependency order. What has to come back before what: identity systems first. Five minutes with a whiteboard photo beats nothing, and it beats trying to remember it live during an incident.
  4. Confirm your backups can’t be encrypted along with production. Isolated, immutable, or otherwise protected from the same blast radius as your primary environment, with its own access controls and its own MFA, not just inherited from whatever identity system might be compromised in the same incident. If you can log into your backup storage with the same credentials that just got compromised, you don’t actually have a backup: you have a second copy of the ransomware’s target. And if that immutable copy lives off-site (which it should, that’s the “1” in 3-2-1), don’t stop at confirming it’s safe. Find out exactly how long it takes to pull it back. Hours? Days? Weeks? Does that number change if half the region is trying to restore at the same time you are? And separately: what does your off-site provider actually do for you in an emergency: is there a real human on a real phone line at 2am, or a support ticket queue with a business-hours SLA? A safe backup you can’t retrieve fast enough, from a provider that goes quiet when you need them most, is a plan with a hole in it you won’t find until you’re helplessly standing in it.
  1. Use tiered recovery, not one-size-fits-all. Your order-processing system and your archive file server don’t deserve the same RTO, and pretending they do wastes money on the trivial stuff and leaves you exposed on the critical stuff. Match the recovery tier to what the system is actually worth to the business.
  2. Have a one-page communication plan. Who gets notified, in what order, and what they’re told, the moment “this might be a real incident” becomes true.
  3. Loop in security, not just infrastructure, when you build the plan. A DR plan built in isolation from your security team tends to assume a clean, cooperative failure: a server dying, a data center outage. Ransomware doesn’t cooperate: it actively defies your assumptions. Identity revalidation, isolating what might still be compromised, deciding whether you’re even allowed to restore from a given backup yet: those are security calls as much as infrastructure ones, and they belong in the plan from the start, not bolted on after the tabletop exercise reveals the gap.
  4. Check where compliance already tells you what “good” looks like. If you’re already accountable to a framework: NIST, ISO 27001, PCI, SOC 2, NERC/FERC, CIS Controls, or whatever your industry requires, it almost certainly already has expectations for backup and DR maturity baked in. That’s a free starting checklist, not extra homework. Your cyber insurance policy is another one worth reading closely: insurers increasingly spell out specific backup requirements (immutability, offline copies, tested restores) as a condition of coverage, and finding that out during a claim instead of during the policy review is a bad way to learn it.
  5. If you don’t have this expertise in-house, go get it before you need it. This isn’t a sales pitch so much as thirty years of watching people learn it the hard way: a real DR plan takes work most internal teams don’t have the bandwidth or specialized experience to do well on top of their day jobs. Bringing in people who build and test these plans for a living: my own company, Long View Systems, does exactly this for organizations of every size, is how you end up with something viable, tested, and verifiable instead of a document that sounded good in a meeting. Great Googly Moogly, don’t wait to figure this out while it’s happening! Be prepared before you need to be.

That’s a real DR posture. Not a binder. A tested, prioritized, honest picture of what happens when things go sideways.

The takeaway

Backups answer “do I still have my data.” DR answers “how fast is my business actually functioning again, and in what order.” Treat them as one item on a checklist and you’ll find out the hard way that green backup reports don’t run payroll, and don’t answer the phone when a customer wants to know why the site’s down.

Test the restore before you need it. That’s the whole difference between a bad day and a bad month. (Or even whether you’ll still have a company to run once it’s all said and done.)

Skip the hype. Get the good stuff.

Practical AI, security, and cloud insights from a Principal Architect's desk: sent only when there's something worth your time.

We don’t spam! Read our privacy policy for more info.

Written by

Ken Gebhart

Real-world technology insights from an architect's perspective. Hi, I'm Ken, and on High Tech Yeti I share my passion for technology, artificial intelligence, cybersecurity, cloud computing, automation, and the latest innovations shaping our future. Drawing from years of experience in enterprise IT and solution architecture, I'll break down complex topics into practical, easy-to-understand content while exploring the coolest tech, gadgets, tools, and trends along the way. Topics include: • Artificial Intelligence • Cybersecurity & Governance • Azure & Cloud Technologies • Automation • Technology Trends & News • Reviews of Geeky Tech & Gadgets Technology Explained. Solutions That Work. Value That Lasts. Subscribe and join the HighTechYeti community!

More about HighTechYeti →

Leave a Reply

Your email address will not be published. Required fields are marked *