Skip to content

Business Continuity Planning: The Third Thing Backup and DR Don’t Cover

I sat in on a recovery once where every system came back exactly on schedule. Servers up. Data restored. Identity revalidated, clean. The IT team should have been taking a well-earned victory lap. Instead the building was still dead quiet, because nobody could take a customer call: the phone system routed through the data center that had just gone down, and not one person on the floor knew the backup number to forward calls to, because that number lived in a document on a file share that also hadn’t come back yet, on a system that also required the same restoration steps. Technically, IT had done everything right. The business was still not open.

That’s the gap I flagged in the last piece and promised to come back to: backup gets your data back, DR gets your systems back, and business continuity planning is the one that gets your business back. It’s not a footnote to the other two. It’s a separate discipline, with its own owner, its own document, and its own set of questions, and it’s the one most organizations skip because it doesn’t feel like an IT problem. It’s not, entirely. That’s exactly why it gets missed.

What BCP is actually answering

Backup answers “do I still have the data.” DR answers “how fast are my systems back, and in what order.” BCP answers a completely different question: if your systems, and your people, can’t do what they normally do, for an hour, a day, a week, how does the business keep functioning anyway?

That’s a people-and-process question first, and a technology question second. A BCP that only talks about servers is really just a DR plan with a new cover page.

The scope is broader than most people expect going in:

  • Which functions genuinely can’t stop, even briefly, and which can tolerate a delay. Payroll on the 1st and 15th behaves differently than restocking the supply closet.
  • Where people work if the office, the data center, or a whole region is unreachable.
  • Manual workarounds for the systems that are down. Can you take an order with a paper form and a phone instead of the CRM? Can you process a payment without the card terminal?
  • Who’s actually in charge if the people normally in charge can’t be reached: a real delegation-of-authority list, not an assumption that someone will figure it out.
  • Vendors and supply chain: the dependencies outside your walls that you don’t control but absolutely rely on.
  • Communication: to employees, customers, and sometimes regulators, on a timeline measured in minutes, not “whenever someone gets around to drafting a statement.”

None of that shows up in a restore-time report. All of it shows up in whether the business is actually open the next morning.

The BIA: the one piece of jargon worth learning

If you take one term away from this article, make it BIA: Business Impact Analysis. It’s the exercise that turns “we should probably have a plan” into an actual, prioritized plan, and it’s simpler than it sounds.

A BIA walks through each business function and asks three things:

  1. What happens if this stops? Not vaguely, specifically. Lost revenue, missed compliance deadlines, safety risk, reputational damage.
  2. How long can it be down before that damage becomes serious? This is your MTPD, or Maximum Tolerable Period of Disruption. Some functions tolerate a day. Some tolerate about twenty minutes before someone’s calling you in a panic.
  3. What does this function actually depend on? People, systems, physical space, specific vendors, specific data: the full dependency chain, not just “the app.”

Do this honestly for your handful of truly critical functions. Not all forty things the business does, just the ones that would actually hurt if they stopped. Now you’ve got the prioritization spine the rest of the plan hangs off of. Skip it, and you end up building a beautifully detailed plan for the process that matters least, because that’s the one someone happened to volunteer to write up first.

The part that isn’t about computers at all

Here’s where BCP genuinely splits off from anything IT normally owns.

Alternate work locations. If the office is unusable (fire, flood, an extended power outage, a building that’s suddenly a crime scene, riots or civil disorder nearby), where does the team actually work from, and does everyone know that answer without having to ask? “Work from home” isn’t a plan; it’s an assumption that half your team has a reliable home setup and the other half hasn’t thought about it once.

Manual workarounds. This is the one I see skipped most often, because it feels almost embarrassingly low-tech next to a DR runbook full of failover automation. But when the system’s down and the MTPD clock is running, a manual workaround is the thing that keeps revenue moving while IT does its job. A paper order form. A pre-printed script for the phone team. A way to run a credit card manually. These feel like relics until the exact hour you need one, at which point they’re the only thing standing between “degraded but open” and “closed until further notice.”

Delegation of authority. Who makes the call if the person who normally makes the call is unreachable: on a plane, in a hospital, or simply asleep in a different time zone when the incident starts at 2 AM their time? Write the actual names down, in order, with how to reach each one. Make sure they understand their responsibilities and what that means for their piece of the plan. Saying “someone senior will handle it” is not a plan anyone can execute under pressure.

Vendor and supply chain dependencies. Your BCP doesn’t stop at your own walls. If a key supplier, a payment processor, or a logistics partner goes down, that’s your problem too, even though you don’t control their recovery. Know which vendors are load-bearing for your critical functions, and know if they have their own continuity plan you’ve actually seen, not one they’ve just told you exists.

Communication. This deserves more than the one-page notification plan I mentioned in the DR piece; that plan tells people an incident is happening. BCP’s communication piece covers what customers are told, what regulators need to hear and by when if you’re in a regulated industry, and who’s authorized to say anything publicly at all before legal or leadership has weighed in. Getting this wrong in the first hour can cost you more in trust than the outage itself.

How this plays out when identity and systems actually come back clean

Go back to the scenario I opened with. Say the DR plan in the last article had been followed to the letter: identity revalidated first, systems restored in the right dependency order, RTOs met. Good DR execution, genuinely. And the business was still effectively closed, because nobody had planned for the fact that “systems are back” and “the business can resume operating right at this moment” aren’t the same instant. There’s a gap between those two points, sometimes minutes, sometimes hours, and BCP is what you’re operating on during that gap.

A well-run DR effort with no BCP behind it gets you a beautifully restored, fully staffed office where nobody can do their job yet. That’s not a hypothetical; it’s the most common failure mode I see in orgs that treated “we have backup and DR” as the whole answer.

So what do you actually do about it

Same approach as last time: you don’t need a binder nobody will read. Start here.

  1. Run a lightweight BIA on your true critical functions. Not all of them, just the handful that would genuinely hurt the business if they stopped. Assign each one an MTPD.
  2. Write down the manual workaround for your top three critical functions. If there isn’t one, that’s the finding, not a reason to skip the exercise.
  3. Build a real delegation-of-authority list. Names, order, contact info. Test it by actually calling person #2 sometime when nothing’s on fire, just to confirm the number still works.
  4. Pick an alternate work location and tell people where it is before you need it, not during.
  5. List your load-bearing vendors and find out, actually ask, whether they have a continuity plan of their own.
  6. Draft the customer and regulatory communication piece, including who’s authorized to speak publicly, before an incident forces you to figure that out live.
  7. Test it with the actual people, not just IT. A DR tabletop exercises the technical recovery. A BCP tabletop exercises whether the front-desk person, the ops manager, and the person who normally approves purchase orders all know what they’re supposed to do. Run both; they’re testing different failure modes.

That’s the whole gap between “we have backup and DR” and “we can actually stay open.” According to a U.S. Chamber of Commerce Foundation survey, 94% of businesses believe they’d recover from a serious disruption, but only about 26% actually have a documented plan in place. That’s a wide gap between confidence and preparation, and it’s usually the BCP piece, not the backup or DR piece, sitting in that gap.

The takeaway

Backup answers “do I still have my data.” DR answers “how fast are my systems back.” BCP answers the question that actually determines whether you have a business to run once the systems are back: can the people keep working, in the meantime, without them?

Test the restore. Then test whether anyone can actually answer the phone while you’re doing it.

Skip the hype. Get the good stuff.

Practical AI, security, and cloud insights from a Principal Architect's desk: sent only when there's something worth your time.

We don’t spam! Read our privacy policy for more info.

Written by

Ken Gebhart

Real-world technology insights from an architect's perspective. Hi, I'm Ken, and on High Tech Yeti I share my passion for technology, artificial intelligence, cybersecurity, cloud computing, automation, and the latest innovations shaping our future. Drawing from years of experience in enterprise IT and solution architecture, I'll break down complex topics into practical, easy-to-understand content while exploring the coolest tech, gadgets, tools, and trends along the way. Topics include: β€’ Artificial Intelligence β€’ Cybersecurity & Governance β€’ Azure & Cloud Technologies β€’ Automation β€’ Technology Trends & News β€’ Reviews of Geeky Tech & Gadgets Technology Explained. Solutions That Work. Value That Lasts. Subscribe and join the HighTechYeti community!

More about HighTechYeti →

Leave a Reply

Your email address will not be published. Required fields are marked *