Business continuity planning is the organisation-wide discipline that keeps critical business functions operating during and after a disruption. It is a business programme rather than a technical one, and it runs through five phases in a fixed order: initiation, a business impact analysis, recovery strategy development, plan development, and testing with ongoing maintenance. Each phase consumes what the phase before it produced, which is why the order is a constraint rather than a convention.
Continuity sits across two CISSP domains and is easy to answer technically when the question is not technical. The governance and analysis work belongs to Domain 1, Security and Risk Management, and the operational recovery and testing work belongs to Domain 7, Security Operations. A candidate is expected to hold both halves at once: to know that senior management owns the programme, and to know which testing method a given situation calls for.
This guide walks the lifecycle and then the six testing methods in order of the disruption they cause. Its companion, the business impact analysis guide, goes inside phase 2, the phase every later decision is judged against.
What is business continuity planning?
Business continuity planning is the proactive, business-driven programme that ensures an organisation’s critical functions continue through a disruptive event and resume fully afterwards. The disruption might be a fire, a ransomware attack, the loss of a sole-source supplier, or a building becoming inaccessible for a week. The programme is indifferent to which of those happens, because it is built around consequence rather than cause.
Continuity is a business programme, not an IT project
The most consequential misunderstanding in this topic is treating continuity as something the IT department does. It covers people, premises, suppliers, processes and communications, and information systems are one dependency among several. An organisation whose payroll runs but whose staff cannot reach the building has not achieved continuity, and no amount of server availability fixes it. That is why the team assembled in phase 1 is cross-functional, drawing on business units, IT, legal, human resources and facilities.
Continuity, recovery and crisis management: how the terms divide
Business continuity keeps critical functions delivering during and after the disruption. Disaster recovery restores the technology they depend on. Crisis management directs the response while it is happening, coordinating decisions and stakeholder messaging. A fourth term, incident management, is narrower again: the detection, containment and resolution of a security event, which may or may not escalate into something the plan is invoked for. A scenario about deciding whether to activate the plan is describing continuity governance, not incident handling.
BCP vs DRP: what each plan is responsible for
BCP is the umbrella and DRP is one component beneath it. The two are presented as a pair so routinely that they read as alternatives, which they are not: an organisation does not choose between them, it has a continuity programme that contains a recovery plan.
| Business continuity planning (BCP) | Disaster recovery planning (DRP) | |
|---|---|---|
| Scope | The whole organisation and its mission | IT systems, infrastructure and data |
| Question it answers | How do we keep serving customers through a disruption? | How do we get our systems back online? |
| Primary focus | Business resilience and continuity of operations | Technical recovery of technology |
| Who owns it | Senior management, business driven | IT, executing under the continuity programme |
| Relationship to the other plan | The umbrella programme, the parent | One component underneath BCP, the child |
| Who takes part in a test | Business leaders and cross-functional teams | Technical recovery teams |
CISSP Exam Note
The distinction is most often tested by what a scenario asks about rather than by what it calls itself. A question describing continued service to customers is describing the BCP. A question describing servers, databases or a data centre coming back is describing the DRP. A purely technical answer to a continuity question misses the scope of what was asked.
Who is accountable for the business continuity programme?
Senior management is accountable for business continuity. It authorises the scope, allocates the budget, and carries the consequences when the programme fails. A coordinator manages day-to-day execution under delegated authority, and business function owners and IT supply the facts and build the capability. None of them owns the programme, because responsibility for a task can be delegated while accountability for the outcome cannot. Maintaining the programme is a standing due care obligation, and it is not discharged by naming someone else to do the work.
The failure mode is worth recognising in a scenario. Hand the whole programme to IT and continuity quietly becomes systems recovery, with the scope shrinking to whatever IT is able to restore. The functions that depend on people, premises or suppliers fall out of the plan entirely, and nobody notices until those are the things that fail.
What are the five phases of the BCP lifecycle?
The lifecycle is a structured sequence, and its ordering carries the argument. NIST SP 800-34 Rev. 1, the Contingency Planning Guide for Federal Information Systems, sets out a comparable seven-step process: the contingency planning policy statement, the business impact analysis, preventive controls, contingency strategies, the plan itself, testing and training, and plan maintenance. That document is an information system contingency guide rather than a business continuity standard, but the shape of the sequence is the same, and the analysis sits in the same place in both.
| BCP lifecycle phase | What happens in it | What it produces |
|---|---|---|
| Phase 1: project initiation and scoping | Senior management authorises the programme, defines its scope, appoints a coordinator and assembles a cross-functional team | A mandate, a budget and a team |
| Phase 2: business impact analysis | Critical functions are identified and the consequence of losing each one is measured over time | Criticality rankings, recovery targets and dependency maps |
| Phase 3: recovery strategy development | Strategies are selected that can meet the tolerances the analysis established | Chosen recovery sites, suppliers, workarounds and communications routes |
| Phase 4: plan development and implementation | The plan is written, distributed and made activatable | Procedures, contact lists, resource requirements and activation criteria |
| Phase 5: testing, training and maintenance | The plan is exercised, staff are trained, and the document is kept current | Evidence the plan works, and a plan that still matches reality |
Phase 1: project initiation and scoping
Getting the scope right here is what stops the programme quietly narrowing later, because the boundaries set at this point determine which functions are even eligible for analysis in phase 2.
Phase 2: the business impact analysis
The analysis produces the maximum tolerable downtime for each function, along with the recovery time objective and data loss tolerance that follow from it. It is the phase most likely to be skipped under time pressure, and skipping it is what produces plans that protect the wrong things carefully.
Phase 3: recovery strategy development
The constraint is the same for every strategy: restoring the function, and then doing the work that makes it usable again, has to fit inside the maximum tolerable downtime with margin left over. Site selection is the decision this phase is best known for, and the trade is cost against speed, from a hot site that fails over almost immediately, through a warm site that takes longer, to a cold site that is bare space. Which one is right follows entirely from the numbers phase 2 produced, and the options are compared in the disaster recovery sites guide.
Phase 4: plan development and implementation
A plan held in somebody’s head is not a plan, and neither is one nobody can reach. Storage is the detail most often overlooked: if the plan lives only on the systems it is meant to recover, a disruption that takes out the network takes the plan with it.
Phase 5: testing, training and maintenance
New systems, personnel, locations, suppliers and threats all age a plan, and one that is not maintained stops matching reality without anyone deciding that it should. Maintenance and testing are themselves sequenced, in that order: testing a stale plan measures an obsolete document, which is worse than not testing at all, because it produces evidence of readiness that is not true.
The six BCP and DRP testing methods, from least to most disruptive
Testing methods are arranged from least to most disruptive, and the reason to climb the ladder rather than jump to the top is that each rung buys a different class of evidence at a different level of operational risk. The cheap problems get caught cheaply, and confidence is built before anything is put at stake.
| Testing method | Disruption to live operations | What it validates | Its main limitation |
|---|---|---|---|
| Checklist review (read-through) | None | Accuracy of contact details, roles and resource lists, and roles left vacant by staff who have moved on | No coordination between teams is tested |
| Tabletop exercise | None | Decision-making, escalation paths, communication chains and the plan’s underlying logic | No procedure is physically executed |
| Structured walkthrough | Minimal | Whether documented steps are genuinely executable, and where teams compete for the same resources | No time pressure |
| Simulation test | Low | Performance under time pressure, against non-critical functions or a test environment | Production systems are never exercised |
| Parallel test | Low | That the recovery environment can carry the real workload | Production stays live, so a genuine cutover is untested |
| Full interruption test | High | End-to-end recovery under real conditions, using documented procedures only | A failed recovery is a real outage with real business impact |
Checklist review and tabletop exercise: no operational risk
A checklist review, also called a read-through, distributes the plan so each stakeholder confirms their own section. Handing it out also reveals roles left unfilled by people who have moved on. What it cannot show is whether the plan works as a coordinated effort.
A tabletop exercise gathers key personnel to talk through a scenario the moderator withholds until the session begins. No systems are activated, and it routinely surfaces conflicting assumptions nobody noticed while the plan was being written. Because it validates the plan’s logic without touching anything, it is the usual starting point for an organisation that has never tested.
Structured walkthrough and simulation test: rehearsal and pressure
A structured walkthrough has each team physically rehearse its procedures, at the locations and with the tools it would actually use, which exposes resource conflicts a discussion cannot reveal. A simulation test adds the clock, revealing bottlenecks and slow handoffs that only appear against a deadline. Alternate locations may be activated and partial data recovered, but scope stays on non-critical functions or a test environment so production is untouched.
Parallel test vs full interruption test: the safety net
These two are the pair most easily confused, and the distinction is the highest-value point in the topic. Both activate the recovery site with real systems. The difference is whether production keeps running.
In a parallel test, staff relocate, the site-activation procedures run, and recovery duties are carried out on backup systems processing alongside live production. The safety net is intact, so a failure during the test does not take the business down. In a full interruption test, the primary site is deliberately shut down and operations shift across. There is no safety net, which is why it gives the highest confidence and why a failed recovery during one is a real outage with real business impact.
CISSP Exam Note
Where a scenario describes an organisation that has never tested its plan, the expected answer is usually the tabletop exercise, because the plan’s logic is validated before anything is risked. Where it describes a plan that has not been updated through a migration or a change of leadership, the expected answer is to bring the plan current first: testing an obsolete document produces misleading evidence of readiness.
Worked example: matching a recovery strategy to a 48 hour MTD
An organisation completes its business impact analysis and finds that email has a maximum tolerable downtime of 48 hours. The IT director proposes a hot site.
The proposal is not wrong because a hot site would fail. It is wrong because the analysis has already told the organisation it does not need one: a hot site carries the maintenance cost of duplicated live infrastructure, for a function that can be down for two days. Overspending on recovery is as much a misalignment between capability and requirement as underspending, and the analysis output is what judges both.
The proportionate answer is a warm site, and the check that decides it is not the tolerable downtime alone. Restoring the systems is only part of the elapsed time, and validating data and resuming normal processing afterwards has to fit inside the same window. Where recovery is measured in hours to days, that margin needs looking at rather than assuming. The metrics are worked through in the recovery metrics guide.
Where business continuity plans fail
- Treating continuity as IT’s job. The programme narrows to systems recovery and loses the business-wide view that defines it.
- Selecting a strategy before the analysis. Buying a recovery site before knowing what is critical and how urgently inverts the lifecycle and protects the wrong things well.
- A plan that is never tested, or never maintained. An untested plan is an assumption, and a plan nobody updates stops matching reality without any decision being taken.
- A plan stored only on the systems it must recover. If the disruption takes out the network, the plan is unreachable at the moment it is needed.
- Omitting communications from the test. A recovery that restores systems but fails to notify the right people in time is still a failed response, and regulated organisations often face mandatory reporting windows.
- No after-action follow-up. If the gaps a test reveals are never fed back into the plan, the exercise achieved nothing beyond documenting them.
Conclusion
Business continuity planning is a governance obligation carried out through a sequence, and both halves of that sentence matter. Senior management owns the programme and cannot delegate its accountability, and the five phases run in an order where each consumes what the last produced. The analysis has to precede strategy selection because a strategy is an answer, and until the analysis has asked the question there is nothing for it to be an answer to.
The part most easily deferred is the last one. Testing and maintenance never feel urgent until a disruption arrives, which is exactly when it is too late to start. A plan’s value has never been that it exists. It is the confidence, demonstrated through exercises that someone actually ran, that it will work when it is invoked.
If you want to test that understanding under exam conditions, the LSM CISSP practice tests are built around exactly these scenarios: which lifecycle phase comes next, who is accountable when a plan fails, and which testing method a described situation actually calls for, with full explanations for every answer.
Quick reference for the CISSP exam
The lifecycle in order
Project initiation and scoping, business impact analysis, recovery strategy development, plan development and implementation, then testing, training and maintenance. Each phase consumes the previous phase’s output. The analysis always precedes strategy selection.
BCP against DRP
BCP is the umbrella: organisation-wide, business-driven, owned by senior management, concerned with continuing to serve customers. DRP is one component beneath it: IT-focused, concerned with restoring systems, infrastructure and data. BCP is the parent, DRP is the child.
The testing ladder
Checklist review, tabletop exercise, structured walkthrough, simulation test, parallel test, full interruption test. The first three carry no operational risk. The parallel test keeps production live; the full interruption test does not, which is the whole distinction between them.
Ownership
Senior management is accountable and cannot delegate that. The coordinator executes. Business function owners and IT supply facts and build capability. An option that has IT deciding what the business can survive without is one to treat with suspicion.
Spotting the topic in a scenario
Continued service to customers, cross-functional teams, or the question of who authorises the programme point at BCP. Servers, data centres and restoration point at DRP. An organisation that has never tested points at the tabletop exercise. An organisation whose plan predates a significant change points at maintenance before testing.



