7.12 Disaster Recovery Testing

CISSP Domain 7 · Security Operations

7.12 Disaster Recovery Testing

A Disaster Recovery Plan can look perfect on paper and still fail when it is actually needed.

People may not know their roles. Backups may not restore. Contact details may be outdated. A recovery site may lack capacity. DNS may be forgotten. A supplier may not respond. The documented RTO may simply be impossible.

Disaster Recovery testing provides evidence that the organisation can activate, coordinate and perform recovery under realistic conditions.

The objective is not merely to prove that a document exists.

It is to discover weaknesses: before the disaster discovers them for you.

📋

Exercise the Plan

Verify that roles, decisions, procedures and communications make sense.

CAN WE FOLLOW IT?
🧪

Test the Capability

Verify that technology, backups, failover and recovery mechanisms actually work.

CAN WE RECOVER?
📈

Improve

Record weaknesses, assign corrective actions and retest.

CAN WE DO IT BETTER?
Current CISSP 7.12 Scope

Test Disaster Recovery Plans

The current CISSP Exam Outline explicitly identifies:

Read-Through / Tabletop

Review and discuss the plan and a recovery scenario.

Walkthrough

Step through recovery procedures in greater operational detail.

Simulation

Respond to a realistic simulated disaster without deliberately interrupting normal production.

Parallel

Recover operations using alternate capability while the production environment continues operating.

Full Interruption

Stop or transfer actual production activity and rely on the recovery capability.

Communications

Coordinate stakeholders, test status and required external communications such as regulators where applicable.

Official 7.12 Topics

READ Read-through / tabletop
WALK Walkthrough
SIMULATE Realistic scenario
PARALLEL Recover beside production
INTERRUPT Use recovery for real
COMMUNICATE Coordinate everyone
Critical Domain 7 Distinction

Strategy · Process · Test

7.10 Recovery Strategies

What recovery capability has been designed?

Backups Alternate Sites Replication High Availability

BUILD IT

7.11 DR Processes

How will the organisation actually execute recovery?

Respond Assess Coordinate Restore

USE IT

7.12 DR Testing

Does the recovery plan and capability really work?

Exercise Observe Measure Improve

PROVE IT

7.10 · 7.11 · 7.12

7.10 Build recovery capability
7.11 Perform recovery
7.12 Prove recovery works

Why Test Disaster Recovery?

Validate the Plan

Confirm procedures are complete, understandable and usable.

Validate People

Confirm personnel understand their responsibilities.

Validate Technology

Confirm backups, recovery sites, infrastructure and failover capability operate correctly.

Validate Dependencies

Discover systems or suppliers the plan forgot.

Validate Communications

Confirm responders and stakeholders can communicate during disruption.

Validate Objectives

Determine whether RTO and RPO can actually be achieved.

Train Personnel

Give responders practical familiarity with recovery activities.

Find Weaknesses

Identify problems while the organisation still has time to correct them safely.

The purpose of a DR test is not to get a perfect score. Finding weaknesses during a controlled exercise is a successful outcome.
Testing Lifecycle

Plan → Exercise → Measure → Improve → Retest

Recovery Requirement → Define Objectives
Objectives → Select Test Type
Test Type → Plan Safely
Plan → Execute
Exercise → Observe & Measure
Results → Identify Gaps
Gaps → Correct
Correction → Retest

DR Testing Lifecycle

OBJECTIVE What are we proving?
EXERCISE Perform the test
MEASURE What happened?
IMPROVE Correct weaknesses
RETEST Prove correction

Start With a Test Objective

A useful exercise should answer a defined question.

Weak objective

"Test disaster recovery."

Better objectives
Restore Customer Database Within 2 Hours Validate Alternate-Site Capacity Test Emergency Call Tree Verify Cloud Region Failover Restore From Immutable Backup Validate Supplier Activation
If you do not know what the test is supposed to prove, you cannot meaningfully decide whether it succeeded.
Before Testing

Define Success Criteria

Recovery Time

Did the service recover within its RTO?

Recovery Point

Was data restored within the required RPO?

Functionality

Did the recovered service perform its required function?

Security

Were required controls functioning?

Capacity

Could the recovery environment support required workload?

Communications

Were required stakeholders reached?

Test success should be defined before the test starts.
CISSP Test Spectrum

From Discussion to Real Interruption

Read-Through / Tabletop → Discuss
Walkthrough → Step Through
Simulation → React to Scenario
Parallel → Recover Separately
Full Interruption → Rely on Recovery
Generally, greater operational realism provides stronger evidence but introduces greater cost, complexity and potential business risk.

DR Test Types at a Glance

TestWhat Happens?Production ImpactAssurance
Read-Through / TabletopDiscuss plan and scenarioVery lowPlan / people
WalkthroughStep through proceduresLowPlan / operational detail
SimulationRespond to realistic simulated eventUsually lowDecision / coordination
ParallelRecover alternate environment while production continuesLimitedTechnical recovery
Full InterruptionProduction relies on recovery capabilityHighestHighest operational realism
Official Test Type

Read-Through

Participants review the DR documentation to determine whether the plan is:

Complete Understandable Current Consistent Operationally Plausible
During review

Recovery procedure says:

"Contact the backup administrator."

But: no administrator or contact details are identified.

Read-through testing can identify documentation defects cheaply and safely before more complex exercises occur.
Official Test Type

Tabletop Exercise

Participants discuss how they would respond to a hypothetical disaster.

A facilitator introduces a scenario and asks participants to explain decisions and actions.

Scenario

At 09:15: primary data centre loses power.

Generator: fails to start.

Estimated power restoration: 12 hours.

Payment platform RTO: 2 hours.

Facilitator Questions

Who Declares DR? Who Is Contacted? Which Site Activates? Which Systems Recover First? Who Communicates With Customers? What If the Hot Site Fails?

Tabletop

SCENARIO Facilitator presents event
DISCUSS Participants explain actions
DISCOVER Find plan gaps
💬 Tabletop Success Does Not Prove Technical Recovery Talking through a restore is not the same as restoring
Tabletop answer

Database team says:

"We would restore the database from backup."

Unknown:

Is Backup Valid? How Long Restore Takes? Are Keys Available? Is Capacity Sufficient?
Tabletop validates thinking and procedures. Technical testing validates technical capability.
Official Test Type

Walkthrough

A walkthrough steps through the recovery plan in operational sequence.

Participants may examine:

Recovery Procedures Recovery Facilities Equipment Contact Information Access Requirements Dependencies
Walkthrough

DR team travels to: alternate recovery facility.

They verify:

Access Badges Work Recovery Documentation Available Network Equipment Present Supplier Contact Valid Workspace Exists
Walkthrough asks: can we follow the recovery path in the real environment?
🚶 Tabletop vs Walkthrough Discuss the response vs step through the process
Tabletop

Participants primarily discuss: what they would do.

TALK IT THROUGH

Walkthrough

Participants step through: how the process would actually proceed.

WALK IT THROUGH

Official Test Type

Simulation

A simulation creates a realistic disaster scenario and requires personnel to respond as though the event were real.

Production usually remains protected from deliberate interruption.

Simulation

Scenario: ransomware disables corporate identity infrastructure.

Participants must:

Activate Crisis Team Use Alternate Communications Prioritise Services Request Backup Recovery Engage Suppliers Prepare Regulatory Communications
Simulation adds realistic decision pressure without necessarily deliberately taking production offline.
🎭 Scenario Injects Make the exercise evolve instead of following a predictable script
Initial scenario

Primary cloud region: unavailable.

Later injects:

Secondary Region Capacity Only 60% DR Manager Unreachable Customer Calls Increase Regulator Requests Update Backup Restore Fails DNS Provider Also Degraded
Good simulations test decision-making when the scenario does not go exactly according to plan.
Official Test Type

Parallel Test

In a parallel test, recovery systems are activated while the normal production environment continues operating.

Production → Continues Serving Users
Recovery Environment → Activated Separately
Data → Restored / Replicated
Application → Started
Tests → Validate Recovery Capability
Parallel testing provides significant technical assurance while avoiding deliberate shutdown of the primary production environment.
🖥️ Parallel Recovery Example Recover the service without moving real users
Production

Customer payment system: continues normally in London.

DR Test

Recovery team restores: the payment platform in the alternate site.

Test transactions: are executed against the recovered environment.

Real customer traffic remains: on the primary system.

Parallel = recovery environment runs beside production.

What Parallel Testing Still Does Not Prove

A successful parallel test provides strong evidence, but some aspects of a real failover may remain untested.

Real User Redirection Actual DNS Cutover Full Production Load Real Customer Behaviour Production Shutdown True Failback
A parallel test can prove that the alternate environment works without necessarily proving the complete production cutover process.
Official Test Type

Full Interruption Test

A full interruption test deliberately stops or transfers actual production activity and requires the organisation to use its recovery capability.

Primary Production → Stop / Disconnect / Transfer
DR Plan → Activate
Alternate Environment → Take Over
Real Workload → Run From Recovery Capability
Validation → Confirm Business Service
Highest operational risk

A failed full-interruption test can become a real business outage. It therefore requires careful planning, authorisation, rollback and business involvement.

🔌 Full Interruption Example The organisation genuinely relies on the recovery environment
Planned DR exercise

Customer service: moved from Data Centre A to Data Centre B.

Production traffic: actually redirected.

Users: operate against the DR environment.

Full interruption provides the strongest realism because recovery capability becomes the live service.
Test Governance

Do Not Cause the Disaster You Are Testing

Higher-impact recovery tests should themselves be risk managed.

Authorisation

Ensure appropriate business and technical authority approves the exercise.

Scope

Clearly identify what systems and services are included.

Maintenance Window

Choose an appropriate time where disruption can be managed.

Rollback

Define how normal service will be restored if the test fails.

Stop Criteria

Define circumstances requiring the exercise to be terminated.

Monitoring

Observe both primary and recovery environments throughout the test.

Testing recovery should reduce uncertainty - not introduce unmanaged production risk.
🛑 Define Stop Criteria Know when the exercise should be terminated
Full interruption test

Stop the exercise if:

Customer Data Integrity at Risk Safety Issue Occurs Rollback Becomes Impossible Unexpected Major Outage Recovery Environment Becomes Unstable
A controlled test should always retain appropriate control of the test itself.

Protect Data During Testing

DR tests may involve copies of real production information.

Confidentiality

Protect sensitive data in recovery environments.

Access Control

Restrict access to authorised recovery personnel.

Test Data

Use suitable masked, synthetic or protected data where appropriate.

Cleanup

Remove test copies when no longer required according to policy.

A recovery test should not create a new confidentiality breach.
Technical Testing

Restore Testing

One of the most valuable DR tests is simple:

actually restore the backup.

Backup Job → Successful
Restore → Attempt
Data → Validate
Application → Use Recovered Data

Backup Test

BACKUP Did it run?
RESTORE Can we recover?
VALIDATE Is data usable?
Successful backup ≠ successful recovery.
Measure Reality

Test the RTO

Documented requirement

RTO: 2 hours.

Parallel recovery test

Actual recovery: 5 hours 40 minutes.

Result:

the organisation has evidence that its current recovery capability cannot meet the business requirement.

An untested RTO is an objective. A measured RTO is evidence.

Test the RPO

Required RPO

15 minutes.

Recovered database

Latest available transaction: 55 minutes before simulated failure.

Recovery may technically succeed while still failing the required RPO.
Common DR Finding

Test the Dependency Chain

Application recovery

Application servers: restored successfully.

Database: restored successfully.

Users: still cannot log in.

Cause: identity provider not included in the DR exercise.

Testing often reveals dependencies that documentation failed to identify.

Test Recovery Capacity

Hot site

All applications start: successfully.

Simulated workload: 20% of production.

Actual DR requirement: 100% of critical workload.

Recovery environment existing ≠ recovery environment having enough capacity.
People

Test Personnel Too

Role Knowledge

Do responders know what they are responsible for?

Alternate Personnel

Can another person perform the task if the primary specialist is unavailable?

Access

Do recovery personnel have the credentials and physical access they need?

Contactability

Can the recovery team actually be reached?

The strongest recovery technology still fails if nobody knows how to operate it.
👥 Test the Alternate Person Do not always let the expert perform the exercise
Normal recovery exercise

Senior database engineer: always performs the restore.

Next exercise:

Assume senior engineer is: unavailable.

Alternate administrator must recover using: the documented procedure.

A good DR test can validate whether documented knowledge is actually transferable.
Official 7.12 Topic

Communications During DR Testing

Communications are explicitly part of the current CISSP 7.12 objective.

Participants

Know the exercise scope, timing and responsibilities.

Stakeholders

Understand whether observed disruption is real or part of the test.

Management

Receive appropriate exercise status and outcomes.

Service Owners

Understand potential impact to their systems.

Customers / Users

Receive appropriate advance or exercise-related communication where required.

Regulators

May require communication or awareness depending on applicable requirements and the nature of the exercise.

📢 Clearly Identify Test Status A simulation should not accidentally trigger a real-world panic
Exercise message

Participant receives:

"Critical payment systems compromised. Customer accounts affected."

If exercise communications are poorly controlled, somebody may:

Contact Media Escalate to Regulators Trigger Real Incident Response Inform Customers Incorrectly
Exercise communications should make test status clear to the appropriate recipients while preserving realism for participants where required.
Communication Scenario

Teams Is Down

DR exercise assumes: corporate identity platform has failed.

DR plan tells responders to coordinate using: Microsoft Teams.

Teams requires: the unavailable identity service.

The exercise has discovered a circular recovery dependency.
External Dependencies

Include Critical Suppliers

Recovery plans can depend on:

Cloud Providers Telecommunications Providers Hot-Site Providers SaaS Providers Managed Service Providers Hardware Vendors
Plan assumption

Telecom provider will activate: alternate WAN circuit within 30 minutes.

Test result: actual activation takes four hours.

Test assumptions that depend on external organisations whenever practical and authorised.
☎️ Supplier Contact Test Even a simple call can reveal important problems
Recovery plan

Emergency number: listed.

During test: number disconnected.

A low-risk communications test can expose a high-risk recovery weakness.

Announced vs Unannounced Exercises

Organisations may use different levels of participant notice depending on the objective and risk of the exercise.

Announced

Participants know the exercise will occur.

Useful for: learning, preparation and lower-risk validation.

Limited-Notice / Unannounced

Can provide more realistic evidence of readiness.

But requires: careful governance and safety controls.

Do not confuse realism with recklessness

Surprise should never create unmanaged safety, legal or operational risk.

Evaluation

Use Observers & Evaluators

Participants perform the recovery.

Evaluators observe what actually happens.

Decision Times Recovery Times Communication Delays Procedure Deviations Unexpected Dependencies Technical Failures Successful Actions
Capture facts during the exercise rather than relying entirely on memory afterwards.

Build an Exercise Timeline

TimeEvent
10:00Simulated outage begins
10:12Operations identifies major disruption
10:30DR formally activated
11:05Alternate infrastructure available
12:45Database restored
13:15Application available
13:40Business validation complete

If the RTO was: 3 hours

and usable service required: 3 hours 40 minutes,

the test has identified: a recovery capability gap.

CISSP Mindset

A Failed DR Test Can Be Valuable

Test result

Database recovery: failed.

Cause: backup encryption key inaccessible at alternate site.

The correct conclusion is not:

"The exercise was a failure and should be hidden."

The organisation has discovered a serious recovery defect: before the actual disaster.

A test that safely reveals a weakness has performed one of its most important functions.
After the Exercise

After-Action Review

What Worked?

Preserve effective procedures.

What Failed?

Identify recovery gaps.

What Was Slow?

Identify delays affecting the RTO.

What Was Missing?

Identify forgotten dependencies or resources.

What Was Confusing?

Improve documentation and responsibilities.

What Should Change?

Define corrective actions.

Turn Findings Into Actions

Observation → Backup Restore Took 6 Hours
Requirement → RTO 2 Hours
Root Cause → Insufficient Recovery Bandwidth
Action → Increase Recovery Throughput
Owner → Infrastructure Team
Verification → Retest
Observation → cause → action → owner → retest.

Improvement Loop

FIND Identify gap
ANALYSE Understand cause
ASSIGN Give ownership
FIX Correct weakness
RETEST Prove improvement
Continuous Improvement

Update the DR Plan

Testing often reveals that the documented environment no longer matches reality.

New Systems Retired Systems Changed Suppliers New Phone Numbers Changed Network Architecture New Cloud Regions Different Recovery Priorities Updated RTO / RPO
Test → discover → update.
🔄 Production Change Should Trigger DR Review Recovery capability must follow the environment it protects
Production change

Application moves from: on-premises database

to: managed cloud database.

Existing DR plan still says:

"Restore database from tape at alternate data centre."

Significant change can invalidate previously successful recovery testing.

How Often Should DR Be Tested?

There is no single universal frequency that is correct for every organisation and every system.

Frequency should reflect:

Business Criticality Risk Regulation Contractual Requirements Technology Change Previous Findings Recovery Complexity
Higher criticality and greater change generally justify stronger and more frequent assurance.
Compliance

Regulatory & Contractual Requirements

Some organisations may have external requirements governing:

Test Frequency Test Scope Evidence Retention Regulatory Notification Third-Party Participation Recovery Objectives
DR exercises should account for applicable legal, regulatory, contractual and organisational requirements.

Keep Evidence of the Test

Scope

What was tested?

Participants

Who participated?

Scenario

What disruption was simulated?

Timeline

When did important events occur?

Results

Which objectives succeeded or failed?

Findings

Which weaknesses were identified?

Corrective Actions

What must change and who owns it?

Evidence

Logs, screenshots, reports and recovery measurements as appropriate.

Useful Distinction

Training · Exercise · Test

Training

Develop the person's ability to perform a role.

CAN THE PERSON DO IT?

Exercise

Practise plans, coordination and decision-making under a scenario.

CAN THE TEAM RESPOND?

Technical Test

Verify the operability of recovery components and systems.

DOES THE TECHNOLOGY WORK?

Mature DR assurance normally needs people, process and technology to be evaluated together over time.
Tabletop Scenario

Ransomware at 03:00

Facilitator announces:

Domain controllers, file servers and backup-management servers show ransomware activity.

Participants discover:

DR Lead Is on Holiday Corporate Email Unavailable Newest Backups May Be Compromised Customer Service Cannot Authenticate
Tabletop exercises are excellent for exposing decision, communication and dependency problems without deliberately damaging production.
Walkthrough Scenario

The Recovery Facility

DR team physically walks through the alternate site procedure.

They discover:

Two Access Badges Expired Firewall Documentation Outdated Backup Appliance Model Changed Emergency Phone Missing
None of these problems required a real disaster to discover.
Parallel Scenario

The Database Restore

Production database: continues serving customers.

Recovery team restores: the database to the alternate site.

Documented RTO: 90 minutes.

Actual restore: 4 hours.

Parallel testing has proven that the current recovery process cannot meet the RTO without risking a production outage.
Full Interruption Scenario

Weekend Data Centre Failover

Organisation deliberately: moves live production to the recovery site.

Recovery works, but:

Authentication Is Slow Batch Jobs Fail Monitoring Missing Recovery Site Reaches 95% Capacity
A technically successful failover can still reveal important operational and security weaknesses.
Measurement Scenario

"The Servers Came Up"

Infrastructure team reports:

"DR test successful. All 40 servers started."

But:

Customers Cannot Log In Payment API Cannot Reach Supplier SIEM Receives No Logs
Infrastructure recovery ≠ business-service recovery.
Supplier Scenario

The 30-Minute Network Recovery

DR documentation states:

"Carrier will activate alternate circuit in 30 minutes."

Exercise discovers:

current contract specifies four-hour activation.

Testing validates assumptions against reality.
Documentation Scenario

The Server That No Longer Exists

DR walkthrough instructs:

"Restore APP-SRV-27 first."

APP-SRV-27 was: retired 18 months ago.

Application is now: containerised in the cloud.

Recovery plans must evolve with production architecture.
RPO Scenario

The Successful Restore That Lost Too Much Data

Database restoration: successful.

Recovery time: within RTO.

Business RPO: 15 minutes.

Recovered data: two hours old.

The test failed the RPO even though the technology successfully restored.
Personnel Scenario

The Expert Is Unavailable

Exercise inject:

primary storage administrator cannot participate.

Alternate administrator attempts the documented procedure.

Result: procedure omits three critical commands.

The exercise has simultaneously tested personnel resilience and documentation quality.
Communications Scenario

The Test Notification Mistake

Exercise simulates: large customer-data breach.

A participant forwards an exercise message externally without the exercise context.

Recipient interprets it as: a real incident.

DR test communication requires control just like real crisis communication.
🎓 CISSP Scenarios Recognise the Disaster Recovery testing method or principle
Scenario 1

Personnel review the DR document for completeness without performing recovery.

Which method?

Read-through.

Scenario 2

A facilitator describes a flood and asks recovery teams what they would do.

Which exercise?

Tabletop.

Scenario 3

Does a successful tabletop prove backup restoration works?

Answer?

No.

Scenario 4

Personnel step through recovery procedures and inspect the alternate facility.

Which method?

Walkthrough.

Scenario 5

The team discovers its recovery-site access badges no longer work.

Which exercise could easily reveal this?

Walkthrough.

Scenario 6

A realistic ransomware scenario is presented and participants respond to evolving events.

Which method?

Simulation.

Scenario 7

Production remains operational while the alternate environment is restored and tested.

Which method?

Parallel test.

Scenario 8

Real customer traffic remains on production during a parallel test.

Primary advantage?

Strong recovery testing with lower production interruption risk.

Scenario 9

The primary environment is deliberately shut down and users rely on the recovery site.

Which test?

Full interruption.

Scenario 10

Which of the listed CISSP DR tests generally creates the greatest operational risk?

Answer?

Full interruption.

Scenario 11

A full-interruption test fails and creates a real outage.

Why should such tests be carefully governed?

The test itself can affect production.

Scenario 12

Management wants more assurance than a tabletop but does not want to shut production down.

Strong option?

Parallel recovery testing.

Scenario 13

A tabletop discovers that nobody knows who can declare a disaster.

Was the exercise useful?

Yes. It identified a governance gap.

Scenario 14

A simulation produces several unexpected problems.

Does this automatically mean the exercise failed?

No. Identifying weaknesses is a purpose of testing.

Scenario 15

A test objective says only "test DR."

Primary weakness?

No clear measurable objective or success criteria.

Scenario 16

Actual recovery takes six hours against a two-hour RTO.

What has the test demonstrated?

Current recovery capability does not meet the RTO.

Scenario 17

Service recovers within RTO but restored data is three hours old against a 30-minute RPO.

Did recovery meet requirements?

No. The RPO was missed.

Scenario 18

Backup jobs report success every night.

What should DR assurance also include?

Actual restore testing.

Scenario 19

A restore succeeds but the application cannot use the recovered database.

Primary lesson?

Restore validation must include service usability.

Scenario 20

A test restores application and database but forgets DNS.

What did the test reveal?

An undocumented recovery dependency.

Scenario 21

Recovery site works for 50 test users but production has 40,000 users.

What else should be validated?

Recovery capacity.

Scenario 22

The only administrator capable of restoring storage performs every exercise.

What should a future test consider?

Testing alternate personnel.

Scenario 23

An alternate administrator cannot follow the documented recovery procedure.

What does this indicate?

Training and/or documentation weakness.

Scenario 24

The DR contact list contains disconnected numbers.

Which capability failed?

Recovery communications.

Scenario 25

Corporate email is unavailable in the scenario and no alternate communication channel exists.

What has been identified?

A communications dependency gap.

Scenario 26

A test message is mistaken for a genuine regulatory notification.

Primary problem?

Exercise communications were not adequately controlled.

Scenario 27

A recovery provider's emergency number is never tested.

What risk remains?

Supplier activation capability is unverified.

Scenario 28

A telecommunications contract promises four-hour recovery while the plan assumes 30 minutes.

What has the test found?

A recovery assumption that does not match reality.

Scenario 29

Why should observers record events during an exercise?

Best answer?

To create accurate evidence for evaluation and improvement.

Scenario 30

A DR test identifies a weakness but nobody is assigned to correct it.

Primary problem?

The finding may never result in improvement.

Scenario 31

A weakness is corrected after an exercise.

What provides stronger assurance that it is resolved?

Retesting.

Scenario 32

A plan is updated after every test but technical weaknesses are never fixed.

Is that sufficient?

No. Process and technical corrective actions may both be required.

Scenario 33

A major production architecture change occurs after the last successful DR test.

What should be considered?

Reviewing and retesting affected recovery capability.

Scenario 34

Does one successful DR test prove recovery capability forever?

Answer?

No. Systems, people and dependencies change.

Scenario 35

What should primarily determine how often recovery capability is tested?

Best answer?

Risk, business requirements and applicable obligations.

Scenario 36

A test uses copied production customer data in an unsecured lab.

Primary concern?

The test creates a confidentiality risk.

Scenario 37

A full-interruption exercise has no stop criteria.

Primary concern?

The organisation may lose control of test risk.

Scenario 38

Participants practise decision-making but no technology is exercised.

What should this not be mistaken for?

Proof that technical recovery works.

Scenario 39

Engineers successfully recover every server but the business owner says the service is unusable.

Was DR successful?

Not necessarily. Business-service recovery must be validated.

Scenario 40

A recovery test works but security logging is absent in the DR environment.

What was identified?

A security-control recovery gap.

Scenario 41

What is the distinction between 7.11 and 7.12?

Answer?

7.11 executes DR processes; 7.12 validates them through testing.

Scenario 42

Which test gives the greatest realism without necessarily shutting the primary system down?

Common CISSP answer?

Parallel testing.

Scenario 43

Which test relies on the actual recovery capability instead of normal production?

Answer?

Full interruption.

Scenario 44

Which test is generally safest and easiest for initial review of a DR plan?

Answer?

Read-through / tabletop.

Scenario 45

Management asks for the main objective of DR testing.

Best answer?

Demonstrate that recovery plans, personnel, communications and technical capabilities can meet business recovery requirements, identify weaknesses and drive measurable improvement.

CISSP Exam Perspective

Recognise the Clue Words

Review the Document

Lowest complexity.

Read-Through

Discuss a Scenario

Facilitated discussion.

Tabletop

Step Through Procedures

Operational review.

Walkthrough

Realistic Fake Disaster

Decision pressure.

Simulation

Recovery Runs Beside Production

Technical assurance.

Parallel

Production Actually Transfers

Highest realism.

Full Interruption

Lowest Operational Risk

Discussion-based.

Read-Through / Tabletop

Highest Operational Risk

Live dependency.

Full Interruption

Primary Keeps Running

Recovery tested separately.

Parallel

Did Backup Restore?

Technical assurance.

Restore Test

Actual Recovery Time

Compare requirement.

RTO Validation

Recovered Data Too Old

Data requirement.

RPO Failure

Hidden DNS Dependency

Architecture gap.

Dependency Finding

Recovery Site Too Small

Scale.

Capacity Failure

Expert Unavailable

Human resilience.

Alternate Personnel

Emergency Number Invalid

Official topic.

Communications Test

Test Message Mistaken as Real

Control.

Exercise Communications

Finding Has No Owner

No improvement.

Corrective Action

Fix Applied

Prove it.

Retest

Architecture Changed

Old test may be invalid.

Review / Retest DR
⚠️ Common CISSP Mistakes Testing is about evidence, not paperwork
DR Plan Exists ≠ DR Plan Works

Test it.

Tabletop ≠ Technical Recovery Test

Discussion validates decisions and procedures, not actual system operability.

Walkthrough ≠ Full Interruption

Stepping through procedures does not require real production shutdown.

Simulation ≠ Real Disaster

It creates realistic decision pressure without necessarily disrupting production.

Parallel ≠ Full Interruption

Production normally continues while recovery capability operates separately.

Full Interruption ≠ Safest Test

It provides high realism but introduces substantial operational risk.

More Realistic ≠ Automatically Better

Select the method according to the assurance objective and business risk.

Test Failure ≠ Useless Exercise

Finding a weakness before disaster is valuable.

Backup Job Success ≠ Restore Success

Test restoration.

Restore Success ≠ RTO Met

Recovery can work but take too long.

RTO Met ≠ RPO Met

Time and recoverable data point are separate requirements.

Servers Started ≠ Business Recovered

Validate end-to-end service.

Functional ≠ Secure

Logging, EDR, IAM and other controls also need recovery validation.

Alternate Site Exists ≠ Capacity Proven

Test expected workload.

Dependency Documented ≠ Dependency Tested

Include important supporting services.

Expert Can Recover ≠ Team Can Recover

Test alternate personnel.

Contact List Exists ≠ Contacts Work

Validate communication paths.

Supplier Contract Exists ≠ Supplier Activation Tested

External dependencies require assurance too.

Exercise Message ≠ Ordinary Message

Control communications so simulated events are not mistaken for real ones.

Finding Recorded ≠ Finding Fixed

Corrective actions require ownership.

Finding Fixed ≠ Finding Verified

Retest where appropriate.

Last Year's Successful Test ≠ Today's Recovery Assurance

Architecture, people and suppliers change.

Testing ≠ Training

Training develops people. Testing evaluates capability.

7.11 ≠ 7.12

7.11 performs DR. 7.12 tests DR.

Quick Reference

If you see...Think...
Review documentRead-Through
Discuss hypothetical disasterTabletop
Step through proceduresWalkthrough
Realistic disaster scenarioSimulation
Recovery environment runs beside productionParallel
Production deliberately depends on DRFull Interruption
Can backup actually recover?Restore Test
Did recovery meet required time?RTO Validation
Did recovered data meet required point?RPO Validation
Can alternate environment handle workload?Capacity Test
Can another person recover?Personnel Resilience
Can supplier be reached?Communications Test
Finding discoveredCorrective Action
Correction implementedRetest
Major production changeReview DR Capability

Test Types Memory Aid

READ Review the plan
TALK Tabletop discussion
WALK Step through process
SIMULATE Act through scenario
PARALLEL Recover beside production
INTERRUPT Rely on recovery

Realism Memory Aid

TABLETOP Talk
WALKTHROUGH Step
SIMULATION React
PARALLEL Recover
FULL Depend on DR

More realism usually means more assurance - and more operational risk.

7.12 Master Memory Aid

DEFINE Objective + success
PLAN Scope + safety
EXERCISE Test recovery
OBSERVE Collect evidence
MEASURE RTO + RPO + capability
CORRECT Fix weaknesses
RETEST Prove improvement

Exercise → Measure → Improve → Retest

The DR Test Leader's Questions

WHY? What are we proving?
SCOPE? What is included?
METHOD? Which test type?
RISK? Could testing affect production?
STOP? When do we terminate?
PEOPLE? Who participates?
COMMUNICATION? Who must know?
RTO? How quickly did we recover?
RPO? How current was the data?
CAPACITY? Can DR handle demand?
SECURE? Did security controls recover?
FINDINGS? What did we learn?
OWNER? Who fixes each problem?
RETEST? Can we prove the fix?

Key Takeaways

CISSP 7.12 is: Test Disaster Recovery Plans.

The current ISC2 outline explicitly covers: read-through/tabletop, walkthrough, simulation, parallel, full interruption and communications.

DR testing determines whether recovery plans, personnel, communications and technology actually support the organisation's recovery requirements.

A DR plan existing does not prove a DR capability exists.

Testing can validate documentation, personnel, technical systems, dependencies, suppliers and recovery objectives.

Testing should begin with clearly defined objectives.

The organisation should know what the exercise is expected to prove.

Success criteria should be established before the exercise.

Measures may include RTO, RPO, functionality, security, capacity and communications.

Objective → exercise → measure → improve → retest.

A read-through reviews the plan for completeness, accuracy and usability.

It is low risk and can identify documentation problems before more complex exercises occur.

A tabletop exercise uses a hypothetical disaster scenario.

Participants discuss how they would respond and make decisions.

Tabletop exercises are particularly useful for testing roles, responsibilities, escalation, communication and decision-making.

Tabletop = talk it through.

A successful tabletop does not prove that backups or alternate systems actually work.

A walkthrough steps through the recovery process in greater operational detail.

Participants may verify facilities, equipment, access, procedures and contacts.

Tabletop = discuss it. Walkthrough = step through it.

A simulation places participants into a realistic disaster scenario and requires them to respond as if the event were real.

Simulations can introduce changing conditions and unexpected events to test decision-making.

Production usually remains protected from deliberate interruption.

A parallel test activates and tests the recovery environment while normal production continues operating.

It provides stronger technical evidence without intentionally making production dependent on the recovery site.

Parallel = DR runs beside production.

Parallel testing may still leave some activities untested, such as real user redirection or complete production cutover.

A full-interruption test deliberately makes actual operations depend on the recovery capability.

This creates the highest realism but also the greatest operational risk among the test types listed by ISC2.

Full interruption = use DR for real.

Full-interruption testing requires careful risk assessment, authorisation, monitoring, rollback and stop criteria.

The test should not accidentally become the disaster.

More realistic testing can provide greater assurance, but more realism is not automatically appropriate.

The organisation should select the test method that provides sufficient assurance while managing business risk.

Restore testing is especially important.

Backup software reporting success does not prove that data can actually be recovered.

Backup success ≠ restore success.

DR testing should measure actual recovery time against the documented RTO.

A service with a two-hour RTO that takes six hours to recover has a measurable recovery gap.

DR testing should also validate the RPO.

A technically successful restore may still fail requirements if recovered data is older than the business can tolerate.

RTO tests time. RPO tests recoverable data.

Dependencies should be included in DR exercises.

Recovering an application does not help if DNS, identity, databases, networking or required suppliers remain unavailable.

Testing is one of the best methods for finding undocumented dependencies.

Recovery capacity should also be tested.

An alternate environment that supports a handful of test users may not be able to support production demand.

DR testing should include people as well as technology.

Organisations should consider whether alternate personnel can perform important procedures if the primary specialist is unavailable.

Expert knows how ≠ organisation knows how.

Communications are an explicit part of current CISSP 7.12.

Exercises should consider stakeholders, test status and any applicable external communications.

Alternate communication channels should be tested where normal communications could fail during the scenario.

Exercise messages should be controlled so simulated events are not accidentally interpreted as genuine incidents.

Critical third-party dependencies should also be incorporated into recovery assurance where practical.

A contract saying that a supplier can recover within a particular time is weaker evidence than a successfully tested activation process.

DR exercises should have observers or evaluators where appropriate.

Recovery times, decisions, communication delays, errors and unexpected dependencies should be recorded.

Testing should produce evidence rather than relying solely on participant memory.

A failed recovery test can be valuable.

Discovering that a backup cannot be restored during a controlled exercise is far better than discovering it during a real disaster.

Finding weakness during testing is not failure of the testing process. It is one of its purposes.

Exercises should be followed by evaluation.

Findings should identify what failed, why it failed and what needs to change.

Corrective actions should have owners and target completion.

Significant corrections should then be retested.

Finding → cause → owner → correction → retest.

DR documentation should be updated when exercises reveal inaccurate information.

Major changes to production architecture can also invalidate previously tested recovery procedures.

A successful DR test from last year does not automatically prove today's environment remains recoverable.

Test frequency should reflect organisational risk, business criticality, applicable requirements, technological change and prior findings.

There is no single universal test interval appropriate for every system and organisation.

Test environments and recovery copies must also remain secure.

Recovery testing should not expose sensitive production data or create uncontrolled access.

Training, exercises and technical testing should not be confused.

Training develops the person's capability.

Exercises evaluate coordination and use of plans.

Technical tests validate the operability of systems and recovery mechanisms.

NIST refers to these collectively as Test, Training and Exercise activities.

The three surrounding CISSP objectives can be remembered as:

7.10 = build recovery capability. 7.11 = execute recovery. 7.12 = prove recovery works.

The central CISSP principle is:

test Disaster Recovery at a level appropriate to the organisation's risk, measure actual performance against business recovery requirements, treat weaknesses as useful findings, assign corrective actions and retest until the organisation has evidence that recovery capability works.

📚 Sources & Further Reading Current Disaster Recovery testing and exercise references