7.10 Recovery Strategies

CISSP Domain 7 ยท Security Operations

7.10 Recovery Strategies

Recovery is not something an organisation should design after systems have already failed.

Recovery strategies determine in advance how technology, data and critical services can continue or be restored when normal operations are disrupted.

Different systems need different strategies.

A payroll archive might tolerate several hours of downtime. A payment platform may require almost continuous service. A ransomware event may require recovery from an older trusted backup rather than simply failing over to another live copy.

The correct recovery strategy therefore depends on: business requirements, acceptable downtime, acceptable data loss, system dependencies, threat scenarios and cost.

๐Ÿ’พ

Protect Data

Maintain recoverable copies of important information and system components.

WHAT CAN WE RESTORE?
๐Ÿข

Recover Services

Provide alternate facilities, systems and processing capability.

WHERE CAN WE RUN?
โ™ป๏ธ

Remain Resilient

Design systems that continue operating despite component failures and disruptions.

HOW DO WE KEEP GOING?
Current CISSP 7.10 Scope

Implement Recovery Strategies

The current CISSP Exam Outline explicitly identifies four areas.

Backup Storage Strategies

Including examples such as:

Cloud Storage Onsite Offsite
Recovery Site Strategies

Including examples such as:

Cold Sites Hot Sites Resource Capacity Agreements
Multiple Processing Sites

Distribute processing capability across more than one physical or logical location.

Resilience & Availability
System Resilience High Availability Quality of Service Fault Tolerance

7.10 Official Topics

BACKUP Protect the data
SITE Provide somewhere to recover
MULTI-SITE Distribute processing
RESILIENCE Keep operating

The Big Idea

Business Requirement โ†’ BIA
BIA โ†’ RTO + RPO + MTD
Recovery Requirements โ†’ Recovery Strategy
Recovery Strategy โ†’ Technology + Sites + Backups + People
Implementation โ†’ Testing
Testing โ†’ Proven Recoverability
Recovery technology should be selected from business recovery requirements - not the other way around.
Essential CISSP Concepts

RTO ยท RPO ยท MTD

RTO - Recovery Time Objective

The target for how quickly a system or service must be restored after disruption.

Think:

HOW LONG CAN RECOVERY TAKE?

RPO - Recovery Point Objective

The point in time to which data must be recovered.

It represents the amount of recent data the organisation can tolerate losing.

Think:

HOW MUCH DATA CAN WE LOSE?

MTD - Maximum Tolerable Downtime

The maximum period that the business process can be disrupted before the disruption causes unacceptable or significant harm.

Think:

WHEN DOES THE BUSINESS DAMAGE BECOME UNACCEPTABLE?

RTO vs RPO vs MTD

RTO Time to restore
RPO Data point to restore
MTD Maximum tolerable disruption

Visualising RTO & RPO

09:00 โ†’ Last Recoverable Data
10:00 โ†’ Failure Occurs
12:00 โ†’ Service Restored
RPO Example

Failure: 10:00

Recoverable point: 09:00

Potential data loss: 1 hour.

RTO Example

Failure: 10:00

Required restoration: 12:00

Recovery objective: 2 hours.

RPO looks backward from the disruption. RTO looks forward toward restoration.
โฑ๏ธ RTO Must Fit Inside Business Tolerance The recovery objective must support the business requirement
Business

Maximum tolerable disruption: 8 hours.

IT recovery strategy

Expected recovery: 14 hours.

The technical recovery capability does not satisfy the business requirement.

Recovery strategy must be capable of meeting the required recovery objectives.
Risk & Economics

Recovery Speed Has a Cost

Longer RTO โ†’ Usually Simpler Recovery Capability
Short RTO โ†’ More Ready Capacity
Near-Zero RTO โ†’ Active Resilient Architecture
The business should not pay for instantaneous recovery where a long outage is genuinely acceptable - nor choose cheap recovery where minutes of downtime would be catastrophic.
Official 7.10 Topic 1

Backup Storage Strategies

A backup is a recoverable copy of information or system components maintained so they can be restored if the original is lost, corrupted or unavailable.

Onsite Backup

Backup stored at or very near the primary operating location.

Advantage: fast local access.

Risk: same physical disaster may affect both.

Offsite Backup

Backup stored at a separate geographic location.

Advantage: reduces common physical-site risk.

Cloud Backup

Backup data stored using cloud infrastructure or backup services.

Consider:

Connectivity Identity Encryption Provider Resilience Recovery Throughput

Full ยท Incremental ยท Differential

Backup TypeWhat Is Copied?Backup Speed / SizeTypical Restore Requirement
FullAll selected dataLargest / longestFull backup
IncrementalChanges since the previous backupUsually smallest / fastestFull + subsequent incrementals
DifferentialChanges since the last full backupGrows until next fullFull + latest differential

Backup Types

FULL Everything
INCREMENTAL Since last backup
DIFFERENTIAL Since last full
๐Ÿ“… Backup Example Follow what each backup contains
DayChangeIncrementalDifferential
SundayFull backupFullFull
MondayFile A changesAA
TuesdayFile B changesBA + B
WednesdayFile C changesCA + B + C
Incremental resets its reference after each backup. Differential continues accumulating changes since the full backup.
Backup Trade-Off

Backup Speed vs Restore Complexity

Incremental

Advantages:

Smaller Backups Shorter Backup Window

Restore may require: multiple backup sets.

Differential

Backup grows larger throughout the cycle.

Restore commonly requires: full backup + latest differential.

Full

Simplest recovery chain.

But: requires more storage and backup time.

Backup design should consider recovery requirements, not merely how quickly backups complete.

Snapshot vs Backup

Snapshot

Captures system or storage state at a point in time.

Often useful for:

Fast Rollback Short-Term Recovery Testing
Independent Backup

Maintains a recoverable copy that can be separated from the original environment or failure domain.

Often stronger for: major disaster or destructive attack recovery.

Snapshot โ‰  automatically independent backup

A snapshot stored on the same compromised or failed storage platform may disappear with the original environment.

Critical Distinction

Replication โ‰  Backup

Primary Database โ†’ Replica Database
Good Change โ†’ Replicated
Malicious Deletion โ†’ Also Replicated
Replication

Keeps another copy relatively current.

Excellent for: availability and fast failover.

Backup

Preserves recoverable historical state.

Useful when: current state itself is corrupted or malicious.

Replication protects availability. Historical backup helps recover from bad state.
๐Ÿ”’ Ransomware Changes the Backup Question Can attackers destroy the recovery copy too?
Environment

Production administrators can:

delete every backup.

Attacker compromises: production administrator account.

Result:

production + backups destroyed.

Useful protections can include:

Immutable Backups Offline Copies Separate Credentials Separate Administrative Plane Strong MFA Deletion Protection Geographic Separation
A backup that the attacker can destroy with production credentials may not provide meaningful ransomware recovery.

Immutable & Offline Backups

Immutable Backup

Backup cannot be modified or deleted during a defined retention period according to the storage controls used.

Useful against: malicious or accidental alteration.

Offline / Air-Gapped Copy

Recovery copy is not continuously accessible from the production environment.

Useful against: attacks spreading through connected systems.

Separate the recovery capability from the failure or compromise you are trying to survive.
3๏ธโƒฃ Common 3-2-1 Backup Heuristic A useful operational memory aid rather than an ISC2 7.10 sub-bullet

3-2-1

3 Copies of important data
2 Different storage types or failure characteristics
1 Copy separated from the primary location
Modern environments may go further

Immutability, offline copies and separate administrative credentials can provide additional protection against destructive attacks.

Backup Security

Backups Contain Valuable Data Too

Confidentiality

Encrypt sensitive backup data where appropriate.

Integrity

Protect backup contents against unauthorised modification.

Access Control

Restrict who can read, restore or delete backups.

Retention

Maintain recovery points according to business and regulatory needs.

Monitoring

Detect unexpected deletion or modification of recovery copies.

Key Management

Ensure encryption keys needed for recovery remain available and protected.

Perfectly encrypted backup + lost encryption key = unavailable backup.
Most Important Backup Test

Can You Restore It?

Dashboard

Backup jobs: 100% successful.

Disaster

Restore attempted.

Result: backup archive corrupt.

Backup Assurance

CREATE Backup
PROTECT Backup
VERIFY Integrity
RESTORE Test it
Backup success โ‰  recovery success.
Official 7.10 Topic 2

Recovery Site Strategies

A recovery site provides an alternate location from which technology or business operations can be restored after the primary environment becomes unavailable.

Cold Site

Facility with basic environmental and physical capabilities but without the full computing environment ready for immediate operation.

CHEAPER ยท SLOWER RECOVERY

Warm Site

Partially equipped alternate facility with some systems and telecommunications capability already available.

MIDDLE GROUND

Hot Site

Highly prepared alternate processing environment with required hardware and software available for recovery.

MORE EXPENSIVE ยท FASTER RECOVERY

Recovery Sites

COLD Facility ready
WARM Partially ready
HOT Technology ready

Cold vs Warm vs Hot

SiteReadinessRecovery SpeedRelative Cost
ColdBasic facilitySlowestLowest
WarmPartially equippedIntermediateIntermediate
HotHighly equippedFastest of the threeHighest
Generally: more readiness โ†’ faster recovery โ†’ greater cost.
๐Ÿ”ฅ Hot Site โ‰  Automatically Zero Downtime The site may be ready while data or applications still require activation

A hot site may have hardware, software and connectivity available.

Recovery can still require:

Data Synchronisation Application Start-Up DNS Changes Routing Changes Validation User Redirection
Site readiness and data currency are separate recovery questions.
Capacity Agreements

Reserved & Reciprocal Capacity

Recovery capability does not always require the organisation to own an entire idle recovery facility.

Provider Capacity Agreement

A service provider agrees to make defined recovery resources or capacity available when needed.

Consider:

Guaranteed Capacity Activation Time Priority Testing Rights Geography
Reciprocal Agreement

Two organisations agree to support one another's recovery capability when disruption occurs.

Advantages: potentially lower cost.

Challenges: capacity, compatibility and simultaneous disaster.

๐Ÿค Reciprocal Agreements Need Reality Checks Good intentions are not recovery capacity
Agreement

Company A and Company B agree to: host one another during disaster.

Five years later:

Different Cloud Platforms Different Applications No Spare Capacity Agreement Never Tested
Contractual recovery capability should be technically feasible and tested.

Geographic Separation

Alternate sites should reduce exposure to the same disruption affecting the primary site.

Poor geographic diversity

Primary data centre: Building A.

Recovery site: Building B next door.

Shared:

Electricity Substation Flood Zone Telecommunications Route
Different address โ‰  different failure domain.
Official 7.10 Topic 3

Multiple Processing Sites

Rather than operating from one primary environment and activating another only after disaster, organisations can distribute processing capability across multiple sites.

Active / Passive

One environment normally serves production while another remains ready to take over.

PRIMARY + STANDBY

Active / Active

Multiple sites actively process production workload.

MULTIPLE LIVE SITES

Multi-Site

ACTIVE / PASSIVE One runs ยท one waits
ACTIVE / ACTIVE Both run
๐ŸŸข Active / Passive Standby capacity is activated when the primary fails
Site A โ†’ Processes Production
Site B โ†’ Standby
Site A Fails โ†’ Failover
Site B โ†’ Processes Production
Standby does not guarantee readiness. It must still be synchronised, maintained and tested.
๐ŸŸข Active / Active Production workload is distributed across multiple environments
Users โ†’ Traffic Distribution
Traffic โ†’ Site A + Site B
Site A Fails โ†’ Site B Continues
Higher resilience can mean greater complexity

Active-active architecture may require careful coordination of data, application state, traffic routing and capacity.

Can the Remaining Site Handle the Load?

Normal operation

Site A: 50% workload.

Site B: 50% workload.

After Site A failure

Site B must potentially support: 100% workload.

If each site has capacity for only 60% of total demand, failover may work technically while service performance becomes unacceptable.

Redundant site โ‰  sufficient recovery capacity.
Modern Architecture

Availability Zones & Regions

Cloud platforms can provide logically or geographically separated deployment locations that support resilient architectures.

Multiple Zones

Protect against some local infrastructure or facility failures.

Multiple Regions

Can provide broader geographic separation for larger-scale disruptions.

Multi-zone โ‰  every disaster solved

Common dependencies such as identity, DNS, configuration, cloud account, deployment pipelines or application defects may still affect all locations.

โš ๏ธ Common-Mode Failure Redundant systems can still depend on the same thing
Architecture

Site A: redundant.

Site B: redundant.

Both use:

Same Identity Provider Same DNS Service Same Admin Account Same Software Release Same Network Carrier

Failure of the shared dependency can still: disable both sites.

Redundancy is useful only when the redundant components do not share the same critical failure path.
Data Across Sites

Synchronous vs Asynchronous Replication

Synchronous Replication

A write is committed across participating locations before the transaction is treated as complete.

Advantage: very small potential data gap.

Trade-off: latency and distance can affect performance.

Asynchronous Replication

Primary operation completes before the secondary copy necessarily receives every change.

Advantage: supports greater distance and lower latency impact.

Trade-off: recent transactions may be lost during sudden failover.

Replication

SYNC Lower data loss ยท higher coordination
ASYNC More distance ยท possible data lag
๐Ÿง  Split-Brain Risk What if both sides believe they are primary?
Cluster

Site A loses communication with Site B.

Site A believes: Site B failed.

Site B believes: Site A failed.

Both begin accepting: independent writes.

Result: data inconsistency or corruption.

Distributed resilience requires mechanisms for coordination, quorum and authoritative ownership.
Official 7.10 Topic 4

System Resilience

Resilience is broader than simply restoring a failed system.

A resilient system is designed to:

Disruption โ†’ Absorb
Failure โ†’ Continue Essential Operations
Reduced Capability โ†’ Adapt
Recovery โ†’ Return to Effective Operation

Resilience

WITHSTAND Failure
CONTINUE Essential service
ADAPT Degraded conditions
RECOVER Normal capability
โ™ป๏ธ Resilience vs Recovery One tries to keep capability available while the other restores it
Resilience

Continue important capability despite adverse conditions where possible.

KEEP GOING

Recovery

Restore capability after disruption.

BRING IT BACK

Good architecture may use both: tolerate what it can and recover what it cannot.
High Availability

High Availability - HA

High availability uses redundancy and failover mechanisms to reduce service interruption when components fail.

Component A โ†’ Serving Traffic
Component A Fails โ†’ Detect Failure
Component B โ†’ Take Over
Clustering Load Balancing Redundant Links Redundant Power Multiple Servers Automatic Failover

High Availability

REDUNDANCY Another component exists
DETECTION Failure recognised
FAILOVER Replacement takes over

High Availability vs Disaster Recovery

High Availability

Primarily aims to minimise interruption from component or local service failure.

KEEP THE SERVICE RUNNING

Disaster Recovery

Restores capability after a significant disruptive event has affected normal operations.

RECOVER AFTER MAJOR DISRUPTION

HA โ‰  DR

A highly available system located entirely within one data centre may still fail when the entire facility becomes unavailable.

โš–๏ธ Load Balancing Distribute demand across multiple service instances
Users โ†’ Load Balancer
Load Balancer โ†’ Server A
Load Balancer โ†’ Server B
Server A Fails โ†’ Traffic to Server B
Load balancing can support both performance and availability when multiple healthy service instances exist.
๐Ÿ”— Clustering Multiple systems cooperate to provide a service

Cluster designs can provide:

Redundancy Failover Shared Workload Maintenance Flexibility
Cluster โ‰  automatically resilient if every node shares the same single point of failure.
Fault Tolerance

Continue Despite Failure

Fault tolerance is the ability of a system to continue correct operation despite failure of one or more components within its designed tolerance.

Redundant Power Supplies

One power supply can fail while another continues supplying power.

Redundant Network Links

Alternative communication path remains available.

Disk Redundancy

Some storage designs continue operating when a disk fails.

Redundant Components

Multiple components remove reliance on a single device.

Fault Tolerance

FAULT OCCURS Component fails
SYSTEM CONTINUES Service remains correct
๐Ÿ” High Availability vs Fault Tolerance Closely related but not identical ideas
High Availability

Designed to minimise downtime.

A brief failover interruption may still occur.

Fault Tolerance

Designed so the system continues correct operation despite a component fault within its tolerance.

Greater emphasis on: continuity through failure.

HA = minimise interruption. Fault tolerance = continue despite fault.

Disk Redundancy

RAID and similar storage technologies can provide resilience against some disk failures.

Disk failure

RAID configuration may allow: service to continue while failed storage is replaced.

RAID โ‰  Backup

RAID may protect against disk failure.

It does not inherently protect against:

File Deletion Ransomware Database Corruption Site Destruction
Quality of Service

QoS - Protect Important Traffic

Quality of Service provides mechanisms for controlling or assuring characteristics of network service such as bandwidth, latency, priority, packet loss or jitter.

Disruption

Network capacity reduced: 50%.

Competing traffic:

Critical Voice Calls Payment Transactions Software Downloads Video Streaming

QoS can prioritise: the services most important to continued operations.

QoS

CAPACITY LIMITED Resources constrained
PRIORITISE Important traffic
MAINTAIN Required service quality
๐Ÿ“ก QoS โ‰  Redundancy QoS manages available capacity - it does not create another network
Network

WAN link: completely unavailable.

QoS cannot prioritise traffic across: a link that no longer exists.

QoS manages service quality. Redundancy provides alternative capability.

Graceful Degradation

Resilient systems do not always need to provide every feature during disruption.

They may deliberately preserve: essential functionality while reducing less important capability.

Banking service under capacity pressure

Preserve:

Payments Balance Access Fraud Controls

Temporarily reduce:

Historical Analytics Marketing Content Non-Critical Reports
Resilience can mean operating in a reduced but still useful state.
Remove Single Points of Failure

Redundancy

Power

UPS, generators, multiple power feeds.

Network

Multiple links, routers, switches and carriers.

Compute

Multiple servers or cluster nodes.

Storage

Replication and redundant storage components.

Sites

Multiple processing locations.

People

Avoid dependency on one individual with unique operational knowledge.

Find the single component whose failure stops everything. Then decide whether the business can tolerate that risk.
Architecture Scenario

The Redundant Data Centre With One Router

Application has:

8 Web Servers 4 Database Servers Redundant Storage Dual Power

But every external connection uses:

one network router.

Router fails.

Result: entire service unavailable.

Resilience is limited by the remaining single point of failure.
Critical Concept

Recover the Dependency Chain

Applications rarely operate alone.

Customer Application โ†’ Identity Provider
Identity Provider โ†’ Directory
Application โ†’ Database
Everything โ†’ DNS + Network + Encryption Keys
Failure

Application recovered: successfully.

DNS unavailable: users still cannot access it.

Recovering one component does not recover the business service if its dependencies remain unavailable.
๐Ÿ”ข Recovery Order Matters Dependencies may need restoration before the application
1 โ†’ Core Infrastructure
2 โ†’ Identity / Network / DNS
3 โ†’ Database / Middleware
4 โ†’ Application
5 โ†’ Business Validation
Recovery priority and recovery sequence are not always the same thing.

Recovery Systems Still Need Security

Organisations sometimes weaken controls during disaster recovery because restoring availability becomes urgent.

Authentication Logging Network Segmentation Encryption Privileged Access Monitoring
Bad recovery

Production application restored quickly.

Temporary recovery environment uses:

Shared Admin Password No SIEM Logging Flat Network
Recovery should restore an acceptable security posture, not merely turn the service back on.
Cyber Recovery

Known-Good Recovery

Recovery from cyberattack creates a problem that traditional hardware failure does not always create:

the newest copy may already be compromised.

Ransomware detected

Monday: files encrypted.

Investigation finds attacker entered: three weeks earlier.

Restoring Sunday's backup may restore: attacker persistence.

Most recent recovery point โ‰  safest recovery point.

Before Cyber Recovery

CLEAN? Is the recovery point trustworthy?
PATCHED? Has the exploited weakness been addressed?
CREDENTIALS? Have compromised secrets been rotated?
PERSISTENCE? Has attacker access been removed?
MONITORED? Can recurrence be detected?

Recovery Does Not Always Mean Technology

Some business processes can continue temporarily using alternative or manual methods while technology is restored.

System unavailable

Automated approval platform cannot operate.

Temporary process: documented manual approval procedure.

Recovery strategy should preserve the business capability - not necessarily recreate the normal technology immediately.
Third Parties

Recovery Depends on Suppliers Too

Cloud Provider

What resilience and recovery capability is actually provided?

Network Carrier

Are alternate communication paths available?

SaaS Provider

What are its recovery commitments and data-backup capabilities?

Recovery Facility Provider

Is capacity guaranteed during a widespread disruption?

Your recovery objective cannot be shorter than a critical supplier's capability unless you have another strategy.
๐Ÿ“„ SLAs & Recovery Commitments Contractual wording should support the required business outcome

Recovery-related agreements can address:

Service Availability Recovery Time Data Recovery Capacity Notification Testing Support
Mismatch

Business RTO: 2 hours.

Critical provider commitment: recovery within 24 hours.

The organisation's recovery plan must account for the real capability of its dependencies.
Assurance

A Recovery Strategy Must Be Tested

A recovery architecture existing on paper does not prove that it can achieve the required RTO and RPO.

Backup Restore

Can data actually be restored?

Failover

Does the secondary environment take over?

Capacity

Can the recovery environment handle required workload?

Dependencies

Are DNS, identity, networking and integrations recoverable?

Security

Are security controls still functioning after recovery?

Timing

Did recovery actually meet the RTO?

Designed recovery capability โ‰  demonstrated recovery capability.
Critical Domain 7 Distinction

7.10 Recovery Strategy vs 7.11 Disaster Recovery Process

7.10 Recovery Strategies

Determine the capabilities available for recovery.

Backups Recovery Sites Replication HA Fault Tolerance

HOW WILL RECOVERY BE POSSIBLE?

7.11 Disaster Recovery Processes

Covers what the organisation does when executing disaster recovery.

HOW DO WE PERFORM THE RECOVERY?

7.10 vs 7.11

7.10 Recovery capability
7.11 Recovery execution
๐Ÿข Business Continuity vs Disaster Recovery Keep the business operating vs restore disrupted technology
Business Continuity

Broader capability to continue important business processes during disruption.

KEEP BUSINESS OPERATING

Disaster Recovery

Focuses more directly on restoring technology and information capabilities after major disruption.

RESTORE TECHNOLOGY

Practical Scenario

Payment Platform

Business requirements:

RTO: 15 Minutes RPO: Near-Zero 24/7 Service

Strategy:

Site A โ†’ Active
Site B โ†’ Active
Data โ†’ Highly Current Replication
Traffic โ†’ Load Balanced
Backups โ†’ Separate Historical Recovery
Very short recovery objectives generally require capabilities already operating before the disaster occurs.
Cost Scenario

Monthly Reporting System

Business requirements:

RTO: 48 Hours RPO: 24 Hours Low Operational Criticality

A continuously active second data centre might provide: far more recovery capability than the business requires.

Recovery strategy should be economically proportionate to business requirements.
Site Scenario

The Two Data Centres in One Flood Zone

Primary site and recovery site are: five kilometres apart.

Both depend on: the same river flood defence.

Geographic distance alone does not prove geographic risk separation.
Cloud Scenario

Multi-Region But One Compromised Account

Application is deployed across: two cloud regions.

Both environments are administered through: one privileged cloud account.

Attacker compromises the account and deletes: resources in both regions.

Geographic redundancy does not protect against every logical or administrative failure mode.
Ransomware Scenario

The Perfectly Replicated Ransomware

Primary file system: encrypted by ransomware.

Replication: working perfectly.

Result: encrypted files rapidly copied to secondary site.

Availability replication can faithfully replicate destructive changes. Historical recovery copies are still needed.
Cryptography Scenario

The Backup Nobody Can Decrypt

Organisation maintains: excellent encrypted offsite backups.

Disaster destroys: the only key-management server containing the decryption key.

Recovery planning must include the dependencies required to recover the backup itself.
Recovery Site Scenario

The Hot Site With No Capacity

Recovery site has:

Servers Network Applications Current Data

But it can support: 30% of normal user demand.

A technically functional recovery site may still fail the business if capacity is insufficient.
Network Scenario

Reduced Network Capacity

One WAN circuit fails.

Remaining circuit can support: 60% of normal traffic.

QoS prioritises:

Critical Transactions Voice Incident Communications

Lower-priority bulk transfer is: deprioritised.

QoS can help maintain important services in a degraded environment.
Testing Scenario

The Recovery Site Nobody Maintained

Recovery site was built: three years ago.

Production has since received:

12 Application Releases Database Upgrade New IAM Integration Network Architecture Change

Recovery environment: was never updated.

Recovery capability must evolve as production changes.
๐ŸŽ“ CISSP Scenarios Recognise the recovery strategy being tested
Scenario 1

A business process can tolerate two hours before the system must be available again.

Which objective?

RTO.

Scenario 2

The business can tolerate losing at most 30 minutes of transactions.

Which objective?

RPO.

Scenario 3

A business process cannot remain disrupted beyond eight hours without significant harm.

Which concept?

MTD.

Scenario 4

RTO is 12 hours but MTD is 6 hours.

Problem?

The recovery objective does not satisfy business tolerance.

Scenario 5

A backup is stored in the same data centre as the production system.

Primary risk?

Both may be lost in the same site disaster.

Scenario 6

A backup copy is maintained at a geographically separate location.

Which strategy?

Offsite backup.

Scenario 7

All selected data is copied during every backup.

Which backup?

Full.

Scenario 8

Only changes since the previous backup are copied.

Which backup?

Incremental.

Scenario 9

All changes since the last full backup are copied.

Which backup?

Differential.

Scenario 10

Which generally requires a full plus multiple subsequent backup sets during restoration?

Answer?

Incremental backup strategy.

Scenario 11

A storage snapshot resides on the same failed array as the production volume.

Primary lesson?

A snapshot is not automatically an independent backup.

Scenario 12

Deletion from the primary database immediately appears on the secondary replica.

Which lesson?

Replication is not the same as historical backup.

Scenario 13

Ransomware encrypts production and the encryption is copied to the replica.

What additional capability is important?

Protected historical backups.

Scenario 14

Attackers cannot modify or delete a recovery copy during its retention period.

Which concept?

Immutable backup.

Scenario 15

A recovery copy is disconnected from normal production access.

Which concept?

Offline / air-gapped backup.

Scenario 16

Backup dashboard says every job succeeded but restores have never been tested.

Primary concern?

Recoverability is unproven.

Scenario 17

Encrypted backups survive the disaster but decryption keys do not.

Result?

The backups may be unavailable for recovery.

Scenario 18

An alternate facility has power and environmental controls but no computing hardware installed.

Which site?

Cold site.

Scenario 19

An alternate facility is partially equipped with computing and telecommunications capability.

Which site?

Warm site.

Scenario 20

An alternate facility has hardware and software available for rapid recovery.

Which site?

Hot site.

Scenario 21

Which site generally costs the least but takes longest to prepare?

Answer?

Cold site.

Scenario 22

Does having a hot site automatically mean zero data loss?

Answer?

No. Site readiness and data replication are separate issues.

Scenario 23

Two organisations agree to support each other's processing requirements during disruption.

Which arrangement?

Reciprocal agreement.

Scenario 24

A recovery provider promises resources but cannot guarantee enough servers during a regional disaster.

Primary concern?

Recovery capacity.

Scenario 25

The primary and recovery sites use the same power substation.

Primary concern?

Common-mode failure.

Scenario 26

One processing site runs production while the second waits to take over.

Which model?

Active-passive.

Scenario 27

Two geographically separate sites simultaneously process customer traffic.

Which model?

Active-active.

Scenario 28

Each of two active sites handles 50% of traffic but can support only 60% alone.

What is the concern?

Insufficient failover capacity.

Scenario 29

Two regions rely on the same compromised privileged account.

What does this demonstrate?

Geographic redundancy can still share a logical failure domain.

Scenario 30

A transaction is not committed until the secondary location confirms the write.

Which replication?

Synchronous.

Scenario 31

The secondary site may lag slightly behind the primary.

Which replication?

Asynchronous.

Scenario 32

Which replication approach can create potential loss of the most recent writes during sudden failure?

Answer?

Asynchronous replication.

Scenario 33

Two isolated cluster nodes both believe they are primary and accept writes.

Which problem?

Split brain.

Scenario 34

A server fails and another automatically assumes its workload after a short interruption.

Which capability?

High availability / failover.

Scenario 35

A system continues correct service without interruption when one redundant component fails.

Which concept?

Fault tolerance.

Scenario 36

A degraded network prioritises voice and critical transaction traffic.

Which concept?

Quality of Service.

Scenario 37

The only WAN circuit completely fails.

Can QoS restore the missing circuit?

No. Redundancy is required for an alternate path.

Scenario 38

A RAID array survives one disk failure.

Does this mean the organisation no longer needs backups?

No.

Scenario 39

A service continues operating with reduced non-critical functionality during disruption.

Which principle?

Graceful degradation / resilience.

Scenario 40

The application server is restored before DNS and identity services.

Primary problem?

Recovery dependency sequence was not considered.

Scenario 41

The newest backup contains attacker persistence.

Should it automatically be restored?

No. Select a sufficiently trusted recovery point.

Scenario 42

A SaaS provider has a 24-hour recovery commitment while the business needs the service within one hour.

Primary problem?

Supplier recovery capability does not support organisational RTO.

Scenario 43

An alternate site exists but has not been updated after years of production changes.

Primary risk?

Recovery environment may no longer be compatible or usable.

Scenario 44

Management asks what 7.10 primarily determines.

Best answer?

The capabilities and architecture used to make recovery possible.

Scenario 45

Management asks for the central principle of recovery strategy.

Best answer?

Select and maintain backup, alternate processing and resilience capabilities that satisfy business recovery objectives and have been validated through testing.

CISSP Exam Perspective

Recognise the Clue Words

How Long to Restore?

Recovery time.

RTO

How Much Data Loss?

Recovery point.

RPO

Maximum Business Disruption

Absolute tolerance.

MTD

Same Building

Fast but shared risk.

Onsite Backup

Separate Geography

Site resilience.

Offsite Backup

Everything Copied

Backup type.

Full

Since Previous Backup

Backup type.

Incremental

Since Last Full

Backup type.

Differential

Point-in-Time State

Fast recovery.

Snapshot

Live Copy

Availability.

Replication

Cannot Be Modified

Recovery protection.

Immutable Backup

Disconnected Copy

Isolation.

Offline Backup

Building Only

Slow recovery.

Cold Site

Partially Equipped

Middle recovery.

Warm Site

Highly Ready

Fast recovery.

Hot Site

Organisations Support Each Other

Capacity arrangement.

Reciprocal Agreement

One Live ยท One Standby

Processing sites.

Active / Passive

Both Live

Processing sites.

Active / Active

Writes Confirmed Both Places

Data currency.

Synchronous Replication

Replica Can Lag

Data gap.

Asynchronous Replication

Both Think They Are Primary

Distributed data problem.

Split Brain

Minimise Downtime

Availability.

HA

Continue Despite Fault

Continuity.

Fault Tolerance

Prioritise Network Traffic

Service quality.

QoS

Reduced but Useful Service

Adaptation.

Graceful Degradation

Another Component Exists

Resilience.

Redundancy

One Component Stops Everything

Architecture risk.

Single Point of Failure

Restore One System But Service Fails

Architecture.

Dependency Problem

Newest Copy May Be Infected

Cyber recovery.

Known-Good Recovery Point

Does It Actually Restore?

Assurance.

Recovery Testing
โš ๏ธ Common CISSP Mistakes Recovery capability must match the business requirement
RTO โ‰  RPO

RTO concerns recovery time. RPO concerns recoverable data point.

RTO โ‰  MTD

RTO is a recovery objective. MTD represents maximum tolerable disruption.

Backup Exists โ‰  Backup Recoverable

Restore testing is required.

Backup Success โ‰  Restore Success

A successful backup job does not prove the recovery process works.

Onsite Backup โ‰  Disaster Separation

The same physical event may destroy production and the backup.

Cloud Backup โ‰  Automatically Resilient

Provider, identity, region and connectivity dependencies still matter.

Snapshot โ‰  Always Independent Backup

It may share the production failure domain.

Replication โ‰  Backup

Corruption and malicious deletion can also replicate.

Newest Backup โ‰  Known-Good Backup

Attacker persistence may already exist.

Immutable โ‰  No Access Control Needed

Backup infrastructure still requires strong security.

Encrypted Backup โ‰  Recoverable Without Keys

Key availability is part of recovery planning.

Cold Site โ‰  Hot Site

Cold sites require considerably more preparation after activation.

Hot Site โ‰  Zero RTO

Activation, routing and validation may still take time.

Hot Site โ‰  Zero RPO

Data currency depends on backup or replication strategy.

Different Building โ‰  Different Failure Domain

Sites may share power, telecoms, geography or other dependencies.

Reciprocal Agreement โ‰  Guaranteed Technical Compatibility

Capacity and capability require validation.

Second Site โ‰  Enough Capacity

Determine whether it can support required recovery workload.

Multi-Region โ‰  No Common Dependencies

Identity, DNS and administrative systems may remain shared.

Synchronous โ‰  Always Best

Performance, latency and distance constraints matter.

Asynchronous โ‰  Zero Data Loss

Secondary data can lag behind the primary.

High Availability โ‰  Disaster Recovery

HA may handle component failure while remaining vulnerable to a facility-wide event.

HA โ‰  Fault Tolerance

HA minimises interruption. Fault tolerance emphasises continued correct operation despite fault.

RAID โ‰  Backup

Storage redundancy does not provide historical recovery.

QoS โ‰  Redundant Network

QoS manages existing capacity rather than creating an alternate path.

Application Restored โ‰  Service Restored

Dependencies such as DNS, databases and identity may still be unavailable.

Recovery Priority โ‰  Recovery Order

A critical application may depend on infrastructure that has to be recovered first.

Availability Restored โ‰  Security Restored

Recovery environments still require appropriate security controls.

Recovery Strategy Exists โ‰  Recovery Capability Proven

Test it.

Recovery Site Built Once โ‰  Recovery Site Ready Forever

Production and recovery environments must remain aligned.

7.10 โ‰  7.11

7.10 focuses on recovery capabilities and strategy. 7.11 focuses on implementing the DR process.

Quick Reference

If you see...Think...
How quickly must service return?RTO
How much recent data can be lost?RPO
Maximum tolerable disruptionMTD
All data copiedFull Backup
Changes since previous backupIncremental
Changes since full backupDifferential
Point-in-time storage stateSnapshot
Current copy at another siteReplication
Protected against modificationImmutable Backup
Disconnected recovery copyOffline Backup
Building and utilities onlyCold Site
Partially equipped siteWarm Site
Highly prepared recovery facilityHot Site
Two organisations support each otherReciprocal Agreement
Primary + standbyActive / Passive
Both sites processingActive / Active
Both locations confirm writeSynchronous Replication
Replica slightly behindAsynchronous Replication
Two nodes both become primarySplit Brain
Reduce downtime using failoverHigh Availability
Continue despite component faultFault Tolerance
Prioritise critical network trafficQoS
Operate with reduced capabilityGraceful Degradation
Another component can take overRedundancy
One failure takes down everythingSingle Point of Failure
Does backup actually restore?Recovery Testing
Most recent copy may contain attackerKnown-Good Recovery Point

RTO ยท RPO ยท MTD Memory Aid

RTO How long until service returns?
RPO How far back must data recover?
MTD How long until disruption becomes intolerable?

Backup Memory Aid

FULL Everything
INCREMENTAL Since previous backup
DIFFERENTIAL Since last full
REPLICATION Current copy
BACKUP Recoverable history

Recovery Site Memory Aid

COLD Space ready
WARM Some technology ready
HOT Technology ready

Colder = cheaper + slower ยท Hotter = costlier + faster

Availability Memory Aid

REDUNDANCY Another component
HA Minimise downtime
FAULT TOLERANCE Continue through fault
QoS Prioritise service quality
RESILIENCE Withstand + adapt + recover

7.10 Master Memory Aid

OBJECTIVES RTO ยท RPO ยท MTD
BACKUP Protect recoverable data
SITE Provide alternate location
REPLICATE Maintain another processing copy
REDUNDANCY Remove single points of failure
RESILIENCE Continue and recover
TEST Prove the strategy

Requirements โ†’ Backup โ†’ Alternate Capacity โ†’ Resilience โ†’ Test

The Recovery Architect's Questions

CRITICAL? Which business service matters?
RTO? How quickly must it return?
RPO? How much data can be lost?
MTD? When does harm become unacceptable?
BACKUP? What recoverable copies exist?
SAFE? Can attackers destroy them?
SITE? Where will processing occur?
CAPACITY? Can the alternate environment handle demand?
DEPENDENCIES? What must recover first?
REDUNDANCY? Where are single points of failure?
SECURE? Will recovery preserve security controls?
TESTED? Can we prove it works?

Key Takeaways

CISSP 7.10 focuses on implementing recovery strategies.

The current CISSP outline explicitly covers backup storage strategies, recovery site strategies, multiple processing sites, system resilience, high availability, Quality of Service and fault tolerance.

Recovery strategy should begin with business requirements.

A Business Impact Analysis helps identify critical processes and the recovery requirements needed to support them.

RTO describes how quickly service should be recovered.

RPO describes the point in time to which data should be recovered.

MTD represents how long disruption can continue before significant business harm occurs.

RTO = time to restore. RPO = data to restore. MTD = maximum tolerable disruption.

Recovery capability must be capable of satisfying the business recovery objectives.

Faster recovery generally requires more ready capacity and therefore greater investment.

Recovery architecture should therefore be proportionate to business impact.

Backup strategies can use onsite, offsite and cloud storage.

Onsite backups can provide convenient and rapid access but may share physical risk with production.

Offsite copies provide greater separation from site-level disasters.

Cloud backup can provide scalable remote storage, but connectivity, identity, provider resilience, encryption and restoration throughput still matter.

Full backups copy all selected data.

Incremental backups copy changes since the previous backup.

Differential backups copy changes since the last full backup.

Full = everything. Incremental = since previous backup. Differential = since last full.

Incremental strategies can reduce backup time and storage requirements but often create a longer restore chain.

Backup design should therefore consider restoration requirements rather than optimising only backup speed.

Snapshots can provide valuable point-in-time recovery.

However, a snapshot is not automatically an independent backup if it remains within the same storage or administrative failure domain.

Replication and backup are also different.

Replication maintains another relatively current copy and can support availability.

Historical backups allow recovery to an earlier state.

Replication = current copy. Backup = recoverable history.

Replication can faithfully reproduce ransomware encryption, corruption or malicious deletion.

Historical recovery points are therefore still required even in highly replicated environments.

Ransomware recovery may require backups that attackers cannot easily modify or delete.

Immutable or offline recovery copies, separate administrative credentials and strong access controls can strengthen recovery resilience.

Backups themselves contain sensitive information.

Their confidentiality, integrity and availability must therefore be protected.

Encryption can protect backup confidentiality, but recovery planning must also protect and preserve the keys required to decrypt that information.

Encrypted backup + unavailable key = unavailable recovery.

The most important measure of a backup is whether the organisation can successfully restore from it.

Backup success โ‰  restore success.

Recovery site strategies commonly include cold, warm and hot sites.

Cold sites provide basic facility capability and require substantial technology installation before operation.

Warm sites provide some technology and telecommunications capability.

Hot sites are much more prepared for rapid technology recovery.

Generally:

colder = cheaper but slower. hotter = more expensive but faster.

A hot site does not automatically guarantee zero downtime.

Routing, application activation, data synchronisation and validation may still be required.

A hot site also does not automatically guarantee zero data loss.

Data recovery depends on the backup and replication architecture.

Recovery capacity can also be obtained through contractual arrangements with providers or through reciprocal arrangements between organisations.

Such arrangements should define and validate capacity, compatibility, activation time and responsibilities.

Different physical sites should reduce exposure to common-mode failures.

Two facilities may still share electricity, telecommunications, flood zones or administrative systems.

Different location โ‰  different failure domain.

Multiple processing sites can use active-passive or active-active architectures.

Active-passive designs keep one environment operating while another waits to take over.

Active-active designs distribute live production workload across multiple locations.

Recovery capacity must be considered.

If two active sites normally each process half of the workload, each site may need enough remaining capacity to support the required service when the other fails.

Cloud availability zones and regions can support multi-site recovery architectures.

They do not eliminate common dependencies such as identity, DNS, configuration, privileged access or application defects.

Data replication can be synchronous or asynchronous.

Synchronous replication keeps participating copies very closely aligned but introduces coordination and latency considerations.

Asynchronous replication allows the secondary copy to lag the primary.

This can support greater geographic separation but may permit loss of recent transactions during sudden failover.

Distributed recovery designs should also guard against split-brain conditions where multiple systems incorrectly believe they are authoritative at the same time.

System resilience is broader than recovery.

Resilient systems aim to withstand disruption, continue important functions where possible, adapt to degraded conditions and recover effectively.

High availability uses redundancy and failover to minimise service interruption.

Fault tolerance aims to allow correct operation to continue despite component failure within the system's designed tolerance.

HA = minimise interruption. Fault tolerance = continue despite fault.

High Availability and Disaster Recovery are not the same.

A highly available cluster inside one building may survive server failure but still disappear when the entire building fails.

Load balancing and clustering can support availability by distributing workloads and enabling failover.

Their effectiveness depends on avoiding shared single points of failure.

RAID and similar disk-resilience technologies can protect against some storage-device failures.

RAID โ‰  backup.

Disk redundancy does not preserve historical state against ransomware, deletion or application corruption.

Quality of Service can prioritise important network traffic when resources are constrained.

QoS may help critical applications continue operating during degraded network conditions.

QoS does not create an alternate network connection.

QoS = prioritise available capacity. Redundancy = provide alternative capacity.

Graceful degradation is another resilience approach.

During disruption, less important functionality can be reduced so essential services remain available.

Recovery architectures should identify single points of failure across power, compute, network, storage, sites, suppliers and people.

Recovery planning must also understand application dependencies.

An application can be operational while remaining unusable because DNS, identity, networking, databases or encryption-key services are still unavailable.

Application recovered โ‰  business service recovered.

Recovery order must therefore reflect dependencies.

The service with the highest business priority may rely on lower-level infrastructure that must technically recover first.

Cyber recovery requires additional attention to data integrity.

The newest backup may not be the safest backup if an attacker had already established persistence.

Most recent recovery point โ‰  known-good recovery point.

Recovery environments should retain appropriate security controls.

Logging, identity controls, segmentation, encryption and monitoring should not be casually discarded simply because restoration is urgent.

Suppliers and cloud providers can become critical recovery dependencies.

Their real recovery capabilities should align with the organisation's own recovery requirements.

A business RTO of two hours cannot be reliably achieved through a critical supplier whose recovery capability takes 24 hours unless another strategy exists.

Recovery capability must be maintained as the production environment changes.

An alternate site that has not been updated after years of application and architecture changes may no longer provide meaningful recovery.

Recovery strategies should therefore be regularly tested.

Testing should validate backups, failover, data integrity, capacity, dependencies, security and actual recovery time.

Designed recovery โ‰  demonstrated recovery.

CISSP 7.10 and 7.11 are closely connected but should not be confused.

7.10 focuses on the strategies and capabilities that make recovery possible.

7.11 focuses on performing the Disaster Recovery processes when those capabilities are needed.

7.10 = How will recovery be possible? 7.11 = How do we execute the recovery?

The central CISSP principle is:

determine business recovery requirements first, then implement sufficiently independent backup, alternate-processing, redundancy and resilience capabilities to meet those requirements - and test them before the organisation has to depend on them.

๐Ÿ“š Sources & Further Reading Recovery, resilience and contingency-planning references