7.10 Recovery Strategies
7.10 Recovery Strategies
Recovery is not something an organisation should design after systems have already failed.
Recovery strategies determine in advance how technology, data and critical services can continue or be restored when normal operations are disrupted.
Different systems need different strategies.
A payroll archive might tolerate several hours of downtime. A payment platform may require almost continuous service. A ransomware event may require recovery from an older trusted backup rather than simply failing over to another live copy.
The correct recovery strategy therefore depends on: business requirements, acceptable downtime, acceptable data loss, system dependencies, threat scenarios and cost.
Protect Data
Maintain recoverable copies of important information and system components.
WHAT CAN WE RESTORE?Recover Services
Provide alternate facilities, systems and processing capability.
WHERE CAN WE RUN?Remain Resilient
Design systems that continue operating despite component failures and disruptions.
HOW DO WE KEEP GOING?Implement Recovery Strategies
The current CISSP Exam Outline explicitly identifies four areas.
Including examples such as:
Including examples such as:
Distribute processing capability across more than one physical or logical location.
7.10 Official Topics
The Big Idea
RTO ยท RPO ยท MTD
The target for how quickly a system or service must be restored after disruption.
Think:
HOW LONG CAN RECOVERY TAKE?
The point in time to which data must be recovered.
It represents the amount of recent data the organisation can tolerate losing.
Think:
HOW MUCH DATA CAN WE LOSE?
The maximum period that the business process can be disrupted before the disruption causes unacceptable or significant harm.
Think:
WHEN DOES THE BUSINESS DAMAGE BECOME UNACCEPTABLE?
RTO vs RPO vs MTD
Visualising RTO & RPO
Failure: 10:00
Recoverable point: 09:00
Potential data loss: 1 hour.
Failure: 10:00
Required restoration: 12:00
Recovery objective: 2 hours.
โฑ๏ธ RTO Must Fit Inside Business Tolerance The recovery objective must support the business requirement
Maximum tolerable disruption: 8 hours.
Expected recovery: 14 hours.
The technical recovery capability does not satisfy the business requirement.
Recovery Speed Has a Cost
Backup Storage Strategies
A backup is a recoverable copy of information or system components maintained so they can be restored if the original is lost, corrupted or unavailable.
Backup stored at or very near the primary operating location.
Advantage: fast local access.
Risk: same physical disaster may affect both.
Backup stored at a separate geographic location.
Advantage: reduces common physical-site risk.
Backup data stored using cloud infrastructure or backup services.
Consider:
Full ยท Incremental ยท Differential
| Backup Type | What Is Copied? | Backup Speed / Size | Typical Restore Requirement |
|---|---|---|---|
| Full | All selected data | Largest / longest | Full backup |
| Incremental | Changes since the previous backup | Usually smallest / fastest | Full + subsequent incrementals |
| Differential | Changes since the last full backup | Grows until next full | Full + latest differential |
Backup Types
๐ Backup Example Follow what each backup contains
| Day | Change | Incremental | Differential |
|---|---|---|---|
| Sunday | Full backup | Full | Full |
| Monday | File A changes | A | A |
| Tuesday | File B changes | B | A + B |
| Wednesday | File C changes | C | A + B + C |
Backup Speed vs Restore Complexity
Advantages:
Restore may require: multiple backup sets.
Backup grows larger throughout the cycle.
Restore commonly requires: full backup + latest differential.
Simplest recovery chain.
But: requires more storage and backup time.
Snapshot vs Backup
Captures system or storage state at a point in time.
Often useful for:
Maintains a recoverable copy that can be separated from the original environment or failure domain.
Often stronger for: major disaster or destructive attack recovery.
A snapshot stored on the same compromised or failed storage platform may disappear with the original environment.
Replication โ Backup
Keeps another copy relatively current.
Excellent for: availability and fast failover.
Preserves recoverable historical state.
Useful when: current state itself is corrupted or malicious.
๐ Ransomware Changes the Backup Question Can attackers destroy the recovery copy too?
Production administrators can:
delete every backup.
Attacker compromises: production administrator account.
Result:
production + backups destroyed.
Useful protections can include:
Immutable & Offline Backups
Backup cannot be modified or deleted during a defined retention period according to the storage controls used.
Useful against: malicious or accidental alteration.
Recovery copy is not continuously accessible from the production environment.
Useful against: attacks spreading through connected systems.
3๏ธโฃ Common 3-2-1 Backup Heuristic A useful operational memory aid rather than an ISC2 7.10 sub-bullet
3-2-1
Immutability, offline copies and separate administrative credentials can provide additional protection against destructive attacks.
Backups Contain Valuable Data Too
Encrypt sensitive backup data where appropriate.
Protect backup contents against unauthorised modification.
Restrict who can read, restore or delete backups.
Maintain recovery points according to business and regulatory needs.
Detect unexpected deletion or modification of recovery copies.
Ensure encryption keys needed for recovery remain available and protected.
Can You Restore It?
Backup jobs: 100% successful.
Restore attempted.
Result: backup archive corrupt.
Backup Assurance
Recovery Site Strategies
A recovery site provides an alternate location from which technology or business operations can be restored after the primary environment becomes unavailable.
Facility with basic environmental and physical capabilities but without the full computing environment ready for immediate operation.
CHEAPER ยท SLOWER RECOVERY
Partially equipped alternate facility with some systems and telecommunications capability already available.
MIDDLE GROUND
Highly prepared alternate processing environment with required hardware and software available for recovery.
MORE EXPENSIVE ยท FASTER RECOVERY
Recovery Sites
Cold vs Warm vs Hot
| Site | Readiness | Recovery Speed | Relative Cost |
|---|---|---|---|
| Cold | Basic facility | Slowest | Lowest |
| Warm | Partially equipped | Intermediate | Intermediate |
| Hot | Highly equipped | Fastest of the three | Highest |
๐ฅ Hot Site โ Automatically Zero Downtime The site may be ready while data or applications still require activation
A hot site may have hardware, software and connectivity available.
Recovery can still require:
Reserved & Reciprocal Capacity
Recovery capability does not always require the organisation to own an entire idle recovery facility.
A service provider agrees to make defined recovery resources or capacity available when needed.
Consider:
Two organisations agree to support one another's recovery capability when disruption occurs.
Advantages: potentially lower cost.
Challenges: capacity, compatibility and simultaneous disaster.
๐ค Reciprocal Agreements Need Reality Checks Good intentions are not recovery capacity
Company A and Company B agree to: host one another during disaster.
Five years later:
Geographic Separation
Alternate sites should reduce exposure to the same disruption affecting the primary site.
Primary data centre: Building A.
Recovery site: Building B next door.
Shared:
Multiple Processing Sites
Rather than operating from one primary environment and activating another only after disaster, organisations can distribute processing capability across multiple sites.
One environment normally serves production while another remains ready to take over.
PRIMARY + STANDBY
Multiple sites actively process production workload.
MULTIPLE LIVE SITES
Multi-Site
๐ข Active / Passive Standby capacity is activated when the primary fails
๐ข Active / Active Production workload is distributed across multiple environments
Active-active architecture may require careful coordination of data, application state, traffic routing and capacity.
Can the Remaining Site Handle the Load?
Site A: 50% workload.
Site B: 50% workload.
Site B must potentially support: 100% workload.
If each site has capacity for only 60% of total demand, failover may work technically while service performance becomes unacceptable.
Availability Zones & Regions
Cloud platforms can provide logically or geographically separated deployment locations that support resilient architectures.
Protect against some local infrastructure or facility failures.
Can provide broader geographic separation for larger-scale disruptions.
Common dependencies such as identity, DNS, configuration, cloud account, deployment pipelines or application defects may still affect all locations.
โ ๏ธ Common-Mode Failure Redundant systems can still depend on the same thing
Site A: redundant.
Site B: redundant.
Both use:
Failure of the shared dependency can still: disable both sites.
Synchronous vs Asynchronous Replication
A write is committed across participating locations before the transaction is treated as complete.
Advantage: very small potential data gap.
Trade-off: latency and distance can affect performance.
Primary operation completes before the secondary copy necessarily receives every change.
Advantage: supports greater distance and lower latency impact.
Trade-off: recent transactions may be lost during sudden failover.
Replication
๐ง Split-Brain Risk What if both sides believe they are primary?
Site A loses communication with Site B.
Site A believes: Site B failed.
Site B believes: Site A failed.
Both begin accepting: independent writes.
Result: data inconsistency or corruption.
System Resilience
Resilience is broader than simply restoring a failed system.
A resilient system is designed to:
Resilience
โป๏ธ Resilience vs Recovery One tries to keep capability available while the other restores it
Continue important capability despite adverse conditions where possible.
KEEP GOING
Restore capability after disruption.
BRING IT BACK
High Availability - HA
High availability uses redundancy and failover mechanisms to reduce service interruption when components fail.
High Availability
High Availability vs Disaster Recovery
Primarily aims to minimise interruption from component or local service failure.
KEEP THE SERVICE RUNNING
Restores capability after a significant disruptive event has affected normal operations.
RECOVER AFTER MAJOR DISRUPTION
A highly available system located entirely within one data centre may still fail when the entire facility becomes unavailable.
โ๏ธ Load Balancing Distribute demand across multiple service instances
๐ Clustering Multiple systems cooperate to provide a service
Cluster designs can provide:
Continue Despite Failure
Fault tolerance is the ability of a system to continue correct operation despite failure of one or more components within its designed tolerance.
One power supply can fail while another continues supplying power.
Alternative communication path remains available.
Some storage designs continue operating when a disk fails.
Multiple components remove reliance on a single device.
Fault Tolerance
๐ High Availability vs Fault Tolerance Closely related but not identical ideas
Designed to minimise downtime.
A brief failover interruption may still occur.
Designed so the system continues correct operation despite a component fault within its tolerance.
Greater emphasis on: continuity through failure.
Disk Redundancy
RAID and similar storage technologies can provide resilience against some disk failures.
RAID configuration may allow: service to continue while failed storage is replaced.
RAID may protect against disk failure.
It does not inherently protect against:
QoS - Protect Important Traffic
Quality of Service provides mechanisms for controlling or assuring characteristics of network service such as bandwidth, latency, priority, packet loss or jitter.
Network capacity reduced: 50%.
Competing traffic:
QoS can prioritise: the services most important to continued operations.
QoS
๐ก QoS โ Redundancy QoS manages available capacity - it does not create another network
WAN link: completely unavailable.
QoS cannot prioritise traffic across: a link that no longer exists.
Graceful Degradation
Resilient systems do not always need to provide every feature during disruption.
They may deliberately preserve: essential functionality while reducing less important capability.
Preserve:
Temporarily reduce:
Redundancy
UPS, generators, multiple power feeds.
Multiple links, routers, switches and carriers.
Multiple servers or cluster nodes.
Replication and redundant storage components.
Multiple processing locations.
Avoid dependency on one individual with unique operational knowledge.
The Redundant Data Centre With One Router
Application has:
But every external connection uses:
one network router.
Router fails.
Result: entire service unavailable.
Recover the Dependency Chain
Applications rarely operate alone.
Application recovered: successfully.
DNS unavailable: users still cannot access it.
๐ข Recovery Order Matters Dependencies may need restoration before the application
Recovery Systems Still Need Security
Organisations sometimes weaken controls during disaster recovery because restoring availability becomes urgent.
Production application restored quickly.
Temporary recovery environment uses:
Known-Good Recovery
Recovery from cyberattack creates a problem that traditional hardware failure does not always create:
the newest copy may already be compromised.
Monday: files encrypted.
Investigation finds attacker entered: three weeks earlier.
Restoring Sunday's backup may restore: attacker persistence.
Before Cyber Recovery
Recovery Does Not Always Mean Technology
Some business processes can continue temporarily using alternative or manual methods while technology is restored.
Automated approval platform cannot operate.
Temporary process: documented manual approval procedure.
Recovery Depends on Suppliers Too
What resilience and recovery capability is actually provided?
Are alternate communication paths available?
What are its recovery commitments and data-backup capabilities?
Is capacity guaranteed during a widespread disruption?
๐ SLAs & Recovery Commitments Contractual wording should support the required business outcome
Recovery-related agreements can address:
Business RTO: 2 hours.
Critical provider commitment: recovery within 24 hours.
A Recovery Strategy Must Be Tested
A recovery architecture existing on paper does not prove that it can achieve the required RTO and RPO.
Can data actually be restored?
Does the secondary environment take over?
Can the recovery environment handle required workload?
Are DNS, identity, networking and integrations recoverable?
Are security controls still functioning after recovery?
Did recovery actually meet the RTO?
7.10 Recovery Strategy vs 7.11 Disaster Recovery Process
Determine the capabilities available for recovery.
HOW WILL RECOVERY BE POSSIBLE?
Covers what the organisation does when executing disaster recovery.
HOW DO WE PERFORM THE RECOVERY?
7.10 vs 7.11
๐ข Business Continuity vs Disaster Recovery Keep the business operating vs restore disrupted technology
Broader capability to continue important business processes during disruption.
KEEP BUSINESS OPERATING
Focuses more directly on restoring technology and information capabilities after major disruption.
RESTORE TECHNOLOGY
Payment Platform
Business requirements:
Strategy:
Monthly Reporting System
Business requirements:
A continuously active second data centre might provide: far more recovery capability than the business requires.
The Two Data Centres in One Flood Zone
Primary site and recovery site are: five kilometres apart.
Both depend on: the same river flood defence.
Multi-Region But One Compromised Account
Application is deployed across: two cloud regions.
Both environments are administered through: one privileged cloud account.
Attacker compromises the account and deletes: resources in both regions.
The Perfectly Replicated Ransomware
Primary file system: encrypted by ransomware.
Replication: working perfectly.
Result: encrypted files rapidly copied to secondary site.
The Backup Nobody Can Decrypt
Organisation maintains: excellent encrypted offsite backups.
Disaster destroys: the only key-management server containing the decryption key.
The Hot Site With No Capacity
Recovery site has:
But it can support: 30% of normal user demand.
Reduced Network Capacity
One WAN circuit fails.
Remaining circuit can support: 60% of normal traffic.
QoS prioritises:
Lower-priority bulk transfer is: deprioritised.
The Recovery Site Nobody Maintained
Recovery site was built: three years ago.
Production has since received:
Recovery environment: was never updated.
๐ CISSP Scenarios Recognise the recovery strategy being tested
A business process can tolerate two hours before the system must be available again.
Which objective?
RTO.
The business can tolerate losing at most 30 minutes of transactions.
Which objective?
RPO.
A business process cannot remain disrupted beyond eight hours without significant harm.
Which concept?
MTD.
RTO is 12 hours but MTD is 6 hours.
Problem?
The recovery objective does not satisfy business tolerance.
A backup is stored in the same data centre as the production system.
Primary risk?
Both may be lost in the same site disaster.
A backup copy is maintained at a geographically separate location.
Which strategy?
Offsite backup.
All selected data is copied during every backup.
Which backup?
Full.
Only changes since the previous backup are copied.
Which backup?
Incremental.
All changes since the last full backup are copied.
Which backup?
Differential.
Which generally requires a full plus multiple subsequent backup sets during restoration?
Answer?
Incremental backup strategy.
A storage snapshot resides on the same failed array as the production volume.
Primary lesson?
A snapshot is not automatically an independent backup.
Deletion from the primary database immediately appears on the secondary replica.
Which lesson?
Replication is not the same as historical backup.
Ransomware encrypts production and the encryption is copied to the replica.
What additional capability is important?
Protected historical backups.
Attackers cannot modify or delete a recovery copy during its retention period.
Which concept?
Immutable backup.
A recovery copy is disconnected from normal production access.
Which concept?
Offline / air-gapped backup.
Backup dashboard says every job succeeded but restores have never been tested.
Primary concern?
Recoverability is unproven.
Encrypted backups survive the disaster but decryption keys do not.
Result?
The backups may be unavailable for recovery.
An alternate facility has power and environmental controls but no computing hardware installed.
Which site?
Cold site.
An alternate facility is partially equipped with computing and telecommunications capability.
Which site?
Warm site.
An alternate facility has hardware and software available for rapid recovery.
Which site?
Hot site.
Which site generally costs the least but takes longest to prepare?
Answer?
Cold site.
Does having a hot site automatically mean zero data loss?
Answer?
No. Site readiness and data replication are separate issues.
Two organisations agree to support each other's processing requirements during disruption.
Which arrangement?
Reciprocal agreement.
A recovery provider promises resources but cannot guarantee enough servers during a regional disaster.
Primary concern?
Recovery capacity.
The primary and recovery sites use the same power substation.
Primary concern?
Common-mode failure.
One processing site runs production while the second waits to take over.
Which model?
Active-passive.
Two geographically separate sites simultaneously process customer traffic.
Which model?
Active-active.
Each of two active sites handles 50% of traffic but can support only 60% alone.
What is the concern?
Insufficient failover capacity.
Two regions rely on the same compromised privileged account.
What does this demonstrate?
Geographic redundancy can still share a logical failure domain.
A transaction is not committed until the secondary location confirms the write.
Which replication?
Synchronous.
The secondary site may lag slightly behind the primary.
Which replication?
Asynchronous.
Which replication approach can create potential loss of the most recent writes during sudden failure?
Answer?
Asynchronous replication.
Two isolated cluster nodes both believe they are primary and accept writes.
Which problem?
Split brain.
A server fails and another automatically assumes its workload after a short interruption.
Which capability?
High availability / failover.
A system continues correct service without interruption when one redundant component fails.
Which concept?
Fault tolerance.
A degraded network prioritises voice and critical transaction traffic.
Which concept?
Quality of Service.
The only WAN circuit completely fails.
Can QoS restore the missing circuit?
No. Redundancy is required for an alternate path.
A RAID array survives one disk failure.
Does this mean the organisation no longer needs backups?
No.
A service continues operating with reduced non-critical functionality during disruption.
Which principle?
Graceful degradation / resilience.
The application server is restored before DNS and identity services.
Primary problem?
Recovery dependency sequence was not considered.
The newest backup contains attacker persistence.
Should it automatically be restored?
No. Select a sufficiently trusted recovery point.
A SaaS provider has a 24-hour recovery commitment while the business needs the service within one hour.
Primary problem?
Supplier recovery capability does not support organisational RTO.
An alternate site exists but has not been updated after years of production changes.
Primary risk?
Recovery environment may no longer be compatible or usable.
Management asks what 7.10 primarily determines.
Best answer?
The capabilities and architecture used to make recovery possible.
Management asks for the central principle of recovery strategy.
Best answer?
Select and maintain backup, alternate processing and resilience capabilities that satisfy business recovery objectives and have been validated through testing.
Recognise the Clue Words
How Long to Restore?
Recovery time.
RTOHow Much Data Loss?
Recovery point.
RPOMaximum Business Disruption
Absolute tolerance.
MTDSame Building
Fast but shared risk.
Onsite BackupSeparate Geography
Site resilience.
Offsite BackupEverything Copied
Backup type.
FullSince Previous Backup
Backup type.
IncrementalSince Last Full
Backup type.
DifferentialPoint-in-Time State
Fast recovery.
SnapshotLive Copy
Availability.
ReplicationCannot Be Modified
Recovery protection.
Immutable BackupDisconnected Copy
Isolation.
Offline BackupBuilding Only
Slow recovery.
Cold SitePartially Equipped
Middle recovery.
Warm SiteHighly Ready
Fast recovery.
Hot SiteOrganisations Support Each Other
Capacity arrangement.
Reciprocal AgreementOne Live ยท One Standby
Processing sites.
Active / PassiveBoth Live
Processing sites.
Active / ActiveWrites Confirmed Both Places
Data currency.
Synchronous ReplicationReplica Can Lag
Data gap.
Asynchronous ReplicationBoth Think They Are Primary
Distributed data problem.
Split BrainMinimise Downtime
Availability.
HAContinue Despite Fault
Continuity.
Fault TolerancePrioritise Network Traffic
Service quality.
QoSReduced but Useful Service
Adaptation.
Graceful DegradationAnother Component Exists
Resilience.
RedundancyOne Component Stops Everything
Architecture risk.
Single Point of FailureRestore One System But Service Fails
Architecture.
Dependency ProblemNewest Copy May Be Infected
Cyber recovery.
Known-Good Recovery PointDoes It Actually Restore?
Assurance.
Recovery Testingโ ๏ธ Common CISSP Mistakes Recovery capability must match the business requirement
RTO concerns recovery time. RPO concerns recoverable data point.
RTO is a recovery objective. MTD represents maximum tolerable disruption.
Restore testing is required.
A successful backup job does not prove the recovery process works.
The same physical event may destroy production and the backup.
Provider, identity, region and connectivity dependencies still matter.
It may share the production failure domain.
Corruption and malicious deletion can also replicate.
Attacker persistence may already exist.
Backup infrastructure still requires strong security.
Key availability is part of recovery planning.
Cold sites require considerably more preparation after activation.
Activation, routing and validation may still take time.
Data currency depends on backup or replication strategy.
Sites may share power, telecoms, geography or other dependencies.
Capacity and capability require validation.
Determine whether it can support required recovery workload.
Identity, DNS and administrative systems may remain shared.
Performance, latency and distance constraints matter.
Secondary data can lag behind the primary.
HA may handle component failure while remaining vulnerable to a facility-wide event.
HA minimises interruption. Fault tolerance emphasises continued correct operation despite fault.
Storage redundancy does not provide historical recovery.
QoS manages existing capacity rather than creating an alternate path.
Dependencies such as DNS, databases and identity may still be unavailable.
A critical application may depend on infrastructure that has to be recovered first.
Recovery environments still require appropriate security controls.
Test it.
Production and recovery environments must remain aligned.
7.10 focuses on recovery capabilities and strategy. 7.11 focuses on implementing the DR process.
Quick Reference
| If you see... | Think... |
|---|---|
| How quickly must service return? | RTO |
| How much recent data can be lost? | RPO |
| Maximum tolerable disruption | MTD |
| All data copied | Full Backup |
| Changes since previous backup | Incremental |
| Changes since full backup | Differential |
| Point-in-time storage state | Snapshot |
| Current copy at another site | Replication |
| Protected against modification | Immutable Backup |
| Disconnected recovery copy | Offline Backup |
| Building and utilities only | Cold Site |
| Partially equipped site | Warm Site |
| Highly prepared recovery facility | Hot Site |
| Two organisations support each other | Reciprocal Agreement |
| Primary + standby | Active / Passive |
| Both sites processing | Active / Active |
| Both locations confirm write | Synchronous Replication |
| Replica slightly behind | Asynchronous Replication |
| Two nodes both become primary | Split Brain |
| Reduce downtime using failover | High Availability |
| Continue despite component fault | Fault Tolerance |
| Prioritise critical network traffic | QoS |
| Operate with reduced capability | Graceful Degradation |
| Another component can take over | Redundancy |
| One failure takes down everything | Single Point of Failure |
| Does backup actually restore? | Recovery Testing |
| Most recent copy may contain attacker | Known-Good Recovery Point |
RTO ยท RPO ยท MTD Memory Aid
Backup Memory Aid
Recovery Site Memory Aid
Colder = cheaper + slower ยท Hotter = costlier + faster
Availability Memory Aid
7.10 Master Memory Aid
Requirements โ Backup โ Alternate Capacity โ Resilience โ Test
The Recovery Architect's Questions
Key Takeaways
CISSP 7.10 focuses on implementing recovery strategies.
The current CISSP outline explicitly covers backup storage strategies, recovery site strategies, multiple processing sites, system resilience, high availability, Quality of Service and fault tolerance.
Recovery strategy should begin with business requirements.
A Business Impact Analysis helps identify critical processes and the recovery requirements needed to support them.
RTO describes how quickly service should be recovered.
RPO describes the point in time to which data should be recovered.
MTD represents how long disruption can continue before significant business harm occurs.
RTO = time to restore. RPO = data to restore. MTD = maximum tolerable disruption.
Recovery capability must be capable of satisfying the business recovery objectives.
Faster recovery generally requires more ready capacity and therefore greater investment.
Recovery architecture should therefore be proportionate to business impact.
Backup strategies can use onsite, offsite and cloud storage.
Onsite backups can provide convenient and rapid access but may share physical risk with production.
Offsite copies provide greater separation from site-level disasters.
Cloud backup can provide scalable remote storage, but connectivity, identity, provider resilience, encryption and restoration throughput still matter.
Full backups copy all selected data.
Incremental backups copy changes since the previous backup.
Differential backups copy changes since the last full backup.
Full = everything. Incremental = since previous backup. Differential = since last full.
Incremental strategies can reduce backup time and storage requirements but often create a longer restore chain.
Backup design should therefore consider restoration requirements rather than optimising only backup speed.
Snapshots can provide valuable point-in-time recovery.
However, a snapshot is not automatically an independent backup if it remains within the same storage or administrative failure domain.
Replication and backup are also different.
Replication maintains another relatively current copy and can support availability.
Historical backups allow recovery to an earlier state.
Replication = current copy. Backup = recoverable history.
Replication can faithfully reproduce ransomware encryption, corruption or malicious deletion.
Historical recovery points are therefore still required even in highly replicated environments.
Ransomware recovery may require backups that attackers cannot easily modify or delete.
Immutable or offline recovery copies, separate administrative credentials and strong access controls can strengthen recovery resilience.
Backups themselves contain sensitive information.
Their confidentiality, integrity and availability must therefore be protected.
Encryption can protect backup confidentiality, but recovery planning must also protect and preserve the keys required to decrypt that information.
Encrypted backup + unavailable key = unavailable recovery.
The most important measure of a backup is whether the organisation can successfully restore from it.
Backup success โ restore success.
Recovery site strategies commonly include cold, warm and hot sites.
Cold sites provide basic facility capability and require substantial technology installation before operation.
Warm sites provide some technology and telecommunications capability.
Hot sites are much more prepared for rapid technology recovery.
Generally:
colder = cheaper but slower. hotter = more expensive but faster.
A hot site does not automatically guarantee zero downtime.
Routing, application activation, data synchronisation and validation may still be required.
A hot site also does not automatically guarantee zero data loss.
Data recovery depends on the backup and replication architecture.
Recovery capacity can also be obtained through contractual arrangements with providers or through reciprocal arrangements between organisations.
Such arrangements should define and validate capacity, compatibility, activation time and responsibilities.
Different physical sites should reduce exposure to common-mode failures.
Two facilities may still share electricity, telecommunications, flood zones or administrative systems.
Different location โ different failure domain.
Multiple processing sites can use active-passive or active-active architectures.
Active-passive designs keep one environment operating while another waits to take over.
Active-active designs distribute live production workload across multiple locations.
Recovery capacity must be considered.
If two active sites normally each process half of the workload, each site may need enough remaining capacity to support the required service when the other fails.
Cloud availability zones and regions can support multi-site recovery architectures.
They do not eliminate common dependencies such as identity, DNS, configuration, privileged access or application defects.
Data replication can be synchronous or asynchronous.
Synchronous replication keeps participating copies very closely aligned but introduces coordination and latency considerations.
Asynchronous replication allows the secondary copy to lag the primary.
This can support greater geographic separation but may permit loss of recent transactions during sudden failover.
Distributed recovery designs should also guard against split-brain conditions where multiple systems incorrectly believe they are authoritative at the same time.
System resilience is broader than recovery.
Resilient systems aim to withstand disruption, continue important functions where possible, adapt to degraded conditions and recover effectively.
High availability uses redundancy and failover to minimise service interruption.
Fault tolerance aims to allow correct operation to continue despite component failure within the system's designed tolerance.
HA = minimise interruption. Fault tolerance = continue despite fault.
High Availability and Disaster Recovery are not the same.
A highly available cluster inside one building may survive server failure but still disappear when the entire building fails.
Load balancing and clustering can support availability by distributing workloads and enabling failover.
Their effectiveness depends on avoiding shared single points of failure.
RAID and similar disk-resilience technologies can protect against some storage-device failures.
RAID โ backup.
Disk redundancy does not preserve historical state against ransomware, deletion or application corruption.
Quality of Service can prioritise important network traffic when resources are constrained.
QoS may help critical applications continue operating during degraded network conditions.
QoS does not create an alternate network connection.
QoS = prioritise available capacity. Redundancy = provide alternative capacity.
Graceful degradation is another resilience approach.
During disruption, less important functionality can be reduced so essential services remain available.
Recovery architectures should identify single points of failure across power, compute, network, storage, sites, suppliers and people.
Recovery planning must also understand application dependencies.
An application can be operational while remaining unusable because DNS, identity, networking, databases or encryption-key services are still unavailable.
Application recovered โ business service recovered.
Recovery order must therefore reflect dependencies.
The service with the highest business priority may rely on lower-level infrastructure that must technically recover first.
Cyber recovery requires additional attention to data integrity.
The newest backup may not be the safest backup if an attacker had already established persistence.
Most recent recovery point โ known-good recovery point.
Recovery environments should retain appropriate security controls.
Logging, identity controls, segmentation, encryption and monitoring should not be casually discarded simply because restoration is urgent.
Suppliers and cloud providers can become critical recovery dependencies.
Their real recovery capabilities should align with the organisation's own recovery requirements.
A business RTO of two hours cannot be reliably achieved through a critical supplier whose recovery capability takes 24 hours unless another strategy exists.
Recovery capability must be maintained as the production environment changes.
An alternate site that has not been updated after years of application and architecture changes may no longer provide meaningful recovery.
Recovery strategies should therefore be regularly tested.
Testing should validate backups, failover, data integrity, capacity, dependencies, security and actual recovery time.
Designed recovery โ demonstrated recovery.
CISSP 7.10 and 7.11 are closely connected but should not be confused.
7.10 focuses on the strategies and capabilities that make recovery possible.
7.11 focuses on performing the Disaster Recovery processes when those capabilities are needed.
7.10 = How will recovery be possible? 7.11 = How do we execute the recovery?
The central CISSP principle is:
determine business recovery requirements first, then implement sufficiently independent backup, alternate-processing, redundancy and resilience capabilities to meet those requirements - and test them before the organisation has to depend on them.
๐ Sources & Further Reading Recovery, resilience and contingency-planning references
- ISC2 - CISSP Certification Exam Outline
View the current CISSP Exam Outline - NIST SP 800-34 Rev. 1 - Contingency Planning Guide for Federal Information Systems
View NIST contingency-planning guidance - NIST SP 800-184 - Guide for Cybersecurity Event Recovery
View NIST cybersecurity recovery guidance - NIST SP 800-53 Rev. 5 - Security and Privacy Controls for Information Systems and Organizations
View NIST SP 800-53 - NIST - Recovery Time Objective
View NIST RTO terminology - NIST - Recovery Point Objective
View NIST RPO terminology - NIST - High Availability
View NIST HA terminology - NIST - Fault Tolerance
View NIST fault-tolerance terminology - NIST - Quality of Service
View NIST QoS terminology
