Build a disaster recovery plan in eight steps: set governance and scope, assess your disaster risks, run a business impact analysis, map system and vendor dependencies, choose recovery strategies that meet your RTO and RPO targets, document procedures with named owners, test through escalating exercises, and maintain the plan on defined triggers.
On June 19, 2024, BlackSuit ransomware took down CDK Global, the dealer management platform behind roughly 15,000 North American dealerships, and a second strike a day later crushed the first recovery attempt. Dealers wrote deals on paper for about two weeks.
Anderson Economic Group put collective dealer losses at $1.02 billion, and CNN reported CDK likely paid a $25 million ransom to get its systems back. Every dollar of that bill is an argument for a rehearsed, RTO-driven recovery capability instead of improvisation.
|
Disaster Recovery Plan: Key Takeaways |
|
A disaster recovery plan documents how technology, data, and critical operations come back after a disruption, with named owners, tested procedures, and recovery clocks measured as RTO and RPO. |
|
CDK Global’s June 2024 ransomware outage put 15,000 dealerships on paper for two weeks and cost them an estimated $1.02 billion, the going rate for recovery that has never been rehearsed. |
|
Uptime Institute’s 2026 analysis prices the exposure: 57% of the most recent major outages cost over $100,000, and one in five topped $1 million for the second straight year. |
|
Build in eight steps: govern, assess risks, run the business impact analysis, map dependencies, choose RTO-matched strategies, write the plan, exercise it, and maintain it on triggers. |
|
Recovery strategies trade money for speed: backup-and-restore recovers in days, warm sites in hours, active-active in minutes, and NIST SP 800-34 frames the tiers. |
|
An untested disaster recovery plan is an assumption. Exercise quarterly to annually by tier, score against the RTO, and log fixes to closure the way FFIEC examiners expect. |
The eight-step build below follows NIST SP 800-34 and ISO 22301 discipline, prices the strategy choices honestly, and treats testing as the step that separates working recovery from confident stationery. Skim the tables; each one is a working artifact you can reuse.
What a Disaster Recovery Plan Covers
A disaster recovery plan is the technology chapter of your resilience library. Where the business continuity plan keeps priority operations running by any workable means, the DR plan restores the systems, data, and infrastructure those operations depend on, inside defined recovery clocks.
|
Plan |
What it restores |
Typical trigger and owner |
|
Disaster recovery plan |
Systems, data, connectivity, and facilities within RTO and RPO targets |
Technology outage or data loss; CIO or IT operations |
|
Business continuity plan |
Priority business activities, by workaround if needed |
Any disruption breaching impact thresholds; continuity manager |
|
Contingency plan |
One named scenario, end to end |
That scenario materializes; the process owner |
|
Incident response plan |
Security containment and eradication first |
Confirmed attack; security lead, feeding DR |
The boundaries matter operationally, because CDK-style events fire all four at once. Our guides to disaster recovery versus continuity planning, the contingency plan, and the incident response walkthrough cover the neighbors; everything below builds the DR chapter itself.
Why Recovery Speed Is a Financial Line Item
Recovery speed has a price curve, and so does its absence. Uptime Institute’s 2026 outage analysis found 57% of organizations’ most recent major outages cost over $100,000, with one in five clearing $1 million, the second consecutive year at that level.

Figure 1. Uptime Institute’s 2026 analysis: outage costs stay high, power still leads, and connectivity failures are lengthening.
Per-minute pricing sharpens the argument. ITIC’s downtime surveys put the median above $9,000 a minute for large enterprises, and IBM’s breach report adds an average $10.22 million bill for a US data breach, before regulators and class-action plaintiffs even arrive.

Figure 2. CDK Global, June 2024: two weeks of manual operations across 15,000 dealerships, and a ten-figure collective bill.
Ransomware is the scenario forcing the issue, which is why CISA’s StopRansomware program pushes offline backups so hard, and public companies now disclose material incidents on the SEC’s four-business-day clock. A recovery capability you can evidence is fast becoming a governance expectation, consistent with the wider business resilience agenda.
Before You Write: Risk Assessment and Business Impact Analysis
Two analyses feed every good DR plan, and skipping them is how plans end up protecting the wrong systems. A risk assessment identifies which disasters are plausible for your footprint, from ransomware and regional power loss to flood and vendor failure, using structured identification techniques.
The business impact analysis then ranks systems by the damage their absence causes over time, which converts into the two clocks that drive every downstream choice. The RPO and RTO distinction deserves precision, because vague targets buy the wrong architecture:
|
Clock |
Question it answers |
Example target |
It drives |
|
RTO |
How long until the system is back? |
4 hours for order processing |
Strategy tier and cost |
|
RPO |
How much data can we lose? |
15 minutes of transactions |
Backup and replication frequency |
|
MTD |
How long before damage is irreversible? |
24 hours for billing |
The ceiling RTO must beat |
|
Work recovery time |
How long to validate and catch up? |
2 hours after restore |
The gap RTO must leave |
Set the targets per system, not per company, and pressure-test them in a scenario exercise before committing budget. A continuity risk assessment workbook speeds the scoring and keeps on file the evidence a regulator or auditor will eventually ask to see.
How to Build the Disaster Recovery Plan in Eight Steps
With risks ranked and clocks set, the build follows a sequence a mid-size IT team can complete in a quarter. The steps compress NIST SP 800-34’s contingency process and ISO 22301’s planning clauses into deliverables, and each produces a dated artifact for the program file.
|
Step |
What you do |
Deliverable on record |
|
1. Govern |
Charter the program, name the DR lead, set scope and budget authority |
One-page charter with roles |
|
2. Assess risks |
Score plausible disaster scenarios for likelihood and impact |
Ranked scenario register |
|
3. Run the BIA |
Rank systems by time-based impact; set RTO, RPO, MTD per system |
Signed-off recovery targets |
|
4. Map dependencies |
Chart systems, data flows, vendors, and single points of failure |
Dependency map and vendor list |
|
5. Choose strategies |
Match each tier of systems to a recovery approach it can afford |
Strategy sheet with costs |
|
6. Write the plan |
Document activation, procedures, contacts, and runbooks per system |
The versioned DR plan |
|
7. Exercise |
Walk through, simulate, and fail over against the clocks |
Exercise reports with gaps |
|
8. Maintain |
Review on triggers, track fixes, retrain owners |
Change log and action tracker |
Step six is where most teams start and why most plans disappoint. Write it last: procedures documented before targets and strategies exist describe the infrastructure you have, never the recovery you need. A disaster recovery plan template accelerates the writing once steps one through five are real.
The plan document itself stays lean, because responders read it at 2am by flashlight, under pressure, and possibly without the network it normally lives on. Long prose fails that reader completely. The core contents every version of the plan carries:
- Activation criteria and who declares, with deputies named for every role
- Contact tree: staff, vendors, insurers, and regulators, verified quarterly
- System-by-system runbooks ordered by recovery priority, with credentials access noted
- Communication templates for staff, customers, and leadership, pre-approved by legal
- Return-to-normal criteria, data validation steps, and post-incident review requirements
Store copies where the disaster cannot reach them: printed, offline, and in a second cloud region. CDK’s dealers learned that a plan living inside the system that just died is a diary entry, and Ready.gov’s IT recovery guidance makes offline accessibility a baseline requirement.
Choosing Recovery Strategies That Match Your RTO
Strategy selection is arithmetic once the BIA exists: each tier of your systems gets the cheapest approach that still beats its RTO. Overshooting wastes budget on speed nobody needs, while undershooting quietly converts your published recovery targets into comfortable fiction.

Figure 3. The strategy ladder: each rung up buys recovery speed with standing infrastructure cost.
|
Strategy |
Typical RTO |
Relative cost |
Fits |
|
Backup and restore |
1 to 3 days |
$ |
Tier-3 systems, archives |
|
Cold site |
2 to 3 days |
$$ |
Facility loss cover, rarely alone |
|
Warm site |
Hours |
$$$ |
Core business applications |
|
Hot site / replica |
Under an hour |
$$$$ |
Revenue-critical platforms |
|
Active-active |
Minutes |
$$$$$ |
Systems that must never stop |
|
DRaaS |
Hours, contracted |
$$ to $$$$ |
Teams without a second site |
Cloud shifted the economics without repealing them: replication across regions makes warm and hot tiers affordable for mid-market firms, but egress fees, untested failback, and concentration in one provider are the new failure modes. ISO/IEC 27031 covers ICT readiness, and supply chain continuity planning handles the vendor half of the dependency map.
Whatever the tier, the ransomware era adds one non-negotiable: at least one backup copy that is offline, immutable, and tested for restore integrity. Attackers now target backups first, a pattern Verizon’s DBIR documents yearly, so a backup you can encrypt is a backup you do not have.
Exercise the Plan Until It Stops Surprising You
Strategy on paper proves nothing; the exercise ladder is where the plan becomes capability. The FFIEC’s continuity booklet expects escalating test complexity with findings tracked to closure, and its standard is worth adopting even where no examiner will ever ask.
|
Exercise |
Cadence |
What it proves |
|
Plan walkthrough |
Quarterly |
Owners know their runbooks; contacts still current |
|
Tabletop scenario |
Twice yearly |
Decisions, communications, and declaration authority work |
|
Restore test |
Quarterly per critical system |
Backups actually restore, with data integrity verified |
|
Failover drill |
Annually per tier-1 system |
The alternate site carries the load within RTO |
|
Full-scale exercise |
Every 12 to 24 months |
End-to-end recovery meets the clocks under realistic pressure |
Score every exercise against the RTO and RPO it claims to meet, and treat a miss as a finding with an owner and a date. Ready.gov’s exercise guidance is free and sufficient for the first year, while a maturity model gives the board a scale for progress.
We have watched restore tests expose backup jobs that had silently failed for months, and failover drills reveal DNS records nobody owned. Wire the plan into monitoring between exercises with recovery-focused indicators, such as backup success rates, replication lag, and restore-test age, so drift shows up on a dashboard before it shows up mid-disaster.
Common Disaster Recovery Plan Questions Practitioners Ask
What should a disaster recovery plan include?
Seven essentials: activation criteria with a named declaration authority, a verified contact tree, system-by-system recovery runbooks in priority order, RTO and RPO targets per system, pre-approved communication templates, return-to-normal and validation steps, and a version log. Keep copies offline and in a second location the disaster cannot touch.
What is the difference between a disaster recovery plan and a business continuity plan?
The disaster recovery plan restores technology: systems, data, and connectivity within defined recovery clocks. The business continuity plan keeps priority business operations running through any disruption, including by manual workaround while technology is down. Continuity is the umbrella program; disaster recovery is its technology chapter, and each depends on the other.
How often should you test a disaster recovery plan?
Walk through the plan quarterly, run restore tests quarterly on critical systems, tabletop twice a year, fail over tier-1 systems annually, and run a full-scale exercise every 12 to 24 months. Add an out-of-cycle test after any major system change, and score every exercise against its stated RTO.
Who should be on the disaster recovery plan team?
A named DR lead with declaration authority, system owners for each critical platform, network and infrastructure engineers, a security lead for attack scenarios, a communications owner, and an executive sponsor who controls spend. Name a deputy for every role, because disasters show no respect for vacation schedules.
How much does a disaster recovery plan cost to build?
The plan itself is mostly staff time: a quarter of part-time effort for a mid-size firm covering assessment, documentation, and first exercises. The strategy tier drives real cost, from low thousands yearly for backup-and-restore to six figures for active-active, which is why RTO-matched selection precedes any purchase.
What are RTO and RPO in a disaster recovery plan?
RTO, recovery time objective, is the maximum time a system can stay down before unacceptable damage, and it drives your strategy tier. RPO, recovery point objective, is the maximum data loss you can absorb, measured backward from the failure, and it drives backup and replication frequency. Set both per system.
Seven Traps That Derail Recovery Programs
DR programs rarely fail during the disaster; they fail quietly in the months before it, and the failure patterns repeat across every post-incident review we read. The table pairs the seven most common with corrections that fit inside normal operations.
|
Trap |
How it shows |
Correction |
|
Plan stored only in the systems it protects |
Nobody can open the runbook mid-outage |
Printed, offline, and second-region copies |
|
Backups never restore-tested |
Silent job failures for months |
Quarterly restore tests with integrity checks |
|
Company-wide RTO instead of per-system |
Budget overspent on some systems, exposed on others |
Tier the estate through the BIA |
|
Vendor dependencies unmapped |
A CDK-style provider outage has no play |
Dependency map with vendor recovery SLAs |
|
Exercises always pass |
Scenarios designed not to fail |
Score against clocks; misses become findings |
|
One person holds the knowledge |
The plan works only when they answer the phone |
Deputies named and rotated through drills |
|
Plan frozen since approval |
Systems changed, runbooks did not |
Maintenance triggers on every major change |
Emerging Threats Your Program Isn’t Ready For
Three exposures will stress DR programs through 2027. Ransomware crews now attack backups and strike twice, as CDK’s second-day hit showed, so immutability and staged restoration move from best practice to baseline, with CISA’s hardening guidance the free starting point for most teams.
Concentration is the second. SaaS and cloud providers have become single points of failure for entire industries, and Uptime’s data shows connectivity outages lengthening, so dependency mapping and provider-failure playbooks belong in every plan, from an operational risk lens as much as a technical one.
Third, evidence expectations keep rising. SEC disclosure clocks, insurer questionnaires, and customer resilience audits all now ask for tested recovery proof, and crisis governance reviews increasingly treat an unexercised plan as no plan, a standard McKinsey’s resilience work echoes at board level.
If you own technology recovery for a US mid-market firm, start with the eight steps and the program structure around them, or move faster with the policy and strategy templates. Then review our services and write to us through the contact page; bring your RTO list and we will find the gaps before an outage does.

Chris Ekai is a Risk Management expert with over 10 years of experience in the field. He has a Master’s(MSc) degree in Risk Management from University of Portsmouth and is a CPA and Finance professional. He currently works as a Content Manager at Risk Publishing, writing about Enterprise Risk Management, Business Continuity Management and Project Management.