At 11:48 PM Pacific on 19 October 2025, a latent race condition in the DynamoDB DNS automation wrote an empty record for the US-EAST-1 regional endpoint, and Amazon Web Services did not declare the event closed until 2:20 PM the following afternoon. The company’s own post-event summary puts the disruption at 14 hours and 32 minutes.

Snapchat, Venmo, Roblox, Ring and United Airlines booking all degraded behind that single record, and network telemetry from ThousandEyes traced the cascade through more than a dozen dependent services, among them EC2, Lambda and Network Load Balancer. Business resilience is the discipline that decides whether a morning like that one is a costly inconvenience or an existential event.

Business Resilience: Key Takeaways
Business resilience is the ability to deliver critical operations through disruption from any hazard and restore normal service inside a tolerance the board has approved, which is the definition the OCC, the Federal Reserve and the FDIC adopted in 2020.
The AWS US-EAST-1 disruption of 19 to 20 October 2025 ran 14 hours and 32 minutes from a single DNS race condition, which is longer than the recovery time objective most firms had written down for their cloud-hosted services.
Uptime Institute found 57 percent of operators put their most recent major outage above $100,000 and one in five above $1 million, so business resilience gaps price themselves quickly.
ISO 22301 supplies the certifiable management system, ISO 22316 supplies the organizational attributes, and NIST SP 800-34 supplies the technical contingency planning. A business resilience program needs all three layers, not one.
The United States absorbed 23 billion-dollar weather and climate disasters in 2025 costing $115 billion, with the January Los Angeles wildfires alone at $61.2 billion, so physical business resilience still deserves budget.
Business resilience becomes measurable only when impact tolerances are set per service, tested against severe but plausible scenarios, and reported to the board with the same discipline as capital or liquidity.

Business resilience is the capacity to keep delivering critical operations through disruption from any hazard, then restore normal service inside a tolerance the board has formally approved. Everything else in the field is machinery for making that one sentence testable rather than aspirational, and most programs never get there.

What Business Resilience Means, and What It Does Not

The definition worth using comes from supervisors rather than vendors. In October 2020 the OCC, the Federal Reserve and the FDIC jointly issued Sound Practices to Strengthen Operational Resilience, describing it as the ability to deliver operations through a disruption from any hazard.

That supervisory phrasing matters for two reasons. It is hazard-agnostic, so a single business resilience program covers cyber attacks, severe weather, vendor failure and pandemic under one roof, and it is outcome-based, judging the firm on service delivered to customers rather than on paperwork completed on time.

Business Resilience Compared With Continuity and Disaster Recovery

Practitioners still trip over the vocabulary, and boards notice. The cleanest way to hold the three apart is by what each one owns and what it produces, which is also how the difference between continuity and disaster recovery gets settled in most audit findings we review.

Discipline Primary question Owns Typical artifact
Business resilience Can the firm keep serving customers through any severe disruption? The end-to-end service, including people, premises, technology and third parties Impact tolerances per critical service, tested against severe but plausible scenarios
Business continuity management How do we run the business while normal operations are unavailable? Workarounds, alternate sites, manual processes and recovery sequencing Business continuity plan, business impact analysis, exercise reports
Disaster recovery How do we restore the technology that failed? Infrastructure, data replication, failover and backup restoration Runbooks, RTO and RPO records, failover test evidence
Crisis management Who decides and who speaks while the event is live? Command structure, escalation, stakeholder and regulatory communication Crisis playbook, call tree, holding statements, incident log

Business resilience sits above the other three disciplines and consumes their outputs rather than replacing them. A firm can hold a certified business continuity management system and still fail a business resilience test, because certification proves the process exists while resilience proves the customer-facing service actually survives.

The Four Capabilities Inside Business Resilience

The original four-part model of prevention, protection, response and recovery still holds, though the modern version adds adaptation. ISO 22316 frames resilience as the ability to absorb and adapt in a changing environment, which is a higher bar than bouncing back to yesterday’s operating model.

Capability What it does Evidence that it works
Anticipate Horizon scanning, scenario analysis and risk assessment that name the disruptions worth planning for A scenario library refreshed annually, with at least one severe but plausible scenario per critical service
Protect Controls, redundancy, segregation and diversification that reduce the chance or blast radius of failure Second supplier or second region live and tested, not merely contracted
Respond Command, escalation, decision rights and communication while the event is unfolding Timed exercise results showing declaration within the agreed threshold
Recover Restoration of technology, data and process to an agreed service level Failover evidence against recovery time and recovery point objectives
Adapt Post-incident change that alters design, contracts or capacity rather than just the plan document Board-tracked actions closed with design changes, not with training refreshers

Why Business Resilience Became a Board Agenda Item

Three forces moved business resilience from the facilities budget to the board pack in under a decade. The first is concentration, because a single vendor defect now propagates across entire sectors within minutes rather than days, and no contract clause slows that down.

The 19 July 2024 CrowdStrike Falcon update is the reference case. Parametrix estimated $5.4 billion in direct losses for US Fortune 500 companies alone, with cyber insurance covering only 10 to 20 percent, leaving the balance on shareholders and customers.

Business resilience chart showing healthcare and banking absorbed more than half the Fortune 500 loss from a single vendor update

Figure 1. Healthcare and banking absorbed more than half the Fortune 500 loss from a single vendor update.

The second force is physical, and it has not softened with time. Climate Central recorded 23 billion-dollar weather and climate disasters across the United States in 2025, costing $115 billion and 276 lives, with the January Los Angeles wildfires alone reaching $61.2 billion in damage.

That followed 28 events costing $92.9 billion in 2023 and 27 events costing $182.7 billion in 2024, the fourth-costliest year on record. Business resilience budgets built on the assumption that insurance has already solved physical risk are reading the wrong decade of data.

Business resilience data on US billion-dollar climate disaster counts and damage costs from 2023 to 2025

Figure 2. Event counts eased slightly while the damage bill stayed above $90 billion every year.

The third force is disclosure. Since the SEC adopted its cybersecurity incident disclosure rule in July 2023, material incidents must reach investors on Form 8-K within four business days, which converts a resilience failure into a securities filing with almost no time to compose one.

The Standards and Rules That Define Business Resilience

Nobody needs to invent a business resilience framework from scratch, and firms that try usually rebuild ISO badly. The useful move is to map each layer of the program to the standard or rule that already governs it, then hold the map current.

ISO Standards Behind Business Resilience

Three ISO documents carry most of the weight inside a business resilience program. ISO 22301:2019 is the certifiable management system for business continuity, ISO 22316:2017 sets out organizational resilience principles and attributes, and ISO 31000:2018 supplies the risk management architecture sitting underneath both of them.

Technology teams add two more. NIST SP 800-34 Revision 1 governs contingency planning for federal information systems and travels well into private practice, while the recover function inside NIST Cybersecurity Framework 2.0 gives cyber resilience a common language with the rest of the business.

Our practical guidance is to certify against ISO 22301 only when a customer contract or a regulator asks for the certificate. Otherwise treat the ISO 22301 requirements as a business resilience design checklist, and spend the audit budget on exercising the plans instead.

US Supervisory Expectations for Business Resilience

Regulated firms face a denser set of rules, and the supervisory drafting is far more specific than most internal policies. The table below maps the sources examiners actually cite when they challenge a business resilience program, together with who each one binds.

Source Who it binds What it expects
OCC, FRB and FDIC Sound Practices, October 2020 Banks above $250 billion in assets, or above $100 billion with significant cross-jurisdictional activity Governance, scenario-informed risk management, secure and resilient systems, surveillance and reporting
FFIEC Business Continuity Management booklet All federally supervised financial institutions and their service providers Enterprise-wide resilience process covering third parties, testing and board reporting
Basel Committee Principles for Operational Resilience, March 2021 Internationally active banks Tolerance for disruption on critical operations, mapped interconnections, scenario testing
SEC cybersecurity disclosure rule, effective December 2023 SEC registrants Form 8-K Item 1.05 filing within four business days of a materiality determination
EU Digital Operational Resilience Act, applying 17 January 2025 US firms with regulated EU financial entities or ICT contracts ICT risk management, incident reporting, threat-led penetration testing, third-party register
CISA critical infrastructure guidance Owners and operators across the 16 critical infrastructure sectors Continuity of the national critical functions the organization supports

Two of those deserve a second look before any firm assumes they do not apply to it. The FFIEC booklet reaches service providers, and the Basel principles have shaped supervisory business resilience language well beyond the banks they formally bind.

Firms with European subsidiaries also inherit the Digital Operational Resilience Act, which has applied since 17 January 2025 and reaches their US-based technology contracts. We wrote a side-by-side comparison of DORA and NIS2 for business resilience teams working out which regime bites first.

How to Build a Business Resilience Program in Six Steps

A business resilience program is a sequence, not a document set, and the order matters because each step consumes the output of the one before it. Skipping the second step is the most common reason the fifth step produces theatre.

Step What you do What proves it is finished
1. Scope Name the critical services a customer or regulator would miss within 24 hours, in business language rather than system names A signed list of critical services with a named accountable executive per service
2. Map Trace each critical service end to end across people, process, technology, premises, data and third parties A dependency map that identifies the single points of failure by name
3. Set tolerance Agree the maximum tolerable disruption per service, expressed in time, volume or customer harm Board-approved impact tolerances stated as numbers, not as adjectives
4. Close gaps Fund redundancy, alternate suppliers, manual workarounds and recovery capability where the map shows exposure Funded remediation plan with dates, owners and a residual position accepted in writing
5. Test Exercise against severe but plausible scenarios that assume the primary control has already failed Exercise reports showing whether tolerance was breached, plus timestamps
6. Report and adapt Report resilience position to the board on a fixed cycle and change design where tests failed Quarterly board reporting and closed actions that changed architecture or contracts

Step two is where programs quietly break. Most firms can name their critical services but cannot trace them past the first vendor, which is why a third-party risk management framework and a resilience program have to be built by the same team.

Step three separates real programs from paper ones. An impact tolerance says something like four hours for card authorization or one business day for claims intake, and a risk appetite statement that cannot produce those numbers is not yet operational.

The supporting documents follow from that sequence rather than leading it, which is the reverse of how most programs begin. Teams starting from zero can lift structure from a business continuity policy template and a continuity strategy template, then adapt the content rather than adopt it wholesale.

Metrics That Prove Business Resilience to a Skeptical Board

Business resilience reporting fails when it counts activity rather than capability. A board learns nothing from the number of plans updated last quarter, and a great deal from whether the payments service recovered inside its stated tolerance during the most recent unannounced test.

Metric Definition Why boards should ask for it
Recovery time objective Target elapsed time to restore a service after disruption Sets the engineering standard, though it means little until tested under load
Recovery point objective Maximum tolerable data loss measured in time Exposes replication design and the real cost of cheap backup tiers
Impact tolerance Maximum disruption the firm will accept before customer or market harm becomes unacceptable The only metric expressed from the customer side rather than the system side
Tolerance breach rate Share of exercises and live incidents where tolerance was exceeded Shows whether stated capability survives contact with a scenario
Mean time to declare Elapsed time from first alert to formal incident declaration Declaration lag is often larger than technical recovery time
Single points of failure closed Count of mapped dependencies with no tested alternative, trending Turns the dependency map into a funded work programme

Recovery objectives deserve particular scrutiny because they are so often aspirational rather than evidenced. Our walkthrough on setting and validating RTO and RPO covers the arithmetic, and the honest business resilience test is whether that objective has ever been met during an unannounced exercise.

Cost gives the business resilience argument its teeth in a budget meeting. Uptime Institute’s Annual Outage Analysis 2026 reports that 57 percent of operators put their most recent major outage above $100,000, and that one in five put it above $1 million.

Business resilience cost chart showing one in five operators priced their last major outage above $1 million

Figure 3. One in five operators priced their last major outage above $1 million.

Cyber incidents run higher still. IBM’s Cost of a Data Breach research put the United States average at a record $10.22 million, roughly 2.3 times the global figure of $4.44 million, which reframes resilience spending as loss avoidance rather than overhead.

Indicators keep the picture current between formal reports. Firms already running operational risk key risk indicators can extend the same set to business resilience by adding declaration lag, failover test age and supplier concentration to the dashboard the committee already reads.

Business Resilience Across Four Families of Disruption

Hazard-agnostic does not mean hazard-blind. A business resilience program still needs different playbooks for the four families that account for almost every material incident we see, because the failure mechanics differ sharply even when the recovery goal stays the same.

Family Typical trigger Resilience lever that works Lever that disappoints
Technology and cyber Ransomware, faulty update, cloud region failure Tested failover to a second region or provider, immutable backups, rehearsed manual mode Insurance alone, or a recovery plan never run end to end
Physical and climate Wildfire, hurricane, flood, extended power loss Site diversity, remote-capable processes, pre-agreed logistics and mutual aid A single alternate site inside the same hazard zone
Third party and concentration Vendor outage, insolvency, sanctioned counterparty Substitutability testing, exit plans, contractual resilience clauses with evidence rights Questionnaires answered once a year and never verified
Supply chain and market Component shortage, tariff shift, logistics disruption Dual sourcing, buffer inventory sized to lead time, mapped tier-two suppliers A supplier list mistaken for a supply chain map

Business resilience view of disruption costs across cloud, climate and cyber between 2024 and 2026

Figure 4. The price of disruption across cloud, climate and cyber between 2024 and 2026.

Concentration is the family that has changed most since this article first ran in 2022. Concentration risk in third-party relationships now behaves like systemic risk, because thousands of firms share the same handful of cloud regions, endpoint security agents and payment rails.

Supply chain resilience rewards mapping over paperwork. Teams using ISO 28000 for supply chain security management tend to find their exposure sits two tiers below the suppliers they contract with, which is exactly where supply chain key risk indicators earn their reporting slot.

Climate exposure needs its own arithmetic rather than a paragraph buried inside the continuity plan. A structured climate risk assessment tells you whether your alternate site shares a hazard zone with the primary one, which is the first question an insurer or an examiner asks.

Cyber sits closest to the board because the disclosure clock runs fastest there. Pairing an incident response plan with a ransomware-specific business impact analysis keeps the technical and commercial responses from diverging under pressure, which is where business resilience programs most often lose control.

Where Business Resilience Programs Stall

Business resilience programs rarely collapse outright. They stall in a handful of recognizable places, and the table below collects the seven patterns we meet most often in review work, each paired with the correction that costs least to apply and least to defend in front of a committee.

Pitfall Root cause Remedy
Plans exist for systems, not for services The program was built by IT and inherited an application inventory Rebuild the scope around customer-facing services and name an executive owner for each
Impact tolerances written as adjectives Nobody wanted to defend a number in front of the board Force a time or volume figure per service, then let testing correct it
Exercises assume the control works Scenario design starts from the plan rather than from the failure Open every exercise with the primary control already unavailable
Third parties assessed once at onboarding Procurement owns the vendor file and risk owns nothing after signature Move to continuous monitoring with contractual evidence rights and exit testing
Recovery objectives never validated Failover testing is scheduled, announced and run in daylight Run at least one unannounced failover per critical service each year
Post-incident actions close as training Root causes are recorded as human error rather than design weakness Require an architecture, contract or capacity change for any tolerance breach
Resilience reported only after incidents No fixed reporting cycle exists in the governance calendar Put a standing quarterly resilience item in the risk committee agenda

The pattern behind most of these stalls is simple optimism about untested capability. PwC’s global crisis and resilience research found leaders consistently overestimate their readiness, and our own business resilience review findings match that gap almost every time we score a program against evidence.

Scoring helps because it converts opinion into a trend the committee can actually track. A business continuity maturity model gives them a business resilience number that can move, and a library of exercise scenarios removes the excuse that designing a good test takes too long.

Common Business Resilience Questions Practitioners Ask

What is business resilience in simple terms?

Business resilience is the ability of an organization to keep delivering its critical services through a serious disruption and to restore normal operations within an agreed time. It covers cyber attacks, weather events, vendor failures and pandemics under a single hazard-agnostic approach rather than separate plans.

How is business resilience different from business continuity?

Business continuity is one component of business resilience. Continuity answers how the firm operates while normal processes are unavailable, whereas resilience judges whether the customer-facing service survived at all and holds accountability for the end-to-end chain including third parties and premises.

Who owns business resilience inside an organization?

Accountability belongs to a named executive for each critical service, with a central resilience function coordinating standards, testing and reporting. Boards retain oversight and approve impact tolerances, while the enterprise risk management function integrates resilience into the wider risk profile.

What standards should a business resilience program follow?

Start with ISO 22301 for the management system, ISO 22316 for organizational attributes and ISO 31000 for risk architecture. Technology teams add NIST SP 800-34 for contingency planning and NIST Cybersecurity Framework 2.0 for cyber recovery, while regulated firms overlay their supervisory rules.

How often should business resilience plans be tested?

Test each critical service at least annually, with severe but plausible scenarios and at least one unannounced exercise per year for the highest-priority services. Any material change to architecture, suppliers or premises should trigger a further test rather than a document review.

What does a business resilience program cost?

Cost scales with the number of critical services and the tolerance you set, since tighter tolerances demand duplicated capacity. Compare that spend against the Uptime Institute finding that 57 percent of major outages exceed $100,000, and the business case usually settles itself.

Can small businesses build meaningful business resilience?

Yes, and the sequence is identical at a smaller scale. Federal guidance for smaller firms at Ready.gov’s business preparedness resources covers assessment, planning and testing, and a two-page plan that has been exercised beats a fifty-page plan that has not.

Where Business Resilience Is Heading Through 2028

Concentration will keep setting the business resilience agenda for the rest of this decade. The World Economic Forum’s Global Risks Report 2026 puts geoeconomic confrontation at the top of its two-year outlook, and half of respondents expect the period to be turbulent or stormy, which points at supply routes and vendor availability rather than at weather alone.

Expect supervisors to keep converging on impact tolerance as the organizing metric. The Basel principles and the European rulebook already speak that language, and US examiners increasingly ask for the number rather than the plan when they test operational risk management in banking.

Watch the business resilience tooling market consolidate around evidence rather than document storage. Buyers now compare business continuity management software, operational resilience platforms and crisis management tools on whether they can produce dependency maps and test evidence on demand, not on how many templates ship in the box.

Build the capability that regulators have not asked for yet. Deloitte’s organizational resilience research and CISA’s critical infrastructure guidance both point the same direction, toward firms that can reconfigure operations mid-disruption instead of waiting to restore the version that failed.

 

Test Your Business Resilience Before the Next Outage Does

Operations leaders in banking, healthcare and logistics come to us with a plan library and no evidence. We map critical services, set defensible impact tolerances and run the unannounced test that tells you the truth. Review our advisory services, then start a conversation about your riskiest service.

Teams that would rather start alone, without an adviser in the room, can begin with a business impact analysis, build out the business continuity plan behind it, and check the result against banking continuity planning practice or an enterprise risk management framework for governance fit. Business resilience earns nothing on paper until somebody sets out to break it.

One last note for anyone weighing where to start. Pick the single service whose failure would put your firm in a regulatory filing, map it end to end this quarter, and let that one exercise set the standard for operational risk management across everything else you own.

Index