The incident management process is the sequence an organization follows to detect, declare, triage, respond to, resolve, and review any event that disrupts normal operations. Six steps, four command roles, and a set of regulatory reporting clocks turn a scramble into a repeatable routine that restores service quickly and feeds lessons back into the risk register.
At 11:20 UTC on November 18, 2025, Cloudflare’s network started failing to deliver core traffic, and within eight minutes customers across the internet were seeing error pages. CEO Matthew Prince’s post-incident review, published the same day, called it the company’s worst outage since 2019.
The trigger was mundane. A database permission change at 11:05 made a query return duplicate rows, which doubled the size of a Bot Management feature file past a hard-coded limit of 200 features. The software that routes traffic panicked, and it took until 13:37 for engineers to pin the failure on that file.
| Incident Management Process: Key Takeaways |
| Cloudflare’s November 18, 2025 outage ran from a database permission change at 11:05 UTC to all-clear at 17:06, with the root cause found only at 13:37. |
| PagerDuty’s March 2026 survey of 1,000 leaders found 68% lose more than $300,000 per hour of incident, and 8% lose more than $1 million. |
| Uptime Institute’s May 2026 analysis: 57% say their last major outage cost over $100,000, and one in five paid more than $1 million. |
| NIST SP 800-61 Revision 3 (April 2025) folds incident response into CSF 2.0: Govern, Identify, and Protect prepare; Detect, Respond, and Recover execute. |
| Regulatory clocks now run from 4 hours (DORA) through 36 hours (US bank regulators) and 72 hours (CIRCIA) to four business days (SEC Form 8-K). |
| Six steps, four command roles, and three metrics (MTTD, MTTA, MTTR) are enough to make the incident management process auditable. |
A corrected file went out globally at 14:30, and every system was declared normal at 17:06. Nearly six hours, one small config file, and a lesson every risk practitioner can use: the quality of the incident management process decides how long a bad Tuesday lasts.
What the Incident Management Process Covers
Start with the definition the service-management world agrees on. ITIL 4 defines an incident as an unplanned interruption to a service, or a reduction in its quality, and the incident management practice as the work of restoring normal service as fast as possible. Risk practitioners widen that lens to safety, security, and reputation events, but the shape stays the same.
NIST rebuilt its own guidance around the same idea. SP 800-61 Revision 3, published April 3, 2025, retires the old four-phase handling guide and maps incident work onto the six CSF 2.0 functions. Govern, Identify, and Protect are preparation; Detect, Respond, and Recover are the live lifecycle.
| Term | What it means | Where it belongs |
| Incident | An unplanned event that interrupts or degrades a service, process, or safety condition | This process; the incident management workflow |
| Problem | The underlying cause of one or more incidents, worked after service is restored | Problem management and root cause analysis |
| Major incident | An incident whose scale triggers the full command structure and executive updates | The major incident management runbook |
| Crisis | An event threatening the organization’s survival, reputation, or licence to operate | Crisis management and the executive team |
| Disaster | Loss of a site, region, or platform that forces recovery elsewhere | Disaster recovery and continuity plans |
Those boundaries decide who is in the room. An incident that cannot be contained hands off to the major incident runbook, a threat to survival escalates into crisis management, and a lost site invokes the disaster recovery plan. The incident management process is the front door to all three.
The practitioner’s job is to make that front door consistent. The incident management workflow fixes the order of steps, the incident response plan template fixes the paperwork, and the incident management team supplies the people. None of the three works without the other two.
Why Process Quality Sets the Bill for an Outage
Definitions matter because the meter runs the whole time. PagerDuty’s 2026 State of AI-First Operations report, released March 17, 2026 from a survey of 1,000 director-level leaders across seven markets, found 68 percent lose more than $300,000 for every hour an incident continues. A third lose at least $500,000, and 8 percent lose more than $1 million.

Figure 1. Two-thirds of surveyed organizations lose more than $300,000 an hour while an incident runs, per PagerDuty.
The damage spreads beyond the invoice. The same survey has 52 percent citing brand damage, 50 percent recovery costs, 48 percent lost productivity, and 42 percent developer burnout. Katherine Calvert, PagerDuty’s chief marketing officer, summed it up: the financial risk of major incidents now makes operational resilience a board-level priority.
| Consequence of a slow incident management response | Share citing it | Which process step limits it |
| Revenue lost per hour above $300,000 | 68% | Detection and triage speed |
| Damaged brand reputation | 52% | Communication cadence during response |
| Recovery and remediation costs | 50% | Containment choices in the respond step |
| Reduced productivity | 48% | Clear roles, fewer people pulled in |
| Developer or responder burnout | 42% | Handoffs and post-incident review |
Uptime Institute’s Annual Outage Analysis 2026, released May 13, 2026, shows the same pattern from the data-center side. Fifty-seven percent of respondents say their most recent major outage cost more than $100,000, and one in five report a bill above $1 million for the second year running.
The airline lobby on October 29, 2025 made that arithmetic visible. An inadvertent configuration change in Azure Front Door took down Microsoft 365, Xbox Live, and the Alaska and Hawaiian Airlines websites for roughly eight hours. Alaska Airlines told guests to see an agent at the airport for a boarding pass and allow extra time in the lobby.
Microsoft’s own recovery followed the textbook order. Engineers blocked new customer configuration changes, rolled back to the last known good configuration, and told customers to fail over through Azure Traffic Manager.
That sequence, freeze, revert, communicate, is what a mature incident management process rehearses before it is needed, and operational risk examples like this one belong in every register.
The Six Steps of the Incident Management Process
With the stakes set, here is the routine itself. Competing guides count five, seven, or eight steps; the six below are the ones an auditor can test, because each ends with a dated artefact. The Cloudflare timeline in Figure 2 shows how they play out under real pressure.
| Step | What happens | Owner | Evidence it produced |
| 1. Detect | Monitoring, user reports, or a control breach flags an anomaly | Service desk, SOC, or line staff | Alert or ticket with timestamp |
| 2. Declare | Someone with authority confirms it is an incident and assigns severity | On-call lead or incident manager | Severity call logged with time |
| 3. Triage | Scope, affected services, and business impact are assessed | Incident manager with SMEs | Impact statement and priority |
| 4. Respond | Contain, work around, and communicate on a fixed cadence | Operations lead and comms lead | Status updates, change records |
| 5. Resolve | Restore normal service, verify, and close the ticket | Operations lead | Closure record and verification |
| 6. Review | Root cause, timeline, and corrective actions within 30 days | Incident manager and risk team | Post-incident review and register update |

Figure 2. Cloudflare’s November 18, 2025 outage: 2 hours 17 minutes from first errors to root cause, then 3.5 hours more to all-clear.
Detection and declaration are where the clock starts, and where most time leaks. Cloudflare’s engineers first suspected a hyper-scale DDoS attack because the company’s own status page, hosted elsewhere, failed at the same moment. The misdiagnosis cost two hours, which is why the declare step should record a working hypothesis and an explicit time to revisit it.
Severity is the decision that drives everything after. Atlassian’s severity scale is the most copied model: Sev 1 for a full outage or data loss, Sev 2 for major degradation with no workaround, Sev 3 for a partial issue with a workaround. The number describes current customer impact, never how hard the fix looks.
| Severity | Definition | Response target | Review required |
| Sev 1 | Full outage, data loss, safety event, or confirmed breach | Page on-call now; acknowledge in 15 minutes | Yes, within 5 business days |
| Sev 2 | Major degradation for many users, no clean workaround | Page on-call; acknowledge in 30 minutes | Yes, within 10 business days |
| Sev 3 | Partial or minor issue, workaround available | Business hours; respond in 2 to 4 hours | Optional, at manager’s call |
| Sev 4 | Cosmetic or single-user issue | Next sprint or queue | No |
The same incident management process holds outside IT. Suppose a payroll provider’s file transfer fails at 06:10 on the morning wages are due. Detection is the missing confirmation receipt, declaration is the finance on-call lead calling it Sev 2 at 06:25, and triage confirms 1,400 employees will not be paid on time without action.
Response means a manual bank file by 08:00 and a staff notice by 08:30, resolution means confirming credits landed, and the review asks why one transfer had no alert. Six steps, one page of incident management records, and a corrective action the crisis management software comparison would call table stakes.
Respond and resolve are the visible steps, and the incident management system post covers the MTTD, MTTA, and MTTR mechanics behind them. Review is the step organizations skip, and it is the only one that lowers the cost of the next incident. Essential incident management tools exist to automate the first five so people have energy left for the sixth.
Who Does What: Command Roles and Escalation Triggers
Steps need owners, and the owners need boundaries. Google’s site reliability engineers set out the cleanest model in the SRE book chapter on managing incidents. It names an incident commander who holds the overall state, an operations lead who is the only person changing systems, a communications lead, and a planning lead for handoffs and logistics.
| Role | Owns | Never does |
| Incident commander | Overall picture, priorities, decisions, delegation | Hands-on fixes while commanding |
| Operations lead | Every change to production during the incident | Public or executive communication |
| Communications lead | Status cadence to customers, executives, regulators | Speculating on root cause |
| Planning lead | Shift handoffs, records, rollback list, follow-up tickets | Blocking the operations lead |
Two SRE rules do most of the work. The first is recursive separation of responsibilities: when one role is overloaded, split it rather than let it blur. The second is an explicit handoff, spoken aloud and logged, so that no incident ever has two commanders or none at all.
Escalation is the moment the incident management process stops being a service-desk routine and becomes a major incident. Write the triggers down in advance, because judgment degrades under pressure. The list below is the one we use when designing runbooks for clients and it holds up across IT, safety, and security events:
- Any Sev 1 declaration, or a Sev 2 that passes 60 minutes without a confirmed workaround
- Customer-facing impact for more than 500 users, or any impact to a regulated service
- A safety injury, a confirmed data breach, or evidence of ransomware
- Media, regulator, or major-customer enquiry about the event
- The incident commander asks for it, no justification required
Escalation also changes the audience. Executives want impact, timeline, and decision points; regulators want the facts in a prescribed format; customers want honesty on a schedule. Effective incident management teams pre-write those three message templates so the communications lead fills in blanks instead of drafting from scratch at 2 a.m.
Reporting Clocks the Incident Management Process Must Beat
Roles handle the inside of the incident; regulators now dictate the outside. The tightest clock in force is the EU’s Digital Operational Resilience Act. Its Article 19 requires financial entities to send an initial notification within four hours of classifying an ICT incident as major, an intermediate report within 72 hours, and a final report within a month.

Figure 3. Reporting deadlines run from four hours under DORA to four business days for a material SEC filing.
US rules are looser but multiplying. Federally supervised banks have had a 36-hour notification requirement since the Federal Reserve, OCC, and FDIC finalized it in November 2021. The clock runs from the moment a bank determines that a notification incident has occurred, which makes the declare step a compliance event.
Public companies must file a Form 8-K Item 1.05 within four business days of deciding a cyber incident is material. Debevoise counted 29 Item 1.05 filers and 50 voluntary Item 8.01 filers in the rule’s first two years, and most of those 8.01 disclosures never became a 1.05 filing.
| Regime | Clock | Starts when | Who it covers |
| DORA Art. 19 (EU) | 4 hours, then 72 hours, then 1 month | Incident classified as major | EU financial entities and ICT providers |
| Bank notification rule (US) | 36 hours | Bank determines a notification incident occurred | FRB, OCC, and FDIC supervised banks |
| CIRCIA (US) | 72 hours; 24 hours for ransom payments | Reasonable belief a covered incident occurred | 16 critical infrastructure sectors |
| SEC Form 8-K Item 1.05 (US) | 4 business days | Materiality determination | SEC registrants |
The next US clock is close. CISA’s CIRCIA rule will require covered critical-infrastructure entities to report substantial cyber incidents within 72 hours and ransom payments within 24.
CISA missed its October 2025 statutory deadline and, per Hunton’s July 2026 note, now targets a final rule in September 2026, though the statutory clocks cannot be softened by rulemaking.
Every one of these clocks starts at a classification decision, which puts the declare step of the incident management process on the compliance critical path. Our DORA incident classification guide walks through the thresholds, and the compliance requirements primer explains how to map each regime to an owner. A four-hour clock nobody starts is a breach in waiting.
Metrics That Prove the Process Works
Clocks measure the outside; three timings measure the inside. Mean time to detect, acknowledge, and resolve are the incident management metrics boards understand, and each maps to a step. Uptime Institute also reports that failure to follow established procedures now leads human-error outages, which makes procedure adherence a fourth metric worth tracking.
| Incident management metric | How to calculate it | Process step | Practical target |
| MTTD | Time from first symptom to detection, averaged per quarter | Detect | Under 5 minutes for monitored services |
| MTTA | Time from alert to a named human acknowledging | Declare | Under 15 minutes for Sev 1 and 2 |
| MTTR | Total downtime divided by number of incidents | Respond and resolve | Falling quarter on quarter |
| Repeat rate | Share of incidents whose root cause was seen before | Review | Under 10 percent |
| Review completion | Sev 1 and 2 reviews closed within 30 days | Review | 100 percent |
| Procedure adherence | Runbook steps followed as written, sampled from reviews | All | Above 95 percent |
Treat these as key risk indicators, with thresholds and an owner, and they belong on the quarterly risk report next to the risk register. A rising repeat rate is the clearest signal that reviews are being written but not acted on.
Trend the numbers against the business impact analysis, because a 90-minute MTTR is fine for an internal wiki and catastrophic for a payments gateway with a 15-minute recovery objective. Report the gap between the two, per critical service. That single table is the most useful page in most quarterly packs:
- Recovery time objective from the BIA, per critical service
- Actual MTTR for that service over the last four quarters
- Number of Sev 1 and Sev 2 incidents, with the share reviewed on time
- Open corrective actions older than 90 days, with owners named
Post-Incident Review: What Cloudflare Did Next
Metrics show where the process leaks; the review step is where it gets fixed. Cloudflare’s Code Orange: Fail Small plan, published December 19, 2025 by Dane Knecht, shows a review turning into controls. It followed the November 18 outage and a second event on December 5 that hit 28 percent of applications for about 25 minutes.
| Lesson from the outages | Control Cloudflare committed to | Transferable to any organization as |
| A config file propagated globally in minutes | Health Mediated Deployment for all production configuration | Staged rollout for changes, not just code |
| One oversized file crashed the core proxy | Failure handling designed for every interface | Graceful degradation, tested in drills |
| Responders locked out of their own tools | Break-glass emergency access with safeguards | Out-of-band access to runbooks and comms |
| Status page failed with the platform | Independent hosting for status communication | Communication channel with no shared dependency |
Knecht’s framing is the one to borrow: assume failure will occur between each interface and handle it in the most reasonable way possible. That sentence is a design principle, a test case, and a review question at once. Reviews that produce sentences like it are worth the meeting.
A blameless review needs a fixed agenda: the timeline in the responders’ own words, the decision points, what surprised people, and at most five corrective actions with owners and dates. Uptime’s finding that procedural failures drive most human-error outages adds one more question for every review: was the runbook followed, and if not, why was it easier to ignore?
Close the loop into the wider risk management lifecycle. Each corrective action becomes a control in the register, each near-miss adjusts a likelihood score, and each review feeds the next tabletop exercise. That loop is what separates a risk culture that learns from one that only apologizes.
Frequently Asked Questions About the Incident Management Process
What is the incident management process in simple terms?
The incident management process is the agreed routine for handling anything that interrupts normal operations: detect it, declare it with a severity, triage the impact, respond and communicate, restore service, then review what happened. ITIL 4 defines the goal as restoring normal service as fast as possible while limiting business impact.
What are the steps of the incident management process?
Six steps cover the incident management process end to end: detect, declare, triage, respond, resolve, and review. Some frameworks split response into containment and recovery or fold declaration into detection, which produces five- to eight-step versions. The count matters less than each step ending with a timestamped record someone owns.
How does the incident management process differ from incident response?
Incident response is the respond step inside the wider incident management process, and in cybersecurity it often carries its own plan under NIST SP 800-61. Incident management adds the detection and declaration before response and the resolution and review after it, plus the governance that keeps all six steps consistent.
Who owns the incident management process?
A named process owner, usually the head of IT operations or the operational risk lead, owns the design and the metrics. Individual incidents are owned by an incident commander appointed at declaration, supported by operations, communications, and planning leads. The risk function owns the review step’s link into the register.
Which metrics show an incident management process is working?
Track mean time to detect, mean time to acknowledge, and mean time to resolve for each critical service, then compare MTTR with the recovery time objective from the business impact analysis. Add repeat rate and on-time review completion. Falling MTTR with a stable repeat rate means reviews are producing real controls.
How often should the incident management process be tested?
Run a tabletop exercise at least twice a year and after any material change to systems, suppliers, or reporting rules. Test the four-hour DORA notification or the 36-hour bank notification clock as part of the drill, because the classification step is where the regulatory clock starts and where rehearsal pays off most.
Where Incident Programs Break Down
The failures below come from post-incident reviews we have read or written, and each one has a cheap fix that survives contact with a real outage. Most are governance gaps rather than technical ones, which is why they persist in organizations with excellent tooling.
| Failure pattern | What it costs | Fix |
| Nobody empowered to declare | Hours lost debating severity while the meter runs | Named on-call authority with a written declare threshold |
| Commander also fixing the system | Lost overall picture, missed communications | Enforce the SRE role split; commanders do not type |
| Status page shares platform dependencies | Customers learn from social media instead | Host status communication independently |
| Reviews written, actions unassigned | Repeat incidents, rising repeat rate | Maximum five actions, each with owner and date |
| Regulatory clocks unmapped | Missed 4-hour, 36-hour, or 72-hour notices | Clock table per regime inside the runbook |
| Runbooks untested since the last reorg | Procedures ignored under pressure | Tabletop twice a year, with the real on-call roster |
The Road to 2027: Faster Clocks, Smarter Tooling
September 2026 is the month to circle. If CISA lands the CIRCIA final rule on its stated timetable, a 72-hour reporting duty arrives for 16 critical-infrastructure sectors. The declare step then becomes a legal act for hospitals, utilities, and water systems that never ran a formal incident management process before.
AI is moving into the response step faster than into the review step. PagerDuty reports 59 percent of organizations already use AI in digital operations, and 75 percent of adopters saw resilience improve against 66 percent of non-adopters. Expect gains in MTTD and MTTA first, while root cause analysis stays a human job for now.

Figure 4. Outage frequency is falling for a fifth year, but the cost of the ones that still happen is not, per Uptime Institute.
Concentration is the trend to plan against. Cloudflare and Azure each took a measurable share of the internet down in a single quarter, and neither event was an attack. Put shared-dependency scenarios into the next business continuity plan review, alongside the cybersecurity risk scenarios that already get the attention.
Organizations that want their incident management process designed, tested, or benchmarked against ISO 22301 and NIST SP 800-61r3 can hand the work to us. We draft the runbook, the severity matrix, and the regulatory clock table, then run the first tabletop with your on-call roster.
Compare the options on our services page, or send a short brief through the contact form and expect a scoped reply within five working days. The runbook you rehearse this quarter is the one that shortens the next bad Tuesday on your calendar.

Chris Ekai is a Risk Management expert with over 10 years of experience in the field. He has a Master’s(MSc) degree in Risk Management from University of Portsmouth and is a CPA and Finance professional. He currently works as a Content Manager at Risk Publishing, writing about Enterprise Risk Management, Business Continuity Management and Project Management.