02 — Capability · BC-740.20
Incident & Problem Management
Restore service quickly when it breaks, coordinate the response when it breaks badly, and find and remove the causes so the same failure is not restored twice.
- Incident Management
- Major Incident
- Problem Management
In scope
- Incident logging with categorization and routing
- Major incident coordination and communication
- Root cause investigation and known errors
- Trend analysis across incidents
Out of scope
- Security incident response (see BC-760)
- Crisis management beyond technology (see BC-830)
Realized by · 2
Used in · 4
Build it · 1
Decomposes into · 4
- BC-740.20.10Incident ResolutionCapture, prioritize, route and resolve service disruptions against agreed targets, keeping the user informed along the way.
- BC-740.20.20Major Incident CoordinationRun the bridge, the communications and the decisions during a high-impact outage, then close with a timeline everyone can learn from.
- BC-740.20.30Problem InvestigationTrace recurring or significant incidents to their root cause and drive the fix through to the team that owns it.
- BC-740.20.40Known Error ManagementDocument causes that cannot be fixed yet, with their workarounds, so the service desk can resolve the symptom in minutes.