02 — Capability · BC-1230
Site Reliability
Keep the product available and performant to the level promised to customers, detect when it is not before they do, restore it quickly, and learn from every failure. This is reliability of the product the company sells, not of the systems employees use.
- SRE
- Production Operations
- Service Reliability
- Product Operations Engineering
In scope
- Service level objectives, error budgets and the decisions they drive
- Observability of production — metrics, logs, traces and alerting
- Incident response for the product, on-call, and blameless postmortems
- Resilience engineering, capacity under load and disaster recovery for the product
Out of scope
- Incident and problem management for internal IT services (see BC-740)
- Corporate infrastructure operations (see BC-730)
- Enterprise-wide business continuity planning (see BC-830)
Realized by · 0
- No product in the catalog yet.
Used in · 2
- Commit-to-Production · 04 Release Progressively via BC-1230.20
- Commit-to-Production · 05 Operate & Observe via BC-1230.10
Build it · 16
- pattern Circuit Breaker
- pattern Retry with Backoff
- pattern Bulkhead via BC-1230.40
- pattern Throttling via BC-1230.40
- pattern Canary Release via BC-1230.20.20
- pattern Distributed Tracing via BC-1230.20.10
- pattern Structured Logging via BC-1230.20.10
- pattern SLO and Error Budget via BC-1230.10
- stack Open observability stack via BC-1230.20
- book Release It!
- book Building Microservices via BC-1230.40
- book The Art of Scalability via BC-1230.40
- book Observability Engineering via BC-1230.20
- book Site Reliability Engineering via BC-1230.10
- book Chaos Engineering via BC-1230.40
- book Building Secure and Reliable Systems via BC-1230.50
Decomposes into · 5
- BC-1230.10Service Level ManagementDefine what reliable means for each product surface in measurable terms, agree how much unreliability is acceptable, and let that budget decide when to ship and when to harden.
- BC-1230.20ObservabilitySee what production is doing from the outside in — metrics, logs, traces and synthetic checks — well enough to answer new questions during an incident without shipping new code.
- BC-1230.30Incident ResponseDetect, declare, coordinate and resolve production incidents with clear roles and timely customer communication, then learn from them without assigning blame.
- BC-1230.40Resilience & Capacity EngineeringDesign and prove that the product survives component failure, traffic spikes and regional loss, and that it has the headroom for next quarter's customers.
- BC-1230.50Operational ReadinessMake sure a service is ready to be run before it is run — owned, monitored, documented and recoverable — and keep it that way through reviews and runbooks.