Waking someone at 3 a.m. is cheap. A sev1 that sat for an hour because nobody was in charge is not.
When something breaks, two clocks start: how bad is it and who is doing what. The first is severity — a fast read of impact × scope × trend that decides who you page and how loudly. The second is structure: one incident commander who coordinates but doesn’t touch the keyboard, with operations, communications, and investigation running in parallel.
The classic failure isn’t a lack of effort. It’s five smart people all firefighting at once — colliding on the same change, nobody deciding, nobody telling customers anything. Structure is what turns effort into progress.
1 · classify severity fast
Read three dials. Severity should fall out in seconds — you can always revise it up.
2 · the first hour — assign roles, then watch the clock
You have five responders for a sev1: checkout is failing for everyone. Give each a role, then run the hour. Watch how the three tracks move together — or don’t.
Same five people, same broken checkout. The only variable is how you organise them. Here is the arithmetic the simulator runs each minute:
# fix rate = base x ops-effort x coordination
# ops-effort: 1 person = 1.0, then +0.5 per extra hand (diminishing)
# coordination: one commander = 1.0 ; no commander = 0.5 ; two = 0.7
# a commander also activates the prepared failover at min 5 (+14%)
# no commander + 2+ ops -> colliding changes knock -10% at min 12/24/36/48
Textbook split (1 commander, 2 ops, 1 comms, 1 scribe)
rate = 1.55 x 1.5 x 1.0 = 2.3%/min ; +14% failover at min 5
stable ~37 min ; 4 stakeholder updates ; timeline captured
Everyone firefights (0 commander, 5 ops, 0 comms, 0 scribe)
rate = 1.55 x 3.0 x 0.5 = 2.3%/min # same raw horsepower!
minus collisions at 12/24/36/48, no failover
still 99% at 60 min ; 0 updates ; no timeline
Notice the raw fix rate is identical — five uncoordinated hands equal two coordinated ones. Coordination doesn’t just prevent the collisions; it frees people to keep customers informed and to write down what happened.
| Stand up the full structure when… | The trade-off |
|---|---|
| It’s a sev1/sev2 — broad impact, or worsening, or touching data/safety. | Roles cost people. For a small, contained sev3 one owner and a ticket is enough. |
| More than one person is responding, or more than one team is involved. | Structure adds a little ceremony up front — worth it past two responders. |
| Stakeholders (customers, execs, support) need updates while you fix. | You spend one person on comms who isn’t fixing — almost always the right call. |
An interviewer asks: “Checkout starts throwing 500s. Walk me through your first ten minutes.” A strong answer moves in order. Classify: core feature broken × all users × worsening → sev1. Escalate: page the on-call and the eng manager now — “waking someone is cheap.” Structure: “I take incident commander and stay off the keyboard. One person on operations to run the rollback, one on comms posting to the status page and #incidents, one scribe timestamping everything.” Then contingency over improvisation: “First move is the prepared rollback of the 3:40 deploy, not a live debug of prod.” If the interviewer says “another engineer jumps in and starts restarting nodes,” you’d name the risk — that’s a colliding change — and route it through the IC.
Check yourself
Twenty minutes into a sev1, the fix is close but customers have heard nothing. You have four responders, all deep in the fix. What’s the disciplined move?
A teammate says “this is only hitting one region and it’s recovering — don’t bother anyone.” But impact is data being written wrong. Your read?