Incident command: severity, roles & the first hour

Waking someone at 3 a.m. is cheap. A sev1 that sat for an hour because nobody was in charge is not.

The idea

When something breaks, two clocks start: how bad is it and who is doing what. The first is severity — a fast read of impact × scope × trend that decides who you page and how loudly. The second is structure: one incident commander who coordinates but doesn’t touch the keyboard, with operations, communications, and investigation running in parallel.

The classic failure isn’t a lack of effort. It’s five smart people all firefighting at once — colliding on the same change, nobody deciding, nobody telling customers anything. Structure is what turns effort into progress.

1 · classify severity fast

Read three dials. Severity should fall out in seconds — you can always revise it up.

SEV1

2 · the first hour — assign roles, then watch the clock

You have five responders for a sev1: checkout is failing for everyone. Give each a role, then run the hour. Watch how the three tracks move together — or don’t.

Ada
Ben
Cy
Dev
Eli
ops progress & sent updates colliding changes (setback) time cursor
Elapsed
0 min
Mitigation
0%
Updates sent
0
Timeline
—
Pick roles or tap a preset, then press run the hour.

How it works

Same five people, same broken checkout. The only variable is how you organise them. Here is the arithmetic the simulator runs each minute:

# fix rate = base x ops-effort x coordination
#   ops-effort: 1 person = 1.0, then +0.5 per extra hand (diminishing)
#   coordination: one commander = 1.0 ; no commander = 0.5 ; two = 0.7
# a commander also activates the prepared failover at min 5  (+14%)
# no commander + 2+ ops -> colliding changes knock -10% at min 12/24/36/48

Textbook split  (1 commander, 2 ops, 1 comms, 1 scribe)
  rate  = 1.55 x 1.5 x 1.0 = 2.3%/min ; +14% failover at min 5
  stable ~37 min ; 4 stakeholder updates ; timeline captured

Everyone firefights  (0 commander, 5 ops, 0 comms, 0 scribe)
  rate  = 1.55 x 3.0 x 0.5 = 2.3%/min   # same raw horsepower!
  minus collisions at 12/24/36/48, no failover
  still 99% at 60 min ; 0 updates ; no timeline

Notice the raw fix rate is identical — five uncoordinated hands equal two coordinated ones. Coordination doesn’t just prevent the collisions; it frees people to keep customers informed and to write down what happened.

When to use it

Stand up the full structure when…The trade-off
It’s a sev1/sev2 — broad impact, or worsening, or touching data/safety.Roles cost people. For a small, contained sev3 one owner and a ticket is enough.
More than one person is responding, or more than one team is involved.Structure adds a little ceremony up front — worth it past two responders.
Stakeholders (customers, execs, support) need updates while you fix.You spend one person on comms who isn’t fixing — almost always the right call.

Watch out for

Worked example

An interviewer asks: “Checkout starts throwing 500s. Walk me through your first ten minutes.” A strong answer moves in order. Classify: core feature broken × all users × worsening → sev1. Escalate: page the on-call and the eng manager now — “waking someone is cheap.” Structure: “I take incident commander and stay off the keyboard. One person on operations to run the rollback, one on comms posting to the status page and #incidents, one scribe timestamping everything.” Then contingency over improvisation: “First move is the prepared rollback of the 3:40 deploy, not a live debug of prod.” If the interviewer says “another engineer jumps in and starts restarting nodes,” you’d name the risk — that’s a colliding change — and route it through the IC.

Check yourself

Twenty minutes into a sev1, the fix is close but customers have heard nothing. You have four responders, all deep in the fix. What’s the disciplined move?

A teammate says “this is only hitting one region and it’s recovering — don’t bother anyone.” But impact is data being written wrong. Your read?