The first hour of a crisis: stabilize, learn, communicate

When it’s on fire, the order of your moves matters more than any single move — contain first, then learn, then speak.

The idea

“Walk me through your next thirty minutes” isn’t asking for the root cause. It’s testing whether you can stay calm and sequence. The reliable loop: stop the bleed, split facts from unknowns, set a communication , and make the reversible calls now while parking the irreversible ones until the facts are in.

Two things move underneath every incident: the damage that accrues while the problem is uncontained, and the trust of everyone watching how you handle it. Speaking before you know, or promising an update and going quiet, spikes one and drains the other.

Sequence your first 30 minutes

Damage and trust over the first 30 minutes. Two lines respond to the order in which you take actions: damage rises while the incident is uncontained, and trust falls during silence or after a premature public statement. 100 75 50 25 0 0 10 20 30 min
damage (lower is better) trust (higher is better)
Peak exposure
—
Standing trust
—

Drag to reorder, or use the arrows. Each action takes a few minutes — you can only do one at a time.

The irreversible call: announce permanent refunds & purge the affected data
Reversible moves you can make now. This one you can’t take back — park it until the facts are confirmed.

How it works

  1. 1Contain. Stop the bleed before anything else — roll back, flip the flag, take the path offline. Damage accrues every minute the cause is live, so the fastest safe containment is worth more than a perfect diagnosis.
  2. 2Learn. Assign someone to fact-finding and keep two lists: what you know and what you don’t. You can only communicate honestly, and decide safely, from that split.
  3. 3Communicate on a cadence. Acknowledge early — even “we’re aware, investigating, next update in 15 minutes” buys calm. Then hit that time, every time, even if the update is “still working on it.”
  4. 4Decide reversibly. Make the reversible calls now; explicitly park the one-way doors (public blame, permanent deletion, legal statements) until the facts are confirmed.

A calm 30-minute schedule tends to look like this:

min  0 ── contain: roll back the bad deploy      (stop the bleed)
min  5 ── learn:   assign fact-finding           (facts vs unknowns)
min  9 ── comms:   internal note + "update in 15"(set the cadence)
min 12 ── align:   brief the exec on facts so far
min 15 ── comms:   public statement, once facts are in
        ── parked: refunds / data purge          (irreversible -> later)

External statements wait for facts; internal acknowledgement does not. That single distinction protects trust more than eloquence ever will.

When to use it

FitsThe trade-off
Outages, security incidents, a bad deploy, a data mix-up, a PR flare-up — any “it just broke, what now” prompt.This is a bias-to-action loop for the first hour, not a root-cause investigation. The post-mortem comes later — say so, don’t skip it.

Watch out for

Worked example

“A deploy at 2:00 starts double-charging customers on payment retries. Walk me through your next thirty minutes.” A strong answer sequences it: 2:00 — roll the deploy back immediately; the bleeding stops even before I know the exact cause. 2:05 — I put someone on fact-finding: how many customers, over what window, which charges. 2:09 — internal note to support and leadership: “retries double-charged between 1:50 and 2:04, rolled back, refund plan by 2:20, next update in 15.” 2:15 — once we’ve scoped it, a public status update and proactive refunds to the affected list. What I don’t do at 2:01 is email every customer an apology or purge the charge records — those wait for the facts.

Check yourself

A deploy is actively corrupting orders. What’s your first move?

You promised an update in 15 minutes. It’s minute 15 and you still don’t know the root cause. What do you do?

Prep Room · the interviewer is listening for order and composure, not a miracle fix.