The goal isn’t to find who broke it — it’s to learn what let it break, and change that.
People make mistakes; systems decide how much those mistakes cost. A blameless review starts from facts and a timeline, describes what happened in systems language — “the deploy lacked a guard,” not “Sam forgot” — and ends in action items that are specific, owned, and dated. Blame makes people hide the next problem. Systems language makes it safe to tell the truth, which is the only way you find the real cause.
Interactive · the postmortem editor
A draft has blame baked in. Rewrite each line into systems language and watch how safe the room feels to keep telling the truth.
0 of 4 lines rewritten
Now the action items · make one real
“Improve monitoring” is a wish, not an action item. Toggle each piece until it’s specific, owned, and dated.
Not yet an action item — it’s missing all three.
Three moves turn a blame session into a review that changes something:
1. Facts + timeline first — what happened, in order, before any "why"
2. Systems language — describe the system, not the person
"Jamie skipped the step" -> BLAME (person)
"the checklist had no gate, -> SYSTEM (mechanism)
so the step could be skipped"
3. Action items: specific + owned + dated
"improve monitoring" -> not an action item
"add a saturation alert on checkout's job queue, -> action item
owner: Priya, due Aug 15"
Test each item: can someone start it Monday,
is exactly one person accountable,
and is there a date on the calendar?
| Fits | Keep in mind |
|---|---|
| Incident reviews, outage debriefs, near-misses | Blameless is not consequence-free — it’s just not about punishment |
| Project retros and launch post-mortems | Facts first; the “why” is a hypothesis, not a verdict |
| “Tell me about a failure” interview answers | Show what the system learned, not who you blamed |
Interviewer: “A deploy took down checkout for 20 minutes. Walk me through the postmortem.” A strong answer resists naming a culprit: “First the timeline — deploy at 4:02, error rate spikes at 4:05, paged at 4:14, rolled back at 4:22. Then the framing: the deploy could reach production without a staging check, and the alert paged a channel no one watched — that’s why time-to-ack was nine minutes. So the action items are specific, owned, and dated: ‘add a required staging gate to the deploy pipeline, owner Marco, due next sprint’ and ‘route checkout alerts to the on-call pager, owner Priya, due Friday.’” Not one sentence about who pushed the button — and two fixes that make the next person’s slip harmless.
Which of these is a real action item?