Blameless postmortems: separating the person from the system

“Human error” is where a good investigation starts — the real question is what made the mistake so easy to make.

The idea

When something breaks, the fastest story is “someone messed up.” It feels like an answer, but it’s a full stop where the analysis should begin. Blame someone and the investigation closes — while every trap that caught them stays set for the next person.

Blameless doesn’t mean no one is accountable. Accountability is owning the outcome and the next steps, in public. Blame is assigning shame. One makes the system safer; the other just makes people hide the next mistake.

The launch that broke checkout

The incident
A launch pushed a feature-flag config pointing at a dead endpoint. Checkout errored for 40 minutes.
if you only blame (traps stay set) your current path
Projected similar incidents / year
—
Fixes owned & committed
0 of 3
Choose how the team responds. Blame closes the case; asking what made it easy opens the whys — and each committed fix bends the recurrence curve down.

How it works

A blameless review is a short discipline, not a personality. It moves the question from the person to the conditions, and it only counts if it ends in fixes someone owns:

1. Start at "human error" -- then keep going.
   Treat it as the first clue, not the verdict: what made this
   mistake easy, likely, or invisible?

2. Ask why until you reach a cause you can design out.
   why shipped?  -> no config review gate
   why unseen?   -> alert buried under alert fatigue
   why an outage?-> config service = single point of failure

3. Separate accountability from blame.
   accountability = own the outcome + commit the next step (public)
   blame          = assign shame (private, and it makes people hide)

4. Model the payoff -- fixes remove the trap:
   base recurrence            4.0 / yr   (blame only: x1.00 -> 4.0)
   + config review gate       x0.40  -> 1.6
   + alert prioritization     x0.60  -> 0.96
   + fallback + canary        x0.50  -> 0.48 / yr

Same incident, same person. The only thing that decides whether it happens again is whether you fixed the system or just named the human.

When to use it

Reach for a blameless review when…The trade-off / limit
An incident had a human action right at the surface.It only works if people trust they won’t be punished for being honest.
You want the same failure never to recur.Fixing the system costs real engineering time — more than filing a “be careful” ticket.
Several people could have made the same slip.Blameless doesn’t erase accountability — the team still owns the fixes in public.

Watch out for

Worked example

An interviewer asks: “Tell me about a time something you shipped broke.” A weak answer either grovels (“I felt terrible, I’m usually careful”) or deflects (“QA should’ve caught it”). A strong one owns the outcome plainly — “my config change took checkout down for 40 minutes” — then immediately moves to the system: “what made it easy was no review gate on config, an alert buried under noise, and a single point of failure. I drove three fixes; we haven’t seen that class of incident since.” That’s accountable and blameless: you owned it, you named the traps, you removed them — without shaming yourself or a teammate.

Check yourself

In the postmortem, which is the most useful statement of root cause?

A teammate worries that “blameless” means letting people off the hook. Your response?