Root cause: five whys

Keep asking why until you reach something you can actually change.

The idea

When something breaks, the first explanation is usually a symptom, not a cause. Ask "why did that happen?" a few times and you travel from the thing you noticed to the thing you can fix. The trap is stopping too early — especially at "human error," which is where an investigation starts, never where it ends.

Real failures rarely have one cause. Separate the proximate cause (what triggered it) from the contributing factors and the latent ones (the quiet, systemic gaps that were waiting).

The incident. A transfer pump seized overnight and took the line down. Click "ask why?" to drive from the symptom toward a cause — then choose how far to go.

Start at the symptom, then ask why to extend the chain.
Depth: 0 whys  ·  proximate 0 · contributing 0 · latent 0
Layers of defense — where the holes lined up

A solid layer stops the failure. The operator layer can be reminded, but a human check is never solid — it always keeps a hole.

How it works

Five whys is a ladder, not a magic number — sometimes it's three, sometimes seven. The framework:

  1. Start at the symptom, stated as an observable fact ("the pump seized"), not a guess.
  2. Ask why, answer with evidence, and make that answer the next thing to explain. Branch when there's more than one cause — that's the fishbone view.
  3. Stop at a cause you can change and own. "Metal filings clogged the lube line" is real, but the fixable answer is one layer deeper: why was there nothing to catch them?
  4. Sort the causes. Proximate (the trigger), contributing (made it worse or likelier), latent (the systemic gap that sat there quietly). The latent ones are usually the cheap, durable fixes.
  5. Never stop at "human error." Ask what made the error easy to make and hard to catch — that's the design or process gap you can actually close.
symptom      pump seized
  why?  -->  bearing overheated          (proximate)
  why?  -->  no lubricant reached it     (proximate)
  why?  -->  auto-lube line was clogged  (proximate)
  why?  -->  no filter on the fill port  (contributing: design)
  why?  -->  PM checklist never inspects the lube line (contributing: procedure)
  why?  -->  no one owns keeping the checklist current  (latent: organization)
             ^ fixable, and it stays fixed

When to use it

Fits wellKey limitation
Single incidents with a fairly linear cause chain (an outage, a defect, a missed SLA).Real failures branch — use a fishbone to hold several "why" threads at once, not one straight line.
Fast, low-cost first pass before a heavier investigation.It's only as good as the evidence; unverified "whys" just encode a guess.
Getting a team past blame toward a systemic fix.Complex, safety-critical failures need formal methods too (fault trees, timelines).

Watch out for

Worked example

Interviewer: "A nightly job silently stopped writing to the warehouse for a week before anyone noticed. Walk me through the root cause."

Start at the symptom — no data landed for seven days. Why? The job exited early on an error. Why did no one notice? There was no alert on zero-rows-written. Why not? The alert existed once but was removed during a migration. Why did that slip through? No one owned the monitoring config after the team reorganized. The proximate cause is the failing job; the latent cause is ownerless monitoring — and that's the fix that stops the next silent failure too. If you'd stopped at "the on-call engineer should have checked," you'd have blamed a person and left the hole wide open.

Check yourself

Post-incident, someone writes the root cause as "the technician forgot to reset the valve." Best response?