Keep asking why until you reach something you can actually change.
When something breaks, the first explanation is usually a symptom, not a cause. Ask "why did that happen?" a few times and you travel from the thing you noticed to the thing you can fix. The trap is stopping too early — especially at "human error," which is where an investigation starts, never where it ends.
Real failures rarely have one cause. Separate the proximate cause (what triggered it) from the contributing factors and the latent ones (the quiet, systemic gaps that were waiting).
The incident. A transfer pump seized overnight and took the line down. Click "ask why?" to drive from the symptom toward a cause — then choose how far to go.
A solid layer stops the failure. The operator layer can be reminded, but a human check is never solid — it always keeps a hole.
Five whys is a ladder, not a magic number — sometimes it's three, sometimes seven. The framework:
symptom pump seized
why? --> bearing overheated (proximate)
why? --> no lubricant reached it (proximate)
why? --> auto-lube line was clogged (proximate)
why? --> no filter on the fill port (contributing: design)
why? --> PM checklist never inspects the lube line (contributing: procedure)
why? --> no one owns keeping the checklist current (latent: organization)
^ fixable, and it stays fixed
| Fits well | Key limitation |
|---|---|
| Single incidents with a fairly linear cause chain (an outage, a defect, a missed SLA). | Real failures branch — use a fishbone to hold several "why" threads at once, not one straight line. |
| Fast, low-cost first pass before a heavier investigation. | It's only as good as the evidence; unverified "whys" just encode a guess. |
| Getting a team past blame toward a systemic fix. | Complex, safety-critical failures need formal methods too (fault trees, timelines). |
Interviewer: "A nightly job silently stopped writing to the warehouse for a week before anyone noticed. Walk me through the root cause."
Start at the symptom — no data landed for seven days. Why? The job exited early on an error. Why did no one notice? There was no alert on zero-rows-written. Why not? The alert existed once but was removed during a migration. Why did that slip through? No one owned the monitoring config after the team reorganized. The proximate cause is the failing job; the latent cause is ownerless monitoring — and that's the fix that stops the next silent failure too. If you'd stopped at "the on-call engineer should have checked," you'd have blamed a person and left the hole wide open.
Post-incident, someone writes the root cause as "the technician forgot to reset the valve." Best response?