Prevention that sticks: CAPA & the hierarchy of controls

A fix you have to remember is a fix that decays — the strongest controls don’t depend on anyone behaving.

The idea

Corrective action repairs this occurrence. Preventive action removes the whole class of cause so it can’t happen the same way again. The two are different jobs, and closing an incident with only the corrective half is how the same failure comes back next quarter.

When you choose the preventive action, rank the options by the hierarchy of controls: eliminate, substitute, engineer, administrate, retrain. The higher you reach, the less the fix relies on a tired human doing the right thing — and “we told them to be more careful” is the weakest rung there is.

Recurrence simulator · pick a control, run 12 months

A failure just happened. Choose a preventive control from the ladder — strongest at the top — then fast-forward a year and watch the monthly recurrence risk.

monthly recurrence risk near miss caught failure recurred do-nothing baseline
Month
0 / 12
Recurrence risk now
—
Near misses caught
0
Failures suffered
0
Control holds?
—
Effort to place
—
Pick a control from the ladder, then press play 12 months. Try the weakest rung first, then compare it against an engineering control.

How it works

Two questions, always asked in order:

1 · Corrective, then preventive. First stop the bleeding on the current occurrence. Then ask the harder question: what class of cause let this happen, and how do I remove it? A CAPA that only answers the first question isn’t finished.

2 · Rank the preventive options by the hierarchy. Reach as high up the ladder as is feasible. Effectiveness climbs as human dependence falls:

strongest  ELIMINATE     remove the hazard         no human needed
           SUBSTITUTE    swap for a safer design   almost none
           ENGINEER      guard / interlock         built into the tool
           ADMINISTRATE  procedure, checklist      relies on adherence
weakest    RETRAIN       "be more careful"         relies on memory

recurrence risk stays FLAT for eliminate / substitute / engineer
recurrence risk CREEPS BACK for administrate / retrain
  -> people forget, staff turn over, the shortcut returns

The bottom two rungs still matter — but only as a layer on top of a stronger control, never as the whole fix. And every near miss is free failure data: it’s the warning that a weak control is drifting, arriving before the real failure does.

When to use it

Reach up the ladder when…The trade-off
The same failure has recurred, or a near miss shows it’s about to.Higher controls cost more up front — design change, tooling, downtime.
The consequence is severe or the population large.Elimination isn’t always feasible; you settle for the highest rung you can build.
You can’t rely on perfect human behaviour at 3am.Retraining is fast and cheap — but it decays, so treat it as a supplement.

Watch out for

Worked example

Interview prompt: “An engineer deployed a config that took down checkout because they skipped the staging step. What’s your fix?”

Weak answer stops at: “I’d roll back and remind the team to always test in staging.” That’s a corrective action plus the weakest control — the next new hire will skip staging again. A strong answer climbs the ladder: “Corrective: roll back and restore checkout. Preventive: I’d engineer it out — make the deploy pipeline physically refuse a production push that hasn’t passed staging, so the unsafe path no longer exists. As a backing layer, a checklist and clearer ownership. And I’d watch the near-miss rate — blocked deploys that would have skipped staging — as the signal the control is working.” That answer shows you know corrective from preventive, and that you reach for a control that doesn’t decay.

Check yourself

One question

A pharmacy keeps mixing up two look-alike vials. Which preventive action sits highest on the hierarchy?