Checklists, runbooks & the pre-shift briefing

High-pressure work doesn’t run on heroics — it runs on a page of writing that catches the errors experts make under load.

The idea

Skill isn’t the problem. Under load, even experts skip a step they’ve done a thousand times. A checklist is a memory backstop for exactly those moments. A runbook goes further — it turns a failure you’ve seen before into a rehearsed if this, then that so nobody improvises during the fire. And a tight pre-shift briefing — today’s risks, who owns what, when to escalate — beats a long one, because people actually remember it.

The catch: a checklist only helps if it holds the load-bearing items. Spend your slots on cosmetics and it’s theater.

The interactive · opening a café before the rush

You open a café before the morning rush. There are twelve prep tasks — but under load you’ll forget some. Put your six most load-bearing on a checklist, then run the rush and watch order wait times.

Build your six-item checklist

0 / 6 chosen
If the card reader drops → then switch to the backup reader and manual tickets. Rehearsing this turns a scramble into a 20-second swap.
from memory with your checklist min 00
peak wait · memory
—
peak wait · checklist
—
items still on memory
—

from memory — what slipped

  • Run the rush to see.

with your checklist

  • Run the rush to see.
Pick your six most load-bearing items above, then press run the rush. The same random day is dealt to both runs, so you’re seeing what the checklist actually changed.

How it works

The rush is a simple queue: orders arrive, baristas serve them. When a prep item is skipped, that station slows down — service falls behind arrivals and the wait climbs. Here is the logic the simulator runs each minute:

capacity(t) = base_speed - (sum of drags from every prep item that got skipped)
queue(t+1) = queue(t) + orders_arriving(t) - orders_served(t)
wait(t)     = queue(t) / capacity(t)

# A skipped item's drag is active for the whole rush — one slip at 7am
# is still costing you at 8am. That's why prep, not hustle, sets the day.

from memory : each item is remembered ~55% of the time under load
with a list : listed items are done every time; the rest still ride on memory
a runbook   : doesn't prevent the failure — it shrinks its drag from 0.50 to 0.12

Notice the runbook line: it doesn’t stop the card reader from dropping. It makes the response cheap, because you rehearsed it. That is the whole point of a runbook.

When to use it

Reach for…when…trade-off
A checkliststeps are easy to skip under load and the cost of skipping is high.Too long and people skip the checklist itself. Keep it to what’s load-bearing.
A runbookthe same failure recurs and speed of response matters.Goes if the system changes and nobody updates it.
A briefingthe shift has specific risks or handoffs today.A long briefing is forgotten. Name today’s risks, owners, and escalation triggers — then stop.

Watch out for

Worked example

An interviewer asks: “Your on-call keeps getting paged for the same disk-full alert at 3am. What do you do?” A weak answer says “I’d be more careful.” A strong one reaches for these three tools. First, a runbook: if disk > 90%, then rotate logs on host X and page the owner if it’s still climbing in ten minutes — so the half-asleep responder executes instead of investigates. Then a checklist line in the deploy that verifies log rotation is on, so the alert stops recurring. And a note in the shift briefing: “host X is close to full tonight, here’s the runbook.” Same three moves, whether it’s a café or a data center.

Check yourself

Your checklist keeps growing — it’s now 30 items and people have started skipping it. Best fix?

A runbook exists for the card-reader failure, but nobody has ever practiced it. When it fails at peak…