Bottlenecks & queues: why waits explode near capacity

A line moves only as fast as its slowest step — and near full capacity, small load increases cause huge waits.

The idea

Every process has one step that is slower than the rest. That step — the constraint — sets the pace for the whole line, no matter how fast everything else runs. Making a non-constraint faster just leaves it idle sooner; it does not add output.

The second, less obvious truth: as a busy step approaches 100% utilization, the wait in front of it does not rise gently — it curves toward vertical. That is how a team that looks “fully staffed” can still be drowning.

See it work

A four-station line — find the constraint

4.0/min
Constraint
Station 2
WIP at constraint
0
Total produced
0

Press play. Work flows left to right; each station runs at its slider speed. Watch where the pile grows.

One busy step — utilization vs. wait

Utilization
60%
Avg wait
1.5×

At 60% utilization, a job waits about 1.5× its own service time. Comfortable and stable.

How it works

Two simple relationships do most of the work.

Throughput of a line = the rate of its slowest station. Upstream stations that are faster just fill a queue in front of the constraint; downstream stations that are faster sit starved, waiting for it.

Wait grows with utilization, not linearly. For a single busy step where work arrives irregularly (the M/M/1 model), the average time a job spends waiting in the queue, measured in units of its own service time, is:

utilization  ρ = arrival rate / capacity

wait in queue  Wq = ρ / (1 - ρ)   (in service-time units)

ρ = 0.50  →  0.50 / 0.50 =  1.0×
ρ = 0.80  →  0.80 / 0.20 =  4.0×
ρ = 0.90  →  0.90 / 0.10 =  9.0×
ρ = 0.95  →  0.95 / 0.05 = 19.0×
ρ = 0.98  →  0.98 / 0.02 = 49.0×

The denominator (1 - ρ) shrinks toward zero as you approach full load, so the wait blows up. Kingman’s formula generalizes this: more variability in arrivals or service times makes the same curve even steeper.

The playbook from the theory of constraints: find the constraint, exploit it (keep it never idle, never starved), subordinate everything else to its pace, then elevate it (add capacity). And expect the constraint to move to the next-slowest step once you do.

When to use it

Reach for it whenTrade-off to remember
Diagnosing why a pipeline’s output is stuck despite “enough” peopleImproving anything but the constraint is wasted effort — sometimes worse, since it feeds the pile faster
Deciding where a support / ops team is at risk of collapseRunning any step near 100% trades a little idle time for wildly unstable waits
Sizing capacity for spiky, uncertain demandTargeting high utilization looks efficient on paper but leaves no for variability

Watch out for

Worked example

You run a document-review pipeline: intake, first pass, senior review, publish. Intake and first pass are quick; senior review handles 4 files/hour; publish handles 12. Throughput of the whole line is 4 files/hour — the senior-review rate — and a backlog piles up in front of it while publish sits mostly idle. Hiring another intake clerk does nothing. Adding a second senior reviewer lifts the line to the next constraint. And when leadership pushes each reviewer to 95% booked, the average turnaround stops being predictable: one sick day and the queue takes weeks to drain. In an interview, naming the constraint and the utilization risk in one breath is what senior sounds like.

Check yourself

Your line runs at 4 units/min because station 2 is the slowest. You double the speed of station 3 (which was already faster). What happens to output?

A busy queue goes from 90% to 95% utilization. Roughly what happens to the average wait?

Powering your career growth.