Debugging as hypothesis testing

A bug is not a mystery to stare at — it is a search space to cut in half.

The idea

Good debugging is not luck or intuition. It is a loop: reproduce the failure reliably, ask what changed, form a hypothesis that a cheap probe can prove false, and split the suspect space roughly in half with each probe. The trick is that a well-placed probe removes half the candidates whether it passes or fails — so the number of probes grows with the logarithm of the search space, not its size.

The other half of the discipline: change one thing at a time. The moment you vary two things together, a failure can no longer be blamed on one of them, and the trail goes cold.

Interactive · find the fault

input16-stage pipeline →output (failing)
suspects probes →
your probes bisection ideal
suspect stages
16
probes used
0
bisection guarantee
4
The pipeline reproduces its failure every run — good, a reliable repro is step one. The fault sits in one of 16 stages. Click a stage to probe its output, or let bisection pick.

A probe on stage k tells you if the data is still healthy there. Healthy → the fault is later. Broken → the fault is at or before k.

How it works

Model the pipeline as a boundary you are hunting: every stage before the fault produces healthy output; every stage from the fault onward is broken. Finding the fault means finding that boundary — and binary search finds a boundary in a list of n items in about log2(n) steps.

suspect range = [1 .. 16]          # 16 stages could hold the fault

while range has more than one stage:
    k = middle of the range         # the cheapest 50/50 probe
    if output at k is healthy:
        fault is after k   -> keep the upper half   [k+1 .. hi]
    else:                           # output at k is broken
        fault is at or before k -> keep the lower half [lo .. k]

# each probe halves the range:
16 -> 8 -> 4 -> 2 -> 1              # 4 probes, worst case
log2(16) = 4

The same logic is what "walk me through your investigation" is really asking for. git bisect is this over commits; splitting a request path is this over services; commenting out half a config is this over settings.

When to use it

Fits wellThe catch
A regression with a known-good and known-bad point (a passing build, a healthy deploy, a good input).You need a reliable repro and a fast, honest pass/fail test at each probe.
A long linear chain: commits, pipeline stages, request , migration steps.If failures are flaky or intermittent, a single probe can lie — repeat it before trusting it.
Any space you can cut cleanly in half without side effects.Bisection assumes one boundary. Two independent faults break the "healthy below, broken above" shape.

Watch out for

Worked example

Checkout starts 500-ing on Monday; Friday's build was fine. That is your known-good and known-bad — 40 commits between them. Instead of reading all 40, you git bisect: test the commit in the middle, mark it good or bad, and each answer discards 20. Six probes (log2(40) ≈ 5.3) land you on the exact commit — a currency-formatting change that throws on a locale you now see in prod. In the interview, the strong version is naming it out loud: "I reproduced it, found my good and bad boundary, bisected, and confirmed the one change — I never touched two variables at once."

Check yourself

You have narrowed the fault to 8 possible stages. How many more probes does bisection guarantee?

Your fix works, but the same commit also bumps a library version. Best next step before you call it solved?