Code RoomSleep fix masks root cause
HardPrep Room Coding #4995

Sleep fix masks root cause

Vibe & agenticAlgorithms & data structuresSenior–Staff~22 min

A test was flaky — failing ~1 in 20 CI runs. You asked an AI agent to fix it; it added a `sleep(500)` before the assertion and the test stopped failing in 50 local runs. The team wants to merge. As the reviewer, how do you determine whether this fix addresses the root cause or just lowers the failure rate — and what would you require before approving?

Implement
review_flaky_fix(patch_lines: list[str], green_run_count: int, baseline_one_failure_in_runs: int) → list[str]
Examples
in[["--- a/tests/test_checkout.py","+++ b/tests/test_checkout.py","- assert order.status == 'paid'","+ sleep(500)","+ assert order.status == 'paid'"],50,20]out["wall_clock_sleep","no_condition_wait","insufficient_runs"]
in[["+ wait_until(order_is_paid, timeout_ms=5000)","+ assert order.status == 'paid'"],600,20]out["approve"]
in[["+ poll_until(order_is_paid)","+ assert order.status == 'paid'"],50,20]out["insufficient_runs"]
What a strong answer looks like

Treat the AI’s output as a draft to verify, not an answer to trust. Name the specific flaw and the input that triggers it, say how you’d catch it (tests, edge cases, reading critically), and how you’d re-prompt or decompose to get it right.

0:00 of about 22 min

Vibe & agentic: describe the solution in plain language (or narrate it) and the coach grades your approach.

Which questions mattered is sealed until you submit. Telling you now would just be handing over the edge cases.