Failure evidence rots on a clock — grab the fastest-fading thing first, then piece the story back together.
When something breaks, the truth is scattered across sources that are disappearing at very different rates. A log ring overwrites itself in minutes. A witness’s memory blurs and picks up other people’s versions by the hour. A fracture surface will still be there next week.
So you by volatility: capture what vanishes fastest first, photograph and tag physical evidence before anyone moves it, and only then reconstruct a timeline — reconciling clocks that disagree and witnesses who contradict the sensors.
Part 1 · collect in order of volatility
Tap the four sources in the order you would collect them. Each capture takes time, and everything not yet captured keeps draining while you work. Then press run collection.
Part 2 · reconcile the clocks, then read the timeline
The captured log and the captured sensor trace share three real moments — but the sensor’s clock is skewed. Slide the offset until the paired events line up and the true sequence appears.
Two separate disciplines, run in this order — preserve, then reconstruct.
1 · Triage by volatility. List every source, estimate how fast each one degrades, and collect fastest-first. The arithmetic the simulator runs is just: whatever is still uncollected keeps decaying while you spend time on the current item.
# decay rate (per min) and capture cost (min)
log ring buffer 10 %/min 2 min (overwrites)
debris field 5 %/min 3 min (staff clear it)
witness memory 4 %/min 3 min (fades, contaminates)
fracture surface 0.5%/min 4 min (stable metal)
Best order — most volatile first
log captured at t=0 -> 100 %
debris captured at t=2 -> 100 - 5*2 = 90 %
witness captured at t=5 -> 100 - 4*5 = 80 %
fracture captured at t=8 -> 100 - .5*8 = 96 %
preserved = (100+90+80+96)/4 = 91 %
Worst order — stable first
fracture captured at t=0 -> 100 %
witness captured at t=4 -> 100 - 4*4 = 84 %
debris captured at t=7 -> 100 - 5*7 = 65 %
log captured at t=10 -> 100 - 10*10 = 0 % (gone)
preserved = (100+84+65+0)/4 = 62 %
2 · Reconstruct with reconciled clocks. Every source stamps time on its own clock. Find events that appear in two sources, measure the constant offset between them, apply it, and the merged sequence becomes defensible. In the widget the three shared moments are each 6 s apart, so the misalignment 3 × |6 − offset| only reaches zero at +6 s.
| Reach for this when… | The trade-off |
|---|---|
| A live incident is still producing volatile state (logs, RAM, open scenes). | Capturing costs minutes you might feel pressure to spend on the fix. |
| Multiple sources will be cited later — sensors, logs, people, physical marks. | Reconciling clocks and contradictions is slow, careful work. |
| The finding must survive scrutiny (, regulator, litigation). | Over-collecting buries the signal; you still have to prioritise. |
You’re asked in an interview: “A payment service threw a burst of errors at 02:14 and recovered on its own. Walk me through your first ten minutes.”
Strong answer: “Before touching anything, I’d preserve the perishable evidence in volatility order — snapshot the pod’s memory and in-flight requests, then pull the rolling application logs before they overwrite, then the load-balancer and database traces. Only then would I interview the on-call engineer while it’s fresh, kept separate from the channel chatter. With that captured, I’d reconcile clocks: our app logs in UTC, the payment gateway’s callbacks were 6 seconds behind, and once I corrected that, the ‘error before request’ ordering resolved — the gateway timeout actually preceded our retries, which changes the root cause entirely.” That answer shows both instincts: save it before it’s gone, then don’t trust a timeline until the clocks agree.
One question
Service is degraded. You have a fracture-analysis part in a bag, a still-running server whose logs roll every few minutes, and a witness heading home. What do you secure first?