Preserve the evidence, reconstruct the sequence

Failure evidence rots on a clock — grab the fastest-fading thing first, then piece the story back together.

The idea

When something breaks, the truth is scattered across sources that are disappearing at very different rates. A log ring overwrites itself in minutes. A witness’s memory blurs and picks up other people’s versions by the hour. A fracture surface will still be there next week.

So you by volatility: capture what vanishes fastest first, photograph and tag physical evidence before anyone moves it, and only then reconstruct a timeline — reconciling clocks that disagree and witnesses who contradict the sensors.

Part 1 · collect in order of volatility

Tap the four sources in the order you would collect them. Each capture takes time, and everything not yet captured keeps draining while you work. Then press run collection.

Clock elapsed
0 min
Sources captured
0 / 4
Evidence preserved
—
Build a collection order above, or press sort by volatility to let the sim order them for you.

Part 2 · reconcile the clocks, then read the timeline

The captured log and the captured sensor trace share three real moments — but the sensor’s clock is skewed. Slide the offset until the paired events line up and the true sequence appears.

+0 s
Total misalignment
18 s
Sequence status
contradictory
At offset +0 s the sensor’s flow reading looks like it happened before the valve command — a false story born of clock skew.
Reconstructed timeline (log & sensor reconciled)

    How it works

    Two separate disciplines, run in this order — preserve, then reconstruct.

    1 · Triage by volatility. List every source, estimate how fast each one degrades, and collect fastest-first. The arithmetic the simulator runs is just: whatever is still uncollected keeps decaying while you spend time on the current item.

    # decay rate (per min) and capture cost (min)
    log ring buffer   10 %/min   2 min   (overwrites)
    debris field       5 %/min   3 min   (staff clear it)
    witness memory     4 %/min   3 min   (fades, contaminates)
    fracture surface  0.5%/min   4 min   (stable metal)
    
    Best order  — most volatile first
      log      captured at t=0   -> 100 %
      debris   captured at t=2   -> 100 - 5*2  =  90 %
      witness  captured at t=5   -> 100 - 4*5  =  80 %
      fracture captured at t=8   -> 100 - .5*8 =  96 %
      preserved = (100+90+80+96)/4 = 91 %
    
    Worst order — stable first
      fracture captured at t=0   -> 100 %
      witness  captured at t=4   -> 100 - 4*4  =  84 %
      debris   captured at t=7   -> 100 - 5*7  =  65 %
      log      captured at t=10  -> 100 - 10*10 =  0 %  (gone)
      preserved = (100+84+65+0)/4 = 62 %

    2 · Reconstruct with reconciled clocks. Every source stamps time on its own clock. Find events that appear in two sources, measure the constant offset between them, apply it, and the merged sequence becomes defensible. In the widget the three shared moments are each 6 s apart, so the misalignment 3 × |6 − offset| only reaches zero at +6 s.

    When to use it

    Reach for this when…The trade-off
    A live incident is still producing volatile state (logs, RAM, open scenes).Capturing costs minutes you might feel pressure to spend on the fix.
    Multiple sources will be cited later — sensors, logs, people, physical marks.Reconciling clocks and contradictions is slow, careful work.
    The finding must survive scrutiny (, regulator, litigation).Over-collecting buries the signal; you still have to prioritise.

    Watch out for

    Worked example

    You’re asked in an interview: “A payment service threw a burst of errors at 02:14 and recovered on its own. Walk me through your first ten minutes.”

    Strong answer: “Before touching anything, I’d preserve the perishable evidence in volatility order — snapshot the pod’s memory and in-flight requests, then pull the rolling application logs before they overwrite, then the load-balancer and database traces. Only then would I interview the on-call engineer while it’s fresh, kept separate from the channel chatter. With that captured, I’d reconcile clocks: our app logs in UTC, the payment gateway’s callbacks were 6 seconds behind, and once I corrected that, the ‘error before request’ ordering resolved — the gateway timeout actually preceded our retries, which changes the root cause entirely.” That answer shows both instincts: save it before it’s gone, then don’t trust a timeline until the clocks agree.

    Check yourself

    One question

    Service is degraded. You have a fracture-analysis part in a bag, a still-running server whose logs roll every few minutes, and a witness heading home. What do you secure first?