Metric anomaly triage: real move or artifact?

Before you believe a number moved, prove it isn’t the calendar, the noise, or the pipeline lying to you.

The idea

A dashboard drops 12% overnight and the room panics. Slow down. The vast majority of “overnight” moves are artifacts — a holiday, ordinary noise, a deploy that double-logged or dropped events, a pipeline still catching up, or two dashboards that define the metric differently.

So triage in a fixed order, cheapest checks first: is this even unusual (variance, )? then, did we break the measurement (instrumentation)? and only when both are clean do you conclude real users actually changed. Real behavior is what’s left after you rule the artifacts out.

Triage game · spend checks to find the culprit

Round 1 / 5 Tokens left 6 Cases solved 0

—

Change by segment (last day vs prior)

Your verdict

Spend a check to reveal a clue — lead with the cheap ones. When you’re confident, call the culprit.

How it works

Run the checks in ascending cost. Each one either closes the case or rules out a whole class of cause. The order matters because the cheap checks resolve most cases — you rarely need the expensive segment split.

  1. Is it even unusual? Compare against the normal band and the weekly / holiday pattern. Inside the band or on a known seasonal trough → probably nothing to explain.
  2. Did we break the measurement? Check the deploy log around the change, and sanity-check event count against user count. A shared, uniform jump right after a release is the fingerprint of an instrumentation bug.
  3. Only then, is it real? If the size beats noise, no deploy explains it, events per user look sane, and the move is broad across segments, you’re likely looking at real behavior — escalate to why.

The event-vs-user check is the highest-yield artifact test. Here is the arithmetic behind a “purchases up 40%” scare:

purchase events logged yesterday:   4,400
distinct paying users yesterday:    2,000
events per paying user:  4,400 / 2,000 = 2.2   (should sit near 1.1)

users did not double. the logging did.
"+40% purchases" was a duplicate-event artifact, not demand.

When to use it

Reach for triage-first when…The trade-off / limit
A metric moves sharply “overnight” and someone wants an answer now.Costs a few minutes before you react — but reacting to an artifact costs far more.
The move lines up with a deploy, a date, or a dashboard change.Correlation with a deploy is a strong hint, not proof — confirm with events-per-user.
You can slice by segment and compare against history.Segment splits are the most expensive check — save them for when cheap checks disagree.

Watch out for

Worked example

An interviewer says: “iOS revenue just fell 25% overnight. What do you do?” A weak answer starts guessing at churn. A strong one triages: first the cheap checks — no holiday, and the move is far outside the normal band, so it’s real-sized. Then the deploy log: an iOS app release shipped yesterday with an analytics SDK bump. Events-per-user shows iOS purchase events collapsed to near zero while App Store payouts held steady — events are being lost, not sales. The segment split confirms Android and web are flat; only iOS fell. Verdict: an instrumentation artifact from the broken SDK, not a revenue event. You’d roll back the tracking, backfill, and tell the room the money is fine before anyone reprices anything.

Check yourself

Signups look down 12% since yesterday. You have one minute. What’s the first check?

Purchases jumped 40% right after a 2am checkout deploy. Events per paying user went from 1.1 to 2.2. Your read?