Signal vs noise: is this spike real?

Every process breathes within a band of ordinary noise — the skill is knowing when a bad day has actually left it.

The idea

Numbers move every day even when nothing changed. That ordinary wobble is common-cause variation — the natural noise of a stable process. Reacting to it (a stern email after one bad Tuesday) is called tampering, and it usually makes things worse.

A special cause is different: a real shift you can trace to a reason. A control chart tells them apart. You set a center line and limits at plus/minus three sigma from a stable baseline, then watch. Inside the band: noise, leave it. Outside the band, or a long run on one side: a signal, go investigate.

Control chart · daily complaint counts

noise band (±3σ) within band — noise broke a rule — signal
12 / day
σ = 3.0
Your call:
Center line
12
Control limits ±3σ
3 – 21
Verdict
—
A stable month. Every point is common-cause noise inside the band. Drag the sliders to feel the band widen with noise — then inject an episode and judge it before you press your call.

How it works

The limits come from a stable baseline — a stretch when nothing special was happening. Freeze them, then judge new days against them:

baseline mean (center line) .......... CL = 12 complaints/day
baseline spread .................... σ  =  3
upper limit .................. UCL = 12 + 3×3 = 21
lower limit .................. LCL = 12 - 3×3 =  3

signal rules (either one is enough):
  1. any point above 21 or below 3          -> special cause
  2. eight days in a row on one side of 12   -> special cause
otherwise -> common cause. leave the process alone.

Why three sigma? For stable data, only about 0.3% of days fall outside the band by chance — so a point outside is far more likely a real change than a fluke. A run of eight on one side has odds near 1 in 250, so it flags a quiet drift the single-point rule misses.

Isolating a change when several things moved at once: compare the recent window to the frozen baseline, and hold the confounds still. If you changed the return policy the same week a holiday hit, you can’t credit the policy until you compare like weeks (year-over-year, or a region that didn’t get the change).

When to use it

Reach for a control chart when…The trade-off
You watch a metric over time and keep reacting to single bad days.Needs a genuinely stable baseline — garbage limits from a chaotic month mislead.
You need to tell a real regression from normal churn (, defects, complaints).Assumes roughly independent points; strong weekly needs charting by day-type.
You want a pre-agreed trigger so the team stops debating every wobble.3σ is deliberately slow — it trades a few missed small shifts for very few false alarms.

Watch out for

Worked example

An interviewer says: “Complaints jumped 30% on Tuesday — what do you do?” The strong answer resists the reflex. First, plot it: is Tuesday outside the ±3σ band you’d built from a stable stretch, or just a loud day inside it? If it’s inside, it’s common cause — you note it and don’t reorganise the team over noise. If it’s outside or part of a run, it’s a special cause worth a root-cause hunt. And before you blame the new checkout flow that shipped Monday, you check what else moved — a promo, an outage, a holiday — and compare against a window where only the checkout changed. That sequence, baseline → compare → control for confounds, is what separates a calm operator from a fire-fighter.

Check yourself

Your defect rate sits at 4/day (σ = 1, so the band is roughly 1 to 7). Today you logged 6. What’s the disciplined read?

No single day has crossed a limit, but the last ten days have all landed above the center line. Signal or noise?