Precision vs recall in product terms

One dial — the threshold — trades blocking good customers against letting bad ones through. There is no setting that does neither.

The idea

A model gives every case a score. You pick a threshold: act on everything above it, ignore everything below. That single choice creates four outcomes — the confusion matrix.

Precision answers “when we act, how often are we right?” Recall answers “of all the bad cases, how many did we catch?” Move the threshold and they trade: catch more fraud, block more good customers — and vice versa. Accuracy, the single overall number, quietly lies when one class is rare.

drag the threshold across a scored population

actually fraud actually legit left of the line = allowed · right = flagged
0.50
—caught fraud · TP
—blocked good customer · FP
—missed fraud · FN
—allowed good customer · TN
precision
—
recall
—
Accuracy
—
 
Product cost
—
FP × $50  +  FN × $300
Slide the threshold. Watch precision and recall pull against each other — and switch to 1% fraud to see accuracy stay high while the product quietly fails.

How it works

Every threshold sorts the population into four boxes. Precision reads across the “we acted” row; recall reads down the “actually positive” column.

flag a case when score >= threshold

                   actually fraud     actually legit
   flagged            TP                 FP        -> precision = TP / (TP + FP)
   allowed            FN                 TN

   recall   = TP / (TP + FN)     "of all fraud, how much did we catch?"
   accuracy = (TP + TN) / total  "how many calls were right, overall?"

lower the threshold -> more flags -> recall up, precision down  (block more good)
raise the threshold -> fewer flags -> precision up, recall down  (miss more fraud)

the imbalance trap  (fraud is 1% of traffic):
   a model that flags NOTHING:  TP=0  FP=0  FN=5  TN=495
   accuracy = 495 / 500 = 99%     looks excellent
   recall   = 0 / 5     = 0%      catches no fraud
   -> pick the metric that matches the cost of being wrong, and read it per slice.

Because the classes carry different costs, you don’t chase one number — you price the four boxes. Here a missed fraud (FN) costs six times a blocked good customer (FP), so the cheapest threshold leans toward catching fraud even at some precision.

When to use which

Optimise for…When the costly mistake is…
Recall (catch more, tolerate false alarms)Missing a positive — fraud, disease screening, safety alerts. A miss is far worse than a false alarm.
Precision (act only when confident)Acting wrongly — blocking paying customers, auto-suspending accounts, sending a costly intervention.
A cost-weighted blend (F-beta / expected cost)Both errors hurt but unequally — price each box and minimise expected cost, per segment.

Watch out for

Worked example

An interviewer asks: “Your fraud model is 99% accurate. Ship it?” The strong move is to refuse the single number. Fraud is maybe 1% of transactions, so a model that approves everything is also 99% accurate and catches zero fraud — accuracy is meaningless here. You’d ask for precision and recall at the operating threshold, and for the cost of each error: a missed fraud is a chargeback plus goodwill; a blocked good customer is a lost order and a support ticket. With those costs you pick the threshold that minimises expected cost, not the one that maximises accuracy — and you’d check precision and recall separately for new customers, where the base rate and the stakes differ.

Check yourself

Fraud is 1% of transactions. A stakeholder is thrilled the model is “99% accurate.” Your read?

Support says too many legitimate customers are being blocked. Which single move most directly reduces that, on the same model?