One dial — the threshold — trades blocking good customers against letting bad ones through. There is no setting that does neither.
A model gives every case a score. You pick a threshold: act on everything above it, ignore everything below. That single choice creates four outcomes — the confusion matrix.
Precision answers “when we act, how often are we right?” Recall answers “of all the bad cases, how many did we catch?” Move the threshold and they trade: catch more fraud, block more good customers — and vice versa. Accuracy, the single overall number, quietly lies when one class is rare.
drag the threshold across a scored population
Every threshold sorts the population into four boxes. Precision reads across the “we acted” row; recall reads down the “actually positive” column.
flag a case when score >= threshold
actually fraud actually legit
flagged TP FP -> precision = TP / (TP + FP)
allowed FN TN
recall = TP / (TP + FN) "of all fraud, how much did we catch?"
accuracy = (TP + TN) / total "how many calls were right, overall?"
lower the threshold -> more flags -> recall up, precision down (block more good)
raise the threshold -> fewer flags -> precision up, recall down (miss more fraud)
the imbalance trap (fraud is 1% of traffic):
a model that flags NOTHING: TP=0 FP=0 FN=5 TN=495
accuracy = 495 / 500 = 99% looks excellent
recall = 0 / 5 = 0% catches no fraud
-> pick the metric that matches the cost of being wrong, and read it per slice.
Because the classes carry different costs, you don’t chase one number — you price the four boxes. Here a missed fraud (FN) costs six times a blocked good customer (FP), so the cheapest threshold leans toward catching fraud even at some precision.
| Optimise for… | When the costly mistake is… |
|---|---|
| Recall (catch more, tolerate false alarms) | Missing a positive — fraud, disease screening, safety alerts. A miss is far worse than a false alarm. |
| Precision (act only when confident) | Acting wrongly — blocking paying customers, auto-suspending accounts, sending a costly intervention. |
| A cost-weighted blend (F-beta / expected cost) | Both errors hurt but unequally — price each box and minimise expected cost, per segment. |
An interviewer asks: “Your fraud model is 99% accurate. Ship it?” The strong move is to refuse the single number. Fraud is maybe 1% of transactions, so a model that approves everything is also 99% accurate and catches zero fraud — accuracy is meaningless here. You’d ask for precision and recall at the operating threshold, and for the cost of each error: a missed fraud is a chargeback plus goodwill; a blocked good customer is a lost order and a support ticket. With those costs you pick the threshold that minimises expected cost, not the one that maximises accuracy — and you’d check precision and recall separately for new customers, where the base rate and the stakes differ.
Check yourself
Fraud is 1% of transactions. A stakeholder is thrilled the model is “99% accurate.” Your read?
Support says too many legitimate customers are being blocked. Which single move most directly reduces that, on the same model?