Calibration: defending ratings against drift and bias

The rating a manager walks in with is a hypothesis — the evidence, read against a shared bar, is what settles it.

The idea

A calibration meeting exists to make ratings comparable. Left alone, each manager grades on their own private curve, and predictable distortions creep in: one manager runs generous, the person you see every day looks stronger than the one who's remote, and the quiet work that keeps everything running gets no credit.

The discipline is simple to name and hard to hold: anchor every rating to evidence against the shared , not to volume, visibility, or the loudest advocate. Below is a board that came in from three managers. Some numbers are drifting. Drag each rating to what the evidence supports — defending the ones that are right, conceding the ones that aren't.

2 of 8 ratings anchored to evidence
the shared scale — incoming proposal (warm tick) and current rating (dot); green once it matches the evidence

Six ratings are drifting off the evidence. Adjust each one to what the rubric supports; the bias behind each drift is flagged on the card.

Bias watch
Leniency drift. Manager A’s proposals run about half a point high across the board — correct Priya and Lena to the evidence.
Visibility bias. Marcus and Owen are highly visible; exposure inflated their scores above delivered impact.
Glue-work blindness. Dana and Nadia do invisible, load-bearing work; it was undercounted, not overcounted.

How it works

Calibration is a loop, run person by person:

for each employee:
    1. state the rating and read the EVIDENCE behind it
    2. test it against the SHARED RUBRIC, not other people's numbers
    3. name any distortion pulling it off the bar:
         leniency drift    a manager consistently high or low
         visibility bias   seen often  != contributed more
         glue-work blind   invisible, load-bearing work undercounted
         volume != impact  activity is not achievement
    4. move the rating to the evidence — defend it if it already
       holds, concede it if the evidence doesn't support the number
    5. settle only when rating == what the evidence supports

Note the direction cuts both ways. Bias doesn't only inflate; glue-work blindness and low visibility quietly deflate. “Defending a rating” means holding a number the evidence supports against pressure to move it — including moving a strong-but-quiet contributor up.

When to use it

SituationWhat calibration buys you
Ratings feed pay, promotion, or stack decisionsComparability — a 4 from one manager means the same as a 4 from another
Distributed / hybrid teamsA structured check against the in-office visibility advantage
New or lenient managers in the roomDrift gets caught as a pattern, not litigated case by case
Trade-offIt's slow and can feel adversarial; without evidence discipline it just averages opinions

Watch out for

Worked example

In a performance-management interview: “A peer manager insists their report is a 4 because they ‘carried the team all quarter.’ You've seen the record and it reads like a solid 3. How do you handle it?” The strong answer stays on evidence: acknowledge the specific contributions, then ask which of them clear the Exceeds bar in the rubric versus meeting expectations well. If the evidence caps at Meets, you concede nothing to volume or advocacy and hold the 3 — while flagging that if a genuinely load-bearing, low-visibility contributor is sitting at a 3, that's the rating you should be arguing up. Same tool, both directions.

Check yourself

Dana is remote. She carried on-call, unblocked four teams, and wrote the runbooks nobody sees. Her manager proposed a 3. What does calibration ask you to do?