The rating a manager walks in with is a hypothesis — the evidence, read against a shared bar, is what settles it.
A calibration meeting exists to make ratings comparable. Left alone, each manager grades on their own private curve, and predictable distortions creep in: one manager runs generous, the person you see every day looks stronger than the one who's remote, and the quiet work that keeps everything running gets no credit.
The discipline is simple to name and hard to hold: anchor every rating to evidence against the shared , not to volume, visibility, or the loudest advocate. Below is a board that came in from three managers. Some numbers are drifting. Drag each rating to what the evidence supports — defending the ones that are right, conceding the ones that aren't.
Six ratings are drifting off the evidence. Adjust each one to what the rubric supports; the bias behind each drift is flagged on the card.
Calibration is a loop, run person by person:
for each employee:
1. state the rating and read the EVIDENCE behind it
2. test it against the SHARED RUBRIC, not other people's numbers
3. name any distortion pulling it off the bar:
leniency drift a manager consistently high or low
visibility bias seen often != contributed more
glue-work blind invisible, load-bearing work undercounted
volume != impact activity is not achievement
4. move the rating to the evidence — defend it if it already
holds, concede it if the evidence doesn't support the number
5. settle only when rating == what the evidence supports
Note the direction cuts both ways. Bias doesn't only inflate; glue-work blindness and low visibility quietly deflate. “Defending a rating” means holding a number the evidence supports against pressure to move it — including moving a strong-but-quiet contributor up.
| Situation | What calibration buys you |
|---|---|
| Ratings feed pay, promotion, or stack decisions | Comparability — a 4 from one manager means the same as a 4 from another |
| Distributed / hybrid teams | A structured check against the in-office visibility advantage |
| New or lenient managers in the room | Drift gets caught as a pattern, not litigated case by case |
| Trade-off | It's slow and can feel adversarial; without evidence discipline it just averages opinions |
In a performance-management interview: “A peer manager insists their report is a 4 because they ‘carried the team all quarter.’ You've seen the record and it reads like a solid 3. How do you handle it?” The strong answer stays on evidence: acknowledge the specific contributions, then ask which of them clear the Exceeds bar in the rubric versus meeting expectations well. If the evidence caps at Meets, you concede nothing to volume or advocacy and hold the 3 — while flagging that if a genuinely load-bearing, low-visibility contributor is sitting at a 3, that's the rating you should be arguing up. Same tool, both directions.
Dana is remote. She carried on-call, unblocked four teams, and wrote the runbooks nobody sees. Her manager proposed a 3. What does calibration ask you to do?