Competency and Culture
Calibration Isn't Bureaucracy. It's the Only Honest Part of Your Review.
More raters will not fix unreliable ratings. Evidence-based calibration will. What a good session looks like, and how to avoid forced ranking.

About This Series
This is the sixth post in our seven-part Competency and Culture series. In our first post, we showed that most of a performance rating reflects the rater, and we promised to come back to the fix. This is that post. Calibration is where a company checks its ratings before they become pay, promotion, or exit decisions. Done well, it is the most honest hour in the review cycle. Done badly, it becomes either a rubber stamp or forced ranking with a friendlier name. Get this right, and every earlier post in this series has something reliable to stand on.
Most founders hear the word calibration and picture a long meeting, a spreadsheet, and managers arguing over decimals. It sounds like process for the sake of process, the kind of thing a growing company adds when it starts to feel like a corporation.
That picture is understandable, and it is wrong. If most of a rating reflects the person giving it, then calibration is not bureaucracy. It is the only step in the whole review process where anyone checks whether the rating is true.
The obvious alternative is to skip it and add more opinions instead: peer reviews, 360s, continuous ratings from everyone. It feels fairer and more modern. As one European company showed, it can make things worse.
If one manager's rating is unreliable, why don't more ratings solve it?
Because every rater brings their own ruler. Adding raters adds rulers, not accuracy.
The landmark study on rating reliability, by Scullen, Mount, and Goff in 2000, is useful here too. It looked at 4,492 managers, each rated from six angles: two bosses, two peers, two direct reports. If the vantage point mattered most, bosses would agree with bosses and peers with peers. But the differences between perspectives explained only a small share of the variation. The biggest share, about 62%, came from each individual rater's own tendencies, whatever their position.
In other words, a peer is not automatically more accurate than a manager. They are simply another person with their own standards, blind spots, and relationships. Averaging six biased opinions gives you a smoother number. It does not give you a truer one.
More raters also bring a new problem. When ratings come from everyone, all the time, and feed directly into pay, people start rating strategically. They reward allies, protect friends, and mark down competitors. The system meant to reduce bias gives people new ways to act on it.
What happened when Zalando let employees rate each other continuously?
In November 2019, Germany's Hans-Böckler-Stiftung published a study by sociologists Philipp Staab and Sascha-Christopher Geschke of Humboldt University on Zonar, the performance software then used at Zalando's Berlin head office.
Zonar worked a little like product reviews on an online shop, except employees were rating each other. Staff gave continuous feedback on colleagues across departments, and an algorithm turned those ratings into scores that sorted people into Low, Good, and Top Performers. Those categories shaped promotions and pay.
The employees the researchers interviewed described "a culture of total, one-sided transparency." The study reported that in some departments only 2 to 3% of people were classed as Top Performers, the only tier with meaningful raises, and that the system created competitive pressure and stress rather than better performance.
Fairness matters here, so here is the other side. The study was based on interviews with 10 employees, plus group discussions and training materials. Zalando said the study was not representative, pointed to high satisfaction in its own internal surveys, and defended its feedback culture as transparent. Both positions are on the public record.
We are not using Zonar to say peer feedback is bad. Peer input is valuable. The lesson is narrower and more useful: more ratings, fed straight into pay without a human check on the evidence, do not remove bias. They can scale it. What Zonar's design lacked, at least as the study described it, was a step where people looked at the evidence behind the scores before the scores decided anything.
What this looks like in practice
use peer and upward feedback as evidence for a conversation, not as votes in a formula. The number from a 360 should never reach a pay decision without someone asking, "What did they actually see?"
What does a calibration session actually do?
It replaces opinions with evidence, one rating at a time, before any decision is made. That is all it needs to do. Everything else is detail.
A good calibration session follows one rule: no score without evidence. Every rating comes with at least two written examples of behaviour from the review period, tied to the behaviour descriptions for that person's level. A manager is never asked to defend a number. They are asked to show what they saw.
| Step | What happens | Why it matters |
|---|---|---|
| 1. Prepare | Each manager submits ratings with two behaviour examples per person | Forces the rater to check their own score before anyone else does |
| 2. Compare by team | All ratings for a team sit on one page, beside each rater's average and the other perspectives on each person | Shows rater patterns (strict, generous, bunched) before individual cases |
| 3. Discuss the gaps | Time goes to ratings where perspectives disagree, or where a rater sits far from their peers | Puts effort where error is most likely |
| 4. Test the evidence | Each disputed rating is checked against its examples and the behaviour descriptions | Turns "I think" into "here is what happened" |
| 5. Record the reason | Any changed rating is noted with the evidence that changed it | Makes the outcome explainable to the employee later |
Notice what is missing from that table: a target distribution. Good calibration does not start from how many people should be in each box. It starts from the evidence and lets the distribution fall where it falls.
How did calibration fix two very different rating errors?
In our first post, we described two rating errors in two different teams. Here is how calibration dealt with each, in two separate sessions.
From our experience
In the first team, we did not ask the strict manager to defend a number. We asked for two written examples of behaviour from the review period. Those examples described a solid, reliable engineer, and the manager's scores sat below colleagues across the whole team. The rating moved.
In the second team, the skip manager's examples all came from quarterly updates, while the direct manager and peers had specific examples of missed handoffs. That rating moved too, in the other direction.
Afterwards, we changed the process: any rating that sat well away from the other perspectives on the same person needed written evidence before calibration.
Two things made the difference. The evidence rule meant neither conversation became a debate about who was right. And looking at the whole team, not just one person, showed that the first problem belonged to the rater, not the engineer. Without that team view, the strict manager's score would have looked like an honest assessment, because it was one. It was just measured with a different ruler.
What this looks like in practice
before your next calibration, ask every manager to submit two behaviour examples per rating. You will find that some ratings change before the meeting even starts, because the manager could not find the examples.
When does calibration become forced ranking by another name?
The moment it starts from a quota instead of from evidence.
In June 2025, Meta asked managers to place 15 to 20% of employees in its "below expectations" category at mid-year, up from 12 to 15% the previous year, according to reporting on an internal memo. Whatever the business reasons, the mechanism is clear: the distribution was decided before the evidence was reviewed. That is not calibration. It is ranking with a target attached.
This is the line every growing company needs to hold.
| Evidence-based calibration | Forced ranking |
|---|---|
| Starts from what each person did | Starts from how many people must be in each box |
| Ratings move when the evidence says so | Ratings move to hit the target |
| A strong team can be rated strongly | A strong team must still produce low performers |
| Managers bring examples | Managers bargain over names |
| Builds trust in the outcome | Teaches people to compete with each other |
Our earlier post in the Talent Density series, Your Performance Review System Was Designed for a Different Stage, looked at what forced ranking did to collaboration at scale. The short version: when your rating depends on a colleague being rated lower, you stop helping colleagues.
None of this means ratings should drift upwards unchecked. If every team is rated above expectations, calibration should challenge that too, with the same rule. Show the evidence. Evidence-based calibration can move ratings down as easily as up. It just does not decide the answer in advance.
What this looks like in practice
if your calibration meeting starts with a slide showing the target distribution, remove the slide. Start with the evidence. Review the distribution at the end, and ask why it looks the way it does.
How does this connect to the rest of the series?
Calibration is where every idea in this series meets.
It answers our first post, which showed that most of a rating reflects the rater. Calibration is the step that separates the rater from the rating.
It makes our second and third posts workable. Scoring behaviour as well as results, and describing values as observable actions, only helps if those scores are checked against evidence in the same way across teams.
It protects our fifth post. Once pay decisions move out of the review meeting, they depend on ratings that have been calibrated. Calibration is what makes it safe to separate the two conversations.
And it leads to our final post. Calibration is only as good as the framework it calibrates against. If that framework has 25 competencies, no calibration session can check them all. In our last post, we look at why a growing company needs fewer than ten, and how to choose them.
The reframe underneath all of it is simple. A review without calibration is a collection of opinions. A review with calibration is a set of decisions you can explain. That is not bureaucracy. That is the honest part.
Frequently asked questions
- If one manager's rating is unreliable, why don't more ratings fix it?
- Because more ratings from the same biased vantage points just average the same distortions. The Scullen, Mount and Goff study of 4,492 managers rated from six angles found the biggest share of variation, about 62%, still came from each rater's own tendencies rather than the employee's work.
- What does a calibration session actually do?
- It puts managers in a room to compare ratings against shared standards and real evidence, so a 4 from a lenient manager and a 4 from a strict one come to mean the same thing. It corrects inflation and harshness at the same time.
- When does calibration become forced ranking?
- The moment a distribution is decided before the evidence is. If a session opens with a target such as 15 to 20% below expectations, as Meta's 2025 mid-year memo required, it is ranking with a quota attached, not calibration, and it drives out the people you most want to keep.
