---
title: "Your Performance Reviews Measure Your Managers, Not Your People"
description: "62% of a performance rating reflects the person giving it. Here is why ratings drift, and how calibration makes them fair again."
url: "https://sageo.ai/blog/competency-and-culture/reviews-measure-managers"
updated: "2026-10-05T00:00:00.000Z"
author: "Deepti Gupta"
section: "Competency and Culture"
image: "https://sageo.ai/blog/covers/reviews-measure-managers.png"
source: "Sageo (sageo.ai)"
---


# Your Performance Reviews Measure Your Managers, Not Your People

_62% of a performance rating reflects the person giving it. Here is why ratings drift, and how calibration makes them fair again._

By Deepti Gupta.

## About This Series

This is the first post in our seven-part Competency and Culture series. The series makes one argument: culture is what you measure and reward, not what you write on the wall. We start with the measurement itself, because every people decision that follows (who gets promoted, who gets paid more, who is asked to leave) rests on a rating. If the rating tells you more about the manager than about the person, everything built on top of it inherits the error. Get this right, and the posts that follow, on brilliant jerks, values, culture fit, pay conversations, calibration, and competency design, become problems you are solving on solid ground.

Two managers can watch the same person do the same work for a year and hand in scores two points apart. Neither of them is lying. Neither of them is careless. They are simply measuring against different rulers, and nobody has checked the rulers.

Most founders assume a performance rating is a measurement of the employee. The research says otherwise. The largest share of any rating describes the person giving it: their standards, their habits, their blind spots, and how close they sit to the actual work. The employee's real performance is a smaller part of the number than almost anyone expects.

That matters more in a growing company than anywhere else. At 50 or 150 people, one manager's rating can decide a promotion, a pay increase, or an exit, and there are rarely enough other data points to catch the mistake.

## How much of a performance rating is actually about the employee?

The most cited answer comes from [a study published in the *Journal of Applied Psychology* in 2000](https://www.researchgate.net/publication/12202559_Understanding_the_Latent_Structure_of_Job_Performance_Ratings) by Steven Scullen, Michael Mount, and Maynard Goff. They looked at 4,492 managers, each rated by two bosses, two peers, and two direct reports. Because every person had six raters, the researchers could separate how much of each score came from the person being rated and how much came from the person doing the rating.

The result has been repeated in boardrooms ever since. About 62% of the variation in ratings came from the individual rater's own tendencies, what researchers call the idiosyncratic rater effect. Actual performance accounted for about 21%.

| What drives a performance rating | Share of variation |
| --- | --- |
| The rater's own tendencies (standards, habits, biases) | About 62% |
| The employee's actual performance | About 21% |
| Everything else (vantage point, noise, context) | The remainder |

_Source: [Scullen, Mount and Goff, Journal of Applied Psychology, 2000](https://www.researchgate.net/publication/12202559_Understanding_the_Latent_Structure_of_Job_Performance_Ratings)._

The study is from 2000, and it remains the landmark because nothing since has contradicted it. More recent data shows the problem has not gone away. In 2023, [Gallup surveyed 135 chief people officers](https://www.gallup.com/workplace/644717/chros-think-performance-management-system-works.aspx) at Fortune 500 companies and 18,665 employees. Only 2% of the people officers strongly agreed that their performance system inspires employees to improve. Only 22% of employees strongly agreed that their review process is fair and transparent.

Those two numbers are connected. When most of a score reflects the rater, employees notice. They compare notes. They see that the same work earns a 3 in one team and a 4 in another. The review stops feeling like feedback and starts feeling like luck.

And the raters themselves are under strain. [Gallup's 2026 State of the Global Workplace report](https://www.gallup.com/workplace/349484/state-of-the-global-workplace.aspx) found that manager engagement fell to 22% in 2025, down nine points since 2022. The people handing in ratings are, on average, less engaged than they have been in years. That does not make them worse people. It makes their ratings less reliable.

> **What this looks like in practice**
>
> if you have ever looked at a promotion list and thought "that team always gets rated higher," you were probably right. You were not looking at a stronger team. You were looking at a more generous manager.

## Why do strict managers and distant managers get it wrong in opposite directions?

Rater error is not random. It tends to come from two places, and they push ratings in opposite directions.

The first is the **strict manager**. Some managers carry a personal standard that is simply higher than everyone else's. They are often excellent operators, and they mean well. But their 3 is another manager's 4, and nobody has adjusted for it. The people who report to them are marked down, cycle after cycle, for having the wrong boss.

The second is the **distant manager**. Skip-level managers and senior leaders often rate people they see only in reviews, demos, and leadership updates. They see the presentation, not the preparation. They see the result, not the three handoffs that were missed on the way. Their ratings drift upward for people who present well, and downward for quiet people who hold the work together.

We have seen both, in two different teams.

> **From our experience**
>
> In one team, a senior engineer received a noticeably lower rating from their direct manager than from their skip-level manager, their peers, and their own direct reports. Nothing in their delivery record explained the gap. The manager's other ratings did: this manager scored everyone on the team lower than colleagues leading comparable teams. The manager was not being unfair to one person. They were strict with everyone.
> In a completely different team, we saw the opposite. A skip-level manager rated a team lead as exceptional, while the direct manager and peers described missed handoffs and a team that had learned to work around them. The skip manager saw the quarterly updates, not the day-to-day reality.

In both cases the rating was honest, and in both cases it was wrong. The first would have cost a strong engineer a promotion. The second would have promoted someone on the strength of a good presentation, and handed their team a bigger version of the same problem.

Neither problem shows up if you look at one rating at a time. Both show up the moment you put the rating next to the other perspectives on the same person, and next to how the same rater scores everyone else.

> **What this looks like in practice**
>
> when a direct manager's rating and a skip manager's rating disagree by a full point or more, do not average them. Ask each for two specific examples of behaviour from the review period. The rater with examples from the actual work usually wins.

## Which biases sit inside every rating?

Strictness and distance are the two you can see. Underneath them sit a handful of well-documented biases that affect every rater, including the good ones. None of them require bad intent. They are how human judgement works when nobody gives it a structure.

| Bias | What it looks like in a review | What it costs you |
| --- | --- | --- |
| Leniency and severity | The same work is "exceptional" to one manager and "meets expectations" to another | Pay and promotion depend on which team someone joined |
| Halo and horn | One strong launch, or one bad incident, colours every other score | People are rated on their most visible moment, not their year |
| Similarity | Managers rate people who think, talk, and work like them more highly | Teams slowly fill with one type of person |
| Recency | The last six weeks outweigh the first ten months | Steady contributors lose to people who finish strong |
| Central tendency | Everyone gets a 3 because it is safe | Your best and weakest people become invisible |

The uncomfortable part is that experience does not cure these. A manager who has run fifty review cycles has simply had fifty cycles to settle into their own pattern. What cures them is structure: clear behaviour descriptions for each rating, [more than one perspective on each person](https://sageo.ai/guides/best-360-feedback-software), and a moment where ratings are compared before they become decisions.

This is also why asking managers to "be more objective" never works. You cannot instruct your way out of a bias that people cannot see in themselves. One practical fix, first made famous by [Deloitte's redesign of its reviews](https://www.physicianleaders.org/articles/reinventing-performance-management), is to stop asking managers to rate a person's qualities and ask instead what they would *do*: would they give this person the highest possible increase, would they always want them on their team, are they ready for promotion today? People are far more reliable reporting their own intentions than judging someone else's character.

> **What this looks like in practice**
>
> look at the spread of ratings for each manager before calibration. A manager whose team is all 3s, or all 4s and 5s, is telling you something about their rating style, not about their team.

## Why does calibration across a whole team matter more than any single review?

If most of a rating reflects the rater, the fix is not a better form. It is a second look. Calibration is that second look: a structured conversation where ratings are compared across people, perspectives, and managers before they turn into pay, promotion, or exit decisions.

The two cases above were calibrated separately, because they were very different teams with different managers and different work. But the method was the same in both. Every score was compared with the other perspectives on the same person, and with how that rater scored everyone else on the team. Nobody was asked to defend a number. Everyone was asked for evidence.

In the first team, the strict manager's own examples described a solid, reliable engineer, which did not match the score they had given. In the second, the skip manager's examples all came from quarterly updates, while the direct manager and peers had specific examples from the day-to-day work. Both ratings changed to match the evidence, one up and one down.

That is the real value of calibration across a team. A single review can only tell you what one person thinks. A calibrated team view tells you three things a single review cannot:

1. **Whether the rater is the pattern.** If one manager's scores sit consistently below or above their peers, the gap belongs to the manager, not the team.
2. **Whether the perspectives agree.** When a direct manager, a skip manager, peers, and direct reports all see something different, that disagreement is the most useful information in the review.
3. **Whether the rating matches the evidence.** A score without two concrete examples of behaviour is an opinion. Calibration turns opinions back into evidence, or exposes them.

One warning. Calibration can go wrong too. If a calibration meeting starts with "we need 10% in the bottom box," it is not calibration. It is forced ranking with a friendlier name. We looked at what forced ranking does to the people you most want to keep in [Your Performance Review System Was Designed for a Different Stage](https://sageo.ai/blog/talent-density/why-performance-reviews-fail), and we come back to it in the [sixth post of this series](https://sageo.ai/blog/competency-and-culture/calibration-done-right).

> **What this looks like in practice**
>
> calibrate by team, not just by company. Put every rating for a team on one page, next to the rater and the other perspectives on each person, before anyone discusses individual cases.

## How can you check your own ratings before they cost you a good employee?

You do not need a full calibration process to spot a rater problem. You need two numbers per manager: their average rating, and how spread out their ratings are. Put those next to the rest of the organisation and the pattern usually jumps out.

We built a simple self-check to make that quick. It is not a verdict on any manager, and it does not store anything you enter. It is a prompt for the conversation that calibration should have.

_[Interactive tool: rater-bias-check. Available on the web version of this page.]_

If the check flags a gap, that is not an accusation. Some teams are stronger than others. But the burden of proof shifts: a manager whose average sits well above or below their peers should be ready to show the behaviour examples behind it.

Founders can run the same check across the whole company. Line up every manager's average and spread on one page. If one manager rates half a point below everyone else, it is worth asking three questions: are their people being paid less than peers doing similar work, do their strongest performers have a reason to leave, and has anyone told the manager?

> **What this looks like in practice**
>
> before the next review cycle closes, ask every manager for their average rating and the lowest and highest score they gave. It takes five minutes, and it will tell you where calibration needs to start.

## How does this connect to the rest of the series?

Every idea in this series rests on the rating being right. If most of a score reflects the rater, then everything built on top of it inherits the error.

It connects to [our next post, on the brilliant jerk tax](https://sageo.ai/blog/competency-and-culture/cost-of-toxic-high-performers). When ratings only capture *what* someone delivered, a high performer who damages their team can score well for years. Separating the *what* from the *how* only works if both scores are calibrated.

It connects to [values](https://sageo.ai/blog/competency-and-culture/values-fail-in-the-review). A value like "ownership" gives raters nothing to hold on to unless it is described as behaviour at each level. Vague values are where rater bias does most of its work.

It connects to culture fit, pay conversations, and competency design, because each of them depends on a rating that means the same thing in every team. And it leads directly to [our sixth post, on calibration itself](https://sageo.ai/blog/competency-and-culture/calibration-done-right): how to run it well, and how to stop it from becoming forced ranking by another name.

The reframe underneath all of it is simple. A rating is not a measurement of a person. It is a measurement taken by a person. Until you check the instrument, you cannot trust the reading. Calibration, clear behaviour descriptions, and more than one perspective are how you check it, and none of them require a large people team to get started.

In the next post, we look at what happens when the numbers are good and the behaviour is not: why founders protect their most expensive hires, and how to see the cost before your team pays it.

## Frequently asked questions

### How much of a performance rating actually reflects the employee?

Less than most leaders assume. The landmark Scullen, Mount and Goff study found about 62% of the variation in ratings came from the individual rater's own tendencies, the idiosyncratic rater effect, while actual performance accounted for only about 21%. A rating often tells you as much about the manager as the person being rated.

### Does calibration make reviews fairer?

Yes, when it is done to reconcile standards rather than hit a quota. Comparing ratings across a team surfaces the strict manager and the lenient one and pulls both toward a shared bar, so an employee's score depends on their work rather than on which manager they happened to report to.

### Why do performance reviews still feel unfair?

Because the rater problem has not gone away. In Gallup's 2023 research only 2% of chief people officers strongly agreed their system inspires employees to improve, and only 22% of employees felt their review process was fair and transparent. Structure and calibration, not better intentions, are what move those numbers.
