Your best performer might be your weakest engineer
Pull up your last performance calibration and find the name everyone agrees is a star. Now answer a harder question: could that person do the same job on a different team, in a different market, without the account they inherited three years ago?
If you can't answer that with evidence, you don't actually know what you think you know. A performance rating measures a result — and results are produced by far more than the person. Your top performer might be genuinely, provably brilliant. Or they might be an average operator riding a strong team, a soft territory and a product that sells itself. From the dashboard, the two are indistinguishable. From the org chart, they get promoted identically. Only one of them survives contact with a harder job.
Performance tells you what happened. It does not tell you who made it happen.
Why results lie
A result is a team sport wearing an individual's name. Any of these can inflate a rating without proving a single underlying skill:
- The territory. A rep on the enterprise heartland out-bills a better rep stuck in a dead patch — every quarter, on skill they never had to use.
- The team. Strong colleagues, a great manager, good tooling. The output is collective; the credit is individual.
- Tenure and relationships. Inherited accounts and internal goodwill compound over years. That's position, not capability.
- The tailwind. A rising market, a category on fire, a product that hit product-market fit before they arrived.
- The easy comparison. A strong-looking number graded on a curve against a weak peer set.
None of this means top performers aren't skilled. It means a high rating is consistent with high skill and equally consistent with a favourable context — and a number that can't tell those two apart isn't a measurement. It's a guess with a decimal point.
Two people, one rating
Here's the trap in a single table: two "top 10%" ratings that mean completely different things — and a performance system that cannot see the difference.
Promote both on the strength of the rating and you've made one great bet and one expensive mistake — and you won't find out which is which until the mistake is running a bigger team.
The moment it breaks
Context-inflated performance is a debt, and it comes due at the worst possible time — precisely when you lean on the rating:
- Promotion. You move your best "performer" into a role that needs the exact skill the context was hiding. This is the Peter Principle with a mechanism: people rise on borrowed results until the borrowing stops.
- Succession. Your bench is a list of high performers. But a bench is a bet on capability under new conditions — the one thing a performance rating never measured.
- Reorg or market turn. Strip away the team, the territory and the tailwind, and the borrowed performers crater while the quietly-capable ones hold the line.
Three questions that separate capability from context
Before you bet a promotion or a succession slot on a rating, interrogate it:
- Would it travel? If you moved this person to a neutral context tomorrow, what would you expect to happen — and on what evidence?
- Can they show it, not just tell it? Can they explain the how, reproduce the work, and pass a task that isn't cushioned by their current setup?
- Does evidence exist independent of the result? A graded assessment, a shipped artifact, an independent review — something that would still be true if you deleted the scoreboard.
That last one is the whole game. If the only evidence for a skill is the result, you haven't measured the skill — you've measured the situation.
Measure capability directly
The fix isn't to distrust your best people. It's to stop asking one number to answer two questions. Keep performance for what it's genuinely good at — pay, recognition, accountability. And measure capability separately, on evidence: what each person can actually do, at what level, with how much confidence — independent of whether this quarter's context flattered them.
That's exactly what GoMeasure does. We measure skills from evidence — a graded assessment, a live interview, real work, an independent review — and report two numbers, never one: how good (a level) and how sure (confidence). A skill only counts once it crosses from claim to proof, so a rating inflated by an easy quarter can't quietly masquerade as capability. Here's exactly how we measure it.
Do that, and the top of your calibration and the top of your skill map can finally be laid side by side. The names that appear on both are your real stars. The names on one but not the other — the borrowed performers and the blocked talent — are the most valuable thing you'll learn about your workforce all year.
The takeaway
- A result is a team sport. Territory, team, tenure and tailwind inflate a rating without proving a single skill.
- Two identical ratings can hide opposite realities — one skill-driven, one context-driven — and performance management can't tell them apart.
- The bill comes due on promotion, succession and reorg, exactly when you're betting on capability you never actually measured.
- If the only evidence for a skill is the result, you measured the situation. Measure capability on independent evidence — a level and a confidence, claim turned to proof.
Curious where your organisation stands? The 2-minute Skills Readiness check shows how much of what you "know" about your people is proof — and how much is a rating in a trench coat. For the fuller split between the two, read performance and skills: why you need both.
The whitepaper, the playbooks and new research — by email.
Skills intelligence for HR and business leaders. We'll send the “Jobs to Skills” whitepaper and share new frameworks as we publish them. No noise.
One email now, occasional research later. Unsubscribe anytime.