How we measure

A skill score you can audit, not just trust.

Verified skills intelligence built on one rule: the evidence decides the score, and a fixed formula — not a model's opinion — computes it.

5 proof sources
2 axes: level & confidence
1 append-only ledger
0 black boxes
Live readoutSkill · Retrieval-Augmented Generation
0Level / 5
confidence0.00
independent sources3
maturityverified
Assessment 360 review Manager review
Why the number needs a method 01

Most skill scores are opinions in a trench coat.

Ask an AI to rate someone and it will — fluently, instantly, and differently every time you ask. A single test is honest, but it's a snapshot: one room, one hour, one format. Neither tells a hiring manager the thing they actually need: how good, and how sure.

Our method treats a score as a claim that has to be defensible. Every point of it traces back to something that happened — an artifact shipped, an answer graded, a peer's confirmation — and the path from that evidence to the number is fixed, inspectable and reproducible.

If you removed the person's self-description, would the evidence still exist? That question runs through everything below.
Principle I · Evidence-centered 02

Every score traces back to evidence.

Nothing enters a score without a source. Each piece of evidence — a graded assessment, a merged pull request, a manager's rating, a peer's endorsement — is written once to an append-only ledger with its origin, timestamp and a snippet of what was observed. The ledger is the record; the score is a reading of it.

Because the trail is complete and immutable, any score can be opened up: this level, because these pieces of evidence, weighted this way. No orphan numbers, no “the model felt strongly.”

observable work  →  typed evidence  →  deterministic score  →  a level you can defend
Principle II · The wall 03

AI reads the language. A formula does the math.

We put a hard wall between interpretation and scoring. Language models are extraordinary at reading messy human work and unreliable at being consistent about a number — so we only let them do the first job.

AI · interpretation
  • Reads an artifact, answer or message
  • Decides what kind of evidence it is
  • Emits a typed tag — never a number
  • Can be re-run or corrected
the wall
Formula · scoring
  • Takes the typed evidence as input
  • Applies fixed weights and rules
  • Same evidence → same score, always
  • No temperature, no prompt, no drift
The consequence buyers care about: your score doesn't change because a model had a different day. It changes only when the evidence changes or the published formula does — and both are versioned.
Principle III · Claims → proof 04

Did they do the thing, or just talk about it?

The biggest gap in any skill signal is between a claim — self-reported, unverified — and proof — something that was actually observed or graded. Every source sits somewhere on that line and carries a trust weight to match. A claim only starts to count once it crosses the proof line.

SourceWhat it isTrust weight
Self-declared“I have this skill.” No artifact behind it.
0.20
Profile / résuméListed, not shown.
0.25
Work sampleAttached, but unverified.
0.30
Manager noteAn informal vouch.
0.40
Certificate (claimed)Uploaded, unchecked.
0.50
The proof line — evidence becomes verifiable
Manager reviewAn expert grades the actual proof.
0.60
CertificationA verified, checkable credential.
0.70
360 reviewIndependent peers corroborate.
0.80
InterviewLive, structured, scored.
0.80
AssessmentA proctored, graded task they sat.
0.90
A claim

A résumé line reading “expert in retrieval-augmented generation.” Remove the person and nothing is left behind to check.

Proof

A retrieval pipeline merged to production with a benchmark showing a measurable lift. The work exists whether or not they mention it.

Principle IV · Corroboration 05

Independent sources converge into one measurement.

No single lens sees a whole person. We fuse every piece of evidence for a skill into one reading — and corroboration rewards independence, not volume. Three messages from one channel are one source agreeing with itself. An assessment, a shipped artifact and a manager's rating are three different vantage points landing on the same answer — and that is what earns confidence.

Lead
One uncorroborated signal. We hold a confidence, but withhold the level until something confirms it.
Indicative
Some evidence, not yet independent. A provisional read, labelled as such.
Verified
Independent sources agree. The measurement is decision-grade.
every score is stamped with its maturity — so you always know how it was earned
What the number reports 06

How good, and how sure — kept separate.

A score is never one number pretending to answer two questions. We report both, and never let one contaminate the other.

Level · 1–5
How good?

The proficiency the evidence supports, on a defined scale from novice to expert.

Confidence · 0–1
How sure?

How much independent, corroborating evidence stands behind that level.

The invariant we never break: absence of evidence lowers confidence, never level. A quiet month makes us less certain — it does not make someone worse at their job.
Assessed × Proven 07

Where capability and application meet.

When we hold a graded-assessment signal and proof from real work, the pair tells a story neither could alone. This is the read a manager acts on.

Proven in real work →
Practical competence

Delivers in the real world; the formal test undersold them. Trust the work.

Decision-grade

Capability and application agree. The strongest signal we produce.

Development need

Neither signal is there yet. A clear, honest place to start learning.

Latent capability

Knows it, hasn't shown it in the work yet. A coaching opportunity, not a gap.

Assessed in a graded task →
Principle V · Freshness 08

Proof has a shelf life.

A score has to behave the way skill actually behaves over time. Proof carries a validity window — a certification expires, an assessment ages — and once it lapses we flag it stale and prompt a re-verification rather than quietly trusting old evidence.

But ageing evidence only softens confidence. The level stays at its last measured reading until new evidence moves it — a slow month lowers how sure we are, not how good someone is.

level & confidence52 weeks
level — the last measured reading confidence — decays as proof ages, recovers on re-verify
Making it actionable 09

Measured against the level the role requires.

A level on its own answers how good. Read against the level a role needs, it answers the question a manager actually has: are they there yet? Every skill is measured against its required level — meets, approaching or below — so a score becomes a staffing, hiring or development decision, not just a rating.

On the roadmap: role-based norms that turn a level into a percentile — “ahead of 88% of practitioners we've measured” — the step from a scoring engine to a full psychometric one.
Reading it honestly 10

What our score is — and what it is not.

A score is

  • Evidence, weighted. A reading of what was actually observed.
  • Two answers. How good, and how sure — always together.
  • Reproducible. The same evidence yields the same number.
  • Openable. Every point traces to its sources.

A score is not

  • A verdict on a person. A low score can simply mean “not yet observed.”
  • Comparable across gaps. Different evidence coverage isn't a fair contest.
  • An automated decision. It informs a human — it doesn't replace one.
  • A model's guess. The formula, not the language model, sets it.
The short version

Evidence decides. A formula computes. A human acts.

That's the whole method — and the reason a GoMeasure score holds up in the room where it matters.