A skill score you can audit, not just trust.
Verified skills intelligence built on one rule: the evidence decides the score, and a fixed formula — not a model's opinion — computes it.
Most skill scores are opinions in a trench coat.
Ask an AI to rate someone and it will — fluently, instantly, and differently every time you ask. A single test is honest, but it's a snapshot: one room, one hour, one format. Neither tells a hiring manager the thing they actually need: how good, and how sure.
Our method treats a score as a claim that has to be defensible. Every point of it traces back to something that happened — an artifact shipped, an answer graded, a peer's confirmation — and the path from that evidence to the number is fixed, inspectable and reproducible.
Every score traces back to evidence.
Nothing enters a score without a source. Each piece of evidence — a graded assessment, a merged pull request, a manager's rating, a peer's endorsement — is written once to an append-only ledger with its origin, timestamp and a snippet of what was observed. The ledger is the record; the score is a reading of it.
Because the trail is complete and immutable, any score can be opened up: this level, because these pieces of evidence, weighted this way. No orphan numbers, no “the model felt strongly.”
AI reads the language. A formula does the math.
We put a hard wall between interpretation and scoring. Language models are extraordinary at reading messy human work and unreliable at being consistent about a number — so we only let them do the first job.
- Reads an artifact, answer or message
- Decides what kind of evidence it is
- Emits a typed tag — never a number
- Can be re-run or corrected
- Takes the typed evidence as input
- Applies fixed weights and rules
- Same evidence → same score, always
- No temperature, no prompt, no drift
Did they do the thing, or just talk about it?
The biggest gap in any skill signal is between a claim — self-reported, unverified — and proof — something that was actually observed or graded. Every source sits somewhere on that line and carries a trust weight to match. A claim only starts to count once it crosses the proof line.
| Source | What it is | Trust weight |
|---|---|---|
| Self-declared | “I have this skill.” No artifact behind it. | 0.20 |
| Profile / résumé | Listed, not shown. | 0.25 |
| Work sample | Attached, but unverified. | 0.30 |
| Manager note | An informal vouch. | 0.40 |
| Certificate (claimed) | Uploaded, unchecked. | 0.50 |
| The proof line — evidence becomes verifiable | ||
| Manager review | An expert grades the actual proof. | 0.60 |
| Certification | A verified, checkable credential. | 0.70 |
| 360 review | Independent peers corroborate. | 0.80 |
| Interview | Live, structured, scored. | 0.80 |
| Assessment | A proctored, graded task they sat. | 0.90 |
A résumé line reading “expert in retrieval-augmented generation.” Remove the person and nothing is left behind to check.
A retrieval pipeline merged to production with a benchmark showing a measurable lift. The work exists whether or not they mention it.
Independent sources converge into one measurement.
No single lens sees a whole person. We fuse every piece of evidence for a skill into one reading — and corroboration rewards independence, not volume. Three messages from one channel are one source agreeing with itself. An assessment, a shipped artifact and a manager's rating are three different vantage points landing on the same answer — and that is what earns confidence.
How good, and how sure — kept separate.
A score is never one number pretending to answer two questions. We report both, and never let one contaminate the other.
The proficiency the evidence supports, on a defined scale from novice to expert.
How much independent, corroborating evidence stands behind that level.
Where capability and application meet.
When we hold a graded-assessment signal and proof from real work, the pair tells a story neither could alone. This is the read a manager acts on.
Delivers in the real world; the formal test undersold them. Trust the work.
Capability and application agree. The strongest signal we produce.
Neither signal is there yet. A clear, honest place to start learning.
Knows it, hasn't shown it in the work yet. A coaching opportunity, not a gap.
Proof has a shelf life.
A score has to behave the way skill actually behaves over time. Proof carries a validity window — a certification expires, an assessment ages — and once it lapses we flag it stale and prompt a re-verification rather than quietly trusting old evidence.
But ageing evidence only softens confidence. The level stays at its last measured reading until new evidence moves it — a slow month lowers how sure we are, not how good someone is.
Measured against the level the role requires.
A level on its own answers how good. Read against the level a role needs, it answers the question a manager actually has: are they there yet? Every skill is measured against its required level — meets, approaching or below — so a score becomes a staffing, hiring or development decision, not just a rating.
What our score is — and what it is not.
A score is
- Evidence, weighted. A reading of what was actually observed.
- Two answers. How good, and how sure — always together.
- Reproducible. The same evidence yields the same number.
- Openable. Every point traces to its sources.
A score is not
- A verdict on a person. A low score can simply mean “not yet observed.”
- Comparable across gaps. Different evidence coverage isn't a fair contest.
- An automated decision. It informs a human — it doesn't replace one.
- A model's guess. The formula, not the language model, sets it.
Evidence decides. A formula computes. A human acts.
That's the whole method — and the reason a GoMeasure score holds up in the room where it matters.