Applied AI

AI Evaluation Analyst

₹2–7 LPA
Entry level · 0–1 years · Remote · Gig or permanent · Indicative CTC, India
In demandGrowing fast from a small base. This role did not exist as a graduate hire five years ago, so intakes are smaller than the established tracks — but it is the fastest-growing part of entry-level demand, and competition is lighter because fewer students know it exists.AI-exposedAI is changing what this role does day to day. The work is not disappearing, but what employers screen for is shifting — from producing output to verifying it and owning the decision.
Capability
OversightRubric applicationCritical reading
Disposition & behaviour
ScepticismOpenness to experienceJudgement under uncertainty
Take the qualification testQualify once. Your verified profile goes to employers hiring for this role.

GoMeasure Platform runs GoMeasure Campus — a talent network for final-year students and new graduates, where what you can do is measured from real work rather than claimed on a CV.

We’re looking for AI Evaluation Analysts to join this network ahead of the placement cycle. Judge model output against rubrics, find where it fails and explain why it failed.

You’ll take one qualification test covering oversight, rubric application, critical reading, build a verified profile, and go to the employers hiring into this track — typically ₹2–7 LPA at entry level.

Before you apply — a warning about this market. This corner of the market attracts scams, and students are the target. A genuine employer never asks you to pay to apply, to pay for training or an activation fee, or to pay to withdraw your earnings. Impersonation of known annotation platforms is common. Before committing time to an unpaid assessment, check the platform's recent reviews — unpaid tests and sudden account suspensions are a recurring complaint even on legitimate-looking sites.

Role summary

Judge model output against rubrics, find where it fails and explain why it failed.

What you will do

In the first six months

  • Judge model output against a rubric, consistently, across a lot of items
  • Write the reasoning for a judgement so someone else could reach the same one
  • Find the failure patterns hiding inside individually-acceptable outputs
  • Feed what you find back to the people building the system

By twelve to eighteen months

  • Shape the rubric rather than only applying it
  • Specialise into a domain where your judgement is worth more — medical, legal, financial, linguistic
  • Move toward evaluation design, data quality or product work

Required skills

SkillWhat good looks like at entry level
Critical readingSpotting that a fluent, confident answer is wrong. The entire job, and harder than it sounds.
Rubric applicationApplying the same standard on item two hundred as on item one. Inconsistent evaluation produces data nobody can act on.
Written reasoningExplaining why something failed, specifically enough to be actionable.
Domain knowledgeDepth in any field — medicine, law, finance, a language — is worth more here than in almost any other entry track.
Attention to detailThe errors that matter are usually small and plausible rather than obvious.

Disposition & traits

Skills describe what someone can do when they try hardest. Disposition describes what they typically do — and over a first year, that second question predicts as much as the first. For this track the dispositions that matter most are:

  • Scepticism
  • Openness to experience
  • Judgement under uncertainty
How to read these. Higher is not automatically better — a disposition that helps in one role works against another. These are not a pass mark, not trainable in the way a skill is, and never reported on their own.

What the hiring bar looks like

Compiled by GoMeasure from publicly available accounts, October 2026.

Employer typeWhat they screen for
AI labs & product companiesA practical evaluation exercise using real output with errors planted in it. Structured roles here are the better-paid end, but many require prior paid evaluation or annotation experience as a stated condition.
Data services & BPOVolume hiring with a screening task. The usual entry point, often on short initial contracts.

Eligibility — who actually gets to apply

Be clear-eyed about the two tiers. Entry-level annotation and evaluation work commonly pays ₹10,000–18,000 a month, frequently on contract — materially below the other tracks in this domain. The structured ₹4–7 LPA evaluation roles typically list prior paid annotation or RLHF experience as a requirement rather than a preference. The honest route is to treat the entry tier as a paid way to earn that experience, not as a destination. Degree requirements are genuinely open, including non-technical disciplines.

Eligibility is set per drive and your placement cell’s notice is what actually applies on your campus. One trap worth knowing: a 6.0 CGPA is not always 60%. Where a university converts with (CGPA − 0.75) × 10, a 6.0 is 52.5% — below most employers’ floor. Check which formula yours uses before assuming you qualify.

And on the package: CTC is not take-home. A ₹4 LPA offer lands nearer ₹28,000–32,000 a month once provident fund, gratuity and tax come out, and offers with a large variable or joining-bonus component differ more again. Compare offers on fixed monthly pay, not on the headline. On-campus mass recruiters rarely move off a standard package; startups, GCCs and mid-size firms hiring off-campus often have some room.

Who should apply

Graduates whose degree and interests line up with the work above. Eligibility aside, employers in this track screen on demonstrated capability more than on which campus you attended.

This role is probably not for you if you need variety week to week, or you would find sustained detailed review draining. Both are reasonable; neither survives this job.

Evaluation notice

Results are shared with the student and, with consent, with employers hiring into this track. Scores carry the evidence and the assessment date. Practice and assessed sessions are clearly distinguished before either begins.

What you get back

The outcome of qualifying is a report, not a pass mark. Three layers, each scored and reported separately — there is no single number, deliberately, because a composite hides the trade-off an employer actually needs to see.

01 · Capability

What you can do at your best

Role-specific skills and real work, placed on a proficiency level with the evidence attached.

OversightRubric applicationCritical reading
02 · Disposition

What you typically do

How you tend to work, and the judgement you show in realistic situations. Not a pass mark, and higher is not automatically better.

ScepticismOpenness to experienceJudgement under uncertainty
03 · Alignment

What you are optimising for

What you want from work, and whether a role supplies it. Read as fit rather than quality.

Fit, not good or bad
Who sees it

Every layer carries its own score, its weighting and the date it was assessed — and you see the same report the employer does. The report is shared with employers hiring into this track, with your consent, and you can withdraw it. Individual employers are named on the drive itself, once that employer is participating.

Retaking the test

One qualification attempt per track, with a retake available after one to two months. The wait is deliberate: a retake a week later measures how well you remember the test, not whether anything changed. The gap is long enough for preparation against your reported gaps to actually show.

Get placed in this track

One qualification test puts you in the talent pool for this role, with a verified profile that employers can act on — instead of a CV that looks like every other CV in the stack.

Join the talent pool