Primwright Enterprise Product

Rubric evaluation

Pilot-scoped service · Anchored scoring · English / Arabic

“A beats B” doesn’t tell you whether A is good enough to ship. We score your AI’s outputs on anchored rubrics — every scale point defined in words with examples — and report agreement per dimension, including the dimensions that don’t work.

01The problem

Rankings don’t set thresholds.

“A beats B” doesn’t tell a team whether A is good enough to ship, whether the judge agrees with humans on the dimensions that matter, or whether last month’s “4.2” means the same thing as this month’s “4.2.” Absolute scoring on unanchored scales produces numbers that drift with rater mood, guideline version, and time — and teams make ship, model-switch, and vendor decisions on them anyway.

You need scores that mean the same thing across raters and across time — with the noise reported per dimension instead of hidden inside an average.

02What it is

Anchors, not numbers.

Every scale point is defined in words with examples — what makes a 3 a 3 is written down, versioned, and signed. This is the specific craft most internal rubrics skip and most vendor decks gloss over, and it is where evaluations quietly fail.

The rubric is a reusable instrument you keep: the anchored rubric plus the calibration protocol becomes your standing acceptance spec — for vendor contracts, for regression rounds, for judge calibration. Round one builds the instrument; later rounds score against the locked rubric, so repeats cost less by design.

03Expected inputs

What you provide

  • Items to score: prompts with single model responses (or other text artifacts), agreed schema, stable unit IDs.
  • A scoring rubric: yours, or one drafted with you. It must define dimensions, anchored scales with written scale-point descriptions, and how dimension scores combine — if they do.
  • Gold and anchor examples per dimension per scale point, with written rationales — sealed from raters during qualification.
  • Output schema: the exact shape for scores, agreed before production. Plus a data-sensitivity declaration — no regulated or sensitive content.

04Workflow

How a pilot works

  1. Intake & sensitivity screen

    We verify schema, ID uniqueness, and item completeness, and run the sensitivity screen. Out-of-scope inputs stop here with a written intake report.

  2. Rubric lock

    The anchored rubric is signed by you and versioned. Every scale point has a written anchor with examples. Changes pause production; affected items are re-scored.

  3. Rater qualification

    Blind qualification against sealed golds, with per-dimension agreement measured. Below-threshold raters are not assigned — absolute scoring is harder to calibrate than preference, and the bars reflect that.

  4. Calibration batch

    Double-scored items; per-dimension disagreements adjudicated; anchors tightened through versioned patches where ambiguity recurs.

  5. Production scoring

    Raters score each item on every dimension per the anchors — flagging ambiguous items, guideline gaps, and data defects instead of guessing.

  6. Agreement & sampling review

    A defined subset is double-scored and per-dimension agreement is computed and reported. An independent reviewer re-scores a stratified sample blind; deviation rates are reported per dimension.

  7. Adjudication

    Disagreements go to an adjudicator — never the producer of the disputed units — with rulings logged.

  8. Delivery & acceptance

    Score dataset, rubric version, agreement figures, adjudication log, and delivery memo — then your acceptance review.

05Deliverables

What you receive

  • Score dataset in the agreed schema: unit ID, per-dimension scores, flags, rater, rubric version.
  • The anchored rubric, versioned with a changelog — the reusable instrument you keep.
  • QA report: per-dimension inter-rater agreement, reviewer sampling deviation rates, qualification results, flag rates — including any dimensions reported as unreliable.
  • Complete adjudication log and delivery memo.

06Quality controls

How we keep it honest

  • Anchored scales only — no bare-number scoring; every scale point defined in words with examples.
  • Sealed golds for qualification, with per-dimension gold agreement reported.
  • Per-dimension agreement on a double-scored subset — reported per dimension, never blended into a single number.
  • Reviewer sampling with per-dimension deviation rates; adjudication of all disagreements; flag-rather-than-guess discipline.
  • Unreliable-dimension policy: if agreement on a dimension cannot be raised, the dimension is reported as unreliable rather than shipped as if precise. A stated limitation, not a hidden one.
  • Rubric patching: recurring ambiguity triggers a versioned anchor tightening, never silent convention drift.

07Scope and limits

Stated up front

What this product does not include — stated here, so there are no surprises in scoping:

  • Scores are structured human judgments, not objective measurements of model quality — dimension-level noise is reported, not hidden.
  • No invented thresholds or scoring. Calibration can fail honestly — a round may find that your dimensions don’t work, and that finding is reported, not discounted.
  • No safety or red-team evaluation; no regulated or sensitive data or content.
  • No specialist medical, legal, or financial scoring without separately credentialed reviewers.
  • No claim that rubric scores predict downstream model performance or business outcomes.
  • English and Arabic only. No guaranteed score-accuracy figures; acceptance criteria are agreed per pilot, in writing, before work begins.

→Next step

Start with a scoped pilot.

Tell us about your project. We review every inquiry against our scope — if it’s a fit, we schedule a scoping conversation.