Synthetic demonstration

Interactive synthetic demonstration · 3 responses · 1 dimension, before and after anchoring

This demonstration uses entirely synthetic data and simulated annotators. It shows how our evaluation process works — it does not show results on real data, real annotators, or real client work. No human annotators participated. Agreement figures shown are internal process measurements, not product accuracy, not guarantees, and not transferable to any real engagement. Primwright Enterprise does not currently offer these services commercially.

Primwright Enterprise Demonstration 6 of 7

A rubric is built.

“Rate actionability 1–5” — score three synthetic responses against that instruction alone, then against the anchored version with worked examples. Feel the difference an anchor makes, then see the versioning discipline that keeps a rubric honest over time.

01Score, anchor, re-score

Build the dimension with us.

Step 1 The dimension, unanchored

Score these three responses.

Actionability

Rate how actionable the response is, from 1 (not actionable) to 5 (very actionable).

That is the entire instruction — the way most scoring rubrics actually ship. Score each response 1–5. There is no score for you; the point is to feel the ambiguity.

R1 · My invoice shows a charge I don't recognize.

I've opened a billing review for invoice INV-2041. A specialist will investigate and email you the findings within 2 business days; if the charge is an error, the reversal is automatic and you'll see it on your next statement.

Your score for R1 (unanchored)

R2 · Can you change the email on my account?

Sure — I can update that for you. Which email address would you like to use going forward?

Your score for R2 (unanchored)

R3 · The app keeps logging me out on my phone.

Sorry you're running into that. Our team is aware of login issues and working on improvements to the mobile experience.

Your score for R3 (unanchored)

0 of 3 scored.

Step 2 The dimension, anchored

Now score them again — with anchors.

Step 3 The versioning discipline

What happens when calibration finds a gap.

◈Method, stated honestly

What this demonstrates — and what it doesn’t.

It demonstrates

  • Why rubrics are built, not just written: an unanchored dimension admits multiple defensible readings, and the spread is where the work is.
  • The mechanism by which anchors constrain interpretation — shown as spread narrowing across simulated judges.
  • The versioning discipline: calibration findings become prospective guideline amendments, never silent edits.

It does not claim

  • That the spread narrowing is a human calibration result — the judges are simulated personas.
  • That anchored rubrics eliminate disagreement — they constrain it, and the residual spread is reported.
  • Any accuracy or quality figure for real evaluation work.

→Next step

Rubrics are agreed before scoring begins.

In a real pilot, the rubric is approved by you before a single output is scored. Start with a scoping conversation.