Synthetic demonstration

Interactive synthetic demonstration · 4 pairs · fictional Northwind Desk policy sheet

This demonstration uses entirely synthetic data and simulated annotators. It shows how our evaluation process works — it does not show results on real data, real annotators, or real client work. No human annotators participated. Agreement figures shown are internal process measurements, not product accuracy, not guarantees, and not transferable to any real engagement. Primwright Enterprise does not currently offer these services commercially.

Primwright Enterprise Demonstration 5 of 7

Pairwise preference, judged twice.

Two responses, one prompt: which is better — and why? Vote on four synthetic pairs, then reveal the method view to see how the procedure records each preference, separates defect-driven from style-driven judgments, and resolves a both-defective pair as a tie.

01Vote, then reveal

Cast your four votes.

0 of 4 pairs voted — vote on each pair, then reveal the method view to see how the procedure handles it.

Pair 1 of 4Synthetic

Customer promptMy order arrived with the wrong item. What are my options?

Response A

I'm sorry about the mix-up. You can request a replacement within 14 days of delivery, and we'll cover return shipping for the wrong item. If you'd prefer a refund instead, that is also available within the same window. A support agent will confirm the details shortly.

Response B

Sorry for the error. You can get a refund within 30 days of delivery. Return shipping is on you, and replacement isn't an option for mis-shipped orders.

Which response do you prefer? (This is not a test — there is no score for you.)

Pair 2 of 4Synthetic

Customer promptHow long do you keep my support chat transcripts?

Response A

We retain support transcripts for 90 days, then they are deleted. You can request earlier deletion at any time by contacting support.

Response B

Your transcripts are stored indefinitely so we can improve our service. They are kept for 90 days per our retention policy.

Which response do you prefer? (This is not a test — there is no score for you.)

Pair 3 of 4Synthetic

Customer promptI was charged twice for my subscription this month. Please fix it.

Response A

I've flagged the duplicate charge for review. A billing specialist will reverse the second charge within 3 business days and confirm by email. Nothing else you need to do.

Response B

I can see the duplicate charge and I'm really sorry about the hassle this caused — I know billing surprises are stressful. A billing specialist will reverse the duplicate charge within 3 business days and confirm by email.

Which response do you prefer? (This is not a test — there is no score for you.)

Pair 4 of 4Synthetic

Customer promptCan I get a refund for a service I cancelled last week?

Response A

Refunds are available within 14 days of purchase, so yes — you are within the window. I have started the refund; expect it within 5 business days.

Response B

Since you cancelled last week you are within the 14-day window. Your account will be credited immediately and the refund is guaranteed within 24 hours.

Which response do you prefer? (This is not a test — there is no score for you.)

Policy sheet for this fictional scenario: 14 days refund window · 90 days retention · 4-hour first response first response. Simulated judges: J1 (steady), J2 (tone-sensitive), J3 (noisy). Run seed 20261002.

◈Method, stated honestly

What this demonstrates — and what it doesn’t.

It demonstrates

  • How a preference judgment is recorded: the vote plus its basis — defect-driven vs style-driven — because the two are not interchangeable evidence.
  • How a split panel escalates to adjudication, and how the dominance order resolves a both-defective pair as a tie.
  • That disagreement is treated as signal: the procedure is built to surface and resolve it, not to hide it.

It does not claim

  • That the simulated panel’s votes resemble any human judgment — the judges are simulated personas.
  • That any preference shown here says anything about a real model’s quality.
  • That the planted gold is “objectively correct” — it is authored in the build module to demonstrate the procedure.

→Next step

This is the procedure a pilot would run.

Pairwise preference evaluation is one of our evaluation services. Every pilot starts with a written rubric and a scoping conversation — never a handshake.