Primwright Enterprise Product

Model-output evaluation

Pilot-scoped service · Human review · English / Arabic

You are about to put a chatbot or AI feature in front of customers. We independently review what it produces — scored against criteria you approve, defects classified with measured frequencies — so ship, delay, and cut decisions rest on evidence instead of hope.

01The problem

What breaks, how often, and where?

The team is about to ship an LLM-powered feature and cannot answer the only question that matters. Automated evals grade the model’s homework with another model; engineers “eyeball 200 responses” with no methodology; public benchmarks don’t cover your prompts, your policy, or your tone. You are deciding ship, delay, or cut on measurement you privately don’t trust.

When the question is comparative — “is response A or B better?” — head-to-head judging with hidden model identities and mandatory rationales is available inside this service as a method, never as a separate product.

02What it is

The human-agreement layer.

We are the layer your automated eval stack is missing. We don’t compete with eval tooling on per-evaluation cost — we sell what that tooling doesn’t measure: whether human reviewers agree with each other and with your judge, reported per criterion, with the noisy dimensions shown first instead of averaged away.

The deliverable is a ranked list of failure modes with frequency data and example evidence per class — the artifact a product team can actually act on: fix the top defect class, then re-test. Where a head-to-head comparison answers your question better than absolute scoring, we run it that way and say why.

03Expected inputs

What you provide

  • Model-output items: prompts with complete responses, agreed schema (JSONL / CSV), stable unit IDs — and the generating model identified per item where relevant.
  • Evaluation criteria: yours, or a rubric co-drafted with you — written, versioned, and signed before review begins.
  • Defect taxonomy: proposed by us, approved by you — adapted to your output type.
  • Sample definition: full review or a pre-agreed sampling plan. Plus a data-sensitivity declaration — no regulated or sensitive content.

04Workflow

How a pilot works

  1. Intake & sensitivity screen

    We verify schema, ID completeness, and ID uniqueness, and run the sensitivity screen. Malformed or out-of-scope inputs stop here with a written intake report.

  2. Criteria lock

    Evaluation criteria or rubric signed and versioned. Mid-project changes pause production and need a change order.

  3. Reviewer qualification & calibration

    Reviewers qualify against sealed golds with written rationales; a calibration batch with adjudication locks the criteria interpretation.

  4. Blind review

    Reviewers assess each output against the criteria without seeing any client-supplied “expected” judgments. Every assessment records unit ID, per-criterion judgments, defect classifications, rationale, reviewer, and criteria version.

  5. Sampling review & adjudication

    An independent reviewer re-checks a stratified sample; the overturn rate is reported. All disagreements go to an adjudicator — never the producer of the disputed units — and rulings are logged.

  6. Delivery & acceptance

    QA report, defect evidence pack, adjudication log, and the criteria version used — then your acceptance review.

05Deliverables

What you receive

  • QA report: per-criterion results, defect taxonomy with frequencies, estimated defect rate, systematic patterns, and remediation recommendations.
  • Defect evidence pack: anonymized examples per defect class, within NDA limits.
  • Complete adjudication log: every disagreement, every ruling.
  • The criteria / rubric version used, with its changelog.

06Quality controls

How we keep it honest

  • Blind review — reviewers never see client-supplied expected judgments while assessing.
  • Sealed golds for qualification, with reviewer gold performance reported.
  • Inter-reviewer agreement on a double-reviewed subset, reported per criterion — rubrics fail dimension by dimension, so we never report a blended average alone.
  • Independent sampling review with a reported overturn rate; flag-rather-than-guess discipline throughout.
  • Criteria discipline: recurring judgment ambiguity triggers a versioned criteria patch, never silent drift.

07Scope and limits

Stated up front

What this product does not include — stated here, so there are no surprises in scoping:

  • Audits outputs; does not fix, rewrite, or regenerate them. Remediation is advice in the report, not rework.
  • No safety or red-team evaluation — outputs needing adversarial safety testing are declined with a documented reason.
  • No regulated or sensitive data or content; no specialist medical, legal, or financial output evaluation without separately credentialed reviewers.
  • No claim about the underlying model’s general capability beyond what the reviewed outputs evidence.
  • Findings do not generalize beyond the reviewed outputs and the stated sampling methodology.
  • English and Arabic only. No guaranteed defect figures and no model certification; acceptance criteria are agreed per pilot, in writing, before work begins.

→Next step

Start with a scoped pilot.

Tell us about your project. We review every inquiry against our scope — if it’s a fit, we schedule a scoping conversation.