Primwright Enterprise Product
Model-output evaluation
Pilot-scoped service · Human review · English / Arabic
You are about to put a chatbot or AI feature in front of customers. We independently review what it produces — scored against criteria you approve, defects classified with measured frequencies — so ship, delay, and cut decisions rest on evidence instead of hope.
01The problem
What breaks, how often, and where?
The team is about to ship an LLM-powered feature and cannot answer the only question that matters. Automated evals grade the model’s homework with another model; engineers “eyeball 200 responses” with no methodology; public benchmarks don’t cover your prompts, your policy, or your tone. You are deciding ship, delay, or cut on measurement you privately don’t trust.
When the question is comparative — “is response A or B better?” — head-to-head judging with hidden model identities and mandatory rationales is available inside this service as a method, never as a separate product.
02What it is
The human-agreement layer.
We are the layer your automated eval stack is missing. We don’t compete with eval tooling on per-evaluation cost — we sell what that tooling doesn’t measure: whether human reviewers agree with each other and with your judge, reported per criterion, with the noisy dimensions shown first instead of averaged away.
The deliverable is a ranked list of failure modes with frequency data and example evidence per class — the artifact a product team can actually act on: fix the top defect class, then re-test. Where a head-to-head comparison answers your question better than absolute scoring, we run it that way and say why.
03Expected inputs
What you provide
- Model-output items: prompts with complete responses, agreed schema (JSONL / CSV), stable unit IDs — and the generating model identified per item where relevant.
- Evaluation criteria: yours, or a rubric co-drafted with you — written, versioned, and signed before review begins.
- Defect taxonomy: proposed by us, approved by you — adapted to your output type.
- Sample definition: full review or a pre-agreed sampling plan. Plus a data-sensitivity declaration — no regulated or sensitive content.
04Workflow
How a pilot works
Intake & sensitivity screen
We verify schema, ID completeness, and ID uniqueness, and run the sensitivity screen. Malformed or out-of-scope inputs stop here with a written intake report.
Criteria lock
Evaluation criteria or rubric signed and versioned. Mid-project changes pause production and need a change order.
Reviewer qualification & calibration
Reviewers qualify against sealed golds with written rationales; a calibration batch with adjudication locks the criteria interpretation.
Blind review
Reviewers assess each output against the criteria without seeing any client-supplied “expected” judgments. Every assessment records unit ID, per-criterion judgments, defect classifications, rationale, reviewer, and criteria version.
Sampling review & adjudication
An independent reviewer re-checks a stratified sample; the overturn rate is reported. All disagreements go to an adjudicator — never the producer of the disputed units — and rulings are logged.
Delivery & acceptance
QA report, defect evidence pack, adjudication log, and the criteria version used — then your acceptance review.
05Deliverables
What you receive
- QA report: per-criterion results, defect taxonomy with frequencies, estimated defect rate, systematic patterns, and remediation recommendations.
- Defect evidence pack: anonymized examples per defect class, within NDA limits.
- Complete adjudication log: every disagreement, every ruling.
- The criteria / rubric version used, with its changelog.
06Quality controls
How we keep it honest
- Blind review — reviewers never see client-supplied expected judgments while assessing.
- Sealed golds for qualification, with reviewer gold performance reported.
- Inter-reviewer agreement on a double-reviewed subset, reported per criterion — rubrics fail dimension by dimension, so we never report a blended average alone.
- Independent sampling review with a reported overturn rate; flag-rather-than-guess discipline throughout.
- Criteria discipline: recurring judgment ambiguity triggers a versioned criteria patch, never silent drift.
07Scope and limits
Stated up front
What this product does not include — stated here, so there are no surprises in scoping:
- Audits outputs; does not fix, rewrite, or regenerate them. Remediation is advice in the report, not rework.
- No safety or red-team evaluation — outputs needing adversarial safety testing are declined with a documented reason.
- No regulated or sensitive data or content; no specialist medical, legal, or financial output evaluation without separately credentialed reviewers.
- No claim about the underlying model’s general capability beyond what the reviewed outputs evidence.
- Findings do not generalize beyond the reviewed outputs and the stated sampling methodology.
- English and Arabic only. No guaranteed defect figures and no model certification; acceptance criteria are agreed per pilot, in writing, before work begins.
Related services
→Next step
Start with a scoped pilot.
Tell us about your project. We review every inquiry against our scope — if it’s a fit, we schedule a scoping conversation.