Primwright Enterprise Product

Guideline development

Pilot-scoped service · Taxonomies & annotation guidelines · English / Arabic

Your annotators disagree and your eval numbers wobble because nobody wrote down what “good” looks like. We design the taxonomy and annotation guideline — definitions, near-miss counterexamples, edge-case rulings — and validate that independent raters can actually apply it.

01The problem

“We don’t trust our own eval numbers.”

Labels are inconsistent, eval scores wobble between runs, and nobody can say whether a “bad” score means the model is bad or the instructions were vague. Teams make product decisions — ship, retrain, switch model — on top of measurement noise they cannot see.

The usual fixes don’t fix it: a hurried one-pager drifts within a week; a vendor’s bundled task design commits you to their production volume; a consultancy deck ends in advice, not an instrument anyone can judge against.

02What it is

The instrument before the measurement.

We design the measurement instrument before you spend on measurement: the taxonomy or anchored rubric, written definitions, inclusion and exclusion rules, positive examples — and, crucially, near-miss counterexamples and a standing edge-case table. The counterexamples are the load-bearing part; they are what hurried internal guidelines always skip.

Then we try to break the draft ourselves before you ever see it, and — if commissioned — run a small blind pilot that measures whether independent raters can actually apply the guideline. If they can’t, the guideline is at fault, not the raters, and the report says so.

03Expected inputs

What you provide

  • Your task description and the downstream decision the labels or scores will drive — a taxonomy that doesn’t serve the decision is decoration.
  • A representative sample of real items (or a synthetic stand-in, flagged as provisional in writing).
  • Existing categories or criteria, and the disagreement points you already know about.
  • A named contact with sign-off authority, and a data-sensitivity declaration — no regulated or sensitive data.

04Workflow

How a pilot works

  1. Discovery

    Structured interviews on the decision the guideline must serve, who judges today, and where they disagree. Notes are filed, not kept in anyone’s head.

  2. Draft taxonomy / rubric

    Category definitions or anchored dimensions, inclusion and exclusion rules, open questions marked explicitly — not smoothed over.

  3. Counterexamples & edge cases

    For every category and scale anchor: real positives from your sample and boundary near-misses. A standing edge-case table records ambiguity → ruling → version.

  4. Adversarial self-review

    We attempt to break the draft — finding items the definitions don’t decide — and write the missing rulings before client review.

  5. Labeled sample

    A small sample labeled against the draft, delivered with the guideline so you can see the taxonomy in action.

  6. Client review & sign-off

    You challenge the draft; disagreements are resolved in the text. Version 1.0 is signed — the single source of truth.

  7. Pilot validation (optional, recommended)

    A small blind pilot — qualification plus calibration discipline — measures whether independent raters can apply the guideline. Findings feed a v1.1 patch; a below-threshold pilot triggers revision, not silent acceptance.

  8. Delivery

    Versioned guideline document with changelog, labeled sample, sign-off record, and — if commissioned — the pilot validation report with honestly reported agreement figures.

05Deliverables

What you receive

  • Annotation guideline document: task definition, output schema, taxonomy or anchored rubric, positive examples, near-miss counterexamples, edge-case table, flag definitions, and changelog.
  • Labeled sample demonstrating the guideline in action.
  • Client sign-off record for the signed v1.0.
  • If pilot validation is commissioned: pilot validation report with measured agreement figures — reported honestly, including failure.

06Quality controls

How we keep it honest

  • Adversarial self-review before client review — we break the draft first.
  • Counterexample coverage check: every category boundary illustrated by a near miss, not just clean positives.
  • Independent-applicability test: the pilot validation measures whether two independent raters can reach the stated agreement — if not, the guideline is revised.
  • Version discipline: every post-sign-off change is a versioned patch with a changelog entry; raters work from the signed version only.
  • Acceptance bar: every category ships with a definition, at least two positives, and at least one counterexample; the edge-case table covers your listed ambiguities.

07Scope and limits

Stated up front

What this product does not include — stated here, so there are no surprises in scoping:

  • Design only — no production annotation volume. Labeling beyond a small demonstration sample is a separately scoped authorization.
  • No guidelines for safety or red-team work, regulated data, or specialist medical, legal, or financial evaluation without separately credentialed reviewers.
  • English and Arabic only — with worked examples in the item language. Dialect scope is stated explicitly; items outside documented rater competence are excluded, never forced through.
  • No legal, medical, financial, or compliance advice embedded or implied.
  • No guarantee of any downstream agreement or accuracy figure. A guideline built without a representative sample is flagged as provisional in writing — never silently shipped.

→Next step

Start with a scoped pilot.

Tell us about your project. We review every inquiry against our scope — if it’s a fit, we schedule a scoping conversation.