Primwright Enterprise Product

Dataset QA

Pilot-scoped service · Labeled-data audits · English / Arabic

You spent serious money on labels and the model still isn’t improving. We audit your labeled dataset — stratified sample, blind re-annotation, defect taxonomy, error-rate estimate with the method attached — so you know whether the data is the problem before you spend again.

01The problem

Is the data bad?

“We spent six figures on labels and the model isn’t improving — is the data bad?” When labeled data underperforms, teams face three expensive forks: re-label everything (and repeat the original sin — the guideline was never written down), trust the vendor’s QA (the vendor grades its own homework), or fly blind while the doubt compounds.

A dataset audit answers the cheapest decisive question first: what is the actual error rate, where is it concentrated, and what would fix it — with the method attached so the answer is auditable.

02What it is

The audit, not the labels.

We sell the audit, not the labels. We have no annotation volume business, so unlike managed labeling providers we have no incentive to find your dataset “fine.” The independence is structural, not claimed.

The audit is itself auditable: a stratified sampling plan published in the deliverable, blind re-annotation of the sample, a defect taxonomy with measured frequencies, and an error-rate estimate with the sampling methodology attached. We report bad news as the product — a finding of “badly broken” is as acceptable as “fine,” and what is never acceptable is an unauditable estimate.

03Expected inputs

What you provide

  • Labeled dataset with stable unit IDs and an agreed schema (CSV / JSONL).
  • Your original annotation guideline, if one exists — used to distinguish genuine defects from guideline drift.
  • Your written question and scope (for example: “estimate the label error rate,” “identify systematic defect patterns”).
  • A data-sensitivity declaration. We take no regulated or sensitive data — that screen is binary.

04Workflow

How a pilot works

  1. Intake & sensitivity screen

    We verify dataset completeness, ID uniqueness, and schema, and run the sensitivity screen. Malformed or out-of-scope inputs stop here, with a written intake report.

  2. Scope lock

    You sign the sampling methodology, the strata, the defect taxonomy, and the definition of done — before any review begins.

  3. Blind re-annotation

    Qualified reviewers re-annotate the agreed sample without seeing your original labels. Sealed gold items calibrate every reviewer.

  4. Defect categorization

    Every reviewer-vs-original disagreement is classified under the agreed taxonomy: guideline ambiguity, annotator error, systematic bias, edge-case miss, data defect.

  5. Error-rate estimation

    Per stratum and overall, with the sampling method attached — the estimate can be checked by a third party reading the report alone.

  6. Delivery & acceptance

    Audit report, defect evidence pack, adjudication log, and methodology appendix — then your acceptance review.

05Deliverables

What you receive

  • Audit report: estimated error rate (overall and per stratum), defect taxonomy with frequencies, systematic patterns, and prioritized remediation recommendations.
  • Defect evidence pack: anonymized examples per defect class, within NDA limits.
  • Complete adjudication log: every disagreement, every ruling.
  • Methodology appendix: sampling plan, strata, sample sizes, reviewer qualification results, agreement figures.

06Quality controls

How we keep it honest

  • Blind re-annotation — reviewers never see the original labels during the sample pass.
  • Sealed gold items: reviewer qualification against keys with written rationales; gold performance reported per reviewer.
  • Inter-reviewer agreement measured on a double-annotated subset and reported — including poor figures.
  • Every disagreement adjudicated by someone other than the producer of the disputed units; rulings logged.
  • The sampling plan ships inside the deliverable. If a systematic defect pattern emerges mid-audit, we flag it immediately; expanding the sample needs your agreement.

07Scope and limits

Stated up front

What this product does not include — stated here, so there are no surprises in scoping:

  • Audits labels; does not relabel. Corrections are a separately scoped follow-on, never folded silently into the audit.
  • No pre-committed findings — we do not promise your error rate is below any threshold. The estimate is produced by the work.
  • Text only. No image, audio, or video annotation; no safety or red-team evaluation.
  • No regulated or sensitive data — no health, financial, legal-privileged, or otherwise regulated datasets.
  • Findings do not generalize beyond the sampled units and the stated method.
  • English and Arabic only. No guaranteed accuracy figures; acceptance criteria are agreed per pilot, in writing, before work begins.

→Next step

Start with a scoped pilot.

Tell us about your project. We review every inquiry against our scope — if it’s a fit, we schedule a scoping conversation.