Synthetic demonstration

Interactive synthetic demonstration · 60 audited units · filterable by batch, defect type, reviewer

This demonstration uses entirely synthetic data and simulated annotators. It shows how our evaluation process works — it does not show results on real data, real annotators, or real client work. No human annotators participated. Agreement figures shown are internal process measurements, not product accuracy, not guarantees, and not transferable to any real engagement. Primwright Enterprise does not currently offer these services commercially.

Primwright Enterprise Demonstration 7 of 7

The audit that finds the bad news.

A QA audit of a synthetic labeled dataset — filter it by batch, defect type, and reviewer. Every number carries its n and its method, and the failures stay on the record: one batch failed the overturn gate, and one reviewer was wrong two times out of three.

01Filter the audit

Inspect the findings.

Dataset

Northwind Desk labeled bank (synthetic)

Audited sample

60 of 480 units

Defects found

24 [ACTUAL — internal, simulated; synthetic data]

Sampling

Stratified · blind re-annotation

Batch

Defect type

Reviewer

Showing: All batches · n=60 audited of 480 synthetic units. Synthetic data — simulated annotators.

Defects by type

Wrong class6Boundary misread5Guideline drift3Missed defect (eval)4Overturned label3Provenance gap1Language-slice error2
Defect counts by type (the data behind the chart)
Defect typeDefinitionCount
Wrong classLabel contradicts the taxonomy's precedence rules.6
Boundary misreadCorrect neighborhood, wrong side of a documented edge case.5
Guideline driftLabel follows an older guideline version than the batch's governing version.3
Missed defect (eval)For evaluation units: a planted factual defect the judge did not cite.4
Overturned labelAdjudication reversed the production label.3
Provenance gapUnit missing its guideline-version or judge record.1
Language-slice errorLabel correct for one language slice but wrong for the other.2

Error-rate estimate by batch

Defect rate per batch with Wilson 95% interval. The method is stated so the estimate is checkable — the audit is itself auditable.
BatchAuditedDefectsRateWilson 95% CI
Batch A20525%11.2% – 46.9%
Batch B20735%18.1% – 56.7%
Batch C201260%38.7% – 78.1%

The bank was seeded with defects to demonstrate the audit — this error rate describes the synthetic bank, not any real dataset. Wide intervals are honest: n=20 per batch cannot support a precise estimate, and the dashboard says so.

Overturn gate — the bad news, on the record

Overturn gate: provisional bar 10% [PROPOSED — unvalidated standard]. Overturn rate measures reviewer–majority disagreement, not reviewer correctness.
BatchOverturn rate (n=20)Result
Batch A5.0%Pass
Batch B10.0%Pass
Batch C15.0%FAIL

Batch C failed the gate

Batch C exceeded the provisional 10% overturn bar. Per procedure the batch was re-sampled at double depth before delivery — the re-sample is shown, and the original failure stays on the record. Failing gates are reported, not hidden.

Reviewers — including the fallible one

Q1 · QA reviewer [simulated]

2 overturns raised · 2 upheld

Raised 2 overturns; adjudication upheld both.

Q2 · QA reviewer [simulated]

3 overturns raised · 1 upheld

Disagreed with the majority 3 times; right only once when checked against gold. Reviewers are fallible — the audit says so.

Q3 · QA reviewer [simulated]

1 overturn raised · 1 upheld

Raised 1 overturn; upheld.

Example findings (6 shown)

  • F-01minorBatch A · Boundary misread

    Billing-triggered lockout labeled billing_inquiry; precedence rules place it in account_access.

    Disposition: Overturned to account_access; example added to the edge-case appendix.

  • F-02minorBatch B · Guideline drift

    Two units labeled under guideline G1.0 after G1.1 took effect for the batch.

    Disposition: Re-labeled under G1.1; provenance records corrected.

  • F-03majorBatch B · Missed defect (eval)

    Judge did not cite the self-contradictory retention claim in a response evaluation unit.

    Disposition: Unit re-scored; judge re-briefed on contradiction defects.

  • F-04majorBatch C · Wrong class

    Security-relevant phrasing labeled technical_issue; precedence rules require security_concern.

    Disposition: Overturned; batch C flagged for re-sampling (see overturn gate).

  • F-05minorBatch C · Provenance gap

    One unit missing its judge-identity record — provenance gap.

    Disposition: Record reconstructed from the batch log; gap logged as a process defect.

  • F-06minorBatch C · Language-slice error

    Arabic item labeled with the English-slice reading of a boundary example.

    Disposition: Overturned; in-language qualification rule restated in the batch briefing.

The sampling method, in full

Stratified random sample: 20 units per batch (60 of 480 synthetic units, 12.5%), stratified by class and language slice. Blind re-annotation by two independent simulated judges; disagreements adjudicated against sealed gold. Error-rate estimate uses the Wilson 95% interval on the sampled defect count. The audit is itself auditable: sample IDs, strata, and the gold seals are listed in the manifest.

Honest notes

  • Overturn rate measures reviewer-majority disagreement, not reviewer correctness — a distinction made on every report.
  • The audit measures quality; it does not certify it.
  • Findings generalize only to the sampled units under the stated method.
  • Batch C failed the provisional overturn gate and was re-sampled; the failure remains on the record.

◈Method, stated honestly

What this demonstrates — and what it doesn’t.

It demonstrates

  • An audit that is itself auditable: stratified sampling, blind re-annotation, adjudication against sealed gold, and an error-rate estimate with the method attached.
  • The habit of reporting bad results: a failed gate stays on the record, and a reviewer is shown fallible against gold.
  • That overturn rate measures reviewer–majority disagreement — not reviewer correctness — a distinction made on every report.

It does not claim

  • That this error rate describes any real dataset — the synthetic bank was seeded with defects to demonstrate the audit.
  • That the audit certifies quality — it measures quality. Findings generalize only to the sampled units under the stated method.
  • Any staffing, scale, or turnaround figure.

→Next step

We sell the habit of reporting bad results.

A QA audit measures your data; it does not certify it. Start with a scoping conversation — we will tell you what we can and cannot audit.