Q1 · QA reviewer [simulated]
2 overturns raised · 2 upheld
Raised 2 overturns; adjudication upheld both.
Primwright Enterprise Demonstration 7 of 7
A QA audit of a synthetic labeled dataset — filter it by batch, defect type, and reviewer. Every number carries its n and its method, and the failures stay on the record: one batch failed the overturn gate, and one reviewer was wrong two times out of three.
01Filter the audit
Dataset
Northwind Desk labeled bank (synthetic)
Audited sample
60 of 480 units
Defects found
24 [ACTUAL — internal, simulated; synthetic data]
Sampling
Stratified · blind re-annotation
Batch
Defect type
Reviewer
Showing: All batches · n=60 audited of 480 synthetic units. Synthetic data — simulated annotators.
| Defect type | Definition | Count |
|---|---|---|
| Wrong class | Label contradicts the taxonomy's precedence rules. | 6 |
| Boundary misread | Correct neighborhood, wrong side of a documented edge case. | 5 |
| Guideline drift | Label follows an older guideline version than the batch's governing version. | 3 |
| Missed defect (eval) | For evaluation units: a planted factual defect the judge did not cite. | 4 |
| Overturned label | Adjudication reversed the production label. | 3 |
| Provenance gap | Unit missing its guideline-version or judge record. | 1 |
| Language-slice error | Label correct for one language slice but wrong for the other. | 2 |
| Batch | Audited | Defects | Rate | Wilson 95% CI |
|---|---|---|---|---|
| Batch A | 20 | 5 | 25% | 11.2% – 46.9% |
| Batch B | 20 | 7 | 35% | 18.1% – 56.7% |
| Batch C | 20 | 12 | 60% | 38.7% – 78.1% |
The bank was seeded with defects to demonstrate the audit — this error rate describes the synthetic bank, not any real dataset. Wide intervals are honest: n=20 per batch cannot support a precise estimate, and the dashboard says so.
| Batch | Overturn rate (n=20) | Result |
|---|---|---|
| Batch A | 5.0% | Pass |
| Batch B | 10.0% | Pass |
| Batch C | 15.0% | FAIL |
Batch C failed the gate
Batch C exceeded the provisional 10% overturn bar. Per procedure the batch was re-sampled at double depth before delivery — the re-sample is shown, and the original failure stays on the record. Failing gates are reported, not hidden.
Q1 · QA reviewer [simulated]
2 overturns raised · 2 upheld
Raised 2 overturns; adjudication upheld both.
Q2 · QA reviewer [simulated]
3 overturns raised · 1 upheld
Disagreed with the majority 3 times; right only once when checked against gold. Reviewers are fallible — the audit says so.
Q3 · QA reviewer [simulated]
1 overturn raised · 1 upheld
Raised 1 overturn; upheld.
F-01minor
Billing-triggered lockout labeled billing_inquiry; precedence rules place it in account_access.
Disposition: Overturned to account_access; example added to the edge-case appendix.
F-02minor
Two units labeled under guideline G1.0 after G1.1 took effect for the batch.
Disposition: Re-labeled under G1.1; provenance records corrected.
F-03major
Judge did not cite the self-contradictory retention claim in a response evaluation unit.
Disposition: Unit re-scored; judge re-briefed on contradiction defects.
F-04major
Security-relevant phrasing labeled technical_issue; precedence rules require security_concern.
Disposition: Overturned; batch C flagged for re-sampling (see overturn gate).
F-05minor
One unit missing its judge-identity record — provenance gap.
Disposition: Record reconstructed from the batch log; gap logged as a process defect.
F-06minor
Arabic item labeled with the English-slice reading of a boundary example.
Disposition: Overturned; in-language qualification rule restated in the batch briefing.
Stratified random sample: 20 units per batch (60 of 480 synthetic units, 12.5%), stratified by class and language slice. Blind re-annotation by two independent simulated judges; disagreements adjudicated against sealed gold. Error-rate estimate uses the Wilson 95% interval on the sampled defect count. The audit is itself auditable: sample IDs, strata, and the gold seals are listed in the manifest.
◈Method, stated honestly
→Next step
A QA audit measures your data; it does not certify it. Start with a scoping conversation — we will tell you what we can and cannot audit.