10 to the 23 AI logo

Practice Library

Governance patternstructural

Content-aware decontamination

Clean the record system by reading what you remove — deleting unread does not reduce contamination, it concentrates it.

What it changes

dampenedRecord contamination pressure(content-aware audit-and-correct only; bounded)

Who can pull it

Deploying organizationHarness builderData-protection officer

What it looks like institutionally

The instinct to "clear out old records" treats deletion as cleanup. In published PAN runs it is the opposite: removing records without reading them raises the contaminated SHARE of the record system, because the benign entries being deleted were diluting the bad ones. A content-blind purge strips the dilution and leaves a denser pool of error behind — a backfire, not a scrub.

Only content-aware decontamination — audit each record, correct or retire the ones that are actually wrong — reliably reduces contamination. And even that is bounded: an optimistic anchor near 96% token accuracy falls away on hard content, so decontamination is a partial control that must be planned as a range, not a guarantee.

The practice, then, is disciplined: never delete by age or volume alone; read before you remove; label what you cannot yet verify (see Provenance labeling) rather than purging it blind.

Addresses: Content-blind record removal · Contaminated-share inflation. Test a version of this lever in the PAN Lab.

Deciding whether this lever fits your deployment?

Which patterns matter, and in what order, depends on your system's actual shape. Ranking your options on evidence, with what can backfire stated, is engagement work.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

ScenarioIn the published runs, deleting records without reading them raised the contaminated share by stripping out be…

In the published runs, deleting records without reading them raised the contaminated share by stripping out benign entries; only content-aware cleanup reliably reduced it.

From the published runs: PAN governance-lever audit.

EmpiricalRecord audit-and-correct shares the detection-ceiling family: an optimistic anchor near 96% token accuracy fal…

Record audit-and-correct shares the detection-ceiling family: an optimistic anchor near 96% token accuracy falls away on hard content, so decontamination is bounded rather than total.

halludetectlegaldomainGroundingPreprintSave

HalluDetect legal-domain (arXiv:2509.11619) — best mitigation architecture reaches ~96% token accuracy in a FAVORABLE, retrieval-grounded legal setting (optimistic end). https://arxiv.org/abs/2509.11619

https://arxiv.org/abs/2509.11619

Grounds: empirical cap: decontaminate (max)

samedetectionaccuracyliteratGroundingPeer-reviewedSave

Same detection-accuracy literature as catch_at_generation (FaithBench arXiv:2410.13210; arXiv:2508.08285); audit-time detection is bounded by the same hallucination-detection ceiling.

Grounds: empirical cap: catch_at_generation (max); empirical cap: decontaminate (max)

theillusionofprogressGroundingPreprintSave

'The Illusion of Progress' (arXiv:2508.08285) — LLM-as-Judge Precision 0.736 / Recall 0.957 / F1 0.832 vs human consensus on QA. https://arxiv.org/abs/2508.08285

https://arxiv.org/abs/2508.08285

Grounds: empirical cap: catch_at_generation (max)