10 to the 23 AI logo

Practice Library

Governance patternstructural

Improve the model

The default instinct — buy or build a better model — is a real lever with an honest, limited reach.

What it changes

decreasedAI modelPAN Lab model result: ≈6% of the harm removed by a model upgrade alone in the same published runs that credited a verifier with ≈46%.

Who can pull it

DeveloperVendorDeploying organization

What it looks like institutionally

Reducing the model's base error rate helps every downstream pathway a little. It is also the lever institutions reach for first, because it requires no organizational change: procurement instead of governance.

Its honest limits: base error has an empirical floor (no current system reaches zero in demanding domains), and in propagation terms a better model shrinks the source while leaving every loop — adoption, records, retrieval, peer spread — untouched. PAN Lab runs across hundreds of deployment structures found system-side levers outperforming equal-effort model improvements in the overwhelming majority of cases; the ledgered scenario results carry the specifics and their caveats.

Use it, but use it last-alone: pair model improvements with the structural levers that govern what happens to the errors that remain.

Ledgered PAN-run results used above

In the published runs, over a supervised-plus-agent scenario, adding a verifier to the autonomous agent removed roughly 46% of the harm that persists and a coordinated governance package roughly 43%, while upgrading the model alone removed only about 6%.[]

In the published runs, fixing the surrounding system out-leveraged an equal-effort model upgrade in nearly every case tested, and by several times the margin - a better model helps least where the system, not the model, does the damage.[]

Addresses: High base error. Test a version of this lever in the PAN Lab.

Deciding whether this lever fits your deployment?

Which patterns matter, and in what order, depends on your system's actual shape. Ranking your options on evidence, with what can backfire stated, is engagement work.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

ScenarioIn the published runs, over a supervised-plus-agent scenario, adding a verifier to the autonomous agent remove…

In the published runs, over a supervised-plus-agent scenario, adding a verifier to the autonomous agent removed roughly 46% of the harm that persists and a coordinated governance package roughly 43%, while upgrading the model alone removed only about 6%.

From the published runs: PAN social-work governance guidance, lever-ranking comparison.

ScenarioIn the published runs, fixing the surrounding system out-leveraged an equal-effort model upgrade in nearly eve…

In the published runs, fixing the surrounding system out-leveraged an equal-effort model upgrade in nearly every case tested, and by several times the margin - a better model helps least where the system, not the model, does the damage.

From the published runs: PAN baseline analysis.

EmpiricalModel error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors ru…

Model error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors run roughly 1.6–11.6% in frontier evaluations and 4–86% across domains.

xuetal2024GroundingPreprintSave

Xu et al. (2024), 'Hallucination is Inevitable: An Innate Limitation of LLMs' — formal proof that hallucination cannot be eliminated.

Grounds: empirical cap: model_error_base (min)

karpowicz2025GroundingPreprintSave

Karpowicz (2025) — three independent mathematical frameworks (auction theory, proper scoring, log-sum-exp) all conclude no LLM inference mechanism can be simultaneously truthful, etc.

Grounds: empirical cap: model_error_base (min)

halogenGroundingPeer-reviewedSave

HALoGEN (arXiv:2501.08292) — best models hallucinate 4%-86% of generated facts depending on domain. https://arxiv.org/abs/2501.08292

https://arxiv.org/abs/2501.08292

Grounds: empirical cap: model_error_base (min)

openai2025GroundingFrontier labSave

OpenAI (2025), 'Why Language Models Hallucinate' — next-token training plus IDK-penalizing benchmarks push models to bluff; explains the persistent nonzero floor.

Grounds: empirical cap: model_error_base (min)

llmstats2026GroundingIndustry evaluationSave

llm-stats.com failure-focused eval (2026) — FactsGrounding 89.1% accuracy => ~10.9% failure on a relatively easy grounded benchmark.

Grounds: empirical cap: model_error_base (min)

suprmindbenchmarkdigest2026GroundingIndustry evaluationSave

Suprmind benchmark digest (2026) — production ChatGPT ~4.8% major-incorrect with reasoning vs ~11.6% without; HealthBench 3.6%->1.6% with GPT-5 thinking.

Grounds: empirical cap: model_error_base (min)