10 to the 23 AI logo

Practice Library

Governance patternprocedural

Verification training

Train staff to actually check AI output — verification as a taught skill, deepest where checking is hardest: conceptual errors.

What it changes

increasedPeople or agents using it(verification skill — trained catch capacity, deepest on conceptual errors)PAN Lab model result: direction from the documented catch pattern (N=2,784): people caught surface errors at ~82% but conceptual errors at only ~31%, and catch rates were governed by verification effort, trust calibration and error legibility — the trainable levers — not by incentives; a collaborative-AI metacognition measure predicted benefit beyond general metacognition, making verification skill the documented counter to over-reliance.
cappedOperator deference drift(you defer less to what you can check yourself)PAN Lab model result: direction only: trained verification skill is the documented counter to over-reliance; magnitude is an authored step, not a measured delta.

Who can pull it

Deploying organizationUsers & agents

What it looks like institutionally

Checking AI output is a skill, and the evidence says it is unevenly distributed in a very particular way. In a randomized study with objective ground truth (N=2,784), people caught surface errors in AI suggestions about 82% of the time — and conceptual errors, the ones requiring domain judgment, only about 31% of the time. Catch rates were governed by verification effort, prior trust calibration, and error legibility; financial incentives and extra time did not move them. What moves them is training.

Verification training teaches the checking itself: check against source rather than for plausibility, look for the conceptual failure the fluent draft hides, and know which error classes the tool actually produces. A collaborative-AI metacognition measure predicted who benefits from AI beyond general metacognition — the skill is specific, learnable, and distinct from general competence. The strong form puts recurring drills with ground-truth feedback on protected time, because a skill nobody exercises decays back into trust.

This pattern BUILDS the checking skill; its sibling, deskilling arrest, PRESERVES the underlying domain skill the checking depends on. They are deliberately mirror-weighted in the Lab — train verification where staff must judge machine output; keep no-AI practice where the machine would otherwise become the only practitioner. And what you can check yourself, you defer to less: trained verification is the documented counter to over-reliance.

Ledgered PAN-run results used above

In a randomized study (N=2,784) with objective ground truth, humans accepted incorrect AI suggestions about a third of the time, and their rate of catching AI errors was governed by verification effort, prior trust in AI, and error legibility — surface errors were caught ~82% of the time versus ~31% for errors requiring conceptual judgment — not by financial incentives or time spent.[]

A validated collaborative-AI metacognition scale (planning, monitoring, evaluation of one's own reliance) predicted collaboration benefits incrementally beyond general metacognition — verification-skill training, not generic AI knowledge, is the calibrated counter to over-reliance.[2]

A validated collaborative-AI metacognition scale (planning, monitoring, evaluation of one's own reliance) predicted collaboration benefits incrementally beyond general metacognition — verification-skill training, not generic AI knowledge, is the calibrated counter to over-reliance.[2]

Addresses: Conceptual errors accepted as checked · Verification assumed rather than taught · Over-reliance on fluent output. Test a version of this lever in the PAN Lab.

Deciding whether this lever fits your deployment?

Which patterns matter, and in what order, depends on your system's actual shape. Ranking your options on evidence, with what can backfire stated, is engagement work.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalA validated collaborative-AI metacognition scale (planning, monitoring, evaluation of one's own reliance) pred…

A validated collaborative-AI metacognition scale (planning, monitoring, evaluation of one's own reliance) predicted collaboration benefits incrementally beyond general metacognition — verification-skill training, not generic AI knowledge, is the calibrated counter to over-reliance.

sidra2025AcademicSave

Sidra, & Mason, C. (2025). Generative AI in Human-AI Collaboration: Validation of the Collaborative AI Literacy and Collaborative AI Metacognition Scales for Effective Use. International Journal of Human-Computer Interaction. https://doi.org/10.1080/10447318.2025.2543997

doi.org/10.1080/10447318.2025.2543997

Appears in: Evidence reverification (2026)

Topics: ai-safety, human-ai-interaction

bucinca2021AcademicSave

Bucinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188. https://doi.org/10.1145/3449287

doi.org/10.1145/3449287

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction

EmpiricalIn a randomized study (N=2,784) with objective ground truth, humans accepted incorrect AI suggestions about a …

In a randomized study (N=2,784) with objective ground truth, humans accepted incorrect AI suggestions about a third of the time, and their rate of catching AI errors was governed by verification effort, prior trust in AI, and error legibility — surface errors were caught ~82% of the time versus ~31% for errors requiring conceptual judgment — not by financial incentives or time spent.

beck2026AcademicSave

Beck, J., Eckman, S., Kern, C., & Kreuter, F. (2026). Bias in the Loop: How Humans Evaluate AI-Generated Suggestions. Harvard Data Science Review, 8(2). https://hdsr.mitpress.mit.edu/pub/nrcn4h7d/release/1

https://hdsr.mitpress.mit.edu/pub/nrcn4h7d/release/1

Appears in: PAN framework development

Topics: algorithmic-fairness, human-ai-interaction