What it changes
Who can pull it
What it looks like institutionally
Checking AI output is a skill, and the evidence says it is unevenly distributed in a very particular way. In a randomized study with objective ground truth (N=2,784), people caught surface errors in AI suggestions about 82% of the time — and conceptual errors, the ones requiring domain judgment, only about 31% of the time. Catch rates were governed by verification effort, prior trust calibration, and error legibility; financial incentives and extra time did not move them. What moves them is training.
Verification training teaches the checking itself: check against source rather than for plausibility, look for the conceptual failure the fluent draft hides, and know which error classes the tool actually produces. A collaborative-AI metacognition measure predicted who benefits from AI beyond general metacognition — the skill is specific, learnable, and distinct from general competence. The strong form puts recurring drills with ground-truth feedback on protected time, because a skill nobody exercises decays back into trust.
This pattern BUILDS the checking skill; its sibling, deskilling arrest, PRESERVES the underlying domain skill the checking depends on. They are deliberately mirror-weighted in the Lab — train verification where staff must judge machine output; keep no-AI practice where the machine would otherwise become the only practitioner. And what you can check yourself, you defer to less: trained verification is the documented counter to over-reliance.
Ledgered PAN-run results used above
In a randomized study (N=2,784) with objective ground truth, humans accepted incorrect AI suggestions about a third of the time, and their rate of catching AI errors was governed by verification effort, prior trust in AI, and error legibility — surface errors were caught ~82% of the time versus ~31% for errors requiring conceptual judgment — not by financial incentives or time spent.[†]
A validated collaborative-AI metacognition scale (planning, monitoring, evaluation of one's own reliance) predicted collaboration benefits incrementally beyond general metacognition — verification-skill training, not generic AI knowledge, is the calibrated counter to over-reliance.[2]
A validated collaborative-AI metacognition scale (planning, monitoring, evaluation of one's own reliance) predicted collaboration benefits incrementally beyond general metacognition — verification-skill training, not generic AI knowledge, is the calibrated counter to over-reliance.[2]
Addresses: Conceptual errors accepted as checked · Verification assumed rather than taught · Over-reliance on fluent output. Test a version of this lever in the PAN Lab.