Domain Atlas / Clinical documentation copilots (ambient scribes)
Ambient AI scribe at scale (2.5M encounters)
Explore this deployment in the PAN Lab ↗
The largest documented ambient-scribe deployment ran a 10-week pilot at an integrated medical group and then scaled to 7,260 physicians and 2,576,627 patient encounters over fourteen months, with roughly 16,000 hours of documentation time saved and sustained physician support measured along the way. The system records the visit and drafts the clinical note; the clinician edits and signs, and the model-to-record write is gated both by that clinician review and by a standing internal quality-assurance program over the AI output — a real subsystem with a real cost, because the drafted note becomes a permanent record that later clinicians and later tools read as fact.[2]
What happened
This is the largest documented ambient-scribe deployment: a 10-week pilot at Kaiser Permanente's integrated medical group, then scale-up to 7,260 physicians and 2,576,627 patient encounters over fourteen months, with roughly 16,000 hours of documentation time saved and sustained physician support measured along the way. The system records the visit and drafts the clinical note; the clinician edits and signs.
What makes this the well-governed pole of the domain is where the AI's output goes and what stands between it and the permanent record. The model-to-record write is gated twice: by the clinician who reviews, edits, and signs each note, and by a standing internal quality-assurance program that samples the AI output across the deployment. That QA program is a real subsystem with a real cost — not a default, not an assumed practice, but a designed function that exists because the drafted note does not stay a draft. Today's AI-generated note becomes tomorrow's copied-forward clinical "fact": later clinicians read it as ground truth, and later tools ingest it as training or context. The review gate and the QA program are therefore not politeness; they are the contamination controls on a permanent record.
The honesty boundaries matter as much as the headline. The deployment's numbers are the deployer's own measurements — peer-reviewed, but first-party — and the papers do not name the vendor. A multisite study of 8,581 clinicians across five health systems tempers the magnitude: on the order of 13 to 16 fewer minutes per day, with no meaningful after-hours relief, a far more modest picture than any single system's dashboard. And at the product-class level, a validated per-note evaluation found hallucinations in about 31 percent of ambient-generated notes under structured review, versus about 20 percent of physician-written gold-standard notes — ambient notes are more thorough but less accurate. So the scale numbers here should be read as this system's dashboard, not the product class's guarantee, and the review-and-QA gate is what stands between a one-in-three hallucination rate and a contaminated record.
The sociotechnical reading
This domain moves the governable object from the alert to the record. In the alerting cases, the AI produces a signal a human acts on and the signal evaporates; here the AI produces a durable artifact — a clinical note — that is written into a permanent record and read forward as fact. That single difference relocates where governance has to live. The load-bearing control is not the model's fluency, which is high, but the two gates on the write into the record: the clinician's review before signing, and a standing quality-assurance program that samples the output across the whole deployment. This deployment is the well-governed pole because both gates are real and resourced, and the case exists to show what that costs and why it is not optional.
The reason it is not optional is the store's own dynamics. A drafted note does not stay a draft; it becomes the thing later clinicians copy forward and later tools ingest, so an error that survives the review does not just sit in one record — it propagates as inherited truth. Against a documented class-level hallucination rate near one in three, a review that has degraded into a rubber stamp under time pressure is not a small failure; it is a contamination source with a long half-life. The QA program is the answer to exactly that: it is the reconciliation that catches what individual reviews miss, and it is a payroll, not a policy. The honest counterweights are built into the case: the benefit is real but first-party-measured and smaller in independent multisite data, the vendor is unnamed, and the coding accuracy that improves alongside the notes lives next door to a coding-intensity arms race. The map's instruction is that when AI generates into a permanent record, the write is the surface to govern — with a human review that stays a real edit and a standing QA program over the output — because the record is read forward as fact whether or not anyone checked it. The honest boundary throughout: no care outcome is modeled on the Lab diagram. The patients whose visits are transcribed are boundary-only; the notes, edits, and QA findings are institutional signals, and the time-saved figures and hallucination rates live in the case file, never on any network.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.