10 to the 23 AI logo

Domain Atlas / Clinical documentation copilots (ambient scribes)

Case fileUnited States (academic health system; commercial SaaS ambient scribe)large deployment

Ambient scribe RCT + monitoring playbook

Explore this deployment in the PAN Lab ↗

The strongest causal evidence in the ambient-scribe family comes from a 24-week stepped-wedge, individually randomized trial of an ambient scribe across 66 practitioners and 71,487 notes (38 percent AI-generated), which found work exhaustion significantly reduced, professional fulfillment unchanged (a recorded null), roughly 22 minutes per day of documentation time returned, and diagnostic coding accuracy improved. Unlike the larger first-party deployment reports, this is a randomized estimate of the well-being and time effects — though it measured practitioner well-being and time, not per-note error rates.[]

What happened

This deployment is the causal-evidence anchor of the ambient-scribe family. The scribe was evaluated in a 24-week stepped-wedge, individually randomized trial across 66 practitioners and 71,487 notes, 38 percent of them AI-generated. The trial found work exhaustion significantly reduced, professional fulfillment unchanged — a recorded null, not a spun positive — roughly 22 minutes per day of documentation time returned, and diagnostic coding accuracy improved. Unlike the large first-party deployment reports elsewhere in this domain, this is a randomized estimate of the well-being and time effects, the strongest causal claim the family has.

What makes the organization structurally distinctive is that its monitoring apparatus is itself a published artifact. The same team released an open operations playbook for safety and effectiveness monitoring of ambient AI in production — built around this deployment. In most deployments the org-side monitoring function is assumed rather than shown; here it exists as a citable, designed subsystem. That matters for the whole domain, because it converts "there is oversight" from an assertion into a documented practice that can be inspected, copied, and held to.

The honesty boundaries are specific. The trial measured practitioner well-being and time, not per-note error rates — the class-level hallucination profile is carried by the domain's other cases, not re-claimed here. The vendor relationship is a commercial SaaS agreement with the vendor named in the trial record. And the improved coding accuracy sits inside a documented system-level risk: a coding arms race. Better AI documentation raises coding intensity; payers recalibrate in response; the escalation grows, and the clinician who attests to a code the AI suggested carries the liability. Better coding and upcoding pressure are neighbors, and the monitoring playbook is the org's stated answer to exactly that adjacency.

The sociotechnical reading

If the well-governed generation-into-record case shows what the write-gate and QA cost, this case shows two things that are rarer still: the strongest causal estimate of the benefit, and monitoring as a designed, documented subsystem rather than an assumed one. The trial is the anchor. A stepped-wedge individually randomized design is not a first-party dashboard; it is a causal claim, and it says the scribe reduced work exhaustion and returned time — while also recording a null on professional fulfillment, which is the mark of an honest evaluation rather than a marketed one. The Lab's service term for the whole family is most defensible here, and its limits are visible here too: well-being and time were measured; per-note accuracy was not.

The structurally important move is the monitoring playbook. Across this Atlas, the org-side oversight lever is usually an assertion — "a human reviews," "there is a QA process" — that the deployment does not evidence. This case is the counterexample: the monitoring function is published, so oversight is a thing you can point at, inspect, and reproduce, not a claim you take on faith. That is the honest form of the oversight lever, and it is what the other cases in this domain should be measured against. The counterweight the case builds in is the coding arms race: the trial's improved coding accuracy is not free, because better documentation raises coding intensity, invites payer recalibration, and shifts attestation liability onto the clinician — and the playbook exists precisely because better coding and upcoding pressure are neighbors that have to be watched together. The map's instruction is that the strongest evidence and the strongest monitoring in a domain still come with a named system-level risk, and the governable move is to keep the monitoring pointed at that risk, not just at the note. The honest boundary throughout: no care outcome is modeled on the Lab diagram. The patients whose visits are transcribed are boundary-only; the notes, edits, and monitoring findings are institutional signals, and the trial estimates and the coding-arms-race risk live in the case file, never on any network.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

afshar2025aGroundingPeer-reviewedSave

Afshar, M., Baumann, M.R., Resnik, F., et al. (2025). A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being. NEJM AI. https://doi.org/10.1056/AIoa2500945 https://pubmed.ncbi.nlm.nih.gov/41625485/

doi.org/10.1056/AIoa2500945

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)

afshar2025bGroundingPeer-reviewedSave

Afshar, M., et al. (2025). A Novel Playbook for Pragmatic Trial Operations to Monitor and Evaluate Ambient Artificial Intelligence in Clinical Practice. NEJM AI. https://doi.org/10.1056/AIdbp2401267 https://ai.nejm.org/doi/full/10.1056/AIdbp2401267

doi.org/10.1056/AIdbp2401267

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)

dai2025aGroundingPeer-reviewedSave

Dai, T., Kvedar, J.C., & Polsky, D. (2025). Policy brief: ambient AI scribes and the coding arms race. npj Digital Medicine. https://doi.org/10.1038/s41746-025-02272-z https://www.nature.com/articles/s41746-025-02272-z

doi.org/10.1038/s41746-025-02272-z

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalThe strongest causal evidence in the ambient-scribe family comes from a 24-week stepped-wedge, individually ra…

The strongest causal evidence in the ambient-scribe family comes from a 24-week stepped-wedge, individually randomized trial of an ambient scribe across 66 practitioners and 71,487 notes (38 percent AI-generated), which found work exhaustion significantly reduced, professional fulfillment unchanged (a recorded null), roughly 22 minutes per day of documentation time returned, and diagnostic coding accuracy improved. Unlike the larger first-party deployment reports, this is a randomized estimate of the well-being and time effects — though it measured practitioner well-being and time, not per-note error rates.

afshar2025aGroundingPeer-reviewedSave

Afshar, M., Baumann, M.R., Resnik, F., et al. (2025). A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being. NEJM AI. https://doi.org/10.1056/AIoa2500945 https://pubmed.ncbi.nlm.nih.gov/41625485/

doi.org/10.1056/AIoa2500945

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)

EmpiricalThe same team that ran the ambient-scribe trial released an open operations playbook for safety and effectiven…

The same team that ran the ambient-scribe trial released an open operations playbook for safety and effectiveness monitoring of ambient AI in production — the rare case where the organization-side monitoring function exists as a citable, designed subsystem rather than an assumed practice. That monitoring is the org's stated answer to a documented system-level risk of the technology: a coding arms race, in which better AI documentation raises coding intensity, payers recalibrate in response, and clinician attestation liability grows — so the improved coding accuracy the trial measured sits next door to an upcoding pressure the monitoring is meant to watch.

afshar2025bGroundingPeer-reviewedSave

Afshar, M., et al. (2025). A Novel Playbook for Pragmatic Trial Operations to Monitor and Evaluate Ambient Artificial Intelligence in Clinical Practice. NEJM AI. https://doi.org/10.1056/AIdbp2401267 https://ai.nejm.org/doi/full/10.1056/AIdbp2401267

doi.org/10.1056/AIdbp2401267

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)

dai2025aGroundingPeer-reviewedSave

Dai, T., Kvedar, J.C., & Polsky, D. (2025). Policy brief: ambient AI scribes and the coding arms race. npj Digital Medicine. https://doi.org/10.1038/s41746-025-02272-z https://www.nature.com/articles/s41746-025-02272-z

doi.org/10.1038/s41746-025-02272-z

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)