10 to the 23 AI logo

Domain Atlas / Clinical decision support & deterioration alerting

Case fileUnited States (proprietary model in a widely used EHR; external validation at an academic health system)giant deployment

Proprietary EHR sepsis model (external validation)

Explore this deployment in the PAN Lab ↗

A widely implemented proprietary sepsis-prediction model shipped inside a common electronic-health-record platform and switched on across hundreds of hospitals was externally validated in 2021 across 38,455 hospitalizations at an academic health system: it achieved an area under the curve of 0.63, identified only 33 percent of sepsis cases, and had a positive predictive value of about 12 percent, generating roughly 109 alerts for every true sepsis case — a real-world performance the vendor had not fully examined before selling the model, and which an investigation attributed in part to undisclosed features such as antibiotic-order data that inflated internal validation.[2]

What happened

A proprietary sepsis-prediction model shipped inside a widely used electronic-health-record platform and was switched on across hundreds of hospitals. Unlike the developer-led evaluations elsewhere in this domain, its defining evidence is an independent one. A 2021 external validation across 38,455 hospitalizations at an academic health system found the model achieved an area under the curve of 0.63 (a measure of how well a score ranks a real case above a non-case, where 0.5 is a coin flip), identified only about 33 percent of sepsis cases, and had a positive predictive value near 12 percent — generating roughly 109 alerts for every true case of sepsis. An accompanying investigation reported that the vendor had not fully examined the model's real-world performance before selling it, and that the gap between the vendor's strong internal numbers and the weak external ones was explained in part by undisclosed features, including antibiotic-order data, that let the model partly predict a treatment clinicians had already started rather than the condition itself.

Two failure modes compound here. The first is alert fatigue: at 109 alerts per true case, the overwhelming majority of firings are false, and a clinician who is interrupted that often learns to dismiss the alert reflexively — the signal degrades into noise the workflow routes around. The second is vendor opacity: the model was shielded behind a corporate firewall that made independent scrutiny difficult, so the hospitals switching it on could not readily inspect what it did or how well it worked, and the external validation that exposed the gap came from researchers, not the vendor or the deploying institutions.

The arc has a correction, and the correction is the lesson. After the external criticism, the vendor overhauled the model — retraining it, changing the sepsis-onset definition, and reducing its reliance on antibiotic-order features. A 2026 multicenter prospective validation of the updated model across 227,091 encounters reported an area under the curve of 0.82 to 0.92 with positive predictive value of 0.13 to 0.26, and substantial between-site variability, with its authors urging local validation and alert-silencing strategies rather than trusting the model out of the box. The model got better — but only after independent scrutiny forced the issue, years after it was already running at scale, and the retuned model's own authors still say it cannot be trusted without local validation.

The sociotechnical reading

This is the domain's failure arc, and it inverts every other case in it. TREWS (Targeted Real-time Early Warning System), Advance Alert Monitor (AAM), and Sepsis Watch are deployments where the AI helped and the governable question was about the human loop or the staffing or the authority around it. The Epic-class sepsis model is the case where the tool was switched on across hundreds of hospitals before anyone independent had checked whether it worked — and when they did, it caught a third of sepsis at a hundred-plus false alarms per true case. The failure is not subtle model bias; it is deployment at scale ahead of validation, sold by a vendor who had not fully examined the real-world performance and shielded the model from the scrutiny that would have surfaced it.

The governable surfaces are the ones this deployment skipped. The first is validation-before-scale: an independent check that the model works on the patients it will sort, demanded before it becomes the default at hundreds of sites — the exact check that, when it finally ran, showed the model failing, and whose absence let the model run for years first. The second is the alert flood itself: 109 alerts per true case is not a tuning detail, it is the mechanism by which a nominally-helpful tool becomes actively harmful, training clinicians to ignore it and burying the rare true alert in noise — the honest correction is aggressive tiering and alert-silencing, which the retuned model's own authors now recommend. The third is vendor opacity: a model shielded behind a firewall cannot be governed by the institutions deploying it, because they cannot see what it does; the independent validation that exposed the gap is the thing the firewall was preventing. This case is why the domain's other governance questions matter — it is what deployment looks like when the confirmation loop is flooded, the validation is absent, and the vendor's internal number is the only one anyone has until an outsider checks. The honest boundary throughout: no patient or sepsis outcome is computed on the Lab diagram. The patients being scored are boundary-only; the alerts, overrides, and validation studies are institutional signals, and the accuracy numbers, the alert ratio, and the retune live in the case file, never on any network.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

wong2021cGroundingAcademicSave

Wong, A., Otles, E., et al. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307

https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307

Appears in: PAN framework development

Grounds: domain grounding: clinical AI (deterioration, imaging, documentation); model org: epic_sepsis_model_michigan

Topics: complexity-science

statnews2021GroundingInvestigativeSave

STAT News (2021, July 26). Epic's AI algorithms, shielded from scrutiny by a corporate firewall, are delivering inaccurate information on seriously ill patients. https://www.statnews.com/2021/07/26/epic-hospital-algorithms-sepsis-investigation/

https://www.statnews.com/2021/07/26/epic-hospital-algorithms-sepsis-investigation/

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage); model org: epic_sepsis_model_michigan

statnews2022GroundingInvestigativeSave

STAT News (2022, Oct 3). Epic overhauls popular sepsis algorithm criticized for faulty alarms. https://www.statnews.com/2022/10/03/epic-sepsis-algorithm-revamp-training/

https://www.statnews.com/2022/10/03/epic-sepsis-algorithm-revamp-training/

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage); model org: epic_sepsis_model_michigan

wong2026aGroundingPeer-reviewedSave

Wong, A., Currey, D., Schwinne, M., et al. (2026). Multicenter Prospective Validation of an Updated Proprietary Sepsis Prediction Model. JAMA Network Open. https://doi.org/10.1001/jamanetworkopen.2026.0181 https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2845595

doi.org/10.1001/jamanetworkopen.2026.0181

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage)

Topics: complexity-science

felisberto2024aGroundingPeer-reviewedSave

Felisberto, M., dos Santos Lima, G., Celuppi, I.C., et al. (2024). Override rate of drug-drug interaction alerts in clinical decision support systems: A brief systematic review and meta-analysis. Health Informatics Journal. https://doi.org/10.1177/14604582241263242 https://journals.sagepub.com/doi/10.1177/14604582241263242

doi.org/10.1177/14604582241263242

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage)

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalA widely implemented proprietary sepsis-prediction model shipped inside a common electronic-health-record plat…

A widely implemented proprietary sepsis-prediction model shipped inside a common electronic-health-record platform and switched on across hundreds of hospitals was externally validated in 2021 across 38,455 hospitalizations at an academic health system: it achieved an area under the curve of 0.63, identified only 33 percent of sepsis cases, and had a positive predictive value of about 12 percent, generating roughly 109 alerts for every true sepsis case — a real-world performance the vendor had not fully examined before selling the model, and which an investigation attributed in part to undisclosed features such as antibiotic-order data that inflated internal validation.

wong2021cGroundingAcademicSave

Wong, A., Otles, E., et al. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307

https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307

Appears in: PAN framework development

Grounds: domain grounding: clinical AI (deterioration, imaging, documentation); model org: epic_sepsis_model_michigan

Topics: complexity-science

statnews2021GroundingInvestigativeSave

STAT News (2021, July 26). Epic's AI algorithms, shielded from scrutiny by a corporate firewall, are delivering inaccurate information on seriously ill patients. https://www.statnews.com/2021/07/26/epic-hospital-algorithms-sepsis-investigation/

https://www.statnews.com/2021/07/26/epic-hospital-algorithms-sepsis-investigation/

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage); model org: epic_sepsis_model_michigan

EmpiricalAfter external criticism, the vendor overhauled the sepsis model — retraining it, changing the sepsis-onset de…

After external criticism, the vendor overhauled the sepsis model — retraining it, changing the sepsis-onset definition, and reducing its reliance on antibiotic-order features. A 2026 multicenter prospective validation of the updated model across 227,091 encounters reported an area under the curve of 0.82 to 0.92 with positive predictive value of 0.13 to 0.26 and substantial between-site variability, and its authors urged local validation and alert-silencing strategies rather than trusting the model out of the box — a correction that arrived only after independent scrutiny of a model that had already been deployed at scale behind a corporate firewall shielding it from outside inspection.

statnews2022GroundingInvestigativeSave

STAT News (2022, Oct 3). Epic overhauls popular sepsis algorithm criticized for faulty alarms. https://www.statnews.com/2022/10/03/epic-sepsis-algorithm-revamp-training/

https://www.statnews.com/2022/10/03/epic-sepsis-algorithm-revamp-training/

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage); model org: epic_sepsis_model_michigan

wong2026aGroundingPeer-reviewedSave

Wong, A., Currey, D., Schwinne, M., et al. (2026). Multicenter Prospective Validation of an Updated Proprietary Sepsis Prediction Model. JAMA Network Open. https://doi.org/10.1001/jamanetworkopen.2026.0181 https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2845595

doi.org/10.1001/jamanetworkopen.2026.0181

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage)

Topics: complexity-science