Domain Atlas / Clinical decision support & deterioration alerting
TREWS sepsis early-warning system
Explore this deployment in the PAN Lab ↗
The Targeted Real-Time Early Warning System (TREWS), a machine-learning sepsis early-warning model, was evaluated prospectively across five hospitals of an academic health system covering 590,736 monitored patients — the largest prospective study of an ML sepsis system on record. Its central finding was conditional on the human loop: sepsis patients whose alert was evaluated and confirmed by a provider within three hours had a 3.3 percentage-point absolute and 18.7 percent relative adjusted reduction in in-hospital mortality, with less organ failure and shorter stays, while the alert on its own did not; a companion study found provider uptake varied with experience, unit culture, and alert context.[2]
What happened
The Targeted Real-Time Early Warning System (TREWS) is a machine-learning model that continuously scores hospitalized patients for sepsis from their electronic health records and raises an alert when risk crosses a threshold, for a provider to evaluate. It was deployed across the five hospitals of Johns Hopkins Medicine and evaluated in a prospective, multi-site study covering 590,736 monitored patients — the largest prospective evaluation of a machine-learning sepsis system on record, published in Nature Medicine in 2022.
The study's central finding is conditional, and the condition is the human loop. Among patients who went on to have sepsis, those whose TREWS alert was evaluated and confirmed by a provider within three hours had a 3.3 percentage-point absolute and 18.7 percent relative adjusted reduction in in-hospital mortality, along with less organ failure and shorter length of stay. The benefit did not come from the alert firing; it came from the alert being confirmed and acted on in time. A companion paper on adoption found that whether providers evaluated and confirmed alerts varied with their experience of the system, unit culture, and the context in which the alert arrived — so the confirmation step that carried the entire benefit was itself uneven, driven by trust and workflow rather than automatic.
Two honesty boundaries are built into what this evidence can claim. First, the evaluation is observational, not randomized: the mortality benefit is an association between confirmation and outcome, and providers who engaged with alerts may differ systematically from those who did not — in caseload, in the patients they saw, in attentiveness — in ways the statistical adjustment cannot fully capture. Confirmation may partly mark the patients who were already going to do better, not only cause the improvement. Second, the evaluation is developer-led: TREWS was built at the deploying institution and commercialized through a company founded by its principal investigator, so the strongest numbers in the record were produced, on the developer's own patients, by the party with the greatest stake in the result. The study is prospective and peer-reviewed — genuinely stronger evidence than most deployed clinical AI can show — but it is not independent, and at the time of the evaluation no outside group had replicated the mortality effect.
The sociotechnical reading
TREWS is the case where the AI works, the human loop works, and the benefit is measured — and the governable surface is still not the alert. Almost every failure elsewhere in this Atlas is a broken human loop: an override no one exercises, a review the workload has hollowed out, a flag buried in a flood. TREWS inverts that. The alert on its own does nothing; the mortality benefit exists only because providers confirm alerts within three hours and act. The load-bearing part of the system is the confirmation step, which means the thing an institution must protect is not the model's accuracy but the clinician's capacity and willingness to evaluate what it surfaces in time. A benefit that runs entirely through a human confirmation is a benefit you can lose without touching the model at all — just by letting the confirmation step erode under caseload, or by deploying at a site where trust and unit culture make providers stop looking. The companion adoption finding is the warning: the confirmation rate was already uneven, so the measured benefit is not a fixed property of the tool but a property of the workflow around it, one site at a time.
The second governable surface is independence. This is the strongest evidence base for a deployed clinical AI in the Atlas — prospective, multi-site, peer-reviewed, more than half a million patients — and it is still developer-led and observational. The people who built and commercialized the model produced the numbers, on their own patients, and the mortality effect is an association a randomized trial has not confirmed. Nothing here is fabricated or hidden; the study openly is what it is. But the honest reading is that "peer-reviewed and prospective" is not the same as "independently validated," and the check that is missing — a replication of the mortality effect by a group with no stake in the result — is exactly the check that developer-led evidence cannot supply about itself. The map's instruction is to treat the confirmation workflow as the resourced, monitored thing it has to be, and to treat even excellent developer-produced evidence as awaiting the independent replication that would turn a strong association into a settled effect. The honest boundary throughout: no patient outcome is modeled on the Lab diagram. Patients being scored are not in the dynamics; the alerts, confirmations, and record writes are institutional signals, and the mortality finding lives in the case file, not on any network.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.