Clinical decision support & deterioration alerting
Machine-learning early-warning models that flag hospitalized patients for sepsis or clinical deterioration — the domain where an AI's measured benefit is real but runs entirely through the human loop it interrupts. The same alert that saves a life when a clinician confirms it in time becomes a source of fatigue when it fires a hundred times per true case; what separates the two is whether the confirmation workflow is resourced, whether the model was validated independently of the vendor who sells it, and whether anyone reconciles the alerts against the outcomes they were meant to change. The Lab networks here model only the deploying hospital — its models, clinicians, and records; the patients being scored sit outside the dynamics, and no clinical outcome is ever computed on a diagram.
Use cases
What AI is doing here
Sepsis & deterioration early-warning alerting
PredictiveMachine-learning models that continuously score hospitalized patients' electronic-health-record data and raise an alert when sepsis or clinical deterioration is predicted, for a clinician to evaluate and confirm — a signal whose measured benefit is contingent on a resourced human confirmation step.
Rapid-response & virtual-nurse escalation
PredictiveDeterioration alerts routed through a mediating escalation tier — a regional virtual-nurse desk or a rapid-response-team nurse — who screens the score and mobilizes bedside care, embedding the model inside a staffed workflow whose hidden coordination labor is what makes the alert actionable.
Proprietary EHR-embedded risk scores
PredictiveVendor-built risk models shipped inside a widely used electronic-health-record platform and switched on across many hospitals at once, where the model's real-world accuracy and alert burden may not be independently validated before deployment and the vendor's internal evaluation is shielded from outside scrutiny.
Case files
What has gone wrong and right
Documented deployments, presented as model organizations calibrated to the evidence, with full citations.
TREWS sepsis early-warning system
United States (Johns Hopkins Medicine; five hospitals across Maryland and Washington, DC)A machine-learning sepsis alert deployed across five hospitals of an academic health system, evaluated on 590,736 patients — the largest prospective study of its kind. Its mortality benefit was real but conditional: it accrued only to patients whose alert a provider confirmed within three hours, and the alert alone did nothing. The evaluation was prospective and peer-reviewed, but observational and run by the developer who commercialized the model.
Explore this deployment in the PAN Lab →Advance Alert Monitor (AAM) deterioration model
United States (Kaiser Permanente Northern California; 21 hospitals)An in-hospital deterioration model running around the clock across 21 hospitals, firing about twelve hours ahead of predicted deterioration. Its NEJM-measured mortality benefit is inseparable from where the alert goes: not to the bedside, but to a dedicated regional tier of critical-care virtual nurse consultants who screen every alert before escalating. The governance here is not a checkbox — it is a staffed, 24/7 subsystem with a payroll.
Explore this deployment in the PAN Lab →Sepsis Watch deep-learning detection system
United States (academic hospital; registered clinical trial NCT03655626)A deep-learning sepsis detector scoring every emergency-department patient every five minutes, fronted by rapid-response nurses on treatment-bundle timers. Its fault line is an authority split: the nurse who receives the alert is not the physician empowered to act on it. An independent ethnography found the system worked only because nurses did hidden, undervalued repair work to make a risk score actionable across a professional hierarchy.
Explore this deployment in the PAN Lab →Proprietary EHR sepsis model (external validation)
United States (proprietary model in a widely used EHR; external validation at an academic health system)A proprietary sepsis-prediction model shipped inside a common EHR and switched on across hundreds of hospitals — then externally validated to catch only a third of sepsis cases at roughly 109 alerts per true case, a real-world performance the vendor had not fully examined before selling it, shielded behind a firewall from outside scrutiny. This is the domain's failure arc: deployment at scale ahead of independent validation, alert fatigue, and vendor opacity — corrected only after outside criticism forced a retune.
Explore this deployment in the PAN Lab →System map
Who is in the system and what pushes on it
Who is in the system
- Frontline workers. Caseworkers, screeners, eligibility staff — the operator network whose judgment the system augments or erodes.
- Supervisors & QA. The institutional correction layer: overrides, second reads, quality review.
- Agency leadership. Owns procurement, policy, and the authority map; answers for the system publicly.
- Served people & families. Those the decisions land on. Deliberately outside the PAN dynamics — their outcomes are measured, never simulated.
- Vendors. Build and update the systems; hold the information asymmetry procurement must govern.
- Regulators & oversight bodies. Boards, auditors, data-protection officers, inspectorates — external correction capacity.
Dominant pressures
- Caseload surge. Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck. One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity. The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure. Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Data & policy drift. The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
Governance
Questions leaders should be asking
- 1. This model's measured benefit ran entirely through providers confirming its alerts in time — so is the confirmation step actually resourced, or is the benefit being claimed for a review the workload cannot sustain?
- 2. Who validated this model, and were they independent of the party selling it? A prospective, peer-reviewed evaluation run by the developer is still the strongest number produced by the most interested party.
- 3. How many alerts fire per true case, and who measures whether the clinicians being interrupted have started tuning the alarm out — the fatigue that turns a working tool into background noise?
- 4. Is anyone reconciling the alerts and the confirmations against the outcome the system exists to change, or only against how often it fired — and would a model that drifted out of calibration be caught before or after a bad patient outcome?
For the actions behind these questions, see the Practice Library.
Seeing your organization in this domain? Mapping its actual pathways, pressures, and correction capacity is engagement work.
Work With 1023AI