Domain Atlas / Behavioral-health & crisis triage
Vanderbilt VSAIL suicide-risk alert
In a single-center randomized trial across three Vanderbilt neurology clinics (August 2022 to February 2023), an EHR suicide-risk model flagged 596 of 7,732 encounters (about 8%) at a 2%-or-higher 30-day-risk threshold; making the identical alert interruptive rather than passive led clinicians to elect a suicide-risk screen in 42% of encounters (121/289) versus 4% (12/307) for a passive chart icon, an adjusted odds ratio of 17.70 (95% CI 6.42–48.79). Screening remained fully advisory: about 58% of interruptive and 96% of passive alerts produced no screening.[3]
What happened
Vanderbilt University Medical Center (VUMC) built an electronic-health-record suicide-risk model in-house — branded VSAIL (Vanderbilt Suicide Attempt and Ideation Likelihood) in the institution's communications and trial registry, though the peer-reviewed papers describe it plainly as an EHR-based, real-time suicide-risk model. It is a random-forest model that estimates a patient's probability of a suicide attempt within 30 days from routine structured records — diagnoses, medications, five years of visit history, demographics, and a neighborhood-deprivation index by ZIP — and requires no PHQ-9 or other questionnaire. The machine-learning method was first published in 2017, distinguishing 3,250 suicide-attempt cases (identified via self-injury claim codes among 5,167 adult patients) at an AUC of about 0.84. In a prospective silent-mode study running June 2019 to April 2020, the model generated 115,905 predictions on 77,973 patients without triggering any alert, reporting a c-statistic of 0.797 for attempt and 0.836 for ideation center-wide — but far lower (0.544 for attempt) in behavioral-health settings, where suicidality is most concentrated — and it required manual recalibration after early miscalibration. In the highest-risk quantile the number-needed-to-screen was 23 for ideation and 271 for attempt, and those figures differed by subgroup (attempt number-needed-to-screen 176 for Black versus 373 for White patients, 256 for men versus 323 for women) — differences the source reports but does not adjudicate as bias versus base-rate.
The distinctive part is a single-center randomized trial (NCT05312437, deployed as "Vanderbilt Safecourse"). From August 2022 to February 2023, across three VUMC ambulatory neurology divisions, the model flagged 596 of 7,732 encounters (about 8%) at a 2%-or-higher 30-day-risk threshold and randomized them one-to-one inside the EHR to two alert modalities: an interruptive pop-up advisory that had to be dismissed to proceed, or a passive chart icon. With the identical score, threshold, and embedded screen, the interruptive alert led clinicians to elect a suicide-risk screen in 42% of encounters (121/289) versus 4% (12/307) for the passive icon — an adjusted odds ratio of 17.70 (95% CI 6.42–48.79) — and documented Columbia Suicide Severity Rating Scale assessments rose to 22% against a prior-year baseline of 8%. Screening stayed fully advisory: the pop-up could be dismissed with no required action, and about 58% of interruptive and 96% of passive alerts produced no screening. No suicidal ideation or attempts were documented in either arm during 30-day follow-up, and the trial was explicitly not powered for clinical outcomes, so it showed more screening but could not speak to reduced harm. The researchers named alert fatigue as the central tradeoff and noted clinicians preferred the less-effective passive alert; the trial's consent waiver was justified partly to protect clinicians who might credibly disagree with the tool and decline to screen. Governance was comparatively rigorous — IRB approval with embedded ethicists and legal consultation, prospective registration, and a published silent-mode study before any live alerting — and the model was applied separately to Navy primary care (260,583 patients at Naval Medical Center Portsmouth: 0.77 AUC applied directly, 0.92 after retraining). It is not an independently cleared device, and live interruptive alerting is documented only within the single-center trial, not as routine center-wide deployment.
The sociotechnical reading
Almost every tool in this Atlas puts the governance fight on the model — its accuracy, its bias, its inputs. VSAIL moves the fight one step downstream, to the channel between the model and the human. The trial is a rare clean measurement: hold the score, the threshold, the population, and the screen all constant, change only whether the alert interrupts the clinician or waits in the chart, and screening moves roughly tenfold. The lever is not the prediction; it is the delivery surface and the attention it commands. That is the reframing the case offers — the same signal can be nearly inert or highly active depending on a design choice about interruption that has nothing to do with how good the model is.
But the channel is double-edged, and the case documents both edges. Turn the alert off (passive) and a real flag reaches no engaged human — the failure is a quiet miss, not a wrong action. Turn it up (interruptive) on a flag that is mostly false positives — the highest-risk quantile still needed 271 screens to find one attempt — and you invite alert fatigue, the reflex to dismiss a familiar interruption, which hollows out the very screen the alert exists to prompt. The load-bearing structures are therefore the attention channel itself (kept alive, not blunted into habituation), the deliberate bound on how often it fires (the roughly 8% flag cap that makes targeted screening survivable), the protected right of the clinician to be right when the flag is wrong (the consent-waiver logic), and the slow silent-monitoring loop that is the only thing watching a single in-house model drift. The lesson the Atlas draws is that where a model is purely advisory, the decisive governance object can be the modality of the prompt, not the merits of the prediction — and that the same interruption that rescues a signal from neglect can, overused, manufacture the neglect it was meant to cure. It is also the Atlas's sharpest reminder that a process win is not an outcome win: screening rose roughly eighteenfold and the trial still could not show one prevented harm.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.