Domain Atlas / Behavioral-health & crisis triage
REACH VET
The U.S. Department of Veterans Affairs' REACH VET program has run a monthly suicide-risk model across the Veterans Health Administration since 2017, scoring about 6.28 million patients and flagging the top 0.1% at each facility (roughly 6,300 to 6,700 veterans a month, more than 130,000 since 2017); an independent re-analysis of 2018 data found the top-0.1% flag has a positive predictive value near 0.05% and a false-negative rate of about 98% for death by suicide, and a 2024 investigation reported that the model treated being a white man as a stronger risk signal than factors specific to women and excluded military sexual trauma and intimate-partner violence from its variables, a characterization VA has contested by framing the excluded factors as less predictive.[4]
What happened
REACH VET (Recovery Engagement and Coordination for Health — Veterans Enhanced Treatment) is a nationwide Veterans Health Administration suicide-prevention program, deployed across VHA in April 2017, that runs a predictive statistical model monthly on veterans' electronic health records to flag those at highest risk. The operational model is a penalized regression (Lasso) using 61 variables in six categories — demographics (including age, gender, marital status, and race/ethnicity), diagnoses, medications, utilization, prior suicide attempt, and interaction terms such as marital-status-by-gender — reduced from a 381-variable 2015 proof-of-concept because the larger model was computationally unwieldy; Harvard's Ronald Kessler refined it and added machine-learning methods. Each month the model scores every patient with VHA inpatient or outpatient care in the prior two years — about 6.28 million people — and produces the top 0.1% highest-risk list at each VA facility, roughly 6,300 to 6,700 veterans a month; more than 130,000 have been identified since 2017. When a veteran is flagged, a facility Suicide Prevention Coordinator receives the name on an internal dashboard and notifies the clinician, who re-evaluates suicide risk and treatment and conducts non-scripted outreach offering enhanced care, safety planning, increased monitoring, and coping support. The flag is advisory: VA frames the model as a prompt for clinician review, not an automated decision, and states that AI will "never replace human intervention."
Two independent re-analyses put hard numbers on what the flag does and does not catch. For predicting death by suicide, the top-0.1% tier concentrates risk — it dies by suicide at roughly 19 to 30 times the overall VHA rate — yet against about 179 suicides a month across 6.28 million scored patients, an accuracy re-analysis of 2018 data found the flag has a positive predictive value near 0.05% and a false-negative rate of about 98% (sensitivity near 2%), meaning it misses roughly 98% of suicides; for the broader suicide-attempt-or-death outcome the positive predictive value was higher, about 5%. A separate retraining experiment moved the combined-outcome predictive value only slightly and not significantly, evidence that the ceiling is the rare-event base rate rather than a fixable modeling flaw. On process outcomes, the VA's own 2021 effectiveness evaluation (a triple-differences design across 141 facilities and 173,313 veterans) found REACH VET associated with more completed outpatient appointments, more new safety plans, fewer mental-health admissions, and fewer documented suicide attempts — but not with reduced death by suicide or all-cause mortality. A 2025 mortality follow-up (266,246 observations) replicated that null, with every confidence interval crossing one. The program reliably improved the things it could measure quickly, while the outcome it was built to move did not move in either study.
In May 2024 an investigation by The Markup and The Fuller Project reported that the model treated being a white man as a stronger indicator of suicide risk than factors affecting women, giving preference to veterans who were "divorced and male" or "widowed and male" but to no female group; it reported that military sexual trauma and intimate-partner violence — both linked to elevated suicide risk in women veterans — were excluded from the model, and that the model did not account for LGBTQ+ identity despite VA research that transgender veterans die by suicide at about twice the cisgender rate. VA's suicide-prevention executive framed the exclusions as a predictive-strength judgment, saying military sexual trauma "was not among the most powerful for us to be able to predict suicide risk," so the equity finding is best read as documented but contested rather than a VA-confirmed design intent. The concern coincided with a measured demographic divergence: the female-veteran suicide rate rose about 24% between 2020 and 2021, roughly four times the increase among male veterans. In October 2024 VA announced REACH VET 2.0, adding military sexual trauma, intimate-partner-violence, and women-specific medical factors, with a stated commitment to evaluate the new model "for performance and bias before it is deployed," after Senator Jon Tester introduced legislation requiring the changes. By December 2025, trade-press reporting said 2.0 had launched in 2025, had added the new factors, and had removed race and ethnicity as variables — though no peer-reviewed 2.0 documentation or independent bias audit had been published, so 2.0's actual performance and equity are not yet independently verified. A 2022 GAO review, mandated by the 2019 Hannon Act, had described the program and noted that VHA planned but had not yet completed studies of how the model's performance varied by age, sex, and race; it made no REACH-VET-specific recommendations. Congress has continued to back the approach, with the FY2026 Military Construction-VA appropriations bill (signed November 2025) allocating funds for suicide prevention and encouraging expanded predictive modeling.
The sociotechnical reading
Almost every failure in this Atlas is a failure of the human loop — a screener who defers, an override that cannot be exercised, a flag buried under thousands of others. REACH VET is the case where the human loop works and is not the problem. The flag is advisory, the coordinator-to-clinician handoff is a real staffed review, clinicians keep their discretion, and veterans can decline outreach. And the program does what a good triage program should do on the measures nearest to hand: more appointments kept, more safety plans written, fewer documented attempts. This is the case that asks the harder question — what remains ungoverned when the human-in-the-loop is genuine?
Three answers, and none of them is accuracy or the override. The first is a rare-event ceiling: with about 179 suicides a month among 6.28 million scored patients, even a tier concentrated at 19-to-30 times the base rate captures only about 2% of suicides, so the vast majority happen among the un-flagged — a mathematical property no model improvement climbs, which is why a retraining experiment barely moved it. The map's honest instruction here is a tiering that respects real scarcity and never reads a clear flag as an all-clear. The second is a surrogate-versus-target gap. The program improved the proximal metrics it could see in the EHR and showed no effect on the outcome it exists to prevent, in two evaluations; the true outcome returns only with lag, in a separate mortality repository, reconciled only by periodic external study. A program can be working on its dashboard and not on its purpose, and the only way to know is to reconcile the surrogate against the target on a cadence — a check drawn empty at baseline here. The third, and sharpest, is that a fixed 0.1% budget of scarce help is allocated by predicted risk, which makes the model's variable set an allocation decision, not merely an accuracy one. Leaving out military sexual trauma, intimate-partner violence, and LGBTQ+ identity did not just make the model less accurate for women — it decided that factors specific to women would carry less weight in who reaches the front of the line. That is why the corrective that mattered was not a better fit but a different variable set (REACH VET 2.0), and why it arrived years after an outside investigation rather than from an internal check: the pre-deployment subgroup and bias validation that would have caught the miss was, per GAO, planned and not completed. The lesson this case adds to the Atlas is that a working human loop and improving numbers can coexist with missing the outcome you care about and mis-allocating scarce care by demographic — and that the governable surfaces are upstream and downstream of the human loop, not in it: the choice of features (an equity decision disguised as a modeling one), the reconciliation of surrogate against target, and the subgroup check that never ran. The honest boundary throughout: none of this models suicide, and the veterans being ranked are not in the diagram. The error structure, the outcome gap, and the contested demographic finding are documented outside any system map, and a monthly flag is an institutional signal, never a life.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.