Behavioral-health & crisis triage
Risk scores and triage rankers deciding whose crisis is seen first — where the rare event is nearly impossible to predict reliably, the flag moves a proxy more surely than the outcome, and a score can quietly gate access to care.
Use cases
What AI is doing here
Suicide-risk prediction & outreach
PredictiveEHR- or registry-based models that flag patients at elevated suicide risk to trigger clinician review or outreach — rare-event scoring in which the large majority of flags are false positives.
Crisis-line severity triage
PredictiveMachine-learning severity rankers that reorder crisis-line contacts so the most-at-risk are reached first.
Conversational mental-health intake & support
GenerativeConversational systems that handle self-referral, triage, or supportive self-help contact for mental-health care before, between, or in place of clinician time.
Substance & overdose risk scoring
PredictiveProprietary risk scores embedded in prescribing and pharmacy workflows that can gate a person's access to controlled medications.
Clinician fidelity & care-quality scoring
PredictiveAI that scores the clinician's own practice from session or call audio — coverage rising from small hand-review samples toward every encounter — where the instrument measures the workforce, and the governance question is who calibrates the measurement people are managed by.
Passive safety-surveillance monitoring
PredictiveAlways-on monitoring of people in institutional care or custody of a duty-bearing body — school-account scanning, ward sensor systems — flagging risk from ambient activity rather than a clinical encounter, typically under contested consent and without published error rates.
Case files
What has gone wrong, and right
Documented deployments, presented as model organizations calibrated to the evidence, with full citations.
REACH VET
United States (federal; Department of Veterans Affairs, Veterans Health Administration; nationwide)A national VA model flags the 0.1% of patients at highest suicide risk each month — a working, coordinator-mediated program that improved appointments and safety plans without reducing suicide deaths in two evaluations, and that a 2024 investigation found under-flagged women by leaving their risk factors out of the model.
Stress-test this shape in the PAN Lab →Vanderbilt VSAIL suicide-risk alert
Nashville, Tennessee, USA (Vanderbilt University Medical Center)A suicide-risk model that scores routine hospital records made clinicians far more likely to screen a flagged patient — but only when its alert was made interruptive rather than passive, in a trial too small to show a single prevented suicide.
Stress-test this shape in the PAN Lab →Kaiser Permanente Suicide-Risk Model
United States (California; Kaiser Permanente Northern California integrated health system)An EHR-embedded machine-learning score that flags high suicide-risk patients within about 30 minutes of a virtual mental-health intake and routes them into the same assessment-and-outreach workflow as a positive self-report screen — a carefully co-designed integration whose main effect is to add a very-low-PPV sensor to an existing human screen, published as a feasibility design with no outcome-effectiveness evaluation.
Stress-test this shape in the PAN Lab →Crisis Text Line & Loris.ai
United States (national nonprofit; also operates via affiliates in Canada, the UK, and Ireland)A national crisis line built an in-house model to rank its most-at-risk texters and reorder who gets a counselor first — while the same anonymized conversations became a corpus that was routed to a for-profit it owned a stake in, to train commercial software. A 2022 exposé ended the data-sharing in three days; the deeper failure was that the review body that should have gated it was reportedly never asked, and that consent was gathered at the moment of crisis via a Terms of Service.
Stress-test this shape in the PAN Lab →NarxCare
United States (state Prescription Drug Monitoring Programs)A proprietary overdose-risk score embedded in most US prescription-monitoring programs assigns patients secret 000-999 numbers that can gate their access to pain medication — never independently validated, not regulated by the FDA, uncontestable by the patient, and, when independent researchers rebuilt a version of it in 2026, far less accurate than the vendor claims.
Stress-test this shape in the PAN Lab →Limbic Access (NHS Talking Therapies)
England, United Kingdom (NHS Talking Therapies for Anxiety and Depression, formerly IAPT)A self-referral triage chatbot that genuinely widened access to NHS talking therapies — with the largest gains among under-served groups — while nearly all the evidence that it works was measured, published, and acted on by the company that built it.
Stress-test this shape in the PAN Lab →Woebot (a governed app wind-down)
United States (Woebot Health, San Francisco, California; direct-to-consumer and enterprise or health-system distribution)A rule-based CBT chatbot used by roughly 1.5 million people was deliberately retired on a published schedule — with a transcript-access window and scheduled data anonymization — because its maker judged the FDA marketing-authorization economics unsustainable, not because the tool failed clinically. The Atlas's reference point for what a responsible wind-down looks like.
Stress-test this shape in the PAN Lab →LyssnCrisis counselor QA at ProtoCall Services (988)
United States (ProtoCall Services, Portland, Oregon: a SAMHSA-contracted national 988 backup provider and the primary 988 line for New Mexico; tool built by Lyssn.io, Inc.)An AI quality-assurance tool that scores 988 crisis counselors on their own call practice — not the callers — expanding measured review from under 3% of calls toward nearly all of them. The domain's positive, governed counterexample: the AI sits on the supervision link over the operators, with strong peer-reviewed reliability evidence but a vendor-authored, still-unpublished trial of whether its feedback actually helps.
Stress-test this shape in the PAN Lab →The discontinuation that wasn't: a school communication scanner, swapped not stopped
United States (national vendor, roughly 1,500 districts and about 6 million students); governance events centered in Lawrence USD 497, Kansas and Vancouver Public Schools, Washington; litigation in the U.S. District Court for the District of KansasAn AI service scans everything students write on school-issued accounts — email, homework, art, the student newspaper — around the clock, flagging anything that might signal self-harm. Used by roughly 1,500 US districts covering about 6 million students, it works as a multi-hop chain: a machine flags, an off-site vendor reviewer triages, a counselor gets an alert, and an after-hours emergency flag can send police to a home. In Vancouver, Washington nearly 1 in 10 students triggered alerts in a year; in Lawrence, Kansas about two-thirds of 1,200-plus flags were deemed nonissues — and the archive of flagged suicide notes and coming-out essays was accidentally released, nearly 3,500 unredacted documents at once. The atlas's discontinuation-that-wasn't case: when nine students sued, Lawrence quit the vendor mid-lawsuit and quietly swapped in an equivalent one with no board vote, and a federal judge found the district broke open-records law trying to withhold the swap's paperwork from the student journalists it had been flagging.
Stress-test this shape in the PAN Lab →Oxevision camera monitoring on NHS mental health wards
England, United Kingdom (NHS mental health inpatient wards; early adopter Essex Partnership University NHS Foundation Trust, documented deployments also at Oxford Health NHS Foundation Trust and others; tool built by Oxehealth Ltd, rebranded LIO / LIO Health in August 2025)Camera-based, contact-free vital-signs monitoring installed in NHS psychiatric bedrooms. In March 2026 the health ombudsman partly upheld a complaint that a trust never sought a patient's consent, gave her no information, and would not switch the camera off when she asked — and found even the revised procedure still lets staff film a refusing, capacitous patient. The defining defect is vendor capture of the evaluation link: the company selling the tool authored the trust's business case and shaped the evidence meant to vouch for it, then rebranded mid-inquiry. Contested throughout; an ICO investigation is open.
Stress-test this shape in the PAN Lab →ODMAP overdose spike alerts on a drug-enforcement-housed store
United States (all 50 states, the District of Columbia, and Puerto Rico; ODMAP is operated by the Washington/Baltimore High Intensity Drug Trafficking Area, a federal Office of National Drug Control Policy program)A nationwide, rule-based early-warning network that fires overdose spike alerts from aggregate, unnamed event counts rather than scoring any individual. Its algorithm is a plain threshold, not machine learning; the hard questions live off the model and on the links - the shared store sits inside a federal drug-enforcement program where public-health and law-enforcement read access coexist, the threshold learns from the store's own reporting, and every figure is program-self-published with no independent evaluation.
Stress-test this shape in the PAN Lab →System map
Who is in the system, and what pushes on it
Who is in the system
- Frontline workers. Caseworkers, screeners, eligibility staff — the operator network whose judgment the system augments or erodes.
- Supervisors & QA. The institutional correction layer: overrides, second reads, quality review.
- Agency leadership. Owns procurement, policy, and the authority map; answers for the system publicly.
- Served people & families. Those the decisions land on. Deliberately outside the PAN dynamics — their outcomes are measured, never simulated.
- Vendors. Build and update the systems; hold the information asymmetry procurement must govern.
- Regulators & oversight bodies. Boards, auditors, data-protection officers, inspectorates — external correction capacity.
- Advocates & community organizations. Surface harms institutions do not see; historically the earliest accurate signal.
Dominant pressures
- Caseload surge. Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck. One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity. The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure. Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Compliance over substance. Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
Governance
Questions leaders should be asking
- 1. Does the evidence show this system changed the outcome it exists to change, or only an easier-to-measure proxy like screening or contact rates — and did anyone independent of the builder produce that evidence?
- 2. Under a rare event, most flags will be false — so what does each flag cost a clinician's attention, and what does the alert displace when it interrupts a caseload already at capacity?
- 3. Can a person see and contest a score that shapes their access to care — and does anyone check whether it reads some groups as higher-risk for reasons unrelated to their actual need?
- 4. When a mental-health tool is withdrawn — by its vendor, a regulator, or a lapsed budget — what happens to the people mid-care and to the intimate data they entrusted to it?
For the actions behind these questions, see the Practice Library.
Seeing your organization in this domain? Mapping its actual pathways, pressures, and correction capacity is engagement work.
Work With 1023AI →