10 to the 23 AI logo

Domain Atlas / Behavioral-health & crisis triage

Case fileUnited States (ProtoCall Services, Portland, Oregon: a SAMHSA-contracted national 988 backup provider and the primary 988 line for New Mexico; tool built by Lyssn.io, Inc.)large deployment

LyssnCrisis counselor QA at ProtoCall Services (988)

Work with this case in the PAN Lab ↗

An AI quality-assurance tool deployed on a national 988 backup line scores crisis counselors' own call practice rather than callers, expanding measured review from the under-3% of calls that had been reviewed by hand toward nearly all of them; a peer-reviewed reliability study of 476 labeled calls reported agreement with human ratings at 98 percent of human interrater agreement for detecting any risk assessment, with average F1 of about 0.86 at call level and 0.66 at statement level, and its authors include four holders of equity in the vendor.[3]

What happened

LyssnCrisis is an AI quality-assurance system that scores crisis-counselor practice on 988 and backup crisis calls, deployed by the software company Lyssn.io with ProtoCall Services, a Portland, Oregon contractor that serves as a SAMHSA-contracted national backup provider for 988 and runs the primary 988 line for New Mexico. The direction is the defining fact and is worth stating plainly: the tool scores the counselor, not the caller. It transcribes call audio and classifies the counselor's own conduct against a hand-coded fidelity scheme — whether a suicide risk assessment was performed and its components, empathy, active listening — and returns fidelity scores, transcripts, and dashboards to counselors and their supervisors within minutes. It never interacts with a caller, issues no caller-facing decision, and takes no automated action.

The problem it targets is a measured coverage gap. 988-participating call centers are required to review 3% of their calls for quality assurance; STAT News reported in June 2023 that ProtoCall, handling more than 1,000 calls a day (about 560,000 in the prior twelve months), was reviewing under 3%, mostly selected at random, so the large majority of crisis-counseling quality went unmeasured. Against that, guidance recommends a suicide risk assessment in every conversation, and Lyssn's own grant announcement cites prior research finding counselors assessed suicide risk in only about 50% of calls — a figure the announcement attributes to research from 2007, which predates 988 entirely and should be read as vendor-cited context rather than a measured ProtoCall baseline. The tool's proposition is to expand measured QA from that thin manual sample toward the full call stream.

The work was funded by a fast-track NIMH Small Business Innovation Research grant (project R44MH133517, PI David C. Atkins), whose three fiscal-year awards total exactly $2,113,800 (FY2023 $275,373, FY2024 $998,118, FY2025 $840,309), with a project end date of January 31, 2026; the vendor's stated "$2.1M" matches this total. The peer-reviewed evidence to date is a reliability study (Psychiatric Services, epub July 2024) that fully labeled 476 calls (193,257 statements) and fine-tuned a transformer-based model; it reported agreement with human ratings at 98% of human interrater agreement for detecting any risk assessment, with average F1 of 0.86 at the call level and 0.66 at the statement level, the lower figure driven by low-base-rate risk labels. That reliability evidence is genuine but vendor-authored: the study's conflict-of-interest statement discloses that four authors hold Lyssn equity and three are cofounders, and a ProtoCall clinician is a co-author; there is no independent replication.

Whether the scores actually change counselor behavior — the effectiveness question — was put to a registered trial. NCT06299384 enrolled 81 call-takers at ProtoCall in a randomized crossover design (a four-week baseline, then twelve weeks of AI-based coding and feedback versus supervision-as-usual, with crossover), ran from April 2, 2024 to October 31, 2025 (primary completion September 5, 2025), and is listed as Completed. Its registered outcomes measure the feedback loop directly: AI-generated fidelity scores on empathy, active listening, and seven key risk-assessment questions, a nine-item caller post-call survey, and implementation measures. As of mid-2026, however, no results are posted to ClinicalTrials.gov, a PubMed sweep finds no results paper, and participant-level data sharing is marked unavailable for proprietary reasons — so every counselor-skill-improvement claim remains, for now, a vendor claim. Lyssn's academic-papers index continued to list the ProtoCall deployment-and-evaluation collaboration as active as of July 2026 (under a "Data Collection" phase label), which corroborates an ongoing collaboration but is a vendor-index inference, not an independent status report.

The people quoted around the tool frame both its promise and its risk. ProtoCall's chief clinical officer, Brad Pendergraft, described the underlying supervision problem the tool addresses — "People can burn out in this work. They can stop doing things that are more emotionally difficult for them" — and doubted SAMHSA would mandate AI QA tools given the burden on smaller providers. A psychologist quoted by STAT called the technology "a potential gamechanger" for identifying under-performing staff, which the dossier is careful to treat as aspirational commentary rather than an evaluated practice. SAMHSA was aware of and supportive of the exploration but did not fund or mandate it; providers obtain caller consent, and the platform is described as HIPAA-compliant with data-removal options. So the honest reading of this case is a split one: strong, peer-reviewed reliability evidence that the AI measures roughly what a human coder would at the call level (noisier per statement), paired with an unpublished, vendor-run effectiveness trial and a governance surround that is, beyond the trial, the ordinary vendor-customer relationship.

The sociotechnical reading

Almost every other place the Atlas meets AI on a crisis line, the model is pointed at the caller: a severity ranker reordering the queue, an e-triage chatbot doing intake. This case is the deliberate inversion, and the inversion is the lesson. Here the model is placed on the supervision-and-quality-assurance link, over the operator network — it scores the counselor's own practice and feeds the result back to the counselor and their supervisor. On the system map that is a specific and, in an important sense, a safer shape: the model does not sit on any client-facing edge, issues no decision, and takes no action; structurally it is a measurement and state-feedback channel on operator behavior, the same family of object as a vigilance control, not a decision channel over vulnerable people. That relocation is a genuine design achievement. It attacks a directly measured constraint — hand review reached under 3% of calls, against a standard of assessing risk in every one — and it was subjected to a registered randomized trial before its effect was claimed, which is more governance discipline than most cells in this library ever saw.

But moving the AI to the safe side of the system does not retire the governance question; it moves it. A measurement placed over the operators is only as trustworthy as the independent check on the measurement itself — and that check is exactly what is missing. The reliability study is real, but it is authored by people who own the tool, and the trial built to test whether the scores are right and whether the feedback actually helps has not published its results. So the tool watches the counselor, and no one independent watches the tool. That is a different failure surface from anything else in the Atlas, and it is second-order: it never touches a caller directly. Its two channels are a miscalibrated fidelity score contaminating the supervision memory the tool trains on (the statement-level agreement is materially noisier than the headline call-level number, and worst on the rare risk labels that matter most), and score-driven performance management — "identifying under-performing staff" — running ahead of the evidence that the scores mean what they are taken to mean. A counselor learns to produce whatever the dashboard rewards; if the dashboard is an uncalibrated instrument, the practice bends toward the instrument's blind spots, and the harder, more emotionally difficult techniques a burned-out counselor drops first are precisely what the checklist can miss.

The distinct lesson the Atlas draws here is about who calibrates the calibrator. The productive move — putting the AI on the supervision link rather than the client-facing one — is real and worth copying, but it comes with an obligation the client-facing cases make obvious and this one makes easy to forget: an instrument you manage people by has to be independently calibrated, on a cadence, by someone who does not own it, before its numbers are allowed to carry weight. The governing levers that matter are therefore not aimed at the scorer's raw accuracy but at the measurement's trustworthiness: keep a human actually reading the scores as coverage climbs (a standing vigilant channel), keep each fidelity number legible as a measured estimate rather than a ground-truth verdict (provenance labeling), require the independent calibration the trial was designed to provide before scaling the scores into how staff are judged (a scheduled challenge to the scores, an oversight cadence on the regime, a vendor gate securing independent test access), and protect the counselor's own skill so the score becomes a mirror rather than a target (deskilling arrest). The honest boundary throughout, and a hard rule for a behavioral-health cell: served callers and the counselors themselves are not modeled here, no suicide, crisis, or clinical outcome is computed from anything in this reading, and every effectiveness figure remains a vendor claim until the trial publishes.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

clinicaltrialsgovusnationall2026GroundingGovernmentSave

ClinicalTrials.gov (U.S. National Library of Medicine), Voice-Based AI to Scale Evaluation of Crisis Counseling in 988 Rollout (NCT06299384) (2026) https://clinicaltrials.gov/study/NCT06299384

https://clinicaltrials.gov/study/NCT06299384

Grounds: model org: lyssn_protocall_988

nihreporternationalinstitute2025GroundingGovernmentSave

NIH RePORTER (National Institutes of Health), Voice-based AI to scale evaluation of crisis counseling in 988 rollout (R44MH133517) (2025) https://reporter.nih.gov/project-details/10983779

https://reporter.nih.gov/project-details/10983779

Grounds: model org: lyssn_protocall_988

imel2024GroundingAcademicSave

Imel, Pace, Pendergraft, Pruett, Tanana, Soma, Comtois, Atkins, Machine Learning-Based Evaluation of Suicide Risk Assessment in Crisis Counseling Calls (Psychiatric Services, 2024;75(11):1068-1074) https://pubmed.ncbi.nlm.nih.gov/39026467/

https://pubmed.ncbi.nlm.nih.gov/39026467/

Grounds: model org: lyssn_protocall_988

lyssn2026GroundingVendorSave

Lyssn.io, Academic Papers: Deployment and evaluation of Lyssn's risk and safety assessment tool at a national crisis and 988 call center (research index, 2026) https://www.lyssn.io/resources/academic-papers/

https://www.lyssn.io/resources/academic-papers/

Grounds: model org: lyssn_protocall_988

Topics: ai-safety

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalAn AI quality-assurance tool deployed on a national 988 backup line scores crisis counselors' own call practic…

An AI quality-assurance tool deployed on a national 988 backup line scores crisis counselors' own call practice rather than callers, expanding measured review from the under-3% of calls that had been reviewed by hand toward nearly all of them; a peer-reviewed reliability study of 476 labeled calls reported agreement with human ratings at 98 percent of human interrater agreement for detecting any risk assessment, with average F1 of about 0.86 at call level and 0.66 at statement level, and its authors include four holders of equity in the vendor.

imel2024GroundingAcademicSave

Imel, Pace, Pendergraft, Pruett, Tanana, Soma, Comtois, Atkins, Machine Learning-Based Evaluation of Suicide Risk Assessment in Crisis Counseling Calls (Psychiatric Services, 2024;75(11):1068-1074) https://pubmed.ncbi.nlm.nih.gov/39026467/

https://pubmed.ncbi.nlm.nih.gov/39026467/

Grounds: model org: lyssn_protocall_988

nihreporternationalinstitute2025GroundingGovernmentSave

NIH RePORTER (National Institutes of Health), Voice-based AI to scale evaluation of crisis counseling in 988 rollout (R44MH133517) (2025) https://reporter.nih.gov/project-details/10983779

https://reporter.nih.gov/project-details/10983779

Grounds: model org: lyssn_protocall_988

EmpiricalThe registered randomized crossover trial of the tool's counselor feedback (81 call-takers) completed on Octob…

The registered randomized crossover trial of the tool's counselor feedback (81 call-takers) completed on October 31, 2025, but as of mid-2026 no results were posted to the trial registry or found in the peer-reviewed literature and participant-level data were marked unavailable for proprietary reasons, so reported counselor-skill-improvement effects remain vendor claims pending independent publication.

clinicaltrialsgovusnationall2026GroundingGovernmentSave

ClinicalTrials.gov (U.S. National Library of Medicine), Voice-Based AI to Scale Evaluation of Crisis Counseling in 988 Rollout (NCT06299384) (2026) https://clinicaltrials.gov/study/NCT06299384

https://clinicaltrials.gov/study/NCT06299384

Grounds: model org: lyssn_protocall_988

lyssn2026GroundingVendorSave

Lyssn.io, Academic Papers: Deployment and evaluation of Lyssn's risk and safety assessment tool at a national crisis and 988 call center (research index, 2026) https://www.lyssn.io/resources/academic-papers/

https://www.lyssn.io/resources/academic-papers/

Grounds: model org: lyssn_protocall_988

Topics: ai-safety