10 to the 23 AI logo

Domain Atlas / Behavioral-health & crisis triage

Case fileUnited States (state Prescription Drug Monitoring Programs)giant deployment

NarxCare

Work with this case in the PAN Lab ↗

NarxCare is a proprietary clinical-decision-support platform built by Bamboo Health that layers over state Prescription Drug Monitoring Programs and returns three Narx Scores plus a composite Overdose Risk Score (each 000-999) into the electronic health record, the PDMP portal, or pharmacy software, often in the patient header alongside vitals and allergies; adoption figures vary by what is counted (more than 40 states and territories run their PDMPs on Bamboo technology and five of the top six pharmacy chains use NarxCare, while the scoring module itself is switched on in more than 20 states). The vendor states the scores are intended to aid, not replace, clinical judgment and should never be sole justification for providing or refusing medication, but clinician and patient-advocacy sources document de facto determinative use — denials, forced tapers, and pharmacy refusals — driven by automation bias and fear of regulatory and criminal liability; patients cannot see, challenge, or correct their scores, the algorithm is proprietary and has not been independently validated for clinical care, and the FDA has not regulated it as a Software-as-a-Medical-Device, so contestation has instead run through FDA citizen petitions (one rejected on procedural grounds in 2023 and a second, docket FDA-2025-P-0701, pending since 2025 with more than 1,000 public comments).[8]

What happened

NarxCare is a proprietary clinical-decision-support platform built by Bamboo Health (formerly Appriss Health) that layers on top of state Prescription Drug Monitoring Programs (PDMPs) and automatically returns three substance-specific Narx Scores plus a composite Overdose Risk Score (ORS) for a patient, delivered into the PDMP web portal, the electronic health record, or pharmacy management software — frequently placed in the patient header alongside vital signs and allergies. Each Narx Score is a three-digit number from 000 to 999: the first two digits express the patient's controlled-substance exposure as a percentile against the rest of that state's PDMP population and the last digit counts active prescriptions, computed as a weighted average of scaled values across four overlapping time windows (recent two months, six months, 180 days, 365 days), with morphine-milligram-equivalents (MME) weighted half of the total. The ORS is a logistic-regression classifier of unintentional overdose death using nine ranked PDMP-derived inputs — total 365-day MME weighted highest at roughly a quarter to a third, then pharmacy and prescriber counts and high-daily-MME prescriptions, with the last three inputs weighted under 1% — originally trained on a case-control set of more than 5,000 autopsy-adjudicated overdose deaths matched to 500,000 controls, drawn from a single Midwestern state over 2013 to 2016. Adoption depends on what is being counted: KFF Health News and the Associated Press reported that more than 40 states and territories use Bamboo Health technology to run their PDMPs and that five of the top six pharmacy chains use NarxCare, while the 2026 npj study describes the NarxCare scoring module as switched on in statewide PDMPs in more than 20 states; a clinician perspective in the Journal of General Internal Medicine described the platform as reaching across 45-plus states and, by a reach estimate the authors flag as such, positioned to influence over a billion patient encounters a year.

The tool's accuracy is contested, and its opacity is why. On its own 2013-2016 training and validation data, Bamboo Health reported an ORS precision of about 75%, recall 57%, specificity 81%, and per-band odds ratios of unintentional overdose death rising from 1.0 (scores 000-199) to 12.4 (500-599) to 29.3 (800-999). Those figures are vendor self-reports and have not been independently reproduced; the vendor's own external-validation set (a different state, 2017-2023) showed sharp degradation, with precision falling to about 52%, which Bamboo attributes to the rise of illicit fentanyl — untracked in PDMPs — and wider use of medications for opioid use disorder. In a 2026 study in npj Digital Medicine, researchers who reconstructed the ORS on California's CURES prescription database (about 17.9 million training observations) and on commercial claims data could not reproduce the vendor's precision, obtaining 0.01 to 0.32 instead across logistic regression, random forest, gradient boosting, and neural-network models. That reconstruction is not a strict like-for-like refutation: because overdose-death labels were not available to the independent researchers, they trained on proxy outcomes — initiation of medication for opioid use disorder in one dataset, opioid-related adverse events in the other — so the result is best read as evidence that proprietary opacity prevents anyone outside the vendor from assessing the model's accuracy, fairness, or safety, rather than as a same-target head-to-head. The through-line, made across the npj paper, the Oliva law-review article, and the clinician perspective, is that NarxCare has been deployed at national scale but never independently validated for clinical care.

The vendor's documentation states more than a dozen times that the scores are "intended to aid, not replace" medical decision-making and should "never" be sole justification for providing or refusing medication. In practice, clinician and patient-advocacy sources document de facto determinative use: the score appears in the patient header next to vitals and allergies, its basis is invisible, and clinicians fearing DEA scrutiny and criminal liability treat a high number as a stop sign — producing forced tapers, patient abandonment, and pharmacy refusals, which themselves can raise overdose risk. Documented harms include a Michigan graduate student whose two dogs' controlled-substance prescriptions, filled under her name at the veterinarian, inflated her profile and contributed to a denial of emergency pain care, and a patient told before an MRI, "Your Narx Score is so high, I can't give you any narcotics." Patients cannot see, challenge, or correct their scores, and NarxCare does not track deprescribing or tapering outcomes, so the downstream harm to a denied patient — including a push toward the illicit market — is invisible to the record the model reads and is never fed back to correct it. Other evidence shows the scores drive real clinical stratification: a 2022 JAAOS Global study of 1,136 knee-replacement patients found higher admission Narx Narcotic Scores independently predicted 30-day readmission (odds ratios 2.46 for scores 300-499 and 3.98 for 500 and above versus opioid-naive patients). Jennifer Oliva's California Law Review article argues that PDMP risk platforms were designed for law-enforcement surveillance, were never validated for clinical care, and likely produce inflated scores for women and for Black, poor, uninsured, and rural patients through proxies such as cash payment and distance traveled — a documented but contested equity concern, not a measured per-subgroup rate; economist Angela Kilby, reconstructing a comparable opioid-risk model, found that predicted risk does not correlate with the individual causal benefit or harm of receiving opioids, an objective-function mismatch. The FDA currently treats NarxCare as clinical decision support outside its Software-as-a-Medical-Device premarket review, even though (as the npj authors note) its own guidance says software providing a risk score for a disease would be a regulated device function. Contestation has run through the regulator: a 2023 citizen petition asking the FDA to declare NarxCare a misbranded device and order a recall was rejected on procedural grounds, and a second citizen petition (regulations.gov docket FDA-2025-P-0701) has been pending since 2025, drawing more than 1,000 public comments.

The sociotechnical reading

Most systems in this Atlas fail at a channel that exists but does not work: a screener who defers, an override that cannot be exercised, an appeal a claimant cannot win. NarxCare is the case where the correction channel does not exist at all, for anyone — and that absence, not the accuracy dispute, is the point. Three doors that governance normally reaches through are each held shut at once. The subject cannot see or contest the score, so the person with the most reason to challenge an error has no way in. No outsider can audit the deployed model, because the algorithm and its engineered features are proprietary — the strongest independent check to date is a simulation reconstruction, not access to the real thing, which is why its striking result (a measured precision of 0.01 to 0.32 against the vendor's self-reported 0.75) documents opacity as much as inaccuracy. And no external authority has validated it, because the FDA declines to regulate a risk score that gates medication as the medical device it functions as. With all three shut, error has no channel to reach any correcting hand: the deployer configures a box it cannot open, the regulator stands down, the courts have not yet bound it, and the only actors left are outside the system entirely — reconstruction researchers and citizen petitioners.

That closure is what makes the rest self-sealing. The harm from a wrong score — a denied patient pushed toward the illicit market — is invisible to the very prescription record the model reads, so the feedback loop that would expose a false positive never closes; the system cannot learn from the deaths it may help cause. Meanwhile the record-to-score loop runs perversely in the other direction: because recovery-treatment medication carries high MME, being in treatment can raise a patient's own overdose-risk score, and as illicit fentanyl (untracked by any PDMP) came to dominate overdose deaths, the model quietly decayed while still reading as authoritative on screen. The map's instruction is that contestability is a precondition every other control needs, and it is bought upstream and around the model, not inside it: procurement terms that force the box open (an audit right, disclosed performance, independent test access), a validation gate that makes the score earn its place before it can steer prescribing, an oversight cadence that re-authorizes it as the world it scores changes, and vigilance that keeps an advisory number from hardening into a determinative one. Improving the score's accuracy is the move that changes least — under opacity you could not even verify the improvement, and none of the three shut doors is an accuracy problem. The honest boundary holds throughout: none of this models overdose, and the patients the score sorts are not in the diagram. The contested accuracy, the documented denials, and the discrimination concern are recorded outside any system map; a score here is an institutional signal, never a person's care.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

wang2026GroundingAcademicSave

Wang, Stofer, Chu, Huang, Li, Algorithmic opacity in opioid risk scoring and the need for transparent AI regulation (npj Digital Medicine, 2026; DOI 10.1038/s41746-026-02491-y) https://www.nature.com/articles/s41746-026-02491-y

https://www.nature.com/articles/s41746-026-02491-y

Grounds: model org: narxcare

Topics: algorithmic-fairness

kilby2021GroundingAcademicSave

Kilby, Algorithmic Fairness in Predicting Opioid Use Disorder using Machine Learning (Northeastern University working paper, 2021) https://angelakilby.com/pdfs/AKilbyFairness_2021-01.pdf

https://angelakilby.com/pdfs/AKilbyFairness_2021-01.pdf

Grounds: model org: narxcare

Topics: algorithmic-fairness

buonora2023GroundingAcademicSave

Buonora, Axson, Cohen, Becker, Paths Forward for Clinicians Amidst the Rise of Unregulated Clinical Decision Support Software: Our Perspective on NarxCare (Journal of General Internal Medicine, 2023) https://pmc.ncbi.nlm.nih.gov/articles/PMC11043299/

https://pmc.ncbi.nlm.nih.gov/articles/PMC11043299/

Grounds: model org: narxcare

admissionnarxcarenarcoticsco2022GroundingAcademicSave

Admission NarxCare Narcotic Scores Are Associated With Increased Odds of Readmission and Prolonged Length of Hospital Stay After Primary Elective Total Knee Arthroplasty (JAAOS Global Research and Reviews, 2022) https://pmc.ncbi.nlm.nih.gov/articles/PMC9726283/

https://pmc.ncbi.nlm.nih.gov/articles/PMC9726283/

Grounds: model org: narxcare

aiincidentdatabaseresponsibl2024aGroundingReferenceSave

AI Incident Database (Responsible AI Collaborative), Incident 172: NarxCare's Risk Score Model Allegedly Lacked Validation and Trained on Data with High Risk of Bias (2024) https://incidentdatabase.ai/cite/172/

https://incidentdatabase.ai/cite/172/

Grounds: model org: narxcare

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalNarxCare is a proprietary clinical-decision-support platform built by Bamboo Health that layers over state Pre…

NarxCare is a proprietary clinical-decision-support platform built by Bamboo Health that layers over state Prescription Drug Monitoring Programs and returns three Narx Scores plus a composite Overdose Risk Score (each 000-999) into the electronic health record, the PDMP portal, or pharmacy software, often in the patient header alongside vitals and allergies; adoption figures vary by what is counted (more than 40 states and territories run their PDMPs on Bamboo technology and five of the top six pharmacy chains use NarxCare, while the scoring module itself is switched on in more than 20 states). The vendor states the scores are intended to aid, not replace, clinical judgment and should never be sole justification for providing or refusing medication, but clinician and patient-advocacy sources document de facto determinative use — denials, forced tapers, and pharmacy refusals — driven by automation bias and fear of regulatory and criminal liability; patients cannot see, challenge, or correct their scores, the algorithm is proprietary and has not been independently validated for clinical care, and the FDA has not regulated it as a Software-as-a-Medical-Device, so contestation has instead run through FDA citizen petitions (one rejected on procedural grounds in 2023 and a second, docket FDA-2025-P-0701, pending since 2025 with more than 1,000 public comments).

wang2026GroundingAcademicSave

Wang, Stofer, Chu, Huang, Li, Algorithmic opacity in opioid risk scoring and the need for transparent AI regulation (npj Digital Medicine, 2026; DOI 10.1038/s41746-026-02491-y) https://www.nature.com/articles/s41746-026-02491-y

https://www.nature.com/articles/s41746-026-02491-y

Grounds: model org: narxcare

Topics: algorithmic-fairness

buonora2023GroundingAcademicSave

Buonora, Axson, Cohen, Becker, Paths Forward for Clinicians Amidst the Rise of Unregulated Clinical Decision Support Software: Our Perspective on NarxCare (Journal of General Internal Medicine, 2023) https://pmc.ncbi.nlm.nih.gov/articles/PMC11043299/

https://pmc.ncbi.nlm.nih.gov/articles/PMC11043299/

Grounds: model org: narxcare

EmpiricalOn its own 2013-2016 training and validation data Bamboo Health reported an Overdose Risk Score precision of a…

On its own 2013-2016 training and validation data Bamboo Health reported an Overdose Risk Score precision of about 75% (self-reported, never independently reproduced), and its own external-validation set from 2017-2023 showed precision falling to about 52%, which the vendor attributed to rising illicit fentanyl (untracked by prescription-monitoring programs) and wider use of opioid-use-disorder treatment medication. A 2026 npj Digital Medicine study that reconstructed the model on California's CURES prescription database (about 17.9 million observations) and on commercial claims data obtained a precision of only 0.01 to 0.32 across several model architectures; because overdose-death labels were unavailable to the independent researchers, that reconstruction was trained on proxy outcomes rather than the score's actual overdose-death target, so it is best read as evidence that proprietary opacity prevents anyone outside the vendor from assessing the deployed model's accuracy, fairness, or safety, rather than as a strict like-for-like refutation of the vendor's figure.

wang2026GroundingAcademicSave

Wang, Stofer, Chu, Huang, Li, Algorithmic opacity in opioid risk scoring and the need for transparent AI regulation (npj Digital Medicine, 2026; DOI 10.1038/s41746-026-02491-y) https://www.nature.com/articles/s41746-026-02491-y

https://www.nature.com/articles/s41746-026-02491-y

Grounds: model org: narxcare

Topics: algorithmic-fairness