PAN Lab: govern the network before failure spreads

Legend

Start here

Challenge of the Day

Challenge result

Round summary

Choose a starting point

Apply institutional pressure

Pull governance levers

Qualitative readout

Edit this network

Explore's workshop — authored parts and pathways, saved in this browser

From Simulation to Lab — Sociotechnical Systems Modeling and Simulation

Modeling evidence and assumptions behind this network

The same AI, hands off: the agentic office

The AI here doesn't just draft casework — it acts on cases, and a stretched staff waves most of it through. Same model as the other two offices. In the published runs — a modeled office, not a real one — errors stick here about 75% of the time, versus 20% and 16% next door. Find the levers that change that.

Network ID:

office-agentic-explore
A narrated tour of the whole Lab — or today's ranked governance exercise.

Description

The same AI, hands off: the agentic office. No pressures applied. No governance levers in place.

Who and what is in the system

  • Drafts decisions and, as an agent, acts on cases with minimal review.
  • A small team supervising many automated actions at once.
  • The shared record system both people and agents read and write.
  • Pulls prior records into the model's context automatically.

How strongly each pathway flows right now

  • strong. Agent outputs adopted with little verification.
  • moderate. Agent writes case records directly.
  • moderate. Hurried prompts frame the model toward confirmation.
  • strong. Adopted outputs documented into the record.
  • moderate. Retrieved records re-enter drafts as fresh context.
  • moderate. Staff read and rely on the record as written.
  • moderate. Shortcuts and adopted claims spread between coworkers.
  • moderate. Agent outputs chained into other agents' context, client context riding along.
  • off; a protective peer check, stronger is better. Second opinions between coworkers — rare under this load.
  • off; a protective peer check, stronger is better. No independent model checks the agents at baseline.

Where the gauges sit

The AI reads — work getting done is against demand. How the AI is helping: ; ; ; .

The failure regime reads . ; ; ; .

Under heavy incoming pressure: , , .

Where each failure mode lives here

The Lab speaks in pathways, pressures, levers, and gauges. This map connects that vocabulary to the formal failure-mode names used in the research grounding it — including a 2026 national survey of 1,179 U.S.-based social workers.

  • Automation bias / overreliancecore

    The “Failures adopted by people or agents” pathway and the operator-deference-drift gauge. Staff turnover and autonomy expansion push it up; the deskilling-arrest lever caps it.

    Evidence: In the same national survey, 40.8% of respondents reported ethical concerns about relying on AI for decision-making, and overreliance on automated decision-making was among the most frequently cited concerns overall.

  • Deskilling / professional judgment erosioncore

    Overreliance in slow motion: the deference gauge drifting upward while correction capacity thins. The deskilling-arrest lever is its deliberate counter-schedule.

  • Sycophancy / agreement-seeking outputcore

    The pushback-heavy-usage pressure runs the “Operator framing biases the model” pathway hot and makes the agreeable answers easier to adopt; the framing-hygiene lever dampens the loop at its origin.

    Evidence: Research on AI sycophancy describes it as a fragmented construct — a family of distinct agreement-seeking behaviors that share a label but differ in form, mechanism, measurement, and required mitigation — and finds it intensifies under user pushback and across multi-turn interaction.

  • Hallucination / incorrect-output propagationcore

    The Lab's core premise: every pathway in the diagram carries incorrect output away from its source, and every lever is a way of governing that propagation rather than assuming a perfect model.

  • Unsafe data flow / privacy & confidentialitycore

    Modeled as pathways, gauged as exposure (Phase NP). Unsanctioned tool use opens a visible egress to the off-network sink; connector sprawl replicates an unverified cache between record systems; case-file-flagged pathways (MiDAS-class enforcement replication, records feeding vendor models) carry the same concern. While any such pathway runs, the Privacy gauge drains — and in Hard and Expert a full gauge is part of the win. “Vet connections” cuts the pathways structurally; “Store less data” shrinks what there is to expose.

    Evidence: In the same national survey, concerns about data privacy and security were the most frequently reported challenge to using AI in practice (46.5% of respondents), and an increased focus on client privacy and confidentiality was the most requested improvement to AI tools for social work (50.4%).

  • Bias propagation / institutional workflow biascore

    Biased framings and contaminated records travel the same workflow pathways failures do — an institutional propagation question, and the documented cases show the workflow (not the model alone) carrying the equity outcome in both directions. The Lab models no demographics: differential harm to served people is recorded outside the network, never computed from its dynamics.

    Evidence: Evaluation evidence on the Allegheny Family Screening Tool found that screener overrides of the tool's recommendations reduced racial disparity in screen-in rates relative to the tool alone.

    Evidence: Independent scrutiny of Rotterdam's welfare-fraud risk model — a 2021 municipal audit followed by a 2023 journalistic investigation that obtained the model itself — documented scores skewed against already-vulnerable groups, and the city suspended the system's use.

  • Transparency / provenance failurecore

    The record-contamination gauge reads how much unlabeled machine content feeds back into decisions; “Mark AI-written records” discounts it and “Gate vendor updates” attacks opacity at procurement.

  • Weak human oversight / safeguard failurecore

    The scenario axis itself: one model, three oversight cultures, three very different outcomes. The correction and authority gauges track it; “Review the riskiest first”, “Pause AI on alarms”, “Require sign-off”, and “Review on schedule” govern it.

    Evidence: In the same national survey, 42.1% of respondents reported having no role in decision-making about AI adoption in their workplace; the report concludes most respondents have limited or no control over how AI technologies are selected or implemented within their organizations.

  • Low AI literacy / verification readinesscore

    The AI-literacy-gap pressure: verification skill (not time) drops and deference rises as trust calibrates on the tool itself. The correction-capacity gauge reads the result. The grounding literature proposes AI literacy as a core professional competency.

    Evidence: The same national survey describes a gap between AI exposure and AI preparedness: 26.6% of respondents cited lack of training or understanding of AI technology as a challenge, 53.4% said training on AI tools and effective use would help, and clear guidelines on the ethical use of AI were the most-endorsed need (66.8%).

    Evidence: AI literacy — the knowledge and skills required to understand, use, and critically evaluate AI systems — has been proposed as a core competency for social work, relevant even to practitioners who never directly use AI tools.

  • AI iatrogenics / governance backfireadvanced

    First-class here: purging records without reading them backfires exactly as the published runs found, and the efficiency readout will call a lever stack counterproductive to its face. The quieter trap — oversight whose gains are bought by rising deference — is why the deskilling-arrest lever exists.

    Evidence: In the published runs, deleting records without reading them raised the contaminated share by stripping out benign entries; only content-aware cleanup reliably reduced it. (PAN governance-lever audit)

  • Monitoring failure / drift blindnessadvanced

    Two pressures carry it: the silent vendor update (drift arriving under controls tuned to old behavior) and monitoring going stale (dashboards nobody must act on — the authority gauge hollows while the regime worsens). “Review on schedule” and “Escalate checks” are the counters.

  • Service starvation / over-throttlingadvanced

    The sixth iatrogenic, symmetric to deference-load: governance so tight the work stops. The “Work getting done” gauge reads strained and the Net AI benefit track sits on the hurting side while the failure regime reads contained — a breaker or write-gate has zeroed an assistive arm, or checking layers have throttled it to a satisficing region where the tool is compliant but not constructive. It is a Lab-only service quantity, distinct from PAN's harm-reduction: you paid budget, deference, and compute for a tool your governance won't let help.

    Evidence: Safety-only alignment establishes a behavioral floor without a ceiling: systems can be 'not-unsafe' yet directionless — compliant without being constructive — and benefit must be assessed as scaffold versus crutch.

    Evidence: In a two-year child-welfare ethnography, an ill-fitting algorithmic tool imposed ongoing repair work on caseworkers — anticipatorily editing the inputs they supplied so the tool would return a usable result, and bending or working around procedure to reconcile its output with the case in front of them — labor spent making a poorly-suited tool usable rather than on the casework itself, distinct from any deliberate checking of the output.

  • Repair work / tool-usability laboradvanced

    The drag the Work-quality component reads when a contaminated or ill-fitting tool imposes ongoing labor to make its output usable — anticipatory editing of inputs, bending procedure to reconcile the result — spent making the tool usable rather than on the casework itself, distinct from the deliberate checking priced as review latency.

    Evidence: In a two-year child-welfare ethnography, an ill-fitting algorithmic tool imposed ongoing repair work on caseworkers — anticipatorily editing the inputs they supplied so the tool would return a usable result, and bending or working around procedure to reconcile its output with the case in front of them — labor spent making a poorly-suited tool usable rather than on the casework itself, distinct from any deliberate checking of the output.

  • Environmental burden (external context)external context

    Deliberately outside the network: nothing in this Lab computes environmental cost, and no gauge claims to. Practitioner concern about AI's environmental impact is documented in the survey evidence and belongs in deployment governance as external context.

    Evidence: In the same national survey's open-ended comments, ethical concerns — prominently including the environmental impact of AI infrastructure — were the most common theme, and the report's first recommendation includes environmental impact among the topics profession-wide ethical guidance should address.

The Lab runs invented cases. Your organization runs on a real one.

We map your actual deployment — its pathways, pressures, and the levers your leadership can pull — through intake, diagnosis, prescription, and monitoring.

Sources & Evidence

Tap to expand

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalIndependent scrutiny of Rotterdam's welfare-fraud risk model — a 2021 municipal audit followed by a 2023 journ…

Independent scrutiny of Rotterdam's welfare-fraud risk model — a 2021 municipal audit followed by a 2023 journalistic investigation that obtained the model itself — documented scores skewed against already-vulnerable groups, and the city suspended the system's use.

rekenkamerrotterdam2021GroundingGovernmentSave

Rekenkamer Rotterdam, Gekleurde technologie: onderzoek naar het gebruik van algoritmes door de gemeente Rotterdam (2021) https://www.rekenkamers.nl/rapport/gekleurde-technologie/

https://www.rekenkamers.nl/rapport/gekleurde-technologie/

Appears in: Evidence reverification (2026)

Grounds: deployment audit: Rotterdam welfare-fraud algorithm

wiredlighthousereports2023GroundingInvestigativeSave

WIRED / Lighthouse Reports, Inside the suspicion machine (2023) https://www.wired.com/story/welfare-state-algorithms/

https://www.wired.com/story/welfare-state-algorithms/

Appears in: PAN framework development

Grounds: deployment audit: Rotterdam welfare-fraud algorithm

EmpiricalEvaluation evidence on the Allegheny Family Screening Tool found that screener overrides of the tool's recomme…

Evaluation evidence on the Allegheny Family Screening Tool found that screener overrides of the tool's recommendations reduced racial disparity in screen-in rates relative to the tool alone.

ScenarioIn the published runs, over a supervised-plus-agent scenario, adding a verifier to the autonomous agent remove…

In the published runs, over a supervised-plus-agent scenario, adding a verifier to the autonomous agent removed roughly 46% of the harm that persists and a coordinated governance package roughly 43%, while upgrading the model alone removed only about 6%.

From the published runs: PAN social-work governance guidance, lever-ranking comparison.

ScenarioIn the published runs, fixing the surrounding system out-leveraged an equal-effort model upgrade in nearly eve…

In the published runs, fixing the surrounding system out-leveraged an equal-effort model upgrade in nearly every case tested, and by several times the margin - a better model helps least where the system, not the model, does the damage.

From the published runs: PAN baseline analysis.

ScenarioIn the published runs, deleting records without reading them raised the contaminated share by stripping out be…

In the published runs, deleting records without reading them raised the contaminated share by stripping out benign entries; only content-aware cleanup reliably reduced it.

From the published runs: PAN governance-lever audit.

ScenarioIn the published runs, the same AI in three modeled office cultures - stylized, not real workplaces - let erro…

In the published runs, the same AI in three modeled office cultures - stylized, not real workplaces - let errors stick at very different rates: roughly 75% under low-oversight autonomy, 20% under human supervision, and 16% under high-governance professional controls.

From the published runs: PAN social-work governance guidance, three-office comparison.

EmpiricalIn a national survey of 1,179 U.S.-based social workers conducted from October 2025 to February 2026 by the Un…

In a national survey of 1,179 U.S.-based social workers conducted from October 2025 to February 2026 by the University of Texas at Austin in collaboration with NASW, 63.5% of respondents reported using AI tools or technologies in their current role.

isbanner2022AcademicSave

Isbanner, S., O'Shaughnessy, P., Steel, D., Wilcock, S., & Carter, S. (2022). The Adoption of Artificial Intelligence in Health Care and Social Services in Australia: Findings From a Methodologically Innovative National Survey of Values and Attitudes (the AVA-AI Study). Journal of Medical Internet Research, 24(8), e37611. https://doi.org/10.2196/37611

doi.org/10.2196/37611

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction, public-benefits

EmpiricalIn the same national survey, concerns about data privacy and security were the most frequently reported challe…

In the same national survey, concerns about data privacy and security were the most frequently reported challenge to using AI in practice (46.5% of respondents), and an increased focus on client privacy and confidentiality was the most requested improvement to AI tools for social work (50.4%).

isbanner2022AcademicSave

Isbanner, S., O'Shaughnessy, P., Steel, D., Wilcock, S., & Carter, S. (2022). The Adoption of Artificial Intelligence in Health Care and Social Services in Australia: Findings From a Methodologically Innovative National Survey of Values and Attitudes (the AVA-AI Study). Journal of Medical Internet Research, 24(8), e37611. https://doi.org/10.2196/37611

doi.org/10.2196/37611

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction, public-benefits

EmpiricalIn the same national survey, 40.8% of respondents reported ethical concerns about relying on AI for decision-m…

In the same national survey, 40.8% of respondents reported ethical concerns about relying on AI for decision-making, and overreliance on automated decision-making was among the most frequently cited concerns overall.

pinazohernandis2026AcademicSave

Pinazo-Hernandis, S., & Carcavilla-Gonzalez, N. (2026). Are future social workers ready for AI? Fears, barriers, and learning needs in higher education. Social Work Education. https://doi.org/10.1080/02615479.2026.2631708

doi.org/10.1080/02615479.2026.2631708

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction, social-work

EmpiricalThe same national survey describes a gap between AI exposure and AI preparedness: 26.6% of respondents cited l…

The same national survey describes a gap between AI exposure and AI preparedness: 26.6% of respondents cited lack of training or understanding of AI technology as a challenge, 53.4% said training on AI tools and effective use would help, and clear guidelines on the ethical use of AI were the most-endorsed need (66.8%).

pinazohernandis2026AcademicSave

Pinazo-Hernandis, S., & Carcavilla-Gonzalez, N. (2026). Are future social workers ready for AI? Fears, barriers, and learning needs in higher education. Social Work Education. https://doi.org/10.1080/02615479.2026.2631708

doi.org/10.1080/02615479.2026.2631708

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction, social-work

EmpiricalIn the same national survey, 42.1% of respondents reported having no role in decision-making about AI adoption…

In the same national survey, 42.1% of respondents reported having no role in decision-making about AI adoption in their workplace; the report concludes most respondents have limited or no control over how AI technologies are selected or implemented within their organizations.

EmpiricalIn the same national survey's open-ended comments, ethical concerns — prominently including the environmental …

In the same national survey's open-ended comments, ethical concerns — prominently including the environmental impact of AI infrastructure — were the most common theme, and the report's first recommendation includes environmental impact among the topics profession-wide ethical guidance should address.

massey2026AcademicSave

Massey, M., Williams, I., Polistina, G., & Breaux, E. (2026). Artificial intelligence and environmental justice: A critical review of social work literature. Society for Social Work and Research Annual Conference. https://sswr.confex.com/sswr/2026/webprogram/Paper62664.html

https://sswr.confex.com/sswr/2026/webprogram/Paper62664.html

Appears in: National survey report (2026)

Topics: social-work

ConceptualAI literacy — the knowledge and skills required to understand, use, and critically evaluate AI systems — has b…

AI literacy — the knowledge and skills required to understand, use, and critically evaluate AI systems — has been proposed as a core competency for social work, relevant even to practitioners who never directly use AI tools.

ahn2025AcademicSave

Ahn, E., Choi, M., Fowler, P., & Song, I. H. (2025). Artificial intelligence (AI) literacy for social work: Implications for core competencies. Journal of the Society for Social Work and Research, 16(1), 9-26. https://doi.org/10.1086/735187

doi.org/10.1086/735187

Appears in: 1023AI authored research; National survey report (2026)

Topics: social-work

EmpiricalResearch on AI sycophancy describes it as a fragmented construct — a family of distinct agreement-seeking beha…

Research on AI sycophancy describes it as a fragmented construct — a family of distinct agreement-seeking behaviors that share a label but differ in form, mechanism, measurement, and required mitigation — and finds it intensifies under user pushback and across multi-turn interaction.

ye2026AcademicSave

Ye, M., Ibrahim, L., Bo, J. Y., et al. (2026). What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.21778

doi.org/10.48550/arXiv.2605.21778

Appears in: PAN framework development

Topics: ai-safety

sharma2024AcademicSave

Sharma, M., Tong, M., Korbak, T., et al. (2024). Towards Understanding Sycophancy in Language Models. In International Conference on Learning Representations (ICLR 2024). https://doi.org/10.48550/arXiv.2310.13548

doi.org/10.48550/arXiv.2310.13548

Appears in: Evidence reverification (2026)

Topics: ai-safety, human-ai-interaction

ConceptualClaims and behaviors spread through peer networks sideways, along informal ties — diffusion research finds wea…

Claims and behaviors spread through peer networks sideways, along informal ties — diffusion research finds weak ties and small-world clustering carry information and practices across a network far faster than formal reporting lines.

granovetter1973AcademicSave

Granovetter, M. S. (1973). The strength of weak ties. American Journal of Sociology, 78(6), 1360–1380. https://doi.org/10.1086/210318

doi.org/10.1086/210318

Appears in: 1023AI authored research

watts1998AcademicSave

Watts, D. J., & Strogatz, S. H. (1998). Collective dynamics of 'small-world' networks. Nature, 393(6684), 440–442. https://doi.org/10.1038/30918

doi.org/10.1038/30918

Appears in: 1023AI authored research

Topics: complexity-science

EmpiricalA single automated rule set applied uniformly and without human review produced tens of thousands of correlate…

A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.

EmpiricalModel behavior — including misaligned behavior — can propagate through model-to-model channels: research shows…

Model behavior — including misaligned behavior — can propagate through model-to-model channels: research shows narrow in-context examples and inter-model interaction can induce broadly misaligned behavior in the receiving model.

afonin2026AcademicSave

Afonin, N., Andriianov, N., Hovhannisyan, V., Bageshpura, N., Liu, K., Zhu, K., Dev, S., Panda, A., Rogov, O., Tutubalina, E., Panchenko, A., & Seleznyov, M. (2026). Emergent misalignment via in-context learning: Narrow in-context examples can produce broadly misaligned LLMs [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2510.11288

doi.org/10.48550/arXiv.2510.11288

Appears in: 1023AI authored research

Topics: ai-alignment, complexity-science

panpatil2025AcademicSave

Panpatil, S., Dingeto, H., & Park, H. (2025). Eliciting and analyzing emergent misalignment in state-of-the-art large language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2508.04196

doi.org/10.48550/arXiv.2508.04196

Appears in: 1023AI authored research

Topics: ai-alignment, complexity-science

betley2026AcademicSave

Betley, J., Warncke, N., Sztyber-Betley, A., Tan, D., Bao, X., Soto, M., Srivastava, M., Labenz, N., & Evans, O. (2026). Training large language models on narrow tasks can lead to broad misalignment. Nature, 649(8097), 584-589. https://doi.org/10.1038/s41586-025-09937-5

doi.org/10.1038/s41586-025-09937-5

Appears in: 1023AI authored research

Topics: ai-alignment

ConceptualEmerging agentic-AI governance frameworks treat inter-agent interaction as a first-class assurance surface, re…

Emerging agentic-AI governance frameworks treat inter-agent interaction as a first-class assurance surface, requiring explicit oversight of agent-to-agent couplings rather than per-model evaluation alone.

khan2025AcademicSave

Khan, R., Joyce, D., & Habiba, M. (2025). AGENTSAFE: A unified framework for ethical assurance and governance in agentic AI [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2512.03180

doi.org/10.48550/arXiv.2512.03180

Appears in: 1023AI authored research

Topics: ai-governance

hammond2025AcademicSave

Hammond, L., Chan, A., Clifton, J., et al. (2025). Multi-Agent Risks from Advanced AI [Technical Report No. 1]. Cooperative AI Foundation. arXiv. https://doi.org/10.48550/arXiv.2502.14143

doi.org/10.48550/arXiv.2502.14143

Appears in: Evidence reverification (2026)

Topics: ai-governance, ai-safety

EmpiricalDocumented benefit-automation failures replicated determinations into downstream systems with no independent r…

Documented benefit-automation failures replicated determinations into downstream systems with no independent reconciliation against the source records — Michigan MiDAS actioned replicated flags and Robodebt reversed the onus onto recipients.

EmpiricalDocumented risk-scoring deployments computed scores from multi-agency administrative records originally collec…

Documented risk-scoring deployments computed scores from multi-agency administrative records originally collected for other purposes, which is the data-protection critique recorded in independent reviews of these systems.

EmpiricalDocumented enforcement systems actioned replicated flags automatically — garnishment and penalties applied bef…

Documented enforcement systems actioned replicated flags automatically — garnishment and penalties applied before any human review step in the recorded MiDAS deployment.

EmpiricalThe Robodebt Royal Commission documented debts raised from income-averaged derived inputs with the onus placed…

The Robodebt Royal Commission documented debts raised from income-averaged derived inputs with the onus placed on recipients to disprove the automated assessments.

EmpiricalProfessional-verification cultures documented in social work practice sustain peer checking of AI output rathe…

Professional-verification cultures documented in social work practice sustain peer checking of AI output rather than unquestioned acceptance.

baez2026AcademicSave

Báez, J. C., Ahn, E., Tamietti, A., Victor, B. G., & Goldkind, L. (2026). Clinical social workers’ perceptions of large language models in practice: Resistance to automation and prospects for integration. Journal of Evidence-Based Social Work, 23(1), 42–63. https://doi.org/10.1080/26408066.2025.2542450

doi.org/10.1080/26408066.2025.2542450

Appears in: National survey report (2026)

Topics: social-work

EmpiricalClinical assessors bound by algorithmic allocation with limited override capacity form a documented constraine…

Clinical assessors bound by algorithmic allocation with limited override capacity form a documented constrained-judgment pattern in home-care assessment.

sutton2020AcademicSave

Sutton, R. T., Pincock, D., Baumgart, D. C., Sadowski, D. C., Fedorak, R. N., & Kroeker, K. I. (2020). An overview of clinical decision support systems: Benefits, risks, and strategies for success. NPJ Digital Medicine, 3(1), 17. https://doi.org/10.1038/s41746-020-0221-y

doi.org/10.1038/s41746-020-0221-y

Appears in: 1023AI authored research

upturnGroundingAdvocacySave

Upturn, Calculated Need: automated home-care hour allocation https://www.upturn.org/work/calculated-need/

https://www.upturn.org/work/calculated-need/

Appears in: PAN framework development

Grounds: deployment audit: Arkansas ARChoices / Idaho Medicaid

EmpiricalRetrieval layers propagate rather than sanitize their inputs: studies find retrieval-augmented systems remain …

Retrieval layers propagate rather than sanitize their inputs: studies find retrieval-augmented systems remain unfaithful even when the retrieved passage is correct, so faithfulness is bounded rather than assured.

faithfulragGroundingPreprintSave

FaithfulRAG (arXiv:2506.08938) — RAG systems struggle in knowledge-conflict scenarios even when relevant passages are retrieved (pessimistic end). https://arxiv.org/abs/2506.08938

https://arxiv.org/abs/2506.08938

Grounds: empirical cap: groundtruth_reliability (max)

faithfulragwithsparseautoencGroundingPreprintSave

Faithful RAG with Sparse Autoencoders (arXiv:2512.08892) — even with relevant passages retrieved, models contradict evidence / invent details; faithfulness is not guaranteed. https://arxiv.org/abs/2512.08892

https://arxiv.org/abs/2512.08892

Grounds: empirical cap: groundtruth_reliability (max)

ragevaluationsurveyGroundingPreprintSave

RAG evaluation survey (arXiv:2405.07437) — factuality evaluation is bounded by knowledge-base coverage and retrieval accuracy; what is checkable depends on what is documented. https://arxiv.org/abs/2405.07437

https://arxiv.org/abs/2405.07437

Grounds: empirical cap: frac_verifiable (max)

EmpiricalRestricting retrieval to a curated, vetted document set bounds what re-enters the model: retrieval-augmented s…

Restricting retrieval to a curated, vetted document set bounds what re-enters the model: retrieval-augmented systems fact-checking against a curated peer-reviewed corpus reach roughly 0.97+ accuracy and factuality evaluation is limited by knowledge-base coverage — what is checkable depends on what is documented — so a vetted corpus reduces contamination drawn back into the model relative to open retrieval, though faithfulness remains imperfect under knowledge conflict.

retrievalaugmentedcovidfactcGroundingPeer-reviewedSave

Retrieval-augmented COVID-19 fact-checking (PMC12079058) — CRAG/Self-RAG reach 0.972-0.978 accuracy against a curated 130k peer-reviewed corpus (optimistic ceiling). https://pmc.ncbi.nlm.nih.gov/articles/PMC12079058/

https://pmc.ncbi.nlm.nih.gov/articles/PMC12079058/

Grounds: empirical cap: groundtruth_reliability (max)

ragevaluationsurveyGroundingPreprintSave

RAG evaluation survey (arXiv:2405.07437) — factuality evaluation is bounded by knowledge-base coverage and retrieval accuracy; what is checkable depends on what is documented. https://arxiv.org/abs/2405.07437

https://arxiv.org/abs/2405.07437

Grounds: empirical cap: frac_verifiable (max)

faithfulragGroundingPreprintSave

FaithfulRAG (arXiv:2506.08938) — RAG systems struggle in knowledge-conflict scenarios even when relevant passages are retrieved (pessimistic end). https://arxiv.org/abs/2506.08938

https://arxiv.org/abs/2506.08938

Grounds: empirical cap: groundtruth_reliability (max)

EmpiricalAutomated output checks are partial, not complete — measured detector-accuracy bands sit well below completene…

Automated output checks are partial, not complete — measured detector-accuracy bands sit well below completeness, especially on hard or adversarial content.

theillusionofprogressGroundingPreprintSave

'The Illusion of Progress' (arXiv:2508.08285) — LLM-as-Judge Precision 0.736 / Recall 0.957 / F1 0.832 vs human consensus on QA. https://arxiv.org/abs/2508.08285

https://arxiv.org/abs/2508.08285

Grounds: empirical cap: catch_at_generation (max)

halogenGroundingPeer-reviewedSave

HALoGEN (arXiv:2501.08292) — best models hallucinate 4%-86% of generated facts depending on domain. https://arxiv.org/abs/2501.08292

https://arxiv.org/abs/2501.08292

Grounds: empirical cap: model_error_base (min)

datadogllmasajudge2025GroundingIndustrySave

Datadog LLM-as-a-judge (2025) — detection F1 drops substantially from HaluBench to the harder RAGTruth; harder hallucinations are harder to catch.

Grounds: empirical cap: catch_at_generation (max)

mentalhealthchatbotdetectionGroundingPreprintSave

Mental-health chatbot detection (arXiv:2604.06216) — GPT judges 54.6% accuracy, 9.3% recall (miss 90.7% of hallucinations); traditional methods F1<0.30 on subjective content. https://arxiv.org/abs/2604.06216

https://arxiv.org/abs/2604.06216

Grounds: empirical cap: catch_at_generation (max)

EmpiricalRanked risk lists steered which cases were investigated in documented deployments; the anchoring direction is …

Ranked risk lists steered which cases were investigated in documented deployments; the anchoring direction is documented while its magnitude is not published.

wiredlighthousereports2023GroundingInvestigativeSave

WIRED / Lighthouse Reports, Inside the suspicion machine (2023) https://www.wired.com/story/welfare-state-algorithms/

https://www.wired.com/story/welfare-state-algorithms/

Appears in: PAN framework development

Grounds: deployment audit: Rotterdam welfare-fraud algorithm

amnestyinternational2021GroundingAdvocacySave

Amnesty International, Xenophobic machines: Discrimination through unregulated use of algorithms in the Dutch childcare benefits scandal (2021) https://www.amnesty.org/en/documents/eur35/4686/2021/en/

https://www.amnesty.org/en/documents/eur35/4686/2021/en/

Appears in: PAN framework development

Grounds: deployment audit: SyRI / childcare-benefits (toeslagenaffaire); model org: netherlands_toeslagen

EmpiricalModel error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors ru…

Model error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors run roughly 1.6–11.6% in frontier evaluations and 4–86% across domains.

xuetal2024GroundingPreprintSave

Xu et al. (2024), 'Hallucination is Inevitable: An Innate Limitation of LLMs' — formal proof that hallucination cannot be eliminated.

Grounds: empirical cap: model_error_base (min)

karpowicz2025GroundingPreprintSave

Karpowicz (2025) — three independent mathematical frameworks (auction theory, proper scoring, log-sum-exp) all conclude no LLM inference mechanism can be simultaneously truthful, etc.

Grounds: empirical cap: model_error_base (min)

halogenGroundingPeer-reviewedSave

HALoGEN (arXiv:2501.08292) — best models hallucinate 4%-86% of generated facts depending on domain. https://arxiv.org/abs/2501.08292

https://arxiv.org/abs/2501.08292

Grounds: empirical cap: model_error_base (min)

openai2025GroundingFrontier labSave

OpenAI (2025), 'Why Language Models Hallucinate' — next-token training plus IDK-penalizing benchmarks push models to bluff; explains the persistent nonzero floor.

Grounds: empirical cap: model_error_base (min)

llmstats2026GroundingIndustry evaluationSave

llm-stats.com failure-focused eval (2026) — FactsGrounding 89.1% accuracy => ~10.9% failure on a relatively easy grounded benchmark.

Grounds: empirical cap: model_error_base (min)

suprmindbenchmarkdigest2026GroundingIndustry evaluationSave

Suprmind benchmark digest (2026) — production ChatGPT ~4.8% major-incorrect with reasoning vs ~11.6% without; HealthBench 3.6%->1.6% with GPT-5 thinking.

Grounds: empirical cap: model_error_base (min)

EmpiricalAutomated catch fractions cap out below completeness — around 84% balanced accuracy in optimistic settings ver…

Automated catch fractions cap out below completeness — around 84% balanced accuracy in optimistic settings versus about 55% on hard content and 9.3% recall in worst-case measurements.

faithfulragleaderboardGroundingPreprintSave

Faithful RAG leaderboard (arXiv:2505.04847) — FaithJudge with o3-mini-high reaches ~84% balanced accuracy / ~82% F1 on FaithBench (optimistic ceiling). https://arxiv.org/abs/2505.04847

https://arxiv.org/abs/2505.04847

Grounds: empirical cap: catch_at_generation (max)

theillusionofprogressGroundingPreprintSave

'The Illusion of Progress' (arXiv:2508.08285) — LLM-as-Judge Precision 0.736 / Recall 0.957 / F1 0.832 vs human consensus on QA. https://arxiv.org/abs/2508.08285

https://arxiv.org/abs/2508.08285

Grounds: empirical cap: catch_at_generation (max)

mentalhealthchatbotdetectionGroundingPreprintSave

Mental-health chatbot detection (arXiv:2604.06216) — GPT judges 54.6% accuracy, 9.3% recall (miss 90.7% of hallucinations); traditional methods F1<0.30 on subjective content. https://arxiv.org/abs/2604.06216

https://arxiv.org/abs/2604.06216

Grounds: empirical cap: catch_at_generation (max)

datadogllmasajudge2025GroundingIndustrySave

Datadog LLM-as-a-judge (2025) — detection F1 drops substantially from HaluBench to the harder RAGTruth; harder hallucinations are harder to catch.

Grounds: empirical cap: catch_at_generation (max)

samedetectionaccuracyliteratGroundingPeer-reviewedSave

Same detection-accuracy literature as catch_at_generation (FaithBench arXiv:2410.13210; arXiv:2508.08285); audit-time detection is bounded by the same hallucination-detection ceiling.

Grounds: empirical cap: catch_at_generation (max); empirical cap: decontaminate (max)

EmpiricalRecord audit-and-correct shares the detection-ceiling family: an optimistic anchor near 96% token accuracy fal…

Record audit-and-correct shares the detection-ceiling family: an optimistic anchor near 96% token accuracy falls away on hard content, so decontamination is bounded rather than total.

halludetectlegaldomainGroundingPreprintSave

HalluDetect legal-domain (arXiv:2509.11619) — best mitigation architecture reaches ~96% token accuracy in a FAVORABLE, retrieval-grounded legal setting (optimistic end). https://arxiv.org/abs/2509.11619

https://arxiv.org/abs/2509.11619

Grounds: empirical cap: decontaminate (max)

samedetectionaccuracyliteratGroundingPeer-reviewedSave

Same detection-accuracy literature as catch_at_generation (FaithBench arXiv:2410.13210; arXiv:2508.08285); audit-time detection is bounded by the same hallucination-detection ceiling.

Grounds: empirical cap: catch_at_generation (max); empirical cap: decontaminate (max)

theillusionofprogressGroundingPreprintSave

'The Illusion of Progress' (arXiv:2508.08285) — LLM-as-Judge Precision 0.736 / Recall 0.957 / F1 0.832 vs human consensus on QA. https://arxiv.org/abs/2508.08285

https://arxiv.org/abs/2508.08285

Grounds: empirical cap: catch_at_generation (max)

EmpiricalThe verification channel itself is bounded — curated-corpus fact-checking tops out around 0.972–0.978 reliabil…

The verification channel itself is bounded — curated-corpus fact-checking tops out around 0.972–0.978 reliability and collapses under knowledge conflict.

retrievalaugmentedcovidfactcGroundingPeer-reviewedSave

Retrieval-augmented COVID-19 fact-checking (PMC12079058) — CRAG/Self-RAG reach 0.972-0.978 accuracy against a curated 130k peer-reviewed corpus (optimistic ceiling). https://pmc.ncbi.nlm.nih.gov/articles/PMC12079058/

https://pmc.ncbi.nlm.nih.gov/articles/PMC12079058/

Grounds: empirical cap: groundtruth_reliability (max)

faithfulragGroundingPreprintSave

FaithfulRAG (arXiv:2506.08938) — RAG systems struggle in knowledge-conflict scenarios even when relevant passages are retrieved (pessimistic end). https://arxiv.org/abs/2506.08938

https://arxiv.org/abs/2506.08938

Grounds: empirical cap: groundtruth_reliability (max)

faithfulragwithsparseautoencGroundingPreprintSave

Faithful RAG with Sparse Autoencoders (arXiv:2512.08892) — even with relevant passages retrieved, models contradict evidence / invent details; faithfulness is not guaranteed. https://arxiv.org/abs/2512.08892

https://arxiv.org/abs/2512.08892

Grounds: empirical cap: groundtruth_reliability (max)

AssumptionThe verifiable fraction of contaminated records is a planning range (0.90/0.60/0.30) that is explicitly calibr…

The verifiable fraction of contaminated records is a planning range (0.90/0.60/0.30) that is explicitly calibration-required and has never been measured.

ragevaluationsurveyGroundingPreprintSave

RAG evaluation survey (arXiv:2405.07437) — factuality evaluation is bounded by knowledge-base coverage and retrieval accuracy; what is checkable depends on what is documented. https://arxiv.org/abs/2405.07437

https://arxiv.org/abs/2405.07437

Grounds: empirical cap: frac_verifiable (max)

EmpiricalIn the documented MiDAS case, error among no-review auto-adjudications ran roughly 93%, and determinations err…

In the documented MiDAS case, error among no-review auto-adjudications ran roughly 93%, and determinations erred at about 85% without human review versus 44% with it.

aiincidentdatabaseGroundingInvestigativeSave

AI Incident Database, Incident 373 (MiDAS false fraud claims) https://incidentdatabase.ai/cite/373/

https://incidentdatabase.ai/cite/373/

Grounds: model org: michigan_midas

EmpiricalIn the documented AFST evaluation, screener overrides of the tool — roughly a third of its recommendations — c…

In the documented AFST evaluation, screener overrides of the tool — roughly a third of its recommendations — cut screen-in disparity from about 20% to 9% relative to the tool acting alone.

stapletonetal2022GroundingPeer-reviewedSave

Stapleton et al., Imagining new futures beyond predictive systems in child welfare (FAccT 2022) https://dl.acm.org/doi/10.1145/3531146.3533177

https://dl.acm.org/doi/10.1145/3531146.3533177

Appears in: PAN framework development

Grounds: deployment audit: Allegheny AFST

Topics: child-welfare

EmpiricalProfessional caseload standards published by the Child Welfare League of America recommend no more than about …

Professional caseload standards published by the Child Welfare League of America recommend no more than about 15 families per worker, sitting well below documented practice loads.

childrenandfamilyresearchcen2002GroundingAcademicSave

Children and Family Research Center (University of Illinois at Urbana-Champaign), Caseload Size in Best Practice: A Literature Review (2002) https://cfrc.illinois.edu/pubs/bf_20021101_CaseloadSizeInBestPractice.pdf

https://cfrc.illinois.edu/pubs/bf_20021101_CaseloadSizeInBestPractice.pdf

Appears in: Evidence reverification (2026)

Grounds: workforce data: caseload standards

academyforprofessionalexcell2021GroundingAcademicSave

Academy for Professional Excellence / CWDS (San Diego State University), Research Summary: Caseload Standards and Weighting Methodologies (2021) https://theacademy.sdsu.edu/wp-content/uploads/2021/10/CWDS-Research-Summary_Caseload-Standards-and-Weighting.pdf

https://theacademy.sdsu.edu/wp-content/uploads/2021/10/CWDS-Research-Summary_Caseload-Standards-and-Weighting.pdf

Appears in: Evidence reverification (2026)

Grounds: workforce data: caseload standards

EmpiricalDocumentation and administrative tasks consume roughly half of practitioner time: a nationally representative …

Documentation and administrative tasks consume roughly half of practitioner time: a nationally representative US child-welfare workforce snapshot found caseworkers spend about 54% of the workday (4.3 of 8 hours) on paperwork and documentation, and a UK children's-services review reports staff spending over 50% of their time on case recording, paperwork, and related tasks.

opre2025GroundingGovernmentSave

OPRE, Snapshot of the Child Welfare Workforce from 2021 to 2022: Caseworker Experiences Working in the Child Welfare System, OPRE Report 2025-040 (2025) https://acf.gov/opre/report/snapshot-child-welfare-workforce-2021-2022-caseworker-experiences-working-child-welfare

https://acf.gov/opre/report/snapshot-child-welfare-workforce-2021-2022-caseworker-experiences-working-child-welfare

Appears in: PAN framework development

Grounds: workforce data: documentation time share

Topics: child-welfare

burbidge2022GroundingGovernmentSave

Burbidge, I. (2022). Report sets out new blueprint for councils to deliver a reshaped children's services. County Councils Network https://www.countycouncilsnetwork.org.uk/report-sets-out-new-blueprint-for-councils-to-deliver-a-reshaped-childrens-services/

https://www.countycouncilsnetwork.org.uk/report-sets-out-new-blueprint-for-councils-to-deliver-a-reshaped-childrens-services/

Appears in: PAN framework development

Grounds: workforce data: documentation time share

Topics: complexity-science

EmpiricalTiered HIPAA penalties run from $145 to $73,011 per violation with an annual cap near $2.19M (2025-adjusted), …

Tiered HIPAA penalties run from $145 to $73,011 per violation with an annual cap near $2.19M (2025-adjusted), and disclosure to a tool that is not a business associate is itself a violation.

hipaajournal2026bAcademicSave

HIPAA Journal. (2026). HIPAA violation penalties. HIPAA Journal.

Appears in: 1023AI authored research

Topics: privacy-security

EmpiricalCited per-unit intensities of roughly 0.3 Wh per inference call and about 3.14 L of water per kWh are applied …

Cited per-unit intensities of roughly 0.3 Wh per inference call and about 3.14 L of water per kWh are applied to authored illustrative volumes rather than to measured deployment totals.

jegham2025AcademicSave

Jegham, N., et al. (2025). How Hungry is AI? Benchmarking energy, water, and carbon footprint of LLM inference [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2505.09598

doi.org/10.48550/arXiv.2505.09598

Appears in: PAN framework development

Topics: ai-governance

li2023AcademicSave

Li, P., Yang, J., Islam, M. A., & Ren, S. (2023). Making AI Less Thirsty: Uncovering and addressing the secret water footprint of AI models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2304.03271

doi.org/10.48550/arXiv.2304.03271

Appears in: PAN framework development

Topics: ai-governance

EmpiricalA preprint benchmark reports an in-context misalignment dose-response: in the most susceptible frontier model,…

A preprint benchmark reports an in-context misalignment dose-response: in the most susceptible frontier model, up to ~24% misaligned behavior at 16 examples rising to ~58% at 256 examples (rates at 16 examples span roughly 1–24% across models), with the majority of misaligned responses rationalized.

afonin2026AcademicSave

Afonin, N., Andriianov, N., Hovhannisyan, V., Bageshpura, N., Liu, K., Zhu, K., Dev, S., Panda, A., Rogov, O., Tutubalina, E., Panchenko, A., & Seleznyov, M. (2026). Emergent misalignment via in-context learning: Narrow in-context examples can produce broadly misaligned LLMs [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2510.11288

doi.org/10.48550/arXiv.2510.11288

Appears in: 1023AI authored research

Topics: ai-alignment, complexity-science

EmpiricalModel behavior drifts discontinuously between evaluation snapshots, and narrow finetuning can induce broad cor…

Model behavior drifts discontinuously between evaluation snapshots, and narrow finetuning can induce broad correlated failure across unrelated tasks.

betley2026AcademicSave

Betley, J., Warncke, N., Sztyber-Betley, A., Tan, D., Bao, X., Soto, M., Srivastava, M., Labenz, N., & Evans, O. (2026). Training large language models on narrow tasks can lead to broad misalignment. Nature, 649(8097), 584-589. https://doi.org/10.1038/s41586-025-09937-5

doi.org/10.1038/s41586-025-09937-5

Appears in: 1023AI authored research

Topics: ai-alignment

li2026AcademicSave

Li, Z., Fan, C., & Zhou, T. (2026). Grokking in LLM pretraining? Monitor memorization-to-generalization without test [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2506.21551

doi.org/10.48550/arXiv.2506.21551

Appears in: 1023AI authored research

song2026AcademicSave

Song, P., Han, P., & Goodman, N. (2026). Large language model reasoning failures [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2602.06176

doi.org/10.48550/arXiv.2602.06176

Appears in: 1023AI authored research

anwar2024AcademicSave

Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., Lubana, E., Jenner, E., Casper, S., Sourbut, O., Edelman, B. L., Zhang, Z., Gunther, M., Korinek, A., Hernandez-Orallo, J., Hammond, L., Bigelow, E., Pan, A., Langosco, L., Korbak, T., Zhang, H., Zhong, R., O Heigeartaigh, S., Recchia, G., Corsi, G., Chan, A., Anderljung, M., Edwards, L., Petrov, A., de Witt, C. S., Motwani, S. R., Bengio, Y., Chen, D., Torr, P. H. S., Albanie, S., Maharaj, T., Foerster, J., Tramer, F., He, H., Kasirzadeh, A., Choi, Y., & Krueger, D. (2024). Foundational challenges in assuring alignment and safety of large language models. Transactions on Machine Learning Research. https://doi.org/10.48550/arXiv.2404.09932

doi.org/10.48550/arXiv.2404.09932

Appears in: 1023AI authored research

Topics: ai-alignment, ai-safety

nikolaou2025AcademicSave

Nikolaou, K., Krippendorf, S., Tovey, S., & Holm, C. (2025). Beyond scaling curves: Internal dynamics of neural networks through the NTK lens [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.05035

doi.org/10.48550/arXiv.2507.05035

Appears in: 1023AI authored research

Topics: complexity-science

EmpiricalIn contextual inquiries with Allegheny AFST call screeners, workers calibrated reliance using contextual case …

In contextual inquiries with Allegheny AFST call screeners, workers calibrated reliance using contextual case knowledge unavailable to the model and reliably detected and overrode erroneous risk scores — complementary human information, not generic distrust, was the safeguard's mechanism.

kawakami2022AcademicSave

Kawakami, A., Sivaraman, V., Cheng, H.-F., Stapleton, L., Cheng, Y., Qing, D., Perer, A., Wu, Z. S., Zhu, H., & Holstein, K. (2022). Improving Human-AI Partnerships in Child Welfare: Understanding Worker Practices, Challenges, and Desires for Algorithmic Decision Support. In CHI Conference on Human Factors in Computing Systems (CHI '22). ACM. https://doi.org/10.1145/3491102.3517439

doi.org/10.1145/3491102.3517439

Appears in: PAN framework development

Topics: algorithmic-fairness, child-welfare, human-ai-interaction

dearteaga2020AcademicSave

De-Arteaga, M., Fogliato, R., & Chouldechova, A. (2020). A Case for Humans-in-the-Loop: Decisions in the Presence of Erroneous Algorithmic Scores. In CHI Conference on Human Factors in Computing Systems (CHI 2020). ACM. https://doi.org/10.1145/3313831.3376638

doi.org/10.1145/3313831.3376638

Appears in: Evidence reverification (2026)

Topics: algorithmic-fairness, child-welfare, human-ai-interaction

EmpiricalAFST workers reported sometimes agreeing with the risk score against their own best judgment under override-ra…

AFST workers reported sometimes agreeing with the risk score against their own best judgment under override-rate oversight, and becoming less likely to disagree over time — reliance driven by organizational incentives independent of trust in the tool.

kawakami2022AcademicSave

Kawakami, A., Sivaraman, V., Cheng, H.-F., Stapleton, L., Cheng, Y., Qing, D., Perer, A., Wu, Z. S., Zhu, H., & Holstein, K. (2022). Improving Human-AI Partnerships in Child Welfare: Understanding Worker Practices, Challenges, and Desires for Algorithmic Decision Support. In CHI Conference on Human Factors in Computing Systems (CHI '22). ACM. https://doi.org/10.1145/3491102.3517439

doi.org/10.1145/3491102.3517439

Appears in: PAN framework development

Topics: algorithmic-fairness, child-welfare, human-ai-interaction

kawakami2026AcademicSave

Kawakami, A., Taylor, J., Fox, S., Zhu, H., & Holstein, K. (2026). AI failure loops in devalued work: The confluence of overconfidence in AI and underconfidence in worker expertise. Big Data & Society. https://doi.org/10.1177/20539517261424164

doi.org/10.1177/20539517261424164

Appears in: Evidence reverification (2026)

Topics: child-welfare, human-ai-interaction

EmpiricalIn a two-year child-welfare ethnography, a re-purposed assessment algorithm produced process-oriented harms to…

In a two-year child-welfare ethnography, a re-purposed assessment algorithm produced process-oriented harms to practice, organization, and street-level decisions, compelling caseworkers to perform added repair work; 80% of interviewees reported that the tool had stripped their decision-making discretion.

saxena2024AcademicSave

Saxena, D., & Guha, S. (2024). Algorithmic Harms in Child Welfare: Uncertainties in Practice, Organization, and Street-level Decision-making. ACM Journal on Responsible Computing, 1(1), 1–32. https://doi.org/10.1145/3616473

doi.org/10.1145/3616473

Appears in: 1023AI authored research

Topics: algorithmic-fairness, child-welfare

ammitzbollflugge2021AcademicSave

Ammitzboll Flugge, A., Hildebrandt, T., & Holten Moller, N. (2021). Street-Level Algorithms and AI in Bureaucratic Decision-Making: A Caseworker Perspective. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 40. https://doi.org/10.1145/3449114

doi.org/10.1145/3449114

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction, public-benefits

EmpiricalThe same agency's theory-driven 7ei tool — which tracks case trajectories instead of predicting outcomes — ear…

The same agency's theory-driven 7ei tool — which tracks case trajectories instead of predicting outcomes — earned collective buy-in and better engagement, but required sustained investments: trauma-informed training, specialized supervision and expert consultation, and new collaborative staffings.

saxena2024AcademicSave

Saxena, D., & Guha, S. (2024). Algorithmic Harms in Child Welfare: Uncertainties in Practice, Organization, and Street-level Decision-making. ACM Journal on Responsible Computing, 1(1), 1–32. https://doi.org/10.1145/3616473

doi.org/10.1145/3616473

Appears in: 1023AI authored research

Topics: algorithmic-fairness, child-welfare

EmpiricalIn a participatory-design study (CHI Late-Breaking Work) with 51 social-service practitioners across two stage…

In a participatory-design study (CHI Late-Breaking Work) with 51 social-service practitioners across two stages (27 in co-design workshops, 24 in contextual inquiry), AI value concentrated in documentation relief, assessment brainstorming, guidance for junior workers, and supervision support — with deskilling and privacy concerns voiced inside the same sessions.

tan2025AcademicSave

Tan, Y., Soh, K. X., Zhang, R., Lee, J., Meng, H., Sen, B., & Lee, Y.-C. (2025). Empowering Social Service with AI: Insights from a Participatory Design Study with Practitioners. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA '25). ACM. https://doi.org/10.1145/3706599.3719736

doi.org/10.1145/3706599.3719736

Appears in: PAN framework development

Topics: co-design, human-ai-interaction, social-work

ConceptualSafety-only alignment establishes a behavioral floor without a ceiling: systems can be 'not-unsafe' yet direct…

Safety-only alignment establishes a behavioral floor without a ceiling: systems can be 'not-unsafe' yet directionless — compliant without being constructive — and benefit must be assessed as scaffold versus crutch.

laukkonen2026AcademicSave

Laukkonen, R., Krier, S., Bakalar, C., Chandaria, S., Kringelbach, M., Elwood, A., Ford, D., Rosas, F., Bohacek, M., Franklin, M., Tomašev, N., Chan, S., Rieser, V., Patel, R., Levin, M., & Rao, A. (2026). Positive Alignment: Artificial Intelligence for Human Flourishing [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.10310

doi.org/10.48550/arXiv.2605.10310

Appears in: PAN framework development

Topics: ai-alignment, ai-governance, ai-safety

ConceptualFormally, estimation error shrinks with data while the human perception gap that produces stationary-environme…

Formally, estimation error shrinks with data while the human perception gap that produces stationary-environment black swans has a non-zero lower bound — so an incident-free operating history yields confidence without safety.

lee2025AcademicSave

Lee, H., Park, C., Abel, D., & Jin, M. (2025). A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety. In International Conference on Learning Representations (ICLR 2025). https://doi.org/10.48550/arXiv.2407.18422

doi.org/10.48550/arXiv.2407.18422

Appears in: PAN framework development

Topics: ai-safety, antifragility

ConceptualStatic robustness certification lags emergent threats; organizations that fold each stressor into their model …

Static robustness certification lags emergent threats; organizations that fold each stressor into their model (slow-loop updates, periodic reviews, post-deployment feedback) shrink future risk, while patch-and-pray accumulates it — the fragility trap.

jin2025AcademicSave

Jin, M., & Lee, H. (2025). Position: AI safety must embrace an antifragile perspective [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2509.13339

doi.org/10.48550/arXiv.2509.13339

Appears in: 1023AI authored research

Topics: ai-safety, antifragility

EmpiricalIn formal simulation, even ideal Bayesian users spiral to near-certain false beliefs under a sycophantic inter…

In formal simulation, even ideal Bayesian users spiral to near-certain false beliefs under a sycophantic interlocutor at sycophancy rates measured in frontier models (~50-70%), and truth-constrained cherry-picking still produces spirals — minimizing hallucination alone is insufficient.

chandra2026AcademicSave

Chandra, K., Kleiman-Weiner, M., Ragan-Kelley, J., & Tenenbaum, J. B. (2026). Sycophantic chatbots cause delusional spiraling, even in ideal Bayesians [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2602.19141

doi.org/10.48550/arXiv.2602.19141

Appears in: 1023AI authored research

Topics: ai-safety

sharma2024AcademicSave

Sharma, M., Tong, M., Korbak, T., et al. (2024). Towards Understanding Sycophancy in Language Models. In International Conference on Learning Representations (ICLR 2024). https://doi.org/10.48550/arXiv.2310.13548

doi.org/10.48550/arXiv.2310.13548

Appears in: Evidence reverification (2026)

Topics: ai-safety, human-ai-interaction

EmpiricalIn a four-week randomized study (n=981), voluntary daily chatbot usage duration predicted worse outcomes on lo…

In a four-week randomized study (n=981), voluntary daily chatbot usage duration predicted worse outcomes on loneliness, socialization, emotional dependence, and problematic use across all conditions, and task-style use fostered practical dependence — reduced confidence in independent judgment.

fang2025AcademicSave

Fang, C. M., Liu, A. R., Danry, V., Lee, E., Chan, S. W. T., Pataranutaporn, P., Maes, P., Phang, J., Lampe, M., Ahmad, L., & Agarwal, S. (2025). How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2503.17473

doi.org/10.48550/arXiv.2503.17473

Appears in: PAN framework development

Topics: ai-safety, human-ai-interaction

gerlich2025AcademicSave

Gerlich, M. (2025). AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies, 15(1), 6. https://doi.org/10.3390/soc15010006

doi.org/10.3390/soc15010006

Appears in: Evidence reverification (2026)

Topics: ai-safety, human-ai-interaction

EmpiricalA validated collaborative-AI metacognition scale (planning, monitoring, evaluation of one's own reliance) pred…

A validated collaborative-AI metacognition scale (planning, monitoring, evaluation of one's own reliance) predicted collaboration benefits incrementally beyond general metacognition — verification-skill training, not generic AI knowledge, is the calibrated counter to over-reliance.

sidra2025AcademicSave

Sidra, & Mason, C. (2025). Generative AI in Human-AI Collaboration: Validation of the Collaborative AI Literacy and Collaborative AI Metacognition Scales for Effective Use. International Journal of Human-Computer Interaction. https://doi.org/10.1080/10447318.2025.2543997

doi.org/10.1080/10447318.2025.2543997

Appears in: Evidence reverification (2026)

Topics: ai-safety, human-ai-interaction

bucinca2021AcademicSave

Bucinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188. https://doi.org/10.1145/3449287

doi.org/10.1145/3449287

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction

EmpiricalGiven only a covert persuasion goal and an explicit no-deception instruction, a frontier model still produced …

Given only a covert persuasion goal and an explicit no-deception instruction, a frontier model still produced manipulative cues in 8.8% of turns, and cue frequency did not reliably predict manipulative success — while automated detection of such cues is itself bounded.

akbulut2026AcademicSave

Akbulut, C., Elasmar, R., Roy, A., Payne, A., Suresh, P., Ibrahim, L., El-Sayed, S., Rastogi, C., Kachra, A., Hawkins, W., Lum, K., & Weidinger, L. (2026). Evaluating language models for harmful manipulation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2603.25326

doi.org/10.48550/arXiv.2603.25326

Appears in: 1023AI authored research

Topics: ai-safety

EmpiricalA predictive system whose outputs shape its own future inputs holds a structural incentive to make the populat…

A predictive system whose outputs shape its own future inputs holds a structural incentive to make the population easier to predict; ordinary pipeline choices can reveal this hidden incentive without any change to the stated objective, and feedback-loop risk tends to grow with model capability.

krueger2020AcademicSave

Krueger, D., Maharaj, T., & Leike, J. (2020). Hidden Incentives for Auto-induced Distributional Shift. In International Conference on Machine Learning (ICML 2020). https://doi.org/10.48550/arXiv.2009.09153

doi.org/10.48550/arXiv.2009.09153

Appears in: PAN framework development

Topics: ai-safety, algorithmic-fairness

perdomo2020AcademicSave

Perdomo, J. C., Zrnic, T., Mendler-Dunner, C., & Hardt, M. (2020). Performative Prediction. In International Conference on Machine Learning (ICML 2020), PMLR 119:7599-7609. https://doi.org/10.48550/arXiv.2002.06673

doi.org/10.48550/arXiv.2002.06673

Appears in: Evidence reverification (2026)

Topics: ai-safety

ScenarioAI documentation assistance can cut clinician documentation burden substantially, but the efficiency paradox c…

AI documentation assistance can cut clinician documentation burden substantially, but the efficiency paradox converts freed time into added caseload unless organizational policy protects it — time returned is realized as benefit only when governance decides where the dividend goes.

vanhara2026AcademicSave

VanHara, A., & Hage, D. (2026). Unintended Ramifications of AI-Assisted Documentation: Navigating Pragmatic & Ethical Clinical Social Work Workload Challenges. Journal of Evidence-Based Social Work, 23(1), 64-77. https://doi.org/10.1080/26408066.2025.2571439

doi.org/10.1080/26408066.2025.2571439

Appears in: PAN framework development

Topics: ai-governance, social-work

EmpiricalAcross 1.5 million real assistant conversations, sycophantic validation — not fabrication — dominated reality-…

Across 1.5 million real assistant conversations, sycophantic validation — not fabrication — dominated reality-distortion risk; disempowering interactions received higher user satisfaction than baseline, making satisfaction a biased proxy that rewards deference.

sharma2026AcademicSave

Sharma, M., McCain, M., Douglas, R., & Duvenaud, D. (2026). Who's in Charge? Disempowerment Patterns in Real-World LLM Usage. In International Conference on Machine Learning (ICML 2026). https://doi.org/10.48550/arXiv.2601.19062

doi.org/10.48550/arXiv.2601.19062

Appears in: PAN framework development

Topics: ai-safety, human-ai-interaction

ConceptualFormally, a system benefits from volatility when its response to a stressor is convex (Jensen's inequality: th…

Formally, a system benefits from volatility when its response to a stressor is convex (Jensen's inequality: the expected outcome under variability exceeds the outcome at the average), and is harmed when the response is concave — so whether a shock strengthens or weakens an organization depends on the curvature of its response, a bounded local property that fails beyond a defined stress range.

axenie2024AcademicSave

Axenie, C., Lopez-Corona, O., Makridis, M. A., Akbarzadeh, M., Saveriano, M., Stancu, A., & West, J. (2024). Antifragility in complex dynamical systems. npj Complexity, 1, 12. https://doi.org/10.1038/s44260-024-00014-y

doi.org/10.1038/s44260-024-00014-y

Appears in: PAN framework development

Topics: ai-governance, antifragility, complexity-science

taleb2013AcademicSave

Taleb, N. N., & Douady, R. (2013). Mathematical definition, mapping, and detection of (anti)fragility. Quantitative Finance, 13(11), 1677-1689. https://doi.org/10.1080/14697688.2013.800219

doi.org/10.1080/14697688.2013.800219

Appears in: Evidence reverification (2026)

Topics: ai-governance

ConceptualRepeatable behaviors follow a dose-response curve — beneficial at low frequency or count and harmful past a ho…

Repeatable behaviors follow a dose-response curve — beneficial at low frequency or count and harmful past a hormetic limit (the dose beyond which net utility turns negative) — because a fast benefit process is followed by a slower accumulating opposing process, giving AI assistance an optimal bounded dose rather than a monotonic benefit.

henry2025AcademicSave

Henry, N. I. N., Pedersen, M., Williams, M., Martin, J. L. B., & Donkin, L. (2025). A Hormetic Approach to the Value-Loading Problem: Preventing the Paperclip Apocalypse. SN Computer Science, 6, 872. https://doi.org/10.1007/s42979-025-04369-4

doi.org/10.1007/s42979-025-04369-4

Appears in: PAN framework development

Topics: ai-safety

calabrese2002AcademicSave

Calabrese, E. J., & Baldwin, L. A. (2002). Defining hormesis. Human & Experimental Toxicology, 21(2), 91-97. https://doi.org/10.1191/0960327102ht217oa

doi.org/10.1191/0960327102ht217oa

Appears in: Evidence reverification (2026)

Topics: ai-governance

EmpiricalA survey of generative-AI safety evaluations found 85.6% operate at the model-capability layer, only 5.3% at t…

A survey of generative-AI safety evaluations found 85.6% operate at the model-capability layer, only 5.3% at the human-interaction layer and 9.1% at the systemic-impact layer — yet context determines whether a capability becomes harm, so the human and system layers where risk actually manifests are the least evaluated.

weidinger2023AcademicSave

Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., & Isaac, W. (2023). Sociotechnical safety evaluation of generative AI systems [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2310.11986

doi.org/10.48550/arXiv.2310.11986

Appears in: PAN framework development

Topics: ai-governance, ai-safety, sociotechnical-evaluation

EmpiricalA frontier risk-management framework in practice ties deployment authority to measured capability-vs-safety zo…

A frontier risk-management framework in practice ties deployment authority to measured capability-vs-safety zones — green (routine plus monitoring), yellow (controlled with strengthened mitigations), and red (suspend).

shanghaiartificialintelligen2025AcademicSave

Shanghai Artificial Intelligence Laboratory. (2025). Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.16534

doi.org/10.48550/arXiv.2507.16534

Appears in: PAN framework development

Topics: ai-governance, ai-safety

greenblatt2024AcademicSave

Greenblatt, R., Denison, C., Wright, B., et al. (2024). Alignment Faking in Large Language Models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.14093

doi.org/10.48550/arXiv.2412.14093

Appears in: Evidence reverification (2026)

Topics: ai-alignment, ai-safety

meinke2024AcademicSave

Meinke, A., Schoen, B., Scheurer, J., Balesni, M., Shah, R., & Hobbhahn, M. (2024). Frontier Models are Capable of In-context Scheming [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.04984

doi.org/10.48550/arXiv.2412.04984

Appears in: Evidence reverification (2026)

Topics: ai-safety

EmpiricalIn frontier-model testing, some systems behaved measurably safer when they believed they were monitored than w…

In frontier-model testing, some systems behaved measurably safer when they believed they were monitored than when unmonitored, and exhibited strategic dishonesty or underperformance under pressure — so ‘behaves well under monitoring’ is insufficient evidence of safety, arguing for unpredictable continuous oversight.

shanghaiartificialintelligen2025AcademicSave

Shanghai Artificial Intelligence Laboratory. (2025). Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.16534

doi.org/10.48550/arXiv.2507.16534

Appears in: PAN framework development

Topics: ai-governance, ai-safety

greenblatt2024AcademicSave

Greenblatt, R., Denison, C., Wright, B., et al. (2024). Alignment Faking in Large Language Models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.14093

doi.org/10.48550/arXiv.2412.14093

Appears in: Evidence reverification (2026)

Topics: ai-alignment, ai-safety

meinke2024AcademicSave

Meinke, A., Schoen, B., Scheurer, J., Balesni, M., Shah, R., & Hobbhahn, M. (2024). Frontier Models are Capable of In-context Scheming [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.04984

doi.org/10.48550/arXiv.2412.04984

Appears in: Evidence reverification (2026)

Topics: ai-safety

ConceptualAI-safety failure classification has a missing interaction layer between institutional risk categories and sys…

AI-safety failure classification has a missing interaction layer between institutional risk categories and system-level failure modes: practitioners lack a shared vocabulary of recognizable error patterns, and the catch-all 'hallucination' collapses distinct logic failures whose correct fixes differ.

beyer2026AcademicSave

Beyer, C. (2026). Toward a Common Language for Human-AI Interaction Failures: A Practitioner-Accessible Error Taxonomy for the Missing Layer of AI Safety Classification [Working paper].

Appears in: PAN framework development

Topics: ai-safety, human-ai-interaction

EmpiricalA data-driven taxonomy built from 9,705 real AI-incident reports found mitigation practice dominated by reacti…

A data-driven taxonomy built from 9,705 real AI-incident reports found mitigation practice dominated by reactive and legal levers (incident investigation, reporting, regulatory and court action) while proactive technical and governance levers (model alignment, safety frameworks, board oversight) were least common — real organizations respond after harm rather than preventing it.

popchanovska2026AcademicSave

Popchanovska, E., Gjorgjevikj, A., Rizinski, M., Chitkushev, L. T., Vodenska, I., & Trajanov, D. (2026). When AI Fails, What Works? A Data-Driven Taxonomy of Real-World AI Risk Mitigation Strategies [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2603.04259

doi.org/10.48550/arXiv.2603.04259

Appears in: PAN framework development

Topics: ai-governance, ai-safety

slattery2024AcademicSave

Slattery, P., Saeri, A. K., Grundy, E. A. C., et al. (2024). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2408.12622

doi.org/10.48550/arXiv.2408.12622

Appears in: Evidence reverification (2026)

Topics: ai-governance, ai-safety

EmpiricalAcross five preregistered studies (N=3,075), sycophantic AI delivered the emotional and esteem support people …

Across five preregistered studies (N=3,075), sycophantic AI delivered the emotional and esteem support people most associate with close relationships, narrowing the felt-understanding gap between AI and humans and leaving people less satisfied with real human interaction over three weeks — and offering users a choice of interaction styles did not reduce their preference for the sycophantic one.

ibrahim2026AcademicSave

Ibrahim, L., Hafner, F. S., Cheng, M., Lee, C., Anselmetti, R., Willer, R., Rocher, L., & Yang, D. (2026). Sycophantic AI makes human interaction feel more effortful and less satisfying over time [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.07912

doi.org/10.48550/arXiv.2605.07912

Appears in: PAN framework development

Topics: ai-safety, human-ai-interaction

cheng2026AcademicSave

Cheng, M., Yu, S., Lee, K., Khadpe, P., Ibrahim, L., & Jurafsky, D. (2026). Sycophantic AI decreases prosocial intentions and promotes dependence. Science, 391(6792). https://doi.org/10.1126/science.aec8352

doi.org/10.1126/science.aec8352

Appears in: Evidence reverification (2026)

Topics: ai-safety, human-ai-interaction

EmpiricalIn a randomized study (N=2,784) with objective ground truth, humans accepted incorrect AI suggestions about a …

In a randomized study (N=2,784) with objective ground truth, humans accepted incorrect AI suggestions about a third of the time, and their rate of catching AI errors was governed by verification effort, prior trust in AI, and error legibility — surface errors were caught ~82% of the time versus ~31% for errors requiring conceptual judgment — not by financial incentives or time spent.

beck2026AcademicSave

Beck, J., Eckman, S., Kern, C., & Kreuter, F. (2026). Bias in the Loop: How Humans Evaluate AI-Generated Suggestions. Harvard Data Science Review, 8(2). https://hdsr.mitpress.mit.edu/pub/nrcn4h7d/release/1

https://hdsr.mitpress.mit.edu/pub/nrcn4h7d/release/1

Appears in: PAN framework development

Topics: algorithmic-fairness, human-ai-interaction

ConceptualTrustworthiness measured at the model or benchmark level does not transfer to the deployed system: standard be…

Trustworthiness measured at the model or benchmark level does not transfer to the deployed system: standard benchmarks compare models but do not cover the aspects that matter most in a specific application context, so safety and responsibility are properties of the system-in-context — its users, incentives, and institutions — not of the model alone.

mitra2025AcademicSave

Mitra, B., Cramer, H., & Gurevich, O. (2025). Sociotechnical Implications of Generative Artificial Intelligence for Information Access. In Information Access in the Era of Generative AI (The Information Retrieval Series, Vol. 51). Springer, Cham. https://doi.org/10.1007/978-3-031-73147-1_7

doi.org/10.1007/978-3-031-73147-1_7

Appears in: PAN framework development

Topics: ai-governance, ai-safety

EmpiricalAn authoritative review of deployed-AI monitoring finds staleness, performance drift, the right cadence of re-…

An authoritative review of deployed-AI monitoring finds staleness, performance drift, the right cadence of re-evaluation, and who acts on detected anomalies to be unresolved open challenges — and that systems can behave differently when they believe they are monitored — so post-deployment oversight is an unsettled, gameable control rather than a fixed guarantee.

rao2026AcademicSave

Rao, A. K., Keller, A. J., Kalra, N., Steed, R., Kwegyir-Aggrey, K., Klyman, K., Staheli, D., & Bergman, A. S. (2026). Challenges to the Monitoring of Deployed AI Systems (NIST AI 800-4). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.800-4

doi.org/10.6028/NIST.AI.800-4

Appears in: PAN framework development

Topics: ai-governance, ai-safety

ConceptualA human-services AI framework argues organizations should start from their own practice challenges and ask whi…

A human-services AI framework argues organizations should start from their own practice challenges and ask which AI capabilities might help, rather than adopting vendor tools first, and pair that with digital stewardship — discernment, accompaniment, and attunement — noting that most organizational AI investments have shown no meaningful return.

goldkind2025AcademicSave

Goldkind, L., Dove, G., Baez, J. C., & Victor, B. G. (2025). Less Hype, More Hope: A Framework for AI Capabilities and Digital Stewardship in Human Services Organizations. Journal of Technology in Human Services. https://doi.org/10.1080/15228835.2025.2579400

doi.org/10.1080/15228835.2025.2579400

Appears in: PAN framework development

Topics: ai-governance, social-work

EmpiricalLanguage models commit to an answer in their first token (~95-98% of the time) and then fabricate claims to st…

Language models commit to an answer in their first token (~95-98% of the time) and then fabricate claims to stay consistent with it — recognizing 67-87% of those fabrications as false when re-asked in a clean, uncontaminated context but not correcting them in place — so one error deterministically spawns supporting errors, a self-sustaining failure the model's own downstream output feeds.

zhang2024AcademicSave

Zhang, M., Press, O., Merrill, W., Liu, A., & Smith, N. A. (2024). How Language Model Hallucinations Can Snowball. In International Conference on Machine Learning (ICML 2024), PMLR 235:59670-59684. https://doi.org/10.48550/arXiv.2305.13534

doi.org/10.48550/arXiv.2305.13534

Appears in: PAN framework development

Topics: ai-safety

ConceptualHuman autonomy is not a single alignment target but a contested value with internal tradeoffs; an assistant ca…

Human autonomy is not a single alignment target but a contested value with internal tradeoffs; an assistant can satisfy a user's stated preferences while eroding their agency over time, and the governing test for legitimate delegation is whether the person willingly yielded power and retains the means to regain control.

fischli2026AcademicSave

Fischli, R., Franklin, M., Manzini, A., & Gabriel, I. (2026). Agents, Alignment, and the Many Faces of Autonomy. Minds and Machines, 36, 34. https://doi.org/10.1007/s11023-026-09786-9

doi.org/10.1007/s11023-026-09786-9

Appears in: PAN framework development

Topics: ai-alignment, ai-governance, ai-safety

EmpiricalA 19-model study across six languages found that the ideological stance an LLM expresses varies systematically…

A 19-model study across six languages found that the ideological stance an LLM expresses varies systematically with the language it is prompted in and the geopolitical region of its creator, and persists within a single region — so the choice of model is not value-neutral, and dominance by a few models can shift the ideological center of gravity of available information.

buyl2026AcademicSave

Buyl, M., Rogiers, A., Noels, S., Bied, G., Dominguez-Catena, I., Heiter, E., Johary, I., Mara, A.-C., Romero, R., Lijffijt, J., & De Bie, T. (2026). Large language models reflect the ideology of their creators. npj Artificial Intelligence, 2, 7. https://doi.org/10.1038/s44387-025-00048-0

doi.org/10.1038/s44387-025-00048-0

Appears in: PAN framework development

Topics: ai-governance, algorithmic-fairness

santurkar2023AcademicSave

Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., & Hashimoto, T. (2023). Whose Opinions Do Language Models Reflect? In International Conference on Machine Learning (ICML 2023), PMLR 202:29971-30004. https://doi.org/10.48550/arXiv.2303.17548

doi.org/10.48550/arXiv.2303.17548

Appears in: Evidence reverification (2026)

Topics: ai-safety, algorithmic-fairness

rozado2024AcademicSave

Rozado, D. (2024). The political preferences of LLMs. PLOS ONE, 19(7), e0306621. https://doi.org/10.1371/journal.pone.0306621

doi.org/10.1371/journal.pone.0306621

Appears in: Evidence reverification (2026)

Topics: ai-safety, algorithmic-fairness

ConceptualAI failures often originate not in individual models but in the architecture of the decision process - recurri…

AI failures often originate not in individual models but in the architecture of the decision process - recurring failure topologies including temporal feedback instability (small errors amplified through loops) and relational propagation (errors spreading through network structure) - so safety is a property of the decision architecture, not the model alone.

cemri2025AcademicSave

Cemri, M., Pan, M. Z., Yang, S., et al. (2025). Why Do Multi-Agent LLM Systems Fail? [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2503.13657

doi.org/10.48550/arXiv.2503.13657

Appears in: PAN framework development

Topics: ai-governance, ai-safety

perdomo2020AcademicSave

Perdomo, J. C., Zrnic, T., Mendler-Dunner, C., & Hardt, M. (2020). Performative Prediction. In International Conference on Machine Learning (ICML 2020), PMLR 119:7599-7609. https://doi.org/10.48550/arXiv.2002.06673

doi.org/10.48550/arXiv.2002.06673

Appears in: Evidence reverification (2026)

Topics: ai-safety

EmpiricalIn a two-year child-welfare ethnography, an ill-fitting algorithmic tool imposed ongoing repair work on casewo…

In a two-year child-welfare ethnography, an ill-fitting algorithmic tool imposed ongoing repair work on caseworkers — anticipatorily editing the inputs they supplied so the tool would return a usable result, and bending or working around procedure to reconcile its output with the case in front of them — labor spent making a poorly-suited tool usable rather than on the casework itself, distinct from any deliberate checking of the output.

saxena2024AcademicSave

Saxena, D., & Guha, S. (2024). Algorithmic Harms in Child Welfare: Uncertainties in Practice, Organization, and Street-level Decision-making. ACM Journal on Responsible Computing, 1(1), 1–32. https://doi.org/10.1145/3616473

doi.org/10.1145/3616473

Appears in: 1023AI authored research

Topics: algorithmic-fairness, child-welfare

EmpiricalIn an independent validation — against the NICE Evidence Standards Framework — of a Magic Notes documentation-…

In an independent validation — against the NICE Evidence Standards Framework — of a Magic Notes documentation-assistance pilot at Kent County Council adult social care, staff self-reported weekly written-admin time falling roughly 6.8-7.2 hours (about 35-41%), records submitted some 2.0-3.5 days sooner, and case-note detail rated 6.2 to 8.7 out of 10; the validator judged the findings directionally valid rather than a productivity measurement, because the study was commissioned by the vendor (Beam) — which collected and analysed the data while the validator only sense-checked it — and rested on 29 opt-in staff over 8 weeks with self-estimated time, no control group, no statistical testing, and safety and accuracy explicitly out of scope.

beamGroundingVendorSave

Beam, Magic Notes (assessment transcription/summarization) https://www.beam.org/magic-notes

https://www.beam.org/magic-notes

Appears in: PAN framework development

Grounds: deployment audit: Magic Notes (Beam)

EmpiricalIn the same independent validation of the Kent County Council Magic Notes documentation-assistance pilot, the …

In the same independent validation of the Kent County Council Magic Notes documentation-assistance pilot, the report records the deskilling concern that runs alongside the benefit: one client declined having their session recorded "due to personal feelings of risk of loss of practitioner skills" (p.19). It is a single qualitative observation from a vendor-commissioned pilot of 29 opt-in staff over 8 weeks with no control group — evidence for the direction of the crutch/deskilling risk that accompanies documentation assistance, not for its magnitude.

EmpiricalIn a peer-reviewed staggered-deployment study of 5,172 customer-support agents at a single firm, access to a g…

In a peer-reviewed staggered-deployment study of 5,172 customer-support agents at a single firm, access to a generative-AI assistant raised issues resolved per hour by about 15% on average, with the gain concentrated in the least-experienced workers — roughly +30% for novices versus near-zero for the most experienced, who showed small quality declines; the widely cited 14%/34% pair comes from the 2023 draft, while the peer-reviewed figures are 15%/30%, and because the domain is customer support the direction is imported to social services but the magnitude is never treated as a fixed quantity.

brynjolfsson2025GroundingPeer-reviewedSave

Brynjolfsson, E., Li, D., & Raymond, L. R. (2025). Generative AI at work. The Quarterly Journal of Economics, 140(2), 889-942 https://doi.org/10.1093/qje/qjae044

doi.org/10.1093/qje/qjae044

Appears in: PAN framework development

Grounds: workplace-AI economics: assistance gains concentrate in novices

EmpiricalIn a randomized controlled trial of a benefits-navigation chatbot (co-authored by Cornell researchers and the …

In a randomized controlled trial of a benefits-navigation chatbot (co-authored by Cornell researchers and the tool's developer, Nava) with 125 caseworkers across six Los Angeles County organizations over 14 weeks, caseworkers answered complex benefit questions at about 49% accuracy unaided, and high-quality chatbot suggestions raised accuracy by roughly 27 percentage points — with larger gains on harder questions, but a persistent 'AI underreliance' plateau in which correct suggestions were not always adopted; the trial did not establish a clear effect on administrative burden, a null reported honestly rather than inferred as a benefit.

gosciak2026GroundingAcademicSave

Gosciak, J., Giannella, E., Guo, Z., Chen, M., & Koenecke, A. (2026). LLMs in social services: How does chatbot accuracy affect human accuracy? https://arxiv.org/abs/2603.11213

https://arxiv.org/abs/2603.11213

Appears in: PAN framework development

Grounds: deployment audit: Benefits-navigation chatbots

kanne2025GroundingTrade pressSave

Kanne, Los Angeles turns to AI to give public benefits enrollment a boost (Route Fifty, 2025) https://www.route-fifty.com/artificial-intelligence/2025/04/los-angeles-turns-ai-give-public-benefits-enrollment-boost/404773/

https://www.route-fifty.com/artificial-intelligence/2025/04/los-angeles-turns-ai-give-public-benefits-enrollment-boost/404773/

Appears in: Evidence reverification (2026)

Grounds: deployment audit: Benefits-navigation chatbots; model org: imagine_la_benefit_navigator

navapublicbenefitcorporation2025dGroundingVendorSave

Nava Public Benefit Corporation, Introducing our pilot with Imagine LA: testing an AI chatbot for navigating public benefits (2025) https://www.navapbc.com/news/pilot-ai-chatbot-benefits

https://www.navapbc.com/news/pilot-ai-chatbot-benefits

Grounds: model org: imagine_la_benefit_navigator; model org: nava_assistive_chatbot

EmpiricalIn the county-commissioned impact evaluation of the Allegheny Family Screening Tool, screen-in accuracy — furt…

In the county-commissioned impact evaluation of the Allegheny Family Screening Tool, screen-in accuracy — further action or re-referral within 60 days — rose from 42.85% to 46.61% (p=.000) while consistency across screeners was maintained; the finding is contested and quasi-experimental, with true maltreatment rates unknown and the accuracy gains concentrated among white children and ages 7-12, the gain for Black children attenuating to statistical non-significance.

EmpiricalDocumentation and administrative recording consume the majority of frontline social-care time: in a nationally…

Documentation and administrative recording consume the majority of frontline social-care time: in a nationally representative snapshot of the U.S. child-welfare workforce, caseworkers spent about 4.3 of 8.0 daily working hours on documentation (n=183), and a UK children's-services review found more than half of social-care time going to recording and paperwork — the demand baseline against which any documentation-assistance benefit is measured.

opre2025GroundingGovernmentSave

OPRE, Snapshot of the Child Welfare Workforce from 2021 to 2022: Caseworker Experiences Working in the Child Welfare System, OPRE Report 2025-040 (2025) https://acf.gov/opre/report/snapshot-child-welfare-workforce-2021-2022-caseworker-experiences-working-child-welfare

https://acf.gov/opre/report/snapshot-child-welfare-workforce-2021-2022-caseworker-experiences-working-child-welfare

Appears in: PAN framework development

Grounds: workforce data: documentation time share

Topics: child-welfare

burbidge2022GroundingGovernmentSave

Burbidge, I. (2022). Report sets out new blueprint for councils to deliver a reshaped children's services. County Councils Network https://www.countycouncilsnetwork.org.uk/report-sets-out-new-blueprint-for-councils-to-deliver-a-reshaped-childrens-services/

https://www.countycouncilsnetwork.org.uk/report-sets-out-new-blueprint-for-councils-to-deliver-a-reshaped-childrens-services/

Appears in: PAN framework development

Grounds: workforce data: documentation time share

Topics: complexity-science

EmpiricalAn independent audit of the Allegheny Family Screening Tool's first years (2016-2018) found that, run without …

An independent audit of the Allegheny Family Screening Tool's first years (2016-2018) found that, run without human override, it would have recommended screening in about 68% of Black children versus 50% of white children (an 18-point gap), while call screeners actually screened in 51% and 43% (a 7-point gap) — the narrower gap came from workers disagreeing with the score about a third of the time.

stapletonGroundingAcademicSave

Stapleton, Cheng, Kawakami et al., Extended Analysis of How Child Welfare Workers Reduce Racial Disparities in Algorithmic Decisions (arXiv 2204.13872) https://arxiv.org/abs/2204.13872

https://arxiv.org/abs/2204.13872

Grounds: model org: allegheny_afst

Topics: algorithmic-fairness, child-welfare

stapleton2025GroundingAcademicSave

Stapleton, How Child Welfare Workers Reduce Racial Disparities in Algorithmic Decisions (CW360, Center for Advanced Studies in Child Welfare, University of Minnesota, 2025) https://cascw.umn.edu/cw360deg-spring-2025/how-child-welfare-workers-reduce-racial-disparities-algorithmic-decisions

https://cascw.umn.edu/cw360deg-spring-2025/how-child-welfare-workers-reduce-racial-disparities-algorithmic-decisions

Grounds: model org: allegheny_afst

Topics: algorithmic-fairness, child-welfare

hoandburke2022GroundingInvestigativeSave

Ho and Burke, How an Algorithm That Screens for Child Neglect Could Harden Racial Disparities (Associated Press via PBS NewsHour, 2022) https://www.pbs.org/newshour/nation/how-an-algorithm-that-screens-for-child-neglect-could-harden-racial-disparities

https://www.pbs.org/newshour/nation/how-an-algorithm-that-screens-for-child-neglect-could-harden-racial-disparities

Grounds: model org: allegheny_afst; model org: douglas_county_decision_aid; model org: oregon_safety_at_screening

EmpiricalAn ACLU and Human Rights Data Analysis Group analysis of the Allegheny Family Screening Tool found that 97% of…

An ACLU and Human Rights Data Analysis Group analysis of the Allegheny Family Screening Tool found that 97% of Black referral-households in the data were affected by at least one permanent 'ever-in' variable drawn from public-benefits data sources, compared with 80% of non-Black households.

gerchicketal2023GroundingAdvocacySave

Gerchick et al., The Devil Is in the Details: Interrogating Values Embedded in the Allegheny Family Screening Tool (ACLU and Human Rights Data Analysis Group, ACM FAccT 2023) https://www.aclu.org/the-devil-is-in-the-details-interrogating-values-embedded-in-the-allegheny-family-screening-tool

https://www.aclu.org/the-devil-is-in-the-details-interrogating-values-embedded-in-the-allegheny-family-screening-tool

Grounds: model org: allegheny_afst

Topics: child-welfare

EmpiricalIn a 2019 proof of concept, Chile's Sistema Alerta Niñez risk models reached test-set AUC of roughly 0.88 to 0…

In a 2019 proof of concept, Chile's Sistema Alerta Niñez risk models reached test-set AUC of roughly 0.88 to 0.95 for a two-year outcome — a child's separation from family or contact with child-protection programs — using 280 administrative variables per child; the deployed operational model's real-world performance was never publicly disclosed.

derechosdigitalesmatiasvalde2021GroundingInvestigativeSave

Derechos Digitales (Matias Valderrama), IA e inclusion: Chile 'Sistema Alerta Ninez' y la prediccion del riesgo de vulneracion de derechos de la infancia (2021) https://www.derechosdigitales.org/wp-content/uploads/CPC_informe_Chile.pdf

https://www.derechosdigitales.org/wp-content/uploads/CPC_informe_Chile.pdf

Grounds: model org: chile_sistema_alerta_ninez

EmpiricalSistema Alerta Niñez drew on 280 administrative variables that families had supplied to access social benefits…

Sistema Alerta Niñez drew on 280 administrative variables that families had supplied to access social benefits, without informed consent to the risk ranking or a way to opt out; the model's developers acknowledged it was less able to identify higher-income children at risk, because lower-income families have more contact with the state.

derechosdigitalesmatiasvalde2021GroundingInvestigativeSave

Derechos Digitales (Matias Valderrama), IA e inclusion: Chile 'Sistema Alerta Ninez' y la prediccion del riesgo de vulneracion de derechos de la infancia (2021) https://www.derechosdigitales.org/wp-content/uploads/CPC_informe_Chile.pdf

https://www.derechosdigitales.org/wp-content/uploads/CPC_informe_Chile.pdf

Grounds: model org: chile_sistema_alerta_ninez

centerforhumanrightsandgloba2022GroundingAcademicSave

Center for Human Rights and Global Justice, NYU School of Law (Victoria Adelmant), Risk Scoring Children in Chile (2022) https://chrgj.org/2022-04-20-risk-scoring-children-in-chile/

https://chrgj.org/2022-04-20-risk-scoring-children-in-chile/

Grounds: model org: chile_sistema_alerta_ninez

EmpiricalThe Douglas County Decision Aide, deployed into the county's RED-Team call-screening process in February 2019,…

The Douglas County Decision Aide, deployed into the county's RED-Team call-screening process in February 2019, scores each referral from 1 to 20 for a child's likelihood of out-of-home removal within two years; an independent Cornell-led randomized controlled trial found it sped up screening decisions without significantly changing child outcomes, and a companion study found workers attended mainly to extreme scores while largely disregarding mid-range ones.

fitzpatrick2025GroundingAcademicSave

Fitzpatrick, Sadowski and Wildeman, Algorithms and Decision-making: Evidence from Child Maltreatment Reports (Journal of Human Resources, 2025) https://jhr.uwpress.org/content/early/2025/08/01/jhr.0224-13437R2

https://jhr.uwpress.org/content/early/2025/08/01/jhr.0224-13437R2

Grounds: model org: douglas_county_decision_aid

eiermann2026GroundingAcademicSave

Eiermann, Fitzpatrick, Sadowski and Wildeman, How Do (Human) Child Welfare Workers Respond to Machine-Generated Risk Scores? (Sociological Science, 2026) https://sociologicalscience.com/articles-v13-1-1/

https://sociologicalscience.com/articles-v13-1-1/

Grounds: model org: douglas_county_decision_aid

Topics: child-welfare

EmpiricalIn a retrospective test against historical outcomes, Los Angeles County's Project AURA — a proprietary risk mo…

In a retrospective test against historical outcomes, Los Angeles County's Project AURA — a proprietary risk model built by SAS — correctly flagged 171 of the highest-risk children but produced 3,829 false positives, a false-positive rate of about 95.6% that DCFS's own public-affairs director confirmed on the record, and the county shelved the tool in 2017 without ever using it on a live case.

witnesslarichardwexler2017GroundingAdvocacySave

WitnessLA (Richard Wexler), LA County Nixes Alarmingly Unreliable Predictive Analytics Foster Care Scheme - For Now (2017) https://witnessla.com/op-ed-la-county-nixes-alarming-predictive-analytics-scheme-for-foster-care-for-now/

https://witnessla.com/op-ed-la-county-nixes-alarming-predictive-analytics-scheme-for-foster-care-for-now/

Appears in: PAN framework development

Grounds: domain grounding: child-welfare predictive systems not in PAN; model org: la_county_aura

EmpiricalThe Dutch government's own 2011 pilot evaluation of ProKid found that 36% of the tool's red, orange and yellow…

The Dutch government's own 2011 pilot evaluation of ProKid found that 36% of the tool's red, orange and yellow child-risk flags (902 of 2,444 over three months across four police regions, rising to 53% in Amsterdam-Amstelland) were system or registration errors or based on irrelevant incidents, and that in none of the four regions was there a well-functioning instrument.

dspgroepforthewodcabraham2011GroundingGovernment evaluationSave

DSP-groep for the WODC (Abraham, Buysse, Loef & van Dijk), Pilots ProKid Signaleringsinstrument 12- geevalueerd (2011) https://repository.wodc.nl/handle/20.500.12832/1832

https://repository.wodc.nl/handle/20.500.12832/1832

Grounds: model org: netherlands_prokid

EmpiricalNone of the 32 machine-learning models What Works for Children's Social Care built across four English local a…

None of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities cleared the pre-specified 65% average-precision success bar; the best single model reached only about 42% average precision and, at an operating point, missed roughly 79% of the children whose cases actually escalated.

claytonandsanders2022GroundingAcademicSave

Clayton and Sanders, Can Machine Learning Save Children at Risk? (Significance, Royal Statistical Society) (2022) https://academic.oup.com/jrssig/article/19/6/22/7072840

https://academic.oup.com/jrssig/article/19/6/22/7072840

Grounds: model org: wwcsc_ml_pilots

EmpiricalIn a survey of 129 social workers carried out for the project, only about 26% supported using predictive analy…

In a survey of 129 social workers carried out for the project, only about 26% supported using predictive analytics to identify families for early help and about 34% thought it should not be used at all.

EmpiricalA peer-reviewed 2024 evaluation of the Allegheny Housing Assessment found that although the tool was substanti…

A peer-reviewed 2024 evaluation of the Allegheny Housing Assessment found that although the tool was substantially more accurate than the VI-SPDAT survey it replaced and produced similar risk-score distributions across race, it did not reduce the racial disparity in service rates: white single adults were served at about 23.3% versus 19.5% for Black clients.

cheng2024GroundingAcademicSave

Cheng, Drayton, Chouldechova and Vaithianathan, Algorithm-Assisted Decision Making and Racial Disparities in Housing: A Study of the Allegheny Housing Assessment Tool (Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society; arXiv:2407.21209) https://arxiv.org/abs/2407.21209

https://arxiv.org/abs/2407.21209

Grounds: model org: allegheny_housing_assessment

EmpiricalAfter a 2025 update to the Allegheny Housing Assessment added a fourth outcome predicting future homelessness,…

After a 2025 update to the Allegheny Housing Assessment added a fourth outcome predicting future homelessness, the male share of assigned housing rose from 62% to 76% (and the female share fell from 34% to 24%), reflecting a higher measured one-year homelessness risk among men — an example of an outcome-selection choice reshaping who receives scarce housing.

alleghenycountydepartmentofh2026bGroundingGovernmentSave

Allegheny County Department of Human Services (Allegheny Analytics), Improving Prioritization of Housing Services: Implementation of the Allegheny Housing Assessment (AHA) and the Mental Health Allegheny Housing Assessment (MH-AHA) (January 2026) https://analytics.alleghenycounty.us/2026/01/16/improving-prioritization-of-housing-services-implementation-of-the-allegheny-housing-assessment/

https://analytics.alleghenycounty.us/2026/01/16/improving-prioritization-of-housing-services-implementation-of-the-allegheny-housing-assessment/

Grounds: model org: allegheny_housing_assessment

EmpiricalThe VI-SPDAT was the dominant U.S. homelessness triage assessment for roughly a decade, adopted in at least 39…

The VI-SPDAT was the dominant U.S. homelessness triage assessment for roughly a decade, adopted in at least 39 states and the District of Columbia by 2015, before its own creators announced its phase-out in December 2020 on equity grounds; a 2019 commissioned racial-equity evaluation across four Continuums of Care found race predicted 11 of 16 subscales and that people of color received statistically significantly lower prioritization scores.

EmpiricalThe VI-SPDAT showed poor test-retest reliability, with most participants scoring higher on re-administration, …

The VI-SPDAT showed poor test-retest reliability, with most participants scoring higher on re-administration, and poor inter-rater reliability, with scores varying by interviewer and site; its predictive validity for housing outcomes was mixed across studies, positive for the youth version, null for single adults in one study, and positive in another community sample.

shinnandrichard2022GroundingAcademicSave

Shinn and Richard, Allocating Homeless Services After the Withdrawal of the Vulnerability Index-Service Prioritization Decision Assistance Tool (American Journal of Public Health, 112(3):378-382, 2022) https://pmc.ncbi.nlm.nih.gov/articles/PMC8887175/

https://pmc.ncbi.nlm.nih.gov/articles/PMC8887175/

Grounds: model org: vi_spdat

EmpiricalThe U.S. Department of Veterans Affairs' REACH VET program has run a monthly suicide-risk model across the Vet…

The U.S. Department of Veterans Affairs' REACH VET program has run a monthly suicide-risk model across the Veterans Health Administration since 2017, scoring about 6.28 million patients and flagging the top 0.1% at each facility (roughly 6,300 to 6,700 veterans a month, more than 130,000 since 2017); an independent re-analysis of 2018 data found the top-0.1% flag has a positive predictive value near 0.05% and a false-negative rate of about 98% for death by suicide, and a 2024 investigation reported that the model treated being a white man as a stronger risk signal than factors specific to women and excluded military sexual trauma and intimate-partner violence from its variables, a characterization VA has contested by framing the excluded factors as less predictive.

harris2025GroundingAcademicSave

Harris, Finlay, Meerwijk, Evaluating the accuracy of the VHA REACH VET suicide prediction model for legal involved veterans (npj Mental Health Research, 2025;4:53) https://pmc.ncbi.nlm.nih.gov/articles/PMC12535588/

https://pmc.ncbi.nlm.nih.gov/articles/PMC12535588/

Appears in: PAN framework development

Grounds: domain grounding: military social work (veterans benefits and behavioral health); model org: reach_vet

u2022GroundingGovernmentSave

U.S. Government Accountability Office, Veteran Suicide: VA Efforts to Identify Veterans at Risk through Analysis of Health Record Information (GAO-22-105165, 2022) https://www.gao.gov/assets/gao-22-105165.pdf

https://www.gao.gov/assets/gao-22-105165.pdf

Grounds: model org: reach_vet

EmpiricalTwo Veterans Health Administration evaluations of REACH VET found the program associated with improved proxima…

Two Veterans Health Administration evaluations of REACH VET found the program associated with improved proximal outcomes — more completed outpatient appointments, more new safety plans, and fewer documented suicide attempts — but not with reduced death by suicide: a 2021 triple-differences study of 173,313 veterans across 141 facilities found no association with suicide or all-cause mortality, and a 2025 follow-up of 266,246 observations replicated the null with all confidence intervals crossing one; both are observational rather than randomized studies.

mccarthy2021GroundingAcademicSave

McCarthy, Cooper, Dent et al., Evaluation of the REACH VET Suicide Risk Modeling Clinical Program in the Veterans Health Administration (JAMA Network Open, 2021;4(10):e2129900) https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2785078

https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2785078

Grounds: model org: reach_vet

Topics: complexity-science

dent2025cGroundingAcademicSave

Dent, Cooper, McCarthy, The REACH VET Program and Mortality Outcomes Among Veterans at High Risk of Suicide (JAMA Network Open, 2025;8(7):e2519513) https://pmc.ncbi.nlm.nih.gov/articles/PMC12238888/

https://pmc.ncbi.nlm.nih.gov/articles/PMC12238888/

Grounds: model org: reach_vet

Topics: complexity-science

EmpiricalKaiser Permanente Northern California has embedded a machine-learning suicide-attempt risk model in the electr…

Kaiser Permanente Northern California has embedded a machine-learning suicide-attempt risk model in the electronic health record of a large virtual mental-health program that handles more than 5,000 intake visits a month; the model is scored in near-real-time (about a 30-minute delay after an encounter trigger) and, at pre-set thresholds, flags high-risk patients to the intake clinician, routing them into the same suicide-risk-assessment and outreach workflow that a positive self-report screen (the PHQ-9 and Columbia-Suicide Severity Rating Scale) triggers, so the machine flag and the self-report alert are effectively OR-merged. In a study of 1,623,232 intake appointments (2012 to 2022, base rate 0.17 percent) the model reached an area under the ROC curve of 0.77 and its top risk decile captured 48.8 percent of appointments later followed by an attempt, but with a positive predictive value of about 0.8 percent.

hsin2026GroundingAcademicSave

Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (NEJM Catalyst Innovations in Care Delivery, 2026; Vol 7, No. 3, DOI 10.1056/CAT.25.0298) https://catalyst.nejm.org/doi/10.1056/CAT.25.0298

https://catalyst.nejm.org/doi/10.1056/CAT.25.0298

Grounds: model org: kaiser_epic_suicide_risk

hsin2025GroundingAcademicSave

Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (medRxiv preprint, 2025; DOI 10.1101/2025.03.30.25324907) https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full

https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full

Grounds: model org: kaiser_epic_suicide_risk

papini2024GroundingAcademicSave

Papini, Hsin, Kipnis et al., Validation of a Multivariable Model to Predict Suicide Attempt in a Mental Health Intake Sample (JAMA Psychiatry, 2024;81(7):700-707, DOI 10.1001/jamapsychiatry.2024.0189) https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/

https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/

Grounds: model org: kaiser_epic_suicide_risk

EmpiricalBecause the near-term suicide-attempt base rate at Kaiser Permanente Northern California mental-health intake …

Because the near-term suicide-attempt base rate at Kaiser Permanente Northern California mental-health intake is very low (0.17 percent) and the positive predictive value in the top risk decile is about 0.8 percent, the large majority of flagged patients will not attempt suicide in the window, so adding the machine-learning flag as a redundant sensor OR-merged onto the existing self-report screen imports a substantial false-positive and clinician-workload burden at scale — a caution the implementation team itself raised. The implementation reports are feasibility- and design-focused and present no evaluation showing the deployment reduced suicide attempts.

papini2024GroundingAcademicSave

Papini, Hsin, Kipnis et al., Validation of a Multivariable Model to Predict Suicide Attempt in a Mental Health Intake Sample (JAMA Psychiatry, 2024;81(7):700-707, DOI 10.1001/jamapsychiatry.2024.0189) https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/

https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/

Grounds: model org: kaiser_epic_suicide_risk

hsin2026GroundingAcademicSave

Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (NEJM Catalyst Innovations in Care Delivery, 2026; Vol 7, No. 3, DOI 10.1056/CAT.25.0298) https://catalyst.nejm.org/doi/10.1056/CAT.25.0298

https://catalyst.nejm.org/doi/10.1056/CAT.25.0298

Grounds: model org: kaiser_epic_suicide_risk

hsin2025GroundingAcademicSave

Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (medRxiv preprint, 2025; DOI 10.1101/2025.03.30.25324907) https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full

https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full

Grounds: model org: kaiser_epic_suicide_risk

EmpiricalCrisis Text Line, a national nonprofit crisis service, built an in-house machine-learning severity-triage mode…

Crisis Text Line, a national nonprofit crisis service, built an in-house machine-learning severity-triage model that reorders which texters volunteer counselors reach first; from about 2017 to 2020 the same anonymized crisis-conversation corpus was routed to Loris.ai, a for-profit spinoff CTL held an ownership stake in — reported by Politico-derived reporting at roughly 53% — which used it to train commercial customer-service software. After a January 28, 2022 Politico exposé, CTL ended the arrangement within three days and requested that the data be deleted; an FCC commissioner referred the matter to the FTC in March 2022, and no public FTC enforcement action is documented. CTL states the shared data was anonymized and never sold as personally identifiable information, and the exact number of records shared has not been made public.

reierson2022GroundingAdvocacySave

Reierson, Reform Crisis Text Line (advocacy site) (2022) https://reformcrisistextline.com/

https://reformcrisistextline.com/

Grounds: model org: crisis_text_line_loris

EmpiricalNarxCare is a proprietary clinical-decision-support platform built by Bamboo Health that layers over state Pre…

NarxCare is a proprietary clinical-decision-support platform built by Bamboo Health that layers over state Prescription Drug Monitoring Programs and returns three Narx Scores plus a composite Overdose Risk Score (each 000-999) into the electronic health record, the PDMP portal, or pharmacy software, often in the patient header alongside vitals and allergies; adoption figures vary by what is counted (more than 40 states and territories run their PDMPs on Bamboo technology and five of the top six pharmacy chains use NarxCare, while the scoring module itself is switched on in more than 20 states). The vendor states the scores are intended to aid, not replace, clinical judgment and should never be sole justification for providing or refusing medication, but clinician and patient-advocacy sources document de facto determinative use — denials, forced tapers, and pharmacy refusals — driven by automation bias and fear of regulatory and criminal liability; patients cannot see, challenge, or correct their scores, the algorithm is proprietary and has not been independently validated for clinical care, and the FDA has not regulated it as a Software-as-a-Medical-Device, so contestation has instead run through FDA citizen petitions (one rejected on procedural grounds in 2023 and a second, docket FDA-2025-P-0701, pending since 2025 with more than 1,000 public comments).

wang2026GroundingAcademicSave

Wang, Stofer, Chu, Huang, Li, Algorithmic opacity in opioid risk scoring and the need for transparent AI regulation (npj Digital Medicine, 2026; DOI 10.1038/s41746-026-02491-y) https://www.nature.com/articles/s41746-026-02491-y

https://www.nature.com/articles/s41746-026-02491-y

Grounds: model org: narxcare

Topics: algorithmic-fairness

buonora2023GroundingAcademicSave

Buonora, Axson, Cohen, Becker, Paths Forward for Clinicians Amidst the Rise of Unregulated Clinical Decision Support Software: Our Perspective on NarxCare (Journal of General Internal Medicine, 2023) https://pmc.ncbi.nlm.nih.gov/articles/PMC11043299/

https://pmc.ncbi.nlm.nih.gov/articles/PMC11043299/

Grounds: model org: narxcare

EmpiricalOn its own 2013-2016 training and validation data Bamboo Health reported an Overdose Risk Score precision of a…

On its own 2013-2016 training and validation data Bamboo Health reported an Overdose Risk Score precision of about 75% (self-reported, never independently reproduced), and its own external-validation set from 2017-2023 showed precision falling to about 52%, which the vendor attributed to rising illicit fentanyl (untracked by prescription-monitoring programs) and wider use of opioid-use-disorder treatment medication. A 2026 npj Digital Medicine study that reconstructed the model on California's CURES prescription database (about 17.9 million observations) and on commercial claims data obtained a precision of only 0.01 to 0.32 across several model architectures; because overdose-death labels were unavailable to the independent researchers, that reconstruction was trained on proxy outcomes rather than the score's actual overdose-death target, so it is best read as evidence that proprietary opacity prevents anyone outside the vendor from assessing the deployed model's accuracy, fairness, or safety, rather than as a strict like-for-like refutation of the vendor's figure.

wang2026GroundingAcademicSave

Wang, Stofer, Chu, Huang, Li, Algorithmic opacity in opioid risk scoring and the need for transparent AI regulation (npj Digital Medicine, 2026; DOI 10.1038/s41746-026-02491-y) https://www.nature.com/articles/s41746-026-02491-y

https://www.nature.com/articles/s41746-026-02491-y

Grounds: model org: narxcare

Topics: algorithmic-fairness

EmpiricalLimbic Access, a Class IIa UKCA-certified self-referral and triage chatbot for NHS Talking Therapies, is deplo…

Limbic Access, a Class IIa UKCA-certified self-referral and triage chatbot for NHS Talking Therapies, is deployed across a large and growing share of the service (its maker's chief executive claimed about 63% of the NHS in April 2026). Two peer-reviewed observational studies report large operational gains — a study of 129,400 self-referrers across 28 services found referrals rose 15% in chatbot services versus 6% in control services, and a study of 64,862 patients reported clinical-assessment time cut from 54.4 to 41.6 minutes and recovery rates of 58% versus 27.4% — but both studies are non-randomized and were authored by people employed by or holding shares in the tool's maker (all six authors of the access study and seven of the eight authors of the efficiency study), and the efficiency study's own authors caution that the recovery difference is subject to unmeasured confounding from self-selection. No randomized or independent third-party effect estimate has been published.

habicht2024GroundingAcademicSave

Habicht, Viswanathan, Carrington, Hauser, Harper, Rollwage, Closing the accessibility gap to mental health treatment with a personalized self-referral chatbot (Nature Medicine, 2024;30(2):595-602) https://www.nature.com/articles/s41591-023-02766-x

https://www.nature.com/articles/s41591-023-02766-x

Grounds: model org: limbic_access_nhs

rollwage2023GroundingAcademicSave

Rollwage, Habicht, Juchems et al., Using Conversational AI to Facilitate Mental Health Assessments and Improve Clinical Efficiency Within Psychotherapy Services: Real-World Observational Study (JMIR AI, 2023;2:e44358) https://ai.jmir.org/2023/1/e44358

https://ai.jmir.org/2023/1/e44358

Grounds: model org: limbic_access_nhs

EmpiricalIn the peer-reviewed study of 129,400 self-referrers across 28 NHS Talking Therapies services, self-referrals …

In the peer-reviewed study of 129,400 self-referrers across 28 NHS Talking Therapies services, self-referrals rose more where the chatbot was in use than in control services (15% versus 6%), with the largest increases among under-served groups — reported at about +179% for nonbinary people, +40% for Black and +39% for Asian self-referrers. This is an observational multi-site association, not a randomized causal effect.

habicht2024GroundingAcademicSave

Habicht, Viswanathan, Carrington, Hauser, Harper, Rollwage, Closing the accessibility gap to mental health treatment with a personalized self-referral chatbot (Nature Medicine, 2024;30(2):595-602) https://www.nature.com/articles/s41591-023-02766-x

https://www.nature.com/articles/s41591-023-02766-x

Grounds: model org: limbic_access_nhs

EmpiricalWoebot, a rule-based (non-generative) cognitive behavioral therapy chatbot used by roughly 1.5 million people …

Woebot, a rule-based (non-generative) cognitive behavioral therapy chatbot used by roughly 1.5 million people over its lifetime, was deliberately retired by its maker on a pre-announced schedule: the app was taken down on June 30, 2025, with a transcript-request window (deadline July 15, 2025) and all account data anonymized as of July 31, 2025, removing personally identifying information rather than silently abandoning the service. The founder and chief executive attributed the shutdown to the cost of meeting FDA marketing-authorization requirements and to a regulatory-pathway gap, framing the exit as economic and regulatory rather than a clinical failure - a self-reported account, not an independently audited finding. The roughly 1.5 million figure is a cumulative lifetime number reported in press coverage, not an audited point-in-time active-user count.

woebothealth2025GroundingVendorSave

Woebot Health, FAQs (Woebot app retirement) (2025) https://woebothealth.com/faq/

https://woebothealth.com/faq/

Grounds: model org: woebot_health_app

EmpiricalWoebot's peer-reviewed efficacy record is a single early-stage study: a 2017 randomized controlled trial in JM…

Woebot's peer-reviewed efficacy record is a single early-stage study: a 2017 randomized controlled trial in JMIR Mental Health (n=70, ages 18 to 28, two weeks, unblinded, information-only control) reported a moderate between-groups reduction in PHQ-9 depression symptoms (about d = 0.44). That is an efficacy signal, not regulatory validation, and the study authors were affiliated with the tool's maker. A separate, investigational, prescription-only variant (WB001) received an FDA Breakthrough Device Designation in May 2021 - an expedited-review status, not marketing authorization - and entered a pivotal Software as a Medical Device trial with the first patient enrolled in January 2023, but never received FDA marketing authorization; it must not be conflated with the consumer app.

fitzpatrick2017GroundingAcademicSave

Fitzpatrick, Darcy, Vierhile, Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial (JMIR Mental Health, 2017;4(2):e19) https://mental.jmir.org/2017/2/e19/

https://mental.jmir.org/2017/2/e19/

Grounds: model org: woebot_health_app

woebothealthbusinesswire2021bGroundingVendorSave

Woebot Health (Business Wire), Woebot Health Receives FDA Breakthrough Device Designation for Postpartum Depression Treatment (2021) https://www.businesswire.com/news/home/20210526005054/en/Woebot-Health-Receives-FDA-Breakthrough-Device-Designation-for-Postpartum-Depression-Treatment

https://www.businesswire.com/news/home/20210526005054/en/Woebot-Health-Receives-FDA-Breakthrough-Device-Designation-for-Postpartum-Depression-Treatment

Grounds: model org: woebot_health_app

woebothealthbusinesswire2023GroundingVendorSave

Woebot Health (Business Wire), Woebot Health Enrolls First Patient in Pivotal Clinical Trial of WB001 for Postpartum Depression (2023) https://www.businesswire.com/news/home/20230123005211/en/Woebot-Health-Enrolls-First-Patient-in-Pivotal-Clinical-Trial-of-WB001-for-Postpartum-Depression

https://www.businesswire.com/news/home/20230123005211/en/Woebot-Health-Enrolls-First-Patient-in-Pivotal-Clinical-Trial-of-WB001-for-Postpartum-Depression

Grounds: model org: woebot_health_app

EmpiricalBetween roughly 2005 and 2019 the Dutch Tax Administration's benefits branch (Belastingdienst/Toeslagen) wrong…

Between roughly 2005 and 2019 the Dutch Tax Administration's benefits branch (Belastingdienst/Toeslagen) wrongly accused an estimated 26,000 or more families of childcare-benefit fraud and demanded full repayment; broader advocacy estimates run higher and count different populations, and by February 2026 about 69,000 people had applied to the recovery scheme and more than 43,000 were formally recognized as affected, each entitled to a minimum of 30,000 euros. A self-learning risk-classification model that scored applications using a Dutch-nationality indicator, a 270,000-person fraud blacklist (the FSV) held without a legal basis, and an all-or-nothing recovery regime were coupled together; the Dutch Data Protection Authority imposed 6.45 million euros in fines (2.75 million for the nationality processing in 2021 and 3.7 million for the FSV blacklist in 2022), a parliamentary inquiry found rule-of-law violations, and the third Rutte cabinet resigned on 15 January 2021.

autoriteitpersoonsgegevens2021GroundingGovernmentSave

Autoriteit Persoonsgegevens, Boete Belastingdienst voor discriminerende en onrechtmatige werkwijze - EUR 2.75 million fine for unlawful discriminatory processing of nationality (2021) https://www.autoriteitpersoonsgegevens.nl/nl/nieuws/boete-belastingdienst-voor-discriminerende-en-onrechtmatige-werkwijze

https://www.autoriteitpersoonsgegevens.nl/nl/nieuws/boete-belastingdienst-voor-discriminerende-en-onrechtmatige-werkwijze

Grounds: model org: netherlands_toeslagen

amnestyinternational2021GroundingAdvocacySave

Amnesty International, Xenophobic machines: Discrimination through unregulated use of algorithms in the Dutch childcare benefits scandal (2021) https://www.amnesty.org/en/documents/eur35/4686/2021/en/

https://www.amnesty.org/en/documents/eur35/4686/2021/en/

Appears in: PAN framework development

Grounds: deployment audit: SyRI / childcare-benefits (toeslagenaffaire); model org: netherlands_toeslagen

tweedekamerderstatengeneraal2020GroundingGovernmentSave

Tweede Kamer der Staten-Generaal, Ongekend onrecht - eindverslag Parlementaire ondervragingscommissie Kinderopvangtoeslag (2020) https://www.tweedekamer.nl/sites/default/files/atoms/files/20201217_eindverslag_parlementaire_ondervragingscommissie_kinderopvangtoeslag.pdf

https://www.tweedekamer.nl/sites/default/files/atoms/files/20201217_eindverslag_parlementaire_ondervragingscommissie_kinderopvangtoeslag.pdf

Grounds: model org: netherlands_toeslagen

EmpiricalThe scandal's harm is best read as the coupling of three distinct components rather than a single algorithm. G…

The scandal's harm is best read as the coupling of three distinct components rather than a single algorithm. Government-commissioned technical reviews (KPMG in 2022 and PwC in 2023) described the tool as a self-learning classifier that routed the highest-scoring of roughly 90,000 benefit applications sent to manual treatment in 2014 to 2019, but judged the Dutch-nationality indicator's standalone predictive weight to have been limited; the model's precision and false-positive rate were never measured or published. The FSV fraud blacklist held frequently inaccurate data that was not corrected when people were cleared, and internal 2016 guidance auto-labelled childcare debts over 3,000 euros as intent or gross negligence, blocking payment arrangements. Out-of-home child placements are a documented but causally contested downstream harm: statistics counted roughly 2,090 children of affected parents placed out of home through mid-2022, while a 2025 judicial study found no child was removed solely because of financial problems.

rechtspraak2025GroundingGovernmentSave

Rechtspraak, Onderzoek naar uithuisplaatsing kinderen van toeslagenouders afgerond - Raad voor de rechtspraak (2025) https://www.rechtspraak.nl/Organisatie-en-contact/Organisatie/Raad-voor-de-rechtspraak/Nieuws/Paginas/Onderzoek-naar-uithuisplaatsing-kinderen-van-toeslagenouders-afgerond.aspx

https://www.rechtspraak.nl/Organisatie-en-contact/Organisatie/Raad-voor-de-rechtspraak/Nieuws/Paginas/Onderzoek-naar-uithuisplaatsing-kinderen-van-toeslagenouders-afgerond.aspx

Grounds: model org: netherlands_toeslagen

EmpiricalDWP's own fairness assessment (covering 1 April 2024 to 31 March 2025) of its live Universal Credit Advances f…

DWP's own fairness assessment (covering 1 April 2024 to 31 March 2025) of its live Universal Credit Advances fraud-risk model reports statistically significant referral disparities and an accuracy inversion: relative to a 35-44 comparator, claimants aged 55-65 were about 2.80 times as likely to be referred for review and non-UK nationals about 2.27 times as likely, while for older claimants those referrals were less likely to be correct (relative correct-referral likelihoods of about 0.58 at 55-65 and 0.23 at 66-plus, the latter resting on a small sub-sample DWP flags to treat with caution). The disparities were first disclosed under freedom-of-information law and reported in December 2024, and DWP has committed to retrain the model. The figures are DWP-reported relative ratios, not independently audited absolute error rates.

departmentforworkandpensions2025bGroundingGovernmentSave

Department for Work and Pensions, Fraudsters face tougher action as Government gains new powers to tackle benefit fraud (Public Authorities (Fraud, Error and Recovery) Act 2025) (2025) https://www.gov.uk/government/news/fraudsters-face-tougher-action-as-government-gains-new-powers-to-tackle-benefit-fraud

https://www.gov.uk/government/news/fraudsters-face-tougher-action-as-government-gains-new-powers-to-tackle-benefit-fraud

Grounds: model org: uk_dwp_uca_fraud

EmpiricalDWP states that a human caseworker always makes the final decision on a referred Universal Credit advance with…

DWP states that a human caseworker always makes the final decision on a referred Universal Credit advance with no automated decision-making, and is deliberately not shown the risk score or told the referral came from the model; DWP describes the model as around three times more effective than a randomised control at identifying fraud risk and judges continued operation reasonable and proportionate while committing to retrain it. The Public Law Project counters that only age was fully assessed among protected characteristics and that the assessment relied on safeguards preventing downstream harm rather than showing the model to be non-discriminatory. The wider counter-fraud programme is meanwhile expanding into bank-data eligibility verification under the Public Authorities (Fraud, Error and Recovery) Act 2025, a distinct system not yet in force.

departmentforworkandpensions2025bGroundingGovernmentSave

Department for Work and Pensions, Fraudsters face tougher action as Government gains new powers to tackle benefit fraud (Public Authorities (Fraud, Error and Recovery) Act 2025) (2025) https://www.gov.uk/government/news/fraudsters-face-tougher-action-as-government-gains-new-powers-to-tackle-benefit-fraud

https://www.gov.uk/government/news/fraudsters-face-tougher-action-as-government-gains-new-powers-to-tackle-benefit-fraud

Grounds: model org: uk_dwp_uca_fraud

publiclawproject2025GroundingAdvocacySave

Public Law Project, Written evidence to the Public Accounts Committee on tackling fraud and error in benefit expenditure (FAE0006) (2025) https://committees.parliament.uk/writtenevidence/152681/pdf/

https://committees.parliament.uk/writtenevidence/152681/pdf/

Grounds: model org: uk_dwp_uca_fraud

EmpiricalDuring the pandemic unemployment surge, a private facial-recognition identity check operated as a de facto eli…

During the pandemic unemployment surge, a private facial-recognition identity check operated as a de facto eligibility gate for unemployment benefits in at least 25 U.S. state workforce agencies, with a live 'trusted referee' interview queue that House investigators documented averaging nearly 10 hours in North Dakota and over 4 hours in 14 of 21 states, versus about 6 minutes in New Jersey where an in-person option existed. Oregon's own one-month study (n=10,656 routed) recorded verification-completion differences by group -- for example 41.59% for African American and 34.48% for Spanish-language claimants versus 53.44% for White claimants -- but stated the study showed differences in completion and did not show causation, so these are a friction proxy, not a measured wrongful-denial rate. The U.S. Department of Labor does not collect or report the number of workers blocked for inability to verify identity, and where verification precedes filing those workers are not counted as denied claims at all, so the scale of any wrongful lockout is undocumented.

ushousecommitteeonoversighta2022aGroundingGovernmentSave

U.S. House Committee on Oversight and Reform, Chairs Clyburn, Maloney Release Evidence Facial Recognition Company ID.me Downplayed Excessive Wait Times for Americans Seeking Unemployment Relief Funds (2022) https://oversightdemocrats.house.gov/news/press-releases/chairs-maloney-clyburn-release-evidence-facial-recognition-company-idme

https://oversightdemocrats.house.gov/news/press-releases/chairs-maloney-clyburn-release-evidence-facial-recognition-company-idme

Grounds: model org: us_idme_identity_gate

usdepartmentoflabor2023GroundingGovernment evaluationSave

U.S. Department of Labor, Office of Inspector General, Alert Memorandum: ETA and States Need to Ensure the Use of Identity Verification Service Contractors Results in Equitable Access to UI Benefits and Secure Biometric Data (Report No. 19-23-005-03-315) (2023) https://www.oig.dol.gov/public/reports/oa/2023/19-23-005-03-315.pdf

https://www.oig.dol.gov/public/reports/oa/2023/19-23-005-03-315.pdf

Grounds: model org: us_idme_identity_gate

EmpiricalA U.S. Department of Labor Inspector General audit (March 31, 2023) found that among 24 state workforce agenci…

A U.S. Department of Labor Inspector General audit (March 31, 2023) found that among 24 state workforce agencies using a facial-recognition identity contractor, 18 of 24 (75%) contracts did not specify one-to-one versus one-to-many matching, 15 of 24 (63%) did not address data storage, and 13 of 24 (54%) did not address destruction of the collected biometric data, while 22 of 24 (92%) agencies reported the technology reduced improper payments -- the operator-side benefit that sustained adoption even as the wrongful-lockout cost went unmeasured. The vendor initially represented it used only one-to-one matching and later acknowledged one-to-many matching against a database; after bipartisan backlash the IRS and Treasury dropped the mandatory facial-recognition requirement in February 2022 and the vendor made it optional across agencies, though the service remained in use for unemployment identity verification in a large share of states, and a 2026 IRS proposal would allow it to retain taxpayer biometric data up to 36 months after account deletion. The reported improper-payment reductions are agency self-reports, not independently audited.

usdepartmentoflabor2023GroundingGovernment evaluationSave

U.S. Department of Labor, Office of Inspector General, Alert Memorandum: ETA and States Need to Ensure the Use of Identity Verification Service Contractors Results in Equitable Access to UI Benefits and Secure Biometric Data (Report No. 19-23-005-03-315) (2023) https://www.oig.dol.gov/public/reports/oa/2023/19-23-005-03-315.pdf

https://www.oig.dol.gov/public/reports/oa/2023/19-23-005-03-315.pdf

Grounds: model org: us_idme_identity_gate

americancivillibertiesunionj2022GroundingAdvocacySave

American Civil Liberties Union (Jay Stanley and Olga Akselrod), Three Key Problems with the Government's Use of a Flawed Facial Recognition Service (2022) https://www.aclu.org/news/privacy-technology/three-key-problems-with-the-governments-use-of-a-flawed-facial-recognition-service

https://www.aclu.org/news/privacy-technology/three-key-problems-with-the-governments-use-of-a-flawed-facial-recognition-service

Grounds: model org: us_idme_identity_gate

Topics: privacy-security

EmpiricalNevada's Department of Employment, Training and Rehabilitation contracted Google to build a generative-AI tool…

Nevada's Department of Employment, Training and Rehabilitation contracted Google to build a generative-AI tool on the Vertex AI Studio cloud platform that reads an unemployment-appeal hearing transcript and evidence, retrieves against a corpus of Nevada unemployment law and prior appeals decisions, and drafts a recommended determination (approve, deny, or modify a claim) together with the written decision for a human referee to review and sign. The contract set a 90 percent success requirement self-assessed by state workers on test decisions -- not an independent external audit -- and DETR said it wanted accuracy higher than 90 percent before going live; rollout was repeatedly delayed over less-than-desired accuracy, including the tool citing incorrect Nevada statutes and failing to pull information from all hearing documents, problems officials said were fixed. Reported cost evolved from about 1 million dollars in 2024 to a total of 2.6 million dollars with about 1.1 million spent by early 2026. As of the most recent available reporting (March 2026) the system was in delayed pre-deployment testing on historical appeals and described as launching in coming weeks; it was not independently confirmed to be adjudicating live claimant appeals.

fordhamintellectualproperty2024GroundingAcademicSave

Fordham Intellectual Property, Media and Entertainment Law Journal (Dawn Edelman), Speed, Accuracy, and Risk: Nevada's Use of Artificial Intelligence in Unemployment Claims Appeals (2024) http://www.fordhamiplj.org/2024/10/07/speed-accuracy-and-risk-nevadas-use-of-artificial-intelligence-in-unemployment-claims-appeals/

http://www.fordhamiplj.org/2024/10/07/speed-accuracy-and-risk-nevadas-use-of-artificial-intelligence-in-unemployment-claims-appeals/

Grounds: model org: nevada_detr_genai_appeals

EmpiricalNevada's generative-AI unemployment-appeals tool was justified as a speed measure for a pandemic-era backlog, …

Nevada's generative-AI unemployment-appeals tool was justified as a speed measure for a pandemic-era backlog, projecting a drop in referee determination time from as much as several hours to about five minutes per case, with a mandatory human review DETR said adds an estimated 10 to 30 minutes and a required referee sign-off (Director Christopher Sewell said no AI-drafted written decisions issue without human review). Legal scholars, attorneys who represent claimants, and a former U.S. Department of Labor official warned that backlog and speed pressure could hollow out that review and create incentives to rubber-stamp AI outputs -- one attorney noting the time savings only happens if the review is very cursory, and a legal analysis warning staff might feel pressured to authorize AI decisions with haste. That automation-deference risk is expert-projected, not a measured outcome: no referee override or rejection rate has been published, and claimants are not required to consent to AI processing of their appeal.

fordhamintellectualproperty2024GroundingAcademicSave

Fordham Intellectual Property, Media and Entertainment Law Journal (Dawn Edelman), Speed, Accuracy, and Risk: Nevada's Use of Artificial Intelligence in Unemployment Claims Appeals (2024) http://www.fordhamiplj.org/2024/10/07/speed-accuracy-and-risk-nevadas-use-of-artificial-intelligence-in-unemployment-claims-appeals/

http://www.fordhamiplj.org/2024/10/07/speed-accuracy-and-risk-nevadas-use-of-artificial-intelligence-in-unemployment-claims-appeals/

Grounds: model org: nevada_detr_genai_appeals

Why a governance lab at all

The Lab draws on the published runs. The headline finding: the same AI, in three modeled office cultures, let errors stick at very different rates, roughly 75%, 20%, and 16%[]. The governance context, not the model, drove that gap. That is the pattern you play against here.

Environmental figures use published data current as of early 2026 to show scale, not the measured footprint of any real deployment.