Skip to content

Evidence · The claim ledger

Other534

Every ledgered claim this site makes in this evidence area, with the sources that ground it — or, for a PAN-simulation-derived claim, the run it comes from. Source keys link back to the full reference lists on the Evidence Registry.

EmpiricalAn independent audit of the Allegheny Family Screening Tool's first years (2016-2018) found that, run without human over…

An independent audit of the Allegheny Family Screening Tool's first years (2016-2018) found that, run without human override, it would have recommended screening in about 68% of Black children versus 50% of white children (an 18-point gap), while call screeners actually screened in 51% and 43% (a 7-point gap) — the narrower gap came from workers disagreeing with the score about a third of the time.

Sources: stapleton, stapleton2025, hoandburke2022

Appears on: /domains/cases/allegheny-afst, /pan-lab

EmpiricalAn ACLU and Human Rights Data Analysis Group analysis of the Allegheny Family Screening Tool found that 97% of Black ref…

An ACLU and Human Rights Data Analysis Group analysis of the Allegheny Family Screening Tool found that 97% of Black referral-households in the data were affected by at least one permanent 'ever-in' variable drawn from public-benefits data sources, compared with 80% of non-Black households.

Sources: gerchicketal2023

Appears on: /domains/cases/allegheny-afst, /pan-lab

EmpiricalThe U.S. Department of Justice's Civil Rights Division was reported to be scrutinizing the Allegheny Family Screening To…

The U.S. Department of Justice's Civil Rights Division was reported to be scrutinizing the Allegheny Family Screening Tool after civil-rights complaints filed in fall 2022 raised concerns that its use of disability, mental-health, and Supplemental Security Income data may discriminate against parents with disabilities; families are not shown their scores, and no public findings or enforcement have been reported.

Sources: associatedpress2023, hoandburke2023

Appears on: /domains/cases/allegheny-afst

EmpiricalIn Allegheny County's Hello Baby program, the top-tier roughly 5% of newborns by predictive risk score accounted for abo…

In Allegheny County's Hello Baby program, the top-tier roughly 5% of newborns by predictive risk score accounted for about 54% of children later removed from the home by age three, at roughly twenty times the removal risk of other newborns (methodology relative risk 22.24, 95% CI 17.50-28.25); the model reported an AUC of about 0.93 on holdout data.

Sources: centreforsocialdataanalytics2020, vaithianathan2025

Appears on: /domains/cases/allegheny-hello-baby

EmpiricalAn external evaluation of Hello Baby covering birth cohorts 2016-2024, controlling for COVID-19, found the program assoc…

An external evaluation of Hello Baby covering birth cohorts 2016-2024, controlling for COVID-19, found the program associated with fewer first child-maltreatment investigations and first substantiated investigations, but no reduction in out-of-home foster-care placements - the outcome the predictive model was built to estimate.

Sources: lery2025, centreforsocialdataanalytics2020

Appears on: /domains/cases/allegheny-hello-baby

EmpiricalIn a 2019 proof of concept, Chile's Sistema Alerta Niñez risk models reached test-set AUC of roughly 0.88 to 0.95 for a …

In a 2019 proof of concept, Chile's Sistema Alerta Niñez risk models reached test-set AUC of roughly 0.88 to 0.95 for a two-year outcome — a child's separation from family or contact with child-protection programs — using 280 administrative variables per child; the deployed operational model's real-world performance was never publicly disclosed.

Sources: derechosdigitalesmatiasvalde2021, derechosdigitalesmatiasvalde2022, centreforsocialdataanalytics2019b

Appears on: /domains/cases/chile-alerta-ninez, /pan-lab

EmpiricalSistema Alerta Niñez drew on 280 administrative variables that families had supplied to access social benefits, without …

Sistema Alerta Niñez drew on 280 administrative variables that families had supplied to access social benefits, without informed consent to the risk ranking or a way to opt out; the model's developers acknowledged it was less able to identify higher-income children at risk, because lower-income families have more contact with the state.

Sources: derechosdigitalesmatiasvalde2021, derechosdigitalesmatiasvalde2022, centerforhumanrightsandgloba2022

Appears on: /domains/cases/chile-alerta-ninez, /pan-lab

EmpiricalThe Douglas County Decision Aide, deployed into the county's RED-Team call-screening process in February 2019, scores ea…

The Douglas County Decision Aide, deployed into the county's RED-Team call-screening process in February 2019, scores each referral from 1 to 20 for a child's likelihood of out-of-home removal within two years; an independent Cornell-led randomized controlled trial found it sped up screening decisions without significantly changing child outcomes, and a companion study found workers attended mainly to extreme scores while largely disregarding mid-range ones.

Sources: vaithianathanetalcentreforso2019, fitzpatrick2025, eiermann2026

Appears on: /domains/cases/douglas-county-decision-aid, /pan-lab

EmpiricalEckerd's Rapid Safety Feedback spread from Hillsborough County, Florida to child-welfare agencies in several states — pr…

Eckerd's Rapid Safety Feedback spread from Hillsborough County, Florida to child-welfare agencies in several states — promoted on the vendor's own reported gains and highlighted as “innovative” in a 2016 federal commission report — years before an independent 2022 peer-reviewed evaluation found the process did not lower repeat high-severity maltreatment among children identified as high risk (a joint odds ratio of about 1.05).

Sources: eckerdconnects2016, routefiftygovernmentexecutiv2016, parker2022, floridaschildrenfirst2012

Appears on: /domains/cases/eckerd-florida-rsf-origin

EmpiricalGladsaxe's early-detection project (DTO) was a decision-tree model over about 44 risk indicators, meant to score, for ev…

Gladsaxe's early-detection project (DTO) was a decision-tree model over about 44 risk indicators, meant to score, for every child aged 0 to 6 rather than only families already receiving help, the estimated probability that the child was living in vulnerability; per a university-run Danish public-sector AI catalogue it was to be trained on roughly 173,000 notifications the authorities received between April 2016 and December 2017, but only about 117 usable historical cases existed, and it was halted in its development phase in 2019 without ever running on live decisions, after a national media storm and an unrelated data breach that exposed about 20,000 citizens' personal identification numbers.

Sources: offentligaiuniversityrundanind, kennethkristensensamfundsled2022, helenefriisratnerandkasperel2023, katarinafastlappalainen2021, tvkosmopolformerlytvlorry2018

Appears on: /domains/cases/gladsaxe-denmark

EmpiricalHackney paid the analytics firm Xantura £361,400 over four years to run an Early Help Profiling System that flagged fami…

Hackney paid the analytics firm Xantura £361,400 over four years to run an Early Help Profiling System that flagged families for preventive intervention from council data, but scrapped the pilot in 2019 after finding that, despite flagging about 350 families, it surfaced only 7 children previously unknown to the council and the available data was too limited and variable to justify continuing.

Sources: hackneycouncilpayskpoundstod2018, townhalldropspilotprogrammep2019

Appears on: /domains/cases/hackney-early-help

EmpiricalFamilies whose data Hackney's Early Help Profiling System processed were not informed directly: reporting describes fami…

Families whose data Hackney's Early Help Profiling System processed were not informed directly: reporting describes families profiled without their knowledge, given notice only through a general online privacy notice, with no option to opt out recorded in the system's impact assessment and the method withheld as commercially sensitive; the council argued that disclosing the system could prejudice potential interventions.

Sources: townhalldropspilotprogrammep2019, reddenj2020, hackneycouncilpayskpoundstod2018

Appears on: /domains/cases/hackney-early-help

EmpiricalInternal DCFS tracking data released under Illinois public-records law showed the Rapid Safety Feedback tool flagged mor…

Internal DCFS tracking data released under Illinois public-records law showed the Rapid Safety Feedback tool flagged more than 4,100 children at a 90-percent-or-higher probability of death or serious injury within two years, including 369 children under age 9 assigned a 100-percent probability, while children who died in cases already known to the system — among them 17-month-old Semaj Crosby, found dead after at least ten DCFS investigations — were not flagged as top-risk; the roughly $366,000 program was ended in 2017.

Sources: chicagotribune2017, governmenttechnologya

Appears on: /domains/cases/illinois-rapid-safety-feedback

EmpiricalIllinois brought in the Eckerd/MindShare Rapid Safety Feedback program under DCFS director George Sheldon through a no-b…

Illinois brought in the Eckerd/MindShare Rapid Safety Feedback program under DCFS director George Sheldon through a no-bid arrangement the state classified as a grant; a July 2017 joint report by the Illinois Office of Executive Inspector General and the DCFS Inspector General found this classification to be mismanagement because it avoided state bidding-transparency requirements.

Sources: chicagotribune2017, sunshinestatenews2017

Appears on: /domains/cases/illinois-rapid-safety-feedback

EmpiricalBristol's Think Family Database drew on roughly 30 to 35 fused council, police and other datasets covering about 55,000 …

Bristol's Think Family Database drew on roughly 30 to 35 fused council, police and other datasets covering about 55,000 families (some 170,000 residents in 2021 reporting), and its child sexual and criminal exploitation risk models were quietly withdrawn in 2023 as 'not fit for operational use' after an independent evaluation judged the risk-scoring models the weakest element and staff reported victims of exploitation scoring below people involved in burglary; FOI responses indicate no record was kept of why the models were switched off, and auditors could not locate their source code or variable lists.

Sources: seanmorrison2026a, markwildingandmattburgess2026, seanmorrison2026b, bristolcitycouncil2025, jakehurfurtbigbrotherwatch2021

Appears on: /domains/cases/insight-bristol

EmpiricalReporting and FOI responses on Bristol's Think Family Database indicate the exploitation models' source code and variabl…

Reporting and FOI responses on Bristol's Think Family Database indicate the exploitation models' source code and variable lists could not be located when auditors sought them, and that an ethics committee advising the police analytics reportedly did not revisit the analytics after 2017; a 2021 review warned that data gathered through 'legal gateways' meant 'legality is not the same as legitimacy.'

Sources: markwildingandmattburgess2026, seanmorrison2026a, seanmorrison2026b

Appears on: /domains/cases/insight-bristol

EmpiricalIn a retrospective test against historical outcomes, Los Angeles County's Project AURA — a proprietary risk model built …

In a retrospective test against historical outcomes, Los Angeles County's Project AURA — a proprietary risk model built by SAS — correctly flagged 171 of the highest-risk children but produced 3,829 false positives, a false-positive rate of about 95.6% that DCFS's own public-affairs director confirmed on the record, and the county shelved the tool in 2017 without ever using it on a live case.

Sources: theimprintdanielheimpel2015, childprotectiveservicesdefen2015, nccprrichardwexler2017, witnesslarichardwexler2017

Appears on: /domains/cases/la-county-aura, /pan-lab

EmpiricalThe Dutch government's own 2011 pilot evaluation of ProKid found that 36% of the tool's red, orange and yellow child-ris…

The Dutch government's own 2011 pilot evaluation of ProKid found that 36% of the tool's red, orange and yellow child-risk flags (902 of 2,444 over three months across four police regions, rising to 53% in Amsterdam-Amstelland) were system or registration errors or based on irrelevant incidents, and that in none of the four regions was there a well-functioning instrument.

Sources: dspgroepforthewodcabraham2011, dimitritokmetzissargasso2012

Appears on: /domains/cases/netherlands-prokid, /pan-lab

EmpiricalNew Zealand's Ministry of Social Development commissioned a child-maltreatment risk-modelling tool that, on a 2012 devel…

New Zealand's Ministry of Social Development commissioned a child-maltreatment risk-modelling tool that, on a 2012 development sample of 57,986 children and 132 selected variables, reported an area under the ROC curve of 76% and a top risk decile in which 47.8% had a substantiated maltreatment finding by age five; those figures come from development data rather than field performance, the tool was never operationally deployed, and a proposed two-year study that would have scored about 60,000 newborns was halted by the incoming Social Development Minister, who annotated the briefing papers 'Not on my watch! These are children not lab rats.'

Sources: vaithianathan2013, nzherald2015, otagodailytimes2015, mordaunt2026

Appears on: /domains/cases/nz-msd-prm

EmpiricalOregon's Department of Human Services stopped using its Safety at Screening tool at the end of June 2022 and replaced it…

Oregon's Department of Human Services stopped using its Safety at Screening tool at the end of June 2022 and replaced it with a non-algorithmic Structured Decision Making process, telling staff the change was meant to reduce disparities; the move followed Associated Press reporting on racial disparity in the Allegheny tool it was derived from and a racial-bias inquiry from a U.S. senator.

Sources: associatedpress2022, willametteweek2022, hoandburke2022, nprap2022

Appears on: /domains/cases/oregon-safety-at-screening

EmpiricalOregon's 2019 report describes a post-processing fairness correction — group-specific thresholds selected under an 'erro…

Oregon's 2019 report describes a post-processing fairness correction — group-specific thresholds selected under an 'error rate balance' criterion — applied to a dual-outcome risk model built only on the state's own child-welfare administrative records.

Sources: orrai2019, associatedpress2022

Appears on: /domains/cases/oregon-safety-at-screening

EmpiricalNone of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities…

None of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities cleared the pre-specified 65% average-precision success bar; the best single model reached only about 42% average precision and, at an operating point, missed roughly 79% of the children whose cases actually escalated.

Sources: claytonandsanders2022, communitycareturner2020a, childhubterredeshommes2020

Appears on: /domains/cases/wwc-uk-ml-pilots, /pan-lab

EmpiricalIn a survey of 129 social workers carried out for the project, only about 26% supported using predictive analytics to id…

In a survey of 129 social workers carried out for the project, only about 26% supported using predictive analytics to identify families for early help and about 34% thought it should not be used at all.

Sources: communitycareturner2020a

Appears on: /domains/cases/wwc-uk-ml-pilots, /pan-lab

EmpiricalIn a Los Angeles County pilot, 335 people who enrolled in the voluntary Homelessness Prevention Unit were reported to be…

In a Los Angeles County pilot, 335 people who enrolled in the voluntary Homelessness Prevention Unit were reported to be 71% less likely than a regression-adjusted comparison group of 1,285 eligible non-enrollees to enter a homeless shelter or have street-outreach contact within 18 months; the California Policy Lab describes this as an association not yet shown to be causal, pending a randomized controlled trial with results expected in 2027.

Sources: blackwell2025, countyoflosangeles2025, uclanewsroom2025

Appears on: /domains/cases/la-homelessness-prevention

EmpiricalThe Homelessness Prevention Unit's own November 2024 equity audit, on a test population of 47,582 individuals eligible t…

The Homelessness Prevention Unit's own November 2024 equity audit, on a test population of 47,582 individuals eligible to be scored, reported false-negative rates ranging from about 56% for Black individuals to roughly 63 to 65% for other groups: the model misses a majority of the people who later become homeless, while performing roughly consistently across race, ethnicity, and gender and identifying Black individuals slightly more strongly.

Sources: californiapolicylab2024, foxsowell2025

Appears on: /domains/cases/la-homelessness-prevention

EmpiricalXantura's OneView integrates more than 15 multi-agency data feeds into a single household view and flags residents as li…

Xantura's OneView integrates more than 15 multi-agency data feeds into a single household view and flags residents as likely to become homeless months ahead. In Maidstone's pilot year it produced 650-plus alerts that a single financial-inclusion officer could contact only about 260 of. Its headline effectiveness figures - a reported 40 percent fall in homelessness, savings and an ROI over 600 percent, and the widely quoted contrast between contacted and uncontacted households - are vendor- and council-reported pre/post numbers from one COVID-affected pilot year; the contact-versus-no-contact contrast reflects capacity-driven selection rather than a randomised comparison, and the independent randomised controlled trial commissioned to test the causal claim was still in progress into 2026.

Sources: crisisuk2023, xantura2023, governmenttransformationmaga2023, ministryofhousing2024, centreforhomelessnessimpact2024

Appears on: /domains/cases/xantura-oneview-housing

EmpiricalOneView's single view of vulnerability is built by integrating sensitive multi-agency records - including offending, hea…

OneView's single view of vulnerability is built by integrating sensitive multi-agency records - including offending, health, benefits and debt data - under a statutory Digital Economy Act 2017 data-sharing agreement with named public-body controllers and processors. An independent ethnography of an early deployment (its fieldwork centered on children's social care and the COVID-19 response) found frontline staff could not see which factors drove the tool's alerts and were not all convinced it was as accurate as described, and a separate NGO investigation characterised the vendor's COVID-era model as operating without residents' knowledge.

Sources: digitaleconomyactregister2023, adalovelaceinstitute2024, bigbrotherwatch2021

Appears on: /domains/cases/xantura-oneview-housing

EmpiricalLondon, Ontario's CHAI is a live, caseworker-facing machine-learning model that flags people in the city's shelter syste…

London, Ontario's CHAI is a live, caseworker-facing machine-learning model that flags people in the city's shelter system as at risk of chronic homelessness (more than 180 shelter days in a year) about six months ahead; it provides intelligence to prevention caseworkers and does not itself make service decisions. Its widely repeated '93 percent accuracy' is a builder-reported, testing-phase figure from 10-fold cross-validation on historical HIFIS records, never independently validated after deployment; the same technical work reports recall of about 0.921 but precision of only about 0.651, implying substantial false positives under a low base rate.

Sources: wray2020, vanberlo2009, lebel2023, govlaunchstories2020

Appears on: /domains/cases/chai-london-ontario

EmpiricalCHAI is consent-based: it draws on de-identified HIFIS records pooled from roughly 20 to 24 London homelessness-support …

CHAI is consent-based: it draws on de-identified HIFIS records pooled from roughly 20 to 24 London homelessness-support organizations and lets individuals opt out of inclusion, and it was built with reference to GDPR principles, Canada's Directive on Automated Decision-Making, and local feature-attribution explanations for caseworkers. Because HIFIS captures people who use public shelters, an independent review and reporting at launch note it can under-represent or miss groups who avoid them - including many women, families, new immigrants, some Indigenous people, and private-shelter users; academic researchers situating the tool raise related fairness and inequality concerns. So the population the model can score is a selected sample of actual need, and the opt-out self-selects it further.

Sources: wray2020, lebel2023, lamberink2020, redden2026

Appears on: /domains/cases/chai-london-ontario

EmpiricalIn a 2025 Los Angeles County pilot evaluated by Nava Labs with academic partners at Cornell University and Georgetown Un…

In a 2025 Los Angeles County pilot evaluated by Nava Labs with academic partners at Cornell University and Georgetown University's Better Government Lab, a generative-AI assistive chatbot for Imagine LA's Benefit Navigator was estimated to improve benefits-navigation answer accuracy by an average of about 40% in a randomized controlled trial of 125 caseworkers answering hypothetical client questions, alongside a fourteen-week field pilot with 61 caseworkers across six organizations; the evaluation was co-authored by the tool builder rather than independently replicated, the accuracy figure is a decision-support contrast on hypothetical questions rather than a live-caseload eligibility audit, and time-savings and administrative-burden effects were reported as promising but inconclusive (published March 2026).

Sources: navapublicbenefitcorporation2026, chen2026

Appears on: /domains/cases/imagine-la-benefit-navigator

EmpiricalThe same evaluation reported that the chatbot's accuracy gains were largest on the most difficult client questions and a…

The same evaluation reported that the chatbot's accuracy gains were largest on the most difficult client questions and among the newest, least-experienced staff (a directional finding, not a quantified breakdown), that about 65% of caseworkers with access used it at an average of about 14 prompts each and a modest, low-positive satisfaction (a Net Promoter Score of 11), that usage tended to decline over time without sustained engagement, and that answers averaged a tenth-to-twelfth-grade reading level against college-level source manuals.

Sources: navapublicbenefitcorporation2026, chen2026

Appears on: /domains/cases/imagine-la-benefit-navigator

EmpiricalIn a single-center randomized trial across three Vanderbilt neurology clinics (August 2022 to February 2023), an EHR sui…

In a single-center randomized trial across three Vanderbilt neurology clinics (August 2022 to February 2023), an EHR suicide-risk model flagged 596 of 7,732 encounters (about 8%) at a 2%-or-higher 30-day-risk threshold; making the identical alert interruptive rather than passive led clinicians to elect a suicide-risk screen in 42% of encounters (121/289) versus 4% (12/307) for a passive chart icon, an adjusted odds ratio of 17.70 (95% CI 6.42–48.79). Screening remained fully advisory: about 58% of interruptive and 96% of passive alerts produced no screening.

Sources: walshetal2025, aitestedforalertingclinician2025, suicidepreventionmorefeasibl2025

Appears on: /domains/cases/vsail-vanderbilt

EmpiricalIn a separate 2021 prospective silent-mode study (115,905 predictions on 77,973 patients, June 2019 to April 2020), the …

In a separate 2021 prospective silent-mode study (115,905 predictions on 77,973 patients, June 2019 to April 2020), the model reported a c-statistic of 0.797 for suicide attempt and 0.836 for ideation center-wide but only 0.544 for attempt in behavioral-health settings, and in the highest-risk quantile the number-needed-to-screen was 271 for attempt and 23 for ideation. In the 2022 to 2023 trial no suicidal ideation or attempts were documented in either arm during 30-day follow-up, and the trial was explicitly not powered for clinical outcomes, so it measured a process outcome (screening) rather than reduced harm.

Sources: walshetal2021, walshetal2025, suicidepreventionmorefeasibl2025

Appears on: /domains/cases/vsail-vanderbilt

EmpiricalBetween roughly 2005 and 2019 the Dutch Tax Administration's benefits branch (Belastingdienst/Toeslagen) wrongly accused…

Between roughly 2005 and 2019 the Dutch Tax Administration's benefits branch (Belastingdienst/Toeslagen) wrongly accused an estimated 26,000 or more families of childcare-benefit fraud and demanded full repayment; broader advocacy estimates run higher and count different populations, and by February 2026 about 69,000 people had applied to the recovery scheme and more than 43,000 were formally recognized as affected, each entitled to a minimum of 30,000 euros. A self-learning risk-classification model that scored applications using a Dutch-nationality indicator, a 270,000-person fraud blacklist (the FSV) held without a legal basis, and an all-or-nothing recovery regime were coupled together; the Dutch Data Protection Authority imposed 6.45 million euros in fines (2.75 million for the nationality processing in 2021 and 3.7 million for the FSV blacklist in 2022), a parliamentary inquiry found rule-of-law violations, and the third Rutte cabinet resigned on 15 January 2021.

Sources: wikipedia2026, autoriteitpersoonsgegevens2021, autoriteitpersoonsgegevens2022, amnestyinternational2021b, tweedekamerderstatengeneraal2020, rijksoverheid2026

Appears on: /domains/cases/nl-toeslagenaffaire, /pan-lab

EmpiricalThe scandal's harm is best read as the coupling of three distinct components rather than a single algorithm. Government-…

The scandal's harm is best read as the coupling of three distinct components rather than a single algorithm. Government-commissioned technical reviews (KPMG in 2022 and PwC in 2023) described the tool as a self-learning classifier that routed the highest-scoring of roughly 90,000 benefit applications sent to manual treatment in 2014 to 2019, but judged the Dutch-nationality indicator's standalone predictive weight to have been limited; the model's precision and false-positive rate were never measured or published. The FSV fraud blacklist held frequently inaccurate data that was not corrected when people were cleared, and internal 2016 guidance auto-labelled childcare debts over 3,000 euros as intent or gross negligence, blocking payment arrangements. Out-of-home child placements are a documented but causally contested downstream harm: statistics counted roughly 2,090 children of affected parents placed out of home through mid-2022, while a 2025 judicial study found no child was removed solely because of financial problems.

Sources: kpmg2022, pwc2023, autoriteitpersoonsgegevens2022, statisticsnetherlandscbs2022, rechtspraak2025, wikipedia2026

Appears on: /domains/cases/nl-toeslagenaffaire, /pan-lab

EmpiricalOn 5 February 2020 the District Court of The Hague ruled that the legislation authorising SyRI, the Dutch state's secret…

On 5 February 2020 the District Court of The Hague ruled that the legislation authorising SyRI, the Dutch state's secret cross-database welfare-fraud risk-profiling system, violated Article 8 of the European Convention on Human Rights, and it ordered the system's use stopped; the State did not appeal. The ruling is widely described as one of the first times a court anywhere halted a digital welfare-fraud technology on human-rights grounds. Across its two executed neighbourhood projects SyRI was reported to have produced no confirmed fraud cases, and in one municipality 62 of 113 risk notifications were reported to be false positives.

Sources: districtcourtofthehague2020, vanbekkum2021, unofficeofthehighcommissione2020, algorithmwatch2020a, pontdataprivacyprivacywebnl2019, publicinterestlitigationproj2020

Appears on: /domains/cases/nl-syri

EmpiricalThe District Court of The Hague found that the SyRI framework provided no duty to notify people that their data had been…

The District Court of The Hague found that the SyRI framework provided no duty to notify people that their data had been processed or that a risk report had been filed, so a flagged person generally could not know about, access, or contest the notification; notifications were retained in a register for up to two years. The court held that a risk notification carried significant effect for the person even though it lacked formal legal effect, and it faulted the scheme for a lack of transparency and for breaching data-minimisation and purpose-limitation principles.

Sources: districtcourtofthehague2020, vanbekkum2021

Appears on: /domains/cases/nl-syri

EmpiricalFrance's family-benefits fund (CNAF) computes a monthly benefit-fraud suspicion score, on a 0-to-1 scale, for every bene…

France's family-benefits fund (CNAF) computes a monthly benefit-fraud suspicion score, on a 0-to-1 scale, for every benefit-receiving household — analysing the data of about 32 million people and producing more than 13 million scores each month, close to half of France's population; the highest scores route households into fraud controls, up to the most invasive on-site checks. An analysis by Le Monde and Lighthouse Reports of an extracted production model (a logistic regression of about 33 variables) found that markers of economic vulnerability raised the score: a stable-income family averaged about 0.33, while a person working while receiving the disability allowance (AAH) averaged about 0.66. The model's target was an overpayment (indu) above a threshold, which is frequently unintentional administrative error rather than proven intentional fraud, and the score itself is not disclosed to the person and cannot be appealed directly. CNAF disputed the discrimination framing, describing the tool as a neutral decision-aid that only prioritises which files to check; a coalition that grew to 25 organisations challenged the model before the Conseil d'État, and as of this writing no court had ruled.

Sources: lighthousereports2023b, lighthousereports2023a, laquadraturedunet2023, laquadraturedunet2026a, amnestyinternational2024b, generationnt2026

Appears on: /domains/cases/france-cnaf

EmpiricalIn an internal simulation study by CNAF's own statistics department (DSER), reported in October 2025 by Le Monde and La …

In an internal simulation study by CNAF's own statistics department (DSER), reported in October 2025 by Le Monde and La Quadrature du Net, recipients of the RSA minimum-income benefit were about 13% of beneficiaries but 39 to 41% of the highest-scoring 5%, and single mothers were about 14% of beneficiaries but 37 to 40% of that top bracket; households including a foreign national scored higher on average even after the nationality variable was removed. The full study is not public, and false-positive rates by protected group have not been released. The French ombudsperson (Défenseur des droits) told the Conseil d'État that a presumption of indirect discrimination appeared established because the differential treatment rests on beneficiaries' economic vulnerability; CNAF disputed the characterisation, and no court had ruled.

Sources: laquadraturedunet2026b, generationnt2026, laquadraturedunet2026a

Appears on: /domains/cases/france-cnaf

EmpiricalAnalysing the Swedish Social Insurance Agency (Forsakringskassan) 2017 outcome data, Lighthouse Reports and Svenska Dagb…

Analysing the Swedish Social Insurance Agency (Forsakringskassan) 2017 outcome data, Lighthouse Reports and Svenska Dagbladet reported on 27 November 2024 that the agency's in-house machine-learning risk profile for the temporary parental allowance (VAB) selected women (more than 1.5x), people of a foreign background (about 2.5x), below-median earners (2.97x), and people without a university degree (3.31x) for fraud investigation more often than comparison groups by demographic parity, and wrongly flagged those groups at higher false-positive rates (about 1.7x for women and 2.4x for people of a foreign background); in the agency's paired random-control sample, 20.2 percent of applications contained at least one day incorrectly paid, an unbiased base error rate. These are outcome computations under specific fairness definitions from a single obtained year of data, not confirmed model internals; the agency disputed the framing and did not release the model. The data-protection regulator IMY closed its GDPR supervision on 18 November 2025 for mootness after the agency withdrew the system, and no court or regulator issued a discrimination or GDPR penalty.

Sources: lighthousereports2024, lighthousereportsb, lighthousereportsc, integritetsskyddsmyndigheten2025a, integritetsskyddsmyndigheten2025b

Appears on: /domains/cases/sweden-forsakringskassan

EmpiricalThe Swedish Social Insurance Agency (Forsakringskassan) did not disclose the machine-learning risk profile it used to se…

The Swedish Social Insurance Agency (Forsakringskassan) did not disclose the machine-learning risk profile it used to select temporary-parental-allowance recipients for fraud investigation: its algorithm class, features, and precision were never released, and the agency resisted freedom-of-information disclosure for roughly three years on fraud-prevention grounds. In 2018 the audit inspectorate ISF found the risk-based profiling substantially more accurate than alternative controls while warning that it raised legal-certainty and equal-treatment concerns, and cautioning that an accurate model can still be inequitable when two groups err equally but only one is followed up. Amnesty International reported that a former agency data protection officer warned in 2020 that the operation breached European data-protection rules. The system was decommissioned in 2025 during the regulator's supervision, before any court or regulator ruled on it.

Sources: lighthousereports2024, lighthousereportsb, inspektionenforsocialforsakr2018a, inspektionenforsocialforsakr2018b, amnestyinternational2024d

Appears on: /domains/cases/sweden-forsakringskassan

EmpiricalDenmark's Udbetaling Danmark (UDK), administered by ATP, runs a data-driven welfare-fraud operation that as of 2019 used…

Denmark's Udbetaling Danmark (UDK), administered by ATP, runs a data-driven welfare-fraud operation that as of 2019 used up to about 60 AI and machine-learning models to score benefit recipients into a 'wonderlist' of high-risk people, which a human control team filters into control cases for investigation. In UDK's own 2023 control statistics (three documented models), the 'Model Abroad' foreign-affiliation model sent 511 cases for control but recovered money in only 36 -- about 7%, with roughly nine in ten resulting in no further action -- and UDK confirmed that 54% of the 'Really Single' household-outlier cases its unit opened were in fact legitimate. Those 'revenue' outcomes conflate deliberate fraud with honest error, which UDK does not separate, so they are not pure fraud rates. Amnesty International characterised the system as mass surveillance and prohibited social scoring under the EU AI Act; UDK, ATP and the ministry (STAR) rejected that characterisation, the system was not suspended, and as of this writing no court had ruled.

Sources: amnestyinternationalalgorith2024, amnestyinternational2024a, fortuneeurope2024, bablai2024

Appears on: /domains/cases/denmark-udbetaling

EmpiricalUdbetaling Danmark's 'Joint Data Unit' merges and links the personal data of millions of residents from around ten natio…

Udbetaling Danmark's 'Joint Data Unit' merges and links the personal data of millions of residents from around ten national registers -- civil registration (CPR), buildings and dwellings (BBR), business, income, tax (R75), health, VAT, cash and sickness benefits, education grants and the motor-vehicle register -- alongside a 'Joint Data Unit Abroad' that pulls data from foreign authorities; in 2021 UDK paid about DKK 241 billion to roughly 2.4 million recipients. Amnesty International documents this as mass surveillance and argues the design carries a discrimination risk: 'Model Abroad' scores a relative strength of ties to non-EEA countries with citizenship as a direct input, and 'Really Single' treats statistically atypical households as suspicious. That harm is a design-level risk rather than a measured outcome, because UDK and ATP denied all requests for the demographic data needed to test the models for bias, so no disparate-impact figure exists in the record. Oversight is thin: the Danish Data Protection Authority (Datatilsynet) can generally act only on complaints (GDPR Art. 57) with no proactive power, and because flagged people rarely learn an algorithm selected them, complaints are rare. UDK rejects the discrimination-by-design and social-scoring findings; no court has ruled.

Sources: amnestyinternationalalgorith2024, amnestyinternationaldanmark2024, bablai2024

Appears on: /domains/cases/denmark-udbetaling

EmpiricalIn judgment STS 1119/2025 of 11 September 2025, the Third Section of Spain's Supreme Court (Sala de lo Contencioso-Admin…

In judgment STS 1119/2025 of 11 September 2025, the Third Section of Spain's Supreme Court (Sala de lo Contencioso-Administrativo) ordered the government to give the transparency foundation Civio access to the source code of BOSCO, the software that determines eligibility for the electricity social bonus (bono social electrico). Applying the Transparency Law (Ley 19/2013) together with Article 42 of the EU Charter of Fundamental Rights and Article 105.b of the Spanish Constitution, the Court held that access to public information is a constitutional right and that neither intellectual property nor national security is an automatic shield, dismissing the government's secrecy claims as a 'mere risk' of eventual harm to be assessed case by case under a proportionality test. Civio and legal commentators describe an 'error multiplier': because BOSCO decides automatically and gives no reasons, one systematic error can propagate to thousands of eligible people at once. As of May 2026, roughly eight months after the ruling, the source code had still not been delivered and Civio had filed for judicial enforcement.

Sources: consejogeneraldelpoderjudici2025, fundacionciudadanacivio2025b, fundacionciudadanacivio2025a, derechoadministrativoyurbani2025, fundacionciudadanacivio2025c, fundacionciudadanacivio2026, freesoftwarefoundationeurope2026

Appears on: /domains/cases/spain-bosco

EmpiricalThe transparency foundation Civio documented, by reconstructing BOSCO's behaviour from partial technical specifications …

The transparency foundation Civio documented, by reconstructing BOSCO's behaviour from partial technical specifications and functional test cases, two systematic ways the software denied the electricity social bonus to people who qualified: when a pensioner ticked the 'pensioner' box the application could return an 'imposibilidad de calculo' (impossibility of calculation) error and be rejected without properly evaluating income; and large families, entitled to the bonus regardless of income, were denied whenever a household member withheld authorization to consult income data, although income was not a regulatory requirement for that category. After a 2017-2018 overhaul required all beneficiaries to re-apply by 31 December 2018, enrollment fell from roughly 2.4 to 2.5 million under the prior scheme to 1,111,958 as of January 2019 (later cited around 1.5 million), against an estimated 4.5 to 5.5 million eligible people, and more than half a million applicants were rejected. No audited per-decision error rate is public, because the source code and verification-test results were withheld; one academic analysis records errors in both directions, but the documented net effect is under-inclusion.

Sources: fundacionciudadanacivio2019, algorithmwatchnicolaskayserb2019, freesoftwarefoundationeurope2026, xatakaenriqueperez2024, rebootdemocracyjoseluismarti2025

Appears on: /domains/cases/spain-bosco

EmpiricalSerbia's Social Card (Socijalna karta) registry, given a statutory basis by the Law on the Social Card in force from 1 M…

Serbia's Social Card (Socijalna karta) registry, given a statutory basis by the Law on the Social Card in force from 1 March 2022 and financed in part by an 82.6 million euro World Bank public-sector loan, cross-links roughly 130 to 135 categories of data from other state registers to verify social-assistance eligibility and flag suspected undeclared income or assets. After the law, named sources report the caseload falling by a range of tens of thousands: government figures cited by Amnesty International show about 35,000 fewer recipients by August 2023, A11 counts at least 44,000 people having lost assistance by early 2024, and the UN Working Group on Business and Human Rights reported over 60,000 without assistance by October 2025. These are largely net caseload declines rather than audited counts of system-caused removals, and the government attributes part of the fall to a stronger economy. Roma are reported among the most affected because informal earnings are misclassified as income, but the registry records no ethnicity, so this is inferred rather than officially disaggregated. As of the latest reporting, Constitutional Court, World Bank Inspection Panel, and UN scrutiny were pending or active, with no court or panel yet ordering changes.

Sources: ainitiativeforeconomicandsoc2024, amnestyinternational2023b, contextthomsonreutersfoundat2023, unworkinggrouponbusinessandh2025, worldbankinspectionpanel2024, chinaceeinstitute2024

Appears on: /domains/cases/serbia-social-card

EmpiricalUnder Serbia's Social Card system, a removed beneficiary has 15 days to appeal and must wait three months to reapply reg…

Under Serbia's Social Card system, a removed beneficiary has 15 days to appeal and must wait three months to reapply regardless of changed circumstances, and removal letters frequently reference only unspecified data from the electronic database. A11's Request for Inspection to the World Bank Inspection Panel alleges that, because the system is semi-automated, social workers cannot correct errors recorded in it. Over roughly two years the Ministry processed more than 100,000 notifications of suspected income or asset increases, while beneficiaries filed only 361 appeals against Centers for Social Work rulings; because the two figures cover different populations, the gap illustrates how rarely flags were contested rather than a measured appeal rate. Documented misclassifications include a one-off funeral donation read as income and long-scrapped cars still counted as assets.

Sources: amnestyinternational2023b, ainitiativeforeconomicandsoc2024, worldbankinspectionpanel2024, chinaceeinstitute2024, contextthomsonreutersfoundat2023

Appears on: /domains/cases/serbia-social-card

EmpiricalSamagra Vedika, an entity-resolution system built by the Telangana government, decided welfare eligibility by matching r…

Samagra Vedika, an entity-resolution system built by the Telangana government, decided welfare eligibility by matching residents across thirty-plus government databases into a consolidated profile; between 2014 and 2019 more than 1.86 million ration cards were cancelled and 142,086 fresh applications were rejected without notice. Its core error was entity-resolution false-positive matching, in which a similarly-named third party's asset was attributed to the applicant and silently flipped the eligibility flag. After the Supreme Court of India ordered field re-verification in April 2022, a partial re-verification found roughly 7.5 percent wrongful rejection (at least 15,471 approved of 205,734 re-processed cases), a lower bound from an incomplete review; the system is proprietary and closed and an independent technical audit could not be completed, with no source code or accuracy data released. The government cited a self-reported 95 percent fraud-filtering efficiency, which measures spurious-application filtering rather than the wrongful-exclusion rate.

Sources: amnestyinternational2024c, tapasya2024, tusharvsharma2026, sumitjha2024, kumarsambhav2020, pulitzercenteraiaccountabili2024

Appears on: /domains/cases/india-samagra-vedika

EmpiricalUnder Samagra Vedika, exclusions were silent and there was no statutory route to contest an algorithmic decision, so the…

Under Samagra Vedika, exclusions were silent and there was no statutory route to contest an algorithmic decision, so the burden of proof fell on the excluded person: reporting describes officials who, though formally able to override the algorithm with evidence, deferred to it and declined to overturn its verdict, treating errors as backend technical issues. Documented individual harms include a 67-year-old widow denied rations for more than seven years after the system linked her deceased husband to a car owned by a similarly-named third person, and a family rejected for allegedly owning a four-wheeler that was declared eligible only after a Telangana High Court ruling. Corrections came through individual litigation and did not systematically feed back into the model, and the same entity-resolution technology was reused to issue new ration cards in 2024-2025.

Sources: tapasya2024, thereporterscollective2024, tusharvsharma2026, amnestyinternational2024c, sumitjha2024

Appears on: /domains/cases/india-samagra-vedika

EmpiricalIn A.M.C. v. Smith (No. 3:20-cv-00240, M.D. Tenn.), a federal court held after a five-day bench trial that Tennessee's D…

In A.M.C. v. Smith (No. 3:20-cv-00240, M.D. Tenn.), a federal court held after a five-day bench trial that Tennessee's Deloitte-built TEDS automated Medicaid eligibility system, operational statewide since March 19, 2019 for a program covering roughly 1.7 million residents, produced wrongful terminations, wrong-household assignments, and misleading or missing notices that violated the Medicaid Act, the Fourteenth Amendment's Due Process Clause, and the Americans with Disabilities Act; the 116-page opinion, issued August 26, 2024 by Judge Waverly D. Crenshaw Jr., ordered mediation before considering an injunction.

Sources: statescoopkeelyquinlan2024, stotlerhayesgroupllcerinsail2024, georgetownuniversitycenterfo2024, nationalhealthlawprogram2024

Appears on: /domains/cases/tennessee-tenncare-teds

EmpiricalThe UK government built its own AI meeting scribe for council caseworkers and piloted it through a cohort of 25 selected…

The UK government built its own AI meeting scribe for council caseworkers and piloted it through a cohort of 25 selected councils (22 active, more than 400 users) under one shared pooled-assurance record, then open-sourced it and adapted it to enlist around 500 housing and homelessness workers by June 2026; the cohort published a multi-council governance dataset but no transcription-accuracy or error-rate evaluation, and standard risk controls such as penetration testing and certification had not been completed on the alpha at pilot time.

Sources: localgovernmentassociation2025b, localgovernmentassociation2025a, ministryofhousing2026, trendall2026, incubatorforartificialintell2026

Appears on: /domains/cases/minute-local-ai

EmpiricalIndependent research by the Ada Lovelace Institute on AI transcription in social work, based on interviews with 39 socia…

Independent research by the Ada Lovelace Institute on AI transcription in social work, based on interviews with 39 social workers across 17 local authorities in England and Scotland, reported that local authorities focus their evaluations on efficiency rather than impact on people who draw on care and that perceptions of reliability and the need for human oversight vary significantly among workers; the research covers such tools sector-wide, not this tool specifically.

Sources: adalovelaceinstitute2026b, bruff2026, adalovelaceinstitute2026a

Appears on: /domains/cases/minute-local-ai

EmpiricalThe Ministry of Justice built an in-house AI transcription and summarisation copilot, Justice Transcribe, for probation …

The Ministry of Justice built an in-house AI transcription and summarisation copilot, Justice Transcribe, for probation staff in England and Wales, scaling it from a pilot to more than 1,000 officers in October 2025 and to every probation officer by June 2026, with official transparency data recording more than 800,000 supervision meetings summarised between 7 October 2025 and 2 June 2026; the reported time-savings are the ministry's own and rest on an operating assumption the department itself labels illustrative, and no transcription-accuracy rate, officer correction rate, or independent evaluation of the tool has been published.

Sources: justiceaiunit2026, ministryofjustice2025, ministryofjusticeandhmprison2025b, ministryofjusticeanddsit2025, ministryofjustice2026

Appears on: /domains/cases/justice-transcribe-probation

EmpiricalCopilot-written probation case records sit upstream of high-volume algorithmic risk assessment over the same record ecos…

Copilot-written probation case records sit upstream of high-volume algorithmic risk assessment over the same record ecosystem: reporting places the ministry's OASys-based reoffending-risk prediction at more than 1,300 people a day, drawing on probation and prison caseload systems and the Police National Computer, with a successor tool rolling out during 2026, and the ministry's own validation found lower predictive validity for all Black, Asian and Minority Ethnic groups for non-violent reoffending and for Black and Mixed ethnicity offenders for violent reoffending - a property of the downstream risk model, not the copilot; peer-reviewed commentary raises the erosion of professional judgment and the unresolved accountability for algorithm-influenced decisions as structural concerns, and no published source documents a named data pipeline from the copilot's output into the risk tools.

Sources: statewatch2025, phillips2026, nellis2026

Appears on: /domains/cases/justice-transcribe-probation

EmpiricalThe US Social Security Administration requires decision writers to run fully favorable disability decisions through its …

The US Social Security Administration requires decision writers to run fully favorable disability decisions through its in-house Insight verifier before issuance, with narrow documented exceptions, and the 2025 federal AI inventory records the tool computing 43 quality flags. In the agency's internal five-month study of roughly 50,000 appeals-level cases, reported through the 2019 Inspector General audit, analysts who used Insight logged about 0.9 errors per case against 0.7 for non-users, saw processing time fall about 4.7 days per case, and had about 12.6 percent of their cases returned for quality issues against 21.5 percent for non-users. These are internal, non-randomized comparisons among self-selected voluntary users, and the same audit found the agency stopped tracking performance after the first five months and could not determine any effect on remands.

Sources: ussocialsecurityadministrati2019, ussocialsecurityadministrati2026, engstrom2020

Appears on: /domains/cases/ssa-insight

EmpiricalIn a review issued April 30, 2026 (report 25-00153-47), the Department of Veterans Affairs Office of Inspector General f…

In a review issued April 30, 2026 (report 25-00153-47), the Department of Veterans Affairs Office of Inspector General found that at least 8,000 of an estimated 8,100 automated Dependency and Indemnity Compensation (survivor-benefit) granting decisions issued from September 2023 through August 2024 - nearly all - contained at least one legal or procedural deficiency, such as incomplete evidence summaries and omitted favorable findings, with most rating decisions listing only the death certificate as evidence. The OIG separately found that at least 2 percent of the decisions (at least 190) carried monetary-impact legal errors totaling at least 2.7 million dollars (2,727,764 dollars in questioned costs); the roughly 98 percent figure is the share with any legal or procedural defect, not the monetary-error rate. The system, phased in beginning May 2020, extracts data from scanned documents and applies predefined encoded rules to grant service-connected death claims end to end with no human involvement when the rules are met; the OIG describes it as rules-based automation and document extraction, not machine learning, and its figures are outcome statistics from a statistical sample rather than a per-interaction rate.

Sources: departmentofveteransaffairso2026, nieberg2026, weston2026

Appears on: /domains/cases/va-claims-automation

EmpiricalThe Office of Inspector General reported that VA's internal correction channels did not catch the automated survivor-ben…

The Office of Inspector General reported that VA's internal correction channels did not catch the automated survivor-benefit deficiencies and that the external audit was, empirically, the only channel that changed behavior. In April 2020 a VBA analyst reported through the internal defect-tracking system that automated decisions listed only the death certificate as evidence, and the Pension and Fiduciary Service closed the defect without action; the same deficiency was central to the 2026 findings, and VA removed the long-form guidance from its manual only in March 2025, immediately after the OIG's preliminary briefing - roughly five years later, and the OIG's full public report did not follow until 2026, roughly six years after the ticket. The OIG found the quality-review checklist for automated claims was less rigorous than the review traditional claims receive, and that the PACT Act section 701(b) modernization plan to Congress did not fully disclose that VBA grants these claims end to end without human intervention. Errors persisted as the program expanded: the VA Secretary announced expanded DIC automation in May 2025, and 20 additional automated decisions from September and October 2025 showed similar errors as of November 2025, with one recommendation still open and VBA concurring only in part.

Sources: departmentofveteransaffairso2026

Appears on: /domains/cases/va-claims-automation

EmpiricalIn Trelleborg, Sweden, the first municipality to fully automate social-assistance decisions, peer-reviewed analysis repo…

In Trelleborg, Sweden, the first municipality to fully automate social-assistance decisions, peer-reviewed analysis reports that about 30 percent of digital reapplications are decided entirely by rules-based software with no human review and about 85 percent receive at least partial automated handling; decision time on reapplications fell from roughly two days to under a minute, and a human caseworker re-enters the path only by exception, when a routing rule detects significantly changed circumstances, a missing activity plan or job-seeking documentation, or a complex or negative case. No error, override, exception-routing, or appeal-rate figures for the automated path have been published, so the fraction of automated decisions that ever reaches a human cannot be established from the record.

Sources: algorithmwatch2020b, ranerupandhenriksen2022, europeancommissionjointresea2021

Appears on: /domains/cases/trelleborg-rpa

EmpiricalThe City of Amsterdam spent roughly five years and an estimated EUR 535,000 building a deliberately fair, explainable we…

The City of Amsterdam spent roughly five years and an estimated EUR 535,000 building a deliberately fair, explainable welfare-fraud screening model with nearly every recommended pre-deployment safeguard in place - a bias audit, training-data reweighting that approximately equalized wrongful-flag rates on retrospective data, a data-protection assessment and a human-rights assessment, external and academic review, a citizen panel, and dual algorithm-register transparency - and discontinued it after a 2023 live pilot on nearly 1,600 applications. In the investigating journalists' analysis of aggregate data the city provided, the group disparities re-emerged inverted on the live pilot, now more likely to wrongly flag Dutch nationals, women, and applicants with children, with the tool flagging more applications than the analog process and no better than caseworkers at finding genuine cases. The Dutch national algorithm register records the deployment ending September 2023 and lists it out of use, and the responsible alderman announced the halt in November 2023.

Sources: braun2025, lighthousereportsa, algoritmeregisterdutchnation2023

Appears on: /domains/cases/amsterdam-slimme-check

EmpiricalIn a developer-reported randomised controlled trial of more than 1,000 adviser support requests, an adviser-facing benef…

In a developer-reported randomised controlled trial of more than 1,000 adviser support requests, an adviser-facing benefits copilot at Citizens Advice returned supervisor-checked answers in about four minutes, roughly half the previous response time, with about 80 percent of its drafts approved by supervisors without revision; these figures are reported by the tool's builders and have not been independently replicated.

Sources: varotsis2025, departmentforscience2025a, stanfordlegaldesignlabjustic

Appears on: /domains/cases/caddy-citizens-advice, /what-ai-can-do

EmpiricalAdvisers given access to the copilot were reported to be more than twice as likely to say they felt confident giving adv…

Advisers given access to the copilot were reported to be more than twice as likely to say they felt confident giving advice than a control group, a self-reported measure from post-call in-chat surveys rather than a client-outcome or accuracy measure.

Sources: varotsis2025, stanfordlegaldesignlab2025

Appears on: /domains/cases/caddy-citizens-advice, /what-ai-can-do

EmpiricalIn 2023 Singapore's GovTech began a whole-of-government retire-and-replace of its scripted Ask Jamie chatbots, embedded …

In 2023 Singapore's GovTech began a whole-of-government retire-and-replace of its scripted Ask Jamie chatbots, embedded since 2014 on 70-plus (a vendor case study claims 80) agency websites as independent per-agency answer engines, migrating government chatbots onto centrally provided large-language-model engines; the stated aim was to convert all 88 chatbots and retire the scripted engine by end 2023, the verified snapshot is 21 of 88 converted as of September 2023 (migration completion not independently documented), and by the VICA product page updated 29 April 2026 the successor platform hosts over 100 chatbots for 60-plus agencies at an average of over 800,000 monthly queries, figures that are all government self-reported.

Sources: hirdaramani2023, govtechsingapore2026, govtechsingapore2019

Appears on: /domains/cases/singapore-chatbot-fleet-refresh

EmpiricalThe Singapore government benefits-navigation surface is documented as scope-limited to information and estimates rather …

The Singapore government benefits-navigation surface is documented as scope-limited to information and estimates rather than adjudication: the Ministry of Finance Support For You Calculator turns self-declared inputs into benefit estimates that are explicitly estimates and not entitlement decisions, and the Chat.Gov.SG (Beta) explainer hosted on the SupportGoWhere domain states the assistant summarises information from official government websites and does not assess eligibility, make decisions, submit applications, or complete transactions, and warns users not to share personal or sensitive information.

Sources: publicservicedivisionsingapo2026, mustsharenews2024

Appears on: /domains/cases/singapore-chatbot-fleet-refresh

EmpiricalA June 2026 Treasury Inspector General for Tax Administration performance audit (Report Number 2026-308-029) reported th…

A June 2026 Treasury Inspector General for Tax Administration performance audit (Report Number 2026-308-029) reported that the IRS expanded its Automated Collection System chatbot and live-chat program and made live chat permanent while having no performance measures for it, despite a Taxpayer First Act requirement for metrics and benchmarks, and that management's claim the bots reduced telephone demand could not be substantiated; the statistical reports the IRS did collect were deemed unreliable, in one instance showing a single assistor apparently working 603 chats at once against a systemic cap of three, attributed partly to a miscalculated handle-time metric the vendor had not resolved as of December 2025.

Sources: treasuryinspectorgeneralfort2026, bracken2026, bramwell2026

Appears on: /domains/cases/irs-acs-chatbots

EmpiricalIn the same audit, of a judgmental sample of 40 IRS ACS live assistors, 24 (60%) were found working multiple chats concu…

In the same audit, of a judgmental sample of 40 IRS ACS live assistors, 24 (60%) were found working multiple chats concurrently and 12 of those 24 had at least one authenticated chat open while working another, which TIGTA reported as raising the risk of disclosing taxpayer information to the wrong taxpayer; the audit also reported 635,684 resolution codes against 613,056 chats (a mismatch management knew of but did not investigate) and, in March 2025 hand-testing, 14% of chatbot process flows deficient and 83% of tested keywords unrecognized or insufficient, with the figures drawn from a nonprobability sample and data the audit itself characterized as unreliable and not projectable to the full assistor population.

Sources: treasuryinspectorgeneralfort2026, bramwell2026, cohn2026

Appears on: /domains/cases/irs-acs-chatbots

EmpiricalIn a spring-2025 randomized pilot inside a large consumer EBT app, the vendor reports that 53% of eligible SNAP recipien…

In a spring-2025 randomized pilot inside a large consumer EBT app, the vendor reports that 53% of eligible SNAP recipients took up in-app AI help for missed deposits and that treated users were restored faster and more often in the same month than a control group, with every AI dead-end escalated to a named human; all outcome figures are vendor-published and the effect magnitudes were not disclosed.

Sources: propelincpropelinsights2025b, guarino2025a

Appears on: /domains/cases/propel-snap-assistant

EmpiricalBy the vendor's own account of the design, the assistant grounds on a state-verified deposit record it reads but does no…

By the vendor's own account of the design, the assistant grounds on a state-verified deposit record it reads but does not write to, and steers recipients to act on the state system of record rather than acting for them.

Sources: propelincpropelinsights2025b

Appears on: /domains/cases/propel-snap-assistant

EmpiricalBetween 2019 and 2025 more than 70% (about 73% per its ten-year retrospective) of California's online SNAP applications …

Between 2019 and 2025 more than 70% (about 73% per its ten-year retrospective) of California's online SNAP applications were submitted through GetCalFresh, a deterministic, structured-workflow application assister built and operated by the nonprofit Code for America, which reports helping 6.2 million people obtain more than $12.8 billion in food benefits from 2017 to 2025 (organization-published figures that are not independently audited); the node made no eligibility determinations, and in 2024 and 2025 the California Department of Social Services coordinated a dated, phased transfer of its functions into the state-owned BenefitsCal portal.

Sources: codeforamerica2024b, codeforamerica2025a, californiadepartmentofsocial2025

Appears on: /domains/cases/getcalfresh, /what-ai-can-do

EmpiricalA randomized controlled trial of roughly 65,000 Los Angeles GetCalFresh applicants (Giannella, Homonoff, Rino, and Somer…

A randomized controlled trial of roughly 65,000 Los Angeles GetCalFresh applicants (Giannella, Homonoff, Rino, and Somerville, American Economic Journal: Economic Policy 16(4), 2024) found that access to applicant-initiated flexible interviews increased SNAP approvals by about 6 percentage points, doubled early approvals, and raised long-term participation by over 2 percentage points, identifying the intake interview as a key procedural-denial barrier; Code for America separately reported an in-house experiment lifting renewal-form submissions among prior non-responders from about 1.5% to roughly 12% (organization-published, without sample sizes or confidence intervals).

Sources: giannella2024, codeforamerica2021, codeforamerica2024a

Appears on: /domains/cases/getcalfresh

EmpiricalIn June 2024 the board of Benefits Data Trust, a Philadelphia benefits-navigation nonprofit that reported helping more t…

In June 2024 the board of Benefits Data Trust, a Philadelphia benefits-navigation nonprofit that reported helping more than 120,000 people access about $182 million in benefits in 2023, voted unanimously to wind the organization down within a self-imposed 60-day window, citing only 'a perfect storm of circumstances'; the organization closed on August 24, 2024, laying off 273 employees, despite roughly $12 million in unrestricted reserves at the end of 2023 and about $32 million in projected 2024 revenue.

Sources: brubaker2024a, brubaker2024b, wink2024, mosbruckergarza2024

Appears on: /domains/cases/benefits-data-trust-winddown

EmpiricalThe closure left active government partnerships without a designated successor, including a Pennsylvania Department of A…

The closure left active government partnerships without a designated successor, including a Pennsylvania Department of Aging workload of nearly 48,000 applications from 27,018 households in the final year and a Philadelphia BenePhilly call-center contract the organization was reported to be exceeding through mid-2024; the navigation function fragmented to higher-friction channels, with the work redistributed across partner agencies and a subcontractor and referral waits reported as several months, which a Pew analyst described as a 'cascading effect.'

Sources: brubaker2024d, burnley2024, mosbruckergarza2024

Appears on: /domains/cases/benefits-data-trust-winddown

EmpiricalLondon's Strategic Insights Tool for Rough Sleeping probabilistically links records from three separately governed syste…

London's Strategic Insights Tool for Rough Sleeping probabilistically links records from three separately governed systems - CHAIN street-outreach contacts, In-Form charity casework, and H-CLIC borough statutory applications - into a single rough-sleeping journey per person that is read, in aggregate form only, across all 33 London local authorities; the tool makes no individual-level determinations, and after the build vendor's data-processor contract ended on 2 February 2024 the Greater London Authority contracted Homeless Link, which also operates the CHAIN source system, for its ongoing hosting, management, and maintenance.

Sources: techuk2024, loti2023, londonofficeoftechnologyandi2023

Appears on: /domains/cases/london-rough-sleeping-sit

EmpiricalThe Strategic Insights Tool's matcher accepts an association only above an 85% probability threshold chosen to minimise …

The Strategic Insights Tool's matcher accepts an association only above an 85% probability threshold chosen to minimise false positives, and the project's own Phase 2 Data Protection Impact Assessment reports 91% recall - conceding that roughly 9 in 100 true cross-system matches are missed so that 'numbers subsequently appear lower in places where they should be higher' and that recall varies as new data of varying quality is ingested; no false-positive rate is published, the accuracy figures are self-reported by the delivery team, and no independent evaluation of the tool's decision impact exists.

Sources: loti2023, lotiannahumplebyandfacultyja2025

Appears on: /domains/cases/london-rough-sleeping-sit

EmpiricalSan Jose's vehicle-mounted computer-vision pilot, described by city officials and national housing advocates as the firs…

San Jose's vehicle-mounted computer-vision pilot, described by city officials and national housing advocates as the first US experiment training AI to recognize tents and lived-in vehicles, reported sharply class-asymmetric accuracy in the city's own staff-ground-truthed evaluation — 97% for potholes and 88% for trash, but only 70% for RVs (unable to distinguish a lived-in RV from an empty one) and 12.5% for lived-in vehicles, with a March 2024 official interview bracketing the habitation figures at 70–75% for RVs and 10–15% for lived-in cars against a 70% goal; no detection ever generated an operational dispatch, and after investigative exposure and structured engagement the city removed every habitation-detection use case, its March 2025 status report declining to recommend implementing AI object detection in city operations at this time.

Sources: feathers2024, cityofsanjoseinformationtech2025, usdepartmentoftransportation2025

Appears on: /domains/cases/san-jose-encampment-detection

EmpiricalThe pilot's published data-usage protocol declares that the footage cannot be actively monitored for law-enforcement pur…

The pilot's published data-usage protocol declares that the footage cannot be actively monitored for law-enforcement purposes while preserving a police request path to it — verbatim, 'Law enforcement may request access to previously stored footage. Law enforcement is not actively monitoring any data collected' — and requires de-identification or deletion within one month; the CIO stated data was not shared with police during the pilot, yet public-records reporting documented that one vendor's system ran optical character recognition of license plate numbers despite the city's no-identification claim, so the declared authority rule and the feasible data flows diverged, a gap surfaced by journalists rather than by any standing audit, and the no-law-enforcement-use clause is city protocol language rather than statute.

Sources: feathers2024, cityofsanjoseinformationtech2024, varian2024

Appears on: /domains/cases/san-jose-encampment-detection

EmpiricalThe Los Angeles Coordinated Entry System replaced the VI-SPDAT survey for single adults with the Los Angeles Housing Ass…

The Los Angeles Coordinated Entry System replaced the VI-SPDAT survey for single adults with the Los Angeles Housing Assessment Tool, a 19-item self-report score whose weights were derived by a regression on 71,747 historical assessments linked to county records; where the CESTTRR research estimated the VI-SPDAT scored near chance (AUC 0.54) with racial false-negative gaps up to 8.5 percentage points, the equity-adjusted successor was deliberately traded down in overall accuracy (AUC 0.60, from an accuracy-only 0.64) to close those gaps to under one percentage point, and every such figure is a pre-deployment estimate on 2015 to 2018 held-out data rather than an observed post-launch outcome.

Sources: rice2023, losangeleshomelessservicesau2025a

Appears on: /domains/cases/lahsa-triage-revision

EmpiricalDuring the dual-tool transition the two instruments' PSH-consideration thresholds were 8-plus on the VI-SPDAT and 17-plu…

During the dual-tool transition the two instruments' PSH-consideration thresholds were 8-plus on the VI-SPDAT and 17-plus on the LA HAT, and by LAHSA's account initial quantitative data and provider feedback showed participants were more likely to obtain an eligible score under the VI-SPDAT, so direct-service providers opted to administer it, a trend LAHSA states 'perpetuated the racial bias of the VI-SPDAT in the System'; on April 22, 2026 the CES Policy Council lowered the LA HAT threshold to 12-plus, ruled the most recent LA HAT score supersedes a coexisting VI-SPDAT score, and forced deactivation of new VI-SPDAT completions (for LA HAT-access programs on May 1, 2026 and system-wide on June 30, 2026), though LAHSA has not released the underlying eligibility-rate numbers.

Sources: losangeleshomelessservicesau2025a, losangeleshomelessservicesau2026b, losangeleshomelessservicesau2026a

Appears on: /domains/cases/lahsa-triage-revision

EmpiricalIn a registered randomized controlled trial of 1,263 imminent-risk applicants (514 treatment, 749 control) run by the Un…

In a registered randomized controlled trial of 1,263 imminent-risk applicants (514 treatment, 749 control) run by the University of Notre Dame's evaluation lab, households offered flexible emergency financial assistance averaging about 2,000 dollars, typically one to two months of back rent, through Santa Clara County's homelessness-prevention system were reported 81 percent less likely to become homeless within six months and 73 percent within twelve months; the peer-reviewed article's abstract states the assistance reduced homelessness by 3.8 percentage points from a 4.1 percent base rate, and the researchers conservatively estimated 2.47 dollars in community benefits per net dollar spent.

Sources: phillipsandsullivan2025, universityofnotredamenews2023, phillipsandsullivan2021

Appears on: /domains/cases/santa-clara-prevention

EmpiricalBecause becoming homeless is statistically rare even among at-risk applicants - about 96 percent of the trial's control …

Because becoming homeless is statistically rare even among at-risk applicants - about 96 percent of the trial's control group never became homeless without assistance - the program's own co-author cautions that prevention resources can flow to households that would have stayed housed anyway, making screening precision on a low base rate the binding constraint; as of February 2026 the model is being replicated across about ten heterogeneous US jurisdictions under a 77-million-dollar initiative, with the same evaluation lab as the common evidence partner assessing each site.

Sources: kendall2026a, phillipsandsullivan2025, destinationhome2026b, universityofnotredamenews2026

Appears on: /domains/cases/santa-clara-prevention

EmpiricalAt the Calgary Drop-In Centre, a University of Calgary engineering group and the NGO shelter operator built deliberately…

At the Calgary Drop-In Centre, a University of Calgary engineering group and the NGO shelter operator built deliberately interpretable screening for chronic and episodic shelter use - explicit stay-count thresholds (for example 81 or more stays in a 90-day window) and database-queryable rules derived from the shelter's own administrative records, reported to flag candidate clients at a median of about 98 days versus 285 days under the Government of Canada definition and 365 under the Alberta definition - and, rather than surface a risk score, deployed a co-designed data-navigation interface that shows frontline staff raw client histories; no fetched source confirms the thresholds running as an automated production screener, and the deployed, studied artifact is the raw-history interface.

Sources: messier2021, arulesearchframeworkfortheea2022, masrani2025

Appears on: /domains/cases/calgary-drop-in-shelter-ml

EmpiricalAcross a 2022 to 2024 embedded deployment study of the interface (16 staff across 7 role categories; 29.5 hours of quali…

Across a 2022 to 2024 embedded deployment study of the interface (16 staff across 7 role categories; 29.5 hours of qualitative data; five committee observations; three deployed versions), the participant-research team documented a stakes-dependent 'data-outsourcing continuum': staff were reluctant to outsource high-stakes barring decisions, treating the data as a starting point for collaborative discussion, while reporting more willingness to accept automated data-driven recommendations for lower-stakes housing triage; the finding is the staff's own articulated practice rather than a measured override or agreement rate, all deployment evidence is authored by the embedded research team, and no independent evaluation, usage logs, or decision volumes are published.

Sources: masrani2025, thehumanbehindthedatareflect2023

Appears on: /domains/cases/calgary-drop-in-shelter-ml

EmpiricalAn AI quality-assurance tool deployed on a national 988 backup line scores crisis counselors' own call practice rather t…

An AI quality-assurance tool deployed on a national 988 backup line scores crisis counselors' own call practice rather than callers, expanding measured review from the under-3% of calls that had been reviewed by hand toward nearly all of them; a peer-reviewed reliability study of 476 labeled calls reported agreement with human ratings at 98 percent of human interrater agreement for detecting any risk assessment, with average F1 of about 0.86 at call level and 0.66 at statement level, and its authors include four holders of equity in the vendor.

Sources: imel2024, aguilar2023, nihreporternationalinstitute2025

Appears on: /domains/cases/lyssn-protocall-988

EmpiricalThe registered randomized crossover trial of the tool's counselor feedback (81 call-takers) completed on October 31, 202…

The registered randomized crossover trial of the tool's counselor feedback (81 call-takers) completed on October 31, 2025, but as of mid-2026 no results were posted to the trial registry or found in the peer-reviewed literature and participant-level data were marked unavailable for proprietary reasons, so reported counselor-skill-improvement effects remain vendor claims pending independent publication.

Sources: clinicaltrialsgovusnationall2026, lyssn2026

Appears on: /domains/cases/lyssn-protocall-988

EmpiricalGaggle's student-communication safety monitoring, used by roughly 1,500 US districts covering about 6 million students a…

Gaggle's student-communication safety monitoring, used by roughly 1,500 US districts covering about 6 million students as of a March 2025 AP and Seattle Times investigation, scans school-issued accounts around the clock and routes flags through a multi-hop chain (a machine flag, an off-site vendor reviewer, district safety staff, and, for imminent-danger after-hours alerts, occasional police welfare checks); in Vancouver Public Schools nearly 2,200 students (about 10% of enrollment) triggered alerts in one year, in Lawrence USD 497 more than 1,200 incidents were logged in ten months with about two-thirds deemed nonissues by officials (a figure the plaintiffs drew from district records), and the archive of flagged documents was accidentally released to reporters as nearly 3,500 unredacted files through unprotected links, while a 2023 RAND review found only scant evidence of either benefit or risk and the vendor publishes no accuracy figures.

Sources: bryanandlurye2025, associatedpress2025, lawrencejournalworld2025

Appears on: /domains/cases/gaggle-school-monitoring

EmpiricalAfter nine Lawrence, Kansas students sued their district in early August 2025 over its use of AI communication monitorin…

After nine Lawrence, Kansas students sued their district in early August 2025 over its use of AI communication monitoring, court filings revealed the district had ceased using Gaggle mid-litigation and substituted a different monitoring vendor with no board vote or public disclosure — surfacing only as a line in a check register — and the plaintiffs' amended complaint argued the swap does not moot the case because the core practice of suspicionless scanning, flagging, and seizure of student speech continues; on April 10, 2026 a federal judge found the district violated the Kansas Open Records Act in withholding the substitution and phase-out records, and on June 4, 2026 ordered it to pay the students' attorney fees, characterizing the conduct as drawn out, hollow and perplexing, with a jury trial on the surviving constitutional claims set for January 2027.

Sources: heimsoth2025, heimsoth2026a, heimsoth2026b

Appears on: /domains/cases/gaggle-school-monitoring

EmpiricalIn March 2026 the UK Parliamentary and Health Service Ombudsman partly upheld a complaint that an NHS mental health trus…

In March 2026 the UK Parliamentary and Health Service Ombudsman partly upheld a complaint that an NHS mental health trust installed camera-based, contact-free bedroom monitoring on a psychiatric ward without seeking a patient's consent, gave her no information about it, and did not switch it off when she asked; the case documentation and investigative reporting describe an internal clinical evaluation that the vendor is reported to have authored the business case for and shaped, a rebrand of the vendor during a statutory inquiry, and an open data-protection investigation, while the tool's own outcome-reduction figures are vendor claims contested by a campaign-linked meta-analysis and its adoption share across NHS mental health trusts is reported only as a contested range.

Sources: parliamentaryandhealthservic2026, williamson2026a, williamson2026b, nationalsurvivorusernetwork2025, stopoxevision2026

Appears on: /domains/cases/oxevision-nhs-wards

EmpiricalThe ombudsman's report on the case (decision 27 March 2026) found the trust did not seek or revisit the patient's consen…

The ombudsman's report on the case (decision 27 March 2026) found the trust did not seek or revisit the patient's consent for the bedroom monitoring, did not turn the camera off when she asked, gave her no information about it, and kept no record of how staff used it, and that even the trust's revised 2025 procedure still permits overriding a capacitous patient's refusal on clinically-safe grounds with multidisciplinary-team approval; on the separate question of over-reliance the ombudsman found on balance, cross-referencing observation charts, a nurse-adviser review and door key-card data, that in-person observations had continued and did not uphold that part of the complaint.

Sources: parliamentaryandhealthservic2026

Appears on: /domains/cases/oxevision-nhs-wards

EmpiricalIn 2025 ODMAP's pre-set county thresholds - a rolling 24-hour count against a threshold each agency sets or accepts, rec…

In 2025 ODMAP's pre-set county thresholds - a rolling 24-hour count against a threshold each agency sets or accepts, recommended by the system as two standard deviations above the county's own previous 90-day mean, a deterministic rule rather than a machine-learning model - fired 74,805 advisory spike-alert notifications from 498,003 suspected, unconfirmed overdose events that only about 1,362 of its 5,605 approved agencies actually submitted; these figures are self-published by the program in its own annual report and manuals, and ODMAP states its data are suspected, incomplete, not a system of record, and should not be generalized beyond participating agencies.

Sources: washingtonbaltimorehidta2025b, washingtonbaltimorehidta2026c, washingtonbaltimorehidta2025a

Appears on: /domains/cases/odmap-overdose-spike-alerts

EmpiricalODMAP's shared overdose store is housed inside a federal drug-enforcement program, and its operating policies both state…

ODMAP's shared overdose store is housed inside a federal drug-enforcement program, and its operating policies both state that ODMAP is neither an intelligence sharing database nor a pointer index records system and grant the host permission to use the data as the HIDTA sees fit, including combining it with other databases it manages for law enforcement and public health products; a 2024 peer-reviewed stakeholder study documented divergent public-health versus public-safety data-privacy standards, and a 2025 peer-reviewed analysis argues the integration risks racialized surveillance and criminalization of people who experience overdose, a contested scholarly critique of the link structure rather than a documented misuse incident.

Sources: washingtonbaltimorehidta2022, syvertsen2025, allen2024

Appears on: /domains/cases/odmap-overdose-spike-alerts

EmpiricalThe Targeted Real-Time Early Warning System (TREWS), a machine-learning sepsis early-warning model, was evaluated prospe…

The Targeted Real-Time Early Warning System (TREWS), a machine-learning sepsis early-warning model, was evaluated prospectively across five hospitals of an academic health system covering 590,736 monitored patients — the largest prospective study of an ML sepsis system on record. Its central finding was conditional on the human loop: sepsis patients whose alert was evaluated and confirmed by a provider within three hours had a 3.3 percentage-point absolute and 18.7 percent relative adjusted reduction in in-hospital mortality, with less organ failure and shorter stays, while the alert on its own did not; a companion study found provider uptake varied with experience, unit culture, and alert context.

Sources: adams2022a, henry2022a

Appears on: /domains/cases/johns-hopkins-trews, /pan-lab

EmpiricalThe TREWS mortality-benefit evaluation was prospective and peer-reviewed but observational and developer-led: it was bui…

The TREWS mortality-benefit evaluation was prospective and peer-reviewed but observational and developer-led: it was built at the deploying institution and commercialized through a company founded by its principal investigator, and confirmation-associated benefit is an observational association rather than a randomized effect of the algorithm — providers who engaged with alerts may differ from those who did not in ways the adjustment does not capture. The strongest numbers in the record therefore come from the party with the strongest interest in them, and no independent replication of the mortality effect had been published.

Sources: adams2022a

Appears on: /domains/cases/johns-hopkins-trews, /pan-lab

EmpiricalThe Advance Alert Monitor is an in-hospital deterioration model running around the clock across 21 hospitals of an integ…

The Advance Alert Monitor is an in-hospital deterioration model running around the clock across 21 hospitals of an integrated health system, scoring inpatients hourly and firing roughly twelve hours before predicted deterioration; a 2020 New England Journal of Medicine evaluation associated its alert-driven rapid-response workflow with lower mortality. Its defining feature is where the alert goes: not to the bedside, but to a dedicated regional tier of critical-care virtual quality nurse consultants who screen every alert around the clock, work up the chart, and only then escalate to the on-site rapid-response team — so the measured benefit is priced against the whole two-tier staffing topology, not the model alone.

Sources: escobar2020a, thekaiserpermanentenorthernc2022

Appears on: /domains/cases/kaiser-aam-deterioration, /pan-lab

EmpiricalSepsis Watch is a deep-learning sepsis-detection system scoring every emergency-department patient every five minutes ov…

Sepsis Watch is a deep-learning sepsis-detection system scoring every emergency-department patient every five minutes over 86 variables, deployed at an academic hospital under a registered clinical trial, with alerts fronted by rapid-response-team nurses who track treatment-bundle completion on three- and six-hour timers. Its structural fault line is an authority split: the operator who receives the alert (the nurse) is not the operator empowered to act on it (the physician who holds treatment authority), so the correction runs through a peer-persuasion edge. An independent ethnography found the system worked because nurses performed hidden repair work — mediating the professional hierarchy and doing the emotional labor of communicating a risk score upward — labor that was structurally necessary, largely invisible to the deployment's formal description, and undervalued.

Sources: sendak2020a, elish2020

Appears on: /domains/cases/duke-sepsis-watch, /pan-lab

EmpiricalA widely implemented proprietary sepsis-prediction model shipped inside a common electronic-health-record platform and s…

A widely implemented proprietary sepsis-prediction model shipped inside a common electronic-health-record platform and switched on across hundreds of hospitals was externally validated in 2021 across 38,455 hospitalizations at an academic health system: it achieved an area under the curve of 0.63, identified only 33 percent of sepsis cases, and had a positive predictive value of about 12 percent, generating roughly 109 alerts for every true sepsis case — a real-world performance the vendor had not fully examined before selling the model, and which an investigation attributed in part to undisclosed features such as antibiotic-order data that inflated internal validation.

Sources: wong2021c, statnews2021

Appears on: /domains/cases/epic-sepsis-michigan, /pan-lab

EmpiricalAfter external criticism, the vendor overhauled the sepsis model — retraining it, changing the sepsis-onset definition, …

After external criticism, the vendor overhauled the sepsis model — retraining it, changing the sepsis-onset definition, and reducing its reliance on antibiotic-order features. A 2026 multicenter prospective validation of the updated model across 227,091 encounters reported an area under the curve of 0.82 to 0.92 with positive predictive value of 0.13 to 0.26 and substantial between-site variability, and its authors urged local validation and alert-silencing strategies rather than trusting the model out of the box — a correction that arrived only after independent scrutiny of a model that had already been deployed at scale behind a corporate firewall shielding it from outside inspection.

Sources: statnews2022, wong2026a

Appears on: /domains/cases/epic-sepsis-michigan, /pan-lab

EmpiricalThe largest documented ambient-scribe deployment ran a 10-week pilot at an integrated medical group and then scaled to 7…

The largest documented ambient-scribe deployment ran a 10-week pilot at an integrated medical group and then scaled to 7,260 physicians and 2,576,627 patient encounters over fourteen months, with roughly 16,000 hours of documentation time saved and sustained physician support measured along the way. The system records the visit and drafts the clinical note; the clinician edits and signs, and the model-to-record write is gated both by that clinician review and by a standing internal quality-assurance program over the AI output — a real subsystem with a real cost, because the drafted note becomes a permanent record that later clinicians and later tools read as fact.

Sources: tierney2024a, tierney2025a

Appears on: /domains/cases/kaiser-tpmg-scribe, /pan-lab

EmpiricalThe scale numbers from a single ambient-scribe deployment are the deployer's own first-party measurements and should be …

The scale numbers from a single ambient-scribe deployment are the deployer's own first-party measurements and should be read as that system's dashboard rather than a guarantee of the product class: a multisite study of 8,581 clinicians across five health systems found more modest effects — on the order of 13 to 16 fewer minutes per day with no meaningful after-hours relief — and a validated per-note evaluation found hallucinations in about 31 percent of ambient-generated notes under structured review, versus about 20 percent of physician-written gold-standard notes, making ambient notes more thorough but less accurate. The clinician review and quality-assurance program are the controls that stand between that error rate and a contaminated permanent record.

Sources: rotenstein2026a, palm2025a

Appears on: /domains/cases/kaiser-tpmg-scribe, /pan-lab

EmpiricalThe strongest causal evidence in the ambient-scribe family comes from a 24-week stepped-wedge, individually randomized t…

The strongest causal evidence in the ambient-scribe family comes from a 24-week stepped-wedge, individually randomized trial of an ambient scribe across 66 practitioners and 71,487 notes (38 percent AI-generated), which found work exhaustion significantly reduced, professional fulfillment unchanged (a recorded null), roughly 22 minutes per day of documentation time returned, and diagnostic coding accuracy improved. Unlike the larger first-party deployment reports, this is a randomized estimate of the well-being and time effects — though it measured practitioner well-being and time, not per-note error rates.

Sources: afshar2025a

Appears on: /domains/cases/uw-health-abridge-scribe, /pan-lab

EmpiricalThe same team that ran the ambient-scribe trial released an open operations playbook for safety and effectiveness monito…

The same team that ran the ambient-scribe trial released an open operations playbook for safety and effectiveness monitoring of ambient AI in production — the rare case where the organization-side monitoring function exists as a citable, designed subsystem rather than an assumed practice. That monitoring is the org's stated answer to a documented system-level risk of the technology: a coding arms race, in which better AI documentation raises coding intensity, payers recalibrate in response, and clinician attestation liability grows — so the improved coding accuracy the trial measured sits next door to an upcoding pressure the monitoring is meant to watch.

Sources: afshar2025b, dai2025a

Appears on: /domains/cases/uw-health-abridge-scribe, /pan-lab

EmpiricalA peer-reviewed evaluation of an ambient documentation platform at a large multi-specialty system found note time per ap…

A peer-reviewed evaluation of an ambient documentation platform at a large multi-specialty system found note time per appointment reduced (6.2 to 5.3 minutes) and NASA-TLX cognitive load reduced — but the burnout change (42.1 to 35.1 percent) was not statistically significant, the domain's honest null bound of cognitive-load relief without a demonstrated burnout effect.

Sources: stults2025a

Appears on: /domains/cases/sutter-ambient-scribe, /pan-lab

EmpiricalIn the same evaluation, benefit varied sharply by clinician group: 85.8 percent of primary-care physicians reported impr…

In the same evaluation, benefit varied sharply by clinician group: 85.8 percent of primary-care physicians reported improved satisfaction against 36.4 percent of medical specialists — the same tool, in the same system, under the same workflow, helping one operator class and largely failing another, so any uniform service term overstates the effect for the group it helps least.

Sources: stults2025a

Appears on: /domains/cases/sutter-ambient-scribe, /pan-lab

EmpiricalCompany-run, pre-registered, peer-reviewed randomized rollouts of a commercial code-completion assistant across 4,867 de…

Company-run, pre-registered, peer-reviewed randomized rollouts of a commercial code-completion assistant across 4,867 developers at three enterprises found a pooled 26.08 percent increase in completed tasks, with gains concentrated among less-experienced developers. An independent randomized study of 16 experienced open-source maintainers on 246 tasks in familiar repositories bounded the expert tail from the other direction: those developers were about 19 percent slower with the AI while believing themselves about 20 percent faster — a measured perception-reality gap that means a uniform productivity number overstates the effect for senior engineers.

Sources: cui2025a, becker2025a

Appears on: /domains/cases/msft-accenture-copilot, /pan-lab

EmpiricalIndividual coding-assistant gains do not automatically compose to organization-level delivery outcomes: a cross-industry…

Individual coding-assistant gains do not automatically compose to organization-level delivery outcomes: a cross-industry research program measured a roughly 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability for every 25 percent increase in AI adoption, evidence that the churn the assistant adds must be absorbed by code-review and testing gates or the individual speed-up degrades the organization's delivery performance.

Sources: googleclouddora2024

Appears on: /domains/cases/msft-accenture-copilot, /pan-lab

EmpiricalAn in-house machine-learning code-completion system built, deployed, and measured by a company's own platform organizati…

An in-house machine-learning code-completion system built, deployed, and measured by a company's own platform organization for more than 10,000 internal developers reported, against a control group, a 25 to 34 percent suggestion-acceptance rate, a 6 percent reduction in coding iteration time versus control, and 3 percent of new code characters coming from the model at the time of measurement. The measuring party, the building party, and the deploying party were the same organization, and the numbers were published as an engineering-blog self-report rather than a peer-reviewed or independent evaluation.

Sources: tabachnyk2022

Appears on: /domains/cases/google-internal-completion, /pan-lab

EmpiricalThe same company's cross-industry research program reported that AI-assisted software development amplifies an organizat…

The same company's cross-industry research program reported that AI-assisted software development amplifies an organization's existing strengths and weaknesses rather than substituting for them, with policy clarity and platform investment identified as the levers that determine whether AI adoption improves or degrades delivery — evidence that the individual coding gains do not compose to organization-level outcomes on their own, and that the deploying organization's existing gates and platform quality are what decide the result.

Sources: googleclouddora2025

Appears on: /domains/cases/google-internal-completion, /pan-lab

EmpiricalA regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before scaling a co…

A regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before scaling a commercial coding assistant to roughly 1,000 engineers, publishing its own measurement of the rollout. The bank's engineers reported productivity and code-quality improvements — and recorded the security impact as explicitly inconclusive, a real gating decision taken and documented under uncertainty rather than resolved by assertion, with the honestly recorded unknown carried forward into the scaled deployment.

Sources: chatterjee2024a, theregister2024

Appears on: /domains/cases/anz-bank-copilot, /pan-lab

EmpiricalWhat the bank's inconclusive security finding leaves open is not hypothetical: an independent security assessment of cod…

What the bank's inconclusive security finding leaves open is not hypothetical: an independent security assessment of code generated by a widely used assistant found that about 40 percent of generated programs contained vulnerabilities across scenarios spanning the CWE top-25 weaknesses, and separate research documents developers accepting insecure suggestions with overconfidence — so the security unknown a deployment carries forward unresolved sits against a class-level literature in which insecure generation is common.

Sources: pearce2022a

Appears on: /domains/cases/anz-bank-copilot, /pan-lab

EmpiricalA mid-size enterprise ran a systematic four-phase evaluation-to-rollout of a commercial coding assistant across more tha…

A mid-size enterprise ran a systematic four-phase evaluation-to-rollout of a commercial coding assistant across more than 400 developers, publishing acceptance telemetry (a 33 percent suggestion-acceptance rate, with 20 percent of suggested lines accepted), a 72 percent satisfaction figure, documented per-language variation, and stated limitations. Its evaluation instrument is acceptance-rate telemetry — which the productivity literature identifies as the measure most correlated with perceived productivity rather than outcome, and perception is measured to be miscalibrated for experienced developers, so acceptance telemetry captures adoption feel, not delivered output.

Sources: bakal2025a, ziegler2024a

Appears on: /domains/cases/zoominfo-copilot, /pan-lab

EmpiricalThe deployment report stated its limitations but reported no security evaluation at all — an unrecorded unknown, one ste…

The deployment report stated its limitations but reported no security evaluation at all — an unrecorded unknown, one step less honest than a deployment that runs a security check and records the result as inconclusive, because an absence no one has written down is not a governed object and cannot be carried forward or resolved. The value of the case is the documentation quality of an ordinary, competent adoption — phase gates, telemetry definitions, per-language deltas, and stated limitations by the deployer itself — with the missing security question priced as the one thing even that documentation did not name.

Sources: bakal2025a

Appears on: /domains/cases/zoominfo-copilot, /pan-lab

EmpiricalA global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-laundering p…

A global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-laundering product as its primary monitoring system in key markets, reporting two to four times more confirmed suspicious activity with roughly 60 percent fewer alerts. Every one of those numbers is a vendor-and-customer self-report with no independent audit — which is itself the honest structure of the domain, because a peer-reviewed deployment-scale benefit measurement inside a named financial-crime operation does not publicly exist, and the alert-volume reduction the vendor advertises is precisely the lever a regulator scrutinizing an under-monitoring risk would question.

Sources: googlecloud2023

Appears on: /domains/cases/hsbc-aml, /pan-lab

EmpiricalTwo structural dynamics govern fraud and financial-crime detection. Under extreme base rates, detection precision is dom…

Two structural dynamics govern fraud and financial-crime detection. Under extreme base rates, detection precision is dominated by the false-alarm rate rather than by accuracy, so at realistic prevalence a threshold change moves the burden of alerts rather than the truth of them (the base-rate fallacy). And the labels the model learns from are the investigators' own dispositions: only a small set of flagged transactions is ever verified, and models are retrained on the analysts' calls, so a rise in 'confirmed' activity is partly a measure of what the system taught its reviewers to confirm rather than an independent ground truth (the label-feedback loop).

Sources: axelsson2000a, dalpozzolo2018a

Appears on: /domains/cases/hsbc-aml, /pan-lab

EmpiricalA neobank's fraud algorithms — triggered heavily by pandemic-era government benefit deposits — froze and closed the acco…

A neobank's fraud algorithms — triggered heavily by pandemic-era government benefit deposits — froze and closed the accounts of legitimate customers at scale, holding their balances for thirty to more than ninety days, and the company admitted some of the closures were mistakes. The false-positive tail here lands on real people as immediate hardship, concentrated among benefit-deposit recipients and low-balance households for whom a frozen account means no access to funds for weeks.

Sources: kessler2021

Appears on: /domains/cases/chime-fraud, /pan-lab

EmpiricalA 2024 federal consent order priced the downstream operational failure rather than the model: thousands of consumers wai…

A 2024 federal consent order priced the downstream operational failure rather than the model: thousands of consumers waited weeks to months for their balances after account closure, and the order imposed a 3.25 million dollar civil penalty plus at least 1.3 million dollars in consumer redress for the delayed refunds. The harm ran through three stages inside the organization's control — the scoring model's false positives, the operations backlog that turned a freeze into months without funds, and the refund process whose delay drew the regulator — and the enforcement attached to the last stage, the backlog, not to the model that started it.

Sources: consumerfinancialprotectionb2024

Appears on: /domains/cases/chime-fraud, /pan-lab

EmpiricalA Nordic bank's rules-based legacy fraud system ran at roughly 40 percent detection with a 99.5 percent false-positive r…

A Nordic bank's rules-based legacy fraud system ran at roughly 40 percent detection with a 99.5 percent false-positive rate — a measured pre-machine-learning baseline whose badness is the most credible datum in the record, since a 99.5 percent false-positive rate is not a marketing claim. The vendor-published rollout of a deep-learning engine scoring transactions in real time (under 300 milliseconds) claims false positives cut by about 60 percent and true-positive detection raised by about 50 percent; those figures are an organization-named, trade-press-covered vendor case study, entered here as claimed magnitudes against that legacy baseline because they were not independently audited.

Sources: teradata2017, groenfeldt2017

Appears on: /domains/cases/danske-fraud, /pan-lab

EmpiricalThe same institution that improved its in-line fraud scoring later ranked worst among UK banks for reimbursing victims o…

The same institution that improved its in-line fraud scoring later ranked worst among UK banks for reimbursing victims of authorized-push-payment scams in the regulator's bank-by-bank performance data — better detection and worse victim-outcome performance coexisting in one organization. And it was a rule, not a model, that moved the institutional behavior: the regulator's mandatory-reimbursement regime raised sector reimbursement from roughly two-thirds to about 89 percent, demonstrating that detection quality and the justice of the disposition are different levers held by different actors, and that the victim-outcome lever is a regulatory rule rather than a better classifier.

Sources: ukpaymentsystemsregulator2023

Appears on: /domains/cases/danske-fraud, /pan-lab

EmpiricalAn internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five stars per role…

An internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five stars per role and location — trained on ten years of the company's own hiring decisions, a period whose hires were predominantly male. The models learned that history: they penalized the word 'women's' and downgraded graduates of women's colleges, reading gender proxies as negative signal. The team patched the identified terms but concluded that term-level fixes could not guarantee neutrality against unknown proxies, because the model had learned the pattern rather than the words, and the company scrapped the project around 2017; per the company, recruiters saw the tool's recommendations but it was never used as a sole ranking.

Sources: dastin2018

Appears on: /domains/cases/amazon-resume-engine, /pan-lab

EmpiricalTraining a screener on an organization's past hiring decisions imports the past's selection function: research on hiring…

Training a screener on an organization's past hiring decisions imports the past's selection function: research on hiring as exploration finds that models trained on prior hires raise hire rates but replicate historical selection, and that a screener which values exploration rather than only exploitation breaks that lock-in loop. Two governance lessons follow — the patch lever has a documented ceiling, since removing named proxies does not remove a learned correlation, and abandonment can itself be a governance outcome, taken here before any external harm was documented rather than after an adjudication.

Sources: li2020a

Appears on: /domains/cases/amazon-resume-engine, /pan-lab

EmpiricalAn applicant-tracking platform whose AI screening and recommendation features operate inside thousands of employers' hir…

An applicant-tracking platform whose AI screening and recommendation features operate inside thousands of employers' hiring pipelines at once is the subject of a live federal collective action testing whether the vendor is directly liable as the employers' agent. On the litigation record, the court sustained the agent theory at the dismissal stage in 2024 and preliminarily certified a nationwide age-discrimination collective in 2025, covering applicants forty and over since September 2020, on a record in which the lead plaintiff reported more than one hundred rejections across employers using the platform. The litigation is ongoing and nothing here is an adjudicated finding of discrimination; these are allegations and procedural rulings, not a verdict.

Sources: mobleyvworkday2024

Appears on: /domains/cases/workday-screening, /pan-lab

EmpiricalA 2026 discovery ruling in the same matter held the vendor's internal bias-testing data privileged because counsel had c…

A 2026 discovery ruling in the same matter held the vendor's internal bias-testing data privileged because counsel had curated it — meaning the testing record exists and is legally unreachable, a configuration in which audit opacity is not the absence of testing but testing shielded from external verification. The case surfaces two further structural facts: a single vendor's screening model multiplied across many employer boundaries, so one learned defect can propagate as widely as the platform, and accountability diffusion between deployer and vendor, each holding part of the governance the other points to, against a survey backdrop showing assessment vendors' validation and bias-mitigation claims are often unverifiable from outside.

Sources: raghavan2020a, u2023b

Appears on: /domains/cases/workday-screening, /pan-lab

EmpiricalA graduate-hiring pipeline chained a games-based assessment with automated video-interview scoring, and the deployer rep…

A graduate-hiring pipeline chained a games-based assessment with automated video-interview scoring, and the deployer reports roughly a 90 percent reduction in time-to-hire (from about four months to about four weeks), around 50,000 candidate interview hours saved, about one million pounds in annual savings, and a 16 percent improvement in diversity. Every one of those figures is company- or vendor-reported and none is independently audited, so they are the deployer's own dashboard rather than an external measurement — which is exactly what the family's service regime looks like from inside.

Sources: bestpracticeai

Appears on: /domains/cases/unilever-hiring, /pan-lab

EmpiricalBoth vendors' audit machinery is on the public record in an honest but partial form. The games vendor underwent a cooper…

Both vendors' audit machinery is on the public record in an honest but partial form. The games vendor underwent a cooperative academic audit with source-code access, in which its four-fifths-rule de-biasing pipeline was found faithfully implemented — with the independence caveat that vendor staff were co-authors — and the video vendor retired its facial-analysis input under scrutiny after internal research found visual features added only about 0.25 percent predictive power, publicizing a narrow-scope external audit. The family's structural blind spot applies in full: rejected candidates never re-enter the outcome data, so the claimed quality and diversity effects are measured on hires only.

Sources: wilson2021a, maurer2021a

Appears on: /domains/cases/unilever-hiring, /pan-lab

EmpiricalA machine-learning underwriting and pricing platform using education and other alternative data operated for five years …

A machine-learning underwriting and pricing platform using education and other alternative data operated for five years under a regulator's no-action letter with a reporting obligation, and the regulator published the access results: 27 percent more applicants approved than a traditional model at 16 percent lower average APRs, with near-prime applicants (FICO 620 to 660) approved at roughly twice the rate, and gains across the tested demographic segments. This is the lending family's only regulator-verified service term. Underwriting is fully automated with no per-application human review, so the organizational levers are all upstream — model choice, the testing regime, the search for alternatives, and the reporting channel to the regulator.

Sources: consumerfinancialprotectionb2019

Appears on: /domains/cases/upstart-underwriting, /pan-lab

EmpiricalThe same deployment carries the family's most detailed public fair-lending testing record: four reports from an independ…

The same deployment carries the family's most detailed public fair-lending testing record: four reports from an independent monitorship agreed with civil-rights organizations found no close protected-class proxies quantitatively, but identified approval disparities for Black applicants, flagged a likely viable less-discriminatory alternative model, and ended in a documented methodological impasse over how hard the law requires an organization to search for such an alternative. Independently of the disparity question, adverse-action notices must give specific, accurate principal reasons for a denial regardless of the model's complexity — a governed explanation duty a complex model does not discharge by being accurate.

Sources: relmancolfaxpllc2021, consumerfinancialprotectionb2022b

Appears on: /domains/cases/upstart-underwriting, /pan-lab

EmpiricalA bank's automated credit-decisioning for a widely used consumer card was investigated by a state regulator after viral …

A bank's automated credit-decisioning for a widely used consumer card was investigated by a state regulator after viral allegations of gender bias in credit-line assignment. The regulator analyzed roughly 400,000 in-state applicants and found no unlawful discrimination on a prohibited basis — the model was cleared on the numbers. But the same investigation documented failures of explanation, customer service, and perceived transparency: applicants could not learn why they received the terms they did, front-line staff could not explain the decisions, and the resulting opacity destroyed consumer trust even though the underwriting itself was found lawful. This is the domain's cleared-but-faulted case: a statistically clean model paired with a failed duty to explain.

Sources: newyorkstatedepartmentoffina2021

Appears on: /domains/cases/goldman-apple-card, /pan-lab

EmpiricalThe lesson the cleared-but-faulted outcome carries is that a lawful, statistically clean model does not discharge the se…

The lesson the cleared-but-faulted outcome carries is that a lawful, statistically clean model does not discharge the separate duty to explain a decision. Regulators have made explicit that adverse-action notices must give specific, accurate principal reasons regardless of how complex the model is, and that a model being a black box is not a defense — checking the nearest sample-form box does not comply. The explanation and customer-service channel is therefore a distinct, separately-resourced surface that can fail on its own: an organization can pass its fair-lending testing and still fail the people it decides on by being unable to tell them why.

Sources: consumerfinancialprotectionb2022b

Appears on: /domains/cases/goldman-apple-card, /pan-lab

EmpiricalA state attorney general reached a $2.5 million settlement with a student-loan lender over its AI underwriting. The docu…

A state attorney general reached a $2.5 million settlement with a student-loan lender over its AI underwriting. The documented conduct is the domain's cleanest failure-then-mandated-governance arc: the model used a cohort-default-rate feature — a school's aggregate default rate priced into an individual applicant's terms — that disparately impacted Black and Hispanic applicants, and an immigration-status rule that automatically denied certain non-citizen applicants, while the organization ran no disparate-impact testing and gave inadequate adverse-action notices. The remedy did not fine-and-close: it mandated the missing program — model governance, disparate-impact testing, documentation, and reporting controls — so the enforcement action wrote the governance the deployment had never built.

Sources: officeofthemassachusettsatto2025

Appears on: /domains/cases/earnest-ai-underwriting, /pan-lab

EmpiricalThe mechanism the case turns on is the facially-neutral aggregate feature: a cohort default rate is a property of a scho…

The mechanism the case turns on is the facially-neutral aggregate feature: a cohort default rate is a property of a school, not of the applicant, and no input names a protected class — yet pricing a group's aggregate history into an individual's terms can carry protected-class impact, which is exactly what disparate-impact testing exists to catch. Here that testing was not done, so the impact went unmeasured until an enforcement action found it. The remedy installed the program the deployment lacked, which is the governable reading: an aggregate feature can look neutral input-by-input and still produce a disparity only outcome testing would reveal, and the absence of that testing is itself the failure.

Sources: officeofthemassachusettsatto2025, consumerfinancialprotectionb2022b

Appears on: /domains/cases/earnest-ai-underwriting, /pan-lab

EmpiricalThe strongest field evidence for an agent-assist copilot in customer service comes from a staggered randomized rollout o…

The strongest field evidence for an agent-assist copilot in customer service comes from a staggered randomized rollout of a generative-AI assistant to roughly 5,000 customer-support agents at a large software firm. Measured against a control group, the copilot raised issues resolved per hour by about 15 percent on average, and it also improved customer sentiment and agent retention. The gain, however, was sharply uneven: novice and low-skill agents improved by roughly 30 to 34 percent, agents with two months of experience performed like agents with six months and no AI, and the most experienced agents gained close to nothing, with some evidence of slight quality degradation. This is the contact-centre domain's cleanest measured benefit, and it is a distribution rather than a single number.

Sources: brynjolfsson2025a, brynjolfsson2023

Appears on: /domains/cases/fortune500-agent-copilot, /pan-lab

EmpiricalThe lesson the randomized evidence carries is skill compression: an agent-assist copilot mostly raises the floor. Becaus…

The lesson the randomized evidence carries is skill compression: an agent-assist copilot mostly raises the floor. Because almost the entire measured gain accrues to less-experienced agents and the most experienced gain close to nothing, an average productivity number overstates the effect for the agents who least need it and hides that the tool does little for the experienced while possibly costing a small amount of quality there. The governable reading is that the benefit must be measured as a distribution across agent skill, not reported as a scalar — a copilot that helps novices a great deal and experts not at all is a real and specific benefit, and describing it with one average misstates who it helps.

Sources: brynjolfsson2025a

Appears on: /domains/cases/fortune500-agent-copilot, /pan-lab

EmpiricalAn airline's customer-facing website chatbot told a customer they could claim a bereavement fare retroactively — a polic…

An airline's customer-facing website chatbot told a customer they could claim a bereavement fare retroactively — a policy that did not exist. The customer relied on the chatbot's statement, bought a ticket, and was then refused the fare by the airline's human staff. A civil-resolution tribunal found the airline liable for negligent misrepresentation and awarded damages, and in doing so rejected the airline's argument that the chatbot was a separate legal entity responsible for its own actions. The tribunal held that the organization is responsible for all the information on its website, whether it comes from a static page or a chatbot, and that a customer has no way to know which source to trust. This is the contact-centre domain's cleanest accountability ruling: the bot is a tool the company answers for, not an entity that answers for itself.

Sources: moffattv2024, sookman2024

Appears on: /domains/cases/air-canada-chatbot, /pan-lab

EmpiricalThe duty the ruling establishes is that an organization must take reasonable care that its chatbot's representations are…

The duty the ruling establishes is that an organization must take reasonable care that its chatbot's representations are accurate, because the chatbot is a tool it deploys rather than a separate entity that answers for itself. A hallucinated policy or a wrong rule stated by the bot is therefore the organization's own misrepresentation, and a posture that treats the AI as speaking only for itself does not transfer that responsibility away. The governable reading is that a customer-facing chatbot is a channel the organization is accountable for exactly as it is accountable for a page on its own website — so the accuracy control on what the bot states, and the ownership of what it says, are the organization's to build, not the bot's to carry.

Sources: sookman2024, moffattv2024

Appears on: /domains/cases/air-canada-chatbot, /pan-lab

EmpiricalAn organization published striking first-month numbers for its customer-facing AI assistant: it handled about two-thirds…

An organization published striking first-month numbers for its customer-facing AI assistant: it handled about two-thirds of customer-service chats (some 2.3 million conversations), was described as doing the equivalent work of about 700 full-time agents, cut average resolution time from about 11 minutes to under 2, was said to match human customer satisfaction, and was projected to improve profit by tens of millions. Every one of those figures was self-reported and not independently audited. Roughly a year later the same organization reversed course on quality grounds — its chief executive said cost had become too predominant an evaluation factor and the result was lower quality — and committed to always keeping a human available to customers who want one. This is the contact-centre domain's cleanest benefit-then-cost arc: the deflection numbers and the walk-back come from the same deployment.

Sources: klarnabankab2024, ivanova2025

Appears on: /domains/cases/klarna-ai-assistant, /pan-lab

EmpiricalThe lesson the benefit-then-cost arc carries is that deflection is not resolution. A published deflection number reports…

The lesson the benefit-then-cost arc carries is that deflection is not resolution. A published deflection number reports how many contacts the AI handled, not whether it handled them well, and a figure that is impressive on cost can hide a quality cost that only shows up later — which is what the organization's own reversal described. The survey backdrop sharpens it: most customers say they would rather not meet AI in service and fear it makes reaching a human harder, and industry analysts expect a large share of organizations to abandon plans to reduce their customer-service workforce with AI. The governable reading is to measure resolution and repeat contact against deflection rather than counting deflection as a win by itself, and to protect the path to a human as the safety valve a deflection-maximizing design tends to erode.

Sources: ivanova2025, gartner2025b

Appears on: /domains/cases/klarna-ai-assistant, /pan-lab

EmpiricalA video platform ran an unintended natural experiment on automated content moderation. When the pandemic sent its human …

A video platform ran an unintended natural experiment on automated content moderation. When the pandemic sent its human reviewers home, the platform said it would rely more on automated removal and deliberately chose over-enforcement rather than let harmful content stay up. The result, from the platform's own transparency reporting, was that automated removals more than doubled in a single quarter (to about 11.4 million videos), appeals roughly doubled, and the reinstatement rate on appeal jumped from about 25 percent to about 50 percent. The platform also withheld strikes where no human had reviewed the removal, treating the automated decision as provisional. The doubling of the reinstatement rate is the finding: it is direct evidence that the automation was making roughly twice the rate of catchable errors, and that the human review and appeals path was the loop catching them.

Sources: youtubegoogle2020

Appears on: /domains/cases/youtube-covid-enforcement, /pan-lab

EmpiricalThe lesson the natural experiment carries is that the human review and appeals path is the error-correction loop for aut…

The lesson the natural experiment carries is that the human review and appeals path is the error-correction loop for automated enforcement, not an optional add-on. Automated moderation makes errors at scale, and a doubling of the reinstatement rate when human review thinned is a measurement of those errors — they were always being made at that rate, and were visible only because the appeals queue surfaced them. Two things follow. Over-enforcement versus under-enforcement is a chosen trade-off: with review capacity cut, the organization decided which error to make, and that was a governance decision. And proactive removal acts before anyone sees the content, so an over-broad takedown is invisible unless appealed — and some removals are irreversible, as when automated systems destroyed documentation of war crimes with archival access declined, leaving no correction loop at all.

Sources: youtubegoogle2020, humanrightswatch2020

Appears on: /domains/cases/youtube-covid-enforcement, /pan-lab

EmpiricalA platform enforces its content standards with automated classifiers at a scale no human team could match, backed by a l…

A platform enforces its content standards with automated classifiers at a scale no human team could match, backed by a layered correction structure: an internal appeals process, and above it an external oversight board that issues binding decisions on the individual cases it takes and non-binding policy recommendations to the platform. In one year the board overturned the platform's original decision in around 90 percent of the cases it decided, and the platform reported implementing, in progress on, or already aligned with the large majority of the board's cumulative recommendations. This is the moderation domain's most built-out, institutionalized correction structure — layered appeals rising to an independent-adjacent external body that publishes its reasons.

Sources: metaplatforms, oversightboard2024

Appears on: /domains/cases/meta-content-enforcement, /pan-lab

EmpiricalThe reach of the correction structure is the governable limit. The roughly 90 percent overturn rate is measured on selec…

The reach of the correction structure is the governable limit. The roughly 90 percent overturn rate is measured on selected cases — the board chooses emblematic disputes to set precedent, so the figure is evidence that escalated decisions were often wrong, not a random error rate, and the overwhelming majority of automated enforcement decisions never reach the board at all. The board is funded through a platform-established trust, which makes it independent-adjacent rather than fully independent, and its policy recommendations are non-binding. The honest reading is that this correction structure is real and genuinely better than most, and its reach is bounded to the tiny fraction of cases that escalate — so the governing question is whether the correction reaches the scale of the enforcement it is meant to check.

Sources: oversightboard2024, oversightboard2025

Appears on: /domains/cases/meta-content-enforcement, /pan-lab

EmpiricalA media outlet published AI-drafted finance explainers under a human-sounding staff byline without disclosing to readers…

A media outlet published AI-drafted finance explainers under a human-sounding staff byline without disclosing to readers that the articles were machine-written. When the practice came to light, the outlet's own audit found it had to issue corrections on a majority of the AI-written articles — on the order of 41 of 77. A byline implies a human review that the reader trusts, and a correction rate that high is a direct measurement that the review the byline implied was not actually performed before publication. A later and sharper case saw another outlet publish articles under entirely fabricated author personas presented as real people, so the failure ran from undisclosed AI drafting to invented human bylines.

Sources: bonifacic2023, harrisondupre2023

Appears on: /domains/cases/cnet-ai-drafting, /pan-lab

EmpiricalEditorial AI moves the failure from a takedown to a publication, but the governable structure is the same as in moderati…

Editorial AI moves the failure from a takedown to a publication, but the governable structure is the same as in moderation: the byline is the accountability object, and it stands for a review that either happened or did not. Two things are owed to the reader — disclosure that AI was involved, and an editorial check that actually took place — and this deployment gave neither, publishing under a staff byline that implied both. When a large share of AI-drafted articles needs correction, the review was not performed, and the byline misrepresented who did the work. The governable reading is that a human byline on machine-drafted content is a claim about review and disclosure, and a high correction rate is the evidence that the claim was false.

Sources: bonifacic2023

Appears on: /domains/cases/cnet-ai-drafting, /pan-lab

EmpiricalA large automaker deployed in-line AI inspection at production scale: camera and acoustic systems that detect defects du…

A large automaker deployed in-line AI inspection at production scale: camera and acoustic systems that detect defects during assembly and feed real-time flags to the line worker via a smart device, on a line running on the order of a thousand-plus vehicles a day at a takt of under a minute per station. The system has been established as a company standard and is being extended to suppliers. The governing design is that the AI flags and a human on the line responds — the inspection is wired into a resourced response loop, including the ability to stop the line, so the benefit runs through the human response the flag triggers rather than through the model alone. The documented facts here are the system's function, the worker-interaction model, the scale, and the standardization; the deployment's benefit is reported through corporate and trade channels, and defect-rate deltas from a primary source are not public.

Sources: bmwgrouppressclub2025, metrologyandqualitynews2026, leanenterpriseinstitute

Appears on: /domains/cases/bmw-aiqx-inspection, /pan-lab

EmpiricalThe lesson the deployment carries is that an in-line inspection AI is only as good as the human-response loop it trigger…

The lesson the deployment carries is that an in-line inspection AI is only as good as the human-response loop it triggers, and that loop is the governable object. When the AI flags a defect, a resourced response — a worker with the time to check the flag and the authority to stop the line — is what turns a detection into a caught defect; without it, the flag is just a decision no one acts on. This is why the failure modes in this domain are matters of the loop's calibration rather than the model's raw accuracy: too many false alarms and operators stop responding, too much trust and they stop checking. The honest boundary is that the benefit is reported through corporate and trade channels, and no named manufacturer has publicly attributed a shipped-defect escape to its AI inspection, so the response loop is drawn as the resourced strength and its calibration as the thing to govern, not as a claim about defects that did or did not ship.

Sources: bmwgrouppressclub2025, leanenterpriseinstitute

Appears on: /domains/cases/bmw-aiqx-inspection, /pan-lab

EmpiricalA peer-reviewed heavy-industry predictive-maintenance case study achieved a large, measured reduction in false alarms — …

A peer-reviewed heavy-industry predictive-maintenance case study achieved a large, measured reduction in false alarms — on the order of 90 percent — through a closed operator-feedback loop: the maintenance crews investigated the alerts, labeled which were real, and the model retrained on those labels, so the false-alarm rate fell sharply over successive rounds. This is the industrial-QA domain's best-measured quantitative benefit, and it comes from an anonymized study site rather than a named-manufacturer press release, which is the pattern across this domain — the peer-reviewed magnitudes are at anonymized or smaller sites, while the named deployments report their benefit through corporate and trade channels.

Sources: hermansa2021a

Appears on: /domains/cases/heavy-industry-pdm, /pan-lab

EmpiricalThe lesson the case carries is that the same closed feedback loop that produced the benefit is the thing that can break …

The lesson the case carries is that the same closed feedback loop that produced the benefit is the thing that can break it, because the loop depends on the crews continuing to engage with the alerts — investigating them, labeling them, responding — and that engagement fails in two opposite directions. Alert fatigue: if too many false alarms arrive before the loop has tuned them down, crews stop trusting the alerts and stop responding, so the feedback the model needs to improve never arrives and the loop stalls. Automation bias: if crews defer to the alerts and stop applying their own judgment, the labels the model retrains on become an echo of its own calls rather than an independent check. Either way the loop degrades, so the measured benefit is contingent on the loop staying calibrated — enough trust that crews respond, enough independence that their labels still carry real judgment.

Sources: romeo2025a, wittbold2026

Appears on: /domains/cases/heavy-industry-pdm, /pan-lab

EmpiricalMachine-learning automated visual inspection of filled injectable drug products flags particulate and cosmetic defects t…

Machine-learning automated visual inspection of filled injectable drug products flags particulate and cosmetic defects that manual inspection or fixed-rule cameras would otherwise judge. In this safety-critical, regulated manufacturing setting the error trade-off is asymmetric and deliberate: a false accept — a missed defect in an injectable that reaches a patient — is a patient-safety failure, while a false reject — scrapping a good vial — is a cost, so the system is tuned to over-reject rather than risk a miss. Because the inspection sits inside a validated pharmaceutical quality process, the AI cannot simply be switched on; it must be qualified within that process, and a regulator is actively developing the framework for how AI in drug manufacturing should be validated and monitored.

Sources: veillon2023a, usfda2023

Appears on: /domains/cases/pharma-avi-inspection, /pan-lab

EmpiricalTwo governable surfaces follow from putting AI inside a regulated inspection. First, qualification: an AI in a validated…

Two governable surfaces follow from putting AI inside a regulated inspection. First, qualification: an AI in a validated quality process is not simply deployed but must be qualified and monitored for drift, and because the regulator's AI-specific framework is still developing, the qualification of the model's behavior over time is an emerging, not-yet-settled check rather than a solved one. Second, the human backstop: the manual inspector is what catches the false accepts the over-reject tuning is meant to avoid, so if inspectors come to defer to the AI and stop scrutinizing, that backstop erodes exactly where it matters most — the missed defect the asymmetric tuning was designed to prevent. The governable reading is that the over-reject tuning lowers the visible risk without removing it, and the qualification and the human backstop are what keep the residual risk covered.

Sources: veillon2023a, usfda2023

Appears on: /domains/cases/pharma-avi-inspection, /pan-lab

EmpiricalA state's statewide dropout early-warning system used ensemble machine learning to label every grade 6 to 9 student's ri…

A state's statewide dropout early-warning system used ensemble machine learning to label every grade 6 to 9 student's risk of not graduating on time and delivered the label to school staff through dashboards for about a decade. An independent, decade-scale audit found the system was wrong roughly 74 percent of the time when it predicted a student would not graduate, produced higher false-alarm rates for Black and Hispanic students, and that the deployer's own internal equity research had gone unpublished — while a survey of districts found administrators reporting no training on how to interpret a 'high risk' label. The state stopped publishing the dashboards in 2023 and said it was evaluating the system's future. The deployment is the education domain's clearest case of a risk label whose error and group disparity entered how students were seen rather than the help they received.

Sources: feathers2023, knowles2015a, wisconsindepartmentofpublici2023

Appears on: /domains/cases/wisconsin-dews, /pan-lab

EmpiricalThe lesson the case carries is that a risk label is only as good as the intervention it triggers and the training of the…

The lesson the case carries is that a risk label is only as good as the intervention it triggers and the training of the human who reads it. A label that is wrong most of the time, delivered to staff with no guidance on interpreting it, imports the model's error and its group disparity into how students are perceived rather than into a resourced response — the flag becomes a lens on the student rather than a trigger for help. Set against this, a large district's transparent, low-tech on-track indicator, built on interpretable research and paired with real intervention, accompanied a rise in graduation to a record level. The contrast locates the benefit in the intervention the indicator makes legible enough for staff to act on well, not in the sophistication of the prediction — an interpretable indicator that drives help can outperform an opaque model that only labels.

Sources: allensworth2007, feathers2023

Appears on: /domains/cases/wisconsin-dews, /pan-lab

EmpiricalA public university required students to pan their webcam around their home before an online exam, using remote-proctori…

A public university required students to pan their webcam around their home before an online exam, using remote-proctoring software that flags suspected cheating from the video. A federal court held that the pre-exam room scan was an unreasonable search under the Fourth Amendment — a first-of-its-kind ruling that a routine proctoring practice violated a student's constitutional rights in their own home. Separately, peer-reviewed measurement of automated proctoring found the software produced more face-detection failures, more red flags, and higher priority scores for darker-skinned and Black students, with no corresponding difference in actual cheating. The deployment is the education domain's clearest case of surveillance-based integrity AI whose costs — a rights violation and a demographic burden of suspicion — are each independently established.

Sources: ogletreev2022a, yoderhimes2022a

Appears on: /domains/cases/cleveland-state-proctoring, /pan-lab

EmpiricalThe lesson the case carries is that surveillance-based integrity AI is not a free default: it carries a rights cost that…

The lesson the case carries is that surveillance-based integrity AI is not a free default: it carries a rights cost that can be independently adjudicated and a demographic burden that can be measured, and both are owed a reckoning before the surveillance is imposed, not after a court or an audit finds the harm. A room scan of a student's home was held to be an unreasonable search, so the surveillance has a rights dimension a court can rule on regardless of the integrity goal. And because the software flags darker-skinned and Black students more often with no more actual cheating, and a flag is an accusation the student must answer, a disparate flag rate is a disparate burden of suspicion. The governable surfaces are the proportionality of the surveillance to the integrity problem it is trying to solve, and the measured flag rate by group.

Sources: yoderhimes2022a, ogletreev2022a

Appears on: /domains/cases/cleveland-state-proctoring, /pan-lab

EmpiricalA parcel carrier's route-optimization system is a documented operations-research success: it re-optimizes delivery route…

A parcel carrier's route-optimization system is a documented operations-research success: it re-optimizes delivery routes across the fleet and was reported to save on the order of 100 million miles and about 10 million gallons of fuel a year, a genuine and peer-reviewed efficiency gain. The same system that computes the efficient route also dictates it to the driver and monitors adherence through vehicle telematics, so the efficiency is enforced through workplace surveillance — the optimization and the monitoring are one system, and the driver's discretion over how to run the route is what it replaces. The benefit is real and measured in miles and fuel; the cost is the driver autonomy the enforcement removes and the surveillance the enforcement requires.

Sources: holland2017, levy2023

Appears on: /domains/cases/ups-orion-routing, /pan-lab

EmpiricalThe lesson the case carries is that an optimization which manages the worker executing it couples the efficiency gain to…

The lesson the case carries is that an optimization which manages the worker executing it couples the efficiency gain to a cost the efficiency metric does not see: the worker's autonomy, and the surveillance required to enforce the plan. The system measures miles and fuel, not whether the pace it sets is feasible for a person or whether the monitoring it requires is proportionate — so the governable surfaces are whether the optimization internalizes the human executing it, meaning a route that is feasible and humane rather than merely optimal on paper, and whether the surveillance that enforces it is governed rather than treated as a free byproduct of routing. An optimization is a success on its own terms and can still externalize a cost onto the worker that never appears in the miles-and-fuel number it reports.

Sources: levy2023, holland2017

Appears on: /domains/cases/ups-orion-routing, /pan-lab

EmpiricalA warehouse operation's algorithmic management pairs a genuine, peer-reviewed human-robot picking benefit — robots and w…

A warehouse operation's algorithmic management pairs a genuine, peer-reviewed human-robot picking benefit — robots and workers collaborating to raise throughput, documented in the operations-research literature — with a documented injury-productivity trade-off. When the algorithm sets the pace of the physical work, a federal safety regulator cited the operation for exposing workers to ergonomic hazards, and a legislative inquiry tied the speed the system demands to warehouses it described as uniquely dangerous. The productivity gain and the worker-injury risk are therefore coupled: the same pace that raises units per hour is the pace regulators and the inquiry connected to injury. The benefit is real and the injury cost is separately documented, one in the OR literature and one in safety-inspection findings and a legislative report.

Sources: allgor2023, ussenatecommitteeonhealth2024, usdepartmentoflabor2023a

Appears on: /domains/cases/amazon-fulfillment-management, /pan-lab

EmpiricalThe lesson the case carries is that when an algorithm sets the pace of physical work, the productivity metric it optimiz…

The lesson the case carries is that when an algorithm sets the pace of physical work, the productivity metric it optimizes — units per hour — cannot see the cost the pace imposes on the body executing it. The injury shows up in safety-inspection data and a legislative inquiry, not on the throughput dashboard, so a productivity number can rise while the cost accumulates unrecorded on the metric that reports success. The governable question is whether the pace-setting internalizes the worker's safety, treating a sustainable rate as part of what 'optimal' means, or externalizes it as an injury the metric never records. Ethnographic research describes this algorithmic management as a 'game' whose rules the worker cannot change, which is what makes the pace a management decision the organization owns rather than a fact of the work.

Sources: cheon2025, ussenatecommitteeonhealth2024

Appears on: /domains/cases/amazon-fulfillment-management, /pan-lab

EmpiricalA federal asylum agency uses dialect-recognition AI to estimate an applicant's country or region of origin from a short …

A federal asylum agency uses dialect-recognition AI to estimate an applicant's country or region of origin from a short speech sample, as one input into the credibility assessment of their claimed origin. The tool's reliability is limited: government-reported recognition is around 80 percent for Arabic — roughly a 20 percent error rate — and computational linguists judge separating some closely related language varieties close to hopeless. The agency's own caseworkers describe the tool as only a rough compass, too imprecise to resolve the hard cases, and its outputs as clues rather than determinations. Used honestly as one clue among several it is defensible; the documented risk is that an imprecise output acquires more authority than its accuracy supports, in a determination where the state's tool is set against the applicant's own account of who they are.

Sources: lulamae2022a, scheel2024a

Appears on: /domains/cases/bamf-dias-dialect, /pan-lab

EmpiricalTwo governable surfaces follow from putting a low-reliability signal into a high-stakes credibility determination. First…

Two governable surfaces follow from putting a low-reliability signal into a high-stakes credibility determination. First, whether the tool's documented imprecision actually bounds the weight it carries: a rough compass treated as one is honest, but the same output can harden into a credibility finding it cannot support once a phrase like the software indicates a particular origin enters the record and confronts the applicant. Second, whether the applicant can see and contest the signal: in asylum determinations the person with the most at stake and the most knowledge of their own origin is often unable to see or challenge the AI's estimate, so the correction that would catch an error is severed on exactly the side that holds the truth. The governable reading is that reliability must bound authority, and the affected person must be able to contest a signal used against them.

Sources: scheel2024a, lulamae2022a

Appears on: /domains/cases/bamf-dias-dialect, /pan-lab

EmpiricalA government's immigration-enforcement triage algorithm identifies and recommends people for enforcement actions — retur…

A government's immigration-enforcement triage algorithm identifies and recommends people for enforcement actions — returns, bail conditions, casework — drawing on sensitive data including detention, health, vulnerability, and location-monitoring records. Uncovered through roughly a year of freedom-of-information litigation, its training materials show an asymmetric override design: officials must record a justification for rejecting a recommendation but not for accepting one. That design builds a rubber-stamping incentive into the workflow — accepting the algorithm is frictionless, overriding it requires work — so the human in the loop is nominal rather than a real check. It is the corpus's clearest documented instance of automation bias engineered into an agency workflow, in one of the highest-stakes enforcement settings a state operates.

Sources: privacyinternational2024b, privacyinternational2024a

Appears on: /domains/cases/home-office-ipic, /pan-lab

EmpiricalThe lesson the case carries is that nominal human oversight is not real oversight. An asymmetric override — where accept…

The lesson the case carries is that nominal human oversight is not real oversight. An asymmetric override — where accepting the algorithm's recommendation is frictionless and rejecting it requires a recorded justification — engineers automation bias into the process by making deference the path of least resistance, so the claim that a human makes the final decision can be true and empty at once. Two governable surfaces follow. Whether the review is genuinely symmetric: the official as free and as prompted to reject as to accept, so an error is as likely to be caught as waved through. And whether the affected person is told the AI is used and can contest it: applicants are frequently not told, which severs the correction on the side that could challenge the recommendation, so the one check that survives the asymmetric override — the person it is about — is cut out too.

Sources: privacyinternational2024c, privacyinternational2024a

Appears on: /domains/cases/home-office-ipic, /pan-lab

EmpiricalA large automaker developed an in-house deep-learning system to detect hairline cracks in pressed sheet-metal parts, tra…

A large automaker developed an in-house deep-learning system to detect hairline cracks in pressed sheet-metal parts, trained on several terabytes of images drawn from seven presses at its home plant plus several sister plants, in development since mid-2016 and tested for series deployment. The documented change is a generational replacement: the system takes over an inspection duty previously performed by manual visual checks plus fixed-rule camera systems, rather than augmenting a human inspector's judgment on each part. The record — a reprint of the manufacturer's own press material with its CIO quoted — documents the development lineage, the data scale, and what the system replaced; it publishes no quantitative defect-rate figures, so the deployment's benefit magnitude is a corporate claim, not an audited measurement.

Sources: justauto2018

Appears on: /domains/cases/audi-press-shop-inspection, /pan-lab

EmpiricalThe governance shape of this deployment is inheritance rather than assistance: by replacing the manual visual check and …

The governance shape of this deployment is inheritance rather than assistance: by replacing the manual visual check and the fixed-rule camera generation, the learned system inherits the whole inspection duty for the defect class it covers, so there is no per-part human judgment running alongside it to catch what it misses. Its training data is pooled across presses and plants, which means one model's blind spots are correlated across every line it inspects. The failure regime is mechanism-level — drift as dies wear and parts change, complacency over an inspection nobody re-performs — because no named manufacturer, including this one, has publicly attributed a shipped-defect escape to its AI inspection.

Sources: justauto2018

Appears on: /domains/cases/audi-press-shop-inspection, /pan-lab

EmpiricalA rail service provider operates sensor-based predictive maintenance on a high-speed fleet under a priced availability c…

A rail service provider operates sensor-based predictive maintenance on a high-speed fleet under a priced availability contract: roughly 300 sensors per train read at five-minute intervals (on the order of a million readings per train-year), overlaid with human-written failure reports, maintained against a promise that refunds the full fare if a journey is delayed more than fifteen minutes. The documented results are only one noticeably delayed journey in 2,300 (by five minutes) and discovered failure signatures such as an engine-temperature pattern preceding failure by three days. The fleet outcome is documented in a vendor-side trade case study; the analytics' own precision is not published, so the model-level figures remain unstated while the operational outcome is on the record.

Sources: rcrwirelessnews2016

Appears on: /domains/cases/siemens-renfe-velaro-pdm, /pan-lab

EmpiricalThe governance shape of this deployment is uptime-as-contract: the party that operates and tunes the analytics is the se…

The governance shape of this deployment is uptime-as-contract: the party that operates and tunes the analytics is the service provider who pays for misses under the refund promise, so the incentive to prevent a delay is priced into the same organization that holds the model levers - an alignment the domain's other deployments lack. Two documented dependencies temper it: the failure signatures were discovered by joining sensor streams to human-written failure reports, so the discovery loop runs on documentation crews write for their own purposes; and a continuous sensor stream is the analytics' only view of the machine, so a failing sensor and a failing train arrive looking the same until someone goes and looks.

Sources: rcrwirelessnews2016

Appears on: /domains/cases/siemens-renfe-velaro-pdm, /pan-lab

EmpiricalA large urban school district operationalized a transparent ninth-grade indicator - course credits earned plus no more t…

A large urban school district operationalized a transparent ninth-grade indicator - course credits earned plus no more than one core-course failure - from consortium research showing it predicts high-school graduation with about 85 percent accuracy, and wired it to school-level attention rather than to an opaque score. District graduation rates subsequently rose to record highs. The indicator is a rule anyone can read: a teacher can explain to a student exactly why they are off-track and exactly what would change it, so the contest-and-correction loop that opaque early-warning deployments sever is open by construction.

Sources: allensworth2007

Appears on: /domains/cases/cps-freshman-ontrack, /pan-lab

EmpiricalThe documented limits are as instructive as the result. The indicator's accuracy and the district's graduation rise are …

The documented limits are as instructive as the result. The indicator's accuracy and the district's graduation rise are associational at district scale - no randomized trial assigns schools to use it - and the benefit mechanism runs through the intervention, not the flag: an indicator wired to attention still depends on the attention being resourced, and the research base's central finding is that what predicted graduation was a condition schools could act on (freshman-year course performance), not a fixed trait of the student. The rule's power is that it points at something changeable, and the district's practice is what changed it.

Sources: allensworth2007

Appears on: /domains/cases/cps-freshman-ontrack, /pan-lab

EmpiricalA storied sports outlet published product-review articles under entirely fabricated author personas - invented names, AI…

A storied sports outlet published product-review articles under entirely fabricated author personas - invented names, AI-generated headshots, fictional biographies - produced by a third-party content contractor, with the AI involvement disclosed to no reader. An investigation surfaced the personas by reading the public site; the articles were then deleted rather than corrected, the outlet attributed the content to the contractor, and the parent company's chief executive was subsequently fired. The corroborated record documents the fabrication, the deletion, the contractor attribution, and the executive consequence.

Sources: harrisondupre2023, npr2023

Appears on: /domains/cases/sports-illustrated-advon, /pan-lab

EmpiricalThe governance failure ran across an organizational seam: the drafting, the bylines, and the personas were produced by a…

The governance failure ran across an organizational seam: the drafting, the bylines, and the personas were produced by a contractor, and the outlet's editorial function demonstrably did not operate across that boundary - the fabrication was discovered by outside investigation, not by any internal check, and the accountability afterward ran through contract and employment rather than through any editorial process. A byline makes two claims to the reader - that a person produced this, and that the outlet's review stands behind it; this deployment fabricated the first and vacated the second, and the deletion afterward removed the evidence rather than correcting the record.

Sources: harrisondupre2023, npr2023

Appears on: /domains/cases/sports-illustrated-advon, /pan-lab

EmpiricalA parcel firm's customer-facing support chatbot, after a system update, was prompted by a customer into swearing and int…

A parcel firm's customer-facing support chatbot, after a system update, was prompted by a customer into swearing and into composing a poem calling its own operator the worst delivery firm in the world. The firm attributed the behavior to the update and disabled the AI element immediately. The documented governance facts are exactly two: the update preceded the behavior, and the off switch worked - the firm learned of the incident from the customer's viral post rather than from any release gate, but the disablement was immediate once it knew.

Sources: itvnews2024

Appears on: /domains/cases/dpd-uk-chatbot, /pan-lab

EmpiricalThe failure shape is a change-management regression, not a wrong policy or a deflection metric: guardrails that had held…

The failure shape is a change-management regression, not a wrong policy or a deflection metric: guardrails that had held in production stopped holding after a change, publicly, within hours, in a channel that talks to anyone. What the deployment lacked was a release gate between the update and the public - the constraint layer's behavior after the change was tested by a customer with a prompt, not by the firm with a suite - and the discovery path ran through screenshots of one conversation going viral.

Sources: itvnews2024

Appears on: /domains/cases/dpd-uk-chatbot, /pan-lab

EmpiricalSince May 2018, New York City's Administration for Children's Services has scored every open child-protection investigat…

Since May 2018, New York City's Administration for Children's Services has scored every open child-protection investigation at day 10 with an in-house machine-learning model (documented as the ASAP Tool / Severe Harm Predictive Model), rank-ordering cases by predicted likelihood of substantiated physical or sexual abuse within 24 months to fill a quality-assurance review worklist of about 3,000 of roughly 50,000 investigations a year; the LL35 register states scores are not shared with staff in the QA unit or the investigative unit, families and their attorneys are not told when a case is flagged, ACS told state auditors there would be 'no basis for a complaint' about a predictive model on an individual case, and the agency's own internal audit acknowledged the training data likely included implicit and systemic biases, that geographic variables may act as partial proxies for race, and that flag predictions are more likely to be incorrect than correct.

Sources: nycofficeoftechnologyandinno2026, lecher2025, newyorkstatecomptroller2023

Appears on: /domains/cases/nyc-acs-qa-risk-algorithm

EmpiricalThe May 2026 'Access Denied' report by the NYC Department of Investigation — the Charter-mandated inspector general for …

The May 2026 'Access Denied' report by the NYC Department of Investigation — the Charter-mandated inspector general for ACS — documents that five provisions of NY Social Services Law, as applied by state OCFS, routinely deny, limit, or delay DOI's access to ACS child-welfare records, with unfounded-report and CARES records prohibited entirely; DOI was barred from the full case history in 17 of the 18 child fatalities with prior ACS involvement reported to it in 2025 (13 of 16 in 2024; 19 of 25 in 2023). The report does not mention the algorithm — the linkage is a topological inference — but in 2025 ACS expanded the model to score CARES alternative-response cases, the record class state law prohibits DOI from accessing; the NYC Council's GUARD Act (passed unanimously November 25, 2025) legislated an Office of Algorithmic Data Accountability whose implementation status remains unverified as of mid-2026.

Sources: newyorkcitydepartmentofinves2026, nycofficeoftechnologyandinno2026, statescoop2025

Appears on: /domains/cases/nyc-acs-qa-risk-algorithm

EmpiricalThe ACS severe-harm model trains on and scores exclusively from the agency's own administrative records (including hotli…

The ACS severe-harm model trains on and scores exclusively from the agency's own administrative records (including hotline call counts and durations, prior involvement, and geography), and the intervention its flags trigger — extra interviews, collateral contacts, service referrals, consults, and QA documentation follow-ups — writes new activity into those same records, enriching the prior-involvement features of any future report on the family; ACS has never studied what effect the flag-triggered extra review has on downstream case outcomes, including whether a child is ultimately removed from the home (while telling state auditors it produces quarterly internal reports tracking the model's use in the QA program), kept no logs of model performance evaluations or updates as of the 2019-2022 state-audit fieldwork, and its claim that the model outperformed experienced caseworkers with fewer false positives and more race/ethnicity equity is an agency self-claim with no published methodology or independent verification.

Sources: lecher2025, newyorkstatecomptroller2023

Appears on: /domains/cases/nyc-acs-qa-risk-algorithm

EmpiricalColorado's legislature built a distinctive oversight topology — a standing independent statutory ombudsman (created 2010…

Colorado's legislature built a distinctive oversight topology — a standing independent statutory ombudsman (created 2010; an independent judicial-department agency since 2016 under C.R.S. 19-3.3) was handed a funded, nine-criteria audit mandate (HB 24-1046, signed May 28, 2024; $109,392 appropriation) over the statewide Family Safety and Family Risk Assessment tool suite used by 64 county-administered agencies — and the resulting ICF audit (176 pages, dated February 27, 2026, released March 2, 2026, with 50 recommendations across nine directives) found the tools well-aligned with policy on paper but inconsistently implemented: only 32% of surveyed caseworkers complete assessments in real time, vignette agreement fell to 59–63% on neglect and poverty scenarios versus 91–96% on abuse scenarios, and the audit recommends redefining safety and risk with observable behavior-based criteria and revising or replacing the actuarial risk tool.

Sources: icfincorporated2026, steffen2026, coloradogeneralassembly2024a, coloradorevisedstatutes2024

Appears on: /domains/cases/colorado-safety-risk-tools-audit

EmpiricalColorado's statutory audit loop fired end to end on the way up — the Child Protection Ombudsman's complaint-stream obser…

Colorado's statutory audit loop fired end to end on the way up — the Child Protection Ombudsman's complaint-stream observations became testimony to the 2023 Child Welfare System Interim Study Committee (including its brief stating the Family Safety Assessment had never been validated since its 1999 inception), the testimony became HB 24-1046 requiring the ombudsman rather than the child welfare agency to procure a third-party audit, and ICF's competitively procured audit returned to named legislative committees by the March 1, 2026 statutory deadline with four public information sessions following — but as of July 2026 the loop has produced statute, budget, a public audit artifact, and public information only: no follow-up legislation has been enacted, no formal CDHS response has been identified, and the tools remain in statewide operation unchanged, a decade after the state auditor's October 2014 performance audit flagged overlapping assessment and documentation weaknesses without producing redesign.

Sources: coloradogeneralassembly2024, childprotectionombudsmanofco2026, steffen2026, officeofcoloradoschildprotec2023, coloradoofficeofthestateaudi2014, icfincorporated2026

Appears on: /domains/cases/colorado-safety-risk-tools-audit

EmpiricalThe ICF audit found Colorado's actuarial Family Risk Assessment 'heavily weights prior reports, which inflate risk score…

The ICF audit found Colorado's actuarial Family Risk Assessment 'heavily weights prior reports, which inflate risk scores for incidents that occurred long ago, and disproportionately impacts minority families and perpetuates systemic bias' — with historical risk factors that 'disproportionately elevate scores for families of color' and a DV history item that codes any past incident, even decades old, as recent, penalizing survivors — while the record substrate degrades the oversight signal meant to catch exactly this: 68% of surveyed caseworkers cannot complete assessments in real time (documentation is back-filled into Trails after decisions), and race and ethnicity are so inconsistently documented that the audit says the disproportionality analysis requested by the legislature was limited; these are structural and qualitative audit findings, as no administrative-data disparity analysis has been published.

Sources: icfincorporated2026

Appears on: /domains/cases/colorado-safety-risk-tools-audit

EmpiricalFamily-Match, a proprietary two-sided 'relational fit' adoption matching algorithm built for Adoption-Share by former eh…

Family-Match, a proprietary two-sided 'relational fit' adoption matching algorithm built for Adoption-Share by former eharmony researchers, produced 1 known adoption in Virginia's two-year test (an official's statement to the AP; VDSS said the tool 'had not proven effective'; the pilot's end is undated in the record) and 2 adoptions in Georgia's year-long pilot ended October 2022; caseworkers in Florida, Georgia, and Virginia said it often led them to unwilling families, and Virginia social workers were perplexed that the algorithm seemed to match all the children with the same group of parents. In Florida the vendor's own quarterly report claimed 603 placements yielding 431 adoptions over five years — figures partner agencies could not verify: FamiliesFirst Network's records showed 76 Family-Match placements with no documented adoption plus 3 failed trial placements since 2019, and Children's Network of Southwest Florida counted 22 matches and 8 adoptions in five years while making hundreds of matches and hundreds of adoptions without the tool over the same period.

Sources: hoandburke2023a, fortuneassociatedpressrepubl2023

Appears on: /domains/cases/family-match-adoption-share

EmpiricalThe outcome evidence for Family-Match lived in a vendor-owned, credit-claiming data store: Virginia officials said that …

The outcome evidence for Family-Match lived in a vendor-owned, credit-claiming data store: Virginia officials said that once families' data was entered 'Adoption Share owned the data,' assistant director Traci Jones said 'We did not have access to the algorithm even after it was requested,' and agencies 'couldn't explain Family-Match's self-reported data.' The vendor's April 2023 'confidential' user guide instructed caseworkers not to delete cases matched outside the tool but to document them in the system 'so that Adoption-Share could refine its algorithm and follow up with the families'; Miami's Citrus Family Care Network said vendor staff asked social workers to have parents register in the tool even when it played no role in the adoption, and Georgia's spokesperson said Family-Match could claim credit for pairings already in its system. In Georgia the vendor-owned store holds whether foster youth have been sexually abused, the gender of their abuser, criminal records, and whether they identify as LGBTQIA — data typically restricted to secured child protective services case files — and two Florida agencies fed the system data with no contract and would not say how children's data was secured.

Sources: hoandburke2023a, fortuneassociatedpressrepubl2023

Appears on: /domains/cases/family-match-adoption-share

EmpiricalFamily-Match's four-state lifecycle turned on procurement authority actions made while states were structurally dependen…

Family-Match's four-state lifecycle turned on procurement authority actions made while states were structurally dependent on the vendor's own ledger for outcome evidence: Georgia ended its pilot on null results in October 2022, then — after Ramirez met with the governor's office and lobbied a statehouse committee — signed a new agreement in July 2023 for free; Virginia dropped the matching pilot as 'not proven effective' and by 2022 awarded Adoption-Share a larger contract for the Faster Families Highway recruitment portal ($188,000 budgeted / $212,546 spent SFY2023, $246,100 planned SFY2024, all 120 local departments enrolled by December 31, 2022, with policies and procedures — including family removal and response-time expectations — drafted only from February 2023, after statewide enrollment); Florida converted a philanthropy-funded rollout to a $350,000 DCF contract in October 2023 and added a Florida DOH contract for a medically-complex-children algorithm; and Tennessee, the only state whose reviewers formally questioned pre-deployment why the tool needed certain sensitive data points and how they influenced the match score, scrapped the rollout before it began. No state has ever published an evaluation of the matching pilot.

Sources: hoandburke2023a, virginiadepartmentofsocialse2023

Appears on: /domains/cases/family-match-adoption-share

EmpiricalFive US states (Michigan 2001, Minnesota 2001/2006, Maryland 2009, Texas 2013, Missouri 2021) operate 'Birth Match' prog…

Five US states (Michigan 2001, Minnesota 2001/2006, Maryland 2009, Texas 2013, Missouri 2021) operate 'Birth Match' programs in which an identity record-join between new birth registrations and a registry of parents with prior terminations of parental rights, serious-harm findings, or specified convictions fires a mandatory child-protective response with no risk score, no threshold, and no screening discretion: Michigan's current policy (PSM 712-2, reissued 2026-04-01) permits screen-out only for an inaccurate identity match or an already-open case and otherwise requires the referral to be screened in and assigned for investigation, and Maryland's 2025 response policy limits screen-out to a non-qualifying conviction or an adopted child while requiring a 24-hour intake coded 'Risk of Harm: birth match' — a mandatory response with no screening discretion, formally structured in Maryland as a voluntary non-CPS assessment families may refuse, with tracing and escalation duties attached. The only tunable parameters anywhere are registry scope and lookback: 2 years in Texas, 10 in Maryland and Missouri, unlimited in Michigan (records to 1978) and Minnesota.

Sources: cohen2022, michigandepartmentofhealthan2026, marylanddepartmentofhumanser2025, minnesotarevisorofstatutes2021

Appears on: /domains/cases/us-birth-match

EmpiricalThe Birth Match trigger is self-arming: the registry that fires it is populated by the system's own outputs, so a birth-…

The Birth Match trigger is self-arming: the registry that fires it is populated by the system's own outputs, so a birth-match response that ends in a new termination of parental rights writes the parent back onto the match list — permanently in Michigan and Minnesota, which have no lookback limit or retirement mechanism — and every subsequent birth re-fires with strictly increasing history; Michigan additionally adds manual list entries for severe cases without any termination. Maryland law accelerates the loop by letting the agency skip reasonable reunification efforts on account of a prior involuntary termination, and the 2025-26 Maryland legislative fight (HB 944 of 2025, died without a committee vote; HB 48 of 2026, heard 2026-01-29 and dead in committee without a vote when the session adjourned sine die on April 13, 2026) targeted that waiver, not birth-match repeal — with Civil Rights Corps testifying that the combination is 'functionally state sterilization' and coerces parents into 'voluntary' relinquishment, an advocacy characterization in the legislative record.

Sources: cohen2022, richardsonandsteckelcivilrig2026, marylandgeneralassembly2026, michigandepartmentofhealthan2025

Appears on: /domains/cases/us-birth-match

EmpiricalThe Birth Match automation runs without ownership or evaluation: match volumes fell by more than half in Michigan (1,186…

The Birth Match automation runs without ownership or evaluation: match volumes fell by more than half in Michigan (1,186 in FY2019 to 515 in FY2021) and Maryland (243 in CY2019 to 124 in CY2020) without the operating agencies noticing or being able to explain the change; fewer than 10 percent of matched families in any state with data received services (Maryland: 5 of 124 matched families in the source's mixed FFY2020/CY2020 pairing); Maryland's unanimous 2018 expansion statute ordered an independent evaluation of the match's sensitivity, specificity, and predictive value that a public-records request found was never implemented; and the only oversight that demonstrably changed the system — child fatality review teams auditing the matcher's misses — widened the trigger both documented times. Nearly all of these figures flow through one proponent researcher's ad hoc public-records requests (AEI 2022), which the author herself flags as limited and of dubious accuracy; no state publishes birth-match data in any regular report.

Sources: cohen2022, cohen2022a, marylandgeneralassembly2018, baltimorecitychildfatalityre2017

Appears on: /domains/cases/us-birth-match

EmpiricalDC's Child and Family Services Agency published a 16-page pre-deployment AI Values Alignment Report (one of only three d…

DC's Child and Family Services Agency published a 16-page pre-deployment AI Values Alignment Report (one of only three district-wide under Mayor's Order 2024-028) for CORA, a staff-facing policy chatbot launched June 16, 2025 inside its STAAND case-management system, whose centerpiece guardrail is a read-side exclusion — 'the chatbot does not draw from confidential case files' and has no access to STAAND case data; by December 2025 the agency's own tip sheets documented a CORA phone app that ingests case documents, photos of handwritten notes, and voice dictation and saves AI-drafted contact notes into the STAAND case record after worker approval — the write direction the no-case-files rule never governed, superseding the May 2025 report's dated statement that generated text is 'guidance, not text to be used by the employee as part of case documentation.'

Sources: dcchildandfamilyservicesagen2025b, dcchildandfamilyservicesagen2025, dcofficeofthechieftechnology2026

Appears on: /domains/cases/dc-cfsa-cora-chatbot

EmpiricalCFSA's May 2025 AI Values Alignment Report concedes that per-output human validation of CORA is impossible — asked wheth…

CFSA's May 2025 AI Values Alignment Report concedes that per-output human validation of CORA is impossible — asked whether humans can review and approve AI outputs before they are enacted, it answers 'No,' with validation occurring before content enters the knowledge source (SME validation of every document, annual committee re-review of the corpus) and afterward via quality assurance (monthly internal accuracy reports and a real-time answer-flagging channel, none published); by December 2025 the tool's own landing card carried the in-product warning 'The Ask Feature is under construction and answers may not be fully accurate. Consult your supervisor as needed,' delegating per-query vigilance to workers and supervisors six months after launch.

Sources: dcchildandfamilyservicesagen2025b, dcchildandfamilyservicesagen2025a

Appears on: /domains/cases/dc-cfsa-cora-chatbot

EmpiricalCORA operates as a single-agency closed loop with no external examination of the running system: its knowledge authors, …

CORA operates as a single-agency closed loop with no external examination of the running system: its knowledge authors, tool owners, operators, and oversight committees are all CFSA units; no OIG or auditor review, no published accuracy or usage data, and no DC Council finding exists as of mid-2026; the deployment sits in CFSA's first era without a court monitor in three decades (LaShawn A. oversight ended 2021, final closure 2022); the 25,000 queries per month figure is a budget ceiling, not observed usage; and every quantitative outcome figure (45 minutes saved per intake report, one to four hours per case, three-week feature delivery at a claimed 20x lower cost) is an unaudited vendor claim from a customer story that never names CORA, so CORA-specific claims rest on agency sources alone.

Sources: dcchildandfamilyservicesagen2025b, microsoft2025, dcchildandfamilyservicesagen2021

Appears on: /domains/cases/dc-cfsa-cora-chatbot

EmpiricalIn Richmond, Virginia, Predict-Align-Prevent's open-source place-based model (built with Urban Spatial; report byline Ke…

In Richmond, Virginia, Predict-Align-Prevent's open-source place-based model (built with Urban Spatial; report byline Ken Steif, Matthew D. Harris, and Sydney Goldstein, 2019) reported its highest risk tier capturing about 70% of held-out child-maltreatment events against about 35% for a kernel-density baseline, while covering about 10% of city land area holding roughly 48,500 residents including about 8,200 children — builder-generated figures from a commissioned report — and the builders' own published fairness audit stated verbatim that the meta-model 'generalizes well across neighborhoods of varying poverty rates, but does not generalize well across neighborhoods of varying race,' despite race and income being excluded from the feature set.

Sources: steif2019, urbanspatial2019

Appears on: /domains/cases/predict-align-prevent-geospatial

EmpiricalNew Hampshire's 2018-2023 federal Community Collaborations cooperative agreement is the only documented operational plan…

New Hampshire's 2018-2023 federal Community Collaborations cooperative agreement is the only documented operational planner use of Predict-Align-Prevent's maps: state narratives confirm the mapping (funded by Casey Family Programs, using address-level inputs accessed through stewarded databases) was completed for all three service areas before sites received implementation resources, and the March 2024 ACF/OPRE grantee profile documents that Community Implementation Teams used PAP-identified areas of need to target family outreach and issued funded requests for proposals with stipends — documented resource-routing behavior change, with no independent outcome evaluation of maltreatment effects in any deployment.

Sources: administrationforchildrenand2024, stateofnewhampshire2020

Appears on: /domains/cases/predict-align-prevent-geospatial

EmpiricalThe 2016 Fort Worth study (Daley et al., Child Abuse & Neglect), trained on 2013 state child-welfare substantiations and…

The 2016 Fort Worth study (Daley et al., Child Abuse & Neglect), trained on 2013 state child-welfare substantiations and Fort Worth police data, reported the top 10% of grid cells capturing 52% of 2014 maltreatment cases against 43% for a conventional hotspot model — figures independently corroborated verbatim by the Marchment and Gill (2021) Crime Science systematic review, which also found it to be the single child-maltreatment application of risk terrain modeling in the reviewed literature; Fort Worth remained a published study plus local advocacy, with no evidence of operational resource allocation by the maps.

Sources: daley2016, marchmentandgill2021

Appears on: /domains/cases/predict-align-prevent-geospatial

EmpiricalIn December 2025 the Massachusetts Department of Transitional Assistance piloted an Accenture-built tool that transcribe…

In December 2025 the Massachusetts Department of Transitional Assistance piloted an Accenture-built tool that transcribes SNAP eligibility calls in real time and generates a structured, caseworker-editable summary that is saved into BEACON, the state's benefits eligibility system of record; full transcripts are not retained, only the summaries, according to technical documentation reviewed by The Shoestring, and the summarization prompt was withheld as proprietary. About 400 calls had been processed by the April 2026 reporting, against roughly 45,000 calls connected to staff per month, in a record system that fed 31,390 SNAP application dispositions in December 2025 alone. No independent evaluation, inspector-general audit, or completed privacy impact assessment of the tool is on record.

Sources: theshoestring2026, massachusettsdepartmentoftra2025

Appears on: /domains/cases/massachusetts-dta-call-summaries

EmpiricalOversight of the DTA call summarizer ran through the labor channel: SEIU Local 509, representing DTA call-center workers…

Oversight of the DTA call summarizer ran through the labor channel: SEIU Local 509, representing DTA call-center workers among roughly 9,000 state employees, reached an agreement with DTA — pre- or early-deployment; the record does not establish it preceded the December 2025 rollout — that made worker use voluntary and protected jobs, without addressing record provenance or client-side safeguards. The formal privacy apparatus sat empty: none of the nine AI use cases Massachusetts disclosed from its internal inventory of at least 40, the DTA summarizer included, reported a completed privacy impact assessment; the tool's interaction-data plan was listed as still to be defined while it was in production; and details on the other 31 use cases were withheld until the Supervisor of Records ordered them submitted for in camera review.

Sources: theshoestring2026, massachusettsexecutiveoffice2026

Appears on: /domains/cases/massachusetts-dta-call-summaries

EmpiricalThe record store the AI summaries enter was itself contested while the pilot ran: in California v. USDA, No. 3:25-cv-063…

The record store the AI summaries enter was itself contested while the pilot ran: in California v. USDA, No. 3:25-cv-06310 (N.D. Cal.), a 21-state-plus-DC coalition including Massachusetts obtained preliminary injunctions on October 15, 2025 and February 27, 2026 blocking USDA from cutting SNAP funding over states' refusal to hand over personal SNAP applicant and recipient data, the court holding the proposed data protocol would likely permit sharing beyond the entities allowed under 7 U.S.C. 2020(e)(8). Both orders are preliminary and the litigation is live; whether AI-generated call summaries held in BEACON fall within the demanded data is unresolved, and the Massachusetts AG's office declined to comment on that question.

Sources: massachusettsattorneygeneral2026, jurist2026, theshoestring2026

Appears on: /domains/cases/massachusetts-dta-call-summaries

EmpiricalIllinois DCFS rolled out a vendor natural-language-processing layer statewide over its legacy SACWIS case-management sys…

Illinois DCFS rolled out a vendor natural-language-processing layer statewide over its legacy SACWIS case-management system beginning in 2023, provisioning roughly 7,500 users with a DCFS-estimated 20 percent administrative-time saving; the agency's federal Annual Progress and Services Report states that all notes in SACWIS can be mined, rendered as per-case Risks and Strengths overviews with click-through to source notes. By the vendor CEO's November 2025 congressional testimony, more than 6,000 Illinois staff use the tool and it saves each social worker about five hours per week — vendor figures delivered in a hearing whose majority summary records no AI-risk scrutiny — and no independent accuracy evaluation, published usage metric, or audit of the deployment has been located.

Sources: elisco2025, illinoisdepartmentofchildren2024, governmenttechnology2023, housewaysandmeanscommittee2025

Appears on: /domains/cases/illinois-dcfs-augintel

EmpiricalFrom the 2023 announcement onward the same platform carried a second, management-facing channel reading the same note st…

From the 2023 announcement onward the same platform carried a second, management-facing channel reading the same note store: early-warning signs, caseworker safety, practice-model fidelity, and statewide trends per the launch release; compliance-related data collected in the background from the notes workers already document, service-outcome mining tied to funding accountability, and vendor-analytics deep dives informing where programs should be established or eliminated, per the vendor CEO's 2025 testimony; and a named Chapin Hall partnership using machine learning over case notes to measure practitioners' use of Motivational Interviewing. The associated documentation-behavior effect on the workers whose notes are mined is a structural inference, not a documented finding.

Sources: augintel2023, elisco2025, chapinhallattheuniversityofc2025

Appears on: /domains/cases/illinois-dcfs-augintel

EmpiricalIn December 2017 the same agency, Illinois DCFS, terminated its Eckerd Rapid Safety Feedback predictive-analytics pilot …

In December 2017 the same agency, Illinois DCFS, terminated its Eckerd Rapid Safety Feedback predictive-analytics pilot — a 366,000-dollar sole-source arrangement that scored more than 4,100 children at 90-percent-plus probability of death or serious injury while assigning low risk to two children who died — after a joint OEIG/DCFS-OIG report found the arrangement had been misclassified as a grant rather than a no-bid contract. That shutdown is the only oversight event at this agency documented to have changed system behavior, and it preceded the current deployment's deliberately score-free extraction design.

Sources: theimprint2017, governmenttechnologya

Appears on: /domains/cases/illinois-dcfs-augintel

EmpiricalCalifornia's tax agency ran the first state generative AI call-center assistant through a complete governed procurement …

California's tax agency ran the first state generative AI call-center assistant through a complete governed procurement arc: a competitive sandbox in which two vendors were each paid one dollar to test for six months in a secure environment, a 10-month pilot in a simulated environment, and production rollout to roughly 375 agents completed August 22, 2025 under a 12-month, 445,000 dollar contract. The pilot projected a minimum 1.5 percent per-call time saving, and by May 2026 the agency had declined to renew the contract, with the official who launched the project saying the working system 'didn't quite save as much time as we had hoped.'

Sources: stateofcaliforniagenaiportal2024, officeofthegovernorofcalifor2025, governmenttechnologyindustry2025, melhado2026

Appears on: /domains/cases/california-cdtfa-genai-call-center

EmpiricalThe benefit figures for the CDTFA call-center assistant are pilot projections from a simulated environment, self-reporte…

The benefit figures for the CDTFA call-center assistant are pilot projections from a simulated environment, self-reported by the agency through trade press and echoed in vendor marketing: the call-center chief said the simulated environment 'did show some potential to save at least one and a half percent time on some of those calls,' extrapolated to roughly 100,000 minutes per year and capacity for about 10,000 additional calls annually across roughly 800,000 yearly inquiries. No published production-environment measurement, methodology, or independent assessment report has been located.

Sources: governmenttechnologyindustry2025, symsoftsolutionsviabusinessw2025

Appears on: /domains/cases/california-cdtfa-genai-call-center

EmpiricalBy May 2026 CDTFA had declined to renew the 12-month, 445,000 dollar contract for its custom call-center assistant and s…

By May 2026 CDTFA had declined to renew the 12-month, 445,000 dollar contract for its custom call-center assistant and shifted to a similar off-the-shelf tool from Amazon Web Services under an existing larger contract, with Government Operations Secretary Nick Maduros saying the solutions 'worked in practice' but 'didn't quite save as much time as we had hoped.' The decision was a procurement authority action inside a broader portfolio review of eight executive-order-driven generative AI pilots totaling more than 5.8 million dollars, in which three other projects also ended while three continued.

Sources: melhado2026

Appears on: /domains/cases/california-cdtfa-genai-call-center

EmpiricalNew Jersey built and hosts its own generative drafting assistant for state employees, with a state-owned interface, host…

New Jersey built and hosts its own generative drafting assistant for state employees, with a state-owned interface, hosting and logs and a hosted commercial frontier-model service as the one external dependency in the serving path, an ownership arrangement that leaves the underlying model replaceable without reprocuring an application or moving employees onto a different product. The state reports that roughly 20,000 employees had used the tool across more than 300,000 sessions and more than 1,000,000 prompts by February 2026 at about one dollar per user per month, that access onboarding routes through a responsible-AI course whose curriculum is used by 25 or more states, and that unemployment-insurance staff rewriting claimant emails in plain language saw claimants respond 35 percent faster. Every outcome figure is state self-reported; the 35 percent figure has no published methodology and predates the statewide launch, and no inspector-general evaluation, external audit or peer-reviewed causal study of the tool was located.

Sources: njofficeofinnovation2026a, routefifty2025, statescoop2024, njofficeofinnovation2024, sofi2026

Appears on: /domains/cases/nj-ai-assistant

EmpiricalNew Jersey's joint policy circular 25-OIT-001 requires human review of all AI-generated content for accuracy, bias, comp…

New Jersey's joint policy circular 25-OIT-001 requires human review of all AI-generated content for accuracy, bias, completeness, accessibility and style; permits sensitive personal information only inside state-approved tools, naming the NJ AI Assistant, with Agency CIO approval; and requires State Chief Technology Officer clearance plus registration for resident-facing or decisional generative systems, a gate the staff-facing assistant has not been reported to pass. The same circular's text says all state employees 'should' take the responsible-AI course, a should-language obligation stated in the governing document itself, while the focal unemployment-insurance agency is reported to have reached full training and tool coverage of its staff. A statewide public-workforce survey preceded the November 2024 AI Task Force recommendations to the Governor, and user surveys and interviews drove the March 2026 rebuild, which added a visible reasoning section intended to help employees catch errors and hallucinations.

Sources: njofficeofinformationtechnol2025, stateofnewjerseyaitaskforce2024, njofficeofinnovation2026, routefifty2025

Appears on: /domains/cases/nj-ai-assistant

EmpiricalNew Jersey's unemployment-insurance and TDI/FLI teams built plain-language glossaries, reusable prompt libraries and qua…

New Jersey's unemployment-insurance and TDI/FLI teams built plain-language glossaries, reusable prompt libraries and quality-evaluation rubrics for AI-assisted translation into Spanish and Haitian Creole, with human review by professional translators, subject-matter experts and seven community organizations; the state reports a higher share of Spanish-language unemployment-insurance applications and better follow-through, without quantifying either. US Digital Response gave the department a 2025 SEED Award and republished the materials for reuse by other states, the associated responsible-AI curriculum is used by 25 or more states and local partners, and the March 2026 rebuild moved training guidance inside the tool itself.

Sources: newjerseydepartmentoflaboran2025, usdigitalresponse2025, routefifty2025, njofficeofinnovation2026, innovateus2024

Appears on: /domains/cases/nj-ai-assistant

EmpiricalThe Social Security Administration deployed a conversational question-and-answer chatbot on its national 800-number in A…

The Social Security Administration deployed a conversational question-and-answer chatbot on its national 800-number in April 2025, answering 74 frequently asked questions before a caller reaches an employee. Automation on that line went from roughly 300,000 handled calls a month in fiscal 2024 to roughly 2.9 million a month in fiscal 2025, peaking at 5.1 million automated calls in March 2025, while the agency served 68 million callers, a 65 percent increase over the prior year, with a workforce that fell about 10 percent net from roughly 57,000 to roughly 51,400 and about 1,000 field office employees reassigned onto 800-number duty. About 25 million fiscal 2025 calls ended with no service, abandoned in queue or met with a busy signal, and the busy rate spiked to 29.4 percent in March 2025 during the Social Security Fairness Act surge, which affected 3.2 million beneficiaries.

Sources: ssaofficeoftheinspectorgener2025, kffhealthnewsdariustahir2025, aarp2025

Appears on: /domains/cases/ssa-800-number-ai-assistant

EmpiricalThe agency's inspector general found the published telephone metrics arithmetically accurate as computed and documented …

The agency's inspector general found the published telephone metrics arithmetically accurate as computed and documented what they cover: the headline average speed of answer, 12.7 minutes in October 2024, a 29.7 minute peak in January 2025 and 7.0 minutes in September 2025, counts a caller who accepts a callback as a zero wait, and the roughly 25 million calls a year ending in a hang-up or a busy signal are excluded. The 23.8 million callers who accepted callbacks in fiscal 2025 waited an average of 61.5 to 151.8 minutes depending on the month, and the 9.3 million who declined and held waited an average of 18.8 to 100.1 minutes. A separate 87 percent first-contact-resolution figure comes from a post-call survey that is not offered to callers served only by automation, and no audit has published a misrouting or wrongful-disconnection rate for the chatbot itself.

Sources: ssaofficeoftheinspectorgener2025, nextgovfcw2025b

Appears on: /domains/cases/ssa-800-number-ai-assistant

EmpiricalDeployment decisions on this line have followed changes of leadership rather than published evaluation results. The chat…

Deployment decisions on this line have followed changes of leadership rather than published evaluation results. The chatbot was developed and tested under one administration and shelved as not ready, with the former chief information officer saying the team wanted to ensure the automation produced consistent and accurate answers and that this would take more time; the next administration deployed it in April 2025 and targeted extension to roughly 1,200 field offices by August 2025. A separate phone anti-fraud AI check was added in April 2025 and its three-day claim holds removed in mid-May 2025 after it flagged 2 of more than 110,000 claims while slowing retirement claim processing by 25 percent, a different tool whose figures do not merge with the chatbot's. The predecessor Next Generation Telephony Project, built by Verizon Business Network Services under a February 2020 contract, was abandoned on August 22, 2024 after about ten months and more than 160 million dollars paid, with an audit finding that contract lacked performance-based quality standards; the vendor of the current cloud platform is not named in the public audit. Five commissioners or acting commissioners served during fiscal 2025, and the metrics published on the agency's performance website were added and removed according to what each leadership believed were the most important metrics for the public.

Sources: kffhealthnewsdariustahir2025, nextgovfcw2025a, officeofsenatorelizabethwarr2025, ssaofficeoftheinspectorgener2025b, ssaofficeoftheinspectorgener2025

Appears on: /domains/cases/ssa-800-number-ai-assistant

EmpiricalCalifornia's Employment Development Department runs a two-tier chat assistant delivered under the Integrated Contact Cen…

California's Employment Development Department runs a two-tier chat assistant delivered under the Integrated Contact Center work stream of EDDNext, the state's roughly $1.258 billion modernization of its unemployment, disability and paid family leave systems. The unauthenticated public-site tier became available around the clock in the state's top eight working-age languages in May 2025 and served 554,792 unique customers across 2,103,782 messages between January 1 and June 30, 2025. A live agent chat channel for unemployment customers, which the department dates to July 2025, lets a customer escalate to a person on weekdays between 9 a.m. and 2 p.m. after identity verification, with account details passed to the agent, real-time machine translation in six non-English languages, a redacted transcript saved to the account and a post-chat survey. On May 8, 2026 a second, authenticated tier launched inside the customer portal, answering a signed-in unemployment customer's own claim status, payment and eligibility questions for claims filed in the past three years; the department reported more than 25,000 uses and nearly 18,000 fully self-service interactions in its first two weeks. The platform is documented as intent-based conversational AI; the public sources do not establish generative language modeling. All usage figures are agency self-reported, and the two tiers' counts belong to different systems.

Sources: californiaemploymentdevelopm2023, californiaemploymentdevelopm2025a, californiaemploymentdevelopm2025d, californiaemploymentdevelopm2025c, californiaemploymentdevelopm2026

Appears on: /domains/cases/california-edd-virtual-assistant

EmpiricalNo body oversees the EDD chat assistant as such: its accountability is inherited two levels up, from legislative and exe…

No body oversees the EDD chat assistant as such: its accountability is inherited two levels up, from legislative and executive scrutiny of the EDDNext programme it is a deliverable of. That scrutiny is substantial and has demonstrably changed programme behaviour — the 2026-27 Governor's Budget reverts $70.6 million of unused modernization funding early as a budget solution, and the core claims-system replacement was resequenced to do disability and paid family leave first with unemployment integration designated 'mandatory optional' to reduce risk for the state — and it is also documented as partial, since the adopted 2025-26 Budget Act retained extended spending authority against the Legislative Analyst's Office recommendation to drop it. What the oversight record engages with is budget, schedule and procurement: the 2026 analyst-office questions on the record concern the core project's vendor and the rationale for removing unemployment from it, and the February 2026 handout flags that new front-end functionality is linked to the legacy claims mainframe through 'informal and untested data bridges and custom-built interfaces' that have not been stress tested — a finding stated generically about new functionality, without naming the chat assistant. No inspector-general review, state-auditor evaluation or academic study of this assistant's answer accuracy or translation fidelity has been located in the public record.

Sources: californialegislativeanalyst2025, californialegislativeanalyst2025c, californialegislativeanalyst2026, californialegislativeanalyst2026a, californiastatesenate2026

Appears on: /domains/cases/california-edd-virtual-assistant

EmpiricalThe success measures published for the EDD chat assistant are deflection-shaped, and the department both produces them a…

The success measures published for the EDD chat assistant are deflection-shaped, and the department both produces them and reports them upward. Since July 2025 the agency states that 23 percent more customers resolve questions through self-service and 47 percent fewer need to speak with an agent after using it, alongside more than 2.1 million self-service actions since November 2024, more than 830,000 customers using callback since May 2024, more than 29,900 customers served by live-chat agents, and an in-chat issue-resolution rate above 80 percent. Those figures measure channel exit rather than verified problem resolution, and they are agency self-reported; the human channel they are measured against runs weekdays from 9 a.m. to 2 p.m. and reached nearly 6,000 unemployment customers a month as of October 2025. Separately, the integrator published on its own marketing blog that the bot stack reduced call volume by 35,000 calls a day, cut live-agent volume by 3,800 calls a day, saved constituents 684 hours daily and generated $8.4 million in annual savings; those are vendor claims with no independent verification.

Sources: californiaemploymentdevelopm2025b, californiaemploymentdevelopm2025c, intervisionsystems2025

Appears on: /domains/cases/california-edd-virtual-assistant

EmpiricalMassachusetts built a retrieval-grounded generative-AI Virtual Assistant in house in 83 days and launched it on motor-ve…

Massachusetts built a retrieval-grounded generative-AI Virtual Assistant in house in 83 days and launched it on motor-vehicle registry pages in April 2025 ahead of the May 7, 2025 federal identification deadline, describing it as a state-owned platform that replaced a vendor-managed rule-based chatbot; by 2026 the assistant's own support page listed motor-vehicle, toll, unemployment-assistance, tax, child-support, transitional-assistance, family-and-medical-leave and state-login content. The state reports more than 1,500 conversations a day and roughly 200,000 conversations since launch, a chat open rate rising from 1.29 to 2.92 percent, positive feedback rising from 10 to 56 percent, negative feedback down 60 percent, page abandonment down 40 percent, around-the-clock availability in English, Spanish and Portuguese, a knowledge base refreshed nightly from published state content, and a custom evaluation tool that flags answer issues within hours. Every one of those figures is an agency self-report; the technology agency withheld the cost and usage reports and the technical logs that would confirm its claims, and the underlying foundation model and hosting stack have not been publicly identified.

Sources: massachusettsdigitalservice2025, executiveofficeoftechnologys2026, executiveofficeoftechnologys2025, theshoestring2026, massachusettsdigitalservice2026

Appears on: /domains/cases/massgov-virtual-assistant

EmpiricalMassachusetts policy AI.001, effective January 31, 2025, requires human fact-checking of generative output, conspicuous …

Massachusetts policy AI.001, effective January 31, 2025, requires human fact-checking of generative output, conspicuous labeling of AI content, Chief Technology Officer approval for generative procurement and a generative-AI inventory, while routing privacy review through consultation with legal and security teams; the phrase privacy impact appears in it zero times. The Enterprise Privacy Office told the legislature in February 2025 that it has been using a Privacy Impact Assessment and was overseeing a pilot program with its risk and security teams to assess privacy risks during the contracting and development stages. Records obtained by an independent investigation showed at least forty AI use cases in the agency's internal survey with thirty-one withheld, and of the nine entries released not one recorded a completed privacy impact assessment, the fields left blank without explanation including for tools processing Social Security numbers and Medicaid data, with the agency spokesperson declining to answer questions about them; a count across all forty is therefore an inference the agency has not contradicted rather than a verified number, and the assessment is an internal office practice rather than a statutory mandate. After months of negotiation the Supervisor of Records ordered the agency to produce the withheld records for in camera review, allowing ten business days to produce and up to fifteen business days to review; no final determination was found as of July 20, 2026.

Sources: theshoestring2026, commonwealthofmassachusettse2025, executiveofficeoftechnologys2025, massachusettsexecutiveoffice2026

Appears on: /domains/cases/massgov-virtual-assistant

EmpiricalThe Massachusetts assistant answers from a knowledge base rebuilt nightly out of published state content, and the state …

The Massachusetts assistant answers from a knowledge base rebuilt nightly out of published state content, and the state reports that those published pages were updated and simplified to suit the assistant while chat analytics drive prompt tuning, content updates and enhancements, so the authoritative public statement of program rules is adapted to the channel that reads it. On the human-channel side, an independent institute brief reporting state figures records motor-vehicle registry calls falling by about 1,000 a day and emails by about 200 a week after launch, counts about twenty AI use cases publicly reported with three facing the public against forty in the internal survey, and recommends a formal public AI inventory and structured user feedback loops; the assistant offers three languages where the predecessor rule-based chatbot offered twenty. The state technology secretary committed publicly that there will be a human reviewing output before public distribution and framed AI as relieving stretched agencies without adding headcount, while a public-employee union representing about 9,000 state employees argued that staffing rather than AI tooling is what would let the transitional-assistance agency serve more clients.

Sources: massachusettsdigitalservice2025, pioneerinstitute2026, commonwealthbeacon2026, theshoestring2026, executiveofficeoftechnologys2025

Appears on: /domains/cases/massgov-virtual-assistant

EmpiricalMyFriendBen is an advisory multi-state benefits screener: an anonymous survey of about six minutes returns a report of p…

MyFriendBen is an advisory multi-state benefits screener: an anonymous survey of about six minutes returns a report of programs a household is likely eligible for, with estimated dollar values, and submits nothing, routing people instead to separate government application channels that make every determination independently. All of its eligibility math is computed by PolicyEngine, a separate nonprofit whose open-source codebase encodes federal and state statute; MyFriendBen deliberately built no proprietary rules logic, and the same substrate sits under each of its state deployments, so an encoding error would misestimate benefits in every state simultaneously and a correction would propagate to every state simultaneously. That substrate's scale is indicated by a $300,000 NSF POSE Phase I award announced on August 18, 2025. The shared-substrate structure is documented in both organizations' technical materials; no incident of a cross-state encoding error appears in the public record. The accuracy figure attached to it, 'over 90%', is a builder, infrastructure-organization and operator claim with no published methodology and no independent audit.

Sources: policyengine2025a, githubmyfriendbenorg2026, policyengine2025b, codethedreammyfriendbenncben2025

Appears on: /domains/cases/myfriendben

EmpiricalThe gap between benefits identified and benefits received is measured only from outside the system. Independent journali…

The gap between benefits identified and benefits received is measured only from outside the system. Independent journalism reported that in all of 2023 in Colorado about 5,500 households were screened, about $30 million in benefits was identified and about $5 million was estimated to have actually been obtained, with the median user reporting an income just over $8,200 a year and a household size of two. The deployment's own impact totals are self-published or funder claims and are mutually inconsistent, ranging across 20,000+ Coloradans, 55,000+ households, 100,000+ households, 115,000+ households nationally and 65,000+ families, with dollar figures of $33 million, $52 million, $58 million and $1.2 billion identified, on denominators that shift between identified, applied for, unlocked and delivered. The one commissioned evaluation, by the Urban Institute, is described by a funder's post as currently under way; its interim figures from 301 users cover discovery, intent and satisfaction rather than calculation accuracy or verified enrollment, and its primary report was not located in open search. The Aspen Institute's Financial Security Program finds that no systematic evaluation exists of whether benefits screeners increase benefits access.

Sources: thecoloradosun2024, deltafund2026, prnewswiremyfriendbencpalrel2026, aspeninstitutefinancialsecur2024

Appears on: /domains/cases/myfriendben

EmpiricalThe deployments are operated by independent local anchors rather than by one organization: Code the Dream with NC 211 in…

The deployments are operated by independent local anchors rather than by one organization: Code the Dream with NC 211 in North Carolina, Benefit Illinois's Illinois Benefit Hub, the MASSCAP community-action network in Massachusetts and Child Poverty Action Lab in Texas, each configuring its own program catalog on a shared white-label backend that carries per-state feature flags and admits a program only if it is worth at least $300 a year. The front-line users are employed by those partner agencies and others: 2-1-1 Colorado piloted the tool with its own call staff in autumn 2023, and 100 to 130 individuals at one Dallas employer use it weekly, with one partner reporting that it serves about 4,000 individuals a year through it. The builder's stated reason for navigator uptake is that training a case manager on a single program's rules otherwise takes months. Screening data has also flowed outward into policy: a Colorado Child Tax Credit Calculator derived from the same code identified $177 million, and the builder credits the surrounding ecosystem with unlocking an $810 million state child tax credit through HB24-1311, a rule the shared substrate then had to encode. The catalog, threshold, uptake and policy-loop figures are builder, operator and joint-release claims.

Sources: garycommunityventures2025, prnewswiremyfriendbencpalrel2026, githubmyfriendbenorg2026, masscapmassachusettsassociat2026

Appears on: /domains/cases/myfriendben

EmpiricalSafeRent Solutions (formerly CoreLogic Rental Property Solutions) sold landlords a Registry ScorePLUS tenant-screening m…

SafeRent Solutions (formerly CoreLogic Rental Property Solutions) sold landlords a Registry ScorePLUS tenant-screening model returning a single 200-800 'lease performance risk' score plus an accept/decline/conditional recommendation measured against a landlord-chosen cutoff (500 in the complaint's worked example, with an optional 450-499 conditional band), built from credit bureau reports and scores including non-tenancy debt, bankruptcy records, past-due accounts, payment performance, and eviction and landlord-tenant court records, weighted 'according to their statistical significance in predicting lease performance'; factor weights were undisclosed to landlords, applicants and the public, and the company's marketing told buyers a landlord cannot change the screening algorithm. The plaintiffs' pleaded theory concerns a missing input: as of 2021, HUD data cited in the complaint shows Black and Hispanic voucher holders in Massachusetts paying an average of $423/month toward rent and utilities while public housing authorities paid landlords an average of $1,159/month directly (at least 73.26% of the expected payment), on tenancies averaging over 21 years in the same unit, and that subsidy is not a model input. These are plaintiff-pleaded figures accepted by the court as allegations at the motion-to-dismiss stage only.

Sources: louisv2022, louisv2023, findlawcaselaw2023

Appears on: /domains/cases/saferent-score-voucher-screening

EmpiricalThe authority over a SafeRent-screened tenancy decision was documented as inverted: the vendor authored the weights and …

The authority over a SafeRent-screened tenancy decision was documented as inverted: the vendor authored the weights and returned the answer without any housing relationship to the applicant, property management picked the cutoff 'in consultation with SafeRent' without knowing how scores are computed, and the leasing desk that signed the lease received only the bottom-line score ('The Leasing Manager does not receive the detailed credit information at the time of running the applicant screening'; 'CoreLogic sends us a number, and if it is above the predetermined approved number, we move forward... We do not know why they were denied') and disclaimed override authority to a named plaintiff in writing ('we do not accept appeals and cannot override the outcome of the Tenant Screening'). On July 26, 2023 Judge Angel Kelley denied the motions to dismiss the FHA 3604(a),(b) and Massachusetts c.151B race and source-of-income claims on the reasoning that SafeRent 'effectively controls' approval decisions because it alone builds and conceals the algorithm, while dismissing the c.93A consumer counts. The single documented pre-settlement reversal ran outside the screening system entirely: a July 22, 2021 denial, a personal appeal rejected August 26, and a second appeal drafted by the tenant-advocacy organization City Life/Vida Urbana delivered September 1-8, 2021 — roughly six weeks of organized advocacy for one reversal, from which no reversal rate can be generalized.

Sources: louisv2022, louisv2023, unitedstatesdepartmentofjust2023a, civilrightslitigationclearin2025

Appears on: /domains/cases/saferent-score-voucher-screening

EmpiricalThe Louis v. SafeRent settlement (executed March 28, 2024; finally approved November 20, 2024; $2.275M total, $1.175M fu…

The Louis v. SafeRent settlement (executed March 28, 2024; finally approved November 20, 2024; $2.275M total, $1.175M fund with no reversion) remedied the case on the output channel rather than in the model: for five years SafeRent may return no SafeRent Score, no other tenant screening score and no accept/decline recommendation on a report for a voucher-holder application, providing a report of underlying information instead (3.5.2); for the 'market' and 'no-credit' products the score is suppressed by default unless the landlord affirmatively certifies the applicant is not a voucher recipient (3.5.3); any tenant screening score may re-enter the voucher channel only if 'found to be valid when used for voucher-holders by the National Fair Housing Alliance' or another organization mutually agreed by class counsel and SafeRent (3.5.5(ii)(1)); third-party credit scores (FICO, VantageScore) may still be passed through with source disclosure (3.5.5(iii)); customer training is required (3.5.4); and the court retains continuing exclusive enforcement jurisdiction for five years from SafeRent's certification (6.9), approximately through 2029-2030. The geographic reach of the practice-change terms is ambiguous: 3.5.2 and 3.5.3 are not expressly limited to Massachusetts, while the classes, the notice population and the plaintiffs' framing are Massachusetts-centered, and some coverage characterizes the changes as company-wide. Suppression therefore shifts rather than eliminates the decision input. SafeRent settled without admitting liability and states it 'continues to believe the SRS Scores comply with all applicable laws' (vendor claim); as of 2026-07-20 no public post-settlement compliance report, enforcement motion or validating-organization event had been located, so the gate exists on court-approved terms and has never been observed operating. The >18,000 figure is Massachusetts applications scored below the housing provider's accept threshold between May 25, 2020 and September 27, 2023 (the records-pull window), an upper-bound flow figure that does not identify voucher status or race, paired with SafeRent's own October 2023 estimate of 3,300-4,200 putative class members.

Sources: louisv2024, epiqclassactionservices2024, associatedpressviafortune2024, greaterbostonlegalservices2024

Appears on: /domains/cases/saferent-score-voucher-screening

EmpiricalThe United States and ten plaintiff states allege in United States v. RealPage, Inc. (M.D.N.C., filed 23 August 2024, am…

The United States and ten plaintiff states allege in United States v. RealPage, Inc. (M.D.N.C., filed 23 August 2024, amended 7 January 2025) that competing landlords contractually fed a single vendor nonpublic, competitively sensitive information — executed new-lease rents, renewal offers and rates, lease terms and occupancy signals — to train and run a common pricing algorithm that recommended rents back to all of them: the complaint alleges at least 80 percent of the commercial revenue-management software market for multifamily housing and data agreements reaching over 16 million units nationwide including units of landlords who were not customers, describes an 'Auto Accept' setting that implemented daily recommendations with no human review, a 'Governor' feature alleged to constrain price decreases more than increases, and vendor pricing advisors who monitored client acceptance and pushed property managers toward compliance, and reports a national acceptance rate of 40 to 50 percent for new leases across January 2017 to June 2023 against internal analysis finding nearly 60 percent of final floor-plan prices within 2.5 percent of the recommendation and more than 85 percent within 5 percent — allegations only, with no defendant having admitted wrongdoing and no liability adjudicated.

Sources: usdepartmentofjusticeantitru2024, usdepartmentofjusticeofficeo2024, paul2025

Appears on: /domains/cases/realpage-rent-algorithm

EmpiricalThe consent decree the Department of Justice proposed on 24 November 2025 edits the deployment's structure rather than i…

The consent decree the Department of Justice proposed on 24 November 2025 edits the deployment's structure rather than its accuracy: competitor data used in models must be at least 12 months old and not drawn from active leases, the existing demand and supply models may not be trained with a geographic variable narrower than a state, the 'Governor' feature must treat increases and decreases symmetrically, automatic-acceptance ranges must be user-set and off by default, vendor-hosted meetings of competing landlords are barred, and a court-appointed monitor holds sweeping oversight for three years under a seven-year term with inspection rights, a written antitrust compliance program and a cooperation obligation — with no fine and no admission of liability; the Proposed Final Judgment and Competitive Impact Statement were published on 5 December 2025 (90 FR 56286), the stipulation was entered 26 March 2026, and the Department responded to eight public comments on 8 May 2026, with final public-interest entry by the court still pending as of July 2026 and no compliance findings published by the monitor.

Sources: usdepartmentofjusticeofficeo2025b, federalregister2026, hoganlovells2025, wilsonsonsini2025

Appears on: /domains/cases/realpage-rent-algorithm

EmpiricalThe White House Council of Economic Advisers estimated in December 2024 that rental pricing algorithms cost United State…

The White House Council of Economic Advisers estimated in December 2024 that rental pricing algorithms cost United States renters more than $3.8 billion in 2023, roughly $70 per month per unit in algorithm-priced buildings, with the software pricing at least 10 percent of all US rental units and nearly one in four multifamily rental units — model-based counterfactual estimates the council explicitly framed as a lower bound because non-participating landlords also raised rents in response, and never measured overcharge; separately, the Middle District of Tennessee preliminarily approved 26 private settlements involving 27 landlord defendants totalling $141.8 million on 21 November 2025 for a class of renters of covered properties between 18 October 2018 and 21 November 2025, over objections from five state attorneys general, with a second batch of roughly $218 million announced in 2026.

Sources: whitehousecouncilofeconomica2024, americanbarassociationantitr2025, multifamilydive2026

Appears on: /domains/cases/realpage-rent-algorithm

EmpiricalNew York City's Homebase homelessness-prevention program has since June 2012 routed applicant households on a 15-item Ri…

New York City's Homebase homelessness-prevention program has since June 2012 routed applicant households on a 15-item Risk Assessment Questionnaire scored 0-25 (each answer worth 1-3 points), distilled by backward elimination from a Cox proportional-hazards model of shelter entry fitted to 11,105 families who applied October 2004 - June 2008 (12.8% entered shelter within three years; decile risk 1% to 37%); a total at or above 7 points routed a household to 'full' services (financial assistance, case management, legal and mediation referrals) rather than a 'brief' one-or-two-visit contact. It is a regression-derived additive point screener, not a machine-learning system: the city's statutory algorithmic-tool register files it under computation type 'Scoring', purpose 'Resource allocation', autonomy 'Monitored', frequency 'Daily', with 'Vendor(s): None' in the CY2025 entry. Against caseworker judgment, which had deemed 66.5% of applicants eligible, the instrument would have increased correct targeting of families entering shelter by 26% and cut misses by almost two-thirds at equivalent false-alarm rates. Services are delivered by seven contracted nonprofit providers across 26 neighborhood offices, and The Department of Homeless Services states the network serves more than 25,000 at-risk households a year. The causal effect evidence for the program comes from outside the deploying agency's own research office: a randomized controlled trial by Abt Associates, commissioned by the Department of Homeless Services (2010-2013, 295 families with children analyzed across eleven sites), found the share spending at least one night in shelter falling from 14.5% to 8.0%, the share applying for shelter falling from 18.2% to 9.3%, and average shelter nights falling by 22.6; an independent community-district difference-in-differences study estimated roughly 5-11% fewer family shelter entries (11.2 log points, 95% CI 3.6-18.8), a $14.2M annual budget avoiding an estimated $20-44M of shelter expenditure. The separately claimed 'prevention rate' of around 90-97% is a city performance metric with no counterfactual and is not an effect size. The threshold of 7 describes the documented pre-2023 configuration; no public source states the threshold in force after the 2023 item revision.

Sources: shinnm2013, nycofficeoftechnologyandinno2026, rolstonh2013, goodmans2016, instituteforchildrenpovertya2024, newyorkcitydepartmentofhomel2025

Appears on: /domains/cases/nyc-shelter-entry-prediction

EmpiricalThe deploying agency's own Office of Research & Policy Innovation published a peer-reviewed re-examination (Housing Poli…

The deploying agency's own Office of Research & Policy Innovation published a peer-reviewed re-examination (Housing Policy Debate, 2022) of 48,450 deduplicated families with children applying 2013-2016 (58,674 family-years, over a period in which the caseload rose from about 600 cases a month in 2013 to more than 1,500 a month in some 2015 months), and documented two distortions against itself. First, score clustering at the eligibility cutoff: 4,269 cases scored 6, 10,634 scored exactly 7, and 7,756 scored 8, which its researchers wrote 'suggests that Homebase staff may be focused on getting families to that threshold so they qualify for full services' — while the agency separately asserts that the override valve reduces incentives for workers to misreport data to ensure eligibility. Second, override outcomes: 6.1% of the 58,674 applications departed from the score with mandatory supervisor approval (4.8% up to full services, 1.4% down to brief), and only 3.7% of the families moved up later applied to shelter against 25.8% of the families moved down, the paper concluding that worker judgment is less accurate than the RAQ on average; a ten-case note review attributed many downward decisions to needs beyond the program's capacity, such as families needing an apartment immediately with no funds, rather than to a judgment about risk. The same paper reported 73.9% of applications at or above the threshold, shelter application within two years at 13.7% above the cutoff against 5.9% below it (chi-square 699.98, p<.001), an area under the curve of 0.7387 for the deployed instrument against 0.7389 for the revised one, and a simulated alternative threshold of 5 raising precision from 13.7% to 15.2% at similar enrollment volume. These are agency-authored figures, published in a peer-reviewed venue; the paper's own caveats are that the observed score distribution is confounded by the clustering it documents and by low-scoring self-selection out of intake, and that outcomes compared by service tier are contaminated by the overrides being measured.

Sources: mullenej2022, newyorkcityofficeoftechnolog2024, farrelldc2023

Appears on: /domains/cases/nyc-shelter-entry-prediction

EmpiricalLocal Law 35 of 2022 (passed by the NYC Council 2021-12-15, lapsed into law unsigned 2022-01-14) added Admin. Code sec. …

Local Law 35 of 2022 (passed by the NYC Council 2021-12-15, lapsed into law unsigned 2022-01-14) added Admin. Code sec. 3-119.5, requiring every city agency to report by 31 December every algorithmic tool it used one or more times during the prior calendar year — expressly including tools that 'generate risk scores' or 'determine what resources are allocated to particular groups or individuals' — with six mandatory disclosure elements, compiled by the Office of Technology and Innovation into a public report delivered to the mayor and Council speaker each 31 March. DSS filed the Homebase RAQ in all four cycles to date (CY2022-CY2025), each entry disclosing in the agency's own words the June 2012 start date, the 2004-2008 training data analyzed with academic researchers, the input factors, the points-and-threshold eligibility mechanism and the worker-override-with-supervisor-permission rule. The loop demonstrably fired: after initial publication of the CY2023 report the register carries the note 'Update 3/27/2024 - The Department of Social Services updated their reporting to include changes made to the Homebase Risk Assessment Questionnaire after initial publication of the report', disclosing the 2023 item revision and adding its peer-reviewed citation; the Council's Committee on Technology held an oversight hearing on the regime on 2024-10-28, at which OTI described coordinating 45 agencies plus 24 further offices. Register entries are unaudited agency self-reports. Separately, three government audits have examined this program and none examined the risk model: the city comptroller's MG12-125A (2013-06-27) found no written monitoring policies, no records of initial ineligibility determinations and all provider risk assessments announced in advance, with the unannounced-visit recommendations rejected; the city comptroller's January 2020 audit of HRA's oversight of the then-$53M-a-year program found 80 of 240 required provider case-file reviews performed, 2,661 of 24,938 FY2018 households (11%) returning one to four times within twelve months, $2,271,797 in provider advances unrecouped some sixteen months after closeout, and 5 of 28 visited client homes not habitable (4 never fixed), issuing 19 recommendations; and the state comptroller's audit 2023-N-8 (issued 2026-01-07, scope July 2021 - July 2025) examined the downstream CityFHEPS rental-subsidy channel for DSS Homebase clients (57,888 new cases and 123,762 individuals housed since 2018 through March 2025; spending $176M in FY2019 rising to $834M in FY2024) and found units with hazardous violations approved, 30 of 75 sampled case records without evidence of income verification, and rents averaging $525 a month above comparables in eleven of thirty sampled statewide cases. None of the three may be cited as an audit of the algorithm.

Sources: newyorkcitycouncil2022, newyorkcityofficeoftechnolog2024, newyorkcitycouncilcommitteeo2024, officeofthenewyorkcitycomptr2013, officeofthenewyorkcitycomptr2020, officeofthenewyorkstatecompt2026

Appears on: /domains/cases/nyc-shelter-entry-prediction

EmpiricalCrimSAFE, a criminal-record tenant-screening product sold by CoreLogic Rental Property Solutions, computes no risk score…

CrimSAFE, a criminal-record tenant-screening product sold by CoreLogic Rental Property Solutions, computes no risk score: it is a deterministic record-matching and filtering engine that matches applicant identity data against a database of court and arrest records aggregated from more than 800 US jurisdictions, classified into three primary categories (Crimes Against Property, Crimes Against Persons, Crimes Against Society) with sub-classifications, and applies filter criteria the HOUSING PROVIDER configures: offense type, disposition and severity across felony and non-felony convictions and charges, and a lookback period configurable from 0 to 99 years for convictions and 0 to 7 years for charges (federal consumer-reporting law permits reporting non-conviction records for seven years). The output is a report carrying a lease decision driven by the provider's criteria plus a credit score, a Record(s) Found flag, message text the provider authors, full record detail for the users the provider authorizes, and an optional provider-customizable adverse-action letter template. Every new CrimSAFE user is by default authorized to receive full record data, with no cap on how many users get full access; a provider must affirmatively change configuration settings to restrict full reports to senior managers. In April 2016 Carmen Arroyo's application to move within ArtSpace in Windham, Connecticut, so her son Mikhail could live with her after a 2015 injury, was denied on 26 April 2016 after the screen returned Record(s) Found; the only matched record was a pending Pennsylvania shoplifting charge, later withdrawn in April 2017. Mikhail's report carried the vendor's default message: 'Please verify the applicability of these records to your applicant and proceed with your community's screening policies.'

Sources: connecticutfairhousingcenter2026, connecticutfairhousingcenter2023, nationalhousinglawproject2018

Appears on: /domains/cases/corelogic-crimsafe

EmpiricalThe location of decision authority in this deployment was the contested question, and the Second Circuit resolved it com…

The location of decision authority in this deployment was the contested question, and the Second Circuit resolved it component by component on 20 February 2026 (Cabranes, Wesley, Menashi; opinion by Menashi; Nos. 23-1118(L), 23-1166(XAP)). WinnResidential had suppressed full reports from its own on-site staff so that leasing decisions involving criminal records would be made 'by someone in a more elevated position,' out of concern about leasing-commission incentives, so the on-site agent saw only a Record(s) Found flag and told Arroyo the application was denied without individualized review; answering the state commission's complaint, WinnResidential said it did not know 'the facts behind the criminal background findings' because it had 'trust' in CoreLogic's reports. After a ten-day bench trial Judge Vanessa L. Bryant ruled on 20 July 2023 that CrimSAFE does not disqualify applicants because the housing provider decides what records matter and whether to deny; the earlier August 2020 summary-judgment characterizations that the companies 'acted hand-in-glove' and that CoreLogic 'was an integral participant' belong to that posture and did not survive trial. The Second Circuit rejected the threshold reasoning that screening companies sit categorically outside the Fair Housing Act — a point the United States had urged as amicus on 24 November 2023 — but affirmed on proximate cause, holding the denial came after a chain of the provider's discretionary decisions (configuration, record relevance, staff access, adverse-action letters, final approval) and quoting the district court's own sentence with approval: 'No housing provider who uses CrimSAFE could reasonably believe that CoreLogic makes housing decisions for them.' It rejected the 'cat's paw' theory because the screening policies applied were the provider's own, rejected liability for failing to restrict lawfully reportable non-conviction records as extending liability beyond the first step, and dismissed the Connecticut Fair Housing Center's own claim for lack of Article III standing under a 2024 organizational-injury doctrine unrelated to tenant screening, vacating rather than deciding its merits. Disparate impact was never proven: the race and national-origin claim failed at the prima facie causation step, so the underlying statistics were never adjudicated.

Sources: connecticutfairhousingcenter2026, connecticutfairhousingcenter2023, connecticutfairhousingcenter2020, unitedstatesdepartmentofjust2023b

Appears on: /domains/cases/corelogic-crimsafe

EmpiricalThe correction loop in this deployment ran through records the subject could not see at the point of harm. CoreLogic's A…

The correction loop in this deployment ran through records the subject could not see at the point of harm. CoreLogic's Authentication Procedure Guide listed only a notarized power of attorney as third-party authorization and escalated 'any scenarios not covered' to a supervisor; staff demanded a power of attorney that Mikhail Arroyo, a conservatee, was legally incapable of executing — a demand the trial court called an 'impossible condition' — across a blocked window running from the 24 June 2016 request with a conservatorship certificate to mid-November 2016, when a 1 November call escalated to CoreLogic's legal department and the company agreed about two weeks later that a conservatorship certificate with a visible probate seal would suffice; the resubmitted copy again lacked a visible seal and the disclosure was never completed. That escalation is the only documented change in vendor behavior in the record. The family learned which record had caused the denial in December 2016, roughly eight months after the denial, from the housing provider rather than the vendor, and the adverse-action letter that should have triggered the correction loop was sent but never received; correction ultimately happened at the original source when Arroyo petitioned the Pennsylvania court and the charge was withdrawn in April 2017. Arroyo's was the first and only conservator file-disclosure request CoreLogic had ever received (478 F. Supp. 3d at 282), a finding both courts used to defeat the disability disparate-impact claim. The Connecticut Commission on Human Rights and Opportunities held an evidentiary hearing on 13 June 2017 and WinnResidential approved the move-in ten days later — about fourteen months after the denial, and the only oversight action documented to have changed an outcome. The sole liability finding, a willful consumer-reporting violation carrying $1,000 statutory and $3,000 punitive damages awarded in July 2023, was reversed on 20 February 2026; final vendor liability was zero, and no injunction, consent decree or policy change was ever ordered.

Sources: connecticutfairhousingcenter2026, connecticutfairhousingcenter2020, connecticutfairhousingcenter2023, courtlistenerandthefreelawpr2023

Appears on: /domains/cases/corelogic-crimsafe

EmpiricalNew York City launched NYC Teenspace on November 15, 2023 as a three-year, $26M contract buying population access to Tal…

New York City launched NYC Teenspace on November 15, 2023 as a three-year, $26M contract buying population access to Talkspace's existing consumer teletherapy platform for all city residents aged 13-17, with unlimited asynchronous messaging (therapist replies five days per week) plus one 30-minute live session per month from NY-licensed clinicians; inside the therapy chat a proprietary NLP model scans teen-authored messages for suicide and self-harm language and fires an urgent real-time alert to the teen's own treating therapist, which never acts autonomously — escalation (child-protective-services referral, intensive therapy, hospitalization) is entirely the clinician's; in the first six months the scan flagged roughly 50 teens as moderate-to-high suicide risk (under 1% of about 6,800 users) while therapists navigated 36 high-risk events including suicide attempts, a child-abuse report and a drug overdose; enrollment ran from about 6,800 (May 2024) to 19,000+ (DOHMH letter, Dec 1 2024) to a vendor-reported 45,000+ (Feb 2026), with about 80% of early registrants Black, Hispanic, AAPI, bi-racial or Native American, about 70% female, and more than half resident in the city's priority TRIE neighborhoods.

Sources: nycofficeofthemayor2023, chalkbeatnewyork2024, newyorkcitydepartmentofhealt2024, newyorkcitydepartmentofhealt2026

Appears on: /domains/cases/nyc-teenspace-talkspace

EmpiricalA September 2024 advocacy tracker audit of the NYC Teenspace pages counted 15 ad trackers and 34 cookies sharing teen vi…

A September 2024 advocacy tracker audit of the NYC Teenspace pages counted 15 ad trackers and 34 cookies sharing teen visitors' personally identifiable information with recipients including Facebook/Meta, Amazon, Google and Microsoft, alongside sensitive intake data (name, date of birth, address, school, gender, mental-health screening answers) collected before parental consent was secured; DOHMH's December 18, 2024 letter (General Counsel Landau, Chief Privacy Officer Elcock) recorded that Talkspace removed all social-media and advertising trackers as of December 11, 2024, that the sign-up flow was minimized to age and zip code only, and that the contract's Data Security Rider would be amended to ban marketing use and trackers outright; advocates then documented continuing tracker-based disclosure on pages beyond the landing page (including the page hosting the revised privacy policy) on January 9 and February 12, 2025 with no DOHMH response to their follow-up letter after more than a month, and February 2025 technical testing found the NYC landing page clean as of January 24 while Talkspace's Seattle and Baltimore teen pages still transmitted visitor IP addresses to TikTok, Meta, Snapchat, Google, X, Reddit, LinkedIn, Spotify and Quora until the reporter's inquiry — several recipients being companies NYC had sued in February 2024 over teen mental-health harm. Talkspace's Chief Privacy Officer stated no personal medical information was transmitted; advocates dispute the completeness of the remediation, so the leakage is identifier-level and contested-scope.

Sources: parentcoalitionforstudentpri2024, newyorkcitydepartmentofhealt2024, parentcoalitionforstudentpri2025, gizmodo2025

Appears on: /domains/cases/nyc-teenspace-talkspace

EmpiricalThe only published performance figure for the Teenspace suicide-alert algorithm is Talkspace's own claim of 83% accuracy…

The only published performance figure for the Teenspace suicide-alert algorithm is Talkspace's own claim of 83% accuracy versus a human expert, resting on a 2020 Psychotherapy Research study of ADULT platform data rather than any measurement of the teen population it reads, and the independent evaluation DOHMH said in a September 12, 2024 email that it was planning had still not been published as of mid-2026, with no OIG, comptroller or FTC action specific to Teenspace located; outcome figures (65% 'reported improvement' at six months; 66% 'measurable clinical improvement' among 45,000+ enrollees) are vendor or city self-reports with undisclosed instruments and no independent verification; meanwhile the record store the program writes into is described by Talkspace's CEO as 8 billion words, 140 million messages and 6.2 million assessments, is subpoenable (a Talkspace user's complete therapy history was obtained by her former employer and used in court in 2026 — an adult employer-benefit member, not a Teenspace teen), is stated by the vendor to be used for training behavioral-health LLMs, and transfers intact under the pending $835M Universal Health Services acquisition announced March 9, 2026 and expected to close in Q3 2026, while the NYC contract is still live.

Sources: talkspace2023, kdive2024, proofnews2026, prnewswireanduniversalhealth2026

Appears on: /domains/cases/nyc-teenspace-talkspace

EmpiricalCharacter.AI built its crisis-safety stack in four dated increments, each within days to weeks of a specific external ev…

Character.AI built its crisis-safety stack in four dated increments, each within days to weeks of a specific external event: a self-harm phrase screen referring users to the 988 Lifeline (vendor blog dated 22 October 2024, the day the Garcia wrongful-death complaint was filed in M.D. Fla., publicized the 23rd), a separate more restrictive under-18 model (December 2024, after two Texas family suits and a 15-company Texas Attorney General SCOPE Act investigation), a weekly parental usage summary carrying time spent and top characters while deliberately excluding chat content (25 March 2025), and — one week after an investigation found dozens of harmful personas live and seven weeks after FTC 6(b) orders reached seven companies — the removal of open-ended chat for under-18 users, announced 29 October 2025 and effective 25 November, with an interim two-hour daily cap ramping down and a two-stage age-assurance stack; on 7 January 2026 the company, both founders, and Google agreed to settle the Garcia case and four others in New York, Colorado, and Texas, terms undisclosed and approval still required, so causation was never adjudicated. Each retrofit is a company authority action temporally associated with an external event; the operator's own announcement cited several pressures at once, and no sole cause is asserted.

Sources: characterai2024, characterai2025a, cnnbusiness2026

Appears on: /domains/cases/character-ai-crisis-response

EmpiricalThe only direct measurement of Character.AI's self-harm detector is a single adversarial test published 29 October 2024:…

The only direct measurement of Character.AI's self-harm detector is a single adversarial test published 29 October 2024: across sixteen conversations with mental-distress-focused bots the 988-hotline pop-up appeared three times, triggered by two exact phrasings ('I am going to commit suicide' and 'I will kill myself right now') while many equally explicit statements including 'I want to end my life' did not trigger it, dismissable and non-blocking so the chat continued, with firing frequency increasing after the publication contacted the company — a sample of sixteen, and therefore anecdotal rather than a measured rate. No independent evaluation of the sensitivity or specificity of the crisis classifier, the two-stage age-assurance classifier, or the separate under-18 model exists in either direction; every efficacy claim about the retrofits is a vendor claim. The first compelled quantitative record of this detector class is prospective: California SB 243, signed 13 October 2025 and operative 1 January 2026, requires a crisis protocol issuing referral notifications and, from 1 July 2027, an annual report to the state Office of Suicide Prevention of the number of notifications issued, posted publicly — a count of referrals, not a measure of what they achieved.

Sources: futurism2024, californialegislativeinforma2025

Appears on: /domains/cases/character-ai-crisis-response

EmpiricalCharacter.AI's crisis pathway terminates in an automated referral to an external hotline with no staff review, no escala…

Character.AI's crisis pathway terminates in an automated referral to an external hotline with no staff review, no escalation call, and no capture of what followed; the human roles the record does document sit elsewhere — trust-and-safety moderators taking down user-authored characters reactively, parents receiving a weekly usage summary with no transcript access and only where the teen adds them, and, after November 2025, third-party identity reviewers adjudicating contested age determinations. The external review tier is by contrast unusually crowded and repeatedly behavior-forcing: two federal district courts, the Texas Attorney General twice, FTC 6(b) compulsory orders to seven companies on 11 September 2025, California SB 243, and a Senate subcommittee hearing, alongside investigative journalism that on 22 October 2025 — a year into the litigation — found dozens of harmful bots live including a 'Bestie Epstein' persona that had logged almost 3,000 chats, a gang simulator, school-shooter personas, and a 'doctor' giving antidepressant-tapering instructions. None of these reviewers ever operated inside the decision loop; what they moved was the structure itself.

Sources: thebureauofinvestigativejour2025, ftc2025b, characterai2025

Appears on: /domains/cases/character-ai-crisis-response

EmpiricalFour days after a labor board certified its paid helpline staff's union election, a national eating-disorder nonprofit t…

Four days after a labor board certified its paid helpline staff's union election, a national eating-disorder nonprofit told those staff they were terminated and that a chatbot would replace a helpline that had fielded nearly 70,000 contacts in 2022 with six paid staff, about two supervisors, and up to roughly 200 trained volunteers; the chatbot was suspended on May 30, 2023, two days before it was to become the sole channel, after testers published screenshots of it recommending a 500 to 1,000 calorie daily deficit, a 1 to 2 pound weekly loss, a 2,000 calorie cap, regular weigh-ins, and where to buy skinfold calipers, and the helpline closed on June 1 as scheduled while the chatbot was already offline.

Sources: wells2023, kffhealthnews2023, nprshots2023, picchi2023, npr2023a

Appears on: /domains/cases/neda-tessa-chatbot

EmpiricalThe deployed chatbot had two layers with different evidence status: a closed, pre-scripted program the vendor and the re…

The deployed chatbot had two layers with different evidence status: a closed, pre-scripted program the vendor and the research team describe as unable to depart from its authored content, and a generative question-and-answer feature the operating vendor added in what its chief executive called a systems upgrade covered by the client's contract, a reading the client's chief executive denies by saying the organization was never advised of the changes and would not have approved them; the vendor's public account of the harmful outputs was mixed, saying it was still trying to determine how a closed system allowed such content, so the attribution of the May 2023 advice to the generative layer is the reported and attributed explanation rather than an established mechanism, and the contract scope stays contested with no adjudicated breach.

Sources: nprshots2023, kffhealthnews2023, thewraprepublishedonyahoo2023

Appears on: /domains/cases/neda-tessa-chatbot

EmpiricalThe earliest documented external warning about harmful chatbot responses came in October 2022 from the executive directo…

The earliest documented external warning about harmful chatbot responses came in October 2022 from the executive director of a peer eating-disorder organization, and the specific language she flagged was quickly removed after she reported it, with no documented systemic review, output monitoring, or reassessment of the deployment following; the vendor's chief executive said the flagged language was part of the pre-scripted content rather than the generative layer, which the research team denies, so its provenance is contested, and system-level action arrived roughly seven months later when public screenshots circulated.

Sources: nprshots2023

Appears on: /domains/cases/neda-tessa-chatbot

EmpiricalThe Veterans Health Administration built an opioid overdose and suicide risk model in-house, with no commercial vendor, …

The Veterans Health Administration built an opioid overdose and suicide risk model in-house, with no commercial vendor, on its own electronic health record and Corporate Data Warehouse: fitted on 1,135,601 patients with an opioid prescription in fiscal 2010 against 23,790 overdose-related or suicide-related events among them in fiscal 2011 (a 2.1% base rate), reporting an area under the curve above 0.80 in training and test sets, and refreshed nightly as a continuous one-year risk estimate binned into percentile tiers on a population-management dashboard that shows each patient's risk factors and the guideline-recommended mitigation actions rather than a bare number. Its predictors are administrative - demographics, pharmacy records including opioid type and dose and co-prescribed sedatives, mental health and substance use disorder diagnoses, prior overdose-related and suicide-related events, detoxification episodes, and emergency department and other utilization history - so no structured risk questionnaire and no clinician-scored instrument feeds the score. VHA Notice 2018-08, issued 8 March 2018 and effective 18 April 2018, required all 140 VHA medical centers to convene interdisciplinary teams and case-review every patient in the very-high-risk tier, initially the top 1% of scores at a threshold of 0.166. The mandate compels an assessment of risk, prescription appropriateness and mitigation options and never an action: there is no mandated taper and no automatic prescription cutoff tied to the score, and every clinical decision remains with the treatment team. Across 44,042 patients in the top 1% to 5% band from April 2018 to March 2020 the mandate raised the odds of receiving a case review 5.1-fold (95% CI 3.64-7.23) and added 0.498 risk-mitigation strategies per patient (95% CI 0.39-0.61). Facility completion had a median of 71% (interquartile range 48-95%), about one facility in five met the 97% target, and each of the 89 surveyed facilities used a median of 23 distinct implementation strategies. Team composition and workflow were left to each facility.

Sources: oliva2017, strombotne2023, minegishi2022, rogal2020

Appears on: /domains/cases/va-storm-opioid-risk

EmpiricalThe Veterans Health Administration randomized two features of its own opioid case-review policy across its 140 medical c…

The Veterans Health Administration randomized two features of its own opioid case-review policy across its 140 medical centers, and the two randomizations are separate experiments whose findings do not combine. The first was timing: when a facility's mandated review tier widened from the top 1% of risk scores to the top 5%, executed as a stepped wedge with waves on 12 February 2019 and 13 August 2019, which the primary trial paper places at study months 11 and 17. The second was language: whether a facility's copy of the policy notice carried an accountability paragraph naming a 97% case-review completion target, with quarterly reporting to the national Office of Mental Health and Suicide Prevention and technical assistance and action plans for facilities below it. Seventy facilities received that paragraph and seventy did not. Across 16,272 very-high-risk patients (8,734 in the accountability arm, 7,538 outside it), about 57% received a case review overall against a pre-mandate baseline of 6.6%, and there was no difference between arms in opioid-related serious adverse events (hazard ratio 1.03, 95% CI 0.97-1.08) or mortality (hazard ratio 1.00, 95% CI 0.91-1.09) - but patients at accountability-arm facilities were less likely to receive a case review at all (hazard ratio 0.91, 95% CI 0.87-0.95). The implementation tracking study reports the same result at facility level: median completion 71% (interquartile range 48-95%), 18 of 89 surveyed facilities (20%) meeting the 97% target, and facilities given the plain mandate meeting it more often than facilities given the accountability language, 30% against 11% (p=0.04). The practice associated with higher completion was regular self-monitoring and adaptation inside the facility (adjusted incidence rate ratio 1.40) rather than reporting upward from it; dashboard use was reported by 97% of surveyed facilities and local opinion leaders by 80%, while patient-engagement strategies were used by 13%. The oversight-backfire finding belongs to the language experiment alone.

Sources: minegishi2022, strombotne2023, rogal2020

Appears on: /domains/cases/va-storm-opioid-risk

EmpiricalThe benefit and the harm signals from the Veterans Health Administration's mandated opioid case-review policy travel tog…

The benefit and the harm signals from the Veterans Health Administration's mandated opioid case-review policy travel together and neither may be reported alone. The widely quoted mortality result - four-month all-cause mortality odds of 0.78 (95% CI 0.65-0.94) - is an EXPLORATORY endpoint of a trial whose pre-specified primary composite of nine serious-adverse-event categories did not move (odds ratio 0.995, 95% CI 0.875-1.132). Among patients NEWLY DIAGNOSED with opioid use disorder during the trial (28,251 analyzed, estimated off the stepped-wedge threshold-expansion randomization rather than the accountability-language arm) the mandate was associated with 90-day all-cause mortality odds of 1.74 (95% CI 1.06-2.87) with no significant change in serious adverse events; a post-hoc subgroup with an opioid prescription before but not after diagnosis showed 5.87 (95% CI 1.85-18.58). Nearly every quantitative source on this deployment is a department-affiliated research-operations partnership: this is peer-reviewed agency self-evaluation that published its own null, backfire and harm findings, and it is not third-party replication.

Sources: strombotne2023, auty2023

Appears on: /domains/cases/va-storm-opioid-risk

EmpiricalInternal records subpoenaed by the U.S. Senate Permanent Subcommittee on Investigations show that early-2021 testing of …

Internal records subpoenaed by the U.S. Senate Permanent Subcommittee on Investigations show that early-2021 testing of an auto-authorization model inside UnitedHealthcare produced faster handle times together with an increase in adverse determination rate - attributed to finding contraindicated evidence missed in original review - and the internal committee voted to tentatively approve the model at the following meeting; the April 2021 approval of 'Machine Assisted Prior Authorization' was paired with testing that removed six to ten minutes from the average review while the reviewing doctor or nurse still had to verify that the primary evidence is acceptable. Over the same period the insurer's post-acute prior authorization denial rate went from 10.9 percent (2020) to 16.3 percent (2021) to 22.7 percent (2022), and its 2019 skilled-nursing-facility denial rate was nine times lower than its 2022 rate. A January 2022 vendor presentation shows a naviHealth care coordinator completing nH Predict to determine optimal post-acute placement while the patient is still hospitalized, and an April 2022 vendor instruction told call handlers not to guide providers on the questions used to collect the information determinations are made from.

Sources: ussenatepermanentsubcommitte2024, caseyrossandbobherman2023

Appears on: /domains/cases/nh-predict-utilization-review

EmpiricalAcross Medicare Advantage in 2022, of 46.2 million prior authorization determinations, 3.4 million (7.4 percent) were de…

Across Medicare Advantage in 2022, of 46.2 million prior authorization determinations, 3.4 million (7.4 percent) were denied in whole or in part; 9.9 percent of denials were appealed; and 83.2 percent of appeals resulted in the initial decision being overturned (KFF analysis of the plans' own federal reporting). The figures are program-wide, not plan- or service-line-specific, and the overturn rate is conditioned on the self-selected minority of denials that were appealed, so it overstates what the same review would correct if applied to every denial.

Sources: kff2022

Appears on: /domains/cases/nh-predict-utilization-review

EmpiricalWhether nH Predict was used to make coverage determinations is contested and unadjudicated. The vendor's public statemen…

Whether nH Predict was used to make coverage determinations is contested and unadjudicated. The vendor's public statement in STAT's March 2023 series opener was that the tool 'is not used to make coverage determinations' and 'is used as a guide'; the Lokken order (D. Minn., Feb 13, 2025) recites - as pleaded allegations taken as true on a motion to dismiss - an allegation about the share of claim denials reversed on appeal, and records in the same paragraph that 'UHC denies any use of nH Predict'; the Barrows order (W.D. Ky., Aug 14, 2025) recites Humana's use of the tool as pleadings only. Both class actions survived motions to dismiss in narrowed form - contract and implied-covenant counts in both, plus unjust enrichment and common-law fraud in the Humana action and are in active discovery as of August 2026, with no dispositive merits ruling. CMS's contract-year-2024 rule and its February 6, 2024 FAQ constrain the order of operations: an algorithm may assist a coverage determination, the plan remains responsible, an algorithm determining coverage from a larger data set instead of the individual patient's medical history, physician recommendations or clinical notes would not comply with 42 CFR 422.101(c), and a predicted length of stay alone cannot be the basis to terminate post-acute care services - only re-assessing the individual patient's condition can. The naviHealth brand was retired in Q1 2024 into Optum 'Home & Community Care'; the pipeline continues under that name.

Sources: estateofgeneb2025, centersformedicaremedicaidse2024, cfr2024, cbsnews, caseyrossandbobherman2023a, barrows2025, bobhermanandcaseyross2023, georgetownlaw2026a, georgetownlaw2026

Appears on: /domains/cases/nh-predict-utilization-review

EmpiricalThe closest ground-truth audit of this decision class is HHS OIG evaluation OEI-09-18-00260 (April 2022): reviewing a st…

The closest ground-truth audit of this decision class is HHS OIG evaluation OEI-09-18-00260 (April 2022): reviewing a stratified random sample of 250 prior authorization denials and 250 payment denials issued by 15 of the largest Medicare Advantage organizations during one week of June 2019, health-care coding experts and physician reviewers found 13 percent of the prior authorization denials and 18 percent of the payment denials met Medicare coverage rules, with identified causes including internal clinical criteria applied beyond Medicare rules, insufficient-documentation findings the reviewers judged unfounded, manual processing errors and system programming failures, and with stays in post-acute facilities among the report's own examples. The evaluation predates naviHealth's management of this benefit and pools fifteen organizations, so it orders - and does not measure - an error term for this deployment.

Sources: u2022

Appears on: /domains/cases/nh-predict-utilization-review

EmpiricalA commercial population-health risk score used to target a scarce high-risk care management program at a large academic …

A commercial population-health risk score used to target a scarce high-risk care management program at a large academic hospital (studied 2013-2015) was trained to predict total medical expenditure in the following year while the deployment's stated purpose was to identify the patients with the greatest health needs, with race excluded from its features. The independent evaluation found the model well calibrated for cost across race — roughly equal realized following-year costs at every level of predicted risk, 5,147 versus 4,995 dollars at the median score — while at the 97th-percentile auto-identification threshold Black patients carried 26.3 percent more active chronic conditions than White patients at the same score (4.8 versus 3.8, P < 0.001). Patients above the 97th percentile were automatically identified for enrollment and those above the 55th were referred to their primary care physician; enrollment covered 1.3 percent of observations, and commercial tools of this class are applied to roughly 200 million people in the US each year per industry estimates cited in the study.

Sources: obermeyer2019a, obermeyer2019

Appears on: /domains/cases/cost-proxy-care-stratification

EmpiricalThe deployment's designed human check — physician referral above the 55th percentile — was measured: realized program en…

The deployment's designed human check — physician referral above the 55th percentile — was measured: realized program enrollment was 19.2 percent Black against 11.9 percent Black in the whole sample, while simulated race-blind sampling within bins of the same score would have produced 18.3 percent (P = 0.8348 against the observed figure) and sampling on predicted chronic-condition count within bins would have produced 26.9 percent. The study's abstract states that remedying the disparity would increase the percentage of Black patients receiving additional help from 17.7 to 46.5 percent — a figure describing the composition of the helped cohort under a counterfactual ranking, not a rate of selection for any subpopulation.

Sources: obermeyer2019a, obermeyer2019

Appears on: /domains/cases/cost-proxy-care-stratification

EmpiricalPer the Science study, the manufacturer independently replicated the analysis on its national dataset of 3,695,943 comme…

Per the Science study, the manufacturer independently replicated the analysis on its national dataset of 3,695,943 commercially insured patients, confirmed the finding, and jointly rebuilt the predictor with the study's authors using the same sample, the same features with race still excluded, and the same training process, changing only the training label to an index combining health prediction with cost prediction. One measure of predictive bias — excess active chronic conditions in Black patients conditional on risk score — fell from 48,772 to 7,758, an 84 percent reduction in that bias measure (counts of excess conditions, not error rates). The relabelled predictor was experimental and holdout at publication, with no production role in enrollment decisions, and no independent evaluation of it by anyone other than its authors has been located. The study's authors later generalized the remediation practice in the Algorithmic Bias Playbook (Chicago Booth Center for Applied AI, 2021), which centers label-choice bias and, like the study, does not name the product.

Sources: obermeyer2019a, obermeyer2021

Appears on: /domains/cases/cost-proxy-care-stratification

EmpiricalThe Science study anonymized the product and manufacturer. On its publication day, October 25, 2019, the Superintendent …

The Science study anonymized the product and manufacturer. On its publication day, October 25, 2019, the Superintendent of the New York State Department of Financial Services and the Commissioner of the New York State Department of Health jointly wrote to UnitedHealth Group's chief executive, naming 'Optum's data analytics program, Impact Pro', footnoting the study, and demanding that the company immediately investigate and demonstrate the algorithm is not racially discriminatory or cease using Impact Pro or any similar program, asserting insurers' responsibility under New York law regardless of who built the algorithm; contemporaneous press reports carried the same identification and the company's statement that the cost model 'was highly predictive of cost, which is what it was designed to do.' No publicly documented resolution of the New York inquiry has been located as of August 2026. Separately, 45 CFR 92.210 (adopted in the 2024 HHS section 1557 final rule) prohibits discrimination through the use of patient care decision support tools and imposes ongoing duties to identify tools employing inputs measuring race, color, national origin, sex, age, or disability and to mitigate the resulting risk — a standing duty that postdates the studied deployment and is carried for the duty it imposes, not as an adjudication of this case.

Sources: newyorkstatedepartmentoffina2019, cfr, snowbeck2019, ledford2019

Appears on: /domains/cases/cost-proxy-care-stratification

EmpiricalThe Child Abuse Clinical Decision Support system (CA-CDS), built by a consortium at UPMC Children's Hospital of Pittsbur…

The Child Abuse Clinical Decision Support system (CA-CDS), built by a consortium at UPMC Children's Hospital of Pittsburgh, combines a five-item nurse-administered screen for every child under 13, a free-text scan of the chief complaint and nursing assessment, physician orders and discharge diagnoses into a rule-and-text trigger — no learned risk score — that raises a dashboard icon and a pop-up recommending a guideline-aligned physical-abuse order set. Implementation measurably moved volume: identification roughly quadrupled at two general emergency departments (P < .001), and triggering roughly doubled at both disseminated health systems (2.4% to 3.5% of children at the University of Wisconsin on Epic; 1.1% to 1.9% at Northwell Health on Allscripts). Its published performance — sensitivity 96.8%, specificity 98.5%, positive predictive value 26.5% — was measured against the hospital child protection team's chart assessment as reference standard: a concordance with an expert record judgment, not detection of abuse in any child. The documented failure mode is a string match: 'burn' in a 3-month-old's chief complaint is a trigger and so is 'burning up with fever'; 'broke' in a young infant triggers on 'the father speaks broken English'; 70% (33 of 47) of audited overtriggers came from the free-text scan, and 81% (195 of 242) of triggers were judged appropriate on chart review.

Sources: berger2019, feldstein2023, rosenthal2019

Appears on: /domains/cases/pediatric-abuse-detection-cds

EmpiricalThe CA-CDS deployment's working end product is a mandated report that crosses a one-way organisational boundary. An indi…

The CA-CDS deployment's working end product is a mandated report that crosses a one-way organisational boundary. An individual clinician, discharging a personal statutory duty under each state's mandatory-reporting laws pursuant to the federal Child Abuse Prevention and Treatment Act, transmits the report to the state child protection agency: health-care workers are designated reporters in 46 states plus DC and five territories, no institutional policy relieves the individual duty in 17 states plus DC and the Virgin Islands, and the standard is suspicion with no burden of proof. Implementation raised reports to child protective services from 0.6% to 0.9% of children at one health system (P = .03), with 17% of triggered children reported against 0.26% of non-triggered. The receiving agency must screen every referral; in federal fiscal year 2023 states screened in 2,107,473 referrals against an estimated 2,292,000 screened out, on 5,936 intake workers and 21,739 investigation and alternative-response workers completing 66 responses each per year at a mean first-contact time of 102 hours, with 28 states reporting increases and citing staff shortages and turnover. The statutory scheme establishes no duty to return the screening decision or disposition to the reporting clinician's record, and no source documents a record-level write-back onto the hospital chart; the report joins a durable agency record whose prior-involvement and central-registry checks are standard elements of the investigation that follows the next report about the same family.

Sources: childwelfareinformationgatew2023, childwelfareinformationgatew, u2023, feldstein2023

Appears on: /domains/cases/pediatric-abuse-detection-cds

EmpiricalThe CA-CDS record measures both its check and the gap over it. The physical-abuse order set produced 100% full guideline…

The CA-CDS record measures both its check and the gap over it. The physical-abuse order set produced 100% full guideline compliance when used at the originating hospital (43 uses in a seven-month trial; partial compliance fell from 10% to 3%, P = .04) and 96% (22 of 23 uses) at Northwell Health — while 58% of surveyed practitioners reported not using it at all. Of 71 practitioners analysed across 19 UPMC general emergency departments in February 2020, 69% did not recognise the dashboard icon indicating a trigger, 54% did not know they could view the screen result, 27% could not recall seeing an alert, and 65% were uncertain which tests to order; only 4.5% (3 of 66) disagreed with the recommendations, 75% said the tool raised awareness, 72% discussed alerts face-to-face with the child's nurse, and 54% named lack of social work or ancillary support as a barrier. Full compliance with the guideline evaluation at the moment of reporting ran 80% to 75% at Wisconsin and 33% to 50% at Northwell, and site acceptance diverged sharply on nominally identical software: 81% of Wisconsin respondents wanted continued use against 3% at Northwell.

Sources: suresh2018, feldstein2023, peterson2024

Appears on: /domains/cases/pediatric-abuse-detection-cds

EmpiricalThe only deployment-level equity measurement for the CA-CDS is a null: screening rates did not differ by patient or hosp…

The only deployment-level equity measurement for the CA-CDS is a null: screening rates did not differ by patient or hospital characteristics across 13 UPMC general emergency departments. No source publishes trigger, report or compliance rates disaggregated by subpopulation for this deployment. The domain's best-known disparity measurement — 22.5% of white versus 52.9% of minority children reported for suspected abuse among 388 children under 3 with acute fractures, and skeletal surveys ordered at 8.75 adjusted odds for minority toddlers — comes from a different institution (an urban academic children's hospital in Philadelphia) and era (1994-2000), and is the measured clinician variability a universal, instrument-driven screen is meant to standardise: context for this case, never a measurement of this system.

Sources: berger2019, lane2002

Appears on: /domains/cases/pediatric-abuse-detection-cds

EmpiricalStopNCII.org (operated by the Revenge Porn Helpline within the UK charity SWGfL, developed with Meta, launched December …

StopNCII.org (operated by the Revenge Porn Helpline within the UK charity SWGfL, developed with Meta, launched December 2021) and Take It Down (operated by the US National Center for Missing & Exploited Children, launched February 2023) run one on-device hash-removal mechanism in two configurations: the person who holds the material generates a hash on their own device, only the hash leaves the device, the original is never uploaded or stored, third-party submission is refused and eligibility is self-attested, and participating platforms match the hash against uploads on public or unencrypted surfaces, reviewing any match under their own policies. StopNCII states its algorithms as PDQ and PhotoDNA for photos and MD5 for videos; independent researchers verified by inspecting the Take It Down web client that it runs PDQ. The operators' own volume reports — advocacy-tier counts of reporting behaviour, never incidence — were 2 million images protected across more than 785,000 cases by November 25, 2025 on the StopNCII side (a reported 97 percent increase over 2024, 17 industry partners) and 130,000+ submissions covering 273,000+ images and videos in 2025 on the Take It Down side, up from 83,000+ submissions in 2024. As fetched 2026-08-27 the two partner rosters overlap and diverge (YouTube on the minor index only, X on the adult index only, Discord on neither), encrypted surfaces are outside both by design, and Google announced on September 17, 2025 that it will use StopNCII hashes in Search — an announcement scoped to Search results, with Google absent from the StopNCII partners page as fetched the same day.

Sources: stopncii, nationalcenterformissingexpl, nationalcenterformissingexpl2025, hawkes2024, swgfl2025, google2025a

Appears on: /domains/cases/stopncii-hash-removal

EmpiricalA peer-reviewed independent evaluation (IEEE Security & Privacy Magazine 2024) reconstructed recognizable pre-images — h…

A peer-reviewed independent evaluation (IEEE Security & Privacy Magazine 2024) reconstructed recognizable pre-images — hair colour and length, face shape, other facial features, some background — from PDQ, PhotoDNA, NeuralHash and aHash hashes using an off-the-shelf conditional image-to-image generative network trained on 1,000 public celebrity-face images on 2015-era consumer hardware, with mean perceptual similarity of 60.10 percent for PDQ and 74.04 percent for PhotoDNA, measured on a public celebrity-face benchmark and never on any reporter's material; the authors concluded the hashes should be treated as sensitive in the same way as the original images. The same team quoted Take It Down's FAQ answer that material 'cannot be reverse engineered or created from the hash values shared with NCMEC', reported that answer to be wrong, wrote to the operator twice between August and December 2023, received no reply, and recorded that the website did not change; the sentence was still on the FAQ page on 2026-08-27. A separate USENIX Security 2023 evaluation demonstrated efficient targeted second-preimage and detection-avoidance attacks against PhotoDNA and PDQ, concluding existing perceptual hash functions are likely insufficiently robust for adversarial settings. No operator in this class publishes an audited error rate for its deployed system.

Sources: hawkes2024, prokos2023, nationalcenterformissingexpl

Appears on: /domains/cases/stopncii-hash-removal

EmpiricalThe two configurations diverge on the victim channel by design, and both operators document the propagation surface a wi…

The two configurations diverge on the victim channel by design, and both operators document the propagation surface a withdrawal does not reach. StopNCII gives a reporter a case number, PIN and status page (updates can take 3 to 5 days) and services whole-case withdrawal against credentials the operator states it does not store and cannot reset; Take It Down is anonymous by construction, with no status channel, no notification of matches, and no withdrawal mechanism at all. StopNCII states that hashes persist after the reporter deletes the image and are shared with new partners as they join, and its FAQ states that participating companies 'reserve the right to continue enforcing their policies once they've acquired knowledge of the hash' — so an entry's reach grows after submission and a withdrawal retracts the bank's copy but not partner-held copies. A hash match obliges nobody: matching platforms review the content against their own policies before any action, platform-side matching is deployed self-hosted or by vendor API, and every documented failure of this protective system is a protection gap — a non-participating platform, an encrypted surface, a re-encoded copy that no longer matches, a partner copy a withdrawal cannot reach — never a false accusation.

Sources: stopncii, nationalcenterformissingexpl, thorn

Appears on: /domains/cases/stopncii-hash-removal

EmpiricalA University of Michigan audit study (ACM CSCW 2026) posted 50 synthetic-persona deepfake nude images to X and reported …

A University of Michigan audit study (ACM CSCW 2026) posted 50 synthetic-persona deepfake nude images to X and reported half through X's non-consensual-nudity mechanism and half as DMCA copyright violations: the DMCA reports achieved 100 percent removal within about 25 hours (mean 20.3 hours), while the non-consensual-nudity reports achieved zero removals over 21 days. The study tested X's victim-facing report channels, not hash matching, and X is a StopNCII partner — independent evidence that a platform's participation in a hash program can coexist with a non-functioning victim-facing report channel, which is why the match channel and the platform report channel are modeled separately.

Sources: zhangandcolleagues2026

Appears on: /domains/cases/stopncii-hash-removal

EmpiricalThe voluntary hash indexes acquired a mandatory legal overlay during 2024-2026, on separate tracks that operate independ…

The voluntary hash indexes acquired a mandatory legal overlay during 2024-2026, on separate tracks that operate independently of hash-program membership. In the US, the TAKE IT DOWN Act's Section 3 notice-and-removal duty took effect May 19, 2026: covered platforms must remove reported non-consensual intimate imagery and known identical copies within 48 hours of a valid request, enforced by the FTC, which opened a complaint portal and sent compliance letters to fifteen major companies including Alphabet, Discord and X — companies that do not all participate in either voluntary hash index. In the UK, NCII sharing became a priority offence under the Online Safety Act via 2024 regulations, a creation offence covering purported intimate images including deepfakes came into force February 6, 2026 under the Data (Use and Access) Act 2025, Ofcom's hash-matching code measures were expected in force from summer 2026, and a Crime and Policing Bill amendment announced February 19, 2026 would require 48-hour takedown with penalties up to 10 percent of worldwide turnover. The statutes mandate responding to removal requests; they do not require joining either hash index, and they leave the indexes' unauditability untouched.

Sources: federaltradecommission2026, fladgatellp2024, osborneclarke2026

Appears on: /domains/cases/stopncii-hash-removal

EmpiricalUnder Digital Services Act Articles 15(1)(e) and 42(2) — which require indicators of the accuracy and possible rate of e…

Under Digital Services Act Articles 15(1)(e) and 42(2) — which require indicators of the accuracy and possible rate of error of automated content moderation, broken down by each official Member State language — X published, for the period 1 October 2024 to 31 March 2025, appeal and overturn rates for automated hateful-conduct visibility filtering across twenty named EU languages: appeal rates from 0.0 to 14.0 percent and overturn rates from 0.0 percent (Hungarian) to 72.2 percent (Swedish), with English at 40.4 percent. The independent audit of the sector's harmonised filings — a different set of reports from the April 2025 table above — found that of the eight largest EU platforms, four reported language-wise classification metrics (Facebook and Instagram across all 24 official languages, LinkedIn 22, X 7), three substituted country, and one reported none. Those are self-reported classification metrics; this atlas, not the audit, reads the April 2025 table as the only language-indexed error signal of its kind, because it indexes an outcome of the correction channel — appeals and overturns — rather than a platform's own accuracy figure. The figures are the platform's own legally mandated self-report, filed under Commission scrutiny, and no independent measurement of the deployment's per-language enforcement accuracy exists in the record.

Sources: regulation2022, xinternetunlimitedcompany2025, trujillo2026

Appears on: /domains/cases/dsa-multilingual-hate-speech

EmpiricalAs of March 2025, X named 1,352 content moderators with primary-language professional proficiency covering 8 of the 24 o…

As of March 2025, X named 1,352 content moderators with primary-language professional proficiency covering 8 of the 24 official EU languages — 1,197 of them (88.5 percent) in English — plus a secondary list of 185 people, stated to be not distinct from the first, adding three more languages; thirteen official languages appear in neither list. For those, the company states: 'In situations where we need additional language support, we use translation services and/or machine translation tools, to investigate and address challenges in additional languages.' Flagged content is either human-reviewed before action or auto-actioned on the model's historical accuracy, and the models are trained on labels generated by the company's own trained moderators. Civil-society review found the same reviewer-language asymmetry sector-wide (for example Meta at 1 named Maltese reviewer against 3,110 Spanish) and concluded there is 'no clarity on moderation outcomes per country or per language.'

Sources: xinternetunlimitedcompany2025a, internationalnetworkagainstc2025

Appears on: /domains/cases/dsa-multilingual-hate-speech

EmpiricalThe published per-language error indicator is an overturn rate conditional on appeal: its denominator, the appeal rate, …

The published per-language error indicator is an overturn rate conditional on appeal: its denominator, the appeal rate, varies by language from 0.0 to 14.0 percent, and the same metric family exceeds 100 percent for other policies in the same report, marking it as a period ratio rather than a probability or population error rate. The three languages with 0.0 percent appeal rates (Bulgarian, Greek, Lithuanian) carry no overturn rate at all — no error measurement exists for them — and four official languages (Croatian, Irish, Maltese, Slovak) are absent from the table entirely. The platform's own footnote on blank cells ('Cells that are blank mean that there was no enforcement. For cells containing 0.0% value, there were no cases of successful appeals or overturns') is ambiguous, and the reading used here — blank overturn cells aligning with the three zero-appeal columns — is the internally consistent one. The published overturn rates do not sort by reviewer staffing (Swedish, with no named reviewer, at 72.2 percent; Hungarian, also with no named reviewer, at 0.0; English, with 1,197 reviewers, at 40.4), and no monotone relation between staffing and measured error is supported by the record.

Sources: xinternetunlimitedcompany2025, xinternetunlimitedcompany2025a

Appears on: /domains/cases/dsa-multilingual-hate-speech

EmpiricalIn the reporting period immediately following the twenty-language table (1 April to 30 June 2025), X published the same …

In the reporting period immediately following the twenty-language table (1 April to 30 June 2025), X published the same visibility-filtering accuracy indicators broken down by country rather than by language, and its first harmonised-template filing covered language-wise metrics for 7 of the 24 official languages per the independent audit; independent analysis of the sector's reports separately found their figures disconnected and their category vocabularies inconsistent across platforms. Every moderation decision files a statement of reasons to the Commission's public DSA Transparency Database, whose landing figure records roughly 42 percent of submitted decisions across all platforms as fully automated. Enforcement against the operator is active but distinct from these tables: the Commission's formal proceedings opened 18 December 2023 — content moderation and dissemination of illegal content among the grounds — remain open with no published findings and were extended on 26 January 2026 (recommender systems, plus a new investigation into the platform's integrated generative AI assistant), and the EUR 120 million fine of 5 December 2025, the first DSA non-compliance decision, concerns transparency obligations (deceptive verified-checkmark design under Article 25(1), the advertising repository under Article 39, researcher data access under Article 40(12)) — not moderation accuracy, hate-speech enforcement, or the language tables.

Sources: xinternetunlimitedcompany2025a, trujillo2026, europeancommission2023, europeancommission2025, eucrim2026, ohnesorge2025, europeancommissiona

Appears on: /domains/cases/dsa-multilingual-hate-speech

EmpiricalM-Shwari, launched November 2012 by the Commercial Bank of Africa (now NCBA) in partnership with Safaricom over the M-PE…

M-Shwari, launched November 2012 by the Commercial Bank of Africa (now NCBA) in partnership with Safaricom over the M-PESA rails, scores customers with fully automated credit approval rules run on six months of the mobile network operator's transaction and airtime record: a score assigned at account opening irrespective of when the customer borrows, never disclosed to the customer, expressed only as a credit limit that is zero below the cutoff and grows on timely repayment. Loans in the evaluated era started at KSh 100 (the minimum has been KSh 2,000 since August 2020, with smaller borrowing redirected to the Fuliza overdraft), run 30 days at a 7.5 percent facilitation fee unchanged from 2012 through 2026, roll over automatically with a further 7.5 percent, and are reported to a credit reference bureau after 120 days of non-payment.

Sources: suri2021, cook2015, cytonninvestments2020

Appears on: /domains/cases/alternative-data-digital-credit

EmpiricalA peer-reviewed regression-discontinuity evaluation of M-Shwari found that 34 percent of eligible households took a loan…

A peer-reviewed regression-discontinuity evaluation of M-Shwari found that 34 percent of eligible households took a loan, eligible households were 6.3 percentage points less likely to forego expenses in response to a negative shock, and digital loans did not substitute for other credit; the same evaluation's administrative data show no significant difference in first-loan default between borrowers just above and just below the score cutoff — a coefficient of 0.007 against a control mean of 0.066 — reported by the authors as a check on their study design, local to the cutoff, and the only published measurement in this deployment family of what the score separates at the margin where it decides.

Sources: suri2021

Appears on: /domains/cases/alternative-data-digital-credit

EmpiricalAt market level — findings about the Kenyan digital-credit market and its app-based lender class, not about M-Shwari — n…

At market level — findings about the Kenyan digital-credit market and its app-based lender class, not about M-Shwari — nationally representative survey evidence (FinAccess 2019, borrower n = 2,194) puts digital-credit default at 13.8 percent against 6.4 percent formal, negative bureau listing at 5.3 percent of digital borrowers against 1.4 percent formal, 38.4 percent of digital defaults reported to a bureau against 22.7 percent of formal ones, and roughly 90 percent of all blacklistings attributed to digital credit, on a digital cohort the source documents as lower-income and more vulnerable than the formal cohort; the Competition Authority of Kenya's market inquiry with Innovations for Poverty Action (a 793-user phone survey plus an audit of provider transaction and account-level data) found a mean effective APR of 280.5 percent and median 96.5 percent, 77 percent of mobile loan users unable to repay at least once, 40 percent able to state their last loan's cost within five percent, and 27 percent aware of other providers' charges.

Sources: johnen2021, competitionauthorityofkenyaa2021

Appears on: /domains/cases/alternative-data-digital-credit

EmpiricalThe Central Bank of Kenya's Credit Reference Bureau Regulations of 8 April 2020 barred negative listings at or below one…

The Central Bank of Kenya's Credit Reference Bureau Regulations of 8 April 2020 barred negative listings at or below one thousand shillings, forced a one-off mass delisting below that floor, and required the credit score in every credit report, while non-bank digital lenders were barred from the credit information system until licensing; the Digital Credit Providers Regulations of 18 March 2022 brought digital lenders under CBK licensing, made two-way credit information sharing mandatory with the KSh 1,000 floor carried forward, required pre- and post-listing notice, capped recovery on non-performing loans, required ability-to-repay assessment and a complaints mechanism with 30-day resolution, prohibited phone-book access and shaming in collection, and made product and pricing-parameter changes subject to prior written approval; the Office of the Data Protection Commissioner's December 2023 guidance documents app-class providers taking 'phone contacts, call logs, pictures, location, messages, all mobile phone applications and everything in the phone' stored 'as a collateral' and sharing customer data with marketing companies without consent, directing collection of transaction detail only. By 14 April 2026 the CBK had licensed 227 digital credit providers from more than 800 applications, with licensed providers having granted 7.5 million loans worth KSh 133.5 billion as of February 2026; the Business Laws (Amendment) Act 2024 recategorised digital credit as non-deposit-taking credit business, and draft successor regulations were exposed for comment in August 2025 — a draft, not a replacement.

Sources: centralbankofkenya2020, centralbankofkenya2022, competitionauthorityofkenyaa2021, officeofthedataprotectioncom2023, centralbankofkenya2026, centralbankofkenya2025

Appears on: /domains/cases/alternative-data-digital-credit

EmpiricalProPublica and The Capitol Forum (Patrick Rucker, Maya Miller, David Armstrong), computing from internal Cigna records a…

ProPublica and The Capitol Forum (Patrick Rucker, Maya Miller, David Armstrong), computing from internal Cigna records and interviews with former employees, reported on March 25, 2023 that Cigna's PxDx review flags claims where the billed procedure code does not pair with the diagnosis code on a payer-authored list, and that company medical directors then sign the flagged denials in batches without opening patient files: over 300,000 payment requests denied through this method across a two-month period in 2022, at an average of 1.2 seconds of physician attention per case, with individual medical directors signing between roughly 60,000 and 121,000 denials in one to two months. A former Cigna doctor told the reporters, 'We literally click and submit. It takes all of 10 seconds to do 50 at a time.' The same records, per the investigation, show the match list being extended on cost grounds: adding autonomic-nervous-system testing in 2014 carried an internal projection of roughly 2.4 million dollars a year in savings, and the executive credited with developing the process said it had 'undoubtedly saved billions of dollars.' Cigna publicly disputes the article's characterization of the process - a spokesperson called a complaint built on it 'based on an article riddled with factual errors and misinformation' - and has published no substitute figures; the underlying documents are now in discovery. The investigation received the April 2023 Sidney Award. These figures are the investigation's computation and are not adjudicated fact.

Sources: propublica2023, sidneyhillmanfoundation2023, healthcaredive2023

Appears on: /domains/cases/cigna-pxdx

EmpiricalCigna's own published account of PxDx, which is vendor-tier evidence and is corroborated in outline by the investigation…

Cigna's own published account of PxDx, which is vendor-tier evidence and is corroborated in outline by the investigation, states that the review is 'procedure to diagnosis' code matching applied to roughly 50 common, relatively low-cost tests and procedures; that 94 percent of the claims subject to it are automatically approved and paid; that denials issued through it are 'less than 1 percent of our total volume of claims'; that the review 'occurs after the patient has received treatment and once their physician bills for the treatment'; that it 'does not involve algorithms, artificial intelligence, or machine learning'; and that in-network patients should not be billed for services denied this way. Two elements of that account are structural facts rather than contested framing: no care is gated by this review, because it runs after treatment has been delivered, so the decision allocates payment rather than access; and the screen is a deterministic list lookup rather than a learned system, which both sides of the dispute agree on. The wrong the record supports concerns the emptiness of the physician-review layer above the match and the authorship of the match list, not a model erring.

Sources: thecignagroupnewsroom2023, propublica2023

Appears on: /domains/cases/cigna-pxdx

EmpiricalThe correction channel in this record is documented in two stages with different positions, and its reach and its per-it…

The correction channel in this record is documented in two stages with different positions, and its reach and its per-item effect are separate facts. A denial appealed inside the plan goes to a different Cigna doctor; beyond that, an independent review organisation outside the plan can be reached. In the one patient arc the record follows end to end - a roughly 350 dollar vitamin-D blood test denied in autumn 2021 as not medically necessary - the internal appeal upheld the denial and the external independent reviewer reversed it roughly seven months after the denial was issued. What gates the channel is a design expectation rather than a measured rate: ProPublica reports from company records that Cigna internally estimated only about 5 percent of people would appeal a denial, on a set of claims selected for being low-dollar. That figure is an internal expectation about appeal propensity and must not be read as an observed appeal rate. Cigna has published no internal overturn rate for this review. The appeal and overturn statistics quoted in the May 2023 congressional correspondence (roughly one in five denials appealed, about 80 percent of appeals overturned) are Medicare Advantage prior-authorization figures the committee used as an analogy; they measure a different programme and are not measurements of this review.

Sources: propublica2023, rucker2023, u2023c

Appears on: /domains/cases/cigna-pxdx

EmpiricalOn October 8, 2025 the California Department of Managed Health Care fined Cigna HealthCare of California, Inc. 500,000 d…

On October 8, 2025 the California Department of Managed Health Care fined Cigna HealthCare of California, Inc. 500,000 dollars for improperly denying providers' claims as not medically necessary. The Department found that the plan 'reviewed and denied claims without physicians conducting clinical reviews of the claims prior to issuing denials' and that it used a different review process than the policy it had filed with the Department. Cigna agreed to pay the fine and to corrective actions including re-reviewing denials issued under the non-compliant process going back two years and revising and refiling its review policy. Department Director Mary Watanabe said the stability of the health care delivery system is impacted when health plans wrongly deny the payment of claims for health care services. These are regulator findings agreed to by the regulated entity and are stated as fact. The scope caveat is load-bearing: the Department's release names neither PxDx nor the investigation, so this action is described as addressing the plan's claims-review practice, consistent with the documented pattern, and never as a PxDx fine; the respondent is the California-regulated plan entity rather than the national group. No report of the two-year re-review's completion or its results has been located as of August 2026. The statutory standard the parallel state-law claim rests on is California Health and Safety Code section 1367.01(e), under which no individual other than a licensed physician or a licensed health care professional competent to evaluate the specific clinical issues may deny or modify requests for authorization for reasons of medical necessity.

Sources: californiadepartmentofmanage2025, kistingleung2025

Appears on: /domains/cases/cigna-pxdx

EmpiricalThe federal class litigation is live, mixed and unadjudicated on the merits. Kisting-Leung v. Cigna Corp. (E.D. Cal., No…

The federal class litigation is live, mixed and unadjudicated on the merits. Kisting-Leung v. Cigna Corp. (E.D. Cal., No. 2:23-cv-01477-DAD-CSK) was filed July 24, 2023. On March 31, 2025 Judge Dale A. Drozd granted in part and denied in part the motion to dismiss the third amended complaint: the ERISA section 1132(a)(1)(B) denial-of-benefits claim was dismissed with leave to amend, and plaintiffs elected on April 11, 2025 not to replead it; the ERISA section 1132(a)(3) fiduciary-duty claim proceeds; and the California unfair-competition claim proceeds on the licensed-physician-review theory. Three of the six original plaintiffs, including the named lead plaintiff, were dismissed for lack of standing after Cigna's Rule 12(b)(1) factual attack, supported by a declaration stating there were no PxDx denial letters associated with their claims - which establishes from the defense's own evidence that not every denial by this insurer runs through this review. Assuming arguendo the most deferential standard, the court held that reading the plan term requiring medical-necessity determinations by a medical director 'as allowing an algorithm to make the decision so long as a medical director pushes the button' would conflict with the plain language of the plan and constitute an abuse of discretion. That is a pleading-stage interpretation ruling on allegations assumed true, not a factual finding about what the company did. Cigna answered May 2, 2025; one further plaintiff was voluntarily dismissed by stipulation on August 13, 2026, leaving two. Per the parties' August 24, 2026 stipulation, Cigna has produced roughly 2.1 million pages and depositions are coordinated with Snyder v. The Cigna Group (D. Conn., No. 3:23-cv-1451-OAW, filed November 2, 2023), described there as another class action involving Cigna's PxDx process, so Cigna witnesses sit once for both actions; fact discovery closes in autumn 2026 and class-certification briefing begins October 29, 2026. No class has been certified and the trial setting will move. The House Energy and Commerce Committee's letter of May 16, 2023 demanded PXDX process documents, legality memoranda, the list of plans subject to the review, per-medical-director denial records and 2022 review, denial, appeal and overturn counts by May 30, 2023; no public committee findings, hearing record or released production has been located as of August 2026, and nothing here implies the inquiry concluded or found anything.

Sources: kistingleung2025, kistingleungv2026, georgetownlaw2026b, u2023c, healthcaredive2023

Appears on: /domains/cases/cigna-pxdx

EmpiricalEviCore by Evernorth (eviCore healthcare MSI, LLC), a Tennessee-domiciled utilization-review entity owned by The Cigna G…

EviCore by Evernorth (eviCore healthcare MSI, LLC), a Tennessee-domiciled utilization-review entity owned by The Cigna Group since 2018, performs delegated prior-authorization review for more than 100 client insurers — including UnitedHealthcare, Aetna, Blue Cross Blue Shield plans, Medicare and Medicaid contractors, and its own parent — covering about 100 million people, roughly one in three insured Americans. The mechanism is common ground between the company and its critics: an artificial-intelligence-backed algorithm scores each submitted request with a probability of approval, requests above the operating threshold are approved with no clinical review, requests below it route to in-house nurses and then to physicians, and only a physician may issue a denial. The algorithm approves or routes; it denies nothing. On October 23, 2024 ProPublica and The Capitol Forum (T. Christian Miller, Patrick Rucker, David Armstrong; co-published by CNN on November 7), working from internal documents, corporate data and interviews including five former employees, reported that this routing threshold is adjustable and that insiders called it the dial: 'If EviCore wants more denials, it can send on for review anything that scores lower than a 95%,' one former employee said, and a former executive said of the review rate, 'We could control that. That's the game we would play.' The same investigation reports strictness sold as a product feature — a marketed return of three dollars of medical spend avoided per dollar of fees, sales staff touting denial increases of up to 15 percent, insurers including Aetna and Cigna requesting 'high touch' configurations that send more cases to review, and 'risk' contracts under which the vendor keeps or splits what it saves below a client's baseline spending target. EviCore and Cigna dispute this characterization, stating that the company 'uses the latest evidence-based medicine' and that its algorithms exist 'ONLY to accelerate approval of appropriate care and reduce the administrative burden on providers.' The dial and everything hanging off it are the investigation's findings, resting on internal documents and largely unnamed former employees, and are not adjudicated fact.

Sources: miller2024, coalitiontostrengthenamerica2024

Appears on: /domains/cases/evicore-prior-auth-dial

EmpiricalThe one enforcement action on this vendor's record is procedural, small and admitted. Between September 19, 2023 and Jan…

The one enforcement action on this vendor's record is procedural, small and admitted. Between September 19, 2023 and January 29, 2024 the Connecticut Insurance Department conducted a market conduct examination of 196 of eviCore's calendar-2021 Connecticut files. Its report and stipulation and consent order of February 5, 2024, Docket MC 24-15, records 77 adverse-determination letters that omitted the mandatory notice of the 120-day external appeal; urgent-care, appeal and retrospective determinations issued past their statutory 48-or-72-hour, 30-day and 60-day clocks; three files whose documentation could not support regulatory review; two denials not reviewed by an appropriate clinical peer; and one appeal decided by the same physician who made the initial denial. The violations are of Conn. Gen. Stat. sections 38a-591b and 38a-591d and Regulation 38a-591-8. eviCore admitted the allegations, paid a $16,000 fine and undertook to file a corrective-action report within 90 days. These are utilization-review compliance findings and are stated here as fact because they were admitted. Their scope is load-bearing and is stated with them: the examination read FILES, it made no finding about the algorithm, the routing threshold or any denial rate, and presenting the fine as algorithmic enforcement would overstate the only enforcement loop on this record that ever closed. No regulator on this record has examined a routing threshold. (ProPublica's shorthand of '77 violations found in a review of 196 files' compresses the order: 77 is the external-appeal-language category alone, and the total findings span seven categories.)

Sources: stateofconnecticutinsuranced2024, miller2024

Appears on: /domains/cases/evicore-prior-auth-dial

EmpiricalThree measured signals sit around the contested threshold, each with its own limit. First, ProPublica's ANALYSIS of data…

Three measured signals sit around the contested threshold, each with its own limit. First, ProPublica's ANALYSIS of data that Arkansas publishes about eviCore's book of business found prior-authorization requests turned down in full or in part almost 20 percent of the time since 2021, against roughly 7 percent for Medicare Advantage plans overall in 2022. The ~20 percent is the reporters' analysis and not a finding Arkansas published, and the ~7 percent comparator covers a different population and a different programme, so the pair is an order-of-magnitude contrast and never a controlled comparison. Second, Vermont Medicaid materials showed quarterly denial rates under this vendor's review ranging from 6.1 percent to almost 15 percent, and a 2019 Vermont presentation recorded advanced-radiology requests falling 16 percent, to 3,629, and cardiology requests falling 38 percent after review began — the demand suppression the industry markets as the 'sentinel effect', whose governance property is that a request never submitted appears in no denial statistic taken by anyone. Third, a former eviCore physician reviewer, maternal-fetal medicine specialist Gail Miller, told the investigation she was required to decide at least 15 cases an hour, one every four minutes, frequently outside her own specialty, and left after nine months. All three are the ProPublica/Capitol Forum investigation's reporting, disputed in characterization by eviCore and Cigna. None of them is a measurement of the algorithm: every denial rate here sits downstream of human review and cannot be decomposed into a model error, and no deployment error rate for this system is published anywhere. Neither is the share of requests approved outright rather than routed to review, which is quantified nowhere public at all. The HHS OIG evaluation of Medicare Advantage prior-authorization denials and the Senate Permanent Subcommittee on Investigations majority staff report of October 17, 2024 document insurer-side prior-authorization automation raising denial rates, but both examine INSURERS and not this vendor; they are domain context, and nothing in them is attributed to eviCore.

Sources: miller2024, u2022, ussenatepermanentsubcommitte2024

Appears on: /domains/cases/evicore-prior-auth-dial

EmpiricalThe scoring layer and the human reviewers both work against clinical criteria that eviCore authors itself, and the recor…

The scoring layer and the human reviewers both work against clinical criteria that eviCore authors itself, and the record contains one dated instance of that arrangement failing. Per the investigation, citing the audit, a 2018 CMS audit found that outdated eviCore cancer guidelines led to inappropriate denials for 30 patients at the Blue Cross insurer HCSC; eviCore retrained staff in response. Two properties of that event matter more than its size. It reached this vendor only through one of its insurer CLIENTS rather than through any regulator of the vendor itself, and it is the single documented demonstration that an outside reader can adjudicate this vendor's determinations against its own criteria — the check that the public record otherwise shows nobody performing. No audit cadence for the criteria library is published, and no independent technical evaluation of the scoring layer exists in the public record. An adjacent settlement suggests the incentive belongs to the market structure rather than to one firm: Carelon, formerly AIM and the utilization-management arm of Elevance, paid $13 million in 2022 over wrongful-denial techniques, without admitting fault.

Sources: miller2024

Appears on: /domains/cases/evicore-prior-auth-dial

EmpiricalThe vendor's published position is the other side of this record and is vendor-tier evidence throughout. An Evernorth ar…

The vendor's published position is the other side of this record and is vendor-tier evidence throughout. An Evernorth article updated October 22, 2024 — the eve of the investigation's publication — states 'We aren't in the denial business, we're in the approval business', and reports more than 500 board-certified physicians across more than 60 specialties, more than 1,200 nurses and other clinical specialists, roughly two-thirds of medical decisions rendered in real time, and 90 percent of approvals completed within one business day, with peer-to-peer consultations framed as educational. Every one of those figures is a company claim with no independent verification, and the auto-approve versus routed-to-review split they imply is published nowhere. Two later company statements are operator-tier and are dated rather than characterized. On its Q1 2026 earnings call, held April 30, 2026, Cigna reported removing hundreds of tests, procedures and services from prior authorization altogether and cutting medical prior-authorization volume by about 15 percent, citing the 2025 industrywide voluntary prior-authorization commitments; on the same call, incoming chief executive Brian Evanko announced a strategic review of alternatives for eviCore, contemplating partnership or combination with other industry participants, citing scalability, management attention and industrywide standardization, with no transaction underway per the company. Any reading of this deployment should therefore avoid assuming stable Cigna ownership going forward: the owner of the adjustable parameter may change hands.

Sources: evernorththecignagroup2024, thecignagroupqearningscall2026

Appears on: /domains/cases/evicore-prior-auth-dial

EmpiricalIBM's Watson for Oncology, a treatment-recommendation clinical decision support product trained by Memorial Sloan Ketter…

IBM's Watson for Oncology, a treatment-recommendation clinical decision support product trained by Memorial Sloan Kettering Cancer Center physicians, was in use at roughly 50 hospitals across five continents by September 2017, with only two named US adopters and primary markets in India, South Korea, China, Thailand and Mongolia; more than 70 Chinese institutions adopted it amid published reliability concerns, and some adopting hospitals marketed it to patients. It was sold as machine reading of the medical literature. It read no local health record: at each site a person abstracted the case by hand into a structured form of 13 to 17 attributes — 17 at a Danish pilot, 13 in Korean studies — and at Jupiter Medical Center in Florida a nurse spent roughly 90 minutes a week doing so, with the treating oncologist finding the output largely redundant with the plan already in hand. Its recommendation logic was undisclosed to buyers and IBM shipped multiple software versions without transparent change documentation, so evaluations run at different times were measuring different software.

Sources: ross2017, tupasela2020, strickland2019, jncijournalofthenationalcanc2017

Appears on: /domains/cases/ibm-watson-oncology

EmpiricalThe system's recommendations were computed from a curated knowledge base built on a small number of SYNTHETIC — hypothet…

The system's recommendations were computed from a curated knowledge base built on a small number of SYNTHETIC — hypothetical, hand-written — cancer cases authored by Memorial Sloan Kettering physicians working with IBM engineers, together with literature and guidelines those physicians selected, resting on the expertise of a few specialists for each cancer type rather than on guidelines or evidence; the engineering post-mortem records that synthetic cases were adopted as a workaround after the system could not learn from the medical literature. Internal IBM Watson Health slide decks of June and July 2017, presented by the division's deputy chief health officer to management and later obtained by STAT, recorded 'multiple examples of unsafe and incorrect treatment recommendations' and stated that these raised 'serious questions about the process for building content and the underlying technology.' That phrasing is IBM's own internal-document language as reported by STAT; IBM publicly contested STAT's characterisation, the examples came largely from testing and training exercises, and no patient harm from a recommendation is documented anywhere in this record. Global sales continued for years afterwards, and no correction of the knowledge base following the finding has been located.

Sources: ross2018, strickland2019, ross2017

Appears on: /domains/cases/ibm-watson-oncology

EmpiricalNo published study measured whether Watson for Oncology improved patient outcomes; the published evidence base is concor…

No published study measured whether Watson for Oncology improved patient outcomes; the published evidence base is concordance — how often clinicians chose what the system chose — and it is tiered rather than uniform. The flagship study of 638 breast-cancer cases at Manipal, published in Annals of Oncology in 2018, reported 93 percent agreement, a figure reached after a blinded tumour-board re-review of the non-concordant cases lifted agreement from 73 percent, pooling the 'recommended' and 'for consideration' output tiers, on an author list carrying four IBM Watson Health-affiliated co-authors. The strongest independent evaluation, at Gachon Gil Medical Center in Korea with no IBM involvement and declared conflicts of none, measured agreement at the strict 'recommended' level in 41.5 percent of 65 advanced gastric cancer patients (87.7 percent counting 'for consideration'), attributing the divergence to unaccounted patient history, US-centric regimens outdated or non-standard locally, national insurance not covering recommended agents, and the S-1 regimen — standard practice in Korea — being absent from the knowledge base. A Danish pilot found roughly one-third agreement, its physician citing overweighting of American studies, and the hospital declined to adopt; at UB Songdo Hospital in Mongolia, which lacked oncology specialists, clinicians followed the recommendations approximately 100 percent of the time. A peer-reviewed methodological critique concludes that concordance cannot establish safety or benefit — agreement is compatible with both parties being wrong, and disagreement cannot distinguish system error from clinician error — and documents the MSK-trained system functioning as a de facto ground truth against which other countries' practice was scored.

Sources: somashekhar2018, canadianjournalofgastroenter2019, tupasela2020, ross2017

Appears on: /domains/cases/ibm-watson-oncology

EmpiricalThe MD Anderson Oncology Expert Advisor was a separate product — an MD Anderson-owned build on IBM Watson technology, di…

The MD Anderson Oncology Expert Advisor was a separate product — an MD Anderson-owned build on IBM Watson technology, distinct from Watson for Oncology — which never reached clinical use. IBM support ended 1 September 2016; IBM and the university agreed the system was 'not ready for human investigational or clinical use, and its use in the treatment of patients is prohibited,' and it was 'not in clinical use and has not been piloted outside of MD Anderson.' The University of Texas System Administration Audit Office's special review, reported publicly in February 2017, found more than 62 million dollars paid to IBM and PricewaterhouseCoopers as of 31 August 2016 (approximately 39 to 40 million to IBM and 21 to 23 million to PwC across contemporaneous accounts), an 11.59 million dollar deficit spend against donations not yet received, fees consistently set just below the amount that would have required Board approval, invoices paid regardless of delivery, and procurement and IT-governance processes bypassed — while expressly disclaiming any opinion on the scientific basis or functional capabilities of the system. The figure is spend to date, not a fine or an audit-assessed loss, and the review is a procurement audit, not a scientific verdict on either product. No regulator enforcement action, medical-device action, or product-liability litigation over either system's recommendations was located as of August 2026, and no medical-device regulator reviewed Watson for Oncology's recommendations before it was marketed globally. IBM announced the sale of Watson Health's data and analytics assets to Francisco Partners on 21 January 2022; the sale closed on 30 June 2022, launching Merative around six named product families with the oncology treatment-recommendation products not among them. No formal discontinuation of Watson for Oncology was announced.

Sources: universityoftexassystemadmin2016, theregister2017, jncijournalofthenationalcanc2017, ibmnewsroom2022

Appears on: /domains/cases/ibm-watson-oncology

EmpiricalOn April 11, 2018 the US Food and Drug Administration granted De Novo request DEN180001, received January 12, 2018 under…

On April 11, 2018 the US Food and Drug Administration granted De Novo request DEN180001, received January 12, 2018 under Breakthrough Device review, permitting IDx LLC of Coralville, Iowa to market IDx-DR — in the agency's own words the first device authorized for marketing that provides a screening decision without the need for a clinician to also interpret the image or results, and the first autonomous diagnostic authorized in any field of medicine. The software is locked and deterministic, paired by the authorization to one nonmydriatic fundus camera model, and returns exactly one of two messages: more than mild diabetic retinopathy detected, refer to an eye care professional; or negative for more than mild diabetic retinopathy, rescreen in 12 months. In the pivotal trial (Abramoff et al., npj Digital Medicine 2018 — funded by the manufacturer, with company-affiliated authors) 900 adults with diabetes were enrolled at 10 primary care sites and 819 were fully analyzable, yielding 87.2% sensitivity (95% CI 81.8-91.2), 90.7% specificity (95% CI 88.3-92.7) and 96.1% imageability against a reading centre's widefield stereo photography and macular imaging, exceeding pre-specified endpoints of 85% and 82.5%; the FDA's own announcement states 87.4% and 89.5% for the same trial, an analysis-population difference. Camera operators were existing clinic staff who attested they had never performed ocular imaging, after a single four-hour standardized training.

Sources: abramoff2018, u2018a, u2026

Appears on: /domains/cases/idx-dr-autonomous-screening

EmpiricalThe autonomy is fenced by the labeling rather than by a reviewer. The device is indicated for adults 22 and over with di…

The autonomy is fenced by the labeling rather than by a reviewer. The device is indicated for adults 22 and over with diagnosed diabetes who have not previously been diagnosed with diabetic retinopathy, used with the paired camera; patients are not to be screened if pregnant — retinopathy can progress rapidly in pregnancy — or if they report persistent vision loss, blurred vision or floaters, or have previously been diagnosed with macular edema, severe non-proliferative, proliferative or radiation retinopathy or retinal vein occlusion, or have had retinal laser treatment, intraocular injections or retinal surgery. It detects diabetic retinopathy and no other condition, exposes no severity gradations and issues no confidence figure. That eligibility and exclusion screen is administered by clinic staff before any image is taken, which places the surviving human judgment in this workflow before the camera rather than after the result. The escalation runs the same way: when image quality is insufficient the operator re-images, with pharmacologic dilation as the labeled next step — 23.6% of pivotal-trial participants required it — and current labeling warns that a patient who still yields no result after dilation may have vision-threatening diabetic retinopathy.

Sources: digitaldiagnostics, u2018a, abramoff2018

Appears on: /domains/cases/idx-dr-autonomous-screening

EmpiricalThe measured benefit of this deployment is replicated, and every study of it carries a sponsorship or investigator-affil…

The measured benefit of this deployment is replicated, and every study of it carries a sponsorship or investigator-affiliation label. In the ACCESS randomised trial at two Johns Hopkins pediatric diabetes clinics (Wolf et al., Nature Communications 2024; investigator ties to the vendor disclosed; ages 8 to 21, investigational use below the device's cleared adult indication) 164 youth were randomised, and eye-exam completion within six months was 100% (81 of 81) in the autonomous-AI arm against 22% under scripted specialist referral, with follow-through after an abnormal result 64% against 22%, in a cohort that was 35% Black and 47% Medicaid-insured, and no adverse events. At system scale (Huang, Channa, Wolf et al., npj Digital Medicine 2024; same affiliation caveat; 30-plus primary care sites, roughly 17,600 diabetes patients a year) sites that deployed the system raised diabetic-eye-exam adherence from 46.1% to 54.5% between 2019 and 2021 while comparison sites moved -0.3 points, with Black patients gaining 12.2 points and the Asian-to-Black adherence gap narrowing from 15.6 to 3.5 points. A five-health-system programme covering roughly 151,000 diabetes patients (Journal of CME 2025, an industry-supported education-journal evaluation used for texture rather than headline figures) recorded 20,160 screenings, 72% suitable for diagnosis, 24% of suitable screens positive, ophthalmology referrals up 23% and anti-VEGF treatment up 27%, with one system reporting a 118% rise in screening rate, against barriers of physician hesitancy and workflow fragmentation. What scaled the deployment was payment rather than the authorization: the AMA CPT Editorial Panel created code 92229 for point-of-care automated retinal analysis effective January 1, 2021, and CMS's CY2022 Physician Fee Schedule final rule of November 2, 2021 set the first national Medicare payment for it — approximately 45 to 47 dollars nationally at launch by crosswalk to CPT 92325, effective January 1, 2022, with later-year rates drifting down.

Sources: wolf2024, huang2024, journalofcme2025, digitaldiagnostics2021

Appears on: /domains/cases/idx-dr-autonomous-screening

EmpiricalThe one place trial performance and field performance part company is the image-quality gate rather than the classifier.…

The one place trial performance and field performance part company is the image-quality gate rather than the classifier. In the only fully vendor-independent evaluation located (Hunfeld et al., Scientific Reports 2026; Karlsburg Diabetes Hospital, Germany; 875 patients, February 2020 to November 2021) 26.1% of patients' images could not be analyzed by the device and 10.5% of patients yielded no image at all, against 96.1% imageability in the sponsor-funded pivotal trial; the measured drivers were capture-side and demographic — mean pupil diameter 2.65mm against 4.28mm, mean age 63.9 years against 46.4, impaired retinal view with cataract documented in 61% of non-analyzable right eyes, and variability between examiners. Among images the device could analyze it held up, at 94.4% sensitivity and 90.5% specificity for severe disease, with exact grade agreement against an ophthalmologist at 54.2%. A five-health-system US programme found 28% of screenings not suitable for diagnosis, corroborating the direction. An earlier external evaluation of the pre-authorization European version of the software (van der Heijden et al., Acta Ophthalmologica 2018; Hoorn Diabetes Care System, Netherlands; 1,415 patients with type 2 diabetes, 898 of sufficient image quality; the founder a co-author) adds a separate caution about measurement itself: sensitivity for referable disease was 68% under one human grading scheme and 91% under another on the same eyes, with the human reference graders agreeing among themselves only 40% to 61% of the time.

Sources: hunfeld2026, abramoff2018, vanderheijden2018, journalofcme2025, digitaldiagnostics

Appears on: /domains/cases/idx-dr-autonomous-screening

EmpiricalThe compensating controls that replaced the physician overread are regulatory and contractual, and their record is thin …

The compensating controls that replaced the physician overread are regulatory and contractual, and their record is thin in the way an empty docket is thin. Direct queries of the FDA's public device databases on August 28, 2026 return no adverse-event reports for this device under either product name, no recalls for the product code or for the firm, and exactly one adverse-event report across the entire retinal-diagnostic-software class, for a competitor device; no product-liability or other litigation over the system was located. That is an empty reported-event docket after eight years, not a measurement of clinical safety: a screening false negative that waits out its 12-month rescreen interval would rarely generate a device adverse-event report at all. The De Novo created a durable device class — retinal diagnostic software, 21 CFR 886.1100, product code PIB, with special controls — through which seven follow-on clearances across four firms have since passed, including this device's own version 2.3 in 2021 and 2022, the documented change-control channel for a locked algorithm. IDx renamed itself Digital Diagnostics Inc. on August 19, 2020 alongside an acquisition, and the product was later renamed LumineticsCore. Academic legal literature records that the manufacturer carries medical malpractice liability insurance for the system and assumes liability for injuries arising from it; trade reporting scopes that assumption contractually to the system's diagnostic output rather than downstream care management, and no case has tested it.

Sources: u2026, researchhandbookonhealth2024, digitaldiagnosticsviaprnewsw2020

Appears on: /domains/cases/idx-dr-autonomous-screening

EmpiricalThe object corrected here is a clinical equation, not a learned system. The standard formulas for estimating glomerular …

The object corrected here is a clinical equation, not a learned system. The standard formulas for estimating glomerular filtration rate from a serum creatinine result — the 1999 MDRD equation and then the 2009 CKD-EPI creatinine equation — applied a coefficient that raised the estimated kidney function of any patient identified as Black; in the 2009 equation that coefficient was 1.159 (95 percent CI 1.144 to 1.170). The stated biological rationale, higher average muscle mass, treated race as a biological rather than a social category, and the MDRD equation behind it was derived from roughly 1,400 White and fewer than 200 Black patients. Inker and colleagues reported in the New England Journal of Medicine on 23 September 2021 that the race-including equation overestimated measured GFR in Black patients by a median of 3.7 mL/min/1.73m2, and concluded that race in eGFR equations is a social and not a biologic construct; OPTN's own patient materials state in a different register that the race variable automatically increased all Black patients' eGFR values, by as much as 16 percent. A kidney transplant candidate begins accruing waiting time when the estimate reaches 20 mL/min/1.73m2 or lower, so an estimate raised above that line delayed the clock; two studies cited in the nephrology commentary put the lost time at 1.3 and 1.9 years for affected Black candidates.

Sources: inker2021, pavlakis2023, hrsaoptn

Appears on: /domains/cases/optn-egfr-race-correction

EmpiricalThe correction arrived in two board actions. After a public comment period running 27 January to 23 March 2022, the OPTN…

The correction arrived in two board actions. After a public comment period running 27 January to 23 March 2022, the OPTN Board of Directors unanimously approved 'Establish OPTN Requirement for Race-Neutral eGFR Calculations' on 27 June 2022, requiring every kidney transplant program to use an eGFR formula without a Black-race variable from 27 July 2022; the National Kidney Foundation, whose joint task force with the American Society of Nephrology had recommended a race-free equation, called the vote an important first step and said there is no place for race-based variables in evaluating organs offered through the allocation system. Because a prohibition runs only forward, the Board then unanimously approved the retroactive remedy on 5 December 2022, effective 5 January 2023: each kidney program had to assess its waiting list, identify Black candidates disadvantaged by race-inclusive eGFR, establish whether a race-neutral calculation would have qualified them sooner, and apply to OPTN to backdate the qualifying date. Eligibility required documentation that the candidate's eGFR was over 20 mL/min under the race-inclusive calculation and 20 mL/min or lower without it. Programs had until 3 January 2024 to complete assessments, submit every qualifying modification, notify all kidney candidates before and after assessment regardless of race, and file an attestation; programs that did not comply would be referred to the Membership and Professional Standards Committee.

Sources: hrsaoptn, nationalkidneyfoundation2022, unos2023, hrsaoptna, pavlakis2023

Appears on: /domains/cases/optn-egfr-race-correction

EmpiricalThe undo ran by hand, program by program. Trade guidance written for transplant programs describes the working method: i…

The undo ran by hand, program by program. Trade guidance written for transplant programs describes the working method: identify candidates through the OPTN custom reporting tool's 'Current waitlisted African American Candidates' query, pull historical laboratory results from the electronic medical record, from external and reference laboratories, from dialysis centers and through record-retrieval and health-information-exchange systems, compare the race-inclusive and race-neutral figures against the 20 mL/min threshold, submit an eGFR Waiting Time Modification Form in UNet with supporting documentation, and send two notifications — before and after assessment — to every kidney candidate regardless of race. The same guidance records the correction's own second wave as a documented operational pressure: some patients were getting a lot of time back, bouncing them to the top of the list to start receiving offers, which it describes as a significant operational impact on many transplant centers, especially in ensuring that a patient had been re-evaluated recently. Accrued waiting time is a term in the ranking donor kidneys are offered down, so a backdated qualifying date changes a candidate's position in the live deceased-donor offer sequence directly rather than serving as a note on a file.

Sources: theallianceorgandonationalli2023, hrsaoptna, hrsaoptn

Appears on: /domains/cases/optn-egfr-race-correction

EmpiricalOPTN published its own counts twice, and each figure carries the date its data was cut. The early monitoring report (pub…

OPTN published its own counts twice, and each figure carries the date its data was cut. The early monitoring report (published 4 December 2023, data through 5 July 2023) recorded more than 6,100 Black candidates with modified waiting times at a median of 1.7 years, of whom 491 had received a deceased-donor transplant and 15 a living-donor transplant; at that date 12 of 232 active kidney programs had submitted attestations. The one-year report (published 8 May 2024, data through 4 January 2024) recorded 14,701 waiting-time modifications processed at a median of 1.7 years, with roughly half of modified registrations receiving between one and three years, 2,709 modified candidates transplanted from deceased donors and 158 from living donors, and all 230 active kidney programs having submitted attestations confirming that lists were reviewed, candidates notified and required modifications submitted. These are counts of modifications and of registrations at specific snapshot dates rather than counts of distinct people, and they are cumulative figures that grew between the two reports.

Sources: hrsaoptn2023, hrsaoptn2024

Appears on: /domains/cases/optn-egfr-race-correction

EmpiricalAn independent national evaluation measured what the programme's own counts could not. Schold and colleagues, publishing…

An independent national evaluation measured what the programme's own counts could not. Schold and colleagues, publishing in the Journal of the American Society of Nephrology on 6 May 2025, assessed 44,912 Black candidate kidney waitlist registrations nationally and found that 32 percent (14,419) received an eGFR waiting-time modification worth a median of about 610 priority days, and that modified candidates had an adjusted hazard ratio of 2.85 (95 percent CI 2.7 to 3.02) for deceased-donor transplantation against candidates who were not modified. The same study found modification rates varying significantly by candidate characteristics and by transplant center; the investigators read that variability as indicating there was still mixed use of the policies and that the requirement was not articulated clearly at the beginning. The study's primary record was behind a paywall and a consent wall at verification time on 28 August 2026, and these figures are carried as corroborated through the trade report of the same study rather than as directly retrieved from the primary.

Sources: schold2025, healionephrology2025

Appears on: /domains/cases/optn-egfr-race-correction

EmpiricalThe measured unevenness produced a further governance action rather than a closed file. The Membership and Professional …

The measured unevenness produced a further governance action rather than a closed file. The Membership and Professional Standards Committee, having observed programs implementing the requirements in various ways, referred a follow-on project to the OPTN Minority Affairs Committee; after public comment from 21 January to 19 March 2025 the OPTN Board approved 'Monitor Ongoing eGFR Modification Policy Requirements' at its June 2025 meeting, effective 10 September 2025. The update converts the one-time legacy audit into a standing per-candidate obligation: every registered kidney candidate must be assessed for eligibility, and each program must maintain written protocols and document compliance in three areas — confirming a candidate's race, fulfilling the notification requirements, and seeking supporting documentation, naming at minimum which sources will be reviewed. The notification requirements (education, eligibility and outcome) apply to candidates registered on or after 4 January 2024, and the update removes the superseded 3 January 2024 attestation language. Programs must complete the strengthened requirements by 11 September 2026, a date that had not passed as of 28 August 2026.

Sources: hrsaoptn2025, hrsaoptna

Appears on: /domains/cases/optn-egfr-race-correction

EmpiricalThe equity framing around this programme belongs to the commentary and policy materials that used it, and the same comme…

The equity framing around this programme belongs to the commentary and policy materials that used it, and the same commentary records the critique. Pavlakis, writing in the Journal of the American Society of Nephrology in 2023, describes the waiting-time modification as a restorative justice project in kidney allocation and also records that the policy has been criticised as being unfair to people suffering under other inequities besides Black or African American race, and as both unfairly too broad and too narrow, leaving open the broader question of whether all patients should accrue predialysis waiting time. The remedy's scope is bounded on the face of the policy: it reaches registered kidney candidates whose documentation meets the eligibility rule, it addresses the eGFR-driven delay and no other source of delay, and it can only advance a qualifying date, never delay one. No litigation and no enforcement action appears anywhere in this record as of 28 August 2026; the record is affirmative governance — a prohibition, a retroactive modification programme, published monitoring, an independent peer-reviewed evaluation, and a tightening in response to measured variance.

Sources: pavlakis2023, hrsaoptna, hrsaoptn2025

Appears on: /domains/cases/optn-egfr-race-correction

EmpiricalPractice Fusion, Inc., a free ad-supported cloud electronic health record used by tens of thousands of provider practice…

Practice Fusion, Inc., a free ad-supported cloud electronic health record used by tens of thousands of provider practices, admitted in a stipulated Statement of Facts that it solicited and received $959,700 from an opioid manufacturer's marketing department for a clinical decision support alert — $144,600 for a retrospective analysis and $815,100 for the alert work, under a statement of work effective 1 March 2016. It admitted that it had modeled the sponsor's return on investment at 5.8 to 7.8 times cost, with a projected 'patient gain' of 2,777 and $8,458,232 to $11,277,643 in additional opioid revenue, and had deliberately kept that model out of the written proposal; that the sponsor's Director of eMarketing proposed edits to the alert's workflow and treatment-option list and Practice Fusion's chief medical officer approved them, neither the closing account director nor the approving officer having experience treating pain or prescribing schedule II narcotics; and that an early patient-safety concept, screening patients for opioid-abuse risk, was discussed and dropped. On 27 January 2020 the U.S. Attorney for the District of Vermont charged the company by a two-count felony information — one count of soliciting and receiving kickbacks under 42 U.S.C. 1320a-7b(b)(1) and one count of conspiracy under 18 U.S.C. 371 — and resolved it by deferred prosecution agreement, in what the Department of Justice and trade press reported as the first criminal action against an electronic health records vendor. The company was never convicted and entered no plea. The global resolution was $145,000,000: a $25,398,300 criminal fine, $959,700 in forfeiture equal to the payment, and a $118,642,000 civil settlement.

Sources: statementoffacts2020, deferredprosecutionagreement2020, healthcaredive2020

Appears on: /domains/cases/practice-fusion-opioid-cds

EmpiricalThe Pain CDS was not a statistical model. It was an authored cascade of three chained alerts over chart data: a prompt t…

The Pain CDS was not a statistical model. It was an authored cascade of three chained alerts over chart data: a prompt to record a pain score; a suggestion to complete a Brief Pain Inventory for patients with two or more pain scores at or above 4 out of 10 within three months or a chronic-pain diagnosis; and a prompt to create a pain follow-up plan, firing where pain at or above 4 was recorded twice within four months or an inventory was completed. It terminated in a drop-down of nine treatment options placed on equal footing, including 'Opioid Therapy (short-acting, long-acting/extended release)' alongside biofeedback, non-opioid analgesics, nonpharmacologic care, referral, surgery and 'pain resolved'. The stipulated record is that the CDC opioid-prescribing guideline published 15 March 2016 — start with immediate-release opioids, lowest effective dose, non-opioid therapy preferred — was circulated among the designers at both companies during development and was not incorporated, and that as built the alert offered extended-release opioids to opioid-naive patients and to patients whose pain was not chronic, contrary to that guideline, to the applicable clinical quality measure and to the sponsor's own approved product labeling. One documented sponsor edit is a 29 January 2016 change adding an 'Extended Release Opioid initiated' checkbox to trigger re-assessment. The alert named no drug brand at any point: unbranded clinical messaging was an explicit design feature and the admitted mechanism was steering between treatment categories rather than toward a product.

Sources: statementoffacts2020

Appears on: /domains/cases/practice-fusion-opioid-cds

EmpiricalThe Pain CDS ran from 6 July 2016 to the spring of 2019 and, per the stipulated Statement of Facts, 'alerted more than a…

The Pain CDS ran from 6 July 2016 to the spring of 2019 and, per the stipulated Statement of Facts, 'alerted more than approximately 230,000,000 times' — a count of alert displays, not of prescriptions and not of patients. Through 30 November 2016 alone it had fired during 21 million patient visits involving 7.5 million patients and 97,000 healthcare providers. Practice Fusion's own program analytics recorded that providers who received the alerts prescribed extended-release opioids at a higher rate than those who did not, with a general shift from immediate-release toward extended-release largest in emergency medicine, orthopedics and pain medicine: a comparison the platform made on its own platform data, not an independent causal study. The same analytics answered the clinical question in the other direction. Presented to the sponsor at its headquarters on 14 December 2016, they reported extended-release opioids as the least effective of the listed options at lowering pain — 39.17 percent of patients treated with them had lower pain — and second-least effective among chronic-pain patients. That analysis was delivered to the sponsor as the program reporting it had contracted for and not to the prescribers being alerted, and no channel returning observed outcomes into the alert's content is documented. A sponsor attorney at that meeting expressed reservations and considered pausing the program; it continued.

Sources: statementoffacts2020, americanmedicalassociation2021

Appears on: /domains/cases/practice-fusion-opioid-cds

EmpiricalThe 2020 deferred prosecution agreement built a governance control where none had existed and the record then shows that…

The 2020 deferred prosecution agreement built a governance control where none had existed and the record then shows that control failing in operation. Its forward-looking terms: a three-year term extendable to a maximum of five, with the government obliged to seek dismissal with prejudice within 30 days of expiry; an independent Oversight Organization required to review and approve any sponsored clinical decision support before implementation; a public online repository of the documents underlying the conduct, hosted at Practice Fusion's expense with the sponsor's identity, employees and drug brands redacted; a compliance program separating clinical from commercial activities; and an obligation to report evidence of kickbacks by other record vendors. The Oversight Organization subsequently resigned. A U.S. Attorney's Office notice letter of 25 August 2021 alleged that the company had failed to retain a replacement, to give the organization adequate access to information and witnesses, and to pay certain of its expenses. A letter agreement of 17 March 2022, filed on the criminal docket, settled those allegations for $200,000 with an express no-admission clause and extended the agreement's term by eleven weeks, to 13 April 2023. The criminal docket shows a termination date of 9 May 2023; the dismissal order itself is PACER-gated and was not retrieved for this record.

Sources: deferredprosecutionagreement2020, settlementandreleaseagreemen2022, unitedstatesv

Appears on: /domains/cases/practice-fusion-opioid-cds

EmpiricalThe civil settlement reached conduct far beyond the single criminal count, and every part of it is a government allegati…

The civil settlement reached conduct far beyond the single criminal count, and every part of it is a government allegation resolved with an express no-admission clause. The United States alleged that Practice Fusion entered fourteen separate sponsored clinical decision support arrangements with various pharmaceutical manufacturers, first entered between 11 November 2013 and 17 August 2017, in which paying sponsors participated in designing the alerts — selecting the guidelines an alert cited, setting its trigger criteria, and in some cases drafting its language — with claims alleged tainted from April 2014 to April 2019. The Department of Justice's press release describes thirteen arrangements other than the criminally charged one; the settlement agreement itself enumerates fourteen in total, one contract having contained two. The United States separately alleged that the company obtained 2014 Edition ONC certification for software that disabled data export and lacked required SNOMED CT and LOINC support, causing false meaningful-use attestations between 2014 and 2017. The civil total was $118,642,000, of which $113,374,952 was federal — half of it, $56,687,476, as restitution — and $5,267,048 was escrowed for state Medicaid settlements; Connecticut announced its own share at $336,087.89. Allscripts Healthcare acquired Practice Fusion on or about 13 February 2018, after the conduct.

Sources: civilsettlementagreementbetw2020, stateofconnecticutofficeofth2020, healthcaredive2020

Appears on: /domains/cases/practice-fusion-opioid-cds

EmpiricalThe Practice Fusion resolution papers pseudonymize the sponsor as 'Pharma Co. X' and the court-ordered public repository…

The Practice Fusion resolution papers pseudonymize the sponsor as 'Pharma Co. X' and the court-ordered public repository was required to redact its identity, so the sponsor is identified here only on the strength of the sponsor's own criminal case and later Department of Justice releases — never on the Practice Fusion documents. Purdue Pharma L.P. pleaded guilty on 24 November 2020 in the District of New Jersey to three felonies, among them conspiracy to violate the Anti-Kickback Statute through its payments to an electronic health records company to install prompts intended to cause prescribing of its extended-release opioids, as part of a resolution reported at more than $8 billion. Later Department of Justice releases in the related obstruction case state that Purdue paid Practice Fusion almost one million dollars in exchange for altering its physician-facing user interface to generate more opioid prescriptions. Separately, Steven Mack, Practice Fusion's former Director of National Accounts on that account, pleaded guilty on 8 March 2021 to attempting to obstruct the grand-jury investigation by deleting hundreds of files from his company laptop, and was sentenced on 13 May 2024 to one year of probation, a $20,000 fine and forty hours of community service involving people suffering from drug addiction. In June 2021 the American Medical Association's House of Delegates, citing this case, adopted policy opposing direct-to-prescriber pharmaceutical promotional content in electronic health records and e-prescribing software.

Sources: usdepartmentofjustice2020, vtdigger2024, americanmedicalassociation2021

Appears on: /domains/cases/practice-fusion-opioid-cds

EmpiricalOn 13 February 2018 the FDA granted De Novo request DEN170073 for Viz.AI's ContaCT, creating the radiological computer-a…

On 13 February 2018 the FDA granted De Novo request DEN170073 for Viz.AI's ContaCT, creating the radiological computer-aided triage and notification device class (21 CFR 892.2080, product code QAS) that later stroke-AI entrants reach by 510(k). The FDA decision summary describes the device as a notification-only, parallel workflow tool: it analyses CT angiograms and alerts a neurovascular specialist 'in parallel to standard of care image interpretation', identification of suspected findings is not for diagnostic use beyond notification, and the benefit-risk section concludes 'there are no major risks for the device because the device operates in parallel to the current usual standard of care'. The three risks the order does list are triage risks: deprioritization of other patients' images, inappropriate use for primary interpretation, and delayed management if prioritization fails. In the pivotal standalone study of 300 CT angiograms from two US sites the algorithm measured 87.8% sensitivity (95% CI 81.2-92.5), 89.6% specificity (83.7-93.9) and an area under the curve of 0.91; in the same record notification arrived a mean 51.4 minutes earlier than the standard radiology pathway (7.3 against 58.7 minutes; medians 5.6 against 51.5; earlier in 42 of 44 cases) — a figure the FDA's own summary flags as potentially biased because it was computed on the 44 true-positive cases that happened to have documented standard-of-care notification times, and which this atlas therefore carries as that study's measurement rather than a field-wide constant.

Sources: usfoodanddrugadministration2018, u2018

Appears on: /domains/cases/viz-lvo-stroke-triage

EmpiricalIndependent multi-site evaluation of this detector measured accuracy materially below the pivotal figures, and lowest on…

Independent multi-site evaluation of this detector measured accuracy materially below the pivotal figures, and lowest on the distal occlusions. In a large integrated hub-and-spoke stroke network handling more than 6,000 code strokes a year, 3,851 patients were screened and 220 (5.7%) had a neuroradiologist-confirmed internal-carotid or MCA-M1 occlusion: sensitivity 78.2% (95% CI 72-83), specificity 97% (96-98), positive predictive value 61% (55-67), negative predictive value 99% (98-99), with sensitivity falling to 60.6% once MCA-M2 occlusions were included and false negatives driven by non-main-vessel occlusions, segmentation errors and difficult anatomy; those authors state that the software 'does not allow the clinician to confidently rule out' an occlusion on a negative result. A prospective study of 1,822 consecutive stroke-code CT angiograms (May 2019 to October 2020) at a three-tiered network found 190 occlusions and measured 93.8% sensitivity with 99.7% negative predictive value for internal-carotid-terminus and M1 occlusions, dropping to 74.6% sensitivity and 97.6% negative predictive value once M2 was included, with per-site detection of 100% at the carotid terminus, 93% at M1 and 49% at M2. A positive predictive value of 61% means roughly two proximal alerts in five are false. Both groups publish inside networks with substantial vendor-adjacent authorship, and the prospective study's site is a frequent vendor co-publisher.

Sources: karamchandani2023, matsoukas2022, usfoodanddrugadministration2018

Appears on: /domains/cases/viz-lvo-stroke-triage

EmpiricalIn the FY2021 hospital inpatient prospective payment final rule (85 FR 58432, 18 September 2020) CMS approved a new-tech…

In the FY2021 hospital inpatient prospective payment final rule (85 FR 58432, 18 September 2020) CMS approved a new-technology add-on payment for ContaCT of up to $1,040 per case — 65% of an applicant-estimated $1,600 cost — projecting roughly 12,700 cases and about $20.6 million in FY2021, billed through ICD-10-PCS 4A03X5D, with the newness period anchored to 1 October 2018; CMS renewed it for FY2022. The characterization of this as the first add-on payment granted for AI software is the applicant's, peer-reviewed commentary's and the press's framing: the rule text approves the technology without declaring a first, and in the same rule CMS questioned whether AI-based workflow streamlining is a unique mechanism of action, raised what 'new' means when an algorithm updates or a competitor's algorithm performs better, and reserved the category question for future rulemaking. What CMS accepted was the applicant's asymmetry argument — that a false positive causes only an earlier image review while a false negative leaves the patient in the unchanged standard-of-care pathway. The add-on is temporary by construction and its post-FY2022 status was not verified in rule text; this atlas does not assert that it is still being paid. Adoption counts for this deployment are vendor-tier and point-in-time: 'over 800 U.S. hospitals' at the August 2021 renewal release, and 1,400-plus hospitals and health systems across the US and Europe on the vendor's current undated page.

Sources: centersformedicaremedicaidse2021, hassan2021, viz2020

Appears on: /domains/cases/viz-lvo-stroke-triage

EmpiricalThe add-on payment left a per-use administrative trace, and that trace is the only system-wide measurement of this deplo…

The add-on payment left a per-use administrative trace, and that trace is the only system-wide measurement of this deployment anyone holds. A 2026 Harvey L. Neiman Health Policy Institute study in the American Journal of Neuroradiology read 2,116 Medicare inpatient acute-ischemic-stroke episodes at 1,076 facilities between October 2020 and December 2023: add-on-billed use of the tool peaked at 21% of eligible episodes in 2022 and then declined, 14.8% across the whole window. Use concentrated at comprehensive stroke centres and facilities of 1,000 or more beds, with roughly twice the odds in the Stroke Belt, and showed no differences by patient demographics or stroke severity — facility resources, not patient factors, predicted use, which the lead author summarised as access depending 'more on where a patient is treated than on their clinical needs'. Two smaller before-and-after studies measured pathway compression with their designs attached and are not outcome evidence: in a hub-and-spoke network (n=43, 28 before and 15 after) the median interval from angiography at the primary centre to door-in at the comprehensive centre fell from 132.5 to 110 minutes, a 22.5-minute reduction (p=0.047), with reductions in overall and neuro-ICU length of stay; and in a transferred thrombectomy cohort at one academic centre (n=55) the median door-to-team-notification interval fell from 40.0 to 25.0 minutes (p=0.01) while the 25-minute door-to-puncture improvement did not reach significance (p=0.15). No randomised trial of this deployment exists, several of these authors hold vendor-affiliated work, and no outcome or mortality benefit is asserted.

Sources: pelzl2026, harveyl2026, hassan2020

Appears on: /domains/cases/viz-lvo-stroke-triage

EmpiricalAfter a ten-day ERISA bench trial in October 2017, the United States District Court for the Northern District of Califor…

After a ten-day ERISA bench trial in October 2017, the United States District Court for the Northern District of California issued 106 pages of Findings of Fact and Conclusions of Law (28 February 2019; public redacted version 5 March 2019) holding that the 2011-2017 editions of United Behavioral Health's Level of Care Guidelines and Coverage Determination Guidelines were significantly and pervasively more restrictive than generally accepted standards of care. The operative 3 February 2026 judgment declares eight specific deviations, among them excessive emphasis on acute crisis stabilization, no effective treatment of co-occurring conditions, failure to err toward a higher level of care where the indicated level is ambiguous, no coverage to maintain function, motivation-based exclusions in the 2014-2017 editions, no child-and-adolescent-specific criteria, an overbroad custodial-care exclusion paired with a narrow active-treatment requirement, and mandatory prerequisites in place of a multidimensional assessment; the editions also omitted the ASAM residential levels 3.1, 3.3 and 3.5. The court found the criteria operated as binding rules rather than guidance: only a physician or doctoral-level psychologist could issue a clinical non-coverage determination, such a reviewer typically spent about thirty minutes talking to the requesting physician and writing up conclusions, every denial letter had to cite the specific guideline relied on, and the company's testimony that reviewers could deviate from the guidelines on clinical judgment was found not credible. No machine-learning system is involved; the guidelines are codified decision criteria applied by human reviewers.

Sources: witv2019, witv2026

Appears on: /domains/cases/wit-ubh-guidelines

EmpiricalThe court's central finding is about who wrote the criteria and under what pressure. In the words of the Findings of Fac…

The court's central finding is about who wrote the criteria and under what pressure. In the words of the Findings of Fact at paragraph 180: 'The Court finds that the financial incentives discussed above have, in fact, infected the Guideline development process. In particular, instead of insulating its Guideline developers from these financial pressures, UBH has placed representatives of its Finance and Affordability Departments in key roles in the Guidelines development process throughout the class period.' Named Finance and Affordability representatives sat as members of the approving committees — the Behavioral Policy and Analytics Committee from 2011 to 2016 and the successor Utilization Management Committee from 2016 — which reviewed and reissued the guidelines at least annually; proposed changes were modelled for their benefit-expense impact before adoption and proposed liberalizations are recorded awaiting a 'green light' from finance. The court also found that the company prepared detailed benefit-expense forecasts and targets, tracked monthly trends, and took action to address benefit expenses exceeding its projections, and that a 2014 internal presentation named continued use of concurrent review to ensure appropriate utilization as the 'Mitigation Strateg[y]' for the 2008 Parity Act's removal of day and visit limits. Because the guidelines were kept uniform across fully-insured plans, where the company bears the benefit-expense risk, and self-funded plans, where it does not, the court held the conflict tainted its decision-making as to both categories.

Sources: witv2019, witv2026

Appears on: /domains/cases/wit-ubh-guidelines

EmpiricalFour states required these determinations to use external professional-society criteria rather than payer-authored ones,…

Four states required these determinations to use external professional-society criteria rather than payer-authored ones, and the court adjudicated violations of all four: Connecticut, where ASAM criteria have been required since 1 October 2013, violated throughout; Illinois, required from 18 August 2011, violated until 1 January 2016; Rhode Island, required from 10 July 2015, violated through the class period; and Texas, under the criteria of its Department of Insurance, violated throughout. The operative judgment also declares that the company's 2013 and 2015 crosswalks told Connecticut regulators that all three ASAM residential levels were included in its admission criteria and that, at the time these statements were made to Connecticut regulators, the company knew them to be false. United Behavioral Health did not appeal this portion of the judgment, and the Ninth Circuit recorded that it therefore remains intact — so the state-mandate rulings survived every appellate cycle unchanged. California later answered legislatively: SB 855 (Stats. 2020 ch. 151, effective 1 January 2021) requires commercial plans to make mental-health and substance-use medical-necessity determinations using the current criteria of the relevant nonprofit clinical specialty association, naming ASAM, LOCUS/CALOCUS and CASII/ECSII, and forbids applying different, additional, conflicting or more restrictive utilization review criteria.

Sources: witv2026, witv2023, californiasb2020

Appears on: /domains/cases/wit-ubh-guidelines

EmpiricalThe class-wide wrongful-denial theory did not survive. In its amended opinion of 22 August 2023 (Wit III, 79 F.4th 1068)…

The class-wide wrongful-denial theory did not survive. In its amended opinion of 22 August 2023 (Wit III, 79 F.4th 1068) the Ninth Circuit affirmed Article III standing and affirmed certification of the three classes for the fiduciary-duty claim, but reversed certification of the denial-of-benefits classes under the Rules Enabling Act; held that the district court erred to the extent it determined that the ERISA plans required the guidelines to be coextensive with generally accepted standards of care; held that reprocessing was not appropriate equitable relief under 29 U.S.C. section 1132(a)(3); and remanded the exhaustion question. When the district court's scope-of-remand order tried to preserve more than the mandate allowed, the Ninth Circuit granted mandamus on 4 September 2024 (No. 24-242) and directed entry of judgment for United Behavioral Health on the denial-of-benefits claim, writing that in its thorough analysis of the spirit of the mandate, the district court lost the letter. On remand the district court held on 5 August 2025 that the fiduciary claim survives Wit III insofar as it rests on the duties of loyalty and due care — entering judgment for the company on the duty-to-follow-plan-terms theory — and that the surviving statutory claim requires no administrative exhaustion, alternatively excused as futile on the trial findings. The roughly 67,000-request reprocessing ordered in November 2020 was stayed on 12 February 2021, held unavailable on appeal, and vacated; no coverage request was ever reprocessed under it. This is a bench record: findings of fact by a judge, no jury and no damages, with relief declaratory and injunctive.

Sources: witv2023, witv2025, witv2026, witv2020, manatt2021, congressionalresearchservice2022

Appears on: /domains/cases/wit-ubh-guidelines

EmpiricalThe operative remedy is the Amended Remedies Order of 3 February 2026 (dkt. 695), which vacated the 3 November 2020 Reme…

The operative remedy is the Amended Remedies Order of 3 February 2026 (dkt. 695), which vacated the 3 November 2020 Remedies Order in its entirety and superseded it. It declares that the company's misconduct in developing and adopting the guidelines was willful and systematic, that the adjudicated editions are irreparably tainted by the company's disloyalty and lack of care, and that the company breached the ERISA duties of loyalty and care under 29 U.S.C. sections 1104(a)(1)(A) and (B) and violated the Connecticut, Illinois, Rhode Island and Texas criteria mandates. It permanently enjoins use of those editions to implement plan terms about generally accepted standards of care, and orders that for five years — through 3 February 2031, with jurisdiction retained — any criteria the company adopts for that purpose shall accurately reflect those standards as established in the court's Findings of Fact and the requirements of any applicable state law. The vacated 2020 order's features are not part of it: there is no special master, no supervised retraining programme, no court-specified external criteria catalogue, and no reprocessing. The ten-year injunction belonged to that vacated order and never took effect; the 2031 endpoint belongs to the new five-year mandate. Attorney-fee litigation was reported ongoing in mid-2026 and appellate review of the 2026 order remained possible as of August 2026, so the injunction posture should be re-verified before republication. Class-scale figures are recorded litigation facts — roughly 67,000 coverage determinations for roughly 50,000 people, because a member can have more than one denial — and the characterization that about half of those people were children or adolescents is plaintiff-side reporting.

Sources: witv2026, witv2020, thekennedyforum2026

Appears on: /domains/cases/wit-ubh-guidelines

EmpiricalAmazon's fulfilment centres time every task an associate performs and feed the result into structured discipline process…

Amazon's fulfilment centres time every task an associate performs and feed the result into structured discipline processes that generate written warnings, final warnings and terminations. A letter from an attorney for Amazon to the National Labor Relations Board dated 4 September 2018, obtained and published in April 2019, described a system in which 'Amazon's system tracks the rates of each individual associate's productivity, and automatically generates any warnings or terminations regarding quality or productivity without input from supervisors', and stated that hundreds of workers had been terminated at a single facility between August 2017 and September 2018: about 300 full-time workers at the Baltimore fulfilment centre over that roughly thirteen-month window, representing roughly ten percent of that site's workforce. Amazon disputed the characterisation on the record the day the documents were published — 'It is absolutely not true that employees are terminated through an automatic system' — adding that it would not dismiss an employee without ensuring they received coaching, that managers can intervene in the process, and that terminations can be appealed; a spokesperson described the same period as about 300 employees of productivity-related turnover at that site. Amazon contested the adjective and not the count. Six years later a U.S. Senate committee majority, working from internal documents produced to it and from two letters by Amazon's outside counsel in 2024, found in its own voice that 'When workers cannot keep up, Amazon uses automated systems to initiate disciplinary procedures. These disciplinary procedures progress in severity and eventually result in termination.' That is the finding of a committee majority, disputed by Amazon to the committee, and no court or regulator has adjudicated whether any individual termination issued without a human decision-maker. The tracker behind the 2019 documents was reported as ADAPT, the Associate Development and Performance Tracker; Amazon has confirmed that name in no source verified for this record, and the process names it put on the record with Congress are Structured Productivity Performance Review and Structured Quality Performance Review.

Sources: lecher2019, businesshumanrightsresourcec2019, cbsnews2019, mittechnologyreview2019, ussenatecommitteeonhealth2024

Appears on: /domains/cases/amazon-adapt-productivity-termination

EmpiricalThe rule is a percentile of the governed population, and both sides of the record describe it that way. Amazon's counsel…

The rule is a percentile of the governed population, and both sides of the record describe it that way. Amazon's counsel told the Senate committee that its speed-related discipline process compares 'each eligible [worker's] performance in a given week to the performance of other employees doing the same work at the same facility', that the slowest five percent may be disciplined, and that as of 2020 the system identified for potential discipline the bottom five percent of performers whose actual rate was 50 percent or less of the expected rate — with the committee recording that it does not know whether that 50 percent threshold remains policy. Eligibility is gated and quantified: speed-related discipline 'applies to only a minority of Amazon [workers] who work at fulfillment centers, specifically Tier 1 (entry level) [workers] who have worked in an eligible process path for at least five hours in a given week and for at least 160 hours over the course of the [worker's] tenure' — roughly sixteen shifts. Speed runs through Structured Productivity Performance Review and quality through Structured Quality Performance Review, driven by counted defects such as scanning an incorrect item, with a third stream running off 'unknown idle time' or time off task; internal charts reviewed by the committee show speed-related write-ups as by far the most common form of discipline, quality second. Amazon's public answer to the California citation describes the same design from the other side: 'The truth is, we don't have fixed quotas. At Amazon, individual performance is evaluated over a long period of time, in relation to how the entire site's team is performing.' Nothing in this arithmetic is machine learning: there is no learned model, no score and no prediction in it.

Sources: ussenatecommitteeonhealth2024, abclosangeles2024

Appears on: /domains/cases/amazon-adapt-productivity-termination

EmpiricalThe documented human step sits at the exemption of accrued time rather than at the decision. Internal Amazon documents p…

The documented human step sits at the exemption of accrued time rather than at the decision. Internal Amazon documents provided to the committee track unknown idle time 'in some instances, down to the second' — one warning cites 97.68 minutes of unknown idle time on a named date — and that accumulator fills from conveyor breakdowns, manager conversations, pallet problems and restroom trips alike. A manager holds a 'seek to understand conversation' about the accrual and may exempt part of it: in the instance the committee published, the manager exempted 14 of 48 minutes for the worker's travel to and from a restroom in a one-million-square-foot warehouse and issued a first written warning for the remaining 34 minutes as a violation of Amazon's Standards of Conduct. The committee found instances of more than ten days between an alleged time infraction and the delivery of the disciplinary consequence, and identified that delay as making it difficult for workers to defend themselves and increasing the likelihood that they are wrongly disciplined. Above the manager sits an escalation authority: a low-level manager reported that management had discretion over whether to terminate workers after a given number of write-ups and would not terminate when headcount was low, but only if Human Resources permitted the deviation from protocol. Throughput is coupled to labour demand as well as to performance — Amazon told the committee that during the peak holiday period 'non-automated warnings, reprimands, write-ups, and improvement plans are paused', and an internal August 2020 chart shows write-ups dipping through the October-to-New-Year peak of 2019 and rising sharply in the first week of January once peak ended. When speed-related write-ups were paused at the start of the pandemic, an internal Amazon team observed that warehouse managers increased their use of other write-up types: behavioural, attendance and safety. Amazon has published no termination rate, no appeal rate, no override rate, no exemption rate and no false-positive rate, and says only that 'the rate of termination is very low.'

Sources: ussenatecommitteeonhealth2024

Appears on: /domains/cases/amazon-adapt-productivity-termination

EmpiricalA second automated termination path runs off a different metric and is where the record shows a human reconciliation ste…

A second automated termination path runs off a different metric and is where the record shows a human reconciliation step being removed. The Senate committee found terminations for 'job abandonment' generated when automated time-tracking failed to account for workers on approved medical leave and registered large negative unpaid-time-off balances. A Human Resources employee described having, before automation, to review a daily report of workers with negative leave balances — sometimes finding workers on leave with hundreds of thousands of hours of negative time — and personally removing them so they would not be flagged for discipline. One worker recovering from a foot injury was terminated by email a week before her scheduled return. The two paths share the timekeeping and the termination surface and do not share a trigger. The live vehicle on this path is a pleading and nothing in it is adjudicated: Lyster v. Amazon.com Services LLC, a class action over the operator's workplace absence practices filed 12 November 2025 in federal court in New York. Amazon says claims that it does not follow federal and state law 'are simply not true' and that its accommodations team reviews each request individually.

Sources: ussenatecommitteeonhealth2024, cbsnews2025

Appears on: /domains/cases/amazon-adapt-productivity-termination

EmpiricalWhether the governed worker can read the ledger that judges them is contested on the record, and the gap between the two…

Whether the governed worker can read the ledger that judges them is contested on the record, and the gap between the two accounts is the case. Amazon's position, stated in response to the California citation, is that 'Employees can - and are encouraged to - review their performance whenever they wish. They can always talk to a manager if they're having trouble finding the information.' The congressional record describes the same channel differently: disciplinary consequences arriving more than ten days after the conduct they cite, and workers running their own parallel timekeeping in response — 'I keep a timer on my watch to keep track of everything', another worker carrying a notebook — because they cannot see or contest the system's own ledger in time. Six state legislatures have since built a statutory version of the same channel, and what they built is a read right rather than a limit on the rule: a written description of each quota and of the discipline attached to it, provided on hire and in the worker's primary language with notice of changes within two business days, plus a right to request the quota description and the most recent 90 days of the worker's own personal work-speed data, answered in New York within 14 calendar days at no cost together with aggregate data for similar workers at the same site. The worker-kept timer or notebook is therefore a record store that exists because the primary one is not legible to its subject in time to contest an entry.

Sources: ussenatecommitteeonhealth2024, abclosangeles2024, assemblybill2021, newyorkstatedepartmentoflabo2023

Appears on: /domains/cases/amazon-adapt-productivity-termination

EmpiricalTwo independent state labour regulators have cited this design, and both cited the notice rather than the threshold. The…

Two independent state labour regulators have cited this design, and both cited the notice rather than the threshold. The California Labor Commissioner's Office cited Amazon.com Services LLC $5,901,700 for 59,017 violations of the state Warehouse Quotas law at the Moreno Valley and Redlands fulfilment centres, the violations occurring between 20 October 2023 and 9 March 2024 and penalised at $100 per violation under Labor Code section 2699(f); the inspection opened on 22 September 2022 and the Warehouse Worker Resource Center assisted. Labor Commissioner Lilia Garcia-Brower stated that 'The peer-to-peer system that Amazon was using in these two warehouses is exactly the kind of system that the Warehouse Quotas law was put in place to prevent. Undisclosed quotas expose workers to increased pressure to work faster and can lead to higher injury rates and other violations by forcing workers to skip breaks.' Amazon appealed: 'We disagree with the allegations made in the citations and have appealed.' Minnesota reached the same conclusion under its own statute: after an October 2023 inspection at the Shakopee facility, Minnesota OSHA issued two serious citations in April 2024 totalling $10,500, one of them for the violation that 'warehouse employees who were expected to meet a quota of selecting, stowing and packaging products were not provided a written copy of the quota before they were expected to meet the quota.' Amazon contested the citations. Both are agency determinations rather than judicial findings and neither was resolved in any source located for this record; the per-facility split of the California total that circulates in secondary coverage is stated by no primary source and is not carried here.

Sources: californiadepartmentofindust2024, minnesotadepartmentoflaboran2024, abclosangeles2024

Appears on: /domains/cases/amazon-adapt-productivity-termination

EmpiricalThe correction channel the law actually built for this decision surface is a disclosure-and-data right, not a limit on t…

The correction channel the law actually built for this decision surface is a disclosure-and-data right, not a limit on the rule. California Labor Code sections 2100 to 2112, effective 1 January 2022, define a quota as a work standard under which an employee is required to perform at a specified productivity speed or handle a quantified amount of material within a defined time period and under which the employee may suffer an adverse employment action for failing to meet it; require a written description of each quota on hire; bar adverse employment action for failing to meet an undisclosed quota; give employees the right to request the written quota description and a copy of the most recent 90 days of their own personal work-speed data; and create a rebuttable presumption of retaliation for adverse action within 90 days of such a request or complaint. New York's Warehouse Worker Protection Act adds a fourteen-calendar-day no-cost response deadline, notice of quota changes within two business days, provision in the worker's primary language, and a right to aggregate speed data for similar workers at the same site. Six states now carry such statutes — California (2021, effective January 2022), New York (19 June 2023), Minnesota (1 July 2023), Washington (31 May 2024), Oregon (1 January 2025) and Connecticut (signed March 2026, effective 1 July 2026 with notices to current employees due by 1 August 2026). The Senate committee majority's own prescription aims one step further in, and is a proposal rather than law: the No Robots Bosses Act, which would 'prevent employers from exclusively relying on automated systems to make decisions about disciplining or firing workers' and 'require employers using automated decision-making systems to tell workers how the system works and how workers can appeal system decisions', recommended on the committee's finding that 'Amazon subjects workers to discipline based on automated systems that are prone to errors, including firing workers who are on medical leave'; and the Stop Spying Bosses Act, recommended on the finding that Amazon 'closely tracks workers' movements and actions throughout the workday, and uses this information to make disciplinary decisions.'

Sources: assemblybill2021, newyorkstatedepartmentoflabo2023, littlermendelson2026, ussenatecommitteeonhealth2024

Appears on: /domains/cases/amazon-adapt-productivity-termination

EmpiricalThe litigation record on this mechanism is thinner than the volume of coverage suggests, and every item in it stops shor…

The litigation record on this mechanism is thinner than the volume of coverage suggests, and every item in it stops short of the question. The only merits ruling to date on the quota machinery went on pleading specificity: in January 2023 a U.S. magistrate judge in the Northern District of California dismissed a proposed class action alleging that hourly quotas of roughly 150 to 300 items discriminate against older workers, holding the allegations too vague and writing that 'simply because physical strength declines with age does not automatically mean that older workers are more likely to get injured or fail to keep up with the quotas.' That is reasoning about a discrimination theory and about the sufficiency of a pleading; it is not a holding that the quotas do not exist. The live federal vehicle, filed 12 November 2025 in the Southern District of New York, pleads the adjacent automated absence-control path and is wholly unadjudicated.

Sources: wiessner2023, cbsnews2025

Appears on: /domains/cases/amazon-adapt-productivity-termination

EmpiricalThe governed side of this deployment has been quantified once, in a weighted opt-in national survey rather than a probab…

The governed side of this deployment has been quantified once, in a weighted opt-in national survey rather than a probability sample, and it is labelled as such wherever it is used. The Center for Urban Economic Development at the University of Illinois Chicago surveyed 1,484 frontline Amazon warehouse workers between April and August 2023, drawing respondents from 42 states and 451 facilities, recruited through targeted advertising with CAPTCHA screening, a fake-facility-code trap and fraud-cluster removal, and weighted to Amazon's 2021 workforce demographics. In it, 77 percent said the technology can tell whether they are actively engaged in their work always or most of the time, against 47 percent of warehouse workers industry-wide; 72 percent said how fast they work is measured in detail by company technology, against 58 percent; and 58 percent said their pace is always or most of the time ranked and compared with the pace of their coworkers, against 46 percent. Asked what the electronic monitoring is used for, 45 percent said it is mainly used to control or discipline workers and 36 percent said it is mainly used to develop workers' skills and abilities, with the control-or-discipline share rising to 52 percent among workers of more than three years. 45 percent said keeping up with the pace of work is hard (47 percent at fulfilment centres against 31 percent at sortation centres), 41 percent feel pressure to work faster always or most of the time, and 53 percent report always or most of the time feeling watched or monitored.

Sources: centerforurbaneconomicdevelo2023

Appears on: /domains/cases/amazon-adapt-productivity-termination

EmpiricalThe largest governance instrument attached to these facilities does not reach this decision surface, and that is what se…

The largest governance instrument attached to these facilities does not reach this decision surface, and that is what separates this case from the pace-and-injury case that shares the same buildings. On 19 December 2024 the U.S. Department of Labor announced a corporate-wide settlement with Amazon requiring Site Ergonomics Leads who review corporate risk assessments and prepare annually updated site-level assessments, together with multiple employee channels — including anonymous ones — for raising ergonomic concerns, across fulfilment centres, sortation centres and delivery stations in federal OSHA jurisdiction, for a $145,000 penalty. The settlement contains no quota provision, no pace-setting provision and no discipline provision. The instruments that do reach this decision surface are the state warehouse-quota-notice statutes and the two contested citations issued under them, whose remedy is written disclosure of the rule and access to the worker's own speed data. Nothing in the record adjudicates whether the threshold itself is correct.

Sources: usdepartmentoflabor2024

Appears on: /domains/cases/amazon-adapt-productivity-termination

EmpiricalAmazon Flex, a first-party last-mile delivery programme launched in September 2015, rates its United States contract dri…

Amazon Flex, a first-party last-mile delivery programme launched in September 2015, rates its United States contract driver fleet into four standing tiers — Fantastic, Great, Fair and At Risk — computed from arrival punctuality at delivery stations, completion of routes inside the reserved block window, compliance with customer special requests, and delivery-quality signals. A Bloomberg investigation published on 28 June 2021, read here through the syndication carrying its full text, reports that algorithms scan incoming performance data and decide which drivers get more routes and which are deactivated, that human feedback is rare, and that terminations arrive by automated email; it interviewed fifteen drivers, four of whom said they were wrongly terminated, together with former Amazon managers and a former engineer, and documents deactivations following circumstances the input set cannot represent — locked apartment gates on predawn routes, malfunctioning lockers, hour-long waits at understaffed stations, a nail in a tire, snowbound rural roads and failed selfie identity checks. Former insiders told the investigation that the programme's benefits far outweigh the collateral damage and that the company decided it was cheaper to trust the algorithms than to pay people to investigate mistaken firings. Amazon disputes the characterisation: its spokesperson called the driver accounts anecdotal and unrepresentative and said the company has invested heavily in technology and resources to provide drivers visibility into their standing and eligibility to continue delivering, and investigates all driver appeals. No court or regulator has ever ruled on how standing is computed or on whether any particular deactivation was substantively correct, and no scoring internals, error rate, termination rate or reversal rate has been published by anyone. Scale figures are the operator's own and count downloads rather than active drivers — approximately 4 million globally and 2.9 million in the United States, with more than 660,000 in one quarter of 2021, up about 21 percent year over year; the only hard official count in this record is the 140,128 individual drivers a federal regulator paid tip refunds to for a single 2016-to-2019 window.

Sources: soper2021, thespokesmanreview2021, dent2021, aiincidentdatabase2021, federaltradecommission2021

Appears on: /domains/cases/amazon-flex-driver-rating

EmpiricalThe correction channel around the automated decision is slow, templated and priced. As reported in June 2021, a deactiva…

The correction channel around the automated decision is slow, templated and priced. As reported in June 2021, a deactivated Flex driver has ten days to appeal by email; first replies arrive the next day, read as machine-generated and are typically generic rather than specific to the incident, signed with a support agent's first or full name; a promised six-day review is often exceeded, and in the two cases the investigation follows end to end the final answer lands on day 26 and on day 11, with no pay in the interim. The only escalation beyond email under the platform terms then in force carried a $200 arbitration filing fee — verbatim, drivers 'pay $200 to take their dispute to arbitration, but few do, seeing it as a waste of time and money' — against roughly $80 of net pay on a typical four-hour block, a fee legal commentary describes as an unrealistic prospect for workers barely making minimum wage. Drivers also report that standing takes months to recover from delays outside their control while standing feeds route and block allocation. The figure is date-stamped: Amazon's widely reported 2021 removal of mandatory arbitration applied to consumer terms rather than to the Flex driver terms, and the present fee level was not independently verified. One consolidating source overstates the position by saying drivers had no opportunity to contest or appeal; the primary record documents a channel that is automated, slow and fee-gated rather than absent, and no appeal-outcome or reversal rate exists from any side.

Sources: soper2021, thespokesmanreview2021, dent2021, bajgiran2022, aiincidentdatabase2021

Appears on: /domains/cases/amazon-flex-driver-rating

EmpiricalSince 1 January 2025 the deactivation channel itself is regulated in one jurisdiction. Seattle's App-Based Worker Deacti…

Since 1 January 2025 the deactivation channel itself is regulated in one jurisdiction. Seattle's App-Based Worker Deactivation Rights Ordinance (Seattle Municipal Code 8.40, Ordinance 126878; administrative rules SHRR Chapter 260 effective 24 June 2025) covers network companies with 250 or more app-based workers worldwide, Amazon Flex among them, and requires a published deactivation policy reasonably related to safe and efficient operations; fourteen days' notice before deactivation except for egregious misconduct or where law requires immediacy; a written statement of the reasons and the specific incidents together with all records relied on and considered; investigation of the alleged violation to a more-likely-than-not standard before deactivating; penalties applied consistently and in proportion to the violation with the circumstances of the work considered; and an internal challenge procedure invocable within ninety days, with a private right of action after the company's response or fourteen days after the challenge. A worker is covered at twenty-five percent of completed or cancelled-with-cause offers in the prior 180 days performed in the city, or by a single incident there, and the Office of Labor Standards enforces procedural compliance only until 1 June 2027, not whether a deactivation was substantively warranted. That office's published resolved-investigations record for October to December 2025 shows two informal resolutions under this ordinance against Amazon Logistics, Inc. doing business as Amazon Flex, each returning $1,245.70 to one worker and requiring Amazon Flex to restart that worker's deactivation process, alongside equivalent resolutions the same quarter against three other network companies. In Uber Technologies, Inc. v. City of Seattle, Nos. 25-228 and 25-231 (9th Cir. 4 March 2026), a panel of Graber, Clifton and Bennett affirmed the denial of a preliminary injunction sought by Uber and Instacart, holding that the ordinance regulates nonexpressive conduct — the unwarranted deactivation of worker accounts — that any compelled disclosure would in any event be commercial speech surviving Zauderer review, and that 'reasonably related to safe and efficient operations' is not unconstitutionally vague, with Judge Bennett dissenting in part. That is a preliminary-injunction affirmance rather than a final merits judgment, and Amazon was not a party.

Sources: cityofseattle2025, seattleofficeoflaborstandard2025, ubertechnologies2026, komonews2025

Appears on: /domains/cases/amazon-flex-driver-rating

EmpiricalThe independent-contractor classification underneath the deactivation regime is contested and has repeatedly gone agains…

The independent-contractor classification underneath the deactivation regime is contested and has repeatedly gone against the operator in adjacent forums, while no forum has reached the rating rule. A Wisconsin Department of Workforce Development audit of more than 1,000 Flex drivers covering 2016 to 2018 found the vast majority employees for unemployment-insurance purposes with an assessment of about $205,000; that determination was upheld by the state Court of Appeals in 2023 and left standing when the Wisconsin Supreme Court dismissed Amazon's appeal as improvidently granted on 26 March 2024. The New Jersey Department of Labor and Workforce Development sued Amazon on 20 October 2025 in Essex County Superior Court alleging Flex misclassification since at least 2017 with millions of dollars in annual losses to state benefit funds; that suit is pending and entirely unadjudicated. In Rittmann v. Amazon.com, Inc., No. 19-35381 (9th Cir. 19 August 2020), the court held that Flex delivery workers are transportation workers engaged in interstate commerce and exempt from the Federal Arbitration Act under 9 U.S.C. § 1 even without crossing state lines, and that the arbitration provision with its choice-of-statute clause was unenforceable under federal or Washington law, affirming the denial of the motion to compel arbitration of the wage claims. That holding concerns the WAGE-claims clause; the deactivation-dispute channel continued to route through individual arbitration as reported in 2021, and the two facts are distinct.

Sources: pbswisconsinassociatedpress2024, rittmannv2020

Appears on: /domains/cases/amazon-flex-driver-rating

EmpiricalWhat the scorer measures, and what it has no field for, are both documented. Human Rights Watch's May 2025 cross-platfor…

What the scorer measures, and what it has no field for, are both documented. Human Rights Watch's May 2025 cross-platform study — 95 workers interviewed across 13 states including Flex drivers, plus a 127-worker Texas survey — records that Flex times a worker from arrival at the warehouse through completion of the delivery and detects whether the worker is driving, walking or running, while noting that Flex is the one studied platform paying a posted flat hourly block rate of $18 to $25 rather than an opaque per-job wage algorithm; the study's $5.12 hourly after-expenses median is a cross-platform Texas figure and is not a Flex figure. Mandatory selfie identity checks are part of the same surface, and failed checks are among the documented contributors to adverse action. A Harvard Law School labour-law commentary of January 2022 additionally characterises the reporting record as including monitoring of seatbelt use, acceleration and screen touches while driving; that characterisation mixes Flex with Amazon's separate Delivery Service Partner van programme, whose inward-facing camera telematics does not apply to Flex drivers, who use their own vehicles. What none of these inputs can represent is the reason a stop went wrong — a locked building, a closed office, a station queue, weather, road conditions or a vehicle failure — which the investigation records arriving in the measurement as the driver's own shortfall.

Sources: humanrightswatch2025, bajgiran2022, soper2021, thespokesmanreview2021

Appears on: /domains/cases/amazon-flex-driver-rating

EmpiricalAon Consulting, Inc., the human-capital arm of Aon plc, designs, markets and administers a suite of pre-hire assessments…

Aon Consulting, Inc., the human-capital arm of Aon plc, designs, markets and administers a suite of pre-hire assessments sold into other companies' hiring pipelines. Three of them are at issue: ADEPT-15, a computer-adaptive forced-choice personality test scoring fifteen constructs from item pairs matched for social desirability out of a bank of more than a thousand statements, with no skip option and no agree-with-both option, whose outputs are construct scores, job-fit scores and generated interview scripts for the employer; gridChallenge, a gamified working-memory test of nine items each pairing a timed target-icon memorisation task with a symmetry or shape-judgment distractor task; and vidAssess-AI, an asynchronous video product that transcribes a candidate's spoken answers with a commercial speech-to-text service and scores phrases in that transcript as far-extreme positive or negative indicators of the same ADEPT-15 constructs. On Aon's own marketing, quoted in the ACLU's complaint, the company administers more than 30 million assessments a year in over 40 languages across 90 countries, and third-party market data quoted in the same paragraph places it at the second-largest share of the global pre-hire assessment market. The assessments are marketed to screen out applicants prior to more resource-intensive hurdles such as resume screening and interviews, and Aon's own case-study marketing is cited for one client that screened out 62 percent of applicants on the assessment alone; for the population removed at that step there is no human review, and the assessment is the first and the final decision point. Every scale and screen-out figure here is Aon's own marketing and none is independently audited.

Sources: americancivillibertiesunionf2024, americancivillibertiesunion2024, hrdive2024

Appears on: /domains/cases/aon-assessment-suite

EmpiricalThe figures at the centre of this record are Aon's own, published in Aon's technical documentation and quoted at table a…

The figures at the centre of this record are Aon's own, published in Aon's technical documentation and quoted at table and page level in the ACLU's complaint to the Federal Trade Commission; the numbers are Aon's and the reading of them is the ACLU's. On ADEPT-15, the reliability table shows ten of fifteen constructs producing two-week test-retest coefficients below the 0.7 floor Aon's own documents tell buyers to demand, the lowest being 0.44 on 'Awareness', which the complaint identifies as the construct with the greatest overlap with autism diagnostic screeners. On gridChallenge, the gridChallenge and GAME test documentation at Table 43 reports Black assessment-takers scoring on average 0.48 standard deviations below white assessment-takers, with two-or-more-ethnicities at 0.38, Hispanic or Latino at 0.37 and Asian at 0.35; Aon's documentation for sibling cognitive tools reports Black test-takers averaging 0.86 standard deviations lower on scales clues and 1.21 lower on switchChallenge, which are 'large' effects under the Cohen's-d heuristic Aon itself applies in the same documents. An average score difference is NOT an adjudicated adverse impact: the complaint concedes at paragraph 62 that actual selection impact depends on the cut-off scores an employer sets and on how the assessment is combined with other tools, which Aon does not control. No error rate, accuracy figure or independent evaluation exists for any of the three products, from any source. The manuals these figures come from were removed from public access after the EEOC charges were filed and could not be independently re-verified; the ACLU states it knows of no other public source and cites archived copies. Aon disputes the complaint and states that its assessment solutions follow industry best practices as well as the Equal Employment Opportunity Commission, legal and professional guidelines and are intended to be fair to everyone.

Sources: americancivillibertiesunionf2024, hrdive2024

Appears on: /domains/cases/aon-assessment-suite

EmpiricalThe ACLU's disability theory runs through construct overlap rather than through any measured screen-out. Statements the …

The ACLU's disability theory runs through construct overlap rather than through any measured screen-out. Statements the EEOC charging party recalled encountering on ADEPT-15 — among them 'I tend to avoid large groups of people' and 'I have difficulty determining how someone feels by looking at their face' — are set in the complaint's Table 1 beside near-parallel items on the clinical autism screeners AQ and RAADS-R; the statement wordings are the charging party's recollections and precise wording may differ. Aon's own interpretation guides flag extreme-end scorers as potential problems for employers. Meta-analyses cited in the complaint show autistic people scoring significantly lower on working-memory measures of the kind gridChallenge uses. And the design record is the load-bearing part: across a 166-page ADEPT-15 technical documentation, the only disability-bias measure the complaint identifies is an outside-attorney review of item wording that flagged physical-disability language such as 'seeing', 'speaking' and 'hearing' — with no clinical or neurodiversity sensitivity review and no study on disabled populations anywhere in the programme. What the complaint alleges is a high RISK of screening out autistic and disabled applicants, together with one charging party's individual account. No adjudicated or statistically demonstrated screen-out of disabled applicants exists in this record. In August 2025 Bloomberg Law reported the EEOC charge as filed and unresolved and recorded that an Aon spokesperson declined to comment.

Sources: americancivillibertiesunionf2024, bloomberglaw2025

Appears on: /domains/cases/aon-assessment-suite

EmpiricalTwo federal forums were opened on one product suite and neither has publicly moved. In late 2023 the ACLU, with co-couns…

Two federal forums were opened on one product suite and neither has publicly moved. In late 2023 the ACLU, with co-counsel Winston Cooks, LLC, filed class-wide charges of discrimination with the Equal Employment Opportunity Commission under the Americans with Disabilities Act and Title VII against Aon AND a client employer, on behalf of a biracial Black and white autistic job applicant with mental-health disabilities; the employer is publicly identified only as a mid-sized company headquartered in the United States. Those charges cover ADEPT-15 and gridChallenge. On 30 May 2024 the ACLU filed a complaint and request for investigation with the Federal Trade Commission against Aon Consulting, Inc., alleging that marketing the assessments as bias-free, fair, having no adverse impact and able to improve diversity is a deceptive act or practice under Section 5 of the FTC Act, and that selling assessments alleged to discriminate while failing to assess disability harms is an unfair practice injuring both workers and employers; that complaint covers all three products including vidAssess-AI. The relief requested asks the Commission to open an investigation, enjoin the claims, require truthful risk information, and require Aon to pause sale or administration of the assessments until discrimination is eliminated or discontinue products where it cannot be. Management-side analysis of the filings records the procedural register: an advocacy complaint to the Commission initiates agency review and triggers no automatic enforcement, and an EEOC charge is confidential while pending and is commonly a precursor to private litigation. The enforcement environment contracted during the pendency: in January 2025 the EEOC removed its artificial-intelligence hiring technical-assistance guidance following the administration change, and an executive order of 23 April 2025 directed federal agencies to deprioritize disparate-impact liability enforcement, the theory underlying the race claims; legal-press analysis in March 2025 listed the Aon matters among pending ones with no indication of closure. As of 28 August 2026 no public FTC action, no public EEOC determination, settlement or follow-on lawsuit, and no product withdrawal has been located. Because charge proceedings are confidential and the Commission never docketed a public matter, that silence establishes nothing in either direction: it means no PUBLIC action, and an unannounced staff inquiry can be neither ruled in nor ruled out.

Sources: americancivillibertiesunion2024, americancivillibertiesunionf2024, fisherphillips2024, bloomberglaw2025a

Appears on: /domains/cases/aon-assessment-suite

EmpiricalThe contested link in this record is between two sets of documents the same company wrote. The ACLU's complaint quotes A…

The contested link in this record is between two sets of documents the same company wrote. The ACLU's complaint quotes Aon's marketing as describing the assessments as 'fair[] for all: no adverse impact', 'scientifically proven not to have bias', 'culturally-agnostic', and as excluding prejudices, and alleges that selling on those sentences while Aon's own technical documentation reported the group score differences above is a deceptive act under Section 5. Aon responded through trade press in June 2024 that its assessment solutions 'follow industry best practices as well as the Equal Employment Opportunity Commission, legal and professional guidelines' and are 'intended to be fair to everyone'; that is a vendor statement and no adjudication of any of it exists. The record contains one documented act by Aon after the EEOC charges were filed, and it went to evidence rather than to claims: the public link to the assessment technical documentation on Aon's Norwegian support portal was disabled, and the ACLU states it knows of no other public source and cites archived copies. The marketing did not change. When Aon's pre-hire talent assessment page was fetched on 28 August 2026 it advertised 'Fair and unbiased assessments' and asserted that 'Assessment tools that are scientifically proven not to have bias are essential' — twenty-seven months after those claims were formally alleged to be deceptive, with all three products still on sale. That page is recorded as a vendor claim, its content can change, and the fetch date is the citation of record.

Sources: americancivillibertiesunionf2024, aonplc2026, hrdive2024

Appears on: /domains/cases/aon-assessment-suite

EmpiricalTwo documented properties of this deployment decide what an assessed person can do about a score. First, the position of…

Two documented properties of this deployment decide what an assessed person can do about a score. First, the position of the instrument: the assessments are marketed to screen out applicants prior to more resource-intensive hurdles, and for the population screened out there is no human review at all — the assessment is the first and the final decision point, and Aon's own case-study marketing is cited for one client that screened out 62 percent of applicants on the assessment alone. Second, what the candidate is told before responding: Aon's test-taker guide states that 'there is no possibility of bias — for or against specific candidates — either during the test or during the evaluation of answers', and discusses accommodations only for physical disabilities. The ACLU argues that a candidate who reads that does not make the accommodation request that would generate the only disability-relevant signal anywhere in the system, and that the absence of the signal then reads as evidence that nothing was needed; invoking the request costs the candidate a disclosure of disability. The quoted text is Aon's and the alleged suppressing effect is the ACLU's. The employer, not Aon, holds the decision that converts a score into an outcome: which constructs are used, where the cut-off falls, and what the assessment is combined with. No cut-off or vetting practice for any named client appears in this record, and the client employer named as an EEOC co-respondent is publicly unnamed.

Sources: americancivillibertiesunionf2024, americancivillibertiesunion2024

Appears on: /domains/cases/aon-assessment-suite

EmpiricalCheckr, Inc. is a consumer reporting agency founded in 2014 whose automated background-check platform sits one company u…

Checkr, Inc. is a consumer reporting agency founded in 2014 whose automated background-check platform sits one company upstream of the gig platforms that buy from it: it retrieves criminal items from county, state and national sources, matches them to a named applicant, normalises the charge data, applies each platform customer's own eligibility matrix, and furnishes a report, after which the platform executes its own access decision. Every scale, speed and automation figure for it is the vendor's own marketing, fetched from its product pages on 28 August 2026 and independently audited by nobody: 90 percent of gig economy background checks run on Checkr; 140,000-plus businesses already run on Checkr; 260 million-plus real identities mapped, covering 96 percent of US adults; instant criminal checks that screen users in under a second; 89 percent of national criminal record checks complete within one hour; Checkr AI normalises criminal charge data; automated adjudication tools reduce manual workflows by up to 80 percent; and continuous checks are offered as a standing-monitoring product. Those claims are self-serving in both directions — they are capability marketing and they are also the clearest available statement of how unmanned the generation path is. Lyft, Hyer and GigSmart are named as gig clients on the vendor's own pages; the Uber relationship is documented litigation-side, where a defence-bar analysis records Uber moving its New York City driver checks to Checkr in mid-2017.

Sources: checkr2026, huntonandrewskurthllp2021

Appears on: /domains/cases/checkr-gig-background-screening

EmpiricalThe federal consumer regulator wrote down what the Fair Credit Reporting Act requires of automated background screeners,…

The federal consumer regulator wrote down what the Fair Credit Reporting Act requires of automated background screeners, and its statements are about the industry rather than about any one company. In an advisory opinion of 4 November 2021 (Fair Credit Reporting; Name-Only Matching Procedures, 86 FR 62468) the Consumer Financial Protection Bureau held that name-only matching does not assure maximum possible accuracy under 15 U.S.C. 1681e(b), stated that when background screening companies and their algorithms carelessly assign a false identity to applicants for jobs and housing they are breaking the law, warned that the risk of mistaken identities from name-only matching is likely to be greater among Hispanic, Black and Asian communities because there is less surname diversity in those populations, and observed that because of the sheer scale of background screening activity even ostensibly low error rates can harm significant numbers of consumers. In two companion advisory opinions of 11 January 2024 (Background Screening, 89 FR 4171; File Disclosure, 89 FR 4167) it stated that screeners must prevent the reporting of expunged, sealed or legally restricted records, ensure dispositions accompany any reported arrest or charge, prevent duplicative reporting and honour per-item obsolescence windows such as the seven-year bar on non-conviction arrests, and that a consumer's file disclosure must include both the originating sources and any intermediary or vendor sources. An independent legal-bar reading of the January 2024 opinions records the same four accuracy-procedure requirements and the same file-disclosure source rule. No error rate, accuracy figure or per-cohort disparity measurement for Checkr's pipeline has ever been published by anyone.

Sources: consumerfinancialprotectionb2021, consumerfinancialprotectionb2024a, ballardspahrllp2024

Appears on: /domains/cases/checkr-gig-background-screening

EmpiricalThe correction channel that reaches this deployment is statutory rather than discretionary, and the record documents it …

The correction channel that reaches this deployment is statutory rather than discretionary, and the record documents it being outrun. Under 15 U.S.C. 1681i a consumer reporting agency must reinvestigate a disputed item and correct the file, within 30 days and 45 with an extension; 15 U.S.C. 1681k imposes a currency-or-contemporaneous-notice duty on public-record items reported for employment purposes; and 15 U.S.C. 1681b(b)(1) bars furnishing an employment-purpose report at all unless the user certifies compliance with the notice and pre-adverse-action duties. Because the report is assembled and held at one upstream agency, a dispute that succeeds corrects the record for every platform that reads it. Against that clock, the 2021 class complaint pleads that Checkr furnished Uber the report and that Uber deactivated Job Golightly one day later without any notice, process or communication. The same complaint records the recurrence conditions: approximately 80,000 TLC-licensed rideshare vehicles in New York City, the majority on Uber, and an Uber policy of re-checking current drivers at least every two years through Checkr. The vendor separately markets continuous checks, a standing-monitoring product that re-furnishes on new record events. No dispute volume, reinvestigation resolution rate or reversal rate for this vendor appears in any public source.

Sources: classactioncomplaint2021, consumerfinancialprotectionb2021, checkr2026

Appears on: /domains/cases/checkr-gig-background-screening

EmpiricalOn 12 May 2025 the Consumer Financial Protection Bureau withdrew all four of the Fair Credit Reporting Act advisory opin…

On 12 May 2025 the Consumer Financial Protection Bureau withdrew all four of the Fair Credit Reporting Act advisory opinions covering this deployment class — Name-Only Matching (86 FR 62468), Permissible Purposes (87 FR 41243), Background Screening (89 FR 4171) and File Disclosure (89 FR 4167) — inside a mass withdrawal of 67 guidance documents, confirmed on the Bureau's own withdrawn-guidance list. The statutes those opinions interpreted were untouched: 15 U.S.C. 1681e(b), 1681i, 1681k and 1681b(b) all remain in force and privately enforceable, and courts may still reach the same results. What was withdrawn is the regulator's stated reading of them, not the duties themselves, and no source in this record describes the accuracy duty as repealed or ended. The withdrawal is a documented contraction of the interpretive layer over a deployment class that continued operating unchanged.

Sources: consumerfinancialprotectionb2025, consumerfinancialprotectionb2024a

Appears on: /domains/cases/checkr-gig-background-screening

EmpiricalTwo class vehicles have reached this pipeline and neither has been adjudicated on the merits. Golightly v. Uber Technolo…

Two class vehicles have reached this pipeline and neither has been adjudicated on the merits. Golightly v. Uber Technologies, Inc. and Checkr, Inc., No. 1:21-cv-03005 (S.D.N.Y., filed 8 April 2021), pleads five counts of which four are against Uber — the New York City Fair Chance Act, disparate racial impact, and the federal and New York consumer reporting statutes — and one is against Checkr: furnishing employment-purpose reports without obtaining the user's compliance certification under 15 U.S.C. 1681b(b)(1). It alleges that Job Golightly, a Black Bronx driver who had driven for Uber since 2014 averaging about $1,500 a week, was deactivated on or around 28 August 2020 one day after Checkr furnished his report, over a single 2013 Virginia speeding offence at 22 miles per hour over the limit — a misdemeanour under Virginia law and a civil infraction under New York law — with no pre-adverse-action notice, no copy of the report, no Article 23-A individualised analysis and no three-business-day window, and that he learned the reason months later. No misreporting by Checkr is pleaded in that case; the record was accurate as pleaded and the failures alleged are process failures. The complaint estimates classes of at least 1,000 and at least 300 since 11 January 2020, and records, citing research from The New School's Center for New York City Affairs, that 87 percent of New York City transportation independent contractors are persons of colour, 81 percent lack a college degree and 90 percent are foreign-born. On 21 December 2022 Judge Lewis J. Liman granted the motion to compel individual arbitration and stayed the claims; the order itself could not be read from the environment that verified this record, its specific treatment of the claims against Checkr is not asserted, and no public resolution has been located as of August 2026. Davis v. Checkr, Inc., No. 0:26-cv-60088 (S.D. Fla., filed 26 January 2026), pleads a different mechanism: a putative nationwide class under 15 U.S.C. 1681e(b) alleging that a Checkr report included criminal case records plainly unassociated with the consumer and that the company knowingly and willfully maintains deficient procedures because reporting more information is more profitable, covering consumers who within two years received a Checkr report that incorrectly included criminal case records belonging to someone else. It is pending and pre-certification. All of the above are allegations.

Sources: classactioncomplaint2021, mobilizationforjustice2021, huntonandrewskurthllp2021, opinionandorder2022, topclassactions2026

Appears on: /domains/cases/checkr-gig-background-screening

EmpiricalFederal enforcement in neighbouring screening markets sets the posture the Fair Credit Reporting Act establishes for thi…

Federal enforcement in neighbouring screening markets sets the posture the Fair Credit Reporting Act establishes for this class of automation, and none of it is an action against Checkr. On 12 October 2023 the Federal Trade Commission and the Consumer Financial Protection Bureau settled with TransUnion's rental-screening arm for $15 million — $11 million in redress and a $4 million penalty — over duplicate eviction entries, misreported outcomes, sealed records that were not removed, and undisclosed third-party sources. On 11 September 2023 the Commission took $5.8 million from TruthFinder and Instant Checkmate over background reports marketed as most accurate that were assembled from sources disclaiming accuracy, with a Flag as Inaccurate button that triggered no investigation; the Commission's position there was that report assemblers marketing for employment and tenant screening are consumer reporting agencies bound by the maximum-possible-accuracy and permissible-purpose duties. No public Commission or Bureau enforcement action against Checkr itself was located as of 28 August 2026. The one completed money resolution in this record is platform-side: Aguilera v. Uber Technologies, Inc. d/b/a Uber Eats, No. 509275/2023 (N.Y. Sup. Ct., Kings County), settled for $3.35 million over allegations that Uber Eats used a flawed criminal background check process that unfairly denied individuals access to the platform and failed to provide the legally required disclosures, covering New York City couriers denied between 24 October 2015 and 28 July 2021, split into serious-conviction and less-job-related groups, with automatic payments and a final approval hearing on 8 February 2024. Checkr was not a defendant there and no admission of liability was located.

Sources: ftccfpb2023, federaltradecommission2023, topclassactions2023

Appears on: /domains/cases/checkr-gig-background-screening

EmpiricalAround January 2021 a Milton, Massachusetts resident applied for a CVS supply chain position, sat a HireVue video interv…

Around January 2021 a Milton, Massachusetts resident applied for a CVS supply chain position, sat a HireVue video interview and was not hired. As recited in the court's published opinion from the amended complaint, the interview asked integrity-framed questions — what integrity means to the applicant, and a time the applicant acted with integrity — and HireVue uploaded the recordings to Affectiva, a Boston affect-analysis firm spun out of the MIT Media Lab, whose AI 'analyzes candidates' facial expressions, eye contact, voice intonation, and inflection' to draw conclusions about the applicant's degree of cultural fit; HireVue then provided CVS with employability scores. The amended complaint lists the affect features as smiles, surprise, contempt, disgust and smirks, and describes the score as rating traits including conscientiousness and responsibility and an innate sense of integrity and honor. Every mechanical element of that description is an allegation accepted as true for the purpose of a motion to dismiss; none of it was ever adjudicated, and no error rate, accuracy figure or independent evaluation of this screen exists from any source.

Sources: bakerv2024, hrdivecrist2024, thebostonglobejohnston2023

Appears on: /domains/cases/cvs-hirevue-integrity-screen

EmpiricalMassachusetts General Laws chapter 149, section 19B has banned lie detector tests in employment since 1959, and a 1985 a…

Massachusetts General Laws chapter 149, section 19B has banned lie detector tests in employment since 1959, and a 1985 amendment added both a private civil action and a mandatory notice. The statute defines the banned instrument by PURPORTED function: it covers 'any test utilizing a polygraph or any other device, mechanism, instrument or written examination' used 'for the purpose of purporting to assist in or enable the detection of deception, the verification of truthfulness, or the rendering of a diagnostic opinion regarding the honesty of an individual.' Every Massachusetts employment application must carry, in clearly legible print, one sentence: 'It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.' Remedies run to a civil action within three years by any person aggrieved, not less than 500 dollars in damages per violation, treble damages for lost wages or benefits, costs and reasonable attorney fees, with criminal penalties of 300 to 1,000 dollars for a first violation and up to 1,500 dollars and 90 days thereafter. The notice provision then went roughly forty years without being enforced, until filings began in 2023. The 500-dollar figure is a statutory minimum, not a payment made to anyone in this case.

Sources: massachusettsgenerallaws1959, agencychecklists2025, morganlewislawflashengelman2025

Appears on: /domains/cases/cvs-hirevue-integrity-screen

EmpiricalOn 16 February 2024 Judge Patti B. Saris denied both of CVS's motions in their entirety in Baker v. CVS Health Corporati…

On 16 February 2024 Judge Patti B. Saris denied both of CVS's motions in their entirety in Baker v. CVS Health Corporation, No. 1:23-cv-11483 (D. Mass.) — the failure-to-state-a-claim motion aimed at the notice count and the separate Article III standing motion. No count and no defendant was dismissed. The court held that the statute's private right of action for any person aggrieved reaches notice violations, and that denial of information to which a plaintiff has a legal right can be a concrete injury in fact: the required notice 'would have specifically informed Baker that the HireVue Interview was a lie detector test,' and he participated without that warning. CVS did not challenge the sufficiency of the counts alleging that the screen itself violated the lie-detector prohibition, so the court accepted that characterization as plausibly pleaded rather than deciding it. This is a pleading-stage holding. There is no merits finding that the HireVue interview was a lie detector test, and no finding of any kind about the screen's accuracy.

Sources: bakerv2024, arentfoxschiffdavidson2024

Appears on: /domains/cases/cvs-hirevue-integrity-screen

EmpiricalLiability in this case landed on the employer while the two companies that built and ran the contested inference sat out…

Liability in this case landed on the employer while the two companies that built and ran the contested inference sat outside it. The defendants were CVS Health Corporation and CVS Pharmacy, Inc. only; HireVue and Affectiva were never parties and no claim was ever adjudicated against either. On the pleaded account the inference executed two organizational hops from the party bearing the statutory duty: the recordings were captured on the interview platform, uploaded onward to the affect-analysis firm, composed into an employability score at the platform, and returned to CVS recruiters. The statutory duties of section 19B bind the employer, not the vendors. The applicant-facing surface was the recorded interview alone: the complaint pleads three deprivations — no notice, no opt-out, and no ability to challenge the assessment results — and the record documents no applicant access to the recording, the affect analysis or the score, and no feedback of hiring outcomes into retraining. Because neither vendor's internal pipeline was ever put in evidence, the court's recitation of the pleaded mechanics is the authoritative public account of how this system worked.

Sources: bakerv, bakerv2024, hrdivecrist2024

Appears on: /domains/cases/cvs-hirevue-integrity-screen

EmpiricalThe case ended without deciding anything about the technology. A settlement notice was filed 17 July 2024; the public do…

The case ended without deciding anything about the technology. A settlement notice was filed 17 July 2024; the public docket shows the case terminated 22 July 2024; a stipulation of voluntary dismissal with prejudice followed on 20 September 2024. The settlement was individual and confidential and was reached before any class-certification ruling: no monetary terms, no practice changes and no admission of liability were disclosed, and no class member received anything. The named claim was extinguished and the pleaded class went unrepresented. Both docket dates are carried here rather than a single 'settled' date. The mirrored opinion does not recite an original filing court or removal path, so the filing history is stated only as reported by the Boston Globe on 22 May 2023 and on the federal docket from 30 June 2023.

Sources: hrdivecrist2024, bakerv, thebostonglobejohnston2023

Appears on: /domains/cases/cvs-hirevue-integrity-screen

EmpiricalAfter the February 2024 ruling, litigation under section 19B surged: more than twenty class actions in roughly a year, l…

After the February 2024 ruling, litigation under section 19B surged: more than twenty class actions in roughly a year, largely by a single New York firm, often with the same individuals suing multiple employers, including Procter & Gamble. Most allege ONLY that an employment application lacked the statutory notice, with no actual screening device involved — so what propagated was the notice theory rather than any finding about AI screening. Compliance advisories now direct every Massachusetts employer to print the statutory notice in the application itself rather than in a policy filed elsewhere, and illustrate the exposure arithmetically: an employer processing 200 Massachusetts applications a year faces roughly 100,000 dollars of annual notice-violation exposure at the 500-dollar statutory minimum. That figure is an illustration of the statutory structure, not a measured exposure for CVS or any other employer, and no CVS Massachusetts application volume or class size appears in the verified record. The system-level correction now visible — notice lines appearing on Massachusetts employment applications — is driven by per-application statutory damages and copycat litigation pressure rather than by any adjudicated finding about the technology.

Sources: morganlewislawflashengelman2025, agencychecklists2025

Appears on: /domains/cases/cvs-hirevue-integrity-screen

EmpiricalTwo things in this record cut against the pleaded account and both are carried rather than resolved. HireVue removed fac…

Two things in this record cut against the pleaded account and both are carried rather than resolved. HireVue removed facial analysis from new assessments in March 2020 and announced the change in January 2021 alongside a commissioned algorithmic audit, retaining speech-and-language analysis; it told the Boston Globe that visual and audio analysis 'have since been eliminated,' and its chief data scientist rejected the deception-detection characterization, saying the assessments measure work competencies 'statistically linked' to job success using 'validated industrial organizational psychology.' The application in this case was made around January 2021, and whether Affectiva affect analysis in fact ran on it, or on any class member's interview, was never adjudicated. Separately, the scientific question the 1959 statute was answering was never reached either: Brandeis psychologist Leonard Saxe told the Globe that there is no neurological signal of deception and 'no way for an automated system to distinguish a falsehood from the truth.' That is an expert view recorded in a newspaper, not a finding, and it is a statement about the construct rather than a measurement of this or any deployment.

Sources: maurer2021b, thebostonglobejohnston2023

Appears on: /domains/cases/cvs-hirevue-integrity-screen

EmpiricalHireVue's video-assessment platform is a vendor-layer deployment: one scoring engine behind hundreds of separate employe…

HireVue's video-assessment platform is a vendor-layer deployment: one scoring engine behind hundreds of separate employers' hiring pipelines. Candidates record answers to a structured question set, or play game-based assessments; per-assessment models score verbal and paraverbal features of the responses into competency scores that place each applicant in a Bottom, Middle or Top tier, and the client employer chooses where to cut — one deploying employer's published bias audit records it evaluating at Top-plus-Middle against Bottom. The client also chooses whether to use algorithmic scoring at all: as of January 2021 the vendor reported that approximately 20 percent of its customers used the predictive-analytics feature and that the rest used the platform for human review of recorded video. Every scale figure is vendor-reported and none is independently audited: more than 19 million video interviews and more than 700 customers as of January 2021, and more than 33 million interviews, 200 million chat-based candidate engagements and more than 800 customers as of January 2023. The models are built on historical applicant data pooled across employer implementations, and the mandated bias audits are computed on that same pooled store. A rejected candidate generates no outcome data anywhere in that loop, so the store that trains and audits the models cannot observe the population the models screened out.

Sources: fortune2021, maurer2021b, dciconsultinggroup2023, hirevue2021

Appears on: /domains/cases/hirevue-video-assessment

EmpiricalThe vendor removed visual and facial analysis from new assessment models in early 2020 — the Society for Human Resource …

The vendor removed visual and facial analysis from new assessment models in early 2020 — the Society for Human Resource Management reports the discontinuation as March 2020 — and announced the decision publicly on 12 January 2021, together with the results of an algorithmic audit it had commissioned from O'Neil Risk Consulting and Algorithmic Auditing. Its stated reason was that advances in language analysis had left visual features contributing little: internal research put the nonverbal visual contribution at about 0.25 percent of the model's predictive power in most job models and about 4 percent for high-customer-contact roles, figures given by the vendor's chief data scientist and never independently audited. Its chief executive said it was not worth the concern it was causing people. The sequence is complaint in November 2019, removal in about March 2020, public announcement with the audit in January 2021; causation by the complaint is an inference the record does not make, and the vendor's own stated reasons were the low measured contribution and rising public concern. The removal is a rare documented instance of an input dropped as its measured value approached zero while the cost of scrutiny rose, and it was applied platform-wide by a party no single client employer could have compelled.

Sources: maurer2021b, fortune2021, hirevue2021, schellmann2021

Appears on: /domains/cases/hirevue-video-assessment

EmpiricalTwo audit regimes reach this deployment and their perimeters are documented. The 2020 commissioned audit, announced 12 J…

Two audit regimes reach this deployment and their perimeters are documented. The 2020 commissioned audit, announced 12 January 2021, examined one representative pre-built early-career assessment use case, did not examine the tool's technical design or its training data, proceeded largely by structured stakeholder interviews, and its report is published on the vendor's own site only behind a nondisclosure agreement; the phrase that the assessments work as advertised with regard to fairness and bias is the vendor's characterization of exactly that scope, and the recorded criticisms are that its depth compared unfavourably with a contemporaneous source-code audit of a competitor, that an auditor paid by the audited party carries a conflict risk, and that neither audit addressed whether the products improve hiring at all. The audit did produce recommendations, on investigating accent bias and on the flagging of candidates who give brief answers. From January 2023 the vendor engaged DCI Consulting Group for the annual bias audits New York City Local Law 144 requires, covering competency-based and game-based algorithms across race, gender and intersectional groups. A summary produced 5 July 2023 and published by Pfizer as a deploying employer reports the mechanics: nationwide applicant data from January 2021 to December 2022, pooled across employer implementations, analysed per implementation and then aggregated; on the Communication assessment for intern and new-college-graduate jobs, 20,060 male against 9,121 female applicants, selection rates of 0.66 and 0.65 for Top-plus-Middle against Bottom, a gender impact ratio of 0.98, race and ethnicity ratios from 0.87 to 0.96 and intersectional ratios down to 0.82; on the Adaptability assessment for the same population, 7,161 male against 3,884 female applicants, a gender ratio of 0.95 and intersectional ratios from 0.79 to 1.01. A peer-reviewed study of all 116 publicly available Local Law 144 bias audits published between July 2023 and November 2024 found that DCI conducted 20 percent of them, that 54 percent of audits carried at least one impact ratio above 1 — the majority of those being DCI audits of HireVue tools deployed by JetBlue, Citizens, Pfizer or Burlington, an artifact of the aggregation and comparator-group method — and that every identified 'silent duplicate', meaning identical quantitative results republished across or within reports, appeared in DCI-conducted audits, all but one describing HireVue tools. Its authors could not determine the cause of all the duplicates. These are observations about a measurement regime; they are not findings of discrimination and not findings of audit fraud. An impact ratio is a ratio of selection rates over the applicants an audit could classify, and it is not an error rate — no error rate for this engine has ever been published by anyone.

Sources: schellmann2021, maurer2021a, dciconsultinggroup2023, gerchick2025

Appears on: /domains/cases/hirevue-video-assessment

EmpiricalThree external channels reached the vendor and produced three different kinds of result. On 6 November 2019 the Electron…

Three external channels reached the vendor and produced three different kinds of result. On 6 November 2019 the Electronic Privacy Information Center filed a complaint with the Federal Trade Commission alleging that the AI-based candidate assessments constituted unfair and deceptive practices under Section 5, that the company falsely denied using facial recognition, and that its results were biased, unprovable and not replicable; no public FTC enforcement action against the company is on the record as of August 2026, so what the record supports is that the vendor was the subject of an advocacy complaint and nothing stronger. Six Illinois residents filed Deyerler v. HireVue, Inc. on 27 January 2022 in the Northern District of Illinois (No. 1:22-cv-01284), alleging that the software collected facial geometry and voice data during virtual job interviews without the disclosures and written consent the Illinois Biometric Information Privacy Act requires; on 26 February 2024 Judge Jeremy C. Daniel granted in part and denied in part the motion to dismiss, letting claims under sections 15(a), (b) and (d) proceed, dismissing the section 15(c) profit claim on the ground that selling software is not selling biometric identifiers, and rejecting the argument that the Illinois Artificial Intelligence Video Interview Act precludes BIPA claims — holding the two statutes impose different but concurrent obligations. A claim that survives a motion to dismiss is an allegation held plausible and never a finding of violation. On 25 June 2026 the Circuit Court of Lake County, Illinois (No. 2026LA00000141, Hon. Daniel L. Jasica) granted PRELIMINARY approval of a 3,750,000 dollar class settlement covering an estimated 91,305 people who completed an interview involving the challenged voice and facial biometrics technology while in Illinois between 27 January 2017 and 25 June 2026, with an estimated 150 dollars per valid claimant subject to pro rata reduction, a claims deadline of 13 October 2026 and a final approval hearing on 28 October 2026. Nothing has been paid, the class was conditionally certified for settlement purposes only, the settlement is expressly no admission of wrongdoing, and the vendor denies that it collected or possessed biometrics or any other information subject to BIPA. On 19 March 2025 the ACLU of Colorado filed charges with the Colorado Civil Rights Division and the EEOC on behalf of a Deaf, Indigenous employee, alleging that an automated video interview relying on automated speech recognition disadvantaged her and that a request for human-generated captioning was denied; both companies dispute the charges and the vendor's chief executive called the complaint entirely without merit and stated that the employer did not use one of its AI-based assessments. No court and no regulator has ever found this vendor violated the biometric statute, the FTC Act, or any discrimination law.

Sources: electronicprivacyinformation2019, memorandumorder2024, noticeofproposedclassactions2026, aclu2025a

Appears on: /domains/cases/hirevue-video-assessment

EmpiricalBoth statutes that reach this deployment place their duties on EMPLOYERS rather than on the vendor, and that placement i…

Both statutes that reach this deployment place their duties on EMPLOYERS rather than on the vendor, and that placement is the structural finding of the case. The Illinois Artificial Intelligence Video Interview Act (820 ILCS 42, effective 1 January 2020) requires employers using AI analysis of video interviews to notify applicants before the interview, explain how the AI works and what general types of characteristics it uses to evaluate them, and obtain consent to be evaluated; it restricts sharing of the video, requires deletion within 30 days of an applicant's request, and from 2022 requires employers relying solely on AI analysis to report applicant race and ethnicity data annually. Its text states no express enforcement mechanism and no private right of action — which is why the Deyerler plaintiffs sued under the biometric statute instead, and why the February 2024 holding that the video-interview statute does not displace the biometric one mattered: it kept open the only channel in this record with teeth. New York City Local Law 144 likewise places its annual independent-bias-audit and publication duty on employers. The vendor holds the pooled data and engages the auditor; the employers hold the duty and publish the result — which is how four named employers came to publish the same vendor-level numbers as their own audits. The vendor commissioned the audits voluntarily and argued publicly, in announcing the January 2023 engagement, that vendors should also bear audit responsibility. Nothing in this record describes either statute as regulating the vendor directly.

Sources: artificialintelligencevideoi2020, memorandumorder2024, gerchick2025, hirevue2021

Appears on: /domains/cases/hirevue-video-assessment

EmpiricalIndependent scholarship bounds what automated video assessment can validly claim, and it measures the construct class ra…

Independent scholarship bounds what automated video assessment can validly claim, and it measures the construct class rather than any one product. Hickman and colleagues, in the Journal of Applied Psychology in 2022, investigated automated video-interview personality assessments across a development sample of 1,073 and a retest sample of 99: models trained on self-reports showed little evidence of reliability or validity, while models trained on interviewer reports performed better but with mixed cross-sample reliability, and the authors cautioned vendors and adopting organizations accordingly. Raghavan and colleagues, at ACM FAT* in 2020, analysed what algorithmic pre-employment vendors publicly claim about validation and bias mitigation and found those claims largely unverifiable from what vendors disclose. Against that, a technology-press review of both audit regimes reaching this deployment recorded that neither addressed whether the products improve hiring at all. So the record contains no measurement of this engine's accuracy from any source: the mandated audits report selection-rate ratios, the commissioned audit did not examine the technical design or the training data, and no error rate has ever been published.

Sources: hickman2022, raghavan2020a, schellmann2021

Appears on: /domains/cases/hirevue-video-assessment

EmpiricalOn 19 March 2025 the ACLU, the ACLU of Colorado, Public Justice and Eisenberg & Baum LLP filed a Complaint of Discrimina…

On 19 March 2025 the ACLU, the ACLU of Colorado, Public Justice and Eisenberg & Baum LLP filed a Complaint of Discrimination with the Colorado Civil Rights Division and the U.S. Equal Employment Opportunity Commission on behalf of D.K., a Deaf Pawnee woman who communicates in English with a deaf accent and in ASL, against BOTH Intuit, Inc. and HireVue, Inc. It alleges violations of the Colorado Anti-Discrimination Act (Colo. Rev. Stat. § 24-34-402), the Americans with Disabilities Act and Title VII in the denial of her promotion to Seasonal Manager, and pleads HireVue as an employment agency, an agent of the employer, an indirect employer and an aider and abettor under state law. None of those theories has been tested. Everything the charge asserts is an allegation and no probable-cause determination, dismissal, right-to-sue notice, court filing or settlement has been made public.

Sources: complaintofdiscriminationaga2025, aclu2025a, publicjustice2025

Appears on: /domains/cases/intuit-hirevue-promotion-screen

EmpiricalThis is an internal-promotion case rather than a point-of-hire case, and the employer held years of direct performance e…

This is an internal-promotion case rather than a point-of-hire case, and the employer held years of direct performance evidence about the person the screen was assessing. Per the complaint and D.K.'s own published account she was hired as a seasonal Tax Associate in late 2019, became a Tax Expert Lead supporting a team of about four hundred Tax Associates, held that role for three seasons, earned a bonus every year, was praised by supervisors for her communication, joined Intuit's Accessibility Team in 2023, and was encouraged to apply for Seasonal Manager by her own manager, who sat on the hiring team. The promotion ladder ran Tax Associate to Tax Expert Lead to Seasonal Manager, with Tax Expert Leads trained for the next step and, in the complaint's words, many being promoted each year. She applied in spring 2024 through the employee dashboard and was invited on 21 June 2024 into an assessment of roughly three hours: about a dozen timed recorded video questions, mostly management scenarios with a few tax-law questions, plus essay and multiple-choice sections, with instructions and questions delivered audibly by a recording of a person speaking.

Sources: complaintofdiscriminationaga2025, acluofoklahoma2025

Appears on: /domains/cases/intuit-hirevue-promotion-screen

EmpiricalThe accommodation sequence as alleged: the interview invitation offered only a technical-support email address and carri…

The accommodation sequence as alleged: the interview invitation offered only a technical-support email address and carried no accommodations information, so D.K. initiated a request herself from prior familiarity with Intuit's process. She asked for Communication Access Real-time Translation (CART) — human-generated real-time captioning — deliberately not ASL interpretation, because interpreters she had worked with previously could not handle tax concepts, and she did not ask to be excused from the video interview because it was never offered and she feared that asking would count against her. The request was denied and she was told the platform's built-in subtitles could be enabled. When she began the assessment, she recounts, no subtitle option existed, and she completed roughly three hours of it relying on her browser's automatic captions, which she describes as less accurate and particularly poor on complicated words. The alleged failure therefore includes a control that was pointed to and was not present, rather than only a degraded substitute channel. Intuit's on-record position is that the allegations are entirely without merit and that it provides reasonable accommodations to all candidates.

Sources: complaintofdiscriminationaga2025, acluofoklahoma2025, aclu2025b

Appears on: /domains/cases/intuit-hirevue-promotion-screen

EmpiricalThe complaint alleges Intuit was on notice twice over, through two channels of its own. First, after the 2023 season D.K…

The complaint alleges Intuit was on notice twice over, through two channels of its own. First, after the 2023 season D.K. — by then a member of Intuit's Accessibility Team — told the Team's Chair that the HireVue step was inaccessible and could exclude deaf applicants; she was told it would be looked into, no action was ever reported back, and Intuit required the same step of her the following year. Second, in 2020 Intuit's OWN automated call-monitoring system, which scored how closely agents followed scripts from speech-recognition transcripts of customer calls, allegedly read her deaf accent as deviation from script and produced one artificially low indicator against otherwise strong customer-satisfaction, resolution-count and response-speed measures. Intuit's alleged response was to reassign her from phone calls to the chat channel rather than to correct the metric; the low indicator remained in her record, and the complaint pleads it may also have been considered in the denial of her promotion. That call-monitoring system is Intuit's, not HireVue's, and the charge uses it as evidence of awareness rather than as the promotion screen.

Sources: complaintofdiscriminationaga2025, acluofoklahoma2025

Appears on: /domains/cases/intuit-hirevue-promotion-screen

EmpiricalOn 13 August 2024 D.K. received an automated rejection stating that Intuit had decided to move forward with other candid…

On 13 August 2024 D.K. received an automated rejection stating that Intuit had decided to move forward with other candidates. On 25 September 2024 she received a feedback email — which she believed from its return address and generic language to have been generated by HireVue's automated analysis — advising her to give more concise and direct answers, to adapt her communication style to different audiences, and to practise active listening; the complaint says those recommendations map directly onto her being Deaf. On 31 October 2024 she made a personnel-file request under Colorado law, and the complaint filed the following March alleges that only minimal documents were produced. The complaint pleads HireVue's platform mechanism in a three-stage form — speech recognition of the spoken answers, a second system interpreting the transcript, and machine-learning scoring against job competencies — but pleads the scoring allegations about her specific interview on information and belief, inferred from that feedback email. That inference is exactly what HireVue disputes.

Sources: complaintofdiscriminationaga2025, proskauerrosellp2025, aclu2025a

Appears on: /domains/cases/intuit-hirevue-promotion-screen

EmpiricalThe technical premise the charge builds on is a published measurement of general-purpose speech recognition, and it belo…

The technical premise the charge builds on is a published measurement of general-purpose speech recognition, and it belongs to that study rather than to any deployment. A 2025 paper in The Laryngoscope (135:191-197) measured four widely used commercial speech-recognition services against a twenty-four-speaker corpus and reported a mean word error rate of 52.6 percent for d/Deaf and hard-of-hearing speakers versus 5.0 percent for normal-hearing speakers — per service 45.1 to 57.3 percent against 3.8 to 5.9 percent, which the authors describe as ten times higher — rising to 85.9 percent for speakers with the lowest speech-intelligibility classification, 80.5 percent for prelingual hearing-loss onset and 70.2 percent for sign-primary communicators. The complaint cites that study, together with peer-reviewed findings of higher speech-recognition error rates for Black speakers and for ethnicity-related dialects, and extends them on information and belief to speakers of Indigenous dialects of English. The study measured four general-purpose services. It did not measure HireVue's system or this interview, and the inference from that literature to this platform is contested.

Sources: zhao2025a, complaintofdiscriminationaga2025

Appears on: /domains/cases/intuit-hirevue-promotion-screen

EmpiricalBoth respondents dispute the charges, and the vendor's dispute goes to the factual premise rather than to the legal conc…

Both respondents dispute the charges, and the vendor's dispute goes to the factual premise rather than to the legal conclusions. HireVue's chief executive Jeremy Friedman stated that the complaint 'is entirely without merit and is based on an inaccurate assumption about the technology used in the interview. Intuit did not use a HireVue AI-based assessment.' An Intuit spokesperson stated: 'The allegations in the complaint are entirely without merit. We provide reasonable accommodations to all candidates.' Whether any automated assessment ran on this interview at all is therefore in dispute, and the candidate-side evidence bearing on it is the return address and register of one feedback email. As of 28 August 2026 the matter remained a pending, non-public administrative investigation: no Colorado Civil Rights Division or EEOC determination, no right-to-sue notice, no court complaint and no settlement had been publicly reported, and an April 2026 legal-commentary treatment of the case records no post-filing developments. Administrative charge proceedings are confidential, so that silence establishes no public development rather than that nothing occurred.

Sources: aclu2025a, undergraduatelawreviewatflor2026, aclu2025b

Appears on: /domains/cases/intuit-hirevue-promotion-screen

EmpiricalThe U.S. Equal Employment Opportunity Commission alleged that iTutorGroup, Inc., Shanghai Ping'An Intelligent Education …

The U.S. Equal Employment Opportunity Commission alleged that iTutorGroup, Inc., Shanghai Ping'An Intelligent Education Technology Co., Ltd. and Tutor Group Limited — three integrated companies providing English-language tutoring to students in China through United-States-based tutors working fully remotely from their homes — had programmed their tutor application software to automatically reject female applicants aged 55 or older and male applicants aged 60 or older, rejecting more than 200 qualified United-States applicants because of their age. After conciliation failed the Commission sued under the Age Discrimination in Employment Act on 5 May 2022, No. 1:22-cv-02565 (E.D.N.Y.); then-Chair Charlotte Burrows framed the filing for the agency's algorithmic-enforcement agenda with the words 'Even when technology automates the discrimination, the employer is still responsible.' The parties filed a joint notice of settlement with a proposed consent decree on 9 August 2023 and the court approved the decree in September 2023, announced by the Commission on 11 September 2023. iTutorGroup pays $365,000 to be distributed among the more than 200 rejected applicants through a claims process, without admitting liability; per-claimant amounts were not made public. iTutorGroup denied the allegations and disputed that the tutors were employees at all, characterizing them as independent contractors, a question the settlement resolved without adjudication.

Sources: u2023a, u2022c, hrdive2023

Appears on: /domains/cases/itutorgroup-age-screening

EmpiricalThe screen at issue was an authored decision boundary rather than anything learned, and the distinction is the case's do…

The screen at issue was an authored decision boundary rather than anything learned, and the distinction is the case's doctrinal value. The EEOC's own releases describe 'tutor application software' that was 'programmed to automatically reject female applicants aged 55 or older and male applicants aged 60 or older' — a deterministic, sex-differentiated threshold computed from the birthdate field collected on the application form, with no score, no training data and no published error rate anywhere in the record. It sat at the top of the hiring funnel, upstream of any human reviewer: rejection was automatic at application intake, no human review point between the rule and the rejection notice is documented, and the employer's own hiring staff saw only applicants who came through the screen. Rejected applicants received no statement of the reason. The settlement was widely described in the legal press as the EEOC's first workplace artificial-intelligence settlement, a framing that belongs to that press and not to the agency, whose releases never use the term; the case's significance is that the regulator treated an automated screen as ordinary actionable age discrimination regardless of mechanism, which is why it proceeded as intentional disparate treatment. All of the conduct described here is the Commission's allegation, resolved by a decree carrying no admission of liability.

Sources: u2023a, u2022c, greenbergtraurigllp2023

Appears on: /domains/cases/itutorgroup-age-screening

EmpiricalThe practice surfaced through a single applicant's experiment on the input rather than through any internal control. The…

The practice surfaced through a single applicant's experiment on the input rather than through any internal control. The charging party applied with her real birthdate and was rejected immediately, then re-applied about a day later with an application identical in every respect except a more recent birthdate, and was offered an interview. That comparison required no inside access, one changed field and roughly a day, and it is the only comparison of a rejection against a counterfactual anywhere in the public record. The decree's reapplication invitations cover applicants rejected in March and April 2020, fixing the documented rejection window at two months, and roughly three and a half years then ran from that window to an enforceable remedy: rejections March to April 2020, a charge and failed conciliation, suit on 5 May 2022, joint notice of settlement 9 August 2023, decree approved September 2023. No public record identifies who authored the rule, why the thresholds differed by sex, or whether any internal review approved or missed it.

Sources: hrdive2023, greenbergtraurigllp2023, u2023a

Appears on: /domains/cases/itutorgroup-age-screening

EmpiricalThe consent decree acted on the rule's input and on the oversight structure rather than on any decision logic. Its terms…

The consent decree acted on the rule's input and on the oversight structure rather than on any decision logic. Its terms: injunctions against hiring discrimination based on age or sex; a prohibition on requesting applicants' birth dates before an offer, which removes from the intake form the field the rule computed on; a new anti-discrimination policy and an internal memo; multiple anti-discrimination trainings for those involved in hiring tutors; written notice to the EEOC of discrimination complaints, converting a formerly internal channel into a regulator-visible one; invitations to the applicants rejected in March and April 2020 to reapply; and, because iTutorGroup had already ceased hiring tutors in the United States, an obligation to notify and interview those applicants if it resumes United States operations, so part of the decree's machinery has never had occasion to operate. The Commission's own release and Seyfarth's 2024 recap give the duration as monitoring compliance 'for at least the next five years or longer if iTutorGroup resumes hiring tutors in the United States' — a floor with a conditional tail, not a flat five-year term. No post-decree enforcement activity in this case was identified.

Sources: u2023a, greenbergtraurigllp2023, seyfarthshawllp2024

Appears on: /domains/cases/itutorgroup-age-screening

EmpiricalThe federal enforcement channel this decree instantiated changed measurably after the decree was entered, and the change…

The federal enforcement channel this decree instantiated changed measurably after the decree was entered, and the change is about the agency rather than about this case. In January 2025 the EEOC removed its Artificial Intelligence and Algorithmic Fairness Initiative content along with its May 2023 Title VII technical assistance on artificial intelligence and its May 2022 guidance on the Americans with Disabilities Act, following Executive Order 14179; the Department of Labor and the Office of Federal Contract Compliance Programs made parallel removals. Removing guidance repeals no law: Title VII, the Age Discrimination in Employment Act and the Uniform Guidelines on Employee Selection Procedures are unchanged, and no federal safe harbor was created. Executive Order 14281 of 23 April 2025 directs that 'All agencies shall deprioritize enforcement of all statutes and regulations to the extent they include disparate-impact liability' and orders a ninety-day evaluation of existing disparate-impact consent judgments; this decree rests on intentional disparate treatment, so that directive does not reach its theory, and nothing in the public record suggests the decree is threatened. The Commission lacked a quorum from January 2025 until 7 October 2025, when Commissioner Brittany Panuccio was confirmed; Andrea Lucas was designated Chair on 5 November 2025, and the restored commission's published 2026 priorities centre on investigations of diversity programmes, religious accommodation and national-origin cases, with artificial intelligence in hiring absent from the stated agenda.

Sources: cooleyllp2025, executiveorder2025, hollandknightllp2025

Appears on: /domains/cases/itutorgroup-age-screening

EmpiricalMcHire is the chatbot-first hiring front end for approximately 90 percent of McDonald's franchisees, powered by Paradox.…

McHire is the chatbot-first hiring front end for approximately 90 percent of McDonald's franchisees, powered by Paradox.ai's conversational assistant Olivia. An applicant supplies contact details and shift preferences in natural-language chat, answers screening questions, is routed into a personality assessment administered by a further third party, Traitify.com, presented as agree-or-disagree phrase items, and self-schedules an interview before advancing to a human review stage where franchisee owners and managers manage them through the hiring stages. The employment-decision surface here is pre-screening and flow control rather than final selection: the assistant screens and schedules, and a person decides. Paradox is a conversational-hiring vendor whose client list extends well beyond this brand to other high-volume frontline employers including Aramark, Lockheed Martin, Lowe's and Pepsi, and in acquisition coverage Wendy's, 7-Eleven and General Motors. No error rate, completion rate or complaint rate for the assistant has been published by anyone, and no independent evaluation of this deployment exists; what the record does contain is public complaints that the assistant was answering nonsensically, which is what prompted two outside researchers to examine the platform at all.

Sources: carroll2025, bleepingcomputer2025, krebsonsecurity2025, workday2025

Appears on: /domains/cases/mchire-paradox-chatbot

EmpiricalEvery conversation and every form field on McHire persisted into a single vendor-held lead record. Each record contained…

Every conversation and every form field on McHire persisted into a single vendor-held lead record. Each record contained the applicant's name, email address, phone number, home address, shift preferences, the application's status and state-change history, the full raw chat transcript of the exchange with the assistant, and an authentication token permitting login to that applicant's own consumer interface — an impersonation surface stored alongside the record it identifies. Records accumulated without any purge described in the record: identifiers could be walked backward across the platform's history. The researchers' own test application received a lead_id of approximately 64,185,742, so application records numbering to roughly 64 million were REACHABLE through the flawed endpoint. That figure is the size of an identifier space and BleepingComputer states it represents the total number of job applications on the platform rather than unique applicants; it is a derivation from an identifier and is not a count of records exfiltrated or of people harmed. What makes this store a hiring object rather than a generic personal-data store is its content: who applied where, what they told a recruiting agent in their own words, the assessment step they were routed through, and their application status history. No documented model-retraining loop reads from these records.

Sources: carroll2025, bleepingcomputer2025, csoonline2025

Appears on: /domains/cases/mchire-paradox-chatbot

EmpiricalSecurity researchers Ian Carroll and Sam Curry found two flaws in what their writeup describes as a cursory review, prom…

Security researchers Ian Carroll and Sam Curry found two flaws in what their writeup describes as a cursory review, prompted by public complaints that the assistant was answering nonsensically. First, the McHire administration login for restaurant owners accepted the default credentials 123456:123456 on a Paradox test-restaurant account, granting administration-panel access for that instance including all in-progress conversations. Paradox states that this legacy test account had not been logged into since 2019 and, in its own words, should have been decommissioned. Second, the endpoint PUT /api/lead/cem-xhr performed no authorization check on its lead_id parameter, returning any applicant's unmasked record to a caller who simply named the identifier; decrementing the identifier returned other applicants' records. Before 30 June 2025 nothing in the record surfaced either flaw: no employment regulator, no privacy regulator, no contractual security review, no audit and no internal review is documented examining the custody surface, McDonald's is not documented exercising its principal-to-vendor authority over the platform before the disclosure email arrived, and franchisees running hiring on the platform had no visibility into vendor credential practice and no documented channel through which to acquire any.

Sources: carroll2025, csoonline2025, paradox2025

Appears on: /domains/cases/mchire-paradox-chatbot

EmpiricalThe disclosure-to-remediation sequence is documented to the minute. The researchers emailed Paradox.ai and McDonald's at…

The disclosure-to-remediation sequence is documented to the minute. The researchers emailed Paradox.ai and McDonald's at 5:46 PM Eastern on 30 June 2025; McDonald's acknowledged at 6:24 PM, 38 minutes later; the default 123456 credentials were disabled by 7:31 PM the same evening, under two hours after the report; and Paradox confirmed the insecure direct object reference (IDOR) fix at 10:18 PM Eastern on 1 July 2025, approximately 29 to 30 hours after disclosure. On 9 July 2025 Paradox published a Responsible Security Update accepting responsibility in terms ('We take responsibility for this issue. Full stop.'), stating that five candidate records containing personally identifiable information — names, emails, phone numbers and IP addresses, all US-based candidates — plus two chat records containing no candidate information had been viewed and exclusively by the two researchers, that the store contained no Social Security numbers, that only the one client instance was affected and that no candidate information was leaked online, and announcing a bug bounty programme and a dedicated security contact at security@paradox.ai. That vendor position is uncontradicted and has never been verified by any independent forensic report, so potential exposure and actual access remain two separate quantities with nothing reconciling them. McDonald's statement, given to Wired and quoted verbatim in trade coverage, placed the failure with its supplier: 'We're disappointed by this unacceptable vulnerability from a third-party provider, Paradox.ai. As soon as we learned of the issue, we mandated Paradox.ai to remediate the issue immediately, and it was resolved on the same day it was reported to us.' The governance channel that fired was the security-research responsible-disclosure norm, operated by two people holding no contract, no mandate, no statutory standing and no access; detection and correction were both external to the deployment.

Sources: carroll2025, paradox2025, bleepingcomputer2025, csoonline2025

Appears on: /domains/cases/mchire-paradox-chatbot

EmpiricalNo litigation, enforcement action or state attorney-general breach-notification filing tied to the McHire exposure was l…

No litigation, enforcement action or state attorney-general breach-notification filing tied to the McHire exposure was located as of 28 August 2026. That is stated as verified absence and never as exoneration, and the limits of the check are stated with it: the Maine attorney general's breach portal was offline at check time, reporting an apparent abuse of its data-breach reporting system, and the California attorney general's published breach list returned no entries, so the notification-registry finding rests on searches plus those partial registry checks rather than an exhaustive sweep. One search-engine summary asserted a Maine attorney-general filing; no underlying record could be found and it is recorded as unsubstantiated. Paradox's position — that only the two researchers viewed a handful of records and that nothing was leaked online — is the stated factual predicate under which broad breach-notification duties would not attach, and no independent adjudication of that position exists. A separate thread is kept separate: two weeks after the disclosure, Krebs on Security reported that infostealer malware on Paradox administrator and developer machines in Vietnam, identified as Nexus Stealer, had exposed weak seven-digit numeric passwords reused across multiple customer accounts along with Okta single-sign-on and Atlassian tokens valid into December 2025, and that a second developer compromise in late 2024 involved GitHub credentials. Paradox disputed the currency of the exposed passwords in part, attributing them to legacy password-manager migrations and stating that few remained active. That reporting concerns the company's credential hygiene generally rather than the McHire authorization chain specifically, and the dispute travels with it.

Sources: krebsonsecurity2025, paradox2025, csoonline2025

Appears on: /domains/cases/mchire-paradox-chatbot

EmpiricalWorkday, Inc. announced a definitive agreement to acquire Paradox.ai on 21 August 2025, seven weeks after the disclosure…

Workday, Inc. announced a definitive agreement to acquire Paradox.ai on 21 August 2025, seven weeks after the disclosure, and completed the acquisition on 1 October 2025, folding the Olivia assistant into its talent-acquisition suite. Workday's own completion release describes the assistant as a conversational candidate-experience agent handling applications, screening support, self-scheduling and round-the-clock chat for frontline high-volume roles, and discloses no price. Trade coverage of the transaction carried the McHire incident and Paradox's dispute of its scope as due-diligence context, and recorded Paradox clients including Wendy's, 7-Eleven and General Motors. The effect on this case is that custody of the same applicant records moved inside a much larger human-capital software company three months after the exposure was disclosed. That acquirer is separately the subject of a different deployment record in this atlas on an unrelated question, and the two are not merged: this case is documented from the June and July 2025 record, when Paradox was independent.

Sources: workday2025, hrdive2025

Appears on: /domains/cases/mchire-paradox-chatbot

EmpiricalMeta Platforms' advertising system decides, for every job ad, which users inside the advertiser's eligible audience actu…

Meta Platforms' advertising system decides, for every job ad, which users inside the advertiser's eligible audience actually receive impressions, optimising predicted relevance and engagement against auction economics — so the allocation runs before anyone applies and the people it does not reach generate no application, no rejection and no adverse-action record anywhere. The platform's own published data, as recorded in the December 2022 class charge and the December 2023 joinder release, gives the scale: roughly 239 million US Facebook users, more than 30,000 US job ads published daily, women 54 percent of users interested in job hunting and people 55 and over more than 28 percent of them. The governance record runs on two layers that do not coincide: a restricted advertiser portal created by private settlement in March 2019, and a delivery-optimisation stage that independent peer-reviewed audits measured weeks later.

Sources: realwomenintruckingv2022, aarpfoundation2023, nationalfairhousingalliance2019, ali2019

Appears on: /domains/cases/meta-job-ad-delivery

EmpiricalOn 19 March 2019 Facebook settled five coordinated legal actions — brought by the National Fair Housing Alliance and thr…

On 19 March 2019 Facebook settled five coordinated legal actions — brought by the National Fair Housing Alliance and three regional fair-housing organizations with Emery Celli Brinckerhoff & Abady LLP, the Communications Workers of America, Outten & Golden LLP, the American Civil Liberties Union and individual plaintiffs — by agreeing to structural platform changes recorded in a joint statement the company co-signed: housing, employment and credit ads confined to a separate restricted portal; no gender, age or multicultural-affinity targeting; a fifteen-mile minimum geographic radius and no postal-code targeting; Lookalike Audiences for those ads stripped of gender, age, religion, postal-code and group-membership inputs; detection and rerouting of covered ads created outside the portal; advertiser anti-discrimination certification; plaintiff testing rights and regular implementation meetings; and a commitment to engage researchers on algorithmic bias, with implementation by 30 September 2019. The joint statement discloses no monetary terms; the roughly five-million-dollar figure attached to these settlements comes from contemporaneous press reporting rather than from the document.

Sources: nationalfairhousingalliance2019

Appears on: /domains/cases/meta-job-ad-delivery

EmpiricalIn April 2019, weeks after the settlements took effect, Ali, Sapiezynski, Bogen, Korolova, Mislove and Rieke published a…

In April 2019, weeks after the settlements took effect, Ali, Sapiezynski, Bogen, Korolova, Mislove and Rieke published a peer-reviewed study (CSCW 2019) showing that paired employment and housing ads carrying identical neutral targeting were delivered to strongly gender- and race-skewed audiences by the delivery stage itself, driven by the platform's own relevance predictions plus budget and market effects — with the budget channel operating independently of the content channel, because under a budget constraint the demographics that are cheaper to reach take more of the impressions. The method used ordinary advertiser accounts and was reproducible by outside auditors. In 2021 Imana, Korolova and Heidemann (WWW 2021) controlled for qualification distributions by running paired ads for the same job at companies with different de facto workforce gender mixes, confirmed gender skew in Facebook's job-ad delivery that qualification differences do not justify, and found no comparable skew on LinkedIn — making the behaviour platform-specific rather than labour-market-inevitable. Both studies measured the platform in the 2019-to-2021 era, before any remedial delivery controller existed even for housing ads.

Sources: ali2019, imana2021

Appears on: /domains/cases/meta-job-ad-delivery

EmpiricalOn 21 June 2022 the United States sued Meta Platforms under the Fair Housing Act over housing-ad delivery, following a H…

On 21 June 2022 the United States sued Meta Platforms under the Fair Housing Act over housing-ad delivery, following a HUD Secretary-initiated complaint and charge, and on 27 June 2022 the court entered a settlement in United States v. Meta Platforms, Inc., No. 22-cv-5187 (S.D.N.Y.): Meta stopped using the Special Ad Audience tool, paid the $115,054 civil penalty that was then the statutory maximum, and agreed to build a Variance Reduction System to shrink the gap between an ad's eligible audience and its actual audience by sex and estimated race or ethnicity, under a compliance-metrics agreement of 9 January 2023, four-monthly reporting, an independent third-party reviewer and court oversight through 27 June 2026. The system samples eligible-audience demographics, estimates race and ethnicity by Bayesian Improved Surname Geocoding at a 50 percent threshold with differential-privacy noise added to the aggregates, retains no individual-level race data, and adjusts ongoing delivery in flight. Guidehouse Inc. issued five verification reports between June 2023 and 30 October 2024 (the fifth updated 19 December 2024); in its period of 1 May to 31 August 2024, for housing ads above 1,000 impressions, 95.1 percent met the ten-percent sex-variance threshold against a 91.7 percent requirement and 85.5 percent met the ten-percent estimated race and ethnicity threshold against an 81.0 percent requirement, with the reviewer's independent recomputation matching Meta's reported coverage at a 0.0 percent difference. All of the court-supervised metrics and all published verification cover HOUSING ads. Meta reports in its own newsroom having extended the same system to US employment and credit ads by October 2023, an operator-reported extension with no published compliance metrics and no third-party verification. The court-oversight term expired on 27 June 2026; the Department of Justice case page was last updated 21 January 2025 and lists no verification report after 30 October 2024, and no public post-expiry disposition was located as of 28 August 2026.

Sources: ua, guidehouseinc2024, metaplatforms2023

Appears on: /domains/cases/meta-job-ad-delivery

EmpiricalPLEADING, NOT A FINDING. On 1 December 2022 Real Women in Trucking filed a class charge of discrimination with the Equal…

PLEADING, NOT A FINDING. On 1 December 2022 Real Women in Trucking filed a class charge of discrimination with the Equal Employment Opportunity Commission (No. 570-2023-00655), through Gupta Wessler PLLC and Upturn, Inc., alleging that Meta's ad-delivery algorithm discriminates by sex and age in deciding who receives job ads and pleading disparate treatment, disparate impact, the advertising-discrimination provisions and an employment-agency theory under Title VII and the ADEA and under state and local law; the charge alleges the post-2019 era inverted the mechanism, so that where job ads were once unlawful because advertisers expressly excluded older people and women, they are alleged to violate federal law since 2020 because the delivery algorithm itself relies on gender and age to limit who sees them. Its exhibits are tables from Facebook's own public ad library for ads whose eligible audiences were all-gender and all adult ages: a commercial-driving ad shown 94 percent to men and 5 percent to women, a mechanic ad above 99 percent to men, another blue-collar ad at 1 percent women and 3 percent users 55 and over, a school administrative-assistant ad at 2 percent men and 6 percent users 55 and over, a health-facility hospitality ad at 83 percent women, and a correctional-nursing ad at 22 percent men and 18 percent users 55 and over. AARP Foundation joined the charge on 19 December 2023 on behalf of older workers, citing more than 75 job postings as the evidentiary corpus. Meta's public response is that 'Addressing fairness in ads is an industry-wide challenge,' citing collaboration with civil-rights groups, academics and regulators. No tribunal has found Meta's delivery algorithm unlawful in employment. Counsel's live case page, accessed 28 August 2026, lists the charge as pending with no public EEOC determination. Its enforcement environment changed in 2025: the executive order of 23 April 2025, 'Restoring Equality of Opportunity and Meritocracy,' directed agencies to deprioritise disparate-impact liability, and a reported internal EEOC directive ordered disparate-impact-only charges administratively closed by 30 September 2025 with right-to-sue letters by 31 October 2025, from which charges also alleging intentional discrimination reportedly proceed on that theory alone.

Sources: realwomenintruckingv2022, upturn2022, aarpfoundation2023, prflawpeterromerfriedmanlaw2026, franczekp2025

Appears on: /domains/cases/meta-job-ad-delivery

EmpiricalFacebook's public ad library publishes per-ad demographic delivery tables only for ads classified as 'Issues, elections …

Facebook's public ad library publishes per-ad demographic delivery tables only for ads classified as 'Issues, elections or politics.' Employment ads receive no such disclosure unless they carry that classification, so neither employers nor job seekers can routinely observe the demographic composition of a job ad's delivery — and the December 2022 charge's own per-ad evidence exists precisely because some job ads were so classified, as does the corpus of more than 75 postings the AARP Foundation cited when it joined in December 2023. The consequence recorded across this case's sources is that every delivery-layer finding in the record originated outside the operator: paired-ad academic audits run on ordinary advertiser accounts, and advocate harvesting of the ad library.

Sources: realwomenintruckingv2022, aarpfoundation2023, ali2019, imana2021

Appears on: /domains/cases/meta-job-ad-delivery

EmpiricalATTRIBUTED TO THE EEOC, AND AGAINST ADVERTISERS. On 25 September 2019 the American Civil Liberties Union announced that …

ATTRIBUTED TO THE EEOC, AND AGAINST ADVERTISERS. On 25 September 2019 the American Civil Liberties Union announced that the Equal Employment Opportunity Commission had issued reasonable-cause determinations against seven employers — Capital One, Edward Jones, Enterprise Holdings, Drive Time Auto, Nebraska Furniture Mart, Renewal by Andersen and Sandhills Publishing — finding they unlawfully excluded women and older workers from the audiences of their Facebook job ads. The determinations form part of roughly 66 charges filed since 2018 by the Communications Workers of America, the ACLU and Outten & Golden (56 age-only, 10 age and gender), addressed advertiser TARGETING choices rather than the platform's delivery algorithm, produced no platform remedy, and moved to conciliation with no public outcomes since. Reasonable cause is an administrative finding, not a court judgment, and the verified release does not itself date the determinations' issuance — 25 September 2019 is the date of the announcement.

Sources: americancivillibertiesunion2019

Appears on: /domains/cases/meta-job-ad-delivery

EmpiricalIn March 2020 Northeastern University and pymetrics, inc. signed a sponsored research agreement that fixed, before any w…

In March 2020 Northeastern University and pymetrics, inc. signed a sponsored research agreement that fixed, before any work began, the scope of a source-code audit of pymetrics' candidate-screening model-generation pipeline, the auditors' remuneration, their publication rights, a non-compete and a thirty-day responsible-disclosure window. Four Northeastern auditors conducted the audit in summer 2020 inside a pymetrics-provisioned AWS virtual machine, with access to source code, eight Jupyter notebooks, one engagement's full training and evaluation data, confidential technical, fairness-testing and job-analysis documents, staff time and a closing live demonstration of a production model trained, tested and deployed. The independence instruments were real and are documented: pymetrics paid $104,465 to the university, $64,813 of it salaries for the team, structured as a grant and paid in full BEFORE findings were delivered so it could not be conditioned on results, and supplied the audit compute at its own expense; the auditors kept their test methods secret from pymetrics throughout; and they held final editorial discretion including the contractual right to publish results that might reflect negatively on pymetrics. The contract, the audit and data-sharing protocol, the non-compete, the budget spreadsheet and the final report with section 3.2 redacted for proprietary information were posted publicly by the auditors. The paper was published at ACM FAccT in March 2021 with eight authors, four of them Northeastern auditors and four pymetrics employees including the company's chief executive as last author. The auditors coined the term 'cooperative audit' for the arrangement and explicitly placed it outside the existing internal/external taxonomy: they were not pymetrics employees, yet held privileged access to source code, data and staff, so by the field's own prior definitions the audit was not an external one. The transparency address printed in the 2021 paper no longer resolves; the live page is at cbw.sh/research/audits/, verified 2026-08-28.

Sources: wilson2021a, wilson2021b, wilson2021, schellmann2021

Appears on: /domains/cases/pymetrics-cooperative-audit

EmpiricalThe control the audit examined is a pre-deployment four-fifths gate with abandon-if-noncompliant semantics. Per client r…

The control the audit examined is a pre-deployment four-fifths gate with abandon-if-noncompliant semantics. Per client role, pymetrics built a screening model from an in-group of 50 to 100 high-performing incumbents against an out-group sampled from a player database of more than two million, then tested the candidate model against a held-out 'bias group' — typically more than 10,000 players engineered to equal proportions across the seven EEOC categories, drawn from a pool of more than 600,000 players with self-reported demographics at over 75% survey completion per the vendor. Recommendation tiers sit at the 50th and 70th score percentiles, and deployment requires an impact ratio of at least 0.8 at BOTH thresholds on the bias group. If no compliant model is found, the engagement is reworked or no model ships. The audited code used support-vector-machine models over 64 features produced by 12 core games, filled gaps by median imputation, and dropped players missing more than two games.

Sources: wilson2021a, wilson2021b

Appears on: /domains/cases/pymetrics-cooperative-audit

EmpiricalThe auditors found pymetrics' code correctly implemented four-fifths testing over the seven EEOC categories; that demogr…

The auditors found pymetrics' code correctly implemented four-fifths testing over the seven EEOC categories; that demographic data was not used as a training feature and no overt proxies such as zip codes were present; that they could not construct manipulated incumbent data producing a biased model without it being flagged, because every control-flow path reached the adverse-impact tests; that the 100-plus-item completion checklist per model plus the mandatory second-data-scientist review before production was a reasonable safeguard against negligence and against a single malicious insider, leaving collusion between two insiders as the residual path they explicitly named; and that median imputation did not substantively alter adverse-impact results despite statistically significant demographic differences in missing data. The auditors also recorded that no programmatic back-end re-check of adverse impact occurs after the data-scientist stage. Their own summary is that pymetrics 'passed the audit, subject to the qualifications and limitations we state'. No litigation, regulator enforcement action, consent decree or documented harm involving pymetrics or the audit was identified anywhere in this record.

Sources: wilson2021a, wilson2021b

Appears on: /domains/cases/pymetrics-cooperative-audit

EmpiricalThe audit's exclusions were pre-agreed with the audited party in the contracted scope before the work began, not discove…

The audit's exclusions were pre-agreed with the audited party in the contracted scope before the work began, not discovered as gaps afterwards. Excluded by that agreement: the choice of fairness objective and metric, the EEOC category set, intersectional groups, construct validity of the games, the newer reasoning games, post-deployment back-testing, customised client processes, security, and privacy compliance. Whether the games measure anything job-relevant — the assessment's reason for existing — was not assessed; the auditors state it was 'beyond our capabilities' as computer scientists. Intersectionality was excluded because such groups are not EEOC-recognised and all parties agreed the legal risk of optimising to a non-regulatory standard barred it, a reason the record states rather than adjudicates; pymetrics itself expressed interest in it. The audit is a summer-2020 snapshot of the standard, non-customised pipeline, and the auditors disclaim any claim about earlier practice, later practice or customised client engagements.

Sources: wilson2021a, wilson2021b

Appears on: /domains/cases/pymetrics-cooperative-audit

EmpiricalAt ACM FAccT 2022, Young, Katell and Krafft made this engagement the lead case of what they term 'publishing and certify…

At ACM FAccT 2022, Young, Katell and Krafft made this engagement the lead case of what they term 'publishing and certifying corporate apologia': four of the paper's eight authors were pymetrics employees including the chief executive as last author; the company funded the work and set the scope, 'precluding various scenarios from scrutiny'; and 'it seems unlikely that this paper would have been submitted had the results negated the firm's validation claims.' They also documented the downstream loop: pymetrics had publicised the engagement in a company post of 13 May 2020, before any results existed, and after publication marketed the product as 'independently audited' — a term the audit paper's own cited definitions do not support, since the auditors classify the arrangement as cooperative and place it outside both the internal and external categories. They further recorded that a co-author sat on the FAccT Executive Committee when the paper was accepted and that 2021 FAccT review did not mandate funding disclosure. These are the critics' characterisations and arguments in a peer-reviewed position paper, attributed to them; no court, regulator, board or conference process has adjudicated them. The 'audited' framing carried into the acquisition: Harver announced its acquisition of pymetrics on 11 August 2022 with terms undisclosed, describing the product as 'mitigating multiple forms of bias through an audited AI platform'.

Sources: young2022, harverviaprnewswire2022

Appears on: /domains/cases/pymetrics-cooperative-audit

EmpiricalThe audited practice migrated into a statutory channel. Under New York City Local Law 144, in effect from July 2023, the…

The audited practice migrated into a statutory channel. Under New York City Local Law 144, in effect from July 2023, the same game-based assessment platform is audited annually by a paid third-party auditor, BABL AI Inc., and the summary is published by the party deploying the tool: a V2.0 summary dated 29 June 2023, addressed to 'pymetrics inc. (Harver)', was posted by the employer Paramount, and a V1.0 summary dated 17 July 2025 for what is now Harver's Soft Skills Platform was posted by Harver. The platform retains the same three recommendation tiers and a 50th-percentile selection threshold. The 2025 summary reports every calculated impact ratio at or above 0.8, the lowest non-intersectional ratio being 0.914 for Asian candidates, computed across 164,014 male and 119,242 female candidates with known gender. It also reports what the calculation excluded: 191,455 candidates of unknown gender, 317,513 of unknown race or ethnicity — roughly twice the included known-race population — and 331,177 with at least one unknown demographic, with groups under two percent of the sample reported as not applicable. These are recorded external observations from a mandated annual audit: selection-rate ratios over candidates who disclosed demographics, not harm findings, and not evidence of platform-wide compliance. What the annual audits verify is bias-testing assertions and reported statistics, not the continued existence or exact mechanics of the pre-deployment abandon-if-noncompliant gate.

Sources: bablaiinc2023, bablaiinc2025

Appears on: /domains/cases/pymetrics-cooperative-audit

EmpiricalAn ACLU-led study published at ACM FAccT 2025 analysed all 116 publicly available Local Law 144 bias audits from July 20…

An ACLU-led study published at ACM FAccT 2025 analysed all 116 publicly available Local Law 144 bias audits from July 2023 to November 2024 and found the mandated audits to be incomplete evaluations of bias: missing demographic data, opaque aggregation, problematic test data, and metrics that do not represent how the tools are actually used. It warns of audit washing, corporate capture of auditors, and 'discrimination-hacking', records that the four-fifths rule is a rule of thumb rather than the legal standard, and shows that tools reporting four-fifths compliance could be in violation once missing-demographic impacts are considered. It cites the pymetrics cooperative audit among the proposed audit standards the field built on, and the corporate-capture critique among its warnings. The finding characterises the REGIME across all 116 audits and not this product's audits specifically, which were among the more complete in disclosing intersectional tables and unknown-group counts.

Sources: gerchick2025

Appears on: /domains/cases/pymetrics-cooperative-audit

EmpiricalThe baseline this engagement stood out against is documented in the prior literature: a survey of 18 algorithmic pre-emp…

The baseline this engagement stood out against is documented in the prior literature: a survey of 18 algorithmic pre-employment assessment vendors published at ACM FAT* 2020 found their publicly stated validation and bias-mitigation claims largely unverifiable from outside. The pymetrics engagement was the first publicly documented case of an assessment vendor opening its source code to outside auditors under pre-negotiated publication rights.

Sources: raghavan2020a, raghavan2020b, wilson2021a

Appears on: /domains/cases/pymetrics-cooperative-audit

EmpiricalThe audited pipeline ran at scale: more than 600 active client engagements between January and October 2020, covering 16…

The audited pipeline ran at scale: more than 600 active client engagements between January and October 2020, covering 16 of 23 major O*NET occupation groups, with players from 191 countries and roughly 40% of them in the United States. pymetrics also open-sourced its adverse-impact testing framework as the audit-ai library, cited in the audit paper, which implements the four-fifths rule alongside Fisher's exact test, a z-test, chi-squared and Cochran-Mantel-Haenszel tests; the repository remains public with low activity after the acquisition. No error rate, accuracy figure or defect rate for the model-generation pipeline is published in any source in this record, and no measurement compares this deployment with a human counterfactual on the same candidates.

Sources: wilson2021a, wilson2021b, pymetrics

Appears on: /domains/cases/pymetrics-cooperative-audit

EmpiricalArshon Harper, an African-American information-technology professional in Detroit, applied for approximately 150 positio…

Arshon Harper, an African-American information-technology professional in Detroit, applied for approximately 150 positions with Sirius XM Radio, LLC through the company's iCIMS-powered application platform between November 2023 and 21 November 2024, and was rejected for all but one. The single exception was a thirty-minute interview for an IT Desktop Support role in late 2023, which also ended in rejection, so the pleaded record is zero offers from about 150 applications and 149 without an interview. The complaint states that arithmetic as a 99.3 percent rejection rate and offers it as its pattern evidence. His pleaded qualification profile is a 2019 B.S. in Business Administration from Wayne State University and more than ten years of information-technology experience, including Tier 2 support to about 4,500 employees at Wayne State's computing and information technology unit between 2018 and 2021 and data-collection work at the Detroit Department of Transportation since 2008; his named targets were IT Desktop Support, Software Engineer and Technical Support Specialist. The employer is pleaded to receive thousands of applications annually. Every figure here is a pleading. It is one applicant's personal rejection rate rather than a measured disparity across applicants, no class-wide or agency-collected data exists in this record, and whether an individual probe record can support a systemic claim is precisely what the pending motion for judgment on the pleadings puts in issue.

Sources: classactioncomplaint2025, fisherphillipsllp2025

Appears on: /domains/cases/siriusxm-icims-screening

EmpiricalThe screening pipeline pleaded in the complaint, part of it drawn from the vendor's own published product documentation:…

The screening pipeline pleaded in the complaint, part of it drawn from the vendor's own published product documentation: an applicant submits a resume through the iCIMS-powered careers platform; the applicant tracking system parses it on submission and extracts name, contact information, skills, work history and education; the platform generates its own skills list from the full resume text rather than from the list the applicant wrote, and that generated list is what recruiters search when sourcing candidates; AI and machine-learning candidate matching, shortlisting and sourcing features rank or filter candidates for recruiter consumption; and rejected applicants receive automated rejections. The complaint cites iCIMS's own 'Understanding iCIMS Talent Cloud AI' documentation for the product mechanics and pleads the employer's use of the AI and machine-learning features on information and belief. As pleaded, human review sits after the algorithmic gate: recruiters see what the matcher surfaces, and 149 of the plaintiff's roughly 150 applications allegedly never reached a human decision point. No accuracy figure, error rate or check on the parse appears in any source, the applicant never sees what was extracted, and the employer's specific configuration, thresholds and human touchpoints are undiscovered.

Sources: classactioncomplaint2025, fisherphillipsllp2025, icims2026

Appears on: /domains/cases/siriusxm-icims-screening

EmpiricalThe complaint alleges that Sirius XM, by and through the iCIMS platform, evaluates applicants 'based on data points (e.g…

The complaint alleges that Sirius XM, by and through the iCIMS platform, evaluates applicants 'based on data points (e.g., educational institutions, employment history, zip codes) that proxy for race' — a triple stated in the introduction, repeated at paragraph 20 as 'data points correlated with race' and carried into Counts One and Three — and that the tools were built on historical hiring data, importing prior human selection into an automated screen and applying it at scale. It further pleads a memory mechanism inside the screen: when a Sirius XM official questioned the plaintiff's use of multiple email addresses, he explained that he had used them to circumvent suspected algorithmic penalties for repeat applications, and he pleads that this circumvention produced his only interview. The complaint alleges the tools 'intentionally and disproportionately reject African-American applicants' and are not job-related and lack business necessity. All of this is allegation, much of it expressly on information and belief; Sirius XM answered on 6 January 2026 denying liability with affirmative defenses; no court has ruled on any of it; and the EEOC expressly made no determination. Independent management-side analyses at Fisher Phillips, Metz Lewis and Lathrop GPM record the same proxy framing and the same historical-bias theory as the case's theory rather than as findings.

Sources: classactioncomplaint2025, metzlewisbrodmanmustokeefell2025, lathropgpmllp2025

Appears on: /domains/cases/siriusxm-icims-screening

EmpiricalNo employer-side bias audit, validation study or monitoring of the Sirius XM screening configuration appears anywhere in…

No employer-side bias audit, validation study or monitoring of the Sirius XM screening configuration appears anywhere in the public record reviewed, and the complaint's assertion that the tools are not job-related and lack business necessity has drawn no employer-side check to answer it. The administrative channel functioned as a pass-through rather than a merits check: the EEOC charge, No. 520-2025-01266, was filed 22 November 2024, and on 6 May 2025 the Commission's Newark Area Office issued a Determination and Notice of Rights stating that it 'will not proceed further with its investigation and makes no determination about whether further investigation would establish violations' — about five and a half months from charge to closure, with no view of the deployment recorded either way. Suit followed within the ninety-day window. The would-be strongest instrument, discovery into how the tool actually behaves, has not opened, and no independent evaluation of this deployment exists. The vendor's published programme describes processes for its product line and discloses no audit results.

Sources: classactioncomplaint2025, pacermonitor2026, icims2026

Appears on: /domains/cases/siriusxm-icims-screening

EmpiricalHarper v. Sirius XM Radio, LLC, No. 2:25-cv-12403, was filed on 4 August 2025 in the United States District Court for th…

Harper v. Sirius XM Radio, LLC, No. 2:25-cv-12403, was filed on 4 August 2025 in the United States District Court for the Eastern District of Michigan, Southern Division, before Judge Terrence G. Berg with referral to Magistrate Judge Anthony P. Patti. It pleads three counts — Title VII disparate treatment under 42 U.S.C. 2000e-2(a), Title VII disparate impact under 2000e-2(k), and intentional race discrimination under 42 U.S.C. 1981 — on behalf of a proposed class of all African-American individuals who applied for employment with Sirius XM through its iCIMS platform since 27 January 2024 and were rejected or screened out, seeking Rule 23(b)(2) and/or (b)(3) certification or issue certification under Rule 23(c)(4). The relief sought includes a declaration, a permanent injunction prohibiting continued discrimination and requiring reforms, backpay, front pay, and compensatory and punitive damages. The complaint was signed by Harper himself with Winston Cooks, LLC of counsel signing the civil cover sheet; the $405 filing fee went unpaid until 1 October 2025; plaintiff-side appearances followed on 26 September and 8 October 2025 and defense appearances on 4 and 5 November 2025. On 6 January 2026 Sirius XM answered with affirmative defenses and moved for judgment on the pleadings under Rule 12(c) the same day; the motion was fully briefed on 12 March 2026 after an opposition on 26 February, and was noticed for determination without oral argument. On the docket record verified 28 August 2026 there is no ruling of any kind, no class-certification motion and no settlement. That record is an aggregator's transcription rather than the court's own system, which was not reachable from the verifying environment, so a ruling issued after its last crawl cannot be fully excluded.

Sources: classactioncomplaint2025, pacermonitor2026

Appears on: /domains/cases/siriusxm-icims-screening

EmpiricalThe structural feature of this case is that the deployer stands alone. Sirius XM Radio, LLC is the sole defendant for ou…

The structural feature of this case is that the deployer stands alone. Sirius XM Radio, LLC is the sole defendant for outcomes its licensed third-party tool is alleged to produce; iCIMS, Inc. is not a party, is not accused as a party, and faces no identified enforcement action. Independent legal analyses frame the case as the employer-side companion to the vendor-side litigation over a different applicant-screening platform, in which the vendor is the defendant under a theory that it acts as the employers' agent — that action's second motion to dismiss was denied in July 2024 and a nationwide collective was conditionally certified in May 2025. The proposition Harper is understood to test is the ordinary employer-liability route: as one analysis puts it, employers cannot contract out or bypass their Title VII liability simply by outsourcing the technology. The two cases answer the same domain question from opposite seams — race under two statutes with a single-firm class here, age under one statute with a cross-employer collective there — and nothing about either case's posture is evidence about the other.

Sources: metzlewisbrodmanmustokeefell2025, lathropgpmllp2025, fisherphillipsllp2025

Appears on: /domains/cases/siriusxm-icims-screening

EmpiricalVENDOR CLAIMS ONLY, and none of it is tested by this litigation or is evidence about the Sirius XM configuration. iCIMS,…

VENDOR CLAIMS ONLY, and none of it is tested by this litigation or is evidence about the Sirius XM configuration. iCIMS, Inc. publishes an account of a responsible-AI program built on six pillars — human-led, transparent, private and secure, inclusive and fair, technically robust and safe, and accountable — describing processes for bias audits, transparency reporting and human oversight that reference New York City Local Law 144's automated-employment-decision-tool obligations, alignment claims to the National Institute of Standards and Technology's AI Risk Management Framework, the OECD AI Principles and ISO 42001, and a 'human in the loop at appropriate times' positioning. It states that it obtained TrustArc's TRUSTe Responsible AI certification in March 2025 and claims to be the first enterprise recruiting software provider to do so. The blog discloses no audit results. The litigation record contains no employer-side audit of this deployment, and no source connects the vendor's programme to what this employer configured or ran. No parallel litigation over this platform and no enforcement action against the vendor was identified as of the run date, which is stated as none identified rather than none existing.

Sources: icims2026

Appears on: /domains/cases/siriusxm-icims-screening

EmpiricalCommunity Notes, launched as Birdwatch in January 2021 and deployed worldwide from 11 December 2022, is a crowd-annotati…

Community Notes, launched as Birdwatch in January 2021 and deployed worldwide from 11 December 2022, is a crowd-annotation system on X in which volunteer contributors write contextual notes on specific posts and other contributors rate those notes, with an open-source matrix-factorization model deciding which notes appear. An independent parse of the operator's own daily public corpus for 23 January 2021 to 23 January 2025 counts 227,702 unique contributors writing 1,614,743 notes on 1,016,673 distinct posts, detected in 103 languages, with the program available in more than 60 countries. Participation is highly unequal and largely monolingual: the top 10 percent of contributors wrote 58 percent of all notes, a Gini coefficient of 0.68; one apparently automated account wrote 33,186 notes; and only about 16 percent of note authors ever wrote in more than one language. Contributor admission is by three published criteria — an account at least six months old, a verified phone number from a trusted carrier not associated with another Community Notes account, and no recent notice of violations of the platform's rules — with random selection from country-specific waitlists where applicants exceed available slots. All contributions are pseudonymous and publicly visible. Around 22 to 26 May 2025 notes stopped appearing in users' feeds for several days following a data-centre outage, with the operator's engineering account acknowledging continuing issues and the Community Notes account stating on 26 May that it was working to get notes appearing normally.

Sources: mohammadi2026, razuvayevskaya2025, thejournal2025

Appears on: /domains/cases/community-notes-x

EmpiricalThe enforcement action here is an addition, not a subtraction, and the operator's own documentation states it. Its publi…

The enforcement action here is an addition, not a subtraction, and the operator's own documentation states it. Its published FAQ says notes rated helpful by enough contributors from different points of view will appear directly on posts, and that beyond that, notes do not affect display of posts or enforcement of X's Rules. Its introduction page states that X does not write, rate or moderate notes, except where a note itself violates the platform's rules, and that the mechanism does not work by majority rules: a note requires agreement between contributors who have sometimes disagreed in their past ratings. Nothing is removed, downranked or restricted by the note itself, and the entire measured effect runs through readers changing their own behaviour. There is no appeal body, no human reviewer of last resort and no escalation queue; a post author's only documented channel is to request additional review of a note or report it, and the remedy for a wrong or missing note is more ratings rather than adjudication. One documented consequence sits outside the display channel and in tension with the operator documentation: on 29 to 30 October 2023 the platform's owner announced that posts carrying a Community Note become ineligible for creator ad-revenue sharing, framing it as maximizing the incentive for accuracy over sensationalism and asserting that attempts to weaponize notes to demonetize people would be immediately obvious because the code and data are open. Both statements are carried here and the tension between them is flagged rather than resolved. Claims of large algorithmic reach penalties for noted posts appear only on low-tier marketing pages, are supported by no credible source, and are not asserted anywhere in this atlas.

Sources: xcorpc, bell2023, renault2024

Appears on: /domains/cases/community-notes-x

EmpiricalA note is displayed only when a published consensus gate is cleared, and most notes never clear it. The operator's ranki…

A note is displayed only when a published consensus gate is cleared, and most notes never clear it. The operator's ranking documentation gives the rule: a matrix factorization fits a global intercept, a per-rater intercept and factor, and a per-note intercept and factor; intercept terms are regularized at 0.15 against 0.03 for the factor terms, five times more strongly, which the documentation says is what requires that notes are rated by raters with diverse factors before a note gets a label. A note is Currently Rated Helpful at intercept 0.40 or above with latent factor magnitude under 0.50, raised to 0.50 by a tag-outlier filter; at least five ratings are needed to leave the Needs More Ratings state; the model is re-trained from scratch every hour, with status changes deliberately delayed so as not to influence independent raters. Measured against the whole public corpus for 2021 to 2025, 87.7 percent of notes remained in Needs More Ratings and 8.3 percent reached Helpful, and only 13.55 percent of posts with at least one proposed note ever received a Helpful note. Independent samples put the note-level rate at 11.3 percent, at 11.5 percent, and at about 10 percent and declining, and a newspaper analysis counted roughly 79,000 of more than 900,000 notes written in 2024 shown publicly, under 9 percent; these denominators are notes rather than posts and are not interchangeable. The gate is weakest where the stakes are highest. The Center for Countering Digital Hate reported on 30 October 2024 that 209 of 283 sampled misleading US-election posts, 74 percent, had accurate notes that were never shown to all users; that is advocacy-organisation research on a purposive, non-random sample rather than a population estimate, its publisher's site is captcha-blocked to automated requests, and the figure is carried through trade coverage quoting it verbatim. An archival analysis of more than 1.8 million notes finds 69 percent of noted posts receiving conflicting classifications from contributors, and about 68 percent annotated as not needing a note at all. A regression-discontinuity analysis finds that having a note published raises the retention of first-time contributors, so a low and declining publication rate erodes the contributor population that would raise it.

Sources: xcorpa, mohammadi2026, razuvayevskaya2025, arjmandilari2025, centerforcounteringdigitalha2024

Appears on: /domains/cases/community-notes-x

EmpiricalThe deployment's documented failure mode is timing, and there is no single latency figure for it. Published central esti…

The deployment's documented failure mode is timing, and there is no single latency figure for it. Published central estimates of the delay between a post and a note appearing beneath it are: a mean of 15.5 hours and a median of 14.3 hours in a 2021-2023 difference-in-differences sample; quartile boundaries of 12, 23 and 47 hours from post creation to note attachment in a March-June 2023 synthetic-control sample, implying a median near 23 hours; means of 2.85 days after the US rollout and 2.23 days after the worldwide rollout, with the shortest display delay observed anywhere in that dataset being 80.2 minutes; an average of 26 hours in a whole-corpus parse; and a mean of 65.7 hours across 1.8 million notes. These differ by clock — post to note creation, post to display, or note creation to first status — by sample, and by period, and each use must carry its clock and its sample or state the range. Against that clock, roughly 50 percent of a post's reposts occur in its first 5 hours and 80 percent within 16, and the median half-life of a post's impressions is about 79.5 minutes. The dose-response is measured and monotone: notes attached in the 1-12, 12-23, 23-47 and 47-plus hour quartiles are associated with lifetime repost reductions of 24.9, 12.3, 4.3 and 0.1 percent, the last statistically indistinguishable from nothing, and in that slowest quartile view growth rose 13.6 percent and reply growth 27.0 percent, consistent with a late note drawing attention back to a stale post. The earliest rigorous evaluation, a difference-in-differences and regression-discontinuity analysis of the US and worldwide roll-outs, found no evidence that introducing Community Notes reduced aggregate engagement with misleading posts and attributed the null to display latency exceeding the diffusion half-life.

Sources: renault2024, slaughter2025, chuai2024, mohammadi2026, razuvayevskaya2025

Appears on: /domains/cases/community-notes-x

EmpiricalOnce a note is attached the mechanism works, and every effect figure must carry both its conditional and its uncondition…

Once a note is attached the mechanism works, and every effect figure must carry both its conditional and its unconditional counterpart because the three causal studies measure different estimands and their headline numbers must never be pooled. A synthetic-control study of 40,078 posts for which notes were proposed between 16 March and 23 June 2023, of which 6,757 (16.9 percent) received a Helpful note, estimated post-attachment GROWTH reductions of 46.1 percent in reposts, 44.1 percent in likes, 21.9 percent in replies and 13.5 percent in views, and WHOLE-LIFESPAN reductions of 11.6, 13.3, 6.9 and 5.5 percent respectively; it also found noted content's repost cascades becoming less deep and less structurally viral than matched counterfactuals, and it measured no deletion outcome at all. An independent difference-in-differences study of 237,180 fact-checked cascades reposted more than 431 million times estimated a 61.2 percent reduction in subsequent spread and a 94.3 percent increase in the ODDS that the author deletes the post, against a SYSTEM-WIDE reduction of 14.9 percent in total engagement with misleading posts, stating that notes often appear too late to intervene in the early and most viral stage of diffusion, and reporting the effect as significantly weaker for posts from influential accounts and for political content. A working paper on about 285,000 notes reports 49.1 percent fewer retweets by difference-in-differences and 52.4 percent by pre-treatment outcome matching, with overall reductions of 16.34 percent in retweets, 11.75 percent in replies and 16.87 percent in quotes once publication delay is accounted for, and a deletion rate of 15.8 percent just above the 0.4 helpfulness threshold against 8.6 percent just below it — a relative gap across a threshold, not a causal probability increase. The notes themselves are usually right: in a randomly sampled set of notes on popular COVID-19 vaccine posts evaluated with an infectious-disease physician and a virologist, 97.5 percent were entirely accurate, 2 percent partially accurate and 0.5 percent inaccurate, with 49 percent citing highly credible sources and 44 percent moderately credible ones. That accuracy measurement covers one topic in one period and says nothing about the notes that were never displayed. Operator-reported pilot figures — that people who saw bridging-selected annotations were 25 to 34 percent less likely to like or repost, and 20 to 40 percent less likely to agree with the substance of a potentially misleading post — are vendor-tier claims from platform-run experiments and are superseded for causal purposes by the independent studies.

Sources: slaughter2025, chuai2026, renault2024, allen2024a

Appears on: /domains/cases/community-notes-x

EmpiricalThe decision rule and the whole operating record are public, and that is why an independent evidence base for this deplo…

The decision rule and the whole operating record are public, and that is why an independent evidence base for this deployment exists. The scoring code, its documentation and its note-writer interface template sit in a public repository under the Apache-2.0 licence, and five data files — notes, ratings, note status history, contributor enrolment and note requests — are released daily on a best-effort basis, cumulative, containing only items created up to 48 hours before release, with each participant carrying a Community-Notes-specific pseudonymous identifier that is stable across handle changes; deleted content disappears from the downloads while the note status history retains participant identifiers and status timelines. Three independent research teams produced causal estimates of the system's effects from exactly that data without the operator's permission, a civil-society organisation measured coverage failure on a purposive sample of election claims, a university team parsed four years of the corpus, and an agent-based study probed the manipulation surface of the deployed rule itself rather than a guess at it, reporting that under polarization and in-group rating preference the published scorer suppresses a substantial fraction of genuinely helpful notes and that a coordinated minority of 5 to 20 percent of raters could strategically suppress targeted helpful notes — a simulation on synthetic data, not a measurement that manipulation occurred. The operator's own published list of challenges names four risks it watches: coordinated manipulation as a crucial risk for open rating systems, outcomes dominated by a simple majority or biased by the distribution of contributors, harassment of contributors, and rater burden from high volumes of low-quality notes. Speed and coverage — the two properties every independent measurement identifies as the deployment's failure mode — are not among them. The openness is also what made the mechanism copyable: on 7 January 2025 Meta announced it would end third-party fact-checking in the United States in favour of a community-notes-style model, and on 18 March 2025 it began testing on Facebook, Instagram and Threads with about 200,000 signed-up contributors, stating it would use X's open source algorithm as the basis of its rating system and that, unlike the fact-check labels it replaced, notes would provide extra context but would not impact who can see the content or how widely it can be shared. That adoption is context about another company and is described as begun and under review rather than complete or global.

Sources: xcorp, xcorpb, truong2025, kaplan2025

Appears on: /domains/cases/community-notes-x

EmpiricalSince 1 July 2025 the writing side of the mechanism has admitted automated agents while the deciding side has stayed hum…

Since 1 July 2025 the writing side of the mechanism has admitted automated agents while the deciding side has stayed human. The operator's published interface documentation states that automated writers propose notes while humans still decide what is helpful enough to show, and that ratings come from regular contributors, that is humans, whose input ultimately determines which notes show. Admission is earned in test mode against an automated evaluator that screens URL validity and whether a note addresses a claim rather than an opinion: at least 95 percent of the candidate's most recent 50 test notes must score high on URL validity and at least 30 percent high on the claim-versus-opinion measure. Daily writing limits start at 10 and scale between 2 and 500 with measured helpfulness. A think-tank evaluation computed from the operator's public downloads for September 2025 to March 2026 counts 27 automated writer accounts enrolled and 24 active, submitting 31,464 notes: 7.4 percent of note volume but 13.9 percent of displayed notes, reaching Currently-Rated-Helpful at 18.0 percent against 8.9 percent for human writers, rated Not Helpful at 2.3 percent against 4.1 percent, and consuming roughly 304 ratings per displayed note against 908 for human-written ones, with monthly output growing from 93 notes in September to 8,109 in February. Its median time-to-verdict of 6.0 hours against 6.3 for humans is measured from note creation to first non-Needs-More-Ratings status and is therefore not comparable to the post-to-display latencies measured elsewhere. The evaluation explicitly measured no engagement outcome, so nothing follows from it about whether automated notes reduced spread, and it records that the enrolment criteria for automated accounts are undisclosed. The operator's interface documentation does not state that automated notes are visibly labelled to readers, although press coverage of the July 2025 launch reported that they would be; that labelling is therefore press-attributed rather than operator-confirmed.

Sources: xcorpd, purnell2026

Appears on: /domains/cases/community-notes-x

EmpiricalOne regulator has named this mechanism and none has published a finding on it. The European Commission's first formal pr…

One regulator has named this mechanism and none has published a finding on it. The European Commission's first formal proceedings under the Digital Services Act, opened 18 December 2023 against a platform designated a Very Large Online Platform on 25 April 2023 with 112 million monthly active users in the EU, list among their grounds the effectiveness of measures taken to combat information manipulation on the platform, notably the effectiveness of X's so-called Community Notes system in the EU and the effectiveness of related policies mitigating risks to civic discourse and electoral processes, against Articles 34(1), 34(2) and 35(1). That limb remains open. The Commission's preliminary findings of July 2024 were limited to the blue-checkmark design, the advertisement repository and researcher data access, so no preliminary finding has ever issued on the Community Notes limb; and the Commission's first non-compliance decision against the platform, a 120 million euro fine of 5 December 2025, rests on those same three grounds and concerns this mechanism nowhere. That fine must never be attached to Community Notes. A Commission spokesperson declined to comment on the mechanism in May 2025 because of the ongoing proceedings while confirming the investigation into its effectiveness. No litigation of any kind concerns this deployment. On 26 March 2026 the Oversight Board issued a policy advisory opinion on Meta's plans to expand community notes beyond the United States, concluding that delays in note publication, the limited number of published notes and the model's dependence on the broader information environment's reliability raise serious doubts about the extent to which community notes can meaningfully address misinformation linked to harm, and recommending that countries with repressive human-rights records, imminent major elections, histories of coordinated disinformation networks, active crises or conflicts, unsupported language complexity or persistent internet-access obstacles be omitted or delayed. That opinion is non-binding, it is that company's own body, it addresses that company's expansion plans rather than X's system, and its evidence is X's published performance figures. It is context here and is not a ruling about this deployment.

Sources: europeancommission2023, europeancommission2025, thejournal2025

Appears on: /domains/cases/community-notes-x

EmpiricalThe GIFCT Hash-Sharing Database is a cross-company index of terrorist and violent extremist content: a member platform t…

The GIFCT Hash-Sharing Database is a cross-company index of terrorist and violent extremist content: a member platform that has found such material on its own service judges it against that platform's own policy, judges it a second time against the Global Internet Forum to Counter Terrorism's taxonomy, converts it to a perceptual hash, attaches labels and publishes the hash to a shared store that every other integrated member queries against its own uploads. On GIFCT's own figures in its 2025 Annual and Transparency Report, the store held approximately 2.4 million hashes at the end of 2025 — an increase of about 123,500 during the year — covering approximately 408,000 unique and distinct items, comprising about 329,000 visually distinct images, 79,000 visually distinct videos and 200 textually distinct PDF items. Membership stood at 39 platforms with 23 named as integrated or integrating, against 12 companies with access in the 2022 report and four founders in 2017 (Facebook, Microsoft, Twitter and YouTube); GIFCT counts over 5 billion net monthly active users across all members. Access is gated by GIFCT membership plus a signed information-sharing agreement plus the hash-sharing database code of conduct, and GIFCT stated in 2022 that governments and other non-tech company organizations do not have access to the database. There are three inclusion pathways: association with an entity on the United Nations Security Council 1267 Consolidated Sanctions List, satisfaction of the behavioural inclusion criteria added by the July 2021 taxonomy expansion, or an activation of the Incident Response Framework. GIFCT operates no platform, holds no source content, and states that it does not own or store any source data or personally identifiable information of users associated with member platforms. Every quantitative figure here is GIFCT's own and none has been independently verified.

Sources: globalinternetforumtocounter2026b, globalinternetforumtocounter2025, globalinternetforumtocounter2022, globalinternetforumtocounter2026a

Appears on: /domains/cases/gifct-hash-sharing

EmpiricalOnly the contributing member may remove its own hash from the GIFCT Hash-Sharing Database. GIFCT publishes four removal …

Only the contributing member may remove its own hash from the GIFCT Hash-Sharing Database. GIFCT publishes four removal grounds: the contributing platform's own review, another member's feedback, new information or evolving context, and data availability — the case where the contributor no longer retains the underlying content and so can no longer verify the hash. GIFCT itself may create hashes for inclusion and may add an alternative opinion to a record. In the Year 4 Hash Sharing Working Group review published in December 2024, Dr Sean Doody and Dr Michael Jensen of the National Consortium for the Study of Terrorism and Responses to Terrorism recorded that GIFCT 'is only allowed to add additional alternative opinions to records and lacks the ability to modify the labels added to hashed content by members' and that 'as it currently stands, GIFCT itself cannot directly remediate labeling mistakes for hashes submitted to the HSDB'. They recommended that GIFCT 'should be endowed with the proper authority to directly audit, quality control, validate, and fix labeling errors', noting that this 'would almost certainly require members to provide GIFCT access to pre-hashed content'. GIFCT's 2025 Annual and Transparency Report does not record that authority being granted. That review is independent in authorship and was commissioned, framed, hosted and published by GIFCT through its own working-group programme, and its authors had no access to the hashed content.

Sources: doody2024, globalinternetforumtocounter2026b, globalinternetforumtocounter2025

Appears on: /domains/cases/gifct-hash-sharing

EmpiricalIndependent review of the GIFCT Hash-Sharing Database is architecturally obstructed rather than merely withheld, and thr…

Independent review of the GIFCT Hash-Sharing Database is architecturally obstructed rather than merely withheld, and three separate sources say so in nearly the same terms. Gavin Sullivan, writing in the London Review of International Law in 2025, states that 'Despite widespread agreement that third-party reviews of the hash-sharing database are necessary, it is not yet clear how such reviews can be carried out given the hash-sharing database itself has no content', and records that the member disagreement mechanism 'is a means for platforms to signal disagreement with a hash's inclusion in the database' and is 'only open to GIFCT members', and that GIFCT participants 'consider themselves one step removed from the human rights impact' of the system. Courtney Radsch wrote for Just Security on 30 September 2020 that 'none of the associated content is available for independent review or audit, either by regulators or researchers'. GIFCT gave the same reason from the inside in its 2022 Transparency Report, explaining that it 'is neither a tech company nor a social media platform and does not have any access to source content to determine what the hash corresponds to' — which is why the only quality review ever published was conducted by members on their own submissions. In that 2022 exercise, members randomly sampled three strata of hashes they had themselves submitted (sanctions-list-derived, incident-derived, and hashes carrying another member's disagreement feedback), reported no significant quality errors and no accuracy difference between strata, and found 'a very small number of hashes were incorrectly labeled and an even smaller number did not meet the taxonomy for inclusion', with labels corrected and out-of-scope hashes removed; no denominator was given. Two commissioned reviews exist — a Business for Social Responsibility human rights impact assessment reviewed between December 2020 and May 2021 and published in July 2021, and the December 2024 database review — and neither had access to the hashed material. No false-positive rate, false-negative rate, precision or recall has ever been published for the index or for any member's matcher.

Sources: sullivan2025, radsch2020, globalinternetforumtocounter2022, doody2024

Appears on: /domains/cases/gifct-hash-sharing

EmpiricalThe largest category in the GIFCT Hash-Sharing Database is also its least determinate, and the labels the index depends …

The largest category in the GIFCT Hash-Sharing Database is also its least determinate, and the labels the index depends on are documented as unreliable by GIFCT's own reviewers. GIFCT defines 'Glorification of Terrorist Acts' as content that 'glorifies, praises, condones, or celebrates attacks after the fact'. As a share of BEHAVIOURALLY LABELLED hashes it stood at 75.62 percent at the end of 2025 and 75.94 percent at the end of 2024, with graphic violence against defenceless people at 16.02 and 16.65 percent, recruitment and instruction at 6.29 and 5.27 percent, and imminent credible threat at 2.07 and 2.14 percent; approximately 92 percent of hashes carried behavioural labels at end-2025 and approximately 91 percent at end-2024. The 2022 and 2023 reports state their shares on a DIFFERENT basis — of total hashes — at 65.23 and 62 percent for glorification, so the series is not comparable across that boundary without stating the denominator; normalised to the labelled subset the recent years run 78.0, 72.1, 75.94 and 75.62 percent. Angel Diaz of the Brennan Center for Justice read GIFCT's first transparency report in September 2019 as showing 85.5 percent glorification against 0.4 percent imminent credible threats, and criticised 'glorification', 'praise' and 'condone' as 'notoriously imprecise' terms that 'will almost inevitably capture expressions of general sympathy or an understanding for certain viewpoints, not to mention news reporting', alongside the absence of appeals processes, redress mechanisms and third-party audits assessing error rates. The commissioned December 2024 review reports an internal GIFCT finding that 'content in the HSDB frequently lacks labels or is sometimes labeled inconsistently or incorrectly' and that 'labeling errors have accumulated'; that it was 'especially difficult to determine if content advocates for, or is making a call to, violence'; that more clarity was needed on what constitutes a hate-based ideology; that ideology labels were never made mandatory and 'are missing for most hashes'; and that members 'primarily share TVEC from designated entities', with a small number of the largest members responsible for most of the activity and some members contributing nothing, because improving the representativeness of the database 'is not currently a priority for them'.

Sources: globalinternetforumtocounter2026b, globalinternetforumtocounter2025, globalinternetforumtocounter2024, globalinternetforumtocounter2022, diaz2019, doody2024

Appears on: /domains/cases/gifct-hash-sharing

EmpiricalA hash in the GIFCT Hash-Sharing Database compels no action anywhere. GIFCT states in its 2024 Annual and Transparency R…

A hash in the GIFCT Hash-Sharing Database compels no action anywhere. GIFCT states in its 2024 Annual and Transparency Report that 'Adding hashes does not prompt any direct or automatic action on another member's platform, such as removing content. Each member can use the hashes provided through the HSDB to identify content on their respective platform that matches known terrorist or violent extremist content. Each member also independently determines what potential action to take.' The flow diagram in its 2025 report routes a match to 'Platform B flags the content for human review' and then to 'Platform B can confirm the hash is a match and takes action against the content in line with its own policies'. The modelled pipeline is therefore a contributor's moderation decision, a hash with labels, the shared store, a consuming member's automated comparison, that member's human review, and that member's enforcement — two independent human policy judgements in two different companies bracketing one automated comparison. Matching is automatic; action is not. GIFCT's own reviewers add that 'The HSDB is not meant to be the final authoritative source on what constitutes TVEC, and tech companies are always free to remove content according to their own moderation' policies, and GIFCT states that its taxonomy 'represents a selection of high-severity content that seeks to capture areas of strong consensus among members' and can be 'more limited than individual member company's policies'. The accurate structural claim is that one company's classification automatically enters every other member's review queue, not that it removes anything.

Sources: globalinternetforumtocounter2025, globalinternetforumtocounter2026b, doody2024, globalinternetforumtocounter2022

Appears on: /domains/cases/gifct-hash-sharing

EmpiricalNo published channel connects a person whose content was matched to the GIFCT Hash-Sharing Database. GIFCT's membership …

No published channel connects a person whose content was matched to the GIFCT Hash-Sharing Database. GIFCT's membership criteria require each member to have 'the ability to receive, review, and act on reports of activity that is illegal and/or violates terms of service and user appeals', so every member runs its own user-appeal process — and that appeal reaches the acting platform's own policy decision, not the shared index. The only disagreement mechanism that touches the index runs between companies: members may indicate agreement or disagreement with a hash's labelling and inclusion, with all feedback visible to GIFCT and to participating members. Uptake was 34,014 hashes across approximately 9,000 distinct items, or 1.63 percent, at the 2022 report, and approximately 2 percent in each report since; in the 2023 breakdown the vast majority was agreement, 5 percent of feedback-carrying hashes disagreed about descriptive labels while still agreeing the item belonged, and disagreement that the content met the taxonomy at all ran to less than 0.01 percent of feedback-carrying hashes. GIFCT warns that this feedback 'should be treated with caution and not be considered statistically significant'. Gavin Sullivan records that the mechanism is 'only open to GIFCT members'. As fetched on 28 August 2026, GIFCT's public hash-sharing database explainer page describes no appeal, review or redress process for a user whose content is hashed and publishes no participating-company list, and GIFCT's Human Rights Policy update of 18 May 2026 describes due-diligence tooling and Independent Advisory Committee oversight without describing any remedy, grievance or appeals mechanism for affected users. The Global Network Initiative argued in October 2025 that individuals 'should be able to challenge wrongful takedowns or account suspensions and have their content re-evaluated' and that oversight of such databases 'should be independent, with stakeholder participation from affected communities, researchers, and human rights bodies'. Against this sit GIFCT's own mitigations, each with its documented reach: member-side appeal duties reach the acting platform, the feedback channel reaches other members, the commissioned reviews reach GIFCT's governance, and the Independent Advisory Committee advises without operating the database.

Sources: globalinternetforumtocounter2025, globalinternetforumtocounter2026b, globalinternetforumtocounter2026a, globalinternetforumtocounter2026, sullivan2025, globalnetworkinitiative2025

Appears on: /domains/cases/gifct-hash-sharing

EmpiricalCivil society has measured lawful documentation of violence disappearing from platforms at scale, and none of it is attr…

Civil society has measured lawful documentation of violence disappearing from platforms at scale, and none of it is attributed to the GIFCT Hash-Sharing Database. Human Rights Watch reported on 10 September 2020 that it had reviewed 5,396 pieces of content cited in 4,739 of its own reports since 2007 and found 619 of them — 11 percent — removed, and that the Syrian Archive found 361,061 of the YouTube videos it had preserved, 21 percent, no longer accessible; the Brennan Center recorded that over 100,000 of the Syrian Archive's videos were removed from YouTube through the use of automated tools. Human Rights Watch also recorded that YouTube removed 6.1 million videos in the first quarter of 2020 with 49.9 percent taken down before any user saw them, and that Facebook's automated systems flagged 99.3 percent of terrorist-propaganda content before any user report; and it recorded that civil society could not establish what the shared database contained — then over 300,000 unique hashes as of July 2020 — or whether its contents matched any individual platform's definition of terrorism. Every one of those removal figures measures the platform-side automated removal environment that the shared index feeds. Not one of them is attributed to a hash match, and no such attribution exists anywhere in the public record: GIFCT publishes no count of content removed, demoted or blocked because of a match, platform appeal statistics do not separate hash-matched actions, and no mechanism exists by which a person whose content was matched learns that a shared index was involved. GIFCT's own commissioned reviewers took the position in December 2024 that bystander, survivor and journalistic footage of an attack is out of scope for hashing, having 'no core hate-based ideology or extremist identifier associated with the producer of the content'; that position is a reviewer's recommendation rather than a published change to the inclusion criteria, and no mechanism exists to check whether it is followed.

Sources: humanrightswatch2020, diaz2019, doody2024, globalinternetforumtocounter2026b

Appears on: /domains/cases/gifct-hash-sharing

EmpiricalGIFCT reported in its 2025 Annual and Transparency Report that 'X (formerly Twitter), a founding member of GIFCT, conclu…

GIFCT reported in its 2025 Annual and Transparency Report that 'X (formerly Twitter), a founding member of GIFCT, concluded its GIFCT membership in 2025 to focus on its internal trust and safety efforts'. X appears among the hash-sharing-database-integrated members listed in the 2024 report and is absent from the 2025 report's list. That statement is GIFCT's characterisation of a member's decision, and nothing published says what became of the hashes that member had already contributed — a question that matters because only a contributing member may remove its own hashes, so a departure leaves the entries in place with no party identified as able to correct them. Governance moved in the other direction over the same period: GIFCT's 2025 report names Meta, Microsoft and YouTube as holding the founding Operating Board seats, with Discord and Twitch elected to two at-large seats for 2026 — the first non-founder board seats — which GIFCT attributes to recommendations in the 2021 human rights impact assessment it commissioned from Business for Social Responsibility. Membership grew from 33 platforms at the end of 2024 to 39 at the end of 2025, and GIFCT activated its Incident Response Framework fourteen times across seven countries in 2025 against seven times in 2024. GIFCT is not a regulated entity anywhere: no regulator supervises the consortium as such, and the instruments that bind — Regulation (EU) 2021/784, applicable from 7 June 2022 and requiring hosting service providers to remove terrorist content within one hour of a national authority's removal order, the EU Digital Services Act, and the UK Online Safety Act — fall on member platforms individually. The Christchurch Call, launched on 15 May 2019 and supported by 55 governments plus the European Commission and 19 online service providers, names 'the expansion and use of shared databases of hashes and URLs' among its industry commitments and is non-binding.

Sources: globalinternetforumtocounter2026b, globalinternetforumtocounter2025, businessforsocialresponsibil2021, europeancommission2021, christchurchcall2019

Appears on: /domains/cases/gifct-hash-sharing

EmpiricalIn February 2021 a father in San Francisco photographed his toddler son's swollen, painful groin because an advice nurse…

In February 2021 a father in San Francisco photographed his toddler son's swollen, painful groin because an advice nurse asked for images ahead of an emergency telehealth consultation during pandemic-era remote care; the doctor used the photographs to diagnose the infection and prescribed antibiotics, which cleared it up. The images auto-uploaded from an Android phone to Google Photos, and two days later his entire Google Account was disabled for 'harmful content' that was 'a severe violation of the company's policies and might be illegal'. Google reported the material to the National Center for Missing & Exploited Children's CyberTipline, and San Francisco police served search warrants on Google and on his internet service provider within a week of the photographs, seeking 'everything in Mark's Google account: his internet searches, his location history, his messages and any document, photo and video'. He learned of it in December 2021, when an envelope arrived containing the warrants and a letter telling him he had been investigated. THE POLICE CLEARED HIM: the investigator, with access to everything Google held, concluded that 'the incident did not meet the elements of a crime and that no crime occurred'. A near-identical case ran in parallel in Houston, where a father photographed his toddler's genital infection at a pediatrician's request, the images auto-backed up and were sent to his wife over a Google messaging service, and his decade-old paid account was locked while he was in the middle of buying a house; he too was cleared, quickly, after showing a detective his correspondence with the pediatrician. Both men appealed with the exculpatory material in hand: 'A few days after Mark filed the appeal, Google responded that it would not reinstate the account, with no further explanation.' Google publicly stood by the decisions and has never conceded error in either case. Asked directly in December 2025, the reporter who broke the story said neither parent had recovered his account, though one had been able to retrieve some account data that was turned over to police; that answer is relayed second-hand by the writer who asked her, and no first-party statement of the outcome exists.

Sources: hill2022, bhuiyan2022, mullin2022, gizmodo2022, heer2025

Appears on: /domains/cases/google-csam-account-closure

EmpiricalGoogle describes its child-safety detection as two technologies used in combination and augmented by human review. Hash …

Google describes its child-safety detection as two technologies used in combination and augmented by human review. Hash matching compares uploads against verified sets of previously confirmed material, with hashes drawn from the Internet Watch Foundation, the National Center for Missing & Exploited Children and content Google itself confirms, each independently verified before deployment. Separately, machine-learning classifiers trained on confirmed material 'flag new content that is very similar to patterns of previously confirmed CSAM'. Specialist reviewers with backgrounds in law enforcement, child advocacy and social work confirm both hash matches and classifier hits before action. THE DOCUMENTED FLAGS CAME FROM THE CLASSIFIER, NOT FROM HASH MATCHING: the photographs in both 2021 cases were newly created, so no hash of them could have existed, and the New York Times reported that the images were flagged by the artificial intelligence and that 'a human content moderator for Google would have reviewed the photos after they were flagged by the artificial intelligence to confirm they met the federal definition of child sexual abuse material'. Two trade outlets attribute the flag to Microsoft PhotoDNA hash matching, one of them in its own headline slug, and that attribution is wrong for these photographs. Google's spokesperson line describing 'a combination of hash matching technology and artificial intelligence' describes the whole system, not these flags. The same classifier layer is distributed to other platforms through Google's Child Safety Toolkit: a Content Safety API that prioritises never-before-seen images and video for partner review and a video hash-matching tool, with Google stating that partners 'process billions of files' and naming NCMEC, Adobe, Yahoo, Nextdoor, Scribd and Substack among them, and the 2023 reporting adding that the new-material classifier was made available to other companies including Meta and TikTok. Volume, from different periods and counting bases that are not chained: over 600,000 CyberTipline reports and more than 270,000 accounts disabled in 2021; over one million reports in the first half of 2022; more than two million across 2022; 1,470,958 reports in 2023 as counted by NCMEC, about 4 percent of all platform tips; and approximately 270,000 accounts suspended for this ground annually as an operator round number.

Sources: jasper2022, hill2022, google, hill2023, grossman2024

Appears on: /domains/cases/google-csam-account-closure

EmpiricalNo error rate exists for the detection channel that produced the documented cases, and the reason is a definition rather…

No error rate exists for the detection channel that produced the documented cases, and the reason is a definition rather than an omission. Google's regulated filings under Regulation (EU) 2021/1232 publish measured error counts for hash matching alone: 18 content items incorrectly flagged in 2023 and 10 in 2024, in both years caught by human review during detection with nothing removed, nothing reported externally and no account access lost; the 2025 filing reports 1,604 items automatically flagged as known material, 335 subject to human review, and a reported error rate of zero. For classifiers, Google states that because they only sort and prioritise content for human confirmation, 'there is no risk of false positives by reason of this technology alone'. Under that definition the medical-photo cases are by construction not counted as detection errors in any published figure. This is a measurement-scope fact and not an accusation that any published number is false. The scope of those filings is narrower still: they cover Google's messaging and mail services in the European Union and exclude Google Photos and YouTube, which are the services in which the documented account closures happened. Independent analysis converges on the shape rather than the rate. The peer-reviewed scanning literature states that 'false positives are inevitable: some innocuous content will be flagged as targeted', that error likelihood rises where training and deployment distributions differ, and gives the scale arithmetic explicitly. A Stanford trust-and-safety fellow told the New York Times in 2023 that adjudicating value judgements at this scale is 'just a very, very hard-to-solve problem' and that 'when you roll the dice that many times, you are going to roll snake eyes'. The Electronic Frontier Foundation recorded comparative evidence in 2022 that most flagged accounts are non-malicious: Facebook found 75 percent of accounts reported for alleged CSAM had sent 'non-malicious' images, and LinkedIn reported 75 accounts to EU authorities with manual review confirming CSAM in 31. A trade analysis at the time cautioned that most people affected by this shape will never speak publicly, which cuts both ways on any attempted rate.

Sources: googleirelandlimited2024, abelson2024, hill2023, mullin2022

Appears on: /domains/cases/google-csam-account-closure

EmpiricalThe unit of enforcement in this deployment is the account rather than the item. The San Francisco father lost more than …

The unit of enforcement in this deployment is the account rather than the item. The San Francisco father lost more than a decade of Gmail, Google Drive documents, Google Photos including the entire photographic record of his son's first years, contacts for friends and former colleagues, and his Google Fi phone service, which required obtaining service and a new number from another carrier. The Houston father, a paying customer, lost a decade-old account in the middle of buying a house, disrupting the transaction. The December 2023 case shows the downstream blast radius on a third account: work-schedule messages, bank statements, and third-party applications signed in with the Google Account; February 2026 trade reporting adds gig-work platforms, banking and home-security services. Google's own current help text sets the shape of the remedy: disablement reasons include child sexual abuse and exploitation; 'For some policy violations, Google will review up to 2 appeals'; data download is unavailable for certain violations 'including but not limited to: Valid legal requests, Account hijacking, Egregious content violations'; and if an appeal is not approved 'your entire Google Account will remain unavailable... your account will be permanently disabled and considered for deletion'. Google's own transparency-centre description of its appeals estate lists product-specific appeal forms for advertising, video, applications, maps and search and a general account-restoration path, and lists no child-safety-specific redress instrument; the child-safety entry point is a reporting form rather than a redress form. No partial remedy appears anywhere in the record.

Sources: hill2022, gizmodo2022, hill2023, googlea, piunikaweb2026

Appears on: /domains/cases/google-csam-account-closure

EmpiricalGoogle's own regulated filings state what an appeal overturns, in the same construction three years running: a reinstate…

Google's own regulated filings state what an appeal overturns, in the same construction three years running: a reinstatement 'was not due to an error in detection or a content-level false positive, but rather a reinstatement based on contextual information identified during the appeal process, which indicated that the content was correctly identified but did not appear to be possessed or shared with intent to harm, abuse, or exploit children'. The appeal re-reads intent; the record does not describe it re-reading the image. The measured yields, under Regulation (EU) 2021/1232 and, from 2025, Regulation (EU) 2024/2916: in 2023, 635 accounts identified by automated technologies, 734 CyberTipline reports, 1,558 content items, 297 accounts appealed and 10 reinstated; in 2024, 503 accounts, 508 reports, 1,824 content items, 216 appealed and 19 reinstated; in 2025, on a new Commission standard form, 380 known-material reports covering 1,419 images and 92 videos, 483 content items removed, 114 accounts suspended, 254 complaints lodged with the internal mechanism, 7 accounts restored, and 6 instances where an initial content verdict was overturned on review with the file made available to the user for download. Judicial complaints in all three years: zero. SCOPE, stated every time these figures are used: they are European Union only, cover Google's messaging and mail services and, in 2023, a historic chat product, and exclude Google Photos and YouTube, which are the services in which the documented closures happened. The 2025 denominators are complaints rather than appealing accounts and its suspension count (114) is smaller than its complaint count (254), so the three years are a trend rather than a series. Google also describes the quality regime behind the verdicts these appeals contest: weekly quality audits of reviewer verdicts, precision and recall monitored and reported monthly against an agreed target (its own example is 95 percent) with root-cause analysis and corrective action below it, and reporting automated only for content matching a hash previously confirmed as CSAM by a manual reviewer, with automated verdicts sampled rather than each manually reviewed before reporting.

Sources: googleirelandlimited2024

Appears on: /domains/cases/google-csam-account-closure

EmpiricalThe referral path carries no return leg, and Google says so in its own regulated filing: 'While Google's reports to the …

The referral path carries no return leg, and Google says so in its own regulated filing: 'While Google's reports to the NCMEC CyberTipline are one-way reporting, the information sharing and collaboration with NCMEC and NGOs provide the necessary feedback loop to continuously improve Google's detection technology.' Stanford's 2024 ecosystem study, based on dozens of interviews with platforms, NCMEC and law enforcement, found independently that platforms 'rarely get either' outcome information or report-quality feedback from law enforcement, and that NCMEC built a structured law-enforcement outcome field into the report flow which law enforcement rarely fills in. A platform respondent described the resulting posture: without that feedback 'you are stuck in a system where turning over anything is better than trying to think through how to do this well.' The study's proposed remedy names this deployment's exact mechanism: NCMEC should 'publish a negative hash set of images that have been reported as CSAM but have been verified to not be violative... This would allow platforms to stop reports (and automated processes such as account termination) on known legal content.' NCMEC's April 2024 response appreciated the analysis, disputed nothing specifically and said it would explore the recommendations; no such published negative hash set was located as of this run. The scale of the channel on the receiving side: NCMEC's CyberTipline received 21.3 million reports in 2025, of which 21,181,300 (99.2 percent) came from electronic service providers and 170,193 (0.8 percent) from the public, with five providers accounting for more than 75 percent of reports; NCMEC referred more than 18.8 million reports to law enforcement, designated more than 4.5 million as informational, issued over 172,000 removal notices with a 2.6-day average takedown, and distributes to task forces in all fifty US states and to law enforcement in 170 countries. In 2023, 245 platforms reported at all, 41 percent of them submitting 20 or fewer reports, and NCMEC escalated 63,892 tips as urgent or imminent-danger. So an investigator's conclusion that no crime occurred exists as paper in a cleared person's hand and has no documented input anywhere in the enforcement system.

Sources: googleirelandlimited2024, grossman2024, nationalcenterformissingexpl2024, nationalcenterformissingexpla

Appears on: /domains/cases/google-csam-account-closure

EmpiricalThe channel with a demonstrated success rate in this record has no formal standing in the process. On 28 October 2022 Go…

The channel with a demonstrated success rate in this record has no formal standing in the process. On 28 October 2022 Google's VP of Trust and Safety Operations published the company's account of the pipeline and said Google was 'actively working on ways to increase transparency' about suspension reasons and to improve the appeals experience. On 30 December 2022 the New York Times reported the resulting change: users flagged for child-safety violations now receive a more specific reason and a path to supply context, and a mother in Colorado recovered her account after four months, following a Times inquiry. The same piece summarised the two 2022 fathers: 'The police determined that the fathers had committed no crime, but the company still deleted their accounts.' In December 2023 the Times documented a third shape: a mother in Australia lost her whole Google Account after her seven-year-old uploaded a video to YouTube; the upload was flagged within minutes, her repeated appeals were denied even on a paid support channel, and the account was restored one day after a Times reporter asked about it. Google's statement was that 'we understand that the violative content was not uploaded maliciously', and the company 'had no response for how to escalate a denial of an appeal beyond emailing a Times reporter.' The policy change is real and partial: it altered what a person is told and what they may submit, and it did not alter the sanction, the two-appeal cap, the one-way referral, or the absence of any port for an exculpatory finding produced outside the platform. February 2026 trade reporting describes a fresh wave of Google Photos false-positive account bans, with one appeal rejected in about ten minutes against a stated review window of up to two days and one account restored after 24 hours on a second appeal explicitly requesting human review; that source aggregates user posts with no operator comment and no independent verification, and is carried here only for the claim that the pattern persists.

Sources: jasper2022, heer2025, hill2023, piunikaweb2026

Appears on: /domains/cases/google-csam-account-closure

EmpiricalThe legal and political forces on this pipeline run hardest in the direction of detecting and reporting more. 18 U.S.C. …

The legal and political forces on this pipeline run hardest in the direction of detecting and reporting more. 18 U.S.C. section 2258A requires a provider to report an apparent violation 'as soon as reasonably possible after obtaining actual knowledge', sets out what a report may include (subscriber identity, upload and transmission timestamps and time zones, geographic data, the visual depictions and the complete communication), and prices a failure to report from $600,000 to $1,000,000 depending on the offence and the provider's user base; the same section, at subsection (f)(3), disclaims any duty to 'monitor any user' or to 'affirmatively search, screen, or scan'. 47 U.S.C. section 230(c)(2)(A) immunises voluntary good-faith restriction of objectionable material. On 4 March 2026 a United States senator opened an investigation into Google for failing to remove child sexual abuse material and assist survivors, demanding by 18 March ten categories of documents including internal detection and removal policies, victim removal-request response times since January 2020, annual CyberTipline reports broken out by product, every case where content was not removed within 48 hours, Trust and Safety staffing levels and budgets, and any decision LIMITING deployment of CSAM detection technology. Those are a senator's allegations and demands and no adjudicated finding, and no Google response was located. They are recorded here for what they demonstrably are: a measure of the direction and intensity of pressure on the detection side in the same period in which the redress available to a wrongly closed account remained up to two appeals and a text box. Independent legal commentary makes the same point from the other end, observing that regulatory pressure to do 'more' against this material tightens filters and raises false positives.

Sources: ub, officeofu2026, googlea

Appears on: /domains/cases/google-csam-account-closure

EmpiricalNeither documented father sued, and that absence is recorded as an absence rather than as evidence that the conduct was …

Neither documented father sued, and that absence is recorded as an absence rather than as evidence that the conduct was lawful or that remedies exist. No class action or regulatory enforcement over these facts was located. Two United States decisions on this deployment shape do exist and both went Google's way on different questions, neither involving either man. In Baker v. Google LLC, No. 1:23-cv-02013, 2024 WL 3551878 (D.D.C. 26 July 2024), Judge Kollar-Kotelly dismissed a self-represented plaintiff's challenge to a CSAM-based account termination at the pleading stage: the contract claim because 'Plaintiff does not allege any facts indicating that Defendant was contractually prohibited from removing her Google account', the fraud claim, and the constitutional claim because 'Defendant Google is a private business, not a state actor'. That is a pleading-stage dismissal of one complaint, not a general holding that such terminations are lawful in every circumstance. In State v. Rauch Sharak, 2026 WI 4 (Wis. 24 February 2026), No. 2024AP469-CR, the Wisconsin Supreme Court held unanimously that Google 'acted as a private actor — not as an instrument or agent of the government — when it scanned Rauch Sharak's files and an employee opened and viewed files flagged as CSAM', reasoning from section 2258A(f)(3)'s disclaimer that searches are not required and from section 230(c) being 'entirely passive', and collecting the federal courts of appeals in agreement. THAT CASE INVOLVES A CONVICTED DEFENDANT, has no connection to the medical-photo cases, and is cited only for the legal architecture. The practical consequence for the correction channel is the point: the doctrine that keeps the scan outside the Fourth Amendment is the same doctrine that makes a platform wary of taking direction, or evidence, from law enforcement. A March 2026 legal round-up places both decisions in context and confirms the current posture, that Section 230 continues to immunise suspension decisions while granting no incentive to scan.

Sources: bakerv2024a, statev2026a, grossman2024

Appears on: /domains/cases/google-csam-account-closure

EmpiricalMeta's cross-check programme is an exemption tier bolted on top of at-scale content enforcement, and it inverts the usua…

Meta's cross-check programme is an exemption tier bolted on top of at-scale content enforcement, and it inverts the usual order of moderation: for entities on Meta's lists, content its own systems identify as violating is NOT removed as it would be for an ordinary user, but is left fully accessible pending additional human review. Meta disclosed the surrounding scale to the Oversight Board directly — about 100 million enforcement attempts on content every day at the time of the Board's briefings, from which the Board's own arithmetic is that 99 per cent accuracy would still leave a million mistakes a day — and disclosed that approximately 0.01 per cent of all content identified as needing enforcement was escalated through cross-check to reviewers empowered to apply context-specific policies and allowances. The programme has two pathways. Early Response Secondary Review, the entity-list pathway, renamed Secondary Sensitive Entity Review effective 25 April 2024, commits a listed entity's content to human review through as many as five successive layers: initial automated or at-scale human identification; the Regional Market Team, whose staff and contractors have language and market knowledge; the Early Response Team, the first layer that may authorise enforcement and the holder of the escalation-only policies and the newsworthiness and spirit-of-policy allowances, which is not required to have language or regional expertise and works from Market Team notes and translations; an enhanced Early Response Team review with subject-matter experts plus Public Policy, Communications and Legal; and a discretionary global senior-leadership review. The process stops and the content stays up if any reviewer finds it non-violating. General Secondary Review, built in 2021 and extended to content from all users in early 2022, is the content-ranked pathway: an automated ranker scores a flagged item on topic sensitivity, enforcement severity, false-positive probability, predicted reach and entity sensitivity — topic sensitivity and entity sensitivity being the most heavily weighted — and either suspends enforcement and queues the item or lets the original action stand. The decisive engineering fact, which Meta disclosed, is that both pathways draw on the SAME Regional Market Team capacity, with the entity-list pathway guaranteed and served first, and the same teams also carry non-cross-check work.

Sources: oversightboard2022a, oversightboard2022, metaplatforms2024, metaplatforms2021

Appears on: /domains/cases/meta-cross-check

EmpiricalThe harm this programme generates is a latency rather than a wrong verdict, and Meta disclosed the measurements. Its int…

The harm this programme generates is a latency rather than a wrong verdict, and Meta disclosed the measurements. Its internal goal for Market Team cross-check review was 12 to 120 hours by severity. The actual mean time to a final Early Response Secondary Review decision was more than five days; approximately 12 days on average for content posted from the United States; about 17 days on average for Afghanistan and Syria; and the longest single piece of content in the queue in the data Meta supplied had sat 222 days. Meta's own stated position is that most views happen when content is fresh, so violating content on this pathway stays up through its peak-virality window. The general pathway carries the mirror image of the same shortfall. Content ranked high by the cross-check ranker has its enforcement suspended and enters a queue that Market Teams serve with residual capacity after the guaranteed pathway is satisfied; if reviewers do not reach it, it times out after two to four days and the original enforcement action is applied unreviewed. That pathway's overturn rate was about 80 per cent at the February 2022 briefings and about 70 per cent for March to May 2022, from which the Board concluded that when General Secondary Review content times out 'it is highly likely that Meta is enforcing a significant number of false positives'. Meta told the Board it 'never intended to operate with a consistent backlog' but that 'operational capacity constraints and increasing volumes have led to a backlog in Early Response Secondary Review'. Meta's own research identified the drivers of users seeing violating content as incorrect overturns and the delay of enforcement for non-overturns slowed by the secondary review process. These overturn figures are the programme's internal efficacy metric on the content it routed — a different quantity from the Oversight Board's overturn rate on the emblematic cases it selects, and the two are never combined.

Sources: oversightboard2022a, oversightboard2022

Appears on: /domains/cases/meta-cross-check

EmpiricalEntity selection was tied to the cost of a mistake to Meta rather than to the value of the speech, and the Oversight Boa…

Entity selection was tied to the cost of a mistake to Meta rather than to the value of the speech, and the Oversight Board's opinion records the criteria. Entitlement 'is also determined by an assessment of the impact a potential enforcement mistake would have on the company in terms of the level of company leadership that would be involved in finding a solution', with 'escalation at the highest levels (CEO, COO)' corresponding to an 'extremely high severity' cross-check tag; inclusion criteria referenced advertising spend or revenue across the family of apps, enforcement history and follower reach; and at the time of the opinion a single Meta employee could add an entity to a cross-check list with no required review of that decision. Listed entities carry tags in six categories — civic and government; significant world events; media organisations, businesses, communities and creators, including advertisers; historically over-enforced; legal and regulatory; and entities whose content is under review — which generally expire after a year, with civic-and-government tags carrying no default expiry. The people on the list are not told they are enrolled. Meta disclosed the resulting allocation: 42 per cent of content reviewed through the entity-list pathway originated from the United States or Canada and 20 per cent of listed entities corresponded to those two countries, against 9 per cent of Facebook monthly active people. Alongside cross-check sits a harder exemption Meta calls 'technical corrections': automatic bars that override almost all enforcement attempts for a preselected entity-and-violation pair, checked before cross-check eligibility is assessed. Meta applied about a thousand a day and stated four remained active as of 21 September 2022, primarily on spam or inauthentic behaviour and impersonation, and acknowledged that a past 'lack of governance over practices [...] inadvertently resulted in some entities not receiving many enforcement actions'. Public reporting had described the practice as allowlisting or whitelisting, and the Board agreed the label is apt. Separately, the Board found that a user reporting a listed entity's content is not told that special procedures and longer timelines apply, and Meta declined the recommendation that they be told.

Sources: oversightboard2022a, metaplatforms2024, horwitz2021

Appears on: /domains/cases/meta-cross-check

EmpiricalThe Oversight Board is not a court and not a regulator. It is a quasi-judicial body Meta established and funds through a…

The Oversight Board is not a court and not a regulator. It is a quasi-judicial body Meta established and funds through an irrevocable trust: its decisions on the individual cases it takes bind, its policy recommendations do not, and it holds no audit power, no subpoena power and no way to compel anything. It learned of cross-check only because it asked a question during the Trump case, which Meta had not disclosed in its referral; after The Wall Street Journal's September 2021 reporting the Board found that 'the team within Facebook tasked to provide information has not been fully forthcoming in its responses on cross-check', and Meta requested a policy advisory opinion days later. During the opinion Meta refused the Board's repeated requests for the entity list itself, citing user-privacy obligations, and almost five months later supplied only aggregate fields — entity type, self-selected country and language, a civic flag and a partner flag — and for a quarter of the listed Instagram entities disclosed only that they existed. Meta answered 58 of the Board's 74 questions fully, 11 partially, and 5 not at all. The Board's opinion of 6 December 2022 found four shortcomings — unequal treatment of users, delayed removal of violating content, failure to track core metrics, and lack of transparency — and concluded that while Meta told the Board cross-check advances its human rights commitments, 'the program appears more directly structured to satisfy business concerns'. On unequal access to the rulebook it wrote that 'Meta has repeatedly told the Board and the public that the same set of policies apply to all users. Such statements and the public-facing content policies are misleading.' It issued 32 recommendations, the largest set it had issued at once. Two further findings define the governance shape. Meta 'did not provide the Board with information showing that it tracks data about the accuracy of decisions made through its cross-check system', and had no statistically significant data distinguishing account-level penalties applied to cross-checked versus non-cross-checked entities — so the programme's founding claim, that the exception path is more accurate than ordinary enforcement, was untested by its own operator. And the Board's own finding on the remedy it obtained is that internal auditing without external oversight falls short: there is no external audit of cross-check anywhere in the record. The Board's reach into the programme was itself tiered: for May and June 2022 an average 35 per cent of cross-check content could not be escalated to the Board at all, so the highest-reach accounts' content was systematically the least appealable to the external reviewer.

Sources: oversightboard2022a, oversightboard2022, horwitz2021

Appears on: /domains/cases/meta-cross-check

EmpiricalMeta responded publicly to the opinion on 6 March 2023, stating in its Q1 2023 quarterly update that it had 'responded p…

Meta responded publicly to the opinion on 6 March 2023, stating in its Q1 2023 quarterly update that it had 'responded publicly to all 33 of the board's cross-check recommendations, committing to implementing 82% either in part or in full' — a count of 33 against the 32 the Board's opinion and annex enumerate and the 32 Meta's own tracker page enumerates. Both counts are stated here and neither is silently reconciled, and no aggregate implementation tally is asserted, because the tallies on Meta's tracker are unstable across renderings; only per-recommendation statuses that reproduced consistently and are corroborated by the downloaded quarterly-update PDF are used. What Meta implemented is recorded independently by the Board's own Q2 2023 transparency report of 26 October 2023: Meta 'has cleared all outstanding backlogs in its cross-check review queues dedicated to potentially violating content from entities on its lists', 'producing a 96% decrease in resolution time (time taken for review and any subsequent enforcement) for 90% of the jobs created in the first half of 2023, compared with the second half of 2022'; and the new technical-corrections approach 'led to an immediate decrease in the overall size of the technical corrections list by more than half (55%)'. Meta also established add-and-remove criteria, time-bound cross-check tags, multi-person approval and internal audit over the lists, and committed to staffing cross-check decisions with reviewers who speak the language and have regional expertise. What Meta declined is equally specific: its own tracker records recommendations 5, 6, 12, 13 and 29 as 'No Further Action' — an open, criteria-based application route into the programme; an explicit rules re-commitment at enrolment; publicly marking the accounts of state actors, political candidates, business partners, media actors and commercially included public figures; telling a user who reports such an account's content that special procedures and longer timelines apply; and publishing metrics quantifying the adverse effects of delayed enforcement, such as views accrued on content left up during enhanced review and later found violating. Meta cited targeting and gamification risk for the two marking-and-notice recommendations, and pointed to a promised cross-check-specific report under recommendation 30 in place of the harm metric. The Board's five-year retrospective of 4 December 2025 cites this work as a flagship impact, in a document that is the body assessing its own effect.

Sources: oversightboard2023, metaplatforms2023a, metaplatforms2024, oversightboard2025a

Appears on: /domains/cases/meta-cross-check

EmpiricalThe transparency the Board asked for has not arrived in the form it asked for, and the honest statement is an absence fo…

The transparency the Board asked for has not arrived in the form it asked for, and the honest statement is an absence found by search rather than an abandonment. Under recommendation 30 Meta says it will produce 'an annual report containing metrics on the functionality and impact of cross-check' and describes this as a long-term effort; no such report was located as published as of 28 August 2026. Meta's cross-check recommendation tracker was last updated 3 October 2024. Its H2 2025 bi-annual report on the Oversight Board, published 19 March 2026 and covering 326 recommendations responded to as of 31 December 2025, contains no cross-check reporting. Meanwhile the programme continued and grew: the entity-list pathway was renamed Secondary Sensitive Entity Review effective 25 April 2024, and in March 2025 Meta's cross-check teams sought the Board's input on expanding coverage to more users, with the result including further investment in a Dynamic Multi-Review system intended to reduce over-enforcement at scale while keeping sensitive activism and journalism content with specialised reviewers. A population-level over-enforcement metric did arrive, but not the exemption-path one: Meta began publishing global enforcement precision in 2025, reporting around 91 per cent on Facebook and around 92 per cent on Instagram at the end of H1 2026, and reported roughly a 50 per cent reduction in United States enforcement mistakes between Q4 2024 and Q1 2025 following its 7 January 2025 policy overhaul, in which it said one to two of every ten December 2024 enforcement actions may have been mistakes, ended third-party fact-checking in the United States, and narrowed automated enforcement to illegal and high-severity violations while requiring user reports for less severe ones. None of those figures is disaggregated for the cross-check pathway the Board asked about.

Sources: metaplatforms2024, metaplatforms2026, metaplatforms2025a, metaplatforms, metaplatforms2025, kaplan2025

Appears on: /domains/cases/meta-cross-check

EmpiricalNo court and no regulator has adjudicated cross-check. The nearest regulatory pressure is Digital Services Act-shaped an…

No court and no regulator has adjudicated cross-check. The nearest regulatory pressure is Digital Services Act-shaped and adjacent rather than about the programme: on 24 October 2025 the European Commission issued PRELIMINARY findings that Facebook and Instagram appear not to provide a user-friendly, easily accessible notice-and-action mechanism for illegal content and appear to use dark patterns in it; that their appeal mechanisms appear not to allow users to provide explanations or supporting evidence; and that Meta and TikTok both breached researcher data-access obligations. The investigation was conducted with Coimisiun na Mean, the Irish Digital Services Coordinator. Preliminary findings expressly do not prejudge the outcome; if confirmed, exposure runs to fines of up to 6 per cent of total worldwide annual turnover. The reason this belongs beside cross-check is a structural adjacency rather than a legal one, and it is stated as such: the reporting channel and the appeal channel the Commission is examining are the same two channels the Oversight Board found cross-check quietly bypasses, since a user reporting a listed entity's content is not told that special procedures and longer timelines apply, and an average 35 per cent of cross-check content could not be escalated to the Board at all in May and June 2022. No DSA systemic-risk finding, proceeding or risk-assessment document naming cross-check was located.

Sources: europeancommission2025c, oversightboard2022a

Appears on: /domains/cases/meta-cross-check

EmpiricalThe CyberTipline is the single congressionally authorised reporting mechanism for online child sexual exploitation in th…

The CyberTipline is the single congressionally authorised reporting mechanism for online child sexual exploitation in the United States, built by the National Center for Missing & Exploited Children in March 1998, when it received 2,772 reports in its first calendar year. The congressionally mandated transparency report to the appropriations committees gives the recent series: 36,210,368 reports in calendar 2023, 20,512,803 in 2024 and 21,351,493 in 2025, the 2025 arrivals carrying 61,833,177 files. Of the 2025 total, 21,181,300 came from electronic service providers and 170,193 from members of the public, a 99.2 to 0.8 per cent split, and the public channel carried more than 5,700 reports directly from the person depicted. More than 2,000 providers are registered, just over 300 submitted any report in 2025, and five accounted for more than 75 per cent. The automated element at the centre is not a classifier: it resolves where a report belongs and matches it against entities already in the record across fields such as electronic mail addresses and network addresses, with analyst review on the matching, and no accuracy figure for it is published by anyone. The matching is exact-match; fuzzy matching that would catch a suspended account's near-identical new handle is not implemented, and material attached to a report is not automatically scanned for matches. The clearinghouse's own resolution failures are published as a series: reports whose location could not be determined ran 1,368,404 (3.8 per cent) in 2023, 1,957,640 (9.5 per cent) in 2024 and 3,044,434 (14.3 per cent) in 2025, and reports whose location cannot be resolved are made available to United States federal law enforcement by default, so a failure of resolution is itself a routing rule.

Sources: officeofjuvenilejusticeandde2026, grossman2024, nationalcenterformissingexpla, nationalcenterformissingexpl2025a

Appears on: /domains/cases/ncmec-cybertipline-triage

EmpiricalFederal law makes reporting mandatory, detection voluntary and report content discretionary, and no instrument anywhere …

Federal law makes reporting mandatory, detection voluntary and report content discretionary, and no instrument anywhere sets a standard for what a report must contain. 18 U.S.C. section 2258A(a) requires a provider to report an apparent violation 'as soon as reasonably possible after obtaining actual knowledge', a duty extended in May 2024 to minor sex trafficking under section 1591 and enticement under section 2422(b). Section 2258A(b) says the report 'may, at the sole discretion of the provider, include' the identity of the person involved, the historical reference showing when the material was uploaded, the geographic location including the network address, the visual depictions themselves and the complete communication. Section 2258A(f) says nothing in the section requires a provider to monitor any user or to 'affirmatively search, screen, or scan' for violations. Section 2258A(e) sets failure-to-report penalties of $600,000 to $850,000 for a first violation and $850,000 to $1,000,000 for subsequent ones, scaled by whether the provider has at least 100 million monthly active users. The clearinghouse concedes it lacks authority to make platforms change their reporting, and it declines to publish written guidance to senders: its staff told researchers that if there were a written document, 'defense attorneys would characterize this in criminal cases as NCMEC is advising companies what to report', and that it preferred best practices to come from an industry body composed solely of private companies. It also has no authority to require feedback from any receiving agency. So the only price signal any sender faces is a penalty for failing to report, which is one-sided.

Sources: ub, publiclaw2024, grossman2024, officeofjuvenilejusticeandde2026

Appears on: /domains/cases/ncmec-cybertipline-triage

EmpiricalThe clearinghouse's legal status is contested, and the contest decides what its own operators may examine. In United Sta…

The clearinghouse's legal status is contested, and the contest decides what its own operators may examine. In United States v. Ackerman, 831 F.3d 1292 (10th Cir. 2016), the court held that NCMEC qualifies as a governmental entity in light of its authorising statutes and the functions Congress gave it, and in the alternative acted as a government agent, so its opening and viewing of the reported files was a warrantless search. In United States v. Wilson, 13 F.4th 961 (9th Cir. 2021), the court held that the private-search exception does not cover files the platform never opened. NCMEC disagrees with the first holding, describes itself as a private non-profit and, in its staff's words to researchers, 'merely a middleman', prints a disclaimer that it is not an agent or instrumentality of the government at the foot of every report it sends and in its law-enforcement tooling, and complies anyway: it opens only the files a platform employee recorded as viewed, for reports bound for United States law enforcement. The mechanism is one indication added to the reporting form at the start of 2014. United States v. Lowers, 170 F.4th 134 (4th Cir. 2026), records it operating: a platform's hashing flagged 156 files uploaded to one account, a reviewer at the platform opened 31 of them, the report identified which had been viewed and which had not, and 'An employee at NCMEC received that CyberTip and opened and viewed the same 31 images as the Google Reviewer. The NCMEC employee did not open any of the remaining 125 unreviewed files.' The same rule cuts both ways: the no-opening practice applies to reports bound for United States law enforcement, and the clearinghouse is able to open a file with the indication absent where the report will go to law enforcement outside the United States, which in 2025 was 77.1 per cent of reports. This is voluntary compliance with a contested holding, and this atlas does not describe NCMEC as a government agency.

Sources: unitedstatesv2016, unitedstatesv2021, unitedstatesv2026, grossman2024

Appears on: /domains/cases/ncmec-cybertipline-triage

EmpiricalThe clearinghouse is required to forward everything, and its only lever over the downstream load is a label defined by t…

The clearinghouse is required to forward everything, and its only lever over the downstream load is a label defined by the absence of information rather than by the seriousness of the conduct. 18 U.S.C. section 2258A(c) provides that NCMEC 'shall make available each report' to the relevant agencies, and NCMEC states in its own report to the appropriations committees that it is required by law to make available every CyberTipline report it receives, and that the label is applied based on the information a reporting party voluntarily chose to include. Its definitions, verbatim: 'An actionable report contains information indicative of a suspected prior, ongoing, or planned child sexual exploitation incident. An informational report contains severely limited information in which there is no apparent child sexual exploitation nexus; or so little information was provided by the reporting party that it is impossible to identify a location to refer the report to; or contains frequently seen child sexual exploitation or abuse material that has been shared in a non-malicious context, such as for inappropriate comedic effect or moral outrage or concern for the child depicted.' Field-study respondents report that United States law enforcement typically read the informational label as meaning a report can be set aside, and that not every report which could be set aside carries it. The second, smaller channel is urgency: more than 53,000 reports in 2025 escalated as urgent or involving imminent danger, identified through sender notifications, internal alerts and manual review, with the clearinghouse able to decline to escalate a report a platform escalated. THREE ACTIONABILITY FIGURES IN THIS RECORD SIT ON THREE DIFFERENT BASES AND ARE NEVER BLENDED: the operator's 2022 report figure of 49 per cent of all reports; the mandated per-recipient-agency tables, where a report may be counted against more than one agency and the 2025 rows sum to 19,286,122 actionable and 4,719,166 informational against a 21,351,493 report total; and the operator's public page figures of more than 18.8 million referred and more than 4.5 million informational. De-duplication of multi-agency reports is the likeliest reconciliation of the last two, the operator does not say so, and this atlas does not assert it.

Sources: officeofjuvenilejusticeandde2026, ub, grossman2024, nationalcenterformissingexpl2025

Appears on: /domains/cases/ncmec-cybertipline-triage

EmpiricalA primary appellate record measures this pipeline end to end and shows what a routing failure costs. In United States v.…

A primary appellate record measures this pipeline end to end and shows what a routing failure costs. In United States v. Lowers (4th Cir. 2026): files uploaded to one account on 20 September 2019 and hash-flagged by the platform within days; a report to NCMEC on 23 September; an NCMEC employee reading the network address as one Virginia county and forwarding the report there on 29 October 2019; that office, in the court's account, letting it sit for half a year; an investigator subpoenaing the internet service provider on 16 April 2020, learning the address was in a different city, and closing the file on 13 May 2020; and the receiving city's detective viewing three previously unopened files without a warrant and applying for a warrant on 27 May 2020. Eight months from upload to warrant, with one wrong jurisdiction and one half-year queue wait, and under the 90-day preservation rule then in force a preservation window that had lapsed twice. The Fourth Circuit held that 'a hashing algorithm, which reveals nothing about a given file but a non-descriptive serial number, does not frustrate a defendant's expectation of privacy in his unopened files', and that unless someone visually inspects the contents of a file containing apparent material before law enforcement does, the private-search doctrine is inapplicable; it aligned itself with the Second and Ninth Circuits and expressly recognised that this 'puts us at odds with the Fifth and Sixth Circuits'. It AFFIRMED the conviction on attenuation, so it is not a suppression win. It also recorded what the record did not contain: 'The record does not reveal how Google trains Google Reviewers on interpreting and applying the federal CSAM definition. Nor is there any record evidence indicating how accurate or reliable Google Reviewers are at actually identifying apparent CSAM. Similarly, there is no record evidence demonstrating how accurate Google's proprietary hashing algorithm is in practice.' The split is live: the Supreme Court of Wisconsin decided State v. Gasper 5-2 on 14 January 2026 on the other side, and a certiorari petition, No. 25-1191, was filed on 14 April 2026 and distributed on 17 June 2026 for the conference of 28 September 2026, neither granted nor denied as of 28 August 2026.

Sources: unitedstatesv2026, statev2026

Appears on: /domains/cases/ncmec-cybertipline-triage

EmpiricalThe return channel that would let anyone learn which reports were worth investigating is built, well designed and largel…

The return channel that would let anyone learn which reports were worth investigating is built, well designed and largely unused, and the operator publishes the counts. The structured schema records case status (conviction, arrest, ongoing investigation, referred, closed), whether a child victim was identified on arrest, ten named closure reasons (unable to locate subject, provider legal response does not contain information, no crime committed, no prosecutorial merit, alleged child is an adult, age of child victim unable to be determined, false report, unfounded, person or user reported is deceased, other), and a direct question on whether the information NCMEC provided was useful, with a stale-information option. NCMEC states that agencies 'are not generally required by law to provide feedback on CyberTipline reports, and NCMEC has no authority to require such feedback be submitted' and that 'most agencies provide little or no feedback.' The measured uptake in calendar 2025: task force units returned 549,584 feedback instances against 1,932,435 reports received; federal law enforcement returned 7,085 against 3,435,257; local agencies returned 156; international recipients returned 265,079. The consequence is stated by the only field study of the system: it is unknown what share of reports, if fully investigated, would reveal hands-on abuse, and no empirical prioritisation rule exists anywhere in the pipeline. Respondent estimates of the share of reports leading to a United States arrest range from 5 per cent, given in congressional testimony in September 2023, to 7.6 per cent from one officer's 2023 figure, both for reports sent to task forces and neither covering the federal stream; one officer estimated that in 2022, 3.8 per cent of reports in his state led to a child being reached. Those are respondent figures on one stream, not system measurements.

Sources: officeofjuvenilejusticeandde2026, grossman2024

Appears on: /domains/cases/ncmec-cybertipline-triage

EmpiricalCapacity on both sides of this pipeline is set by appropriation and by salary rather than by arrivals. NCMEC has more th…

Capacity on both sides of this pipeline is set by appropriation and by salary rather than by arrivals. NCMEC has more than 400 employees across five programme areas, of which the CyberTipline is one of two core exploitation programmes; no CyberTipline analyst headcount is published, and no queue depth, review time, backlog or per-analyst caseload figure has ever been published for this deployment by anyone. The field study records the staffing constraint in the operator's own terms: analysts are constantly recruited away by industry trust and safety teams, and the organisation described a never ending cycle of trying to replace the workforce; asked what it would do with more resources, it said it would build out its technology team; and a federal department employee described the position as 'The house is flooding, they're bailing water, and we're asking them to build a drainage system at the same time. You can't stop bailing, otherwise you'll drown.' Downstream, the Office of Juvenile Justice and Delinquency Prevention funded the 61 Internet Crimes Against Children task forces, a network of more than 6,200 federal, state, local and Tribal agencies, at $33,976,146 in fiscal 2025 across a competition with 61 expected awards and a published award maximum of $1,042,765 with no published minimum; in 2025 the network conducted nearly 347,000 investigations leading to more than 17,000 arrests and trained approximately 73,000 professionals, while NCMEC made 1,932,435 reports available to those units. In fiscal 2023 NCMEC received $41.4 million for its fifteen programmes and the 61 task forces received $40.8 million between them, and the field study reports a perception among participants that a larger share for one means less for the other. Two of the study's technical recommendations were unshipped capacity rather than new ideas: a commissioned interface matching report network addresses against peer-to-peer file-sharing data was completed in autumn 2020 and, as of 2024, had not been integrated, and an offer of cloud translation capacity for recipients abroad had not been taken up. NCMEC's declared answer is a $10 million, three-year CyberTipline Modernization Initiative supported by three cloud and analytics companies with further corporate investors; it is an announced programme with no published outcome measurement.

Sources: grossman2024, officeofjuvenilejusticeandde2026, officeofjuvenilejusticeandde2025, nationalcenterformissingexpl2026a

Appears on: /domains/cases/ncmec-cybertipline-triage

EmpiricalReport counts are not victim counts, and reporting volume is not a safety metric in either direction. Roughly 35 per cen…

Report counts are not victim counts, and reporting volume is not a safety metric in either direction. Roughly 35 per cent of the 2025 file volume is exact or near duplicate of something already held: of 29,408,181 images submitted, 19,091,252 were unique by exact hash and 13,994,568 distinct under visual-similarity matching; of 26,324,863 videos, 15,144,788 and 7,304,334. Hundreds of reports may concern one person, and the Belgian Federal Police reported receiving over 500 distinct reports about a single offender in five months. A large share of the material is older material recirculating, NCMEC does not break out reports where the child is already known and safe, and its own 2022 figure was that 49 per cent of reports were actionable. The inference does not run backwards either: every interviewee in the field study with a view believed the underlying threat is understated rather than overstated, and one respondent said 'We aren't doing a good enough job of selling the threat... The number gets trotted out to justify everything, and then people wonder why they don't get resources.' Sender volume is equally unsafe to read as a signal: NCMEC's own tables show one company's volume rising more than thirty-six-fold between 2024 and 2025 and the largest sender's roughly halving, and the four accounts of the 2024 decline disagree. NCMEC's chief legal officer attributed the fall of 15.7 million reports almost entirely to default end-to-end encryption on one platform's messaging surfaces, said her first question was whether a company had stopped reporting or gone out of business and that there was nothing like that, and said unbundling to count every incident still leaves a 7 million-report gap; that platform attributes part of the fall to a report-bundling feature it partnered on and says it maintains safety measures inside encryption; two other companies claimed consolidation and an NCMEC spokesperson said any such changes were 'not via the official feature in the CyberTipline reporting pipeline'. The generative-artificial-intelligence load arrived in the same period on two bases: NCMEC counted more than 400,000 2025 reports with such a nexus, more than 182,000 involving offenders possessing, generating or attempting to generate the material, and more than 158,000 files so categorised, while the figure released through Senate Judiciary oversight for the same year was 1.5 million reports with such a connection, including over 12,000 reports of the material found in training data.

Sources: officeofjuvenilejusticeandde2026, nationalcenterformissingexpla, goggin2025, grossman2024, u2026a

Appears on: /domains/cases/ncmec-cybertipline-triage

EmpiricalOversight of this deployment is layered and is weak in the one direction that would change the input. Congress authorise…

Oversight of this deployment is layered and is weak in the one direction that would change the input. Congress authorises and funds the programmes and, since the fiscal-2022 appropriations act's joint explanatory statement, requires an annual transparency report to the appropriations committees specifying de-duplication, victim-identification and series counts; that document is the count of record for almost every quantity in this case file. Two Senate offices have run direct oversight of the SENDERS through the clearinghouse's own data: on 30 April 2025 the author of the REPORT Act opened an inquiry with four companies over what her office called a sharp decline in reports since the Act's passage, citing testimony by NCMEC's president and chief executive; and on 16 March 2026 the Senate Judiciary Chairman put an oversight letter to NCMEC whose answers the committee released on 9 April 2026, covering eight companies that submitted over 17 million 2025 reports, 81 per cent of the total. As characterised in that committee release, NCMEC told the committee that one artificial intelligence service's more than 1.1 million reports contained zero per cent actionable information because the service was designed not to collect user or content data; that over 80 per cent of one messaging platform's more than 752,000 reports were deemed inactionable by law enforcement for insufficient information; that over 90 per cent of another sender's more than 135,000 reports were originally inactionable, improving after intervention; that one platform supplied location information in 4 per cent of its 2025 reports against 35 per cent in 2024; that another routinely submitted unrelated content; and that a fifth's omissions of location or account information rendered reports inactionable. THESE ARE OVERSIGHT FINDINGS ABOUT DATA COMPLETENESS AND NOT ENFORCEMENT FINDINGS: they are second-hand from NCMEC through a committee majority release, the companies were pressed for responses, and none has been found to have violated the statute on this record. Academic oversight is a single field study, published 22 April 2024 on interviews with 66 individuals plus three days of on-site observation with the operator's cooperation; NCMEC published a roughly 300-word response the same day that appreciated the study's 'thorough consideration of the inherent challenges', called the recommendations 'creative', disputed no finding and gave no number. Three weeks later contemporaneous reporting described the Stanford Internet Observatory as being dismantled, with child-safety work continuing under another Stanford laboratory; Stanford disputed the characterisation, saying 'The important work of SIO continues under new leadership'. Both are carried. The only other known study of this system, commissioned in 2021 by a federal science directorate, was never made public. No regulator supervises this clearinghouse, and no court or regulator has ever ordered it to do anything.

Sources: officeofjuvenilejusticeandde2026, u2026a, officeofu2025, grossman2024, nationalcenterformissingexpl2024, newton2024

Appears on: /domains/cases/ncmec-cybertipline-triage

EmpiricalThe governed subsystem at the Samasource Kenya EPZ Limited (Sama) Nairobi delivery centre was the outsourced human revie…

The governed subsystem at the Samasource Kenya EPZ Limited (Sama) Nairobi delivery centre was the outsourced human review layer itself rather than any scoring model, and the arrangement's defining feature is a split principal. Automated detection and user reports upstream of the vendor fed a queue; items that stage could not dispose of were routed to reviewers employed by Sama, who applied the client's written policy through the client's review tool and returned an action. TIME reported approximately 200 reviewers covering roughly eleven African languages for a sub-Saharan Africa queue in February 2022; the Employment and Labour Relations Court's ruling of 2 June 2023 records approximately 260 moderators affected by the January 2023 redundancy. The party that set the policy, supplied the tool, composed the queue and set both review targets — Meta Platforms, Inc. and Meta Platforms Ireland Limited — is not the party that employed, insured or medically supported the reviewers, and its controlling legal position throughout the litigation is that it is not an employer at all. Everything the pipeline exists to filter out passed through the reviewers' eyes by construction; the exposure classes named in the court record and the filed medical assessments include gruesome killings, self-harm and suicide, sexual violence, explicit sexual content, child physical and sexual abuse, mutilated bodies, and conflict footage from the Ethiopia-Tigray war. No enforcement error rate for this deployment is published in any source read for this record.

Sources: perrigo2022, arendseothersvmetaplatforms2023, stockwell2024, metaplatforms2023b

Appears on: /domains/cases/sama-nairobi-moderation-workforce

EmpiricalReviewers were metered on two axes at once, and every figure in this claim is attributed to a single investigative sourc…

Reviewers were metered on two axes at once, and every figure in this claim is attributed to a single investigative source reporting worker accounts and documents rather than to any court finding. TIME's investigation of 14 February 2022, working from payslips and worker statements, reports an average handling time target of fifty seconds per ticket, a quality or accuracy score requirement of at least eighty-four per cent audited against the client's policy, and shifts of up to nine hours with monitored screen time, from which it computes an implied quota of roughly 580 items per reviewer per day. The same investigation reports guidance to watch only the first fifteen seconds of a video before actioning it where the title and the surrounding comments appeared innocuous — a sampling rule inside the review step rather than an incidental practice. It reports take-home pay of about $1.46 an hour for Kenyan staff and about $2.20 an hour, roughly $440 a month, for non-Kenyan staff; a moderator interviewed in December 2023 gave a monthly salary of about $429 with non-Kenyan staff receiving an additional $200 three times a year, and a four-year moderator reported about $600 a month in May 2023. Sama's position is that moderators earned about triple the Kenyan minimum wage, which is the vendor's own account. None of these operational figures is confirmed in any of the six Kenyan judgments read for this record. The structural consequence is that the two targets are enforced against the same person from opposite directions, so an ambiguous or distressing item is costly to dwell on, and the reviewer holds action discretion within a ticket and no process discretion over queue composition, either target, the tool's defaults or the policy the audit scores against.

Sources: perrigo2022, siele2023

Appears on: /domains/cases/sama-nairobi-moderation-workforce

EmpiricalThis deployment instrumented throughput continuously and psychological load not at all, and the one measurement of the s…

This deployment instrumented throughput continuously and psychological load not at all, and the one measurement of the second quantity was made from outside, years later, for litigation. Throughput and accuracy were recorded per item with monitored screen time and fed performance management of the individual reviewer. On 4 December 2024 medical reports for 144 of the 185 claimants who volunteered for assessment were filed at the Nairobi Employment and Labour Relations Court by the claimants' advocates; the head of mental health services at Kenyatta National Hospital classed 81 per cent of those assessed as suffering severe post-traumatic stress disorder, with generalised anxiety disorder and major depressive disorder also diagnosed, and at least 40 were reported as misusing alcohol or drugs. THESE ARE FILED MEDICAL OPINIONS IN SUPPORT OF A PENDING CLAIM AND NOT ADJUDICATED FINDINGS, and the causal attribution of the diagnoses to the work is a pleaded allegation. Meta declined to comment on the reports because of the ongoing litigation, stating that it takes moderator support seriously, that its contracts with third-party firms set expectations on counselling, training and fair pay, and that moderators can customise the content-review tool so that graphic content appears blurred or in black and white; Samasource did not respond to the same request. Both are the parties' own positions and neither is tested by any independent source read for this record. No published mechanism carries moderator-health data back into the queue-routing, staffing or target-setting decisions that generated the exposure. The comparative baseline predates this case: the same outsourced structure was documented at the United States sites in 2019 and studied in the peer-reviewed literature on moderator psychological well-being in 2021.

Sources: stockwell2024, newton2019, steiger2021

Appears on: /domains/cases/sama-nairobi-moderation-workforce

EmpiricalThe one relief channel in this deployment existed and its access was held by the party whose objective was throughput. S…

The one relief channel in this deployment existed and its access was held by the party whose objective was throughput. Sama provided wellness counsellors on site and one hour of wellness break a week. A former counsellor told TIME that managers, rather than counsellors, held the final say over whether a break was granted, and frequently refused on productivity grounds. Sama's vice-president has stated separately that the company provides on-site licensed mental health professionals that employees can access at any time, which is the vendor's own account. In Constitutional Petition E052 of 2023 the petitioners pleaded that the wellness counsellors were not qualified psychiatrists or psychologists and that the insurance provided was inadequate; those are pleadings and not findings. On 2 June 2023 Justice B Ongaya granted thirteen interim orders including a requirement that the respondents provide proper medical, psychiatric and psychological care for the petitioners and other Facebook content moderators in place of wellness counselling, together with regularisation of the immigration status of non-Kenyan moderators and directions to named state bodies to review occupational-safety and employment law for virtual and digital work. ON 20 SEPTEMBER 2024 THE COURT OF APPEAL SET THAT RULING ASIDE IN ITS ENTIRETY TOGETHER WITH ALL CONSEQUENTIAL ORDERS, so the care requirement no longer stands and nothing in the current record obliges anyone to provide it.

Sources: perrigo2022, siele2023, arendseothersvmetaplatforms2023, samasourceepzlimitedtasamavm2023

Appears on: /domains/cases/sama-nairobi-moderation-workforce

EmpiricalCollective voice in this deployment arrived twice and both times after it could govern the capacity it was formed over. …

Collective voice in this deployment arrived twice and both times after it could govern the capacity it was formed over. In July 2019 a group of more than a hundred Sama moderators organised as the Alliance and petitioned for a doubling of wages; the drive's organiser was suspended and dismissed on 20 August 2019 on grounds recorded as bullying, harassment and coercion said to have placed the relationship with the client at risk. He alleges the dismissal was union-busting, and that allegation has not been adjudicated. On 1 May 2023, five weeks after the terminations took effect, more than 150 current and outsourced workers moderating for three different platform clients met in Nairobi and voted to register a content moderators union, the first such body on the continent; it is reported under two names across the coverage and no source read for this record confirms that registration with the Kenyan labour office was completed, so this record says voted to register rather than formed or registered. The organiser of the 2019 drive addressed that meeting and is the lead petitioner in Petition E071 of 2022, filed on 10 May 2022 on his own behalf and on behalf of current and former Facebook content moderators, alleging poor working conditions, unfair labour practices and violation of fundamental rights, with the pleaded case also including forced-labour and human-trafficking allegations. Every one of those allegations remains an allegation.

Sources: perrigo2022, perrigo2023, motaungvsamasourcekenyaepzlt2022, siele2023

Appears on: /domains/cases/sama-nairobi-moderation-workforce

EmpiricalThe capacity in this arrangement proved portable and the chronology of its removal is on the record to the day. On 10 Ja…

The capacity in this arrangement proved portable and the chronology of its removal is on the record to the day. On 10 January 2023 Sama announced it was leaving content moderation to concentrate on computer-vision data annotation and would not renew the Meta contract, which ran to the end of March 2023; redundancy notices issued on 10 January with a last working day of 28 February, a revised notice of 18 January moved that to 31 March, and termination letters issued on 8 February 2023, with approximately 260 moderators affected. The work moved to Majorel. The dismissed reviewers brought Constitutional Petition E052 of 2023 alleging that the redundancy was retaliation for the earlier petition and for complaints about pay and conditions, and alleging that the successor vendor had been instructed not to hire former Sama moderators; the 2 June 2023 interim orders included a prohibitory order restraining refusal to recruit qualified moderators on the ground of prior engagement through Sama, and that order was set aside on appeal on 20 September 2024 along with the rest of that ruling. Both the retaliation and the blacklisting claims remain allegations. On 11 May 2023 the court directed Sama to pay April salaries to the 184 former moderators, still outstanding at the time of reporting. On 16 April 2026 Sama issued 1,108 redundancy notices at the Nairobi delivery centre after Meta terminated its remaining contract; that work was data annotation rather than content moderation and is three years and one line of business away from the 2023 redundancy. The claimant count differs across the record and no bare number is asserted: 43 in the June 2023 caption, 183 in the December 2023 caption, 184 in May 2023 reporting and the salary order, 185 in the medical-evidence reporting, and 186 or 187 in the Court of Appeal captions.

Sources: njanja2023, arendseothersvmetaplatforms2023, siele2023, ndege2026, arendseothersvmetaplatforms2023a

Appears on: /domains/cases/sama-nairobi-moderation-workforce

EmpiricalThe litigation is active with no merits determination after four years, and the single most misreported fact about it is…

The litigation is active with no merits determination after four years, and the single most misreported fact about it is that two Court of Appeal judgments issued on the same day went in opposite directions. On 6 February 2023 Justice JK Gakeri disallowed the Meta entities' strike-out application as inopportune at that stage and directed compliance with the rule on service outside the jurisdiction, expressly leaving weighty outstanding issues to be determined — a refusal to strike out at an interlocutory stage, not a holding that Kenyan law governs the client's conduct. On 7 December 2023 Justice MN Nduma dismissed both contempt applications, holding that placing employees on paid leave was not an action that constituted willful or deliberate disobedience and that the electronic evidence was not sufficient to prove that the respondents had replaced the petitioners, contempt requiring a near-criminal standard of proof. On 20 September 2024, in Civil Appeal E595 of 2023 consolidated with E602 and E615, the Court of Appeal held that the trial judge had impermissibly and dangerously delved into contested issues of fact and law and that orders extending expired contracts and compelling medical and psychological care have the effect of final orders, ordered that the 2 June 2023 ruling is set aside in its entirety together with all consequential orders arising therefrom, and substituted an order dismissing the moderators' application. ON THE SAME DAY, in Civil Appeal E232 and E445 of 2023 consolidated, it dismissed the Meta entities' jurisdiction appeals with costs, holding that whether the appellants are engaged in virtual business in Kenya and whether the pleaded violations occurred in Kenya are contested questions of fact best resolved in a full hearing as opposed to an interlocutory application. Court-encouraged mediation before a former Chief Justice began in August 2023 and collapsed in October 2023. Leave to serve outside the jurisdiction was granted on 23 January 2024. On 26 May 2025 Justice MN Nduma dismissed the application to stay the consolidated trial pending Supreme Court certification. Rulings expected on 12 February 2026 were not delivered and the court adjourned on notice without fixing a date. No Kenyan court has made any merits finding against any respondent, and no damages figure attaches to this litigation — the $1.6 billion figure circulating in coverage belongs to a separate Kenyan High Court petition about the amplification of hateful content during the Ethiopia conflict. The United States comparator, a $52 million class settlement for moderators with post-traumatic stress disorder in 2020, is a different case in a different legal system settled without any admission or finding of liability.

Sources: motaungvsamasourcekenyaepzlt2022, samasourceepzlimitedtasamavm2023, metaplatforms2023b, arendseothersvmetaplatforms2023a, motaungvsamasourcekenyaepzli2022, capitalfm2026, nothias2026, allyn2020

Appears on: /domains/cases/sama-nairobi-moderation-workforce

EmpiricalThe European Commission has two formal Digital Services Act proceedings open against TikTok and neither names content-mo…

The European Commission has two formal Digital Services Act proceedings open against TikTok and neither names content-moderation staffing as a ground. The first, opened 19 February 2024, covers protection of minors, advertising transparency, data access for researchers, and the risk management of addictive design and harmful content, naming suspected infringements of Articles 34(1), 34(2), 35(1), 28(1), 39(1) and 40(12), and records TikTok's designation as a very large online platform on 25 April 2023 at 135.9 million EU monthly active recipients; the release states that the DSA sets no legal deadline for bringing formal proceedings to an end. The second, opened 17 December 2024, concerns election-integrity systemic risk following the annulled Romanian presidential first round of 24 November 2024 and is limited to two grounds — recommender systems including coordinated inauthentic manipulation, and policies on political advertisements and paid-for political content — under Articles 34(1), 34(2) and 35(1), with Coimisiún na Meán, the Irish Digital Services Coordinator, associated to the case and a retention order of 5 December 2024 preceding it. A third and earlier proceeding, on the TikTok Lite Rewards programme, was opened on 22 April 2024 and closed on 5 August 2024 when the Commission made binding TikTok's commitment to withdraw the programme from the EU permanently and not to launch a circumventing programme: the first DSA case closed and the first commitments accepted. As of 28 August 2026 the February 2024 proceeding has produced four sets of preliminary findings — the advertisement repository on 15 May 2025, researcher data access on 24 October 2025, addictive design on 6 February 2026 and minors' account settings on 24 July 2026 — and one closure by binding commitments, on advertising transparency, on 5 December 2025, with the rabbit-hole effect of the recommender systems and the risk of age misrepresentation still under investigation; the December 2024 proceeding has produced no preliminary findings. Preliminary findings are not findings of breach and do not prejudge the outcome, and TikTok said of the addictive-design set that 'The Commission's preliminary findings present a categorically false and entirely meritless depiction of our platform, and we will take whatever steps are necessary to challenge these findings.' There is no non-compliance decision and no fine against TikTok under the Digital Services Act.

Sources: europeancommission2024a, europeancommission2024, europeancommission2025c

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalThe European Commission has stated in writing that the Digital Services Act gives it no rule about content-moderation re…

The European Commission has stated in writing that the Digital Services Act gives it no rule about content-moderation resourcing. Asked in European Parliament written question E-002454/2024, submitted 6 November 2024 by Kim Van Sparrentak, whether a platform can comply with Articles 16, 20 and 34 to 35 after firing an entire national moderation team of 300 people in the Netherlands, the Commission answered, in a reply last updated 15 January 2025: 'The DSA does not prescribe any specific rules about the resources to be dedicated to content moderation.' It added that platforms must enforce their moderation rules 'in a diligent, objective and proportionate manner', that 'qualified staff' must ensure fair and unbiased decision-making in internal complaint handling, that 'it is important that designated companies put in place adequate content moderation processes and dedicate enough resources for diligent content moderation', and that it 'is closely monitoring TikTok's compliance with the DSA and will follow up with formal enforcement steps if appropriate'. No formal enforcement step naming moderation staffing has followed as of 28 August 2026. The question itself put the staffing change in Digital Services Act terms — Article 16 notice and action, Article 20 internal complaint handling, which requires reasoned decisions taken under the supervision of appropriately qualified staff and not solely on the basis of automated means, and Articles 34 and 35 on systemic risk taking into account specific regional or linguistic aspects — and cited 5.7 million Dutch monthly users. The boundary this establishes is the structure of the case: the regulator holds a diligence lever and not a headcount lever.

Sources: europeanparliamentwrittenque2024

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalTikTok's EU content-moderation workforce fell across 2024 to 2026 while its EU audience grew, on its own legally mandate…

TikTok's EU content-moderation workforce fell across 2024 to 2026 while its EU audience grew, on its own legally mandated disclosures. Its second DSA transparency report recorded 'More than 6k moderators are dedicated to the moderation of content in the European Union as of the end of December 2023'. Its fifth report, for January to June 2025, gives 4,596 people dedicated to EU content moderation at the end of June 2025, of whom 247 are not language-specific, with an Annex D per-language table — as corrected on 15 April 2026 — of English 1,552, German 567, French 525, Spanish 437, Italian 331, Polish 144, Portuguese 143, Romanian 103, Dutch 100, Swedish 63, Hungarian 37, Bulgarian 34, Czech 31, Greek 29, Finnish 28, Slovenian 26, Slovak 25, Danish 19, Latvian 11, Croatian 10, Estonian 10, Lithuanian 5, Irish 0 and Maltese 0; TikTok notes that Czech, Slovak and Slovenian are one team, that Croatian moderators also cover Serbian, and that the totals also include Arabic, Catalan, Hindi, Pashto, Persian, Turkish, Ukrainian, Norwegian, Russian and Icelandic capacity. EUobserver, reading the sixth report for July to December 2025, records 91 employed staff against 3,583 contracted human moderators. Over the same span TikTok reported 169 million EU monthly active recipients in the first half of 2025 against the 135.9 million declared at designation in April 2023; Social Media Today computes the September-2023-to-June-2025 change as an audience up about 25 percent and a moderation workforce down about 26 percent, which is an analyst's arithmetic over TikTok's disclosures rather than a TikTok statement or an audited figure. The site-level events: the entire 300-person Netherlands moderation team closed in September 2024, as stated in European Parliament question E-002454/2024; fewer than 500 Malaysian roles were cut in October 2024; about 300 Dublin trust-and-safety roles were notified in March 2025 and about 300 more proposed on 1 July 2026 out of more than 2,000 Irish staff; about 150 Berlin trust-and-safety and TikTok Live roles were announced on 10 August 2025; and approximately 430 London roles were put at potential risk with notices on 22 August 2025. TikTok stressed to Parliament that no changes had yet taken effect, and site-level figures vary between sources — Berlin is reported as both about 150 and 160 of roughly 400 staff, and London as over 400, approximately 430 and 439 — so each figure is pinned to its source and its date and none are summed.

Sources: tiktoktechnologylimited2025a, euobserver2026, socialmediatoday2025, letterfromtiktoktothechairof2025a, europeanfederationofjournali2025, rteandtheirishexaminer2026

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalTikTok's own account of the London reduction is that two thirds of it is not front-line moderation but the functions tha…

TikTok's own account of the London reduction is that two thirds of it is not front-line moderation but the functions that maintain moderation. Writing to Dame Chi Onwurah MP, Chair of the House of Commons Science, Innovation and Technology Committee, on 7 November 2025, TikTok stated: 'In the UK, there are approximately 430 roles at potential risk under this proposal', and that 'a third of these roles are in teams involved in the labelling of data for AI model training. Progress in the development of these models has significantly decreased the need for this kind of manual labelling. Another significant proportion of those potentially affected are in ancillary roles, for example training teams, whose duties include activities such as training moderators on our Community Guidelines... Around a third of those impacted are front line moderation teams.' Its earlier letter of 20 October 2025 named the labelling team as the AI Data Service and Operations team and said that 'the majority of those potentially affected are not in front line moderation roles'. The same 7 November letter describes a structural re-partition rather than only a reduction: TikTok is 'moving from a region-based structure of generalised moderators to one based on different types of products or types of risk, known as verticals' — harassment, misinformation, fraud — consolidated into fewer sites and supplemented by third-party specialists offering 'greater ability to rapidly expand' and 'greater levels of language-specific, follow-the-sun coverage', which it says is 'not a like-for-like replacement'. TikTok's published Year 3 risk assessment describes the same functions from the other side: safety-topic experts and local-market experts write and update the keyword lists and detection rules the classifiers execute, and moderators apply a Moderation Policy Framework. The language-indexed staffing table the DSA requires TikTok to publish therefore describes a partition the operator states it is leaving.

Sources: letterfromtiktoktothechairof2025a, letterfromtiktoktothechairof2025, tiktoktechnologylimited2025, tiktoktechnologylimited2025a

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalTikTok's published Year 3 systemic risk assessment describes the moderation pipeline in its own words. 'All video, photo…

TikTok's published Year 3 systemic risk assessment describes the moderation pipeline in its own words. 'All video, photo and text-based content uploaded to the Platform are subject to a real time, technology-based automated review. While a video is undergoing this review, it is visible only to the uploading user/creator.' Detection uses 'vision-based, audio-based, text-based and LLM-based' technologies together with keyword lists and natural-language processing; no model, vendor or product is identified anywhere in the record. Automated removal is 'applied when violations are the most clear-cut', and otherwise the item is routed to a human queue; high-view content may be routed for additional human review. Specialist misinformation moderators work against a repository of previously fact-checked claims from IFCN-accredited partners, and separate lanes handle illegal content, advertising and marketplace listings. A strikes policy escalates to account bans, and every decision is appealable. The assessment also records the internal governance chain: each risk module goes to a specialist Risk Assessment Review Group of senior internal stakeholders with Compliance input, then to the Online Safety Oversight Committee, a cross-functional leaders' steering group, then to the Board of Directors of TikTok Technology Ireland for review and approval, across twelve risk modules under four categories. It names as a standing moderation risk 'The risk that TikTok's content moderation systems and human moderators may: (i) over-moderate... or (ii) under-moderate'.

Sources: tiktoktechnologylimited2025, tiktoktechnologylimited2025a

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalThe enforcement, notice and appeal volumes for TikTok in the European Union, January to June 2025, all from its own DSA …

The enforcement, notice and appeal volumes for TikTok in the European Union, January to June 2025, all from its own DSA transparency report: 24,534,707 Community Guidelines removals of which 17,729,896 were automatic, 2,470,592 advertising removals and 829,861 TikTok Shop removals; 169,527,678 content restrictions, roughly seven times the removal layer; 2,781,470 service restrictions; and 4,906,735 account bans or suspensions of which 871,819 were automatic. The highest-volume removal policies were Regulated Goods and Commercial Activities at 9,485,450 (7,729,860 automatic), Sensitive and Mature Themes at 8,552,437 (6,305,376), Youth Safety and Well-Being at 5,944,993 (4,612,550), Mental and Behavioral Health at 5,581,036 (5,026,613, or 90.1 percent automatic) and Safety and Civility at 4,352,589 (2,855,504, or 65.6 percent, the lowest automation share of the major policies). On the notice side: 308,755 illegal-content reports from EU users covering 151,354 unique items, of which 26,512 were actioned as unlawful and 15,365 as policy breaches — 27.7 percent of unique reported items actioned at all — at a median decision time under 17 hours on policy grounds and under 21 hours on legal grounds; 3,976 removal orders from Member State authorities with a median action time under 3 hours, and 782 such requests in the following period led by France and Romania; and 82 trusted-flagger reports under Article 22, against 22,429 user reports of illegal hate speech and only five trusted-flagger hate-speech reports in the previous period on a civil-society transcription. In the July to December 2025 period EUobserver records approximately 112 million removals with 99.3 percent removed before any user report and 714,000 user-reported incidents. Every figure here is TikTok's own, published because the Digital Services Act compels it, and the independent auditor found the controls over the data behind the transparency report not sufficient and appropriate.

Sources: tiktoktechnologylimited2025a, euobserver2026, internationalnetworkagainstc2025, kpmgadvisoryn2025

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalTikTok's own automation figures rose across the substitution period under changing definitions, and they are three separ…

TikTok's own automation figures rose across the substitution period under changing definitions, and they are three separately sourced statements rather than a trend. Its DSA transparency report for January to June 2025 gives 17,729,896 automatic removals of 24,534,707 total Community Guidelines removals, or 72.3 percent. Its letter to the House of Commons Science, Innovation and Technology Committee of 20 October 2025 states that '86% of the content we remove is now removed by automation', attributed to an April-to-June 2025 transparency report from a different report family. Its sixth DSA transparency report, for July to December 2025, states that 'Automated systems actioned 93.8% of all violating content without human review', with '97.6% of automated enforcement decisions being confirmed as correct', against approximately 112 million pieces of violating content — and that period was the first to include comment enforcement volumes, so the denominator is not the previous periods'. TikTok itself cautions that the later report 'captures a broader range of automated enforcement actions, including automated LIVE enforcement, when compared with our previous reports'. Social Media Today reads the same disclosures as automated detection for mental and behavioural health concerns rising from 49 percent in 2023 to 90 percent, and for youth safety from 38 percent to 77 percent. The three headline percentages measure different things over different denominators and are never presented here as a time series.

Sources: tiktoktechnologylimited2025a, letterfromtiktoktothechairof2025, tiktok2026, euobserver2026, socialmediatoday2025

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalTikTok's contest machinery under the Digital Services Act runs on two levels with very different force, and both are mea…

TikTok's contest machinery under the Digital Services Act runs on two levels with very different force, and both are measured in its own January-to-June 2025 report. Internally, 3,075,758 appeals came from uploaders and advertisers and 1,054,432 from users who had reported content — about 22,700 a day against 4,596 moderators — producing 1,359,823 pieces of content reinstated or unrestricted and 61,095 items removed after a reporter's appeal, at a median decision time under two hours on both tracks; TikTok warns the reinstatement count does not map one to one onto the period's appeals. Externally, under Article 21, out-of-court dispute settlement bodies received 1,121 complaints and closed 498 in period: the body agreed with TikTok in 113 cases, disagreed in 106, and 189 closed without a formal decision. Of the 106 decisions that went against TikTok, TikTok implemented the decision in 29 — 27.4 percent — at a median handling time of about 26 days against the internal channel's under two hours. That gap is the difference between a body that can decide and a body that can bind. Article 20 requires the internal channel to produce reasoned decisions under the supervision of appropriately qualified staff and not solely on the basis of automated means, and the Commission repeated that requirement in its written answer on moderation resourcing. The independent auditor's conclusion on Article 20(4) is negative: complaint records could not be retrieved 'due to limitations in documentation retention', so it 'was unable to confirm that all complaints were handled in a timely, non-discriminatory, diligent, and non-arbitrary manner'.

Sources: tiktoktechnologylimited2025a, kpmgadvisoryn2025, europeanparliamentwrittenque2024

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalThe independent audit the Digital Services Act itself mandates returned a NEGATIVE, qualified opinion on TikTok Technolo…

The independent audit the Digital Services Act itself mandates returned a NEGATIVE, qualified opinion on TikTok Technology Limited for the year 1 July 2024 to 30 June 2025. KPMG Advisory N.V.'s Article 37 assurance report, dated 29 August 2025, covers 90 specified requirements and reaches four negative conclusions, all inside the moderation record. On Article 16(6) it identified notices where TikTok did not perform moderation actions and could not evidence the monitoring controls over the interface between notice intake and the moderation systems, so it 'could not obtain sufficient assurance to support the completeness of the total population of notices'. On Article 20(4) complaint records could not be retrieved 'due to limitations in documentation retention', so KPMG 'was unable to confirm that all complaints were handled in a timely, non-discriminatory, diligent, and non-arbitrary manner'. On Article 24(5) duplicate statements of reasons were transmitted to the Commission's DSA Transparency Database, producing 'more records... than the actual number of decisions taken', until a remediation in May 2025, and post-remediation sampling still found statements of reasons that were never transmitted. And the advertisement repository was found defective under Article 39(3). Six requirements were DISCLAIMED because they sit under the Commission's open proceedings — Articles 28(1), 34(1), 34(2), 35(1), 39(1) and 40(12) — with material observations on two. On Article 42(2), the article that mandates the per-language human-resources figures and the accuracy indicators themselves, the conclusion is 'Positive with comments', the comment being that 'internal controls concerning data accuracy and completeness monitoring, between the various source systems and Transparency Report are not sufficient and appropriate'. The auditor holds no remedy power of any kind.

Sources: kpmgadvisoryn2025, tiktoktechnologylimited2025a

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalTikTok restated the moderator counts in Annex D of its January-to-June 2025 DSA transparency report on 15 April 2026, ne…

TikTok restated the moderator counts in Annex D of its January-to-June 2025 DSA transparency report on 15 April 2026, nearly eight months after publication, noting that the values 'have been updated with the correct values'. The mandated record of how many people moderate content in each EU official language was wrong when it was first published, and the independent auditor's comment on the article that mandates that record is that the internal controls over data accuracy and completeness monitoring between the source systems and the transparency report are not sufficient and appropriate. The consequence for anyone reading the series is that headcount figures from different report editions are not safely comparable: the July-to-December 2024 per-language figures reach this file through a civil-society transcription by INACH rather than the original document — English 1,524, German 532, Spanish 531, French 509, Italian 290, Portuguese 160, Polish 146, Dutch 99, Romanian 99, Swedish 72, Czech 53, Hungarian 51, Greek 50, Bulgarian 38, Slovenian 37, Slovak 33, Finnish 31, Croatian 29, Latvian 22, Lithuanian 19, Estonian 17 and Danish 15, totalling 4,357 across 22 named languages — one of the two editions was restated, and TikTok groups several language teams and folds several non-EU languages into its totals. The aggregate is nearly flat between the two while the smallest columns fall hard and the largest rise; the direction and the shape of that distribution are usable and the differences are not. INACH's own judgement of the sector's reports is that TikTok 'has almost no data available specifically on hate speech as a separate category' and that none of the reports reviewed provides a breakdown of hate-speech enforcement by country or by protected characteristic.

Sources: tiktoktechnologylimited2025a, internationalnetworkagainstc2025, kpmgadvisoryn2025

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalThe House of Commons Science, Innovation and Technology Committee asked TikTok on 28 October 2025 six questions about th…

The House of Commons Science, Innovation and Technology Committee asked TikTok on 28 October 2025 six questions about the UK trust-and-safety reductions, including the total job losses, the core responsibilities of the roles and how they support moderation, whether a risk assessment of the job losses for UK user safety had been conducted and what its outcome was, whether third-party moderation teams would replace the responsibilities, Ofcom's response, and how UK users would be safeguarded by staff in other countries. The Chair's letter quoted TikTok's own earlier written evidence back to it — 'tens of thousands' of safety professionals working 24/7 and 'In 2024, we invested over $2 billion in our Trust and Safety efforts' — and its oral evidence of 25 February 2025 distinguishing what automation handles well, 'pornographic material, blood and that kind of thing', from content referred to 'human moderators who have to use their nuance, skills and training to be able to rule on other elements that can include hateful behaviour and misinformation'. TikTok replied on 7 November 2025 describing an internal analysis expecting improvements in 'speed of moderation' and 'efficacy of moderation (i.e. how often the moderation decision is the correct one)', and supplied no data. On 13 November the Committee published the reply, stating that 'TikTok did not share its data or risk assessment that justified this in its reply to the Committee Chair', with Chair Dame Chi Onwurah commenting that 'TikTok's response represents a commitment to reducing staffing levels in favour of increasing the use of AI to moderate content on its platform. But TikTok have come up empty to show that this transition to AI won't lead to more harms for its users', and that 'TikTok refers to evidence showing that their proposed staffing cuts and changes will improve content moderation and fact-checking - but at no point do they present any credible data on this to us.' The Committee's instrument is scrutiny and publication; it has no authority over the moderation design. TikTok confirmed in the same correspondence that it had given Ofcom prior notice of the London proposals and that it has been regulated by Ofcom since 2021 under the Video Sharing Platform regime and now under the Online Safety Act; Ofcom's response is not on the public record.

Sources: ukhouseofcommonsscience2025, letterfromtiktoktothechairof2025a, letterfromtiktoktothechairof2025

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalTwo labour disputes over the reductions are live and neither has been adjudicated. In Berlin, ver.di held five one-day s…

Two labour disputes over the reductions are live and neither has been adjudicated. In Berlin, ver.di held five one-day strikes in July 2025, beginning 23 July, and a four-day strike from 23 September 2025, over the announcement of 10 August 2025 that the Trust and Safety and TikTok Live teams of about 150 people would close; its demands were a collective agreement with severance worth three years' salary and a twelve-month notice extension, under the slogan 'We trained your machines, pay us what we deserve!'. Jacobin, reporting worker accounts, gives 160 of roughly 400 Berlin staff affected, moderators reviewing 800 to 1,000 videos a day and a typical tenure of three to four years before psychological strain ends it — testimony relayed by a partisan outlet and the only per-head capacity referents anywhere in this record, since TikTok publishes none. TikTok's response, carried in the Business & Human Rights Resource Centre tracker, is that the changes would 'streamline workflows and improve efficiency' with 'full commitment to protecting safety and integrity'. Dismissal cases went to the Berlin Labour Court and no outcome is on the public record. In London, the redundancy notices of 22 August 2025 preceded a scheduled union-recognition ballot with UTAW, a branch of the Communication Workers Union, and on 19 December 2025 two moderators supported by Foxglove and UTAW and represented by Leigh Day sent a pre-action letter alleging unlawful detriment and automatic unfair dismissal, citing internal TikTok documents of May 2025 referencing the 'complexity and volume of certain categories of moderation that require human judgment for safety and compliance, rather than automated tools'. TikTok told Parliament the union-timing claims are 'categorically untrue', that the decisions were made globally, and that it had written to the CWU expressing 'regret about these timescales' while remaining open to re-engaging after consultation. These are allegations and denials, not findings; no tribunal has ruled. In Ireland the Communications Workers' Union objected to the July 2026 Dublin proposal on user-safety grounds, arguing that 'only quality jobs can provide the level of rapid policy response, moderation oversight, and overall safety that are required' while TikTok is under Commission and national investigation; that is a union contention about a future effect, not a measurement.

Sources: europeanfederationofjournali2025, jacobin2025, foxglove2025, letterfromtiktoktothechairof2025, rteandtheirishexaminer2026

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalTikTok's published Year 3 systemic risk assessment, dated 28 August 2025 — eighteen days after the Berlin closure announ…

TikTok's published Year 3 systemic risk assessment, dated 28 August 2025 — eighteen days after the Berlin closure announcement and six days after the London redundancy notices — records no change to its Fundamental Rights inherent risk score, which stands at Medium-High and Likely and is described as 'consistent with TikTok's score in Year 2', and contains no reference to the trust-and-safety workforce reduction, restructuring or redundancies. A keyword sweep of the published report for redundancy, restructuring, headcount, layoff, workforce, union and labour returns nothing on the reduction announced in the same month. This is an observation about what the published text contains and not proof that the matter was never assessed internally: the document is headed 'Confidential' and what is public is its published version. The assessment does name over-moderation and under-moderation by 'content moderation systems and human moderators' as a standing moderation risk, and it records the internal approval chain through specialist Risk Assessment Review Groups, the Online Safety Oversight Committee and the Board of Directors of TikTok Technology Ireland.

Sources: tiktoktechnologylimited2025

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalNo published independent measurement of a moderation-quality effect from TikTok's staffing substitution exists, and the …

No published independent measurement of a moderation-quality effect from TikTok's staffing substitution exists, and the three available quality signals are each weak in a different way. The first is TikTok's own indicator, which is an overturn-rate complement rather than a population error rate: it defines accuracy as 'the proportion of content where the original enforcement decision was upheld or maintained' and error as 'the proportion... overturned', reporting automated accuracy of 99.2 percent and error of 0.8 percent for January to June 2025 against 99.12 percent the previous half, with per-Member-State error running from 0.3 percent in Slovakia and Bulgaria to 1.6 percent in Austria and Germany and France at 1.4 percent, and 97.6 percent of automated enforcement decisions 'confirmed as correct' for July to December 2025. That statistic is conditional on somebody appealing and is computed by the party that made the original decision, and in the same six months 3,075,758 uploader appeals produced 1,359,823 reinstatements or lifted restrictions on a different denominator that cannot be turned into an error rate. The second is the independent Article 37 auditor, which could not establish the completeness of the notice population or of the complaint population from which that sample is drawn. The third is negative: on 24 October 2025 the European Commission preliminarily found that TikTok and Meta may have put in place burdensome procedures and tools for researchers to request access to public data, 'often leav[ing] them with partial or unreliable data' — so the outside parties who could measure a quality effect independently are the ones the regulator says cannot get reliable data. Nothing in this file attributes any change in moderation quality to the staffing substitution, because no such measurement has been made. TikTok's claim to Parliament that its platform has 'the lowest error rates and highest accuracy rates among all major platforms' is an operator claim across disclosures that civil-society and academic reviewers describe as mutually incomparable, and the ranking is not repeated here.

Sources: tiktoktechnologylimited2025a, tiktok2026, kpmgadvisoryn2025, europeancommission2025c, ukhouseofcommonsscience2025, internationalnetworkagainstc2025

Appears on: /domains/cases/tiktok-dsa-moderation

EmpiricalWikipedia's edit-scoring service publishes a damage or revert-risk probability for essentially every edit as it is saved…

Wikipedia's edit-scoring service publishes a damage or revert-risk probability for essentially every edit as it is saved and has no path of its own by which it can act on one. Aaron Halfaker and R. Stuart Geiger, writing the system up for CSCW in 2020, describe it as built to decouple four activities normally performed by the same engineers: curating training data, building models, auditing predictions, and building the interfaces or bots that act on predictions. The operator did not gate access either, on its own stated reasoning that 'Given the open API, there is no barrier where we can selectively decide who can request a score from a classifier.' Acting on the score is done by separately governed agents: a volunteer-run bot on English Wikipedia, an operator-built agent local administrators switch on, tool-assisted patrollers working score-ranked queues, and ordinary editors with score-driven filters enabled in their own preferences. At the time of that paper the estate ran roughly 110 classifiers across 44 languages in four families; the ORES infrastructure has since been retired and what runs today is Lift Wing, with ores-legacy.wikimedia.org as a compatibility endpoint in front of it, serving a revert-risk family that replaced the older edit-quality models. Verified live on 28 August 2026: the legacy host answers, self-identifies as the 'ORES legacy service' and still returns scores, and the Lift Wing endpoint returns a revert-risk score for an arbitrary revision without any credential.

Sources: halfaker2020, albon2023

Appears on: /domains/cases/wikipedia-ores

EmpiricalThe scores, the models, the training-data provenance and the deliberation about all three are public here, and that open…

The scores, the models, the training-data provenance and the deliberation about all three are public here, and that openness is what makes every outside measurement in this case possible. Scores are queryable for any revision by anyone without a credential, and were from 2015; the models and their code ship under the Apache 2.0 licence and the training pipelines were built as reproducible Makefiles so a third party can rebuild an equivalent model. The operator publishes model cards as a standing institutional practice, written to the Mitchell et al. framework, on a public index of proposed, production and deprecated models; each card is asked to state why the model was made, its proper and improper uses, and its evaluation scores especially for marginalised groups, and each card page carries a talk page, so the documentation is also the venue where it is argued about. Every operating point's precision and recall is published in patroller-facing help text before anyone chooses among them. Against that openness sit two documented limits: historical predictions were retained only until the end of 2019, and because no operator adjudicates anything, a false-positive report is a public wiki page rather than a case with a disposition, so no reversal rate on contested reverts is published the way an overturn rate appears in a mandated transparency report.

Sources: halfaker2020, wikimediafoundationb, wikimediafoundation, levonian2024

Appears on: /domains/cases/wikipedia-ores

EmpiricalEvery operational lever in this deployment is held by the governed community rather than by the model's operator, and th…

Every operational lever in this deployment is held by the governed community rather than by the model's operator, and the record shows each of them being exercised. A bot may not edit at all before English Wikipedia's Bot Approvals Group has approved it and it has run a trial, and any administrator may block one that malfunctions. The rollback right that the fastest patrolling tool requires is granted and revoked by administrators, and that tool's own documentation states it 'is not intended for new Wikipedia users' and that misuse 'may result in revocation of rollback permissions or being blocked from editing'; the cross-wiki patrolling tool gates its global queue on global rollback, steward or global sysop rights, or at least 1,000 global edits and no active block, and its documentation records no scoring integration. Automoderator, the operator's own reverting agent, 'will not begin running until a local administrator turns it on', its caution level is written to MediaWiki:AutoModeratorConfig.json through Special:CommunityConfiguration — an ordinary watchlistable wiki page — and creating a false-positive reporting page is a required step of deployment. It is live on twelve Wikipedias (Indonesian, Turkish, Ukrainian, Vietnamese, Afrikaans, Bengali, Azerbaijani, Chinese, Spanish, Italian, Dutch and Albanian, first Turkish in June 2024) and not on English, whose project page records the team's own position that 'If English Wikipedia editors don't want to use Automoderator, that's fine!' When a volunteer wired a score directly to automatic reversion on Spanish Wikipedia, the community stopped it without the operator: PatruBOT was crowd-audited on ordinary wiki pages, judged to be erring too often, and had its account blocked by an administrator — an episode the ORES team recorded as 'entirely a community governed activity that required no intervention of our team or the Wikimedia Foundation staff', with a successor, SeroBOT, later resuming at a higher confidence threshold.

Sources: halfaker2020, englishwikipediab, wikimediafoundationa

Appears on: /domains/cases/wikipedia-ores

EmpiricalThe acting agent's control variable is a human-chosen error budget rather than a score, and both halves of the trade it …

The acting agent's control variable is a human-chosen error budget rather than a score, and both halves of the trade it buys are published. ClueBot NG's documentation states that 'The threshold is not randomly chosen by a human, but is instead calculated to match a given false positive rate... A human selects a false positive rate, which is the percentage of constructive edits incorrectly classified as vandalism.' At the current 0.1 percent setting the bot catches approximately 40 percent of vandalism; at the previous 0.25 percent setting it caught approximately 55 percent, so roughly fifteen points of catch rate were given up to halve the wrongful-reversion rate. Chosen instead for total accuracy, it classifies over 90 percent of edits correctly. Post-processing filters on self-reverts, edit count and warning share cut actual false positives below the stated rate before any revert happens. These figures are the volunteer maintainers' own, computed on a held-out, human-reviewed slice of their own dataset: methodologically described and not independently audited. The agent's scale and standing are separately verifiable: registered 20 October 2010, 6,682,890 edits as of 28 August 2026, holding the bot, reviewer, rollbacker and autoconfirmed user groups, all community-granted, and stoppable by any administrator editing a run page to 'False'. The operator's own agent applies its own published carve-outs, never reverting administrators, global sysops, stewards or bots, self-reverts, reverts of its own actions, or new page creations.

Sources: englishwikipedia, wikimediafoundationa

Appears on: /domains/cases/wikipedia-ores

EmpiricalThe deployed classifier's own bias against anonymous editors was measured independently and is large. Mykola Trokhymovyc…

The deployed classifier's own bias against anonymous editors was measured independently and is large. Mykola Trokhymovych, Muniza Aslam, Ai-Jou Chou, Ricardo Baeza-Yates and Diego Saez-Trumper reported at KDD in 2023 a Disparate Impact Ratio of 20.02 for the deployed ORES model against a base-rate ratio of 7.93 in the same data, meaning it flagged anonymous editors well out of proportion even to their genuinely higher revert rate; the successor multilingual model reaches 9.54 with the same user features and 1.98 to 3.08 without them. The deployed model's area under the curve was 0.84 with precision at recall 0.75 of 0.22 on an unbalanced holdout, against 0.75 and 0.07 for a deliberately unfair baseline that simply reverts every anonymous edit. Model accuracy correlates negatively with a language edition's share of anonymous editors in both systems. The operator states the same limitation in its own documentation: the model card for the language-agnostic revert-risk model — the model its own reverting agent uses — says the model 'may exhibit bias against edits from new users, temporary accounts, or IP edits' and recommends the multilingual model instead for anonymous edits in the languages that model covers, while forbidding use of its predictions as ground truth for training other models and forbidding scoring a page's first revision.

Sources: trokhymovych2023, wikimediafoundationb

Appears on: /domains/cases/wikipedia-ores

EmpiricalPublishing the flag demonstrably changed the human decision, and the measured direction was toward more equal treatment …

Publishing the flag demonstrably changed the human decision, and the measured direction was toward more equal treatment — a separate finding from the classifier's own bias, at a different layer, and the two must be carried together. Nathan TeBlunthuis, Benjamin Mako Hill and Aaron Halfaker ran a regression discontinuity across 23 Wikipedia language editions from January 2019 to March 2020, using the RCFilters threshold cutoffs as the discontinuity and revisions within 0.03 of a cutoff. At the 'maybe damaging' cutoff, being flagged raised revert probability from 13.5 to 19.2 percent for unregistered editors and from 4.6 to 14.3 percent for registered ones; at 'likely damaging', from 33.5 to 50.2 percent and from 15.5 to 44.5 percent respectively. Because flagging moved the under-scrutinised group more than the over-scrutinised one, aggregated across thresholds it INCREASED demographic parity between registered and unregistered editors. It also lowered the odds that a revert was itself controversial for unregistered editors, from 3.08 to 2.81 percent at 'likely damaging' and 3.33 to 2.92 percent at 'very likely damaging', which the authors read as evidence that flagging lowered the decision system's false-positive rate. The same authors state plainly in the same paper that 'ORES encodes biases against unregistered editors and - to a lesser extent - against editors without user pages'. The honest reading of the two findings together is a biased classifier whose published score reduced a larger human bias, and neither half stands alone.

Sources: teblunthuis2021, trokhymovych2023

Appears on: /domains/cases/wikipedia-ores

EmpiricalThe costs this deployment was designed against were measured on Wikipedia itself before the scoring service existed, and…

The costs this deployment was designed against were measured on Wikipedia itself before the scoring service existed, and are never a measured effect of it. Aaron Halfaker, R. Stuart Geiger, Jonathan T. Morgan and John Riedl reported in 2013 that the share of good-faith newcomers whose first-session edit was reverted rose from 6.1 percent in the first half of 2006 to 18.2 percent in the first half of 2007, that two-month survival of those newcomers fell from 25.6 percent to 11.7 percent within a year and did not recover, and that tool-mediated rejection of them rose from about 0 percent in 2006 to about 40 percent in 2010, with both rejection and tool-mediated rejection significant negative predictors of newcomer survival. The harm sat in the interaction as much as in the classification: reciprocation of a reverted newcomer's attempt to open a discussion averaged 7 percent for editors using Huggle, about 30 percent for Rollback, 53 percent for Twinkle and 56 to 67 percent for manual reverters, and 2,250 discussion attempts from 918 registered editors were addressed to an algorithmic editor that could not reply. The 2020 ORES paper concedes that after this research 'the often-hostile quality control processes that were designed over a decade ago remain largely unchanged'. Against that, the same score is also used prosocially: the 'Very likely good' filter, about 99 percent precise at over 90 percent recall, is documented as a way to find good-faith newcomers to thank, and a Wiki Education tool asks the article-quality model to re-score a student's draft with one more citation, header or image in order to recommend the most productive next edit.

Sources: halfaker2013, halfaker2020, wikimediafoundation

Appears on: /domains/cases/wikipedia-ores

EmpiricalAuditing this deployment is a tooled activity for the governed population rather than a privilege of the operator, and i…

Auditing this deployment is a tooled activity for the governed population rather than a privilege of the operator, and its limit is a retention decision. ORES-Inspect, described by Zachary Levonian, Lauren Hagen, Lu Li, Jada Lilleboe, Solvejg Wastvedt, Aaron Halfaker and Loren Terveen at the 2024 Wiki Workshop, is an open-source Toolforge interface that lets any editor sample two disagreement quadrants — 'Unexpected Reverts', where the model called an edit fine and the community reverted it, and 'Unexpected Consensus', where the model called an edit damaging and the community left it standing — across the 35.6 million non-bot English Wikipedia edits of 2019, with the prediction as it was made at the time, and turn a single noticed misclassification into a quantified false-positive or false-negative rate for a chosen slice such as newcomers, LGBT-history pages or stubs. The design targets exactly the loop the deployment carries: reverts become training labels for the next model, so an over-flagging threshold can teach itself. The same paper records that historical predictions were retained only until the end of 2019, so an audit of what the model said at the moment of a past edit cannot run past that year even though the live scores have always been public.

Sources: levonian2024, halfaker2020

Appears on: /domains/cases/wikipedia-ores

EmpiricalThe label definition itself is contestable by the governed population here, and the training data is collected in the op…

The label definition itself is contestable by the governed population here, and the training data is collected in the open. Tzu-Sheng Kuo, Aaron Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein and Haiyi Zhu reported at CHI 2024 on Wikibench, a system that puts AI evaluation-data curation through Wikipedia's ordinary talk-page and consensus machinery; the authors report that datasets curated this way 'can effectively capture community consensus, disagreement, and uncertainty' and that participants used it to refine label definitions, set data inclusion criteria and author data statements. The acting agent's own training labels come from a public Toolforge review interface whose maintainers state their aim plainly — 'We need volunteers to help review edits and classify them as either vandalism or constructive. We hope to eventually completely replace our current dataset with a random sampling of edits, reviewed and classified by volunteers' — with each edit in the trial slice behind their published statistics reviewed by at least two humans, and with the same documentation conceding that 'Our current dataset has some degree of bias, as well as some inaccuracies.' The older edit-quality models were trained on per-wiki volunteer labelling campaigns; the successor language-agnostic model was trained on published MediaWiki History and Wikitext History tables for January 2022 to January 2023 excluding bot edits on a 70/30 split, and the multilingual model on 8.6 million revisions from January to July 2022, sampled up to 300,000 per language, with a 17 percent unregistered-edit rate and an 8 percent revert rate in the training data.

Sources: kuo2024, englishwikipedia, wikimediafoundationb, halfaker2020

Appears on: /domains/cases/wikipedia-ores

EmpiricalLanguage coverage is the honest limit of this deployment's transparency, and the inequality runs the wrong way. The bett…

Language coverage is the honest limit of this deployment's transparency, and the inequality runs the wrong way. The better-calibrated multilingual revert-risk model covers 47 languages; the language-agnostic model runs on any of more than 250 language editions and is the model whose own card warns about bias against new users, temporary accounts and unregistered editors — and it is the model the operator's reverting agent actually uses. Across both systems, model accuracy correlates negatively with a language edition's share of anonymous editors, so the communities that lean most on anonymous contribution get the least accurate scoring. The estate also narrowed: roughly 110 classifiers in four families across 44 languages have given way to the revert-risk family. Precision figures are per-wiki and per-threshold and travel badly, which the operator's own help page demonstrates: on Polish Wikipedia the 'Likely have problems' filter captures 91 percent of problem edits against 34 percent for the corresponding English filter, and Polish Wikipedia therefore 'does not need - or have' the broader, noisier filter English Wikipedia relies on. English Wikipedia meanwhile keeps an agent trained on English Wikipedia alone, and the operator's own project page names running both as one of three choices open to that community.

Sources: wikimediafoundationb, trokhymovych2023, wikimediafoundationa, wikimediafoundation

Appears on: /domains/cases/wikipedia-ores

EmpiricalFour outages of the fastest automated tier in the first half of 2011 give this deployment a measured counterfactual, and…

Four outages of the fastest automated tier in the first half of 2011 give this deployment a measured counterfactual, and what it measures is latency rather than coverage. R. Stuart Geiger and Aaron Halfaker, at WikiSym 2013, analysed ClueBot NG's downtime on 15 to 18 February, 13 to 17 March, 29 March to 7 April and 15 April to 1 May 2011. Comparing only Wednesdays and Thursdays in order to control for the weekly editing rhythm, median time-to-revert rose from 744 seconds with the bot running to 1,286 seconds with it down, and the geometric mean from 941 to 1,674 seconds. The proportion of edits that were reverts fell significantly during downtime (chi-square 115.9, p<0.001), but the proportion of revisions EVENTUALLY reverted did not differ (chi-square 0.64, p=0.43): the human tiers absorbed the work at a slower rate rather than losing it. The authors were careful not to read this as the bot being dispensable, asking instead what the workaround cost the editors who performed it. The same paper documents the reviewer tiers whose different clocks make that absorption possible: fully automated bots reverting within seconds, tool-assisted humans mostly within a minute, manual browser reverts between a minute and a day, and batch scripts on an idiosyncratic scatter. The measurement is excellent evidence for the shape of the effect and weak evidence for its present magnitude: it concerns one bot on one wiki at a scale and tooling mix that no longer obtain.

Sources: geiger2013

Appears on: /domains/cases/wikipedia-ores

EmpiricalThe transparency mechanism the peer-reviewed record identifies as the core of participatory machine learning here was re…

The transparency mechanism the peer-reviewed record identifies as the core of participatory machine learning here was removed in an infrastructure migration as an unused feature, and nobody decided against it. ORES exposed threshold optimisations in a machine-readable format so a wiki's tool developer could ask for the maximum filter rate at a stated recall for their own wiki and their own model version; the published worked example on English Wikipedia was a threshold of 0.32 giving a filter rate of 0.89, a false-positive rate of 0.087, precision of 0.23 and recall of 0.75. Probed live on 28 August 2026, that query returns: 'model_info query parameter is not supported by this endpoint anymore.' Chris Albon announced on the wikitech-l list on 3 August 2023 that the ORES API endpoint would move onto Lift Wing by 30 September 2023, that the Foundation wanted zero traffic on the old endpoint by January 2024, that 'The servers that run ORES are at the end of their planned lifespan and so to save cost we are going to shut them down in early 2024', and that 'The ores-legacy endpoint is not a 100% replacement for ores, we removed some very old and not used features.' The migration remains an open programme: a Phabricator task opened on 5 March 2026 records the deprecation guidance as scattered across three wiki pages and 'difficult to find and to maintain', and the MediaWiki modernization page carries its own warning that it 'contains outdated information that may not accurately reflect the current state of Wikimedia ML systems'. The published-score channel survived the migration; the published-fitness-statistics channel did not survive on this endpoint.

Sources: albon2023, halfaker2020

Appears on: /domains/cases/wikipedia-ores

EmpiricalThe correction channel for a wrongly reverted editor is an ordinary public wiki page rather than a ticket nobody outside…

The correction channel for a wrongly reverted editor is an ordinary public wiki page rather than a ticket nobody outside can see, and its openness is exactly what makes it unmeasured. The acting agent's documentation tells anyone reverted in error to redo the edit, remove the warning and report the false positive, and points at a page that also holds the full public list of reported false positives; that page was reachable when checked on 28 August 2026. The operator's own agent makes creating such a page a required step of deployment and links it from the talk-page message, from the page history and from the user's contributions beside the ordinary Undo and Thank actions, in a default message that reads: 'Because the model I use is not perfect, it sometimes reverts good edits. If you believe the change you made was constructive, please report it here.' The agent's team states an intention to investigate retraining on reported false positives. What does not exist is a disposition: because no operator adjudicates anything, a false-positive report is a wiki page rather than a case, so this deployment publishes no reversal rate on contested reverts and cannot supply the overturn statistic that a mandated platform transparency report supplies. Its correction channel is more open and less measured than a regulated one.

Sources: englishwikipedia, englishwikipediaa, wikimediafoundationa

Appears on: /domains/cases/wikipedia-ores

EmpiricalThe scale this deployment exists to address is published, and so is the labour arithmetic behind it. Queried from the Wi…

The scale this deployment exists to address is published, and so is the labour arithmetic behind it. Queried from the Wikimedia Analytics REST API on 28 August 2026, English Wikipedia alone took between 2.59 and 2.94 million non-bot edits to content pages per month across the twelve months to July 2026 — 2,839,865 in July 2026 — on the order of 85,000 to 95,000 human content edits a day on one of more than 250 language editions the models serve, against an operator-stated base rate of fewer than 5 problem edits in 100. The ORES paper's all-editions figure was about 290,000 edits a day, and its arithmetic is that reviewing that at an aggressive ten revisions a minute is about 483 volunteer labour hours daily, which a model filtering ninety percent of the stream reduces to about 48.3 — turning 240 volunteers at two hours a day into 24, and, for a small wiki, turning the task into one or two part-time volunteers. The service that does the filtering ran at 50 to 125 external requests a minute in steady state with bursts to 400 to 500 a second, precaching requests roughly an order of magnitude higher because a scoring job starts for nearly every edit, an approximately 80 percent cache hit rate, and most predictions computed in about a second. The team behind it 'never had more than 3 paid staff and 3 volunteers at any time, and no more than 2 requests typically in progress simultaneously', against roughly 66,000 monthly active English Wikipedia editors at the time.

Sources: wikimediafoundation2026, halfaker2020, wikimediafoundation

Appears on: /domains/cases/wikipedia-ores

EmpiricalThere is no litigation, no regulator, no court, no consent order and no statutory transparency mandate anywhere in this …

There is no litigation, no regulator, no court, no consent order and no statutory transparency mandate anywhere in this deployment's record, and that absence is a structural finding rather than an absence of controversy. Litigation posture: None. Every accountability artefact here — the open scoring interface, the published operating points, the model cards stating their own biases, the public false-positive pages, the community audit tooling — exists because the operator and the self-governing volunteer communities chose it, and could be withdrawn the same way; the removal of the machine-readable threshold-statistics query in the 2023 to 2024 infrastructure migration is a small, dated instance of exactly that. Two further honesty notes belong with any description of the strength of this record. First, this is not a solved or harm-free deployment: its own operators concede the hostile quality-control processes documented in 2013 remain largely unchanged, the deployed model's card warns it may be biased against new users, temporary accounts and unregistered editors, and the acting agent's maintainers concede their dataset carries bias and inaccuracies. Second, the peer-reviewed evidence base leans on a small overlapping author group — Halfaker and Geiger appear on the ORES paper, the outage paper and the newcomer-decline paper, and Halfaker also co-authors the flagging-fairness paper and the audit probe, having been a Wikimedia Foundation employee for most of the period covered — with the KDD evaluation and the audit probe the clearest checks outside that lineage, and the KDD paper the one that measures the deployed model unfavourably.

Sources: halfaker2020, wikimediafoundationb, englishwikipedia, trokhymovych2023, levonian2024, albon2023

Appears on: /domains/cases/wikipedia-ores

EmpiricalYouTube's Content ID is a fingerprint-matching copyright claiming system whose deciding party is an outside rights-holde…

YouTube's Content ID is a fingerprint-matching copyright claiming system whose deciding party is an outside rights-holder rather than the platform. Rights-holders admitted through an eligibility gate deliver reference files; YouTube derives fingerprints and compares every upload against the reference store; on a match the partner's pre-set match policy fires automatically — block, monetize or track — and the policy can differ country by country on the same video, with no case-by-case human decision on the claiming side. YouTube states that it 'is not in a position to mediate this type of dispute as we are not a court of law', and that when a matter reaches a legal removal request 'the ownership issue has exited the Content ID claim and dispute system built by YouTube, and enters the legal removal and remediation process defined by the DMCA and similar applicable laws'. The volume, on the deployer's own published reporting: 2,502,941,368 Content ID claims in calendar 2025, up 14 per cent on approximately 2.2 billion in calendar 2024, against 722,649,569 in the first half of 2021. TWO DISTINCT 99-PER-CENT FIGURES appear in these reports and mean different things. Content ID's share of ALL copyright actions taken on the platform was 99.43 per cent in 2024 and 99.48 per cent in 2025. The share of Content ID's OWN claims generated by automated matching rather than by a partner's manual claiming feature was 'over 99 per cent' in every reported period, with manual claiming at 0.4 per cent in the first half of 2021, 'fewer than 0.5 per cent' in the second half of 2022, and 0.31 per cent — about 6.9 million claims — in 2024. YouTube states the system cannot assess fair use: 'it's impossible for matching technology to take into account complex legal considerations like fair use or fair dealing.' Every quantitative figure here is the deployer's own accounting of its own system, published voluntarily in the United States, and none has been independently verified.

Sources: youtubegoogle2021, torrentfreak2025, youtubehelp2026, u2020

Appears on: /domains/cases/youtube-content-id

EmpiricalAbout half a per cent of Content ID claims are ever disputed, the rate is stable across five years and a tripling of vol…

About half a per cent of Content ID claims are ever disputed, the rate is stable across five years and a tripling of volume, and it is not an error rate. Verified from YouTube's four machine-readable Copyright Transparency Report editions: 3,698,019 disputes on 722,649,569 claims in the first half of 2021 (0.512 per cent) with 38,864 copyright removals originating from disputes; 3,810,395 on 759,540,199 in the second half of 2021 (0.502 per cent) with 43,198 removals; 3,690,786 on 757,993,607 in the first half of 2022 (0.487 per cent) with 24,931 removals; and 826,242,639 claims in the second half of 2022. Four years later the 2025 edition reports 12,840,608 disputes on 2,502,941,368 claims, 0.51 per cent. The share of disputes resolved in the uploader's favour was 'over 60 per cent' in the first half of 2021 and the second half of 2022, 'over 65 per cent' in the 2024 edition and 67.42 per cent in the 2025 edition — and YouTube's own definition counts a dispute as resolved for the uploader when the claimant 'either voluntarily released the claim or did not respond within the 30-day window', so a large share of those outcomes are claimant non-responses rather than determinations. Two independent trade readings of the same 2024 edition give 'over 65 per cent' and '70 per cent'; the discrepancy is recorded rather than averaged. YouTube's own tier comparison is why the dispute rate cannot be read as accuracy: counter-notifications ran at over 5 per cent of removals through the open public webform in the first half of 2021 and over 4 per cent in the second half of 2022, against fewer than 2 per cent in the limited-access tools and under 1 per cent against Content ID claims — pushback rising as access broadens, the inverse of what an error signal would do. YouTube also reports repeatedly that manual claims are more than twice as likely to be disputed as automated ones: under 0.6 per cent against over 1 per cent in the first half of 2021, under 0.5 against over 0.9 in the second half of 2022, and 0.54 against 1.13 in 2024. Paul Keller of the Communia Association, writing for infojustice in December 2021, derived a FLOOR from the first edition's published numbers — 729.3 million copyright actions in six months, 3.7 million disputes, roughly 60 per cent resolved for the uploader, therefore at least 2.2 million confirmed unjustified actions in half a year — and argued the true figure is necessarily higher because most affected uploaders never complain, concluding that 'over-enforcement (both unjustified blocking and unjustified demonetisation) is a very real issue that affects the rights of a substantial number of uploaders on a regular basis'. That is his derivation from the deployer's own numbers, not a finding by anyone. YouTube publishes a caveat that cuts the other way too: dispute and counter-notification counts are trailing events that keep accruing after a period closes, so it snapshots them three months after period end and any rate read from a freshly closed period is an undercount by construction. No figure exists anywhere for claims that were wrong and were never disputed.

Sources: youtubegoogle2021, torrentfreak2025, keller2021, googletransparencyreport2022

Appears on: /domains/cases/youtube-content-id

EmpiricalMoney rather than removal is the dominant outcome of a Content ID claim, which is why its error surface is nearly invisi…

Money rather than removal is the dominant outcome of a Content ID claim, which is why its error surface is nearly invisible. Over 90 per cent of Content ID claims are monetized rather than blocked, on YouTube's own reporting for the second half of 2022 and for calendar 2024: the claimed video stays up and the advertising revenue goes to the claimant, so an over-broad claim usually produces no takedown, no strike and no visible trace, only a diverted revenue stream that the uploader must notice and contest to reverse. The revenue clock is published and it turns on speed rather than correctness. Revenue on a claimed video is held while the claimant reviews a dispute, but held from the CLAIM date only if the uploader disputes within five days of the claim; if the uploader disputes later, the hold runs only from the dispute date; and if the uploader takes no action within those five days, the revenue accrued in that window is paid to the CLAIMANT regardless of how the dispute is later resolved. Revenue data is also suppressed in the uploader's analytics while a claim is active. Cumulative Content ID payouts to rights-holders reached 5.5 billion United States dollars from advertising as of December 2020, 9 billion as of December 2022, and over 12 billion as of December 2024, of which approximately 3 billion in 2024 alone. The Electronic Frontier Foundation's 2020 study argues that this pricing produces pre-emptive self-censorship rather than contested claims: because Content ID cannot assess fair use and each rung of the ladder risks deplatforming or lost income, creators cut clips to a few seconds, re-edit videos as the matcher changes, and surrender revenue on uses copyright law would permit, being in that study's words 'so afraid of being deplatformed or losing that income' that the loop goes unused. That is the study's analysis, attributed to it.

Sources: youtubegoogle2021, youtubehelp2026, torrentfreak2025, trendacosta2020

Appears on: /domains/cases/youtube-content-id

EmpiricalThe Content ID objection ladder is documented, asymmetrically priced at every rung, and its deterrent effect is measurab…

The Content ID objection ladder is documented, asymmetrically priced at every rung, and its deterrent effect is measurable in the deployer's own integers. An uploader may dispute a claim; the claimant has 30 days to respond and non-response releases the claim automatically. If the claimant reinstates, the uploader may appeal; the claimant then has 7 days, cut from 30 in September 2022 when an 'Escalate to Appeal' route was introduced. At that point the claimant may no longer reinstate and must either release the claim or file a legal removal request. A Content ID claim by itself carries no copyright strike; a legal removal request does. Three strikes in 90 days terminates the account and all associated channels, and strikes expire after 90 days if the uploader completes YouTube's Copyright School while holding fewer than three. After a counter-notification the claimant has 10 business days to show it has initiated court action or the content is reinstated — that window is statutory rather than YouTube's. THE MEASURED DETERRENCE: of 45,724 failed appeals in the second half of 2022, 13,841 (just over 30 per cent) resulted in a copyright removal, and the remaining 31,883 ended because the uploader cancelled the appeal or deleted the video rather than accept the strike risk. Roughly seven in ten uploaders who had already lost twice abandoned the matter rather than proceed. THE FUNNEL'S TAIL: in the same half-year YouTube accepted fewer than 25 per cent of the counter-notifications submitted to it and fewer than 1 per cent of counter-notifications resulted in a lawsuit, against 826,242,639 claims — six orders of magnitude of attrition from claim to court. Perel and Elkin-Koren documented in 2016 that appeal eligibility historically depended on the account being in 'good standing', so a prior strike could remove the ability to appeal the next claim; current documentation places Content ID appeal behind advanced-feature verification. This is the COPYRIGHT strike ledger, a different rule and a different count from the community-guidelines strike the same platform applies to its own policy enforcement.

Sources: youtubegoogle2021, youtubehelp2026, perel2016

Appears on: /domains/cases/youtube-content-id

EmpiricalAccess to Content ID is rationed, and YouTube's own tier comparison shows the ration suppressing abuse and concentrating…

Access to Content ID is rationed, and YouTube's own tier comparison shows the ration suppressing abuse and concentrating enforcement authority at the same time. In calendar 2025, 7,626 entities held Content ID access and 4,454 actively used it; in calendar 2024 the figures were 7,703 and 4,564; earlier editions give 'over 9,000 partners' as of December 2022 and the U.S. Copyright Office recorded over 9,000 rights-holders as of 2020. Against that, 295,531 claimants used the public copyright webform in 2025 and 173,338 used the Copyright Match Tool, with over 4 million channels holding Copyright Match Tool access as of December 2025 — up from over 2 million in July 2021 and over 2.5 million in December 2022. So roughly seven and a half thousand entities generate 99.48 per cent of all copyright actions on the platform. The stated criterion is exclusive rights to 'a substantial body of original material that is frequently uploaded by the YouTube creator community', plus demonstrated need and capacity, with categories excluded by rule: mashups, compilations and remixes; video game footage and software visuals; unlicensed media; licensed content without exclusive rights; and recordings of performances, concerts, events and speeches. A refused applicant may respond once with additional information. THE GATE'S MEASURED EFFECT, from the deployer's own reporting: videos requested for removal through the open public webform that YouTube's review team deemed 'a likely false assertion of copyright ownership' ran at over 8 per cent in the first half of 2021, over 5 per cent in the second half of 2022 and over 6 per cent in 2025, against 0.2 per cent or lower in the limited-access tools in the first half of 2021 and 0.5 per cent or lower in the second half of 2022; the 2025 edition describes the webform abuse rate as more than ten times that of all other copyright removal tools. YouTube states it terminates 'tens of thousands of accounts each year that attempt to abuse our copyright tools' and that claimants who repeatedly make erroneous Content ID claims can have Content ID access disabled and their partnership terminated — and publishes no count of partners actually de-accessed for erroneous claiming, the one funnel number absent from every edition. No headcount or review capacity is published for any of the copyright teams; the figure YouTube gives is 'hundreds of millions of dollars' invested in the Copyright Management Suite, which is investment rather than capacity.

Sources: torrentfreak2025, youtubegoogle2021, youtubehelp2026, u2020, keller2021

Appears on: /domains/cases/youtube-content-id

EmpiricalThe reference store, not the comparison, is where Content ID's errors scale, and YouTube publishes both the mechanism an…

The reference store, not the comparison, is where Content ID's errors scale, and YouTube publishes both the mechanism and a worked example. Its own words: 'Just one bad copyright webform notice can result in a handful of videos being temporarily removed from YouTube. In Content ID the impact is multiplied due to its automated nature; one bad reference file can impact hundreds or even thousands of videos across the site.' The example the report gives is its own — a news channel uploaded public-domain NASA Mars-rover footage as a reference file and made claims against every other channel using the same footage, including NASA's own channel. There is a dedicated correction loop over the store, upstream of and independent from the per-claim dispute loop: a dedicated team plus automated systems detect bad or low-quality reference files; the partner may exclude the offending segment, remove the whole reference file, or ask for re-review; and if the partner does not respond the reference file is marked invalid and removed and ALL claims associated with it are released at once. YouTube names the recurring causes: partners delivering non-exclusive content, public-domain material, licensed-but-not-owned clips, or reference files capturing indistinct sound effects and nature sounds. Content ID also holds a queue of PENDING claims where reference files carry flawed or conflicting ownership data and the system is uncertain whether a claim should be made at all, with the conflict resolved between partners rather than against the uploader — an explicit abstain-and-hold path inside an otherwise fully automated channel. A DOCUMENTED FALSE-CLAIM CASE ON NON-COPYRIGHTABLE AUDIO: a ten-hour white-noise recording uploaded in 2015 had drawn five separate Content ID claims by January 2018, at least two of them matching other white-noise recordings held by a single company; all five claimants chose to MONETIZE rather than block, so the error's only visible effect was a diversion of advertising revenue, and the claims were released after press attention rather than through the dispute process. Nothing about a released claim propagates back into the reference file unless the integrity team independently flags it, so a reference file that should not have been admitted keeps generating claims against every future matching upload.

Sources: youtubegoogle2021, torrentfreak2025, electronicfrontierfoundation2018

Appears on: /domains/cases/youtube-content-id

EmpiricalClaimant-side fraud through Content ID is an adjudicated criminal fact, and it exposes the delegability of the access ga…

Claimant-side fraud through Content ID is an adjudicated criminal fact, and it exposes the delegability of the access gate rather than any inaccuracy in the matching. Two principals of MediaMuv L.L.C. were indicted on thirty counts in the District of Arizona on 16 November 2021 for conspiracy, wire fraud, money laundering and aggravated identity theft arising from false Content ID ownership claims; both pleaded guilty, one in April 2022 and the other in February 2023, and one was sentenced in June 2023 to 70 months in prison. On the charging record they falsely claimed ownership of over 50,000 Latin music recordings and monetized them through Content ID via a third-party rights administrator, obtaining $20,776,517.31 by the indictment's count and approximately $23.4 million by the plea. The indictment describes the method: staff found unmonetized music on the platform, downloaded and re-uploaded it, and asserted ownership through the content management system, presenting the administrator with a contract stating they were the 'writer, author, publisher, copyright holder and creator' of the catalogue, backed by forged letters from artists — assertions the administrator accepted without ownership verification. Individual artist losses recorded on the charging record run to $132,702, $128,339 and $102,626, and the scheme ran roughly four years before it was stopped. THE STRUCTURAL POINT: an approved Content ID partner can present claims for a catalogue the platform never assessed, so the eligibility gate is delegable and the vetting failure in this case was at the intermediary as much as at the platform. THE BOUNDARY: this is fraud against RIGHTS-HOLDERS committed through the claiming tools, not over-claiming against uploaders, and it is not evidence about the comparison's accuracy. The individual defendants are not named here and the third-party rights administrator is identified in the charging documents by initials only and is not named or guessed.

Sources: usdepartmentofjustice2021

Appears on: /domains/cases/youtube-content-id

EmpiricalNo court and no regulator has found Content ID unlawful, ordered it changed or sanctioned it anywhere as of 28 August 20…

No court and no regulator has found Content ID unlawful, ordered it changed or sanctioned it anywhere as of 28 August 2026, and the one United States case that attacked its access structure produced no finding of any kind. Schneider et al. v. YouTube, LLC (N.D. Cal. 3:20-cv-04423) alleged that Content ID was reserved for powerful copyright owners and unavailable to ordinary creators. Class certification was DENIED on 22 May 2023 on the ground that classwide copyright ownership 'will entail individualized proof that precludes certification', with the court adding that 'the takedown of content in response to a DMCA notice is miles away from substantive proof of copyright ownership or infringement'. On 12 June 2023 — the day trial was scheduled to begin, and after YouTube's 25 May withdrawal of its safe-harbour defence — the parties stipulated to dismissal WITH PREJUDICE of all claims raised or that could have been raised. There was no trial and no verdict. The eligibility complaint is nonetheless on the federal record: the U.S. Copyright Office's 2020 Section 512 Report describes Content ID as a voluntary filtering system beyond section 512's requirements, quotes the participation criterion, and reproduces commenters objecting that it 'unfairly excludes smaller copyright owners', that 'every artist should be entitled to this service', and — from the party who would later sue — 'basically, that means the little guy need not apply. That's wrong.' The same report records the OPPOSITE complaint from rights-holders that Content ID misses a significant share of unauthorized uploads, one commenter reporting a contractor identifying 1,488,035 infringing copies since December 2012 that Content ID had not caught, and user-advocacy comments that the system is 'prone to false positives and cannot properly take fair use considerations into account'. Both error directions sit on the same federal record from opposing parties. Separately, in YouTube, LLC v. Christopher L. Brady (D. Neb. 8:19-cv-00353, filed 19 August 2019) YouTube itself brought an action under 17 U.S.C. 512(f), alleging the defendant sent dozens of false takedown notices and threatened to trigger a third strike, terminating channels, unless creators paid him; the case settled in October 2019 without adjudication. Those are allegations attributed to YouTube as the pleading party, and they concern the public webform and strike channel rather than Content ID, though they terminate in the same strike ledger the Content ID appeal ladder feeds. THE STRUCTURAL CHARACTERISATION, attributed: Maayan Perel and Niva Elkin-Koren wrote in the Stanford Technology Law Review in 2016 that Content ID welds ex ante algorithmic blocking onto DMCA-style ex post removal and 'has turned algorithmic copyright enforcement into a private-financial model' protecting owners 'beyond the basic removal process provided by the DMCA', proposing transparency, due process and public oversight as the accountability frame.

Sources: goldman2023, u2020, complaintforviolationofthedi2019, perel2016, youtubegoogle2021

Appears on: /domains/cases/youtube-content-id

EmpiricalIn Consent Order 2023-CFPB-0013, issued 8 November 2023 against Citibank, N.A., the Consumer Financial Protection Bureau…

In Consent Order 2023-CFPB-0013, issued 8 November 2023 against Citibank, N.A., the Consumer Financial Protection Bureau found that from at least 1 January 2015 through 31 December 2021 employees performing Citi Retail Services 'Judgmental Review' — defined in the order as the process by which a Citi employee or agent manually underwrites a credit-card application and approves, denies, or otherwise makes a credit decision — used an applicant's last name ending in -ian or -yan, especially with an address in or around Glendale, California, to identify applications they associated with Armenian national origin and to treat those applicants as presenting high fraud risk. The flag drove five documented actions: denial or approval on less favourable terms; additional scrutiny including verification of income or assets; a block or hold on the account; and referral to Citi's fraud prevention units for potential account freeze, credit-line decrease or account closure. The order finds that Citi took corrective action against employees who FAILED to identify and deny such applications, including action that could affect an agent's performance rating, pay and authority to approve future applications. The order names no model, score, algorithm, rule engine or screening vendor; the Bureau's FY2023 Fair Lending Report records that Glendale is home to approximately 15 percent of the Armenian American population of the United States. Citi executed the stipulation on 3 November 2023 without admitting or denying any finding of fact or conclusion of law.

Sources: consentorder2023, consumerfinancialprotectionb2024c

Appears on: /domains/cases/citi-armenian-surname-flags

EmpiricalThe order's second count is that the explanation channel itself was falsified: a failure to provide an accurate and adeq…

The order's second count is that the explanation channel itself was falsified: a failure to provide an accurate and adequate statement of the specific reasons for adverse action when an applicant was denied on a prohibited basis, under 15 U.S.C. 1691(d) and 12 C.F.R. 1002.9(a)-(b). Paragraph 25 documents one instance at the keystroke level — in 2016 an underwriter with approve and deny authority messaged a colleague asking for reasons to use for a decline, the colleague supplied several apparently pretextual reasons, and one second later the first employee recorded the application as declined due to possible credit abuse. The order states that Citi identified no legitimate, non-discriminatory explanation for the disparities and that any reasons it identified were pretextual justifications. The doctrinal frame is the Bureau's own: Circular 2022-03 states that the specific-reason requirement admits no exception for opaque or complex decision processes, and Circular 2023-03 states that a creditor may not rely on checklist sample-form reasons that do not reflect the actual basis for the decision. This record instantiates the second proposition in its deliberate form — the decision-maker knew the actual reason and entered a different one — so the applicant's only route to the real basis was closed by the notice itself.

Sources: consentorder2023, consumerfinancialprotectionb2022c, consumerfinancialprotectionb2022b

Appears on: /domains/cases/citi-armenian-surname-flags

EmpiricalThe order finds that concealment was taught rather than improvised: Citi supervisors and trainers instructed employees t…

The order finds that concealment was taught rather than improvised: Citi supervisors and trainers instructed employees to conceal their reliance on surname and address in the credit decision, including by telling employees not to discuss it in writing or on recorded phone lines, and the Bureau's FY2023 Fair Lending Report describes supervisors who conspired to hide the discrimination and employees who lied about the bases of denial by providing false reasons to denied applicants. An internal escalation reached management once and did not stop the practice: in 2018 a Citi employee emailed a group manager of Citi Retail Services and others asking how to document adverse-action reasons, writing that customers could not be told they were being declined because they are in Glendale, and the order records that the practice persisted after that concern was raised. What did reach the practice was external and statistical — regression analyses of Citi Retail Services credit-card data from 2015 through 2021, restricted to applications referred for Judgmental Review, showing denial-rate disparities for -ian and -yan surnames, larger when combined with a Glendale-area address, that were statistically significant.

Sources: consentorder2023, consumerfinancialprotectionb2024c

Appears on: /domains/cases/citi-armenian-surname-flags

EmpiricalThe order imposed $24,500,000 in civil money penalty to the Bureau's Civil Penalty Fund and $1,400,000 reserved for cons…

The order imposed $24,500,000 in civil money penalty to the Bureau's Civil Penalty Fund and $1,400,000 reserved for consumer redress, both payable within ten days, and built a five-year governance regime around the surface that failed: monitoring of the written AND oral communications and training materials of Judgmental Review personnel; portfolio-wide statistical analysis of judgmental credit decisions at least annually; a 60-day root-cause-and-corrective-action window on any flag; at-least-quarterly reporting of every instance with its affected population, root cause and corrective action; at least annual ECOA and Regulation B training for all Covered Personnel and for affiliate credit decision-makers, with training within 30 days for newly assigned personnel; and a corrective menu that names correcting inaccurate or inadequate adverse action notices and correcting inaccurate consumer reporting alongside remunerating consumers and extending credit previously denied, with the Board holding ultimate responsibility and the Chief Executive Officer and Board reviewing every submission. On 16 October 2025, roughly 23 months into that five-year term, the Bureau terminated the order under 12 U.S.C. 5563(b)(3), stating that Citibank had fulfilled certain obligations and that the Bureau also waives any alleged noncompliance therewith — a waiver of alleged noncompliance rather than a finding of compliance. A bicameral congressional letter of 23 April 2026 demanded the record by 7 May 2026; the reply reported on 14 May 2026 gave the first public redress accounting, $1,370,207.16 paid to 573 individuals with 126 checks uncashed and redistributed among the 447 who cashed theirs, characterized the conduct as rogue conduct by a few underwriters at one location, and asserted that more than 100 recipients were not Armenian, naming surnames such as Christian and Bryan. That characterization is the Acting Director's, is in tension with paragraph 21 of the 2023 order, and comes from an official defending his own termination decision; 573 is the identified-and-payable cohort rather than an estimate of how many people were affected.

Sources: consentorder2023, orderterminatingtheconsentor2023, enforcementactionpagecitiban2023, officeofu2026a, bankingdive2026, yahoofinance2026

Appears on: /domains/cases/citi-armenian-surname-flags

EmpiricalThree positions on this conduct sit in the public record and none is treated as settled. The Bureau's 2023 findings, whi…

Three positions on this conduct sit in the public record and none is treated as settled. The Bureau's 2023 findings, which Citi neither admitted nor denied, rest on portfolio-wide statistically significant disparities and locate the concealment instruction with supervisors and trainers; then-Director Rohit Chopra said at announcement that Citi stereotyped Armenians as prone to crime and fraud and illegally fabricated documents to cover up its discrimination. Citi's own position is that in trying to thwart what it calls a well-documented Armenian fraud ring operating in certain parts of California a few employees took impermissible actions, that basing credit decisions on national origin is unacceptable, and that it apologises to any applicant evaluated unfairly by the small number of employees who circumvented its fraud detection protocols. Private claims followed immediately and were largely diverted out of court: Marine Grigorian v. Citibank, N.A., No. 2:23-cv-09519 (C.D. Cal., filed 10 November 2023, Judge Michael W. Fitzgerald) and Smbatian et al. v. Citibank, N.A. et al., No. 2:23-cv-09811 (C.D. Cal., filed 17 November 2023) were among several proposed class actions alongside state-court mass-tort filings reported to involve hundreds of customers, and in April 2024 the court compelled arbitration in Grigorian, rejecting the argument that the card agreement's arbitration clause was unenforceable under California's McGill rule on public injunctive relief. All allegations in those cases are allegations and not findings, including the plaintiff-side claim that applications with apparently Armenian surnames were routed to a special manual-review unit, which does not appear in the consent order. No public resolution of the arbitration, of Smbatian, or of the state-court filings was located. Separately, on 11 March 2025 the Los Angeles Civil Rights Department, working with the city attorney's office and the state attorney general, opened a call for complaints about anti-Armenian banking discrimination across major national banks; it is not a proceeding against Citi and has published no findings.

Sources: consentorder2023, thearmenianweekly2023, ababankingjournal2025a, dailyjournal2024, classaction2023, losangelestimes2025

Appears on: /domains/cases/citi-armenian-surname-flags

EmpiricalCredit Acceptance Corporation is a publicly traded indirect subprime vehicle lender: it does not lend across its own cou…

Credit Acceptance Corporation is a publicly traded indirect subprime vehicle lender: it does not lend across its own counter but buys retail instalment contracts from dealerships enrolled in its program, which pay a monthly fee for access to its Credit Approval Processing System and its servicing. The pleaded scale is that approximately 1.9 million consumers obtained loans through the operator and its affiliated dealers between 2 November 2015 and 30 April 2021, that the network exceeded 12,000 affiliated dealerships, and that consumers obtained more than $4.9 billion in operator-financed loans in 2020 alone, with New York among its top five state markets. The operator's own quarterly report for the period ended 30 June 2026 gives the current shape: 11,004 active dealers in the quarter, a record and up 3.3 per cent year over year, with 1,456 new dealers enrolled; 84,615 consumer-loan unit assignments worth $1.0 billion in the quarter; an $8.0 billion average loan portfolio balance; $90.6 million of dealer holdback paid in the first half of 2026; and a $599 monthly per-dealer program fee. The pleaded borrower population had a median credit-bureau score of 546 and a gross annual income of approximately $35,000.

Sources: complaint2023, creditacceptancecorporation2026

Appears on: /domains/cases/credit-acceptance-predicted-default

EmpiricalThe mechanics of the scored artifact are common ground between the parties. When a dealership submits an application thr…

The mechanics of the scored artifact are common ground between the parties. When a dealership submits an application through the operator's origination software, the operator scores the proposed transaction — the applicant's credit-bureau attributes, the application data, the deal structure of term, monthly payment, down payment and trade-in, and the vehicle — and returns a Score from 0 to 100 representing its estimate of the percentage of total amounts owed it expects to collect over the life of the contract, inclusive of scheduled payments, late fees, post-repossession auction proceeds, post-default collection recoveries, deficiency judgments and wage garnishment. The Score prices the payment the operator makes to the dealership; it does not decide whether financing is offered; it does not set the borrower's interest rate; and it is never disclosed to the borrower. The defendant's own revised dismissal memorandum states that the Score 'does not affect the terms of a Contract' and describes it as 'an internal scoring metric that Credit Acceptance uses to calculate the CAC Payment to the Dealer'. The plaintiffs allege, and it is an allegation, that the pleaded portfolio-level forecast averaged 64 nationwide and 66 in New York — an expectation of collecting roughly 64 to 66 cents per dollar of total amounts owed, interest included — and that for more than 39 per cent of loans nationwide and about 25 per cent of New York loans the projection fell below the amount financed, which is the stated principal rather than the total of payments. They further allege that for New York loans originated between 2015 and 2020 whose projected collections were below the amount financed, nearly 70 per cent were sixty or more days past due, repossessed, or sold at auction. The operator denied the complaint in full and moved to dismiss it in its entirety, and no court ruled on any of it.

Sources: complaint2023, memorandumoflawinsupportofcr2024

Appears on: /domains/cases/credit-acceptance-predicted-default

EmpiricalThe plaintiffs alleged that the pricing runs to the dealership rather than to the borrower. They pleaded that the operat…

The plaintiffs alleged that the pricing runs to the dealership rather than to the borrower. They pleaded that the operator offered a dealership on average about 72 per cent of the Score-derived projected net collections, that the nationwide average dealer payment was about 22 per cent less than the amount financed, and that recovering roughly 78 per cent of the amount financed therefore sufficed to exceed the cash the operator had actually put at risk — a gap they put at nearly $2,500 per New York loan. They alleged that the dealership's back-end share was rarely paid, with under 12 per cent of new loans nationwide in dealer pools receiving any earnout and total earnout under 2 per cent of the value received by New York dealers. They alleged that because the interest rate does not vary with borrower risk — New York contracts disclosed 22.99 or 23.99 per cent against a 25 per cent state criminal usury cap, and nationwide disclosed rates averaged about 22 per cent — the risk premium was relocated into the amount financed, where it is invisible to comparison shopping and accrues interest. They alleged that the operator's stated origination guardrail was proof of income together with a monthly payment not exceeding 25 per cent of gross monthly income, 30 per cent with manager approval, with no collection of recurring debt obligations, housing cost, food, healthcare or childcare cost, no debt-to-income ratio, no residual-income calculation and no adjustment for the number of dependents; that its stated pricing bound permitted a vehicle to be priced at up to 115 per cent of the highest published book value for that make and model without inspecting the vehicle's condition, against a measured median disclosed selling price about 77 per cent over wholesale book value and slightly more than 50 per cent over reported dealer cost; that 90 per cent of loans carried an operator-approved add-on product, at an average vehicle service contract retail cost of $1,545 and an average guaranteed-asset-protection cost of about $782, generating approximately $250 million of add-on revenue nationwide in 2020 and a flat dealership commission of about $385 per vehicle service contract; that the origination software displayed to the dealership, in real time, the projected dealership profit for each vehicle for that applicant and how it changed with each product added, in one pleaded screen moving from $1,605.31 to $4,051.82; and that the operator's internal dealer rating fed back into the size of the dealer advance. The plaintiffs' recomputation of the cost of credit — treating the difference between the disclosed deal cost excluding interest and the dealership's compensation as a concealed finance charge, and on that basis alleging that nearly 90 per cent of New York loans exceeded the 25 per cent cap with a median recomputed rate of about 34 per cent — is their construction: the defendant argued the cash-price proxy behind it was 'invented ... for this litigation' and incompatible with the Truth in Lending Act's defined terms, and the court never resolved it.

Sources: complaint2023, memorandumoflawinsupportofcr2024, officeofthenewyorkstateattor2023

Appears on: /domains/cases/credit-acceptance-predicted-default

EmpiricalThe information asymmetry is the pleaded governed surface, and it is manufactured inside one system. The plaintiffs alle…

The information asymmetry is the pleaded governed surface, and it is manufactured inside one system. The plaintiffs alleged that consumers 'do not know about, and certainly do not have access to, the extensive predictive data' the operator was using, and pleaded in three state-law counts, as a failure to disclose, that the operator did not tell the borrower that its algorithms and the Score had forecast projected collections far below the total amounts owed under the loan agreement, and that the projection potentially included amounts collected through late fees, repossession and wage garnishment. They alleged that the same software instance that shows the dealership a live per-vehicle profit projection then generates the borrower's contract and disclosures. The defendant answered that the plaintiffs had abandoned any standalone claim over non-disclosure of the Score, and that as an indirect lender it has no contact with the consumer until after the dealership's contract is executed and assigned to it. The disclosure regime either side points to — the federal Truth in Lending Act and New York's motor vehicle retail instalment regime — governs the contract form; no disclosure rule in this record requires telling a borrower what a lender's model predicts about them. There is no adverse-action notice, no explanation, no contest path and no borrower-facing channel of any kind attaching to the forecast anywhere in the record. None of this was adjudicated.

Sources: complaint2023, memorandumoflawinsupportofcr2024

Appears on: /domains/cases/credit-acceptance-predicted-default

EmpiricalThe plaintiffs pleaded a quantified recovery apparatus rather than a general complaint about collections. They alleged t…

The plaintiffs pleaded a quantified recovery apparatus rather than a general complaint about collections. They alleged that within days of a missed payment the operator would, as a matter of policy, disable the vehicle through a GPS starter-interrupt device, a practice they state was used on New York-financed vehicles until late 2018; that contract terms ran 60 to 72 months while repossessed-and-resold vehicles averaged under two years from origination to auction with a quarter auctioned within one year; that a majority of the roughly 1.9 million contracts became delinquent at some point and more than half of borrowers were delinquent within the first year; that the operator repossessed more than a quarter of financed vehicles nationwide and resold about 20 per cent at auction, with approximately 44 per cent repossessed in New York and 21 per cent of repossessed New York vehicles repossessed more than once; that auction proceeds satisfied on average only 29 per cent of remaining amounts owed nationwide and under 28 per cent in New York, leaving an average post-auction deficiency of about $8,500; that the operator made more than 138,000 referrals to debt-collection attorneys nationwide and obtained judgments against more than one in six New York borrowers whose loans reached maturity by May 2021, with New York judgment counts running between roughly 2,500 and 4,200 a year from 2013 to 2019 and more than 7,000 default judgments obtained in 2017 and 2018 in a window when 40 of the thousands of New York borrowers sued had counsel; and that, excluding support personnel, the operator employed nearly twice as many people in servicing and collections as in origination. Separately from every allegation in the case, the operator's own quarterly report for the period ended 30 June 2026 discloses that an AI-enabled call-centre agent handled 67 per cent of inbound customer-service and account-solutions calls in June 2026, up from 27 per cent in March 2026, and that the capability is integrated into core servicing workflows; that is a 2026 disclosure about a channel postdating the pleaded conduct and is not part of it. The pleaded statistics were never adjudicated.

Sources: complaint2023, officeofthenewyorkstateattor2023, creditacceptancecorporation2026

Appears on: /domains/cases/credit-acceptance-predicted-default

EmpiricalThe operator publishes the forecast's own accuracy, which is unusual and is the most modelling-relevant artifact in this…

The operator publishes the forecast's own accuracy, which is unusual and is the most modelling-relevant artifact in this record. Its quarterly report for the period ended 30 June 2026 tabulates, for each assignment year, the current total-loan forecast collection percentage against the percentage forecast when the loans were assigned: 2017 at 64.8 against 64.0 initial; 2018 at 65.6 against 63.6; 2019 at 67.3 against 64.0; 2020 at 68.1 against 63.4; 2021 at 64.1 against 66.3; 2022 at 59.3 against 67.5; 2023 at 62.9 against 67.5; 2024 at 65.1 against 67.2; 2025 at 66.9 against 67.0; and 2026 at 67.1 against 67.2. The company states that it monitors credit quality monthly by comparing current forecast collection rates to initial expectations and periodically adjusts the statistical pricing model for trends identified through that review, and that 'since all known, significant credit quality indicators have already been factored into our forecasts and pricing, we are not able to use any specific credit quality indicators to predict or explain variances in actual performance from our initial expectations.' Accurate forecasting of loan performance is listed first among the company's three published critical success factors. This series is a calibration record of the operator's own recovery against its own expectation. It is not a deployment error rate, no source in this record publishes a per-decision defect rate, an override rate or a correction rate for this deployment, and no independent evaluation of it exists.

Sources: creditacceptancecorporation2026

Appears on: /domains/cases/credit-acceptance-predicted-default

EmpiricalThe matter was resolved by a settlement in principle without any adjudication. On 4 January 2023 the Consumer Financial …

The matter was resolved by a settlement in principle without any adjudication. On 4 January 2023 the Consumer Financial Protection Bureau and the People of the State of New York jointly filed a 59-page complaint in the Southern District of New York pleading seven causes of action, following state subpoenas running from 2016, a federal civil investigative demand of April 2019, a multistate investigation expanded in August 2020 to 41 further states plus the District of Columbia, notice-of-intent letters in November 2020 and August 2022 and a federal notice-and-opportunity-to-respond letter in December 2021. The case was stayed on 7 August 2023 pending the Supreme Court's decision in Consumer Financial Protection Bureau v. Community Financial Services Association of America and the stay was lifted on 1 July 2024, with a revised motion to dismiss filed 14 August 2024 and fully briefed by 29 October 2024. On 24 April 2025 the Bureau filed a consented motion under Federal Rule of Civil Procedure 21 to be dropped as a plaintiff and to withdraw its counsel's appearances; the defendant consented, New York did not object, and the court granted it on 29 April 2025, leaving New York as sole plaintiff. The Bureau's supporting memorandum offers no substantive or merits-based reason and makes no representation about the merits of the withdrawn claims; it is a litigation decision recorded on a docket and not an adjudication or a determination that the claims lacked merit. Contemporaneous trade reporting placed the withdrawal within a broader 2025 pattern of agency enforcement dismissals and carried the operator's statement, through its Chief Legal Officer, that the case 'never should have been brought'; that context is press attribution rather than an agency statement. The case was reassigned to Judge Jesse M. Furman on 28-29 January 2026, and on 6 February 2026 the court deferred ruling on the fully briefed motion to dismiss and terminated it to facilitate settlement, stating it would restore the motion if settlement failed. On 5 June 2026 the court, advised that all claims had been settled in principle, ordered the action dismissed and discontinued without costs and without prejudice to a right to reopen within sixty days, and noted that it would not retain jurisdiction to enforce a settlement agreement unless the agreement were made part of the public record. On 23 July 2026 the New York Attorney General's office reported that the parties had agreed on all settlement terms with documentation to be executed and obtained a forty-five day extension of the reopen deadline. The operator's quarterly report of 4 August 2026 states that it offered $45.0 million in September 2025 to settle this action jointly with the parallel multistate investigation, that in January 2026 it reached preliminary alignment on material terms including a potential cash payment of $75.5 million covering both matters, and that until the matter is fully and finally resolved it intends to defend itself. As of the docket's last update no executed settlement, consent judgment or public statement of terms appears. There are no findings of fact and no adjudicated liability.

Sources: docket, memoranduminsupportofconsent2025, orderdismissinganddiscontinu2026, creditacceptancecorporation2026, civilrightslitigationclearin, americanbanker2025

Appears on: /domains/cases/credit-acceptance-predicted-default

EmpiricalOne version of these theories against this operator did produce a governed remedy. On 1 September 2021 the Massachusetts…

One version of these theories against this operator did produce a governed remedy. On 1 September 2021 the Massachusetts Attorney General announced a $27.2 million settlement in Suffolk Superior Court, described by that office as the largest of its kind, resolving allegations that the operator made high-interest subprime loans 'it knew or should have known many borrowers would be unable to repay', imposed hidden finance charges violating the state's 21 per cent usury cap, engaged in unlawful collection practices, and failed to tell investors that higher-risk loans were placed into securitisation pools. Relief covered more than 3,000 Massachusetts borrowers and comprised cash, debt forgiveness, credit-bureau deletion of related negative marks, and required changes to loan-handling practices; the required practice changes were not enumerated publicly. The operator did not admit wrongdoing, stated that the suit had been 'vigorously contested' and said it looked forward to continuing to serve customers in the Commonwealth through its financing programs. The settlement resolved allegations and is not a finding that the conduct occurred; it is recorded here as evidence that these theories are not novel and that one jurisdiction's version of them landed a remedy.

Sources: officeofthemassachusettsatto2021, americanbanker2021

Appears on: /domains/cases/credit-acceptance-predicted-default

EmpiricalIn United States v. Dave, Inc. and Jason Wilk, No. 2:24-cv-09566-MRA-AGR (C.D. Cal.), the operative First Amended Compla…

In United States v. Dave, Inc. and Jason Wilk, No. 2:24-cv-09566-MRA-AGR (C.D. Cal.), the operative First Amended Complaint filed 30 December 2024 alleges that in the first fourteen months after Dave began advertising cash advances of 'up to $500', it offered a $500 advance to a new user about 0.002 per cent of the time — fewer than one determination in forty-five thousand; that only about 0.13 per cent of new users were offered even half of the advertised amount; that the most common offer, when an offer was made, was $25; that more than three-quarters of the time no advance was offered at all; and that on average more than 40 per cent of new users obtained no offer in a calendar month. Of new users who did receive offers, about 0.009 per cent of offers were for $500 and about 0.56 per cent were for at least $250; for existing users over the same window, on average more than a third were offered no advance in a calendar month and a $500 advance was offered less than 1 per cent of the time. THE DEFENDANTS DENY PARAGRAPHS 34, 35 AND 36 IN THEIR ENTIRETY and state in their dismissal brief that 'Dave can, does and did provide $500 advances'. The pleading never dates the fourteen-month window. Independently of the disputed tail, and undisputed because the operator publishes it, the advertised ceiling is $500 while Dave's own SEC-filed average advance was $170 in fiscal 2024 and $205 in fiscal 2025.

Sources: firstamendedcomplaintforperm2024, complaintforpermanentinjunct2024, answertoamendedcomplaintforp2025, memorandumofpointsandauthori2025, daveinc2026a

Appears on: /domains/cases/dave-extracash-advance

EmpiricalThe enforcement record says nothing whatever about the decision system. The words algorithm, artificial intelligence, ma…

The enforcement record says nothing whatever about the decision system. The words algorithm, artificial intelligence, machine learning, model and underwriting appear nowhere in the original complaint, the operative amended complaint, the defendants' motion-to-dismiss memorandum, the government's opposition, the defendants' answer, or the court's 34-page order; the pleading describes only that Dave 'uses its access to consumers' bank accounts to analyze their finances and banking history' and 'uses this information to make decisions about how much (if any) to advance the consumer'. Every statement about the engine therefore comes from the operator's own investor-facing disclosures. Dave's Form 10-K for fiscal 2025 states that it uses 'our proprietary AI-powered underwriting system, CashAI' to 'analyze a Member's checking account transaction data to determine eligibility and set the bank's credit approval amount', in a 'fully automated process' that 'requires no credit check and does not rely on FICO or credit bureau data', drawing on 'hundreds of data points — including income patterns, spending behavior, and transaction history', and that the system has 'leveraged insights from over 180 million ExtraCash originations and billions of bank transactions'. The same filing discloses three further automated components in the same deployment: a real-time behavioural fraud-mitigation system with user-level controls, an income-and-expense prediction component feeding both underwriting decisions and the member-facing budgeting feature, and an automated support chatbot. In September 2025 the company announced CashAI v5.5, described as trained on more than 7 million recent originations that had reached full maturity, nearly doubling the prior feature set and optimized for the current fee structure, with claimed improvements in risk ranking, approval amounts, conversion, delinquency and loss. ALL OF THAT IS THE OPERATOR'S OWN, UNAUDITED CLAIM. No regulator, court or auditor has examined, described or characterized the engine, and no model documentation, validation report, fairness assessment or independent evaluation of it exists in the public record.

Sources: daveinc2026a, daveinc2025, firstamendedcomplaintforperm2024, memorandumofpointsandauthori2025, civilminutesgeneral2025, answertoamendedcomplaintforp2025

Appears on: /domains/cases/dave-extracash-advance

EmpiricalThree of the five counts are Restore Online Shoppers' Confidence Act counts and they concern the order in which things h…

Three of the five counts are Restore Online Shoppers' Confidence Act counts and they concern the order in which things happen rather than the amount. The pleading alleges that a monthly membership fee was charged to every consumer who linked a bank account, whether or not any advance was ever offered, and that from at least August 2021 through November 2022 no in-app mechanism existed to stop that charge for consumers who also held a Dave bank account — which, from early 2022, Dave required of new consumers who wanted advances. It alleges at least nine separate in-app steps from the main screen to complete cancellation, diversion from cancellation for consumers who select the most prominent option, identity checks demanded to cancel including date of birth, sign-up phone number, mailing address, the last four digits of a Social Security number and details of the last two transactions on the external bank account, a July 2020 customer-service instruction that only consumers with no open advance and no pending advance payment were eligible to pause, and one consumer who required 27 days and nine messages to support and a threat to contact the Better Business Bureau (the court's order recites 29 days). A fourth count concerns the historic tip mechanic: a default charge of 15 per cent behind a large green 'Thank you!' button above imagery of a cartoon child and boxes reading '10 Healthy Meals', '15 Healthy Meals', '20 Healthy Meals', with the custom-tip alternative rendered white on white at about half the width and the child replaced by an empty plate at a zero tip; the pleading alleges Dave donated ten cents per percentage point of tip, usually $1.50 or less per advance, and kept the rest. Dave admits that tipping 'was formerly a revenue source', that members were 'presented with the option of providing an optional tip after the ExtraCash overdraft was sent', and that it 'donated a portion of each tip', and denies the remainder. The FTC's press release of 5 November 2024 states — citing Dave's own SEC filings and NOT the complaint — that Dave reported more than $149 million in revenue from these tips from 2022 through the first six months of 2024.

Sources: firstamendedcomplaintforperm2024, civilminutesgeneral2025, answertoamendedcomplaintforp2025, federaltradecommission2024

Appears on: /domains/cases/dave-extracash-advance

EmpiricalThe two loops in this deployment run at different orders of magnitude, and both speeds are documented. On the operator's…

The two loops in this deployment run at different orders of magnitude, and both speeds are documented. On the operator's side, Dave's Form 10-K for fiscal 2025 states that the approximately eleven-day average term of an advance 'creates rapid feedback loops, enabling iterative model refinement', and the company publicly versions the result. The pleading alleges a comparable apparatus on the interface: an experiment removing the 'Healthy Meals' content for some consumers, after which the percentage of new users charged a tip fell by about a third and overall tip revenue fell by almost a quarter, followed by an internal analysis recommending the content resume for all users; and a second experiment removing a three-box screen that likewise reduced both the number of consumers charged and the amounts. The pleading further alleges that dissatisfaction was measured with precision and answered without correction: an internal analysis of customer-service data naming 'Low advance amount', 'Low advance limits and approval' and 'Advance request denied' among the top drivers of contact; an internal survey naming 'Not enough money' a top source of dissatisfaction; thousands of monthly cancellation contacts of which 'most don't qualify for an advance or get a smaller than expected advance'; hundreds of monthly contacts on the topic 'What is the $1 charge?'; an internal analysis of Better Business Bureau complaints flagging 'inability to cancel easily within the app'; and an internal presentation stating that on the Express Fee screen 'what we promised is not what they see' and recommending Dave 'set expectations much earlier on the true cost of the money they are borrowing'. The pleading alleges each corresponding recommendation went unimplemented; the defendants refer the court to the documents in their entirety and deny mischaracterizations. On the other side, the consumer-protection loop ran on a civil investigative demand served in January 2023, suit in November 2024, referral in December 2024, a dismissal ruling in September 2025, contested discovery through mid-2026, and a final pretrial conference set for 9 November 2026, with no ruling on the merits at any point.

Sources: firstamendedcomplaintforperm2024, answertoamendedcomplaintforp2025, daveinc2026a, docket2024, federaltradecommission2024

Appears on: /domains/cases/dave-extracash-advance

EmpiricalThe disclosure moved while the litigation ran, and the before-and-after is directly observable without discovery. The am…

The disclosure moved while the litigation ran, and the before-and-after is directly observable without discovery. The amended complaint at paragraph 22 quotes Dave's website shortly after the November 2024 filing as carrying 'Get up to $500 in 5 minutes or less' with a fine-print footnote stating only that 'the average advance is $170' and that 'enrollment and initial qualification [are] typically completed in 5 minutes', and alleges that even that footnote failed to disclose that many consumers who give Dave bank-account access will be offered no advance at all. The same site as displayed on 28 August 2026 still leads with 'Up to $500 in 5 min or less', and its footnote now reads: 'ExtraCash amounts range from $25-$500, typically authorized within 5 minutes, with an overdraft fee equal to the greater of $5 or 5%. Multiple overdrafts may be required. Not all members qualify for ExtraCash and few qualify for $500.' The two concessions now present — that not all members qualify at all, and that few qualify for $500 — are precisely the two omissions the amended complaint pleads. The price surfaces moved too, while liability was denied: members onboarded from 4 December 2024 were placed on a structure without optional tips or express fees, and in February 2025 Dave completed a transition to a mandatory 5 per cent overdraft service fee with a $5 minimum, with tip revenue falling 89 per cent from $67.6 million in fiscal 2024 to $7.5 million in fiscal 2025 and subscription revenue rising 51 per cent to $37.2 million. ONE DETAIL IS DELIBERATELY LEFT UNRESOLVED: the fiscal 2025 annual report reports a $3 monthly membership fee for new members from mid-2025 while the live site in August 2026 states an 'Up to $5 monthly membership fee', one reporting period apart and one of them a ceiling rather than a rate, so no single current subscription price is asserted.

Sources: firstamendedcomplaintforperm2024, daveinc2026, daveinc2026a, bankingdive2025b

Appears on: /domains/cases/dave-extracash-advance

EmpiricalOn 12 September 2025 Judge Monica Ramirez Almadani denied the motion to dismiss in full in a 34-page order, holding that…

On 12 September 2025 Judge Monica Ramirez Almadani denied the motion to dismiss in full in a 34-page order, holding that 'numerous courts have found that "up to" representations can materially mislead reasonable consumers where the defendant does not or cannot provide the good or service as represented, especially when the representation references a particular, quantified amount', that 'the government has plausibly alleged that it was exceedingly rare for Dave to offer the maximum amount of the cash advance advertised or even amounts approaching the maximum', and that a fine-print 'Terms apply' disclaimer in two banner advertisements did not cure the net impression because 'a disclaimer does not automatically exonerate deceptive activities'. THAT IS A PLAUSIBILITY RULING ON THE PLEADINGS AND ESTABLISHES NOTHING FACTUAL, and the order's footnote 2 records the dispute it did not resolve: 'Defendants contend that the government's method of calculating cash advances is wrong and that the data it used is [in]complete. ... Such a factual dispute cannot be resolved on a Rule 12(b)(6) motion to dismiss.' Dave and Wilk answered on 10 October 2025 with a general denial and seven affirmative defences: lack of fair notice of the government's interpretation of ROSCA; that ROSCA is unconstitutionally vague as applied; standing and mootness; good faith, resting in part on the assertion that 'the Consumer Financial Protection Bureau opened and closed an investigation — and declined to recommend an enforcement action against Dave' (an assertion by the defendant, with no agency document confirming it located); offsets; the statute of limitations; and that the penalties sought are unconstitutionally excessive. Publicly the company called the amended complaint 'a continued example of government overreach' resting on 'numerous allegations that are based on various inaccuracies', said it believes it has 'always acted within the law', and pledged to 'vigorously defend itself'. A Civil Trial Order of 14 November 2025 set a final pretrial conference for 9 November 2026; a stipulated protective order was entered in January 2026; contested discovery ran from March through June 2026; the last docket activity as of 6 August 2026 is counsel withdrawals. THERE IS NO SETTLEMENT, NO CONSENT ORDER AND NO ADJUDICATION ON THE MERITS, and the operator's Form 10-Q filed 5 August 2026 states it is 'unable to reasonably predict the possible outcome' and records a $9.7 million aggregate legal-contingency accrual across its three pending consumer matters. The 2025 change of FTC leadership did not thin the case: both authorizing Commission votes were 4-1 with Commissioner Melissa Holyoak voting no, the United States became the real party in interest in December 2024, and the Department of Justice litigated the matter through 2025 and 2026. Two private actions run alongside and are likewise unadjudicated: a Military Lending Act and Truth in Lending Act putative class action naming Dave and Evolve Bank & Trust, in which dismissal and arbitration were both denied on 12 December 2025 and which is on appeal to the Ninth Circuit with district proceedings stayed, and a suit by the Mayor and City Council of Baltimore under a municipal consumer-protection ordinance.

Sources: civilminutesgeneral2025, answertoamendedcomplaintforp2025, docket2024, daveinc2026b, federaltradecommission2024, bankingdive2025b

Appears on: /domains/cases/dave-extracash-advance

EmpiricalIn two consent orders four years apart, the Consumer Financial Protection Bureau found that Enova International, Inc. an…

In two consent orders four years apart, the Consumer Financial Protection Bureau found that Enova International, Inc. and its subsidiaries branded CashNetUSA and NetCredit debited or attempted to debit consumers' bank accounts without authorisation, and attributed the conduct to defects in Enova's own servicing and payment-processing software. Consent Order 2019-BCFP-0003, issued 25 January 2019, carried a $3,200,000 civil money penalty, four permanent conduct prohibitions and a five-year term, and applied the statutory unfairness test under 12 U.S.C. 5531(c)(1) directly to a software defect: 'The injury was caused by an error in Enova's software. The cost of debugging software would not have been significant and the erroneous practice did not confer any benefit to consumers or competition.' Consent Order 2023-CFPB-0014, issued 15 November 2023, found that Enova had VIOLATED four named paragraphs of the 2019 order and, because an order prescribed by the Bureau is federal consumer financial law, that each violation was itself a violation of the Consumer Financial Protection Act. It carried a $15,000,000 civil money penalty, a seven-year ban on Covered Loans, a ban on using or selling the associated consumer information, five behavioural prohibitions and a seven-year term. The Bureau's own headline described Enova as a repeat offender. The 2023 order enumerates eleven separately-described defect classes, each with its date range, consumer count and remediation figure, and names their machine causes: 'a coding error with its internal systems', 'flaws in Respondent's system logic', 'Enova's code failed to register', 'Respondent's automatic payment generation system', 'Respondent's internal systems did not accurately record'. Neither order uses the words algorithm, model, score, machine learning or artificial intelligence. The Bureau's aggregate for the 2023 order is 'violations of federal consumer protection law that involved over 111,000 consumers' — a figure dominated by two non-debit cohorts, the 50,565 unique consumers who did not receive a proper authorisation copy and the 43,359 customers charged incorrect amounts due, with the consumers debited from unauthorised accounts by the discrete code defects numbering in the low thousands. Both orders were entered without Enova admitting or denying the findings.

Sources: consumerfinancialprotectionb2023b, consumerfinancialprotectionb2019b, americanbanker2023

Appears on: /domains/cases/enova-lending-algorithm-errors

EmpiricalThe 2019 order found that from 2010 onward Enova overwrote existing customers' bank-account records with account details…

The 2019 order found that from 2010 onward Enova overwrote existing customers' bank-account records with account details taken from applications it had purchased from third-party lead generators, and then debited the substituted accounts, affecting 5,520 consumers. Enova stopped overwriting records on newly purchased applications in June 2014, but 'After June 2014, Enova continued to debit or attempt to debit 265 consumers' bank accounts that had already been overwritten, at least 6,425 times,' continuing until December 2018 for any of those consumers who still held an outstanding line of credit — so a repair applied at the ingest point left the contaminated store producing unauthorised debits for four and a half more years. The 2023 order found the same pathway carrying 356 further consumers in July and August 2020 and over $79,000, and answered with a prohibition the 2019 order had not contained: no debiting an account using information received from a Lead Generator 'without directly obtaining the consumer's express informed consent'. Separately, the 2019 order requires an electronic fund transfer authorisation to be signed or similarly authenticated with a copy returned to the consumer identifying the specific account; the 2023 order found 57,310 instances affecting 50,565 unique CashNetUSA consumers where that copy was not provided or did not name the account, and found that failure to violate the earlier order. The 2023 order also bans Enova from using the associated consumer information to market any consumer financial product and from selling or transferring it. The lead generators are named in the record by role only, and no reliable identification of any of them exists in the public sources.

Sources: consumerfinancialprotectionb2019b, consumerfinancialprotectionb2023b

Appears on: /domains/cases/enova-lending-algorithm-errors

EmpiricalThe 2023 order describes a fully specified machine rule that revoked a promise Enova had already confirmed in writing: '…

The 2023 order describes a fully specified machine rule that revoked a promise Enova had already confirmed in writing: 'Respondent's internal systems would compare the current account balance on the day before the loan extension was to be funded to the account balance at the time of the extension approval. A mismatch between the two balances resulted in an automatic cancelation of the approved loan extension. Interim partial payments would create such a mismatch.' Notification went out only after that check ran, 'which occurred the day before the due date of the original payday loan', telling the consumer to apply for a new extension to avoid a full-balance debit — 'But that was often insufficient time for consumers to act.' Between 2011 and 2020 about 3,500 granted extensions were cancelled, over 2,500 consumers were debited the full loan balance instead of the extension fee, and over $1 million was recorded as remediated on Enova's own representation. The rule appeared in no consumer-facing document: 'The loan extension contract did not inform consumers that Respondent would cancel the loan extension if the consumer made any partial payment on their loan balance before the original due date... and until 2020, neither did the online portal or the confirmation email' — which had told consumers 'You have successfully extended your loan'. The Bureau found the extension conduct both unfair and deceptive on that basis. The same order attributes the 296-consumer self-service due-date harm to two causes at once: 'First, Enova's code failed to register certain consumers' modification of their payment due date. Second, Enova's online portal allowed consumers to select the SSDDA option even if their upcoming minimum payment had already been irreversibly generated.'

Sources: consumerfinancialprotectionb2023b

Appears on: /domains/cases/enova-lending-algorithm-errors

EmpiricalThe 2019 order records a correction loop failing at every hop, with dates: 'Consumers first notified Enova about this is…

The 2019 order records a correction loop failing at every hop, with dates: 'Consumers first notified Enova about this issue in September 2013. In November 2013, Enova identified a coding error as the source of the problem. It implemented a coding fix in January 2014. When the fix failed ten days later, however, Enova manually disabled it. Enova did not re-enable the fix until May 2014, and did not run daily checks in the interim to ensure that the Flash Cash extension issue had been resolved.' Affected consumers were not told that full loan payments rather than extension fees had been taken from their accounts until April 2015. The 2023 order records the relapse and states the deployment-gate absence directly: after the 2019 order Enova 'continued to obtain consumer bank account information from Lead Generators and launched a project to route its Leads through a newly developed proprietary framework that was intended to, among other outcomes, generate more profitable decisions on extending loan offers to consumers. At the time the new process launched in 2020, no one at Enova had checked to determine whether the new process would overwrite existing consumer bank account information, as it had before.' It did, for 356 consumers, roughly eighteen months after a federal order permanently enjoined that conduct. NOTED AS AN ARGUMENT FROM ABSENCE, with the orders' full text as the evidence: neither the 2019 order nor the 2023 order contains any requirement about code review, change control, regression testing, deployment gating or automated payment-integrity monitoring.

Sources: consumerfinancialprotectionb2019b, consumerfinancialprotectionb2023b

Appears on: /domains/cases/enova-lending-algorithm-errors

EmpiricalThe 2023 order measures detection latency by channel in its own words. Self-detection where instrumented was fast: a cod…

The 2023 order measures detection latency by channel in its own words. Self-detection where instrumented was fast: a coding error in August 2019 that caused Enova to debit or attempt to debit 156 CashNetUSA consumers in Idaho one to three times more than they had authorised was self-identified 'within one week of it first occurring', and the 2020 lead-generator relapse was also self-identified. Complaint-driven detection was slow: on the self-service due-date feature Enova 'received complaints about these unauthorized debits relatively soon after the feature began' but 'it took three months for Enova to identify that it was a systemic problem', and on skip statements consumers complained 'in the first weeks' and 'it took Enova four months'. Complaint-driven detection could fail entirely: on the third-party debit-card processing issue running from as early as 2013 to 2019, 'Respondent received consumer complaints about this issue, but it treated these complaints as isolated and failed to identify it as a systemic problem impacting 1,378 consumers' — a window of up to six years between an arriving signal and its correct classification, on a defect where the vendor reported a payment as failed when it had in fact succeeded and representatives reprocessed it 'in some cases up to four additional times, until there was no error message'. Enova's own characterization is materially softer and is attributed rather than adopted: it says the 2019 issues 'impacted less than 0.2% of total payments processed', that they were 'identified and self-disclosed to the CFPB in 2014', that the 2023 matters 'resulted from unintended computer and system errors', and that 'the majority of items were self-reported by Enova to the CFPB'. The orders record self-identification expressly for two classes and find the opposite on the vendor issue. Neither the vendor nor any lead generator is identified anywhere in the record.

Sources: consumerfinancialprotectionb2023b, enovainternational2019, enovainternational2023

Appears on: /domains/cases/enova-lending-algorithm-errors

EmpiricalThe 2023 order built a control stack far richer than the 2019 order's and terminated early. It required a Compliance Pla…

The 2023 order built a control stack far richer than the 2019 order's and terminated early. It required a Compliance Plan within 90 days with dated implementation steps and a mechanism for apprising the Board of progress; gave the Board 'the ultimate responsibility for ensuring that Respondent complies with this Consent Order' and required the Chief Executive Officer, with the Board, to review all plans, reports and submissions before they went to the Bureau and to authorise corrective actions; required a sworn Compliance Report at one year; required an unaffiliated qualified third-party consulting firm retained within 60 days to determine from Enova's business records whether redress had reached all Affected Consumers, with the Enforcement Director holding non-objection over the firm AND over its sampling protocol with a revise-and-resubmit loop, and a Redress Plan where gaps were found; imposed a seven-year ban on Covered Loans and a ban on using or selling the associated consumer information; and required, first of its kind in Bureau practice, that executive compensation agreements taking effect after the order consider the actions the executive took to ensure compliance, with an annual Executive Compensation Report to the Bureau. Self-provided remediation across the enumerated defect classes sums to roughly $3.8 million, recorded in the orders as Respondent's own representations, against penalties of $3.2 million and then $15 million. On 2 September 2025 the Bureau terminated the order — written to run to at least November 2030 — reciting that the penalty had been paid, the consultant retained, further redress provided to consumers 'whom the third-party consultant determined had not previously received complete redress', and steps taken to implement the conduct provisions, and then stating: 'the Bureau hereby terminates this Consent Order. The Bureau also waives any alleged non-compliance by Enova with the Consent Order.' Enova reports the termination as 'Effective August 29, 2025'. The Covered Loans ban and the executive-compensation provision fell with the order roughly five years early, and a waiver of alleged non-compliance is not a finding of compliance. Enova's own stated remedial steps are architectural: centralised payment processing implemented in 2021, enhanced processes to identify and address customer impacts quickly, and sunsetting its single-payment product in 2022.

Sources: consumerfinancialprotectionb2023b, consumerfinancialprotectionb2025d, consumerfinancialprotectionb2025b, enovainternational2023, enovainternational2026, consumerfederationofamerica2025

Appears on: /domains/cases/enova-lending-algorithm-errors

EmpiricalThe layer neither consent order examines is the one Enova markets, and every statement about it here is the operator's o…

The layer neither consent order examines is the one Enova markets, and every statement about it here is the operator's own, taken from its Annual Report on Form 10-K for the fiscal year ended 31 December 2025 and attributed rather than adopted. Enova describes 'a fully integrated decision engine that evaluates and rapidly makes credit and other determinations throughout the customer relationship, including automated decisions regarding marketing, fraud, underwriting, customer contact and collections that leverage artificial intelligence and machine learning-enabled models', states that the engine 'currently handles more than 100 algorithms and over 1,000 variables', reports approximately 90 data and analytics professionals supporting it, and says its fraud models identify fraudulent applications 'with a very low false positive rate' — a rate Enova does not publish and no regulator has measured. The same filing states the release posture: 'Our software development life cycle is rapid and iterative to increase the efficiency of our platform,' with integration systems designed to 'launch new products rapidly, modify our business operations quickly and account for complex regulatory requirements imposed in the jurisdictions in which we operate.' No regulator, court or auditor has examined the decision engine, and nothing in that description appears in either consent order; no claim about model quality, disparate impact or the false-positive rate is made here. The filing also supplies the operating scale: approximately $7.8 billion in credit or financing extended in 2025, consumer lending in 37 US states plus Brazil, small-business financing in 49 states and the District of Columbia, 1,836 employees, and information collected from nearly 100 million credit reports during 2025. As of July 2026 a coalition of 20 state attorneys general had urged the Federal Reserve to reject Enova's approximately $369 million acquisition of Grasshopper Bank on the ground that nonbank acquisition of banks would let high-cost lenders bypass state usury caps; that is advocacy and a pending application, not a finding.

Sources: enovainternational2026, americanbanker2026

Appears on: /domains/cases/enova-lending-algorithm-errors

EmpiricalEquifax's Online Model Server is the legacy on-premise platform that takes a consumer credit file, derives credit attrib…

Equifax's Online Model Server is the legacy on-premise platform that takes a consumer credit file, derives credit attributes from it, executes third-party scoring models over those attributes, and returns a score — and under some contracts the attributes themselves — to the requesting lender or credit reseller. Many of the attributes are date-relative: whether a consumer has ever been sixty days late on a credit card, the age of the oldest tradeline, the number of inquiries within one month. On 17 March 2022 a code change entered that production environment. The Consumer Financial Protection Bureau found, in consent order 2025-CFPB-0002, that Equifax introduced test code in a production environment in a scoring model server, and that certain scoring models thereafter computed date-based attributes against a fixed reference date rather than the then-current date. Every model downstream then produced arithmetically correct outputs over silently wrong inputs. There was no defect in the scoring models themselves: FICO stated publicly that the problem was an Equifax issue and not a FICO issue. This is therefore not an artificial-intelligence failure, not a discriminatory-design failure, and not an automated decision by Equifax at all — the scoring algorithms behaved correctly and the input pipeline feeding them did not, which is why no model-level audit, fairness test or explainability review would have caught it. The scale is the operator's own, cited to its published material: credit scores maintained on more than 200 million United States consumers, and more than 2.8 billion consumer credit files delivered to United States lenders in 2021.

Sources: consumerfinancialprotectionb2025e, officeofthenewyorkstateattor2025a, nationalmortgagenews2022, warren2022

Appears on: /domains/cases/equifax-score-delivery-error

EmpiricalThe window and the magnitude both come in two framings, and both must be carried. THE WINDOW. Equifax opened an internal…

The window and the magnitude both come in two framings, and both must be carried. THE WINDOW. Equifax opened an internal investigation into the issue on 22 March 2022, five days after the code change. The New York Attorney General's Assurance of Discontinuance No. 24-102 records that the issue was partially resolved on 6 April 2022 and fully resolved on 8 April 2022; the Bureau's order states the error persisted until 8 April 2022; and the settlement class period runs from 17 March to 8 April 2022. Equifax's own public statements say the issue took place between 17 March and 6 April and that the fix was put in place on 6 April. The outer window is therefore twenty-two days, with seventeen of them running after the operator had an investigation open. No source located identifies what triggered the 22 March investigation. THE MAGNITUDE, OPERATOR TIER, from the statement of 2 August 2022: there was no shift in the majority of scores; fewer than 300,000 consumers experienced a score shift of 25 points or more; a score shift does not necessarily mean that a consumer's credit decision was negatively impacted; and, in the consumer-facing statement of 4 August 2022, only a small number of consumers may have received a different credit decision. As of 2026 Equifax still says the vast majority of scores during the three-week period did not change and a large number had a positive shift. THE MAGNITUDE, REGULATOR TIER, reciting Equifax's own score-shift analysis: more than 600,000 consumers were underscored by 10 or more points and 139,000 consumers saw a score decrease of 25 points or more. The two reconcile because the Bureau counts only decreases, and the operator's public number is the one that reads smaller. Equifax told mortgage clients that approximately 12 percent of credit scores calculated from Equifax data during the window may have been impacted — an operator representation, independently confirmed as an Equifax representation by Freddie Mac's notice of 2 June 2022 rather than independently measured.

Sources: officeofthenewyorkstateattor2025a, consumerfinancialprotectionb2025e, equifaxinc2022a, freddiemac2022

Appears on: /domains/cases/equifax-score-delivery-error

EmpiricalThe corrupted values were transient computations rather than stored file contents, and both Equifax and the Bureau recor…

The corrupted values were transient computations rather than stored file contents, and both Equifax and the Bureau record that the contents of consumer credit reports were not changed. That is not a mitigating detail; it is the mechanism that removed the consumer's remedy. There was no wrong entry to dispute, nothing visible in a consumer's own copy of their report, and a dispute channel that processes approximately 765,000 disputes per month had no purchase on a score that was computed wrong at delivery and correctly thereafter. Equifax never notified affected consumers. Its consumer-facing statement of 4 August 2022 told anyone who attempted to obtain credit in the window and thought their decision may have been impacted to reach out to the lender for more information — placing discovery on the applicant, who had no way to learn that the number behind a denial had been wrong. Neither regulator's remedy imposes a duty to notify an affected consumer, and the Bureau's order provides no consumer redress specific to the coding error. The Bureau's unfairness reasoning is squarely about that unobservability: consumers could not avoid the coding and system errors or the method and speed with which the company responded to them, and there is no benefit to consumers or competition of system changes or upgrades without appropriate safeguards that resulted in inaccurate credit scores. Senators Elizabeth Warren and Mark Warner and Representative Raja Krishnamoorthi, and separately House Financial Services Chair Maxine Waters, asked Equifax how the error was detected, how long the company knew before alerting lenders, whether any protected class was disproportionately affected, and whether affected consumers had been identified, notified and made whole. No public response by Equifax to either letter was located, and no source establishes disparate impact by race, income or geography in either direction.

Sources: consumerfinancialprotectionb2025e, equifaxinc2022, warren2022, waters2022

Appears on: /domains/cases/equifax-score-delivery-error

EmpiricalCorrection ran operator to operator rather than to the person. Equifax reissued updated scores and data to lenders, and …

Correction ran operator to operator rather than to the person. Equifax reissued updated scores and data to lenders, and gave lenders updated data for their own custom scores; the Bureau found that when Equifax sold incorrect attributes, other scores generated by third parties may also have failed to reflect a consumer's credit profile, so the contamination propagated into lenders' own custom scorecards. The incident became visible outside Equifax through the customer channel: Equifax began telling lenders and credit resellers in May 2022, National Mortgage Professional published the first press account on 27 May 2022, and the Wall Street Journal published on 2 August 2022, the same day Equifax issued its first public statement. On 2 June 2022 Freddie Mac and Fannie Mae notified their sellers and investors and required affected loans in process to be resubmitted through Loan Product Advisor or Desktop Underwriter with corrected credit data, or to be manually underwritten, with already-sold loans corrected through post-purchase processes under the life-of-loan representation and warranty for data inaccuracies; the Federal Housing Finance Agency worked with both enterprises to determine impacts. That is the only place in the record where correction was mandatory rather than discretionary, and it reaches loans rather than people. Outside the enterprise channel a lender could reprice a loan or invite a denied applicant to reapply and was under no legal obligation to do either. Two per-lender observations reach this record at second hand, through a congressional letter quoting a paywalled article that was not read directly: one large auto lender was told several thousand of its applicants in the window saw changes of 25 points or more, and one large bank reported 18 percent of its applicants were given incorrect scores with an average swing of 8 points. Roughly 2.5 million credit scores were pulled by mortgage lenders from the national bureaus during the three-week window, which is a channel-volume denominator and not a count of affected consumers.

Sources: freddiemac2022, fanniemae2022, nationalmortgagenews2022, warren2022

Appears on: /domains/cases/equifax-score-delivery-error

EmpiricalA separate March 2022 coding error at the same company duplicated disputed collection tradelines in 46,400 consumer file…

A separate March 2022 coding error at the same company duplicated disputed collection tradelines in 46,400 consumer files. The code was remediated on 12 April 2022, but the duplicates themselves were not fully removed until at least November 2022, and the cleanup surfaced approximately 10,000 further system-generated duplicates dating to at least 1 January 2020. American Banker independently reports the same 46,400-account defect and its roughly seven-month cleanup. The pair of defects is the structural content of this record rather than an incidental coincidence: the delivery defect wrote nothing durable and therefore left the consumer no remedy, while the duplication defect wrote durable rows into consumer files, which is exactly the kind of object the statutory dispute machinery and ordinary file-integrity checks can reach. The same dispute capacity that could do nothing about a wrong score could act on a wrong row, and it still took seven months between stopping the write and clearing what the write had already produced.

Sources: consumerfinancialprotectionb2025e, americanbanker2025a

Appears on: /domains/cases/equifax-score-delivery-error

EmpiricalTwo regulators acted three days apart in January 2025 on the same conduct through different instruments, and both aimed …

Two regulators acted three days apart in January 2025 on the same conduct through different instruments, and both aimed at the change-control pipeline rather than at any model. On 14 January 2025 the New York Attorney General accepted Assurance of Discontinuance No. 24-102 under Executive Law section 63(12) and General Business Law sections 349 and 350: Equifax pays $725,000 as restitution and penalties, does not admit any negligence, wrongdoing or violation of law, and agrees to prospective relief comprising Change Advisory Board review of system changes, pre-deployment code review consistent with industry standards, a Fair Credit Reporting Act module in developer training, and at least weekly monitoring of incident reports filed by Equifax's own customers to identify issues with the potential to adversely affect scores. On 17 January 2025 the Consumer Financial Protection Bureau issued consent order 2025-CFPB-0002 against Equifax Inc. and Equifax Information Services LLC. The Online Model Server coding error is ONE of five findings in that order; it is held to violate Fair Credit Reporting Act section 607(b) and to constitute an unfair act or practice under the Consumer Financial Protection Act. Equifax pays a $15,000,000 civil money penalty into the victims relief fund, which covers the whole order rather than the coding error alone and carries no consumer redress specific to it. The order requires policies addressing potential consumer impact in the development, testing and implementation of system changes, including Change Advisory Board review of any change reasonably anticipated to materially impact consumer files or reports once in production, systems to monitor the results of such changes, a senior-executive committee including the Chief Compliance Officer meeting at least quarterly and reporting to the Board, a compliance plan reviewed annually, and developer training on accuracy obligations. The Bureau's enforcement action record for docket 2025-CFPB-0002 reads Post Order / Post Judgment with no subsequent termination or vacatur listed. Neither instrument imposes a duty to notify an affected consumer.

Sources: officeofthenewyorkstateattor2025a, consumerfinancialprotectionb2025e, consumerfinancialprotectionb2025a, consumerfinancialprotectionb2025c, officeofthenewyorkstateattor2025, americanbanker2025a

Appears on: /domains/cases/equifax-score-delivery-error

EmpiricalThe private litigation is live and its posture must be stated exactly. Roughly nine suits filed in August and September …

The private litigation is live and its posture must be stated exactly. Roughly nine suits filed in August and September 2022 were consolidated before Judge Leigh Martin May as In re: Equifax Fair Credit Reporting Act Litigation, No. 1:22-cv-03072-LMM-CCB in the Northern District of Georgia, the lead case being Jenkins v. Equifax, Inc., filed 3 August 2022. On 11 September 2023 the court largely denied Equifax's motion to dismiss: the willful section 1681e(b) claim and the class allegations proceeded, the Georgia common-law negligence claim and the demand for injunctive relief were dismissed, and the court rejected the argument that section 1681e(b) does not reach credit scores. That is a pleading-stage decision on the sufficiency of allegations and NOT a finding that Equifax acted willfully. On 21 June 2024 the court denied Equifax's motion to dismiss for lack of subject-matter jurisdiction, its request to certify a question for interlocutory appeal, and its motion to stay discovery. Equifax agreed in principle in June 2026 to settle nationwide and class-wide, and its Form 10-Q for the quarter ended 30 June 2026 records $100.0 million accrued, a $60.0 million insurance receivable and a $40.0 million net charge, tying that figure to this matter and no other. An unopposed motion for preliminary approval was filed 12 August 2026 by class representatives Sarah Hunter, Maurice Moore and Michael Rodela; preliminary approval was granted 17 August 2026; the final fairness hearing is set for 22 January 2027; the class is approximately four million United States residents whose affected scores or attributes were reported to a third party between 17 March and 8 April 2022; the fund is non-reversionary and no proof of injury is required to claim. Equifax denies liability in the settlement and characterises it as a compromise of disputed claims. Nothing has been paid, the settlement is not final, and no court has adjudicated the merits. Class counsel describe it as the largest FCRA class settlement in history, which is advocacy rather than an established fact. The named plaintiff's account — a report delivered to a lender with a score approximately 130 points below her actual score, a denied auto loan she had been approved for at about $350 a month, and financing obtained elsewhere at an alleged cost about $2,352 a year higher — is a pleading allegation that no court has found.

Sources: inreequifaxfaircreditreporti2022, duanemorrisllp2023, equifaxinc2026, classaction2026, cbsnews2022

Appears on: /domains/cases/equifax-score-delivery-error

EmpiricalHello Digit's automated-savings tool is an algorithm with a money-movement actuator rather than a decision about a perso…

Hello Digit's automated-savings tool is an algorithm with a money-movement actuator rather than a decision about a person: a member grants access to a checking account through a third-party financial-technology platform the consent order does not name, and a proprietary algorithm analyses that account's transaction data to determine when and how much to save, then initiates automated clearing house transfers out of the account into pooled accounts held in the company's name at partner banks. There is no human in the transfer loop, no per-transfer consumer confirmation, no applicant, no eligibility decision, no score and no adverse-action notice. The Consumer Financial Protection Bureau's consent order of 10 August 2022 records that since its inception the company knew its algorithm was not perfect and had limitations that hampered its ability to predict precisely how much to withdraw, and that it was aware of the underlying reasons its algorithm triggers overdrafts, naming transactions that may be stale or inaccurate due to time lags in the banking system or inaccurate data provided by third-party partners, and the impossibility of predicting all customer, bank and automated clearing house behaviour. Settlement takes one to three days, widened to as much as five business days in some cases by the current terms. Two published overdraft rates exist and both are the operator's own, at different denominators: the order records the company's own estimate that approximately 1 to 2 percent of USERS experience overdrafts as a result of using the service, and the acquirer publicly countered that a Digit Save transaction caused an overdraft fee in less than 0.008 percent of CASES. Neither is an audited measurement, and neither may be given as the error rate without its unit.

Sources: consumerfinancialprotectionb2022a, bankingdive2022, oportun2025a

Appears on: /domains/cases/hello-digit-autosave-algorithm

EmpiricalThe failure the record documents is a failure to honour the undertaking after the fact, and the channel it ran through w…

The failure the record documents is a failure to honour the undertaking after the fact, and the channel it ran through was measured and rationed rather than measured and refused. The consent order records nearly 70,000 overdraft-reimbursement requests since 2017 and complaints about overdrafts received daily, against over 7,200 requests declined — roughly one in ten, so most of what reached the channel was paid. The order itemises the declines by reason: 728 caused by Digit's own subscription-fee debit, 1,823 attributed to consumer-initiated manual saves, 600 because the user did not reconnect a checking account after unsubscribing, 1,347 where Digit had already reimbursed the consumer twice, and 1,852 deemed unrelated to Digit activity. Until at least mid-2020 the policy was to reimburse no more than two instances per lifetime, and a consumer overdrawn a third time was refused, which the order says happened over a thousand times since 2017. Reimbursement also required an active account and a still-connected checking account, so a consumer who disconnected or cancelled after being overdrawn was refused unless and until they reactivated and reconnected in the app; the order's only outright behavioural prohibition beyond the misrepresentation ban forbids requiring a consumer to connect a third-party bank account in order to obtain reimbursement, and requires payment by direct deposit or paper check. The order also records that the error signal was fully visible inside the company: a support agent logged a complaint as a user incurring overdrafts after being misled by the no-overdraft-guarantee language at sign-up, another wrote that the language was misleading and should be changed, and an internal document listed both the two-instance maximum and users declining to reconnect among common issues that receive user pushback. The order's recordkeeping clause requires retention of all consumer complaints and reimbursement requests, the results of Digit's auto-reimbursement tool, and any responses, which is the record's evidence that the adjudication was itself at least partly automated. Nothing in the record shows overdraft outcomes being fed back into the transfer decision.

Sources: consumerfinancialprotectionb2022a, consumerfinancialprotectionb2022

Appears on: /domains/cases/hello-digit-autosave-algorithm

EmpiricalThe instrument is a consented administrative order, not an adjudicated finding, and what it does not require is as load-…

The instrument is a consented administrative order, not an adjudicated finding, and what it does not require is as load-bearing as what it does. Hello Digit, LLC consented without admitting or denying any of the findings of fact or conclusions of law, except the facts necessary to establish the Bureau's jurisdiction, which it admits; the separate stipulation of 9 August 2022 phrases the consent as being without admitting or denying any wrongdoing. All three legal conclusions are deception counts under Consumer Financial Protection Act sections 1031(a) and 1036(a)(1)(B) — about the savings tool's accuracy, about the reimbursement of overdraft fees, and about the retention of interest on consumer balances. There is no unfairness count and no abusiveness count. The order imposes a civil money penalty of $2,700,000 and requires a reserve of not less than $68,145 for redress to consumers whose reimbursement requests were denied between 1 January 2017 and the effective date because they exceeded the reimbursement cap or did not reconnect their account, with any shortfall paid to the Bureau; the penalty is therefore roughly forty times the redress floor, the redress class covers only two of the five documented denial reason codes, and $68,145 is a floor to be reserved rather than a measure of the harm caused. Trade reporting, and not the order, is the single located source for the distribution figure of $68,145 across 1,947 consumers, about $35 each. On governance the order installs a compliance plan and a redress plan subject to Enforcement-Director non-objection, review of every submission by Oportun's board or its Audit and Risk Committee with reporting back at least quarterly, compliance reports at 90 days and one year approved by the board and sworn under penalty of perjury, an acknowledgment within 7 days, a distribution list within 90 days, five years of distribution to new officers and any successor entity, and a standing Bureau right to interview employees and compel documents. It imposes no model-accuracy requirement, no validation obligation, no cap on overdrafts caused, no independent monitor and no ongoing public reporting of the overdraft rate, and requires no change to the algorithm. An independent legal analysis published the following day reads the counts the same way — messaging deception rather than a finding that the algorithm was unlawful — and adds a defence-side observation, attributed and not adopted here, that the Bureau's press-release framing leaned harder on the word algorithm than the counts did.

Sources: consumerfinancialprotectionb2022a, consumerfinancialprotectionb2022f, consumerfinancemonitor2022, bankingdive2022

Appears on: /domains/cases/hello-digit-autosave-algorithm

EmpiricalThe deployment was not wound down after the order; it was rebranded, disclaimed and expanded, and the order remains live…

The deployment was not wound down after the order; it was rebranded, disclaimed and expanded, and the order remains live. The product now runs as Oportun Set & Save, and the terms last updated 22 July 2025 replace the undertaking with a Disclaimer of Warranty of Purpose: each prediction merely represents an attempt to predict the relevant facts regarding the linked bank account, accuracy is not guaranteed, the service might not predict an overdraft that actually occurs, the service is provided as is and as available, and sophisticated algorithms cannot anticipate everything. Responsibility for avoiding an overdraft is routed to member-set controls including a Safe Saving Level, with the terms placing on the member the duty to monitor those controls. The same terms set weekly limits of $2,000 for automatic saves and $2,000 for manual saves and a $5,000 cap per instant withdrawal, authorise a debit of the linked bank account where a balance is negative for more than five days, assert a right of setoff across the member's goal balances, permit discretionary closure where a linked account is persistently overdrawn, state that the operator may earn commission or interest on members' funds deposited with its partner banks while the member earns none, name Wells Fargo, JP Morgan Chase and Citi as deposit destinations, and carry a binding arbitration provision with an express class-action and jury waiver. The acquirer reports 1.1 billion algorithmic transfers over ten-plus years, more than $12.5 billion set aside since 2015, restated as over $12.8 billion in July 2026, about $1,800 per member per year, and interest on member accounts of $17.414 million in fiscal 2025 against subscription revenue of $19.465 million. On 27 July 2026 it launched Smart Bills, described as an artificial-intelligence extension of the same engine that identifies recurring obligations and reserves for them ahead of the due date. Every figure in that release is an unaudited operator claim. A present-day reimbursement policy capping save-caused fees at four per calendar month within a two-day window on a consumer-supplied screenshot is corroborated from a help-centre article whose live page returns a single-sign-on shell with no article body, and is treated as corroborated rather than directly verified.

Sources: oportun2025a, oportunfinancialcorporation2026b, oportunfinancialcorporation2026a, oportun2025

Appears on: /domains/cases/hello-digit-autosave-algorithm

EmpiricalPosture, attribution and status, each stated as the record states it. The operator disputes the substance while carrying…

Posture, attribution and status, each stated as the record states it. The operator disputes the substance while carrying the order: Oportun told trade press that it disagreed with the Bureau but wanted to resolve the matter and countered the faulty-algorithm framing with the per-transaction figure, and its securities filings state that it believes the business practices of the company, including Digit's, have been in full compliance with applicable laws and that it agreed to the order in the interest of resolving the matter. Attribution runs one way: the conduct described in the order occurred at Hello Digit — the product launched in February 2015, the civil investigative demand issued in June 2020, the marketing changed in mid-2020 and late 2021 — and Oportun Financial Corporation acquired Hello Digit, Inc. on 22 December 2021 with the investigation disclosed during acquisition diligence, merging it into Hello Digit, LLC, the respondent named in the order. Oportun is therefore the acquirer that inherited a disclosed investigation and now runs the deployment, and the conduct is not attributed to it as actor. Status as of the run date of 28 August 2026: the order is in force and unterminated. The Bureau's enforcement action page lists the matter as Post Order / Post Judgment; the administrative adjudication docket for 2022-CFPB-0007 holds only Document 001, the consent order, and Document 002, the stipulation, both filed 10 August 2022, with no termination entry; and Oportun's Form 10-K filed 27 February 2026 and Form 10-Q filed 6 August 2026 both still carry the order in their risk factors. That finding is stated rather than assumed because the Bureau terminated a number of consent orders of this vintage early during 2025 — trade reporting documents early terminations for U.S. Bank, Apple, Bank of America and Navy Federal Credit Union, plus a move to terminate a 2021 Trustmark redlining order — and this one is not among the terminations located. It rests on negative evidence, an unchanged two-entry docket plus continued securities disclosure, rather than on an affirmative Bureau statement. By its own terms the order runs to 10 August 2027. A minor date discrepancy is carried rather than resolved silently: the stipulation is dated 9 August 2022, the order and docket 10 August 2022, and the acquirer's filings say 11 August 2022; 10 August 2022 is used.

Sources: consumerfinancialprotectionb2022e, consumerfinancialprotectionb2022d, oportunfinancialcorporation2026b, oportunfinancialcorporation2026c, bankingdive2025a, bankingdive2022

Appears on: /domains/cases/hello-digit-autosave-algorithm

EmpiricalNavy Federal Credit Union underwrites residential mortgages through what the pleadings and the Fourth Circuit both descr…

Navy Federal Credit Union underwrites residential mortgages through what the pleadings and the Fourth Circuit both describe as an 'at-least semi-automated underwriting process' built on a 'proprietary underwriting algorithm' whose variables and weights are, in the complaint's words, 'entirely up to Navy Federal,' which 'maintains secrecy' over both; no third-party underwriting vendor is identified anywhere in the verified record. CNN reported on 14 December 2023 that for 2022 conventional home purchase mortgages the credit union approved 77 per cent of White applicants, 69 per cent of Asian applicants, 56 per cent of Latino applicants and 48 per cent of Black applicants — a spread of nearly 29 percentage points, the widest of the fifty lenders that originated the most mortgages that year — and that holding more than a dozen variables constant Black applicants were more than twice as likely to be denied as White applicants and Latino applicants roughly 85 per cent more likely. Navy Federal disputes that analysis, saying the statistics 'do not appear to have considered several key credit criteria,' and adds that it ranks first among large lenders in the share of mortgages made to Black borrowers, made $3.5 billion in 2022 mortgages to Black borrowers, and counts roughly one in four members as Black. Separately, and unconditionally, the CFPB's own HMDA Data Browser queried directly for this institution (LEI 5493003GQDUH26DNNH17) across all reportable loan types and purposes gives denial shares of originated-plus-denied applications, White against Black, of 23.8 and 45.0 per cent in 2018, 31.0 and 50.8 in 2022, 34.5 and 56.6 in 2023 and 27.4 and 46.8 in 2025; those counts control for nothing whatsoever, and in particular not for credit score, which the public loan-level file excludes by rule. No court and no regulator has ever found that Navy Federal discriminated in mortgage underwriting.

Sources: unitedstatescourtofappealsfo2026, tolan2023, consumerfinancialprotectionb2026, navyfederalcreditunion2023

Appears on: /domains/cases/navy-federal-mortgage-underwriting

EmpiricalOne loan file generates two records with different contents, and the difference decides who could say what. The full Hom…

One loan file generates two records with different contents, and the difference decides who could say what. The full Home Mortgage Disclosure Act submission that goes to the regulator carries the applicant credit score, the debt-to-income ratio and the loan-to-value; the public modified loan-level file suppresses the applicant credit score by rule and coarsens other fields. CNN published that limit on its own instrument in the same article that carried its findings, and the Fourth Circuit's opinion recites it: applicant credit score, available cash deposits and relationship history with the lender are not available in the public mortgage data. The NCUA Office of the Chief Economist's research note documents the same split from the other side, describing which credit-risk fields the regulatory file contains and applying a federal logit model to them. The consequence is structural rather than rhetorical: an analysis built on the published file can be answered, accurately and unfalsifiably, by pointing at the variable it could not observe, and that answer cannot itself be checked, because the file it rests on is not published either. When Congressional Black Caucus members asked Navy Federal in February 2024 for 'at minimum, aggregate data ... regarding credit scores or any other non-public variable that Navy Federal has suggested serves as an explanation' for the gap, the March 2024 congressional letter records that 'Navy Federal failed to provide this information.'

Sources: unitedstatescourtofappealsfo2026, nationalcreditunionadministr2022, cleaver2024, consumerfinancialprotectionb2026

Appears on: /domains/cases/navy-federal-mortgage-underwriting

EmpiricalThree analyses of this lending exist, each run from a different slice of the data, and no forum ever set one against ano…

Three analyses of this lending exist, each run from a different slice of the data, and no forum ever set one against another. CNN's was built on the public file with the credit score removed by rule, and published its own caveats. Navy Federal's was commissioned from Debo P. Adegbile and reported complete on 21 March 2024: the credit union announced 'no race-based decision making' once 'all non-public underwriting factors are accounted for — including credit score, income verification, debt-to-income ratio, and incomplete credit applications,' and Adegbile stated that 'when all relevant factors are controlled for, which CNN did not do, the difference in approval rates between Black and White borrowers falls to less than 1%,' with the remainder 'explained by legitimate, non-race factors like income verification and incomplete credit applications.' The report, the model and the underlying data have never been released; Adegbile is a partner at WilmerHale, which is Navy Federal's defense counsel in the same class action and argued the appeal for it; plaintiffs' counsel called the arrangement 'a classic conflict of interest'; and the credit union nonetheless described the work as an 'external review.' The third analysis is the NCUA Office of the Chief Economist's, applying the FDIC's Popick logit model to 2020 and 2021 HMDA data including the credit-score, debt-to-income and loan-to-value fields the public file suppresses, and reporting credit union denial odds against White applicants of about 1.60 for Black applicants on conventional purchase, 1.55 for Hispanic applicants and 1.60 for Asian applicants, 1.54 on rate-and-term refinance and 1.81 on cash-out refinance, with average marginal effects of two to four percentage points and contract-rate premiums of eight to thirteen basis points — published as an INDUSTRY-WIDE aggregate that names no institution and separately estimates none, and carrying the caveat that such differences 'should not be interpreted as evidence of discrimination ... as such results may reflect unobserved factors.' Congress supplied volume and no compulsion: at least four letter campaigns, a chief-executive meeting requested by forty members, and a ten-question demand of 1 March 2024 asking what matters requiring attention had issued in five years, how fair lending findings had affected ratings, whether a referral to the Department of Justice had been made, and how each agency ensures a lender searches for a less discriminatory alternative — a test Consumer Reports had put to the Bureau ten weeks earlier — with answers requested by 5 April 2024. No public response has been located.

Sources: navyfederalcreditunion2024, navyfederalcreditunion2024a, wral2024, nationalcreditunionadministr2022, cleaver2024, consumerreports2023

Appears on: /domains/cases/navy-federal-mortgage-underwriting

EmpiricalOliver v. Navy Federal Credit Union, No. 1:23-cv-01731-LMB-WEF, was filed in the Eastern District of Virginia on 17 Dece…

Oliver v. Navy Federal Credit Union, No. 1:23-cv-01731-LMB-WEF, was filed in the Eastern District of Virginia on 17 December 2023; related suits were consolidated and the operative complaint of 20 February 2024 named nine plaintiffs pleading the Fair Housing Act, the Equal Credit Opportunity Act, 42 U.S.C. section 1981 and state analogues. On 30 May 2024 Judge Leonie Brinkema dismissed the disparate-treatment theory because 'the Complaint has failed to allege plausible direct or circumstantial evidence of discriminatory intent,' dismissed four further counts, preserved the disparate-impact theory on the ground that at the motion-to-dismiss stage 'the statistical disparities reveal a disparate impact among non-white loan applicants and the underwriting algorithm and process is alleged to have caused the disparity,' and struck all class allegations before any discovery. On 9 February 2026 the Fourth Circuit decided No. 24-1656 in a published two-to-one opinion: Rule 23(c)(1)(A) rather than Rule 12(f) or Rule 23(d)(1)(D) is the source of a district court's authority to decide certification, and before discovery a court may deny certification only if the class allegations fail as a matter of law on their face. It affirmed the denial of a Rule 23(b)(3) damages class and vacated the denial of a Rule 23(b)(2) injunctive class, holding the complaint made 'a sufficient prima facie showing' of commonality and that the district court 'acted prematurely,' and it reserved the merits expressly: 'Of course, discovery in this case might show that the [complaint's] allegations ... are false.' Judge Richardson, concurring in part and dissenting in part, would have affirmed the strike, arguing the complaint uses the plural elsewhere, says nothing about how the algorithm and loan officers interact, and that the disparity 'may be caused not by any underwriting algorithm, but by the individual loan officers exercising their discretion in differing ways.' On 25 August 2026 all nine named plaintiffs filed a notice of dismissal with prejudice, which the district court so-ordered the same day; the 28 August status conference was cancelled. The docket text states no reason and discloses no settlement. The correct posture is: concluded by voluntary dismissal with prejudice, merits never reached, no class ever certified, and no finding of discrimination by any court.

Sources: unitedstatescourtofappealsfo2026, u2026b, oliverv, oliverandjacobv2023, ababankingjournal2026

Appears on: /domains/cases/navy-federal-mortgage-underwriting

EmpiricalThe supervisory test moved while the dispute was live. Executive Order 14281 of 23 April 2025 directed federal agencies …

The supervisory test moved while the dispute was live. Executive Order 14281 of 23 April 2025 directed federal agencies to eliminate the use of disparate-impact liability, and on 4 September 2025 the National Credit Union Administration issued Letter to Credit Unions 25-CU-04 removing every reference to disparate impact from its Fair Lending Guide and other issuances and stating that its 'examination and supervision processes will no longer include reviews for disparate impact,' while continuing HMDA analysis and examinations for disparate treatment. The theory the district court dismissed in Oliver is the one the supervisor kept; the theory that survived dismissal and was revived for class treatment on appeal is the one the supervisor stopped examining for. In the same window the consumer-compliance supervisor's posture toward this institution reversed on an entirely separate matter: the Consumer Financial Protection Bureau ordered Navy Federal on 7 November 2024 to pay more than $95 million — $80.6 million in redress and a $15 million penalty — over surprise overdraft fees charged between 2017 and 2022, and on 1 July 2025 the Bureau terminated that order and waived any alleged non-compliance, a decision House Financial Services Democrats questioned in an August 2025 letter to the credit union's chief executive, which records Navy Federal saying it 'firmly believe[s] the CFPB's decision to terminate the order was appropriate.' The overdraft matter concerns deposit-account fees and is not evidence about mortgage lending; it is recorded here as evidence about a supervisor's posture and capacity.

Sources: nationalcreditunionadministr2025, nationalcreditunionadministr2025a, consumerfinancialprotectionb2024b, democraticstaff2025

Appears on: /domains/cases/navy-federal-mortgage-underwriting

EmpiricalOportun Financial Corporation (Nasdaq: OPRT), founded in 2005 as Progreso Financiero and certified by the U.S. Treasury …

Oportun Financial Corporation (Nasdaq: OPRT), founded in 2005 as Progreso Financiero and certified by the U.S. Treasury as a community development financial institution since 2009, lends $300 to $10,000 against alternative data to borrowers with little or no credit history, and accepts an Individual Taxpayer Identification Number in place of a Social Security number. It is a first-party creditor operating every stage of the deployment in house: no third-party model vendor, no third-party collections-litigation vendor and no debt buyer appears anywhere in this record. At the documented time it had 793,254 active customers and originated 726,964 loans worth $2.05 billion in 2019, lending in twelve states through about 80 Texas retail locations and 213 California storefronts, with roughly 3.9 million loans worth about $9 billion extended cumulatively by the end of that year; contact-centre servicing ran from Mexico, Colombia and Jamaica, two of those outsourced. Its FY2025 annual report gives the current shape: $1.957 billion originated in 2025, a 4.9 per cent thirty-day-plus delinquency rate, a 12.0 per cent annualised net charge-off rate against a stated strategic target of nine to eleven per cent, 126 retail locations, 1,580 employees in Mexico including two contact centres, and $21.8 billion extended across more than 8.0 million loans and cards in nineteen years. The portfolio's identification profile is documented at the defendant level: of the 467 borrowers the company sued in one Texas county in June 2020, fewer than half had a Social Security number on its file. The operator's own counterweight to the litigation record, offered in the same coverage and unverified, is that it sued on fewer than 6 per cent of loans over the preceding five years and that 92 per cent of customers historically repaid on time and in full.

Sources: propublica2020a, oportunfinancialcorporation2026b, oportunfinancialcorporation2021a, communitydevelopmentfinancia2026

Appears on: /domains/cases/oportun-collections-machine

EmpiricalThe filing counts in this case are court-records analyses by parties outside the deployment, and each is attributed to t…

The filing counts in this case are court-records analyses by parties outside the deployment, and each is attributed to the analysis that produced it. ProPublica and The Texas Tribune assembled 1.45 million debt-claim records from 62 justice of the peace courts in nine of Texas's ten largest counties — scraping seven counties' online dockets, filing public-records requests with more than a dozen individual courts in three others, with one county refusing — standardised more than seventy spellings of the plaintiff's name, and found that Oportun had sued borrowers more than 47,000 times from May 2016 through July 2020. The reporters state in print that imperfect plaintiff-name matching makes that an UNDERCOUNT. Nearly 10,000 of those suits were filed in the first half of 2020 alone, more than half of them after the World Health Organization declared a pandemic in mid-March, against more than 9,000 distinct borrowers; across the nine counties the company was the most litigious personal-loan company in Texas and the second-most litigious company of any kind in that window, and a top filer in eight of the nine. Separately and with a different scope, The Guardian analysed records available in 20 of California's 58 counties and found more than 30,000 collections suits in 2019 and at least 14,000 in the first half of 2020, with more than 15,000 Los Angeles County filings in 2019 — about one for every 667 residents — and at least 15 per cent of ALL California small-claims filings between mid-2017 and mid-2018. The Center for Responsible Lending independently analysed California's ten most-populous counties and reached compatible figures of at least 23,500 cases in 2019 and over 13,000 in 2020, with a Los Angeles County series showing this first-party consumer lender out-filing Midland Funding and Portfolio Recovery Associates, the two largest national debt buyers, in each of 2018, 2019 and 2020. The three analyses have different scopes, are not additive with one another, and no national total exists. A date seam in the published methodology is carried rather than resolved: the note describes the collected dataset as running from January 2015 to 30 June 2020 while the headline count is stated for May 2016 through July 2020, and no source reconciles the two ranges.

Sources: propublica2020a, propublica2020, thetexastribune2020, theguardianhosseini2020, centerforresponsiblelendingb2021

Appears on: /domains/cases/oportun-collections-machine

EmpiricalThe mechanism that produced the volume is documented as a threshold and a staffing model, not as a model in the machine-…

The mechanism that produced the volume is documented as a threshold and a staffing model, not as a model in the machine-learning sense, and this distinction is the record's own. Borrowers were contractually in default after a single missed payment; The Guardian describes the operator running 'a robust process for filing small claims actions against customers who fall 60 days behind in their payments'; ProPublica describes non-attorney employees titled 'legal collections specialists', some straight out of college, filing suits en masse in justice of the peace courts. No source in this record — journalistic, advocacy, regulatory or corporate — describes an algorithm, classifier or automated engine selecting who gets sued, and no source describes any review step between the delinquency counter and the courthouse or any documented authority to decline a referral. Describing the filing pipeline as automated or algorithmic is an inference from throughput and is not made here. The machine learning the operator does claim sits elsewhere: its own securities filings describe traditional and alternative data ingested as 'billions of data points' to build underwriting, pricing, marketing, fraud and servicing models, with a claim to score 100 per cent of applicants, and describe the post-2020 collections apparatus as expanded digital and telephony contact, broadened eligibility for payment-difficulty tools, self-enrolment in the app and on the web, 'supported by a new collections strategy system that enables centralized, faster, and more-targeted application of strategies.' Servicing models are claimed; a litigation-selection model is claimed by nobody. Every one of those capability statements is the operator's own description in its own filings and none has been examined or validated by any regulator, auditor or researcher.

Sources: theguardianhosseini2020, propublica2020a, oportunfinancialcorporation2026b

Appears on: /domains/cases/oportun-collections-machine

EmpiricalThe forums were chosen for the asymmetry they create and the outcome distributions follow from it. Texas justice of the …

The forums were chosen for the asymmetry they create and the outcome distributions follow from it. Texas justice of the peace courts cap a claim at $10,000, charge about $50 to file and permit a non-lawyer to file; California small-claims courts cap a high-volume filer at $2,500, do not provide a guaranteed interpreter to a defendant population for whom Spanish is often the first language, and bar legal counsel on both sides — which in practice reaches only the defendant, because the plaintiff arrives with a trained specialist and a stack of identical filings. The median claim in Harris and Dallas counties in 2020 was approximately $1,400. Roughly one in three Los Angeles County cases in a 1,165-case sample from June and December 2019 ended in a DEFAULT judgment, the defendant never having appeared; in Tulare County the operator won uncontested judgments in 38 per cent of the 755 suits it filed in calendar 2018. A default judgment supports wage garnishment and accrues interest for at least ten years. The one channel that reliably stopped a case was a lawyer: of about 7,600 Harris County defendants in 2019, 105 obtained counsel and 96 per cent of those cases were dismissed, with a Dallas consumer attorney reporting that the company dismissed cases as soon as it learned a defendant was represented. That is a review channel whose measured effect when exercised is close to total and whose exercise rate is about 1.4 per cent. Roughly two-thirds of all filed Texas suits were eventually dropped without judgment — a filing pattern that a law professor quoted in the investigation characterised as harassment and intimidation rather than litigation, which is his characterisation and not a finding. The Center for Responsible Lending argues that adverse credit histories and related lawsuits count against immigrants applying for permanent residency or citizenship, which would make a $1,400 judgment a categorically heavier event in this borrower population; that is attributed advocacy reasoning rather than any tribunal's finding, and it is the reason the harm surface here cannot be measured in dollars alone.

Sources: propublica2020a, theguardianhosseini2020, centerforresponsiblelendingb2021, centerforresponsiblelending2022

Appears on: /domains/cases/oportun-collections-machine

EmpiricalThe sequence of the intervention is the structural point of this case and it runs in the order the sources establish. Th…

The sequence of the intervention is the structural point of this case and it runs in the order the sources establish. The Guardian put its California court-records findings to the company. Four days later, on 28 July 2020 — and BEFORE either investigation published, the Guardian's on 2 August and ProPublica and the Texas Tribune's on 31 August — Oportun announced four things: an all-in 36 per cent annual-percentage-rate cap on new originations nationwide, fully implemented by mid-August; immediate dismissal of all pending legal-collection cases; suspension of all new filings, for an unstated period; and a commitment to reduce future filings by more than 60 per cent. The forcing function was private observation by an outsider, not public exposure. The announcement was explicitly a reduction rather than an exit: the chief executive's own words in the release were 'As we continue to provide affordable unsecured loans, legal collections remains necessary, but we are committing to the development of new tools and approaches that better reflect who we are', alongside a separate line describing the prior position as one that 'does not reflect our objectives as a mission-driven company'. The company's own account of why it looked at all is on the record: it had not, on its own account, compared its filing counts to those of its peers until 'recent media inquiries' prompted it, and the chief executive said that when he did he found the company 'near the top ... and in some counties, we were the top'. Oportun acknowledged the causal chain in its own SEC risk factors, and resented it: the July 2020 changes were 'partially the result of inquiries we received from certain consumer advocates and media outlets', after which 'certain media outlets and consumer advocates chose to highlight and have continued to highlight the very past practices that we had already modified' — a sentence still present, in slightly edited form, in the FY2025 annual report six years later.

Sources: oportunfinancialcorporation2020, theguardianhosseini2020, propublica2020a, oportunfinancialcorporation2021a, oportunfinancialcorporation2026b

Appears on: /domains/cases/oportun-collections-machine

EmpiricalThe reform overshot the promise on the operator's own books and undershot it in the courts, and both are documented. On …

The reform overshot the promise on the operator's own books and undershot it in the courts, and both are documented. On the audited books it did more than it had promised: the FY2020 Form 10-K records severance for 'ceasing of legal collections' and a $3.6 million expense decrease 'related to ceasing legal collection on default loans beginning in August 2020', and the FY2021 Form 10-K states that the company 'dismissed all pending small claims court filings and suspended all new legal collection actions and have not restarted legal collections programs'. A promised 60 per cent cut appears in the audited expense line as the shutdown of a function. Audited independently, the dismissals were less complete: in a random sample of 106 California cases filed in 2020 the Center for Responsible Lending found only 52 per cent dismissed WITH prejudice, 36 per cent dismissed without prejudice and therefore refilable, 4 per cent already producing default judgments and 9 per cent still pending, and the Legal Aid Society of San Diego found 477 of 500 cases filed in 2020 still pending as of 26 January 2021. The price half of the commitment held and is measurable: the FY2025 annual report states 'We have capped the APR for newly originated loans at 36% since August 2020', with a weighted average APR at origination of 35.2 per cent and a weighted average term of 38 months at 31 December 2025, against a book whose stated average rate at the documented time had run around 34 to 36 per cent with rates reaching 66.99 per cent in Texas and California, and where about 43 per cent of the operator's 374,488 sub-$2,500 unsecured California loans in a 2018 regulatory filing had carried APRs between 40 and 69.9 per cent. That cap must not be read as wholly voluntary. California's Fair Access to Credit Act had already capped consumer loans of $2,500 to $10,000 at 36 per cent plus the federal funds rate, operative 1 January 2020, with existing licensees transitioned by 1 July 2020 — three and a half weeks before the announcement, in the company's largest market. What the voluntary cap actually bound was Texas and the California loans under $2,500. The commitment was real and it has held for six years; it was also, in part, an announcement of compliance.

Sources: oportunfinancialcorporation2021a, oportunfinancialcorporation2022, centerforresponsiblelendingb2021, oportunfinancialcorporation2026b, californiadepartmentoffinanc, oportunfinancialcorporation2020

Appears on: /domains/cases/oportun-collections-machine

EmpiricalThere is no adverse legal outcome against Oportun anywhere in this record. The Consumer Financial Protection Bureau serv…

There is no adverse legal outcome against Oportun anywhere in this record. The Consumer Financial Protection Bureau served a civil investigative demand on 3 March 2021 as part of a broader small-dollar-lending inquiry, and follow-up requests narrowed it to the company's 'legal collection practices from 2019 to 2021 and hardship treatments offered to members during the COVID-19 pandemic'. On 15 September 2022 Bureau enforcement staff sent a Notice and Opportunity to Respond and Advise letter — the closest this record ever comes to a government allegation — stating that it was considering whether to recommend legal action based on 'failure to timely dismiss certain lawsuits and the hardship treatments offered during the COVID-19 pandemic, including credit reporting related thereto'. Oportun disputed the allegations in writing on 14 October 2022. On 28 March 2023 the company announced, and its next quarterly report recorded, that the Bureau had completed the investigation and that its Office of Enforcement staff would not recommend pursuing an enforcement action. There is no finding, no consent order, no penalty and no admission in this matter. The parallel channels ended the same way. Oportun applied to the Office of the Comptroller of the Currency for a national bank charter on 23 November 2020; more than forty organisations including LULAC, UnidosUS, the National Consumer Law Center and Consumer Reports objected on 22 December 2020 citing 'egregious debt collection practices'; nearly two dozen groups asked the Acting Comptroller in August 2021 to hold the application until the Bureau finished; and the company withdrew the application voluntarily on 8 October 2021, saying it intended to amend and refile. It is still not a bank. Advocacy groups asked Treasury to revoke the community-development certification in December 2020; it was not revoked, and Oportun, Inc. appears on the CDFI Fund's list of currently certified CDFIs dated 14 August 2026 under CDFI number 131CE011959. The only sanction this deployment ever incurred was reputational. Separately and to keep two records disjoint: Oportun completed its acquisition of Hello Digit, Inc. on 22 December 2021, after the conduct in this matter and after the conduct behind that company's own federal matter; the two share a corporate parent from that date and nothing else, and the consent order in that matter is that company's, on an automated savings product, and is not Oportun's.

Sources: oportunfinancialcorporation2022, oportunfinancialcorporation2023, oportunfinancialcorporation2023a, oportunfinancialcorporation2023b, propublicawiththetexastribun2021, oportunfinancialcorporation2021, bankingdive2021, woodstockinstitute2020, communitydevelopmentfinancia2026, consumerfinancialprotectionb2022

Appears on: /domains/cases/oportun-collections-machine

EmpiricalThe record's currency gap is stated rather than papered over. No legal-collections disclosure of any kind survives in Op…

The record's currency gap is stated rather than papered over. No legal-collections disclosure of any kind survives in Oportun's FY2023, FY2024 or FY2025 annual reports — not the suspension, not the possibility of resumption, not the investigation. The last affirmative statement on the record is the FY2022 report's line that the company 'temporarily suspended our legal collections process, which may be resumed in the future'; after the investigation closed in March 2023 the subject left the filings. No independent court-records analysis of Oportun's filings has been published for any year after 2020. Absence of disclosure is not evidence of non-resumption, and whether the company files collection suits today is not established in either direction: the defensible statement is that the practice has not been independently examined since 2020. What the FY2025 report does describe is the replacement: expanded digital and telephone contact, broadened eligibility for payment-difficulty tools, self-enrolment in the app and on the web, and a new collections strategy system the company describes as enabling centralised, faster and more-targeted application of strategies — an operator description that no regulator, auditor or researcher has examined and for which no performance figure is published. Two further facts are reported side by side and are not joined: the operator's annualised net charge-off rate rose from 6.8 per cent in 2020 to 12.0 per cent in 2024 and 2025 while legal collections stayed off, and the company attributes that rise to new-borrower mix and cost-of-living pressure over a period that covers a credit cycle. No source establishes that losses rose because the courthouse channel closed. Finally, Raul Vazquez, the chief executive who made the July 2020 commitments and gave the executive quotations in the 2020 coverage, announced his departure on 21 January 2026, so no present-tense corporate intent is attributed to him.

Sources: oportunfinancialcorporation2023, oportunfinancialcorporation2026b, oportunfinancialcorporation2026

Appears on: /domains/cases/oportun-collections-machine

EmpiricalThe visibility surfaces of this deployment are asymmetric in a way the sources document precisely, and the asymmetry is …

The visibility surfaces of this deployment are asymmetric in a way the sources document precisely, and the asymmetry is what made the case both possible and invisible for four years. What was NOT public: Texas justice courts do not post petitions, so what tens of thousands of suits alleged is not publicly readable; the state's central electronic filing system excludes justice courts, so no state actor held the aggregate; and the Texas Office of Consumer Credit Commissioner, the licensing authority nearest the conduct, held the operator's annual lending-activity reports and refused to release them as confidential, so the supervisor closest to the practice was the one actively withholding the denominator. What WAS public: the docket line — plaintiff, defendant, date, court and case type — carrying the defendant's own name, city and often street address, machine-readable in seven of the nine counties studied and aggregable by anyone with a scraper. That was exactly enough. Two newsrooms and an advocacy research organisation reconstructed the operator's conduct from those rows alone, using a scraper, a set of public-records requests and a question. The same rows are a permanent public exposure for every person named on one, and a judgment entered on one becomes a derogatory credit entry read by every other lender for at least a decade. Oportun created every one of those rows, held them in its own systems throughout, and — on the chief executive's own account — did not compare them against its peers until journalists asked.

Sources: propublica2020a, propublica2020, centerforresponsiblelending2022

Appears on: /domains/cases/oportun-collections-machine

EmpiricalSantander Consumer USA Inc. was an indirect subprime vehicle lender: it bought retail installment contracts from franchi…

Santander Consumer USA Inc. was an indirect subprime vehicle lender: it bought retail installment contracts from franchise and independent dealerships rather than lending across its own counter. Its FY2020 annual report describes the deployment in its own words: robust historical data on both organically originated and acquired loans is used to perform advanced loss forecasting, and each applicant is automatically assigned a risk score using information from credit bureau and credit application, placing the applicant in one of multiple pricing tiers, which the company continuously maintains and adjusts to reflect market and risk trends, with interest rate, down payment and loan-to-value named as the material components of risk-based pricing. A manual underwriting team is retained for manual review, consideration of exceptions, and review of deal structures with dealers; the record describes no per-application human underwriter. Scale at the last full public year: 5,576 employees; 1,938,764 retail installment receivables outstanding at 31 December 2020 against 1,810,973 a year earlier; $32.9 billion of retail installment contracts held for investment; $26.6 billion of total originations; average origination credit-bureau score 626 and average annual percentage rate 14.1 per cent on retained contracts, against 598 and 16.3 per cent in 2019. The correction leg ran at the same scale: 285,661 repossessions in 2019, 15.7 per cent of average receivables outstanding, and 177,639 in 2020 at 9.3 per cent, a figure suppressed by a nationwide suspension of involuntary repossession at the pandemic's onset. The company was taken private on 31 January 2022 and stopped filing, so FY2020 is the last full public operating record.

Sources: santanderconsumerusaholdings2020, santanderholdingsusa2022

Appears on: /domains/cases/santander-subprime-auto-scoring

EmpiricalThe executed consent judgments give the scored artifact its operative name and its operative thresholds: the 'loss forec…

The executed consent judgments give the scored artifact its operative name and its operative thresholds: the 'loss forecasting score', with the remedy banded at 401 or below, 501 or below, 502 to 600, and 601 or above. California's complaint, the only pleading in the multistate group that opens the architecture, describes two chained models: one ingesting the consumer's borrowing history together with the applied-for deal's loan-to-value, debt-to-income, payment-to-income, mileage and term and emitting a probability that the consumer becomes severely delinquent within a defined window, converted into a scaled score on a proprietary, FICO-like scale; and a separate life-of-the-loan model mapping a given proprietary score to a probability of default before the end of the term. The complaint alleges that for at least part of the time period examined by the People, Santander projected that consumers with the lowest proprietary scores had a greater than 70 per cent likelihood of default over the life of the loan; that figure is a pleading allegation, is hedged in the pleading itself, is scoped to the lowest scores rather than to subprime borrowers generally, and is a modelled probability from the second model rather than an observed default rate. The coalition's shared formulation is softer and appears verbatim across state releases: that Santander, through its use of sophisticated credit scoring models to forecast default risk, knew that certain segments of its population were predicted to have a high likelihood of default. The judgment never describes the architecture and the complaint never uses the name; treating them as one artifact is a reasonable inference from their shared function, scale shape and settlement context rather than a documented identity.

Sources: officeoftheillinoisattorneyg2020, complaintforcivilpenalties2020, southcarolinaattorneygeneral2020

Appears on: /domains/cases/santander-subprime-auto-scoring

EmpiricalCalifornia's complaint alleges that the operator's detection apparatus for dealer abuse existed and did not bind. It des…

California's complaint alleges that the operator's detection apparatus for dealer abuse existed and did not bind. It describes a problematic-dealer tracking process running since as early as 2010, and internal tension at Santander between punishing problematic dealers and retaining Santander's market share, with reluctance to act against flagged dealers so long as enough of their paper was profitable; a preferred-lender arrangement with a manufacturer under which flagged dealers were allowed to participate; and a stated-income policy rolled out without barring dealers with a history of misstating income, which it says led to a significant spike in the number of early payment defaults. Massachusetts, settling an earlier and separate action on 29 March 2017 for $22 million alongside Delaware for up to about $4 million, alleged that the operator funded loans without a reasonable basis to believe borrowers could afford them, predicted that a large portion of the loans would default, knew dealer-reported incomes were often inflated, kept lending through dealers it had internally flagged including a group it called fraud dealers, and that its own internal audit had concluded its dealer oversight was inadequate. The operator's own FY2020 annual report describes the same apparatus without the adverbs: early-payment-default monitoring used to identify dealers subject to more extensive documentation requirements or exclusion, and dealer agreements capping the finance charge a dealer may retain. All of the characterisations above are allegations in actions resolved without admission and without adjudication of any fact or law.

Sources: complaintforcivilpenalties2020, officeofthemassachusettsatto2017, stateofdelaware2017, santanderconsumerusaholdings2020

Appears on: /domains/cases/santander-subprime-auto-scoring

EmpiricalOn 19 May 2020 a coalition of 34 attorneys general — 33 states plus the District of Columbia — announced parallel consen…

On 19 May 2020 a coalition of 34 attorneys general — 33 states plus the District of Columbia — announced parallel consent judgments carrying approximately $550 million, effective 1 May 2020. The package is heterogeneous and must be broken out: approximately $65 million in cash restitution to a settlement-administrator trust; $5 million to the multistate working group for fees and costs; up to $2 million for administration, reverting if unspent; up to $45 million of loan forgiveness letting defaulted consumers scoring 401 or below who had not yet been repossessed keep the vehicle and the title; and approximately $433 million of immediate deficiency forgiveness on defaulted loans the operator still owned, plus further waivers on loans it had to attempt to buy back. The going-forward regime is the substantive part. No purchase of a loan where the sole obligor's residual income at origination — gross monthly income less monthly debt obligations, less a reasonable estimate of basic living expenses, less a reasonable estimate of payroll taxes — is zero or negative; a reasonable debt-to-income threshold re-evaluated at least annually with no purchase above it, and quarterly statistically relevant sampling for calculation accuracy and threshold compliance. Santander shall not substantially change its loss forecasting score formula without sixty days' advance notice to a Monitoring Committee describing the change and its potential impact on the back-test. An income reasonability model, due by 31 December 2020, using historical consumer, third-party and geographic data to score confidence in stated income and route low-confidence applications to additional manual review, with annual reassessment of its assumptions and two-year document retention for every update. Mandatory additional Treatments — screens, documentation requirements, stipulations — for any dealer known or reasonably suspected of income inflation, expense deflation or power booking, which may not be waived or excepted until the dealer has demonstrably fixed the problem. And a quarterly back-test for four years: for every future default, recompute residual income at origination, and where it was zero or negative waive the deficiency and request deletion of the tradeline, with the qualifying time-to-default window widening as the original forecast worsened — eighteen months for a score of 501 or below, twelve for 502 to 600, six for 601 or above. Most injunctive subparagraphs run seven years from implementation, with a contingent restart across all 34 jurisdictions if a court in any coalition state adjudges a violation. The judgments were entered without the taking of proof, without trial or adjudication of any fact or law, without any admission of liability, and state that they do not constitute approval of the operator's business practices.

Sources: officeoftheillinoisattorneyg2020, southcarolinaattorneygeneral2020, officeoftheattorneygeneralfo2020

Appears on: /domains/cases/santander-subprime-auto-scoring

EmpiricalThe outbound credit-reporting channel is the one surface in this record where a federal regulator made findings rather t…

The outbound credit-reporting channel is the one surface in this record where a federal regulator made findings rather than alleging. On 22 December 2020 the Consumer Financial Protection Bureau issued a consent order, Docket 2020-BCFP-0027, finding that between January 2016 and August 2019 Santander Consumer USA Inc. furnished consumer loan information to credit reporting agencies that it knew or reasonably should have known was inaccurate, failed to promptly correct it, omitted dates of first delinquency, and lacked reasonable written policies and procedures for accuracy, and imposing a $4.75 million civil money penalty under the Fair Credit Reporting Act and Regulation V. That is a furnishing and record-keeping failure rather than an underwriting one. The same channel is what the multistate remedy uses to deliver: for every consumer receiving a deficiency waiver or vehicle-and-title relief, the operator must notify each credit reporting agency it reports to and request deletion of the tradeline, and for loans that defaulted between 1 January 2010 and 31 December 2012 it may neither collect the deficiency nor sell the loan.

Sources: consumerfinancialprotectionb2020, officeoftheillinoisattorneyg2020

Appears on: /domains/cases/santander-subprime-auto-scoring

EmpiricalCalifornia's complaint locates the alleged defect in the inputs rather than in the estimator, and says so in terms: alth…

California's complaint locates the alleged defect in the inputs rather than in the estimator, and says so in terms: although Santander has sophisticated models that forecast consumer default, Santander's policies with respect to stated income and expenses allow it to underestimate default risk in important ways, and Santander employs modeling that makes use of housing costs that are based on faulty information and therefore likely incorrect. The pleaded practices are that the operator generally let applicants state mortgage and rent expenses without proof and had no apparent measure against falsified housing figures; that where housing cost was not stated it assumed a default amount that would not be reasonably sufficient to pay for mortgage or rent in the vast majority of localities; and that from early 2013 it made an aggressive push to waive proof of income on most applications. An independent measurement of the same gap was reported on 22 May 2017 by Bloomberg News from a Moody's Investors Service report of 17 May 2017 using newly available asset-backed issuer data: income was verified on 8 per cent of borrowers in a Santander asset-backed deal against 64 per cent for a contemporaneous AmeriCredit deal, and loans combining low or no credit score, no co-signer and no income verification were about 9 per cent of Santander's pool balance against under 1 per cent of AmeriCredit's. A separate Moody's report put roughly 42 per cent of Santander's 2009 to 2014 subprime loans written through dealers identified as high-risk in the Massachusetts and Delaware settlements as having defaulted or being expected to. The operator's contemporaneous response through its treasurer was that the income-verification practice had been consistent over time though lower than competitors', that the higher losses in the loans backing the bonds had been visible to investors, and that bondholders were protected by loss cushioning in the bonds.

Sources: complaintforcivilpenalties2020, scully2017

Appears on: /domains/cases/santander-subprime-auto-scoring

EmpiricalThe securitisation question has to be stated narrowly or it is wrong. The multistate complaint pleads no securitisation …

The securitisation question has to be stated narrowly or it is wrong. The multistate complaint pleads no securitisation theory at all: it pleads origination of loans known to be highly likely to fail, non-verification of income and expenses, dealer-abuse blindness, and deceptive servicing. What is verified is different and narrower. The multistate civil investigative demands sought documents on the operator's underwriting, securitization, servicing and collection of nonprime vehicle loans; a Department of Justice civil subpoena under the Financial Institutions Reform, Recovery, and Enforcement Act sought documents on the underwriting and securitization of nonprime auto loans since 2007 and was last disclosed as open in the FY2019 annual report, with no public resolution located; Securities and Exchange Commission subpoenas from October 2014 opened an investigation into securitization practices that resolved on 17 December 2018 as an accounting and internal-controls case about the credit loss allowance for certain impaired loans, with a $1.5 million civil penalty and no admission or denial; and Massachusetts and Delaware settled a funding-and-securitisation theory on 29 March 2017. The 2020 remedy is itself bounded by the funding structure: the judgment defines 'Owns' as on the company's balance sheet and not part of a securitization, deficiency waivers are owed on owned loans or on sold loans the operator can repurchase at or below its sale price using best efforts within 150 days of the effective date, and the prospective back-test reaches securitised loans only to the extent permitted by the relevant securitization documents. What is not supported is that the loss was transferred to investors and that this is why the forecast did not restrain origination: the operator's securitisations were largely on-balance-sheet secured financings with approximately $26 billion outstanding, its treasurer stated publicly that the higher losses were visible to investors and that bondholders were protected by loss cushioning, and the rating agency made no claim that noteholders were at risk. The honest formulation is that the predicted loss was priced and funded rather than avoided, and that the party it landed on was the borrower.

Sources: officeoftheillinoisattorneyg2020, santanderconsumerusaholdings2019, u2018b, scully2017, santanderconsumerusaholdings2020

Appears on: /domains/cases/santander-subprime-auto-scoring

EmpiricalOversight of this deployment is layered and almost entirely external, and it changed the deployment rather than only des…

Oversight of this deployment is layered and almost entirely external, and it changed the deployment rather than only describing it. The multistate investigation opened in March 2015 after the Illinois Attorney General's office recorded an increase in consumer complaints, with a six-state executive committee — California, Illinois, Maryland, New Jersey, Oregon and Washington — serving civil investigative demands in October 2014, May 2015, July 2015 and February 2017 on behalf of a group the operator's own FY2019 annual report describes as 33 state attorneys general and the District of Columbia. The 2020 judgments created a Monitoring Committee whose authority is specific and whose cadence is sparse: a compliance report on written request, no more than annually unless a report shows non-compliance, in which case a remediation plan comes to the Committee, which objects or does not object within thirty days; sixty days' advance notice of any substantial change to the loss forecasting score formula; and a right to make written requests about specific dealers, with records kept at least three years. Adjacent authorities acted on separate surfaces: the Federal Reserve Bank of Boston entered a written agreement in March 2017 requiring enhanced compliance risk management and board and senior-management oversight, closed in February 2021; Mississippi, never a coalition member, sued in January 2017 and settled on 21 July 2021 for $3.7 million including $1.8 million of consumer restitution; Massachusetts returned on 18 February 2022 with a $5.56 million assurance of discontinuance for more than 1,000 borrowers over insufficient disclosure of how post-repossession deficiency balances were calculated; two Department of Justice servicemember consent orders in February 2015 and October 2021 concerned repossession process and lease administration and contain no model; and on 3 June 2026 the New York State Department of Financial Services settled for a $400,000 penalty plus more than $275,000 of restitution over undisclosed recurring monthly extension fees where the disclosure documents showed only a single $25 fee. As of 28 August 2026 no public compliance report under the 2020 judgments and no adjudicated violation of them was located.

Sources: officeoftheillinoisattorneyg2020, santanderconsumerusaholdings2019, santanderconsumerusaholdings2020, magnoliatribune2021, officeofthemassachusettsatto2022, newyorkstatedepartmentoffina2026, autoremarketing2015

Appears on: /domains/cases/santander-subprime-auto-scoring

EmpiricalFrom 2002 TransUnion LLC, one of three nationwide consumer reporting agencies, sold subscribers an add-on to the ordinar…

From 2002 TransUnion LLC, one of three nationwide consumer reporting agencies, sold subscribers an add-on to the ordinary credit report — marketed and litigated variously as 'OFAC Advisor', 'OFAC Name Screen' and, in TransUnion's own filings with the Securities and Exchange Commission, 'the OFAC Alert service'. On a subscriber credit pull carrying the append, TransUnion passed the consumer's first and last name, and nothing else, to third-party software and data held by a vendor, Accuity, Inc., which compared it against the U.S. Treasury Department's Specially Designated Nationals and Blocked Persons list; a match was written into the 'SPECIAL MESSAGES' section on the front page of the report sold to the subscriber. The Ninth Circuit found that TransUnion introduced the product on the belief that it was exempt from the Fair Credit Reporting Act because the data sat in the vendor's file rather than its own database, that on that basis 'TransUnion did not follow its normal procedures to ensure accuracy', and that it adopted a policy of not disclosing OFAC matches to consumers who requested their own reports. It summarised the result as name-only searches 'for more than a decade, resulting in thousands of false positives and not a single known actual match identified' — an evidentiary absence in one case over one class window rather than a proven zero. On 27 February 2011 Sergio Ramirez was refused the sale of a car at a Dublin, California dealership after a credit check returned an OFAC alert and a salesman told him his name was on a 'terrorist list'; his wife bought the car in her own name.

Sources: unitedstatescourtofappealsfo2020, supremecourtoftheunitedstate2021, unitedstatescourtofappealsfo2010, transunion2016

Appears on: /domains/cases/transunion-name-screen

EmpiricalThe comparison used two fields and the disambiguating fields existed on both sides of it. Before November 2010 it fired …

The comparison used two fields and the disambiguating fields existed on both sides of it. Before November 2010 it fired on names that were 'either identical or similar' — the Ninth Circuit's worked example is that 'Cortez' would match 'Cortes'; from November 2010 an exact first-and-last-name match was required, which the court records as cutting the false-positive rate from about five per cent to about half a per cent (a litigation-record figure with no published methodology, denominator or independent audit, and order of magnitude only). No date of birth, middle initial, Social Security number, citizenship or address was compared at any point. Meanwhile TransUnion held the consumer's date of birth and Social Security number in its own CRONUS database and, the Third Circuit found, required a creditor to supply at least a name AND an address to retrieve from it, while sending only a name to Accuity 'even though Trans Union may have more information about the person who is the subject of the inquiry'; it 'neither compares the OFAC information to other information about a given consumer already in its files, nor does it compare it to any information provided by the creditor/subscriber'. The sanctions records themselves carried the listed persons' first, middle and last names, dates of birth and passport information, and TransUnion reprinted exactly those fields in the letter it mailed the consumer. The Ninth Circuit recorded that for tax liens and bankruptcy judgments TransUnion used at least one identifier besides the name, and that 'OFAC information was the only consumer-report data that TransUnion collected using name alone'. The Third Circuit called the failure 'to take the utmost care in ensuring the information's accuracy — at the very least, comparing birth dates when they are available' reprehensible; the district court reasoned that with a birth-date comparison 'none of the class members would be even a potential match'.

Sources: unitedstatescourtofappealsfo2020, unitedstatescourtofappealsfo2010, unitedstatesdistrictcourtfor2017, supremecourtoftheunitedstate2021

Appears on: /domains/cases/transunion-name-screen

EmpiricalThe corrective loop was closed at both ends and the two halves must be read together. From 2002 the consumer-facing copy…

The corrective loop was closed at both ends and the two halves must be read together. From 2002 the consumer-facing copy of the report did not show the OFAC alert, and TransUnion's own witnesses acknowledged that the personal credit reports it gives consumers never show any information or alerts from the OFAC product it provides to creditors. The Fair Credit Reporting Act's dispute channel was switched off for this data class by policy: the Third Circuit recorded that 'once Trans Union receives the OFAC information it does not check or confirm its accuracy; in fact, Trans Union has a policy of never reinvestigating disputes involving OFAC alerts', and TransUnion's call centre told both named plaintiffs there was no alert on their report while an alert sat on the version being sold. Against that, the district court found that TransUnion 'removed the OFAC Alert of each class member who contacted Trans Union following receipt of the OFAC letter' — a removal rate of one hundred per cent on the contacting subset — and cited it against TransUnion's own contention that correction was technically infeasible. Between 1 January and 26 July 2011 TransUnion sent two envelopes a day apart: the credit report with the OFAC alert redacted plus the statutory summary of rights, then a separate 'OFAC Letter' that named the potential match, reprinted the matched sanctions records with their dates of birth and passport information, gave no dispute instructions, omitted the summary of rights, and never stated that the alert appeared on the version sold to third parties. TransUnion stopped that practice in July 2011 and began putting the alerts directly on consumer-facing reports. Treasury's OFAC publishes consumer guidance describing the same loop and routing removal back through the Fair Credit Reporting Act dispute process at the bureau — the channel this operator's policy had closed.

Sources: unitedstatescourtofappealsfo2020, unitedstatescourtofappealsfo2010, unitedstatesdistrictcourtfor2017, uc

Appears on: /domains/cases/transunion-name-screen

EmpiricalIn one seven-month window in 2011 the product labelled 8,185 people, and the law treated them differently according to s…

In one seven-month window in 2011 the product labelled 8,185 people, and the law treated them differently according to something none of them could observe. A class of 8,185 was certified in July 2014; on 21 June 2017 a jury awarded $984.22 statutory and $6,353.08 punitive damages per class member, about $60 million in total, and in February 2020 the Ninth Circuit affirmed liability and standing while cutting punitive damages to $3,936.88 per member. On 25 June 2021 the Supreme Court held 5-4 that only the 1,853 class members whose reports were provided to third-party businesses had suffered a concrete harm and had standing on the reasonable-procedures claim, and only the named plaintiff on the two mailing claims: 'the mere existence of inaccurate information, absent dissemination, traditionally has not provided the basis for a lawsuit in American courts', the harm being 'roughly the same, legally speaking, as if someone wrote a defamatory letter and then stored it in her desk drawer.' THE HOLDING IS ABOUT ARTICLE III STANDING AND NOT ABOUT ACCURACY: footnote 5 states that the parties assumed TransUnion violated the statute even as to those whose alerts were never disseminated and that the Court 'take[s] no position on that issue'; the judgment was reversed and remanded and the case then settled, so no merits judgment survives in the principal action. Justice Thomas, dissenting with three colleagues, wrote that 'in a 7-month period, it is undisputed that nearly 25 percent of the class had false OFAC-flags sent to potential creditors... If 25 percent is insufficient, then, pray tell, what percentage is?' TransUnion's Form 10-K states that the ruling left 'only approximately 23% of the class' with concrete harm. Final approval of a class settlement was entered on 19 December 2022 (a 15 December 2022 final-approval hearing date appears in legal trade press) and TransUnion paid on 20 January 2023; the amount is not stated in any primary document located, with a $9 million fund, $4.2 million in class-counsel fees, a $75,000 service award and an estimated recovery above $2,000 per claimant reported by legal trade press. The Illinois Appellate Court, taking judicial notice of the federal record, recorded that the settlement covered 'the 1,853 class members determined to have standing, as well as an additional 147 class members who demonstrated standing pursuant to a claims process' — 2,000 of the original 8,185. On 23 January 2023 the remaining members' claims were dismissed without prejudice; the same day one of them refiled in the Circuit Court of Cook County, Illinois, a forum with no concreteness requirement, and on 31 March 2025 the Illinois Appellate Court affirmed dismissal as time-barred, declining equitable tolling.

Sources: supremecourtoftheunitedstate2021, unitedstatescourtofappealsfo2020, illinoisappellatecourt2025, transunion2023, transunion2020, consumerfinancialserviceslaw2023

Appears on: /domains/cases/transunion-name-screen

EmpiricalThe vendor boundary was the architecture and the legal theory at once, and the contractual export of responsibility was …

The vendor boundary was the architecture and the legal theory at once, and the contractual export of responsibility was rejected twice. The Third Circuit held in August 2010 that OFAC alerts are part of the consumer report and subject to the maximum-possible-accuracy duty: 'We do not believe that Congress intended to allow credit reporting companies to escape the disclosure requirement... by simply contracting with a third party to store and maintain information that would otherwise clearly be part of the consumer's file.' It also rejected the subscriber addendum as a defence — 'We are not persuaded that Trans Union's private contractual arrangements with its clients can alter the application of federal law' — and the district court likewise rejected the argument that contractual human review discharged the statutory duty. In October 2010 officials at Treasury's own OFAC wrote to TransUnion saying they continued to hear from its customers and from individual consumers adversely affected by false OFAC alerts, warning that a product 'that does not include rudimentary checks to avoid false positive reporting can create more confusion than clarity and cause harm to innocent consumers', and that they were 'particularly worried' by alerts 'disseminated broadly in conjunction with credit reports'; the letter carried no enforcement power. The Ninth Circuit found that TransUnion then 'made surprisingly few changes': the November 2010 wording change and exact-match tightening, with further software enhancements requested from Accuity 'not implemented until 2013', and name-only matching continuing until then. TransUnion's own position on appeal — OPERATOR ADVOCACY, not a finding — is that reporting potential matches 'even though other information (like date of birth) could disprove an actual match, is part of the trade-off inherent in the credit-check process' and that 'lenders have a strong interest in an OFAC product that casts a wide initial net and then relies on a lender's human judgment'. A TransUnion master agreement dated March 2015 and filed with the Securities and Exchange Commission in 2020 carries clause 4.6 'OFAC Name Screen': the name screened is the one 'supplied by Subscriber to TransUnion on input and not as may be found on TransUnion's database(s)', and the subscriber 'shall be solely responsible for taking any action that may be required... and shall not deny or otherwise take any adverse action against any consumer which is based, in whole or in part, on TransUnion's OFAC Name Screen services.' What the matcher compares today is NOT ESTABLISHED by any verified source: a 2020 complaint alleges TransUnion represented it gained the ability to consider dates of birth in 2013 and continued to disregard available birth dates thereafter, but that case was voluntarily dismissed with prejudice on 7 March 2022 with no class certified and nothing adjudicated, and TransUnion's Form 10-K for fiscal year 2025 makes no mention of OFAC or Ramirez. Separately and unrelatedly, an October 2023 Federal Trade Commission and Consumer Financial Protection Bureau settlement required TransUnion Rental Screening Solutions, Inc. and Trans Union LLC to pay $15 million over eviction and criminal-record accuracy in TENANT screening reports; it is a different product, is not about OFAC name screening, and must not be cited as such.

Sources: unitedstatescourtofappealsfo2010, unitedstatescourtofappealsfo2020, unitedstatesdistrictcourtfor2017, transunionllc2020, transunionllcandupstartnetwo2015, alshaikliv2020, transunion2026, ftccfpb2023

Appears on: /domains/cases/transunion-name-screen

EmpiricalThe public class-certification order of 5 August 2025 in In re Wells Fargo Mortgage Discrimination Litigation recites, o…

The public class-certification order of 5 August 2025 in In re Wells Fargo Mortgage Discrimination Litigation recites, on facts the plaintiffs did not dispute, a three-part home-lending stack that is not an artificial-intelligence system. CORE is a front-end workflow tool that 'guides the flow of the loan origination process' and 'retrieves, stores, and displays' application data, Risk Engine outputs and final lending decisions, and it 'is not an automated underwriting system and it does not make lending decisions or calculate or assign Credit Risk Classes.' A separate application, the Risk Engine, assigns Credit Risk Classes using business rules, business services and data attributes; it 'currently executes tens of thousands of separately identifiable business rules' grouped into 16 business services, the largest of which, Get Risk Decision, 'contains 14,000 rules' and '1,300+ data attributes.' ECS comprises exactly two scorecard models, 11419 for government loans and 11960 for conventional loans, each of which per the bank's declaration 'is a simple scorecard whose logic can be fully stated on 2-3 sheets of paper,' 'neither utilizes artificial intelligence,' and each is 'simple enough that an applicant's score could be calculated by hand'; they map credit-bureau attributes to a score and thence to a Credit Risk Class and also generate risk insight messages displayed to underwriters. Wells Fargo's position, accepted as undisputed for certification purposes, is that these are distinct applications and that 'there is no such thing as CORE/ECS.' Underwriters 'are instructed that the Risk Engine never supersedes the judgment of the underwriter' and '[u]nderwriting is done by human underwriters.' Separately, Bloomberg News reported on 11 March 2022, from an analysis of Home Mortgage Disclosure Act data covering roughly 8 million 2020 refinance applications, that Wells Fargo approved 47 per cent of Black homeowners' completed 2020 refinance applications, 53 per cent of Hispanic or Latino homeowners' and 72 per cent of white homeowners' — the largest racial gap among major lenders and the only major lender to reject more Black refinance applicants than it approved. Those figures are carried here because a Senate letter of 16 March 2022 and a Senate Banking Committee release of 17 March 2022 independently restate them; the article itself was not retrievable for verification and no claim is made about what its analysis controlled for. Wells Fargo did not dispute the arithmetic of the counts, attributed the gap to 'additional, legitimate, credit-related factors' including credit scores, home appraisals and broader economic inequity, and a spokesman said the analysis was 'designed to present a skewed picture of our lending efforts.' That attribution has never been tested by any adjudicator, and no court and no regulator has ever found that Wells Fargo discriminated in refinance underwriting.

Sources: unitedstatesdistrictcourtfor2025, donnan2022, warren2022a, ussenatecommitteeonbanking2022, fortune2022

Appears on: /domains/cases/wells-fargo-refi-underwriting

EmpiricalThe evidence that could settle causation in this matter sits on both sides of a wall, and there are two walls. On the pu…

The evidence that could settle causation in this matter sits on both sides of a wall, and there are two walls. On the public side, the Consumer Financial Protection Bureau's March 2023 report on the Home Mortgage Disclosure Act rule states that 'Credit score and free form text fields used to report the name and version of credit scoring models are excluded from the public loan-level data,' so no public-data analysis of any lender can control for the one variable this bank says explains the gap. The explanatory ceiling on that public record is measured rather than argued: a 2025 defence-side analysis of ten large lenders' 2024 filings found public-data models never exceeding about 56 per cent explanatory power and argued that sampled manual loan-file review remains necessary to confirm any statistical indication, matching the interagency fair-lending examination procedure in which regression is the scoping step and file review is the confirmation step; and on the confidential regulator-held version that DOES contain credit score, a Federal Reserve study of roughly nine million applications reports explanatory power of about 39.8 per cent, leaving most of the variation in loan outcomes unexplained. That defence-side analysis is consulting work published by a firm that serves defendants in fair-lending litigation; its factual points are corroborated by the two government sources and its framing is not neutral. On the private side, the internal rule base and the two scorecards were produced only in discovery under a stipulated protective order, and the class-certification and summary-judgment records are extensively sealed, in a proceeding that stopped before any merits ruling. The consequence is structural rather than rhetorical: an analysis built on the published record can be answered, accurately and unfalsifiably, by pointing at the variable it could not observe, and that answer cannot itself be checked, because the object it rests on is under seal.

Sources: consumerfinancialprotectionb2023c, bhutta2022, chernin2025, inrewellsfargomortgagediscri2026

Appears on: /domains/cases/wells-fargo-refi-underwriting

EmpiricalThe two testifying statisticians converged on the measurement and diverged on causation, and the case turned on that gap…

The two testifying statisticians converged on the measurement and diverged on causation, and the case turned on that gap rather than on either analysis. As recited in the class-certification order, the plaintiffs' expert Dr. Amanda Kurzendoerfer ran a regression controlling for key underwriting factors and found statistically significant approval-rate disparities favouring white applicants that 'cannot be explained by legitimate underwriting factors.' At a concurrent expert hearing held on 12 February 2025, Wells Fargo's expert Dr. Marsha J. Courchane 'did not substantially disagree with the logistic analysis ... or with her findings of a statistical disparity along racial lines'; her objection was causal, that without individual file review one cannot tell whether a denial was wrongful, and that 'you will almost always find an underwriting disparity, on average, for minorities.' Neither expert report, neither specification and none of the underlying data is public, because the expert record is sealed. The burden allocation is the other half of the shape. The order records the plaintiffs' characterisation of Wells Fargo's own documents as showing that it 'could not itself identify what was driving the disparity' and that it 'found that there were three potential proxies for race in its ECS Model — major derogatories, average months in file, and recent inquiries — which could be causing the disparity'; the court held that this cut AGAINST the plaintiffs, because the burden of showing class-wide causation was theirs. In May 2021, six members of Wells Fargo's own Corporate Model Risk group had published a paper warning that historical data skew and automated feature engineering can 'miss the potential for correlated surrogate variables causing proxy discrimination,' that black-box algorithms have 'potential for serious harm' in consumer lending, and that models 'must be continually monitored for disparate impact testing'; the plaintiffs quoted that paper back at the bank in the operative complaint. Every allegation in that complaint — the 'digital redlining' theory, race imputation by surname and geocoded location, uncorrected appraisal values, credit-score overlays above agency minimums, the A1/A2/C1/C2 decision states, the point-of-sale intake platform, the processor caseloads rising from about 30 applications a month to more than 50 and sometimes nearly 100, the systematic disincentive to check the work, the monthly internal report circulating the racial breakdown of lending, and the account of nine months of lost paperwork followed by approval the day after a federal housing agency was notified — is an allegation that was never adjudicated and never conceded.

Sources: unitedstatesdistrictcourtfor2025, zhou2021, inrewellsfargomortgagediscri2023

Appears on: /domains/cases/wells-fargo-refi-underwriting

EmpiricalSix private actions filed in early 2022 were consolidated on 18 January 2023 before Judge James Donato as In re Wells Fa…

Six private actions filed in early 2022 were consolidated on 18 January 2023 before Judge James Donato as In re Wells Fargo Mortgage Discrimination Litigation, No. 3:22-cv-00990-JD (N.D. Cal.), pleading a nationwide putative class under the Equal Credit Opportunity Act, the Fair Housing Act, 42 U.S.C. section 1981, California's Unruh Civil Rights Act and the California Unfair Competition Law. There was never a motion to dismiss in the consolidated case: Wells Fargo Bank, N.A. answered the Amended and Consolidated Class Action Complaint on 17 May 2023 and the case went straight to discovery, expert work and Rule 23 briefing. On 5 August 2025 class certification was DENIED for failure of Rule 23(a)(2) commonality. Numerosity had been conceded on the plaintiffs' estimate of 'at least 119,100' members for classes narrowed to minority applicants approved by an external automated underwriting system or by Wells Fargo's own ECS and ultimately denied. The court wrote that the plaintiffs 'did not present any classwide evidence whatsoever of robust causality' and had 'focused like a laser on the statistical disparity in application denial rates, without anything in the way of explanatory factors,' and that 'the fact that Wells Fargo potentially analyzed 1,300+ data attributes pursuant to 14,000 rules indicates that commonality with respect to denial of a mortgage is not at all obvious here'; it relied on Wal-Mart Stores v. Dukes and treated the undisputed presence of human underwriter discretion as bringing the case within the line of authority refusing certification of discretionary-policy lending classes. It reached no other Rule 23 element and made no merits finding. Both petitions for permission to appeal, Ninth Circuit Nos. 25-5262 and 25-5267, were denied on 8 January 2026. The case then resolved without any ruling: a Notice of Settlement was filed on 7 May 2026 covering four named plaintiffs and any parties represented by that counsel who were potential class members, the court granted dismissal with prejudice and relieved interim lead counsel on 14 May 2026, three further plaintiffs were dismissed with prejudice between 8 May and 1 June 2026, and the last remaining individual claim was stayed on 28 July 2026 through 25 September 2026 with the court noting that further requests to continue the stay are not likely to be granted. Settlement terms are confidential and no amount is public; the notice extinguishes claims without any concession. The pending summary-judgment motion was never decided. No publicly announced investigation, enforcement action or finding by any federal regulator concerning Wells Fargo refinance underwriting discrimination has been located, and the bank's Form 10-K for fiscal 2025 records that 'In August 2025, the district court denied class certification and plaintiffs' interlocutory appeal of the decision was denied in January 2026' while disclosing no government proceeding or investigation concerning mortgage-lending discrimination and disclosing several unrelated government matters. One month after publication, on 13 April 2022, Wells Fargo announced a $150 million Special Purpose Credit Program to lower rates and refinancing costs for Black homeowners it already serviced plus $60 million in WORTH grants aimed at roughly 40,000 homeowners of colour in eight markets through 2025; those are operator-reported, unaudited figures for a programme that concedes no disparity in underwriting.

Sources: unitedstatesdistrictcourtfor2025, inrewellsfargomortgagediscri2026, inrewellsfargomortgagediscri2026a, unitedstatesdistrictcourtfor2026, inrewellsfargomortgagediscri2026b, wellsfargocompany2026, wellsfargocompany2022, ababankingjournal2025, bankingdive2025c

Appears on: /domains/cases/wells-fargo-refi-underwriting