10 to the 23 AI logo

Domain Atlas / Child welfare & family services

Case fileEngland, United Kingdommedium deployment

What Works for Children's Social Care ML pilots

None of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities cleared the pre-specified 65% average-precision success bar; the best single model reached only about 42% average precision and, at an operating point, missed roughly 79% of the children whose cases actually escalated.[3]

What happened

What Works for Children's Social Care — a What Works Centre funded by the Department for Education — ran an approximately 18-month project building machine-learning models in-house with four English local authorities (the programme appears to have recruited up to five, with one withdrawing). Reported in the September 2020 report "Machine learning in children's services: does it work?", the work built 32 models across eight escalation outcomes — such as referral to statutory services within 12 months of early help, escalation to a child protection plan, or a child becoming looked-after — testing decision trees, logistic regression and gradient boosting, and in some builds adding pseudonymised free-text case notes processed with natural-language techniques. The team pre-registered a public 65 percent average-precision success threshold before the work began; none of the 32 models cleared it. The best single model reached only about 42 percent average precision and, at an operating point, missed roughly 79 percent of the children whose cases actually escalated, while on aggregate the models missed around four in five children genuinely at risk and were wrong roughly six in ten times they flagged a child. The authors attributed the poor performance to the rarity of the outcomes, similarity between sibling cases, limited historical training data and the complexity of care pathways. A survey of 129 social workers carried out for the project found low professional support, and executive director Michael Sanders concluded it was time for the sector to "stop and reflect." A companion ethics review by the Alan Turing Institute and the Rees Centre (January 2020) recommended national standards and found systems that devalue person-centred approaches "not ethically permissible." The models were never placed in front of practitioners in live casework.

The sociotechnical reading

Several systems in this Atlas were stopped before deployment — by a minister's refusal, a missing legal basis, a base-rate test. This one was stopped by a rule written down in advance: a pre-registered effectiveness threshold, committed to in public before the models were built and scored against real outcomes, that the tool had to clear to go live. It did not, so it never reached a caseworker. Two lessons distinguish it. First, you cannot upgrade your way past a rare-outcome prediction problem: when serious escalation is a needle in a haystack, accuracy has a ceiling no model change climbs, and the honest governance move is the evidence gate that refuses deployment — not another model. Second, transparency is itself a control: publishing a negative result, and pressing the operationally deployed tools the programme could only ask to disclose their effectiveness, is what makes an evidence bar enforceable across a market. The feedback-loop concern the companion ethics review flagged — that models trained on past intervention decisions learn recorded practice, not underlying risk — is the reason "does it work?" has to be asked against real outcomes, in advance, by someone willing to answer no. In map terms, this is the conformity-gate and vendor-gate patterns exercised before go-live rather than after harm: prove it, publish it, and be willing to not deploy.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

claytonandsanders2022GroundingAcademicSave

Clayton and Sanders, Can Machine Learning Save Children at Risk? (Significance, Royal Statistical Society) (2022) https://academic.oup.com/jrssig/article/19/6/22/7072840

https://academic.oup.com/jrssig/article/19/6/22/7072840

Grounds: model org: wwcsc_ml_pilots

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalNone of the 32 machine-learning models What Works for Children's Social Care built across four English local a…

None of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities cleared the pre-specified 65% average-precision success bar; the best single model reached only about 42% average precision and, at an operating point, missed roughly 79% of the children whose cases actually escalated.

claytonandsanders2022GroundingAcademicSave

Clayton and Sanders, Can Machine Learning Save Children at Risk? (Significance, Royal Statistical Society) (2022) https://academic.oup.com/jrssig/article/19/6/22/7072840

https://academic.oup.com/jrssig/article/19/6/22/7072840

Grounds: model org: wwcsc_ml_pilots

EmpiricalIn a survey of 129 social workers carried out for the project, only about 26% supported using predictive analy…

In a survey of 129 social workers carried out for the project, only about 26% supported using predictive analytics to identify families for early help and about 34% thought it should not be used at all.