Domain Atlas / Child welfare & family services
What Works for Children's Social Care ML pilots
None of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities cleared the pre-specified 65% average-precision success bar; the best single model reached only about 42% average precision and, at an operating point, missed roughly 79% of the children whose cases actually escalated.[3]
What happened
What Works for Children's Social Care — a What Works Centre funded by the Department for Education — ran an approximately 18-month project building machine-learning models in-house with four English local authorities (the programme appears to have recruited up to five, with one withdrawing). Reported in the September 2020 report "Machine learning in children's services: does it work?", the work built 32 models across eight escalation outcomes — such as referral to statutory services within 12 months of early help, escalation to a child protection plan, or a child becoming looked-after — testing decision trees, logistic regression and gradient boosting, and in some builds adding pseudonymised free-text case notes processed with natural-language techniques. The team pre-registered a public 65 percent average-precision success threshold before the work began; none of the 32 models cleared it. The best single model reached only about 42 percent average precision and, at an operating point, missed roughly 79 percent of the children whose cases actually escalated, while on aggregate the models missed around four in five children genuinely at risk and were wrong roughly six in ten times they flagged a child. The authors attributed the poor performance to the rarity of the outcomes, similarity between sibling cases, limited historical training data and the complexity of care pathways. A survey of 129 social workers carried out for the project found low professional support, and executive director Michael Sanders concluded it was time for the sector to "stop and reflect." A companion ethics review by the Alan Turing Institute and the Rees Centre (January 2020) recommended national standards and found systems that devalue person-centred approaches "not ethically permissible." The models were never placed in front of practitioners in live casework.
The sociotechnical reading
Several systems in this Atlas were stopped before deployment — by a minister's refusal, a missing legal basis, a base-rate test. This one was stopped by a rule written down in advance: a pre-registered effectiveness threshold, committed to in public before the models were built and scored against real outcomes, that the tool had to clear to go live. It did not, so it never reached a caseworker. Two lessons distinguish it. First, you cannot upgrade your way past a rare-outcome prediction problem: when serious escalation is a needle in a haystack, accuracy has a ceiling no model change climbs, and the honest governance move is the evidence gate that refuses deployment — not another model. Second, transparency is itself a control: publishing a negative result, and pressing the operationally deployed tools the programme could only ask to disclose their effectiveness, is what makes an evidence bar enforceable across a market. The feedback-loop concern the companion ethics review flagged — that models trained on past intervention decisions learn recorded practice, not underlying risk — is the reason "does it work?" has to be asked against real outcomes, in advance, by someone willing to answer no. In map terms, this is the conformity-gate and vendor-gate patterns exercised before go-live rather than after harm: prove it, publish it, and be willing to not deploy.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.