Domain Atlas / Child welfare & family services
ProKid (Netherlands)
The Dutch government's own 2011 pilot evaluation of ProKid found that 36% of the tool's red, orange and yellow child-risk flags (902 of 2,444 over three months across four police regions, rising to 53% in Amsterdam-Amstelland) were system or registration errors or based on irrelevant incidents, and that in none of the four regions was there a well-functioning instrument.[2]
What happened
ProKid Signaleringsinstrument 12-, built by the Gelderland-Midden police, sorted children under 12 who appeared anywhere in police records into four colour risk categories — white, yellow, orange, red — and handed the flags to police quality controllers, who decided which children to refer to the Youth Care Agency. It scored children from up to twelve years of police records: incidents in which a child was a suspect, a witness, or a victim, plus reports registered at the home address and its co-residents — so a child logged only as a victim could accrue risk, and the risk of the address could escalate the child's own. The government-commissioned pilot evaluation (DSP-groep for the WODC, running 2009 to 2011 across four regions) found that of 2,444 red, orange and yellow flags in one three-month window, 902 — 36% — were system or registration errors or based on irrelevant incidents (53% in Amsterdam-Amstelland), and concluded that in none of the four regions was there a well-functioning instrument, with the "scientific" colour categories carrying little weight in practice and showing no difference in the intensity of help children needed. A later successor, ProKid Plus — an automated police-data model whose validation reported ROC AUC values around 0.83 for predicting future violent or property offending — was reported to parliament in December 2022 to have been used exactly once, as a pilot input to Amsterdam's Top400 programme in July 2016, and no longer used, with a further successor not yet built.
The sociotechnical reading
ProKid is the Atlas's record-construction case — the leverage sits in how a child's record is assembled, retained, and corrected, not in who acts on the score. The override that governance checklists ask for was present: quality controllers re-checked the criteria before referring anyone. But the colour categories "carried little weight" — the evaluation found no outcome difference across the risk bands — so the score was never the operative thing, and re-checking a flag ran case by case. What the override could not reach sat upstream, in the record the score was computed from: which children get profiled, from which records, kept how long. A child's colour was assembled from up to twelve years of police history — their own, their address's, their household's — and referrals fed new records back into it. The evaluation found 36% of the flags erroneous or based on irrelevant incidents; reporting raised the question the pilot could not close — whether such an entry stays in the system, uncorrected, to score the next child. In system-map terms the leverage is not on the model-to-operator pathway but on the store-to-model loop: what data may feed the score, how long it is kept, and whether a wrong record is ever corrected before it re-enters. The children's-rights critique names what the accuracy debate obscured: a tool can be "valid" at predicting future police contact — the successor's AUC around 0.83 — while profiling victims as potential perpetrators and treating a child's whole recorded life as evidence. And it ended not in court but in quiet disuse. The governance question ProKid poses is the one accuracy metrics never answer: not "is the score right?" but "should this record have been built at all?"
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.