Domain Atlas / Child welfare & family services
Oregon Safety at Screening
Work with this case in the PAN Lab ↗
Oregon's child-welfare agency dropped its AFST-derived Safety at Screening tool in 2022, citing equity concerns amid national scrutiny of racial disparity in child-welfare algorithms.[2]
What happened
Oregon's child-welfare agency built its Safety at Screening tool in-house through ORRAI, its analytics office, deriving it from the Allegheny approach but deliberately narrowing it: the model drew only on the state's own child-welfare administrative records — no call text or voice — and used a dual-outcome design to give hotline screeners a four-tier risk score, shown only after they had finished entering a report. Oregon also engineered in a post-processing fairness correction, using group-specific thresholds under an "error rate balance" criterion to balance error rates across race and ethnicity groups, and its 2019 report described automation-bias mitigations layered on top. In 2022 the agency stopped using the tool, telling staff it wanted to reduce disparities and pursue equity, and replaced it with a non-algorithmic Structured Decision Making process — a department spokesman also noting the aggregate-data risk score simply could not be folded into SDM's family-specific method. The move came weeks after Associated Press reporting that the parent Allegheny tool had flagged a disproportionate share of Black children for "mandatory" investigation, and amid a racial-bias inquiry from a U.S. senator; Oregon's own tool was never publicly shown to reproduce those disparities.
The sociotechnical reading
The Atlas's most heavily documented harms ended with courts or commissions. Oregon is the counter-example: an internal actor exercising the authority to stop, before a documented crisis forced it — a far cheaper control than the litigation that ended systems on either side of it here. Two details make it unusually instructive. The agency built its fairness fix INTO a deployed model — group-specific thresholds meant to counteract, not sever, the loop in which a score trained on the system's own administrative history helps write the next record it will train on. And its stated reasons for stopping sat in tension: an equity rationale beside a mechanical "the score doesn't fit the replacement process" one — a reminder that the record of WHY a control was exercised is itself governance data. Whatever one concludes about the tool, the governance event is the point: discontinuation was available, exercised, and rehearsable — and much of its value came before the stop, in a design that kept screener discretion real.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.