A risk label most students never needed
Explore this deployment in the PAN Lab ↗
A state's statewide dropout early-warning system used ensemble machine learning to label every grade 6 to 9 student's risk of not graduating on time and delivered the label to school staff through dashboards for about a decade. An independent, decade-scale audit found the system was wrong roughly 74 percent of the time when it predicted a student would not graduate, produced higher false-alarm rates for Black and Hispanic students, and that the deployer's own internal equity research had gone unpublished — while a survey of districts found administrators reporting no training on how to interpret a 'high risk' label. The state stopped publishing the dashboards in 2023 and said it was evaluating the system's future. The deployment is the education domain's clearest case of a risk label whose error and group disparity entered how students were seen rather than the help they received.[3]
What happened
The Wisconsin Department of Public Instruction built and ran a statewide Dropout Early Warning System (DEWS): an ensemble machine-learning model that scored every student in grades 6 through 9 for their risk of not graduating on time, and delivered that risk label to school staff through a dashboard. It ran for roughly a decade. The system was built in earnest — its own designer published a detailed account of the methodology and the goal of catching at-risk students early enough to help them — and it is exactly the kind of predictive tool that sounds unarguable: find the students who need help before it is too late.
An independent, decade-scale audit is what complicates the picture, and its findings are the case's content. The system was wrong about 74 percent of the time when it predicted that a student would not graduate — that is, the large majority of students it labeled high-risk went on to graduate anyway. Its false-alarm rate was higher for Black and Hispanic students than for white students, so the burden of the wrong labels fell unevenly by race. The department had conducted internal equity research but had not published it. And a survey of districts found administrators reporting that they had received no training on how to interpret a "high risk" label — no guidance on what it did or did not mean, or what to do about it.
Put those together and the failure is specific. A risk label is not an intervention; it is a signal that reaches a human who then does something, or nothing, with it. When the signal is wrong most of the time, arrives with a group disparity, and lands on staff who were never taught how to read it, the most likely thing it changes is not the help a student receives but the way the student is seen — a "high risk" tag that lowers expectations or hardens attention, applied disproportionately to Black and Hispanic students, and wrong far more often than right. The model's error and its disparity did not stay in the model; they were imported into the perception of students, which is the one place a school most needs to get right. The state stopped publishing the dashboards in 2023 and said it was evaluating the system's future.
The counter-case is what keeps this from being a story about prediction being useless. A large urban district used a transparent, low-tech "on-track" indicator — built on interpretable research about which early signals actually predict graduation, and paired with a real program of intervention for the students it flagged — and its graduation rate rose to a record level over the years it was used. The difference is not that one district had a better model; it is that the on-track indicator was legible enough for staff to understand and act on, and it was wired to an intervention rather than to a dashboard. The benefit lived in what the indicator made educators do, not in the sophistication of the prediction.
The honest reading is that an early-warning model is only as good as the intervention it triggers and the training of the human who reads it, and that an opaque model that merely labels — especially one wrong most of the time and unevenly by race, handed to untrained staff — can do harm precisely where it was meant to help. The governable variables are the accuracy-by-group check the department left unpublished, the training that would let staff read a label correctly, and the intervention that would turn a flag into help rather than a lens.
The sociotechnical reading
This case is the education domain's early-warning anchor, and it makes a point the map cares about across every predictive-flag domain: a risk label is not a decision or an intervention — it is a signal handed to a human, and its value is set by what the human is trained to do with it and what help it is wired to. Here the signal was wrong about 74 percent of the time on its positive predictions, disparate by race in its false alarms, and delivered to staff who reported no training on interpreting it. So the most likely effect was not help delivered but perception changed: a "high risk" tag that shifts how a student is seen, applied more often and more wrongly to Black and Hispanic students. The model's error and disparity were imported into perception, which is the worst place for them to land.
The two governable surfaces are drawn as the latent checks. The first is the accuracy-by-group check the department ran internally and did not publish — the equity audit that would have surfaced the false-alarm disparity and the overall wrongness as facts to act on rather than an outside journalist's finding a decade in. The second is the training-and-intervention loop: a label reaching untrained staff, wired to a dashboard rather than to a resourced response, is a lens; the same label reaching trained staff, wired to an intervention, is help. The map reads the absence of both as the reason a well-intentioned predictive tool became a mechanism for mislabeling students by race.
The counter-case sharpens the instruction rather than softening it. A transparent, low-tech on-track indicator, interpretable by design and paired with real intervention, accompanied a record graduation rate elsewhere — so the lesson is not "prediction does not work in schools" but "the benefit lives in the intervention the indicator makes legible, not in the sophistication of the model." An interpretable indicator wired to help outperforms an opaque model wired to a dashboard, because staff can understand the first and act on it well, and the second only labels. Interpretability here is not a nicety; it is what lets the human loop turn a signal into help.
The Lab network models only the deploying institution: its early-warning model, the school staff who read its labels, and its student records. No student outcome is computed on any diagram. The students being scored are boundary-only; the audit's accuracy and disparity findings, the missing training, the unpublished equity research, and the interpretable-indicator counter-case are institutional signals that live in this case file, never on any network. The audit figures are the independent investigation's own analysis, and the disparity is drawn as a recorded external finding, never a computed harm. The map's instruction is to treat an early-warning label as valuable only through the training and intervention behind it, to make the accuracy-by-group check a published standing obligation, and to prefer an interpretable indicator wired to help over an opaque model that hands untrained staff a lens on the students it most often mislabels.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.