Education AI
AI in education spans early-warning prediction, remote-exam proctoring, tutoring assistants, and automated grading, and its subjects are students — often minors — so the deploying institution's duty of care and the training of the staff who read the AI's output are load-bearing. Two deployment patterns anchor the governance. In predictive early warning, a model labels a student's risk of not graduating and delivers it to school staff, and the label is only as good as the intervention it triggers and the training of the human who reads it: a statewide dropout early-warning system was found by an independent audit to be wrong most of the time when it predicted a student would not graduate, to produce higher false-alarm rates for Black and Hispanic students, and to reach staff who reported no training on how to interpret a 'high risk' label — so the model's error and its group disparity were imported into how students were seen rather than into help they received, and the system was withdrawn. Against that, a large district's transparent, low-tech on-track indicator, paired with real intervention, accompanied a rise in graduation to a record level — locating the benefit in the intervention the indicator makes legible, not in the sophistication of the prediction. In remote proctoring, surveillance-based AI carries a rights cost that can be independently adjudicated: a public university's requirement that students pan their webcam around their home before an exam was held to be an unreasonable search, and peer-reviewed measurement found proctoring software flags darker-skinned and Black students more often with no corresponding difference in actual cheating. The Lab networks model only the deploying institution — its model or proctoring tool, the staff and proctors who read its output, and its student records; the students being scored or watched sit outside the dynamics, and no student outcome is computed on any diagram.
Use cases
What AI is doing here
Dropout & on-track early warning
PredictiveModels that flag a student's risk of not graduating and deliver the label to school staff — where the label is only as good as the intervention it triggers and the training of the human who reads it, and a model wrong most of the time and less accurate by race imports its error and disparity into how students are seen.
Remote-exam proctoring
PredictiveSurveillance-based AI that watches students during remote exams — webcam room scans, behavior flags — where the surveillance carries a rights cost (a pre-exam room scan was held an unreasonable search) and peer-reviewed measurement finds more flags for darker-skinned and Black students with no more actual cheating.
Tutoring & grading copilots
GenerativeGenerative assistants that tutor students or draft grades and feedback — where the benefit depends on the teacher or tutor who reviews the output, and AI-text detectors used to police the same tools are documented to be biased against non-native English writers.
Case files
What has gone wrong and right
Documented deployments, presented as model organizations calibrated to the evidence, with full citations.
A risk label most students never needed
United States (a state department of public instruction's statewide K-12 early-warning deployment)Wisconsin's statewide Dropout Early Warning System (DEWS) used ensemble (several models combined into one score) machine learning to label every grade 6 to 9 student's risk of not graduating and delivered the label to school staff through dashboards for about a decade. An independent, decade-scale audit found it was wrong roughly 74 percent of the time when it predicted a student would not graduate, produced higher false-alarm rates for Black and Hispanic students, and that the deployer's own internal equity research had gone unpublished — while a survey of districts found administrators reporting no training on how to interpret a 'high risk' label. The state stopped publishing the dashboards in 2023. Against it stands a large district's transparent, low-tech on-track indicator, paired with real intervention, which accompanied a rise in graduation to a record level — locating the benefit in the intervention the indicator makes legible, not in the sophistication of the prediction.
Explore this deployment in the PAN Lab →A search of the home and a suspicion by group
United States (a public university's remote-proctoring requirement; a federal Fourth Amendment ruling)Cleveland State University required students to pan their webcam around their home before an online exam, using remote-proctoring software that flags suspected cheating from the video. A federal court held the pre-exam room scan was an unreasonable search under the Fourth Amendment — a first-of-its-kind ruling that a routine proctoring practice violated a student's constitutional rights in their own home. Separately, peer-reviewed measurement of automated proctoring found more face-detection failures, more red flags, and higher priority scores for darker-skinned and Black students, with no corresponding difference in actual cheating. The case is the education domain's clearest portrait of surveillance-based integrity AI whose costs are each independently established: a rights violation a court can rule on, and a demographic burden of suspicion an audit can measure.
Explore this deployment in the PAN Lab →System map
Who is in the system and what pushes on it
Who is in the system
- Frontline workers. Caseworkers, screeners, eligibility staff — the operator network whose judgment the system augments or erodes.
- Supervisors & QA. The institutional correction layer: overrides, second reads, quality review.
- Agency leadership. Owns procurement, policy, and the authority map; answers for the system publicly.
- Served people & families. Those the decisions land on. Deliberately outside the PAN dynamics — their outcomes are measured, never simulated.
- Regulators & oversight bodies. Boards, auditors, data-protection officers, inspectorates — external correction capacity.
- Advocates & community organizations. Surface harms institutions do not see; historically the earliest accurate signal.
Dominant pressures
- Reviewer bottleneck. One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Austerity & recovery incentives. Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Vendor opacity. The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift. The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance. Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
Governance
Questions leaders should be asking
- 1. A dropout-risk label reaches a teacher who acts on it — so if the model is wrong most of the time and less accurate for some groups, and staff get no training on interpreting a 'high risk' flag, is the label importing the model's error and disparity into how students are seen rather than into help they receive?
- 2. A transparent, low-tech on-track indicator paired with real intervention accompanied a record graduation rate, while an opaque statewide ML system was withdrawn — so is the benefit being sought in the sophistication of the prediction, or in the intervention the indicator makes legible enough for staff to act on well?
- 3. A public university's pre-exam webcam room scan was held to be an unreasonable search — so is a proctoring deployment weighing the rights cost of surveillance in a student's home against the integrity problem it is trying to solve, or treating the surveillance as a free default?
- 4. Proctoring software flags darker-skinned and Black students more often with no corresponding difference in actual cheating — so is anyone measuring the flag rate by group before the flags become accusations a student has to answer?
For the actions behind these questions, see the Practice Library.
Seeing your organization in this domain? Mapping its actual pathways, pressures, and correction capacity is engagement work.
Work With 1023AI