Caseworker documentation & copilots
Generative and rules-based assistants that transcribe, summarize, triage, and increasingly draft the records and decisions institutions run on — meeting scribes, mail-triage filters, evidence summarizers, ruling drafters, and automation that removes the caseworker from the routine path entirely. The human review step is the load-bearing control, and the throughput that justifies the tool is the same pressure that erodes it — with a verifier that checks human work at one edge of the range and deployments halted after the harm was measured at the other.
Use cases
What AI is doing here
Generative case documentation
GenerativeTranscription and drafting of assessments and case notes, reviewed by workers before entering the record.
General staff copilots
GenerativeOffice assistants (drafting, summarizing, search) used informally across casework and administration.
Evidence summarization for decision-makers
GenerativeGenerative compression of case evidence — interview transcripts, policy or country-guidance notes, medical and administrative records — into the summary a caseworker reads before deciding, standing in for the underlying reading it replaces.
Correspondence & document triage
PredictiveLanguage-model classification of inbound letters, forms, and evidence to flag vulnerability or route work, reordering the queue ahead of a human without itself touching the entitlement decision.
Generative decision drafting
GenerativeAI drafting of the decision itself — proposed orders, determinations, or notification letters — for a human to review and adopt, where the machine draft becomes the record once accepted.
Rules-based decision automation & verification
PredictiveRules and NLP systems operating directly on the eligibility or benefit decision — automating recurring determinations end-to-end with a human exception path, or flagging a human-drafted decision against enumerated quality rules before it issues.
Case files
What has gone wrong, and right
Documented deployments, presented as model organizations calibrated to the evidence, with full citations.
Magic Notes (Beam)
United Kingdom (local authorities)A generative documentation assistant spreading through UK social care — the live test of human-in-the-loop at scale.
Stress-test this shape in the PAN Lab →Minute / Local Transcribe
United Kingdom (England; cross-council cohort)A state-built AI meeting scribe piloted through a governed council cohort and scaled nationally - an exemplary published governance record sitting on top of no published accuracy evaluation.
Stress-test this shape in the PAN Lab →Justice Transcribe
England and Wales, United Kingdom (HM Prison and Probation Service / Ministry of Justice)An in-house probation copilot that transcribes supervision sessions into the case record - scaled to every probation officer in England and Wales before any accuracy evaluation was published, and sitting at the head of a pipeline whose downstream node scores reoffending risk from the records it helps write.
Stress-test this shape in the PAN Lab →GDS Microsoft 365 Copilot cross-government experiment
United Kingdom (cross-government)The largest published government copilot evaluation anywhere — about 20,000 UK civil servants across 12 organisations — found a self-reported 26 minutes saved per working day and most users unwilling to go back, while a companion departmental evaluation found no robust evidence the time savings improved productivity and about a fifth of its users caught the tool hallucinating.
Stress-test this shape in the PAN Lab →DWP Whitemail Insights and Vulnerability Scanner
United Kingdom (the transparency record lists the region as England and Wales; DWP correspondence operations cover Great Britain)A language-model scanner reads roughly 25,000 paper letters to the UK welfare agency every day and decides which claimants surface on the potentially-vulnerable shortlist. Its transparency record names precision, recall and F1-score as its metrics but discloses no values, and the claimants were never told it exists — the impact assessment said letter writers do not need to know.
Stress-test this shape in the PAN Lab →UK Home Office asylum AI copilots: interview summarisation and policy search
United Kingdom (Home Office; national asylum casework)The Home Office piloted and then rolled out two AI copilots for asylum casework — one that compresses the interview transcript into a summary and one that summarises country policy. In the Home Office's own pilot the summariser saved a measured 23 minutes a case, 9% of its summaries were removed as inaccurate before any caseworker saw them, and 23% of users were not fully confident in the rest; the summaries carried no source references, no requirement to verify them exists, and claimants are never told AI touched their case.
Stress-test this shape in the PAN Lab →Learned Hand AI clerk pilot (LA and Riverside courts)
Los Angeles County and Riverside County, California, USAThe largest trial court in the United States is piloting an AI clerk that drafts rulings in each judge's own writing style, under a contract with a roadmap into criminal, family, and probate cases — and because California's disclosure rule is triggered only by documents written entirely by AI, the litigants whose cases it touches are never told.
Stress-test this shape in the PAN Lab →SSA Insight
United States (federal); Social Security Administration, Office of Hearings Operations and Office of Appellate Operations; nationwideA federal disability-adjudication copilot that runs the other way — it verifies human-drafted decisions against 43 automated quality flags instead of drafting them — and the atlas's clearest example of a governed, precision-gated checker, along with the one failure mode a checker has.
Stress-test this shape in the PAN Lab →VA claims automation (automated survivor-benefit decisions)
United States (federal, nationwide; Department of Veterans Affairs, Veterans Benefits Administration)Rules-based automation grants veterans' survivor-benefit claims end to end with no human when the rules match — and a federal watchdog on a multi-year clock was the only correction channel that ever changed anything.
Stress-test this shape in the PAN Lab →Trelleborg's Welfare Robot
Trelleborg Municipality (Skåne), Sweden; the model later spread to other Swedish municipalities including KungsbackaThe first Swedish municipality to fully automate social-assistance decisions — a rules-based robot that decides about one in three recurring reapplications with no human at all, and re-inserts a caseworker only when a single routing rule fires — and the atlas's clearest example of automation replacing discretion by design rather than overriding it in practice.
Stress-test this shape in the PAN Lab →Amsterdam Smart Check
City of Amsterdam, Netherlands (municipal welfare department, Werk Participatie en Inkomen; Handhaving enforcement and Income Services)The fairness-first welfare-fraud screener a city built with nearly every recommended safeguard — bias audits, debiasing, human-rights and data-protection assessments, a citizen panel, public transparency registers — and then discontinued itself after a live pilot showed the bias had re-emerged inverted, now wrongly flagging the majority group, women, and parents, with no gain over caseworkers. The atlas's governed-exit case.
Stress-test this shape in the PAN Lab →System map
Who is in the system, and what pushes on it
Who is in the system
- Frontline workers. Caseworkers, screeners, eligibility staff — the operator network whose judgment the system augments or erodes.
- Supervisors & QA. The institutional correction layer: overrides, second reads, quality review.
- Agency leadership. Owns procurement, policy, and the authority map; answers for the system publicly.
- Served people & families. Those the decisions land on. Deliberately outside the PAN dynamics — their outcomes are measured, never simulated.
- Vendors. Build and update the systems; hold the information asymmetry procurement must govern.
- Regulators & oversight bodies. Boards, auditors, data-protection officers, inspectorates — external correction capacity.
- Advocates & community organizations. Surface harms institutions do not see; historically the earliest accurate signal.
Dominant pressures
- Caseload surge. Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck. One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity. The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure. Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover. Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Compliance over substance. Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
Governance
Questions leaders should be asking
- 1. What fraction of AI-drafted or AI-triaged records is substantively changed before a human accepts it — and does anyone measure that edit rate, or is 'a human reviews everything' asserted without evidence it survives the workload?
- 2. How much of the work now reaches a record or a decision without a human reading it at all — and who set the rule that decides which cases take the no-human path?
- 3. When the affected person is never told a system touched their case, the operator's own vigilance is the only thing between an error and its consequence — is anyone auditing what that vigilance misses, and would a live bias or error surface before an outside body forced it to?
- 4. Are AI-written entries labeled where they land, so a later reader, an auditor, or a downstream risk tool knows the record was machine-drafted and not independently established — and if one shared tool writes everyone's records, who notices when a single error mode is repeating across all of them?
For the actions behind these questions, see the Practice Library.
Seeing your organization in this domain? Mapping its actual pathways, pressures, and correction capacity is engagement work.
Work With 1023AI