Benefits navigation & public-facing chat
Conversational systems standing between the public and their benefits — public-facing chatbots, adviser-gated copilots, federated whole-of-government fleets, and navigation intermediaries whose quiet removal is itself the harm. An authoritative wrong answer is indistinguishable, to its victim, from policy; what sets the exposure is whether a professional gates the answer, whether anyone measures accuracy at all rather than mere deflection, and whether the whole channel rests on a single actor who can switch it off. The Lab networks in this domain model only what happens inside the operating organization — its operators, engines, and knowledge stores; the members of the public asking the questions sit outside the dynamics, and harm to them is documented in each case file, never computed on a diagram.
Use cases
What AI is doing here
Benefits-navigation chatbots
GenerativeConversational guidance about eligibility and process, for applicants directly or for navigators assisting them.
Retrieval-grounded appeal drafting
GenerativeRAG systems drafting determinations or appeal responses from policy corpora for adjudicator review.
Supervisor-gated adviser copilots
GenerativeRetrieval-grounded assistants that draft benefits answers for a professional adviser — often behind a supervisor approval gate — so AI output passes a human check before it ever reaches the member of the public.
Digital navigation intermediaries
GenerativeThird-party front doors — guided online intake, call centers, and recipient-side apps, increasingly with embedded AI help — that carry large shares of a program's applications outside the agency's own systems, making the intermediary node itself a point of dependency.
Whole-of-government chatbot fleets
GenerativeFederated or centrally provisioned networks of public-sector chatbots sharing a common engine or knowledge store, so a single defect, content error, or fix propagates across many agencies' answers at once.
Case files
What has gone wrong, and right
Documented deployments, presented as model organizations calibrated to the evidence, with full citations.
Caddy adviser copilot at Citizens Advice
England and Wales, United Kingdom (Citizens Advice network; Greater Manchester pilot; scaled with the UK Government Incubator for AI)A benefits-advice copilot that never talks to the public: every AI draft is routed to a separate supervisor to approve, edit, or reject before it reaches even the adviser — the domain's positive counterexample, carrying a developer-run RCT and a live question about whether the every-message gate survives scaling to a whole national network.
Stress-test this shape in the PAN Lab →GOV.UK Chat
United Kingdom (national)Britain's biggest public test of generative AI in government: over roughly 18 months and two gated pilots, more than 10,000 people asked a retrieval-grounded assistant 26,000 questions about tax, benefits and visas — and GDS published the whole learning curve, from a 2023 version it held back for missing its accuracy bar to a self-reported 76 percent and then 90 percent measured accuracy, then launched it in the GOV.UK app in May 2026. The staged-gate counterexample to public chatbots that skipped the gates — with the honest catch that the operator grades its own accuracy.
Stress-test this shape in the PAN Lab →Frida (NAV Norway)
Norway (national)The most heavily studied welfare-agency chatbot anywhere: Frida is the anonymous front door to Norway's national labour-and-welfare agency, and its payload is the moment a citizen can ask for a human. About four in five conversations end without one — but that is a completion rate, not an accuracy rate, and the escalation itself is a dial the agency turns, not the citizen: about one in five conversations reached a person under free choice, and about 30 percent when NAV removed the explicit human-chat option. The conversations that cross the boundary arrive degraded, and no per-answer accuracy rate has ever been published. The handover-boundary counterexample: the governance question is who sets how easy it is to reach a person, and whether anyone reads the conversations that ended without one.
Stress-test this shape in the PAN Lab →Burokratt
Estonia (national); Information System Authority (RIA) under the Ministry of Justice and Digital Affairs; deployed across public-sector institutions including the Tax and Customs Board, the Police and Border Guard Board, the Health Insurance Fund, Statistics Estonia, the National Library, and the municipality of Rae ParishEstonia's national network of public-sector chatbots — a federated 'chatbot of chatbots' in which each institution runs its own assistant, a central classifier routes between them, and one shared knowledge module feeds them all — and the atlas's clearest real instance of a governance question that is about the shared link, not the node.
Stress-test this shape in the PAN Lab →Singapore's chatbot fleet refresh: eighty scripted engines retired for a shared LLM platform
Singapore (national; whole-of-government, 60-plus agencies), operated by GovTech with individual agencies as content operators and the Ministry of Finance authoring the SupportGoWhere calculator rulesA whole-of-government retire-and-replace: Singapore decommissioned Ask Jamie, the decade-old scripted assistant embedded on 70-plus agency websites as independent per-agency answer engines, and migrated government chatbots onto a small number of centrally provided large-language-model engines - a single decision that swapped roughly 80 uncorrelated error sources for one shared generative layer now fielding, per the government's own figures, over 800,000 citizen queries a month across 60-plus agencies. The domain's clean natural experiment in error-correlation structure, and its retire-and-replace lifecycle case.
Stress-test this shape in the PAN Lab →IRS collection chatbots: expanded and made permanent with no performance measures
United States (federal); Internal Revenue Service, Small Business/Self-Employed Division; Automated Collection System chat applications audited at the Brookhaven (Holtsville, NY) and Philadelphia campuses and the Office of Online Services (Lanham, MD)A federal collection agency deployed a scripted chatbot, human live chat, and voice bots to deflect balance-due inquiries off its phone line, then expanded them and made live chat permanent — while, an inspector-general audit found, having no performance measures for the program at all and telemetry so unreliable it reported one assistor handling 603 chats at once against a system cap of three: the domain's monitoring-absence case, where the danger is not a wrong decision but an unmeasured one, and the only functioning oversight is an external, episodic audit.
Stress-test this shape in the PAN Lab →Albert France Services
France (national DINUM/ANCT pilot; about 48 France Services counters across six departments, launched at Sceaux)France's flagship sovereign AI assistant for one-stop benefits counters gave the Prime Minister a wrong answer at its own launch, was quietly shelved within about eighteen months, and was formally denied generalization in January 2026 in its current form. Its distinctive feature is an absence: no error rate, usage or override figure was ever published, so the error-detection channel was the workforce and the unions, not a metric, and a cost-accounted review now gates its successor.
Stress-test this shape in the PAN Lab →Propel in-app SNAP benefits assistant
United States (nationwide consumer app; California for the CalFresh triage-flow test)A consumer benefits app grounds its AI help on a state-verified deposit record it reads but never writes to, and escalates every AI dead-end to a human — a read-only design that severs the memory loop, on the strength of vendor-published pilot results only.
Stress-test this shape in the PAN Lab →GetCalFresh: the nonprofit front door that carried most of California's online SNAP intake
California, USA (statewide, all 58 counties from May 2019; originated 2014 as a Code for America pilot with the San Francisco County Human Services Agency)A nonprofit-built web form - not a government system, not a scoring model - that came to carry more than 70% of California's online SNAP applications while making no eligibility decisions at all: the domain's positive anchor for a critical navigation node whose safety came from a statutory determination floor it never touched, and whose single point of dependency was resolved by a planned, dated transfer into a state-owned portal rather than an abrupt collapse.
Stress-test this shape in the PAN Lab →Benefits Data Trust wind-down
United States (headquartered Philadelphia, Pennsylvania; state and city contracts across several states)A twenty-year benefits-navigation nonprofit dissolved by its own board in 60 days — a node deletion that orphaned data-sharing links and left referral partners with nowhere to hand off.
Stress-test this shape in the PAN Lab →System map
Who is in the system, and what pushes on it
Who is in the system
- Frontline workers. Caseworkers, screeners, eligibility staff — the operator network whose judgment the system augments or erodes.
- Agency leadership. Owns procurement, policy, and the authority map; answers for the system publicly.
- Served people & families. Those the decisions land on. Deliberately outside the PAN dynamics — their outcomes are measured, never simulated.
- Vendors. Build and update the systems; hold the information asymmetry procurement must govern.
- Regulators & oversight bodies. Boards, auditors, data-protection officers, inspectorates — external correction capacity.
- Advocates & community organizations. Surface harms institutions do not see; historically the earliest accurate signal.
Dominant pressures
- Caseload surge. Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck. One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity. The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift. The world, the intake process, and the rules change under a system trained on how things used to be.
- Compliance over substance. Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
Governance
Questions leaders should be asking
- 1. Is the system answering the public directly, or drafting for a professional who gates every answer before it reaches anyone? The same technology carries radically different exposure depending on which side of that line it sits.
- 2. What is the retrieval corpus, who curates it, and how quickly does it track a rule change — and now that many agency bots share one engine or one knowledge store, does a single stale page or content error surface in every answer at once?
- 3. How would a confident wrong answer be detected — by an accuracy measure the operator actually tracks, or only when a union, an auditor, or a journalist surfaces it? Many of these systems ship with no measured accuracy at all, only a containment or usage figure that counts deflection rather than correctness.
- 4. If this channel comes to carry most of the traffic, what happens to the people who depend on it when it is defunded, decommissioned, or wound down — and does an accountable public body hold that switch, or a board and a funding model no partner can see or veto?
For the actions behind these questions, see the Practice Library.
Seeing your organization in this domain? Mapping its actual pathways, pressures, and correction capacity is engagement work.
Work With 1023AI