10 to the 23 AI logo

Domain Atlas

Benefits navigation & public-facing chat

Conversational systems standing between the public and their benefits — public-facing chatbots, adviser-gated copilots, federated whole-of-government fleets, and navigation intermediaries whose quiet removal is itself the harm. An authoritative wrong answer is indistinguishable, to its victim, from policy; what sets the exposure is whether a professional gates the answer, whether anyone measures accuracy at all rather than mere deflection, and whether the whole channel rests on a single actor who can switch it off. The Lab networks in this domain model only what happens inside the operating organization — its operators, engines, and knowledge stores; the members of the public asking the questions sit outside the dynamics, and harm to them is documented in each case file, never computed on a diagram.

Use cases

What AI is doing here

Benefits-navigation chatbots

Generative

Conversational guidance about eligibility and process, for applicants directly or for navigators assisting them.

Retrieval-grounded appeal drafting

Generative

RAG systems drafting determinations or appeal responses from policy corpora for adjudicator review.

Supervisor-gated adviser copilots

Generative

Retrieval-grounded assistants that draft benefits answers for a professional adviser — often behind a supervisor approval gate — so AI output passes a human check before it ever reaches the member of the public.

Digital navigation intermediaries

Generative

Third-party front doors — guided online intake, call centers, and recipient-side apps, increasingly with embedded AI help — that carry large shares of a program's applications outside the agency's own systems, making the intermediary node itself a point of dependency.

Whole-of-government chatbot fleets

Generative

Federated or centrally provisioned networks of public-sector chatbots sharing a common engine or knowledge store, so a single defect, content error, or fix propagates across many agencies' answers at once.

Case files

What has gone wrong, and right

Documented deployments, presented as model organizations calibrated to the evidence, with full citations.

System map

Who is in the system, and what pushes on it

Who is in the system

  • Frontline workers. Caseworkers, screeners, eligibility staff — the operator network whose judgment the system augments or erodes.
  • Agency leadership. Owns procurement, policy, and the authority map; answers for the system publicly.
  • Served people & families. Those the decisions land on. Deliberately outside the PAN dynamics — their outcomes are measured, never simulated.
  • Vendors. Build and update the systems; hold the information asymmetry procurement must govern.
  • Regulators & oversight bodies. Boards, auditors, data-protection officers, inspectorates — external correction capacity.
  • Advocates & community organizations. Surface harms institutions do not see; historically the earliest accurate signal.

Dominant pressures

  • Caseload surge. Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck. One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity. The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift. The world, the intake process, and the rules change under a system trained on how things used to be.
  • Compliance over substance. Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

Governance

Questions leaders should be asking

  1. 1. Is the system answering the public directly, or drafting for a professional who gates every answer before it reaches anyone? The same technology carries radically different exposure depending on which side of that line it sits.
  2. 2. What is the retrieval corpus, who curates it, and how quickly does it track a rule change — and now that many agency bots share one engine or one knowledge store, does a single stale page or content error surface in every answer at once?
  3. 3. How would a confident wrong answer be detected — by an accuracy measure the operator actually tracks, or only when a union, an auditor, or a journalist surfaces it? Many of these systems ship with no measured accuracy at all, only a containment or usage figure that counts deflection rather than correctness.
  4. 4. If this channel comes to carry most of the traffic, what happens to the people who depend on it when it is defunded, decommissioned, or wound down — and does an accountable public body hold that switch, or a board and a funding model no partner can see or veto?

For the actions behind these questions, see the Practice Library.

Seeing your organization in this domain? Mapping its actual pathways, pressures, and correction capacity is engagement work.

Work With 1023AI