Domain Atlas / Benefits navigation & public-facing chat
Caddy adviser copilot at Citizens Advice
Work with this case in the PAN Lab ↗
In a developer-reported randomised controlled trial of more than 1,000 adviser support requests, an adviser-facing benefits copilot at Citizens Advice returned supervisor-checked answers in about four minutes, roughly half the previous response time, with about 80 percent of its drafts approved by supervisors without revision; these figures are reported by the tool's builders and have not been independently replicated.[3]
What happened
Caddy is an adviser-facing drafting copilot used inside Citizens Advice contact centres in England and Wales, built in-house by Citizens Advice Stockport, Oldham, Rochdale & Trafford (CASORT) and scaled with the UK Government's Incubator for AI (i.AI, Cabinet Office and later the Department for Science, Innovation and Technology). Its defining design decision is what it refuses to be: the team decided any AI solution "cannot be client-facing," to avoid putting vulnerable, upset, or confused clients into an automated "chatbot loop of hell," and instead inserted the tool on the internal adviser-to-supervisor escalation link — the flow they identified as the organisation's biggest pain point after remote work made timely supervisor answers hard to get. CASORT had trialled and abandoned an earlier chatbot in 2019; it built the generative prototype in about four weeks in 2023, then hardened it over roughly twelve months of volunteer-driven work, and i.AI announced the co-developed tool on 28 March 2024, beginning at two Greater Manchester centres, with the code released open-source under an MIT licence (the original i-dot-ai/caddy-chatbot repository, last updated in November 2024, was archived read-only on 21 October 2025; no public 2.0 successor repository has been located).
The workflow is a live verification policy. When a frontline adviser — the staff are heavily trainees and volunteers — asks Caddy a benefits question, the system retrieves from a closed corpus of two trusted source families, GOV.UK and Citizens Advice's internal knowledge base (AdviserNet), using retrieval-augmented generation and semantic routing tuned with experienced advisers, and drafts a short cited answer (about 400 words) at a lowered generation temperature. The draft is not shown to the adviser. It is routed to an experienced supervisor whose approve, edit, or reject decision gates delivery back to the adviser, who then relays the answer to the client in their own words. The client only ever talks to the human adviser, and Caddy makes no eligibility or entitlement determination; CASORT frames the supervisor validation as what keeps the tool compliant with advice-accreditation standards, since Caddy itself never advises.
i.AI built evaluation in from the start as what it describes as a randomised controlled trial: over roughly three months, more than 1,000 adviser support requests were invisibly randomly assigned either to Caddy or to the pre-Caddy human-supervisor pathway, with in-chat surveys after every call in both arms. The developers reported (i.AI blog, March 2025) that about 80% of Caddy's drafts were of high enough quality to pass to advisers without revision; that messages were generated, validated, and returned in about four minutes, roughly half the roughly ten-minute baseline (with up to 60% of supervisor time saved per query); and that advisers with Caddy access were more than twice as likely to report confidence giving advice and more than 1.5 times as likely to report resolving the client's issue than the control group (the DSIT Knowledge Hub entry states these last two as "twice" and "1.5 times" without the "more than"). The RCT also found that new and trainee advisers felt more comfortable asking extra questions of Caddy than they would of a busy human supervisor — a change in the escalation channel's social dynamics, not just its latency. Every one of these figures is reported by the builders, i.AI and CASORT, in blog and conference form; no peer-reviewed trial report, preregistration, or raw data was located, i.AI itself wrote that the trial "wasn't without issues," and an independent Stanford Legal Design Lab write-up describes the six-office evaluation as a four-to-six-week pilot rather than detailing the randomisation — so the RCT is developer-design-confirmed but not independently audited, the confidence and resolution numbers are adviser self-reports rather than client-outcome or ground-truth-accuracy measures, and no absolute post-supervision error rate is published.
The 20% of drafts not approved unchanged were decomposed into two buckets — the adviser's prompt lacked information, or the retrieval corpus lacked coverage — each feeding a distinct repair loop (an agentic prompt-support assistant for the former, negotiated corpus extension, for example with the Child Poverty Action Group, for the latter), which the team expects to raise the approval rate. Governance around the pilot included a citizen-participation review by Manchester Metropolitan University's People's Panel for AI (with Manchester City Council), which the team reported gave "a ringing endorsement," plus consequence scanning. By 2026 i.AI listed Caddy as Live: supporting more than 40 local offices with a self-reported 90% adviser approval rating, over 70 offices on a waiting list, 100+ signed up during the open beta, and an ambition of full rollout across all roughly 250 Citizens Advice offices in England and Wales (a network that reported helping 2.5 million-plus people one-to-one in 2022-23; a Stanford account gives materially different figures of 270 local organisations across 2,540 locations advising 2.8 million people in 2024, so the "250 offices" ambition is i.AI's own framing rather than a measured full footprint). Office counts vary by date and should be read time-indexed, not blended, and the office and approval figures on the 2026 page are unaudited developer claims. Caddy 2.0, announced on 4 November 2025 with an open beta the following month, adds an automated verification engine that deconstructs answers into verifiable claims cross-referenced against the trusted sources "to streamline the supervisor checks," learns from responses that supervisors reject, and includes an automated tool that strips personal information from the adviser's query; national rollout was aimed for 2026. i.AI also announced, in March 2025, an intention to extend Caddy into UK government-department frontline pilots (third-party write-ups name immigration and tax); no fetched 2026 source confirms those pilots are ongoing. This case is distinct from the Atlas's earlier US benefits-navigation copilots (Nava's assistive chatbot and the Imagine LA Benefit Navigator), which place verify-before-use inside the same caseworker who uses the answer: Caddy is UK, routes every draft to a separate supervisor before the adviser sees it, and carries a developer-designed RCT evidence base.
The sociotechnical reading
The Atlas already carries two careful benefits copilots — Nava and the Imagine LA Benefit Navigator — and both locate the safeguard in the same place: a habit inside the person using the answer, who is supposed to read the citation before relaying it. Caddy does something structurally different, and that difference is the lesson. It routes every draft to a separate person — a supervisor whose whole job in that moment is to check — before the answer reaches even the adviser, let alone the client. Verification here is a role, not a habit. On the system map that is a specific shape: the model does not sit on the operator's own adoption edge at all; its output lands first on a dedicated operator-to-operator check that is drawn strong at baseline, and only the approved answer is handed down. Moving the check off the person under time pressure and onto a second party is a genuine structural gain over an individual habit that erodes quietly under load — which is exactly why this is the domain's governed counterexample, and why the RCT could report advisers' confidence doubling and trainees asking the tool more freely than they would a busy supervisor. The floor rose because the checking was someone else's dedicated job.
But externalizing verification onto a role relocates the failure mode rather than removing it. The question is no longer "will the individual keep checking?" — the Nava and Imagine LA question — but "will the institution keep requiring that every message pass the gate?" And that is a different kind of parameter. A habit decays in a thousand private, unlogged moments; a gate is a written policy, and it erodes in the open, in the language an organisation uses about itself. The signs are already legible in this case's own record: the 2.0 verification engine exists explicitly to "streamline the supervisor checks," and the 2026 programme page describes routing responses "for human checks when needed" — a quiet drift from the original every-message gate as volume climbs toward a whole national network. Each step is individually defensible; together they are the difference between a gate every draft passes and a gate that samples. So the governance instrument that matters is not a lever aimed at any one adviser's diligence but the verification policy itself — an institutional setting, measurable by the share of drafts approved without revision and, more importantly, the share of messages actually gated. That is the productive corner: keep the deference channel from silently widening (a standing vigilant check), keep the confident, plain-language draft legible as generated-versus-cited so the supervisor weighs it for what it is (provenance labeling), protect the review capacity the gate depends on rather than letting surge collapse it into rubber-stamping (deskilling arrest), and authorize each proposed automation of the check one grant at a time, on the record, rather than letting "streamline" quietly become "skip." The distinct lesson the Atlas draws here: a dedicated verifier is a real upgrade over a personal habit, because it trades a control that erodes invisibly for one that erodes in writing — and a control that erodes in writing is the only kind you can actually govern, provided someone is reading the policy as closely as the supervisor is meant to read the draft. The honest boundary throughout: served people, and the advice they do or do not ultimately receive, are not modeled here; the Lab reads institutional propagation only, and every headline number in this case is developer-reported and not independently replicated.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.