10 to the 23 AI logo

Domain Atlas / Benefits navigation & public-facing chat

Case fileUnited Kingdom (national)large deployment

GOV.UK Chat

Work with this case in the PAN Lab ↗

The UK Government Digital Service ran what it called the government's biggest public test of generative AI to date: across two gated public pilots (a late-2024 web pilot of 10,136 users asking 23,838 questions, and an autumn-2025 GOV.UK app pilot of 641 users asking 2,670 questions in four weeks), more than 10,000 people asked GOV.UK Chat about 26,000 questions on tax, benefits and visas. Its first 2023 version was held back in findings published January 18, 2024 because, GDS reported, answers did not reach the highest level of accuracy demanded for a site like GOV.UK, including a few cases of hallucination. GDS reports measured answer accuracy rising from 76 percent (its earliest benchmark) to 90 percent by the autumn 2025 pilot, assessed by subject-matter experts plus automated evaluation, an 88 percent answer rate for in-scope questions after a clarifying-questions feature was added, and that 508 attempts to jailbreak the system across the pilots were all prevented by its guardrails; it soft-launched to all GOV.UK app users on March 26, 2026 and officially launched on May 14, 2026. Nearly every one of these figures is self-reported by GDS, the system's operator, and the accuracy denominators and sampling frames are unpublished.[3]

What happened

GOV.UK Chat is a public-facing generative AI assistant built in-house by the Government Digital Service (GDS, part of the Department for Science, Innovation and Technology) to answer members of the public's questions about government services. It is a retrieval-augmented system: a user's free-text question is matched against a vectorised store of curated official GOV.UK guidance, and a commercial large language model hosted on a commercial cloud AI platform synthesises an answer only from the retrieved official content, with an explicit instruction to ignore its training data. It is an information-provision tool, not a decision system — GDS states it "does not attempt to provide advice" and "clearly signposts when users should check the original guidance" — and every answer links back to the GOV.UK source pages with a reminder to verify.

The distinctive feature is a staged pilot-gate cadence with a published, phase-by-phase record. GDS's first generative AI experiment, in 2023, put a retrieval-grounded chatbot over GOV.UK content through internal red-teaming phases, a roughly dozen-user test, then a live private experiment with 1,000 invited users; nearly 70% found the responses useful and just under 65% were satisfied (a small survey, n=157), but in findings published on January 18, 2024 GDS reported that "answers did not reach the highest level of accuracy demanded for a site like GOV.UK," including "a few cases of hallucination," and it explicitly flagged the sociotechnical risk that a user's misunderstanding of how generative AI works "could lead to users having misplaced confidence in a system that could be wrong some of the time." GDS framed its posture as "not moving fast and breaking things." In November 2024 it opened a waiting-list private beta on selected GOV.UK business pages, citing a consistent improvement in accuracy since 2023, with a five-step pipeline (question, GOV.UK retrieval, answer generation, safety check, answer with source links) grounded exclusively in GOV.UK content and guardrails including answer checking, jailbreak red-teaming, and subject-matter-expert scoring conducted with HMRC.

On October 7, 2025 GDS published a formal Algorithmic Transparency Record documenting the architecture: retrieval over roughly 700,000 vectorised chunks (36.9 GB, about 100,000 pages) of GOV.UK content; the instruction to ignore training data; regular-expression filtering that rejects queries containing phone numbers, email addresses or card numbers; source-page links plus a verify-your-answer reminder on every response; an admin system for monitoring answer quality and malicious use; LLM-as-a-Judge evaluation metrics covering factual precision, factual recall, relevancy and groundedness; jailbreak assessments run with the AI Security Institute (AISI) alongside the record's own concession that "it's not possible to guarantee no jailbreaking attempts will be successful"; and 12-month encrypted-at-rest retention of question data. In a pilot-completion post on March 16, 2026, GDS reported the pilot totals across its two public pilots — the first (late-2024 web pages) drew 10,136 users asking 23,838 questions, and the autumn-2025 GOV.UK app pilot drew 641 users asking 2,670 questions in four weeks, together more than 10,000 users and about 26,000 questions, which GDS called "the government's biggest public test of generative AI to date." It reported measured answer accuracy rising from 76% ("our earliest benchmark") to 90% by the app pilot, assessed by a combination of subject-matter experts and automated evaluation; an 88% answer rate for in-scope questions after a clarifying-questions feature was added; and that there were 508 attempts to jailbreak GOV.UK Chat during the two pilots, all of which its safety guardrails prevented. A follow-up survey of GOV.UK app users found 73% useful and 64% satisfied; the average response time was 10.7 seconds, satisfaction increased when faster answer speeds were simulated, and GDS noted the "latest versions of frontier models have been more powerful but slower," with accuracy deliberately prioritised over speed. Independent trade-press coverage (March 2026) reported the same 76-to-90% record and, per the transparency record, GDS's claim that for government-related questions GOV.UK Chat scores higher than widely-used consumer AI assistants and is "in line with industry benchmarks" — an operator self-comparison.

GOV.UK Chat soft-launched to all GOV.UK app users on March 26, 2026 and officially launched on May 14, 2026; by that point more than 7,800 people had used it, asking more than 15,000 questions since the soft launch, with strongest demand in tax, driving and transport, and benefits. GDS's December 2025 vision post made the content-dependency loop explicit — "GOV.UK Chat can only be as good as the content published on GOV.UK by departmental teams. Clear and accurate content will result in better answers" — positioned Chat as joining guidance across departments (for example a new parent asking "what help can I get?" spanning HMRC, DWP and the Department for Education), noted the tool was selected as a Prime Minister's AI Exemplar, and said the team is experimenting with agentic AI for future transactions; the GOV.UK app it rides in entered public beta in July 2025. Availability on the GOV.UK website, as opposed to the app, was still being explored as of mid-2026. Nearly every quantitative figure here — the 76-to-90% accuracy, the 88% answer rate, the 508 blocked jailbreak attempts, the satisfaction rates — is self-reported by GDS, the system's operator; there is no independent audit of the accuracy methodology, the accuracy denominators and sampling frames are unpublished, and the phase labels are inconsistent across documents (the transparency record describes a private beta capped at 2,000 users over four weeks, distinct from the 10,136-user first public pilot), so per-phase figures belong to their specific source. The specific hosted model stack is disclosed in the government's transparency record and deliberately omitted here.

The sociotechnical reading

The Atlas already carries the case this one answers. The NYC MyCity business chatbot is the same structural shape — a public-facing generative adviser, the public asking directly with no professional intermediary, error delivered in an official government voice — and its lesson was that exposure is not correction: the errors were published in a detailed investigation and again in a formal audit, and the owner kept the tool online and rejected all seven recommendations until an unrelated budget cut ended it. GOV.UK Chat is that same shape with the opposite governance move, and that is its whole point. Where MyCity was told it was wrong, in public, and kept answering anyway, GOV.UK Chat built a gate that could say "not yet" — and used it. The 2023 version was held back for not reaching the accuracy demanded of a site like GOV.UK, and rather than bury that, GDS published the whole learning curve, gate by gate: a held version, a private beta, a transparency record, two pilots with rising numbers, then a launch. The distinct lesson this case adds is that a gate is only a control if it can fire, and the strongest governance signal a public generative system can send is a version it declined to ship with the hold recorded next to the launch. That is a presence, not an absence — the rare Atlas case whose payload is a working accountability step rather than a missing one.

Two quieter properties reinforce it, and both cut against the copilots. First, the store coupling is protective rather than contaminating. The model is instructed to answer only from a curated corpus of official guidance and does not write its generated answers back into it, so — unlike the horizontal copilots whose drafts become memory the next query reads as fact — a wrong line here does not entrench. The one deliberate feedback loop runs the benign direction: GDS tells departments that Chat can only be as good as the content published on GOV.UK, so observed answer failures become pressure to fix the underlying official guidance, and every answer routes the user back to the authoritative source page rather than substituting for it. A retrieval system that points at the record instead of replacing it is doing something the failed public chatbots did not. Second, and less flattering, the gate grades its own homework. Nearly every number the cadence turns on — the 76-to-90% accuracy, the 88% answer rate, the 508 blocked jailbreak attempts — is scored first-party, the denominators are unpublished, and the one genuinely external evaluation in the record is a security assessment with the AI Security Institute, not an accuracy audit. So the honest reading is not that the gate is theatre — it plainly fired on the 2023 version — but that a self-run gate is exactly as trustworthy as the metric it reads, and no one outside has audited that metric. That reframes where the remaining leverage sits: not on a sharper model (a newer one can regress on the exact questions the last gate cleared, which is a reason to keep the gate, not to trust past it), and not on more measurement (the measurement is unusually good), but on an independent look at the number the gate turns on, a standing vigilant channel over a public-facing answer that has no caseworker to hold its verify habit, and provenance kept legible on every answer — labeled machine-generated, linked to the record it paraphrases — so the authoritative page, not the paraphrase, stays the thing people act on. The governance question this case makes precise is the one the domain page already asks: how are confident wrong answers detected before a journalist detects them — and here the answer is a real gate, whose one unmet condition is that someone other than its operator reads the dial.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

departmentforscience2025bGroundingGovernmentSave

Department for Science, Innovation and Technology (GOV.UK Algorithmic Transparency Recording Standard), GOV.UK Chat Algorithmic Transparency Record (2025) https://www.gov.uk/algorithmic-transparency-records/dsit-gov-dot-uk-chat

https://www.gov.uk/algorithmic-transparency-records/dsit-gov-dot-uk-chat

Grounds: model org: govuk_chat

Topics: algorithmic-fairness

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalThe UK Government Digital Service ran what it called the government's biggest public test of generative AI to …

The UK Government Digital Service ran what it called the government's biggest public test of generative AI to date: across two gated public pilots (a late-2024 web pilot of 10,136 users asking 23,838 questions, and an autumn-2025 GOV.UK app pilot of 641 users asking 2,670 questions in four weeks), more than 10,000 people asked GOV.UK Chat about 26,000 questions on tax, benefits and visas. Its first 2023 version was held back in findings published January 18, 2024 because, GDS reported, answers did not reach the highest level of accuracy demanded for a site like GOV.UK, including a few cases of hallucination. GDS reports measured answer accuracy rising from 76 percent (its earliest benchmark) to 90 percent by the autumn 2025 pilot, assessed by subject-matter experts plus automated evaluation, an 88 percent answer rate for in-scope questions after a clarifying-questions feature was added, and that 508 attempts to jailbreak the system across the pilots were all prevented by its guardrails; it soft-launched to all GOV.UK app users on March 26, 2026 and officially launched on May 14, 2026. Nearly every one of these figures is self-reported by GDS, the system's operator, and the accuracy denominators and sampling frames are unpublished.

EmpiricalGOV.UK Chat is a retrieval-augmented assistant that, per its Algorithmic Transparency Record published October…

GOV.UK Chat is a retrieval-augmented assistant that, per its Algorithmic Transparency Record published October 7, 2025, answers only from roughly 700,000 vectorised chunks (36.9 GB) of curated official GOV.UK guidance, is instructed to ignore its training data, rejects questions containing phone numbers, emails or card numbers, links every answer back to its GOV.UK source pages with a reminder to verify, and retains question data encrypted for 12 months; GDS states it does not attempt to provide advice and makes no automated decision. GDS's December 2025 vision post frames a content-dependency loop, stating that GOV.UK Chat can only be as good as the content published on GOV.UK by departmental teams. The record's independent evaluation is a jailbreak (security) assessment conducted with the AI Security Institute, alongside the record's own caveat that it is not possible to guarantee no jailbreaking attempts will succeed; there is no independent audit of the accuracy methodology, and GDS's claim that for government-related questions the tool scores higher than widely-used consumer AI assistants is the operator's own comparison.

departmentforscience2025bGroundingGovernmentSave

Department for Science, Innovation and Technology (GOV.UK Algorithmic Transparency Recording Standard), GOV.UK Chat Algorithmic Transparency Record (2025) https://www.gov.uk/algorithmic-transparency-records/dsit-gov-dot-uk-chat

https://www.gov.uk/algorithmic-transparency-records/dsit-gov-dot-uk-chat

Grounds: model org: govuk_chat

Topics: algorithmic-fairness