10 to the 23 AI logo

Domain Atlas / Caseworker documentation & copilots

Case fileUnited Kingdom (cross-government)giant deployment

GDS Microsoft 365 Copilot cross-government experiment

Work with this case in the PAN Lab ↗

The Government Digital Service ran a cross-government experiment with Microsoft 365 Copilot from September 30 to December 31, 2024, with about 20,000 employees across 12 organisations, and published the findings report on June 2, 2025. Participants self-reported saving an average of about 26 minutes per working day (the report extrapolates this to roughly 13 days a year from the median values of six reported time-savings ranges; independent coverage recomputed it to about 4.6 days on a 253-working-day basis), 17% reported no clear savings, adoption held near 80% after peaking at about 83%, and 82% said they would not want to return to working without it. The experiment measured adoption and self-reported time rather than output quality: the report recorded no audited error rate, flagged significant accuracy concern for low-verifiability tasks such as grievance handling and performance evaluations, noted external web data was used without built-in verification, and documented a provenance failure in which the tool struggled to identify which documents generated a response.[3]

What happened

The Government Digital Service ran a cross-government experiment with Microsoft 365 Copilot — a horizontal generative-AI productivity layer that drafts documents, summarises and transcribes meetings, searches across files, and writes email, rather than adjudicating any single case — from September 30 to December 31, 2024, with about 20,000 employees across 12 organisations (including DWP, HMRC, the Home Office, and MoJ), the largest deployment of the tool in UK government to date. The findings report, published June 2, 2025, recorded self-reported average savings of about 26 minutes per working day, which the report extrapolated to roughly 13 days a year from the median values of six reported time-savings ranges; the distribution was uneven, with 17% reporting no clear savings and over a third saving 30 or more minutes daily, and task-level means of 24 minutes for document drafting, 19 for presentation creation, and 9 for communication and scheduling. Adoption peaked at about 83% in the first month and held near 80%, with Copilot in Teams most used (71%) while Excel (23%) and PowerPoint (24%) lagged; 82% of users said they would not want to return to working without it, overall satisfaction was 7.7 out of 10, and 85% agreed the tool provided good value. The evaluation drew on 7,115 survey responses, anonymised usage telemetry on 14,500 users, and five profession-based focus groups, and found that professions and grades with the lowest AI familiarity saw the least benefit. Crucially, it measured adoption and time, not output quality: it flagged "significant concern about the accuracy of M365 Copilot's outputs" for domain-heavy, low-verifiability tasks such as grievance handling and performance evaluations, where "any errors could lead to reputational risks"; noted the tool "struggles with nuanced or context-heavy data requiring human judgement" and used external web data "without built-in verification"; and recorded a provenance failure in which the tool "struggled to identify the exact documents used to generate responses." Independent coverage stressed that the headline figures are self-reported and that the report could not determine how the saved time was actually spent, and recomputed the 13-day extrapolation to about 4.6 days on a 253-working-day basis.

Companion departmental evaluations sharpened the picture. The Department for Business and Trade's evaluation (1,000 licences, October to December 2024; a diary study plus observed-task exercises), published August 28, 2025 and announced on the department's Digital Trade blog on September 25, 2025, reported 72% user satisfaction but concluded it "did not find robust evidence to suggest that time savings are leading to improved productivity." In its observed tasks, Copilot users completed spreadsheet data analysis more slowly and to worse quality and accuracy than non-users, and produced PowerPoint slides over 7 minutes faster on average but to worse quality and accuracy that then required corrective action — a speed-for-quality trade, not an unambiguous gain. In its diary study, 22% of respondents said they had identified hallucinations in the output, 43% detected none, 11% were unsure, and a further roughly one in five did not answer the question (with 3% not applicable), so the figure measures user-detected hallucination rather than audited incidence; users averaged 1.14 Copilot actions per working day, and the tool suited HR and Commercial roles better than Policy and Legal. The Department for Work and Pensions' extended trial (3,549 licences, October 2024 to March 2025), published January 29, 2026, measured 19 minutes a day saved across eight routine tasks against a comparison group (statistically significant, 95% confidence interval 17 to 22 minutes), found 85% rating meeting-note accuracy "Very Good" or "Good," reported that users consistently reviewed outputs before use with particular caution on sensitive content, and measured job satisfaction rising 0.56 points on a 7-point scale; it concluded the tool is "complementary" to human expertise and requires "consistent human oversight."

HMRC's three-phase evaluation, published July 9, 2026, added a randomly allocated cohort (3,000 licences plus 500 volunteer and reasonable-adjustment licences, September to December 2024, covering HMRC and the Valuation Office Agency, with use restricted to material below Official Sensitive) to earlier volunteer phases. It reported self-reported savings of about 2 to 3 percent of the working week (roughly 60 minutes), discounted about 20% for non-usage; found 46% of non-users citing security and data-privacy concerns; and carried an analyst projection that deploying up to 50,000 licences until March 2028 might generate around 50 million pounds a year in "capacity generating, net productivity benefits" — a projection, not a funded commitment. By April 2026 HMRC had, per independent reporting, rolled out about 28,000 licences, with its chief AI officer stating an ambition to make it "the most AI-enabled tax authority on the planet"; the same coverage framed the trust-and-reliance tension bluntly, describing a tool that "works well enough to rely on, not quite well enough to trust, and far too embedded to switch off." Publicly cited per-licence figures (about 19 pounds per user per month) are a reference to the consumer Microsoft Copilot Pro subscription price rather than the government's negotiated cost, with commercial licence prices reported in a 4.90-to-18.10-pound range. This is a horizontal productivity layer, not a casework-adjudication tool: it makes no eligibility determination, and every time-savings figure it produced is self-reported and unverified.

The sociotechnical reading

The Atlas already carries two kinds of copilot, and this is their horizontal cousin. Magic Notes is a vertical documentation tool that rests its whole safety case on one human review gate on the error-to-record pathway; the Nava and Imagine LA benefits navigators are verify-before-use tools in which a caseworker reads a cited answer before relaying it. This deployment is different in kind: a single copilot inserted across every operator class at once — drafting, summarising, transcribing, searching, emailing — and it teaches something none of the vertical cases do, because it is the best-measured copilot deployment anywhere and what it measures is the trade, not the harm. A horizontal copilot is locally positive and globally uncertain. Task-level time is saved, self-reported and real enough that most users would not revert, yet the companion evaluation could find no robust department-level productivity gain — because the minutes saved are the minutes not spent verifying. The benefit concentrates in high-verifiability tasks (drafting, transcription) while the accuracy concern concentrates in low-verifiability, high-stakes tasks (grievances, performance evaluations); the tool is trusted most exactly where a wrong line is hardest to catch. The clean empirical signature is adoption lock-in outrunning verification capacity: 82% would not go back, and 22% caught it hallucinating — but only the ones vigilant enough to look, with no systemic backstop behind them.

Two edges on the system map carry the case, and both are quiet. The first is the write-back into shared memory. Unlike the verify-before-use copilots that barely write anything back, this layer writes generated text broadly into shared documents, minutes, and mail that the same tool later retrieves and re-summarises — a direct memory-write, memory-read loop — and it could not identify which document an answer came from, so a contaminated line is untraceable once written. The second is the missing output-level check. For all its rigor — sample sizes, confidence intervals, a comparison group, a randomised cohort, honest negative findings — the experiment measured adoption and self-reported time and not per-output quality; there was no output-level quality-assurance layer, only the individual reviewer whose checking time the tool consumes. That is the distinct lesson this case adds to the collection: measurement is not control. A deployment can be evaluated more carefully than almost any other here and still have no gate on what generated text enters the shared record. So the governance question is not "does it save time?" — it does, self-reported — but "what is the saved time no longer spent on, and where does the unverified text go?" The productive levers follow directly from the two edges: a write-gate on what enters shared memory, provenance labels so a re-summarised line carries its origin, the independent output-level check the evaluations themselves noted was absent, and protecting verification capacity as adoption locks in. The benefit and the safety cost here are the same transaction, which is why the honest reading is not that the copilot failed but that a locally-measured gain and a diffuse, unmeasured contamination of the shared record can be, precisely, the same 26 minutes.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

governmentdigitalservicedsit2025GroundingGovernment evaluationSave

Government Digital Service (DSIT), Microsoft 365 Copilot Experiment Cross-Government Findings Report (HTML) (2025) https://www.gov.uk/government/publications/microsoft-365-copilot-experiment-cross-government-findings-report/microsoft-365-copilot-experiment-cross-government-findings-report-html

https://www.gov.uk/government/publications/microsoft-365-copilot-experiment-cross-government-findings-report/microsoft-365-copilot-experiment-cross-government-findings-report-html

Grounds: model org: gds_m365_copilot_experiment

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalThe Government Digital Service ran a cross-government experiment with Microsoft 365 Copilot from September 30 …

The Government Digital Service ran a cross-government experiment with Microsoft 365 Copilot from September 30 to December 31, 2024, with about 20,000 employees across 12 organisations, and published the findings report on June 2, 2025. Participants self-reported saving an average of about 26 minutes per working day (the report extrapolates this to roughly 13 days a year from the median values of six reported time-savings ranges; independent coverage recomputed it to about 4.6 days on a 253-working-day basis), 17% reported no clear savings, adoption held near 80% after peaking at about 83%, and 82% said they would not want to return to working without it. The experiment measured adoption and self-reported time rather than output quality: the report recorded no audited error rate, flagged significant accuracy concern for low-verifiability tasks such as grievance handling and performance evaluations, noted external web data was used without built-in verification, and documented a provenance failure in which the tool struggled to identify which documents generated a response.

governmentdigitalservicedsit2025GroundingGovernment evaluationSave

Government Digital Service (DSIT), Microsoft 365 Copilot Experiment Cross-Government Findings Report (HTML) (2025) https://www.gov.uk/government/publications/microsoft-365-copilot-experiment-cross-government-findings-report/microsoft-365-copilot-experiment-cross-government-findings-report-html

https://www.gov.uk/government/publications/microsoft-365-copilot-experiment-cross-government-findings-report/microsoft-365-copilot-experiment-cross-government-findings-report-html

Grounds: model org: gds_m365_copilot_experiment

EmpiricalA companion Department for Business and Trade evaluation of Microsoft 365 Copilot (1,000 licences, October to …

A companion Department for Business and Trade evaluation of Microsoft 365 Copilot (1,000 licences, October to December 2024; published August 28, 2025) reported 72% user satisfaction but concluded it did not find robust evidence that time savings were leading to improved productivity; in observed tasks its users completed spreadsheet data analysis more slowly and to worse quality and accuracy than non-users, and produced presentation slides over 7 minutes faster on average but to worse quality and accuracy that then needed correction. In its diary study, 22% of respondents said they had identified hallucinations, 43% detected none, and 11% were unsure, with a further roughly one in five not answering, so the figure reflects user-detected hallucination rather than audited incidence. A Department for Work and Pensions evaluation (3,549 licences; published January 29, 2026) measured 19 minutes a day saved across eight routine tasks against a comparison group (95% confidence interval 17 to 22 minutes), found 85% rating meeting-note accuracy good or very good, reported that users consistently reviewed outputs before use, and concluded the tool is complementary to human expertise and requires consistent human oversight.