Domain Atlas / Caseworker documentation & copilots
Minute / Local Transcribe
The UK government built its own AI meeting scribe for council caseworkers and piloted it through a cohort of 25 selected councils (22 active, more than 400 users) under one shared pooled-assurance record, then open-sourced it and adapted it to enlist around 500 housing and homelessness workers by June 2026; the cohort published a multi-council governance dataset but no transcription-accuracy or error-rate evaluation, and standard risk controls such as penetration testing and certification had not been completed on the alpha at pilot time.[5]
What happened
Minute is an AI meeting scribe the UK government built for itself, rather than bought: it transcribes public sector meetings and drafts customisable, standardised summaries, and it was developed by the Incubator for AI inside the Government Digital Service in the Department for Science, Innovation and Technology. After a demonstration to more than 150 local-government officers in February 2025, 25 councils were selected from 52 applicants for a six-week alpha pilot, each capped at 25 users and given the tool free of charge on one shared central instance without a formal contract. 22 councils remained active (three withdrew citing capacity, digital fatigue, and cloud-policy conflicts), and testing reached more than 400 users across adult social care, children's services, planning, housing and democratic services, in meeting types including safeguarding reviews, supervisions and legal case reviews. The pilot identified more than 40 distinct social-care use cases in which workers must complete lengthy forms, reports or assessments; on the shared infrastructure it cost about 0.50 pounds to transcribe a meeting, and councils rated the tool 8.5 out of 10 for recommendation.
What is unusual about the record is how much of it is governance and how little of it is accuracy. Assurance was pooled across the cohort: the Local Government Association and the London Office of Technology and Innovation facilitated a central data-processing template reviewed by the department as designated data processor, shared impact-assessment, equality-assessment and consent templates, a council readiness toolkit, and bi-weekly cohort calls that most councils rated useful. Internal assurance timelines varied widely (14 per cent of councils completed it in under two weeks, 41 per cent took more than a month), and standard risk controls such as penetration testing and security certification had not been completed on the alpha tool. The codebase was open-sourced in October 2025, the Ministry of Justice forked it into a probation tool, and by June 2026 the Ministry of Housing, Communities and Local Government had adapted it into Local Transcribe and enlisted around 500 housing and homelessness workers to pilot it, with a design flow of transcribe, then a standardised summary, then user review and submission. The department's AI director publicly justified a central tool on the grounds that fragmented transcription adoption produces duplicate spend and duplicative assessment and assurance, which matter when officers are transcribing critical interactions with vulnerable groups.
The gap runs through the whole record: there is no published transcription-accuracy evaluation, no error-rate audit, and no reviewer-override data for the tool, and the time-savings figures cited for it - some users reporting halved note-taking, one council estimating up to a 90 per cent reduction in recap time, a government early-testing claim of about one hour saved per one-hour meeting - are self-reported by users or asserted by government, not independent measurements. Independent scrutiny exists, but at sector level: the Ada Lovelace Institute's research on AI transcription in social work (interviews with 39 social workers across 17 local authorities in England and Scotland) found that local authorities focus their evaluations on efficiency rather than impact on people who draw on care, that risks such as bias and hallucination are not being fully assessed, and that perceptions of reliability and the need for human oversight vary significantly among workers. That research covers such tools generally, not this one specifically, and no decommissioning, litigation or scandal has been reported.
The sociotechnical reading
Minute is the Atlas's cleanest case that governance maturity and evidence of accuracy are different axes, and a system can be exemplary on one and blank on the other. Almost everything usually missing from these cases is present here: a published multi-council governance dataset, a shared data-processing template with a named data processor, pooled impact and equality assessments, a readiness toolkit, and a real community-of-practice check in the bi-weekly cohort calls. What is absent is the one thing the map cares about most - any standing measurement of whether the summaries are right. On the diagram that shows up as a thorough assurance node wired only to the process pathways, with the accuracy-check edges left latent; the productive question is not "is this well governed" (it plainly is, on process) but "governed for what", and the levers that matter are the ones that switch the latent accuracy check on.
Two structural features sharpen the reading. First, the coupling is one-to-many: every council runs the same centrally hosted instance, so a single tool-level defect - a systematic mis-transcription mode, a template bias - would not be a local error but the same error written into many councils' records at once, and the modeled danger is correlation rather than incidence. That is the opposite failure geometry from a per-deployment tool, and it is why an independent second read on the shared instance is a different control from a better model. Second, the very pooling that makes assurance efficient can concentrate the scrutiny that might catch the defect: consolidating twenty-two assessments into one is a real saving and also a single point where an accuracy question, once not asked, stays not asked everywhere. The state-built posture removes one classic exposure entirely - there is no commercial vendor and no data-processing agreement to an outside processor, so the vendor-egress pathway that dominates the bought-copilot case simply does not exist here - which is exactly what leaves the accuracy question, not the data-sharing question, as the load-bearing one. The lesson the Field Guide's evidence and monitoring material turns on: a beautifully documented assurance record is a claim about process, not a claim about output, and until someone measures the error rate the safeguard the whole design rests on - a human who reviews before submitting - is being asked to catch a failure no one has characterized.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.