10 to the 23 AI logo

Domain Atlas / Caseworker documentation & copilots

Case fileUnited Kingdom (Home Office; national asylum casework)large deployment

UK Home Office asylum AI copilots: interview summarisation and policy search

Work with this case in the PAN Lab ↗

In the UK Home Office's own pilot of an AI tool that summarises asylum interview transcripts for decision-makers, 9% of the generated summaries were deemed inaccurate or incomplete and removed by a pre-use filter before any caseworker saw them, and 23% of users reported not being fully confident in the rest; the summaries carried no source references back to the transcript. The official evaluation, published April 29, 2025, measured a 23-minute-per-case time saving (a 32% reduction) for the summariser and about 37 minutes for a companion policy-search tool, and Home Office Calibre quality-assurance reviews found no statistically significant difference in decision quality on small pilot samples. The evaluation recommended addressing the identified limitations before a full rollout, continuous monitoring in early rollout, and a larger-scale evaluation after deployment; the Home Office announced expansion the same day. By January 2026 the policy-search tool had been rolled out to all asylum decision-makers, and per trade-press reporting the summarisation tool entered national rollout in April 2026.[3]

What happened

The Home Office piloted two generative-AI tools in asylum casework: the Asylum Case Summarisation tool, which summarises asylum interview transcripts for decision-makers, and the Asylum Policy Search tool, a chat-interface assistant that finds and summarises Country Policy Information. Both were framed as aids that "do not, and cannot, replace any part of the decision-making process." The summarisation tool ran two eight-week pilot phases (May to June 2024 with 60 test and 15 comparison decision-makers; September to October 2024 with 45 test and 15 comparison), logging 334 AI-assisted cases against 95 comparison cases; the policy-search tool ran an eight-week pilot from October to December 2024 (50 test and 30 comparison decision-makers; 270 versus 214 cases).

The official evaluation, published April 29, 2025, found that summarisation-tool test-group decision-makers reviewed transcripts about 23 minutes quicker per case on average — a 32% time saving — and that the policy-search tool saved about 37 minutes per case (roughly 12 minutes pre-interview and 25 minutes at the decision-writing stage). Crucially, it also recorded that 9% of the summariser's generated summaries were deemed inaccurate or incomplete and removed by a pre-use filter before any caseworker saw them, that 23% of users were not fully confident in the summaries, and that the summaries lacked source references back to the transcript. For the policy-search tool, 79% of users found its representations accurate, 82% reported it provided sufficient information, 54% would continue using it, and 5% were not confident. Home Office Calibre quality-assurance reviews found neither tool had a statistically significant positive or negative impact on decision quality — a finding drawn from small samples in a self-evaluated pilot, not an independent validation. The evaluation's own stated caveats included small scale, non-representative unit selection, and reliance on self-reporting; it recommended that the identified limitations be addressed before a full rollout, that any further rollout involve continuous monitoring and data capture in the early stages, and that a full evaluation on a larger scale follow after deployment. It announced no rollout decision. On the same day the evaluation was published, the Home Office announced it would expand the tools' use, framing AI as helping caseworkers make swifter decisions on asylum claims.

By January 2026 the policy-search tool had been rolled out to all asylum decision-makers, and per public-sector trade-press reporting the summarisation tool entered national rollout to asylum caseworkers in April 2026 (reported as "will be rolled out" that month), with ministers stating the tools operate on a "human in the loop" principle and cannot by themselves decide a claim. Ministers described deployment safeguards including mandatory "AI for all" staff training in 2025, caseworker training, continued subject-matter-expert testing of policy-search outputs with the Country Policy and Information Team, and a dedicated inbox for decision-makers to report problems; whether a Data Protection Impact Assessment would be published was not confirmed. As of mid-2026 the rollout had proceeded without any published post-deployment evaluation or continuous-monitoring data, and — per Open Rights Group — without a published Data Protection Impact Assessment or Equality Impact Assessment, without an Algorithmic Transparency Recording Standard entry, and with the tools' prompts withheld after a Freedom of Information refusal; caseworkers are not required to verify AI outputs. Open Rights Group's January 14, 2026 report "Automating the Hostile Environment" documented these governance gaps and reported that asylum claimants are not informed AI is used on their cases, with over 62,000 asylum applications awaiting an initial decision at end-September 2025.

In a May 2026 written parliamentary answer to Labour MP Kate Osamor, minister Alex Norris confirmed that claimants are not told about the AI tools used in their cases: "No process and/or tooling details are currently released to asylum claimants — this has not changed with the incorporation of AI elements into caseworking." That answer postdates Article 22C of UK GDPR (in force February 5, 2026) on automated-decision-making transparency. A legal opinion by Robin Allen KC and Dee Masters (Cloisters) with Joshua Jackson (Doughty Street Chambers), published March 16, 2026, concluded the Home Office's use of the tools is likely unlawful on procedural-fairness and data-protection grounds, citing the 9% flawed-summary rate and the absence of any correction opportunity for applicants — a contested legal position, not a court ruling; no litigation outcome exists as of mid-2026. Advocacy and legal critics (Open Rights Group, the Helen Bamber Foundation) argue the tools risk omitting crucial case details, embedding bias via undisclosed prompts, and eroding scrutiny in decisions where an error can mean wrongful refusal of protection, with vulnerable applicants serving as a testing ground for the technology. The Home Office's continued rollout is its implicit answer to the legal challenge. Underlying model and vendor details are reported by Open Rights Group, not confirmed in the official evaluation, which describes only a "Large Language Model" and names no vendor; they are paraphrased here, and no model identifier is used.

The sociotechnical reading

The Atlas already carries several copilots, and this is the one where the stakes are highest and the correction channels thinnest. Magic Notes is a vertical documentation tool that rests its whole safety case on one human review gate; the Nava and Imagine LA benefits navigators are verify-before-use tools in which a caseworker reads a cited answer before relaying it; the GDS cross-government experiment is a horizontal layer whose lesson is that measurement is not control. This case is a vertical compression edge: a tool inserted between the primary evidence — the asylum interview transcript, the claimant's own account — and the human judgment node, in a decision that is asymmetric and near-irreversible. The distinct lesson it adds is about who is left to catch an error, and it is this: at a compression edge under throughput pressure, a deployment can close both of the channels that could surface a mistake, at once.

The first channel is the operator's own verification, and the deployment does not remove it so much as price it out. The tool's measured value is a 23-minute-per-case time saving, and those 23 minutes exist only insofar as the decision-maker does not redo the reading the summary replaced. There is no requirement to check a summary against the transcript, and the pilot summaries carried no source references, which is what would have made checking cheap. So full verification is quietly self-defeating: to verify is to give back the entire benefit. The second channel is the affected person, and here the deployment closes it by design. The one party with first-hand knowledge of whether the summary of their own account is right is the claimant — and claimants are never told AI touched their case, so they cannot flag a summary that dropped the detail their claim turned on. A compression error is a subtraction: it is the sentence that is no longer there, and the person best placed to notice a missing sentence about their own life is the person kept unaware the summary exists. Between these two, a third would-be backstop is also thin: the pilot's pre-use filter, the layer that removed 9% of summaries as inaccurate before caseworkers saw them, is the source of the only measured error figure this system has — and it has no documented production equivalent, so the number that makes the case alarming is also the number that no longer exists at scale.

On the system map the case reads as a squeeze with the exits closed. The interview transcript is read and compressed by the summariser; the summary flows forward into case familiarisation and can propagate into the refusal letter and the permanent record; and the two checks that could interrupt it — an independent verification against the transcript, and a standing accuracy filter — are both drawn latent, because in the record neither exists at production scale. That is why the productive levers are the ones that reopen a correction channel rather than sharpen the model: an independent check against the transcript, a standing review cadence that keeps measuring after the pilot filter is gone, a gate on what the summary writes into the record, source-reference provenance that makes verification cheap again, and protecting the verification habit as the saved minutes are spent elsewhere. The honest reading is not that the tools failed — the pilot found no measured quality difference on small samples, and compressing evidence before judgment is a legitimate aid — but that a locally-measured efficiency and an unmeasured, near-invisible subtraction of evidence can be the same 23 minutes, in a decision where the missing sentence can be the difference between protection and refusal. The governance question is not "does it save time?" — it does, measurably — but "when a summary is wrong, who is still positioned to notice, and have you told them the summary exists?"

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

ukhomeofficegovuk2025GroundingGovernment evaluationSave

UK Home Office (GOV.UK), Evaluation of AI trials in the asylum decision making process (2025) https://www.gov.uk/government/publications/evaluation-of-ai-trials-in-the-asylum-decision-making-process/evaluation-of-ai-trials-in-the-asylum-decision-making-process

https://www.gov.uk/government/publications/evaluation-of-ai-trials-in-the-asylum-decision-making-process/evaluation-of-ai-trials-in-the-asylum-decision-making-process

Grounds: model org: home_office_asylum_summarisation

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalIn the UK Home Office's own pilot of an AI tool that summarises asylum interview transcripts for decision-make…

In the UK Home Office's own pilot of an AI tool that summarises asylum interview transcripts for decision-makers, 9% of the generated summaries were deemed inaccurate or incomplete and removed by a pre-use filter before any caseworker saw them, and 23% of users reported not being fully confident in the rest; the summaries carried no source references back to the transcript. The official evaluation, published April 29, 2025, measured a 23-minute-per-case time saving (a 32% reduction) for the summariser and about 37 minutes for a companion policy-search tool, and Home Office Calibre quality-assurance reviews found no statistically significant difference in decision quality on small pilot samples. The evaluation recommended addressing the identified limitations before a full rollout, continuous monitoring in early rollout, and a larger-scale evaluation after deployment; the Home Office announced expansion the same day. By January 2026 the policy-search tool had been rolled out to all asylum decision-makers, and per trade-press reporting the summarisation tool entered national rollout in April 2026.

ukhomeofficegovuk2025GroundingGovernment evaluationSave

UK Home Office (GOV.UK), Evaluation of AI trials in the asylum decision making process (2025) https://www.gov.uk/government/publications/evaluation-of-ai-trials-in-the-asylum-decision-making-process/evaluation-of-ai-trials-in-the-asylum-decision-making-process

https://www.gov.uk/government/publications/evaluation-of-ai-trials-in-the-asylum-decision-making-process/evaluation-of-ai-trials-in-the-asylum-decision-making-process

Grounds: model org: home_office_asylum_summarisation

EmpiricalThe Home Office's asylum interview-summarisation tool inserts a compression step whose measured value is a 23-…

The Home Office's asylum interview-summarisation tool inserts a compression step whose measured value is a 23-minute-per-case time saving that exists only insofar as the decision-maker does not redo the reading the summary replaced: caseworkers are not required to verify summaries against transcripts, and the pilot summaries carried no source references that would make checking cheap. The correction loop is also severed from the other side. In a May 2026 written parliamentary answer, minister Alex Norris confirmed that asylum claimants are not told about the AI tools used in their cases, so the one party with first-hand knowledge of their own account cannot surface a summary error; this postdates Article 22C of UK GDPR (in force February 5, 2026). As of mid-2026 the rollout had proceeded without a published post-deployment evaluation or continuous-monitoring data and, per Open Rights Group, without a published Data Protection Impact Assessment, Equality Impact Assessment, or Algorithmic Transparency Recording Standard entry, with prompts withheld under a Freedom of Information refusal. A March 16, 2026 commissioned legal opinion argues the use is likely unlawful on procedural-fairness and data-protection grounds; that is a contested legal position, not a court ruling.

ukhomeofficegovuk2025GroundingGovernment evaluationSave

UK Home Office (GOV.UK), Evaluation of AI trials in the asylum decision making process (2025) https://www.gov.uk/government/publications/evaluation-of-ai-trials-in-the-asylum-decision-making-process/evaluation-of-ai-trials-in-the-asylum-decision-making-process

https://www.gov.uk/government/publications/evaluation-of-ai-trials-in-the-asylum-decision-making-process/evaluation-of-ai-trials-in-the-asylum-decision-making-process

Grounds: model org: home_office_asylum_summarisation