Domain Atlas / Immigration & asylum AI
A rough compass in the hardest place to be wrong
Explore this deployment in the PAN Lab ↗
A federal asylum agency uses dialect-recognition AI to estimate an applicant's country or region of origin from a short speech sample, as one input into the credibility assessment of their claimed origin. The tool's reliability is limited: government-reported recognition is around 80 percent for Arabic — roughly a 20 percent error rate — and computational linguists judge separating some closely related language varieties close to hopeless. The agency's own caseworkers describe the tool as only a rough compass, too imprecise to resolve the hard cases, and its outputs as clues rather than determinations. Used honestly as one clue among several it is defensible; the documented risk is that an imprecise output acquires more authority than its accuracy supports, in a determination where the state's tool is set against the applicant's own account of who they are.[2]
What happened
Germany's federal asylum agency, the BAMF, deployed dialect-recognition software that analyzes a short speech sample from an applicant and estimates the country or region their speech is most consistent with. It is used as one input into a step of the asylum process that is genuinely hard and genuinely important: assessing whether an applicant's claimed origin is credible, which can bear on whether they qualify for protection. The intent is defensible — origin is often contested and hard to verify, and a linguistic signal is one more piece of information a caseworker can weigh.
The tool's reliability is the first fact that shapes the governance, and it is modest. Government figures put recognition around 80 percent for Arabic — meaning roughly one in five judgments is wrong — and computational linguists have said that separating some closely related varieties (vernacular Persian, Dari, Pashto) is close to hopeless, because the varieties genuinely overlap and a short sample cannot resolve them. This is not a tool that is usually right with occasional errors; it is a tool that is often uncertain by the nature of what it is trying to do. The agency's own caseworkers, in fieldwork, describe it accordingly: a "rough compass," "too imprecise to solve the problematic cases," producing "clues" rather than answers.
That candid framing is, in fact, the responsible one, and it is worth crediting. A low-reliability signal used explicitly as one clue among several, by a caseworker who knows its limits and weighs it against everything else, is a legitimate use — the tool is not automating the decision, and the people running it are not pretending it is precise. The governance question is whether that honest framing actually holds in practice, or whether the imprecision erodes under pressure. The documented risk, drawn from the fieldwork, is subtle: the tool's output can strengthen the caseworker's epistemic authority in the credibility confrontation — "the software indicates your speech is not consistent with your claimed origin" is a hard thing for an applicant to rebut, and it can carry more weight in the room and in the record than a 20-percent error rate warrants. A rough compass pointed at a person, in a setting where the state's instrument meets the applicant's own testimony, does not stay merely rough if it is allowed to tip a credibility finding.
The second fact is the severed correction loop, and it is structural to the setting. The person best placed to catch an error in an origin estimate is the applicant, who actually knows where they are from. But in asylum procedures the applicant is frequently not shown the AI's role or its output in a form they can contest, so the one party with both the most at stake and the most knowledge of the truth is cut out of the correction. An error in a rough-compass reading that the applicant could immediately explain — a childhood spent across a border, an education in a second dialect, a family that spoke differently from the region — may never surface, because the loop that would surface it is closed on the applicant's side.
The honest reading is that this is a defensible tool honestly described by the people who use it, deployed in the setting where the consequences of overweighting it are most severe, with two governable surfaces the honesty does not by itself secure: whether the tool's documented imprecision actually bounds the weight it carries in the decision and the record, and whether the applicant can see and contest a signal used to assess them. The stakes are what make these non-negotiable: a 20-percent error rate is one thing in a recommendation and another when it helps decide whether a person is returned to a country they fled.
The sociotechnical reading
This case opens the immigration-asylum domain by pairing an honestly-described tool with the setting where honesty is hardest to keep. The agency's own caseworkers call the dialect tool a rough compass and use it as one clue — which is the responsible framing, and the map credits it. The governance question is whether that framing survives contact with a credibility confrontation in which the state's instrument is set against an applicant's account, and the documented risk is that it does not: an imprecise output can acquire authority in the room and in the record that its roughly one-in-five error rate cannot support. The map's reading is that in the highest-stakes settings, reliability must actively bound authority, because the honest description of a tool's limits does not automatically constrain the weight the tool is given.
The first latent check is exactly that binding: a control that ties the weight an origin signal carries to its measured reliability, so a rough compass is recorded and used as one, and cannot harden into a credibility finding it cannot support. This is not a demand that the tool be more accurate — its imprecision may be irreducible for genuinely overlapping dialects — but that its imprecision be load-bearing in how it is used, rather than acknowledged and then set aside when a caseworker needs a tiebreaker.
The second latent check is the applicant's ability to see and contest the signal. The person who knows their own origin is the best error-corrector any such system could have, and asylum procedures routinely sever that loop by not disclosing the AI's role or output in a contestable form. The map reads this as the domain's severed correction loop on the side that holds the truth — a structural inversion of good practice, because the correction is cut off at precisely the point where it would be most reliable. Disclosing the signal and giving the applicant a real chance to rebut it is the check, and it is owed most where the stakes are highest.
The Lab network models only the deploying agency: its dialect tool, the caseworkers who weigh its outputs, and its case records. No asylum outcome and no applicant's credibility is computed on any diagram. The applicant is boundary-only; the tool's reliability figures, the linguists' judgments, the caseworkers' rough-compass framing, and the disclosure gap are institutional signals that live in this case file, never on any network — reliability figures and fieldwork are recorded external findings, and nothing here adjudicates any individual claim. The map's instruction is to credit the honest one-clue framing, to make reliability bound the authority the signal carries, and to reopen the correction loop to the applicant, because a rough compass in the hardest place to be wrong is safe only if its roughness is enforced and the person it points at can answer it.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.