Domain Atlas / Public benefits & eligibility
Forsakringskassan VAB fraud-selection profile (Sweden)
Work with this case in the PAN Lab ↗
Analysing the Swedish Social Insurance Agency (Forsakringskassan) 2017 outcome data, Lighthouse Reports and Svenska Dagbladet reported on 27 November 2024 that the agency's in-house machine-learning risk profile for the temporary parental allowance (VAB) selected women (more than 1.5x), people of a foreign background (about 2.5x), below-median earners (2.97x), and people without a university degree (3.31x) for fraud investigation more often than comparison groups by demographic parity, and wrongly flagged those groups at higher false-positive rates (about 1.7x for women and 2.4x for people of a foreign background); in the agency's paired random-control sample, 20.2 percent of applications contained at least one day incorrectly paid, an unbiased base error rate. These are outcome computations under specific fairness definitions from a single obtained year of data, not confirmed model internals; the agency disputed the framing and did not release the model. The data-protection regulator IMY closed its GDPR supervision on 18 November 2025 for mootness after the agency withdrew the system, and no court or regulator issued a discrimination or GDPR penalty.[5]
What happened
Forsakringskassan, the Swedish Social Insurance Agency, used an in-house machine-learning risk profile — a "riskbaserad urvalsprofil" — to score recipients of the temporary parental allowance (tillfallig foraldrapenning, or VAB, paid to parents who stay home to care for a sick child) and select the highest-scored for targeted fraud and error control investigations. The system was, per Amnesty International, in use since at least 2013. It was a targeting step, not an automated denial: flagged cases were routed to the agency's control department, where human investigators worked them before any sanction. In 2017 about 6,129 people were selected for investigation out of roughly 977,730 applicants — 1,047 chosen at random and 5,082 chosen by the model. The agency also ran that random arm as a deliberate design feature, and in the random sample 20.2 percent of applications contained at least one day incorrectly paid, giving an unbiased ground-truth base error rate for the benefit.
On 27 November 2024, Lighthouse Reports, working with the Swedish newspaper Svenska Dagbladet, published "Sweden's Suspicion Machine." Having obtained the agency's 2017 outcome data (including whether errors were actually found) but not the model itself, the journalists tested the selections against six fairness definitions. By demographic parity the model over-selected women by more than 1.5 times, people of a foreign background by about 2.5 times, below-median earners by 2.97 times, and people without a university degree by 3.31 times, relative to comparison groups. False-positive rates — the rate of being wrongly flagged — ran higher for the same groups: about 1.7 times for women, 2.4 times for people of foreign background, and more than 3 times for below-median earners and for those without a degree. Notably, on the precision measure the model was not less accurate for those groups: it was slightly more precise for people of foreign background, without a degree, and below the median income than for the corresponding advantaged groups, and only marginally less precise for women. The reported harm was therefore concentrated in who was wrongly suspected, not in a lower hit rate. These are outcome computations under specific fairness definitions, drawn from a single year of obtained data; the agency did not release the model, its features, or its precision, and it disputed the discrimination framing. It had resisted disclosure for roughly three years on fraud-prevention grounds. (Agency-reported figures for a different year, 2022, recorded 5,520 fraud investigations, 1,686 police referrals, and 166 convictions among 1,054 reported prosecutorial outcomes; these are a separate cohort and are not comparable with the 2017 fairness data.)
The system had been examined before. In 2018 the audit inspectorate ISF (Inspektionen for socialforsakringen) published two reviews. One, "Profilering som urvalsmetod for riktade kontroller," found the risk-based controls "betydligt mer traffsakra" (substantially more accurate) than alternative control methods, while raising legal-certainty (rattssakerhet) and equal-treatment concerns and recommending safeguards. The other, "Riskbaserade urvalsprofiler och likabehandling," warned that an accurate model can still be inequitable — its example was that "two groups err to an equal extent but only one is followed up" — and that correlation-based models are vulnerable to confounding; Amnesty later characterized ISF as finding the algorithm "does not meet equal treatment." Amnesty also reported that in 2020 a former data protection officer at Forsakringskassan had warned that the operation violated European data-protection rules. When the investigation broke, Amnesty called on 27 November 2024 for the system to be "immediately discontinued," invoking the EU AI Act's high-risk and social-scoring provisions; two days later Sweden's Equality Ombudsman (DO), through chief of staff Samuel Engblom, urged people who felt discriminated against to file complaints so DO could investigate.
The binding scrutiny never came. In June 2025 the privacy regulator IMY (Integritetsskyddsmyndigheten, the Swedish Authority for Privacy Protection) opened a supervision case into whether the risk-based selection complied with the GDPR. In its response dated 5 September 2025, Forsakringskassan confirmed it had used a machine-learning risk profile but had taken it out of service about a month earlier, with no plans to resume it. On 18 November 2025 IMY closed the case, reasoning that because the system was no longer used and the potential risks had ceased, no further investigation was warranted. That closure is for mootness — not a ruling that the system was lawful or non-discriminatory. No court judgment and no regulatory fine were issued. The deployed model's exact algorithm class, feature set, and precision remain undisclosed, and no AI model identifier is documented.
The sociotechnical reading
Forsakringskassan is the Atlas's case of the refused audit. Its nearest neighbour, Rotterdam, is the mirror image: there, investigative journalists obtained the live model and its training data, measured the bias directly, and the city suspended the tool — the audit-as-actor succeeding because it got inside. Here the same outlet reached the same kind of finding, but from the outside only. Denied the model through a roughly three-year information refusal, it had to reconstruct the skew from an obtained year of outcome data, the agency disputed the methodology, and the system was retired before any of it could bind. Where SyRI was struck down by a court reading a statute, and where the Dutch childcare-benefits reckoning came from a parliamentary inquiry, this agency simply out-waited its scrutiny: the data-protection regulator closed for mootness once the tool was gone, so no court or regulator ever ruled.
The map's lesson is narrower and sharper than "it was biased." This system possessed, by design, exactly the instrument an audit needs — a paired random-control arm that each year produced bias-free ground truth about who was actually erring. An agency that has ground truth and refuses to reconcile its model against it, or to let anyone outside see the model, is not auditable, however much data it holds. Auditability is not the same as having the data; it is the willingness to turn the check on the model and to be seen doing it. That is why, in the Lab, the sharpest levers are the disclosure and external-audit ones and the reconciliation of the profile against the random arm it already ran — not a better classifier. It also matters what the harm was and was not. The audit inspectorate itself judged the profiling comparatively accurate, and the external investigation's own precision numbers do not show the model performing worse for the groups it over-flagged; the injury it documented was a heavier false-positive burden landing on women, migrants, low earners, and the less-educated. This is the counterpart to the accuracy-inversion cases elsewhere in this domain: here the model can be accurate on its own terms and still concentrate wrongful suspicion, exactly the inequity ISF warned of in 2018 and exactly the reason a hit-rate defence answers the wrong question. The feedback loop keeps it turning — past control outcomes are the training labels, so the over-investigated are selected again — and with the reconciliation edge closed, that loop tightens where no one can watch it.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.