Domain Atlas / Benefits navigation & public-facing chat
Frida (NAV Norway)
Work with this case in the PAN Lab ↗
Frida is the chatbot at the front line of the Norwegian Labour and Welfare Administration's (NAV) anonymous contact-center chat channel; NAV states it launched in summer 2018 and, as of 2026, that citizens first meet Frida (open 24 hours a day) and can ask it for a human advisor on weekdays between 9:00 and 15:00, with the channel anonymous and no personal information visible to NAV. During the COVID-19 lockdown NAV reported a roughly 250 percent surge in inquiries; the platform vendor's case study reports the chatbot answered more than 270,000 coronavirus-related inquiries and that about 80 percent of enquiries were resolved without escalating to a human, and NAV's own funded research report records nearly 11,000 inquiries in Frida on some days between March and May 2020 with a week-13-2020 peak equal to the capacity of about 230 human advisors, where the vendor and the peer-reviewed EJIS study state about 220. These pandemic figures originate substantially in the vendor's marketing case study and are reported here as vendor claims with the 220-versus-230 source tension left unresolved; the roughly 80 percent containment is a completion or non-escalation rate, not a measure of answer accuracy.[4]
What happened
Frida is the chatbot at the front line of the Norwegian Labour and Welfare Administration (NAV), Norway's national agency for pensions, child support, unemployment benefits and sick-leave benefits. It sits on the NAV Kontaktsenter (NKS) chat channel on nav.no: a citizen types a free-text question and Frida returns a scripted answer about welfare rules and services. In the documented period it is not a generative system but an intent-classification conversational agent built on a commercial no-code conversational-AI platform (the vendor, boost.ai), with a curated library of intents and answers maintained in-house by a small NAV team of about six non-technical employees — "AI trainers" — who mine live conversations to update the content. NAV states the chatbot launched in summer 2018. The channel is explicitly anonymous: as of 2026 the nav.no contact page tells citizens they "will first meet chatbot Frida, who is open 24 hours a day," that on weekdays between 9:00 and 15:00 they can ask Frida to chat with a human advisor, and that "you can be anonymous, and we can't see any personal information about you." It is a navigation-tier tool — it gives general guidance, not benefit decisions, and it writes back to no individual case record.
The event that made Frida a research object was the COVID-19 lockdown. NAV faced a roughly 250 percent surge in citizen inquiries; the platform vendor's case study reports the chatbot answered more than 270,000 coronavirus-related inquiries, and the NAV-funded research report records nearly 11,000 inquiries in Frida on some days between March and May 2020. At peak — week 13 of 2020 — the volume corresponded to the capacity of roughly 220 human advisors per the peer-reviewed European Journal of Information Systems (EJIS) study and the vendor case study, or 230 advisors per the NAV-funded report; the small numeric inconsistency across sources is a documented tension, not a resolved number. Most conversations were completed by the chatbot alone. About one in five was transferred to a live chat with a human service agent — the vendor phrases this as "80% of enquiries resolved without escalating to a human representative." That containment figure is the number most often quoted about Frida, and it is a completion or non-escalation rate, not an accuracy rate.
The distinctive, best-documented property of this deployment is the chatbot-to-human handover boundary, and it turns out to be a governance dial. The NAV-funded Frida@work project (November 2020 to October 2021; NTNU, the University of Agder and the University of Oslo) recorded that when NAV removed the explicit choice between the chatbot and human chat, "citizens received sufficient help from the chatbot in most cases," and only about 30 percent of dialogues were transferred to a human — a directly measured intervention on the handover link. The escalation fraction, in other words, is regime-specific: about one in five under free channel choice, about 30 percent when the human-chat option was hidden, and any statement about it has to name the interface in force. The citizen holds the escalation lever — asking Frida for an advisor — but NAV controls how visible that lever is.
What crosses the handover is imperfect. The Frida@work project found that citizens transferred to a human advisor are often unsure whether they are chatting with a human or a machine, because the chat window is near-identical; that general trust and pre-use expectations strongly shape whether people perceive the chatbot as competent; and it delivered eleven design principles plus a specific recommendation to build a separate internal chatbot to assist chat employees at handover — mapping context, building familiarity, and preparing responses. A University of Agder master's thesis studied the Frida-to-NKS handover from an affordances perspective and found chat employees developed their own understanding of affordances beyond the chatbot's intended use; its peer-reviewed continuation by the thesis supervisors carried the handover topic into the IFIP TDIT 2022 proceedings (a short paper that frames the handover problem generically). Independent chat-log research characterized how Frida actually answers: a Scandinavian Journal of Information Systems study documented irrelevant answers and omitted important information, and an Electronic Participation (ePart 2020) study identified three categories of domain-knowledge failure — vocabulary gaps, uncertainty about which regulations apply, and citizens' misinterpretation of rules — with the most critical failures occurring when a misunderstanding goes undetected; it recommended avoiding human-like avatars and making the tool's limitations transparent. (The ePart paper anonymizes the agency's chatbot under the pseudonym "Anna"; the system identification — the Norwegian Labour and Welfare Administration's chatbot — is explicit and correct, but the paper does not use the name Frida.) A 2023 University of Oslo case-study thesis found that question complexity affected answer accuracy, that citizens asked about their own specific cases while the chatbot answered in general terms, that the chatbot lacked contextual understanding, and that citizens gave positive feedback when they were escalated to a human advisor.
Two later developments frame the current picture. NAV's own analysis unit published a quantitative channel-use study in 2025 (Arbeid og velferd nr. 2-2025). It reported that the chatbot icon was prominently visible on nav.no from autumn 2020 until it was hidden roughly two years later; that inquiries to the contact center have declined steadily since 2019 (apart from the early pandemic); that the chatbot's visibility appears not to influence contact-center inquiry volumes; and that chatbot use correlates with use of the nav.no search engine, while average call duration increased. Crucially, the analysis attributes the declining inquiry volumes not to the chatbot but to a bundle of causes — self-service improvements, new application systems, SMS notifications and changed NKS practices — and the chatbot's own contribution is not separable; this is a null result on substitution, so Frida cannot be presented as demonstrably reducing human workload outside the crisis peak. Separately, NAV piloted an internal generative-AI assistant for NKS advisors in autumn 2024 that answers advisors' natural-language questions with references restricted to the internal knowledge base (a distinct system, named here without any model identifier) — in effect realizing the Frida@work recommendation of an employee-facing assistant at the handover boundary. As of 2026 Frida remains operational as NAV's anonymous chat front door.
The evidence here is unusually deep for a navigation-tier system — a NAV-commissioned clinical-inquiry study in a peer-reviewed journal, a NAV-funded three-university project reporting directly to the agency, multiple independent chat-log studies and theses, and the agency's own statistical analysis — but two honest limits define what can be said. The headline pandemic figures (270,000-plus inquiries, 80 percent containment, the 220-full-time-employee equivalence, the 250 percent surge) originate in the vendor's marketing case study and are labeled as such; and no per-answer accuracy or error rate has ever been published, so the roughly 80 percent containment must be read as a completion rate, not a measure of how often the answers were right.
The sociotechnical reading
Almost every case in this domain asks whether the answer is right. Frida is the case that asks a different question, and it is the one the others quietly depend on: when the machine cannot help, how easy is it to reach a person — and who decides? The atlas already carries public-facing chatbots whose lesson is about the answer. The NYC MyCity business bot was documented-wrong and kept answering ("exposure is not correction"); GOV.UK Chat built a staged gate that could say "not yet." Frida's payload is neither the answer nor a gate in front of it. It is the handover boundary itself — the single best-documented chatbot-to-human escalation link in any welfare agency — and the striking finding is that the boundary is a governance dial. Under free channel choice, about one in five conversations reached a human; when NAV removed the explicit human-chat option, about 30 percent did. Those are regime-specific numbers, not a clean before-and-after, and the atlas reports them as such — but the structural point survives the caveat: the escalation fraction, one of the very few directly measured, directly governable quantities in this whole library, is set by the agency's interface design, not by the citizen and not by the model. The citizen holds the lever; NAV decides how visible it is.
That reframes where the risk lives, and it is the distinct lesson this case adds: containment is not resolution. The number everyone quotes about Frida — roughly four in five conversations "resolved" without a human — is a completion rate, and the log studies are unanimous that the most dangerous conversations are inside that four-fifths, not outside it. The signature failure the researchers named is the undetected misunderstanding: a citizen asks about their own specific situation, the chatbot answers in general terms, something is quietly wrong or omitted, and nothing in the interaction tells the citizen to escalate. A gate that could say "not yet" does not help here, because there is no release to hold — the harm is a single contained conversation, invisible unless someone reads it. And no one is required to. This is the paradox the Lab draws as a latent check that starts closed: the oversight around Frida is extraordinary — a commissioned clinical inquiry, a three-university project, independent log studies, the agency's own analysis unit — and yet not one of them is a per-answer accuracy audit that reads the conversations the chatbot completed. Frida is heavily *studied* and not once *gated on its own answers*. So a governance move that looks efficient — raise containment by making the human option less prominent — can silently widen exactly the population no one is auditing.
The handover has a second defect the copilot cases do not: what crosses it is degraded. The verify-before-use benefits copilots (Nava, the Benefit Navigator) and Caddy put a human *between* the model and the affected person by design; Frida puts no one in that seat, and reaches a human only by escalation — into a transcript whose context, the NAV-funded project found, survives imperfectly, arriving with the citizen often unsure whether they are now talking to a person at all. The advisors had to invent their own workarounds at the seam; NAV eventually saw the problem clearly enough to build an advisor-facing assistant aimed squarely at it. So the human who is the case's real safety layer inherits a half-legible conversation, which is how a system that looks like it has a human backstop can quietly erode the backstop's ability to help. That is why the leverage this case points to is not a sharper chatbot — a better model still hands degraded context across a handover the citizen cannot distinguish from a person, still answers everyone identically, and still leaves the contained-but-wrong conversation unread — but three things the record shows are missing: a standing check on a public answer that has no caseworker to hold its verify habit, a review cadence that samples the *contained* population and not just the escalated one, and an independent read of the answers the machine completes. Frida's most useful contribution to the domain is to make the boundary explicit: the governance question is not only "is the answer right?" but "who sets how easy it is to reach a person, does anyone detect the conversations that ended wrong without one, and does the person on the other side inherit a conversation they can actually read?"
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.