Domain Atlas / Housing & homelessness services
Imagine LA Benefit Navigator copilot
Work with this case in the PAN Lab ↗
In a 2025 Los Angeles County pilot evaluated by Nava Labs with academic partners at Cornell University and Georgetown University's Better Government Lab, a generative-AI assistive chatbot for Imagine LA's Benefit Navigator was estimated to improve benefits-navigation answer accuracy by an average of about 40% in a randomized controlled trial of 125 caseworkers answering hypothetical client questions, alongside a fourteen-week field pilot with 61 caseworkers across six organizations; the evaluation was co-authored by the tool builder rather than independently replicated, the accuracy figure is a decision-support contrast on hypothetical questions rather than a live-caseload eligibility audit, and time-savings and administrative-burden effects were reported as promising but inconclusive (published March 2026).[2]
What happened
Nava Labs, a division of Nava Public Benefit Corporation, built and piloted a generative-AI assistive chatbot layered onto Imagine LA's Benefit Navigator platform, the digital tool that benefit navigators and case managers use to connect low-income Los Angeles County families to public benefits (Imagine LA is a 501(c)(3) with more than eighteen years of direct services, and Amplifi handled platform integration). The copilot answers a caseworker's free-text questions about eligibility, application steps, documentation, and how programs interact across roughly forty LA County benefit programs and tax credits; it returns a generated summary answer plus direct-quote citations pulled from a curated corpus of approved government policy documents, is multilingual, and can restate an answer in simpler language or in Spanish. It is explicitly assistive and human-in-the-loop: caseworkers use it while on the phone with a client and verify the generated answer against its cited source before relaying it, and it does not make eligibility determinations on its own. The $3.6 million Gates Foundation grant announced in December 2024 funded the pilot; a subsequent $1.5 million Google.org grant in June 2025 funded a move toward AI agents that would navigate portals and complete applications under caseworker supervision, with Imagine LA and First 5 Riverside County. The evaluation, run with Cornell University and Georgetown University's Better Government Lab, combined a randomized controlled trial with 125 caseworkers (measuring the accuracy of being shown AI-generated responses to hypothetical client questions built from real experiences) and a fourteen-week real-world pilot with 61 caseworkers across six organizations in LA County. Use of the chatbot was estimated to improve answer accuracy by an average of about 40%, with the largest gains on the most difficult questions and among newer, less-experienced staff (a directional finding, not a quantified breakdown). About 65% of caseworkers with access used it, submitting an average of about 14 prompts each, at a modest, low-positive satisfaction — a Net Promoter Score of 11, with 40% classified as promoters; usage tended to decline over time and varied by site, so sustained adoption was reported to require ongoing organizational support. Answers averaged a tenth-to-twelfth-grade reading level, more accessible than the college-level source manuals but still short of a plain-language goal. Results on administrative burden and time savings were reported as promising but inconclusive and not statistically significant, owing to sample-size limits and a lower response rate. The work was published in March 2026 as a Nava case study and, with the same team, as the academic report "Helping the Helpers"; its authors were Allison Koenecke and Jennah Gosciak (Cornell), Eric Giannella and Zhaowen Guo (Georgetown Better Government Lab), and Michael Chen and Martelle Esposito (Nava), so the headline was co-authored by the tool builder together with its academic partners rather than independently replicated. Two figures often cited nearby belong to the underlying Benefit Navigator platform and not to the chatbot: an earlier 2023-2024 platform pilot reached more than 500 case managers, assisted over 10,000 beneficiaries, and secured on average an additional $10,869 in benefits per household. This case is distinct from the Atlas's earlier Nava assistive-benefits-chatbot entry, which covers Nava's earlier exploratory benefits-navigation work rather than this specific, co-evaluated Imagine LA pilot.
The sociotechnical reading
The Atlas already carries a Nava case as the cautious corner of the design space — retrieval-grounded, verify-before-use, a professional between the system and the affected person. This one is that same careful design put to an actual randomized test, and the test surfaces a subtler lesson the earlier entry could not. The measured benefit and the deference risk turn out to be co-located. The accuracy lift was biggest on the hardest questions and for the newest, least-experienced staff — which is exactly the population and the cases with the least independent capacity to catch the tool when it is wrong. A floor-raiser is a wonderful thing right up until the floor becomes load-bearing for people who cannot yet feel it shift. The system's only real safeguard is a single maintained habit, read the citation before you relay the answer, and the evaluation's own honest findings show that habit is not a settled property: adoption reached about two-thirds, then decayed without sustained engagement, and the fluent tenth-to-twelfth-grade answers sound more certain than the underlying policy warrants. So the governance work here is not the usual fight over who was wrongly flagged — this tool flags no one and denies no one. It is keeping a behavioral control alive: protecting the verify habit (deskilling arrest), calibrating trust for the staff who lean hardest (AI literacy), labeling the generated from the cited (provenance), and standing a vigilant channel where the checking would otherwise fade. There is a second, quieter lesson about evidence itself. The headline 40% is real, but it is a builder-and-academic-co-authored contrast on hypothetical questions, not an independently replicated audit of live eligibility decisions, and the time-savings everyone wants to claim were explicitly inconclusive. A good number authored by the people who built the tool is decision-support evidence, not a downstream-outcome promise, and treating it as the latter is its own governance failure. Finally, the case has a clock. The funded next step moves from an assistive chatbot a caseworker can second-guess on the phone toward agents that would complete applications under supervision — the moment the human gate that bounds error out of a real benefit decision is the thing being removed. The distinct lesson the Atlas draws here: when a tool's help concentrates where verification is weakest and its safety rests on a habit that decays, governance is the maintenance of a behavioral state, not a one-time design choice — and the most dangerous upgrade is the one that quietly retires the human who was doing the checking.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.