Domain Atlas / Benefits navigation & public-facing chat
Propel in-app SNAP benefits assistant
In a spring-2025 randomized pilot inside a large consumer EBT app, the vendor reports that 53% of eligible SNAP recipients took up in-app AI help for missed deposits and that treated users were restored faster and more often in the same month than a control group, with every AI dead-end escalated to a named human; all outcome figures are vendor-published and the effect magnitudes were not disclosed.[2]
What happened
Propel Inc., a private Brooklyn company, builds a free consumer app that more than five million people use to check their EBT (SNAP), WIC, and TANF balances by connecting to state benefit portals with the user's permission; NPR independently reports the roughly five-million-user scale, and the company states that over 25% of EBT cardholders nationwide use it (Apple's App Store listed about 574,000 ratings averaging 4.9 out of 5 as of July 2026). About 200,000 of those monthly users experience an interruption in their benefit deposits each month — driven chiefly by periodic reports due about six months after application and by annual recertification, the "program churn" failure mode Propel's AI tools target (a vendor-published figure).
In spring 2025 Propel ran two deliberately small pilots, each over a single deposit cycle. A multi-step self-diagnosis triage flow served about 1,300 CalFresh (California SNAP) recipients; its code was AI-generated and then manually edited before launch. A nationwide AI chat tool — built on a third-party generative-AI customer-support platform — assisted about 1,000 SNAP recipients with missed-deposit issues. Propel reports that 53% of eligible users offered the help tools used them, and that randomized testing against a control group showed "modest but meaningful" improvements on two pre-specified outcomes: days to the next deposit (faster restoration) and the share of users restored in the same month rather than churning and reapplying. The exact effect magnitudes were not published, and all of these figures are vendor-published, with no independent evaluation, replication, or third-party coverage of the pilots located. A documented escalation protocol was in place: when the chat detected it could not help, it handed off to a named Propel staff member who contacted the user directly, and the pilots were held to roughly 1,000-user scale precisely so the team could monitor results and absorb every escalation by hand. The chat generated state-specific responses, correctly recognized state-specific shorthand, and switched languages mid-conversation when a user began replying in Haitian Creole — multilingual behavior the team did not explicitly design in advance. Users could rate conversations; most did not, and among those who did there were very few negative ratings.
Propel positions these tools inside a documented practice of pre-deployment evaluation. An earlier prototype, published in January 2025, used a commercial language model with an empathetic, sixth-grade-reading-level prompt to explain SNAP notices, with PII redaction under consideration and off-ramps to human help; it remained in expert-feedback alpha and was never deployed to recipients. A September 2025 internal benchmark tested about forty-five commercial models on a state-dependent SNAP asset-limit question and found their answers improved from incorrect and discouraging in early 2024 to correctly identifying that 37 states plus the District of Columbia had eliminated asset limits by 2025. On September 29, 2025 the company launched an "AI Residency" R&D program, philanthropically funded into 2026, hiring former senior benefits administrators — a former California Medicaid Director, a former USDA Food and Nutrition Service senior advisor, and a former Colorado SNAP compliance officer — to help state agencies implement safety-net AI; this September announcement, not the July 2025 pilot write-up, is where the company said it is "ramping up" its investments. The pilot write-up itself avoided scaling commitments and recommended starting with small pilots.
The system's largest documented use came during the November 2025 federal shutdown that delayed SNAP payments to about 42 million recipients. Propel deployed an AI-assisted pipeline that monitored all fifty state SNAP agency websites hourly, using a language model to filter meaningful updates from formatting noise before human review; about 2.8 million users saw the official-state-updates button and more than 370,000 accessed state agency information through it over roughly two weeks, reaching, the company says, more than one in four SNAP households nationwide. In the same period Propel described a donation effort with partners including Robin Hood, GiveDirectly, and Babylist: NPR reported $10 million as the amount needed to send $50 payments to about 230,000 identified high-need users, with roughly $6 million raised at the time of publication (including $1 million from Propel itself). As of mid-2026 no source documents a full-scale production rollout of the AI chat assistant to the multimillion-user base; the deployments on record remain the capped pilots plus the shutdown pipeline.
The sociotechnical reading
Almost every cell in this Atlas watches an error flow from a model into a record and then back out into someone's decision — the contamination loop. This deployment is the Atlas's clearest instance of that loop being severed at the source, and it is worth naming why. The help surface grounds its central claim — that a deposit did not arrive — on a state-verified deposit record it reads but is architecturally forbidden to write. The one fact that carries the interaction is close to ground truth, externally checkable (the next deposit either lands in the state-verified EBT account or it does not), and the model's output cannot alter the authoritative record: it can only steer a recipient's own administrative action on the state system that owns it. Where other cells fight to keep a record from rotting under generated text, this one never lets the model touch the record at all. That read-only boundary, not any accuracy number, is the strongest control on the map.
Two things follow, and both are governance questions rather than accuracy questions. First, the safety case rests on a capacity-limited human channel. When the AI reaches a dead-end it hands off to a person, and the pilots were sized — about a thousand users — so that people could absorb every hand-off. That is a vigilance controller run by staffing, and its failure mode is not a wrong answer but a queue: at consumer scale the escalations outrun the humans, and the channel that made the pilot safe becomes a hiring plan. Second, the containment is almost entirely internal and discretionary. No state agency, regulator, or independent auditor governs the tool; the generative half runs on a third-party commercial platform outside the recipient-owned surface; and the only published evidence that any of this works is the company's own, with the effect magnitudes withheld. The distinctive lesson is the mirror image of the Atlas's usual one: when a system is built to read ground truth it does not control and to write nothing back, its risks migrate off the record loop and onto the two things scale and secrecy erode fastest — the human channel that catches what the model misses, and the outside scrutiny no one has yet been given. The honest boundary is that none of this measures the recipients themselves: the deposit that does or does not arrive, and any difference in who the tools help, sit outside a diagram like this one.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.