Domain Atlas / Public benefits & eligibility
NYC MyCity business chatbot
Work with this case in the PAN Lab ↗
New York City launched the MyCity Business chatbot in 2023 on Microsoft Azure AI as a public-facing generative-AI adviser for business owners. A March 29, 2024 investigation by The Markup with THE CITY and Documented NY found it confidently and repeatably wrong on legal obligations, advising businesses in ways that would break the law, including that employers could take a cut of workers' tips, that landlords need not accept Section 8 vouchers or source-of-income tenants (illegal in New York City), that stores could go cashless against a 2020 city law, and that funeral-price disclosure could be concealed against the federal funeral rule; when ten staffers asked the housing-voucher question they received the same wrong answer, which had changed from an earlier correct one, showing the tool was non-deterministic. The 2024 findings are qualitative, based on specific tested questions rather than a sampled error rate. The city relabeled the tool a beta product with a disclaimer and applied a scope-narrowing patch rather than withdrawing it, kept it online for roughly two years, and shut it down in early 2026 as a budget cut rather than an accuracy fix.[4]
What happened
New York City launched the MyCity Business chatbot in 2023 — publicly announced in October, though the city's own December 2025 audit records a September 2023 launch — as a Microsoft Azure AI service meant to give business owners actionable, trusted guidance on city rules and processes. It produced free-text answers rather than adjudicating any eligibility decision, so its output was guidance, not a determination. On March 29, 2024, an investigation by The Markup with THE CITY and Documented NY found the chatbot routinely advised businesses in ways that would break the law: that employers could take a cut of workers' tips, that landlords need not accept Section 8 vouchers or source-of-income tenants (illegal in NYC), that tenant lockouts were permissible, that there were "no restrictions" on rent charges, that stores could go cashless against a 2020 city law, and that funeral-price disclosure could be concealed against the federal funeral rule. The reporting was qualitative — specific tested questions, not a sampled error rate. The bot was also non-deterministic: when The Markup had ten staffers ask the housing-voucher question, all ten received the same wrong answer, and the inconsistency was that an earlier test had returned the correct one, so the answer had changed over time rather than the ten disagreeing with each other.
Mayor Eric Adams declined to take the tool offline, framing it as a pilot ("It's wrong in some areas, and we've got to fix it"; "we're going to have the best chatbot system on the globe"). The city relabeled it a beta product, warned it might be inaccurate, and told users to double-check via provided links and not to treat responses as legal or professional advice — but the bot itself, when asked, still said it could be used for professional business advice, and the Office of Technology and Innovation asserted it had "already provided thousands of people with timely, accurate answers," a claim not independently corroborated. Microsoft said it was working with the city on fixes. About two weeks after the tips-theft error surfaced, the city applied a patch that redirected out-of-scope questions to NYC.gov, and by March 2025 then-CTO Matt Fraser said the city still planned to expand the chatbot beyond small business to cover all 311 content and that a post-error patch had reduced hallucinations "exponentially" — an unquantified claim.
A December 30, 2025 performance audit of the MyCity system, issued under Comptroller Brad Lander, found the chatbot "appears to be unable to provide accurate or consistent information." Its figures require care: an internal weekly production report reproduced in the audit showed the chatbot did not answer 23 of 48 tested government questions; of the more than 2,200 questions asked in July and August 2025, the 70 users who left thumbs-up-or-down feedback were 71.4% negative (50 of 70), a share the city disputes as roughly 2.25% of all responses (the audit rebuts that denominator), and the audit team's own testing documented the same question answered two different ways and outputs that changed with trivial edits like "NYC" versus "nyc." The audit also found the wider MyCity system had cost over $100 million across more than 120 agreements with about 50 vendors, lacked a system development plan, and had not delivered the promised single-form access to city benefits; OTI disagreed with all seven recommendations, including one to conduct AI red-teaming. Mayor Zohran Mamdani announced the chatbot's shutdown in late January 2026 as a budget cut, calling it "functionally unusable" — the roughly $600,000 development cost is corroborated in 2024 reporting, while the roughly $500,000-a-year maintenance figure and the "waste" framing come from the successor administration. As of early 2026 the chat.nyc.gov endpoint states "The Chatbot beta test has ended." A separate MyCity component, a childcare-benefits application flow, is not this chatbot; its documented figures (including a roughly 46% application-ineligibility rate) belong to that eligibility-automation portal, not to the adviser, and "deemed ineligible" is an application outcome, not proof of algorithmic error.
The sociotechnical reading
The Atlas already carries two careful benefits chatbots — the Nava-class assistive tool and the Imagine LA Benefit Navigator — and this is deliberately their opposite pole. Those two put a professional between the machine and the affected person: a caseworker reads the model's cited answer and verifies it before relaying, and their entire safety case is that one maintained habit. The MyCity chatbot removed the professional. The operator here is the public itself — a business owner, a landlord, an employer — asking directly and acting directly, in an official municipal voice with nothing in between. That single structural change is the whole lesson. There is no verify-before-use habit to protect because there is no professional to hold it, so the levers the sibling cases lean on (deskilling arrest, AI literacy, a vigilant channel over the caseworker) have no purchase. And the harm does not land on the person in the conversation. A landlord who is wrongly told he may refuse a voucher does not suffer; the tenant who is refused does, and the tenant was never in the room. This is the case where error externalizes onto third parties, so the usual governance question — who was wrongly flagged, and can they appeal — does not even apply. The chatbot flagged no one and denied no one; it handed out instructions, and the legal risk of following them fell on people outside the interaction with no channel to contest anything.
The sharper, distinct lesson is about what happened after the error was known. Every model is wrong sometimes; the defining failure here is institutional, not technical. The errors were published in a detailed investigation, catalogued in an incident database, and then documented again in a formal performance audit that made seven recommendations — and the owner rejected all seven, kept the tool online, and planned to expand it, until an unrelated budget cut ended it two years later. Exposure is not correction. An oversight signal that lands and is declined is not a control; it is a control the map draws as present and leaves running dry. That reframes where the leverage sits. Because there is no operator to train and no downstream appeal to strengthen, the productive moves are all upstream of the answer and above the owner's discretion: gate a generated answer against an authoritative legal source before it reaches the public (the check this deployment never had), label a confident free-text answer as generated and out of scope so its municipal authority is legible, hold a pre-committed take-down trigger so a documented-wrong adviser is paused on the evidence rather than defended until it is defunded, and make the external audit bind rather than advise. The monoculture makes each of these matter more, not less: one public-facing model answered everyone identically, so a single confident error was not scattered noise but the same wrong instruction served to every person who asked — the reason the map's model-to-model self-loop and its missing legal-truth check are the two edges this case turns on. The Practice Library's provenance, conformity-gate, and pre-authorized-halt patterns all trace back to a system that was wrong out loud, in public, for two years, with everyone able to see it and no one positioned — or willing — to stop it.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.