10 to the 23 AI logo

Domain Atlas / Benefits navigation & public-facing chat

Case fileUnited States (federal); Internal Revenue Service, Small Business/Self-Employed Division; Automated Collection System chat applications audited at the Brookhaven (Holtsville, NY) and Philadelphia campuses and the Office of Online Services (Lanham, MD)large deployment

IRS collection chatbots: expanded and made permanent with no performance measures

A June 2026 Treasury Inspector General for Tax Administration performance audit (Report Number 2026-308-029) reported that the IRS expanded its Automated Collection System chatbot and live-chat program and made live chat permanent while having no performance measures for it, despite a Taxpayer First Act requirement for metrics and benchmarks, and that management's claim the bots reduced telephone demand could not be substantiated; the statistical reports the IRS did collect were deemed unreliable, in one instance showing a single assistor apparently working 603 chats at once against a systemic cap of three, attributed partly to a miscalculated handle-time metric the vendor had not resolved as of December 2025.[3]

What happened

The Automated Collection System (ACS) function within the IRS Small Business/Self-Employed (SB/SE) Division — which resolves balance-due and delinquent-return inquiries — was the first IRS function to pilot chat applications, with the stated goal of directing taxpayers to online self-help tools instead of the toll-free phone line. The deployment came in stages: an unauthenticated live-chat pilot in November 2017; authenticated live chat in June 2019 (where, after identity verification under Internal Revenue Code Section 6103, assistors can take account actions such as setting up payment plans); a scripted ACS chatbot in December 2021 (English and Spanish, preprogrammed responses, explicitly non-generative); an unauthenticated voice bot in January 2022; and an authenticated voice bot in June 2022. The live-chat pilot was made permanent in June 2025.

On June 15, 2026, the Treasury Inspector General for Tax Administration (TIGTA) issued a performance audit, "Opportunities Exist to Improve the Quality of Chat Applications" (Report Number 2026-308-029), covering ACS chat statistical reports from February 2023 through December 2024, with fieldwork from August 2024 to September 2025; voice bots were explicitly excluded from testing. Its central finding is a governance absence rather than a decision error: the IRS had not analyzed its chat applications and had no performance measures for them, despite the Taxpayer First Act's requirement for metrics and benchmarks — yet the program was expanded and made permanent anyway. SB/SE officials' internal claim (in a Q3 FY2023 Business Performance Review) that chatbots reduced telephone demand could not be substantiated.

The statistics the IRS did collect were unreliable. Reports showed a single live assistor working as many as 603 chats concurrently, although a systemic control caps concurrent chats at three; more than 16,000 records showed potential concurrency and 193 records showed 4 to 603 concurrent chats. IRS IT attributed this partly to a miscalculated handle-time metric — chats left open auto-close after four hours — and the vendor's root-cause investigation was still unresolved as of December 2025; the defect affects all IRS chat programs on the platform, including the Automated Underreporter Program Live Chat. Resolution-code records did not reconcile with chat counts either: 635,684 resolution codes against 613,056 live chats (a 22,628-record excess, from assistors closing chats with no code or multiple codes), a mismatch ACS management knew about but did not investigate. Only 290,181 (46%) of 635,684 reported chats were classified as resolved; 345,503 (54%) went unresolved — 136,838 abandoned, 127,288 disconnected, and 81,377 referred to the ACS phone line — with abandoned chats roughly tripling from 33,781 (CY2023) to 103,057 (CY2024). Because reasons for non-resolution are not documented, the IRS cannot determine why chats go unresolved. An April 2025 platform update restricted assistors to one chat at a time, but September–November 2025 reports still showed 181 records of 23 assistors apparently working 2 to 11 concurrent chats — the reports will keep showing false values until the handle-time calculation is fixed.

A second finding couples operator load to disclosure risk. Of a judgmental sample of 40 live assistors, 24 (60%) worked multiple live chats concurrently, and 12 of those 24 (50%) had at least one authenticated chat open while working another — raising the risk of inappropriately disclosing taxpayer information to the wrong taxpayer. Management believed authenticated chats were worked one at a time; two floor assistors interviewed during walkthroughs at two different ACS locations said they were allowed to work multiple, authenticated or not. Chatbot content testing in March 2025 found 29 of 206 process flows (14%) needing improvement (broken hyperlinks, forms without instructions) and 44 of 53 tested keywords (83%) — including "Debit Card" and "Tax Pro Account" — unrecognized or producing insufficient responses, in some cases returning raw computer programming code; the Spanish-language chatbot lacks the free-text search box entirely, limiting Spanish speakers to prepopulated options.

The feedback loops that would have caught any of this were dead. No survey has been offered after live chats since September 2022; the two-question chatbot survey tells taxpayers "We are collecting the information to help improve our service," but IRS officials stated they do not collect or review the responses, and a replacement live-chat survey was, as of December 2025, still awaiting submission to the Office of Management and Budget for clearance — leaving the IRS noncompliant with OMB Circular A-11's customer-experience requirements for designated services. Human oversight was similarly notional: ACS management reviewed only one chat per assistor per month, nonevaluatively (the reviews cannot be used in performance appraisals), on the rationale that live chat was a "pilot," though no policy prohibits evaluative reviews during pilots, and the Internal Revenue Manual had no guidance for managing live chat. TIGTA made nine recommendations — validate data accuracy and completeness; set quantitative performance standards; fix the handle-time programming error; enforce single non-blank resolution codes; implement and assess a live-chat survey; add written feedback to the chatbot survey; conduct evaluative assistor reviews; issue management policies; and fix the chatbot's links and word bank and add Spanish keyword search. IRS management agreed with all nine, describing several (such as evaluative reviews and live-chat management guidance) as already implemented and the rest as planned; the audit's own wording is hedged and does not enumerate which were complete.

The context is one of contraction, not growth. The IRS spent about $7.2 million of Inflation Reduction Act funds on chatbot and live-chat applications in FY2024–25, even as IRA funding overall was cut from $79.4 billion to about $26 billion; between January and May 2025 the IRS workforce fell from about 103,000 to 77,000 (25%), and ACS live assistors fell about 20% (159 to 127) as of July 2025. The agency's long-term AI strategy is on hold indefinitely after its Transformation & Strategy Office was eliminated in March 2025, pending a new Treasury Department strategy. Two provenance boundaries frame the record. First, the IRS's own early performance claims — a September 2022 "A Closer Look" essay (by SB/SE Deputy Commissioner Darren Guillot; the page is marked historical, last reviewed or updated February 3, 2026) reporting 450,000-plus chatbot interactions at 42% resolved without escalation, 4.8 million unauthenticated and 1 million-plus authenticated voice-bot calls at about 40% containment, and 7,600 installment agreements covering over $50 million — are unaudited agency self-claims, and TIGTA specifically found the IRS could not substantiate its containment claims. Second, an earlier independent voice had already flagged the shape: the IRS Advisory Council's November 2024 annual report found the bots "were not designed to provide the taxpayer a direct answer" but referred taxpayers to general information, and that multiple bot vendors create "parallel bot channels, multiple entry points, [that] may create taxpayer confusion," recommending single entry points, live-agent escalation, and testing with non-English speakers and users with disabilities. The chat-platform vendor is unnamed in every public source.

The sociotechnical reading

Most of the Atlas is about decisions: a fraud score, an eligibility cutoff, a risk tier, a generated answer someone might act on. The IRS ACS chatbot decides nothing — it is scripted, explicitly non-generative, and its live-chat channel is staffed by humans. That is exactly what makes it the instructive inversion, and the lesson is about criticality without measurement. Here is a federal-scale deflection funnel — hundreds of thousands of live chats, millions of claimed voice-bot calls — expanded and made permanent while the agency had, in the inspector-general's finding, no performance measures for it at all. On the system map, the failure is not a bad edge but a severed one: the feedback link from the platform's telemetry to the people who govern it is drawn in place but functionally cut. And the case's sharpest point is that a link which is formally present but functionally severed is more dangerous than one that is visibly missing, because its corrupt readings get treated as fact. Telemetry that reports a single worker handling 603 simultaneous chats against a hard cap of three is not merely useless; it launders an ungoverned expansion as measured success — management read those numbers, knew they did not reconcile, and did not investigate. A blank space invites a question. A confident wrong number answers it.

The domain gives the contrast in the cleanest possible form. The Atlas's other public-facing government chatbot ran the same surface the opposite way: a staged pilot gate that held an early version below its accuracy bar and published its metrics before any wider release — measurement as the thing that governs deployment. The IRS cell is that mechanism's negative image: the gate absent, the metrics unreliable, the surveys unread, the expansion decided anyway. And where the domain's positive navigation node was made safe by an external floor that is continuous and load-bearing — a statutory determination made on every case — the only functioning check here is external and episodic: an inspector-general audit that fires once in years and had to discard most of the very data it was auditing to hand-verify a narrow band. An audit is a fire alarm, not a smoke detector; a system that can only be corrected when an outside body happens to look is a system running blind between glances.

There is a second dynamic the map carries, and it is the one that shows why the missing measurement is not a paperwork problem. Automation deployed to deflect demand instead reloaded the human channel: most sampled assistors worked several chats at once, and half of those had an authenticated session open alongside another, raising the documented risk of surfacing one taxpayer's account inside a session opened for someone else — federal tax data, under statutory confidentiality, exposed by concurrency the broken telemetry could not even see. So the harm the automation created was precisely the harm the severed feedback link hid. The governance instruments that matter for a cell like this are therefore not the ones the rest of the domain reaches for. No verifier, no risk tiering, no smarter model touches a system that makes no decision; a sharper chatbot would not reconcile a report or unload an assistor. What bites is the restoration of measurement itself: a review cadence someone can be held to, telemetry labeled as unreliable so no one plans on it, corrupt writes gated at the record, the derived reports reconciled against the sessions that actually happened, and the authenticated session bound to one taxpayer at a time. The distinct lesson the Atlas draws here: the most dangerous failure in a deployed system is not the wrong answer but the unmeasured one, and a monitoring channel that reports confident nonsense is worse than none — because you cannot govern what you have decided, in effect, not to see, and you will not even know what it is costing you. The honest boundaries throughout: served taxpayers, and the obligations they do or do not resolve, are not modeled in the paired Lab; the concurrency figures are computed from data the audit itself calls unreliable and a nonprobability sample it says cannot be projected; voice bots were untested; the 2022 volume claims are unaudited agency self-reports; and this was never a generative-AI or scoring system — any reading that implies algorithmic decision-making misreads it.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalA June 2026 Treasury Inspector General for Tax Administration performance audit (Report Number 2026-308-029) r…

A June 2026 Treasury Inspector General for Tax Administration performance audit (Report Number 2026-308-029) reported that the IRS expanded its Automated Collection System chatbot and live-chat program and made live chat permanent while having no performance measures for it, despite a Taxpayer First Act requirement for metrics and benchmarks, and that management's claim the bots reduced telephone demand could not be substantiated; the statistical reports the IRS did collect were deemed unreliable, in one instance showing a single assistor apparently working 603 chats at once against a systemic cap of three, attributed partly to a miscalculated handle-time metric the vendor had not resolved as of December 2025.

EmpiricalIn the same audit, of a judgmental sample of 40 IRS ACS live assistors, 24 (60%) were found working multiple c…

In the same audit, of a judgmental sample of 40 IRS ACS live assistors, 24 (60%) were found working multiple chats concurrently and 12 of those 24 had at least one authenticated chat open while working another, which TIGTA reported as raising the risk of disclosing taxpayer information to the wrong taxpayer; the audit also reported 635,684 resolution codes against 613,056 chats (a mismatch management knew of but did not investigate) and, in March 2025 hand-testing, 14% of chatbot process flows deficient and 83% of tested keywords unrecognized or insufficient, with the figures drawn from a nonprobability sample and data the audit itself characterized as unreliable and not projectable to the full assistor population.