10 to the 23 AI logo

Domain Atlas / Caseworker documentation & copilots

Case fileUnited States (federal); Social Security Administration, Office of Hearings Operations and Office of Appellate Operations; nationwidegiant deployment

SSA Insight

The US Social Security Administration requires decision writers to run fully favorable disability decisions through its in-house Insight verifier before issuance, with narrow documented exceptions, and the 2025 federal AI inventory records the tool computing 43 quality flags. In the agency's internal five-month study of roughly 50,000 appeals-level cases, reported through the 2019 Inspector General audit, analysts who used Insight logged about 0.9 errors per case against 0.7 for non-users, saw processing time fall about 4.7 days per case, and had about 12.6 percent of their cases returned for quality issues against 21.5 percent for non-users. These are internal, non-randomized comparisons among self-selected voluntary users, and the same audit found the agency stopped tracking performance after the first five months and could not determine any effect on remands.[3]

What happened

Insight is an in-house Social Security Administration decision-support suite that verifies human-drafted disability decisions rather than drafting them. Devised and pitched in 2015 by Kurt Glaze, an SSA attorney with a computer-science interest who became its product owner, it grew from an "under the desk" proof of concept into an enterprise system supporting thousands of adjudicators. The direction is reversed relative to a drafting assistant: a decision writer or appeals analyst writes a decision from templates and adjudicator instructions, then runs Insight, which applies natural-language information extraction to the written text, joins structured case-management data (claim type, dates of birth, onset dates, claim history), and evaluates the decision through a hybrid of rule-based heuristics and probabilistic supervised classifiers. Its output is a set of real-time quality flags on potential inconsistencies, omissions, or policy-noncompliant conditions; SSA's 2025 federal AI inventory records 43 such flags, up from "over 30" around 2020. Insight is explicitly assistive: per the insider account co-written by its creator, it decides no element of a decision and advises no specific remedy, and its interface reminds users the content is "a jumping off point for further analysis." Each flag is admitted into the tool only if it exceeds an 80% accuracy threshold — an enumerable per-rule precision gate — and an internal two-flag study measured precision above 86%. The tool is classical predictive machine learning, not a generative model, and per the 2025 inventory it uses PII but no demographic features and was developed in-house with an authorization to operate.

The deployment scaled through 2017 and 2018: a voluntary Appeals Council pilot in January 2017 (46 analysts, 83% of those given access, on 660 hearing-level decisions, with more than 70% saying they would keep using it), a phased hearings rollout from September 26, 2017, and national completion to all decision writers in 2018. SSA requires decision writers to run Insight on fully favorable disability decisions before issuance, with narrow documented exceptions, while appeals-level use stays voluntary. The empirical anchor is a 2019 Inspector General audit (A-12-18-50353). In an internal five-month study of roughly 50,000 appeals-level cases, analysts used Insight 55% to 60% of the time and adjudicators about 25%; users logged 0.9 errors per case against 0.7 for non-users, saw processing time fall about 4.7 days per case after filtering confounds, and had 12.6% of their cases returned by appellate judges for quality issues against 21.5% for non-users. These are internal, non-randomized comparisons with self-selected voluntary users, reported by the agency rather than independently. Of 274 survey respondents who had used Insight, 74% found it useful and valuable and 71 commented on flag-accuracy problems. The audit's central governance finding was that SSA had stopped regularly tracking whether Insight met its goals — the Office of Appellate Operations stopped updating performance reports after the first five months because the work was manual — could not determine whether Insight reduced remands (the agency told the Inspector General in February 2019 that "Insight cannot guarantee a reduction in remands"), and, as of the audit, the agency's own policy office had not reviewed the flags for consistency with disability policy. SSA agreed with the recommendations to develop metrics, decide whether appeals-level use should be mandatory, and have the policy component review flag accuracy.

Two scholarly readings sit alongside the audit. The 2020 Administrative Conference of the United States report "Government by Algorithm" documents the technical anatomy (over 30 flags, the 80% admission gate, dependence on structured Findings Integrated Templates, OCR failure cascades) and names the specific failure mode of this reversed design: under high caseloads, adjudicators "may review cases solely to pass Insight quality flags, progressively ignoring errors that evade automated detection," and it proposes deactivating the tool for a random hold-out set as prospective benchmarking. The 2022 Oxford Handbook chapter, co-authored by Insight's creator and a former Deputy Executive Director of the Office of Appellate Operations, concludes that "AI governance is quality assurance" — continuous evaluation rather than one-off IT acquisition — while conceding that formal evaluations of Insight's impact on accuracy and remand rates have been limited. Insight remains listed as deployed and actively utilized in SSA's 2025 inventory, and SSA has kept layering AI onto the same hearings pipeline; a March 13, 2025 press release states a separate hearing-recording and automated-transcription platform was scheduled to complete nationwide rollout by March 17, 2025.

The sociotechnical reading

Almost every other copilot in this atlas points its model at the record: it drafts, and the human is the gate between generated text and the case file. Insight inverts that geometry. The human drafts, and the model is the gate — a verification layer standing on the adjudicator's own output edge, an institutionalized version of exactly the vigilant check the Field Guide keeps asking deployments to build. That makes it the atlas's clearest positive control, and it earns the label honestly: the 80% per-flag admission gate is the single most transferable idea in the collection, because it makes trust enumerable. A flag is not "the AI's opinion"; it is a named rule that cleared a stated accuracy bar, and a system that adds capability one accuracy-gated rule at a time is doing the thing most deployments only claim to do. The measured lift is consistent with that design working: the adjudicators who actually engaged the flags caught more and were returned less.

But a verifier has a different failure mode than a drafter, and the distinct lesson here is about that difference. A drafter fails by commission — a wrong sentence written into the record. A verifier fails by omission: not the flag that fires wrong, but the error class the flag-set never learned to see, which a trusted checklist renders invisible precisely because the tool is trusted. The ACUS warning is the sign flip of the whole loop — a mandate to RUN a check quietly becoming a mandate to SATISFY it, so the un-flagged errors stop being anyone's job, worst under exactly the caseload pressure this apparatus lives with. So governing a checker is not "is the model accurate?" It is three other questions the record answers plainly. First, audit the coverage of the flag-set, not just the precision of each flag — which is exactly the review the agency's policy office had not done. Second, keep the check a check: protect the reading-past-the-flags habit that produced the benefit, because it is the first casualty of a full docket. Third, and most quietly, do not let the evaluation that shows the tool is working lapse after launch — the creators' own thesis is that governance here IS continuous quality assurance, and the audit found the tracking went dark inside a year. There is a second-order tail the map keeps visible: every run also manufactures a new memory — structured flag and usage exhaust that today feeds management reporting and could tomorrow feed a model no one decided to train. The honest boundary is that none of this measures anything about the disabled people whose decisions are being checked; what a flag moves here is the quality of an institution's own paperwork, and whether it changed a single remand was, by the agency's own admission, never established.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

ussocialsecurityadministrati2019GroundingGovernment evaluationSave

US Social Security Administration, Office of the Inspector General, The Social Security Administration's Use of Insight Software to Identify Potential Anomalies in Hearing Decisions (A-12-18-50353) (2019) https://oig-files.ssa.gov/audits/full/A-12-18-50353.pdf

https://oig-files.ssa.gov/audits/full/A-12-18-50353.pdf

Grounds: model org: ssa_insight

glaze2022GroundingAcademicSave

Glaze, Ho, Ray, Tsang, Artificial Intelligence for Adjudication: The Social Security Administration and AI Governance (in The Oxford Handbook of AI Governance, Oxford University Press, 2022) https://dho.stanford.edu/wp-content/uploads/SSA.pdf

https://dho.stanford.edu/wp-content/uploads/SSA.pdf

Grounds: model org: ssa_insight

Topics: ai-governance

engstrom2020GroundingAcademicSave

Engstrom, Ho, Sharkey, Cuellar, Government by Algorithm: Artificial Intelligence in Federal Administrative Agencies (Administrative Conference of the United States, Stanford Law School, NYU School of Law, 2020) https://www.law.stanford.edu/wp-content/uploads/2020/02/ACUS-AI-Report.pdf

https://www.law.stanford.edu/wp-content/uploads/2020/02/ACUS-AI-Report.pdf

Grounds: model org: ssa_insight

stanfordreglab2022GroundingReferenceSave

Stanford RegLab, Artificial Intelligence for Adjudication: The Social Security Administration and AI Governance (publication page) (2022) https://reglab.stanford.edu/publications/artificial-intelligence-for-adjudication-the-social-security-administration-and-ai-governance/

https://reglab.stanford.edu/publications/artificial-intelligence-for-adjudication-the-social-security-administration-and-ai-governance/

Appears in: PAN framework development

Grounds: domain grounding: disability benefits adjudication and care allocation; model org: ssa_insight

Topics: ai-governance

ussocialsecurityadministrati2025GroundingGovernmentSave

US Social Security Administration, Social Security Announces AI Enhancements for Hearings Recordings (press release, 2025) https://www.ssa.gov/news/en/press/releases/2025-03-13.html

https://www.ssa.gov/news/en/press/releases/2025-03-13.html

Grounds: model org: ssa_insight

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalThe US Social Security Administration requires decision writers to run fully favorable disability decisions th…

The US Social Security Administration requires decision writers to run fully favorable disability decisions through its in-house Insight verifier before issuance, with narrow documented exceptions, and the 2025 federal AI inventory records the tool computing 43 quality flags. In the agency's internal five-month study of roughly 50,000 appeals-level cases, reported through the 2019 Inspector General audit, analysts who used Insight logged about 0.9 errors per case against 0.7 for non-users, saw processing time fall about 4.7 days per case, and had about 12.6 percent of their cases returned for quality issues against 21.5 percent for non-users. These are internal, non-randomized comparisons among self-selected voluntary users, and the same audit found the agency stopped tracking performance after the first five months and could not determine any effect on remands.

ussocialsecurityadministrati2019GroundingGovernment evaluationSave

US Social Security Administration, Office of the Inspector General, The Social Security Administration's Use of Insight Software to Identify Potential Anomalies in Hearing Decisions (A-12-18-50353) (2019) https://oig-files.ssa.gov/audits/full/A-12-18-50353.pdf

https://oig-files.ssa.gov/audits/full/A-12-18-50353.pdf

Grounds: model org: ssa_insight

engstrom2020GroundingAcademicSave

Engstrom, Ho, Sharkey, Cuellar, Government by Algorithm: Artificial Intelligence in Federal Administrative Agencies (Administrative Conference of the United States, Stanford Law School, NYU School of Law, 2020) https://www.law.stanford.edu/wp-content/uploads/2020/02/ACUS-AI-Report.pdf

https://www.law.stanford.edu/wp-content/uploads/2020/02/ACUS-AI-Report.pdf

Grounds: model org: ssa_insight