The Lab speaks in pathways, pressures, levers, and gauges. This map connects that vocabulary to the formal failure-mode names used in the research grounding it — including a 2026 national survey of 1,179 U.S.-based social workers.
Automation bias / overreliancecore
The “Failures adopted by people or agents” pathway and the operator-deference-drift gauge. Staff turnover and autonomy expansion push it up; the deskilling-arrest lever caps it.
Evidence: In the same national survey, 40.8% of respondents reported ethical concerns about relying on AI for decision-making, and overreliance on automated decision-making was among the most frequently cited concerns overall.
Deskilling / professional judgment erosioncore
Overreliance in slow motion: the deference gauge drifting upward while correction capacity thins. The deskilling-arrest lever is its deliberate counter-schedule.
Sycophancy / agreement-seeking outputcore
The pushback-heavy-usage pressure runs the “Operator framing biases the model” pathway hot and makes the agreeable answers easier to adopt; the framing-hygiene lever dampens the loop at its origin.
Evidence: Research on AI sycophancy describes it as a fragmented construct — a family of distinct agreement-seeking behaviors that share a label but differ in form, mechanism, measurement, and required mitigation — and finds it intensifies under user pushback and across multi-turn interaction.
Hallucination / incorrect-output propagationcore
The Lab's core premise: every pathway in the diagram carries incorrect output away from its source, and every lever is a way of governing that propagation rather than assuming a perfect model.
Unsafe data flow / privacy & confidentialitycore
Modeled as pathways, gauged as exposure (Phase NP). Unsanctioned tool use opens a visible egress to the off-network sink; connector sprawl replicates an unverified cache between record systems; case-file-flagged pathways (MiDAS-class enforcement replication, records feeding vendor models) carry the same concern. While any such pathway runs, the Privacy gauge drains — and in Hard and Expert a full gauge is part of the win. “Vet connections” cuts the pathways structurally; “Store less data” shrinks what there is to expose.
Evidence: In the same national survey, concerns about data privacy and security were the most frequently reported challenge to using AI in practice (46.5% of respondents), and an increased focus on client privacy and confidentiality was the most requested improvement to AI tools for social work (50.4%).
Bias propagation / institutional workflow biascore
Biased framings and contaminated records travel the same workflow pathways failures do — an institutional propagation question, and the documented cases show the workflow (not the model alone) carrying the equity outcome in both directions. The Lab models no demographics: differential harm to served people is recorded outside the network, never computed from its dynamics.
Evidence: Evaluation evidence on the Allegheny Family Screening Tool found that screener overrides of the tool's recommendations reduced racial disparity in screen-in rates relative to the tool alone.
Evidence: Independent scrutiny of Rotterdam's welfare-fraud risk model — a 2021 municipal audit followed by a 2023 journalistic investigation that obtained the model itself — documented scores skewed against already-vulnerable groups, and the city suspended the system's use.
Transparency / provenance failurecore
The record-contamination gauge reads how much unlabeled machine content feeds back into decisions; “Mark AI-written records” discounts it and “Gate vendor updates” attacks opacity at procurement.
Weak human oversight / safeguard failurecore
The scenario axis itself: one model, three oversight cultures, three very different outcomes. The correction and authority gauges track it; “Review the riskiest first”, “Pause AI on alarms”, “Require sign-off”, and “Review on schedule” govern it.
Evidence: In the same national survey, 42.1% of respondents reported having no role in decision-making about AI adoption in their workplace; the report concludes most respondents have limited or no control over how AI technologies are selected or implemented within their organizations.
Low AI literacy / verification readinesscore
The AI-literacy-gap pressure: verification skill (not time) drops and deference rises as trust calibrates on the tool itself. The correction-capacity gauge reads the result. The grounding literature proposes AI literacy as a core professional competency.
Evidence: The same national survey describes a gap between AI exposure and AI preparedness: 26.6% of respondents cited lack of training or understanding of AI technology as a challenge, 53.4% said training on AI tools and effective use would help, and clear guidelines on the ethical use of AI were the most-endorsed need (66.8%).
Evidence: AI literacy — the knowledge and skills required to understand, use, and critically evaluate AI systems — has been proposed as a core competency for social work, relevant even to practitioners who never directly use AI tools.
AI iatrogenics / governance backfireadvanced
First-class here: purging records without reading them backfires exactly as the published runs found, and the efficiency readout will call a lever stack counterproductive to its face. The quieter trap — oversight whose gains are bought by rising deference — is why the deskilling-arrest lever exists.
Evidence: In the published runs, deleting records without reading them raised the contaminated share by stripping out benign entries; only content-aware cleanup reliably reduced it. (PAN governance-lever audit)
Monitoring failure / drift blindnessadvanced
Two pressures carry it: the silent vendor update (drift arriving under controls tuned to old behavior) and monitoring going stale (dashboards nobody must act on — the authority gauge hollows while the regime worsens). “Review on schedule” and “Escalate checks” are the counters.
Service starvation / over-throttlingadvanced
The sixth iatrogenic, symmetric to deference-load: governance so tight the work stops. The “Work getting done” gauge reads strained and the Net AI benefit track sits on the hurting side while the failure regime reads contained — a breaker or write-gate has zeroed an assistive arm, or checking layers have throttled it to a satisficing region where the tool is compliant but not constructive. It is a Lab-only service quantity, distinct from PAN's harm-reduction: you paid budget, deference, and compute for a tool your governance won't let help.
Evidence: Safety-only alignment establishes a behavioral floor without a ceiling: systems can be 'not-unsafe' yet directionless — compliant without being constructive — and benefit must be assessed as scaffold versus crutch.
Evidence: In a two-year child-welfare ethnography, an ill-fitting algorithmic tool imposed ongoing repair work on caseworkers — anticipatorily editing the inputs they supplied so the tool would return a usable result, and bending or working around procedure to reconcile its output with the case in front of them — labor spent making a poorly-suited tool usable rather than on the casework itself, distinct from any deliberate checking of the output.
Repair work / tool-usability laboradvanced
The drag the Work-quality component reads when a contaminated or ill-fitting tool imposes ongoing labor to make its output usable — anticipatory editing of inputs, bending procedure to reconcile the result — spent making the tool usable rather than on the casework itself, distinct from the deliberate checking priced as review latency.
Evidence: In a two-year child-welfare ethnography, an ill-fitting algorithmic tool imposed ongoing repair work on caseworkers — anticipatorily editing the inputs they supplied so the tool would return a usable result, and bending or working around procedure to reconcile its output with the case in front of them — labor spent making a poorly-suited tool usable rather than on the casework itself, distinct from any deliberate checking of the output.
Environmental burden (external context)external context
Deliberately outside the network: nothing in this Lab computes environmental cost, and no gauge claims to. Practitioner concern about AI's environmental impact is documented in the survey evidence and belongs in deployment governance as external context.
Evidence: In the same national survey's open-ended comments, ethical concerns — prominently including the environmental impact of AI infrastructure — were the most common theme, and the report's first recommendation includes environmental impact among the topics profession-wide ethical guidance should address.