Customer service & contact-centre AI
AI in the contact centre takes two forms, and they fail in different ways. Agent-assist copilots draft suggested responses for a human agent, and the strongest field evidence shows their benefit is real and unevenly distributed: a large randomized rollout raised issues-resolved-per-hour by roughly 15 percent on average, but almost the entire gain went to novice and low-skill agents while the most experienced gained close to nothing — so an average productivity number overstates the effect for the agents who least need it and hides that a copilot mostly raises the floor. Customer-facing chatbots answer the person directly, and here the governance facts are sharper. The organization is accountable for what its chatbot tells a customer: a tribunal has held that a chatbot is a tool the company is responsible for, not a separate entity that answers for itself, so a hallucinated policy or a wrong fare rule is the company's misrepresentation. Deflection is not resolution: routing a contact away from a human can suppress the very signal the customer needed to send, and surveys show most customers would rather not meet AI in service at all and fear it makes reaching a human harder — which makes the escalation path to a person the system's real safety valve. And a published deflection number is not a settled result: the same organization that reported handling two-thirds of its chats with AI later walked back on quality grounds and committed to keeping a human available. The Lab networks model only the deploying organization — its agents, its QA and escalation functions, and its interaction records; the customers being served sit outside the dynamics, and no customer outcome is computed on any diagram.
Use cases
What AI is doing here
Agent-assist copilot
GenerativeA generative copilot that drafts suggested responses for a human support agent — where the strongest field evidence shows the productivity gain concentrated in novice and low-skill agents and close to zero for the most experienced, so an average overstates the effect and the agent who accepts a suggestion is also its reviewer of record.
Customer-facing service chatbot
GenerativeA chatbot that answers the customer directly and deflects contacts from human agents — where the organization is accountable for what the bot says (a tribunal has held the bot is a tool, not a separate entity), a hallucinated policy is the company's own misrepresentation, and deflection is not the same as resolution.
QA, coaching & escalation to a human
PredictiveThe quality-assurance, coaching, and escalation-desk functions around a deployed contact-centre AI — the human safety valve customers say they want, where the governance question is whether the path to a person is protected and whether deflection is measured against resolution and repeat contact rather than counted as a win by itself.
Case files
What has gone wrong and right
Documented deployments, presented as model organizations calibrated to the evidence, with full citations.
The copilot that raised the floor
United States (large software firm's customer-support operation; peer-reviewed randomized field study)The contact-centre domain's strongest measured benefit comes from a staggered randomized rollout of a generative-AI agent-assist copilot to roughly 5,000 customer-support agents at a large software firm. Against a control group, the copilot raised issues resolved per hour by about 15 percent on average and improved customer sentiment and agent retention. The gain was sharply uneven: novice and low-skill agents improved by roughly 30 to 34 percent, an agent with two months of experience performed like one with six months and no AI, and the most experienced agents gained close to nothing — with some evidence of slight quality degradation. The benefit is real and it is a distribution: a copilot that mostly raises the floor, so an average overstates the effect for the agents who least need it.
Explore this deployment in the PAN Lab →The organization answers for what its chatbot says
Canada (British Columbia Civil Resolution Tribunal; decided ruling with published reasons)An airline's customer-facing website chatbot told a customer they could claim a bereavement fare retroactively — a policy that did not exist. The customer relied on the statement, bought a ticket, and was then refused the fare by human staff. A civil-resolution tribunal found the airline liable for negligent misrepresentation and awarded damages, rejecting the airline's argument that the chatbot was a separate legal entity responsible for its own actions. The tribunal held the organization responsible for all the information on its website, whether from a static page or a chatbot, since a customer has no way to know which source to trust. The contact-centre domain's cleanest accountability ruling: the bot is a tool the company answers for, not an entity that answers for itself.
Explore this deployment in the PAN Lab →The deflection numbers and the walk-back that followed
Multinational (company-published deployment figures and a later company-reported reversal)Klarna published striking first-month numbers for its customer-facing AI assistant: it handled about two-thirds of chats (some 2.3 million conversations), was described as doing the work of about 700 full-time agents, cut average resolution time from about 11 minutes to under 2, was said to match human customer satisfaction, and was projected to improve profit by tens of millions — all self-reported and not independently audited. About a year later the same organization reversed on quality grounds — its chief executive said cost had become too predominant an evaluation factor and the result was lower quality — and committed to always keeping a human available. The contact-centre domain's cleanest benefit-then-cost arc: the deflection numbers and the walk-back come from the same deployment.
Explore this deployment in the PAN Lab →System map
Who is in the system and what pushes on it
Who is in the system
- Frontline workers. Caseworkers, screeners, eligibility staff — the operator network whose judgment the system augments or erodes.
- Supervisors & QA. The institutional correction layer: overrides, second reads, quality review.
- Agency leadership. Owns procurement, policy, and the authority map; answers for the system publicly.
- Served people & families. Those the decisions land on. Deliberately outside the PAN dynamics — their outcomes are measured, never simulated.
- Vendors. Build and update the systems; hold the information asymmetry procurement must govern.
- Regulators & oversight bodies. Boards, auditors, data-protection officers, inspectorates — external correction capacity.
Dominant pressures
- Reviewer bottleneck. One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Austerity & recovery incentives. Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Vendor opacity. The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift. The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance. Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
Governance
Questions leaders should be asking
- 1. The measured benefit of an agent-assist copilot is concentrated in novice agents and near-zero for experienced ones — so is the deployment reporting an average that overstates the effect for the people it helps least, and is anyone watching for the slight quality cost the same studies flag for experts?
- 2. A tribunal has held that a company is responsible for what its chatbot tells a customer — the bot is a tool, not a separate entity — so when the chatbot states a policy or a price, is the organization treating that as its own representation, with the accuracy controls that implies, or as something the vendor's model said?
- 3. Deflecting a contact is not the same as resolving it, and a deflected customer may have needed to send exactly the signal the bot absorbed — so is deflection being measured against resolution and repeat contact, and is the path to a human protected as the safety valve rather than designed to be hard to reach?
- 4. The same organizations that publish striking deflection numbers have later reversed on quality and put humans back — so is the deflection figure being read as a settled result or as a claim that has to survive a look at quality, escalation, and whether customers wanted the AI at all?
For the actions behind these questions, see the Practice Library.
Seeing your organization in this domain? Mapping its actual pathways, pressures, and correction capacity is engagement work.
Work With 1023AI