Domain Atlas / Customer service & contact-centre AI
The deflection numbers and the walk-back that followed
Explore this deployment in the PAN Lab ↗
An organization published striking first-month numbers for its customer-facing AI assistant: it handled about two-thirds of customer-service chats (some 2.3 million conversations), was described as doing the equivalent work of about 700 full-time agents, cut average resolution time from about 11 minutes to under 2, was said to match human customer satisfaction, and was projected to improve profit by tens of millions. Every one of those figures was self-reported and not independently audited. Roughly a year later the same organization reversed course on quality grounds — its chief executive said cost had become too predominant an evaluation factor and the result was lower quality — and committed to always keeping a human available to customers who want one. This is the contact-centre domain's cleanest benefit-then-cost arc: the deflection numbers and the walk-back come from the same deployment.[2]
What happened
Klarna deployed a customer-facing AI assistant across its support operation and, in its first month, published a set of numbers that made the deployment famous. The assistant was said to have handled about two-thirds of the organization's customer-service chats — some 2.3 million conversations — to do the equivalent work of about 700 full-time agents, to have cut average resolution time from about eleven minutes to under two, to match human customer satisfaction, to reduce repeat inquiries by about a quarter, and to be worth tens of millions in profit improvement. It is important to be precise about what these figures are: they are the organization's own report, self-published and not independently audited. They are the deployer's dashboard, and they enter the record as exactly that.
About a year later the same organization walked the story back, and the reversal is what makes the case valuable. Its chief executive said publicly that cost had become too predominant an evaluation factor in the AI push, and that "what you end up having is lower quality" — and the organization committed to always keeping a human available to customers who want one, piloting a return to human agents. Nothing here says the original deflection numbers were false; the assistant very likely did handle two-thirds of chats. What the walk-back says is that handling two-thirds of chats was the wrong thing to have optimized, because the metric that was maximized — deflection, on cost — was not the metric that mattered, which was whether customers were well served.
That is the structural lesson: deflection is not resolution. A deflection number reports how many contacts the AI absorbed, not whether it resolved them well, and the two can move in opposite directions. An assistant that deflects more contacts at lower cost can be doing so by giving faster, cheaper, worse answers — and because the deflection number looks like a win, the quality cost stays invisible until something surfaces it. Here what surfaced it was the organization's own later judgment that quality had fallen, not a measurement built into the deployment from the start.
The survey backdrop makes this more than one organization's course-correction. Most customers say they would rather not meet AI in customer service at all, and a majority fear it makes reaching a human harder; industry analysts expect a large share of organizations to abandon plans to reduce their customer-service workforce with AI. Against that backdrop, the honest reading of this deployment is that the benefit numbers were real to the organization and reported in good faith, that they measured deflection and cost rather than resolution and quality, and that the safety valve the walk-back restored — always a human available — is exactly the thing a deflection-maximizing design tends to erode. The governable move is to measure resolution and repeat contact against deflection, and to protect the path to a person, before a reversal has to teach it.
The sociotechnical reading
This case is the contact-centre domain's benefit-then-cost anchor, and its instruction is to read a published deflection number as a claim that has to survive a look at quality, not as a settled result. The organization's first-month figures were striking and self-reported; its later reversal, on its own account, was about quality falling because cost had become the dominant metric. Both come from the same deployment, which is what makes the arc clean: the same numbers that looked like a decisive win were, a year later, described by the deployer as the wrong thing to have optimized.
The transferable lesson is that deflection is not resolution. Deflection measures how many contacts the AI absorbed; resolution measures whether the customer's problem was actually solved; and an AI can raise the first while lowering the second — faster, cheaper, worse. Because the deflection number reads as success, the quality cost is invisible until something external surfaces it, so the governable design measures resolution and repeat contact alongside deflection rather than counting deflection by itself. The check drawn latent here is exactly that: the resolution-and-quality measurement that a deflection headline does not contain.
The escalation path is the second surface, and it is the one the walk-back restored. The organization's commitment to always keep a human available is the safety valve customers say they want, and a deflection-maximizing design tends to erode it — the more the goal is to keep contacts away from people, the harder the path to a person becomes, which surveys show customers both notice and resent. The map's instruction is that the human escalation route is not a failure of the AI deployment but a designed part of it, and that protecting it is a governance choice made before a reversal forces it, not after.
The Lab network models only the deploying organization: its assistant, its human agent function, and its interaction records. No customer outcome is computed on any diagram. The customers being served are boundary-only; the deflection figures, the self-reported satisfaction parity, the later quality reversal, and the survey backdrop are institutional signals that live in this case file, never on any network. Because the benefit figures are the deployer's own self-report and the reversal is the deployer's own later judgment, both are drawn as claims the deployment made about itself, not as measurements the diagram computes. The map's instruction is to read the deflection numbers as a real but unaudited dashboard, to measure resolution against them, and to keep the path to a human open as the safety valve the arc shows a deflection-first design tends to lose.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.