Domain Atlas / Customer service & contact-centre AI
The copilot that raised the floor
Explore this deployment in the PAN Lab ↗
The strongest field evidence for an agent-assist copilot in customer service comes from a staggered randomized rollout of a generative-AI assistant to roughly 5,000 customer-support agents at a large software firm. Measured against a control group, the copilot raised issues resolved per hour by about 15 percent on average, and it also improved customer sentiment and agent retention. The gain, however, was sharply uneven: novice and low-skill agents improved by roughly 30 to 34 percent, agents with two months of experience performed like agents with six months and no AI, and the most experienced agents gained close to nothing, with some evidence of slight quality degradation. This is the contact-centre domain's cleanest measured benefit, and it is a distribution rather than a single number.[2]
What happened
A large software firm rolled out a generative-AI agent-assist copilot to its customer-support agents, and because the rollout was staggered it could be studied as a randomized field experiment against a control group — the strongest causal evidence the contact-centre domain has. The headline result is a genuine benefit: agents with the copilot resolved about 15 percent more issues per hour on average, and the deployment also improved customer sentiment and agent retention. On its own that is a clear win, and it should be read as one.
The result that matters for governance is the distribution beneath the average. Almost the entire gain accrued to less-experienced agents: novice and low-skill agents improved by roughly 30 to 34 percent, and an agent with two months on the job performed like an agent with six months and no AI. The most experienced agents, by contrast, gained close to nothing, and the study found some evidence of slight quality degradation for them. The copilot, in other words, mostly raised the floor. It compressed the skill distribution by pulling novices up toward the level experienced agents had already reached on their own, and it did comparatively little — possibly slightly negative — for the agents at the top.
That is why the average is the wrong number to govern by. A 15 percent mean improvement describes almost no individual agent: it overstates the effect for the experienced agents who least need it, understates it for the novices who benefit most, and hides the small quality cost at the top entirely. An organization that reports the mean and staffs against it is managing a benefit it has not actually measured, because the thing it deployed does different things to different agents, and only a distribution shows that.
The honest reading is that this is a real and specific benefit, not a uniform one. A copilot that helps new agents a great deal and helps veterans not at all is worth deploying — it shortens ramp time, it raises the experience of being a customer talking to a new agent, and it may improve retention precisely because it makes hard early months easier. But describing it as "15 percent more productive" misstates who it helps, and treating that average as the deployment's value invites two mistakes at once: expecting gains from experienced agents that will not come, and missing the slight quality cost the same evidence flags for them. The benefit is a distribution across skill, and it has to be measured as one.
The sociotechnical reading
This case is the contact-centre domain's anchor because it supplies the rare thing — randomized causal evidence of an AI benefit — and then immediately complicates it in the governable direction. The benefit is real: about 15 percent more issues resolved per hour, better customer sentiment, better retention. But it is a distribution, not a scalar, and the distribution is the whole lesson. Almost all of the gain went to novice agents; the most experienced gained close to nothing and may have lost a little quality. The average describes no one.
The generalizable instruction is that an agent-assist copilot mostly raises the floor, and a floor-raising tool has to be measured where the floor is. Reporting a mean productivity figure overstates the effect for experienced agents, understates it for novices, and hides the small quality cost at the top — so the honest measurement is the benefit's distribution across agent skill, tracked over time, not a single headline number. This is the same shape the map sees elsewhere when a benefit is uneven: the operator-heterogeneity reading, where one tool produces two outcomes and a uniform service term overstates for the group it helps least. Here the two groups are novice and experienced agents, drawn with identical structure, because the difference is a measured outcome, not a structural one.
The escalation and quality-assurance functions are where this becomes actionable. If the copilot raises novices toward veteran performance, the organization's quality risk shifts: the floor is higher, but the ceiling is unchanged and possibly slightly lower, and the agent who most needs a second look may now be a fluent-sounding novice whose draft reads like an expert's. The check drawn latent here is the benefit-distribution measurement — the function that could tell the organization which agents the tool helps, which it does not, and where a small quality cost is accumulating — rather than a single average that reports a win and stops looking.
The Lab network models only the deploying organization: its copilot, its agents, its interaction records, and the QA function that could measure the benefit's spread. No customer outcome is computed on any diagram. The customers being served are boundary-only; the productivity figures, the skill-compression finding, and the slight quality degradation at the top are institutional signals that live in this case file, never on any network. The map's instruction is to read the copilot as a real, floor-raising benefit and to govern it by its distribution across skill, not by the average that hides who it helps.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.