Domain Atlas / Software engineering AI (coding assistants)
Ordinary competent coding-assistant rollout (400+ devs)
Explore this deployment in the PAN Lab ↗
A mid-size enterprise ran a systematic four-phase evaluation-to-rollout of a commercial coding assistant across more than 400 developers, publishing acceptance telemetry (a 33 percent suggestion-acceptance rate, with 20 percent of suggested lines accepted), a 72 percent satisfaction figure, documented per-language variation, and stated limitations. Its evaluation instrument is acceptance-rate telemetry — which the productivity literature identifies as the measure most correlated with perceived productivity rather than outcome, and perception is measured to be miscalibrated for experienced developers, so acceptance telemetry captures adoption feel, not delivered output.[2]
What happened
This case is the mid-size-enterprise playbook of the software-engineering domain: ZoomInfo's systematic four-phase evaluation-to-rollout of GitHub Copilot across more than 400 developers, published both as a preprint and on the company's engineering blog. It is not a randomized experiment — the domain's anchor case carries that — and not a regulated gate — the bank carries that. It is the ordinary, competent adoption, and its value is exactly the documentation quality of an average case: phase gates, telemetry definitions, per-language deltas, satisfaction surveys, and stated limitations, published by the deployer itself.
The reported numbers are an acceptance story: a 33 percent suggestion-acceptance rate, 20 percent of suggested lines accepted, 72 percent satisfaction, with variation documented across programming languages. That is where the case earns its place, because acceptance rate is precisely the measure the domain's honesty literature bounds. In the productivity-measurement framework, acceptance rate is the metric most correlated with perceived productivity — how the tool feels — rather than with delivered output, and perception is measured to be miscalibrated for experienced developers, who can feel faster while being slower. So a deployment whose evaluation instrument is acceptance telemetry is measuring adoption feel, not outcome, however carefully it defines and reports that telemetry.
The sharpest datum is what the report does not contain. Its stated limitations are real and creditable, but no security evaluation is reported at all. That is an unrecorded unknown — one step less honest than the bank's recorded-inconclusive finding, not because the deployer acted in bad faith, but because an absence no one has written down is not a governed object: it cannot be carried forward, assigned, or resolved, and it does not appear on any list of things still to answer. The honesty boundaries of the case are that it is self-reported and not peer-reviewed, the telemetry measures adoption rather than outcome, and the security question is missing from an otherwise well-documented account — priced here as an unrecorded unknown, one honesty step below a recorded one.
The sociotechnical reading
This case is the baseline the others are measured against: a competent, well-documented, ordinary adoption. It does most things right — phased rollout, defined telemetry, per-language breakdowns, stated limitations — and precisely because it is careful, its two gaps are instructive rather than damning. The first is that its evaluation instrument measures the wrong thing well. Acceptance rate is defined, tracked, and reported with care, but acceptance rate is adoption feel, not delivered output — the measure the productivity literature ties to perceived rather than real productivity, on a population where perception is documented to be miscalibrated. A number can be rigorously measured and still be the wrong number; the governable move is to measure outcome, not acceptance, and to treat a satisfaction figure as a report on how the tool feels rather than what it produced.
The second gap is the one that ranks this case against the bank's, and it is the whole point of putting them side by side. The bank ran a security check and recorded the result as inconclusive; this deployment reported no security evaluation at all. Both end without a resolved security answer, but they are not equivalent: a recorded inconclusive finding is a governed unknown — named, carried forward, resolvable — while an unrecorded absence is nothing at all, invisible on every list of open questions, impossible to assign or close. The honesty ladder in this domain has three rungs, and this case sits one below the bank: resolve the unknown (best), record it unresolved (honest), or leave it unrecorded (the ordinary default). The governable instruction is to climb the ladder — run the check, and if it cannot yet conclude, write the unknown down so it becomes something the organization owes an answer on. The honest boundary throughout: no product outcome is modeled on the Lab diagram. The business customers are boundary-only; suggestions, acceptances, and merges are institutional signals, and the acceptance and satisfaction figures, and the missing security evaluation, live in the case file, never on any network.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.