10 to the 23 AI logo

Domain Atlas / Software engineering AI (coding assistants)

Case fileUnited States and multinational (three enterprises; pre-registered peer-reviewed field experiments)giant deployment

Randomized coding-assistant field experiments

Explore this deployment in the PAN Lab ↗

Company-run, pre-registered, peer-reviewed randomized rollouts of a commercial code-completion assistant across 4,867 developers at three enterprises found a pooled 26.08 percent increase in completed tasks, with gains concentrated among less-experienced developers. An independent randomized study of 16 experienced open-source maintainers on 246 tasks in familiar repositories bounded the expert tail from the other direction: those developers were about 19 percent slower with the AI while believing themselves about 20 percent faster — a measured perception-reality gap that means a uniform productivity number overstates the effect for senior engineers.[2]

What happened

The software-engineering domain's causal anchor is a set of company-run randomized rollouts of a commercial code-completion assistant across 4,867 developers at two named enterprises and an anonymized third. The experiments were pre-registered and peer-reviewed, and they found a pooled 26.08 percent increase in completed tasks. The central fact is not the average but its shape: the gains were concentrated among less-experienced developers, and near zero for experts.

The expert tail is bounded from the other direction by an independent randomized study of 16 experienced open-source maintainers working 246 tasks in repositories they already knew well. Those developers were about 19 percent slower with the AI than without it — while believing themselves about 20 percent faster. That perception-reality gap is the domain's most important honesty datum: the measure most deployments track (how the tool feels, via acceptance rates) is exactly the measure that is miscalibrated for the engineers the tool helps least. A uniform productivity number, quoted without the seniority distribution, reports the novice result and lets it stand for everyone.

Two structural features define the deployment. First, the operator and the verifier collapse into one person: the engineer who accepts a suggestion is also its reviewer of record, so there is no independent check between an accepted suggestion and the shared repository unless the organization builds one. Second, the shared repository the assistant writes into is the same corpus later engineers and later assistants read as fact, so an accepted error does not sit in one place — it propagates as inherited code. The organization-level literature closes the loop: a cross-industry program measured a 7.2 percent decrease in delivery stability for every 25 percent increase in AI adoption, evidence that the individual speed-up does not compose to the organization's delivery outcome unless code-review and testing gates absorb the churn the assistant adds. Honesty boundaries: several authors are affiliated with the assistant's vendor and the experiments were run by the firms themselves — the peer-reviewed venue and the pre-registration are the checks that make the numbers credible despite that.

The sociotechnical reading

This case moves the Atlas out of social services and clinical care and into an organization's own workforce, but the governance grammar is the same — and it exposes two traps that are specific to generation into a shared codebase. The first is that the benefit is a distribution masquerading as an average. A pre-registered randomized experiment is strong evidence, and it says the tool delivers a real, large lift — to novices. For experts, an independent randomized study measured a slowdown paired with a confident feeling of speed-up. So the honest service term here is not a single percentage; it is a curve over seniority, and a deployment that reports the pooled number is reporting the group it helps most. The governable move is to measure who the tool actually helps rather than how fast it makes people feel.

The second trap is the collapse of the operator into the verifier. In the clinical documentation domain the clinician reviews an AI draft before it enters the record; here the engineer who accepts the suggestion is the review, so there is no second read between the machine's output and the shared repository unless the organization inserts one. That matters because the repository is not a private notebook — it is the corpus later engineers copy from and later assistants train and contextualize on, so an accepted error becomes inherited code with a long half-life, exactly as an unreviewed clinical note becomes inherited fact. And the organization-level evidence says the individual speed-up actively degrades delivery stability as adoption climbs unless the code-review and testing gates are resourced to absorb the extra volume. The map's instruction is that a coding assistant's real governance surfaces are downstream of the keystroke: an independent review between acceptance and merge, a reconciliation of machine-written code against a standard before it becomes the corpus, and gates sized to the churn. The honest boundary throughout: no product or software outcome is modeled on the Lab diagram. The downstream users of the software are boundary-only; suggestions, acceptances, and merges are institutional signals, and the productivity numbers, the expert slowdown, and the delivery-stability bound live in the case file, never on any network.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

cui2025aGroundingPeer-reviewedSave

Cui, Z.K., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2025). The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science. https://doi.org/10.1287/mnsc.2025.00535 https://pubsonline.informs.org/doi/10.1287/mnsc.2025.00535

doi.org/10.1287/mnsc.2025.00535

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

becker2025aGroundingIndustrySave

Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR. https://doi.org/10.48550/arXiv.2507.09089 https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

doi.org/10.48550/arXiv.2507.09089

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

googleclouddora2024GroundingReferenceSave

Google Cloud DORA (2024). Accelerate State of DevOps Report 2024. https://dora.dev/research/2024/dora-report/

https://dora.dev/research/2024/dora-report/

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review); model org: msft_accenture_copilot_experiments

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalCompany-run, pre-registered, peer-reviewed randomized rollouts of a commercial code-completion assistant acros…

Company-run, pre-registered, peer-reviewed randomized rollouts of a commercial code-completion assistant across 4,867 developers at three enterprises found a pooled 26.08 percent increase in completed tasks, with gains concentrated among less-experienced developers. An independent randomized study of 16 experienced open-source maintainers on 246 tasks in familiar repositories bounded the expert tail from the other direction: those developers were about 19 percent slower with the AI while believing themselves about 20 percent faster — a measured perception-reality gap that means a uniform productivity number overstates the effect for senior engineers.

cui2025aGroundingPeer-reviewedSave

Cui, Z.K., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2025). The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science. https://doi.org/10.1287/mnsc.2025.00535 https://pubsonline.informs.org/doi/10.1287/mnsc.2025.00535

doi.org/10.1287/mnsc.2025.00535

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

becker2025aGroundingIndustrySave

Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR. https://doi.org/10.48550/arXiv.2507.09089 https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

doi.org/10.48550/arXiv.2507.09089

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

EmpiricalIndividual coding-assistant gains do not automatically compose to organization-level delivery outcomes: a cros…

Individual coding-assistant gains do not automatically compose to organization-level delivery outcomes: a cross-industry research program measured a roughly 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability for every 25 percent increase in AI adoption, evidence that the churn the assistant adds must be absorbed by code-review and testing gates or the individual speed-up degrades the organization's delivery performance.

googleclouddora2024GroundingReferenceSave

Google Cloud DORA (2024). Accelerate State of DevOps Report 2024. https://dora.dev/research/2024/dora-report/

https://dora.dev/research/2024/dora-report/

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review); model org: msft_accenture_copilot_experiments