Software engineering AI (coding assistants)
AI code-completion assistants deployed to an organization's own engineers — the domain where the AI's benefit is real, large, and radically uneven: a big lift for less-experienced developers and near zero (or negative) for experts, who in one randomized study were about 19% slower while believing themselves 20% faster. Two structural traps define the domain. The engineer who accepts a suggestion is also its reviewer of record, so the operator and the verifier collapse into one person; and the shared repository the assistant writes into is the same corpus later engineers and later assistants read as fact. Individual speed-ups do not automatically compose to the organization's delivery outcomes — measured delivery stability falls as adoption rises unless code-review and testing gates absorb the added churn. The Lab networks model only the deploying organization — its engineers, its repository, its gates; the downstream users of the software sit outside the dynamics, and no product outcome is computed on any diagram.
Use cases
What AI is doing here
AI code-completion assistant
GenerativeAI that suggests code an engineer accepts, edits, or rejects inline — where the engineer who accepts the suggestion is also its reviewer of record, and the measured benefit varies sharply with the developer's experience.
Code-review & testing gates
PredictiveThe code-review and continuous-integration gates that must absorb the added churn AI-generated code introduces — the control on which whether individual coding speed-ups compose to organization-level delivery outcomes depends.
Rollout evaluation & acceptance telemetry
PredictiveThe acceptance-rate telemetry and phased-rollout evaluation deployments use to judge a coding assistant — a measure most correlated with perceived productivity, which is documented to be miscalibrated on experienced developers.
Case files
What has gone wrong and right
Documented deployments, presented as model organizations calibrated to the evidence, with full citations.
Randomized coding-assistant field experiments
United States and multinational (three enterprises; pre-registered peer-reviewed field experiments)The software-engineering domain's causal anchor: pre-registered, peer-reviewed randomized rollouts of a commercial code assistant across 4,867 developers found a pooled 26% increase in completed tasks — concentrated among less-experienced developers. An independent study of 16 experienced maintainers bounds the other tail: about 19% slower with the AI while believing themselves 20% faster. The engineer who accepts a suggestion is also its reviewer, and individual gains do not compose to org delivery unless gates absorb the added churn.
Explore this deployment in the PAN Lab →In-house code completion (one org owns every node)
United States (large technology company; internal platform deployment)An in-house machine-learning code-completion system built, deployed, and measured by Google's own platform organization for 10,000+ developers, with a control group: 25–34% acceptance, a 6% reduction in coding iteration time versus control, 3% of new code characters from the model. The structural point is unification — one organization owns the model, the monorepo, the review gates, and the telemetry, so it can tune the whole loop; the corresponding risk is that the measuring, building, and deploying party are the same, so no external check exists at all.
Explore this deployment in the PAN Lab →Gated coding-assistant rollout at a regulated bank
Australia (regulated bank; internal engineering)A regulated bank ran a structured six-week experiment with ~100 of its 5,000 engineers before scaling a commercial coding assistant to ~1,000, and published its own measurement. Engineers reported productivity and code-quality gains — and recorded the security impact as explicitly inconclusive: a real gating decision taken under uncertainty, with the unknown honestly carried forward rather than resolved by assertion. That recorded unknown is the case's most valuable datum, because the class-level security literature says the unknown is not hypothetical.
Explore this deployment in the PAN Lab →Ordinary competent coding-assistant rollout (400+ devs)
United States (mid-size enterprise; internal engineering)ZoomInfo's systematic four-phase evaluation-to-rollout of GitHub Copilot across 400+ developers, well documented: 33% acceptance, 20% of suggested lines accepted, 72% satisfaction, per-language variation, stated limitations. Its evaluation instrument is acceptance-rate telemetry — the measure most correlated with perceived, not real, productivity. And it reported no security evaluation at all: an unrecorded unknown, one step less honest than a recorded inconclusive one. The value here is the documentation quality of an ordinary, competent adoption.
Explore this deployment in the PAN Lab →System map
Who is in the system and what pushes on it
Who is in the system
- Frontline workers. Caseworkers, screeners, eligibility staff — the operator network whose judgment the system augments or erodes.
- Supervisors & QA. The institutional correction layer: overrides, second reads, quality review.
- Agency leadership. Owns procurement, policy, and the authority map; answers for the system publicly.
- Vendors. Build and update the systems; hold the information asymmetry procurement must govern.
- Regulators & oversight bodies. Boards, auditors, data-protection officers, inspectorates — external correction capacity.
Dominant pressures
- Caseload surge. Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck. One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity. The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure. Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Data & policy drift. The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
Governance
Questions leaders should be asking
- 1. The benefit is large for novices and near zero — or negative — for experts, who are measured to feel faster while being slower; so does the rollout track who it actually helps, or does a single headline productivity number let the novice result stand for everyone?
- 2. The engineer who accepts a suggestion is also its reviewer of record — operator and verifier collapsed into one person; so what independent check stands between an accepted suggestion and the shared repository, or is 'a human reviewed it' the same human who accepted it?
- 3. The assistant writes into the same repository that later engineers and later assistants read as fact — so does an accepted error propagate as inherited code, and is anyone reconciling machine-written code against an independent standard before it becomes the corpus?
- 4. Individual speed-ups do not automatically compose to delivery outcomes — measured stability falls as adoption rises unless gates absorb the churn; so are the code-review and testing gates resourced to the added volume, and did anyone record what the evaluation did not establish, such as security?
For the actions behind these questions, see the Practice Library.
Seeing your organization in this domain? Mapping its actual pathways, pressures, and correction capacity is engagement work.
Work With 1023AI