10 to the 23 AI logo

Domain Atlas / Software engineering AI (coding assistants)

Case fileAustralia (regulated bank; internal engineering)giant deployment

Gated coding-assistant rollout at a regulated bank

Explore this deployment in the PAN Lab ↗

A regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before scaling a commercial coding assistant to roughly 1,000 engineers, publishing its own measurement of the rollout. The bank's engineers reported productivity and code-quality improvements — and recorded the security impact as explicitly inconclusive, a real gating decision taken and documented under uncertainty rather than resolved by assertion, with the honestly recorded unknown carried forward into the scaled deployment.[2]

What happened

This case is the gated-adoption pole of the software-engineering domain: a regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before scaling a commercial coding assistant to roughly 1,000 engineers, and published its own measurement of the rollout. The trial-then-gate sequence is the governance topology worth modeling — a bounded pilot population, an explicit evaluation phase, a scale decision, and a documented result carried into the larger deployment.

The bank's engineers reported productivity and code-quality improvements. The case's most valuable datum, though, is what the bank recorded as unresolved: the security impact was explicitly inconclusive. That is a real gating decision taken and documented under uncertainty, and it is the honest form of a governance finding — an unknown named and carried forward, rather than a risk asserted away or quietly omitted. A deployment that says "we could not establish the security effect" has done something a deployment that says nothing about security has not: it has priced what its evaluation did not settle.

What the inconclusive finding leaves open is not hypothetical. An independent security assessment of code generated by a widely used assistant found that about 40 percent of generated programs contained vulnerabilities across scenarios spanning the Common Weakness Enumeration (CWE) top-25 weaknesses, and separate research documents developers accepting insecure suggestions with overconfidence. So the security unknown the bank carried forward sits against a class-level literature in which insecure generation is common. The honesty boundaries of the case cut both ways: the strength is that the deploying organization, not a vendor or an academic team, is the measuring party and documented its own gate; the boundary is that the report is a preprint authored by the bank's own engineers, not peer-reviewed, reporting the bank's own success — with the security-inconclusive finding as the honest counterweight inside its own account.

The sociotechnical reading

This case is about the shape of a good governance decision, and the honesty of recording what it did not resolve. The bank did the thing the composition literature implies: it did not switch the assistant on for everyone and measure later, it ran a bounded pilot, evaluated, and then made a scale decision — trial, then gate. That sequence is the governable structure, and it is worth modeling precisely because most deployments in this domain skip it. But the datum that makes this case valuable is not the productivity number; it is the security finding recorded as inconclusive. An organization that writes down "we could not establish this" has converted an unknown into a governed object — something named, carried forward, and available to be resolved later — instead of an absence no one has to answer for.

The instruction the case adds to the Atlas is that the honest treatment of an unknown is to record it, not to resolve it by assertion, and that a recorded unknown is one governance step ahead of an unrecorded one. The security question is the sharpest version here because the class-level evidence says insecure generation is common — roughly 40 percent of generated programs vulnerable in one assessment, with overconfident acceptance documented — so an inconclusive security finding is not a formality; it is a live, named risk sitting in a deployment scaled to a thousand engineers. The governable move is to keep resolving it: a security check that actually runs against generated code before it merges, the gate the pilot could not yet close. And the honesty boundary is that the measuring party is the deployer, reporting its own success in a non-peer-reviewed preprint — which is exactly why the self-recorded inconclusive finding matters, because it is the organization documenting a limit on its own good news. The honest boundary throughout: no banking or product outcome is modeled on the Lab diagram. The bank's customers are boundary-only; suggestions, merges, and trial metrics are institutional signals, and the productivity claims, the security-inconclusive finding, and the class-level vulnerability rate live in the case file, never on any network.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

chatterjee2024aGroundingIndustrySave

Chatterjee, S., Liu, C.L., Rowland, G., & Hogarth, T. (2024). The Impact of AI Tool on Engineering at ANZ Bank: An Empirical Study on GitHub Copilot within Corporate Environment [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2402.05636 https://www.theregister.com/2024/02/10/anz_bank_github_copilot/

doi.org/10.48550/arXiv.2402.05636

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

pearce2022aGroundingPeer-reviewedSave

Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., & Karri, R. (2022). Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions. In 43rd IEEE Symposium on Security and Privacy (SP 2022). https://doi.org/10.48550/arXiv.2108.09293 https://arxiv.org/abs/2108.09293

doi.org/10.48550/arXiv.2108.09293

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

Topics: privacy-security

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalA regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before sc…

A regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before scaling a commercial coding assistant to roughly 1,000 engineers, publishing its own measurement of the rollout. The bank's engineers reported productivity and code-quality improvements — and recorded the security impact as explicitly inconclusive, a real gating decision taken and documented under uncertainty rather than resolved by assertion, with the honestly recorded unknown carried forward into the scaled deployment.

chatterjee2024aGroundingIndustrySave

Chatterjee, S., Liu, C.L., Rowland, G., & Hogarth, T. (2024). The Impact of AI Tool on Engineering at ANZ Bank: An Empirical Study on GitHub Copilot within Corporate Environment [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2402.05636 https://www.theregister.com/2024/02/10/anz_bank_github_copilot/

doi.org/10.48550/arXiv.2402.05636

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

EmpiricalWhat the bank's inconclusive security finding leaves open is not hypothetical: an independent security assessm…

What the bank's inconclusive security finding leaves open is not hypothetical: an independent security assessment of code generated by a widely used assistant found that about 40 percent of generated programs contained vulnerabilities across scenarios spanning the CWE top-25 weaknesses, and separate research documents developers accepting insecure suggestions with overconfidence — so the security unknown a deployment carries forward unresolved sits against a class-level literature in which insecure generation is common.

pearce2022aGroundingPeer-reviewedSave

Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., & Karri, R. (2022). Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions. In 43rd IEEE Symposium on Security and Privacy (SP 2022). https://doi.org/10.48550/arXiv.2108.09293 https://arxiv.org/abs/2108.09293

doi.org/10.48550/arXiv.2108.09293

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

Topics: privacy-security