10 to the 23 AI logo

Domain Atlas / Security operations & fraud detection

Case fileMultinational (global bank; machine-learning AML as primary monitoring in key markets)giant deployment

ML anti-money-laundering as primary monitoring

Explore this deployment in the PAN Lab ↗

A global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-laundering product as its primary monitoring system in key markets, reporting two to four times more confirmed suspicious activity with roughly 60 percent fewer alerts. Every one of those numbers is a vendor-and-customer self-report with no independent audit — which is itself the honest structure of the domain, because a peer-reviewed deployment-scale benefit measurement inside a named financial-crime operation does not publicly exist, and the alert-volume reduction the vendor advertises is precisely the lever a regulator scrutinizing an under-monitoring risk would question.[]

What happened

A global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-laundering product as its primary monitoring system in key markets, reporting two to four times more confirmed suspicious activity with roughly 60 percent fewer alerts. Those are the headline numbers, and the first honest thing to say about them is what they are: a vendor-and-customer self-report with no independent audit. A peer-reviewed, deployment-scale benefit measurement inside a named financial-crime operation does not publicly exist, so the model enters these as claimed magnitudes and does not pretend otherwise — which is itself the honest structure of this domain, where almost every benefit figure is a marketing artifact rather than an audited result.

Two structural dynamics underneath the numbers are the domain's foundations. The first is the base-rate fallacy. Money laundering is rare relative to the transaction volume monitored, and under extreme base rates the precision of a detector is dominated by its false-alarm rate rather than by its accuracy. That means a threshold change moves the burden of alerts — how many analysts have to look at how many cases — far more than it moves the truth of what is caught, so an advertised "60 percent fewer alerts" is first of all a statement about workload, not about detection. The second is the label-feedback loop. Only a small fraction of flagged transactions is ever verified against a real outcome, and models are retrained on the investigators' own dispositions, so the labels the system learns from are the analysts' calls. A reported rise in "confirmed suspicious activity" is therefore partly a measure of what the system taught its reviewers to confirm, not an independent ground truth — the model and its labels can drift together, agreeing more over time without becoming more correct.

The governance texture is that this deployment sits inside a compliance function operating under long-standing regulatory scrutiny of exactly the failure the numbers touch: the alert-volume reduction the vendor advertises is precisely the lever a regulator worried about under-monitoring would question. The honest reading is not that the tool does not help — it may help substantially — but that the benefit is claimed rather than audited, the alert reduction is a workload figure that a base rate can make look like accuracy, and the "confirmed" rate is entangled with the labels the system generated for itself.

The sociotechnical reading

This domain is where the arithmetic sets the terms, and this case is its service-led anchor. The two numbers a fraud or AML deployment advertises — more caught, fewer alerts — are both true and both misleading unless you hold the base rate in mind. When the target is rare, precision is governed by the false-alarm rate, so "fewer alerts" is a workload claim first and an accuracy claim only if something independent confirms it; and "more confirmed activity" is entangled with the label-feedback loop, because the model retrains on the investigators' dispositions and only a sliver of flags is ever verified against reality. The governable insight is that the two levers everyone reaches for — the detection threshold and model improvement — mostly move burden, not truth, under these base rates, and that the check that would separate the two is an independent verification of ground truth that this domain almost never has.

That absence is the honest structure worth naming. Almost every deployment-scale benefit number here is a vendor self-report; peer-reviewed measurement inside a named financial-crime operation does not exist, so the map's instruction is to treat the magnitudes as claimed and to make the missing independent audit the thing you build. Two governable surfaces follow. The label-feedback loop is the first: reconciling the model's "confirmed" outputs against verified outcomes, rather than against the analysts' own prior calls, is what keeps the model and its labels from drifting into agreement without becoming more correct. The independent audit is the second: in a compliance function under standing regulatory scrutiny, the advertised alert reduction is exactly the lever an outside party should test, because a workload win and an under-monitoring risk look identical from the inside. The honest boundary throughout: no customer or enforcement outcome is modeled on the Lab diagram. Account holders and flagged parties are boundary-only; alerts, dispositions, and suspicious-activity filings are institutional signals, and the claimed magnitudes and the regulatory history live in the case file, never on any network.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

googlecloud2023GroundingVendorSave

Google Cloud (2023, June 21). Google Cloud Launches AI-Powered Anti Money Laundering Product for Financial Institutions (with HSBC-reported results). https://www.googlecloudpresscorner.com/2023-06-21-Google-Cloud-Launches-AI-Powered-Anti-Money-Laundering-Product-for-Financial-Institutions

https://www.googlecloudpresscorner.com/2023-06-21-Google-Cloud-Launches-AI-Powered-Anti-Money-Laundering-Product-for-Financial-Institutions

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring); model org: hsbc_aml_ai

axelsson2000aGroundingPeer-reviewedSave

Axelsson, S. (2000). The Base-Rate Fallacy and the Difficulty of Intrusion Detection. ACM Transactions on Information and System Security, 3(3), 186-205. https://doi.org/10.1145/357830.357849 https://dl.acm.org/doi/10.1145/357830.357849

doi.org/10.1145/357830.357849

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring)

dalpozzolo2018aGroundingPeer-reviewedSave

Dal Pozzolo, A., Boracchi, G., Caelen, O., Alippi, C., & Bontempi, G. (2018). Credit Card Fraud Detection: A Realistic Modeling and a Novel Learning Strategy. IEEE Transactions on Neural Networks and Learning Systems, 29(8), 3784-3797. https://doi.org/10.1109/TNNLS.2017.2736643 https://dalpozz.github.io/static/pdf/TNNLS_2017.pdf

doi.org/10.1109/TNNLS.2017.2736643

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring)

Topics: complexity-science

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalA global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-la…

A global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-laundering product as its primary monitoring system in key markets, reporting two to four times more confirmed suspicious activity with roughly 60 percent fewer alerts. Every one of those numbers is a vendor-and-customer self-report with no independent audit — which is itself the honest structure of the domain, because a peer-reviewed deployment-scale benefit measurement inside a named financial-crime operation does not publicly exist, and the alert-volume reduction the vendor advertises is precisely the lever a regulator scrutinizing an under-monitoring risk would question.

googlecloud2023GroundingVendorSave

Google Cloud (2023, June 21). Google Cloud Launches AI-Powered Anti Money Laundering Product for Financial Institutions (with HSBC-reported results). https://www.googlecloudpresscorner.com/2023-06-21-Google-Cloud-Launches-AI-Powered-Anti-Money-Laundering-Product-for-Financial-Institutions

https://www.googlecloudpresscorner.com/2023-06-21-Google-Cloud-Launches-AI-Powered-Anti-Money-Laundering-Product-for-Financial-Institutions

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring); model org: hsbc_aml_ai

EmpiricalTwo structural dynamics govern fraud and financial-crime detection. Under extreme base rates, detection precis…

Two structural dynamics govern fraud and financial-crime detection. Under extreme base rates, detection precision is dominated by the false-alarm rate rather than by accuracy, so at realistic prevalence a threshold change moves the burden of alerts rather than the truth of them (the base-rate fallacy). And the labels the model learns from are the investigators' own dispositions: only a small set of flagged transactions is ever verified, and models are retrained on the analysts' calls, so a rise in 'confirmed' activity is partly a measure of what the system taught its reviewers to confirm rather than an independent ground truth (the label-feedback loop).

axelsson2000aGroundingPeer-reviewedSave

Axelsson, S. (2000). The Base-Rate Fallacy and the Difficulty of Intrusion Detection. ACM Transactions on Information and System Security, 3(3), 186-205. https://doi.org/10.1145/357830.357849 https://dl.acm.org/doi/10.1145/357830.357849

doi.org/10.1145/357830.357849

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring)

dalpozzolo2018aGroundingPeer-reviewedSave

Dal Pozzolo, A., Boracchi, G., Caelen, O., Alippi, C., & Bontempi, G. (2018). Credit Card Fraud Detection: A Realistic Modeling and a Novel Learning Strategy. IEEE Transactions on Neural Networks and Learning Systems, 29(8), 3784-3797. https://doi.org/10.1109/TNNLS.2017.2736643 https://dalpozz.github.io/static/pdf/TNNLS_2017.pdf

doi.org/10.1109/TNNLS.2017.2736643

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring)

Topics: complexity-science