Domain Atlas / Security operations & fraud detection
Deep-learning fraud scoring and the reimbursement gap
Explore this deployment in the PAN Lab ↗
A Nordic bank's rules-based legacy fraud system ran at roughly 40 percent detection with a 99.5 percent false-positive rate — a measured pre-machine-learning baseline whose badness is the most credible datum in the record, since a 99.5 percent false-positive rate is not a marketing claim. The vendor-published rollout of a deep-learning engine scoring transactions in real time (under 300 milliseconds) claims false positives cut by about 60 percent and true-positive detection raised by about 50 percent; those figures are an organization-named, trade-press-covered vendor case study, entered here as claimed magnitudes against that legacy baseline because they were not independently audited.[2]
What happened
This is the two-sided portrait of the fraud domain. A Nordic bank's rules-based legacy fraud system ran at roughly 40 percent detection with a 99.5 percent false-positive rate. That legacy baseline is, paradoxically, the most credible number in the whole domain: a 99.5 percent false-positive rate is not a figure anyone markets, so it is the measured reality against which the improvement is claimed. The bank rolled out a deep-learning engine scoring transactions in real time, under 300 milliseconds, and the vendor-published case study claims false positives cut by about 60 percent and true-positive detection raised by about 50 percent. Those figures are organization-named and covered by independent trade press, but they are a vendor case study, not an independent audit, so the model enters them as claimed magnitudes against the one baseline the record establishes credibly.
The two-sidedness is what makes the case valuable, and it is a lesson about which lever moves which outcome. The same institution that improved its in-line fraud scoring later ranked worst among UK banks for reimbursing victims of authorized-push-payment scams, in the regulator's bank-by-bank performance data. Better detection and worse victim-outcome performance coexisted in one organization — because they are different things. Detection quality is a property of the model and the fraud-operations function; the justice of the disposition, once a customer has been defrauded, is a property of the reimbursement policy, held by a different part of the organization and ultimately by the regulator.
And the decisive evidence is what moved the victim outcome: not a better model, but a rule. The UK regulator's mandatory-reimbursement regime raised sector reimbursement from roughly two-thirds to about 89 percent. A classifier (a model that sorts each case into a category) improvement did not close the reimbursement gap; a regulatory mandate did. The honest reading is that a bank can genuinely improve its fraud detection and still fail the people the fraud reaches, because detection and reimbursement are different levers held by different actors — and the lever that moved the victim outcome was a rule change, which no amount of model improvement substitutes for.
The sociotechnical reading
Every other case in this domain lives on one side of the score — better detection, or the false-positive tail. This one holds both sides at once and draws the line between them. The bank improved its in-line fraud scoring, plausibly a lot; the legacy 99.5 percent false-positive baseline it replaced is the domain's most credible datum precisely because no one markets a number that bad. But the same bank ranked worst at reimbursing scam victims, and that is not a contradiction — it is the point. Detection quality and disposition justice are different levers held by different actors: the model and the fraud-operations team own detection; the reimbursement policy, and above it the regulator, own what happens to a customer the fraud reached. A better classifier moves the first and leaves the second untouched.
The governable insight is that the lever which moved the victim outcome was a rule, not a model. The regulator's mandatory-reimbursement mandate raised sector reimbursement from about two-thirds to 89 percent — a change no accuracy gain produced or could produce, because the reimbursement gap was never a detection problem. For an organization, this reframes what "improving fraud AI" can and cannot buy: it can buy detection, and detection is worth buying, but it cannot buy the justice of the disposition, which is a policy and regulatory surface a model does not reach. The map's instruction is to keep the two levers distinct and to name who holds each: measure detection and the victim outcome separately, resist letting a real detection gain stand in for a reimbursement record it does not touch, and recognize that some of the harms in this domain are moved only by rules held by actors outside the deploying organization. The honest boundary throughout: no customer or scam-victim outcome is computed on the Lab diagram. Customers and victims are boundary-only; scores, dispositions, and reimbursement decisions are institutional signals, and the detection figures, the reimbursement ranking, and the rule-change effect live in the case file, never on any network.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.