Skip to content
KEDBYTE
How Money Moves
Chapter
51

Fraud Detection

Part V · Trust, Failure and the Law|8,448 words|about 37 min read|Volume 5
Fast-moving material. Figures, model names, prices and version numbers in this chapter were verified in August 2026. Claims are separated into established fact, active research and marketing claim. Re-check anything you intend to rely on.

51.0 What this chapter gives you#

  1. You will be able to explain why a fraud system makes two kinds of mistake and why only one of them ever generates an incident report.
  2. You will be able to compute the decline threshold for a given basket from the loss given fraud and the margin forgone, and say why it lands near one in three rather than nine in ten.
  3. You will be able to show how a rule with seventy-five per cent recall and five per cent precision can make a business nearly eight times worse off while appearing all the while to work.
  4. You will be able to distinguish precision, recall and specificity, and explain why at a fraud rate of 0.4 per cent almost all the leverage sits on the false positive rate.
  5. You will be able to name the four different problems the word fraud covers, and say which of them no device signal will ever help with.
  6. You will be able to separate signals supplied by the party being judged from evidence that party does not produce, and layer the two accordingly.
  7. You will be able to say who bears the loss at each of the five decision points in a card transaction, and how United Kingdom strong customer authentication moves it.
  8. You will be able to explain why a fraud rate below the reference rate is a commercial asset under the transaction risk analysis exemption, and what happens after two consecutive quarters above it.
  9. You will be able to size a manual review queue by economics rather than headcount, and say when a human is worth roughly two pounds an order.
  10. You will be able to explain why a model trained only on approved traffic grows confident about a world it has stopped observing, and what a holdout costs and buys.

The plain version#

Imagine a doorman on a busy Friday night. He has three answers, not two. He can wave you in, he can turn you away, or he can say “hang on a second” and check something. That third answer is the interesting one, and we will come back to it.

He cannot ask everybody for identification, because the queue would stretch round the block and half of it would go somewhere else. He cannot let everybody in, because after enough bad nights the venue loses its licence. So he reads people. He notices that you arrived alone but keep glancing back at a group, that this is the fourth time in twenty minutes somebody has come to the door in the same brand-new jacket, that you have been here on eleven previous Saturdays and never caused trouble. None of those observations is proof. Every one shifts the odds.

Now here is the part everybody gets wrong. The doorman makes two completely different kinds of mistake and only one is visible. If he lets a troublemaker in, there is a fight and everyone remembers. If he turns away an ordinary person who came out for a quiet drink, nothing happens at all. No incident, no report, no smashed glass. That person goes somewhere else, tells their friends the place is unwelcoming, and never comes back. The mistake is real, expensive and completely silent.

Fraud detection is that doorman, wired into a computer, running thousands of times a second. Now put money on it.

Nadia sells running shoes online from a unit outside Leicester, about ten thousand orders a month at an average of eighty pounds. The shoes cost her forty-eight pounds a pair, so she keeps thirty-two pounds of gross margin on a typical order.

About forty orders a month are placed with stolen card details. She ships them, and six to ten weeks later the real cardholder notices the charge and disputes it. Three things then go wrong at once. The eighty pounds is taken back out of her account. The shoes are gone. And her payment provider charges her an administration fee for handling the dispute, somewhere around fifteen pounds, set by her contract rather than by any regulator.

So one successful fraud costs Nadia forty-eight pounds of stock plus fifteen pounds of fee, which is sixty-three pounds. Forty of those a month is two thousand five hundred and twenty pounds. That is her fraud problem, whole and entire. One wrongly refused genuine customer costs her thirty-two pounds, the margin she did not make.

Those two numbers are the whole chapter. Sixty-three against thirty-two. Stopping one fraud is worth roughly two lost customers. Not two hundred. Two.

Nadia’s first instinct is the one everybody has. Most of the fraudulent orders go to a delivery address that does not match the billing address on the card, so she writes a rule: if the two differ, refuse the order. It works, in the sense that it catches fraud. Thirty of the forty frauds are stopped, one thousand eight hundred and ninety pounds of loss prevented, and if that were the only number on the report she would think she had had a very good month.

But the rule does not only flag fraudsters. It flags the man buying trainers for his nephew’s birthday, the student whose card is registered to her parents’ house, everybody who has moved recently, everybody shipping to an office. Six hundred genuine orders a month trip the same rule. Six hundred times thirty-two pounds is nineteen thousand two hundred pounds of margin, destroyed silently, with no incident report anywhere.

Add it up. Before the rule, fraud cost her two thousand five hundred and twenty pounds a month. After it, she still loses six hundred and thirty pounds to the ten frauds that got through and has additionally thrown away nineteen thousand two hundred pounds of good business. The rule made her situation nearly eight times worse, and it felt the whole time as though it were working, because the only number anybody looked at was the fraud she stopped.

There are two words for what went wrong and a business reader should know both. Precision asks: of everything I stopped, how much was actually bad? Nadia’s rule stopped six hundred and thirty orders and thirty were fraud. Precision of about five per cent, meaning nineteen out of every twenty people she refused had done nothing wrong. Recall asks: of all the bad things out there, how much did I stop? She stopped thirty of forty, so recall of seventy-five per cent. Her rule had good recall and catastrophic precision, and precision was where the money was.

Now watch what happens when she asks four questions at once. She keeps the address mismatch, but only acts on it when three other things are also true: the account was created less than ten minutes before the order; the same card number was submitted more than once in the session with different expiry dates; and that delivery postcode has been used by three or more different accounts in the last thirty days.

Each is innocent alone. New accounts are made constantly, people mistype expiry dates, blocks of flats share postcodes. It is the coincidence that is not innocent. The combination fires on forty-five orders a month and twenty-eight are fraud.

Precision has gone from five per cent to sixty-two, recall has slipped from seventy-five to seventy. She now prevents one thousand seven hundred and sixty-four pounds of fraud while destroying five hundred and forty-four pounds of good business, so her total monthly cost is one thousand three hundred pounds against two thousand five hundred and twenty if she did nothing. She is making money on the arrangement for the first time, and she is also letting twelve frauds through every month on purpose.

That is not failure. That is the answer.

Because there is a way to get fraud to exactly zero and everybody knows what it is: refuse every order. Nadia would lose three hundred and twenty thousand pounds of margin a month and never be defrauded again. Zero fraud is trivially available and nobody wants it. So the real question was never how to stop fraud. It was where, on the dial between letting everything through and letting nothing through, the total cost bottoms out.

The dial has a setting, it falls out of those two numbers, and it is genuinely surprising. Nadia should refuse an order when she thinks there is more than about a one in three chance it is fraudulent. Not a nine in ten chance. Not “probably fraud”. Thirty-four per cent. At that point the sixty-three pounds she saves two-thirds of the time exactly balances the thirty-two pounds she throws away one-third of the time. Above it she should refuse; below it she should ship the shoes and accept the losses, because refusing costs more than being defrauded. Cheap shoes on a fat margin move the dial one way and expensive shoes on a thin margin move it the other. It is a business decision wearing a technical costume.

So what does the machine look at, if not the address?

Velocity. Not “how many”, but “how many, in how long, from the same thing”. One card tried once is a customer; one card tried nineteen times in four minutes across eleven accounts is not. One delivery address on two orders is a household; one delivery address on sixty orders in a week, all to different names, is a drop. Velocity is the oldest and cheapest signal there is and still one of the most useful, because criminals have to repeat themselves to make money and honest people mostly do not.

Device fingerprinting. When your browser talks to Nadia’s shop it says a great deal about itself without being asked: which fonts are installed, how big the screen is, which time zone, how the graphics chip draws a particular curve. Combine forty of those and you get something often unique. It is not you. It is the machine you are sitting at, which matters because a fraudster with a thousand stolen cards is often sitting at one machine.

Behavioural signals. How you type, not what you type. Whether you typed your own name fluently, the way somebody does who has written it ten thousand times, or slowly, the way somebody does who is copying it off a screen. Whether you were on a call the whole time you were paying, which is one of the strongest warnings in the business, because someone talking a stranger through a bank transfer is usually stealing from them.

The network. Nadia alone sees one strange order. Nadia plus four hundred other shops on the same platform sees the same device try eleven of them in ninety seconds, a pattern invisible to any one of them. The biggest advantage a large processor has over a small one is not better mathematics. It is that it has seen the card before.

And finally the doorman’s third answer. Some orders go into a pile for a person to look at. A trained reviewer costs an employer something like forty thousand pounds a year once you count national insurance, pension, a desk and a supervisor, about two pounds per order at a dozen an hour. Two pounds is not nothing when the order only makes thirty-two, so the rule for the review pile is simple and almost never followed: only put an order in front of a human if a human might change the answer.

Where the plain version stops being true#

The doorman can see the person. The machine can only see what the person chooses to send it. Every signal above arrives over a wire from the party being judged. The device fingerprint, the browser headers, the claimed time zone, the typing rhythm: all of it is client-supplied, and a determined attacker can influence any of it. This is not a reason to abandon the signals, which remain useful because spoofing consistently is much harder than spoofing once. It is a reason to notice which signals do not come from the client at all. The issuer’s view of the account, the network’s view of the card across thousands of merchants, a cryptogram generated inside a chip, a passkey bound to hardware: these are expensive to forge precisely because the attacker is not producing them. The defence is layering: cheap client signals to sort the ordinary from the odd, expensive non-client evidence to decide the odd cases.

“Fraud” is at least four different problems and the arithmetic above fits only one. Nadia’s case is third-party fraud: a criminal using a real cardholder’s details. Beside it sit three others with different physics. First-party misuse, sometimes called friendly fraud, is where the genuine cardholder makes a genuine purchase and then disputes it, and no device signal will ever help because everything about the transaction is authentic. Account takeover is where the criminal is inside the real customer’s account, so the history that normally exonerates now vouches for the attacker. And authorised push payment scams invert the problem: the victim makes the payment themselves, correctly authenticated, fully intending to, while being manipulated by someone on the phone. In the United Kingdom that fourth category is now comparable in scale to the first. UK Finance’s Annual Fraud Report 2026, published on 15 June 2026 and covering calendar year 2025, reported unauthorised fraud losses of £703.4 million against authorised push payment losses of £576.4 million, the second figure rising nineteen per cent year on year. In APP fraud the model is not looking for a criminal in the session. It is looking for a victim. The signal set, the intervention and the legal consequences are all different.

The plain version put the whole loss on the merchant, and in the United Kingdom that is frequently wrong. Who pays depends on who authenticated, and it moved substantially with strong customer authentication. Under the Payment Services Regulations 2017, as they stood at the time of writing in August 2026, regulation 76 requires a payment service provider to refund an unauthorised transaction no later than the end of the business day following the day it becomes aware of it, unless it has reasonable grounds to suspect fraud by the customer and has notified the authorities under the tipping-off provisions of the Proceeds of Crime Act 2002. Regulation 77, again as at August 2026, caps the payer’s own liability at a maximum of thirty-five pounds for losses arising from a lost, stolen or misappropriated payment instrument, and removes even that where strong customer authentication was required but the payer’s provider did not apply it, or where the instrument was used in connection with a distance contract. So a false negative and a false positive have different owners at every point in the chain, and “the loss ratio” means one thing to a merchant, another to an acquirer and another to a card issuer. Any conversation about tuning that has not established whose money is at stake is a conversation about nothing.

A score is not a probability, and the threshold is not an engineering decision. Nadia’s thirty-four per cent rule only works if the model’s output is a genuine probability, meaning that among all orders scored at thirty-four per cent, roughly thirty-four per cent really do turn out to be fraud. Most models out of training do not have that property, and a raw score of eight hundred and fifty on a scale to a thousand means nothing until somebody has mapped it onto reality. Worse, the ground truth arrives late and incomplete: a dispute may not surface for months, and blocked orders never generate any outcome at all, so the system learns only from traffic it approved and becomes confident about a world it has stopped observing.

The technical version#

On dating#

Everything below is stated as accurate at the time of writing, in August 2026, and this is among the fastest-moving corners of the subject. Three things shifted in the eight months before it was written and are named in place: the FCA’s removal of the prescribed contactless limits, effective March 2026; the replacement of Article 22 of the UK GDPR with new Articles 22A to 22D on 5 February 2026; and the continuing revision of Visa’s acquirer monitoring thresholds. Scheme programmes are distributed to acquirers rather than published, so treat every scheme figure here as a prompt to check your acquirer’s current bulletin.

Where the decision actually gets made#

There is no single fraud check. There are five decision points in a card transaction, each with different data, a different latency budget and a different party carrying the loss.

Point Who decides Data available Latency Who bears the loss
Pre-authorisation screening Merchant or gateway Basket, history, device, session behaviour Tens to hundreds of milliseconds Merchant, in lost margin
Authentication Issuer, via EMV 3-D Secure Merchant-supplied 3DS data plus issuer’s own view Seconds, with a challenge Shifts on successful authentication
Authorisation Issuer Card, account, balance, issuer models, network signals Milliseconds in the network round trip Issuer, or merchant if unauthenticated
Fulfilment hold Merchant Everything above, plus time and manual work Minutes to hours Merchant
Dispute Issuer, then scheme arbitration The full record, months later Days Per scheme rules and liability shift

EMVCo publishes the authentication specification, the EMV 3-D Secure Protocol and Core Functions Specification, which has run through version 2.3.1.1, covered by EMVCo Specification Bulletin 279 of 11 August 2025. The point of 3DS 2.x is that it carries merchant-supplied risk data into the issuer’s decision, permitting a frictionless outcome in most cases rather than a password screen for everybody.

Rules, scorecards and learned models#

Practitioners sometimes talk as though machine learning replaced rules. It did not, and a stack with no rules layer usually cannot implement policy. Rules carry what are not statistical judgements: sanctions and blocklist hits, hard regulatory constraints, contractual prohibitions, an instruction from an investigator to freeze a specific device. These must be deterministic, auditable and changeable at three in the morning by somebody who is not a data scientist. A rule engine edited under change control, with every version retained, is a compliance artefact as much as a detection tool.

Gradient-boosted decision tree ensembles trained on tabular features remain the workhorse for scoring the grey zone. They handle mixed data types and missing values, they are fast enough at inference, and their feature attributions are legible enough to survive a model validation review.

Sequence and graph models sit above both. A sequence model treats activity as an ordered series rather than a bag of counters, which lets it notice that a session is shaped wrongly rather than that a single field is wrong. A graph model works over the entity network, with cards, devices, addresses and beneficiaries as nodes and transactions as edges, and earns its keep on organised fraud and mule networks, where the individual transactions are unremarkable and only the topology gives it away.

They fail differently: rules by going stale, tree ensembles by drifting as the population moves, graph models by being expensive and lagging real time. Running all three, with the rules able to override the models, is defence in depth.

Velocity, entity resolution and the feature store#

A velocity feature is a count or sum over a sliding window, keyed on an entity. The design questions are which entity, which window, which aggregate. Cards per device in the last hour. Distinct billing postcodes per card in the last day. Declined authorisations per merchant per BIN range in the last five minutes.

Two engineering problems dominate. The first is entity resolution: a counter is only as good as the identity it is keyed on, and identities are messy, since the same card is a different token at every merchant and the same device changes after an update. Real systems maintain an identity graph stitching these together with explicit confidence, and the quality of that graph sets a ceiling on everything downstream.

The second is point-in-time correctness. If a training feature is computed using data that would not have been available at the moment of the decision, the model learns from the future and looks superb in backtest and mediocre in production. This is leakage, and it is the commonest reason a fraud model performs worse live than on paper.

Card testing, also called enumeration, is the canonical velocity problem. From the defender’s side it presents as a burst of low-value authorisation attempts with an abnormally high decline rate, often against a narrow range of card numbers, at merchants with a low-friction checkout. The controls exist because of it. Visa’s Account Attack Intelligence produces a VAAI Score used to confirm enumerated authorisations, and Visa’s Acquirer Monitoring Program carries a separate enumeration ratio. Mastercard’s Security Rules and Procedures, Merchant Edition, in the edition dated 11 February 2025 which was current at the time of writing, provides at section 3.12.4 that in the case of a BIN attack at a card-not-present merchant or payment facilitator, the acquirer, its service providers or the affected merchant must begin to capture and transmit CVC 2 values in all card-not-present authorisation request messages within seventy-two hours of detection or notification.

Device identification in a browser draws on the properties a client exposes: user agent and header ordering, characteristics of the TLS handshake, screen and viewport metrics, font enumeration, canvas and WebGL rendering differences, declared time zone and language, and the coherence between them. On mobile the signal set is richer and more stable, because an app can read hardware and operating system attributes a browser cannot.

Two caveats belong here. First, the entropy available in browsers is deliberately shrinking as privacy changes across the major engines reduce the distinguishing power of these signals, so fingerprint-based identification is a depreciating asset. Second, this is regulated processing. In the United Kingdom, at the time of writing in August 2026, storing information on or gaining access to information stored on a user’s terminal equipment engages regulation 6 of the Privacy and Electronic Communications (EC Directive) Regulations 2003, which requires consent subject to a narrow exemption where the access is strictly necessary to provide a service the subscriber has explicitly requested. Firms generally rely on that exemption for payment fraud checks, and on legitimate interests as the lawful basis under Article 6(1)(f) of the UK GDPR, supported by recital 47, which states expressly that processing strictly necessary for the purposes of preventing fraud constitutes a legitimate interest. Neither position survives inspection if it has not been documented, scoped and assessed.

Behavioural signals are a distinct family. Passive behavioural biometrics observe keystroke dynamics, pointer or touch trajectories, hesitation and correction patterns, and device motion from the accelerometer and gyroscope. The interpretive value is comparative: does this session look like this account’s previous sessions, and like a person performing a familiar task or an unfamiliar one. In APP scam detection this family carries the load, because there is nothing wrong with the credentials or the device, and the signals that matter indicate coercion or coaching: a long dwell on the payee details, evidence of an active call during the session, hesitation inconsistent with the customer’s baseline. UK Finance’s 2026 report quoted BioCatch’s global advisory director to this effect, that detection must move to identifying social engineering in real time before funds leave the account. One caution fraud teams underweight: behavioural biometrics can incidentally reveal information about a person’s health, motor control or disability, and where processing reveals data concerning health it becomes special category data under Article 9 of the UK GDPR, which matters both for the lawful basis and for the automated decision rules below.

The network effect and the mechanics of sharing#

The largest structural advantage in fraud detection is population. A model trained on one merchant’s traffic sees that merchant’s fraud. A model trained across a consortium sees the same device, card or beneficiary as it moves between merchants, which is where organised fraud is visible. Sharing runs through four mechanisms in the United Kingdom and they are not interchangeable.

Scheme-level merchant risk lists are the oldest. Mastercard operates MATCH, the Mastercard Alert to Control High-risk Merchants system, now MATCH Pro, into which acquirers must add terminated merchants under defined reason codes and against which they must query prospective ones; Visa operates an equivalent. These are compliance obligations under the schemes’ security rules, not optional intelligence products, and being listed is close to fatal to a merchant’s ability to obtain acquiring.

Industry fraud databases operate at consumer and account level. Cifas maintains the National Fraud Database, into which members file confirmed fraud markers and against which they screen, and UK Finance operates fraud intelligence sharing among its members. Pay.UK’s Confirmation of Payee, mandated by the Payment Systems Regulator in phases, checks the name a payer types against the name on the destination account before the payment is sent, a control aimed squarely at APP fraud and invoice redirection.

Statutory gateways sit underneath. The Economic Crime and Corporate Transparency Act 2023 created a direct disclosure gateway allowing certain businesses to share customer information to prevent economic crime, with protection from civil liability for breach of confidence where the conditions are met. That addresses the confidentiality problem, not the data protection one, which still requires a lawful basis, purpose limitation analysis, transparency and an impact assessment. Commercial consortium models are what most merchants buy, and the question usually left unasked in procurement is what happens to the merchant’s own data on exit.

Precision and recall, properly#

The confusion matrix has four cells and every one has a price.

Model says fraud Model says legitimate
Actually fraud True positive: loss avoided False negative: chargeback, goods gone, fee
Actually legitimate False positive: lost margin and customer True negative: normal business

Precision is true positives divided by everything flagged. Recall, also called the detection rate, is true positives divided by all actual fraud. Specificity, the true negative rate, is rarely quoted and is where the money hides.

The reason is the base rate, and it deserves a worked example because business readers are routinely misled by accuracy figures. Take ten thousand orders with forty frauds, a fraud rate of 0.4 per cent, and a detector with ninety per cent recall and a one per cent false positive rate. It flags thirty-six real frauds and one hundred false alarms, giving precision of twenty-six per cent: three out of four people you refuse have done nothing wrong. Now improve only the false positive rate, from one per cent to 0.2 per cent, changing nothing about recall. The flagged set becomes thirty-six real frauds and twenty false alarms, and precision jumps to sixty-four per cent.

Nothing about the model’s ability to detect fraud changed. All the leverage was on the negative class, because the negative class is two hundred and fifty times bigger. At low base rates, small movements in the false positive rate dominate everything else, and any metric averaging over both classes will hide it. That is also why area under the receiver operating characteristic curve is a poor headline metric for fraud, since it is computed across the full range of false positive rates, almost all of which are commercially unusable. Precision-recall curves, or simply recall at a fixed operating point such as detection rate at a one per cent flag rate, tell you what you need to know.

Two refinements separate professional measurement from amateur. Count-weighted and value-weighted metrics diverge, so a system catching ninety per cent of fraud attempts by count while missing the three largest by value has failed; value detection rate belongs beside the count figures on every report. And the labels are wrong. Chargebacks arrive with a lag of weeks to months, and blocked transactions produce no outcome at all, so the false positive rate cannot be measured without deliberately letting some flagged traffic through as a holdout. A team that has never run one does not know its own false positive rate and should not claim to.

Tuning to a loss ratio#

The industry expresses fraud performance in basis points of processed volume: fraud losses divided by transaction value, times ten thousand. The operating discipline hangs off that ratio rather than off an absolute number of incidents.

The context for what counts as a normal ratio comes from the regulators. The joint EBA and ECB Report on Payment Fraud published on 1 August 2024, the first in that series, found EEA payment fraud of €4.3 billion in 2022 and €2.0 billion in the first half of 2023, with card fraud on EEA-issued cards running at 0.031 per cent of total value and 0.015 per cent of total number in the first half of 2023. It also found fraud rates on card payments around ten times higher where the counterparty was outside the EEA, where strong customer authentication is not legally required, and that seventy-one per cent of card fraud by value in that period was cross-border. Later editions report the same direction with a wider extra-EEA gap.

That ratio is not merely a management preference. It is a regulatory threshold in two directions. The first is the transaction risk analysis exemption. Under Article 18 of the UK version of the strong customer authentication regulatory technical standards, as they stood at the time of writing in August 2026 and as reflected in the FCA’s Technical Standards chapter last updated on 19 March 2026, a payment service provider may decline to apply strong customer authentication to a remote payment it has identified as low risk, provided its fraud rate for that transaction type is at or below the reference fraud rate in the Appendix, the value is within the corresponding exemption threshold value, and a real-time risk analysis has not identified abnormal spending or behaviour, unusual device or software information, malware in the authentication session, a known fraud scenario, an abnormal payer location or a high-risk payee location. That clause turns a fraud team’s performance into a commercial asset: below the reference rate you may waive authentication on low-risk payments, which materially improves conversion; above it you may not.

The reference fraud rates in the equivalent EU instrument, Commission Delegated Regulation (EU) 2018/389, remain as follows at the time of writing, with the UK Appendix carrying sterling exemption threshold values rather than euro ones.

Exemption threshold value Reference fraud rate, remote card Reference fraud rate, remote credit transfer
500 0.01% 0.005%
250 0.06% 0.01%
100 0.13% 0.015%

Article 19, at the time of writing, defines the calculation precisely: the total value of unauthorised or fraudulent remote transactions, whether or not funds were recovered, divided by the total value of all remote transactions of that type, whether authenticated or executed under any exemption, on a rolling ninety-day basis, subject to the audit review required by Article 3(2). Article 20 sets the consequence: a provider must report immediately to the FCA when a monitored fraud rate exceeds the applicable reference rate, must cease using the exemption for that transaction type and threshold band if the rate is exceeded for two consecutive quarters, and may not resume until the rate has been at or below the reference rate for one quarter and it has notified the FCA with evidence. Article 21 requires quarterly monitoring of fraud values, fraud rates, average transaction values and exemption usage.

Two other UK figures under the same instrument are worth stating with their dates. The low-value remote transaction exemption in Article 16, at the time of writing in August 2026, applies where the individual transaction does not exceed £25 and either the cumulative value since the last strong customer authentication does not exceed £85 or there have been no more than five consecutive such transactions. And the contactless exemption in Article 11 no longer carries a prescribed limit at all: the FCA announced on 19 December 2025, in a press release accompanying Handbook Notice 136, that it was removing the fixed contactless limits and allowing providers with strong fraud controls to set their own, with the rule changes taking effect in March 2026. The Article 11 text as it stood in August 2026 simply permits the exemption where the provider identifies the transaction as posing a low level of risk. The FCA expected most providers to keep their existing limits, so a firm’s operational limit and the regulatory limit are now separate facts to be checked separately. UK Finance reported contactless fraud losses of £46.8 million in 2025, up eight per cent.

The second regulatory direction is the schemes’ own monitoring, which binds merchants and acquirers rather than issuers.

Visa’s Acquirer Monitoring Program consolidated the older fraud and dispute programmes into a single count-based ratio. The numerator combines card-not-present fraud reports, filed as TC40 records, with non-fraud disputes filed as TC15 records; the denominator is settled card-not-present transactions. A separate enumeration ratio measures enumerated authorisation attempts, approved and declined, against the total count of authorisation attempts, approved and declined. The programme identifies both acquirer portfolios and individual merchants, applies minimum monthly counts before a party enters it, and assesses fees per fraudulent or disputed transaction levied on the acquirer, which passes them on by contract. The thresholds themselves are published by Visa rather than confined to acquirer bulletins: its fact sheet “Visa Acquirer Monitoring Program Overview”, on the corporate Visa site, puts an acquirer portfolio above standard at a VAMP ratio of fifty basis points and excessive at seventy, sets the excessive merchant threshold in Asia-Pacific, Canada, the European Union and the United States at two hundred and twenty basis points falling to one hundred and fifty on 1 April 2026, requires fifteen hundred combined fraud and dispute events in a month before a party is identified at all, and puts the enumeration criteria at a ratio of two thousand basis points on a minimum of three hundred thousand enumerated authorisations. Those figures have been revised more than once since the programme launched and are reported inconsistently by third parties, so check them against Visa’s current fact sheet and against your acquirer’s own bulletin before relying on any of them.

Mastercard operates an Excessive Chargeback Program and an Excessive Fraud Merchant programme of similar structure, keyed on fraud chargeback basis points, absolute fraud value and, distinctively, the merchant’s take-up of EMV 3-D Secure. Separately and verifiably, the Mastercard Security Rules and Procedures, Merchant Edition, in the 11 February 2025 edition current at the time of writing, provides at section 3.12.4 that an acquirer must ensure that each merchant which has exceeded one hundred gross basis points in fraudulent card-not-present transactions for two consecutive calendar months captures and transmits the CVC 2 value for all mail order and telephone order transactions, and for e-commerce either captures and transmits CVC 2 or becomes Mastercard Identity Check enabled, within two months following the second trigger month.

So there are two floors under the loss ratio. There is the commercial floor, where fraud losses plus lost good business plus operating cost are minimised, and a regulatory or scheme floor, above which you lose exemptions, lose the frictionless authentication path and start paying per-transaction penalties. The second is often higher, which means a merchant optimising purely on expected value can find itself out of compliance while making money.

The expected-value calculation is the one from the plain version, stated properly. Let the calibrated probability of fraud be p, the loss given fraud be L, and the contribution margin forgone on a wrongly declined order be M. Decline when pL exceeds (1 − p)M, that is when p exceeds M divided by (M + L). With the shoe retailer’s L of £63 and M of £32 the optimal threshold is 32 divided by 95, approximately 33.7 per cent. Three consequences follow and all three are routinely missed. The threshold moves with the basket, not just the customer: a high-margin, low-cost digital good has a small L and a large M, so the threshold rises. Calibration is not optional, because if the score is not a probability the threshold is meaningless, and calibration decays faster than discrimination. And L is not just the goods; it includes the dispute fee, the operational cost of the case, the effect on the scheme monitoring ratio and any downstream penalty. Firms modelling L as the transaction value alone systematically under-block.

Manual review queues and their economics#

A review queue is a third decision bucket and should be sized by economics rather than headcount, which is the reverse of how it usually happens. A reviewer on a thirty-thousand-pound salary costs an employer roughly forty thousand fully loaded once employer national insurance, pension, workspace, tooling and supervision are counted. Against roughly seventeen hundred productive hours a year that is about twenty-four pounds an hour, and at twelve cases an hour the marginal cost of a review is about two pounds.

Now test whether the queue earns it. Take three hundred orders a month in the grey band, of which twenty are fraud and two hundred and eighty genuine, on the shoe retailer’s numbers. Auto-approving the band costs twenty times sixty-three pounds, or £1,260. Auto-declining it saves that but destroys two hundred and eighty times thirty-two pounds, a net loss of £7,700, so declining is clearly wrong. Reviewing all three hundred at a reviewer accuracy of ninety per cent on fraud and ninety-seven per cent on genuine orders costs six hundred pounds of review time, plus two missed frauds at £126, plus about eight wrongly declined genuine orders at £269, totalling roughly £995. The queue saves about two hundred and sixty-five pounds a month. It is worth running, barely. Raise the average order value to four hundred pounds on the same margin percentage and the same queue saves thousands, which is why review is standard in travel, jewellery and electronics and marginal in low-value retail: the economics are driven overwhelmingly by order value.

Four operating principles follow. Queue only where a human has information the model lacks: a reviewer seeing the same features on the same screen is an expensive, less consistent copy of the model, whereas one who can make an outbound call or verify a document earns their cost. Measure reviewer precision and recall as you would a model’s, by seeding the queue with blind duplicates of resolved cases; reviewers drift, tire and develop idiosyncratic heuristics, and on cases where the model was already confident they are frequently worse than it is. Treat latency as a cost, because conversion decays measurably with delay. And do not let capacity set the threshold: the commonest failure in real operations is a threshold chosen so the queue is exactly as long as the team can clear, which means the risk appetite of the business is set by its recruitment budget.

Fraud decisions refuse people service, sometimes close accounts, and sometimes place a marker other institutions will see. That engages data protection law directly, and the UK position changed recently enough that many internal policies had not caught up at the time of writing.

Section 80 of the Data (Use and Access) Act 2025 replaced Article 22 of the UK GDPR with new Articles 22A to 22D, brought into force on 5 February 2026 by the Data (Use and Access) Act 2025 (Commencement No. 6 and Transitional and Saving Provisions) Regulations 2026. The old general prohibition on solely automated decisions with legal or similarly significant effects has become a safeguards regime. Under Article 22A a decision is based solely on automated processing where there is no meaningful human involvement, and is significant where it produces a legal effect or a similarly significant one. Under Article 22B a significant decision based entirely or partly on special category data may not be taken solely automatically unless explicit consent applies, or the decision is necessary for a contract or required or authorised by law and the substantial public interest condition in Article 9(2)(g) is met. Under Article 22C, where a significant decision is taken solely automatically, the controller must have safeguards giving the individual information about the decision and enabling them to make representations, to obtain human intervention and to contest it.

For a fraud team this converts into four obligations: know which decisions are significant, know which are solely automated, demonstrate that any human in the loop is meaningfully involved rather than clicking through, and have a working route for a customer to contest an outcome. A rubber-stamp review queue is not a compliance control; after these provisions it is arguably worse than no queue at all, because it asserts a human involvement that does not exist.

Two adjacent points. In the European Union the AI Act classifies AI systems used to evaluate the creditworthiness of natural persons as high risk under Annex III, but with an express exception: systems used for the purpose of detecting financial fraud are carved out. The carve-out attaches to the fraud detection purpose, not to any model a fraud team happens to run. And in the United Kingdom, banks, building societies and PRA-designated investment firms holding internal model approval for regulatory capital have been subject to the PRA’s Supervisory Statement SS1/23 on model risk management principles since it took effect on 17 May 2024, whose five principles on identification and tiering, governance, development and use, independent validation and mitigants apply across model types rather than only to capital models. Firms outside that scope adopt the framework voluntarily, because the alternative when a fraud model misbehaves is having no answer to who validated it.

What breaks#

Fraud models degrade for reasons ordinary models do not. The data-generating process is an adversary who observes your decisions and adapts, so performance decay is not drift but response. Champion-challenger deployment, where a shadow model scores the same traffic without acting on it, is the minimum; holdout groups are the only way to measure true precision, and their cost is the price of knowing.

Two failure modes are structural rather than technical. The first is the feedback loop: a model trained on approved traffic loses sight of the population it declines, grows more confident about a shrinking world, and drifts without any metric moving. The second is metric substitution: a team measured on fraud losses will reduce fraud losses, and the only thing stopping it destroying the business is that somebody else is measured on approval rate. Fraud detection is one of the few disciplines where the correct governance structure is two people with opposed incentives and a shared number.

It is worth closing on what all this exists to change. The Payment Systems Regulator’s dashboard updated on 30 July 2026 recorded that in the eighteen months from 7 October 2024 to 31 March 2026, eighty-eight per cent of in-scope APP scam losses, some £316 million, were reimbursed to victims, with eighty-two per cent of claims closed within five business days and the maximum level of reimbursement set at £85,000 per claim under Specific Direction 20, following the PSR’s Policy Statement PS24/7. Those figures are current at the time of writing in August 2026 and will be superseded. The structure they describe, in which detection quality decides both how much is stolen and how much a firm must hand back, will not be.

51.98 Common wrong ideas#

Wrong: The goal is to stop fraud. Right: Zero fraud is trivially available by refusing every order, and nobody wants it; the goal is the setting of the dial at which fraud losses plus destroyed good business plus operating cost are lowest.

Wrong: A rule that catches most of the fraud is a good rule. Right: Nadia’s address-mismatch rule had seventy-five per cent recall and about five per cent precision, and it destroyed nineteen thousand two hundred pounds of margin to prevent one thousand eight hundred and ninety pounds of loss.

Wrong: You should decline an order when it is probably fraudulent. Right: Decline when the calibrated probability exceeds M divided by (M + L), about thirty-four per cent on the shoe retailer’s numbers, because refusing a good order costs money too.

Wrong: The loss given fraud is the value of the transaction. Right: It is the stock, the dispute fee, the operational cost of the case, the effect on the scheme monitoring ratio and any downstream penalty, and firms modelling it as transaction value alone systematically under-block.

Wrong: A score of eight hundred and fifty out of a thousand means the order is very likely fraud. Right: A score is not a probability until somebody has mapped it onto outcomes, and calibration decays faster than discrimination.

Wrong: Machine learning replaced rules. Right: Rules carry what are not statistical judgements, sanctions and blocklist hits, hard regulatory constraints, an investigator’s freeze, and they must stay deterministic, auditable and changeable at three in the morning by somebody who is not a data scientist.

Wrong: Device fingerprinting identifies the person. Right: It identifies the machine, its available entropy is deliberately shrinking as browser privacy changes land, and in the United Kingdom it engages regulation 6 of the Privacy and Electronic Communications Regulations and needs a documented lawful basis.

Wrong: The merchant bears the fraud loss. Right: Who pays depends on who authenticated, and regulations 76 and 77 of the Payment Services Regulations 2017 place much of it on the payer’s provider, so any tuning conversation that has not established whose money is at stake is a conversation about nothing.

Wrong: You know your false positive rate. Right: Blocked orders generate no outcome at all, so it cannot be measured without deliberately letting some flagged traffic through as a holdout, and a team that has never run one does not know its own rate.

Wrong: A manual review queue is a safety net worth having. Right: It earns its cost only where the human holds information the model lacks, and since Articles 22A to 22D took effect a rubber-stamp queue is arguably worse than none, because it asserts a human involvement that does not exist.

51.99 Chapter summary in 20 lines#

  1. A fraud system is a doorman with three answers, wave in, turn away, or hold for a look, running thousands of times a second with money on it.
  2. It makes two kinds of mistake, and only the fraud it lets through leaves a trace; the good customer it refuses simply goes elsewhere and says nothing.
  3. For Nadia the two numbers are sixty-three pounds lost to a successful fraud and thirty-two pounds of margin lost to a wrongly refused order.
  4. Stopping one fraud is therefore worth about two lost customers, not two hundred, and every design decision in the chapter follows from that ratio.
  5. Her address-mismatch rule caught three-quarters of the fraud and left her nearly eight times worse off, because nineteen of every twenty orders it stopped were innocent.
  6. Precision asks how much of what you stopped was actually bad, recall asks how much of the bad you stopped, and the money was in precision.
  7. Combining four signals that are each innocent alone lifted precision from five per cent to sixty-two and turned a loss into a gain, while deliberately letting twelve frauds a month through.
  8. Zero fraud is available to anyone willing to refuse every order, so the real question was always where the total cost bottoms out.
  9. The optimal decline threshold is M divided by (M + L), about thirty-four per cent here, and it moves with the basket rather than only with the customer.
  10. The workhorse signals are velocity, device fingerprinting, behavioural patterns and the network, and the largest structural advantage in the field is simply having seen the card before.
  11. Every client-supplied signal can be influenced by the party being judged, so the defence is layering cheap client signals against expensive evidence the attacker does not produce.
  12. Fraud is at least four problems, third-party fraud, first-party misuse, account takeover and authorised push payment scams, the last now comparable in United Kingdom scale to the first.
  13. There are five decision points in a card transaction, each with different data, latency and loss-bearer, and United Kingdom law decides where the loss lands at each.
  14. Rules, tree ensembles and graph models fail differently, so running all three with the rules able to override the models is defence in depth.
  15. Velocity features are only as good as the entity resolution beneath them, and point-in-time correctness is what separates a model that works live from one that merely backtests well.
  16. At a fraud rate of 0.4 per cent, moving the false positive rate from one per cent to 0.2 per cent lifts precision from twenty-six to sixty-four while changing nothing about detection.
  17. Labels arrive late and incomplete, so value-weighted metrics belong beside count-weighted ones and a holdout is the only honest measure of false positives.
  18. The loss ratio has two floors, the commercial one where total cost is minimised and a regulatory and scheme one below which exemptions, frictionless authentication and freedom from per-transaction penalties are retained.
  19. A review queue is a third bucket to be sized by arithmetic, and the commonest operational failure is letting the recruitment budget set the risk appetite.
  20. Fraud models decay because the data-generating process is an adversary that watches your decisions, and the correct governance is two people with opposed incentives and a shared number.

Sources, all consulted in August 2026: the FCA Handbook Technical Standards on Strong Customer Authentication, Chapter 3, as updated 19 March 2026; the FCA press release of 19 December 2025 on contactless limits and Handbook Notice 136; Commission Delegated Regulation (EU) 2018/389; the Payment Services Regulations 2017 and the Data (Use and Access) Act 2025 section 80 with SI 2026/82; the PSR’s APP scams reimbursement dashboard of 30 July 2026, PS24/7 and Specific Direction 20; UK Finance’s Annual Fraud Report 2026; the joint EBA and ECB Report on Payment Fraud of 1 August 2024; Mastercard Security Rules and Procedures, Merchant Edition, 11 February 2025; EMVCo Specification Bulletin 279; PRA Supervisory Statement SS1/23; and the EU AI Act Annex III.