Skip to content
KEDBYTE
How Money Moves
Chapter
54

What Breaks and What It Costs

Part V · Trust, Failure and the Law|8,488 words|about 37 min read|Volume 5
Fast-moving material. Figures, model names, prices and version numbers in this chapter were verified in August 2026. Claims are separated into established fact, active research and marketing claim. Re-check anything you intend to rely on.

54.0 What this chapter gives you#

  1. You will be able to explain why a forty-minute terminal outage produces three weeks of consequences, and why the two numbers are not related to each other.
  2. You will be able to do the arithmetic that turns 12,400 uncertain items into roughly 413 hours of manual research, and give senior management a recovery estimate they will not like but cannot argue with.
  3. You will be able to explain why a partial batch is more dangerous than a total failure, and why resubmitting one converts a partial failure into a duplicate.
  4. You will be able to name the five phases of an incident and say which of them firms are habitually bad at.
  5. You will be able to list the clocks that start at once in a United Kingdom incident — four hours from first detection, seventy-two hours for a personal data breach, the end of the next business day for an unauthorised transaction, five business days for a scam claim — and say what triggers each.
  6. You will be able to explain why the regulators fined TSB for the governance around the decision to go live rather than for the fault itself.
  7. You will be able to say which rails can be pulled back and which cannot, and explain that settlement finality is a deliberately created legal state rather than a property of a message.
  8. You will be able to explain why batch file submission has no idempotency key, and why duplicate protection is therefore a control the submitter must build rather than a service the scheme provides.
  9. You will be able to distinguish operational resilience, the critical third parties regime and incident reporting, and explain why the three are one policy.
  10. You will be able to specify a control that alerts on the absence of an expected file, which is the cheapest fix for the most expensive silent failure in the industry.

Every chapter of this book so far has described something working. Cards authorising, files clearing, settlement landing, mandates honoured, disputes adjudicated. That is the right way round, because you cannot understand a failure in a system you do not understand working. But it leaves an impression that needs correcting before we finish, which is that payments is a set of mechanisms that occasionally break.

It is not. Payments is a set of mechanisms that break continually at a low rate, and the entire apparatus around them — the reconciliation, the exception queues, the indemnities, the redress rules, the regulators, the two o’clock in the morning phone trees — exists because that is the normal condition rather than the exceptional one. A payment system is not a machine that works and sometimes fails. It is a machine that fails at a known rate, wrapped in enough correction to make the failures survivable.

The published record is unambiguous about the rate. In March 2025 the House of Commons Treasury Committee published correspondence from nine of the largest banks and building societies operating in the United Kingdom, and the aggregate was at least 803 hours of unplanned technology and systems outage — more than thirty-three days — across at least 158 separate incidents between January 2023 and February 2025. Nine well-capitalised, heavily supervised institutions, in a mature market, with mature controls. That number is not a scandal. It is the baseline.

This chapter is about what happens inside those hours and, much more importantly, in the weeks after them. The thing practitioners learn early and outsiders almost never see is that the outage is the small part. What outlasts it is the reconciliation, and the reconciliation is where the money actually is.

The plain version#

Start with a water main.

A pipe bursts under a road at seven in the morning. The water company gets a crew there by nine and the leak is stopped by half past ten. Three and a half hours, and the problem in the technical sense is over. Nobody thinks the story ends there. The road is dug up for a month. The trench has to be backfilled, the tarmac replaced, the two shops whose cellars flooded dealt with, the claims from people whose cars were damaged assessed one at a time, and somebody has to work out whether the pipe that burst was the one that was supposed to have been replaced in 2019.

Payments incidents are exactly like this, and people consistently misjudge them for exactly the same reason. The visible event — the app is down, the card is declined, the wages have not arrived — is the leak. Fixing the leak is a few hours of engineering. The road is the ledger, and the ledger takes weeks.

Here is a real-shaped example with real arithmetic.

Riverside Court is a care home. It employs eighty-four people and pays them monthly, and the total that leaves its bank account on payday is £176,400, an average of £2,100 each. The way it pays them is old and reliable: on the Wednesday, the payroll clerk sends a list to the bank containing eighty-four lines, each line saying who to pay, how much, and which account. Three working days later the money lands.

One month, the confirmation screen does not load. The clerk waits, refreshes, sees nothing, and sensibly assumes the list did not go. So she sends it again. Both lists were correctly signed, both were validly formatted, and the bank has no way of knowing that the second one is not a genuine second run of wages. On the Friday, £352,800 leaves the account. Every single member of staff is paid twice.

Now watch how long each part takes.

The cause takes about four seconds to explain and about four minutes to fix: the software gains a check that warns you when you are about to send a list identical to one sent in the last twenty-four hours. That is the whole technical remedy. It could be written on a Monday morning.

Detection takes three days, and this is the first surprise. Nobody notices on Friday. The staff certainly do not ring up to complain about being paid twice. The bank does not stop it, because nothing about the second list was invalid. The alarm is not an alarm at all: on Monday the finance manager opens the bank statement for an unrelated reason, sees a balance £176,400 lower than she expected, and has to work backwards from a number to a cause. In payments, this is the ordinary way things are found. Not a red light. A figure that is wrong.

Then comes the part that takes eleven weeks.

The money has gone into eighty-four accounts belonging to eighty-four adults, and it is now legally theirs to argue about. Riverside writes to all of them on the Tuesday. Sixty send it back within a fortnight: £126,000 recovered. Fifteen more come back over the following six weeks, some awkwardly, because they had already spent part of it — one had cleared a credit card, one had paid a deposit on a holiday. That is another £31,500, some of it in instalments. Six agree to repay over four months. Three never pay at all: one has left the country, one disputes being overpaid, one simply stops answering. £6,300 is written off.

Meanwhile the books do not balance, in a specific and irritating way. The payroll system recorded one payroll. The bank recorded two payments to each person. So for eleven weeks there is an unexplained hole that starts at £176,400 and shrinks unevenly as money dribbles back, and every returned payment has to be matched to the right person by hand, because people send money back with references like “wages” and “sorry”. The cash flow forecast was wrong for two months. The overdraft was used, costing £940 in interest and fees. The auditor asks about it the following March, so somebody writes the whole thing up again seven months later.

Add it up. Four minutes of code. Roughly sixty-five hours of finance and payroll time at Riverside. £6,300 written off, £940 of borrowing costs, one very uncomfortable staff meeting, and a company that will be nervous about payday for a year.

The same shape appears when nothing goes out at all. A café’s card terminal stops taking payments at half past twelve on a Saturday and is fixed in forty minutes, and the café loses perhaps £600 of lunchtime trade that never comes back. But some of the failed taps were not clean failures. A handful of customers were charged, saw a decline on the terminal, and paid again with a different card. Sorting out fourteen double charges — finding them, matching them to receipts, refunding them, answering the two customers who go to their banks instead and start a formal dispute — takes the owner three weeks of intermittent effort and one furious online review.

Forty minutes of outage, three weeks of consequences. If you remember one sentence from this chapter, make it that one: the outage is measured in hours, the recovery is measured in weeks, and the two numbers are not related.

One last thing the plain version needs to establish, because it is why all of this is so laborious. When money moves and settles, it is genuinely gone. It is not sitting somewhere temporary waiting to be recalled. Somebody else now owns it, and getting it back requires either their agreement or a court. That is not a design flaw. It is the entire point of a payment: to end the argument about who owes what. A payment you could quietly reverse would not be a payment, it would be a loan with extra steps. The price of finality — the thing that lets the shopkeeper hand over the goods — is that a mistake is expensive to undo.

Where the plain version stops being true#

The clock that matters does not start when you understand the problem. The Riverside story has detection on the Monday and everything flowing from there. In a regulated firm, the regulatory clock is deliberately not tied to your understanding. At the time of writing in August 2026, a UK payment service provider that becomes aware of a major operational or security incident must notify the Financial Conduct Authority without undue delay under regulation 99 of the Payment Services Regulations 2017, with the timing currently specified through the European Banking Authority’s Guidelines on incident reporting under PSD2, EBA/GL/2017/10 as issued on 27 July 2017, applied by SUP 15.14 of the FCA Handbook. That guidance requires an initial report within four hours of first detection. Not from root cause, not from confirmed impact, not from the incident bridge agreeing what to call it. Firms must therefore be able to file a report saying, in effect, something is wrong, we do not yet know what, here is what we do know. The instinct to wait until you can give a clean answer is the instinct that produces a breach.

Duplicate is easier to explain than partial, and partial is far more dangerous. The Riverside example is a duplicate, which is tidy: you know exactly what happened, twice, to exactly whom. Real incidents are more often partial, and partial is worse in a way that is not intuitive. If a batch of 40,000 payments fails completely, you know the state of the world: nothing happened, resend. If 31,600 succeeded, 2,900 failed cleanly, and 5,500 returned no answer at all, you do not know the state of the world and cannot find out from your own records. You have to reconcile against somebody else’s. Until you have, you cannot safely resubmit, because resubmitting the ambiguous 5,500 might duplicate them and not resubmitting them might leave 5,500 people unpaid. A clean total failure is a good day. Ambiguity is the expensive one.

The cost is almost never the outage. Riverside’s direct loss was small. The pattern in published cases is that the operational loss is a rounding error against the redress, the remediation, the enforcement and the supervisory tail. TSB’s platform migration in April 2018 gives the clearest published ratio: the Financial Conduct Authority and Prudential Regulation Authority stated when fining the bank a combined £48,650,000 on 20 December 2022 that all of TSB’s branches and a significant proportion of its then 5.2 million customers were affected by the initial issues, that the problems were not fully resolved for eight months, and that TSB had by then paid £32.7 million in redress. Against that, the engineering hours to restore service were trivial. The asymmetry runs the other way too: the Bank of England’s Real-Time Gross Settlement system was down for roughly nine hours on 20 October 2014, delaying 142,759 CHAPS payments worth £289.3 billion, and the Bank’s published response of 25 March 2015 records that it settled nine compensation claims totalling £4,056.89. A national settlement outage produced about four thousand pounds of claims; a retail migration produced tens of millions. Where the harm lands, not how big the system is, determines the bill.

“The technical fix is the short part” has a real exception, and it is the migration case. When the failure is a design decision rather than a fault — a new platform that does not behave as the old one did, a data model that lost information in transit, a capacity assumption wrong by an order of magnitude — there is no short fix, because there is nothing to restore to. TSB’s disruption ran for eight months. Deloitte’s review found the root cause of the 2014 RTGS outage to be defects introduced in functionality changes made in April 2013 and May 2014, meaning the fault had been latent for around eighteen months before it expressed itself. Incidents caused by a change you deliberately made are a different species from incidents caused by a component that failed, and they behave differently in every dimension: longer, more expensive, more likely to attract enforcement, and much more likely to be attributed to governance rather than technology.

The technical version#

On dating#

Every figure, deadline, threshold and rule in this section is stated as accurate at the time of writing, in August 2026. Operational resilience is one of the fastest-moving areas of financial regulation in the United Kingdom and the European Union, and three specific things were in motion as this was written: the FCA’s and PRA’s new operational incident and third-party reporting regime had been finalised but did not come into force until 18 March 2027; HM Treasury had made its first designations under the critical third parties regime only weeks earlier; and the EU’s Digital Operational Resilience Act had been applicable for less than two years, with its supervisory practice still forming. Verify every number below against the primary source before relying on it.

The anatomy of an incident#

An incident has five phases, and firms tend to be good at the second and third and bad at the first, fourth and fifth.

Detection. The commonest detection mechanism in payments is not an error alert. It is a volume anomaly: throughput that has departed from its expected shape for the time of day, the day of the week and the point in the month. This matters because the failure modes that hurt most produce no errors at all. A file that is never picked up throws nothing. A downstream system that silently truncates a field returns success. A duplicate submission is, by construction, two valid events. Error-rate monitoring finds the loud failures; throughput and reconciliation monitoring find the quiet ones, and the quiet ones cost more. Mature operations therefore alert on the absence of expected traffic — no Bacs input report by 23:00, no settlement file by 06:00, credit volumes 40 per cent below the same weekday last week — rather than only on the presence of errors.

Classification. Classification is a regulatory act, not an engineering one, because it starts clocks. The problem is that it demands an impact estimate at the moment when impact is least knowable. The discipline that works is to classify on plausible worst case and revise downwards later, because every reporting regime tolerates an overstated report far better than a late one.

Containment. A genuine tension runs through every serious incident: stopping the bleeding usually destroys evidence. Failing over destroys the state of the failed node. Restarting a stuck process discards the queue you needed to count. Reprocessing a file overwrites the log that told you what the first attempt did. Good runbooks resolve this in advance by specifying what must be captured before each containment action, so the decision is not being made by a tired engineer at four in the morning who will later be asked, by a skilled person under section 166 of the Financial Services and Markets Act 2000, why there is no record of what happened.

Recovery. Restoring service is not the same as restoring correctness, and conflating them is the classic mistake. Service is restored when payments flow again. Correctness is restored when every payment that should have happened has happened exactly once, and every payment that should not have happened has been reversed or paid for. The gap between those two moments is where the reconciliation lives: typically two to eight weeks for a material incident.

Aftermath. Redress, reporting, root cause analysis, remediation, and the supervisory conversation, which continues long after the engineering is closed.

Command structure is where competent firms differ visibly from incompetent ones. There is one incident commander, who does not fix anything. There is a separate communications lead, because the person diagnosing must not also be the person drafting the customer statement. There is a scribe keeping a timestamped log, because the log is the evidence and cannot be reconstructed afterwards. And there is a named accountable senior manager: in the United Kingdom, operational resilience typically sits with the Chief Operations function, SMF24, under the Senior Managers and Certification Regime, which means the failure has a person’s name attached to it in a way that predates the incident.

The clocks#

The single hardest operational fact about a modern incident is that several unrelated deadlines start at once, they are triggered by different events, and they are owed to different people.

Obligation Trigger Deadline Basis, as at August 2026
UK major payments incident, initial First detection 4 hours PSRs 2017 reg. 99, SUP 15.14, EBA/GL/2017/10
UK personal data breach Awareness 72 hours UK GDPR Art. 33
EU major ICT incident, initial Classification as major 4 hours, and no later than 24 hours from becoming aware DORA Art. 19(4)(a)
EU major ICT incident, intermediate Initial notification 72 hours DORA Art. 19(4)(b)
EU major ICT incident, final Intermediate report 1 month DORA Art. 19(4)(c)
Refund of unauthorised transaction Notification by payer End of the next business day PSRs 2017 reg. 76
APP scam reimbursement Claim 5 business days, subject to stop the clock PSR reimbursement requirement, from 7 Oct 2024
UK incident report, from 18 Mar 2027 Threshold determination Ordinarily 24 hours, but 4 hours from first detection for PSPs FCA PS26/2, FG26/3; PRA PS7/26, SS1/26

Two of those deserve elaboration.

DORA — Regulation (EU) 2022/2554, applicable since 17 January 2025 — sets its initial four-hour clock from classification rather than detection, but caps it by requiring notification no later than twenty-four hours from becoming aware of the incident. Classification itself must be done without undue delay, against the criteria in the Commission Delegated Regulation on classification of major ICT-related incidents. The design intent is to stop firms deferring the clock by deferring the decision.

The United Kingdom is replacing its arrangements. On 18 March 2026 the FCA published PS26/2, with finalised guidance FG26/3 on operational incident reporting and FG26/4 on material third-party reporting, alongside the PRA’s PS7/26 and supervisory statement SS1/26. The rules come into force on 18 March 2027. They create a single definition of an operational incident, a single template and a single submission route through FCA Connect for firms that today report separately to two or three authorities. An operational incident is a single event or a series of linked events which disrupt the firm’s operations such that they disrupt delivery of a service to an end user external to the firm, or affect the availability, integrity, authenticity or confidentiality of that end user’s data. Reporting is split into standard and enhanced tiers, with payment service providers in the enhanced tier and, crucially, retaining the accelerated four-hour deadline from first detection while other enhanced firms work to twenty-four hours from threshold determination. The EBA guidelines will be disapplied for UK PSPs, who will discharge regulation 99 solely through the new regime. Final reports are due within thirty working days of resolution, or no later than sixty with reasons for the delay.

Failure modes, and why some are worse than others#

Duplicate submission and duplicate settlement. The commonest serious payments incident, and the one with the worst recovery profile, because a duplicate is by definition a valid instruction. The largest published UK example is Santander’s Christmas Day 2021 incident, in which approximately £130 million was paid out in around 75,000 duplicate transactions originating from about 2,000 business accounts. Recovering a duplicate credit means recovering funds from third parties who have no contractual relationship with the party that made the mistake, which is why the money often does not come back.

Stuck file. The instruction was created but never delivered, or delivered and never picked up. Zero errors, zero alarms, zero payments. Only a positive control catches this: a reconciliation that expects a specific file, of roughly a specific size, by a specific time, and alerts on its absence.

Partial batch. Some items processed, some did not, and the boundary is unknown. This is the failure mode that most often produces the wrong recovery action, because the temptation to resubmit the whole file is enormous and doing so converts a partial failure into a duplicate.

Silent corruption. Field-level truncation and encoding faults are especially nasty in the file-based world because the legacy formats are fixed-width and unforgiving. In the Bacs Standard 18 format, the fields carrying the counterparty’s name and the payment reference are eighteen characters. Anything that overruns is cut, and a cut reference is a payment that arrives but cannot be applied. The sender sees success. The receiver sees an unidentifiable credit and a reconciliation break.

Idempotency, and why batch payments lack it. Modern payment APIs are built around an idempotency key: a client-generated identifier that lets the server recognise a retry and return the original result rather than performing the action twice. That solves the Riverside problem completely. Batch file submission has no equivalent primitive. A Bacs submission is authenticated and non-repudiable, and is therefore trusted; it is not deduplicated, because two genuine payroll runs to the same eighty-four people for the same amounts on consecutive days is a legitimate business event no scheme can distinguish from an error. Duplicate protection in the file world is a control submitters must build themselves.

Reversal asymmetry#

Whether a payment can be pulled back is not a question about technology. It is a question about which rail it went through, and the answers differ sharply.

Rail Reversal mechanism Practical character, as at August 2026
Bacs Direct Credit No unilateral reversal; recall by request to the receiving PSP, requiring the beneficiary’s consent Weak. Success depends on goodwill and on the money still being there
Bacs Direct Debit Direct Debit Indemnity Claim under the Direct Debit Guarantee Strong for the payer. The paying PSP refunds first and recovers from the originator afterwards, with a challenge process available to the service user
Faster Payments Recall request; no guaranteed return Weak, and fast money is fast in both directions
CHAPS Irrevocable between direct participants once settled None, absent the beneficiary’s agreement
Card Chargeback under scheme rules, within scheme-defined windows Strong for the cardholder, structured, and adversarial

Underneath all of this sits settlement finality. The Financial Markets and Insolvency (Settlement Finality) Regulations 1999, S.I. 1999/2979, allow payment systems to be designated so that transfer orders effected through them are protected from being unwound by insolvency law. CHAPS was designated in 2000 and is a recognised interbank payment system under Part 5 of the Banking Act 2009. Finality is not a technical property of a message; it is a legal state, deliberately created, that exists precisely so that a settled payment cannot be reopened by later events.

The consequences of that design are best shown by the largest published misdirected-payment case of recent years. On 11 August 2020, Citibank, acting as administrative agent on a Revlon term loan, sent lenders the full outstanding principal when it had intended to send only an interest payment. Some returned the money. Others did not. The sum in dispute — principal and accrued interest owed to the lenders who declined to return it — was $558,558,375.74. The United States District Court for the Southern District of New York held in February 2021 that the discharge-for-value rule gave those lenders a defence against restitution, because they were owed exactly what they received and had no notice of the mistake. The Second Circuit vacated that ruling on 8 September 2022 in In re Citibank August 11, 2020 Wire Transfers, No. 21-487, holding among other things that the debt was not then due and that the recipients had constructive notice. The point for a practitioner is not the eventual outcome. It is that recovering an erroneous payment of over half a billion dollars from sophisticated counterparties took two years and two courts, and at the halfway mark the money was gone. Payment systems are built to make payments stick. They succeed.

The reconciliation aftermath#

This is the part of an incident that no press release describes and every practitioner remembers.

Reconciliation after an incident is a three-way problem, not a two-way one. There is the scheme’s or acquirer’s view, expressed in settlement and clearing files; the bank’s view, expressed as movements on the settlement or client account; and the firm’s own ledger, its record of what it believed it did. In normal operation these three agree to within a small and stable break rate. After an incident they do not, and the work is to reduce the disagreement to zero, item by item, with each item explained.

Breaks fall into recognisable categories, and the category determines the remedy:

Break type Meaning Typical remedy
In the ledger, not at the bank We think we paid; no money moved Reissue after confirming non-payment
At the bank, not in the ledger Money moved; we have no record of instructing it Investigate as potential duplicate or unauthorised
Matched, wrong amount Both agree it happened, disagree on value Adjust, then find the truncation or rounding fault
Matched, wrong date Value-dating disagreement Usually interest and reporting impact only
Unidentified receipt Money arrived, cannot be applied Suspense, then manual research, then escalating attempts to contact
Duplicate Two records, one intended event Recall attempt, then recovery, then write-off

Unmatched items go to a suspense account, which is not a solution but a holding pen with a clock on it. Two disciplines matter: ageing, because a two-day break is operations and a ninety-day break is a provision, and named ownership, because unowned breaks are immortal.

The reason this outlasts the outage is arithmetic. Suppose an incident leaves 12,400 items in an uncertain state. Automated matching on amount, date, counterparty and reference will typically resolve 85 to 95 per cent of them, so call it 1,240 items left. Manual research on a payments break — pulling the original instruction, the scheme record and the bank movement, and forming a view — averages fifteen to twenty-five minutes when straightforward and runs to hours when it involves contacting a beneficiary. At twenty minutes each, 1,240 items is roughly 413 hours: eleven weeks for one full-time person, or three weeks for a team of four who also have their normal work to do. Then there is the tail, the items requiring a third party’s cooperation, which arrive back over months, and the small residue that never resolves and becomes a write-off with an audit trail.

That is why a forty-minute outage produces a six-week recovery, and why the honest answer to “when will this be finished?” on the day of an incident is a number that senior management will not like.

Redress in the United Kingdom is not discretionary generosity. Several distinct regimes apply at once, and they attach to different failures.

For unauthorised transactions, regulation 76 of the Payment Services Regulations 2017 requires the payer’s payment service provider to refund the amount as soon as practicable and in any event no later than the end of the business day following the day on which it becomes aware of, or is notified of, the transaction, and to restore the account to the state it would have been in. There are exceptions where the provider has reasonable grounds to suspect fraud and reports them appropriately. Regulations 91 to 95 deal with non-execution, defective execution and late execution, allocating liability between the payer’s and payee’s providers and requiring the payment to be traced.

For authorised push payment scams, the Payment Systems Regulator’s reimbursement requirement has applied to eligible Faster Payments claims since 7 October 2024, with the Bank of England setting the equivalent for CHAPS. Reimbursement is due within five business days, subject to a stop-the-clock provision where the sending provider needs more information, and the maximum mandatory level is £85,000 per claim, a figure the PSR settled on having previously proposed £415,000. Firms may reimburse more, and losses above it can be taken to the Financial Ombudsman Service.

For everything else — the distress and inconvenience of an outage, the consequential losses, the missed mortgage completion, the declined card at the till — there is complaint handling under DISP in the FCA Handbook, the Financial Ombudsman Service, and since July 2023 the Consumer Duty, which requires firms to act to deliver good outcomes for retail customers and specifically covers consumer support. At the time of writing, the Ombudsman’s award limits rose on 1 April 2026 to £455,000 for complaints about acts or omissions on or after 1 April 2019, and £205,000 for earlier ones.

The scale of voluntary distress payments is now public record because the Treasury Committee asked. Barclays confirmed that during its outage from 31 January to 2 February 2025, 56 per cent of online payments failed as a result of severe degradation of mainframe processing performance, and that it expected to pay between £5 million and £7.5 million for inconvenience or distress, taking its total potential payout over the two-year period to as much as £12.5 million. The second-highest figure among the nine firms was £350,000, from Bank of Ireland. One very large number, then a cliff: these payments are driven by the severity and timing of individual events, not by aggregate downtime. The Barclays outage landed on payday.

Four published incidents, and what each one teaches#

Bank of England RTGS, 20 October 2014. Roughly nine hours of outage in the system that settles CHAPS. Operating hours were extended to 20:00 and all submitted payments, 142,759 of them worth £289.3 billion, settled the same day. Deloitte’s independent review, published with the Bank’s response on 25 March 2015, found the root cause to be defects introduced as part of functionality changes made in April 2013 and May 2014. The Bank aims for 99.95 per cent availability of RTGS settlement services to CHAPS, and had achieved 100 per cent from 2008 to 2013. Nine compensation claims were settled, totalling £4,056.89. The lessons the Bank accepted were almost entirely about governance, testing and crisis management rather than software: reconstituting the RTGS board under a Deputy Governor, deferring non-routine changes, separating test and pre-production environments, and reducing the barriers to invoking the MIRS contingency solution so that it could switch and fix in parallel rather than treating failover as a last resort. That last point is the most transferable. A contingency you are afraid to invoke is not a contingency.

Visa Europe, 1 June 2018. Charlotte Hogg, then Visa’s European chief executive, wrote to the Treasury Committee explaining that the disruption ran from 14:35 on 1 June to 00:45 the following day, that 51.2 million transactions were initiated across Europe in that window and 5.2 million failed, and that 2.4 million UK transactions failed to process. The cause was a very rare partial failure of a component within a switch at the primary of two data centres, either of which was designed to handle 100 per cent of European volume alone. The partial nature of the failure is the whole lesson: because the primary had not cleanly failed, the backup switch did not activate, and the malfunctioning primary kept trying to synchronise messages with the secondary site, creating a backlog that degraded the secondary’s own processing. It took nearly five hours to fully deactivate the failing system. Redundancy protects against a component that stops. It does not automatically protect against a component that continues, badly.

TSB, April 2018. A platform migration in which, as the regulators put it, the data itself migrated successfully but the platform immediately experienced technical failures. On 20 December 2022 the FCA fined TSB £29,750,000 and the PRA £18,900,000, a combined £48,650,000 after a 30 per cent early-resolution discount, for operational risk management and governance failures including management of outsourcing risk. Note what was penalised. Not the fault. The governance around the decision to go live.

CrowdStrike, 19 July 2024. At 04:09 UTC a faulty content configuration update to the Falcon sensor caused an out-of-bounds memory read in the Windows sensor client, crashing machines into a boot loop. The update was reverted at 05:27 UTC, seventy-eight minutes later, and machines that booted after the reversion were unaffected. Microsoft estimated roughly 8.5 million Windows devices were affected, less than one per cent worldwide. Recovery took days, because many machines needed individual physical intervention and could not be fixed remotely by the very software that had broken them. This is the chapter’s central ratio at global scale: the vendor’s fix took seventy-eight minutes and the world took a week. It is also why the concentration of financial firms on a handful of shared suppliers became a regulatory priority rather than a procurement one.

The regulatory frame around resilience#

Three regimes now sit on top of the failure modes described above, and they are worth distinguishing because firms routinely conflate them.

Operational resilience is about outcomes, not systems. Under the FCA’s PS21/3 and the corresponding PRA rules, firms must identify their important business services, set an impact tolerance for each — the maximum tolerable level of disruption, expressed as a duration alongside other relevant metrics, beyond which further disruption would cause intolerable harm to consumers or risk to market integrity — and map and test to the point where they can remain within those tolerances in severe but plausible scenarios. Services had to be identified and tolerances set by 31 March 2022, and the transition period for being able to stay within them ended on 31 March 2025. The conceptual shift is from availability targets, which are about systems, to impact tolerances, which are about customers and which assume disruption will happen rather than trying to prevent it.

Critical third parties. Section 312L of the Financial Services and Markets Act 2000, as amended by the Financial Services and Markets Act 2023, allows HM Treasury to designate a third-party service provider as a critical third party where, in its opinion, failure or disruption of its services to firms could threaten the stability of, or confidence in, the UK financial system. The joint rules of the Bank of England, PRA and FCA, published on 12 November 2024 as PRA PS16/24 and FCA PS24/16, took effect on 1 January 2025 but bite on a provider only when a designation order comes into force. HM Treasury announced its first designations on 10 July 2026, effective 13 July 2026: Microsoft Ireland Operations Limited, Google Cloud EMEA Limited, Amazon Web Services EMEA SARL and Oracle Corporation UK Limited. Two features are commonly misunderstood. Oversight extends only to the systemic services provided to the financial sector, not to the provider’s wider business. And designation transfers nothing: firms remain fully accountable for the risks in their own third-party arrangements, and a designated provider is not thereby safer than an undesignated one.

Incident and third-party reporting is the data layer that makes the other two supervisable. PS26/2 and PS7/26 also require in-scope firms to notify material third-party arrangements before commitments are finalised and to submit an annual register of them, explicitly so that the regulators can identify concentration risk and inform future critical third party designations. The three regimes are one policy.

What good operational practice actually looks like#

Strip away the frameworks and the practices that distinguish firms which recover well are unglamorous and specific.

Design for reconciliation before designing for throughput. Every payment instruction should carry an identifier that survives every hop and appears in every record on both sides, so that three-way matching is a lookup rather than an investigation. Retrofitting this after an incident is close to impossible, which is why it has to be a build-time decision.

Make duplicate prevention a control, not a hope. Hash every outbound file and compare it against submissions from the preceding days. Maintain a submission register that operations staff, not just engineers, can read. Require a second authoriser for any submission matching a recent one in value and item count. Set value limits with the sponsoring bank so that an anomalous total cannot be submitted at all. Riverside’s four minutes of code is the entire fix, and almost nobody writes it until after the first time.

Alert on absence, not only on errors. The expected file that never arrives is the most expensive silent failure in the industry.

Rehearse the failover you are afraid of. The Bank of England’s own accepted recommendation after 2014 was to reduce the barriers to invoking its contingency solution so that it could switch and fix in parallel. A failover that has never been executed in anger, or that requires a decision nobody wants to make, is a document rather than a capability. This is what the impact-tolerance regime is really testing.

Pre-agree the recovery paperwork: recall templates, beneficiary letters, the customer notice, the redress matrix saying what is paid for what category of harm. Drafting these during an incident guarantees they will be slow, inconsistent and later criticised.

Separate the restore-service team from the restore-correctness team on day one. They have different skills, timescales and definitions of done, and if they are the same people the reconciliation starts two weeks late because everyone is exhausted.

Write a blameless post-incident review with a real timeline, publish the actions with owners and dates, and then check them. The commonest finding in enforcement notices is not that a firm failed to identify a weakness. It is that the firm identified it, wrote it down, and did not fix it in time. Keep the incident log as if a regulator will read it, because one may: contemporaneous, timestamped, factual, including the decisions not taken and why.

Returning to the tap#

This book opened in a shop in Leeds, with a card touched against a terminal, and with the observation that almost nobody’s description of what happens next is correct. Money did not move. No object travelled. A record was edited, and everyone accepted the edit because everyone accepted the record-keeper.

Fifty-three chapters later, you know what the edit involves. That the terminal runs an application which must satisfy a kernel specification, that the token in the phone is not the card number, that the authorisation travels through an acquirer to a scheme to an issuer and back inside a second or two, and that it is a promise rather than a payment. That the movement happens hours later in a settlement file, and finality later still across accounts at the Bank of England. You know who is liable if the card was stolen, who is liable if the shopkeeper never delivers, what the interchange fee is and why it is capped, why the merchant’s money arrives on Wednesday, and which regulator would care if it did not.

You also now know the thing the tap is designed to conceal. Behind that gesture sit several dozen organisations, any of which can fail; a legal framework whose entire purpose is to make the result stick even when somebody has made a mistake; a reconciliation apparatus that quietly resolves millions of small disagreements every day; and people whose job, when the disagreements stop being small, is to spend six weeks matching numbers to numbers until the two sides of a ledger agree again.

In 2024, according to UK Finance’s UK Payment Markets report, nearly 49 billion payment transactions were made in the United Kingdom, of which 18.9 billion were contactless. Pay.UK’s figures for 2025 show approximately 5.03 billion Direct Debits and 1.83 billion Bacs Direct Credits, including around 378 million payroll payments and 267 million state pension payments, running through a system whose cycle is still three working days and whose input files are still eighteen characters wide in the fields that matter. The overwhelming majority of those payments worked. The ones that did not were caught, in most cases, by somebody looking at a number that was wrong.

That is what money moving actually is. Not a substance travelling, but a very large number of records being written, checked against each other, and corrected. Ada’s notebook, with more lawyers.

The tap takes half a second. Everything in this book is what makes the half-second safe to trust, and everything in this chapter is what happens on the days when it is not. Both halves are the system. A reader who understands only the first half understands payments the way somebody who has only watched aircraft take off understands aviation.

Now you have seen the landings too.

54.98 Common wrong ideas#

Wrong: The regulatory clock starts when you understand what has gone wrong. Right: The initial report is due within four hours of first detection, not from root cause or confirmed impact, so a firm must be able to file a report that says something is wrong, we do not yet know what, here is what we do know.

Wrong: A complete failure is the worst outcome. Right: A clean total failure tells you the state of the world and can be resent; ambiguity is the expensive one, because items that returned no answer at all can be neither safely resubmitted nor safely abandoned.

Wrong: The cost of an incident is the outage. Right: The operational loss is usually a rounding error against redress, remediation, enforcement and the supervisory tail — TSB’s combined fine was £48,650,000 against £32.7 million already paid in redress, while nine hours of national settlement outage in 2014 produced nine claims totalling £4,056.89.

Wrong: The bigger the system that failed, the bigger the bill. Right: Where the harm lands determines the bill, which is why a retail migration cost tens of millions and an RTGS outage cost four thousand pounds.

Wrong: The technical fix is always the short part. Right: Where the failure is a design decision rather than a fault there is nothing to restore to, which is why TSB’s disruption ran for eight months and why the 2014 RTGS defects had been latent for around eighteen months.

Wrong: The incident is over when service is restored. Right: Service is restored when payments flow again; correctness is restored when every payment that should have happened has happened exactly once, and the gap between those two moments is typically two to eight weeks.

Wrong: Redundancy protects against the failure of a component. Right: It protects against a component that stops, not one that continues badly — Visa’s primary switch in June 2018 failed partially, so the backup never activated and the malfunctioning primary degraded the secondary site as well.

Wrong: Error-rate monitoring will find the incident. Right: The failure modes that cost most produce no errors at all, so mature operations alert on the absence of expected traffic rather than only on the presence of errors.

Wrong: A settled payment sent by mistake can be pulled back. Right: Finality exists precisely to stop that, which is why recovering Citibank’s $558,558,375.74 of mis-sent Revlon principal took two years and two courts, with the money gone at the halfway mark.

Wrong: Designating a cloud provider as a critical third party makes the firms relying on it safer. Right: Oversight extends only to the systemic services provided to the financial sector, designation transfers nothing, and firms remain fully accountable for the risks in their own third-party arrangements.

54.99 Chapter summary in 20 lines#

  1. A payment system is not a machine that works and sometimes fails, but a machine that fails at a known rate wrapped in enough correction to make the failures survivable.
  2. Nine of the largest United Kingdom banks reported at least 803 hours of unplanned outage across at least 158 incidents between January 2023 and February 2025, which is the baseline rather than a scandal.
  3. The visible event is the leak and the ledger is the road, and the road takes weeks after the leak is stopped.
  4. Riverside Court’s duplicated payroll of £176,400 needed four minutes of code to prevent, took three days to detect and eleven weeks to reconcile, and ended in £6,300 written off plus £940 of borrowing costs.
  5. Detection in payments is ordinarily a wrong number noticed for an unrelated reason, not a red light.
  6. Once money has moved and settled it is genuinely gone, and getting it back requires the recipient’s agreement or a court, because ending the argument about who owes what is the whole purpose of a payment.
  7. Several unrelated regulatory clocks start at once, triggered by different events and owed to different people, and the United Kingdom four-hour clock runs from first detection.
  8. Classification is a regulatory act rather than an engineering one, because it starts those clocks at the moment impact is least knowable.
  9. Containment usually destroys evidence, so good runbooks specify in advance what must be captured before each containment action.
  10. Restoring service is not restoring correctness, and conflating them is the classic mistake.
  11. Partial failures are worse than total ones because the state of the world cannot be established from the firm’s own records, and the temptation to resubmit the whole file converts a partial failure into a duplicate.
  12. Batch file rails have no idempotency key, because two identical payroll runs on consecutive days are a legitimate business event no scheme can distinguish from an error, so duplicate prevention is a control submitters must build themselves.
  13. Whether a payment can be reversed depends on the rail: Direct Debit and cards give strong structured remedies, Bacs credits and Faster Payments give only a request, and CHAPS gives none once settled.
  14. Post-incident reconciliation is a three-way problem between the scheme’s view, the bank’s view and the firm’s own ledger, and every break must be categorised, aged and given a named owner.
  15. The arithmetic is unforgiving: 12,400 uncertain items, ninety per cent matched automatically, twenty minutes each on the remainder, is about 413 hours before the tail and the write-offs even begin.
  16. Redress is not discretionary generosity but a stack of regimes — regulation 76 for unauthorised transactions, regulations 91 to 95 for defective execution, the £85,000 scam reimbursement requirement, DISP, the Ombudsman and the Consumer Duty.
  17. Barclays expected to pay between £5 million and £7.5 million in distress payments for one outage that landed on payday, while the next highest figure among nine firms was £350,000, because severity and timing drive these payments rather than aggregate downtime.
  18. The four published incidents teach four distinct lessons: a contingency you are afraid to invoke is not a contingency, redundancy does not protect against a component that continues badly, regulators penalise the governance of the decision rather than the fault, and a seventy-eight-minute vendor fix can take the world a week.
  19. Operational resilience, the critical third parties regime and the new incident and third-party reporting rules coming into force on 18 March 2027 are three parts of one policy, the last being the data layer that makes the other two supervisable.
  20. Money moving is not a substance travelling but a very large number of records being written, checked against each other and corrected, and this chapter is what happens on the days when the correcting becomes visible.

Sources, all consulted in August 2026: the Payment Services Regulations 2017, regulations 76, 91 to 95 and 99, and FCA Handbook SUP 15.14; EBA Guidelines EBA/GL/2017/10; FCA PS26/2 with FG26/3 and FG26/4, and PRA PS7/26 with SS1/26, published 18 March 2026 and in force 18 March 2027; FCA PS21/3 and the FCA’s operational resilience insights pages; PRA PS16/24 and FCA PS24/16 on critical third parties, with HM Treasury’s designation announcement of 10 July 2026; section 312L of the Financial Services and Markets Act 2000 as amended by FSMA 2023; Regulation (EU) 2022/2554 (DORA), Article 19; the Financial Markets and Insolvency (Settlement Finality) Regulations 1999, S.I. 1999/2979; the Bank of England’s response to the independent review of the RTGS outage of 20 October 2014, published 25 March 2015; the Bank of England and FCA announcement of the TSB fine, 20 December 2022; Visa’s letter to the Treasury Committee of 15 June 2018; the Treasury Committee’s publication of bank IT failure correspondence, 6 March 2025; Payment Systems Regulator PS24/7 and its APP reimbursement pages; Financial Ombudsman Service award limits for 1 April 2026; Pay.UK’s Bacs glossary and Bacs annual processing statistics 2025; UK Finance UK Payment Markets 2025; Microsoft’s statement of 20 July 2024 and CISA’s alert of 19 July 2024; and In re Citibank August 11, 2020 Wire Transfers, No. 21-487 (2d Cir., 8 September 2022).