Skip to content
KEDBYTE
How Money Moves
Chapter
33

The Physical Network

Part III · The Networks|8,631 words|about 38 min read|Volume 3

33.0 What this chapter gives you#

  1. You will be able to name the five roles a card authorisation passes through, and say which two of them are scheme members and therefore the only parties that can be fined or hold a settlement position.
  2. You will be able to account for the 790 milliseconds of a contactless payment segment by segment, and explain why the card network is about five per cent of it while the phone and the reader are almost two thirds.
  3. You will be able to separate the parts of a latency budget that are physics from the parts that are software, and say why an issuer processor in Virginia costs a London transaction about seventy-five milliseconds that no amount of money buys back.
  4. You will be able to explain why a payment HSM is a computer that refuses rather than a safe that hides, and why zeroising one makes every key the institution holds permanently unreadable.
  5. You will be able to trace a PIN block through translation at the acquirer, the scheme and the issuer, and say how many times the clear PIN exists and where.
  6. You will be able to explain why a timeout is a financial event rather than a performance event, and what the 0400 reversal is for.
  7. You will be able to say why partial failure is more dangerous than total failure, using Visa Europe’s ten hours and ten minutes on Friday 1 June 2018 as the worked example.
  8. You will be able to answer a due diligence question about HSM certification by naming the validated module and its certificate rather than the appliance.
  9. You will be able to state the three phases of PCI PIN Security Requirement 18-3 and their dates, and say whether an estate still shipping bare encrypted keys to terminals is compliant.
  10. You will be able to explain why moving scheme endpoints into the public cloud improves one firm’s resilience while concentrating everybody’s risk, and name the regime that now oversees that concentration.

At twenty to nine on a Tuesday morning a woman buys a sausage roll and a coffee in a bakery on Kirkstall Road in Leeds. Four pounds twenty. She holds her phone against the reader, the reader beeps, a green tick appears, and she is out of the door before the receipt has finished printing. The whole thing takes about a second.

Everything in this volume so far has described what was in that second as messages. A 0100 authorisation request with a bitmap and a set of data elements; a 0110 response carrying an approval in DE39 and a six-character code in DE38. That description is true and it is the one you need in order to reason about the system. It is also, in a specific sense, weightless. Messages do not travel. Optical and electrical signals travel, along glass, through switches, into racks in buildings with concrete walls and diesel tanks.

This chapter is about the buildings. It is about which company owns which machine, what happens inside the box that no engineer is permitted to open, where the second actually goes, and what a merchant sees when the whole apparatus stops working on a Friday afternoon.

The plain version#

A chain of desks#

Imagine that the bakery cannot approve the payment itself, so it has to ask. The question travels down a chain of desks, and each desk does one job and hands it on.

At the first desk is the card reader on the counter. Its job is to read the phone, work out that the shop wants £4.20, write all of that down on a slip of paper in a very compressed shorthand, and hand it on.

The second desk belongs to the company the bakery pays to handle its card money: a building full of computers somewhere in England. It checks that the slip is properly filled in, works out which card network the payment belongs to, translates the shorthand into the shorthand that network prefers, and passes it on.

The third desk is the card network itself. It has no opinion about whether this payment is a good idea. Its job is to look at the front of the card number, work out which bank in the world issued it, and route the slip there. It is a sorting office, not a decision maker.

The fourth desk belongs to the bank’s card system, and this is where the decision happens. Is this really her phone? Does the cryptographic signature the phone produced check out? Has she got £4.20 available? Has anything about this purchase looked strange in the last five minutes?

The fifth desk is the bank’s actual accounts system: the ledger with her name on it and a balance next to it. Then the answer comes back down the chain in reverse, and the reader beeps.

The stopwatch#

Now put real numbers on it, because the numbers are the point. Every figure below is in thousandths of a second — milliseconds.

Step Time
Phone held against the reader, chip does its cryptographic sum 500
Reader to the payments company’s building 25
Payments company checks and translates the slip 10
Payments company to the card network 5
Card network sorts and routes 20
Card network to the bank’s card system 10
Bank decides 120
Bank’s card system back to the card network 10
Card network back 20
Card network to the payments company 5
Payments company to the reader 25
Reader shows the tick and beeps 40
Total 790

Just under eight tenths of a second, and look at how it is spent. Almost two thirds of it is the phone and the reader talking to each other before anything leaves the shop at all. The card network — the thing everybody pictures when they imagine “the payment system” — accounts for forty of those eight hundred milliseconds, or five per cent. The travelling itself, all six legs of it, is seventy milliseconds. The single biggest slice after the phone is the bank thinking. So when a shopkeeper says the card machine is slow, the card network is almost never what they are complaining about.

Why there is spare time in the budget#

People in the industry talk about a payment needing to complete in about two seconds. Eight hundred milliseconds is well inside that. The remaining twelve hundred are not waste; they are margin, and the margin exists because of specific, nameable things that go wrong. The bank’s computers might be busy. The shop’s broadband might be having a bad minute and a packet might have to be sent twice. The bank might not be in Britain.

That last one matters more than you would think, and it is the one thing in the entire chain that money cannot fix. Light in a glass fibre travels at about two hundred thousand kilometres per second, slower than light in a vacuum because glass slows it down. Leeds to London and back is roughly three milliseconds of pure travel. London to Virginia and back, which is where a great many card systems physically live, is about seventy-five. If the bank’s computers are in Virginia, seventy-five milliseconds of your budget is gone before anybody has done any work, and no amount of money buys it back, because you cannot buy a faster universe.

The locked box#

There is one desk in the chain where the person does not do the work themselves.

When the bank checks the customer’s PIN, or checks the cryptographic signature the chip produced, it needs secret keys. Those keys are the crown jewels: anyone holding them could forge cards for millions of customers. So the bank does not keep them in an ordinary computer. It keeps them in a metal box bolted into a rack in a locked cage inside a locked room.

The box works like this. You post a question through a slot: “here is an encrypted PIN and here is the customer’s account number — is the PIN right?” The box does the sum inside itself and posts out a single word: yes or no. What it will never do, no matter how you ask, is post out the PIN. There is no command for that. The engineers who built it deliberately left it out.

And if anybody tries to open the box, drill it, freeze it, or lift it off the rack, it does not resist. It erases itself, and what the thief carries away is an expensive paperweight. That is the opposite of how people imagine security. The box is not strong. The box is suicidal.

Two of everything#

Because all of this is a chain, and because a chain stops if any link stops, everybody in it builds two of everything. Two card readers behind the counter. Two internet connections into the payments company. Two data centres, in two towns, each big enough to carry the whole load on its own.

And here is the thing a twelve-year-old works out faster than an adult: the frightening failure is not one building burning down. If a building burns down, everyone knows, and the other building takes over. The frightening failure is one building going half wrong. Still answering the phone. Still saying it is fine. Still sending confused messages to the other building. Nobody is sure whether to switch over, because switching over is itself dangerous, and while everybody argues, the queue at the bakery gets longer.

That is not hypothetical. It is what happened to Visa in Britain on Friday 1 June 2018, and we will walk through it.

Where the plain version stops being true#

The first correction is that the five desks are five jobs, not five companies, and often not five buildings. The chain above is a chain of roles, and roles do not map one-to-one onto companies. Adyen and Stripe are the gateway, the licensed acquirer and the acquiring processor in one legal and technical stack, so three of my five hops become function calls between processes in the same estate. Conversely, one apparent hop routinely conceals six: what I drew as “the payments company checks the slip” may be a gateway, a tokenisation vault, a risk engine, a switch, a scheme endpoint and a message-repository write, each with a separate failure mode. Do not count boxes. Count trust boundaries — the places where data crosses from one organisation’s liability into another’s — and count clocks, because every boundary has a timer attached to it.

The second correction is that “two seconds” is a convention, not a specification. No card scheme publishes a rule saying an authorisation must complete within two seconds. What exists instead is a nest of separate timers with separate owners expiring at different moments: the terminal’s timer waiting on its acquirer, the acquirer’s on the scheme, the scheme’s on the issuer, and the entirely unregulated timer running behind the customer’s eyes. These disagree on purpose, and the consequences of each expiring differ. When the innermost one expires you do not get a slow approval; you get a different transaction — a reversal, a stand-in decision made by somebody who is not your bank, or a decline. Timeouts in payments are financial events, not performance events.

The third correction is that the locked box is a computer that refuses, not a safe that hides. The security of a hardware security module is not principally about concealment; it is about a deliberately impoverished instruction set. The device will take an encrypted PIN block and re-encrypt it under a different key, because that operation is necessary; it will not decrypt one and hand you the result, because that command was never implemented. Related, and routinely misunderstood: most keys are not stored inside the box. A large issuer’s millions of card keys live in an ordinary database, each encrypted under a master key, and only that master key lives inside the tamper boundary. The box is not a vault holding your keys; it is the only place where your keys can be turned back into usable form, and only during an operation whose output you are permitted to see. Erasing it does not destroy your keys. It makes every copy of them, everywhere, permanently unreadable, which is worse.

The fourth correction is that “we have two data centres” is a claim about capacity, and the thing that fails is the switching. Building a second site that can carry the full load is expensive but conceptually simple. Detecting that the first site is unwell, distinguishing “unwell” from “briefly busy”, and committing to a cutover that will itself cause disruption, is neither. The industry’s own six-nines language conceals this. Availability of 99.9999 per cent sounds like a promise; it is 31.6 seconds of downtime in a year. The Visa Europe outage described later ran for ten hours and ten minutes: roughly one thousand one hundred and fifty years of a six-nines budget consumed in one afternoon. Such a number states design intent about steady-state operation. It says nothing about the tail, and the tail is what puts queues in bakeries.

The technical version#

The five roles, stated exactly#

The card industry’s role names are used loosely in conversation and precisely in contracts. Here they are precisely.

Role What it is Licensed by a scheme? Physical artefact
Terminal estate POI devices on merchant counters, unattended units, softPOS on phones No, but devices are PCI PTS approved The PIN entry device
Payment gateway Accepts the merchant’s transaction, formats it for an acquirer No Servers, usually cloud
Acquirer Holds the merchant agreement, is the scheme member, bears settlement risk Yes, principal member A licence and a balance sheet
Acquiring processor Runs the switch that speaks to the scheme No, acts for the acquirer Data centres, switch, HSM farm
Scheme network Routes, applies scheme rules, calculates interchange, clears and settles It is the scheme Global data centres
Issuer processor Authorises on the issuer’s behalf: cryptogram, PIN, risk, limits No, acts for the issuer Data centres, HSM farm
Issuer Holds the cardholder account, owns the credit decision Yes, principal member A licence and a balance sheet
Core banking The ledger of record for the customer’s account No Mainframe or modern equivalent

Two of these are legal facts and the rest are engineering facts. Only the acquirer and the issuer are scheme members; only they can be fined, sanctioned or expelled, and only they can hold settlement positions. Everyone else in the chain, however large, is somebody’s agent. This matters when things break, because the party a merchant shouts at is usually the processor and the party that owes the merchant money is always the acquirer. In Britain the acquiring names are Worldpay, Barclaycard, Global Payments, Elavon, Adyen and Stripe; the issuing-processor names include Fiserv, FIS, Thredd, Marqeta and Enfuce; the core banking names include Temenos, Finastra, Oracle, TCS BaNCS, Infosys Finacle, Thought Machine and Mambu. Ownership of these brands changes every few years. Their positions in the chain do not.

Inside an acquiring processor#

An acquiring processor’s estate is more than a switch, and a due diligence questionnaire will ask about each part of it by name.

A terminal management system, which knows every deployed device by serial number, pushes configuration and application updates, and injects or rotates terminal keys. An estate of tens of thousands of terminals is an estate of tens of thousands of key holders.

The switch itself: the process that receives an inbound transaction, validates it, applies merchant-level rules, routes it to the right scheme or domestic network, converts between message dialects, and applies the acquirer’s own timers. Chapter 24 established that ISO 8583 is a family rather than a format, and the switch is where that family becomes a problem: the acquirer’s internal dialect, Visa’s, Mastercard’s and a domestic scheme’s are all different, and the switch owns the translation table.

A store-and-forward facility for transactions the terminal completed offline, which must be delivered later without being duplicated. A capture and clearing subsystem, which accumulates approved authorisations into the day’s batch and generates the clearing files described in Chapter 30 — historically Visa BASE II and Mastercard IPM formats. An HSM farm, discussed below.

And finally the endpoint: the connection to each scheme. This is not a single wire. A processor typically holds multiple endpoints per scheme across multiple sites, each with its own institution identifier, its own key set and its own independent sequence of system trace audit numbers. DE11 is only unique within an endpoint within a day, which is precisely why Chapter 28 had to introduce four different identifiers rather than one.

The scheme switch#

Visa’s authorisation system began as BASE I, built by Bank of America and introduced in 1973 as the first real-time electronic card authorisation system; the batch clearing and settlement counterpart was BASE II. Both were later subsumed into V.I.P., the VisaNet Integrated Payment system. Mastercard’s equivalents were INAS for authorisation and INET for clearing, carried over Banknet, and the modern architecture retains the split between a dual-message system, where authorisation and clearing are separate messages, and a single-message system, where one message does both. That split determines whether a refund can be a reversal or must be a fresh credit, and it is why debit and credit behave differently on the same terminal.

The routing decision itself is unglamorous. The scheme maintains an account range table mapping leading digits of the PAN to an issuer endpoint, the switch looks up the range, and the message goes there. Everything interesting is bolted around that lookup: inline fraud scoring, tokenisation lookups, currency conversion, and stand-in processing when the issuer does not answer.

The scale is worth stating with sources, because it is routinely exaggerated. Visa’s stated figures as of October 2025 are 322 billion transactions a year, more than $16 trillion of payments volume, more than 150 million merchant locations, 4.8 billion credentials, seven independent data centres, availability of 99.9999 per cent, and a network processing up to 83,000 transaction messages per second worldwide. Its network operations centre stress-test capacity exceeds 65,000 messages per second, the figure Visa’s public fact sheets have carried since 2017.

Divide 322 billion by the number of seconds in a year and the average is about 10,200 transactions per second. Set that against a peak of 83,000 messages per second and you have the number that actually governs the engineering: a peak-to-mean ratio of roughly eight to one. Card networks are not sized for the average. They are sized for the Friday before Christmas, and they spend most of the year at an eighth of capacity, which is why redundancy in this industry is cheaper than it is elsewhere.

Visa’s Ashburn, Virginia site, which the company describes publicly, occupies 44 acres, is entered by only about 75 cleared employees, has 18-inch concrete walls rated to 170 mph winds, holds 500 petabytes of stored data, and is backed by 28 megawatts of standby power with 100,000 gallons of fuel for five days of autonomy. Those numbers are the answer to “what does a payment network physically look like”.

The issuer side, and why authorisation does not touch the ledger#

The issuer processor performs, in order: message validation; cryptogram verification, meaning it recomputes the ARQC the chip produced in DE55 and compares; cardholder verification, meaning online PIN verification if a PIN block is present; risk and fraud scoring; and an availability check.

That last step is where most people’s mental model is wrong. The authorisation does not debit the customer’s account. It writes to an authorisation record — variously the shadow balance, the open-to-buy or the memo post — that reduces available funds without altering the ledger balance. The ledger moves later, when the clearing record arrives, which is the mechanism behind every “pending transaction” a customer has ever queried and behind the whole of Chapter 30.

Behind the issuer processor sits core banking, which in a large bank is very often exactly what its reputation suggests: COBOL on IBM z/OS, batch-oriented, with a nightly cycle during which the ledger is closed for updating. The newer platforms are built around continuous processing precisely to remove that window. Whether the window exists is the single biggest determinant of what a bank can promise about timing, and it is a property of the ledger, not of the card system in front of it.

Why key operations happen inside tamper-responsive hardware#

The threat being defended against is specific. ISO 9564 requires that a PIN never exist in cleartext outside a secure cryptographic device, because a PIN is four digits: an attacker who obtains encrypted PIN blocks and any oracle that will decrypt them can brute-force the space trivially. Equally, an issuer’s card master keys can generate valid cryptograms for every card in a portfolio. Neither secret can be permitted to exist in the addressable memory of a general-purpose computer, however well patched, because a general-purpose computer’s whole design goal is that its memory is addressable.

Three terms are used loosely and mean different things. Tamper-evident means an attack leaves visible traces. Tamper-resistant means an attack is made difficult by physical construction. Tamper-responsive means the device actively detects an attack in progress and reacts, and in a payment HSM the reaction is zeroisation: erasure of the local master key, after which every key the institution holds becomes undecryptable ciphertext. The sensors typically cover case opening, drilling, temperature excursions in both directions, voltage manipulation and radiation.

The market is small: Thales’s payShield line and Utimaco’s Atalla and CryptoSec families cover most of it, alongside a handful of others. Take the current Thales device as a concrete specimen, since generalities are useless here. The payShield 10K, in its PS10-S, PS10-D and PS10-F variants, appears on the PCI Security Standards Council’s approved devices list under approval number 4-40266, evaluated against version 3.x of the PCI HSM Security Requirements, with an approval expiry of 30 April 2028. Its listed approved usage is Restricted, meaning the approval holds only when the device is deployed in an environment meeting at least the security requirements of a Controlled Environment as defined in the PCI key management requirements and the device’s own security policy. It supports ISO Format 4 (AES) PIN blocks and is listed as supporting Remote Administration.

Separately, and this is the distinction practitioners get wrong, its FIPS validation is not of the whole appliance. NIST certificate #3610 covers the Thales Advanced Security Platform, a multi-chip embedded module inside the payShield 10000 family, validated to FIPS 140-2 overall Level 3 by UL Verification Services in January 2020 and now moved to the historical list as part of the general sunsetting of FIPS 140-2. If a supplier questionnaire asks you whether your HSM is “FIPS 140-2 Level 3 certified”, the honest answer names the module and the certificate, not the box.

Published specifications for the same device: Triple DES at 112 and 168 bits, AES at 128, 192 and 256, RSA up to 4096 bits, and the FIPS 186-3 curves P-256, P-384 and P-521; host connectivity over TCP/IP and UDP on dual 1 Gbps or 10 Gbps Ethernet ports, with an optional single FICON port for mainframe attachment; and conformance claims against ISO 9564, 10118, 11568, 13491 and 16609, ANSI X3.92, X9.8, X9.9 and X9.17, ASC X9 TR-31 and X9.143, TR-34 and X9.139, and APACS 40 and 70. The presence of FICON on a 2020s product tells you what these devices are still plugged into.

The key hierarchy, named#

Key Full name Where it lives What it protects
LMK Local master key Inside the HSM tamper boundary only Every other key the institution holds
ZMK / KEK Zone master key, key encrypting key Exchanged between two institutions The transport of working keys
ZPK Zone PIN key Under LMK at rest, under ZMK in transit PIN blocks between two parties
BDK Base derivation key Under LMK; never leaves the HSM Derives per-terminal, per-transaction keys
PVK PIN verification key Under LMK PIN verification values
CVK Card verification key Under LMK CVV, CVC, iCVV, dCVV values
IMK Issuer master key Under LMK Derives per-card keys for EMV
UDK Unique DEA key Derived per card Generates and verifies the ARQC

Only the LMK is physically inside the device. Everything else is a database column holding ciphertext, meaningful only when handed back through the HSM’s slot. That is what makes the architecture affordable, and what makes zeroisation catastrophic.

DUKPT, Derived Unique Key Per Transaction, standardised in ANSI X9.24, solves a specific problem: a terminal on a counter is physically accessible to the attacker. So the terminal is never given a long-lived key. It is injected with an initial key derived from the acquirer’s base derivation key and its own device identifier, and derives a fresh key for every transaction, discarding the previous one irreversibly. Compromising a terminal therefore yields the transactions after the compromise and none before it, and nothing about any other terminal.

Key blocks, and a compliance date you must know#

For decades, keys were shared between institutions as bare encrypted values with the key’s purpose recorded separately, in documentation, by convention. This is an obvious flaw: an attacker who can persuade a system to use a PIN key as if it were a data key can extract secrets the design never intended to expose. Key blocks fix this by cryptographically binding a key’s permitted usage to the key itself, so that the binding cannot be altered without invalidating the block. The relevant standards are ASC X9 TR-31 and its successor ANSI X9.143.

PCI PIN Security Requirement 18-3 mandates key blocks in three phases, and the dates are exact. Phase 1, internal connections and key storage within service provider environments including all applications and databases connected to an HSM, took effect on 1 June 2019. Phase 2, external connections to associations and networks, took effect on 1 January 2023. Phase 3, extending key blocks to all merchant hosts, point-of-sale devices and ATMs, took effect on 1 January 2025. Anyone still transporting bare encrypted keys to a terminal estate as of this writing is out of compliance and has been for over a year and a half.

PIN translation, hop by hop#

This is the single most instructive walk in the chapter, because it shows the tamper boundary doing real work.

A customer enters a PIN on a PIN entry device. The device constructs a PIN block — a fixed-length structure combining the PIN with part of the PAN so that identical PINs on different cards do not produce identical blocks — and encrypts it under a key derived by DUKPT for that transaction only. The clear PIN exists in the PED’s secure processor for a few milliseconds and is then gone.

The encrypted block travels to the acquiring processor. Its HSM is given the block, the key serial number and the target key, and performs a translation: internally it derives the terminal’s transaction key, decrypts the block, re-encrypts it under the zone PIN key shared with the scheme, and outputs the new block. The clear PIN existed for microseconds, inside the tamper boundary, and never entered the memory of the switch. At the scheme the same operation happens again: decrypt under the acquirer’s ZPK, re-encrypt under the issuer’s ZPK. At the issuer processor the block is finally verified rather than translated, decrypted inside the HSM and checked against a PIN verification value computed with the PVK, or against an offset, with only a yes or no coming back out.

The PIN therefore crosses three organisations and exists in the clear four times, each time for microseconds, each time inside a tamper-responsive device that has no command capable of revealing it.

The block formats are specified in ISO 9564-1. Formats 0 to 3 are built for Triple DES’s 64-bit block. Format 4, standardised in 2015, is built for AES’s 128-bit block and, unlike its predecessors, includes randomness so that the same PIN, PAN and key produce a different block every time. PCI’s device requirements now force the migration: POI devices at version 5 and above supporting online PIN, and HSMs at version 4 and above supporting PIN processing, are required to support ISO Format 4.

The latency budget, restated with owners#

The plain version’s table was honest but anonymous. Here it is again with the two questions a performance engineer asks: who owns this slice, and is it physics or software?

Segment Typical Owner Nature
Card or phone in the contactless field up to ~500 ms EMVCo kernel, card, reader Protocol and silicon
Terminal to acquirer over broadband or 4G 20-40 ms Merchant’s network, acquirer edge Mostly software and access network
Acquirer switch processing 5-20 ms Acquiring processor Software
Acquirer to scheme endpoint 2-10 ms Private circuit or Edge connection Physics plus equipment
Scheme routing and inline scoring 10-40 ms Scheme Software
Scheme to issuer processor 2-80 ms Depends entirely on geography Physics
Issuer decision, including HSM calls 50-250 ms Issuer processor and core banking Software and database
Return path mirrors outbound as above as above
Terminal display and receipt 30-80 ms Terminal Software and printer

The contactless figure is the one hard number in the table: the EMV contactless specifications set the requirement that a card need not remain in the reader’s field for more than about 500 milliseconds. Everything else is a range observed in production, not a published limit, and it moves with load.

Two conclusions follow. In a domestic transaction the network is a rounding error and the issuer’s database is the story, so performance work belongs in the issuer processor rather than the wires. In a cross-border transaction the network is not a rounding error at all, and no amount of software engineering will help, because the constraint is geometry.

What physics charges you#

Light in a vacuum travels 299,792,458 metres per second. In single-mode fibre, with a refractive index of about 1.468, it travels roughly 204,200 kilometres per second. The rule of thumb used by optical engineers is 4.9 microseconds per kilometre, and it is accurate enough for any budget you will ever build.

Route Approximate fibre path Round-trip floor Typically observed
Leeds to London ~330 km ~3.2 ms 8-15 ms
London to Frankfurt ~750 km ~7.4 ms 12-20 ms
London to Ashburn, Virginia ~6,500 km ~64 ms 70-80 ms
London to Singapore ~11,000 km ~108 ms 160-200 ms

The gap between the floor and the observed figure is equipment: optical amplifiers, regenerators, routers, firewalls and the serialisation delay of putting bits onto a wire. It is bounded, and it is the part you can buy your way out of. The floor is not.

There is a second physics-adjacent cost engineers forget. A TLS 1.2 full handshake costs two round trips before any application data flows; TLS 1.3 reduces that to one. On a London-to-Virginia link that is 150 milliseconds of handshake for TLS 1.2 and 75 for TLS 1.3, on top of the TCP handshake. This is why payment links are never opened per transaction. They are long-lived, pooled, persistent sessions, kept warm and monitored, and a transaction that has to open a fresh connection has already lost more time than the entire issuer decision would have taken.

The heartbeat, and what it is for#

Keeping a session warm requires knowing whether it is alive, which is the purpose of the 8xx message class introduced in Chapter 24. A 0800 network management request with the appropriate function code is an echo test; the peer answers 0810. The same class carries sign-on, sign-off and key change. Endpoints exchange echoes on a fixed interval, and after a configured number of consecutive failures the link is marked down and traffic redistributed to the surviving endpoints.

That is why partial failure is so much more dangerous than total failure. A dead link fails its echoes and is removed automatically in seconds. A link that answers every echo correctly while mishandling the transactions behind it will never be removed by any automatic process, because by the only test the system applies, it is healthy.

Timers, reversals and stand-in#

When the issuer does not answer within the scheme’s response time parameters, something must happen, and what happens is defined rather than improvised.

The scheme may substitute its own decision. Both major schemes operate stand-in processing, in which the network authorises on the issuer’s behalf against limits the issuer has pre-registered. Visa’s published action code 91 covers the case where stand-in is not available: “Issuer or switch inoperative and STIP not applicable or not available for this transaction; time-out when no stand-in”. Issuer processors document a related behaviour in which Visa forces single-request stand-in and converts a 91 or 96 into an N0. Chapter 29 covered what these codes mean to a merchant; the point here is that they are emitted by network conditions rather than by any bank’s opinion of the cardholder.

When the acquirer’s own timer expires, it must assume the issuer may have approved and reserved funds it does not know about, and must therefore send a reversal — the 0400 reversal request of Mastercard’s Transaction Processing Rules and its equivalents. The reversal is what stops a network hiccup from silently holding a customer’s money for a week, and unwinding un-reversed timeouts is a large share of the reconciliation work described in Volume V.

Redundancy in practice#

Serious payment estates use one of two patterns. Active-active runs both sites live, splitting traffic, so a site loss is a capacity loss rather than an outage and the failover has been exercised continuously. Active-passive runs one site live and one warm, which is cheaper and has the fatal property that the failover path is exercised only in the emergency.

Underneath sits the replication decision. Synchronous replication does not acknowledge a write until the second site has it, giving a recovery point objective of zero at the cost of adding the inter-site round trip to every write. That is why paired sites sit tens of kilometres apart rather than hundreds: at 4.9 microseconds per kilometre, a 50 km separation costs about half a millisecond per write and a 500 km separation costs five. Asynchronous replication removes that cost and accepts a non-zero recovery point objective. In payments that window is not an abstraction; it is a set of authorisations that happened, that customers were told about, and that the surviving site has never heard of.

The genuinely hard problem is neither of these. It is the decision. A cutover is disruptive in itself, so nobody wants to trigger one on a false positive; but the detection signal in a partial failure is ambiguous by definition. Firms therefore write runbooks with explicit trigger thresholds and rehearse them. Since 31 March 2025, when the transition period for the FCA’s PS21/3 operational resilience regime ended, UK firms have been required to have identified their important business services, set impact tolerances for each, mapped the resources supporting them, and demonstrated by testing that they can stay within those tolerances in severe but plausible scenarios.

Friday 1 June 2018, narrated#

The best public account of a scheme outage in Britain is Visa Europe’s own, given to the House of Commons Treasury Committee by its chief executive Charlotte Hogg in a letter dated 15 June 2018. What follows is from that account and the contemporaneous reporting of it.

Visa Europe operated two data centres in the United Kingdom, either of which could independently handle 100 per cent of Visa’s European transactions. In normal operation the systems were synchronised and either centre could take over from the other immediately.

At 14:35 on Friday 1 June 2018, a component within a switch in the primary data centre suffered what Hogg described as a very rare, partial failure. Because the failure was partial, the backup switch did not activate.

This is the exact scenario the seam warned about. The malfunctioning system at the primary site did not stop. It continued attempting to synchronise messages with the secondary site, building a backlog there which in turn degraded the secondary’s own ability to process incoming transactions. The redundant capacity was not merely unused; it was being poisoned by the failing primary. Isolating the primary took nearly five hours, and full correct processing was not restored until 00:45 the following morning: ten hours and ten minutes end to end.

The numbers Visa gave Parliament are worth recording precisely. Across Europe, 51.2 million transactions were attempted during the affected period; about 10 per cent were affected, and 5.2 million failed to process. In the United Kingdom 2.4 million failed, affecting 1.7 million credit and debit cards; a further 2.8 million failed elsewhere in Europe. Disruption was not constant. There were two peaks — roughly ten minutes just after 15:00, and fifty minutes between 17:40 and 18:30 — during which 35 per cent of transactions failed. Outside those peaks the UK failure rate was closer to 7 per cent.

That last detail is the one merchants remember, because it describes what an outage actually feels like on a shop floor. Not “the card machines are down”. Instead: one payment in fourteen fails for no visible reason, then briefly one in three, then one in fourteen again, with no announcement, no error message worth reading, and a customer at the front of the queue re-presenting the same card three times until it works. Staff cannot tell a scheme outage from a bad card, because on the counter they look identical. Visa engaged EY to conduct an independent review and said it would migrate European processing onto its global VisaNet platform, which it described as more resilient in detecting and recovering from partial malfunctions of exactly that type.

Three lessons generalise. Partial failure is worse than total failure. A hot standby coupled to a sick primary is a liability rather than an asset, because the coupling is the failure path. And the recovery clock is dominated not by repair but by diagnosis and by the decision to isolate.

From leased lines to TLS over the public internet#

The connectivity story runs in four eras and the last one is still arriving.

Dedicated circuits, 1973 to the mid-1990s. BASE I ran over leased lines between member banks and Visa’s centres. Terminals dialled: a countertop unit picked up a phone line, dialled a modem pool and took fifteen to thirty seconds, which is why floor limits and offline authorisation, covered in Volume II, were not an optimisation but a necessity. Behind the scenes ran X.25, and in bank environments IBM’s SNA and SDLC.

Private packet networks, the 1990s and 2000s. Frame relay and then MPLS replaced point-to-point circuits with managed private clouds, and the scheme endpoint became a physical appliance installed in the customer’s data centre — a Visa access point, or Mastercard’s Interface Processor, the MIP. The scheme shipped you a box, and the box was the border.

IP with cryptographic protection, the 2010s. Terminals moved to ADSL, then Ethernet, then 4G, and the transport became the ordinary internet with TLS carrying the security a private circuit had previously carried by being private. PCI DSS Requirement 4 governs this directly: strong cryptography must safeguard the primary account number during transmission over open, public networks, with TLS 1.2 or higher the practical floor. Version 4 added Requirement 4.2.1, that certificates used to protect PAN in transit are confirmed valid and not expired or revoked, and 4.2.1.1, that an inventory of those trusted keys and certificates is maintained. The second catches firms out, because certificates expire on a schedule nobody is monitoring until the morning they do.

Cloud and colocation endpoints, now. The physical box in the customer’s data centre is being retired. Mastercard publishes three connectivity flavours side by side: Mastercard Edge for customers in their own on-premises data centres, Co-Lo Edge for those in third-party colocation facilities, and Cloud Edge for those in the public cloud, delivered in collaboration with cloud providers including AWS. A processor that runs no data centre of its own can now hold a scheme endpoint, which was not true a decade ago. Merchant-side, the same arc ends in softPOS: a commodity phone acting as the contactless reader, with no dedicated hardware between the card and the internet at all.

Concentration, and the critical third parties regime#

Moving payment infrastructure to a small number of cloud providers improves the resilience of any single firm and concentrates the risk of all of them.

On 20 October 2025 that argument stopped being theoretical. Beginning at approximately 06:48 UTC, a latent race condition in DynamoDB’s automated DNS management in Amazon Web Services’ us-east-1 region produced an empty DNS record for the regional DynamoDB endpoint, which the automation could not repair. DNS resolution was restored by 09:40 UTC, but the cascade through dependent services — around 140 AWS services, including EC2 — meant full recovery took over fifteen hours. Firms whose control planes lived in that region discovered that their own multi-availability-zone architecture had not made them independent of it.

Regulators had already moved. The Critical Third Parties regime was created by the Financial Services and Markets Act 2023, and its rules were finalised jointly by the Bank of England, the PRA and the FCA on 12 November 2024, published as PS16/24 by the Bank and PRA and PS24/16 by the FCA. Designation is HM Treasury’s decision, not the regulators’. On 10 July 2026 the Treasury made the first designations, effective 13 July 2026: Microsoft Ireland Operations Limited, Google Cloud EMEA Limited, Amazon Web Services EMEA SARL and Oracle Corporation UK Limited. Oversight applies only to the systemic services those firms provide to the financial sector, and the regime does not transfer responsibility: firms remain accountable for managing the risks in their own third-party arrangements. In the European Union the parallel instrument is the Digital Operational Resilience Act, which has applied since 17 January 2025.

Design rules that follow#

Count trust boundaries, not companies. Every boundary has a timer, a key relationship and a different party to telephone at three in the morning, and your runbook needs all three written down before you need them.

Budget latency by segment and label each segment physics or software. You can optimise software; you cannot optimise distance, so decide where the issuer processor physically sits before you write any code. And never open a TLS session per transaction: pool, persist, keep warm and echo.

Treat a timeout as a financial event. Every one must produce a reversal or a reconciliation exception, and an exception nobody works is a customer complaint with a delay fuse.

Assume partial failure. Health checks must exercise the actual work rather than answer a ping, because a component that passes its heartbeat while corrupting its output will never be evicted by any automatic mechanism. Then rehearse the cutover, and rehearse the decision to cut over, separately. The technical failover is the easy half.

Keep the key inventory and the certificate inventory as first-class operational assets, and when you make a compliance claim about an HSM, name the validated module and its certificate rather than the appliance. On a due diligence questionnaire, precision reads as competence.

33.98 Common wrong ideas#

Wrong: the five desks are five companies in five buildings. Right: they are five roles, and Adyen or Stripe may hold three of them in one stack while a single apparent hop conceals a gateway, a token vault, a risk engine, a switch, a scheme endpoint and a repository write.

Wrong: the card network is what makes a card machine slow. Right: the scheme accounts for about forty of the eight hundred milliseconds; the phone-to-reader exchange and the issuer’s decision are where the time actually goes.

Wrong: the schemes publish a two-second rule for authorisation. Right: no such specification exists, only a nest of separate timers with separate owners — terminal, acquirer, scheme, issuer and the customer — expiring at different moments with different consequences.

Wrong: a timeout produces a slow approval. Right: it produces a different transaction: a reversal, a stand-in decision made by somebody who is not your bank, or a decline.

Wrong: an HSM is a safe that hides your keys. Right: almost every key lives as ciphertext in an ordinary database and only the local master key sits inside the tamper boundary; the security comes from a deliberately impoverished instruction set with no command capable of revealing a PIN.

Wrong: zeroisation protects your keys. Right: it makes every copy of them, everywhere, permanently unreadable, which is worse than losing the box.

Wrong: having two data centres means you have failover. Right: capacity is the easy half; detecting that a site is unwell, distinguishing unwell from briefly busy, and committing to a cutover that is itself disruptive is the part that fails.

Wrong: six-nines availability means outages are negligible. Right: 99.9999 per cent is 31.6 seconds a year, and the Visa Europe outage consumed roughly one thousand one hundred and fifty years of that budget in a single afternoon.

Wrong: authorisation debits the customer’s account. Right: it writes an authorisation record — shadow balance, open-to-buy or memo post — that reduces available funds, and the ledger moves later when the clearing record arrives.

Wrong: “our HSM is FIPS 140-2 Level 3 certified” describes the appliance. Right: NIST certificate #3610 covers the Thales Advanced Security Platform, a module inside the payShield 10000 family, and an honest answer names the module and the certificate rather than the box.

33.99 Chapter summary in 20 lines#

  1. This book has described a payment as messages, but messages do not travel; optical and electrical signals travel, through glass and switches into racks in buildings with concrete walls and diesel tanks.
  2. An authorisation passes through five jobs — the terminal, the merchant’s payments company, the scheme network, the issuer’s card system and the bank’s ledger — and the answer returns down the same chain.
  3. Timed end to end, a typical contactless approval takes about 790 milliseconds, comfortably inside the two seconds the industry talks about.
  4. Almost two thirds of that is the phone and the reader talking before anything leaves the shop, the scheme is about five per cent, and all six network legs together are seventy milliseconds.
  5. The remaining margin exists for busy issuers, retransmitted packets and geography, and geography is the only one money cannot fix.
  6. Light in single-mode fibre travels about 204,200 kilometres per second, roughly 4.9 microseconds per kilometre, so London to Ashburn and back has a floor near sixty-four milliseconds however much you spend.
  7. Because a full TLS handshake costs one or two round trips on top of that, payment links are long-lived, pooled and kept warm rather than opened per transaction.
  8. The five desks are roles rather than companies, so the honest unit of analysis is the trust boundary, and every boundary carries a timer, a key relationship and a different party to telephone at three in the morning.
  9. Only the acquirer and the issuer are scheme members, which is why the party a merchant shouts at is usually the processor while the party that owes the merchant money is always the acquirer.
  10. The scheme switch itself is unglamorous — an account range table maps the leading digits of the card number to an issuer endpoint — and fraud scoring, tokenisation, currency conversion and stand-in are bolted around that lookup.
  11. Visa’s published 322 billion transactions a year against a peak of 83,000 messages a second implies a peak-to-mean ratio of about eight to one, which is why these networks are sized for the Friday before Christmas and why redundancy is comparatively cheap.
  12. On the issuer side, authorisation verifies the chip’s cryptogram, the PIN and the risk position and then writes a shadow balance rather than touching the ledger, which is what every pending transaction really is.
  13. A PIN is four digits and an issuer master key can forge a whole portfolio, so neither may exist in the addressable memory of a general-purpose computer, and key operations happen inside tamper-responsive hardware that erases itself when attacked.
  14. Only the local master key lives inside that hardware; everything else is a database column of ciphertext, which is what makes the architecture affordable and zeroisation catastrophic.
  15. DUKPT gives each terminal a key that changes every transaction, so compromising a counter-top device yields nothing before the compromise and nothing about any other device.
  16. Key blocks bind a key’s permitted usage to the key itself, and PCI PIN Security Requirement 18-3 has extended that obligation to merchant hosts, point-of-sale devices and ATMs since 1 January 2025.
  17. A PIN block is translated rather than revealed at each hop, existing in the clear for microseconds inside three organisations’ tamper boundaries and never entering the memory of any switch.
  18. Redundancy fails at the decision rather than the capacity: an 8xx echo evicts a dead link in seconds, but a link that answers every heartbeat correctly while mishandling the traffic behind it will never be evicted by any automatic mechanism.
  19. That is exactly what happened to Visa Europe on 1 June 2018, when a partial switch failure poisoned the healthy site, took nearly five hours to isolate, and left 5.2 million European transactions unprocessed with UK failure rates oscillating between 7 and 35 per cent.
  20. The rules that follow are to count trust boundaries rather than companies, to label every latency segment physics or software, to treat every timeout as a financial event that must produce a reversal, to health-check the real work rather than a ping, and to rehearse the decision to cut over separately from the cutover itself.

Sources: Visa corporate publications on VisaNet and the Ashburn data centre (October 2025) and Visa Europe fact sheet; Visa Europe’s letter of 15 June 2018 to the Treasury Committee, with reporting by Finextra and The Guardian; Visa Developer response code documentation; Mastercard Edge Connectivity and Transaction Processing Rules; PCI SSC bulletin on PIN Security Requirement 18-3 and the PTS approved devices listing for approval 4-40266; NIST CMVP certificate #3610; Thales payShield 10K and Utimaco payment HSM specifications; EMVCo contactless specifications; FCA PS21/3 and the joint Bank of England, PRA and FCA policy statements PS16/24 and PS24/16; HM Treasury’s designation announcement of 10 July 2026; AWS post-event summary and ThousandEyes analysis of the us-east-1 incident of 20 October 2025.