Skip to content
KEDBYTE
How Identity Works
Chapter
5

Names Are Not Identifiers

Part I · What Identity Is|11,995 words|about 52 min read|Volume 1

5.0 What this chapter gives you#

  1. You will be able to take any form that asks for a name and say, for each box on it, which group of human beings that box excludes.
  2. You will be able to explain why “first name” is the wrong label even in English, and give the replacement that works everywhere without translation.
  3. You will be able to write a real person’s name into the strip of capital letters at the bottom of a passport by the actual international rule, including what happens when it is too long.
  4. You will be able to say what it means for two names that look identical on screen to be different data, and show the exact numbers behind the letters.
  5. You will be able to choose between the four ways a computer can tidy a name before storing it, and say which two silently destroy information.
  6. You will be able to explain why turning a name into capital letters is lossy, and name the language in which it corrupts an ordinary word.
  7. You will be able to design a record that survives a marriage, a divorce, a gender recognition certificate and a deed poll without losing history and without leaking it.
  8. You will be able to pick a length for a name column and defend the number with real cases where a smaller one caused real harm.
  9. You will be able to specify the storage shape argued for here, one required full name plus optional structured parts, in a real schema.

There is a moment in the life of every identity system when somebody walks up to it and does not fit, and usually a name does it. The name is too long, or it is one word, or two words that are both surnames, or it contains a letter the keyboard cannot make, or it is spelled one way on a birth certificate and another on a passport because two clerks in two decades guessed differently about the Latin alphabet. The system then does one of three things: it refuses the person, it quietly mangles the name and carries on, or it stores the name faithfully and fails later when it compares that name with a copy held elsewhere. Only the first failure is visible.

This chapter is a catalogue of every assumption you are carrying about human names, and a demonstration that each one is false somewhere. That sounds like trivia. It is not. Each false assumption maps to a group of people your system cannot serve, and the groups are not small. Roughly one in five South Koreans is surnamed Kim. Tens of millions of Indonesians have historically used a single name. Every Icelander’s surname changes with a parent’s first name rather than being inherited. Every Spanish citizen carries two surnames. A form with two boxes labelled first and last has already excluded several hundred million people from being recorded correctly, before a line of validation code is written.

We will build up from a shopkeeper’s ledger, take it apart where the simple picture stops being honest, and then do it properly with real specification sections, real code points, real statutes and real people whose names broke real systems. Three neighbours to mark and leave. What a person record is, and how a database of people should be shaped, is chapter 4. What a passport physically is, and what its chip and check digits do, is chapter 6; here we take only the rules governing how a name is written into it. Deciding whether two differently spelled names belong to one human being is chapter 8, and this chapter stops one step short of it.

The plain version#

A shopkeeper with a ledger and two columns#

Picture a small shop in a market town in 1955. The owner keeps a ledger for customers who buy on credit, ruled into columns, the first two headed “Christian name” and “Surname”. Under them she writes the people she knows: John Marsh, Mary Whitcombe, Peter Gray. The ledger works perfectly for thirty years, because everybody who walks in has exactly the shape of name the columns expect. Then the town changes, as towns do, and over one summer eight new customers open accounts, every one of whom breaks the ruling of the page.

The first says her name is Björk Guðmundsdóttir. The shopkeeper writes Guðmundsdóttir in the surname column, and it is wrong, because that is not a family name at all. It means “Guðmundur’s daughter”. Her brother’s surname is Guðmundsson, and her father’s is made from his own father’s first name. Nobody in that family shares a surname, and all of them would be baffled to be asked for one. The filing puts her under G, and when she returns the shopkeeper cannot find her, because in Iceland you look people up under their first name.

The second says his name is Suharto. Just that. One word. The shopkeeper hovers over the two columns, writes Suharto in the surname column and leaves the other blank, because the filing runs on surnames. The following week an assistant tidying the ledger finds a blank Christian name, assumes an error, and writes in “Mr”. There is now a customer called Mr Suharto in the books, and nobody in the town by that name.

The third is María José Cañón Iglesias. Two given names the shopkeeper can just about handle; two surnames she cannot. Cañón came from her father, Iglesias from her mother, and both are permanently hers. She is not Mrs Iglesias and not Mrs Cañón. The shopkeeper writes “Cañón Iglesias” in one column, worries, crosses it out and writes “Iglesias”, which is the one part of her name the shop will never use to address her.

The fourth is Mohammed bin Rashid bin Saeed. The middle word means “son of”. His name is a chain: himself, his father, his grandfather. There is no family name in it anywhere, and asked for a surname the honest answer is that the question does not apply.

The fifth introduces herself as Wu Chen Doris, or Doris Wu Chen, depending on the language she is speaking, because in her own language the family name comes first and in English she has learned to flip it. The shopkeeper writes down the first word she hears under Christian name. It is the family name. From that day the ledger has her backwards, and so does every letter the shop sends.

The sixth is Robert Marsh Junior, whose father Robert Marsh has had an account here for twenty years. There is no column for the Junior, so it is dropped. Two men now share one line. The debts of one become the debts of the other, and it takes four years and a solicitor to undo.

The seventh is a woman whose surname is Null. That word means nothing in the shop and something destructive inside a computer, and we will come back to her.

The eighth is Janice Keihanaikukauakahiheʻekahaunaele, whose surname does not fit the column. The shopkeeper writes as much as she can and runs off the edge of the page.

The rule that the ledger was really enforcing#

Look at what the two columns were actually doing. They were not recording a name. They were recording a claim about the structure of a name: that every human being has exactly two meaningful parts, that one is shared with their relatives, that the shared one comes second, and that both fit a fixed amount of space. All four claims are false, each in a different part of the world.

That is the whole of this chapter in one idea. A form is a law. When you draw two boxes and label them you are legislating, and the people who do not fit your law will be turned away, renamed, or recorded wrongly. The shopkeeper intended none of it. She was ruling a page.

The one-box ledger#

Suppose the shopkeeper starts again with a new book, and this time she rules only one wide column, headed simply “Name”. She writes exactly what each customer tells her, letter for letter, in their own spelling, however long it is.

Almost everything gets better at once. Björk Guðmundsdóttir is correct. Suharto is correct, and no assistant is tempted to fill in a blank. María José Cañón Iglesias is correct and complete. The chain of sons is correct. Nothing is flipped, dropped or truncated.

Two things get worse, and a book that told you otherwise would be lying. The first is filing: with one string of letters she does not know which part to file under. Filing by the first word is right for Iceland and wrong for Spain; filing by the last is right for England and wrong for China. So she rules a small extra column headed “File under” and asks each customer what to write in it. Björk says Björk. María José says Cañón. Doris says Wu. The customer always knew; the shopkeeper never did. The second is address, because knowing a full name does not tell you what to call somebody at the counter, so she rules a second small column headed “Call them” and asks that too.

That is the design this chapter argues for, and it is worth saying plainly before any technical detail, because the detail only supports it. Store one required full name, exactly as the person writes it, in a generous amount of space. Then, only if you genuinely need them, add optional extra parts and ask the person to fill them in rather than guessing. Everything else here is evidence for that sentence or a warning about ignoring it.

One name, carried all the way through#

We will follow one person through the whole chapter, so every idea lands on the same real value. Her name is María José Cañón Iglesias. She was born in Madrid, holds a Spanish passport, and works for a company with offices in Istanbul and Reykjavík.

Her name has four words. María and José are her given names; in Spain a woman is quite ordinarily called José as a second given name, which is the first of your assumptions this example breaks. Cañón and Iglesias are her surnames, her father’s first surname and her mother’s first surname in that order, both permanently hers. She has no middle name in the English sense and no maiden name, because Spanish women do not change their names on marriage. Asked in Britain for her “first name and last name”, the honest answer is that she has two of each and neither pair works the way the question assumes.

Now watch those four words travel.

On her national identity card in Spain it is the whole thing: María José Cañón Iglesias.

In the strip of capital letters at the bottom of her passport, the strip a machine reads, it becomes something that does not look like a name at all. That strip may hold only the twenty-six capital letters A to Z, the ten digits, and one filler symbol used as a kind of blank: no accents, no tildes, no lower case, no punctuation, no spaces. So the tilde and the accents are thrown away, the surnames go first, and every gap is packed with filler. The result is a block of thirty-nine characters reading CANON, a filler, IGLESIAS, two fillers, MARIA, a filler, JOSE, and then filler to the end.

Notice what has happened. Her passport now carries two spellings of her name: the correct one printed above, and a stripped one below that a machine will copy into an airline system, a border database and a hotel register. From that moment there are two of her.

Written into a computer’s memory, something stranger happens. The letter n with a tilde over it, the ñ in Cañón, can be stored in two completely different ways, and both are correct. The computer can hold it as a single item meaning “the letter n-with-a-tilde”, or as two items meaning “the letter n” followed by “put a tilde on the thing before”. On the screen they are identical. Character for character, they are not equal. If one system stores her name one way and another stores it the other way, and later they compare the two, the comparison says the names are different, and María José Cañón Iglesias is told she does not exist.

Written into a system that decides to be tidy and put everything in capitals, her name becomes MARÍA JOSÉ CAÑÓN IGLESIAS, and that system cannot put it back, because the tidying threw away which letters were capitals to begin with. In Turkey, where her company has an office, the same tidying turns one letter of an ordinary name into a different letter, and in 2008 that exact substitution, in a text message rather than a database, is reported to have ended in a killing.

The general shape of the problem#

A name is not an identifier. An identifier is something a system assigns to point at exactly one thing, and it has properties the system controls: unique, unchanging, fixed in shape, meaningless on its own. A name has none of them. It is not unique, it changes, it has no fixed shape, and it is loaded with meaning about family, religion, region, gender, marital status and history.

Treating a name as an identifier is the commonest error in the field. It shows up as a name used as a database key, a name used to match two records, a name used to decide whether a document belongs to its bearer, and a stripped machine-readable name treated afterwards as a true spelling. The identifier is what the system uses to keep track; the name is what the person uses to be a person. You owe them an accurate copy of the second and must never run the system on it.

Where the plain version stops being true#

One box does not mean no rules#

The honest version: “just store what they typed” is a good instruction for the entry field and an incomplete one for the system, because a real system also prints a name on a card, sorts a list, and checks it against a document. What you store is exactly what they typed; what you compare, sort and print is derived from it by rules you choose deliberately, write down and apply identically everywhere. The difference between a system that works and one that does not is almost never the storage. It is whether those derived forms exist and are computed the same way in every component.

The person is not always the best authority on the spelling#

We told the shopkeeper to ask the customer what to file under, and that is right. But the plain version quietly assumed there is always one answer the person can give, and there is not.

A person with a name in a non-Latin script may have three or four Latin spellings in circulation, none of which they chose: one on a passport, one written the way a bank clerk heard it, one in a scholarly system with marks over the letters, one produced by a consulate’s software. All four are that person’s name, and asked which is correct they will reasonably say “all of them”, which is not an answer a column can take.

Nor is the legal position clean. In some countries the birth register entry is definitive and everything else is a copy; in others there is no official spelling at all. Some keep a list of permitted names and refuse the rest, as Iceland does through its naming committee. So “the person’s real name” is not one fact you can look up. It is a claim made by a particular authority at a particular time, and different authorities make different ones.

English-speaking readers usually carry an idea that everybody has one legal name, recorded somewhere official, and that changing it is a formal act. That is a local custom, and not even reliably true in England.

In England and Wales you may change your name simply by using a new one consistently. A deed poll evidences that change; it does not cause it. An unenrolled deed poll may be made from the age of sixteen, and from eighteen you may apply to have the change entered on the public record at the High Court, which as of August 2026 costs 53.05 pounds according to GOV.UK. Neither step is what makes the new name yours. Marriage there changes no name automatically either; a marriage certificate is simply evidence that organizations conventionally accept.

Elsewhere the opposite holds: the registered name is the only name that exists for official purposes, changing it is a court matter, and using another is an offence. Both models are normal, and a system that assumes either one is broken in half the world.

The name on the document is not the name of the person#

The plain version treated the passport strip as a slightly damaged copy of a name. Be harder than that. The stripped, capitalized form is not a copy of her name at all; it is a machine-readable token derived from it by a lossy rule so that a scanner in poor light gets the same answer twice. It has a fixed alphabet, a fixed length, and it may be truncated. When an airline copies it into a booking, a border system copies the booking into a watchlist query, and the watchlist compares it with a token derived under another country’s transliteration rules, the whole chain is comparing tokens rather than names, and it is very good at producing both false matches and false misses. Chapter 8 covers what to do about that; chapter 6 covers what the document itself is worth.

Unique is not almost-unique#

The honest version: a name is evidence whose strength depends entirely on the population you are searching. In the 2015 South Korean census 10,689,959 people had the surname Kim, 21.5 per cent of the population, with Lee at 14.7 and Park at 8.4, so nearly forty-five people in every hundred carried one of three surnames. In a village a name is nearly an identifier; in a national database it narrows the field and settles nothing. Chapter 4 argues for a real key; this chapter adds only that a name is not one.

Correct storage does not make comparison safe#

The last place the plain version breaks is the most technical, and the second half of this chapter grows out of it. Two names a human reads as identical can be different data in at least four independent ways. They can use different underlying representations of the same accented letter. They can use visually identical letters borrowed from different alphabets. They can differ in capitalization in a way no simple rule reverses. And they can differ in invisible characters that occupy no space on the screen.

None of those is exotic; each shows up in ordinary European and Asian names on ordinary web forms, and each makes an equality test return the wrong answer. The remedy is not to clean the stored name, which destroys it, but to compute a separate comparison form alongside it. That distinction, between the name you keep and the key you compare, is the most useful thing in this chapter.

The technical version#

The falsehoods, taken one at a time#

The canonical statement of the problem is Patrick McKenzie’s essay “Falsehoods Programmers Believe About Names”, published on his Kalzumeus blog in June 2010. It lists forty assumptions, each stated flatly as something a programmer believes and each false, and it endures because every entry corresponds to a class of production defect rather than to a curiosity. Reproduce the shape of the argument rather than the list, because a list is easy to nod at and hard to use: the forty group into seven families, each with a named counter-example.

Cardinality, falsehoods 1 to 5: that a person has exactly one name, or exactly N name parts for some fixed N. Refuted by the mononym, until recently the ordinary pattern for much of the Indonesian population, by Arabic chains running to five or six components, and by anybody whose professional, legal and social names differ and are all real.

Space, falsehood 6: that names fit a defined amount of room. Refuted below by the Hawaii driving licence and the passport strip. Stability, falsehoods 7 and 8: that names do not change, or change only at an enumerable set of events. Refuted by marriage, divorce, adoption, religious conversion, gender recognition, witness protection, transliteration corrections, and jurisdictions where a person may change their name at will and repeat the exercise.

Character set, falsehoods 9 to 11: that names are ASCII, or in one character set, or that every name maps to Unicode code points. Refuted by Chinese personal names using rare characters, by Japanese variant forms of family-name characters that are legally distinct, and by the fact that Unicode is still adding characters found in existing legal records. Case, falsehoods 12 and 13: that names are case sensitive, or that they are case insensitive. The pair is deliberately contradictory, because both are believed and both are wrong.

Structure, falsehoods 14 and 18 to 20: that names have a reliable order, that first and last names differ, that family names exist and are shared with relatives, and that prefixes and suffixes can be ignored. Refuted by Iceland, Spain, China and Hungary, and by the fact that a generational suffix can be the only thing separating father from son. Uniqueness, falsehoods 21 to 23: refuted by the Korean census figures.

The list ends with two entries that are the whole point. Falsehood 39: “People whose names break my system are weird outliers. They should have had solid, acceptable names, like 田中太郎.” Falsehood 40, the last: “People have names.” That is not flippant. There are living people with no name recorded anywhere, and there are records that must be created before a name is known, which is why a hospital system needs a way to register a newborn who has not been named yet.

McKenzie’s own experience shows how relative all this is. He is an American living in Japan, and told the BBC in March 2016 that his surname, at eight characters, exceeds the printed space on Japanese forms, so he files his taxes as “McKenzie P”. When his bank dropped support for the katakana rendering of his name, a paper request had to travel from his branch to corporate IT to have the database edited by hand before he could use the bank’s website. There is no such thing as a normal name; there is only a name that happens to match the assumptions of the system in front of you.

The shapes a name actually takes#

Six patterns cover most of the world, and a system that handles all six handles nearly everything.

A mononym is a single-word name with no second component, not a nickname and not an abbreviation, and handling it requires that no field be mandatory except the full name. The United States Department of State’s Foreign Affairs Manual, at 8 FAM 403.1, gives the passport rule: where an applicant has only one name, “the name must be recorded in the last name field”, with a caret in the first and middle name fields. The same manual covers documents bearing the placeholders “No First Name” or “First Name Unknown”, abbreviated NFN and FNU, and treats such a person as having one name. That abbreviation is why many people from single-name traditions travel with the given name “Fnu” printed on their visas, having acquired a first name from a data-entry convention.

A patronymic derives the second element from the father’s given name, and a matronymic from the mother’s. Iceland’s Personal Names Act, No. 45 of 17 May 1996, in force from 1 January 1997, states the rule in Article 8: surnames are of two types, patronymics or metronymics and family surnames, and the first kind is formed by placing the parent’s name in the genitive case after the person’s given name, with the suffix “son” for a man and “dóttir” for a woman. The same article forbids the adoption of new family surnames in Iceland altogether, and Article 4 permits no more than three given names. Names not already on the Personal Names Register go to the Personal Names Committee, mannanafnanefnd, whose rulings under Article 22 cannot be appealed. On 31 January 2013 a Reykjavík court overruled the committee for a girl registered only as “Stúlka”, meaning “girl”, holding that her parents’ chosen name Blær could be a woman’s name as well as a man’s. Since the Gender Autonomy Act of 2019 a given name need not match the bearer’s registered gender, and people registered as neither male nor female may use the suffix “bur” in place of “son” or “dóttir”.

Russian names carry a patronymic as a distinct middle element with its own grammatical gender: Boris Nikolayevich Yeltsin, where Nikolayevich means “son of Nikolai”, while a daughter’s patronymic is Nikolayevna. Arabic naming may chain several generations with “ibn” or “bin” and “bint”, and may add a teknonym naming a person as somebody’s father or mother. Malay names use “bin” and “binti” the same way, so Isa bin Osman is Isa, son of Osman, and Osman is not a family name and is not shared with Isa’s children.

Generational suffixes are the part everyone drops. Junior, Senior, II, III and IV are not decoration; they are often the only element distinguishing two living people who share a household, an address and a bank. Systems that push “Jr” into the surname field produce a person named “Marsh Jr” who does not match “Marsh”; systems that drop it merge father and son. The travel document standard takes the strict line: prefixes and suffixes including titles, professional and academic qualifications, honours, awards and hereditary status are excluded from the machine readable zone unless the issuing State considers them legally part of the name, in which case they become components of the secondary identifier.

Honorifics are a separate axis and should never share a field with a name. Dr, Prof, Sir, Dame, Rev, and the Malaysian and Thai title systems with their several ranks, attach to a person rather than forming part of a name and change over a lifetime, while Mr, Mrs, Miss and Ms encode gender and marital status that most services have no business collecting. The GOV.UK Design System advises against collecting a title at all for exactly that reason, and a free text input rather than a fixed list where one is genuinely needed.

Multiple family names are standard across the Spanish- and Portuguese-speaking world. A Spanish citizen has two surnames, one from each parent. Until 30 June 2017 the paternal surname came first by default; from that date parents registering a birth in Spain must decide the order themselves, and if they cannot agree the civil registry decides. Both surnames are part of the name and neither is “the” surname; filing conventionally uses the first.

The table below is the compressed version. It is not a lookup table you can code against, because no such table exists; it is a list of shapes your fields must not forbid.

Pattern Example What breaks
Mononym Suharto Required given name
Patronymic Bjork Gudmundsdottir Shared family name
Chain Isa bin Osman Surname field
Two surnames Canon Iglesias One surname field
Generational Robert Marsh Jr Suffix dropped
Family name first Wu Chen Doris Order assumed

Order, and why “first name” is the wrong label#

There are two dominant orders. Given-family, used across most of Europe, the Americas, part of South Asia and the Middle East. Family-given, used in China, Japan, Korea, Vietnam, Hungary and parts of Africa. In family-given order the family name is written first and is the first name in the literal sense, which is why “first name” is a defect rather than a simplification. The label also fails for mononyms, where there is no first and no second, and for anybody whose display order differs from their filing order, which is almost everyone who has ever written “Marsh, John” on a form.

The replacement costs nothing: label the fields “given name” and “family name” and never use position words. That is the recommendation in the W3C internationalization article “Personal names around the world”, which advises avoiding “first name” and “last name” in non-localized forms and asks the prior question of whether separate fields are needed at all.

That this is more than a designer’s preference is shown by a government changing its practice over it. On 1 January 2020 the Japanese government began writing Japanese personal names family-name-first when using the Roman alphabet on official documents, reversing decades of flipping them for Western readers, after public advocacy in 2019 by the then foreign minister, Tarō Kōno, who asked foreign media to write Abe Shinzō. If your system stores an order-dependent full-name string and assumes it was built given-family, every Japanese record created since that date may be stored backwards relative to older ones, with nothing in the data to say which.

Three resolutions follow. Name parts by role, not position. Store the person’s own preferred full form and use it for display. And store a sort key separately, because sorting is a different question: Iceland sorts by given name, Spain by the first surname, China by the family name that already comes first.

Transliteration, transcription and romanisation#

Three words are used loosely and mean different things. Transliteration is a strict, letter-for-letter mapping between scripts, designed to be reversible. Transcription is a looser mapping representing how a name sounds in the target language, and is not reversible. Romanisation is the general term for either when the target script is Latin.

Doc 9303, the travel document standard published by the International Civil Aviation Organization, is unusually candid about the confusion. In Part 3, Appendix B, it observes that although the standard requires a “transliteration” where the national script is not Latin, what is commonly supplied is a phonetic equivalent and “should be more correctly termed a transcription”. The distinction matters because a transcription is many-to-one: it discards information, and different clerks produce different answers.

The scale shows in the Korean census. Of the 10.7 million people surnamed Kim in 2015, 99.3 per cent romanize it as “Kim”; the other 0.7 per cent, some seventy-five thousand people, use Gim, Ghim or Kin. One Korean surname, four strings in your database. Arabic is worse: Doc 9303 Part 3 notes that a single classical Arabic name can generate 256 plausible Latin variants from ordinary spelling choices alone, before anybody makes a mistake.

Doc 9303 Part 3, Section 6, titled “Transliterations recommended for use by States”, gives the mapping tables States should use for the most common Latin, Cyrillic and Arabic characters. A sample of the Latin table shows the character of the compromise, and shows that even here there is no single answer.

Code point Character Recommended
00C4 A diaeresis AE or A
00C5 A ring above AA or A
00D1 N tilde N or NXX
00D6 O diaeresis OE or O
00D8 O stroke OE
00DE Thorn (Iceland) TH
1E9E Capital sharp s SS

Two entries deserve comment. Thorn, the Icelandic letter, becomes two letters, so the name grows. And the capital sharp s of German becomes SS, which means a name containing it grows by one character every time it is written into a travel document.

The “NXX” alternative for N-with-tilde is the interesting one. Doc 9303 Part 3 explains in Appendix B that nine characters may be transliterated with an X as an escape marker, so the original letter can be recovered for database searching. Its worked example is Térèsa Cañón, which becomes CANXXON, two fillers, TERESA. The standard concedes the result “appears unaesthetic (and may lead to complaints)” and defends it because the zone exists for machines and this way CAÑÓN stays distinct from CANON. Both options are permitted, so two States may encode one name differently and both be conformant.

The machine readable zone: name rules, field by field#

The strip of capital letters at the bottom of a passport is the machine readable zone, and its name rules are the sharpest example anywhere of what a fixed-width field does to human names. The relevant edition as of August 2026 is Doc 9303, Eighth Edition, 2021, carrying amendments of 14 November 2022 and 20 March 2024 in Part 3 and 20 March 2024 and 20 February 2026 in Part 4. Three form factors, three sizes.

Form Layout Name chars
TD3 passport 2 lines of 44 39
TD2 card 2 lines of 36 31
TD1 card 3 lines of 30 30

The rules are in Doc 9303 Part 3, Section 4.6, “Convention for writing the name of the holder”, and are worth stating exactly, because almost every practitioner misremembers at least one.

The name is split into a primary identifier and a secondary identifier, the standard’s terms, chosen because “surname” and “given name” do not travel. The primary is whatever the issuing State treats as the predominant component; the secondary is the rest; either may be empty. The primary is written first, followed by two filler characters, and the secondary begins immediately after them. Within either, components are divided by a single filler, and all remaining positions are packed with fillers to the end of the field.

Punctuation is not permitted, and Section 4.6 disposes of each kind. An apostrophe is deleted and the parts joined, so D’ARTAGNAN becomes DARTAGNAN. A hyphen becomes a single filler, so MARIE-ELISE becomes MARIE, filler, ELISE, and can never be reconstructed as hyphenated. A comma separating primary from secondary becomes two fillers, a comma between components one. Every other mark is deleted with nothing inserted, and numeric characters are forbidden outright.

Here is our worked example. Spain treats both of María José Cañón Iglesias’s surnames as the primary identifier; the diacritics go under the Section 6 tables, the space between surnames becomes one filler, the boundary between surnames and given names two, and the space between the given names one.

Printed above:  Canon Iglesias, Maria Jose
Primary:        CANON<IGLESIAS        (14 characters)
Secondary:      MARIA<JOSE            (10 characters)
Joined:         CANON<IGLESIAS<<MARIA<JOSE   (26 of 39)
Fillers needed: 13

The full upper line of a TD3 passport puts a two-character document code and a three-character State code in front of the thirty-nine. Under Doc 9303 Part 4 Section 4.4 an ordinary national passport carrying a harmonized secondary document code uses PP.

         12345678901234567890123456789012345678901234
line 1:  PPESPCANON<IGLESIAS<<MARIA<JOSE<<<<<<<<<<<<<
         | |  |                                     |
         | |  +---- name field, columns 6 to 44 ----+
         | +---- issuing State, columns 3 to 5
         +------ document code, columns 1 to 2

Had Spain chosen the escape convention from Appendix B, the same line would carry CANXXON where this one carries CANON, and a system that did not know the convention would read her surname as “Canxxon”. Both encodings conform, so two States may write the same name two different ways and both be right. That is the state of the art.

Truncation, and the letter that means nothing survived#

Thirty-nine characters is not many. Doc 9303 Part 4, Section 4.2.3, “Truncation of names in the MRZ”, sets out what happens when a name does not fit. Characters are removed from components of the primary identifier until three positions are freed, so that the two fillers and at least the first character of the secondary identifier fit; further truncation of the primary is permitted to let more of the secondary in. There is a signal: the last character of the name field, position 44, must be alphabetic, and the standard says this “indicates that truncation may have occurred”.

The specification’s own examples teach it best. Two of them share the primary identifier NILAVADHANANANDA and neither fits: Nilavadhanananda Chayapa Dejthamrong Krasuang, and Nilavadhanananda Arnpol Petch Charonguang. The standard permits both of these outcomes.

a) trailing component cut to an initial
   PPUTONILAVADHANANANDA<<CHAYAPA<DEJTHAMRONG<K

b) trailing component simply cut short
   PPUTONILAVADHANANANDA<<ARNPOL<PETCH<CHARONGU

Both lines are exactly forty-four characters. In the first a whole given name has become the letter K; in the second it ends mid-word at CHARONGU. Neither can be undone, and anybody scanning that passport is reading a name that does not exist.

Now the sting, also in the specification, at Section 4.2.3.4. Consider Jonathon Warren Trevor Papandropoulous, whose name comes to exactly thirty-nine characters and is not truncated at all.

   PPUTOPAPANDROPOULOUS<<JONATHON<WARREN<TREVOR

The last character is a letter because the name happens to fill the field. The standard’s note is explicit: even though the name has not been truncated, “it must be assumed that it has been truncated”. The signal is one-way. A letter in the last position means the name may be incomplete; it never means it is complete. Any downstream system that treats a scanned name as authoritative is therefore, on the specification’s own admission, sometimes working from a fragment and can never tell which times.

The consequence is direct. The machine readable zone is not a source of truth for a person’s name; it is a scanning aid. The name of record must come from the visual zone, the chip, or the issuing authority, and chapter 6 covers what each of those is worth.

Unicode: what a letter actually is#

Underneath all of this sits the question of what a computer holds when it holds a letter. Unicode assigns every character a number, called a code point, written in the form U+00F1. The current version as of August 2026 is Unicode 17.0, released in September 2025.

The complication for names is that the same visible letter can often be encoded in more than one way, because Unicode inherited precomposed characters from older standards as well as a general mechanism for combining marks. The letter ñ can be U+00F1, a single code point meaning “Latin small letter n with tilde”, or U+006E followed by U+0303, meaning “Latin small letter n” then “combining tilde”. The two are canonically equivalent: they mean the same thing and must render the same. They are not the same bytes.

Unicode Standard Annex 15, “Unicode Normalization Forms”, resolves this. As of August 2026 it is at revision 57, dated 30 July 2025, matching Unicode 17.0. It defines four normalization forms, and the difference between two of them is the difference between a working system and a corrupted one.

Form Operation Safe for names
NFD Canonical decompose Yes
NFC Decompose, recompose Yes, preferred
NFKD Compatibility decompose No
NFKC Compat., then compose No

NFC and NFD are lossless with respect to meaning. They rearrange how a letter is spelled internally, and converting between them changes no character’s identity. NFC is the form to use: the W3C Character Model for the World Wide Web recommends Normalization Form C for all content, and UAX #15 says so explicitly.

Here is the worked example measured exactly. Her name in the two canonical forms, with counts.

Maria Jose Canon Iglesias, with real accents:

NFC:  25 code points, 29 bytes in UTF-8
      ... 0043 0061 00F1 00F3 006E ...   (C a n-tilde o-acute n)

NFD:  29 code points, 33 bytes in UTF-8
      ... 0043 0061 006E 0303 006F 0301 006E ...

Identical on screen. Not equal as strings.

Four extra code points and four extra bytes for the same name. That is where a column holding twenty-five characters fails: it takes the composed form and rejects or silently truncates the decomposed one. Worse, byte-oriented truncation of the decomposed form can cut between a base letter and its combining mark, or inside a multi-byte character, producing a string that is not valid text at all.

The K forms, NFKC and NFKD, are a different animal and must be kept away from names. The K stands for compatibility, and these forms deliberately destroy distinctions that Unicode considers merely presentational. Under NFKC the ligature fi becomes the two letters f and i; the Roman numeral Ⅸ becomes I and X; full-width Japanese forms collapse to half-width; the single character meaning “kabushiki gaisha” expands to four. Every one of those is irreversible, and names contain exactly these characters. A person whose name is legally registered with a ligature or a compatibility character will find NFKC quietly rewriting it. The rule is short: normalize names to NFC, never to NFKC, and use NFKC if at all for a separate search index rather than for the stored name.

One property makes the rule safe to apply everywhere. Unicode’s normalization stability policy fixes the composition version at Unicode 3.1.0 and guarantees that a string normalized under one version stays normalized under every future version, so choosing NFC is not a decision you will have to revisit when Unicode adds characters.

Homoglyphs: two names that look the same and are not#

Canonical equivalence handles two encodings of the same character. The harder case is two genuinely different characters that look identical.

The Latin letter a is U+0061; the Cyrillic letter a is U+0430; in almost every typeface they are the same shape. A name written with one of each is two strings no normalization will ever reconcile, because they really are different letters. Unicode Technical Standard #39, “Unicode Security Mechanisms”, at revision 32 dated 4 September 2025, addresses this. It publishes a data file of visual confusables and defines a “skeleton” transformation: map every character to a prototype form, and two strings are confusable when their skeletons are equal. It also defines restriction levels for identifiers, from ASCII-Only through Single Script, Highly Restrictive, Moderately Restrictive and Minimally Restrictive to Unrestricted, and a mixed-script rule based on intersecting the script sets of every character in a string. A string whose resolved script set is empty is mixed-script, and mixed script inside one word is the signal that something is being spoofed.

The demonstration everyone remembers came from domain names. On 14 April 2017 the researcher Xudong Zheng published a working proof of concept in which a registered domain rendered in the address bar as “apple.com” while consisting entirely of Cyrillic letters. Because every character came from one script, the browser’s mixed-script defence did not fire. He had reported it to Chrome and Firefox on 20 January 2017; the fix landed in the Chrome trunk on 24 March and shipped in Chrome 58.

For names the consequence is subtler than phishing and just as real. A sanctions screening run, a duplicate check or a fraud rule that compares raw strings will treat a Cyrillic-a name and a Latin-a name as unrelated. That is exploitable at enrolment and also an ordinary accident, because somebody copying their own name from one document into a form can easily carry a stray character across. Compute and store a confusable skeleton alongside the name, and flag mixed-script single words for review rather than rejecting them, because genuine mixed-script names exist. Invisible characters belong in the same discussion: a zero-width joiner inside a name occupies no visual space and changes every comparison, so strip the defined default-ignorable characters in the comparison form and not in the stored name.

Case: folding, the dotless i, and why not to upper-case a name#

Case looks like the easiest problem here and does the most damage, because the operation feels reversible and is not.

Unicode holds case data in three files. UnicodeData.txt carries the simple one-to-one mappings; SpecialCasing.txt carries the one-to-many mappings, such as the uppercase of the German sharp s; CaseFolding.txt carries the mappings used for case folding. Folding is not case mapping: it exists for caseless comparison and is deliberately language-neutral, while mapping exists to produce text for display and is not.

Three facts follow, and each breaks a common assumption.

Case mapping loses information. The German sharp s, ß, uppercases by default to the two letters SS, so the name Straßer becomes STRASSER and never comes back; STRASSER lowercases to strasser. A capital form has existed since Unicode 5.1.0 in April 2008 added U+1E9E, Latin capital letter sharp s, and German orthography accepted it when the Council for German Orthography updated the ruleset on 29 June 2017: paragraph 25 E3 states that in all-capitals one writes SS, and that the capital ẞ is also possible. Both are correct, both appear in real records, and the default Unicode mapping still produces SS.

Case mapping is language-dependent. Turkish and Azerbaijani have four letters where English has two.

Character Code point Turkish pair
I U+0049 lower is U+0131
i U+0069 upper is U+0130
dotless i U+0131 upper is U+0049
dotted I U+0130 lower is U+0069

Under the default, language-neutral mappings, uppercase of i is I and lowercase of I is i. Under Turkish rules, uppercase of i is İ and lowercase of I is ı, so a round trip fails. Take the Turkish given name Işık, which begins with a capital dotless I and contains a lowercase dotless ı. Uppercase it and you get IŞIK, correct under both conventions. Lowercase that with the default rules and you get “işik”, with dotted letters where dotless ones belong. The name is now spelled wrong and nothing downstream will notice.

Isik  ->  upper  ->  ISIK  ->  lower  ->  isik
(dotless i's)                        (dotted i's)

Round trip does not return the original.

That is the famous “Turkish I problem”, and it is not confined to names. Its worst recorded consequence was in a text message. In April 2008 the Turkish press reported, and the linguist Mark Liberman discussed on Language Log on 22 April 2008, a case in which a man sent his estranged wife a message containing the word “sıkışınca”, with a dotless ı, from a handset that did not have the character. Her phone displayed the word spelled with a dotted i, which is an obscenity. The misreading led to a confrontation in which the woman was fatally stabbed, and the sender later killed himself in prison. The mechanism is exactly the substitution above, applied to a word instead of a name.

Case folding is not lowercasing, and full folding is not simple folding. CaseFolding.txt marks entries C for common, F for full, S for simple and T for the Turkic variants. Full folding maps ß to the two characters ss, so Maße and MASSE fold alike; simple folding keeps a single character where it can, for systems that cannot change string length. Choosing one and using it everywhere matters more than which.

Three conclusions, and they are absolute. Never store an upper-cased name. Never round-trip a name through case. And never call your language’s default case operations on a name in a way that depends on the machine’s locale setting, because the same line of code then gives different answers on different servers.

Length: the numbers that hurt people#

Field lengths are chosen carelessly, encoded in a hundred systems, and paid for by the person on the other end.

The clearest case is Janice Lokelani Keihanaikukauakahiheʻekahaunaele of Hawaii, whose surname is thirty-six characters long, counting the ʻokina as the letter it is in the Hawaiian alphabet. Hawaii’s driving licence system allowed thirty-five characters across the name fields, so her licence carried her surname with the final character cut off and omitted her first and middle names entirely: an official identity document bearing a name that was not hers and not even complete. The case became public in 2013, the state’s Department of Transportation changed its rules to forty characters for the last name, forty for the first and thirty-five for the middle, and she received a licence with her full name on 30 December 2013.

The second case is Christopher Null, who wrote about it in Wired in November 2015, and Jennifer Null, who spoke to the BBC in March 2016. “Null” is the word many programming languages and databases use to mean “no value here”, so careless software converts the surname into the absence of a surname and the booking system reports that the field was left blank. Jennifer Null described being unable to buy plane tickets online, unable to enter her details on a government tax site, and unable to use the online system that notified her of substitute teaching shifts, so that she arranged every shift by telephone. Told the reason, airline staff replied “there’s no way that’s true”. The problem is old: the New York Times ran “Why, O Why, Doesn’t That Name Compute?” on 28 August 1991, about apostrophes in Irish surnames.

The third case is a length rule imposed by a government rather than suffered from one. Indonesian Minister of Home Affairs Regulation number 73 of 2022, issued on 21 April 2022, requires a name in population documents to consist of at least two words and no more than sixty characters including spaces, drawn only from the Indonesian alphabet, without abbreviations, numerals or punctuation. It does not force existing single-name holders to change, but it ends the mononym for newly registered children, and it exists in large part because the rest of the world’s systems could not cope with Indonesian mononyms. Rather than fix the software, a country changed the names of its citizens.

The fourth is a uniqueness assumption acting like a length limit. On 23 July 2007 the CBC reported that Citizenship and Immigration Canada had for about ten years required applicants surnamed Singh or Kaur to adopt a different surname, and published a letter from the Canadian High Commission in New Delhi stating that “the names Kaur and Singh do not qualify for the purpose of immigration to Canada”. A departmental spokeswoman gave the reason as “because it is so common”, and confirmed that no comparable policy applied to other common surnames. Singh is given to every baptized Sikh man and Kaur to every baptized Sikh woman. The policy was reversed within days of becoming public.

Limit System Consequence
35 chars total Hawaii licence Name cut, given names lost
39 chars Passport MRZ Silent truncation
The word Null Many Surname read as empty
60 chars, 2 words Indonesia 2022 Mononyms ended

There is no defensible small number. The floor is that storage must comfortably exceed the longest name any document you accept can carry, in the decomposed encoding, measured in code points rather than bytes: in practice a variable-length text column with no small fixed limit and a validation maximum in the low hundreds, chosen to stop abuse rather than to stop names. The German standard DIN 91379, published in August 2022 as “Characters and defined character sequences in Unicode for the electronic processing of names and data exchange in Europe”, exists to give European systems a normative character repertoire for names rather than leaving each one to guess.

Names change, and the record has to hold both facts#

A name is a time-varying attribute, and a record must answer two questions: what is this person called now, and what were they called when this event happened. A system storing only the current name cannot reconstruct its own history; a system showing every historical name to every operator commits a serious harm. The design must do both.

Marriage in England and Wales changes nobody’s name by operation of law; it produces a certificate a person may use as evidence of a change they have chosen. In Spain it changes nothing at all. Elsewhere a change is automatic or compulsory, and divorce reverts nothing automatically anywhere. A system that models “maiden name” as a fixed field has encoded one country’s custom and one gender’s experience; GOV.UK recommends “previous name”, because it is not only women who change their family names. The FHIR HumanName type makes the same move with a “use” code whose permitted values are usual, official, temp, nickname, anonymous, old and maiden, so a former name is another name with a different use and a validity period rather than a special field.

Gender recognition shows what a well-designed record looks like, because the United Kingdom legislated the data model explicitly. The Gender Recognition Act 2004, chapter 7, provides at section 9(1) that where a full gender recognition certificate is issued, “the person’s gender becomes for all purposes the acquired gender”, subject at section 9(2) to the rule that this does not affect things done or events occurring before the certificate is issued.

The registration machinery is in section 10 and Schedule 3. Rather than altering the original birth entry, the Registrar General maintains a separate Gender Recognition Register and makes traceable the connection between the new entry and the original one. Both the register and the information kept to maintain that connection are closed to public inspection and search, and a certified copy “must not disclose the fact that the entry is contained in the Gender Recognition Register”, so the certificate a person receives is indistinguishable from an ordinary birth certificate.

Section 22 completes the design by making disclosure an offence. Where a person has acquired “protected information” in an official capacity, it is an offence to disclose it, punishable on summary conviction by a fine not exceeding level 5 on the standard scale. Protected information covers the fact of an application and the person’s gender before it became the acquired gender. Subsection (4) lists the exceptions: information that does not identify the person, the person’s consent, a discloser who did not know a certificate had been issued, court orders, legal proceedings, the prevention or investigation of crime, disclosure to the registrars general, and social security and pension purposes.

Read as a specification rather than as law, that is a precise data architecture. Keep the link, because the record must remain correct. Close the link, because exposure causes harm. Enumerate the lawful readers. Make unlawful reading an offence rather than a policy breach. Few systems built without a statute behind them achieve that shape, and it is the shape to aim at for every sensitive name change.

The surrounding facts carry dates. The application fee was cut from 140 pounds to 5 pounds on 4 May 2021; as of August 2026 the GOV.UK service page states that it costs 6 pounds, and legislation.gov.uk records the consolidated Act as up to date to 13 August 2026 with pending amendments outstanding. Volume V handles the wider legal position; the point here is the record design.

The principle that falls out of all of these cases is one line. Never overwrite a name. Insert a new one, close the period on the old one, and mark which is current. Then decide separately, and restrictively, who may see the closed ones.

What to store#

Everything above converges on one storage design, and the difficulty is only in resisting the urge to model names more finely than you can.

Store one required field, the full name, as the person writes it. That is what vCard does. RFC 6350, the vCard Format Specification of August 2011, makes the FN property, the formatted name, mandatory with cardinality one or more, and the structured N property optional with cardinality zero or one. The full name is the fact; the parts are a convenience. RFC 9554 of May 2024 extended N with two more components, secondary surname and generation, which is the Spanish problem and the Junior problem arriving thirteen years late.

Store optional structured parts only when a real requirement needs them, and get them from the person rather than by splitting the full name. The modern reference model is RFC 9553, “JSContact: A JSON Representation of Contact Data”, of May 2024, whose Name object holds a “full” string plus an optional list of components, each with a kind drawn from an explicit list: title, given, given2, surname, surname2, credential, generation and separator. Note what that gets right. It has “given2” rather than “middle name”, so it can hold a patronymic; “surname2” for a second family name; “generation” as its own kind rather than glued to the surname; and a separator component, so ordering and punctuation are data rather than assumptions.

Store a display form and a sort key separately, as derived values, and never compute the sort key by taking the last word.

Store a comparison form separately again, and never show it to anyone; this is where normalization, case folding and confusable skeletons belong. The model to copy is the PRECIS framework: RFC 8265 of October 2017, “Preparation, Enforcement, and Comparison of Internationalized Strings Representing Usernames and Passwords”, which obsoletes RFC 7613. Its UsernameCaseMapped profile applies, in order, a width mapping rule, a case mapping rule that maps uppercase and titlecase code points to lowercase, a normalization rule applying NFC, and a directionality rule; UsernameCasePreserved is the same without the case mapping. Those are profiles for usernames, not for human names, and the reason to know them is that they show what a defensible comparison pipeline looks like: an ordered list of transformations, specified once, applied identically everywhere.

Finally, know what the identity protocols hand you, because you will often be receiving names rather than collecting them. OpenID Connect Core 1.0, in the version incorporating errata set 2 dated 15 December 2023, defines the standard claims at section 5.1: “name” is “End-User’s full name in displayable form including all name parts, possibly including titles and suffixes, ordered as the End-User would normally display them”, while “given_name” warns that “in some cultures, people can have multiple given names; all can be present, with the names being separated by space characters”, and “family_name” says the same. The protocol will hand you a single string in a field you were planning to treat as one word.

CREATE TABLE person_name (
  name_id       uuid PRIMARY KEY,
  person_id     uuid NOT NULL REFERENCES person(person_id),
  full_name     text NOT NULL,          -- as written, NFC
  name_use      text NOT NULL,          -- official|usual|old|...
  given         text[],                 -- optional, from the person
  surname       text[],                 -- optional, may be 2
  generation    text,                   -- Jr, III
  credential    text,                   -- letters after
  sort_key      text NOT NULL,          -- chosen by the person
  script        text,                   -- ISO 15924, e.g. Latn
  valid_from    date NOT NULL,
  valid_to      date,                   -- null means current
  CONSTRAINT one_current EXCLUDE USING gist
    (person_id WITH =, daterange(valid_from, valid_to) WITH &&)
    WHERE (name_use = 'official')
);

Five things there are deliberate. The full name is the only mandatory text. The structured parts are arrays, because both given names and surnames come in pluralities. The sort key is stored rather than computed. The script tag lets the same person hold a name in Devanagari and a name in Latin script as two rows rather than one mangled row. And the validity period with its exclusion constraint enforces at most one official name at any instant while keeping every previous one, which is the shape the Gender Recognition Act’s Schedule 3 reaches by another route.

   entry            storage                comparison
   -----            -------                ----------
   what the   ->    full_name    ->    NFC normalize
   person           (exact,            case fold
   typed            NFC only)          strip ignorables
                        |              confusable skeleton
                        |                     |
                        v                     v
                    display,              match, dedupe,
                    printing              screening
                    (chapter 6)           (chapter 8)

   One arrow is lossless. The other is lossy on purpose.
   Never let the lossy one write back into storage.

A validation policy is only useful if it is short enough to apply. Accept every Unicode character except control characters, and normalize to NFC on receipt. Do not strip accents, change case, remove apostrophes or hyphens, or reject digits outright. Do not require a minimum of two characters; single-letter surnames exist, and the W3C guidance warns against assuming a single letter is an initial. Do not require more than one word, and do not require a family name at all. Set the maximum generously and measure it in code points after normalization rather than in bytes.

Label fields by role: “given name” and “family name”, never “first” and “last”, both optional when a full name is present. Do not ask for a title unless legally obliged to, and use “previous name” rather than “maiden name”. Where a document standard forces a stripped form, keep that form as an attribute of the document, not as a correction to the person. And when a name does not fit, do not truncate silently: fail loudly to an operator who can widen the field, because a silently truncated name is a person your system has quietly renamed.

5.98 Common wrong ideas#

Wrong: Splitting a name into first and last is a reasonable simplification that works for most people. Right: It encodes four false claims at once: that everybody has exactly two meaningful name parts, that one is inherited, that the inherited one comes second, and that both fit a fixed width. Mononyms break the first, Icelandic patronymics the second, Chinese and Hungarian order the third, Hawaiian and Thai lengths the fourth, and those groups run to hundreds of millions of people rather than a handful of outliers.

Wrong: The name in the machine readable strip of a passport is the holder’s name in a standard form. Right: It is a lossy token derived for scanning, limited to A to Z, the digits and one filler, capped at 39 characters on a TD3 passport, 31 on a TD2 card and 30 on a TD1 card, and truncated when it does not fit. Doc 9303 Part 4 section 4.2.3.4 states that even an untruncated name filling the field must be assumed truncated, so the strip is never an authoritative spelling.

Wrong: Normalizing a name with NFKC makes matching more reliable without real loss. Right: NFKC applies compatibility mappings that are irreversible: ligatures split, Roman numeral characters become letters, full-width and half-width forms collapse, single characters expand into several. Those transformations hit real registered names, particularly Japanese ones. NFC is the correct normalization for stored names, and it is what UAX #15 and the W3C Character Model recommend.

Wrong: Upper-casing a name is harmless because you can lower-case it again. Right: Case mapping loses information in both directions. The German sharp s uppercases by default to SS and never comes back. Turkish has four i-letters where English has two, so the default mappings turn dotless letters into dotted ones and Işık round-trips to işik. Case folding, defined in CaseFolding.txt, is a separate operation intended for comparison only and must never be written back to storage.

Wrong: Two names that look identical on screen are the same name. Right: They may differ in normalization form, in which alphabet the letters came from, or in invisible characters. The Latin a at U+0061 and the Cyrillic a at U+0430 are indistinguishable in most typefaces and are different characters that no normalization will ever merge. UTS #39 exists for this, and defines a skeleton transformation and mixed-script detection to catch it.

Wrong: A person has one legal name, and changing it is a formal legal act. Right: That is a local convention. In England and Wales a name changes by use; a deed poll evidences the change rather than causing it, an unenrolled deed poll can be made from age sixteen, and enrolment at the High Court is optional and as of August 2026 costs 53.05 pounds. Marriage changes no name automatically there, and changes nothing at all in Spain, while other jurisdictions treat the registered name as the only one that exists.

Wrong: Keeping every previous name on the record is good practice because history matters. Right: History matters and exposure harms. The Gender Recognition Act 2004 shows the shape: Schedule 3 keeps the connection between the Gender Recognition Register and the original birth entry, closes both to public inspection, requires that certified copies not reveal their source, and section 22 makes disclosure of protected information acquired in an official capacity an offence with a fine up to level 5 on the standard scale.

Wrong: Field lengths are an engineering detail with no human consequences. Right: Hawaii’s 35-character allowance put a document into a woman’s hand that omitted her given names and cut the last letter off her 36-character surname, corrected only after public attention with a new licence on 30 December 2013. Indonesia’s Regulation 73 of 2022 met the opposite pressure by requiring at least two words and no more than 60 characters, ending the mononym for newly registered children.

Wrong: A full name plus a date of birth is close enough to unique to act on. Right: In the 2015 South Korean census 21.5 per cent of the population was surnamed Kim, 14.7 per cent Lee and 8.4 per cent Park, so nearly forty-five in a hundred share one of three surnames, and given names come from a narrow pool as well. A name narrows a search and settles nothing; chapter 8 covers the shortlist.

Wrong: If we store what the user typed, byte for byte, we have solved the problem. Right: Storage is half the problem, because a system also compares, sorts and prints. Store the exact string, then derive a display form, a sort key chosen by the person, and a comparison key from one written-down pipeline, each in its own column so it can be rebuilt when the pipeline changes.

5.99 Chapter summary in 20 lines#

  1. A name is not an identifier, because an identifier must be unique, stable, fixed in shape and meaningless, and a name is none of those.
  2. A form with two name boxes is a law about what a person may be called, and it excludes hundreds of millions.
  3. Patrick McKenzie’s 2010 essay lists forty false beliefs about names, each a class of production defect.
  4. Mononyms, patronymics, name chains, double surnames, generational suffixes and honorifics are ordinary shapes a two-field form cannot hold.
  5. Iceland’s Personal Names Act No. 45 of 17 May 1996 builds surnames from a parent’s given name with the suffixes son and dóttir.
  6. “First name” is a defective label, because in China, Japan, Korea, Vietnam and Hungary the first name written is the family name.
  7. Japan’s own government adopted family-name-first Romanization on official documents from 1 January 2020.
  8. The correct labels are “given name” and “family name”, both optional, with the full name the only required field.
  9. Transliteration is reversible and letter-based, transcription is phonetic and lossy, and Doc 9303 admits documents carry the second.
  10. Doc 9303 Part 3 section 6 recommends transliterations in which A-diaeresis becomes AE or A, thorn becomes TH and the capital sharp s becomes SS.
  11. The passport machine readable zone allows 39 characters for the name on a TD3 book, 31 on a TD2 card and 30 on a TD1 card.
  12. It deletes apostrophes, turns hyphens into fillers, forbids digits and excludes titles and suffixes unless a State treats them as part of the name.
  13. A name that does not fit is truncated, signalled by a letter in the final position, but the signal is one-way and a name filling the field must be assumed truncated.
  14. One visible letter can be encoded more than one way, so a name is 25 code points composed and 29 decomposed, and the two are not equal.
  15. NFC and NFD are lossless, with NFC recommended by UAX #15, while NFKC and NFKD destroy information and must never touch a stored name.
  16. Two names may be different strings that look identical because their letters come from different scripts, which UTS #39 catches with confusable skeletons.
  17. Upper-casing is lossy: the German sharp s becomes SS, and default mappings turn Turkish dotless letters dotted, so Işık does not survive a round trip.
  18. Case folding is for comparison, not storage, and locale-dependent case mapping makes one line of code give different answers on different servers.
  19. Names change through marriage, divorce, deed poll, gender recognition and correction, so a name needs a validity period and must never be overwritten.
  20. The design argued for here is one required full name stored exactly as written, optional structured parts supplied by the person, and separate derived columns for display, sorting and comparison.

Chapter sources: Patrick McKenzie, “Falsehoods Programmers Believe About Names”, Kalzumeus, June 2010; Chris Baraniuk, “These unlucky people have names that break computers”, BBC Future, 25 March 2016, for the Jennifer Null and McKenzie interviews, with Christopher Null in Wired, November 2015, and the New York Times of 28 August 1991, “Why, O Why, Doesn’t That Name Compute?”; ICAO Doc 9303, Machine Readable Travel Documents, Eighth Edition 2021 — Part 3 section 4.6 on writing the holder’s name, Part 3 section 6 “Transliterations recommended for use by States”, Part 3 Appendix B for the X-escape convention and the Arabic variant count, Part 4 sections 4.2.2, 4.2.3 and 4.2.3.4 for the 39-character field and the truncation examples, Part 5 for TD1 and Part 6 for TD2, with Part 3 amendments of 14 November 2022 and 20 March 2024 and Part 4 amendments of 20 March 2024 and 20 February 2026; Unicode Standard Annex #15, Unicode Normalization Forms, revision 57 of 30 July 2025 for Unicode 17.0, with the normalization stability policy and the W3C Character Model recommendation of Form C; Unicode Technical Standard #39, Unicode Security Mechanisms, revision 32 of 4 September 2025, with Xudong Zheng’s Unicode domain demonstration of 14 April 2017 fixed in Chrome 58; the Unicode case mapping FAQ with UnicodeData.txt, SpecialCasing.txt and CaseFolding.txt, the addition of U+1E9E in Unicode 5.1.0 in April 2008 and the Council for German Orthography’s update of 29 June 2017 at paragraph 25 E3; Mark Liberman on Language Log, 22 April 2008, for the Turkish dotless i case; Iceland’s Personal Names Act No. 45 of 17 May 1996, articles 4, 5, 6, 8, 21 and 22, the Reykjavík decision of 31 January 2013 in the Blær case, and the Gender Autonomy Act of 2019; the Gender Recognition Act 2004 chapter 7, sections 9, 10 and 22 and Schedule 3, on legislation.gov.uk up to date to 13 August 2026, with the fee cut of 4 May 2021 and the 6 pound fee on GOV.UK as of August 2026; GOV.UK guidance on deed polls, including the 53.05 pound enrolment fee as of August 2026; the GOV.UK Design System pattern “Ask users for names” and the W3C article “Personal names around the world”; United States Foreign Affairs Manual 8 FAM 403.1, revision of 8 November 2021, for one-word names and the NFN and FNU placeholders; Indonesian Minister of Home Affairs Regulation 73 of 2022, of 21 April 2022; the Spanish surname order change of 30 June 2017; Japan’s family-name-first Romanization from 1 January 2020; the 2015 South Korean census figures for Kim, Lee and Park; CBC News, 23 July 2007, on the Singh and Kaur letter; the Hawaii licence issued to Janice Lokelani Keihanaikukauakahiheʻekahaunaele on 30 December 2013; DIN 91379:2022-08; RFC 6350 of August 2011, RFC 9553 and RFC 9554 of May 2024, and RFC 8265 of October 2017; HL7 FHIR release 5 HumanName and its NameUse value set; and OpenID Connect Core 1.0 with errata set 2, 15 December 2023, section 5.1.