Skip to content
KEDBYTE
Site navigation
How Data Works
Chapter
7

Files, Folders and Data Formats

Part B · From Records to Databases|8,385 words|about 36 min read|Volume B

7.0 What this chapter gives you#

  1. You will separate the location of a file from the identity of its contents and the identity of the records inside it.
  2. You will follow bytes through decoding, parsing and validation without treating those as the same operation.
  3. You will read a CSV field containing a comma, a quotation mark or a line break without accidentally creating extra records.
  4. You will distinguish a JSON number, string, Boolean, null, object and array, then apply a separate business contract.
  5. You will explain what a binary format must specify and why changing a filename does not convert the contents.
  6. You will design a bounded import that accounts for accepted rows, repeated deliveries and rejected material.
  7. You will know when a parser must stop because it cannot safely identify the next record.

Part A established what the shop’s records mean. Part B asks how those meanings travel between programs and become enforceable structures. We begin with files because the simplest exchange is often an exported document, not a database connection. Mira sends Dev a file containing agreed order lines. The canonical orders remain O-1042 and O-1043: four lines, six units and a combined value of 28,650 paise. The deliberately defective rows in this chapter are separate synthetic import inputs; they do not change either original order.

7.1 Paths and file identity#

7.1.1 PLAIN — in simple words#

  1. A filename tells a program where to look. It does not, by itself, tell the program what it will find there.
  2. Imagine that Mira saves orders.csv on Monday and replaces it on Tuesday. The same name can now lead to different bytes. Conversely, two differently named files can contain exactly the same bytes.
  3. The records inside a file have another kind of identity. Order O-1042 remains the same identified order whether its representation arrives in an email attachment, an export folder or a database response.
  4. Therefore ask three separate questions: “Which location?”, “Which captured contents?” and “Which business records?” A useful import log answers all three rather than using the filename as a substitute for everything else.
  5. A folder is a way of organising names. The folder called approved does not prove that a human approved its contents. Naming conventions can help people work, but their meaning needs an actual process behind it.
  6. A suffix such as .csv is a useful hint. It is not an independent validation of the data. A program should still check that the contents satisfy the format and contract it expects.

7.1.2 PLAIN — a picture in your head#

  1. Picture a numbered shelf slot containing an envelope. The shelf address is the path. The sheets in the envelope are the file contents. The order numbers written on the sheets identify the shop’s orders.
  2. Move the envelope to another shelf and its contents need not change. Replace the envelope in the original slot and the address remains familiar even though the contents are new.
  3. Photocopy the sheets into a second envelope and the two physical envelopes are different, while the copied text is the same. A digital byte-for-byte copy creates the analogous distinction between file objects and contents.
  4. Now imagine a sign saying “use the envelope on shelf B instead”. This is a useful starting picture for a symbolic link: resolving one name can lead the program to another location. [S42]
  5. Where the comparison breaks: real filesystems have rules about links, access, open handles and replacement that depend on the platform. A paper shelf does not model those rules. In particular, a path checked a moment ago need not still name the same object when opened later.
  6. The practical lesson is not to memorise shelf metaphors. It is to stop treating a convenient label as permanent evidence of identity or authority.

7.1.3 PLAIN — a worked example#

  1. Mira has the following four captured inputs. The letters A and B stand for two different byte sequences in this exercise; they are not actual checksum values.
Captured input Name at capture time Contents Business records claimed
I-01 exports/orders.csv A Four canonical order lines
I-02 archive/orders-copy.csv A The same four lines
I-03 exports/orders.csv B A changed export requiring review
I-04 exports/orders.json C A JSON representation of the four lines
  1. I-01 and I-02 have different captured names but identical contents. Importing both without checking record identity could apply the same data twice.
  2. I-01 and I-03 have the same captured name but different contents. “I already processed orders.csv” is therefore not a complete explanation of what the system processed.
  3. I-01 and I-04 can describe equal domain records even though their bytes differ. A byte checksum alone cannot identify this kind of semantic equality.
  4. For each input, the proposed log records an import-run identifier, the displayed source name, the captured byte length, a digest, the selected contract and the outcome. Source authentication, where required, is a separate concern.
  5. If someone can replace both a file and its accompanying digest, matching the two only shows agreement between those supplied objects. It does not establish who produced the file. This follows directly from the assumed attacker’s ability to replace both.

7.1.4 PLAIN — what is really happening inside#

  1. A program asks the operating system to open a path. The operating system resolves names, applies its access rules and returns a handle through which the program can work with the opened object. Python exposes this through file objects. [S43] [S64]
  2. A relative path is interpreted from some starting location. If a script assumes that the current working directory is its own folder, running it from another directory can make it open the wrong file or find no file at all.
  3. An absolute path names a location from a filesystem root or other platform-specific anchor. It avoids that particular ambiguity, but it does not make the underlying contents immutable. [S42]
  4. Some filesystem identifiers can help recognise an object within their defined scope. They are not universal business identifiers and should not replace order keys. They may also be reused after an object is removed. [S64]
  5. Our educational importer avoids a filesystem race by accepting a byte string that the caller has already captured. Every parse, checksum and diagnostic in one call refers to that same byte string.
  6. This does not solve safe capture on every operating system. A production capture procedure must also handle a file being modified while it is read, permissions, size limits and the trustworthiness of the source. The exercise starts after that capture boundary.

7.1.5 TECHNICAL — the engineer’s version#

  1. Separate pathname, filesystem object, byte representation, logical record identity and provenance in the design. Each answers a different question; equality at one layer does not imply equality at all other layers.
  2. In Python, PurePosixPath and PureWindowsPath perform lexical path operations without opening files. Path adds filesystem operations for the current platform. A lexical check is not a security decision about a subsequently opened object. [S42]
  3. The following example manipulates a synthetic path only. It neither creates directories nor reads a real export.
from pathlib import PurePosixPath

p = PurePosixPath("exports/2026-09/orders.csv")
assert p.name == "orders.csv"
assert p.suffix == ".csv"
assert str(p.parent) == "exports/2026-09"
  1. A content digest is useful for identifying a captured byte sequence with an explicitly chosen algorithm. It should travel with the byte length and capture metadata. Do not call it proof that a sender was authorised or that a sale happened.
  2. A byte-oriented fingerprint also changes when insignificant formatting changes. Whitespace or object-member order can change JSON bytes while the application interprets the same fields. Comparing validated domain values is a different operation with its own equality rule. [S07]
  3. Implementation detail: filesystem case sensitivity, supported link operations and rename behaviour vary. Our exercises do not assume that Orders.csv and orders.csv are necessarily distinct on every machine. [S42]
  4. Chosen contract: this lab processes an already captured, at-most-one-mebibyte input in memory. It is not a general directory-walking service, an upload server or a complete defence against malicious filesystem races. Extending that boundary requires a new design and new tests.

7.1.6 WORDS — remember these#

  1. Path: instructions for locating something in a filesystem — a name interpreted under platform-specific pathname-resolution rules.

  1. Relative path: a location described from a starting point — a pathname whose interpretation depends on a working directory or another supplied base.

  1. File extension: the familiar suffix on a name — a naming convention that does not independently validate the represented format.

  1. Content digest: a compact fingerprint of captured bytes — the result of applying a specified hash function to an exact byte sequence.

  1. Capture boundary: the point at which this exercise obtains its input — the boundary after which one immutable byte string is used for parsing and diagnostics.

  1. Provenance: where information came from and how it was handled — recorded source, capture and transformation context, not automatic proof of truth.

7.2 Bytes, records and boundaries#

7.2.1 PLAIN — in simple words#

  1. A file normally gives the reader bytes. The reader needs rules to turn those bytes into useful values.
  2. When the file is text, the first rule says how bytes represent characters. This is the character encoding introduced in Chapter 3. A second rule says how those characters form records and fields.
  3. A third rule says what the fields mean and which values the application accepts. Reading a number successfully does not establish that the number is a valid order quantity.
  4. These are separate gates. “The bytes are valid UTF-8”, “the text is valid JSON” and “the record satisfies our order-line contract” are three different statements.
  5. A record needs a boundary so the reader knows where it stops. A format might use a delimiter, a fixed size, a length written before the content, or a structured grammar. The reader must use the actual rule, not a guess based on what looks tidy.
  6. A displayed line is not necessarily a record. Text wrapping is presentation, and some formats allow a real newline inside a field. Counting screen lines cannot reliably count business records. [S08]

7.2.2 PLAIN — a picture in your head#

  1. Imagine several parcels arriving in a delivery bag. The bag is the file. Each parcel is a record, and the labelled compartments inside a parcel are fields.
  2. The delivery worker needs a way to identify each parcel’s boundary. Separate boxes make it obvious. A roll of unmarked paper containing every address one after another does not.
  3. One possible rule is “every parcel begins with a label saying how many centimetres of wrapping belong to it”. That resembles a length-prefixed representation.
  4. Another possible rule is “a certain marker ends the parcel, unless it occurs inside an explicitly protected area”. That resembles quoting and delimiters in text formats.
  5. Where the comparison breaks: a parser has no human intuition about where a torn parcel ought to end. If a damaged length or quotation mark destroys the boundary, confidently skipping to the next visible line can manufacture records that were never present.
  6. The safe response may be to stop the whole import, retain the exact input and report a file-level error. A smaller amount of accepted data is not worth a false claim that the remainder was understood.

7.2.3 PLAIN — a worked example#

  1. Consider the text consisting of the letter A, the rupee sign and a newline. It contains three Unicode code points, but its UTF-8 representation contains five bytes.
Displayed meaning      A     rupee sign     newline
Unicode code points    0041  20B9           000A
UTF-8 bytes            41    E2 82 B9       0A
Total                  3 code points; 5 bytes
  1. The calculation is one byte plus three bytes plus one byte. A byte limit of four cannot hold the complete representation. Cutting after the first two bytes would also cut inside the rupee sign.
  2. Now define an original toy format for two ASCII words. Each word starts with a two-byte unsigned length, most significant byte first. The length counts payload bytes, not letters seen on screen.
00 03  70 65 6E
length 3; the bytes for "pen"

00 08  6E 6F 74 65 62 6F 6F 6B
length 8; the bytes for "notebook"
  1. The complete two-record representation uses 2 + 3 + 2 + 8 = 15 bytes. The format carries four bytes of length information and eleven bytes of word content.
  2. If the second length says eight but only six payload bytes remain, the reader has an incomplete record. It must not invent the missing two bytes.
  3. Conversely, if the contract permits exactly two records but more bytes follow, the reader must decide according to the contract whether those bytes are an error, another record or a defined extension. Ignoring them accidentally is not a specification.

7.2.4 PLAIN — what is really happening inside#

  1. A text decoder translates a byte sequence according to an encoding. It can reject an impossible sequence. Replacing damaged bytes with a replacement character may be useful for display, but it loses evidence and can change identifiers. [S43]
  2. A parser then follows a grammar. It distinguishes punctuation that structures a record from punctuation that belongs inside a value.
  3. The application applies validation after parsing. It can ask whether the expected fields exist, whether the version is supported and whether the quantity is a positive whole number.
  4. A streaming reader may receive only part of a record in one read. The amount returned by an input operation is an I/O boundary, not necessarily a record boundary. A parser must retain incomplete state or request additional bytes. [S43]
  5. Our lab uses a complete captured input instead of teaching network streaming at the same time. That keeps parsing and record validation visible. The two-word framing example still demonstrates why a known boundary matters.
  6. The stages should report errors at the stage where evidence supports them. Invalid UTF-8 is a decoding error. Missing CSV fields are a structural or contract error. An unknown product is a relationship error, not a broken byte encoding.

7.2.5 TECHNICAL — the engineer’s version#

  1. Use an explicit pipeline: capture → decode → parse → validate fields → check relationships and identity → publish. The exact partition can vary, but the distinctions must remain visible in results.
  2. Syntactic validity means conformance to a representation grammar. Domain validity means conformance to stated application rules. Neither establishes the truth of a real-world claim; that limitation carries forward from Chapter 2.
  3. Count lengths in the unit specified by the format. Bytes, Unicode code points and grapheme clusters are not interchangeable. The rupee example is a fixed illustration, not a universal one-byte-per-character rule. [S14] [S16]
  4. For fixed-width fields and length prefixes, specify byte order, widths, signedness, permitted range and the response to truncation. Python’s struct format prefixes make native versus explicitly specified byte layouts distinguishable. [S46]
  5. Do not use a lossy text-decoding mode in an evidence-preserving import unless that loss is an explicit, separately recorded transformation. This lab decodes UTF-8 strictly and retains the original bytes on failure. [S43]
  6. Chosen framing contract: the two-word example requires a two-byte length followed by exactly that many bytes for each payload. Its format is invented for teaching. It is not Avro, a network protocol or a recommendation to build a production file format.
  7. Separate an input’s size from the size of objects produced by parsing it. A small encoded input can produce many runtime objects, and nested structures can require substantial processing. This lab is bounded and deliberately small; its limit is not a complete resource-isolation system. [S45]

7.2.6 WORDS — remember these#

  1. Decoding: turning bytes into characters — interpreting a byte sequence under a specified character encoding.

  1. Parsing: finding the structure in a representation — recognising records, fields and values under a grammar.

  1. Framing: marking where messages or records begin and end — a boundary convention independent of the meaning of each payload.

  1. Length prefix: a size written before the content — a field giving the amount of following data in a stated unit.

  1. Truncation: an input ending before its promised content is complete — insufficient bytes to satisfy a declared structure or length.

  1. Syntactic validity: being written in the right shape — conformance to a format’s grammar, distinct from business rules.

7.3 CSV and quoting#

7.3.1 PLAIN — in simple words#

  1. CSV is a way to write rows of fields as text. The letters stand for comma-separated values. It is useful because many different programs can read and write it.
  2. The simple picture is “commas separate fields and line endings separate records”. The important correction is that a quoted field can itself contain commas and line endings. A real CSV parser understands that distinction. [S08]
  3. A quoted field is not automatically a special kind of business value. Quoting is part of how text is represented; the application still decides whether the resulting value is a product name, quantity or identifier.
  4. The column headings do not supply every missing rule. A column called amount is still ambiguous until the contract specifies currency, unit, scale and meaning.
  5. Different exports can use different delimiters and conventions. A file using semicolons is not repaired by pretending that its semicolons are ordinary content and its commas are separators.
  6. For the shop’s exercise we choose a precise CSV profile rather than trying to guess every possible export. Inputs that do not match the profile are reported, not silently “cleaned up”.

7.3.2 PLAIN — a picture in your head#

  1. Imagine writing a shopping list in a notebook and using a vertical line to separate each field. That works until a product name contains the same character.
  2. You add a rule: text inside a pair of quotation marks belongs to one field even when it contains a separator. The quotation marks describe the boundary; they are not normally part of the resulting field value.
  3. What happens when the product name itself includes a quotation mark? In the CSV convention used here, a doubled quotation mark inside a quoted field represents one literal quotation mark. [S08]
  4. Where the comparison breaks: the reader cannot rely on colour, handwriting or visual alignment. It follows the character sequence and its parsing rules exactly. A missing closing quote can make the boundary ambiguous for the rest of the input.
  5. Also, CSV quotation marks are not protection against spreadsheet formulas. A spreadsheet may interpret a field’s contents as a formula after the CSV layer has been decoded. Representation safety and spreadsheet behaviour are separate questions. [S63]
  6. That is why the lab inspects values as data in Python and does not automatically open unfamiliar exports in a spreadsheet.

7.3.3 PLAIN — a worked example#

  1. This independent two-column CSV illustration contains three logical records, including the header. The second data field contains a real line ending inside quotation marks.
product_id,description
P-PEN,"Blue, fine-tip pen"
P-NOTE,"Notebook with ""draft"" label
and plain pages"
  1. The first data record has two fields: P-PEN and Blue, fine-tip pen. The comma inside the description does not create a third field.
  2. The second data record also has two fields. Its description contains quotation marks around the word draft and a newline before the word and. The input has more physical lines than logical data records.
  3. In this example, line.split(',') gives the wrong answer. Splitting first on every newline and then on commas also gives the wrong answer.
  4. The following Python code is a compact, original parsing exercise. The string is complete before it is read.
import csv
import io

text = 'product_id,description\r\nP-PEN,"Blue, fine-tip pen"\r\n'
rows = list(csv.reader(io.StringIO(text, newline=""),
                       strict=True))
assert rows[1] == ["P-PEN", "Blue, fine-tip pen"]
  1. The parser returns strings. Turning a field such as 7550 into the integer number of paise is a later step under our explicit import contract. A field such as 007 used as an identifier must not be converted to seven merely because it looks numeric.

7.3.4 PLAIN — what is really happening inside#

  1. The reader tracks whether it is inside a quoted field. It interprets a separator or newline according to that state rather than simply splitting at every occurrence.
  2. A dialect describes representation choices such as the delimiter, quote character and escape behaviour. Python’s CSV module provides explicit reader and writer settings. [S44]
  3. A header row maps positions to names. Our profile requires the exact eight distinct headings in the agreed order. Duplicate, missing, unexpected or whitespace-altered headings cause a file-level rejection.
  4. Each later row must have eight fields. A valid CSV parser can still return a row with too many or too few fields, so the importer checks width separately.
  5. Only after that check does it convert the integer fields and call the existing validate_line function from Chapters 1–3. The domain function remains unchanged.
  6. A row with an empty quantity is rejected with its raw fields and reason retained. It is not converted to zero, and the importer does not shift the remaining fields left to conceal the problem.
  7. If the CSV parser cannot safely finish parsing the input, the lab withholds all candidates from that file. Earlier rows may help diagnose the failure, but they do not become a quietly accepted partial business import.

7.3.5 TECHNICAL — the engineer’s version#

  1. Documented format: RFC 4180 describes a common CSV format and registers text/csv; it is an Informational RFC, not a complete universal definition of every file called CSV. Our profile uses its familiar quoting rules with an explicit UTF-8 choice. [S08]
  2. Python recommends opening CSV text files with newline='', allowing the CSV module to handle embedded newlines. strict=True raises errors for certain malformed inputs, but is not a schema validator or a guarantee that the file obeys every stricter exchange profile. [S44]
  3. Chosen import profile: header names are record_type, schema_version, order_id, line_no, product_id, quantity, unit_price_minor, currency. The encoding is UTF-8 without a byte-order mark. Input LF and CRLF record endings are accepted by the Python reader; the supplied writer emits CRLF. This chosen adapter rejects a bare carriage return anywhere, including inside a field; embedded LF or CRLF remains supported.
  4. Integer fields use canonical ASCII decimal text: 0 or a nonzero digit followed by digits. Leading signs, outer whitespace, leading zeros on multi-digit numeric fields, fractions and exponent notation are rejected. The narrower text grammar is an adapter decision, not a change to the earlier integer domain model.
  5. The separate identifier fields remain text. Their leading zeros, case and punctuation are not normalised automatically. The domain contract continues to reject empty identifiers and outer whitespace.
  6. CSV has no universal built-in representation for SQL NULL, types, units or foreign keys. A metadata contract is needed to carry such semantics; W3C’s tabular-data model explicitly separates textual cell representations from their semantic values and annotations. [S49]
  7. Quoting a spreadsheet-sensitive field is not a general formula-injection defence. Keep archival machine-exchange data separate from a deliberately prepared spreadsheet-view export; the display transformation must be tested against the intended consuming software. The lab implements neither spreadsheet execution nor a universal sanitiser. [S63]
  8. The source bytes are the authoritative diagnostic artifact for this exercise. Parsed physical-line ranges and field lists are navigation aids. They do not replace the retained original representation.

7.3.6 WORDS — remember these#

  1. CSV: a text representation of tabular records — comma-separated values with a specified dialect and quoting rules.

  1. Delimiter: the mark used to separate fields — a structural character interpreted according to the parser’s current state.

  1. Quoting: marking a region as one field — a representation mechanism that allows otherwise structural characters inside a value.

  1. Dialect: the particular set of parsing conventions — settings governing delimiters, quotation marks, escaping and related behaviour.

  1. Logical record: one complete parsed record — a structure that may span more than one physical text line.

  1. Import profile: the exact agreement for an exchange — a bounded specification combining encoding, syntax, fields and application rules.

7.4 JSON and nested data#

7.4.1 PLAIN — in simple words#

  1. JSON gives names and values a written structure. An object groups named members; an array holds an ordered sequence of values. Objects and arrays can contain further objects and arrays. [S07]
  2. That makes JSON convenient for an order containing several lines. The order can be one object with an array called lines inside it.
  3. Convenience does not mean automatic correctness. The JSON parser does not know whether the array should contain one line or five, whether a price is in paise, or whether the order already exists.
  4. The value 2 is a number, "2" is text, and true is a Boolean. A receiving program must not quietly treat them as equivalent merely because all could be described informally as “something positive”.
  5. A member containing null is present with a null value. An absent member is not present at all. The application may assign different meanings to those cases. [S07]
  6. Reading JSON successfully means that a representation was recognised. It does not authenticate the sender, authorise an action or verify a transaction in the physical shop.

7.4.2 PLAIN — a picture in your head#

  1. Picture a folder with labelled pockets. One pocket says order identifier, another says currency, and a third contains an ordered stack of line cards.
  2. Each line card has its own labelled spaces for product, quantity and agreed price. This nested structure resembles a JSON object containing an array of objects.
  3. An empty pocket is not the same as a missing pocket. Nor is a pocket containing the text “unknown” automatically the same as the application’s null value.
  4. Where the comparison breaks: JSON defines representation kinds, not a full model of your organisation. A folder has physical ownership and access boundaries; a JSON object does not automatically acquire either.
  5. The analogy also does not justify relying on object-member order. An array has a defined order, while a JSON object’s members are not a suitable substitute for an ordered business sequence. [S07]
  6. When order matters, make it explicit with a suitable structure or field. The line number from the earlier chapters remains meaningful even when the surrounding object is printed in a different order.

7.4.3 PLAIN — a worked example#

  1. Here is the existing canonical first order line in its version-1 JSON representation. It describes two notebooks at 7,550 paise each.
{
  "record_type": "agreed_order_line",
  "schema_version": 1,
  "order_id": "O-1042",
  "line_no": 1,
  "product_id": "P-NOTE",
  "quantity": 2,
  "unit_price_minor": 7550,
  "currency": "INR"
}
  1. The arithmetic is unchanged: 2 × 7,550 = 15,100 paise. JSON did not discover this meaning; the field contract supplied it.
  2. Changing the quantity to the text "2" produces a JSON document that can still be parsed. Our integer-based line validator rejects it.
  3. Changing the quantity to true also produces parseable JSON. Python’s Boolean type has a relationship to integers, but our domain validator explicitly rejects Booleans for integer fields. [S22]
  4. Removing currency creates a different contract error: a required field is absent. Setting it to null leaves the field present, but still fails the INR-only rule.
  5. This separate counterexample is dangerous to interpret casually:
{"quantity": 2, "quantity": 9}
  1. There are two members with the same name. Rather than allowing an implementation’s default to choose a winner, our parser rejects the input before any domain decision. The reported conflict is about representation, not evidence that nine items were sold.

7.4.4 PLAIN — what is really happening inside#

  1. A JSON parser recognises brackets, braces, quoted strings, numbers and the literal values true, false and null, then constructs corresponding program values. [S07]
  2. Different languages represent numbers differently. A number that fits comfortably in one runtime’s integer type may lose precision in a consumer using a limited-precision floating-point representation. Representation compatibility must be checked across the actual participants.
  3. Python’s standard JSON decoder exposes hooks for object-member pairs and numeric conversions. The lab uses the pair hook to detect repeated object names before they collapse into an ordinary mapping. [S45]
  4. The lab also rejects nonstandard constants such as NaN and Infinity. Python’s default JSON behaviour can accept them, so relying on the default would make our stated profile less strict than intended. [S45]
  5. After the representation checks, the resulting object goes to the same order-line validator used by the CSV adapter. This keeps the business rules consistent while allowing different formats to reach them.
  6. Nested JSON is introduced conceptually here, but the executable version-1 line importer deliberately accepts one line object, not a complete nested order. A nested-order contract would need separate rules for the envelope, its lines and their relationships.
  7. Importers should make this boundary visible. “Can parse nested JSON” is not the same capability as “can validate and publish a complete order without partial failure”.

7.4.5 TECHNICAL — the engineer’s version#

  1. RFC 8259 recommends unique object-member names and describes divergent behaviour when names repeat. Our profile makes uniqueness mandatory. This is an explicit interoperability choice rather than a claim that the JSON grammar alone resolves every duplicate-member case. [S07]
  2. Python’s object_pairs_hook receives each object as ordered pairs, allowing a duplicate-name check before dictionary construction. parse_constant can reject the nonstandard numeric constants accepted by the default decoder. [S45]
  3. A compact duplicate-name hook can be understood without hiding its decision:
def unique_members(pairs):
    obj = {}
    for name, value in pairs:
        if name in obj:
            raise ValueError("duplicate JSON member: " + name)
        obj[name] = value
    return obj
  1. Type validation still follows. In our chosen contract, a parsed floating-point 2.0 is not accepted as an integer quantity even though it describes a mathematically integral value. The exchange rule could be different in another system, but must not vary accidentally between import paths.
  2. Version 1 rejects unknown fields instead of discarding them. A future optional field therefore requires a considered compatibility policy, not an assumption that a permissive old reader will preserve information it ignores.
  3. JSON is a data-interchange format, not a command language for the importer. Do not evaluate its text as program code. Python’s pickle, in contrast, can execute code during deserialisation and must not be treated as a safe substitute for untrusted JSON. [S62]
  4. The lab checks byte size and catches parsing or recursion failures, but it is not a complete hostile-input sandbox. Production services also need externally enforced CPU, memory, nesting and request-rate limits appropriate to their threat model.
  5. Source of truth: the schema version describes the meaning and permitted shape of our message. A package version, a Python version and a JSON syntax specification identify different things; no one number replaces the others.

7.4.6 WORDS — remember these#

  1. JSON object: a group of named values — a collection of name/value members within a JSON representation.

  1. JSON array: a sequence of values — an ordered collection whose elements may themselves be structured.

  1. Duplicate member: a name repeated within one object — an interoperability hazard rejected by this book’s input profile.

  1. Deserialisation: rebuilding values from a stored representation — decoding and interpreting data into runtime objects under a specific format.

  1. Schema version: the named edition of a record contract — an identifier for the supported structure and meaning, distinct from a software release.

  1. Unknown field: a member the contract does not recognise — data requiring an explicit reject, ignore, preserve or extension policy.

7.5 Binary formats and versioning#

7.5.1 PLAIN — in simple words#

  1. Every file is stored as bytes. “Binary format” usually means that its organisation is not simply a human-readable text grammar.
  2. A binary format can represent a number directly in a fixed number of bytes instead of writing its decimal digits. To read it, the receiver must know how many bytes belong to the number and how to interpret their order.
  3. Saving fewer bytes is not the only reason to choose a format. Efficient scanning, nested structure, compatibility, supported tools and future readability can matter more than a tiny size difference.
  4. Some formats carry a description of their structure. That helps a reader decode values, but it still does not tell the reader every business rule or whether the recorded facts are true.
  5. A new version should say what changed and which readers can understand it. Writing the number two on a file is not a compatibility strategy by itself.
  6. The shop’s first requirement is not “use the most advanced format”. It is “preserve these particular meanings between these particular writers and readers, with errors that we can explain”.

7.5.2 PLAIN — a picture in your head#

  1. Think of a custom tray containing exactly sized compartments. A four-unit compartment holds a price, another holds a quantity, and a label says which tray design is being used.
  2. A reader with the correct tray diagram knows where to find each value. A reader using a different diagram may take part of the price as part of the quantity.
  3. A versioned diagram therefore matters as much as the tray. Changing the compartment layout without changing the agreement is how neatly stored bytes acquire the wrong meaning.
  4. Where the comparison breaks: binary formats need not use fixed compartments. They may use variable-length encodings, compression, repeated structures, dictionaries and metadata. The tray models explicit layout, not every storage technique.
  5. Also, a compact tray is not necessarily easier to work with. If Mira’s partner cannot read it, the few saved bytes do not compensate for an unusable exchange.
  6. The right design is the one that matches a documented workload and an actual compatibility requirement, not the one with the most impressive file extension.

7.5.3 PLAIN — a worked example#

  1. The decimal price 7,550 can be represented as a four-byte unsigned integer. With the most significant byte first, those bytes are 00 00 1D 7E.
  2. Reversing the byte-order assumption changes the interpretation dramatically. The same four bytes read least-significant-first produce 2,115,829,760, not 7,550.
  3. Python makes the convention visible:
import struct

encoded = struct.pack(">I", 7550)
assert encoded.hex() == "00001d7e"
assert struct.unpack(">I", encoded)[0] == 7550
assert struct.unpack("<I", encoded)[0] == 2115829760
  1. The symbols > and < choose the byte order. The I in this standard-size layout represents an unsigned four-byte integer. These are struct conventions, not letters stored as part of the integer itself. [S46]
  2. Our separate toy envelope begins with four marker bytes, one version byte and a two-byte payload length. The payload is an existing JSON line. It is intentionally simple enough to inspect by hand.
4 bytes   marker: KDBF
1 byte    envelope version: 1
2 bytes   payload byte length, big-endian
N bytes   UTF-8 JSON payload, under the line-v1 contract
  1. The envelope therefore occupies 7 + N bytes. A different marker, unsupported version, short payload or unexpected trailing bytes is rejected. A matching marker only selects a parser; it does not authenticate the file or validate its business meaning.

7.5.4 PLAIN — what is really happening inside#

  1. The envelope parser checks its fixed header before interpreting the payload. It then compares the declared payload length with the actual remaining bytes.
  2. Only a complete, supported envelope reaches the JSON reader. The JSON reader still performs its own checks, and the resulting values still reach the domain validator.
  3. These nested checks form layers rather than replacements. A correct envelope does not rescue malformed JSON, and valid JSON does not rescue a negative quantity.
  4. A more capable format can organise data for a particular access pattern. Apache Parquet, for example, is column-oriented and designed for efficient bulk storage and retrieval. That design is useful context for the later analytical chapters, not a reason to replace every small file. [S47]
  5. Apache Avro specifies schemas and rules for resolving data written under one schema against a reader’s schema. That illustrates how compatibility can be defined as concrete rules rather than optimism. [S48]
  6. The lab does not implement or benchmark Parquet or Avro. Its invented envelope exists only to expose marker, version, length and payload boundaries. Use mature, appropriate formats for actual interchange rather than promoting the toy to a standard.

7.5.5 TECHNICAL — the engineer’s version#

  1. Specify a format’s syntax independently of the domain schema. The toy envelope has version 1, and the enclosed order-line message separately has schema_version: 1. Either layer could evolve without the other changing.
  2. Python’s native struct mode can depend on platform byte order, size and alignment. The examples deliberately use explicitly ordered standard-size fields instead of native layout. [S46]
  3. Define backward compatibility operationally as a new reader handling an older writer’s representation, and forward compatibility as an older reader handling a newer writer’s representation. State the direction because terminology is sometimes used inconsistently in project discussions.
  4. Consider a proposed v2 field discount_minor. An old reader that silently ignores it could compute a different payable amount. Syntactic tolerance would not establish semantic compatibility. This is an original hypothetical extension; discounts remain outside our actual v1 arithmetic.
  5. Under Avro’s schema-resolution rules, a reader can use a declared default for a field absent from a writer’s record. That is a specified interpretation, not evidence that the writer originally recorded the default value. [S48]
  6. Parquet’s project documentation cautions that different implementations support different features. Select and test a concrete writer/reader combination; a shared format name is not sufficient evidence of identical feature coverage. [S47]
  7. Compression, checksums, encryption and authentication answer different questions. A shorter payload is not thereby secret. A checksum agreement does not, by itself, identify an authorised producer. Our envelope supplies none of those additional security properties.
  8. The lab’s maximum payload is limited by an unsigned two-byte length, 65,535 bytes. Its other import bounds can be stricter. A bound should be checked before the parser treats a declared size as permission to allocate or read arbitrarily much data.

7.5.6 WORDS — remember these#

  1. Binary format: a byte layout read under explicit rules — a representation not restricted to a plain-text grammar.

  1. Byte order: which end of a multi-byte number comes first — the endianness used to encode a numeric value.

  1. Magic marker: an initial clue about the representation — a fixed signature used to select or validate a format boundary, not authenticate a producer.

  1. Envelope: the wrapper around a payload — metadata and framing interpreted before the enclosed message.

  1. Compatibility: the ability of specified participants to exchange usable data — a tested relationship between writer, reader and semantic contracts.

  1. Column-oriented format: a layout that groups values by column within its storage structure — an organisation suited to some selective and analytical reads, not a universal performance guarantee.

7.6 Import validation and quarantine#

7.6.1 PLAIN — in simple words#

  1. An import is not complete merely because a program reached the end of a file. The program must explain what it understood, what it rejected and what it actually made available for use.
  2. Rejected material needs a place to wait for review. We call that quarantine. In this chapter it means retained diagnostic data, not a secure storage product.
  3. A useful rejection includes the original input, the location of the problem where known and a reason that someone can act on. “Invalid” alone is usually not enough.
  4. Repeated rows also need an explicit outcome. Ignoring them without recording that decision makes a row-count reconciliation misleading.
  5. A conflict is different from an exact repeat. If two rows claim the same line identity with different quantities, choosing the last one simply because it came later is an invented policy.
  6. Our batch importer waits until it has examined the complete parseable input. It withholds all members of a conflicting identity group instead of choosing a winner. This is a new batch policy, distinct from the in-memory event-delivery example in Chapter 5.
  7. Finally, acceptance by the file adapter is not automatic publication into a database. The destination still needs its own relationship, permission, conflict and transaction checks.

7.6.2 PLAIN — a picture in your head#

  1. Picture a receiving desk with three labelled trays: candidates ready for the next check, repeated copies already accounted for, and material needing review.
  2. The receiving clerk also retains the delivery bag and a manifest. This lets someone check that nothing was silently thrown away while sorting the contents.
  3. If two differently completed forms claim the same order-line number, the clerk places the whole conflicting group in review. Arrival order is not allowed to decide which form is true.
  4. Where the comparison breaks: storing rejected customer data can create privacy and access responsibilities. “Quarantine” does not mean “keep everything forever in an open folder”. Retention and permissions need a separately defined policy.
  5. Our laboratory contains only synthetic bytes and reports retained in memory. It neither creates a production quarantine service nor proves compliance with any retention obligation.
  6. The transferable idea is accountable handling: a rejected input remains visible as rejected, and a human correction becomes a new input rather than a secret edit to the evidence.

7.6.3 PLAIN — a worked example#

  1. The supplied untidy CSV fixture contains eight logical data rows after its header. Four rows reproduce the canonical order lines. One repeats the first line exactly. Three are defective new rows.
Row Claimed line Relevant feature Outcome in this fixture
1 O-1042 / 1 Two notebooks at 7,550 paise Candidate
2 O-1042 / 2 One pen at 2,000 paise Candidate
3 O-1043 / 1 One notebook at 7,550 paise Candidate
4 O-1043 / 2 Two pens at 2,000 paise Candidate
5 O-X1 / 1 Quantity is empty Rejected
6 O-X2 / 1 Price text is 75.50, not integer paise Rejected
7 O-X3 / 1 Currency is USD, not INR Rejected
8 O-1042 / 1 Exact repeat of row 1 Duplicate
  1. The accounting identity is eight data rows = four candidates + three rejected rows + one duplicate. The header is not counted as a data row.
  2. The four candidates still describe two orders, four lines and six units. Their total remains 28,650 paise; the duplicate contributes no additional candidate line.
  3. The report retains all eight parsed rows, their outcomes and the captured input bytes. Rejection of rows 5–7 is visible rather than concealed in a lower output count.
  4. In a separate conflict fixture, two individually valid rows both identify O-1042 / 1 but give quantities two and three. Both are quarantined. There are zero candidates, two rejected rows and no invented “correct” quantity.
  5. A third fixture ends inside an open quoted field. The result is a file-level parsing error with zero candidates. The importer retains the raw input and does not pretend that it knows the business identity of every byte after the damaged boundary.

Figure 7.1 — An import keeps captured bytes, parses under a named profile, checks values and identity, and separates candidates, duplicate deliveries and rejected material. File-level syntax failure publishes no candidates.

Figure 7.1 — An import keeps captured bytes, parses under a named profile, checks values and identity, and separates candidates, duplicate deliveries and rejected material. File-level syntax failure publishes no candidates.

7.6.4 PLAIN — what is really happening inside#

  1. The importer first checks the byte bound and the declared text profile. It calculates a digest of the captured input and attempts strict decoding.
  2. It validates the header, then parses complete records. Each parsed record retains its fields and physical-line range before any numeric conversion.
  3. Field conversion and the existing line validator identify individually unacceptable rows. A rejected row’s missing quantity is not filled from a neighbouring record.
  4. For individually valid rows, the importer groups by (order_id, line_no). Equal validated values form one candidate plus recorded duplicates. Different values under the same key make the complete group conflicting.
  5. Only after the file has been parsed successfully does the function return candidates. Structural file errors, a record-limit breach or an undecodable input withhold the entire candidate set.
  6. The next stage may publish an explicitly accepted subset into a destination transaction. It must not claim that the source file was accepted in full when the report contains rejected rows.
  7. A corrected resubmission should be separately captured and linked to its predecessor by the surrounding workflow. The lab does not implement that persistent workflow; it leaves the necessary distinction visible.

7.6.5 TECHNICAL — the engineer’s version#

  1. The lab’s report separates file_error from a successfully parsed file with row-level findings. A candidate is a validated, non-conflicting domain object at this stage, not a committed database record.
  2. A normal parse has one decision per logical input data row. An exact duplicate retains its relationship to a representative row. A conflicting group retains all its members with an explicit conflict reason.
  3. The chosen identity comparison is over validated OrderLine values, not raw JSON text or a universal canonicalisation algorithm. Representation differences and application equality must remain separate.
  4. File-level failure retains raw bytes and any trustworthy diagnostics collected before the stop. It does not fabricate decisions for records whose boundaries were never established. Thus a file-level failure cannot use the same complete-row reconciliation claim as a fully parsed file.
  5. CSV numeric spellings are limited to 18 digits before integer conversion. Each parsed field is limited to 4,096 characters and the complete input to 10,000 logical data records. These are explicit adapter bounds, not universal CSV limits. The one-mebibyte input bound prevents this exercise from accepting arbitrarily large byte strings, but the caller still has to capture data safely before invoking it.
  6. Publication must account for destination state. A candidate whose parent order does not exist can fail a foreign-key check even though its input fields are well formed. Chapter 9 supplies that relational stage. [S59]
  7. A transaction should make the chosen publication unit all-or-nothing at the database boundary. It does not make a file-level rejection disappear, and it does not undo an email, payment or physical dispatch outside the transaction. [S01] [S25]
  8. The worked report is an original design demonstration. It is not a claim that every organisation should retain rejected bytes indefinitely, use the same conflict policy or accept partial batches. Those are explicit product decisions to be reviewed against the actual workflow.

7.6.6 WORDS — remember these#

  1. Quarantine: retaining rejected material for controlled review — a handling state with separately specified access and retention rules.

  1. Candidate: a record ready for the next acceptance stage — a validated, non-conflicting input that has not necessarily been published.

  1. Reconciliation: explaining how the input relates to the output — accounting for candidates, repeats, rejections and file-level limits without silently losing material.

  1. File-level error: a failure affecting interpretation of the input as a whole — an outcome that withholds publication when trustworthy complete parsing is unavailable.

  1. Conflict group: records claiming one identity with incompatible content — a set for which this batch policy refuses to choose a winner.

  1. Publication boundary: the point where accepted work becomes available in the destination — a separately controlled step beyond decoding and input validation.

7.97 Practice and worked answers#

  1. Question: two exports have different names but the same captured bytes. Does that prove two different orders occurred? Answer: no. Filenames identify locations or labels, not business events. Inspect the record identities and the intended delivery policy.
  2. Question: a five-byte UTF-8 input contains A, the rupee sign and a newline. How many code points does it contain? Answer: three. The rupee sign needs three of the five bytes. A byte count is not a character count.
  3. Question: a CSV description contains a comma inside quotes. Should the importer create another column? Answer: no. Parse according to the dialect; the comma is part of that field.
  4. Question: JSON parsing succeeds, but quantity is "2". Is the line accepted? Answer: not by this contract. The value is text, and the line validator requires an integer that is not a Boolean.
  5. Question: two valid rows have the same line key but different quantities. Which wins in the batch lab? Answer: neither. The entire conflicting identity group is withheld. The event-consumer policy in Chapter 5 is a different, explicitly scoped exercise.
  6. Question: an eight-row parse returns four candidates, one duplicate and three rejections. What total do the canonical candidates describe? Answer: 28,650 paise across four lines and two orders. The file was not accepted in full, and no candidate is a committed destination row merely because it passed parsing.
  7. Question: the parser encounters an unfinished quoted field after two apparently good records. May it claim those records were published? Answer: not in this lab. File-level parsing failure returns zero candidates and retains the original input. A different partial-publication policy would have to be designed and disclosed separately.
  8. Run the companion exercise: python3 data_bridge_lab.py prints the canonical fixture outcomes, including the untidy-file accounting and relational examples developed in Chapters 8–9. python3 -m unittest discover -v runs the supplied test suite. All fixtures are synthetic; read the practice appendix before adapting anything to actual data.

7.98 Common wrong ideas#

  1. Wrong: the filename uniquely identifies the data. Right: a name can be reused for changed bytes, and equal bytes can appear under different names.
  2. Wrong: one newline always ends one CSV record. Right: quoted fields can contain line endings; physical lines and logical records differ.
  3. Wrong: a successful parser has validated the business data. Right: decoding, grammar, field rules, relationships and publication are distinct boundaries.
  4. Wrong: quote every CSV cell and spreadsheet execution becomes harmless. Right: quoting handles CSV structure, not all downstream interpretation.
  5. Wrong: JSON cannot contain ambiguous repeated names because dictionaries have unique keys. Right: the representation can repeat a member name; reject or handle it before information is collapsed by a runtime mapping.
  6. Wrong: a binary format automatically makes data smaller, safer and faster. Right: layout, workload, implementation and security properties need separate evaluation.
  7. Wrong: rejected rows may be discarded because they are useless. Right: their diagnostic value and their retention obligations require an explicit policy.
  8. Wrong: the importer must always recover the next row after an error. Right: some errors destroy trustworthy boundaries. Stopping is more accurate than guessing.

7.99 Chapter summary in 20 lines#

  1. A path names a location; it is not a permanent identity for contents or business records.
  2. The same name can lead to changed bytes, and different names can lead to equal bytes.
  3. A digest identifies captured bytes under an algorithm; it does not authenticate an unknown sender by itself.
  4. Bytes, Unicode code points and business records are different units.
  5. Decoding, parsing, domain validation and publication answer different questions.
  6. Framing makes record boundaries explicit through a defined rule.
  7. A partial input must not be mistaken for a complete record.
  8. CSV quoting lets one field contain delimiters and line endings.
  9. A CSV header does not supply all missing meaning, units or types.
  10. Our CSV adapter uses a named, deliberately narrow input profile.
  11. Identifiers remain text instead of being normalised as convenient numbers.
  12. JSON objects, arrays, strings, numbers, Booleans and null have distinct representation roles.
  13. Parseable JSON can still violate every important rule of an order-line contract.
  14. This profile rejects duplicate JSON member names and nonstandard numeric constants.
  15. Binary interpretation requires an explicit layout and byte-order agreement.
  16. A format marker selects a representation; it does not certify trust or truth.
  17. Compatibility must be stated for particular writer and reader contracts.
  18. A normal batch accounts for every parsed row as a candidate, duplicate or rejection.
  19. Conflicting identities and file-level parsing failures are not silently resolved by arrival order.
  20. Candidates still need destination checks and an explicit publication boundary.

Return to contents