Skip to content
KEDBYTE
Site navigation
How Data Works
Chapter
3

Numbers, Text, Time and Missing Values

Part A · Meaning Before Machines|4,972 words|about 22 min read|Volume A

3.0 What this chapter gives you#

A database cannot make a field meaningful if the application has chosen the wrong kind of value. This chapter explains the everyday decisions hidden inside “just store it”: whether something is a count or a label, whether a fraction must be exact, what a character means, which kind of time we are recording, and how to preserve the difference between zero and not known.

You will calculate integer limits, inspect a floating-point surprise, compare two representations of the same visible letter, read an offset timestamp and reason about SQL NULL. The included laboratory repeats the numerical and textual examples with synthetic inputs in an isolated environment.

Code blocks are optional on the first reading. Each one is accompanied by an explanation of what it demonstrates. A printed result is not a promise that every language or database implements the same behaviour.

3.1 Choose a type before choosing a format#

3.1.1 PLAIN — in simple words#

  1. A type describes the kind of value a field can carry and the operations that make sense for it. A number of notebooks is a count. An order identifier is a label. Both may contain digits, but adding one to them has different meanings.

  1. For the count, adding one can mean another notebook. For the label 0017, converting it to the number 17 may throw away a distinction the sending system needs. A field should not become numeric merely because its current examples contain only digits.

  1. Start by asking what the value represents. Then choose the representation and the rules that preserve that meaning.

3.1.2 PLAIN — a picture in your head#

  1. Think of drawers marked “quantities,” “names” and “dates.” Putting a name in the quantity drawer does not make arithmetic on the name useful. Labels help prevent inappropriate operations before they happen.

  1. Where the comparison stops: programming languages differ in how strongly they enforce types and how readily they convert values. The drawer analogy describes the purpose of a type, not a guarantee that every language locks the drawers.

3.1.3 PLAIN — a worked example#

Value Meaning in an example Suitable starting model
12 Counted notebooks on a shelf Non-negative integer count
0017 Identifier supplied as a four-character label Text with a documented identifier rule
7550 Unit price in paise for our fixture Integer amount plus currency and scale
2026-09-01 A calendar date Date, not a made-up midnight instant
09:30 at offset +05:30 A local clock reading with an offset Needs a date to identify the example’s instant
Not yet counted Absence of an observation Missing-value state, not the number zero
  1. The same visible digits can participate in different models. A product named “100” is still a product label. A genuine measured quantity may require fractions rather than a whole-number type. The meaning decides the direction of the design.

3.1.4 PLAIN — what is really happening inside#

  1. A program reads bytes and interprets them as values. It may convert text to a number, validate the range and store a representation chosen by its engine. Each conversion is a place where information can be preserved, rejected or lost.

  1. An import that turns 0017 into 17 cannot reconstruct the original four-character spelling from 17 alone without another rule. If the field’s contract treats the spelling as significant, the conversion has lost information.

  1. Conversely, storing every value as text postpones rather than solves the problem. A later calculation still needs to know whether “12” is a quantity, whether “12.0” is allowed and whether an empty string means missing or invalid.

3.1.5 TECHNICAL — the engineer’s version#

  1. Separate the domain type from the storage type. OrderId, StockCount and UnitPriceInPaise are domain concepts. text and integer are possible storage types. Several domain concepts can share a storage type while having different valid operations.

  1. For our fixture, StockCount permits zero, while an agreed order line’s quantity must be positive. Both are integer-valued, but their domains differ. A type alone may therefore need a constraint. PostgreSQL documents column and table constraints as a way to express additional restrictions. [S02]

  1. Conversions should have explicit failure behaviour. Reject a quantity of two, distinguish an unsupported format from an out-of-range value, and do not truncate a fractional quantity to make it fit a whole-unit contract.

  1. SQLite’s ordinary dynamic typing is a product-specific reminder not to infer all runtime behaviour from a declared type name. The laboratory uses explicit validation for its teaching contract rather than claiming that every engine will enforce identical types. [S21]

3.1.6 WORDS — remember these#

  1. Type: the kind of value; technically, a classification that governs representation and permitted or defined operations.

  1. Domain type: the business meaning of the value; technically, a model-specific type whose rules may exceed the underlying storage type.

  1. Storage type: the engine’s representation category; technically, a supported type or storage class with documented operations and limits.

  1. Lossy conversion: a change that forgets information; technically, a mapping from which the original value cannot always be reconstructed.

3.2 Integers and the units they count#

3.2.1 PLAIN — in simple words#

  1. An integer is a whole number: zero, one, two, and their negative counterparts. Integers are useful for counting indivisible units and for representing amounts in a deliberately chosen smallest unit.

  1. Whole does not mean unlimited. A fixed-size representation has only so many available patterns. You must know both what each step counts and how large the value can become.

  1. For our order prices, one step counts one paisa. For notebook stock, one step counts one notebook. They should not be added to each other merely because both are integers.

3.2.2 PLAIN — a picture in your head#

  1. Imagine a mechanical counter with three decimal wheels. It has room for 000 through 999. A fourth digit cannot appear without a design change or some defined response to exceeding the limit.

  1. Where the comparison stops: computers use binary patterns and different signed representations. Some languages can allocate more space for larger integers; some fixed-size operations reject an excessive value, and some have other specified overflow behaviour. Always name the actual environment.

3.2.3 PLAIN — a worked example#

  1. A bit has two possible values. With eight bits there are 2 × 2 × 2 × 2 × 2 × 2 × 2 × 2 = 256 patterns. An unsigned interpretation can use them for 0 through 255: 256 distinct integers, not 255.

  1. A 32-bit signed two’s-complement interpretation has the range -2,147,483,648 through 2,147,483,647. PostgreSQL’s four-byte integer type has that range and rejects values outside its range. [S10]

  1. If a positive price amount uses that maximum as a count of paise, it represents INR 21,474,836.47. The calculation is 2,147,483,647 / 100. Changing the unit changes the represented amount without changing the bit limit.

  1. For O-1042, the calculation remains comfortably within this example range:

2 × 7550 + 1 × 2000 = 17100 paise
17100 paise / 100 = INR 171.00

3.2.4 PLAIN — what is really happening inside#

  1. The representation stores a pattern interpreted under a rule. “Signed” means the chosen interpretation includes negative values. The unit comes from the surrounding contract, not from a special “money” property in each bit.

  1. Multiplication and addition can produce results larger than their inputs. Choosing a type that holds one unit price is not enough if the same type must also hold the total for many items or a long reporting period.

  1. The review question is therefore about the maximum intermediate result as well as the maximum stored input. In a toy fixture, the answer may be simple. In a real workload, it needs a stated bound and behaviour when that bound is exceeded.

3.2.5 TECHNICAL — the engineer’s version#

  1. For an unsigned n-bit representation, the range is 0 through 2^n - 1. For an n-bit signed two’s-complement representation, it is -2^(n-1) through 2^(n-1) - 1. These expressions count finite bit patterns; they do not claim that all language integer types use a fixed n.

  1. Python’s integers in the laboratory can grow beyond a fixed 32-bit range, subject to available resources. That means a passing arithmetic example in Python does not test whether a PostgreSQL integer column would accept the same enormous result. Our explicit range demonstration calculates the bound rather than claiming to exercise a PostgreSQL server. [S22]

  1. The unit should be carried by naming, documentation and, where practical, domain-specific types. Converting a minor-unit amount to display text should not change the stored value. Combining currencies is outside the fixture’s permitted operations; an amount tagged INR is not interchangeable with the same integer tagged another currency.

  1. Integer minor units are not a universal answer to all price calculations. Proportions and allocations can produce fractions of the chosen unit. The next section introduces the rounding policy needed at that boundary.

3.2.6 WORDS — remember these#

  1. Integer: a whole number; technically, a value in the set of signed whole numbers, represented with implementation-specific limits.

  1. Range: the smallest through largest supported values; technically, the domain bounds of a representation or declared type.

  1. Overflow: exceeding a representation’s capacity; technically, an operation whose mathematical result lies outside the representable range.

  1. Scale: what one stored step represents; here, the fixed relationship between an integer amount and its displayed unit.

3.3 Decimals, fractions and rounding#

3.3.1 PLAIN — in simple words#

  1. Some fractions fit a number system neatly and others do not. In decimal notation, one half is 0.5, but one third continues as 0.333… without ending. Binary arithmetic has a different set of neatly fitting fractions.

  1. This is why a computer may show a tiny surprise when adding familiar decimals with a binary floating-point type. The machine is operating on nearby representable values, not on an infinitely precise version of every number you typed.

  1. The solution is not “never use floating point.” It is to choose a representation and an error or rounding policy suitable for the task.

3.3.2 PLAIN — a picture in your head#

  1. A ruler marked only in whole millimetres cannot directly express every possible distance. A point halfway between two marks needs an approximation or a more detailed measuring method.

  1. Where the comparison stops: floating-point spacing changes with magnitude, unlike the evenly spaced marks on that ruler. The analogy explains limited representation, not the exact layout of floating-point numbers.

3.3.3 PLAIN — a worked example#

  1. In the Python environment used by the supplied laboratory:

print(0.1 + 0.2)
  1. The displayed result is:

0.30000000000000004
  1. Python’s floating-point tutorial explains why common decimal fractions are approximated in binary and why display formatting can hide or reveal that approximation. This is a representation issue, not evidence that addition is randomly unreliable. [S12]

  1. Now construct decimal values directly from decimal text:

from decimal import Decimal

print(Decimal("0.1") + Decimal("0.2"))
  1. The result in the supplied example is 0.3. Constructing the values from strings is deliberate: it avoids first introducing a binary floating-point approximation. Decimal arithmetic still has a precision context and rounding rules; it does not make every fraction finitely representable. [S13]

3.3.4 PLAIN — what is really happening inside#

  1. The chosen numeric representation determines which values are exact and how operations produce rounded results. Formatting determines what part of the result is shown. These are different layers.

  1. Displaying a total with two decimal places does not prove that the underlying calculation used an appropriate representation or rounding policy. It can merely hide a discrepancy.

  1. Consider dividing 100 paise equally among three recipients. Each mathematical share is 33 and one third paise. Under our whole-paisa output rule, three displayed shares of 33 paise add to only 99 paise. One paisa remains to allocate. An explicit rule might give the extra paisa to the first recipient in a defined ordering, producing 34, 33 and 33. That is a selected allocation policy, not an arithmetic law.

3.3.5 TECHNICAL — the engineer’s version#

  1. Distinguish representation error, measurement error and business rounding. Representation error comes from storing or computing with an approximation. Measurement error concerns how closely a measured value reflects the relevant quantity. Business rounding is a deliberate policy for mapping an amount to an allowed unit. Treating all three as “rounding” conceals different causes.

  1. PostgreSQL provides exact numeric/decimal types as well as floating-point types. Declared precision and scale affect what is stored. A field’s type choice should be tied to its required calculations and boundaries rather than copied from an unrelated system. [S10]

  1. A robust calculation specification names the input units, intermediate precision, rounding rule, rounding stage and reconciliation rule. Rounding each line and then summing is not always the same operation as summing exact intermediates and rounding once.

  1. The laboratory’s allocate_minor_units(100, 3) example uses integer division and remainder to produce [34, 33, 33]. It verifies that the shares sum to 100 and differ by at most one. The rule selects recipients by list position; it does not claim to be the correct policy for every allocation problem.

3.3.6 WORDS — remember these#

  1. Floating point: a number representation with a movable scale; technically, a finite significand and exponent representation with defined rounding behaviour.

  1. Decimal arithmetic: arithmetic based on decimal representation; technically, operations over decimal values under a precision and rounding context.

  1. Quantisation: fitting values to allowed steps; technically, mapping values to a discrete set of representable or permitted outputs.

  1. Rounding policy: the chosen way to handle non-exact results; technically, the rule and stage used to select an allowed output value.

3.4 Text, Unicode and bytes#

3.4.1 PLAIN — in simple words#

  1. A computer stores a representation of text, not tiny printed letters. A text system needs an agreement about which abstract characters are represented and how to turn that sequence into bytes.

  1. Unicode provides a shared character framework. UTF-8 is one way to encode Unicode text as bytes. They are related concepts, not two names for the same layer. A Unicode scalar value takes from one to four bytes in UTF-8. [S14]

  1. The number of bytes is therefore not generally the number of visible characters. That distinction matters when storing names, limiting input length or cutting a string into pieces.

3.4.2 PLAIN — a picture in your head#

  1. Imagine a catalogue assigning a number to each symbol, and a packing rule for sending those numbers in small envelopes. The catalogue is one layer; the packing rule is another. Some symbols need more envelope space than others.

  1. Where the comparison stops: a visible character can be formed from several Unicode code points. The reader’s “one character” is not always one catalogue entry, and rendering may combine or shape entries according to the writing system.

3.4.3 PLAIN — a worked example#

  1. The Latin letter A is one code point, U+0041. The rupee sign ₹ is one code point, U+20B9. In UTF-8 their byte sequences differ in length:

Text Code point UTF-8 bytes in hexadecimal Byte count
A U+0041 41 1
₹ U+20B9 E2 82 B9 3
  1. The supplied laboratory checks these sequences directly. Hexadecimal, or base 16, is just a compact way to write the byte values; it does not change the text.

  1. Now compare the visible letter “é.” It can be represented by U+00E9, or by U+0065 followed by U+0301, a combining acute accent. The first sequence has one code point and the second has two. Unicode defines canonical equivalence and normalisation forms for cases such as this. [S15]

3.4.4 PLAIN — what is really happening inside#

  1. A decoder turns valid encoded bytes into a text representation. A renderer draws the resulting text using fonts and shaping rules. These are not the same operation: being able to decode a character does not guarantee that a particular font draws it well.

  1. A program comparing code-point sequences can find our two forms of “é” unequal even when they look alike. Applying an appropriate normalisation policy can make canonically equivalent forms comparable. It does not translate a name, decide whether two people are the same, or make all similar-looking characters identical.

  1. When a user asks for “20 characters,” the product should decide what that limit is for. A storage byte limit, a code-point count and a user-visible editing limit are different requirements. Unicode’s text-segmentation guidance defines grapheme-cluster boundaries relevant to user-perceived characters. [S16]

Figure 3.1. The rupee sign passes through three different descriptions: visible symbol, Unicode code point and UTF-8 byte sequence.

Figure 3.1. The rupee sign passes through three different descriptions: visible symbol, Unicode code point and UTF-8 byte sequence.

3.4.5 TECHNICAL — the engineer’s version#

  1. The laboratory uses Python strings and UTF-8 encoding. len("₹") is 1 in that environment, while len("₹".encode("utf-8")) is 3. Neither result alone defines a universal user-interface character-count policy.

  1. For the two acute-e forms, Python reports code-point counts of one and two. Their UTF-8 lengths are two and three bytes. Normalising each with NFC makes the two resulting strings equal in the demonstrated case. The equality before normalisation is false; the equality after it is true. [S15]

  1. Normalisation is not the same operation as case folding, language-sensitive sorting or identity matching. A system that needs all of them should specify their order and purpose rather than hide them behind a vague “clean text” step.

  1. At an import boundary, invalid encoding needs a deliberate policy. Silent replacement can make a file displayable while losing the original distinction that an identifier depended on. The laboratory uses strict decoding behaviour for its bounded example and verifies that an invalid UTF-8 byte sequence is rejected.

3.4.6 WORDS — remember these#

  1. Code point: an assigned position in the Unicode code space; technically, a numeric value used in the character model, not a guarantee of one visible glyph.

  1. Encoding: the mapping into bytes; technically, a defined representation of a character sequence for storage or interchange.

  1. Grapheme cluster: a unit closer to what a reader perceives as a character; technically, a text-segmentation unit defined by applicable Unicode rules and possible tailoring.

  1. Normalisation: a specified way to standardise equivalent text sequences; technically, conversion to a Unicode normalisation form, not general identity verification.

3.5 Dates, instants and durations#

3.5.1 PLAIN — in simple words#

  1. Not every value involving time is the same kind of thing. A birthday is a calendar date. A completed sale has an occurrence instant. “The shop opens at 09:00” is a local schedule. “The upload took seven minutes” is a duration.

  1. Treating all four as an arbitrary text field creates ambiguity. Treating all four as midnight in a time zone can create a different ambiguity: a birthday was not necessarily meant to identify an instant at all.

  1. Choose the temporal meaning before choosing the storage representation.

3.5.2 PLAIN — a picture in your head#

  1. A calendar, a wall clock and a stopwatch answer different questions. A calendar tells you which date. A wall clock gives a local time of day. A stopwatch measures elapsed time.

  1. Where the comparison stops: computers can combine these concepts and convert between representations, but a conversion cannot recover context that was never supplied. A bare 09:30 does not identify one instant in history.

3.5.3 PLAIN — a worked example#

  1. Read this timestamp from left to right:

2026-09-01T09:30:00+05:30
  1. It states the date 1 September 2026, local clock time 09:30:00, and a numeric offset five hours and thirty minutes ahead of UTC. For this representation, subtracting that offset gives:

2026-09-01T04:00:00Z
  1. The two strings describe the same instant. The Z form uses UTC. The syntax is within the timestamp format described by RFC 3339. [S09]

  1. Now compare occurrence at 09:30 and recording at 09:37 with the same date and offset. Under the example’s comparable-clock assumption, the interval is 420 seconds. The laboratory verifies the conversion and subtraction for these inputs.

3.5.4 PLAIN — what is really happening inside#

  1. The program interprets the date, clock fields and offset, then can compare the resulting instants. A named time zone adds rules associated with a place; a numeric offset only describes an offset for the representation in front of you.

  1. For a historical event, an instant representation may be sufficient for ordering and duration calculations, with additional source context retained when needed. For a recurring local opening time, “09:00 in the shop’s local time zone” is a different requirement from “the same UTC time forever.”

  1. Two source clocks can disagree. Extra decimal places on a timestamp do not establish that the clock was accurate to that precision. A latency calculation between independent systems therefore needs assumptions about the clocks or another measurement method.

3.5.5 TECHNICAL — the engineer’s version#

  1. PostgreSQL distinguishes date, timestamp without time zone and timestamp with time zone. For the latter, input is normalised to an instant and output is displayed using the session’s time-zone setting; the original named zone is not retained in that value. The type’s name should not be read as a promise to preserve all original zone context. [S11]

  1. Use a date for a date-only concept. Use an instant-capable representation for an occurrence when an actual instant is required. Preserve a named zone separately when the application needs that original context or future local scheduling rules. A local recurring schedule is not fully represented by one past offset.

  1. Duration arithmetic also has a domain boundary. “Add 24 elapsed hours” and “same local time tomorrow” need not be interchangeable in every zone and calendar situation. Our 420-second example avoids those complications by using two explicit offsets on the same date.

  1. The Python laboratory checks the stated examples using its standard datetime facilities. It is not a complete RFC 3339 conformance test, a time-zone database certification or a clock-synchronisation experiment.

3.5.6 WORDS — remember these#

  1. Date: a calendar position; technically, a calendar day without necessarily specifying a time of day or an instant.

  1. Instant: one point on a timeline; technically, a temporal position distinguishable from its local display representation.

  1. UTC offset: the difference from UTC in a representation; technically, a signed hours-and-minutes offset, not a complete named-zone rule set.

  1. Duration: elapsed amount of time; technically, a measured or specified interval whose units and arithmetic semantics must be stated.

3.6 Missing is not zero#

3.6.1 PLAIN — in simple words#

  1. Zero is a value. “We have not counted yet” is a lack of a value. Replacing the second with the first changes the claim from “not known” to “known to be none.”

  1. An empty text value, a field that was omitted, an explicit missing marker and the word “unknown” are also different representations. A system needs a policy for them rather than assuming they all mean the same thing.

  1. The policy should match the purpose. Sometimes a missing value should reject a record. Sometimes it should be preserved so the record remains honest about what is not known.

3.6.2 PLAIN — a picture in your head#

  1. Imagine four envelopes labelled “refund amount.” One contains a slip saying zero. One contains a slip saying 100. One is empty. One has a note saying the envelope was never requested.

  1. You cannot add the empty envelope as zero without making an extra assumption. Where the comparison stops: SQL NULL is a specific database marker with defined expression behaviour. It does not automatically preserve every possible reason for absence. A separate status may be needed when the reason matters.

3.6.3 PLAIN — a worked example#

  1. Suppose four fictional observations contain values 100, 200, missing and 0. The sum of the known values is 300. There are four rows, but only three known values. The mean of the known values is 300 / 3 = 100.

  1. Replacing missing with zero would produce an average of 300 / 4 = 75. That is not another spelling of the same answer. It is a result under a different assumption.

  1. In the laboratory’s SQLite example, the query is:

SELECT
    COUNT(*) AS row_count,
    COUNT(value) AS known_count,
    AVG(value) AS mean_known
FROM observations;
  1. For the supplied fixture, the result is (4, 3, 100.0). The same distinction between counting rows and counting non-null values is documented for PostgreSQL’s aggregate functions. [S19]

3.6.4 PLAIN — what is really happening inside#

  1. In ordinary SQL comparisons, a NULL operand generally produces an unknown result rather than a true equality or inequality. Asking value = NULL is not the way to find missing values; use value IS NULL. PostgreSQL documents this distinction explicitly. [S17]

  1. Think of an unknown sealed-box quantity. You cannot say that it equals five, but you also cannot conclude that it does not equal five. The comparison is unknown. A filter keeps rows whose condition is true, so those unknown comparison results do not behave like true matches.

  1. This also explains why changing value > 0 to NOT (value > 0) does not automatically include missing values. Negating unknown still does not establish a known true condition. The missing case needs to be handled according to the question being asked. [S18]

3.6.5 TECHNICAL — the engineer’s version#

  1. The relevant three-valued logic has true, false and unknown. Here are selected operations; “unknown” is shown as a word rather than confused with a stored text value.

Expression Result
NOT unknown unknown
true AND unknown unknown
false AND unknown false
true OR unknown true
false OR unknown unknown
  1. These operations are consistent with PostgreSQL’s documented logical tables. They do not imply that evaluation order is a safe substitute for controlling side effects in expressions. [S18]

  1. A check constraint and a missing-value rule are also distinct. In PostgreSQL, CHECK (quantity > 0) alone does not reject NULL: a check is satisfied by true or null. Add NOT NULL when absence is forbidden. [S02]

  1. JSON has a null literal, while an absent object member is structurally different. Neither representation carries a universal business explanation for missingness. Our line validator requires all its fields; it does not use null as a substitute for an agreed quantity. [S07]

  1. For reporting, publish the denominator. “Mean among three known observations, one missing” is more informative than showing 100 with no population or missingness context. The arithmetic cannot decide whether the missing case should be excluded, estimated or treated as an error; that is part of the analytical contract.

3.6.6 WORDS — remember these#

  1. NULL: a specific SQL missing/unknown marker; technically, a marker with defined behaviour in comparisons, logic, constraints and aggregates.

  1. Empty string: text containing no characters; technically, a zero-length string, distinct from SQL NULL in the engines discussed here.

  1. Denominator: what the division is over; technically, the quantity that defines the base of a ratio or average.

  1. Missingness policy: what absence means and how it is handled; technically, the domain rules for rejecting, preserving, classifying or analysing unavailable values.

3.97 Practice, then inspect the evidence#

Exercise 1: the identifier. A source contract defines 0017 as a four-character order label. An import turns it into the integer 17. What was lost?

Worked answer. The original representation’s leading zeros were lost. Whether that changes business identity depends on the source contract, but this contract explicitly required the four-character label. The receiver should not perform that conversion silently.

Exercise 2: the allocation. Split 100 paise among three recipients using the book’s rule: integer base shares, then give one extra paisa to the earliest recipients until the remainder is exhausted.

Worked answer. Integer division gives a base share of 33 and a remainder of one. The shares are 34, 33 and 33. They sum to 100. A different recipient ordering may allocate the extra paisa differently, so the ordering must be part of the policy.

Exercise 3: the timestamp. Do 2026-09-01T09:30:00+05:30 and 2026-09-01T04:00:00Z identify different instants?

Worked answer. No. The first is five and a half hours ahead of UTC in its local representation. Subtracting the offset yields the second.

Exercise 4: the average. Four rows contain 100, 200, NULL and 0. Explain both 100 and 75 without calling them equivalent.

Worked answer. The mean among three known values is 100. Replacing NULL with zero produces a mean of 75 over four values. The latter requires an additional assumption that the missing value should count as zero. The source rows alone do not justify it.

Exercise 5: the text. Two strings look like “é,” but one has one code point and the other two. Does this prove a corruption?

Worked answer. No. They may be the canonically equivalent forms demonstrated in the chapter. Inspect their code points and apply the relevant text policy rather than judging identity only by appearance.

3.98 Common wrong ideas#

“Digits mean the field should be numeric.” Identifiers can contain digits while remaining labels with significant spelling and scope.

“Integers cannot have numerical limits.” Fixed-size representations have finite ranges. A language with growing integers does not remove a database column’s limit.

“Showing two decimals makes a calculation correct.” Display formatting does not define representation, intermediate precision or rounding policy.

“Decimal arithmetic makes every fraction exact.” Some fractions still require rounding under a finite precision context.

“One visible character is one byte.” Text has distinct byte, code-point and user-perceived segmentation layers.

“Store every time as UTC and all temporal questions disappear.” Dates, instants, local schedules and original zone context remain different requirements.

“Unknown means zero or false.” Missingness and three-valued logic must be handled according to the actual domain and query.

3.99 Chapter summary in 20 lines#

  1. Choose a value’s meaning before choosing its representation.
  2. An identifier containing digits is not automatically a quantity.
  3. Domain types can be more specific than storage types.
  4. A conversion can lose information while appearing convenient.
  5. Integers count whole steps under a stated unit.
  6. Fixed-size integer representations have finite ranges.
  7. Intermediate calculations need range analysis too.
  8. Minor-unit amounts still need currency and scale context.
  9. Binary floating point cannot represent every familiar decimal exactly.
  10. Decimal arithmetic also has precision and rounding boundaries.
  11. Formatting is not a substitute for a calculation policy.
  12. Allocations need a rule for indivisible remainders.
  13. Unicode and UTF-8 describe different layers of text representation.
  14. Bytes, code points and grapheme clusters are different counting units.
  15. Normalisation is not general identity verification.
  16. A calendar date is not the same concept as an instant.
  17. A numeric UTC offset is not a complete named-zone rule set.
  18. Missing, zero and empty text must not be silently equated.
  19. SQL NULL changes comparison, logical and aggregate behaviour.
  20. Clear types preserve meaning before a system ever becomes large.

Return to contents