Skip to content
KEDBYTE
Site navigation
How Data Works
Chapter
2

What a Record Really Means

Part A · Meaning Before Machines|4,681 words|about 20 min read|Volume A

2.0 What this chapter gives you#

This chapter teaches you to look at a record and ask the questions that a screen full of neat columns can hide. What is being claimed? What does one row stand for? Which thing does an identifier refer to? Who supplied the information, and which moment does its time describe?

You will turn an ambiguous line of text into an explicit teaching record. You will also learn why keeping more fields is not the same as keeping better evidence, and why a correction needs a meaning as well as a new value.

The running example remains Mira’s Corner. O-1042 is the order for two notebooks and one pen, totalling INR 171.00 under the simplified pricing rules established in Chapter 1. We will not quietly change those assumptions halfway through an example.

2.1 The claim inside a record#

2.1.1 PLAIN — in simple words#

  1. Consider this line: 1042, notebook, 2, 75.50, 09:30. It looks informative, but important parts of its meaning are missing. Is 1042 an order, a customer or a delivery? Does two mean two notebooks or two boxes? Is 75.50 the price of one item or the amount for the whole line? Which day does 09:30 belong to?

  1. A record becomes useful when the reader can recover the intended claim without guessing. A useful sentence might be: “Order O-1042, line 1, records an agreed quantity of two individual notebooks at INR 75.50 per notebook.”

  1. That sentence is already doing design work. It separates the order, the line, the item, the quantity, the unit and the price basis.

2.1.2 PLAIN — a picture in your head#

  1. Imagine receiving a parcel with only the number “2” on its label. It might mean two items, a second attempt, a shelf number or a delivery priority. A better label does not change the parcel; it tells the next person how to treat it.

  1. Where the comparison stops: a data field’s meaning can depend on rules stored somewhere else, such as a schema or data dictionary. You do not need to repeat a long explanation in every row, but the explanation must be available and connected to the row’s version.

2.1.3 PLAIN — a worked example#

  1. We can make the first line of O-1042 explicit:

Field Example value Meaning in this book
order_id O-1042 The shop’s identifier for this order
line_no 1 This line’s number within the order
product_id P-NOTE The notebook product in our catalogue
quantity 2 Two individual saleable units
unit_price_minor 7550 Agreed price per unit in paise
currency INR The currency attached to the amount
  1. For this case study, one rupee is represented by 100 paise. The line amount is 2 × 7550 = 15100 paise, displayed as INR 151.00. Storing 7550 without stating “paise per unit” would not preserve the full claim.

  1. The record says what was agreed for the order line. It does not, on its own, say that two notebooks were handed over or that a payment completed. Those require their own definitions and evidence.

2.1.4 PLAIN — what is really happening inside#

  1. The producer and consumer of a record share a contract. That contract says how to read the fields, which values are allowed and what business meaning the combination has.

  1. A field name can help, but names are not a complete contract. amount is vague. unit_price_minor is better, but it still needs a currency, a definition of the minor unit and a rule about when the price was agreed.

  1. The data dictionary is where we gather these meanings. W3C’s data guidance distinguishes structural metadata from descriptive information that helps people understand and use data. In our book, the dictionary supplies both where appropriate. [S06]

2.1.5 TECHNICAL — the engineer’s version#

  1. Model the example as a predicate: AgreedOrderLine(order_id, line_no, product_id, quantity, unit_price_minor, currency). Read that as a statement with named arguments, not as runnable code.

  1. The predicate’s interpretation contains facts the type system alone cannot express. For example, quantity counts saleable individual units in this exercise. A later product sold by mass would need a different quantity-and-unit policy; reusing the integer field without revisiting its meaning would be a modelling error.

  1. Separate the syntactic contract from the semantic contract. Syntax tells a parser where a field begins and ends. Semantics tells an application what the value means and which operations are meaningful. A JSON number can parse correctly while being interpreted in the wrong unit. [S07]

  1. For each field, document its domain, unit or scale, missing-value policy, source, mutability and version boundary. These are proposed design requirements for our case study, not properties acquired merely by choosing JSON or SQL.

2.1.6 WORDS — remember these#

  1. Field: one named part of a record; technically, an attribute position with an agreed interpretation.

  1. Domain: the allowed kind of value; technically, the set of permitted values and associated meaning for an attribute.

  1. Syntax: the shape of a representation; technically, the grammar used to recognise valid constructions.

  1. Semantics: what a representation means; technically, the interpretation and permitted consequences assigned to it.

2.2 Fields, values and the contract around them#

2.2.1 PLAIN — in simple words#

  1. A label and a value travel together, but they are not the whole story. The same text can mean different things under different rules. A date written as 03/04/2026 may be read as 3 April or 4 March. Neither reader has necessarily made an arithmetic mistake; they may be using different conventions.

  1. A safe exchange tells the receiver which convention applies. When that information is missing, “pick the interpretation that looks likely” is not a reliable general rule.

  1. The goal is to make the contract explicit before a guess turns into a stored fact.

2.2.2 PLAIN — a picture in your head#

  1. Picture a board game with pieces but no instructions. You can see the pieces clearly and still have no idea whether moving diagonally is allowed. A schema describes much of the board and the pieces; the full application contract also describes the permitted play.

  1. Where the comparison stops: a schema may itself encode some rules, while others live in application logic or procedures. There is no universal split that makes every important rule belong in exactly one place.

2.2.3 PLAIN — a worked example#

  1. Here is the first order line as a JSON object. JSON is a text format for exchanging structured values. The braces enclose an object; each quoted name is followed by its value. The structure below is an original teaching example, not a complete production message specification. [S07]

{
  "record_type": "agreed_order_line",
  "schema_version": 1,
  "order_id": "O-1042",
  "line_no": 1,
  "product_id": "P-NOTE",
  "quantity": 2,
  "unit_price_minor": 7550,
  "currency": "INR"
}
  1. Suppose another program sends the same field names but uses 75.50 to mean rupees. It has not followed our contract. Silently accepting the value and treating it as paise would produce a different price.

  1. The receiver should reject or explicitly transform that message under a documented input contract. It should not decide the unit by whether a number contains a decimal point.

2.2.4 PLAIN — what is really happening inside#

  1. A receiver first recognises the representation, then checks its structure and values. In our example, it expects a supported record type and schema version. It checks required fields and rejects values outside the contract.

  1. A version number is useful only when it points to an actual definition. Writing schema_version: 1 does not make an undocumented format self-explanatory. Nor should a receiver automatically accept version 2 because it understood version 1.

  1. When a producer changes a field’s meaning, readers need a compatibility plan. Changing an amount from paise to rupees under the same name and version would break that agreement even if every message still parsed.

2.2.5 TECHNICAL — the engineer’s version#

  1. Define four outcomes for an incoming teaching record: accepted without transformation; accepted through a documented transformation; rejected with a precise reason; or held for review because its meaning is unresolved. The review path must not quietly feed uncertain values into ordinary totals.

  1. For quantity, our contract requires a positive integer, not just something a language can coerce to one. The included Python validator deliberately rejects booleans as quantities, even though Python’s Boolean type has an integer relationship. This is our validator’s domain policy, demonstrated by a test. [S22]

  1. Unknown fields also need a policy. An extensible analytics feed may preserve them. A tightly bounded command may reject them. Our small validator rejects unexpected field names so a misspelling such as unit_prcie_minor does not silently disappear.

  1. The contract should also define how invalid inputs are reported. “Bad data” is not as useful as “quantity must be a positive integer.” The latter helps the sender repair the actual defect without exposing unrelated records.

2.2.6 WORDS — remember these#

  1. Schema: the declared structure; technically, a definition of fields, types, relationships and applicable structural rules.

  1. Schema version: which contract applies; technically, an identifier for a defined revision of the schema and its interpretation.

  1. Parser: a shape reader; technically, software that recognises a representation according to its grammar.

  1. Coercion: automatic conversion; technically, a transformation between types that may preserve or change intended meaning.

2.3 One row means one what?#

2.3.1 PLAIN — in simple words#

  1. Before counting rows, finish this sentence: “One row in this dataset represents one ___.” This is the row’s grain.

  1. An order row and an order-line row are not interchangeable. A product row describes a catalogue item. A stock-count row describes an observation of an item at a location and time. All can mention the same product without representing the same kind of fact.

  1. When different grains are mixed, a table can look tidy while its totals become misleading.

2.3.2 PLAIN — a picture in your head#

  1. Imagine counting a classroom from two lists. The attendance list has one line per pupil. The timetable has one line per lesson attended by a pupil. Counting timetable lines does not count pupils unless you deliberately remove the repetition.

  1. Where the comparison stops: repetition is not always a defect. The timetable genuinely needs several records for the same pupil because each record stands for a different lesson. The design question is whether repetition matches the intended grain.

2.3.3 PLAIN — a worked example#

  1. Our order-line dataset contains these four rows:

Order Line Product Units Line amount, paise
O-1042 1 P-NOTE 2 15100
O-1042 2 P-PEN 1 2000
O-1043 1 P-NOTE 1 7550
O-1043 2 P-PEN 2 4000
  1. There are four rows because the grain is one agreed order line. There are two orders because only O-1042 and O-1043 appear as order identifiers. There are six units because 2 + 1 + 1 + 2 = 6.

  1. Now imagine that each row also repeats its complete order total. Adding that repeated column would count each two-line order twice: 17100 + 17100 + 11550 + 11550 = 57300 paise. The correct combined order value is 28650 paise.

  1. The stored numbers need not be individually false for the calculation to be wrong. The error comes from aggregating at the wrong grain.

2.3.4 PLAIN — what is really happening inside#

  1. A query works with the rows it receives. If a combination operation produces several rows for one order, a later sum sees several rows. The engine does not automatically know that a repeated order total should count only once.

  1. The repair is not a universal instruction to remove duplicates. Two different orders can legitimately have the same total. Deduplicating the amount itself could remove a real order from the calculation.

  1. Instead, decide which identity defines the things being counted and at which grain each amount is meaningful. We will use separate order-level and line-level structures when the relational model is built in later chapters.

2.3.5 TECHNICAL — the engineer’s version#

  1. The sample line identity is (order_id, line_no). The quantity and agreed unit price describe that line. An order total describes the order, so its natural aggregation boundary is different.

  1. A functional dependency says that one set of attributes determines another under the data model. For our case study, the order-line identifier determines the values for that accepted line version. It does not imply that product identity alone determines an agreed historical price.

  1. A product’s current catalogue price may change tomorrow. Replacing yesterday’s agreed line price with today’s catalogue price would change the meaning of the historical order. Whether to preserve a copied price, reference a versioned price record or use another design is a later modelling decision; the requirement is to retain the correct historical meaning.

  1. A basic test can expose the grain mistake: compute totals from the four line amounts and compare them with totals from the two order headers. Agreement is a useful check in this fixture, not a guarantee that every other report is correct.

2.3.6 WORDS — remember these#

  1. Grain: what one row represents; technically, the unit of observation or business fact represented at the dataset’s chosen level.

  1. Aggregation: combining several records into an answer; technically, a calculation over a defined group of inputs.

  1. Functional dependency: one part determines another; technically, a constraint under which equal determining attributes imply equal dependent attributes.

  1. Historical value: a value tied to an earlier event or agreement; technically, information whose interpretation must not be replaced by an unrelated current-state value.

2.4 Identity without guesswork#

2.4.1 PLAIN — in simple words#

  1. An identifier lets us refer to a particular thing without redescribing it every time. It answers “which one?” It does not automatically answer “is this genuine?” or “is this the same real-world person?”

  1. Two orders can contain the same items and amounts and still be different orders. One order can be delivered twice as a message and still represent only one order. Matching visible values does not settle the distinction.

  1. We need to define the thing being identified and the scope within which its identifier is unique.

2.4.2 PLAIN — a picture in your head#

  1. A seat number is useful inside a particular theatre and performance. “Seat 12” alone may not identify where you should sit. You may also need the row, venue and performance.

  1. Where the comparison stops: a software identifier can be designed to have a broader scope. Even then, a broad identifier does not authenticate the person presenting it or prove that two separately created records describe different real-world entities.

2.4.3 PLAIN — a worked example#

  1. At Mira’s single shop, O-1042 identifies our first order. Later a second branch starts its own sequence and also creates O-1042. If both identifiers were only promised to be unique within their own branch, the collision is a failure of the combined interpretation, not necessarily a broken local sequence.

  1. One solution is to identify the order by the pair (branch_id, order_id). Another is to introduce an identifier allocated across the whole system. We will compare designs later; for now, either requires an explicit scope.

  1. For the current one-shop fixture, line 1 of O-1042 and line 1 of O-1043 are different lines. line_no alone is not their full identity.

2.4.4 PLAIN — what is really happening inside#

  1. The system can enforce a uniqueness rule over a chosen set of stored values. That rule is valuable, but its meaning depends on the columns chosen.

  1. If every incoming message receives a fresh database identifier, the database may contain perfectly unique rows for the same repeated business event. The uniqueness rule succeeded at row identity while the business deduplication rule was never supplied.

  1. Conversely, collapsing all records with the same product, quantity and amount could merge two genuine purchases. Duplicate detection requires a definition of “same event,” not simply a desire for fewer rows.

2.4.5 TECHNICAL — the engineer’s version#

  1. Distinguish row identity, business identity and delivery identity. Row identity selects a stored record. Business identity selects the real or administrative occurrence represented. Delivery identity selects a particular transmission or attempt. Our design may intentionally map several deliveries to one business event.

  1. In PostgreSQL, a primary key can span several columns and requires unique, non-null values for the key. That supports the composite line identity used in our model. It does not infer whether two different keys refer to the same underlying purchase. [S02]

  1. A deduplication decision also needs a conflict path. Two deliveries claiming the same business identifier but carrying different quantities should not automatically be treated as identical replays. The system must apply a documented rule or retain the conflict for review.

  1. For now, the safe mental model is that an identifier is a reference under a contract. Its format, scope, allocation and lifecycle all belong in that contract.

2.4.6 WORDS — remember these#

  1. Identifier: a reference to a particular thing; technically, a value or tuple interpreted within a defined namespace and lifecycle.

  1. Composite key: an identifier made from several parts; technically, a key consisting of more than one attribute.

  1. Namespace: where a name has its meaning; technically, the scope within which identifiers are interpreted and uniqueness is defined.

  1. Deduplication: recognising repeated representations of the same thing; technically, a policy-driven process whose identity assumptions and conflict handling must be explicit.

2.5 Source, event time and recording time#

2.5.1 PLAIN — in simple words#

  1. A record has a history. Somebody or something produced it, possibly from earlier records. Knowing that history can help you assess what it supports and how to reproduce it.

  1. Time is part of that context. The time an event happened may differ from the time its record reached the system. An offline till can record a sale at the counter and upload it later.

  1. Calling both moments simply time makes later questions unnecessarily difficult.

2.5.2 PLAIN — a picture in your head#

  1. A letter has a date written inside and a date stamped when it reaches an office. The first describes what the sender says about timing; the second describes an observed step in delivery. You may need both to answer a question fairly.

  1. Where the comparison stops: neither timestamp guarantees a perfectly accurate clock. The date inside may also describe a plan rather than an occurrence. The field’s meaning and the clock’s limitations still matter.

2.5.3 PLAIN — a worked example#

  1. For our fictional order, assume the till records an occurrence time of 2026-09-01T09:30:00+05:30. The receiving system records arrival at 2026-09-01T09:37:00+05:30.

  1. Under the example’s shared, accurate-clock assumption, the difference is seven minutes. A report asking when the order occurred uses the first timestamp. A report asking how long the feed was delayed compares the two.

  1. A third report might ask what the central office knew at 09:35. At that point the order had occurred in the story but had not yet arrived in the central dataset. Rebuilding that historical view requires more than the occurrence time alone.

2.5.4 PLAIN — what is really happening inside#

  1. The receiver can preserve an event-time field supplied by the source and create a separate recording-time field itself. It can also record the source identifier and import batch so a suspicious record can be traced.

  1. This is a practical form of provenance: information about how a record or result came to exist. W3C’s PROV model distinguishes entities, activities and agents, and represents relationships such as derivation. Our small shop metadata borrows that distinction without claiming to implement the complete PROV standard. [S05]

  1. A source label is a starting point for investigation, not a certificate of truth. A wrongly configured till can send consistently wrong times with a perfectly consistent source identifier.

Figure 2.1. One occurrence, two relevant times. The timeline assumes the clocks are accurate and comparable; the record itself does not prove that assumption.

Figure 2.1. One occurrence, two relevant times. The timeline assumes the clocks are accurate and comparable; the record itself does not prove that assumption.

2.5.5 TECHNICAL — the engineer’s version#

  1. Use explicit field semantics such as occurred_at, recorded_at and source_system. Their exact names are a case-study convention. recorded_at should state which receiving boundary it describes: initial receipt, validation completion or database commit are not necessarily the same moment.

  1. Distinguishing event time from recording time introduces the need for multiple temporal perspectives. Two timestamp columns alone do not implement a full bitemporal database or preserve every revision needed for historical queries.

  1. The numeric offset in our timestamp identifies a relationship to UTC for that representation. RFC 3339 defines a widely used Internet timestamp syntax with such offsets. It does not make a source clock accurate, and a numeric offset is not the same thing as a named time zone with rules. [S09]

  1. For a derived report, provenance should connect the output to its input set and transformation version. An output label such as “sales report” is not enough to reconstruct which records were included or which rule produced the result.

2.5.6 WORDS — remember these#

  1. Event time: when the described occurrence happened; technically, the timestamp assigned to the event under its source’s semantics.

  1. Recording time: when a defined system boundary recorded the information; technically, a timestamp whose acquisition point must be specified.

  1. Provenance: the history behind a record; technically, information about entities, activities, agents and derivations involved in producing it.

  1. Lineage: how one result came from other data; technically, the dependency path between inputs, transformations and outputs.

2.6 Corrections and the limits of evidence#

2.6.1 PLAIN — in simple words#

  1. Suppose a quantity was entered incorrectly. Replacing the value can make the current record more useful, but it may erase the explanation for why yesterday’s report differed from today’s.

  1. A correction should say what claim changed and why the change was accepted. It should not pretend that the corrected record was always what the system knew.

  1. This does not mean every system must retain every old value forever. Keeping history has a purpose and a lifecycle. The requirement is to choose the policy deliberately and describe its limits honestly.

2.6.2 PLAIN — a picture in your head#

  1. On a paper form, a correction might cross out one value, write the replacement and add a note. The visible change helps a later reader understand the discrepancy.

  1. Where the comparison stops: digital records can have access controls, version histories and multiple derived copies. Preserving an old value in one place does not ensure that every report or export has been corrected, and indefinite preservation can itself be inappropriate.

2.6.3 PLAIN — a worked example#

  1. This is a separate correction exercise, not a change to the canonical O-1042 fixture. Imagine a draft order D-2001 with quantity three, later verified as quantity two. At a unit price of 7550 paise, the draft line amount was 22650 paise and the corrected amount is 15100 paise.

  1. The difference is 7550 paise. If a report already included the draft amount, simply changing the current row does not tell us whether that report refreshed. We need to identify which outputs depend on the changed record.

  1. A useful correction note might record the affected draft, old and new quantities, the reason “entry corrected against the retained source slip,” the approving role, and the time of the decision. The note describes evidence used in this hypothetical example; it does not prove the source slip was itself accurate.

2.6.4 PLAIN — what is really happening inside#

  1. The system applies a chosen correction policy. It might update a draft in place while keeping an authorised revision history. It might append a correcting event. It might prevent direct editing after a particular business transition and require a separate reversal workflow.

  1. Which policy is appropriate depends on what the record means. Correcting a spelling mistake in a draft label is not necessarily the same operation as reversing an accepted commercial obligation.

  1. Derived outputs need their own response: refresh, correction notice, versioned replacement or an explicit statement that a previous export cannot be recalled. A deletion from the current screen does not tell us what happened to every retained copy.

2.6.5 TECHNICAL — the engineer’s version#

  1. Separate current-state correctness, historical explainability and retention policy. A design may improve one while damaging another. For example, overwriting a value simplifies the current row but can destroy the information needed to explain an earlier output.

  1. The case-study correction model will record the identity of the subject, the scope of the change, the asserted reason, the permitted actor and the affected revision. It will avoid treating an unverified reason code as established causation.

  1. Provenance relationships can connect a corrected output with its predecessor and the activity that produced it. The PROV model provides concepts for such relationships; it does not dictate our business authorisation or retention policy. [S05]

  1. The acceptance test must match the claim. A test that verifies the latest row says two does not also prove that historical reports are reproducible, permissions are correct or all external copies have changed. Those require separate observations.

2.6.6 WORDS — remember these#

  1. Revision: a particular state after a change; technically, an identifiable version of a record or artefact.

  1. Correction: a change to repair a stated defect; technically, an authorised transition whose scope, evidence and consequences should be recorded.

  1. Audit trail: information used to explain actions; technically, a record of relevant operations and context, with its own integrity and retention requirements.

  1. Retention: how long and why information is kept; technically, a policy governing continued storage and eventual disposition for defined record classes.

2.97 A record-design exercise#

You receive this export from a fictional third-party till:

id,item,qty,amount,time
1042,Notebook,2,151,09:30
1042,Pen,1,20,09:30

Task. Write down what can be read directly and what remains unknown. Do not fill the gaps from the familiar numbers in our running story; in this exercise the sender has not supplied that context.

Worked answer. The file contains two data rows with the same id, different item labels, quantities two and one, amounts 151 and 20, and matching time strings. That describes the representation, not its full business interpretation.

We do not yet know what id identifies, whether the amount is per unit or per line, which currency and scale apply, what units the quantities count, which date and offset accompany the time, or whether the records describe requests, accepted orders or completed handovers. We also lack a source record identifier and a duplicate-delivery policy.

A suitable repair is to obtain the sender’s contract, then transform the records under that contract while retaining appropriate import provenance and rejected-input evidence. It is not to guess INR because the values happen to resemble our previous example.

Second task. Design a field dictionary for an agreed order line.

Worked answer. State the line grain and composite identity first. Define product reference, quantity unit, agreed unit price, currency, schema version and missing-field rules. Define source and temporal fields according to the actual workflow. A field is not required merely because it is fashionable; it is required when the stated use needs it.

2.98 Common wrong ideas#

“Clear field names eliminate the need for a contract.” Names help, but they do not fully define units, timing, scope, missingness or permitted changes.

“A unique database ID proves there are no duplicate business events.” Separate stored rows can refer to the same event.

“Repeated values should always be removed.” Repetition can be legitimate at the chosen grain. Two different orders can have the same total.

“A timestamp tells us everything about time.” It needs a defined meaning and a clock context. Occurrence, receipt and commit are different boundaries.

“An audit record proves the recorded explanation.” It preserves an asserted explanation and relevant context. The explanation still needs suitable evidence.

2.99 Chapter summary in 20 lines#

  1. A useful record makes an interpretable claim.
  2. Values without units and context can be ambiguous.
  3. Field names support a contract but do not replace it.
  4. Syntax describes form; semantics describes meaning.
  5. A data dictionary connects values to their intended interpretation.
  6. A schema version must identify a real definition.
  7. Parsing successfully does not prove domain validity.
  8. A transformation should be explicit rather than silently guessed.
  9. The row grain states what one record represents.
  10. Counting lines is not the same as counting orders.
  11. Aggregating repeated order totals can multiply the answer.
  12. Historical prices must retain their historical meaning.
  13. Identifiers operate within defined scopes.
  14. Row identity and business-event identity can differ.
  15. A repeated delivery and a repeated purchase require different treatment.
  16. Event time and recording time answer different questions.
  17. Provenance helps explain where information came from.
  18. Provenance does not automatically certify truth.
  19. Corrections need meaning, authority and an explicit history policy.
  20. The next chapter explains the values that fields can carry.

Return to contents