Skip to content
KEDBYTE
Site navigation
How Data Works
Chapter
30

Durability From Application to Device

Part E · Storage and Recovery|3,693 words|about 16 min read|Volume E

30.0 What this chapter gives you#

  1. “Saved” can mean copied into an application buffer, accepted by the operating system, acknowledged by a storage device or preserved across a specified failure. Those meanings are not interchangeable.
  2. This chapter follows the acknowledgement chain and teaches you to state exactly which failures a durability claim covers. It separates process failure, operating-system failure, power loss, device loss and correlated loss of copies.
  3. The supplied exercises do not cut power, reset hardware or modify a user’s storage configuration. They test models and bounded local file/database operations, with their limits stated explicitly.

30.1 The acknowledgement chain#

30.1.1 PLAIN — in simple words#

  1. A user sees a success message only after several layers have exchanged requests and responses. Each response confirms something at its own boundary.
  2. A server saying “accepted into memory” is different from a database saying “committed under this durability setting.” Both can be useful contracts, but the interface must not describe one as the other.
  3. A lost response creates uncertainty. The operation may have completed before the reply disappeared, so the caller’s silence cannot be treated as proof of rollback.

30.1.2 PLAIN — a picture in your head#

  1. A shop hands a document to a clerk, the clerk hands it to a filing department, and the department places it in a protected archive. Each handover can produce a receipt.
  2. The first receipt does not prove that the archive step finished unless the organization explicitly waits for that boundary before issuing it.
  3. Where the comparison breaks: software layers can buffer, reorder and batch work. The chain is a protocol, not necessarily one person waiting motionless at each step.

30.1.3 PLAIN — a worked example#

  1. Define three API statuses for a hypothetical service: queued, locally committed and externally confirmed. Queued means the request is retained under the queue’s stated contract; locally committed means the database accepted the transaction; externally confirmed means a named external participant acknowledged its defined action.
  2. A request can reach locally committed while the client never receives that response. It can also be locally committed while the outbox delivery remains pending.
  3. The API should expose the original request identity so the caller can reconcile these states rather than creating an accidental second operation.
  4. The status names are a proposed teaching contract, not claims about KedByte’s live website or a deployed service.

30.1.4 PLAIN — what is really happening inside#

  1. The application maps lower-level results into business responses. Incorrect mapping can weaken a strong database guarantee by showing success too early or can hide an already committed result behind a generic failure.
  2. Database settings decide which lower-level persistence conditions COMMIT waits for. Replication settings can add remote acknowledgement requirements, each with its own meaning.
  3. Document the exact success boundary and make tests observe it. A diagram saying “disk” without naming the promised failure survival is incomplete.

30.1.5 TECHNICAL — the engineer’s version#

  1. A durability contract should name the acknowledged operation, persistence boundary, fault model and assumptions about storage correctness. PostgreSQL’s reliability documentation explicitly discusses the cache layers between memory and non-volatile storage. [S130]
  2. Asynchronous commit intentionally changes when acknowledgement occurs relative to WAL persistence. It is not equivalent to ordinary durable commit merely because both statements eventually return success. [S136]
  3. Request idempotency and outcome lookup address uncertain completion, not the physical persistence of the underlying storage. Both protocols are needed when the service requires both properties.

30.1.6 WORDS — remember these#

  1. Acknowledgement chain: the sequence of accepted handovers — responses from application, database, operating system and storage layers with distinct meanings. Success boundary: the event a positive response promises has occurred — the precise acceptance or persistence condition represented to the caller. Durability contract: what accepted data is promised to survive — a guarantee defined against specified failures and storage assumptions.

30.2 Process memory and operating-system caches#

30.2.1 PLAIN — in simple words#

  1. A program can hold data in its own memory before passing it to the operating system. The operating system can then hold data in a cache before writing it to lower storage.
  2. Ending one program does not necessarily erase the operating system’s cache. Therefore a process-kill test is not the same as a sudden power-loss test.
  3. Conversely, a successful buffered write call may only mean that bytes were copied into an intermediate buffer. The program must understand the interface it used.

30.2.2 PLAIN — a picture in your head#

  1. Mira’s notebook, the office’s outgoing tray and the archive cabinet are three different places. Losing Mira’s notebook does not destroy papers already in the office tray; losing the whole office’s power-protected storage is another event.
  2. Testing only the notebook’s loss does not establish the archive’s behaviour in every disaster.
  3. Where the comparison breaks: memory and cache failures do not follow simple room boundaries. Hardware, kernel, filesystem and controller behaviour all matter.

30.2.3 PLAIN — a worked example#

  1. A Python file object receives text in a user-space buffer. Calling its flush() pushes buffered output through the underlying stream interface; it is not by itself a universal guarantee that the device has durably stored the bytes.
  2. On a supported ordinary file, an application can request a lower persistence boundary using the platform’s synchronization API, such as os.fsync() after flushing the language buffer.
  3. Even then, inspect errors and the platform contract. The API cannot promise survival of the complete loss of the only storage device holding the data.
  4. This is an explanation of layers, not a replacement for a database engine’s carefully designed logging and file-management protocol. [S43] [S64] [S141]

30.2.4 PLAIN — what is really happening inside#

  1. Buffered interfaces combine small writes for efficiency. A lower write can be partial or fail later, so robust software cannot assume that every requested byte was accepted merely because it attempted one call.
  2. Operating-system writeback may occur after the application has continued. Delayed errors need to be surfaced at the relevant synchronization or closing boundary according to the platform.
  3. Database engines already implement these details under their supported configurations. Application authors should not bypass the engine by editing its files directly.

30.2.5 TECHNICAL — the engineer’s version#

  1. Linux write() can return fewer bytes than requested. Successful return does not, by itself, guarantee that the data has been committed to persistent storage. [S142]
  2. Language-buffer flushing and file synchronization are separate operations. A correct low-level file protocol accounts for partial writes, synchronization errors, file metadata and directory entry persistence where required. [S141]
  3. Tests must identify whether they terminate only a process, crash the operating system or remove power. The surviving cache layers differ between those fault models.

30.2.6 WORDS — remember these#

  1. User-space buffer: data waiting inside the program or library — memory used to batch output before passing it to the operating system. Page cache: operating-system memory holding file data — a cache between application I/O and lower storage devices. Partial write: fewer bytes were accepted than requested — an I/O result requiring explicit handling rather than an assumption of complete transfer.

30.3 Flush requests#

30.3.1 PLAIN — in simple words#

  1. A flush request asks a layer to advance pending work toward a lower boundary. Different APIs use the word for different guarantees.
  2. Synchronizing file contents does not always synchronize the directory entry that makes a newly created or renamed file discoverable after a crash.
  3. Atomic visibility and durability are also different. A rename can appear all at once to concurrent readers while still requiring a persistence protocol for the promised crash behaviour.

30.3.2 PLAIN — a picture in your head#

  1. The archive stores a document securely, but its catalogue card naming the document has not yet been secured. After an interruption, the document and its discoverable name can have different states.
  2. A reliable filing procedure considers both the contents and the catalogue.
  3. Where the comparison breaks: filesystem rules differ, and not every platform accepts the same directory-synchronization operations. Copying a Linux recipe into another operating system without verification can be wrong.

30.3.3 PLAIN — a worked example#

  1. A conceptual safe-publication workflow writes a new file, verifies all intended bytes were written, synchronizes the file under the platform contract, installs its name using the appropriate atomic operation, and synchronizes required directory metadata.
  2. If the process crashes before the name installation, readers should still use the previous installed version under the chosen design. If the installation completes durably, readers can use the new complete version.
  3. The exact sequence and guarantees depend on the filesystem, operating system and overwrite semantics. This paragraph is a protocol outline, not executable cross-platform recovery code.
  4. Database backup and storage tools should use their documented procedures rather than a homemade file replacement routine applied to a live database.

30.3.4 PLAIN — what is really happening inside#

  1. The filesystem coordinates data and metadata, while the device may have its own caches and flush commands. The application observes a returned result, not the physical movement of each electron.
  2. Error handling is part of correctness. A failed sync is not a warning to ignore while still reporting the promised durable success.
  3. Repeated sync calls do not fix a fundamentally incompatible or dishonest lower storage contract. Verify the supported stack and its documented assumptions.

30.3.5 TECHNICAL — the engineer’s version#

  1. Linux fsync() synchronizes file data and associated metadata and waits for the device to report completion. Its manual explicitly notes that synchronizing a file does not necessarily synchronize its containing directory entry. [S141]
  2. fdatasync() has a related but narrower metadata requirement: it synchronizes metadata needed for subsequent data retrieval, rather than all unrelated metadata. This distinction is platform-specific API detail, not a generic database setting. [S141]
  3. Do not confuse an I/O synchronization guarantee with application-level transaction atomicity. Persisting two independent files successfully does not make their updates one atomic business operation.

30.3.6 WORDS — remember these#

  1. File synchronization: request persistence of pending file state — an operating-system operation with documented data, metadata and completion semantics. Directory durability: preserve the file’s naming relationship — persistence of directory metadata needed to find the intended file after failure. Atomic visibility: observers do not see an intermediate replacement state — a concurrency property distinct from surviving a crash or power loss.

30.4 Device promises#

30.4.1 PLAIN — in simple words#

  1. A storage device can acknowledge a write while retaining data in a volatile cache unless the protocol and configuration require a stronger boundary.
  2. Power-loss protection can help preserve certain cached state, but its scope must be understood. It does not protect against every device fault or the loss of the entire storage location.
  3. A database depends on the lower stack honouring its promises. No database setting can make a failed or dishonest device become reliable merely by naming it durable.

30.4.2 PLAIN — a picture in your head#

  1. A courier says a parcel is secured, but it is actually still in an unlocked van. Whether the receipt is trustworthy depends on what “secured” promised and whether the courier fulfilled it.
  2. A backup power supply may keep the van’s lock working, but it does not prevent the van itself from being destroyed.
  3. Where the comparison breaks: device persistence has precise command semantics and firmware behaviour. Consumer metaphors cannot substitute for tested hardware and supported filesystem configurations.

30.4.3 PLAIN — a worked example#

  1. A database writes an 8,192-byte page. Suppose the storage path can fail after only part of that page reaches its final location. The surviving page may contain a mixture of old and new portions.
  2. This is a torn-page risk, not merely the loss of the entire latest page. Recovery needs sufficient information to reconstruct or detect the damaged state under its design.
  3. PostgreSQL’s full-page-image mechanism addresses partial-page-write recovery within its documented assumptions. A page checksum can detect some damage but does not itself contain the original page needed to repair it. [S130] [S137]
  4. The numerical page size here matches the common PostgreSQL page size; it does not claim that all devices provide 8,192-byte atomic writes.

30.4.4 PLAIN — what is really happening inside#

  1. Controllers, devices and filesystems can reorder or cache work. Flush and barrier semantics constrain that behaviour when correctly implemented and configured.
  2. A battery-backed cache has maintenance requirements. An exhausted battery or an unverified firmware behaviour can invalidate assumptions that previously held.
  3. Avoid casually disabling synchronization or barriers to obtain a better benchmark. First identify exactly which guarantee would change and whether the application can tolerate that loss.

30.4.5 TECHNICAL — the engineer’s version#

  1. PostgreSQL’s reliability guidance discusses operating-system, controller and device caches, full-page writes and the need for a trustworthy storage subsystem. Those assumptions are part of its durability story. [S130]
  2. Checksums support detection, not general reconstruction or authentication. A checksum-valid page can still encode a logically wrong business change. [S137]
  3. The educational tests do not validate a device’s flush implementation, controller battery, firmware, atomic-write unit or power-loss protection. Such claims require separate hardware-specific evidence.

30.4.6 WORDS — remember these#

  1. Volatile write cache: acknowledged data still depends on power — intermediate storage that can lose pending writes when its power disappears. Torn page: only part of a page update survives — a mixed old/new physical page caused by an interrupted multi-part write. Power-loss protection: preserve specified device state during power interruption — a hardware capability whose scope and health must be verified.

30.5 Failure domains#

30.5.1 PLAIN — in simple words#

  1. Two copies are useful only to the extent that the failures you care about do not destroy both together.
  2. Two files on one disk are separate filenames but share the disk’s failure. Two disks in one machine share other risks, such as a mistaken deletion command or a stolen machine.
  3. Replication can spread copies across machines while faithfully spreading an accidental deletion. Recovery requires both independent failure planning and a way to return to an earlier valid state.

30.5.2 PLAIN — a picture in your head#

  1. Keeping two photocopies in the same folder helps if one sheet is smudged. It does not help if the folder is lost.
  2. Keeping another copy elsewhere changes the risk, but someone must still verify that the remote copy is readable and current enough for its purpose.
  3. Where the comparison breaks: digital copies can fail through shared credentials, software defects and automation, not only physical disasters. Geography alone does not make failures independent.

30.5.3 PLAIN — a worked example#

  1. A hypothetical system has a primary database and a replica in the same building. It tolerates some individual-machine failures but not necessarily a building-wide outage.
  2. Add a backup in a separate administrative and storage boundary. This can improve recovery options, but its access credentials, encryption keys, retention and restore procedure must also survive the incident.
  3. If the same compromised administrator can erase the primary, replica and backup, the copies share a destructive authority domain despite their different locations.
  4. The example is a design exercise, not a claim that any named provider guarantees independence between its products.

30.5.4 PLAIN — what is really happening inside#

  1. Failure domains include processes, hosts, devices, racks, sites, regions, identities, software versions and operational procedures.
  2. Independence is an assumption to test and document, not something proved by drawing two boxes on a diagram.
  3. A complete recovery plan also includes dependencies such as configuration, secrets, keys and application versions. Restored database bytes can remain unusable if the only decryption key was lost.

30.5.5 TECHNICAL — the engineer’s version#

  1. Build a fault matrix listing which copies and dependencies survive each proposed failure. Distinguish availability during the fault from recoverability after it.
  2. Include logical corruption and operator error as well as hardware loss. A replica of the latest state is not necessarily a restore point preceding a bad transaction. [S131]
  3. Avoid multiplying independent-failure probabilities when the independence assumption is unsupported. Shared infrastructure and authority can dominate the practical risk.

30.5.6 WORDS — remember these#

  1. Failure domain: components that can fail together — a shared physical, software, administrative or operational dependency boundary. Correlated failure: one cause affects several copies — a failure pattern that invalidates an assumption of independent redundancy. Recovery dependency: something required to make restored data usable — keys, configuration, software, access and other prerequisites beyond database bytes.

30.6 Testing a durability claim#

30.6.1 PLAIN — in simple words#

  1. A test should state the failure it injects, the data acknowledged before that failure and the evidence checked afterward.
  2. A clean shutdown is not a crash test. A process kill is not a power cut. A restored in-memory copy is not an off-site disaster rehearsal.
  3. Small tests are still useful when their claims stay small. The problem is not limited evidence; it is describing limited evidence as proof of a larger untested guarantee.

30.6.2 PLAIN — a picture in your head#

  1. Testing whether a door closes is useful. It does not prove that the building survives a flood.
  2. A test report becomes more useful when it says exactly which door, under which condition and with what observed result.
  3. Where the comparison breaks: data systems can hide failure until later reads or reconciliation. “The program restarted” is only one observation in a recovery test.

30.6.3 PLAIN — a worked example#

  1. Prepare a disposable database with known records and a separate record of acknowledged request identities. Run a bounded operation, induce the approved fault, then recover into an isolated inspection environment.
  2. Check structural integrity, required relationships, acknowledged outcomes, request replay behaviour and business totals. Record missing or uncertain evidence explicitly.
  3. The book’s actual companion performs in-memory rollback, backup/restore and logical surviving-prefix tests. It does not execute the hardware-fault procedure above.
  4. A real power-loss test must use authorized disposable equipment and a reviewed safety procedure; this book does not ask readers to interrupt power to their ordinary computer or production database.

30.6.4 PLAIN — what is really happening inside#

  1. A test harness must avoid becoming the source of misleading evidence. Recording acknowledgements only in the same volatile memory that the test destroys can leave no trustworthy oracle afterward.
  2. Failure injection timing and reproducibility matter. Randomly killing a process once may miss the vulnerable window entirely.
  3. Preserve logs and configuration together with the outcome. A test that passed under stronger synchronization settings does not validate a later weakened configuration.

30.6.5 TECHNICAL — the engineer’s version#

  1. Separate safety properties, such as “no acknowledged durable transaction is lost under the stated fault,” from liveness and recovery-time properties. A test can establish one observation without proving all schedules or failures.
  2. Record the exact engine/runtime, filesystem, device, settings, fault point, acknowledged set, recovered set and reconciliation procedure. Compare identities and values, not just row counts.
  3. PostgreSQL and SQLite document storage and recovery assumptions that must remain part of the claim. Successful educational checks are not independent hardware certification. [S130] [S53] [S139]

30.6.6 WORDS — remember these#

  1. Failure injection: deliberately create a specified fault — a controlled experiment against an authorized disposable system. Recovery oracle: independent evidence of the expected result — identities and values used to judge what should remain after the tested failure. Durability evidence: observations supporting a bounded persistence claim — a test record tied to configuration, fault model and verification procedure.

30.97 Practice and worked answers#

  1. Read the response contract. An API acknowledges “queued in memory.” Answer: this is not a promise of survival after process or power loss unless another documented mechanism establishes it.
  2. Separate buffers. Python flush() succeeds. Answer: language-buffer progress alone is not a universal device-persistence guarantee.
  3. Handle a short write. A low-level write returns 600 for a 1,000-byte request. Answer: only 600 bytes were reported accepted; the remaining bytes and error conditions require explicit handling.
  4. Find missing metadata. A new file’s data is synchronized but its directory entry is not covered by the platform’s persistence procedure. Answer: content durability and discoverable-name durability may differ.
  5. Classify a torn page. Half an updated page survives. Answer: recovery needs appropriate reconstruction information; a checksum alone can detect some damage but cannot recreate the lost original.
  6. Identify correlated authority. One compromised credential can erase all three copies. Answer: the copies share a destructive administrative failure domain.
  7. Limit a process-kill test. The process dies but the operating system continues. Answer: lower caches can survive; this is not a power-loss test.
  8. Choose a recovery oracle. Compare only row counts after restore. Answer: counts can match while identities or values differ. Check the acknowledged set, relationships and defined business totals.

30.98 Common wrong ideas#

  1. Wrong: every success response means the same kind of saved. Right: name the acknowledgement and failure-survival boundary.
  2. Wrong: a process kill and a power cut test the same thing. Right: different cache and storage layers survive.
  3. Wrong: successful write means all bytes are permanently stored. Right: partial writes and persistence semantics must be handled explicitly.
  4. Wrong: atomic rename automatically proves crash durability everywhere. Right: platform data and metadata synchronization rules still matter.
  5. Wrong: power-loss protection prevents every storage failure. Right: its scope is narrower than total device or site survival.
  6. Wrong: three copies are three independent failure domains. Right: infrastructure and authority can correlate their loss.
  7. Wrong: a checksum repairs damaged data. Right: detection and reconstruction require different evidence.
  8. Wrong: passing bounded local tests certifies production durability. Right: the tested configuration and fault model define the claim.

30.99 Chapter summary in 20 lines#

  1. Saved data passes through several acknowledgement boundaries.
  2. The application must map those boundaries into honest business statuses.
  3. A lost response can leave a committed operation unresolved to its caller.
  4. User-space buffers, operating-system caches and device caches are separate layers.
  5. Language-buffer flushing is not the same as device synchronization.
  6. Low-level writes can be partial and require error handling.
  7. File data and directory metadata have distinct persistence requirements.
  8. Atomic visibility does not by itself establish durability.
  9. Synchronization depends on the lower stack honouring its documented contract.
  10. Volatile controller and device caches affect power-loss assumptions.
  11. Partial-page writes require appropriate recovery protection.
  12. Checksums detect some damage but do not reconstruct missing content.
  13. Power-loss protection has a specific scope and maintenance requirements.
  14. Copies can share physical, software and administrative failure domains.
  15. Replication does not replace a recoverable earlier version.
  16. Keys, configuration and application dependencies belong in recovery planning.
  17. Tests must state the exact fault they inject.
  18. Preserve an independent record of acknowledged identities and expected values.
  19. Reconcile recovered meaning as well as structural validity.
  20. Report bounded evidence without extending it to untested hardware or failures.

Return to contents