Skip to content
KEDBYTE
Site navigation
How Data Works
Chapter
31

Backups, Restore and Point-in-Time Recovery

Part E · Storage and Recovery|3,556 words|about 15 min read|Volume E

31.0 What this chapter gives you#

  1. A backup is a copy kept for a recovery purpose. Its value depends on whether the required state can actually be restored, validated and made usable within the intended limits.
  2. This chapter distinguishes logical exports, physical backups, snapshots, log archives and restore rehearsals. You will calculate recovery objectives and identify dependencies that a database file alone does not contain.
  3. The actual lab restores synthetic SQLite data into a fresh isolated destination. PostgreSQL point-in-time recovery is explained from its documentation, not claimed as an executed server restore.

31.1 Copies with a recovery purpose#

31.1.1 PLAIN — in simple words#

  1. A backup is not simply any second copy. It has a defined source, time or recovery boundary, retention policy and procedure for turning it back into usable data.
  2. A replica usually follows current changes, including mistakes. A backup can preserve an earlier state needed after accidental deletion or corruption.
  3. An export can be useful without containing everything required for a complete service recovery. Know what it includes and what it omits.

31.1.2 PLAIN — a picture in your head#

  1. A photocopy of yesterday’s ledger can help recover from today’s mistaken erasure. A second clerk copying today’s erasure immediately may leave two equally damaged ledgers.
  2. The recovery copy needs a date, an identity and enough context to interpret its contents.
  3. Where the comparison breaks: digital backups may contain physical engine structures, logical SQL or incremental changes. They are not interchangeable sheets that can be opened by any database version.

31.1.3 PLAIN — a worked example#

  1. Mira exports the four canonical order lines. Their total is 28,650 paise across six units. That export can help reconstruct or verify those lines.
  2. It does not automatically include customers, permissions, product metadata, outbox state, schema definitions, secrets or the application version needed to run the service.
  3. A complete recovery inventory lists every required component and how it is acquired. Some dependencies should be backed up separately under stronger access restrictions rather than bundled into a public teaching archive.
  4. The canonical export contains invented data only. No live customer records or credentials are included in this book’s files.

31.1.4 PLAIN — what is really happening inside#

  1. Logical backups describe objects and values in a form the target can reconstruct. Physical backups preserve the engine’s storage representation under its supported consistency procedure.
  2. The two approaches have different portability, granularity and recovery characteristics. A logical export can support selected-object restoration, while a physical base backup can participate in a matching WAL-replay recovery chain.
  3. Backups themselves need monitoring. A scheduled job that stopped producing usable copies weeks ago is not a current recovery capability.

31.1.5 TECHNICAL — the engineer’s version#

  1. PostgreSQL distinguishes SQL dumps from filesystem-level base backups. pg_dump output is logical and is not a physical base for continuous WAL replay. [S67] [S131]
  2. SQLite’s online backup API supports making a consistent database copy through the engine. Copying only a live main file without the appropriate journal/WAL coordination can produce an invalid or incomplete backup. [S55] [S139]
  3. A backup manifest should name source identity, engine version, acquisition method, completion state, checksums and required accompanying material. A filename containing a date is not sufficient provenance.

31.1.6 WORDS — remember these#

  1. Backup: a copy retained for a defined recovery purpose — data and metadata acquired under a procedure that supports a stated restoration objective. Logical backup: reconstruct objects from their logical definitions and values — an export such as SQL or another supported logical representation. Physical backup: preserve an engine’s storage representation — a copy requiring the matching consistency, version and recovery procedure.

31.2 Consistent snapshots#

31.2.1 PLAIN — in simple words#

  1. A recovery copy must represent a state the database can interpret consistently. Copying files one after another while they change can mix incompatible moments.
  2. A supported snapshot or backup protocol establishes the required boundary, sometimes using recovery logs to reconcile files copied over an interval.
  3. A storage snapshot is not automatically an application-consistent service snapshot. Other databases, object stores and external actions may have separate boundaries.

31.2.2 PLAIN — a picture in your head#

  1. Copying the first half of a ledger before a transfer and the second half afterward can create an impossible combined total.
  2. A controlled copying procedure either fixes the view or preserves the instructions needed to reconstruct the matching state.
  3. Where the comparison breaks: database backups can be consistent even when file copying spans time, provided the engine’s backup and log protocol supplies the required recovery evidence. They need not all be instantaneous photographs.

31.2.3 PLAIN — a worked example#

  1. A toy service stores an order in one database and its attachment in an object store. The database backup records attachment key ATT-7, but the object-store snapshot predates ATT-7’s upload.
  2. Each component can be internally valid while the restored service contains a missing attachment reference.
  3. The recovery plan must define the relationship: coordinate acquisition, retain referenced object versions or verify and reconcile missing objects under an explicit policy.
  4. Replacing the missing attachment with a blank file would conceal the failure rather than restore the original evidence.

31.2.4 PLAIN — what is really happening inside#

  1. Recovery consistency has layers: file-format validity, transaction consistency, cross-store references and business meaning.
  2. Engine-supported backup procedures specify which logs, metadata and completion markers are required. A partially completed backup must not be labelled usable merely because many files exist.
  3. The destination should be isolated while it is validated. Starting ordinary outbound workers immediately after restore can repeat historical notifications or other external effects.

31.2.5 TECHNICAL — the engineer’s version#

  1. PostgreSQL continuous archiving combines a supported physical base backup with the required WAL sequence. The base copy can span time because replay resolves the corresponding physical inconsistencies under the documented procedure. [S131]
  2. SQLite’s backup API operates through database connections rather than treating a changing file as an uncoordinated byte source. The companion uses this API for its actual synthetic restore test. [S55] [S86]
  3. Define application-consistency checks beyond engine startup: foreign keys, expected objects, request/outbox relationships, attachment references and reconciliation totals.

31.2.6 WORDS — remember these#

  1. Consistent backup boundary: the state the recovery procedure can validly reconstruct — a snapshot or log-supported acquisition point with matching metadata. Application consistency: related service components agree — correctness of references and workflows beyond one engine’s internal file validity. Restore isolation: keep the recovered copy from affecting live systems — a controlled environment with outbound effects and production access disabled during validation.

31.3 Log retention#

31.3.1 PLAIN — in simple words#

  1. Point-in-time recovery starts from an earlier supported base and replays the required changes only as far as the chosen target.
  2. The necessary log chain must be continuous. A missing required segment can prevent reaching the target even if newer segments remain available.
  3. Retaining a base backup without its required logs, or retaining logs without a usable matching base, can leave a collection of files that cannot meet the promised recovery objective.

31.3.2 PLAIN — a picture in your head#

  1. A verified ledger at Monday morning plus every accepted correction through Thursday can reconstruct Thursday’s state. Missing Tuesday’s corrections cannot be repaired by reading Friday’s entries alone.
  2. The base and the journal belong to one history, including any later branches created by recovery.
  3. Where the comparison breaks: real log segments have engine-specific identities, timelines and validation rules. Sorting filenames lexically is not a recovery algorithm.

31.3.3 PLAIN — a worked example#

  1. A synthetic recovery chain has a valid base at 10:00 and complete archived changes through 10:30. An incident occurs at 10:40.
  2. The chain can support an appropriate target no later than its verified coverage, subject to the engine’s transaction boundaries. It cannot claim recovery through 10:40 merely because the primary generated changes after 10:30.
  3. If the intended target is 10:25, the plan must preserve every required part from the base through that target. A missing segment at 10:15 cannot be ignored.
  4. These times illustrate coverage; actual WAL recovery targets and inclusive/exclusive settings require the documented engine procedure.

31.3.4 PLAIN — what is really happening inside#

  1. Archiving is a transfer and retention workflow with its own failures. Generated WAL is not the same as successfully archived, verified and retrievable WAL.
  2. Recovery can create a new timeline or history branch. Later backups and log references must identify which history they continue.
  3. Retention deletion should evaluate dependencies before removing a base or log range. Deleting the oldest-looking file can break several otherwise retained restore points.

31.3.5 TECHNICAL — the engineer’s version#

  1. PostgreSQL PITR requires a continuous sequence of archived WAL extending back at least to the required start of the base backup. Its documentation separately explains recovery targets and timelines. [S131]
  2. Record archive success, failures, coverage gaps and restore-tested ranges. A counter saying that some archive command succeeded does not prove every desired target remains recoverable. [S144]
  3. Do not replay a physical log against a different logical export or unrelated cluster. Match base identity, version compatibility and timeline evidence before attempting recovery.

31.3.6 WORDS — remember these#

  1. Point-in-time recovery: restore a supported earlier state — replay from a matching base to a chosen engine-supported recovery target. Archive coverage: the recovery history actually retained and retrievable — verified continuity of required log material over a stated range. Recovery timeline: a named branch of physical recovery history — metadata distinguishing histories after recovery or promotion under the engine’s rules.

31.4 Recovery objectives#

31.4.1 PLAIN — in simple words#

  1. Recovery point objective, or RPO, describes the tolerated loss of recent data measured against a defined incident boundary. Recovery time objective, or RTO, describes how quickly the required service should be restored.
  2. An objective is a requirement, not evidence that the current backup system achieves it. A restore rehearsal measures actual performance under stated conditions.
  3. The service is not necessarily recovered when file copying finishes. Validation, application startup, access checks and controlled return to users can consume substantial time.

31.4.2 PLAIN — a picture in your head#

  1. Mira says, “We can tolerate losing at most fifteen minutes of accepted work, and we need usable service within one hour.” These are two different targets.
  2. Finding a recent ledger quickly does not help if nobody can unlock the cabinet or run the software that reads it.
  3. Where the comparison breaks: recovery objectives need precise definitions for accepted work, incident start, degraded service and completion. Everyday phrases such as “back online” can hide incompatible meanings.

31.4.3 PLAIN — a worked example#

  1. An incident occurs at 10:40. The verified recoverable state reaches 10:30, so the recent-data gap is ten minutes under this simplified timeline. That fits a fifteen-minute RPO only if the coverage and acceptance definitions truly match.
  2. A rehearsal takes 32 minutes to restore, 18 minutes to validate and 12 minutes to return the required service safely. Total recovery time is 62 minutes.
  3. A one-hour RTO is missed by two minutes. Reporting only the 32-minute copy time would conceal the operational result.
  4. These invented values show how to calculate the measures; they are not a promise about the book’s lab or any hosting plan.

31.4.4 PLAIN — what is really happening inside#

  1. Restore time depends on data volume, retrieval bandwidth, decompression, replay, index work, validation and operational coordination.
  2. A backup in low-cost archival storage can have a retrieval delay that dominates the recovery timeline. Keys or approvals can also become bottlenecks.
  3. Objectives can differ by service function. A read-only emergency view may return before full write service, but the completion definition must state that distinction honestly.

31.4.5 TECHNICAL — the engineer’s version#

  1. A lower-bound transfer calculation for 1 TB decimal at a sustained 100 MB/s is 10,000 seconds, about 2.78 hours, before replay, validation or contention. It is not an end-to-end restore estimate.
  2. Define RPO relative to acknowledged operations and verified recoverable history, not merely backup-job start times. Define RTO’s start and finish events before measuring it.
  3. Keep objectives, measured rehearsal results and residual assumptions in separate fields. A green backup-job status is not a measured RTO.

31.4.6 WORDS — remember these#

  1. RPO: how much recent accepted work may be lost — a recovery-point requirement defined against a specified incident and acceptance boundary. RTO: how long required service restoration may take — a recovery-time requirement with explicit start and completion events. Restore rehearsal: actually exercise the recovery procedure — a controlled test measuring whether the required state and service can be recovered.

31.5 Restore rehearsals#

31.5.1 PLAIN — in simple words#

  1. A backup that has never been restored is an untested recovery input. Reading its filename and checking that it is large enough is not a substitute for trying the procedure.
  2. Rehearse in isolation, verify the restored state and record the results. Keep outbound side effects disabled until their replay risks are understood.
  3. Tests should include failure cases: missing files, wrong keys, incompatible versions and a target beyond the retained log coverage.

31.5.2 PLAIN — a picture in your head#

  1. An emergency key is useful only if it opens the right door. A rehearsal tries it before the emergency rather than admiring the key’s label.
  2. The rehearsal also checks what is behind the door and whether the staff can perform the required work.
  3. Where the comparison breaks: restoring historical state can reactivate pending jobs, old permissions or stale request records. A restored service can create new harm if connected to live systems without review.

31.5.3 PLAIN — a worked example#

  1. The companion creates the canonical synthetic shop, makes an engine-supported SQLite backup into a fresh destination and checks that order totals remain O-1042=17,100 and O-1043=11,550 paise.
  2. It also checks the combined 28,650-paise total, foreign-key violations and structural integrity under the available SQLite checks.
  3. A later separate mutation of the source should not silently mutate the already completed backup destination. The test verifies separation under this local setup.
  4. This is an actual bounded backup/restore exercise, not an off-site recovery rehearsal, encrypted-archive test or PostgreSQL PITR run.

31.5.4 PLAIN — what is really happening inside#

  1. Structural verification and business verification answer different questions. A readable database can still contain the wrong target time or missing expected records.
  2. Compare identities and values, not just totals: two offsetting errors can leave a sum unchanged. Include representative application workflows and boundary tests.
  3. Record the exact backup input, software versions, procedure, timings and checks. A later successful restore from another file does not retroactively validate an earlier failed backup.

31.5.5 TECHNICAL — the engineer’s version#

  1. PostgreSQL pg_verifybackup checks a base backup against its manifest and performs defined WAL-related validation. Its documentation explicitly states that this does not replace test restores and application-data verification. [S138]
  2. A restore acceptance report should include acquisition identity, hashes, target boundary, isolated environment, executed commands, checks, timing and unresolved findings.
  3. Restored outbox and idempotency records require special review before external dispatch resumes. A backup can precede actions already performed outside the restored database; reconciliation must prevent unsafe replay.

31.5.6 WORDS — remember these#

  1. Backup verification: check the retained input against defined expectations — structural, manifest or checksum validation that may precede a full restore. Restore acceptance: decide whether the recovered state meets its objective — a documented result based on actual recovery and relevant application checks. Replay review: inspect historical pending work before reactivation — reconciliation of restored tasks with effects that may already exist outside the backup boundary.

31.6 Access and retention of backups#

31.6.1 PLAIN — in simple words#

  1. Backups can contain sensitive information long after the live application has changed or deleted it. They need access controls, retention rules and a way to prevent old data from quietly returning to normal use.
  2. Encryption helps protect stored bytes, but recovery also depends on keeping the required keys available to authorized responders.
  3. Retention should preserve the required restore points without claiming that every copy lasts forever. Storage cost, purpose and deletion obligations need an explicit policy.

31.6.2 PLAIN — a picture in your head#

  1. An old ledger in a locked archive is still a ledger containing people’s information. Locking it does not make its contents disappear.
  2. Destroying the only key may prevent recovery, while leaving a spare key on the door defeats the access control.
  3. Where the comparison breaks: digital deletion and encryption-key lifecycle are complex across replicas, caches and services. A local file deletion is not proof that every retained copy was erased.

31.6.3 PLAIN — a worked example#

  1. A teaching policy retains daily backups for 30 days and monthly backups for a separately defined period. The policy is an example, not legal advice or a universal retention recommendation.
  2. A deletion request affects the live system today, while an older restricted backup still contains the record until its permitted retention expires.
  3. A restore procedure must reapply the relevant deletion or suppression decisions before that old record becomes available to ordinary users again, under the organization’s approved policy.
  4. Keep evidence of which stores were checked and which retained exceptions remain. Do not report universal erasure based only on the live-table DELETE.

31.6.4 PLAIN — what is really happening inside#

  1. Backup credentials should be scoped to their tasks. Separate the ability to create backups from the ability to erase every historical copy where the storage system and operational model support that separation.
  2. Retention automation needs dependency checks so it does not delete logs required by retained base backups. It also needs monitoring for failures and unexpectedly retained data.
  3. Recovery access should be rehearsed under realistic restrictions. A plan that works only with an unavailable administrator’s personal credentials is operationally fragile.

31.6.5 TECHNICAL — the engineer’s version#

  1. Record backup classes, purpose, retention, encryption-key dependencies, allowed operators and restore-time suppression procedures. Chapter 48 develops lineage and deletion boundaries in more detail.
  2. Protect manifests and integrity metadata against unauthorized replacement. A checksum stored beside a file can detect accidental damage but does not authenticate the file if an attacker can rewrite both.
  3. The book does not prescribe a jurisdiction-specific retention period. Legal and contractual requirements must be checked for the actual organization and data classes; the technical protocol should expose those decisions rather than invent them.

31.6.6 WORDS — remember these#

  1. Backup retention: how long recovery copies remain available — a policy governing restore-point preservation and eventual retirement. Key dependency: recovery requires particular cryptographic material — an access and availability requirement for decrypting retained data. Restore-time suppression: prevent retired data from re-entering ordinary use — reapplication of approved deletion or access decisions after historical recovery.

31.97 Practice and worked answers#

  1. Classify an export. A CSV contains only four order lines. Answer: it is a useful logical data extract, not automatically a complete service backup.
  2. Choose a PITR base. Can a PostgreSQL SQL dump be used directly as the physical base for WAL replay? Answer: no; use the supported matching physical-backup procedure.
  3. Find cross-store inconsistency. The database references ATT-7, absent from the object-store restore. Answer: component validity does not establish application consistency; reconcile the missing dependency.
  4. Calculate the data gap. Incident at 10:40, verified coverage through 10:30. Answer: ten minutes under the stated acceptance and coverage assumptions.
  5. Calculate recovery time. Restore 32 minutes, validate 18, return service 12. Answer: 62 minutes, missing a one-hour objective by two minutes.
  6. Compute a transfer floor. One TB decimal at 100 MB/s. Answer: 10,000 seconds, about 2.78 hours before other recovery work.
  7. Interpret manifest success. pg_verifybackup reports success. Answer: perform an actual test restore and application checks; the tool does not establish every recovery property.
  8. Review restored workers. Historical outbox rows appear pending. Answer: keep dispatch isolated until external effects and replay identities are reconciled.

31.98 Common wrong ideas#

  1. Wrong: any second copy is a usable backup. Right: provenance, consistency and a tested restore procedure matter.
  2. Wrong: replication protects against every mistaken deletion. Right: replicas can copy the mistake.
  3. Wrong: individually valid component snapshots form a valid service snapshot. Right: cross-store references can disagree.
  4. Wrong: newer WAL can compensate for missing required older segments. Right: the recovery chain must be continuous.
  5. Wrong: RPO and RTO are achieved because they are written in a policy. Right: objectives need measured evidence.
  6. Wrong: copying finished means recovery finished. Right: validation and safe service return also take time.
  7. Wrong: matching row counts prove correct restoration. Right: identities, values and business relationships must be checked.
  8. Wrong: deleting live data proves every backup copy is gone. Right: retained copies and restore-time controls need separate accounting.

31.99 Chapter summary in 20 lines#

  1. A backup has a defined recovery purpose and provenance.
  2. Logical exports and physical backups support different procedures.
  3. A PostgreSQL logical dump is not a physical WAL-replay base.
  4. Supported backup interfaces coordinate changing data safely.
  5. Cross-store application consistency extends beyond one database’s validity.
  6. Keep restored environments isolated while validating them.
  7. Point-in-time recovery requires a matching base and continuous required logs.
  8. Generated log data is not automatically archived and retrievable coverage.
  9. Recovery timelines distinguish branches of physical history.
  10. Retention deletion must preserve dependencies of retained restore points.
  11. RPO defines tolerated recent-data loss under a stated boundary.
  12. RTO defines the allowed time to restore required service.
  13. Objectives are different from measured rehearsal results.
  14. Transfer, replay, validation and controlled service return all consume time.
  15. Manifest verification does not replace a test restore.
  16. Check identities, values, constraints and business reconciliation after recovery.
  17. Historical outbox tasks need replay review before external dispatch resumes.
  18. Backups require access, encryption-key and retention controls.
  19. Restoring old data must not silently undo approved deletion or suppression decisions.
  20. Report exactly which recovery procedure was executed and which remains untested.

Return to contents