Replication: More Than One Copy
Introductions, exercises and summaries stay visible.
33.0 What this chapter gives you#
- A second copy of a database can help answer questions, survive some failures or move data closer to readers. It also creates a new question: which copy knows about which changes?
- You will distinguish physical from logical replication, receiving from applying, synchronous from asynchronous acknowledgement, and a useful replica from a recoverable historical backup.
- The examples use a separate fictional replication register, not a live shop cluster. Positions such as 40 and 41 are teaching sequence numbers, not PostgreSQL WAL addresses. The companion model demonstrates prefixes and stale reads; it does not implement a network replication protocol.
33.1 Why copies exist#
33.1.1 PLAIN — in simple words#
- Keeping one copy makes the meaning of “current” relatively straightforward, but it puts much responsibility on one machine. If that machine is unavailable, readers may have nowhere else to ask.
- A replica is another maintained copy. It might support recovery from a machine failure, serve read-only reports or keep frequently requested information near a distant branch.
- Those purposes are not identical. A server configured for delayed reporting may be a poor emergency replacement. A nearby replacement may share the same power, building or administrative failure as the original.
33.1.2 PLAIN — a picture in your head#
- Mira keeps an order notebook and asks Dev to maintain a second notebook from numbered change slips. Dev can answer some questions while Mira serves another customer.
- If slips are delayed, Dev’s notebook is not necessarily wrong about the past; it is simply behind the latest accepted changes. Problems begin when that older view is presented as current.
- Where the comparison breaks: a database replica follows an engine-specific protocol with transaction and recovery rules. It is not an independent human witness confirming that a sale physically occurred.
33.1.3 PLAIN — a worked example#
- Suppose a hypothetical service receives 900 read requests and 100 write requests each second. Two followers could share some of the read work while the leader remains responsible for writes.
- Splitting 900 reads equally would assign 300 to each of three machines. That arithmetic is a proposed allocation, not a measured threefold speedup: the machines also spend resources transferring and applying changes.
- If 200 of the reads must immediately observe a just-completed write, routing those indiscriminately to lagging followers can violate the application requirement. Read traffic must be classified before it is distributed.
- The benefit therefore depends on the workload, the freshness contract and the failure being addressed. “Three servers” alone specifies none of them.
33.1.4 PLAIN — what is really happening inside#
- A replication design chooses an authoritative source or conflict protocol, captures changes, transports them and brings each destination to a defined state.
- Additional readers can increase useful capacity only if the cost of serving them is not replaced by a larger bottleneck in change generation, transmission or replay.
- The operator must monitor both the original and the copies. A silently disconnected replica can look healthy to a simple process check while holding increasingly old data.
33.1.5 TECHNICAL — the engineer’s version#
- Separate high availability, disaster recovery, read scaling and geographic placement as requirements. Their recovery objectives, consistency needs and failure-domain assumptions should be explicit.
- PostgreSQL 17 distinguishes warm standby operation from hot standby read-only queries. Streaming and log shipping maintain a recovery-capable history under their documented configuration and retention requirements. [S147]
- Replication adds ongoing work and additional privileged data paths. Capacity planning must include replay, retained logs, connection management and recovery rehearsals, not just foreground queries. [S144] [S147]
33.1.6 WORDS — remember these#
Replica: another maintained copy — a data instance updated through a defined replication mechanism. Read scaling: sharing question-answering work — distributing eligible read operations across resources without assuming identical freshness. Failure domain: things that can fail together — a boundary such as process, host, power supply, site or administrative authority.
33.2 Physical and logical approaches#
33.2.1 PLAIN — in simple words#
- One way to maintain a copy is to reproduce changes in the database’s physical storage machinery. Another is to describe changes to named records and apply them to a compatible destination.
- Physical replication closely follows the source engine’s representation. Logical replication works with data objects and their identities, allowing more choice about which records are copied and how the destination is organized.
- Neither label means “copy absolutely everything.” You must inspect the mechanism’s actual coverage, including schemas, sequences, permissions and objects that are not ordinary table rows.
33.2.2 PLAIN — a picture in your head#
- Copying each amended notebook page resembles the physical approach. Sending “change the quantity on order line X” resembles the logical approach.
- Page copying expects a compatible notebook layout. Record instructions require both people to agree on what identifies X and what each field means.
- Where the comparison breaks: physical replication commonly transports recovery records, not photographs or repeated whole-file copies. Logical replication may preserve transaction boundaries rather than sending unrelated individual edits.
33.2.3 PLAIN — a worked example#
- A source contains
order_id,line_no,product_id,quantityandunit_price_paise. The agreed line identity is the pair(order_id, line_no). - A logical update naming only
product_idwould not identify an order line: the same product occurs in many orders. Replication needs an appropriate identity for the action it is applying. - Suppose an analytical subscriber intentionally receives only order lines and products. A row count matching those tables does not prove that customers, roles or sequence state have been copied.
- In PostgreSQL 17, built-in logical replication does not automatically replicate DDL or sequence state. A promotion or schema-change plan must account for those omissions rather than inferring completeness from successfully copied rows. [S164]
33.2.4 PLAIN — what is really happening inside#
- A destination normally needs an initial consistent data boundary and a stream of changes beginning at a compatible position. A gap between the initial copy and the change stream loses work; uncontrolled overlap can repeat work.
- The apply process interprets the stream using its configured schema and identity rules. Incompatible changes or local writes can stop application or create conflicts.
- Physical formats place tighter compatibility requirements on engine versions and storage layout. Logical formats avoid some of those constraints but introduce their own schema, identity and semantic requirements.
33.2.5 TECHNICAL — the engineer’s version#
- PostgreSQL 17 logical replication uses publication/subscription and replication identity. Its transactional consistency guarantee is scoped to publications within a single subscription, not arbitrary independent subscriptions observed at unrelated positions. [S148]
- Physical standby compatibility and logical object coverage are product-specific. Do not generalize a logical subscriber into a complete physical failover target without a separate assessment. [S147] [S164]
- Record initial-snapshot boundaries, schema versions, identity configuration and apply errors. These are evidence needed to diagnose a divergent or stalled subscriber; the transport connection alone is insufficient.
33.2.6 WORDS — remember these#
Physical replication: following the storage engine’s changes — replication tied to physical representation and recovery information. Logical replication: following changes to data objects — replication based on record identity and logical operations rather than identical physical placement. Replication identity: the handle for locating a changed row — the key or supported old-row representation used to apply updates and deletions.
33.3 Synchronous and asynchronous paths#
33.3.1 PLAIN — in simple words#
- An asynchronous path can report success before a remote copy has reached the required point. That can reduce waiting, but a later failure may expose a gap between what was acknowledged and what the survivor knows.
- A synchronous path waits for specified remote confirmation. The important word is “specified”: confirmation of receipt, durable storage and application to readable state are different milestones.
- Waiting for one follower does not imply waiting for every follower. A reader sent to a different, slower machine may still see older data.
33.3.2 PLAIN — a picture in your head#
- Mira can finish with a customer after putting a change slip in her own tray, after Dev says “received,” or after Dev says “entered in my notebook.” Each acknowledgement means something different.
- If Dev has merely received the slip, a query against his notebook can still return the old quantity. The message arrived, but the visible state has not changed.
- Where the comparison breaks: durable confirmation relies on storage and recovery guarantees, not a person’s verbal assurance. A remote operating-system write and a durable flush are not interchangeable.
33.3.3 PLAIN — a worked example#
- At a synthetic instant, the leader has accepted sequence 41. Follower A has received and durably stored 41 but has applied only through 40. Follower B has applied through 39.
- A policy waiting for A’s durable storage can be satisfied while A’s read view is still at 40. A policy waiting for A to apply 41 cannot be satisfied yet.
- Even after A applies 41, a request sent to B can remain stale. The acknowledgement scope must match the read-routing decision.
- For a hypothetical remote path with 35 ms round-trip delay, a commit that must await that remote response cannot complete that confirmation in 2 ms. This is a lower-bound illustration, not a benchmark of a particular service.
33.3.4 PLAIN — what is really happening inside#
- The leader tracks progress reported by each configured follower. A transaction waits until its selected acknowledgement condition is met.
- Slow or missing followers can therefore delay success responses. Changing the policy to stop waiting may restore responsiveness while also changing the protection promised to future requests.
- Configuration is part of the data contract. A per-transaction override can matter even when a global setting appears strict.
33.3.5 TECHNICAL — the engineer’s version#
- With synchronous standbys configured, PostgreSQL 17
synchronous_commit=onwaits for their required durable WAL confirmation;remote_writewaits for a weaker remote file-system write;remote_applyadditionally waits for replay visibility.localdoes not wait for replication. [S149] - When no synchronous standby is configured, these names do not create a remote guarantee. Confirm the active standby-selection policy and actual transaction setting. [S149]
- Synchronous replication is not, by itself, a complete failover, consensus or multi-record application-correctness protocol. Those claims require their own mechanisms and failure assumptions.
33.3.6 WORDS — remember these#
Asynchronous replication: remote progress may follow acknowledgement — a configuration that does not wait for a selected remote replication milestone before reporting success. Synchronous acknowledgement: success waits for specified remote progress — a policy whose exact receipt, flush or apply condition must be named. Apply position: how far changes affect the readable replica — the prefix already replayed into the destination’s visible state.
33.4 Lag and staleness#
33.4.1 PLAIN — in simple words#
- Lag is a gap in progress. Staleness describes how old an answer is relative to the information the application requires. The two are related but not identical.
- A replica can be many bytes behind because one large update is pending, or only a small number of bytes behind while missing a very important cancellation.
- A time since the last replayed transaction can also be misleading when the source has been idle. An old last-change time is not proof that new changes are waiting.
33.4.2 PLAIN — a picture in your head#
- Dev’s last change slip is an hour old. If Mira made no changes during that hour, his notebook may be fully caught up.
- If Mira wrote a crucial new slip ten seconds ago, Dev can be behind even though his last update sounds recent.
- Where the comparison breaks: real systems expose several progress counters with different update intervals and clock assumptions. A single wall-clock subtraction is not a universal freshness measure.
33.4.3 PLAIN — a worked example#
- Suppose a fictional stream has generated 120 MB, sent 115 MB, remotely stored 110 MB and applied 100 MB at one observation boundary.
- The unsent gap is 5 MB; the sent-but-not-stored gap is 5 MB; the stored-but-not-applied gap is 10 MB. The total generated-to-applied difference is 20 MB.
- At a net catch-up rate of 4 MB/s, clearing a fixed 20 MB backlog would take 5 seconds. If the source generates another 4 MB/s while the receiver only applies 4 MB/s, the backlog does not shrink.
- These decimal-byte calculations assume a consistent stream and compatible observation positions. They do not identify which business records are missing or provide a guaranteed catch-up deadline.
33.4.4 PLAIN — what is really happening inside#
- Separate generation, transmission, receipt, flush and apply bottlenecks. More network bandwidth will not resolve a replay process blocked by an incompatible schema or resource shortage.
- Retained log data allows a disconnected destination to catch up only while the required history remains available. Keeping everything indefinitely can exhaust the source’s disk.
- Monitoring therefore needs both freshness evidence and retention-pressure evidence. A slot that preserves history can protect a follower while threatening the primary’s remaining storage.
33.4.5 TECHNICAL — the engineer’s version#
- PostgreSQL exposes WAL positions and replication statistics for distinguishing progress stages. Interpret counters within the same history and account for observation timing rather than subtracting unrelated positions. [S144] [S147]
- Replication slots can retain WAL for disconnected consumers. Limits and monitoring are necessary because retained WAL can consume the available storage; losing required history may force reinitialization. [S147]
- An application freshness requirement should identify a minimum causal position, acceptable age under stated clock assumptions, or authoritative read path. “Lag looks low” is not a complete correctness condition.
33.4.6 WORDS — remember these#
Replication lag: one replication stage is behind another — a difference measured using compatible positions, bytes or carefully interpreted timing evidence. Stale read: an answer lacks required newer information — a read that does not meet the application’s freshness or consistency contract. Catch-up rate: how quickly a backlog shrinks — destination progress minus continuing source generation over the measured interval.
33.5 Read routing#
33.5.1 PLAIN — in simple words#
- Choosing where to read is part of application correctness. A recently changed order should not appear unchanged merely because the next screen used a slower copy.
- Some questions can tolerate an older view. Others need to observe the user’s own completed action or a particular confirmed state before they proceed.
- The design must state what happens when the preferred copy is unavailable: wait, use an authoritative alternative, show explicitly older data, or decline the operation. Quietly weakening the requirement hides the problem.
33.5.2 PLAIN — a picture in your head#
- Mira gives a customer a receipt for change slip 41. The customer asks Dev about that change and shows the receipt number.
- Dev can answer only after his notebook includes at least slip 41, or he can direct the customer to someone who can. Guessing from slip 39 does not satisfy the request.
- Where the comparison breaks: a progress token must be meaningful for the actual history, tenant and operation. A bare number from an unrelated stream is not evidence of causal inclusion.
33.5.3 PLAIN — a worked example#
- Start a separate teaching register at position 40 with quantity 5. A committed change at 41 sets quantity to 2. The follower remains at 40.
- An unrestricted follower read returns 5. A read requiring position 41 must not report that value as satisfying the requirement.
- After replaying the contiguous entry at 41, the follower can return 2 with applied position 41. The model rejects a missing predecessor rather than pretending that receiving position 42 establishes an uninterrupted history.
- The companion tests exercise these states as a pure model. They do not prove a real deployment’s failover-safe session consistency or cross-database token handling.
33.5.4 PLAIN — what is really happening inside#
- A session may carry a minimum observed position, or the application may route selected reads to the current authority. Each approach requires explicit handling of role changes and unavailable dependencies.
- Waiting on a follower must be bounded operationally. A timeout should remain visible; it must not silently convert a required-fresh read into an old answer labelled current.
- Transaction snapshots also matter. A session that already holds an older snapshot may not see a newly applied change merely because another observer reports that replay has advanced.
33.5.5 TECHNICAL — the engineer’s version#
- Read-your-writes and monotonic-read requirements are session-level contracts, distinct from a generic claim that replicas eventually converge. Track the scope and validity of any causal token.
- PostgreSQL
remote_applycan support useful visibility conditions on the selected synchronous standbys, but arbitrary follower selection or an older transaction snapshot can invalidate a broader inference. [S149] [S66] - Reader routing, transaction isolation and failover authority must be reviewed together. Routing to a machine called “leader” is insufficient if stale authority is possible; Chapter 34 develops that boundary.
33.5.6 WORDS — remember these#
Read-your-writes: seeing your own completed changes — a session guarantee that later eligible reads include the session’s prior acknowledged writes. Monotonic reads: not going backward in observed progress — a session constraint preventing later reads from using an older relevant state. Causal token: evidence of a required prior position — a scoped marker used to ensure that a read includes specified earlier work.
33.6 Replication is not a restore plan#
33.6.1 PLAIN — in simple words#
- A replica normally follows changes, including mistaken deletions and harmful updates. Copying a mistake faithfully is successful replication, not successful recovery.
- A backup preserves a state for a recovery purpose. Retained logs may let that state be advanced to a chosen point before a mistake, subject to the supported recovery procedure.
- Both mechanisms can be valuable. Neither replaces an inventory of failure scenarios, protected copies and tested restoration steps.
33.6.2 PLAIN — a picture in your head#
- If Mira crosses out the wrong page and Dev copies the crossing-out, both notebooks now contain the same mistake.
- An earlier protected copy may help reconstruct what existed before the error. A second notebook that always follows the first does not provide that historical boundary by itself.
- Where the comparison breaks: restoring a database also involves transaction histories, credentials, external effects and compatible software. It is not merely undoing ink on a page.
33.6.3 PLAIN — a worked example#
- A synthetic deletion is accepted at sequence 42. Both leader and follower eventually apply 42 and no longer contain the deleted record.
- A protected snapshot at 40 plus complete retained changes through 41 can support a separate recovery exercise stopping before 42. A snapshot without the needed log or encryption key may not.
- If the restored service replays old outbox messages without controls, it can repeat downstream actions even when the database restoration itself is correct.
- A recovery plan therefore verifies data, identity, access and external-effect policy before reconnecting the restored instance to ordinary traffic.
33.6.4 PLAIN — what is really happening inside#
- Replication reduces some recovery delays by maintaining a prepared copy. Backup retention protects selected historical states against different failures.
- Separate administrative and storage failure domains where the requirements demand it. A single compromised authority able to erase the primary, replicas and all backups defeats mere copy counting.
- Rehearse both failover and restoration. Their procedures, evidence and acceptable data-loss windows are not the same.
33.6.5 TECHNICAL — the engineer’s version#
- Distinguish replica promotion from point-in-time recovery. PostgreSQL’s continuous-archiving procedure requires compatible base data and sufficient WAL coverage for the selected recovery target. [S131]
- Backup verification should include actual restore rehearsals, not only manifest checking. A healthy replication stream does not establish that a historical recovery point is usable. [S138]
- State the tested failure scope. The book’s prefix model and local backup tests are educational evidence, not a certification of multi-host durability, operational failover or off-site disaster recovery.
33.6.6 WORDS — remember these#
Promotion: making a replica an active authority — a role transition requiring a supported procedure and protection against competing writers. Historical recovery point: a selected earlier usable state — a state reconstructible from compatible retained recovery material. Recovery rehearsal: trying the recovery procedure in isolation — a test of restoration steps, evidence and operational assumptions before an incident.
33.97 Practice and worked answers#
- A follower has stored position 41 but applied 40. Can it necessarily answer a query requiring 41? No. Durable receipt and query visibility are separate milestones.
- A standby’s last replay timestamp is an hour old. Is it definitely behind? No. First establish whether the source produced any later work and compare compatible progress evidence.
- A 20 MB backlog is processed at 8 MB/s while the source adds 4 MB/s. What is the simple catch-up estimate? The net rate is 4 MB/s, so approximately 5 seconds under these fixed-rate assumptions.
- The leader waited for follower A. May a later query use follower B without another check? Only if B independently satisfies the required read contract. A’s acknowledgement is not B’s progress.
- Does copying table rows through PostgreSQL 17 logical replication copy every sequence and DDL operation? No. Those boundaries require a separate compatibility and promotion plan.
- A mistaken deletion appears on all replicas. Did replication fail? Not necessarily. It may have reproduced the accepted operation correctly; recovery from the mistake is a different responsibility.
- What should happen if a required-fresh follower read reaches its deadline? Return the documented timeout, route to an authorized alternative that meets the requirement, or expose an explicitly weaker permitted mode. Do not conceal the downgrade.
- What does the companion model prove? Its contiguous-prefix, stale-read and minimum-position rules behave as tested. It does not execute a PostgreSQL cluster or model every network and storage failure.
33.98 Common wrong ideas#
- Wrong: more copies automatically mean more truth. Right: replicas can faithfully repeat incorrect source data.
- Wrong: received means readable. Right: receipt, durable storage and application are distinct stages.
- Wrong: synchronous means every replica is current. Right: acknowledgement follows a configured scope and milestone.
- Wrong: low byte lag proves every important record is fresh. Right: a small missing change may be operationally critical.
- Wrong: a connected replica is caught up. Right: it may be blocked or accumulating backlog.
- Wrong: logical replication copies every database object. Right: inspect supported object and schema coverage.
- Wrong: a replica eliminates backups. Right: current copies and protected recovery points serve different purposes.
- Wrong: a successful local model certifies failover. Right: real authority, routing, retention and failure handling remain separate validation work.
33.99 Chapter summary in 20 lines#
- Replication maintains additional copies for explicitly chosen purposes.
- Read scaling, high availability and historical recovery are different requirements.
- Physical replication follows engine-level representation and recovery rules.
- Logical replication follows identified data objects and changes.
- Initial copies and change streams need a compatible boundary.
- Schema and object coverage must be inspected rather than assumed.
- Asynchronous acknowledgement may precede remote progress.
- Synchronous acknowledgement waits for a named remote milestone.
- Receipt, flush and application are different milestones.
- Waiting for one follower does not establish another follower’s freshness.
- Lag can be measured at several stages of the replication path.
- Time since the last change is not a complete lag measurement.
- Catch-up requires progress faster than continuing input.
- Retained logs protect catch-up but consume storage.
- Read routing is part of the application’s consistency contract.
- Minimum-position reads require compatible history and snapshot handling.
- Timeouts must not conceal a weaker freshness guarantee.
- Replicas can faithfully reproduce harmful accepted operations.
- Backups and retained logs support separately tested historical recovery.
- State exactly which replication and recovery properties were demonstrated.