Caches, Queues and What They Do Not Guarantee
Introductions, exercises and summaries stay visible.
39.0 What this chapter gives you#
- Caches reduce repeated work by keeping derived copies. Queues separate the arrival of work from its processing. Both can improve a system while also introducing new states that must be interpreted correctly.
- You will trace stale cache fills, expiry, message acknowledgement, redelivery, ordering and backpressure. You will learn why fast retrieval or reliable transport is not the same as end-to-end business correctness.
- The companion exercises use deterministic cache and delivery models together with the local outbox/inbox example. They do not operate Redis or RabbitMQ servers or prove production messaging guarantees.
39.1 Derived copies#
39.1.1 PLAIN — in simple words#
- A cache stores an answer or copy so the system can avoid repeating expensive work. It is useful only when the cached result is still suitable for the next request.
- Suitability includes more than age. The result must belong to the right tenant, user permissions, query parameters and interpretation of the data.
- A cache hit means a matching cache entry was found. It does not prove that the underlying information is current, authorized or correct.
39.1.2 PLAIN — a picture in your head#
- Mira keeps a short list of frequently requested product descriptions beside the till. Reading the list is quicker than opening the full catalogue each time.
- The list must identify which catalogue edition it came from. A customer’s private discount or another branch’s restricted information must not be served simply because the product name matches.
- Where the comparison breaks: digital cache keys can include structured identity and version information. Their correctness depends on exactly which inputs the implementation includes, not on human recognition of the right paper.
39.1.3 PLAIN — a worked example#
- Suppose a cached report key contains only
daily-sales. Tenant A requests a report, and the cache stores A’s result under that key. - Tenant B then requests its own report and receives the same entry. The cache can be working exactly as implemented while the application violates the tenant boundary.
- A safer key specification includes the authorized tenant scope, report definition, date range, currency and relevant permission or data-version conditions. The exact fields follow the report contract.
- Adding a tenant identifier supplied freely by an untrusted caller does not solve authorization. The scope must be derived from or checked against trusted access context.
39.1.4 PLAIN — what is really happening inside#
- Cache-aside logic checks the cache, loads from the source on a miss and stores the result for later use. Each step can race with updates or invalidation.
- Cached data remains a derived view. The authoritative source and the policy for unavailable or stale entries must remain explicit.
- Cache failures can increase load on the source suddenly. A system sized only for a perfect hit rate may overload when the cache is cold or unavailable.
39.1.5 TECHNICAL — the engineer’s version#
- A cache key is part of the semantic and security contract. Include every input that changes the eligible result, or use a design that validates those differences before reuse.
- Cache-aside is a loading and invalidation pattern, not an automatic consistency guarantee. Its benefit depends on hit rate, access cost and tolerated staleness. [S161]
- The authoritative path must retain validation and authorization. A cache should not become an accidental alternate API that bypasses those checks.
39.1.6 WORDS — remember these#
Cache: a retained result or copy used to avoid repeated work — a derived store with explicit validity, scope and eviction rules. Cache hit: a lookup finds an eligible entry — an event that establishes presence, not automatically freshness or correctness. Cache-aside: the application loads missing data — a pattern in which code checks a cache and populates it from another source.
39.2 Invalidation and expiry#
39.2.1 PLAIN — in simple words#
- Invalidation tells the system not to reuse an old entry. Expiry removes eligibility after a specified interval. Neither should be confused with proof that every returned value is current.
- A delayed source read can arrive after an invalidation and refill the cache with an old value. The order of asynchronous events matters.
- A lifetime measured from cache insertion also does not establish the age of the source information. An already stale value can be freshly inserted.
39.2.2 PLAIN — a picture in your head#
- Dev copies a price from an old catalogue. While he is walking back to the till, Mira updates the catalogue and removes the old quick-reference card.
- Dev arrives afterwards and places his old copy into the now-empty slot. The invalidation happened, but a stale refill undid its intended effect.
- Where the comparison breaks: a correct implementation can use generation checks or ordering protocols. The walking story identifies the race; it does not supply a complete cache-coherence algorithm.
39.2.3 PLAIN — a worked example#
- Begin a separate cache fixture with source value 10 and cache generation 4. Reader R starts loading value 10 under generation 4.
- The source changes to 11, and invalidation advances the cache generation to 5. R’s delayed response then tries to store value 10.
- A generation-checked fill rejects that insertion because its captured generation 4 no longer matches 5. A naive unconditional fill would store the stale value.
- This prevents the specific stale-refill race in the model. It does not guarantee that R’s own already obtained response is fresh; the caller’s read contract may require a retry or authoritative check as well.
39.2.4 PLAIN — what is really happening inside#
- The cache needs a rule linking in-flight loads with invalidation. A placeholder or generation can let a delayed result detect that its right to populate the cache has expired.
- Lost invalidation connections create another uncertainty. A client that cannot know which entries changed must follow its documented flush, reconnect and freshness policy.
- Expiry bounds cache residence under its timer assumptions. It is useful but is not equivalent to linearizability, authorization freshness or a guaranteed bound on the age of an already stale source replica.
39.2.5 TECHNICAL — the engineer’s version#
- Redis’s client-side caching documentation explicitly describes invalidation arriving before a delayed GET response and a placeholder-based way to avoid retaining that obsolete response. It also discusses flushing on loss of the invalidation connection. [S160]
- The book’s generation model is an original simplification: invalidation increments a local generation, and a fill succeeds only for the captured current generation. It models one cache boundary, not Redis’s full protocol.
- A TTL, invalidation signal or generation check must be evaluated against the promised read semantics. Preventing one cache race does not prove a complete distributed freshness guarantee.
39.2.6 WORDS — remember these#
Invalidation: mark a cached result unusable — a signal or state transition preventing further eligible reuse. Time to live: an entry’s configured remaining lifetime — an expiry policy that does not by itself prove source-data age. Stale refill: an old in-flight result repopulates a cache — a race occurring after an update or invalidation unless the fill protocol prevents it.
39.3 Queue acknowledgement#
39.3.1 PLAIN — in simple words#
- A queue holds work for later processing. Sending a message, the broker accepting it, a worker receiving it and the business action completing are different events.
- An acknowledgement covers a particular boundary. A publisher confirmation does not automatically mean that a consumer completed the task.
- A worker should acknowledge according to the intended processing contract. Acknowledging before a required durable change can lose work if the worker then fails.
39.3.2 PLAIN — a picture in your head#
- Mira hands a delivery request to a dispatch desk. The desk stamps “received.” That stamp does not mean the courier has delivered the parcel.
- The courier’s later completion report answers a different question. Confusing the two can turn “accepted for processing” into a false “delivered” message.
- Where the comparison breaks: brokers have precise acknowledgement modes, durability settings and queue types. A stamp is not a substitute for their documented guarantees.
39.3.3 PLAIN — a worked example#
- Message M1 asks a consumer to create a local fulfillment task. In unsafe schedule A, the consumer acknowledges M1 first and crashes before committing the task. The broker may now consider the delivery handled even though no task exists.
- In schedule B, the consumer commits its inbox record and local task together, then acknowledges. If it crashes after commit but before acknowledgement, M1 may be delivered again.
- The unique inbox identity lets the consumer recognize that repeat and avoid creating a second task. The two schedules trade lost work for replay that the application must handle correctly.
- The local task is not a physical delivery. Its existence does not prove that a courier acted or a customer received anything.
39.3.4 PLAIN — what is really happening inside#
- The publisher and consumer interact with different broker boundaries. A lost publisher confirmation can lead to repeated publication; a lost consumer acknowledgement can lead to repeated delivery.
- Durable queue configuration and message persistence need to match the required failure model. A successfully written client socket is not a broker durability acknowledgement.
- The application must decide what “processed” means before acknowledging. It may mean a durable local task, not completion of every later external effect.
39.3.5 TECHNICAL — the engineer’s version#
- RabbitMQ documents publisher confirms and consumer acknowledgements as separate, unaware mechanisms. Confirms cover publishing to the broker; consumer acknowledgements cover broker-to-consumer delivery processing. [S158]
- Consumer delivery tags are scoped to their channel and are not stable business request identities. Use an application message identity for idempotent processing across redelivery and reconnects. [S158]
- An inbox record and local effect can share one database transaction. The broker acknowledgement remains outside that local transaction, so the consumer must tolerate the corresponding replay window.
39.3.6 WORDS — remember these#
Publisher confirm: the broker acknowledges publication — evidence about the publisher-to-broker boundary rather than consumer completion. Consumer acknowledgement: a worker reports a delivery handled — a protocol action whose timing must match the application’s processing contract. Inbox record: retained evidence of an accepted incoming message — a local identity and outcome record used to recognize duplicates and conflicts.
39.4 Redelivery and order#
39.4.1 PLAIN — in simple words#
- Reliable messaging often includes the possibility of another delivery after failure or uncertainty. A repeated message is not necessarily a second business event.
- Ordering also has a scope. Publishing order, delivery order and completion order can differ, especially with several workers or retries.
- An application that depends on sequence must enforce or validate that sequence rather than assume that the word “queue” guarantees every kind of ordering.
39.4.2 PLAIN — a picture in your head#
- Two workers take tasks numbered 1 and 2. Worker 2 finishes quickly while worker 1 is delayed. The desk handed them out in order, but completion happened in the opposite order.
- If task 1 is later returned for retry, it can appear again after other work. The workers need identities and state rules, not only a line to stand in.
- Where the comparison breaks: brokers expose specific ordering and consumer features. The analogy shows why those features must be checked, not that every queue behaves identically.
39.4.3 PLAIN — a worked example#
- A synthetic entity receives event E1, “create reservation,” followed by E2, “cancel reservation.” Worker A takes E1 and pauses; worker B takes E2 and runs first.
- A consumer that treats “reservation missing” as “cancellation completed forever” can later create the reservation when E1 resumes. The final state contradicts the intended event order.
- Possible designs include per-entity ordered processing, a state machine that retains pending cancellation evidence, or version checks that reject obsolete changes. The appropriate choice depends on the contract.
- A separate duplicate test delivers the same message identity twice. It should produce one local effect, while two distinct valid event identities remain two events even if their payloads look similar.
39.4.4 PLAIN — what is really happening inside#
- Concurrency can preserve enqueue order while changing completion order. Redelivery can reintroduce older work after newer work has already run.
- Product features such as single-active-consumer modes or ordered streams have documented limits. They do not automatically order all database actions or external effects performed by an application.
- A message identity reused with a different semantic payload is a conflict to investigate, not a duplicate to ignore silently. The identity contract must protect meaning as well as count.
39.4.5 TECHNICAL — the engineer’s version#
- RabbitMQ documents ordering within particular publishing and queue conditions and explains how multiple consumers, priorities and redelivery affect effective order. [S159]
- Its acknowledgement documentation states that unacknowledged deliveries can be requeued after channel or connection closure. Consumers must be designed for redelivery. [S158]
- Sequence guarantees should be scoped by entity, partition, channel or stream as appropriate. Application state transitions and external effects must still preserve the required causality.
39.4.6 WORDS — remember these#
Redelivery: a message is delivered again — a transport event that must be distinguished from a new business event. Completion order: the sequence in which effects finish — an order that may differ from publishing or delivery order. Per-entity ordering: related changes follow one entity’s sequence — a scoped requirement rather than a claim of global ordering across all work.
39.5 Backpressure#
39.5.1 PLAIN — in simple words#
- A queue can absorb a temporary burst, but it cannot make a permanently overloaded worker fast enough. If work arrives faster than it leaves, the backlog grows.
- Backpressure makes that pressure visible upstream through limits, slowing, rejection or controlled admission. It prevents an unbounded pile from silently consuming all available memory or disk.
- Retry traffic also counts as work. A failing dependency can trigger enough retries to worsen the overload that caused the failures.
39.5.2 PLAIN — a picture in your head#
- A dispatch desk receives 1,000 parcels each minute but can process only 600. A larger room delays the moment it fills; it does not close the 400-parcel gap.
- The desk needs more sustainable capacity, less admitted work or a clear policy for excess demand. Hiding the queue behind a door changes none of those facts.
- Where the comparison breaks: digital work varies in size and cost, and some can be safely combined or discarded under a stated policy. A parcel count alone may not represent resource demand.
39.5.3 PLAIN — a worked example#
- A synthetic queue receives 1,000 messages per second while consumers finish 600. The net growth is 400 messages per second.
- At 1,000 bytes per message, payload alone grows by 400,000 bytes per second, or 24 MB per minute using decimal units. Broker metadata, replication and indexing add further storage.
- After ten minutes, the simple backlog is 240,000 messages and 240 MB of payload, assuming fixed rates and no other losses or processing changes.
- A queue length limit can bound storage, but its overflow policy must be explicit. Silently dropping required reservation work is not a valid capacity solution.
39.5.4 PLAIN — what is really happening inside#
- Prefetch or in-flight limits cap how much unacknowledged work a consumer receives. Too much can create memory pressure and long recovery delays; too little can leave capacity idle.
- Admission control should consider deadlines and downstream health. Work that cannot possibly finish before its useful deadline may need an explicit rejection rather than a hidden hour-long wait.
- Monitor queue age as well as length. Ten expensive messages can be more urgent than a thousand tiny ones, and the oldest pending request may reveal starvation hidden by average throughput.
39.5.5 TECHNICAL — the engineer’s version#
- In a simple stable-rate model, backlog growth is arrival rate minus service completion rate when that difference is positive. Real workloads require distributions, cost variation and retry accounting.
- RabbitMQ prefetch limits constrain outstanding unacknowledged deliveries under their documented scope; they are not a complete admission-control or business-deadline policy. [S158]
- Combine bounded queues, retry budgets, idempotency and observability. Record any overflow, dead-letter or expiration disposition so the application can distinguish delayed work from permanently unprocessed work.
39.5.6 WORDS — remember these#
Backpressure: downstream limits affect upstream admission — a mechanism preventing work from accumulating without bound. Prefetch: how much unacknowledged work a consumer may receive — a broker/client flow-control setting with a defined scope. Queue age: how long work has waited — a latency and starvation signal distinct from the number of queued messages.
39.6 End-to-end correctness#
39.6.1 PLAIN — in simple words#
- A database, cache and queue each provide limited guarantees. The application must connect them into a valid story from request to user-visible outcome.
- A message can be durably queued while a cache shows an old state. A consumer can finish a local task while the external action remains pending. These states need honest labels.
- The final question is not whether every component is individually “reliable,” but whether the complete operation preserves its intended meaning across every boundary.
39.6.2 PLAIN — a picture in your head#
- A parcel journey includes acceptance at the desk, loading onto a vehicle, arrival at a branch and delivery to the customer. Each stamp is evidence for one stage.
- The tracking screen should not label the first stamp “delivered.” Nor should a missing final stamp automatically trigger shipment of a second identical parcel.
- Where the comparison breaks: the system needs explicit identities, state transitions and reconciliation, not merely a longer list of status words.
39.6.3 PLAIN — a worked example#
- A local reservation and outbox message commit together. The publisher sends the message; the consumer atomically records its inbox identity and creates one local task.
- The publisher then fails before marking the outbox row delivered. On retry, the same message is published again. The consumer recognizes its identity and returns the prior result without a second local task.
- A cache invalidation can still be delayed. The status endpoint must use an eligible authoritative or sufficiently fresh path rather than interpreting a stale cache miss as proof that the reservation never existed.
- The companion test records two delivery attempts and one local task. It does not claim exactly-once email, physical dispatch or execution at an uncontrolled external provider.
39.6.4 PLAIN — what is really happening inside#
- Define each boundary: request acceptance, transaction commit, publication, consumer commit, external acknowledgement and user-visible confirmation.
- Test failures between boundaries rather than only the uninterrupted happy path. Unknown outcomes, repeated messages and stale reads are normal design cases.
- Keep a reconciliation path that can connect the request, outbox, inbox and resulting task without exposing unnecessary private data. Evidence should explain progress, not merely generate large logs.
39.6.5 TECHNICAL — the engineer’s version#
- The outbox pattern records local intent atomically with business state, while idempotent consumers limit repeated local effects. Their scope must remain explicit. [S125] [S126]
- Cache coherence, broker acknowledgement and database atomicity are independent mechanisms. None can be promoted into an end-to-end guarantee without checking how the application composes them. [S158] [S160]
- The next part follows data into analytical and search systems. Those derived systems inherit the same obligations: identity, progress boundaries, replay handling and honest interpretation of incomplete results.
39.6.6 WORDS — remember these#
End-to-end correctness: the complete operation preserves its contract — a property spanning components rather than inferred from any single successful API call. Boundary evidence: proof about a particular stage — an acknowledgement or record whose scope must not be exaggerated. Dead-letter disposition: work is moved aside after unsuccessful processing — a recorded outcome requiring review rather than automatic proof that the business task finished.
39.97 Practice and worked answers#
- Why is a cache key containing only a report name dangerous? Different tenants, parameters or permissions can produce different eligible results. Reusing one entry can leak or misrepresent data.
- A delayed value loaded under generation 4 arrives after invalidation to 5. Should it populate the model cache? No. Its generation is obsolete.
- Does rejecting that fill prove the original caller received current data? No. The caller’s read result needs its own freshness policy.
- What does a publisher confirm prove about consumer completion? Nothing by itself. It covers a separate publisher-to-broker boundary.
- Why commit an inbox record and local effect before acknowledging? A later redelivery can then be recognized without losing the local task or duplicating it.
- Can FIFO delivery guarantee FIFO completion with two workers? No. Processing duration and retries can change completion order.
- What is the payload backlog after ten minutes at a 400-message/s deficit and 1,000 bytes/message? 240,000 messages and 240 MB, excluding overhead, under the fixed-rate assumptions.
- What does the repeated outbox delivery test establish? Two attempts produce one local consumer task under the tested inbox contract. It does not establish exactly-once external effects.
39.98 Common wrong ideas#
- Wrong: a cache hit proves freshness. Right: it proves an entry was found under the implemented eligibility rule.
- Wrong: TTL alone gives linearizable reads. Right: expiry is not a complete coherence or source-freshness protocol.
- Wrong: invalidation cannot be undone by a delayed read. Right: stale refill is a real ordering hazard.
- Wrong: broker acceptance means the task finished. Right: publication, delivery and processing have separate acknowledgements.
- Wrong: a repeated message is always a new event. Right: message identity and semantic intent determine replay handling.
- Wrong: queue order means every effect completes in order. Right: parallel consumers and retries can change completion order.
- Wrong: a large queue solves permanent overload. Right: it only postpones resource exhaustion when arrival exceeds service.
- Wrong: reliable components automatically compose into exactly-once business behaviour. Right: the complete boundary and recovery design must establish the intended outcome.
39.99 Chapter summary in 20 lines#
- Caches retain derived results to avoid repeated work.
- Cache validity includes scope, permissions, parameters and freshness.
- A hit is not proof that the source information is current.
- Invalidation and expiry are different mechanisms.
- Delayed reads can create stale refills.
- Generation or ordering checks can prevent specific refill races.
- Losing an invalidation channel requires explicit recovery policy.
- Queue publication, delivery and processing are separate stages.
- Publisher confirms do not establish consumer completion.
- Consumer acknowledgements must match the intended durable processing boundary.
- Inbox identity and local effects can share one transaction.
- Redelivery must be distinguished from a new business event.
- Publishing, delivery and completion order can differ.
- Per-entity sequence rules require explicit enforcement.
- Backlogs grow when arrival exceeds sustainable completion.
- Prefetch and admission limits bound different kinds of outstanding work.
- Queue age and overflow dispositions must remain observable.
- Outbox retries can coexist with one idempotent local consumer effect.
- External actions and cache freshness remain separate obligations.
- End-to-end correctness requires evidence across the whole request path.