Skip to content
KEDBYTE
Site navigation
How Data Works
Chapter
50

Observability, Capacity and Cost

Part H · Running Data Systems|3,922 words|about 17 min read|Volume H

50.0 What this chapter gives you#

  1. A running system produces signals about its behaviour. Useful observability connects those signals to questions: Are users completing their work? Where is time being spent? Which resource is approaching a limit? What changed before the failure?
  2. You will distinguish logs, metrics and traces; calculate a capacity forecast with explicit assumptions; and separate workload measurements from guesses based on server labels.
  3. Every rate, storage size and cost in the worked examples is invented for arithmetic. None is a measurement of KedByte’s live products, a current hosting price or a capacity promise for a particular machine.
  4. This chapter connects Chapter 20’s measurement discipline with Chapter 46’s metric definitions. A precise dashboard can still answer the wrong question when it misses failed attempts or mixes incompatible populations.

50.1 Signals tied to user outcomes#

50.1.1 PLAIN — in simple words#

  1. Start with the work the service is supposed to complete. A healthy-looking processor graph does not prove that an order was recorded correctly or that a customer can retrieve it.
  2. Define success, failure and unknown outcomes. A request that timed out may have committed or may not have reached the database; the timeout alone does not settle the stored result.
  3. Measure the experience at a useful boundary. Server processing time excludes some network and client delays, while a browser measurement may include them.
  4. An objective is a target, not an observation. The system must collect evidence to determine whether the target was met.

50.1.2 PLAIN — a picture in your head#

  1. Mira’s delivery van has a working engine and a full fuel tank. Those facts are useful, but customers care whether the correct parcels arrive on time.
  2. Engine temperature helps explain a delay; delivery completion tells her whether the service actually fulfilled its purpose.
  3. Where the comparison breaks: digital work may complete invisibly after a timeout. Outcomes need stable identities and reconciliation, not just a single caller’s immediate impression.

50.1.3 PLAIN — a worked example#

  1. Define a synthetic request objective: at least 99.5% of 20,000 eligible read requests return an authorised, valid response within a stated threshold during one observation window.
  2. Under that definition, the allowance for nonconforming requests is 0.5% × 20,000 = 100. If 60 are nonconforming, the measured success proportion is 19,940 / 20,000 = 99.7%, with 40 requests of the illustrative allowance remaining.
  3. These calculations depend on the eligibility and observation rules. Excluding failed requests because they lack a successful response record would change the denominator and could make the result misleading.
  4. Record unknown outcomes separately. A claim of 99.7% cannot silently treat unobserved attempts as successful merely because no error reached the dashboard.

50.1.4 PLAIN — what is really happening inside#

  1. A service-level indicator operationalises a user-facing property through measured events or values. A service-level objective selects a target for that indicator over a defined window.
  2. Collection occurs at a particular boundary. Load balancers, application handlers and database drivers may each see different populations and durations.
  3. Reconcile those populations when interpreting a failure. An attempt can disappear before application logging, and a database commit can occur before its acknowledgement is lost.

50.1.5 TECHNICAL — the engineer’s version#

  1. Google’s SRE text distinguishes indicators, objectives and agreements and emphasises choosing indicators that represent useful service behaviour. The book’s 20,000-request example is our own arithmetic, not Google’s reliability target. [S197]
  2. Document event eligibility, correctness checks, measurement boundary, threshold, window, missingness and aggregation. Include the specification version in the metric definition.
  3. Error-budget arithmetic does not permit violating a data-integrity invariant. A latency objective and a rule against cross-tenant disclosure are different kinds of requirement, not interchangeable percentages.

50.1.6 WORDS — remember these#

  1. Service-level indicator: a measured service property — a defined observation of availability, latency, correctness or another relevant outcome. Service-level objective: the target for an indicator — a specified level over a stated population and window. Error budget: the allowed nonconformance under an objective — an arithmetic planning quantity whose definition does not override other requirements.

50.2 Logs, metrics and traces#

50.2.1 PLAIN — in simple words#

  1. A log records an event with details. A metric summarises quantities over time or categories. A trace connects steps belonging to a particular operation across components.
  2. Each helps answer different questions. A count shows how often something happened; a trace can show where one request waited; a log can explain a particular rejection.
  3. Collect enough context to connect evidence without copying unnecessary private data. Request identifiers are often more useful than complete request bodies.
  4. Missing signals need explanation. An absent log can mean nothing happened, collection failed, the event was sampled out or the retention window expired.

50.2.2 PLAIN — a picture in your head#

  1. Mira has a daily counter of orders, a diary of unusual incidents and a route sheet following one parcel through packing and delivery.
  2. The counter cannot explain one damaged parcel. The route sheet cannot by itself show the whole month’s failure rate. The tools complement each other.
  3. Where the comparison breaks: digital signals can be sampled, buffered, delayed and dropped. Their apparent timing and completeness depend on the collection system.

50.2.3 PLAIN — a worked example#

  1. A synthetic request takes 240 milliseconds end to end. Its trace records 20 milliseconds before the application, 40 milliseconds in a queue, 130 milliseconds in database work and 50 milliseconds in other application/response work. The stated non-overlapping intervals sum to 240.
  2. Real spans can overlap, so blindly adding every displayed span duration can double-count time. The example’s disjoint intervals are an explicit teaching assumption.
  3. The associated metric counts one eligible request and its outcome. A rejection log stores the scoped request identifier and reason category, not the secret credential or full customer note.
  4. If only one in ten traces is retained, the retained traces cannot automatically supply an exact count of all requests. Use a suitable unsampled counter or a documented estimator instead.

50.2.4 PLAIN — what is really happening inside#

  1. Instrumentation creates events and measurements in running code. Collectors transport and process them; storage and query systems retain and expose them.
  2. Correlation identifiers let investigators connect related signals. Their propagation needs care at trust boundaries, because external callers can supply misleading identifiers unless the system distinguishes trusted and untrusted context.
  3. Labels determine how metrics are divided. Adding every order ID as a metric label can create a huge number of distinct time series; detailed per-order evidence may belong in a different store with appropriate access and retention.

50.2.5 TECHNICAL — the engineer’s version#

  1. OpenTelemetry describes signals including traces, metrics and logs and their associated data models. Adopting its vocabulary or library does not automatically establish complete instrumentation or causal diagnosis. [S198]
  2. Record sampling, buffering, export failure and clock assumptions. A timestamp difference across unsynchronised hosts is not necessarily a precise duration; use suitable local duration measurements and trace semantics.
  3. Google’s monitoring guidance identifies latency, traffic, errors and saturation as useful general signals. Our observation design must still connect those signals to the shop’s actual operations and declared failure conditions. [S103]

50.2.6 WORDS — remember these#

  1. Log: a recorded event — timestamped or otherwise ordered details describing a particular occurrence. Metric: a measured quantity — an aggregated or sampled numeric signal with units and dimensions. Trace: connected steps of an operation — spans and context describing work across a request path. Cardinality: how many distinct labelled combinations exist — a property that can drive metric storage and query cost.

50.3 Storage and growth#

50.3.1 PLAIN — in simple words#

  1. Storage demand depends on how much arrives, how large each retained representation is and how long it remains. Indexes, logs, copies and temporary work add to the primary records.
  2. A count of students, shops or users does not determine storage by itself. One user can create thousands of records, large attachments or almost no activity.
  3. Distinguish logical payload size from physical occupied storage. Compression, row overhead, fragmentation and retention can change the latter substantially.
  4. A forecast should expose its assumptions and define when measurements will replace them. A neat multiplication is not a measured capacity result.

50.3.2 PLAIN — a picture in your head#

  1. Mira estimates shelf space from boxes arriving per day, the space per box and the days each box remains. Counting suppliers alone is not enough.
  2. She also needs room for labels, aisles, returns and temporary packing. Those are not new sales, but they still occupy the building.
  3. Where the comparison breaks: database compression and indexes do not scale exactly like cardboard boxes. Updates and maintenance can create temporary or retained representations whose size depends on engine behaviour.

50.3.3 PLAIN — a worked example#

  1. Invent a workload of 12,000 records per day, each with 600 bytes of logical payload, retained for 365 days. The payload-only estimate is 12,000 × 600 × 365 = 2,628,000,000 bytes, or 2.628 decimal GB.
  2. Under a deliberately hypothetical 1.8× combined table/index overhead multiplier, the active representation estimate becomes 4.7304 GB. Three active copies would total 14.1912 GB under the same simplifying assumptions.
  3. This does not include backups, logs, exports, attachments, temporary sort space or free capacity for maintenance. Do not add them through an unexplained universal percentage; estimate their mechanisms separately.
  4. If a host has 100 GB usable and 70 GB is the chosen intervention threshold, starting at 20 GB with 2 GB/day net growth reaches that threshold in (70−20)/2 = 25 days. The 70 GB threshold is an example policy, not a universal safe limit.

50.3.4 PLAIN — what is really happening inside#

  1. Physical storage grows through inserts, updates, index maintenance and retained history. Deletion may make space reusable without immediately shrinking the file.
  2. Copies have different lifetimes. A follower, nightly backup and exported report cannot be modelled accurately as three identical continuously updated files.
  3. Measure representative data and maintenance cycles. A fresh empty database can understate steady-state overhead; an unusual bulk load can overstate typical daily growth.

50.3.5 TECHNICAL — the engineer’s version#

  1. Keep byte units explicit: decimal GB is 10^9 bytes and GiB is 2^30 bytes. The forecast above uses decimal units throughout.
  2. PostgreSQL’s statistics views expose specific counters and activity information under documented visibility and timing rules. They are observations with semantics, not a direct prediction of next year’s data size. [S200]
  3. A useful forecast records measured baseline, growth distribution, retention action, replication factor, backup policy and uncertainty. Update it when the workload mix changes rather than treating the first spreadsheet as permanent truth.

50.3.6 WORDS — remember these#

  1. Logical payload: the data’s application representation — the meaningful values before engine and operational overhead. Net growth: increase after relevant additions and removals — a rate whose measurement window and storage boundary must be stated. Intervention threshold: the point to act before exhaustion — a chosen policy leaving time and resources for a controlled response.

50.4 Saturation and headroom#

50.4.1 PLAIN — in simple words#

  1. A system is saturated when demand pushes a relevant resource toward the point where more work mostly creates waiting, rejection or failure.
  2. The bottleneck might be CPU, storage latency, memory, a connection pool, a lock or a slow external dependency. Buying more of a different resource may not help.
  3. Headroom is spare usable capacity under the chosen workload and failure assumptions. It is not simply the percentage of CPU currently idle.
  4. Repeated retries can increase demand while the system is already struggling. A recovery policy needs limits, not an instruction to retry everything immediately forever.

50.4.2 PLAIN — a picture in your head#

  1. A counter can serve customers at a finite rate. Once arrivals exceed that rate, the queue grows even though the assistant is working continuously.
  2. Asking waiting customers to rejoin the queue every few seconds makes the crowd larger without producing more completed orders.
  3. Where the comparison breaks: software can parallelise some work and has multiple coupled queues. A single service-rate number is a model, not a universal description of the whole architecture.

50.4.3 PLAIN — a worked example#

  1. A synthetic worker completes 100 jobs per second, while 120 arrive per second for 30 seconds. Ignoring variability and starting empty, the queue grows by (120−100) × 30 = 600 jobs.
  2. After arrivals stop, draining those 600 jobs at 100 per second takes 6 seconds under the same assumptions. If arrivals instead continue at 90 per second, net drain is only 10 per second and the backlog takes 60 seconds.
  3. For a stable observation window, Little’s-law arithmetic relates average work in the system L, throughput λ and average time W as L=λW, under the relevant steady-state conditions. At 80 jobs/second and 0.25 seconds average residence, L=20 jobs. This is not a p99 latency estimate.
  4. If retries add 50 attempts per second to the original 120 while capacity stays 100, the backlog grows faster. Backoff, retry budgets and admission control address this feedback, but each needs an explicit business failure policy.

50.4.4 PLAIN — what is really happening inside#

  1. Requests wait when a needed resource is unavailable. Queueing changes latency distributions and can consume additional memory or connection capacity.
  2. Rejecting or shedding some work can protect the rest, but the service must decide which operations may be rejected and how callers learn the outcome. An order with unknown commit status cannot be retried as a brand-new order without considering identity.
  3. Capacity during one-node failure can be lower than normal capacity. Headroom calculations should include the failure state the service claims to tolerate.

50.4.5 TECHNICAL — the engineer’s version#

  1. Google’s overload discussion covers load shedding and mechanisms for avoiding overload amplification. The queue arithmetic here is an original deterministic model, not an experiment on Google’s systems. [S199]
  2. Throughput and latency should be measured together under a specified arrival process. Closed-loop clients that wait for responses can reduce their offered load as the server slows, hiding a production arrival-rate problem.
  3. Report accepted, completed, rejected, timed-out and unknown attempts separately. A benchmark that silently discards failed requests may make overload appear faster rather than less reliable.

50.4.6 WORDS — remember these#

  1. Saturation: demand meeting a limiting resource — a condition producing increased waiting, rejection or reduced useful throughput. Headroom: usable spare capacity — margin under stated workload and failure assumptions, not a generic machine percentage. Admission control: deciding which work to accept — protection of bounded resources through explicit entry policies. Retry amplification: repeated attempts adding load during trouble — feedback that can worsen the condition callers are trying to recover from.

50.5 Cost attribution#

50.5.1 PLAIN — in simple words#

  1. Cost belongs to resources and work over time. A monthly server price alone does not describe the full cost of a data service.
  2. Include storage, backups, transfer, monitoring, support and operational effort where they matter. State what is excluded rather than hiding it in an apparently precise cost per user.
  3. Shared infrastructure needs an allocation rule. Equal division, usage-based division and capacity-reservation division answer different questions.
  4. A cheaper architecture is not necessarily cheaper for the required outcome. A design that saves storage but makes restoration too slow may fail its purpose.

50.5.2 PLAIN — a picture in your head#

  1. Mira shares a delivery van with another shop. Dividing the bill equally is simple, but one shop may use most of the route time.
  2. Charging by parcel count can still be misleading when one shop’s parcels are much larger. The allocation method needs to match the decision being made.
  3. Where the comparison breaks: computing has fixed commitments, burst usage, tiered prices and shared bottlenecks. A single per-request charge may not represent marginal cost or reserved capacity.

50.5.3 PLAIN — a worked example#

  1. Invent monthly amounts of 6,000 currency units for compute, 1,000 for storage, 800 for backups and 1,200 for observability. The sum is 9,000 before any omitted transfer, tax, support or labour costs.
  2. With 300,000 eligible completed operations, the allocated average is 9,000 / 300,000 = 0.03 units per completed operation under this chosen boundary.
  3. If only 200,000 operations complete while the fixed bill remains 9,000, the same allocated measure becomes 0.045. That does not mean each additional operation physically costs 0.045; average and marginal cost differ.
  4. A report should state the currency, accounting window, included charges, denominator and allocation method. These invented values are not quotes from a cloud provider or estimates for the user’s website.

50.5.4 PLAIN — what is really happening inside#

  1. Resource meters measure selected consumption. Billing rules convert it into charges, and an attribution model allocates those charges to tenants, workflows or products.
  2. A shared fixed cost cannot be uniquely divided by arithmetic alone; an allocation policy is required. The policy should be recorded so comparisons remain consistent.
  3. Optimisation should preserve required outcomes. Removing all backups lowers one bill line while removing the recovery property the system was supposed to provide.

50.5.5 TECHNICAL — the engineer’s version#

  1. Separate invoice reconciliation, allocated unit cost, marginal cost and opportunity cost. They answer different management questions and can legitimately produce different numbers.
  2. Evaluate changes using a fixed workload definition and correctness constraints. A query that becomes cheap by excluding required records is not an optimisation of the original question.
  3. The companion cost function checks nonnegative finite amounts and an explicit nonzero denominator. It does not connect to billing accounts or supply current market prices.

50.5.6 WORDS — remember these#

  1. Allocated cost: a share assigned under a rule — a reporting quantity based on an explicit apportionment policy. Marginal cost: the change caused by an additional unit of work — distinct from dividing the total bill by current volume. Cost boundary: which charges are counted — the stated inclusions and exclusions behind a reported amount.

50.6 Operational reviews#

50.6.1 PLAIN — in simple words#

  1. Review a data service as an ongoing responsibility, not a machine purchased once. Workload, dependencies, retention and permissions change.
  2. A useful review compares actual outcomes with objectives, checks forecasts against measurements and assigns concrete actions for gaps.
  3. Alerts need an owner and a response. A dashboard containing hundreds of red boxes but no decision path is not a complete operating method.
  4. Keep a record of what changed and what followed. An incident is a chance to improve the system’s evidence and controls, not an excuse to replace uncertainty with a confident story.

50.6.2 PLAIN — a picture in your head#

  1. Mira checks the van’s service record, delivery outcomes and next month’s expected workload. She does not wait for the engine to fail before discovering that no replacement driver is available.
  2. When a delivery goes wrong, she records the observed sequence and tests plausible explanations rather than blaming the nearest person automatically.
  3. Where the comparison breaks: distributed incidents can have several interacting causes. One graph changing first does not automatically establish it as the root cause.

50.6.3 PLAIN — a worked example#

  1. A synthetic weekly review contains: request success under its metric definition, p50/p95/p99 latency with sample counts, oldest queued job age, restore rehearsal age, growth versus forecast and unresolved lifecycle tasks.
  2. A storage forecast says the intervention threshold is 25 days away. Procurement or migration takes 14 days and the team reserves 7 days for uncertainty. That leaves only 4 days before the planned action should begin under these assumptions.
  3. A backup exists, but the last verified restore is old and uses a different application schema. The review records a new rehearsal task; it does not treat backup presence as current recovery proof.
  4. An alert for growing outbox age includes a runbook: inspect worker health and failures, check destination availability, preserve message identities and avoid blind replay with new keys.

50.6.4 PLAIN — what is really happening inside#

  1. A review closes the loop between measurements and decisions. Forecast errors change the next forecast; recurring incidents change tests, capacity or architecture.
  2. Operational evidence should retain its context: revision, environment, workload and observation time. Comparing two dashboards without those details can confuse a workload change with an implementation improvement.
  3. A runbook proposes bounded observations and safe actions. Destructive changes and broad privilege escalation need their own authority, not an implicit permission hidden in an alert.

50.6.5 TECHNICAL — the engineer’s version#

  1. Tie each alert to an actionable condition and owner. Separate symptoms affecting users from diagnostic signals that help explain them; both are valuable but serve different roles. [S103]
  2. Maintain explicit unknowns: untested failover, incomplete telemetry, obsolete restore evidence or unidentified growth sources. Assigning an owner makes an unknown manageable; renaming it green does not.
  3. The capstone’s operational worksheet uses original synthetic arithmetic. Its production extension requires measurements from the actual service rather than estimates inferred from user count alone.

50.6.6 WORDS — remember these#

  1. Runbook: a documented response path — observations, decisions, bounded actions and escalation conditions for an operational situation. Forecast error: the difference between prediction and observation — evidence used to update assumptions rather than hide uncertainty. Operational owner: the accountable responder — the person or team responsible for a signal, policy or unresolved task.

50.97 Practice and worked answers#

  1. Question: What is the nonconformance allowance at 99.5% over 20,000 eligible requests? Answer: 100 under the example’s exact definition.
  2. Question: Can all trace span durations simply be added? Answer: Not if spans overlap; the worked 240 ms decomposition explicitly uses disjoint intervals.
  3. Question: How much payload do 12,000 daily records of 600 bytes create over 365 days? Answer: 2,628,000,000 bytes before physical and operational overhead.
  4. Question: At 20 GB used, 2 GB/day growth and a 70 GB intervention threshold, how much time remains? Answer: 25 days under constant net growth.
  5. Question: How quickly does a 600-job backlog drain with capacity 100/s and ongoing arrivals 90/s? Answer: In 60 seconds under the deterministic assumptions, not 6 seconds.
  6. Question: Does L=λW estimate p99 latency? Answer: No. It relates suitable averages under steady-state conditions.
  7. Question: What is the allocated average of 9,000 units over 300,000 completed operations? Answer: 0.03 units per operation within the stated cost boundary.
  8. Question: Does an existing backup establish current recoverability? Answer: No. A relevant restore procedure must be tested and its outcome checked.

50.98 Common wrong ideas#

  1. Wrong: Low CPU proves healthy service. Right: User outcomes and other bottlenecks may disagree.
  2. Wrong: Missing errors mean all requests succeeded. Right: Observation coverage and unknown outcomes must be checked.
  3. Wrong: Logs, metrics and traces are interchangeable. Right: They preserve different kinds of evidence.
  4. Wrong: Number of users determines capacity. Right: Workload, data size, concurrency and objectives determine the relevant demand.
  5. Wrong: More retries always improve reliability. Right: Retries can amplify overload and duplicate effects.
  6. Wrong: A cost per operation is its marginal cost. Right: Allocated averages and incremental resource costs differ.
  7. Wrong: A forecast is a specification. Right: It is an assumption-dependent prediction that needs updating.
  8. Wrong: A dashboard replaces an operating team. Right: Signals need ownership, interpretation and bounded response procedures.

50.99 Chapter summary in 20 lines#

  1. Start observability with the user’s intended outcome.
  2. Define success, failure and unknown states.
  3. Keep indicator, objective and agreement distinct.
  4. Preserve the eligibility denominator and observation boundary.
  5. Use logs, metrics and traces for their different strengths.
  6. Correlate operations without copying unnecessary secrets or personal data.
  7. Account for sampling, delayed export and missing signals.
  8. Avoid unbounded metric-label cardinality.
  9. Forecast storage from rates, representation sizes and lifetimes.
  10. Separate payload from indexes, logs, copies and temporary work.
  11. Keep GB and GiB explicit.
  12. Set intervention thresholds early enough to act.
  13. Identify the actual bottleneck before buying resources.
  14. Measure throughput and latency under a declared arrival process.
  15. Bound retries and protect the service from overload amplification.
  16. State the failure assumptions behind headroom.
  17. Make cost inclusions and allocation rules explicit.
  18. Distinguish average cost from marginal cost.
  19. Turn forecast errors and incidents into revised actions and tests.
  20. Give every important unresolved operational condition an owner.

Return to contents