Measuring Performance Without Fooling Yourself
Introductions, exercises and summaries stay visible.
20.0 What this chapter gives you#
- A performance result is an observation about a defined workload, environment and measurement boundary. Without those definitions, a fast number can be true and still misleading.
- You will specify a benchmark, calculate latency percentiles under an explicit convention, separate throughput from latency, and report errors and limitations instead of selecting only attractive runs.
- The companion experiment is a local educational microbenchmark, not a capacity test of KedByte, a comparison of database brands or a claim about production hardware.
20.1 Define the workload#
20.1.1 PLAIN — in simple words#
- Before timing, say what the system must do correctly. Two queries that return different populations are not competing implementations of the same task.
- Specify data size and distribution, query parameters, selected fields, concurrency and the mix of reads and writes. “One million rows” alone leaves most of the workload undefined.
- Decide where the timer starts and stops. Measuring only server execution differs from measuring connection setup, result transfer, decoding and application work as well.
20.1.2 PLAIN — a picture in your head#
- Timing a runner requires a marked course and finish line. One runner cannot stop after 80 metres while another completes 100 and still be compared as though they ran the same race.
- A race with hurdles is also different from an empty track, even when the distance is equal.
- Where the comparison breaks: database workloads can include many simultaneous operations and changing state. There may be no single finish line for the entire service, so define per-operation and experiment-level boundaries separately.
20.1.3 PLAIN — a worked example#
- Our query experiment compares the same fixed product-like lookup on separately generated teaching rows before and after a candidate index. It checks the complete ordered result against an independent expected list.
- Dataset creation and index construction are recorded separately from repeated query execution. The query timer includes fetching all result rows into the local client process.
- Rare, common and absent keys form separate cases. A combined summary, when used, must specify their weighting rather than silently letting the fastest case dominate.
- The four canonical order lines remain a correctness fixture. They are not expanded by pretending the same four sales happened thousands of times in the story.
20.1.4 PLAIN — what is really happening inside#
- A workload generator determines arrivals and parameters. The database executes work, while the client observes successes, failures and durations within its timer boundary.
- A closed-loop generator waits for one response before issuing more work at a fixed number of clients. An open-loop generator schedules arrivals independently of previous completion. Under overload, they expose different queueing behaviour.
- If the generator slows down whenever the system slows down, it can omit the waiting that independently arriving users would experience. State the arrival model rather than claiming that every benchmark measures the same service demand.
20.1.5 TECHNICAL — the engineer’s version#
- A benchmark contract includes schema, indexes, data-generation code and seed, parameter distribution, operation mix, concurrency, arrival process, warm-up, timer boundary, sample count and correctness oracle.
- Use a monotonic timer for elapsed intervals. Python’s
perf_counter_ns()avoids floating-point timestamp subtraction and returns integer nanoseconds; its unit is not an accuracy guarantee. [S102] - Retain exact runtime and library versions. A reproducible dataset does not make two machines’ execution conditions identical, so reproducibility of inputs and portability of timing results are separate claims.
20.1.6 WORDS — remember these#
Benchmark contract: the rules of the measurement — a specification of inputs, operations, environment, timing boundaries and success criteria. Closed-loop load: clients wait before issuing their next request — a workload in which completion controls later arrivals. Open-loop load: arrivals follow an independent schedule — a workload whose offered demand does not automatically slow with each response.
20.2 Latency distributions#
20.2.1 PLAIN — in simple words#
- Latency is not one fixed property. The same operation can complete quickly most of the time and occasionally take much longer.
- An average compresses those differences. Percentiles describe positions in an ordered sample, but their calculation convention and sample size matter.
- Keep failed and timed-out attempts visible. Reporting only successful durations can make an overloaded system look faster by discarding its worst experiences.
20.2.2 PLAIN — a picture in your head#
- Nineteen customers wait ten seconds and one waits a thousand. Saying “the average wait was 59.5 seconds” is mathematically correct but does not describe either group’s experience well.
- Sorting the waits lets us ask what a chosen fraction of customers experienced.
- Where the comparison breaks: a small observed sample does not establish the long-run distribution. A percentile estimated from twenty requests is not reliable evidence about rare one-in-ten-thousand delays.
20.2.3 PLAIN — a worked example#
- Take an invented sample of twenty latencies: nineteen at 10 ms and
one at 1,000 ms. The mean is
(19×10+1,000)/20 = 59.5 ms. - This lab uses the nearest-rank convention: for fraction p with
0 < p <= 1, select sorted positionceil(pN), counting from one. The median under this convention is 10 ms; p95 selects position 19 and is 10 ms; p99 selects position 20 and is 1,000 ms. - Another percentile interpolation convention can give a different numerical p95. Label the convention instead of treating one library’s output as the only mathematically possible percentile.
- These are calculated teaching values, not measured database latencies. The benchmark’s raw observed samples are stored separately in its evidence output.
20.2.4 PLAIN — what is really happening inside#
- Latency combines active work and waits within the chosen boundary. Cache misses, locks, scheduling, storage and result size can create different parts of the distribution.
- Histograms and raw samples preserve more shape than a single mean. Histogram bucket boundaries limit the precision of derived quantiles; raw samples still have sampling uncertainty.
- A timeout is a censored observation: the client knows it waited at least until its deadline without obtaining the desired completion. Recording the deadline as though it were the actual completion time understates what is unknown.
20.2.5 TECHNICAL — the engineer’s version#
- State whether a percentile describes all attempts, successful attempts or another population. Retain attempt counts, success counts, timeout counts and error classes alongside duration summaries.
- Quantile convention, sample size and workload stationarity affect interpretation. Do not average p99 values from separate groups and call the result the combined p99; combine compatible underlying observations or histogram counts instead.
- User-oriented monitoring benefits from latency distributions together with traffic, errors and saturation, rather than one average alone. [S103]
20.2.6 WORDS — remember these#
Percentile: a position in an ordered population or sample — a quantile defined under a stated convention, such as nearest rank. Tail latency: the unusually slow end of the distribution — high-percentile delays whose meaning depends on the measured population and sample size. Censored observation: completion time is only partly known — an observation limited by a timeout or stopping rule rather than a measured final duration.
20.3 Throughput and saturation#
20.3.1 PLAIN — in simple words#
- Throughput counts completed work per unit time. Latency measures how long an operation takes. A system can complete many operations per second while each operation waits in a long queue.
- As demand approaches a bottleneck’s capacity, waiting can grow. Adding clients does not guarantee proportionally more completed work.
- Count useful, correct completions separately from errors or retries. A server returning failures very quickly has high response throughput but may provide little successful service.
20.3.2 PLAIN — a picture in your head#
- One checkout counter can serve customers only so quickly. More people joining the queue increase waiting, not necessarily the number served each minute.
- Opening a second counter helps only if the shared payment terminal, stock lookup or packing station is not the real bottleneck.
- Where the comparison breaks: computer workloads can overlap CPU, storage and network work, and requests have different costs. One checkout’s fixed service time is an intentionally simplified model.
20.3.3 PLAIN — a worked example#
- In a stable hypothetical system, 100 completed requests per second
spend an average 0.2 seconds within a defined service boundary.
Little’s-law reasoning gives an average
100 × 0.2 = 20requests inside that boundary, under its steady-state accounting assumptions. - That count includes any waiting represented by the latency boundary. It is not automatically twenty actively executing CPU tasks or twenty database connections.
- Suppose offered load increases to 200 requests per second while sustainable completion remains 100 and nothing rejects or slows arrivals. The queue grows by roughly 100 requests per second in this simplified interval.
- A bounded system must instead apply some policy: admission limits, backpressure, deadlines, rejection or additional capacity. Infinite queues are not a service-quality strategy.
20.3.4 PLAIN — what is really happening inside#
- Bottlenecks can move as load changes. CPU may dominate one workload, storage another, and locks or a connection pool another.
- Retry traffic can amplify overload. Counting retries as independent fresh business operations can also overstate useful throughput and create correctness problems.
- Plot or tabulate offered demand, successful completions, errors, queueing and latency together. A throughput plateau with growing latency suggests saturation, but further evidence is needed to identify the constrained resource.
20.3.5 TECHNICAL — the engineer’s version#
- Under suitable steady-state conditions and consistent boundaries,
Little’s law is
L = λW: average number in the system equals average effective arrival/completion rate times average time in the system. It is an accounting relation, not a universal capacity forecast for transient overload. - Distinguish offered rate, admitted rate, completed-attempt rate and successfully completed business-operation rate. Retries and rejected requests separate these quantities.
- Report saturation signals alongside user outcomes. Resource utilisation alone can miss queueing at a logical bottleneck, and high utilisation is not automatically harmful when latency and reliability requirements remain satisfied. [S103]
20.3.6 WORDS — remember these#
Throughput: completed work per unit time — a rate whose counted operation and success condition must be defined. Saturation: a limiting resource is at or near effective capacity — a condition in which additional demand tends to create waiting, rejection or deterioration rather than proportional useful work. Backpressure: make upstream work respect downstream limits — a mechanism that slows or bounds production when consumers cannot keep up.
20.4 Warm versus cold state#
20.4.1 PLAIN — in simple words#
- A repeated operation can benefit from work and data retained by earlier runs. This warm state may be realistic for frequent requests, but it must be named.
- Cold state is not one universal condition. A new connection can still use database, operating-system or device caches warmed by another process.
- Do not clear caches on a live machine merely to manufacture a benchmark. Use an isolated test environment and record what was actually reset.
20.4.2 PLAIN — a picture in your head#
- Dev’s first trip finds the right cupboard and brings folders to the desk. A second question about those folders avoids the trip.
- Starting a new notebook for timings does not move the folders back to a remote archive.
- Where the comparison breaks: there are several cache layers, each with its own lifetime and replacement policy. A single “desk empty” analogy cannot establish the state of all of them.
20.4.3 PLAIN — a worked example#
- In the educational microbenchmark, both variants receive a stated warm-up phase before measured repetitions. The comparison is labelled warm, in-process query execution on generated data.
- Reopening an in-memory SQLite database with newly inserted rows includes a different preparation history from reopening a file already read by the operating system. Neither should be called a device-cold test without further evidence.
- Alternate the order of variants or use a recorded random order to reduce systematic first-run advantage. Save all samples, including the first measured one.
- Index construction remains a separately reported preparation cost. Excluding it from steady-state query timing is acceptable only when that boundary is explicit.
20.4.4 PLAIN — what is really happening inside#
- Warm state can include pages, prepared statements, code paths, allocator state, connection establishment and storage caches. It can also include undesirable accumulated state such as a growing queue.
- Background work can occur between runs: checkpoints, cleanup or other tenants’ activity. A warm-up phase does not make the rest of the environment stationary.
- Compare the use case you need. Startup latency, infrequent cold retrieval and sustained warm service are different workloads, each potentially important.
20.4.5 TECHNICAL — the engineer’s version#
- Define reset scope precisely: process, connection, database buffer pool, filesystem cache, device state or dataset reconstruction. Do not claim a wider reset than the procedure demonstrates.
- In-memory SQLite timing omits network transport and does not exercise persistent-device durability. That can be appropriate for a teaching comparison but cannot establish server or storage capacity.
- Counterbalancing run order reduces one source of bias; it does not eliminate all thermal, scheduling or shared-host effects. Record those limitations rather than reporting false environmental control.
20.4.6 WORDS — remember these#
Warm state: earlier work remains useful — cached data or prepared resources available to later operations under a specified scope. Reset scope: what was actually returned to a starting condition — the exact layers or resources reinitialised before a measurement. Counterbalancing: vary which option runs first — an experimental ordering technique intended to reduce systematic sequence effects.
20.5 Fair comparisons#
20.5.1 PLAIN — in simple words#
- Compare equal work with equal correctness requirements. Disabling durability, dropping validation or returning fewer fields changes the service being measured.
- Use enough repetitions to expose variation, but do not confuse many nearly identical local samples with evidence about every production condition.
- Preserve unsuccessful runs and reasons for exclusion. A result selected after seeing which measurements look attractive is not a neutral comparison.
20.5.2 PLAIN — a picture in your head#
- Two delivery services cannot be compared fairly if one includes packaging and insurance while the other only moves an unwrapped box across the room.
- The cheaper service may still be useful, but the difference must be stated rather than hidden inside the word “faster.”
- Where the comparison breaks: software guarantees can be subtle. Two functions returning the same value once may differ under retries, concurrent access or failure, so equal happy-path output is necessary but not sufficient.
20.5.3 PLAIN — a worked example#
- Before timing a query variant, compare its ordered rows with an independently computed expected result. A variant that loses one valid duplicate fails correctness even if its duration is lower.
- Record the same generated dataset, parameter cases and selected fields for both variants. Keep schema and settings differences limited to the proposed intervention unless additional differences are explicitly part of the comparison.
- Use a fixed number of warm-ups and measured repetitions decided before inspecting the results. Summarise the samples with the same convention.
- A result such as “variant B was faster in these warm local samples” is narrower and more defensible than “B scales to ten thousand users.” The latter requires a different workload and capacity experiment.
20.5.4 PLAIN — what is really happening inside#
- A benchmark harness itself consumes resources. Data generation, logging, hashing results and printing can dominate a tiny operation if placed inside its timer boundary.
- Conversely, timing only a cursor creation can omit the work performed when rows are fetched. Ensure the measurement includes the actual completion condition.
- A trusted correctness check can be outside the repeated timing loop when the workload is deterministic, but it must still run for every variant and relevant parameter case. Record where checks occur.
20.5.5 TECHNICAL — the engineer’s version#
- Predefine the estimand: the quantity the comparison aims to estimate. Examples include median local fetch latency, successful throughput under a fixed arrival process or p99 end-to-end latency at a stated admitted load.
- Avoid a durability or isolation mismatch between variants. Different guarantees can be compared as different services, but not silently labelled equivalent implementations.
- Do not impose a universal speed assertion in a functional test. The educational suite checks result equality and valid measurement records; host scheduling can change timing without making the algorithm incorrect.
20.5.6 WORDS — remember these#
Estimand: the exact quantity a study seeks — the defined performance or outcome measure to which observations are intended to speak. Harness overhead: work caused by the measurement setup — time and resources consumed by generation, instrumentation and reporting rather than the target operation alone. Correctness oracle: a separate expected-answer check — a mechanism that rejects a fast but semantically wrong variant.
20.6 Reporting limitations#
20.6.1 PLAIN — in simple words#
- A useful performance report makes its boundaries easy to see. Readers should know what was measured, what was excluded and what cannot be concluded.
- Include actual versions, dataset construction, sample counts, result checks and failures. Label simulated or calculated figures as such.
- Do not turn a local educational run into a claim about production readiness, security, recovery or the number of schools a server can safely host.
20.6.2 PLAIN — a picture in your head#
- A laboratory label says which material was tested, under what load and for how long. It does not simply say “strong.”
- Performance evidence needs the same discipline: the label travels with the result.
- Where the comparison breaks: software and workloads change quickly. A benchmark’s input files and version identifiers are often necessary to understand a result later, even when the machine’s model name remains unchanged.
20.6.3 PLAIN — a worked example#
- A minimal record includes: Python and SQLite versions; generated row count and distribution; query text and parameter case; schema/index variant; warm-up and repetition counts; timer boundary; raw nanosecond samples; ordered result digest; and any errors.
- The result digest is a convenient reproducibility check, not an authentication or semantic proof by itself. The independently expected rows remain the correctness basis.
- Add an interpretation such as: “Warm single-process SQLite query comparison; no network, multi-user load, persistent-device failure or PostgreSQL server tested.”
- This explicit limitation does not make the result worthless. It tells the reader exactly which small question was answered and which larger experiments remain.
20.6.4 PLAIN — what is really happening inside#
- Reproducibility records separate stable inputs from variable observations. The same code and seed can reproduce the data while observed timings differ on another run.
- Keep raw measurements so summaries can be recalculated under another convention. Preserve original failures rather than overwriting them with a later passing run.
- A benchmark informs a decision together with correctness, maintainability, operational costs and service requirements. It is not the entire decision.
20.6.5 TECHNICAL — the engineer’s version#
- The companion benchmark emits a machine-readable record and makes no host-independent speed guarantee. Functional tests validate its schema, sample count, non-negative durations and result equivalence, not a predetermined winner.
- Report measurement uncertainty and environmental limits at the level the experiment supports. A narrow microbenchmark does not warrant an invented confidence interval for production throughput.
- Future capacity work should connect latency, successful throughput, errors and saturation under realistic demand, including failure and retry behaviour. [S103]
20.6.6 WORDS — remember these#
Raw sample: an individual retained observation — an unaggregated measurement from which summaries can be recomputed. Measurement boundary: what the timer includes — the explicitly chosen start and completion conditions of an observed operation. External validity: whether a result transfers beyond its test — the justified scope of applying observations to other workloads, environments or populations.
20.97 Practice and worked answers#
- Question: Nineteen requests take 10 ms and one takes 1,000 ms. Find the mean and nearest-rank p99. Answer: The mean is 59.5 ms; p99 selects the twentieth observation and is 1,000 ms.
- Question: Why can two libraries disagree on p95 for the same twenty samples? Answer: Quantile interpolation conventions differ. State the convention and retain raw samples.
- Question: At 100 completions per second and 0.2 seconds average time in a stable boundary, what average population does Little’s law imply? Answer: Twenty operations within that boundary, not necessarily twenty active CPU workers.
- Question: Can reopening a connection prove a cold-device benchmark? Answer: No. Other cache layers can remain warm.
- Question: A variant is faster because it stops fetching after the first row. Is it comparable to a full-result variant? Answer: Only if the required task is explicitly first-row retrieval. Otherwise the completion boundaries differ.
- Question: Should failed attempts disappear from a successful-latency chart? Answer: They may be excluded from that specifically labelled distribution, but their counts and outcomes must remain visible alongside it.
- Question: Why should a functional test avoid asserting that an index is always faster? Answer: Timing depends on host and workload conditions. The test can verify correct results and valid measurement records without inventing a universal performance bound.
- Question: What does the local in-memory benchmark not establish? Answer: Networked service capacity, multi-host behaviour, persistent-device durability, production isolation or a supported user count.
20.98 Common wrong ideas#
- Wrong: one average describes every user’s experience. Right: distributions and failures reveal important differences.
- Wrong: percentiles have one universal calculation convention. Right: label the method and sample population.
- Wrong: more clients always increase useful throughput. Right: saturation can increase waiting and errors instead.
- Wrong: a new connection means everything is cold. Right: reset scope must be demonstrated.
- Wrong: a faster result is valid even when guarantees changed. Right: compare equal services or disclose the difference.
- Wrong: the benchmark harness has no cost. Right: setup, fetching and instrumentation boundaries matter.
- Wrong: enough local repetitions prove production scalability. Right: the workload and environment determine transferability.
- Wrong: limitations weaken an otherwise strong result. Right: they prevent the result being used to support a claim it never tested.
20.99 Chapter summary in 20 lines#
- Define the required correct operation before timing it.
- Record data distribution, parameters, concurrency and operation mix.
- A timer boundary determines what the duration means.
- Closed-loop and open-loop arrivals expose different overload behaviour.
- Use a monotonic timer for elapsed intervals.
- Nanosecond units do not imply nanosecond accuracy.
- Latency is a distribution rather than one constant.
- Percentile convention and sample size affect interpretation.
- Keep timeouts and errors visible.
- Throughput needs a defined successful unit of work.
- Queueing can grow while completion rate stops increasing.
- Little’s law requires consistent boundaries and suitable steady-state assumptions.
- Warm state exists at several layers.
- Reset only what the procedure can legitimately establish.
- Compare identical result and guarantee contracts.
- Account for harness overhead and complete result retrieval.
- Retain raw samples and predefined exclusion rules.
- Do not build a universal speed assertion into a functional test.
- Report actual versions, inputs, observations and limitations together.
- A bounded experiment answers a bounded question, not every deployment decision.