Metrics, Bias and Honest Analysis
Introductions, exercises and summaries stay visible.
46.0 What this chapter gives you#
- A metric is a precisely defined question expressed as a number. The calculation can be flawless while the definition, population or interpretation is wrong.
- You will inspect denominators, missing observations, selection effects and causal claims. You will also calculate an uncertainty interval while keeping clear what the interval’s assumptions do and do not cover.
- All operational comparisons in this chapter are invented teaching examples. They are not observations of KedByte customers, published business performance or evidence that a particular real process causes an improvement.
46.1 A metric is a definition#
46.1.1 PLAIN — in simple words#
- “Sales,” “success” and “active user” are labels, not complete calculations. A useful metric specifies which events count, which period they belong to and what is excluded.
- A number should carry its unit and grain. Orders, order lines, items, money and customers are different quantities even when they come from the same database.
- Changing a definition can change the result without any change in the underlying business. Such a change should be recorded, not presented as an unexplained improvement or decline.
46.1.2 PLAIN — a picture in your head#
- Mira asks how busy the shop was. Dev can count customers, orders, items or hours of queueing. Each answer describes a different part of “busy.”
- Before comparing two weeks, they agree to count the same thing under the same rules.
- Where the comparison breaks: digital dashboards can hide definitions behind attractive labels and automatically refreshed charts. The visible number needs a traceable specification, not just a tooltip saying it is accurate.
46.1.3 PLAIN — a worked example#
- The canonical fixture contains two orders, four order lines, six items and 28,650 paise of agreed line amount. None of those counts can be substituted for another without changing the question.
- Average agreed amount per order is 28,650 / 2 = 14,325 paise. Per line it is 28,650 / 4 = 7,162.5 paise. Per item it is 28,650 / 6 = 4,775 paise.
- Calling 28,650 “cash received” would assert information the fixture does not contain. The database stores agreed order facts, not evidence of payment settlement.
- A metric card should therefore state: agreed line amount, INR paise, these four canonical lines, no taxes/discounts/refunds/shipping included, and the selected source and transformation versions.
46.1.4 PLAIN — what is really happening inside#
- A query operationalizes a definition using filters, joins and aggregates. Every step can change the population or multiplicity of contributing records.
- A metric registry connects a human definition to the query, owner, version and test cases. It helps prevent two teams from using one label for incompatible calculations.
- Reconciliation should link the displayed value back to its inputs. A dashboard total is more trustworthy when its contributing keys and exclusions can be inspected under appropriate permissions.
46.1.5 TECHNICAL — the engineer’s version#
- Specify numerator, denominator, eligibility, time basis, unit, grain, exclusions, missingness policy and aggregation level. Preserve the distinction between an observation and an inference.
- SQL aggregates operate on the rows produced by prior relational operations. A wrong join or filter can create a coherent but unintended statistic. [S58] [S80]
- Version metric definitions and record restatements. Comparing an old definition’s historical series with a new definition’s current value requires an explicit bridge or a recomputed comparable series.
46.1.6 WORDS — remember these#
Metric definition: the exact question behind a number — rules for population, units, timing, inclusion and calculation. Operationalization: turning an idea into a measurable procedure — implementing a definition through observable records and calculations. Restated series: historical results recalculated under a new rule — a versioned view distinct from what was previously published.
46.2 Denominators and populations#
46.2.1 PLAIN — in simple words#
- A percentage needs a denominator: successful out of which opportunities? Changing that population can change the rate substantially.
- Combining percentages requires attention to group sizes. The average of two group percentages is not generally the percentage across both groups.
- A comparison also needs comparable populations. If one process handles mostly simple cases and another mostly difficult cases, their overall rates can reflect the case mix rather than the process itself.
46.2.2 PLAIN — a picture in your head#
- Mira compares two counters. One handles ninety simple enquiries and ten difficult ones; the other handles twenty simple and eighty difficult ones.
- Looking only at each counter’s overall completion rate can hide the difference in work assigned to them.
- Where the comparison breaks: “simple” and “difficult” must be defined before the outcome and applied consistently. Inventing categories afterward can create another misleading comparison.
46.2.3 PLAIN — a worked example#
- In a synthetic two-method example, method A completes 81 of 90 easy cases and 1 of 10 hard cases. Its subgroup rates are 90% and 10%, and its overall rate is 82/100 = 82%.
- Method B completes 19 of 20 easy cases and 16 of 80 hard cases. Its subgroup rates are 95% and 20%, while its overall rate is 35/100 = 35%.
- B’s observed rate is higher within both defined subgroups, but A’s overall rate is higher because A handled a much larger share of easy cases. This is a constructed aggregation reversal, often called Simpson’s paradox.
- Under a hypothetical common mix of half easy and half hard, the standardized rates are 50% for A and 57.5% for B. These are reweighted descriptive calculations, not proof that changing methods would cause those outcomes.
46.2.4 PLAIN — what is really happening inside#
- Overall rates weight subgroup rates by each subgroup’s share of the denominator. When those shares differ, aggregate comparisons answer a different question from within-group comparisons.
- Standardization applies a chosen common population mix. It can clarify one source of difference, but the selected weights and subgroup definitions must be disclosed.
- Not every variable should automatically be adjusted for. A variable affected by the process or a selection rule can change the causal question. Descriptive reweighting is not a substitute for a justified causal design.
46.2.5 TECHNICAL — the engineer’s version#
- For subgroup counts, compute a pooled rate as
sum(successes) / sum(eligible_count), not an unweighted mean of rates unless equal group weighting is the actual question. - Preserve numerator and denominator in stored aggregates so later regrouping remains possible. An average without its count may be insufficient to compute a correct combined average.
- The Simpson example is an original arithmetic construction. It demonstrates sensitivity to population composition; it does not estimate a treatment effect or validate the easy/hard classifier for real operations.
46.2.6 WORDS — remember these#
Denominator: the population a rate is out of — the eligible count or exposure giving a numerator its meaning. Case mix: the composition of the observed population — subgroup proportions that can influence aggregate comparisons. Standardization: applying a common comparison population — reweighting subgroup results under explicitly chosen weights.
46.3 Missingness#
46.3.1 PLAIN — in simple words#
- Missing information is not automatically zero, failure or success. It means the required value is unavailable under the current observation process.
- Ignoring missing records changes the denominator. That may be reasonable for a clearly labelled observed-only statistic, but it does not automatically describe the entire eligible population.
- The reason information is missing matters. If difficult cases are less likely to be recorded, the observed subset can systematically look better than the full population.
46.3.2 PLAIN — a picture in your head#
- Mira receives delivery-status slips for only ninety of one hundred eligible deliveries. Nine of the returned slips say late and eighty-one say on time.
- The ten absent slips could change the full-population rate. They cannot be treated as on time merely because no complaint was filed.
- Where the comparison breaks: missingness can arise at several technical stages: capture, transmission, parsing, joins or permission filters. An apparently empty cell does not reveal which stage failed.
46.3.3 PLAIN — a worked example#
- There are 100 eligible deliveries, 81 known on time, 9 known late and 10 with unknown status. The observed-only rate is 81 / 90 = 90%.
- If all unknown statuses are late, the full-population on-time rate is 81%. If all are on time, it is 91%. The data alone therefore bounds the rate between 81% and 91% under the stated binary-status assumption.
- This range is not a statistical confidence interval. It is a deterministic sensitivity bound over the unresolved records.
- A dashboard can report “90% among 90 observed statuses; 10 of 100 statuses missing.” That statement preserves both the useful observed statistic and its limitation.
46.3.4 PLAIN — what is really happening inside#
- SQL aggregates often ignore NULL values in specific ways.
COUNT(*)counts rows, while counting a nullable expression counts its non-null values. The denominator must match the definition. - Joins can create or hide missingness. An inner join can silently remove records lacking a matching status row; an outer join retains them so a policy can be applied explicitly.
- Imputation fills missing values under a model or rule. It can be useful, but the filled values are inferred, not newly observed facts, and their uncertainty should not disappear from the analysis.
46.3.5 TECHNICAL — the engineer’s version#
- Separate unknown, not applicable, not yet observed, invalid and deliberately withheld states when the workflow needs those distinctions. A single NULL field may be insufficient without status metadata.
- Missingness assumptions affect inference. An observed-only rate represents the observed subset directly; extending it to unobserved cases requires justification beyond the division itself.
- The companion checks compare observed-only rates with explicit best/worst bounds and reject impossible counts. They do not implement a general missing-data model or assert that an imputation rule recovers the true values. [S80]
46.3.6 WORDS — remember these#
Missingness: required information is unavailable — a property of the observation process whose causes affect interpretation. Observed-only statistic: a result using records with available values — a measure whose population excludes unresolved observations. Sensitivity bound: a range under alternative assumptions — an explicit calculation showing how unknown information could change a conclusion.
46.4 Selection effects#
46.4.1 PLAIN — in simple words#
- The records available for analysis may not represent the population we care about. People, systems or filters can determine which cases become visible.
- Studying only successful requests cannot reveal the experience of requests that failed before logging. Studying only customers who stayed cannot fully explain those who left.
- More records do not automatically remove this problem. A huge systematically selected dataset can support a very precise answer to the wrong population question.
46.4.2 PLAIN — a picture in your head#
- Mira asks only customers who return next week whether they liked last week’s service. Customers who decided never to return are absent from the survey.
- The collected opinions are real, but they describe a selected group rather than all customers served.
- Where the comparison breaks: selection can be implemented invisibly by query filters, telemetry sampling and retention rules. The analyst may not notice it by reading the final table alone.
46.4.3 PLAIN — a worked example#
- Suppose 1,000 requests are attempted. The application logs only the 900 that reach its success handler. Of those, 891 complete within the chosen time bound.
- The logged-only rate is 891 / 900 = 99%. Treating all 1,000 attempts as the required denominator gives 891 / 1,000 = 89.1% known timely successes, with the other outcomes needing classification.
- A dashboard based only on the success table can therefore look excellent while hiding 100 attempts. The arithmetic is not the defect; the observation and eligibility boundaries differ.
- The remedy is to instrument the complete request lifecycle and reconcile attempts, successes, failures and unknown outcomes under a consistent identity scheme, not simply to rename 99% as reliability.
46.4.4 PLAIN — what is really happening inside#
- Data reaches an analytical table through capture, sampling, filtering and retention. Each stage can select cases based on properties related to the outcome.
- Duplicate logs can bias the other direction by giving some events extra weight. The intended statistical unit might be a request, user or day rather than one stored log row.
- Evaluation data can also be contaminated by information unavailable at the intended decision time. A predictor using a field written after completion may appear accurate while being unusable when the real decision must be made.
46.4.5 TECHNICAL — the engineer’s version#
- Document the sampling frame and all inclusion/exclusion stages. Trace a record’s route into the analysis and inspect cases lost before the final query.
- For temporal evaluation, preserve the feature-availability boundary: the model or rule should use only information available at the relevant prediction time. Keep development and evaluation decisions separate to reduce overfitting.
- Repeated observations from one subject or time series can be dependent. NIST’s autocorrelation discussion illustrates why order and dependence matter; treating every row as an independent draw can overstate information. [S179]
46.4.6 WORDS — remember these#
Selection bias: the observed group systematically differs from the target — distortion caused by how cases enter or remain in the dataset. Sampling frame: the cases available for selection — the actual population from which observations can be drawn. Information leakage: using knowledge unavailable at the intended decision point — contamination that can make evaluation results unrealistically strong.
46.5 Association versus cause#
46.5.1 PLAIN — in simple words#
- Two quantities changing together establishes an observed association, not automatically that one caused the other.
- Another factor can influence both, or the direction can run the other way. The data-collection process can also create an apparent relationship.
- A causal question asks what would change under an intervention. Answering it requires a justified design and assumptions, not just a chart with a sloping line.
46.5.2 PLAIN — a picture in your head#
- Mira notices that busy days have longer queues. She should not conclude that deliberately making people queue will create more demand.
- Demand may contribute to queue length, while staffing, promotions or holidays influence both. Several stories can fit the same observed pattern.
- Where the comparison breaks: causal reasoning can be formalized with experiments and explicit models. The point is not that causes are unknowable, but that the evidence must distinguish plausible explanations.
46.5.3 PLAIN — a worked example#
- Consider a proposed new checkout workflow. Comparing this month’s rate with last month’s rate also changes season, staffing, customer mix and perhaps measurement rules.
- Randomly assigning eligible cases to old and new workflows can help separate assignment from those pre-existing differences, provided the experiment is properly designed and implemented.
- Randomization does not guarantee identical finite groups, and shared queues can create interference between them. Missing outcomes, noncompliance and changed logging can still damage the interpretation.
- The synthetic A/B case-mix table in section 46.2 was not randomized. Its subgroup and standardized percentages remain descriptive calculations, not experimentally established causal effects.
46.5.4 PLAIN — what is really happening inside#
- Observational data records what happened under existing selection and assignment processes. Counterfactual outcomes under another action are not directly present in the same record.
- Experiments define interventions, assignment, outcomes and analysis rules. Observational causal studies require additional assumptions about confounding, timing and selection.
- Good reporting separates the observed difference from the causal interpretation and states the design limitations. It does not dismiss useful associations, but it does not promote them into stronger claims without evidence.
46.5.5 TECHNICAL — the engineer’s version#
- A scatter plot can reveal patterns and possible relationships but does not establish cause and effect. NIST explicitly distinguishes this diagnostic use from causal proof. [S180]
- Define the estimand: the population, intervention, outcome and contrast being sought. Choosing covariates or subgroups should follow that causal question rather than automatically adjusting for every available field.
- The book does not estimate a real treatment effect. Its examples teach the separation between arithmetic, descriptive association and causal inference, leaving a real study’s assumptions and review to its actual context.
46.5.6 WORDS — remember these#
Association: quantities vary together in observations — a statistical relationship that does not by itself establish intervention effects. Confounder: a factor affecting the causal comparison — a variable related to both assignment/exposure and outcome in a relevant causal structure. Estimand: the precise quantity a study seeks — a defined population-level contrast under specified interventions or conditions.
46.6 Communicating uncertainty#
46.6.1 PLAIN — in simple words#
- A reported number can be uncertain because only a sample was observed, measurements are imperfect or records are missing. These are different sources of uncertainty.
- An interval calculated from a sampling model covers only the uncertainty represented by that model. It does not repair biased selection, incorrect labels or a changing process.
- Explain the result with its denominator, assumptions and practical limits. Extra decimal places should not create the appearance of knowledge the evidence does not contain.
46.6.2 PLAIN — a picture in your head#
- Mira checks a sample of packages rather than every package. She can calculate how much the observed proportion might vary under repeated comparable sampling.
- If she checked only the easiest packages to reach, a narrow interval around that selected sample does not solve the selection problem.
- Where the comparison breaks: statistical coverage is a property of a procedure under assumptions. It is not a promise that every particular interval contains the fixed unknown quantity.
46.6.3 PLAIN — a worked example#
- In a separate idealized sample, 90 of 100 independent, comparably sampled cases satisfy a binary condition. The observed proportion is 0.9.
- Using a Wilson interval with z = 1.96 gives an approximate two-sided 95% interval from 0.826 to 0.945. This calculation assumes a suitable binomial sampling model and is not the same as the missing-status bound from section 46.3.
- The Wilson centre is
(p + z²/(2n)) / (1 + z²/n). Its half-width isz × sqrt(p(1-p)/n + z²/(4n²)) / (1 + z²/n). Substituting p = 0.9 and n = 100 produces the stated rounded limits. - The companion code checks the arithmetic, edge cases and valid range. It does not establish that real operational observations satisfy independence, stable probability or representative sampling.
46.6.4 PLAIN — what is really happening inside#
- Sampling intervals describe repeated-sampling performance under a model. Measurement uncertainty instead concerns how values are attributed from instruments and observations; missingness bounds explore unresolved cases.
- Dependence reduces the information obtained from repeated observations compared with an ideal independent sample of the same row count. A thousand correlated readings are not automatically a thousand independent pieces of evidence.
- Report uncertainty together with decision relevance. A small numerical difference may be operationally unimportant, while a modest uncertainty range near a safety or capacity limit may matter greatly.
46.6.5 TECHNICAL — the engineer’s version#
- NIST presents the Wilson interval and discusses alternatives for small samples or few failures. Avoid the naive interpretation that a confidence interval assigns a 95% posterior probability to a fixed parameter without a Bayesian model. [S183]
- A confidence procedure’s nominal coverage is conditional on its statistical assumptions and can be approximate. Selection, dependence, repeated unplanned testing and definition changes can invalidate a simplistic interpretation. [S179]
- Preserve raw counts, calculation method, source version and rounding policy. Label unresolved measurement and data-quality limitations separately instead of folding them into an unjustified single “confidence” percentage.
46.6.6 WORDS — remember these#
Confidence interval: a range from a sampling procedure — an interval method with stated repeated-sampling coverage under assumptions. Wilson interval: a binomial-proportion interval — a score-based method whose limits remain within the valid probability range. Statistical unit: the entity providing an observation — the request, person, item or time unit whose independence and weighting need definition.
46.97 Practice and worked answers#
- Question: What does 28,650 paise mean in the canonical fixture? Answer: The agreed amount of the four included order lines, not verified cash collection.
- Question: Why does averaging subgroup percentages often fail? Answer: It ignores subgroup denominator sizes unless equal group weighting is the intended question.
- Question: What explains the A/B aggregate reversal? Answer: Different easy/hard case mixtures. The constructed arithmetic does not establish a causal effect.
- Question: With 81 on time, 9 late and 10 unknown out of 100, what is observed-only performance? Answer: 90% among the 90 observed statuses; the full-population binary bound is 81% to 91%.
- Question: Is that 81%–91% range a confidence interval? Answer: No. It is a sensitivity bound over missing outcomes.
- Question: Why can success-only logs mislead? Answer: They omit attempts that never reached the success handler, changing the denominator and population.
- Question: Does a sloping scatter plot establish cause? Answer: No. It shows association; causal interpretation needs additional design and assumptions.
- Question: What does the Wilson calculation fail to verify? Answer: Whether the real observations satisfy the sampling, independence, measurement and stability assumptions required for its interpretation.
46.98 Common wrong ideas#
- Wrong: A familiar metric name is a complete definition. Right: Population, timing, units and exclusions must be explicit.
- Wrong: A correct formula guarantees a correct conclusion. Right: The formula may answer the wrong question.
- Wrong: Group percentages can always be averaged directly. Right: Denominator weights matter.
- Wrong: Missing means failure or zero. Right: Missingness needs a declared interpretation and may require separate status.
- Wrong: More rows automatically remove bias. Right: Systematic selection can persist at any scale.
- Wrong: Association is evidence enough for intervention effects. Right: Causal conclusions require stronger assumptions or design.
- Wrong: Every repeated measurement is independent. Right: Dependence can reduce effective information.
- Wrong: A confidence interval covers every source of uncertainty. Right: It covers only what the selected statistical model represents.
46.99 Chapter summary in 20 lines#
- A metric is a defined question, not only a label.
- Preserve the distinction between orders, lines, items and money.
- State the numerator, denominator and eligible population.
- Keep units and time basis explicit.
- Version definition changes and historical restatements.
- Pool counts before calculating a pooled rate.
- Case mix can reverse aggregate comparisons.
- Standardization is a declared reweighting, not automatic causal proof.
- Missing observations do not become known outcomes by convenience.
- Report observed-only statistics with their coverage.
- Sensitivity bounds and confidence intervals answer different questions.
- Selection can occur before data reaches the final table.
- Success-only logs can hide unsuccessful attempts.
- Statistical units are not always individual stored rows.
- Association does not by itself establish cause.
- Define the intervention and estimand for causal questions.
- Sampling intervals rely on explicit assumptions.
- Dependence and selection are not cured by extra decimal places.
- Preserve counts, methods and unresolved limitations.
- Communicate the strongest conclusion the evidence actually supports.