Stats 2 · Inference & Experiments
Formulas are drawn by KaTeX, which loads from the internet. You seem to be offline, so formulas show as plain text for now.
Statistics for ML · Guide 2 of 4

Estimation, Inference & Experiments

How to learn about the world from a sample, and how sure you can be. Estimators and their errors, maximum likelihood, regularization as a prior, standard errors, hypothesis tests, confidence intervals, A/B testing and its traps, causal thinking, regression and GLMs, and the tools for looking at many variables at once (PCA, t-SNE). Every idea is explained in plain English first, then with numbers, then with plots you can play with.

What is this guide about, in one sentence? Going backwards: from the data you saw to the process that made it, with an honest statement of how uncertain that conclusion is.

Why care? Your A/B framework is Bayesian, but every reviewer, product manager and interviewer will compare it with the classical approach: p-values, confidence intervals, power and sample size. And your forecasting model is, at its heart, a regression with a likelihood: MLE, MAP, regularization and GLMs are exactly the ideas it is built from.

Three ways to say it:

  • Picture: you only see one handful of marbles; inference is guessing what is in the whole bag, and how wrong your guess might be.
  • Numbers: 60 of 500 converted with the new checkout vs 50 of 500 with the old one. Is the 2-point gap real, or luck?
  • Slogan: an estimate without its uncertainty is only half an answer.

Before you start. This guide uses Guide 1 (Probability & Data) all the time: random variables, expectation and variance, the Normal, Binomial and Student-t distributions, the CLT. Whenever one appears, you get a one-line reminder and a link. Linear regression and PCA also lean on the Linear Algebra guide (projections, eigenvectors, SVD), and MLE and regularization lean on the Optimization guide.

The four guides

1 · Probability & Data

Probability, random variables, distributions, LLN and CLT, descriptive statistics, correlation, KDE, Q-Q plots, transformations. Syllabus Parts 0–4.

2 · Estimation, Inference & Experiments (you are here)

Estimators, MLE and MAP, regularization, sampling, standard errors, tests, confidence intervals, A/B testing, causal thinking, regression, GLMs, PCA, t-SNE. Parts 5–9, 26, 27.

3 · Bayesian Modeling & Computation

Priors and posteriors, conjugacy, credible intervals, hierarchical models, MCMC, NUTS, variational inference, SVI, JAX. Parts 10, 11, 19–22, 24, 25.

4 · Time Series & Bayesian Forecasting

Autocorrelation, stationarity, classical forecasting, the Prophet-style model, forecast evaluation, diagnostics, production, capstone. Parts 12–18, 23, 28–30.

Built around your two projects

The Bayesian A/B testing framework

The frequentist toolkit you will be compared against (5.6–5.9), how real experiments are designed and how they go wrong (5.10–5.11), why randomization lets you claim causation and how CUPED cuts variance (5.12), and MAP vs the full posterior (5.2–5.3).

The Prophet-style forecasting model

The model is a regression with structured columns (5.13), a likelihood family with a link (5.14), and a Laplace prior that acts like an L1 penalty only at the MAP (5.3). Correlated regressors and the covariance ideas behind full-rank guides live in 5.13 and 5.15.

How every topic is taught

Same steps as in every guide: Intuition (said three ways) → Example (every step) → Definition → Why · Where · How → picture and play → Careful (✗ wrong idea / ✓ right idea) → In your projects and Say it right → Write this in your notebook → quick check. Each chapter ends with a recap, Python code, quizzes and practice problems.

Try your first interactive

Here is the whole idea of hypothesis testing in one widget, using the checkout example from your hypothesis-testing notes.

The old checkout converted 50 of 500 users (10%); the new one converted 60 of 500 (12%). If the new checkout did nothing, the "converted" labels would be shuffled at random between the two groups. Press Shuffle once to deal the 110 conversions randomly into two groups of 500 and see the gap you get by pure luck. Press Shuffle 1000 times: the grey histogram shows the gaps luck alone produces. The fraction of shuffles with a gap at least as big as the real one (red tails) is the p-value. Here it comes out near 0.36 (the formula version below gives 0.31; small differences between methods are normal), so luck explains this gap easily.

The roadmap

What matters most (your P0 list). Estimators, bias, variance and MSE (5.1), sampling distributions and standard errors (5.5), hypothesis testing, p-values and confidence intervals (5.6–5.8), Type I/II errors and power (5.7), A/B testing (5.10–5.11), and MAP vs full posterior (5.2–5.3). PCA (5.16) is P0 too.

How to read the symbols

SymbolSay it asMeaning
$\theta$, $\hat\theta$"theta", "theta hat"The unknown parameter, and an estimate of it computed from data.
$E[\hat\theta]-\theta$"bias"How far the estimator is off on average.
$SE(\hat\theta)$"standard error"How much the estimate would wobble across repeated samples.
$H_0$, $H_1$"H-nought", "H-one"The null hypothesis ("no effect") and the alternative.
$\alpha$, $\beta$, $1-\beta$"alpha", "beta", "power"False-positive rate, false-negative rate, and the chance of detecting a real effect.
$p$-value"p-value"If $H_0$ were true, how often you would see data at least this extreme.
$L(\theta)$, $\ell(\theta)$"likelihood", "log-likelihood"How well each parameter value explains the observed data.
$\arg\max_\theta$"the arg max over theta"The value of $\theta$ that makes the expression largest.
$y = X\beta + \epsilon$"y equals X beta plus epsilon"The linear regression model in matrix form.
$g(E[Y]) = X\beta$"g of the mean equals X beta"A generalized linear model with link function $g$.
$\Sigma$"capital sigma"A covariance matrix (careful: not the summation sign $\sum$).

If a section feels too hard, a word from earlier is usually fuzzy: "statistic", "estimator", "standard error" and "likelihood" are the usual suspects. Find it in the sidebar or the glossary and reread it. Statistics takes two passes for everyone.

Chapter 5.1 · Syllabus Module 10

Estimators: bias, variance, MSE, consistency, efficiency

Every number you report from data (a conversion rate, an average order size, a noise level, a slope) is a guess about something you cannot see directly. This chapter teaches you how to judge a guessing rule: is it right on average, how much does it wobble, does it get better with more data, and does it waste information? These five questions decide which recipe you should trust.

  • Use the four words parameter, statistic, estimator, estimate correctly, and explain why an estimator is a random variable
  • Compute the bias and the variance of an estimator, and see both as features of its sampling distribution
  • Derive MSE = Bias² + Variance and use it to compare recipes
  • See the bias–variance trade-off in action: a deliberately biased, shrunk estimate can beat the honest average
  • Tell unbiased apart from consistent, and know an example of each without the other
  • Explain efficiency: why the mean beats the median for Normal data (relative efficiency ≈ 2/π ≈ 0.64) and loses for heavy-tailed data

What we need from earlier chapters: expected value and variance, and the rules $E[aX+b] = aE[X]+b$ and $Var(aX) = a^2Var(X)$ (Chapter 4.5); the Law of Large Numbers (Chapter 4.13); Chebyshev's and Markov's inequalities (Chapter 4.12); mean, median and trimmed mean (Chapter 4.14). Notation for this guide: a parameter is a Greek letter such as $\theta$ ("theta"); a hat means "our guess of it": $\hat\theta$ ("theta hat"). Data are $D = \{x_1, \dots, x_n\}$, the sample mean is $\bar x$, the sample standard deviation is $s$; $\mu$ and $\sigma$ are the population mean and standard deviation. We write $N(\mu, \sigma^2)$ with the variance inside; NumPy, SciPy and NumPyro take the standard deviation $\sigma$ instead.

Four words: parameter, statistic, estimator, estimate core

A cook wants to know how salty a big pot of soup is. She cannot drink the whole pot, so she takes a spoonful and tastes it.

  • The true saltiness of the whole pot is the parameter. It is one fixed number, but nobody can see it directly.
  • The spoonful is the sample. Any number she works out from the spoonful alone is a statistic.
  • Her tasting method ("take three sips from different places and average them") is the estimator. It is a recipe, decided before she tastes anything.
  • The verdict she reaches today ("7 out of 10 salty") is the estimate: the number the recipe produced on this particular spoonful.

Tomorrow, with a new spoonful, the same method gives a slightly different verdict. The pot did not change. Only the spoonful did.

Three ways to say it:

  • Picture: the estimator is the recipe card; the estimate is the dish you cooked from it today.
  • Numbers: the rule "conversions ÷ users" is the estimator; $53/500 = 0.106$ is the estimate.
  • Slogan: the parameter stays still; the estimate moves with every new sample.

A checkout test. 500 users saw a new checkout page and 53 of them bought something.

  1. Parameter: $\theta$ = the true purchase rate of this checkout for all users who could ever see it. Unknown, fixed.
  2. Statistics: anything computed from the 500 rows alone. For example the count $53$, the proportion $53/500$, or "the number of purchases in the first hour". A statistic may not use $\theta$, because we do not know $\theta$.
  3. Estimator 1: the recipe $\hat\theta_1 = \dfrac{K}{n}$, where $K$ is the number of buyers and $n$ the number of users.
  4. Estimate 1: apply it to today's data: $\hat\theta_1 = 53/500 = 0.106$.
  5. Estimator 2 (a different recipe for the same parameter): "add one pretend buyer and one pretend non-buyer", $\hat\theta_2 = \dfrac{K+1}{n+2}$.
  6. Estimate 2: $\hat\theta_2 = 54/502 \approx 0.1076$.

Two recipes, two estimates, one parameter. Which recipe is better is not a question about today's numbers. It is a question about how each recipe behaves over all the samples we could have drawn. That is what the rest of this chapter measures.

Suppose the data $X_1, \dots, X_n$ come from a process described by an unknown number $\theta$.

  • A parameter $\theta$ is a fixed number that describes the population or the data-generating process (a rate, a mean, a standard deviation, a slope).
  • A statistic is any function of the data alone: $T = g(X_1, \dots, X_n)$. It must not contain unknown parameters. (Here "statistic" means "a number computed from the sample", not the school subject.)
  • An estimator $\hat\theta = g(X_1, \dots, X_n)$ is a statistic that we use to guess $\theta$. Because the $X_i$ are random, $\hat\theta$ is a random variable.
  • An estimate is the value the estimator takes on the data we actually observed: $\hat\theta = g(x_1, \dots, x_n)$, one fixed number.

Capital letters ($X_i$) are the random quantities before we look; small letters ($x_i$) are the values we saw (Chapter 4.4). People usually write $\hat\theta$ for both the estimator and the estimate; the sentence tells you which one is meant.

Why do we need it?

Without these words, "the conversion rate" can mean the unknown truth, the rule you used, or today's number, and arguments go in circles. Separating them lets us ask the right question: not "is 0.106 correct?" (we can never know) but "is the rule that produced it a good rule?".

Where is it used?

Every report of a metric: conversion rates and lifts in A/B tests, average order value, the slope of a regression, the noise level $\sigma$ in a forecasting model, the weights of a trained neural network (they are estimates of the best weights).

How is it used?

Name the parameter first ("the true purchase rate of variant B"). Then write the recipe you will apply (for example k / n, a posterior mean, or an MLE from Chapter 5.2). Then judge the recipe with bias, variance and MSE, before trusting the number it gives.

Population / process parameter θ fixed, never seen draw Sample D x₁, x₂, …, xₙ 500 users, 53 bought recipe g Estimate θ̂ = 53/500 = 0.106 one number, today Estimator = the recipe g(·) "buyers ÷ users", chosen in advance A new sample gives a new estimate; θ and the recipe stay the same.
The parameter lives in the process and never changes. The sample is random. The estimator is the fixed recipe that turns any sample into a number; the estimate is what it produced on the sample you got.

Seven days of orders (blue dots). Each coloured line is a different estimator of "the typical number of orders per day", applied to the same data. Drag one day far to the right (or press One wild day): the mean and the midrange jump, the median does not move at all, the trimmed mean barely moves. Press Lopsided week: now every recipe gives a different answer. Same data, different rules, different estimates.

"An estimator is a number."

An estimator is a rule (a function of the data). The number it gives on your data is the estimate. The rule can be judged; a single number cannot.

"A statistic is something from statistics class."

A statistic is any number computed from the sample alone: the mean, the maximum, the count of zeros, a test statistic. It may not contain unknown parameters.

"When I collect more data, the parameter changes."

The parameter is a property of the process and stays fixed. Your estimate changes, because you now apply the recipe to a different sample.

In an A/B framework like yours, $\theta_A$ (the true conversion rate of variant A) is the parameter. The observed rate $k_A/n_A$ is one estimator of it (it is the maximum likelihood estimator, Chapter 5.2). The posterior mean of a Beta-Binomial model, $(\alpha + k_A)/(\alpha + \beta + n_A)$, is another estimator of the same parameter: a different recipe applied to the same data. In your forecasting model the trend slope, each changepoint adjustment $\delta_j$ and the noise scale $\sigma$ are parameters, and the values your SVI fit reports for them are estimates.

"My estimator is 0.106."

"My estimate is 0.106. My estimator is the observed proportion $K/n$, which is unbiased with standard error about $\sqrt{\theta(1-\theta)/n}$."

Model answer: "An estimator is a function of the random sample, so it has a distribution: it has a bias, a variance and an MSE. An estimate is the value that function took on my data. I judge the estimator, and I report the estimate with its uncertainty."

Parameter $\theta$: fixed, unknown. Statistic: any number computed from the sample. Estimator $\hat\theta = g(X_1..X_n)$: a statistic used to guess $\theta$, a random variable. Estimate: its value on my data.

Different recipes give different estimates of the same parameter.

Trap: you judge the recipe, never a single number.

Quick check: is "$\bar x - \mu$" a statistic?

No. It contains $\mu$, the unknown population mean, so you cannot compute it from the sample alone. $\bar x$ is a statistic; $\bar x - 50$ (with the number 50 fixed in advance) is also a statistic.

An estimator is a random variable: its sampling distribution core

Before you run an experiment, you do not know which users will arrive. So you do not know what the estimate will be. If you could rerun the same experiment many times (same page, same kind of users, same sample size), you would get a different estimate each time.

Collect all those imaginary estimates and draw their histogram. That histogram is the sampling distribution of the estimator. Every question in this chapter (bias, variance, consistency, efficiency) is a question about the shape and position of this histogram.

Three ways to say it:

  • Picture: one sample gives one brick; many repeated samples build a pile of bricks, and the pile is the sampling distribution.
  • Numbers: with a true rate of 10% and 500 users, about 95% of the estimates from repeated experiments land between 0.074 and 0.126.
  • Slogan: your estimate is one draw from the estimator's distribution.

A tiny world we can list completely. A shop sells 2, 4 or 6 boxes a day, each equally likely (so the true mean is $\mu = 4$). We take a sample of $n = 2$ days and use the recipe "sample mean" $\bar X = (X_1 + X_2)/2$.

  1. There are $3 \times 3 = 9$ equally likely samples: (2,2), (2,4), (2,6), (4,2), (4,4), (4,6), (6,2), (6,4), (6,6).
  2. Their means: 2, 3, 4, 3, 4, 5, 4, 5, 6.
  3. So $\bar X$ takes the value 2 with probability $1/9$, 3 with $2/9$, 4 with $3/9$, 5 with $2/9$, 6 with $1/9$. This list is the sampling distribution of $\bar X$.
  4. Its centre: $(2\cdot1 + 3\cdot2 + 4\cdot3 + 5\cdot2 + 6\cdot1)/9 = 36/9 = 4 = \mu$.
  5. Its spread: $Var(\bar X) = \big(1\cdot(2-4)^2 + 2\cdot(3-4)^2 + 3\cdot 0 + 2\cdot(5-4)^2 + 1\cdot(6-4)^2\big)/9 = 12/9 = 4/3$.
  6. Compare one single day $X_1$: values 2, 4, 6 with probability $1/3$ each, variance $(4 + 0 + 4)/3 = 8/3$. The average of two days has half the variance: $\tfrac{4/3}{8/3} = \tfrac12 = \tfrac1n$.

The sampling distribution of an estimator $\hat\theta$ is the probability distribution of $\hat\theta = g(X_1, \dots, X_n)$ over repeated samples drawn from the same process with the same sample size $n$.

  • Its mean $E[\hat\theta]$ tells us where the estimator is centred (this gives the bias, next section).
  • Its standard deviation is called the standard error $SE(\hat\theta)$: the typical size of the estimator's wobble. For the sample mean of iid data, $SE(\bar X) = \sigma/\sqrt n$. Standard errors get their own chapter (Chapter 5.5).
  • It depends on three things: the data-generating process, the sample size $n$, and the recipe $g$.
Why do we need it?

A single estimate cannot tell you how much to trust it. The sampling distribution can: if repeated experiments would scatter widely, today's number might be far from the truth. It is the bridge from "a number" to "a number with uncertainty".

Where is it used?

Standard errors, confidence intervals and p-values (Chapters 5.5–5.8) are all read off a sampling distribution. Power calculations for A/B tests, the bootstrap, and simulation studies that compare estimators all work with it directly.

How is it used?

Either work it out with theory (for the mean: centre $\mu$, standard error $\sigma/\sqrt n$, roughly Normal by the CLT) or simulate it: draw many fake samples from a believable process, apply the recipe to each, and look at the histogram of the results.

Daily orders come from a process with true mean $\mu = 50$ (green line) and standard deviation 10. Press Draw one sample a few times: the top strip shows the $n$ values of the new sample, and its estimate drops into the histogram as one brick. Then press Draw 100 samples. The pile is the sampling distribution. Switch the recipe to first value: the pile is much wider. Raise $n$: the pile of means gets narrower, roughly like $10/\sqrt n$.

"The sampling distribution is the distribution of my data."

The data's distribution describes single observations (one day's orders). The sampling distribution describes the estimate across repeated samples. For the mean, it is about $\sqrt n$ times narrower than the data's distribution.

"I only have one sample, so my estimator has no distribution."

The distribution is about the samples you could have drawn. You can still learn its shape: from theory (like $\sigma/\sqrt n$) or by resampling your one sample with the bootstrap (Chapter 5.5).

Estimator = random variable. Its distribution over repeated samples (same process, same $n$) = sampling distribution.

Centre → bias. Width (its sd) → standard error; for the mean $SE = \sigma/\sqrt n$.

Trap: do not confuse the spread of the data ($\sigma$) with the spread of the estimate ($\sigma/\sqrt n$).

Quick check: in the 2-4-6 shop, what is the sampling distribution of the recipe "first day only" for $n = 2$?

The first day is 2, 4 or 6 with probability $1/3$ each, whatever the second day is. So the recipe's sampling distribution is the population itself: centre 4 (unbiased), variance $8/3$, twice the variance of $\bar X$. It ignores half the data.

Bias: is the recipe right on average? core

A bathroom scale that always reads 1 kg too heavy is biased. Weigh yourself a hundred times and average the readings: the average is still 1 kg too heavy. Random noise cancels out when you average; a bias does not.

An estimator can have the same problem. Imagine running your experiment again and again and averaging all the estimates. If that long-run average sits exactly on the truth, the recipe is unbiased. If it sits off to one side, the gap is the recipe's bias.

Three ways to say it:

  • Picture: the centre of the pile of estimates is shifted away from the green "truth" line.
  • Numbers: dividing by $n$ instead of $n-1$ makes the sample variance too small on average by exactly $\sigma^2/n$.
  • Slogan: bias is the error that averaging can never remove.

A world small enough to list every sample. A value $X$ is 0 or 2, each with probability $\tfrac12$. So $\mu = 1$ and $\sigma^2 = \tfrac12(0-1)^2 + \tfrac12(2-1)^2 = 1$. Take samples of size $n = 2$.

  1. The four equally likely samples are (0,0), (0,2), (2,0), (2,2), each with probability $\tfrac14$.
  2. Their means are 0, 1, 1, 2. The squared distances to the sample mean add up to $0$, $1 + 1 = 2$, $2$, $0$.
  3. Recipe A, divide by $n = 2$: the four estimates are $0, 1, 1, 0$. Their average is $(0+1+1+0)/4 = 0.5$.
  4. So $E[\hat\sigma^2_A] = 0.5$ while the truth is $\sigma^2 = 1$. Bias $= 0.5 - 1 = -0.5$. This equals $-\sigma^2/n = -1/2$.
  5. Recipe B, divide by $n - 1 = 1$: the estimates are $0, 2, 2, 0$, with average $1$. Bias $= 1 - 1 = 0$: unbiased.

Recipe A measures distances to the sample mean, which always sits in the middle of the sample, so the distances come out too small. Dividing by $n-1$ exactly makes up for that (the full story is in Chapter 4.5).

The bias of an estimator $\hat\theta$ of a parameter $\theta$ is

$$Bias(\hat\theta) = E[\hat\theta] - \theta,$$

where the expectation is over the sampling distribution (repeated samples from the same process, same $n$).

  • $\hat\theta$ is unbiased if $E[\hat\theta] = \theta$ for every possible value of $\theta$.
  • It is asymptotically unbiased if the bias goes to $0$ as $n \to \infty$.

Standard examples (iid data):

Estimatorof$E[\hat\theta]$Bias
sample mean $\bar X$$\mu$$\mu$$0$
first value $X_1$$\mu$$\mu$$0$ (but very noisy)
$\hat\sigma^2_n = \frac1n\sum(X_i-\bar X)^2$$\sigma^2$$\frac{n-1}{n}\sigma^2$$-\sigma^2/n$
$s^2 = \frac1{n-1}\sum(X_i-\bar X)^2$$\sigma^2$$\sigma^2$$0$
$s = \sqrt{s^2}$ (Normal data)$\sigma$$c_4\,\sigma$, with $c_4 \lt 1$ (0.94 at $n=5$)negative
largest value, data Uniform$(0,\theta)$$\theta$$\frac{n}{n+1}\theta$$-\frac{\theta}{n+1}$
$K/n$ (conversion rate)$p$$p$$0$
$(K+1)/(n+2)$$p$$\frac{np+1}{n+2}$$\frac{1-2p}{n+2}$ (pulls toward $\tfrac12$)
Why do we need it?

A systematic error does not shrink when you collect more of the same kind of data, and it does not show up in a standard error. If a recipe leans one way, every decision built on it leans the same way. Bias is how we detect that lean.

Where is it used?

The $n-1$ in np.var(x, ddof=1) and in pandas' .var(); bias corrections in survey weighting; the bias of the MLE of a Normal variance (Chapter 5.2); and the deliberate, useful bias of ridge regression, shrinkage and partial pooling (later in this chapter).

How is it used?

Work out $E[\hat\theta]$ with the rules of expectation, or simulate: generate many samples from a process where you know $\theta$, apply the recipe, and compare the average estimate with $\theta$. Then either correct the recipe, or accept the bias if it buys a large drop in variance.

The widget draws many samples from a process where the truth is known (green line), applies the chosen recipe to each, and draws the pile of estimates. The purple line is their average; the red arrow is the bias. Start with variance, divide by n at $n = 2$: the average sits at about half the truth. Slide $n$ up and watch the bias shrink like $\sigma^2/n$. Switch to divide by n − 1: the arrow disappears. Then try standard deviation s: unbiased $s^2$ does not give unbiased $s$.

"Unbiased means my estimate is right."

Unbiased means right on average over repeated samples. A single estimate from an unbiased recipe can still be far from the truth; the first-value recipe is unbiased and terrible.

"$s^2$ is unbiased for $\sigma^2$, so $s$ is unbiased for $\sigma$."

The square root is a curved function, and $E[\sqrt{V}] \lt \sqrt{E[V]}$ (Jensen's inequality, Chapter 4.5). So $s$ underestimates $\sigma$ on average. Bias does not survive non-linear transformations.

"A biased estimator is a bad estimator."

Bias is only one part of the error. A small, deliberate bias can buy a large cut in variance and give a smaller total error. That is exactly what shrinkage, ridge regression and partial pooling do (see the trade-off section below).

$Bias(\hat\theta) = E[\hat\theta] - \theta$. Unbiased: $E[\hat\theta] = \theta$ for every $\theta$.

$\frac1n\sum(x_i-\bar x)^2$ has bias $-\sigma^2/n$; dividing by $n-1$ fixes it. $s$ is still biased for $\sigma$.

Trap: unbiased ≠ accurate, and biased ≠ bad.

Quick check: data are Uniform(0, θ) and $n = 1$. What is the bias of "the largest value" as an estimator of θ?

With one value, the largest value is that value, whose mean is $\theta/2$. Bias $= \theta/2 - \theta = -\theta/2$. The formula agrees: $-\theta/(n+1) = -\theta/2$. Doubling it, $2X_1$, is unbiased.

Variance: how much does the recipe wobble? core

Think of a dartboard. The bullseye is the true parameter. Each dart is the estimate from one imaginary repeat of the experiment. Bias asks: is the cloud of darts centred on the bullseye? Variance asks a different question: how tightly bunched is the cloud?

A thrower can be centred but scattered (unbiased, high variance), or tightly grouped but off to one side (biased, low variance). You need both questions to judge a thrower, and both to judge an estimator.

Three ways to say it:

  • Picture: variance is the width of the pile of estimates, wherever that pile is centred.
  • Numbers: with $\sigma = 10$, one day's orders wobbles with variance 100, but the mean of 25 days wobbles with variance only 4.
  • Slogan: bias is where you aim; variance is how steady your hand is.

Three unbiased recipes for the mean daily orders $\mu$. Days are independent with standard deviation $\sigma = 10$, and we have $n = 25$ days.

  1. First day only, $X_1$: $Var(X_1) = \sigma^2 = 100$, so it typically misses by about $\sqrt{100} = 10$ orders.
  2. Mean of the first 5 days: $Var\big(\tfrac15(X_1 + \dots + X_5)\big) = \tfrac{1}{5^2}\,(5 \times 100) = \tfrac{500}{25} = 20$. We used $Var(aX) = a^2 Var(X)$ and, because the days are independent, "the variance of a sum is the sum of the variances". Typical miss $\sqrt{20} \approx 4.47$.
  3. Mean of all 25 days: $Var(\bar X) = \tfrac{1}{25^2}(25 \times 100) = \tfrac{100}{25} = 4$. Typical miss $\sqrt 4 = 2$.
  4. All three are centred on $\mu$ (zero bias). Using all the data cuts the variance from 100 to 4.

The variance of an estimator is the variance of its sampling distribution:

$$Var(\hat\theta) = E\big[(\hat\theta - E[\hat\theta])^2\big].$$
  • It measures spread around the estimator's own centre $E[\hat\theta]$, not around the truth $\theta$.
  • Its square root is the standard error, $SE(\hat\theta) = \sqrt{Var(\hat\theta)}$ (Chapter 5.5).
  • For the mean of $n$ independent values with variance $\sigma^2$: $Var(\bar X) = \dfrac{1}{n^2}\sum_{i=1}^n Var(X_i) = \dfrac{\sigma^2}{n}$. This needs independence; with positively correlated values the true variance is larger.
Why do we need it?

An unbiased recipe can still be useless if it wobbles wildly. Variance tells us how much one estimate can be trusted and how much more data would help: it falls like $1/n$ for averages.

Where is it used?

Standard errors and confidence intervals, sample-size planning for A/B tests (noisier metrics need more users), variance-reduction methods like CUPED (Chapter 5.12), and the "variance" half of the bias–variance trade-off in model selection.

How is it used?

Compute it with the variance rules ($\sigma^2/n$ for a mean), estimate it from data by replacing $\sigma$ with $s$, or simulate many repeats and take the variance of the estimates. Then compare recipes: lower variance at the same bias is better.

low bias · low variance accurate and steady low bias · high variance right on average, wobbly high bias · low variance steady but off target high bias · high variance off target and wobbly green = truth θ · orange = estimates from repeated samples · purple ring = their average · red = bias
Bias is the distance from the bullseye to the centre of the darts; variance is how scattered the darts are around their own centre. They are two separate properties.

Each orange dart is the estimate from one imaginary repeat of the experiment; the green bullseye is the truth. Drag the purple handle to move where the recipe aims (its average, $E[\hat\theta]$), and use the slider to change how steady it is. Watch the readout: the average squared distance to the bullseye (MSE) always equals the squared distance from bullseye to the cloud's centre (bias²) plus the cloud's own spread (variance). Real estimators are usually single numbers; two dimensions just make the picture easier to see.

"Low variance means accurate."

Low variance means steady. A recipe can be very steady and steadily wrong (the third dartboard). Accuracy needs low bias and low variance; MSE (next section) combines them.

"The variance of the sample mean is $\sigma^2$."

$\sigma^2$ is the variance of one observation. The sample mean of $n$ independent observations has variance $\sigma^2/n$.

"$\sigma^2/n$ works for any data."

It assumes independent observations. Repeated sessions from the same user, users in the same city, or consecutive days of a time series are correlated, and then the true variance of the mean is bigger than $\sigma^2/n$ (Chapter 5.4, Chapter 5.10).

$Var(\hat\theta) = E[(\hat\theta - E\hat\theta)^2]$: spread around its own centre. $SE = \sqrt{Var}$.

For iid data: $Var(\bar X) = \sigma^2/n$.

Trap: low variance ≠ accurate (it can be steadily wrong); $\sigma^2/n$ needs independence.

Quick check: you average the first 4 of 16 independent days instead of all 16. By what factor is the variance bigger?

$\sigma^2/4$ versus $\sigma^2/16$: four times bigger. The standard error is $\sqrt4 = 2$ times bigger. Throwing away data costs variance.

Mean squared error: one score that combines bias and variance core

We now have two ways an estimator can be bad: it can aim in the wrong place (bias) or wobble (variance). To compare recipes we want one score that says how far off an estimate typically is from the truth.

The natural score: take each estimate's miss, $\hat\theta - \theta$, square it (so misses to the left and to the right both count), and average over all the samples you could have drawn. That is the mean squared error. Beautifully, it splits into exactly two pieces: the squared bias plus the variance. Nothing else.

Three ways to say it:

  • Picture: on the dartboard, the average squared distance to the bullseye = (distance from bullseye to the cloud's centre)² + (spread of the cloud).
  • Numbers: a recipe that is 1 too high on average and wobbles with variance 2 has MSE $1^2 + 2 = 3$.
  • Slogan: total error = aim error² + wobble.

The truth is $\theta = 10$. A recipe gives one of four estimates, 9, 11, 13 or 11, each equally likely (one per possible sample).

  1. Average estimate: $E[\hat\theta] = (9 + 11 + 13 + 11)/4 = 44/4 = 11$.
  2. Bias: $11 - 10 = 1$, so $Bias^2 = 1$.
  3. Variance (spread around 11): $\big((9-11)^2 + 0^2 + (13-11)^2 + 0^2\big)/4 = (4 + 0 + 4 + 0)/4 = 2$.
  4. MSE directly (misses from the truth 10): $\big((-1)^2 + 1^2 + 3^2 + 1^2\big)/4 = (1 + 1 + 9 + 1)/4 = 12/4 = 3$.
  5. Check: $Bias^2 + Var = 1 + 2 = 3$ ✓.

The mean squared error of $\hat\theta$ is $MSE(\hat\theta) = E\big[(\hat\theta - \theta)^2\big]$, and

$$MSE(\hat\theta) = Bias(\hat\theta)^2 + Var(\hat\theta).$$

Derivation. Write $m = E[\hat\theta]$ for the centre of the estimator. Add and subtract $m$ inside the miss:

  1. $\hat\theta - \theta = (\hat\theta - m) + (m - \theta)$. The first part is the wobble; the second part, $m - \theta$, is the bias, a fixed number.
  2. Square: $(\hat\theta-\theta)^2 = (\hat\theta - m)^2 + 2(\hat\theta - m)(m-\theta) + (m-\theta)^2$.
  3. Take expectations term by term: $E[(\hat\theta-m)^2] = Var(\hat\theta)$; $\;E[2(\hat\theta-m)(m-\theta)] = 2(m-\theta)\,E[\hat\theta - m] = 2(m-\theta)\cdot 0 = 0$; $\;E[(m-\theta)^2] = (m-\theta)^2 = Bias^2$.
  4. So $MSE = Var(\hat\theta) + 0 + Bias(\hat\theta)^2$. ∎
  • For an unbiased estimator, $MSE = Var$.
  • The root mean squared error $RMSE = \sqrt{MSE}$ is in the same units as $\theta$.
  • MSE is one choice of score. Others exist (mean absolute error, for example); MSE is popular because it splits so cleanly.
Why do we need it?

Bias alone and variance alone can each be gamed: the constant "always guess 5" has zero variance, and the first observation has zero bias. MSE scores the total error, so recipes that make different trade-offs can be compared on one scale.

Where is it used?

Choosing between estimators in simulation studies; choosing the ridge penalty or the amount of shrinkage; comparing forecasting models (MSE and RMSE of forecasts, Chapter 7.15); the bias–variance decomposition of prediction error in model selection (Chapter 7.18).

How is it used?

Compute bias and variance (by algebra or by simulation), add $Bias^2 + Var$, and prefer the recipe with the smaller MSE for the sample sizes you will actually have. Report RMSE when you want a number in the original units.

truth θ E[θ̂] bias spread → variance one estimate θ̂ its miss θ̂ − θ MSE = average of (miss)² = bias² + variance
The blue curve is the sampling distribution. Its centre sits a "bias" away from the truth, and its width is the variance. The average squared miss of single estimates splits exactly into these two parts.

Three recipes for the variance $\sigma^2 = 4$ of Normal data: divide the sum of squares by $n-1$ (unbiased), by $n$, or by $n+1$. Each bar is the recipe's MSE, split into bias² (red, bottom) and variance (blue, top). At $n = 5$ the unbiased recipe has the largest MSE; dividing by $n+1$, the most biased, is the most accurate. Slide $n$: the gap shrinks but never flips. Turn on Check by simulation to see the purple dots (simulated MSE) land on the bar tops.

"MSE is just the variance."

Only for unbiased estimators. In general $MSE = Bias^2 + Var$; a recipe with a small variance can still have a big MSE if it aims in the wrong place.

"The unbiased estimator always has the smallest error."

For Normal data, $s^2$ (divide by $n-1$) has a larger MSE than dividing by $n$ or $n+1$. Unbiasedness is a nice property, not a guarantee of accuracy.

"The MSE of an estimator and the MSE of a model's predictions are the same thing."

They share the idea. The MSE of predictions of new data points adds a third piece, the irreducible noise of the new observations, which no estimator can remove (Chapter 7.18).

$MSE(\hat\theta) = E[(\hat\theta-\theta)^2] = Bias^2 + Var$ (the cross term has mean 0).

$RMSE = \sqrt{MSE}$, in the units of $\theta$. Unbiased ⇒ $MSE = Var$.

Trap: the unbiased recipe is not automatically the most accurate ($s^2$ loses to ÷$(n+1)$ for Normal data).

Quick check: recipe A has bias 0 and variance 9; recipe B has bias 2 and variance 3. Which has the smaller MSE?

A: $0 + 9 = 9$. B: $2^2 + 3 = 7$. B is more accurate overall, even though it is biased.

The bias–variance trade-off: buy a little bias, save a lot of variance core

A small customer segment had 4 users in your test, and their average lift came out at +7. You know from experience that lifts are usually small. Do you believe +7?

A sensible person would say: "probably real lift is smaller; with 4 users the average jumps around a lot". So they pull the estimate toward a sensible default (zero, or the average over all segments). That pull adds a little bias (if the true lift really is big, we under-report it), but it removes a lot of the wobble. When the data are noisy, the trade is worth it.

Three ways to say it:

  • Picture: the MSE curve is a valley; the honest average sits on one wall, "always say zero" sits on the other, and the bottom is in between.
  • Numbers: halving a noisy average can cut its MSE from 4 to 2.
  • Slogan: when data are noisy, shrink toward something sensible.

Each user's lift is Normal with true mean $\theta = 2$ and standard deviation $\sigma = 4$. We have $n = 4$ users. Compare the shrunk recipes $\hat\theta_c = c\,\bar X$ for a few values of $c$ (keep the fraction $c$ of the average).

  1. Variance of the raw mean: $\sigma^2/n = 16/4 = 4$.
  2. For any $c$: $E[c\bar X] = c\theta$, so $Bias = c\theta - \theta = (c-1)\theta$, and $Var(c\bar X) = c^2\sigma^2/n = 4c^2$.
  3. $c = 1$ (the raw mean): $Bias^2 = 0$, $Var = 4$, $MSE = 4$.
  4. $c = 0.5$: $Bias = -0.5 \times 2 = -1$, $Bias^2 = 1$; $Var = 4 \times 0.25 = 1$; $MSE = 1 + 1 = 2$. Half the error of the raw mean.
  5. $c = 0$ (always say 0): $Bias^2 = 2^2 = 4$, $Var = 0$, $MSE = 4$. As bad as the raw mean, from the other side.
  6. $c = 0.75$ or $c = 0.25$: $MSE = 0.25 + 2.25 = 2.5$ and $2.25 + 0.25 = 2.5$. The best value is in the middle.

For the shrunk mean $\hat\theta_c = c\,\bar X$ (iid data, mean $\theta$, variance $\sigma^2$):

$$MSE(c) = \underbrace{(1-c)^2\theta^2}_{Bias^2} + \underbrace{c^2\,\frac{\sigma^2}{n}}_{Var}.$$

Set the derivative to zero: $-2(1-c)\theta^2 + 2c\,\sigma^2/n = 0$, which gives

$$c^* = \frac{\theta^2}{\theta^2 + \sigma^2/n}, \qquad MSE(c^*) = \frac{\theta^2\,\sigma^2/n}{\theta^2 + \sigma^2/n} \lt \min\!\Big(\theta^2, \frac{\sigma^2}{n}\Big).$$
  • Noisy data ($\sigma^2/n$ large compared with $\theta^2$): $c^*$ is small, shrink hard.
  • Plenty of data ($\sigma^2/n$ small): $c^* \approx 1$, barely shrink.
  • The catch: $c^*$ contains the unknown $\theta$. Real methods choose the amount of shrinkage another way: a prior (MAP and posterior means, Chapter 5.2), many groups learning it together (hierarchical models, Chapter 6.6), or cross-validation (ridge and lasso, Chapter 5.3). A famous result (Stein's paradox) shows that when you estimate three or more Normal means at once (with known noise level), the James–Stein recipe, which shrinks them all toward a fixed point by an amount computed from the data, has a smaller expected total squared error than the raw means, whatever the true means are. (Shrinking toward their own common average works the same way from four means on.)

The bias–variance trade-off is the general pattern: a knob that makes a method more flexible lowers bias and raises variance, and the best MSE is usually at a compromise.

Why do we need it?

With little data per quantity (small segments, rare holidays, many changepoints), raw estimates are so noisy that they are worse than a sensible default. Accepting some bias is the only way to get usable numbers.

Where is it used?

Ridge and lasso regression, partial pooling in hierarchical models, empirical Bayes, Laplace priors on changepoint slopes, smoothing a histogram or a KDE bandwidth, the Fourier order and the number of changepoints in a forecasting model, early stopping in neural networks.

How is it used?

Find the knob (prior strength, penalty λ, number of components). Estimate the error for several settings, by theory, simulation or held-out data, and choose the setting with the lowest total error, not the lowest bias.

The x-axis is the fraction $c$ of the average you keep ($c = 1$ is the raw mean). Red is bias², blue is variance, purple is their sum, the MSE, all measured in units of the raw mean's MSE (so the raw mean sits at 1). Start at $c = 1$ and slide $c$ down toward the green dot: the MSE falls to half. Now raise $\theta$ to 6 or $n$ to 50: the green dot moves right toward $c = 1$, so shrinking helps less. Press Check by simulation to test one point with real random samples.

Shrinking toward zero is a toy. In practice we shrink many noisy estimates toward their common average, and the data themselves tell us how hard to pull. The next widget previews that idea; you will meet it properly as partial pooling in Chapter 6.6.

Sixteen segments each have a true mean (green tick) and only $n$ users, so each raw average (blue) is noisy. The orange dots keep a fraction $c$ of each segment's distance from the overall average (purple line). With the ideal $c$, the total squared error usually drops a lot. Press New segments several times. Then raise $n$ to 60 or the true spread $\tau$ to 4: the ideal $c$ moves toward 1 and shrinking matters less. Set $c = 0$ to see "complete pooling" (everyone gets the overall average).

"The best $c^*$ is a recipe I can use."

$c^*$ needs the true $\theta$, which we do not know. It shows that shrinkage can help. Usable methods get the amount of shrinkage from a prior, from many groups at once (hierarchical models), or from held-out data.

"Shrinkage always helps."

Only the right amount helps. Shrink too hard (toward 0 when the truth is large, or with lots of data) and the bias dominates: try $\theta = 6$, $n = 50$, $c = 0.3$ in the first widget.

"Bias and variance are fixed properties of a method."

They depend on the sample size and on the truth. The same shrinkage that helps a 4-user segment can hurt a 4 000-user segment.

Both of your projects use deliberate, helpful bias. In an A/B framework like yours, hierarchical partial pooling pulls each segment's estimate toward the overall average, more strongly for segments with little data: exactly the second widget, except that the model learns $\tau$ (and so the right amount of pulling) from the data (Chapters 6.5–6.6). In your forecasting model, the Laplace prior $\delta_j \sim Laplace(0, b)$ pulls every changepoint slope adjustment toward 0; a small scale $b$ means strong shrinkage (low variance, risk of missing a real trend change), a large $b$ means weak shrinkage (flexible, risk of chasing noise). Choosing $b$ is a bias–variance decision (Chapters 5.3, 7.10, 7.18).

"We should always use unbiased estimators."

"We should use the estimator with the smallest error for the decisions we make. With small, noisy samples, a biased, shrunk estimator often has a much smaller MSE."

Model answer: "MSE equals bias squared plus variance. Shrinking a noisy estimate toward a sensible value adds a little bias but can remove a lot of variance. That is why ridge regression, partial pooling and shrinkage priors work: they accept bias on purpose to reduce total error."

$\hat\theta_c = c\bar X$: $MSE(c) = (1-c)^2\theta^2 + c^2\sigma^2/n$; best $c^* = \theta^2/(\theta^2 + \sigma^2/n)$.

Noisy data → shrink more. Lots of data → shrink less. The amount needs a prior, pooling, or validation.

Trap: "unbiased" is not the goal; low total error is. But too much shrinkage hurts.

Quick check: $\theta = 1$, $\sigma = 2$, $n = 4$. What are $c^*$ and the two MSEs?

$\sigma^2/n = 4/4 = 1$. $c^* = 1/(1+1) = 0.5$. Raw mean: $MSE = 1$. Shrunk: $(1-0.5)^2\cdot1 + 0.25\cdot1 = 0.25 + 0.25 = 0.5$. Half the error.

Consistency: does more data lead you to the truth? core

Bias and variance describe a recipe at one sample size. Consistency asks about the future: if you could keep collecting data forever, would the recipe close in on the truth?

A consistent recipe is like a navigation app that keeps getting more GPS signals: the more signals, the closer it puts you to where you really are, until the error is as small as you like. An inconsistent recipe is like a map with the wrong scale: no amount of extra driving fixes it.

Three ways to say it:

  • Picture: as $n$ grows, the pile of estimates gets narrower and slides onto the truth, until almost all of it sits inside any small band around $\theta$.
  • Numbers: the chance that a sample mean misses by more than 2 falls from 69% ($n = 4$) to 0.006% ($n = 400$).
  • Slogan: consistent = enough data makes the error as small as you like, with probability close to 1.

Daily orders are Normal with standard deviation $\sigma = 10$. We use $\bar X$ and ask how likely it is to miss $\mu$ by more than $\varepsilon = 2$ orders.

  1. $\bar X$ is Normal with centre $\mu$ and standard error $SE = 10/\sqrt n$.
  2. A miss bigger than 2 means $|Z| \gt 2/SE$, where $Z$ is standard Normal. So $P(\text{miss}) = 2\,(1 - \Phi(2/SE))$, with $\Phi$ the standard Normal CDF.
  3. $n = 4$: $SE = 5$, $2/5 = 0.4$, $P = 2(1 - \Phi(0.4)) \approx 0.689$.
  4. $n = 25$: $SE = 2$, $2/2 = 1$, $P \approx 0.317$.
  5. $n = 100$: $SE = 1$, $2/1 = 2$, $P \approx 0.0455$.
  6. $n = 400$: $SE = 0.5$, $2/0.5 = 4$, $P \approx 0.000063$.

The probability of missing by more than 2 heads to zero. The same is true for any $\varepsilon$, however small: that is consistency. (Without assuming Normal data, Chebyshev's inequality still gives $P \le \sigma^2/(n\varepsilon^2) = 25/n$, which also goes to 0.)

A sequence of estimators $\hat\theta_n$ (one for each sample size $n$) is consistent for $\theta$ if, for every $\varepsilon \gt 0$,

$$P\big(|\hat\theta_n - \theta| \gt \varepsilon\big) \;\to\; 0 \quad \text{as } n \to \infty.$$

We write $\hat\theta_n \xrightarrow{p} \theta$ ("converges in probability to $\theta$").

  • A handy sufficient condition: if $MSE(\hat\theta_n) \to 0$ (bias → 0 and variance → 0), the estimator is consistent. Proof in one line with Markov's inequality applied to $(\hat\theta_n - \theta)^2$: $P(|\hat\theta_n - \theta| \gt \varepsilon) \le MSE/\varepsilon^2 \to 0$ (Chapter 4.12).
  • $\bar X$ is consistent for $\mu$ whenever the mean exists: this is the Law of Large Numbers (Chapter 4.13).
  • Consistency is a promise about large $n$; it says nothing about how good the recipe is at your $n$.
Why do we need it?

It is the minimum we ask of a recipe: if even infinite data would not reveal the truth, the recipe is broken. It also separates "fixable by collecting more data" problems (variance) from "not fixable" ones (an inconsistent recipe, a wrong model).

Where is it used?

Justifying the sample mean, the sample variance, maximum likelihood estimators and posterior means as sensible recipes; checking that a fitted model recovers known parameters from simulated data ("parameter recovery" tests of a NumPyro model); proofs that confidence intervals work for large $n$.

How is it used?

Check that bias and variance both shrink to 0 as $n$ grows (formulas or simulation). For a new model, simulate data with known parameters at growing $n$ and confirm the estimates close in on the truth.

unbiased biased consistent not consistent sample mean X̄ centred on θ, narrows to a spike Σx/(n+1), ÷n variance off-centre, but bias → 0 and width → 0 first value X₁ centred on θ, but never narrows 0.8 × X̄ narrows onto the wrong value 0.8θ
Green line = truth θ. Dashed curve = sampling distribution at small n; solid shaded curve = at large n. Unbiased and consistent are different properties: every combination happens.

Four recipes for a mean $\theta = 5$ (green line), data with $\sigma = 4$. Each row is a recipe's sampling distribution (scaled to the same height); red shading is the probability of missing by more than $\varepsilon$ (outside the green band). Slide $\log_{10} n$ from 0 to 3 (that is $n$ from 1 to 1000). The mean and $\Sigma x/(n+1)$ squeeze into the band: their miss probability goes to 0. The first value never narrows. $0.8\bar X$ narrows, but onto 4, outside the band, so its miss probability goes to 1.

"Unbiased means consistent."

The first value $X_1$ is unbiased for every $n$ but never improves, so it is not consistent.

"Consistent means unbiased."

$\frac1n\sum(X_i - \bar X)^2$ and $\Sigma X_i/(n+1)$ are biased at every finite $n$, yet both are consistent because their bias and variance shrink to 0.

"My estimator is consistent, so my estimate is good."

Consistency is about $n \to \infty$. At your $n$ the estimate can still be noisy or biased. It also relies on assumptions: for very heavy tails (Cauchy data) the sample mean is not consistent, and dependent data (one user's repeated sessions, a time series) make the sample mean much noisier than $\sigma/\sqrt n$ suggests, and very persistent series settle even more slowly.

"Unbiased and consistent mean the same thing."

"Unbiased is about the centre of the sampling distribution at a fixed $n$. Consistent is about the whole distribution collapsing onto the truth as $n$ grows."

Model answer: "$X_1$ is unbiased but not consistent: its spread never shrinks. The divide-by-$n$ variance is biased but consistent: its bias $-\sigma^2/n$ and its variance both go to 0. A sufficient condition for consistency is that the MSE goes to 0."

Consistent: $P(|\hat\theta_n - \theta| \gt \varepsilon) \to 0$ for every $\varepsilon \gt 0$. Sufficient: $MSE \to 0$ (Markov).

Unbiased but not consistent: $X_1$. Biased but consistent: $\frac1n\sum(x_i-\bar x)^2$.

Trap: consistency is a large-$n$ promise that needs assumptions (finite mean, enough independence).

Quick check: is $\bar X + 1/n$ a consistent estimator of $\mu$? Is it unbiased?

Its bias is $1/n$, which goes to 0, and its variance $\sigma^2/n$ goes to 0, so the MSE goes to 0: consistent. But $E = \mu + 1/n \ne \mu$, so it is biased at every finite $n$.

Efficiency: who squeezes the most out of the data? core

Two unbiased recipes, same data. If one wobbles less, it is getting more information out of the same observations. We call it more efficient. The less efficient recipe needs more data to reach the same precision, and data cost money and time.

The classic contest is the mean against the median for the centre of a symmetric distribution. For Normal data the mean wins: the median throws away information about how far points are from the middle. But for heavy-tailed data, where occasional wild values appear, the mean gets dragged around by the wild values and the median wins.

Three ways to say it:

  • Picture: two piles of estimates, both centred on the truth; the narrower pile belongs to the more efficient recipe.
  • Numbers: for Normal data the median needs about 157 observations to match the precision of the mean with 100 (relative efficiency $2/\pi \approx 0.64$).
  • Slogan: efficiency = precision per data point; which recipe wins depends on the tails.

Normal data. Daily orders are Normal with $\sigma = 10$; we have $n = 100$ days.

  1. Mean: $Var(\bar X) = \sigma^2/n = 100/100 = 1$.
  2. Median: for large $n$ and Normal data, $Var(\text{median}) \approx \dfrac{\pi}{2}\cdot\dfrac{\sigma^2}{n} = 1.571 \times 1 = 1.571$.
  3. Relative efficiency of the median to the mean: $\dfrac{Var(\bar X)}{Var(\text{median})} = \dfrac{1}{1.571} = \dfrac{2}{\pi} \approx 0.637$.
  4. Meaning: the median with $n$ points is as precise as the mean with $0.637n$ points. To match the mean's precision with 100 days, the median needs $100 \times 1.571 \approx 157$ days.

Heavy-tailed data (Student-t with 3 degrees of freedom, scale 1, large $n$):

  1. Mean: $Var(\bar X) = \dfrac{Var(X)}{n} = \dfrac{3/(3-2)}{n} = \dfrac{3}{n}$.
  2. Median: $Var(\text{median}) \approx \dfrac{1}{4 n f(0)^2}$, where $f(0) \approx 0.3676$ is the density at the centre. So $\approx \dfrac{1}{4 \times 0.1351\,n} = \dfrac{1.85}{n}$.
  3. Relative efficiency of the median to the mean: $\dfrac{3/n}{1.85/n} \approx 1.62$. Now the median is the efficient one: the mean would need about 62% more data.

For two unbiased estimators $\hat\theta_1$ and $\hat\theta_2$ of the same $\theta$, the relative efficiency of $\hat\theta_1$ with respect to $\hat\theta_2$ is

$$\text{eff}(\hat\theta_1, \hat\theta_2) = \frac{Var(\hat\theta_2)}{Var(\hat\theta_1)}.$$

Above 1 means $\hat\theta_1$ is more precise. (For biased estimators, compare MSEs instead.) The limit as $n \to \infty$ is the asymptotic relative efficiency (ARE). It reads as a data ratio: an ARE of 0.64 means "needs about $1/0.64 \approx 1.57$ times as much data".

  • Median versus mean (large $n$, symmetric data): $Var(\text{median}) \approx 1/(4 n f(m)^2)$, with $f(m)$ the density at the median $m$. ARE of median to mean $= 4 f(m)^2\,\sigma^2$: $2/\pi \approx 0.64$ for Normal data, $2$ for Laplace data, about $1.62$ for $t_3$, and it crosses 1 at about $\nu = 4.7$ degrees of freedom.
  • The best possible: under regularity conditions (roughly: a smooth model, and a true value not on the edge of the allowed range), no unbiased estimator can have variance below the Cramér–Rao lower bound $1/(n\,I(\theta))$, where $I(\theta)$ is the Fisher information of one observation (how sharply the likelihood peaks; Chapter 5.2). An unbiased estimator that reaches the bound is called efficient. For Normal data, $I(\mu) = 1/\sigma^2$, the bound is $\sigma^2/n$, and $\bar X$ reaches it.
  • Maximum likelihood estimators reach the bound approximately for large $n$ (they are "asymptotically efficient") when the model is correct.
Why do we need it?

An inefficient recipe silently wastes data: your experiment needs more users or more days to reach the same confidence. Efficiency also explains when "robust" recipes are worth their small cost under Normal data, because they save you under heavy tails.

Where is it used?

Choosing mean versus median or trimmed mean for skewed revenue metrics; robust regression; Student-t likelihoods for heavy-tailed noise in forecasting; the justification of maximum likelihood (asymptotically efficient); variance-reduction tricks such as CUPED, which make the same experiment more efficient.

How is it used?

Simulate both recipes on data shaped like yours (or use the formulas), compare their variances, and read the ratio as "how much data does the loser need?". If you are unsure about the tails, a trimmed mean or a Student-t model gives up a little under Normal data and gains a lot under heavy tails.

Top: the piles of simulated means (blue) and medians (orange) for samples of size $n$; the bars above show ± one standard error. Start with Normal: the blue pile is narrower. Switch to Student-t₃: now the orange pile is narrower. Bottom: the large-$n$ efficiency of the median relative to the mean for Student-t data, as a function of the degrees of freedom $\nu$ (smaller $\nu$ = heavier tails). Above the line at 1, the median wins; the crossing is near $\nu = 4.7$. Try Cauchy and press New sample a few times: the mean's spread jumps around wildly.

"The median is always a worse estimator than the mean."

Only for light-tailed data such as the Normal. For heavy tails (Student-t with few degrees of freedom, Laplace, contaminated data) the median, or a trimmed mean, is more efficient.

"An efficient estimator is one that runs fast."

Statistical efficiency is about variance (precision per data point), not computing time.

"The median and the mean estimate the same thing."

Only for symmetric distributions. For skewed data (revenue, waiting times) the population median and mean are different numbers, so comparing their variances is comparing answers to different questions. Decide the target first.

Your forecasting model offers a Student-t likelihood next to the Normal one, and your A/B framework supports Student-t for continuous metrics. This section is the reason: when residuals or revenue values have heavy tails, a Normal model behaves like the sample mean and gets pulled around by extreme values, while a Student-t model behaves more like a robust estimator. Say it precisely: Student-t does not remove outliers; it assigns more probability to extreme values, so they exert less influence on the fit (Chapter 7.13).

$\text{eff}(\hat\theta_1,\hat\theta_2) = Var(\hat\theta_2)/Var(\hat\theta_1)$: a data ratio.

Median vs mean (large $n$): $2/\pi \approx 0.64$ for Normal, $\approx 1.62$ for $t_3$, $2$ for Laplace; tie near $\nu \approx 4.7$.

Cramér–Rao: $Var(\hat\theta) \ge 1/(nI(\theta))$ for unbiased recipes; $\bar X$ reaches it for Normal data. Trap: which recipe is efficient depends on the tails.

Quick check: recipe A has variance 2/n and recipe B has variance 3/n (both unbiased). How many observations does B need to match A with 200?

A with 200: $2/200 = 0.01$. B needs $3/n = 0.01$, so $n = 300$. Efficiency of B relative to A $= (2/n)/(3/n) = 2/3$; B needs $1.5\times$ the data.

Putting it together: the estimator race

We now have a full scorecard: bias, variance and their sum, the MSE. Let's put five recipes on the same track and race them on different kinds of data. The lesson is surprising at first and obvious afterwards: no recipe wins everywhere. The winner depends on the shape of the data, the sample size and even the true value.

Three ways to say it:

  • Picture: a bar for each recipe, red for bias² and blue for variance; the shortest bar wins.
  • Numbers: with Normal data the shrunk mean scores 1.18 against the mean's 1.60, but with heavy tails the trimmed mean scores 0.88.
  • Slogan: choose the recipe for the data you expect, not for the data you wish you had.

Normal data, true mean $\theta = 2$, $\sigma = 4$, $n = 10$. Exact scores for three recipes:

  1. Mean: bias 0, variance $16/10 = 1.6$, MSE $= 1.6$.
  2. First value: bias 0, variance $16$, MSE $= 16$.
  3. Shrunk mean $0.8\bar X$: bias $= -0.2 \times 2 = -0.4$, so $Bias^2 = 0.16$; variance $= 0.8^2 \times 1.6 = 1.024$; MSE $= 0.16 + 1.024 = 1.184$.
  4. The median and the 20% trimmed mean have no simple exact formula at $n = 10$; a simulation gives MSE ≈ 2.21 and ≈ 1.81.
  5. Ranking: shrunk (1.18) < mean (1.60) < trimmed (1.81) < median (2.21) < first value (16). Change the data to Student-t₃ with the same standard deviation and the ranking changes: trimmed ≈ 0.88 and median ≈ 0.96 now beat the mean's 1.60.

A fair comparison of estimators fixes four things: the target parameter $\theta$, the data-generating process, the sample size $n$, and the score (here MSE). With $R$ simulated samples and estimates $\hat\theta^{(1)}, \dots, \hat\theta^{(R)}$:

$$\widehat{Bias} = \frac1R\sum_{r}\hat\theta^{(r)} - \theta, \qquad \widehat{Var} = \frac1R\sum_r\big(\hat\theta^{(r)} - \bar{\hat\theta}\big)^2, \qquad \widehat{MSE} = \widehat{Bias}^2 + \widehat{Var}.$$

These are Monte Carlo estimates (estimates made by simulation): they have their own small random error, which shrinks like $1/\sqrt R$.

Why do we need it?

Formulas exist only for simple recipes. A simulation race lets you compare any recipes, including medians, trimmed means, shrinkage and full Bayesian models, on data that look like yours, before you trust one of them with real decisions.

Where is it used?

Simulation studies in statistics papers; choosing a metric summary for a heavy-tailed revenue metric; "parameter recovery" checks of a NumPyro model (simulate data with known parameters, fit, compare); choosing between Normal and Student-t likelihoods.

How is it used?

Write a function that simulates one dataset from a believable process, apply every recipe, repeat a few thousand times, and tabulate bias, variance and MSE. Repeat for a few plausible processes (light tails, heavy tails, skew) and prefer a recipe that is never badly beaten.

Each bar is one recipe's MSE: red = bias², blue = variance; ★ marks the winner. Start with Normal data: the shrunk mean wins at $\theta = 2$. Now slide $\theta$ to 6: shrinking toward 0 now costs too much bias, and the mean wins. Switch to Student-t₃ or wild values: the trimmed mean and the median overtake the mean. Switch to Skewed: the median gets a big red bias part, because for skewed data the median is not the mean. Press New sample to see the simulation wobble.

"The simulation's numbers are exact."

They are Monte Carlo estimates. Press New sample and the second decimal moves. Use enough repetitions, and do not crown a winner on a difference smaller than the wobble.

"Pick the recipe that gives the nicest number on my dataset."

That is choosing after seeing the answer, and it biases your results. Choose the recipe before looking, based on how it performs on data like yours.

"The shrunk mean is the best recipe."

It wins only when the truth is near the point you shrink toward (here 0). Move $\theta$ away and it loses. A recipe that is never badly beaten (the mean for light tails, a trimmed mean when tails might be heavy) is often the safer choice.

Revenue per user in an A/B framework like yours is typically skewed and heavy-tailed: most users spend 0, a few spend a lot. Run this race in your head before choosing a summary. The raw mean is unbiased for the mean but noisy; a trimmed mean or a median is steadier but targets a different number; a Student-t or a hierarchical model changes both the bias and the variance. A simulation race on data shaped like yours is the honest way to decide.

Compare recipes by MSE = bias² + variance, at your $n$, on data like yours (simulate $R$ samples).

Normal data: mean ≈ best unbiased; shrinkage can win if the truth is near the target. Heavy tails: trimmed mean / median win. Skewed: decide the target first.

Trap: no recipe wins everywhere; Monte Carlo numbers wobble.

Quick check: why does the median get a large red (bias²) bar for the skewed data?

For a right-skewed distribution the population median is below the population mean (here by about 1.2). The sample median estimates the median, so as an estimator of the mean it is systematically too low: that is bias, and more data will not remove it.

Recap, cheat sheet and practice

  • A parameter is a fixed unknown number; a statistic is any number computed from the sample; an estimator is a statistic used as a recipe to guess the parameter; an estimate is the number the recipe gave on your data.
  • An estimator is a random variable. Its distribution over repeated samples is the sampling distribution; its standard deviation is the standard error.
  • Bias $= E[\hat\theta] - \theta$ (where the pile is centred). Variance = spread of the pile around its own centre.
  • $MSE = Bias^2 + Var$. The unbiased recipe is not always the most accurate: for Normal data, dividing the sum of squares by $n+1$ beats $n-1$.
  • The bias–variance trade-off: shrinking a noisy estimate toward a sensible value adds bias but can remove much more variance. Ridge, partial pooling and Laplace priors all do this.
  • Consistent: the estimate closes in on the truth as $n \to \infty$. Unbiased and consistent are different properties ($X_1$; $\frac1n\sum(x_i-\bar x)^2$).
  • Efficiency compares variances: median vs mean is $2/\pi \approx 0.64$ for Normal data, but the median wins for heavy tails (about 1.62 for $t_3$).

Cheat sheet

IdeaFormulaIn words
Bias$E[\hat\theta] - \theta$how far the centre of the pile is from the truth
Variance$E[(\hat\theta - E\hat\theta)^2]$how wide the pile is
Standard error$\sqrt{Var(\hat\theta)}$; $\sigma/\sqrt n$ for $\bar X$typical wobble of the estimate
MSE$E[(\hat\theta-\theta)^2] = Bias^2 + Var$average squared miss
Divide-by-$n$ variance$E = \frac{n-1}{n}\sigma^2$biased low by $\sigma^2/n$; consistent
Shrunk mean $c\bar X$$MSE = (1-c)^2\theta^2 + c^2\sigma^2/n$, $c^* = \frac{\theta^2}{\theta^2+\sigma^2/n}$noisy data → shrink more
Consistency$P(|\hat\theta_n-\theta| \gt \varepsilon) \to 0$; enough: $MSE \to 0$enough data finds the truth
Relative efficiency$Var(\hat\theta_2)/Var(\hat\theta_1)$data ratio between two recipes
Median vs mean (large $n$)$4f(m)^2\sigma^2$: $2/\pi$ (Normal), $\approx1.62$ ($t_3$), 2 (Laplace)tails decide the winner
Cramér–Rao bound$Var(\hat\theta) \ge 1/(nI(\theta))$ (unbiased)the best possible precision
Code it · Python

import numpy as np

rng = np.random.default_rng(0)

# 1) Bias, variance and MSE of five recipes, by simulation (Monte Carlo)
theta, sigma, n, R = 2.0, 4.0, 10, 200_000
X = rng.normal(theta, sigma, size=(R, n))       # R fake datasets, each of size n
xs = np.sort(X, axis=1)
recipes = {
    "mean":       X.mean(axis=1),
    "median":     np.median(X, axis=1),
    "trim 20%":   xs[:, 2:n-2].mean(axis=1),     # drop 2 smallest and 2 largest
    "first":      X[:, 0],
    "0.8 * mean": 0.8 * X.mean(axis=1),
}
for name, est in recipes.items():
    bias = est.mean() - theta
    var = est.var()                              # spread of the estimates
    mse = np.mean((est - theta) ** 2)            # = bias**2 + var
    print(f"{name:10s} bias={bias:+.3f}  var={var:.3f}  mse={mse:.3f}")
# mean       bias=+0.004  var=1.597  mse=1.597
# median     bias=+0.004  var=2.209  mse=2.209
# trim 20%   bias=+0.004  var=1.809  mse=1.809
# first      bias=-0.000  var=15.995  mse=15.995
# 0.8 * mean bias=-0.397  var=1.022  mse=1.180    <- biased, yet the most accurate here

# 2) Dividing by n-1, n or n+1: the unbiased recipe has the LARGEST MSE (Normal data)
n, R = 5, 200_000
X = rng.normal(0, 2, size=(R, n))                # sigma^2 = 4
ss = ((X - X.mean(axis=1, keepdims=True)) ** 2).sum(axis=1)
for d in (n - 1, n, n + 1):
    est = ss / d
    print(f"divide by {d}: bias={est.mean() - 4:+.3f}  mse={np.mean((est - 4) ** 2):.3f}")
# divide by 4: bias=-0.003  mse=8.006     (theory: 0 and 8)
# divide by 5: bias=-0.802  mse=5.767     (theory: -0.8 and 5.76)
# divide by 6: bias=-1.335  mse=5.341     (theory: -1.333 and 5.333)
# Library conventions: np.var(x) divides by n (ddof=0); np.var(x, ddof=1) and pandas .var() divide by n-1

# 3) Consistency: P(|mean - mu| > 2) shrinks as n grows (sigma = 10)
for n in (4, 25, 100, 400):
    means = rng.normal(50, 10, size=(20_000, n)).mean(axis=1)
    print(n, np.mean(np.abs(means - 50) > 2))
# 4 0.68975   25 0.31695   100 0.0431   400 0.0   (theory: 0.689, 0.317, 0.046, 0.00006)

# 4) Efficiency flip: median vs mean, Normal vs Student-t with 3 df (n = 25)
for name, draw in [("normal", lambda s: rng.normal(size=s)),
                   ("t3", lambda s: rng.standard_t(3, size=s))]:
    X = draw((100_000, 25))
    print(name, "Var(mean)/Var(median) =", round(X.mean(axis=1).var() / np.median(X, axis=1).var(), 3))
# normal Var(mean)/Var(median) = 0.649   (large-n theory 2/pi = 0.637: the mean wins)
# t3 Var(mean)/Var(median) = 1.614       (large-n theory 1.62: the median wins)
Test yourself

1. 500 users saw a page and 53 bought. Which of these is the estimator?

θ is the parameter, 0.106 is the estimate, the 500 users are the sample. The estimator is the recipe K/n, a function of the random data.

2. A recipe has bias −1 and variance 3. Its MSE is…

$MSE = Bias^2 + Var = (-1)^2 + 3 = 4$. The sign of the bias does not matter because it is squared.

3. Which estimator of a mean μ is unbiased but not consistent?

$E[X_1] = \mu$ at every $n$, but its variance stays $\sigma^2$ forever, so it never closes in on μ. $\bar X$ is unbiased and consistent; $\Sigma X_i/(n+1)$ is biased but consistent; $0.8\bar X$ is biased and inconsistent.

4. For large samples of Normal data, the efficiency of the median relative to the mean is about…

$Var(\text{median}) \approx \frac{\pi}{2}\frac{\sigma^2}{n}$, so the ratio is $2/\pi \approx 0.64$: the median needs about 57% more data. 1.62 is the value for Student-t with 3 degrees of freedom, and 2 is the value for Laplace data.

5. Why can the shrunk estimate $c\bar X$ with $c \lt 1$ have a smaller MSE than $\bar X$?

$MSE(c) = (1-c)^2\theta^2 + c^2\sigma^2/n$. When $\sigma^2/n$ is large compared with $\theta^2$, the drop in variance outweighs the added bias². This is the bias–variance trade-off.

6. $s^2$ is unbiased for $\sigma^2$. What can you say about $s$ as an estimator of $\sigma$?

The square root is concave, so $E[\sqrt{s^2}] \lt \sqrt{E[s^2]} = \sigma$ (Jensen). For Normal data $E[s] = c_4\sigma$ with $c_4 \approx 0.94$ at $n = 5$; the bias shrinks with $n$ but is never exactly 0.

Practice problems

A. Data are Uniform(0, θ) and you see 2, 7, 5, 3. Compute the estimates from $2\bar X$ and from $\frac{n+1}{n}\max$. Both recipes are unbiased; which has the smaller variance?

$\bar x = 17/4 = 4.25$, so $2\bar x = 8.5$. $\max = 7$, so $\frac54 \times 7 = 8.75$.

Variances: one value has $Var(X) = \theta^2/12$, so $Var(2\bar X) = 4\cdot\frac{\theta^2}{12n} = \frac{\theta^2}{3n} = \frac{\theta^2}{12}$ at $n = 4$. For the max recipe, $Var = \frac{\theta^2}{n(n+2)} = \frac{\theta^2}{24}$. The max-based recipe has half the variance: it is twice as efficient here (and the gap grows with $n$, since it falls like $1/n^2$ instead of $1/n$).

B. Show that $E\big[\sum(X_i - \bar X)^2\big] = (n-1)\sigma^2$ for iid data.
  1. Expand: $\sum(X_i - \bar X)^2 = \sum X_i^2 - 2\bar X\sum X_i + n\bar X^2 = \sum X_i^2 - n\bar X^2$ (because $\sum X_i = n\bar X$).
  2. $E[X_i^2] = Var(X_i) + \mu^2 = \sigma^2 + \mu^2$, so $E[\sum X_i^2] = n\sigma^2 + n\mu^2$.
  3. $E[\bar X^2] = Var(\bar X) + \mu^2 = \sigma^2/n + \mu^2$, so $E[n\bar X^2] = \sigma^2 + n\mu^2$.
  4. Subtract: $n\sigma^2 + n\mu^2 - \sigma^2 - n\mu^2 = (n-1)\sigma^2$. Dividing by $n-1$ gives an unbiased recipe; dividing by $n$ gives $E = \frac{n-1}{n}\sigma^2$.
C. For a conversion rate with $n = 10$ users, compare the MSE of $K/n$ and $(K+1)/(n+2)$ when $p = 0.5$ and when $p = 0.05$.

$K/n$ is unbiased with $MSE = p(1-p)/n$: $0.025$ at $p = 0.5$ and $0.00475$ at $p = 0.05$.

$(K+1)/(n+2)$ has bias $\frac{1-2p}{n+2}$ and variance $\frac{np(1-p)}{(n+2)^2}$. At $p = 0.5$: bias 0, variance $2.5/144 \approx 0.0174$: it wins (0.0174 vs 0.025). At $p = 0.05$: bias $0.9/12 = 0.075$, bias² $\approx 0.0056$, variance $0.475/144 \approx 0.0033$, MSE $\approx 0.0089$: it loses (0.0089 vs 0.00475). Shrinking toward ½ helps only when the truth is near ½; for rare conversions it hurts. The pull should point toward a sensible value for your metric, which is what a well-chosen Beta prior does (Chapter 6.3).

D. Is "the average of the first 10 observations" a consistent estimator of μ? Is it unbiased?

Unbiased: yes, $E = \mu$. Consistent: no. However large $n$ gets, it uses only 10 values, so its variance stays $\sigma^2/10$ and the probability of missing by more than a small $\varepsilon$ does not go to 0.

E. (Interview) "Why do we divide by $n-1$, and does it make the standard deviation unbiased?"

"The squared distances are measured from the sample mean, which is pulled toward the data, so they are too small on average: $E[\sum(x_i-\bar x)^2] = (n-1)\sigma^2$. Dividing by $n-1$ makes the variance unbiased. It does not make the standard deviation unbiased, because the square root is curved (Jensen): $s$ is slightly too small on average. And unbiased is not the same as most accurate: dividing by $n+1$ gives a smaller MSE for Normal data. For large $n$ none of this matters much."

F. A heavy-tailed metric looks like Student-t with 3 degrees of freedom. With the median, you need 10 000 observations for a target precision of the centre. Roughly how many would the mean need?

The large-$n$ efficiency of the median relative to the mean is about 1.62 for $t_3$, so the mean needs about $1.62 \times 10\,000 \approx 16\,200$ observations for the same variance. (For Normal data it would be the other way round: the median would need about $1.57\times$ as many as the mean.) Both estimate the same centre here because the distribution is symmetric.

Chapter 5.2 · Syllabus Module 10

Maximum likelihood and MAP

Where do estimators come from? Mostly from one idea: choose the parameter value that best explains the data you actually saw. That is maximum likelihood. Add what you believed before seeing the data and you get MAP. Along the way you will discover that squared error, cross-entropy and absolute error are not arbitrary losses: each one is a likelihood in disguise.

  • Read the likelihood $L(\theta) = p(D \mid \theta)$ as a function of $\theta$ with the data fixed, and explain why it is not a probability distribution over $\theta$
  • Use the log-likelihood (sums instead of products, no underflow, same peak)
  • Derive the maximum likelihood estimates for a coin ($k/n$), a Normal ($\bar x$ and a variance that divides by $n$) and a Poisson ($\bar y$)
  • Read precision from the curvature of the log-likelihood, and know the large-sample properties of the MLE
  • See that minimizing a negative log-likelihood gives the familiar losses: Gaussian ↔ squared error, Bernoulli ↔ cross-entropy, Laplace ↔ absolute error
  • Compute MAP estimates (prior × likelihood), including $(k+\alpha-1)/(n+\alpha+\beta-2)$ for a Beta prior; compare MAP with the posterior mean; and explain why MAP, unlike the MLE, changes when you reparameterize

What we need from earlier chapters: estimators, bias, MSE and efficiency (Chapter 5.1); the Binomial, Poisson, Normal, Laplace and Beta distributions (Chapters 4.7–4.11); Bayes' theorem $P(A\mid B) = P(B\mid A)P(A)/P(B)$ (Chapter 4.3); derivatives and setting them to zero to find a maximum (Calculus 2.3, Optimization 3.2). Notation: $P(\cdot)$ is a probability of an event; $p(\cdot)$ is a probability mass function (for counts) or a density (for continuous values). $p(D\mid\theta)$ reads "the probability (or density) of the data $D$ if the parameter were $\theta$". "log" always means the natural logarithm $\ln$.

The likelihood: how well does each θ explain the data? core

A detective has one piece of evidence, and it cannot change: it is what happened. For each suspect she asks: "If this person did it, how likely is it that we would see exactly this evidence?" Suspects whose story makes the evidence likely move up the list; suspects whose story makes it very unlikely move down.

The likelihood is that score, worked out for every possible value of the parameter. The data are the evidence (fixed); the parameter values are the suspects (we try them all). A high likelihood means "this value of θ explains the data well". It does not mean "this value of θ is probably true": that would need a prior, and it is a different question.

Three ways to say it:

  • Picture: a curve over all possible θ, high where θ explains the data well and low where it does not.
  • Numbers: after 7 successes in 10 tries, θ = 0.7 gives the data probability 0.267 and θ = 0.5 gives 0.117, so 0.7 explains the data about 2.3 times better.
  • Slogan: probability fixes θ and varies the data; likelihood fixes the data and varies θ.

A new onboarding flow. 10 test users tried it and 7 completed it. Let $\theta$ be the true completion rate. If the users are independent, the number of completions is Binomial, so

$$p(k = 7 \mid \theta) = \binom{10}{7}\theta^7(1-\theta)^3 = 120\,\theta^7(1-\theta)^3.$$
  1. $\theta = 0.5$: $120 \times 0.5^{10} = 120/1024 \approx 0.117$.
  2. $\theta = 0.7$: $0.7^7 \approx 0.08235$ and $0.3^3 = 0.027$, so $120 \times 0.08235 \times 0.027 \approx 0.267$.
  3. $\theta = 0.9$: $0.9^7 \approx 0.4783$ and $0.1^3 = 0.001$, so $120 \times 0.4783 \times 0.001 \approx 0.0574$.
  4. Likelihood ratio of 0.7 against 0.5: $0.267/0.117 \approx 2.28$. The data are about 2.3 times as probable under $\theta = 0.7$.
  5. Add up (integrate) the curve over all $\theta$ from 0 to 1: $\int_0^1 120\,\theta^7(1-\theta)^3\,d\theta = 1/11 \approx 0.091$. Not 1. The likelihood is not a probability distribution over $\theta$.

Let $p(D\mid\theta)$ be the model: the probability (or density) of data $D$ when the parameter is $\theta$. The likelihood function is the same formula read the other way round:

$$L(\theta) = L(\theta; D) = p(D \mid \theta), \quad \text{viewed as a function of } \theta \text{ with } D \text{ fixed}.$$

For independent observations $x_1, \dots, x_n$ the probabilities multiply:

$$L(\theta) = \prod_{i=1}^n p(x_i \mid \theta).$$
  • $L(\theta)$ is not a probability distribution over $\theta$: it need not add or integrate to 1, and it says nothing about how plausible $\theta$ was before the data.
  • Only ratios $L(\theta_1)/L(\theta_2)$ matter. Factors that do not depend on $\theta$ (like $\binom{10}{7}$) can be dropped without changing any comparison.
  • For continuous data, $p(x_i\mid\theta)$ is a density, so likelihood values can be bigger than 1. That is fine: they are scores, not probabilities.
Why do we need it?

We need a way to say which parameter values are supported by the data and which are not. The likelihood turns any probability model into such a score, for every possible θ, using nothing but the model and the data.

Where is it used?

Maximum likelihood estimation (logistic regression, GLMs, most of classical statistics); likelihood-ratio tests; the "likelihood" term in Bayes' theorem, posterior ∝ likelihood × prior, which every NumPyro model computes; model comparison (AIC, BIC).

How is it used?

Write down the probability of one observation given θ, multiply over observations (in practice, add logs), and treat the result as a function of θ. Then maximize it (MLE), combine it with a prior (MAP, Bayes), or compare it between models.

Fix θ = 0.5, vary the data k Fix the data k = 7, vary θ 0 5 10 k = number of completions bars add up to 1 (a distribution over data) peak at θ = 0.7 (L = 0.267) same number 0.117 0 0.5 1 θ = completion rate area under the curve = 1/11, not 1 (not a distribution over θ)
One formula, $p(k\mid\theta) = \binom{10}{k}\theta^k(1-\theta)^{10-k}$, read two ways. Left: θ fixed, a probability distribution over possible data. Right: data fixed, the likelihood, a curve over θ whose area is not 1. The highlighted bar and the orange dot are the same number.

The big picture is the whole table $p(k\mid\theta)$ for $n = 10$: each row is a possible result $k$, each column a possible θ, darker = more probable. Drag the purple handle. The vertical slice through it (fixed θ) gives the left plot: a distribution over $k$ that always adds up to 1. The horizontal slice (fixed $k$) gives the right plot: the likelihood over θ, whose area is always $1/11$. Same table, two different questions.

Set the data with the two sliders. The curve shows $L(\theta)$ divided by its highest value, so the peak is always at height 1 (only ratios matter). Drag the purple point along the curve to read the likelihood ratio of any θ against the best one. Notice: the peak is always at $k/n$, and with $n = 100$ the curve is much narrower than with $n = 10$, even at the same proportion. More data rule out more values of θ.

"$L(0.7) = 0.267$ means there is a 26.7% chance that θ = 0.7."

It is the probability of the data if θ were 0.7. A statement about the chance of θ itself needs a prior and Bayes' theorem (the posterior, MAP below and Chapter 6.1).

"The likelihood curve integrates to 1, like a density."

Here its area is $1/11$. In general the area can be anything; only ratios of likelihoods are meaningful.

"$p(D\mid\theta)$ and $p(\theta\mid D)$ are the same thing."

They are different conditional probabilities (Chapter 4.3). The first is the likelihood; the second is the posterior, and getting from one to the other needs the prior.

"The likelihood is the probability of the parameter given the data."

"The likelihood is the probability (or density) of the observed data given the parameter, read as a function of the parameter."

Model answer: "For fixed θ, $p(D\mid\theta)$ is a distribution over possible datasets and sums to 1 over the data. For fixed data, the same expression as a function of θ is the likelihood; it is not a distribution over θ and need not integrate to 1. Only likelihood ratios are meaningful. To get a distribution over θ I multiply by a prior and normalize: that is the posterior."

$L(\theta) = p(D\mid\theta)$ with $D$ fixed; iid: $L(\theta) = \prod_i p(x_i\mid\theta)$.

A score for how well θ explains the data. Only ratios matter; constants in θ can be dropped.

Trap: the likelihood is not a probability distribution over θ (area $1/11$ in the coin example).

Quick check: a coin gives 3 heads in 3 tosses. What is $L(\theta)$, and which θ maximizes it?

$L(\theta) = \theta^3$. It increases all the way to $\theta = 1$, so the likelihood is largest at θ = 1. Three heads are "best explained" by a coin that always lands heads. With so little data that is clearly over-confident, which is one motivation for MAP later in this chapter.

The log-likelihood: sums instead of products core

With many observations, the likelihood is a product of many probabilities, each smaller than 1 (or many densities). The product gets astonishingly small: with a thousand data points it can be smaller than any number a computer can store, and the computer silently rounds it to 0. Every θ then looks equally (im)possible.

The cure is the logarithm. The log turns a product into a sum of manageable numbers. And because the log only ever goes up when its input goes up, the θ with the highest likelihood is also the θ with the highest log-likelihood. We lose nothing and gain stability.

Three ways to say it:

  • Picture: the log squashes a tall, thin spike into a gentle hill with the same top.
  • Numbers: $0.05^{1000} \approx 10^{-1301}$ rounds to 0 on a computer, but $1000 \times \log 0.05 \approx -2995.7$ is a perfectly ordinary number.
  • Slogan: take logs, add instead of multiply; the peak does not move.

The onboarding data again (7 completions of 10):

  1. $\ell(\theta) = \log L(\theta) = \log 120 + 7\log\theta + 3\log(1-\theta)$.
  2. At $\theta = 0.7$: $\log 120 \approx 4.787$, $\;7\log 0.7 \approx -2.497$, $\;3\log 0.3 \approx -3.612$.
  3. Sum: $4.787 - 2.497 - 3.612 = -1.321$, and $e^{-1.321} \approx 0.267$ ✓ (the likelihood we found before).

Underflow. 1000 observations, each with density about 0.05:

  1. Product: $0.05^{1000} = 10^{1000\log_{10}0.05} \approx 10^{-1301}$.
  2. The smallest positive number a standard 64-bit float can hold is about $5 \times 10^{-324}$. So the computer stores 0.
  3. Sum of logs: $1000 \times \log 0.05 \approx 1000 \times (-2.9957) = -2995.7$. No problem at all.

The log-likelihood is

$$\ell(\theta) = \log L(\theta) = \sum_{i=1}^n \log p(x_i\mid\theta) \quad\text{(iid data)}.$$
  • Because $\log$ is strictly increasing, $\arg\max_\theta \ell(\theta) = \arg\max_\theta L(\theta)$: the same best θ.
  • The negative log-likelihood $NLL(\theta) = -\ell(\theta)$ turns "maximize" into "minimize", the usual direction for optimizers and loss functions.
  • The derivative $\ell'(\theta)$ is called the score; at an interior maximum it is 0.
  • Differences of log-likelihoods are logs of likelihood ratios: $\ell(\theta_1) - \ell(\theta_2) = \log\big(L(\theta_1)/L(\theta_2)\big)$.
Why do we need it?

Products of thousands of probabilities underflow to zero, and products are awkward to differentiate. Sums of logs are numerically safe, easy to differentiate term by term, and easy to split into minibatches.

Where is it used?

Every fitting library works in logs: scipy.stats.*.logpdf, PyTorch's cross-entropy, NumPyro's log_prob on every distribution, the log-sum-exp trick, the log-likelihood term inside the ELBO (Chapter 6.12), AIC/BIC.

How is it used?

Never multiply probabilities in code: add logpdf or logpmf values instead. Optimize the (negative) log-likelihood. Compare models on the same data by differences in log-likelihood, never by raw likelihood values.

The data are $n$ tosses with 70% successes. Top: the likelihood computed the naive way, as a direct product of $n$ probabilities, scaled so its highest point is 1. Bottom: the log-likelihood (relative to its peak), computed as a sum of logs. Slide $n$ up. Just above $n = 1170$ the top curve turns jagged (the numbers are so tiny that the computer has lost most of their digits), and from about $n = 1230$ the direct product underflows to 0 for every θ: the curve collapses. The bottom curve stays perfectly smooth and still peaks at 0.7.

"Taking the log changes which θ is best."

The log is strictly increasing, so it never changes the order of values: the maximizer is the same. (It does change the shape, which is why gradients and curvature look different on the log scale.)

"A log-likelihood must be negative; a positive one is a bug."

For discrete data, each $\log p \le 0$, so $\ell \le 0$. For continuous data, densities can exceed 1 (a narrow Normal), so $\ell$ can be positive. Its size depends on $n$ and on the units of the data.

"Model A has log-likelihood −500 and model B has −800 on another dataset, so A is better."

Log-likelihoods are only comparable on the same data. More data points always make the sum more negative.

Both of your projects live in log space. NumPyro computes log_prob for every sample statement and adds them up; the expected log-likelihood is the "fit" term of the ELBO that your SVI loop maximizes (Chapter 6.12). If you ever compute a likelihood yourself, for example a posterior predictive check on thousands of forecast days, add logs (and use log-sum-exp when you must add probabilities); a direct product would underflow exactly as in the widget.

$\ell(\theta) = \sum_i \log p(x_i\mid\theta)$; same maximizer as $L$; $NLL = -\ell$.

Products of many probabilities underflow (below ~$10^{-308}$, and to 0 below ~$5\times10^{-324}$); sums of logs do not.

Trap: compare log-likelihoods only on the same data.

Quick check: for 7 of 10, which is bigger, $\ell(0.6)$ or $\ell(0.8)$? Use $\ell(\theta) = 7\log\theta + 3\log(1-\theta)$ (dropping $\log 120$).

$\ell(0.6) = 7(-0.511) + 3(-0.916) = -3.576 - 2.749 = -6.325$. $\ell(0.8) = 7(-0.223) + 3(-1.609) = -1.562 - 4.828 = -6.390$. So $\ell(0.6)$ is slightly bigger: 0.6 explains 7/10 a little better than 0.8 does (ratio $e^{0.065} \approx 1.07$).

Maximum likelihood: pick the θ that explains the data best core

Once you can score every θ by how well it explains the data, the obvious recipe is: pick the θ with the best score. Walk along the likelihood curve and stop at the top of the hill. That top is the maximum likelihood estimate, the MLE.

To find the top you can use the oldest trick in calculus: at the top of a smooth hill the ground is flat, so the slope (the derivative) is zero. Write the log-likelihood, take its derivative, set it to zero, solve.

Three ways to say it:

  • Picture: the MLE is the highest point of the likelihood hill.
  • Numbers: 7 completions out of 10 gives $\hat\theta = 7/10 = 0.7$.
  • Slogan: the MLE makes the data you saw as unsurprising as possible.

The coin (or completion rate), derived. $k$ successes in $n$ independent tries.

  1. Likelihood: $L(\theta) = \binom nk\theta^k(1-\theta)^{n-k}$.
  2. Log-likelihood: $\ell(\theta) = \log\binom nk + k\log\theta + (n-k)\log(1-\theta)$.
  3. Derivative (the score): $\ell'(\theta) = \dfrac{k}{\theta} - \dfrac{n-k}{1-\theta}$. (The constant $\log\binom nk$ disappears.)
  4. Set it to zero: $\dfrac{k}{\theta} = \dfrac{n-k}{1-\theta} \;\Rightarrow\; k(1-\theta) = (n-k)\theta \;\Rightarrow\; k - k\theta = n\theta - k\theta \;\Rightarrow\; k = n\theta$.
  5. Solve: $\hat\theta_{MLE} = k/n$. For 7 of 10: $0.7$.
  6. Check it is a maximum: $\ell''(\theta) = -\dfrac{k}{\theta^2} - \dfrac{n-k}{(1-\theta)^2} \lt 0$ everywhere, so the hill has one top. At $\theta = 0.7$: $\ell'' = -7/0.49 - 3/0.09 \approx -47.6$.

Reading the slope: at $\theta = 0.3$, $\ell'(0.3) = 7/0.3 - 3/0.7 \approx 23.33 - 4.29 = 19.05 \gt 0$: uphill is to the right. At $\theta = 0.9$, $\ell'(0.9) = 7.78 - 30 = -22.2 \lt 0$: uphill is to the left.

The maximum likelihood estimator is

$$\hat\theta_{MLE} = \arg\max_\theta\, p(D\mid\theta) = \arg\max_\theta\, \ell(\theta) = \arg\min_\theta\, \big(-\ell(\theta)\big).$$

($\arg\max$ means "the value of θ at which the maximum is reached", not the maximum value itself.)

The recipe: (1) write the probability of the data given θ; (2) take logs; (3) differentiate and set to zero; (4) solve; (5) check it is a maximum, and check the edges of the allowed range. When step 4 has no formula (logistic regression, most real models), minimize $-\ell(\theta)$ with a numerical optimizer such as gradient descent, Newton's method or L-BFGS (Optimization 3.3).

  • If $k = 0$ or $k = n$, the maximum is at the edge ($\hat\theta = 0$ or $1$), where the derivative is not zero. Always check edges.
  • The MLE is a recipe (an estimator in the sense of Chapter 5.1), so it has a bias, a variance and an MSE like any other.
Why do we need it?

It gives one general way to build an estimator for any probability model, instead of inventing a recipe for each problem. For large samples it is also hard to beat (consistent and efficient, see the curvature section).

Where is it used?

Linear regression (least squares is the Gaussian MLE), logistic and Poisson regression, fitting distributions (scipy.stats.norm.fit), Box-Cox's λ, training neural networks with cross-entropy, the "observed conversion rate" $k/n$ in A/B tests.

How is it used?

Write the negative log-likelihood as a function of the parameters and hand it to an optimizer (scipy.optimize.minimize, gradient descent, or IRLS, iteratively reweighted least squares, for GLMs). Check convergence, check edge cases, and report the estimate with an uncertainty from the curvature.

1 · model p(D | θ) 2 · take logs ℓ(θ) = Σ log p 3 · slope = 0 ℓ′(θ) = 0 4 · solve θ̂ = k/n 5 · check ℓ″ < 0, edges no formula in step 4? → minimize −ℓ(θ) numerically (gradient descent, Newton, L-BFGS)
The maximum likelihood recipe. Steps 3–4 work by hand for simple models; real models go straight to a numerical optimizer on the negative log-likelihood.

The blue curve is the log-likelihood (relative to its top). Drag the purple point anywhere: the orange line is the slope there, $\ell'(\theta)$. Positive slope means "uphill is to the right". Press Take a step uphill repeatedly: each step moves in the direction the slope points, and the steps get shorter as the slope flattens near the top. At the top the slope is exactly 0. Try $k = 0$: the top is at the edge θ = 0, where the slope is not 0.

"The MLE is the most probable value of θ."

The MLE makes the data most probable. "The most probable θ given the data" is a statement about the posterior and needs a prior; its peak is the MAP estimate (later in this chapter).

"Maximum likelihood estimates are always unbiased."

$k/n$ happens to be unbiased, but the MLE of a Normal variance is biased (next section), and many MLEs are biased at small $n$. Their bias usually shrinks like $1/n$.

"Derivative = 0 proves I found the best θ."

It finds flat points. Check that it is a maximum (second derivative negative), that it is the highest of several bumps, and that the edges of the parameter range are not higher (like $k = 0$).

"The MLE gives the most likely parameter."

"The MLE is the parameter value under which the observed data are most likely."

Model answer: "MLE maximizes $p(D\mid\theta)$ over θ. It treats θ as a fixed unknown and has no prior. If I want the most probable θ, I need $p(\theta\mid D) \propto p(D\mid\theta)p(\theta)$, whose maximizer is the MAP estimate; with a flat prior on θ the two coincide."

$\hat\theta_{MLE} = \arg\max_\theta \ell(\theta)$: set $\ell'(\theta) = 0$, solve, check $\ell'' \lt 0$ and the edges.

Coin: $\ell'(\theta) = k/\theta - (n-k)/(1-\theta) = 0 \Rightarrow \hat\theta = k/n$.

Trap: MLE maximizes the probability of the data, not of θ.

Quick check: 0 conversions out of 40 visitors. What is the MLE, and what is odd about it?

$\hat\theta = 0/40 = 0$, at the edge, where $\ell'(\theta) = -40/(1-\theta)$ is never 0. It claims the page can never convert, which nobody believes. Small or extreme data are where a prior (MAP or a full posterior) helps.

MLE for a Normal: the mean, and a variance that divides by n core

Now the data are numbers on a line, and the model is a bell curve with two knobs: its centre $\mu$ and its width $\sigma$. Which bell explains the points best?

  • Slide the bell left and right: it fits best when it sits in the middle of the points, at their mean.
  • Make the bell wider or narrower: too narrow and the points far from the centre get almost no probability; too wide and every point gets only a little. The best width is the typical distance of the points from the mean, measured by the average squared distance.

The catch: "average" here divides by $n$, not $n-1$. So the MLE of the variance is the slightly-too-small recipe from Chapter 5.1.

Three ways to say it:

  • Picture: on a map of all (μ, σ) pairs, the log-likelihood is a single hill whose top sits at (mean, root-mean-square distance).
  • Numbers: data 4, 6, 8, 10, 12 give $\hat\mu = 8$ and $\hat\sigma^2 = 40/5 = 8$, while $s^2 = 40/4 = 10$.
  • Slogan: the Normal MLE is the sample mean plus the divide-by-$n$ variance.

Five days of residual errors (or orders): 4, 6, 8, 10, 12.

  1. $\hat\mu = \bar x = (4+6+8+10+12)/5 = 40/5 = 8$.
  2. Squared distances from 8: $16, 4, 0, 4, 16$; sum $= 40$.
  3. $\hat\sigma^2_{MLE} = 40/5 = 8$, so $\hat\sigma = \sqrt8 \approx 2.83$. (The unbiased $s^2 = 40/4 = 10$.)
  4. The highest log-likelihood: $\ell(\hat\mu, \hat\sigma) = -\tfrac n2\log(2\pi\hat\sigma^2) - \tfrac n2 = -2.5\log(16\pi) - 2.5 \approx -9.79 - 2.5 = -12.29$.

Data $x_1, \dots, x_n$ iid $N(\mu, \sigma^2)$. Each point has density $\frac{1}{\sqrt{2\pi}\sigma}\exp\!\big(-\frac{(x_i-\mu)^2}{2\sigma^2}\big)$, so

$$\ell(\mu, \sigma) = -n\log\sigma - \frac n2\log(2\pi) - \frac{1}{2\sigma^2}\sum_{i=1}^n (x_i - \mu)^2.$$
  1. Derivative in $\mu$: $\dfrac{\partial\ell}{\partial\mu} = \dfrac{1}{\sigma^2}\sum(x_i - \mu) = 0 \Rightarrow \sum x_i = n\mu \Rightarrow \hat\mu = \bar x$. (Whatever σ is: maximizing over μ is the same as minimizing the sum of squared errors.)
  2. Derivative in $\sigma$: $\dfrac{\partial\ell}{\partial\sigma} = -\dfrac n\sigma + \dfrac{1}{\sigma^3}\sum(x_i-\mu)^2 = 0 \Rightarrow \sigma^2 = \dfrac1n\sum(x_i - \mu)^2$.
  3. Plug in $\hat\mu$: $\hat\sigma^2_{MLE} = \dfrac1n\sum(x_i - \bar x)^2$.
  • $E[\hat\sigma^2_{MLE}] = \frac{n-1}{n}\sigma^2$: biased low by $\sigma^2/n$, but consistent (the bias vanishes as $n$ grows).
  • $\hat\mu = \bar x$ is unbiased.
  • Library check: scipy.stats.norm.fit(x) returns $(\bar x, \hat\sigma_{MLE})$ (divide by $n$); np.std(x) also divides by $n$; np.std(x, ddof=1) gives $s$.
Why do we need it?

The Normal model is the default for continuous noise. Knowing its MLE explains why least squares is the standard fitting rule and where the $n$ versus $n-1$ difference comes from.

Where is it used?

Ordinary least squares regression (the Gaussian MLE of the coefficients, Chapter 5.13), scipy.stats.norm.fit, Kalman filters, Gaussian noise models in forecasting, and the residual standard deviation reported by many fitting tools.

How is it used?

Take the mean as the centre and the root-mean-square residual as the noise level. If you need an unbiased variance (for example to build a t-interval), multiply by $n/(n-1)$. In a model with a mean that depends on inputs, the same logic gives least squares for the mean parameters.

Top strip: five data points (drag them). Bottom: every possible pair (μ, σ) coloured by its log-likelihood (darker = better), with contour lines 1, 3 and 8 below the top. The green dot is the MLE $(\bar x, \hat\sigma)$; the orange ring is $(\bar x, s)$, which sits slightly above the top because $s$ divides by $n-1$. Drag the purple handle to any (μ, σ) and read how much worse it is. Then drag one data point far away: the hill moves right and stretches upward (a wider σ is needed).

"The MLE of $\sigma^2$ is the sample variance $s^2$."

The MLE divides by $n$: $\hat\sigma^2 = \frac1n\sum(x_i - \bar x)^2 = \frac{n-1}{n}s^2$. It is biased low by $\sigma^2/n$; $s^2$ is the unbiased correction.

"A biased MLE is a wrong MLE."

The MLE is what maximizes the likelihood; bias is a separate property of that recipe. Here the bias is small and disappears as $n$ grows.

"dist.Normal(mu, sigma**2) in NumPyro."

We write $N(\mu, \sigma^2)$ with the variance, but NumPyro's Normal(loc, scale) and SciPy's norm(loc, scale) take the standard deviation σ.

In your forecasting model with a Normal likelihood, $y_t \sim N(\mu_t, \sigma^2)$ where $\mu_t = g(t) + s(t) + h(t) + X_t\beta$. Maximizing over the components of $\mu_t$ alone is exactly least squares on the residuals $y_t - \mu_t$. And if you maximized over σ alone with no prior, the answer would be the root-mean-square residual $\sqrt{\frac1n\sum(y_t-\mu_t)^2}$. Your Bayesian fit instead puts a prior on σ and reports a posterior, but this is the number the data alone point to.

Normal MLE: $\hat\mu = \bar x$, $\hat\sigma^2 = \frac1n\sum(x_i-\bar x)^2$ (biased low by $\sigma^2/n$; $s^2$ divides by $n-1$).

Maximizing the Normal likelihood over the mean = minimizing squared error.

Trap: NumPyro/SciPy Normal takes σ (scale), not σ².

Quick check: data 1, 3. What are $\hat\mu$, $\hat\sigma^2_{MLE}$ and $s^2$?

$\hat\mu = 2$. Squared distances 1 and 1, sum 2. $\hat\sigma^2_{MLE} = 2/2 = 1$, $s^2 = 2/1 = 2$. With $n = 2$ the MLE is only half of the unbiased value.

MLE for a Poisson: the rate is the average count

Orders arrive at random; you count how many came in each hour. The Poisson model has one knob, the rate λ (the expected count per hour). Which λ explains your counts best?

Each hour "votes" for λ: an hour with 6 orders prefers λ near 6, an hour with 2 prefers λ near 2. Adding the votes (the log-likelihoods) gives a hill whose top is the average count.

Three ways to say it:

  • Picture: one small hill per observation, each peaked at its own count; their sum peaks at the mean.
  • Numbers: counts 2, 4, 3, 6, 5 give $\hat\lambda = 20/5 = 4$.
  • Slogan: for a Poisson, the best rate is the average count.

Orders in five hours: 2, 4, 3, 6, 5. The Poisson pmf is $p(y\mid\lambda) = \lambda^y e^{-\lambda}/y!$.

  1. $\ell(\lambda) = \sum_i\big(y_i\log\lambda - \lambda - \log y_i!\big) = (\textstyle\sum y_i)\log\lambda - n\lambda - \sum\log y_i!$.
  2. Here $\sum y_i = 20$ and $n = 5$: $\ell(\lambda) = 20\log\lambda - 5\lambda - \text{const}$.
  3. Derivative: $\ell'(\lambda) = 20/\lambda - 5 = 0 \Rightarrow \hat\lambda = 20/5 = 4$.
  4. Second derivative $-20/\lambda^2 \lt 0$: a maximum.
  5. Values: $\ell(3) \approx -10.06$, $\ell(4) \approx -9.30$, $\ell(5) \approx -9.84$. The top is at 4.

For iid counts $y_1, \dots, y_n \sim Poisson(\lambda)$:

$$\ell(\lambda) = \Big(\sum_i y_i\Big)\log\lambda - n\lambda - \sum_i \log y_i!, \qquad \hat\lambda_{MLE} = \bar y.$$
  • $\hat\lambda$ is unbiased, with variance $\lambda/n$.
  • If every count is 0, $\hat\lambda = 0$ (at the edge of the allowed range).
  • Assumptions: independent counts with the same rate. The Poisson model also says variance = mean. Real counts often have variance much bigger than the mean (overdispersion, Chapter 4.8). The estimate $\bar y$ of the mean is still sensible then, but the Poisson model's statements about uncertainty become too confident.
Why do we need it?

Counts (orders per hour, clicks per session, support tickets per day) are everywhere, and the Poisson is the simplest model for them. Its MLE gives the rate and is the starting point for Poisson regression.

Where is it used?

Poisson regression and its log link (Chapter 5.14), count metrics in A/B tests, arrival-rate estimates in queueing and staffing, and as the baseline that Negative Binomial models extend when counts are overdispersed.

How is it used?

Estimate the rate by the average count (per unit of exposure if the intervals differ in length). Then compare the sample variance with the mean: if the variance is much larger, switch to a Negative Binomial model before trusting any uncertainty statements.

Top: five hourly counts (drag them; they snap to whole numbers). Bottom: each thin grey curve is one count's own log-likelihood over λ, peaked at that count. The thick orange curve is their sum, the full log-likelihood, and it peaks at the average (green). Press One busy hour: a single 14 drags the peak to the right, and the variance-to-mean ratio jumps above 1, a hint of overdispersion.

"If the Poisson MLE is the mean, the Poisson model is fine for any counts."

The point estimate of the mean is fine, but a Poisson model also claims variance = mean. With overdispersed counts its standard errors and predictive intervals are too narrow; a Negative Binomial likelihood fixes that (Chapter 5.14, Chapter 7.13).

"Hours of different lengths can just be averaged."

If hour $i$ has exposure $t_i$ (say, minutes open), the model is $y_i \sim Poisson(\lambda t_i)$ and the MLE is $\hat\lambda = \sum y_i / \sum t_i$: total count over total exposure.

Your A/B framework uses Poisson likelihoods for count metrics: its rate parameter plays the role of λ here, and the data's "vote" is driven by the total count and the number of units. Your forecasting model offers a Negative Binomial likelihood, the usual choice when daily demand counts are overdispersed: the mean is still estimated sensibly, but a Poisson would claim far too little day-to-day variation.

Poisson: $\ell(\lambda) = (\sum y_i)\log\lambda - n\lambda + \text{const}$, $\hat\lambda = \bar y$ (with exposure: $\sum y_i/\sum t_i$).

Check $s^2$ against $\bar y$: a Poisson needs variance ≈ mean.

Trap: overdispersion leaves $\hat\lambda$ fine but makes Poisson uncertainty too small.

Quick check: 3 days with 0, 0, 9 support tickets. What is $\hat\lambda$, and should you trust a Poisson model here?

$\hat\lambda = 9/3 = 3$. The sample variance is $\big((0-3)^2 + (0-3)^2 + (9-3)^2\big)/2 = (9+9+36)/2 = 27$, nine times the mean. That is strong overdispersion (or a mix of quiet and busy days): the Poisson would understate the uncertainty.

Curvature: how sure is the MLE? (Fisher information and large-sample behaviour)

Two hills can have their tops at the same place and still be very different. One is a sharp mountain peak: step a little to the side and you drop fast, so the data clearly rule out nearby values. The other is a broad, flat plateau: many values explain the data almost equally well, so the data barely tell them apart.

The curvature at the top (how fast the slope changes) measures this sharpness. Sharp peak → precise estimate, small standard error. Flat top → vague estimate, big standard error. More data make the peak sharper.

Three ways to say it:

  • Picture: the width of the log-likelihood hill near its top is the uncertainty of the MLE.
  • Numbers: 7 of 10 gives a standard error of about 0.145; 70 of 100 gives about 0.046, $\sqrt{10}$ times smaller.
  • Slogan: precision comes from the curvature of the peak, not from its height.

Coin data, two sample sizes, same proportion 0.7. From the MLE section, $\ell''(\theta) = -\frac{k}{\theta^2} - \frac{n-k}{(1-\theta)^2}$. At $\hat\theta = k/n$ this simplifies to $-\frac{n}{\hat\theta(1-\hat\theta)}$.

  1. $n = 10$: curvature $J = -\ell''(0.7) = 10/(0.7 \times 0.3) = 10/0.21 \approx 47.6$.
  2. Standard error $\approx 1/\sqrt J = 1/\sqrt{47.6} \approx 0.145$.
  3. $n = 100$: $J = 100/0.21 \approx 476$, so $SE \approx 1/\sqrt{476} \approx 0.046$.
  4. Ten times the data → curvature ten times bigger → standard error $\sqrt{10} \approx 3.16$ times smaller.
  5. Note that $1/\sqrt{J} = \sqrt{\hat\theta(1-\hat\theta)/n}$: exactly the familiar standard error of a proportion (Chapter 5.5).
  • Observed information: $J(\hat\theta) = -\ell''(\hat\theta)$, the curvature of the log-likelihood at its peak.
  • Fisher information of one observation: $I(\theta) = -E\!\left[\dfrac{\partial^2}{\partial\theta^2}\log p(X\mid\theta)\right]$. For $n$ iid observations the total is $n\,I(\theta)$. Coin: $I(\theta) = \dfrac{1}{\theta(1-\theta)}$; Normal mean: $I(\mu) = 1/\sigma^2$.
  • Standard error from curvature: $SE(\hat\theta) \approx 1/\sqrt{J(\hat\theta)} \approx 1/\sqrt{n\,I(\hat\theta)}$.

Large-sample properties of the MLE (when the model is correct, the true θ is inside the allowed range, the model is smooth and identifiable, meaning different values of θ give different data distributions, and the number of parameters stays fixed as $n$ grows):

  1. Consistent: $\hat\theta_n \to \theta$ (Chapter 5.1).
  2. Approximately Normal: $\hat\theta \approx N\!\big(\theta,\ 1/(n I(\theta))\big)$ for large $n$.
  3. Asymptotically efficient: its variance approaches the Cramér–Rao bound $1/(nI(\theta))$, the smallest possible for unbiased recipes (Chapter 5.1).
  4. Invariant: the MLE of $g(\theta)$ is $g(\hat\theta)$ (more on this at the end of the chapter).

These are large-$n$ statements. At small $n$ the MLE can be biased (the Normal variance) and the curvature-based SE can be poor, especially near the edges of the parameter range.

Why do we need it?

An estimate without an uncertainty is half an answer. The curvature gives standard errors for almost any MLE for free, even when no formula exists, because an optimizer can compute the second derivative (the Hessian) at the peak.

Where is it used?

Standard errors printed by statsmodels for regressions and GLMs (from the curvature, or information, matrix), Wald tests and intervals (estimate ± z × SE), the Laplace approximation of a posterior (a Gaussian whose width comes from the curvature at the mode, Chapter 6.9), experimental design (more information = more precise).

How is it used?

Fit by minimizing $-\ell$, then take the Hessian of $-\ell$ at the optimum. With several parameters, the inverse Hessian is the approximate covariance matrix; the square roots of its diagonal are the standard errors.

Blue: the log-likelihood for an observed rate $k/n$, relative to its top. Orange dashed: the parabola with the same curvature at the top, $-\tfrac12 J(\theta - \hat\theta)^2$. The purple bar is $\hat\theta \pm 1.96\,SE$ with $SE = 1/\sqrt J$. Switch $n$ from 10 to 400: the hill becomes a needle and the bar shrinks like $1/\sqrt n$. With large $n$ the blue curve and the parabola agree; with $n = 10$ and a rate of 0.1 they do not, and the bar spills below 0, a sign that the curvature approximation is breaking down near the edge.

"A high peak means a precise estimate."

The height of $L(\hat\theta)$ depends on $n$ and the units; it says nothing about precision. Precision comes from the curvature (width) of the peak.

"$1/\sqrt J$ is the exact standard error."

It is a large-sample approximation. It breaks down for small $n$ and near edges: with $k = 0$ the coin formula $\sqrt{\hat\theta(1-\hat\theta)/n}$ gives $SE = 0$, which is absurd. Chapter 5.8 shows better intervals for proportions (Wilson).

"The MLE's nice properties hold for any model."

They need a correct (or nearly correct) model and regularity conditions. If the model is wrong, the MLE converges to the parameter value that makes the model closest to the truth, which may not mean what you hoped.

The curvature idea reappears in Bayesian computation. A Gaussian placed at the posterior mode with its width taken from the curvature there is the Laplace approximation (Chapter 6.9). Your SVI guides are also Gaussians, but they choose their means and covariances by maximizing the ELBO rather than by reading curvature at one point; the full-rank versus low-rank choice (Chapter 6.13) is about how much of the parameter-to-parameter curvature structure (the correlations) the guide can represent.

$J(\hat\theta) = -\ell''(\hat\theta)$; $SE \approx 1/\sqrt J \approx 1/\sqrt{nI(\theta)}$. Coin: $SE = \sqrt{\hat\theta(1-\hat\theta)/n}$.

Large $n$, correct model: MLE is consistent, ≈ Normal, efficient (reaches Cramér–Rao), invariant.

Trap: precision = curvature, not height; the approximation fails at small $n$ and near edges.

Quick check: for Normal data with known σ, $\ell(\mu) = -\frac{1}{2\sigma^2}\sum(x_i-\mu)^2 + \text{const}$. What is $J$, and what SE does it give?

$\ell'(\mu) = \frac{1}{\sigma^2}\sum(x_i - \mu)$ and $\ell''(\mu) = -n/\sigma^2$. So $J = n/\sigma^2$ and $SE = 1/\sqrt J = \sigma/\sqrt n$: the standard error of the mean, exactly.

Every familiar loss is a negative log-likelihood core

In machine learning you choose a loss: squared error, absolute error, cross-entropy. In statistics you choose a noise model: Normal, Laplace, Bernoulli. These are the same choice. Write the negative log-likelihood of a noise model and drop the constants: out comes a familiar loss.

This matters because a loss quietly states an assumption about your data. Squared error assumes Normal-looking noise with a constant spread, so it punishes a big residual enormously. Absolute error assumes Laplace noise, with heavier tails, so it punishes big residuals less. A Student-t noise model goes further: its loss bends over for large residuals.

Three ways to say it:

  • Picture: a parabola (Normal), a V (Laplace), and a curve that flattens out (Student-t).
  • Numbers: for the data 1, 2, 3, 4, 20, squared error picks 6 (the mean), absolute error picks 3 (the median).
  • Slogan: choosing a loss = choosing a noise model.

Gaussian → squared error. One observation $y$ with prediction $\mu$ and noise $N(0,\sigma^2)$:

  1. $-\log p(y\mid\mu) = -\log\Big(\frac{1}{\sqrt{2\pi}\sigma}e^{-(y-\mu)^2/(2\sigma^2)}\Big) = \frac{(y-\mu)^2}{2\sigma^2} + \log\sigma + \tfrac12\log 2\pi$.
  2. With σ fixed, the last two terms are constants. Summing over data, minimizing the NLL = minimizing $\sum(y_i - \mu_i)^2$: least squares.

Bernoulli → cross-entropy (log loss). Outcome $y \in \{0, 1\}$, predicted probability $q$:

  1. $p(y\mid q) = q^y(1-q)^{1-y}$, so $-\log p = -\big[y\log q + (1-y)\log(1-q)\big]$: binary cross-entropy.
  2. If $q = 0.8$ and the user converted ($y = 1$): loss $= -\log 0.8 \approx 0.223$. If the user did not ($y = 0$): $-\log 0.2 \approx 1.609$.

Laplace → absolute error. $p(y\mid\mu) = \frac{1}{2b}e^{-|y-\mu|/b}$, so $-\log p = |y-\mu|/b + \log 2b$: minimizing gives $\sum|y_i - \mu|$, minimized by the median. For 1, 2, 3, 4, 20: the median is 3 while the mean is 6.

For iid data, $\hat\theta_{MLE} = \arg\min_\theta \sum_i -\log p(y_i\mid\theta)$. Dropping constants gives:

Noise model (likelihood)Loss per point (constants dropped)Common nameBest constant fit
Normal $N(\mu, \sigma^2)$, σ fixed$(y-\mu)^2/(2\sigma^2)$squared error, MSEmean
Laplace$(\mu, b)$, b fixed$|y-\mu|/b$absolute error, MAEmedian
Bernoulli$(q)$$-[y\log q + (1-y)\log(1-q)]$binary cross-entropy, log lossproportion of 1s
Categorical$(q_1..q_K)$$-\log q_{y}$cross-entropy (softmax loss)class frequencies
Poisson$(\lambda)$$\lambda - y\log\lambda$Poisson loss / deviancemean
Student-t$(\nu, \mu, s)$$\frac{\nu+1}{2}\log\!\big(1 + \frac{(y-\mu)^2}{\nu s^2}\big)$a robust lossno formula (outliers count less)

When the location depends on inputs, $\mu_i = \mathbf{x}_i^\top\boldsymbol\beta$, the same table gives linear regression (Normal), least absolute deviations regression (Laplace), logistic regression (Bernoulli) and Poisson regression (Chapters 5.13–5.14). If σ is also learned, the $\log\sigma$ term stays in the loss and matters.

Why do we need it?

It tells you what your loss assumes, so you can choose it on purpose. It also explains how to add uncertainty (learn σ), how to handle outliers (Laplace, Student-t) and why classification uses cross-entropy instead of squared error.

Where is it used?

Every training loop: MSELoss, L1Loss, BCEWithLogitsLoss and CrossEntropyLoss in PyTorch are negative log-likelihoods (up to constants and scaling); scikit-learn's log_loss; robust regression; and every NumPyro model, where the likelihood is the loss.

How is it used?

Start from the data: what values can they take, how does the spread behave, are there outliers? Pick the matching likelihood, write its negative log, and minimize. Read your existing loss backwards to see the noise model you have been assuming.

−8 0 8 residual r = y − μ (in units of the scale) loss = −log p(r) Normal: r²/2 Laplace: |r| Student-t₃: 2·log(1 + r²/3)
The negative log-density of three noise models. The Normal's parabola makes a residual of 8 cost 32; the Laplace's V makes it cost 8; the Student-t's curve flattens and charges only about 6.2. Flatter growth means outliers pull the fit less.

Top: seven days of orders (drag them). Bottom: the total negative log-likelihood of each noise model as a function of the centre μ (each shifted so its minimum is 0); the dots mark the best μ. Blue (Normal) is a smooth parabola with its minimum at the mean; teal (Laplace) is made of straight pieces with its minimum at the median; pink (Student-t) has its minimum near the bulk of the data. Press One wild day: the Normal answer jumps, the Laplace answer does not move, the Student-t answer moves only a little.

Ten users: 7 converted, 3 did not (change the count with the slider). Drag the predicted conversion probability $q$. Each converter costs $-\log q$ and each non-converter costs $-\log(1-q)$; the curve is the average cost, the log loss. Its lowest point is exactly the proportion $k/n$, the maximum likelihood estimate. Push $q$ toward 0 or 1 and watch the loss explode: confident wrong predictions are punished very hard.

"Squared error is just a convenient default."

It is the Gaussian negative log-likelihood. Using it assumes roughly Normal noise with constant variance; with heavy-tailed noise or outliers it lets a few points dominate the fit.

"Cross-entropy and maximum likelihood are different things."

Average cross-entropy on the training data is the negative log-likelihood divided by $n$. Minimizing one is maximizing the other.

"Student-t removes outliers."

It removes nothing. It assigns more probability to extreme residuals, so they are less surprising and exert less influence on the fit than under a Normal likelihood.

Your forecasting model offers Normal, Student-t and Negative Binomial likelihoods: by this section, three different losses. The Normal makes the fit chase every big residual (squared error); the Student-t lets occasional spikes (an unmodelled promotion, a data glitch) pull the trend and seasonality much less; the Negative Binomial is the right "loss" for counts whose spread grows with their level. In the A/B framework, the Binomial likelihood for conversions is, up to constants, a cross-entropy loss on the conversion probability.

"Student-t likelihood is robust because it ignores outliers."

"Student-t assigns more probability to extreme residuals, so they exert less influence than under a Normal likelihood."

Model answer: "Minimizing a loss is maximizing a likelihood: squared error is Gaussian, absolute error is Laplace, log loss is Bernoulli. The Student-t negative log-likelihood grows only logarithmically in the residual, so its gradient for a huge residual is small: outliers still count, but much less."

MLE = $\arg\min \sum -\log p(y_i\mid\theta)$. Normal ↔ squared error (mean); Laplace ↔ absolute error (median); Bernoulli ↔ cross-entropy; Poisson ↔ $\lambda - y\log\lambda$.

Choosing a loss = choosing a noise model.

Trap: "Student-t removes outliers" is wrong; it gives them less influence.

Quick check: why does a classifier trained with log loss get punished so hard for predicting 0.99 when the answer is 0?

The loss for that example is $-\log(1 - 0.99) = -\log 0.01 \approx 4.6$, compared with about 0.01 for a correct confident prediction. Under the Bernoulli model, saying "99%" and being wrong makes the observed data very improbable, and the negative log of a tiny probability is large.

MAP: maximum a posteriori, the peak of prior × likelihood core

A brand-new variant converts 3 of its first 3 users. The MLE says the conversion rate is 100%. Nobody believes that: you know from hundreds of past tests that conversion rates are almost never near 100%.

MAP is the recipe that uses that knowledge. Before the data you have a prior: a distribution saying which values of θ are plausible. The data give the likelihood. Multiply them, value by value, and you get (up to a constant) the posterior: how plausible each θ is after seeing the data. MAP picks the highest point of the posterior. ("A posteriori" is Latin for "after the fact", here: after seeing the data.)

Three ways to say it:

  • Picture: multiply the prior curve by the likelihood curve; the peak of the product is the MAP.
  • Numbers: 3 of 3 with a Beta(2, 2) prior gives MAP = 4/5 = 0.8 instead of the MLE's 1.0.
  • Slogan: MAP = MLE plus a sensible pull from what you already knew.

3 conversions out of 3 users, prior Beta(2, 2) (a gentle belief that the rate is "somewhere in the middle"; its density is $\propto \theta(1-\theta)$).

  1. Likelihood: $p(D\mid\theta) = \theta^3$ (3 successes, 0 failures).
  2. Prior: $p(\theta) \propto \theta^{2-1}(1-\theta)^{2-1} = \theta(1-\theta)$.
  3. Posterior $\propto$ likelihood × prior $= \theta^3 \cdot \theta(1-\theta) = \theta^4(1-\theta)^1$. That is the shape of a Beta(5, 2) distribution.
  4. Find the peak: $\frac{d}{d\theta}\big[4\log\theta + \log(1-\theta)\big] = \frac4\theta - \frac{1}{1-\theta} = 0 \Rightarrow 4(1-\theta) = \theta \Rightarrow \theta = 4/5$.
  5. $\hat\theta_{MAP} = 0.8$. The MLE was $1.0$. The posterior mean of Beta(5, 2) is $5/7 \approx 0.714$.

7 of 10 with the same prior: posterior $\propto \theta^{8}(1-\theta)^{4}$ = Beta(9, 5), MAP $= 8/12 \approx 0.667$, a little below the MLE 0.7.

By Bayes' theorem, $p(\theta\mid D) = \dfrac{p(D\mid\theta)\,p(\theta)}{p(D)}$. The bottom, $p(D)$, does not depend on θ, so it does not move the peak:

$$\hat\theta_{MAP} = \arg\max_\theta\, p(\theta\mid D) = \arg\max_\theta\, \big[\log p(D\mid\theta) + \log p(\theta)\big] = \arg\min_\theta\, \big[\underbrace{-\log p(D\mid\theta)}_{\text{loss}} \;\underbrace{-\,\log p(\theta)}_{\text{penalty}}\big].$$

Beta prior with Binomial data. Prior Beta(α, β) and $k$ successes in $n$ tries give the posterior Beta$(\alpha + k,\ \beta + n - k)$ (the prior and posterior are in the same family, which is called a conjugate prior; the update is derived in Chapter 6.3). Its peak:

$$\hat\theta_{MAP} = \frac{k + \alpha - 1}{n + \alpha + \beta - 2}\quad(\text{when } \alpha + k \gt 1 \text{ and } \beta + n - k \gt 1).$$
  • Pseudo-counts: it is the MLE after adding $\alpha - 1$ imaginary successes and $\beta - 1$ imaginary failures to the data.
  • Flat prior Beta(1, 1): MAP = $k/n$ = MLE.
  • Lots of data: as $n$ grows the pseudo-counts become negligible and MAP → MLE.
  • The penalty view: $-\log p(\theta)$ acts like a regularizer. A Gaussian prior on regression weights gives ridge (L2); a Laplace prior gives lasso (L1). That story is Chapter 5.3.
Why do we need it?

With little data the MLE can be absurd (100% from 3 of 3, 0% from 0 of 40) or wildly noisy. A prior adds sensible information and stabilizes the estimate; MAP is the simplest way to use a prior when you only want one number.

Where is it used?

Regularized regression (ridge and lasso are MAP estimates), smoothed rates ("add-one" or Laplace smoothing in naive Bayes and n-gram models), weight decay in neural networks, and point-estimate fits of Bayesian models (NumPyro's AutoDelta guide gives a MAP-style point estimate).

How is it used?

Write the negative log-likelihood, add the negative log-prior, and minimize the sum with any optimizer. For a Beta prior on a rate, just use $(k + \alpha - 1)/(n + \alpha + \beta - 2)$. Check how much the answer moves when you change the prior.

× ∝ prior Beta(2, 2) likelihood θ³ (3 of 3) posterior Beta(5, 2) MAP 0.8 MLE 1.0 01 01 01 θ = conversion rate on each axis; each curve scaled to the same height
Multiply the prior by the likelihood at every θ: the product has the shape of the posterior. Its peak (the MAP, 0.8) sits between the prior's centre (0.5) and the MLE (1.0).

Teal: the prior. Blue: the likelihood of your data. Orange: their product, the posterior (each curve scaled to the same height). Start at "3 of 3": the posterior peak (MAP, purple) is pulled back from the MLE at 1.0. Press Strong prior: Beta(20, 20) holds the estimate near 0.5. Press Lots of data: the prior hardly matters and MAP ≈ MLE. Press Flat prior: with Beta(1, 1), MAP = MLE exactly.

"MAP is the Bayesian answer."

MAP is a single point. The Bayesian answer is the whole posterior distribution, with its spread, its skew and its tails. MAP throws all of that away.

"MAP is the posterior mean."

MAP is the posterior's peak; the mean is its balance point. For 3 of 3 with a Beta(2, 2) prior: MAP 0.8, mean 0.714. They agree only for symmetric posteriors (next section).

"The MAP formula works for any Beta prior."

It needs $\alpha + k \gt 1$ and $\beta + n - k \gt 1$. With a prior like Beta(0.5, 0.5) and $k = 0$, the posterior density shoots up at θ = 0 and the "peak" is at the edge.

Your A/B framework puts Beta priors on conversion rates. For a Beta-Binomial model the posterior is exactly Beta$(\alpha + k, \beta + n - k)$ (Chapter 6.3), whether your code uses that closed form or approximates it with SVI. The framework then answers questions such as $P(\theta_B \gt \theta_A \mid D)$ from the whole posterior; a MAP alone could not, because a probability like that needs the spread of both posteriors. In your forecasting model, minimizing (negative log-likelihood) + (negative log of the Laplace prior on the changepoint adjustments $\delta_j$) would be a MAP fit, and that sum is exactly an L1-penalized loss (Chapter 5.3). Your SVI fit goes further and approximates the full posterior.

$\hat\theta_{MAP} = \arg\max[\log p(D\mid\theta) + \log p(\theta)]$ = MLE with a penalty $-\log p(\theta)$.

Beta(α, β) prior, $k$ of $n$: MAP $= \frac{k+\alpha-1}{n+\alpha+\beta-2}$ (pseudo-counts); flat prior → MLE; large $n$ → MLE.

Trap: MAP is one point of the posterior, not the posterior.

Quick check: 0 conversions out of 40 users, prior Beta(2, 20). What is the MAP?

$(0 + 2 - 1)/(40 + 2 + 20 - 2) = 1/60 \approx 0.0167$. Instead of the MLE's impossible 0, MAP gives a small positive rate, as if we had seen 1 extra conversion and 19 extra non-conversions.

MAP versus the posterior mean (and median)

A posterior is a whole hill of plausibility. If you must give one number, you can report its peak (the MAP), its middle (the median: half the plausibility on each side), or its balance point (the mean: where the hill would balance on a finger).

For a symmetric hill these are the same point. For a lopsided hill they spread apart: the long tail drags the balance point toward it. Rare conversion rates give exactly such lopsided posteriors, squeezed against 0 with a tail to the right.

Three ways to say it:

  • Picture: peak, middle and balance point of a skewed hill, in that order from left to right.
  • Numbers: 1 conversion in 50 users (flat prior): MAP 0.020, median 0.033, mean 0.038.
  • Slogan: which single number is "best" depends on how you will be penalized for being wrong.

1 conversion in 50 users, flat prior Beta(1, 1). The posterior is Beta(1 + 1, 1 + 49) = Beta(2, 50).

  1. MAP (peak): $(a - 1)/(a + b - 2) = 1/50 = 0.020$. With a flat prior this is also the MLE.
  2. Posterior mean (balance point): $a/(a + b) = 2/52 \approx 0.0385$.
  3. Posterior median (middle): no simple formula; numerically $\approx 0.0327$.
  4. The mean is almost twice the MAP. Reporting "2%" or "3.8%" is a real difference for a business decision.
  5. Expected squared error of each guess $a$: $E[(\theta - a)^2\mid D] = Var(\theta\mid D) + (E[\theta\mid D] - a)^2$. Guessing the mean gives $\approx 0.00070$; guessing the MAP gives $\approx 0.00104$. Under squared error, the mean is the better report.

Three point summaries of a posterior $p(\theta\mid D)$, each the best guess under a different penalty for being wrong:

  • Posterior mean $E[\theta\mid D]$ minimizes the expected squared error $E[(\theta - a)^2 \mid D]$. (Because $E[(\theta-a)^2] = Var + (E\theta - a)^2$, which is smallest at $a = E\theta$.)
  • Posterior median minimizes the expected absolute error $E[\,|\theta - a|\,\mid D]$.
  • MAP (posterior mode) is the single most plausible value: the answer when only "exactly right" counts (it is the limit of a 0–1 loss with a shrinking tolerance).

For a Beta$(a, b)$ posterior: mean $\frac{a}{a+b}$, mode $\frac{a-1}{a+b-2}$. For symmetric, single-peaked posteriors all three coincide; with more data, posteriors become nearly Normal and the three converge. Chapter 6.4 covers posterior summaries and credible intervals in full.

Why do we need it?

Dashboards and reports show one number. Knowing which summary you show, and how far apart the summaries are, prevents quiet misreporting, especially for rare events, small samples and positive quantities with long right tails.

Where is it used?

Posterior means from MCMC or SVI draws (the usual default in NumPyro workflows), posterior medians for skewed quantities like rates and scales, MAP from optimization-based fits (ridge, lasso, AutoDelta), and point forecasts taken from a predictive distribution.

How is it used?

With posterior draws: samples.mean(), np.median(samples); the MAP needs an optimizer or a density estimate. Pick the summary that matches the cost of errors, say which one you report, and show an interval next to it.

The orange hill is the posterior for a conversion rate after $k$ conversions in $n$ users (flat prior). Purple = MAP (peak), teal = median, pink = mean. Start at 1 of 50: the three lines are clearly apart, with the mean pulled right by the tail. Slide $k$ up to 25: the hill becomes symmetric and the lines merge. Turn on expected squared error: the dashed curve shows the average squared miss of every possible guess, and its lowest point sits exactly on the mean.

"The MAP is the most probable value, so it is the best number to report."

"Best" depends on the cost of errors. Under squared error the posterior mean wins; under absolute error the median wins. The MAP is best only if nothing but an exact hit counts.

"The peak is where most of the posterior's mass is."

Not necessarily. In many dimensions, or for skewed posteriors, most of the probability can sit away from the peak. In hierarchical models the joint posterior density can even grow without limit as the group spread τ shrinks to 0, so a joint MAP can collapse to a degenerate answer (the narrow neck of the funnel in Chapter 6.7 is the same geometry).

If your A/B framework reports a conversion rate or a lift from posterior draws (check which summary your code uses), that is a posterior mean or median, not a MAP. For low-conversion metrics the posterior is skewed, so mean, median and mode can differ noticeably; say which one a dashboard shows. In your forecasting model the same choice appears for point forecasts: the mean or the median of the predictive distribution (Chapter 7.14), and they differ whenever the predictive distribution is skewed, as it is for Negative Binomial counts.

"MAP and posterior mean are just two names for the Bayesian estimate."

"They are different summaries of the same posterior, optimal under different losses."

Model answer: "The posterior mean minimizes expected squared error and the median minimizes expected absolute error; the MAP is the mode. For symmetric unimodal posteriors they agree; for skewed ones, such as rare conversion rates, the mean is pulled toward the tail. MAP also comes from optimization, not integration, so it carries no uncertainty and it depends on the parameterization."

Mean ↔ squared error; median ↔ absolute error; MAP = mode (peak).

Beta(a, b): mean $a/(a+b)$, mode $(a-1)/(a+b-2)$. 1 of 50, flat prior: MAP 0.020, median 0.033, mean 0.038.

Trap: skewed posteriors spread the three apart; always say which one you report.

Quick check: a posterior is Normal with mean 3 and sd 0.5. What are its MAP, median and mean?

All three are 3: a Normal is symmetric and has a single peak, so peak, middle and balance point coincide.

Change the ruler: the MLE stays put, the MAP moves core

You can describe a conversion rate as a probability $p$ (between 0 and 1) or as log-odds $\eta = \log\frac{p}{1-p}$ (any real number; logistic regression and many NumPyro models work in this scale). Same quantity, different ruler.

The likelihood just asks "how well does this value explain the data?". Moving to a new ruler relabels the values but does not change the scores, so the best value is the same point: the MLE "moves with the ruler" and stays the same answer.

A density is different. It is probability per unit length, and changing the ruler stretches some regions and squeezes others. Where the ruler stretches, the density must get lower (the same probability spread over more length); where it squeezes, the density gets higher. So the posterior's peak can move to a different value of θ in the new ruler. The MAP depends on the ruler.

Three ways to say it:

  • Picture: stretch a rubber sheet with a hill painted on it; the top of the painted hill shifts.
  • Numbers: for a Beta(9, 5) posterior the MAP is $p = 0.667$, but the MAP in log-odds corresponds to $p = 0.643$.
  • Slogan: likelihoods relabel; densities stretch. MLE is invariant, MAP is not.

7 of 10 with a Beta(2, 2) prior: posterior Beta(9, 5).

  1. MAP on the $p$ ruler: $(9 - 1)/(9 + 5 - 2) = 8/12 \approx 0.667$. Moved to log-odds: $\log(0.667/0.333) = \log 2 \approx 0.693$.
  2. Change ruler: $p = \sigma(\eta) = 1/(1 + e^{-\eta})$, and $dp/d\eta = p(1-p)$. Densities pick up this stretching factor: $p_\eta(\eta) = p_p\big(\sigma(\eta)\big)\cdot p(1-p) \propto p^{8}(1-p)^{4}\cdot p(1-p) = p^{9}(1-p)^{5}$.
  3. Maximize $9\log p + 5\log(1-p)$: $9/p = 5/(1-p) \Rightarrow p = 9/14 \approx 0.643$, so the MAP on the log-odds ruler is $\eta = \log(9/5) \approx 0.588$.
  4. $0.588 \ne 0.693$: the two MAPs are different points. (Curiously, $9/14$ is the posterior mean of $p$.)
  5. MLE: $\hat p = 0.7$. On the log-odds ruler the likelihood is the same function relabelled (no stretching factor), so its peak is at $\hat\eta = \log(0.7/0.3) \approx 0.847$: exactly $\hat p$ moved over.
  6. The median also moves cleanly: median of $p$ ≈ 0.650, and the median of $\eta$ is $\log(0.650/0.350) \approx 0.618$, because "half the probability on each side" survives any increasing change of ruler.

An extreme case: 2 of 2 with a flat prior gives Beta(3, 1). Its MAP on the $p$ ruler is $p = 1$ (at the edge, log-odds $+\infty$); its MAP on the log-odds ruler is at $p = 3/4$.

Let $\eta = g(\theta)$ be a one-to-one change of parameter.

  • MLE is invariant: $\hat\eta_{MLE} = g(\hat\theta_{MLE})$. The likelihood in the new parameter is $L(g^{-1}(\eta))$, the same values relabelled, so the maximizer maps across.
  • MAP is not invariant: the posterior density transforms with a Jacobian (the stretching factor), $p_\eta(\eta\mid D) = p_\theta\big(g^{-1}(\eta)\mid D\big)\,\Big|\dfrac{d g^{-1}(\eta)}{d\eta}\Big|$, and the extra factor moves the peak. In general $\hat\eta_{MAP} \ne g(\hat\theta_{MAP})$.
  • Posterior quantiles (median, interval endpoints) are invariant under increasing transformations. The posterior mean is not: $E[g(\theta)] \ne g(E[\theta])$ in general (Jensen).
  • Consequence: "MAP with a flat prior = MLE" is true only on the ruler where the prior is flat. A prior flat in $p$ is not flat in log-odds (Chapter 6.2).
Why do we need it?

Models are routinely fitted on a transformed scale (log σ, logit p, log λ) for numerical convenience. If you report a MAP, the number depends on that hidden choice; knowing this prevents confusing, inconsistent reports.

Where is it used?

Logistic regression (log-odds), positive parameters optimized as $\log\sigma$, NumPyro's transforms to unconstrained space for SVI and HMC, the choice of priors written on σ versus log σ, and the classic argument for reporting posterior medians and intervals rather than modes.

How is it used?

Report MLEs freely on any scale (transform the estimate). For Bayesian results, prefer medians and quantile intervals, which transform cleanly; if you report a MAP, say on which parameter scale it was computed.

flat density on p same probability on η = log(p / (1 − p)) 00.51 −505 0.5–0.6 → a narrow strip 0.9–1.0 → a strip stretched to infinity each coloured strip holds probability 0.1 in both pictures
Changing the ruler from p to log-odds squeezes the middle strips and stretches the outer ones. The probability in each strip is kept, so the density must rise where strips are squeezed: a flat density with no special peak becomes a bell with a peak at η = 0. Densities, and therefore MAPs, depend on the ruler.

Top: the posterior of a rate $p$. Bottom: the same posterior on the log-odds ruler $\eta = \log(p/(1-p))$. Purple solid = the MAP computed on each ruler; purple dashed (bottom) = the top MAP carried over to log-odds. They do not line up. The teal median and the blue MLE carry over exactly. Press 2 of 2, flat prior: the top MAP sits at the edge $p = 1$ (log-odds $+\infty$), while the bottom MAP corresponds to $p = 0.75$. Add data and the gap shrinks.

"MAP with a flat prior is always the MLE."

Only on the ruler where the prior is flat. Flat on $p$ means not flat on log-odds (the figure), so the "flat prior MAP" depends on which parameter you called flat.

"If I fit on log σ and transform the MAP back, I get the MAP of σ."

You get the image of the log σ MAP, which is generally a different number from the MAP computed on σ. For the MLE the back-transform is fine; for the MAP it is not.

"The posterior mean is the safe, invariant summary."

The mean is not invariant either: $E[e^{\eta}] \ne e^{E[\eta]}$. Quantiles (median, interval endpoints) are the summaries that carry over exactly under increasing transformations.

NumPyro runs SVI and NUTS on an unconstrained scale: a positive scale such as your forecasting model's σ is handled as $\log\sigma$, and a probability as a log-odds-like value, then transformed back. Draws transformed back are fine for means, medians and intervals of the original parameters. But a Gaussian guide such as AutoNormal is Gaussian on the unconstrained scale, so its centre, transformed back, is the guide's median on the original scale, not its mean and not its mode. That is exactly the "quantiles carry over, modes and means do not" rule; NumPyro's autoguides have a median() method that returns this transformed centre. (For point estimates, NumPyro's documentation notes that AutoDelta does MAP inference in the constrained space, that is, on the parameters as you wrote them in the model.)

"MAP is invariant to reparameterization, like the MLE."

"The MLE is invariant; the MAP is not, because a posterior density picks up a Jacobian factor when the parameter is transformed, and that factor moves the mode."

Model answer: "The likelihood is a function of the parameter, not a density over it, so relabelling the parameter just relabels the function and the maximizer maps across. The posterior is a density over the parameter; under $\eta = g(\theta)$ it is multiplied by $|d\theta/d\eta|$. For a Beta(9, 5) posterior, the mode in $p$ is 0.667 but the mode in log-odds corresponds to 0.643. Medians and quantiles are invariant, which is one reason to report them."

MLE invariant: $\widehat{g(\theta)} = g(\hat\theta)$. MAP not: $p_\eta(\eta) = p_\theta(\theta)\,|d\theta/d\eta|$ moves the mode.

Beta(a, b): mode in $p$ = $(a-1)/(a+b-2)$; mode in log-odds ↔ $p = a/(a+b)$.

Trap: "flat prior" depends on the ruler. Quantiles transform cleanly; modes and means do not.

Quick check: the MLE of a Poisson rate is λ̂ = 4. What is the MLE of log λ, and of the probability of a zero count, $e^{-\lambda}$?

By invariance, the MLE of $\log\lambda$ is $\log 4 \approx 1.386$, and the MLE of $P(Y = 0) = e^{-\lambda}$ is $e^{-4} \approx 0.0183$. No new optimization needed.

Recap, cheat sheet and practice

  • The likelihood $L(\theta) = p(D\mid\theta)$ scores how well each θ explains the fixed data. It is not a probability distribution over θ (the coin example has area 1/11); only ratios matter.
  • Work with the log-likelihood $\ell(\theta) = \sum\log p(x_i\mid\theta)$: same maximizer, no underflow.
  • MLE = $\arg\max\ell$: coin $k/n$; Normal $\bar x$ and $\frac1n\sum(x_i-\bar x)^2$ (biased low); Poisson $\bar y$. Check edges ($k = 0$).
  • Curvature $J = -\ell''(\hat\theta)$ gives $SE \approx 1/\sqrt J$. For large $n$ and a correct model, the MLE is consistent, approximately Normal, efficient and invariant.
  • Minimizing a negative log-likelihood = minimizing a familiar loss: Gaussian ↔ squared error, Laplace ↔ absolute error, Bernoulli ↔ cross-entropy, Poisson ↔ $\lambda - y\log\lambda$.
  • MAP = $\arg\max[\ell(\theta) + \log p(\theta)]$; with a Beta prior, $(k+\alpha-1)/(n+\alpha+\beta-2)$. The negative log-prior acts as a penalty (Chapter 5.3).
  • MAP is the posterior's peak; the posterior mean (squared error) and median (absolute error) are different summaries, far apart for skewed posteriors.
  • The MLE is invariant to reparameterization; the MAP is not (densities carry a Jacobian). Quantiles carry over; modes and means do not.

Cheat sheet

IdeaFormulaRemember
Likelihood$L(\theta) = \prod_i p(x_i\mid\theta)$data fixed, θ varies; not a distribution over θ
Log-likelihood$\ell(\theta) = \sum_i \log p(x_i\mid\theta)$same peak; never multiply probabilities in code
MLE$\arg\max_\theta \ell(\theta)$; solve $\ell'(\theta) = 0$check $\ell'' \lt 0$ and the edges
Coin / Normal / Poisson$k/n$; $\ \bar x,\ \frac1n\sum(x_i-\bar x)^2$; $\ \bar y$Normal σ² MLE divides by $n$
Standard error$SE \approx 1/\sqrt{-\ell''(\hat\theta)}$coin: $\sqrt{\hat\theta(1-\hat\theta)/n}$
Losses$-\log p$: Normal → $(y-\mu)^2$, Laplace → $|y-\mu|$, Bernoulli → log lossa loss is a noise model
MAP$\arg\max[\log p(D\mid\theta) + \log p(\theta)]$MLE + penalty $-\log p(\theta)$
Beta-Binomial MAP$\frac{k+\alpha-1}{n+\alpha+\beta-2}$; mean $\frac{k+\alpha}{n+\alpha+\beta}$pseudo-counts $\alpha-1$, $\beta-1$
Posterior summariesmean ↔ squared error; median ↔ absolute error; mode = MAPskewed posterior: they differ
Invariance$\widehat{g(\theta)}_{MLE} = g(\hat\theta)$; $p_\eta(\eta) = p_\theta(\theta)\,|d\theta/d\eta|$MLE invariant, MAP not
Code it · Python

import numpy as np
from scipy import stats, optimize

# 1) Coin: the likelihood on a grid; its peak is k/n and its area is NOT 1
k, n = 7, 10
theta = np.linspace(0, 1, 1001)
L = stats.binom.pmf(k, n, theta)              # p(k | theta), read as a function of theta
print(round(theta[np.argmax(L)], 3), round(L.max(), 4))   # 0.7 0.2668
print(round(np.trapezoid(L, theta), 4))       # 0.0909 = 1/11  (np.trapz in NumPy < 2.0)

# 2) Underflow: multiply 1000 densities vs add 1000 logs
x = np.full(1000, 0.05)
print(np.prod(x), round(np.log(x).sum(), 2))  # 0.0 -2995.73

# 3) Normal MLE by numerical optimization (optimize log(sigma) so sigma stays positive)
y = np.array([4, 6, 8, 10, 12.0])
def nll(params):
    mu, log_sigma = params
    return -stats.norm.logpdf(y, loc=mu, scale=np.exp(log_sigma)).sum()   # scale = sigma, not sigma**2
res = optimize.minimize(nll, x0=[0.0, 0.0])
mu_hat, sigma_hat = res.x[0], np.exp(res.x[1])
print(round(mu_hat, 4), round(sigma_hat**2, 4), round(-res.fun, 4))   # 8.0 8.0 -12.2933
print(np.var(y), np.var(y, ddof=1))           # 8.0 10.0   (MLE divides by n; ddof=1 by n-1)
print([round(float(v), 4) for v in stats.norm.fit(y)])   # [8.0, 2.8284]  SciPy's fit = the MLE

# 4) Poisson MLE = the average count
counts = np.array([2, 4, 3, 6, 5])
res = optimize.minimize_scalar(lambda lam: -stats.poisson.logpmf(counts, lam).sum(),
                               bounds=(0.01, 20), method="bounded")
print(round(res.x, 4), counts.mean())         # 4.0 4.0

# 5) Standard error from the curvature of the log-likelihood (numerical second derivative)
ll = lambda t: stats.binom.logpmf(7, 10, t)
t0, h = 0.7, 1e-4
J = -(ll(t0 + h) - 2 * ll(t0) + ll(t0 - h)) / h**2
print(round(J, 2), round(1 / np.sqrt(J), 4), round(np.sqrt(0.7 * 0.3 / 10), 4))   # 47.62 0.1449 0.1449

# 6) MAP with a Beta(2, 2) prior after 3 of 3, vs MLE, posterior mean and median
a, b, k, n = 2, 2, 3, 3
post = stats.beta(a + k, b + n - k)           # Beta(5, 2); SciPy's beta(a, b) = NumPyro Beta(concentration1=a, concentration0=b)
print((k + a - 1) / (n + a + b - 2), k / n, round(post.mean(), 4), round(post.median(), 4))   # 0.8 1.0 0.7143 0.7356

# 7) The MAP moves when you change the ruler (p -> log-odds); the MLE does not
A, B = 9, 5                                   # posterior after 7 of 10 with a Beta(2, 2) prior
eta = np.linspace(-4, 4, 80001)
p = 1 / (1 + np.exp(-eta))
log_dens_eta = A * np.log(p) + B * np.log(1 - p)   # Beta(9,5) density x Jacobian p(1-p), up to a constant
print(round((A - 1) / (A + B - 2), 4), round(p[np.argmax(log_dens_eta)], 4))   # 0.6667 0.6429
Test yourself

1. For 7 successes in 10 tries, the likelihood curve $L(\theta)$ has area $1/11$ under it. What does that tell you?

The likelihood is $p(D\mid\theta)$ read as a function of θ. It sums to 1 over possible data for fixed θ, not over θ. Only ratios of likelihoods are meaningful.

2. Data 1, 3, 5 from a Normal. The maximum likelihood estimate of $\sigma^2$ is…

$\bar x = 3$; squared distances 4, 0, 4 sum to 8. The MLE divides by $n = 3$: $8/3 \approx 2.67$. The unbiased $s^2$ divides by 2 and gives 4.

3. Fitting a constant by minimizing the sum of absolute errors is maximum likelihood under which noise model?

The Laplace density is $\propto e^{-|y-\mu|/b}$, so its negative log is $|y-\mu|/b$ plus a constant. The minimizer is the median. Squared error is the Normal; log loss is the Bernoulli.

4. Prior Beta(3, 3), data 2 conversions in 10 users. The MAP estimate is…

$(k+\alpha-1)/(n+\alpha+\beta-2) = (2+2)/(10+4) = 4/14 \approx 0.286$. The prior adds 2 pseudo-successes and 2 pseudo-failures, pulling the MLE 0.2 toward 0.5.

5. You re-express a parameter on a new scale (for example log-odds instead of probability). Which estimate is guaranteed to map across unchanged?

The likelihood is not a density over the parameter, so relabelling does not change where it peaks: $\widehat{g(\theta)} = g(\hat\theta)$. Posterior densities pick up a Jacobian, which moves the MAP; means change by Jensen's inequality. (Posterior quantiles also map across under increasing transformations.)

6. A posterior for a rare conversion rate is Beta(2, 50), skewed to the right. Which ordering is correct?

MAP $= 1/50 = 0.020$, median ≈ 0.033, mean $= 2/52 \approx 0.038$. The long right tail pulls the mean furthest to the right; the peak stays on the left.

Practice problems

A. Waiting times between orders (minutes): 2, 4, 6, modelled as Exponential with rate λ, density $\lambda e^{-\lambda x}$. Find the MLE of λ and of the mean waiting time $1/\lambda$.
  1. $\ell(\lambda) = \sum(\log\lambda - \lambda x_i) = n\log\lambda - \lambda\sum x_i = 3\log\lambda - 12\lambda$.
  2. $\ell'(\lambda) = 3/\lambda - 12 = 0 \Rightarrow \hat\lambda = 3/12 = 0.25$ per minute. $\ell''(\lambda) = -3/\lambda^2 \lt 0$: a maximum.
  3. By invariance, the MLE of the mean waiting time $1/\lambda$ is $1/0.25 = 4$ minutes, which is just $\bar x$.
B. Show that minimizing $\sum|x_i - \mu|$ gives the median, using the data 1, 2, 3, 4, 20.

The slope of $\sum|x_i - \mu|$ in μ is (number of points below μ) − (number of points above μ). It is negative while more points lie above, positive once more lie below, so the minimum is where the counts balance: the median, 3. Check: $\sum|x_i - 3| = 2+1+0+1+17 = 21$, while at the mean 6, $\sum|x_i - 6| = 5+4+3+2+14 = 28$. The outlier 20 drags the mean but not the median.

C. Show that the Beta-Binomial MAP approaches the MLE as $n$ grows (with the observed rate $k/n$ held fixed).

$\frac{k+\alpha-1}{n+\alpha+\beta-2} - \frac kn$. Divide top and bottom of the first fraction by $n$: $\frac{k/n + (\alpha-1)/n}{1 + (\alpha+\beta-2)/n}$. As $n \to \infty$ the extra terms vanish and the fraction tends to $k/n$. The prior's pseudo-counts become negligible next to real data, which is also why priors matter most for small segments.

D. (Interview) "Why is least squares linear regression a maximum likelihood method, and what does it assume?"

"If $y_i = \mathbf{x}_i^\top\boldsymbol\beta + \varepsilon_i$ with $\varepsilon_i$ iid $N(0, \sigma^2)$, the negative log-likelihood is $\frac{1}{2\sigma^2}\sum(y_i - \mathbf{x}_i^\top\boldsymbol\beta)^2 + n\log\sigma + \text{const}$. For any fixed σ, minimizing it over β is minimizing the sum of squared errors, so OLS is the Gaussian MLE. The assumptions are independent errors with constant variance and light, Normal-like tails. With heavy tails a Laplace or Student-t likelihood (absolute error or a robust loss) is more appropriate; with counts, a Poisson or Negative Binomial GLM."

E. A shop was open 8 hours on Monday (20 orders) and 4 hours on Tuesday (14 orders). With $y_i \sim Poisson(\lambda t_i)$, what is the MLE of the hourly rate λ?

$\ell(\lambda) = \sum[y_i\log(\lambda t_i) - \lambda t_i] + \text{const}$, so $\ell'(\lambda) = \sum y_i/\lambda - \sum t_i = 0$ and $\hat\lambda = \sum y_i/\sum t_i = 34/12 \approx 2.83$ orders per hour. Averaging the two daily rates, $(2.5 + 3.5)/2 = 3.0$, would wrongly give the short day as much weight as the long one.

F. (Interview) "Is the MAP estimate invariant to reparameterization? Give an example."

"No. Take a coin with 2 heads in 2 tosses and a flat prior: the posterior is Beta(3, 1) with density $\propto p^2$, whose peak is at $p = 1$. On the log-odds scale the density picks up the Jacobian $p(1-p)$, giving $\propto p^3(1-p)$, whose peak is at $p = 3/4$. Same posterior, two different MAPs. The MLE ($p = 1$ on any scale) and the posterior quantiles do map across, which is why I report medians and intervals for Bayesian results."

Chapter 5.3 · Syllabus Module 11

Regularization as a prior: Ridge, Lasso, Gaussian and Laplace priors

A model with many knobs and little data will twist itself to fit the noise. Regularization adds a price for big knob settings. This chapter shows that the two most famous prices, ridge (L2) and lasso (L1), are exactly what you get from a Gaussian or a Laplace prior when you take the single most probable answer (the MAP). Then it shows the part most people get wrong: the full posterior under a Laplace prior is not sparse. Only the MAP is. That distinction sits right inside your forecasting model, where every changepoint slope δⱼ has a Laplace prior.

  • Explain why we regularize, and read the objective loss + λ × penalty, with λ as a dial between "fit the data" and "keep the coefficients small"
  • Compute ridge (L2) estimates, see why they shrink every coefficient but never make one exactly 0, and why they calm down collinear features
  • Compute lasso (L1) estimates with soft-thresholding, and explain with the diamond vs circle picture why L1 produces exact zeros
  • Read coefficient paths as λ grows, for ridge and for lasso
  • Derive: Gaussian prior ↔ ridge with $\lambda = \sigma^2/\tau^2$, and Laplace prior ↔ lasso with $\lambda \propto \sigma^2/b$, both at the MAP
  • State the precise distinction: under a Laplace prior the MAP can be exactly 0, but the full posterior has no point mass at 0, and its mean and median are almost never exactly 0
  • Explain why $\delta_j \sim Laplace(0, b)$ gives "sparse-ish" changepoints in a Prophet-style trend

What we need from earlier chapters: the bias–variance trade-off and MSE (Chapter 5.1); the likelihood, MLE, MAP and "loss = negative log-likelihood" (Chapter 5.2); the Normal and Laplace distributions (Chapter 4.9); vector norms, the L1 diamond and the L2 circle (Linear Algebra 1.2); penalties, soft-thresholding and coordinate descent as optimization tools (Optimization 3.11 and 3.13). This chapter adds the statistical meaning: a penalty is a prior in disguise. Notation: $\beta = (\beta_1, \dots, \beta_p)$ are regression coefficients, $X$ is the $n \times p$ table of features, $y$ the targets, $\mathrm{RSS}(\beta) = \|y - X\beta\|^2 = \sum_i (y_i - x_i^\top\beta)^2$ is the residual sum of squares. $N(\mu, \sigma^2)$ is written with the variance; NumPyro and SciPy take the standard deviation.

Why regularize? Loss + λ × penalty core

Imagine fitting a curve through 12 noisy points with a model that has 10 adjustable coefficients (the numbers the model multiplies its features by). The model can bend almost anywhere. Plain least squares only asks "how close is the curve to the points?", so it bends through every point, noise included. The curve looks perfect on the data you have and silly everywhere else. That is overfitting: learning the noise as if it were the signal.

Regularization changes the question. It asks: "how close is the curve to the points, plus how big are the coefficients?" Big coefficients now cost something, so the model only uses them when the data really insist. A dial called λ (lambda) sets the price: λ = 0 is plain least squares; a huge λ makes every coefficient so expensive that the model gives up and predicts a flat line.

Three ways to say it:

  • Picture: a rubber band ties every coefficient to 0; the data pull them away; λ is how stiff the band is.
  • Numbers: a wild fit with residual sum of squares 4.0 and coefficients (3, −2.5) loses to a calm fit with 4.6 and (0.4, 0.3) as soon as λ is above 0.04.
  • Slogan: fit the data, but pay for every unit of complexity.

Two candidate coefficient vectors for the same data. Candidate A fits a little better; candidate B is much calmer.

  1. A: $\beta = (3, -2.5)$, residual sum of squares $\mathrm{RSS} = 4.0$. B: $\beta = (0.4, 0.3)$, $\mathrm{RSS} = 4.6$.
  2. No penalty (λ = 0): A wins, because $4.0 \lt 4.6$.
  3. Ridge penalty $\sum_j \beta_j^2$: A pays $3^2 + 2.5^2 = 9 + 6.25 = 15.25$; B pays $0.4^2 + 0.3^2 = 0.16 + 0.09 = 0.25$.
  4. With λ = 1: A scores $4.0 + 1 \times 15.25 = 19.25$; B scores $4.6 + 1 \times 0.25 = 4.85$. B wins.
  5. Where do they tie? $4.0 + 15.25\lambda = 4.6 + 0.25\lambda$ gives $15\lambda = 0.6$, so $\lambda = 0.04$. Even a small price flips the choice.
  6. Lasso penalty $\sum_j |\beta_j|$ instead: A pays $3 + 2.5 = 5.5$, B pays $0.7$. With λ = 1: A scores $9.5$, B scores $5.3$. Same verdict.

A regularized estimate minimizes a data-fit loss plus a penalty on the size of the coefficients:

$$\hat\beta_\lambda = \arg\min_{\beta}\; \Big[\, \underbrace{\mathrm{Loss}(\beta; D)}_{\text{fit to the data}} \;+\; \lambda \cdot \underbrace{\mathrm{Penalty}(\beta)}_{\text{size of the coefficients}} \,\Big], \qquad \lambda \ge 0.$$
  • Ridge (L2): $\mathrm{Penalty} = \|\beta\|_2^2 = \sum_j \beta_j^2$. Lasso (L1): $\mathrm{Penalty} = \|\beta\|_1 = \sum_j |\beta_j|$. "L2" and "L1" are the names of these two ways of measuring size (norms).
  • $\lambda$ is the regularization strength: λ = 0 gives the unpenalized estimate (least squares, or the MLE in general); λ → ∞ pushes every penalized coefficient to 0.
  • Shrinkage means "pulled toward 0 compared with the unpenalized estimate". Sparse means "many coefficients exactly 0".
  • The intercept (the baseline level) is usually not penalized, and features are usually standardized first (the penalty compares coefficient sizes, and sizes depend on the units of each feature).
  • λ is not learned by minimizing the same objective (that would always choose λ = 0). It is chosen by cross-validation, by a rule, or, as you will see, by a prior.
Why do we need it?

With many coefficients and limited data, unpenalized estimates have huge variance: they chase noise, swing wildly between samples and predict badly on new data. A penalty trades a little bias for a big drop in variance (Chapter 5.1), which usually lowers the total error.

Where is it used?

Ridge and lasso regression (scikit-learn Ridge, Lasso, ElasticNet; R's glmnet), weight decay in neural networks, penalized logistic regression, and every Bayesian model with a prior on its coefficients: the Normal priors on seasonality and holiday effects and the Laplace prior on changepoint slopes in a Prophet-style model.

How is it used?

Standardize the features, leave the intercept unpenalized, fit the model for a range of λ values, and pick λ by cross-validation (RidgeCV, LassoCV) or set it through a prior scale. Then look at the coefficients: how much they shrank, and (for the lasso) which ones became exactly 0.

fit to the data RSS = Σ (yᵢ − ŷᵢ)² + λ the dial × size of the coefficients Σ βⱼ² (ridge) or Σ |βⱼ| (lasso) → minimize over β λ = 0 plain least squares chases the noise λ "just right" smooth, sensible fit lowest error on new data λ huge all coefficients ≈ 0 ignores the data
Every regularized objective has the same three parts: a data-fit term, a dial λ, and a size penalty. Turning the dial moves you from overfitting (left) to underfitting (right).

Twelve noisy points (blue) come from a smooth wave (green, dashed). The model is a polynomial: an intercept plus 9 coefficients for $u, u^2, \dots, u^9$ (with $u = 2x - 1$); the bottom panel shows the 9 coefficients. Press λ tiny: the orange/teal fit wiggles through the points and the coefficients explode (bars past the edge show their value). Press Best λ for this sample: the fit becomes smooth and close to the truth. Press λ huge: the fit flattens to the average. Switch to Lasso and watch some bars become exactly 0 (marked "0"). Press New sample to see that the wiggles at tiny λ change completely from sample to sample: that is high variance.

"A smaller training error always means a better model."

At λ = 0 the training error is smallest and the error on the true curve is often the worst. The penalty deliberately accepts a worse fit to the sample you have, to do better on data you have not seen.

"Penalize the intercept too; it is just another coefficient."

The intercept sets the overall level. Penalizing it would drag every prediction toward 0 for no reason. Libraries leave it unpenalized by default.

"The penalty does not care about units."

It does. A price feature in dollars needs a coefficient 100 times bigger than the same feature in cents, so it pays a different penalty. Standardize the features first, or choose a prior scale that matches each feature's units.

$\hat\beta_\lambda = \arg\min\, [\mathrm{Loss} + \lambda \cdot \mathrm{Penalty}]$. Ridge: $\sum\beta_j^2$. Lasso: $\sum|\beta_j|$.

λ = 0: unpenalized (overfits); λ → ∞: everything → 0 (underfits). Choose λ by cross-validation or through a prior.

Trap: standardize features and do not penalize the intercept.

Quick check: with λ = 0.02, which candidate in the example wins under the ridge penalty?

A scores $4.0 + 0.02\times15.25 = 4.305$; B scores $4.6 + 0.02\times0.25 = 4.605$. A still wins, because λ = 0.02 is below the tie point 0.04.

Ridge regression (L2): shrink everything, smoothly core

Ridge charges each coefficient its square. A square is tiny for small numbers ($0.1^2 = 0.01$) and huge for big ones ($10^2 = 100$). So ridge hates big coefficients and barely notices small ones. The result: every coefficient is pulled toward 0 by a fraction of its size, like a spring. Springs pull, but they never snap a coefficient all the way to exactly 0.

Ridge shines when features are near-copies of each other (collinear, for example "price" and "price including tax"). Least squares cannot tell which twin deserves the credit, so it gives one a huge positive weight and the other a huge negative weight that almost cancel. Ridge finds that silly and splits the credit evenly.

Three ways to say it:

  • Picture: a spring from every coefficient to 0; the farther out, the harder it pulls.
  • Numbers: with one feature, the estimate $13/14 \approx 0.93$ becomes $13/(14 + \lambda)$: $0.62$ at λ = 7, $0.46$ at λ = 14.
  • Slogan: ridge shrinks, it does not select.

One feature, no intercept: $x = (1, 2, 3)$, $y = (1, 3, 2)$. Minimize $\sum_i (y_i - \beta x_i)^2 + \lambda\beta^2$.

  1. Set the derivative to zero: $-2\sum_i x_i(y_i - \beta x_i) + 2\lambda\beta = 0$, so $\beta = \dfrac{\sum x_i y_i}{\sum x_i^2 + \lambda}$.
  2. $\sum x_i y_i = 1 + 6 + 6 = 13$ and $\sum x_i^2 = 1 + 4 + 9 = 14$.
  3. λ = 0 (least squares): $\beta = 13/14 \approx 0.929$.
  4. λ = 7: $\beta = 13/21 \approx 0.619$. λ = 14: $\beta = 13/28 \approx 0.464$, exactly half of the least-squares value.
  5. However large λ gets, $13/(14 + \lambda)$ is never exactly 0. It only gets closer.

The ridge estimate minimizes the residual sum of squares plus λ times the squared L2 norm:

$$\hat\beta_{\text{ridge}} = \arg\min_\beta\, \|y - X\beta\|^2 + \lambda\|\beta\|_2^2 = (X^\top X + \lambda I)^{-1}X^\top y.$$
  • For λ > 0 the matrix $X^\top X + \lambda I$ is always invertible, even when features are perfectly collinear or there are more features than rows. Ridge always has one unique answer.
  • If the features are orthonormal ($X^\top X = I$), every coefficient is simply divided: $\hat\beta_j^{\text{ridge}} = \hat\beta_j^{\text{OLS}}/(1 + \lambda)$. In general the shrinkage is strongest along directions where the data carry little information (small singular values $d$: factor $d^2/(d^2 + \lambda)$, see Optimization 3.11).
  • Ridge is biased toward 0 but has lower variance than least squares; for a good λ the MSE is lower (Chapter 5.1).
  • Library convention: scikit-learn's Ridge(alpha=λ) minimizes exactly $\|y - X\beta\|^2 + \lambda\|\beta\|^2$.
Why do we need it?

Least squares breaks down when features are strongly correlated or when there are many features for few rows: the coefficients swing wildly from sample to sample, or the answer is not unique. Ridge makes the problem stable and the predictions more reliable.

Where is it used?

Linear and logistic regression with many correlated features (scikit-learn Ridge, LogisticRegression with its default L2 penalty), weight decay in neural networks, kernel ridge regression, and, as a Normal prior, the seasonality and holiday coefficients of a Prophet-style model.

How is it used?

Standardize features, fit RidgeCV(alphas=...) over a grid of λ, and check that the coefficients of correlated features became similar and stable. Expect every coefficient to be non-zero: if you need feature selection, ridge is the wrong tool.

Forty separate datasets, each with 30 rows and two features that are near-copies (correlation about 0.995). The truth is $\beta = (1, 1)$ (green). Each blue dot is the least-squares estimate from one dataset; each teal dot is the ridge estimate from the same dataset. The blue dots spread far along the dashed line $\beta_1 + \beta_2 = 2$: the data pin down the sum of the twins but not how to split it. Slide λ up from 0 and watch the teal cloud collapse toward the truth. Press New samples for 40 fresh datasets.

"Ridge removes useless features."

Ridge pulls the coefficients toward 0 (the whole set always gets smaller as λ grows; with strongly correlated features one coefficient can briefly move the other way), but none becomes exactly 0 (for λ < ∞). Every feature stays in the model. For exact zeros you need the lasso (next section).

"Ridge is unbiased because it is still least squares."

The penalty pulls the estimate toward 0, so on average it falls short of the truth: it is biased. You accept the bias because the variance drops much more.

"Correlated twins get one big and one tiny weight under ridge."

Ridge tends to split the weight evenly between near-identical features. The lasso tends to pick one twin and zero the other.

$\hat\beta_{\text{ridge}} = (X^\top X + \lambda I)^{-1}X^\top y$; one feature: $\sum x_iy_i/(\sum x_i^2 + \lambda)$.

Shrinks the coefficients smoothly toward 0, never exactly to 0; stabilizes collinear features (splits the credit).

Trap: biased but lower variance; it does not select features.

Quick check: with orthonormal features, the least-squares coefficients are (4, −2, 0.5). What are the ridge coefficients for λ = 1?

Each is divided by $1 + \lambda = 2$: $(2, -1, 0.25)$. All three are still non-zero; the small one shrank by the same fraction as the big one.

Lasso (L1): shrink, and switch small effects off core

The lasso charges each coefficient its absolute value: the same price per unit, whether the coefficient is small or big. Think of a toll booth: every coefficient must pay a fixed toll λ to be allowed to move away from 0. A coefficient whose evidence is weaker than the toll cannot afford it and stays at exactly 0. A coefficient with strong evidence pays the toll and moves, but it ends up λ smaller than it would have been.

That is why the lasso does two jobs at once: it selects features (the zeros are switched off) and it shrinks the survivors. A model with many exact zeros is called sparse.

Three ways to say it:

  • Picture: a dead zone around 0: estimates that land inside it are snapped to 0; the others are moved toward 0 by a fixed step.
  • Numbers: with the same one-feature data as ridge, the estimate is $(13 - \lambda)/14$: $0.43$ at λ = 7, and exactly $0$ from λ = 13 on.
  • Slogan: ridge shrinks by a fraction; lasso subtracts a constant, and stops at zero.

Same data: $x = (1, 2, 3)$, $y = (1, 3, 2)$. Minimize $\tfrac12\sum_i (y_i - \beta x_i)^2 + \lambda|\beta|$.

  1. Expand: $\tfrac12\sum (y_i - \beta x_i)^2 = \tfrac12(14\beta^2 - 26\beta + 14)$, because $\sum x_i^2 = 14$, $\sum x_iy_i = 13$, $\sum y_i^2 = 14$.
  2. Try $\beta \gt 0$, where $|\beta| = \beta$: the derivative is $14\beta - 13 + \lambda$. Setting it to 0 gives $\beta = (13 - \lambda)/14$, which is positive only if $\lambda \lt 13$.
  3. λ = 7: $\beta = 6/14 \approx 0.429$ (ridge gave $0.619$ at the same λ).
  4. λ = 13 or more: the formula would give $\beta \le 0$, contradicting $\beta \gt 0$. Trying $\beta \lt 0$ fails the same way. The only candidate left is the corner of $|\beta|$: $\beta = 0$ exactly.
  5. In one line: $\hat\beta = \dfrac{\mathrm{sign}(13)\max(13 - \lambda, 0)}{14}$. This "subtract, and stop at 0" rule is called soft-thresholding.

The lasso estimate (least absolute shrinkage and selection operator) is

$$\hat\beta_{\text{lasso}} = \arg\min_\beta\, \tfrac12\|y - X\beta\|^2 + \lambda\|\beta\|_1, \qquad \|\beta\|_1 = \sum_j|\beta_j|.$$
  • There is no general closed form. With orthonormal features ($X^\top X = I$) each coefficient is soft-thresholded: $\hat\beta_j = S_\lambda(z_j) = \mathrm{sign}(z_j)\max(|z_j| - \lambda, 0)$, where $z_j$ is the least-squares estimate. With one feature: $\hat\beta = S_\lambda(x^\top y)/x^\top x$.
  • In general it is solved by coordinate descent (soft-threshold one coefficient at a time) or proximal gradient (ISTA): see Optimization 3.13.
  • Coefficients can be exactly 0: the lasso does feature selection. With correlated features it tends to keep one and drop the others.
  • Scaling conventions differ. Some books write $\|y - X\beta\|^2 + \lambda\|\beta\|_1$ (no ½), which doubles the λ for the same answer. scikit-learn's Lasso(alpha) minimizes $\tfrac{1}{2n}\|y - X\beta\|^2 + \alpha\|\beta\|_1$, so $\alpha = \lambda/n$ in the convention above. Always check which one your code uses.
  • The elastic net mixes both penalties ($\lambda_1\|\beta\|_1 + \lambda_2\|\beta\|_2^2$) to get zeros and sensible sharing between correlated features.
Why do we need it?

When you suspect that most candidate features do nothing, you want a model that switches them off by itself, so it is simpler to read, cheaper to run and less likely to chase noise. Ridge cannot do this; the lasso can.

Where is it used?

Feature selection in linear and logistic models (scikit-learn Lasso, LassoCV, LogisticRegression with an L1 penalty (it needs a solver that supports L1, such as liblinear or saga; check your version's documentation for how the penalty is set), glmnet), compressed sensing, sparse signal recovery, and as the optimization twin of the Laplace prior on changepoint slopes in Prophet-style models.

How is it used?

Standardize the features, run LassoCV to choose λ, then list the non-zero coefficients. Treat the selected set with care: with correlated features a different sample may select a different twin. Refit or report uncertainty before claiming "feature 3 does not matter".

Orthonormal features, so each coefficient is handled on its own. The horizontal axis is the least-squares estimate $z$; the curves show what each method reports. Drag the blue handle to choose $z$, and move λ. Blue dashed: no penalty (report $z$). Teal: ridge divides, $z/(1 + \lambda)$: a straight line through 0 that only gets flatter. Orange: lasso soft-thresholds, $S_\lambda(z)$: flat at exactly 0 inside the shaded dead zone $|z| \le \lambda$, then parallel to the blue line, λ below it. Put $z$ inside the dead zone: only the lasso gives 0.

"Lasso and ridge differ only in how much they shrink."

They differ in shape. Ridge multiplies by a factor below 1 (never reaching 0). Lasso subtracts a constant and stops at 0, so it produces exact zeros.

"A coefficient the lasso set to 0 has no effect."

It means the evidence in this sample, at this λ, was below the toll. With correlated features the lasso may keep a twin instead; a slightly different sample may swap them. Zero means "not selected", not "proved useless".

"The lasso estimate of a selected coefficient is unbiased."

Survivors are shifted toward 0 by λ, so even large effects are underestimated. That is why people sometimes refit least squares on the selected features ("relaxed lasso").

$\hat\beta_{\text{lasso}} = \arg\min \tfrac12\|y - X\beta\|^2 + \lambda\|\beta\|_1$. Orthonormal case: $S_\lambda(z) = \mathrm{sign}(z)\max(|z| - \lambda, 0)$.

Dead zone $|z| \le \lambda$ → exactly 0 (selection); outside it, shrink by λ.

Trap: conventions differ (½ or not; scikit-learn's $\alpha = \lambda/n$); zero ≠ "proved useless".

Quick check: orthonormal features with least-squares estimates (3, −2, 0.5) and λ = 1. What does the lasso report?

Soft-threshold each: $3 \to 2$, $-2 \to -1$, $0.5 \to 0$ (because $|0.5| \le 1$). So $(2, -1, 0)$: the smallest effect is switched off and the other two move 1 unit toward 0.

The picture: why the L1 diamond gives zeros and the L2 circle does not core

A penalty can also be read as a budget: "find the best-fitting coefficients whose total size is at most $t$". With two coefficients, the allowed region is a shape around the origin. For L2 it is a circle ($\beta_1^2 + \beta_2^2 \le t^2$). For L1 it is a diamond ($|\beta_1| + |\beta_2| \le t$), a square standing on its corner, with its four corners exactly on the axes.

The loss (the misfit) has its lowest point at the unpenalized estimate, and equal-loss lines around it are ellipses. Inflate an ellipse from that point until it first touches the allowed region: the touching point is the answer. A diamond's corners poke out toward the ellipse, so the first touch is very often a corner, where one coefficient is exactly 0. A circle is round everywhere, so the first touch is almost never exactly on an axis.

Three ways to say it:

  • Picture: a balloon growing from the OLS point bumps into the diamond's sharp corner first.
  • Numbers: OLS (2, 0.5), budget 1.5: the diamond gives (1.5, 0); the circle gives about (1.46, 0.36).
  • Slogan: corners on the axes make zeros.

Orthonormal features, so the loss is $\tfrac12\|\beta - z\|^2$ with least-squares point $z = (2, 0.5)$: its equal-loss lines are circles around $z$.

  1. L1 budget $t = 1.5$. The lasso answer is $S_\lambda(z)$ for the λ that spends exactly the budget. With $\lambda \ge 0.5$ the second coordinate is 0 and the first is $2 - \lambda$; spending the budget needs $2 - \lambda = 1.5$, so $\lambda = 0.5$ and $\hat\beta = (1.5, 0)$: a corner of the diamond.
  2. L2 budget $t = 1.5$. The ridge answer points from 0 straight toward $z$: $\hat\beta = z/(1 + \lambda)$. Its length must be 1.5, and $\|z\| = \sqrt{4 + 0.25} \approx 2.062$, so $1 + \lambda = 2.062/1.5 \approx 1.374$ and $\hat\beta \approx (1.455, 0.364)$.
  3. Same budget size, same data: the diamond kills the small coefficient, the circle keeps it.

The penalized and budget (constrained) forms give the same answers:

$$\min_\beta\, \mathrm{Loss}(\beta) + \lambda\,\mathrm{Penalty}(\beta) \quad\Longleftrightarrow\quad \min_\beta\, \mathrm{Loss}(\beta)\ \text{ subject to }\ \mathrm{Penalty}(\beta) \le t.$$
  • Each λ corresponds to some budget $t$ (bigger λ ↔ smaller $t$). The exact match depends on the data, so there is no fixed formula linking them.
  • For least squares, $\mathrm{RSS}(\beta) = \mathrm{RSS}(\hat\beta_{\text{OLS}}) + (\beta - \hat\beta_{\text{OLS}})^\top X^\top X(\beta - \hat\beta_{\text{OLS}})$: equal-loss lines are ellipses centred at the OLS point, tilted when features are correlated.
  • L1 ball: a diamond (in more dimensions, a cross-polytope) with corners on the axes and flat faces. L2 ball: a circle (a sphere). The solution sits where the smallest loss ellipse touches the ball.
Why do we need it?

The algebra of soft-thresholding only covers orthonormal features. The picture explains sparsity for any correlated features, and shows when the lasso keeps a coefficient (the ellipse touches an edge) and when it drops it (the ellipse touches a corner).

Where is it used?

It is the standard explanation of lasso sparsity in textbooks (ESL, ISLR) and interviews; the same "corners make zeros" idea explains L1 in compressed sensing, sparse PCA, and why L∞ constraints make many values hit the limit together.

How is it used?

When you choose a penalty, picture its unit ball: corners on the axes → exact zeros (L1, elastic net); round → smooth shrinkage (L2). When an interviewer asks "why does the lasso give sparse solutions?", draw this picture.

Lasso: budget |β₁| + |β₂| ≤ t (a diamond) Ridge: budget β₁² + β₂² ≤ t² (a circle) β₁ β₂ OLS β₂ = 0 exactly β₁ β₂ OLS both ≠ 0 grey ellipses = equal loss around the OLS point · purple = where the growing ellipse first touches the budget
Grow the loss ellipse outward from the OLS point until it first touches the allowed region. The diamond's corners stick out along the axes, so the first touch is often a corner, where one coefficient is exactly 0. The circle has no corners, so the touching point almost never lies on an axis.

Drag the blue OLS point (the unpenalized estimate). Grey ellipses are equal-loss lines; the dark ellipse is the one that just touches the allowed region; the purple dot is the penalized answer; the dashed purple curve shows how the answer moves as λ goes from 0 to large. With L1, drag the OLS point around and notice how often the answer sits on a corner (one coordinate exactly 0), and how the path runs along an axis. With L2 the answer slides smoothly and lands on an axis only if the OLS point is already on it. Change the feature correlation to tilt the ellipses.

"The lasso always produces zeros."

Only when the touching point is a corner. If the ellipse touches a flat edge, both coefficients stay non-zero. Small λ (a big diamond) often keeps everything; the zeros appear as λ grows.

"Ridge can give an exact zero if λ is large enough."

Only in the limit λ → ∞, or by coincidence when the OLS coefficient was already 0. The circle has no corners on the axes to catch the ellipse.

Penalized form ⇔ budget form ($\mathrm{Penalty} \le t$). Loss contours are ellipses around the OLS point.

L1 ball = diamond with corners on the axes → first touch is often a corner → exact zeros. L2 ball = circle → no zeros.

Trap: zeros appear only when the touch is at a corner; not guaranteed.

Quick check: the OLS point is exactly on the β₁ axis, at (2, 0), and features are uncorrelated. Does ridge give β₂ = 0?

Yes, but only because the unpenalized β₂ was already 0: ridge scales the OLS point toward the origin along the axis. This is the coincidence case; if the OLS point moves even slightly off the axis, ridge's β₂ becomes non-zero, while the lasso's β₂ stays 0 for a whole range of positions.

Coefficient paths: watch every coefficient as λ grows

One λ gives one set of coefficients. A coefficient path shows them all: for every λ from tiny to huge, plot each coefficient. It is like a time-lapse of the penalty getting stronger.

The two pictures look very different. Ridge paths are smooth curves that glide toward 0 together and never touch it. Lasso paths are made of straight pieces; one by one, the weakest coefficients hit 0 and stay there, until only the strongest survive and finally they too reach 0. Reading the lasso path from right to left tells you the order in which features "enter" the model.

Three ways to say it:

  • Picture: ridge = curves fading toward 0; lasso = lines that drop to 0 one at a time.
  • Numbers: estimates (3, −2, 0.5): the lasso drops 0.5 at λ = 0.5, −2 at λ = 2, and 3 at λ = 3.
  • Slogan: the lasso path is a ranking of features; the ridge path is a dimmer switch.

Orthonormal features with least-squares estimates $z = (3, -2, 0.5)$.

  1. Ridge path: $z/(1 + \lambda)$. At λ = 1: $(1.5, -1, 0.25)$. At λ = 9: $(0.3, -0.2, 0.05)$. Never exactly 0.
  2. Lasso path: $S_\lambda(z)$. At λ = 0.5 the third coefficient hits 0. At λ = 1: $(2, -1, 0)$.
  3. At λ = 2 the second hits 0: $(1, 0, 0)$. At λ = 3 all are 0.
  4. Order of entry when λ decreases: feature 1 first (at λ = 3), then feature 2 (λ = 2), then feature 3 (λ = 0.5): the order of $|z_j|$.

The regularization path is the function $\lambda \mapsto \hat\beta_\lambda$ for all $\lambda \ge 0$, usually plotted against $\log\lambda$.

  • Ridge paths are smooth and reach 0 only as λ → ∞.
  • Lasso paths are piecewise linear in λ (computed exactly by the LARS algorithm, or approximately on a grid with warm-started coordinate descent). For the convention $\tfrac12\mathrm{RSS} + \lambda\|\beta\|_1$ with centred data, every coefficient is 0 once $\lambda \ge \lambda_{\max} = \max_j|x_j^\top y|$.
  • λ is then chosen from the path, usually by K-fold cross-validation (fit on K − 1 folds, measure error on the held-out fold, average, pick the λ with the lowest error, or the largest λ within one standard error of it).
Why do we need it?

λ is unknown in advance. The path shows the whole range of models at once: which features are robust (they survive large λ), which enter only when the penalty is weak, and how stable each coefficient is.

Where is it used?

sklearn.linear_model.lasso_path, enet_path, LassoLarsCV; glmnet's default plots; feature-importance reports in credit scoring and genomics; choosing how strong a prior should be in a Bayesian regression.

How is it used?

Fit the path on standardized features, plot it against log λ, overlay the cross-validated error, choose λ, and read off the selected features. Be suspicious of features that enter and leave the lasso path repeatedly: they are usually correlated with others.

Simulated daily demand with six standardized features: promo (true effect 3), price (−2), temperature (1.5), ads (0.6) and two pure-noise features (true effect 0). Top: ridge paths; bottom: lasso paths, both against log λ (each with its own λ range). Move the penalty strength slider and read the coefficients at the purple line. In the lasso panel the noise features usually die first and the true ones survive longest; in the ridge panel nothing ever reaches 0. Press New sample to see which parts of the picture are stable.

"The order in which the lasso drops features is their order of importance."

It is the order of their evidence in this sample, given the other features. Correlated features can swap places between samples; a useful but correlated feature can be dropped early because a twin already explains it.

"Pick λ where the training error is lowest."

Training error is always lowest at λ = 0. Choose λ with held-out data (cross-validation) or with a prior you can defend.

Path = $\hat\beta_\lambda$ for all λ (plot vs log λ). Ridge: smooth, never 0. Lasso: piecewise linear; coefficients hit 0 one by one; all 0 once $\lambda \ge \max_j|x_j^\top y|$.

Choose λ by cross-validation.

Trap: drop order ≠ true importance when features are correlated.

Quick check: on the lasso path, a coefficient is 0 for all λ above 4 and non-zero below. What does that tell you?

The feature "enters" the model at λ = 4: below that toll its evidence is strong enough to pay for a non-zero coefficient. With orthonormal features this would mean $|z_j| = 4$; with correlated features it is its evidence after the other active features are accounted for.

A Gaussian prior is ridge in disguise (at the MAP) core

So far λ was just a dial. Bayesian statistics gives it a meaning. Before seeing data, you might believe: "each coefficient is probably small, somewhere around 0, and very large values are rare". A bell curve centred at 0, the Normal distribution $N(0, \tau^2)$, says exactly that. This belief is the prior; $\tau$ (tau) is its spread: how big you think coefficients typically are.

Bayes' rule combines the prior with the likelihood: posterior ∝ likelihood × prior (Chapter 5.2; in depth in Chapter 6.1). The MAP (maximum a posteriori) estimate is the single θ where the posterior is highest. Taking minus the log turns "× prior" into "+ penalty", and minus the log of a bell curve is a parabola, $\beta^2/(2\tau^2)$: the ridge penalty. So "ridge regression" and "the most probable coefficients under a Normal prior" are the same calculation.

Three ways to say it:

  • Picture: a bell-shaped prior, flipped upside down by the minus log, is the ridge parabola.
  • Numbers: noise σ = 2 and prior spread τ = 1 give λ = σ²/τ² = 4.
  • Slogan: ridge = MAP with a Normal prior; λ = noise variance ÷ prior variance.

One unknown effect θ (for example a holiday uplift). Four observations with noise sd σ = 2, average $\bar y = 2$ (so $\sum y_i = 8$). Prior $\theta \sim N(0, \tau^2)$ with τ = 1.

  1. Ridge view: minimize $\sum_i (y_i - \theta)^2 + \lambda\theta^2$ with $\lambda = \sigma^2/\tau^2 = 4/1 = 4$. Derivative: $-2\sum(y_i - \theta) + 2\lambda\theta = 0$, so $\hat\theta = \sum y_i/(n + \lambda) = 8/(4 + 4) = 1$.
  2. Bayes view: the data part has precision (1/variance) $n/\sigma^2 = 4/4 = 1$; the prior has precision $1/\tau^2 = 1$. Posterior precision $= 1 + 1 = 2$, so the posterior variance is $1/2$ (sd ≈ 0.707).
  3. Posterior mean = precision-weighted average of the data mean and the prior mean 0: $(1 \times 2 + 1 \times 0)/2 = 1$.
  4. The posterior is Normal, $N(1, 0.5)$, so its peak (MAP), mean and median are all 1, the same as ridge.
  5. The difference: ridge reports only "1". The posterior also says "give or take 0.71".

Model: $y \mid \beta \sim N(X\beta, \sigma^2 I)$ (σ known) and prior $\beta_j \sim N(0, \tau^2)$, independently. Then

$$-\log p(\beta \mid y) = \frac{\|y - X\beta\|^2}{2\sigma^2} + \frac{\|\beta\|^2}{2\tau^2} + \text{const}.$$

Multiplying by $2\sigma^2$ does not move the minimum, so

$$\hat\beta_{\text{MAP}} = \arg\min_\beta\, \|y - X\beta\|^2 + \frac{\sigma^2}{\tau^2}\|\beta\|^2 = \hat\beta_{\text{ridge}}\Big(\lambda = \frac{\sigma^2}{\tau^2}\Big).$$
  • Noisy data (big σ) or a confident prior (small τ) → big λ → strong shrinkage. A very wide prior (τ → ∞) → λ → 0 → least squares, the MLE.
  • Special to the Gaussian case: the whole posterior is Normal, $\beta \mid y \sim N\big((X^\top X + \lambda I)^{-1}X^\top y,\ \sigma^2(X^\top X + \lambda I)^{-1}\big)$. Its mean, median and mode coincide, so here MAP = posterior mean = ridge. That coincidence fails for the Laplace prior.
  • Ridge returns only the centre. The Bayesian answer also has a covariance: uncertainty you can use for intervals and decisions.
Why do we need it?

It turns the arbitrary dial λ into a statement you can reason about ("effects are usually within ±2τ"), lets you set λ from prior knowledge or a hierarchical model instead of only cross-validation, and adds uncertainty to a ridge fit for free.

Where is it used?

Bayesian linear regression; Gaussian process regression; the Normal priors on Fourier seasonality, holiday and regressor coefficients in a Prophet-style model (Prophet's seasonality_prior_scale and holidays_prior_scale are such τ's); weight decay in neural networks, read as a Gaussian prior on the weights.

How is it used?

Write beta = numpyro.sample("beta", dist.Normal(0, tau)) (NumPyro takes the sd τ, not τ²) on standardized features. If you only need a point estimate with known σ, ridge with λ = σ²/τ² gives the same MAP; for intervals, use the full posterior.

likelihood × prior −log posterior (+ const) same minimum as y ~ N(Xβ, σ²) βⱼ ~ N(0, τ²) Normal prior ‖y − Xβ‖² / (2σ²) + ‖β‖² / (2τ²) ×2σ² RSS + λ‖β‖² (ridge) λ = σ² / τ² y ~ N(Xβ, σ²) βⱼ ~ Laplace(0, b) Laplace prior ‖y − Xβ‖² / (2σ²) + ‖β‖₁ / b ×σ² ½RSS + λ‖β‖₁ (lasso) λ = σ² / b Only the location of the lowest point is shared. The posterior itself is the whole curve e^(−…), not just its lowest point.
Minus the log of "likelihood × prior" is "data misfit + penalty". A Normal prior produces the squared (ridge) penalty, a Laplace prior the absolute-value (lasso) penalty. Multiplying by a positive constant changes λ's scale, not the minimizer.

Grey dashed: the prior $N(0, \tau^2)$. Blue: the likelihood of the data (peak at the data mean $\bar y$; drag the blue handle). Orange: the posterior. All three are scaled to the same height so you can compare shapes. The purple line marks the posterior peak (MAP), which here is also its mean and its median. Increase $n$: the likelihood narrows and the posterior moves toward $\bar y$ (λ matters less). Shrink τ: the prior narrows, λ = σ²/τ² grows and the answer is pulled toward 0. The readout checks that the ridge formula gives the same number.

"A Normal prior and ridge regression are the same thing."

They share one number: ridge's answer equals the posterior's peak (and, for the Normal case, its mean). The Bayesian model also gives a full distribution, with a spread you can use; ridge throws that away.

"λ = σ²/τ² holds whatever the scaling of the objective."

It holds for the objective RSS + λ‖β‖². If your library divides RSS by $n$ or by 2, λ changes by that factor. Write down the objective before converting.

"dist.Normal(0, tau) has variance τ."

NumPyro's Normal(loc, scale) takes the standard deviation. The variance is τ², which is the τ² in λ = σ²/τ².

Prior $\beta_j \sim N(0, \tau^2)$, noise $N(0, \sigma^2)$ ⇒ $\hat\beta_{\text{MAP}} = \hat\beta_{\text{ridge}}$ with $\lambda = \sigma^2/\tau^2$.

Gaussian case only: the posterior is Normal, so MAP = mean = median.

Trap: ridge gives the centre only; the posterior also gives the spread.

Quick check: σ = 3 and you believe effects are typically within ±1 (take τ = 0.5). What λ does that imply for RSS + λ‖β‖²?

$\lambda = \sigma^2/\tau^2 = 9/0.25 = 36$. Noisy data and a tight prior: strong shrinkage.

A Laplace prior is the lasso in disguise, but only at the MAP core

A different belief: "most coefficients are essentially zero, but a few may be large". The Laplace distribution (Chapter 4.9), $p(\beta) = \frac{1}{2b}e^{-|\beta|/b}$, says exactly that: a sharp peak at 0 plus tails heavier than a Normal's. Its scale $b$ sets how big a typical effect is (the average $|\beta|$ is $b$).

Minus its log is $|\beta|/b$: a V shape, which is the lasso penalty. The V has a sharp corner (a kink) at 0. Add the smooth data parabola to it, and the lowest point of the sum often sits right in that corner, exactly at 0. That is soft-thresholding seen from the Bayesian side.

Three ways to say it:

  • Picture: a smooth bowl (the data) plus a V (the prior); if the bowl's slope at 0 is gentler than the V's, the bottom is stuck in the V's corner.
  • Numbers: data mean 0.8 with s = 1 and b = 1: the data slope at 0 is 0.8, the V's corner is ±1, so the MAP is exactly 0.
  • Slogan: Laplace prior + "take the peak" = lasso.

Same setup as before: $n = 4$ observations, σ = 2, so the data mean $\bar y$ has standard error $s = \sigma/\sqrt n = 1$. Prior $\theta \sim Laplace(0, b = 1)$.

  1. Minus log posterior (dropping constants): $\dfrac{(\theta - \bar y)^2}{2s^2} + \dfrac{|\theta|}{b} = \dfrac{(\theta - \bar y)^2}{2} + |\theta|$.
  2. $\bar y = 2$: try $\theta \gt 0$. Derivative $(\theta - 2) + 1 = 0$ gives $\theta = 1$, which is positive, so MAP = 1 $= \bar y - s^2/b$.
  3. $\bar y = 0.8$: try $\theta \gt 0$: $\theta = 0.8 - 1 = -0.2$, not positive, contradiction. Try $\theta \lt 0$: $(\theta - 0.8) - 1 = 0$ gives $\theta = 1.8$, not negative, contradiction. So the minimum is at the kink: MAP = 0 exactly.
  4. The rule: MAP $= S_{s^2/b}(\bar y)$, soft-thresholding at $s^2/b = 1$.
  5. Lasso check: minimize $\tfrac12\sum_i(y_i - \theta)^2 + \lambda|\theta|$ with $\lambda = \sigma^2/b = 4$. Its threshold is $\lambda/n = 4/4 = 1 = s^2/b$. Same answer.

Model: $y \mid \beta \sim N(X\beta, \sigma^2 I)$ and prior $\beta_j \sim Laplace(0, b)$, independently. Then

$$-\log p(\beta\mid y) = \frac{\|y - X\beta\|^2}{2\sigma^2} + \frac{\|\beta\|_1}{b} + \text{const} \quad\Longrightarrow\quad \hat\beta_{\text{MAP}} = \arg\min_\beta\, \tfrac12\|y - X\beta\|^2 + \frac{\sigma^2}{b}\|\beta\|_1.$$

So the MAP is the lasso with $\lambda = \sigma^2/b$ in the $\tfrac12$RSS convention. In other conventions the constant changes, which is why the rule is usually quoted as $\lambda \propto \sigma^2/b$:

Objectiveλ that matches the Laplace(0, b) MAP
$\tfrac12\|y - X\beta\|^2 + \lambda\|\beta\|_1$ (ESL, this chapter)$\lambda = \sigma^2/b$
$\|y - X\beta\|^2 + \lambda\|\beta\|_1$$\lambda = 2\sigma^2/b$
scikit-learn Lasso: $\tfrac{1}{2n}\|y - X\beta\|^2 + \alpha\|\beta\|_1$$\alpha = \sigma^2/(n\,b)$
  • One coefficient: MAP $= S_{s^2/b}(\bar y) = \mathrm{sign}(\bar y)\max(|\bar y| - s^2/b,\ 0)$ with $s^2 = \sigma^2/n$. It is exactly 0 whenever $|\bar y| \le s^2/b$.
  • Small $b$ (confident that effects are tiny) or noisy data → large threshold → more exact zeros at the MAP.
  • This equivalence is about the location of the posterior's highest point only. The next section shows what the rest of the posterior looks like.
Why do we need it?

It explains where the lasso's λ "comes from" (noise variance over prior scale), tells you how to translate between a Bayesian prior scale and a frequentist penalty, and sets up the precise statement about sparsity that interviewers probe.

Where is it used?

The changepoint slope changes in Prophet-style trends ($\delta_j \sim Laplace(0, b)$, Prophet's changepoint_prior_scale is that $b$), the "Bayesian lasso", sparse Bayesian regression, and any MAP fit (for example NumPyro with an AutoDelta guide) of a model with Laplace priors.

How is it used?

In NumPyro: dist.Laplace(0.0, b) (loc, scale = b). If you fit by MAP, expect lasso-like behaviour: many coefficients pushed to (or very near) 0. If you fit the full posterior (NUTS, or SVI with a Normal-family guide), expect shrinkage toward 0 but no exact zeros.

Normal(0, τ²) prior density p(θ) −log p(θ) + constant θ²/(2τ²): a parabola → ridge (L2) -2 2 Laplace(0, b) prior density p(θ) −log p(θ) + constant |θ|/b: a V with a kink → lasso (L1) -2 2
Top: a Normal and a Laplace prior with the same variance. The Laplace is sharper at 0 and has heavier tails. Bottom: minus their logs. The Normal gives a smooth parabola (the ridge penalty); the Laplace gives a V with a sharp corner at 0 (the lasso penalty). That corner is what lets the MAP sit exactly at 0.

Blue: the data part of minus the log posterior, $(\theta - \bar y)^2/(2s^2)$, a bowl centred at $\bar y$ (drag the blue handle). Grey: the prior part $|\theta|/b$, a V with its corner at 0. Orange: their sum; the purple dot is its lowest point, the MAP. Start at $\bar y = 0.8$: the bottom is stuck in the corner at exactly 0. Drag $\bar y$ past the threshold $s^2/b$ and the bottom leaves the corner and slides away, always $s^2/b$ behind $\bar y$. Make $b$ small (a sharper V) and the corner holds the MAP for a wider range of data.

"A Laplace prior is L1 regularization."

Taking the peak of the posterior under a Laplace prior (with a Gaussian likelihood) solves the lasso problem. The prior itself is a distribution, and the posterior it produces is a distribution, not a penalized point estimate. The equivalence lives at the MAP only.

"The Laplace prior puts extra probability exactly at 0."

It is a continuous density: $P(\beta = 0) = 0$. It only has a high, sharp peak near 0. The exact zeros of the MAP come from the kink in $-\log p$, not from any point mass.

"Same variance means same shrinkage."

A Normal and a Laplace with equal variance shrink very differently: the Normal shrinks every effect by the same fraction; the Laplace shrinks small effects a lot (to 0 at the MAP) and large effects relatively little.

Prior $\beta_j \sim Laplace(0, b)$ ⇒ $-\log p = |\beta|/b$ ⇒ $\hat\beta_{\text{MAP}} = $ lasso with $\lambda = \sigma^2/b$ (½RSS form); in general $\lambda \propto \sigma^2/b$.

One coefficient: MAP $= S_{s^2/b}(\bar y)$, exactly 0 when $|\bar y| \le s^2/b$.

Trap: the equivalence is for the MAP only; the zeros come from the kink, not from a point mass.

Quick check: σ = 2, n = 16, b = 0.5. How large must $|\bar y|$ be for the MAP to be non-zero?

$s^2 = \sigma^2/n = 4/16 = 0.25$, threshold $s^2/b = 0.25/0.5 = 0.5$. The MAP is non-zero only when $|\bar y| \gt 0.5$, and then it equals $\bar y \mp 0.5$.

The precise distinction: the MAP is sparse, the full posterior is not core

The posterior is a whole landscape: for every value of θ it says how plausible that value is after seeing the data. The MAP is only the highest point of that landscape. With a Laplace prior, the highest point is often exactly at 0, because the landscape has a sharp peak there (the kink). But a peak is a single point, and a single point holds no area. The posterior's probability is spread over the ground on both sides of the peak.

So the same model gives two different stories. The MAP says "θ = 0: switch this effect off". The full posterior says "θ is most likely small, probably positive, and could be anywhere from about −0.7 to 1.7". Neither is wrong; they answer different questions. What is wrong is saying "the Bayesian model with a Laplace prior is sparse". Only its MAP is.

Three ways to say it:

  • Picture: a tent whose pole stands exactly at 0; the pole's position is 0, but the canvas covers ground on both sides.
  • Numbers: data mean 0.8, s = 1, b = 1: MAP = 0 exactly, yet the posterior mean is 0.39, the median 0.33, $P(\theta \gt 0 \mid D) = 0.70$ and $P(\theta = 0 \mid D) = 0$.
  • Slogan: the Laplace prior makes the MAP sparse, not the posterior.

Same numbers: $\bar y = 0.8$, $s = 1$, prior $Laplace(0, 1)$. The posterior density is $p(\theta\mid D) \propto \exp\!\big(-(\theta - 0.8)^2/2 - |\theta|\big)$.

  1. MAP: soft-thresholding at $s^2/b = 1$; $|0.8| \le 1$, so the MAP is exactly 0 (previous section).
  2. Normalize: evaluate the formula on a fine grid of θ values, add up (integrate), and divide so that the total area is 1. (Or use the exact formula in the definition below.)
  3. Posterior mean $= \int\theta\,p(\theta\mid D)\,d\theta \approx 0.394$. Posterior median (the θ with half the area on each side) $\approx 0.326$. Neither is 0.
  4. Sign: area to the right of 0 is $P(\theta \gt 0\mid D) \approx 0.703$; to the left $\approx 0.297$.
  5. Exactly zero: $P(\theta = 0\mid D) = 0$, because a density gives zero area to a single point. Even the small window $|\theta| \lt 0.05$ holds only about 6.4% of the probability.
  6. Spread: 90% of the posterior lies between about −0.71 and 1.68.
  7. Repeat with $\bar y = 2$: MAP $= 1$, posterior mean $\approx 1.16$, median $\approx 1.11$. With a strong signal the summaries move close together; with a weak signal they disagree most.

Normal likelihood for one coefficient ($\bar y$ with standard error $s$) and a $Laplace(0, b)$ prior give the posterior

$$p(\theta \mid D) \propto \exp\!\Big(-\frac{(\theta - \bar y)^2}{2s^2} - \frac{|\theta|}{b}\Big),$$

which is a mixture of two Normals cut at 0: on $\theta \gt 0$ a $N(\bar y - s^2/b,\ s^2)$ piece with weight $\propto e^{-\bar y/b}\,\Phi\big((\bar y - s^2/b)/s\big)$, and on $\theta \lt 0$ a $N(\bar y + s^2/b,\ s^2)$ piece with weight $\propto e^{\bar y/b}\,\Phi\big(-(\bar y + s^2/b)/s\big)$ ($\Phi$ = standard Normal CDF).

  • MAP (posterior mode) $= S_{s^2/b}(\bar y)$: exactly 0 whenever $|\bar y| \le s^2/b$.
  • No point mass: the posterior is a continuous density, so $P(\theta = 0 \mid D) = 0$. Posterior draws (MCMC, or SVI with a continuous guide) are never exactly 0.
  • Posterior mean and median have the sign of $\bar y$ and are non-zero whenever $\bar y \ne 0$. They shrink toward 0 smoothly, without a dead zone.
  • The full posterior under a Laplace prior is sometimes called the Bayesian lasso (Park and Casella, 2008). It shrinks; it does not select.
  • Exact zeros in the posterior require a prior that itself puts probability on exactly 0, such as a spike-and-slab prior. Continuous "shrinkage priors" such as the horseshoe concentrate more mass near 0 than the Laplace but still give no exact zeros.
  • Remember from Chapter 5.2: the MAP also depends on the parameterization, while posterior probabilities do not.
Why do we need it?

Reports built on posterior draws (means, intervals, probabilities) and reports built on a MAP fit can disagree about which effects are "zero". Knowing which one you computed prevents wrong claims such as "the Bayesian model found that 20 of 25 changepoints are inactive".

Where is it used?

Prophet (which by default reports a MAP fit, and runs MCMC only when you ask for samples), Prophet-style models fitted with NUTS or SVI in NumPyro, Bayesian lasso regressions, and A/B lift estimates with sparsity-inducing priors.

How is it used?

State which summary you report. With posterior draws, decide "active or not" by a rule you choose (for example, a 90% interval that excludes 0, or $P(|\delta_j| \gt c\mid D) \gt 0.9$ for a practically meaningful $c$), not by looking for zeros. If you need posterior probabilities of exactly 0, use a spike-and-slab prior.

-2 -1 0 1 2 3 mean 0.39 peak (MAP) at 0 P(θ < 0) = 0.30 P(θ > 0) = 0.70 90% of the posterior: −0.71 to 1.68 What the full posterior says P(θ = 0 exactly) = 0: a single point has no area What the MAP reports -2 -1 0 1 2 3 θ̂_MAP = 0 exactly one number, no spread: "this coefficient is switched off"
Same data, same prior (ȳ = 0.8, s = 1, Laplace prior with b = 1). Left: the MAP is a single number, exactly 0. Right: the full posterior. Its highest point is indeed at 0, but its area spreads over both sides: 70% of the probability says the effect is positive, the mean is 0.39, and the probability of "exactly 0" is zero.

Top: prior (grey dashed), likelihood (blue; drag the handle to change the data mean $\bar y$) and posterior (orange; lighter left of 0, darker right of 0), each scaled to the same height. Purple line: the MAP. Pink dashed: the posterior mean. Teal dashed: the posterior median. Bottom: what each summary reports for every possible $\bar y$; the purple band is the MAP's dead zone. Start at $\bar y = 0.8$: the MAP is pinned at 0 while the mean and median are clearly positive. Drag $\bar y$ slowly from 0 to 3: the purple MAP curve is flat at 0 and then bends away, but the pink and teal curves leave 0 immediately. Press Tight prior to widen the dead zone.

4,000 draws from the posterior (s = 1, b = 1), as an MCMC sampler would give you. With the Laplace prior, look at the readout: even when the MAP is exactly 0, none of the draws is exactly 0; they pile up near 0 and spread to both sides. Switch to the spike-and-slab prior (half its prior mass exactly at 0, half spread as a Normal with the same variance as the Laplace): now a real share of draws is exactly 0 (the purple spike), and that share is the posterior probability that the effect is truly absent. Move $\bar y$ to see both respond to the data.

"With a Laplace prior, the posterior mean of a weak effect is exactly 0, like the lasso."

The lasso equals the posterior mode. The posterior mean and median are non-zero whenever the data mean is non-zero (0.39 and 0.33 in the example). They shrink, but they do not snap to 0.

"My NUTS run shows most δⱼ near 0, so the model selected a few changepoints."

The posterior says most δⱼ are small. "Selected" needs a decision rule you choose (an interval excluding 0, or a threshold on $P(|\delta_j| \gt c\mid D)$). The Laplace prior did not make that decision for you.

"MAP by gradient descent gives exact zeros."

The exact MAP has exact zeros, but plain gradient methods (Adam in SVI with an AutoDelta guide, L-BFGS) step across the kink and end up hovering near 0 (for example 0.001). Coordinate descent and proximal methods (soft-thresholding steps) land exactly on 0.

"MAP and posterior mean are two equally good point estimates; just pick one."

They answer different questions: the MAP is the single most probable value; the posterior mean minimizes expected squared error. And only the posterior (not the MAP) tells you the uncertainty.

In your forecasting model the slope changes have $\delta_j \sim Laplace(0, b)$. If you fit it with SVI using a Normal-family guide (mean-field, full-rank or low-rank Gaussian), the approximate posterior of each δⱼ is a Gaussian piece: its draws are never exactly 0, and its mean is small but non-zero. If you fit a MAP (an AutoDelta guide, or Prophet's default optimizer), many δⱼ are pushed to, or very near, 0. In the A/B framework the same logic applies to any lift or segment effect with a sparsity-type prior: report $P(\text{lift} \gt \delta \mid D)$ and intervals from the posterior, not "the effect is exactly zero".

"A Laplace prior is the same as L1 regularization, so the Bayesian model is sparse."

"The L1 equivalence holds at the MAP. The full posterior under a Laplace prior has no point mass at 0, so it is shrunk but not sparse."

Model answer: "With a Gaussian likelihood and a Laplace(0, b) prior, minus the log posterior is the squared error over 2σ² plus the L1 norm over b, so the posterior mode solves the lasso with λ = σ²/b, up to the scaling convention. That mode can be exactly zero because of the kink. But the posterior is a continuous density: the probability of exactly zero is zero, and the posterior mean and median are non-zero whenever the data push away from zero. So my changepoint slopes are 'sparse-ish': most are small, none is exactly zero in posterior draws. If I needed true posterior sparsity I would use a spike-and-slab prior."

Laplace prior: MAP $= S_{s^2/b}(\bar y)$ (exactly 0 in the dead zone) but $P(\theta = 0\mid D) = 0$; mean, median ≠ 0 when $\bar y \ne 0$.

Example ($\bar y = 0.8, s = b = 1$): MAP 0; mean 0.39; median 0.33; $P(\theta \gt 0\mid D) = 0.70$.

Trap: "Laplace prior ⇒ sparse posterior" is false. Exact posterior zeros need a point mass (spike-and-slab).

Quick check: a colleague runs NUTS on a model with Laplace priors and finds no posterior draw exactly equal to 0. Is the sampler broken?

No. The posterior is a continuous density, so the probability of drawing exactly 0 is zero. Only the MAP can sit exactly at 0 (because of the kink). The sampler is behaving correctly.

Preview: Laplace priors on changepoint slopes give "sparse-ish" trends

A Prophet-style trend is a chain of straight pieces. You place many candidate changepoints (dates where the slope is allowed to change), and give each one a slope change $\delta_j$. You believe that most candidates are not real changes (δⱼ ≈ 0) and that a few are (δⱼ large). That is precisely the "many tiny, a few big" belief of the Laplace prior: $\delta_j \sim Laplace(0, b)$.

Compared with a Normal prior of the same variance, the Laplace prior produces more near-zero changes and more large ones: trends with long straight stretches and a few clear bends, instead of many medium wiggles. But it is "sparse-ish": no prior draw and no posterior draw is exactly 0. Exact zeros appear only if you report the MAP.

Three ways to say it:

  • Picture: a ruler with a few sharp bends, not a gently wobbling ruler.
  • Numbers: with 25 candidates, on average about 10 Laplace draws are tiny (|δ| < b/2) and about 1.2 are large (|δ| > 3b); a Normal with the same variance gives about 7 and 0.85.
  • Slogan: many small slope changes, a few big ones, none exactly zero (except at the MAP).

25 candidate changepoints with $\delta_j \sim Laplace(0, b)$, $b = 0.05$, compared with $\delta_j \sim N(0, 0.0707^2)$, which has the same variance $2b^2 = 0.005$.

  1. Laplace: $P(|\delta| \lt t) = 1 - e^{-t/b}$. Tiny changes, $|\delta| \lt b/2 = 0.025$: $1 - e^{-0.5} \approx 0.393$, so about $25 \times 0.393 \approx 9.8$ of the 25.
  2. Large changes, $|\delta| \gt 3b = 0.15$: $e^{-3} \approx 0.050$, about $25 \times 0.050 \approx 1.2$ of the 25.
  3. Normal with sd 0.0707: $P(|\delta| \lt 0.025) = 2\Phi(0.354) - 1 \approx 0.276$ (about 6.9 tiny) and $P(|\delta| \gt 0.15) = 2(1 - \Phi(2.12)) \approx 0.034$ (about 0.85 large).
  4. So the Laplace prior favours "mostly flat, occasionally a big bend". The average size of a change is $E|\delta| = b = 0.05$.

Prophet-style piecewise-linear trend with changepoints $s_1 \lt \dots \lt s_S$ (full treatment in Chapters 7.8–7.10):

$$g(t) = \Big(k + \sum_{j:\, s_j \le t}\delta_j\Big)\,t + \Big(m + \sum_{j:\, s_j \le t}\gamma_j\Big), \qquad \gamma_j = -s_j\delta_j, \qquad \delta_j \sim Laplace(0, b).$$
  • $k$ is the starting slope, $m$ the offset; $\delta_j$ is the change in slope at $s_j$; $\gamma_j$ keeps the line connected.
  • Small $b$: strong shrinkage, a stiff trend (risk of underfitting real changes). Large $b$: weak shrinkage, a trend that bends to chase noise (overfitting).
  • At the MAP, many δⱼ are exactly 0 (lasso behaviour). In the full posterior none is; most concentrate near 0, and the posterior may spread one real change over neighbouring candidates when it is unsure exactly where the bend is.
Why do we need it?

We do not know where the trend changes, so we offer many candidates. Without shrinkage the fit would bend at every candidate and chase noise; with a Laplace prior it bends only where the data insist.

Where is it used?

Prophet (changepoint_prior_scale is $b$, applied after Prophet scales the series), Prophet-style NumPyro models such as your forecasting model, and other trend-filtering and changepoint models with L1-type penalties.

How is it used?

Put candidates on a grid (and/or from a detector such as PELT), give each δⱼ a dist.Laplace(0.0, b) prior on a scaled series, tune $b$ by checking holdout forecasts and prior predictive trends, and report changes with posterior intervals for δⱼ or for the slope itself.

0.4 0 −0.5 −1 −1.5 −2 0 s1 0 s2 0 s3 s4 s5 0 s6 0 s7 0 s8 true change −2 at s5 Slope changes δⱼ at 8 candidate changepoints (one simulated series) MAP (6 of 8 exactly 0) posterior mean and 90% interval (NUTS, 4000 draws; 0 draws exactly 0)
A real fit (60 noisy points, true slope change −2 at candidate 5, Laplace prior with b = 0.1 on every δⱼ, noise sd known). The MAP switches 6 of the 8 candidates off exactly. The full posterior (NUTS) keeps every δⱼ small but non-zero, and it is unsure whether the bend happens at s4, s5 or s6, so it shares the change between neighbours. Same model, same prior: the zeros belong to the MAP, not to the posterior.

Each press of New draws draws 25 slope changes from each prior and builds the trend they imply (top: starting slope 1; orange = Laplace prior, teal = Normal prior with the same variance). Bottom: the 25 changes, in units of $b$ (orange and teal bars side by side, dashed lines at ±b/2 and ±3b). Press it several times. The orange trend usually has long straight stretches and one or two sharp bends; the teal one wobbles more evenly. Look at the readout: no draw is ever exactly 0, even in the prior. Moving $b$ only rescales the trend; the shape of the prior is the same for every $b$.

"The Laplace prior tells the model that most changepoints are exactly zero."

It says most are probably small. It puts zero probability on exactly zero. Exact zeros come only from taking the MAP.

"A bigger $b$ is always safer because it lets the data speak."

A bigger $b$ lets the trend bend at every candidate to follow noise, which hurts forecasts (the extrapolated slope becomes unstable). $b$ is a bias–variance knob: tune it on held-out forecasts.

"If the posterior mean of δⱼ is small, nothing happened at sⱼ."

When a real bend falls between candidates, the posterior often shares it among neighbours (as in the figure: s4, s5, s6). Look at the slope over a window, not only at one δⱼ.

In your forecasting model, the candidates come from a Prophet-like grid plus PELT detection, and each slope change has $\delta_j \sim Laplace(0, b)$. The prior encourages most adjustments to stay small, giving a sparse-ish changepoint representation. With your SVI fit (a full-rank or low-rank Gaussian guide), every δⱼ in the approximate posterior is a continuous distribution: none is exactly 0, and the honest summary of "where did the trend change?" is the posterior of the slope over time, with intervals. If someone compares your model with Prophet's default output, remember that Prophet reports a MAP fit, where many δⱼ can come out as (near) zero. Sensitivity to $b$ is covered in Chapter 7.10.

"The Laplace prior makes the changepoints sparse."

"The Laplace prior makes the changepoint slopes sparse-ish: it shrinks most of them toward zero and allows a few large ones. Exact sparsity happens only at the MAP."

Model answer: "δⱼ ~ Laplace(0, b) has a sharp peak at zero and heavier tails than a Normal, so most slope changes are shrunk to tiny values and only well-supported changes stay large. Fitting by MAP is like a lasso, so some δⱼ are exactly zero; fitting the posterior with SVI or NUTS gives no exact zeros, just concentration near zero. b controls the trade-off between a stiff trend and an overfit one."

$\delta_j \sim Laplace(0, b)$: many tiny changes, a few big ones; $E|\delta_j| = b$.

MAP: some δⱼ exactly 0. Posterior (SVI/NUTS): none exactly 0; real changes may be shared among neighbours.

Trap: "sparse" is true only for the MAP. $b$ is a bias–variance knob.

Quick check: with b = 0.05, what fraction of Laplace prior draws exceed 0.1 in absolute value?

$P(|\delta| \gt 0.1) = e^{-0.1/0.05} = e^{-2} \approx 0.135$, so about 13.5% (about 3.4 of 25 candidates).

Recap, cheat sheet and practice

  • Regularization = loss + λ × penalty. It trades a little bias for a big drop in variance. Standardize features; do not penalize the intercept; choose λ by cross-validation or a prior.
  • Ridge (L2): $(X^\top X + \lambda I)^{-1}X^\top y$. Shrinks the coefficients smoothly toward 0, never exactly to 0, and stabilizes collinear features.
  • Lasso (L1): soft-thresholding $S_\lambda(z)$ in the orthonormal case. Exact zeros (selection) plus shrinkage of the survivors. Diamond corners on the axes explain the zeros; paths show features dropping out one by one.
  • Gaussian prior ↔ ridge at the MAP, λ = σ²/τ². Here the posterior is Normal, so MAP = mean = median.
  • Laplace prior ↔ lasso at the MAP, λ = σ²/b (½RSS form; λ ∝ σ²/b in general).
  • The precise distinction: under a Laplace prior the MAP can be exactly 0, but the full posterior is a continuous density: $P(\theta = 0\mid D) = 0$, and the posterior mean and median are non-zero whenever the data are. Exact posterior zeros need a point-mass (spike-and-slab) prior.
  • Changepoints: $\delta_j \sim Laplace(0, b)$ gives sparse-ish trends: most slope changes tiny, a few large, none exactly 0 in posterior draws.

Cheat sheet

IdeaFormulaMeaning / trap
Regularized estimate$\arg\min\, \mathrm{Loss} + \lambda\,\mathrm{Penalty}$λ = 0: unpenalized; λ → ∞: all → 0
Ridge$(X^\top X + \lambda I)^{-1}X^\top y$; orthonormal: $z/(1 + \lambda)$shrinks, never 0; splits collinear twins
Lasso$\tfrac12\mathrm{RSS} + \lambda\|\beta\|_1$; orthonormal: $S_\lambda(z)$exact zeros; scikit-learn $\alpha = \lambda/n$
Soft-thresholding$S_\lambda(z) = \mathrm{sign}(z)\max(|z| - \lambda, 0)$dead zone $|z| \le \lambda$
Geometrybudget $\|\beta\|_1 \le t$ vs $\|\beta\|_2 \le t$diamond corners on axes → zeros
Gaussian prior → ridge$\lambda = \sigma^2/\tau^2$posterior Normal: MAP = mean
Laplace prior → lasso$\lambda = \sigma^2/b$ (½RSS form)MAP only
One-coefficient MAP, Laplace$S_{s^2/b}(\bar y)$, $s^2 = \sigma^2/n$exactly 0 if $|\bar y| \le s^2/b$
Full posterior, Laplace$\propto e^{-(\theta - \bar y)^2/2s^2 - |\theta|/b}$$P(\theta = 0\mid D) = 0$; mean, median ≠ 0
Changepoint prior$\delta_j \sim Laplace(0, b)$, $E|\delta_j| = b$sparse-ish; b = bias–variance knob
Code it · Python

import numpy as np
from scipy import optimize, integrate
from sklearn.linear_model import Ridge, Lasso

# 1) Ridge: minimize ||y - Xb||^2 + lam * ||b||^2  ->  b = (X'X + lam I)^-1 X'y
rng = np.random.default_rng(0)
n, p = 50, 5
X = rng.normal(size=(n, p))
X[:, 1] = X[:, 0] + 0.01 * rng.normal(size=n)          # feature 1 is a near-copy of feature 0
beta_true = np.array([2.0, 0.0, -1.0, 0.0, 0.5])
y = X @ beta_true + rng.normal(size=n)
lam = 10.0
b_ridge = np.linalg.solve(X.T @ X + lam * np.eye(p), X.T @ y)
print(np.allclose(b_ridge, Ridge(alpha=lam, fit_intercept=False).fit(X, y).coef_))   # True: same objective
b_ols = np.linalg.lstsq(X, y, rcond=None)[0]
print("OLS  ", b_ols.round(2))       # [-7.04  9.12 -0.84  0.01  0.57]: the twins get huge opposite-sign weights
print("ridge", b_ridge.round(2))     # [ 0.9   0.9  -0.71 -0.02  0.53]: the twins share; nothing is exactly 0

# 2) Lasso. Textbook: 1/2 ||y - Xb||^2 + lam * ||b||_1.  scikit-learn: 1/(2n) ||y - Xb||^2 + alpha * ||b||_1
lam = 20.0
b_lasso = Lasso(alpha=lam / n, fit_intercept=False).fit(X, y).coef_   # so alpha = lam / n
print("lasso", b_lasso.round(2), int((b_lasso == 0).sum()), "exact zeros")   # [ 0.  1.49 -0.46 -0.  0.34] 2: keeps one twin

# 3) One feature: ridge shrinks, lasso soft-thresholds
x1, y1 = np.array([1.0, 2, 3]), np.array([1.0, 3, 2])
sxx, sxy = x1 @ x1, x1 @ y1                                        # 14, 13
for lam in [0, 7, 13]:
    ridge = sxy / (sxx + lam)
    lasso = np.sign(sxy) * max(abs(sxy) - lam, 0) / sxx
    print(lam, round(ridge, 3), round(lasso, 3))                   # 0: 0.929 0.929 | 7: 0.619 0.429 | 13: 0.481 0.0

# 4) MAP vs FULL posterior under a Laplace prior. Data: n = 4, noise sd sigma = 2, prior Laplace(0, b = 1)
y4 = np.array([2.9, -1.6, 0.5, 1.4]); sigma, b = 2.0, 1.0
ybar, s = y4.mean(), sigma / np.sqrt(len(y4))                     # 0.8, 1.0
neg_log_post = lambda t: np.sum((y4 - t) ** 2) / (2 * sigma**2) + abs(t) / b
map_numeric = optimize.minimize_scalar(neg_log_post, bounds=(-5, 5), method="bounded", options={"xatol": 1e-10}).x
map_formula = np.sign(ybar) * max(abs(ybar) - s**2 / b, 0)        # soft-thresholding at s^2 / b = 1
print("MAP:", round(map_numeric, 6), map_formula)                 # 0.0 (to optimizer tolerance) and exactly 0.0

t = np.linspace(-8, 8, 160001)
log_post = -(t - ybar) ** 2 / (2 * s**2) - np.abs(t) / b
post = np.exp(log_post - log_post.max()); post /= np.trapezoid(post, t)
cdf = integrate.cumulative_trapezoid(post, t, initial=0)
print("posterior mean  ", round(np.trapezoid(t * post, t), 3))    # 0.394
print("posterior median", round(np.interp(0.5, cdf, t), 3))       # 0.326
print("P(theta > 0 | D)", round(1 - np.interp(0, t, cdf), 3))     # 0.703
print("90% interval    ", np.interp([0.05, 0.95], cdf, t).round(2))   # [-0.71  1.68]
print("P(|theta| < 0.05 | D)", round(np.interp(0.05, t, cdf) - np.interp(-0.05, t, cdf), 3))   # 0.064; P(theta = 0) is exactly 0

# 5) The same in NumPyro: NUTS draws are never exactly 0; MAP by gradient steps (AutoDelta) only gets near 0
import jax, numpyro, numpyro.distributions as dist
from numpyro.infer import MCMC, NUTS, SVI, Trace_ELBO
from numpyro.infer.autoguide import AutoDelta

def model(y):
    theta = numpyro.sample("theta", dist.Laplace(0.0, b))         # Laplace(loc, scale = b)
    numpyro.sample("y", dist.Normal(theta, sigma), obs=y)          # Normal(loc, scale = sd)

mcmc = MCMC(NUTS(model), num_warmup=500, num_samples=2000, progress_bar=False)
mcmc.run(jax.random.PRNGKey(0), y=y4)
draws = np.asarray(mcmc.get_samples()["theta"])
print("NUTS mean", draws.mean().round(2), "exact zeros:", int((draws == 0).sum()), "of", draws.size)   # about 0.38-0.39; 0 of 2000

svi = SVI(model, AutoDelta(model), numpyro.optim.Adam(0.01), Trace_ELBO())
res = svi.run(jax.random.PRNGKey(1), 3000, y=y4, progress_bar=False)
print("AutoDelta MAP:", round(float(res.params["theta_auto_loc"]), 4))   # about -0.0013: near 0, not exactly 0 (Adam hovers at the kink)
Test yourself

1. Orthonormal features. A coefficient's least-squares estimate is 0.6 and λ = 1. What do ridge ($z/(1 + \lambda)$) and the lasso ($S_\lambda(z)$) report?

Ridge divides: $0.6/2 = 0.3$. The lasso subtracts λ and stops at 0: $|0.6| \le 1$, so it is exactly 0.

2. Noise variance σ² = 4 and prior $\beta_j \sim N(0, 0.25)$ (variance 0.25). Which λ in RSS + λ‖β‖² gives the same answer as the MAP?

$\lambda = \sigma^2/\tau^2 = 4/0.25 = 16$.

3. A single coefficient has a Laplace(0, b) prior and a Normal likelihood. The data mean is $\bar y = 0.5 \ne 0$ and the MAP is exactly 0. Which statement is true?

The posterior is a continuous density with its peak at the kink. A single point has no probability, and more area lies on the side of $\bar y$, so the mean and median are positive.

4. Why does the lasso produce exact zeros while ridge does not?

Exact zeros come from the non-differentiable corner of |β| at 0 (the diamond's corners). The L2 penalty is smooth and round, so the optimum lands on an axis only by coincidence.

5. Your notes use $\tfrac12\|y - X\beta\|^2 + \lambda\|\beta\|_1$ with λ = 50 and n = 100 rows. What alpha gives the same fit in scikit-learn's Lasso?

scikit-learn minimizes $\tfrac{1}{2n}\|y - X\beta\|^2 + \alpha\|\beta\|_1$; dividing your objective by $n$ gives $\alpha = \lambda/n = 50/100 = 0.5$.

6. You fit your forecasting model with NUTS, with $\delta_j \sim Laplace(0, b)$ on 25 candidate changepoints. What fraction of the posterior draws of δⱼ will be exactly 0?

The posterior is continuous, so exact zeros have probability 0. Most draws will be near 0 for inactive candidates, but none exactly 0. Only a MAP (or a spike-and-slab prior) gives exact zeros.

Practice problems

A. One feature, no intercept: $x = (1, 2, 2)$, $y = (2, 3, 5)$. Compute the least-squares, ridge (λ = 4.5, RSS + λβ²) and lasso (λ = 4.5, ½RSS + λ|β|) estimates. For which λ does the lasso give exactly 0?

$\sum x_iy_i = 2 + 6 + 10 = 18$, $\sum x_i^2 = 1 + 4 + 4 = 9$. Least squares: $18/9 = 2$. Ridge: $18/(9 + 4.5) = 18/13.5 \approx 1.333$. Lasso: $(18 - 4.5)/9 = 13.5/9 = 1.5$. The lasso is exactly 0 for every $\lambda \ge 18$; ridge is never exactly 0.

B. Noise sd σ = 1.5, prior sd τ = 0.5. Find λ for ridge. Then for a Laplace(0, b = 0.25) prior, find λ in the ½RSS lasso form, in the RSS form, and scikit-learn's alpha with n = 90.

Ridge: $\lambda = \sigma^2/\tau^2 = 2.25/0.25 = 9$. Lasso, ½RSS form: $\lambda = \sigma^2/b = 2.25/0.25 = 9$; RSS form: $2\sigma^2/b = 18$; scikit-learn: $\alpha = \lambda/n = 9/90 = 0.1$.

C. One coefficient: $\bar y = 1.5$, s = 1, prior Laplace(0, 1). Compute the MAP. The posterior mean is about 0.81 and the median about 0.73. Why are they bigger than the MAP?

MAP $= S_1(1.5) = 1.5 - 1 = 0.5$. The posterior on $\theta \gt 0$ is a Normal piece centred at $\bar y - s^2/b = 0.5$, cut at 0; cutting off the left part pushes the mean and median to the right of 0.5, and there is also a small piece on $\theta \lt 0$. Overall the posterior is skewed: a sharp peak with a long right side. The mode (MAP) sits at the peak, while the mean and median are pulled toward the long side. (About 85% of the posterior is above 0.)

D. Explain to an interviewer: "Prophet's changepoint prior is L1 regularization, right?"

"Partly. Prophet puts δⱼ ~ Laplace(0, changepoint_prior_scale) on the slope changes. With a Gaussian likelihood, minus the log posterior is squared error over 2σ² plus |δ|/b, so the MAP, which Prophet computes by default, is an L1-penalized fit with λ = σ²/b in the ½RSS convention. That is why a MAP fit can switch some changepoints off. But the prior is a distribution, not a penalty: the full posterior (Prophet with MCMC, or a NumPyro model with NUTS or SVI) has no point mass at zero, so the δⱼ are shrunk toward zero but not exactly zero. The precise statement: Laplace prior ↔ L1 at the MAP only."

E. Orthonormal features with least-squares estimates $z = (2.5, -0.4, 1.2, -1.8)$ and λ = 1. Give the ridge and lasso estimates. Which features does the lasso select?

Ridge: $z/2 = (1.25, -0.2, 0.6, -0.9)$, all non-zero. Lasso: $S_1(z) = (1.5, 0, 0.2, -0.8)$: features 1, 3 and 4 are selected; feature 2 ($|{-0.4}| \le 1$) is switched off.

F. Spike-and-slab prior: θ = 0 with probability 0.5, otherwise θ ~ N(0, 2). Data: $\bar y = 0.8$ with s = 1. Compute $P(\theta = 0\mid D)$ and contrast it with the Laplace(0, 1) posterior.

Under the spike, $\bar y \sim N(0, 1)$: density $\varphi(0.8) \approx 0.2897$. Under the slab, $\bar y \sim N(0, 1 + 2)$: density $\frac{1}{\sqrt{2\pi\cdot3}}e^{-0.64/6} \approx 0.2303 \times 0.8988 \approx 0.2070$. So $P(\theta = 0\mid D) = \frac{0.5\times0.2897}{0.5\times0.2897 + 0.5\times0.2070} \approx 0.583$. With the Laplace prior (same variance), $P(\theta = 0\mid D) = 0$ even though the MAP is 0. A posterior probability of "exactly no effect" exists only if the prior allowed it.

Chapter 5.4 · Syllabus Module 12

Sampling and its biases

Every number you estimate comes from a sample, and a sample can mislead you in two very different ways. It can be unlucky (random sampling variation, which shrinks as the sample grows) or it can be unfair (bias: the way units got into your data favours some of them, and no amount of extra data fixes that). This chapter teaches both, the classic biases (selection, non-response, survivorship), the two clever designs (stratified and cluster sampling), and the assumption behind almost every formula in this guide: IID.

  • Tell apart the population, the sampling frame and the sample, and say what makes a simple random sample
  • See sampling variation and take a first look at the sampling distribution of an estimate
  • Recognize sampling bias and selection bias, and explain why bias does not shrink with more data, including the A/B trap of analysing only users selected after treatment
  • Quantify non-response bias and spot survivorship bias in product data
  • Use stratified sampling to reduce variance, and explain why cluster sampling increases it (the design effect $1 + (m-1)\rho$)
  • State exactly what independence and IID mean, and when they fail: time series, repeated users, clusters, drift

What we need from earlier chapters: population vs sample and "a statistic is a number computed from a sample" (Chapter 4.1); independence and conditional probability (Chapter 4.3); variance rules and the law of total variance (Chapters 4.5–4.6); estimators, bias and MSE = bias² + variance (Chapter 5.1). Standard errors are owned by Chapter 5.5; here you only get a first look. Notation: $N$ = population size, $n$ = sample size, $p$ = a population proportion (for example a conversion rate), $\hat p$ = its estimate from the sample.

Population, sampling frame, sample, and random sampling core

A cook tastes one spoonful of soup to judge the whole pot. That works only if the spoonful is like the pot. If the salt sank to the bottom and she tastes only from the top, her spoon lies to her, however carefully she tastes. The fix is simple and powerful: stir, then take a random spoonful. Random sampling is the statistical way of stirring.

Three words keep the picture straight. The population is everything you want to learn about (the whole pot: all your users). The sampling frame is the list you can actually pick from (the users you can reach, for example those with an email address). The sample is the units you end up measuring (the spoonful). Trouble enters at every step between them.

Three ways to say it:

  • Picture: stir the pot, then taste: every spoonful position equally likely.
  • Numbers: 30% of users are mobile (5% convert) and 70% desktop (12% convert): the true rate is 9.9%; a list that is 90% desktop suggests 11.3%.
  • Slogan: a sample can only speak for the people it could have contained.

A product has 10,000 users: 30% on mobile (conversion 5%) and 70% on desktop (conversion 12%).

  1. True population conversion rate: $0.3 \times 0.05 + 0.7 \times 0.12 = 0.015 + 0.084 = 0.099$, i.e. 9.9%.
  2. A simple random sample of 200 users: every user has the same chance $200/10000 = 2\%$ of being picked. On average the sample is 30% mobile, and the expected number of conversions is $200 \times 0.099 = 19.8$.
  3. A "convenient" sample from the newsletter list, which happens to be 90% desktop: expected rate $0.1 \times 0.05 + 0.9 \times 0.12 = 0.005 + 0.108 = 0.113$, i.e. 11.3%.
  4. The convenient sample overstates conversion by $11.3 - 9.9 = 1.4$ points, not because of bad luck, but because of who could get in.
  • Population: the full set of units (users, days, orders) you want a conclusion about. A parameter (like the true rate $p$) describes it.
  • Unit: one member of the population (one user, one day).
  • Sampling frame: the list or mechanism from which the sample is actually drawn. If the frame misses part of the population (users without email, days before logging started), that part can never appear: coverage error.
  • Sample: the $n$ units actually measured. A statistic (like $\hat p$) is a number computed from it (Chapter 4.1).
  • Simple random sample (SRS) of size $n$: every possible set of $n$ units from the frame is equally likely to be chosen. Then every unit has inclusion probability $n/N$, and $\hat p$ is an unbiased estimator of the frame's $p$.
  • Probability sample: every unit has a known, non-zero chance of selection (SRS, stratified, cluster sampling are all examples). A convenience sample (whoever is easiest to reach) has unknown chances, so its bias cannot be measured or corrected.
Why do we need it?

Measuring everyone is usually too slow, too costly or impossible (future users). Random sampling is what lets a small sample speak for a large population, with an error you can quantify.

Where is it used?

Surveys and polls, quality checks on a sample of orders, labelling a random sample of data for an ML model, train/test splits (train_test_split is a random sample), random assignment of users to A/B variants, and down-sampling huge event logs.

How is it used?

Define the population, check the frame covers it, then draw with a seeded random generator: df.sample(n=200, random_state=0) or rng.choice(N, size=n, replace=False). Write down the frame and the method, so readers know who the sample can speak for.

Population all your users Sampling frame who you can reach Sample who you pick Analysed data who answers / remains coverage error parts never listed selection / sampling bias unequal chances to be picked non-response, survivorship who drops out depends on y Random sampling variation (luck) appears at the "pick" step and shrinks as n grows. The red biases are systematic: they do not shrink as n grows.
The road from the population to the data you analyse. Each arrow is a place where the data can stop representing the population. Only the randomness of the "pick" step is fixed by a bigger sample.

Each dot is one user, in sign-up order (left to right, top to bottom). Big dots converted. Press Draw a sample: blue rings mark the users picked. With Simple random, the estimate lands near the truth, sometimes above, sometimes below. Switch to First n sign-ups or Newsletter volunteers and draw several times: the estimate is too high again and again. Tick Reveal segments to see why: early sign-ups and newsletter readers are mostly desktop users (pink), who convert more than mobile users (teal).

"A big sample is a representative sample."

Size and fairness are different things. A million newsletter readers still over-represent desktop users. Representativeness comes from how units were chosen, not from how many.

"Random means haphazard: I just picked whoever came to mind."

Random means chosen by a chance mechanism with known probabilities (a seeded random number generator). Human "random" choices are biased in ways we cannot see.

"My sample is random, so it represents everyone in the company's market."

It represents the frame it was drawn from. If the frame is "users who logged in this month", conclusions are about them, not about churned users or future users.

Population (what you want) → frame (what you can reach) → sample (what you measure).

SRS: every set of $n$ units equally likely; inclusion probability $n/N$; $\hat p$ unbiased for the frame.

Trap: big ≠ representative; a sample speaks only for its frame.

Quick check: you survey a random 500 of the users who opened the app yesterday. Who can your results speak for?

For users who opened the app yesterday (the frame). Infrequent users are under-covered and churned users are not covered at all, so you cannot claim the result holds for "all customers".

Sampling variation and the sampling distribution (first look) core

Take three random spoonfuls from the same well-stirred pot and they will taste slightly different. Take three random samples of 100 users and you might see 8, 12 and 10 conversions. Nothing is wrong: different random samples contain different people. This unavoidable wobble is sampling variation.

If you could repeat the sampling thousands of times and draw a histogram of all the estimates, you would get the sampling distribution of the estimate. For a random sample it is centred on the truth, and it gets narrower as $n$ grows. Its width tells you how far a single estimate is typically off (you will compute it as the standard error in Chapter 5.5).

Three ways to say it:

  • Picture: a cloud of estimates around the true value, tighter for bigger samples.
  • Numbers: true rate 9.9%, samples of 100: estimates typically within about ±3 points; samples of 400: about ±1.5.
  • Slogan: one sample, one estimate; another sample, another estimate.

True conversion rate $p = 0.099$. Three random samples of 100 users gave 8, 12 and 10 conversions.

  1. Estimates: $8/100 = 0.08$, $12/100 = 0.12$, $10/100 = 0.10$. All three are honest; they differ only by luck.
  2. Their average, $(0.08 + 0.12 + 0.10)/3 = 0.10$, is close to the truth: random sampling is right on average.
  3. How big is a typical miss? The number of conversions in $n$ users is Binomial, so the spread of $\hat p$ is $\sqrt{p(1-p)/n} = \sqrt{0.099 \times 0.901/100} = \sqrt{0.000892} \approx 0.030$: about 3 points.
  4. Quadruple the sample to 400: $\sqrt{0.099 \times 0.901/400} \approx 0.015$. Four times the data halves the typical miss (the $\sqrt n$ rule, derived in Chapter 5.5).
  • Sampling variation: the differences between estimates computed from different random samples of the same population.
  • Sampling distribution of a statistic $\hat\theta$: the probability distribution of $\hat\theta$ over all the samples the sampling method could have produced. It is a distribution of estimates, not of the data (Chapter 5.1).
  • Its centre minus the truth is the bias; its standard deviation is the standard error (Chapter 5.5). For an SRS proportion: centre $p$, standard error $\sqrt{p(1-p)/n}$ (ignoring the small finite-population correction when $n \ll N$).
  • By the CLT the sampling distribution of a mean or proportion is roughly Normal for large $n$ (Chapter 4.13).
Why do we need it?

Without it you cannot tell whether a difference between two numbers is real or luck. Every error bar, p-value and confidence interval in this guide is a statement about a sampling distribution.

Where is it used?

Standard errors and confidence intervals (5.5, 5.8), hypothesis tests and A/B readouts (5.6–5.11), the bootstrap (which simulates a sampling distribution by resampling), and judging how much a validation-set metric would change with another split.

How is it used?

Ask "what would happen if I repeated the data collection?". Simulate it when you can (draw many samples, compute the statistic each time, look at the histogram), or use the formula for its standard error. Report every estimate with that spread.

The true conversion rate is 9.9% (green line). Press Draw 1 sample a few times: each sample of $n$ users gives one estimate (blue line) and adds one block to the histogram. Then press Draw 1,000 samples: the histogram is the sampling distribution. Move $n$ from 25 to 400 and draw again: the histogram stays centred on the truth but becomes much narrower. The readout compares the spread with $\sqrt{p(1-p)/n}$.

"My estimate differs from last month's, so something changed."

Two random samples of the same population differ by luck alone. Compare the difference with the sampling spread before concluding anything (Chapters 5.5–5.6).

"The sampling distribution is the distribution of my data."

The data distribution is how individual users vary (convert or not). The sampling distribution is how the estimate would vary across repeated samples. It is much narrower, and narrows further as $n$ grows.

Sampling variation = differences between estimates from different random samples (luck).

Sampling distribution = distribution of the estimate over repeated samples; for an SRS proportion centred at $p$ with spread $\sqrt{p(1-p)/n}$.

Trap: it describes estimates, not data; 4× the data halves the spread.

Quick check: true rate 20%, samples of 100. About how far is a typical estimate from 20%?

$\sqrt{0.2 \times 0.8/100} = \sqrt{0.0016} = 0.04$: about 4 percentage points. Estimates of 16% or 24% are perfectly ordinary.

Sampling bias and selection bias: errors that more data cannot fix core

A bathroom scale that always reads 2 kg too heavy is biased. Weighing yourself a hundred times and averaging does not help: the average is still 2 kg too heavy. Averaging only removes random wobble. Sampling works the same way. If the way people get into your data favours some of them (heavy users, desktop users, happy customers), the estimate is pushed in one direction, and a bigger sample just gives you a more precise wrong answer.

Selection bias is the general name: the units you analyse were selected by a process that is related to the thing you measure. Sampling bias is the special case where the sampling method itself gives some units a bigger chance. A sneaky version hides inside A/B tests: if you analyse only users selected by something the treatment itself changed, the comparison stops being fair.

Three ways to say it:

  • Picture: a target shooter whose sight is off: the shots cluster tightly, in the wrong place.
  • Numbers: a desktop-heavy sample says 11.3% instead of 9.9%, with 100 users or with 10,000.
  • Slogan: more data shrinks the noise, not the bias.

Part 1: bias does not shrink. Truth $p = 0.099$; the biased method's estimates centre on $0.113$ (bias $+0.014$).

  1. Random sampling, $n = 100$: typical miss $\sqrt{0.099 \times 0.901/100} \approx 0.030$. Biased method: total error $\sqrt{\text{bias}^2 + \text{spread}^2} = \sqrt{0.014^2 + 0.113 \times 0.887/100} \approx \sqrt{0.000196 + 0.001002} \approx 0.035$.
  2. $n = 10{,}000$: random sampling misses by about $0.003$. The biased method: $\sqrt{0.000196 + 0.00001} \approx 0.0144$. Almost all of that is the bias, $0.014$, which no sample size removes.

Part 2: selecting after treatment. An A/B test with 10,000 users per arm. Control: 20% add something to the cart (2,000 users) and 25% of those buy (500): overall 5.0%. The new "quick add" button brings an extra 1,000 low-intent users to the cart, who buy at 10%.

  1. Treatment buyers: $2000 \times 0.25 + 1000 \times 0.10 = 500 + 100 = 600$. Overall conversion: $600/10000 = 6.0\%$, up from 5.0%. The button helps.
  2. "Conversion among users who added to cart": control $500/2000 = 25\%$; treatment $600/3000 = 20\%$. It looks like the button hurts.
  3. The cart group in treatment contains different kinds of people (extra low-intent users). Selecting on a variable the treatment changed broke the fairness that randomization created.
  • Sampling bias: the sampling method gives units different chances of selection, and those chances are related to the outcome. Example: sampling from a newsletter list when newsletter readers convert more.
  • Selection bias (broader): the units in the analysis were selected through a process related to the outcome, or, in a comparison, related to both the group and the outcome. It includes sampling bias, self-selection (volunteers, opt-in surveys), non-response, survivorship, attrition, and selecting on a post-treatment variable.
  • Bias does not depend on $n$: $\text{MSE} = \text{bias}^2 + \text{variance}$ (Chapter 5.1). The variance shrinks like $1/n$; the bias² stays.
  • Remedies: a probability sample from a frame that covers the population; when the selection probabilities are known, weight each unit by $1/\text{(its inclusion probability)}$ (inverse-probability weighting); in experiments, compare groups only on subsets defined before treatment (or by a trigger that is the same in both arms).
Why do we need it?

Bias is the error that hides: confidence intervals and p-values only account for random variation, so a biased analysis looks precise and confident while being wrong. Knowing the forms of selection bias is the only defence.

Where is it used?

Survey design, opt-in feedback analysis, training data for ML models (labels only for items users chose to rate), A/B analyses restricted to "engaged" or "triggered" users, and observational product analytics ("users who used feature X retain better").

How is it used?

For every dataset ask: who could get in, and does that depend on the outcome? In experiments, define segments using pre-treatment information only (device, country at assignment), analyse everyone who was randomized, and treat any post-treatment filter as a red flag.

Each histogram shows 400 repeated estimates. Blue: simple random samples (centred on the true 9.9%, green line). Orange: samples from the desktop-heavy list (centred on 11.3%). Move the sample size from 25 to 10,000. Both histograms get narrower, but the orange one tightens around the wrong value: at large $n$ it no longer even touches the truth. The readout shows that the total error of the biased method stops falling at about the size of the bias.

10,000 users per arm. In control, 20% reach the cart and 25% of them buy. The new button pushes extra, lower-intent users into the cart. Leave the defaults and compare the two pairs of bars: overall conversion (fair: all randomized users) goes up, while "conversion among cart users" goes down. Set the extra users' purchase rate to 25% (as good as everyone else): now both comparisons agree. The cart filter is only fair when the treatment does not change who passes it.

"With a million rows, bias does not matter any more."

With a million rows the random error is tiny and the bias is all that is left. Big biased data gives confident wrong answers.

"Randomization protects every subgroup comparison."

It protects comparisons of all randomized users, and of subgroups defined by information fixed before assignment. Subgroups defined by behaviour after assignment (reached the cart, stayed active, opened the email) can differ between arms because of the treatment itself.

"Selection bias and sampling variation are two words for error."

Sampling variation is random and averages out; it is what standard errors measure. Selection bias is systematic and does not average out; standard errors do not include it.

In an A/B framework like yours, the posterior for $\theta_A$ and $\theta_B$ is only as fair as the users fed into the likelihood. Randomization makes the two arms comparable as assigned; filtering to "users who reached step X" before fitting the Beta-Binomial model can flip a conclusion, exactly as in the cart example, and a narrow posterior will make the wrong answer look certain. Segments for the hierarchical model should be defined from pre-assignment attributes (device or country at assignment), not from post-treatment behaviour.

Selection bias: who enters the analysis depends on the outcome (or on group and outcome). Sampling bias: the sampling method causes it.

MSE = bias² + variance: more data shrinks variance only.

Trap: never compare A/B arms on a subset defined after treatment.

Quick check: users can rate a movie only if they finished it. The average rating is 4.4 stars. Is that the average opinion of everyone who started the movie?

No. People who disliked it tend to stop early and never rate, so the raters are selected by their opinion. 4.4 is the average among finishers, biased upward as an estimate of all viewers' opinion.

Non-response bias: the people who do not answer

You email a satisfaction survey to 1,000 randomly chosen customers. The sample was perfect. But only some answer, and the happy ones are more willing to click a survey link than the annoyed ones. The answers you receive are from a group that leans happy, so they overstate satisfaction. That is non-response bias.

The surprising part: a low response rate alone does not cause bias. If happy and unhappy customers answer at the same low rate, the respondents still look like the sample. Bias appears when the chance of answering is related to the answer.

Three ways to say it:

  • Picture: a sieve that lets satisfied customers through more easily than unsatisfied ones.
  • Numbers: 60% are satisfied; satisfied answer 30% of the time, unsatisfied 10%: respondents are 82% satisfied.
  • Slogan: it is not how many answer, it is who answers.

1,000 sampled customers: 600 satisfied, 400 not. Satisfied customers respond with probability 0.30, unsatisfied with 0.10.

  1. Expected responses: $600 \times 0.30 = 180$ satisfied and $400 \times 0.10 = 40$ unsatisfied; 220 in total (response rate 22%).
  2. Estimate from respondents: $180/220 \approx 0.818$, i.e. 81.8% satisfied, against the true 60%. Bias $\approx +21.8$ points.
  3. Formula check: the average response rate is $\bar\rho = 0.6 \times 0.3 + 0.4 \times 0.1 = 0.22$. The covariance between "responds" and "satisfied" is $E[\rho y] - \bar\rho\,E[y] = 0.6 \times 0.3 - 0.22 \times 0.6 = 0.18 - 0.132 = 0.048$. Bias $= 0.048/0.22 \approx 0.218$. ✓
  4. If both groups responded at 10% (response rate only 10%!), we would expect 60 satisfied and 40 unsatisfied: 60%, no bias.

Non-response bias arises when selected units fail to provide data and the failure is related to the outcome. If unit $i$ responds with probability $\rho_i$ and has outcome $y_i$, the respondent mean estimates

$$\frac{\sum_i \rho_i y_i}{\sum_i \rho_i} = \bar y + \frac{\mathrm{Cov}(\rho, y)}{\bar\rho},$$

where the covariance and $\bar\rho$ are taken over the population. So the bias is (correlation between responding and the outcome) × (their spreads) ÷ (average response rate): zero when responding is unrelated to $y$, and larger when response rates are low.

  • Unit non-response: the whole unit is missing (no survey returned). Item non-response: some answers are missing.
  • Remedies: follow up with non-respondents, weight respondents by $1/\hat\rho$ using known characteristics (age, plan, usage), and compare respondents with the frame on variables you know for everyone.
Why do we need it?

Most real data collection has missing responses: surveys, ratings, NPS, opt-in tracking, app-store reviews. Without this idea you would trust averages that describe only the people who chose to answer.

Where is it used?

Customer satisfaction and NPS surveys, product feedback widgets, polling, ratings-based recommenders (missing-not-at-random ratings), and event logging where some clients block tracking.

How is it used?

Report the response rate, compare respondents and non-respondents on known variables (plan, tenure, usage), reweight if they differ, and say clearly that results describe respondents unless you corrected for it.

Green dots are satisfied customers (60%), red dots are not (40%). Filled dots answered the survey; faint dots did not. Start with the defaults (satisfied answer 30%, unsatisfied 10%): the respondents look far happier than the customers. Press Equal response rates: the response rate drops to 10%, yet the estimate is now about right. Then set the unsatisfied rate higher than the satisfied rate (angry customers love to complain): the bias flips sign. Press New survey to see the random wobble on top of the bias.

"A 10% response rate makes the survey useless; 60% makes it reliable."

Bias depends on whether responding is related to the answer. A 10% response rate with unrelated non-response can be unbiased; a 60% rate where unhappy customers stay silent can be badly biased (though high response rates do limit how large the bias can get).

"Non-response only matters for surveys."

Any missing outcome counts: users who never rate, devices that drop tracking events, customers who delete their accounts. If missingness depends on the outcome, the observed data are biased.

Respondent mean ≈ $\bar y + \mathrm{Cov}(\rho, y)/\bar\rho$ (ρ = response probability).

Example: 60% satisfied, response 30% vs 10% → respondents 81.8% satisfied.

Trap: low response rate ≠ bias; response related to the outcome = bias.

Quick check: in the example, what response rate for unsatisfied customers would make the survey unbiased?

0.30, the same as for satisfied customers. Then Cov(ρ, y) = 0 and the respondents (expected 180 satisfied, 120 unsatisfied) are 60% satisfied.

Survivorship bias: only the survivors are in the data

In the Second World War, engineers studied the bullet holes on bombers that came back from missions and planned to add armour where the holes were thickest. The statistician Abraham Wald pointed out the flaw: they were only looking at planes that survived. Planes hit in the engines did not come back, so engine hits were missing from the data. The armour belonged where the surviving planes had no holes.

Product data is full of the same trap. "Average spend per active user grows every month" can happen even if nobody's spend changes: low spenders churn faster, so the survivors are increasingly the big spenders. The data contain only the units that survived some filter, and the filter is related to what you measure.

Three ways to say it:

  • Picture: you only ever see the planes that came home.
  • Numbers: half the users spend 10, half spend 50; low spenders churn 30% a month, high spenders 5%: average spend of active users goes 30 → 33.0 → 35.9 → 38.6 with no one changing.
  • Slogan: the missing ones are missing for a reason.

A cohort of users: 50% spend 10 per month and 50% spend 50. Each month 30% of the low spenders and 5% of the high spenders churn. Individual spending never changes.

  1. Month 0: average spend $= 0.5 \times 10 + 0.5 \times 50 = 30$.
  2. Month 1: low spenders left $0.5 \times 0.7 = 0.35$, high $0.5 \times 0.95 = 0.475$. Average spend of active users $= (0.35 \times 10 + 0.475 \times 50)/(0.35 + 0.475) = (3.5 + 23.75)/0.825 \approx 33.0$.
  3. Month 2: $0.245$ and $0.451$ remain; average $= (2.45 + 22.56)/0.696 \approx 35.9$.
  4. Month 3: $0.172$ and $0.429$; average $\approx 38.6$.
  5. A dashboard of "spend per active user" shows steady growth. The truth: the mix of survivors changed, not the behaviour.

Survivorship bias is selection bias in which the data contain only units that passed a survival filter (still active, still in business, still running, came back), and survival is related to the variable being studied.

  • Typical sources: analysing only current customers, only funds or companies that still exist, only experiments that shipped, only products still in the catalogue, only users who completed onboarding.
  • Fix: start from a cohort defined at the beginning (everyone who signed up in January) and follow all of them, including those who left; compare within the same units over time.
Why do we need it?

Survivorship creates trends and correlations out of nothing: rising averages, "successful companies all did X", "our long-time users love feature Y". Recognizing it stops you from acting on patterns made by the filter.

Where is it used?

Cohort and retention analysis, finance (fund performance databases), churn modelling, forecasting with a catalogue that drops discontinued items, and model training on approved applications only (loans, fraud reviews).

How is it used?

Define cohorts at a starting point and keep everyone in them (also those who churned, with their zeros); compare the same units over time; when only survivors are available, model the survival process explicitly or state the limitation.

engine: no holes holes seen on returning planes (wings, tail, body) Missing data: planes hit in the engine never came back The places with no holes in the surviving data are the places where hits were fatal.
Wald's bomber problem, the classic picture of survivorship bias. The data contain only returning planes, so a lack of holes in the engines is evidence that engine hits were deadly, not that engines were rarely hit.

Half the cohort spends 10 per month, half spends 50, forever. Set how fast each group churns. Top: average spend per active user over 12 months (orange) and the average over the whole original cohort (green, flat at 30). Bottom: who the active users are each month. With different churn rates the orange line climbs while every individual stays the same; set the two churn rates equal and the illusion disappears. Drag the month slider to read the arithmetic.

"Our long-term users all use feature X, so feature X keeps users."

Maybe, or maybe users who would stay anyway are the kind who explore features. Only the survivors are in the "long-term" group. A cohort analysis from sign-up, or an experiment, is needed for a causal claim (Chapter 5.12).

"An increasing average per active user means users are spending more."

It can come purely from the changing mix of who is still active. Track the same users over time, or report the cohort total including churned users.

For your forecasting model, survivorship hides in the history: if the training data include only stores or products that still exist today, the discontinued ones (often the declining ones) are missing, and the trends look healthier than reality. In the A/B framework, computing a metric only over users still active at the end of the test is a survivorship filter; if one variant changes who stays, the comparison is biased (attrition, Chapter 5.11).

Survivorship bias = analysing only units that passed a survival filter related to the outcome.

Example: churn 30% (spend 10) vs 5% (spend 50) → spend per active user 30 → 33.0 → 35.9 → 38.6 with no one changing.

Trap: changing mix ≠ changing behaviour. Fix: follow a cohort from the start, including leavers.

Quick check: a list of "the 20 best-performing startups of the decade" shows that 18 had a founder who dropped out of university. Should you drop out?

No conclusion is possible from survivors alone: you would need the rate of success among all dropouts and all non-dropouts, including the many failed startups that never make such lists.

Stratified sampling: sample inside every group core

Suppose 80% of your customers are small shops spending about 100 a month and 20% are big shops spending about 500. A simple random sample of 100 might, by luck, contain 12 big shops or 28. Each extra big shop moves the average a lot, so the estimate jumps around. That jumping is caused by luck in the mix, and you can remove it.

Stratified sampling splits the population into groups called strata (here: small and big shops), draws a separate random sample inside each one, and combines the group averages using the known group sizes. The mix is now fixed by design (exactly 20% big shops in the formula), so only the variation within each group is left.

Three ways to say it:

  • Picture: taste every layer of a layered cake, then weight each taste by the layer's thickness.
  • Numbers: the same 100 shops: standard error 16.7 with simple random sampling, 4.8 with stratification, 3.6 with the best split.
  • Slogan: do not leave the mix to luck.

Small shops: share $W_1 = 0.8$, mean 100, sd 20. Big shops: $W_2 = 0.2$, mean 500, sd 100. Sample size $n = 100$.

  1. Population mean: $0.8 \times 100 + 0.2 \times 500 = 80 + 100 = 180$.
  2. Total variance (law of total variance, Chapter 4.6): within $= 0.8 \times 20^2 + 0.2 \times 100^2 = 320 + 2000 = 2320$; between $= 0.8 \times (100 - 180)^2 + 0.2 \times (500 - 180)^2 = 5120 + 20480 = 25600$; total $= 27920$.
  3. Simple random sample: $SE = \sqrt{27920/100} \approx 16.7$.
  4. Stratified, proportional allocation (80 small, 20 big): $\hat\mu = 0.8\,\bar y_1 + 0.2\,\bar y_2$, with $Var = 0.8^2 \times 400/80 + 0.2^2 \times 10000/20 = 3.2 + 20 = 23.2$, so $SE \approx 4.8$. The huge "between" part (25,600) is gone.
  5. Neyman allocation puts more of the sample where the spread is: $n_h \propto W_h\sigma_h$, i.e. $0.8 \times 20 = 16$ vs $0.2 \times 100 = 20$, so about 44 small and 56 big. Then $Var = (\sum_h W_h\sigma_h)^2/n = 36^2/100 = 12.96$ and $SE = 3.6$.

Split the population into non-overlapping strata $h = 1, \dots, H$ with known shares $W_h = N_h/N$. Take an independent simple random sample of size $n_h$ in each, with sample means $\bar y_h$. The stratified estimator of the population mean is

$$\hat\mu_{st} = \sum_h W_h\,\bar y_h, \qquad Var(\hat\mu_{st}) = \sum_h W_h^2\,\frac{\sigma_h^2}{n_h} \quad\text{(ignoring finite-population corrections)}.$$
  • Proportional allocation $n_h = nW_h$ gives $Var = \frac{1}{n}\sum_h W_h\sigma_h^2$: only the within-stratum variance remains. It is never worse than SRS (apart from the tiny finite-population term).
  • Neyman (optimal) allocation $n_h \propto W_h\sigma_h$ gives the smallest variance for a fixed total $n$: $(\sum_h W_h\sigma_h)^2/n$.
  • The gain is large when strata means differ a lot (big between-strata variance) and small when they are similar.
  • If strata are sampled at different rates, an unweighted average of all sampled units is biased; always weight by $W_h$.
Why do we need it?

It gives the same precision with a much smaller sample when groups differ, guarantees that small but important groups appear in the sample, and gives a separate estimate for each group.

Where is it used?

National surveys and polls, audit sampling, train_test_split(..., stratify=y) to keep class proportions in ML splits, stratified K-fold cross-validation, and stratified (blocked) randomization or post-stratification by segment in A/B tests.

How is it used?

Choose strata that predict the outcome (segment, size, region), sample within each (df.groupby("segment").sample(frac=0.05, random_state=0)), estimate each stratum's mean, and combine with the population shares $W_h$, not the sample shares.

Each histogram shows the error (estimate − true mean) of 1,000 repeated samples of 100 shops. Blue: simple random sample. Orange: stratified sample with the allocation you choose. With the defaults the orange histogram is about three times narrower. Push the big-shop mean down toward 100 (the strata become alike): the advantage nearly vanishes. Try Neyman allocation with a large big-shop spread, and compare it with Equal allocation.

"Stratifying always gives a big improvement."

It removes only the between-strata part of the variance. For conversion rates of 5% (mobile) vs 12% (desktop), the between part is tiny: stratifying cuts the variance by only about 1%. It pays off when the strata means really differ, as with shop sizes.

"Pool all the sampled units and take the plain average."

With unequal sampling rates (Neyman, equal allocation, oversampling a small group) the plain average is biased. Weight each stratum mean by its population share $W_h$.

Segments in an A/B framework like yours (for example device, country or plan) play the role of strata. Randomizing within each segment, or re-weighting segment results by their traffic shares, removes the luck of the segment mix from the overall comparison. The same split is what your hierarchical model sees as groups; there the goal is different (estimating every segment's own effect with partial pooling, Chapter 6.6), but the "between vs within" variance picture is the same.

$\hat\mu_{st} = \sum_h W_h\bar y_h$; $Var = \sum_h W_h^2\sigma_h^2/n_h$. Proportional: within-variance only. Neyman: $n_h \propto W_h\sigma_h$.

Example: SE 16.7 (SRS) → 4.8 (proportional) → 3.6 (Neyman).

Trap: gains only when strata means differ; weight by population shares.

Quick check: two strata with equal shares, the same mean, and sds 10 and 30. Does stratification help, and how should you allocate 100 units?

Proportional stratification gives no gain over SRS, since there is no between-strata variance. Neyman allocation still helps a little: $n_h \propto W_h\sigma_h = 5 : 15$, i.e. 25 and 75 units. Variance $(0.5\times10 + 0.5\times30)^2/100 = 4$, versus $(0.5\times100 + 0.5\times900)/100 = 5$ for proportional allocation.

Cluster sampling: cheaper, but each extra unit tells you less core

Visiting 200 random customers spread over 50 stores is expensive. It is much cheaper to pick 5 stores and survey all 40 customers in each. That is cluster sampling: you sample whole groups (clusters). The catch: customers of the same store are alike (same neighbourhood, same manager, same stock). The 40 answers from one store partly repeat each other, so they carry less information than 40 independent answers.

Stratified and cluster sampling use groups in opposite ways. Strata: sample inside every group, so the differences between groups cancel out. Clusters: sample a few whole groups, so the differences between groups dominate the error. The design effect says how much bigger the variance is than for a simple random sample of the same size.

Three ways to say it:

  • Picture: 40 people from one street are like a few voices, not 40.
  • Numbers: 5 stores × 40 customers with a within-store correlation of 0.05: the 200 answers are worth only about 68 independent ones.
  • Slogan: similar units repeat each other; count clusters, not rows.

50 stores with 40 customers each. Design A: a simple random sample of 200 customers. Design B: 5 random stores, all 40 customers in each (also 200 customers).

  1. Suppose the intra-cluster correlation (how alike two customers of the same store are) is $\rho = 0.05$.
  2. Design effect: $\text{DEFF} = 1 + (m - 1)\rho = 1 + 39 \times 0.05 = 2.95$. Design B's variance is about 2.95 times design A's.
  3. Effective sample size: $200/2.95 \approx 68$. The standard error is $\sqrt{2.95} \approx 1.72$ times larger.
  4. With $\rho = 0.2$: $\text{DEFF} = 1 + 39 \times 0.2 = 8.8$ and the 200 customers are worth only $200/8.8 \approx 23$ independent ones.
  5. Even a small $\rho$ matters when clusters are big, because it is multiplied by $m - 1$.

In cluster sampling the population is divided into clusters (stores, households, schools, users with many sessions); a random sample of clusters is drawn and all (or some) units inside them are measured.

  • Write each value as cluster effect + individual noise: $y_{ij} = \mu + u_j + e_{ij}$ with $Var(u_j) = \tau^2$ (between clusters) and $Var(e_{ij}) = \sigma^2$ (within). The intra-cluster correlation is $\rho = \tau^2/(\tau^2 + \sigma^2)$: the share of the total variance that lies between clusters, and the correlation between two units of the same cluster.
  • For $k$ clusters of equal size $m$ ($n = km$): $Var(\bar y) = \dfrac{\tau^2 + \sigma^2/m}{k} = \dfrac{\tau^2 + \sigma^2}{n}\,\big[1 + (m - 1)\rho\big]$.
  • Design effect $\text{DEFF} = 1 + (m - 1)\rho$; effective sample size $n_{\text{eff}} = n/\text{DEFF}$. For $\rho \gt 0$ cluster sampling has larger variance than an SRS of the same size; its advantage is cost.
  • Analysis must respect the clusters: standard errors that treat the $n$ units as independent are too small by a factor $\sqrt{\text{DEFF}}$ (use cluster-robust SEs, aggregate to cluster level, or a hierarchical model).
Why do we need it?

Clustered data are everywhere, sampled on purpose or not: sessions within users, users within companies, days within stores. Ignoring the clustering makes confidence intervals too narrow and p-values too small.

Where is it used?

Household and school surveys, field audits, A/B tests whose randomization unit is larger than the analysis unit (randomize users, analyse page views), cluster-randomized experiments (by city or store), GroupKFold cross-validation, and cluster-robust standard errors in statsmodels.

How is it used?

Estimate ρ (or just compute standard errors at the cluster level), compute DEFF and the effective sample size when planning, and in the analysis use cov_type="cluster" in statsmodels, aggregate to one row per cluster, or fit a hierarchical model.

Stratified: a few units from EVERY group Cluster: ALL units from a FEW groups between-group differences cancel → smaller variance between-group differences dominate → larger variance blue = sampled units · each box = one group (segment, store, user)
The same groups, used in opposite ways. Stratified sampling takes some units from every group, so the groups' differences do not add noise. Cluster sampling takes whole groups, so which groups you happened to pick adds noise; the price is paid in variance, the gain is in cost.

A fixed population: 30 stores × 40 customers. Each design surveys 120 customers, repeated 250 times; the histograms show the error of each design's estimate of the average. Blue: simple random sample of customers. Orange: stratified by store (4 customers from every store). Red: cluster sample (3 whole stores). Raise the intra-cluster correlation ρ (how much stores differ): the red histogram spreads out dramatically, the orange one gets narrower relative to blue, and at ρ = 0 all three are about the same.

Set the intra-cluster correlation ρ, the cluster size $m$ (customers per store, sessions per user) and the total number of observations $n$. The curve shows the effective sample size as a fraction of $n$ for every cluster size; faint curves show ρ = 0.01, 0.05 and 0.2 for reference. Notice how fast big clusters lose value: with ρ = 0.05, clusters of 100 keep only about 17% of the information.

"I have 10,000 sessions, so my standard error is $\sigma/\sqrt{10000}$."

If those sessions come from 500 users and sessions of the same user are alike, the effective sample size is far smaller. Count independent units (users), or use cluster-robust standard errors.

"A small intra-cluster correlation like 0.02 can be ignored."

It is multiplied by $m - 1$. With 200 sessions per user, $\text{DEFF} = 1 + 199 \times 0.02 \approx 5$: the standard error is more than twice as large as the naive one.

$\rho = \tau^2/(\tau^2 + \sigma^2)$; $\text{DEFF} = 1 + (m-1)\rho$; $n_{\text{eff}} = n/\text{DEFF}$; naive SE too small by $\sqrt{\text{DEFF}}$.

Strata (some units from every group) reduce variance; clusters (a few whole groups) increase it but save cost.

Trap: count independent units, not rows.

Quick check: 300 users with 10 sessions each, ρ = 0.1. What is the effective sample size of the 3,000 sessions?

$\text{DEFF} = 1 + 9 \times 0.1 = 1.9$, so $n_{\text{eff}} = 3000/1.9 \approx 1579$. Standard errors that treat all 3,000 sessions as independent are too small by $\sqrt{1.9} \approx 1.38$.

Independence and the IID assumption: what it means and when it fails core

Almost every formula in this guide quietly assumes that the observations are IID: independent and identically distributed. Picture rolling the same fair die again and again. "Identically distributed": every roll uses the same die (the same distribution, no drift, no different machines mixed in). "Independent": each roll ignores the previous ones (knowing one result tells you nothing new about another).

Real data break this all the time. Yesterday's orders predict today's (time dependence). The same user's sessions resemble each other (repeated measures). Customers of the same store are alike (clusters). A price change shifts everything after it (not identical any more). When IID fails, the usual standard errors, tests and intervals can be badly wrong, usually too confident.

Three ways to say it:

  • Picture: the same die, rolled fresh each time.
  • Numbers: 1,000 sessions from 100 users with ρ = 0.3 behave like about 270 independent observations, so the naive standard error is about 1.9 times too small.
  • Slogan: same recipe, no memory.

An analyst computes the average session value from 1,000 sessions: 100 users with 10 sessions each. Sessions of the same user are correlated with $\rho = 0.3$.

  1. Treating sessions as IID: $SE = s/\sqrt{1000}$.
  2. The sessions are clustered by user: $\text{DEFF} = 1 + (10 - 1) \times 0.3 = 3.7$.
  3. Effective sample size $1000/3.7 \approx 270$. The real standard error is $\sqrt{3.7} \approx 1.92$ times the naive one.
  4. So a "95%" interval built from the naive SE covers the truth much less often than 95% of the time (a ±1.96 naive-SE interval is only a ±1.02 real-SE interval: coverage about 69%).
  5. Fix: average per user first (100 independent user means), or use cluster-robust standard errors by user.

Random variables $X_1, \dots, X_n$ are IID if

  • identically distributed: each $X_i$ has the same distribution $F$ (same mean, same variance, same shape);
  • independent: their joint distribution factorizes, $p(x_1, \dots, x_n) = \prod_i p(x_i)$: learning some of them does not change the distribution of the others (Chapter 4.3).

IID is what gives $Var(\bar X) = \sigma^2/n$, the LLN and CLT in their simple forms, and a likelihood that is a product $\prod_i p(x_i\mid\theta)$. It says nothing about the shape of $F$ (IID data need not be Normal).

SituationWhat failsTypical consequence / fix
Time series (daily orders)independence (autocorrelation); identical (trend, seasonality)SEs too small; model the dependence, check residual ACF (Chapter 7.3)
Repeated users (many sessions each)independence within useranalyse per user or cluster-robust SEs
Clusters (stores, companies)independence within clusterdesign effect; hierarchical model
Drift / regime change (price change, new app version)identical distributionsplit the periods; model the change
Interference (one user's treatment affects another)independence across unitscluster randomization; SUTVA (Chapter 5.11)

A weaker, very common version: conditionally IID. In a Bayesian model $y_i \mid \theta \sim p(y\mid\theta)$ independently: the observations are independent given θ, but dependent marginally, because each one carries information about the shared θ.

Why do we need it?

It is the hidden assumption behind standard errors, tests, confidence intervals, cross-validation and likelihoods. Checking it is how you know whether those numbers can be trusted.

Where is it used?

Every textbook test (z, t, chi-square), the bootstrap (resamples IID units), random K-fold CV (invalid for time series and grouped data: use TimeSeriesSplit or GroupKFold), and every likelihood written as a product over observations.

How is it used?

Before analysing, ask: what is the independent unit? Is anything shared between rows (user, store, day)? Is the process stable over the period? Then aggregate to the independent unit, model the dependence (time-series terms, random effects), or use robust SEs; check residual autocorrelation after fitting.

experiment user 1 (A) user 2 (B) user 3 (A) independent correlated Randomize users → users are (approximately) independent units. Sessions of one user share that user: they are not.
The randomization unit decides what is independent. In a user-randomized A/B test, the independent units are users; their sessions are clustered inside them and must not be counted as independent observations.

Pick a dataset. The plot shows its values in the order they were collected (standardized), and the readout gives the lag-1 autocorrelation (correlation of each value with the next one; near 0 for independent data) and the means of the first and second halves (similar for identically distributed data). Decide for yourself, then read the verdict. Press New data to see which features are stable.

100 users, each with $k$ sessions; ρ is the share of the variance that belongs to the user (how alike one user's sessions are). The histogram shows the average session value from 250 repeated experiments: its real spread. The red band is the ±1.96 naive SE band (treating all sessions as independent); the green band uses the per-user SE. Raise ρ or $k$: the histogram spills far outside the red band, so a naive 95% interval would miss far more than 5% of the time. At ρ = 0 the two bands agree.

"IID means the data are Normal."

IID is about independence and a common distribution. That distribution can be skewed, discrete or heavy-tailed. Normality is a separate assumption.

"My A/B test has a million page views, so the observations are IID."

If users were randomized, the independent units are users. Page views of the same user are correlated; treating them as independent gives standard errors that are too small (this "unit mismatch" is studied in Chapter 5.10).

"Shuffling the rows makes the data IID."

Shuffling changes the order, not the dependence: sessions of the same user are still alike, and shuffling a time series destroys the information in its order (Chapter 7.1). Random K-fold on such data leaks information.

In an A/B framework like yours, the Beta-Binomial model treats each user's conversion as a Bernoulli draw given the variant's rate θ: conditionally IID users. That is justified by randomizing users and counting each user once; feeding sessions or page views into the same likelihood would overstate the data and make the posterior too narrow. Segments break "identically distributed" across the whole population, which is exactly why your hierarchical model gives each segment its own parameter. In your forecasting model the daily observations are clearly not IID (trend, seasonality, holidays, autocorrelation); the model's job is to explain that structure so the remaining residuals are close to independent noise, and the residual ACF (Chapter 7.17) is how you check.

"The data are IID, so the test is valid." (said about a page-view table from a user-randomized test)

"The independent unit is the user, because that is what was randomized. I aggregate to one row per user, or use standard errors clustered by user."

Model answer: "IID means every observation comes from the same distribution and none carries information about another. It fails with time series, repeated measurements of the same user, clustered units and drift. When it fails, the variance of an average is not σ²/n; with clusters of size m and intra-cluster correlation ρ it is inflated by 1 + (m − 1)ρ, so naive intervals are too narrow. I match the analysis unit to the randomization unit, or model the dependence."

IID = identically distributed (same $F$) + independent ($p(x_1..x_n) = \prod p(x_i)$). Gives $Var(\bar X) = \sigma^2/n$ and product likelihoods.

Fails with time series, repeated users, clusters, drift, interference. Naive SE too small by $\sqrt{1 + (m-1)\rho}$ for clusters.

Trap: IID ≠ Normal; the independent unit is the randomization unit.

Quick check: are the monthly revenues of one company over 5 years IID?

No. Revenue usually has a trend and seasonality (the distribution changes over time, so not identical) and one month's level predicts the next (not independent). You need time-series methods (Guide 4), not formulas that assume IID.

Recap, cheat sheet and practice

  • Population → frame → sample → analysed data. A sample speaks only for its frame; a simple random sample gives every unit the same chance and an unbiased estimate.
  • Sampling variation is luck: estimates differ between random samples. Their distribution over repeated samples is the sampling distribution; its spread shrinks like $1/\sqrt n$ (standard errors: Chapter 5.5).
  • Selection bias (including sampling bias, non-response, survivorship, post-treatment filters) is systematic: MSE = bias² + variance, and more data shrinks only the variance.
  • Non-response bias ≈ Cov(response probability, outcome) / mean response probability: who answers matters, not how many.
  • Survivorship bias: filtering to survivors changes the mix and creates trends from nothing; follow cohorts from the start.
  • Stratified sampling samples inside every group and removes the between-group variance; cluster sampling samples whole groups and inflates the variance by $1 + (m-1)\rho$.
  • IID = same distribution + independence. It fails for time series, repeated users, clusters, drift and interference; then naive standard errors are too small. Analyse at the level of the randomization unit.

Cheat sheet

IdeaFormula / ruleMeaning / trap
Simple random sampleevery set of $n$ equally likely; inclusion prob. $n/N$unbiased for the frame
Sampling spread of $\hat p$$\sqrt{p(1-p)/n}$4× data → half the spread
Error decomposition$\text{MSE} = \text{bias}^2 + \text{variance}$bias does not shrink with $n$
Non-responsebias $= \mathrm{Cov}(\rho, y)/\bar\rho$low response rate ≠ bias
Selection after treatmentcompare arms on pre-treatment subsets only"cart users" flipped 5% → 6% into 25% → 20%
Stratified estimator$\sum_h W_h\bar y_h$, $Var = \sum_h W_h^2\sigma_h^2/n_h$weight by population shares
Neyman allocation$n_h \propto W_h\sigma_h$more sample where the spread is
Intra-cluster correlation$\rho = \tau^2/(\tau^2 + \sigma^2)$how alike units of one cluster are
Design effect$1 + (m-1)\rho$; $n_{\text{eff}} = n/\text{DEFF}$naive SE too small by $\sqrt{\text{DEFF}}$
IIDsame $F$, $p(x_1..x_n) = \prod p(x_i)$not about Normality
Code it · Python

import numpy as np
import pandas as pd
import statsmodels.formula.api as smf

rng = np.random.default_rng(0)

# 1) A population of 100,000 users: 30% mobile (convert 5%), 70% desktop (convert 12%)
N = 100_000
mobile = rng.random(N) < 0.30
converted = rng.random(N) < np.where(mobile, 0.05, 0.12)
pop = pd.DataFrame({"mobile": mobile, "converted": converted})
print("true rate", round(pop.converted.mean(), 4))                      # 0.0985 (expected 0.099)

# 2) Simple random sample vs a desktop-heavy "newsletter" sample (desktop 4x more likely to be picked)
w = np.where(pop.mobile, 1.0, 4.0)
for n in [100, 10_000]:
    srs = [pop.converted.sample(n, random_state=s).mean() for s in range(300)]
    biased = [pop.converted.sample(n, weights=w, random_state=s).mean() for s in range(300)]
    print(n, "SRS mean", round(np.mean(srs), 4), "sd", round(np.std(srs), 4),
          "| biased mean", round(np.mean(biased), 4), "sd", round(np.std(biased), 4))
# 100:    SRS mean 0.0996, sd 0.0306 | biased mean 0.1124, sd 0.0328
# 10000:  SRS mean 0.0984, sd 0.0031 | biased mean 0.1124, sd 0.0029: the bias (+0.014) does not shrink

# 3) Non-response: respondent mean = true mean + Cov(rho, y) / mean(rho)
sat = rng.random(N) < 0.60
rho = np.where(sat, 0.30, 0.10)                                         # response probabilities
responds = rng.random(N) < rho
print("respondents satisfied:", round(sat[responds].mean(), 3))         # 0.819 (truth 0.60)
print("formula:", round(sat.mean() + np.cov(rho, sat, bias=True)[0, 1] / rho.mean(), 3))   # 0.818

# 4) Stratified sampling: small shops (80%, mean 100, sd 20) and big shops (20%, mean 500, sd 100)
shops = pd.DataFrame({"big": rng.random(N) < 0.2})
shops["spend"] = np.where(shops.big, rng.normal(500, 100, N), rng.normal(100, 20, N))
W = shops.big.value_counts(normalize=True)                              # population shares of the strata
mu = shops.spend.mean()
srs_err, strat_err = [], []
for s in range(1000):
    srs_err.append(shops.spend.sample(100, random_state=s).mean() - mu)
    st = shops.groupby("big").sample(frac=100 / N, random_state=s)      # about 80 small + 20 big
    strat_err.append((st.groupby("big").spend.mean() * W).sum() - mu)   # weight stratum means by W
print("SE SRS", round(np.std(srs_err), 1), "| SE stratified", round(np.std(strat_err), 1))   # 16.6 vs 4.6 (formulas: 16.7 and 4.8)

# 5) Cluster sampling: 50 stores x 40 customers, intra-cluster correlation 0.2
K, m, icc = 50, 40, 0.2
tau, sig = np.sqrt(icc), np.sqrt(1 - icc)                               # total variance 1
store = np.repeat(np.arange(K), m)
y = (tau * rng.normal(size=K))[store] + sig * rng.normal(size=K * m)
est_srs = [y[rng.choice(K * m, 200, replace=False)].mean() for _ in range(3000)]
est_clu = [y[np.isin(store, rng.choice(K, 5, replace=False))].mean() for _ in range(3000)]
print("variance ratio cluster/SRS:", round(np.var(est_clu) / np.var(est_srs), 1),
      "| DEFF formula 1 + (m-1)*icc =", 1 + (m - 1) * icc)               # 10.2 vs 8.8: these 50 stores happen to differ
# a bit more than average (variance of store means 0.27 vs 0.22 expected); over many populations it averages near 8.8

# 6) Repeated users: naive SE vs SE clustered by user (100 users x 10 sessions, rho = 0.3)
users = np.repeat(np.arange(100), 10)
val = (np.sqrt(0.3) * rng.normal(size=100))[users] + np.sqrt(0.7) * rng.normal(size=1000)
df = pd.DataFrame({"user": users, "value": val})
naive = smf.ols("value ~ 1", df).fit()
clustered = smf.ols("value ~ 1", df).fit(cov_type="cluster", cov_kwds={"groups": df.user})
print("naive SE", round(naive.bse.iloc[0], 4), "| clustered SE", round(clustered.bse.iloc[0], 4),
      "| ratio", round(clustered.bse.iloc[0] / naive.bse.iloc[0], 2), "vs sqrt(DEFF) =", round(np.sqrt(1 + 9 * 0.3), 2))
# naive SE 0.0321 | clustered SE 0.0643 | ratio 2.01 vs sqrt(DEFF) = 1.92: the naive SE is about half the honest one
Test yourself

1. A sampling method gives estimates centred 2 points above the truth. You increase the sample from 1,000 to 100,000. What happens?

Random spread falls like $1/\sqrt n$ (√100 = 10). Bias comes from how units are selected and is unaffected by $n$. The LLN makes the estimate converge to the centre of its distribution, which here is the wrong value.

2. 50% of customers are satisfied. Satisfied customers answer a survey with probability 0.4, unsatisfied ones with 0.2. What share of respondents is satisfied, on average?

Respondents: $0.5 \times 0.4 = 0.2$ satisfied and $0.5 \times 0.2 = 0.1$ unsatisfied, so $0.2/0.3 \approx 66.7\%$.

3. Your sample is grouped by store. Which design removes the between-store differences from the error of the overall average?

Stratifying by store fixes the store mix by design, so only within-store variation remains. Cluster sampling does the opposite: which stores you happen to pick drives the error.

4. 3,000 observations come in clusters of 21 with intra-cluster correlation 0.1. What is the effective sample size?

$\text{DEFF} = 1 + 20 \times 0.1 = 3$, so $n_{\text{eff}} = 3000/3 = 1000$.

5. "Users who have been with us for 2+ years spend twice as much as new users, so loyalty programmes raise spending." What is the most direct problem?

Only users who stayed are in the 2+ years group, and staying is related to spending. The difference can arise even if nobody's spending ever changes.

6. Which dataset is closest to IID?

Distinct randomized users in one arm are (approximately) independent draws from the same distribution. Page views repeat users; daily sales have trend, seasonality and autocorrelation; pooling before and after a redesign mixes two distributions.

Practice problems

A. 40% of users are mobile (convert 4%) and 60% desktop (convert 10%). Your sample is 20% mobile and 80% desktop. Compute the true rate, the sample's expected rate, and fix it by re-weighting the segments.

True: $0.4 \times 0.04 + 0.6 \times 0.10 = 0.016 + 0.06 = 7.6\%$. Sample: $0.2 \times 0.04 + 0.8 \times 0.10 = 0.008 + 0.08 = 8.8\%$ (bias +1.2 points). Re-weight each segment's rate by its population share (post-stratification): $0.4 \times 0.04 + 0.6 \times 0.10 = 7.6\%$. Equivalently, weight each mobile user by $0.4/0.2 = 2$ and each desktop user by $0.6/0.8 = 0.75$. This works because the bias came from a segment mix you can measure.

B. 70% of customers are satisfied. Satisfied ones respond with probability 0.5, unsatisfied ones with 0.25. Find the respondent estimate and check it with the covariance formula.

$\bar\rho = 0.7 \times 0.5 + 0.3 \times 0.25 = 0.35 + 0.075 = 0.425$. Respondent share satisfied: $0.35/0.425 \approx 82.4\%$ (bias +12.4 points). Formula: $\mathrm{Cov}(\rho, y) = 0.7 \times 0.5 - 0.425 \times 0.7 = 0.35 - 0.2975 = 0.0525$; $0.0525/0.425 \approx 0.124$. ✓

C. Two equal-size strata with means 20 and 80, each with sd 10. Compare the SE of the mean for an SRS of 50 and for a proportional stratified sample (25 + 25).

Population mean 50. Within variance $= 100$; between $= 0.5 \times 30^2 + 0.5 \times 30^2 = 900$; total 1000. SRS: $SE = \sqrt{1000/50} = \sqrt{20} \approx 4.47$. Stratified: $Var = 0.5^2 \times 100/25 + 0.5^2 \times 100/25 = 1 + 1 = 2$, $SE \approx 1.41$. The between-strata variance, 90% of the total, is removed.

D. A survey interviews all 25 students in each of 20 randomly chosen classes (500 students). The intra-class correlation is 0.15. What are the design effect, the effective sample size and the SE inflation?

$\text{DEFF} = 1 + 24 \times 0.15 = 4.6$; $n_{\text{eff}} = 500/4.6 \approx 109$; standard errors are $\sqrt{4.6} \approx 2.14$ times those of a simple random sample of 500 students. An analysis that ignores the classes would report intervals less than half as wide as they should be.

E. Explain to an interviewer: "Our A/B analysis used 2 million page views as rows and got p = 0.001. What would you check first?"

"What was randomized. If users were randomized, the independent units are users, and page views of the same user are correlated, so the page-view-level standard error is too small by roughly $\sqrt{1 + (m-1)\rho}$, where m is page views per user and ρ their correlation. With m around 20 and ρ = 0.1, that is a factor of about 1.7, which can turn p = 0.001 into something unremarkable. I would aggregate to one row per user (or use the delta method or cluster-robust standard errors by user) and rerun. I would also check for a sample ratio mismatch and for any filter defined after assignment."

F. A demand-forecasting dataset contains only products still in the catalogue today. Why can that bias the model, and what would you do?

Discontinued products are usually the ones whose demand declined; removing them is a survivorship filter, so the training data over-represent stable or growing products. Trends and the typical size of declines are underestimated, and forecasts for new products may be too optimistic. Fix: rebuild the history as it was known at each point in time (include products that were later discontinued), or at least evaluate on all products that existed at each forecast origin.

Chapter 5.5 · Syllabus Module 13

Standard error and sampling distributions

You ran one experiment and got one number: a 12% conversion rate, an average of 50 orders a day. If you could run it again, the number would come out a little different. How different? The standard error answers exactly that question. Every p-value, every confidence interval, every sample-size plan and every "is this lift real?" argument is built on top of it.

  • Tell the standard deviation (how much single data values spread) from the standard error (how much an estimate would spread over repeated samples), and never mix them up again
  • Derive $SE(\bar X) = \sigma/\sqrt n$ from the variance rules of Guide 1, and use the √n law: four times the data for half the error
  • Compute the SE of a proportion, $\sqrt{p(1-p)/n}$, and of a difference, $\sqrt{SE_1^2 + SE_2^2}$, for your checkout example
  • Understand the estimated SE ($s$ in place of $\sigma$) and why it adds extra uncertainty when $n$ is small
  • See the full sampling distribution behind the SE: its centre, its width and its shape
  • Use the bootstrap to get an SE from one sample when no formula exists, and know when it fails

What we need from earlier chapters: variance and its rules $Var(aX) = a^2Var(X)$ and "variances of independent pieces add" (Chapter 4.5); the Bernoulli variance $p(1-p)$ (Chapter 4.7); the Central Limit Theorem (Chapter 4.13); the words statistic, estimator and estimate and the idea that an estimator has a sampling distribution (Chapter 5.1); random samples and IID data (Chapter 5.4). Reminder of the words: a statistic is any number computed from the sample (it has nothing to do with the school subject); an estimator is a statistic used to guess an unknown number $\theta$; an estimate is the value it gave on your data.

SD vs SE: the spread of the data and the spread of an estimate core

Think about the heights of people in a city. Pick one random person: they could easily be 155 cm or 185 cm. Single people vary a lot. Now pick 100 random people and take their average height. Do it again with another 100 people. The two averages will be very close, maybe 169.8 cm and 170.3 cm. Averages vary much less than single people.

So there are two different "wobbles", and they get two different names:

  • The standard deviation (SD) measures how much single data values wobble around their mean.
  • The standard error (SE) measures how much an estimate (like the average of 100 people) would wobble if you repeated the whole sampling again and again.

The word "error" here does not mean a mistake. It means "how far off the estimate typically is, just because of which sample you happened to get".

Three ways to say it:

  • Picture: the SD is the width of the crowd; the SE is the width of the pile of averages you would get from many crowds.
  • Numbers: daily orders have SD 10; the average of 25 days has SE $10/\sqrt{25} = 2$.
  • Slogan: the SD describes your data; the SE describes your estimate.

Orders per day. A shop's daily orders come from a process with mean $\mu = 50$ and standard deviation $\sigma = 10$. You record $n = 25$ days and report the average.

  1. Question 1: how different is one day from a typical day? That is the SD of the data: about 10 orders. Roughly 95% of single days land within $50 \pm 2\times 10$, so between 30 and 70 orders.
  2. Question 2: if I recorded another 25 days, how different would the new average be? That is the SE of the average. (Next section shows why.) $SE = \sigma/\sqrt n = 10/\sqrt{25} = 10/5 = 2$ orders.
  3. Roughly 95% of 25-day averages land within $50 \pm 2\times 2$, so between 46 and 54 orders.
  4. Same process, same data, two very different answers: 10 for a day, 2 for an average. The SE is $\sqrt{25} = 5$ times smaller than the SD.

The standard deviation describes one random value $X$: $\;SD(X) = \sqrt{Var(X)} = \sigma$. From data we estimate it by the sample standard deviation $s$.

The standard error of an estimator $\hat\theta$ is the standard deviation of its sampling distribution (the distribution of $\hat\theta$ over repeated samples of the same size from the same process):

$$SE(\hat\theta) = \sqrt{Var(\hat\theta)}.$$
  • For the sample mean of $n$ independent values: $SE(\bar X) = \sigma/\sqrt n$ (derived in the next section).
  • The SD does not shrink when you collect more data (more data only gives a better estimate $s$ of the same $\sigma$). The SE does shrink, because averages of more values are steadier.
  • Every estimator has its own SE: the SE of a mean, of a proportion, of a difference, of a median, of a regression slope. "The SE" alone is incomplete; always say "the SE of what".
  • We usually do not know $\sigma$, so we compute an estimated SE such as $s/\sqrt n$ (later in this chapter).
Why do we need it?

A single estimate cannot tell you how much to trust it. The SE can: it is the typical distance between your estimate and the truth caused by the luck of the sample. Without it, a 2-point lift could be a breakthrough or pure noise, and you could not tell which.

Where is it used?

Every z-test and t-test divides by an SE (Chapter 5.6); every confidence interval is "estimate ± about 2 SE" (Chapter 5.8); sample-size planning for A/B tests; error bars on dashboards; the "std err" column of every regression table.

How is it used?

Compute the estimate, then its SE (a formula like s / np.sqrt(n), or the bootstrap). Report both: "50.3 ± 2.0 orders (SE)". Compare differences with their SE: a gap of 1 SE is ordinary noise, a gap of 3 SE is rare luck.

μ = 50 3040506070 orders SD = 10 (one day) SE = 10/√4 = 5 (average of 4 days) single days:wide averages of 4 days:half as wide, twice as tall
Blue: how single days are spread (SD = 10). Orange: how the average of 4 days would be spread over many repeats (SE = 5). Both curves are centred on the same μ; averaging only makes the spread narrower.

Top: how single days are spread (blue curve, SD = 10) and the $n$ days of your current sample (blue dots). Bottom: the averages from many repeated samples of $n$ days (orange histogram). Move the slider from 1 to 64 and watch the bottom pile get narrower while the top curve does not change at all. At $n = 16$ the SE is $10/4 = 2.5$; at $n = 64$ it is $10/8 = 1.25$. Switch to skewed days: the days are lopsided, but the averages still form a narrow, nearly symmetric pile.

"If I collect more data, the standard deviation goes down."

The SD of the data is a property of the process (how different days are). More data gives a more accurate estimate $s$ of it, but $\sigma$ itself does not shrink. What shrinks is the SE of your estimates.

"SE = 2 means my estimate is off by 2."

It means a typical repeat of the experiment would land about 2 away. Your actual error on this sample is unknown: it could be 0.3 or 4.

"These error bars show the uncertainty." (without saying which kind)

Always say whether bars are ±SD (spread of the data) or ±SE (uncertainty of the mean). With 100 points they differ by a factor of 10.

In an A/B framework like yours, a revenue-per-user metric has a large SD (most users spend 0, a few spend a lot), but the SE of each variant's average revenue is that SD divided by $\sqrt{n}$: this is why big experiments can detect small changes in a very noisy metric. In your forecasting model, $\sigma$ in the likelihood is an SD (how much a single day scatters around the fit), while the posterior standard deviation of a coefficient such as a holiday effect plays the role of an SE: how uncertain the estimate of that effect is.

"The standard error of the data is 10."

"The standard deviation of the data is 10. The standard error of the mean of 25 values is $10/\sqrt{25} = 2$."

Model answer: "The SD describes how individual observations vary. The SE is the standard deviation of an estimator's sampling distribution: how much the estimate would vary across repeated samples. For a mean of $n$ independent values, $SE = \sigma/\sqrt n$, so the SE shrinks with more data while the SD does not."

SD = spread of single values (data). SE = spread of an estimate over repeated samples = SD of its sampling distribution.

For a mean: $SE(\bar X) = \sigma/\sqrt n$. More data: SE shrinks, SD does not.

Trap: always say "SE of what", and label error bars as SD or SE.

Quick check: daily orders have SD 12. What is the SE of a 36-day average? And the SD of the orders on day 37?

SE $= 12/\sqrt{36} = 12/6 = 2$ orders. Day 37 is a single day, so its spread is still the SD: about 12 orders. Averaging helps the estimate, not single days.

Why $SE(\bar X) = \sigma/\sqrt n$: noise partly cancels when you average core

Imagine adding up 4 days of orders. Some days are above 50, some below. When you add them, the ups and downs partly cancel. The total does get more spread out than one day, but not 4 times more: only $\sqrt 4 = 2$ times more, because independent surprises do not all point the same way.

Then you divide the total by 4 to get the average. Dividing shrinks the spread by a full factor of 4. Grows by 2, shrinks by 4: overall the average is $2/4 = 1/2$ as spread out as one day. In general: grows by $\sqrt n$, shrinks by $n$, so the net effect is $\sqrt n / n = 1/\sqrt n$.

Three ways to say it:

  • Picture: $n$ people pull a rope in random directions; the total pull grows only like $\sqrt n$ because the pulls partly cancel.
  • Numbers: with $\sigma = 10$ and $n = 4$: the total has SD $20$, the average has SD $20/4 = 5 = 10/\sqrt 4$.
  • Slogan: variances add, so the noise of a total grows like $\sqrt n$; dividing by $n$ leaves $1/\sqrt n$.

Four days of orders, each with $\sigma = 10$ (variance $10^2 = 100$), independent of each other.

  1. Variance of the total $T = X_1 + X_2 + X_3 + X_4$: independent variances add, so $Var(T) = 100 + 100 + 100 + 100 = 400$.
  2. SD of the total: $\sqrt{400} = 20$. (Not $4 \times 10 = 40$: SDs do not add, variances do.)
  3. The average is $\bar X = T/4$. Dividing by 4 multiplies the variance by $1/4^2 = 1/16$: $Var(\bar X) = 400/16 = 25$.
  4. $SE(\bar X) = \sqrt{25} = 5$, which is exactly $10/\sqrt 4 = 10/2$.

Check with the tiny shop of Chapter 5.1 (2, 4 or 6 boxes, equally likely): one day has variance $8/3$, and listing all 9 samples of 2 days gave $Var(\bar X) = 4/3$. The rule says $\sigma^2/n = (8/3)/2 = 4/3$. ✓

Let $X_1, \dots, X_n$ be independent random variables, each with the same variance $\sigma^2$ (for example IID draws). Using the rules of Chapter 4.5:

$$Var\Big(\sum_{i=1}^n X_i\Big) = \sum_{i=1}^n Var(X_i) = n\sigma^2, \qquad Var(\bar X) = Var\Big(\frac{1}{n}\sum_i X_i\Big) = \frac{1}{n^2}\, n\sigma^2 = \frac{\sigma^2}{n},$$ $$SE(\bar X) = \sqrt{Var(\bar X)} = \frac{\sigma}{\sqrt n}.$$
  • Needed: independence (so the covariance terms are zero) and a finite variance $\sigma^2$.
  • Not needed: Normal data. The formula holds for any shape. (That the shape of $\bar X$'s distribution becomes Normal is a separate fact, the CLT.)
  • With correlation: $Var(\sum X_i) = n\sigma^2 + \sum_{i \ne j} Cov(X_i, X_j)$. Positive correlation makes the true SE larger than $\sigma/\sqrt n$.
Why do we need it?

It tells us exactly how much an average can be trusted, before we even collect the data. It turns "more data is better" into a precise rule we can plan with: how many users, days or runs we need for a given precision.

Where is it used?

The z-test and t-test denominators, confidence intervals for means, A/B test sample-size formulas, Monte Carlo error of any simulated average (for example an ELBO estimate from $S$ samples, whose noise falls like $1/\sqrt S$), and mini-batch gradient noise in SGD.

How is it used?

Estimate $\sigma$ (by $s$), divide by $\sqrt n$: x.std(ddof=1) / np.sqrt(len(x)), which is also scipy.stats.sem(x). Before trusting it, check independence: are the rows separate users, or the same user many times, or consecutive days?

σ² σ² ⋮ σ² n days add up total: Var = nσ² SD = σ√n (grows) divide by n Var × 1/n² average: Var = σ²/n SE = σ/√n (shrinks) needs independent days
Variance bookkeeping. Independent variances add, so the total's variance is $n\sigma^2$. Dividing by $n$ divides the variance by $n^2$. What is left is $\sigma^2/n$, and its square root is the standard error.

Each simulated "experiment" adds up $n$ days whose orders wobble with $\sigma = 10$. Top: the total minus its expected value ($n \times 50$). Bottom: the average minus 50. Move $n$ from 1 to 25: the top pile widens like $10\sqrt n$ (from 10 to 50) while the bottom pile narrows like $10/\sqrt n$ (from 10 to 2). The readout shows the variance bookkeeping at each step.

"$SE = \sigma/n$: the average of 100 values is 100 times more precise."

It is $\sigma/\sqrt n$: 100 values make the average only $\sqrt{100} = 10$ times more precise. Noise cancels, but slowly.

"The formula $\sigma/\sqrt n$ needs Normal data."

It needs independence and a finite variance, nothing about shape. The Normal shape of the sampling distribution comes from the CLT and needs a large enough $n$.

"I have 30 days of 1 000 sessions each, so $n = 30\,000$."

Only if the 30 000 rows are independent. Sessions of the same user, or days next to each other, are positively correlated, and then the true SE is larger than $\sigma/\sqrt{30\,000}$. The rule needs the right $n$: the number of independent units (Chapter 4.13; unit mismatch in Chapter 5.10).

In your forecasting model the residuals of neighbouring days are often correlated (yesterday's surprise tells you something about today's). Then the variance of an average residual is bigger than $\sigma^2/n$, and any SE computed as if days were independent is too small (autocorrelation is the topic of Chapter 7.3). In an A/B framework like yours, the safe unit is the user (the unit of randomization): computing an SE over page views treats repeated views of one person as independent, which they are not.

$Var(\sum X_i) = n\sigma^2$ (independent) $\Rightarrow$ $Var(\bar X) = \sigma^2/n$ $\Rightarrow$ $SE(\bar X) = \sigma/\sqrt n$.

Needs: independence + finite variance. Does not need: Normal data.

Trap: positively correlated rows (same user, nearby days) make the real SE bigger than $\sigma/\sqrt n$.

Quick check: one day has variance 64. What are the SD of a 16-day total and the SE of a 16-day average?

Total: $Var = 16 \times 64 = 1024$, SD $= 32$ ($= 8\sqrt{16}$). Average: $Var = 1024/16^2 = 4$, SE $= 2$ ($= 8/\sqrt{16}$).

The √n law: four times the data for half the error core

Because the SE is $\sigma/\sqrt n$ and not $\sigma/n$, precision is expensive. To make your estimate twice as precise (half the SE), you do not need twice the data. You need four times the data. To make it ten times as precise, you need a hundred times the data.

It is like climbing a hill that gets steeper and steeper: the first users you add help a lot, later users help less and less.

Three ways to say it:

  • Picture: the SE curve drops fast at first and then flattens; every halving step is four times as long as the one before.
  • Numbers: 400 users give SE 1.5 points; 1 600 users give 0.75; 6 400 users give 0.375.
  • Slogan: half the noise costs four times the data.

Conversion at about 10%. One user is a 0/1 value with SD $\sqrt{0.1 \times 0.9} = 0.3$ (the reason is the next section). So the SE of the conversion rate is $0.3/\sqrt n$.

  1. $n = 100$: $SE = 0.3/\sqrt{100} = 0.3/10 = 0.030$, i.e. 3 percentage points.
  2. $n = 400$ (4×): $SE = 0.3/20 = 0.015$, i.e. 1.5 points. Half.
  3. $n = 1600$ (4× again): $SE = 0.3/40 = 0.0075$, i.e. 0.75 points. Half again.
  4. $n = 6400$: $SE = 0.3/80 = 0.00375$, i.e. 0.375 points.
  5. Turn it around: to reach a target SE, solve $\sigma/\sqrt n = SE$ for $n$: $n = (\sigma/SE)^2$. For a target of 0.5 points ($0.005$): $n = (0.3/0.005)^2 = 60^2 = 3600$ users.

Because $SE(n) = \sigma/\sqrt n$, multiplying the sample size by $k$ divides the SE by $\sqrt k$:

$$SE(kn) = \frac{\sigma}{\sqrt{kn}} = \frac{SE(n)}{\sqrt k}, \qquad n_{\text{needed}} = \left(\frac{\sigma}{SE_{\text{target}}}\right)^2 .$$
  • $k = 4$ halves the SE; $k = 100$ divides it by 10; $k = 2$ divides it only by $\sqrt 2 \approx 1.41$ (to about 71%).
  • The needed $n$ grows with the square of the precision you ask for. This is why detecting a lift half as big needs about four times as many users (the full sample-size formula is in Chapter 5.7).
  • The rule assumes independent units and the same $\sigma$. It says nothing about bias: a biased sample stays biased however big it gets.
Why do we need it?

Data costs time and money: every extra day of an experiment delays a decision. The √n law tells you what extra precision you buy with extra data, so you can decide whether waiting is worth it.

Where is it used?

Sample-size and test-duration planning for A/B tests (Chapter 5.10), choosing how many Monte Carlo samples to draw (posterior predictive draws, ELBO samples), how many bootstrap resamples to use, and how many backtest windows to average over.

How is it used?

Decide the precision you need (a target SE), estimate $\sigma$ from past data, and compute $n = (\sigma / SE_{target})^2$. Then divide by daily traffic to get a duration. If the answer is huge, reduce $\sigma$ (for example with CUPED, Chapter 5.12) instead of waiting longer.

SE 3 pts 1.5 pts 0.75 pts 0.375 pts n = 100 n = 400 n = 1 600 n = 6 400 ×4 users ×4 users ×4 users each step: four times the users, half the standard error
The √n law for a 10% conversion rate (one user's SD is 0.3). Every halving of the SE costs four times as many users as the step before.

The blue curve is $SE = 0.3/\sqrt n$ for a 10% conversion rate. Drag the blue dot along the curve to choose $n$. The green dot is at $4n$: its SE is always exactly half. Notice how the halving steps get longer and longer to the right. Then move the target SE slider: the purple line shows the precision you want, and the readout gives the users you need, $n = (0.3/SE)^2$. Try a target of 0.3 points: 10 000 users.

"Doubling the users halves the noise."

Doubling divides the SE by $\sqrt 2 \approx 1.41$: the SE drops to about 71% of its old value. Halving needs four times the users.

"With enough data the SE goes to zero, so the estimate becomes correct."

The SE measures only the luck of the draw. A bias (a sampling bias, a broken tracking event, a non-random split) does not shrink with $n$ (Chapter 5.1, Chapter 5.4). A huge sample just gives you a very precise wrong answer.

In an A/B framework like yours the same law governs the posterior: with a weak prior, the posterior standard deviation of a conversion rate also shrinks like $1/\sqrt n$. So a segment with 4 times fewer users has a posterior about twice as wide. This is one reason hierarchical partial pooling helps small segments: it borrows strength from the others instead of waiting for 4× more data (Chapter 6.6).

$SE(kn) = SE(n)/\sqrt k$: 4× the data → half the SE; 100× → a tenth.

Planning: $n = (\sigma / SE_{target})^2$.

Trap: more data shrinks noise, never bias.

Quick check: a metric has SE 2.0 with 900 users. How many users give SE 0.5?

You want the SE 4 times smaller, so you need $4^2 = 16$ times the users: $16 \times 900 = 14\,400$. (Check: $\sigma = 2.0 \times \sqrt{900} = 60$ and $(60/0.5)^2 = 120^2 = 14\,400$.)

The SE of a proportion: $\sqrt{p(1-p)/n}$ core

A conversion rate looks like a new kind of number, but it is just an average of 0s and 1s. Write 1 for every user who bought and 0 for every user who did not. The average of that list is exactly the share of buyers. So everything we know about averages applies: the SE is "SD of one user" divided by $\sqrt n$.

And the SD of one 0/1 user is known from Chapter 4.7: a Bernoulli variable has variance $p(1-p)$. It is largest at $p = 0.5$ (a coin flip is the most uncertain) and small near 0 or 1 (almost everyone does the same thing).

Three ways to say it:

  • Picture: a conversion rate is the average height of a row of 0- and 1-blocks.
  • Numbers: 50 buyers out of 500: $\hat p = 0.10$, one user's SD $= \sqrt{0.1 \times 0.9} = 0.3$, $SE = 0.3/\sqrt{500} = 0.0134$.
  • Slogan: a proportion is a mean, so its SE is $\sqrt{p(1-p)}/\sqrt n$.

Your checkout test. Old checkout (A): 50 of 500 users bought. New checkout (B): 60 of 500 bought.

  1. Group A: $\hat p_A = 50/500 = 0.10$. One user's variance: $0.10 \times 0.90 = 0.09$, so one user's SD is $\sqrt{0.09} = 0.30$.
  2. $SE(\hat p_A) = \sqrt{0.09/500} = \sqrt{0.00018} = 0.0134$, i.e. 1.34 percentage points.
  3. Group B: $\hat p_B = 60/500 = 0.12$. Variance: $0.12 \times 0.88 = 0.1056$; SD $= \sqrt{0.1056} = 0.325$.
  4. $SE(\hat p_B) = \sqrt{0.1056/500} = \sqrt{0.0002112} = 0.0145$, i.e. 1.45 points.
  5. Meaning: if you reran the old checkout with another 500 users, its rate would typically land about 1.3 points away from this 10%: maybe 8.7% or 11.4%.

Let $X_1, \dots, X_n$ be independent 0/1 values with $P(X_i = 1) = p$, and $\hat p = \frac1n\sum X_i$ the sample proportion. Since $Var(X_i) = p(1-p)$,

$$Var(\hat p) = \frac{p(1-p)}{n}, \qquad SE(\hat p) = \sqrt{\frac{p(1-p)}{n}}, \qquad \widehat{SE}(\hat p) = \sqrt{\frac{\hat p(1-\hat p)}{n}}.$$
  • $p$ is the unknown true rate; in practice we plug in $\hat p$ (the hat on SE means "estimated").
  • $n$ is the number of trials (users who could convert), not the number of conversions.
  • The largest SE for a given $n$ is at $p = 0.5$: $\sqrt{0.25/n} = 0.5/\sqrt n$.
  • Relative SE $= SE/p = \sqrt{(1-p)/(np)}$: for rare events (small $p$) it is large, even though the absolute SE is small.
  • The Normal approximation to $\hat p$ needs enough successes and failures (a common rule of thumb: $np \ge 10$ and $n(1-p) \ge 10$). Near 0 or 1, use better intervals (Wilson, Chapter 5.8).
Why do we need it?

Most product metrics are rates: conversion, click-through, retention, churn. Without the SE of a rate you cannot say whether 10% vs 12% is a real difference or a lucky draw, and you cannot plan how many users an experiment needs.

Where is it used?

The two-proportion z-test of A/B testing (Chapter 5.6), Wald confidence intervals for rates, sample-size calculators for conversion experiments, error bars on click-through dashboards, and polling margins of error ("± 3 points").

How is it used?

Count successes $k$ and trials $n$, compute p = k / n and se = np.sqrt(p * (1 - p) / n). Report "10.0% ± 1.3 points (SE)". For planning, use the expected rate before the test.

Top: the curve $p(1-p)$, the variance of one 0/1 user. Drag the purple dot to choose the true rate $p$: the noise is biggest at 50% and small near 0% and 100%. Bottom: 1 000 simulated experiments of $n$ users each; the orange pile is the spread of their observed rates $\hat p$, and the bracket is the formula's SE. Move $n$ from 500 to 2 000: the pile halves. Press Rare event (1%): the absolute SE is tiny, but the readout's relative SE is large.

"A 1% conversion rate is less noisy than a 10% one, because its SE is smaller."

The absolute SE is smaller, but relative to the rate it is much bigger. With 500 users, 1% ± 0.44 points is a 44% relative wobble; 10% ± 1.34 points is only 13%. Rare events need many more users to measure well.

"0 conversions out of 40 gives $\hat p = 0$ and SE $= 0$, so the rate is surely 0."

The plug-in SE $\sqrt{\hat p(1-\hat p)/n}$ breaks down at 0 and 1. Zero out of 40 is perfectly possible with a true rate of 3%. Use a Wilson interval (Chapter 5.8) or a Beta posterior.

"$n$ is the number of conversions."

$n$ is the number of users who could convert (the trials). 50 conversions out of 500 users means $n = 500$.

In a Beta-Binomial model like yours, a flat Beta(1, 1) prior with 50 conversions out of 500 gives the posterior Beta(51, 451). Its standard deviation is $0.01347$, almost exactly the classical $SE = \sqrt{0.1 \times 0.9/500} = 0.01342$. This is not a coincidence: with a weak prior and plenty of data, the posterior sd of a rate is close to its standard error. They answer different questions (spread of your belief about $\theta$ vs spread of $\hat p$ over repeated experiments), but they are about the same size, which is why Bayesian and classical A/B results usually agree for large tests.

A proportion is a mean of 0/1 values: $SE(\hat p) = \sqrt{p(1-p)/n}$ (plug in $\hat p$).

Checkout: $\sqrt{0.1 \times 0.9/500} = 0.0134$ and $\sqrt{0.12 \times 0.88/500} = 0.0145$.

Trap: rare events have small absolute but large relative SE; the plug-in SE fails at $\hat p = 0$ or $1$.

Quick check: 200 of 1 000 users clicked. What is the SE of the click rate?

$\hat p = 0.2$, so $\sqrt{0.2 \times 0.8 / 1000} = \sqrt{0.00016} = 0.0126$, about 1.26 percentage points.

The SE of a difference: $\sqrt{SE_1^2 + SE_2^2}$ core

In an A/B test you do not care about each group's rate on its own. You care about the gap between them: $\hat p_B - \hat p_A$. Both rates wobble, and they wobble independently (different users). So the gap wobbles too, and more than either rate alone: sometimes A's luck and B's luck push the gap the same way.

But the two wobbles do not simply add up, because sometimes they push in opposite directions and partly cancel. The rule is the same as for sums in the √n section: variances add, standard errors do not. Geometrically, the two SEs are the two short sides of a right triangle, and the SE of the gap is the long side.

Three ways to say it:

  • Picture: walk 1.34 steps east and 1.45 steps north; you end up 1.98 steps from where you started, not 2.79.
  • Numbers: $\sqrt{1.34^2 + 1.45^2} = \sqrt{1.80 + 2.11} = 1.98$ points.
  • Slogan: square, add, take the root (Pythagoras for noise).

The checkout gap. From the previous section: $SE(\hat p_A) = 0.01342$ and $SE(\hat p_B) = 0.01453$. The observed gap is $0.12 - 0.10 = 0.02$ (2 points).

  1. Square each SE (turn them into variances): $0.01342^2 = 0.000180$ and $0.01453^2 = 0.000211$.
  2. Add the variances: $0.000180 + 0.000211 = 0.000391$.
  3. Take the square root: $SE(\hat p_B - \hat p_A) = \sqrt{0.000391} = 0.0198$, i.e. 1.98 points.
  4. Wrong way, for comparison: $0.0134 + 0.0145 = 0.0279$. Adding SEs overstates the noise by about 40%.
  5. Read it: the observed gap (2 points) is about $0.02/0.0198 \approx 1.01$ SEs. A gap of about one SE is exactly the size that luck produces all the time. (This ratio is the test statistic of Chapter 5.6.)

If $\hat\theta_A$ and $\hat\theta_B$ are independent estimates (computed from different, independent units), then $Var(\hat\theta_B - \hat\theta_A) = Var(\hat\theta_B) + Var(\hat\theta_A)$, so

$$SE(\hat\theta_B - \hat\theta_A) = \sqrt{SE(\hat\theta_A)^2 + SE(\hat\theta_B)^2}.$$
  • Two proportions: $SE(\hat p_B - \hat p_A) = \sqrt{\dfrac{\hat p_A(1-\hat p_A)}{n_A} + \dfrac{\hat p_B(1-\hat p_B)}{n_B}}$ (the "unpooled" SE, used for confidence intervals).
  • Two means: $SE(\bar x_B - \bar x_A) = \sqrt{s_A^2/n_A + s_B^2/n_B}$ (the SE behind Welch's t-test, Chapter 5.9).
  • The minus sign does not matter: $Var(B - A) = Var(B) + Var(A)$, because $(-1)^2 = 1$.
  • If the estimates are correlated, $Var(B - A) = Var(A) + Var(B) - 2\,Cov(A, B)$. Positive correlation (the same users measured twice, as in a paired design) makes the difference less noisy.
  • In the hypothesis test of Chapter 5.6 you will meet a second version, the pooled SE, which assumes both groups share one rate. For the checkout data it is also 0.0198; the next chapter explains why the two agree here and when they do not.
Why do we need it?

Decisions are about differences: treatment minus control, this week minus last week, model B's error minus model A's. To judge whether a gap is real, you must compare it with the noise of the gap, not with the noise of either group alone.

Where is it used?

Two-sample z- and t-tests, confidence intervals for a lift, A/B sample-size formulas (the $p_A(1-p_A) + p_B(1-p_B)$ term), comparing two models' backtest errors, and difference-in-differences estimates (Chapter 5.12).

How is it used?

Compute each group's SE, then np.sqrt(se_a**2 + se_b**2). Divide the observed gap by it to see how many SEs the gap is. If the groups share users (paired), do not use this formula: work with per-user differences instead.

SE_A = 1.34 SE_B = 1.45 SE of gap = 1.98 in percentage points: ✓ √(1.34² + 1.45²) = √3.91 = 1.98 ✗ 1.34 + 1.45 = 2.79 (SEs do not add) variances add: 1.80 + 2.11 = 3.91 then take the square root
The SE of a difference of independent estimates is the long side of a right triangle whose short sides are the two SEs. Adding the SEs directly (red) overstates the noise.

Drag the purple corner: its horizontal distance is $SE_A$ (blue leg) and its vertical distance is $SE_B$ (orange leg). The purple side is the SE of the gap. The purple arc swings it down onto the axis so you can compare it with the red ring, the wrong answer $SE_A + SE_B$. Press One group much bigger: when one SE is tiny, the gap's SE is almost the other SE. The two are never more than $\sqrt 2 \approx 1.41$ times the bigger one.

Suppose the truth is exactly 10% (old) vs 12% (new), a real 2-point improvement (green line). Each simulated rerun gives a different observed gap; the orange pile shows 1 000 reruns. Its spread is the SE of the gap, $\approx 1.98$ points for 500 users per group. Look at the red part: in about 16% of reruns the new checkout looks no better or worse, even though it really is better. Raise the users per group to 4 000 and watch the red part almost vanish.

"The SE of the difference is $SE_A + SE_B$."

Variances add, not SEs: $\sqrt{SE_A^2 + SE_B^2}$. Adding the SEs only works if the two estimates move perfectly together, which independent groups never do.

"Subtracting two noisy numbers cancels their noise."

For independent groups the opposite is true: the difference is noisier than either part. Noise only cancels in a difference when the two parts are positively correlated (paired data: the same users before and after).

"Each group's rate has a small SE, so the lift is precisely measured."

The lift is a difference of two rates and is about $\sqrt 2$ times noisier than each rate (for equal groups). And the relative lift $(\hat p_B - \hat p_A)/\hat p_A$ is noisier still.

Your Bayesian A/B framework works with the same quantity in a different language: the posterior of $\theta_B - \theta_A$. With independent posteriors for the two arms, its standard deviation is $\sqrt{sd_A^2 + sd_B^2}$, the posterior twin of this formula. A decision like $P(\theta_B \gt \theta_A \mid D)$ is essentially asking how many posterior standard deviations the gap's centre lies above zero (Chapter 6.4).

Independent estimates: $SE(B - A) = \sqrt{SE_A^2 + SE_B^2}$ (variances add; Pythagoras).

Checkout: $\sqrt{0.0134^2 + 0.0145^2} = 0.0198$ = 1.98 points; the 2-point gap is only about 1 SE.

Trap: never add SEs; paired (correlated) data need the per-user differences instead.

Quick check: two independent averages have SEs 3 and 4. What is the SE of their difference? Of their sum?

Both are $\sqrt{3^2 + 4^2} = \sqrt{25} = 5$. For independent pieces the variance of a sum and of a difference are the same: $9 + 16 = 25$.

The estimated SE: using $s$ when $\sigma$ is unknown core

The formula $\sigma/\sqrt n$ has a problem: we almost never know $\sigma$, the true spread of the population. So we do the natural thing: we plug in $s$, the spread we measured in our own sample. The result $s/\sqrt n$ is the estimated standard error.

But $s$ is itself computed from the sample, so it is also a statistic that wobbles from sample to sample. With a big sample, $s$ is close to $\sigma$ and nothing much changes. With a tiny sample (3 people!), $s$ can be far too small or far too big, and so can your SE. We pay for this extra uncertainty later, by using the slightly wider t distribution instead of the Normal.

Three ways to say it:

  • Picture: you measure the noise with a ruler that is itself a little noisy.
  • Numbers: heights 160, 170, 180 give $s = 10$ and $\widehat{SE} = 10/\sqrt 3 = 5.77$ cm; another 3 people could give 1.7 or 9.8.
  • Slogan: an estimated SE is an estimate too.

Your height sample. Three randomly chosen people: 160, 170 and 180 cm.

  1. Mean: $\bar x = (160 + 170 + 180)/3 = 510/3 = 170$ cm.
  2. Distances from the mean: $-10, 0, +10$. Squares: $100, 0, 100$. Sum $= 200$.
  3. Sample variance (divide by $n - 1 = 2$): $s^2 = 200/2 = 100$, so $s = 10$ cm.
  4. Estimated SE: $\widehat{SE}(\bar x) = s/\sqrt n = 10/\sqrt 3 = 10/1.732 = 5.77$ cm.
  5. Meaning: another random group of 3 people would typically give an average about 6 cm away from 170. Three people tell you very little about a city's average height.
  6. The catch: with only 3 values, $s = 10$ is itself shaky. If heights in the town really have $\sigma = 10$, other samples of 3 could easily give $s = 3$ or $s = 17$, and then $\widehat{SE}$ would be 1.7 or 9.8.

When $\sigma$ is unknown, the estimated standard error of the mean is

$$\widehat{SE}(\bar x) = \frac{s}{\sqrt n}, \qquad s = \sqrt{\frac{1}{n-1}\sum_{i=1}^n (x_i - \bar x)^2}.$$
  • It is a statistic, so it has its own sampling distribution. Its relative wobble is large for small $n$ and small for large $n$.
  • Because $s$ is on average a little smaller than $\sigma$ (Chapter 4.5), $\widehat{SE}$ tends to be a little too small, and it is below the true SE in more than half of small samples.
  • The consequence: for Normal data, $\dfrac{\bar X - \mu}{s/\sqrt n}$ follows a Student-t distribution with $n - 1$ degrees of freedom (heavier tails than the Normal, Chapter 4.9), not a standard Normal. For $n = 3$ the 95% multiplier is 4.30 instead of 1.96; for $n = 30$ it is 2.05. This is the one-sample t-test (Chapter 5.6, Chapter 5.9).
  • Other estimated SEs work the same way: $\sqrt{\hat p(1-\hat p)/n}$ plugs $\hat p$ into the proportion formula; regression software plugs $\hat\sigma$ into the coefficient formula.
Why do we need it?

The true $\sigma$ is unknown in real life, so every SE you will ever compute from data is an estimated one. Knowing that it wobbles explains why small-sample tests use the t distribution and why "SE from 5 data points" deserves little trust.

Where is it used?

One- and two-sample t-tests, Welch's test, t-based confidence intervals, the "std err" column of OLS output (Chapter 5.13), scipy.stats.sem, and every dashboard error bar computed from data.

How is it used?

se = x.std(ddof=1) / np.sqrt(len(x)) (or scipy.stats.sem(x)). With small $n$, pair it with t critical values, not 1.96. If $n$ is tiny, say so, and consider prior information (a Bayesian model) or more data.

Imagine the town's heights are Normal with $\sigma = 10$ cm, so the true SE of an $n$-person average is $10/\sqrt n$ (green line). Each simulated sample computes its own $\widehat{SE} = s/\sqrt n$; the blue pile shows thousands of them. At $n = 3$ the pile is huge and lopsided: your 5.77 (purple dot) happens to equal the truth, but many samples are far off, and most are too small. Move $n$ to 30: the pile tightens around the green line.

"The SE is a known, exact number."

In practice it is estimated from the same data as the estimate. With $n = 3$, the estimated SE is below half of the true one about 1 time in 5, and sometimes far above it.

"$s/\sqrt n$ is unbiased, so small samples are fine."

$s$ is biased slightly low and very variable for small $n$, so $s/\sqrt n$ is too small more often than not. That is exactly why small-sample inference uses the t distribution, with its wider critical values.

"Use $n$ in $s$ and $n - 1$ in $\sqrt n$" (or any other mix).

$s$ divides the squared deviations by $n - 1$ (ddof=1); the SE then divides $s$ by $\sqrt n$. NumPy's np.std defaults to ddof=0, so write ddof=1 explicitly.

In a Bayesian model like yours you do not plug $s$ into a formula: the noise scale is a parameter with its own posterior. In your forecasting model, $\sigma$ (or the Student-t scale) is learned together with the trend and seasonality, and the uncertainty about $\sigma$ flows into the predictive intervals automatically. The lesson carries over though: with little data, the noise scale itself is uncertain, and pretending it is known makes intervals too narrow.

$\widehat{SE}(\bar x) = s/\sqrt n$ (use ddof=1 for $s$).

Heights 160, 170, 180: $s = 10$, $\widehat{SE} = 10/\sqrt 3 = 5.77$ cm.

Trap: the estimated SE wobbles and is usually a bit low for small $n$, hence t (with $n - 1$ df) instead of Normal.

Quick check: 4 values 8, 10, 12, 14. What is the estimated SE of their mean?

$\bar x = 11$; deviations $-3, -1, 1, 3$; squares sum to $9 + 1 + 1 + 9 = 20$; $s^2 = 20/3 = 6.67$, $s = 2.58$; $\widehat{SE} = 2.58/\sqrt 4 = 1.29$.

Sampling distributions: the centre, width and shape behind the SE core

The SE is a single number: the width of the pile of estimates you would get from many repeated samples. But a pile has more than a width. It has a centre (is it sitting on the truth? that is the bias of Chapter 5.1) and a shape (a symmetric bell, or lopsided?). The whole pile is called the sampling distribution of the estimator.

For averages and proportions with enough data, the CLT (Chapter 4.13) makes the pile a bell. Then the SE tells you everything: about 68% of estimates land within 1 SE of the centre and about 95% within 2 SE (1.96 to be exact). That "95% within about 2 SE" is the engine of every test and interval in the next chapters.

Other statistics (the median, the SD, the maximum) also have sampling distributions and SEs, but their piles are not always bells, and they rarely have simple formulas.

Three ways to say it:

  • Picture: every estimator leaves a pile of possible values; the SE is the pile's width, the bias is its offset, and the CLT usually makes it bell-shaped.
  • Numbers: 25-day averages with SE 2 fall within $50 \pm 1.96 \times 2 = [46.1, 53.9]$ in 95% of repeats.
  • Slogan: SE = width, bias = offset, CLT = shape.

Three estimators of "typical orders per day" from $n = 10$ days of bell-shaped data with $\mu = 50$, $\sigma = 10$ (simulated with 200 000 repeats):

  1. Mean: centre 50.0 (unbiased), SE $= 10/\sqrt{10} = 3.16$; the pile is a symmetric bell. 95% of means lie in $50 \pm 1.96 \times 3.16 = [43.8, 56.2]$.
  2. Median: centre 50.0, SE about 3.7, wider than the mean's: for bell-shaped data the median wastes information (the efficiency story of Chapter 5.1; the large-$n$ formula $1.2533\,\sigma/\sqrt n = 3.96$ is only a rough guide at $n = 10$).
  3. Maximum: centre about 65 and SE about 5.9, and the pile is lopsided (long right tail). It does not even aim at a fixed population value, because the population has no largest day.
  4. For the skewed days the mean's pile is still close to a bell at $n = 10$, but the SD's and the maximum's piles are clearly lopsided: about 4.5% of them land above centre $+ 1.96$ SE and almost none below centre $- 1.96$ SE, instead of 2.5% on each side.

The sampling distribution of an estimator $\hat\theta = g(X_1, \dots, X_n)$ is its probability distribution over repeated samples of size $n$ from the same process. It has:

  • a centre $E[\hat\theta]$ (the bias is $E[\hat\theta] - \theta$);
  • a width, the standard error $SE(\hat\theta) = \sqrt{Var(\hat\theta)}$;
  • a shape. For means, proportions, differences of means, and most maximum-likelihood estimates (Chapter 5.2), the shape is approximately Normal for large $n$:
$$\hat\theta \;\dot\sim\; N\big(\theta,\ SE^2\big) \quad\Rightarrow\quad P\big(|\hat\theta - \theta| \le 1.96\,SE\big) \approx 0.95 .$$

The symbol $\dot\sim$ means "is approximately distributed as". The approximation can be poor for small $n$, for rates near 0 or 1, for heavy-tailed data, and for statistics like the maximum.

Why do we need it?

To turn an SE into a probability statement ("an estimate this far from the claimed value would happen less than 5% of the time"), we need the shape of the pile, not only its width. The bell shape is what makes "± 1.96 SE" mean 95%.

Where is it used?

z-tests and t-tests (Chapter 5.6), confidence intervals (Chapter 5.8), power calculations (Chapter 5.7), and simulation studies that compare estimators (which estimator has the narrowest pile centred on the truth?).

How is it used?

For means and rates, assume a bell with the formula SE. For anything else, simulate it (if you can write down a believable process) or bootstrap it (next section), then look at the histogram: if it is lopsided, do not trust symmetric "± 1.96 SE" statements.

centre: E[θ̂] (= θ if unbiased) −1 SE+1 SE −1.96 SE+1.96 SE 68% 95% between the outer marks pile of estimates from many samples
When the sampling distribution is a bell, its width (the SE) tells you how often estimates land near the centre: about 68% within 1 SE and 95% within 1.96 SE.

Each run draws 1 500 samples of $n$ days and computes your chosen statistic on each; the blue pile is its sampling distribution and the orange curve is a bell with the same centre and SE. Start with mean: bell-shaped, SE $= 10/\sqrt n$. Switch to median: wider pile. Switch to SD or max with skewed days: the pile leans right, and the readout shows the two tails beyond $\pm 1.96$ SE are far from 2.5% each. Then raise $n$ and watch which piles become bells.

"The sampling distribution is the histogram of my data."

It is the histogram of the estimate over many imagined repeats of the whole sample. You never see it directly; you derive it (formula), simulate it, or bootstrap it.

"Every estimate is within ±1.96 SE of the truth 95% of the time."

Only if the pile is a bell and centred on the truth. Biased estimators (the SD, the maximum) and lopsided piles (small $n$, skewed data, rare events) break the rule.

"The CLT makes my data Normal when $n$ is large."

The data keep their shape forever. The CLT is about the sampling distribution of the mean (Chapter 4.13).

In your forecasting model, quantities like "the peak demand next month" or "the day with the largest residual" are maximum-type statistics: their uncertainty is lopsided, so report quantiles of the posterior predictive draws rather than "mean ± 2 sd". In your A/B framework, a posterior for a rare-event rate (a few conversions) is also lopsided near 0, which is one reason Beta posteriors are reported with credible intervals rather than ± a standard deviation.

Sampling distribution of $\hat\theta$: centre (bias), width (SE), shape.

Means and rates (enough $n$): bell, so about 95% within $\pm 1.96$ SE.

Trap: medians, SDs, maxima, small samples and rare events can have lopsided piles; then "± 2 SE" is misleading.

Quick check: an unbiased estimator has a bell-shaped sampling distribution with SE 4. Roughly how often is it more than 8 away from the truth?

8 is 2 SE, so roughly 5% of the time (2.5% on each side; exactly 4.6% for 2 SE, 5% for 1.96 SE).

The bootstrap: an SE from one sample, by resampling core

The SE is about repeating the experiment, but you only ran it once. For a mean there is a formula. For a median, a ratio, a 90th percentile or a whole pipeline of steps, there is often no formula at all. What can you do?

The bootstrap's trick: treat your sample as if it were the population. Draw a new sample of the same size from it, with replacement (so some values appear twice and some not at all), and compute your statistic. Do that a thousand times. The spread of those thousand values approximates the SE you would see if you could rerun the real experiment.

It is called the "bootstrap" after the old joke of lifting yourself up by your own bootstraps: the sample is used to judge its own reliability.

Three ways to say it:

  • Picture: put your 12 values in a hat, draw 12 times, putting each one back after drawing; repeat a thousand times.
  • Numbers: 1 000 resample means of the 12 order values have sd about 4.5: that is the bootstrap SE.
  • Slogan: resample the sample to see how much the statistic would wobble.

The height sample again: 160, 170, 180. It is so small that we can list every possible resample.

  1. A resample picks 3 values with replacement. Each pick has 3 choices, so there are $3 \times 3 \times 3 = 27$ equally likely resamples, for example (160, 160, 180), (170, 180, 170), (180, 180, 180).
  2. Their means: 160 appears once (all three picks 160), 163.3 three times, 166.7 six times, 170 seven times, 173.3 six times, 176.7 three times, 180 once. (Check: $1 + 3 + 6 + 7 + 6 + 3 + 1 = 27$.)
  3. The average of the 27 means is 170, the original mean.
  4. Their spread: squared distances from 170 are $100$ (twice), $44.4$ (six times) and $11.1$ (twelve times), total $200 + 266.7 + 133.3 = 600$. Divide by 27: $22.2$. Square root: bootstrap SE $= 4.71$ cm.
  5. Compare the formula: $s/\sqrt n = 5.77$. The bootstrap is smaller because it measures spread with divisor $n$ instead of $n - 1$: $5.77 \times \sqrt{2/3} = 4.71$. With $n = 3$ that matters a lot; with $n = 100$ the factor $\sqrt{99/100}$ is 0.995 and the two agree.
160163.3166.7170173.3176.7180 1367631 all 27 resample means of (160, 170, 180): their sd is the bootstrap SE = 4.71 cm
The complete bootstrap distribution for the three heights. Each dot is the mean of one of the 27 equally likely resamples.

The (nonparametric) bootstrap for the SE of a statistic $\hat\theta = g(x_1, \dots, x_n)$:

  1. For $b = 1, \dots, B$ (typically $B$ = 1 000 to 10 000): draw $n$ values from $x_1, \dots, x_n$ with replacement; call this the resample $x^{*}_b$.
  2. Compute the statistic on it: $\hat\theta^{*}_b = g(x^{*}_b)$.
  3. The bootstrap SE is the standard deviation of $\hat\theta^{*}_1, \dots, \hat\theta^{*}_B$.
  • The histogram of the $\hat\theta^{*}_b$ approximates the sampling distribution of $\hat\theta$ (its spread and shape; a percentile interval uses its quantiles, Chapter 5.8).
  • Assumptions: the rows are IID and the sample represents the population; $n$ is not tiny; the statistic is "smooth" (means, medians, ratios, regression slopes are fine; the maximum is not).
  • A resample leaves out about $(1 - 1/n)^n \approx 37\%$ of the original values (35% for $n = 12$), while repeating others.
  • Dependent data: resample whole blocks of consecutive days (block bootstrap) or whole clusters (all sessions of a user), never single rows.
Why do we need it?

Formulas exist only for simple statistics. Real metrics are often medians, ratios (revenue per session), percent lifts or outputs of a whole pipeline. The bootstrap gives an SE for any of them with nothing but your one sample and a loop.

Where is it used?

SEs and intervals for medians, percentiles and ratio metrics in A/B testing, the uncertainty of a model's test-set accuracy, bagging and random forests (each tree trains on a bootstrap resample), and residual block bootstraps for forecast intervals.

How is it used?

idx = rng.integers(0, n, size=(B, n)); stats = g(x[idx]), then stats.std(). Or scipy.stats.bootstrap((x,), np.median), which also returns a confidence interval. Resample the unit of independence (users, not page views; blocks of days, not single days).

your sample x₁, …, xₙ (the only data) resample 1 → θ̂*₁ resample 2 → θ̂*₂ ⋮ resample B → θ̂*_B n draws WITH replacement each histogram of θ̂* its sd = bootstrap SE
The bootstrap recipe: the sample stands in for the population; resampling it with replacement imitates "running the experiment again"; the spread of the recomputed statistic estimates its SE.

Top row: your 12 order values (blue), with their mean (purple tick). Press Next resample: 12 draws with replacement. Values picked twice or more are stacked (orange); values left out show as hollow rings. The resample's mean (orange line) drops into the pile below. Press +1000: the pile's spread is the bootstrap SE, close to the formula $s/\sqrt{12} = 4.73$. Then switch to median: there is no simple formula, but the bootstrap works the same way.

Here we cheat and use a population we know (skewed days, $40 +$ exponential). Orange: the true sampling distribution of the statistic, from 1 000 fresh samples of the population. Blue: the bootstrap distribution from your one sample. Both are shifted to centre 0 so you can compare widths. At $n = 20$ the widths are similar. Press New sample several times: the bootstrap SE moves around the truth. Drop $n$ to 5: the blue pile becomes lumpy and its SE can be far off.

"The bootstrap creates new data, so it makes a small sample more reliable."

It only reuses the data you have. It cannot add information, cannot remove a bias in the sample, and works poorly when $n$ is tiny (your 3 heights give only 7 different resample means).

"Resample without replacement."

Drawing $n$ values without replacement from $n$ values just gives back the same sample in a different order, so every resample statistic would be identical. Replacement is what creates the variation.

"Resample single days of a time series, or single page views of users."

That breaks the dependence and gives an SE that is too small. Resample blocks of consecutive days, or whole users with all their rows.

"The bootstrap works for any statistic."

It fails for statistics that depend on extreme values, like the sample maximum: a resample can never exceed the observed maximum, so it cannot imitate the real uncertainty.

In an A/B framework like yours, a metric such as "revenue per session" is a ratio of two sums with many sessions per user. A bootstrap that resamples users (each with all their sessions) gives an honest SE for it, and is a good sanity check against the posterior width your Bayesian model reports. In your forecasting model, uncertainty comes from the posterior and the predictive simulation rather than the bootstrap; if you ever bootstrap residuals (for example to compare with a classical model), use blocks of consecutive days so that autocorrelation is kept.

"The bootstrap gives the distribution of the true parameter."

"The bootstrap approximates the sampling distribution of my estimator: how the estimate would vary over repeated samples. The parameter is a fixed number in this frequentist view; a distribution over the parameter is a Bayesian posterior."

Model answer: "I resample my data with replacement many times, recompute the statistic each time, and use the spread of those values as the standard error. It assumes the rows are independent and the sample represents the population; for clustered or time-series data I resample clusters or blocks."

Bootstrap: resample $n$ rows with replacement, recompute $\hat\theta^*$, repeat $B$ times; SE $\approx$ sd of the $\hat\theta^*$.

Heights 160, 170, 180: 27 resamples, bootstrap SE $= 4.71$ (formula 5.77; the bootstrap uses divisor $n$).

Trap: resample the independent unit (users, blocks of days); it fails for tiny $n$ and for the maximum.

Quick check: why does a bootstrap resample of 1 000 users typically leave out about 368 of them?

Each draw misses a given user with probability $1 - 1/1000$; over 1 000 draws the chance of never picking them is $(1 - 1/1000)^{1000} \approx e^{-1} \approx 0.368$.

Recap, cheat sheet and practice

  • The SD measures how much single data values spread; the SE measures how much an estimate would spread over repeated samples. The SE is the standard deviation of the estimator's sampling distribution.
  • For a mean of $n$ independent values, $Var(\bar X) = \sigma^2/n$, so $SE(\bar X) = \sigma/\sqrt n$. It needs independence and a finite variance, not Normal data.
  • The √n law: four times the data halves the SE; the needed sample size is $n = (\sigma/SE_{target})^2$. More data shrinks noise, never bias.
  • A proportion is a mean of 0/1 values: $SE(\hat p) = \sqrt{p(1-p)/n}$. Checkout: 1.34 points (A) and 1.45 points (B).
  • For independent groups, variances add: $SE(B - A) = \sqrt{SE_A^2 + SE_B^2}$ = 1.98 points for the checkout. The observed 2-point gap is only about 1 SE.
  • In practice $\sigma$ is unknown, so we use the estimated SE $s/\sqrt n$ (heights: $10/\sqrt 3 = 5.77$); it wobbles too, which is why small samples use the t distribution.
  • The sampling distribution has a centre (bias), a width (SE) and a shape; for means and rates it is close to a bell, so about 95% of estimates fall within $\pm 1.96$ SE.
  • The bootstrap resamples your sample with replacement to estimate the SE of any statistic; resample the independent unit, and do not trust it for tiny $n$ or for the maximum.

Cheat sheet

QuantityFormulaIn words / checkout or height numbers
SD of data$\sigma$, estimated by $s$ (ddof=1)spread of single values; does not shrink with $n$
SE of a mean$\sigma/\sqrt n$, estimated $s/\sqrt n$heights: $10/\sqrt 3 = 5.77$ cm
√n law$SE(kn) = SE(n)/\sqrt k$; $\;n = (\sigma/SE)^2$4× data → ½ SE
SE of a proportion$\sqrt{p(1-p)/n}$$\sqrt{0.1 \times 0.9/500} = 0.0134$
SE of a difference$\sqrt{SE_A^2 + SE_B^2}$ (independent)$\sqrt{0.0134^2 + 0.0145^2} = 0.0198$
Difference of two rates$\sqrt{\frac{p_A(1-p_A)}{n_A} + \frac{p_B(1-p_B)}{n_B}}$"unpooled" SE (pooled version: Chapter 5.6)
Bell rule$P(|\hat\theta - \theta| \le 1.96\,SE) \approx 0.95$only for bell-shaped, unbiased estimators
Bootstrap SEsd of $\hat\theta^*_b$ over $B$ resamples with replacementany statistic; heights: 4.71 (exact, 27 resamples)
Code it · Python

import itertools
import numpy as np
from scipy import stats

rng = np.random.default_rng(0)

# 1. SD vs SE: single days vs the average of 25 days (sigma = 10)
days = rng.normal(50, 10, size=(100_000, 25))   # 100 000 repeated samples of 25 days
print(days[0].std(ddof=1))                      # 8.73: s of ONE sample (it estimates sigma = 10)
print(days.mean(axis=1).std())                  # 2.004: spread of the 25-day averages = 10/sqrt(25)

# 2. The sqrt(n) law: SE of a 10% conversion rate
for n in [100, 400, 1600, 6400]:
    print(n, np.sqrt(0.1 * 0.9 / n))            # 0.03, 0.015, 0.0075, 0.00375 (4x users -> half)
print((0.3 / 0.005) ** 2)                       # 3600.0 users for a target SE of 0.5 points

# 3. SE of a proportion and of a difference: the checkout test
kA, nA, kB, nB = 50, 500, 60, 500
pA, pB = kA / nA, kB / nB
seA = np.sqrt(pA * (1 - pA) / nA)
seB = np.sqrt(pB * (1 - pB) / nB)
se_diff = np.sqrt(seA**2 + seB**2)              # variances add, not SEs
print(seA, seB, se_diff)                        # 0.01342 0.01453 0.01978
print(seA + seB)                                # 0.02795: the WRONG way (about 40% too big)
print((pB - pA) / se_diff)                      # 1.011: the 2-point gap is about one SE

# 4. Estimated SE for the height sample
h = np.array([160, 170, 180])
print(h.std(ddof=1), h.std(ddof=1) / np.sqrt(3), stats.sem(h))   # 10.0 5.7735 5.7735
print(np.std(h))                                # 8.165: careful, NumPy's default is ddof=0

# 5. Bootstrap SE of the mean and of the median (12 order values)
x = np.array([12, 14, 17, 19, 21, 24, 26, 29, 33, 38, 47, 70])
B = 10_000
idx = rng.integers(0, len(x), size=(B, len(x)))  # B resamples, drawn WITH replacement
boot = x[idx]
print(boot.mean(axis=1).std(), x.std(ddof=1) / np.sqrt(len(x)))  # 4.566 vs formula 4.734
print(np.median(boot, axis=1).std())            # 4.554: bootstrap SE of the median (no simple formula)

# 6. The same with SciPy (it also returns a confidence interval)
res = stats.bootstrap((x,), np.median, n_resamples=10_000, random_state=1)
print(res.standard_error)                       # 4.54

# 7. The exact bootstrap of the 3 heights: all 27 resamples
means = [np.mean(c) for c in itertools.product(h, repeat=3)]
print(len(means), np.std(means))                # 27 4.714 (= 5.77 * sqrt(2/3))
Test yourself

1. Daily orders have SD 10 and days are independent. What is the SE of a 100-day average?

$SE = \sigma/\sqrt n = 10/\sqrt{100} = 10/10 = 1$. (0.1 would be $\sigma/n$, the classic mistake; 10 is the SD of single days, which does not shrink.)

2. Your experiment's SE is too big. To cut it in half you need about…

$SE \propto 1/\sqrt n$, so halving the SE needs $2^2 = 4$ times the data. Twice the users only divides the SE by $\sqrt 2 \approx 1.41$.

3. Group A's rate has SE 0.0134 and group B's (independent) has SE 0.0145. What is the SE of the gap $\hat p_B - \hat p_A$?

Variances add: $\sqrt{0.0134^2 + 0.0145^2} = \sqrt{0.000180 + 0.000211} = 0.0198$. Adding the SEs (0.0279) or subtracting them (0.0011) are both wrong.

4. Which statement is true?

The SE is the spread of the estimate over repeated samples. The SD describes single observations and settles at $\sigma$; the SE keeps shrinking, so they move apart as $n$ grows.

5. How does the (nonparametric) bootstrap estimate an SE?

Resampling with replacement imitates drawing new samples from the population. Without replacement, every resample of size $n$ is the original data reordered, so nothing varies.

6. A 2% conversion rate is measured on 1 000 users. What is its SE, roughly?

$\sqrt{0.02 \times 0.98/1000} = \sqrt{0.0000196} = 0.0044$, i.e. 0.44 points. Small in absolute terms, but 22% of the rate itself: rare events are hard to measure precisely.

Practice problems

A. Five days of orders: 48, 52, 50, 46, 54. Compute the sample SD and the estimated SE of the mean, and say in words what each one means.

Mean $= 250/5 = 50$. Deviations $-2, 2, 0, -4, 4$; squares $4, 4, 0, 16, 16$, sum 40. $s^2 = 40/4 = 10$, $s = 3.16$ orders: a single day is typically about 3 orders from the mean. $\widehat{SE} = 3.16/\sqrt 5 = 1.41$ orders: another 5-day average would typically land about 1.4 orders from this one.

B. Revenue per user has SD 30 dollars. How many users do you need for an SE of the mean of 0.5 dollars? And per group, if you want the SE of the difference between two equal groups to be 0.5?

One group: $n = (30/0.5)^2 = 60^2 = 3600$. For a difference of two equal independent groups, $SE_{diff} = \sqrt 2 \times SE_{group}$, so each group needs $SE = 0.5/\sqrt 2 = 0.354$, i.e. $n = (30/0.354)^2 = 7200$ users per group: twice as many.

C. Interview: "A colleague computed the SE of a click rate from 100 000 page views made by 2 000 users and got a tiny SE. What would you say?"

The formula assumes 100 000 independent trials, but page views of the same user are correlated (some users click a lot, some never). The independent unit is the user, so the effective sample size is much closer to 2 000 than 100 000, and the real SE is larger. Fix: compute per-user metrics, or use a cluster-aware SE (delta method) or a bootstrap that resamples users with all their page views.

D. A test shows 20 of 1 000 conversions in A and 26 of 1 000 in B. Find $SE_A$, $SE_B$, the SE of the gap, and how many SEs the gap is.

$SE_A = \sqrt{0.02 \times 0.98/1000} = \sqrt{0.0000196} = 0.00443$. $SE_B = \sqrt{0.026 \times 0.974/1000} = \sqrt{0.0000253} = 0.00503$. $SE_{gap} = \sqrt{0.0000196 + 0.0000253} = \sqrt{0.0000449} = 0.0067$. Gap $= 0.006$, so $0.006/0.0067 = 0.90$ SE: well within ordinary luck.

E. Show that for two independent groups of size $n$ with the same $\sigma$, $SE(\bar x_B - \bar x_A) = \sigma\sqrt{2/n}$.

$Var(\bar x_A) = Var(\bar x_B) = \sigma^2/n$. Independent, so $Var(\bar x_B - \bar x_A) = \sigma^2/n + \sigma^2/n = 2\sigma^2/n$. Square root: $\sigma\sqrt{2/n}$. The difference is $\sqrt 2 \approx 1.41$ times noisier than either mean.

F. Interview: "What is the bootstrap, and when would you not trust it?"

"I treat the sample as the population, draw many resamples of the same size with replacement, recompute the statistic on each, and use the spread of those values as its SE (or their quantiles for an interval). I would not trust it when $n$ is very small (few distinct resamples), for statistics driven by extremes such as the maximum, or when rows are dependent (time series, many rows per user) unless I resample blocks or clusters. It also cannot fix a biased sample."

Chapter 5.6 · Syllabus Module 14

Hypothesis testing: the logic

The new checkout converted 60 of 500 users (12%); the old one converted 50 of 500 (10%). Is the new one really better, or did it just get luckier users this week? A hypothesis test is a careful way to ask one question: is what I see bigger than the luck I would normally expect? This chapter builds the answer one word at a time, with your checkout numbers all the way through.

  • Explain the logic of a test as "assume nothing is going on, then ask how surprising the data would be"
  • Write the null and alternative hypotheses for a real question, before seeing the data
  • Know exactly what a test statistic is (a number computed from the sample, nothing to do with the school subject) and compute $z \approx 1.01$ for the checkout
  • Understand the pooled proportion (0.11) and why the pooled and unpooled SEs both come out about 0.0198 here
  • See where the null distribution comes from, by simulation, by shuffling and by the bell curve
  • State the p-value precisely ($p \approx 0.31$ here), and say what it is not
  • Use $\alpha$ and the critical value; choose one- or two-sided honestly; say "fail to reject", never "accept"; separate significant from important

What we need from earlier chapters: the SE of a proportion and of a difference (Chapter 5.5, SE of a difference); sampling distributions and the "95% within ±1.96 SE" rule (Chapter 5.5); the estimated SE and the t distribution (Chapter 5.5); the standard Normal and z-scores (Chapter 4.9); conditional probability, and why $P(A \mid B) \ne P(B \mid A)$ (Chapter 4.3). Two everyday words used in a special way: a hypothesis is a claim about the world (about a parameter such as a true conversion rate) that data can be checked against; a statistic is any number computed from the sample.

The logic of a test: could luck alone explain it? core

Your checkout result has two possible stories:

  • Story 1: the new checkout really converts better.
  • Story 2 (the boring story): both checkouts convert at the same rate, and the 2-point gap appeared only because of which 1 000 users happened to arrive this week.

You cannot prove story 1 directly. But you can check story 2: pretend it is true, work out what kind of gaps luck alone produces, and see whether your gap looks normal or strange in that pretend world. If it looks very strange, you start to doubt the boring story. If it looks ordinary, luck is a perfectly good explanation, and you have no evidence for story 1.

A court works the same way: the accused is presumed innocent, and is convicted only if the evidence would be very unlikely for an innocent person.

Three ways to say it:

  • Picture: build a pretend world where the change did nothing, and see whether your result looks at home there.
  • Numbers: in a no-difference world, gaps of 2 points or more (either way) show up in about 31% of experiments this size, so your gap is not surprising.
  • Slogan: a test asks "is this bigger than the luck I would normally expect?"

The checkout, in plain steps (the exact numbers are worked out later in this chapter).

  1. What we saw: old 50/500 = 10%, new 60/500 = 12%. Gap: +2 percentage points.
  2. Assume the boring story: both checkouts really convert at the same rate.
  3. How big are gaps from luck alone? From Chapter 5.5, the gap between two groups of 500 wobbles with an SE of about 1.98 points.
  4. Our gap of 2 points is about $2/1.98 \approx 1$ SE. Luck produces gaps of 1 SE or more all the time: about 31% of experiments would show a gap at least this big, in one direction or the other.
  5. Verdict: luck explains this gap easily. The data do not show that the new checkout is better. (They do not show that it is equal either; more on that later.)

A hypothesis test (also called a significance test) is a recipe with five parts. Each part has its own section in this chapter:

  1. a null hypothesis $H_0$ (the boring claim) and an alternative hypothesis $H_1$;
  2. a test statistic: a number computed from the sample that measures how far the data are from what $H_0$ predicts;
  3. its null distribution: how the test statistic would vary over repeated samples if $H_0$ were true;
  4. the p-value: the probability, computed in the $H_0$ world, of a test statistic at least as extreme as the one observed;
  5. a decision rule: reject $H_0$ if $p \le \alpha$ (the significance level, often 0.05); otherwise fail to reject.

This is the frequentist way to reason (Chapter 4.1): probabilities describe how results would vary over imagined repeats of the experiment, and the true rates are fixed unknown numbers.

Why do we need it?

Every metric moves a little from week to week, even when nothing changed. Without a test, teams ship "winners" that were pure noise. The test gives a standard, agreed way to decide whether a difference is larger than ordinary luck.

Where is it used?

A/B tests on conversion and revenue, clinical trials, regression coefficient tables (is this slope zero?), model comparisons, the Ljung–Box check on forecast residuals, sample-ratio-mismatch checks (Chapter 5.11), and every "p < 0.05" in a paper.

How is it used?

Write $H_0$ and $H_1$ and pick $\alpha$ before looking. Run the experiment, compute the test statistic and its p-value (for two rates: statsmodels proportions_ztest), then report the decision together with the estimate and its uncertainty, not the p-value alone.

1 · Assume H₀"no difference":both rates arethe same 2 · Its worldwhat gaps doesluck produce?null distribution 3 · Your datagap ÷ SE =test statisticz ≈ 1.01 4 · How unusual?share of H₀ worldat least as extremep ≈ 0.31 5 · Decidep ≤ α → reject H₀p > α → fail toreject H₀ checkout: p ≈ 0.31 > 0.05, so we fail to reject "no difference"
The five steps of every classical test. Steps 1 and 2 happen in an imagined world where the change did nothing; step 3 brings in your real data; steps 4 and 5 compare the two.

Grey: 400 simulated experiments in a world where the change does nothing (both checkouts 11%). Orange: 400 experiments in a world where the new checkout is really better (10% vs 12%). The purple line is your gap, +2 points. With 500 users per group the two piles overlap heavily: a +2 gap is at home in both worlds, so the data cannot tell them apart. Move the slider to 6 000 users: the piles separate, and now a +2 gap would be rare in the no-effect world.

"A test tells me whether the new checkout is better."

A test tells you whether the data would be surprising if there were no difference. That is a different, narrower question. Whether B is better, and by how much, needs the estimate and its uncertainty too.

"The result is not surprising under $H_0$, so $H_0$ is true."

The widget shows that a +2 gap is also perfectly ordinary in the real-effect world. "Not surprising under $H_0$" means "no evidence against $H_0$", not "evidence for $H_0$".

"We should test the exciting hypothesis directly."

"B is better" does not say how much better, so it does not tell us what data to expect. "No difference" is one exact value, which is what lets us compute the pretend world.

Your Bayesian A/B framework asks a different question with the same data. A test asks "how surprising is a gap like this if there is no difference?" Your framework asks "given the data, how probable is it that B is better?" With flat Beta(1, 1) priors, the checkout gives $P(\theta_B \gt \theta_A \mid D) \approx 0.84$. The one-sided p-value is about 0.16, close to $1 - 0.84$: for large samples and flat priors the two numbers often land near each other, but they mean different things. An interviewer will want you to say both questions clearly (Chapter 6.4).

Test logic: assume $H_0$ (no difference) → what does luck alone produce? → is my result unusual there?

Checkout: gap 2 points ≈ 1 SE; luck gives gaps that big about 31% of the time → no evidence of a real difference.

Trap: "not surprising under $H_0$" ≠ "$H_0$ is true".

Quick check: with 50 000 users per group, the same 10% vs 12% rates give an SE of the gap of about 0.2 points. Would a 2-point gap be surprising under "no difference"?

Yes: 2 points would be about 10 SEs. Luck almost never produces a gap that large, so the boring story would be very hard to believe. Same gap, very different verdict: the SE decides how surprising a gap is.

The null and alternative hypotheses core

A test needs two clear claims, written down before you look at the data:

  • The null hypothesis $H_0$ ("H-nought") is the boring claim: no difference, no effect, the claimed value is right. "Null" means "nothing". It is the claim the test puts on trial.
  • The alternative hypothesis $H_1$ (also written $H_a$) is what you suspect instead: there is a difference (or a difference in one particular direction).

Why is the null the boring one? Because it is precise. "The two rates are equal" pins down one exact situation, so we can compute what data it would produce. "The new checkout is better" could mean 0.1 points better or 5 points better; it does not tell us what to expect.

Three ways to say it:

  • Picture: $H_0$ is a single point on the number line; $H_1$ is everything else (or everything on one side).
  • Numbers: $H_0: p_B - p_A = 0$ against $H_1: p_B - p_A \ne 0$.
  • Slogan: the null is the "nothing happening" claim we try to find evidence against.
  1. Checkout. Parameter: the true gap $\Delta = p_B - p_A$ between the two checkouts' conversion rates. $H_0: \Delta = 0$ (same rate). $H_1: \Delta \ne 0$ (different rate, better or worse).
  2. Height claim (your earlier example). Someone claims the average height is 175 cm. Parameter: the true mean $\mu$. $H_0: \mu = 175$. $H_1: \mu \ne 175$. Your sample mean 170 is not part of the hypotheses; it is the evidence.
  3. Forecast bias. Parameter: the true mean forecast error $\mu_e$. $H_0: \mu_e = 0$ (the forecast is unbiased). $H_1: \mu_e \ne 0$.
  4. A one-direction question. If, before the test, the only question that matters is "is the new checkout better?", write $H_0: \Delta \le 0$ against $H_1: \Delta \gt 0$. (One- vs two-sided gets its own section below.)

For a parameter $\theta$ and a fixed value $\theta_0$ chosen before seeing the data:

$$H_0: \theta = \theta_0 \qquad\text{against}\qquad H_1: \theta \ne \theta_0 \quad\text{(two-sided)},$$ $$\text{or}\quad H_0: \theta \le \theta_0 \;\text{ vs }\; H_1: \theta \gt \theta_0, \qquad H_0: \theta \ge \theta_0 \;\text{ vs }\; H_1: \theta \lt \theta_0 \quad\text{(one-sided)}.$$
  • Hypotheses are statements about parameters (the population, the process), never about statistics you can already see.
  • $H_0$ and $H_1$ do not overlap, and together they cover all the possibilities you care about.
  • For a one-sided null like $\theta \le \theta_0$, the calculations are done at the boundary $\theta = \theta_0$ (the value most favourable to $H_1$ among those in $H_0$).
  • $H_0$ is not "what we believe". It is a reference assumption that we test.
Why do we need it?

Without a precise $H_0$ there is nothing to compute the "luck alone" world from. Writing both hypotheses first also stops you from bending the question after seeing the data, which quietly inflates false positives.

Where is it used?

The pre-registration or experiment plan of an A/B test, every regression output ("$H_0$: this coefficient is 0"), goodness-of-fit and independence tests, the Ljung–Box test ("$H_0$: no autocorrelation up to lag $h$"), sample-ratio-mismatch checks ("$H_0$: the split is 50/50").

How is it used?

Name the parameter in words ("the true conversion rate gap"), choose $\theta_0$ (usually 0 for a difference), choose two-sided unless only one direction could ever change your decision, and write all of it in the experiment plan before the data arrive.

two-sided H₁: Δ ≠ 0H₀: Δ = 0 right-sided H₁: Δ > 0H₀: Δ ≤ 0 left-sided H₁: Δ < 0H₀: Δ ≥ 0 B worseB better
The true gap Δ = pB − pA on a number line. Purple: what the null hypothesis allows. Orange: what the alternative allows. The test's calculations always happen at the purple dot, Δ = 0.

Choose a scenario and the question you care about. The purple dot is the exact value the null hypothesis claims; the orange region is what the alternative allows. Try the height claim: notice that $H_0$ is about the population mean $\mu$ (175), not about your sample mean (170). Then switch between different?, bigger? and smaller? and read how the symbols change.

"$H_0: \bar x = 175$."

You already know $\bar x$ (it is 170). Hypotheses are about the unknown parameter: $H_0: \mu = 175$.

"I'll decide whether to test 'better' or 'worse' after I see which way the data went."

Choosing the direction after looking doubles your false-positive rate (simulated in the one- vs two-sided section). Write $H_1$ first.

"The null hypothesis is what I believe."

It is a reference assumption, put on trial. Often you suspect it is false; the test checks whether the data give enough evidence to say so.

Hypotheses hide inside tools you already use. In your forecasting work, the Ljung–Box test on residuals has $H_0$: "no autocorrelation up to lag $h$" (Chapter 7.17); a classical regression table tests $H_0$: "this regressor's coefficient is 0" for every column. In your A/B framework there is no $H_0$ at all: the Bayesian model gives a posterior for $\theta_B - \theta_A$, and the decision is a probability statement about it. Knowing what $H_0$ a classical tool assumes is how you translate between the two.

$H_0$: the precise "nothing happening" claim about a parameter ($\Delta = 0$, $\mu = 175$). $H_1$: what you suspect instead ($\ne$, or $\gt$ / $\lt$).

Write both (and the direction) before looking at the data.

Trap: never put a sample statistic ($\bar x$, $\hat p$) in a hypothesis.

Quick check: a team asks "did the redesign change average order value (currently 40 dollars)?" Write the hypotheses.

Parameter: $\mu$, the true mean order value with the redesign. $H_0: \mu = 40$; $H_1: \mu \ne 40$ (two-sided, because "change" includes both directions). If you compare against a control group instead: $H_0: \mu_B - \mu_A = 0$ vs $H_1: \mu_B - \mu_A \ne 0$.

The test statistic: how far from $H_0$, measured with an SE-ruler core

First, the word. "Statistic" here does not mean the school subject "statistics". A statistic is simply a number you calculate from your sample. From the heights 160, 170, 180: the sample mean 170 is a statistic, the sample SD 10 is a statistic, even the sample size 3 is a statistic. From the checkout: 50/500 = 10% is a statistic.

A test statistic is a statistic built for one job: to measure how far your data are from what $H_0$ predicts. Raw distances are hard to judge: is a 2-point gap big? It depends on how much the gap normally wobbles. So we measure the distance with a special ruler whose unit is one standard error. A test statistic of 1 means "one ordinary wobble away"; a test statistic of 3 means "three wobbles away: rare".

Three ways to say it:

  • Picture: lay an SE-ruler next to the gap and count how many ruler-units long it is.
  • Numbers: checkout: $(0.12 - 0.10 - 0)/0.0198 = 1.01$; heights: $(170 - 175)/5.77 = -0.87$.
  • Slogan: test statistic = (what I saw − what $H_0$ says) ÷ the usual luck.

(a) Your height claim. $H_0: \mu = 175$. Sample 160, 170, 180.

  1. What I saw: $\bar x = 170$. What $H_0$ says: 175. Distance: $170 - 175 = -5$ cm.
  2. The usual luck for a mean of 3 people: $\widehat{SE} = s/\sqrt n = 10/\sqrt 3 = 5.77$ cm (Chapter 5.5).
  3. Test statistic: $t = -5/5.77 = -0.87$. The sample mean is 0.87 rulers below the claim. Less than one ordinary wobble: not unusual.
  4. It is called $t$ (not $z$) because the ruler was built from $s$, an estimate; its null distribution is the t distribution with $n - 1 = 2$ degrees of freedom.

(b) The checkout. $H_0: p_B - p_A = 0$.

  1. What I saw: $\hat p_B - \hat p_A = 0.12 - 0.10 = 0.02$. What $H_0$ says: 0. Distance: 0.02.
  2. The usual luck for this gap, computed in the $H_0$ world: $SE_0 = 0.0198$ (the next section shows how, with the pooled proportion).
  3. Test statistic: $z = 0.02/0.0198 = 1.01$. The gap is about one ruler long: an everyday amount of luck.

A statistic is any function of the sample data alone (Chapter 5.1). A test statistic $T$ is a statistic chosen so that (1) it measures disagreement with $H_0$, and (2) its distribution is known when $H_0$ is true. The most common form is

$$T = \frac{\hat\theta - \theta_0}{SE(\hat\theta)} = \frac{\text{estimate} - \text{value claimed by } H_0}{\text{standard error of the estimate}}.$$
  • Sign: which side of $\theta_0$ the data fell on. Size $|T|$: how unusual. Under $H_0$ it is usually between −2 and 2.
  • It is called a z-statistic when its null distribution is the standard Normal $N(0, 1)$ (known $\sigma$, or large samples, as for proportions), and a t-statistic when the SE uses $s$ and the null distribution is Student-t.
  • Other tests use other test statistics with the same job: $\chi^2$ for counts in tables, $F$ for ANOVA, $U$ for Mann–Whitney (Chapter 5.9).
  • "z" is the same z as in z-scores (Chapter 4.9): a distance measured in standard deviations, here the SD of the estimate, i.e. the SE.
Hard wordPlain English
statistica number computed from the sample (mean, SD, count, rate)
test statistica statistic that says how far the data are from $H_0$, in SE units
z-statistica test statistic whose $H_0$ distribution is the standard bell $N(0,1)$
t-statisticthe same idea when the ruler (SE) was estimated with $s$; uses the t curve
Why do we need it?

A 2-point gap with 500 users and a 2-point gap with 50 000 users are very different evidence. Dividing by the SE puts every result on the same scale ("how many ordinary wobbles?"), so one table of the bell curve works for every test.

Where is it used?

The z in a two-proportion A/B test, the t in a t-test, the "t value" column next to each coefficient in a regression table, Welch's t, the z of a CUPED-adjusted metric, and z-scores in anomaly detection on forecast residuals.

How is it used?

Compute the estimate, subtract the value $H_0$ claims, divide by the SE: z = (p_b - p_a - 0) / se0. Then look it up in the null distribution to get a p-value (next sections). A quick mental check: $|z|$ above about 2 is the zone where results start to look unusual.

a statistic = any number from the sample mean 170 · SD 10 · n = 3 50/500 = 10% · 60/500 = 12% (not the school subject!) estimate − value claimed by H₀ standard error (the ruler) T = checkout: (0.02 − 0) / 0.0198 = 1.01 heights: (170 − 175) / 5.77 = −0.87 a test statistic = a statistic built to measure distance from H₀ in units of ordinary luck
Left: a statistic is just a number from the sample. Right: a test statistic is a special statistic, the distance between the data and the null claim, measured in standard errors.

The purple handle is your result; the grey ruler below has one tick per standard error, starting at the value $H_0$ claims. Drag the handle and read the test statistic: it is just "how many ruler units away". In Checkout mode, keep the gap at +2 points and move the users slider from 500 to 4 000: the ruler's ticks shrink, so the same gap becomes more ruler units long. Switch to Height claim to see your $t = -0.87$.

"A test statistic is something from statistics class."

A statistic is any number computed from the sample. A test statistic is the particular statistic used to test $H_0$: usually (estimate − null value) / SE.

"The test statistic is the gap."

The gap (2 points) is the raw distance. The test statistic is the gap divided by its SE (1.01). The same 2-point gap is $z = 1.01$ with 500 users per group and $z \approx 2.9$ with 4 000.

"A bigger test statistic means a bigger effect."

It means a bigger effect relative to the noise. A tiny effect with millions of users can have a huge $z$ (see significance vs importance).

"The test statistic tells me the probability that my result is due to chance."

"The test statistic is a number computed from the sample: how many standard errors my estimate is from the value under the null. Its probability interpretation comes only after comparing it with the null distribution, which gives the p-value."

Model answer: "For the checkout, the estimate is a 2-point lift, the null says 0, and the SE under the null is 1.98 points, so $z = 1.01$: the lift is about one standard error from zero, which is well within what sampling noise produces."

Statistic = any number computed from the sample (NOT the school subject).

Test statistic $= \dfrac{\text{estimate} - \text{null value}}{SE}$: distance from $H_0$ in SE units. Checkout $z = 0.02/0.0198 = 1.01$; heights $t = -5/5.77 = -0.87$.

Trap: the test statistic grows with $n$ for the same gap; it is not the effect size.

Quick check: a regression slope is estimated as 0.6 with SE 0.25. What is the test statistic for $H_0$: slope = 0?

$(0.6 - 0)/0.25 = 2.4$. The estimate is 2.4 standard errors above zero (it is the "t value" a regression table would print).

The pooled proportion: one shared rate in the "no difference" world core

To build the SE-ruler for the checkout test, we need to know how much the gap wobbles in the $H_0$ world, the pretend world where both checkouts convert at the same rate. But what is that one shared rate? We do not know it. We have to estimate it.

In the $H_0$ world, the labels "old" and "new" do not matter: all 1 000 users are just users of one checkout with one rate. So the best estimate of that rate uses everyone: pour both groups into one pool and count. 50 + 60 = 110 conversions out of 500 + 500 = 1 000 users, so the shared rate is 11%. That is the pooled proportion. "Pooled" simply means "poured together into one pool".

Then the SE of the gap in that world uses 11% for both groups.

Three ways to say it:

  • Picture: pour the old jar (50 of 500) and the new jar (60 of 500) into one big jar: 110 of 1 000.
  • Numbers: $\hat p_{pool} = (50 + 60)/(500 + 500) = 0.11$, and $SE_0 = \sqrt{0.11 \times 0.89 \times (1/500 + 1/500)} = 0.0198$.
  • Slogan: in the "no difference" world there is only one rate, so estimate it from everybody.

The checkout, step by step.

  1. Pool the counts: $\hat p_{pool} = \dfrac{50 + 60}{500 + 500} = \dfrac{110}{1000} = 0.11$.
  2. Variance of one user in the $H_0$ world: $0.11 \times (1 - 0.11) = 0.11 \times 0.89 = 0.0979$.
  3. Two groups of 500, both with that variance: $Var(\hat p_B - \hat p_A) = 0.0979/500 + 0.0979/500 = 0.0979 \times 0.004 = 0.000392$.
  4. $SE_0 = \sqrt{0.000392} = 0.0198$, i.e. 1.98 points. (The subscript 0 means "computed assuming $H_0$".)
  5. Test statistic: $z = (0.12 - 0.10)/0.0198 = 1.01$.
  6. The puzzle of the two SEs. In Chapter 5.5 we computed the SE of the gap with the two separate rates: $\sqrt{0.1 \times 0.9/500 + 0.12 \times 0.88/500} = \sqrt{0.000180 + 0.000211} = 0.0198$. Also 0.0198! It is not one SE computed twice. They are two different SEs that happen to agree here, because 10% and 12% are both close to 11% and the groups are the same size.

For $k_A$ conversions out of $n_A$ and $k_B$ out of $n_B$, the pooled proportion is

$$\hat p_{pool} = \frac{k_A + k_B}{n_A + n_B} = \frac{n_A\,\hat p_A + n_B\,\hat p_B}{n_A + n_B},$$

a weighted average of the two rates (bigger groups count more). The SE of the gap under $H_0$ and the two-proportion z-statistic are

$$SE_0 = \sqrt{\hat p_{pool}(1 - \hat p_{pool})\left(\frac{1}{n_A} + \frac{1}{n_B}\right)}, \qquad z = \frac{\hat p_B - \hat p_A}{SE_0}.$$
  • Test ("pretend world"): $H_0$ says one shared rate, so the test uses the pooled $SE_0$. This keeps the test consistent with the hypothesis it is testing.
  • Interval ("real world"): a confidence interval for the gap does not assume $H_0$, so it uses the unpooled SE with the two separate rates (Chapter 5.8).
  • They give nearly the same number when the rates are close and the groups are similar in size. They differ when the rates are far apart (4% vs 20% with 500 each: 0.0206 pooled vs 0.0199 unpooled) or the groups are very unequal (10/100 vs 600/5000: 0.0328 vs 0.0304).
  • The estimate of the lift itself still uses the separate rates: $\hat p_B - \hat p_A$. Pooling is only for the ruler.
Why do we need it?

The test statistic must be computed in the world $H_0$ describes. In that world the rates are equal, so we need one estimate of the common rate to build the SE-ruler. Using everyone gives the most stable estimate.

Where is it used?

The standard two-proportion z-test (the default in statsmodels' proportions_ztest), the classical A/B sample-size formula (its $\sqrt{2\bar p(1-\bar p)}$ term), the chi-square test of a 2×2 table (without continuity correction it equals $z^2$ of the pooled test), and pooled-variance (Student) t-tests for means.

How is it used?

p_pool = (k_a + k_b) / (n_a + n_b), then se0 = sqrt(p_pool * (1 - p_pool) * (1/n_a + 1/n_b)) and z = (p_b - p_a) / se0. For an interval around the lift, switch to the unpooled SE.

old: 50 of 500 10% new: 60 of 500 12% pour together pool: 110 of 1 000 = 11% in the H₀ world both groups share this one rate SE₀ = √(0.11 × 0.89 × (1/500 + 1/500)) = 0.0198
The pooled proportion. Under H₀ the two checkouts are one checkout, so their users can be poured into one pool to estimate the shared rate, 11%. The SE of the gap in that world is built from this one rate.

Bar widths are proportional to group sizes; the purple line is the pooled rate, a weighted average that sits closer to the bigger group. Start with the checkout: pooled and unpooled SEs both round to 0.0198. Press Rates far apart: the two SEs now differ by a few percent. Press Very unequal groups: the pooled rate is pulled almost all the way to group B, and the SEs differ by about 8%. That is when it matters which one you use.

"The pooled rate is the average of the two rates, (10% + 12%)/2."

It is the total conversions over the total users, a weighted average. It equals the simple average only when the groups are the same size. With 10/100 and 600/5 000 it is 610/5 100 = 11.96%, not 11%.

"The pooled SE and the unpooled SE are one SE computed twice."

They live in different worlds: the pooled SE assumes $H_0$ (one shared rate) and is used for the test; the unpooled SE uses the separate rates and is used for confidence intervals. They agree only when the rates are close.

"Pooling means we mix the groups and estimate one lift from everyone."

The lift is still $\hat p_B - \hat p_A$ from the separate groups. Only the ruler (the SE in the $H_0$ world) uses the pool.

Careful with the word "pooling" in your own project. Your hierarchical A/B model uses partial pooling: each segment's rate is pulled toward a shared group mean by an amount the model learns (Chapter 6.6). That is a modelling choice about the estimates. The pooled proportion of a z-test is a different thing: a single-use estimate of the common rate under $H_0$, used only to build the test's SE. Same word, different idea.

"We pool because the two groups are basically the same."

"We pool because the test is computed assuming $H_0$, and under $H_0$ the two groups share one rate. The pooled proportion is the best estimate of that shared rate."

Model answer: "For a two-proportion z-test I estimate the common rate under the null as total conversions over total users, 110/1 000 = 0.11, and use it to get the null standard error, 0.0198. For a confidence interval for the lift I use the unpooled SE, because the interval should not assume the null. Here both are about 0.0198 because the rates are close."

$\hat p_{pool} = \dfrac{k_A + k_B}{n_A + n_B}$ (checkout: 110/1 000 = 0.11). $SE_0 = \sqrt{\hat p_{pool}(1-\hat p_{pool})(1/n_A + 1/n_B)} = 0.0198$.

Test → pooled SE (assumes $H_0$). Interval → unpooled SE (separate rates).

Trap: both are 0.0198 here only because 10% and 12% are close and $n_A = n_B$.

Quick check: 30 of 200 converted in A (15%) and 60 of 300 in B (20%). Compute the pooled rate, $SE_0$ and $z$.

$\hat p_{pool} = 90/500 = 0.18$. $SE_0 = \sqrt{0.18 \times 0.82 \times (1/200 + 1/300)} = \sqrt{0.1476 \times 0.00833} = \sqrt{0.00123} = 0.0351$. $z = (0.20 - 0.15)/0.0351 = 1.43$.

The null distribution: what pure luck looks like core

We have one test statistic, $z = 1.01$. Is that big? To answer, we need to know which values of $z$ luck alone produces. Imagine running the checkout experiment thousands of times in the $H_0$ world (same rate for both), computing $z$ each time, and piling up the results. That pile is the null distribution: the sampling distribution of the test statistic when $H_0$ is true.

There are three ways to get it, and they agree:

  • Simulate the $H_0$ world (possible because $H_0$ pins down the rates).
  • Shuffle the labels: if the checkout made no difference, which users were called "old" and "new" is arbitrary, so reshuffling the labels shows what gaps luck makes (a permutation test; the shuffle widget in the guide's introduction does this for the checkout).
  • Theory: by the CLT, $z$ is close to a standard Normal $N(0, 1)$ in the $H_0$ world.

Three ways to say it:

  • Picture: the pile of test statistics from a world where nothing changed.
  • Numbers: in that world, $|z| \ge 1.01$ in about 31% of experiments and $|z| \ge 1.96$ in about 5%.
  • Slogan: the null distribution is "luck", written in the units of your test statistic.

A permutation test small enough to do by hand. Four users saw the old checkout and bought 3, 5, 2, 4 items; four saw the new one and bought 6, 4, 7, 5 items.

  1. Means: old $= 14/4 = 3.5$, new $= 22/4 = 5.5$. Observed difference (new − old) $= 2.0$ items. This difference is our test statistic.
  2. $H_0$: the page has no effect on items bought. Then each user would have bought the same number whichever page they saw, and the labels "old"/"new" are just a random split of the 8 numbers.
  3. How many ways are there to choose which 4 of the 8 users are called "new"? $\binom{8}{4} = 70$. Each is equally likely under $H_0$.
  4. For every one of the 70 splits compute (mean of "new" − mean of "old"). The 70 differences are: −2.5 (once), −2.0 (4 times), −1.5 (5), −1.0 (9), −0.5 (10), 0 (12), +0.5 (10), +1.0 (9), +1.5 (5), +2.0 (4), +2.5 (once). This list is the exact null distribution.
  5. Splits at least as extreme as ours (difference ≤ −2 or ≥ +2): $1 + 4 + 4 + 1 = 10$. So $p = 10/70 = 0.143$ (two-sided). A 2-item gap from 4 users per group is quite possible by luck.

The null distribution of a test statistic $T$ is its probability distribution over repeated samples when $H_0$ is true. It is the sampling distribution of $T$ in the $H_0$ world.

  • For the two-proportion z-test (and other large-sample z-tests): approximately $N(0, 1)$, by the CLT.
  • For the one-sample t-test on Normal data: exactly Student-t with $n - 1$ degrees of freedom.
  • For a permutation test: the distribution of $T$ over all relabelings of the data (or many random ones). It needs only that, under $H_0$, the labels are exchangeable (swapping them does not change the distribution of the data), which randomization guarantees.
  • The approximations can be poor for small samples or rare events; permutation and simulation then give more honest answers.
Why do we need it?

A test statistic on its own is just a number. The null distribution turns it into a statement about luck: how often would a value this extreme appear if nothing were going on? Without it there is no p-value and no critical value.

Where is it used?

The Normal table behind every z-test, the t table behind t-tests, chi-square and F tables (Chapter 5.9), permutation tests for odd metrics (medians, ratios), and A/A simulations that check an experimentation platform's false-positive rate (Chapter 5.7).

How is it used?

For standard tests, the library uses the known curve (scipy.stats.norm, t). For anything unusual, simulate it: shuffle the labels many times and recompute the statistic (scipy.stats.permutation_test), then compare your observed value with the pile.

Each experiment gives both groups the same true rate (11%) and computes the pooled $z$. Press Run 1000: the grey pile settles onto the orange bell, the standard Normal $N(0,1)$. The red bars are experiments with $|z| \ge 1.01$, as extreme as your checkout: about 31% of them. Then switch to 20 per group: with so few conversions $z$ can only take a few values, the pile is lumpy, and the bell is only a rough guide.

Top rows: items bought by 4 users of the old page (blue) and 4 of the new page (orange); drag any dot. Bottom: the differences (new − old) for all 70 ways of choosing which 4 users are called "new"; red bars are at least as extreme as the real difference. Move the shuffle number slider (or press Random shuffle) to see one relabeling in the third row and where its difference sits in the pile. Drag the new page's dots further right: the real difference grows, fewer splits beat it, and the p-value drops. With 4 users per group the smallest possible two-sided p is 2/70 = 0.029.

"The null distribution is the distribution of my data."

It is the distribution of the test statistic, in the imagined world where $H_0$ is true. Your data give one point on it.

"The bell curve is always the right null distribution for $z$."

It is a large-sample approximation. With few conversions (try 20 per group in the widget) the true null distribution is lumpy; exact or permutation methods are safer.

"A permutation test can find significance in any sample."

With 3 users per group there are only 20 splits, so the smallest possible two-sided p-value is 2/20 = 0.10. Tiny samples simply cannot produce strong evidence.

The simulation habit transfers directly to your NumPyro work. To check how often a decision rule in your Bayesian A/B framework would declare a winner when there is no difference, simulate many "A/A" datasets with $\theta_A = \theta_B$, run the model on each, and count the false "wins". That pile of decisions is the Bayesian cousin of a null distribution, and it is the honest way to learn your rule's false-positive rate (more in Chapter 5.7 and Chapter 5.11).

Null distribution = distribution of the test statistic if $H_0$ is true. Get it by theory ($z \approx N(0,1)$, $t_{n-1}$), simulation, or shuffling labels.

Permutation example: 10 of 70 splits as extreme as +2 items, $p = 10/70 = 0.143$.

Trap: the bell is an approximation; small samples and rare events need exact or simulated null distributions.

Quick check: under $H_0$, in what fraction of experiments does $z$ land beyond ±1.96? Beyond ±1.01?

About 5% beyond ±1.96 and about 31% beyond ±1.01, because $z$ follows (approximately) the standard Normal in the $H_0$ world.

The p-value: how surprising would your data be if $H_0$ were true? core

Put your test statistic on the null distribution. Some of the pile lies at your value or farther from 0. The size of that part (its probability, the area under the curve) is the p-value. A big area means "results like mine happen all the time when nothing is going on". A tiny area means "results like mine would be rare if nothing were going on, so maybe something is".

Said four ways, from formal to everyday:

  • Formal: the probability of a test statistic at least as extreme as the observed one, assuming $H_0$ is true.
  • Simple: if the claim is true, how often would we get a result this far away, or farther?
  • Very simple: if the claim is true, how unusual is our result?
  • Everyday: "if nothing changed, would this surprise me?"

Three ways to say it:

  • Picture: the sand in the tails of the bell, beyond your test statistic.
  • Numbers: $z = 1.01$: each tail beyond ±1.01 holds 0.156, so $p = 0.312$.
  • Slogan: small p = "this would be surprising if nothing were going on".

The checkout ($z = 1.01$, two-sided).

  1. "At least as extreme" means $|Z| \ge 1.01$: a gap of 2 points or more in either direction, because $H_1$ says "different".
  2. From the standard Normal: $P(Z \ge 1.01) = 0.156$, and by symmetry $P(Z \le -1.01) = 0.156$.
  3. $p = 0.156 + 0.156 = 0.312$.
  4. In words: if the two checkouts really converted at the same rate, about 31% of experiments like this one would show a gap at least as large as 2 points. Not surprising at all.

How p shrinks as the test statistic grows (two-sided): $|z| = 0 \to p = 1.00$; $1 \to 0.32$; $2 \to 0.046$; $3 \to 0.003$.

Your height claim ($t = -0.87$): with the Normal curve, $p = 2 \times 0.193 = 0.39$. With the correct t curve for $n - 1 = 2$ degrees of freedom (fatter tails, because the SE was estimated from only 3 people), $p = 0.48$. Either way: no evidence against "average = 175".

For a test statistic $T$ with observed value $t_{obs}$:

$$p = P_{H_0}\big(|T| \ge |t_{obs}|\big) \;\;\text{(two-sided)}, \qquad p = P_{H_0}\big(T \ge t_{obs}\big) \;\text{ or }\; P_{H_0}\big(T \le t_{obs}\big) \;\;\text{(one-sided)}.$$

The subscript $H_0$ means "computed in the world where $H_0$ is true". For a z-statistic: two-sided $p = 2\,\Phi(-|z|)$, where $\Phi$ is the standard Normal CDF.

  • The p-value is computed from the data, so it is itself a statistic and changes from sample to sample.
  • When $H_0$ is true (and $T$ is continuous), $p$ is uniformly distributed between 0 and 1: every value is equally likely, so $P(p \le 0.05) = 0.05$. When $H_0$ is false, p-values pile up near 0, more so with more data.
  • It is a conditional probability of the data given $H_0$: $P(\text{data this extreme} \mid H_0)$. It is not $P(H_0 \mid \text{data})$ (Chapter 4.3: these two can be wildly different).
Why do we need it?

It turns "how many SEs away" into a probability on a common 0-to-1 scale, so results from very different tests (z, t, chi-square, permutation) can be read the same way: how compatible are the data with "nothing is going on"?

Where is it used?

Every A/B test report, the "P>|t|" column of regression tables, Ljung–Box and other diagnostic tests, multiple-testing corrections (Bonferroni, Benjamini–Hochberg, Chapter 5.11), and the "p < 0.05" decisions in science.

How is it used?

Compute $z$, then p = 2 * scipy.stats.norm.sf(abs(z)) (two-sided). Compare it with $\alpha$ chosen in advance, and always report it next to the estimate and its interval: "lift +2.0 points, SE 1.98, p = 0.31".

your z = +1.01 mirror −1.01 0.1560.156 0−1+1−2+2 p (two-sided) = 0.156 + 0.156 = 0.312 where z lands if H₀ is true
The p-value is an area. The curve shows where the test statistic usually lands when H₀ is true; the red tails are results at least as extreme as yours (in either direction, for a two-sided test).

The curve is the null distribution of $z$ (the standard bell). Drag the purple handle to your test statistic; the red area is the p-value. Start at the checkout's 1.01: the two tails hold 0.156 each, $p = 0.312$. Drag to 2: $p \approx 0.046$. Drag to 3: $p \approx 0.003$. Switch to one-sided: bigger and the left tail disappears (p halves). Switch to one-sided: smaller with $z = +1.01$: now almost the whole curve counts, $p = 0.84$, because the data went the "wrong" way for that question.

Each run performs 1 000 checkout-style tests and draws the histogram of their p-values. With no real effect, the histogram is roughly flat: every p-value is about equally likely, and about 5% land below 0.05 (red) purely by luck. Switch to real effect, 500 per group (10% vs 12%): p-values lean toward 0, but most are still above 0.05. With 4 000 per group they pile up near 0. One experiment's p-value is one draw from these piles.

"$p = 0.31$ means there is a 31% chance that $H_0$ is true."

The p-value is computed assuming $H_0$ is true, so it cannot be the probability that $H_0$ is true. It is $P(\text{data this extreme} \mid H_0)$, not $P(H_0 \mid \text{data})$.

"If I reran the experiment I would get about the same p-value."

p-values jump around a lot between replications (the widget above): under $H_0$ any value from 0 to 1 is equally likely, and even with a real effect one run can give 0.01 and the next 0.4.

"The p-value measures how big the effect is."

It mixes effect size and sample size. A tiny effect with huge $n$ can have $p \lt 0.001$; a large effect with small $n$ can have $p = 0.3$.

"The p-value is the probability that the result happened by chance."

"The p-value is the probability, computed assuming the null hypothesis is true, of getting a test statistic at least as extreme as the one I observed."

Model answer: "For the checkout, $p = 0.31$ means: if both checkouts really converted at the same rate, about 31% of experiments of this size would show a gap of 2 points or more in either direction. So the data are quite compatible with no difference. It does not tell me the probability that there is no difference; for that I would need a prior and a Bayesian posterior."

Your Bayesian A/B framework reports the quantity many people wish the p-value was: $P(\theta_B \gt \theta_A \mid D)$, a probability about the parameters given the data. For the checkout with flat priors it is about 0.84. It needs a prior; the p-value does not. When you present results to people used to p-values, say explicitly that 0.84 is a posterior probability that B is better, not "1 − p", and that a p-value of 0.31 is not "a 31% chance of no effect".

$p = P_{H_0}(\text{test statistic at least as extreme as observed})$; two-sided z: $p = 2\Phi(-|z|)$.

Checkout: $z = 1.01 \Rightarrow p = 0.156 + 0.156 = 0.312$. Under $H_0$, p-values are uniform: $P(p \le 0.05) = 0.05$.

Trap: $p \ne P(H_0 \mid \text{data})$; p is not an effect size and is not stable across reruns.

Quick check: a two-sided z-test gives $z = -2.2$. What is the p-value?

$P(Z \le -2.2) = 0.0139$; two-sided doubles it: $p = 0.028$. If $H_0$ were true, about 2.8% of experiments would give $|z| \ge 2.2$.

The significance level $\alpha$ and the critical value core

Sooner or later you must decide: ship or not. So before the experiment you choose a cut-off for "surprising enough". That cut-off is the significance level $\alpha$ (alpha). The usual choice is $\alpha = 0.05$: reject $H_0$ when $p \le 0.05$.

Choosing $\alpha$ is choosing a false-alarm rate. If $H_0$ is really true, a rule "reject when $p \le 0.05$" will still reject in 5% of experiments, purely by luck. You accept that risk in advance.

The same rule can be written on the test-statistic scale. The value of $z$ whose tails hold exactly $\alpha$ is the critical value: 1.96 for a two-sided test at $\alpha = 0.05$. "$p \le 0.05$" and "$|z| \ge 1.96$" are the same rule, said two ways.

Three ways to say it:

  • Picture: paint the outer tails red, holding $\alpha$ of the null distribution; if your $z$ lands in the red, reject.
  • Numbers: $\alpha = 0.05$, two-sided: red beyond ±1.96. Your $z = 1.01$ is in the white middle.
  • Slogan: $\alpha$ is the false-alarm rate you accept; the critical value is its line in the sand.

The checkout at $\alpha = 0.05$, two-sided.

  1. Split $\alpha$ between the two tails: $0.025$ in each.
  2. Critical value: the $z$ with 0.025 above it, $z^* = \Phi^{-1}(0.975) = 1.96$.
  3. Your $|z| = 1.01 \lt 1.96$: not in the rejection region. Same answer from the p-value: $0.312 \gt 0.05$. Decision: fail to reject $H_0$.
  4. In gap units: the smallest significant gap is about $1.96 \times 1.98 = 3.9$ points. Exactly, with A fixed at 50/500, B would need at least 71 conversions out of 500 (14.2%, $z = 2.04$, $p = 0.042$); 70/500 gives $z = 1.95$, $p = 0.052$, just short.
  5. Other common choices: $\alpha = 0.01$ gives $z^* = 2.576$; $\alpha = 0.10$ gives $1.645$. One-sided at 0.05: $z^* = 1.645$.

The significance level $\alpha$ is the probability of rejecting $H_0$ when $H_0$ is true:

$$\alpha = P(\text{reject } H_0 \mid H_0 \text{ true}).$$

This is also called the Type I error rate (the topic of Chapter 5.7).

The rejection region is the set of test-statistic values that lead to rejection; its probability under $H_0$ is $\alpha$. Its boundary is the critical value:

$$z^* = \Phi^{-1}(1 - \alpha/2) \;\text{(two-sided: reject if } |z| \ge z^*), \qquad z^* = \Phi^{-1}(1 - \alpha) \;\text{(one-sided)}.$$
  • The two decision rules always agree: $p \le \alpha \iff |z| \ge z^*$ (two-sided).
  • $\alpha$ must be fixed before seeing the data. 0.05 is a convention, not a law of nature; choose smaller $\alpha$ when false alarms are expensive.
  • A result with $p \le \alpha$ is called statistically significant at level $\alpha$.
Why do we need it?

A decision needs a rule fixed in advance; otherwise people move the line to wherever their result lands. Fixing $\alpha$ also fixes how often the company ships changes that do nothing, which is a real business cost.

Where is it used?

A/B test plans ("two-sided, $\alpha = 0.05$"), sample-size formulas (the $z_{1-\alpha/2}$ term, Chapter 5.7), critical values in t, chi-square and F tables, confidence levels ($95\% = 1 - \alpha$, Chapter 5.8), and multiple-testing corrections that lower $\alpha$ per test.

How is it used?

Write $\alpha$ in the experiment plan. Get the critical value with scipy.stats.norm.ppf(1 - alpha/2). After the test, report the p-value and the decision, and translate the critical value into business units ("we could only detect lifts above about 3.9 points").

+1.96−1.96 reject0.025 reject0.025 fail to reject H₀0.95 of the H₀ world checkout z = 1.01 α = 0.05 split into two tails; critical values ±1.96 are fixed before the data arrive
The rejection region (red) holds α = 0.05 of the null distribution, 0.025 in each tail. The checkout's z = 1.01 falls in the middle, so H₀ is not rejected.

Move the $\alpha$ slider: the red rejection region grows or shrinks, and the critical value moves (1.96 at 0.05, 2.58 at 0.01, 1.64 at 0.10). Drag the purple $z$: whenever it enters the red, both rules say "reject", and the p-value is at most $\alpha$; they never disagree. The last line translates the critical value into the checkout: how many conversions B would need (out of 500, with A at 50/500) to be significant.

"$\alpha = 0.05$ means there is a 5% chance my significant result is a false positive."

$\alpha$ is the rejection rate among experiments where $H_0$ is true, fixed in advance. The chance that a particular significant result is false depends also on how many tested ideas really work and on the power; it can be far above 5% (Chapter 5.7).

"$p = 0.049$ is a real effect and $p = 0.051$ is no effect."

The evidence in the two is almost identical. The cut-off is a convention for making decisions, not a boundary in nature. Report the number, not just the verdict.

"The result was $p = 0.03$, so let's set $\alpha = 0.05$." (or 0.10 after seeing $p = 0.08$)

Choosing $\alpha$ after seeing $p$ makes the false-alarm rate meaningless. Fix it in the plan.

Your Bayesian framework has its own "line in the sand": a decision threshold such as "declare B the winner if $P(\theta_B \gt \theta_A \mid D) \gt 0.95$". That threshold is not $\alpha$, and it does not automatically give a 5% false-positive rate: its false-positive rate depends on the prior, the sample size and how often you look. If someone asks "what is your false-positive rate?", the honest answer comes from simulating A/A experiments through your decision rule.

$\alpha = P(\text{reject } H_0 \mid H_0 \text{ true})$, chosen before the data (usually 0.05).

Critical value: two-sided $z^* = \Phi^{-1}(1 - \alpha/2) = 1.96$; one-sided 1.645. Reject if $|z| \ge z^*$ $\iff$ $p \le \alpha$.

Checkout: needs $|z| \ge 1.96$, i.e. B ≥ 71/500; we had 60 → fail to reject. Trap: α is not "the chance this result is wrong".

Quick check: at $\alpha = 0.01$ two-sided, would $z = 2.3$ be significant? What about at $\alpha = 0.05$?

At 0.01 the critical value is 2.576, and $2.3 \lt 2.576$: not significant ($p = 0.021 \gt 0.01$). At 0.05 the critical value is 1.96, and $2.3 \ge 1.96$: significant ($p = 0.021 \le 0.05$).

One-sided vs two-sided tests core

A two-sided test asks "is B different from A?" It counts surprises in both directions: a big gain or a big loss. A one-sided test asks "is B better?" and counts only surprises in that one direction.

For the same data, the one-sided p-value is half the two-sided one (when the data go in the predicted direction). That makes one-sided tests tempting. But the choice is only honest if it is made before seeing the data, and only if a result in the other direction would truly lead to the same action as "no difference". If you peek first and then pick the side that the data favour, you are secretly running a two-sided test at double the $\alpha$.

Three ways to say it:

  • Picture: two-sided = paint both tails; one-sided = paint one tail, chosen in advance.
  • Numbers: $z = 1.80$: one-sided $p = 0.036$ (significant at 0.05), two-sided $p = 0.072$ (not).
  • Slogan: pick the direction before the data, never after.
  1. Checkout, two-sided: $p = 2 \times 0.156 = 0.312$. One-sided ($H_1$: B better), decided in advance: $p = P(Z \ge 1.01) = 0.156$. Neither is significant.
  2. A borderline case: suppose $z = 1.80$. One-sided: $p = P(Z \ge 1.80) = 0.036 \lt 0.05$, significant. Two-sided: $p = 0.072 \gt 0.05$, not significant. The choice of sides decides the verdict, which is exactly why it must be made first.
  3. The danger of one-sided: you planned "$H_1$: B better" and B turns out much worse, $z = -3$. One-sided $p = P(Z \ge -3) = 0.9987$. This test cannot flag the harm at all. If harm would matter to you, and in product experiments it usually does, use two-sided.
  4. Cheating, in numbers: look at the data, then test in whichever direction it points, at 0.05 one-sided. You reject whenever $|z| \ge 1.645$, which happens 10% of the time under $H_0$, not 5%.
  • Two-sided: $H_1: \theta \ne \theta_0$; $p = P_{H_0}(|T| \ge |t_{obs}|)$; at $\alpha = 0.05$ reject if $|z| \ge 1.96$.
  • One-sided (right): $H_1: \theta \gt \theta_0$; $p = P_{H_0}(T \ge t_{obs})$; at $\alpha = 0.05$ reject if $z \ge 1.645$.
  • One-sided (left): $H_1: \theta \lt \theta_0$; $p = P_{H_0}(T \le t_{obs})$; reject if $z \le -1.645$.
  • For a symmetric null distribution, one-sided $p$ = two-sided $p$ / 2 when the data fall on the predicted side, and $1 -$ (two-sided $p$)/2 when they fall on the other side.
  • A one-sided test has more power for its chosen direction and none for the other. Use it only when that is truly what the decision needs, and write the direction in the plan.
Why do we need it?

The same data can be "significant" or not depending on the sides. Knowing exactly what each version counts protects you from fooling yourself, and from colleagues who quietly halve p-values.

Where is it used?

Two-sided is the default in most A/B platforms and in scipy/statsmodels. One-sided appears in non-inferiority tests ("B is not worse by more than 1 point"), guardrail checks that only care about harm, and some quality-control tests.

How is it used?

Decide in the plan. In code: proportions_ztest(..., alternative='larger') or scipy.stats.ttest_ind(..., alternative='greater'); the default is 'two-sided'. Watch the order of the groups: "larger" means first minus second.

two-sided: "different?" one-sided: "better?" z = 1.80 z = 1.80 p = 0.036 + 0.036 = 0.072 p = 0.036 not significant at 0.05 significant at 0.05
The same z = 1.80 under two different questions. Counting both tails (left) doubles the p-value. Which question you ask must be fixed before you see z.

Each run simulates 2 000 A/A tests (no real difference, 500 users per group) and applies three rules at level $\alpha$. The first two are honest and reject about $\alpha$ of the time. The third peeks at the data, then runs a one-sided test in whichever direction the data point: it rejects about $2\alpha$ of the time. Press Run 2 000 A/A tests a few times, then move $\alpha$.

"One-sided tests are more powerful, so always use them."

They are more powerful only in the direction you chose, and blind in the other. In product experiments a change that hurts conversion is exactly what you must not miss, so two-sided is the safe default.

"$p = 0.07$ two-sided, so it is significant one-sided."

Switching after seeing the data is the "side picked after looking" rule in the widget: about 10% false alarms at a nominal 5%.

A posterior statement like $P(\theta_B \gt \theta_A \mid D)$ in your framework is one-directional by nature: it only asks about "B better". That is fine for choosing a winner, but it says nothing direct about harm. A practical pattern is to pair it with a guardrail in the other direction, such as $P(\theta_B \lt \theta_A - \delta \mid D)$ on key metrics, decided before the experiment, which mirrors why classical A/B tests default to two-sided.

Two-sided: count both tails ($|z| \ge 1.96$ at 0.05). One-sided: one tail chosen in advance ($z \ge 1.645$).

Same $z$: one-sided $p$ = half the two-sided $p$ (if the data go the predicted way). $z = 1.80$: 0.036 vs 0.072.

Trap: choosing the side after seeing the data doubles the false-positive rate (≈ 10% at "5%").

Quick check: a one-sided test ($H_1$: B better) gives $p = 0.97$. What happened?

The data went the other way: B looked worse ($z \approx -1.88$). A one-sided "better" test cannot reject when B is worse; the large p-value here is not "strong evidence of no difference", it is a sign that you asked a question the data contradicted.

Reject vs fail to reject: absence of evidence is not evidence of absence core

A test has only two possible verdicts:

  • Reject $H_0$: the data would be surprising if $H_0$ were true, so we have evidence against it.
  • Fail to reject $H_0$: the data are not surprising enough. We do not have evidence against $H_0$.

Notice what is missing: "accept $H_0$". A court says "not guilty", never "innocent": maybe the person did it, but the evidence was not strong enough. In the same way, a small experiment can easily miss a real effect. Searching a dark room with a weak torch and not seeing the cat does not prove there is no cat.

Three ways to say it:

  • Picture: a dim torch in a dark room: not seeing something is not the same as it not being there.
  • Numbers: with 500 users per group, a real 2-point lift reaches $p \le 0.05$ in only about 17% of experiments.
  • Slogan: "no evidence of an effect" is not "evidence of no effect".

What can we honestly say about the checkout?

  1. $p = 0.31 \gt 0.05$: fail to reject $H_0$. The result is "not statistically significant".
  2. Wrong conclusion: "the new checkout does not work". Right conclusion: "this experiment could not tell whether it works".
  3. Why: the data are compatible with no difference, but also with a sizeable lift. A rough range of plausible gaps is the estimate ± 2 SE: $2 \pm 1.96 \times 1.98$, about $-1.9$ to $+5.9$ points (this is the 95% confidence interval of Chapter 5.8).
  4. A range from "slightly worse" to "6 points better" is inconclusive, not negative. To learn more you need more users (Chapter 5.7 computes how many).
  • Reject $H_0$ when $p \le \alpha$ (equivalently, when the test statistic is in the rejection region). The result is called statistically significant at level $\alpha$.
  • Fail to reject $H_0$ when $p \gt \alpha$. The result is not statistically significant. This is a statement about the strength of the evidence, not about the truth of $H_0$.
  • How informative a "fail to reject" is depends on the test's power, its chance of rejecting when a real effect of a given size exists (Chapter 5.7). A low-power test fails to reject most of the time even when $H_1$ is true.
  • To claim "no important difference", you need a different design: show that the whole confidence interval lies inside a small "doesn't matter" zone (an equivalence or non-inferiority test).
Why do we need it?

Teams regularly kill good ideas because "the test was not significant", when the test was simply too small to see the effect. The precise wording forces the right follow-up question: was the experiment able to detect an effect of the size we care about?

Where is it used?

Experiment readouts and launch reviews, model-comparison reports ("no significant difference in RMSE"), residual diagnostics (Ljung–Box, normality tests), and any regression table where a coefficient is "not significant".

How is it used?

Write "we did not find evidence of a difference (p = 0.31); the data are compatible with effects from −1.9 to +5.9 points". Check the power for the smallest effect worth caring about; if it was low, the experiment was inconclusive by design.

In a courtroom In a hypothesis test presumed innocentassume H₀ (no difference) evidencedata → test statistic → p-value "beyond reasonable doubt"p ≤ α verdict: guiltyreject H₀ verdict: not guiltyfail to reject H₀ (≠ innocent!)(≠ H₀ true!)
The court analogy. "Not guilty" means the evidence was not strong enough, not that innocence was proven. "Fail to reject H₀" means the same for a test.

In every one of these 200 simulated experiments the new checkout is truly 2 points better (green line). Each dot is one experiment's observed gap. Dots inside the grey band fail to reject $H_0$; green dots reject it. With 500 users per group most dots are grey: the effect is real, yet most experiments "fail to reject". Slide the users up to 4 000 and then 8 000 and watch the green take over. Red dots, if any, are significant in the wrong direction.

"$p = 0.31$, so the new checkout has no effect."

The test failed to find evidence. The data are compatible with no effect and with a lift of several points. The experiment was inconclusive.

"We rejected $H_0$, so $H_1$ is proven."

We found evidence against $H_0$. With $\alpha = 0.05$, rejections still happen in 5% of experiments where $H_0$ is true, and many more of them when most tested ideas do nothing (see the last section).

"We accept the null hypothesis."

Say "fail to reject". Accepting would need a test designed to show the effect is small, with enough power.

"The test was not significant, so there is no difference between A and B."

"The test did not detect a difference. Whether that means much depends on the power: with 500 users per group we only had about a 17% chance to detect a real 2-point lift."

Model answer: "Failing to reject the null means the data are compatible with no effect, not that there is no effect. I would look at the confidence interval for the lift: here it runs from about −1.9 to +5.9 points, so the result is inconclusive. If we need to know, we should rerun with a sample size planned for the smallest lift we care about."

In your forecasting work, a Ljung–Box test on the residuals that "fails to reject" does not prove the residuals are white noise; with a short history it has little power to detect mild autocorrelation (Chapter 7.17). Look at the ACF plot as well. The same logic applies when a changepoint or a holiday effect "is not significant": it may be real but poorly measured with the data you have.

Two verdicts only: reject $H_0$ ($p \le \alpha$) or fail to reject ($p \gt \alpha$). Never "accept $H_0$".

Absence of evidence ≠ evidence of absence; how much a "no" means depends on power.

Checkout: fail to reject; plausible gaps about −1.9 to +5.9 points → inconclusive, not negative.

Quick check: a 3-day test with 80 users per group shows a 4-point lift, $p = 0.45$. A colleague says "the feature doesn't work". Reply in one sentence.

"With 80 users per group this test could only detect very large lifts, so $p = 0.45$ means we have not learned much, not that the feature does nothing; the plausible range for the lift is wide and includes both zero and large gains."

Statistical significance is not importance core

In everyday English "significant" means "important". In statistics it means only "unlikely to be pure luck". These are completely different things.

With millions of users, even a microscopic lift (0.1 percentage points) becomes statistically significant, because the SE becomes tiny. With a few hundred users, a huge lift (3 points) can fail to be significant, because the SE is huge. Whether a result matters is a business question: is the lift big enough to be worth shipping, maintaining and the risk? The p-value cannot answer that. The size of the effect and its uncertainty can.

Three ways to say it:

  • Picture: a very precise measurement of a very small thing.
  • Numbers: 10.0% vs 10.1% with 5 million users per group: $z = 5.3$, $p \approx 0.0000001$, but the lift is only 0.1 points.
  • Slogan: significant means detectable, not important.
  1. Tiny but significant. 5 000 000 users per group, 10.0% vs 10.1%. Pooled rate 0.1005; $SE_0 = \sqrt{0.1005 \times 0.8995 \times 2/5\,000\,000} = 0.00019$; $z = 0.001/0.00019 = 5.3$; $p = 1.5 \times 10^{-7}$. Extremely significant, yet if the team needs at least a 0.5-point lift to pay for the change, it is not worth shipping.
  2. Big but not significant. 200 users per group, 10% vs 13% (20 vs 26 conversions). $SE_0 = 0.032$; $z = 0.03/0.032 = 0.94$; $p = 0.35$. A 3-point lift would be great news if real, but this experiment cannot tell it apart from luck.
  3. The fix in both cases: look at the estimate with its uncertainty (estimate ± 1.96 SE) next to the smallest effect that matters for the business.
  • Statistical significance: $p \le \alpha$. It depends on the effect size, the noise and the sample size together.
  • Practical significance (importance): the effect is at least as large as the smallest effect worth acting on, often called $\delta$ (or the minimum detectable effect when planning, Chapter 5.10). It is set by the business, before the test.
  • Report the effect size (absolute lift, relative lift, Chapter 5.7) and its interval (Chapter 5.8) alongside $p$. The four possible situations:
effect ≥ δ (important)effect < δ (unimportant)
significantreal and worth it: shipreal but too small to matter
not significantcould be big, poorly measured: get more datainterval tight around small values: confidently unimportant
Why do we need it?

Big platforms run experiments on millions of users, where almost everything is "significant". Without a practical threshold, teams ship changes that cost more to maintain than they earn, or argue over p-values instead of business value.

Where is it used?

Launch criteria in A/B testing ("lift ≥ 0.5 points and significant"), MDE planning, non-inferiority tests for guardrail metrics, feature selection in regression (a "significant" coefficient that barely changes predictions), and model comparisons on large test sets.

How is it used?

Before the test, agree on $\delta$. After it, report "lift = 0.10 points (95% interval 0.06 to 0.14), p < 0.001; below our 0.5-point threshold, so we will not ship". Decide from the interval and $\delta$, not from $p$ alone.

0 (no effect) δ (worth it) significant and important significant, too small to matter not significant: unknown,get more data not significant and confidentlyunimportant huge n, big lift huge n, tiny lift small n huge n, ~0 lift
Each bar is an estimate ± about 2 SE. "Significant" asks whether the bar avoids 0; "important" asks where it sits relative to the business threshold δ. You need both questions.

Pretend the experiment measured the true lift exactly (blue dot); the blue bar is ± 1.96 SE around it. The grey line is "no effect", the green line is the smallest lift worth shipping. Press the four presets. In Tiny but significant the bar is far from 0 (significant) but far below the green line (unimportant). In Big but not significant the dot is above the green line but the bar reaches below 0. Then move the users slider and see that $p$ changes enormously while the lift does not change at all.

"$p \lt 0.001$: a huge effect!"

A tiny p-value can come from a tiny effect measured on a huge sample. Look at the lift itself.

"Not significant, so the effect is small."

A large effect measured on a small sample is often not significant. The interval tells you whether large effects are still possible.

"Significant means we should ship."

Ship if the lift is large enough to matter (and guardrails are fine). Significance only says the lift is probably not zero.

"The result is highly significant, so the new feature has a big impact."

"The result is statistically significant, meaning a lift this size would be unlikely if there were no real difference. The lift itself is 0.1 points, below our 0.5-point launch threshold."

Model answer: "Statistical significance is about whether an effect is distinguishable from noise; practical significance is about whether it is big enough to matter. With large samples, small effects become significant, so I always report the effect size with its interval and compare it with a threshold agreed before the test."

Your Bayesian A/B framework can build importance directly into the decision: instead of $P(\theta_B \gt \theta_A \mid D)$, compute $P(\theta_B - \theta_A \gt \delta \mid D)$ for a practically meaningful $\delta$ (Chapter 6.4). In your forecasting model, a regressor whose coefficient is clearly non-zero may still improve the holdout error by almost nothing; judge it by forecast accuracy on a time-ordered holdout (Chapter 7.15), not by its significance.

Statistically significant = $p \le \alpha$ = "probably not zero". Practically significant = effect ≥ δ (set by the business).

Huge $n$: tiny effects become significant. Small $n$: big effects may not be.

Always report the effect and its interval next to $p$.

Quick check: a lift of 0.02 points is significant with $p = 0.003$ on 50 million users. The launch threshold is 0.3 points. Ship?

No (on this metric alone). The lift is real but about 15 times smaller than the smallest lift worth shipping. Significance does not make a small effect important.

The whole two-proportion z-test, step by step core

Now every piece has a name. Put together, the classical A/B test of two conversion rates is a short chain, in this order: control → treatment → observed difference → pooled proportion → standard error → z → p-value → decision. Each step answers one small question, and each step uses only the previous ones.

Three ways to say it:

  • Picture: measure the gap, build the luck-ruler in the "no difference" world, count rulers, read the tail area.
  • Numbers: 10% vs 12% → gap 0.02 → pooled 0.11 → SE 0.0198 → $z = 1.01$ → $p = 0.31$ → fail to reject.
  • Slogan: gap, ruler, count, area, decide.
  1. Data. Control (old): $k_A = 50$ of $n_A = 500$. Treatment (new): $k_B = 60$ of $n_B = 500$.
  2. Rates and gap. $\hat p_A = 0.10$, $\hat p_B = 0.12$, observed difference $\hat p_B - \hat p_A = 0.02$.
  3. Hypotheses (written before the data): $H_0: p_A = p_B$, $H_1: p_A \ne p_B$, $\alpha = 0.05$, two-sided.
  4. Pooled proportion (the shared rate under $H_0$): $(50 + 60)/(500 + 500) = 0.11$.
  5. Standard error under $H_0$: $\sqrt{0.11 \times 0.89 \times (1/500 + 1/500)} = \sqrt{0.000392} = 0.0198$.
  6. Test statistic: $z = 0.02/0.0198 = 1.01$.
  7. p-value: $2 \times P(Z \ge 1.01) = 2 \times 0.156 = 0.31$.
  8. Decision: $0.31 \gt 0.05$, fail to reject $H_0$. Report: "+2.0 points (roughly −1.9 to +5.9), p = 0.31: inconclusive."

Two-proportion z-test for $H_0: p_A = p_B$:

$$\hat p_{pool} = \frac{k_A + k_B}{n_A + n_B}, \qquad z = \frac{\hat p_B - \hat p_A}{\sqrt{\hat p_{pool}(1 - \hat p_{pool})\left(\frac{1}{n_A} + \frac{1}{n_B}\right)}}, \qquad p = 2\,\Phi(-|z|).$$

Assumptions (check them, do not just compute):

  • Users were randomly assigned, and each user is counted once and is independent of the others (the unit of analysis is the unit of randomization, Chapter 5.10).
  • Enough conversions and non-conversions in each group for the Normal approximation (a rule of thumb: at least about 10 of each).
  • The sample size was fixed in advance and the data were analysed once, not checked every day until significant (peeking, Chapter 5.11).
Why do we need it?

It is the single most common test in product analytics: almost every conversion experiment is read with it, and every interviewer expects you to compute it by hand and explain each step.

Where is it used?

A/B tests on conversion, click-through, sign-up and retention rates; the default in statsmodels.stats.proportion.proportions_ztest; experimentation platforms; and as the baseline your Bayesian framework's results get compared with.

How is it used?

z, p = proportions_ztest([60, 50], [500, 500]) gives $z = 1.011$, $p = 0.312$. Report the lift, its interval, $p$ and the decision against the pre-registered $\alpha$, and check the assumptions above.

Press Next to build the test in seven steps: data, rates and gap, hypotheses, pooled rate (purple line), the SE-ruler, the test statistic, and the p-value with the decision. Then change the conversion sliders (both groups have 500 users) and step through again: try B = 71 to cross the 0.05 line, or B = 40 to see a significant result in the other direction.

"I computed z, so I am done."

The formula is only valid if users were randomized, counted once, independent, numerous enough, and the data were analysed once at a planned sample size. A perfect z from a broken experiment is worthless.

"z is significant, so B is better."

Check the sign. A two-sided test also flags B being significantly worse (try B = 40 in the widget).

When you present your Bayesian A/B framework, interviewers often ask "what would a classical test say?" For the checkout: $z = 1.01$, two-sided $p = 0.31$, not significant; your model with flat priors says $P(\theta_B \gt \theta_A \mid D) \approx 0.84$. Being able to produce both, and to explain that they answer different questions about the same data, shows you understand both frameworks.

Control → treatment → gap → pooled $\hat p$ → $SE_0$ → $z$ → p → decide.

Checkout: 0.10 vs 0.12 → 0.02 → 0.11 → 0.0198 → 1.01 → 0.31 → fail to reject.

Trap: the formula assumes randomization, independent users, enough counts and no peeking.

Quick check: 120 of 1 000 vs 150 of 1 000. Run the whole test.

Rates 0.12 and 0.15, gap 0.03. Pooled $270/2000 = 0.135$. $SE_0 = \sqrt{0.135 \times 0.865 \times 0.002} = 0.0153$. $z = 0.03/0.0153 = 1.96$. $p = 0.0496$: just below 0.05, so reject at $\alpha = 0.05$, by a whisker. Report the lift with its interval; a result this close to the line deserves caution.

What a p-value is not: the classic misreadings core

The p-value answers one narrow question: if nothing were going on, how surprising would data like mine be? Almost every misreading comes from turning that question around into "given my data, how likely is it that nothing is going on?". Those are different questions, like "how likely is a wet street if it rained?" (very) versus "how likely is it that it rained if the street is wet?" (depends: maybe a cleaning truck went by).

To answer the turned-around question you need extra information the test never uses: how often the ideas you test are actually good (a prior), and how good the test is at spotting them (power). That is exactly what a Bayesian analysis adds.

Three ways to say it:

  • Picture: $P(\text{wet street} \mid \text{rain})$ is not $P(\text{rain} \mid \text{wet street})$.
  • Numbers: if only 10% of tested ideas work and power is 50%, then about 47% of "significant" results are false alarms, even at $\alpha = 0.05$.
  • Slogan: the p-value is about the data given $H_0$, never about $H_0$ given the data.

1 000 experiments, counted as people (natural frequencies, as in Chapter 4.3). Suppose 10% of the ideas a team tests really work, each test has 50% power to detect a working idea, and $\alpha = 0.05$.

  1. 100 ideas really work; 50% power → 50 significant results, 50 missed.
  2. 900 ideas do nothing; $\alpha = 0.05$ → $0.05 \times 900 = 45$ significant results by luck, 855 correctly not significant.
  3. Total significant: $50 + 45 = 95$. Of these, 45 are false alarms: $45/95 = 47\%$.
  4. So "p ≤ 0.05" here means roughly a coin flip that the idea really works. The 5% describes how often a useless idea passes, not how often a passing idea is useless.

The p-value is $P(T \text{ at least as extreme as observed} \mid H_0)$. It is not:

  • the probability that $H_0$ is true, $P(H_0 \mid \text{data})$;
  • the probability that $H_1$ is true, or that B is better ($1 - p$ is not that either);
  • the probability that the result "is due to chance";
  • the probability of a false positive for this result (that is $P(H_0 \mid \text{significant})$, which depends on the prior share of true effects and the power);
  • a measure of effect size or importance;
  • stable: rerunning the same experiment can easily give a very different p-value.

$P(H_0 \mid \text{data})$ needs Bayes' theorem: a prior probability for $H_0$ and the likelihood of the data under each hypothesis (Chapter 6.1).

Why do we need it?

Misread p-values lead to over-confident launches ("95% sure it works") and to killed ideas ("p = 0.3 proves it does nothing"). Interviewers test this more than any formula, because it shows whether you understand what a test can and cannot say.

Where is it used?

Every experiment readout and stakeholder conversation, the replication debates in science, false-discovery reasoning behind Benjamini–Hochberg (Chapter 5.11), and the case for Bayesian A/B testing, which answers $P(\text{B better} \mid \text{data})$ directly.

How is it used?

Say the definition in one sentence that starts with "If there were no difference…". Never write "probability that the null is true". When the team needs "how likely is B better?", use a posterior (your framework) or at least report the estimate with its interval.

1 000 ideas 10%90% 100 really work 900 do nothing power 50%α = 5% 50 significant ✓50 missed 45 significant ✗855 correctly not 95 significant:45 are false(47%)
Why α = 0.05 is not "a 5% chance this significant result is wrong". When most tested ideas do nothing, false alarms make up a large share of the significant results. The share depends on the prior rate of good ideas and on power, which the p-value never sees.

Every card is a sentence someone might say about the checkout result ($z = 1.01$, $p = 0.31$) or about p-values in general. Press True or False, read the explanation, then Next card. Ten cards; your score is kept below. Try to say the corrected sentence out loud before reading it.

"$1 - p$ is the probability that the effect is real."

No such shortcut exists. The probability that the effect is real needs a prior and a posterior. For the checkout, a flat-prior posterior gives $P(\theta_B \gt \theta_A \mid D) \approx 0.84$, which is not $1 - 0.31 = 0.69$.

"A significant result at $\alpha = 0.05$ is wrong only 5% of the time."

That is $P(\text{significant} \mid H_0)$. The share of significant results that are false depends on how many tested ideas are real and on power; in the tree above it is 47%.

"A p-value of 0.03 means there is a 97% probability that our variant is better."

"A p-value of 0.03 means that if there were no real difference, data at least this extreme would occur about 3% of the time. It is evidence against 'no difference', not a probability that the variant is better."

Model answer: "The p-value conditions on the null; it measures how surprising the data are under it. The probability that the null is true given the data is a different quantity, which needs a prior. If the business wants 'the probability that B beats A', a Bayesian posterior gives that directly; that is what our A/B framework reports."

This is the strongest argument for the framework you built: stakeholders almost always want "the probability B is better", and your model gives exactly that, $P(\theta_B \gt \theta_A \mid D)$, along with the full posterior of the lift. The price is that the answer depends on the prior, so be ready to explain your prior choice and show that conclusions are stable under reasonable alternatives (Chapter 6.8).

p = P(data at least this extreme | $H_0$). NOT P($H_0$ | data), not P(B better), not 1 − that, not the chance "it was luck", not an effect size, not stable across reruns.

α = P(significant | $H_0$) ≠ P($H_0$ | significant) (10% good ideas, 50% power → 47% of "wins" are false).

Trap: turning the conditional around. The question the business wants needs a prior (Bayes).

Quick check: if every idea a team tests is useless, what fraction of their significant results (at α = 0.05) are false positives?

All of them: 100%. About 5% of tests will be significant, and every one of those is a false alarm. α controls how often useless ideas pass, not what fraction of passing ideas are useless.

Recap, cheat sheet and practice

  • A test asks one question: is the result bigger than the luck I would normally expect? It assumes the boring claim $H_0$ and checks how surprising the data would be.
  • $H_0$ is a precise "nothing happening" claim about a parameter ($p_A = p_B$, $\mu = 175$); $H_1$ is what you suspect instead. Both, and the direction, are fixed before the data.
  • A statistic is any number computed from the sample. A test statistic measures the distance from $H_0$ in SE units: $(\text{estimate} - \text{null value})/SE$. Checkout $z = 1.01$; heights $t = -0.87$.
  • The pooled proportion (110/1 000 = 0.11) estimates the one shared rate under $H_0$; it builds the test's $SE_0 = 0.0198$. Confidence intervals use the unpooled SE (also 0.0198 here, because the rates are close).
  • The null distribution is the test statistic's distribution if $H_0$ is true: $N(0,1)$ for z, $t_{n-1}$ for t, or by simulation / shuffling (exact permutation: $p = 10/70$).
  • The p-value is the probability under $H_0$ of a result at least as extreme as yours: $2\Phi(-1.01) = 0.31$. Under $H_0$ p-values are uniform.
  • Reject when $p \le \alpha$ $\iff$ $|z| \ge z^*$ (1.96 two-sided at 0.05; 1.645 one-sided). Choose sides and $\alpha$ before looking; picking the side afterwards doubles false alarms.
  • "Fail to reject" ≠ "accept": absence of evidence is not evidence of absence. "Significant" ≠ "important": compare the effect and its interval with a business threshold.
  • p is $P(\text{data} \mid H_0)$, never $P(H_0 \mid \text{data})$; that needs a prior, which is what a Bayesian A/B framework adds.

Cheat sheet

ItemFormula / ruleCheckout (50/500 vs 60/500)
Hypotheses$H_0: p_A = p_B$, $H_1: p_A \ne p_B$two-sided, $\alpha = 0.05$
Pooled proportion$(k_A + k_B)/(n_A + n_B)$110/1 000 = 0.11
SE under $H_0$$\sqrt{\hat p(1-\hat p)(1/n_A + 1/n_B)}$0.0198
Test statistic$z = (\hat p_B - \hat p_A)/SE_0$0.02/0.0198 = 1.01
p-valuetwo-sided $2\Phi(-|z|)$; one-sided $\Phi(-z)$0.312; 0.156
Critical value$\Phi^{-1}(1-\alpha/2)$; one-sided $\Phi^{-1}(1-\alpha)$1.96 (B would need ≥ 71/500)
Decisionreject if $p \le \alpha$; else fail to rejectfail to reject
One-sample t$t = (\bar x - \mu_0)/(s/\sqrt n)$, df $n-1$heights: −0.87, p = 0.48
Code it · Python

import itertools
import numpy as np
from scipy import stats
from statsmodels.stats.proportion import proportions_ztest

# 1. The checkout two-proportion z-test, by hand
kA, nA, kB, nB = 50, 500, 60, 500
pA, pB = kA / nA, kB / nB
p_pool = (kA + kB) / (nA + nB)                      # one shared rate under H0
se0 = np.sqrt(p_pool * (1 - p_pool) * (1 / nA + 1 / nB))
z = (pB - pA) / se0                                 # the test statistic
p_two = 2 * stats.norm.sf(abs(z))                   # two-sided p-value
p_one = stats.norm.sf(z)                            # one-sided (H1: B better), only if chosen in advance
print(p_pool, se0, z, p_two, p_one)                 # 0.11 0.019789 1.0107 0.3122 0.1561

# 2. The same with statsmodels (pooled SE by default; it tests first group minus second)
print(proportions_ztest([kB, kA], [nB, nA]))                          # (1.0107, 0.3122)
print(proportions_ztest([kB, kA], [nB, nA], alternative='larger'))   # (1.0107, 0.1561)

# 3. Critical values
print(stats.norm.ppf(0.975), stats.norm.ppf(0.95), stats.norm.ppf(0.995))   # 1.960 1.645 2.576

# 4. The height claim: one-sample t-test of H0: mu = 175
print(stats.ttest_1samp([160, 170, 180], popmean=175))   # t = -0.866, p = 0.478, df = 2

# 5. Exact permutation test: items bought on the old vs the new page
old = np.array([3, 5, 2, 4])
new = np.array([6, 4, 7, 5])
x = np.concatenate([old, new])
obs = new.mean() - old.mean()
diffs = []
for idx in itertools.combinations(range(8), 4):     # all 70 ways to choose the "new" group
    m = np.zeros(8, dtype=bool)
    m[list(idx)] = True
    diffs.append(x[m].mean() - x[~m].mean())
diffs = np.array(diffs)
print(obs, len(diffs), np.mean(np.abs(diffs) >= abs(obs) - 1e-12))   # 2.0 70 0.1429 (= 10/70)
res = stats.permutation_test((new, old), lambda a, b, axis: a.mean(axis=axis) - b.mean(axis=axis),
                             permutation_type='independent', n_resamples=np.inf, vectorized=True)
print(res.pvalue)                                   # 0.1429: SciPy agrees

# 6. A/A simulation: what p-values look like when H0 is true
rng = np.random.default_rng(0)
R = 100_000
a = rng.binomial(500, 0.11, R)
b = rng.binomial(500, 0.11, R)
pp = (a + b) / 1000
zz = (b - a) / 500 / np.sqrt(pp * (1 - pp) * 2 / 500)
pv = 2 * stats.norm.sf(np.abs(zz))
print((pv <= 0.05).mean())                          # 0.051: the honest rule, about alpha
print((np.abs(zz) >= 1.645).mean())                 # 0.101: "pick the side after looking", about 2 * alpha
Test yourself

1. Which sentence is the correct meaning of $p = 0.31$ for the checkout?

The p-value is computed in the $H_0$ world: $P(\text{result at least this extreme} \mid H_0)$. It is not a probability that $H_0$ or $H_1$ is true.

2. The gap is 0.02 and the SE under $H_0$ is 0.0198. What is the test statistic?

Test statistic = (estimate − null value)/SE = $(0.02 - 0)/0.0198 = 1.01$: the gap measured in standard errors. 0.31 is the p-value; 0.02 is the raw gap.

3. A has 30 of 200 conversions (15%), B has 60 of 300 (20%). What is the pooled proportion?

Pool the counts: $(30 + 60)/(200 + 300) = 90/500 = 0.18$. It is a weighted average; the bigger group B pulls it toward 20%.

4. The checkout test gives $p = 0.31$. Which conclusion is right?

Not significant means "no evidence against $H_0$", not "evidence for $H_0$". With 500 users per group the test had only about 17% power for a 2-point lift.

5. What is the critical value of a two-sided z-test at $\alpha = 0.05$?

Two-sided splits $\alpha$ into 0.025 per tail: $\Phi^{-1}(0.975) = 1.96$. 1.645 is the one-sided value at 0.05; 2.576 is two-sided at 0.01.

6. An analyst always looks at the data first and then runs a one-sided test at 0.05 in the direction the data point. Under $H_0$, how often do they declare significance?

They reject whenever $|z| \ge 1.645$, which has probability 0.10 under $H_0$: a two-sided test at level 0.10 disguised as one-sided at 0.05.

Practice problems

A. A team asks "does free shipping change the average order value?" Write $H_0$ and $H_1$, say one- or two-sided, and name a test statistic.

Parameter: the true difference in mean order value, $\mu_B - \mu_A$ (B = free shipping). $H_0: \mu_B - \mu_A = 0$; $H_1: \mu_B - \mu_A \ne 0$, two-sided because "change" includes both directions (and a drop would matter). Test statistic: $t = (\bar x_B - \bar x_A)/\sqrt{s_A^2/n_A + s_B^2/n_B}$ (Welch, Chapter 5.9).

B. Same rates as the checkout, but 200 of 2 000 vs 240 of 2 000. Run the whole test. What changed, and why?

Rates 0.10 and 0.12, gap 0.02, pooled $440/4000 = 0.11$. $SE_0 = \sqrt{0.11 \times 0.89 \times (2/2000)} = \sqrt{0.0000979} = 0.0099$. $z = 0.02/0.0099 = 2.02$, $p = 0.043$: reject at 0.05. Four times the users halves the SE (the √n law), so the same gap is twice as many SEs. The effect did not change; the evidence did.

C. Interview: "Explain a p-value of 0.03 to a product manager in two sentences."

"If the change truly did nothing, we would see a difference this large or larger in only about 3 of every 100 experiments like this one, so luck alone is an unlikely explanation. It does not tell us how big the improvement is or the probability that it is real; for that, look at the estimated lift and its interval."

D. Interview: "Why does the z-test use a pooled proportion when the confidence interval does not? And why did both SEs come out 0.0198 for the checkout?"

"The test is computed assuming $H_0$, which says both groups share one rate; the best estimate of that rate pools everyone (110/1 000 = 0.11), giving $SE_0$. The interval describes the real difference without assuming $H_0$, so it uses each group's own rate. They agree here (0.01979 vs 0.01978) because 10% and 12% are close and the groups are equal in size; they diverge when rates are far apart or group sizes very unequal."

E. Your height claim: sample 160, 170, 180; $H_0: \mu = 175$. Do the full one-sample t-test and explain why it gives a larger p-value than the z version.

$\bar x = 170$, $s = 10$, $\widehat{SE} = 10/\sqrt 3 = 5.77$, $t = (170 - 175)/5.77 = -0.87$. With $n - 1 = 2$ degrees of freedom, two-sided $p = 0.48$ (the Normal curve would give 0.39). The t curve has fatter tails because the SE itself was estimated from only 3 values, so more extreme t-values happen by luck. Either way: fail to reject; three people tell you almost nothing about 175 vs 170.

F. Interview: "Our test on 10 million users shows a 0.05-point lift with p = 0.0001. Should we ship?"

"Statistically it is almost certainly not zero, but significance is not importance. I would compare the lift and its interval with the smallest lift worth shipping (agreed before the test), check guardrail metrics and costs, and check the experiment's health (sample ratio, novelty effects, Chapter 5.11). If 0.05 points is below the threshold, I would not ship on this metric alone."

Chapter 5.7 · Syllabus Module 14

Errors, power, effect size and sample size

A hypothesis test can be wrong in two different ways: it can shout "winner!" when nothing changed, or stay silent when something really did. This chapter gives both mistakes a name and a number, shows how to measure "how big" an effect is, and turns all of it into the most practical question in experimentation: how many users do I need?

  • Name the two errors (Type I = false positive, Type II = false negative), put them in the 2×2 table, and explain them with a court case
  • Explain what $\alpha$ promises, and check it with 1 000 simulated A/A tests (about 5% come out "significant")
  • Define power $= 1-\beta$ and compute it for the checkout test: was 500 users per group even able to see a 2-point lift?
  • Describe an effect size three ways: absolute lift (points), relative lift (%), and Cohen's $d$ (in standard deviations)
  • Derive the power formula and the sample-size formula $n = 2\sigma^2(z_{1-\alpha/2}+z_{1-\beta})^2/\Delta^2$, and see how effect, $n$, noise and $\alpha$ move power
  • See why low power does double damage: real effects are missed, and the "winners" you do find look bigger than they really are

What we need from earlier chapters: the logic of a test (null hypothesis $H_0$, alternative $H_1$, test statistic, significance level $\alpha$, critical value, p-value) from Chapter 5.6; standard errors, including $SE = \sqrt{p(1-p)/n}$ for a proportion and $\sqrt{SE_1^2 + SE_2^2}$ for a difference (Chapter 5.5); the Normal distribution and its CDF $\Phi$ (Chapter 4.9); the CLT (Chapter 4.13). Notation: $\Phi(z)$ is the probability that a standard Normal is below $z$; $z_q$ is its $q$-quantile, so $z_{0.975} = 1.96$ and $z_{0.8} = 0.84$. Our running example is the checkout test from Chapter 5.6: 50 of 500 converted with the old checkout (A) and 60 of 500 with the new one (B), $z \approx 1.01$, two-sided $p \approx 0.31$.

Two ways to be wrong: Type I and Type II errors core

Think of a court case. The rule is "innocent until proven guilty": the court starts from the boring assumption (no crime) and only drops it when the evidence is strong. The verdict can go wrong in two ways:

  • The court convicts an innocent person. Nothing happened, but the court acted as if it did. This is a Type I error, a false positive, a false alarm.
  • The court lets a guilty person go. Something really happened, but the evidence was too weak to show it. This is a Type II error, a false negative, a miss.

An A/B test is the same court. "Innocent" is the null hypothesis: the new checkout changes nothing. "Convict" is "reject the null and ship the new checkout". A court that demands very strong evidence convicts few innocent people, but it also lets more guilty people walk free. You cannot push both mistakes down at once just by moving the bar.

Three ways to say it:

  • Picture: a smoke alarm that beeps when you make toast (Type I) or stays silent in a real fire (Type II).
  • Numbers: with $\alpha = 0.05$, about 1 in 20 changes that do nothing still look like winners; with 500 users per group, about 83 in 100 real 2-point lifts are missed.
  • Slogan: Type I = crying wolf when there is no wolf; Type II = not seeing the wolf that is there.

A year of experiments. A team tests 100 ideas. Suppose (we can never know this in real life) that 20 ideas really work and 80 do nothing. Every test uses $\alpha = 0.05$ and has power 0.80, which means it detects a real effect 80% of the time (we define power properly in a moment).

  1. The 80 useless ideas: each is wrongly called a winner with probability $\alpha = 0.05$. Expected false alarms: $80 \times 0.05 = 4$ (Type I errors).
  2. So $80 - 4 = 76$ useless ideas are correctly called "no difference".
  3. The 20 real ideas: each is detected with probability 0.80. Expected detections: $20 \times 0.80 = 16$.
  4. So $20 - 16 = 4$ real ideas are missed (Type II errors).
  5. Winners declared: $4 + 16 = 20$. Of these, $4$ are fake: $4/20 = 20\%$ of the "winners" are false alarms.

Notice the last step: $\alpha$ is 5%, but 20% of the declared winners are fake. $\alpha$ is not "the chance that a winner is fake"; that share depends on how many ideas really work and on power. (You will meet this share again as the false discovery rate in Chapter 5.11.)

A test chooses between the null hypothesis $H_0$ ("no effect", e.g. $\theta_B = \theta_A$) and the alternative $H_1$ ("there is an effect", e.g. $\theta_B \ne \theta_A$). The decision is either "reject $H_0$" or "do not reject $H_0$".

  • A Type I error is rejecting $H_0$ when $H_0$ is true. Its probability is the significance level: $\alpha = P(\text{reject } H_0 \mid H_0 \text{ true})$. You choose $\alpha$ before the test (often 0.05).
  • A Type II error is not rejecting $H_0$ when $H_1$ is true. Its probability is $\beta = P(\text{do not reject } H_0 \mid H_1 \text{ true})$. Unlike $\alpha$, $\beta$ is not one number: it depends on how big the true effect is.
  • The power of the test is $1 - \beta = P(\text{reject } H_0 \mid H_1 \text{ true, with a given effect size})$: the chance of catching a real effect of that size.
$H_0$ is true (no effect)$H_1$ is true (real effect)
Reject $H_0$ ("ship it")Type I error, false positive, probability $\alpha$correct detection, probability $1-\beta$ (power)
Do not reject ("keep the old")correct, probability $1-\alpha$Type II error, false negative, probability $\beta$

Both $\alpha$ and $\beta$ are conditional probabilities: each one assumes a particular truth (a column of the table). Careful with the letter: this $\beta$ has nothing to do with regression coefficients $\beta$ or the Beta distribution.

Why do we need it?

"Is the test correct?" is the wrong question: any test with noisy data will sometimes be wrong. The right question is "how often is it wrong, in which direction, and what does each mistake cost?". Shipping a useless change and missing a good one have very different costs.

Where is it used?

Every A/B test and experiment platform (choosing $\alpha$ and power), clinical trials, fraud and spam filters (false alarms vs misses), medical screening, quality control, and the precision/recall trade-off of any binary classifier in ML.

How is it used?

Before the test, write down $\alpha$ (the false-alarm rate you accept) and the power you want for the smallest effect worth finding. Then choose the sample size to deliver both (later in this chapter). After the test, remember which mistake is still possible: a "no" may be a miss.

Truth: H₀ true no effect · "innocent" Truth: H₁ true real effect · "guilty" Reject H₀ "ship it" · "convict" Do not reject "keep old" · "acquit" Type I error false positive, false alarm probability α Correct: detected true positive probability 1 − β = power Correct: no false alarm true negative probability 1 − α Type II error false negative, a miss probability β
The truth picks the column (we never see it); the test picks the row. Each column's two probabilities add up to 1: $\alpha + (1-\alpha)$ on the left, $(1-\beta) + \beta$ on the right.

Each dot is one experiment. In this simulation we know the truth: the top rows are ideas that really work, the rest do nothing. Press Run 100 new experiments a few times and watch the red dots (false alarms) and orange dots (misses) move around. Then slide ideas that really work down to 5%: notice how the share of fake "winners" shoots up even though $\alpha$ stayed at 5%. Lower the power and watch the orange misses multiply.

The blue curve is where the test statistic lands when nothing changed ($H_0$); the orange curve is where it lands when the change really works ($H_1$). You reject $H_0$ when the statistic is to the right of the purple bar (a one-sided test, to keep the picture simple). Drag the purple bar to the right: the red area (Type I, $\alpha$) shrinks but the orange area (Type II, $\beta$) grows. Now raise distance between the curves: both errors can shrink together, which is what more data or a bigger effect buys you.

"α = 0.05 means 5% of my significant results are false positives."

$\alpha = P(\text{reject} \mid H_0 \text{ true})$: 5% of the tests where nothing changed come out significant. The share of significant results that are fake can be much larger (20% in the example above) when few ideas really work or power is low.

"We did not reject $H_0$, so there is no effect."

A Type II error is possible, and with a small sample it is likely. "Not rejected" means "not enough evidence", not "proved equal".

"We can make both errors tiny just by choosing the right α."

With the data fixed, moving the bar trades one error for the other. Only more information (more users, less noise, or a bigger true effect) lowers both at once.

Your A/B framework makes decisions with posterior quantities such as $P(\theta_B \gt \theta_A \mid D)$, not with p-values. But a rule like "ship B when $P(\theta_B \gt \theta_A \mid D)$ is above some threshold" still makes both kinds of mistakes: it can ship a variant that is really no better (a false positive) or keep the old one when B really is better (a miss). The 2×2 table is the language reviewers will use to ask "how often does your rule ship a useless change?". You can answer it by simulation, as the next section shows.

"Type I is when the test is wrong, Type II is when the data are wrong."

Both are wrong decisions caused by random data. Type I: reject a true null (false positive). Type II: keep a false null (false negative).

Model answer: "A Type I error is a false positive: we declare an effect that is not there; its rate is α, which we choose. A Type II error is a false negative: we miss a real effect; its rate β depends on the true effect size, the sample size and the noise. Power is 1 − β. For fixed data, lowering α raises β; to lower both we need more data or less noise."

Type I = reject a true $H_0$ (false positive), rate $\alpha$. Type II = keep a false $H_0$ (miss), rate $\beta$. Power $= 1-\beta$.

$\alpha$ is chosen; $\beta$ depends on the true effect, $n$ and the noise.

Trap: $\alpha$ is not "the chance a winner is fake", and "not significant" is not "no effect".

Quick check: a spam filter flags a real email from your manager as spam. Which error is this, if $H_0$ is "the email is not spam"?

The filter rejected $H_0$ (called it spam) while $H_0$ was true (it was a real email). That is a Type I error, a false positive. Letting a real spam email into the inbox would be a Type II error.

What α promises, checked with 1 000 A/A tests core

Split your users into two random groups, but show both groups the same old checkout. This is an A/A test. Nothing differs between the groups except luck, so every "difference" you measure is pure noise.

If your testing machinery is honest and you use $\alpha = 0.05$, it should call a winner in about 5% of A/A tests: no fewer, no more. That is literally what $\alpha$ promises. An A/A test is a fire drill for your whole pipeline: the random split, the logging, the formula for the standard error.

Three ways to say it:

  • Picture: weigh the same bag of flour twice on the same scale. Any difference is the scale's wobble, never the flour.
  • Numbers: 1 000 A/A tests at $\alpha = 0.05$ give about 50 "significant" results (usually somewhere between 37 and 64).
  • Slogan: when nothing changed, a p-value is just a random number between 0 and 1.

Why exactly 5%? Follow one A/A test, using the two-sided z-test of Chapter 5.6.

  1. Nothing changed, so $H_0$ is true and the test statistic $Z$ is (approximately) standard Normal.
  2. We reject when $|Z| \gt 1.96$. The chance of that under $H_0$ is $P(|Z| \gt 1.96) = 2 \times 0.025 = 0.05$. That is $\alpha$, by construction.
  3. Run 1 000 independent A/A tests. The number of false alarms is Binomial$(1000,\ 0.05)$: mean $1000 \times 0.05 = 50$.
  4. Its standard deviation is $\sqrt{1000 \times 0.05 \times 0.95} = \sqrt{47.5} \approx 6.9$, so counts from about $50 - 2(6.9) \approx 36$ to $50 + 2(6.9) \approx 64$ are normal luck (the exact middle 95% of the Binomial is 37 to 64).
  5. With conversion data (10% baseline, 500 users per group) the counts are whole numbers, so the z-test is only approximately calibrated. Adding up every possible outcome exactly gives a false-alarm rate of $0.0496$: very close to 5%.

If you ran 1 000 A/A tests and saw 150 "winners", you would not have learned anything about your product. You would have learned that your pipeline is broken.

  • The significance level $\alpha$ is the false-positive rate you choose in advance: $\alpha = P(\text{reject } H_0 \mid H_0 \text{ true})$.
  • The size of a test is its actual false-positive rate. A test is calibrated (or "valid") when size $\le \alpha$, ideally $\approx \alpha$. Approximations (the CLT, a wrong standard error, dependence between rows) can make the size differ from $\alpha$.
  • An A/A test is an experiment in which both groups get the identical treatment. Every rejection in an A/A test is a Type I error by definition.
  • The p-value is uniform under $H_0$. For a continuous test statistic, if $H_0$ is true then $P(p \le u) = u$ for every $u$ between 0 and 1. Reason: $p \le u$ happens exactly when $|Z| \ge z_{1-u/2}$, and under $H_0$ that has probability $u$. So rejecting when $p \le \alpha$ gives false alarms with probability $\alpha$. For counts (discrete data) this holds only approximately.
Why do we need it?

The promise "only 5% false alarms" holds only if the test's assumptions hold in your real pipeline. A wrong standard error, duplicated rows or a broken random split silently turn 5% into 20% or more. An A/A test is the cheapest way to catch this.

Where is it used?

Experimentation platforms run A/A tests (real or simulated) before trusting a new metric, a new randomization unit or a new analysis method. Statisticians check any new test by simulating data under $H_0$ and plotting the p-values: they should look flat.

How is it used?

Simulate (or run) many A/A tests with the exact same pipeline as a real test, compute the p-values, and check two things: the share below $\alpha$ is close to $\alpha$, and the histogram of p-values is flat. A pile-up near 0 means the standard error is too small somewhere.

Users arrive random split Group 1 old checkout Group 2 the SAME old checkout z-test on the gap "significant" ≈ α of the time
An A/A test changes nothing, so every significant result is a false alarm. Over many A/A tests, the share of false alarms should match $\alpha$.

Each run simulates 1 000 A/A tests (both groups have the same true conversion rate) and plots the histogram of their p-values. Notice it is flat: under $H_0$ every p-value is equally likely. The red bar is the region $p \lt \alpha$; it holds about $\alpha$ of the tests. Change $n$, the baseline rate and $\alpha$: the false-alarm rate stays near $\alpha$. Then tick broken pipeline (every user's row is logged three times, so the formula thinks there are three times as many users): the p-values pile up near 0 and false alarms jump to about 26%.

"One A/A test came out significant (p = 0.03), so our system is broken."

One significant A/A test in twenty is exactly what $\alpha = 0.05$ promises. Judge the rate over many A/A tests (or many simulated ones), and the shape of the p-value histogram.

"A smaller α is always safer."

A smaller $\alpha$ means fewer false alarms and lower power (more misses) at the same sample size. $\alpha$ is a business choice about which mistake costs more.

"If H₀ is true, the p-value should be large, close to 1."

Under $H_0$ the p-value is spread evenly between 0 and 1. A p-value of 0.9 is no more "proof of no effect" than a p-value of 0.4.

Your A/B framework is Bayesian, but you can (and reviewers often will) ask for its frequentist false-positive rate. Recipe: simulate many datasets with $\theta_A = \theta_B$ (an A/A world), run the full model on each one, and count how often your decision rule (for example "$P(\theta_B \gt \theta_A \mid D)$ above your threshold") fires. That share is the Type I error rate of your Bayesian rule. A pile-up of false alarms in this simulation also catches bugs, such as a preprocessing step (a scaler, a mask) that treats the two arms differently, if your code has one.

$\alpha = P(\text{reject} \mid H_0)$, chosen in advance. A/A test: identical groups, every rejection is a false alarm.

Under $H_0$ the p-value is Uniform(0, 1), so $P(p \le \alpha) = \alpha$. 1 000 A/A tests → about 50 "significant" (37–64 is normal luck).

Trap: a pile-up of small p-values in A/A tests means a broken pipeline (e.g. a standard error that is too small).

Quick check: you simulate 2 000 A/A tests at α = 0.05 and get 230 significant results. Is that plausible as luck?

Expected $2000 \times 0.05 = 100$, with standard deviation $\sqrt{2000 \times 0.05 \times 0.95} \approx 9.7$. 230 is about 13 standard deviations too many: not luck. The false-alarm rate is $230/2000 = 11.5\%$, so the standard error (or something else in the pipeline) is wrong.

Power: the chance of catching a real effect core

A metal detector beeps almost every time for a big gold coin just under the sand. For a tiny coin buried deep, it usually stays quiet, even though the coin is really there. How often it beeps for a coin of a given size is its sensitivity. For a test, this is called power.

Power is always "power to detect a certain effect": big effects are easy to catch, small ones hard. And the sample size decides how sharp the detector is. In Chapter 5.6 the checkout test (500 users per group) gave $p \approx 0.31$, "not significant". The natural next question is: was that test even able to see a 2-point lift?

Three ways to say it:

  • Picture: searching a dark room with a weak torch: you will miss small things that are really there.
  • Numbers: with 500 users per group, a real lift from 10% to 12% is detected only about 17% of the time.
  • Slogan: power = P(the test says "yes" | the effect is real and this big).

Power of the checkout test for a true lift from 10% to 12%, 500 users per group, two-sided $\alpha = 0.05$.

  1. When to reject. Under $H_0$ the pooled rate is about 11%, so the standard error of the gap is $SE_0 = \sqrt{0.11 \times 0.89 \times \tfrac{2}{500}} = 0.01979$, about 1.98 points.
  2. We reject when $|\text{gap}| \gt 1.96 \times SE_0 = 1.96 \times 0.01979 = 0.0388$. So the test only says "yes" if the observed gap is bigger than 3.88 points (in either direction).
  3. What the gap does if the truth is 10% vs 12%. It is approximately Normal with mean $0.02$ and standard error $SE_1 = \sqrt{\tfrac{0.10 \times 0.90}{500} + \tfrac{0.12 \times 0.88}{500}} = 0.01978$.
  4. Chance the gap clears $+3.88$ points: $P\!\left(Z \gt \tfrac{0.0388 - 0.02}{0.01978}\right) = P(Z \gt 0.950) = 0.171$.
  5. Chance of a significant gap in the wrong direction: $P\!\left(Z \lt \tfrac{-0.0388 - 0.02}{0.01978}\right) = P(Z \lt -2.97) = 0.0015$.
  6. Power $\approx 0.171 + 0.0015 = 0.173$. So $\beta \approx 0.83$.

So the answer is no. Even if the new checkout really added 2 points, this test would come out "not significant" about 83% of the time. The $p = 0.31$ of Chapter 5.6 says very little about whether the new checkout works.

For a true effect $\Delta$ (for example $\Delta = \theta_B - \theta_A$), the power of a test is

$$\text{Power}(\Delta) = P(\text{reject } H_0 \mid \text{true effect} = \Delta) = 1 - \beta(\Delta).$$
  • Power depends on four things: the true effect $\Delta$, the sample size $n$, the noise (the standard deviation $\sigma$, or $p(1-p)$ for conversions) and $\alpha$.
  • At $\Delta = 0$ the "power" is just $\alpha$: rejecting when nothing changed is a false alarm.
  • For a two-sided z-test whose estimate has standard error $SE$ (taking $SE$ the same under $H_0$ and $H_1$, as for means): $\text{Power} = \Phi\!\left(\tfrac{\Delta}{SE} - z_{1-\alpha/2}\right) + \Phi\!\left(-\tfrac{\Delta}{SE} - z_{1-\alpha/2}\right)$. For proportions the two standard errors differ slightly (pooled under $H_0$, unpooled under $H_1$), as in the example; the derivation comes in the "four levers" section below.
  • Convention (not a law): plan for 80% (sometimes 90%) power at the smallest effect you care about.
Why do we need it?

A test with low power wastes the experiment: it will probably say "no difference" whatever the truth is. Knowing the power tells you whether a "not significant" result means "probably no big effect" or just "we could not tell".

Where is it used?

Planning A/B tests and clinical trials (power analysis is required in most trial protocols), grant proposals, deciding how long to run an experiment, and judging published results (low-power studies produce unreliable findings).

How is it used?

Before the test: pick the smallest effect worth detecting, compute the power at your planned $n$ (for example with statsmodels.stats.power), and enlarge $n$ until power reaches about 80%. After a "not significant" test: say which effects the test could and could not have detected.

The blue curve shows where the observed gap (B − A, in percentage points) lands if nothing changed; the orange curve shows where it lands if the true lift is $\Delta$. Purple lines are the critical values. Red shading = Type I region ($\alpha$), green shading = power, light orange shading = $\beta$ (misses). Start at the checkout test: most of the orange curve sits inside the "not significant" zone. Now slide users per group up to 3 841: both curves get thinner, they separate, and power reaches 80%. Set the lift to 0: power falls to $\alpha$.

"Our test has 80% power." (full stop)

Power is always for a specific effect size (and $n$, $\alpha$, noise): "80% power to detect a lift from 10% to 12% with 3 841 users per group at $\alpha = 0.05$". Against a smaller true effect, the same test has less power.

"Power is the probability that the alternative hypothesis is true."

Power is a probability about the data, assuming the effect is real: P(reject | effect of size Δ). It says nothing about how likely the effect is to exist.

"The checkout test was not significant, so the new checkout does not help."

With power ≈ 0.17 for a 2-point lift, a non-significant result is what you expect even if the lift is real. The test was too small to tell.

"Power is 1 minus the p-value."

Power is planned before the data: P(reject $H_0$ | a real effect of a stated size). The p-value is computed from the observed data, assuming $H_0$.

Model answer: "Power is the probability that the test rejects the null when a real effect of a given size exists. It is 1 − β, where β is the Type II error rate. It grows with the effect size, the sample size and α, and falls with the noise. We usually design for 80% power at the minimum effect we care about."

$\text{Power}(\Delta) = P(\text{reject} \mid \text{true effect } \Delta) = 1 - \beta$. For a z-test: $\approx \Phi(\Delta/SE - z_{1-\alpha/2})$.

Checkout test (500/group, 10% → 12%): reject only if |gap| > 3.88 pts; power ≈ 0.17.

Trap: always say "power to detect what"; low power makes "not significant" almost meaningless.

Quick check: what is the power of a test when the true effect is exactly zero?

If the effect is zero, $H_0$ is true, so every rejection is a false alarm. The rejection probability is $\alpha$. The power curve starts at $\alpha$ when $\Delta = 0$ and rises toward 1 as $|\Delta|$ grows.

Effect size: how big is the change? Absolute lift, relative lift, Cohen's d

A p-value answers "is there evidence of some difference?". It does not say how big the difference is. The effect size does. The same change can be described in different units, and the choice of unit changes how impressive it sounds:

  • Absolute lift: the plain difference. 10% → 12% is "+2 percentage points".
  • Relative lift: the difference as a share of the starting value. 10% → 12% is "+20%".
  • Standardized effect (Cohen's $d$): the difference measured in units of the natural spread of the data. Two hills of data whose tops are half a hill-width apart have $d = 0.5$.

Three ways to say it:

  • Picture: two hills of data; the effect size is how far apart their tops are, measured in hill widths.
  • Numbers: 10% → 12% and 1% → 1.2% are both "+20% relative", but +2 points vs +0.2 points absolute, and they need about 3 800 vs 43 000 users per group.
  • Slogan: always say "points" or "percent"; never just "a 2% lift".
  1. Absolute lift of the checkout: $0.12 - 0.10 = 0.02$ = 2 percentage points.
  2. Relative lift: $\dfrac{0.12 - 0.10}{0.10} = \dfrac{0.02}{0.10} = 0.20$ = 20%.
  3. The same 20% relative lift at a 1% baseline: $0.012 - 0.010 = 0.002$ = only 0.2 points absolute.
  4. Users needed per group for 80% power at $\alpha = 0.05$ (formula later in this chapter): 3 841 for 10% → 12%, but 42 693 for 1% → 1.2%. Rare events need far more data for the same relative lift.
  5. Cohen's $d$ for an average-order-value test: mean order A = 50 dollars, B = 52 dollars, pooled standard deviation 20 dollars. $d = \dfrac{52 - 50}{20} = 0.1$.
  6. What $d = 0.1$ looks like: pick one random A order and one random B order; the B order is bigger with probability $\Phi(d/\sqrt 2) = \Phi(0.071) \approx 0.53$, barely better than a coin flip. Yet over thousands of users a 2-dollar average gain can be worth a lot.

For two groups with true values $\theta_A$ (control) and $\theta_B$ (treatment):

  • Absolute effect (absolute lift): $\Delta = \theta_B - \theta_A$, in the metric's units (percentage points for rates, dollars for revenue).
  • Relative effect (relative lift): $\dfrac{\theta_B - \theta_A}{\theta_A}$, a percentage of the baseline. Only meaningful when $\theta_A \gt 0$.
  • Cohen's $d$ (for means): $d = \dfrac{\mu_B - \mu_A}{\sigma}$, estimated by $\dfrac{\bar x_B - \bar x_A}{s_p}$ with the pooled standard deviation $s_p = \sqrt{\dfrac{(n_A-1)s_A^2 + (n_B-1)s_B^2}{n_A + n_B - 2}}$. It has no units.
  • For proportions, the standardized version used by many power tools (for example statsmodels.stats.proportion.proportion_effectsize) is Cohen's $h$ $= 2\arcsin\sqrt{p_B} - 2\arcsin\sqrt{p_A}$; for 10% → 12%, $h \approx 0.064$.
  • Cohen's labels $d = 0.2$ "small", $0.5$ "medium", $0.8$ "large" are rules of thumb from behavioural science. In large online experiments real effects are usually far below "small"; judge size by its value to the business, not by the labels.
  • Reading $d$ with two Normal curves of equal spread: the curves overlap by $2\Phi(-|d|/2)$, and a random B value beats a random A value with probability $\Phi(d/\sqrt2)$.
Why do we need it?

With enough data, a tiny, worthless difference becomes "significant". The effect size says whether the change matters. It is also the main input to any sample-size calculation: the smaller the effect you want to catch, the more data you need.

Where is it used?

Reporting A/B results ("+2 points, +20% relative"), power analysis and sample-size planning, meta-analysis (combining studies through $d$), and the minimum detectable effect of an experiment design (Chapter 5.10).

How is it used?

Report the absolute lift with a confidence interval (Chapter 5.8) and the relative lift next to the baseline. For planning, convert "the smallest change worth shipping" into $\Delta$ (or $d = \Delta/\sigma$) and feed it to the sample-size formula.

Keep the relative lift at +20% and slide the baseline rate from 10% down to 1%: the absolute lift shrinks from 2 points to 0.2 points and the users needed (purple dot on the curve, log scale) jump about ten-fold. Then move the baseline up to 30%: the same +20% becomes a large absolute change that is cheap to detect.

Blue is the spread of values in group A, orange in group B (same standard deviation). The purple area is where they overlap. Drag d from 0.1 to 0.8: at $d = 0.1$ the hills almost coincide and a random B user beats a random A user only 53% of the time; at $d = 0.8$ ("large" by Cohen's rule of thumb) the overlap is still about 69%. Watch the users needed fall like $1/d^2$.

"The new page lifted conversion by 2%."

Ambiguous: from 10% to 12% (2 points, 20% relative) or from 10% to 10.2% (0.2 points, 2% relative)? Say "+2 percentage points (10% → 12%), a 20% relative lift".

"Highly significant (p < 0.001), so the effect is big."

Significance is about evidence, not size. With millions of users, a 0.01-point lift can have a tiny p-value. Always report the effect size and its interval.

"Cohen's d = 0.1 is 'negligible', so the change is worthless."

The labels are rough conventions. A $d$ of 0.1 on order value across millions of orders can be very valuable. Judge size in business units.

In your A/B framework the posterior gives you every effect-size version for free: from each pair of posterior draws $(\theta_A, \theta_B)$ compute $\theta_B - \theta_A$ (absolute) and $(\theta_B - \theta_A)/\theta_A$ (relative), and summarize each set of draws. When you state a decision like $P(\theta_B - \theta_A \gt \delta \mid D)$ (Chapter 6.4), say clearly whether $\delta$ is in points or in percent: "+0.5" means very different things in the two units.

"We got a 20% lift."

"Conversion went from 10% to 12%: +2 percentage points absolute, +20% relative, 95% CI for the absolute lift from … to …"

Model answer: "The p-value tells me whether the data are surprising under 'no effect'; the effect size tells me how big the effect is. I report the absolute lift with its confidence interval, give the baseline so the relative lift can be read correctly, and judge it against the smallest change that is worth shipping."

Absolute lift $\theta_B - \theta_A$ (points); relative lift $(\theta_B - \theta_A)/\theta_A$ (%); Cohen's $d = (\mu_B - \mu_A)/\sigma$ (no units).

Same relative lift at a lower baseline = smaller absolute lift = far more users needed.

Trap: "2%" without "points" or "relative" is ambiguous; significant ≠ big.

Quick check: conversion moves from 4% to 5%. Give the absolute and the relative lift.

Absolute: $0.05 - 0.04 = 0.01$ = 1 percentage point. Relative: $0.01/0.04 = 0.25$ = +25%.

The four levers of power, and where the formula comes from core

Imagine trying to hear a friend whisper across a noisy room. You hear them more easily if (1) they speak louder (a bigger effect), (2) the room is quieter (less noise in the metric), (3) you listen for longer (more users), or (4) you are willing to say "I heard something" on weaker evidence (a bigger α, at the price of more false alarms).

All four levers do one thing: they change how far apart the "no effect" curve and the "real effect" curve are, measured in standard errors, or where the bar sits. Power depends only on that distance and on the bar.

Three ways to say it:

  • Picture: two bells; push them apart (bigger effect) or make them slimmer (more data, less noise).
  • Numbers: when the true effect is 2.8 standard errors away from zero, a 5% two-sided test has 80% power.
  • Slogan: power is signal divided by noise, in standard-error units.

Average order value. True gain $\Delta = 2$ dollars per order, standard deviation $\sigma = 20$ dollars, $n = 1000$ orders per group, $\alpha = 0.05$ two-sided ($z_{0.975} = 1.96$).

  1. Standard error of the difference of two means: $SE = \sigma\sqrt{2/n} = 20 \times \sqrt{2/1000} = 20 \times 0.0447 = 0.894$ dollars.
  2. Signal in standard-error units: $\delta = \Delta / SE = 2 / 0.894 = 2.24$.
  3. Power $\approx \Phi(\delta - 1.96) = \Phi(2.24 - 1.96) = \Phi(0.28) = 0.61$. (The other tail, $\Phi(-2.24 - 1.96) = \Phi(-4.2)$, is about 0.00001.)
  4. Double $n$ to 2000: $SE = 0.632$, $\delta = 3.16$, power $= \Phi(1.20) = 0.885$.
  5. Double the noise to $\sigma = 40$ (back at $n = 1000$): $\delta = 1.12$, power $= \Phi(-0.84) + 0.001 \approx 0.20$.
  6. Loosen α to 0.10 ($z_{0.95} = 1.645$): power $= \Phi(2.24 - 1.645) = \Phi(0.59) = 0.72$.
  7. Double the effect to $\Delta = 4$: $\delta = 4.47$, power $= \Phi(2.51) = 0.994$.

Deriving the power of a two-sample z-test (equal group sizes $n$, known common standard deviation $\sigma$):

  1. The estimate $\hat\Delta = \bar x_B - \bar x_A$ has standard error $SE = \sqrt{\sigma^2/n + \sigma^2/n} = \sigma\sqrt{2/n}$ (variances of independent groups add, Chapter 5.5).
  2. The test statistic is $Z = \hat\Delta / SE$. Under $H_0$, $Z \sim N(0, 1)$, and we reject when $|Z| \gt z_{1-\alpha/2}$.
  3. Under $H_1$ with true effect $\Delta$, $\hat\Delta \sim N(\Delta, SE^2)$, so $Z \sim N(\delta, 1)$ with $\delta = \Delta / SE = \dfrac{\Delta\sqrt n}{\sigma\sqrt2}$.
  4. So $P(Z \gt z_{1-\alpha/2}) = P(Z - \delta \gt z_{1-\alpha/2} - \delta) = 1 - \Phi(z_{1-\alpha/2} - \delta) = \Phi(\delta - z_{1-\alpha/2})$, and in the same way $P(Z \lt -z_{1-\alpha/2}) = \Phi(-\delta - z_{1-\alpha/2})$.
$$\text{Power} = \Phi\!\left(\delta - z_{1-\alpha/2}\right) + \Phi\!\left(-\delta - z_{1-\alpha/2}\right), \qquad \delta = \frac{\Delta}{\sigma\sqrt{2/n}}.$$
  • The four levers: $\Delta \uparrow$, $n \uparrow$ and $\sigma \downarrow$ all increase $\delta$; $\alpha \uparrow$ lowers $z_{1-\alpha/2}$. Power grows with $\sqrt n$, not with $n$.
  • With unequal groups, $SE = \sigma\sqrt{1/n_A + 1/n_B}$. For a fixed total this is smallest at a 50/50 split, so equal groups give the most power (when the two variances are equal).
  • When $\sigma$ is estimated from the data, the exact calculation uses the t distribution (TTestIndPower in statsmodels); with hundreds of users per group the answer is almost the same.
Why do we need it?

When an experiment is too weak, you need to know which lever to pull: run longer, pick a less noisy metric, reduce variance (for example CUPED, Chapter 5.12), accept a larger α, or aim for a bigger change. The formula shows how much each lever buys.

Where is it used?

Power calculators (statsmodels NormalIndPower and TTestIndPower, G*Power, R's power.t.test), experiment-planning dashboards, variance-reduction work (CUPED, stratification), and choosing metrics for A/B tests.

How is it used?

Compute $\delta = \Delta/SE$ for your plan. If $\delta$ is well below 2.8 (the 80% power point at $\alpha = 0.05$), you need more data or less noise. Since power depends on $\sqrt n$, doubling $\delta$ needs four times the users; halving $\sigma$ does the same job.

Start H₁ is 2.8 SE away power ≈ 80% Bigger effect Δ curves further apart power ↑ More n or less σ SE = σ√(2/n) shrinks power ↑ Bigger α the bar moves left power ↑ but α ↑ too
Blue = where the statistic lands under $H_0$, orange = under $H_1$, purple line = the bar, red = $\alpha$, green = power. Each lever increases the green area; only the last one also increases the red area.

The purple curve is power as a function of the signal in standard-error units, $\delta = \Delta/SE$. The dot is your experiment. Move Δ, σ and users per group: all three only slide the dot along the same curve. The dashed lines mark 80% power, reached at $\delta \approx 2.8$ when $\alpha = 0.05$. Changing α is different: it reshapes the curve itself (look at where it starts, at $\delta = 0$).

Power vs users: each curve is one true lift on a 10% baseline (users on a log scale). Drag the purple handle along the bottom: at 500 per group (the checkout test) a 2-point lift has power 0.17 and a 3-point lift about 0.32; to reach 80% for 2 points you need about 3 800 per group. Switch to Power vs true lift: each curve is one sample size. All curves pass through $\alpha$ at zero lift, and bigger samples make the valley narrower.

"Twice the users, twice the power."

Power depends on $\delta = \Delta\sqrt n/(\sigma\sqrt2)$: doubling $n$ multiplies $\delta$ by $\sqrt2 \approx 1.41$. To double $\delta$ you need four times the users.

"A 70/30 split puts the new variant in front of more users, so it is just as good."

For a fixed total, unequal groups have a larger $SE = \sigma\sqrt{1/n_A + 1/n_B}$, so less power. Uneven splits can be fine for safety reasons, but they cost power.

"Raising α to 0.20 is a free way to get more power."

It is not free: it quadruples the false-alarm rate compared with 0.05. It is a business trade-off, and it must be decided before the data arrive.

$\text{Power} = \Phi(\delta - z_{1-\alpha/2}) + \Phi(-\delta - z_{1-\alpha/2})$, with $\delta = \Delta/SE = \Delta\sqrt n/(\sigma\sqrt2)$.

Levers: bigger $\Delta$, bigger $n$ (as $\sqrt n$), smaller $\sigma$, bigger $\alpha$. 80% power at $\alpha = 0.05$ needs $\delta \approx 1.96 + 0.84 = 2.8$.

Trap: power grows like $\sqrt n$; α is not a free lever.

Quick check: a test has δ = 1.4. Roughly how many times more users does it need to reach δ = 2.8?

$\delta$ grows like $\sqrt n$. To double $\delta$ (1.4 → 2.8) you need $2^2 = 4$ times as many users.

Sample size: how many users do I need? core

So far we asked "given my sample size, what is my power?". Planning turns the question around: "I want 80% power to catch an effect of this size; how many users do I need?". You decide three things before collecting any data: the smallest effect worth detecting, the false-alarm rate $\alpha$, and the power. The formula then gives $n$.

The key fact is the square in the formula: $n$ grows with $1/\Delta^2$. Looking for an effect half as big needs four times the users, because the standard error only shrinks like $1/\sqrt n$.

Three ways to say it:

  • Picture: choosing the zoom of a microscope before you look: smaller things need a stronger lens.
  • Numbers: 80% power for 10% → 12% at $\alpha = 0.05$ needs about 3 841 users per group; the checkout test used 500.
  • Slogan: halve the effect, quadruple the users.

Means (order value): $\sigma = 20$ dollars, smallest effect worth finding $\Delta = 2$ dollars, $\alpha = 0.05$ two-sided, power 80%.

  1. Look up the two z values: $z_{1-\alpha/2} = z_{0.975} = 1.96$ and $z_{1-\beta} = z_{0.80} = 0.8416$. Their sum is $2.8016$.
  2. $n = \dfrac{2\sigma^2 (z_{1-\alpha/2} + z_{1-\beta})^2}{\Delta^2} = \dfrac{2 \times 400 \times 2.8016^2}{2^2} = \dfrac{800 \times 7.849}{4} = 1569.8$.
  3. Always round up: 1 570 users per group, 3 140 in total.

Proportions (the checkout): $p_A = 0.10$, $p_B = 0.12$, so $\Delta = 0.02$ and the average rate is $\bar p = 0.11$.

  1. Spread under $H_0$: $\sqrt{2\bar p(1-\bar p)} = \sqrt{2 \times 0.11 \times 0.89} = \sqrt{0.1958} = 0.44249$.
  2. Spread under $H_1$: $\sqrt{p_A(1-p_A) + p_B(1-p_B)} = \sqrt{0.09 + 0.1056} = \sqrt{0.1956} = 0.44227$.
  3. Combine: $1.96 \times 0.44249 + 0.8416 \times 0.44227 = 0.86727 + 0.37222 = 1.23949$; squared: $1.53634$.
  4. $n = 1.53634 / 0.02^2 = 1.53634 / 0.0004 = 3840.8$ → 3 841 per group (7 682 in total).
  5. Halve the effect to 1 point (10% → 11%): $n = 14\,751$ per group, about $3.8\times$ more (not exactly 4× because $p(1-p)$ changes a little).

The checkout test used 500 per group: less than one seventh of what 80% power for a 2-point lift requires.

Deriving $n$. Ignore the tiny wrong-direction tail of the power formula. Power $1 - \beta$ means $\Phi(\delta - z_{1-\alpha/2}) = 1 - \beta$, that is $\delta - z_{1-\alpha/2} = z_{1-\beta}$. So the effect must sit $z_{1-\alpha/2} + z_{1-\beta}$ standard errors away from zero (see the figure):

$$\Delta = (z_{1-\alpha/2} + z_{1-\beta})\,\sigma\sqrt{2/n} \quad\Longrightarrow\quad n = \frac{2\sigma^2\,(z_{1-\alpha/2} + z_{1-\beta})^2}{\Delta^2}\ \text{ per group} \;=\; \frac{2\,(z_{1-\alpha/2} + z_{1-\beta})^2}{d^2}, \quad d = \Delta/\sigma.$$
  • For two proportions (pooled standard error under $H_0$, as in LA.stats.sampleSize.twoProp): $n = \dfrac{\left[z_{1-\alpha/2}\sqrt{2\bar p(1-\bar p)} + z_{1-\beta}\sqrt{p_A(1-p_A) + p_B(1-p_B)}\right]^2}{(p_B - p_A)^2}$, with $\bar p = (p_A + p_B)/2$.
  • Handy multipliers $2(z_{1-\alpha/2} + z_{1-\beta})^2$: $\alpha = 0.05$ and 80% power → 15.7; 90% power → 21.0; $\alpha = 0.01$ and 80% power → 23.4. Hence the rule of thumb (Lehr's rule) $n \approx 16\sigma^2/\Delta^2$ per group.
  • Different tools use slightly different versions (pooled or unpooled spread, Cohen's $h$, t instead of z). For the checkout they give 3 835 to 3 841: the differences do not matter, the order of magnitude does.
  • The minimum detectable effect (MDE) is the same formula solved for $\Delta$ at a given $n$: $\Delta_{\min} = (z_{1-\alpha/2} + z_{1-\beta})\,SE$. Turning $n$ into test duration from daily traffic is covered in Chapter 5.10.
Why do we need it?

Too few users and the test is a coin flip that wastes weeks; too many and you hold back a good change (or expose users to a bad one) for no reason. The formula gives the size that answers your question, and nothing more.

Where is it used?

Experiment-planning tools at every company that runs A/B tests, clinical trial protocols, survey design, and statsmodels (NormalIndPower().solve_power, TTestIndPower().solve_power, proportion_effectsize).

How is it used?

Fix $\alpha$, the power and the smallest effect worth shipping (from business value, not from hope). Get the baseline rate or $\sigma$ from historical data. Compute $n$ per group, round up, divide by daily traffic to get the duration, and commit to that $n$ before starting.

α/2 critical value H₀: centre 0 H₁: centre Δ power 1 − β β α/2 z₁₋α/₂ · SE z₁₋β · SE Δ = (z₁₋α/₂ + z₁₋β) · SE
For power $1-\beta$ the bar must sit $z_{1-\alpha/2}$ standard errors above 0 (so only $\alpha$ of the blue curve passes it) and $z_{1-\beta}$ standard errors below $\Delta$ (so $1-\beta$ of the orange curve passes it). Adding the two distances gives $\Delta = (z_{1-\alpha/2} + z_{1-\beta})\,SE$; solving for $n$ inside $SE$ gives the sample-size formula.

Set the baseline conversion rate and the MDE (smallest lift worth detecting, in percentage points). The purple curve shows users per group against the MDE (log scale). Start at 10% and 2 points (3 841 per group), then halve the MDE to 1 point: the dot climbs about four times higher. Try 90% power or $\alpha = 0.01$, and move the daily traffic to see how many days the test would take.

"The MDE is the effect we expect to see."

The MDE is the smallest effect worth detecting. If the true effect is smaller than the MDE, the test will probably miss it; if it is bigger, power is higher than planned.

"The calculator said 3 841, so we need 3 841 users."

Check whether the number is per group or total (here it is per group: 7 682 in total), and whether the inputs were absolute points or relative percent.

"Let's just run until the result is significant, then stop."

Stopping at the first significant look inflates the false-alarm rate far above $\alpha$ (peeking, Chapter 5.11). Fix $n$ in advance, or use a method designed for sequential looks.

For a Bayesian decision rule there is no textbook sample-size formula, but the planning question is the same. Answer it by simulation (sometimes called "Bayesian power" or "assurance"): pick plausible true rates (say 10% and 12%), generate many datasets at a candidate $n$, run your Beta-Binomial update or full model on each, and record how often your rule (for example "$P(\theta_B \gt \theta_A \mid D)$ above your threshold") fires. That fraction is your rule's power at that $n$; the same simulation with $\theta_A = \theta_B$ gives its false-positive rate. Increase $n$ until both look acceptable.

$n = \dfrac{2\sigma^2(z_{1-\alpha/2} + z_{1-\beta})^2}{\Delta^2}$ per group; at $\alpha = 0.05$, 80% power: $n \approx 16\sigma^2/\Delta^2$. For rates use $\sigma^2 \approx p(1-p)$.

10% → 12%: 3 841 per group. Halve the MDE → about 4× the users.

Trap: per group vs total; points vs percent; never "run until significant".

Quick check: with σ = 10 and Δ = 1, use the rule of thumb to estimate n per group (α = 0.05, 80% power).

$n \approx 16\sigma^2/\Delta^2 = 16 \times 100 / 1 = 1600$ per group. (The exact formula gives $15.7 \times 100 = 1570$.)

Low power does double damage: missed effects and exaggerated winners core

An underpowered test does not only miss real effects. When it does call a winner, the winner usually looks bigger than it really is. Why? With a small sample, the observed gap only clears the significance bar when luck pushed it upward. So the significant results are a selected, lucky-high group. This is called the winner's curse.

It is like fishing with a net that has very big holes: you only ever catch the biggest fish, and you go home thinking the lake is full of giants.

Three ways to say it:

  • Picture: a net with big holes keeps only the biggest fish.
  • Numbers: with 500 users per group and a true lift of 2 points, the significant results show about 4.9 points on average: 2.5 times the truth.
  • Slogan: low power → "winners" look bigger than they are.

The checkout test design (500 per group, baseline 10%), with a true lift of 2 points, repeated many times by simulation:

  1. To be significant, the observed gap must exceed 3.88 points (from the power example above).
  2. The truth is 2 points, so every significant positive result overstates it by at least $3.88/2 \approx 1.9$ times.
  3. In 200 000 simulated tests, 17.1% were significant (the formula said 17.3%), and the significant positive gaps averaged 4.9 points: an exaggeration of about 2.5×.
  4. About 0.9% of the significant results even had the wrong sign (B looked significantly worse).
  5. With 3 841 per group (80% power), the significant results averaged about 2.25 points: only about 1.12× the truth.

So a small test that "finds" a +5-point lift has probably found a much smaller real lift, plus luck.

  • Exaggeration ratio (Type M, "magnitude", error): $\dfrac{E\big[\,|\hat\Delta|\ \big|\ \text{significant}\big]}{|\Delta|}$, how much significant estimates overstate the true effect on average. It is close to 1 at high power and grows fast as power falls.
  • Type S ("sign") error: $P(\text{the significant estimate has the wrong sign})$. It is only noticeable at very low power.
  • (These names come from Gelman and Carlin, 2014. The general effect, selecting results because they cleared a bar, is the winner's curse.)
  • Observed (post-hoc) power is power computed by plugging the observed effect in as if it were the true one. For a z-test it equals $\Phi(|z| - z_{1-\alpha/2}) + \Phi(-|z| - z_{1-\alpha/2})$: a function of the p-value alone (about 0.5 when $p = 0.05$, about 0.17 for the checkout's $z = 1.01$). It adds no information beyond $p$.
Why do we need it?

Teams act on the size of a measured lift: they forecast revenue from it and prioritize similar ideas. If small tests systematically inflate the winners, those plans are built on luck. Knowing this, you plan bigger tests and discount surprising wins.

Where is it used?

Interpreting A/B test winners, the replication crisis in psychology and medicine (small studies with huge effects that shrink on replication), "best segment" or "best variant" picks, and the reason Bayesian shrinkage and hierarchical models are popular in experimentation.

How is it used?

Plan for adequate power at a realistic effect size. When a small test shows a surprisingly large win, re-run it or shrink the estimate toward a sensible prior. Report a confidence interval instead of a single number, and never compute "observed power" after the fact.

Every run simulates 1 000 A/B tests in which B really is better by the true lift (green line). Grey bars are non-significant results; orange bars are significant wins; red bars are significant results with the wrong sign. The purple dashed line is the average of the significant wins. At the checkout size (500 per group) the purple line sits far to the right of the green truth. Press 80% power: most tests become significant and the purple line moves close to the green one.

"Our small test found a significant +5-point lift, so the lift is about 5 points."

At low power, significant estimates are inflated (2.5× on average in the checkout design). Expect the true lift to be smaller; replicate the test or shrink the estimate before forecasting from it.

"The test was not significant and the post-hoc power was only 17%, so the effect is probably real; we just lacked power."

Observed power is a re-expression of the p-value: a non-significant result always has low observed power. Use the power you computed before the test for a meaningful effect, and the confidence interval of the effect.

"Significant results are the reliable ones, so we can ignore the rest."

Keeping only the significant results is exactly what creates the bias. Look at all results (and their intervals), not just the winners.

In your A/B framework, hierarchical partial pooling is a natural guard against the winner's curse. When many segments or variants are estimated at once, the small, noisy ones are the most likely to look spectacular by luck. Pooling pulls each group's estimate toward the overall mean, and pulls the noisiest groups the most (Chapter 6.6). It does not remove the problem of picking the "best-looking" group after the fact, but it makes the picked estimate far less inflated.

"The checkout test was not significant, so the new checkout does not work."

"The test could not have told us: with 500 users per group, its power to detect a 2-point lift was only about 17%."

Model answer: "Was this test even able to detect a 2-point lift? No. At 10% baseline and 500 per group, the observed gap had to exceed about 3.9 points to be significant, so power for a 2-point lift was about 17%: an 83% chance of a miss even if the lift is real. The 95% confidence interval for the lift, roughly −1.9 to +5.9 points, says the data are compatible with anything from a small loss to a big win. To detect 2 points with 80% power we would need about 3 800 users per group. And if a test this small had come out significant, I would expect the estimated lift to be inflated."

Low power: most real effects are missed, and the significant ones are exaggerated (Type M) and occasionally wrong-signed (Type S).

Checkout design, true +2 pts: significant wins average about +4.9 pts (×2.5).

Trap: never compute "observed power" after a test; it is just the p-value in disguise.

Quick check: a test with 90% power finds a significant result. Should you expect a big exaggeration?

No. At high power almost every result is significant, so being significant selects very little; the average significant estimate is close to the truth (the exaggeration ratio is near 1). The winner's curse is a low-power problem.

Recap, cheat sheet and practice

  • A test can be wrong in two ways: a Type I error (false positive, rate $\alpha$, chosen) and a Type II error (miss, rate $\beta$, depends on the true effect, $n$ and noise). Power $= 1 - \beta$.
  • $\alpha$ is $P(\text{reject} \mid H_0)$, not "the chance a winner is fake". Under $H_0$ p-values are uniform, so about 5% of A/A tests are significant at $\alpha = 0.05$; more than that means a broken pipeline.
  • Effect size: absolute lift (points), relative lift (%), Cohen's $d = \Delta/\sigma$. Always state the units and the baseline.
  • Power depends on $\delta = \Delta/SE = \Delta\sqrt n/(\sigma\sqrt2)$: $\text{Power} = \Phi(\delta - z_{1-\alpha/2}) + \Phi(-\delta - z_{1-\alpha/2})$. 80% power at $\alpha = 0.05$ needs $\delta \approx 2.8$.
  • Sample size: $n = 2\sigma^2(z_{1-\alpha/2} + z_{1-\beta})^2/\Delta^2 \approx 16\sigma^2/\Delta^2$ per group. Halve the effect → four times the users. Checkout 10% → 12%: 3 841 per group.
  • The checkout test (500 per group) had only about 17% power for a 2-point lift: its "not significant" says little.
  • Low power misses real effects and inflates the winners it finds (about ×2.5 in the checkout design). Never compute observed power after the fact.

Cheat sheet

IdeaFormulaIn words
Type I error rate$\alpha = P(\text{reject } H_0 \mid H_0)$false alarms when nothing changed
Type II error rate$\beta = P(\text{keep } H_0 \mid H_1, \Delta)$misses of a real effect of size Δ
Power$1-\beta = \Phi(\delta - z_{1-\alpha/2}) + \Phi(-\delta - z_{1-\alpha/2})$chance to catch an effect of size Δ
Signal in SE units$\delta = \Delta/SE$, $SE = \sigma\sqrt{2/n}$how far apart the two curves are
Sample size (means)$n = 2\sigma^2(z_{1-\alpha/2}+z_{1-\beta})^2/\Delta^2$per group; ≈ $16\sigma^2/\Delta^2$
Sample size (rates)$\big[z_{1-\alpha/2}\sqrt{2\bar p(1-\bar p)} + z_{1-\beta}\sqrt{p_A q_A + p_B q_B}\big]^2/\Delta^2$$q = 1-p$; 10% → 12%: 3 841
MDE$\Delta_{\min} \approx (z_{1-\alpha/2}+z_{1-\beta})\,SE$smallest effect the design can catch
Effect sizes$\theta_B - \theta_A$; $(\theta_B-\theta_A)/\theta_A$; $d = \Delta/\sigma$points; percent; standard deviations
z values$z_{0.975} = 1.96$, $z_{0.995} = 2.576$, $z_{0.8} = 0.84$, $z_{0.9} = 1.28$for α = 0.05, 0.01; power 80%, 90%
A/A false alarmsBinomial$(m, \alpha)$: mean $m\alpha$, sd $\sqrt{m\alpha(1-\alpha)}$1 000 tests → about 50 ± 14
Code it · Python

import numpy as np
from scipy import stats
from statsmodels.stats.power import NormalIndPower, TTestIndPower
from statsmodels.stats.proportion import proportion_effectsize

z = stats.norm.ppf          # quantile of the standard Normal
Phi = stats.norm.cdf        # its CDF

# 1) Power of the checkout test: 10% -> 12%, 500 users per group, alpha = 0.05 two-sided
p1, p2, n, alpha = 0.10, 0.12, 500, 0.05
pbar = (p1 + p2) / 2
se0 = np.sqrt(2 * pbar * (1 - pbar) / n)               # SE of the gap under H0 (pooled)
se1 = np.sqrt(p1 * (1 - p1) / n + p2 * (1 - p2) / n)   # SE of the gap under H1
crit = z(1 - alpha / 2) * se0                          # gap needed to reject
power = Phi((p2 - p1 - crit) / se1) + Phi((-(p2 - p1) - crit) / se1)
print(f"critical gap = {crit:.4f}, power = {power:.3f}")
# critical gap = 0.0388, power = 0.173
h = proportion_effectsize(p2, p1)                      # Cohen's h (arcsine scale)
print(round(h, 4), round(NormalIndPower().power(effect_size=h, nobs1=n, alpha=alpha), 3))
# 0.064 0.173

# 2) Sample size per group for 80% power
za, zb = z(1 - alpha / 2), z(0.80)
n_needed = (za * np.sqrt(2 * pbar * (1 - pbar)) + zb * np.sqrt(p1*(1-p1) + p2*(1-p2)))**2 / (p2 - p1)**2
print(int(np.ceil(n_needed)), round(NormalIndPower().solve_power(effect_size=h, alpha=alpha, power=0.8)))
# 3841 3835        (two common versions of the formula; both about 3 800)
print(int(np.ceil(2 * (za + zb)**2 * 20**2 / 2**2)))  # means: sigma = 20, Delta = 2
# 1570
print(round(TTestIndPower().solve_power(effect_size=0.1, alpha=alpha, power=0.8)))   # same with the t-test
# 1571

# 3) 1000 A/A tests: about 5% false alarms, p-values flat
rng = np.random.default_rng(0)
def ztest(a, b, n):
    pp = (a + b) / (2 * n)
    se = np.sqrt(pp * (1 - pp) * 2 / n)
    zz = (b - a) / n / se
    return 2 * stats.norm.sf(np.abs(zz)), zz
a = rng.binomial(500, 0.10, size=1000)
b = rng.binomial(500, 0.10, size=1000)                 # same rate: an A/A test
pv, _ = ztest(a, b, 500)
print("false-alarm rate:", np.mean(pv < 0.05))
print("p-value histogram (10 bins):", np.histogram(pv, bins=10, range=(0, 1))[0])
# false-alarm rate: 0.055          (luck range for 1000 tests: about 0.036 to 0.064)
# p-value histogram (10 bins): [114 109 121  92  83 102 102  84  79 114]   (roughly flat)

# 4) 100 000 A/B tests with a REAL 2-point lift: power and the winner's curse
a = rng.binomial(500, 0.10, size=100_000)
b = rng.binomial(500, 0.12, size=100_000)
pv, zz = ztest(a, b, 500)
sig = pv < 0.05
gap = (b - a) / 500
print("simulated power:", round(sig.mean(), 3))
print("avg significant win:", round(gap[sig & (zz > 0)].mean(), 4), "vs truth 0.02")
print("wrong-sign share of significant:", round(np.mean(zz[sig] < 0), 4))
# simulated power: 0.173
# avg significant win: 0.0494 vs truth 0.02      (exaggeration about x2.5)
# wrong-sign share of significant: 0.0079
Test yourself

1. The new checkout does nothing, but your test says "significant, ship it". This is…

$H_0$ (no effect) was true and you rejected it: a false positive. Its long-run rate, when $H_0$ is true, is $\alpha$.

2. You simulate 1 000 A/A tests with a correctly working pipeline at α = 0.05. About how many are significant?

Under $H_0$ the p-value is uniform, so $P(p \lt 0.05) = 0.05$: about $1000 \times 0.05 = 50$ (37 to 64 is normal luck).

3. Which change does not increase the power of a test (everything else fixed)?

A smaller α raises the bar $z_{1-\alpha/2}$, so fewer real effects clear it: power goes down. The other three all increase $\delta = \Delta/SE$.

4. You planned 4 000 users per group for an MDE of 2 points. Product now wants to detect 1 point. Roughly how many users per group?

$n \propto 1/\Delta^2$: halving Δ multiplies $n$ by 4. (For rates the factor is a bit less than 4 because $p(1-p)$ changes slightly.)

5. Conversion goes from 5% to 6%. Which description is correct?

Absolute: $0.06 - 0.05 = 0.01$ = 1 point. Relative: $0.01/0.05 = 0.20$ = 20%.

6. A small, low-power test reports a significant +6-point lift. What should you expect about the true lift?

With low power, only lucky-high estimates clear the significance bar, so significant estimates overstate the effect on average (Type M error, the winner's curse).

Practice problems

A. Compute the power of the checkout design (500 per group, α = 0.05 two-sided) if the true lift is 10% → 13%.
  1. $\bar p = 0.115$; $SE_0 = \sqrt{2 \times 0.115 \times 0.885/500} = \sqrt{0.000407} = 0.02018$; critical gap $= 1.96 \times 0.02018 = 0.03955$.
  2. $SE_1 = \sqrt{0.10 \times 0.90/500 + 0.13 \times 0.87/500} = \sqrt{0.00018 + 0.000226} = 0.02015$.
  3. Power $\approx P\!\left(Z \gt \frac{0.03955 - 0.03}{0.02015}\right) = P(Z \gt 0.474) = 0.318$ (the wrong-direction tail adds about 0.0003).

Even a 3-point lift (a 30% relative lift) would be detected less than a third of the time.

B. Order value has σ = 30 dollars. How many orders per group do you need to detect Δ = 1.5 dollars with 90% power at α = 0.05?

$z_{0.975} + z_{0.90} = 1.96 + 1.2816 = 3.2416$; squared $= 10.51$; times 2 $= 21.01$.

$n = 21.01 \times 30^2 / 1.5^2 = 21.01 \times 900 / 2.25 = 21.01 \times 400 = 8405.9$ → 8 406 per group.

C. A team tests 200 ideas; 10% really work; each test has α = 0.05 and power 0.5. What share of the declared winners is fake?

Useless ideas: 180, false alarms $180 \times 0.05 = 9$. Real ideas: 20, detections $20 \times 0.5 = 10$. Winners $= 19$, of which $9$ are fake: $9/19 \approx 47\%$. Low power and few true ideas make "significant" results unreliable, even with $\alpha = 0.05$.

D. With 500 users per group, a 10% baseline, α = 0.05 and 80% power, what is the minimum detectable effect?

Quick version: $SE \approx \sqrt{2 \times 0.1 \times 0.9 / 500} = 0.0190$, so $\Delta_{\min} \approx 2.80 \times 0.0190 = 0.053$, about 5.3 points. Solving the exact two-proportion formula for $p_B$ (the variance grows a little as $p_B$ rises) gives about 5.9 points. Either way: this design can only reliably see lifts of more than 5 points, more than double the 2 points the team hoped for.

E. A variance-reduction method (CUPED, Chapter 5.12) removes 49% of the variance of your metric. How does the required sample size change?

$n \propto \sigma^2$, so the new $n$ is $0.51$ of the old one. A plan that needed 3 841 per group now needs about $0.51 \times 3841 \approx 1\,959$ per group: the same power from roughly half the users (or the same users and a smaller MDE).

F. (Interview) "We simulated A/A data through our Bayesian A/B model and the decision rule fired in 12% of runs. What does that mean, and what would you do?"

"In an A/A world any 'ship' decision is a false positive, so the rule's Type I error rate is about 12%. That is not automatically wrong: a Bayesian rule is not built to hit 5%. But it should be a conscious choice. I would check the number against the cost of shipping a useless change; if it is too high, raise the posterior-probability threshold or require $P(\theta_B - \theta_A \gt \delta \mid D)$ with a meaningful $\delta$; I would also check for bugs (a much higher rate than expected often means the model is over-confident, e.g. treating correlated rows as independent). Then I would rerun the simulation with a real lift to see what the change does to power."

Chapter 5.8 · Syllabus Module 15

Confidence intervals

A single number such as "12% conversion" is almost certainly a little wrong. A confidence interval replaces it with a range, "between 9.4% and 15.1%", built by a recipe that catches the truth in 95% of samples. This chapter builds intervals for means, proportions and differences, shows exactly what the "95%" does and does not mean, links intervals to tests, and ends with the distinction your syllabus asks you to explain perfectly: confidence interval vs credible interval.

  • Build an interval as estimate ± multiplier × standard error, and see why the multiplier for 95% is 1.96
  • Say precisely what "95% confidence" means: the procedure covers the truth in 95% of repeated samples (watch 100 intervals, with the misses in red)
  • Build intervals for a mean (z when σ is known, t when it is estimated) and see why small samples need the t multiplier
  • Plan the margin of error: four times the data halves the margin
  • Build intervals for a proportion (Wald and Wilson) and see why Wald fails for small $n$ or rates near 0 or 1
  • Build an interval for a difference (the checkout lift: −1.9 to +5.9 points) and avoid the "overlapping intervals" trap
  • Use the duality: a 95% interval is the set of values a 5% test would not reject
  • Explain confidence vs credible interval perfectly: what is random, what the 95% refers to, and when the numbers agree

What we need from earlier chapters: standard errors, $SE(\bar x) = \sigma/\sqrt n$, $SE(\hat p) = \sqrt{p(1-p)/n}$, and $SE$ of a difference $= \sqrt{SE_1^2 + SE_2^2}$ (Chapter 5.5); the logic of tests and p-values (Chapter 5.6) and power (Chapter 5.7); the Normal and Student-t distributions (Chapter 4.9); the CLT (Chapter 4.13); the Beta distribution (Chapter 4.11). Notation: $z_q$ is the $q$-quantile of the standard Normal ($z_{0.975} = 1.96$); $t_{k,\,q}$ is the $q$-quantile of the Student-t distribution with $k$ degrees of freedom. Running examples: daily orders (sample mean 50 over 36 days, σ = 12) and the checkout test (50/500 vs 60/500).

From one number to a range: estimate ± margin core

Try to spear a fish in murky water: you aim at where you think it is, and you almost always miss by a little. Throw a net instead, centred on your best guess, and you catch the fish most of the time. A point estimate (one number) is the spear. A confidence interval is the net.

How wide should the net be? It depends on how much your estimate wobbles from sample to sample, which is exactly the standard error (Chapter 5.5). A wobbly estimate needs a wide net; a precise one needs a narrow net. And if you want to catch the fish more often (99% instead of 95%), you need a wider net.

Three ways to say it:

  • Picture: throw a net centred on your guess instead of a spear.
  • Numbers: 50 orders per day ± 3.9 → from 46.1 to 53.9.
  • Slogan: interval = estimate ± (multiplier × standard error).

Daily orders. Over $n = 36$ days a shop averaged $\bar x = 50$ orders per day. From a long history we know the day-to-day standard deviation is $\sigma = 12$ orders.

  1. Standard error of the mean: $SE = \sigma/\sqrt n = 12/\sqrt{36} = 12/6 = 2$ orders.
  2. By the CLT, $\bar x$ is approximately Normal around the true mean $\mu$ with standard deviation 2. A Normal value lands within $1.96$ standard deviations of its centre 95% of the time.
  3. So in 95% of samples, $|\bar x - \mu| \le 1.96 \times 2 = 3.92$. Turn the sentence around: in 95% of samples, $\mu$ lies within 3.92 of $\bar x$.
  4. 95% interval: $50 \pm 3.92 = [46.08,\ 53.92]$. The number $3.92$ is the margin of error.
  5. 90% interval (multiplier 1.645): $50 \pm 3.29 = [46.71,\ 53.29]$. 99% interval (multiplier 2.576): $50 \pm 5.15 = [44.85,\ 55.15]$. More confidence, wider net.

A confidence interval at level $1-\alpha$ (for example 95%, $\alpha = 0.05$) for a parameter $\theta$ is a pair of statistics $L$ and $U$, both computed from the data, such that

$$P(L \le \theta \le U) = 1 - \alpha \quad \text{for every possible value of } \theta.$$

Here the probability is over repeated samples: $\theta$ is a fixed number and the endpoints $L, U$ are the random parts. This probability is the interval's coverage. (Approximate methods only reach $1-\alpha$ approximately.)

  • When the estimate $\hat\theta$ is approximately Normal (by the CLT), the standard recipe is $\hat\theta \pm z_{1-\alpha/2}\cdot SE(\hat\theta)$.
  • $z_{1-\alpha/2}$ is the critical value or multiplier: 1.645 (90%), 1.96 (95%), 2.576 (99%).
  • $z_{1-\alpha/2}\cdot SE$ is the margin of error; $L$ and $U$ are the lower and upper limits.
  • Assumptions to check: the observations are independent (or the SE accounts for the dependence), the SE formula is right for how the data were collected, and $n$ is large enough for the Normal approximation.
Why do we need it?

A number without its uncertainty invites over-confidence: "conversion is 12%" sounds exact, but with 500 users it could easily be 10% or 15%. The interval shows, in the metric's own units, how precise the estimate is.

Where is it used?

A/B test reports (the interval of the lift), error bars on dashboards and plots, regression coefficient tables (statsmodels prints [0.025 0.975] columns), polls ("± 3 points"), model evaluation (an interval around an accuracy or RMSE), clinical trials.

How is it used?

Compute the estimate and its standard error, pick the level, and report "estimate (95% CI: lower to upper)". Read the width as precision; read whether it excludes 0 (or a business threshold) as the test result.

4446505456 estimate x̄ = 50 lower 46.08 upper 53.92 margin = 1.96 × SE = 3.92 margin 3.92 orders per day · SE = σ/√n = 12/√36 = 2 · level 95% → multiplier 1.96
A 95% confidence interval for the mean number of daily orders: the estimate in the middle, the margin of error (multiplier × standard error) on each side.

The curve shows where the sample mean $\bar x$ lands over many samples (centred on the true mean 50, which you would not normally know). The green middle region holds the chosen share of samples. Below it, the bar is the interval $\bar x \pm z \cdot SE$ built from one sample. Drag the blue dot (your $\bar x$): the interval contains 50 (blue) exactly while the dot stays inside the green region, and misses (red) as soon as it leaves. Press New sample a few times, then change the level and $n$.

"The interval is centred on the true value."

The interval is centred on the estimate, which moves from sample to sample. The true value stays still; the interval moves around it.

"A 99% interval is better than a 95% interval."

It catches the truth more often but it is wider (less precise). The level is a trade-off between "how often right" and "how precise".

"A narrow interval means the estimate is correct."

Width measures precision, not correctness. If the data are biased (a broken split, a non-random sample, Chapter 5.4), the interval can be narrow and still far from the truth.

CI = estimate ± $z_{1-\alpha/2}$ × SE. Multipliers: 1.645 (90%), 1.96 (95%), 2.576 (99%).

Orders: $50 \pm 1.96 \times 12/\sqrt{36} = 50 \pm 3.92 = [46.08, 53.92]$.

Trap: the interval moves, the truth does not; narrow ≠ unbiased.

Quick check: with σ = 12 and n = 144 days, what is the 95% margin of error?

$SE = 12/\sqrt{144} = 12/12 = 1$, so the margin is $1.96 \times 1 = 1.96$ orders: half the margin we had with 36 days, because $n$ is four times bigger.

What "95% confident" really means core

A ring-toss champion rings the peg 95% of the time. That 95% describes the thrower, over many throws. Once a ring has landed, it is either around the peg or not: there is nothing 95% about that one throw. If the peg is hidden behind a curtain, you cannot tell which throws hit; you only know the thrower's long-run record.

A confidence interval is one throw. The 95% describes the recipe that built it: if you repeated the whole study many times, 95% of the intervals would contain the true value. For the one interval you actually have, the true value is either inside or outside, and you do not know which.

Three ways to say it:

  • Picture: 100 nets thrown at a fixed fish; about 95 catch it.
  • Numbers: 100 samples → 100 different intervals → about 95 contain μ and about 5 miss (usually between 1 and 10 miss).
  • Slogan: the confidence is in the procedure, not in this particular interval.

True mean $\mu = 50$ (known only because we are simulating), $\sigma = 12$, $n = 36$, so every 95% interval is $\bar x \pm 3.92$.

  1. Sample 1 gives $\bar x = 51.3$: interval $[47.38,\ 55.22]$. It contains 50. ✓
  2. Sample 2 gives $\bar x = 45.6$: interval $[41.68,\ 49.52]$. Its upper end 49.52 is below 50: a miss. ✗
  3. Repeat 100 times. Each interval misses with probability 0.05, independently, so the number of misses is Binomial(100, 0.05): on average 5, standard deviation $\sqrt{100 \times 0.05 \times 0.95} \approx 2.2$.
  4. In real life you get only one sample, and you cannot see the green line at $\mu$. Your interval might be one of the misses. The 95% is your protection over the long run, not a statement about this interval.
  • Coverage probability: $P_\theta(L \le \theta \le U)$, computed with $\theta$ held fixed and the data random. The nominal level is the advertised 95%; the actual coverage is what really happens, which can be lower when assumptions fail (small samples, wrong SE, dependence).
  • Before the data are seen, "$[L, U]$ will contain $\theta$" is a random event with probability 0.95. After the data are seen, the interval is two fixed numbers $[l, u]$ and $\theta$ is a fixed number, so in the frequentist framework there is no probability left: the statement is simply true or false.
  • A confidence interval is about a parameter (such as the mean). It is not a range that holds 95% of the data (that is a reference or prediction interval), and it does not predict where 95% of future sample means will fall.
Why do we need it?

Most misreadings of statistics come from saying the wrong sentence about an interval. Interviewers test this directly, and a wrong sentence can lead to wrong business decisions ("we are 95% sure the lift is in this range, so let's plan revenue on it").

Where is it used?

Every report that carries an interval: A/B results, regression tables, survey margins, model metrics with error bars, scientific papers. It is also the property that simulation studies check when they evaluate a new interval method ("does it really cover 95%?").

How is it used?

Say: "This interval was built by a method that captures the true value in 95% of repeated samples." Use simulations like the one below to check the actual coverage of any interval recipe before trusting it on your kind of data.

Each horizontal line is the 95% interval from one simulated sample of daily orders; the green line is the true mean. Blue intervals contain it, red ones miss. Press New 100 intervals several times: about 5 red each time (sometimes 2, sometimes 9). Press Add 5 000 intervals to the tally to see the long-run coverage settle near the level. Then choose z with s (wrong) and $n = 4$: plugging the sample standard deviation into the z formula gives too-short intervals, and coverage drops to about 86%.

"There is a 95% probability that the true mean is between 46.08 and 53.92."

In the frequentist framework the true mean is fixed, so it is either in [46.08, 53.92] or not. The 95% is the long-run success rate of the recipe: "95% of intervals built this way contain μ".

"95% of the days had between 46 and 54 orders."

The interval is about the mean. Individual days vary much more (σ = 12 orders). A range for single days is a prediction interval, which is far wider.

"If we repeat the study, 95% of the new sample means will fall inside this interval."

Both the old and the new mean wobble, so the gap between them has standard deviation $\sqrt2 \cdot SE$. A repeat of the same size lands inside the old 95% interval only about 83% of the time on average.

"A 95% confidence interval contains the true value with 95% probability."

"95% of the intervals produced by this procedure, over repeated samples, contain the true value. This particular interval either does or does not."

Model answer: "In frequentist statistics the parameter is fixed and the interval is random. The 95% is a property of the method: if we repeated the experiment many times, 95% of the intervals would cover the true value. Once I have computed one interval, I cannot attach a 95% probability to it without a prior; that probability statement is what a Bayesian credible interval gives."

95% confidence = the recipe covers the fixed true value in 95% of repeated samples.

One computed interval either contains θ or not; misses in 100 intervals ~ Binomial(100, 0.05).

Trap: never say "95% probability θ is in [l, u]" about a confidence interval.

Quick check: you build 20 independent 95% intervals. How many misses do you expect, and what is the chance that all 20 cover?

Expected misses: $20 \times 0.05 = 1$. Chance all cover: $0.95^{20} \approx 0.36$. So about two times in three, at least one of 20 intervals misses.

Interval for a mean: the z version and the t version core

The recipe $\bar x \pm 1.96\,\sigma/\sqrt n$ needs $\sigma$, the true standard deviation. In real life you almost never know $\sigma$; you estimate it with the sample standard deviation $s$. But $s$ is itself a guess that wobbles from sample to sample, and with few data points it is often too small. If you plug $s$ into the z recipe, your net is too often too narrow.

The fix is a slightly bigger multiplier that accounts for the extra uncertainty in $s$: the Student-t multiplier. With 3 data points it is 4.30 instead of 1.96; with 30 points it is 2.05; with hundreds it is practically 1.96 again.

Three ways to say it:

  • Picture: when you are unsure how wide the net should be, make it a bit wider to be safe.
  • Numbers: heights 160, 170, 180: the honest 95% interval is 170 ± 24.8, not 170 ± 11.3.
  • Slogan: σ unknown → use t with $n-1$ degrees of freedom.

Three heights (your example from Chapter 5.5): 160, 170, 180 cm.

  1. Mean: $\bar x = (160 + 170 + 180)/3 = 170$.
  2. Sample standard deviation: squared distances $100, 0, 100$, sum $200$, divided by $n - 1 = 2$ gives $100$, so $s = 10$.
  3. Standard error: $SE = s/\sqrt n = 10/\sqrt3 = 5.774$.
  4. Degrees of freedom $n - 1 = 2$; the t multiplier is $t_{2,\,0.975} = 4.303$ (from a t table or scipy.stats.t.ppf(0.975, 2)).
  5. Margin: $4.303 \times 5.774 = 24.84$. Interval: $170 \pm 24.84 = [145.2,\ 194.8]$ cm.
  6. If you wrongly used 1.96: $170 \pm 11.32 = [158.7,\ 181.3]$. Intervals built this way with $n = 3$ cover the true mean only about 81% of the time, not 95%.

The honest interval is very wide. That is the truth about three data points: they cannot pin down an average height to within a few centimetres.

For data $x_1, \dots, x_n$ with mean $\bar x$ and sample standard deviation $s$:

  • σ known (z interval): $\bar x \pm z_{1-\alpha/2}\,\dfrac{\sigma}{\sqrt n}$.
  • σ unknown (t interval): $\bar x \pm t_{n-1,\,1-\alpha/2}\,\dfrac{s}{\sqrt n}$, where $t_{n-1,\,1-\alpha/2}$ is the quantile of the Student-t distribution with $n-1$ degrees of freedom (the number of independent pieces of information left for estimating the spread after estimating the mean).
  • Why t: if the data are Normal, $\dfrac{\bar X - \mu}{S/\sqrt n}$ follows a $t_{n-1}$ distribution exactly. Its tails are heavier than the Normal's, so its quantiles are larger.
  • 95% multipliers: $n = 3$: 4.303 · $n = 5$: 2.776 · $n = 10$: 2.262 · $n = 30$: 2.045 · $n = 100$: 1.984 · $n \to \infty$: 1.960.
  • Assumptions: independent observations; Normal data, or $n$ large enough for the CLT. For small samples of skewed data (revenue, waiting times) the t interval under-covers; use a transformation or the bootstrap (Chapter 5.5).
  • Code: scipy.stats.t.interval(0.95, df=n-1, loc=xbar, scale=s/np.sqrt(n)); with scipy.stats.sem(x) for $s/\sqrt n$ (it uses ddof=1).
Why do we need it?

Small samples are common: a few days of a new metric, a handful of segments, a few model runs with different seeds. Using 1.96 with an estimated $s$ there gives intervals that look precise but miss the truth far more than 5% of the time.

Where is it used?

The one-sample and Welch t-tests and their intervals (Chapter 5.9), regression coefficient intervals (statsmodels uses t with $n-p$ degrees of freedom), error bars over a few training seeds, and average order value in A/B reports.

How is it used?

Default to the t interval whenever $\sigma$ is estimated: it costs nothing for large $n$ and protects you for small $n$. Look at the data first: if they are very skewed and $n$ is small, prefer a bootstrap interval or a log transform.

The orange curve is the Student-t distribution with $n-1$ degrees of freedom, the blue curve the standard Normal; the markers show where each one cuts off 2.5% in each tail. At $n = 3$ the t cut-off sits at 4.30, far beyond 1.96. Each change runs 1 000 simulated samples and reports how often each interval covered the truth. Slide $n$ up to 30 and watch the two converge. Then switch the data to skewed: even the t interval falls short of 95% for small $n$, because its assumption (Normal data) is broken.

"n is above 30, so I can always use 1.96."

"30" is a rule of thumb. At $n = 30$ the t multiplier is still 2.045, and for strongly skewed data (revenue with a few huge orders) 30 points can be far too few for the CLT. Using t costs nothing.

"The t interval works for any data shape."

It is exact only for Normal data. For small, skewed samples its coverage falls below 95% (about 88% for exponential data with $n = 5$).

"np.std(x) gives the s in the formula."

NumPy's default divides by $n$; use np.std(x, ddof=1) or scipy.stats.sem(x).

σ known: $\bar x \pm z_{1-\alpha/2}\,\sigma/\sqrt n$. σ estimated: $\bar x \pm t_{n-1,\,1-\alpha/2}\,s/\sqrt n$.

Heights 160, 170, 180: $170 \pm 4.303 \times 5.774 = [145.2, 194.8]$.

Trap: z with s under-covers for small $n$; t assumes roughly Normal data.

Quick check: 10 measurements have x̄ = 20 and s = 3. Give the 95% t interval.

$SE = 3/\sqrt{10} = 0.949$; $t_{9,\,0.975} = 2.262$; margin $= 2.262 \times 0.949 = 2.15$. Interval: $[17.85,\ 22.15]$.

Margin of error: how much data for how much precision?

The margin of error is the "±" part of the interval: multiplier × standard error. The standard error shrinks like $1/\sqrt n$, so the margin does too. That has a harsh consequence: to cut the margin in half you need four times the data; to cut it to a tenth you need a hundred times the data.

It also lets you plan backwards: decide how precise the answer must be ("within ±1 point"), then compute the $n$ that delivers it.

Three ways to say it:

  • Picture: a blurry photo sharpens slowly: each extra bit of sharpness costs more data than the last.
  • Numbers: a 10% rate with 500 users is known to ±2.6 points; ±1 point needs about 3 458 users.
  • Slogan: four times the data, half the margin.
  1. Checkout control group: $\hat p = 0.10$, $n = 500$. $SE = \sqrt{0.10 \times 0.90/500} = \sqrt{0.00018} = 0.0134$.
  2. 95% margin: $1.96 \times 0.0134 = 0.0263$, about ±2.6 points.
  3. Target ±1 point: solve $1.96\sqrt{0.09/n} = 0.01$ → $n = 1.96^2 \times 0.09 / 0.01^2 = 3.8415 \times 0.09 / 0.0001 = 3457.3$ → 3 458 users.
  4. A poll with no idea of the rate uses the worst case $p = 0.5$ (where $p(1-p) = 0.25$ is largest). For ±3 points: $n = 3.8415 \times 0.25 / 0.03^2 = 1067.1$ → 1 068 people. This is why many polls survey about 1 000 people and report "± 3 points".
  5. The margin for a difference of two equal-sized groups is $\sqrt2 \approx 1.41$ times each group's margin: two groups of 500 at 10% give about ±3.7 points for the gap.
  • Margin of error: $m = z_{1-\alpha/2} \cdot SE$, half the width of the interval.
  • Mean: $m = z\,\sigma/\sqrt n$, so $n = \left(\dfrac{z\,\sigma}{m}\right)^2$.
  • Proportion: $m = z\sqrt{p(1-p)/n}$, so $n = \dfrac{z^2\,p(1-p)}{m^2}$; with no guess for $p$, use $p = 0.5$ (the most conservative choice).
  • Difference of two independent groups: $m = z\sqrt{SE_1^2 + SE_2^2}$.
  • Planning for a margin is not the same as planning for power (Chapter 5.7): a margin plan controls the width of the interval; a power plan controls the chance of detecting a given effect. For a difference, the "80% power" sample size is about $(2.80/1.96)^2 \approx 2$ times the "margin equals the effect" sample size.
Why do we need it?

"How many users or days do we need to know this metric well enough?" is asked before every survey, dashboard metric and model evaluation. The margin formula answers it in one line.

Where is it used?

Polls and surveys ("± 3 points"), monitoring dashboards (is today's dip outside normal noise?), evaluating a model on a test set (how many labelled examples make an accuracy estimate precise to ±1 point), and A/B test planning.

How is it used?

Pick the precision you need in the metric's units, plug a guess for $p$ (or $\sigma$ from history) into $n = z^2 p(1-p)/m^2$, and round up. If the answer is unaffordable, accept a wider margin or a lower confidence level.

The purple curve is the 95% margin of error for a conversion rate, against the number of users (log scale). Drag the purple handle along the bottom to choose $n$. The green dot marks $4n$: its margin is exactly half. Move the rate to 50% (the worst case) and to 2% (rare events have small absolute margins, but large ones relative to the rate). Set a target margin to see the $n$ that achieves it.

"Doubling the data halves the margin."

The margin falls like $1/\sqrt n$: doubling $n$ cuts it by only about 29% (to $1/\sqrt2 \approx 0.71$). Halving it needs four times the data.

"A ±3-point poll margin also covers the difference between two candidates."

The margin of a difference is larger than the margin of each share (for two shares from the same poll it can be up to twice as large). Compute the SE of the difference itself.

$m = z \cdot SE$; mean: $n = (z\sigma/m)^2$; proportion: $n = z^2 p(1-p)/m^2$ (worst case $p = 0.5$).

±1 point on a 10% rate: 3 458 users; ±3 points worst case: 1 068.

Trap: margin ∝ $1/\sqrt n$, so halving it needs 4× the data.

Quick check: a 95% margin is ±4 points with 600 users. About how many users give ±2 points?

Halving the margin needs 4 times the data: about $4 \times 600 = 2400$ users.

Interval for a proportion: Wald, and why Wilson is safer core

The textbook interval for a conversion rate is the Wald interval: $\hat p \pm 1.96\sqrt{\hat p(1-\hat p)/n}$. It plugs the observed rate into the standard error. With lots of data and a rate that is not close to 0 or 1, it works well. But look what happens with a new feature where 0 of 20 users converted: $\hat p = 0$, so the standard error is 0 and the interval is $[0, 0]$. It claims we are certain the rate is exactly zero after 20 users. That is absurd.

The Wilson interval asks a smarter question: "which true rates $p$ would find my data unsurprising?", using each candidate $p$'s own standard error. It never collapses to zero width and never leaves $[0, 1]$. For 0 of 20 it says the rate could plausibly be anywhere from 0 to 16%.

Three ways to say it:

  • Picture: Wald measures the net with a ruler made from the catch; with no catch, the ruler has length zero.
  • Numbers: 1 of 20: Wald gives −4.6% to 14.6% (impossible negative rates); Wilson gives 0.9% to 23.6%.
  • Slogan: for rare events or small samples, use Wilson (or an exact method), not Wald.

1 conversion out of 20 users, 95% level ($z = 1.96$, $z^2 = 3.8415$).

  1. Wald: $\hat p = 0.05$, $SE = \sqrt{0.05 \times 0.95/20} = 0.0487$, margin $1.96 \times 0.0487 = 0.0955$, interval $[-0.046,\ 0.146]$. The lower end is a negative rate.
  2. Wilson centre: $\dfrac{\hat p + z^2/(2n)}{1 + z^2/n} = \dfrac{0.05 + 0.0960}{1.1921} = 0.1225$. It is pulled from 0.05 toward 0.5.
  3. Wilson half-width: $\dfrac{z\sqrt{\hat p(1-\hat p)/n + z^2/(4n^2)}}{1 + z^2/n} = \dfrac{1.96\sqrt{0.002375 + 0.002401}}{1.1921} = \dfrac{1.96 \times 0.0691}{1.1921} = 0.1136$.
  4. Wilson interval: $0.1225 \pm 0.1136 = [0.009,\ 0.236]$. Lopsided, inside $[0,1]$, and honest about how little 20 users tell us.
  5. With plenty of data the two agree: for 60 of 500, Wald gives $[9.2\%,\ 14.8\%]$ and Wilson $[9.4\%,\ 15.1\%]$.

For $k$ successes in $n$ independent trials, $\hat p = k/n$, multiplier $z = z_{1-\alpha/2}$:

  • Wald: $\hat p \pm z\sqrt{\hat p(1-\hat p)/n}$. Uses the estimated SE. Rule of thumb for when it is acceptable: at least about 10 successes and 10 failures.
  • Wilson (score) interval: all values $p$ with $\dfrac{|\hat p - p|}{\sqrt{p(1-p)/n}} \le z$. Solving this quadratic in $p$ gives $$\frac{\hat p + \frac{z^2}{2n}}{1 + \frac{z^2}{n}} \;\pm\; \frac{z}{1 + \frac{z^2}{n}}\sqrt{\frac{\hat p(1-\hat p)}{n} + \frac{z^2}{4n^2}}.$$
  • Clopper–Pearson ("exact"): built from Beta quantiles, $[\,B_{\alpha/2}(k,\,n-k+1),\ B_{1-\alpha/2}(k+1,\,n-k)\,]$. Its coverage is never below 95%, so it is conservative (wider than needed).
  • Coverage of Wald can be terrible: with $n = 20$ and a true rate of 5%, Wald covers only about 64% of the time; Wilson about 92%. For $n = 100$ and 5%: Wald 88%, Wilson 97%.
  • Code: statsmodels.stats.proportion.proportion_confint(k, n, method='wilson'). Careful: its default is method='normal', which is Wald.
Why do we need it?

Conversion rates are often small (1–5%) and segments or new features often have few users. That is exactly where Wald breaks: zero-width intervals, negative rates, and coverage far below the advertised 95%.

Where is it used?

Conversion and click-through rates in dashboards, error rates of classifiers on small test sets, defect rates in quality control, survey proportions, and rates in small segments of an A/B test.

How is it used?

Default to method='wilson' (or 'beta' for Clopper–Pearson when you must never under-cover). Keep Wald for quick mental arithmetic with large counts. Always check the number of successes, not just $n$.

Choose the number of users and the number who converted. The three bars are the 95% intervals from the three methods; the purple dot is $\hat p$. Start with 1 of 20: Wald (orange) pokes below 0%. Set conversions to 0: Wald shrinks to a single point. Press 60 of 500: all three nearly agree.

For each true rate $p$ (horizontal axis) this computes exactly how often the 95% interval would contain $p$, by adding up the Binomial probabilities of every possible count. The green line is the promised 95%. Wald (orange) collapses near 0% and 100%; Wilson (blue) stays close to 95%. Both zig-zag because counts are whole numbers. Drag the purple handle to read the coverage at a rate, and raise $n$: Wald improves only slowly near the edges.

"0 conversions out of 20, so the 95% interval is [0, 0]."

That is Wald breaking down. Wilson gives [0%, 16.1%]; Clopper–Pearson [0%, 16.8%]. (A quick rule for zero successes, the "rule of three", gives an upper 95% bound of about $3/n = 15\%$.)

"proportion_confint(k, n) gives a good interval."

Its default method is 'normal' (Wald). Pass method='wilson' explicitly.

"n = 1 000 is large, so Wald is fine."

What matters is the number of successes and failures. A 0.2% rate with 1 000 users has only about 2 successes: Wald is unreliable there.

Wald: $\hat p \pm z\sqrt{\hat p(1-\hat p)/n}$; fails for small $n$ or $\hat p$ near 0 or 1 (zero width, negative limits, low coverage).

Wilson: invert the score test; centre pulled toward 0.5; stays in [0, 1]. 1/20 → [0.9%, 23.6%].

Trap: statsmodels' default is Wald; use method='wilson'.

Quick check: why does the Wilson interval's centre move toward 0.5?

Wilson uses each candidate rate's own standard error $\sqrt{p(1-p)/n}$, which is larger for rates nearer 0.5. Rates a bit closer to 0.5 than $\hat p$ are therefore "less surprised" by the data than rates on the other side, so the set of plausible rates leans toward 0.5. The centre is $(\hat p + z^2/2n)/(1 + z^2/n)$, roughly "add about 2 successes and 2 failures" at 95%.

Interval for a difference: the lift and its uncertainty core

In an A/B test the number you care about is the gap between the groups. Each group's rate wobbles, and the gap wobbles more than either one, because the two wobbles add up. The standard error of a difference of independent estimates is $\sqrt{SE_A^2 + SE_B^2}$ (Chapter 5.5). Then the usual recipe applies: gap ± 1.96 × SE.

This interval is the most useful single output of a classical A/B test. It answers both questions at once: "is there evidence of an effect?" (does it exclude 0?) and "how big could it be?" (where are its ends?).

Three ways to say it:

  • Picture: two wobbly rulers; the distance between their tips wobbles more than either tip.
  • Numbers: checkout lift +2.0 points, 95% CI from −1.9 to +5.9 points.
  • Slogan: report the lift with its interval, not just a p-value.

The checkout test: A = 50/500 = 0.10, B = 60/500 = 0.12.

  1. Observed lift: $\hat p_B - \hat p_A = 0.12 - 0.10 = 0.02$.
  2. Variance of each rate: $\dfrac{0.10 \times 0.90}{500} = 0.000180$ and $\dfrac{0.12 \times 0.88}{500} = 0.0002112$.
  3. SE of the difference: $\sqrt{0.000180 + 0.0002112} = \sqrt{0.0003912} = 0.01978$.
  4. Margin: $1.96 \times 0.01978 = 0.0388$.
  5. 95% CI: $0.02 \pm 0.0388 = [-0.0188,\ 0.0588]$, i.e. from −1.9 to +5.9 percentage points.
  6. It contains 0, which matches the test of Chapter 5.6 ($p \approx 0.31$, not significant). But it says more: the data are compatible with a small loss and with a large win. The experiment was too small to decide (its power was about 17%, Chapter 5.7).

Means (Welch): order value A: mean 50, $s = 20$; B: mean 52, $s = 22$; 1 000 orders each. $SE = \sqrt{20^2/1000 + 22^2/1000} = \sqrt{0.884} = 0.940$; with about 1 980 degrees of freedom the t multiplier is 1.961; CI $= 2 \pm 1.843 = [0.16,\ 3.84]$ dollars.

  • Difference of proportions (Wald, unpooled): $(\hat p_B - \hat p_A) \pm z_{1-\alpha/2}\sqrt{\dfrac{\hat p_A(1-\hat p_A)}{n_A} + \dfrac{\hat p_B(1-\hat p_B)}{n_B}}$. With small counts, better versions exist (Newcombe's method, which combines two Wilson intervals; Agresti–Caffo, which adds one success and one failure to each group).
  • Difference of means (Welch): $(\bar x_B - \bar x_A) \pm t_{\nu,\,1-\alpha/2}\sqrt{\dfrac{s_A^2}{n_A} + \dfrac{s_B^2}{n_B}}$, with the Welch degrees of freedom $\nu$ (Chapter 5.9). It does not assume equal variances.
  • Requires the two groups to be independent (separate users). For paired data (the same users before and after), build the interval from the per-user differences instead.
  • Relative lift $(\theta_B - \theta_A)/\theta_A$ needs its own interval (delta method on the log of the ratio, or the bootstrap): dividing the absolute interval by $\hat p_A$ ignores the uncertainty in the baseline.
Why do we need it?

A p-value compresses everything into "significant or not". The interval of the difference keeps the size and the uncertainty, so a product team can compare it with the smallest lift worth shipping and with the cost of being wrong.

Where is it used?

Every A/B test report ("lift +2.0 pts, 95% CI −1.9 to +5.9"), experiment dashboards, treatment-effect estimates in medicine, and comparisons of two ML models on the same test set (there with the paired version).

How is it used?

Compute the gap and its SE, form the interval, then read it three ways: does it exclude 0 (significance)? Does it exclude effects too small to matter (practical significance, Chapter 5.10)? Is it so wide that the test was inconclusive (power)?

Top: each group's own 95% interval (blue = A, orange = B). Bottom: the 95% interval of the difference B − A (purple), with a red line at 0. Start with the checkout numbers: the lift interval crosses 0. Now press Overlap, yet significant (100 vs 130 of 1 000): the two group intervals overlap, yet the interval of the difference excludes 0 and the test gives $p \approx 0.035$. Judge the difference by its own interval, never by eyeballing two separate ones.

"The 95% intervals of A and B overlap, so the difference is not significant."

Two intervals can overlap while the difference is significant (100 vs 130 of 1 000: A [8.3%, 12.0%], B [11.1%, 15.2%] overlap, yet the lift interval is [+0.2, +5.8] points, $p = 0.035$). Build the interval of the difference itself. (The reverse direction holds for the usual symmetric intervals: if two 95% intervals do not overlap at all, the difference is significant at 5%.)

"The lift CI is [−1.9, +5.9] points, so the relative lift CI is [−19%, +59%]."

Dividing by the observed baseline 10% ignores the baseline's own uncertainty. Use the delta method on the log ratio or bootstrap the relative lift.

"The interval contains 0, so there is no effect."

It contains 0 and +5.9 points. The data cannot rule out a large win. Absence of evidence is not evidence of absence.

Your A/B framework reports a posterior for $\theta_B - \theta_A$ instead of a confidence interval. For the checkout data with flat Beta(1, 1) priors the posteriors are Beta(61, 441) and Beta(51, 451); the 95% equal-tailed credible interval of the difference (from posterior draws) is about $[-1.9,\ +5.9]$ points, practically the same numbers as the Wald interval, and $P(\theta_B \gt \theta_A \mid D) \approx 0.84$. Reviewers will notice the agreement; the next sections explain why the numbers agree here and why the sentences you may say about them are still different.

Gap ± $z\sqrt{SE_A^2 + SE_B^2}$. Checkout: $0.02 \pm 1.96 \times 0.01978 = [-1.9, +5.9]$ points.

Means: Welch $\pm\, t_\nu\sqrt{s_A^2/n_A + s_B^2/n_B}$; paired data → interval of the per-user differences.

Trap: overlapping group intervals do not mean "not significant"; a CI containing 0 does not mean "no effect".

Quick check: A = 200/2 000, B = 240/2 000. Is the lift significant at 5%?

$\hat p_A = 0.10$, $\hat p_B = 0.12$, lift 0.02. $SE = \sqrt{0.1 \times 0.9/2000 + 0.12 \times 0.88/2000} = \sqrt{0.0000978} = 0.00989$. Margin $1.96 \times 0.00989 = 0.0194$. CI $= [0.0006,\ 0.0394]$, just above 0: significant, barely. The same 2-point lift that was inconclusive with 500 per group becomes significant with 2 000.

Intervals and tests are two views of the same thing

A test asks about one candidate value: "is the true mean 46? reject or not?". A confidence interval answers that question for every candidate value at once. The 95% interval is simply the list of all values that a 5% two-sided test would not reject: the values the data find "plausible".

So you never need to run a separate test once you have the interval: if the null value (often 0 for a lift) lies outside the 95% interval, the 5% test rejects it; if it lies inside, it does not.

Three ways to say it:

  • Picture: the interval is a guest list; a test checks whether one particular name is on it.
  • Numbers: orders CI [46.08, 53.92]: testing μ = 46 gives p = 0.046 (rejected, outside); μ = 47 gives p = 0.13 (not rejected, inside).
  • Slogan: 95% interval = every value a 5% test would keep.

Orders: $\bar x = 50$, $SE = 2$, 95% CI $[46.08,\ 53.92]$.

  1. Test $H_0: \mu = 46$. $z = (50 - 46)/2 = 2.0$, two-sided $p = 2\,P(Z \gt 2) = 0.0455 \lt 0.05$: reject. And 46 is outside the interval.
  2. Test $H_0: \mu = 47$. $z = (50 - 47)/2 = 1.5$, $p = 0.134$: do not reject. And 47 is inside the interval.
  3. Test $H_0: \mu = 46.08$. $z = 3.92/2 = 1.96$, $p = 0.05$ exactly: the edge of the interval is exactly where the p-value crosses 0.05.
  4. Checkout: the lift interval $[-1.9,\ +5.9]$ points contains 0, and the test of "no lift" gives $p = 0.31$. The two agree.

Duality (inverting a test). For each candidate value $\theta_0$, let a level-$\alpha$ two-sided test decide "reject $H_0: \theta = \theta_0$" or not. Then

$$C = \{\theta_0 : \text{the test does not reject } \theta = \theta_0\}$$

is a $(1-\alpha)$ confidence interval: it contains the true $\theta$ exactly when the test does not wrongly reject it, which happens with probability $1 - \alpha$. Conversely, a $(1-\alpha)$ interval gives a test: reject $\theta_0$ when it lies outside.

  • Matching levels: a 95% two-sided interval ↔ a two-sided test at 5%; a one-sided test at 5% ↔ a one-sided 95% bound (for example "the lift is at least +0.3 points").
  • The match is exact only when the interval and the test use the same standard error. The two-proportion z-test uses the pooled SE (0.01979 for the checkout) and the Wald interval of the difference the unpooled one (0.01978), so in rare cases right at the edge they can disagree. The Wilson interval is the exact inversion of the one-proportion score test.
  • Plotting the two-sided p-value against $\theta_0$ gives the p-value function (or confidence curve); cutting it at height $\alpha$ gives the $(1-\alpha)$ interval.
Why do we need it?

It means one picture (the interval) carries both answers: whether an effect is "significant" and how large it plausibly is. It also explains why reporting intervals is strictly more informative than reporting p-values.

Where is it used?

Reading regression tables (a coefficient is significant at 5% exactly when its 95% interval excludes 0), A/B readouts, non-inferiority and equivalence tests (is the whole interval above −margin?), and building intervals for complicated estimators by inverting tests.

How is it used?

Check whether the null value is inside the interval to get the test decision at the matching level. To test against a business threshold instead of 0 ("is the lift above +0.5 points?"), just compare the threshold with the interval.

The black curve is the two-sided p-value of the test "$H_0: \mu = \mu_0$" for every candidate $\mu_0$, given $\bar x = 50$. The dashed red line is $\alpha$. The 95% interval (purple band) is exactly where the curve is above $\alpha$. Drag the purple handle (your $\mu_0$): inside the band the test does not reject; outside it rejects. Change $\alpha$ and the number of days: the band and the curve change together.

"The 95% CI excludes 0, so p < 0.05; and the 99% CI includes 0, so the result is wrong."

Both can be true at once: $0.01 \lt p \lt 0.05$. A 99% interval matches a 1% test. Match the levels before comparing.

"A value just inside the interval is proven to be plausible, one just outside is proven wrong."

The edge is a convention (where $p$ crosses $\alpha$). Values near the edge are about equally supported. The p-value curve shows that support changes smoothly.

95% CI = all $\theta_0$ that a two-sided 5% test does not reject. Null outside the CI ⇔ $p \lt 0.05$.

Orders: 46 is outside [46.08, 53.92] and $p = 0.046$; 47 is inside and $p = 0.13$.

Trap: match levels (95% ↔ 5%, 99% ↔ 1%); pooled vs unpooled SEs can disagree right at the edge.

Quick check: a regression coefficient has 95% CI [0.3, 1.1]. Is it significantly different from 0 at 5%? From 1?

0 is outside the interval, so yes, it is significantly different from 0 at the 5% level. 1 is inside, so the data do not reject the value 1 at 5%.

Confidence interval vs credible interval: explain it perfectly core

Two schools of statistics answer different questions, and the difference lives in what is treated as random (Chapter 4.1).

  • Frequentist (confidence interval). The true conversion rate θ is one fixed, unknown number. The data are random, so the interval is random. "95%" is the long-run hit rate of the recipe over imagined repeated experiments. After you compute $[9.4\%, 15.1\%]$, θ is simply in it or not.
  • Bayesian (credible interval). You describe your uncertainty about θ with a probability distribution: a prior before the data, a posterior after (Chapter 6.1). The data are fixed (you saw them). A 95% credible interval is a range that holds 95% of the posterior probability, so you may say "given the data and my prior, θ is in this range with probability 0.95".

Often the two give almost the same numbers. They still mean different things, and with small data or a strong prior the numbers differ too.

Three ways to say it:

  • Picture: a confidence interval is a net thrown by a reliable thrower at a fixed fish; a credible interval is a map of where the fish probably is, given everything you know.
  • Numbers: 60 of 500: Wilson CI [9.44%, 15.14%]; flat-prior credible interval [9.44%, 15.15%]. Same numbers, different sentences.
  • Slogan: confidence is about the method over repeats; credibility is about θ, given this data and a prior.

Large data, weak prior: 60 conversions out of 500.

  1. Frequentist: Wilson 95% CI $= [0.0944,\ 0.1514]$.
  2. Bayesian: flat prior Beta(1, 1). By the conjugate update (Chapter 6.3) the posterior is Beta$(1 + 60,\ 1 + 440)$ = Beta(61, 441).
  3. 95% equal-tailed credible interval = the 2.5% and 97.5% quantiles of Beta(61, 441) $= [0.0944,\ 0.1515]$.
  4. The numbers agree to three decimals because the prior is weak and 500 users carry a lot of information.

Small data, informative prior: 3 conversions out of 20, and past experiments suggest rates near 10% (prior Beta(10, 90): mean 10%, worth about 100 users of evidence).

  1. Wilson 95% CI for 3/20: $[0.052,\ 0.360]$. Very wide: 20 users say little.
  2. Posterior: Beta$(10 + 3,\ 90 + 17)$ = Beta(13, 107); 95% credible interval $= [0.059,\ 0.170]$.
  3. Very different numbers. The credible interval is much narrower because it also uses the prior's "100 users" of information. It is only as good as that prior: if past experiments are not like this one, it can be confidently wrong.

A $(1-\alpha)$ credible interval is any interval $[a, b]$ with $P(a \le \theta \le b \mid D) = 1-\alpha$ under the posterior $p(\theta \mid D) \propto p(D \mid \theta)\,p(\theta)$. The common choices are the equal-tailed interval (2.5% and 97.5% quantiles) and the highest-density interval; both, and posterior decisions, are taught in Chapter 6.4.

Confidence intervalCredible interval
What is fixed?θ (an unknown constant)the observed data $D$
What is random?the data, hence the interval's endsθ, in the sense of your uncertainty, described by $p(\theta \mid D)$
What "95%" means95% of intervals from repeated samples contain θ$P(\theta \in [a,b] \mid D) = 0.95$, given the model and prior
Needs a prior?noyes
Guaranteeabout 95% coverage for every θ (if the assumptions hold)a correct probability if model and prior are right; coverage holds on average over θ drawn from the prior, not for every θ
Typical recipeestimate ± z·SE, Wilson, t, bootstrapquantiles of the posterior (or of posterior draws)
The sentence you may say"95% of intervals built this way contain θ""given the data and prior, θ is in [a, b] with probability 0.95"

When do the numbers agree? With plenty of data and a prior that is smooth and not zero near the truth, the posterior becomes approximately Normal around the estimate with standard deviation ≈ SE (the Bernstein–von Mises theorem), so the credible interval ≈ the confidence interval. They differ with small samples, strong priors, parameters near a boundary, or many parameters.

Why do we need it?

The sentence most people want to say ("θ is in this range with 95% probability") is only justified for a credible interval. Mixing the two leads to wrong claims in reviews and interviews, and to silently trusting a prior nobody checked.

Where is it used?

Your Bayesian A/B framework reports credible intervals of rates and lifts; classical experimentation platforms report confidence intervals; forecasting models report posterior predictive intervals (a Bayesian cousin, Chapter 7.14). Comparing the two is a standard interview topic.

How is it used?

State which interval you report, with its correct sentence. With a weak prior and large data, show both: their agreement reassures frequentist reviewers. With small data, show how much the credible interval depends on the prior (a sensitivity check, Chapter 6.8).

Confidence interval Credible interval θ fixed miss data random → interval random 95% of such intervals cover θ posterior p(θ | D) 95% of the area data fixed → θ uncertain P(a ≤ θ ≤ b | D) = 0.95
Left: the truth stands still and the intervals move from sample to sample; the 95% is their hit rate. Right: the data are what they are; the posterior spreads probability over values of θ, and the credible interval holds 95% of it.

Set the data (users and observed rate) and a Beta prior (its mean and its strength in "pretend users"). The orange curve is the posterior; its shaded middle 95% is the credible interval (orange bar); the blue bar is the Wilson confidence interval. With flat prior and 500 users the bars coincide. Now press Small data, strong prior (3 of 20, prior centred on 10% worth 100 users): the credible interval is far narrower, and it is pulled toward the prior. Move the prior mean to 3%: the credible interval follows the prior; the confidence interval does not care.

Here we fix the true rate θ (green) and repeat a 50-user experiment many times. Left: Wilson confidence intervals; right: credible intervals from the chosen prior. Red = misses. The readout gives the exact long-run coverage. With the flat prior both cover about 95%. Choose the wrong prior (centred on 3%, worth 100 users) with θ = 12%: the credible intervals miss most of the time, while Wilson still covers about 95%. Choose the right prior (centred on 10%) and move θ to 20%: a prior is only as good as its match to reality.

"Confidence and credible intervals are the same thing with different names."

They answer different questions and need different assumptions. They often give similar numbers (large data, weak prior) and can give very different numbers (small data, strong prior).

"A credible interval is always better, because it lets me say '95% probability'."

That sentence is only as trustworthy as the prior and the model. With a wrong, strong prior a credible interval can miss the truth most of the time (the widget above: about 14% coverage). Check prior sensitivity.

"A confidence interval tells me nothing about θ after I see the data."

It tells you which values are compatible with the data (the values a test would keep), backed by a guarantee about the method. It just does not attach a probability to θ itself.

In your A/B framework, the interval you report for $\theta_B - \theta_A$ or for the relative lift is a credible interval computed from posterior draws, so you may say "given the data and our priors, the lift is between … and … with 95% probability", and you can also report $P(\theta_B \gt \theta_A \mid D)$, which has no confidence-interval counterpart. Be ready for the reviewer's follow-up: "how much does that depend on your priors?". Answer with the large-data agreement shown above (flat or weak priors give numbers close to the classical interval) and, for small segments, with a prior-sensitivity check. Hierarchical pooling across segments is a prior learned from the other segments, which is why small segments' credible intervals are narrower and pulled toward the overall mean (Chapter 6.6).

"A confidence interval is the range where the parameter lies with 95% probability; a credible interval is the Bayesian name for the same thing."

"A confidence interval's 95% is a property of the procedure over repeated samples, with θ fixed. A credible interval's 95% is a posterior probability about θ, given the observed data and a prior."

Model answer: "In the frequentist view θ is a fixed constant and the data are random, so the interval is random: a 95% confidence interval comes from a recipe that captures θ in 95% of repeated experiments; once computed, it either contains θ or not. In the Bayesian view I treat θ as uncertain and update a prior into a posterior; a 95% credible interval contains 95% of the posterior probability, so I can say θ lies in it with probability 0.95, conditional on my model and prior. With lots of data and weak priors the two intervals nearly coincide numerically (Bernstein–von Mises), but the interpretations differ; with small data or strong priors the numbers differ too, and the credible interval is only as good as the prior. In our A/B framework we report credible intervals and posterior probabilities like P(B better than A), and we check that they agree with the classical interval when the data are large."

Confidence: θ fixed, interval random; 95% = hit rate of the recipe over repeated samples; no prior.

Credible: data fixed, θ uncertain; $P(a \le \theta \le b \mid D) = 0.95$; needs a prior; correct if the prior and model are.

60/500: CI [9.44%, 15.14%] ≈ flat-prior credible [9.44%, 15.15%]. 3/20 with a Beta(10, 90) prior: CI [5.2%, 36.0%] vs credible [5.9%, 17.0%].

Quick check: with a flat prior and 2 000 users, should a credible interval and a Wilson CI for a conversion rate be close? Why?

Yes. With so much data the likelihood dominates the flat prior, and the posterior is nearly Normal around $\hat p$ with standard deviation close to the standard error, so the two intervals almost coincide numerically (Bernstein–von Mises). Their meanings still differ.

Recap, cheat sheet and practice

  • A confidence interval is estimate ± multiplier × SE: the net instead of the spear. Multipliers 1.645 / 1.96 / 2.576 for 90 / 95 / 99%.
  • "95%" describes the recipe: over repeated samples, 95% of the intervals contain the fixed true value. One computed interval either contains it or not.
  • Means: use $t_{n-1}$ instead of $z$ when $\sigma$ is estimated (heights: 170 ± 24.8). The t interval assumes roughly Normal data or a large $n$.
  • The margin falls like $1/\sqrt n$: four times the data halves it. $n = z^2 p(1-p)/m^2$ for a proportion.
  • Proportions: Wald breaks for small $n$ or rates near 0/1 (0 of 20 → [0, 0]); use Wilson (or Clopper–Pearson). statsmodels' default is Wald.
  • Differences: gap ± $z\sqrt{SE_A^2 + SE_B^2}$. Checkout lift: [−1.9, +5.9] points. Overlapping group intervals do not imply "not significant".
  • Duality: a 95% interval contains exactly the values a 5% two-sided test would not reject.
  • Confidence vs credible: confidence = θ fixed, interval random, a promise about the method; credible = data fixed, θ uncertain, $P(\theta \in [a,b] \mid D) = 0.95$ given a prior. Similar numbers with big data and weak priors; different with small data or strong priors.

Cheat sheet

IntervalFormulaNotes
Mean, σ known$\bar x \pm z_{1-\alpha/2}\,\sigma/\sqrt n$orders: 50 ± 3.92
Mean, σ unknown$\bar x \pm t_{n-1,\,1-\alpha/2}\,s/\sqrt n$$t_{2} = 4.30$, $t_{9} = 2.26$, $t_{29} = 2.05$
Proportion (Wald)$\hat p \pm z\sqrt{\hat p(1-\hat p)/n}$only with ≳10 successes and 10 failures
Proportion (Wilson)$\dfrac{\hat p + \frac{z^2}{2n} \pm z\sqrt{\frac{\hat p(1-\hat p)}{n} + \frac{z^2}{4n^2}}}{1 + z^2/n}$stays in [0, 1]; 1/20 → [0.9%, 23.6%]
Difference of rates$(\hat p_B - \hat p_A) \pm z\sqrt{\frac{\hat p_A\hat q_A}{n_A} + \frac{\hat p_B\hat q_B}{n_B}}$$\hat q = 1 - \hat p$; checkout [−1.9, +5.9] pts
Difference of means$(\bar x_B - \bar x_A) \pm t_\nu\sqrt{s_A^2/n_A + s_B^2/n_B}$Welch degrees of freedom $\nu$
Margin and $n$$m = z\cdot SE$; $n = z^2p(1-p)/m^2$, $n = (z\sigma/m)^2$±3 pts worst case → 1 068
DualityCI = $\{\theta_0 : p(\theta_0) \ge \alpha\}$null outside 95% CI ⇔ $p \lt 0.05$
Credible (Beta-Binomial)quantiles of Beta$(a + k,\ b + n - k)$60/500, flat: [9.44%, 15.15%]
Code it · Python

import numpy as np
from scipy import stats
from statsmodels.stats.proportion import proportion_confint

# 1) Mean, sigma known (z) and unknown (t)
print(50 + np.array([-1, 1]) * stats.norm.ppf(0.975) * 12 / np.sqrt(36))
# [46.08 53.92]
x = np.array([160, 170, 180])
print(stats.t.interval(0.95, df=len(x) - 1, loc=x.mean(), scale=stats.sem(x)))   # sem uses ddof=1
# (145.158..., 194.841...)

# 2) One proportion: Wald vs Wilson vs Clopper-Pearson ("beta")
for k, n in [(60, 500), (1, 20), (0, 20)]:
    for m in ["normal", "wilson", "beta"]:          # "normal" (= Wald) is the DEFAULT method!
        lo, hi = proportion_confint(k, n, alpha=0.05, method=m)
        print(f"{k}/{n} {m:7s} [{lo:.4f}, {hi:.4f}]")
# 60/500 normal  [0.0915, 0.1485]   wilson [0.0944, 0.1514]   beta [0.0928, 0.1518]
# 1/20   normal  [0.0000, 0.1455]   (statsmodels clips Wald at 0; unclipped it is -0.0455)
# 1/20   wilson  [0.0089, 0.2361]   beta [0.0013, 0.2487]
# 0/20   normal  [0.0000, 0.0000]   wilson [0.0000, 0.1611]   beta [0.0000, 0.1684]

# 3) Difference of two proportions (checkout), Wald with unpooled SE
kA, nA, kB, nB = 50, 500, 60, 500
pA, pB = kA / nA, kB / nB
se = np.sqrt(pA * (1 - pA) / nA + pB * (1 - pB) / nB)
print(pB - pA, se, (pB - pA) - 1.96 * se, (pB - pA) + 1.96 * se)
# 0.02 0.01978 -0.0188 0.0588

# 4) Coverage simulation: 10 000 intervals for a mean (Normal data, n = 5)
rng = np.random.default_rng(1)
X = rng.normal(50, 12, size=(10_000, 5))
m, s = X.mean(axis=1), X.std(axis=1, ddof=1)
for name, mult in [("t", stats.t.ppf(0.975, 4)), ("z with s", 1.96)]:
    half = mult * s / np.sqrt(5)
    print(name, np.mean((m - half <= 50) & (50 <= m + half)))
# t 0.9538          (theory 0.95; simulation noise is about +-0.002)
# z with s 0.884    (theory 0.878: too short)

# 5) Exact coverage of Wald vs Wilson at n = 20, true p = 0.05
k = np.arange(21); pmf = stats.binom.pmf(k, 20, 0.05)
for m in ["normal", "wilson"]:
    lo, hi = proportion_confint(k, 20, method=m)
    print(m, round(pmf[(lo <= 0.05) & (0.05 <= hi)].sum(), 3))
# normal 0.639    wilson 0.925

# 6) Credible interval from a Beta posterior (flat prior), compared with Wilson
post = stats.beta(1 + 60, 1 + 440)
print(post.ppf([0.025, 0.975]))       # [0.0944 0.1515]  ~ Wilson [0.0944, 0.1514]
post = stats.beta(10 + 3, 90 + 17)    # 3/20 with an informative Beta(10, 90) prior
print(post.ppf([0.025, 0.975]))       # [0.0595 0.1695]  vs Wilson [0.0524, 0.3604]
Test yourself

1. A 95% confidence interval for a mean is [46.1, 53.9]. Which sentence is correct?

The true mean is fixed; the interval is random. "95%" is the long-run success rate of the recipe. The first sentence needs a Bayesian credible interval; the third describes a range for single observations; the fourth is wrong (about 83% on average).

2. Your 95% margin of error is ±4 points. To get ±2 points you need…

The margin is proportional to $1/\sqrt n$; halving it requires $n \times 4$.

3. A new feature: 0 of 25 users converted. Which interval should you report?

Wald collapses to zero width when $\hat p = 0$. Wilson gives $[0,\ 0.133]$ for 0/25, honestly saying the rate could still be over 10%. (Rule of three: about $3/25 = 12\%$.)

4. Group A's 95% interval is [8.3%, 12.0%] and group B's is [11.1%, 15.2%]. What can you conclude about the difference?

Overlapping intervals can still hide a significant difference (here 100 vs 130 of 1 000: lift CI [+0.2, +5.8] points, $p = 0.035$). Judge the difference by its own interval.

5. The 95% interval for a lift is [+0.3, +3.1] points. The two-sided p-value for "lift = 0" is…

By duality, 0 lies outside the 95% interval, so a 5% two-sided test rejects it: $p \lt 0.05$.

6. Which statement about credible intervals is correct?

A credible interval is a posterior probability statement, conditional on the prior. With a wrong, strong prior its coverage at a fixed θ can be far below 95%. With lots of data and a weak prior it nearly coincides with a confidence interval.

Practice problems

A. Orders on five days: 12, 15, 9, 14, 10. Give a 95% interval for the mean daily orders (σ unknown).
  1. $\bar x = 60/5 = 12$. Deviations $0, 3, -3, 2, -2$; squares $0, 9, 9, 4, 4$; sum 26.
  2. $s^2 = 26/4 = 6.5$, $s = 2.55$; $SE = 2.55/\sqrt5 = 1.140$.
  3. $t_{4,\,0.975} = 2.776$; margin $= 2.776 \times 1.140 = 3.17$.
  4. Interval: $12 \pm 3.17 = [8.83,\ 15.17]$.
B. How many people must a survey ask so that a 95% interval for a proportion is at most ±2 points, with no idea of the true proportion?

Worst case $p = 0.5$: $n = 1.96^2 \times 0.25 / 0.02^2 = 3.8415 \times 0.25 / 0.0004 = 2400.9$ → 2 401 people.

C. 2 of 40 users clicked. Compute the Wald and the Wilson 95% intervals.

Wald: $\hat p = 0.05$, $SE = \sqrt{0.05 \times 0.95/40} = 0.0345$, margin $0.0675$: $[-0.018,\ 0.118]$, with an impossible negative end.

Wilson: $1 + z^2/n = 1 + 3.8415/40 = 1.0960$; centre $(0.05 + 0.0480)/1.0960 = 0.0894$; half-width $1.96\sqrt{0.0011875 + 0.0006002}/1.0960 = 1.96 \times 0.04228 / 1.0960 = 0.0756$: $[0.014,\ 0.165]$.

D. A = 120/1 000, B = 160/1 000. Give the 95% interval for the lift and say whether it is significant.

Lift $0.16 - 0.12 = 0.04$. $SE = \sqrt{0.12 \times 0.88/1000 + 0.16 \times 0.84/1000} = \sqrt{0.0001056 + 0.0001344} = \sqrt{0.00024} = 0.01549$. Margin $1.96 \times 0.01549 = 0.0304$. Interval $[0.0096,\ 0.0704]$: from +1.0 to +7.0 points. It excludes 0, so it is significant at 5%; it also shows the lift could be as small as 1 point.

E. A 99% interval for a lift is [−0.2, +4.1] points. What can you say about the two-sided p-value for "no lift", and what is the 95% interval?
  1. 0 is inside the 99% interval, so $p \gt 0.01$.
  2. Centre $= (−0.2 + 4.1)/2 = 1.95$; half-width $2.15 = 2.576 \times SE$, so $SE = 0.835$.
  3. $z = 1.95/0.835 = 2.34$, so $p = 0.019$: significant at 5%, not at 1%.
  4. 95% interval: $1.95 \pm 1.96 \times 0.835 = [0.31,\ 3.59]$, which excludes 0, as the duality says it must.
F. (Interview) Your product manager says: "The 95% CI for the lift is [0.5, 3.0] points, so there is a 95% chance the true lift is in there." How do you respond, and what would let you say that sentence?

"Strictly, that sentence is about a credible interval, not a confidence interval. The confidence interval's 95% means the method captures the true lift in 95% of repeated experiments; this particular interval either contains it or not. In practice, with this much data and no strong prior information, a Bayesian analysis with a weak prior would give almost the same range, and then we could say 'given the data and the prior, the lift is between 0.5 and 3.0 points with 95% probability'. Our Bayesian A/B framework reports exactly that kind of credible interval. The practical message is the same either way: the lift is very likely positive, and it could be as small as half a point."

Chapter 5.9 · Syllabus Module 16

Classical tests and how to choose one

z-test, t-test, Welch, paired, chi-square, ANOVA, Mann–Whitney… The list of names looks like something to memorise. It is not. Every one of them is the same recipe from Chapter 5.6 ("is the signal big compared with the noise luck would make?") with different ingredients. This chapter shows the recipe once, then each test as a small change to it, and ends with the part that matters most: the logic of choosing a test, as an interactive flowchart.

  • See every classical test as signal ÷ noise, read against a reference curve (Normal, t, χ², F), and know why each curve is used
  • Run and explain the z-tests (one-sample, two-sample, proportions) and the t-tests (one-sample, Student, Welch, paired)
  • Explain, with a simulation, why Welch's t-test is the safe default and when Student's t-test gives far too many false alarms
  • Use the three chi-square tests (goodness-of-fit, independence, homogeneity) and say why the last two share their arithmetic but not their design
  • Use one-way ANOVA (between vs within variation, the F ratio) and say why many t-tests are a bad idea
  • Know what the rank-based tests (Mann–Whitney, Wilcoxon signed-rank, Kruskal–Wallis) really test
  • Choose a test by answering five questions, and give the reason for each answer

What we need from earlier chapters: the logic of a test: null hypothesis $H_0$, test statistic, null distribution, p-value, α (Chapter 5.6); Type I and II errors and power (Chapter 5.7); the standard error, $SE(\bar x) = \sigma/\sqrt n$ and $SE$ of a difference $=\sqrt{SE_1^2 + SE_2^2}$ (Chapter 5.5); the Normal and Student-t distributions (Chapter 4.9); the CLT (Chapter 4.13). Reminder of one word: a statistic is any number computed from the sample (a mean, a count, a rank sum). A test statistic is a statistic built so that we know how it behaves when $H_0$ is true. It has nothing to do with "statistics" the school subject.

One recipe behind every test, and its four reference curves core

A friend says her new coffee machine makes coffee faster. You time a few cups. Some are faster, some slower. To decide, you ask two questions: how big is the difference I saw? (the signal) and how big a difference would random cup-to-cup wobble make anyway? (the noise). If the signal is many times the noise, luck is a poor explanation.

Every classical test does exactly this. What changes from test to test is only the ingredients: what you measure (a mean, a share, counts in boxes, ranks), how many groups you compare, and how you estimate the noise. And because statisticians worked out long ago how "signal ÷ noise" behaves under pure luck, each test comes with a ready-made luck curve. There are only four of them: the Normal (z), Student's t, the chi-square (χ²) and the F curve.

Three ways to say it:

  • Picture: a ruler for surprise. You measure your result with it and read off how often pure luck reaches that far.
  • Numbers: checkout 60/500 vs 50/500: signal 0.02, noise 0.0198, ratio 1.01. Luck makes a ratio that big about 31% of the time.
  • Slogan: test statistic = signal ÷ noise, read against a known luck curve.

Is delivery slower than promised? A courier promises 30 minutes on average. You time 25 deliveries: the average is $\bar x = 32$ minutes and the sample standard deviation is $s = 5$ minutes.

  1. Null hypothesis $H_0$: the true average is $\mu_0 = 30$. Alternative: it is not 30.
  2. Signal: the gap between what we saw and what $H_0$ says: $\bar x - \mu_0 = 32 - 30 = 2$ minutes.
  3. Noise: how much $\bar x$ would wobble from sample to sample: $SE = s/\sqrt n = 5/\sqrt{25} = 5/5 = 1$ minute.
  4. Test statistic: signal ÷ noise $= 2/1 = 2.0$. "Our average sits 2 standard errors above the promise."
  5. Reference curve: because we estimated the noise with $s$ (we did not know $\sigma$), the luck curve is Student's t with $n - 1 = 24$ degrees of freedom.
  6. p-value: the chance that luck alone gives $|t| \ge 2.0$ on that curve: $p = 0.057$.
  7. Decision at α = 0.05: $0.057 \gt 0.05$, so we do not reject $H_0$. The data are suggestive, not convincing. (Not rejecting is not the same as proving the promise true: Chapter 5.6.)

Most classical tests have the form

$$\text{test statistic} = \frac{\text{estimate} - \text{value under } H_0}{\text{standard error of the estimate}},$$

or a sum of squares of such standardized gaps. Under $H_0$ (and the test's assumptions) the test statistic follows a known reference distribution (a null distribution with a name):

Reference curveUsed whenWhich tail
Standard Normal $N(0,1)$: zthe noise is known ($\sigma$ known), or $n$ is large (proportions, big A/B tests)both (two-sided)
Student's t with df degrees of freedomthe noise is estimated from the data ($s$ instead of $\sigma$)both
χ² (chi-square) with dfa sum of squared standardized gaps, e.g. counts vs expected countsright only
F with $(df_1, df_2)$a ratio of two variance estimates, e.g. between-group vs within-group (ANOVA)right only
  • Degrees of freedom (df): roughly, the number of independent pieces of information left for estimating the noise. With $n$ values and one estimated mean, $s$ has $n - 1$ degrees of freedom.
  • χ² and F statistics are built from squares, so they are never negative and any kind of difference makes them bigger. That is why only the right tail counts, even though they detect differences in any direction.
  • Shared assumption of all of them: the observations are independent (or independent pairs). Each test adds its own assumptions, listed in its section.
Why do we need it?

Without one recipe, the dozens of named tests feel like unrelated magic. With it, every test becomes three questions: what is my estimate, what is its standard error, and which luck curve describes the ratio? You can then rebuild a test you forgot, or check a library's output by hand.

Where is it used?

Every classical A/B test report (z for conversions, t for revenue), the chi-square check for traffic splits, ANOVA tables in experiment platforms, the t statistics next to each coefficient in a regression summary (Chapter 5.13), and the F test of a whole regression model.

How is it used?

Write down $H_0$. Compute the estimate and its standard error. Divide. Look up the tail area on the right curve (the library does this: scipy.stats returns the statistic and the p-value). Then report the estimate with a confidence interval too, not only the p-value.

Data 25 deliveries Estimate and SE x̄ = 32, SE = 1 Signal ÷ noise t = (32 − 30) / 1 = 2.0 Luck curve p = 0.057 Which luck curve? It depends on how the statistic is built: z: Normal σ known or n large t: fatter tails σ estimated by s χ²: sums of squares counts vs expected F: ratio of variances between ÷ within
Top: the same four steps happen in every classical test. Bottom: the four luck curves. z and t are symmetric (both tails count). χ² and F are built from squares, so they live on the positive side and only big values (the right tail) are surprising.

Pick a curve. The red area is the rejection region: the α share of luck-results that are most extreme. Drag the purple handle (your observed test statistic): the purple area is the p-value. For t, move df from 1 to 40 and watch the tails slim down toward the Normal: the critical value falls from 12.7 to about 2.0. Switch to χ² and F: only the right tail is shaded.

"A test statistic is a number from statistics class."

It is a number computed from your sample (like $t = 2.0$ above), designed so that we know its distribution when $H_0$ is true. That known distribution is what lets us turn it into a p-value.

"χ² and F tests are one-sided, so they only detect differences in one direction."

They use only the right tail, but they detect differences in any direction, because every kind of gap is squared and makes the statistic larger. "Right-tailed" is about the curve, not about the direction of the effect.

"z and t are different ideas."

Same idea, same formula. t replaces the unknown $\sigma$ by $s$ and uses a curve with fatter tails to pay for that extra uncertainty. With large $n$ the two curves are almost identical.

Test statistic = (estimate − value under $H_0$) / SE, or a sum of squared standardized gaps.

Luck curves: z (σ known or n large), t (σ estimated), χ² (counts vs expected), F (ratio of variances). χ² and F: right tail only.

Trap: "right-tailed" does not mean "detects only one direction".

Quick check: in the delivery example, what would change if the 25 deliveries had $s = 10$ instead of 5?

The noise doubles: $SE = 10/5 = 2$, so $t = 2/2 = 1.0$ and $p \approx 0.33$. The same 2-minute gap is now well within what luck produces. The signal alone means nothing until you divide it by the noise.

z-tests: one sample and two samples (means and proportions) core

The z-test is the simplest member of the family. It is used when the noise level is known rather than estimated. That happens in two situations:

  • You have a long, stable history that tells you $\sigma$ (rare in practice, but common in textbooks and in quality control).
  • You test proportions (conversion rates). Here the null hypothesis itself fixes the noise: if the true rate is $p$, one user's outcome has variance $p(1-p)$. With hundreds or thousands of users, the CLT makes the Normal curve an excellent luck curve.

That is why most classical A/B tests on conversion rates are z-tests.

Three ways to say it:

  • Picture: the standard Normal bell is the ruler; z says how many standard errors your result sits from the null value.
  • Numbers: $z = 1.2$ is unremarkable (p ≈ 0.23); $z = 2.3$ happens by luck only about 2% of the time.
  • Slogan: known noise (or lots of data) → z.

(a) One-sample z-test for a mean. Years of data say a shop's order values have standard deviation $\sigma = 20$ (in currency units). After a redesign, 64 orders average $\bar x = 53$. The old average was $\mu_0 = 50$.

  1. $SE = \sigma/\sqrt n = 20/\sqrt{64} = 20/8 = 2.5$.
  2. $z = (\bar x - \mu_0)/SE = (53 - 50)/2.5 = 3/2.5 = 1.2$.
  3. Two-sided $p = 2\,P(Z \ge 1.2) = 2 \times 0.115 = 0.230$. Not significant at 5%.

(b) Two-sample z-test for proportions (the checkout test from Chapter 5.6). Old checkout: 50 of 500 converted; new: 60 of 500.

  1. Rates: $\hat p_A = 50/500 = 0.10$, $\hat p_B = 60/500 = 0.12$. Signal: $0.12 - 0.10 = 0.02$.
  2. Under $H_0$ both groups share one rate, so we pool them: $\hat p = (50 + 60)/(500 + 500) = 110/1000 = 0.11$.
  3. $SE = \sqrt{\hat p(1-\hat p)\left(\tfrac{1}{500} + \tfrac{1}{500}\right)} = \sqrt{0.11 \times 0.89 \times 0.004} = \sqrt{0.0003916} = 0.0198$.
  4. $z = 0.02/0.0198 = 1.01$, two-sided $p = 0.31$.
  5. Same rates, 20 times the users (1 000 of 10 000 vs 1 100 of 10 000): pooled $\hat p = 0.105$, $SE = \sqrt{0.105 \times 0.895 \times 0.0002} = 0.00434$, $z = 0.01/0.00434 = 2.31$, $p = 0.021$. Same 2-point gap, now significant: more data shrinks the noise.

Under $H_0$, each statistic below is approximately $N(0, 1)$:

TestStatisticAssumptions
One-sample z (mean)$Z = \dfrac{\bar x - \mu_0}{\sigma/\sqrt n}$independent values; $\sigma$ known; Normal data or $n$ large
Two-sample z (means)$Z = \dfrac{\bar x_1 - \bar x_2}{\sqrt{\sigma_1^2/n_1 + \sigma_2^2/n_2}}$two independent groups; $\sigma_1, \sigma_2$ known
One-proportion z$Z = \dfrac{\hat p - p_0}{\sqrt{p_0(1-p_0)/n}}$independent yes/no outcomes; $np_0$ and $n(1-p_0)$ both at least about 10 (rule of thumb)
Two-proportion z$Z = \dfrac{\hat p_2 - \hat p_1}{\sqrt{\hat p(1-\hat p)\left(\frac{1}{n_1}+\frac{1}{n_2}\right)}}$, pooled $\hat p = \dfrac{x_1+x_2}{n_1+n_2}$independent users in two groups; at least about 10 successes and 10 failures per group

The pooled SE is used for the test because $H_0$ says the rates are equal. For a confidence interval of the difference we use the unpooled SE $\sqrt{\hat p_1(1-\hat p_1)/n_1 + \hat p_2(1-\hat p_2)/n_2}$ (Chapter 5.8). The squared two-proportion $z$ equals the chi-square statistic of the 2×2 table (you will see this in the chi-square section).

Why do we need it?

Conversion, click-through and retention are all yes/no outcomes. The two-proportion z-test is the standard classical way to ask "is the difference in rates real?", and it is the yardstick any Bayesian conversion analysis will be compared with.

Where is it used?

Classical A/B testing tools for conversion metrics, quality-control charts with a known process σ, opinion polls (one proportion vs 50%), and large-sample tests of means where $s$ is so precise that it acts like a known σ.

How is it used?

Count successes and users per group, then call statsmodels.stats.proportion.proportions_ztest([x2, x1], [n2, n1]), or compute pooled $\hat p$, SE and $z$ by hand. Report the difference with its confidence interval as well as $p$.

Start with the checkout numbers (10% vs 12%, 500 users each): the purple line (observed z = 1.01) sits well inside the bell. Slide users per group to 10 000 without changing the rates: the same 2-point gap now gives z ≈ 4.5. Then set the rates 10% vs 10.5% at 10 000 users: z ≈ 1.2 again. The readout shows every step: pooled rate, SE, z, p.

"I only have 12 values, but I will plug $s$ in for $\sigma$ and use the z-test."

With an estimated $s$ and small $n$, use the t-test (next section). Using the Normal curve there makes p-values too small: with 5 values, a test that claims a 5% false-alarm rate really has about 12%.

"My two-proportion z-test and my 2×2 chi-square test disagree, so one of them is wrong."

Without corrections they are the same test ($z^2 = \chi^2$, same p). The usual cause of disagreement: scipy.stats.chi2_contingency applies Yates' continuity correction to 2×2 tables by default (for the checkout table it gives p = 0.363 instead of 0.312). Pass correction=False to match the z-test.

In an A/B framework like yours, a conversion metric gets a Beta-Binomial model: $\theta_A \sim Beta$, $k_A \sim Binomial(n_A, \theta_A)$, and the same for B. This is the Bayesian counterpart of the two-proportion z-test. With flat Beta(1, 1) priors and large samples the two agree closely: for the checkout data, $P(\theta_B \gt \theta_A \mid D) \approx 0.843$, and $\Phi(z) = \Phi(1.01) \approx 0.844$, which is one minus the one-sided p-value. The numbers match, but the meanings differ: one is a probability about $\theta$ given the data, the other a statement about data under $H_0$ (Chapter 6.4).

One-sample z: $\dfrac{\bar x - \mu_0}{\sigma/\sqrt n}$. Two-proportion z: $\dfrac{\hat p_2 - \hat p_1}{\sqrt{\hat p(1-\hat p)(1/n_1+1/n_2)}}$ with pooled $\hat p$.

Use z when σ is known or n is large (proportions). Same gap + more users → bigger z.

Trap: 2×2 χ² = z², unless a continuity correction is switched on.

Quick check: why does the two-proportion test pool the two groups to compute the SE?

Because the p-value is computed assuming $H_0$ is true, and $H_0$ says both groups have the same rate. The best estimate of that one shared rate uses all the users: $\hat p = (x_1 + x_2)/(n_1 + n_2)$. For a confidence interval we do not assume $H_0$, so there we use each group's own rate.

One-sample t-test: when the noise must be estimated core

Usually nobody hands you $\sigma$. You estimate the noise from the same few values you are testing, with the sample standard deviation $s$. With only a handful of values, $s$ is itself a wobbly guess: sometimes it comes out too small, and then the gap looks more impressive than it really is.

William Gosset (who published as "Student") worked out the fix in 1908: keep the same signal ÷ noise ratio, but read it on a curve with fatter tails than the Normal. The fatter tails pay for the extra uncertainty in $s$. The fewer values you have, the fatter the tails; with many values the t curve becomes the Normal.

Three ways to say it:

  • Picture: the t curve is the Normal bell wearing a heavier coat in the tails.
  • Numbers: to be "surprising at 5%" you need $|t| \gt 2.57$ with 6 values, but only $|z| \gt 1.96$ with known σ.
  • Slogan: estimated noise → t curve; the less data, the more caution.

A shop averaged 50 orders a day for years. After a change, six days give 52, 55, 49, 58, 54, 56. Has the average moved?

  1. Mean: $\bar x = (52+55+49+58+54+56)/6 = 324/6 = 54$. Signal: $54 - 50 = 4$.
  2. Distances from the mean: $-2, 1, -5, 4, 0, 2$. Squares: $4, 1, 25, 16, 0, 4$. Sum $= 50$.
  3. $s^2 = 50/(6-1) = 10$, so $s = \sqrt{10} = 3.162$.
  4. $SE = s/\sqrt n = 3.162/\sqrt 6 = 3.162/2.449 = 1.291$.
  5. $t = 4/1.291 = 3.10$ with $df = 6 - 1 = 5$.
  6. On the t curve with 5 df: two-sided $p = 0.027$ (the 5% critical value is 2.571). Significant at 5%.
  7. If you had wrongly used the Normal curve, you would report $p = 0.002$: fourteen times too small. Overconfidence from pretending $s$ is exact.

For independent values $x_1, \dots, x_n$ from a Normal population with mean $\mu$, the one-sample t statistic for $H_0: \mu = \mu_0$ is

$$T = \frac{\bar x - \mu_0}{s/\sqrt n} \;\sim\; t_{n-1} \quad\text{under } H_0.$$
  • $t_{n-1}$ is Student's t distribution with $n - 1$ degrees of freedom (one degree of freedom is "used up" by estimating the mean inside $s$).
  • If the data are not Normal, $T$ is still approximately $t_{n-1}$ (and approximately Normal) for moderate or large $n$, thanks to the CLT. Strong skew or outliers with small $n$ break this.
  • The matching 95% confidence interval is $\bar x \pm t_{0.975,\,n-1}\, s/\sqrt n$ (Chapter 5.8).
Why do we need it?

Real noise levels are unknown and must be estimated from the data. Without the t correction, small-sample tests would cry wolf far more often than the α they promise.

Where is it used?

Checking a metric against a target (average delivery time vs a promise), the paired t-test (it is a one-sample t-test on differences), the t value printed next to every regression coefficient, and t-based confidence intervals in every analytics report.

How is it used?

scipy.stats.ttest_1samp(x, popmean=50) returns $t$ and $p$. Before trusting it with small $n$, look at the values (a dot plot or Q-Q plot): one huge outlier in six values matters more than any formula.

Each blue dot is one day's orders; the green dashed line is the old average $\mu_0$. Drag a day and watch the mean (purple), the standard error band and the t statistic on the curve below. Press One odd day: a single low day widens $s$ so much that t falls even though five days are high. Move $\mu_0$ to 54 (the sample mean): t becomes 0 and p becomes 1.

Here $H_0$ is true: every test draws $n$ values from a Normal population whose mean really is $\mu_0$. Each blue bar counts the t statistics that landed there. Press Run 2 000 more tests. With $n = 5$ about 12% of the t values fall beyond the Normal rule's ±1.96 (red), not 5%. The orange t curve fits the histogram; the grey Normal curve is too thin in the tails. Raise $n$ to 30: both rules nearly agree.

"The t-test needs Normal data, always."

For small $n$ it does. For larger $n$ the CLT makes $\bar x$ close to Normal and the t-test works well even for non-Normal data. The danger zone is small n plus strong skew or outliers: then use a rank test (later in this chapter) or the bootstrap (Chapter 5.5).

"The t distribution is only for small samples."

It is the right curve whenever $\sigma$ is estimated. For large $n$ it is simply indistinguishable from the Normal, so either gives the same answer.

$T = \dfrac{\bar x - \mu_0}{s/\sqrt n} \sim t_{n-1}$ under $H_0$ (Normal data, or $n$ large).

Fat tails pay for estimating σ; 5% critical value: 2.571 (df 5), 2.262 (df 9), 2.045 (df 29), → 1.96.

Trap: using 1.96 with tiny $n$ gives too many false alarms (≈ 12% at n = 5).

Quick check: with $n = 6$, $\bar x = 54$, $s = 3.162$, what is the 95% confidence interval for $\mu$, and is 50 inside it?

$54 \pm 2.571 \times 1.291 = 54 \pm 3.32$, i.e. from 50.68 to 57.32. The value 50 is just outside, which matches the test: $p = 0.027 \lt 0.05$. A test and its matching interval always agree (Chapter 5.8).

Two independent groups: Student's t vs Welch's t core

Now compare two separate groups, for example revenue per user in variant A and variant B. The signal is the difference of the two means. The noise is the standard error of that difference, and it depends on how noisy each group is.

There are two classical recipes for that noise:

  • Student's t-test assumes both groups share one noise level and averages ("pools") their variances into one.
  • Welch's t-test lets each group keep its own variance.

When the assumption is wrong, Student's pooling can go badly wrong. If the small group is the noisy one, the pooled variance is mostly the calm big group's variance, the SE comes out too small, and the test raises false alarms far more often than 5%. Welch costs almost nothing when variances really are equal, so it is the safe default.

Three ways to say it:

  • Picture: pooling is averaging a loud room and a quiet room, weighted by size; if the loud room is small, you think the whole building is quiet.
  • Numbers: with 8 noisy vs 32 calm users, Student says p = 0.013 and Welch says p = 0.20 on the same data.
  • Slogan: two groups → Welch, unless you have a strong reason not to.

Group A: $n_1 = 8$ users, mean $\bar x_1 = 12$, sd $s_1 = 4$. Group B: $n_2 = 32$ users, mean $\bar x_2 = 10$, sd $s_2 = 1$. Signal: $12 - 10 = 2$.

  1. Student (pooled). Pooled variance: $s_p^2 = \dfrac{(8-1)\cdot 16 + (32-1)\cdot 1}{8 + 32 - 2} = \dfrac{112 + 31}{38} = \dfrac{143}{38} = 3.763$. Notice it is far below A's variance of 16: the big calm group dominates.
  2. $SE = \sqrt{3.763 \times (1/8 + 1/32)} = \sqrt{3.763 \times 0.15625} = \sqrt{0.588} = 0.767$.
  3. $t = 2/0.767 = 2.61$, $df = 38$, $p = 0.013$. "Significant."
  4. Welch (separate). $SE = \sqrt{16/8 + 1/32} = \sqrt{2 + 0.031} = \sqrt{2.031} = 1.425$.
  5. $t = 2/1.425 = 1.40$. Degrees of freedom (formula below): $df = \dfrac{2.031^2}{2^2/7 + 0.031^2/31} = \dfrac{4.126}{0.571} = 7.2$.
  6. $p = 0.20$. Not significant. The honest noise is almost twice what Student assumed, because the mean of only 8 noisy users really is that uncertain.

Two independent groups, means $\bar x_1, \bar x_2$, sample variances $s_1^2, s_2^2$, sizes $n_1, n_2$. $H_0: \mu_1 = \mu_2$.

  • Student's (pooled) t: $s_p^2 = \dfrac{(n_1-1)s_1^2 + (n_2-1)s_2^2}{n_1+n_2-2}$, $\;T = \dfrac{\bar x_1 - \bar x_2}{s_p\sqrt{1/n_1 + 1/n_2}}$, $df = n_1 + n_2 - 2$. Assumes equal population variances.
  • Welch's t: $T = \dfrac{\bar x_1 - \bar x_2}{\sqrt{s_1^2/n_1 + s_2^2/n_2}}$, with the Welch–Satterthwaite degrees of freedom $df = \dfrac{(s_1^2/n_1 + s_2^2/n_2)^2}{\frac{(s_1^2/n_1)^2}{n_1-1} + \frac{(s_2^2/n_2)^2}{n_2-1}}$ (usually not a whole number). Does not assume equal variances.
  • Both assume independent observations, two independent groups, and approximately Normal sample means (Normal data, or large enough $n$).
  • With equal group sizes the two statistics are identical ($n_1 = n_2$ makes the two SE formulas equal); only the df differ, so the p-values are close.
Why do we need it?

Comparing two groups' means is the most common test in experimentation. A treatment often changes the spread as well as the mean, and A/B groups often have different sizes, which is exactly when Student's pooled test misleads.

Where is it used?

Revenue per user, session length or latency in A/B tests, comparing model errors of two systems on separate test sets, comparing two cohorts. R's t.test uses Welch by default; SciPy's ttest_ind does not.

How is it used?

Call scipy.stats.ttest_ind(a, b, equal_var=False) (the False is essential: SciPy's default is Student). Report the difference of means with a Welch confidence interval. Skip the "test variances first" step entirely.

Group A: 8 users, very spread out (s₁² = 16) Group B: 32 users, tightly packed (s₂² = 1) Variances: s₁² = 16 (A alone) s₂² = 1 (B alone) pooled s²ₚ = 3.76: what Student uses for both groups
Pooling weights each group's variance by its size. When the small group is the noisy one, the pooled variance (purple) is close to the calm group's, and Student's t underestimates how uncertain the small group's mean is.

Every run is an A/A test: both groups have the same true mean, so every "significant" result is a false alarm, and a good test should give about 5% (green line). Start with small group noisy: Student's bar shoots to about 30%, Welch stays near 5%. Press big group noisy: now Student is far below 5% (too timid, so it also loses power). With equal sizes both are close to 5%: the trouble needs unequal sizes and unequal variances together.

"First run a test for equal variances (Levene, F-test); if it passes, use Student."

This two-step procedure is not recommended: the variance pre-test has little power with small groups (exactly when it matters) and changes the error rate of the final test. Use Welch from the start.

"scipy.stats.ttest_ind(a, b) is Welch's test."

SciPy's default is equal_var=True, which is Student's pooled test. Write equal_var=False for Welch. (R's t.test defaults to Welch: the two libraries disagree.)

"Welch is only for unequal variances."

Welch is valid whether or not the variances are equal. When they are equal, it loses almost no power compared with Student.

"I used a t-test." (Which one? Why?)

"I used Welch's two-sample t-test, because the groups had different sizes and there was no reason to assume equal variances."

Model answer: "Student's test pools the two variances. If the smaller group has the larger variance, the pooled SE is too small and the false-positive rate can be several times α; if the larger group is noisier, it becomes too conservative and loses power. Welch keeps separate variances and adjusts the degrees of freedom. It costs almost nothing when variances are equal, so it is my default."

Treatments often change the spread of a metric, not just its mean (a new pricing page may create a few very large orders). In a Bayesian A/B framework like yours, giving each variant its own scale parameter in the Normal or Student-t likelihood ($\sigma_A$ and $\sigma_B$) is the Bayesian version of Welch's idea; forcing one shared $\sigma$ is the version of Student's pooling, with the same risk of overconfidence for the noisier group.

Welch: $T = \dfrac{\bar x_1 - \bar x_2}{\sqrt{s_1^2/n_1 + s_2^2/n_2}}$, Welch–Satterthwaite df. Student: pooled $s_p^2$, $df = n_1+n_2-2$.

Small group + big variance → Student's false-alarm rate explodes (≈ 30% in the 8 vs 32, σ ratio 4 case).

Trap: SciPy ttest_ind defaults to Student; pass equal_var=False.

Quick check: with $n_1 = n_2 = 20$, do Student and Welch give the same t statistic?

Yes. With equal sizes, $s_p^2(1/n + 1/n) = \frac{s_1^2 + s_2^2}{2}\cdot\frac{2}{n} = \frac{s_1^2}{n} + \frac{s_2^2}{n}$, which is Welch's squared SE. Only the degrees of freedom differ (38 vs something between 19 and 38), so the p-values are close. That is why the simulation shows little trouble with equal sizes.

Paired t-test: the same units measured twice core

Six stores run a promotion. Before, their weekly sales were 20, 35, 50, 28, 42 and 60 (hundreds of units); after, 22, 38, 51, 31, 43 and 62. Big stores sell a lot, small stores sell a little: that store-to-store difference is huge, and it has nothing to do with the promotion.

If you treat "before" and "after" as two unrelated groups, that huge store-to-store spread becomes the noise, and a small but very consistent improvement disappears in it. But every store was measured twice. Subtract each store's before from its after: the store size cancels out, and what is left is the change, store by store. Then run a one-sample t-test on those changes.

Three ways to say it:

  • Picture: lines joining each store's two values all slope gently upward; that consistency is the evidence.
  • Numbers: paired: $t = 5.48$, $p = 0.003$. Same numbers analysed as two groups: $t = 0.24$, $p = 0.82$.
  • Slogan: compare each unit with itself, and the differences between units vanish.
  1. Differences (after − before): $22-20 = 2$, $38-35 = 3$, $51-50 = 1$, $31-28 = 3$, $43-42 = 1$, $62-60 = 2$.
  2. Mean difference: $\bar d = (2+3+1+3+1+2)/6 = 12/6 = 2$.
  3. Distances from 2: $0, 1, -1, 1, -1, 0$. Squares sum to 4. $s_d^2 = 4/5 = 0.8$, $s_d = 0.894$.
  4. $SE = 0.894/\sqrt 6 = 0.365$. $t = 2/0.365 = 5.48$ with $df = 5$. $p = 0.003$.
  5. Wrong analysis (two independent groups, Welch): the means differ by 2 as well, but the store-to-store sd is about 14.6 and 14.2, so $SE \approx 8.3$, $t = 0.24$, $p = 0.82$. Same data, opposite conclusion.

Paired data: $n$ units, each measured under two conditions, $(x_i, y_i)$. Form the differences $d_i = y_i - x_i$ and test $H_0: \mu_d = 0$ with a one-sample t-test:

$$T = \frac{\bar d}{s_d/\sqrt n} \;\sim\; t_{n-1} \text{ under } H_0.$$
  • Why it helps: $Var(\bar d) = \dfrac{\sigma_x^2 + \sigma_y^2 - 2\,Cov(x, y)}{n}$. When the two measurements of the same unit are positively correlated (big stores stay big), the covariance term removes most of the noise.
  • Assumptions: the pairs are independent of each other; the differences are roughly Normal (or $n$ is large).
  • Pairing is a property of the design (how the data were collected), not something you choose after looking at the data.
Why do we need it?

Before/after studies and within-subject comparisons have huge differences between units. Ignoring the pairing wastes most of the information and can turn a clear effect into "no evidence".

Where is it used?

Before/after metrics for the same stores or users, comparing two ML models on the same test examples, two forecasting models' errors on the same days, A/A checks of the same users across two weeks, and crossover designs.

How is it used?

scipy.stats.ttest_rel(after, before), which is exactly ttest_1samp(after - before, 0). Plot the pairs as joined lines first. If the differences are very skewed with small $n$, use the Wilcoxon signed-rank test instead.

Independent groups: different units Paired: the same units twice group A group B no links: noise = all the unit-to-unit spread before after every line goes up a little
Left: two groups of different units; the only way to judge the gap is against the full unit-to-unit spread. Right: each unit measured twice; the lines show each unit's own change, and the unit-to-unit spread (how high a line sits) no longer matters.

Blue = before, orange = after; grey lines join each store's pair. The lower plot shows only the six differences. Drag any orange dot up or down and watch both readouts. Press Stores of similar size: when units hardly differ from each other, pairing helps much less (the two tests get closer). Press Inconsistent changes: the differences scatter around 0, and even the paired test finds nothing.

"Before/after data from the same users is two groups, so I use the two-sample t-test."

Paired data analysed as independent throws away the pairing and usually the power with it. Use the paired test whenever each unit appears in both conditions.

"My A/B test has equal group sizes, so I can pair user 1 of A with user 1 of B."

No. In a standard A/B test each user is in only one arm; any pairing you invent is arbitrary and the test becomes meaningless. Pairing must come from the design (same unit twice, or units matched before treatment).

Paired: $d_i = y_i - x_i$, $T = \bar d/(s_d/\sqrt n) \sim t_{n-1}$. It is a one-sample t-test on the differences.

Helps when the two measurements of a unit are positively correlated: $Var(\bar d) = (\sigma_x^2 + \sigma_y^2 - 2Cov)/n$.

Trap: pairing comes from the design, never from sorting or matching rows afterwards.

Quick check: you compare two recommendation models on the same 1 000 test users (each user gets a score from both). Paired or independent?

Paired: every user is scored by both models, so compute each user's score difference and test its mean (ttest_rel). Users differ a lot from each other (some are easy, some hard), and pairing removes that.

Chi-square goodness-of-fit: do the counts match the expected shares? core

Now the data are not numbers but categories: which plan each new customer chose (Basic, Pro, Team). You have an expectation for the shares, for example last year's 50% / 30% / 20%. This month's 200 sign-ups give 88, 72 and 40. Has the mix changed, or is this ordinary month-to-month wobble?

Turn the expected shares into expected counts (100, 60, 40), then measure the gap in each category. A gap of 12 matters more when only 60 were expected than when 100 were expected, so each squared gap is divided by its expected count. Add these up: that is the chi-square statistic. Zero means a perfect match; the bigger it is, the worse the fit.

Three ways to say it:

  • Picture: two sets of bars, observed and expected; χ² adds up how badly each pair of bars disagrees, relative to its size.
  • Numbers: gaps of 12, 12 and 0 give $\chi^2 = 1.44 + 2.40 + 0 = 3.84$; luck does at least this 15% of the time.
  • Slogan: observed vs expected counts, squared, scaled, summed.

200 sign-ups: Basic 88, Pro 72, Team 40. $H_0$: the shares are still 50% / 30% / 20%.

  1. Expected counts: $200 \times 0.5 = 100$, $200 \times 0.3 = 60$, $200 \times 0.2 = 40$.
  2. Gaps $O - E$: $88 - 100 = -12$, $72 - 60 = +12$, $40 - 40 = 0$.
  3. Scaled squares $(O-E)^2/E$: $144/100 = 1.44$, $144/60 = 2.40$, $0/40 = 0$.
  4. $\chi^2 = 1.44 + 2.40 + 0 = 3.84$.
  5. Degrees of freedom: 3 categories − 1 = 2 (once two counts are known, the third is fixed by the total of 200).
  6. $p = P(\chi^2_2 \ge 3.84) = 0.147$ (for 2 df this equals $e^{-3.84/2} = e^{-1.92}$). The 5% critical value is 5.99, so we do not reject: the mix may not have changed.

Observed counts $O_1, \dots, O_K$ in $K$ categories, total $N$. $H_0$: the category probabilities are $\pi_1, \dots, \pi_K$ (given in advance). Expected counts $E_k = N\pi_k$.

$$\chi^2 = \sum_{k=1}^{K} \frac{(O_k - E_k)^2}{E_k} \;\approx\; \chi^2_{K-1-m} \text{ under } H_0,$$

where $m$ is the number of parameters you estimated from the same data to get the $\pi_k$ (0 when the shares are fixed in advance; 1 if, say, you fitted a Poisson mean first).

  • Assumptions: independent observations, each falls in exactly one category, and the counts are counts (not percentages or averages).
  • The χ² approximation needs large expected counts. Rule of thumb: every $E_k \ge 5$. Otherwise merge categories or use an exact (multinomial) test.
  • Right tail only; a large $\chi^2$ says "the shares differ somewhere", not where. Look at the individual terms to see which categories drive it.
Why do we need it?

Many checks are about shares of categories: is traffic split as designed, did the mix of plans change, does a count model predict the right number of zeros, ones, twos? Means and t-tests cannot answer these.

Where is it used?

The sample ratio mismatch check in A/B testing (is a 50/50 split really 50/50? Chapter 5.11), checking a die or a random number generator, comparing a fitted Poisson or Negative Binomial to observed count frequencies, fraud checks with Benford's law.

How is it used?

scipy.stats.chisquare(f_obs=[88, 72, 40], f_exp=[100, 60, 40]) gives $\chi^2$ and $p$ (pass ddof=m if you estimated parameters). Check the expected counts first, then read the individual $(O-E)^2/E$ terms.

Blue bars are the observed sign-ups; green outlines are the counts expected under $H_0$. Drag the top of a blue bar. The red gaps feed the χ² sum, shown on the curve below (2 degrees of freedom). Notice that a gap of 12 adds more when the expected count is smaller (Pro) than when it is larger (Basic). Switch the expected shares to equal thirds and watch χ² jump.

"I will run the chi-square test on percentages (44%, 36%, 20%)."

χ² must use counts. The same percentages from 200 people and from 20 000 people carry very different evidence, and only counts know the difference.

"p = 0.15, so the data follow last year's shares."

Not rejecting is not proof of a match: with 200 sign-ups, moderate changes are hard to detect. Absence of evidence is not evidence of absence (Chapter 5.6).

$\chi^2 = \sum_k (O_k - E_k)^2/E_k$, $E_k = N\pi_k$, df $= K - 1 - (\text{estimated parameters})$, right tail.

Counts only; every $E_k \ge 5$ (rule of thumb), else merge or use an exact test.

Trap: percentages instead of counts; reading "not rejected" as "fits".

Quick check: a 50/50 A/B split gives 10 080 vs 9 920 users. Compute χ².

$N = 20\,000$, $E = 10\,000$ each. $\chi^2 = 80^2/10\,000 + 80^2/10\,000 = 0.64 + 0.64 = 1.28$, df 1, $p = 0.26$. No sign of a sample ratio mismatch. (With 11 000 vs 9 000, $\chi^2 = 100 + 100 = 200$: a broken split. More in Chapter 5.11.)

Chi-square tests of independence and homogeneity: one arithmetic, two designs core

Now there are two categorical variables, arranged in a table (a contingency table: rows for one variable, columns for the other, a count in each cell). Example: which variant a user saw (A or B) and which plan they chose (Basic, Pro, Team).

If the variant has nothing to do with the plan, every row should show the same mix as the table as a whole. In the overall total, 52.5% chose Basic; so in a row of 100 users we would expect 52.5 Basic users. That gives an expected count for every cell, and then the same "observed vs expected, squared, scaled, summed" as before.

The same calculation answers two differently worded questions, depending on how the data were collected. That is the only difference between the "independence" and the "homogeneity" test.

Three ways to say it:

  • Picture: if the rows are just copies of the overall mix, the table is "boring"; χ² measures how un-boring it is.
  • Numbers: expected cell = row total × column total ÷ grand total, e.g. $100 \times 105/200 = 52.5$.
  • Slogan: same formula; independence = one sample cross-classified, homogeneity = several groups fixed in advance.

100 users per variant chose a plan:

BasicProTeamRow total
A603010100
B454015100
Column total1057025200
  1. Expected counts: $E_{ij} = (\text{row total}_i \times \text{column total}_j)/N$. For A–Basic: $100 \times 105/200 = 52.5$. Both rows get $52.5, 35, 12.5$.
  2. Gaps: row A: $+7.5, -5, -2.5$; row B: $-7.5, +5, +2.5$.
  3. Terms $(O-E)^2/E$: $56.25/52.5 = 1.071$, $25/35 = 0.714$, $6.25/12.5 = 0.5$, the same three again for row B.
  4. $\chi^2 = 2 \times (1.071 + 0.714 + 0.5) = 4.571$.
  5. df $= (\text{rows} - 1)(\text{columns} - 1) = 1 \times 2 = 2$. $p = e^{-4.571/2} = 0.102$. Not significant at 5%.
  6. The 2×2 link. The checkout table (A: 50 converted, 450 not; B: 60 and 440) has expected counts 55 and 445 in each row, $\chi^2 = 2(25/55 + 25/445) = 1.0215$, df 1, $p = 0.312$. And the two-proportion $z$ was 1.0107: $1.0107^2 = 1.0215$. Same test.

An $r \times c$ table of counts $O_{ij}$, row totals $R_i$, column totals $C_j$, grand total $N$. Expected counts under $H_0$: $E_{ij} = R_i C_j / N$.

$$\chi^2 = \sum_{i=1}^{r}\sum_{j=1}^{c} \frac{(O_{ij} - E_{ij})^2}{E_{ij}} \;\approx\; \chi^2_{(r-1)(c-1)} \text{ under } H_0.$$
  • Test of independence: one sample of $N$ units; both variables are recorded for each unit (e.g. device type and plan for 200 random sign-ups). $H_0$: the two variables are independent, $P(\text{row } i, \text{col } j) = P(\text{row } i)\,P(\text{col } j)$.
  • Test of homogeneity: several groups whose sizes are fixed by the design (e.g. 100 users assigned to A, 100 to B); one variable is recorded. $H_0$: the distribution of that variable is the same in every group.
  • Same statistic, same df, same p-value; only the sampling design and the wording of $H_0$ differ. An A/B test with a categorical metric is, strictly, a homogeneity test, because the arm sizes are set by the experimenter.
  • Assumptions: independent units, counts not percentages, expected counts ≥ 5 (rule of thumb); otherwise use Fisher's exact test (2×2) or a simulated/exact p-value.
  • Effect size: Cramér's $V = \sqrt{\chi^2 / (N(\min(r,c)-1))}$, between 0 (no association) and 1. Here $V = \sqrt{4.571/200} = 0.15$.
Why do we need it?

Many A/B metrics are categorical with more than two outcomes (plan chosen, cancellation reason, star rating as categories). The table test asks whether the whole distribution differs between arms, in one test, instead of many separate comparisons.

Where is it used?

Categorical A/B metrics, checking whether a segment variable (country, device) is balanced across arms, feature-vs-label association checks in feature selection (sklearn.feature_selection.chi2 uses a related statistic), survey cross-tabulations.

How is it used?

scipy.stats.chi2_contingency(table, correction=False) returns $\chi^2$, p, df and the expected table. Check the expected counts, read the cell contributions (or standardized residuals) to see where the difference is, and report Cramér's V.

Independence Homogeneity ONE random sample of 200 sign-ups record device AND plan for each device × plan table row totals are random H₀: device and plan are independent assign 100 to A record plan assign 100 to B record plan variant × plan table row totals fixed by design H₀: plan mix is the same in A and B
The two tests compute exactly the same χ² from the same kind of table. What differs is how the table was produced: one sample cross-classified by two variables (independence), or groups of fixed size compared on one variable (homogeneity, which is what an A/B test is).

Edit the observed counts (rows A and B; columns Basic, Pro, Team). Each cell shows observed / expected; its red shade is its share of χ². Make row B a copy of row A (60, 30, 10): every expected count equals the observed one and χ² = 0. Now move 10 users in row B from Basic to Team: one small expected count (Team) makes that cell dominate. Try a cell of 2 to see the small-count warning.

"Independence and homogeneity are different tests, so I need to pick the right formula."

The formula, df and p-value are identical. Pick the right wording of $H_0$ for your design; the arithmetic does not change.

"χ² is significant, so Pro is the plan that changed."

A significant χ² only says the table is not "boring" somewhere. To see where, look at the cell contributions or standardized residuals, and treat that as exploration (several looks → multiple testing, Chapter 5.11).

"χ² = 40 with p < 0.001: a strong association."

χ² grows with $N$. With a million users, a tiny difference gives a huge χ². Report an effect size (Cramér's V, or the difference in shares with a CI) next to the p-value.

In an A/B framework like yours, a categorical metric (which plan, which reason) is modelled with a Dirichlet-Multinomial: each variant gets a probability vector $\boldsymbol\theta_A, \boldsymbol\theta_B \sim Dirichlet(\boldsymbol\alpha)$ and its counts are Multinomial. That is the Bayesian counterpart of this homogeneity test. Instead of one p-value for "the table is not boring", it gives a posterior for every share and every difference, e.g. $P(\theta_{B,\text{Pro}} \gt \theta_{A,\text{Pro}} \mid D)$ (Chapter 6.3).

$E_{ij} = R_iC_j/N$, $\chi^2 = \sum (O-E)^2/E$, df $= (r-1)(c-1)$.

Independence (one sample, two variables) and homogeneity (fixed groups, one variable): same arithmetic, different design. A/B tests = homogeneity. 2×2: $\chi^2 = z^2$.

Traps: Yates correction by default in SciPy for 2×2; small expected counts → Fisher; big N makes tiny effects significant.

Quick check: a 3×4 table. How many degrees of freedom?

$(3-1)(4-1) = 6$. Once the row and column totals are fixed, only 6 cells can be chosen freely; the rest are forced by the totals.

One-way ANOVA: comparing three or more means at once core

Three page designs, each shown to a few users, give average basket sizes of 5, 7 and 9. Are the designs really different? Two kinds of spread are in play:

  • Between groups: how far the group averages sit from each other (5, 7, 9).
  • Within groups: how much individual users scatter around their own group's average.

If the designs did nothing, the group averages would differ only by the small amount that within-group scatter produces by luck. So we compare the two: F = between-group variation ÷ within-group variation. Around 1 means "the averages differ about as much as luck predicts"; much bigger than 1 means "more than luck".

Why not just run a t-test for every pair? Because each test has its own 5% false-alarm chance. Three groups need 3 tests, five groups need 10; with 5 groups and no real difference, about 29% of experiments would show at least one "significant" pair. ANOVA asks one question with one 5% risk.

Three ways to say it:

  • Picture: are the group centres far apart compared with how wide each group is?
  • Numbers: between-group mean square 16, within-group mean square 2.67, ratio $F = 6.0$.
  • Slogan: ANOVA uses variances to compare means: signal (between) over noise (within).

Basket sizes for three designs, 4 users each: A: 3, 5, 7, 5 · B: 5, 9, 7, 7 · C: 7, 11, 9, 9.

  1. Group means: A $= 20/4 = 5$, B $= 28/4 = 7$, C $= 36/4 = 9$. Grand mean (all 12 values): $84/12 = 7$.
  2. Between sum of squares: each group mean's squared distance from 7, times its group size: $SSB = 4(5-7)^2 + 4(7-7)^2 + 4(9-7)^2 = 16 + 0 + 16 = 32$.
  3. Within sum of squares: each value's squared distance from its own group mean. A: $(-2)^2 + 0 + 2^2 + 0 = 8$; B: $(-2)^2 + 2^2 + 0 + 0 = 8$; C: $8$. $SSW = 24$.
  4. Degrees of freedom: between $k - 1 = 3 - 1 = 2$; within $N - k = 12 - 3 = 9$.
  5. Mean squares: $MSB = 32/2 = 16$, $MSW = 24/9 = 2.667$.
  6. $F = 16/2.667 = 6.0$ with $(2, 9)$ df. $p = 0.022$ (the 5% critical value is 4.26). At least one design differs.

$k$ groups, group $g$ has $n_g$ values with mean $\bar y_g$; grand mean $\bar y$; $N = \sum n_g$. $H_0$: all $k$ population means are equal. $H_1$: at least one differs.

$$SSB = \sum_g n_g(\bar y_g - \bar y)^2,\quad SSW = \sum_g\sum_i (y_{gi} - \bar y_g)^2,\quad F = \frac{SSB/(k-1)}{SSW/(N-k)} \sim F_{k-1,\,N-k} \text{ under } H_0.$$
  • The total spread splits exactly: $SST = \sum (y_{gi} - \bar y)^2 = SSB + SSW$ (here $56 = 32 + 24$).
  • Assumptions: independent observations; roughly Normal values within each group (less important for large groups); equal variances across groups. If spreads differ a lot, use Welch's ANOVA, the same idea as Welch's t-test.
  • With $k = 2$, ANOVA's $F$ equals Student's $t^2$.
  • A significant $F$ says "not all equal", not which groups differ. Follow it with post-hoc comparisons that control the error rate (e.g. Tukey's HSD) or with comparisons planned in advance and corrected (Chapter 5.11).
Why do we need it?

As soon as there are more than two groups (A/B/C/D variants, regions, segments), pairwise t-tests multiply the false-alarm risk. ANOVA gives one honest test of "is there any difference at all?", and its between/within split is the language of hierarchical models.

Where is it used?

Multi-variant experiments, comparing a metric across regions or device types, the F-test of a whole regression model, analysis of designed experiments in manufacturing and agriculture, and the variance decomposition behind random-effects models.

How is it used?

scipy.stats.f_oneway(a, b, c) returns $F$ and $p$; statsmodels prints a full ANOVA table and offers Tukey's HSD (pairwise_tukeyhsd). Plot the groups first; if the spreads differ clearly, use Welch's ANOVA or a rank-based Kruskal–Wallis test.

grand mean 7 A (mean 5) B (mean 7) C (mean 9) orange bars: group mean vs grand mean (BETWEEN) red sticks: value vs its group mean (WITHIN)
The example data. ANOVA compares the orange gaps (how far the group means are from the grand mean) with the red gaps (how far values are from their own group mean). F is the ratio of the two, each averaged per degree of freedom.

Twelve users in three groups. Drag a dot up or down. Orange segments are the group means; the purple dashed line is the grand mean; red sticks are the within-group distances. Press Pull groups apart: the between part grows and F climbs. Press Same means, wide spread: the group means are equal, so the between part is 0 and F = 0, however spread out the values are.

Each simulated experiment has $k$ groups of 10 values with identical true means. We either run a t-test for every pair (red: "at least one pair significant") or one ANOVA (green). Press Run 250 more per k a few times. With 3 groups about 12% of experiments raise a false alarm with pairwise tests; with 8 groups (28 pairs) about half do. ANOVA stays near α. The grey dashed curve is $1 - (1-\alpha)^m$, the rate if the $m$ pairwise tests were independent (they are not, so the truth is a bit lower).

"ANOVA tests whether the variances differ."

ANOVA tests whether the means differ. It uses two variance estimates (between and within) as its tool; the name "analysis of variance" describes the method, not the hypothesis.

"ANOVA was significant, so group C is the best."

A significant F only says the means are not all equal. Which groups differ needs post-hoc comparisons with error control (Tukey's HSD) or pre-planned, corrected contrasts.

"With four variants I will just do all six pairwise t-tests at 5%."

Six tests at 5% give roughly a 20% chance of at least one false alarm when nothing works. Use ANOVA first, or correct the pairwise tests (Bonferroni, Holm: Chapter 5.11).

Between-group vs within-group variation is exactly the split a hierarchical model makes. In a framework like yours with partial pooling across segments, $\tau$ (how much the segment effects differ from each other) plays the role of the between-group spread, and $\sigma$ (noise inside a segment) the within-group spread (Chapter 4.6). Where ANOVA gives one yes/no answer to "do the segments differ?", the hierarchical model estimates how much they differ and shrinks noisy segment estimates toward the overall mean (Chapter 6.6).

$F = \dfrac{SSB/(k-1)}{SSW/(N-k)}$; $SST = SSB + SSW$; $F \approx 1$ under $H_0$; right tail of $F_{k-1, N-k}$.

Assumes independence, roughly Normal groups, equal variances (else Welch ANOVA or Kruskal–Wallis). $k = 2$: $F = t^2$.

Traps: significant F does not say which group; many pairwise t-tests inflate false alarms (≈ 12% at k = 3, ≈ 29% at k = 5).

Quick check: four groups of 10 users each. What are the degrees of freedom of F?

Between: $k - 1 = 3$. Within: $N - k = 40 - 4 = 36$. So compare F with the $F_{3,36}$ curve (its 5% critical value is about 2.87).

Nonparametric alternatives: Mann–Whitney, Wilcoxon signed-rank, Kruskal–Wallis

The t-tests and ANOVA compare means, and with small samples they lean on the Normal shape. One huge order of 5 000 among orders of 20 to 60 can dominate a mean and a standard deviation.

Rank tests take a different view: line all the values up from smallest to largest and replace each value by its position (its rank: 1st, 2nd, 3rd…). The giant order is simply "the largest", rank 10, whether it is 70 or 7 000. Then ask: do one group's ranks tend to be higher than the other's?

Such tests are called nonparametric: they do not assume the data come from a particular family like the Normal. They are robust to outliers and skew. The price: they answer a slightly different question ("does one group tend to produce larger values?") rather than "are the means different?".

Three ways to say it:

  • Picture: forget the ruler, keep only the queue order.
  • Numbers: of the 12 pairs (one value from A, one from B), A wins only 1. That is U = 1.
  • Slogan: ranks ignore how far an outlier is; they only see that it is last in line.

Mann–Whitney by hand. Group A: 2, 4, 7. Group B: 5, 8, 9, 12.

  1. Sort all 7 values and rank them: 2 (A) → 1, 4 (A) → 2, 5 (B) → 3, 7 (A) → 4, 8 (B) → 5, 9 (B) → 6, 12 (B) → 7.
  2. Rank sum of A: $R_A = 1 + 2 + 4 = 7$.
  3. $U_A = R_A - n_A(n_A+1)/2 = 7 - 3 \cdot 4/2 = 7 - 6 = 1$. Meaning: among the $3 \times 4 = 12$ (A, B) pairs, A's value is larger in exactly 1 pair (7 vs 5). $U_B = 12 - 1 = 11$.
  4. Under $H_0$ every ordering of the 7 values is equally likely: there are $\binom{7}{3} = 35$ ways to place A's ranks, and only 2 of them give $U_A \le 1$. Two-sided exact $p = 2 \times 2/35 = 0.114$.
  5. Change B's 12 into 120. The ranks do not change, so U and p do not change. (The Welch t-test's p-value moves from 0.10 to 0.35, because the outlier inflates B's standard deviation.)
TestReplacesStatisticWhat $H_0$ really says
Mann–Whitney U (Wilcoxon rank-sum)two-sample t (independent groups)$U$ = number of (a, b) pairs with $a \gt b$ (ties count ½)$P(X \gt Y) = P(Y \gt X)$: neither group tends to give larger values. A statement about medians only if both groups have the same shape (one is a shifted copy of the other).
Wilcoxon signed-rankpaired t / one-sample trank the $|d_i|$; add the ranks of the positive differencesthe differences are symmetric around 0 (so their median is 0)
Kruskal–Wallis Hone-way ANOVAcompares the average rank of each group; $H \approx \chi^2_{k-1}$all $k$ groups come from the same distribution
  • Assumptions that remain: independent observations (independent pairs for signed-rank), and a continuous or at least ordered outcome. "Nonparametric" does not mean "no assumptions".
  • Efficiency: for Normal data, Mann–Whitney needs about 5% more data than the t-test for the same power (asymptotic relative efficiency $3/\pi \approx 0.955$). For shift alternatives it never needs more than about 16% extra (efficiency ≥ 0.864), and for heavy tails or skew it can need far less.
  • They also work for ordinal data (star ratings, "bad / ok / good"), where a mean is questionable.
Why do we need it?

Small samples with outliers or heavy skew break the t-test's Normal approximation, and ordinal outcomes have no meaningful mean. Rank tests stay valid in both cases, with almost no loss when the data happen to be Normal.

Where is it used?

Latency and load-time comparisons (heavy right tails), star ratings and satisfaction scales, small pilot studies, comparing per-example errors of two models (signed-rank on paired errors), medical trials with skewed outcomes.

How is it used?

scipy.stats.mannwhitneyu(a, b), wilcoxon(after - before), kruskal(a, b, c). State the hypothesis honestly ("B tends to give larger values"), and if you need a size of effect, report the probability $P(B \gt A) = U_B/(n_An_B)$ or the Hodges–Lehmann shift.

Values 12 → 120 Ranks 1 2 3 4 5 6 7 rank 7 whether the value is 12 or 120 blue = group A (2, 4, 7), orange = group B (5, 8, 9, 12)
Top: the raw values; moving B's largest value from 12 to 120 stretches the scale and changes the mean and sd a lot. Bottom: the ranks are evenly spaced and do not change at all. Rank tests only use the bottom line.

Top: the values of group A (blue) and B (orange); drag any dot. Bottom: what a rank test sees (positions 1 to 10). Press Make B's top value huge: B looks even better, yet the Welch t-test's p-value gets worse (the outlier inflates B's sd more than its mean), while the Mann–Whitney result does not change. Then drag A's lowest dot past several B dots and watch U change by one for each B value it passes.

B is A shifted up by the slider amount. Each press runs 200 simulated experiments with $n$ users per group and counts how often each test says "significant" (its power). With Normal data, Welch t wins slightly. Switch to heavy tails or skewed: Mann–Whitney wins clearly (around 0.48 vs 0.37 and 0.71 vs 0.32 at shift 0.8, n = 20). Set the shift to 0: both bars should sit near 5% (the false-alarm rate).

"Mann–Whitney is a test of medians."

It tests whether one group tends to produce larger values ($P(X \gt Y) \ne \tfrac12$). Two groups can have the same median and still differ in this sense, and vice versa. It becomes a test of medians (or of a shift) only if you assume both distributions have the same shape.

"Nonparametric tests have no assumptions, so they are always safer."

They still need independent observations, and they answer a different question than a test of means. If your business cares about the mean (total revenue = users × mean revenue), a rank test can say "B tends to be higher" while the mean is lower (a few huge spenders in A).

"Revenue is skewed, so I used Mann–Whitney to compare the average revenue per user."

"Revenue is skewed. If the decision is about total revenue, I compare means (Welch's t, or a bootstrap CI, which the CLT supports at A/B-test sample sizes, though heavy tails need more data). If the question is whether a typical user spends more, Mann–Whitney answers that, and I say so."

Model answer: "Mann–Whitney tests whether one group's values tend to be larger, $P(X \gt Y)$, not whether the means differ. It is robust and powerful for skewed data, but it can disagree with a comparison of means. I choose the test from the business question first, then check that its assumptions hold."

Both of your projects offer a Student-t likelihood for continuous metrics with outliers. That is the model-based cousin of the rank tests' robustness: extreme values get less pull on the estimated centre. The same caution applies: with skewed data, the location of a Student-t fit follows the bulk of the data and is not the mean, just as Mann–Whitney is not a test of means. Decide first whether the question is about the mean (totals) or about a typical value.

Replace values by ranks. Mann–Whitney (2 independent groups), Wilcoxon signed-rank (paired / one sample), Kruskal–Wallis (k groups).

MW: $U$ = pairs where A > B; $H_0$: $P(X\gt Y) = \tfrac12$. Efficiency vs t under Normal ≈ 0.955; can be much better with heavy tails.

Trap: not a test of medians or means unless the shapes are the same.

Quick check: group A = 1, 2, 3 and group B = 4, 5, 6. What is $U_A$, and what is the smallest possible two-sided exact p-value for these sizes?

A never beats B, so $U_A = 0$. Of the $\binom{6}{3} = 20$ equally likely orderings, only 1 gives $U_A = 0$ (and 1 gives $U_A = 9$), so the two-sided exact p is $2/20 = 0.10$. With 3 per group, a rank test can never reach $p \lt 0.05$: tiny samples cannot give strong evidence, whatever the test.

Choosing a test: the logic, as a flowchart core

You do not need to memorise a list of tests. You need to ask the data a few questions, in order, and each answer has a reason:

  1. What is measured on each unit? A number (revenue, time) can be averaged. An ordered rating (1–5 stars) has an order but uneven gaps. A category (converted or not, plan chosen) can only be counted. This decides the family: means, ranks or counts.
  2. How many groups, and against what? One group against a target, two groups, or three and more. More groups → one overall test first.
  3. Paired or independent? Same units twice → work with differences. Different units → compare two groups.
  4. Can I trust the mean and the Normal approximation? Large $n$ or roughly Normal data → t, z, ANOVA. Small $n$ with strong skew or outliers → ranks (or the bootstrap). For counts the matching question is: are the expected counts big enough?
  5. Is the noise level known? Almost never → t. Known σ → z. (For proportions the null hypothesis fixes the noise, so z is natural.)

Three ways to say it:

  • Picture: a short path through a tree; each fork is a fact about your data, never about the result you hope for.
  • Numbers: conversions, 2 groups, 500 users each, expected counts ≥ 5 → two-proportion z-test.
  • Slogan: data type → groups → design → shape → noise.

Three walks through the questions.

  1. Old vs new checkout, 500 users each, "did they buy?" A category (yes/no) → two groups → independent users → expected counts 55 and 445 per group, all ≥ 5 → two-proportion z-test (identical to the 2×2 χ² test). With only 20 users per group and 2 buyers, expected counts would be tiny → Fisher's exact test.
  2. Load time of the same 12 devices before and after a release, with a few huge values. A number → two conditions → the same devices (paired) → only 12 differences, strongly skewed → Wilcoxon signed-rank test on the differences. (With 500 devices, a paired t-test on the differences would be fine.)
  3. Average basket size in 4 regions, about 200 orders each. A number → 4 groups → independent orders → 200 per group, so the group means are close to Normal by the CLT → one-way ANOVA (Welch's ANOVA if the regions' spreads differ a lot), followed by Tukey comparisons if it is significant.

The full decision table (the figure below draws it):

OutcomeGroups / designMean OK (roughly Normal or $n$ large)Small $n$, strong skew or outliers; or ordered ratings
a numberone group vs a targetone-sample t (z if σ known)Wilcoxon signed-rank
two groups, pairedpaired tWilcoxon signed-rank on the differences
two groups, independentWelch t (z if σ known)Mann–Whitney U
3+ groups, independentone-way ANOVA (Welch ANOVA if spreads differ)Kruskal–Wallis
OutcomeQuestionAll expected counts ≥ 5Some expected counts < 5
a categoryone share vs a targetone-proportion zexact binomial test
K shares vs expected sharesχ² goodness-of-fitexact / simulated p, or merge categories
yes/no in two groupstwo-proportion z (= 2×2 χ²)Fisher's exact test
two variables (r × c table)χ² independence / homogeneityFisher (exact) or simulated p

Rules that sit on top of the table: units must be independent (or correctly paired); the test, the metric and α are chosen before seeing the data; and every test result is reported with an effect size and a confidence interval. "Small $n$" has no sharp cut-off: with strong skew, even a few hundred values can be small for a t-test on the mean.

Why do we need it?

The wrong test gives wrong error rates: Student's t with unequal groups, an independent test on paired data, a t-test on 8 skewed values, or a χ² test with tiny expected counts. Choosing by logic, not by habit or by the smallest p-value, is what makes a result trustworthy.

Where is it used?

Every analysis plan for an A/B test, design reviews ("why this test?"), interview questions about experiments, and the choice between a classical test and a Bayesian model with a matching likelihood in a framework like yours.

How is it used?

Before the experiment, write down the metric type, the design and the planned test in the analysis plan. After collecting the data, check the assumptions you can check (independence, expected counts, a dot plot or Q-Q plot), run the planned test, and report the effect with its CI.

1. What is measured on each unit? IF A NUMBER (an ordered rating → always the right column) 2–3. groups and design 4. mean OK (Normal-ish or n large) 4. small n + skew / outliers one group vs a targetone-sample t(z if σ known)Wilcoxon signed-rank two groups, pairedpaired tsigned-rank on differences two groups, independentWelch t(z if σ known)Mann–Whitney U 3+ groups, independentANOVA(Welch ANOVA)Kruskal–Wallis IF A CATEGORY (converted?, plan chosen) 2. what is the question? 3. expected counts all ≥ 5 3. some expected counts < 5 one share vs a targetone-proportion zexact binomial K shares vs expectedχ² goodness-of-fitexact / simulated p yes/no in two groupstwo-proportion z (= 2×2 χ²)Fisher's exact two variables (r × c table)χ² indep. / homogeneityFisher / simulated p
The whole map on one page. First the kind of outcome (number or category), then the groups and design, then whether the approximation behind the standard test can be trusted (shape and size for numbers, expected counts for categories). Ordered ratings always go to the right-hand (rank) column.

Pick a scenario, or choose "my own data" and answer the questions yourself. Only the questions that matter for your path are shown. Each step of the path says why that answer leads where it does; the black box at the end gives the test, its null hypothesis, the Python call and what to check. Try changing one answer at a time: for example switch "Revenue per user" from mean OK to small n + skew and see Welch become Mann–Whitney.

"First run a normality test (Shapiro–Wilk); if p > 0.05 use the t-test, otherwise Mann–Whitney."

With large $n$, normality tests reject tiny, harmless departures (and the t-test is fine anyway thanks to the CLT); with small $n$ they miss big ones. Choose from what you know about the metric, a plot (Q-Q plot, Chapter 4.17) and the sample size, and above all from the question (mean or typical value?).

"I ran three tests and reported the one with the smallest p-value."

That is p-hacking: trying tests until one "works" raises the false-alarm rate above α. The test belongs in the analysis plan, written before the data arrive.

"The test choice depends on whether the result is significant."

Every fork in the flowchart is a fact about the data and the design (type, groups, pairing, size, counts), never about the outcome.

Your Bayesian A/B framework follows the same first question, in a different language: the kind of outcome picks the likelihood instead of the test. Yes/no conversions → Beta-Binomial (classical twin: two-proportion z-test); a categorical metric → Dirichlet-Multinomial (χ² homogeneity); a continuous metric → Normal or Student-t (Welch's t, or a rank test when outliers dominate); counts → Poisson (a rate comparison). Interviewers often ask "what classical test would you compare this with?": this mapping is the answer, followed by "and the Bayesian version gives $P(\theta_B \gt \theta_A \mid D)$ and the full distribution of the lift instead of a p-value".

"I would use a t-test because that is what A/B tests use."

"It depends on the metric. For conversion (yes/no) I use a two-proportion z-test; for revenue per user, Welch's t-test on the means because the groups are large and the decision is about totals; for a categorical metric, a χ² homogeneity test; and if the unit of analysis differs from the unit of randomization, none of these SEs are valid until I fix that."

Model answer: "I choose the test from the outcome type, the number of groups, whether the design is paired, and whether the sample is large enough for the Normal approximation, and I fix that choice before seeing the data. Then I report the effect size with a confidence interval, not only the p-value."

Choose by: outcome type → groups → paired? → can I trust the mean / are expected counts ≥ 5? → σ known?

Defaults: conversions → two-proportion z; numbers, 2 groups → Welch t; paired → paired t; 3+ → ANOVA; categories table → χ²; small/skewed → rank tests; tiny counts → exact tests.

Trap: choosing after seeing results; normality pre-tests; forgetting that the analysis unit must match the randomization unit.

Quick check: an e-mail test measures "number of clicks per user" (0, 1, 2, …, mostly 0), 50 000 users per arm. Which classical test, and why?

A number per user, two independent groups, very large $n$: Welch's t-test on the mean clicks per user. The data are skewed and full of zeros, but with 50 000 users the mean's sampling distribution is very close to Normal, and the business cares about total clicks, i.e. the mean. (A Bayesian model would use a Poisson or Negative Binomial likelihood for the counts.)

Recap, cheat sheet and practice

  • Every classical test is signal ÷ noise (or a sum of squared standardized gaps) read against a reference curve: z (σ known or $n$ large), t (σ estimated), χ² (counts vs expected), F (ratio of variances). χ² and F use only the right tail but detect differences in any direction.
  • z-tests: one-sample mean with known σ; one and two proportions (the classical conversion test). 2×2 χ² $= z^2$.
  • t-tests: one-sample; Welch for two independent groups (the safe default; Student's pooled test can give ~30% false alarms when the small group is the noisy one); paired = one-sample t on within-unit differences.
  • χ² tests: goodness-of-fit (K shares vs fixed shares, df $K-1$), independence and homogeneity (same arithmetic $E = RC/N$, df $(r-1)(c-1)$; different design). Counts only; expected counts ≥ 5 or use exact tests.
  • One-way ANOVA: $F = MSB/MSW$; one test for "any difference among k means"; many pairwise t-tests inflate false alarms (≈ 29% at k = 5).
  • Rank tests (Mann–Whitney, Wilcoxon signed-rank, Kruskal–Wallis): robust to outliers and skew, good for ratings; they test "one group tends to be larger", not means.
  • Choosing: outcome type → groups → paired? → can I trust the mean / expected counts? → σ known? Decide before seeing the data; report effect sizes and CIs.

Cheat sheet

TestUse whenStatisticReferenceSciPy / statsmodels
One-sample zmean vs target, σ known$(\bar x-\mu_0)/(\sigma/\sqrt n)$$N(0,1)$by hand
Two-proportion zconversion A vs B$(\hat p_2-\hat p_1)/\sqrt{\hat p(1-\hat p)(\frac1{n_1}+\frac1{n_2})}$$N(0,1)$proportions_ztest
One-sample tmean vs target, σ unknown$(\bar x-\mu_0)/(s/\sqrt n)$$t_{n-1}$ttest_1samp
Welch ttwo independent groups$(\bar x_1-\bar x_2)/\sqrt{s_1^2/n_1+s_2^2/n_2}$$t_{\text{Welch df}}$ttest_ind(…, equal_var=False)
Student ttwo groups, equal variances certainpooled $s_p$$t_{n_1+n_2-2}$ttest_ind (default!)
Paired tsame units twice$\bar d/(s_d/\sqrt n)$$t_{n-1}$ttest_rel
χ² goodness-of-fitK counts vs fixed shares$\sum (O-E)^2/E$$\chi^2_{K-1}$chisquare
χ² independence / homogeneityr × c table$\sum (O-E)^2/E$, $E=RC/N$$\chi^2_{(r-1)(c-1)}$chi2_contingency(…, correction=False)
One-way ANOVA3+ independent groups$\frac{SSB/(k-1)}{SSW/(N-k)}$$F_{k-1,N-k}$f_oneway
Mann–Whitney U2 groups, skew / outliers / ratingsU = pairs with a > bexact or Normal approx.mannwhitneyu
Wilcoxon signed-rankpaired or one sample, skewedsum of positive ranks of $|d_i|$exact or Normal approx.wilcoxon
Kruskal–Wallis3+ groups, skewedH from mean ranks$\chi^2_{k-1}$ approx.kruskal
Fisher's exact2×2 with small countsexact table probabilitiesexactfisher_exact
Code it · Python

import numpy as np
from scipy import stats
from statsmodels.stats.proportion import proportions_ztest

# 1) One-sample t: six days vs the old average of 50
x = np.array([52, 55, 49, 58, 54, 56])
r = stats.ttest_1samp(x, popmean=50)
print(round(r.statistic, 3), round(r.pvalue, 4))       # 3.098 0.0269

# 2) Two-proportion z-test (checkout) and the same test as a 2x2 chi-square
z, p = proportions_ztest([60, 50], [500, 500])
print(round(z, 4), round(p, 4))                         # 1.0107 0.3122
table = np.array([[50, 450], [60, 440]])
chi2, p2, dof, exp = stats.chi2_contingency(table, correction=False)
print(round(chi2, 4), round(p2, 4), round(z**2, 4))     # 1.0215 0.3122 1.0215  (chi2 = z^2)
print(round(stats.chi2_contingency(table)[1], 3))       # 0.363  <- Yates correction is ON by default

# 3) Student vs Welch from summary statistics (8 noisy users vs 32 calm users)
s = stats.ttest_ind_from_stats(12, 4, 8, 10, 1, 32, equal_var=True)
w = stats.ttest_ind_from_stats(12, 4, 8, 10, 1, 32, equal_var=False)
print(round(s.statistic, 3), round(s.pvalue, 4), "|", round(w.statistic, 3), round(w.pvalue, 4))
# 2.608 0.0129 | 1.403 0.202     (ttest_ind defaults to equal_var=True = Student!)

# 4) Paired vs (wrongly) independent
before = np.array([20, 35, 50, 28, 42, 60])
after = np.array([22, 38, 51, 31, 43, 62])
pr = stats.ttest_rel(after, before)
ind = stats.ttest_ind(after, before, equal_var=False)
print(round(pr.statistic, 3), round(pr.pvalue, 4), "|", round(ind.statistic, 3), round(ind.pvalue, 3))
# 5.477 0.0028 | 0.24 0.815

# 5) Chi-square goodness-of-fit and independence/homogeneity
g = stats.chisquare([88, 72, 40], f_exp=[100, 60, 40])
print(round(g.statistic, 2), round(g.pvalue, 4))        # 3.84 0.1466
res = stats.chi2_contingency([[60, 30, 10], [45, 40, 15]], correction=False)
print(round(res.statistic, 3), round(res.pvalue, 4), res.dof)   # 4.571 0.1017 2

# 6) One-way ANOVA and its rank-based cousin
A, B, C = [3, 5, 7, 5], [5, 9, 7, 7], [7, 11, 9, 9]
f = stats.f_oneway(A, B, C); k = stats.kruskal(A, B, C)
print(round(f.statistic, 2), round(f.pvalue, 4), "|", round(k.statistic, 2), round(k.pvalue, 4))
# 6.0 0.0221 | 6.41 0.0405

# 7) Mann-Whitney: the outlier does not move the rank test, but moves Welch
a, b, b_out = [2, 4, 7], [5, 8, 9, 12], [5, 8, 9, 120]
print(round(stats.mannwhitneyu(a, b, method="exact").pvalue, 3),
      round(stats.mannwhitneyu(a, b_out, method="exact").pvalue, 3))   # 0.114 0.114
print(round(stats.ttest_ind(a, b, equal_var=False).pvalue, 3),
      round(stats.ttest_ind(a, b_out, equal_var=False).pvalue, 3))     # 0.1 0.35

# 8) A/A simulation: Student's false-alarm rate when the SMALL group is the noisy one
rng = np.random.default_rng(1)
xa = rng.normal(0, 4, (20_000, 8))
xb = rng.normal(0, 1, (20_000, 32))
ps = stats.ttest_ind(xa, xb, axis=1, equal_var=True).pvalue
pw = stats.ttest_ind(xa, xb, axis=1, equal_var=False).pvalue
print((ps < 0.05).mean(), (pw < 0.05).mean())     # 0.298 0.04745  (Student ~30%, Welch ~5%)
Test yourself

1. Revenue per user in two arms of different sizes (2 000 vs 18 000 users). No reason to think the spreads are equal. Which test?

Two independent groups, a numeric outcome, large $n$ (so the mean is fine), unknown and possibly unequal variances: Welch. Student's pooling is risky with unequal sizes; paired needs the same units twice; χ² needs categories.

2. A 2×2 χ² test without continuity correction gives $\chi^2 = 4.0$. What is the two-proportion z statistic (in absolute value)?

For a 2×2 table, $\chi^2 = z^2$, so $|z| = \sqrt 4 = 2$. Same test, same p-value (about 0.046).

3. What does the Mann–Whitney U test actually test?

It compares ranks, so it answers "does one group tend to be larger?". It becomes a test of medians (a shift) only when both distributions have the same shape.

4. Five variants, no real differences. You run all 10 pairwise t-tests at α = 0.05. Roughly how often will at least one be "significant"?

The 10 tests share groups, so they are correlated; simulation gives about 29%, below the independent-tests value $1 - 0.95^{10} = 0.40$ but far above 5%. One ANOVA keeps the overall rate at 5%.

5. Independence vs homogeneity χ² tests: what is different?

Both use $E_{ij} = R_iC_j/N$ and df $(r-1)(c-1)$. Independence: one sample, two variables measured. Homogeneity: group sizes fixed by design (like A/B arms), one variable compared.

6. Six stores measured before and after a promotion. The two-sample t-test gives p = 0.82 and the paired t-test gives p = 0.003. Which is right, and why do they differ?

The design is paired. The independent test treats the store-to-store spread (sd ≈ 14) as noise and drowns a very consistent +2 change; the paired test uses the differences, whose sd is only 0.89.

Practice problems

A. Group 1: $n = 10$, $\bar x = 20$, $s = 6$. Group 2: $n = 10$, $\bar x = 15$, $s = 2$. Compute Welch's t and df, and Student's t. Why are the two t values equal?
  1. Welch: $SE = \sqrt{36/10 + 4/10} = \sqrt{4} = 2$, $t = 5/2 = 2.5$. $df = \dfrac{4^2}{3.6^2/9 + 0.4^2/9} = \dfrac{16}{1.44 + 0.0178} = 10.98$. $p = 0.030$.
  2. Student: $s_p^2 = (9 \cdot 36 + 9 \cdot 4)/18 = 20$, $SE = \sqrt{20 \times (1/10 + 1/10)} = \sqrt 4 = 2$, $t = 2.5$, $df = 18$, $p = 0.022$.
  3. With equal group sizes the two SE formulas coincide; only the df differ. Student's larger df gives a slightly smaller (over-optimistic) p-value.
B. Conversions: A 30 of 200, B 50 of 200. Compute the two-proportion z-test, then the 2×2 χ², and check $\chi^2 = z^2$.
  1. $\hat p_A = 0.15$, $\hat p_B = 0.25$, pooled $\hat p = 80/400 = 0.2$. $SE = \sqrt{0.2 \times 0.8 \times (2/200)} = \sqrt{0.0016} = 0.04$. $z = 0.10/0.04 = 2.5$, $p = 0.0124$.
  2. Table rows A: (30, 170), B: (50, 150). Expected per row: $200 \times 80/400 = 40$ converters and 160 non-converters.
  3. $\chi^2 = 2 \times \left(\frac{10^2}{40} + \frac{10^2}{160}\right) = 2 \times (2.5 + 0.625) = 6.25 = 2.5^2$. Same p = 0.0124.
C. Three groups of 5 have means 10, 12 and 14, and the within-group sum of squares is 48. Build the ANOVA table and decide at 5%.
  1. Grand mean $= 12$. $SSB = 5[(10-12)^2 + 0 + (14-12)^2] = 5 \times 8 = 40$, $df_1 = 2$, $MSB = 20$.
  2. $SSW = 48$, $df_2 = 15 - 3 = 12$, $MSW = 4$.
  3. $F = 20/4 = 5.0$ with (2, 12) df, $p = 0.026$ (critical value 3.89). Reject: not all means are equal. Which ones differ needs a post-hoc test (Tukey).
D. Wilcoxon signed-rank by hand for the store differences 2, 3, 1, 3, 1, 2.

Absolute differences sorted: 1, 1, 2, 2, 3, 3. Tied values share the average rank: the two 1s get 1.5, the two 2s get 3.5, the two 3s get 5.5. All differences are positive, so $W^+ = 1.5+1.5+3.5+3.5+5.5+5.5 = 21$, the largest possible value ($6 \cdot 7/2 = 21$). Under $H_0$ each sign is + or − with probability ½, so "all six positive" has probability $1/2^6 = 1/64$, and the two-sided p is $2/64 = 0.031$ (SciPy: 0.03125). Compare the paired t-test's 0.003: with only 6 pairs the rank test cannot go below 0.031.

E. Interview: "Choose a test for each: (a) share of users who churned, A vs B; (b) satisfaction 1–5, A vs B; (c) checkout time of the same 40 testers on old and new design; (d) is our support-ticket category mix (billing, bug, other) different this month from the long-run mix?"
  • (a) yes/no in two independent groups → two-proportion z-test (Fisher if counts are tiny).
  • (b) ordered rating, two independent groups → Mann–Whitney U (or a χ² homogeneity test on the 5 categories).
  • (c) a number, same testers twice → paired t-test on the differences (Wilcoxon signed-rank if the differences are very skewed).
  • (d) K category counts vs fixed shares → χ² goodness-of-fit (use the counts, check expected counts ≥ 5).
F. Interview: "Why is Welch's t-test the default, and why not test for equal variances first?"

"Student's test pools the two variances. When group sizes differ and the smaller group is noisier, the pooled SE is too small and the false-positive rate can be several times α (about 30% in an 8 vs 32 example with a 4× larger sd); when the larger group is noisier, it becomes too conservative and loses power. Welch keeps separate variances and adjusts the degrees of freedom, and loses almost nothing when variances are equal. A variance pre-test is weak exactly when groups are small, and choosing the test based on it distorts the error rate, so I use Welch from the start."

Chapter 5.10 · Syllabus Module 17

Designing an A/B experiment

Most A/B tests that go wrong were lost before the first user arrived: the coin was flipped for users but the data were counted per session, the metric was too noisy for the traffic, the test could only detect effects three times bigger than anyone expected, or nobody had said what "worth shipping" means. This chapter is about the design: who gets what, what you count, how many users you need, how long to run, and how to read the result against a business threshold. It applies to a classical test and to a Bayesian framework like yours alike.

  • Name the parts of an experiment: control, treatment, treatment assignment, treatment effect (absolute and relative lift)
  • Explain why randomization makes the comparison fair, and how hash-based assignment works in practice
  • Tell the randomization unit from the analysis unit, and show with a simulation why a mismatch makes confidence intervals too narrow (the design effect)
  • Use exposure and triggering to stop diluting the effect with users who could not be affected
  • Choose primary, secondary and guardrail metrics, and see how a metric's noise sets the experiment's size
  • Plan the sample size from α, power and the minimum detectable effect (absolute vs relative), and turn it into a test duration from traffic
  • Judge results by practical significance: compare the whole confidence interval with a "worth it" threshold
  • Plan a Bayesian A/B test by simulating its decision rule

What we need from earlier chapters: standard errors of a proportion and of a difference (Chapter 5.5); sampling, independence and IID (Chapter 5.4); Type I/II errors, power and the sample-size formula (Chapter 5.7); confidence intervals (Chapter 5.8); the two-proportion z-test and Welch's t-test (Chapter 5.9). What comes next: the ways a running experiment goes wrong (sample ratio mismatch, peeking, interference and SUTVA, many metrics) are the subject of Chapter 5.11; why randomization licenses a causal claim is in Chapter 5.12.

Control, treatment and the treatment effect core

An A/B test is a fair race. Users who arrive are split at random into two groups. One group sees the product as it is today: the control (A). The other sees the change: the treatment (B). Everything else (the season, the prices, the bugs, the mood of the internet) hits both groups the same way. So when the groups end up different, the change is the only explanation left, apart from luck.

The treatment effect is how much the change moves the average outcome. We estimate it with the simplest thing possible: the average in B minus the average in A. Luck still makes that estimate wobble, so it always comes with a standard error and a confidence interval.

Three ways to say it:

  • Picture: two identical queues; only one thing differs between them.
  • Numbers: 10.0% vs 11.0% conversion is an effect of +1.0 percentage point (absolute), or +10% (relative).
  • Slogan: randomize, change one thing, compare averages, report the uncertainty.

10 000 users per arm. Control: 1 000 conversions; treatment: 1 100.

  1. Rates: $\hat p_A = 1000/10000 = 0.100$, $\hat p_B = 1100/10000 = 0.110$.
  2. Absolute effect (in percentage points): $0.110 - 0.100 = 0.010$, i.e. +1.0 point.
  3. Relative effect (lift): $0.010/0.100 = 0.10$, i.e. +10%.
  4. Standard error of the difference: $\sqrt{\frac{0.1 \times 0.9}{10000} + \frac{0.11 \times 0.89}{10000}} = \sqrt{0.0000090 + 0.0000098} = \sqrt{0.0000188} = 0.00433$.
  5. 95% CI for the absolute effect: $0.010 \pm 1.96 \times 0.00433 = 0.010 \pm 0.0085$, from +0.15 to +1.85 points. Zero is outside, so the effect is statistically significant at 5%.
  6. Dividing the ends by the control rate gives roughly +1.5% to +18.5% relative. That is only approximate: it ignores the uncertainty in the control rate itself (a ratio needs its own method, e.g. the delta method or a log-ratio interval).
  • Control (A): the current experience. Treatment (B): the changed experience. With more than one treatment, each version is a variant or arm.
  • Treatment assignment: the rule that decides which arm each unit gets (next section).
  • Outcome $Y$: what we measure for each unit (converted or not, revenue, sessions).
  • Average treatment effect: $\tau = E[Y \mid \text{everyone gets B}] - E[Y \mid \text{everyone gets A}]$: the difference between two worlds, only one of which we can see for each user. (The precise language of potential outcomes $Y(1), Y(0)$ is in Chapter 5.12.)
  • Estimator: $\hat\tau = \bar Y_B - \bar Y_A$. Under random assignment it is unbiased for $\tau$, with $SE(\hat\tau) = \sqrt{s_A^2/n_A + s_B^2/n_B}$ when units are independent.
  • Absolute effect $\tau$ (in the metric's units or percentage points) vs relative effect (lift) $\tau/\mu_A$ (a percentage). Always say which one you mean.
Why do we need it?

Without a control group running at the same time, any change in a metric could come from the season, a marketing campaign or a bug. A randomized control is the only clean way to attribute a difference to the change itself.

Where is it used?

Product A/B tests (checkout, ranking, pricing pages), e-mail and notification tests, online evaluation of a new ML model against the current one, clinical trials, and policy experiments. Your Bayesian framework analyses exactly this setup.

How is it used?

Define the two (or more) experiences and the outcome before launch. Assign users at random, log each user's arm and outcome, then estimate $\hat\tau$ with its CI (classical) or the posterior of $\theta_B - \theta_A$ (Bayesian), and report both absolute and relative effects.

Users arrive Eligible? country, app… Random assignment Control (A) current page Treatment (B) new page Outcomes effect = B − A ± CI a coin (in practice a hash of the user id) decides; nobody chooses exposure: did the user really see it?
The life of an A/B test. Eligible users are assigned by chance to control or treatment, they are exposed to their arm, their outcomes are logged, and the treatment effect is estimated as the difference between the arms, with its uncertainty. Each box is a design decision covered in this chapter.

Users arrive in batches and a coin sends each to A or B. The orange line is the estimated effect (B − A, in percentage points) after each batch; the orange band is its 95% CI; the green line is the true effect, which a real analyst never sees. Press Next 1 000 users a few times, then Run to 20 000. Early on the estimate jumps around; later the band narrows like $1/\sqrt n$. Set the true effect to 0 and press New experiment several times: the band sometimes leaves 0 for a while by luck. Deciding the moment it does is "peeking" (Chapter 5.11).

"B converted 11.0% and A 10.0%, so the treatment effect is +1 point."

+1 point is the estimate. The effect itself is unknown; the estimate wobbles around it with standard error 0.43 points here. Report "+1.0 point, 95% CI +0.15 to +1.85".

"A 10% lift" (said without context).

Say "+10% relative (10.0% → 11.0%)" or "+1 percentage point". A "10% lift" on a 10% baseline and a "10-point lift" differ by a factor of ten.

"The treatment effect is +1 point for every user."

It is an average. Some users may gain a lot, some nothing, some may be hurt. Segment-level effects need their own analysis (and their own multiple-testing care, Chapter 5.11).

In an A/B framework like yours, the treatment effect is not one number but a posterior distribution: each posterior draw of $(\theta_A, \theta_B)$ gives an absolute effect $\theta_B - \theta_A$ and a relative effect $(\theta_B - \theta_A)/\theta_A$, and the decision quantities such as $P(\theta_B \gt \theta_A \mid D)$ are averages over those draws. The relative lift's posterior comes for free this way, without the ratio approximation that the classical CI needs.

Control A, treatment B, random assignment. Effect $\tau$ = difference in average outcome; estimate $\hat\tau = \bar Y_B - \bar Y_A$, unbiased under randomization.

Absolute (points) vs relative (%) effect: always say which. $SE = \sqrt{s_A^2/n_A + s_B^2/n_B}$.

Trap: the observed difference is an estimate, not the effect; report its CI.

Quick check: baseline 4%, treatment 4.4%. What are the absolute and relative effects?

Absolute: $4.4 - 4.0 = 0.4$ percentage points. Relative: $0.4/4.0 = 0.10$, i.e. +10%. Same change, two very different-looking numbers.

Treatment assignment: why a coin must decide core

Imagine letting users choose the new checkout ("try our new beta!"). Eager, loyal, heavy buyers opt in much more often. The treatment group then converts more, but because of who is in it, not because of the new page. Any rule that lets a person, a team or the calendar decide who gets the treatment can sneak a hidden difference between the groups.

A coin cannot be eager. If chance alone decides, every hidden trait (loyalty, device, country, time of day) ends up spread evenly between the arms on average, including traits nobody thought of. What remains is ordinary luck, which the standard error already measures.

Three ways to say it:

  • Picture: a coin is the only referee with no favourite team.
  • Numbers: with opt-in, 60% of the treatment group are power users vs 14% of control: a fake "+6 point lift" for a change that does nothing.
  • Slogan: randomize, and the hidden differences cancel on average.

How real systems flip the coin. They do not call a random number generator on every visit (the same user would bounce between arms). They hash the user id together with an experiment name (a "salt"):

  1. Build the string "user_42:checkout-v2".
  2. Hash it to a large integer. With the 32-bit FNV-1a hash used in the widget below, this gives 1 927 872 456.
  3. Take the remainder after dividing by 100: bucket $= 1\,927\,872\,456 \bmod 100 = 56$.
  4. Rule: buckets 0–49 → control, 50–99 → treatment. User 42 gets treatment, on every visit, on every server (the hash is deterministic).
  5. A different experiment uses a different salt ("user_42:search-rank"), which gives an unrelated bucket, so being in treatment in one test says nothing about the other.
  • Treatment assignment: the rule that maps each unit to an arm, $Z_i \in \{A, B\}$.
  • Randomized assignment: $Z_i$ is decided by a chance mechanism (or a pseudo-random hash) that does not depend on anything about the unit, in particular not on how the unit would respond. Consequences: the arms are comparable on average in every trait, observed or not, and $\bar Y_B - \bar Y_A$ is unbiased for the average treatment effect.
  • Allocation ratio: the share of units sent to each arm (50/50 is the most efficient for a fixed total; more in the power section).
  • Hash-based assignment: bucket $= hash(\text{unit id}, \text{salt}) \bmod 100$. It is deterministic (sticky: same user, same arm), reproducible, needs no stored table, and different salts make experiments independent of each other.
  • Randomization balances the arms only on average. In one particular experiment there can be small chance imbalances; that is exactly the "luck" the standard error accounts for.
Why do we need it?

Every other way of choosing who gets the treatment (opt-in, sales picking customers, "launch in one country first", "new users only on Monday") lets a hidden difference masquerade as the treatment effect. Randomization is what turns a correlation into a causal estimate (Chapter 5.12).

Where is it used?

Every online experimentation platform (assignment by hashing user or device ids with an experiment salt), clinical trials (randomization lists), and ML model rollouts (a random share of traffic to the new model).

How is it used?

Pick the unit (user, device, account), hash its id with a per-experiment salt, map buckets to arms with the planned allocation, log the assignment, and check after launch that the arm sizes match the plan (the sample ratio mismatch check, Chapter 5.11).

The change does nothing (true effect 0). 30% of users are hidden "power users" who convert at 20% (others: 6%). Pick an assignment rule and press Run 10 experiments several times. Left: share of power users in each arm (last experiment). Right: each dot is one experiment's estimated effect. With the coin, dots scatter around 0. With opt-in or by sign-up date, they pile up around +5 to +7 points: a fake effect that more data will never remove.

Each cell is one user (ids 1001–1040) with its bucket = hash("user_id:salt") mod 100; blue = control, orange = treatment. Switch the salt (the experiment's name): the users reshuffle, but within one salt a user always lands in the same bucket. Move the treatment share to 10%: only buckets 90–99 get the treatment (a common way to start a risky change small). The readout hashes 10 000 users to check the split and the independence of two experiments.

"Alternating users (A, B, A, B…) or splitting by time of day is random enough."

Systematic rules can line up with hidden patterns (bots arriving in bursts, mornings vs evenings, odd/even ids from different systems). Use a proper random or hashed assignment.

"Randomization guarantees my two groups are identical."

It guarantees they are identical on average over many possible assignments. Your one assignment can show small chance imbalances; the standard error and the CI already include that luck. (Pre-experiment covariates can reduce it further: CUPED, Chapter 5.12.)

"The arms look unbalanced on country, so I will re-assign until they look balanced."

Re-drawing by eye breaks the randomization the analysis relies on. If balance on a key variable matters, use a planned design (stratified randomization) and analyse it accordingly.

The posterior $P(\theta_B \gt \theta_A \mid D)$ in your framework is only a statement about the change if the two arms are comparable, which is what randomized assignment provides. If a "treatment group" were formed by opt-in or by rollout date, the Beta-Binomial model would still produce a confident-looking posterior, but for the difference between two kinds of users, not for the effect of the change. Bayesian modelling does not repair a non-random assignment.

Random assignment ⇒ arms comparable on average in every trait (seen or unseen) ⇒ $\bar Y_B - \bar Y_A$ unbiased for the effect.

In practice: bucket = hash(user id + experiment salt) mod 100 → sticky, reproducible, independent across experiments.

Trap: opt-in, calendar or "pick the best customers" assignments create bias that more data cannot fix.

Quick check: a team launches the new feature to all users in Canada and keeps the US as control. What is the problem?

Assignment is decided by country, not by chance. Any difference between Canadian and US users (prices, seasons, holidays, behaviour) mixes with the feature's effect, and there are effectively only two "units" (two countries), so there is no honest standard error. Randomize users (or, if it must be by region, many regions at random, analysed as clusters: next section).

Randomization unit vs analysis unit core

Suppose you flip the coin per user, but compute the conversion rate per session, treating 50 000 sessions as 50 000 independent observations. A user's five sessions are not five independent opinions: the same person, with the same taste, budget and device, is behind all of them. It is like interviewing one person five times and reporting five votes.

The standard-error formula only knows the row count. It believes you have more independent information than you do, so the SE is too small, the confidence interval is too narrow, and "significant" results appear far more often than α promises, even in an A/A test where nothing changed.

Three ways to say it:

  • Picture: five sessions from one user are five photos of the same person, not a crowd.
  • Numbers: 5 sessions per user, within-user correlation 0.3 → the naive SE is 1.48 times too small, and a "95%" interval covers the truth only about 81% of the time.
  • Slogan: analyse at the level you randomized, or correct the SE for clustering.

1 000 users per arm, each with 5 sessions (5 000 sessions per arm). The intra-class correlation (how alike two sessions of the same user are) is $\rho = 0.3$.

  1. Design effect: $DE = 1 + (m - 1)\rho = 1 + (5 - 1) \times 0.3 = 1 + 1.2 = 2.2$. The true variance of the arm average is 2.2 times what the naive formula assumes.
  2. Effective sample size: $5000/2.2 \approx 2\,273$. The 5 000 correlated sessions carry the information of about 2 273 independent ones.
  3. The naive SE is too small by $\sqrt{2.2} = 1.48$.
  4. A naive "95%" interval $\pm 1.96 \times SE_{naive}$ is really $\pm 1.96/1.48 = \pm 1.32$ true standard errors wide, which covers the truth with probability $2\Phi(1.32) - 1 \approx 0.81$. In an A/A test, about 19% of experiments would look "significant" instead of 5%.
  5. Fix: compute one number per user (e.g. that user's conversion rate or total), then compare the user-level averages: 1 000 independent values per arm, and the SE is right.
  • Randomization unit: the thing the coin is flipped for (user, device, account, session, page view, or a cluster such as a store or a city).
  • Analysis unit: the thing each row of the analysis represents (the denominator of the metric).
  • The usual SE formulas assume the analysis units are independent. Observations that share a randomization unit are typically positively correlated, so the SE must be computed at the randomization unit (or corrected).
  • Intra-class correlation $\rho = \dfrac{\text{between-unit variance}}{\text{between-unit} + \text{within-unit variance}}$, the correlation between two observations of the same unit.
  • Design effect for clusters of equal size $m$: $DE = 1 + (m-1)\rho$; the true variance is $DE \times$ the naive variance; effective sample size $= n_{rows}/DE$.
  • Fixes: aggregate to the randomization unit (per-user metrics); cluster-robust standard errors; the delta method for ratio metrics like "clicks per page view" $= \sum_u \text{clicks}_u / \sum_u \text{views}_u$; or a bootstrap that resamples whole users.
  • If the analysis unit is coarser than the randomization unit there is no such problem; the problem is analysing finer than you randomized.
Why do we need it?

Session-, page-view- and event-level metrics (click-through rate, revenue per session, latency per request) are everywhere, but randomization is almost always per user. Without the correction, an experimentation platform produces systematically overconfident results and too many false wins.

Where is it used?

Ratio metrics in every large experimentation platform (CTR, revenue per session), cluster-randomized tests (by store, city, school, or by advertiser in a marketplace), switchback tests over time slots, and multi-level surveys.

How is it used?

Write down both units in the experiment plan. If they differ, analyse per randomization unit, or use the delta method / cluster-robust SEs (statsmodels OLS with cov_type="cluster"). Validate with A/A tests: a correct pipeline flags about α of them.

Randomized per USER (the coin is flipped once per person) user 1 → A user 2 → B user 3 → A Analysed per SESSION 11 rows, but only 3 independent coin flips: sessions of one user move together Fix one value per user (or cluster-robust SE, delta method)
The coin decides per user; each user then produces several sessions. Rows from the same user are correlated, so the number of rows overstates the information. The SE must count users, not sessions.

Every experiment is an A/A test (no real difference) with 80 users per arm, randomized per user. Each row is one experiment's 95% CI for the difference; red = it misses the truth (0). Left: sessions treated as independent. Right: one value per user. Press Run 40 more a few times. With 5 sessions per user and correlation 0.3, the left side misses about 19% of the time; the right side about 5%. Set sessions per user to 1, or correlation to 0: the two sides agree.

The curves show the real coverage of a "95%" interval computed per session, as the number of sessions per user grows, for several within-user correlations (grey) and the one you choose (purple). Drag the purple dot along the curve. Even a small correlation of 0.05 becomes serious when users have 20 sessions: DE = 1.95.

"We have 5 million page views, so even tiny effects are reliably measured."

Page views from the same users are correlated; the effective sample size is closer to the number of users, possibly far fewer than 5 million. Count the independent randomization units.

"Click-through rate = total clicks / total views, so I can use the binomial SE with n = views."

Views are not independent trials when users are randomized. Use the delta method (or a per-user bootstrap) for ratio metrics; the binomial SE is usually too small.

The same problem exists in Bayesian models. If, in a framework like yours, a metric were modelled per session with an iid Binomial or Normal likelihood while users were the randomized units, the posterior would be too narrow for exactly the reason above: the likelihood would treat correlated sessions as independent evidence. Two honest fixes: aggregate to one observation per user before modelling, or add a per-user random effect (a hierarchical layer, Chapter 6.5), which is the Bayesian version of the design effect.

"We randomized by user and computed the session-level conversion rate with a standard two-proportion test."

"We randomized by user, so we either compute a per-user metric, or use the delta method / cluster-robust SEs for the session-level ratio; otherwise the SE ignores within-user correlation."

Model answer: "The analysis unit has to match the randomization unit, or the variance has to account for clustering. With $m$ observations per user and intra-class correlation $\rho$, the variance is inflated by $1 + (m-1)\rho$; ignoring that makes intervals too narrow and inflates false positives. A/A tests are a good way to check the pipeline."

SE must be computed at the randomization unit. $DE = 1 + (m-1)\rho$; effective $n = n_{rows}/DE$; naive SE too small by $\sqrt{DE}$.

Example: m = 5, ρ = 0.3 → DE = 2.2 → naive "95%" CI covers ≈ 81%.

Fixes: per-user aggregation, delta method for ratios, cluster-robust SE, user-level bootstrap. Trap: counting rows as evidence.

Quick check: users are randomized, and every user has exactly one session. Is there a unit mismatch?

No. With $m = 1$, $DE = 1 + 0 \cdot \rho = 1$: one session per user means sessions are users. The problem appears only when a randomization unit contributes several correlated rows.

Exposure and triggering: counting the users who could be affected

You change the checkout page. But only 20% of visitors ever reach checkout. The other 80% never see the difference between A and B, so their behaviour cannot change. If you analyse all visitors, those 80% add a lot of noise and no signal: the effect gets diluted, and the experiment needs many more users to see it.

Triggering means: analyse only the users who reached the point where A and B actually differ. The rule must be applied identically in both arms; in control, you log the moment a user would have seen the change (they reached checkout, where the old page was shown).

Three ways to say it:

  • Picture: testing a new menu in a restaurant by surveying everybody who walked past the door.
  • Numbers: +10% among the 20% who reach checkout is only +2% across all visitors, and needs about 4.8 times the traffic to detect.
  • Slogan: unaffected users add noise, not signal.

Visitors convert at 5%, whether or not they reach checkout. The new checkout page raises conversion by 10% relative (5% → 5.5%) among users who reach checkout; 20% of visitors reach it ($q = 0.2$). α = 0.05, power 80%.

  1. Triggered analysis (only checkout visitors): 5% vs 5.5% needs 31 234 triggered users per arm. Since only 20% of visitors trigger, that is $31\,234/0.2 = 156\,170$ visitors per arm.
  2. All-visitors analysis: the overall rate moves from 5% to $5\% + 0.2 \times 0.5\text{ points} = 5.1\%$. The effect is diluted to 0.1 points (+2% relative). Detecting 5% vs 5.1% needs 752 703 visitors per arm.
  3. Ratio: $752\,703/156\,170 \approx 4.8$. Same experiment, same users, almost five times more traffic for the diluted analysis.
  4. Reporting: the triggered effect (+0.5 points among checkout visitors) is the effect on the affected users. The site-wide impact is the triggered effect × trigger rate: $0.2 \times 0.5 = 0.1$ points.
  • Exposure: a unit actually experienced its assigned arm (the page rendered, the e-mail was opened, the model's ranking was shown). Assignment and exposure are different events.
  • Trigger condition: the rule that decides which assigned units enter the analysis, e.g. "reached checkout". It must be (i) determined by something the treatment cannot change, and (ii) logged the same way in both arms ("counterfactual logging" in control).
  • Intention-to-treat (ITT): analyse every assigned unit in its assigned arm, exposed or not. Always valid (it keeps the randomization) but diluted.
  • Dilution: if untriggered users are unaffected, the all-users effect is $q \times$ the triggered effect, where $q$ is the trigger rate. Their outcomes still add variance.
  • A trigger based on post-treatment behaviour (something the treatment can change, like "clicked the new button") breaks the randomization: the triggered groups are no longer comparable.
Why do we need it?

Most changes touch only part of the product. Without triggering, experiments on deep pages, rare features or edge cases would need impossible traffic, and good changes would be dismissed as "no effect".

Where is it used?

Checkout and payment changes, search ranking (only users who searched), recommendation widgets (only users who scrolled to them), ML model swaps that only change some predictions, notification tests (only users who were eligible to receive one).

How is it used?

Define the trigger in the experiment plan; log it in both arms at the moment the experiences would diverge; analyse triggered users; check the trigger rate is equal across arms (if the treatment changes it, the trigger is contaminated); translate the result to site-wide impact by multiplying by the trigger rate.

All visitors (assigned at random) 16 of 20 never reach checkout: they see no difference 4 of 20 reach checkout: the TRIGGERED users Effect seen in the analysis triggered users only: all visitors: = 20% of it …while the noise of all 20 users stays In control, log when a user reaches checkout too, so both arms use the same trigger.
Only users who reach the changed page can respond to it. Analysing everyone shrinks the effect to trigger rate × triggered effect, while every user still adds variance, so the signal-to-noise ratio collapses.

Set the share of users who reach the changed page, their baseline conversion, the conversion of everyone else, and the true relative lift among triggered users. The bars show how many visitors per arm each analysis needs for 80% power. At a trigger rate of 20% the all-users analysis needs about 4.8× the traffic. Set the trigger rate to 100%: the two coincide. Set "others convert" to 0%: the gap shrinks, because untriggered users then add no noise (they all have outcome 0).

"Trigger on users who clicked the new 'express pay' button."

That button exists only in treatment, and clicking it is itself a result of the treatment. The triggered groups are no longer comparable. Trigger on something both arms share and the treatment cannot change (reached checkout).

"The triggered effect was +10%, so the site will gain 10%."

Only the triggered users gain. Site-wide impact ≈ trigger rate × triggered effect (here 20% × 10% relative = about 2% relative), assuming untriggered users are unaffected.

Exposure = actually saw the arm. Trigger = the condition where A and B diverge; log it in BOTH arms; never on post-treatment behaviour.

All-users effect = q × triggered effect, but all users add noise → triggering can save several-fold traffic.

Trap: report site-wide impact, not the triggered lift, as the business gain.

Quick check: in treatment, the new page loads faster, so more users reach checkout (trigger rate 22% vs 20% in control). Is "reached checkout" still a safe trigger?

No. The treatment changes who triggers, so the triggered groups differ in composition (treatment's extra 2% may be less committed users). Check trigger rates across arms; if they differ, use an earlier trigger that the treatment cannot affect, or the intention-to-treat analysis.

Primary, secondary and guardrail metrics core

Before launch, decide what "success" means. Otherwise, after the test, someone will find some metric that went up (with 30 metrics, one or two always do by luck) and declare victory.

  • Primary metric: the one number the decision is about (e.g. purchase conversion). One, or a very small pre-declared set.
  • Secondary metrics: help explain why (add-to-cart rate, time on page). They are not the decision.
  • Guardrail metrics: things that must not get worse even if the primary metric improves (page load time, error rate, unsubscribes, refunds, revenue).

The choice of metric also sets the size of the experiment: a noisy metric (revenue, with a few huge spenders) needs far more users than a calm one for the same relative change.

Three ways to say it:

  • Picture: the primary metric is the destination; guardrails are the crash barriers along the road.
  • Numbers: detecting a 5% lift needs about 119 000 users per arm on conversion but about 343 000 on revenue per user.
  • Slogan: one goal, a few explanations, and fences you must not break, all chosen before the data.

Metric sensitivity. A shop converts 5% of visitors; buyers spend on average 60 (sd 80). We want to detect a +5% relative change, α = 0.05, power 80%. A handy formula for the users per arm: $n \approx 2(z_{0.975} + z_{0.8})^2\, CV^2/r^2 = 15.7\, CV^2/r^2$, where $CV = \sigma/\mu$ is the metric's coefficient of variation and $r$ the relative lift.

  1. Conversion (a yes/no per visitor): $\mu = 0.05$, $\sigma = \sqrt{0.05 \times 0.95} = 0.218$, $CV = 0.218/0.05 = 4.36$.
  2. $n \approx 15.7 \times 4.36^2/0.05^2 = 15.7 \times 19.0/0.0025 \approx 119\,300$ per arm.
  3. Revenue per visitor: mean $= 0.05 \times 60 = 3$. Variance $= E[Y^2] - (E[Y])^2 = 0.05 \times (80^2 + 60^2) - 3^2 = 0.05 \times 10\,000 - 9 = 491$, sd $= 22.2$, $CV = 22.2/3 = 7.39$.
  4. $n \approx 15.7 \times 7.39^2/0.05^2 \approx 342\,600$ per arm: about 2.9 times as many as for conversion.
  5. Revenue may still be the right primary metric (it is what the business cares about), but you must budget the traffic for it, or use variance reduction (CUPED, Chapter 5.12; capping extreme values).
  • Primary metric (also called the OEC, overall evaluation criterion): the pre-declared metric whose result decides the launch. It should be aligned with long-term value, sensitive (moves enough to detect) and attributable (the change can plausibly move it).
  • Secondary metrics: diagnostic; they explain the mechanism and are reported, but each extra metric tested at α adds false-alarm risk (Chapter 5.11).
  • Guardrail metrics: must not degrade beyond a tolerance. Often tested as non-inferiority: "the CI for the change in latency lies entirely below +20 ms", rather than "no significant increase".
  • Metric sensitivity: how many units are needed to detect a given relative change; driven by $CV^2$. Metric types: proportions (conversion), means (revenue per user), ratios (CTR, revenue per session; need the delta method), counts.
Why do we need it?

Without a pre-declared primary metric the analysis turns into a search for good news. Without guardrails, a change can win on clicks while slowing the site or annoying users. And without checking sensitivity, an experiment can be doomed from the start.

Where is it used?

Experiment scorecards in every experimentation platform (primary, secondary, guardrail sections), launch reviews, ML model rollouts (accuracy as primary, latency and cost as guardrails), and your framework's choice of which posterior decides.

How is it used?

In the plan: name one primary metric with its MDE, list secondary metrics as "explanatory", list guardrails with tolerances. Compute the sample size for the primary metric. After the test: decide on the primary, check guardrails with non-inferiority intervals, read secondaries as explanation.

PRIMARY (decides) purchase conversion SECONDARY (explain) add-to-cart rate, checkout starts, time to purchase GUARDRAILS (must not break) page load time, errors, refunds, unsubscribes, revenue Decision rule written before launch: ship if the primary improves by at least the threshold AND every guardrail stays within its tolerance
One primary metric decides, a few secondary metrics explain the mechanism, and guardrails protect against hidden damage. The rule that combines them is written before the data arrive.

Set the conversion rate, the buyers' average spend and its spread, and the relative lift to detect. The bars show users per arm needed (α = 0.05, power 80%) for each metric, using $n \approx 15.7\,CV^2/r^2$. Raise the spend sd: revenue gets noisier and much more expensive to test. Halve the lift: both bars grow four times.

"Conversion was flat but time-on-page rose significantly, so the test is a win."

The decision belongs to the pre-declared primary metric. Promoting a secondary metric after the fact is a forking path: with many metrics, something always moves by chance (Chapter 5.11).

"The latency guardrail was not significantly worse, so it is fine."

"Not significant" may just mean "not enough data". A guardrail should be checked with a tolerance: the upper end of the CI for the latency change must be below the allowed limit (non-inferiority).

In a framework like yours, metric type maps directly to the likelihood: conversion → Beta-Binomial, a categorical metric → Dirichlet-Multinomial, revenue or time → Normal or Student-t (Student-t for the heavy tail of spenders), counts → Poisson. The primary/guardrail split carries over as decision rules, e.g. ship if $P(\text{lift}_{primary} \gt \delta \mid D) \gt 0.95$ and $P(\text{latency increase} \gt \text{tolerance} \mid D) \lt 0.05$. Metric sensitivity carries over too: a noisy revenue metric gives a wide posterior unless you have the traffic.

Primary (decides, pre-declared, one), secondary (explain), guardrails (must not break; check with a tolerance / non-inferiority).

Users per arm $\approx 15.7\,CV^2/r^2$ (α 0.05, power 0.8). Conversion CV $= \sqrt{(1-p)/p}$; revenue is usually much noisier.

Trap: choosing the metric after seeing the results; "guardrail not significant" ≠ "guardrail safe".

Quick check: why is "1 − p" inside the CV of conversion, and what happens for rare events?

For a yes/no metric, $\sigma = \sqrt{p(1-p)}$ and $\mu = p$, so $CV = \sqrt{p(1-p)}/p = \sqrt{(1-p)/p}$. For rare events ($p$ small) the CV is large: at $p = 1\%$, $CV = \sqrt{99} \approx 9.9$. Rare events need huge experiments to detect relative changes.

Sample-size planning: power and the minimum detectable effect core

Before running, ask: "If the change really helps by at least X, will my test notice?" A test that is too small will usually miss real improvements and report "no significant difference": a wasted experiment that also teaches the team the wrong lesson.

The minimum detectable effect (MDE) is the smallest true effect the experiment can detect with the chosen reliability (the power, usually 80%) at the chosen α. Pick the MDE from the business side ("we care about +1 point or more"), then compute how many users that needs. The key fact: effects half as big need about four times the users, because noise shrinks only like $1/\sqrt n$.

Three ways to say it:

  • Picture: the MDE is the size of the smallest fish your net can reliably catch; a finer net (more users) catches smaller fish.
  • Numbers: baseline 10%: MDE +1 point needs 14 751 users per arm; MDE +0.5 points needs 57 763.
  • Slogan: half the effect, four times the users.

Baseline conversion $p_1 = 0.10$. We want to detect $p_2 = 0.11$ (MDE = +1 point absolute, +10% relative). Two-sided α = 0.05, power 80%. The formula (Chapter 5.7) is

$$n = \frac{\left(z_{1-\alpha/2}\sqrt{2\bar p(1-\bar p)} + z_{1-\beta}\sqrt{p_1(1-p_1) + p_2(1-p_2)}\right)^2}{(p_2 - p_1)^2}\ \text{ per arm}, \quad \bar p = \tfrac{p_1+p_2}{2}.$$
  1. $z_{0.975} = 1.960$, $z_{0.80} = 0.842$, $\bar p = 0.105$.
  2. $\sqrt{2 \times 0.105 \times 0.895} = \sqrt{0.18795} = 0.4335$; times 1.960 gives 0.8497.
  3. $\sqrt{0.10 \times 0.90 + 0.11 \times 0.89} = \sqrt{0.0900 + 0.0979} = \sqrt{0.1879} = 0.4335$; times 0.842 gives 0.3648.
  4. Sum $= 1.2145$; squared $= 1.4751$; divided by $0.01^2$: $n = 14\,751$ per arm (29 502 in total).
  5. Rule-of-thumb check (Lehr's rule, for α = 0.05 and 80% power): $n \approx 16\,\sigma^2/\delta^2 = 16 \times 0.09/0.0001 = 14\,400$. Close.
  6. Halve the MDE to +0.5 points (5% relative): $n = 57\,763$ per arm, 3.9 times more. Double it to +2 points: $n = 3\,841$.
  • Power $1-\beta$: the probability of a significant result when the true effect equals the MDE. Usual choice 80% (sometimes 90%).
  • MDE $\delta$: the smallest true effect the design detects with that power at level α. Absolute MDE: in points or metric units ($p_2 - p_1$). Relative MDE: $\delta/p_1$. Tools differ in which one they ask for; always check.
  • For means with sd $\sigma$: $n = 2(z_{1-\alpha/2} + z_{1-\beta})^2\sigma^2/\delta^2$ per arm; inverted: $MDE \approx (z_{1-\alpha/2} + z_{1-\beta})\sqrt{2\sigma^2/n}$.
  • $n$ grows with $\sigma^2$ (noisier metric), with smaller $\delta$ ($1/\delta^2$), with smaller α and with higher power.
  • Allocation: with a fraction $f$ of $N$ users in treatment, $Var(\hat\tau) \propto \frac{1}{fN} + \frac{1}{(1-f)N}$, smallest at $f = \tfrac12$. A 90/10 split needs $\frac{1/0.1 + 1/0.9}{4} = 2.78$ times the total traffic of a 50/50 split for the same power.
  • With several treatment arms, each comparison needs its own $n$ per arm, and multiple comparisons usually call for a smaller α per comparison (Chapter 5.11).
Why do we need it?

It tells you before launch whether the experiment can answer the question at all, how much traffic it will cost, and what "no significant difference" will mean afterwards (only "no effect larger than about the MDE, probably").

Where is it used?

Every experiment plan, sample-size calculators in experimentation platforms, clinical-trial protocols, deciding whether a test is worth running on a low-traffic page, and choosing between metrics or variance-reduction methods.

How is it used?

Agree on the MDE with the business, take the baseline and variance from historical data, then use statsmodels.stats.power (e.g. NormalIndPower().solve_power) or the formula. If the $n$ is unaffordable, raise the MDE, pick a less noisy metric, use triggering or CUPED, or do not run the test.

The curve gives users per arm (log scale) needed to detect each relative MDE from the baseline. Drag the purple dot along the curve. From 10% to 5% relative MDE the users needed roughly quadruple; at 2.5% they are about 16 times larger. Lower the baseline to 2%: everything shifts up (rare events are expensive). Raise the power to 90% or lower α to 0.01 and watch the curve lift.

The total traffic is fixed at 29 502 users (the 10% → 11% plan). Drag the purple dot to change the share sent to treatment. Power peaks at 80% with a 50/50 split and falls to about 40% with 90/10. The readout shows how much more total traffic an unequal split needs to get the same power back.

"The MDE is the effect we expect."

The MDE is the smallest effect worth detecting (or the smallest the budget allows). If you set it to an optimistic guess, you only have 80% power for that optimistic effect and much less for realistic smaller ones.

"After the test was not significant, we computed its power from the observed effect: only 20%, so the test was underpowered."

"Observed power" is a one-to-one function of the p-value and adds no information. Power is a planning quantity: compute it for the MDE before the test. After the test, report the CI instead.

"5% MDE" (without saying relative or absolute).

On a 10% baseline, a 5% relative MDE is 0.5 points (57 763 per arm); a 5-point absolute MDE is 10% → 15% (about 700 per arm). Always state which.

"MDE is the minimum effect the test will detect."

"MDE is the smallest true effect the test detects with the planned power (say 80%) at the chosen α. Smaller true effects can still come out significant, just less often; effects at the MDE are missed 20% of the time."

Model answer: "I agree the MDE with the business as the smallest effect worth shipping, take the baseline and variance from history, and compute $n \propto (z_{1-\alpha/2}+z_{1-\beta})^2\sigma^2/\delta^2$. Halving the MDE quadruples the sample. If that is unaffordable, I change the metric, use variance reduction or triggering, or do not run the test."

$n_{arm} = 2(z_{1-\alpha/2} + z_{1-\beta})^2\sigma^2/\delta^2$ (means); two-proportion version above. Rule of thumb $16\sigma^2/\delta^2$.

10% → 11% (α 0.05, power 0.8): 14 751 per arm. Half the MDE ≈ 4× the users. 50/50 is most efficient; 90/10 costs 2.78× traffic.

Traps: MDE ≠ expected effect; relative vs absolute; post-hoc "observed power".

Quick check: a metric's sd doubles (a noisier metric) while the MDE stays the same. What happens to $n$?

$n \propto \sigma^2$, so it quadruples. This is why reducing variance (CUPED, triggering, capping outliers) is worth as much as finding more traffic.

From sample size to test duration

A sample size is a number of users; a test runs in days. Traffic converts one into the other: if 3 000 eligible users arrive per day and the plan needs 29 502, the test needs about 10 days.

Then round up to whole weeks. Monday users are not Saturday users: shopping habits, work patterns and traffic sources change through the week. A test that stops on a Wednesday has seen some weekdays twice and others once, so the users it measured are not a fair picture of a normal week. Two full weeks is a common minimum (a rule of thumb, not a law).

Three ways to say it:

  • Picture: the sample size is the bucket; daily traffic is the tap; duration is how long the bucket takes to fill, rounded up to whole weeks.
  • Numbers: 29 502 users ÷ 3 000 per day = 9.8 days → run 14 days.
  • Slogan: days = users needed ÷ users per day, then round up to full weeks.
  1. From the previous section: 14 751 users per arm, 2 arms → $2 \times 14\,751 = 29\,502$ users.
  2. Eligible traffic: 3 000 new users per day, 100% of them in the experiment.
  3. Days: $29\,502/3\,000 = 9.83$ → 10 days → round up to whole weeks: 14 days.
  4. Bonus: 14 days give $14 \times 3\,000 = 42\,000$ users, 21 000 per arm, so the test can actually detect a smaller effect than planned: MDE ≈ 0.84 points (8.4% relative) instead of 1 point.
  5. If only half the traffic can be given to this test (another test runs in parallel): $29\,502/1\,500 = 19.7$ days → 3 weeks.
  6. With three arms (two treatments): $3 \times 14\,751 = 44\,253$ users → 14.8 days at 3 000/day → 3 weeks (and a correction for two comparisons would raise $n$ further).
$$\text{days} = \left\lceil \frac{k \times n_{arm}}{\text{eligible users per day} \times \text{share in the experiment}} \right\rceil, \quad \text{then round up to whole weeks.}$$
  • $k$ = number of arms; $n_{arm}$ = users per arm from the sample-size formula; $\lceil\cdot\rceil$ = round up.
  • Minimum duration: at least one full weekly cycle, often two (rule of thumb), to cover weekday/weekend patterns and let early novelty effects fade (Chapter 5.11).
  • Maximum duration: long tests suffer from cookie churn (users reappearing as new), seasonal shifts, and the cost of not shipping a good change.
  • Unique users vs visits: returning users mean the number of distinct users grows slower than "daily visitors × days". Use historical counts of unique users over the planned window.
  • The duration is fixed in the plan. Stopping early because the result "looks significant" is peeking, which inflates false positives (Chapter 5.11).
Why do we need it?

Teams need to know when a decision will come, whether the page has enough traffic for the test at all, and how many tests can run in parallel. The duration also protects the result from day-of-week bias.

Where is it used?

Experiment planning and roadmaps, traffic allocation between concurrent tests, deciding between a 2-arm and a 4-arm test, and the stopping rule written into the analysis plan.

How is it used?

Compute $n_{arm}$, multiply by the arms, divide by the daily eligible (unique) users times the allocated share, round up to whole weeks. If the result is too long, go back and raise the MDE, change the metric, or reduce variance.

Set the baseline, the relative MDE you care about, the daily eligible users, the share of traffic this test gets and the number of arms. The readout turns the plan into days and whole weeks. The blue curve shows the MDE the test could detect after each number of days (purple dashed = your target MDE); the dot marks the full-week duration you would actually run. Drop the traffic to 500 per day: the test needs months, and the curve shows what is realistic in 4 weeks.

"We reached the sample size on Wednesday of week 2, so we stop now."

Finish the full week, so every weekday is represented equally. The planned duration (in whole weeks) is part of the design.

"Not significant yet, let's run one more week and check again."

Extending a test because it is not yet significant, and checking repeatedly, is a form of peeking; the false-positive rate climbs above α. If extensions might be needed, plan a sequential design in advance (Chapter 5.11).

A Bayesian analysis like yours does not remove the need for a planned duration. Posterior probabilities are valid summaries of the data you have, but a habit of stopping "as soon as $P(\theta_B \gt \theta_A \mid D) \gt 0.95$" while checking every day ships changes that do nothing more often than a single planned look would. Planning the duration from traffic, and deciding on the full weeks, keeps the decision rule's error rates close to what you designed (simulated in the last section).

days $= \lceil k \cdot n_{arm} / (\text{eligible users/day} \times \text{share}) \rceil$ → round up to whole weeks (≥ 1–2 weeks, rule of thumb).

Example: 29 502 users at 3 000/day → 9.8 days → 14 days (detectable MDE then ≈ 8.4% relative).

Traps: stopping mid-week, extending until significant, counting visits instead of unique users.

Quick check: the plan needs 60 000 users per arm, 2 arms, and the page gets 4 000 eligible users a day. How long should the test run?

$120\,000/4\,000 = 30$ days → round up to whole weeks: 35 days (5 weeks). If that is too long, revisit the MDE or the metric before launch.

Statistical vs practical significance core

"Statistically significant" means only "probably not exactly zero". With millions of users, a change worth +0.03 points can be significant, and still not pay for the engineering, the maintenance and the risk. And with too few users, a change worth +2 points can be "not significant".

Practical significance asks a different question: is the effect big enough to matter? Before the test, the business sets a "worth it" threshold (often equal to the MDE). After the test, compare the whole confidence interval with both zero and that threshold. Four very different situations appear, and a p-value alone cannot tell them apart.

Three ways to say it:

  • Picture: where does the interval sit relative to two lines, zero and the threshold?
  • Numbers: CI [+0.06, +0.44] points with threshold +0.5: real, but too small to be worth shipping.
  • Slogan: significant tells you "not zero"; the CI vs the threshold tells you "worth it?".

Baseline 10%, threshold "worth shipping" = +0.5 points. Four possible results (95% CIs for the effect, in points):

  1. CI [+0.63, +1.37]: whole interval above +0.5 → significant and important: ship.
  2. CI [+0.06, +0.44]: above 0 but entirely below +0.5 → significant but too small: probably not worth it.
  3. CI [−0.78, +1.58]: contains 0 and also values above +0.5 → inconclusive: the test was too small; we cannot rule out a worthwhile effect.
  4. CI [−0.19, +0.19]: contains 0 but lies entirely below +0.5 → confidently no meaningful effect: a useful negative result, not a failure.

Cases 3 and 4 have similar p-values (both "not significant"), yet they lead to opposite actions: "collect more data or move on" vs "this change does not matter".

  • Practical significance threshold $\delta_{min}$: the smallest effect worth acting on, set from costs, risks and strategy before the test. The MDE is usually chosen equal to (or below) it, so the test can detect it.
  • Decision zones by the CI $[L, U]$: $L \ge \delta_{min}$ ship · $0 \lt L$ and $U \lt \delta_{min}$ real but too small · $L \le 0 \le U$ and $U \ge \delta_{min}$ inconclusive · $L \le 0 \le U \lt \delta_{min}$ no meaningful effect · $U \lt 0$ harmful.
  • Equivalence / non-inferiority: formal versions of "confidently small" and "not worse than a margin": test that the CI lies inside $(-\delta, +\delta)$ (equivalence, e.g. the TOST procedure) or above $-\delta$ (non-inferiority, used for guardrails). A TOST at 5% uses the matching 90% interval; checking a 95% interval is a little stricter.
Why do we need it?

Shipping every significant change wastes engineering on trivia, and dropping every non-significant change throws away good ideas tested on too few users. Comparing the interval with a business threshold turns statistics into a decision.

Where is it used?

Launch reviews, experiment scorecards that show CIs against a "minimum worthwhile" line, non-inferiority guardrails (latency, cost), model rollouts ("the new model must be at least 0.5% better to justify its cost"), and Bayesian decisions like $P(\text{lift} \gt \delta \mid D)$.

How is it used?

Write $\delta_{min}$ into the plan and set the MDE accordingly. After the test, draw the CI next to 0 and $\delta_{min}$, name which of the zones it falls in, and act on that, reporting the effect size, not just the p-value.

0 +0.5 (worth it) 1. ship 2. real, too small 3. inconclusive 4. no meaningful effect
Four results judged by where the whole 95% interval sits. Only case 1 clears the practical threshold (purple). Case 2 is significant but small; cases 3 and 4 are both "not significant", but only case 4 actually rules out a worthwhile effect.

Drag the dot (the estimated effect) left and right; change the users per arm to make the interval wider or narrower; move the "worth it" threshold. The readout names the decision zone. Use the four buttons for the textbook cases. Notice how the same estimate of +0.3 points is "inconclusive" with 5 000 users per arm but "real but too small" with 200 000.

"p < 0.001, so the effect is large."

A tiny p-value can come from a tiny effect measured on a huge sample. Size comes from the estimate and its CI, not from p.

"Not significant, so the change has no effect."

Look at the interval: if it still contains worthwhile effects, the honest answer is "we don't know yet" (inconclusive). Only an interval that excludes the threshold supports "no meaningful effect".

The Bayesian version of this section is the decision rule $P(\theta_B - \theta_A \gt \delta \mid D) \gt$ some level, with $\delta$ the practical threshold, which is more useful than $P(\theta_B \gt \theta_A \mid D)$: a change can be almost surely better than control and still almost surely not better by enough to matter (Chapter 6.4). In a framework like yours, both quantities come from the same posterior draws: count the share of draws where the lift exceeds $\delta$.

"The result is statistically significant, so we should ship it."

"The 95% CI for the lift is +0.06 to +0.44 points; our threshold for paying back the engineering cost was +0.5, so it is real but not worth it on its own."

Model answer: "Statistical significance only says the effect is probably not zero. Practical significance compares the effect, with its uncertainty, to the smallest effect that matters for the business. I set that threshold before the test, plan the MDE to match it, and decide from where the whole confidence interval (or the posterior) sits relative to zero and the threshold."

Set $\delta_{min}$ (worth it) before the test; MDE ≤ $\delta_{min}$.

Read the whole CI: above $\delta_{min}$ → ship; in $(0, \delta_{min})$ → real but small; contains 0 and $\delta_{min}$ → inconclusive; contains 0 but below $\delta_{min}$ → no meaningful effect; below 0 → harmful.

Trap: "significant" ≠ "important"; "not significant" ≠ "no effect".

Quick check: CI [+0.1, +2.3] points, threshold +0.5. Which zone, and what would you do?

Significant (0 is outside), but the interval straddles the threshold: the effect is real, and it may or may not be worth it. If the change is cheap, shipping may be fine; otherwise, more data (a longer or follow-up test) would narrow the interval.

Designing a Bayesian A/B test, and the one-page experiment plan core

A Bayesian analysis changes how you read the result: instead of a p-value you get probabilities about the conversion rates, such as $P(\theta_B \gt \theta_A \mid D)$. It does not change the design questions: which unit to randomize, which metric decides, how big an effect matters, how many users and how many weeks.

You can plan a Bayesian test the same way you planned the classical one, by simulation: pretend the truth is "B is better by +1 point", simulate many experiments of the planned size, run your exact decision rule on each, and count how often it ships. Then pretend the truth is "no difference" and count how often it ships anyway. Those two numbers are the Bayesian design's power and false-ship rate.

Three ways to say it:

  • Picture: a flight simulator for your decision rule, before real users are spent.
  • Numbers: 14 750 users per arm, ship if $P(B \gt A) \gt 0.95$: ships about 88% of the time when B is truly +1 point better, and about 5% when B is no better.
  • Slogan: Bayesian analysis, same design discipline.

Flat priors Beta(1, 1) for both rates, baseline 10%, decision rule "ship B if $P(\theta_B \gt \theta_A \mid D) \gt 0.95$", one look at the end, 14 750 users per arm.

  1. Simulate a world where B is truly 11%: draw $k_A \sim Binomial(14\,750, 0.10)$ and $k_B \sim Binomial(14\,750, 0.11)$.
  2. Compute the posteriors $Beta(1 + k_A, 1 + n - k_A)$ and $Beta(1 + k_B, 1 + n - k_B)$ and the probability $P(\theta_B \gt \theta_A \mid D)$ (Monte Carlo draws, or a Normal approximation of the two Betas at this size).
  3. Ship if it exceeds 0.95. Repeat 20 000 times: the rule ships in about 88% of worlds. (Its classical cousin: a one-sided z-test at α = 0.05 has power 87.6% here.)
  4. Simulate a world where B is also 10%: the rule ships in about 5% of worlds, the false-ship rate.
  5. Why so close to the classical numbers? With flat priors and large $n$, $P(\theta_B \gt \theta_A \mid D) \approx \Phi(z)$, so "posterior probability above 0.95" behaves like "one-sided z above 1.645". This match is approximate, and it holds only for a single planned look.
  • Design by simulation (pre-posterior analysis): fix the prior, the decision rule and the sample size; simulate data under chosen "true" scenarios; apply the rule; record its operating characteristics: $P(\text{ship} \mid \text{effect} = \delta)$ (Bayesian power), $P(\text{ship} \mid \text{no effect})$ (false-ship rate), and if you like the expected loss of the decision.
  • Averaging the ship probability over a prior for the true effect (instead of a single assumed effect) gives what is often called assurance.
  • The same design elements as the classical test still apply: randomization unit = analysis unit (or a hierarchical layer), triggering, a primary metric, a practical threshold $\delta$, a planned duration in whole weeks, and a stopping rule.
  • The one-page experiment plan (written before launch): hypothesis · randomization unit and analysis unit · eligibility and trigger · primary, secondary and guardrail metrics · practical threshold and MDE · sample size, allocation and duration · analysis method (test or model, prior, decision rule) · stopping rule · health checks (sample ratio, A/A, guardrails; Chapter 5.11).
Why do we need it?

A posterior threshold like 0.95 has no guaranteed error rate by itself; its behaviour depends on the prior, the sample size and how often you look. Simulation tells you how often your rule ships good and useless changes, which is what stakeholders actually need to know.

Where is it used?

Planning Bayesian A/B tests (Beta-Binomial, Dirichlet-Multinomial, Normal/Student-t models), Bayesian adaptive clinical trials (regulators ask for simulated operating characteristics), and checking any custom decision rule before trusting it.

How is it used?

Write the decision rule exactly as the framework will compute it. Simulate a few hundred to a few thousand experiments under "no effect" and under "effect = MDE"; tune $n$ (or the threshold) until both rates are acceptable; record them in the plan next to the duration.

Experiment plan: new checkout page (written before launch) Hypothesis: the one-page checkout raises purchase conversion. Units: randomize users (hash of user id + "checkout-v2"); analyse per user. Trigger: users who reach checkout, logged in both arms. Metrics: primary = conversion; secondary = checkout starts; guardrails = load time, errors. Threshold / MDE: +1 point (10% → 11%), α = 0.05, power 80%. Size and time: 14 751 per arm, 50/50, 3 000 users/day → 14 days (2 full weeks). Analysis: two-proportion test, or Beta(1,1) priors with "ship if P(lift > 0) > 0.95". Stopping rule: one look at day 14 (no peeking); stop early only if a guardrail breaks. Health checks: sample ratio test, A/A test of the pipeline, guardrails (Chapter 5.11).
Every design decision of this chapter on one page. Writing it before launch is what protects the result from "deciding after seeing the data".

Each simulated experiment has the chosen users per arm, Beta(1, 1) priors and the rule "ship if $P(\theta_B \gt \theta_A \mid D)$ > threshold". Orange: worlds where B is truly better by the chosen lift; blue: worlds where B is no better (baseline 10% in both). The histograms show the posterior probabilities the experiments produced; the purple line is the threshold. Press Simulate 200 more. With 14 750 per arm and +1 point, about 88% of orange worlds ship and about 5% of blue ones do. Cut the users to 3 000: the orange worlds spread out and many good changes are not shipped. Lower the threshold to 0.80: more good changes ship, and more useless ones too.

"Bayesian A/B tests do not need a sample size."

They need one for the same reason classical tests do: too little data gives a wide posterior and a rule that rarely ships good changes. Plan $n$ by simulating the decision rule.

"The 0.95 threshold means only 5% of shipped changes are useless."

The false-ship rate in a no-effect world is about 5% only for flat priors, a single look and large samples. And the share of shipped changes that are useless also depends on how many of your ideas work at all. Simulate to know.

"Bayesian methods are immune to peeking."

The posterior is a valid summary at any time, but a rule "stop the moment $P \gt 0.95$, checking daily" ships no-effect changes far more often than 5%. Peeking is a property of the stopping rule (Chapter 5.11).

This is the design side of your Bayesian experimentation framework. Its Beta-Binomial model and a decision quantity like $P(\theta_B \gt \theta_A \mid D)$ (or, better, $P(\theta_B - \theta_A \gt \delta \mid D)$) can be wrapped in exactly this simulation to report, for any planned traffic, how often the framework would ship a +$\delta$ improvement and how often it would ship a change that does nothing. With partial pooling across segments or more complex likelihoods, the same recipe works; you simulate from the full model and fit it to each simulated dataset (with SVI that can be costly, so a few hundred simulations or a conjugate shortcut for the main metric is common). It is also the clearest answer to the interview question "how do you size a Bayesian test?".

"We use Bayesian A/B testing, so we don't worry about sample size or peeking."

"Our analysis is Bayesian, but the design is planned the same way: we fix the unit, metric, threshold and duration, and we size the test by simulating the decision rule's ship rate under no effect and under the minimum effect we care about."

Model answer: "Bayesian inference gives direct probability statements about the lift, but the frequency with which a decision rule makes wrong calls still depends on sample size, priors and stopping rules. I report those operating characteristics, estimated by simulation, next to the plan."

Bayesian design = simulate: prior + decision rule + n → P(ship | effect = δ) and P(ship | no effect).

Flat priors, large n, one look: "P(B > A) > 0.95" ≈ one-sided z-test at 5% (here ≈ 88% power, ≈ 5% false ships).

The one-page plan: hypothesis, units, trigger, metrics, threshold/MDE, n and duration, analysis and decision rule, stopping rule, health checks. Trap: "Bayesian means no sample size / no peeking problem".

Quick check: in the simulator, what happens to the false-ship rate if you keep the threshold at 0.95 but raise the users per arm to 40 000?

It stays near 5%: under no effect, $P(\theta_B \gt \theta_A \mid D)$ is roughly uniform between 0 and 1 whatever the sample size, so about 5% of experiments exceed 0.95. What changes with more users is the power: the orange histogram moves toward 1 and almost every +1-point world ships.

Recap, cheat sheet and practice

  • An A/B test compares a control and a treatment; the treatment effect is the difference in average outcome, estimated by $\bar Y_B - \bar Y_A$ with its CI. Say whether it is absolute (points) or relative (%).
  • Random assignment (in practice: hash of unit id + experiment salt, mod 100) makes the arms comparable on average in every trait; opt-in or calendar assignment creates bias that more data never removes.
  • Compute the SE at the randomization unit. Several correlated rows per unit inflate the variance by $DE = 1 + (m-1)\rho$; ignoring it gives too-narrow CIs (m = 5, ρ = 0.3: a "95%" CI covers ≈ 81%).
  • Exposure and triggering: analyse users who reached the point where the arms differ, logged identically in both arms; unaffected users dilute the effect (overall = q × triggered) but still add noise.
  • One primary metric decides; secondary metrics explain; guardrails must stay within tolerances. Noisy metrics need more users: $n \approx 15.7\,CV^2/r^2$ per arm.
  • Sample size from α, power and the MDE: 10% → 11% needs 14 751 per arm; half the MDE ≈ 4× the users; 50/50 is the most efficient split.
  • Duration = users needed ÷ eligible users per day, rounded up to whole weeks; fixed in advance.
  • Practical significance: compare the whole CI with 0 and a pre-set "worth it" threshold (ship / real but small / inconclusive / no meaningful effect / harmful).
  • Bayesian tests need the same design; size them by simulating the decision rule's ship rate under "effect = δ" and "no effect".

Cheat sheet

IdeaFormula / ruleIn words
Treatment effect$\hat\tau = \bar Y_B - \bar Y_A$, $SE = \sqrt{s_A^2/n_A + s_B^2/n_B}$difference of arm averages, with its noise
Relative lift$\hat\tau/\bar Y_A$its CI needs a ratio method (delta method, log ratio)
Hash assignmentbucket = hash(id + salt) mod 100sticky, reproducible, independent across tests
Design effect$DE = 1 + (m-1)\rho$, $n_{eff} = n_{rows}/DE$correlated rows carry less information
Dilutionoverall effect = $q \times$ triggered effecttrigger to save traffic
Metric sensitivity$n \approx 2(z_{1-\alpha/2}+z_{1-\beta})^2 CV^2/r^2 = 15.7\,CV^2/r^2$noisy metrics are expensive
Two-proportion n$\frac{(z_{1-\alpha/2}\sqrt{2\bar p\bar q} + z_{1-\beta}\sqrt{p_1q_1+p_2q_2})^2}{(p_2-p_1)^2}$10% → 11%: 14 751 per arm
MDE from n$MDE \approx (z_{1-\alpha/2}+z_{1-\beta})\sqrt{2\sigma^2/n}$half the MDE ⇒ ≈ 4× n
Allocation$Var \propto 1/f + 1/(1-f)$90/10 needs 2.78× the traffic of 50/50
Duration$\lceil k\,n_{arm}/(\text{users/day} \times \text{share})\rceil$ → whole weeks29 502 at 3 000/day → 14 days
Practical significanceCI vs 0 and $\delta_{min}$"significant" ≠ "worth it"
Bayesian designsimulate P(ship | δ), P(ship | 0)flat priors, one look: P(B>A) > 0.95 ≈ one-sided 5% test
Code it · Python

import numpy as np
from scipy import stats
from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize

# 1) Sample size per arm for 10% -> 11% (the two-proportion formula of this chapter)
def n_two_prop(p1, p2, alpha=0.05, power=0.8):
    za, zb = stats.norm.ppf(1 - alpha / 2), stats.norm.ppf(power)
    pbar = (p1 + p2) / 2
    num = za * np.sqrt(2 * pbar * (1 - pbar)) + zb * np.sqrt(p1 * (1 - p1) + p2 * (1 - p2))
    return int(np.ceil(num ** 2 / (p2 - p1) ** 2))

print(n_two_prop(0.10, 0.11), n_two_prop(0.10, 0.105))   # 14751 57763  (half the MDE: ~4x users)
h = proportion_effectsize(0.11, 0.10)                     # statsmodels uses Cohen's h (arcsine scale)
print(round(NormalIndPower().solve_power(h, alpha=0.05, power=0.8)))   # 14744 (a slightly different approximation)

# 2) Duration from traffic, and the MDE you can detect in whole weeks
n_arm, per_day = n_two_prop(0.10, 0.11), 3000
days = int(np.ceil(2 * n_arm / per_day)); run_days = 7 * int(np.ceil(days / 7))
n_run = run_days * per_day // 2
mde = next(d / 10000 for d in range(1, 5000) if n_two_prop(0.10, 0.10 + d / 10000) <= n_run)
print(days, run_days, n_run, mde)       # 10 14 21000 0.0084  (0.84 points = 8.4% relative)

# 3) Metric sensitivity: n per arm ~ 2 (z_a + z_b)^2 CV^2 / r^2
K = 2 * (stats.norm.ppf(0.975) + stats.norm.ppf(0.8)) ** 2
p, mu, sd, r = 0.05, 60, 80, 0.05
cv_conv = np.sqrt((1 - p) / p)
rev_mean, rev_var = p * mu, p * (sd ** 2 + mu ** 2) - (p * mu) ** 2
cv_rev = np.sqrt(rev_var) / rev_mean
print(round(K, 2), round(K * cv_conv ** 2 / r ** 2), round(K * cv_rev ** 2 / r ** 2))   # 15.7 119303 342560

# 4) Unit mismatch: A/A tests randomized per user, analysed per session vs per user
rng = np.random.default_rng(0)
U, m, rho, R = 80, 5, 0.3, 3000
cover_sess = cover_user = 0
for _ in range(R):
    arms = [np.sqrt(rho) * rng.normal(size=(U, 1)) + np.sqrt(1 - rho) * rng.normal(size=(U, m)) for _ in range(2)]
    d = arms[1].mean() - arms[0].mean()
    se_sess = np.sqrt(sum(a.var(ddof=1) / a.size for a in arms))            # pretends 400 independent rows
    se_user = np.sqrt(sum(a.mean(axis=1).var(ddof=1) / U for a in arms))   # one value per user
    cover_sess += abs(d) < 1.96 * se_sess
    cover_user += abs(d) < 1.96 * se_user
de = 1 + (m - 1) * rho
print(de, round(2 * stats.norm.cdf(1.96 / np.sqrt(de)) - 1, 3), cover_sess / R, cover_user / R)
# 2.2 0.814 0.8063333333333333 0.947   (per-session "95%" CIs cover ~81%; per-user ~95%)

# 5) Bayesian decision rule "ship if P(theta_B > theta_A | D) > 0.95", flat Beta(1,1) priors, one look
def ship_rate(pA, pB, n, sims=4000, draws=4000, thr=0.95):
    kA, kB = rng.binomial(n, pA, sims), rng.binomial(n, pB, sims)
    ships = 0
    for a, b in zip(kA, kB):
        tA = rng.beta(1 + a, 1 + n - a, draws)
        tB = rng.beta(1 + b, 1 + n - b, draws)
        ships += (tB > tA).mean() > thr
    return ships / sims

print(ship_rate(0.10, 0.11, 14751), ship_rate(0.10, 0.10, 14751))
# 0.8695 0.055   (Bayesian "power" ~87-88%; false-ship rate ~5%, up to Monte Carlo noise)
Test yourself

1. Users are randomized, but the click-through rate is analysed per page view with a binomial SE over all page views. What goes wrong?

Rows from the same randomization unit are positively correlated: the variance is inflated by the design effect $1 + (m-1)\rho$, which the binomial formula ignores. Use per-user metrics, the delta method or cluster-robust SEs.

2. A plan needs 20 000 users per arm for a 4% relative MDE. Roughly how many for a 2% relative MDE (same baseline, α, power)?

$n \propto 1/\delta^2$: halving the MDE multiplies $n$ by about 4.

3. Which trigger condition is safe for a test of a redesigned checkout page?

A trigger must be defined identically in both arms and must not be affected by the treatment. Clicking a treatment-only button or buying are outcomes of the treatment; "treatment users who saw it" has no control counterpart.

4. 95% CI for the lift: [−0.2, +0.2] points. The "worth it" threshold is +0.5 points. What is the best summary?

0 is inside (not significant), but the whole interval lies below +0.5, so worthwhile effects are ruled out. That is a useful negative result, unlike an interval such as [−0.8, +1.6], which is inconclusive.

5. For a fixed total traffic, which allocation gives the most power for a two-arm test?

$Var(\hat\tau) \propto 1/f + 1/(1-f)$ is smallest at $f = \tfrac12$. A 90/10 split needs about 2.78 times the total traffic for the same power (sometimes worth it for risk, but it costs).

6. Flat Beta(1, 1) priors, a large planned sample, one look at the end, rule "ship if $P(\theta_B \gt \theta_A \mid D) \gt 0.95$". If B is truly no better, how often does the rule ship?

Under no effect, $P(\theta_B \gt \theta_A \mid D)$ is roughly uniform across experiments, so about 5% exceed 0.95; the rule then behaves like a one-sided test at 5%. With daily peeking the rate would be much higher.

Practice problems

A. Baseline conversion 20%. How many users per arm to detect +2 points (20% → 22%) with α = 0.05 (two-sided) and power 80%?
  1. $\bar p = 0.21$. $\sqrt{2 \times 0.21 \times 0.79} = \sqrt{0.3318} = 0.5760$; times 1.960 = 1.1290.
  2. $\sqrt{0.20 \times 0.80 + 0.22 \times 0.78} = \sqrt{0.1600 + 0.1716} = \sqrt{0.3316} = 0.5758$; times 0.8416 = 0.4846.
  3. Sum $= 1.6136$, squared $= 2.6037$, divided by $0.02^2 = 0.0004$: $n = 6\,509.3$ → 6 510 per arm.
  4. Lehr's rule check: $16 \times 0.2 \times 0.8/0.0004 = 6\,400$. Close.
B. With 6 510 per arm (2 arms), 1 200 eligible users per day, and 60% of traffic given to this test, how long should it run?

Users needed: $2 \times 6\,510 = 13\,020$. Users per day in the test: $1\,200 \times 0.6 = 720$. Days: $13\,020/720 = 18.1$ → 19 days → round up to whole weeks: 21 days. If that is too long, give the test 100% of the traffic ($13\,020/1\,200 = 10.9$ → 14 days) or raise the MDE.

C. 2 000 users per arm, 8 page views each, intra-class correlation 0.1. Someone analyses the 16 000 page views per arm as independent. How wrong is the SE, and what is the real coverage of their 95% CI?
  1. $DE = 1 + (8 - 1) \times 0.1 = 1.7$. Effective sample: $16\,000/1.7 \approx 9\,412$ independent page views' worth per arm.
  2. The naive SE is too small by $\sqrt{1.7} = 1.30$.
  3. Real coverage: $2\Phi(1.96/1.30) - 1 = 2\Phi(1.50) - 1 \approx 0.87$. Their "95%" interval is an 87% interval, and an A/A test would flag about 13% of the time.
D. A new search-results layout is shown only to users who search (10% of visitors). Among searchers it raises conversion by 4% relative. If non-searchers are unaffected and convert at the same baseline, what is the site-wide relative lift, and what should be reported?

Site-wide lift ≈ $0.10 \times 4\% = 0.4\%$ relative. Report both: "+4% among searchers (the triggered population), which is about +0.4% site-wide". Analysing only searchers (triggered on "performed a search", logged in both arms) is what makes the 4% detectable at all; analysing all visitors would need roughly ten times more traffic if non-searchers add similar noise.

E. Threshold +0.5 points. Classify: (i) [+0.7, +1.9]; (ii) [−1.0, +2.0]; (iii) [+0.05, +0.35]; (iv) [−0.9, −0.1]; (v) [+0.2, +1.2].
  • (i) significant and above the threshold → ship.
  • (ii) contains 0 and the threshold → inconclusive (underpowered).
  • (iii) significant but entirely below +0.5 → real but too small.
  • (iv) entirely below 0 → significantly harmful.
  • (v) significant, straddles the threshold → real, size unclear; decide by cost or collect more data.
F. Interview: "Design an A/B test for a new 'customers also bought' widget on product pages. Your company uses a Bayesian framework."

"Hypothesis: the widget raises purchases per visitor. Unit: randomize users by hashing the user id with the experiment name; analyse per user (sessions within a user are correlated). Trigger: users who view a product page where the widget would render, logged in both arms. Metrics: primary = purchase conversion of triggered users; secondary = widget clicks, items per order; guardrails = page load time, return rate, revenue per user. Threshold and MDE: agree the smallest worthwhile lift, e.g. +3% relative, from the build and maintenance cost. Size and duration: from the baseline and traffic, compute users per arm, divide by daily triggered users, round up to whole weeks. Analysis: a Beta-Binomial model with weakly informative priors; decision 'ship if $P(\text{lift} \gt 3\% \mid D) \gt 0.9$ and guardrails are not worse than their tolerances'. Before launch I would simulate that rule to check how often it ships a +3% widget and a useless one, and I would fix the duration and avoid peeking. During the test: sample ratio check and guardrail monitoring (Chapter 5.11)."

Chapter 5.11 · Syllabus Module 17

A/B testing pitfalls: SRM, peeking, interference, multiple testing

Chapter 5.10 showed how to design a clean experiment. This chapter is about the ways a real experiment quietly stops being clean: users go missing from one group, groups leak into each other, effects fade with time, users affect each other, people look at the dashboard every morning, and twenty metrics are tested at once. Each trap has a simple picture, a simple check, and a simple fix. Most of them matter just as much for a Bayesian framework like yours as for a p-value.

  • Run a sample ratio mismatch (SRM) check with a chi-square test, and explain why a failed check means "do not read the results"
  • Spot attrition and contamination, and know why intention-to-treat analysis is the safe default
  • Recognise novelty, primacy and carryover effects in a plot of the effect over time
  • State SUTVA precisely, and explain how interference (marketplaces, social networks) biases a naive A/B test
  • Show by simulation why repeated peeking inflates false positives (5% becomes about 28% with 30 daily looks), and how sequential tests fix it
  • Say exactly what peeking does and does not do to a Bayesian decision rule such as "ship B when $P(B \gt A \mid D) \gt 0.95$"
  • Compute the family-wise error rate $1-(1-\alpha)^m$ and apply Bonferroni, Holm and Benjamini–Hochberg

What we need from earlier chapters: the logic of a hypothesis test, the test statistic, the p-value and the significance level $\alpha$ (Chapter 5.6); false positives, power and sample size (Chapter 5.7); the chi-square goodness-of-fit test (Chapter 5.9); control, treatment, randomization unit, exposure and triggering (Chapter 5.10); the "at least one" rule $1-(1-p)^m$ (Chapter 4.2). For the Bayesian parts: a posterior is your updated belief about an unknown after seeing data (Chapter 6.1), and $P(\theta_B \gt \theta_A \mid D)$ is the posterior probability that B's true rate beats A's (Chapter 6.4). Words used all chapter: an arm (or variant) is one of the groups of an experiment; an A/A test is an experiment where both arms get exactly the same experience, so any "winner" it finds is a false positive; "pts" means percentage points (10% → 12% is +2 pts).

Sample ratio mismatch: check the counts before the conversions core

You plan to deal 10,000 users into two piles, half and half, by a fair coin flip for each user. A fair coin never gives exactly 5,000 and 5,000, but it gives something close: 5,050 and 4,950 is normal wobble.

Now imagine the piles come out 5,200 and 4,800. A fair coin almost never does that with 10,000 flips. So either the coin is broken, or something after the coin is eating users from one pile: a slow page, a crash on some phones, a logging bug, a bot filter. And the users who were eaten are rarely a random handful. The two piles are no longer two fair halves of the same crowd, so comparing them is no longer fair.

A sample ratio mismatch (SRM) is exactly this: the split you see does not match the split you planned, by more than chance allows. It is the smoke alarm of experimentation.

Three ways to say it:

  • Picture: two buckets fill from one tap through a fair splitter; if one bucket is clearly lighter, a pipe is leaking, and you do not know what leaked out.
  • Numbers: 5,200 vs 4,800 out of a planned 50/50 gives $\chi^2 = 16$ and $p \approx 0.00006$: chance alone almost never does this.
  • Slogan: check the counts before you check the conversions.

An SRM check by hand. The plan was 50% of users to A and 50% to B. After a week the logs show 5,200 users in A and 4,800 in B (10,000 in total).

  1. Expected counts under the plan: $10{,}000 \times 0.5 = 5{,}000$ in each arm.
  2. Gaps: A has $5{,}200 - 5{,}000 = +200$; B has $4{,}800 - 5{,}000 = -200$.
  3. Squared gap divided by the expected count, for each arm: $200^2/5{,}000 = 40{,}000/5{,}000 = 8$, and the same $8$ for B.
  4. Add them: $\chi^2 = 8 + 8 = 16$. With 2 groups there is $2 - 1 = 1$ degree of freedom.
  5. p-value: the chance that a chi-square variable with 1 degree of freedom is at least 16 is about $0.00006$ (the same as $|z| \ge 4$ for a Normal).
  6. Verdict: a fair 50/50 coin would produce a split this lopsided about 6 times in 100,000 experiments. Something is wrong. Do not read the metrics; find the cause first.

Compare a healthy run: 5,050 vs 4,950. Then $\chi^2 = 50^2/5{,}000 + 50^2/5{,}000 = 0.5 + 0.5 = 1$ and $p \approx 0.32$: normal wobble.

Plan: each unit goes to group $g$ with probability $q_g$ (for example $q_A = q_B = 0.5$). After the experiment we see counts $O_g$, with total $N = \sum_g O_g$. The expected counts are $E_g = N q_g$.

A sample ratio mismatch is a split of units between groups that differs from the planned split by more than chance allows. We test for it with the chi-square goodness-of-fit test (Chapter 5.9):

$$\chi^2 = \sum_{g} \frac{(O_g - E_g)^2}{E_g}, \qquad \text{compared with a chi-square distribution with (number of groups} - 1) \text{ degrees of freedom.}$$
  • The check uses only the counts of units, never the outcomes. You can (and should) run it before looking at any metric.
  • Because the check runs on every experiment, many teams raise the alarm only at a strict threshold such as $p \lt 0.001$, to avoid constant false alarms. This is a rule of thumb, not a law.
  • Typical causes: a bug in assignment; a redirect or extra page load in one arm that loses impatient users; crashes or slow loading on some devices in one arm; bots or fraud filters that treat the arms differently; logging that only fires after something the treatment changes; data joins that drop users with missing fields; changing the split part-way through (ramping from 10/90 to 50/50) and then adding up the counts.
Why do we need it?

If users vanish from one arm for a reason linked to the treatment, the arms stop being comparable, and nothing done afterwards (a better test, a Bayesian model, more data) can repair that. A two-number count check catches many such bugs before anyone makes a decision on broken data.

Where is it used?

Experimentation platforms at large tech companies run it automatically on every experiment and hide or flag the scorecard when it fails (Microsoft and LinkedIn have both published papers on diagnosing SRM). It is also the first item on most "A/B test health" checklists, together with A/A tests.

How is it used?

Before reading any metric, run scipy.stats.chisquare(observed, f_exp=expected) on the user counts per arm. If p is tiny, stop: split the counts by day, browser, country, app version and new vs returning users to find where the loss happens, fix it, and rerun the experiment.

1 · Assignment fair 50/50 coin 2 · Exposure page loads, redirect 3 · Logging events, bot filters 4 · Analysis joined table B's extra redirect loses impatient users B crashes on old phones; bot filter differs by arm join drops users with missing fields Every red leak removes a non-random group of users from one arm only. the coin is rarely the problem
An SRM is usually not a broken coin. It is a leak somewhere between assignment and the analysis table, and the users who leak out are special (impatient, on old phones, bots), so the remaining arms differ.

Type the user counts of A and B, and set the planned share of A. The purple line is the plan; the green zone is how far chance alone usually pushes the observed share (the region where $p \gt 0.001$). Press Classic SRM: the dot lands far outside. Press Tiny gap, huge n: 50.5% vs 49.5% looks harmless, but with a million users chance wobble is tiny, so it is a glaring SRM. Then try Planned 20/80.

Why is an SRM so dangerous? The next widget makes the leak concrete. There is no real effect: A and B are the same page. But a bug in B makes some users on slow devices bounce before they are logged, and those users rarely convert.

20% of users are on slow devices and convert at 4%; the rest convert at 12%. B is identical to A, but B loses some slow-device users before logging (red dashed part). Move lost: B's logged conversion rate rises and a fake lift appears. Then move users per arm down to 1,000: the fake lift stays exactly the same, but the SRM p-value is no longer alarming. The bias does not shrink with less data; only your ability to detect it does.

"The split is 50.5% vs 49.5%. That's close enough."

"Close enough" depends on $N$. With 1,000,000 users, 505,000 vs 495,000 gives $\chi^2 = 100$ and $p \approx 10^{-23}$: certainly a bug. With 1,000 users the same percentages are normal wobble.

"There is an SRM, so I'll reweight the arms back to 50/50 and carry on."

Reweighting fixes the counts, not the people. The users who leaked out are a special subset (slow devices, impatient users, bots), so the remaining arms still differ in who they contain.

"The SRM test is not significant, so the data are clean."

With a small experiment a real leak can hide (try 1,000 users in the widget). A passed check is reassuring, not a proof. Also check the split by day and by segment: a leak in one browser can be invisible in the total.

"The SRM only matters if the metric looks strange."

A leak can push the metric in either direction, including turning a do-nothing change into a convincing winner. Counts first, metrics second, always.

In an A/B framework like yours, the Beta-Binomial model takes the conversions $k$ and the users $n$ of each arm as given and returns $P(\theta_B \gt \theta_A \mid D)$. If 5% of B's users leaked out for a reason linked to the treatment, the posterior is a perfectly computed answer about the wrong groups. A Bayesian model cannot detect or repair a broken assignment, so the SRM count check belongs before the model, for every metric type (Beta-Binomial, Dirichlet-Multinomial, Normal, Student-t, Poisson). It is cheap: one chi-square test on the counts.

"B won with p = 0.01, so we ship."

"Before reading the metric I check that the experiment is valid: SRM on the unit counts, an A/A history for this metric, and that the analysis unit matches the randomization unit. Only then does p = 0.01 mean what it says."

Model answer: "An SRM means the arms no longer hold comparable users. The bias it creates has unknown size and direction and does not shrink with more data, so I treat a failed SRM check as 'results invalid', find the leak, fix it and rerun."

SRM check: $E_g = N q_g$, $\chi^2 = \sum (O_g - E_g)^2/E_g$, df = groups − 1. Tiny p (teams often use $p \lt 0.001$) → results invalid.

Meaning: users leaked from one arm, and leaked users are not random.

Traps: "close enough" depends on $N$; reweighting does not fix it; the bias does not shrink with more data.

Quick check: 20,000 users, planned 50/50, observed 10,100 vs 9,900. SRM?

Expected 10,000 each. $\chi^2 = 100^2/10{,}000 + 100^2/10{,}000 = 1 + 1 = 2$, df = 1, $p \approx 0.16$. No evidence of SRM: this is ordinary coin-flip wobble.

Attrition: comparing only the users who stayed

A gym tries a new, very tough training plan on half its members. Three months later it compares the fitness of the members who are still coming. The tough plan's group looks much fitter. But the tough plan also drove away everyone who struggled. The gym is comparing "people strong enough to survive the tough plan" with "everyone in the normal plan". That is not a fair fight.

Attrition means units leave the experiment (or stop being measured) after they were assigned. It is harmless if the leavers are a random handful in both arms. It is dangerous when the treatment itself changes who leaves, because then the users who remain are different kinds of people in the two arms.

Three ways to say it:

  • Picture: judging a race by the runners who finished, and forgetting the ones the course made drop out.
  • Numbers: among users who stayed, B converts at 10.9% vs 10.0% (+0.9 pts); counting every user who was assigned, B converts at 9.25% vs 10.0% (−0.75 pts).
  • Slogan: analyse everyone you assigned, not only those who stayed.

A slower checkout. 1,000 users per arm. Half are high-intent (they convert 15% of the time), half are low-intent (5%). B's new page is slower. For users who stay, B has no effect at all on conversion, but 30% of low-intent users in B give up before their session is counted as "active". The metric on the dashboard is "conversion per active user". (Numbers are averages, so fractions of a conversion are fine.)

  1. A, everyone stays: $500 \times 0.15 + 500 \times 0.05 = 75 + 25 = 100$ conversions from 1,000 active users: $10.0\%$.
  2. B, who stays: all 500 high-intent users and $500 \times 0.7 = 350$ low-intent users, so 850 active users.
  3. B's conversions: $500 \times 0.15 + 350 \times 0.05 = 75 + 17.5 = 92.5$.
  4. Per active user: $92.5 / 850 = 10.88\%$. Compared with A: $+0.88$ pts. B looks better!
  5. Per assigned user (the 150 who left did not buy): $92.5 / 1{,}000 = 9.25\%$. Compared with A: $-0.75$ pts. B actually lost sales.
  6. Why the two answers disagree: the slow page removed mostly low converters from B's denominator, so B's survivors are a more "elite" crowd than A's users.
  • Attrition (also called dropout or loss to follow-up): units that were assigned to an arm but whose outcome is not observed, or who drop out of the analysis population, after assignment.
  • Differential attrition: the amount or the kind of units that leave differs between arms. Then the remaining units are no longer comparable, and the comparison is biased.
  • Intention-to-treat (ITT) analysis: analyse every unit in the arm it was assigned to, whatever happened afterwards, using an outcome that is defined for everyone (for example "converted within 7 days of assignment", which is 0 for a user who left). Randomization guarantees that the assigned groups are comparable, so the ITT comparison is unbiased for the effect of assigning the treatment.

Attrition is closely related to SRM: if B loses more units from the analysis population than A, an SRM check on the analysed units will often (not always) flag it.

Why do we need it?

Many product metrics are only defined for users who are still around: "revenue per buyer", "sessions per active user", "rating among users who answered the survey". If the treatment changes who is still around, these metrics compare different people and can show a gain that is really a loss.

Where is it used?

Engagement and retention metrics, ratio metrics with a "per active user" or "per buyer" denominator, surveys and ratings answered by only some users, email tests analysed only on openers, and clinical trials, where "loss to follow-up" is reported for every arm.

How is it used?

Fix the analysis population at assignment time. Prefer metrics per assigned user (ITT). Report how many units dropped out of each arm and test that like an SRM. If you must report a per-survivor metric, also report how its denominator changed, so readers can see the shift in who is counted.

Left: the users of each arm (darker = high-intent, lighter = low-intent, red dashed = users who left before being counted). Right: the measured lift with two definitions of the metric. Start with no effect for stayers and 30% of low-intent users leaving B: the survivor metric says "B wins", the ITT metric says "B loses". Set low-intent leave to 0: both agree. Then make high-intent users leave instead: now the survivor metric is pushed the other way.

"I'll drop the users who bounced; they never really used the feature."

Whether a user bounces may itself be caused by the treatment. Dropping them after assignment breaks the randomization. Keep them, in the arm they were assigned to.

"Same number of users left each arm, so attrition is fine."

Equal numbers are not enough: if B loses low-intent users and A loses high-intent users, the counts can match while the kinds of people differ.

"ITT underestimates the effect, so it is the wrong analysis."

ITT answers "what happens if we launch B to everyone?", which is usually the business question. Effects for narrower groups (users who actually saw B) need extra assumptions (see contamination, next, and instrumental variables in Chapter 5.12).

Attrition = assigned units that are no longer measured. Differential attrition breaks comparability.

Fix: intention-to-treat: analyse everyone as assigned, with an outcome defined for everyone (per assigned user).

Trap: "per active user" or "per buyer" metrics can flip sign when the treatment changes who stays.

Quick check: a new onboarding flow makes 10% more users finish onboarding. Among users who finished, 7-day retention is lower in B than in A. Is onboarding B bad for retention?

Not necessarily. B lets more (and probably weaker) users through, so "users who finished onboarding" is a different crowd in B. Compare 7-day retention per assigned user: that comparison is protected by randomization.

Contamination and non-exposure: when the arms leak into each other

A farmer tests a new fertilizer on half of the field. Rain washes some of it into the control rows, so the control plants also grow a bit more. And a few treated rows were never fertilized because the sprayer missed them. At harvest the two halves look more alike than the fertilizer really makes them.

In online experiments the "rain" is: users with two devices or shared family accounts who see both versions, users who clear cookies and get re-assigned, staff who see the new version and change how they help customers. The "missed rows" are treated users who never reach the part of the product that changed. Both push the two arms toward each other.

Three ways to say it:

  • Picture: a drop of the new paint falls into the old can, and part of the new can was never opened; the two colours look closer than they are.
  • Numbers: true lift +2 pts; 20% of control users also get B and 10% of B users never see it; you measure $2 \times (1 - 0.2 - 0.1) = 1.4$ pts.
  • Slogan: leaks usually shrink the difference toward zero, and cost you power.

A new checkout raises conversion from 10% to 12% for anyone who actually sees it (true effect +2 pts). 20% of the control arm also gets the new checkout (shared accounts), and 10% of the treatment arm never reaches checkout, so they see nothing new.

  1. Control arm: $0.8 \times 10\% + 0.2 \times 12\% = 8.0 + 2.4 = 10.4\%$.
  2. Treatment arm: $0.9 \times 12\% + 0.1 \times 10\% = 10.8 + 1.0 = 11.8\%$.
  3. Measured difference: $11.8 - 10.4 = 1.4$ pts, which is $2 \times (0.9 - 0.2) = 2 \times 0.7$.
  4. Power cost: the effect you can see is 0.7 times as big, and the users needed grow like $1/\text{effect}^2$, so you need $1/0.7^2 \approx 2.04$ times as many users to detect it (Chapter 5.7).
  5. If you trust the exposure records, $1.4 / 0.7 = 2.0$ pts recovers the effect for the users whose exposure was changed by the assignment. This is the instrumental-variable idea of Chapter 5.12.
  • Contamination (crossover, spillover between arms): units assigned to one arm receive the other arm's experience, fully or partly.
  • Non-exposure (non-compliance): units assigned to the treatment never actually receive it (they never reach the changed page, or the feature failed to load for them).
  • If $e_T$ is the share of the treatment arm that really gets B, $e_C$ the share of the control arm that gets B, and the effect $\tau$ is the same for everyone, the ITT difference is $$\text{measured} = \tau \,(e_T - e_C).$$ With $e_T = 0.9$, $e_C = 0.2$ the dilution factor is $0.7$.
  • The dilution formula needs an assumption (the effect does not depend on who gets exposed). In real data, the users who cross over are often special (heavy users with many devices), so the bias can also go in other directions.
Why do we need it?

A diluted experiment reports a smaller effect than the change really has, and loses power fast: 30% leakage roughly halves the information. Teams then kill good ideas as "flat" or run tests twice as long as needed.

Where is it used?

Logged-out experiments randomized by cookie (cookie resets, many devices), shared accounts (households, teams in B2B tools), features that only some treated users reach, sales or support staff who see both versions, and clinical trials (patients who do not take the pill, or get it elsewhere).

How is it used?

Randomize at a unit that leaks less (logged-in user or account instead of cookie). Measure exposure in both arms. Analyse by ITT as the main answer; add a triggered analysis (users who reached the changed page, defined the same way in both arms, Chapter 5.10) for sensitivity; plan the sample size for the diluted effect.

Each grid is 100 users of one arm; blue dots got the old experience, orange dots got the new one. Raise control contamination: orange dots appear in A. Raise treatment non-exposure: blue dots appear in B. Watch the measured effect shrink to true × (share of orange in B − share of orange in A), and the "users needed" factor grow like 1/dilution².

"Contamination can only make the effect look smaller, so a significant result is still safe."

Under the simple dilution model, yes. But crossover users are often special (heavy users on many devices), so contamination can also shift the effect in other directions. Measure it rather than assume it.

"I'll analyse only the users who actually saw B, against all of A."

That compares a self-selected group with a random one: users who reach the new page are more engaged. A fair triggered analysis uses the same trigger rule in both arms (users of A who would have seen B at that point).

Contamination: control units get B. Non-exposure: treatment units never get B.

Simple model: measured $= \tau(e_T - e_C)$; users needed $\times 1/(e_T - e_C)^2$.

Main answer: ITT. Prevent leaks with a better randomization unit; check exposure in both arms.

Quick check: 40% of treated users never see the change and nobody in control gets it. The true effect is +3 pts. What do you expect to measure, and how many more users do you need?

Measured $= 3 \times (0.6 - 0) = 1.8$ pts. Users needed grow by $1/0.6^2 \approx 2.8$ times. A triggered analysis on users who reach the changed page (same rule in both arms) can win back much of this power.

Novelty, primacy and carryover: effects that change with time

A shop moves its "special offers" shelf to the entrance. In the first days everyone stops to look, simply because it is new. A month later regulars walk past it like any other shelf. A one-week test would have measured the curiosity, not the lasting effect. That is a novelty effect.

The opposite also happens. Move the milk to a new aisle and regulars are annoyed for a week; they hunt for it and buy less. After they learn the new layout, sales recover and may even rise. That is a primacy effect (also called change aversion).

And sometimes an old experiment leaves fingerprints on a new one: users who lived with last month's treatment still behave differently. That is carryover.

Three ways to say it:

  • Picture: the effect is a moving target; a short test photographs it at one moment, usually the least typical one.
  • Numbers: daily lifts of +5, +3, +2, +1, +1, +1, +1 pts average to +2 pts over a week, but the lasting effect is +1 pt.
  • Slogan: plot the effect over time before you trust its average.

A redesigned home page is tested for 7 days. The daily lift in click-through (B − A), in pts, is: day 1: +5, day 2: +3, day 3: +2, days 4–7: +1 each.

  1. Sum of the daily lifts: $5 + 3 + 2 + 1 + 1 + 1 + 1 = 14$.
  2. Average over the week (equal traffic each day): $14 / 7 = 2.0$ pts. This is what the 7-day test reports.
  3. The lifts settle at $+1$ pt from day 4 on, so the lasting effect is about $+1$ pt: half of what the test reported.
  4. A 28-day test with the same pattern (days 8–28 also at +1) averages $(14 + 21)/28 = 35/28 = 1.25$ pts: closer, but still pulled up by the first days.
  5. The honest summary: "+1 pt after the novelty wears off", read from the plot of lift by day, not from the overall average.
  • Novelty effect: a temporary change in behaviour (usually a boost) because the experience is new. It fades as users get used to it.
  • Primacy effect (change aversion): a temporary drop because users are used to the old experience. It fades as users learn the new one.
  • Both make the short-run effect differ from the long-run effect. Both mostly affect returning users; brand-new users have no old habit, so comparing new and returning users is a useful diagnostic.
  • Carryover: the effect of an earlier treatment persists into a later period or a later experiment. Examples: re-using the same user buckets for the next experiment, so the new "control" contains users changed by the old treatment; in switchback tests (whole market switched between A and B over time), the first hours after a switch still reflect the previous period.
  • Related time effects: day-of-week patterns (run whole weeks so each weekday is equally represented) and seasonality.
Why do we need it?

Launch decisions are about the long run, but experiments are short. If the first days are unusual, the test answers the wrong question: a novelty boost ships a feature that does nothing a month later; change aversion kills a feature that would have won.

Where is it used?

UI redesigns, new recommendation or ranking models (users first explore the new items), notification and email changes (fatigue grows over time), pricing changes, and any experiment that reuses user buckets or switches whole markets over time.

How is it used?

Plot the lift by day and by "days since first exposure"; compare new vs returning users; run full weeks and long enough for the curve to flatten; keep a long-term holdout (a small group that never gets the change) for big launches; re-randomize with a fresh seed between experiments, or leave a washout gap.

time Experiment 1 (weeks 1–2) buckets 0–49: new ranking (B) users learn new habits buckets 50–99: old ranking (A) Experiment 2 (weeks 3–4), same buckets buckets 0–49: control of exp. 2 still changed by exp. 1! buckets 50–99: treatment of exp. 2 carryover Fix: re-randomize with a fresh seed for every experiment, or leave a washout gap.
Carryover through re-used buckets. The users who had the new ranking in experiment 1 still behave differently in experiment 2, so experiment 2's "control" is not a clean control.

The orange curve is the daily lift; the green dashed line is the lasting (long-run) lift; the purple curve is what a test that stops on day $t$ would report (the average of the daily lifts so far). Press Novelty and move test length: a 1-week test reports far more than the lasting lift. Press Primacy: a short test even gets the sign wrong. Press No time effect: then any length gives the right answer.

"The test ran for 7 days and was significant, so the effect is real and lasting."

Significant means "probably not zero during the test". Whether it lasts is a separate question, answered by the shape of the lift over time, by new vs returning users, and sometimes by a long-term holdout.

"Run the test for 5 days; we already have enough users."

Even with enough users, run whole weeks: weekend users often behave differently, and a 5-day test over-weights some weekdays.

"Each experiment can reuse last month's buckets; the assignment is still random."

It is random with respect to the new test, but the buckets carry the old treatment's effects. Re-hash users with a new seed (salt) per experiment.

Novelty: early boost that fades. Primacy: early dip that fades. Short test ≠ long-run effect.

Carryover: an earlier treatment still acts later (re-used buckets, switchback periods).

Check: lift by day and by days-since-exposure, new vs returning users; run full weeks; fresh randomization per experiment.

Quick check: a new feed layout shows −1.5 pts for returning users and +0.8 pts for brand-new users in week 1. What might be going on?

A primacy effect (change aversion): returning users are thrown by the change, while new users, who have no habit to break, already benefit. Expect the returning users' lift to move toward the new users' lift as they adapt; keep the test running and plot the lift by days since first exposure before deciding.

SUTVA and interference: when users affect each other core

Every A/B test silently assumes that my outcome depends only on my group. Usually that is fine: whether you see a blue or a green button does not change what I buy.

But think of a ride-sharing app testing a discount. Treated riders book more rides, so fewer drivers are free, so control riders wait longer and book fewer rides. The control group is hurt by the treatment it never received. The A/B difference now compares "treated riders who grab drivers" with "control riders who lost drivers to them", which is much bigger than what would happen if everyone got the discount (then there is no one to grab drivers from).

This is interference: one unit's treatment changes another unit's outcome. The assumption that it does not happen is part of SUTVA.

Three ways to say it:

  • Picture: two halves of a classroom share one box of pencils; giving half the class "grab faster" lessons makes the other half look worse.
  • Numbers: in a market with limited stock, the naive A/B lift is +3.7 pts, but launching to everyone adds only +1 pt.
  • Slogan: if units share anything (stock, drivers, friends, budget), "my group only" may be false.

A marketplace with limited stock. 1,000 buyers a day; a seller has 110 items a day. With the old page 10% of buyers want to buy; a new "Buy now" button raises that to 14%. The test splits buyers 50/50.

  1. Demand in the test: control $500 \times 0.10 = 50$ buyers, treatment $500 \times 0.14 = 70$ buyers; total $120$.
  2. Only 110 items: each buyer who wants one gets it with probability $110/120 \approx 0.917$.
  3. Measured conversion: control $10\% \times 0.917 = 9.17\%$; treatment $14\% \times 0.917 = 12.83\%$. Naive lift: $+3.67$ pts.
  4. World where everyone gets the old page: demand $100 \le 110$, conversion $10\%$.
  5. World where everyone gets the button: demand $140 \gt 110$, only 110 sales, conversion $11\%$.
  6. The effect of launching is $11\% - 10\% = +1$ pt. The A/B test reported $3.67$ times that, because treated buyers took items that control buyers would otherwise have bought.

Write $Y_i(1)$ for the outcome unit $i$ would have under treatment and $Y_i(0)$ under control. These are potential outcomes; Chapter 5.12 builds them up properly. The Stable Unit Treatment Value Assumption (SUTVA) says that these two numbers are all there is. It has two parts:

  1. No interference: unit $i$'s outcome depends only on unit $i$'s own assignment, not on anyone else's. In symbols, if $\mathbf z = (z_1, \dots, z_N)$ is everyone's assignment, $Y_i(\mathbf z) = Y_i(z_i)$.
  2. No hidden versions of treatment: "treatment" means one well-defined thing. Every treated unit gets the same version (not a fast version for some and a slow, buggy version for others).
  • Interference (spillover) is any violation of part 1. Then the A/B difference no longer estimates the global treatment effect (everyone treated vs no one treated), which is what a launch does.
  • The bias can go either way. Competition for a shared resource (stock, drivers, ad budget, search ranking slots) usually makes the naive test overstate the effect. Positive spillover (a sharing feature in a social network, where control users receive more shares from treated friends) usually makes it understate the effect.
  • Designs that respect interference: cluster randomization (randomize whole groups that mostly interact inside themselves: cities, network communities, schools), switchback experiments (switch a whole market between A and B over time slots), and budget-split designs for ads. The price: far fewer independent units, so wider intervals (Chapter 5.10).
Why do we need it?

Every standard A/B analysis (t-test, z-test, Beta-Binomial posterior) is only an estimate of the launch effect if SUTVA holds. When units share a resource or influence each other, a perfectly randomized test can still give a badly wrong answer, with tight error bars.

Where is it used?

Two-sided marketplaces (ride sharing, food delivery, hotel and holiday-rental booking), ad auctions with shared budgets, search and ranking (items compete for slots), social networks and messaging (features spread between friends), pricing tests (customers compare prices), and B2B tools where a team shares one account.

How is it used?

Before the test, ask: "Can a treated unit change a control unit's outcome?" If yes, randomize at a level where the interaction stays inside the unit (city, cluster, time slot), analyse at that level, and accept fewer units. Sometimes run tests with different treated shares (10%, 50%) to measure how big the spillover is.

By user By cluster (city, community) By time (switchback) red links cross the arms: spillover, SUTVA fails links stay inside a cluster; few clusters → wide intervals ABBA whole market, one hour each everyone shares one arm at a time; watch for carryover after switches blue = control (A) · orange = treatment (B)
When units interact, user-level randomization lets the treatment leak into the control group. Randomizing whole clusters or whole time slots keeps each interaction inside one arm, at the cost of fewer independent units.

Buyers share a limited stock. Left pair of bars: the A/B test (both arms in one shared market). Right pair: two whole worlds, everyone on A or everyone on B. Lower the stock until it runs out: the A/B lift stays large while the launch effect collapses. Raise the stock to 200: no competition, SUTVA holds, both answers agree. Then change the share treated and see that the A/B answer itself depends on how many users you treat, a sure sign of interference.

"The test was perfectly randomized, so the estimate is unbiased."

Randomization makes the arms comparable. It does not stop treated units from changing control units' outcomes. Under interference, the estimate can be precise and still answer the wrong question.

"Interference always inflates the effect."

Competition for a shared resource inflates it; positive spillover (sharing, network effects) shrinks it, because control users also benefit.

"Cluster randomization is free."

It reduces the number of independent units from millions of users to perhaps dozens of cities, so the intervals get much wider. Analyse at the cluster level (Chapter 5.10).

In an A/B framework like yours, the Beta-Binomial (or Normal, Student-t, Poisson) likelihood treats each user's outcome as depending only on that user's arm: that is SUTVA, written as a model. If users interact (shared stock, social features), $P(\theta_B \gt \theta_A \mid D)$ is computed correctly but describes "B in a half-treated world", not "B launched to everyone". The fix is in the design (randomize clusters or time slots), and then the model's unit becomes the cluster or time slot. Hierarchical partial pooling across segments does not repair interference.

"SUTVA means the treatment and control groups are similar."

"SUTVA means the effect is the same for every user."

SUTVA says each unit has exactly one potential outcome per treatment value: (1) no interference (my outcome depends only on my own assignment) and (2) no hidden versions of the treatment. Group similarity is what randomization gives; equal effects is a different (and usually false) assumption.

Model answer: "SUTVA is what lets me write $Y_i(1)$ and $Y_i(0)$ at all. It fails in marketplaces and networks, where one user's treatment changes another user's outcome; then the A/B difference is not the launch effect, and I randomize clusters or time slots instead of users."

SUTVA = (1) no interference: $Y_i(\mathbf z) = Y_i(z_i)$; (2) one well-defined version of treatment.

Interference: competition → naive test overstates; positive spillover → understates.

Fix: cluster or switchback randomization; analyse at that level; expect wider intervals.

Quick check: a messaging app tests "share to group" for 50% of users. Shares received by control users also rise. Which way is the naive estimate of "shares sent per user" biased?

Control users receive more shares from treated friends and some reply or reshare, so control's numbers rise too. The difference shrinks: the naive test likely understates the launch effect (positive spillover). Randomizing by friend clusters would reduce this.

Repeated peeking: every look is another lottery ticket core

Run an A/A test (both arms identical) and plot its z-statistic every day. It does not sit still: it wanders up and down like a drunk walker on a street with a cliff edge at $\pm 1.96$. If you look only once, at the planned end, the walker is over the edge about 5% of the time. That is the 5% the test promised.

But if you look every day and stop the first time he is over the edge, you give him 30 chances instead of one. On one of those days he often happens to be over the edge, and you declare a winner that does not exist.

This is peeking: checking a fixed-sample test again and again while the data come in, and stopping (or deciding) when it looks significant.

Three ways to say it:

  • Picture: a wandering line and a fixed fence; the more often you check, the more likely you catch the line leaning over the fence.
  • Numbers: A/A test, 30 daily looks, stop at the first $p \lt 0.05$: about 28% of tests declare a winner, not 5%.
  • Slogan: every extra look is another lottery ticket for a false positive.

Just two looks. An A/A test is checked halfway and at the end, each time with the usual rule "$|z| \gt 1.96$".

  1. At the halfway look, a false positive happens with probability 5%.
  2. If the two looks were independent, the chance that at least one is a false positive would be $1 - 0.95^2 = 9.75\%$ (the "at least one" rule).
  3. They are not independent: the final z uses all the halfway data plus the same amount of new data, so the two z-values have correlation $\sqrt{1/2} \approx 0.71$.
  4. Working with that correlation (a two-dimensional Normal calculation, or a simulation) gives $8.3\%$: less than 9.75%, but already well above 5%.
  5. With 30 looks: independence would say $1 - 0.95^{30} = 78.5\%$; the true answer, because neighbouring looks share most of their data, is about $28\%$. Still more than five times the promise.

A fixed-sample test promises: "if $H_0$ is true, $P(\text{reject}) = \alpha$" for one analysis at a sample size fixed in advance. Peeking means running that same test at several interim looks and acting on the first significant one. With $K$ equally spaced looks, each at level $\alpha$, the overall false-positive rate is

$$P(\text{reject at some look} \mid H_0) \;=\; P\big(\max_{k \le K} |z_k| \gt 1.96\big) \;\gt\; \alpha .$$
Looks $K$123510203050100
False-positive rate5.0%8.3%10.7%14.2%19.3%24.7%28.0%32.0%37.6%
  • (Two-sided $\alpha = 0.05$, equally spaced looks; values from 400,000 simulated A/A tests, matching the classic table of Armitage, McPherson and Rowe.)
  • With unlimited looks the rate creeps toward 100%: a random walk crosses any fixed fence eventually.
  • Looking is only harmful when it can change what you do: stopping early, shipping, or extending a test because it is "almost significant" are all forms of peeking.
  • Stopping at a lucky high also makes the reported effect too big on average (you stopped because the noise was in your favour): the winner's curse.
Why do we need it?

Dashboards refresh every hour and everyone wants to stop early. Without understanding peeking, a team that "stops when it's significant" ships many changes that do nothing, while believing each one passed a 5% test.

Where is it used?

Every live A/B dashboard showing a p-value or "chance to beat control"; clinical trials, where formal interim analyses were invented for exactly this reason; and online experimentation tools such as Optimizely, which switched to sequential "always-valid" p-values so that customers could watch results continuously.

How is it used?

Fix the sample size or duration in advance with a power calculation (Chapters 5.7 and 5.10) and decide only at the end. If you need interim looks (for safety or to stop early), plan them and use a sequential method with adjusted thresholds (next section).

0% 10% 20% 30% 40% 1 2 5 10 20 50 100 number of looks K (log scale), each with |z| > 1.96 promised: 5% 19.3% at 10 looks 28% at 30 daily looks 37.6% at 100
The real false-positive rate of "check the fixed-sample test at every look and stop at the first p < 0.05", for A/A tests. It keeps rising with the number of looks; the 5% promise holds only for a single planned look.

Each grey line is the z-statistic of one A/A test, day by day (40 of the 500 are drawn). The red dashed lines are $\pm 1.96$. Purple ticks mark the days you look. With look every day, a test is stopped (red, with a dot) the first time it is outside the fence on a look day. Switch to 30 days (1 look): only the final position counts and the false-positive rate drops to about 5%. Press New sample a few times to see the wobble.

"Each daily check is a valid 5% test, so checking daily is fine."

Each check alone has a 5% false-positive rate, but you act on the first success among many checks. The chance that at least one check is a false positive is what matters, and it grows with the number of looks.

"It crossed p < 0.05 on Tuesday, so the effect is real."

Under $H_0$ the z-path wanders across the fence and back regularly. A brief crossing at an unplanned look is weak evidence.

"Peeking is only a problem for small experiments."

The inflation depends on the number of looks, not on the sample size. A test with 10 million users checked 30 times is just as inflated.

"It's almost significant; I'll run it one more week."

Extending because of what you saw is also peeking: it gives the noise another chance. Decide the duration before the test starts.

Peeking: a fixed-sample test checked at $K$ looks, stopping at the first $p \lt \alpha$.

Real false-positive rate (α = 5%): 2 looks 8.3%, 5 looks 14%, 10 looks 19%, 30 looks 28%; → 100% with unlimited looks.

Also: stopping at a lucky high exaggerates the effect (winner's curse). Extending "almost significant" tests is peeking too.

Quick check: a team checks an A/A test at 10 equally spaced times and stops at the first p < 0.05. Roughly how often does it declare a winner?

About 19% of the time, nearly four times the promised 5%. (A single look at the planned end would give 5%.)

Sequential testing: plan the looks, pay for the looks

Think of the 5% false-positive rate as a budget. A single look at the end spends the whole 5% at once. If you want to look five times, you must split the budget across the five looks, so each look is harder to pass. Then the total chance of a false alarm, over all looks together, stays at 5%.

There are different ways to split the budget. Spend it evenly, and every look has the same, stricter fence. Or save most of it for the end: early looks need overwhelming evidence (only stop if the result is spectacular), and the final look is almost as easy as a normal test.

Three ways to say it:

  • Picture: a 5% cake shared among the looks; the more looks, the thinner each slice.
  • Numbers: 5 looks: Bonferroni needs $|z| \gt 2.58$ at each look; Pocock $|z| \gt 2.41$; O'Brien–Fleming starts at $|z| \gt 4.56$ and ends at $2.04$.
  • Slogan: plan the looks, pay for the looks.

Bonferroni over 5 looks. You want to check an experiment after each fifth of its planned users and keep the overall false-positive rate at most 5%.

  1. Split the budget evenly: $\alpha / K = 0.05 / 5 = 0.01$ per look.
  2. Two-sided threshold for $0.01$: $z = \Phi^{-1}(1 - 0.005) = 2.576$. Reject at a look only if $|z_k| \gt 2.576$.
  3. Why it is safe: the chance of at least one false alarm is at most the sum of the chances, $5 \times 0.01 = 0.05$ (the union bound: $P(A_1 \cup \dots \cup A_5) \le \sum P(A_k)$).
  4. Simulated A/A rate: about $3.3\%$, safely below 5%, but wastefully so: the looks are strongly correlated, so the union bound overcounts.
  5. Cost: for an effect that a single final test detects 80% of the time, Bonferroni-over-looks detects it only about 65% of the time (simulation). Pocock gets about 71% and O'Brien–Fleming about 79%, both with exactly 5% false positives.

A group sequential design fixes in advance $K$ looks at fractions $t_k = n_k / N$ of the maximum sample and boundaries $c_1, \dots, c_K$. At look $k$, stop and reject $H_0$ if $|z_k| \ge c_k$; otherwise continue. The boundaries are chosen so that $P(\text{reject at any look} \mid H_0) = \alpha$.

  • Bonferroni over looks: $c_k = \Phi^{-1}(1 - \alpha/(2K))$. Always valid, but conservative.
  • Pocock: the same $c$ at every look, chosen exactly ($c = 2.413$ for $K = 5$, $\alpha = 0.05$). Easier to stop early, harder at the end.
  • O'Brien–Fleming: $c_k = C\sqrt{K/k}$ ($C = 2.040$ for $K = 5$): very strict early, close to 1.96 at the final look, so it costs almost no power at the end.
  • α-spending (Lan–DeMets): instead of fixed looks, choose a spending function $\alpha(t)$ saying how much of the 5% may be used by information fraction $t$. The looks can then happen at times not fixed in advance, as long as the spending function is.
  • Always-valid inference (for example the mixture sequential probability ratio test, mSPRT, and confidence sequences): p-values and intervals that stay valid at any stopping time, even with continuous monitoring. The price: at a fixed planned end they are less powerful (wider) than a fixed-sample test.
Why do we need it?

Teams have good reasons to look early: stop a harmful change fast, or stop a clearly great one and free the traffic. Sequential methods allow this without silently breaking the 5% promise.

Where is it used?

Clinical trials (interim analyses with O'Brien–Fleming-type boundaries are standard practice), online experimentation platforms that offer sequential or always-valid p-values (Optimizely's Stats Engine is a published example), and guardrail monitoring that must be able to stop a harmful launch at any time.

How is it used?

Before the test, choose the number of looks (or a spending function) and the boundary type. Compute the boundaries with a group-sequential library (for example R's gsDesign or rpact) or by simulation. At each look compare $z_k$ with $c_k$; report the effect with methods that account for the early stop.

Where the false-positive budget goes (A/A tests, 5 looks) 5% budget 1.96 every look look 1 2 3 4 5 14.1% in total Bonferroni (2.58) 3.3% (budget left unused) Pocock (2.41) 1 2 3 O'Brien–Fleming 4 5 Bar length = share of A/A tests that falsely stop at each look (segments 1→5).
The naive rule overspends: 14.1% in total. Bonferroni stays safe but leaves budget unused (3.3%). Pocock spends evenly; O'Brien–Fleming saves almost everything for the last looks. Both of those spend exactly 5%.

Grey lines are z-paths of simulated tests, drawn at each look; red marks are the boundaries $\pm c_k$. Choose A/A and compare the four rules: "1.96 every look" overshoots 5% as $K$ grows; the other three stay at or below 5%. Switch to real effect (sized so a single final test has 80% power): O'Brien–Fleming keeps nearly all the power, Bonferroni loses the most, and every rule stops some tests early (green dots).

"We use a sequential test, so we can look whenever we like with the usual 1.96."

A sequential test is valid only with its boundaries (or its spending function / always-valid p-values). Using 1.96 at unplanned looks is ordinary peeking again.

"Bonferroni over looks is the exact correction."

It is a safe upper bound, not exact: with 5 looks it gives about 3.3%, wasting budget and power, because consecutive looks share most of their data. Pocock and O'Brien–Fleming use that correlation.

"Always-valid p-values are a free lunch."

They buy the right to stop at any time; the price is less power (or wider intervals) than a fixed-sample test at the same planned end.

"We stopped early at look 2, so the observed lift is our best estimate."

Early stops happen when the estimate is high, so naive estimates after an early stop are biased upward. Use adjusted estimates, or treat the size of the lift with caution.

Group sequential: planned looks $k = 1..K$, stop if $|z_k| \ge c_k$, with $P(\text{any rejection} \mid H_0) = \alpha$.

5 looks, α = 0.05: Bonferroni 2.576 (actual 3.3%); Pocock 2.413 each; O'Brien–Fleming $2.040\sqrt{5/k}$ = 4.56, 3.23, 2.63, 2.28, 2.04.

α-spending: choose how much α is used by information fraction $t$. Always-valid p-values: stop any time, pay with power.

Quick check: why does O'Brien–Fleming lose almost no power compared with a single final test?

Its early boundaries are so strict (4.56, 3.23, …) that it spends almost none of the 5% early; its final boundary (2.04) is barely above 1.96. So at the end it behaves almost like the fixed-sample test, while still allowing a stop for overwhelming early evidence.

Peeking with a Bayesian decision rule: what changes and what does not core

Your framework does not report p-values. It reports a posterior, for example "$P(B \gt A \mid \text{data}) = 0.93$". Is that safe to watch every day? The honest answer has two halves, and you need both.

Half 1: the number itself is fine. The posterior is "what I should believe now, given my model, my prior and the data so far". The data do not care whether you planned to stop today or were just curious, so today's posterior is the same either way. Looking does not damage it.

Half 2: the habit built on the number is a procedure, and procedures have error rates. "Ship B the first day $P(B \gt A) \gt 0.95$" gives the posterior many chances to wander above 0.95 on a lucky streak, exactly like the z-path in the last two sections. How often that rule ships a useless or worse B depends on how often you look, and on how well your prior matches the real world.

Three ways to say it:

  • Picture: a thermometer that is always accurate, and a rule "open the window the first time it reads above 25°"; check it every minute on a day hovering near 25° and you open the window more often, though the thermometer never lied.
  • Numbers: A/A test, Beta(1, 1) priors, rule "ship when $P(B \gt A) \gt 0.95$": one look at the end ships B 5% of the time; a look every day for 20 days ships it about 21% of the time.
  • Slogan: the posterior is valid at any time; your stopping rule's error rates are not automatic.

A Bayesian A/A test, checked daily. Both arms convert at 10%; 500 users per arm per day; 20 days; Beta(1, 1) priors on each rate; rule: "ship B as soon as $P(\theta_B \gt \theta_A \mid D) \gt 0.95$". (4,000 simulated tests; the Code-it block reproduces this.)

  1. With a nearly flat prior, $P(\theta_B \gt \theta_A \mid D)$ is close to $\Phi(z)$, where $z$ is the usual z-statistic of the difference. In an A/A test $z$ is roughly standard Normal, so at one fixed day $P(B \gt A \mid D)$ is roughly uniform between 0 and 1.
  2. So at a single planned look on day 20, $P(B \gt A \mid D) \gt 0.95$ happens about 5% of the time. Simulation: 4.9–5.1%.
  3. Checked every day and stopped at the first crossing: B is shipped in about 21% of these A/A tests (simulations: 20.5–22.2%).
  4. On the day each of those tests stopped, the posterior really was above 0.95, computed correctly. Nothing was miscalculated.
  5. The 95% is a statement about belief under the prior "each rate is anywhere in [0, 1], independently". It is not a promise that the rule "ship at 0.95, checked daily" is wrong only 5% of the time.

Let $D_t$ be the data up to day $t$ and $\tau$ a stopping time (a day chosen by a rule that only uses the data seen so far).

  • The posterior is unaffected by the stopping rule. The likelihood of the observed data does not depend on the rule that decided when to stop, so $p(\theta \mid D_\tau)$ is the same posterior you would have computed had you planned to stop at $\tau$ all along (the likelihood principle). In this sense the posterior is "valid at any time, given the model".
  • A decision rule has operating characteristics. For a rule such as "ship B at the first look with $P(\theta_B \gt \theta_A \mid D_t) \gt 0.95$", quantities such as $P(\text{ship B} \mid \theta_A = \theta_B)$ (the A/A false-winner rate) depend on the number of looks, just as for p-values.
  • When the prior matches reality (true effects really behave like draws from your prior), the posterior stays calibrated under optional stopping: among the experiments you stop and ship at $P \ge 0.95$, at least about 95% really have B better. This follows from the fact that the posterior probability is the average truth over everything consistent with the data, under the prior.
  • When the prior does not match reality (for example most real changes do nothing, but the prior is a smooth curve that never expects exactly zero), the stated probability is too optimistic, and frequent looking makes the gap larger. Effect sizes at the stopping time are also exaggerated.
Why do we need it?

"Bayesian tests are immune to peeking" and "Bayesian tests are ruined by peeking" are both said in interviews and both are wrong. Knowing exactly which part is safe (the posterior) and which part needs design (the stopping rule and the prior) is what lets you defend a Bayesian experimentation framework.

Where is it used?

Bayesian A/B dashboards that show "probability to beat control" or "expected loss" live; bandit-style allocation; any Beta-Binomial or hierarchical experimentation framework whose users can check results daily; published work on optional stopping for Bayesian tests (for example by Deng and co-authors at Microsoft, and by Rouder in psychology).

How is it used?

Write the decision rule and the look schedule down before the test. Simulate it: A/A tests and realistic effect sizes through the same priors and checks, and measure how often it ships a non-improvement. Use realistic or skeptical priors (ideally learned from past experiments), a minimum run time (whole weeks), and decisions on practical size, $P(\theta_B - \theta_A \gt \delta \mid D)$.

Question 1: what do I believe now? P(θ_B > θ_A | data so far) = 0.93 ✓ valid on any day, given model + prior does not depend on how often you looked before (likelihood principle) Question 2: how often does my rule fool me? "ship the first day P > 0.95" depends on the number of looks A/A: 5% with 1 look, ~21% with 20 and on whether the prior matches the real effects Peeking never changes the answer to question 1. It always changes the answer to question 2.
Two questions that sound alike. The posterior answers the first at any moment. The second is a property of the whole procedure (rule + look schedule + prior vs reality) and must be designed and checked, for Bayesian rules just as for p-values.

Each line is $P(B \gt A \mid D)$ over 20 days for one experiment (40 of 800 drawn): green if B is truly better, red if not. Rule: ship B at the first look where the line is above 0.95 (purple). Start with A/A and change the looks from 1 to 20: the ship rate climbs. Switch to effects follow the prior: now the share of shipped B that are truly better stays close to the posterior's own claim, whatever the number of looks. Switch to most ideas do nothing: the posterior still says about 0.97, but far fewer shipped B are really better, and more looks make it worse.

Two remarks on the widget. First, its A/A ship rate (4.4% → 13.5%) is lower than the 5% → 21% of the example, because the widget's prior Normal(0, 1 pt²) is skeptical: it pulls small observed lifts toward zero. A skeptical prior slows the inflation but does not stop it. Second, in the "most ideas do nothing" world the prior is wrong (it expects lifts of about ±1 pt and never exactly 0). One look already over-promises (64% truly better vs a claimed ~99%); daily looks over-promise more (38% vs ~97%), and with daily looks the posterior mean of the shipped lifts is more than twice the true average lift (0.90 vs 0.39 pts in a large simulation).

"Bayesian A/B tests are immune to peeking."

The posterior is unaffected by peeking. The error rates of a rule like "ship when $P \gt 0.95$" are not: in A/A tests that rule ships B about 5% of the time with one look and about 21% with daily looks (flat priors).

"Bayesian posteriors become invalid if you look early."

The posterior on any day is the correct belief under the model and prior. If the prior matches reality, it even stays calibrated among the experiments you stop. The problems come from the stopping rule meeting an unrealistic prior.

"An expected-loss rule solves peeking completely."

Rules such as "stop when the expected loss of choosing B is below a tiny threshold" answer a different question: they accept "A and B are about equal, either is fine". That can be the right goal, but their behaviour under frequent looks still has to be checked by simulation.

In an A/B framework like yours, if a dashboard shows $P(\theta_B \gt \theta_A \mid D)$ every day, each day's number is a valid posterior under your Beta-Binomial (or Dirichlet-Multinomial, Normal, Student-t, Poisson) model. Whether "ship at 0.95" is a good procedure is a design question you can answer with the framework itself: run simulated A/A tests and tests with realistic effects through the same priors and the same check schedule, and count how often it ships a non-improvement. Hierarchical priors learned from many past experiments make the prior closer to reality, which is exactly the condition under which optional stopping does little harm. Add a minimum duration (whole weeks, for novelty and weekday effects) and decide on $P(\theta_B - \theta_A \gt \delta \mid D)$ for a practically meaningful $\delta$.

"We're Bayesian, so peeking is not a problem for us."

"We're Bayesian, but peeking invalidates our posteriors."

"The posterior is valid whenever I compute it, because the stopping rule does not enter the likelihood. But my decision rule has frequentist operating characteristics, and repeated checking raises its false-winner rate. Those characteristics are only guaranteed to match the posterior's stated probability if my prior matches the distribution of real effects."

Model answer: "So peeking is still a design question for us. We fix the decision rule and the look schedule in advance, use priors informed by past experiments, require a minimum run time, decide on practically meaningful lifts, and simulate the whole procedure on A/A and realistic scenarios to know its error rates."

Posterior $p(\theta \mid D_t)$: valid on any day given model + prior (stopping rule not in the likelihood).

Rule "ship at first $P(B \gt A) \gt 0.95$": A/A ship rate 5% (1 look) → ~21% (20 daily looks), flat priors.

Prior = reality → still calibrated under optional stopping. Prior ≠ reality (many zero effects) → over-promises, worse with more looks. Fix rule + schedule in advance; simulate it.

Quick check: a colleague says "our dashboard shows P(B > A) = 0.97 today, and it was 0.80 yesterday, so today's number is unreliable because we peeked yesterday". Right or wrong?

Wrong about the number: today's posterior does not depend on whether anyone looked yesterday. What peeking affects is the reliability of the habit "stop as soon as it passes 0.95". Whether to ship today should follow a rule fixed in advance whose error rates were checked, not the fact that the number happens to be high right now.

Many metrics, variants and segments: why false positives multiply core

Buy one lottery ticket and you almost surely lose. Buy twenty and your chance of winning something is much higher. Each test at $\alpha = 0.05$ is a ticket in a "false positive lottery": if nothing is going on, it still "wins" 5% of the time.

A real experiment rarely has one test. It has a primary metric, ten secondary metrics, guardrails, three treatment variants, and then someone asks "what about mobile users in Germany?". Every metric, every variant and every segment is another ticket. With enough tickets, a surprising "significant" result is almost guaranteed, even when the change does nothing.

Three ways to say it:

  • Picture: throw enough darts blindfolded and one of them hits the bullseye; the hit says nothing about your aim.
  • Numbers: 20 independent metrics with no real effect, each tested at 5%: $1 - 0.95^{20} = 64\%$ chance that at least one looks significant.
  • Slogan: the more questions you ask the data, the more lucky answers you get.
  1. Metrics. An A/A test is scored on 20 metrics, each tested at $\alpha = 0.05$; assume the metrics are independent. One metric stays quiet with probability $0.95$; all 20 stay quiet with probability $0.95^{20} \approx 0.358$.
  2. So the chance of at least one false "win" is $1 - 0.358 = 0.642$, about 64%. The expected number of false wins is $20 \times 0.05 = 1$.
  3. Variants. An A/B/C/D/E test compares 4 treatments with one control: about $1 - 0.95^4 \approx 18.5\%$ (a little less in reality, because the 4 comparisons share the same control and are correlated).
  4. Segments. Splitting the result by 10 countries and looking for any country with $p \lt 0.05$: $1 - 0.95^{10} \approx 40\%$.
  5. Bound that always holds, whatever the dependence: $P(\text{at least one}) \le m\alpha$ (here $20 \times 0.05 = 1$, which is useless, and that is why we need smarter corrections in the next section).

Run $m$ tests, each at level $\alpha$. Some null hypotheses are true (no effect), some may be false.

  • The family-wise error rate (FWER) is the probability of at least one false positive among the $m$ tests. If all $m$ nulls are true and the tests are independent, $$\text{FWER} = 1 - (1 - \alpha)^m \;\le\; m\alpha .$$
  • The false discovery rate (FDR) is the expected share of false positives among the results you call significant: $\text{FDR} = E\big[V / \max(R, 1)\big]$, where $R$ is the number of discoveries and $V$ the number of false ones.
  • A family is the set of tests you want one guarantee for (for example "all secondary metrics of this experiment"). Which tests form a family is a decision you make before looking.
  • The garden of forking paths: even without formally running many tests, choosing the metric, the segment, the time window or the outlier rule after seeing the data has the same effect.

Standard protection: one pre-registered primary metric decides; guardrail metrics are watched for harm; secondary metrics and segment cuts are labelled exploratory, corrected for multiplicity, and confirmed in a new experiment before anyone acts on them.

Why do we need it?

Experiment scorecards often show dozens of metrics for several variants and many segments. Without accounting for multiplicity, every experiment "finds something", teams chase noise, and segment-level launches are built on lucky cuts.

Where is it used?

Experiment scorecards with many metrics, multi-arm tests (A/B/C/D), segment analyses (country, device, new vs returning), genomics and neuroscience (thousands of tests at once), and feature screening in ML (testing many candidate features or regressors).

How is it used?

Before the test, name the primary metric and the families. Count how many tests each family holds. Apply a correction (Bonferroni or Holm for FWER, Benjamini–Hochberg for FDR) within each family, or use a hierarchical model for segments. Treat any surprising segment or secondary win as a hypothesis for the next experiment.

one A/A test conversionrevenueclicks … mobiledesktop new usersreturning GermanyJapan … "p = 0.03!" week 1 only? without outliers? each choice = another branch With enough branches, some leaf is "significant" even when nothing happened.
The garden of forking paths: metrics × segments × time windows × cleaning rules. Reporting the one leaf that came out significant, without saying how many branches were explored, turns noise into a "finding".

Top: the chance of at least one false positive, $1 - (1 - \alpha)^m$ (red curve), and the simple bound $m\alpha$ (dashed). Move $m$ to 20: about 64%. Bottom: one simulated A/A experiment with $m$ metrics; each dot is a metric's p-value, red if below $\alpha$. Press Run 200 experiments and compare the share of experiments with at least one red dot with the formula.

An A/A test split into segments (rows). Blue bars are 95% intervals for each segment's lift; a red bar misses zero, a "significant segment". Press New sample several times with 10 segments: a red segment appears in about 40% of samples. Turn on partial pooling: orange intervals pull each segment toward the overall lift by an amount learned from how much the segments really differ; in A/A data they collapse together and the false segment winners mostly disappear (over 4,000 simulated A/A samples with 10 segments: at least one red segment in 41% of samples without pooling, 9% with it). Then choose segment 1 truly +2 pts: pooling shrinks it, but a strong real difference usually survives (found in 94% of samples without pooling, 82% with it).

"Our primary metric was flat, but metric 14 improved with p = 0.03, so the change works."

With 20 metrics, one $p \lt 0.05$ is expected from noise alone (expected count $20 \times 0.05 = 1$). A secondary win is a hypothesis to confirm, not a result.

"The test was flat overall, but it won big in Germany on mobile. Let's launch there."

Segment cuts are extra tickets. Unless the segment was planned in advance (and corrected for), rerun the test in that segment before acting.

"The metrics are correlated, so multiple testing does not apply."

Positive correlation makes the FWER smaller than $1 - (1-\alpha)^m$, but it is still above $\alpha$ unless the metrics are essentially the same metric.

In an A/B framework like yours with hierarchical partial pooling across segments, the segment effects share a common prior whose spread is learned from the data. When the segments truly look alike, the learned spread is small and every segment is pulled toward the overall effect, so one noisy segment needs real evidence to stand out. This is the multilevel-model answer to segment fishing (Gelman, Hill and Yajima wrote "Why we (usually) don't have to worry about multiple comparisons" about exactly this). It does not protect you from scanning 20 different metrics, from many separate experiments, or from re-analysing until something crosses 0.95. For categorical metrics, a single Dirichlet-Multinomial model of all categories is one joint analysis, which is better than a separate test per category.

$m$ independent tests at α, all nulls true: FWER $= 1 - (1 - \alpha)^m \le m\alpha$. 20 tests → 64%; 10 → 40%.

FWER = P(≥ 1 false positive); FDR = expected share of false ones among discoveries.

Protect: one primary metric, families fixed in advance, corrections, hierarchical models for segments, confirm surprises in a new test.

Quick check: an A/B/C test (2 variants vs control) is scored on 5 metrics. Roughly how many tests, and what is the chance of at least one false positive if nothing works (treat them as independent)?

$2 \times 5 = 10$ tests. $1 - 0.95^{10} \approx 0.40$: about a 40% chance of a spurious "win" somewhere.

Corrections: Bonferroni, Holm and Benjamini–Hochberg core

If many tests together get too many lucky wins, make each test harder to pass. The question is how much harder, and that depends on what you want to protect.

  • "I want almost never to make even one false claim." That is controlling the family-wise error rate. Bonferroni does it bluntly: split $\alpha$ equally, $\alpha/m$ per test. Holm does it more cleverly: once the strongest result has passed, the next one only has to beat a slightly easier bar, and so on.
  • "I accept a few false claims, as long as they are a small share of everything I claim." That is controlling the false discovery rate. Benjamini–Hochberg (BH) does it, and finds many more real effects when there are many tests.

Three ways to say it:

  • Picture: Bonferroni raises one high bar for everyone; Holm lowers the bar a little after each success; BH lets a whole group through if its weakest member is good enough for its rank.
  • Numbers: ten p-values 0.001, 0.004, 0.006, 0.02, 0.04, …: Bonferroni keeps 2, Holm keeps 3, BH keeps 4, no correction keeps 5.
  • Slogan: FWER: "no false claims"; FDR: "few false claims among my claims".

Ten secondary metrics, $\alpha = 0.05$, sorted p-values: $0.001, 0.004, 0.006, 0.02, 0.04, 0.06, 0.2, 0.3, 0.5, 0.8$ ($m = 10$).

  1. No correction: keep every $p \le 0.05$: the first five.
  2. Bonferroni: bar $\alpha/m = 0.05/10 = 0.005$ for everyone. $0.001 \le 0.005$ ✓, $0.004 \le 0.005$ ✓, $0.006 \gt 0.005$ ✗. Keeps 2.
  3. Holm (step down): compare the $k$-th smallest with $\alpha/(m - k + 1)$ and stop at the first failure. $k=1$: $0.001 \le 0.05/10 = 0.005$ ✓. $k=2$: $0.004 \le 0.05/9 \approx 0.00556$ ✓. $k=3$: $0.006 \le 0.05/8 = 0.00625$ ✓. $k=4$: $0.02 \gt 0.05/7 \approx 0.00714$ ✗, stop. Keeps 3.
  4. Benjamini–Hochberg (step up): bars $k\alpha/m = 0.005k$: $0.005, 0.010, 0.015, 0.020, 0.025, 0.030, \dots$ Find the largest $k$ with $p_{(k)} \le 0.005k$: $p_{(4)} = 0.02 \le 0.020$ ✓; $p_{(5)} = 0.04 \gt 0.025$, $p_{(6)} = 0.06 \gt 0.030$, and all later ones fail too. So $k = 4$: keep the 4 smallest. Keeps 4.
  5. Adjusted p-values (what statsmodels prints) say the same thing: a test is kept when its adjusted p is $\le 0.05$. Bonferroni: $0.01, 0.04, 0.06, \dots$; Holm: $0.01, 0.036, 0.048, 0.14, \dots$; BH: $0.01, 0.02, 0.02, 0.05, 0.08, \dots$

Sort the p-values $p_{(1)} \le p_{(2)} \le \dots \le p_{(m)}$.

  • Bonferroni (controls FWER $\le \alpha$, any dependence): reject $H_i$ if $p_i \le \alpha/m$. Adjusted: $\min(1, m p_i)$.
  • Holm (controls FWER $\le \alpha$, any dependence; never rejects fewer than Bonferroni): for $k = 1, 2, \dots$ reject $H_{(k)}$ while $p_{(k)} \le \alpha/(m - k + 1)$; stop at the first failure.
  • Benjamini–Hochberg (controls FDR $\le \alpha$ when the tests are independent or positively dependent): find the largest $k$ with $p_{(k)} \le k\alpha/m$ and reject $H_{(1)}, \dots, H_{(k)}$. (For arbitrary dependence, the Benjamini–Yekutieli variant divides by an extra factor $\sum_{j=1}^m 1/j$.)
  • FWER control implies FDR control (if you rarely make any false claim, the share of false claims is small too), but not the other way round.
Why do we need it?

The previous section showed that many tests produce false wins. Corrections restore a guarantee you can state: "the chance of any false claim in this family is at most 5%" (FWER) or "on average at most 5% of my claims are false" (FDR).

Where is it used?

Bonferroni/Holm: a handful of guardrail metrics or variant-vs-control comparisons where any false claim is costly. BH: large scorecards with dozens of secondary metrics, segment scans, genomics (thousands of genes), screening many candidate features or regressors.

How is it used?

Collect the family's p-values and call statsmodels.stats.multitest.multipletests(pvals, alpha=0.05, method='holm') (or 'bonferroni', 'fdr_bh'). Report the adjusted p-values and which tests are rejected. Prefer Holm over Bonferroni (same guarantee, more power).

Edit the ten p-values (they start as the example). The plot shows the sorted p-values against their rank, with the three bars: Bonferroni (flat red), Holm (purple, rising slowly at the end) and BH (green line $k\alpha/m$). Change $\alpha$, or set several p-values just below 0.05: BH keeps many more than Holm. Set one p-value to 0.0001 and the rest to 0.3: all three methods agree.

Each simulated experiment has 100 metrics; 10 have a real effect (strength set by the slider, as the expected z-value), 90 have none. Bars show the average number of true discoveries (green) and false ones (red) per experiment for each method. With no correction you find most real effects but also about 4–5 false ones. Bonferroni and Holm almost never make a false claim but miss most real effects. BH sits in between: false claims are a small share of its discoveries (about 5%).

"Benjamini–Hochberg controls the chance of any false positive."

BH controls the expected share of false positives among the discoveries (FDR). With many tests, it will usually make some false claims; in the simulation above, at least one in about a quarter of experiments.

"Bonferroni is always the safe choice."

It is safe but wasteful. Holm gives the same FWER guarantee and never rejects less. With many metrics, FWER control may be so strict that you miss most real effects; FDR control is often the better trade.

"I'll correct only the metrics that turned out significant."

$m$ is the number of tests you ran (the whole family), not the number that looked interesting.

"FDR and FWER are two names for the false positive rate."

FWER = probability of at least one false positive in the family. FDR = expected proportion of false positives among the rejected hypotheses. The per-test false positive rate is α itself.

Model answer: "For a few decision metrics or variant comparisons, where any false claim is costly, I control FWER with Holm. For a scorecard of dozens of exploratory metrics, I control FDR with Benjamini–Hochberg and treat the discoveries as leads to confirm."

Bonferroni: $p_i \le \alpha/m$. Holm: $p_{(k)} \le \alpha/(m-k+1)$, step down, stop at first failure. Both: FWER ≤ α.

BH: largest $k$ with $p_{(k)} \le k\alpha/m$, reject the $k$ smallest. FDR ≤ α (independent / positively dependent tests).

Trap: BH does not control FWER; $m$ = all tests run, not the ones that looked good.

Quick check: m = 4 p-values 0.010, 0.013, 0.020, 0.30 at α = 0.05. How many does Holm reject? BH?

Holm: $0.010 \le 0.05/4 = 0.0125$ ✓; $0.013 \le 0.05/3 \approx 0.0167$ ✓; $0.020 \le 0.05/2 = 0.025$ ✓; $0.30 \gt 0.05$ ✗: 3 rejections. BH: bars $0.0125, 0.025, 0.0375, 0.05$; $p_{(3)} = 0.020 \le 0.0375$ ✓, $p_{(4)} = 0.30 \gt 0.05$: largest $k = 3$, so also 3 rejections.

Recap, cheat sheet and practice

  • SRM: if the observed split differs from the plan by more than chance ($\chi^2$ test on the counts; teams often alarm at $p \lt 0.001$), users leaked from one arm. Results are invalid until the leak is found. The bias does not shrink with more data.
  • Attrition and "per active user" metrics compare different people when the treatment changes who stays. Intention-to-treat (everyone, as assigned) is the safe default.
  • Contamination and non-exposure dilute the effect: measured $= \tau(e_T - e_C)$, and the users needed grow like $1/(e_T - e_C)^2$.
  • Novelty and primacy effects make short tests differ from the long run; carryover lets old treatments leak into new tests. Plot the lift over time; run full weeks; re-randomize.
  • SUTVA = no interference + one version of the treatment. In marketplaces and networks it fails; randomize clusters or time slots (switchbacks).
  • Peeking at a fixed-sample test inflates false positives: 8.3% with 2 looks, 19% with 10, 28% with 30. Sequential designs (Pocock, O'Brien–Fleming, α-spending, always-valid p-values) pay for the looks with stricter boundaries.
  • Bayesian decisions: the posterior is valid at any time given the model; the rule "ship at the first $P(B \gt A) \gt 0.95$" still has error rates that grow with looks (A/A: 5% → about 21% with 20 daily looks, flat priors), and it over-promises when the prior does not match reality. Peeking is a design question.
  • Multiple testing: $m$ independent tests give FWER $= 1 - (1-\alpha)^m$ (64% for 20). Control FWER with Holm (or Bonferroni), FDR with Benjamini–Hochberg; use hierarchical pooling for segments; fix the primary metric in advance.

Cheat sheet

PitfallCheck / formulaFix
Sample ratio mismatch$\chi^2 = \sum (O_g - E_g)^2 / E_g$, df = groups − 1stop, find the leak, rerun
Attritioncompare units lost per arm; per-survivor vs per-assigned metricintention-to-treat
Contaminationmeasured $= \tau(e_T - e_C)$; $n \times 1/(e_T - e_C)^2$better randomization unit; triggered analysis
Novelty / primacy / carryoverlift by day, by days since exposure, new vs returningrun longer, full weeks, fresh seeds, holdouts
Interference (SUTVA)does B change A's outcomes? lift depends on treated share?cluster or switchback randomization
Peeking$K$ looks at 1.96: 5%, 8.3%, 14.2% (5), 19.3% (10), 28% (30)fixed horizon or sequential boundaries
Sequential boundaries ($K = 5$)Bonferroni 2.576; Pocock 2.413; OBF $2.040\sqrt{5/k}$plan looks or a spending function in advance
Bayesian stopping ruleposterior valid any time; rule's error rates depend on looks and prior realismfixed rule + schedule, realistic priors, simulate it
Many testsFWER $= 1 - (1-\alpha)^m \le m\alpha$primary metric, families, corrections
Bonferroni / Holm$p \le \alpha/m$; $p_{(k)} \le \alpha/(m - k + 1)$ step downFWER ≤ α (Holm never worse)
Benjamini–Hochberglargest $k$ with $p_{(k)} \le k\alpha/m$FDR ≤ α
Code it · Python

import numpy as np
from scipy import stats
from statsmodels.stats.multitest import multipletests

rng = np.random.default_rng(0)

# 1) Sample ratio mismatch: planned 50/50, observed 5,200 vs 4,800 users
chi2, p = stats.chisquare([5200, 4800], f_exp=[5000, 5000])
print(f"SRM check: chi2={chi2:.2f}, p={p:.1e}")
# SRM check: chi2=16.00, p=6.3e-05     -> do not read the metrics; find the leak

# 2) Peeking: 10,000 A/A tests, 30 daily looks, "stop when |z| > 1.96"
R, days = 10_000, 30
daily = rng.standard_normal((R, days))                       # each day adds independent noise, no real effect
z = daily.cumsum(axis=1) / np.sqrt(np.arange(1, days + 1))   # z-statistic after each day
print("false positives, one look at day 30:", np.mean(np.abs(z[:, -1]) > 1.96))
print("false positives, a look every day:  ", np.mean((np.abs(z) > 1.96).any(axis=1)))
# false positives, one look at day 30: 0.0535      (theory 0.05)
# false positives, a look every day:   0.2844      (theory about 0.28)

# 3) Bonferroni over 5 looks (days 6, 12, 18, 24, 30): alpha/5 per look
z5 = z[:, 5::6]
c_bonf = stats.norm.ppf(1 - 0.025 / 5)
print("5 looks at 1.96:", np.mean((np.abs(z5) > 1.96).any(axis=1)),
      f"| Bonferroni at {c_bonf:.3f}:", np.mean((np.abs(z5) > c_bonf).any(axis=1)))
# 5 looks at 1.96: 0.1426 | Bonferroni at 2.576: 0.0328     (safe, but below 5%: conservative)

# 4) A Bayesian rule checked daily: "ship B when P(B > A | data) > 0.95"
#    A/A test: both arms convert at 10%, 500 users per arm per day, Beta(1, 1) priors
R, m, p0 = 4000, 500, 0.10
kA = rng.binomial(m, p0, (R, 20)).cumsum(axis=1)
kB = rng.binomial(m, p0, (R, 20)).cumsum(axis=1)
n = m * np.arange(1, 21)
aA, bA, aB, bB = 1 + kA, 1 + n - kA, 1 + kB, 1 + n - kB      # Beta posterior parameters, every day
mean = lambda a, b: a / (a + b)
var = lambda a, b: a * b / ((a + b) ** 2 * (a + b + 1))
# Normal approximation of P(theta_B > theta_A | data); with thousands of users it matches Monte Carlo to ~0.001
prob_B = stats.norm.cdf((mean(aB, bB) - mean(aA, bA)) / np.sqrt(var(aA, bA) + var(aB, bB)))
print("ship B in an A/A test, one look on day 20:", np.mean(prob_B[:, -1] > 0.95))
print("ship B in an A/A test, a look every day:  ", np.mean((prob_B > 0.95).any(axis=1)))
# ship B in an A/A test, one look on day 20: 0.05125
# ship B in an A/A test, a look every day:   0.205    <- every posterior was computed correctly;
#                                                         the daily stopping RULE ships a useless B 4x as often

# 5) Many metrics: Bonferroni, Holm and Benjamini-Hochberg on 10 p-values
pvals = [0.001, 0.004, 0.006, 0.02, 0.04, 0.06, 0.2, 0.3, 0.5, 0.8]
for method in ["bonferroni", "holm", "fdr_bh"]:
    reject, p_adj, _, _ = multipletests(pvals, alpha=0.05, method=method)
    print(f"{method:10s} rejects {reject.sum()}  adjusted p (first 5) = {np.round(p_adj[:5], 3)}")
# bonferroni rejects 2  adjusted p (first 5) = [0.01 0.04 0.06 0.2  0.4 ]
# holm       rejects 3  adjusted p (first 5) = [0.01  0.036 0.048 0.14  0.24 ]
# fdr_bh     rejects 4  adjusted p (first 5) = [0.01 0.02 0.02 0.05 0.08]
Test yourself

1. Planned 50/50; you see 5,200 users in A and 4,800 in B, and B wins on conversion with p = 0.01. What should you do first?

$\chi^2 = 16$, $p \approx 0.00006$: a sample ratio mismatch. Some non-random group of users leaked out of B, so the arms are no longer comparable. Reweighting fixes counts, not composition; running longer does not remove the leak.

2. An A/A test is checked every day for 30 days and stopped at the first p < 0.05. About how often does it declare a winner?

Simulation (and the classic Armitage table) gives about 28%. 78.5% $= 1 - 0.95^{30}$ would be right only if the 30 looks were independent; they share most of their data, so the true rate is lower, but still far above 5%.

3. Your Bayesian dashboard shows $P(\theta_B \gt \theta_A \mid D)$ daily. Which statement is correct?

The stopping rule does not enter the likelihood, so the posterior is unaffected. The rule's operating characteristics are another matter: in A/A tests with flat priors it ships B about 5% of the time with one look and about 21% with 20 daily looks. A skeptical prior (centred on no effect) ships less often in A/A tests, but it does not remove the inflation.

4. Which experiment is most likely to violate SUTVA through interference?

Treated riders book more rides and use up the shared pool of drivers, so control riders wait longer: one rider's treatment changes another rider's outcome. The others have no shared resource or spillover channel worth worrying about.

5. What does the Benjamini–Hochberg procedure control?

BH controls the false discovery rate $E[V/\max(R,1)]$. The probability of at least one false positive is the FWER, controlled by Bonferroni and Holm. With many tests, BH usually makes some false claims, but they are a small share of its discoveries.

6. True effect +3 pts for anyone who sees B. 30% of control users also see B; every treated user sees B. Under the simple dilution model, what do you measure?

Measured $= \tau(e_T - e_C) = 3 \times (1 - 0.3) = 2.1$ pts. You would also need about $1/0.7^2 \approx 2$ times as many users to detect it.

Practice problems

A. A three-arm test (A, B, C) planned 1/3 each. Counts: 10,300, 9,900 and 9,800. Is there an SRM?
  1. $N = 30{,}000$, so $E_g = 10{,}000$ for each arm.
  2. $\chi^2 = 300^2/10{,}000 + 100^2/10{,}000 + 200^2/10{,}000 = 9 + 1 + 4 = 14$.
  3. df $= 3 - 1 = 2$. For 2 degrees of freedom the tail probability is exactly $e^{-\chi^2/2} = e^{-7} \approx 0.0009$.
  4. $p \lt 0.001$: SRM alarm. Look at which arm is off (A has too many users) and split the counts by day and platform to find the leak.
B. Why is the false-positive rate with two equally spaced looks 8.3%, and not $1 - 0.95^2 = 9.75\%$?

$1 - 0.95^2$ assumes the two looks are independent. They are not: the final z-statistic contains all the data of the first look plus an equal amount of new data, so the two z-values have correlation $\sqrt{1/2} \approx 0.71$. When the first look is not significant, the second one starts from a similar position, so it is less likely to be significant than a fresh test would be. The exact two-dimensional Normal calculation gives 8.3%: less than 9.75%, but well above the promised 5%.

C. 2,000 users per arm; 40% are high-intent (convert 20%), 60% low-intent (5%). B has no effect on users who stay, but 25% of low-intent B users leave before being counted as active. Compare "conversion per active user" and ITT.
  1. A: $800 \times 0.20 + 1{,}200 \times 0.05 = 160 + 60 = 220$ conversions of 2,000: $11.0\%$.
  2. B stays: 800 high-intent and $1{,}200 \times 0.75 = 900$ low-intent: 1,700 active users; conversions $160 + 45 = 205$.
  3. Per active user: $205/1{,}700 = 12.06\%$, a "lift" of $+1.06$ pts.
  4. Per assigned user (ITT): $205/2{,}000 = 10.25\%$, a lift of $-0.75$ pts. B lost sales; the survivor metric hid it by dropping low converters from B's denominator.
D. (Interview) "We use a Bayesian A/B framework. Can we look at the results every day?"

"Yes, we can look: each day's posterior is a valid summary given our model and prior, because the stopping rule does not enter the likelihood. What we must not do is improvise the decision. A rule like 'ship the first day $P(B \gt A) \gt 0.95$' is a sequential procedure; in A/A simulations with flat priors it ships B about 21% of the time with 20 daily looks instead of 5%. If our prior matched the real distribution of effects, the shipped variants would still be right about as often as the posterior says; if most real changes do nothing and our prior doesn't know that, it over-promises, and more looks make it worse. So we fix the rule, the minimum duration and the look schedule in advance, use priors informed by past experiments, decide on practically meaningful lifts, and simulate the procedure to know its error rates."

E. Six p-values: 0.003, 0.008, 0.012, 0.03, 0.045, 0.2. At α = 0.05, how many does each of Bonferroni, Holm and BH reject?
  1. Bonferroni: bar $0.05/6 \approx 0.00833$: $0.003$ ✓, $0.008$ ✓, $0.012$ ✗. 2.
  2. Holm: $0.003 \le 0.05/6 = 0.00833$ ✓; $0.008 \le 0.05/5 = 0.01$ ✓; $0.012 \le 0.05/4 = 0.0125$ ✓; $0.03 \gt 0.05/3 \approx 0.0167$ ✗, stop. 3.
  3. BH: bars $k \times 0.05/6$: $0.0083, 0.0167, 0.025, 0.0333, 0.0417, 0.05$. $p_{(4)} = 0.03 \le 0.0333$ ✓; $p_{(5)} = 0.045 \gt 0.0417$; $p_{(6)} = 0.2 \gt 0.05$. Largest $k = 4$: 4.
  4. Check with multipletests: adjusted Holm p = 0.018, 0.04, 0.048, 0.09, …; BH = 0.018, 0.024, 0.024, 0.045, 0.054, 0.2.
F. A food-delivery app wants to test a new courier-assignment algorithm. Why is a user-level A/B test a bad idea, and what would you do instead?

Couriers are a shared resource: if the new algorithm grabs the best couriers for treated orders, control orders get slower deliveries. The user-level difference then mixes "B is better" with "B steals from A", which is not the effect of launching B (SUTVA's no-interference part fails). Better: a switchback design (each city alternates between A and B in time slots, e.g. 1–2 hours, randomized), or randomize whole cities (cluster randomization). Analyse at the level of the slot or city, discard or model the first minutes after each switch (carryover), and expect wider intervals because there are fewer independent units.

Chapter 5.12 · Syllabus Module 80

Causal thinking

"Users who use feature X convert twice as often" is a fact about data. "Feature X makes users convert" is a claim about what would happen if you changed something. This chapter gives you the language to tell the two apart (potential outcomes, counterfactuals, confounding), shows why a randomized experiment is the cleanest way to get causal answers, and teaches the two tools that matter most for your A/B work: regression adjustment and CUPED. At the end you meet three tools for when you cannot randomize: propensity scores, difference-in-differences and instrumental variables.

  • Explain three ways a correlation can appear without causation (reverse causation, a common cause, selection), and why correlation ≠ causation
  • Use potential outcomes $Y_i(1)$, $Y_i(0)$; define the counterfactual and the fundamental problem of causal inference
  • Tell individual, average (ATE) and conditional (CATE) treatment effects apart
  • Show why randomization makes the difference in means an unbiased estimate of the ATE, and what confounding and selection bias do without it (with DAG pictures)
  • Derive CUPED: $\theta = Cov(Y, X)/Var(X)$ and the variance factor $1 - \rho^2$; connect it to regression adjustment
  • Know what propensity scores, difference-in-differences and instrumental variables do, and the one assumption each depends on

What we need from earlier chapters: expectation, variance and their rules (Chapter 4.5); conditional expectation and Simpson's paradox (Chapter 4.6); covariance and correlation $\rho$ (Chapter 4.15); standard errors (Chapter 5.5); power and sample size (Chapter 5.7); experiment design (Chapter 5.10); SUTVA and interference, defined in Chapter 5.11. Regression is taught fully in Chapter 5.13; here we only need "fit a straight line by least squares". Words: the treatment $T$ is the thing we might change (a feature, a discount, an email); $T_i = 1$ if unit $i$ got it and $0$ if not; the outcome $Y$ is what we care about (conversion, spend); a covariate $X$ is any other measured characteristic of a unit (device, last month's spend).

Correlation is not causation: three ways to be fooled core

Cities with more fire fighters at a fire have more damage. Nobody concludes that fire fighters cause damage: big fires cause both. In product data the same trap is everywhere, only less obvious. "Users who use the wishlist convert twice as often" sounds like "build more wishlist features". But engaged shoppers use every feature and buy more. The wishlist may do nothing at all.

A correlation means "these two numbers move together in the data". A causal effect means "if I change one, the other changes". Data passively collected show the first; decisions need the second.

There are three classic ways to get the first without the second: the arrow points the other way (reverse causation), a third thing drives both (a common cause, or confounder), or you only look at a filtered group (selection). Plus plain chance, with small data.

Three ways to say it:

  • Picture: two puppets moving together because one hand holds both strings.
  • Numbers: feature users convert at 19% and non-users at 8% (11 pts apart), but the feature's real effect is only 2 pts; engagement explains the other 9.
  • Slogan: seeing is not doing.

A wishlist that looks magical. Half of the users are highly engaged. Engaged users convert at 20% without the wishlist, others at 5%. Using the wishlist truly adds 2 pts for anyone. But 80% of engaged users use the wishlist, and only 20% of the others do.

  1. Who are the wishlist users? Engaged: $0.5 \times 0.8 = 0.40$ of all users; others: $0.5 \times 0.2 = 0.10$. So $0.40/0.50 = 80\%$ of wishlist users are engaged.
  2. Their conversion: $0.8 \times (20 + 2) + 0.2 \times (5 + 2) = 17.6 + 1.4 = 19.0\%$.
  3. Non-users: engaged $0.5 \times 0.2 = 0.10$, others $0.5 \times 0.8 = 0.40$, so only 20% are engaged. Conversion: $0.2 \times 20 + 0.8 \times 5 = 4 + 4 = 8.0\%$.
  4. The naive comparison says $19 - 8 = 11$ pts.
  5. Inside each group the gap is the real 2 pts: engaged $22 - 20 = 2$; others $7 - 5 = 2$. The other 9 pts come from comparing mostly-engaged users with mostly-unengaged users.
  • Association (correlation): $T$ and $Y$ are statistically related in the data, e.g. $E[Y \mid T = 1] \ne E[Y \mid T = 0]$.
  • Causal effect: changing $T$ (by intervention) changes $Y$. It is a statement about what would happen under different actions, made precise with potential outcomes in the next section.
  • Association without causation arises from:
    1. Reverse causation: $Y$ causes $T$ (users who already decided to buy read the reviews page, so "reading reviews" correlates with buying).
    2. Confounding (common cause): a variable $U$ causes both $T$ and $Y$ (engagement drives feature use and conversion).
    3. Selection: the data only contain units filtered by something related to $T$ and $Y$ (only users who reached checkout).
    4. Chance: small samples and many comparisons (Chapter 5.11).
  • A DAG (directed acyclic graph) draws each variable as a node and each direct cause as an arrow; "acyclic" means no arrow path leads back to where it started. DAGs make these four stories easy to see.
Why do we need it?

Product, marketing and pricing decisions are about changing something. Dashboards full of correlations ("users who do X retain better") lead teams to build things that do nothing, or to stop things that help, unless someone asks "would Y change if we changed X?".

Where is it used?

Feature-adoption analyses, marketing attribution, "users who got the email bought more" reports, observational health studies, policy evaluation, and ML feature importance (a feature can predict well without causing anything, which matters when the model is used to choose actions).

How is it used?

For every correlation, sketch the DAG: list plausible common causes, the possible reverse arrow and any filtering of the data. Then choose a design: randomize if you can (A/B test); otherwise adjust for measured confounders or use a quasi-experiment (difference-in-differences, instrumental variables), and state the assumption you are relying on.

1 · causation T Y feature → purchase 2 · reverse T Y intent to buy → reads reviews 3 · common cause U T Y engagement → both 4 · selection T Y S only rows with S = 1 kept Only diagram 1 means "changing T changes Y". All four produce a correlation between T and Y.
Four stories behind one correlation. Arrows are direct causes. In 2 the arrow is reversed; in 3 a common cause U drives both; in 4 the data only keep units that passed a filter S that depends on both.

Each dot is a user: feature uses per week (x) and spend next month (y). With the hidden variable off, you see one cloud and one steep line (the naive slope). Tick show engagement: the users split into three engagement tiers, and inside each tier the line is much flatter: that flat slope is the real effect. Set true effect to 0: the overall slope stays steep. Set hidden-variable strength to 0: the two slopes agree.

"The correlation is very strong (r = 0.8), so it must be causal."

A strong common cause produces a strong correlation. Strength says nothing about direction or about a hidden third variable.

"We controlled for 30 variables, so it is causal now."

Adjustment removes confounding only by the variables you measured and adjusted for correctly. One unmeasured common cause (motivation, intent) is enough to bias the answer. And adjusting for the wrong variables (a mediator or a collider) can add bias.

In your forecasting model, a regressor's coefficient $\beta$ tells you how the forecast moves with that regressor, given the other columns. It is a predictive association learned from history. If promotions were always run in high-demand weeks, the promotion coefficient mixes "promotions raise demand" with "promotions happen when demand is already high". That is fine for forecasting under the same habits, but not a reliable answer to "what if we run an extra promotion?". In the A/B framework, by contrast, randomization lets $P(\theta_B \gt \theta_A \mid D)$ be read causally (section 4).

"Users who adopt feature X retain 30% better, so feature X increases retention by 30%."

"Adopters retain 30% better, but adopters are self-selected and probably more engaged to begin with. The causal effect of X could be anywhere from 0 to 30%, or even negative. To know, I'd run an experiment that randomizes access or encouragement, or at least adjust for pre-adoption engagement and say which confounders could remain."

Model answer: "Correlation can come from causation, reverse causation, a common cause or selection. Only a design that breaks the other three (randomization is the cleanest) lets me read the difference as a causal effect."

Association: $E[Y \mid T=1] \ne E[Y \mid T=0]$. Causation: changing $T$ changes $Y$.

Fake associations: reverse causation, common cause (confounding), selection, chance.

Habit: draw the DAG before trusting a correlation.

Quick check: hospitals with more doctors per patient have higher death rates. Which story is most likely?

A common cause: hospitals that treat the sickest patients (severity) get more doctors and have more deaths. Severity confounds the comparison; compare hospitals within the same patient-severity level.

Potential outcomes and the fundamental problem core

Imagine every user walking around with two numbers in their pocket. The first is what they would spend if they saw the old page. The second is what they would spend if they saw the new page. The effect of the new page on this user is simply the second number minus the first.

The catch: each user sees only one page. You get to open one pocket, never both. The other number, "what would have happened otherwise", is the counterfactual: real, meaningful, and never observed. This is the fundamental problem of causal inference. Every causal method is a trick for filling in the missing pocket, usually for a group on average rather than for one person.

Three ways to say it:

  • Picture: two parallel worlds for each user; you only live in one of them.
  • Numbers: Ana would spend 7 dollars with the new page and 5 with the old one (effect +2), but you see only one of those numbers.
  • Slogan: every causal question is a comparison with a world you cannot see.

Six users, God's view. Spend next week (dollars) under the old page $Y(0)$ and the new page $Y(1)$:

user123456
$Y_i(0)$538462
$Y_i(1)$739753
$\tau_i = Y_i(1) - Y_i(0)$+20+1+3−1+1
  1. Average effect: $(2 + 0 + 1 + 3 - 1 + 1)/6 = 6/6 = 1$ dollar. Equivalently, mean of $Y(1)$ minus mean of $Y(0)$: $34/6 - 28/6 = 1$.
  2. Real world: users 1, 4, 5 get the new page; users 2, 3, 6 the old one. We see $Y(1)$ = 7, 7, 5 and $Y(0)$ = 3, 8, 2.
  3. Difference in means: $(7 + 7 + 5)/3 - (3 + 8 + 2)/3 = 19/3 - 13/3 = 2.0$. Not 1: this split happened to put the users with big effects in the treated group.
  4. There are 20 ways to choose 3 of 6 users. Their estimates range from $-2$ to $+4$, and their average is exactly $1$, the true average effect. A random split is right on average, not every time.
  • Potential outcomes: for unit $i$, $Y_i(1)$ is the outcome it would have under treatment and $Y_i(0)$ under control. Both exist (conceptually) before treatment is assigned.
  • Observed outcome: $Y_i = T_i Y_i(1) + (1 - T_i) Y_i(0)$. We see exactly one of the two.
  • Counterfactual: the potential outcome under the treatment the unit did not receive.
  • Individual treatment effect: $\tau_i = Y_i(1) - Y_i(0)$.
  • Fundamental problem of causal inference (Holland, 1986): $\tau_i$ is never observed, because $Y_i(1)$ and $Y_i(0)$ are never both observed for the same unit at the same time.

Sub-point: the hidden assumption, SUTVA. Writing just two potential outcomes per unit assumes SUTVA (Chapter 5.11): no interference (unit $i$'s outcome does not depend on other units' treatments) and no hidden versions of treatment. Under interference, user $i$'s outcome would depend on everyone's assignment, $Y_i(z_1, \dots, z_N)$, and "the" effect of the treatment on $i$ is no longer one number.

Why do we need it?

"Effect" is otherwise vague. Potential outcomes say exactly what is being estimated (a difference between two worlds), what is missing (the counterfactual), and what an estimation method must assume to fill the gap. Every method in this chapter is stated in this language.

Where is it used?

The Rubin causal model behind A/B testing, clinical trials and econometrics; uplift modelling in marketing (who changes behaviour because of a coupon); causal ML libraries such as DoWhy and EconML; and off-policy evaluation in recommender systems.

How is it used?

Write down the estimand first ("ATE of the new page on 7-day spend, among users who visit the page"), in terms of $Y(1)$ and $Y(0)$. Then choose a design whose assumptions make the missing potential outcomes estimable on average, and check SUTVA.

The table shows each user's two potential outcomes. In the real world you only see the bold one (the other is the hidden counterfactual). Press Assign at random several times: the difference in means jumps around (orange dot on the strip below). Tick God's view to reveal the hidden column and the individual effects. Press Show all 20 assignments: the estimates scatter from −2 to +4, but their average (purple) is exactly the true average effect, 1.

"The counterfactual is just the control group's outcome."

The counterfactual of a treated user is what that user would have done without treatment. The control group's average stands in for it only when the groups are comparable, which is exactly what randomization buys.

"With enough data we can measure each user's individual effect."

No amount of data reveals both potential outcomes of the same user at the same time. We can estimate averages, and averages within subgroups, but individual effects stay hidden without strong extra assumptions.

$Y_i(1), Y_i(0)$: outcomes in the two worlds; observed $Y_i = T_iY_i(1) + (1-T_i)Y_i(0)$.

$\tau_i = Y_i(1) - Y_i(0)$ is never observed (fundamental problem). The unseen one is the counterfactual.

Needs SUTVA (no interference, one version of treatment), see Chapter 5.11.

Quick check: in the six-user table, what would the estimate be if users 2, 3 and 5 were treated?

Treated $Y(1)$: 3, 9, 5, mean $17/3 \approx 5.67$. Control $Y(0)$ for users 1, 4, 6: 5, 4, 2, mean $11/3 \approx 3.67$. Estimate $= 2.0$. (Different users, same estimate as the example: the 20 estimates take only a few distinct values.)

Individual vs average treatment effects: ITE, ATE and CATE

A new checkout helps some users a lot, does nothing for others, and annoys a few. Each user has their own effect. We can never see one user's effect (previous section), but we can learn the average effect over a group, because an average of differences equals a difference of averages, and averages can be estimated from different people.

Averages can also be taken over smaller groups: "the average effect on mobile users", "on new users". These conditional averages show where the change helps and where it hurts.

Three ways to say it:

  • Picture: a crowd where each person carries a hidden plus or minus number; we can learn the crowd's average, and the average of each corner of the crowd, but not any one person's number.
  • Numbers: effects +3, −1, 0, +2 average to +1 (ATE), but one user out of four is harmed.
  • Slogan: a positive average does not mean everyone benefits.

Four users with individual effects (God's view): user 1 (mobile) +3, user 2 (desktop) −1, user 3 (desktop) 0, user 4 (mobile) +2.

  1. ATE: $(3 - 1 + 0 + 2)/4 = 4/4 = 1$.
  2. CATE for mobile: $(3 + 2)/2 = 2.5$. CATE for desktop: $(-1 + 0)/2 = -0.5$.
  3. Check: the ATE is the share-weighted average of the CATEs: $0.5 \times 2.5 + 0.5 \times (-0.5) = 1.25 - 0.25 = 1$.
  4. Decision: launching to everyone gains on average, but launching to mobile only gains $2.5$ per mobile user and avoids the $-0.5$ on desktop.
  5. Still, inside "mobile", user 1 and user 4 differ; the CATE is an average too.
  • Individual treatment effect (ITE): $\tau_i = Y_i(1) - Y_i(0)$. Never observed.
  • Average treatment effect (ATE): $\tau = E[Y(1) - Y(0)] = E[Y(1)] - E[Y(0)]$, the average over the population of interest. The second form (linearity of expectation) is why it can be estimated: $E[Y(1)]$ from treated units and $E[Y(0)]$ from control units.
  • Conditional average treatment effect (CATE): $\tau(x) = E[Y(1) - Y(0) \mid X = x]$, the ATE inside the subgroup with covariate value $x$ (segment, device, country). The ATE is the average of the CATEs weighted by group size.
  • Average treatment effect on the treated (ATT): $E[Y(1) - Y(0) \mid T = 1]$, the average over the units that actually got the treatment. In a randomized experiment ATT = ATE; with self-selection they can differ.
  • Heterogeneity: effects that vary across units. The share of users harmed, $P(\tau_i \lt 0)$, is not identified from an A/B test alone (it depends on how $Y(1)$ and $Y(0)$ are linked within a person, which we never see).
Why do we need it?

A launch decision usually rests on the ATE, but targeting, personalization and fairness questions need CATEs, and some decisions care about who is harmed. Knowing which quantity you estimate avoids claims like "everyone gains 1 dollar".

Where is it used?

A/B test scorecards (ATE), segment breakdowns (CATE by country or device), uplift models for targeting coupons, heterogeneous-effect methods such as causal forests and meta-learners (EconML, CausalML), and clinical trials (subgroup analyses, with the multiple-testing caveats of Chapter 5.11).

How is it used?

Estimate the ATE as the difference in means of a randomized test. Estimate pre-planned CATEs inside segments (with corrections or partial pooling). Never read a segment estimate or an ATE as an individual's effect, and say "on average" when you report.

The histogram shows the hidden individual effects of 400 users (teal = mobile, pink = desktop). The purple line is the ATE; dashed lines are the two CATEs. Raise the spread with the ATE fixed at +1: the average does not move, but more and more users are harmed. Raise the segment gap: the two CATEs separate. An A/B test estimates the purple line (and, with segments, the dashed lines), never the histogram.

"ATE = +1 dollar, so every user spends 1 dollar more."

The ATE is an average. With the same ATE, effects can be +1 for everyone or +5 for some and −3 for others. A plain A/B test cannot tell these apart.

"Effect in segment X was +3 with p = 0.04, so segment X is where it works."

Segment CATEs are noisier than the overall ATE and are many tests at once (Chapter 5.11). Plan them in advance, correct for multiplicity or pool partially, and confirm.

In an A/B framework like yours with hierarchical partial pooling across segments, each segment's effect parameter is a model of a CATE: $\theta_{B,g} - \theta_{A,g}$ for segment $g$. Partial pooling shrinks noisy segment CATEs toward the overall effect (the ATE-like population parameter) by an amount learned from how different the segments really are. The posterior of each segment's difference answers "what is the average effect in this segment?", never "what is the effect for this user?".

ITE $\tau_i = Y_i(1) - Y_i(0)$ (never seen) · ATE $= E[Y(1)] - E[Y(0)]$ · CATE $\tau(x) = E[Y(1) - Y(0) \mid X = x]$ · ATT = average over the treated.

ATE = size-weighted average of CATEs. A/B tests estimate averages, not individuals.

Trap: positive ATE does not mean no one is harmed; share harmed is not identified.

Quick check: segment A (70% of users) has CATE +2, segment B (30%) has CATE −1. What is the ATE?

$0.7 \times 2 + 0.3 \times (-1) = 1.4 - 0.3 = 1.1$.

Randomization makes the groups comparable core

Why is a coin flip so powerful? Because a coin knows nothing about the users. It cannot prefer engaged users, rich users, or users who were about to buy anyway. So, on average, the two groups it creates contain the same mix of every kind of user, including kinds you never measured.

Then the control group's average outcome is a fair stand-in for "what the treated group would have done without treatment" (their missing counterfactual), and the difference in means estimates the average effect.

If instead users choose the treatment, the choosers are different people (more engaged, more curious), and the difference in means mixes the effect with those differences.

Three ways to say it:

  • Picture: shuffle a deck and deal two hands: each hand has about the same share of aces, even if you never looked at the cards.
  • Numbers: 1,000 users, 300 power users, random 500/500 split: about 150 power users per arm (give or take 7). If users opt in, 63% of the treated are power users vs 10% of controls.
  • Slogan: randomize, and the control group becomes the treated group's missing world.
  1. Random split. 1,000 users, 300 of them power users; 500 are chosen at random for treatment. The number of power users in the treatment arm averages $500 \times 0.3 = 150$, with a standard deviation of about 7.2 (hypergeometric). So typically 143–157, close to the control arm.
  2. Self-selection. Now power users opt in with probability 0.8, others with 0.2. Treated: $300 \times 0.8 = 240$ power users and $700 \times 0.2 = 140$ others, so $240/380 \approx 63\%$ power users. Control: $60$ power users of $620$, about $10\%$.
  3. If power users spend 30 dollars and others 10 dollars without treatment, the self-selected comparison is inflated by roughly $(0.63 - 0.10) \times (30 - 10) \approx 10.7$ dollars, whatever the true effect.
  4. Why randomization is unbiased, in symbols: because $T$ is independent of $(Y(0), Y(1))$, $E[Y \mid T = 1] = E[Y(1) \mid T = 1] = E[Y(1)]$ and $E[Y \mid T = 0] = E[Y(0)]$. Subtracting gives the ATE.

Randomization: treatment is assigned by a chance mechanism that does not depend on the units' characteristics or potential outcomes: $T \perp (Y(0), Y(1))$ ("$\perp$" reads "is independent of").

  • Consequence: $E[\bar Y_T - \bar Y_C] = E[Y(1)] - E[Y(0)] = $ ATE. The difference in means is unbiased, and its standard error is the usual $\sqrt{s_T^2/n_T + s_C^2/n_C}$ (Chapter 5.5).
  • Balance holds on average for every covariate, measured or not. In any one experiment the arms differ a little by chance; that chance imbalance is exactly what the standard error accounts for.
  • Randomization does not fix SUTVA violations, attrition after assignment, or non-compliance (Chapter 5.11); it fixes who gets treated.
  • Without randomization, the difference in means equals ATT plus selection bias $E[Y(0) \mid T=1] - E[Y(0) \mid T=0]$: how different the groups would have been even without treatment.
Why do we need it?

It is the only design that removes confounding by unmeasured variables (motivation, intent, taste) without any modelling. Every other method in this chapter needs an assumption you cannot fully check; a randomized test needs only a working coin and SUTVA.

Where is it used?

A/B tests and multivariate tests in product, marketing and pricing; clinical trials (randomized controlled trials, RCTs); field experiments in economics; holdout groups for measuring ads and recommendations.

How is it used?

Assign by hashing a stable unit id with an experiment-specific salt into buckets (Chapter 5.10). Check balance on pre-experiment covariates and SRM (Chapter 5.11). Compare means (or posteriors) between arms, analysing everyone as assigned.

100 fixed users: 30 power users (base spend about 30 dollars) and 70 others (about 10 dollars); the true average effect is +2 dollars (purple line). Each run assigns treatment again and computes the difference in means; the histogram collects 200 runs. With coin flip the pile is centred on +2: unbiased, with some spread. Switch to users choose: power users opt in, and the whole pile jumps far to the right. More runs never fix that bias.

"Randomization makes the two groups identical."

It makes them identical in distribution, on average over possible randomizations. Any single split has small chance differences; the standard error (or the posterior width) accounts for them. Adjusting for a strong pre-treatment covariate (CUPED, below) reduces their impact.

"We randomized, so any later filtering of the data is fine."

Filtering on something that happened after assignment (clicked, stayed active, reached checkout) can undo the randomization (section 6). Analyse units as assigned.

In an A/B framework like yours, randomization is what turns the posterior statement $P(\theta_B \gt \theta_A \mid D)$ into a causal one: because users were assigned by a coin, $\theta_B - \theta_A$ is the effect of showing B instead of A on the conversion rate, not just a difference between two kinds of users. The Bayesian model adds a careful description of uncertainty; the causal reading comes entirely from the design. Break the design (SRM, attrition, interference) and the same posterior loses its causal meaning.

Randomization: $T \perp (Y(0), Y(1))$ ⇒ $E[\bar Y_T - \bar Y_C] = $ ATE (unbiased).

Without it: difference = ATT + selection bias $E[Y(0) \mid T=1] - E[Y(0) \mid T=0]$.

Balances measured and unmeasured covariates on average; not in every single split.

Quick check: in a randomized test the treatment arm happens to have 3% more iOS users than control. Is the experiment broken?

Not by itself: small chance imbalances are expected, and the standard error covers them. Check it is within chance (a chi-square test, like an SRM check on the iOS share). A large, significant imbalance would suggest an assignment or logging problem. Adjusting for pre-experiment covariates can also reduce the noise such imbalances add.

Confounding: the common cause, drawn as a DAG core

A confounder is a variable that pushes on both the treatment and the outcome. Engagement makes users adopt features and makes them buy. In the picture (a DAG) there are two routes from "feature use" to "conversion": the real causal arrow, and a back route through engagement (feature use ← engagement → conversion). A naive comparison adds up both routes.

There are two ways to close the back route. Randomize the treatment, which deletes the arrow from engagement into feature use. Or adjust: compare users with the same engagement, so engagement cannot differ between the compared groups. Adjusting only works for confounders you have measured.

Three ways to say it:

  • Picture: two pipes from T to Y, one real and one back pipe through U; the meter at Y measures the total flow.
  • Numbers: naive gap 11 pts = 2 pts of real effect + 9 pts that flow through engagement.
  • Slogan: block the back door, or randomize the front door.

Back to the wishlist (section 1): engaged users convert at 20%, others at 5%; the wishlist adds 2 pts; 80% of engaged and 20% of other users use it; half the users are engaged.

  1. Naive: wishlist users 19.0%, non-users 8.0%, gap 11 pts.
  2. Adjusted (stratified): engaged $22 - 20 = 2$; others $7 - 5 = 2$; average weighted by group size $0.5 \times 2 + 0.5 \times 2 = 2$ pts.
  3. Confounding bias $= 11 - 2 = 9$ pts. It equals the difference in engagement mix (80% vs 20% engaged, a gap of 0.6) times engagement's effect on conversion (15 pts): $0.6 \times 15 = 9$.
  4. Randomized alternative: give the wishlist to a random half. Each arm is 50% engaged, so the gap is exactly the 2 pts.
  • Confounder: a variable $U$ that causally affects both the treatment $T$ and the outcome $Y$ (and is not itself caused by $T$).
  • Backdoor path: a path from $T$ to $Y$ that starts with an arrow into $T$ (here $T \leftarrow U \to Y$). Open backdoor paths create association that is not causation.
  • Adjustment (stratification, regression, matching, weighting) on a set of measured variables $X$ that blocks all backdoor paths gives the causal effect: $\text{ATE} = \sum_x P(X = x)\,\big(E[Y \mid T=1, X=x] - E[Y \mid T=0, X=x]\big)$. The assumption that $X$ contains every confounder is called no unmeasured confounding (also "ignorability" or "conditional exchangeability"). It cannot be tested from the data.
  • Do not adjust for a mediator ($T \to M \to Y$: adjusting removes part of the effect you want) or a collider ($T \to S \leftarrow Y$: adjusting creates a fake association; next section).
Why do we need it?

Confounding is the main reason observational comparisons mislead. Naming the confounders and drawing the DAG tells you whether an adjustment can work, what to adjust for, and what must be left alone.

Where is it used?

Observational studies in medicine and economics; product analytics ("feature adopters vs non-adopters"); marketing attribution; DoWhy and similar causal libraries, which take a DAG and tell you the valid adjustment sets; ML models whose predictions are used to choose actions.

How is it used?

Draw the DAG with the domain experts. Find the variables that block every backdoor path and were measured before treatment. Adjust for them (stratify or regress). Report the result as conditional on "no unmeasured confounding" and, if possible, run a sensitivity analysis.

Observational data engagement U feature use T conversion Y +2 back door T ← U → Y is open Randomized experiment engagement U feature T conversion Y coin only the coin decides T: no back door
Left: engagement causes both feature use and conversion, so the naive comparison flows through the back door. Right: in an experiment a coin decides T, the arrow U → T disappears (red cross), and only the causal arrow T → Y connects them.

Set the true effect of the feature and how strongly engagement drives both uptake and conversion. Red: the naive comparison of users vs non-users. Green: the stratified (adjusted) estimate. Blue: a randomized experiment. Set the true effect to 0: the naive bar still shows a big "effect". Make uptake equal in both groups (no confounding): all three agree.

"Adjust for every variable you have, to be safe."

Adjusting for a mediator (e.g. "pages viewed", which the feature itself changes) removes part of the real effect; adjusting for a collider creates bias. Adjust for pre-treatment confounders, chosen with a DAG.

"After adjustment the estimate is causal."

Only if no confounder is missing. That assumption cannot be checked in the data; say it out loud and, where possible, test sensitivity.

"A confounder is any variable correlated with the outcome."

A confounder causes both the treatment and the outcome (a common cause). A variable that only predicts the outcome is a precision covariate: adjusting for it reduces noise (CUPED) but does not remove bias.

Model answer: "Confounding happens when a common cause opens a backdoor path T ← U → Y. Randomization removes it by design; otherwise I need to measure and adjust for a set of pre-treatment variables that blocks every backdoor path, without conditioning on mediators or colliders."

Confounder: $U \to T$ and $U \to Y$. Backdoor path $T \leftarrow U \to Y$.

Adjusted ATE $= \sum_x P(x)\big(E[Y \mid T=1, x] - E[Y \mid T=0, x]\big)$, valid if no unmeasured confounding.

Never adjust for mediators or colliders. Randomization deletes the arrow into T.

Quick check: you want the effect of a "premium" badge on sales. Sellers with more reviews are more likely to get the badge and also sell more. Which variable must you adjust for, and which (if any) should you not adjust for: "number of reviews before the badge" or "page views after the badge"?

Adjust for reviews before the badge: a pre-treatment confounder. Do not adjust for page views after the badge: the badge probably raises views, which raise sales, so it is a mediator, and adjusting would hide part of the badge's effect.

Selection bias: when the filter creates the pattern

Among famous actors, the very good-looking ones seem to be worse at acting. Is beauty bad for talent? No: to become famous you need talent or looks (or both). Among the famous, someone with little talent must have had great looks to get in, and the other way round. The filter "is famous" creates a negative link between two things that are unrelated in the whole population.

The same happens in product data. Analyse only users who reached checkout, only users who clicked an email, only stores that survived a year, and the filter itself can create, hide or flip patterns. When the filter depends on the treatment, it can even undo the randomization of an A/B test.

Three ways to say it:

  • Picture: a sieve that only lets through stones whose length plus width is large; among the stones that pass, long ones tend to be thin.
  • Numbers: among clickers, B converts 20% vs A's 30%; per assigned user, B converts 4% vs A's 3%.
  • Slogan: look at who got into your table before you look at the table.

Filtering on something the treatment changed. 1,000 users per arm. In A, 100 high-intent users click the product page and 30 of them buy. B adds a banner that also makes 100 low-intent users click; the high-intent clickers still buy 30, the low-intent ones buy 10.

  1. A, among clickers: $30 / 100 = 30\%$.
  2. B, among clickers: $(30 + 10) / (100 + 100) = 40 / 200 = 20\%$. "B lowers conversion by 10 pts"?
  3. Per assigned user (ITT): A $30/1{,}000 = 3\%$; B $40/1{,}000 = 4\%$. B raises conversion by 1 pt.
  4. "Clicked" is caused by the treatment (the banner) and related to the outcome (intent). Filtering on it compares different kinds of users.
  • Selection bias: bias from analysing a subset of units chosen by a process related to both the treatment (or exposure) and the outcome. Survivorship bias is a special case (only "survivors" are seen).
  • Collider: a variable caused by two others, $T \to S \leftarrow Y$ (or $A \to S \leftarrow B$). Conditioning on a collider (filtering on it, or adjusting for it) creates an association between its causes even if they are independent.
  • Post-treatment selection: in an experiment, restricting the analysis by a variable measured after assignment (clicked, still active, reached checkout). If the treatment affects that variable, the restricted arms are no longer comparable.
  • In the two-number language of section 4: the comparison picks up $E[Y(0) \mid \text{selected}, T=1] - E[Y(0) \mid \text{selected}, T=0] \ne 0$.
Why do we need it?

Selection is often invisible: the table you are given already contains only some users. Knowing the collider pattern explains surprising negative correlations and stops "per clicker" or "per buyer" analyses from reversing the sign of an effect.

Where is it used?

Funnel metrics (conversion among checkout visitors), email tests analysed on openers, survivorship in cohort retention, ratings from users who chose to rate, hospital-based medical studies, and training data for ML models built only from past approved cases (loan approvals, shown recommendations).

How is it used?

Ask "how did these rows get into my table, and could that depend on the treatment or the outcome?". In experiments, define the population at assignment and use ITT metrics; if you must analyse a funnel step, use a triggering rule applied identically in both arms (Chapter 5.10).

300 products with two independent traits: how attractive the price is (x) and how fast the page loads (y). A product "sells" only if price + speed is above the bar. Blue dots pass the filter, grey dots do not. With the bar at its lowest, everything passes and the correlation is near 0. Raise the bar: among the sellers (blue) a clear negative correlation appears: slow pages "go with" good prices, only because of the filter.

"We'll measure conversion only among users who saw the checkout page; that's the relevant population."

If the treatment changes who reaches checkout, the checkout visitors differ between arms. Use per-assigned-user metrics, or a trigger condition that is applied identically and is not affected by the treatment.

"Adjusting for more variables can only reduce bias."

Adjusting for a collider (a common effect) opens a fake path and adds bias.

Collider: $A \to S \leftarrow B$. Filtering or adjusting on $S$ links $A$ and $B$ even if independent.

Post-treatment filtering in an A/B test (clicked, active, reached checkout) breaks randomization.

Fix: analyse units as assigned (ITT); trigger only on pre-treatment or treatment-independent conditions.

Quick check: a study of restaurants that survived 5 years finds that those with cheaper rents had worse food. Why might that be misleading?

Survival depends on both rent and food quality (a collider). A restaurant with high rent must have had great food to survive; one with cheap rent could survive with average food. Among survivors, rent and food look linked even if they are unrelated among all restaurants.

Regression adjustment: compare like with like, and cut the noise core

In an experiment, users differ a lot before you do anything: some spend 200 dollars a month, some 5. That variety is noise in the comparison. If you know each user's spend last month (before the experiment), you can compare each user with what you would have expected from them, instead of comparing raw totals. The treatment effect is the same; the noise around it is much smaller.

Regression adjustment does this by fitting one line per arm through "outcome vs pre-period covariate" (with the same slope) and reading the effect as the vertical gap between the lines. In observational data the same tool is also used to remove confounding by measured covariates, but then it needs the no-unmeasured-confounding assumption.

Three ways to say it:

  • Picture: two parallel lines through two clouds; the effect is the distance between the lines, not between the clouds' raw averages.
  • Numbers: with correlation 0.8 between pre and post spend, the standard error of the effect falls from 2.0 to about 1.2 dollars.
  • Slogan: explain away the predictable part, keep the effect.

An experiment on spend, 50 users per arm, pre-period spend $X$ with standard deviation 10 dollars, experiment-period spend $Y$ with standard deviation 10 dollars, correlation $\rho = 0.8$ between them, true effect +2 dollars.

  1. Unadjusted: difference in means. Its standard error is $\sqrt{10^2/50 + 10^2/50} = \sqrt{4} = 2.0$ dollars. The effect (+2) is only 1 SE: hard to detect.
  2. Model: $Y_i = \alpha + \tau T_i + \beta (X_i - \bar X) + \varepsilon_i$. The part $\beta(X_i - \bar X)$ explains $\rho^2 = 64\%$ of the variance of $Y$.
  3. Remaining noise: $\sqrt{1 - 0.64} \times 10 = 6$ dollars, so the SE of $\hat\tau$ is about $\sqrt{6^2/50 + 6^2/50} = \sqrt{1.44} = 1.2$ dollars.
  4. Same estimand (the ATE), 40% smaller SE, equivalent to having $1/(1 - 0.64) \approx 2.8$ times as many users.

Regression adjustment (ANCOVA in experiments): fit by least squares

$$Y_i = \alpha + \tau T_i + \beta\,(X_i - \bar X) + \varepsilon_i ,$$

where $X_i$ is a pre-treatment covariate. The coefficient $\hat\tau$ estimates the treatment effect.

  • In a randomized experiment, $\hat\tau$ is consistent for the ATE with or without $X$; including a predictive $X$ shrinks its variance by about the factor $1 - \rho^2$, where $\rho$ is the correlation between $X$ and $Y$. Adding the interaction $T_i (X_i - \bar X)$ (Lin's estimator) protects against the case where the slope differs between arms.
  • In observational data, $\hat\tau$ is causal only if $X$ contains every confounder and the model form is right.
  • $X$ must be measured before treatment (or be unaffected by it). Adjusting for a post-treatment variable can bias the estimate (section 6).
  • Least squares itself is taught in Chapter 5.13.
Why do we need it?

Experiments on noisy metrics (revenue, sessions, time spent) need huge samples. Adjusting for a strong pre-period covariate gives the same answer with far fewer users, or a tighter answer with the same users, at almost no cost.

Where is it used?

ANCOVA in clinical trials (adjusting for baseline measurements), A/B testing platforms (regression adjustment and CUPED, next section), economics field experiments, and observational studies (adjusting for measured confounders).

How is it used?

Choose pre-treatment covariates that predict the outcome (last month's value of the same metric is usually best). Fit smf.ols("y ~ t + x_centered", data) (optionally with t:x_centered), read the coefficient on t and its robust standard error (cov_type="HC1").

Left: one experiment. Blue = control, orange = treatment; x = spend before the experiment, y = spend during it. The two parallel lines are the adjusted fit; their vertical gap is $\hat\tau$. Right: the estimates from 100 repeated experiments: unadjusted (red) and adjusted (green), with the truth +2 (purple). Raise the correlation: the green cloud tightens a lot; the red one does not change. At correlation 0 both are equally wide.

"Regression adjustment changes what the experiment estimates."

In a randomized experiment, adjusting for a pre-treatment covariate keeps the same target (the ATE) and only reduces noise. (In small samples it can add a tiny bias; Lin's interacted version and robust SEs keep it honest.)

"Add this week's page views as a covariate; it predicts revenue really well."

This week's page views may be changed by the treatment (a post-treatment variable). Adjusting for it can remove part of the effect or add bias. Use only pre-treatment covariates.

$Y = \alpha + \tau T + \beta(X - \bar X) + \varepsilon$, $X$ pre-treatment; read $\hat\tau$.

Randomized: same ATE, variance × about $(1 - \rho^2)$. Observational: causal only with all confounders in $X$.

Trap: never adjust for post-treatment variables.

Quick check: pre-period and test-period revenue have correlation 0.5. By how much does adjustment cut the variance of the effect estimate, and what sample-size saving is that?

Variance × $(1 - 0.25) = 0.75$: a 25% cut. The same precision would need $1/0.75 \approx 1.33$ times as many users without adjustment.

CUPED: variance reduction with pre-experiment data core

CUPED (Controlled-experiment Using Pre-Experiment Data) is regression adjustment packaged as a simple recipe. For each user, take the metric during the experiment and subtract the part you could have predicted from the same user's data before the experiment. Heavy spenders stop looking "lucky", light spenders stop looking "unlucky". What is left is less noisy, and because the pre-period data cannot be affected by the treatment, the average difference between the arms does not change.

Three ways to say it:

  • Picture: weigh each person's luggage on the way back and subtract what it weighed on the way out; the difference shows what they bought, without the noise of how big their suitcase was.
  • Numbers: if pre and post spend have correlation 0.7, the variance of the estimate drops to $1 - 0.49 = 51\%$: CUPED with 10,000 users is as precise as no CUPED with about 19,600.
  • Slogan: same signal, less noise, nearly free.

Five users. Pre-period spend $X = 2, 4, 6, 8, 10$; experiment-period spend $Y = 3, 6, 5, 9, 12$ (dollars).

  1. Means: $\bar X = 6$, $\bar Y = 7$. Deviations: $x - \bar x = -4, -2, 0, 2, 4$; $y - \bar y = -4, -1, -2, 2, 5$.
  2. $\sum (x - \bar x)(y - \bar y) = 16 + 2 + 0 + 4 + 20 = 42$; $\sum (x - \bar x)^2 = 16 + 4 + 0 + 4 + 16 = 40$.
  3. $\hat\theta = \widehat{Cov}(Y, X) / \widehat{Var}(X) = 42/40 = 1.05$ (the $n - 1$ cancels).
  4. Adjusted values $Y' = Y - 1.05\,(X - 6)$: $3 + 4.2 = 7.2$; $6 + 2.1 = 8.1$; $5$; $9 - 2.1 = 6.9$; $12 - 4.2 = 7.8$. Their mean is still $7$.
  5. Sample variances: $Var(Y) = (16 + 1 + 4 + 4 + 25)/4 = 12.5$; $Var(Y') = (0.04 + 1.21 + 4 + 0.01 + 0.64)/4 = 5.9/4 = 1.475$.
  6. Ratio $1.475/12.5 = 0.118$, and indeed $\rho^2 = 42^2/(40 \times 50) = 0.882$, so $1 - \rho^2 = 0.118$.

Let $Y$ be the metric during the experiment and $X$ a covariate measured before the experiment (usually the same metric in a pre-period). For any constant $\theta$ define

$$Y^{cv} = Y - \theta\,(X - E[X]).$$

Derivation.

  1. Mean: $E[Y^{cv}] = E[Y] - \theta\,(E[X] - E[X]) = E[Y]$. Because $X$ is pre-treatment, $E[X]$ is the same in both arms, so the difference between arms keeps its expected value: still the ATE.
  2. Variance: $Var(Y^{cv}) = Var(Y) - 2\theta\,Cov(Y, X) + \theta^2 Var(X)$ (rules for $Var(aU + bV)$ from Chapter 4.15).
  3. Minimize over $\theta$: the derivative $-2Cov(Y, X) + 2\theta Var(X)$ is zero at $$\theta^* = \frac{Cov(Y, X)}{Var(X)} \quad(\text{the slope of the regression of } Y \text{ on } X).$$
  4. Plug in: $Var(Y^{cv}) = Var(Y) - \dfrac{Cov(Y, X)^2}{Var(X)} = Var(Y)\,(1 - \rho^2)$, with $\rho = \dfrac{Cov(Y, X)}{\sqrt{Var(Y)\,Var(X)}}$.

The CUPED estimate of the effect is $\hat\Delta^{cv} = \bar Y^{cv}_T - \bar Y^{cv}_C = (\bar Y_T - \bar Y_C) - \hat\theta\,(\bar X_T - \bar X_C)$, with one $\hat\theta$ estimated from both arms pooled. Its variance is about $(1 - \rho^2)$ times that of the plain difference: you need $1 - \rho^2$ times as many users for the same precision. It is essentially regression adjustment with one covariate (previous section).

Why do we need it?

Revenue, sessions and engagement metrics are dominated by stable differences between users. CUPED removes that predictable part, so experiments reach the same power with far fewer users or days: often a 20–50% variance cut on engagement metrics, depending on how predictable they are.

Where is it used?

Introduced at Microsoft (Deng, Xu, Kohavi and Walker, 2013) and now standard in large experimentation platforms; Booking.com has written publicly about using it, and commercial A/B tools offer it as an option. It applies to any metric with a useful pre-period: revenue, sessions, orders, time spent.

How is it used?

For each user, compute the metric in a pre-period of similar length (users with no history get $X$ = the mean, or an indicator). Estimate $\hat\theta$ on both arms pooled; form $Y - \hat\theta(X - \bar X)$; analyse it like the original metric. Report the variance reduction achieved.

Left: users of one arm, pre-period spend (x) vs spend during the test (y), with the line of slope $\theta$. Tick show CUPED values: each user moves to $y - \theta(x - \bar x)$ (orange): the cloud flattens and its spread shrinks. Right: the estimated lift from 400 repeated experiments, plain (blue) vs CUPED (orange); the truth is +1 dollar. Raise $\rho$: the orange pile narrows like $\sqrt{1 - \rho^2}$. At $\rho = 0$ CUPED does nothing (and does no harm).

"CUPED changes the effect being measured."

Because $X$ is measured before assignment, $E[\bar X_T - \bar X_C] = 0$, so the adjusted difference has the same expected value: the ATE. Only the variance changes.

"Use any covariate that correlates with the metric, even one measured during the test."

A covariate measured during the test may be affected by the treatment; subtracting it can remove part of the effect. CUPED needs pre-experiment data (or data provably unaffected by treatment).

"Estimate θ separately in each arm."

Use one $\hat\theta$ for both arms (pooled). Separate thetas mean different transformations per arm, which can bias the comparison.

"CUPED helps every metric."

The gain is $\rho^2$. New users with no history, rare events, and metrics with little week-to-week stability give small $\rho$ and little gain.

In an A/B framework like yours, the CUPED idea can live inside the model instead of being a preprocessing step: for a Normal or Student-t metric, write $y_i \sim \text{Normal}(\alpha + \tau\, t_i + \beta\,(x_i - \bar x), \sigma)$ with the pre-period metric $x_i$ as a centred covariate; the posterior variance of $\tau$ shrinks by roughly the same factor $1 - \rho^2$ (the posterior sd by its square root). For conversion metrics the conjugate Beta-Binomial has no slot for a covariate: CUPED-adjusted values are no longer 0/1, so a Beta-Binomial likelihood no longer fits them. You would need a Bernoulli (logistic) regression with the covariate, or accept the unadjusted model. Either way, the covariate must come from before assignment.

"CUPED is a bias correction."

CUPED is a variance reduction. The plain difference in means is already unbiased in a randomized test; CUPED keeps it unbiased and makes it less noisy.

Model answer: "CUPED subtracts $\theta(X - \bar X)$, where $X$ is pre-experiment data and $\theta = Cov(Y, X)/Var(X)$ minimizes the variance. The mean difference is unchanged because $X$ cannot be affected by the treatment, and the variance shrinks by $1 - \rho^2$. It is equivalent to regression adjustment with one covariate."

$Y^{cv} = Y - \theta(X - \bar X)$, $\theta^* = Cov(Y, X)/Var(X)$, $Var(Y^{cv}) = Var(Y)(1 - \rho^2)$.

$X$ pre-experiment ⇒ same expected difference (ATE); one pooled $\hat\theta$.

Users needed × $(1 - \rho^2)$. ρ = 0.7 → 51% of the users. Trap: post-treatment covariates.

Quick check: a metric has Var(Y) = 400 and its pre-period version has correlation 0.6 with it. What is the CUPED variance, and how many users does a 10,000-user CUPED test "act like"?

$400 \times (1 - 0.36) = 256$. It acts like a plain test with $10{,}000/0.64 \approx 15{,}600$ users.

Propensity scores: re-weighting to imitate a randomized experiment · familiarity level

When users choose the treatment, some kinds of users are over-represented among the treated. Desktop users turn the new feature on often, mobile users rarely. If you knew each user's chance of being treated (their propensity), you could fix the imbalance: give each rarely-treated user who was treated a big weight (they stand in for many similar users who were not), and each often-treated user a small weight. After weighting, the treated group has the same mix as the whole population, and so does the control group, as if a coin had decided.

Three ways to say it:

  • Picture: a survey with too few young people: count each young respondent several times, so the sample looks like the country again.
  • Numbers: mobile users turn the feature on 20% of the time: each treated mobile user counts $1/0.2 = 5$ times.
  • Slogan: weight by one over the chance of what happened.

60% of users are on mobile (conversion 5% without the feature), 40% on desktop (15%). The feature adds +2 pts for everyone. Mobile users turn it on with probability 0.2, desktop users with 0.7.

  1. Treated mix: mobile $0.6 \times 0.2 = 0.12$, desktop $0.4 \times 0.7 = 0.28$. Treated conversion: $(0.12 \times 7 + 0.28 \times 17)/0.40 = (0.84 + 4.76)/0.40 = 14.0\%$.
  2. Control mix: mobile $0.6 \times 0.8 = 0.48$, desktop $0.4 \times 0.3 = 0.12$. Control conversion: $(0.48 \times 5 + 0.12 \times 15)/0.60 = (2.4 + 1.8)/0.60 = 7.0\%$. Naive gap: 7 pts.
  3. Weights: treated mobile $1/0.2 = 5$, treated desktop $1/0.7 \approx 1.43$; control mobile $1/0.8 = 1.25$, control desktop $1/0.3 \approx 3.33$.
  4. Weighted treated mix: mobile $0.12 \times 5 = 0.6$, desktop $0.28 \times 1.43 = 0.4$: back to 60/40. Weighted treated conversion: $0.6 \times 7 + 0.4 \times 17 = 11.0\%$.
  5. Weighted control mix: $0.48 \times 1.25 = 0.6$, $0.12 \times 3.33 = 0.4$. Conversion $0.6 \times 5 + 0.4 \times 15 = 9.0\%$. IPW estimate: $11 - 9 = 2$ pts, the truth.
  • Propensity score: $e(x) = P(T = 1 \mid X = x)$, the probability of being treated given the covariates. In observational data it is unknown and is estimated, usually with a logistic regression of $T$ on $X$ (Chapter 5.14).
  • Balancing property: among units with the same $e(x)$, the covariates $X$ have the same distribution in treated and control units. So one number can replace many covariates for matching or stratifying.
  • Inverse probability weighting (IPW): $\widehat{\text{ATE}} = \frac{1}{n}\sum_i \frac{T_i Y_i}{e(X_i)} - \frac{1}{n}\sum_i \frac{(1 - T_i) Y_i}{1 - e(X_i)}$ (or the normalized version that divides by the sum of weights, as in the example).
  • Assumptions: (1) no unmeasured confounding given $X$; (2) overlap (positivity): $0 \lt e(x) \lt 1$ for every $x$, i.e. every kind of user has some chance of each arm. Near 0 or 1, weights explode and estimates become very noisy; people trim or cap weights.
  • Other uses: propensity-score matching (pair each treated unit with a control of similar $e(x)$) and stratification on $e(x)$. In a randomized experiment $e(x) = 0.5$ for everyone, and IPW reduces to the difference in means.
Why do we need it?

When you cannot randomize (opt-in features, past launches, marketing that targeted some customers), propensity methods turn a lopsided comparison into a balanced one, as long as all the reasons for being treated are measured.

Where is it used?

Observational studies in medicine and economics, opt-in product features, evaluating campaigns sent to targeted customers, off-policy evaluation of recommender systems (logging propensities of the old policy), and survey weighting.

How is it used?

Fit a model for $P(T = 1 \mid X)$ on pre-treatment covariates; check overlap (histograms of $e(x)$ by group); compute weights or matches; check that covariates are balanced after weighting; estimate the effect; state "no unmeasured confounding" as the key assumption.

Left: each segment's users, split into treated (orange) and untreated (blue), with the IPW weight of each part. Right: naive (red), IPW (green) and true (purple) effects. Move the uptake sliders apart: the naive estimate drifts, IPW stays on the truth. Push an uptake to 0.05 or 0.95: IPW is still right here (no sampling noise), but the weights reach 20, so with real data a handful of users would decide the answer: an overlap problem.

"Propensity scores fix confounding, so observational studies are as good as experiments."

They only balance the covariates you put in the propensity model. An unmeasured confounder (motivation, intent) stays unbalanced, exactly as with regression adjustment.

"A propensity model with very high accuracy is a good propensity model."

A propensity model that predicts treatment almost perfectly means poor overlap: some kinds of users are (almost) never in one arm, and weights explode. The goal is balance, not prediction accuracy.

$e(x) = P(T = 1 \mid X = x)$; IPW: treated weight $1/e$, control weight $1/(1 - e)$.

Needs: no unmeasured confounding + overlap $0 \lt e(x) \lt 1$.

Check overlap and covariate balance after weighting; extreme weights = trouble.

Quick check: a treated user had propensity 0.1. How many "users like them" do they represent in IPW, and why?

$1/0.1 = 10$. Only 1 in 10 users like them was treated, so this one treated user stands in for all 10 to recreate the population's mix in the treated group.

Difference-in-differences: use a comparison group's change as the counterfactual · familiarity level

A feature launches in city A only. Orders in A went up after launch. But orders also went up in city B, where nothing launched: maybe it was the season, or a holiday. The fair question is: did A go up more than B?

Difference-in-differences (DiD) takes A's change (after minus before) and subtracts B's change. B's change stands in for what would have happened to A without the launch. This works only if A and B would have moved in parallel without the launch: the parallel trends assumption.

Three ways to say it:

  • Picture: two lines that ran side by side; after the launch one of them jumps up; the jump beyond the other line's movement is the effect.
  • Numbers: A went 100 → 120 (+20), B went 80 → 90 (+10): effect $20 - 10 = +10$ orders a week.
  • Slogan: subtract the change that would have happened anyway.

Average weekly orders before and after a launch in city A (city B never got the feature).

  1. City A: before 100, after 120. Change $+20$.
  2. City B: before 80, after 90. Change $+10$ (season, marketing, weather: everything except the feature).
  3. DiD: $(120 - 100) - (90 - 80) = 20 - 10 = +10$ orders a week.
  4. Implied counterfactual for A: "A would have grown like B": $100 + 10 = 110$. Observed 120, so the effect is $120 - 110 = 10$.
  5. Note: the levels differ (A is bigger); DiD only needs the changes to be comparable, not the levels.

With groups treated (T) and comparison (C), and periods before (pre) and after (post):

$$\widehat{\text{DiD}} = (\bar Y_{T,\text{post}} - \bar Y_{T,\text{pre}}) - (\bar Y_{C,\text{post}} - \bar Y_{C,\text{pre}}).$$
  • Regression form: $Y_{gt} = \beta_0 + \beta_1\,\text{treated}_g + \beta_2\,\text{post}_t + \beta_3\,(\text{treated}_g \times \text{post}_t) + \varepsilon_{gt}$; $\hat\beta_3$ is the DiD estimate.
  • Parallel trends: without treatment, the treated group's average would have changed by the same amount as the comparison group's. It concerns an unobservable counterfactual, so it cannot be proven; checking that the pre-launch trends were parallel makes it more believable.
  • DiD estimates the effect on the treated group (ATT). It also assumes no other change hit only the treated group at the same time, and no spillover between the groups (SUTVA).
Why do we need it?

Many changes cannot be A/B tested: a city-wide launch, a price change for one country, a new law, a marketing campaign in one region. DiD gives a credible estimate from before/after data plus a comparison group, under one clear assumption.

Where is it used?

Geo experiments and regional launches, policy evaluation in economics (the classic minimum-wage study by Card and Krueger), marketing measurement, and staggered feature rollouts across markets.

How is it used?

Pick comparison units that tracked the treated ones before the launch; plot both series and check pre-trends; estimate $\beta_3$ with smf.ols("y ~ treated * post", data), with standard errors clustered by unit; try placebo launch dates as a sanity check.

Blue: comparison city; orange: treated city; the launch is after week 5 (grey line). The dashed orange line is DiD's counterfactual: the treated city moving exactly like the comparison city. With trend gap = 0, DiD recovers the true effect. Make the treated city grow faster even before the launch (trend gap > 0): the pre-launch lines are no longer parallel, and DiD wrongly counts the extra growth as effect. The green dashed line shows the treated city's real no-launch path.

"DiD needs the two groups to have the same level."

It needs parallel changes (trends), not equal levels. A bigger city is fine if it would have grown by the same amount.

"The pre-trends look parallel, so parallel trends is proven."

Parallel pre-trends make the assumption plausible; they cannot prove what would have happened after the launch. Something else may have changed only in the treated city at launch time.

Your forecasting model offers another way to build a counterfactual for a launch that was not A/B tested: fit the model on data up to the launch (with regressors that are not affected by the launch), forecast the post-launch period, and compare the actuals with the forecast distribution. That is the idea behind Bayesian structural time-series tools such as Google's CausalImpact. It relies on the same kind of assumption as DiD: nothing except the launch changed the series' behaviour, and the model's trend, seasonality and holiday terms would have continued.

DiD $= (\bar Y_{T,post} - \bar Y_{T,pre}) - (\bar Y_{C,post} - \bar Y_{C,pre}) = \hat\beta_3$ in $y \sim \text{treated} \times \text{post}$.

Key assumption: parallel trends (no-treatment changes equal). Check pre-trends; use placebo dates.

Estimates the effect on the treated; cluster SEs by unit.

Quick check: store A's sales went 50 → 65 after a redesign; comparison store B went 40 → 48. DiD estimate?

$(65 - 50) - (48 - 40) = 15 - 8 = +7$. Valid if A would otherwise have changed by the same +8 as B.

Instrumental variables: a random nudge reveals the effect · familiarity level

You cannot force users to use a feature, and users who choose it are special (more engaged). But you can randomly send half of them an email that encourages it. The email is random, so it is not tied to engagement. It changes usage (some users try the feature because of it), and, if it has no other effect on buying, any difference in buying between emailed and non-emailed users must have come through the extra usage.

So: effect of the email on buying ÷ effect of the email on usage = effect of usage on buying, for the users the email persuaded. The random nudge is the instrument.

Three ways to say it:

  • Picture: you cannot push the swing yourself, but a random gust of wind pushes it; measure how far the wind moved the swing and how far the swing moved the bell.
  • Numbers: the email raises usage by 30 pts and conversion by 1.5 pts: $1.5/30 = 0.05$, so using the feature adds 5 pts.
  • Slogan: randomize the nudge when you cannot randomize the action.

Half the users are randomly emailed ($Z = 1$). Feature usage: 20% without the email, 50% with it. Conversion: 10.0% without the email, 11.5% with it. Comparing users vs non-users of the feature directly gives a gap of about 8.6 pts (engaged users use it more), which is confounded.

  1. Effect of the email on conversion (intention-to-treat): $11.5 - 10.0 = 1.5$ pts.
  2. Effect of the email on usage (first stage): $50 - 20 = 30$ pts.
  3. Wald (IV) estimate: $1.5 / 30 = 0.05$, i.e. $+5$ pts of conversion for using the feature.
  4. Who is this about? The users who used the feature because of the email (the 30% "compliers"), not those who would use it anyway.
  5. Same idea as the contamination fix in Chapter 5.11: measured effect ÷ difference in exposure.

An instrument $Z$ for the effect of $T$ on $Y$ must satisfy:

  1. Relevance: $Z$ changes $T$ (testable: the first stage is clearly non-zero).
  2. Independence: $Z$ is as good as random (true by design if $Z$ is randomized).
  3. Exclusion: $Z$ affects $Y$ only through $T$ (not testable; the email must not itself make people buy).
$$\hat\tau_{IV} = \frac{\bar Y_{Z=1} - \bar Y_{Z=0}}{\bar T_{Z=1} - \bar T_{Z=0}} \qquad \text{(the Wald estimator).}$$
  • With effects that vary between users, and monotonicity (the nudge never makes anyone less likely to use the feature), it estimates the local average treatment effect (LATE): the average effect among compliers.
  • Weak instruments (tiny first stage) make the ratio very noisy and can bias it; dividing by a number near zero amplifies everything.
  • In regression language this is two-stage least squares (2SLS).
Why do we need it?

Often you can randomize an offer, an invitation or a default, but not the behaviour itself. IV turns the randomized nudge into an estimate of the behaviour's effect, without needing to measure the confounders.

Where is it used?

Encouragement designs in product (emails, banners, defaults), experiments with non-compliance (the treatment-on-the-treated effect), economics (distance to college, lottery-based admissions), and medicine (randomized invitations to screening).

How is it used?

Randomize the nudge; check the first stage is strong; argue for exclusion; compute the Wald ratio (or 2SLS with covariates, e.g. linearmodels.IV2SLS); report it as the effect for compliers, alongside the ITT effect of the nudge itself.

Z: random email(instrument) T: uses feature Y: converts U: engagement relevance effect τ no direct path (exclusion) Z is randomized: no arrow from U to Z
The IV picture. Engagement U confounds T and Y, so comparing users with non-users is biased. The randomized email Z moves T, is unrelated to U, and (by assumption) reaches Y only through T. The ratio of Z's effect on Y to Z's effect on T isolates τ.

Each run simulates 20,000 users: a random half get the email; engaged users often use the feature anyway; the email persuades some others. Histograms over 200 runs: red = naive "users vs non-users", green = IV (Wald) estimate; purple = true effect. Raise engagement's effect: the red pile moves away from the truth, the green one does not. Lower persuaded by the email toward 5%: the instrument gets weak and the green pile spreads out enormously.

"Any variable correlated with the treatment can be an instrument."

It must also be as good as random and affect the outcome only through the treatment. An email that itself contains a discount code breaks exclusion: it changes buying directly.

"The IV estimate is the effect for everyone."

With varying effects it is the effect for compliers (the users the nudge moved). Users who always or never use the feature may have different effects.

IV (Wald): $\hat\tau = \dfrac{\bar Y_{Z=1} - \bar Y_{Z=0}}{\bar T_{Z=1} - \bar T_{Z=0}}$ = ITT on $Y$ ÷ first stage.

Needs relevance, independence (randomized Z), exclusion (Z → Y only via T); + monotonicity → LATE for compliers.

Trap: weak instruments give huge noise; exclusion is untestable.

Quick check: a random "try premium free" banner raises premium trials from 5% to 15% and 30-day retention from 40.0% to 41.5%. IV estimate of the trial's effect on retention?

$(41.5 - 40.0)/(15 - 5) = 1.5/10 = 0.15$: about +15 pts of retention for the users the banner persuaded to start a trial, assuming the banner affects retention only through the trial.

Recap, cheat sheet and practice

  • Correlation ≠ causation: associations also come from reverse causation, common causes (confounding), selection and chance. Draw the DAG.
  • Potential outcomes $Y_i(1), Y_i(0)$: the effect $\tau_i = Y_i(1) - Y_i(0)$ is never observed (fundamental problem); the unseen outcome is the counterfactual. The notation assumes SUTVA (Chapter 5.11).
  • ATE $= E[Y(1)] - E[Y(0)]$ can be estimated because it is a difference of averages; CATE is the ATE inside a subgroup. A positive ATE does not mean nobody is harmed.
  • Randomization makes $T$ independent of the potential outcomes, so the difference in means is unbiased for the ATE, with balance on measured and unmeasured covariates on average.
  • Confounding opens a backdoor path $T \leftarrow U \to Y$; adjust for pre-treatment confounders (if all are measured) or randomize. Colliders and post-treatment filters create bias; do not condition on them.
  • Regression adjustment and CUPED use pre-experiment data to cut variance by about $1 - \rho^2$ without changing the target: $\theta = Cov(Y, X)/Var(X)$.
  • Without randomization: propensity scores (weight by $1/e(x)$; need all confounders + overlap), difference-in-differences (need parallel trends), instrumental variables (need relevance, independence, exclusion; estimate the effect for compliers).
Can you randomize the treatment itself? yes Randomized A/B test difference in means or posterior, + CUPED no Can you randomize a nudge (email, invite)? yes Instrumental variable encouragement design; effect for compliers no Comparison group with parallel pre-trends? yes Difference-in-differences (change treated) − (change comparison) no All confounders measured, with overlap? yes Propensity scores or adjustment assumes no unmeasured confounding no No credible causal estimate: report the association honestly, and design a way to randomize.
A rough guide to choosing a method. Each step down needs a stronger, less checkable assumption: randomization needs almost none; the last box needs "every confounder is measured".

Cheat sheet

IdeaFormulaKey assumption / trap
Potential outcomes$Y_i = T_iY_i(1) + (1-T_i)Y_i(0)$, $\tau_i = Y_i(1) - Y_i(0)$SUTVA; $\tau_i$ never observed
ATE / CATE / ATT$E[Y(1)-Y(0)]$; $E[\cdot \mid X=x]$; $E[\cdot \mid T=1]$averages, not individuals
Randomization$T \perp (Y(0), Y(1))$ ⇒ $E[\bar Y_T - \bar Y_C] = $ ATEdo not filter after assignment
Selection biasnaive = ATT + $E[Y(0) \mid T=1] - E[Y(0) \mid T=0]$why observational gaps mislead
Adjustment$\sum_x P(x)\,(E[Y \mid 1, x] - E[Y \mid 0, x])$no unmeasured confounding; no mediators/colliders
Regression adjustment$Y = \alpha + \tau T + \beta(X - \bar X) + \varepsilon$pre-treatment $X$; variance × $(1 - \rho^2)$
CUPED$Y - \theta(X - \bar X)$, $\theta = Cov(Y,X)/Var(X)$, $Var \times (1-\rho^2)$pre-experiment $X$; one pooled $\theta$
IPWweights $1/e(x)$ (treated), $1/(1-e(x))$ (control)all confounders in $X$; overlap
Difference-in-differences$(\Delta \bar Y_T) - (\Delta \bar Y_C)$ $= \hat\beta_3$ in treated × postparallel trends
Instrumental variable$\dfrac{\bar Y_{Z=1} - \bar Y_{Z=0}}{\bar T_{Z=1} - \bar T_{Z=0}}$relevance, independence, exclusion; LATE
Code it · Python

import itertools
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf

rng = np.random.default_rng(1)

# 1) Potential outcomes: we can only see one of Y(0), Y(1) per user
y0 = np.array([5, 3, 8, 4, 6, 2]); y1 = np.array([7, 3, 9, 7, 5, 3])
print("true ATE:", (y1 - y0).mean())
# true ATE: 1.0
ests = []
for treated in itertools.combinations(range(6), 3):          # all 20 ways to treat 3 of 6 users
    t = np.zeros(6, bool); t[list(treated)] = True
    ests.append(y1[t].mean() - y0[~t].mean())                # what we would SEE under this assignment
print("estimates range:", round(min(ests), 2), "to", round(max(ests), 2), "| average over all 20:", round(np.mean(ests), 3))
# estimates range: -2.0 to 4.0 | average over all 20: 1.0     (unbiased, but any single split can be far off)

# 2) Confounding: engagement drives both feature use and conversion
n = 200_000
high = rng.random(n) < 0.5                                   # engaged users
uses = rng.random(n) < np.where(high, 0.8, 0.2)              # engaged users use the feature more
conv = rng.random(n) < np.where(high, 0.20, 0.05) + 0.02 * uses   # true effect of the feature: +2 pts
naive = conv[uses].mean() - conv[~uses].mean()
strat = np.mean([conv[uses & (high == h)].mean() - conv[~uses & (high == h)].mean() for h in (True, False)])
print(f"naive: {naive:.3f}   stratified by engagement: {strat:.3f}   truth: 0.020")
# naive: 0.111   stratified by engagement: 0.022   truth: 0.020

# 3) CUPED on a simulated experiment (pre-period spend x, experiment-period spend y)
def experiment(n=2000, rho=0.7, effect=1.0):
    x = rng.normal(50, 20, 2 * n)                            # pre-period metric, same distribution in both arms
    t = np.repeat([0, 1], n)
    y = 50 + rho * (x - 50) + np.sqrt(1 - rho**2) * rng.normal(0, 20, 2 * n) + effect * t
    theta = np.cov(y, x)[0, 1] / np.var(x, ddof=1)           # theta = Cov(Y, X) / Var(X), pooled over both arms
    y_cuped = y - theta * (x - x.mean())
    plain = y[t == 1].mean() - y[t == 0].mean()
    cuped = y_cuped[t == 1].mean() - y_cuped[t == 0].mean()
    return plain, cuped, theta, x, y, t
res = np.array([experiment()[:2] for _ in range(2000)])
print("sd of estimate  plain: %.3f  cuped: %.3f  variance ratio: %.3f (theory 1 - 0.7^2 = 0.51)"
      % (res[:, 0].std(), res[:, 1].std(), res[:, 1].var() / res[:, 0].var()))
# sd of estimate  plain: 0.649  cuped: 0.457  variance ratio: 0.496 (theory 1 - 0.7^2 = 0.51)

# 4) CUPED = regression adjustment: the coefficient on t in y ~ t + x
plain, cuped, theta, x, y, t = experiment()
fit = smf.ols("y ~ t + x", data=pd.DataFrame({"y": y, "t": t, "x": x})).fit()
print(f"CUPED estimate {cuped:.3f} vs OLS coefficient on t {fit.params['t']:.3f} (se {fit.bse['t']:.3f})")
# CUPED estimate 1.836 vs OLS coefficient on t 1.837 (se 0.451)   (true effect 1.0: one noisy experiment;
#                                                                    the point is that the two methods agree)

# 5) Difference-in-differences: average weekly orders before/after a launch in city A only
before_A, after_A, before_B, after_B = 100, 120, 80, 90
print("DiD estimate:", (after_A - before_A) - (after_B - before_B))
# DiD estimate: 10
d = pd.DataFrame({"orders": [before_A, after_A, before_B, after_B],
                  "treated": [1, 1, 0, 0], "post": [0, 1, 0, 1]})
print("same number from regression:", smf.ols("orders ~ treated * post", data=d).fit().params["treated:post"])
# same number from regression: 10.0

# 6) Instrumental variable (encouragement design): random email Z, feature use T, conversion Y
n = 1_000_000
z = rng.random(n) < 0.5                                      # randomized email
u = rng.random(n) < 0.5                                      # hidden engagement (confounder)
always = u & (rng.random(n) < 0.4)                           # some engaged users use the feature anyway
use = always | (z & (rng.random(n) < 0.375))                 # the email persuades some others
y = rng.random(n) < 0.05 + 0.10 * u + 0.05 * use             # true effect of using the feature: +5 pts
naive = y[use].mean() - y[~use].mean()
wald = (y[z].mean() - y[~z].mean()) / (use[z].mean() - use[~z].mean())
print(f"naive users-vs-non-users: {naive:.3f}   IV (Wald): {wald:.3f}   truth: 0.050")
# naive users-vs-non-users: 0.087   IV (Wald): 0.052   truth: 0.050
Test yourself

1. Why can we never directly measure one user's treatment effect?

A user is either treated or not at a given time, so we see $Y_i(1)$ or $Y_i(0)$, never both: the fundamental problem of causal inference. Averages over groups get around it; individual effects stay hidden.

2. What does randomizing the treatment guarantee?

Randomization makes treatment independent of the potential outcomes, so $E[\bar Y_T - \bar Y_C] = E[Y(1)] - E[Y(0)]$. Single splits still differ by chance (that is what the SE measures), and it says nothing about interference or heterogeneity.

3. A pre-period covariate has correlation ρ = 0.8 with the metric. CUPED multiplies the variance of the effect estimate by about…

$1 - \rho^2 = 1 - 0.64 = 0.36$: the same precision with 36% of the users (equivalently, the SE shrinks by the factor 0.6).

4. Which covariate is safe to use for CUPED or regression adjustment in an A/B test?

Only pre-treatment covariates are safe: the treatment cannot have changed them, so the adjusted difference keeps the same expected value. The other three can be affected by the treatment.

5. Difference-in-differences relies mainly on which assumption?

Parallel trends: the comparison group's change is a stand-in for the treated group's missing counterfactual change. Levels may differ.

6. An encouragement email is randomized to estimate the effect of using a feature on purchases. Which situation breaks the IV exclusion restriction?

With a discount, the email changes purchases directly, not only through feature use, so the Wald ratio mixes the two. A small first stage (10%) makes the instrument weak but not invalid; confounding of usage is exactly what IV handles; random sending is what makes it an instrument.

Practice problems

A. Four users: $Y(0) = 10, 12, 8, 14$ and $Y(1) = 13, 12, 11, 16$. Find the ATE, the estimate if users 2 and 3 are treated, and the average estimate over all 6 ways to treat 2 of 4.
  1. Effects: $3, 0, 3, 2$; ATE $= 8/4 = 2$.
  2. Users 2, 3 treated: treated mean $(12 + 11)/2 = 11.5$; control (users 1, 4) mean $(10 + 14)/2 = 12$; estimate $-0.5$. An unlucky split.
  3. All six splits: {1,2} 1.5; {1,3} −1; {1,4} 4.5; {2,3} −0.5; {2,4} 5; {3,4} 2.5. Sum 12, average $12/6 = 2$: unbiased.
B. Show that $\theta = Cov(Y, X)/Var(X)$ minimizes $Var(Y - \theta(X - E[X]))$ and that the minimum is $Var(Y)(1 - \rho^2)$.
  1. $Var(Y - \theta X') = Var(Y) - 2\theta Cov(Y, X) + \theta^2 Var(X)$ (the constant $E[X]$ does not change variances).
  2. This is a parabola in $\theta$ opening upward; derivative $-2Cov(Y,X) + 2\theta Var(X) = 0$ gives $\theta^* = Cov(Y,X)/Var(X)$.
  3. Plug in: $Var(Y) - 2\frac{Cov^2}{Var(X)} + \frac{Cov^2}{Var(X)} = Var(Y) - \frac{Cov(Y,X)^2}{Var(X)}$.
  4. Factor: $Var(Y)\Big(1 - \frac{Cov(Y,X)^2}{Var(Y)Var(X)}\Big) = Var(Y)(1 - \rho^2)$.
C. 40% of users are engaged (conversion 25%), 60% are not (5%). A feature adds 3 pts. 90% of engaged and 10% of other users use it. Naive gap? Confounding bias?
  1. Users: engaged $0.36$, others $0.06$, so $0.36/0.42 \approx 85.7\%$ of users are engaged. Their conversion: $0.857 \times 28 + 0.143 \times 8 \approx 25.14\%$.
  2. Non-users: engaged $0.04$, others $0.54$, so $0.04/0.58 \approx 6.9\%$ engaged. Conversion: $0.069 \times 25 + 0.931 \times 5 \approx 6.38\%$.
  3. Naive gap $\approx 18.76$ pts vs a true 3 pts. Bias $= (0.857 - 0.069) \times (25 - 5) \approx 15.76$ pts: the difference in engagement mix times engagement's effect.
D. (Interview) "What is CUPED, why does it work, and when does it not help?"

"CUPED reduces the variance of an A/B estimate using pre-experiment data. For each user I compute $Y - \theta(X - \bar X)$, where $X$ is the same metric before the experiment and $\theta = Cov(Y,X)/Var(X)$, estimated once on both arms pooled. Because $X$ was measured before assignment, it cannot differ between arms in expectation, so the mean difference, and hence the ATE estimate, is unchanged. The variance drops by the factor $1 - \rho^2$, which is equivalent to running with $1/(1-\rho^2)$ times as many users. It is the same as regression adjustment with one covariate. It helps little when $\rho$ is small: new users without history, rare binary events, unstable metrics. And it must never use data that the treatment could have changed."

E. Legal rules forbid forcing users through a new onboarding, but you may invite a random half to it. How do you estimate the onboarding's effect on 30-day retention?

An encouragement design: the random invitation is the instrument $Z$, taking the onboarding is $T$, retention is $Y$. Estimate the ITT effect of the invitation on retention and the first stage (how much the invitation raises onboarding uptake); the Wald ratio ITT / first stage estimates the onboarding's effect for compliers (users who take it because they were invited). Check relevance (a clearly non-zero first stage), and argue exclusion: the invitation itself must not change retention except through onboarding (for example it should not come with a gift).

F. Pick a method: (1) a new ranking model, randomizable by user; (2) a pricing change already rolled out in Spain only; (3) an opt-in feature with rich pre-adoption data; (4) a noisy revenue metric in a running A/B test.
  1. Randomized A/B test: difference in means (or the posterior of the difference), unbiased by design.
  2. Difference-in-differences with comparable countries as the comparison group (check pre-trends), or a forecast-based counterfactual.
  3. Propensity scores or regression adjustment on pre-adoption covariates, stating "no unmeasured confounding" (motivation is a worry), or better, randomize an encouragement and use IV.
  4. CUPED / regression adjustment with last month's revenue to cut the variance.
Chapter 5.13 · Syllabus Module 18

Linear regression

Linear regression is the workhorse of statistics: explain a number $y$ as a weighted sum of other numbers, plus noise. It is also the skeleton of your forecasting model. Trend, changepoints, Fourier seasonality, holidays and external regressors are all just columns of one big table $X$, and the model learns one weight per column. Learn regression properly once and both projects become easier to explain.

  • Read the model $y = X\beta + \epsilon$: what a row, a column, a coefficient and the noise term are, and why "linear" means linear in β, not "straight lines only"
  • Interpret an intercept and a coefficient in the right units, including "holding the other columns fixed"
  • Compute residuals and the least-squares fit, derive the normal equations $X^\top X\hat\beta = X^\top y$, and see them as a projection
  • Read a coefficient table: standard errors, t-values, p-values, confidence intervals and $R^2$
  • Know the assumptions (linearity, independence, constant variance, Normal errors) and exactly which one matters for what
  • Diagnose problems with residual plots (residual vs fitted, Q-Q, residual vs time), handle heteroscedasticity and multicollinearity (VIF)
  • See your forecasting model's trend, changepoints, Fourier terms, holidays and regressors as columns of one design matrix

What we need from earlier chapters: mean, variance and standard deviation (Chapter 4.5); covariance and correlation (Chapter 4.15); Q-Q plots (Chapter 4.17); standard errors, t-tests and confidence intervals (Chapters 5.5–5.8); maximum likelihood and "squared error = Gaussian likelihood" (Chapter 5.2); ridge and lasso as priors (Chapter 5.3). From the Linear Algebra guide: projections (Chapter 1.9) and least squares (Chapter 1.10). Notation: bold-free capital $X$ is the data table (a matrix), $y$ the column of outcomes, $\beta$ ("beta") the column of unknown weights, $\hat\beta$ ("beta hat") our estimate of them, $\epsilon$ ("epsilon") the noise. $N(0,\sigma^2)$ is written with the variance; NumPyro's dist.Normal(loc, scale) takes the standard deviation $\sigma$.

The model: $y = X\beta + \epsilon$ core

A café sells more cold drinks on hot days. Not exactly the same number every hot day, but on average more. If you drew the days on a chart (temperature across, drinks up), the dots would form a cloud that leans upward.

Linear regression says: each day's number = a predictable part + a surprise. The predictable part is a straight-line rule ("18 drinks, plus 1.6 more for every degree"). The surprise is everything the rule does not know about: a football match, a broken freezer, pure chance. We never see the rule directly; we see only the dots, and we try to recover the rule from them.

Three ways to say it:

  • Picture: a tilted line through a cloud of dots; each dot sits a little above or below the line.
  • Numbers: at 25 °C the rule predicts $18 + 1.6\times25 = 58$ drinks; a real day might bring 70 (surprise $+12$) or 50 (surprise $-8$).
  • Slogan: data = signal + noise, and the signal is a weighted sum of the inputs.

Five summer days. Temperature $x$ (°C) and cold drinks sold $y$:

day $i$12345
$x_i$ (°C)1015202530
$y_i$ (drinks)3050407060

Suppose the true rule were $y_i = 18 + 1.6\,x_i + \epsilon_i$ (we will see in a moment that these are exactly the best-fitting numbers).

  1. Day 1: predictable part $18 + 1.6\times10 = 34$. Observed 30, so the surprise is $30 - 34 = -4$.
  2. Day 2: $18 + 1.6\times15 = 42$; surprise $50 - 42 = +8$.
  3. Day 3: $18 + 32 = 50$; surprise $40 - 50 = -10$. Day 4: $18 + 40 = 58$; surprise $+12$. Day 5: $18 + 48 = 66$; surprise $-6$.
  4. Stack all five equations on top of each other and you get one matrix equation: $$\underbrace{\begin{bmatrix} 30 \\ 50 \\ 40 \\ 70 \\ 60 \end{bmatrix}}_{y} = \underbrace{\begin{bmatrix} 1 & 10 \\ 1 & 15 \\ 1 & 20 \\ 1 & 25 \\ 1 & 30 \end{bmatrix}}_{X} \underbrace{\begin{bmatrix} 18 \\ 1.6 \end{bmatrix}}_{\beta} + \underbrace{\begin{bmatrix} -4 \\ 8 \\ -10 \\ 12 \\ -6 \end{bmatrix}}_{\epsilon}.$$
  5. Check row 4: $1\times18 + 25\times1.6 + 12 = 18 + 40 + 12 = 70$ ✓. The column of 1s is how the intercept 18 gets added to every row.

The linear regression model for $n$ observations and $p$ columns is

$$y = X\beta + \epsilon, \qquad\text{one row at a time: } y_i = \beta_0 + \beta_1 x_{i1} + \dots + \beta_{p-1} x_{i,p-1} + \epsilon_i .$$
  • $y$ ($n\times1$): the response (also called target, outcome, dependent variable): the thing we want to explain.
  • $X$ ($n\times p$): the design matrix. One row per observation (a day, a user); one column per input (also called feature, predictor, covariate, regressor). The first column is usually all 1s, for the intercept.
  • $\beta$ ($p\times1$): the unknown coefficients (weights), one per column. $\beta_0$ is the intercept.
  • $\epsilon$ ($n\times1$): the errors (noise): what the columns do not explain. The basic assumptions are $E[\epsilon_i\mid X] = 0$ (the surprises average out to zero at every $x$) and, for the usual standard errors, independent errors with the same variance $\sigma^2$. Often we also assume $\epsilon_i \sim N(0, \sigma^2)$.

"Linear" means linear in β. Each coefficient multiplies a column and the results are added. The columns themselves can be anything you compute from the data: $x^2$, $\log x$, $\sin(2\pi t/7)$, a 0/1 holiday flag. A model like $y = \beta_0 + \beta_1 x + \beta_2 x^2 + \epsilon$ is still linear regression; $y = \beta_0 e^{\beta_1 x} + \epsilon$ is not.

Why do we need it?

We want to know how an outcome moves with its inputs, predict it for new inputs, and say how sure we are. A weighted sum plus noise is the simplest model that does all three, and its maths can be solved exactly.

Where is it used?

Demand and sales models, price-elasticity estimates, A/B analysis with covariates (regression adjustment, CUPED), the mean part of your forecasting model ($g(t) + s(t) + h(t) + X_t\beta$), feature baselines before gradient boosting, and the inner step of every GLM fit (Chapter 5.14).

How is it used?

Build the design matrix (one row per observation, one column per feature, plus a column of 1s), call statsmodels.api.OLS(y, X).fit() or np.linalg.lstsq(X, y), read the coefficients with their standard errors, then check the residuals before trusting anything.

y = X β + ε : the shapes y n × 1 = 1s ← one row = one day one column = one feature X (design matrix) n × p β p × 1 + ε n × 1 one weight per column (β), one surprise per row (ε)
The shapes of the four pieces. Each row of $X$ times the column $\beta$ gives that row's predictable part; adding that row's surprise $\epsilon_i$ gives the observed $y_i$.

You are the "true process". Set the intercept $\beta_0$, the slope $\beta_1$ and the noise level $\sigma$; the green line is the true rule, the blue dots are the days it produces, and the red sticks are the surprises $\epsilon_i$. The orange line is what least squares recovers from the dots alone. Set $\sigma = 0$: the dots sit exactly on the line and the orange line finds the truth perfectly. Raise $\sigma$ and press New sample a few times: the orange line wobbles around the green one. That wobble is what standard errors will measure.

"Linear regression can only fit straight lines."

It is linear in the coefficients. Add a column $x^2$ and it fits a parabola; add $\sin$ and $\cos$ columns and it fits a repeating seasonal shape. Your forecasting model's mean is exactly this kind of "curvy but linear" regression (below).

"Regression assumes $y$ is Normally distributed."

The assumptions are about the errors around the line, not about the overall histogram of $y$. If $x$ is spread out, $y$ can look skewed or bimodal while the errors are perfectly Normal.

"The error $\epsilon_i$ is the distance from the dot to my fitted line."

That distance is the residual $e_i$, measured from the estimated line. The error $\epsilon_i$ is measured from the true line, which we never see (next sections).

"It's called linear regression because the relationship between x and y is a straight line."

"It's linear because the mean of $y$ is a linear combination of the coefficients: $E[y\mid X] = X\beta$. The columns can be non-linear transformations of the raw inputs."

Model answer: "Linear refers to the parameters. Polynomial terms, splines, Fourier terms and indicator variables all keep the model linear in β, so ordinary least squares still applies and has a closed-form solution."

$y = X\beta + \epsilon$: rows = observations, columns = features (first column 1s), $\beta$ = one weight per column, $\epsilon$ = noise with $E[\epsilon\mid X] = 0$.

"Linear" = linear in β; columns may be $x^2$, $\sin$, indicators.

Trap: assumptions are about the errors, not about the histogram of $y$.

Quick check: is $y = \beta_0 + \beta_1 \log x + \beta_2\,\mathbb{1}[\text{weekend}] + \epsilon$ a linear regression? What are its columns?

Yes. It is linear in $\beta_0, \beta_1, \beta_2$. The design matrix has three columns: all 1s, $\log x_i$, and a 0/1 weekend indicator. The non-linear $\log$ is applied to the data before the weights multiply it.

Intercept and coefficients: meaning and units core

A coefficient is a price tag: "one more degree costs you 1.6 more drinks". Like a price tag, it has units (drinks per degree), and if you change the units of the input (degrees Fahrenheit instead of Celsius) the number on the tag changes even though nothing about the café changed.

The intercept is the starting value: the prediction when every input is 0. Sometimes that is meaningful (0 ads, 0 emails). Sometimes it is fiction (a café at 0 °C when you only have summer data). With several inputs, each coefficient is the price tag of its input while the other inputs stay where they are.

Three ways to say it:

  • Picture: the slope is how steep the line is; the intercept is where it crosses the vertical line $x = 0$.
  • Numbers: $\hat\beta_1 = 1.6$ drinks per °C is the same fact as $0.889$ drinks per °F.
  • Slogan: a coefficient = change in the average $y$ per one unit of $x$, others held fixed.

Reading the café fit $\hat y = 18 + 1.6\,x$ (x in °C):

  1. Slope: each extra °C goes with $1.6$ more drinks on average. Units: drinks per °C.
  2. Prediction at 22 °C: $18 + 1.6\times22 = 18 + 35.2 = 53.2$ drinks.
  3. Intercept: at 0 °C the line says 18 drinks. But our data run from 10 to 30 °C, so 0 °C is an extrapolation: the line has never been tested there.
  4. Centering: use $x - 20$ (degrees above the average temperature 20 °C) instead. The fit becomes $\hat y = 50 + 1.6\,(x - 20)$: same slope, and the intercept 50 is now the predicted sales on an average day, which is inside the data.
  5. Change units: in Fahrenheit, $F = 32 + 1.8\,C$, so one °F is $1/1.8$ of a °C. The slope becomes $1.6/1.8 \approx 0.889$ drinks per °F and the intercept becomes $18 - 0.889\times32 \approx -10.4$ (drinks at 0 °F, an even wilder extrapolation). The prediction at 22 °C $= 71.6$ °F is still $-10.4 + 0.889\times71.6 \approx 53.2$ ✓.

In $E[y\mid x] = \beta_0 + \beta_1 x_1 + \dots + \beta_{p-1}x_{p-1}$:

  • $\beta_j$ = the change in the average of $y$ when $x_j$ goes up by one unit and all other columns stay fixed. Units: (units of $y$) per (unit of $x_j$).
  • $\beta_0$ = the average of $y$ when all columns are 0. Meaningful only if "all zero" is a real, observed situation; centering the inputs ($x_j - \bar x_j$) makes it "the average $y$ at average inputs".
  • Rescaling a column by $c$ divides its coefficient (and its standard error) by $c$; the t-value, p-value and every prediction stay the same.
  • Indicator (dummy) column $d_i \in \{0, 1\}$: its coefficient is the difference in average $y$ between the "1" group and the "0" group, holding the other columns fixed. With only a treatment dummy and an intercept, $\hat\beta_1 = \bar y_{\text{treat}} - \bar y_{\text{control}}$.
  • Standardized coefficient: fit on z-scored columns ($(x_j - \bar x_j)/s_j$); then $\beta_j$ is the change per one standard deviation of $x_j$, which makes columns with different units comparable.
Why do we need it?

A coefficient is only useful if you can say what it means: "how many more orders per degree, per holiday, per 1% price cut". Getting the units and the "others held fixed" part wrong is the most common way people misread a regression.

Where is it used?

Price elasticity ("−1.2% demand per +1% price"), holiday lifts in your forecasting model, the treatment coefficient in an A/B regression (= difference in means), feature effects reported to product teams, and checking that standardized regressors (for example with one global scaler) give sensible magnitudes.

How is it used?

Read model.params together with the units of each column. Center inputs when the intercept should mean something. Standardize (or compare t-values) before ranking features. Never quote a coefficient without saying what was held fixed.

Same five café days, same fit, different rulers for temperature. Switch between °C, °F and "tens of °C" and tick center x. Notice: the slope and its standard error change by the same factor, so the t-value never changes; the intercept jumps around (it is the prediction at "x = 0", and 0 means a different temperature on each ruler); and the prediction for a 22 °C day is always 53.2.

Each dot is a day: blue = weekday, orange = weekend. The truth is "1.6 drinks per °C" plus a separate weekend effect. The multiple regression (two parallel lines, one per day type) estimates the temperature slope while holding "weekend" fixed. The purple dashed line is the simple regression on temperature alone. Make weekends hotter and busier (both sliders up): the simple slope climbs well above 1.6, because it quietly credits temperature with the weekend crowd. Set the weekend heat to 0 and the two slopes agree again.

"The coefficient for ad spend is 0.002 and for discount is 5, so discounts matter 2 500 times more."

Coefficients carry units. Spend in rupees and discount in percent cannot be compared by raw size. Compare standardized coefficients (per one standard deviation) or t-values, and think about realistic changes of each input.

"The intercept is always a meaningful baseline."

It is the prediction when every column is 0, which may be impossible (0 °C in a summer dataset, a price of 0). Center the columns if you want it to mean "an average day".

"$\hat\beta_j$ is the effect of $x_j$ on $y$."

It is the association of $x_j$ with $y$ holding the other columns fixed. It is a causal effect only if nothing else that moves with $x_j$ is missing from the model, which randomization guarantees and observational data usually do not (Chapter 5.12). Adding or removing another column can change it a lot (the weekend widget).

Forecasting model. A holiday coefficient means "extra demand on that holiday compared with what trend, seasonality and the other regressors already predict for that day". A regressor coefficient is "change in demand per one unit of that regressor, others fixed"; if your pipeline standardizes regressors with one global scaler (Chapter 4.18), the unit is "one standard deviation", and you must undo the scaling before quoting it in business units. A/B framework. Regress the metric on an intercept and a 0/1 treatment column: the treatment coefficient is exactly $\bar y_B - \bar y_A$. Adding a pre-experiment covariate as another column is regression adjustment, the idea behind CUPED (Chapter 5.12).

"Holding the other variables constant, the coefficient tells us what happens if we change $x_j$."

"It tells us how the average of $y$ differs between observations whose $x_j$ differs by one unit and whose other columns are equal. That is a prediction statement; reading it as 'what happens if we intervene' needs a causal argument."

Model answer: "In observational data, a regression coefficient is a conditional association. It equals a causal effect only when, given the included columns, nothing else that drives $y$ moves with $x_j$. In a randomized A/B test that holds for the treatment indicator by design."

$\beta_j$ = change in average $y$ per +1 unit of $x_j$, others fixed. Units: $y$-units per $x_j$-unit.

$\beta_0$ = average $y$ at all-zero inputs; center $x$ to make it meaningful. Rescaling $x$ changes $\beta$ and SE, not $t$ or predictions.

Trap: dropping a correlated column changes the others (omitted-variable bias); coefficients are not causal effects by default.

Quick check: in an A/B regression $y_i = \beta_0 + \beta_1 d_i + \epsilon_i$ with $d_i = 1$ for treatment, the control mean is 10.0 and the treatment mean is 10.4. What are $\hat\beta_0$ and $\hat\beta_1$?

$\hat\beta_0 = 10.0$ (the average when $d = 0$, the control group) and $\hat\beta_1 = 10.4 - 10.0 = 0.4$ (the difference in means). Least squares with only a dummy fits each group's mean exactly.

Residuals and least squares: the line with the smallest total squared miss core

Any line you draw through the cloud misses each dot by some vertical amount. That miss is the dot's residual: what is left over after the line has made its guess. Positive residual: the dot is above the line. Negative: below.

Which line is best? The classic answer: square every miss and add them up, then pick the line that makes this total as small as possible. Squaring makes every miss count as positive, punishes one big miss more than several small ones, and gives a smooth bowl-shaped total with exactly one lowest point. That line is the least-squares (OLS, "ordinary least squares") fit.

Three ways to say it:

  • Picture: draw a square on every vertical miss; least squares makes the total area of the squares as small as possible.
  • Numbers: the café line $18 + 1.6x$ has squared misses adding to 360; the line $10 + 2x$ adds up to 400; no line does better than 360.
  • Slogan: residual = observed − fitted; least squares = smallest sum of squared residuals.

Café data again ($x$ = 10, 15, 20, 25, 30; $y$ = 30, 50, 40, 70, 60).

  1. Means: $\bar x = 100/5 = 20$, $\bar y = 250/5 = 50$.
  2. Deviations: $x - \bar x = -10, -5, 0, 5, 10$; $\;y - \bar y = -20, 0, -10, 20, 10$.
  3. $S_{xy} = \sum(x_i - \bar x)(y_i - \bar y) = 200 + 0 + 0 + 100 + 100 = 400$. $\;S_{xx} = \sum(x_i - \bar x)^2 = 100 + 25 + 0 + 25 + 100 = 250$.
  4. Slope $\hat\beta_1 = S_{xy}/S_{xx} = 400/250 = 1.6$. Intercept $\hat\beta_0 = \bar y - \hat\beta_1\bar x = 50 - 1.6\times20 = 18$.
  5. Fitted values $\hat y_i = 18 + 1.6x_i$: 34, 42, 50, 58, 66. Residuals $e_i = y_i - \hat y_i$: $-4, 8, -10, 12, -6$. They add up to $0$.
  6. Residual sum of squares $RSS = 16 + 64 + 100 + 144 + 36 = 360$.
  7. A rival line $10 + 2x$: fitted 30, 40, 50, 60, 70; residuals $0, 10, -10, 10, -10$; $RSS = 0 + 100 + 100 + 100 + 100 = 400 \gt 360$. It looks just as plausible by eye, but its squares are bigger.

For any candidate coefficients $b$, the fitted values are $\hat y = Xb$ and the residuals are $e = y - Xb$. The residual sum of squares is

$$RSS(b) = \sum_{i=1}^n (y_i - \hat y_i)^2 = \|y - Xb\|^2 .$$

The ordinary least squares (OLS) estimate is $\hat\beta = \arg\min_b RSS(b)$. With one input and an intercept:

$$\hat\beta_1 = \frac{S_{xy}}{S_{xx}} = \frac{\sum (x_i-\bar x)(y_i - \bar y)}{\sum (x_i - \bar x)^2}, \qquad \hat\beta_0 = \bar y - \hat\beta_1 \bar x .$$
  • Residual vs error: the residual $e_i = y_i - \hat y_i$ is measured from the fitted line and can be computed; the error $\epsilon_i = y_i - E[y_i\mid x_i]$ is measured from the true line and can never be seen. Residuals are our window onto the errors.
  • With an intercept, the OLS residuals always add up to zero, are uncorrelated with every column ($\sum x_i e_i = 0$), and the line passes through $(\bar x, \bar y)$.
  • The slope can also be written $\hat\beta_1 = r\, s_y / s_x$ (correlation times the ratio of standard deviations, Chapter 4.15). It needs spread in $x$: if all $x_i$ are equal, $S_{xx} = 0$ and the slope is undefined.
  • Least squares is maximum likelihood when the errors are iid Normal (Chapter 5.2): the squared error is the negative log of a Gaussian density.
Why do we need it?

We need one agreed recipe that turns a cloud of points into one best line, without arguing by eye. Squared misses give a recipe with an exact formula, a unique answer and a direct link to the Normal likelihood.

Where is it used?

Every OLS fit (statsmodels, scikit-learn's LinearRegression), the inner step of IRLS for GLMs (Chapter 5.14), MSE loss in neural networks, least-squares seasonality fits, and the Gaussian-likelihood part of your forecasting model's objective.

How is it used?

Fit with sm.OLS(y, X).fit(); residuals are .resid, fitted values .fittedvalues, the RSS is .ssr. Always plot the residuals afterwards: they are where a bad model shows itself.

Drag the blue dots. The orange line is the least-squares line; the red squares are its squared residuals. Now switch to Show my squares and move the two sliders to build your own (purple) line: try to make its total square area smaller than the orange one's. You never can; the best you can do is tie by pressing Snap my line. Then drag one dot far away: its square becomes huge and the orange line tilts toward it. Squares make far points very influential.

"Least squares minimizes the perpendicular distance from each point to the line."

It minimizes vertical distances, because the model says only $y$ is noisy. Minimizing perpendicular distances is a different method (total least squares, which is what the first principal component does, Chapter 5.16). Swapping $x$ and $y$ gives a different OLS line.

"One unusual point barely matters among many."

A point far from the others, especially at an extreme $x$, can pull the whole line (its square is huge). Look at residual plots and influence measures before trusting a fit; consider a robust loss or a Student-t likelihood when outliers are expected (Chapter 4.9).

"Residuals adding up to zero proves the model is right."

With an intercept they add up to zero for every dataset, by construction. It proves nothing. Patterns in the residuals (next sections) are what tell you something.

"The residuals are the errors of the model."

"The residuals $e_i = y_i - \hat y_i$ are estimates of the unobservable errors $\epsilon_i$. They are not independent even when the errors are (they must add to zero and be orthogonal to every column), and their variance is slightly smaller than $\sigma^2$."

Model answer: "Errors are deviations from the true regression function, which we never observe; residuals are deviations from the fitted function. That is why we divide the RSS by $n - p$, not $n$, to estimate $\sigma^2$: fitting $p$ coefficients uses up $p$ degrees of freedom."

Residual $e_i = y_i - \hat y_i$; $RSS = \sum e_i^2 = \|y - X\hat\beta\|^2$; OLS minimizes it.

One input: $\hat\beta_1 = S_{xy}/S_{xx}$, $\hat\beta_0 = \bar y - \hat\beta_1\bar x$. With an intercept: $\sum e_i = 0$, $\sum x_i e_i = 0$.

Trap: residual ≠ error; vertical distances, not perpendicular; squares make outliers loud.

Quick check: data $(x, y) = (0, 1), (1, 3), (2, 5)$. Find the least-squares line and its RSS without a computer.

$\bar x = 1$, $\bar y = 3$; $S_{xy} = (-1)(-2) + 0 + (1)(2) = 4$; $S_{xx} = 1 + 0 + 1 = 2$; slope $= 2$; intercept $= 3 - 2\times1 = 1$. The line $1 + 2x$ passes through all three points, so every residual is 0 and $RSS = 0$.

The normal equations, and least squares as a projection core

Stand the whole dataset on its head: instead of 5 dots in a 2D chart, think of the 5 observed $y$ values as one arrow in a 5-dimensional space. Each column of $X$ is also an arrow in that space. Every possible fitted vector $Xb$ is a mix of the column arrows, so all of them lie on one flat sheet: the column space of $X$.

The arrow $y$ usually sticks out of that sheet (the noise pushes it out). The closest point on the sheet is straight "below" $y$, like the shadow of a pole when the sun is directly overhead. That shadow is $\hat y$. The leftover arrow from the shadow up to $y$ is the residual vector, and it stands at a right angle to the sheet: perpendicular to every column. "Perpendicular to every column" written as equations are the normal equations ("normal" is the old word for perpendicular).

Three ways to say it:

  • Picture: $\hat y$ is the shadow of $y$ on the plane of the columns; the residual is the pole sticking straight up.
  • Numbers: café residuals $(-4, 8, -10, 12, -6)$ add to 0 (perpendicular to the 1s column) and $\sum x_i e_i = 0$ (perpendicular to the temperature column).
  • Slogan: least squares leaves nothing in the residuals that the columns could still explain.

The café fit, the matrix way. $X$ has a 1s column and the temperatures; $y$ = (30, 50, 40, 70, 60).

  1. $X^\top X = \begin{bmatrix} n & \sum x_i \\ \sum x_i & \sum x_i^2 \end{bmatrix} = \begin{bmatrix} 5 & 100 \\ 100 & 2250 \end{bmatrix}$ (since $100 + 225 + 400 + 625 + 900 = 2250$).
  2. $X^\top y = \begin{bmatrix} \sum y_i \\ \sum x_i y_i \end{bmatrix} = \begin{bmatrix} 250 \\ 5400 \end{bmatrix}$ (since $300 + 750 + 800 + 1750 + 1800 = 5400$).
  3. Solve $X^\top X\,\hat\beta = X^\top y$. The determinant is $5\times2250 - 100^2 = 1250$, so $(X^\top X)^{-1} = \frac{1}{1250}\begin{bmatrix} 2250 & -100 \\ -100 & 5 \end{bmatrix}$.
  4. $\hat\beta = \frac{1}{1250}\begin{bmatrix} 2250\times250 - 100\times5400 \\ -100\times250 + 5\times5400 \end{bmatrix} = \frac{1}{1250}\begin{bmatrix} 22500 \\ 2000 \end{bmatrix} = \begin{bmatrix} 18 \\ 1.6 \end{bmatrix}$ ✓.
  5. Check perpendicularity, $X^\top e = 0$: $\sum e_i = -4 + 8 - 10 + 12 - 6 = 0$ and $\sum x_i e_i = -40 + 120 - 200 + 300 - 180 = 0$ ✓.

Derivation 1 (calculus). $RSS(b) = (y - Xb)^\top(y - Xb)$. Its gradient is $\nabla_b RSS = -2X^\top(y - Xb)$ (Calculus 2.6). Setting it to zero gives the normal equations

$$X^\top X\,\hat\beta = X^\top y \qquad\Longleftrightarrow\qquad X^\top(y - X\hat\beta) = X^\top e = 0 .$$

Derivation 2 (geometry). $\hat y = X\hat\beta$ is the point of the column space closest to $y$, so the residual must be perpendicular to every column: $X^\top e = 0$. Same equations.

  • If the columns of $X$ are linearly independent (full column rank), $X^\top X$ is invertible and $\hat\beta = (X^\top X)^{-1}X^\top y$ is the unique solution. If not (perfect multicollinearity), infinitely many $b$ give the same fit (below).
  • The fitted values are $\hat y = Hy$ with the hat matrix $H = X(X^\top X)^{-1}X^\top$, the projection onto the column space: $H$ is symmetric, $H^2 = H$ (projecting twice changes nothing), and its trace equals $p$, the number of columns. Its diagonal entries $h_{ii}$ are the leverages: how strongly $y_i$ pulls its own fitted value.
  • In code, never form $(X^\top X)^{-1}$ explicitly: use np.linalg.lstsq (QR/SVD) or a solver. Forming $X^\top X$ squares the condition number and loses accuracy when columns are nearly collinear (Linear Algebra 1.10, 1.15).
Why do we need it?

The normal equations turn "find the best line" into "solve one small linear system": an exact answer in one step, with no learning rate and no iterations. The projection picture explains why residuals are uncorrelated with every feature and where the degrees of freedom go.

Where is it used?

Inside every OLS routine (via QR or SVD), ridge regression ($X^\top X + \lambda I$, Chapter 5.3), each IRLS step of a GLM, Kalman-filter updates, closed-form Fourier seasonality fits, and leverage / Cook's distance diagnostics built from $H$.

How is it used?

Call np.linalg.lstsq(X, y, rcond=None) or sm.OLS(y, X).fit(). Check the rank or condition number of $X$ if coefficients look wild. Use $X^\top e \approx 0$ as a sanity check that a fit converged.

column space of X (every possible Xb) column 1 column 2 y (the data) ŷ = Xβ̂ (the shadow) e = y − ŷ perpendicular to the plane: Xᵀe = 0 Least squares = projection y lives in n dimensions (one per observation)
The geometric view (Linear Algebra 1.9): the fitted vector $\hat y$ is the orthogonal projection of $y$ onto the plane spanned by the columns of $X$; the residual $e$ is perpendicular to that plane, which is exactly $X^\top e = 0$.

Here $n = 3$ observations at $x = 1, 2, 3$, so the data vector $y = (y_1, y_2, y_3)$ is a point in 3D (axis 1 = observation 1, and so on). The orange sheet is the column space: every vector $b_0(1,1,1) + b_1(1,2,3)$. Drag the blue point $y$ (Shift = straight up or down). The orange arrow $\hat y$ is its shadow on the sheet and the red arrow is the residual. Rotate to look along the sheet (try Side and Front): the red arrow always meets it at a right angle. The flat chart below shows the same fit the usual way. Press Put y on the sheet: the residual vanishes and the line hits all 3 dots. Press Make y perpendicular: the shadow is 0 and the best line is flat at 0.

"They are called normal equations because the errors are Normal."

"Normal" here means perpendicular: the equations say the residual is perpendicular to every column. They hold for any least-squares fit, whatever the error distribution.

"To fit a regression in code, compute inv(X.T @ X) @ X.T @ y."

It works on tidy textbook data but squares the condition number and can lose many digits when columns are nearly collinear (Fourier terms plus trend plus many regressors). Use np.linalg.lstsq, a QR-based solver, or statsmodels.

Normal equations: $X^\top X\hat\beta = X^\top y$, i.e. $X^\top e = 0$ (residual ⟂ every column).

$\hat\beta = (X^\top X)^{-1}X^\top y$ if full column rank; $\hat y = Hy$, $H = X(X^\top X)^{-1}X^\top$, trace $H = p$.

Trap: "normal" = perpendicular. Solve with lstsq/QR, not an explicit inverse.

Quick check: why must OLS residuals be uncorrelated with every column of $X$ (in the sample)?

Because $X^\top e = 0$ is the defining condition of the least-squares solution. If some column still had a correlation with the residuals, you could add a bit more of that column to the fit and reduce the RSS, so the fit would not have been the minimum.

How sure are the coefficients? Standard errors, t-values, intervals and $R^2$ core

Five different summers would give five different clouds of dots, and five different least-squares lines. The slope 1.6 is just the one this summer happened to give. The standard error of the slope says how much the slope would typically wobble from one summer to the next (the same idea as the SE of a mean, Chapter 5.5).

Three things make the wobble small: little noise (dots close to the line), many observations, and spread-out x values. The last one surprises people: a line pinned down by days from 10 °C to 30 °C is much steadier than one pinned down by days between 19 °C and 21 °C, just as a plank resting on two far-apart supports wobbles less.

Three ways to say it:

  • Picture: redraw the data many times; the fitted lines form a fan, narrowest in the middle of the data. The SE measures the fan's spread of slopes.
  • Numbers: café slope $1.6 \pm 0.69$ (one SE): with 5 noisy days, a true slope of 0 is not ruled out ($p \approx 0.10$).
  • Slogan: SE = noise ÷ (√n × spread of x), roughly.

The café coefficient table, by hand ($n = 5$ days, $p = 2$ columns, $RSS = 360$, $S_{xx} = 250$).

  1. Noise variance estimate: $\hat\sigma^2 = RSS/(n - p) = 360/3 = 120$, so $\hat\sigma \approx 10.95$ drinks (a typical day misses the line by about 11 drinks).
  2. Slope SE: $SE(\hat\beta_1) = \hat\sigma/\sqrt{S_{xx}} = \sqrt{120/250} = \sqrt{0.48} \approx 0.693$.
  3. t-value: $t = \hat\beta_1/SE = 1.6/0.693 \approx 2.31$, with $n - p = 3$ degrees of freedom.
  4. p-value: $P(|T_3| \ge 2.31) \approx 0.104$. Not below 0.05: five days are too few to be sure the slope is not 0.
  5. 95% confidence interval: the $t_3$ critical value is $3.182$, so $1.6 \pm 3.182\times0.693 = 1.6 \pm 2.205 = [-0.60,\ 3.80]$. It contains 0, matching step 4.
  6. $R^2$: total variation $TSS = \sum(y_i - \bar y)^2 = 400 + 0 + 100 + 400 + 100 = 1000$; $R^2 = 1 - RSS/TSS = 1 - 360/1000 = 0.64$. The line explains 64% of the day-to-day variation in sales.

Assume $E[\epsilon\mid X] = 0$ and errors that are independent with constant variance $\sigma^2$. Then, treating $X$ as fixed,

$$Var(\hat\beta) = \sigma^2 (X^\top X)^{-1}, \qquad \hat\sigma^2 = \frac{RSS}{n - p}, \qquad SE(\hat\beta_j) = \sqrt{\hat\sigma^2\,[(X^\top X)^{-1}]_{jj}} .$$
  • One input: $SE(\hat\beta_1) = \hat\sigma/\sqrt{S_{xx}}$. More spread in $x$ (bigger $S_{xx}$) and more data shrink it.
  • $n - p$ is the residual degrees of freedom: $p$ coefficients were fitted, so dividing by $n - p$ makes $\hat\sigma^2$ unbiased.
  • t-value $t_j = \hat\beta_j/SE_j$ tests $H_0: \beta_j = 0$. It follows a $t_{n-p}$ distribution exactly if the errors are Normal, and approximately (by the CLT) for large $n$ otherwise. CI: $\hat\beta_j \pm t_{n-p,\,0.975}\,SE_j$.
  • Mean of y at a new $x_0$ (confidence band): $SE(\hat y_0) = \hat\sigma\sqrt{x_0^\top (X^\top X)^{-1} x_0}$, which for one input is $\hat\sigma\sqrt{1/n + (x_0 - \bar x)^2/S_{xx}}$: narrowest at $\bar x$. A new day's value (prediction band) adds the day's own noise: $\hat\sigma\sqrt{1 + 1/n + (x_0-\bar x)^2/S_{xx}}$.
  • $R^2 = 1 - RSS/TSS$: the share of the variance of $y$ explained by the fit. It never goes down when you add a column, which is why adjusted $R^2 = 1 - \frac{RSS/(n-p)}{TSS/(n-1)}$ exists. For one input, $R^2 = r^2$.
Why do we need it?

A coefficient without its uncertainty cannot support a decision: "+1.6 drinks per degree" could be anything from −0.6 to 3.8 with five days. Standard errors turn a fit into a statement you can test and put error bars on.

Where is it used?

The statsmodels summary() table, A/B regressions with covariates (the treatment coefficient's SE gives the test), significance of regressors and holiday effects, confidence and prediction bands on forecasts, and power calculations for regression designs.

How is it used?

Read .params, .bse, .tvalues, .pvalues, .conf_int() and .rsquared. Then ask whether the SE formula's assumptions hold; if not, use robust (cov_type='HC3') or clustered standard errors.

Drag the dots and watch the table update like a statsmodels summary. The orange band is the 95% confidence band for the average $y$ at each $x$: notice it is narrowest at $\bar x$ and flares at the edges. Press Bunch x together: same $y$ values, but the slope's SE explodes and its interval swallows 0. Press Spread x out: the SE shrinks. Tick prediction band to see the much wider band for a single new day.

The green line is the truth ($18 + 1.6x$). Press Draw 1 sample a few times: each sample of $n$ days gives its own orange line. Press Draw 200 samples: the right panel collects the 200 slopes, and the purple curve is the Normal shape the formula $SE = \sigma/\sqrt{S_{xx}}$ predicts. Compare the measured spread with the formula. Then shrink the spread of x and watch the fan open up; raise $n$ and watch it close.

"$R^2 = 0.64$, so the model is 64% correct."

$R^2$ is the share of variance explained in this sample. A high $R^2$ does not mean the model is right (a curve can be fitted badly by a line with high $R^2$), causal, or good at forecasting; a low $R^2$ does not mean a coefficient is unimportant. For time series, $R^2$ on trending data is almost always high and almost always uninformative.

"The slope's p-value is tiny, so the effect is big."

A tiny p-value means the slope is clearly not zero given the noise and the sample size. With a million rows, a slope of 0.0001 can have $p \lt 10^{-10}$. Look at the coefficient and its interval in real units.

"The SE formula $\sigma^2(X^\top X)^{-1}$ is always valid."

It assumes independent errors with constant variance. Correlated errors (time series, repeated users) or a funnel of variance make it wrong, usually too small. That is the next section.

A/B framework. Regressing a metric on an intercept and a treatment dummy gives a treatment coefficient whose classical SE equals the pooled two-sample t-test SE (Student's version, Chapter 5.9); with cov_type='HC2' it equals Welch's unpooled SE exactly (HC1 is very close). Adding a pre-period covariate shrinks $\hat\sigma$ and therefore the SE: that is why CUPED works. Forecasting model. Your Bayesian fit reports posterior standard deviations instead of SEs. With a Normal likelihood and very wide priors, the posterior mean of β is close to the OLS $\hat\beta$ and the posterior sd is close to its SE; with real priors (Laplace on the changepoint weights, and whatever you use on the other weights) they shrink: a Laplace prior acts like lasso and a Normal prior like ridge (Chapter 5.3).

$Var(\hat\beta) = \sigma^2(X^\top X)^{-1}$; $\hat\sigma^2 = RSS/(n-p)$; one input: $SE(\hat\beta_1) = \hat\sigma/\sqrt{S_{xx}}$.

$t = \hat\beta_j/SE_j \sim t_{n-p}$; CI $\hat\beta_j \pm t_{n-p,0.975}SE_j$; $R^2 = 1 - RSS/TSS$.

Trap: SE needs independent, constant-variance errors; $R^2$ is not "accuracy", and tiny p ≠ big effect.

Quick check: you double the range of temperatures in your data (same $n$, same noise). What happens to the slope's SE?

$S_{xx}$ grows by a factor of 4 (distances from the mean double, squared distances quadruple), so $SE = \sigma/\sqrt{S_{xx}}$ halves. Spreading the inputs out is as good as quadrupling the sample size, for the slope.

The assumptions, and which ones matter for what core

Regression comes with a short list of promises about the noise. People often learn the list as one block ("check all the assumptions!") and then panic about the wrong one. The useful way to learn it: each promise protects a different part of the output.

If the shape of the mean is wrong (a straight line through a curve), the coefficients and predictions themselves are off. If the noise is correlated or has a changing spread, the line is still fine on average but the error bars lie. If the noise is not Normal, almost nothing breaks for a large sample, except exact small-sample p-values and prediction intervals for single new points.

Three ways to say it:

  • Picture: a map from each assumption to the thing it protects (figure below).
  • Numbers: in our simulation a "95%" interval covered the true slope 95% of the time when all held, 88% with a strong variance funnel, and only 48% with strongly correlated errors.
  • Slogan: wrong mean → wrong answer; wrong noise → wrong error bars.

Break one assumption at a time. True model $y = 2 + 1\cdot x + \epsilon$, $n = 40$, $x$ between 0 and 10, error sd about 2, repeated 4 000 times (statsmodels, seeded; your numbers in the widget will differ slightly):

  1. All assumptions hold: average slope 1.001; the usual 95% interval covered the true slope in 95.1% of runs.
  2. Funnel (error sd grows like $x^2$): average slope 1.000 (still unbiased), but usual coverage only 87.6%; robust (HC1) coverage 93.0%.
  3. Correlated errors ($x$ = time, AR(1) errors with $\phi = 0.8$): average slope 1.005 (unbiased), but coverage 47.6%: the usual SE is far too small because 40 correlated days carry much less information than 40 independent ones.
  4. Heavy-tailed errors (Student-t, 3 df, same sd): coverage 94.9%. Skewed errors (shifted exponential): 95.5%. With $n = 40$ the CLT already protects the coefficient intervals.
  5. Curved truth ($+0.2(x-5)^2$): the line is the wrong shape, so predictions at the edges are biased (at $x = 10$ the line predicts about 13.6 on average when the truth is 17).

The classical assumptions of the linear model, and what each one buys:

  1. Linearity / correct mean: $E[\epsilon\mid X] = 0$, i.e. $E[y\mid X] = X\beta$ with the right columns. Buys: unbiased $\hat\beta$ and unbiased predictions. Broken by a missing curve, a missing interaction, or an omitted variable correlated with the included ones.
  2. Independent errors. Buys: correct standard errors, intervals and p-values. Broken by time series (autocorrelation), repeated measurements of the same user, clustered designs.
  3. Constant variance (homoscedasticity), $Var(\epsilon_i\mid X) = \sigma^2$. Buys: correct usual SEs, and with 1–2, the Gauss–Markov theorem: OLS is the most precise of all linear unbiased estimators ("BLUE"). No Normality is needed for this.
  4. Normal errors, $\epsilon_i \sim N(0,\sigma^2)$. Buys: exact t and F distributions for any $n$, OLS = maximum likelihood, and valid prediction intervals for single new observations. For large $n$ the CLT makes coefficient inference approximately right anyway; prediction intervals still depend on the real error shape.
  5. No perfect multicollinearity (full column rank). Buys: a unique $\hat\beta$.

Assumptions 2–4 are about the errors, and you check them through the residuals (next section).

Why do we need it?

Knowing which promise protects which output tells you what to worry about. A non-Normal residual histogram with 10 000 rows is usually harmless for coefficients; autocorrelated residuals in a daily series are serious even when the histogram looks perfect.

Where is it used?

Model review and interviews ("what are the OLS assumptions?"), choosing robust or clustered SEs in A/B analyses (users with many sessions), deciding between Normal and Student-t likelihoods in forecasting, and residual diagnostics (Chapter 7.17).

How is it used?

Fit, then check in this order: residual vs fitted (mean shape, funnel), residual vs time and its ACF (independence), Q-Q (tails). Fix the mean first (columns, transformations), then fix the error bars (robust or clustered SEs, a better likelihood).

Assumption What it protects 1 · correct mean, E[ε | X] = 0 2 · independent errors 3 · constant variance 4 · Normal errors 5 · full column rank β̂ and predictions unbiased usual SEs, CIs, p-values right OLS most precise (BLUE) exact small-n t/F tests,prediction intervals a unique β̂ exists (everything to the right also assumes 1)
Each assumption protects a different output. Breaking 1 biases the answer itself; breaking 2 or 3 leaves β̂ unbiased but makes the usual error bars wrong (dashed: OLS also stops being the most precise method); breaking 4 mostly matters for small samples and for prediction intervals.

Pick a scenario. The left panel shows one dataset (green = true mean, orange = least-squares line). Press Run 500 experiments: the right panel collects 500 fitted slopes (truth = 1, green line), and the bars show how often the usual and the robust 95% intervals really contained the truth. Compare Correlated errors (slopes centred on 1, intervals badly too narrow) with Curved truth (intervals fine, but the line itself is the wrong shape: look at the edge prediction).

"OLS requires the data (or $y$) to be Normally distributed."

Normality is about the errors, and it is the least important assumption for coefficient inference with a decent sample. Unbiasedness and the Gauss–Markov result need no Normality at all.

"Heteroscedasticity or autocorrelation biases the coefficients."

With a correct mean model, $\hat\beta$ stays unbiased; what goes wrong is the standard error (and efficiency). The fix is in the error bars (robust, clustered or HAC standard errors, which allow for autocorrelation) or a better noise model, not necessarily in the coefficients.

"Robust (HC) standard errors fix correlated errors."

HC errors handle unequal variances but still assume independence. For correlated errors use clustered SEs (groups) or HAC/Newey–West SEs (time series), or model the correlation.

"The assumptions of linear regression are that X and y are Normal and there are no outliers."

"Correct mean specification ($E[\epsilon\mid X]=0$), independent errors, constant error variance, full-rank $X$, and, for exact small-sample inference, Normal errors. The first is about the coefficients; the next two about the standard errors; Normality about exact tests and prediction intervals."

Model answer: "If the mean is misspecified, the estimates are biased and no standard-error fix helps. Heteroscedasticity and autocorrelation leave OLS unbiased but invalidate the usual SEs, so I use robust or clustered SEs or model the noise. Non-Normal errors matter mostly in small samples and for prediction intervals."

In your forecasting model, the residuals over time are often autocorrelated (a busy stretch stays busy). The model's point forecast can still be fine, but its uncertainty bands will be too narrow if the likelihood treats days as independent: this is exactly the "correlated errors" row, and why Chapter 7.17 checks the residual ACF. In the A/B framework, independence is about the units: if you analyse sessions but randomize users, sessions of one user are correlated, and the per-session SE is too small (Chapter 5.10).

Correct mean → unbiased β̂. Independence + constant variance → correct usual SEs and BLUE (Gauss–Markov). Normal errors → exact small-n tests and prediction intervals. Full rank → unique β̂.

Wrong mean = wrong answer; wrong noise = wrong error bars.

Trap: HC SEs fix unequal variance, not correlation (use clustered / HAC).

Quick check: a residual histogram looks clearly skewed, but $n = 5\,000$. Should you worry about the coefficient p-values?

Not much: with 5 000 observations the CLT makes the coefficient estimates approximately Normal, so t-based intervals and p-values are close to right (if errors are independent with constant variance). The skew matters more for prediction intervals of single new observations, and it may hint that a log transform or a different likelihood fits better.

Residual diagnostics: residual vs fitted, Q-Q, residual vs time core

After the model has taken its share, the residuals are the leftovers on the plate. If the model did its job, the leftovers are boring crumbs: no shape, no trend, no pattern, the same amount everywhere. If you can see a shape in the leftovers, the model missed something, and the shape tells you what it missed.

A curve in the leftovers means the mean was the wrong shape. A funnel means the noise grows. Long runs of positive then negative leftovers over time mean the days are connected. A Q-Q plot that bends at the ends means the surprises have heavier tails than a Normal.

Three ways to say it:

  • Picture: a good residual plot is a shapeless horizontal band around 0.
  • Numbers: fit a line to $y = x^2$ at $x = 1..5$ and the residuals are $2, -1, -2, -1, 2$: a U, the signature of a missing curve.
  • Slogan: patterns in residuals are messages from the data about what the model left out.

A missed curve. $x = 1, 2, 3, 4, 5$ and $y = x^2 = 1, 4, 9, 16, 25$; fit a straight line.

  1. $\bar x = 3$, $\bar y = 55/5 = 11$; $S_{xy} = (-2)(-10) + (-1)(-7) + 0 + (1)(5) + (2)(14) = 20 + 7 + 5 + 28 = 60$; $S_{xx} = 10$.
  2. Slope $= 60/10 = 6$; intercept $= 11 - 6\times3 = -7$. Fitted: $-1, 5, 11, 17, 23$.
  3. Residuals: $1-(-1) = 2$, $4 - 5 = -1$, $9 - 11 = -2$, $16 - 17 = -1$, $25 - 23 = 2$.
  4. Plotted against the fitted values: high, low, lower, low, high: a U shape. $R^2$ is 0.96, which looks great, yet the residuals show the line is the wrong model. Adding an $x^2$ column fixes it exactly.

The standard residual plots, and what each one detects:

  • Residuals vs fitted values $(\hat y_i, e_i)$: should be a flat band around 0. A curve → wrong mean shape (missing $x^2$, interaction, transformation). A funnel → variance changes with the level (heteroscedasticity, next section).
  • Normal Q-Q plot of residuals (Chapter 4.17): points on the line → Normal-like errors. Ends bending away (S-shape) → heavy tails; one-sided bend → skew; a few points far off → outliers.
  • Residuals vs time (or observation order) and the ACF of residuals: runs and waves → autocorrelation (independence broken; Chapter 7.3, 7.17).
  • Residuals vs each input, and vs inputs you left out: a pattern against a left-out variable means it belongs in the model.
  • Influence: leverage $h_{ii}$ (an unusual $x$) times a large residual = a point that moves the fit; Cook's distance combines both.

Use standardized residuals $e_i/(\hat\sigma\sqrt{1-h_{ii}})$ when comparing sizes: about 95% should lie within ±2 if the errors are Normal.

Why do we need it?

Every number in the summary table is computed as if the assumptions held. Residual plots are how you find out whether they do, and they point to the fix (a column, a transformation, a likelihood, a correlation model).

Where is it used?

After every regression fit; forecast residual checks in your model (Chapter 7.17); choosing Normal vs Student-t likelihoods from residual Q-Q plots; posterior predictive checks (Chapter 6.8) are the Bayesian cousin.

How is it used?

Plot model.resid against model.fittedvalues, against time, and in sm.qqplot(model.resid, line='45', fit=True); compute plot_acf(model.resid) for time series. Look for shapes, not for perfection: small samples wiggle.

"I made a Q-Q plot of $y$ and it is not straight, so regression is invalid."

Check the residuals, not $y$. $y$ mixes the signal and the noise; only the noise is assumed Normal.

"With 15 points the residual plot wiggles, so the model is wrong."

Small samples always wiggle (the New sample button shows how much under "All good"). Look for clear shapes, and compare with simulated datasets from the fitted model if unsure.

"An outlier in the residuals should be deleted."

First find out why it is there (a data error, a holiday you did not model, a real extreme day). Delete only errors; otherwise model it (an indicator, a heavier-tailed likelihood) or report results with and without it.

In your forecasting model these exact plots answer the syllabus questions: a Q-Q of residuals bending at the ends says "consider the Student-t likelihood"; a funnel against the fitted level says "consider a log transform or the Negative Binomial likelihood, whose variance grows with the mean" (Chapter 5.14); waves in residuals over time say "trend, seasonality and regressors have not captured the day-to-day dependence" (Chapter 7.17). A residual spike on the same date every year is a holiday asking for its own column.

Residual vs fitted: curve = wrong mean; funnel = changing variance. Q-Q: S-shape = heavy tails. Residual vs time / ACF: waves = correlated errors.

Good model = shapeless band, straight Q-Q, no waves.

Trap: check residuals, not $y$; small samples wiggle; don't delete outliers blindly.

Quick check: residuals of a daily demand regression look perfect against fitted values, but are positive every Saturday. What did the model miss?

A weekly pattern: the model has no day-of-week (or weekly Fourier) columns, so Saturday's extra demand ends up in the residuals. Plotting residuals against day of week, or against time, reveals it; adding weekly seasonality columns removes it.

Heteroscedasticity: when the noise grows with the level

A small corner shop sells about 50 items a day, give or take 5. A supermarket sells about 1 000, give or take 100. Both are equally "predictable" in percentage terms, but in raw units the supermarket's surprises are 20 times larger. Put both in one regression and the residual plot opens like a funnel. That is heteroscedasticity: "different scatter" (Greek hetero = different, skedasis = scattering). The opposite, constant scatter, is homoscedasticity.

The line is still aimed correctly. The trouble is the error bars: the usual SE formula assumes every point is equally noisy, so it gets the uncertainty wrong, and OLS wastes information by trusting the noisy points as much as the quiet ones.

Three ways to say it:

  • Picture: a funnel in the residual-vs-fitted plot.
  • Numbers: 50 ± 5 and 1 000 ± 100 are both ±10%; after a log, both become ±0.095.
  • Slogan: unequal noise leaves the line unbiased but makes the usual error bars lie.

Two ways to tame a funnel.

  1. Log transform. If the noise is a percentage of the level, the log turns it into a constant: $\log 55 - \log 50 = \log 1.1 \approx 0.0953$, and $\log 1100 - \log 1000 = \log 1.1 \approx 0.0953$ too. On the log scale both shops wobble by the same amount (Chapter 4.18).
  2. Weights. If one point is twice as noisy as another (sd 10 vs 5), its variance is 4 times bigger, so weighted least squares gives it weight $1/10^2 = 0.01$ against $1/5^2 = 0.04$: a quarter of the influence.
  3. Robust standard errors keep the OLS line and fix only the error bars, by replacing the single $\hat\sigma^2$ with each point's own squared residual $e_i^2$ in the variance formula (below).

Heteroscedasticity: $Var(\epsilon_i\mid x_i) = \sigma_i^2$ that changes across observations (often growing with the mean or with an input).

  • Consequences (mean model correct): $\hat\beta$ is still unbiased; the usual $SE$ from $\hat\sigma^2(X^\top X)^{-1}$ is wrong (too small when the noisiest points also have extreme $x$); OLS is no longer the most precise estimator; prediction intervals have the wrong width (too wide for quiet points, too narrow for noisy ones).
  • Robust ("sandwich", HC) standard errors: $\widehat{Var}(\hat\beta) = (X^\top X)^{-1}\big(\sum_i e_i^2 x_i x_i^\top\big)(X^\top X)^{-1}$ (HC0); HC1 multiplies by $n/(n-p)$; HC3 uses $e_i^2/(1-h_{ii})^2$ and behaves better in small samples. In statsmodels: fit(cov_type='HC3').
  • Weighted least squares: minimize $\sum_i w_i (y_i - x_i^\top b)^2$ with $w_i = 1/\sigma_i^2$; solution $(X^\top W X)^{-1}X^\top W y$. Best when the variance pattern is known or well modelled.
  • Transform $y$ (log, square root, Box-Cox) when the spread grows with the level; or use a likelihood whose variance depends on the mean (Poisson, Negative Binomial: Chapter 5.14); or model the variance explicitly (let $\sigma$ depend on inputs).
  • Detect: a funnel in residual vs fitted; formal tests such as Breusch–Pagan (regress $e_i^2$ on the inputs).
Why do we need it?

Real business data are almost never equally noisy everywhere: big days, big stores and big spenders vary more. Without handling this, p-values and intervals are miscalibrated and forecasts get the wrong uncertainty bands.

Where is it used?

Robust SEs in A/B regressions (revenue per user is wildly unequal), log-demand models, WLS for averaged data with different group sizes, Poisson/NB likelihoods for counts, and variance models in forecasting.

How is it used?

Look at residual vs fitted. If it funnels: try a log (or Box-Cox) on $y$; otherwise report cov_type='HC3' SEs; use WLS if you know the variances; in a Bayesian model choose a likelihood whose spread grows with the mean.

Raise the funnel strength: the noise sd is $2(x/5)^{\gamma}$, so $\gamma = 0$ is constant noise. The right panel shows the residuals opening into a funnel. Press Run 500 experiments: compare how often the usual and the robust 95% intervals contain the true slope 1, and compare how much the OLS slopes and the weighted (WLS, weights $1/\sigma_i^2$) slopes scatter. At $\gamma = 0$ they all agree; at large $\gamma$ the usual interval undercovers and WLS is clearly more precise.

"Heteroscedasticity makes the regression line wrong."

If the mean model is right, the line is still unbiased. What breaks is the uncertainty (SEs, p-values, intervals) and the efficiency.

"Just take logs; that always fixes it."

Logs fix spread that grows proportionally with the level, and need positive data (zeros need log1p or a count likelihood). They also change what the coefficients mean (percent effects) and back-transformed predictions estimate the median, not the mean (Chapter 4.18).

Demand series whose level grows over time usually show a funnel: a ±10% day is ±50 units early on and ±500 later. A Normal likelihood with one constant σ then over-states uncertainty early and under-states it late. Your options mirror the list above: model log-demand, use the Negative Binomial likelihood (variance $\mu + \mu^2/\alpha$ grows with the mean by construction, Chapter 5.14, 7.13), or let σ depend on the level. In A/B tests on revenue, report robust SEs: a few big spenders make the noise very unequal.

Heteroscedasticity: $Var(\epsilon_i)$ varies; residual funnel.

β̂ unbiased, usual SE wrong. Fixes: robust HC SEs (cov_type='HC3'), WLS with $w_i = 1/\sigma_i^2$, log/Box-Cox, a mean-dependent likelihood (Poisson/NB).

Trap: it is an error-bar problem, not a coefficient problem.

Quick check: residual sd is about 5% of the fitted value at every level. Which simple transformation should you try first, and why?

The log of $y$. Noise proportional to the level is multiplicative, and the log turns multiplication into addition: on the log scale every observation has residual sd ≈ 0.05, so the funnel disappears.

Multicollinearity and the variance inflation factor core

Two cooks salt the same soup at the same time, every time. You can taste that the soup is salty, but you can never tell how much salt came from each cook. Any split ("all from cook 1", "half each", "cook 2 added extra and cook 1 took some out") explains the taste equally well.

That is multicollinearity: two (or more) columns of $X$ that move together. The model knows their combined effect very well, but not how to share it between them. So the individual coefficients become wobbly: a tiny change in the data can swap credit from one to the other, even flip a sign, while the predictions hardly move.

Three ways to say it:

  • Picture: in the plane of possible $(\beta_1, \beta_2)$, the good fits form a long thin valley along "$\beta_1 + \beta_2 = $ constant".
  • Numbers: changing one $y$ value by 1 moves the coefficients from (2, 0) to (1, 1) or (3, −1); VIF = 37.
  • Slogan: collinearity hurts the coefficients, not the predictions.

Two ad channels that always rise together. Search spend $x_1 = 1, 2, 3, 4, 5$ and social spend $x_2 = 1, 2, 3, 4, 6$ (thousands); orders $y$ in hundreds. Correlation $r(x_1, x_2) \approx 0.986$.

  1. $y = 2, 4, 6, 8, 10$: least squares gives $\hat y = 0 + 2x_1 + 0x_2$, a perfect fit. "Search does everything, social nothing."
  2. Change only the last day to $y_5 = 11$: now $\hat y = 0 + 1x_1 + 1x_2$, again a perfect fit. "They contribute equally."
  3. Change it to $y_5 = 9$ instead: $\hat y = 0 + 3x_1 - 1x_2$. "Social spend reduces orders?!"
  4. In all three, the fitted values change only on day 5, by exactly the 1 unit we changed. Predictions are stable; the split of credit is not.
  5. Variance inflation factor: $VIF = 1/(1 - r^2) = 1/(1 - 0.986^2) \approx 1/0.027 = 37$. Each coefficient's variance is 37 times what it would be with uncorrelated columns; its SE is $\sqrt{37} \approx 6.1$ times larger.

Multicollinearity: some column of $X$ is (nearly) a linear combination of the others.

  • Perfect collinearity: an exact combination (e.g. an intercept plus a dummy for every category: the "dummy trap"; temperature in °C and in °F together). $X^\top X$ is singular and $\hat\beta$ is not unique.
  • Near collinearity: $\hat\beta$ exists but has huge variance. For column $j$, $$Var(\hat\beta_j) = \frac{\sigma^2}{S_{x_jx_j}}\cdot\frac{1}{1 - R_j^2}, \qquad VIF_j = \frac{1}{1 - R_j^2},$$ where $R_j^2$ is the $R^2$ from regressing $x_j$ on all the other columns. With two inputs, $R_j^2 = r^2$.
  • Symptoms: large SEs; coefficients that change a lot or flip sign when you add or drop a column or a few rows; individual t-values small while the columns are jointly significant (F-test); good predictions within the data range.
  • Rules of thumb (only rules of thumb): VIF above 5 or 10 deserves a look; a large condition number of the standardized $X$ (say above 30) is another warning sign.
  • Remedies, depending on the goal: for prediction, often nothing; combine or drop redundant columns; collect data where they move separately; regularize (ridge, or Normal/Laplace priors, Chapter 5.3), which picks a sensible point in the valley; principal-components regression (Chapter 5.16); center $x$ before forming $x^2$ or interactions.
Why do we need it?

Without knowing about collinearity you will over-interpret wobbly coefficients ("social ads hurt sales!") and be surprised when they flip next month. VIFs tell you which coefficients the data can and cannot pin down.

Where is it used?

Marketing-mix and pricing models (channels move together), forecasting models with many regressors, holidays that coincide with promotions, polynomial and interaction features, and the identifiability discussion for trend vs seasonality vs holidays (Chapter 6.8).

How is it used?

Check the correlation matrix of candidate regressors (Chapter 4.15) and variance_inflation_factor(X, j) from statsmodels. Decide by goal: keep for prediction, combine/drop/regularize for interpretation, and report joint effects instead of individual ones.

Uncorrelated columns Collinear columns (r ≈ 0.99) β₁β₂ good fits: a small round region β₁β₂ β₁ + β₂ ≈ constant (well determined) each βⱼ alone: poorly determined (long valley)
Regions of nearly equally good coefficient pairs (confidence ellipses). With uncorrelated columns, both coefficients are pinned down. With collinear columns, the data only pin down their sum: the region is a long thin valley, so each coefficient alone is very uncertain.

The truth is $y = 1\cdot x_1 + 1\cdot x_2 + \epsilon$ with $n = 50$. Set the correlation $\rho$ between the two inputs and press Draw 300 samples: each blue dot is one sample's $(\hat\beta_1, \hat\beta_2)$, the green dot is the truth. At $\rho = 0$ the cloud is round. Push $\rho$ toward 0.99: the cloud stretches into a thin line along $\beta_1 + \beta_2 = 2$ (purple), many samples even get the wrong sign, yet the readout shows the sum is still estimated precisely.

"High VIF means the model is bad; drop the variable."

It depends on the goal. For prediction within the usual range, collinearity is mostly harmless. For interpreting individual coefficients it is a real problem, but dropping a column that truly matters creates omitted-variable bias in the others. Consider combining, regularizing, or reporting the joint effect.

"No pair of columns has correlation above 0.8, so there is no multicollinearity."

Three or more columns can be nearly dependent with only moderate pairwise correlations (e.g. $x_3 \approx x_1 + x_2$). VIF (regress each column on all the others) catches this; a pairwise correlation matrix may not.

"Collinearity makes the coefficients biased."

OLS stays unbiased; the coefficients just have large variance. Regularization (ridge, priors) deliberately adds a little bias to cut that variance (Chapter 5.1, 5.3).

Your forecasting model is full of potential collinearity: a holiday that always coincides with a promotion regressor (the model cannot split credit between them); a yearly Fourier wave and a trend changepoint when you have barely more than a year of data; day-of-week dummies added on top of weekly Fourier terms (redundant columns). Bayesian priors play the role of ridge/lasso here: they keep the coefficients in a sensible part of the long valley, and posterior correlations between, say, the holiday and promo coefficients reveal the problem. This is the "which component gets the credit?" identifiability question of Chapter 6.8; check a correlation matrix of candidate regressors before adding them (Chapter 4.15).

"Multicollinearity makes my predictions unreliable."

"Multicollinearity inflates the variance of individual coefficients; predictions for inputs like the training data stay stable. It becomes a prediction problem only when the correlated inputs move apart in new data, i.e. extrapolation."

Model answer: "I'd compute VIFs; if they are high, I would not interpret the individual coefficients, but report the combined effect or regularize. For pure forecasting I'd watch for situations where the correlated regressors decouple in the future."

$VIF_j = 1/(1 - R_j^2)$; $SE(\hat\beta_j)$ grows by $\sqrt{VIF_j}$. Rules of thumb: VIF over 5–10.

Symptoms: big SEs, sign flips, unstable when adding columns; predictions fine. Perfect collinearity (dummy trap) → no unique β̂.

Trap: unbiased but high-variance; dropping a true variable causes omitted-variable bias. Fix by goal: combine, regularize, more varied data.

Quick check: you include an intercept and dummy columns for all 7 days of the week. What happens, and how do you fix it?

The 7 dummies add up to the 1s column for every row, so the columns are exactly linearly dependent (the dummy trap): $X^\top X$ is singular and the coefficients are not unique. Drop one day's dummy (it becomes the baseline) or drop the intercept.

Your forecasting model is a regression: the design matrix core

Your forecasting model $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$ looks like four different machines. Look closer and each machine is just a few columns with weights. A rising trend is the column "day number $t$". A changepoint is a column that is 0 until the changepoint and then rises like a ramp. Weekly seasonality is a handful of sine and cosine columns. A holiday is a 0/1 column. An external regressor is its own column.

Line all those columns up side by side and you have one design matrix. The model's mean is that matrix times one long vector of weights. Everything in this chapter (coefficients, least squares, collinearity, residual checks) applies to it directly. What your Bayesian model adds is priors on the weights and a choice of likelihood; the skeleton is regression.

Three ways to say it:

  • Picture: a table with one row per day and column groups for trend, changepoints, seasonality, holidays and regressors.
  • Numbers: 1 intercept + 1 trend + 25 changepoint ramps + 6 weekly Fourier + 20 yearly Fourier + 10 holiday flags + 3 regressors = 66 columns, 66 weights.
  • Slogan: components are column groups; fitting the model is finding one weight per column.

Five days of a design matrix. Days $t = 0, \dots, 4$; one changepoint at $s = 2$; weekly Fourier order 1 (period 7); a holiday on day 3; a temperature-anomaly regressor $r_t$.

$t$1$t$$(t-2)_+$$\sin\frac{2\pi t}{7}$$\cos\frac{2\pi t}{7}$holiday$r_t$
01000.0001.00000.5
11100.7820.6230−1.0
21200.975−0.22300.0
31310.434−0.90111.5
4142−0.434−0.90102.0
  1. $(t - 2)_+$ means "$t - 2$ if positive, else 0": the ramp that switches on at the changepoint.
  2. Weights $\beta = (100,\ 0.4,\ 0.8,\ 6,\ 4,\ 25,\ 5)$: base level 100, slope 0.4 per day, slope change $+0.8$ after day 2, weekly wave amplitudes 6 and 4, holiday lift 25, 5 units per degree of anomaly.
  3. Row $t = 3$: $100 + 0.4\times3 + 0.8\times1 + 6\times0.434 + 4\times(-0.901) + 25\times1 + 5\times1.5$.
  4. $= 100 + 1.2 + 0.8 + 2.603 - 3.604 + 25 + 7.5 = 133.5$. That is the model's mean for day 3: one row times the weight vector.
  5. Slope before the changepoint: 0.4 per day; after it: $0.4 + 0.8 = 1.2$ per day. One column, one weight, one change in trend, and the line stays continuous at $t = 2$ because the ramp starts at 0.

The mean of the additive forecasting model (Chapter 7.7) is a linear model $\mu = X\beta$ with column groups:

  • Trend $g(t)$: columns $1$ and $t$ (weights $m$, $k$) plus one hinge column $(t - s_j)_+$ per candidate changepoint $s_j$ (weight $\delta_j$). Prophet writes the trend as $(k + a(t)^\top\delta)\,t + (m + a(t)^\top\gamma)$ with $a_j(t) = \mathbb{1}[t \ge s_j]$ and $\gamma_j = -s_j\delta_j$; substituting gives exactly $kt + m + \sum_j \delta_j (t - s_j)_+$, linear in $(m, k, \delta)$ (Chapter 7.8).
  • Seasonality $s(t)$: for each period $P$ and order $N$, the $2N$ columns $\sin(2\pi n t/P), \cos(2\pi n t/P)$, $n = 1..N$ (Chapter 7.11).
  • Holidays $h(t)$: 0/1 indicator columns, one per holiday (or per day of a holiday window) (Chapter 7.12).
  • Regressors $X_t\beta$: one column per external variable (often standardized).

Only the weights are estimated; the changepoint locations $s_j$, the periods and the orders are chosen beforehand (grid, PELT, defaults), so the model stays linear in its weights. The Bayesian version keeps this $X$ and adds: priors ($\delta_j \sim \text{Laplace}(0, b)$, which acts like an L1 penalty at the MAP; Normal priors on seasonal, holiday and regressor weights are a common choice and act like ridge, Chapter 5.3; check which priors your code uses), a likelihood (Normal, Student-t, or Negative Binomial through a positive link, Chapter 5.14), and fitting by SVI instead of the normal equations.

Why do we need it?

Seeing the model as a design matrix makes it explainable: every component's effect is "its columns times their weights", collinearity between components becomes visible, and the number of parameters (which drives SVI guide size) is just the number of columns.

Where is it used?

Prophet and Prophet-style NumPyro models, regression with time features as an ARIMA alternative (Chapter 7.6), marketing-mix models, and any "linear model + structured features" baseline you compare against.

How is it used?

Build each block as an array of shape (n_days, n_cols), np.column_stack them, and compute the mean as X @ beta (in NumPyro: jnp.dot(X, beta) with priors on beta). Check the column count, the rank and correlations before fitting.

μ = X β : one row per day, column groups per component 1, ttrend (t − sⱼ)₊changepoints δⱼ sin, cos (2πnt/P)Fourier seasonality 0/1 flagsholidays Xₜregressors t=0 t=n × β (priors: Laplace on δ, Normal on the rest) time runs down the rows; each wiggly line is one column drawn as a time series
The forecasting model's mean as one design matrix. Each column is itself a time series (drawn vertically, time running down): constant and ramp for the trend, hinge ramps that start at each changepoint, sine/cosine waves for seasonality, spikes for holidays, and the external regressors. The model learns one weight per column.

The blue dots are 16 weeks of simulated daily demand (true ingredients: a trend whose slope changes on day 56, a weekly pattern, two holidays and a temperature regressor). Start with only the trend column and add column groups one at a time: the changepoint ramp, the weekly Fourier order, the holiday flag, the regressor. Watch the orange fit improve, the RSS drop and the estimated weights approach the truth. The bottom panel draws every active column as its own little time series: that stack is $X$. Try Fourier order 4 for a surprise.

"Fourier terms and changepoints make the model non-linear, so least-squares ideas do not apply."

The columns are non-linear functions of time, but the mean is linear in the weights. Everything here applies: collinearity between components, leverage of unusual days, the dependence of uncertainty on how spread out the columns are.

"The model estimates where the changepoints are."

In this design the locations $s_j$ are fixed in advance (a grid, or PELT on the training data); only the slope changes $\delta_j$ are weights. Their location uncertainty is not part of the posterior (Chapter 7.9).

"More columns always help."

More columns always lower the training RSS, but each one adds variance (and collinearity). Fourier order, the number of changepoints and their prior scales are bias–variance knobs (Chapter 7.18), and only a holdout tells you whether a column helps.

This is the cleanest interview explanation of your forecasting model: "the mean is a linear model with structured columns: a constant and a ramp for the trend, one hinge column per candidate changepoint chosen by a Prophet-like grid plus PELT, Fourier columns for each seasonality, indicator columns for holidays and one column per external regressor. I put a Laplace prior on the changepoint weights (sparse-ish slope changes) and priors on the others, choose a Normal, Student-t or Negative Binomial likelihood, and fit everything with SVI." It also explains the guide choice: the number of latent weights is the number of columns, which is what decides between a full-rank and a low-rank Gaussian guide (Chapter 6.13).

Trend: $1, t, (t - s_j)_+$; seasonality: $\sin, \cos(2\pi n t/P)$; holidays: 0/1; regressors: their values. Mean $= X\beta$.

Prophet's continuous piecewise trend $= kt + m + \sum_j\delta_j(t - s_j)_+$ (using $\gamma_j = -s_j\delta_j$).

Trap: changepoint locations are fixed, not estimated; columns are non-linear in $t$ but the model is linear in β.

Quick check: yearly Fourier order 10, weekly order 3, 25 candidate changepoints, 8 holidays, 2 regressors, plus intercept and trend. How many columns (weights) does the mean have?

Yearly: $2\times10 = 20$; weekly: $2\times3 = 6$; changepoints 25; holidays 8; regressors 2; intercept and trend 2. Total $20 + 6 + 25 + 8 + 2 + 2 = 63$ columns, so 63 weights for the mean (plus noise parameters such as σ, ν or α).

Recap, cheat sheet and practice

  • Model: $y = X\beta + \epsilon$: rows = observations, columns = features (first column 1s), one weight per column, noise with $E[\epsilon\mid X] = 0$. "Linear" means linear in β; columns may be $x^2$, $\sin$, 0/1 flags.
  • Coefficients: change in the average $y$ per unit of $x_j$, others held fixed; units matter; the intercept is the prediction at all-zero inputs (center to make it meaningful). Not causal by default.
  • Least squares: minimize $RSS = \sum e_i^2$; one input: $\hat\beta_1 = S_{xy}/S_{xx}$, $\hat\beta_0 = \bar y - \hat\beta_1\bar x$. Residual (observable) ≠ error (unobservable).
  • Normal equations $X^\top X\hat\beta = X^\top y$, i.e. $X^\top e = 0$: $\hat y$ is the projection of $y$ onto the column space. Solve with lstsq/QR.
  • Uncertainty: $Var(\hat\beta) = \sigma^2(X^\top X)^{-1}$, $\hat\sigma^2 = RSS/(n-p)$, $t = \hat\beta_j/SE_j$ on $n - p$ df; $R^2 = 1 - RSS/TSS$.
  • Assumptions: correct mean → unbiased; independence + constant variance → correct SEs (and BLUE); Normal errors → exact small-n tests and prediction intervals; full rank → unique β̂.
  • Diagnostics: residual vs fitted (curve, funnel), Q-Q (tails), residual vs time / ACF (correlation). Heteroscedasticity → robust SEs, WLS, log, a mean-dependent likelihood. Multicollinearity → VIF $= 1/(1-R_j^2)$; wobbly coefficients, stable predictions.
  • Your forecasting model is a regression with column groups: $1, t$, hinges $(t - s_j)_+$, Fourier $\sin/\cos$, holiday flags, regressors; plus priors and a likelihood, fitted by SVI.

Cheat sheet

IdeaFormulaRemember
Model$y = X\beta + \epsilon$, $E[\epsilon\mid X] = 0$linear in β, not in x
Simple OLS$\hat\beta_1 = S_{xy}/S_{xx}$, $\hat\beta_0 = \bar y - \hat\beta_1\bar x$needs spread in x
Normal equations$X^\top X\hat\beta = X^\top y$ ⇔ $X^\top e = 0$"normal" = perpendicular
Hat matrix$\hat y = Hy$, $H = X(X^\top X)^{-1}X^\top$projection; trace $= p$; leverage $h_{ii}$
Noise estimate$\hat\sigma^2 = RSS/(n - p)$$p$ counts the intercept
Coefficient SE$\sqrt{\hat\sigma^2[(X^\top X)^{-1}]_{jj}}$; one input $\hat\sigma/\sqrt{S_{xx}}$assumes independent, equal-variance errors
t, CI$t = \hat\beta_j/SE_j$; $\hat\beta_j \pm t_{n-p,0.975}SE_j$rescaling x leaves t unchanged
$R^2$$1 - RSS/TSS$not accuracy, not causality
Robust SE$(X^\top X)^{-1}(\sum e_i^2 x_ix_i^\top)(X^\top X)^{-1}$cov_type='HC3'; not for correlated errors
VIF$1/(1 - R_j^2)$SE grows by $\sqrt{VIF}$; 5–10 rule of thumb
Forecast design$[1, t, (t-s_j)_+, \sin, \cos, \text{holidays}, X_t]$locations $s_j$ fixed; weights δ, β estimated
Code it · Python

import numpy as np
import statsmodels.api as sm
from statsmodels.stats.outliers_influence import variance_inflation_factor
np.set_printoptions(suppress=True)

# 1) The cafe data: OLS with statsmodels
x = np.array([10, 15, 20, 25, 30.0])          # temperature (deg C)
y = np.array([30, 50, 40, 70, 60.0])          # cold drinks sold
X = sm.add_constant(x)                        # column of 1s + temperature
fit = sm.OLS(y, X).fit()
print(fit.params, fit.bse.round(3))           # [18.  1.6] [14.697  0.693]
print(fit.tvalues.round(3), fit.pvalues.round(3))   # [1.225 2.309] [0.308 0.104]
print(fit.conf_int().round(2))                # slope CI [-0.6, 3.8]
print(fit.resid, round(fit.ssr, 6), round(fit.rsquared, 2))   # [ -4.  8. -10.  12.  -6.] 360.0 0.64

# 2) Normal equations vs lstsq (prefer lstsq / QR in real code)
beta_ne = np.linalg.solve(X.T @ X, X.T @ y)
beta_ls = np.linalg.lstsq(X, y, rcond=None)[0]
print(beta_ne, beta_ls, (X.T @ fit.resid).round(10))   # [18. 1.6] [18. 1.6] [0. 0.]  (X^T e = 0)

# 3) Heteroscedasticity: usual vs robust (HC3) standard errors
rng = np.random.default_rng(0)
xh = rng.uniform(1, 10, 200)
yh = 2 + xh + rng.normal(0, 0.1 * xh**2)      # noise sd grows like x^2 (a funnel)
m_usual = sm.OLS(yh, sm.add_constant(xh)).fit()
m_rob = sm.OLS(yh, sm.add_constant(xh)).fit(cov_type="HC3")
print(m_usual.params.round(3), m_usual.bse.round(3), m_rob.bse.round(3))
# [3.019 0.717] [0.929 0.144] [0.604 0.158]: same line, different error bars (slope SE 0.144 vs 0.158)

# 4) Multicollinearity: two channels that move together
x1 = np.array([1, 2, 3, 4, 5.0]); x2 = np.array([1, 2, 3, 4, 6.0])
Xc = sm.add_constant(np.column_stack([x1, x2]))
for y5 in [10, 11, 9]:
    yc = np.array([2, 4, 6, 8, y5])
    print(y5, sm.OLS(yc, Xc).fit().params.round(3))   # (0,2,0) -> (0,1,1) -> (0,3,-1)
print([round(float(variance_inflation_factor(Xc, j)), 1) for j in (1, 2)])   # [37.0, 37.0]

# 5) A forecasting-style design matrix: trend + changepoint ramp + weekly Fourier + holiday + regressor
rng = np.random.default_rng(2)
n = 112; t = np.arange(n)
temp = np.zeros(n)
for i in range(1, n): temp[i] = 0.85 * temp[i - 1] + rng.normal(0, 0.55)
holiday = np.isin(t, [24, 80]).astype(float)
def fourier(t, period, order):
    return np.column_stack([f(2 * np.pi * k * t / period) for k in range(1, order + 1) for f in (np.sin, np.cos)])
y_t = (100 + 0.4 * t + 0.8 * np.maximum(0, t - 56) + 6 * np.sin(2 * np.pi * t / 7) + 4 * np.cos(2 * np.pi * t / 7)
       + 3 * np.sin(4 * np.pi * t / 7) + 25 * holiday + 5 * temp + rng.normal(0, 4, n))
Xd = np.column_stack([np.ones(n), t, np.maximum(0, t - 56), fourier(t, 7, 3), holiday, temp])
print(Xd.shape)                               # (112, 11): one weight per column
fd = sm.OLS(y_t, Xd).fit()
print(fd.params.round(2))                     # truth: 100, 0.4, 0.8, 6, 4, 3, 0, 0, 0, 25, 5
X4 = np.column_stack([Xd, fourier(t, 7, 4)[:, 6:]])   # add the order-4 sin/cos columns
print(X4.shape[1], np.linalg.matrix_rank(X4))  # 13 columns but rank 11: harmonic 4 = harmonic 3 in disguise
Test yourself

1. Which of these is not a linear regression model?

"Linear" means linear in the coefficients. In $\beta_0 e^{\beta_1 x}$ the coefficient $\beta_1$ sits inside an exponential, so no design matrix can express it as $X\beta$. The others are linear in β with transformed columns.

2. You re-measure temperature in tenths of a degree instead of degrees (all values × 10). What happens to the slope's t-value?

The slope and its standard error are both divided by 10, so their ratio (the t-value), the p-value and all predictions stay the same. Only the units of the coefficient change.

3. The residual-vs-fitted plot is a funnel that widens to the right. With a correct mean model, what is the main consequence?

Heteroscedasticity breaks the constant-variance assumption, which the usual SE formula needs. OLS stays unbiased; use robust (HC) SEs, WLS, a transformation, or a likelihood whose variance depends on the mean.

4. Two regressors have correlation 0.9 and there are no other inputs. Their VIF is about…

$VIF = 1/(1 - r^2) = 1/(1 - 0.81) = 1/0.19 \approx 5.26$. Each coefficient's SE is about $\sqrt{5.26} \approx 2.3$ times larger than with uncorrelated inputs.

5. A daily demand regression shows residuals that stay positive for a week, then negative for a week, and so on. Which assumption is broken, and what is the typical effect?

Runs of same-sign residuals over time mean autocorrelation. Correlated days carry less information than independent ones, so the usual SEs (and HC SEs, which also assume independence) are too small. Use HAC SEs or model the dependence.

6. What do the normal equations $X^\top X\hat\beta = X^\top y$ say geometrically?

Rewrite them as $X^\top(y - X\hat\beta) = X^\top e = 0$: every column has zero dot product with the residuals. $\hat y$ is the orthogonal projection of $y$ onto the column space; "normal" means perpendicular.

Practice problems

A. Data $x = 1, 2, 3, 4$ and $y = 2, 3, 5, 6$. Find the least-squares line, the residuals, $\hat\sigma^2$, the slope's SE and its t-value.
  1. $\bar x = 2.5$, $\bar y = 4$. Deviations: $x$: $-1.5, -0.5, 0.5, 1.5$; $y$: $-2, -1, 1, 2$.
  2. $S_{xy} = 3 + 0.5 + 0.5 + 3 = 7$; $S_{xx} = 2.25 + 0.25 + 0.25 + 2.25 = 5$. Slope $= 7/5 = 1.4$; intercept $= 4 - 1.4\times2.5 = 0.5$.
  3. Fitted: 1.9, 3.3, 4.7, 6.1. Residuals: $0.1, -0.3, 0.3, -0.1$ (sum 0). $RSS = 0.01 + 0.09 + 0.09 + 0.01 = 0.2$.
  4. $\hat\sigma^2 = 0.2/(4 - 2) = 0.1$; $SE(\hat\beta_1) = \sqrt{0.1/5} = \sqrt{0.02} \approx 0.141$; $t = 1.4/0.141 \approx 9.9$ on 2 df ($p \approx 0.010$).
B. Show that with an intercept column, the OLS residuals add up to zero and the fitted line passes through $(\bar x, \bar y)$.

The normal equations say $X^\top e = 0$. The row of $X^\top$ that comes from the 1s column gives $\sum_i 1\cdot e_i = 0$. Then $\sum y_i = \sum \hat y_i$, so $\bar y = \bar{\hat y} = \hat\beta_0 + \hat\beta_1\bar x$ (for one input): the line passes through $(\bar x, \bar y)$. Without an intercept column neither property is guaranteed.

C. (Interview) "Your regression residuals fan out as the fitted value grows. What does that do to your estimates, and what would you do?"

"If the mean is correctly specified, the coefficients are still unbiased, but the usual standard errors assume constant variance, so p-values and intervals are off, and OLS is no longer the most efficient estimator. Spread growing with the level suggests multiplicative noise, so I would first try modelling $\log y$. If interpretation on the original scale matters, I'd keep OLS with HC3 robust SEs, or use WLS if I can model the variance. For counts I'd move to a Poisson or Negative Binomial GLM, whose variance grows with the mean by design."

D. In an A/B regression, adding a pre-experiment covariate lowers the residual sd $\hat\sigma$ from 10 to 6. By how much does the treatment coefficient's SE fall, and what sample-size increase would have done the same?

The SE is proportional to $\hat\sigma$ (the treatment column's spread is unchanged), so it falls by the factor $6/10 = 0.6$, a 40% reduction. Since SE $\propto 1/\sqrt n$, the same reduction without the covariate would need $(10/6)^2 \approx 2.8$ times as many users. This is the logic of regression adjustment and CUPED (Chapter 5.12).

E. (Interview) "Explain your forecasting model as a regression. Where are the parameters, and what is not estimated?"

"The mean is $X\beta$ with column groups: a constant and $t$ for the base trend; one hinge column $(t - s_j)_+$ per candidate changepoint, whose weight $\delta_j$ is the slope change there; $2N$ sine/cosine columns per seasonality; 0/1 holiday columns; and external regressors. The weights are the parameters, with a Laplace prior on the $\delta_j$ and priors on the rest; σ (and ν or the NB concentration) belong to the likelihood. The changepoint locations, the periods and the Fourier orders are chosen before fitting (grid plus PELT, defaults), so they are not estimated by the model and their uncertainty is not in the posterior."

F. Regressing a candidate regressor $x_j$ on all the other columns gives $R_j^2 = 0.8$. What are its VIF and its SE inflation, and when should you care?

$VIF = 1/(1 - 0.8) = 5$, so its SE is $\sqrt5 \approx 2.24$ times what it would be if $x_j$ were uncorrelated with the others. Care if you want to interpret or test $\beta_j$ itself (its interval is more than twice as wide); care less if you only want predictions in the range of the training data. Options: combine it with its partners, regularize (priors), or report a joint effect.

Chapter 5.14 · Syllabus Module 19

Generalized linear models

Linear regression assumes the outcome can be any number and the noise is the same size everywhere. Conversions (0 or 1) and order counts (0, 1, 2, …) break both promises. Generalized linear models keep the best part of regression, a weighted sum of columns, and add two things: the right distribution for the data and a link that keeps the mean in its allowed range. Logistic, Poisson and Negative Binomial regression are all the same idea, and so are the likelihood choices in your forecasting model.

  • Explain why a straight line fails for yes/no and count outcomes (range of the mean, variance that depends on the mean, multiplicative effects)
  • Name the three parts of a GLM, $g(E[Y]) = X\beta$: the random component (likelihood), the linear predictor $\eta = X\beta$ and the link function $g$
  • Use the common links (identity, logit, log; probit, cloglog, softplus for awareness) and read what a +1 change in $\eta$ does to the mean
  • Fit and read logistic regression: log-odds, odds ratios, and why an odds ratio is not a relative lift
  • Fit and read Poisson regression: rate ratios and exposure offsets
  • Detect overdispersion, see why it makes Poisson standard errors lie, and fix it with Negative Binomial regression (knowing every library's parameterization)
  • Know how GLMs are fitted (maximum likelihood by IRLS) and read deviance; map all of it onto the likelihoods and links of your two projects

What we need from earlier chapters: linear regression and the design matrix (Chapter 5.13); likelihood, MLE, curvature-based standard errors and "loss = negative log-likelihood" (Chapter 5.2); the Bernoulli, Binomial, Poisson and Negative Binomial distributions (Chapters 4.7–4.8); the two-proportion z-test on the checkout example (Chapter 5.6). Notation: $\mu_i = E[y_i\mid x_i]$ is the mean of observation $i$; $\eta_i = x_i^\top\beta$ is its linear predictor ("eta"); $g$ is the link, $g^{-1}$ its inverse; $\sigma(z) = 1/(1 + e^{-z})$ is the logistic (sigmoid) function, not a standard deviation, in this chapter's logistic sections; "log" is the natural log.

Why a straight line fails for yes/no and count data

Try to predict whether a user converts (1) or not (0) from how many sessions they had, using an ordinary straight line. The line keeps rising forever, so for a heavy user it predicts a conversion "probability" of 1.15, and for a brand-new user it might predict −0.05. Neither is a probability. The same happens with counts: a line for "orders per day" against price eventually predicts −3 orders.

There is a second, quieter problem. The noise is not the same size everywhere. A conversion rate near 50% is very uncertain from user to user; a rate near 2% hardly varies (almost everyone is a 0). A store averaging 50 orders a day swings much more than one averaging 2. Linear regression assumes one constant noise level, so its error bars are wrong too.

Three ways to say it:

  • Picture: a straight line escapes the allowed band (0 to 1, or above 0); an S-curve or an exponential curve stays inside.
  • Numbers: the line $-0.05 + 0.04 \times \text{sessions}$ predicts $1.15$ at 30 sessions; the Bernoulli sd is 0.50 at $p = 0.5$ but only 0.14 at $p = 0.02$.
  • Slogan: bounded data need a bounded mean and a noise level that follows the mean.
  1. Out of range. Suppose a straight-line fit of conversion on sessions is $\hat p = -0.05 + 0.04\,s$. At $s = 0$: $\hat p = -0.05$ (a negative probability). At $s = 30$: $-0.05 + 1.2 = 1.15$ (above 1).
  2. Variance depends on the mean (yes/no). A 0/1 outcome with success probability $p$ has variance $p(1-p)$: at $p = 0.5$ it is $0.25$ (sd 0.50); at $p = 0.02$ it is $0.02\times0.98 = 0.0196$ (sd 0.14).
  3. Variance depends on the mean (counts). A Poisson count has variance equal to its mean: mean 2 → sd $\sqrt2 \approx 1.41$; mean 50 → sd $\sqrt{50} \approx 7.07$.
  4. Effects are often multiplicative. A promotion that lifts a small store from 20 to 30 orders (+10) would lift a big store from 200 to 300 (+100): the same 50% lift, very different additive effects. A straight line on the raw scale must pick one additive number.

Ordinary linear regression models $E[y\mid x] = x^\top\beta$ with constant-variance noise. For many outcome types this breaks in three ways:

  1. Range of the mean: probabilities must lie in $[0, 1]$, counts and rates must be $\ge 0$, but $x^\top\beta$ can be any real number.
  2. Mean–variance link: for 0/1 data $Var = \mu(1-\mu)$, for Poisson counts $Var = \mu$, for overdispersed counts $Var = \mu + \mu^2/\alpha$. The noise level is a function of the mean, not a constant.
  3. Scale of effects: many effects act by multiplying the mean (percent lifts) or the odds, not by adding a fixed amount.

A generalized linear model fixes all three at once: model a transformed mean as the linear part, and use the distribution that matches the data type (next section). The linear probability model (OLS on 0/1 data) is not useless: for a single randomized 0/1 treatment column its slope is exactly the difference in conversion rates. It goes wrong when you predict across a continuous $x$, and its default SEs ignore the changing variance.

Why do we need it?

Most product metrics are not "any real number with constant noise": conversions, clicks, sign-ups, orders, tickets, page views. Using the wrong model gives impossible predictions and miscalibrated uncertainty, exactly where decisions are made.

Where is it used?

Conversion modelling (logistic regression), click-through prediction, count metrics in A/B tests (Poisson), demand forecasting with count likelihoods (Negative Binomial), insurance claims, epidemiology rates, and the likelihood choice in both of your NumPyro projects.

How is it used?

Before fitting, ask two questions about $y$: what values can it take (support), and how does its spread change with its level (variance structure)? The answers choose the distribution and the link; then fit with sm.GLM(y, X, family=...) or a NumPyro model.

Choose a dataset. Blue dots are the data (conversions are drawn at 0 and 1 with a little vertical jitter). The purple line is ordinary least squares; the red zone is impossible values. Notice where the line enters the red zone, especially when you extend it beyond the data. Tick show GLM fit: the orange logistic (or Poisson) curve bends to stay inside the allowed range and flattens or grows exactly where the data say so.

"Just clip the straight line's predictions to [0, 1] (or to ≥ 0)."

Clipping hides the symptom but keeps the wrong shape (effects that should flatten near 0 and 1 keep growing) and the wrong noise model (constant variance). A GLM fixes the shape and the variance together.

"OLS on 0/1 outcomes is always wrong."

For a randomized A/B comparison with only a treatment dummy, the OLS slope is exactly $\hat p_B - \hat p_A$, and with robust SEs the test is fine. The problems appear when you model a continuous input or predict individual probabilities.

Straight lines fail for bounded data: the mean leaves its range; the variance depends on the mean ($p(1-p)$, $\mu$, $\mu + \mu^2/\alpha$); effects are often multiplicative.

Choose a model from the data's support (what values) and variance structure (how spread grows).

Trap: clipping predictions does not fix shape or uncertainty.

Quick check: a count metric has mean 4 on quiet days and 100 on busy days. Under a Poisson model, how do the standard deviations compare?

Poisson sd = $\sqrt{\text{mean}}$: $\sqrt4 = 2$ on quiet days and $\sqrt{100} = 10$ on busy days, five times larger. A constant-variance model would pretend both are equally noisy.

The three parts of a GLM: likelihood, linear predictor, link core

Think of a GLM as three dials. Dial 1, the linear predictor: exactly like regression, a weighted sum of the columns, $\eta = \beta_0 + \beta_1x_1 + \dots$. It can be any number, positive or negative. Dial 2, the link: a translator that turns that unrestricted number into a valid mean: a probability between 0 and 1, or a positive rate. Dial 3, the random component: the distribution that produces the actual data around that mean: Bernoulli for yes/no, Poisson or Negative Binomial for counts, Normal for continuous values.

Linear regression is the special case where the translator does nothing (identity link) and the distribution is Normal. Change the other two dials and you get logistic regression, Poisson regression and Negative Binomial regression, all fitted with the same machinery.

Three ways to say it:

  • Picture: features → weighted sum $\eta$ → link translates $\eta$ to a mean $\mu$ → the distribution scatters data around $\mu$.
  • Numbers: promo day: $\eta = 2.996 + 0.405 = 3.401$, $\mu = e^{3.401} = 30$ orders, and the day's count is a Poisson(30) draw.
  • Slogan: GLM = regression on the right scale, with the right noise.

Orders on promotion days (Poisson regression with a 0/1 promo column). Fitted weights: $\hat\beta_0 = 2.996$, $\hat\beta_1 = 0.405$.

  1. Normal day (promo = 0): linear predictor $\eta = 2.996 + 0.405\times0 = 2.996$.
  2. Link: the log link says $\log\mu = \eta$, so $\mu = e^{2.996} \approx 20.0$ orders.
  3. Promo day (promo = 1): $\eta = 2.996 + 0.405 = 3.401$, $\mu = e^{3.401} \approx 30.0$ orders.
  4. Random component: each day's count is a Poisson draw with that mean, so a normal day has sd $\sqrt{20} \approx 4.5$ and a promo day $\sqrt{30} \approx 5.5$.
  5. The promo multiplies the mean by $e^{0.405} = 1.5$: a 50% lift. On the $\eta$ scale effects add; on the $\mu$ scale they multiply.

A generalized linear model has three parts:

  1. Random component (the likelihood): $y_i\mid x_i$ follows a distribution with mean $\mu_i$, usually from the exponential family (Normal, Bernoulli/Binomial, Poisson, Gamma; Negative Binomial with a fixed dispersion). Its variance is a function of the mean: $Var(y_i) = \phi\,V(\mu_i)$, with variance function $V$ and dispersion $\phi$.
  2. Linear predictor: $\eta_i = x_i^\top\beta$, a row of the design matrix times the weights (Chapter 5.13).
  3. Link function $g$: $g(\mu_i) = \eta_i$, equivalently $\mu_i = g^{-1}(\eta_i)$. It maps the allowed range of the mean onto the whole real line.
$$g\big(E[y_i\mid x_i]\big) = x_i^\top\beta .$$
ModelRandom componentUsual link $g(\mu)$$Var(y)$Coefficient means
Linear regressionNormal$(\mu, \sigma^2)$identity: $\mu$$\sigma^2$ (constant)additive change in the mean
Logistic regressionBernoulli$(\mu)$ / Binomial$(n, \mu)$logit: $\log\frac{\mu}{1-\mu}$$\mu(1-\mu)$ (per trial)$e^\beta$ = odds ratio
Poisson regressionPoisson$(\mu)$log: $\log\mu$$\mu$$e^\beta$ = rate ratio
Negative Binomial regressionNB2$(\mu, \alpha)$log: $\log\mu$$\mu + \mu^2/\alpha$$e^\beta$ = rate ratio

Linear regression as a GLM: with a Normal random component and the identity link, the maximum likelihood estimate is the OLS estimate, and the deviance (below) is the RSS (statsmodels reports the RSS itself; divided by σ² it is the "scaled" deviance). Everything in Chapter 5.13 is the simplest GLM.

Why do we need it?

It gives one recipe for many data types: pick the distribution from the data, pick a link that keeps the mean valid, and reuse everything from regression (design matrices, coefficients, standard errors, tests).

Where is it used?

statsmodels.api.GLM, R's glm, scikit-learn's LogisticRegression and PoissonRegressor, the output layer of neural networks (sigmoid + cross-entropy is a logistic GLM), and every NumPyro model of the form "mean from a linear predictor, then a likelihood".

How is it used?

Write the three parts explicitly: eta = X @ beta; mu = inverse_link(eta); y ~ Distribution(mu, …). In statsmodels: sm.GLM(y, X, family=sm.families.Poisson()).fit() (the family carries its default link).

features x a row of X (promo, price…) linear predictor η = Xβ −∞+∞ inverse link μ = g⁻¹(η) random component y ~ Dist(μ, …) any real number logit⁻¹ → (0,1); exp → (0,∞) Bernoulli, Poisson, NB, Normal Example: η = 2.996 + 0.405·promo → μ = e^η = 20 or 30 → y ~ Poisson(μ) Linear regression = identity link + Normal; logistic = logit + Bernoulli; Poisson / NB = log + count distribution
The GLM pipeline. The linear predictor is ordinary regression and may take any value; the inverse link turns it into a valid mean; the random component scatters the observed data around that mean with the right kind of noise.

Pick a family. The orange curve is the mean $\mu(x) = g^{-1}(\beta_0 + \beta_1 x)$; at five values of $x$ the sideways shapes show the full distribution of $y$ there (bars for counts and yes/no, a bell for Normal); faint blue dots are simulated data. Notice: for Normal every slice has the same width; for Poisson the slices widen as the mean grows; NB widens even faster (lower the concentration α); for Bernoulli the two bars trade height and the curve can never leave 0–1.

"The log link means we take the log of the data and fit a line."

The link transforms the mean, not the data: $\log E[y] = X\beta$. Regressing $\log y$ models $E[\log y]$, a different quantity ($E[\log y] \ne \log E[y]$), cannot handle $y = 0$, and gives back-transformed predictions that estimate a median rather than a mean (Chapter 4.18).

"A GLM assumes Normal errors around the curve."

The random component is whatever distribution you choose: Bernoulli, Poisson, NB… The "errors" are not added Normal noise; the whole distribution of $y$ is centred on $\mu(x)$ with its own variance function.

"A GLM is linear regression on transformed data."

"A GLM models a transformation of the expected value, $g(E[y\mid x]) = x^\top\beta$, together with a distribution for $y$ whose variance depends on the mean. The data themselves are not transformed."

Model answer: "Three parts: a random component from the exponential family, a linear predictor $X\beta$, and a link connecting the mean to the linear predictor. Linear regression is the Normal-identity case; logistic is Bernoulli-logit; Poisson regression is Poisson-log. Because the link acts on the mean, zeros are fine and predictions are means on the original scale."

$g(E[y\mid x]) = x^\top\beta$: random component (likelihood) + linear predictor $\eta = X\beta$ + link $g$ ($\mu = g^{-1}(\eta)$).

Normal/identity (Var σ²), Bernoulli/logit (μ(1−μ)), Poisson/log (μ), NB/log (μ + μ²/α).

Trap: the link transforms the mean, not the data; $\log E[y] \ne E[\log y]$.

Quick check: in a logistic regression, $\eta = -1$ for some user. What is the predicted conversion probability?

$\mu = 1/(1 + e^{1}) = 1/(1 + 2.718) \approx 0.269$. The linear predictor lives on the log-odds scale; the inverse logit turns it into a probability.

Logistic regression: a straight line for the log-odds core

We want the probability of a yes (a conversion, a click) to depend on some features. Probabilities are stuck between 0 and 1, so instead we let a straight line describe the log-odds, a stretched version of the probability that can be any number, and then squash it back with the S-shaped curve. Low log-odds → probability near 0; high log-odds → near 1; log-odds 0 → exactly 50%.

The fitting rule is maximum likelihood: choose the line so that the users who converted get high predicted probabilities and those who did not get low ones. The "badness" being minimized is exactly the cross-entropy (log loss) used to train classifiers.

Three ways to say it:

  • Picture: an S-curve through a band of 0s at the bottom and 1s at the top.
  • Numbers: the checkout test as a logistic regression: $\hat\beta_0 = -2.197$ (10% control), $\hat\beta_1 = 0.205$ (odds × 1.23), $p \approx 0.31$, the same answer as the z-test.
  • Slogan: logistic regression = linear regression for the log-odds, fitted by Bernoulli likelihood.

The checkout A/B test as a logistic regression (Chapter 5.6): control 50 of 500 converted, treatment 60 of 500. Columns: intercept and a 0/1 treatment dummy $d$.

  1. With only a dummy, the fitted probabilities equal the group rates: control $\hat p_A = 0.10$, treatment $\hat p_B = 0.12$.
  2. Intercept = control log-odds: $\hat\beta_0 = \log(0.10/0.90) = \log(50/450) \approx -2.197$.
  3. Slope = difference in log-odds: $\hat\beta_1 = \log(60/440) - \log(50/450) = -1.992 - (-2.197) = 0.205$.
  4. Odds ratio $e^{0.205} = (60/440)/(50/450) = 0.1364/0.1111 \approx 1.227$.
  5. Standard error of a log odds ratio from a 2×2 table: $\sqrt{1/50 + 1/450 + 1/60 + 1/440} = \sqrt{0.0412} \approx 0.203$ (statsmodels gives the same 0.2029).
  6. Wald $z = 0.205/0.203 \approx 1.01$, two-sided $p \approx 0.31$; 95% CI for the odds ratio $e^{0.205 \pm 1.96\times0.203} = [0.82, 1.83]$. Same conclusion as the two-proportion z-test ($z \approx 1.01$, $p \approx 0.31$): not distinguishable from luck.

Logistic regression: $y_i \sim \text{Bernoulli}(p_i)$ (or Binomial$(n_i, p_i)$ for grouped counts), with

$$\log\frac{p_i}{1 - p_i} = x_i^\top\beta \qquad\Longleftrightarrow\qquad p_i = \sigma(x_i^\top\beta) = \frac{1}{1 + e^{-x_i^\top\beta}} .$$
  • Log-likelihood: $\ell(\beta) = \sum_i \big[y_i\log p_i + (1 - y_i)\log(1 - p_i)\big]$. Its negative is the binary cross-entropy (log loss) (Chapter 5.2).
  • There is no closed-form solution; the log-likelihood is concave, so Newton's method / IRLS finds the maximum quickly (it is unique unless the data are separated, see below) (below).
  • Interpretation: $\beta_j$ = change in log-odds per unit of $x_j$; $e^{\beta_j}$ = odds ratio per unit (next section). The intercept is the log-odds when all columns are 0.
  • Separation: if some combination of features perfectly separates 0s from 1s, the MLE does not exist (coefficients run off to ±∞). Regularization or a prior fixes it.
Why do we need it?

It is the standard model for "will it happen?" outcomes that depend on features, giving valid probabilities, interpretable odds ratios and standard errors, and it reduces to the classic two-proportion test when the only feature is a treatment flag.

Where is it used?

Conversion and churn models, click-through prediction, A/B analysis with covariates or segments, propensity scores in causal inference (Chapter 5.12), credit scoring, the last layer of binary neural classifiers, and Bayesian hierarchical conversion models.

How is it used?

smf.logit('converted ~ treat + sessions', df).fit() or sm.GLM(y, X, family=sm.families.Binomial()); read coefficients as log-odds, exponentiate for odds ratios, and convert to probabilities at specific inputs with .predict().

Blue dots: 80 users (0 = did not convert, 1 = converted; jittered a little), plotted against their number of sessions. Purple dots: the observed conversion rate at each session count. Move the sliders to shape the orange S-curve and watch the log-likelihood: higher (closer to 0) is better. Try to beat it, then press Fit by maximum likelihood. Notice how the odds ratio per session, $e^{\beta_1}$, stays the same along the whole curve while the probability step per session does not.

"Logistic regression predicts a class (0 or 1)."

It predicts a probability. Turning it into a class needs a threshold, which is a separate business decision (costs of false positives vs false negatives), not part of the model.

"The software reports enormous coefficients and SEs; the effect must be huge."

That is the classic sign of (quasi-)separation: some feature predicts the outcome perfectly in your sample, so the MLE runs away. Use a prior or penalty (ridge-like), or more data.

Your A/B framework models conversions with a Beta-Binomial: a Beta prior directly on each variant's rate (Chapter 6.3). Logistic regression is the regression cousin: put the treatment flag, segment columns and covariates in $X$, a prior on β, and you get a Bayesian logistic regression; with segments as group effects it becomes a hierarchical logistic model (Chapter 6.5). Decisions such as $P(\theta_B \gt \theta_A\mid D)$ are then computed on the probability scale by pushing each posterior draw of β through the inverse logit. Note the scales differ: a prior on the log-odds is not the same as a uniform prior on the rate.

$\log\frac{p}{1-p} = x^\top\beta$; $p = \sigma(x^\top\beta)$; $\ell = \sum[y\log p + (1-y)\log(1-p)]$ = −cross-entropy.

Treatment-only model: $\hat\beta_0$ = control log-odds, $\hat\beta_1$ = log odds ratio, SE $= \sqrt{1/a + 1/b + 1/c + 1/d}$; checkout: OR 1.23, $p \approx 0.31$.

Trap: coefficients are log-odds, not probability points; separation → infinite MLE.

Quick check: a logistic model has $\hat\beta_0 = -2$ and $\hat\beta_1 = 0.5$ per extra email opened. What is the predicted conversion probability for a user who opened 4 emails?

$\eta = -2 + 0.5\times4 = 0$, so $p = \sigma(0) = 0.5$. Each extra email multiplies the odds by $e^{0.5} \approx 1.65$.

Odds, odds ratios, and why an odds ratio is not a lift

Probability says "out of everyone, how many say yes". Odds say "for every one who says no, how many say yes". A 10% conversion rate is odds of 1 to 9 (0.111); a 50% rate is odds of 1 to 1 (1.0); a 90% rate is odds of 9 to 1 (9.0). Odds run from 0 to infinity, and their log runs over all numbers, which is why logistic regression works on log-odds.

Logistic coefficients become odds ratios when exponentiated, but product teams think in relative lift (risk ratio): "conversion went up 20%". For rare events the two are almost equal. For common events they drift apart, and quoting an odds ratio as a lift overstates the effect.

Three ways to say it:

  • Picture: the same two groups placed on three rulers: probability, odds, log-odds.
  • Numbers: 10% → 12% is a lift of 1.20 and an odds ratio of 1.23; 50% → 60% is also a lift of 1.20 but an odds ratio of 1.50.
  • Slogan: odds ratio ≈ risk ratio only when the event is rare.
  1. Checkout test: $p_A = 0.10$, $p_B = 0.12$. Odds: $0.10/0.90 = 0.111$ and $0.12/0.88 = 0.136$. Odds ratio $= 0.136/0.111 = 1.227$. Risk ratio (relative lift) $= 0.12/0.10 = 1.20$. Absolute difference $= 2$ points.
  2. Common outcome: $p_A = 0.50$, $p_B = 0.60$. Odds $1.0$ and $1.5$; odds ratio $1.50$; risk ratio $1.20$. Same 20% lift, odds ratio much larger.
  3. In between: $0.40 \to 0.48$ (lift 1.20): odds $0.667 \to 0.923$, odds ratio $\approx 1.385$.
  4. Rare outcome: $0.020 \to 0.024$ (lift 1.20): odds $0.0204 \to 0.0246$, odds ratio $\approx 1.204$, nearly the lift.
  • Odds of an event with probability $p$: $\text{odds} = p/(1-p)$; back: $p = \text{odds}/(1 + \text{odds})$. Log-odds = logit$(p)$.
  • Odds ratio $OR = \frac{p_B/(1-p_B)}{p_A/(1-p_A)}$. In logistic regression, $e^{\beta_j}$ is the odds ratio for +1 unit of $x_j$, others fixed.
  • Risk ratio (relative risk, relative lift) $RR = p_B/p_A$; risk difference $p_B - p_A$.
  • $OR = RR\cdot\frac{1-p_A}{1-p_B}$, so OR is always further from 1 than RR (when they are on the same side of 1), and $OR \approx RR$ when both probabilities are small.
  • OR has nice maths: it is symmetric (the OR for "not converting" is $1/OR$), it is what logistic regression estimates, and it can be estimated from case-control samples. RR and risk differences are easier to explain.
Why do we need it?

Logistic regression outputs odds ratios, but decisions are made on lifts and absolute differences. You need to move between the three scales correctly to avoid overstating effects.

Where is it used?

Reporting logistic-regression and GLM results, medical and epidemiology papers, A/B reports that mix "relative lift" and "odds", meta-analyses, and the log-odds parameterization of conversion models in NumPyro.

How is it used?

Exponentiate a logistic coefficient for the OR; for the lift or the difference, predict probabilities for the two scenarios (model.predict on both) and take their ratio or difference. Always state which scale a number is on.

Set the two conversion rates. The blue (A) and orange (B) markers show them on the probability ruler, the odds ruler and the log-odds ruler. Try the presets: for 2% → 2.4% the odds ratio and the lift agree; for 50% → 60% they disagree a lot. On the log-odds ruler the gap between the markers is the logistic coefficient $\beta_1$.

"The treatment's odds ratio is 1.5, so it increases conversion by 50%."

"An odds ratio of 1.5 means the odds are 1.5 times higher. The relative lift in conversion depends on the baseline: from 50% it is 50% → 60%, a 20% lift; only for rare outcomes is the lift close to 50%."

Model answer: "I'd report the odds ratio with its interval as the model output, and translate it into predicted conversion rates for control and treatment at a stated baseline, giving the absolute difference and the relative lift. Those are what the business decision needs."

odds $= p/(1-p)$; $OR = \frac{\text{odds}_B}{\text{odds}_A} = e^{\beta}$; $RR = p_B/p_A$; $OR = RR\,\frac{1-p_A}{1-p_B}$.

Rare events: OR ≈ RR. Common events: OR is further from 1 than RR.

Trap: never present an odds ratio as a relative lift.

Quick check: control converts at 30% and the odds ratio for treatment is 2. What is the treatment's conversion rate?

Control odds $= 0.3/0.7 \approx 0.429$; treatment odds $= 2\times0.429 = 0.857$; treatment rate $= 0.857/1.857 \approx 0.46$. So 30% → 46%: a risk ratio (relative lift) of about 1.54, that is +54%, not 2.

Poisson regression: counts and rate ratios core

Counts (orders per day, tickets per hour, clicks per user) are whole numbers, never negative, and usually more variable when they are bigger. The Poisson distribution has exactly these features: its variance equals its mean. Poisson regression lets the log of the expected count be a straight line in the features.

Because of the log, effects multiply. "Promotion" does not add 10 orders to every store; it multiplies each store's orders by, say, 1.5. That one number, the rate ratio, describes the effect for small and large stores alike, which is usually how business effects behave.

Three ways to say it:

  • Picture: on a log scale the expected count is a straight line; on the normal scale it is an exponential curve.
  • Numbers: $\hat\beta_1 = 0.405$ for the promo flag means $e^{0.405} = 1.5$: 50% more orders on promo days.
  • Slogan: Poisson regression: log of the mean is linear; coefficients exponentiate to rate ratios.

Orders on 5 normal days and 5 promo days. Normal: 14, 25, 20, 16, 25 (total 100, mean 20). Promo: 24, 36, 30, 25, 35 (total 150, mean 30).

  1. Model: $y_i \sim \text{Poisson}(\mu_i)$, $\log\mu_i = \beta_0 + \beta_1\,\text{promo}_i$.
  2. With one 0/1 column, the MLE fits each group's mean: $e^{\hat\beta_0} = 20$, so $\hat\beta_0 = \log 20 \approx 2.996$; $e^{\hat\beta_0 + \hat\beta_1} = 30$, so $\hat\beta_1 = \log(30/20) = \log 1.5 \approx 0.405$.
  3. Standard error: for a log rate ratio of two Poisson totals, $SE = \sqrt{1/100 + 1/150} = \sqrt{0.01667} \approx 0.129$.
  4. 95% CI for the rate ratio: $e^{0.405 \pm 1.96\times0.129} = e^{[0.152,\,0.658]} \approx [1.16, 1.93]$. The promo raises orders by between 16% and 93%.
  5. Check the Poisson assumption: variances 25.5 (normal) and 30.5 (promo) vs means 20 and 30, close. (Pearson dispersion $\approx 1.15$; values near 1 are what Poisson expects.)

Poisson regression: $y_i \sim \text{Poisson}(\mu_i)$ with $\log\mu_i = x_i^\top\beta$, so $\mu_i = e^{\beta_0}\cdot e^{\beta_1x_{i1}}\cdots$

  • Rate ratio: $e^{\beta_j}$ is the factor by which the expected count is multiplied per +1 unit of $x_j$, others fixed. $100(e^{\beta_j} - 1)\%$ is the percent change.
  • Log-likelihood: $\ell(\beta) = \sum_i\big[y_i\log\mu_i - \mu_i - \log y_i!\big]$ (Chapter 5.2). Concave in β; fitted by IRLS.
  • Equidispersion: Poisson assumes $Var(y_i\mid x_i) = \mu_i$. Real counts are often more variable (overdispersion, below), which leaves $\hat\beta$ reasonable but makes the Poisson SEs too small.
  • With an intercept and the canonical log link, the fitted means add up to the observed total: $\sum\hat\mu_i = \sum y_i$.
Why do we need it?

Counts are everywhere in product data, and their effects are usually proportional. Poisson regression gives non-negative predictions, a variance that grows with the mean, and coefficients that read as percent changes.

Where is it used?

Orders, sessions, support tickets and page views per unit; count metrics in A/B tests; event-rate models in epidemiology; the starting point before Negative Binomial demand models; scikit-learn's PoissonRegressor and gradient boosting with Poisson loss.

How is it used?

sm.GLM(y, X, family=sm.families.Poisson()).fit(); exponentiate coefficients and CI limits for rate ratios; then check dispersion (Pearson χ² / df). If it is well above 1, switch to Negative Binomial or quasi-Poisson SEs.

Fifty days of orders against ad spend. Shape the orange mean curve with the sliders and watch the Poisson log-likelihood; the shaded band is mean ± 2√mean, the spread Poisson expects. Press Fit for the maximum-likelihood curve. Tick log scale: the same curve becomes a straight line ($\log\mu = \beta_0 + \beta_1x$), and days with 0 orders appear on the bottom row (their log is −∞). The readout gives the rate ratio per unit of spend.

"$\hat\beta_1 = 0.405$ means the promo adds 0.405 orders."

It adds 0.405 to the log of the expected count, i.e. multiplies the expected count by $e^{0.405} = 1.5$. The absolute gain depends on the baseline (+10 at 20 orders, +100 at 200).

"Poisson regression needs the counts to be Poisson-distributed overall."

It assumes each count is Poisson given its features. The overall histogram can look like anything. What matters is the conditional variance: check dispersion after fitting.

$y \sim \text{Poisson}(\mu)$, $\log\mu = x^\top\beta$; $e^{\beta_j}$ = rate ratio per unit; effects multiply.

Two groups: $\hat\beta_1 = \log(\bar y_1/\bar y_0)$, $SE = \sqrt{1/\sum y_{(0)} + 1/\sum y_{(1)}}$; promo example RR 1.5, CI [1.16, 1.93].

Trap: assumes Var = mean; check Pearson χ²/df; coefficients are on the log scale.

Quick check: a Poisson model of daily orders has a temperature coefficient of 0.02 per °C. What does a 10 °C warmer day do to expected orders?

Multiply by $e^{0.02\times10} = e^{0.2} \approx 1.22$: about 22% more orders. (Not $10\times2\% = 20\%$ exactly, because effects compound multiplicatively.)

Exposure and offsets: compare rates, not raw counts

A shop open 12 hours collects more orders than one open 4 hours, even if it is less popular. Users observed for 30 days click more than users observed for 3. Counts depend on how long (or how much) you watched: the exposure. What we really care about is the rate: orders per hour, clicks per day.

A Poisson regression handles this neatly. Expected count = exposure × rate, so on the log scale log(expected count) = log(exposure) + log(rate). The log(exposure) term is added to the linear predictor with its coefficient fixed at exactly 1. That fixed column is called an offset. The remaining coefficients then describe rates.

Three ways to say it:

  • Picture: two timelines catching events; the longer window catches more, even at the same rate.
  • Numbers: 120 orders in 40 hours (3 per hour) vs 90 orders in 20 hours (4.5 per hour): the region with fewer orders is 1.5 times busier.
  • Slogan: model counts, compare rates: put log(exposure) in as an offset.
  1. Region A: 120 orders in 40 store-hours. Region B: 90 orders in 20 store-hours.
  2. Naive count comparison: $90/120 = 0.75$. "B is 25% worse."
  3. Rates: A $= 120/40 = 3.0$ per hour; B $= 90/20 = 4.5$ per hour. Rate ratio $= 4.5/3.0 = 1.5$. "B is 50% better."
  4. Poisson model with offset: $\log\mu_i = \log(\text{hours}_i) + \beta_0 + \beta_1\,\text{B}_i$. The fit gives $\hat\beta_0 = \log 3 \approx 1.099$ (A's hourly rate) and $\hat\beta_1 = \log 1.5 \approx 0.405$ (rate ratio 1.5), with $SE(\hat\beta_1) = \sqrt{1/120 + 1/90} \approx 0.139$.
  5. Without the offset, the same model gives $e^{\hat\beta_1} = 0.75$: the exposure difference completely flips the conclusion.

If observation $i$ has exposure $t_i$ (time, users, area) and rate $\lambda_i$, then $y_i \sim \text{Poisson}(t_i\lambda_i)$ with $\log\lambda_i = x_i^\top\beta$, so

$$\log\mu_i = \underbrace{\log t_i}_{\text{offset (coefficient fixed at 1)}} + x_i^\top\beta .$$
  • The coefficients are now log rate ratios per unit of exposure.
  • statsmodels: sm.GLM(y, X, family=Poisson(), offset=np.log(t)) or exposure=t. NumPyro: mu = t * jnp.exp(X @ beta).
  • Putting $\log t_i$ in as an ordinary column instead estimates its coefficient (allowing counts to grow less or more than proportionally); the offset assumes proportionality.
  • Do not divide counts by exposure and run OLS: the rate of a short window is far noisier than that of a long one, and the Poisson model weights them correctly.
Why do we need it?

Units are almost never observed for the same amount of time or size. Without an offset, "more exposure" masquerades as "more effect", and comparisons can even flip direction.

Where is it used?

Events per user-day in A/B tests where users join at different times, orders per open hour, incidents per 1 000 sessions, insurance claims per policy-year, disease rates per population, partial weeks or months in aggregated forecasts.

How is it used?

Record the exposure for every row, pass offset=np.log(exposure) (or exposure=) to the Poisson/NB GLM, and interpret $e^{\beta}$ as a ratio of rates. Check that exposure is not itself affected by the treatment.

Region A: 40 hours, rate 3 / hour → 120 orders Region B: 20 hours, rate 4.5 / hour → 90 orders (window ends: half the exposure) Counts: B/A = 0.75. Rates: B/A = 1.5. The offset log(hours) makes the model compare rates.
Exposure. The longer window collects more events even at a lower rate. A Poisson model with offset $\log(\text{hours})$ estimates rates per hour, so the comparison is fair.

Region A's true rate is 3 orders per hour; set region B's true rate ratio and both regions' opening hours. The left panel shows the raw counts, the right panel the counts per hour. Give A many more hours than B: the raw counts say A is "better" even when B's rate is higher. The offset model's rate ratio (which equals the ratio of the per-hour rates) tracks the truth. Press New sample to see Poisson noise.

"Divide each count by its exposure and run an ordinary regression on the rates."

Rates from short windows are much noisier than rates from long ones, and OLS weights them equally. The Poisson (or NB) model with an offset weights each observation by how much information it carries.

"Any variable that differs between groups can be an offset."

An offset assumes the count is exactly proportional to it. Use it for true exposure (time, population, number of sessions shown); otherwise include the variable as an ordinary column and let its coefficient be estimated.

In an A/B framework like yours, a count metric such as "orders per user" is really "orders per user over the days that user was in the experiment". Users who enrolled late have less exposure; if enrolment timing differs between arms, raw counts are biased. A Poisson or NB model with $\log(\text{days exposed})$ as an offset compares rates. In forecasting, the same idea appears when aggregating to weeks or months with unequal lengths (a 28-day vs a 31-day month), or modelling orders per open hour when hours change.

$y_i \sim \text{Poisson}(t_i\lambda_i)$: $\log\mu_i = \log t_i + x_i^\top\beta$; $\log t_i$ is an offset (coefficient fixed at 1).

$e^{\beta}$ = ratio of rates per unit exposure. Example: 120/40 h vs 90/20 h → counts 0.75, rates 1.5.

Trap: never compare raw counts with unequal exposure; don't OLS the rates.

Quick check: users in arm A were observed for 10 days on average and in arm B for 14 days. Arm B has 30% more total clicks. Is B better?

Not necessarily. Per day, B's rate ratio is $1.30/(14/10) = 1.30/1.4 \approx 0.93$: B's users click about 7% less per day. Compare rates with a log(days) offset before deciding.

Overdispersion: when Poisson is too sure of itself core

The Poisson model makes a strong promise: the variance of a count equals its mean. That is true when every day is "the same kind of day" with events arriving at random. Real days are not the same: weather, a competitor's sale, a viral post, a payday. Each day has its own hidden busy-ness, so counts swing much more than Poisson allows. That is overdispersion: variance bigger than the mean.

The danger is quiet. The Poisson fit still finds a sensible mean curve, but it believes the data are much less noisy than they are, so its standard errors are too small and its p-values too small. It starts "finding" effects that are pure noise.

Three ways to say it:

  • Picture: the data spill far outside the narrow band Poisson expects.
  • Numbers: counts 2, 8, 0, 15, 5 have mean 6 but variance 34.5: almost 6 times what Poisson allows.
  • Slogan: overdispersed data + Poisson model = overconfident error bars.

Five days of sign-ups: 2, 8, 0, 15, 5.

  1. Mean $\bar y = 30/5 = 6$. Deviations: $-4, 2, -6, 9, -1$; squares: 16, 4, 36, 81, 1; sum 138.
  2. Sample variance $= 138/4 = 34.5$. Poisson would need variance ≈ 6.
  3. Pearson dispersion statistic: $\sum (y_i - \hat\mu_i)^2/\hat\mu_i = 138/6 = 23$, divided by the residual degrees of freedom $n - p = 5 - 1 = 4$: $\hat\phi = 23/4 = 5.75$. Poisson expects about 1.
  4. Consequence: Poisson SEs are about $\sqrt{5.75} \approx 2.4$ times too small. A "quasi-Poisson" fix multiplies them by 2.4.
  5. Simulation (statsmodels, 500 runs, $n = 100$, mean 5, a feature with no effect, NB counts with concentration $\alpha = 2$): Poisson regression called the useless feature "significant at 5%" in 27% of runs; NB regression in 4.4%, quasi-Poisson in 4.8%.

Overdispersion: $Var(y_i\mid x_i) \gt E[y_i\mid x_i]$ for count data (more generally, more variance than the model's variance function allows).

  • Causes: unobserved differences between units or days (each has its own rate, a mixture), clustering and contagion (one event triggers others), missing covariates, and excess zeros.
  • Diagnosis: Pearson statistic $\hat\phi = \frac{1}{n-p}\sum_i \frac{(y_i - \hat\mu_i)^2}{\hat\mu_i}$; under a correct Poisson model it is close to 1. Values clearly above 1 (a rule of thumb: above about 1.5 with decent $n$) signal overdispersion. A variance-vs-mean plot of binned data shows the same thing.
  • Consequences for Poisson regression: $\hat\beta$ is still consistent if the mean model is right, but SEs are too small by about $\sqrt{\hat\phi}$, so tests reject too often and intervals undercover.
  • Remedies: Negative Binomial regression (next section); quasi-Poisson (multiply SEs by $\sqrt{\hat\phi}$); robust sandwich SEs; random effects (one extra noise term per unit or day, Chapter 6.5).
  • Excess zeros are a related but different problem (more zeros than even an NB with the same mean predicts): zero-inflated or hurdle models (Chapter 4.8, 7.13).
Why do we need it?

Overdispersion is the norm for business counts. Missing it produces false discoveries in A/B tests on count metrics and forecast intervals that are far too narrow.

Where is it used?

Checking any Poisson GLM (statsmodels reports pearson_chi2), choosing between Poisson and NB likelihoods in your forecasting model, A/B tests on orders or sessions per user, and posterior predictive checks of variance (Chapter 6.8).

How is it used?

After a Poisson fit, compute fit.pearson_chi2 / fit.df_resid. If it is well above 1, refit with NB (sm.NegativeBinomial or smf.negativebinomial) or report quasi-Poisson / robust SEs; in a Bayesian model, use an NB likelihood and check the predictive variance.

Each experiment simulates 100 days whose counts do not depend on the feature $x$ at all (true $\beta_1 = 0$), with Negative Binomial noise of concentration α (small α = very overdispersed). Left: one sample, with the band Poisson expects (orange, mean ± 2√mean) and the NB band (purple). Press Run 300 experiments: the bars show how often each method wrongly declares $x$ "significant" at the 5% level. Lower α and watch the Poisson bar climb far above 5%; raise α toward 50 (nearly Poisson) and all three agree.

"Overdispersion biases the Poisson coefficients."

If the mean model is right, the coefficients are fine; the standard errors are too small. That is why the Poisson bar in the widget rises while the fitted line stays flat.

"Lots of zeros means I need a zero-inflated model."

An overdispersed NB with a small mean already produces many zeros. Compare the observed zero fraction with what a fitted NB predicts before adding zero inflation.

In an A/B framework like yours, a Poisson likelihood for a count metric (orders per user) is only safe if the counts are close to equidispersed; heavy users make them overdispersed, and the posterior for the treatment effect becomes too narrow, just like the Poisson bar above. Check the posterior predictive variance against the observed variance (Chapter 6.8), or use an NB / hierarchical model. In forecasting, this is precisely why a Negative Binomial likelihood is offered for count series (Chapter 7.13).

Overdispersion: $Var(y\mid x) \gt E[y\mid x]$. Check $\hat\phi = \frac{1}{n-p}\sum\frac{(y-\hat\mu)^2}{\hat\mu}$ (≈ 1 for Poisson).

Poisson SEs too small by ≈ $\sqrt{\hat\phi}$ → false positives (27% instead of 5% at α = 2, mean 5).

Fixes: NB, quasi-Poisson, robust SEs, random effects. Trap: coefficients OK, error bars not.

Quick check: a Poisson fit has Pearson χ² = 300 on 100 residual degrees of freedom. Its reported SE for a coefficient is 0.05. What is a more honest SE?

$\hat\phi = 300/100 = 3$; quasi-Poisson SE $= 0.05\times\sqrt3 \approx 0.087$. The honest interval is about 1.7 times wider.

Negative Binomial regression: counts with honest spread core

Keep the Poisson idea but admit that every day has its own hidden "busy-ness" multiplier: some days run at 70% of the usual rate, some at 150%. If those multipliers follow a Gamma distribution, the counts follow a Negative Binomial (NB) distribution (Chapter 4.8). Its mean is the same as Poisson's, but its variance has an extra term that grows like the mean squared.

NB regression uses exactly the same mean model as Poisson regression (a log link, coefficients that are rate ratios), plus one extra parameter, the concentration α, learned from the data. Large α: days are alike, NB ≈ Poisson. Small α: days differ a lot, wide spread, many zeros.

Three ways to say it:

  • Picture: the same mean curve as Poisson, with a much wider band around it.
  • Numbers: mean 6 with α = 2: variance $6 + 36/2 = 24$ (Poisson: 6); P(0) = 0.0625 (Poisson: 0.0025).
  • Slogan: NB regression = Poisson regression + a learned extra spread.
  1. Variance. NB2 with mean $\mu = 6$ and concentration $\alpha = 2$: $Var = \mu + \mu^2/\alpha = 6 + 36/2 = 24$, sd $\approx 4.9$. Poisson(6): variance 6, sd 2.45.
  2. Zeros. NB: $P(0) = \big(\tfrac{\alpha}{\alpha + \mu}\big)^\alpha = (2/8)^2 = 0.0625$. Poisson: $e^{-6} \approx 0.0025$, 25 times fewer zeros.
  3. Moment estimate of α from the sign-up data (mean 6, variance 34.5): solve $34.5 = 6 + 36/\alpha$, so $\alpha = 36/28.5 \approx 1.26$.
  4. Same coefficients, different SEs. In the widget below, Poisson and NB fits give nearly the same slope, but NB's SE is larger and its 90% band actually contains about 90% of the days.

NB2 regression: $y_i \sim \text{NB}(\mu_i, \alpha)$ with $\log\mu_i = x_i^\top\beta$ (optionally $+\log t_i$ offset), and

$$E[y_i] = \mu_i, \qquad Var(y_i) = \mu_i + \frac{\mu_i^2}{\alpha}.$$
  • Story: $y_i\mid\lambda_i \sim \text{Poisson}(\lambda_i)$ with $\lambda_i \sim \text{Gamma}(\text{shape } \alpha, \text{rate } \alpha/\mu_i)$: a Gamma-Poisson mixture. $\alpha \to \infty$ gives Poisson.
  • Coefficients: $e^{\beta_j}$ = rate ratios, exactly as in Poisson regression. Fit β and α together by maximum likelihood (statsmodels sm.NegativeBinomial / smf.negativebinomial), or β by IRLS for a fixed α (sm.GLM(..., family=sm.families.NegativeBinomial(alpha=...))).
  • Parameterizations (check which one your code uses):
    LibraryCallDispersion parameterVariance
    NumPyroNegativeBinomial2(mean=μ, concentration=α)α (large = less spread)$\mu + \mu^2/\alpha$
    statsmodelsNegativeBinomial, GLM family=NegativeBinomial(alpha=a)$a = 1/\alpha$ (large = more spread)$\mu + a\mu^2$
    SciPystats.nbinom(n, p)$n = \alpha$, $p = \alpha/(\alpha + \mu)$$n(1-p)/p^2$
    NumPyroNegativeBinomialProbs(total_count=α, probs=μ/(α+μ))total_count = α$\mu + \mu^2/\alpha$
  • The log link is not NB's canonical link, so the fitted means need not add up exactly to the observed total (unlike Poisson).
Why do we need it?

It keeps everything good about Poisson regression (positive means, rate ratios, offsets) while letting the data decide how noisy the counts are, so standard errors, tests and prediction intervals become honest.

Where is it used?

Demand and sales forecasting with count data (your model's NB likelihood), A/B tests on count metrics, RNA-seq (DESeq2), insurance claim counts, web traffic, and as the count likelihood in Stan/NumPyro models.

How is it used?

Fit smf.negativebinomial('y ~ x', df).fit(); read alpha (statsmodels' $1/\alpha$!), exponentiate coefficients for rate ratios, compare with Poisson by AIC or a likelihood-ratio test, and check predictive coverage. In NumPyro: numpyro.sample('y', dist.NegativeBinomial2(mu, conc), obs=y).

mean model μᵢ = exp(xᵢᵀβ) trend, season, promo… hidden day multiplier λᵢ ~ Gamma(α, α/μᵢ) mean μᵢ, spread set by α count yᵢ ~ Poisson(λᵢ) overall: yᵢ ~ NB(μᵢ, α) Var(y) = μ (Poisson part) + μ²/α (day-to-day differences). α → ∞: every day identical, back to Poisson.
Where the Negative Binomial comes from: each day draws its own rate around the model's mean (a Gamma with concentration α), then a Poisson count at that rate. The extra variance $\mu^2/\alpha$ is the day-to-day variation of the hidden rates.

Eighty days of counts that truly follow an NB regression ($\log\mu = 1 + 0.25x$) with concentration α. Both models are fitted by maximum likelihood. The two mean curves are almost identical, but compare the 90% prediction bands: orange (Poisson) is too narrow and misses many days; purple (NB) covers about 90%. Read the slope SEs and the estimated α. Push α up to 100: the data become nearly Poisson and the two models agree.

"statsmodels says alpha = 0.5, so it is the same as NumPyro concentration 0.5."

statsmodels' alpha multiplies $\mu^2$ in the variance ($\mu + a\mu^2$), so $a = 0.5$ means NumPyro concentration = 2. Mixing them up turns "mildly overdispersed" into "wildly overdispersed". Verify with a quick simulation of the mean and variance.

"Switching from Poisson to NB changes the effect estimates a lot."

Usually the mean curves are close (same link, same columns); NB mainly changes the SEs, intervals and predictive bands. Big changes in β hint at influential extreme days, since NB down-weights large counts relative to Poisson.

"We use a Negative Binomial with dispersion 2."

"We use an NB2 likelihood parameterized by its mean $\mu_t$ and a concentration $\alpha$, so $Var = \mu_t + \mu_t^2/\alpha$; in NumPyro that is NegativeBinomial2(mean, concentration), and $\alpha$ is learned."

Model answer: "Poisson forces variance = mean; our counts were overdispersed (Pearson dispersion well above 1), which would make intervals too narrow. NB adds a Gamma-distributed day-level rate, giving variance $\mu + \mu^2/\alpha$. Large α recovers Poisson. The mean model, a log link on a linear predictor, is unchanged, so coefficients are still rate ratios."

Your forecasting model's Negative Binomial likelihood is NB regression with a structured design matrix: $\mu_t$ comes from trend, seasonality, holidays and regressors through a positive link, and the concentration is a learned parameter with a prior. A fitted concentration that is large says the series is nearly Poisson; a small one says days differ a lot. Know the exact class your code calls (NegativeBinomial2 vs NegativeBinomialProbs/Logits), what the prior is placed on (α or 1/α), and check the predictive coverage like the bands above (Chapter 7.13, 7.16).

NB2 regression: $\log\mu = x^\top\beta$, $Var = \mu + \mu^2/\alpha$; Gamma-Poisson mixture; α → ∞ is Poisson.

NumPyro NegativeBinomial2(mean, concentration=α); statsmodels alpha $= 1/\alpha$; SciPy nbinom(n=α, p=α/(α+μ)).

Trap: parameterizations differ; NB changes SEs and bands more than the mean.

Quick check: a daily series has mean 50 and variance 300. Under NB2, what is the concentration α, and what is statsmodels' alpha?

$300 = 50 + 2500/\alpha$, so $\alpha = 2500/250 = 10$ (NumPyro concentration 10). statsmodels' alpha is $1/10 = 0.1$.

How GLMs are fitted: maximum likelihood by IRLS, and deviance

Linear regression has a one-step formula. GLMs do not: the S-curve and the exponential make the equations non-linear. The standard trick is to pretend the problem is a linear regression, solve it, and repeat. Around the current guess, the GLM looks locally like a weighted least-squares problem: each point gets a weight (how much information it carries at its current mean) and a "working" target (where the straight line in $\eta$ should go to fix its error). Solve that weighted regression, update the guess, recompute weights, and repeat. Three to eight rounds usually settle.

This is Newton's method on the log-likelihood for canonical links (Fisher scoring in general; Optimization 3.12), dressed as repeated regressions: iteratively reweighted least squares (IRLS). Its final weight matrix also gives the standard errors, through the curvature of the log-likelihood (Chapter 5.2).

Three ways to say it:

  • Picture: the S-curve jumps toward the data, then makes smaller and smaller corrections.
  • Numbers: 7 successes in 10, starting at log-odds 0: 0 → 0.800 → 0.8469 → 0.8473 = logit(0.7).
  • Slogan: a GLM fit is a short series of weighted linear regressions.

IRLS for an intercept-only logistic model, 7 conversions out of 10 (the answer must be logit(0.7) = 0.8473).

  1. Start $\beta = 0$: $p = \sigma(0) = 0.5$, weight $w = p(1-p) = 0.25$.
  2. Working response for each user: $z_i = \eta + (y_i - p)/w$. Its weighted average (all weights equal) is $0 + (\bar y - p)/w = (0.7 - 0.5)/0.25 = 0.8$. New $\beta = 0.8$.
  3. Now $p = \sigma(0.8) \approx 0.6900$, $w = 0.69\times0.31 \approx 0.2139$. New $\beta = 0.8 + (0.7 - 0.6900)/0.2139 \approx 0.8469$.
  4. Again: $p = \sigma(0.8469) \approx 0.6999$, new $\beta \approx 0.8473$. Converged: $\sigma(0.8473) = 0.700$ ✓.
  5. SE from the final weights: $\sqrt{1/(n\,w)} = \sqrt{1/(10\times0.21)} \approx 0.69$ on the log-odds scale.

Maximum likelihood for GLMs. The estimate $\hat\beta$ maximizes $\ell(\beta) = \sum_i\log p(y_i\mid\mu_i(\beta))$. For canonical links the score equations are

$$X^\top(y - \hat\mu) = 0,$$

the GLM version of "residuals perpendicular to every column" (so with an intercept, $\sum\hat\mu_i = \sum y_i$).

IRLS, repeated until β stops changing:

  1. $\eta = X\beta$, $\mu = g^{-1}(\eta)$.
  2. Weights $w_i = \dfrac{(d\mu_i/d\eta_i)^2}{Var(y_i)}$ (logit: $\mu_i(1-\mu_i)$; Poisson-log: $\mu_i$; NB-log: $\mu_i/(1 + \mu_i/\alpha)$).
  3. Working response $z_i = \eta_i + (y_i - \mu_i)\,\dfrac{d\eta_i}{d\mu_i}$.
  4. Weighted least squares: $\beta \leftarrow (X^\top W X)^{-1}X^\top W z$.

At convergence, $\widehat{Var}(\hat\beta) = (X^\top W X)^{-1}$ (times the dispersion $\phi$ for families that have one): Wald SEs and z-tests.

Deviance $D = 2\big[\ell(\text{saturated}) - \ell(\hat\beta)\big]$, where the saturated model fits every observation exactly. It plays the role of the RSS (for the Normal-identity model statsmodels' deviance is the RSS; the scaled deviance is RSS/σ²). Poisson: $D = 2\sum[y_i\log(y_i/\hat\mu_i) - (y_i - \hat\mu_i)]$. The drop in deviance between nested models is a likelihood-ratio test, approximately $\chi^2$ with (difference in parameters) degrees of freedom. AIC $= -2\ell + 2k$ compares non-nested models such as Poisson vs NB.

Why do we need it?

Knowing how the fit works explains its outputs (SEs from the weights, deviance as the GLM's RSS) and its failure modes (non-convergence under separation, huge SEs, warnings about perfect prediction).

Where is it used?

statsmodels' GLM.fit() and R's glm use IRLS; scikit-learn's GLMs use L-BFGS (by default) on the same likelihood; deviance and LR tests appear in every GLM summary; Newton-type steps also power Laplace approximations (Chapter 6.9).

How is it used?

Read fit.converged and the iteration count; compare nested models with fit_small.deviance - fit_big.deviance against a χ² distribution; compare Poisson and NB with .aic; treat convergence warnings as a signal to check separation or collinearity.

Start from β = (0, 0): a flat 50% line. Press Next step: each step solves one weighted least-squares problem. Dot sizes show the current weights $w_i = p_i(1-p_i)$: points where the curve is near 50% count most, points where it is near 0% or 100% count little. Watch the log-likelihood climb in the small chart and the changes shrink: by step 4 or 5 nothing moves. Back and Reset replay it.

"A small residual deviance proves the model is right."

Deviance measures fit relative to a perfect-fit model; use it to compare nested models (drops in deviance) and, for Poisson with large counts, as a rough dispersion check (deviance/df ≈ 1). It does not prove the link, the columns or the distribution are right; residual and predictive checks still matter.

"The fit did not converge, so increase the iteration limit."

Non-convergence in logistic regression usually means (quasi-)separation, and in Poisson/NB regression extreme collinearity or tiny counts. More iterations only push coefficients further toward infinity. Diagnose the cause; add a penalty or prior if needed.

MLE via IRLS: $w_i = (d\mu/d\eta)^2/Var(y_i)$, $z_i = \eta_i + (y_i - \mu_i)\,d\eta/d\mu$, $\beta \leftarrow (X^\top WX)^{-1}X^\top Wz$; SEs from $(X^\top WX)^{-1}$.

Deviance $D = 2[\ell_{sat} - \ell]$; nested comparison: $\Delta D \sim \chi^2_{\Delta p}$; non-nested: AIC.

Trap: non-convergence = separation/collinearity, not "too few iterations".

Quick check: in the promo example, the intercept-only Poisson model has deviance 19.35 on 9 df and the model with the promo flag has 9.28 on 8 df. Is the promo effect significant?

The drop is $19.35 - 9.28 = 10.07$ on 1 df. The 5% critical value of $\chi^2_1$ is 3.84, and $P(\chi^2_1 \ge 10.07) \approx 0.0015$. Yes: the likelihood-ratio test agrees with the Wald test ($z = 0.405/0.129 \approx 3.14$, $p \approx 0.0017$).

GLM ideas inside your two projects core

Both of your projects are GLMs with priors. Each one answers the same three questions: what kind of numbers is the data (that picks the likelihood), what drives the mean (that builds the linear predictor from columns), and how to keep the mean valid (that picks the link). The Bayesian part adds priors on the weights and replaces IRLS by SVI or MCMC, but the skeleton is this chapter.

The link also decides how components combine. With an identity link (Normal or Student-t likelihood) trend, seasonality and holidays add in demand units. With a log link (typical for the Negative Binomial) they multiply: the weekend effect is a percentage, so it is bigger in absolute terms when the level is higher. A softplus link sits in between: almost additive at high levels, but squeezes the mean to stay positive near zero.

Three ways to say it:

  • Picture: the same weekly wiggle stays the same height under identity, and grows with the trend under log.
  • Numbers: a log-scale weekend effect of 0.2 is ×1.22: +22 orders at level 100, +221 at level 1 000.
  • Slogan: likelihood from the support, link from the range, columns from the story.
  1. Log link (NB forecasting): $\log\mu_t = g(t) + s(t) + h(t) + X_t\beta$, so $\mu_t = e^{g(t)}\cdot e^{s(t)}\cdot e^{h(t)}\cdot e^{X_t\beta}$. A weekend coefficient of 0.2 multiplies demand by $e^{0.2} \approx 1.221$: at level 100 that is 122.1 (+22); at level 1 000 it is 1 221 (+221).
  2. Identity link (Normal/Student-t forecasting): a weekend effect of +22 is +22 at any level. If the real effect is proportional, residuals will show a growing weekly pattern as the level rises.
  3. Softplus link: $\mu = \log(1 + e^\eta)$. At $\eta = 5$, $\mu = 5.007$; at $\eta = -3$, $\mu = 0.049$. Components add (almost exactly) when the mean is large, and a strongly negative component cannot push the mean below 0.
  4. A/B conversion: a Beta-Binomial model puts the prior directly on the rate $p$ (identity-like); a Bayesian logistic regression puts priors on log-odds coefficients. Same data, same likelihood (Binomial), different parameter scale.
Data (support, variance)LikelihoodUsual linkWhere in your projects
0/1 or k of n; Var = np(1−p)Bernoulli / Binomiallogit (or none: Beta prior on p)A/B conversion metrics (Beta-Binomial, 6.3)
one of K categoriesCategorical / Multinomialsoftmax (multinomial logit)A/B categorical metrics (Dirichlet-Multinomial)
counts, Var ≈ meanPoissonlogA/B count metrics
counts, Var > meanNegative Binomial (NB2)log (or softplus)count demand forecasting
real values, light tailsNormalidentitycontinuous A/B metrics; forecasting
real values, heavy tailsStudent-tidentityrobust A/B metrics; forecasting with outliers
  • Student-t is not a classic GLM (not in the exponential family), but it has the same structure: a mean from a linear predictor plus a likelihood with extra parameters (ν, σ). The GLM way of thinking still applies.
  • Bayesian GLM = GLM + priors on β (and on σ, ν or α). The posterior mode with Normal priors is like ridge-penalized IRLS (Chapter 5.3); the full posterior is approximated by SVI or NUTS (Chapters 6.12, 6.10).
  • A positive link is required whenever the likelihood needs a positive mean (Poisson, NB). An identity link with a count likelihood can produce a negative mean, which is invalid (NumPyro then returns NaN log-probabilities, or raises an error if argument validation is switched on).
Why do we need it?

It turns "which likelihood did you use and why?" into a crisp answer: support and variance structure pick the distribution; the required mean range picks the link; and the link decides whether components add or multiply, which changes forecasts at high and low levels.

Where is it used?

Every NumPyro numpyro.sample('y', dist.X(...), obs=y) line in both projects; Prophet's additive vs multiplicative seasonality option (Chapter 7.7); likelihood selection (Chapter 7.13, 7.20).

How is it used?

Write the model as eta = X @ beta; mu = link_inv(eta); sample('y', Likelihood(mu, ...), obs=y). Check residuals or posterior predictive variance against the variance function; if seasonal swings grow with the level, consider the log link (or log-transform for Normal models).

Same three questions, two projects A/B framework conversions, categories, counts, revenue forecasting model daily demand (continuous or counts) Binomial · Multinomial Poisson · Normal · Student-t likelihood from support Normal · Student-t Negative Binomial likelihood from support + variance rate p directly (Beta prior) or logit / log link + hierarchical priors identity: components add log/softplus: positive μ + Laplace / Normal priors
Both projects follow the GLM recipe: the type of data picks the likelihood, the allowed range of the mean picks the link (or a prior placed directly on the rate), and priors on the weights make it Bayesian.

A demand level that grows from about 20 to about 100 over 20 weeks, with a weekend bump and an optional stock-out week. Blue = identity link (components add), orange = log link (components multiply), purple = softplus. With the shock at 0, compare the weekend bumps at the start and the end: blue's stays the same size, orange's grows with the level. Now raise the stock-out size: the identity mean goes below zero (red zone; impossible for counts), while log and softplus stay positive.

"Using a log link is the same as log-transforming the demand and fitting a Normal model."

The log link models $\log E[y]$ with a count likelihood and handles zeros; a Normal model on $\log y$ models $E[\log y]$, cannot take zeros, and its back-transformed forecast is a median-like quantity (Chapter 4.18).

"An NB likelihood with an identity link is fine as long as the trend is positive."

Seasonal dips, negative holiday effects or a negative regressor can push $\mu_t$ below 0 on some days, and the likelihood is then undefined (NaN log-probabilities in NumPyro, which derail SVI). Use exp or softplus, or constrain the components.

How to explain the likelihood choice of your forecasting model in one breath: "The mean is a linear predictor built from trend, Fourier seasonality, holidays and regressors. For continuous demand I use a Normal likelihood with an identity link, or Student-t when residuals have heavy tails. For count demand I use an NB2 likelihood, because counts are overdispersed; it needs a positive mean, so the linear predictor goes through a positive link; with a log link that also makes the components multiplicative." For the A/B framework: "Conversions are Binomial with a Beta prior on the rate; categorical outcomes are Multinomial with a Dirichlet prior; counts are Poisson; revenue-like metrics are Normal or Student-t. Each is the random component of a GLM, and segments get hierarchical priors."

Likelihood ← support + variance; link ← allowed range of the mean; columns ← what drives the mean; priors make it Bayesian.

Identity: components add. Log: components multiply ($e^{0.2} = 1.22$ at every level). Softplus: ≈ additive when large, positive always.

Trap: count likelihoods need a positive mean; log link ≠ log-transforming the data.

Quick check: your NB forecasting model uses a log link and has a December holiday coefficient of 0.5. Demand on an ordinary December day is 400. What does the model expect on the holiday?

$400\times e^{0.5} \approx 400\times1.649 \approx 660$ (about +260). With the same coefficient in a quiet month at level 100, the holiday would add only about 65: the effect is a percentage, not a fixed number of units.

Recap, cheat sheet and practice

  • Straight lines fail for yes/no and count data: the mean leaves its allowed range, the variance depends on the mean, and effects are often multiplicative.
  • A GLM is $g(E[y\mid x]) = x^\top\beta$: a random component (likelihood) + a linear predictor $\eta = X\beta$ + a link $g$. Linear regression is Normal + identity.
  • Links: identity, logit, log (canonical for Normal, Bernoulli, Poisson); probit, cloglog, softplus as alternatives. Coefficients live on the $\eta$ scale; a +1 in $\eta$ changes the mean by different amounts at different levels (logit) or by a constant factor (log).
  • Logistic regression: log-odds linear; $e^\beta$ = odds ratio; with only a treatment dummy it reproduces the two-proportion test (checkout: OR 1.23, $p \approx 0.31$). An odds ratio is not a relative lift unless the event is rare.
  • Poisson regression: log-mean linear; $e^\beta$ = rate ratio; use $\log(\text{exposure})$ as an offset to compare rates.
  • Overdispersion (Var > mean): Poisson SEs too small → false positives; check Pearson χ²/df. NB2 regression adds $Var = \mu + \mu^2/\alpha$; know the parameterization (NumPyro concentration α, statsmodels alpha = 1/α, SciPy $n = \alpha$, $p = \alpha/(\alpha+\mu)$).
  • Fitting: maximum likelihood by IRLS (repeated weighted least squares, Newton's method); SEs from $(X^\top WX)^{-1}$; deviance is the GLM's RSS; LR tests and AIC compare models.
  • Both projects are Bayesian GLMs: likelihood from support and variance, link from the mean's range (log/softplus keep NB means positive and make components multiplicative), priors on the weights.

Cheat sheet

IdeaFormulaRemember
GLM$g(\mu_i) = x_i^\top\beta$, $y_i \sim$ family$(\mu_i)$link acts on the mean, not on $y$
Logistic$\log\frac{p}{1-p} = x^\top\beta$, $p = \frac{1}{1+e^{-x^\top\beta}}$$e^\beta$ = odds ratio
Odds ratio vs lift$OR = RR\cdot\frac{1-p_A}{1-p_B}$OR ≈ RR only for rare events
2×2 log-OR SE$\sqrt{1/a + 1/b + 1/c + 1/d}$checkout: 0.203
Poisson$\log\mu = x^\top\beta$, $Var = \mu$$e^\beta$ = rate ratio
Offset$\log\mu_i = \log t_i + x_i^\top\beta$compare rates, not counts
Dispersion$\hat\phi = \frac{1}{n-p}\sum\frac{(y-\hat\mu)^2}{\hat\mu}$≈ 1 for Poisson; SEs × $\sqrt{\hat\phi}$
NB2$Var = \mu + \mu^2/\alpha$NumPyro concentration α; statsmodels alpha = 1/α
IRLS$\beta \leftarrow (X^\top WX)^{-1}X^\top Wz$logit $w = p(1-p)$; Poisson $w = \mu$; NB $w = \frac{\mu}{1+\mu/\alpha}$
Deviance, LR test$D = 2[\ell_{sat} - \ell]$; $\Delta D \sim \chi^2_{\Delta p}$AIC $= -2\ell + 2k$ for Poisson vs NB
Code it · Python

import numpy as np
import statsmodels.api as sm
from scipy import stats
np.set_printoptions(suppress=True)

# 1) Logistic regression = the checkout A/B test (50/500 control, 60/500 treatment)
treat = np.r_[np.zeros(500), np.ones(500)]
conv = np.r_[np.ones(50), np.zeros(450), np.ones(60), np.zeros(440)]
logit = sm.GLM(conv, sm.add_constant(treat), family=sm.families.Binomial()).fit()
print(logit.params.round(4), logit.bse.round(4))      # [-2.1972  0.2048] [0.1491 0.2029]
print(round(np.exp(logit.params[1]), 3), round(logit.pvalues[1], 3))   # odds ratio 1.227, p 0.313
print(np.exp(logit.conf_int()[1]).round(2))           # OR 95% CI [0.82 1.83]
p_a, p_b = logit.predict([[1, 0], [1, 1]])
print(round(p_b - p_a, 3), round(p_b / p_a, 3))       # difference 0.02, lift (risk ratio) 1.2

# 2) Poisson regression: orders on normal vs promo days
orders = np.array([14, 25, 20, 16, 25, 24, 36, 30, 25, 35])
promo = np.r_[np.zeros(5), np.ones(5)]
pois = sm.GLM(orders, sm.add_constant(promo), family=sm.families.Poisson()).fit()
print(pois.params.round(3), pois.bse.round(3))        # [2.996 0.405] [0.1   0.129]
print(np.exp(pois.params[1]).round(3), np.exp(pois.conf_int()[1]).round(2))   # rate ratio 1.5, CI [1.16 1.93]
print(round(pois.pearson_chi2 / pois.df_resid, 2))    # dispersion 1.15 (close to 1: Poisson is fine)
print(round(pois.null_deviance - pois.deviance, 2))   # 10.07 = likelihood-ratio statistic on 1 df

# 3) Exposure offset: region A = 4 stores x 10 hours (120 orders), region B = 4 stores x 5 hours (90 orders)
counts = np.array([28, 31, 33, 28, 21, 24, 22, 23])
hours = np.r_[np.full(4, 10.0), np.full(4, 5.0)]
regionB = np.r_[np.zeros(4), np.ones(4)]
off = sm.GLM(counts, sm.add_constant(regionB), family=sm.families.Poisson(), offset=np.log(hours)).fit()
print(np.exp(off.params).round(3))                    # [3.  1.5]: 3 per hour in A, rate ratio 1.5
naive = sm.GLM(counts, sm.add_constant(regionB), family=sm.families.Poisson()).fit()
print(np.exp(naive.params[1]).round(3))               # 0.75: ignoring hours flips the conclusion

# 4) Overdispersion: counts with NO effect of x, NB noise (mean 5, concentration 2)
rng = np.random.default_rng(7)
n = 2000
x = rng.normal(size=n)
y = rng.poisson(rng.gamma(shape=2.0, scale=5 / 2.0, size=n))   # Gamma-Poisson = NB2(mean 5, conc 2)
X = sm.add_constant(x)
pfit = sm.GLM(y, X, family=sm.families.Poisson()).fit()
print(round(pfit.pearson_chi2 / pfit.df_resid, 2))    # about 3.5 = 1 + mean/concentration: overdispersed
nbfit = sm.NegativeBinomial(y, X).fit(disp=0)          # NB2; estimates statsmodels' alpha = 1/concentration
print(nbfit.params.round(3))                          # [const, x, alpha]; alpha near 0.5
print(round(1 / nbfit.params[-1], 2))                 # concentration near 2 (NumPyro NegativeBinomial2)
print(pfit.bse.round(4), nbfit.bse[:2].round(4))      # [0.0099 0.01] vs [0.0183 0.0186]: Poisson SEs ~1.9x too small

# 5) One distribution, three parameterizations (mean 6, concentration 2: variance 6 + 36/2 = 24)
mu, conc = 6.0, 2.0
d = stats.nbinom(n=conc, p=conc / (conc + mu))        # SciPy counts failures before the n-th success
print(d.mean(), d.var(), round(d.pmf(0), 4))          # 6.0 24.0 0.0625

# 6) IRLS by hand: intercept-only logistic model, 7 conversions in 10
b = 0.0
for step in range(4):
    p = 1 / (1 + np.exp(-b)); w = p * (1 - p)
    b = b + (0.7 - p) / w                             # weighted LS on z = b + (y - p) / w, all weights equal
    print(step + 1, round(b, 4))                      # 0.8, 0.8469, 0.8473, 0.8473 = log(0.7 / 0.3)
Test yourself

1. Which of these is not one of the three parts of a GLM?

A GLM's randomness comes from its chosen distribution (Bernoulli, Poisson, NB…), centred on $\mu$ with a mean-dependent variance. Only the Normal-identity GLM (linear regression) can be written as "mean + Normal error".

2. A logistic regression gives a treatment coefficient of 0.7. What does it mean?

Logistic coefficients are log odds ratios: $e^{0.7} \approx 2.01$. The change in probability (and the relative lift) depends on the baseline rate, so it must be computed at a stated baseline.

3. In a Poisson regression of daily orders, the promo coefficient is 0.2. On promo days the expected count is…

With a log link, $\mu = e^{\beta_0 + 0.2\cdot\text{promo}}$, so the promo multiplies the mean by $e^{0.2} \approx 1.22$, a 22% lift at any baseline.

4. A Poisson regression has Pearson χ²/df = 4. What is the main problem?

Dispersion 4 means the variance is about 4 times what Poisson assumes; SEs should be about $\sqrt4 = 2$ times larger. Use NB, quasi-Poisson or robust SEs. The mean model can still be fine.

5. statsmodels reports an NB2 alpha of 0.25. Which NumPyro call describes the same distribution for a mean μ?

statsmodels writes $Var = \mu + a\mu^2$, NumPyro writes $Var = \mu + \mu^2/\alpha$. So $\alpha = 1/a = 4$.

6. Users in one arm were observed for twice as many days as users in the other. What should a Poisson model of their event counts include?

Expected count = days × rate, so $\log\mu = \log(\text{days}) + x^\top\beta$. The offset makes the coefficients rate ratios. Dividing by exposure and using OLS ignores that short windows are noisier.

Practice problems

A. Control converts at 4%, treatment at 5%. Compute the absolute difference, the relative lift, the odds ratio and the logistic coefficient for treatment.
  1. Difference: $5\% - 4\% = 1$ percentage point. Lift (risk ratio): $0.05/0.04 = 1.25$.
  2. Odds: $0.04/0.96 = 0.0417$ and $0.05/0.95 = 0.0526$. Odds ratio $= 0.0526/0.0417 \approx 1.263$.
  3. Logistic coefficient $= \log 1.263 \approx 0.234$. Because conversion is rare, OR (1.263) ≈ RR (1.25).
B. A Poisson model says $\log\mu = 1.2 + 0.3\,\text{promo} + 0.05\,\text{temp}$. Find the expected count on a promo day at 10 °C, and the effect of the promo.

$\eta = 1.2 + 0.3 + 0.05\times10 = 2.0$, so $\mu = e^{2} \approx 7.39$. The promo multiplies the expected count by $e^{0.3} \approx 1.35$ (+35%) at any temperature; each extra °C multiplies it by $e^{0.05} \approx 1.05$.

C. (Interview) "Why not just use linear regression for a binary conversion outcome?"

"Three reasons. The fitted line can predict probabilities below 0 or above 1 for extreme inputs. The variance of a 0/1 outcome is $p(1-p)$, which changes with $p$, so OLS's constant-variance standard errors are off. And effects on probabilities naturally flatten near 0 and 1, which a line cannot do. Logistic regression fixes all three with a Bernoulli likelihood and a logit link. That said, for a randomized A/B test with only a treatment dummy, the OLS coefficient is exactly the difference in rates and is fine with robust SEs; the problems appear when modelling continuous covariates or individual probabilities."

D. A Poisson regression reports $\hat\beta = 0.10$ with SE 0.04 for a feature, and Pearson χ² = 450 on 150 df. Recompute the test honestly.
  1. Poisson's own test: $z = 0.10/0.04 = 2.5$, $p \approx 0.012$: "significant".
  2. Dispersion $\hat\phi = 450/150 = 3$. Quasi-Poisson SE $= 0.04\sqrt3 \approx 0.069$.
  3. Honest $z = 0.10/0.069 \approx 1.44$, $p \approx 0.15$: not significant. An NB model would reach a similar conclusion.
E. (Interview) "Your forecasting model has an NB likelihood. Why does it need a link function, and what does the link do to seasonality?"

"The NB mean must be positive, but a sum of trend, Fourier, holiday and regressor terms can be negative on some days, so I pass the linear predictor through a positive link (exp or softplus) before the likelihood. With exp, the model is additive on the log scale and multiplicative on the demand scale: a weekly effect is a percentage, so its absolute size grows with the level, like Prophet's multiplicative seasonality. Softplus is close to additive at high levels and only bends near zero. With Normal or Student-t likelihoods I use the identity link, so components add in demand units. The NB's concentration then controls how much extra day-to-day variance there is beyond Poisson."

F. Adding two holiday columns to a Poisson GLM lowers the deviance by 7.2. Are the holidays jointly significant at 5%?

The models are nested and differ by 2 parameters, so compare 7.2 with $\chi^2_2$: the 5% critical value is 5.99 and $P(\chi^2_2 \ge 7.2) = e^{-3.6} \approx 0.027$. Yes, jointly significant (if the data are not overdispersed; otherwise divide the drop by $\hat\phi$ first, or use an NB model).

Chapter 5.15 · Syllabus Module 20

Multivariate statistics: covariance matrices, ellipses and the multivariate Normal

Real data rarely comes one number at a time. A day has sessions and orders and revenue; a Bayesian model has dozens of parameters that move together. This chapter gives you one object that describes how many quantities wobble together (the covariance matrix), one picture that makes it visible (a tilted ellipse), and the tools to read that picture: eigenvectors, SVD and Mahalanobis distance. It ends with why all of this decides which variational guide your forecasting model can afford.

  • Treat several measurements of one unit as one random vector, and see why the columns alone (the marginals) are not the whole story
  • Build a covariance matrix by hand, read its entries, and use $Cov(AX+b) = A\Sigma A^\top$
  • Explain why every covariance matrix is symmetric and positive semi-definite, and spot an impossible one
  • Turn a covariance matrix into a correlation matrix and know which one changes with units
  • See the multivariate Normal as a hill whose contours are ellipses set by $\Sigma$, and know that the 1-sd ellipse in 2D holds only 39% of the data
  • Read the ellipse's axes as the eigenvectors of $\Sigma$ and their lengths as $\sqrt{\text{eigenvalues}}$; get the same axes from the SVD of the centred data
  • Measure unusualness with the Mahalanobis distance
  • Describe common covariance structures and explain why they decide between mean-field, low-rank and full-rank variational guides

What we need from earlier chapters: covariance and correlation of two variables and the correlation matrix (Chapter 4.15); variance rules such as $Var(aX) = a^2Var(X)$ (Chapter 4.5); the Normal distribution (Chapter 4.9). From the Linear Algebra guide: matrix products and transposes (1.4), eigenvalues and eigenvectors (1.11), quadratic forms and positive definiteness (1.12), Cholesky and SVD (1.13). We do not re-teach those; we use them. Notation for this chapter: $d$ = number of variables, $n$ = number of observations (rows). Vectors are bold: $\mathbf{X}$ is a random vector, $\mathbf{x}_i$ one observed row. $\boldsymbol\mu$ is the mean vector, $\Sigma$ ("capital sigma", not the summation sign) the covariance matrix, $\hat\Sigma$ its estimate from data. For the SVD we write $X_c = USV^\top$, so the letter $S$ means singular values here.

Random vectors: several numbers drawn together core

Think of one day in your shop. You do not get one number; you get a small bundle: how many people visited, how many ordered, how much they spent. These numbers arrive together, from the same day, and they are linked: busy days tend to have more orders.

If you look at each number on its own (only the visits column, then only the orders column), you can learn each one's average and spread. But you lose the most interesting part: how they move together. Two very different worlds can have exactly the same columns. The widget below shows this.

A random vector is simply a bundle of random numbers that are drawn together. One day is one draw of the bundle; your dataset is a stack of draws.

Three ways to say it:

  • Picture: each day is one dot in a scatter plot, not two separate ticks on two separate rulers.
  • Numbers: the day (sessions = 9 thousand, orders = 6 hundred) is one draw $\mathbf{x} = (9, 6)$ of the random vector (sessions, orders).
  • Slogan: the columns tell you about each variable; only the rows tell you how they go together.

Five days of two metrics (we use these numbers all through this chapter and the next). Sessions are in thousands, orders in hundreds.

daysessions $x_{i1}$orders $x_{i2}$
174
296
3106
4118
5136
  1. Each row is one observed vector: $\mathbf{x}_1 = (7, 4)$, $\mathbf{x}_2 = (9, 6)$, and so on. Stacking the 5 rows gives a $5 \times 2$ data matrix $X$ (5 rows = observations, 2 columns = variables).
  2. The mean vector is the column means: sessions $(7+9+10+11+13)/5 = 50/5 = 10$; orders $(4+6+6+8+6)/5 = 30/5 = 6$. So $\bar{\mathbf{x}} = (10, 6)$.
  3. Centre the data by subtracting the mean vector from every row: $(-3,-2),\ (-1,0),\ (0,0),\ (1,2),\ (3,0)$. Each centred row is an arrow from the "average day" to that day.
  4. Each column alone: sessions have sample variance $(9+1+0+1+9)/4 = 5$; orders have $(4+0+0+4+0)/4 = 2$. These are the marginal summaries.
  5. What the columns cannot tell you: on days with more sessions, are there more orders? For that you need the pairs (rows). That is the job of the covariance matrix, in the next section.

A random vector $\mathbf{X} = (X_1, X_2, \dots, X_d)^\top$ is a list of $d$ random variables defined on the same random experiment (the same user, the same day, the same posterior draw).

  • Its joint distribution describes all $d$ numbers together, including how they are related.
  • The marginal distribution of $X_j$ is the distribution of that one coordinate alone, ignoring the others. (The name comes from the totals written in the margins of a table.) Knowing every marginal does not determine the joint distribution.
  • The mean vector is $\boldsymbol\mu = E[\mathbf{X}] = (E[X_1], \dots, E[X_d])^\top$: one mean per coordinate.
  • Data: $n$ observed vectors stacked as the rows of the data matrix $X \in \mathbb{R}^{n\times d}$. The sample mean vector $\bar{\mathbf{x}} = \frac1n\sum_{i=1}^n \mathbf{x}_i$ is the vector of column means. The centred data matrix $X_c = X - \mathbf{1}\bar{\mathbf{x}}^\top$ subtracts each column's mean from that column ($\mathbf{1}$ is a column of $n$ ones).

As in Chapter 4.4, capital $\mathbf{X}$ is the random vector before we look; lowercase $\mathbf{x}_i$ is a value we observed.

Why do we need it?

Relationships between variables live in the joint distribution, not in the columns. Without the idea of a random vector you cannot even ask "do busy days have more orders?" or "do these two model parameters trade off against each other?".

Where is it used?

Every tabular ML dataset (rows = examples), the design matrix of a regression (Chapter 5.13), several metrics measured on the same users in an A/B test, and the parameter vector $\theta$ of a Bayesian model, whose posterior is a distribution over the whole vector.

How is it used?

Store data as an $n\times d$ array with one row per observation. X.mean(axis=0) gives the mean vector, X - X.mean(axis=0) centres it. Always look at a scatter plot or a pair plot (pandas.plotting.scatter_matrix), not only at per-column summaries.

data matrix X (5 × 2) sessionsorders 74 96 106 118 136 day 2 mean 10mean 6 one row = one draw of the random vector x₂ = (9, 6): a dot in the scatter plot one column = one variable alone it shows only the marginal mean vector x̄ = (10, 6) = the column means
Rows are observations (each row is one draw of the random vector); columns are variables. Relationships between variables can only be seen by keeping each row's numbers together.

300 simulated days of (sessions, orders). The grey histograms on the bottom and on the left show each column alone (the marginals). Turn on Shuffle the pairs: the orders are shuffled between days, so both columns keep exactly the same numbers. The two histograms do not change at all, but the cloud loses its tilt and the correlation drops to about 0. Move the $\rho$ slider to build other worlds with the same marginals.

"If I know the mean and sd of every column, I know the data."

Those are only the marginals. The shuffle widget keeps every column identical and still destroys the relationship. You need the joint view: scatter plots, the covariance matrix, the correlation matrix.

"A random vector is a list of numbers."

The list you observed, $\mathbf{x}_i$, is one draw. The random vector $\mathbf{X}$ is the process that produces such lists, with a joint distribution.

"Rows or columns, it does not matter which way I store the data."

It matters for library calls. Most ML code (scikit-learn, pandas) expects rows = observations. NumPy's np.cov expects the opposite by default (next section).

In an A/B framework like yours, one user (or one randomization unit) gives a vector: converted or not, revenue, number of items. Metrics measured on the same users are correlated, so the uncertainty of a combined metric (revenue per converter, a weighted score) depends on the joint distribution, not only on each metric's own variance. In your forecasting model, all the latent quantities (trend slope $k$, offset $m$, every changepoint adjustment $\delta_j$, every regressor coefficient $\beta$, every Fourier coefficient, the noise scale) form one random vector $\theta$ of dimension $d$. The posterior is a distribution over that whole vector, and that whole vector is what your SVI guide approximates.

Random vector $\mathbf{X} = (X_1, \dots, X_d)^\top$: several random numbers drawn together. One observation = one row $\mathbf{x}_i$.

Data matrix $X$: $n \times d$, rows = observations, columns = variables. $\bar{\mathbf{x}}$ = column means; $X_c = X - \mathbf{1}\bar{\mathbf{x}}^\top$.

Trap: the marginals (columns) do not determine the joint distribution.

Quick check: two datasets have the same column means and the same column standard deviations. Must they have the same correlation?

No. Shuffling one column keeps every column summary identical but changes (usually destroys) the correlation. Column summaries describe the marginals only; correlation lives in how the rows pair the values.

The covariance matrix: every variance and covariance in one table core

With two variables there are three numbers to know: how much the first wobbles, how much the second wobbles, and how much they wobble together. With $d$ variables there are $d$ wobbles and many pairs. The covariance matrix puts them all in one square table: the variances on the diagonal, the covariances everywhere else.

Where does it come from? Draw an arrow from the average day to each day (the centred rows). The covariance matrix is the average of "arrow times arrow": for each arrow, multiply its parts with each other in every combination, then average over the arrows. Arrows that point along a tilted direction produce a tilted shape, and the matrix records that tilt.

Three ways to say it:

  • Picture: the covariance matrix is the shape summary of the cloud: how wide, how tall, how tilted.
  • Numbers: for our five days $\hat\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$: sessions vary more (5) than orders (2), and they move together (+2).
  • Slogan: the mean vector says where the cloud is; the covariance matrix says what shape it has.

The five centred days from the previous section: $(-3,-2),\ (-1,0),\ (0,0),\ (1,2),\ (3,0)$.

  1. Sessions with sessions (sum of squares of the first column): $9 + 1 + 0 + 1 + 9 = 20$.
  2. Orders with orders: $4 + 0 + 0 + 4 + 0 = 8$.
  3. Sessions with orders (sum of products): $(-3)(-2) + (-1)(0) + 0\cdot0 + (1)(2) + (3)(0) = 6 + 0 + 0 + 2 + 0 = 8$.
  4. Divide each by $n - 1 = 4$: variance of sessions $20/4 = 5$, variance of orders $8/4 = 2$, covariance $8/4 = 2$.
  5. Arrange them: $\hat\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$. The two off-diagonal entries are the same number, because "sessions with orders" equals "orders with sessions".
  6. The same thing as a sum of arrow-times-arrow tables (outer products). Day 1 gives $\begin{bmatrix} -3 \\ -2 \end{bmatrix}\begin{bmatrix} -3 & -2 \end{bmatrix} = \begin{bmatrix} 9 & 6 \\ 6 & 4 \end{bmatrix}$. Adding all five tables gives $\begin{bmatrix} 20 & 8 \\ 8 & 8 \end{bmatrix} = X_c^\top X_c$, and dividing by 4 gives $\hat\Sigma$ again.

Using it: the variance of "sessions minus orders" is $[1, -1]\,\hat\Sigma\,[1, -1]^\top = 5 + 2 - 2\cdot 2 = 3$. The covariance term matters: if you ignored it you would say $5 + 2 = 7$.

The covariance matrix of a random vector $\mathbf{X}$ with mean $\boldsymbol\mu$ is the $d \times d$ matrix

$$\Sigma = Cov(\mathbf{X}) = E\big[(\mathbf{X}-\boldsymbol\mu)(\mathbf{X}-\boldsymbol\mu)^\top\big], \qquad \Sigma_{ij} = Cov(X_i, X_j), \quad \Sigma_{ii} = Var(X_i).$$

Its estimate from data (the sample covariance matrix) is

$$\hat\Sigma = \frac{1}{n-1}\sum_{i=1}^n (\mathbf{x}_i - \bar{\mathbf{x}})(\mathbf{x}_i - \bar{\mathbf{x}})^\top = \frac{X_c^\top X_c}{n-1}.$$
  • It has $d(d+1)/2$ distinct numbers ($d$ variances plus $d(d-1)/2$ covariances): 3 for $d = 2$, 55 for $d = 10$, 500 500 for $d = 1000$.
  • Entry $\Sigma_{ij}$ has the units of $X_i$ times the units of $X_j$ (thousand sessions × hundred orders), so its size depends on the units.
  • The key rule (linear maps): for a fixed matrix $A$ and vector $\mathbf{b}$, $Cov(A\mathbf{X} + \mathbf{b}) = A\,\Sigma\,A^\top$. Reason: $A\mathbf{X}+\mathbf{b}$ minus its mean is $A(\mathbf{X}-\boldsymbol\mu)$, so its outer product is $A(\mathbf{X}-\boldsymbol\mu)(\mathbf{X}-\boldsymbol\mu)^\top A^\top$, and the expectation goes inside.
  • Special case, one combination $\mathbf{a}^\top\mathbf{X} = a_1X_1 + \dots + a_dX_d$: $\;Var(\mathbf{a}^\top\mathbf{X}) = \mathbf{a}^\top\Sigma\,\mathbf{a}$. For two variables this is the familiar $Var(X - Y) = Var(X) + Var(Y) - 2Cov(X,Y)$ from Chapter 4.15.
Why do we need it?

Any time you add, subtract or weight several correlated quantities (a combined metric, a forecast of a sum, a prediction from several coefficients), its variance needs every covariance. The matrix holds them all and the rule $\mathbf{a}^\top\Sigma\mathbf{a}$ uses them in one step.

Where is it used?

The multivariate Normal, Mahalanobis distance, PCA (Chapter 5.16), the standard errors of regression coefficients $\sigma^2(X^\top X)^{-1}$ (Chapter 5.13), Kalman filters, Gaussian processes, CUPED (Chapter 5.12), and the posterior covariance a full-rank variational guide learns.

How is it used?

np.cov(X, rowvar=False) or df.cov(). Read the diagonal (each variable's spread) and the off-diagonal signs (which pairs move together). Convert to correlations to compare strengths. For any weighted sum, compute a @ S @ a.

sessionsordersrevenue Var(S)Var(O)Var(R) Cov(S,O)Cov(S,R)Cov(O,R) Cov(O,S)Cov(R,S)Cov(R,O) same colour = same number:the matrix is symmetric (Σᵢⱼ = Σⱼᵢ) diagonal (blue): each variable's variance off-diagonal: covariances (sign = directionof co-movement, size depends on units) d = 3 → 3 variances + 3 covariances = 6 numbers
A covariance matrix for three metrics. The diagonal holds the variances; each covariance appears twice, mirrored across the diagonal, so only $d(d+1)/2$ numbers are free.

Five days (blue dots, start = the example). The purple dot is the mean; the thin arrows are the centred rows $\mathbf{x}_i - \bar{\mathbf{x}}$. Drag a day up and to the right of the mean: its arrow adds a positive product, and the covariance grows. Drag it down and to the right: the product is negative. Press Opposite and No link to see the off-diagonal entry change sign and vanish. The shaded ellipse is the shape this matrix describes (you will meet it properly below).

"np.cov(X) on my $n \times d$ data gives the $d\times d$ covariance matrix."

By default np.cov treats each row as a variable and returns an $n\times n$ matrix. Use np.cov(X, rowvar=False) (or np.cov(X.T)). pandas' df.cov() already uses columns.

"np.cov divides by $n$, like np.var."

No: np.cov divides by $n-1$ by default (ddof=None means $n-1$), while np.var divides by $n$. Our example gives 5, 2, 2 with np.cov and 4, 1.6, 1.6 if you pass ddof=0.

"The covariance of 2 between sessions and orders is weak because 2 is small."

A covariance has units, so its size means nothing on its own. Count sessions one by one instead of in thousands and the same relationship becomes 2 000. Compare strengths with correlations (two sections below).

In an A/B framework like yours, conversions and revenue are measured on the same users, so a derived metric such as "revenue minus a cost per conversion" has variance $\mathbf{a}^\top\Sigma\mathbf{a}$, covariance term included. In your forecasting model, the posterior covariance matrix of the latent vector is exactly what the different SVI guides disagree about: a mean-field guide keeps only its diagonal, a full-rank guide learns all of it (last two sections of this chapter).

$\Sigma = E[(\mathbf{X}-\boldsymbol\mu)(\mathbf{X}-\boldsymbol\mu)^\top]$; $\hat\Sigma = X_c^\top X_c/(n-1)$. Diagonal = variances, off-diagonal = covariances, $d(d+1)/2$ free numbers.

$Cov(A\mathbf{X}+\mathbf{b}) = A\Sigma A^\top$; $Var(\mathbf{a}^\top\mathbf{X}) = \mathbf{a}^\top\Sigma\mathbf{a}$.

Trap: np.cov(X, rowvar=False) for rows = observations; covariances depend on units.

Quick check: with $\hat\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$, what is the variance of "sessions + orders"?

$\mathbf{a} = (1, 1)$: $\mathbf{a}^\top\hat\Sigma\mathbf{a} = 5 + 2 + 2 + 2 = 11$ (the two variances plus twice the covariance). Without the covariance term you would wrongly get 7.

Two promises every covariance matrix keeps: symmetric and positive semi-definite core

Not every square table of numbers can be a covariance matrix. Every real one keeps two promises.

  • Symmetric: "sessions with orders" is the same number as "orders with sessions". The table is a mirror image across its diagonal.
  • Positive semi-definite (PSD): pick any direction and flatten the cloud onto a line pointing that way, like a shadow on a wall. The shadow has some spread, maybe zero, but never a negative spread. A variance can never be negative, and every direction is just one more combination of the variables.

If someone hands you a table that would give a negative variance for some combination, no data in the world can have it. It is broken, and code that needs its inverse or its Cholesky factor will crash or return NaNs.

Three ways to say it:

  • Picture: shine a light on the cloud from any side; the shadow is never "less than flat".
  • Numbers: the table $\begin{bmatrix} 4 & 5 \\ 5 & 4 \end{bmatrix}$ would make $Var(X - Y) = 4 + 4 - 2\cdot 5 = -2$. Impossible.
  • Slogan: a covariance matrix can never promise a negative variance.

Test 1: the table $A = \begin{bmatrix} 4 & 5 \\ 5 & 4 \end{bmatrix}$.

  1. It is symmetric. Each variance (4) is positive. So far so good.
  2. Try the combination $\mathbf{a} = (1, -1)$, that is $X - Y$: $\mathbf{a}^\top A\mathbf{a} = 4 - 5 - 5 + 4 = -2 \lt 0$. A negative variance: $A$ is not a valid covariance matrix.
  3. Eigenvalues confirm it: trace $= 8$, determinant $= 16 - 25 = -9$, so $\lambda = 4 \pm 5$, that is $9$ and $-1$. A negative eigenvalue means a direction with negative "variance".
  4. The quick rule for $2\times2$: valid exactly when both variances are $\ge 0$ and $Cov^2 \le Var(X)\,Var(Y)$, i.e. $|\rho| \le 1$. Here $25 \gt 16$, so the implied correlation would be $5/4 = 1.25$.

Test 2: three variables with correlations $r_{12} = 0.9$, $r_{13} = 0.9$, $r_{23} = -0.9$ (all variances 1; the example from Chapter 4.15).

  1. Every pair looks fine on its own: each $|r| \le 1$.
  2. Try $\mathbf{a} = (-1, 1, 1)$: $Var(-X_1 + X_2 + X_3) = 1 + 1 + 1 + 2\big[(-1)(1)(0.9) + (-1)(1)(0.9) + (1)(1)(-0.9)\big] = 3 + 2(-2.7) = -2.4$. Impossible.
  3. The smallest eigenvalue is $-0.8$, with eigenvector $(-1, 1, 1)/\sqrt3$; indeed $-2.4/3 = -0.8$. If $X_1$ is close to $X_2$ and close to $X_3$, then $X_2$ and $X_3$ cannot be strongly opposite.

A $d\times d$ matrix $\Sigma$ is a valid covariance matrix exactly when it is

  • symmetric: $\Sigma = \Sigma^\top$, and
  • positive semi-definite (PSD): $\mathbf{a}^\top\Sigma\,\mathbf{a} \ge 0$ for every vector $\mathbf{a}$. Equivalently, all its eigenvalues are $\ge 0$ (Linear Algebra 1.12).

Why it must hold: $\mathbf{a}^\top\Sigma\mathbf{a} = Var(\mathbf{a}^\top\mathbf{X})$, a variance. For data, $\mathbf{a}^\top\hat\Sigma\mathbf{a} = \|X_c\mathbf{a}\|^2/(n-1) \ge 0$ automatically, so every sample covariance matrix is PSD.

  • Positive definite (PD): $\mathbf{a}^\top\Sigma\mathbf{a} \gt 0$ for every $\mathbf{a}\ne\mathbf{0}$ $\iff$ all eigenvalues $\gt 0$ $\iff$ $\Sigma$ is invertible $\iff$ a Cholesky factor $\Sigma = LL^\top$ exists (1.13).
  • PSD but not PD (a zero eigenvalue): some combination has zero variance, so one variable is an exact linear combination of the others (revenue = 20 × orders when every order costs exactly 20; total = part A + part B). The cloud is flat, $\Sigma$ is singular, and $\Sigma^{-1}$ does not exist.
  • If $d \ge n$, the sample matrix $\hat\Sigma$ has rank at most $n - 1$, so it is always singular.
Why do we need it?

The multivariate Normal density, Mahalanobis distance, sampling and full-rank guides all need $\Sigma^{-1}$ or a Cholesky factor. A matrix that is not PSD (or is singular) makes them fail with errors or NaNs. Knowing the rule lets you diagnose that in seconds.

Where is it used?

Cholesky factors for sampling a multivariate Normal (NumPy's Generator.multivariate_normal offers method="cholesky"; its default is "svd") and NumPyro's MultivariateNormal(scale_tril=...), Gaussian-process kernels, correlation matrices assembled from separate pairwise estimates (often invalid), and the full-rank guide, which stores $\Sigma = LL^\top$ so it can never become invalid.

How is it used?

Check np.linalg.eigvalsh(S).min(). A tiny negative value like $-10^{-12}$ is rounding: add "jitter" (for example $10^{-6} I$; a rule of thumb). A clearly negative value means the matrix is inconsistent: fix the inputs, or project to the nearest PSD matrix.

Set the two variances and the covariance. Left: drag the purple handle to choose a direction $\mathbf{u}$; the green bar is the spread (±1 sd) of the cloud's shadow on that line. Right: the variance $\mathbf{u}^\top\Sigma\mathbf{u}$ for every direction from 0° to 180°. Press Impossible: Cov too big: part of the curve drops below zero (red) and no ellipse can be drawn. Press On the edge: the lowest point touches zero, the cloud is flat (a line), and $\Sigma$ has no inverse.

"Any symmetric table with 1s on the diagonal and entries between −1 and 1 is a valid correlation matrix."

Test 2 above has every entry in range and is still impossible. Validity is a property of the whole matrix (its eigenvalues), not of each pair.

"PSD means all the entries are positive."

Negative covariances are fine: $\begin{bmatrix} 1 & -0.9 \\ -0.9 & 1 \end{bmatrix}$ is positive definite (eigenvalues 1.9 and 0.1). PSD is about directions (eigenvalues), not about the signs of entries.

"An eigenvalue of $-3\times10^{-13}$ means my matrix is broken."

That is floating-point rounding on a matrix that should be singular or nearly so. Add a little jitter to the diagonal. A clearly negative eigenvalue (−0.8) means the numbers are inconsistent.

This is why NumPyro's full-rank guide (AutoMultivariateNormal) learns a lower-triangular scale_tril $L$ and uses $\Sigma = LL^\top$ instead of learning $\Sigma$ directly: every $L$ gives a valid covariance, so the optimizer can move freely without ever stepping onto an impossible matrix. The low-rank guide uses $WW^\top + \text{diag}(\psi)$ with $\psi \gt 0$ for the same reason. In your forecasting model, a regressor that is an exact sum of two others makes the regressor covariance singular: the data cannot tell the three coefficients apart, and the posterior becomes a flat ridge.

"We learn the covariance matrix of the guide directly with gradient descent."

"We learn a Cholesky factor $L$ (lower-triangular, positive diagonal) and form $\Sigma = LL^\top$, which is positive definite by construction."

Model answer: "A covariance matrix must be symmetric positive semi-definite, and a free gradient step on its entries can break that. Parameterizing it as $LL^\top$ removes the constraint: any lower-triangular $L$ with a positive diagonal gives a valid, invertible covariance, and $\log|\Sigma|$ is just $2\sum\log L_{ii}$."

Valid covariance ⟺ symmetric and PSD: $\mathbf{a}^\top\Sigma\mathbf{a} = Var(\mathbf{a}^\top\mathbf{X}) \ge 0$ for all $\mathbf{a}$ ⟺ all eigenvalues $\ge 0$.

PD (all eigenvalues $\gt 0$) ⟺ invertible ⟺ Cholesky exists. Zero eigenvalue = an exact linear relation, flat cloud.

2×2: valid ⟺ variances $\ge 0$ and $Cov^2 \le Var(X)Var(Y)$. Trap: pairwise-valid is not jointly valid.

Quick check: is $\begin{bmatrix} 2 & -1 \\ -1 & 2 \end{bmatrix}$ a valid covariance matrix? And $\begin{bmatrix} 1 & 2 \\ 2 & 1 \end{bmatrix}$?

The first: $Cov^2 = 1 \le 2\cdot2 = 4$, eigenvalues $3$ and $1$, so it is positive definite. The second: $Cov^2 = 4 \gt 1\cdot 1$, eigenvalues $3$ and $-1$; the direction $(1,-1)$ would have variance $1 + 1 - 4 = -2$. Not valid.

From the covariance matrix to the correlation matrix

Covariances carry units, so their sizes cannot be compared across pairs. The fix is the same as for two variables (Chapter 4.15): divide each covariance by the two standard deviations involved. Done for every cell at once, this turns the covariance matrix into the correlation matrix: ones on the diagonal, unit-free numbers between −1 and 1 everywhere else.

Change the units (count sessions one by one instead of in thousands) and every covariance involving sessions changes. The correlation matrix does not move. The shape of the cloud in standard-deviation units is the same.

Three ways to say it:

  • Picture: stretch the axes of the scatter plot as you like; the "tilt in sd units" stays.
  • Numbers: $\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$ gives $r = 2/(\sqrt5\sqrt2) = 0.632$; in raw units it is $\begin{bmatrix} 5\,000\,000 & 200\,000 \\ 200\,000 & 20\,000 \end{bmatrix}$ and still $r = 0.632$.
  • Slogan: covariance = strength × units; correlation = strength only.
  1. Standard deviations from the diagonal of $\hat\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$: $s_1 = \sqrt5 \approx 2.236$, $s_2 = \sqrt2 \approx 1.414$.
  2. Divide the covariance by both: $r_{12} = 2/(2.236 \times 1.414) = 2/\sqrt{10} \approx 0.632$.
  3. The correlation matrix: $R = \begin{bmatrix} 1 & 0.632 \\ 0.632 & 1 \end{bmatrix}$.
  4. Change units: sessions counted one by one (×1000) and orders one by one (×100). The rule $Cov(A\mathbf{X}) = A\Sigma A^\top$ with $A = \text{diag}(1000, 100)$ gives $\begin{bmatrix} 1000^2\cdot5 & 1000\cdot100\cdot2 \\ 1000\cdot100\cdot2 & 100^2\cdot2 \end{bmatrix} = \begin{bmatrix} 5\,000\,000 & 200\,000 \\ 200\,000 & 20\,000 \end{bmatrix}$.
  5. Correlation again: $200\,000/(\sqrt{5\,000\,000}\times\sqrt{20\,000}) = 200\,000/316\,228 \approx 0.632$. Unchanged.

Let $D = \text{diag}(\Sigma_{11}, \dots, \Sigma_{dd})$ be the diagonal matrix of variances. The correlation matrix is

$$R = D^{-1/2}\,\Sigma\,D^{-1/2}, \qquad R_{ij} = \frac{\Sigma_{ij}}{\sqrt{\Sigma_{ii}\Sigma_{jj}}} = \frac{\Sigma_{ij}}{\sigma_i\sigma_j}.$$
  • $R$ is the covariance matrix of the standardized variables $Z_j = (X_j - \mu_j)/\sigma_j$, so it is also symmetric and PSD, with ones on the diagonal.
  • Going back: $\Sigma = D^{1/2} R\, D^{1/2}$ ("sds × correlations × sds"). This split is how many Bayesian models put a prior on a covariance: a prior on the standard deviations and a separate prior on $R$.
  • Rescaling variables ($A$ diagonal with positive entries) changes $\Sigma$ but not $R$. $\Sigma$ and $R$ generally have different eigenvectors (they agree only when all variances are equal). This matters for PCA (Chapter 5.16).
Why do we need it?

To compare how strongly different pairs move together when the variables live in different units (dollars, counts, percentages), and to see the relationship pattern without the scale getting in the way.

Where is it used?

Correlation heatmaps in EDA and regressor screening (4.15), correlation-matrix PCA (5.16), the LKJ prior on correlation matrices in NumPyro (dist.LKJCholesky) for correlated group effects, and reading a fitted posterior covariance as posterior correlations.

How is it used?

np.corrcoef(X, rowvar=False) or df.corr(). From a covariance matrix: sd = np.sqrt(np.diag(S)); R = S / np.outer(sd, sd). Report $R$ for patterns and keep $\Sigma$ for computing variances of combinations.

60 simulated days drawn with the example covariance. Stretch the horizontal or vertical units with the sliders (as if you changed from thousands to some other unit). Watch the cloud and the covariance matrix change in the readout, while the correlation matrix stays fixed. Notice that the tilt of the ellipse in the picture changes too: the direction of the long axis depends on the units.

"A bigger covariance means a stronger relationship."

Only within the same units. Across pairs with different units compare correlations: $R_{ij}$ is unit-free.

"The covariance matrix and the correlation matrix point the cloud in the same directions."

Their eigenvectors differ unless all variances are equal (the widget's long axis moves when you change units, while $R$ does not). Which one you analyse changes the answer of PCA.

$R = D^{-1/2}\Sigma D^{-1/2}$, $R_{ij} = \Sigma_{ij}/(\sigma_i\sigma_j)$; back: $\Sigma = D^{1/2}RD^{1/2}$.

Units change $\Sigma$ (and its eigenvectors), never $R$.

Trap: comparing raw covariances across pairs with different units.

Quick check: $\Sigma = \begin{bmatrix} 9 & -3 \\ -3 & 4 \end{bmatrix}$. What is the correlation?

$\sigma_1 = 3$, $\sigma_2 = 2$, so $r = -3/(3\cdot2) = -0.5$.

The multivariate Normal: a hill with elliptical contours core

In one dimension the Normal is the bell curve: one peak at the mean, falling off on both sides. With two variables the bell becomes a hill standing on the floor of the scatter plot. Cut the hill horizontally at any height and the cut is an ellipse. All these ellipses have the same centre, the same tilt and the same proportions; only their size changes with the height.

The whole hill is fixed by just two things: the mean vector (where the top is) and the covariance matrix (how wide, how tall and how tilted the hill is). Nothing else. That is why the covariance matrix is so central: for Normal data it is the complete description of the shape.

Three ways to say it:

  • Picture: a smooth hill whose footprints are nested, tilted ellipses.
  • Numbers: with $\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$, two points at the same ruler distance $\sqrt5$ from the centre have densities 0.0428 and 0.0053: an 8-fold difference, decided by the tilt.
  • Slogan: mean = where, covariance = shape, and for a Normal that is everything.

Take $\boldsymbol\mu = (0, 0)$ (deviations from the average day) and $\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$. Compare the points $A = (2, 1)$ and $B = (1, -2)$. Both are $\sqrt{4+1} = \sqrt5 \approx 2.24$ from the centre by ruler.

  1. Determinant: $|\Sigma| = 5\cdot2 - 2\cdot2 = 6$. Inverse: $\Sigma^{-1} = \frac16\begin{bmatrix} 2 & -2 \\ -2 & 5 \end{bmatrix}$ (swap the diagonal, flip the signs off the diagonal, divide by the determinant).
  2. The "squared distance" in the exponent, $q = \mathbf{x}^\top\Sigma^{-1}\mathbf{x} = \frac16(2x_1^2 - 4x_1x_2 + 5x_2^2)$. For $A$: $\frac16(8 - 8 + 5) = \frac56 \approx 0.833$. For $B$: $\frac16(2 + 8 + 20) = 5$.
  3. Height at the centre: $\frac{1}{2\pi\sqrt{|\Sigma|}} = \frac{1}{2\pi\sqrt6} \approx 0.0650$.
  4. Density at $A$: $0.0650\times e^{-0.833/2} = 0.0650 \times 0.659 \approx 0.0428$.
  5. Density at $B$: $0.0650 \times e^{-5/2} = 0.0650\times0.0821 \approx 0.0053$.
  6. $A$ lies along the long direction of the cloud (where days often land); $B$ cuts across it. Same ruler distance, 8 times less likely per unit area.

A random vector $\mathbf{X}\in\mathbb{R}^d$ has a multivariate Normal distribution $\mathbf{X}\sim N(\boldsymbol\mu, \Sigma)$ (with $\Sigma$ positive definite) if its density is

$$p(\mathbf{x}) = \frac{1}{(2\pi)^{d/2}\,|\Sigma|^{1/2}}\exp\!\Big(-\tfrac12(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)\Big).$$
  • Contours (points of equal density) satisfy $(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu) = k^2$: ellipses in 2D, ellipsoids in 3D. We call the one with $k = 1$ the "1-sd ellipse".
  • Marginals are Normal: $X_j \sim N(\mu_j, \Sigma_{jj})$. Linear combinations are Normal: $A\mathbf{X}+\mathbf{b} \sim N(A\boldsymbol\mu + \mathbf{b}, A\Sigma A^\top)$.
  • Conditionals are Normal and narrower: for two variables, knowing $X_2 = x_2$ gives $X_1 \mid x_2 \sim N\big(\mu_1 + \frac{\Sigma_{12}}{\Sigma_{22}}(x_2 - \mu_2),\ \Sigma_{11} - \frac{\Sigma_{12}^2}{\Sigma_{22}}\big)$. Example: sessions given orders has variance $5 - 4/2 = 3$ instead of 5.
  • For a multivariate Normal, zero covariance means independence. (For other distributions it does not; see 4.15.)
  • It has $d$ mean parameters and $d(d+1)/2$ covariance parameters. To sample: $\mathbf{x} = \boldsymbol\mu + L\mathbf{z}$ with $\Sigma = LL^\top$ (Cholesky) and $\mathbf{z}$ a vector of independent $N(0,1)$ draws.
Why do we need it?

It is the default model for several continuous quantities that vary together: simple (only $\boldsymbol\mu$ and $\Sigma$), with Normal marginals, Normal conditionals and closed-form formulas for everything. Many approximations end up as one.

Where is it used?

Full-rank and low-rank variational guides, the Laplace approximation of a posterior, the sampling distribution of regression coefficients $\hat\beta \sim N(\beta, \sigma^2(X^\top X)^{-1})$, Gaussian processes, Kalman filters, LDA/QDA classifiers and Gaussian anomaly detectors.

How is it used?

Estimate $\hat{\boldsymbol\mu} = \bar{\mathbf{x}}$ and $\hat\Sigma$; evaluate with scipy.stats.multivariate_normal(mean, cov).logpdf(x); sample with rng.multivariate_normal(mean, cov, size); in NumPyro use dist.MultivariateNormal(loc, scale_tril=L). Work with log-densities to avoid underflow.

Rotate by dragging the background. Drag the purple handle on the floor to move the mean. Raise $\rho$: the hill turns into a tilted ridge. Shrink both $\sigma$'s: the hill gets narrower and taller (the volume under it is always 1). The purple rings are horizontal slices: the 1-sd ellipse is the slice at 61% of the peak height, the 2-sd ellipse the slice at 13.5%. Press Top to see that the slices are ellipses.

Each dot is one draw from $N(\mathbf{0}, \Sigma)$. The three purple ellipses are the 1-, 2- and 3-sd contours. Read the counts: only about 39% of the points fall inside the 1-sd ellipse, not the 68% you know from 1D. Change $\sigma$'s and $\rho$: the percentages do not move, because they depend only on the dimension. Press New sample to see the wobble, and raise $n$ to see $\hat\Sigma$ settle near $\Sigma$.

"The 1-sd ellipse holds 68% of the data, like ±1 sd in one dimension."

In 2D the squared distance $q$ follows a chi-square distribution with 2 degrees of freedom, so the $k$-sd ellipse holds $1 - e^{-k^2/2}$: 39.3% for $k=1$, 86.5% for $k=2$, 98.9% for $k=3$. A 95% ellipse needs $k = 2.45$. In 3D the 1-sd ellipsoid holds only 19.9%. More dimensions, more room outside.

"If each variable is Normal, the pair is a bivariate Normal."

Not necessarily. Take $X\sim N(0,1)$ and $Y = SX$ with $S = \pm1$ a fair coin flip independent of $X$. $Y$ is also $N(0,1)$ and $Cov(X, Y) = 0$, yet $|Y| = |X|$ always: they are strongly dependent and the pair is not bivariate Normal (all the points sit on two diagonal lines).

"SciPy's multivariate_normal takes standard deviations, like norm does."

scipy.stats.norm(loc, scale) takes the sd, but multivariate_normal(mean, cov) takes the covariance matrix (variances on the diagonal). NumPyro is the same: Normal(loc, scale) takes an sd; MultivariateNormal takes covariance_matrix or scale_tril (a Cholesky factor, not a list of sds).

Every Gaussian variational guide is a multivariate Normal over the model's latent vector, in unconstrained space (NumPyro first maps constrained parameters, such as a positive noise scale, to the whole real line, for example with a log). The mean-field guide (AutoNormal) is $N(\boldsymbol\mu, \Sigma)$ with a diagonal $\Sigma$; AutoMultivariateNormal uses a full $\Sigma$; AutoLowRankMultivariateNormal uses a low-rank-plus-diagonal $\Sigma$. So everything in this chapter (ellipses, eigenvalues, Mahalanobis distance) describes the shape of the approximate posterior your SVI loop is fitting.

$N(\boldsymbol\mu,\Sigma)$: $p(\mathbf{x}) \propto |\Sigma|^{-1/2}\exp(-\frac12(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu))$. Contours = ellipses.

Marginals, conditionals and linear combinations are Normal; zero covariance ⟺ independence (for the MVN only).

2D: $k$-sd ellipse holds $1 - e^{-k^2/2}$ (39%, 86%, 99%). Trap: SciPy/NumPyro multivariate versions take the covariance (or its Cholesky), not sds.

Quick check: $\mathbf{X} \sim N(\mathbf{0}, \Sigma)$ with $\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$. What is the distribution of $X_1 + X_2$?

A linear combination of a multivariate Normal is Normal, with mean 0 and variance $\mathbf{a}^\top\Sigma\mathbf{a} = 5 + 2 + 2\cdot2 = 11$ for $\mathbf{a} = (1,1)$. So $X_1 + X_2 \sim N(0, 11)$, standard deviation $\sqrt{11}\approx 3.32$.

Eigenvectors and eigenvalues of $\Sigma$: the axes of the ellipse core

A tilted ellipse has a long axis and a short axis, and they meet at a right angle. Along the long axis the cloud is most spread out; along the short axis it is least spread out. Those two directions are the eigenvectors of the covariance matrix, and the variance along each of them is its eigenvalue.

Better still: if you turn your coordinate axes to line up with these directions, the two new coordinates are uncorrelated. The tilt was only a consequence of looking at the cloud along the "wrong" axes. You already met this picture in the Linear Algebra guide (a symmetric matrix stretches a circle into an ellipse along its eigenvectors, 1.11); here the stretching matrix is $\Sigma^{1/2}$ and the circle is a cloud of independent standard-Normal points.

Three ways to say it:

  • Picture: start from a round cloud, stretch it by $\sqrt{\lambda_1}$ along one direction and by $\sqrt{\lambda_2}$ along the perpendicular one, and you get the tilted ellipse.
  • Numbers: $\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$ has eigenvalues 6 and 1, with directions $(2,1)/\sqrt5$ and $(1,-2)/\sqrt5$: the long axis has sd $\sqrt6 \approx 2.45$, the short one sd 1.
  • Slogan: eigenvectors say where the cloud points; eigenvalues say how much it spreads there.
  1. Characteristic equation: $\det(\Sigma - \lambda I) = (5-\lambda)(2-\lambda) - 2\cdot 2 = \lambda^2 - 7\lambda + 6 = (\lambda - 6)(\lambda - 1) = 0$. So $\lambda_1 = 6$, $\lambda_2 = 1$.
  2. Eigenvector for 6: $(5-6)v_1 + 2v_2 = 0 \Rightarrow v_2 = v_1/2$, so $\mathbf{v}_1 = (2, 1)/\sqrt5 \approx (0.894, 0.447)$. Check: $\Sigma(2,1)^\top = (10 + 2,\ 4 + 2) = (12, 6) = 6\cdot(2, 1)$. ✓
  3. Eigenvector for 1: $\mathbf{v}_2 = (1, -2)/\sqrt5$. Check: $\Sigma(1,-2)^\top = (5-4,\ 2-4) = (1, -2)$. ✓ The two are perpendicular: $2\cdot1 + 1\cdot(-2) = 0$.
  4. Angle of the long axis: $\arctan(1/2) \approx 26.6°$. Half-lengths of the 1-sd ellipse: $\sqrt6 \approx 2.45$ and $\sqrt1 = 1$.
  5. Bookkeeping: $\lambda_1 + \lambda_2 = 7 = 5 + 2$ (the trace: total variance is preserved by the rotation); $\lambda_1\lambda_2 = 6 = |\Sigma|$ (the determinant: the "area" of the cloud).
  6. Compare with the regression line of orders on sessions: its slope is $Cov/Var(\text{sessions}) = 2/5 = 0.4$, while the long axis has slope $1/2 = 0.5$. They are different lines (more on this in the traps).

Because $\Sigma$ is symmetric, the spectral theorem (LA 1.11, 1.13) gives

$$\Sigma = V\Lambda V^\top, \qquad V = [\mathbf{v}_1 \cdots \mathbf{v}_d]\ \text{orthonormal},\quad \Lambda = \text{diag}(\lambda_1 \ge \dots \ge \lambda_d \ge 0).$$
  • The variance of the data along a unit direction $\mathbf{u}$ is $\mathbf{u}^\top\Sigma\mathbf{u}$. It is largest ($\lambda_1$) along $\mathbf{v}_1$ and smallest ($\lambda_d$) along $\mathbf{v}_d$.
  • Rotated coordinates $\mathbf{Y} = V^\top(\mathbf{X}-\boldsymbol\mu)$ have $Cov(\mathbf{Y}) = V^\top\Sigma V = \Lambda$: diagonal, so uncorrelated.
  • The $k$-sd ellipse (ellipsoid) has its axes along the $\mathbf{v}_i$ with half-lengths $k\sqrt{\lambda_i}$.
  • $\text{trace}(\Sigma) = \sum\lambda_i$ = total variance; $|\Sigma| = \prod\lambda_i$ = generalized variance; $\lambda_1/\lambda_d$ (the condition number) measures how elongated the cloud is.
  • For $2\times2$, $\Sigma = \begin{bmatrix} a & b \\ b & c\end{bmatrix}$: $\lambda = \frac{a+c}{2} \pm \sqrt{\big(\frac{a-c}{2}\big)^2 + b^2}$, and the long axis is at angle $\tfrac12\,\text{atan2}(2b,\ a - c)$.
Why do we need it?

It turns a tilted, correlated cloud into independent directions, and tells you which combinations of variables vary a lot and which barely vary at all (a near-zero eigenvalue means an almost exact linear relation).

Where is it used?

PCA (Chapter 5.16), whitening, sampling with $\boldsymbol\mu + V\Lambda^{1/2}\mathbf{z}$, the geometry of posteriors (long thin posteriors slow down MCMC, Chapter 6.10, and break mean-field guides, 6.13), and the conditioning of optimization problems.

How is it used?

lam, V = np.linalg.eigh(S) for a symmetric matrix. It returns eigenvalues in ascending order and eigenvectors as columns; reverse both for "largest first". Draw axes of length $k\sqrt{\lambda_i}$ along V[:, i].

26.6° √λ₁ = √6 ≈ 2.45 √λ₂ = 1 Σ = [[5, 2], [2, 2]] = V Λ Vᵀ v₁ = (2, 1)/√5, λ₁ = 6: long axis v₂ = (1, −2)/√5, λ₂ = 1: short axis axes meet at 90° (v₁ · v₂ = 0) half-length = k · √λ (sd, not variance) trace 7 = 6 + 1 (total variance) det 6 = 6 × 1 (area of the cloud) in rotated axes the two coordinates are uncorrelated: Cov = diag(6, 1)
The 1-sd ellipse of the example. Its axes point along the eigenvectors of $\Sigma$, and each half-length is the square root of the matching eigenvalue (a standard deviation).

Set $\sigma_1$, $\sigma_2$ and $\rho$. The green and teal arrows are the eigenvectors, drawn with length $\sqrt{\lambda}$ (one standard deviation along each axis). Press Stretch a circle to watch a unit circle being stretched along the eigenvectors into the 1-sd ellipse. Try the presets: with $\rho = 0$ the axes line up with $x$ and $y$; with equal $\sigma$'s and $\rho \ne 0$ they sit at 45°; at $\rho = 0.99$ the short axis almost vanishes (nearly singular).

250 draws from a 3-variable Normal (sds 1.4, 1.0, 0.7). Drag the background to rotate; drag the purple handle to move the mean (Shift moves it up and down). The wire ellipsoid is the 2-sd surface; the arrows are the eigenvectors with length $2\sqrt{\lambda}$. Raise all three correlations: the cloud becomes a cigar (one big eigenvalue). Press Impossible correlations (0.9, 0.9, −0.9): the cloud disappears, because no data can have that matrix. Look from Top and Front: each view shows a 2D ellipse.

"The long axis of the ellipse is the regression line."

The regression line of $y$ on $x$ minimizes vertical distances and has slope $Cov/Var(x) = 0.4$ here. The long axis minimizes perpendicular distances (it is the first principal component) and has slope 0.5. They agree only when the cloud is a perfect line.

"The eigenvalues are the standard deviations along the axes."

They are variances. The half-lengths of the ellipse use $\sqrt{\lambda_i}$.

"$\rho = 0.7$ always gives a 45° ellipse."

Only when the two variances are equal. With unequal variances the long axis leans toward the variable with the larger variance.

"np.linalg.eigh gives the biggest eigenvalue first."

It returns them in ascending order (our example prints [1., 6.]), with eigenvectors as columns. Reverse both, and remember that the sign of an eigenvector is arbitrary.

In your forecasting model, parameters that can explain the same feature of the data trade off against each other: the base slope $k$ and the first changepoint adjustments, the offset $m$ and the trend, two overlapping regressors. Their posterior is a long, thin, tilted ellipse. The eigenvector with the large eigenvalue is the poorly determined combination (for two overlapping regressors, roughly their difference); the one with the small eigenvalue is the well determined one (roughly their sum). A large ratio $\lambda_1/\lambda_d$ is what makes SVI and NUTS slow, and it is exactly the tilt a mean-field guide cannot represent.

$\Sigma = V\Lambda V^\top$: columns of $V$ = axis directions; $\lambda_i$ = variance along axis $i$; half-axis = $k\sqrt{\lambda_i}$.

Rotated coordinates $V^\top(\mathbf{x}-\boldsymbol\mu)$ are uncorrelated. trace = total variance, det = product of eigenvalues.

2×2: $\lambda = \frac{a+c}2 \pm\sqrt{(\frac{a-c}2)^2 + b^2}$. Traps: eigenvalues are variances; eigh is ascending; long axis ≠ regression line.

Quick check: $\Sigma = \begin{bmatrix} 9 & 0 \\ 0 & 4 \end{bmatrix}$. What are the axes and half-lengths of the 1-sd ellipse?

The matrix is diagonal, so the eigenvectors are the coordinate axes, with eigenvalues 9 and 4. The ellipse is not tilted: half-length $\sqrt9 = 3$ along $x$ and $\sqrt4 = 2$ along $y$.

The SVD of the centred data matrix: the same axes, a safer road core

So far the road to the axes was: data → covariance matrix → eigenvectors. There is a second road that never builds the covariance matrix at all: take the singular value decomposition (SVD) of the centred data table itself. It gives the same axes, and its "stretch factors" (singular values) give the spreads.

Why bother with a second road? Building $X_c^\top X_c$ multiplies the data by itself, which squares every number's scale and loses floating-point precision for nearly-flat clouds. The SVD works on the data directly. It also copes easily with "wide" data (more variables than rows). That is why scikit-learn's PCA normally uses it (recent versions switch to the covariance road for very tall, narrow data, where it is faster and still accurate).

The SVD itself is taught in the Linear Algebra guide ("rotate, stretch, rotate", 1.13). Here we only need one fact about it.

Three ways to say it:

  • Picture: two roads from the same data table arrive at the same pair of axes.
  • Numbers: the example's centred table has singular values $4.899$ and $2$; squared and divided by $n-1 = 4$: $24/4 = 6$ and $4/4 = 1$, the eigenvalues again.
  • Slogan: eigenvectors of the covariance = right singular vectors of the centred data.

Centred data (from the first section): $X_c = \begin{bmatrix} -3 & -2 \\ -1 & 0 \\ 0 & 0 \\ 1 & 2 \\ 3 & 0 \end{bmatrix}$.

  1. $X_c^\top X_c = \begin{bmatrix} 20 & 8 \\ 8 & 8 \end{bmatrix}$ (the sums of squares and products we computed earlier). Its eigenvalues are 24 and 4 (four times 6 and 1).
  2. The SVD $X_c = USV^\top$ therefore has singular values $s_1 = \sqrt{24} \approx 4.899$, $s_2 = \sqrt4 = 2$ (NumPy prints [4.89897949, 2.]).
  3. The rows of NumPy's Vt are $(0.894, 0.447)$ and $(0.447, -0.894)$: exactly $\mathbf{v}_1 = (2,1)/\sqrt5$ and $\mathbf{v}_2 = (1, -2)/\sqrt5$.
  4. Variances along the axes: $s_i^2/(n-1)$: $24/4 = 6$ and $4/4 = 1$. Same as the eigenvalues of $\hat\Sigma$.
  5. Bonus: $X_c V = US$ gives each day's coordinates along the axes (its "scores"). Day 1: $(-3,-2)\cdot(2,1)/\sqrt5 = -8/\sqrt5 \approx -3.58$ along the long axis.

Thin SVD of the centred $n\times d$ data matrix: $X_c = U S V^\top$, with $U$ ($n\times d$, orthonormal columns), $S = \text{diag}(s_1 \ge \dots \ge s_d \ge 0)$, $V$ ($d\times d$, orthogonal). Then

$$\hat\Sigma = \frac{X_c^\top X_c}{n-1} = \frac{V S U^\top U S V^\top}{n-1} = V\,\frac{S^2}{n-1}\,V^\top,$$

because $U^\top U = I$. This is an eigen-decomposition of $\hat\Sigma$. So:

  • the right singular vectors (columns of $V$) are the eigenvectors of $\hat\Sigma$ (up to sign);
  • $\lambda_i = s_i^2/(n-1)$;
  • the coordinates of the data along the axes are $X_cV = US$;
  • the number of non-zero $s_i$ is the rank of $X_c$, at most $\min(n-1, d)$. With $d \gt n$, most eigenvalues of $\hat\Sigma$ are exactly 0.
Why do we need it?

It is the numerically stable way to get the axes: forming $X_c^\top X_c$ squares the condition number ($\text{cond}(X^\top X) = \text{cond}(X)^2$), which destroys precision for nearly flat clouds. The SVD also works when there are more variables than rows.

Where is it used?

scikit-learn's PCA (full or randomized SVD solvers), latent semantic analysis of text, recommender-system factorization, low-rank approximation (Eckart–Young, LA 1.13), and finding the main directions of a cloud of posterior draws.

How is it used?

Xc = X - X.mean(0); U, s, Vt = np.linalg.svd(Xc, full_matrices=False). The axes are the rows of Vt; the variances are s**2 / (n - 1); the scores are U * s. Centre first.

data Xn × d centre X_ccolumns mean 0 road 1: Σ̂ = X_cᵀX_c/(n−1)then eigh(Σ̂) → V, λ road 2: svd(X_c) = U S Vᵀno X_cᵀX_c needed same axes Vλᵢ = sᵢ²/(n−1)
Two roads to the axes of a data cloud. Road 1 builds the covariance matrix and takes its eigenvectors; road 2 takes the SVD of the centred data. They agree up to the sign of each vector; road 2 is the numerically safer one.

Drag the five days. The thick green arrows come from road 1 (eigenvectors of $\hat\Sigma$, length $\sqrt\lambda$); the dashed orange arrows come from road 2 (right singular vectors of $X_c$, length $s/\sqrt{n-1}$). They always lie on the same lines. Sometimes an orange arrow points the opposite way: the sign of a singular vector is arbitrary. Press Flatten onto a line: $s_2$ and $\lambda_2$ drop to 0 and $\hat\Sigma$ becomes singular.

"U, s, Vt = np.linalg.svd(Xc): the axes are the columns of Vt."

NumPy returns $V^\top$, so the axes are the rows of Vt (or the columns of Vt.T). The covariance eigenvectors from eigh are columns. Mixing these up gives wrong directions silently.

"The singular values are the variances along the axes."

Variances are $s_i^2/(n-1)$. Singular values grow with $\sqrt n$; variances do not.

"SVD of the raw data gives the same axes."

Only of the centred data. Without centring, the first singular vector points from the origin toward the mean of the cloud (Chapter 5.16 shows it).

"PCA is an eigen-decomposition, SVD is something else."

"For centred data they give the same principal directions: the right singular vectors of $X_c$ are the eigenvectors of $\hat\Sigma$, and $\lambda_i = s_i^2/(n-1)$."

Model answer: "$X_c = USV^\top$ implies $X_c^\top X_c = VS^2V^\top$. Libraries mostly use the SVD because it never forms $X_c^\top X_c$, which would square the condition number, and because randomized SVD scales to big, wide data."

$X_c = USV^\top \Rightarrow \hat\Sigma = V\frac{S^2}{n-1}V^\top$: axes = right singular vectors, $\lambda_i = s_i^2/(n-1)$, scores $= X_cV = US$.

Rank of $X_c$ = number of non-zero $s_i \le \min(n-1, d)$.

Traps: centre first; NumPy returns Vt (rows); signs are arbitrary.

Quick check: a centred $101 \times 3$ data matrix has singular values 20, 10 and 0. What are the variances along the three axes, and what does the 0 tell you?

$s^2/(n-1)$ with $n - 1 = 100$: $400/100 = 4$, $100/100 = 1$ and $0$. A zero singular value means the centred cloud lies exactly in a plane: one combination of the three variables never varies, so $\hat\Sigma$ is singular.

Mahalanobis distance: how unusual is this point, given the cloud's shape? core

Two days are both "2.24 away from the average day" by ruler. Day A is far along the direction where days normally vary a lot (many sessions and many orders). Day B is far in a direction where days almost never go (more sessions but fewer orders). Day B is much stranger, but the ruler cannot tell.

The Mahalanobis distance measures distance in units of the cloud's own spread, direction by direction: "how many standard deviations away, along the cloud's axes". Its circles of equal distance are the cloud's own ellipses. In one dimension it is just the absolute z-score.

Three ways to say it:

  • Picture: replace the ruler's circles by the cloud's ellipses.
  • Numbers: A = (12, 7) and B = (11, 4) are both $\sqrt5 \approx 2.24$ from the mean (10, 6) by ruler, but 0.91 vs 2.24 by Mahalanobis.
  • Slogan: a z-score that respects correlations.

Mean $\boldsymbol\mu = (10, 6)$, $\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$, $\Sigma^{-1} = \frac16\begin{bmatrix} 2 & -2 \\ -2 & 5 \end{bmatrix}$.

  1. Deviations: $A - \boldsymbol\mu = (2, 1)$, $B - \boldsymbol\mu = (1, -2)$. Ruler (Euclidean) distances: both $\sqrt{4 + 1} = \sqrt5 \approx 2.236$.
  2. $d_M^2(A) = \frac16(2\cdot2^2 - 4\cdot2\cdot1 + 5\cdot1^2) = \frac16(8 - 8 + 5) = 0.833$, so $d_M(A) \approx 0.913$.
  3. $d_M^2(B) = \frac16(2\cdot1 - 4\cdot1\cdot(-2) + 5\cdot4) = \frac16(2 + 8 + 20) = 5$, so $d_M(B) \approx 2.236$.
  4. Eigen view: A lies exactly on the long axis ($\lambda_1 = 6$): $d_M^2 = 5/6$. B lies exactly on the short axis ($\lambda_2 = 1$): $d_M^2 = 5/1 = 5$. Each squared ruler length is divided by the variance in its direction.
  5. How rare? If days are bivariate Normal, $d_M^2$ follows a chi-square distribution with 2 degrees of freedom, whose tail is $P(d_M^2 \ge c) = e^{-c/2}$. For A: $e^{-0.417} \approx 0.66$ (ordinary). For B: $e^{-2.5} \approx 0.082$ (unusual, but not extreme).

The Mahalanobis distance of a point $\mathbf{x}$ from a distribution with mean $\boldsymbol\mu$ and invertible covariance $\Sigma$ is

$$d_M(\mathbf{x}) = \sqrt{(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)} = \sqrt{\sum_{i=1}^d \frac{\big(\mathbf{v}_i^\top(\mathbf{x}-\boldsymbol\mu)\big)^2}{\lambda_i}}.$$
  • The second form says: rotate to the cloud's axes, divide each coordinate by its standard deviation $\sqrt{\lambda_i}$, then use the ordinary length. That rotate-and-rescale step is called whitening (Chapter 5.16): after it, the cloud is round and Mahalanobis distance is plain Euclidean distance.
  • In 1D: $d_M = |x - \mu|/\sigma = |z|$.
  • It is exactly the quantity in the exponent of the multivariate Normal density. If $\mathbf{X}\sim N(\boldsymbol\mu,\Sigma)$ in $d$ dimensions, $d_M^2 \sim \chi^2_d$. Outlier thresholds come from its quantiles: $\chi^2_{2,\,0.99} = 9.21$ ($d_M \gt 3.03$); $\chi^2_{10,\,0.95} = 18.3$ ($d_M \gt 4.28$).
  • With estimated $\hat{\boldsymbol\mu}$ and $\hat\Sigma$ the chi-square is only approximate, and outliers inflate $\hat\Sigma$, which hides them ("masking"). Robust covariance estimates fix this.
Why do we need it?

"Unusual" depends on the shape of normal variation. A distance that ignores correlations and units flags ordinary days and misses strange combinations. Mahalanobis distance is scale-free and correlation-aware.

Where is it used?

Multivariate outlier and anomaly detection (scikit-learn's EllipticEnvelope, MinCovDet), the exponent of the Gaussian density and so every Gaussian log-likelihood, Hotelling's $T^2$ test, LDA/QDA classifiers, and Mahalanobis matching in causal inference.

How is it used?

scipy.spatial.distance.mahalanobis(u, v, VI) where VI is the inverse covariance, or solve with a Cholesky factor instead of inverting. Compare $d_M^2$ with a $\chi^2_d$ quantile (scipy.stats.chi2(d).ppf(0.99)). Use a robust covariance if outliers may be in the data.

Left: 200 typical days (as deviations from the average day) and two draggable days, A (orange) and B (purple). Grey dashed circles are ruler distances; the solid ellipses through A and B are Mahalanobis distances. Drag A and B to the same ruler circle in different directions and compare. Right: the same data after whitening: the cloud becomes round and each ellipse becomes a circle, so Mahalanobis distance is now just length. Change $\rho$ to see how the shape of normal variation changes the verdict.

"scipy.spatial.distance.mahalanobis(u, v, S) takes the covariance matrix."

Its third argument is VI, the inverse covariance. Passing $\Sigma$ instead of $\Sigma^{-1}$ gives a wrong number with no warning.

"Standardize each column, then use Euclidean distance: that is Mahalanobis."

Only if the variables are uncorrelated. With correlation you must also rotate to the cloud's axes (whiten); standardizing columns alone leaves the ellipse tilted.

"$d_M \gt 2$ is an outlier, like $|z| \gt 2$."

The threshold grows with the dimension, because $d_M^2 \sim \chi^2_d$. In 2D the 95% cut-off is $d_M = 2.45$; with 10 variables it is 4.28. Use chi-square quantiles, not 1D habits.

In an A/B framework like yours you could monitor each day's vector of guardrail metrics (sessions, conversion rate, revenue per user) with a Mahalanobis distance from the recent normal days: a day can be strange as a combination (many sessions but few orders) while every metric alone looks normal. In the forecasting model, $-\tfrac12 d_M^2$ is literally the data-dependent part of the log-density of a full-rank Gaussian guide: the guide scores a parameter vector by its Mahalanobis distance from the guide's mean.

$d_M^2 = (\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu) = \sum_i(\text{coordinate along }\mathbf{v}_i)^2/\lambda_i$. 1D: $|z|$.

For an MVN, $d_M^2\sim\chi^2_d$ (2D tail: $e^{-c/2}$; 99% cut-off 9.21). Euclidean after whitening.

Traps: SciPy wants $\Sigma^{-1}$; thresholds grow with $d$; outliers inflate $\hat\Sigma$ (masking).

Quick check: $\Sigma = \text{diag}(4, 1)$, $\boldsymbol\mu = \mathbf{0}$, $\mathbf{x} = (2, 1)$. Ruler distance and Mahalanobis distance?

Ruler: $\sqrt{4 + 1} = \sqrt5 \approx 2.24$. Mahalanobis: $d_M^2 = 2^2/4 + 1^2/1 = 1 + 1 = 2$, so $d_M \approx 1.41$. The first coordinate is "cheap" because that variable has variance 4.

Covariance structure: which correlations do you allow? core

A full covariance matrix for $d$ variables has $d(d+1)/2$ free numbers: 3 for two variables, 55 for ten, 500 500 for a thousand. You rarely have the data to estimate that many numbers well, or the memory and time to work with them. So in practice we assume, or discover, a pattern in the matrix. That pattern is the covariance structure.

  • Independent (diagonal): no correlations at all.
  • All pairs alike (compound symmetry): every pair has the same correlation, as when all users in one segment share a common segment effect.
  • Blocks: variables form groups that are correlated inside the group but not across groups.
  • Neighbours (banded): each variable is linked mostly to its neighbours, like consecutive days or consecutive changepoints; the link fades with distance.
  • A few shared drivers (low-rank + diagonal, a "factor model"): a few hidden quantities push many variables up and down together, and each variable also has its own private noise.

The eigenvalues reveal the structure. A few shared drivers show up as a few large eigenvalues sticking out above a flat floor.

Three ways to say it:

  • Picture: the pattern of colours in the heatmap of $\Sigma$.
  • Numbers: for $d = 100$: full 5 050 numbers, diagonal 100, one shared driver + diagonal 200.
  • Slogan: structure = which correlations you let the model learn, and how many numbers that costs.

One shared driver. Three metrics each respond to the same hidden "traffic level" $F \sim N(0,1)$ with strengths (loadings) $\mathbf{w} = (2, 1, 1)$, plus independent private noise of variance 1: $X_i = w_iF + e_i$.

  1. Variances: $Var(X_i) = w_i^2 Var(F) + Var(e_i) = w_i^2 + 1$: that is $5, 2, 2$.
  2. Covariances: $Cov(X_i, X_j) = w_iw_j\,Var(F) = w_iw_j$ (the private noises are independent): $Cov(X_1,X_2) = 2$, $Cov(X_1, X_3) = 2$, $Cov(X_2,X_3) = 1$.
  3. So $\Sigma = \mathbf{w}\mathbf{w}^\top + I = \begin{bmatrix} 5 & 2 & 2 \\ 2 & 2 & 1 \\ 2 & 1 & 2 \end{bmatrix}$ (the top-left block is our example matrix).
  4. Eigenvalues: $\mathbf{w}\mathbf{w}^\top$ has one eigenvalue $\|\mathbf{w}\|^2 = 4 + 1 + 1 = 6$ (direction $\mathbf{w}$) and two zeros; adding $I$ adds 1 to each: $\lambda = 7, 1, 1$. One eigenvalue sticks out of a flat floor of 1s: the fingerprint of one shared driver.
  5. Counting: this structure needs $3$ loadings $+\ 3$ private variances $= 6$ numbers, the same as a full $3\times3$ matrix. At $d = 100$ it needs 200 instead of 5 050.

Common structures for a $d\times d$ covariance matrix and their number of free parameters (counting the variances):

StructureFormFree numbersEigenvalue fingerprint
Full (unstructured)any PSD $\Sigma$$d(d+1)/2$anything
Diagonal$\text{diag}(\sigma_1^2,\dots,\sigma_d^2)$$d$the variances themselves
Compound symmetry$\Sigma_{ij} = \rho\,\sigma_i\sigma_j$ for all pairs$d + 1$one large, the rest equal
Block-diagonalindependent groups, full inside eachsum over blocks of $b(b+1)/2$each block's own spectrum
Banded / AR(1)-like$\Sigma_{ij} = \sigma_i\sigma_j\,\phi^{|i-j|}$$d + 1$smooth decay, no gap
Low-rank + diagonal (factor model)$WW^\top + \text{diag}(\boldsymbol\psi)$, $W$ is $d\times r$about $d(r+1)$$r$ large, then a floor

(The low-rank count is "about" because rotating the $r$ columns of $W$ gives the same $WW^\top$, so a few of those numbers are redundant.) Also note: with few rows, the sample eigenvalues are more spread out than the true ones: the largest come out too large and the smallest too small, even when the truth is diagonal.

Why do we need it?

Fewer free numbers means less data needed, less memory and faster computation, and a structure that matches how the data were generated gives better estimates than a noisy full matrix (the bias–variance trade-off of Chapter 5.1).

Where is it used?

Mixed and hierarchical models (a shared group effect gives compound symmetry), time-series models (AR structures, Chapter 7.3), factor models of asset returns, Gaussian-process kernels, Ledoit–Wolf shrinkage (sklearn.covariance.LedoitWolf) when $d$ is large relative to $n$, and variational guides (next section).

How is it used?

Look at the heatmap of the sample correlation matrix and at its eigenvalue spectrum. A few large eigenvalues above a floor suggest low-rank + diagonal; smooth decay along the diagonal suggests banded; blocks suggest groups. With $d$ close to or above $n$, shrink $\hat\Sigma$ toward a structured target instead of trusting it raw.

Pick a structure and a strength. Left: the $8\times8$ correlation matrix (blue = positive). Right: its eigenvalues, largest first (they always add up to 8). Compare one shared driver (one tall bar, then a floor) with neighbours (a smooth slope). Then turn on estimate it from n rows: with $n = 20$, even the independent structure shows fake correlations and its eigenvalues spread out (orange). Raise $n$ to watch the estimate settle.

"A full covariance matrix is always better: it assumes nothing."

With little data, a full $\hat\Sigma$ is noisy: its largest eigenvalues come out too large, its smallest too small, and it shows correlations that are not there. A structured or shrunk estimate is often closer to the truth (bias–variance again).

"If $d \gt n$ I can still use np.cov and invert it."

Then $\hat\Sigma$ has rank at most $n - 1$, so it is singular and has no inverse. Use a structure (diagonal, low-rank + diagonal) or shrinkage.

Structure = the allowed pattern in $\Sigma$. Counts: full $d(d+1)/2$, diagonal $d$, compound symmetry and AR(1) $d+1$, low-rank $r$ + diagonal ≈ $d(r+1)$.

Spectrum fingerprint: $r$ big eigenvalues + flat floor → low-rank + diagonal fits.

Trap: small $n$ spreads the sample eigenvalues and invents correlations.

Quick check: $d = 20$ variables. How many free numbers for a full covariance, a diagonal one, and a rank-2 + diagonal one?

Full: $20\cdot21/2 = 210$. Diagonal: 20. Rank-2 + diagonal: $20\cdot2 + 20 = 60$ (slightly fewer are truly free, because of the rotation freedom of $W$).

Why covariance structure decides between mean-field, low-rank and full-rank guides core

Your SVI loop approximates the posterior of the model's $d$ latent parameters with a multivariate Normal (in unconstrained space). A multivariate Normal is a mean vector plus a covariance matrix. So choosing a Gaussian guide is choosing a covariance structure:

  • Mean-field (AutoNormal): a diagonal covariance. Every parameter gets its own spread; no correlations at all.
  • Full-rank (AutoMultivariateNormal): a full covariance. Any correlation between any two parameters can be represented.
  • Low-rank (AutoLowRankMultivariateNormal): low-rank + diagonal. Correlations along $r$ learned directions (a few "shared drivers"), plus a private spread for each parameter.

Posteriors of real models are correlated: parameters that can explain the same feature of the data trade off. If the guide cannot represent a correlation, it gets the uncertainty wrong: a mean-field guide typically makes each parameter look more certain than it is, and it gets the uncertainty of combinations (sums, differences, forecasts that add many components) wrong. But a full covariance costs $d(d+1)/2$ numbers. The low-rank guide is the middle road.

Three ways to say it:

  • Picture: mean-field = an ellipse that may only stretch along the axes; full-rank = any tilt in any direction; low-rank = a few tilted "sticks" plus an axis-aligned blob.
  • Numbers: for $d = 1000$ latent parameters: mean-field 2 000 numbers, low-rank with $r = 10$: 12 000, full-rank: 501 500.
  • Slogan: give the guide as much covariance structure as the posterior needs and as little as the model size allows.

Parameter counts (checked by initializing each NumPyro guide on a model with $d$ latent dimensions and counting the learned numbers).

  1. $d = 10$, $r = 3$. Mean-field: a location and a scale per parameter, $2d = 20$.
  2. Low-rank: auto_loc ($d = 10$) + auto_cov_factor ($d\times r = 30$) + auto_scale ($d = 10$): $d(r + 2) = 50$.
  3. Full-rank: auto_loc ($d = 10$) + the lower triangle of auto_scale_tril ($d(d+1)/2 = 55$): $65$. (NumPyro stores scale_tril as a full $10\times10$ array, 100 cells, but only the 55 lower-triangular entries are learned.)
  4. Now $d = 1000$, $r = 10$: mean-field $2\,000$; low-rank $1000\times12 = 12\,000$; full-rank $1000 + 500\,500 = 501\,500$. Storing a $1000\times1000$ scale_tril takes 4 MB in float32; at $d = 20\,000$ it would take 1.6 GB.
  5. Work per step: drawing one sample from the full-rank guide multiplies by a $d\times d$ triangular matrix (about $d^2$ operations); the low-rank guide needs about $d\cdot r$. Anything that factorizes or inverts a general $d\times d$ covariance grows like $d^3$.

The three Gaussian guide families, as covariance structures over the unconstrained latent vector $\boldsymbol\theta\in\mathbb{R}^d$:

GuideCovariance of $q$Learned numbers (NumPyro)CapturesMisses
Mean-field AutoNormal$\text{diag}(\boldsymbol\sigma^2)$$2d$each parameter's own spreadevery correlation
Low-rank AutoLowRankMultivariateNormal$WW^\top + \text{diag}(\boldsymbol\psi)$, $W$: $d\times r$$d(r+2)$correlations along $r$ directionspatterns needing more than $r$ directions (e.g. long chains of neighbour correlations)
Full-rank AutoMultivariateNormal$LL^\top$, $L$ lower-triangular$d + d(d+1)/2$any correlation patternnothing Gaussian; still not skew, heavy tails, several modes or funnels

(NumPyro builds the low-rank covariance from cov_factor and scale as $\text{diag}(\mathbf{s})(FF^\top + I)\,\text{diag}(\mathbf{s})$, which is of the form $WW^\top + \text{diag}(\boldsymbol\psi)$.) The parameter counts and fitting of these guides are taught in Chapter 6.13; here the point is the statistics: the guide family is a covariance structure, and the posterior's eigenvalue spectrum tells you which structure is enough.

Why do we need it?

The guide's covariance structure decides which posterior correlations survive into your uncertainty estimates, and it decides memory and speed. Choosing it blindly gives either over-confident intervals (too little structure) or a model that does not fit in memory (too much).

Where is it used?

NumPyro and Pyro autoguides (AutoNormal, AutoLowRankMultivariateNormal, AutoMultivariateNormal), Laplace approximations (a full Gaussian at the mode), and your forecasting model's automatic full-rank vs low-rank choice by model size.

How is it used?

Count $d$ after flattening every latent site. If $d$ is small, full-rank is affordable and safest. If $d$ is large, use low-rank and pick $r$ from how many large eigenvalues the posterior correlation matrix has (fit a small version with full-rank, or run NUTS on a subset, and look at its spectrum: the scree idea of Chapter 5.16).

mean-field low-rank + diagonal full-rank 2d numbersd = 1000: 2 000work ∝ d d(r + 2) numbersd = 1000, r = 10: 12 000work ∝ d·r d + d(d + 1)/2 numbersd = 1000: 501 500memory ∝ d², factorizing ∝ d³
The three Gaussian guides are three covariance structures. Mean-field keeps only the diagonal; low-rank adds a few broad shared patterns ($WW^\top$) on top of the diagonal; full-rank can put a different correlation in every cell.

Pick a "true posterior correlation" for 8 parameters and a rank $r$. The middle heatmap is the best low-rank + diagonal matrix built by a simple recipe (keep the top $r$ eigen-directions, then fix the diagonal); the right one is the error. $r = 0$ is the mean-field picture: every correlation is lost. With one shared driver, $r = 1$ already captures almost everything. With neighbours, you need a high rank, because that pattern has no small set of dominant directions. Read the counts: the price of each choice at $d = 8$ and at $d = 1000$.

"A low-rank guide is a cheaper mean-field guide."

It has more parameters than mean-field ($d(r+2)$ vs $2d$) and sits between mean-field and full-rank: it adds $r$ correlation directions on top of the diagonal.

"A full-rank guide gives the exact posterior."

It is still a Gaussian approximation in unconstrained space, fitted by a noisy optimizer. It cannot represent skewness, heavy tails, several modes or funnel shapes. It only removes the "no correlations" or "only $r$ directions" restriction.

"The guide's correlations are correlations of the parameters as I wrote them."

They live in the unconstrained space the autoguide works in (for example $\log\sigma$ instead of $\sigma$), after flattening all latent sites into one vector.

This is the statistics behind your forecasting model's automatic guide choice. The latent dimension $d$ grows with every changepoint on the grid (one $\delta_j$ each), every Fourier pair ($2N$ coefficients per seasonality), every holiday and every regressor. When $d$ is small, a full-rank guide is affordable and captures every trade-off (slope vs changepoint adjustments, overlapping regressors, seasonality vs holiday effects). When $d$ is large, the $d^2$ memory and work of full-rank becomes the bottleneck, and a low-rank guide keeps the strongest correlation directions at a cost that grows like $d\cdot r$. The size threshold and the rank in your code are design choices; the honest way to defend them is "the posterior correlation spectrum is dominated by a few directions, so rank $r$ captures most of it", checked against a full-rank or NUTS fit on a smaller version.

"Low-rank guides are approximate, full-rank guides are exact."

"Both are Gaussian approximations. They differ in the covariance structure they allow: full-rank any $\Sigma = LL^\top$; low-rank $WW^\top + \text{diag}(\boldsymbol\psi)$ with $r$ directions; mean-field only a diagonal."

Model answer: "A full-rank Gaussian guide learns $d + d(d+1)/2$ numbers and its memory and per-step work grow like $d^2$, so it stops scaling for large models. A low-rank-plus-diagonal guide learns $d(r+2)$ numbers and captures the $r$ most important posterior correlation directions plus each parameter's own spread. I choose by model size: full-rank when $d$ is small, low-rank when $d$ is large, with $r$ set by how many large eigenvalues the posterior correlation has; I validate against full-rank or NUTS on a reduced model."

Guide family = covariance structure: mean-field diag ($2d$); low-rank $WW^\top + \text{diag}$ ($d(r+2)$ in NumPyro); full-rank $LL^\top$ ($d + d(d+1)/2$).

Correlated posteriors → mean-field over-confident; low-rank keeps $r$ directions; full-rank keeps all but costs $\propto d^2$ memory.

Trap: all three are Gaussians in unconstrained space; none is "exact".

Quick check: $d = 50$, $r = 5$. How many numbers does each NumPyro guide learn?

Mean-field: $2\cdot50 = 100$. Low-rank: $50\cdot(5+2) = 350$. Full-rank: $50 + 50\cdot51/2 = 50 + 1275 = 1325$.

Recap, cheat sheet and practice

  • A random vector is several numbers drawn together; the data matrix has one row per draw. The columns (marginals) do not determine how the variables go together.
  • The covariance matrix $\Sigma$ holds every variance (diagonal) and covariance (off-diagonal); $\hat\Sigma = X_c^\top X_c/(n-1)$; $Var(\mathbf{a}^\top\mathbf{X}) = \mathbf{a}^\top\Sigma\mathbf{a}$ and $Cov(A\mathbf{X}+\mathbf{b}) = A\Sigma A^\top$.
  • Every covariance matrix is symmetric and PSD (no direction has negative variance); PD = invertible = Cholesky exists. Pairwise-valid correlations can still be jointly impossible.
  • The correlation matrix $R = D^{-1/2}\Sigma D^{-1/2}$ is unit-free; $\Sigma$ and its eigenvectors change with units.
  • The multivariate Normal is fixed by $\boldsymbol\mu$ and $\Sigma$; its contours are ellipses; in 2D the $k$-sd ellipse holds $1 - e^{-k^2/2}$ (39%, 86%, 99%).
  • Eigenvectors of $\Sigma$ are the ellipse's axes; eigenvalues are the variances along them; the SVD of the centred data gives the same axes with $\lambda_i = s_i^2/(n-1)$.
  • Mahalanobis distance measures distance in the cloud's own standard deviations; $d_M^2\sim\chi^2_d$ for Normal data.
  • Covariance structure (diagonal, low-rank + diagonal, full…) sets how many numbers you must learn; the three Gaussian guides are exactly these three structures.

Cheat sheet

IdeaFormulaIn words
Covariance matrix$\Sigma = E[(\mathbf{X}-\boldsymbol\mu)(\mathbf{X}-\boldsymbol\mu)^\top]$, $\hat\Sigma = X_c^\top X_c/(n-1)$all variances and covariances
Linear map$Cov(A\mathbf{X}+\mathbf{b}) = A\Sigma A^\top$; $Var(\mathbf{a}^\top\mathbf{X}) = \mathbf{a}^\top\Sigma\mathbf{a}$variance of any combination
Valid?symmetric, all eigenvalues $\ge 0$; 2×2: $Cov^2 \le Var_1Var_2$no negative variance anywhere
Correlation matrix$R = D^{-1/2}\Sigma D^{-1/2}$unit-free covariances
Multivariate Normal$p(\mathbf{x}) \propto |\Sigma|^{-1/2}e^{-\frac12(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)}$a hill with elliptical contours
Ellipse coverage (2D)$P(\text{inside } k\text{-sd}) = 1 - e^{-k^2/2}$39%, 86%, 99% for $k = 1, 2, 3$
Eigen-decomposition$\Sigma = V\Lambda V^\top$; half-axes $k\sqrt{\lambda_i}$axes and spreads of the cloud
SVD route$X_c = USV^\top$, $\lambda_i = s_i^2/(n-1)$same axes, numerically safer
Mahalanobis$d_M^2 = (\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)\sim\chi^2_d$a z-score that respects correlations
Guide sizes (NumPyro)$2d$ · $d(r+2)$ · $d + d(d+1)/2$mean-field · low-rank · full-rank
Code it · Python

import numpy as np
from scipy import stats
from scipy.spatial.distance import mahalanobis

# 1) The covariance matrix of five days (sessions in thousands, orders in hundreds)
X = np.array([[7, 4], [9, 6], [10, 6], [11, 8], [13, 6]], dtype=float)   # rows = days
S = np.cov(X, rowvar=False)            # rowvar=False: columns are the variables; divides by n-1
print(S)                               # [[5. 2.] [2. 2.]]
print(np.cov(X).shape)                 # (5, 5)  <- the classic mistake: rows treated as variables
Xc = X - X.mean(axis=0)
print(np.allclose(S, Xc.T @ Xc / (len(X) - 1)))   # True

# 2) Variance of a combination: a^T S a
a = np.array([1, -1])
print(a @ S @ a)                       # 3.0  (= 5 + 2 - 2*2)

# 3) Correlation matrix from the covariance matrix
sd = np.sqrt(np.diag(S))
R = S / np.outer(sd, sd)
print(R.round(3))                      # [[1. 0.632] [0.632 1.]]

# 4) Eigen-decomposition: the axes of the ellipse (eigh: ascending order, vectors in columns)
lam, V = np.linalg.eigh(S)
lam, V = lam[::-1], V[:, ::-1]         # largest first
print(lam, V[:, 0])                    # [6. 1.]  first axis +-(0.894, 0.447): the sign is arbitrary

# 5) The same axes from the SVD of the centred data
U, s, Vt = np.linalg.svd(Xc, full_matrices=False)
print(s, s**2 / (len(X) - 1))          # [4.899 2.] [6. 1.]
print(Vt[0])                           # [0.894 0.447]  (rows of Vt are the axes)

# 6) Is a matrix a valid covariance? Look at the smallest eigenvalue
bad = np.array([[1, .9, .9], [.9, 1, -.9], [.9, -.9, 1]])
print(np.linalg.eigvalsh(bad).min())   # -0.8  -> impossible
print(np.linalg.eigvalsh(S).min())     # 1.0   -> positive definite

# 7) Mahalanobis distance: SciPy wants the INVERSE covariance
mu, VI = X.mean(axis=0), np.linalg.inv(S)
for name, x in [("A", [12, 7]), ("B", [11, 4])]:
    d = mahalanobis(x, mu, VI)
    print(name, round(np.linalg.norm(np.subtract(x, mu)), 3), round(d, 3), round(stats.chi2(2).sf(d**2), 3))
# A 2.236 0.913 0.659
# B 2.236 2.236 0.082

# 8) Multivariate Normal: only ~39% of draws fall inside the 1-sd ellipse in 2D
rng = np.random.default_rng(0)
Z = rng.multivariate_normal(mean=[0, 0], cov=S, size=200_000)   # takes the COVARIANCE, not sds
q = np.einsum("ij,jk,ik->i", Z, np.linalg.inv(S), Z)             # squared Mahalanobis distances
print([float(np.mean(q <= k * k).round(3)) for k in (1, 2, 3)])  # [0.392, 0.865, 0.989]  theory 0.393, 0.865, 0.989
print(stats.multivariate_normal([0, 0], S).pdf([2, 1]).round(4)) # 0.0428

# 9) Guide sizes for d latent parameters (NumPyro counts)
for d, r in [(10, 3), (1000, 10)]:
    print(d, "mean-field", 2 * d, "low-rank", d * (r + 2), "full-rank", d + d * (d + 1) // 2)
# 10 mean-field 20 low-rank 50 full-rank 65
# 1000 mean-field 2000 low-rank 12000 full-rank 501500
Test yourself

1. $\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$ for $(X_1, X_2)$. What is $Var(X_1 - X_2)$?

$\mathbf{a} = (1, -1)$: $\mathbf{a}^\top\Sigma\mathbf{a} = 5 + 2 - 2\cdot2 = 3$. 7 forgets the covariance term; 11 is the variance of the sum.

2. Which of these cannot be a covariance matrix?

$Cov^2 = 25 \gt 4\cdot4 = 16$: the direction $(1, -1)$ would have variance $-2$ (eigenvalues 9 and −1). The second matrix has $Cov^2 = 9 = 9\cdot1$: valid but singular (a flat cloud); the last is valid too (the second variable is constant).

3. For a bivariate Normal, roughly what fraction of draws falls inside the 1-sd ellipse?

The squared Mahalanobis distance is $\chi^2_2$, so $P(d_M^2 \le 1) = 1 - e^{-1/2} \approx 0.393$. The 68% rule is for one dimension.

4. A 2D covariance matrix has eigenvalues 6 and 1. What are the half-lengths of its 2-sd ellipse?

Eigenvalues are variances; a half-length is $k\sqrt{\lambda}$. With $k = 2$: $2\sqrt6 \approx 4.90$ and $2\sqrt1 = 2$.

5. In scipy.spatial.distance.mahalanobis(u, v, VI), what is VI?

SciPy computes $\sqrt{(u - v)^\top VI\,(u - v)}$, so you must pass $\Sigma^{-1}$. Passing $\Sigma$ silently gives a wrong distance.

6. A model has $d = 200$ latent parameters. Which NumPyro guide learns exactly 1 200 numbers?

Low-rank: $d(r+2) = 200\cdot6 = 1200$. AutoNormal learns $2d = 400$; AutoMultivariateNormal $200 + 200\cdot201/2 = 20\,300$; rank 10 gives $2\,400$.

Practice problems

A. Three days: (2, 1), (4, 5), (6, 3). Compute $\hat\Sigma$, the correlation, the eigenvalues and the axes.
  1. Mean: $(12/3, 9/3) = (4, 3)$. Centred rows: $(-2,-2), (0, 2), (2, 0)$.
  2. Sums: $\sum x^2 = 4 + 0 + 4 = 8$, $\sum y^2 = 4 + 4 + 0 = 8$, $\sum xy = 4 + 0 + 0 = 4$. Divide by $n - 1 = 2$: $\hat\Sigma = \begin{bmatrix} 4 & 2 \\ 2 & 4 \end{bmatrix}$.
  3. Correlation: $2/(2\cdot2) = 0.5$.
  4. Eigenvalues: $\frac{4+4}{2} \pm\sqrt{0 + 4} = 4 \pm 2$: 6 and 2. Equal variances, so the axes are at 45°: $(1,1)/\sqrt2$ (variance 6) and $(1,-1)/\sqrt2$ (variance 2).
B. Use "a variance is never negative" to prove that a correlation is always between −1 and 1.

Let $Z_1 = X/\sigma_X$ and $Z_2 = Y/\sigma_Y$, each with variance 1 and covariance $\rho$. Then $Var(Z_1 - Z_2) = 1 + 1 - 2\rho = 2(1 - \rho) \ge 0$, so $\rho \le 1$; and $Var(Z_1 + Z_2) = 2(1 + \rho) \ge 0$, so $\rho \ge -1$. This is the 2×2 PSD condition, written with $\mathbf{a} = (1/\sigma_X, \mp1/\sigma_Y)$.

C. $\Sigma = \begin{bmatrix} 4 & 2 \\ 2 & 1 \end{bmatrix}$. Find the eigenvalues. What is special about this cloud, and what breaks?

Trace 5, determinant $4 - 4 = 0$, so $\lambda = 5$ and $0$. The eigenvector for 0 is $(1, -2)/\sqrt5$: $Var(X_1 - 2X_2) = 4 - 2\cdot2\cdot2 + 4\cdot1 = 0$. So $X_1 - 2X_2$ is a constant: the two variables lie exactly on a line ($\rho = 2/(2\cdot1) = 1$). $\Sigma$ is PSD but not PD: it has no inverse, the Normal density and the Mahalanobis distance are not defined, and Cholesky fails. Drop one variable, or add a small jitter if the exact relation is only a rounding effect.

D. $\boldsymbol\mu = \mathbf{0}$, $\Sigma = \begin{bmatrix} 4 & 2 \\ 2 & 4 \end{bmatrix}$. Compute the Mahalanobis distances of $(2, 2)$ and $(2, -2)$. Which point is more unusual?

$\Sigma^{-1} = \frac1{12}\begin{bmatrix} 4 & -2 \\ -2 & 4 \end{bmatrix}$. For $(2,2)$: $\frac{1}{12}(16 - 16 + 16) = 4/3$, $d_M \approx 1.15$. For $(2,-2)$: $\frac1{12}(16 + 16 + 16) = 4$, $d_M = 2$. Check with eigenvalues: $(2,2)$ lies on the long axis ($\lambda = 6$): $8/6 = 4/3$; $(2,-2)$ on the short axis ($\lambda = 2$): $8/2 = 4$. Both are $\sqrt8 \approx 2.83$ from the centre by ruler. Tail probabilities $e^{-d^2/2}$: 0.51 vs 0.14, so $(2,-2)$ is more unusual.

E. With $\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$ (sessions, orders as deviations, bivariate Normal), a day has orders deviation $+2$. What do you expect for sessions, and how uncertain?

Conditional mean: $\mu_1 + \frac{\Sigma_{12}}{\Sigma_{22}}(x_2 - \mu_2) = 0 + \frac22\cdot2 = 2$. Conditional variance: $\Sigma_{11} - \Sigma_{12}^2/\Sigma_{22} = 5 - 4/2 = 3$ (sd $\approx 1.73$ instead of $\sqrt5 \approx 2.24$). Knowing orders shifts the guess and shrinks the uncertainty; the stronger the correlation, the more it shrinks.

F. (Interview) "Your forecasting model switches from a full-rank to a low-rank Gaussian guide when the model gets big. Explain why, and what you lose."

"Each guide is a Gaussian over the flattened, unconstrained latent vector of dimension $d$; they differ only in covariance structure. Full-rank learns $d + d(d+1)/2$ numbers through a Cholesky factor, so memory and per-step work grow like $d^2$: fine with a few dozen parameters, a bottleneck with thousands (many changepoints, Fourier terms, holidays, regressors). Low-rank learns $WW^\top + \text{diag}$ with $d(r+2)$ numbers, cost about $d\cdot r$. It keeps the $r$ most important correlation directions (the big trade-offs, such as slope vs changepoint adjustments or overlapping regressors) and a private variance per parameter. What I lose: correlation patterns that need more than $r$ directions. I check that the posterior correlation spectrum is dominated by a few eigenvalues, and compare against full-rank or NUTS on a smaller version."

Chapter 5.16 · Syllabus Module 21, Part 26

Principal component analysis (PCA)

You have many columns, and many of them move together. PCA finds the few directions in which your data really varies, lets you describe each row with a handful of numbers instead of dozens, and tells you exactly how much you lose. It is the covariance-matrix picture of Chapter 5.15 put to work: the long axes of the ellipse become new, uncorrelated variables. This chapter builds it from one question ("which direction shows the most spread?"), then covers the practical traps: centring, scaling, choosing $k$, and what PCA cannot do.

  • Explain PCA as finding the direction of maximum variance, and see why that is the same as the direction of smallest reconstruction error
  • Derive that the principal components are the eigenvectors of the covariance matrix and the eigenvalues are their variances; know why the components are orthogonal and their scores uncorrelated
  • Read explained variance, a scree plot and the cumulative curve, and choose $k$ with honest rules of thumb
  • Compute scores, projections, reconstructions and the reconstruction error by hand
  • Know what goes wrong without centring, and when to standardize (covariance vs correlation PCA)
  • Run PCA through the SVD, and read scikit-learn's output, including the sign convention
  • Use PCA for dimensionality reduction, noise removal, multicollinearity (principal component regression) and whitening
  • Explain why PCA is linear and its limits: curved data, variance ≠ predictive importance, sign ambiguity, outliers

What we need from earlier chapters: the covariance matrix, its eigenvectors as the axes of the data ellipse, the SVD of the centred data and Mahalanobis distance (Chapter 5.15); standardization (Chapter 4.18); the correlation matrix (Chapter 4.15). From the Linear Algebra guide: projection onto a line and a subspace (1.9), eigenvectors of symmetric matrices (1.11), SVD and low-rank approximation (1.13), and the short PCA preview in 1.17. Notation: data matrix $X$ ($n\times d$, rows = observations), column means $\bar{\mathbf{x}}$, centred data $X_c$, sample covariance $\hat\Sigma$ with eigenpairs $(\lambda_j, \mathbf{v}_j)$, $\lambda_1 \ge \lambda_2 \ge \dots$; $V_k = [\mathbf{v}_1 \cdots \mathbf{v}_k]$ holds the first $k$ components.

The idea: find the direction in which the data spreads the most core

Hold a torch over a cloud of points and look at the shadow on a wall. Turn the torch around: in some directions the shadow is wide, in others narrow. The wide shadow keeps the most information about how the points differ from each other; the narrow one squashes many different points onto almost the same spot.

PCA asks: which direction gives the widest shadow? That direction is the first principal component (PC1). The best direction at right angles to it is PC2, and so on.

There is a second, equivalent way to see it. Replacing each point by its shadow loses the part that sticks out sideways (the perpendicular distance to the line). By Pythagoras, "shadow spread + sideways loss" is the same for every direction (it is the total spread of the cloud). So the direction with the widest shadow is also the direction with the smallest loss.

Three ways to say it:

  • Picture: turn the line until the shadows of the points on it are as spread out as possible.
  • Numbers: for our five days the shadow variance is 5 along the sessions axis, 5.5 at 45°, and the maximum 6 at 26.6°; the sideways loss is 2, 1.5 and 1: always 7 in total.
  • Slogan: keep the direction where the points differ the most; it is also where you lose the least.

The five days of Chapter 5.15 (sessions, orders) have $\hat\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$, total variance $5 + 2 = 7$. The variance of the shadow on a unit direction $\mathbf{u}$ is $\mathbf{u}^\top\hat\Sigma\mathbf{u}$.

  1. Along the sessions axis, $\mathbf{u} = (1, 0)$: shadow variance $= 5$. Sideways loss (the variance of the perpendicular part) $= 7 - 5 = 2$.
  2. At 45°, $\mathbf{u} = (1, 1)/\sqrt2$: $\frac12(5 + 2\cdot2 + 2) = 5.5$. Loss $= 1.5$.
  3. At 26.6°, $\mathbf{u} = (2, 1)/\sqrt5$: $\frac15(5\cdot4 + 2\cdot2\cdot2\cdot1 + 2\cdot1) = \frac15(20 + 8 + 2) = 6$. Loss $= 1$.
  4. No direction beats 6 (you will prove it in the next section: 6 is the largest eigenvalue). So PC1 $= (2,1)/\sqrt5$, and it keeps $6/7 \approx 85.7\%$ of the total variance.
  5. PC2 must be at right angles: $(1, -2)/\sqrt5$, with variance $7 - 6 = 1$.

For centred data with sample covariance $\hat\Sigma$, the first principal component is the unit vector

$$\mathbf{v}_1 = \arg\max_{\|\mathbf{u}\| = 1}\ \mathbf{u}^\top\hat\Sigma\,\mathbf{u} \;=\; \arg\min_{\|\mathbf{u}\| = 1}\ \sum_{i=1}^n \big\|(\mathbf{x}_i - \bar{\mathbf{x}}) - \big(\mathbf{u}^\top(\mathbf{x}_i - \bar{\mathbf{x}})\big)\mathbf{u}\big\|^2 .$$
  • The two problems have the same answer because, for every unit $\mathbf{u}$ (Pythagoras, point by point): $\|\mathbf{x}_i - \bar{\mathbf{x}}\|^2 = (\text{shadow length})^2 + (\text{sideways distance})^2$, and the left side does not depend on $\mathbf{u}$.
  • The $j$-th component $\mathbf{v}_j$ maximizes the same quantity among unit vectors orthogonal to $\mathbf{v}_1, \dots, \mathbf{v}_{j-1}$.
  • The number $z_{ij} = \mathbf{v}_j^\top(\mathbf{x}_i - \bar{\mathbf{x}})$ (the signed shadow length) is the score of observation $i$ on component $j$. The entries of $\mathbf{v}_j$ are called loadings: how much each original variable contributes to the component.
  • "Sideways distance" means the perpendicular distance to the line, not the vertical distance used by regression.
Why do we need it?

With many correlated columns, much of the data is repetition. PCA finds a few new variables (directions) that keep most of the differences between rows, so you can plot, compress, de-noise or regress on them instead of on dozens of tangled columns.

Where is it used?

Exploratory plots of high-dimensional data, compressing embeddings and images, removing noise from signals, principal component regression for collinear features, face recognition ("eigenfaces"), genetics (population structure), finance (yield-curve factors), and as a first step before clustering or t-SNE (Chapter 5.17).

How is it used?

PCA(n_components=k).fit(X) in scikit-learn, then transform for scores. Look at explained_variance_ratio_ to see how much each direction keeps and at components_ (the loadings) to see what each direction means.

mean a point xᵢ shadow (score) sideways loss for every direction of the line: ‖xᵢ − x̄‖² (fixed, does not depend on the line) = shadow² + sideways² summed over all points: total variance = variance along + loss so: maximum spread along the line = minimum loss (same direction)
Each point splits into a shadow along the line (green) and a sideways part (red) at a right angle. Their squares always add up to the fixed squared distance from the mean, so widening the shadows and shrinking the losses are the same goal.

Drag the purple handle to turn the line through the centre of 40 days. Green dots are the shadows (projections); red segments are the sideways losses. The right plot shows, for every angle, the shadow variance (blue) and the loss (red); their sum (dashed) never changes. Find the angle where blue is highest: red is lowest at exactly the same angle. Press Snap to PC1 to check, and Snap to PC2 for the worst direction.

"PC1 is the regression line through the cloud."

Regression of $y$ on $x$ minimizes vertical errors and treats $y$ as special. PCA minimizes perpendicular distances and treats all variables alike. For our days the regression slope is 0.4, PC1's slope is 0.5.

"PCA picks some of the original columns and drops the others."

Each component is a new variable: a weighted mix of all the columns (the loadings say how much of each). Selecting columns is feature selection, a different idea.

PC1 = $\arg\max_{\|\mathbf{u}\|=1}\mathbf{u}^\top\hat\Sigma\mathbf{u}$ = the line with the smallest perpendicular loss (Pythagoras: along + loss = total).

Score $z_{ij} = \mathbf{v}_j^\top(\mathbf{x}_i - \bar{\mathbf{x}})$; loadings = entries of $\mathbf{v}_j$.

Trap: perpendicular distances, not vertical (PC1 ≠ regression line).

Quick check: with $\hat\Sigma = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$, what is the shadow variance along the orders axis, and the loss?

$\mathbf{u} = (0, 1)$: $\mathbf{u}^\top\hat\Sigma\mathbf{u} = 2$. The loss is $7 - 2 = 5$. That is a poor direction: it keeps only $2/7 \approx 29\%$ of the variance.

Principal components are eigenvectors; eigenvalues are their variances core

We want the direction with the widest shadow. In Chapter 5.15 we saw that the long axis of the data ellipse is an eigenvector of the covariance matrix and that the variance along it is the eigenvalue. It is no coincidence: the "widest shadow" direction is exactly that long axis. So PCA needs nothing new: the principal components are the eigenvectors of $\hat\Sigma$, sorted by eigenvalue.

Two gifts come with it. The components are at right angles to each other (orthogonal), and the scores on different components are uncorrelated: in the new coordinates the ellipse is no longer tilted.

Three ways to say it:

  • Picture: PCA turns the axes of the plot until they line up with the axes of the data ellipse.
  • Numbers: PC1 $= (0.894, 0.447)$ with variance 6, PC2 $= (0.447, -0.894)$ with variance 1; the scores have covariance exactly 0.
  • Slogan: eigenvectors give the new axes, eigenvalues tell you how much each axis is worth.

Why the best direction is an eigenvector (a short derivation with a Lagrange multiplier, Optimization 3.8).

  1. Maximize $f(\mathbf{u}) = \mathbf{u}^\top\hat\Sigma\mathbf{u}$ subject to $\mathbf{u}^\top\mathbf{u} = 1$. Lagrangian: $\mathcal{L} = \mathbf{u}^\top\hat\Sigma\mathbf{u} - \lambda(\mathbf{u}^\top\mathbf{u} - 1)$.
  2. Set the gradient to zero: $2\hat\Sigma\mathbf{u} - 2\lambda\mathbf{u} = \mathbf{0}$, so $\hat\Sigma\mathbf{u} = \lambda\mathbf{u}$. Every candidate is an eigenvector.
  3. At an eigenvector the objective is $\mathbf{u}^\top\hat\Sigma\mathbf{u} = \lambda\,\mathbf{u}^\top\mathbf{u} = \lambda$. So the best one is the eigenvector with the largest eigenvalue, and the variance it captures is that eigenvalue.

The five days, component by component. Centred rows $(-3,-2), (-1,0), (0,0), (1,2), (3,0)$; $\mathbf{v}_1 = (2,1)/\sqrt5$, $\mathbf{v}_2 = (1,-2)/\sqrt5$.

  1. Scores on PC1, $z_1 = (2x_1 + x_2)/\sqrt5$: $-8/\sqrt5, -2/\sqrt5, 0, 4/\sqrt5, 6/\sqrt5 \approx -3.58, -0.89, 0, 1.79, 2.68$.
  2. Scores on PC2, $z_2 = (x_1 - 2x_2)/\sqrt5$: $1/\sqrt5, -1/\sqrt5, 0, -3/\sqrt5, 3/\sqrt5 \approx 0.45, -0.45, 0, -1.34, 1.34$.
  3. Variance of $z_1$: $(64 + 4 + 0 + 16 + 36)/5 = 24$, divided by $n-1 = 4$: $6 = \lambda_1$. Variance of $z_2$: $(1 + 1 + 0 + 9 + 9)/5 = 4$, divided by 4: $1 = \lambda_2$.
  4. Covariance of the scores: $\frac15\big[(-8)(1) + (-2)(-1) + 0 + (4)(-3) + (6)(3)\big] = \frac15(-8 + 2 - 12 + 18) = 0$. Uncorrelated.
  5. Explained variance: PC1 keeps $6/7 = 85.7\%$, PC2 keeps $1/7 = 14.3\%$.

PCA of a centred data matrix: write $\hat\Sigma = V\Lambda V^\top$ with $\lambda_1 \ge \dots \ge \lambda_d \ge 0$. Then

  • the principal components (directions) are the columns $\mathbf{v}_1, \dots, \mathbf{v}_d$ of $V$: unit length, mutually orthogonal ($\mathbf{v}_i^\top\mathbf{v}_j = 0$ for $i\ne j$);
  • the scores are $Z = X_cV$ (row $i$ = coordinates of observation $i$ in the new axes);
  • $Var(z_j) = \lambda_j$ and $Cov(z_i, z_j) = 0$ for $i \ne j$, because $Cov(Z) = V^\top\hat\Sigma V = \Lambda$;
  • $\sum_j\lambda_j = \text{trace}(\hat\Sigma)$: the rotation keeps the total variance, it only redistributes it, as unequally as possible.

Each $\mathbf{v}_j$ is defined only up to its sign: $-\mathbf{v}_j$ is just as good (the scores flip sign too). If two eigenvalues are equal, the directions inside their plane are not unique at all.

Why do we need it?

It turns "search over all directions" into a standard, fast computation (an eigen-decomposition or an SVD), and it delivers new variables that are uncorrelated, which is exactly what makes them easy to keep, drop or regress on.

Where is it used?

Every PCA implementation (scikit-learn, R's prcomp, Spark MLlib), uncorrelated inputs for regression and clustering, the "eigen-portfolios" of finance, and reading the main directions of a posterior covariance (the directions a low-rank guide should capture, 5.15).

How is it used?

lam, V = np.linalg.eigh(np.cov(X, rowvar=False)), reverse to largest first, scores Xc @ V. Read the loadings of each component to name it (for example "overall busyness" for a component with all-positive loadings).

Left: the five days (drag them) with PC1 (green) and PC2 (teal) through the mean. Right: the same days drawn in the new coordinates (score on PC1 across, score on PC2 up). Whatever you do on the left, the right cloud is never tilted: the scores are uncorrelated, and their variances are the eigenvalues. Drag a day far along PC1 and watch only $\lambda_1$ grow; drag it across, and $\lambda_2$ grows.

"The components found today and the components found on new data should have the same signs."

The sign of each eigenvector is arbitrary; different libraries, versions or samples can flip it. Compare components up to sign (for example by the absolute value of $\mathbf{v}^\top\mathbf{v}'$), and never hard-code "positive score means busy" without fixing a sign convention.

"Uncorrelated scores are independent."

PCA only removes linear correlation. The scores are independent only in special cases such as multivariate Normal data. Curved relationships survive the rotation.

If you take posterior draws of your forecasting model's parameters (from NUTS or from a full-rank guide) and run PCA on them, the top components are the parameter combinations the data pins down least (the big trade-offs), and the eigenvalue spectrum tells you how many such directions there are. That is exactly the information a low-rank guide needs: its $r$ columns of $W$ play the role of the top $r$ principal directions of the posterior covariance (Chapter 5.15, Chapter 6.13).

Maximize $\mathbf{u}^\top\hat\Sigma\mathbf{u}$ with $\|\mathbf{u}\| = 1$ ⟹ $\hat\Sigma\mathbf{u} = \lambda\mathbf{u}$: PCs = eigenvectors, sorted by eigenvalue.

$Var(z_j) = \lambda_j$, scores uncorrelated, PCs orthogonal, $\sum\lambda_j$ = total variance.

Traps: signs are arbitrary; uncorrelated ≠ independent.

Quick check: three variables have covariance eigenvalues 8, 1.5 and 0.5. What fraction does PC1 explain, and what is the variance of the PC2 scores?

Total $= 10$, so PC1 explains $8/10 = 80\%$. The PC2 scores have variance $\lambda_2 = 1.5$ (15% of the total).

Explained variance, the scree plot, and how many components to keep core

Each component carries a share of the total variance: its eigenvalue divided by the sum of all eigenvalues. Plot these shares from largest to smallest and you get a scree plot (named after the pile of loose rocks at the foot of a cliff: a steep drop, then rubble). The steep part is structure; the flat rubble is mostly noise.

Add the shares up as you go and you get the cumulative explained variance: "with $k$ components I keep this percentage of the spread". Choosing $k$ is a judgement: there is no single correct rule, only useful habits.

Three ways to say it:

  • Picture: a cliff of tall bars, then a flat scree; keep the cliff.
  • Numbers: eigenvalues 2.8, 1.2, 0.5, 0.3, 0.2 (total 5) explain 56%, 24%, 10%, 6%, 4%; two components keep 80%, three keep 90%.
  • Slogan: keep the components that carry structure, drop the ones that carry noise.

A correlation-matrix PCA of 5 standardized metrics (so the eigenvalues add up to 5) gives $\lambda = 2.8, 1.2, 0.5, 0.3, 0.2$.

  1. Explained shares: $2.8/5 = 56\%$, $1.2/5 = 24\%$, $0.5/5 = 10\%$, $0.3/5 = 6\%$, $0.2/5 = 4\%$.
  2. Cumulative: 56%, 80%, 90%, 96%, 100%.
  3. Rule "keep 90%": $k = 3$. Rule "elbow" (where the bars stop dropping steeply): between PC2 and PC3, so $k = 2$.
  4. Kaiser's rule for correlation PCA (keep eigenvalues above 1, the variance of one original standardized variable): $k = 2$.
  5. The rules disagree (2 or 3). That is normal. The final choice should depend on the use: for a 2D plot $k = 2$; for compression with a quality target, the percentage rule; for prediction, cross-validation of the downstream model.
  • Explained variance ratio of component $j$: $\dfrac{\lambda_j}{\sum_{i=1}^d\lambda_i}$. Cumulative for the first $k$: $\dfrac{\sum_{j\le k}\lambda_j}{\sum_i\lambda_i}$.
  • Scree plot: $\lambda_j$ (or the ratio) against $j$.
  • Common rules for $k$ (all rules of thumb): a cumulative target (80%, 90%, 95%); the elbow of the scree plot; Kaiser's "eigenvalue $\gt 1$" for correlation PCA (known to keep too many components when $d$ is large); and, best when PCA feeds a model, choosing $k$ by cross-validated performance.
  • With noisy data, even pure noise produces unequal sample eigenvalues (5.15), so the tail of the scree plot is never perfectly flat.
Why do we need it?

To decide how many new variables are enough, and to state the price honestly ("these 3 components keep 90% of the variance"). Without it, $k$ is a guess.

Where is it used?

Every PCA report; PCA(n_components=0.9) in scikit-learn chooses $k$ automatically by cumulative ratio; factor analysis; checking how many strong directions a posterior covariance has before picking a low-rank guide's rank.

How is it used?

Fit PCA() with all components, plot explained_variance_ratio_ and its cumsum(), look for an elbow and a cumulative target, then refit with the chosen n_components. If PCA feeds a model, pick $k$ by validation error.

PC1PC2PC3PC4PC5 56%24%10% 90% target → k = 3 cumulative: 56, 80, 90, 96, 100% the "cliff": PC1, PC2 (structure) the "scree": PC3 onward (mostly noise) elbow → k = 2 Kaiser (λ > 1, correlation PCA) → k = 2
A scree plot (blue bars: share of variance per component) with the cumulative curve (orange). Different rules of thumb point to $k = 2$ or $k = 3$; the use of the components decides.

Ten standardized metrics are generated from a few hidden drivers plus noise. Set the true number of drivers and the noise level. With low noise the scree plot has a clear cliff exactly after the true number of drivers; raise the noise and the cliff erodes into the scree. Move the target to see how many components a "keep X%" rule picks, and compare with Kaiser's rule (bars above the dashed 10% line, i.e. eigenvalue above 1). Press New sample to see how much the tail wobbles.

"Keep components until 95% is explained; that is the correct $k$."

All $k$ rules are rules of thumb. With noisy data a high percentage target keeps noise components; with a downstream model, choose $k$ by validation error.

"PC1 explains 85% of the variance, so it explains 85% of my target."

Explained variance is about the spread of the inputs $X$. It says nothing about how well the components predict a target $y$ (see the limitations section).

Explained ratio $\lambda_j/\sum\lambda_i$; cumulative $\sum_{j\le k}\lambda_j/\sum\lambda_i$; scree plot = $\lambda_j$ vs $j$.

Rules of thumb: % target, elbow, Kaiser ($\lambda \gt 1$ on correlation PCA); best: validate the downstream use.

Trap: explained variance of $X$ ≠ explained variance of $y$.

Quick check: eigenvalues 4, 3, 2, 1. How many components for 80% of the variance?

Total 10; cumulative 40%, 70%, 90%, 100%. The first $k$ reaching 80% is $k = 3$.

Projection, reconstruction and reconstruction error core

Keeping $k$ components means describing each row by $k$ scores instead of $d$ original numbers. Geometrically, every point is dropped straight onto the best flat $k$-dimensional sheet through the mean: a line for $k = 1$, a plane for $k = 2$. That drop is a projection.

To go back, rebuild each point from its $k$ scores: start at the mean and walk score × component along each kept direction. The rebuilt point (the reconstruction) lies on the sheet; the gap to the real point is what you threw away. Add up the squared gaps and you get the reconstruction error, and it equals exactly the variance in the components you dropped (times $n - 1$).

Three ways to say it:

  • Picture: squash a thin pancake of points onto the plate it lies on; the error is how thick the pancake was.
  • Numbers: keeping PC1 for the five days, the squared errors are 0.2, 0.2, 0, 1.8, 1.8: total 4 $= (n-1)\lambda_2 = 4 \times 1$.
  • Slogan: what you drop is exactly what you lose.

Five days, keep $k = 1$ (PC1 $= (2,1)/\sqrt5 \approx (0.894, 0.447)$, mean $(10, 6)$).

  1. Day 1 is $(7, 4)$; centred $(-3, -2)$. Score: $z_1 = (-3)(0.894) + (-2)(0.447) = -8/\sqrt5 \approx -3.578$.
  2. Reconstruction: $\hat{\mathbf{x}} = \bar{\mathbf{x}} + z_1\mathbf{v}_1 = (10, 6) + (-3.578)(0.894, 0.447) = (10 - 3.2,\ 6 - 1.6) = (6.8, 4.4)$.
  3. Error: $(7, 4) - (6.8, 4.4) = (0.2, -0.4)$; squared length $0.04 + 0.16 = 0.2$. This equals the dropped score squared: $z_2 = 1/\sqrt5$, $z_2^2 = 0.2$.
  4. All five reconstructions (scikit-learn's inverse_transform prints the same): $(6.8, 4.4)$, $(9.2, 5.6)$, $(10, 6)$, $(11.6, 6.8)$, $(12.4, 7.2)$. Squared errors: $0.2, 0.2, 0, 1.8, 1.8$.
  5. Total squared error $= 4 = (n-1)\lambda_2 = 4\times1$. Mean squared error per day: $4/5 = 0.8$.
  6. Storage: 5 days × 2 numbers became 5 scores + 1 direction + 1 mean. With $d = 100$ columns and $k = 5$, each row shrinks from 100 numbers to 5.

With $V_k = [\mathbf{v}_1\cdots\mathbf{v}_k]$ ($d\times k$):

$$\mathbf{z}_i = V_k^\top(\mathbf{x}_i - \bar{\mathbf{x}}) \in \mathbb{R}^k \quad\text{(scores)}, \qquad \hat{\mathbf{x}}_i = \bar{\mathbf{x}} + V_k\mathbf{z}_i = \bar{\mathbf{x}} + V_kV_k^\top(\mathbf{x}_i - \bar{\mathbf{x}}) \quad\text{(reconstruction)}.$$
  • $P = V_kV_k^\top$ is the orthogonal projection onto the span of the first $k$ components (LA 1.9); the reconstruction error vector $\mathbf{x}_i - \hat{\mathbf{x}}_i$ is perpendicular to that span.
  • Reconstruction error: $\sum_i\|\mathbf{x}_i - \hat{\mathbf{x}}_i\|^2 = (n-1)\sum_{j \gt k}\lambda_j$. The fraction of variance lost is $1 - $ cumulative explained ratio.
  • Optimality (Eckart–Young, LA 1.13): no other $k$-dimensional flat sheet (and no other rank-$k$ approximation of $X_c$) has a smaller total squared error.
Why do we need it?

Scores are the compressed data you actually use; reconstructions let you check what was lost, de-noise, and detect anomalies (a row with a large reconstruction error does not follow the usual pattern).

Where is it used?

Compression of images and embeddings, anomaly detection by reconstruction error (machine monitoring, fraud), filling in or smoothing signals, and as the linear baseline that autoencoders generalize.

How is it used?

Z = pca.transform(X), X_hat = pca.inverse_transform(Z), then ((X - X_hat)**2).sum(axis=1) per row. Plot that per-row error: rows far above the rest deserve a look.

60 days in three variables, shaped like a pancake. Rotate by dragging the background, then press Front and Side until you look along the green sheet: the pancake's thickness is exactly what PCA throws away. Red segments are the losses; green dots are the reconstructions. Switch between k = 2 (plane) and k = 1 (line) and compare the error. Make the pancake thicker with the slider. Drag the two orange extra days far from the sheet: the sheet tilts toward them, because squared errors let a few far points pull hard.

"The reconstruction $\hat{\mathbf{x}}$ is a prediction of the true, noise-free point."

It is the closest point on the kept sheet. It removes the part of the noise that lies in the dropped directions, but it keeps the noise that lies inside the sheet, and it also removes any real signal that happens to live in the dropped directions.

"PCA is robust; one strange row does not matter."

PCA minimizes squared distances, so a few far-away rows can tilt the components (drag the orange days). Look at the per-row reconstruction error and consider robust PCA variants when outliers are likely.

Scores $\mathbf{z} = V_k^\top(\mathbf{x}-\bar{\mathbf{x}})$; reconstruction $\hat{\mathbf{x}} = \bar{\mathbf{x}} + V_k\mathbf{z}$; projection matrix $V_kV_k^\top$.

$\sum_i\|\mathbf{x}_i-\hat{\mathbf{x}}_i\|^2 = (n-1)\sum_{j\gt k}\lambda_j$; PCA's sheet is the best $k$-dim fit (Eckart–Young).

Trap: squared error → outliers tilt the components.

Quick check: 101 rows, eigenvalues 9, 3, 0.5, 0.5. Keeping $k = 2$, what is the total squared reconstruction error?

$(n-1)(\lambda_3 + \lambda_4) = 100 \times (0.5 + 0.5) = 100$. The lost fraction is $1/13 \approx 7.7\%$ of the variance.

Centering: what goes wrong if you forget it core

PCA is about how points spread around their own centre. Variance is measured from the mean. If you skip the centring step and run the SVD on the raw table, the "spread" is measured from the origin $(0, 0)$ instead. When the data sits far from the origin, the biggest "spread" from the origin is simply the distance to the cloud: the first direction points from the origin to the cloud's mean, whatever the cloud's shape.

And the result looks great: the first uncentred direction "explains" 99% of the sum of squares. It is a meaningless number: it only says your data is far from zero.

Three ways to say it:

  • Picture: without centring, the first arrow points at the cloud instead of along it.
  • Numbers: for the five days, the uncentred SVD's first direction is $(0.859, 0.513)$ ("explaining" 99.4%), almost exactly the direction of the mean $(10, 6)$; the real PC1 explains 85.7%.
  • Slogan: subtract the mean first, or PCA measures where the data is instead of how it varies.
  1. Raw data $X$ (five days, not centred): $X^\top X = \begin{bmatrix} 520 & 308 \\ 308 & 188 \end{bmatrix}$ (sums of squares and products from zero).
  2. Its SVD has singular values $26.53$ and $2.03$; the first squared one is $99.4\%$ of the sum of squares.
  3. The first right singular vector is $(0.859, 0.513)$; the unit vector toward the mean $(10, 6)$ is $(10, 6)/\sqrt{136} = (0.857, 0.514)$. Nearly identical: it points at the cloud.
  4. Centred data gives PC1 $= (0.894, 0.447)$ with 85.7%. Here the two directions are only $4°$ apart because, by luck, the cloud's long axis roughly points back at the origin. In the widget below the cloud points the other way, and the uncentred answer is completely wrong.

PCA is the eigen-decomposition of the covariance $\hat\Sigma = X_c^\top X_c/(n-1)$, or equivalently the SVD of the centred matrix $X_c = X - \mathbf{1}\bar{\mathbf{x}}^\top$. The SVD of the raw $X$ decomposes the "second moment about the origin" $X^\top X/(n-1) = \hat\Sigma + \frac{n}{n-1}\bar{\mathbf{x}}\bar{\mathbf{x}}^\top$: the covariance plus a rank-one term pointing at the mean. When $\|\bar{\mathbf{x}}\|$ is large, that term dominates.

  • scikit-learn's PCA centres for you (it stores mean_ and subtracts it in transform).
  • np.linalg.svd and scikit-learn's TruncatedSVD do not centre (TruncatedSVD is meant for sparse data such as word counts, where centring would destroy sparsity).
  • New data must be centred with the training mean, not its own mean.
Why do we need it?

Without centring, the first component describes the location of the data, not its variation, and every explained-variance number is inflated. Centring is the difference between PCA and a meaningless decomposition.

Where is it used?

Every correct PCA pipeline; the mean_ attribute of scikit-learn's PCA; hand-written SVD-based PCA; and any time you reuse a fitted PCA on new rows (forecast-time data, a new experiment) where you must subtract the training mean.

How is it used?

Use PCA, which centres; if you call np.linalg.svd yourself, compute mu = X.mean(axis=0) and use X - mu. Save mu with the components and apply the same mu to new data.

The cloud's long axis runs downhill (green line through the purple mean). The red line is the first direction of the SVD of the uncentred data, always through the origin (black dot). Drag the purple handle to move the whole cloud. Far from the origin, red points at the cloud and claims to "explain" about 99%; move the cloud onto the origin (or press the button) and the two answers agree.

"I used np.linalg.svd(X), so I did PCA."

Only if X was centred first. The same goes for TruncatedSVD, which never centres.

"For new data, centre with the new data's own mean."

Use the mean saved from the training data (scikit-learn does this inside transform). Using the new batch's mean shifts every score and, in time series, can leak future information.

PCA = eigen of $\hat\Sigma$ = SVD of the centred $X_c$. Raw $X^\top X/(n-1) = \hat\Sigma + \frac{n}{n-1}\bar{\mathbf{x}}\bar{\mathbf{x}}^\top$: the extra term points at the mean.

sklearn PCA centres; np.linalg.svd and TruncatedSVD do not.

Trap: forgetting to centre gives a first direction toward the mean and an inflated "explained" share.

Quick check: your uncentred SVD says the first component explains 99.8%. What should you suspect?

That the data is far from the origin and you forgot to centre: the first direction is pointing at the mean. Centre the columns and recompute; the honest explained share will usually be much lower.

Standardization: covariance PCA vs correlation PCA core

PCA chases variance. Variance depends on units. Measure revenue in dollars (spread about 2 000) next to orders per day (spread about 40), and revenue's variance is 2 500 times bigger: PC1 will be "revenue", and it will "explain" 99.97%, simply because of the unit. Switch revenue to thousands of dollars and PC1 becomes "orders" instead. Nothing about the data changed.

The fix is to standardize each column (subtract its mean, divide by its standard deviation) before PCA. Then every column has variance 1, and PCA looks at the correlation matrix instead of the covariance matrix. This is called correlation PCA.

Three ways to say it:

  • Picture: in raw units one axis is a mile long and the other an inch; after standardizing they are the same length.
  • Numbers: revenue (sd 2 000) and orders (sd 40) with correlation 0.6: covariance PCA gives PC1 $\approx (1.000, 0.012)$, 99.97%; correlation PCA gives PC1 $= (0.707, 0.707)$, 80%.
  • Slogan: mixed units → standardize; same units and meaningful sizes → covariance PCA may be fine.
  1. Covariance matrix (revenue in dollars): $\Sigma = \begin{bmatrix} 2000^2 & 0.6\cdot2000\cdot40 \\ 0.6\cdot2000\cdot40 & 40^2 \end{bmatrix} = \begin{bmatrix} 4\,000\,000 & 48\,000 \\ 48\,000 & 1\,600 \end{bmatrix}$.
  2. Eigenvalues: about 4 000 576 and 1 024. PC1 $\approx (0.99993, 0.0120)$: "revenue only", explaining $99.97\%$.
  3. Same data, revenue in thousands of dollars (sd 2): $\Sigma = \begin{bmatrix} 4 & 48 \\ 48 & 1600 \end{bmatrix}$. Now PC1 $\approx (0.030, 0.9995)$: "orders only", 99.84%. The answer flipped because of a unit.
  4. Correlation matrix (either unit): $R = \begin{bmatrix} 1 & 0.6 \\ 0.6 & 1 \end{bmatrix}$. Eigenvalues $1 \pm 0.6 = 1.6, 0.4$; PC1 $= (1, 1)/\sqrt2$: "revenue and orders together", explaining $1.6/2 = 80\%$.
  • Covariance PCA: eigen-decomposition of $\hat\Sigma$ (centred columns, original units). Appropriate when all columns share one unit and their sizes matter (for example pixel intensities, or the same metric measured at 24 hours of the day).
  • Correlation PCA: eigen-decomposition of $R = D^{-1/2}\hat\Sigma D^{-1/2}$, i.e. PCA of the standardized columns $z_{ij} = (x_{ij} - \bar x_j)/s_j$. Appropriate for mixed units. Its eigenvalues add up to $d$.
  • The two give different components, not rescaled versions of each other (5.15).
  • In scikit-learn: make_pipeline(StandardScaler(), PCA()) gives correlation PCA; PCA() alone gives covariance PCA (it centres but does not scale).
Why do we need it?

Without it, the components report your choice of units instead of the structure of the data, and a single large-unit column hijacks PC1.

Where is it used?

PCA of business metrics, survey scores and mixed sensor readings (correlation PCA); PCA of images, spectra or hourly profiles in one unit (covariance PCA); and any model pipeline where standardization must be fitted on the training data only.

How is it used?

Ask: do all columns share a unit, and do their raw sizes matter? If not, put StandardScaler() before PCA() in a pipeline, fit it on the training rows, and apply the same scaler to new rows.

300 simulated days of four metrics: revenue, orders per day, conversion rate (%), minutes per session. The bars show the loadings of PC1 (green) and PC2 (teal). With covariance PCA and revenue in dollars, PC1 is revenue alone and claims about 100%. Switch revenue to thousands: PC1 jumps to orders. Now switch to correlation PCA: the unit switch no longer changes anything, and PC1 becomes a mix of all four metrics.

"Always standardize before PCA."

Standardize when units differ or are arbitrary. When all columns share a meaningful unit (hourly demand in orders, pixel brightness), standardizing inflates tiny, noisy columns to the same importance as the big ones; covariance PCA is then the better choice.

"Correlation PCA components are the covariance PCA components rescaled."

They are genuinely different directions with different explained shares. Report which one you used.

Your A/B framework uses one global scaler across groups for a reason that applies here too (Chapter 4.18): if you ever run PCA on metrics across segments, fit the scaling once on the pooled training data. Standardizing each segment separately would erase the very between-segment differences you want the components to show. In the forecasting model, if you compress a large set of exogenous regressors with PCA, fit the scaler and the PCA on the training window only and apply them unchanged at forecast time; fitting them on the full history leaks future information into the components (Chapter 7.12).

Covariance PCA: eigen of $\hat\Sigma$ (units matter). Correlation PCA: standardize first = eigen of $R$ (unit-free, eigenvalues sum to $d$).

Mixed units → make_pipeline(StandardScaler(), PCA()); one shared meaningful unit → plain PCA() can be right.

Trap: a large-unit column hijacks PC1 and fakes a huge explained share.

Quick check: two standardized variables have correlation $-0.5$. What are PC1, its eigenvalue and its explained share?

$R = \begin{bmatrix} 1 & -0.5 \\ -0.5 & 1 \end{bmatrix}$ has eigenvalues $1.5$ and $0.5$. PC1 $= (1, -1)/\sqrt2$ (the variables move in opposite directions), eigenvalue 1.5, explained share $1.5/2 = 75\%$.

PCA through the SVD, step by step, and how scikit-learn reports it core

In practice nobody maximizes shadows by hand. The recipe is mechanical: centre the table, take its SVD, read off everything. The SVD hands you all of PCA at once: the directions (right singular vectors), the variances (squared singular values divided by $n - 1$) and the scores ($U$ times the singular values). Chapter 5.15 showed why this agrees with the eigenvectors of the covariance matrix.

Three ways to say it:

  • Picture: centre → SVD → keep $k$ → project → (optionally) rebuild.
  • Numbers: the five days: $s = (4.899, 2)$, $\lambda = s^2/4 = (6, 1)$, PC1 keeps 85.7%, the $k = 1$ reconstruction error is 4.
  • Slogan: PCA = SVD of the centred data, plus bookkeeping.

The recipe on the five days (the widget below walks through the same steps).

  1. Centre: subtract $(10, 6)$: $X_c$ has rows $(-3,-2), (-1,0), (0,0), (1,2), (3,0)$.
  2. SVD: $X_c = USV^\top$ with $s = (4.899, 2)$ and rows of $V^\top$: $(0.894, 0.447)$ and $\pm(0.447, -0.894)$.
  3. Variances: $\lambda = s^2/(n-1) = (24/4, 4/4) = (6, 1)$; ratios $85.7\%$, $14.3\%$.
  4. Scores: $Z = X_cV = US$; first column $-3.58, -0.89, 0, 1.79, 2.68$.
  5. Keep $k = 1$ and rebuild: $\hat X = \bar{\mathbf{x}} + Z_{:,1}\mathbf{v}_1^\top$; total squared error $4 = (n-1)\lambda_2$.

scikit-learn prints, for PCA().fit(X): components_ = [[0.894, 0.447], [-0.447, 0.894]], explained_variance_ = [6, 1], explained_variance_ratio_ = [0.857, 0.143], singular_values_ = [4.899, 2], mean_ = [10, 6]. Note the second component: $(-0.447, 0.894)$, the negative of our $(0.447, -0.894)$. Same axis, other sign.

PCA via SVD: $X_c = USV^\top$ ⟹ components = columns of $V$ (rows of NumPy's Vt), $\lambda_j = s_j^2/(n-1)$, scores $Z = X_cV = US$, reconstruction $\hat X = \mathbf{1}\bar{\mathbf{x}}^\top + U_kS_kV_k^\top$.

scikit-learn's PCA conventions (version 1.9 checked here):

  • centres the data (does not scale it); mean_ is stored and reused by transform;
  • components_ has one component per row; explained_variance_ uses $n-1$;
  • n_components can be an integer $k$ or a fraction such as 0.9 (keep enough components for 90% of the variance);
  • the solver is chosen automatically: an SVD of the centred data for most shapes, a randomized SVD for big data when few components are requested, and (in recent versions) an eigen-decomposition of the covariance matrix for very tall, narrow data. All give the same components up to sign;
  • signs are fixed by a convention (currently: the entry with the largest absolute value in each component is made positive); other libraries use other conventions.
Why do we need it?

To compute PCA reliably and fast, and to read library output correctly (rows vs columns, $n$ vs $n-1$, centring vs scaling, signs). Misreading any one of them gives silently wrong components.

Where is it used?

sklearn.decomposition.PCA, IncrementalPCA for data that does not fit in memory, randomized SVD for huge matrices, np.linalg.svd in hand-written pipelines, and the same SVD in latent semantic analysis and recommender systems.

How is it used?

pca = PCA(n_components=0.9).fit(X_train); check pca.n_components_ and explained_variance_ratio_.cumsum(); Z = pca.transform(X_new); X_hat = pca.inverse_transform(Z). Fit on training rows only.

X (n × d)raw table centre(scale if mixedunits) SVDX_c = U S Vᵀ directions: Vvariances: s²/(n−1)scores: U Schoose k (scree) keep k columnsx̂ = x̄ + V_k z save the mean (and the scaler) from training; apply the same ones to new rows
The PCA pipeline. Everything after the SVD is bookkeeping: directions, variances, scores, a choice of $k$, and the optional reconstruction.

Press Next to walk through the recipe on the five days. Watch the picture and the numbers change together: centring moves the origin to the mean, the SVD finds the two axes, the scores are the shadows on PC1, and the last step rebuilds each day from one number and shows the error (total 4). Back and Reset move through the steps.

"pca.components_[:, 0] is the first component."

Components are the rows: pca.components_[0]. A column slice mixes the first loading of every component.

"scikit-learn's PCA standardizes my columns."

It only centres. For correlation PCA, put StandardScaler() in front.

"My PC2 came out with the opposite sign from my colleague's, so one of us has a bug."

Both are correct. Each component is defined up to sign; libraries and versions choose signs by different conventions.

Recipe: centre (→ scale) → SVD $X_c = USV^\top$ → components $V$, $\lambda = s^2/(n-1)$, scores $US$ → choose $k$ → $\hat X = \bar{\mathbf{x}} + Z_kV_k^\top$.

sklearn: centres not scales; components_ rows; explained_variance_ uses $n-1$; n_components=0.9 picks $k$ by share.

Trap: rows vs columns, signs, and fitting on test data.

Quick check: a centred table with 201 rows has singular values 30 and 10. What does explained_variance_ show, and the ratio for PC1?

$s^2/(n-1)$: $900/200 = 4.5$ and $100/200 = 0.5$. Ratio for PC1: $4.5/5 = 90\%$.

Dimensionality reduction and noise removal core

Think of each day's demand as 24 numbers, one per hour. Most days are variations on a few basic shapes: a midday hump, an evening peak, a morning bump. Their sizes change from day to day, but the shapes do not. So 24 numbers per day are mostly repetition: three or so numbers ("how much midday, how much evening, how much morning") would describe a day almost completely. That is dimensionality reduction.

Now add random noise to every hour. The noise is scattered evenly over all 24 directions, but the real shapes live in only a few. Keeping just the top few components keeps nearly all of the shape and only a small slice of the noise. Rebuilding the day from those components gives a cleaner curve than the one you measured: noise removal.

Three ways to say it:

  • Picture: a wobbly 24-point curve rebuilt from 3 smooth basic shapes.
  • Numbers: in one simulation with noise sd 1 per hour, the measured day is off by 1.0 per hour on average (RMS), the 3-component rebuild only by 0.40.
  • Slogan: signal lives in a few directions; noise lives in all of them; keep the few.
  1. Size. 120 days × 24 hours = 2 880 numbers. With $k = 3$: 120 × 3 scores + 3 × 24 component values + 24 means $= 360 + 72 + 24 = 456$ numbers, about 6.3 times smaller.
  2. Noise share. Independent noise with variance $\sigma^2$ in each hour has total variance $24\sigma^2$, spread evenly over all 24 directions. Projecting onto 3 directions keeps only about $3\sigma^2$ of it, so the noise left per hour is about $\sigma\sqrt{3/24} \approx 0.35\sigma$ (if the 3 components were known exactly).
  3. One simulation (noise sd $\sigma = 1$, the components estimated from the same 120 noisy days): RMS error against the clean truth is 1.00 for the raw measurements, 0.65 for $k = 1$, 0.49 for $k = 2$, 0.40 for $k = 3$, 0.48 for $k = 4$, 0.70 for $k = 8$, and back to 1.00 for $k = 24$ (keeping everything rebuilds the noise too).
  4. So there is a best $k$: too few components lose real shape (bias); too many let noise back in (variance). With louder noise (sd 2) the best $k$ in the same simulation drops to 1: weak shapes drown in the noise.
  • Dimensionality reduction: replace each $d$-dimensional row by its $k$ scores $\mathbf{z} = V_k^\top(\mathbf{x}-\bar{\mathbf{x}})$, $k \ll d$, keeping the cumulative share of variance you choose.
  • PCA de-noising: replace each row by its rank-$k$ reconstruction $\hat{\mathbf{x}} = \bar{\mathbf{x}} + V_kV_k^\top(\mathbf{x}-\bar{\mathbf{x}})$. It works when the signal is (nearly) low-dimensional and the noise is spread over many directions with similar variance.
  • The best $k$ is a bias–variance choice (Chapter 5.1): choose it by validation (for example, error on held-out days or held-out entries), not only by the explained share.
Why do we need it?

Fewer columns make models faster and less prone to overfitting, make plots possible, and cut storage; de-noising recovers clean structure from noisy measurements.

Where is it used?

Compressing embeddings before nearest-neighbour search, de-noising sensor and spectral data, image compression, summarizing daily or weekly load profiles in energy and retail, preprocessing before clustering or t-SNE, and gene-expression analysis.

How is it used?

pca = PCA(k).fit(X_train); use pca.transform(X) as features, or pca.inverse_transform(pca.transform(X)) as the de-noised data. Pick $k$ by validation error of whatever you do next.

120 simulated days of hourly orders, each a mix of three basic shapes plus noise. Pick a day and the number of components $k$. Blue dots are what you measured, the green dashed curve is the clean truth (unknown in real life), and the orange curve is the PCA rebuild. Start at $k = 0$ (just the average day), then raise $k$: the rebuild first approaches the truth, then (for large $k$) starts following the noise again. Watch "error vs truth" in the readout fall and rise. Raise the noise and see the best $k$ shrink.

"More components always give a better reconstruction, so keep as many as possible."

More components always fit the measured data better, but beyond the true signal dimension they rebuild the noise. Against the truth (or held-out data) the error is U-shaped in $k$.

"PCA de-noising removes all the noise."

It removes the noise in the dropped directions only; the share inside the kept components stays. It also struggles when the noise variance differs a lot between columns (then the top components may chase the noisiest column).

For your forecasting work, this is a useful lens on seasonality: if the hour-of-day or day-of-week profiles of demand are well described by a few principal shapes, a low-order seasonal model (a few Fourier terms, Chapter 7.11) is enough; if the scree plot shows many real components, the season is more complex. It is also a quick way to spot unusual days: a day with a large PCA reconstruction error does not follow the usual shapes (a holiday, an outage, a data problem). This is an analysis you could add, not something your model needs.

Reduce: keep $k$ scores per row. De-noise: rebuild $\hat{\mathbf{x}} = \bar{\mathbf{x}} + V_kV_k^\top(\mathbf{x}-\bar{\mathbf{x}})$.

Isotropic noise: about a share $k/d$ of its variance survives; the signal survives if it lives in the top $k$ directions.

Trap: error vs truth is U-shaped in $k$; choose $k$ by validation.

Quick check: a 50-column dataset is a 5-dimensional signal plus independent noise of sd 2 in every column. Roughly how much noise sd per column survives a perfect $k = 5$ projection?

About $2\sqrt{5/50} = 2\times0.316 \approx 0.63$ per column (the noise variance kept is about $5/50$ of the total). Real projections, estimated from noisy data, do a bit worse.

Multicollinearity and principal component regression

Temperature and "feels-like" temperature move almost together. Put both in a regression for demand and the data cannot tell which one deserves the credit: one sample says "temperature +3, feels-like −1", the next says "−1 and +3". The sum of their effects is well pinned down; their split is not (4.15, 5.13). This is multicollinearity.

PCA sees this directly: the two regressors form a long, thin cloud. PC1 is "both together" (large variance, well measured); PC2 is "their difference" (tiny variance, mostly noise). Principal component regression (PCR) regresses the target on the top components only. Dropping PC2 removes the direction that makes the coefficients unstable.

Three ways to say it:

  • Picture: estimates of the two coefficients from many samples form a long thin streak (OLS); PCR collapses the streak onto a short segment.
  • Numbers: with correlation 0.95 and 30 rows, each OLS coefficient wobbles with sd about 0.6; with PCR ($k = 1$) about 0.1.
  • Slogan: regress on the directions the data actually measures; drop the ones it cannot tell apart.

Two standardized regressors with correlation $\rho = 0.95$, $n = 30$ rows, noise sd 1, true coefficients $\beta = (1, 1)$. (2 000 simulated datasets.)

  1. Variance inflation: $VIF = 1/(1-\rho^2) = 1/(1 - 0.9025) \approx 10.3$. The variance of each OLS coefficient is about 10 times what it would be with uncorrelated regressors.
  2. OLS: the estimates are centred on $(1, 1)$ (unbiased) but each one has sd about 0.63 (large-sample theory $1/\sqrt{n(1-\rho^2)} \approx 0.58$). Average squared miss $\|\hat\beta - \beta\|^2 \approx 0.80$.
  3. PCR with $k = 1$: PC1 $\approx (1, 1)/\sqrt2$; regress $y$ on the PC1 score, then map back. Estimates centred on $(1, 1)$ with sd about 0.10. Average squared miss about 0.02: 40 times smaller.
  4. The catch: if the truth were $\beta = (2, 0)$ (only temperature matters), PCR with $k=1$ would still return about $(1, 1)$: it can only estimate the "together" direction. Average squared miss about 2.0, worse than OLS (0.75). Dropping PC2 assumed the target does not depend on it.

Principal component regression with $k$ components:

  1. standardize (usually) and centre the regressors; compute PCA on the training rows;
  2. regress $y$ on the scores $Z_k = X_cV_k$ (they are uncorrelated, so the fit is stable): $\hat{\boldsymbol\gamma} = (Z_k^\top Z_k)^{-1}Z_k^\top\mathbf{y}$;
  3. map back to the original regressors: $\hat{\boldsymbol\beta}_{PCR} = V_k\hat{\boldsymbol\gamma}$.

With $k = d$ it equals OLS. With $k \lt d$ it sets the coefficient of every dropped direction to zero: lower variance, but bias whenever $y$ really depends on a dropped (low-variance) direction. Choose $k$ by cross-validation. Ridge regression is a softer cousin: instead of dropping small-variance directions, it shrinks them (Chapter 5.3).

Why do we need it?

Strongly correlated regressors make individual coefficients unstable and uninterpretable, and can make predictions fragile. PCR (or ridge) trades a little bias for a large cut in variance.

Where is it used?

Chemometrics and spectroscopy (hundreds of correlated wavelengths), economics and finance (many correlated indicators), sensor data, and as a baseline before partial least squares (PLS), which picks directions using $y$ as well.

How is it used?

make_pipeline(StandardScaler(), PCA(k), LinearRegression()), with $k$ chosen by cross_val_score or GridSearchCV. Report predictions; be careful interpreting individual coefficients mapped back from components.

Each dot is the pair of coefficients $(\hat\beta_1, \hat\beta_2)$ estimated from one sample of 30 rows. Blue: ordinary least squares. Orange: PCR with $k = 1$. Green: the truth. At $\rho = 0.95$ the blue dots form a long streak (unstable split, stable sum), while the orange dots huddle near the truth. Lower $\rho$ to 0: the streak disappears and PCR no longer helps. Switch the truth to $(2, 0)$: PCR stays tight but in the wrong place (bias), because the truth now has a part along the dropped direction.

"PCR keeps the most important directions for predicting $y$."

It keeps the directions where the regressors vary most, chosen without looking at $y$. If $y$ depends on a low-variance direction, PCR throws that signal away (the $(2, 0)$ case). Validate $k$ on held-out data, or use PLS or ridge.

"After PCR the coefficients mean the same as before."

The mapped-back coefficients are constrained to the span of the kept components; their split between correlated regressors reflects that constraint, not the data.

In your forecasting model, regressors that overlap (temperature and feels-like, two promotion indicators that always fire together) create exactly this long thin posterior ridge for their $\beta$'s. Options, in order of simplicity: keep one of the pair, combine them (their average is the PC1 direction), or let a shrinkage prior handle it. If you replace a block of correlated regressors by their top principal component scores, fit the scaler and the PCA on the training window only and apply them unchanged at forecast time. And remember the trade-off from the widget: a component direction you drop is a direction whose effect you assume is zero.

Multicollinearity = long thin regressor cloud: PC1 well measured, small-variance PCs not. $VIF = 1/(1-\rho^2)$ for two regressors.

PCR: regress $y$ on $Z_k$, $\hat\beta = V_k\hat\gamma$. Less variance, bias if $y$ depends on dropped PCs. $k = d$ ⟹ OLS.

Trap: high-variance directions of $X$ are not necessarily the predictive ones.

Quick check: two regressors have correlation 0.9. What is the VIF, and which combination does PC1 describe?

$VIF = 1/(1 - 0.81) \approx 5.3$. For standardized regressors with positive correlation, PC1 is $(1, 1)/\sqrt2$: their (scaled) sum, the "together" direction, with variance $1.9$; PC2 is their difference, with variance $0.1$.

Whitening: make the cloud round

PCA rotates the cloud so its axes line up with the coordinate axes; the scores are uncorrelated but still have different variances (6 and 1 for our days). Whitening takes one more step: divide each score by its standard deviation $\sqrt{\lambda_j}$. Now every direction has variance 1 and there are no correlations: the cloud is a round ball, like white noise (hence the name).

Distances in the whitened space are Mahalanobis distances in the original space (Chapter 5.15). The price: tiny-variance directions get multiplied by large numbers, so any noise living there is blown up.

Three ways to say it:

  • Picture: rotate the tilted ellipse straight, then squash and stretch it into a circle.
  • Numbers: day 4, centred $(1, 2)$: rotated $(1.789, -1.342)$; divided by $(\sqrt6, 1)$: $(0.730, -1.342)$; squared length $2.333$, its squared Mahalanobis distance.
  • Slogan: rotate, then rescale each axis to variance 1.
  1. Day 4 is $(11, 8)$; centred: $(1, 2)$.
  2. Rotate (PCA scores): $z_1 = (2\cdot1 + 1\cdot2)/\sqrt5 = 4/\sqrt5 \approx 1.789$; $z_2 = (1 - 2\cdot2)/\sqrt5 = -3/\sqrt5 \approx -1.342$.
  3. Rescale: divide by $\sqrt{\lambda_1} = \sqrt6 \approx 2.449$ and $\sqrt{\lambda_2} = 1$: $(0.730, -1.342)$.
  4. Squared length: $0.533 + 1.8 = 2.333$. Check with Mahalanobis: $\frac16(2\cdot1^2 - 4\cdot1\cdot2 + 5\cdot2^2) = \frac16(2 - 8 + 20) = 14/6 = 2.333$. ✓
  5. ZCA whitening rotates back afterwards: $V\,(0.730, -1.342)^\top \approx (0.053, 1.527)$. Same length, but the coordinates stay as close as possible to the original variables.
  • PCA whitening: $\mathbf{w} = \Lambda^{-1/2}V^\top(\mathbf{x}-\bar{\mathbf{x}})$. ZCA (Mahalanobis) whitening: $\mathbf{w} = V\Lambda^{-1/2}V^\top(\mathbf{x}-\bar{\mathbf{x}}) = \hat\Sigma^{-1/2}(\mathbf{x}-\bar{\mathbf{x}})$. Both give $Cov(\mathbf{w}) = I$.
  • $\|\mathbf{w}\|^2 = (\mathbf{x}-\bar{\mathbf{x}})^\top\hat\Sigma^{-1}(\mathbf{x}-\bar{\mathbf{x}})$: the squared Mahalanobis distance.
  • Requires $\lambda_j \gt 0$. Directions with tiny $\lambda_j$ are multiplied by $1/\sqrt{\lambda_j}$, which amplifies noise; in practice drop them or add a small $\epsilon$: $(\Lambda + \epsilon I)^{-1/2}$.
  • scikit-learn: PCA(whiten=True) gives PCA whitening (scores divided by $\sqrt{\lambda_j}$).
Why do we need it?

Some methods assume uncorrelated inputs with equal variance, or work much better with them. Whitening also turns Mahalanobis distance into ordinary distance, so plain Euclidean tools become correlation-aware.

Where is it used?

Independent component analysis (ICA) preprocessing, image preprocessing for some neural networks (ZCA), anomaly detection with distances, the reparameterizations that make samplers and optimizers behave on correlated posteriors, and the "decorrelate then scale" idea behind batch-normalization variants.

How is it used?

PCA(whiten=True).fit_transform(X), or by hand (Xc @ V) / np.sqrt(lam). Drop or regularize near-zero eigenvalues first. Fit on training data and apply the same transform to new data.

Step through the stages with the buttons above the plot. The readout shows the covariance matrix of the cloud at each stage: centring keeps it, rotating makes it diagonal (6 and λ₂), whitening makes it the identity. The three labelled points let you follow individual days. Then lower the smallest variance λ₂ to 0.02 and go to the whitened stage: the thin direction is stretched by $1/\sqrt{\lambda_2}$ (up to 7 times), so any noise in it is amplified just as much.

"Whitening is just standardizing each column."

Standardizing gives each column variance 1 but leaves the correlations; whitening also removes the correlations (it rotates first). Standardized correlated data is still a tilted ellipse.

"Whitening is harmless preprocessing."

It amplifies tiny-variance directions by $1/\sqrt{\lambda}$, which are often pure noise. Drop those components or add a small $\epsilon$ before whitening.

PCA whitening $\mathbf{w} = \Lambda^{-1/2}V^\top(\mathbf{x}-\bar{\mathbf{x}})$; ZCA $= \hat\Sigma^{-1/2}(\mathbf{x}-\bar{\mathbf{x}})$. Both: $Cov = I$, $\|\mathbf{w}\|^2 = d_M^2$.

Standardize ≠ whiten (correlations remain). Trap: $1/\sqrt{\lambda}$ amplifies noise in thin directions.

Quick check: a centred point has PCA scores $(3, 0.5)$ with $\lambda = (9, 0.25)$. What are its whitened coordinates and its Mahalanobis distance?

$(3/3, 0.5/0.5) = (1, 1)$. Squared length $2$, so $d_M = \sqrt2 \approx 1.41$.

Why PCA is linear, and what happens on curved data core

Everything PCA does to a data point is "subtract the mean, then take weighted sums": each score is a fixed weighted sum of the original columns, and the reconstruction is a flat sheet (a line, a plane). That is what linear means here. It is also why PCA is fast, unique (up to signs) and easy to apply to new rows: one matrix multiplication.

But a flat sheet cannot follow a curve. If the data lies along a bent path (a U shape, a horseshoe, a circle), the "one number per point" that would describe it is the position along the curve, and no weighted sum of the columns gives that. PCA will report two or more components where the data really has one, and its first score will put far-apart points (the two ends of a U) at the same value.

Three ways to say it:

  • Picture: PCA lays a straight ruler over the data; it can tilt the ruler but never bend it.
  • Numbers: five points on the parabola $y = x^2$: $(-2,4), (-1,1), (0,0), (1,1), (2,4)$. One number ($x$) describes them exactly, yet PCA needs both components (58% and 42%), and the two ends $(-2, 4)$ and $(2, 4)$ get the same PC1 score.
  • Slogan: PCA finds flat structure; curved structure needs a non-linear method.

Linearity of the score. For the five days, $z_1 = 0.894(\text{sessions} - 10) + 0.447(\text{orders} - 6)$. Day 1 scores $-3.578$, day 4 scores $1.789$. The point halfway between them, $(9, 6)$, scores $0.894\cdot(-1) + 0.447\cdot0 = -0.894 = \frac{-3.578 + 1.789}{2}$: the score of an average is the average of the scores. That is linearity (more precisely, an affine map, because of the mean).

Curved data. The parabola points $(-2,4), (-1,1), (0,0), (1,1), (2,4)$:

  1. Mean $(0, 2)$. Variance of $x$: $(4+1+0+1+4)/4 = 2.5$. Deviations of $y$: $2, -1, -2, -1, 2$, variance $(4+1+4+1+4)/4 = 3.5$.
  2. Covariance: $\big[(-2)(2) + (-1)(-1) + 0 + (1)(-1) + (2)(2)\big]/4 = (-4 + 1 - 1 + 4)/4 = 0$.
  3. So $\hat\Sigma = \text{diag}(2.5, 3.5)$: PC1 is the $y$ axis (58%), PC2 the $x$ axis (42%). PCA sees "two uncorrelated directions", not "one curve".
  4. PC1 scores ($y - 2$): $2, -1, -2, -1, 2$. The two ends of the curve, 4 units apart, get the same score 2. A 1-component PCA summary merges them.

PCA is linear in the following precise sense: once fitted, the map from a data point to its scores is $\mathbf{x}\mapsto V_k^\top(\mathbf{x}-\bar{\mathbf{x}})$ (a matrix times the centred vector), and the reconstruction $\mathbf{z}\mapsto\bar{\mathbf{x}} + V_k\mathbf{z}$ lands on a flat $k$-dimensional affine subspace.

  • It only "sees" the covariance matrix (second moments): any two datasets with the same mean and covariance get the same PCA, whatever their shapes.
  • The fitted directions depend on the data in a non-linear way (through an eigen-decomposition); "linear" refers to how the fitted map treats each point.
  • Non-linear alternatives: kernel PCA, autoencoders, and for visualization t-SNE and UMAP (Chapter 5.17).
Why do we need it?

Knowing PCA is linear tells you when to trust it (roughly flat, elliptical clouds; linear correlations) and when its components are misleading (curves, rings, clusters arranged on a bend).

Where is it used?

Choosing between PCA and non-linear methods (kernel PCA in scikit-learn, autoencoders, t-SNE/UMAP for plots), and explaining in interviews why PCA can be applied to new rows instantly while t-SNE cannot.

How is it used?

Before trusting a PCA summary, plot pairs of variables or the first scores against each other: a curved or folded pattern in the scores is the sign that the structure is non-linear. Then try a non-linear method for exploration, and keep PCA where a linear, reusable transform is needed.

PCA's flat line (one direction) end Aend B PC1 score axis (the vertical direction here) A and B land on the same score the bottom of the curve a straight ruler cannot follow "position along the curve"
On a U-shaped curve the natural one-number summary is the position along the curve. PCA can only use straight directions, so it needs two components, and its first score folds the two ends together.

The points are coloured by their true position along the shape (blue at one end, orange at the other). The green line is PC1; the small marks on it are the points' shadows. For the straight band, the colours stay in order along the line: PC1 is a perfect one-number summary. For the U shape and the horseshoe, blue and orange shadows land on top of each other: PC1 folds the ends together. For the circle, no direction is better than another (about 50% each). Read the rank correlation between true position and PC1 score in the readout.

"PCA failed because I did not keep enough components."

On curved data, more components only rebuild the curve inside a bigger flat space; they never give you the one number (position along the curve) that really describes it. That needs a non-linear method.

"Zero correlation between my two variables means PCA has nothing to find."

The parabola has correlation 0 but is perfectly one-dimensional. PCA (and correlation) miss non-linear structure; plot the data.

"PCA is linear because it uses a linear algebra algorithm."

"PCA is linear because the learned map is linear: each score is a fixed weighted sum of the centred inputs, $\mathbf{z} = V_k^\top(\mathbf{x}-\bar{\mathbf{x}})$, and reconstructions lie on a flat subspace."

Model answer: "PCA finds the best flat $k$-dimensional subspace in squared error, using only the covariance matrix. The fitted transform is a matrix multiply (plus the mean), so it is deterministic, fast, and applies to new rows directly. The cost is that it can only capture linear structure: on curved manifolds it needs extra components and folds distant points together. For non-linear structure I would use kernel PCA or an autoencoder for features, and t-SNE or UMAP for visualization."

Linear: $\mathbf{z} = V_k^\top(\mathbf{x}-\bar{\mathbf{x}})$, reconstructions on a flat sheet; PCA uses only mean + covariance.

Curved data: extra components, folded scores (parabola: correlation 0, still 1-D).

Non-linear options: kernel PCA, autoencoders; t-SNE/UMAP for plots (5.17).

Quick check: a fitted PCA gives scores $z(\mathbf{a}) = 2$ and $z(\mathbf{b}) = -4$ on PC1. What is the PC1 score of the midpoint $(\mathbf{a}+\mathbf{b})/2$?

The score map is affine, so the score of the midpoint is the midpoint of the scores: $(2 + (-4))/2 = -1$.

Limitations: variance is not importance, signs are arbitrary, outliers pull core

PCA ranks directions by how much the inputs vary along them. It never looks at what you want to predict. Imagine users who differ hugely in how many sessions they have, while converters and non-converters differ only slightly, in a quiet direction. PC1 will be "number of sessions"; the direction that separates converters sits in a low-variance component that a "keep 90%" rule would throw away.

Three more habits to keep: a component's sign means nothing (flip it and nothing changes); outliers and a few huge rows can drag the components because errors are squared; and components with nearly equal eigenvalues are unstable: a new sample can rotate them a lot.

Three ways to say it:

  • Picture: the loudest direction is not always the one that tells the classes apart.
  • Numbers: in the example below PC1 explains 95% of the variance and separates the two groups by 0 standard deviations; PC2 explains 5% and separates them by 3.4.
  • Slogan: PCA is unsupervised: big variance ≠ useful.

Two groups of users. Along $x$ (sessions, standardized units) both groups vary widely: sd 3. Along $y$ the groups sit at $-0.6$ and $+0.6$ with a within-group sd of $0.35$.

  1. Variance along $x$: $3^2 = 9$. Variance along $y$ (pooling both groups): within $0.35^2 = 0.1225$ plus between $0.6^2 = 0.36$, total $0.4825$.
  2. PC1 $= x$ axis, explaining $9/9.4825 \approx 94.9\%$; PC2 $= y$ axis, $5.1\%$.
  3. Group separation along PC1: both groups have mean 0 on $x$, so 0 standard deviations.
  4. Along PC2: means differ by $1.2$, within-group sd $0.35$, so $1.2/0.35 \approx 3.4$ standard deviations: an excellent separator.
  5. A "keep 90%" rule keeps only PC1 and throws away the one direction that matters for the groups.

Main limitations of PCA (beyond linearity, previous section):

  • Unsupervised: components are ranked by input variance, not by relevance to any target. Supervised alternatives: partial least squares, or selecting $k$ by validation of the downstream model.
  • Sign (and rotation) ambiguity: $\mathbf{v}_j$ and $-\mathbf{v}_j$ are equally valid; with (near-)equal eigenvalues, any rotation within their plane is too. Never interpret a sign without a convention; compare components across runs up to sign.
  • Outliers and heavy tails: squared errors let a few extreme rows tilt the components; consider robust PCA or a robust covariance.
  • Scale dependence: results depend on units unless you standardize (covariance vs correlation PCA).
  • Interpretability: each component mixes all variables; names like "overall busyness" are interpretations, not facts.
  • Small samples: with $d$ close to $n$, sample eigenvalues are spread out and later components are mostly noise (5.15).
Why do we need it?

To avoid the classic mistakes: dropping the predictive direction, reading meaning into a sign, trusting components driven by one outlier, or comparing components across runs that only differ by a flip.

Where is it used?

Feature engineering for classifiers (where PCA can hurt), interpreting components in reports, monitoring pipelines that refit PCA every week (sign flips break dashboards), and fraud or anomaly data full of outliers.

How is it used?

When PCA feeds a model, choose $k$ by validation, and check the target's relation to each component. Fix a sign convention (sklearn's: largest loading positive) before comparing runs. Look at per-row reconstruction errors and refit without extreme rows to see if the components move.

Blue and orange are two groups of users (for example, non-converters and converters). Left: the data with PC1 (green) and PC2 (teal). Right: the scores of each group on PC1 (top) and on PC2 (bottom). PC1 explains about 95% of the variance but the two groups overlap completely on it; PC2 explains about 5% and separates them. Lower the group gap to 0 and raise it again. Tick flip the sign of PC1: the PC1 histograms mirror, and nothing else changes.

"Keep the components that explain the most variance; they carry the most information about $y$."

They carry the most information about $X$. Relevance to $y$ must be checked separately (validation, PLS, or looking at each component's relation to $y$).

"PC1's loading on revenue went from +0.7 to −0.7 after the retrain, so customer behaviour flipped."

Most likely the whole component's sign flipped, which changes nothing. Compare $|\mathbf{v}^\top\mathbf{v}'|$, or fix a sign convention.

"PCA finds the most important features."

"PCA finds the directions of largest variance in the inputs. Whether they matter for a target is a separate question."

Model answer: "PCA is unsupervised: it ranks directions by input variance, ignoring the label. A low-variance direction can carry all the signal, as when two classes differ along a quiet axis. So when PCA feeds a model I pick $k$ by validation, or use PLS, which uses the target to choose directions. I also standardize mixed units, centre with the training mean, fix a sign convention, and check that outliers are not driving the components."

Unsupervised: variance ≠ predictive importance; validate $k$ on the downstream task.

Signs arbitrary (rotations too when eigenvalues tie); squared error → outliers pull; units matter; components mix all variables.

Trap: dropping a low-variance direction that carries the label.

Quick check: after refitting PCA on new data, the angle between the old and new PC2 is 170°. Has PC2 changed much?

No: 170° is only 10° away from the same axis with the opposite sign. Up to sign, the components differ by 10°. Compare components with $|\mathbf{v}^\top\mathbf{v}'| = |\cos 170°| \approx 0.98$.

Recap, cheat sheet and practice

  • PCA finds the directions of maximum variance; by Pythagoras these are also the directions of minimum reconstruction error (perpendicular, not vertical).
  • The components are the eigenvectors of $\hat\Sigma$; the eigenvalues are the variances of the scores; components are orthogonal, scores uncorrelated.
  • Explained variance $\lambda_j/\sum\lambda$; choose $k$ with a scree plot, a cumulative target, Kaiser's rule (all rules of thumb), or best, validation.
  • Scores $V_k^\top(\mathbf{x}-\bar{\mathbf{x}})$, reconstruction $\bar{\mathbf{x}} + V_k\mathbf{z}$, total error $(n-1)\sum_{j\gt k}\lambda_j$.
  • Centre first (otherwise PC1 points at the mean); standardize mixed units (correlation PCA).
  • In practice PCA = SVD of the centred data: $\lambda = s^2/(n-1)$, scores $US$; scikit-learn centres, does not scale, returns components as rows, with arbitrary signs.
  • Uses: dimensionality reduction, noise removal, PC regression for collinear regressors, whitening (round cloud, Mahalanobis = Euclidean).
  • Limits: linear only (flat sheets), unsupervised (variance ≠ importance), sign ambiguity, sensitive to outliers and units.

Cheat sheet

IdeaFormulaIn words
PC1$\arg\max_{\|\mathbf{u}\|=1}\mathbf{u}^\top\hat\Sigma\mathbf{u}$ ⟹ $\hat\Sigma\mathbf{v} = \lambda\mathbf{v}$widest shadow = smallest loss
Components and variances$\hat\Sigma = V\Lambda V^\top$, $Var(z_j) = \lambda_j$eigenvectors and eigenvalues
Explained share$\lambda_j/\sum_i\lambda_i$; cumulative $\sum_{j\le k}\lambda_j/\sum_i\lambda_i$how much each axis keeps
Scores / reconstruction$\mathbf{z} = V_k^\top(\mathbf{x}-\bar{\mathbf{x}})$, $\hat{\mathbf{x}} = \bar{\mathbf{x}} + V_k\mathbf{z}$compress / rebuild
Reconstruction error$\sum_i\|\mathbf{x}_i-\hat{\mathbf{x}}_i\|^2 = (n-1)\sum_{j\gt k}\lambda_j$what you drop is what you lose
SVD route$X_c = USV^\top$, $\lambda_j = s_j^2/(n-1)$, $Z = US$how libraries compute it
Correlation PCAPCA of standardized columns = eigen of $R$use for mixed units
PCR$\hat\beta = V_k(Z_k^\top Z_k)^{-1}Z_k^\top\mathbf{y}$stable, biased if $y$ needs dropped PCs
Whitening$\Lambda^{-1/2}V^\top(\mathbf{x}-\bar{\mathbf{x}})$ (PCA), $\hat\Sigma^{-1/2}(\mathbf{x}-\bar{\mathbf{x}})$ (ZCA)identity covariance
Code it · Python

import numpy as np
from sklearn.decomposition import PCA, TruncatedSVD
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import cross_val_score

# 1) PCA of the five days (sessions in thousands, orders in hundreds)
X = np.array([[7, 4], [9, 6], [10, 6], [11, 8], [13, 6]], dtype=float)
pca = PCA().fit(X)                      # centres for you (does NOT scale)
print(pca.mean_)                        # [10.  6.]
print(pca.components_.round(3))         # rows = components: [[0.894 0.447] [-0.447 0.894]]
print(pca.explained_variance_)          # [6. 1.]   (divides by n-1)
print(pca.explained_variance_ratio_.round(3))   # [0.857 0.143]
print(pca.transform(X)[:, 0].round(3))  # PC1 scores: [-3.578 -0.894 0. 1.789 2.683]

# 2) Keep k = 1, rebuild, and check: error = (n-1) * dropped eigenvalue
p1 = PCA(n_components=1).fit(X)
X_hat = p1.inverse_transform(p1.transform(X))
print(X_hat.round(2).tolist())          # [[6.8, 4.4], [9.2, 5.6], [10.0, 6.0], [11.6, 6.8], [12.4, 7.2]]
print(((X - X_hat) ** 2).sum().round(6))  # 4.0 = 4 * 1

# 3) The same by hand: SVD of the centred data
Xc = X - X.mean(axis=0)
U, s, Vt = np.linalg.svd(Xc, full_matrices=False)
print(s.round(3), (s ** 2 / (len(X) - 1)).round(3))   # [4.899 2.] [6. 1.]
print(Vt[0].round(3), (U * s)[:, 0].round(3))         # axis + scores (signs can differ from sklearn)

# 4) Forgetting to centre: TruncatedSVD (like np.linalg.svd on raw X) does not centre
ts = TruncatedSVD(n_components=1).fit(X)
print(ts.components_.round(3), (X.mean(0) / np.linalg.norm(X.mean(0))).round(3))
# [[0.859 0.513]] [0.857 0.514]  -> it points at the mean, not along the cloud

# 5) Mixed units: covariance PCA vs correlation PCA
rng = np.random.default_rng(0)
R = np.array([[1, 0.6], [0.6, 1]]); sd = np.array([2000, 40])   # revenue (dollars), orders/day
Y = rng.multivariate_normal([20000, 400], R * np.outer(sd, sd), size=5000)
print(PCA().fit(Y).explained_variance_ratio_.round(4))               # [0.9998 0.0002]: revenue's units win
print(make_pipeline(StandardScaler(), PCA()).fit(Y)[-1].explained_variance_ratio_.round(3))  # [0.81 0.19] (theory 0.8, 0.2)

# 6) Let a variance target choose k (3 hidden drivers + noise in 10 columns)
Z = rng.normal(size=(500, 3)) @ rng.normal(size=(3, 10)) + 0.5 * rng.normal(size=(500, 10))
p90 = PCA(n_components=0.9).fit(Z)
print(p90.n_components_, p90.explained_variance_ratio_.cumsum().round(3))  # 3 [0.489 0.898 0.96]

# 7) PC regression vs OLS on two collinear regressors (5-fold cross-validated R^2)
n = 60
t = rng.normal(size=n)
f = 0.95 * t + np.sqrt(1 - 0.95**2) * rng.normal(size=n)       # "feels-like" ~ temperature
Xr = np.c_[t, f]
y = Xr @ np.array([1.0, 1.0]) + rng.normal(size=n)
for name, model in [("OLS", LinearRegression()),
                    ("PCR k=1", make_pipeline(StandardScaler(), PCA(1), LinearRegression()))]:
    print(name, cross_val_score(model, Xr, y, cv=5, scoring="r2").mean().round(3))
# OLS 0.713 / PCR k=1 0.712: same predictive quality here; PCR's gain is stable coefficients

# 8) Whitening: identity covariance; squared length = squared Mahalanobis distance
W = PCA(whiten=True).fit(X).transform(X)
print(np.cov(W, rowvar=False).round(6))  # identity
print((W[3] ** 2).sum().round(3))        # day 4: 2.333
Test yourself

1. The first principal component of a dataset is…

PC1 maximizes $\mathbf{u}^\top\hat\Sigma\mathbf{u}$ over unit vectors, which gives $\hat\Sigma\mathbf{u} = \lambda\mathbf{u}$ with the largest $\lambda$. It minimizes perpendicular (not vertical) distances; the direction to the mean is what you get when you forget to centre.

2. Covariance eigenvalues are 6, 3 and 1. What share of the variance do the first two components keep?

$(6 + 3)/(6 + 3 + 1) = 9/10 = 90\%$. The reconstruction loses the remaining 10%: total squared error $(n-1)\times1$.

3. Your data sits far from the origin and you run np.linalg.svd on the raw table. The first right singular vector…

The raw second-moment matrix is $\hat\Sigma$ plus a large rank-one term $\frac{n}{n-1}\bar{\mathbf{x}}\bar{\mathbf{x}}^\top$, which dominates. SVD equals PCA only on the centred data.

4. Your columns are revenue in dollars, sessions per day and conversion rate in percent. Before PCA you should…

Covariance PCA follows raw variances, so the dollar column would dominate PC1 because of its unit. Standardizing gives every column variance 1, which is the same as PCA of the correlation matrix.

5. In what sense is PCA "linear"?

The fitted map is $\mathbf{x}\mapsto V_k^\top(\mathbf{x}-\bar{\mathbf{x}})$, a matrix multiply. That is why it is fast and reusable, and why it cannot follow curved structure.

6. A component explains only 3% of the input variance. Which statement is true?

PCA ranks directions by input variance and never looks at the target. Two groups can differ only along a quiet direction. Check relevance by validation (or use PLS).

Practice problems

A. $\hat\Sigma = \begin{bmatrix} 4 & 2 \\ 2 & 4 \end{bmatrix}$ with mean $(4, 3)$. Find the components and shares, then the scores, rebuild ($k = 1$) and error of the point $(6, 3)$.
  1. Eigenvalues $4 \pm 2 = 6, 2$; components $(1,1)/\sqrt2$ and $(1,-1)/\sqrt2$; shares 75% and 25%.
  2. Centred point $(2, 0)$. Scores: $z_1 = 2/\sqrt2 \approx 1.414$, $z_2 = 2/\sqrt2 \approx 1.414$.
  3. Rebuild with $k = 1$: $(4,3) + 1.414\cdot(0.707, 0.707) = (4,3) + (1, 1) = (5, 4)$.
  4. Error vector $(6,3) - (5,4) = (1, -1)$, squared length 2 $= z_2^2$. ✓
B. Show that the PCA scores on different components are uncorrelated.

Scores $Z = X_cV$. Their covariance is $\frac{Z^\top Z}{n-1} = V^\top\frac{X_c^\top X_c}{n-1}V = V^\top\hat\Sigma V = V^\top V\Lambda V^\top V = \Lambda$, because $V^\top V = I$. $\Lambda$ is diagonal: the off-diagonal covariances are 0 and the variances are the eigenvalues.

C. Three points $(9, 11), (11, 9), (10, 10)$. Compare PCA (centred) with the SVD of the raw table.

Centred: mean $(10, 10)$, rows $(-1, 1), (1, -1), (0, 0)$; $\hat\Sigma = \begin{bmatrix} 1 & -1 \\ -1 & 1 \end{bmatrix}$ with eigenvalues 2 and 0: PC1 $= (1, -1)/\sqrt2$, explaining 100% (the points lie on a line going down-right).

Raw: $X^\top X = \begin{bmatrix} 302 & 298 \\ 298 & 302 \end{bmatrix}$, eigenvalues $600$ and $4$; the first direction is $(1, 1)/\sqrt2$, "explaining" $600/604 \approx 99.3\%$. It points at the mean and is perpendicular to the true PC1. Forgetting to centre gives exactly the wrong answer here.

D. Revenue (sd 2 000 dollars) and orders (sd 40) have correlation 0.6. Give PC1 and its share for covariance PCA and for correlation PCA, and explain the difference.

Covariance PCA: $\Sigma = \begin{bmatrix} 4\,000\,000 & 48\,000 \\ 48\,000 & 1\,600 \end{bmatrix}$; PC1 $\approx (0.99993, 0.012)$, share $\approx 99.97\%$. Correlation PCA: $R$ has eigenvalues $1.6, 0.4$; PC1 $= (1,1)/\sqrt2$, share 80%. Covariance PCA is dominated by the unit of revenue (its variance is 2 500 times larger); correlation PCA sees the real structure: the two metrics move together.

E. (Interview) "A colleague reduced 50 candidate regressors for a churn model to 5 principal components, which explain 92% of the variance. What would you ask?"

"Were the columns standardized (mixed units)? Was PCA fitted on the training data only, and is the same mean and scaler applied at prediction time? 92% of the input variance says nothing about churn: was $k$ chosen by validated model performance? Could the churn signal live in a low-variance component that was dropped? Are a few outliers driving the top components? And if anyone reads the loadings, is there a sign convention, since signs can flip between refits?"

F. (Interview) "Why is PCA a linear method, and when would you not use it?"

"After fitting, every point is mapped by $V_k^\top(\mathbf{x}-\bar{\mathbf{x}})$: each score is a fixed weighted sum of the inputs, and reconstructions lie on a flat subspace. PCA depends only on the mean and covariance. So it is fast, deterministic and reusable on new data, but it cannot unfold curved structure: on a U-shaped manifold it needs extra components and folds the two ends together. I would not use it when the structure is clearly non-linear (then kernel PCA or an autoencoder for features, t-SNE or UMAP for pictures), when the target lives in low-variance directions (then PLS or validated $k$), or when heavy outliers dominate (robust PCA)."

Chapter 5.17 · Syllabus Module 22, Part 26

t-SNE and UMAP: non-linear maps for looking at data

PCA gives you a straight shadow of your data. Sometimes the shadow hides what matters: groups that are well separated in 10 dimensions land on top of each other. t-SNE and UMAP draw a different kind of picture: a map that tries to keep every point next to its true neighbours. This chapter builds t-SNE from scratch (you will run a real one in your browser), shows exactly what such a map can and cannot tell you, and explains why it is a tool for looking, not a feature pipeline.

  • Explain the difference between a linear reduction (PCA: one formula, a shadow) and a non-linear neighbour map (t-SNE, UMAP: positions chosen to keep neighbours)
  • Compute t-SNE's high-dimensional similarities: a Gaussian bump around each point, with its own width chosen by perplexity
  • Explain why the map uses a heavy-tailed Student-t kernel (the crowding problem)
  • Say in plain words what the KL divergence objective rewards and punishes, and read its gradient as springs that pull and push
  • Run t-SNE yourself next to PCA, and change perplexity, seed and start
  • Know the big warning: cluster sizes and the distances between far-apart clusters in a t-SNE map are not meaningful
  • Explain why t-SNE is a poor production feature reducer, what UMAP does differently, and when to choose PCA, t-SNE or UMAP

What we need from earlier chapters: PCA as the best straight projection (Chapter 5.16); distances between points and why high dimensions behave strangely (Chapter 4.16, the curse of dimensionality); the Normal and Student-t densities (Chapter 4.9); gradient descent and momentum (Optimization 3.3 and 3.4). The KL divergence gets a short plain definition here; its full story (both directions, variational inference) is in Chapter 6.11. Words: a dimension is one column (feature) of the data; a point in 10 dimensions is a row with 10 numbers. An embedding (or map) is the list of new, low-dimensional positions $y_1, \dots, y_n$ we draw, one per data point $x_1, \dots, x_n$.

Why a non-linear map? A shadow keeps big directions, a map keeps neighbours core

Picture a garden hose wound into a tall coil, like a spring. You want a flat picture of it. Shine a torch from above and trace the shadow: every loop of the coil lands on the same circle. Two bits of hose that are a metre apart along the hose now sit on top of each other. That is what a linear projection such as PCA does: it keeps the directions in which the data spreads most, and it may stack far-apart points on top of each other.

A different idea: unwind the hose and lay it straight on the floor. Now every bit of hose sits next to the same bits it touched before. You lost the coil shape (the "global" picture), but you kept the neighbours. t-SNE and UMAP try to do this kind of unwinding automatically, for data in 10, 50 or 1000 dimensions.

Three ways to say it:

  • Picture: PCA is a shadow on the wall; t-SNE is a hand-drawn map where every house stays next to its neighbours.
  • Numbers: for 8 points on a circle, the shadow puts the top and bottom points (2 apart, the farthest pair) on the same spot; unrolling the circle keeps every neighbour pair except one.
  • Slogan: PCA keeps the big spreads; t-SNE keeps the small neighbourhoods.

Eight points on a circle, squeezed into one dimension. The points sit at angles $0^\circ, 45^\circ, \dots, 315^\circ$ on a circle of radius 1.

  1. Coordinates: point $k$ is at $(\cos\theta_k, \sin\theta_k)$ with $\theta_k = 45^\circ \times k$.
  2. Shadow (linear). A circle spreads equally in every direction, so every line is an equally good "best" line for PCA; take the horizontal one. Each point's shadow is its $x$-coordinate: $1,\ 0.707,\ 0,\ -0.707,\ -1,\ -0.707,\ 0,\ 0.707$.
  3. Collisions: $45^\circ$ and $315^\circ$ both land on $0.707$; $90^\circ$ and $270^\circ$ both land on $0$; $135^\circ$ and $225^\circ$ both land on $-0.707$. The points at $90^\circ$ and $270^\circ$ are $2$ apart on the circle, the largest distance there is, yet their shadows coincide.
  4. Unrolling (non-linear). Cut the circle at $0^\circ$ and straighten it. Each point's new position is the arc length from the cut: $0,\ 0.785,\ 1.571,\ 2.356,\ 3.142,\ 3.927,\ 4.712,\ 5.498$.
  5. Check the neighbours: every point's two neighbours on the circle are still next to it on the line, except across the cut. Nothing lands on top of anything else.
  6. The price: across the cut, $0^\circ$ and $315^\circ$ were only $2\sin(22.5^\circ) \approx 0.765$ apart, and now they are $5.498$ apart. Squeezing data into fewer dimensions always distorts something. A neighbour map chooses to protect small distances and to sacrifice large ones.

Dimensionality reduction turns each data point $x_i \in \mathbb{R}^D$ ($D$ numbers) into a point $y_i \in \mathbb{R}^d$ with far fewer numbers; for a picture, $d = 2$.

  • A linear reduction uses one fixed formula for every point: $y = W^\top (x - \mu)$, with $\mu$ the mean and $W$ a $D \times d$ matrix. PCA picks the columns of $W$ as the top eigenvectors of the covariance matrix (Chapter 5.16). Straight lines stay straight, and a new point is mapped with the same formula.
  • A non-linear reduction is any other rule. t-SNE and UMAP do not even produce a formula: they directly choose the positions $y_1, \dots, y_n$ of the points you gave them. This list of positions is called an embedding.
  • Local structure means "who are each point's nearest neighbours". Global structure means the large-scale layout: how far apart the groups are, how big each group is.
  • A neighbour embedding (t-SNE, UMAP) chooses the positions so that points which are neighbours in $\mathbb{R}^D$ are neighbours in the map. It is built mainly for visualization (making pictures for human eyes).
Why do we need it?

Real data often lies on a curved, folded surface inside many dimensions, or in groups that no single flat shadow can separate. A straight projection can stack separate groups on top of each other, and then the picture lies to you. A neighbour map can still show the groups.

Where is it used?

Pictures of word and sentence embeddings, single-cell RNA data in biology (where t-SNE and UMAP plots are standard), the hidden layers of a neural network, image datasets such as handwritten digits, customer or segment feature vectors, and windows of a time series.

How is it used?

Standardize the features, often reduce to about 50 dimensions with PCA first, run t-SNE or UMAP to 2 dimensions, colour the points by a label you know (segment, device, week) and look. Then check what you see with the original data, never with the picture alone.

0° 90° 270° 315° 8 points in 2-D Shadow (linear, like PCA) 90° and 270° collide 3 pairs land on the same spot Unrolled (non-linear neighbour map) neighbours kept; 0° and 315° torn apart at the cut
The shadow keeps the widest direction of spread but stacks far-apart points (90° and 270°) on one spot. Unrolling keeps every neighbour pair except across the cut, where two close points end up far apart. No flat picture keeps everything.

Sixty points lie on a spiral (top). Colour runs from blue to orange along the spiral, so neighbours have similar colours. The two strips below are two 1-D pictures of the same points: the PCA shadow on the best straight line (purple line) and the unrolled spiral. Raise the number of turns: the shadow strip becomes a jumble of colours (far-apart parts of the spiral land together), while the unrolled strip stays a smooth blue-to-orange ramp. The readout counts how many of each point's 4 nearest neighbours survive in each strip.

"t-SNE is a better PCA."

They answer different questions. PCA is linear, keeps large-scale spread, gives a reusable formula and axes you can interpret (loadings). t-SNE keeps local neighbourhoods only and gives a picture of the points you fed it, nothing more.

"A clever non-linear map can keep all the distances."

In general it cannot. Three points at the corners of an equilateral triangle cannot be placed on a line with all three distances kept, and 10-dimensional data cannot be drawn in 2-D without distortion. Each method decides which distances to protect.

"Non-linear means more accurate."

Non-linear means more flexible. Flexibility lets the map show curved structure, but it also lets it invent or exaggerate structure (you will see clusters appear in pure noise later in this chapter).

Neither of your projects needs t-SNE inside the model. Where it can help is in looking. In an A/B framework like yours, if you have a feature vector per segment or per user (pre-experiment behaviour, device, region), a 2-D map coloured by segment shows whether segments really differ or overlap. In your forecasting model, you could cut the residuals into weekly windows (7 numbers each) and map the windows: weeks that misbehave in the same way land together, which can point to a missing holiday or regressor. Any finding is then checked on the original numbers.

Linear (PCA): one formula $y = W^\top(x-\mu)$, a shadow; keeps big spreads; can stack far-apart points.

Non-linear neighbour map (t-SNE, UMAP): positions chosen directly so neighbours stay neighbours; big distances are sacrificed. Mainly for pictures.

Trap: no 2-D picture keeps every distance. Know which ones your method protects.

Quick check: why can no 1-D map keep all three distances of an equilateral triangle with sides 1?

On a line, one of the three points lies between the other two, so its two distances add up to the third: $a + b = c$. In the triangle all three distances equal 1, and $1 + 1 \ne 1$. So at least one distance must change. Squeezing dimensions forces distortion.

Step 1: who are my neighbours? Similarities in the original space core

Imagine each point is a guest at a party, and each guest has 100% of their attention to share among the other guests. People standing close get a lot of attention; people across the room get almost none. t-SNE writes this down as a table of numbers: for every guest $i$ and every other guest $j$, the share of $i$'s attention that goes to $j$.

The rule for "close gets a lot" is a bell curve (a Normal density) centred on guest $i$. Attention drops off quickly with distance. The width $\sigma_i$ of $i$'s bell says how far $i$'s attention reaches: a narrow bell only notices the very nearest guests; a wide bell spreads attention over many.

Three ways to say it:

  • Picture: every point shines a soft, bell-shaped spotlight around itself; the brighter a neighbour is lit, the more similar it counts as.
  • Numbers: with width 1, a point at distance 1 gets 82% of the attention, a point at distance 2 gets 18%, a point at distance 4 gets almost nothing.
  • Slogan: $p_{j\mid i}$ is the chance that $i$ would pick $j$ as its neighbour.

Point A has three other points around it: B at distance 1, C at distance 2, D at distance 4. Start with width $\sigma = 1$.

  1. Bell height for each neighbour: $e^{-d^2/(2\sigma^2)}$. For B: $e^{-1/2} = 0.6065$. For C: $e^{-4/2} = e^{-2} = 0.1353$. For D: $e^{-16/2} = e^{-8} = 0.000335$.
  2. Add them: $0.6065 + 0.1353 + 0.000335 = 0.7422$.
  3. Divide each height by the sum, so the shares add to 1: $p_{B\mid A} = 0.6065/0.7422 = 0.817$, $p_{C\mid A} = 0.1353/0.7422 = 0.182$, $p_{D\mid A} = 0.000335/0.7422 = 0.00045$.
  4. Now a wider bell, $\sigma = 2$: heights $e^{-1/8} = 0.8825$, $e^{-4/8} = 0.6065$, $e^{-16/8} = 0.1353$; sum $1.6244$; shares $0.543,\ 0.373,\ 0.083$. The wider bell spreads A's attention more evenly.
  5. B has its own bell, so $p_{A\mid B}$ is usually different from $p_{B\mid A}$. t-SNE averages the two directions into one symmetric number: $p_{AB} = \dfrac{p_{B\mid A} + p_{A\mid B}}{2n}$. With $n = 4$ points and, say, $p_{A\mid B} = 0.5$: $p_{AB} = (0.817 + 0.5)/8 = 0.165$.

For data points $x_1, \dots, x_n \in \mathbb{R}^D$, the conditional similarity of $j$ to $i$ is

$$p_{j\mid i} = \frac{\exp\!\big(-\|x_i - x_j\|^2 / 2\sigma_i^2\big)}{\sum_{k \ne i} \exp\!\big(-\|x_i - x_k\|^2 / 2\sigma_i^2\big)}, \qquad p_{i\mid i} = 0,$$

and the joint similarity used by t-SNE is

$$p_{ij} = \frac{p_{j\mid i} + p_{i\mid j}}{2n}.$$
  • $\|x_i - x_j\|$ is the ordinary (Euclidean) distance in the original space; $\sigma_i$ is the bandwidth (bell width) of point $i$, chosen by the perplexity rule in the next section.
  • Each row $p_{\cdot\mid i}$ adds to 1. The $p_{ij}$ are symmetric ($p_{ij} = p_{ji}$) and add to 1 over all $n(n-1)$ ordered pairs, so $P$ is one probability distribution over pairs.
  • Why average the two directions? Every point then gets at least $\frac{1}{2n}$ of the total weight, even an outlier that nobody else notices, and the gradient becomes simpler.
  • Distances depend on units, so features are usually standardized first (z-scores, Chapter 4.18). Otherwise a feature measured in dollars swamps one measured in fractions.
Why do we need it?

The map can only keep neighbours if we first say, in numbers, who is whose neighbour and how strongly. This table $P$ is the target that the whole t-SNE map tries to copy.

Where is it used?

The first step of SNE and t-SNE (scikit-learn TSNE, openTSNE, the R package Rtsne). The same Gaussian "attention" idea appears in kernel density estimation, in the RBF kernel of SVMs and Gaussian processes, and (with a dot product instead of a distance) in softmax attention.

How is it used?

You never compute it by hand: TSNE(perplexity=30) does it. What you control is the input: which features, standardized or not, and whether to run PCA to about 50 dimensions first. Those choices change the distances, so they change $P$ and the final picture.

The purple point is guest $i$. Drag it around. Each blue point gets a green disc whose area shows its share $p_{j\mid i}$ of $i$'s attention; the dashed circles mark distances $\sigma$ and $2\sigma$. Put $i$ inside the tight group on the left: with $\sigma = 1$ its attention is shared among the six tight points and the right group gets almost nothing. Shrink $\sigma$ to 0.3: the 2–3 nearest points take almost everything. Widen $\sigma$ to 3: attention leaks into the sparse group. Move $i$ into the sparse group with $\sigma = 0.3$: one neighbour takes nearly all of it.

"$p_{j\mid i}$ is the probability that $i$ and $j$ are in the same cluster."

It is a similarity score shaped like a probability: the chance that $i$ would pick $j$ if it picked a neighbour with Gaussian-weighted odds. No clusters are involved; t-SNE never sees labels.

"$p_{j\mid i} = p_{i\mid j}$."

Not in general: each point has its own bell width $\sigma_i$ and its own crowd of neighbours. That is why t-SNE averages the two into the symmetric $p_{ij}$.

"Distances are distances; scaling the features does not matter."

Every $p$ is built from distances, and distances change when one feature is measured in larger units. Standardize (or otherwise choose the scale on purpose) before running t-SNE.

$p_{j\mid i} \propto \exp(-\|x_i-x_j\|^2/2\sigma_i^2)$, normalized over $j \ne i$; then $p_{ij} = (p_{j\mid i}+p_{i\mid j})/2n$.

Meaning: a table of "how much is $j$ one of $i$'s neighbours", the target the map must copy.

Trap: built from distances, so the feature scaling and the bell width $\sigma_i$ decide everything.

Quick check: in the example, what happens to $p_{B\mid A}$, $p_{C\mid A}$, $p_{D\mid A}$ when $\sigma$ becomes huge, say 100?

Every bell height becomes almost $e^0 = 1$, so the three shares approach $1/3$ each: A pays equal attention to everyone and the distances stop mattering. When $\sigma$ is tiny, almost all attention goes to the single nearest point B. The width decides how many neighbours count.

Perplexity: how many neighbours each point listens to core

Why not use one bell width for everybody? Because data has crowded places and empty places. In a city, a circle of 1 km holds thousands of people; in the countryside, it may hold nobody. One width would give city points hundreds of "neighbours" and country points none.

So t-SNE asks you a different question: roughly how many neighbours should each point pay attention to? That number is the perplexity. Then, point by point, t-SNE adjusts the width $\sigma_i$ until that point's attention is spread over about that many neighbours: narrow bells in crowded areas, wide bells in empty ones.

Three ways to say it:

  • Picture: every point turns its spotlight dial until it lights up about the same number of neighbours.
  • Numbers: perplexity 30 means each point acts as if it had about 30 equally important neighbours.
  • Slogan: perplexity is a smooth version of the "k" in k-nearest neighbours.

We need one new word. The entropy of a set of shares $p_1, p_2, \dots$ measures how spread out they are: $H = -\sum p_j \log_2 p_j$, in bits. It is the average number of yes/no questions you need to guess which neighbour was picked. The perplexity is $2^H$.

  1. Attention split equally over 4 neighbours, $p = (\tfrac14, \tfrac14, \tfrac14, \tfrac14)$: $H = -4 \times \tfrac14 \log_2 \tfrac14 = 2$ bits, so perplexity $= 2^2 = 4$. Equal shares over $k$ neighbours give perplexity exactly $k$.
  2. Point A from the previous example with $\sigma = 1$: shares $0.817, 0.182, 0.00045$. $H = 0.817 \times 0.291 + 0.182 \times 2.456 + 0.00045 \times 11.1 = 0.238 + 0.448 + 0.005 = 0.691$ bits. Perplexity $= 2^{0.691} = 1.61$: A effectively listens to 1.6 neighbours.
  3. With $\sigma = 2$: shares $0.543, 0.373, 0.083$; $H = 1.308$ bits; perplexity $= 2.48$.
  4. Suppose you asked for perplexity 2. It lies between 1.61 and 2.48, so the right width is between 1 and 2. Halving the gap again and again (a binary search) finds $\sigma_A = 1.40$, where the shares are $0.672, 0.313, 0.015$ and $H = 1.000$ bit exactly.
  5. A has only 3 neighbours, so its perplexity can never pass 3 (that is the equal-shares limit as $\sigma \to \infty$). This is why the perplexity must be smaller than the number of points.

For point $i$ with conditional shares $P_i = (p_{j\mid i})_{j \ne i}$:

$$H(P_i) = -\sum_{j \ne i} p_{j\mid i}\,\log_2 p_{j\mid i}, \qquad \mathrm{Perp}(P_i) = 2^{H(P_i)}.$$
  • You choose one perplexity for the whole dataset. For each $i$, t-SNE finds $\sigma_i$ by binary search so that $\mathrm{Perp}(P_i)$ equals it. This works because a wider bell always gives flatter shares and so a larger perplexity.
  • Typical values are 5 to 50 (a rule of thumb from the original t-SNE paper); scikit-learn's default is 30. It must be smaller than the number of points (scikit-learn raises an error otherwise), and in practice much smaller.
  • scikit-learn's fast (Barnes–Hut) method only looks at each point's $\lfloor 3 \times \text{perplexity} \rfloor + 1$ nearest neighbours, because the bell is almost zero beyond them.
  • (scikit-learn works with natural logs, $\ln$, and compares the entropy with $\ln(\text{perplexity})$. That is the same rule written in different units.)
Why do we need it?

A single bell width cannot fit both crowded and sparse regions. Perplexity gives every point the same "number of neighbours" instead, so dense and sparse parts of the data are both represented. It is the main knob of t-SNE.

Where is it used?

The perplexity argument of scikit-learn's TSNE, openTSNE and Rtsne. UMAP's n_neighbors plays the same role. The same entropy is the "perplexity" used to score language models, where it is the effective number of equally likely next words.

How is it used?

Try a few values (for example 5, 30 and 50, all below the number of points) and compare the maps. Structure that appears at every perplexity is more trustworthy than structure that appears at only one. Small values show fine local detail; larger values show more of the coarse layout.

Forty points: a tight group (blue) and a loose group (orange). Each faint circle has radius $\sigma_i$, the width that point needed to reach the chosen perplexity. Drag the purple probe onto any point to see its binary search. Notice: the same perplexity gives small circles in the tight group and big ones in the loose group. Raise the perplexity to 20 or more (each point has only 19 others in its own group): points must reach into the other group to find enough neighbours, and the tight group's circles suddenly grow a lot.

"Perplexity is the number of clusters."

It is the effective number of neighbours each point listens to. It has nothing to do with how many groups the data contains.

"All points use the same bandwidth."

Each point gets its own $\sigma_i$, chosen so that every point has the same perplexity: small in dense regions, large in sparse ones. A side effect is that t-SNE evens out density, one reason why cluster sizes in the map mean little (see "Reading a t-SNE map" below).

"There is one correct perplexity."

Different values reveal structure at different scales. Look at several, and trust what survives across them. The usual range 5–50 is a rule of thumb.

$\mathrm{Perp}(P_i) = 2^{H(P_i)}$, $H = -\sum_j p_{j\mid i}\log_2 p_{j\mid i}$; $\sigma_i$ is found by binary search to hit the chosen perplexity.

Meaning: the effective number of neighbours per point (a smooth $k$). Default 30; try 5–50; must be below $n$.

Trap: same perplexity, different $\sigma_i$: dense and sparse regions are equalized.

Quick check: you run t-SNE on 25 points with the default perplexity 30. What happens, and why does it make sense?

scikit-learn stops with a ValueError saying that the perplexity must be less than n_samples. With 24 other points, no bell width can make a point behave as if it had 30 equally important neighbours (the maximum is 24, reached when all shares are equal). Use a smaller perplexity, such as 5.

Step 2: similarities in the map, and why they use a heavy-tailed Student-t core

The map needs its own table of similarities, $q_{ij}$, computed from the 2-D positions. The first version of the method (SNE, 2002) used the same bell curve in the map as in the data. It had a problem called crowding.

In 10 dimensions there is a lot of room. You can place 11 points so that every pair is exactly the same distance apart. On a flat sheet, at most 3 points can do that (an equilateral triangle). So when moderately distant points from 10-D are forced onto a sheet, there is not enough space at the right distance, and they get squeezed toward each other. Everything collapses into one crowded blob.

t-SNE's fix: in the map, measure similarity with a curve that falls off slowly (a heavy-tailed Student-t). Then a moderate similarity corresponds to a large map distance, so moderately distant groups can be pushed far apart and gaps appear between them.

Three ways to say it:

  • Picture: 11 guests who all stand equally far from each other fit in a 10-D room, but not around a flat table; the heavy tail lets the table stretch.
  • Numbers: a similarity of 0.018 means distance 2 under a bell curve, but distance 7.3 under the Student-t curve.
  • Slogan: heavy tails give the map room to breathe, and they also make gap sizes meaningless.

Compare the two curves at the same distance $d$: the bell $e^{-d^2}$ and the Student-t curve $1/(1+d^2)$. Then ask: if a pair needs similarity $s$, how far apart must the two points be under each curve?

  1. $d = 0.5$: bell $e^{-0.25} = 0.779$, t-curve $1/1.25 = 0.80$. Almost the same: very close pairs are treated alike.
  2. $d = 1$: bell $e^{-1} = 0.368$, t-curve $1/2 = 0.5$.
  3. $d = 2$: bell $e^{-4} = 0.0183$, t-curve $1/5 = 0.2$. The bell has almost died; the t-curve is still at 0.2.
  4. $d = 3$: bell $e^{-9} = 0.000123$, t-curve $1/10 = 0.1$.
  5. Matching distance: set $1/(1+D^2) = e^{-d^2}$, so $D = \sqrt{e^{d^2} - 1}$. For $d = 0.5$: $D = \sqrt{0.284} = 0.53$. For $d = 1$: $D = 1.31$. For $d = 2$: $D = \sqrt{53.6} = 7.32$. For $d = 3$: $D = \sqrt{8102} = 90$.
  6. So the t-map keeps small distances roughly as they are, and stretches moderate ones enormously (2 becomes 7.3, 3 becomes 90). That stretch opens the gaps between clusters, and it is not proportional to anything: a gap twice as wide in the map does not mean "twice as different".

For map positions $y_1, \dots, y_n \in \mathbb{R}^2$, t-SNE's map similarities are

$$q_{ij} = \frac{\big(1 + \|y_i - y_j\|^2\big)^{-1}}{\sum_{k \ne l} \big(1 + \|y_k - y_l\|^2\big)^{-1}}, \qquad q_{ii} = 0.$$
  • The curve $(1 + d^2)^{-1}$ has the shape of a Student-t density with 1 degree of freedom (also called the Cauchy distribution, Chapter 4.9). Its tails are heavy: it falls like $1/d^2$ (a power), not like $e^{-d^2}$.
  • The $q_{ij}$ add to 1 over all ordered pairs, just like the $p_{ij}$. There is no per-point width in the map: one fixed curve for everybody. (scikit-learn uses $\nu = \max(\text{n\_components} - 1, 1)$ degrees of freedom, which is 1 for a 2-D map.)
  • Two side benefits: no exponential to compute, and far away the curve behaves like $1/d^2$, so a far-away cluster acts almost like a single point.
  • Summary: Gaussian with per-point width in the data, Student-t with fixed width in the map. The "t" in t-SNE is this Student-t; SNE stands for Stochastic Neighbour Embedding.
Why do we need it?

With a bell curve in the map (plain SNE), moderately distant points are squeezed together and clusters merge into one blob. The heavy tail lets the map place them far apart, so separate groups show visible gaps.

Where is it used?

Every t-SNE implementation (scikit-learn, openTSNE, Rtsne, FIt-SNE). UMAP uses a similar heavy-tailed family $1/(1 + a\,d^{2b})$. Heavier tails than 1 degree of freedom (smaller $\nu$) give even more separated clusters; openTSNE exposes this as a dof option.

How is it used?

You do not set it in scikit-learn; it is built in. What you take away is how to read the map: small distances are roughly faithful; large gaps are stretched by a huge, non-linear factor, so compare "which groups exist", not "how far apart they are".

On a flat sheet (2-D) at most 3 points all 1 apart In 10-D 11 points can all be exactly 1 apart (the corners of a 10-D simplex) squeeze Forced into 2-D bell-curve map: crowding heavy tails push them apart
The crowding problem. High-dimensional space has room for many points at the same moderate distance; a flat map does not. A Gaussian map squeezes them into a blob; t-SNE's heavy-tailed kernel lets moderate distances become large map distances.

Drag the purple handle along the bottom to choose a distance $d$ under the bell curve (blue). The dashed line walks sideways to the distance $D$ where the Student-t curve (orange) gives the same similarity. Near $d = 0.5$ the two distances agree. Drag to $d = 2$: the t-curve needs $D \approx 7.3$. Past $d \approx 2.2$, $D$ runs off the chart. This stretch is what separates clusters, and it is why gap sizes in a t-SNE map are not to scale.

Sixty random points fill a cube with $D$ dimensions. The histogram shows all 1 770 pairwise distances, each divided by the average distance (purple line at 1). Slide $D$ from 2 up to 100: the histogram squeezes toward 1. In 100 dimensions almost every pair is within ±10% of the average distance. A flat map cannot show "many points, all at the same distance from each other", which is the crowding problem the Student-t tail is fighting.

"t-SNE uses a Student-t because the data has heavy tails."

The Student-t is used only in the map, to fight crowding. It says nothing about the distribution of your data. The data side uses Gaussian bells.

"A gap twice as wide means the clusters are twice as different."

The heavy tail stretches moderate distances by huge, non-linear factors (2 becomes 7.3, 3 becomes 90 in the example). Gap widths are not proportional to anything in the data.

"The map kernel also gets a per-point width from the perplexity."

Only the data side has per-point widths $\sigma_i$. The map uses one fixed curve $1/(1+d^2)$ for every pair.

$q_{ij} \propto (1 + \|y_i - y_j\|^2)^{-1}$: Student-t with 1 degree of freedom, fixed width, normalized over all pairs.

Why: crowding. Heavy tails let moderate data distances become large map distances, so clusters separate.

Trap: the same stretch makes gap sizes in the map not to scale.

Quick check: two pairs have bell-curve distances 0.5 and 2 (a ratio of 4). Using the matching distances, what is their ratio in a t-SNE-style map?

$D(0.5) = 0.53$ and $D(2) = 7.32$, so the ratio is about $7.32/0.53 \approx 13.8$ instead of 4. Small distances are kept; moderate ones are blown up. Ratios of distances in the map do not match ratios in the data.

The objective: KL divergence, and the springs it creates core

We now have two tables: $P$ ("who is close to whom in the data") and $Q$ ("who is close to whom in the map"). t-SNE moves the map points until $Q$ copies $P$ as well as possible. It needs a score for "how badly does $Q$ copy $P$?". That score is the Kullback–Leibler (KL) divergence.

The KL score has a built-in bias. It charges a lot when two true neighbours are drawn far apart (big $p$, small $q$). It charges very little when two strangers are drawn close (tiny $p$, any $q$), because every pair's cost is weighted by its $p$. So t-SNE works very hard on neighbourhoods and hardly cares where far-apart groups end up.

Moving the points to lower the score feels like a system of springs: each pair is pulled together if the map shows it less similar than the data ($q \lt p$), and pushed apart if the map shows it more similar ($q \gt p$).

Three ways to say it:

  • Picture: strong springs between true neighbours, weak pushing between everyone else.
  • Numbers: the same 10× mistake costs 0.46 for a neighbour pair but only about 0.005 for a stranger pair.
  • Slogan: t-SNE punishes separating friends, and shrugs at seating strangers together.

Each pair contributes $p_{ij}\ln(p_{ij}/q_{ij})$ to the KL score ($\ln$ is the natural log).

  1. Neighbours drawn too far apart: $p = 0.2$ but the map gives $q = 0.02$ (10 times too small). Cost $= 0.2 \times \ln(0.2/0.02) = 0.2 \times \ln 10 = 0.2 \times 2.303 = 0.46$.
  2. Strangers drawn too close: $p = 0.002$ but the map gives $q = 0.02$ (10 times too big). Cost $= 0.002 \times \ln(0.1) = 0.002 \times (-2.303) = -0.0046$. Tiny. (A single term can be negative; the total over all pairs never is.)
  3. Same size of mistake, about 100 times less concern. This is why t-SNE keeps local structure and lets global structure drift.
  4. Direction matters. Take $P = (0.8, 0.15, 0.05)$ and $Q = (0.4, 0.15, 0.45)$. $KL(P\|Q) = 0.8\ln 2 + 0 + 0.05\ln\tfrac19 = 0.5545 - 0.1099 = 0.445$. Swap the roles: $KL(Q\|P) = 0.4\ln\tfrac12 + 0 + 0.45\ln 9 = -0.2773 + 0.9888 = 0.711$. Different numbers: KL is not a distance.

The KL divergence from $Q$ to $P$ (both probability distributions over the same items) is

$$KL(P\,\|\,Q) = \sum_{i \ne j} p_{ij} \ln\frac{p_{ij}}{q_{ij}}.$$
  • In plain words: the average extra "surprise" you get if you use $Q$ to describe things that really follow $P$. It is always $\ge 0$, and it is $0$ only when $Q = P$.
  • It is not symmetric: $KL(P\|Q) \ne KL(Q\|P)$ in general. t-SNE minimizes $KL(P\|Q)$ with the data table $P$ first. Which direction you minimize changes the behaviour; Chapter 6.11 treats this in depth for variational inference.
  • The gradient (the direction in which moving $y_i$ increases the score fastest) is $$\frac{\partial\, KL}{\partial y_i} = 4\sum_{j} (p_{ij} - q_{ij})\,\frac{y_i - y_j}{1 + \|y_i - y_j\|^2}.$$
  • Read it as springs: gradient descent moves $y_i$ against the gradient, so if $p_{ij} \gt q_{ij}$ it moves toward $y_j$ (attraction); if $q_{ij} \gt p_{ij}$ it moves away (repulsion). The factor $1/(1 + d^2)$ makes the force fade at long range.
Why do we need it?

We need one number that says how well the map copies the data's neighbourhoods, so that an optimizer can improve it. KL gives that number and a simple gradient, and its weighting by $p_{ij}$ is exactly what makes t-SNE focus on neighbours.

Where is it used?

The t-SNE objective (scikit-learn reports it as kl_divergence_), variational inference in your NumPyro models (minimizing $KL(q\|p)$, Chapter 6.11), the ELBO, knowledge distillation, and cross-entropy losses (cross-entropy = entropy + KL).

How is it used?

Read tsne.kl_divergence_ after fitting: lower means a closer copy. Compare it between runs on the same data with the same perplexity (for example different seeds). Do not compare it across perplexities, because $P$ itself changes.

Eight points come from two groups in 5-D (1–4 and 5–8). Green lines join pairs that are true neighbours (large $p_{ij}$). The purple arrows are the forces (the negative gradient) on each map point. Press 20 gradient steps a few times: the score falls and the groups separate. Then drag point 1 into the orange group: its arrow points back home and the KL score jumps. Notice how the two groups keep drifting apart slowly: nothing in the score fixes the size of the gap.

"KL divergence is the distance between P and Q."

It is not a distance: it is not symmetric ($KL(P\|Q) \ne KL(Q\|P)$) and it does not obey the triangle rule. It is a "how badly does Q describe P" score that is 0 only when they match.

"Minimizing KL makes the map keep all distances."

Each pair is weighted by $p_{ij}$. Far-apart pairs have $p_{ij} \approx 0$, so the map can put them almost anywhere at almost no cost. Only neighbourhoods are tightly controlled.

"A lower KL value means a better perplexity."

Changing the perplexity changes the target table $P$, so the KL values are scores on different tests. Compare KL only between runs with the same data and the same perplexity.

"t-SNE preserves the structure of the data."

"t-SNE preserves local structure: who is near whom. Its KL objective weights each pair by $p_{ij}$, so putting true neighbours far apart is expensive, but the placement of far-apart groups costs almost nothing."

Model answer: "t-SNE minimizes $KL(P\|Q)$ between Gaussian neighbour probabilities in the data and Student-t similarities in the map. Because the loss is weighted by $p_{ij}$, it mostly enforces that neighbours stay neighbours; large distances and the global arrangement are only weakly constrained, so I read neighbourhoods and clusters from a t-SNE plot, not distances."

$KL(P\|Q) = \sum p_{ij}\ln(p_{ij}/q_{ij}) \ge 0$, $= 0$ iff $P = Q$; not symmetric. Full story: Chapter 6.11.

Gradient $4\sum_j (p_{ij}-q_{ij})(y_i-y_j)/(1+\|y_i-y_j\|^2)$: springs, pull if $p \gt q$, push if $q \gt p$.

Trap: weighted by $p$, so far-apart pairs barely matter: local yes, global no.

Quick check: a pair has $p_{ij} = 0.05$ and the map gives $q_{ij} = 0.01$; the points are at $y_i = (0, 0)$ and $y_j = (2, 0)$. Which way does this pair push $y_i$?

Its term in the gradient is $4 \times (0.05 - 0.01) \times (y_i - y_j)/(1 + 4) = 4 \times 0.04 \times (-2, 0)/5 = (-0.064, 0)$. Gradient descent moves against the gradient: with step size 10, $y_i$ moves by $(+0.64, 0)$, toward $y_j$. Attraction, because the data says they are more similar than the map shows. If instead $q_{ij} = 0.09$, the sign flips and $y_i$ moves $0.64$ away.

The whole algorithm, running live next to PCA core

Put the pieces together. First, measure who is close to whom in the data (the table $P$, with one bell width per point from the perplexity). This is done once. Then start the map with all points huddled near the middle, and repeat many times: compute the map table $Q$, feel the springs (the KL gradient), and nudge every point a little.

One trick helps a lot at the start: for the first 250 moves, all the data similarities are multiplied by 12 (early exaggeration). Friends pull each other very hard, so each group shrinks into a tight ball that can slide past other balls without getting tangled. Then the exaggeration is switched off and the map relaxes into its final shape.

Three ways to say it:

  • Picture: a pile of magnets shaken gently a thousand times until it settles into islands.
  • Numbers: 120 points make 7 140 pairs; every step computes 7 140 small forces; scikit-learn takes 1 000 steps.
  • Slogan: compute $P$ once, then walk the map downhill on the KL score.

The bookkeeping for 120 points with scikit-learn's default settings.

  1. Pairs: $n(n-1)/2 = 120 \times 119/2 = 7\,140$.
  2. Table $P$: for each of the 120 points, a binary search for $\sigma_i$ (a few dozen tries each) to reach perplexity 30. Done once.
  3. Step size ("learning rate", learning_rate='auto'): $\max\!\big(n/(12 \times 4),\ 50\big) = \max(2.5,\ 50) = 50$.
  4. Steps 1–250: every $p_{ij}$ multiplied by 12, momentum 0.5. Steps 251–1000: plain $p_{ij}$, momentum 0.8.
  5. Work with exact forces: $7\,140 \times 1\,000 \approx 7.1$ million pair updates, instant on a laptop. For $n = 100\,000$ there are about $5 \times 10^9$ pairs per step, which is why scikit-learn's default method='barnes_hut' groups far-away points into cells (roughly $n\log n$ work per step).

t-SNE (van der Maaten and Hinton, 2008), as in scikit-learn's TSNE:

  1. Inputs: data $x_1, \dots, x_n$ (usually standardized; often reduced to about 50 dimensions with PCA first) and a perplexity.
  2. Compute $p_{j\mid i}$ with $\sigma_i$ set by binary search, then $p_{ij} = (p_{j\mid i} + p_{i\mid j})/2n$.
  3. Start the map: init='pca' (the default: the PCA picture shrunk so that its first axis has standard deviation $10^{-4}$) or init='random' (tiny random positions, standard deviation $10^{-4}$).
  4. Repeat: compute $q_{ij}$ and the gradient $g$ of $KL(P\|Q)$; update $\Delta \leftarrow m\,\Delta - \eta\,(\text{gains} \odot g)$ and $y \leftarrow y + \Delta$, where $m$ is the momentum, $\eta$ the learning rate, and each coordinate has its own "gain" that grows by 0.2 while the next step would keep going in the same direction as the last move, and shrinks by a factor 0.8 when the direction flips (never below 0.01).
  5. Early exaggeration: $p_{ij} \times 12$ for the first 250 steps. Default total: 1 000 steps (max_iter), stopping earlier if the score stops improving for 300 steps.
  6. Outputs: the map positions (embedding_) and the final score (kl_divergence_).

The widget below runs exactly this with exact forces. It was checked against scikit-learn's method='exact': the same $P$ and the same gradient to many digits; the runs then drift apart slightly (the process is very sensitive to rounding) and end with KL values within about 2% of each other. The widget stops at 500 steps by default because these small maps barely change after that.

Why do we need it?

To actually look at data with more than three features: to see whether groups exist, whether known labels separate, and whether some points sit in an unexpected island. PCA can hide all of that when the structure is curved or spread over many directions.

Where is it used?

sklearn.manifold.TSNE, openTSNE and FIt-SNE for large data, embedding projectors for word vectors, single-cell biology atlases, checks of what a neural network's hidden layer has learned, and exploration of customer or segment features.

How is it used?

Z = TSNE(perplexity=30, random_state=0).fit_transform(X), then scatter-plot Z coloured by a known label. Make several plots (two or three perplexities, two or three seeds) and only trust patterns that appear in all of them.

Data X n points × D Table P (once) Gaussian, σᵢ from the perplexity repeat 1 000 times 1. Q from the map Student-t 1/(1+d²) 2. gradient of KL(P‖Q) 3. move every point momentum + gains first 250: P × 12 Map n × 2 start: PCA picture shrunk to sd 10⁻⁴, or random
t-SNE in one picture. The data side (P) is computed once. The map side is a loop: compute the Student-t similarities Q, take the KL gradient, move the points. The result is a list of 2-D positions, not a formula.

Left: PCA of the data (a straight shadow). Right: t-SNE after 500 steps. Colours show labels that the methods never see. Press Watch from the start to see the map grow: tight balls during early exaggeration (first 250 steps), then relaxing. Switch the data to coiled spring: PCA stacks the two loops on each other, t-SNE unrolls the spring (and sometimes tears it into pieces). Try 2 groups × 3 sub-groups at perplexity 5 and then 50: low perplexity scatters the six sub-groups, higher perplexity shows the two big groups. With random start, New seed gives a different-looking map with almost the same KL.

"t-SNE learns a function from the data space to the plane."

It learns $n$ positions, one per point you gave it. There is no formula you can apply to a new point (more on this below).

"Run it until the KL score reaches 0."

A 2-D map can almost never copy $P$ exactly, so KL stays above 0. Run until it stops improving (scikit-learn: 1 000 steps by default, with an early stop after 300 steps without progress).

"The horizontal axis of a t-SNE plot is like PC1."

The axes mean nothing. Rotating, flipping or shifting the whole map leaves every distance, and so the KL score, unchanged. Do not label t-SNE axes with feature names.

In an A/B framework like yours, suppose each user carries a vector of pre-experiment features (activity, tenure, device, region, past conversion). A t-SNE map coloured by your segment label answers a quick question before you trust a hierarchical model: do the segments really look different, or do they overlap so much that strong partial pooling across them is natural? It can also expose a small island of odd users (for example bots or a tracking bug) worth checking before an experiment. It does not tell you how different two segments are; compare their posterior effects or the raw features for that.

t-SNE: $P$ once (perplexity → $\sigma_i$), then gradient descent on $KL(P\|Q)$ with momentum and gains; early exaggeration ×12 for 250 steps; 1 000 steps; init='pca' by default.

Output: positions only (no formula), axes without meaning. Cost $O(n^2)$ per step exact; Barnes–Hut about $n\log n$.

Trap: one run, one perplexity is one opinion. Make several maps before you believe a pattern.

Quick check: what is early exaggeration for, and what would you expect to see if it were switched off?

Multiplying all $p_{ij}$ by 12 makes the attraction between true neighbours dominate at the start, so each group gathers into a tight ball that can move freely past the others. Without it, groups can get stuck tangled with each other early on and end up split into pieces across the map. You can see the tight-ball phase in the widget with Watch from the start.

Reading a t-SNE map: cluster sizes and gaps are not to scale core

Think of a subway map. It tells you which stations are on which line and in which order, and that is very useful. But every station is the same size, the lines are drawn straight, and two stations that look 1 cm apart may be 500 m or 5 km apart in the city. Nobody measures a subway map with a ruler.

A t-SNE map is a subway map of your data. It is faithful about neighbourhoods (who sits next to whom, which points form a group). It is not faithful about sizes (how spread out a group is) or gaps (how far apart two groups are). Each of the earlier steps caused this: perplexity gives every point the same number of neighbours, which equalizes tight and loose groups; the heavy tail stretches moderate distances by huge factors; and the KL score hardly cares where far-apart groups go.

Three ways to say it:

  • Picture: a subway map, not a street map.
  • Numbers: in the data below, group B is 6.6 times wider than A, and C is 3.3 times farther from A than B is; in the t-SNE map they look the same size (ratio 1.0) and about equally far apart (ratio 1.1).
  • Slogan: read the neighbourhoods, never the ruler.

Three groups of 40 points in 5 dimensions: A is tight, B is wide, C is far away.

  1. In the data: A sits at the origin with spread (root mean squared distance from its centre) $0.66$; B is $12.1$ away with spread $4.35$; C is $40.1$ away from A with spread $2.08$.
  2. True ratios: gap A–C ÷ gap A–B $= 40.07/12.11 = 3.31$. Spread of B ÷ spread of A $= 4.35/0.66 = 6.56$.
  3. PCA map (a straight projection; here both gaps lie in its plane): gap ratio $3.31$, spread ratio $6.54$. The shadow keeps both.
  4. t-SNE map (perplexity 30, random start, seed 1, 500 steps): gaps A–B $= 20.9$ and A–C $= 23.3$, ratio $1.11$; spreads $1.04$ and $1.04$, ratio $1.01$. Three equal-looking blobs at roughly equal distances.
  5. Why the sizes became equal: every point got the bell width that gives it 30 effective neighbours, so a tight group and a wide group of 40 points produce the same pattern of $p_{ij}$.
  6. Why the gaps became equal: all between-group $p_{ij}$ are almost 0, so the KL score only asks the groups not to overlap; it does not care whether they are 20 or 200 units apart.

What a t-SNE map can tell you (with care, and after checking several perplexities and seeds):

  • which points are neighbours of which;
  • that a set of points forms a well-separated group, when that group appears in every run;
  • whether known labels (segments, devices, weeks) separate or mix.

What it cannot tell you:

  • cluster sizes (spread) and densities: they are equalized by the per-point $\sigma_i$;
  • distances between clusters that are far apart: they are barely constrained by KL and stretched by the heavy tail;
  • the axes, the orientation, which side a group is on;
  • that small clumps are real: at low perplexity, pure noise breaks into clumps;
  • the "true" number of clusters from one setting.

Rule: every finding from the picture is checked in the original space (distances, group statistics, a clustering method, or the raw features).

Why do we need it?

t-SNE pictures look convincing, so they are easy to over-read. Knowing exactly which parts of the picture carry information stops you from presenting a stretched gap or an equalized blob as a finding.

Where is it used?

Every time a t-SNE (or UMAP) plot appears in a report: customer segment maps, embedding visualizations, single-cell atlases (where over-reading cluster sizes and distances is a well-known mistake), and model-debugging plots of hidden layers.

How is it used?

Before you say "B is bigger" or "C is farthest", compute those quantities in the original data (group spreads, centre distances) or look at the PCA picture. Use the t-SNE map to spot candidates, the numbers to confirm them.

In the data A B C B 6.6× wider than A · C 3.3× farther than B t-SNE In the t-SNE map A B C same sizes, similar gaps: not to scale
The illusion in one picture. The groups (and their membership) survive, but the map draws them with equal sizes and similar gaps. Sizes and far-apart distances in a t-SNE map are not measurements.

Same data as the example: A tight (blue), B wide (orange), C far away (teal), in 5 dimensions. Left: PCA. Right: t-SNE. The table compares two unit-free ratios in the data, the PCA map and the t-SNE map. At perplexity 30, t-SNE draws three equal blobs at similar distances. Press New seed: the layout changes, the illusion does not. Raise the perplexity above 40 (the group size): more of the true layout returns, but the sizes are still wrong. Drop it to 5: each group breaks into clumps.

These 120 points are pure random noise in 10 dimensions: one round cloud, no groups at all (left, PCA). Pick perplexity 2: the t-SNE map (right) breaks into many little clumps that look like real clusters. At 30 or 50 it becomes one even blob again, which is the honest answer. Press New sample to draw a fresh noise cloud: the clumps move, which tells you they were never real.

"The map shows five islands, so the data has five segments."

The number of islands depends on the perplexity and can even come from noise. Confirm groups in the original data (for example with a clustering method and its own diagnostics) and across several perplexities.

"This island is tiny, so this segment is very homogeneous."

t-SNE equalizes density: tight and loose groups end up similar in size. Measure the spread in the original features.

"Points at the edge of the map are outliers."

Position on the map has no global meaning. Check an outlier in the original space (for example by its distance to its nearest neighbours there).

"These two clusters are far apart in the t-SNE plot, so they are very different." · "This cluster is bigger, so that group is more varied."

"In t-SNE, cluster sizes and the distances between far-apart clusters are not meaningful. I can say these points form a group and which points are neighbours; how big or how far is something I check in the original space."

Model answer: "Perplexity gives every point the same effective number of neighbours, so dense and sparse groups come out the same size. The KL objective barely penalizes the placement of far-apart pairs, and the heavy-tailed Student-t kernel stretches moderate distances non-linearly, so gaps between clusters are not proportional to anything. A t-SNE plot is a map of neighbourhoods; for sizes and distances I use the data itself or PCA."

Meaningful: neighbours, membership of well-separated groups (if stable across perplexities and seeds), whether labels mix.

Not meaningful: cluster sizes, densities, gaps between far-apart clusters, axes, orientation, clumps at low perplexity.

Trap: a beautiful t-SNE plot is a hypothesis, not a result. Confirm in the original data.

Quick check: a colleague says "segment C is the most different, it is the farthest island in the t-SNE plot". What two checks do you ask for?

(1) The distances between segment centres (or a distance between their feature distributions) computed in the original, standardized features, or at least in a PCA picture, which does keep large distances when they lie in its plane. (2) The same t-SNE plot at a few perplexities and seeds: if C's position changes, the "farthest" impression was an artefact. In the example above, the t-SNE map put C at almost the same distance as B although it is 3.3 times farther in the data.

Randomness, seeds and the starting layout

The KL score is a bumpy landscape with many valleys, not one smooth bowl. (In the words of the Optimization guide, the problem is non-convex.) Gradient descent walks downhill from wherever it starts and stops in the nearest good valley. Start somewhere else and you may end in a different valley: the same groups, but arranged differently around each other, rotated or mirrored, often with almost the same score.

Two choices control the start. A random start drops the points at tiny random positions; the seed (the number that fixes the random generator) decides which positions. A PCA start begins from the shrunk PCA shadow, which already contains the coarse layout of the data; it makes runs repeatable and tends to keep big groups in a sensible arrangement.

Three ways to say it:

  • Picture: marbles dropped on a bumpy tray settle in different hollows each time.
  • Numbers: four random seeds on the same data reach KL 0.120, 0.113, 0.111 and 0.113: equally good fits, four different-looking pictures.
  • Slogan: a seed makes a picture repeatable, not correct.

Data: 2 big groups, each made of 3 sub-groups of 15 points (90 points, 10 features). In the data, every sub-group's nearest sub-group belongs to its own big group. Perplexity 10, 500 steps.

  1. Random start, seeds 1, 2, 3, 4: final KL $= 0.120,\ 0.113,\ 0.111,\ 0.113$. The fits are almost equally good.
  2. Count how many of the 6 sub-groups have, in the map, their nearest sub-group from their own big group: $5,\ 3,\ 2,\ 1$. The big two-group layout is kept by some seeds and lost by others.
  3. PCA start: every run gives KL $= 0.116$ and the count $4$. Identical runs, because nothing random is left (scikit-learn's PCA start is also nearly repeatable).
  4. Raise the perplexity to 30: every run, random or PCA, gives KL $\approx 0.011$ and keeps all $6$. More neighbours per point means more of the big picture enters $P$.
  5. Lesson: the seed and the start change the arrangement of groups; the perplexity changes how much large-scale structure the method even tries to keep.
  • t-SNE's objective is non-convex: many local minima with similar KL. The result depends on the start and on the random numbers used.
  • random_state (scikit-learn) fixes the seed, so the same code on the same data gives the same map. It does not make the map "right".
  • init='pca' (the default since scikit-learn 1.2; it used to be 'random') starts from the PCA projection rescaled to standard deviation $10^{-4}$. A 2021 study by Kobak and Linderman showed that the start matters a lot for keeping global structure, in both t-SNE and UMAP.
  • Any rotation, reflection or shift of a whole map gives exactly the same KL. So two maps can be equally good and still look like mirror images.
  • Compare KL values between runs only for the same data and the same perplexity.
Why do we need it?

If you do not control the randomness, you cannot reproduce a plot, and you may mistake a lucky (or unlucky) arrangement for a property of the data. Knowing which parts change between runs tells you which parts to ignore.

Where is it used?

TSNE(random_state=0, init='pca') in scikit-learn; UMAP(random_state=42, init='spectral') in umap-learn (where fixing the seed also turns off parallel speed-ups); any report or notebook that must be re-run and give the same figure.

How is it used?

Fix random_state for the figure you publish, keep the PCA start, and before you draw conclusions run two or three other seeds and perplexities. Say in the report which settings you used. Trust only what is stable.

The same 90 points (cool colours = big group 1, warm colours = big group 2), four t-SNE runs. With random start, each seed arranges the six sub-groups differently, with nearly the same KL. Switch to PCA start: all four runs become identical. Then raise the perplexity to 30: the two big groups come out in every run. Press New seeds to try four other seeds.

"Two seeds gave different pictures, so t-SNE failed."

Different arrangements of the same groups are expected: the objective has many equally good valleys, and rotations or mirror images do not change it. Look at what is stable across runs.

"Setting random_state makes the result trustworthy."

It makes the result repeatable. A repeatable artefact is still an artefact.

"A PCA start guarantees the big layout is right."

It helps, but at low perplexity the widget's PCA start still loses part of the two-group layout. The start and the perplexity together decide how much global structure survives.

Non-convex KL: different seeds/starts → different arrangements with similar KL; rotation and reflection never matter.

random_state = repeatable; init='pca' (default) = repeatable and keeps more coarse layout.

Trap: trust only patterns that survive several seeds and perplexities.

Quick check: run A has KL 0.113 and run B (same data, same perplexity, other seed) has KL 0.120, and the pictures look different. Which picture is "right"?

Neither is the right one in a strong sense. A has a slightly better fit, so if you must publish one, A is a reasonable choice, but the difference is small and both are valid local optima. What you report are the features both share: which groups exist and which points sit together. Their arrangement on the page is not a finding.

Why t-SNE is a poor production feature reducer core

A t-SNE map is like a seating chart drawn by hand for one party. It is a good picture of who sits with whom tonight. But when ten new guests arrive, you do not just add ten chairs: you redraw the whole chart, and the old guests end up in new places. And the chart's "coordinates" (third table from the window) mean nothing to anyone outside the party.

Now imagine a model that learned "guests at the window table buy more". Tomorrow the chart is redrawn and the window table holds different people. The model's input has silently changed meaning. That is what happens if you feed t-SNE coordinates into a production model.

Three ways to say it:

  • Picture: a seating chart for one party, not an address system.
  • Numbers: adding 20 customers to 100 and re-fitting moved the old customers by up to 29% of the map's size (even after the best rotation and scaling); PCA moved them by 0%.
  • Slogan: t-SNE is for eyes, not for models.

100 customers from 4 segments (10 features) get a t-SNE map. Then 20 new customers arrive (5 per segment).

  1. t-SNE has no formula for a new point, so the only option is to re-fit on all 120.
  2. Compare the 100 old customers' positions before and after. To be fair, first shift, rotate, mirror and rescale the new map to match the old one as well as possible (a Procrustes alignment).
  3. Even after that alignment, the old customers moved on average by $0.29$ of the map's radius (seed 1). With seeds 2 and 3 the numbers were $0.07$ and $0.19$: how much things move is itself random.
  4. PCA fitted on the 100: new customers are placed with the same formula $y = W^\top(x - \mu)$, so the old customers move by exactly $0$. Even re-fitting PCA on all 120 moves them by only $0.012$ of the radius.
  5. A model trained on yesterday's t-SNE coordinates would receive features with a different meaning today. A model trained on PCA coordinates would not.

Reasons t-SNE should not be a generic feature-reduction step in a production pipeline:

  1. No transform for new points. scikit-learn's TSNE has fit_transform but no transform. The map exists only for the points it was fitted on. (Extensions exist: openTSNE can place new points into a fixed map, and "parametric t-SNE" trains a neural network to imitate the map. They add complexity and still inherit the next problems.)
  2. Not deterministic, not stable. The result depends on the seed, the start, the perplexity and on every other point; adding data rearranges the map; axes can rotate or flip between fits.
  3. Distorted global structure. Sizes, densities and large distances are wrong by design, so downstream models see misleading geometry.
  4. Built for 2–3 output dimensions and costly. Exact t-SNE costs $O(n^2)$ per step; scikit-learn's Barnes–Hut method only allows fewer than 4 output dimensions.
  5. Leakage risk. Fitting t-SNE on training and test rows together and then using the coordinates as features lets the test rows shape the training features.

Better choices for features: the original (standardized) features, PCA (it has a transform and interpretable loadings), a model's own learned representation, or a method built to embed new points (an autoencoder; UMAP's transform, with care).

Why do we need it?

A beautiful t-SNE plot tempts people to "use those clusters" or "use those 2 coordinates" as inputs. Knowing why this breaks protects a production model from features whose meaning changes every time they are recomputed.

Where is it used?

Code reviews and design reviews of ML pipelines, feature stores, interview questions on dimensionality reduction, and any team that moves a notebook exploration (where t-SNE is fine) into a scheduled job (where it is not).

How is it used?

Keep t-SNE in the exploration notebook. If you need fewer features, use PCA fitted on the training data and apply its transform to new data, or regularize the model. If you found groups with t-SNE, define them with a reproducible rule (a clustering model or a business rule) on the original features.

✓ Exploration features t-SNE map human looks check on theoriginal data ✗ Production feature today's rows re-fit t-SNEnew layout model trained onyesterday's map features changed meaning overnight
t-SNE belongs at the top: a picture for a person, followed by checks on the real data. Used as an input feature (bottom), it changes meaning every time it is re-fitted, and it cannot map new rows at all without re-fitting.

Left: the map fitted on 100 customers. Right: the map after 20 new customers (hollow rings) arrive. The purple numbers 1–4 follow the same four old customers. With t-SNE the whole map is redrawn: the tracked customers and whole segments move. Press New seed a few times: the amount of movement is itself random. Switch to PCA: the new customers are placed with the same formula and the old ones do not move at all.

"t-SNE found 4 clean clusters, so I will use the 2 t-SNE coordinates as features."

The coordinates change meaning with every re-fit and cannot be computed for new rows. Use the original features, PCA, or a reproducible clustering rule defined on the original features.

"I fitted t-SNE on all my data, including the test set, and the classifier scores went up."

The test rows helped shape the training features: that is leakage, and the score is optimistic. Any unsupervised step must be fitted on training data only, and t-SNE cannot then be applied to the test rows at all.

"Just fix random_state and t-SNE is stable."

The seed fixes the random numbers, not the data. As soon as rows are added or removed, the whole map can rearrange.

Keep t-SNE out of both models. In your forecasting model, the regressors $X_t$ must be computable at the forecast origin for every future date with a fixed recipe; t-SNE coordinates fail that test (no transform, and a re-fit changes their meaning). If you have too many correlated regressors, use PCA fitted on the training period and applied with its transform, or shrinkage priors on $\beta$. In an A/B framework like yours, segment definitions must be fixed before the experiment; a cluster seen in a t-SNE map can inspire a segment, but the segment itself should be a written rule on the original features.

"We reduced the features with t-SNE and then trained a classifier on the 2-D output."

"We used t-SNE to look at the data; for the model we used PCA (or the raw features), because t-SNE has no out-of-sample transform, is not stable across fits, and distorts global geometry."

Model answer: "t-SNE optimizes the positions of the given points, not a mapping, so it cannot embed new data without re-fitting, and a re-fit gives a different, arbitrarily rotated layout. It also deliberately distorts densities and large distances. That makes its output a poor and unstable feature. For production I want a fitted transform, like PCA, applied identically to training and new data."

t-SNE in production: no transform, unstable (seed, perplexity, every added row), distorted global geometry, $\le 3$ dims with Barnes–Hut, leakage if fitted on test rows.

Use instead: original features, PCA with transform, a learned representation; define clusters with a reproducible rule.

Trap: "it looked like clean clusters" is not a reason to use the coordinates.

Quick check: can you call TSNE(...).fit(X_train) and then .transform(X_test) in scikit-learn?

No. TSNE only has fit and fit_transform; there is no transform method, because t-SNE optimizes the positions of the training points directly and never learns a mapping. (You can check: hasattr(TSNE(), "transform") is False.)

UMAP, conceptually, and how it differs from t-SNE

UMAP (Uniform Manifold Approximation and Projection, 2018) has the same goal as t-SNE: a 2-D picture that keeps neighbours. It gets there in a different way, more like drawing a social network.

  1. Build a friendship network. Each point links only to its $k$ nearest neighbours. Its closest neighbour always gets full strength 1; farther neighbours get weaker links, measured against that point's own local distances, so crowded and empty regions are treated alike.
  2. Agree on each link. If A counts B as a friend with strength 0.6 and B counts A with 0.3, the link gets $0.6 + 0.3 - 0.6 \times 0.3 = 0.72$: "at least one of them counts the other".
  3. Lay out the network. Linked points pull together; a few randomly chosen unlinked pairs push apart. The pulling curve looks like t-SNE's Student-t, with a shape set by min_dist (how tightly points may pack).

Three ways to say it:

  • Picture: a friendship network drawn with springs.
  • Numbers: neighbours at distances 1.0, 1.5 and 3.0 get link strengths 1, 0.72 and 0.28.
  • Slogan: t-SNE copies probabilities over all pairs; UMAP lays out a sparse neighbour graph.

Point $i$ has its 3 nearest neighbours at distances $1.0$, $1.5$ and $3.0$. In umap-learn this is n_neighbors = 4, because the point counts itself as its own first neighbour.

  1. Local offset: $\rho_i$ = distance to the nearest neighbour $= 1.0$. Subtract it: $0,\ 0.5,\ 2.0$.
  2. Strength of each link: $e^{-(d - \rho_i)/\sigma_i}$. The nearest neighbour gets $e^0 = 1$, whatever $\sigma_i$ is.
  3. Choose $\sigma_i$ so the strengths add up to $\log_2(\text{n\_neighbors}) = \log_2 4 = 2$: solve $1 + e^{-0.5/\sigma} + e^{-2/\sigma} = 2$. A short search gives $\sigma_i = 1.551$.
  4. Strengths: $1$, $e^{-0.5/1.551} = e^{-0.322} = 0.724$, $e^{-2/1.551} = e^{-1.289} = 0.276$. Check: $1 + 0.724 + 0.276 = 2.000$.
  5. Combine directions: with $w_{j\mid i} = 0.6$ and $w_{i\mid j} = 0.3$, $w_{ij} = 0.6 + 0.3 - 0.18 = 0.72$.
  6. In the map, a pair at distance $d$ gets similarity $1/(1 + a\,d^{2b})$. With the default min_dist = 0.1: $a = 1.577$, $b = 0.895$. With $a = b = 1$ it would be exactly t-SNE's curve $1/(1 + d^2)$.
  • Neighbour graph. For each point, its $k$ nearest neighbours (n_neighbors, default 15; it plays the role of perplexity). Weights $w_{j\mid i} = \exp\!\big(-\max(0, d_{ij} - \rho_i)/\sigma_i\big)$ for those neighbours and 0 for everyone else. $\rho_i$ is the distance to the nearest neighbour; $\sigma_i$ is set so that the weights add to $\log_2 k$.
  • Symmetrize with the "fuzzy union": $w_{ij} = w_{j\mid i} + w_{i\mid j} - w_{j\mid i}\,w_{i\mid j}$.
  • Map similarity: $v_{ij} = \big(1 + a\,\|y_i - y_j\|^{2b}\big)^{-1}$, with $a, b$ fitted from min_dist (default 0.1) and spread (default 1).
  • Objective: a cross-entropy over the graph's pairs, $$CE = \sum_{i,j} \Big[ w_{ij}\ln\frac{w_{ij}}{v_{ij}} + (1 - w_{ij})\ln\frac{1 - w_{ij}}{1 - v_{ij}} \Big].$$ The first term pulls neighbours together (like t-SNE's KL). The second term explicitly punishes putting non-neighbours close together, a repulsion that t-SNE's KL has only weakly.
  • Optimization: stochastic gradient descent over the graph edges with negative sampling (a few random pairs pushed apart per edge), no normalization over all $n^2$ pairs, and a spectral start (from eigenvectors of the graph). Number of passes: by default 500 for small data and 200 for large.
  • New points: a fitted UMAP has a transform that places new points into the existing map (approximately).

These are the ideas of the method; the exact recipe lives in the UMAP paper (McInnes, Healy and Melville) and the umap-learn package.

Why do we need it?

t-SNE gets slow on hundreds of thousands of points, cannot place new points, and is usually limited to 2–3 output dimensions. UMAP keeps the "neighbours stay neighbours" idea while being faster, scaling to millions of points, offering a transform, and allowing more output dimensions.

Where is it used?

The umap-learn package (umap.UMAP), single-cell biology (where it largely replaced t-SNE in standard pipelines), embedding visualization for text and images, and as a pre-step for density-based clustering such as HDBSCAN.

How is it used?

umap.UMAP(n_neighbors=15, min_dist=0.1, random_state=0).fit_transform(X). Raise n_neighbors for more global structure, lower it for finer detail; lower min_dist packs clusters tighter. The same reading rules as t-SNE apply: sizes and large gaps are not to scale.

t-SNE UMAP data sideGaussian over all pairsweights on k-nearest-neighbour graph width per pointσᵢ from perplexityρᵢ (nearest) and σᵢ (sum = log₂k) map curve1/(1 + d²)1/(1 + a·d^(2b)), set by min_dist objectiveKL(P‖Q), normalizedcross-entropy: pull + push terms optimizerfull gradient (or Barnes–Hut)SGD + negative sampling default startPCA (scikit-learn ≥ 1.2)spectral (graph eigenvectors) new pointsno transformtransform (approximate) speed / sizeslower; 2–3 output dimsfaster; any number of dims sizes & gapsnot to scalenot to scale either
t-SNE and UMAP side by side. They share the goal (keep neighbours) and the warning (sizes and large gaps are not measurements). UMAP's graph, explicit repulsion term and stochastic optimizer make it faster and give it a transform for new points.

Thirty points: a tight group (left) and a loose group (right). Each line is a link of UMAP's neighbour graph; darker means stronger (after the fuzzy union). Drag the purple probe onto a point to see its own links: the nearest one always has strength 1, the others are scaled by that point's own $\rho_i$ and $\sigma_i$. Change n_neighbors: small values give a sparse graph that can split into pieces; larger values connect the two groups.

The green dashed line is the shape UMAP aims for: similarity 1 up to min_dist, then an exponential fall. The orange curve $1/(1 + a\,d^{2b})$ is the smooth curve fitted to it (this is what UMAP actually uses), and the blue dashed curve is t-SNE's $1/(1 + d^2)$. Pick min_dist = 0: points may sit on top of each other, so clusters pack tightly. Pick 0.5 or 0.8: the curve stays flat longer, so points keep more room and clusters look puffier. The cluster membership does not change, only how the picture packs it.

"UMAP preserves global structure, t-SNE does not."

UMAP pictures often keep more of the coarse layout, but much of that comes from its spectral start; t-SNE with a PCA start keeps more global structure too (Kobak and Linderman, 2021). In both, cluster sizes and large distances remain unreliable.

"UMAP has a transform, so its coordinates are safe production features."

The transform places new points into a fixed map, which fixes one problem. The coordinates still depend on n_neighbors, min_dist and the seed, and still distort densities. Use with care, version the fitted model, and prefer simpler features when they work.

"n_neighbors is the number of clusters."

Like perplexity, it is the size of each point's neighbourhood: small values show fine detail (and can split the graph), large values show more of the big picture.

"UMAP is just a faster t-SNE."

"Both are neighbour embeddings, but they differ in construction: UMAP builds a sparse k-nearest-neighbour graph with locally scaled fuzzy weights, optimizes a cross-entropy that includes an explicit repulsion term, uses SGD with negative sampling and a spectral start, and can transform new points."

Model answer: "I would pick UMAP for large data, when I need to embed new points, or want more than two output dimensions; t-SNE is excellent for a careful 2-D look at moderate data. In either case I read neighbourhoods and cluster membership, not sizes or distances, and I check stability across seeds and the neighbourhood parameter."

UMAP: kNN graph (n_neighbors=15), $w_{j\mid i} = e^{-(d-\rho_i)/\sigma_i}$, weights add to $\log_2 k$; fuzzy union $a + b - ab$; map curve $1/(1+a d^{2b})$ from min_dist=0.1 ($a \approx 1.58, b \approx 0.90$); cross-entropy with pull + push; SGD + negative sampling; spectral start; has transform.

Trap: "UMAP keeps global structure" is only partly true; sizes and gaps are still not to scale.

Quick check: in UMAP's cross-entropy, which term stops two non-neighbours ($w_{ij} = 0$) from being drawn on top of each other?

With $w_{ij} = 0$ the first term vanishes and the second becomes $\ln\frac{1}{1 - v_{ij}}$, which grows without limit as the map similarity $v_{ij} \to 1$ (points on top of each other). That is an explicit push-apart force. In t-SNE, a pair with $p_{ij} \approx 0$ contributes almost nothing directly; repulsion there comes only indirectly through the normalization of $Q$.

PCA vs t-SNE vs UMAP: choosing the right tool core

Think of three ways to show a city. A satellite photo (PCA): honest about big distances and directions, but buildings overlap and side streets hide behind each other. A subway map (t-SNE): perfect for "which stop is next to which", useless with a ruler. A subway map that can add new stations and is quick to redraw for a huge city (UMAP): the same kind of map, built differently.

Which one you want depends on the question, not on which picture looks nicest.

Three ways to say it:

  • Picture: photo for distances and directions, subway map for neighbours.
  • Numbers: in the example of this chapter, PCA kept the true size ratio (6.5) and gap ratio (3.3); t-SNE kept more true neighbours but showed both ratios as about 1.
  • Slogan: features and distances → PCA; pictures of neighbourhoods → t-SNE or UMAP.

You have 50 000 users with 40 standardized features. Three requests arrive.

  1. "Show me whether our segments look different." A picture of neighbourhoods: run PCA to about 30 dimensions (to remove noise and save time), then t-SNE or UMAP to 2-D, coloured by segment. Check what you see with numbers.
  2. "Give the churn model 5 compact features instead of 40." Features for a model: PCA fitted on the training users, then transform for everyone else. Look at the explained variance and the loadings.
  3. "Every morning, put yesterday's new users onto the same picture." New points into a fixed picture: PCA (exact formula) or UMAP's transform; not t-SNE (it has none).
  4. "How far apart are segments A and C?" A distance question: compute it in the original features (or read it from PCA if the two components explain most of the variance), never from t-SNE or UMAP.
PCAt-SNEUMAP
Typelinear projectionnon-linear, positions onlynon-linear, graph layout
Keepslarge-scale variance, distances along the kept directionslocal neighbourhoodslocal neighbourhoods (often a bit more coarse layout)
Deterministicyes (up to the sign of each axis)no (seed, start, perplexity)no (seed, start, n_neighbors)
New pointsexact transformnoneapproximate transform
Axes mean somethingyes: loadings, explained variancenono
Sizes and gapsfaithful in the kept plane (shrunk if variance is lost)not to scalenot to scale
Speedvery fastslow-ish ($n^2$ exact; $n\log n$ Barnes–Hut)fast, scales to millions
Main knobsnumber of components; scalingperplexity, seed, start, stepsn_neighbors, min_dist, seed
Good forfeatures, compression, denoising, whitening, a first lookcareful 2-D pictures of moderate datapictures of big data; embeddings that need new points

A common pipeline uses two of them: standardize → PCA to ~50 dimensions → t-SNE or UMAP to 2-D for the picture.

Why do we need it?

Each method answers a different question. Using t-SNE to answer a distance question, or PCA to look for curved structure, gives confident wrong answers. Matching tool to question is most of the skill.

Where is it used?

Exploratory data analysis notebooks, feature engineering for production models, model debugging (looking at embeddings), single-cell and genomics pipelines (PCA then UMAP), and recommendation and search teams inspecting embedding spaces.

How is it used?

Write the question first ("features?", "neighbours?", "distances?", "new points every day?"), then pick: PCA for features, distances and new points; t-SNE or UMAP for neighbourhood pictures, after PCA, with several settings, followed by checks on the original data.

Pick a goal. The table says which method fits, which is possible with care, and which to avoid, with the reason in one line. Try "Feed compact features to a production model" and "Show whether segments overlap": the answers flip.

"t-SNE separates my classes better than PCA, so t-SNE features will make my classifier better."

Separation in a picture is not predictive power for new rows. t-SNE cannot even compute features for new rows. Compare models with honest validation, using PCA or the raw features.

"PCA failed to show clusters, so there are no clusters."

PCA can stack separate groups when they differ in low-variance directions or lie on curved shapes. Absence in one shadow is not absence in the data.

"Run t-SNE directly on 10 000 raw features."

Standardize, and usually reduce to about 50 dimensions with PCA first: it removes noise and makes the distances, and the run, more reliable and faster.

"PCA is for linear data and t-SNE is for non-linear data; otherwise they do the same thing."

"PCA is a linear map that keeps as much variance as possible and gives a reusable transform; t-SNE optimizes 2-D positions to keep local neighbourhoods and gives only a picture."

Model answer: "PCA finds orthogonal directions of maximum variance; it is deterministic, fast, invertible up to the dropped components, and applies to new data, so I use it for features, compression and a first look. t-SNE matches neighbour probabilities with a heavy-tailed kernel; it reveals clusters and curved structure but distorts sizes and distances, depends on perplexity and seed, and has no transform. I use it, or UMAP, only for visualization, often after PCA."

Features, distances, compression, new rows → PCA. Neighbourhood pictures → t-SNE (moderate data, 2-D) or UMAP (big data, new points).

Pipeline: standardize → PCA (~50) → t-SNE/UMAP (2-D) → check findings on the original data.

Trap: matching the picture to the question matters more than the prettiest picture.

Quick check: a teammate wants "the 2 UMAP coordinates of each store" as regressors in a demand forecast. What do you say?

Ask what question the regressors answer. The coordinates depend on n_neighbors, min_dist and the seed, distort distances, and change if the model is re-fitted, so their effect $\beta$ has no stable meaning. If the goal is fewer, decorrelated regressors, use PCA fitted on the training period (with its transform) or shrinkage priors on $\beta$; if the goal is "stores that behave alike", define groups with a reproducible rule and use them as categorical regressors.

Recap, cheat sheet and practice

  • PCA is a linear shadow: one formula, keeps large-scale variance, can stack far-apart points. t-SNE and UMAP are non-linear neighbour maps: they choose positions so neighbours stay neighbours, mainly for pictures.
  • t-SNE's data side: Gaussian similarities $p_{j\mid i}$ with a per-point width $\sigma_i$ set by binary search so every point has the same perplexity $2^{H}$ (the effective number of neighbours); then $p_{ij} = (p_{j\mid i}+p_{i\mid j})/2n$.
  • t-SNE's map side: Student-t with 1 degree of freedom, $q_{ij} \propto (1+d^2)^{-1}$. Its heavy tail fights the crowding problem and opens gaps between groups, by a stretch that is not to scale.
  • Objective: $KL(P\|Q) = \sum p_{ij}\ln(p_{ij}/q_{ij})$, weighted by $p_{ij}$, so neighbours placed far apart cost a lot and far-apart pairs cost almost nothing: local structure is kept, global structure is weakly constrained. Its gradient is a set of springs (pull if $p \gt q$, push if $q \gt p$).
  • Optimization: gradient descent with momentum and gains, early exaggeration ×12 for 250 steps, 1 000 steps, PCA start by default. The problem is non-convex: seeds and starts change the arrangement.
  • Cluster sizes, densities and distances between far-apart clusters in a t-SNE map are not meaningful; low perplexity can make clumps out of noise. Confirm every finding in the original data.
  • t-SNE is a poor production feature reducer: no transform, unstable across fits and data changes, distorted geometry, leakage risk.
  • UMAP: a fuzzy k-nearest-neighbour graph, a cross-entropy with explicit repulsion, SGD with negative sampling, a spectral start, and a transform. Faster and more scalable; sizes and gaps still not to scale.

Cheat sheet

IdeaFormula / defaultIn words
Data similarity$p_{j\mid i} \propto \exp(-\|x_i-x_j\|^2/2\sigma_i^2)$share of $i$'s attention that goes to $j$
Symmetric similarity$p_{ij} = (p_{j\mid i}+p_{i\mid j})/2n$one distribution over pairs; every point gets weight
Perplexity$2^{H(P_i)}$, $H = -\sum p\log_2 p$; default 30; must be $\lt n$effective number of neighbours; sets each $\sigma_i$
Map similarity$q_{ij} \propto (1+\|y_i-y_j\|^2)^{-1}$Student-t (1 df): heavy tail against crowding
Matching distance$D = \sqrt{e^{d^2}-1}$ (bell $e^{-d^2}$ vs $1/(1+D^2)$)2 → 7.3, 3 → 90: gaps are stretched
Objective$KL(P\|Q) = \sum p_{ij}\ln(p_{ij}/q_{ij}) \ge 0$punishes separated neighbours, ignores far pairs
Gradient$4\sum_j (p_{ij}-q_{ij})(y_i-y_j)/(1+\|y_i-y_j\|^2)$springs: pull if $p \gt q$, push if $q \gt p$
scikit-learn TSNEperplexity 30, exaggeration 12 (250 steps), learning_rate='auto' $=\max(n/48, 50)$, max_iter=1000, init='pca', Barnes–Hutresults: embedding_, kl_divergence_; no transform
Reading rulestrust: neighbours, stable groups; distrust: sizes, gaps, axes, clumps at low perplexitya subway map, not a street map
UMAP weights$e^{-(d_{ij}-\rho_i)/\sigma_i}$ summing to $\log_2 k$; union $a+b-ab$fuzzy neighbour graph, n_neighbors=15
UMAP map curve$1/(1+a\,d^{2b})$; min_dist=0.1 → $a \approx 1.58$, $b \approx 0.90$how tightly points may pack
Check a mapsklearn.manifold.trustworthiness; several seeds and perplexitiesare the map's neighbours true neighbours? is it stable?
Code it · Python

import numpy as np
from scipy.spatial import procrustes
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE, trustworthiness

rng = np.random.default_rng(0)

# 1) Three groups in 5-D, like the widget: A tight, B wide (12 away), C far (40 away)
centres = np.array([[0, 0, 0, 0, 0], [12, 0, 0, 0, 0], [0, 40, 0, 0, 0]], float)
sds = [0.3, 2.0, 1.0]
X = np.vstack([c + s * rng.standard_normal((100, 5)) for c, s in zip(centres, sds)])
y = np.repeat([0, 1, 2], 100)

def ratios(Z):
    """(gap A-C / gap A-B, spread B / spread A): unit-free, so comparable across maps"""
    cen = np.array([Z[y == k].mean(axis=0) for k in range(3)])
    spread = [np.sqrt(((Z[y == k] - cen[k]) ** 2).sum(axis=1).mean()) for k in range(3)]
    gap = lambda a, b: np.linalg.norm(cen[a] - cen[b])
    return round(float(gap(0, 2) / gap(0, 1)), 2), round(float(spread[1] / spread[0]), 2)

# 2) PCA vs t-SNE (scikit-learn defaults: perplexity=30, init='pca', learning_rate='auto', max_iter=1000)
Z_pca = PCA(n_components=2).fit_transform(X)
tsne = TSNE(n_components=2, perplexity=30, random_state=0)
Z_tsne = tsne.fit_transform(X)
print("data  ratios:", ratios(X))
print("PCA   ratios:", ratios(Z_pca))
print("t-SNE ratios:", ratios(Z_tsne), "KL =", round(float(tsne.kl_divergence_), 3), "lr =", tsne.learning_rate_)
# data  ratios: (3.41, 6.2)
# PCA   ratios: (3.41, 6.22)                     <- the shadow keeps gaps and sizes
# t-SNE ratios: (1.22, 1.01) KL = 0.479 lr = 50.0  <- equal sizes, similar gaps: not to scale

# 3) Neighbourhoods: trustworthiness (1.0 = the map's neighbours are true neighbours)
for name, Z in [("PCA", Z_pca), ("t-SNE", Z_tsne)]:
    print(name, "trustworthiness:", round(trustworthiness(X, Z, n_neighbors=10), 3))
# PCA trustworthiness: 0.923
# t-SNE trustworthiness: 0.972                   <- t-SNE wins on local structure

# 4) Perplexity changes the picture (KL values are NOT comparable across perplexities)
for perp in (5, 30, 100):
    t = TSNE(perplexity=perp, random_state=0).fit(X)
    print(f"perplexity {perp:3d}: KL = {t.kl_divergence_:.3f}  ratios = {ratios(t.embedding_)}")
# perplexity   5: KL = 0.758  ratios = (0.95, 0.95)
# perplexity  30: KL = 0.479  ratios = (1.22, 1.01)
# perplexity 100: KL = 0.026  ratios = (3.31, 1.79)   <- more global layout, sizes still wrong

# 5) Seeds and starts: Procrustes disparity (0 = same map up to shift/rotation/flip/scale)
run = lambda init, seed: TSNE(perplexity=30, init=init, random_state=seed).fit_transform(X)
print("random start, seed 0 vs 1:", round(procrustes(run("random", 0), run("random", 1))[2], 4))
print("PCA start,    seed 0 vs 1:", round(procrustes(run("pca", 0), run("pca", 1))[2], 4))
# random start, seed 0 vs 1: 0.0716
# PCA start,    seed 0 vs 1: 0.0                 <- the PCA start makes runs repeatable

# 6) No out-of-sample transform, and perplexity must be below the number of points
print("TSNE has transform:", hasattr(TSNE(), "transform"), "| PCA has transform:", hasattr(PCA(), "transform"))
try:
    TSNE(perplexity=30).fit_transform(X[:20])
except ValueError as e:
    print("ValueError:", e)
# TSNE has transform: False | PCA has transform: True
# ValueError: perplexity (30) must be less than n_samples (20)

# 7) OPTIONAL: UMAP needs `pip install umap-learn` (not installed in this guide's environment,
#    so this part is commented out and its output is not shown).
# import umap
# reducer = umap.UMAP(n_neighbors=15, min_dist=0.1, random_state=0)  # a fixed seed turns off parallelism
# Z_umap = reducer.fit_transform(X)
# Z_new = reducer.transform(X[:5] + 0.1)   # UMAP CAN place new points into the fitted map (approximately)
# print(ratios(Z_umap))                    # read it like t-SNE: sizes and gaps are not to scale
Test yourself

1. What does t-SNE's perplexity control?

Perplexity $= 2^{H(P_i)}$ is a smooth "number of neighbours". t-SNE binary-searches each $\sigma_i$ to hit it, so dense regions get narrow bells and sparse regions wide ones. It is unrelated to the number of clusters.

2. Why does t-SNE use a Student-t with 1 degree of freedom in the map?

The Student-t lives only in the map. There is not enough room in 2-D for all the moderate distances of high-dimensional data; the heavy tail stretches them (distance 2 under a bell matches 7.3 under the t-curve). That stretch is exactly why gaps are not proportional.

3. In a t-SNE plot, cluster A looks twice as wide as cluster B. What can you conclude?

Every point gets the same perplexity, which evens out tight and loose groups; in the chapter's example a group 6.6 times wider than another came out the same size. Sizes in a t-SNE map are not measurements. (The perplexity is one number for the whole map.)

4. Which pair contributes the most to $KL(P\|Q)$?

Contributions $p\ln(p/q)$: $0.2\ln 10 = 0.46$; $0.002\ln 10 = 0.0046$; $0$; and $0.002\ln 0.1 = -0.0046$. Each pair is weighted by $p$, so separating true neighbours dominates. This is why t-SNE keeps local structure.

5. Why is t-SNE a poor choice for reducing features in a production model?

t-SNE optimizes positions of the given points, so new rows need a re-fit, which changes the meaning of the coordinates (and it is non-deterministic and distorts global geometry). It is unsupervised: it never uses labels.

6. Which statement about UMAP is correct?

UMAP's graph, its cross-entropy (whose second term pushes non-neighbours apart), SGD with negative sampling and its transform are the main differences. Sizes and large gaps are still not to scale, and runs vary unless the seed is fixed.

Practice problems

A. A point has three neighbours at distances 1, 1 and 2, and $\sigma = 1$. Compute $p_{j\mid i}$ and the perplexity of this row.
  1. Bell heights $e^{-d^2/2}$: $e^{-0.5} = 0.6065$, $0.6065$, $e^{-2} = 0.1353$. Sum $= 1.3484$.
  2. Shares: $0.6065/1.3484 = 0.450$, $0.450$, $0.1353/1.3484 = 0.100$.
  3. Entropy: $H = 2 \times 0.450 \times \log_2(1/0.450) + 0.100 \times \log_2(1/0.100) = 2 \times 0.450 \times 1.153 + 0.100 \times 3.317 = 1.037 + 0.333 = 1.370$ bits.
  4. Perplexity $= 2^{1.370} = 2.58$: the point behaves as if it had about 2.6 equally important neighbours (at most 3 are possible).
B. Under the bell $e^{-d^2}$, a pair sits at distance $d = 1.5$. At what distance $D$ does the Student-t curve $1/(1+D^2)$ give the same similarity, and what is the stretch factor?

Set $1/(1+D^2) = e^{-2.25}$, so $1 + D^2 = e^{2.25} = 9.488$, $D^2 = 8.488$, $D = 2.91$. The stretch factor is $2.91/1.5 = 1.94$. Compare with $d = 2$, where the factor is $7.32/2 = 3.7$: the farther the pair, the larger the stretch. This is why gaps in the map are not to scale.

C. Compute $KL(P\|Q)$ for $P = (0.5, 0.3, 0.2)$ and $Q = (0.4, 0.4, 0.2)$.

$KL = 0.5\ln\frac{0.5}{0.4} + 0.3\ln\frac{0.3}{0.4} + 0.2\ln 1 = 0.5 \times 0.2231 + 0.3 \times (-0.2877) + 0 = 0.1116 - 0.0863 = 0.0253$. Small and positive: $Q$ is close to $P$. The middle term is negative, but the total cannot be.

D. A pair has $p_{ij} = 0.01$ and $q_{ij} = 0.03$, with $y_i = (1, 1)$ and $y_j = (1, 3)$. Which way does this pair move $y_i$ in one gradient step of size 10?

$\|y_i - y_j\|^2 = 4$, $y_i - y_j = (0, -2)$. Gradient term: $4 \times (0.01 - 0.03) \times (0, -2)/(1 + 4) = 4 \times (-0.02) \times (0, -2)/5 = (0, 0.032)$. The step moves against the gradient: $-10 \times (0, 0.032) = (0, -0.32)$, so $y_i$ goes to $(1, 0.68)$, away from $y_j$. Repulsion, because the map shows them as more similar ($q$) than the data does ($p$).

E. UMAP with n_neighbors = 4: a point's 3 neighbours are at distances 1, 2 and 2. Find $\rho_i$, $\sigma_i$ and the three link strengths.
  1. $\rho_i = 1$ (the nearest distance). Offsets: $0, 1, 1$.
  2. Target sum $\log_2 4 = 2$: $1 + 2e^{-1/\sigma} = 2$, so $e^{-1/\sigma} = 0.5$ and $\sigma_i = 1/\ln 2 = 1.443$.
  3. Strengths: $1$, $0.5$, $0.5$ (sum 2). The nearest neighbour always gets 1; the others are judged relative to this point's own scale.
F. (Interview) A product manager shows a UMAP of customers: "Segment X is far away from everyone else and its island is tiny, so it is a very distinct, very homogeneous group. Let's build a separate model for it." How do you respond?

"The map is useful: it shows that X's customers are neighbours of each other and not mixed with the rest, and that is worth following up. But in UMAP and t-SNE the size of an island and the distance between islands are not measurements: every point gets a similar number of neighbours, which evens out density, and large distances are stretched or squeezed non-linearly. Before deciding, I would compute X's spread and its distance to the other segments in the original standardized features (or a PCA view), check that X stays separate across a few seeds and n_neighbors values, and look at whether the outcome we would model actually behaves differently in X, ideally with a model that pools information across segments rather than a completely separate one."

Appendix A

Glossary

Every important word of this guide in one place, explained in plain English (186 terms). Type in the box to filter: it searches the terms and their explanations. The small numbers after each entry link to the section that teaches it (hover for its title).

All terms, A to Z

No term matches. Try a shorter word, or press / to search the whole guide.

A/A test
An experiment in which both groups get the identical experience, so every significant result is a false positive. At $\alpha = 0.05$ about 5% should come out significant (1 000 tests: about 50 ± 14); a pile-up of small p-values means a broken pipeline. 5.7
Absolute lift (absolute effect)
The difference $\theta_B - \theta_A$ in the metric's own units: 10% → 12% is +2 percentage points. Always say "points" or "percent". 5.7 5.10
Allocation ratio
The share of units sent to each arm. For a fixed total, 50/50 gives the most power; a 90/10 split needs about 2.78 times the traffic for the same power. 5.10 5.10
Alternative hypothesis (H₁)
What you suspect instead of the null: $\theta \ne \theta_0$ (two-sided) or $\theta \gt \theta_0$ / $\theta \lt \theta_0$ (one-sided). Written down, with its direction, before seeing the data. 5.6
ANOVA (one-way) and the F distribution
One test of "are all $k$ group means equal?": $F = \frac{SSB/(k-1)}{SSW/(N-k)}$, between-group over within-group variance, read in the right tail of $F_{k-1,N-k}$. A significant $F$ says "not all equal", not which groups differ; many pairwise t-tests instead inflate false alarms. 5.9
Asymptotic relative efficiency (ARE)
The large-$n$ ratio of two estimators' variances, read as a data ratio. Median vs mean: $2/\pi \approx 0.64$ for Normal data, about 1.62 for $t_3$, 2 for Laplace data. 5.1
Attrition
Assigned units whose outcome is no longer observed. If the amount or kind of leavers differs between arms (differential attrition), the remaining groups are not comparable: analyse by intention-to-treat. 5.11
Average treatment effect (ATE)
$E[Y(1)] - E[Y(0)]$: the average effect over the population. It can be estimated because it is a difference of two averages; under randomization the difference in means is unbiased for it. 5.12 5.10
Average treatment effect on the treated (ATT)
$E[Y(1) - Y(0) \mid T = 1]$: the average effect among the units that actually got the treatment. Equal to the ATE in a randomized experiment; difference-in-differences estimates it. 5.12 5.12
Backdoor path
A path from treatment to outcome that starts with an arrow into the treatment, such as $T \leftarrow U \to Y$. Open backdoor paths create association that is not causation. 5.12
Bayesian power, assurance and operating characteristics
Operating characteristics of a Bayesian decision rule, found by simulating the whole design: how often it ships when there is no effect (false-ship rate) and when the effect is $\delta$ (Bayesian power). Averaging over a prior for the effect gives assurance. 5.10 5.7
Benjamini–Hochberg (BH)
Sort the p-values; find the largest $k$ with $p_{(k)} \le k\alpha/m$ and reject the $k$ smallest. Controls the false discovery rate (for independent or positively dependent tests), not the family-wise error rate. 5.11
Bernstein–von Mises theorem
With plenty of data and a smooth prior that is not zero near the truth, the posterior becomes approximately Normal with sd ≈ SE, so credible and confidence intervals nearly coincide. 5.8
Bias and unbiased estimators
Bias $= E[\hat\theta] - \theta$: how far the centre of the sampling distribution is from the truth; more data does not shrink it. Unbiased means $E[\hat\theta] = \theta$ for every $\theta$; unbiased is not accurate ($X_1$ is unbiased but noisy) and biased is not bad. 5.1 5.4
Bias–variance trade-off
Adding a little bias, by shrinking toward a sensible value, can remove much more variance and lower the MSE. Ridge, partial pooling and Laplace priors all use it; too much shrinkage hurts. 5.1 5.3
Bonferroni correction
Reject only if $p \le \alpha/m$ (adjusted $p = \min(1, mp)$). Controls the family-wise error rate under any dependence, but is conservative. Over $K$ planned looks it gives the boundary $\Phi^{-1}(1 - \alpha/2K)$. 5.11 5.11
Bootstrap
Resample your $n$ rows with replacement many times and recompute the statistic each time; the spread of those values estimates its SE and the shape of its sampling distribution. Resample the independent unit (users, blocks of days); it fails for tiny $n$ and for the maximum. 5.5
CATE (conditional average treatment effect)
The ATE inside a subgroup, $\tau(x) = E[Y(1) - Y(0) \mid X = x]$. The ATE is the size-weighted average of the CATEs; a segment's effect in a hierarchical A/B model is a model of one. 5.12
Centering (for PCA)
Subtracting each column's mean before PCA. Without it, the first direction points at the mean and the "explained" share is inflated. scikit-learn's PCA centres; np.linalg.svd and TruncatedSVD do not. 5.16
Chi-square goodness-of-fit test
$\chi^2 = \sum_k (O_k - E_k)^2/E_k$ with $E_k = N\pi_k$ and df $= K - 1 -$ (parameters fitted). Compares counts with expected shares; the sample-ratio-mismatch check is one. 5.9 5.11
Chi-square test of independence / homogeneity
The same arithmetic on an $r \times c$ table ($E_{ij} = R_iC_j/N$, df $(r-1)(c-1)$) for two designs: one sample with two variables (independence), or groups fixed by design with one variable (homogeneity, the A/B case). For a 2×2 table $\chi^2 = z^2$ without the continuity correction; Cramér's $V$ measures the strength. 5.9
Cholesky factor
A lower-triangular $L$ with $\Sigma = LL^\top$; it exists exactly when $\Sigma$ is positive definite. A full-rank Gaussian guide learns $L$, so its covariance is valid by construction. 5.15 5.15
Cluster sampling
Sampling whole groups (stores, users with many sessions) instead of single units. Cheaper, but each extra unit tells you less: the variance is inflated by the design effect. 5.4
Coefficient (regression)
$\beta_j$: the change in the average of $y$ per one unit of $x_j$ with the other columns held fixed; its units are $y$-units per $x_j$-unit. A conditional association, not a causal effect by default. 5.13
Cohen's d and Cohen's h
$d = (\mu_B - \mu_A)/\sigma$, an effect size in standard deviations (labels 0.2 / 0.5 / 0.8 are behavioural-science rules of thumb; online effects are usually far smaller). For proportions, $h = 2\arcsin\sqrt{p_B} - 2\arcsin\sqrt{p_A}$ (10% → 12%: about 0.064) is what statsmodels' power tools use. 5.7
Collider
A variable caused by two others, $A \to S \leftarrow B$. Filtering or adjusting on it creates a fake association between its causes. 5.12 5.12
Confidence interval
A recipe $[L, U]$ that contains the fixed true parameter in 95% (say) of repeated samples. One computed interval either contains it or not; the 95% describes the method. 5.8 5.8
Confounder
A common cause of the treatment and the outcome. It opens a backdoor path; randomization removes it by design, otherwise you must measure it and adjust. 5.12 5.12
Consistency
$\hat\theta_n$ closes in on $\theta$ as $n \to \infty$: $P(|\hat\theta_n - \theta| \gt \varepsilon) \to 0$. Enough: MSE → 0. A different property from unbiasedness. 5.1
Contamination
Units in one arm receive the other arm's experience. If shares $e_T$ and $e_C$ of the arms really get B, the measured effect is about $\tau(e_T - e_C)$ and the users needed grow like $1/(e_T - e_C)^2$. 5.11
Correlation matrix
$R = D^{-1/2}\Sigma D^{-1/2}$, with $R_{ij} = \Sigma_{ij}/(\sigma_i\sigma_j)$: the covariance matrix of the standardized variables. Unit-free, ones on the diagonal. 5.15
Correlation PCA
PCA of standardized columns, that is the eigen-decomposition of $R$; use it for mixed units. Its eigenvalues add up to $d$. It gives different components from covariance PCA, not rescaled ones. 5.16
Covariance matrix
$\Sigma = E[(\mathbf X - \boldsymbol\mu)(\mathbf X - \boldsymbol\mu)^\top]$: variances on the diagonal, covariances off it; estimated by $X_c^\top X_c/(n-1)$. Rule: $Cov(A\mathbf X + \mathbf b) = A\Sigma A^\top$. 5.15
Covariance structure
The pattern you allow in $\Sigma$ (full, diagonal, compound symmetry, AR(1), block, low-rank + diagonal). It sets how many numbers must be learned: $d(d+1)/2$ for full, $d$ for diagonal. 5.15
Coverage
The probability, over repeated samples, that an interval recipe contains the true value. Nominal = the advertised level; actual coverage can be lower when assumptions fail (Wald at $n = 20$, $p = 5\%$: about 64%). 5.8 5.8
Cramér–Rao lower bound
Under regularity conditions no unbiased estimator has variance below $1/(nI(\theta))$. $\bar X$ reaches it for Normal data; MLEs reach it approximately for large $n$. 5.1 5.2
Credible interval
An interval with $P(a \le \theta \le b \mid D) = 0.95$ under the posterior: a probability statement about $\theta$ given the data, the model and the prior. The Bayesian side is taught in Chapter 6.4. 5.8
Critical value and rejection region
The rejection region is the set of test-statistic values that lead to rejecting $H_0$ (probability $\alpha$ under $H_0$); its boundary is the critical value, $z^* = 1.96$ two-sided at 5%, 1.645 one-sided. $|z| \ge z^*$ exactly when $p \le \alpha$. 5.6
Cross-validation (K-fold)
Fit on $K - 1$ folds, score on the held-out fold, average, repeat for each fold. The usual way to choose λ or the number of PCA components; time series need time-ordered splits instead (Chapter 7.15). 5.3 5.16
CUPED
Variance reduction with a pre-experiment covariate $X$: $Y^{cv} = Y - \theta(X - \bar X)$ with $\theta = Cov(Y,X)/Var(X)$. Same expected effect, variance × $(1 - \rho^2)$; $\rho = 0.7$ needs 51% of the users. Not a bias correction. 5.12
DAG (directed acyclic graph)
A drawing with one node per variable and one arrow per direct cause, with no arrow path leading back to its start. It makes confounders, mediators and colliders visible. 5.12 5.12
Degrees of freedom
Roughly, the number of independent pieces of information left for estimating the noise: $n - 1$ for a one-sample $s$, $n - p$ for a regression with $p$ columns, a fractional Welch value for two groups with unequal variances. 5.9 5.8 5.13
Delta method
A way to get the SE of a function of estimates, such as a ratio metric (clicks per page view) or a relative lift, from their variances and covariances. 5.10 5.8
Design effect and effective sample size
$DEFF = 1 + (m-1)\rho$ for clusters of size $m$ with intra-cluster correlation $\rho$; effective sample size $= n/DEFF$. Naive SEs are too small by $\sqrt{DEFF}$ ($m = 5$, $\rho = 0.3$: a "95%" interval covers about 81%). Count independent units, not rows. 5.4 5.10
Design matrix
$X$: one row per observation, one column per input, first column all 1s. In your forecasting model: 1, $t$, hinge columns $(t - s_j)_+$, Fourier sines and cosines, holiday flags and regressors. 5.13 5.13
Deviance, likelihood-ratio test and AIC
Deviance $D = 2[\ell_{\text{sat}} - \ell(\hat\beta)]$ is the GLM's version of the RSS. The drop in deviance between nested models is approximately $\chi^2$ (likelihood-ratio test); AIC $= -2\ell + 2k$ compares non-nested models such as Poisson vs NB. 5.14
Difference-in-differences (DiD)
(Change in the treated group) − (change in the comparison group); the interaction coefficient in $y \sim \text{treated} \times \text{post}$. Valid under parallel trends; estimates the effect on the treated. 5.12
Duality (interval ↔ test)
A 95% confidence interval contains exactly the values that a two-sided 5% test would not reject: a null value outside the interval ⇔ $p \lt 0.05$ (when both use the same SE). 5.8
Effect size
How big the change is (absolute lift, relative lift, Cohen's $d$), as opposed to how surprising it is (the p-value). Report it with its interval. 5.7
Efficiency (relative)
$Var(\hat\theta_2)/Var(\hat\theta_1)$ for two unbiased estimators of the same parameter: how much more data one needs than the other. Which recipe wins depends on the tails of the data. 5.1
Eigenvalues and eigenvectors of Σ
$\Sigma = V\Lambda V^\top$: the eigenvectors are the axes of the data ellipse, the eigenvalues the variances along them; half-axes of the $k$-sd ellipse are $k\sqrt{\lambda_i}$. 5.15 5.16
Equivalence and non-inferiority tests
Formal versions of "confidently small" (the CI lies inside $(-\delta, +\delta)$, e.g. TOST) and "not worse than a margin" (the CI lies above $-\delta$, used for guardrails; a TOST at 5% uses the matching 90% interval, a 95% interval is a little stricter). 5.10 5.6
Estimate
The value an estimator takes on the data you actually observed: one fixed number, such as 0.106. 5.1
Estimator
A recipe $\hat\theta = g(X_1, \dots, X_n)$ for guessing a parameter. Because the data are random, it is a random variable with a bias, a variance and an MSE. 5.1 5.1
Explained variance and the scree plot
Explained ratio $\lambda_j/\sum_i\lambda_i$; the scree plot shows $\lambda_j$ against $j$. The elbow, a cumulative target (80–95%) or Kaiser's rule are rules of thumb for choosing $k$; validating the downstream use is best. Explained variance of the inputs is not explained variance of a target. 5.16
Exposure
In experiments: a unit actually experienced its assigned arm (assignment ≠ exposure). In count models: the time or number of users over which counts accumulate (see Offset). 5.10 5.14
Fail to reject
The verdict when $p \gt \alpha$: the data are compatible with $H_0$. Never "accept $H_0$": absence of evidence is not evidence of absence, and what a "no" means depends on power. 5.6
False discovery rate (FDR)
The expected share of false positives among the results you call significant, $E[V/\max(R, 1)]$. Controlled by Benjamini–Hochberg. 5.11 5.11
Family-wise error rate (FWER)
The chance of at least one false positive among $m$ tests: $1 - (1 - \alpha)^m$ when all nulls are true and the tests independent (64% for 20 tests at 5%). Controlled by Bonferroni and Holm. 5.11
Fisher information
$I(\theta) = -E[\partial^2\log p(X\mid\theta)/\partial\theta^2]$: how sharply one observation's likelihood peaks. The MLE's SE is about $1/\sqrt{nI(\theta)}$; coin: $I = 1/(\theta(1-\theta))$. 5.2
Fisher's exact test
An exact test for a 2×2 table, used when some expected counts are below about 5. 5.9 5.9
Fundamental problem of causal inference
$Y_i(1)$ and $Y_i(0)$ are never both observed for the same unit, so individual effects are never seen; only averages can be estimated. 5.12
Gauss–Markov theorem (BLUE)
With a correct mean, independent errors and constant variance, OLS is the most precise of all linear unbiased estimators. Normal errors are not needed for this. 5.13
Generalized linear model (GLM)
$g(E[y \mid x]) = x^\top\beta$: a likelihood from the data's support and variance (random component), a linear predictor $\eta = X\beta$, and a link $g$. Linear, logistic, Poisson and Negative Binomial regression are all GLMs. 5.14
Guardrail metric
A metric that must not get worse beyond a tolerance (latency, errors, unsubscribes). Check it as non-inferiority; "guardrail not significant" does not mean "guardrail safe". 5.10
Hash-based assignment
bucket = hash(unit id + experiment salt) mod 100: random-like, sticky (same user, same arm), reproducible, and independent across experiments with different salts. 5.10
Hat matrix and leverage
$H = X(X^\top X)^{-1}X^\top$ projects $y$ onto the column space ($\hat y = Hy$, trace $p$). Its diagonal entries $h_{ii}$ are leverages: how strongly $y_i$ pulls its own fit. High leverage plus a big residual makes an influential point (Cook's distance). 5.13
Heteroscedasticity and weighted least squares
Error variance that changes across observations (a funnel in residual vs fitted). $\hat\beta$ stays unbiased but the usual SEs and bands are wrong. Fixes: robust SEs, weighted least squares with $w_i = 1/\sigma_i^2$, a log transform, or a mean-dependent likelihood such as the NB. 5.13
Holm correction
Step-down Bonferroni: reject $H_{(k)}$ while $p_{(k)} \le \alpha/(m - k + 1)$, stop at the first failure. Same FWER guarantee as Bonferroni, never fewer rejections. 5.11
Hypothesis test
A five-part recipe: $H_0$ and $H_1$, a test statistic, its null distribution, the p-value, and a decision rule (reject if $p \le \alpha$). 5.6
IID and conditionally IID
IID = identically distributed + independent (the joint density factorizes). It gives $Var(\bar X) = \sigma^2/n$ and product likelihoods, says nothing about Normality, and fails for time series, repeated users, clusters, drift and interference. Conditionally IID ($y_i \mid \theta$ iid) is the usual Bayesian-likelihood assumption. 5.4
Indicator (dummy) column
A 0/1 column; its coefficient is the difference in average $y$ between the 1-group and the 0-group, others fixed. With only an intercept and a treatment dummy it equals $\bar y_B - \bar y_A$. An intercept plus a dummy for every category is perfect collinearity (the dummy trap). 5.13 5.13
Instrumental variable (IV) and LATE
An instrument moves the treatment (relevance), is as good as random (independence) and affects the outcome only through the treatment (exclusion, untestable). Wald estimator = effect on $Y$ ÷ effect on $T$; with monotonicity it estimates the local average treatment effect (LATE) among compliers. 5.12
Intention-to-treat (ITT)
Analyse every unit in the arm it was assigned to, exposed or not, with an outcome defined for everyone. It keeps the randomization: diluted, but unbiased for the effect of assignment. 5.11 5.10
Interference (spillover), cluster and switchback designs
One unit's treatment changes another unit's outcome (shared stock, social features): a SUTVA violation. Competition usually makes a naive test overstate the effect, positive spillover understate it. Fix by design: randomize clusters (cities, communities) or switch whole markets over time slots (switchbacks). 5.11
Intra-class correlation (ICC)
$\rho = \tau^2/(\tau^2 + \sigma^2)$: the share of variance between clusters (or users), equal to the correlation between two observations of the same cluster. 5.4 5.10
Inverse probability weighting (IPW)
Weight treated units by $1/e(x)$ and controls by $1/(1 - e(x))$ to imitate a randomized comparison; needs no unmeasured confounding and overlap. With known sampling probabilities, the same idea corrects a biased sample. 5.12 5.4
IRLS
Iteratively reweighted least squares: Newton's method for GLMs. Repeat a weighted least-squares fit, $\beta \leftarrow (X^\top WX)^{-1}X^\top Wz$, with updated weights and working response. 5.14
KL divergence
$KL(P\|Q) = \sum p\ln(p/q) \ge 0$, zero only when $P = Q$, and not symmetric. t-SNE minimizes it; variational inference uses it too (Chapter 6.11). 5.17
Laplace prior
$\beta_j \sim \text{Laplace}(0, b)$: a sharp peak at 0 and heavier tails than a Normal. Its MAP is the lasso with $\lambda = \sigma^2/b$ (½RSS form); its full posterior is shrunk but not sparse. 5.3 5.3 5.3
Lasso (L1) and elastic net
Lasso = least squares + $\lambda\sum_j|\beta_j|$: shrinks and sets small coefficients exactly to 0 (soft-thresholding), so it selects features; its coefficient paths drop to 0 one by one as λ grows. scikit-learn's alpha $= \lambda/n$ in the ½RSS convention. Elastic net adds an L2 term so correlated features share credit. 5.3 5.3
Likelihood
$L(\theta) = p(D\mid\theta)$ read as a function of $\theta$ with the data fixed. Not a probability distribution over $\theta$ (the coin example has area 1/11); only ratios matter. 5.2
Likelihood principle
Inference should depend on the data only through the likelihood. The stopping rule does not enter it, so a posterior is unaffected by peeking, though a decision rule's error rates are not. 5.11
Link function (and canonical link)
$g$ with $g(\mu) = \eta$: it maps the allowed range of the mean onto the whole real line (logit for probabilities, log or softplus for positive means). It transforms the mean, not the data: $\log E[y] \ne E[\log y]$. The canonical link (identity, logit, log) makes the maths simplest. 5.14 5.14
Log-likelihood
$\ell(\theta) = \sum_i\log p(x_i\mid\theta)$: same maximizer as $L$, sums instead of products that underflow; its derivative is the score (zero at an interior maximum) and its negative is the NLL loss. Compare values only on the same data. 5.2
Logistic regression
$\log\frac{p}{1-p} = x^\top\beta$, so $p = 1/(1 + e^{-x^\top\beta})$. $e^{\beta_j}$ is an odds ratio; the negative log-likelihood is the cross-entropy loss. Perfect separation makes the MLE run off to infinity. 5.14
Mahalanobis distance
$d_M = \sqrt{(\mathbf x - \boldsymbol\mu)^\top\Sigma^{-1}(\mathbf x - \boldsymbol\mu)}$: distance in the cloud's own standard deviations, a z-score that respects correlations. For multivariate Normal data $d_M^2 \sim \chi^2_d$. 5.15
Mann–Whitney U test
A rank test for two independent groups; $H_0$: $P(X \gt Y) = P(Y \gt X)$. A test of medians only if both groups have the same shape, and never a test of means. 5.9
MAP (maximum a posteriori)
$\arg\max_\theta[\log p(D\mid\theta) + \log p(\theta)]$: the posterior's peak, i.e. the MLE plus a penalty $-\log p(\theta)$. One point with no uncertainty, and it moves when you reparameterize. 5.2 5.2
Margin of error
$z_{1-\alpha/2} \times SE$: half the width of the interval. It shrinks like $1/\sqrt n$; for a proportion $n = z^2p(1-p)/m^2$. 5.8 5.8
Maximum likelihood estimator (MLE)
$\arg\max_\theta\ell(\theta)$: the parameter value under which the observed data are most likely (not "the most likely parameter"). Invariant to reparameterization; for large $n$ and a correct model, consistent, approximately Normal and efficient. 5.2 5.2
Mean squared error (MSE)
$E[(\hat\theta - \theta)^2] = Bias^2 + Var$: one score combining both kinds of error; RMSE is its square root, in the units of $\theta$. 5.1
Mean-field, low-rank and full-rank guides
Three Gaussian guide families = three covariance structures: diagonal ($2d$ numbers), $WW^\top + \text{diag}(\boldsymbol\psi)$ with $W$ of size $d \times r$ ($d(r+2)$ in NumPyro), and $LL^\top$ ($d + d(d+1)/2$). All are Gaussian approximations; none is exact. 5.15
Mediator
A variable on the causal path, $T \to M \to Y$. Adjusting for it removes part of the effect you want to measure. 5.12
Metric sensitivity
How many units a metric needs to detect a given relative change $r$: about $15.7\,CV^2/r^2$ per arm at $\alpha = 0.05$ and 80% power. Noisy metrics like revenue are expensive. 5.10
Minimum detectable effect (MDE)
The smallest true effect the design detects with the planned power (say 80%) at level $\alpha$: about $(z_{1-\alpha/2} + z_{1-\beta})\,SE$. Effects at the MDE are still missed 20% of the time; halving it needs about 4× the users. 5.10 5.7
Multicollinearity
Columns of $X$ that are (nearly) linear combinations of others. Coefficients get large SEs and can flip sign, while predictions inside the data range stay stable. Measured by the VIF. 5.13 5.16
Multiple testing and the garden of forking paths
Running many tests (metrics, variants, segments, looks) multiplies false positives; so does choosing the metric, segment or window after seeing the data. Fix the families in advance and correct with Bonferroni, Holm or Benjamini–Hochberg. 5.11
Multivariate Normal
$N(\boldsymbol\mu, \Sigma)$: a hill with elliptical contours. Marginals, conditionals and linear combinations are Normal, and zero covariance means independence (true for the MVN only). Sample it as $\boldsymbol\mu + L\mathbf z$. 5.15
Negative Binomial regression (NB2)
Counts with $\log\mu = x^\top\beta$ and $Var = \mu + \mu^2/\alpha$ (a Gamma–Poisson mixture; $\alpha \to \infty$ is Poisson). NumPyro NegativeBinomial2(mean, concentration=α); statsmodels alpha $= 1/\alpha$; SciPy nbinom(n=α, p=α/(α+μ)). 5.14
Non-response bias
Bias when whether a unit answers is related to its outcome: the respondent mean is about $\bar y + Cov(\rho, y)/\bar\rho$. A low response rate on its own is not bias. 5.4
Normal equations
$X^\top X\hat\beta = X^\top y$, i.e. the residuals are perpendicular to every column, so $\hat y$ is the projection of $y$ onto the column space. Solve with lstsq or QR, not an explicit inverse. 5.13
Novelty, primacy and carryover
Temporary boosts (novelty) or dips (change aversion, primacy) because an experience is new, so the short-run effect differs from the long run; carryover is an earlier treatment still acting later (re-used buckets, switchback periods). Plot the lift by day, compare new with returning users, re-randomize. 5.11
Null distribution
The distribution of the test statistic if $H_0$ is true: $N(0,1)$ for z, $t_{n-1}$ for a one-sample t, or built by simulation or by shuffling labels. 5.6
Null hypothesis (H₀)
A precise "nothing happening" claim about a parameter ($p_A = p_B$, $\mu = 175$). A reference assumption to test, not a belief, and never a statement about a statistic. 5.6
Observed (post-hoc) power
Power computed by plugging the observed effect in as if it were true. For a z-test it is a function of the p-value alone (about 0.5 at $p = 0.05$), so it adds nothing. Do not report it. 5.7
Odds ratio
$\frac{p_B/(1-p_B)}{p_A/(1-p_A)} = e^\beta$ in logistic regression; $OR = RR\cdot\frac{1-p_A}{1-p_B}$. Close to the risk ratio only for rare events, so never present it as a relative lift. 5.14
Offset
$\log t_i$ (log exposure) added to a log-link linear predictor with its coefficient fixed at 1, so the model compares rates per unit of exposure instead of raw counts. 5.14
Omitted-variable bias
Leaving out a column that drives $y$ and is correlated with an included column shifts that column's coefficient. Dropping a true variable to cure collinearity causes it. 5.13 5.13
One-sided vs two-sided test
Two-sided counts both tails ($|z| \ge 1.96$ at 5%); one-sided counts one tail chosen in advance ($z \ge 1.645$). Picking the side after seeing the data doubles the false-alarm rate. 5.6
Ordinary least squares (OLS)
The coefficients that minimize $RSS = \|y - Xb\|^2$; maximum likelihood when the errors are iid Normal. One input: $\hat\beta_1 = S_{xy}/S_{xx}$, $\hat\beta_0 = \bar y - \hat\beta_1\bar x$. 5.13
Overdispersion and quasi-Poisson
Count variance larger than the mean (a Poisson forces them equal). Check $\hat\phi = \frac{1}{n-p}\sum\frac{(y - \hat\mu)^2}{\hat\mu}$ (≈ 1 for Poisson); Poisson SEs are too small by about $\sqrt{\hat\phi}$. Fixes: NB regression, quasi-Poisson (SEs × $\sqrt{\hat\phi}$), robust SEs, random effects. 5.14
p-value
The probability, computed assuming $H_0$ is true, of a test statistic at least as extreme as the one observed (checkout: 0.31). Not $P(H_0 \mid \text{data})$, not an effect size, not stable across reruns; uniform under $H_0$. 5.6 5.6
Paired t-test
A one-sample t-test on the within-unit differences: $T = \bar d/(s_d/\sqrt n) \sim t_{n-1}$. It helps when the two measurements of a unit are positively correlated; pairing comes from the design. 5.9
Parallel trends
The difference-in-differences assumption: without treatment, the treated group's average would have changed as much as the comparison group's. Untestable; parallel pre-launch trends make it believable. 5.12
Parameter
A fixed, usually unknown number describing the population or the data-generating process: a true conversion rate, a slope, a noise scale. 5.1
Peeking
Running a fixed-sample test at several interim looks and acting on the first significant one. At $\alpha = 5\%$ the real false-positive rate becomes 8.3% (2 looks), 14% (5), 19% (10), 28% (30). 5.11
Permutation test
Shuffle the group labels many times; the share of shuffles with a gap at least as extreme as the real one is the p-value. Needs exchangeable labels under $H_0$. 5.6
Perplexity
$2^{H(P_i)}$: t-SNE's effective number of neighbours per point; it sets each point's bandwidth $\sigma_i$ by binary search. Default 30, try 5–50, must be below $n$. 5.17
Pocock and O'Brien–Fleming boundaries
Sequential boundaries for 5 looks at $\alpha = 0.05$: Pocock uses 2.413 at every look (easier to stop early); O'Brien–Fleming uses $2.040\sqrt{5/k}$ = 4.56, 3.23, 2.63, 2.28, 2.04 (strict early, almost no power lost at the end). Bonferroni over looks (2.576) is valid but conservative. 5.11
Poisson regression and rate ratios
Counts with $\log\mu = x^\top\beta$; $e^{\beta_j}$ is a rate ratio, the factor that multiplies the expected count per unit of $x_j$ ($100(e^{\beta_j} - 1)\%$ change). It assumes variance = mean. 5.14
Pooled proportion
$(k_A + k_B)/(n_A + n_B)$: the one shared rate under $H_0$ (checkout: 0.11), used only to build the test's SE. Confidence intervals use the unpooled SE. 5.6
Population
The full set of units (users, days, orders) you want a conclusion about. A parameter describes it. 5.4
Positive semi-definite (PSD)
$\mathbf a^\top\Sigma\mathbf a \ge 0$ for every $\mathbf a$ (all eigenvalues ≥ 0): no direction has negative variance. Every covariance matrix is symmetric PSD; positive definite = invertible = a Cholesky factor exists. 5.15
Posterior mean and median
Point summaries of a posterior, best under squared and absolute error. For skewed posteriors they differ from the MAP (the mode): 1 of 50 with a flat prior gives 0.038, 0.033 and 0.020. 5.2
Potential outcomes and counterfactuals
$Y_i(1)$ and $Y_i(0)$: unit $i$'s outcome under treatment and under control; we observe $Y_i = T_iY_i(1) + (1 - T_i)Y_i(0)$. The unseen one is the counterfactual. The notation assumes SUTVA. 5.12
Power
$1 - \beta = P(\text{reject } H_0 \mid \text{true effect } \Delta)$. It grows with $\Delta$, $\alpha$ and $\sqrt n$ and falls with the noise. Plan for 80% at the smallest effect you care about; always say "power to detect what". 5.7 5.7
Practical significance
The effect is at least the smallest effect worth acting on, $\delta_{min}$, agreed before the test. Read the whole interval against 0 and $\delta_{min}$: ship, real but small, inconclusive, no meaningful effect, or harmful. 5.6 5.10
Primary metric (OEC)
The one pre-declared metric whose result decides the launch. Secondary metrics explain the mechanism; guardrails protect against harm. 5.10
Principal component analysis (PCA)
Finds the directions of maximum variance: the eigenvectors of $\hat\Sigma$, sorted by eigenvalue (their variances), computed in practice by the SVD of the centred data. Linear, unsupervised, with a reusable transform. 5.16 5.16 5.16
Principal component regression (PCR)
Regress $y$ on the first $k$ PC scores, then map back: $\hat\beta = V_k\hat\gamma$. Stable under collinearity, but biased if $y$ depends on a dropped direction. $k = d$ gives OLS. 5.16
Probability sample
Every unit has a known, non-zero chance of selection (simple random, stratified, cluster sampling). A convenience sample has unknown chances, so its bias can be neither measured nor corrected. 5.4
Propensity score and overlap
$e(x) = P(T = 1 \mid X = x)$: the probability of treatment given covariates, used to weight, match or stratify observational data (0.5 for everyone in a randomized test). Needs no unmeasured confounding and overlap, $0 \lt e(x) \lt 1$, or the weights explode. 5.12
Random vector and marginal distribution
Several random numbers drawn together (one user, one day, one posterior draw), described by a joint distribution; data are the rows of an $n \times d$ matrix. A marginal is one coordinate on its own; the marginals do not determine the joint. 5.15
Randomization
Assignment by a chance mechanism independent of the units: $T \perp (Y(0), Y(1))$. The arms are then comparable on average in every trait, seen or unseen, and the difference in means is unbiased for the ATE. 5.12 5.10
Randomization unit vs analysis unit
What the coin was flipped for vs what each analysis row represents. Analysing finer than you randomized (sessions of a user-randomized test) makes SEs and posteriors too narrow. Fixes: per-user metrics, the delta method for ratios, cluster-robust SEs, a user-level bootstrap, or a per-user random effect. 5.10
Reconstruction error
$\sum_i\|\mathbf x_i - \hat{\mathbf x}_i\|^2 = (n-1)\sum_{j \gt k}\lambda_j$: what you lose by keeping $k$ components. No other $k$-dimensional flat fit does better (Eckart–Young). 5.16
Regression adjustment (ANCOVA)
Add a pre-treatment covariate to the regression of the outcome on the treatment. In a randomized test it keeps the target (the ATE) and shrinks the variance by about $1 - \rho^2$. 5.12
Regularization
Loss + λ × penalty: trades a little bias for a big drop in variance. Standardize the features, leave the intercept unpenalized, choose λ by cross-validation or through a prior. 5.3
Relative lift
$(\theta_B - \theta_A)/\theta_A$: the effect as a percentage of the baseline (10% → 12% is +20%). It needs its own interval (delta method, bootstrap or posterior draws), not the absolute interval divided by $\hat p_A$. 5.7 5.8
Reparameterization and invariance
Changing the ruler, $\eta = g(\theta)$: the MLE follows ($g(\hat\theta)$), the MAP does not, because a density picks up a Jacobian factor. Quantiles carry over; modes and means do not. 5.2
Residual
$e_i = y_i - \hat y_i$, measured from the fitted line; an observable stand-in for the unobservable error $\epsilon_i$. With an intercept, OLS residuals sum to 0 and are uncorrelated with every column. 5.13
Residual diagnostics
Residual vs fitted (curve = wrong mean, funnel = changing variance), Q-Q (tails), residual vs time and its ACF (correlated errors): how you check a regression's assumptions. 5.13
Ridge (L2)
Least squares + $\lambda\sum_j\beta_j^2$: $\hat\beta = (X^\top X + \lambda I)^{-1}X^\top y$. Shrinks smoothly, never exactly to 0, and stabilizes collinear features; equals the MAP under a Gaussian prior with $\lambda = \sigma^2/\tau^2$. 5.3 5.3
Robust (sandwich, HC) standard errors
SEs that stay valid when the error variance is unequal, e.g. statsmodels fit(cov_type='HC3'). They do not fix correlated errors (use clustered or HAC SEs for those). 5.13 5.13
R² and adjusted R²
$R^2 = 1 - RSS/TSS$: the share of the variance of $y$ explained by the fit ($r^2$ for one input). It never falls when you add a column, which is why adjusted $R^2 = 1 - \frac{RSS/(n-p)}{TSS/(n-1)}$ exists. Neither is accuracy or causality. 5.13
Sample ratio mismatch (SRM)
The observed split of units between arms differs from the plan by more than chance: a $\chi^2$ test on the counts (teams often alarm at $p \lt 0.001$). Results are invalid until the leak is found; the bias does not shrink with more data. 5.11
Sample size (per group)
$n = 2\sigma^2(z_{1-\alpha/2} + z_{1-\beta})^2/\Delta^2 \approx 16\sigma^2/\Delta^2$ at $\alpha = 0.05$ and 80% power (Lehr's rule). Conversions 10% → 12%: 3 841 per group; 10% → 11%: 14 751 per arm. 5.7 5.10
Sampling distribution
The distribution of an estimator over repeated samples from the same process with the same $n$. Its centre gives the bias, its width the SE, and its shape (often a bell) the intervals. 5.1 5.5 5.4
Sampling frame
The list or mechanism the sample is actually drawn from. Parts of the population missing from it can never appear (coverage error); a sample speaks only for its frame. 5.4
Scores and loadings (PCA)
Score $z_{ij} = \mathbf v_j^\top(\mathbf x_i - \bar{\mathbf x})$: the coordinate of observation $i$ along component $j$ (scores are uncorrelated, $Var(z_j) = \lambda_j$). Loadings are the entries of $\mathbf v_j$: how much each original variable contributes. 5.16 5.16
Selection bias and sampling bias
Units enter the analysis through a process related to the outcome (or to group and outcome): biased sampling methods, self-selection, non-response, survivorship, attrition, filters on post-treatment behaviour. More data does not fix it. 5.4 5.12
Sequential testing (group sequential, α-spending, always-valid)
Testing while data arrive, paying for the looks: planned looks with stricter boundaries (group sequential), a spending function that fixes how much $\alpha$ each fraction of the data may use (Lan–DeMets α-spending), or always-valid p-values and confidence sequences (e.g. mSPRT) that allow stopping any time at some cost in power. 5.11
Shrinkage
Pulling an estimate toward a fixed value (usually 0 or a group mean) compared with the unpenalized estimate. It adds bias and removes variance. 5.3 5.1
Significance level (α)
$P(\text{reject } H_0 \mid H_0 \text{ true})$, fixed before the data (often 0.05): the Type I error rate. It is not "the chance that this winner is fake". 5.6 5.7
Simple random sample (SRS)
Every set of $n$ units from the frame is equally likely; each unit has inclusion probability $n/N$, and $\hat p$ is unbiased for the frame's rate. 5.4
Soft-thresholding
$S_\lambda(z) = \text{sign}(z)\max(|z| - \lambda, 0)$: set to exactly 0 inside the dead zone $|z| \le \lambda$, shrink by λ outside it. The lasso and the Laplace-prior MAP in the orthonormal case. 5.3 5.3
Softplus link
$\mu = \log(1 + e^\eta)$: keeps a mean positive while acting almost additively at high levels. Not a classic GLM link, but common for count likelihoods in custom Bayesian models. 5.14 5.14
Sparsity
Many coefficients exactly 0. Lasso fits and Laplace-prior MAPs are sparse; posteriors under continuous priors (Laplace, horseshoe) are not. 5.3 5.3
Spike-and-slab prior
A prior with a point mass at exactly 0 plus a wide "slab". Exact zeros in the posterior need a prior like this; a Laplace prior only shrinks. 5.3
Standard deviation (SD)
The spread of single values, $\sigma$ (estimated by $s$ with ddof=1). It does not shrink when you collect more data. 5.5
Standard error (SE)
The standard deviation of an estimator's sampling distribution: how much the estimate would wobble over repeated samples. For a mean of independent values, $\sigma/\sqrt n$. Always say "the SE of what". 5.5 5.5
Standard error of a difference
$\sqrt{SE_A^2 + SE_B^2}$ for independent estimates (variances add, Pythagoras): 0.0198 for the checkout. Never add SEs; paired data need the per-unit differences. 5.5
Standard error of a proportion
$\sqrt{p(1-p)/n}$ with $\hat p$ plugged in; $n$ counts trials, not conversions. Largest at $p = 0.5$; rare events have a small absolute but a large relative SE. 5.5
Statistic
Any number computed from the sample alone (mean, count, rate). It contains no unknown parameters, and it is not the school subject. 5.1 5.6
Statistical significance
$p \le \alpha$: the data are unlikely under "no effect", so the effect is probably not zero. It depends on the effect, the noise and $n$ together; it is not importance. 5.6
Stratified sampling and Neyman allocation
Sample inside every stratum and weight by the population shares, $\hat\mu_{st} = \sum_h W_h\bar y_h$: the between-strata variance disappears, so with proportional allocation it is never worse than simple random sampling. Neyman allocation $n_h \propto W_h\sigma_h$ gives the smallest variance. 5.4
Student's t-test (pooled)
The two-sample t-test that pools the two variances, assuming them equal. When the small group is the noisy one its false-alarm rate can be several times $\alpha$. It is SciPy's ttest_ind default. 5.9
Survivorship bias
Analysing only units that passed a survival filter related to the outcome (current customers, stores still open). Follow a cohort from the start, leavers included. 5.4
SUTVA
Stable Unit Treatment Value Assumption: (1) no interference, a unit's outcome depends only on its own assignment; (2) one well-defined version of the treatment. It is what lets you write $Y_i(1)$ and $Y_i(0)$. 5.11 5.12
SVD of the centred data
$X_c = USV^\top$: the right singular vectors are the principal directions, $\lambda_i = s_i^2/(n-1)$, the scores are $US$. Libraries use it because it never forms $X_c^\top X_c$. 5.15 5.16
t distribution (as a reference curve)
The heavier-tailed curve for a statistic whose SE uses $s$ instead of $\sigma$, with $n - 1$ degrees of freedom for one sample. 95% multipliers: 4.30 ($n = 3$), 2.26 ($n = 10$), 2.05 ($n = 30$), → 1.96. 5.5 5.8 5.9
t-SNE
A non-linear neighbour map for pictures: Gaussian neighbour probabilities in the data, Student-t (1 df) similarities in the map, minimize $KL(P\|Q)$. Keeps neighbourhoods; cluster sizes and gaps are not to scale; no transform for new points. 5.17 5.17 5.17
Test statistic
A statistic that measures the distance from $H_0$, usually (estimate − null value) / SE, and has a known null distribution. Checkout: $z = 1.01$. It grows with $n$ for the same gap, so it is not an effect size. 5.6
Triggering and dilution
Analysing only users who reached the point where A and B differ (reached checkout), logged the same way in both arms; never trigger on post-treatment behaviour. If untriggered users are unaffected, the all-users effect is $q \times$ the triggered effect ($q$ = trigger rate), yet they still add noise. 5.10
Two-proportion z-test
$z = (\hat p_B - \hat p_A)/\sqrt{\hat p(1-\hat p)(1/n_A + 1/n_B)}$ with the pooled $\hat p$, then $p = 2\Phi(-|z|)$: the classical conversion test. Checkout: $z = 1.01$, $p = 0.31$. 5.6 5.9
Type I and Type II errors
Type I: rejecting a true $H_0$ (false positive), rate $\alpha$, chosen. Type II: keeping a false $H_0$ (a miss), rate $\beta$, which depends on the true effect, $n$ and the noise. Power $= 1 - \beta$. 5.7
Type M and Type S errors
With low power, significant estimates exaggerate the true effect (Type M, magnitude; about ×2.5 in the checkout design) and occasionally have the wrong sign (Type S). One form of the winner's curse. 5.7
UMAP
A neighbour embedding built on a fuzzy k-nearest-neighbour graph (n_neighbors = 15) with a cross-entropy that includes explicit repulsion, fitted by SGD with negative sampling from a spectral start. Scales to big data and has an approximate transform; sizes and gaps are still not to scale. 5.17 5.17
Variance inflation factor (VIF)
$1/(1 - R_j^2)$: how much collinearity inflates the variance of coefficient $j$ ($R_j^2$ from regressing $x_j$ on the other columns). The SE grows by $\sqrt{VIF}$; 5–10 is a rule-of-thumb warning. 5.13
Wald interval (proportion)
$\hat p \pm z\sqrt{\hat p(1-\hat p)/n}$. It fails for small $n$ or rates near 0 or 1 (0 of 20 gives [0, 0]). statsmodels' proportion_confint default. 5.8
Welch's t-test
The two-sample t-test with separate variances and the Welch–Satterthwaite degrees of freedom. The safe default: ttest_ind(..., equal_var=False). 5.9
Whitening
Rotate to the principal axes and divide by $\sqrt{\lambda_j}$ (PCA whitening), or apply $\hat\Sigma^{-1/2}$ (ZCA): the covariance becomes the identity and Euclidean distance becomes Mahalanobis distance. Thin directions get their noise amplified. 5.16
Wilcoxon signed-rank and Kruskal–Wallis tests
Rank tests: Wilcoxon signed-rank for paired data or one sample ($H_0$: differences symmetric around 0), Kruskal–Wallis for $k$ independent groups (the rank version of ANOVA: all groups from the same distribution). 5.9
Wilson interval (and Clopper–Pearson)
Wilson inverts the score test: centre pulled toward 0.5, stays inside [0, 1], good coverage for small $n$ (1 of 20 → [0.9%, 23.6%]); in statsmodels pass method='wilson'. Clopper–Pearson ("exact", from Beta quantiles) never under-covers but is conservative. 5.8
Winner's curse
Results selected because they cleared a bar (significance, the best segment, a lucky high at which you stopped) overstate the truth on average. 5.7 5.11
Appendix B

Formula and test sheet

The key formulas of every chapter on one page, each with what it means in words and the trap that usually goes with it. Then the "which test when" table with the SciPy and statsmodels calls, a one-page A/B planning sheet, and the library defaults that silently change your answer. Notation: $\theta$ a parameter, $\hat\theta$ its estimator, $\bar x$ and $s$ the sample mean and sd (ddof = 1), $\Phi$ the standard Normal CDF, $z_q = \Phi^{-1}(q)$, $N(\mu, \sigma^2)$ written with the variance. The checkout example is 50/500 (A) vs 60/500 (B).

5.1–5.3 · Estimators, likelihood, priors and regularization

5.1 · Estimators: bias, variance, MSE, consistency, efficiency

FormulaIn wordsMain trapSee
$Bias(\hat\theta) = E[\hat\theta] - \theta$Where the pile of estimates is centred, compared with the truth.Unbiased ≠ accurate, biased ≠ bad; bias does not shrink with $n$.5.1
$Var(\hat\theta) = E[(\hat\theta - E\hat\theta)^2]$, $\;SE = \sqrt{Var(\hat\theta)}$How much the recipe wobbles around its own centre. For iid data $Var(\bar X) = \sigma^2/n$.Low variance can be steadily wrong; $\sigma^2/n$ needs independence.5.1
$MSE = E[(\hat\theta - \theta)^2] = Bias^2 + Var$One score for the total error (the cross term has mean 0).The unbiased $s^2$ loses in MSE to dividing by $n + 1$ for Normal data.5.1
$E\big[\tfrac1n\sum(X_i - \bar X)^2\big] = \tfrac{n-1}{n}\sigma^2$The divide-by-$n$ variance is biased low by $\sigma^2/n$; $s^2$ (÷ $n - 1$) is unbiased.$s$ itself is still biased low for $\sigma$.5.1
$MSE(c\bar X) = (1-c)^2\theta^2 + c^2\sigma^2/n$, $\;c^* = \dfrac{\theta^2}{\theta^2 + \sigma^2/n}$Noisy data → shrink more; plenty of data → barely shrink.$c^*$ contains the unknown $\theta$: real methods set the shrinkage with a prior, pooling or validation.5.1
$P(|\hat\theta_n - \theta| \gt \varepsilon) \to 0$; enough: $MSE \to 0$Consistency: more data closes in on the truth.$X_1$ is unbiased but not consistent; the ÷$n$ variance is biased but consistent.5.1
$\text{eff} = Var(\hat\theta_2)/Var(\hat\theta_1)$; median vs mean $= 4f(m)^2\sigma^2$A data ratio: $2/\pi \approx 0.64$ (Normal), $\approx 1.62$ ($t_3$), 2 (Laplace).Which recipe is efficient depends on the tails.5.1
$Var(\hat\theta) \ge 1/(nI(\theta))$Cramér–Rao: the best precision any unbiased recipe can reach.Needs regularity conditions; a biased recipe can still have lower MSE.5.1

5.2 · Maximum likelihood and MAP

FormulaIn wordsMain trapSee
$L(\theta) = \prod_i p(x_i\mid\theta)$The likelihood: data fixed, $\theta$ varies; a score for how well $\theta$ explains the data.Not a probability distribution over $\theta$ (coin example: area 1/11); only ratios matter.5.2
$\ell(\theta) = \sum_i \log p(x_i\mid\theta)$Same maximizer; sums instead of products.Products of many probabilities underflow; compare $\ell$ only on the same data.5.2
$\hat\theta_{MLE} = \arg\max_\theta \ell(\theta)$; coin $k/n$, Poisson $\bar y$Solve $\ell'(\theta) = 0$; check $\ell'' \lt 0$ and the edges ($k = 0$).The MLE maximizes the probability of the data, not of $\theta$.5.2 5.2
Normal: $\hat\mu = \bar x$, $\;\hat\sigma^2 = \frac1n\sum(x_i - \bar x)^2$Maximizing over $\mu$ is least squares.The $\sigma^2$ MLE divides by $n$ (biased low); norm.fit and np.std do too.5.2
$SE \approx 1/\sqrt{-\ell''(\hat\theta)} \approx 1/\sqrt{nI(\theta)}$Precision = curvature at the peak; coin: $\sqrt{\hat\theta(1-\hat\theta)/n}$.Poor at small $n$ and near the edges of the parameter range.5.2
$-\log p$: Normal → $(y-\mu)^2$, Laplace → $|y - \mu|$, Bernoulli → log loss, Poisson → $\lambda - y\log\lambda$Every familiar loss is a negative log-likelihood: choosing a loss = choosing a noise model.Student-t does not remove outliers; it gives them less influence.5.2
$\hat\theta_{MAP} = \arg\max_\theta[\log p(D\mid\theta) + \log p(\theta)]$The posterior's peak = MLE plus a penalty $-\log p(\theta)$.The MAP is one point of the posterior, not the posterior.5.2
Beta$(\alpha, \beta)$ prior, $k$ of $n$: MAP $\frac{k+\alpha-1}{n+\alpha+\beta-2}$, mean $\frac{k+\alpha}{n+\alpha+\beta}$Pseudo-counts; flat prior → MAP = MLE; large $n$ → MLE.Skewed posteriors spread mean, median and mode apart (1 of 50, flat: 0.038, 0.033, 0.020).5.2 5.2
$p_\eta(\eta) = p_\theta(\theta)\,|d\theta/d\eta|$Change the ruler: the MLE follows, the MAP moves (Jacobian)."Flat prior" depends on the ruler; quantiles carry over, modes and means do not.5.2

5.3 · Regularization as a prior

FormulaIn wordsMain trapSee
$\hat\beta_\lambda = \arg\min\,[\text{Loss} + \lambda\cdot\text{Penalty}]$λ = 0: unpenalized (overfits); λ → ∞: everything → 0 (underfits).Standardize features; do not penalize the intercept; choose λ by cross-validation or a prior.5.3
Ridge: $(X^\top X + \lambda I)^{-1}X^\top y$; orthonormal: $z/(1+\lambda)$Shrinks the coefficients smoothly toward 0; always one unique answer.Never exactly 0, so it does not select features.5.3
Lasso: $\tfrac12\|y - X\beta\|^2 + \lambda\|\beta\|_1$; orthonormal: $S_\lambda(z) = \text{sign}(z)\max(|z| - \lambda, 0)$Dead zone $|z| \le \lambda$ → exactly 0; outside it, shrink by λ.Conventions differ (½ or not; scikit-learn alpha $= \lambda/n$); 0 ≠ "proved useless".5.3
All lasso coefficients are 0 once $\lambda \ge \max_j|x_j^\top y|$Paths: ridge smooth and never 0; lasso piecewise linear, zeros one by one.Drop order ≠ importance when features are correlated.5.3
Prior $\beta_j \sim N(0, \tau^2)$ → ridge with $\lambda = \sigma^2/\tau^2$The posterior is Normal, so MAP = mean = median = ridge.Ridge returns only the centre; the posterior also gives the spread.5.3
Prior $\beta_j \sim \text{Laplace}(0, b)$ → lasso at the MAP with $\lambda = \sigma^2/b$ (½RSS form)$2\sigma^2/b$ without the ½; scikit-learn $\alpha = \sigma^2/(nb)$; in general $\lambda \propto \sigma^2/b$.The equivalence is for the MAP only.5.3
$p(\theta\mid D) \propto e^{-(\theta - \bar y)^2/2s^2 - |\theta|/b}$; MAP $= S_{s^2/b}(\bar y)$MAP exactly 0 when $|\bar y| \le s^2/b$, yet $P(\theta = 0\mid D) = 0$ ($\bar y = 0.8$, $s = b = 1$: MAP 0, mean 0.39, median 0.33, $P(\theta \gt 0\mid D) = 0.70$)."Laplace prior ⇒ sparse posterior" is false; exact zeros need a spike-and-slab prior.5.3
$\delta_j \sim \text{Laplace}(0, b)$, $\;E|\delta_j| = b$Sparse-ish changepoint slopes: most tiny, a few large.Sparse only at the MAP; $b$ is a bias–variance knob.5.3

5.4–5.5 · Sampling, standard errors and the bootstrap

5.4 · Sampling and its biases

FormulaIn wordsMain trapSee
SRS: inclusion probability $n/N$Every set of $n$ units equally likely; $\hat p$ is unbiased for the frame.Big ≠ representative; a sample speaks only for its frame.5.4
Sampling spread of $\hat p$: $\sqrt{p(1-p)/n}$Sampling variation is luck; 4× the data halves the spread.It describes estimates, not data.5.4
$MSE = \text{bias}^2 + \text{variance}$Selection bias is systematic.More data shrinks only the variance; never compare arms on a post-treatment subset.5.4
Respondent mean $\approx \bar y + Cov(\rho, y)/\bar\rho$Non-response bias needs response to be related to the outcome (60% satisfied → respondents 81.8%).A low response rate alone is not bias.5.4
$\hat\mu_{st} = \sum_h W_h\bar y_h$, $Var = \sum_h W_h^2\sigma_h^2/n_h$; Neyman $n_h \propto W_h\sigma_h$Sample inside every stratum: the between-strata variance disappears.Weight by population shares; gains only if strata means differ.5.4
$\rho = \tau^2/(\tau^2 + \sigma^2)$, $\;DEFF = 1 + (m-1)\rho$, $\;n_{eff} = n/DEFF$Cluster sampling: cheaper, but each extra unit tells you less.Naive SE too small by $\sqrt{DEFF}$; count independent units, not rows.5.4
IID: same $F$, $\;p(x_1, \dots, x_n) = \prod_i p(x_i)$Gives $\sigma^2/n$ and product likelihoods.IID ≠ Normal; fails for time series, repeated users, clusters, drift, interference.5.4

5.5 · Standard error and sampling distributions

FormulaIn wordsMain trapSee
$SE(\bar X) = \sigma/\sqrt n$ (from $Var(\sum X_i) = n\sigma^2$)Noise partly cancels when you average.Needs independence and a finite variance, not Normal data; correlated rows make the real SE bigger.5.5
$SE(kn) = SE(n)/\sqrt k$; $\;n = (\sigma/SE_{target})^2$The √n law: 4× the data → half the SE.More data shrinks noise, never bias.5.5
$SE(\hat p) = \sqrt{p(1-p)/n}$Checkout: 0.0134 (A), 0.0145 (B).$n$ = trials, not conversions; the plug-in SE fails at $\hat p = 0$ or 1.5.5
$SE(B - A) = \sqrt{SE_A^2 + SE_B^2}$Variances add for independent groups: 0.0198 for the checkout, so the 2-point gap is about 1 SE.Never add SEs; correlated estimates need $-2Cov$.5.5
$\widehat{SE}(\bar x) = s/\sqrt n$, $\;\frac{\bar X - \mu}{s/\sqrt n} \sim t_{n-1}$Heights 160, 170, 180: $s = 10$, $\widehat{SE} = 5.77$.The estimated SE wobbles and is usually a bit low for small $n$: use t.5.5
$P(|\hat\theta - \theta| \le 1.96\,SE) \approx 0.95$For bell-shaped, unbiased estimators.Medians, maxima, small samples and rare events can be lopsided.5.5
Bootstrap SE = sd of $\hat\theta^*_1, \dots, \hat\theta^*_B$Resample $n$ rows with replacement, $B$ times; heights: 4.71.Resample the independent unit (users, blocks); fails for tiny $n$ and for the maximum.5.5

5.6–5.9 · Tests, power, intervals and classical tests

5.6 · Hypothesis testing: the logic

FormulaIn wordsMain trapSee
$T = (\hat\theta - \theta_0)/SE$The test statistic: distance from $H_0$ in SE units (checkout $z = 1.01$)."Statistic" = a number from the sample; $T$ grows with $n$, it is not an effect size.5.6
$\hat p_{pool} = \frac{k_A + k_B}{n_A + n_B}$, $\;SE_0 = \sqrt{\hat p_{pool}(1 - \hat p_{pool})\big(\frac1{n_A} + \frac1{n_B}\big)}$One shared rate under $H_0$ (0.11); $SE_0 = 0.0198$.Tests use the pooled SE, intervals the unpooled one.5.6
$p = P_{H_0}(|T| \ge |t_{obs}|) = 2\Phi(-|z|)$Checkout: $2\Phi(-1.01) = 0.31$. Under $H_0$, p is uniform.$p \ne P(H_0\mid\text{data})$; not an effect size; not stable across reruns.5.6
$\alpha = P(\text{reject}\mid H_0)$; $\;z^* = \Phi^{-1}(1 - \alpha/2) = 1.96$, one-sided 1.645Reject ⇔ $p \le \alpha$ ⇔ $|z| \ge z^*$.Fix α and the side before the data; picking the side afterwards doubles false alarms.5.6 5.6
$P(\text{significant}\mid H_0) \ne P(H_0\mid\text{significant})$10% good ideas and 50% power → about 47% of "wins" are false.Turning the conditional around.5.6
Reject vs fail to rejectOnly two verdicts; report the effect and its interval next to $p$.Absence of evidence ≠ evidence of absence; significant ≠ important.5.6 5.6

5.7 · Errors, power, effect size and sample size

FormulaIn wordsMain trapSee
Type I rate $\alpha$, Type II rate $\beta$, power $1 - \beta$α is chosen; β depends on the true effect, $n$ and the noise.α is not "the chance a winner is fake".5.7
Under $H_0$: $P(p \le u) = u$; A/A false alarms ~ Binomial$(m, \alpha)$1 000 A/A tests → about 50 ± 14 significant.A pile-up of small p-values means a broken pipeline.5.7
$\text{Power} = \Phi(\delta - z_{1-\alpha/2}) + \Phi(-\delta - z_{1-\alpha/2})$, $\;\delta = \Delta/SE = \frac{\Delta\sqrt n}{\sigma\sqrt2}$Levers: bigger Δ, bigger $n$ (as $\sqrt n$), smaller σ, bigger α. 80% power at 5% needs $\delta \approx 2.8$.Always say "power to detect what".5.7 5.7
$n = \frac{2\sigma^2(z_{1-\alpha/2} + z_{1-\beta})^2}{\Delta^2} \approx \frac{16\sigma^2}{\Delta^2}$ per groupMultiplier 15.7 (80%), 21.0 (90%), 23.4 (α = 0.01, 80%).Per group vs total; points vs percent.5.7
Rates: $n = \frac{\left[z_{1-\alpha/2}\sqrt{2\bar p\bar q} + z_{1-\beta}\sqrt{p_Aq_A + p_Bq_B}\right]^2}{(p_B - p_A)^2}$$\bar p = (p_A + p_B)/2$, $q = 1 - p$. 10% → 12%: 3 841 per group.Tools differ slightly (3 835–3 841); halving the MDE ≈ 4× the users.5.7
$\theta_B - \theta_A$; $\;(\theta_B - \theta_A)/\theta_A$; $\;d = \Delta/\sigma$Effect sizes in points, percent and standard deviations."2%" alone is ambiguous; a lower baseline needs far more users for the same relative lift.5.7
Observed power $= \Phi(|z| - z_{1-\alpha/2}) + \Phi(-|z| - z_{1-\alpha/2})$A function of the p-value alone (0.17 for $z = 1.01$).Never report it. Low power also inflates significant winners (about ×2.5 in the checkout design).5.7

5.8 · Confidence intervals

FormulaIn wordsMain trapSee
$\hat\theta \pm z_{1-\alpha/2}\,SE$Multipliers 1.645 / 1.96 / 2.576 for 90 / 95 / 99%.The interval moves, the truth does not.5.8
$P_\theta(L \le \theta \le U) = 1 - \alpha$95% = the hit rate of the recipe over repeated samples.Never "95% probability that θ is in [l, u]" for a confidence interval.5.8
$\bar x \pm t_{n-1,\,0.975}\,s/\sqrt n$$t$ multipliers 4.30 ($n$ = 3), 2.26 (10), 2.05 (30); heights: [145.2, 194.8].z with $s$ under-covers for small $n$; t assumes roughly Normal data.5.8
$m = z\cdot SE$; $\;n = z^2p(1-p)/m^2$; $\;n = (z\sigma/m)^2$±1 point on a 10% rate: 3 458 users; ±3 points worst case: 1 068.A margin plan is not a power plan (the 80%-power $n$ is about twice as big).5.8
Wilson: $\dfrac{\hat p + \frac{z^2}{2n} \pm z\sqrt{\frac{\hat p(1-\hat p)}{n} + \frac{z^2}{4n^2}}}{1 + z^2/n}$Stays inside [0, 1]; 1/20 → [0.9%, 23.6%].Wald fails for small $n$ or rates near 0/1 (0 of 20 → [0, 0]); statsmodels defaults to Wald.5.8
$(\hat p_B - \hat p_A) \pm z\sqrt{\frac{\hat p_A\hat q_A}{n_A} + \frac{\hat p_B\hat q_B}{n_B}}$Checkout lift: [−1.9, +5.9] points.Overlapping group intervals ≠ "not significant"; the relative lift needs its own interval.5.8
CI $= \{\theta_0 : p(\theta_0) \ge \alpha\}$A null value outside the 95% CI ⇔ $p \lt 0.05$.Match levels; pooled vs unpooled SEs can disagree right at the edge.5.8
Credible: $P(a \le \theta \le b\mid D) = 0.95$; Beta-Binomial: quantiles of Beta$(a + k, b + n - k)$60/500, flat prior: [9.44%, 15.15%] vs CI [9.44%, 15.14%].Similar numbers with big data and weak priors, different meaning; with small data or strong priors the numbers differ too.5.8

5.9 · Classical tests

FormulaIn wordsMain trapSee
(estimate − value under $H_0$) / SE, or a sum of squared standardized gapsReference curves: z (σ known or $n$ large), t (σ estimated), χ² (counts), F (variance ratio).χ² and F use only the right tail, yet detect differences in any direction.5.9
$T = \frac{\bar x - \mu_0}{s/\sqrt n} \sim t_{n-1}$One-sample t; 5% critical values 2.571 (df 5), 2.262 (df 9), 2.045 (df 29).Using 1.96 with tiny $n$ gives too many false alarms (≈ 12% at $n = 5$).5.9
Welch: $T = \frac{\bar x_1 - \bar x_2}{\sqrt{s_1^2/n_1 + s_2^2/n_2}}$, Welch–Satterthwaite dfThe safe default for two independent groups.Student's pooled test: ≈ 30% false alarms when a group of 8 has 4× the sd of a group of 32; SciPy defaults to it.5.9
Paired: $T = \bar d/(s_d/\sqrt n)$; $\;Var(\bar d) = (\sigma_x^2 + \sigma_y^2 - 2Cov)/n$Positive within-unit correlation removes noise.Pairing comes from the design, never from matching rows afterwards.5.9
$\chi^2 = \sum_k \frac{(O_k - E_k)^2}{E_k}$, $E_k = N\pi_k$, df $= K - 1 - m$Goodness-of-fit: counts vs fixed shares (the SRM check).Counts, not percentages; every $E_k \ge 5$ (rule of thumb).5.9
$E_{ij} = R_iC_j/N$, df $= (r-1)(c-1)$; 2×2: $\chi^2 = z^2$Independence and homogeneity: same arithmetic, different design.SciPy applies Yates' correction to 2×2 tables by default; small counts → Fisher's exact test.5.9
$F = \frac{SSB/(k-1)}{SSW/(N-k)}$, $SST = SSB + SSW$; $k = 2$: $F = t^2$One-way ANOVA: between vs within variance.A significant F doesn't say which group; pairwise t-tests inflate false alarms (≈ 29% at $k = 5$).5.9
Mann–Whitney $U$ = number of pairs with $a \gt b$Ranks: robust to outliers and skew; efficiency ≈ 0.955 vs t for Normal data.Not a test of means (or of medians unless the shapes match).5.9

5.10–5.12 · Designing experiments, their pitfalls, and causal thinking

5.10 · Designing an A/B experiment

FormulaIn wordsMain trapSee
$\hat\tau = \bar Y_B - \bar Y_A$, $\;SE = \sqrt{s_A^2/n_A + s_B^2/n_B}$Unbiased for the effect under random assignment.Say absolute (points) or relative (%); report the CI.5.10
bucket = hash(unit id + salt) mod 100Sticky, reproducible, independent across experiments.Opt-in or calendar assignment creates bias that no $n$ removes.5.10
$DE = 1 + (m-1)\rho$, $\;n_{eff} = n_{rows}/DE$$m = 5$, $\rho = 0.3$ → 2.2; a naive "95%" CI covers ≈ 81%.Compute the SE at the randomization unit.5.10
All-users effect $= q \times$ triggered effectTriggering removes noise from users who could not be affected.Trigger only on treatment-independent conditions, logged in both arms.5.10
$n \approx 15.7\,CV^2/r^2$ per arm; conversion $CV = \sqrt{(1-p)/p}$Metric sensitivity for a relative change $r$.Revenue metrics are far noisier than conversion.5.10
$MDE \approx (z_{1-\alpha/2} + z_{1-\beta})\sqrt{2\sigma^2/n}$; $\;Var \propto \frac1f + \frac1{1-f}$10% → 11%: 14 751 per arm; a 90/10 split needs 2.78× the traffic of 50/50.MDE ≠ expected effect; half the MDE ≈ 4× the users.5.10
days $= \lceil k\,n_{arm}/(\text{users/day}\times\text{share})\rceil$ → whole weeks29 502 users at 3 000/day → 9.8 days → 14 days.Don't stop mid-week or extend "until significant".5.10
Whole CI vs 0 and $\delta_{min}$Ship · real but small · inconclusive · no meaningful effect · harmful."Significant" ≠ "worth it".5.10
Simulate $P(\text{ship}\mid\delta)$ and $P(\text{ship}\mid 0)$Flat priors, one look: "$P(B \gt A\mid D) \gt 0.95$" ≈ a one-sided 5% z-test (≈ 88% power and ≈ 5% false ships at 14 750 per arm for 10% → 11%)."Bayesian means no sample size" is false.5.10

5.11 · A/B testing pitfalls

FormulaIn wordsMain trapSee
SRM: $E_g = Nq_g$, $\;\chi^2 = \sum_g (O_g - E_g)^2/E_g$, df = groups − 15 200 vs 4 800 → $\chi^2 = 16$, $p \approx 6\times10^{-5}$: results invalid.Reweighting does not fix it; find the leak and rerun.5.11
Measured $= \tau(e_T - e_C)$; users needed × $1/(e_T - e_C)^2$Contamination and non-exposure dilute the effect.Intention-to-treat is the main answer.5.11
$P(\max_{k \le K}|z_k| \gt 1.96\mid H_0)$Peeking: 5%, 8.3%, 14.2%, 19.3%, 28% for $K$ = 1, 2, 5, 10, 30 looks.Extending "almost significant" tests is peeking too.5.11
$K = 5$: Bonferroni 2.576; Pocock 2.413; O'Brien–Fleming $2.040\sqrt{5/k}$Plan the looks, pay for the looks (or use α-spending / always-valid p-values).Bonferroni over looks is conservative (actual 3.3%).5.11
"Ship at first $P(B \gt A\mid D) \gt 0.95$": A/A ship rate 5% (1 look) → about 21% (20 daily looks)The posterior is valid any time; the rule's error rates grow with looks.It over-promises when the prior does not match reality.5.11
FWER $= 1 - (1-\alpha)^m \le m\alpha$20 independent tests → 64%; 10 → 40%.Fix the families before looking.5.11
Bonferroni $p \le \alpha/m$; Holm $p_{(k)} \le \alpha/(m-k+1)$; BH: largest $k$ with $p_{(k)} \le k\alpha/m$FWER control (Bonferroni, Holm) vs FDR control (BH).BH does not control FWER; $m$ = all tests run, not the ones that looked good.5.11

5.12 · Causal thinking

FormulaIn wordsMain trapSee
$Y_i = T_iY_i(1) + (1 - T_i)Y_i(0)$, $\;\tau_i = Y_i(1) - Y_i(0)$Potential outcomes; $\tau_i$ is never observed.The notation assumes SUTVA.5.12
ATE $= E[Y(1)] - E[Y(0)]$; CATE $= E[Y(1) - Y(0)\mid X = x]$; ATT: given $T = 1$Averages can be estimated; individuals cannot.A positive ATE does not mean nobody is harmed.5.12
Naive difference = ATT + $E[Y(0)\mid T{=}1] - E[Y(0)\mid T{=}0]$Selection bias; randomization ($T \perp (Y(0), Y(1))$) makes it 0.Do not filter after assignment.5.12
ATE $= \sum_x P(x)\,\big(E[Y\mid T{=}1, x] - E[Y\mid T{=}0, x]\big)$Adjusting for confounders closes backdoor paths.Needs no unmeasured confounding; never adjust for mediators or colliders.5.12
$Y^{cv} = Y - \theta(X - \bar X)$, $\;\theta = Cov(Y,X)/Var(X)$, $\;Var \times (1 - \rho^2)$CUPED: same target, less noise ($\rho = 0.7$ → 51% of the users).Pre-experiment $X$ only; one pooled θ; not a bias correction.5.12
IPW weights $1/e(x)$ (treated), $1/(1 - e(x))$ (control)Re-weighting imitates a randomized comparison.All confounders measured + overlap.5.12
DiD $= (\bar Y_{T,post} - \bar Y_{T,pre}) - (\bar Y_{C,post} - \bar Y_{C,pre})$$\hat\beta_3$ in $y \sim \text{treated}\times\text{post}$; effect on the treated.Parallel trends; cluster SEs by unit.5.12
IV: $\hat\tau = \frac{\bar Y_{Z=1} - \bar Y_{Z=0}}{\bar T_{Z=1} - \bar T_{Z=0}}$Effect for compliers (LATE).Weak instruments explode the noise; exclusion is untestable.5.12

5.13–5.14 · Linear regression and generalized linear models

5.13 · Linear regression

FormulaIn wordsMain trapSee
$y = X\beta + \epsilon$, $\;E[\epsilon\mid X] = 0$"Linear" means linear in β; columns may be $x^2$, sines, 0/1 flags.The assumptions are about the errors, not the histogram of $y$.5.13
$\hat\beta_1 = S_{xy}/S_{xx}$, $\;\hat\beta_0 = \bar y - \hat\beta_1\bar x$Least squares with one input.Residual ≠ error; vertical distances; needs spread in $x$.5.13
$X^\top X\hat\beta = X^\top y$ ⇔ $X^\top e = 0$; $\;\hat y = Hy$Normal equations: residuals perpendicular to every column (a projection).Solve with lstsq/QR, not an explicit inverse.5.13
$Var(\hat\beta) = \sigma^2(X^\top X)^{-1}$, $\;\hat\sigma^2 = RSS/(n-p)$, $\;t = \hat\beta_j/SE_j \sim t_{n-p}$Coefficient SEs, tests and intervals; one input: $SE(\hat\beta_1) = \hat\sigma/\sqrt{S_{xx}}$.Needs independent, equal-variance errors.5.13
$R^2 = 1 - RSS/TSS$Share of the variance of $y$ explained.Not accuracy, not causality; never falls when you add a column.5.13
Correct mean → unbiased; independence + equal variance → right SEs (BLUE); Normal errors → exact small-$n$ testsWhich assumption buys what.Wrong mean = wrong answer; wrong noise = wrong error bars.5.13
HC SE: $(X^\top X)^{-1}\big(\sum e_i^2x_ix_i^\top\big)(X^\top X)^{-1}$; WLS $w_i = 1/\sigma_i^2$Fixes for heteroscedasticity (residual funnel).HC SEs do not fix correlated errors.5.13
$VIF_j = 1/(1 - R_j^2)$Collinearity inflates $SE(\hat\beta_j)$ by $\sqrt{VIF_j}$.Coefficients wobble, predictions don't; VIF over 5–10 is a rule of thumb.5.13
$[1,\ t,\ (t - s_j)_+,\ \sin,\ \cos,\ \text{holidays},\ X_t]$; trend $= kt + m + \sum_j\delta_j(t - s_j)_+$Your forecasting model's design matrix.Changepoint locations are fixed; only the weights are estimated.5.13

5.14 · Generalized linear models

FormulaIn wordsMain trapSee
$g(E[y\mid x]) = x^\top\beta$Likelihood (from support and variance) + linear predictor + link.The link transforms the mean, not the data: $\log E[y] \ne E[\log y]$.5.14
Logistic: $\log\frac{p}{1-p} = x^\top\beta$, $\;p = \frac{1}{1 + e^{-x^\top\beta}}$$e^\beta$ = odds ratio; checkout OR 1.23, $p \approx 0.31$.Coefficients are log-odds; perfect separation → infinite MLE.5.14
$OR = RR\cdot\frac{1 - p_A}{1 - p_B}$OR ≈ RR only for rare events.Never present an odds ratio as a relative lift.5.14
Poisson: $\log\mu = x^\top\beta$; offset: $\log\mu_i = \log t_i + x_i^\top\beta$$e^\beta$ = rate ratio; compare rates, not raw counts.Assumes variance = mean.5.14 5.14
$\hat\phi = \frac{1}{n-p}\sum\frac{(y - \hat\mu)^2}{\hat\mu}$≈ 1 for a Poisson; clearly above 1 = overdispersion.Poisson SEs too small by about $\sqrt{\hat\phi}$.5.14
NB2: $Var = \mu + \mu^2/\alpha$NumPyro concentration α; statsmodels alpha $= 1/\alpha$; SciPy $n = \alpha$, $p = \alpha/(\alpha + \mu)$.Parameterizations differ: know the exact one your code calls.5.14
IRLS: $\beta \leftarrow (X^\top WX)^{-1}X^\top Wz$; $\;D = 2[\ell_{sat} - \ell]$MLE by repeated weighted least squares; $\Delta D \sim \chi^2_{\Delta p}$; AIC for non-nested models.Non-convergence means separation or collinearity, not "too few iterations".5.14

5.15–5.17 · Covariance matrices, PCA, t-SNE and UMAP

5.15 · Multivariate statistics

FormulaIn wordsMain trapSee
$\Sigma = E[(\mathbf X - \boldsymbol\mu)(\mathbf X - \boldsymbol\mu)^\top]$, $\;\hat\Sigma = X_c^\top X_c/(n-1)$Every variance and covariance; $d(d+1)/2$ free numbers.np.cov(X, rowvar=False) when rows are observations; covariances depend on units.5.15
$Cov(A\mathbf X + \mathbf b) = A\Sigma A^\top$; $\;Var(\mathbf a^\top\mathbf X) = \mathbf a^\top\Sigma\mathbf a$The variance of any combination of metrics.Dropping the covariance terms gives the wrong SE for combined metrics.5.15
Valid ⇔ symmetric and all eigenvalues ≥ 0PD ⇔ invertible ⇔ a Cholesky factor $\Sigma = LL^\top$ exists.Pairwise-valid correlations can still be jointly impossible.5.15
$R = D^{-1/2}\Sigma D^{-1/2}$Unit-free correlation matrix.Units change $\Sigma$ and its eigenvectors, never $R$.5.15
$p(\mathbf x) \propto |\Sigma|^{-1/2}e^{-\frac12(\mathbf x - \boldsymbol\mu)^\top\Sigma^{-1}(\mathbf x - \boldsymbol\mu)}$; 2D $k$-sd ellipse holds $1 - e^{-k^2/2}$39%, 86%, 99% for $k$ = 1, 2, 3.SciPy/NumPyro multivariate Normals take the covariance (or its Cholesky), not sds.5.15
$\Sigma = V\Lambda V^\top$Axes of the ellipse and the variances along them; half-axes $k\sqrt{\lambda_i}$.Eigenvalues are variances; eigh returns them in ascending order.5.15
$X_c = USV^\top$, $\;\lambda_i = s_i^2/(n-1)$, scores $US$Same axes, numerically safer.Centre first; NumPy returns $V^\top$ (rows); signs are arbitrary.5.15
$d_M^2 = (\mathbf x - \boldsymbol\mu)^\top\Sigma^{-1}(\mathbf x - \boldsymbol\mu) \sim \chi^2_d$Distance in the cloud's own sds (2D 99% cut-off: 9.21).Thresholds grow with $d$; outliers inflate $\hat\Sigma$ (masking).5.15
Guide sizes: $2d$ · $d(r+2)$ · $d + d(d+1)/2$Mean-field · low-rank · full-rank (NumPyro).All three are Gaussian approximations; none is exact.5.15

5.16 · Principal component analysis

FormulaIn wordsMain trapSee
PC1 $= \arg\max_{\|\mathbf u\| = 1}\mathbf u^\top\hat\Sigma\mathbf u$ ⇒ $\hat\Sigma\mathbf v = \lambda\mathbf v$Widest shadow = smallest perpendicular loss; $Var(z_j) = \lambda_j$.PC1 is not the regression line; signs are arbitrary.5.16 5.16
$\lambda_j/\sum_i\lambda_i$; cumulative $\sum_{j\le k}\lambda_j/\sum_i\lambda_i$Explained share; pick $k$ by elbow, target or (best) validation.Variance of $X$ ≠ relevance to $y$.5.16
$\mathbf z = V_k^\top(\mathbf x - \bar{\mathbf x})$, $\;\hat{\mathbf x} = \bar{\mathbf x} + V_k\mathbf z$; error $(n-1)\sum_{j \gt k}\lambda_j$Compress and rebuild; Eckart–Young optimal.Squared error lets outliers tilt the components.5.16
$X^\top X/(n-1) = \hat\Sigma + \frac{n}{n-1}\bar{\mathbf x}\bar{\mathbf x}^\top$Why you must centre.np.linalg.svd and TruncatedSVD do not centre.5.16
Correlation PCA = eigen-decomposition of $R$Mixed units → make_pipeline(StandardScaler(), PCA()).A large-unit column hijacks PC1.5.16
PCR: $\hat\beta = V_k(Z_k^\top Z_k)^{-1}Z_k^\top\mathbf y$A stable fit for collinear regressors; $k = d$ gives OLS.Biased if $y$ depends on a dropped direction.5.16
$\Lambda^{-1/2}V^\top(\mathbf x - \bar{\mathbf x})$ (PCA), $\;\hat\Sigma^{-1/2}(\mathbf x - \bar{\mathbf x})$ (ZCA)Whitening: identity covariance; Euclidean = Mahalanobis.$1/\sqrt\lambda$ amplifies noise in thin directions.5.16

5.17 · t-SNE and UMAP

FormulaIn wordsMain trapSee
$p_{j\mid i} \propto e^{-\|x_i - x_j\|^2/2\sigma_i^2}$, $\;p_{ij} = (p_{j\mid i} + p_{i\mid j})/2n$The neighbour table the map must copy.Feature scaling decides the distances; standardize first.5.17
$\text{Perp}(P_i) = 2^{H(P_i)}$Effective number of neighbours; default 30, try 5–50.Must be below $n$; it equalizes dense and sparse regions.5.17
$q_{ij} \propto (1 + \|y_i - y_j\|^2)^{-1}$Student-t with 1 df in the map: heavy tails fight crowding.Gaps are stretched non-linearly: not to scale.5.17
$KL(P\|Q) = \sum p_{ij}\ln\frac{p_{ij}}{q_{ij}}$; gradient $4\sum_j(p_{ij} - q_{ij})\frac{y_i - y_j}{1 + \|y_i - y_j\|^2}$Springs: pull if $p \gt q$, push if $q \gt p$.Weighted by $p$: local structure kept, global barely constrained.5.17
Reading rulesTrust neighbours and groups that are stable across seeds and perplexities.Cluster sizes, densities, far gaps and axes are not meaningful.5.17
UMAP: $w_{j\mid i} = e^{-(d_{ij} - \rho_i)/\sigma_i}$, union $a + b - ab$, map curve $1/(1 + a\,d^{2b})$Fuzzy kNN graph (n_neighbors = 15, min_dist = 0.1); has a transform.Sizes and gaps are still not to scale.5.17

Which test when

Choose the row from the outcome type and the design, before you see the results. Every row assumes independent units at the level of the randomization unit (or correctly paired units). Calls use from scipy import stats, import statsmodels.api as sm, import statsmodels.formula.api as smf; the statsmodels helpers live in statsmodels.stats.proportion, .oneway and .multitest.

Data typeDesignTestSciPy / statsmodels callAssumptions and notes
Yes/no (converted)Two independent groups (A/B)Two-proportion z-test (= 2×2 χ² without correction)proportions_ztest([k_B, k_A], [n_B, n_A])Independent users; at least about 10 successes and 10 failures per group; $n$ fixed in advance, one look. Checkout: $z = 1.01$, $p = 0.31$. 5.6 5.9
Yes/no, small countsTwo groupsFisher's exact teststats.fisher_exact([[k_A, n_A - k_A], [k_B, n_B - k_B]])Independent units; use when expected counts are below about 5. 5.9
Yes/noOne group vs a target rate $p_0$One-proportion z or exact binomial testproportions_ztest(k, n, value=p0, prop_var=p0); stats.binomtest(k, n, p0)Independent trials; $np_0$ and $n(1 - p_0)$ at least about 10 for z. Without prop_var=p0, statsmodels puts $\hat p$ in the SE instead of $p_0$. 5.9
Categories (counts)$K$ counts vs planned shares, e.g. the SRM checkχ² goodness-of-fitstats.chisquare(obs, f_exp=exp) (add ddof=m for $m$ fitted parameters)Counts, not percentages; every expected count ≥ 5 (rule of thumb). 5.9 5.11
Categories (counts)Groups × categories (categorical A/B metric), or two variablesχ² homogeneity / independencestats.chi2_contingency(table, correction=False)Expected counts ≥ 5, else Fisher or a simulated p; SciPy's default Yates correction (2×2) breaks $\chi^2 = z^2$. 5.9
A numberOne group vs a targetOne-sample tstats.ttest_1samp(x, mu0)Roughly Normal data, or $n$ large enough for the CLT. 5.9
A numberTwo independent groupsWelch's tstats.ttest_ind(x_B, x_A, equal_var=False)Sample means roughly Normal; the default equal_var=True is Student's pooled test. 5.9
A numberThe same units measured twicePaired tstats.ttest_rel(after, before)Independent pairs; differences roughly Normal; pairing comes from the design. 5.9
A number3+ independent groupsOne-way ANOVA (Welch ANOVA if spreads differ)stats.f_oneway(*groups); anova_oneway(groups, use_var="unequal")Roughly Normal within groups; the classic F assumes equal variances. Follow up with corrected comparisons. 5.9
A number, skewed or ordinalTwo independent groupsMann–Whitney Ustats.mannwhitneyu(x_B, x_A)Tests $P(X \gt Y) = \tfrac12$, not means; pick it only if that is the business question. 5.9
A number, skewedPaired, or one sample vs a targetWilcoxon signed-rankstats.wilcoxon(after, before)Independent pairs; differences symmetric around 0 under $H_0$. 5.9
A number, skewed3+ independent groupsKruskal–Wallisstats.kruskal(*groups)All groups from the same distribution under $H_0$. 5.9
Counts (orders per user)Groups or covariates, unequal exposurePoisson or NB regression (rate ratios)sm.GLM(y, X, family=sm.families.Poisson(), exposure=t); sm.NegativeBinomial(y, X)Check the dispersion $\hat\phi$; Poisson SEs are too small when counts are overdispersed. 5.14 5.14
A number with a pre-period valueRandomized two groupsRegression adjustment / CUPEDsmf.ols("y ~ treat + x_pre", data=df).fit(cov_type="HC3")The covariate must be measured before assignment; variance × about $1 - \rho^2$. 5.12 5.12
Any statistic (median, ratio metric)ResamplingBootstrap interval / permutation teststats.bootstrap((x,), np.median); stats.permutation_test((x_B, x_A), stat)Resample the independent unit (users with all their sessions, blocks of days). 5.5 5.6
Many p-valuesA family of tests (metrics, segments, variants)Holm (FWER) or Benjamini–Hochberg (FDR)multipletests(pvals, alpha=0.05, method="holm") or method="fdr_bh"Fix the family before looking; $m$ = every test you ran. 5.11

The matching intervals and power calls

WhatCallNote
Interval for a proportionproportion_confint(k, n, method="wilson")The default method="normal" is Wald. 60/500 → [9.44%, 15.14%]. 5.8
Interval for a difference of proportionsconfint_proportions_2indep(k_B, n_B, k_A, n_A)Newcombe's method by default (checkout: [−1.9, +5.9] points); Wald by hand gives almost the same here. 5.8
Interval for a meanstats.t.interval(0.95, df=n-1, loc=x.mean(), scale=stats.sem(x))stats.sem uses ddof = 1. 5.8
Sample size for two ratesNormalIndPower().solve_power(effect_size=proportion_effectsize(0.12, 0.10), alpha=0.05, power=0.8)Uses Cohen's $h$: about 3 835 per group (the pooled formula gives 3 841). 5.7
Sample size for two meansTTestIndPower().solve_power(effect_size=d, alpha=0.05, power=0.8)$d = \Delta/\sigma$ (Cohen's $d$), per group. 5.7

A/B planning cheat sheet

The whole life of one experiment on one page, in the order you do things. All numbers are the chapters' worked examples (two-sided $\alpha = 0.05$, 80% power unless stated). Handy quantiles: $z_{0.975} = 1.96$, $z_{0.95} = 1.645$, $z_{0.995} = 2.576$, $z_{0.8} = 0.84$, $z_{0.9} = 1.28$.

StepFormula or ruleWorked exampleSee
1. Write the planHypothesis · randomization and analysis unit · trigger · one primary metric, guardrails · $\delta_{min}$ and MDE · $n$, split, duration · analysis method and decision rule · stopping rule · health checks.Everything fixed before launch; MDE ≤ $\delta_{min}$.5.10 5.10
2. Sample size (rates)$n = \frac{\left[z_{1-\alpha/2}\sqrt{2\bar p\bar q} + z_{1-\beta}\sqrt{p_Aq_A + p_Bq_B}\right]^2}{(p_B - p_A)^2}$ per arm10% → 12%: 3 841; 10% → 11%: 14 751.5.7 5.10
3. Quick checkLehr's rule $n \approx 16\sigma^2/\Delta^2$, with $\sigma^2 \approx p(1-p)$ for rates; or $15.7\,CV^2/r^2$ for a relative change $r$10% → 11%: $16 \times 0.09/0.01^2 = 14\,400$ (exact 14 751).5.7 5.10
4. MDE for the traffic you have$MDE \approx (z_{1-\alpha/2} + z_{1-\beta})\sqrt{2\sigma^2/n} = 2.80\sqrt{2p(1-p)/n}$5 000 per arm at a 10% baseline: about 1.7 points. Half the MDE ≈ 4× the users.5.10
5. Power of a given design$\Phi(\delta - z_{1-\alpha/2}) + \Phi(-\delta - z_{1-\alpha/2})$, $\;\delta = \Delta/SE$Checkout design (500 per arm, 10% → 12%): about 17%.5.7
6. Split$Var \propto 1/f + 1/(1-f)$: 50/50 is bestA 90/10 split needs 2.78× the total traffic.5.10
7. Durationdays $= \lceil k\,n_{arm}/(\text{eligible users/day}\times\text{share})\rceil$, rounded up to whole weeks29 502 users at 3 000/day → 9.8 → 14 days.5.10
8. Units and variance$DE = 1 + (m-1)\rho$; CUPED variance × $(1 - \rho^2)$; triggered effect × $q$ = overall effect$m = 5$, $\rho = 0.3$ → DE 2.2; CUPED with $\rho = 0.7$ → 51% of the users.5.10 5.12 5.10
9. Health checks before any metricSRM: $\chi^2 = \sum(O_g - E_g)^2/E_g$, df = arms − 1, alarm often at $p \lt 0.001$; A/A history: about 5% significant5 200 vs 4 800 → $\chi^2 = 16$, $p \approx 6\times10^{-5}$: invalid. 5 050 vs 4 950 → $p \approx 0.32$: fine.5.11 5.7
10. LooksOne planned look, or sequential boundaries ($K = 5$: Bonferroni 2.576, Pocock 2.413, O'Brien–Fleming $2.040\sqrt{5/k}$)Peeking at 1.96: 8.3% (2 looks), 14.2% (5), 19.3% (10), 28% (30).5.11 5.11
11. Many metrics or segmentsFWER $= 1 - (1-\alpha)^m$; Holm for decision metrics, BH for scorecards; hierarchical pooling for segments20 tests → 64%. Ten example p-values: Bonferroni rejects 2, Holm 3, BH 4.5.11 5.11
12. Bayesian decision ruleSimulate $P(\text{ship}\mid\text{effect} = \delta)$ and $P(\text{ship}\mid 0)$ under the real priors and look schedule; decide on $P(\theta_B - \theta_A \gt \delta\mid D)$Flat priors, one look, 14 750 per arm, 10% → 11%: ≈ 88% ship, ≈ 5% false ships. Daily looks for 20 days: A/A ship rate about 21%.5.10 5.11
13. Read the resultEffect with its interval (or posterior) against 0 and $\delta_{min}$CI +0.06 to +0.44 points with a +0.5 threshold: real but not worth shipping on its own.5.10

Library defaults that silently change your answer

None of these raise an error. Each one gives a number that looks fine and answers a different question from the one you meant.

WhatConventionSee
Two-sample t-teststats.ttest_ind defaults to equal_var=True (Student's pooled test). Pass equal_var=False for Welch.5.9
2×2 chi-squarestats.chi2_contingency applies Yates' continuity correction to 2×2 tables by default; correction=False makes $\chi^2 = z^2$.5.9
Proportion intervalsproportion_confint defaults to method="normal" (Wald); pass method="wilson".5.8
One-proportion z-testproportions_ztest(k, n, value=p0) uses $\hat p$ in the SE; prop_var=p0 gives the textbook $\sqrt{p_0(1-p_0)/n}$.5.9
Variance divisornp.var, np.std and norm.fit divide by $n$; pandas .var(), .std() and stats.sem use $n - 1$.5.2 5.5
Ridge and lassoRidge(alpha=λ) minimizes $\|y - X\beta\|^2 + \lambda\|\beta\|^2$; Lasso(alpha) minimizes $\frac{1}{2n}\|y - X\beta\|^2 + \alpha\|\beta\|_1$, so $\alpha = \lambda/n$ in the ½RSS convention.5.3 5.3
Negative BinomialNumPyro NegativeBinomial2(mean, concentration=α); statsmodels alpha $= 1/\alpha$; SciPy nbinom(n=α, p=α/(α+μ)); NumPyro NegativeBinomialProbs(total_count=α, probs=μ/(α+μ)). All have $Var = \mu + \mu^2/\alpha$.5.14
Exposure in count GLMsstatsmodels offset=np.log(t) or exposure=t; in NumPyro, mu = t * jnp.exp(X @ beta).5.14
Robust SEsstatsmodels fit(cov_type="HC3"); with a single treatment dummy, "HC2" reproduces Welch's SE exactly. Neither fixes correlated errors.5.13
Covariance and eigenvectorsnp.cov treats rows as variables unless rowvar=False; np.linalg.eigh returns eigenvalues in ascending order; np.linalg.svd returns $V^\top$ (components in rows).5.15 5.15
Multivariate NormalSciPy multivariate_normal(mean, cov) and NumPyro MultivariateNormal take the covariance (or scale_tril), never standard deviations.5.15
PCAscikit-learn's PCA centres but does not scale; components_ holds one component per row; explained_variance_ uses $n - 1$; TruncatedSVD does not centre.5.16 5.16
t-SNEscikit-learn's TSNE: perplexity 30, early exaggeration 12, init="pca", 1 000 iterations, Barnes–Hut; it has fit_transform but no transform.5.17 5.17
Appendix C

Interview question bank

Forty-four questions an interviewer could ask about this guide's material while you walk them through your Bayesian A/B framework or your forecasting model. Most come straight from the chapters' "Say it right" boxes: p-values, confidence vs credible intervals, peeking, MAP vs the posterior under Laplace priors, power, CUPED, PCA vs t-SNE.

How to use this bank. Read the question, answer it out loud in about a minute, and only then open the model answer. A strong answer usually has four parts: (1) a one-sentence definition in plain words, (2) a small number or picture, (3) where it lives in your project, (4) the trap you avoid. If an answer feels shaky, follow its link back to the chapter.

Estimators, likelihood, MAP and priors (5.1–5.3)

1. In your A/B framework, what is the parameter, what is an estimator, and what is the estimate?

The parameter is variant A's true conversion rate $\theta_A$: a fixed, unknown number. An estimator is a recipe applied to the random data, for example the observed rate $K_A/n_A$ (the maximum likelihood estimator) or the Beta-Binomial posterior mean $(\alpha + K_A)/(\alpha + \beta + n_A)$: two different recipes for the same parameter. The estimate is the single number a recipe gave on our data, say 0.106. Because the data are random, an estimator has a sampling distribution with a bias, a variance and an MSE. So I judge the recipe, and I report the estimate together with its uncertainty. Chapter 5.1

2. Should an estimator always be unbiased?

No. What matters is the total error, and $MSE = Bias^2 + Var$. Shrinking a noisy estimate toward a sensible value adds a little bias but can remove much more variance: for $c\bar X$ the best $c^* = \theta^2/(\theta^2 + \sigma^2/n)$ is below 1, and even the unbiased $s^2$ loses in MSE to dividing by $n + 1$ for Normal data. Both my projects use this on purpose: hierarchical partial pooling pulls small A/B segments toward the overall mean, and the Laplace prior pulls changepoint slope changes toward 0. Too much shrinkage hurts too, so the amount should come from a prior, from pooling learned from the data, or from validation. Chapter 5.1

3. What is the difference between "unbiased" and "consistent"? Give an example of each without the other.

Unbiased is about the centre of the sampling distribution at a fixed $n$: $E[\hat\theta] = \theta$. Consistent is about the whole distribution collapsing onto the truth as $n$ grows: $P(|\hat\theta_n - \theta| \gt \varepsilon) \to 0$. The first observation $X_1$ is unbiased for $\mu$ but not consistent, because its spread never shrinks. The divide-by-$n$ variance is biased (by $-\sigma^2/n$) but consistent, because its bias and its variance both go to 0. A handy sufficient condition for consistency is MSE → 0. Chapter 5.1

4. Your forecasting model offers a Student-t likelihood. Why, and does it remove outliers?

Which estimator is efficient depends on the tails. Under Normal data the mean beats the median (relative efficiency $2/\pi \approx 0.64$), but under heavy tails the median wins (about 1.62 for $t_3$). A Normal likelihood behaves like squared error, so one huge residual has a huge pull on the trend and seasonality. The Student-t negative log-likelihood grows only like the log of the residual, so its gradient for a huge residual is small. It does not remove outliers: it assigns more probability to extreme residuals, so they exert less influence on the fit. Chapter 5.1 · Chapter 5.2

5. What is a likelihood? Is it the probability of the parameter?

No. The likelihood is the probability (or density) of the observed data given the parameter, $p(D\mid\theta)$, read as a function of $\theta$ with the data fixed. For a fixed $\theta$ it sums to 1 over possible datasets; as a function of $\theta$ it need not integrate to 1 (in the coin example its area is 1/11), and only likelihood ratios are meaningful. To get a distribution over $\theta$ I multiply by a prior and normalize: that is the posterior. Chapter 5.2

6. "The MLE gives the most likely parameter." Is that right? How is MAP different?

Not quite. The MLE is the parameter value under which the observed data are most likely: it maximizes $p(D\mid\theta)$ and uses no prior. The most probable $\theta$ needs $p(\theta\mid D) \propto p(D\mid\theta)p(\theta)$, and its peak is the MAP: $\arg\max[\log p(D\mid\theta) + \log p(\theta)]$, the MLE plus a penalty $-\log p(\theta)$. With a Beta$(\alpha, \beta)$ prior and $k$ of $n$ conversions the MAP is $(k + \alpha - 1)/(n + \alpha + \beta - 2)$: pseudo-counts. With a flat prior the two coincide, and with lots of data the prior's pull fades. Chapter 5.2 · MAP

7. Your dashboard shows one number for a variant's conversion rate. Is it the MAP, the posterior mean or the median, and does it matter?

From posterior draws it is naturally the mean or the median, not the MAP. They are different summaries of the same posterior, each best under a different loss: the mean under squared error, the median under absolute error, the mode (MAP) when only "exactly right" counts. For a symmetric posterior they agree. For rare conversions the posterior is skewed and they spread apart: 1 conversion in 50 with a flat prior gives MAP 0.020, median 0.033, mean 0.038. So I label which one the dashboard shows. Chapter 5.2

8. Is the MAP invariant to reparameterization? Why does that matter in NumPyro?

The MLE is invariant, the MAP is not. The likelihood is a function of the parameter, so relabelling the parameter just relabels the function and the maximizer maps across. The posterior is a density, and under $\eta = g(\theta)$ it is multiplied by $|d\theta/d\eta|$, which moves the mode: a Beta(9, 5) posterior has its mode at 0.667 on the rate scale, but the mode on the log-odds scale corresponds to 0.643. NumPyro runs SVI and NUTS on an unconstrained scale ($\log\sigma$, log-odds), so a Gaussian guide's centre, mapped back, is the median on the original scale, not the mean or the mode. Quantiles transform cleanly; modes and means do not. Chapter 5.2

9. "A Laplace prior is L1 regularization, so your changepoints are sparse." Do you agree?

Only half. With a Gaussian likelihood and $\delta_j \sim \text{Laplace}(0, b)$, minus the log posterior is the squared error over $2\sigma^2$ plus $\sum|\delta_j|/b$, so the posterior mode is a lasso solution with $\lambda = \sigma^2/b$ (in the ½RSS convention). That mode can be exactly 0 because of the kink. But the posterior is a continuous density: $P(\delta_j = 0\mid D) = 0$, and its mean and median are non-zero whenever the data push away from 0. One coefficient with $\bar y = 0.8$ and $s = b = 1$ has MAP 0 but posterior mean 0.39, median 0.33 and $P(\theta \gt 0\mid D) = 0.70$. My SVI fit approximates the posterior, so the slope changes are "sparse-ish": most tiny, none exactly 0. True posterior sparsity would need a spike-and-slab prior. A Prophet-style MAP fit, by contrast, can return many exact zeros. Chapter 5.3 · changepoints

10. What does ridge regression correspond to in Bayesian terms, and what does the scale b of your Laplace prior control?

Ridge is the MAP under independent $N(0, \tau^2)$ priors with Normal noise, with $\lambda = \sigma^2/\tau^2$: noisy data or a confident prior means strong shrinkage. Here the posterior is Normal, so MAP = mean = median, and the Bayesian version also gives a covariance, which ridge does not. In a Prophet-style forecasting model like mine, Normal priors on seasonal, holiday and regressor weights (if that is what the code uses) act like ridge, and the Laplace prior on the slope changes acts like a lasso at the MAP. Its scale $b$ (with $E|\delta_j| = b$) is a bias–variance knob: small $b$ gives a stiff trend that can miss real changes, large $b$ a trend that chases noise. I would choose it by time-ordered forecast validation and a prior-sensitivity check (Chapters 7.10, 7.15). Chapter 5.3 · Chapter 5.1

Sampling, IID and standard errors (5.4–5.5)

11. You have a table of page views from a user-randomized test. Can you treat the rows as IID?

No. IID means every row comes from the same distribution and no row carries information about another. Page views of the same user are positively correlated, and the unit that was randomized is the user. With $m$ rows per user and intra-class correlation $\rho$ the variance is inflated by $1 + (m - 1)\rho$: with $m = 5$ and $\rho = 0.3$ that is 2.2, and a naive "95%" interval covers only about 81%. I aggregate to one row per user, use the delta method or cluster-robust SEs for ratio metrics, or, in the Bayesian model, add a per-user random effect. Feeding sessions into an iid likelihood makes the posterior too narrow for exactly the same reason. Chapter 5.4 · Chapter 5.10

12. Before fitting your Beta-Binomial model, someone filters to "users who reached checkout". What can go wrong?

That is selection on a post-treatment variable. If the treatment changes who reaches checkout, the filtered arms are no longer comparable, and the comparison can even flip sign (in the chapter's cart example a 5% → 6% improvement turned into 25% → 20%). More data does not help, and a narrow posterior makes the wrong answer look certain. I either analyse all assigned users (intention-to-treat) or trigger on a condition the treatment cannot change, logged the same way in both arms. Chapter 5.4 · Chapter 5.10

13. What is the difference between a standard deviation and a standard error? Which numbers in your models are which?

The SD describes how single values vary; the SE is the standard deviation of an estimator's sampling distribution, how much the estimate would vary over repeated samples. For a mean of $n$ independent values $SE = \sigma/\sqrt n$, so the SE shrinks with more data while the SD does not: an SD of 10 over 25 values gives an SE of 2. In my forecasting model, the likelihood's $\sigma$ is an SD (how much one day scatters around the fit), while the posterior sd of a holiday coefficient plays the role of an SE. In the A/B framework, a flat-prior posterior for 50/500 has sd 0.01347, almost exactly the classical $SE = 0.01342$. Chapter 5.5

14. How would you get an honest standard error for "revenue per session"?

It is a ratio of two sums with many sessions per user, so rows are not independent. I would bootstrap users: resample users with replacement, each with all their sessions, recompute the ratio, and take the sd of the resampled values; or use the delta method at the user level. Resampling single sessions would ignore the within-user correlation and give an SE that is too small. The bootstrap also needs a representative sample and a smooth statistic (it fails for the maximum). The result is a good sanity check on the posterior width my Bayesian model reports. Chapter 5.5 · Chapter 5.10

Tests, p-values, power and intervals (5.6–5.9)

15. Define a p-value precisely, using the checkout example.

The p-value is the probability, computed assuming the null hypothesis is true, of a test statistic at least as extreme as the one observed. Checkout: 50/500 vs 60/500, pooled rate 0.11, SE under $H_0$ 0.0198, $z = 0.02/0.0198 = 1.01$, two-sided $p = 2\Phi(-1.01) = 0.31$. It means: if both checkouts converted at the same rate, about 31% of experiments this size would show a gap of 2 points or more in either direction. So the data are quite compatible with no difference. It is not the probability that there is no difference. Chapter 5.6 · the whole z-test

16. "p = 0.03 means a 97% chance that B is better." Correct this, and say what your framework reports instead.

The p-value conditions on the null: if there were no real difference, data at least this extreme would occur about 3% of the time. That is evidence against "no difference", not a probability about B. $P(\text{data}\mid H_0)$ and $P(H_0\mid\text{data})$ can be very different: with 10% good ideas and 50% power, about 47% of significant "wins" are false. The probability that B is better needs a prior; my Bayesian framework reports it directly as $P(\theta_B \gt \theta_A\mid D)$. For the checkout with flat priors that is about 0.84, close to $1 -$ the one-sided p-value of 0.16 because the sample is large and the prior flat, but it answers a different question. Chapter 5.6

17. What is a "test statistic"?

"Statistic" here means a number computed from the sample, not the school subject. A test statistic is a statistic chosen to measure how far the data are from what $H_0$ predicts, usually (estimate − value under $H_0$) / SE, and whose distribution under $H_0$ is known. For the checkout the lift is about one standard error from zero ($z = 1.01$), well within what sampling noise produces. Its probability meaning only comes after comparing it with the null distribution, which gives the p-value. It grows with $n$ for the same gap, so it is not an effect size. Chapter 5.6

18. The checkout test was not significant. Can you conclude that the new checkout does nothing?

No: absence of evidence is not evidence of absence. With 500 users per group and a 10% baseline, a gap had to exceed about 3.9 points to be significant, so the power to detect a real 2-point lift was only about 17%. The 95% interval for the lift, about −1.9 to +5.9 points, is compatible with anything from a small loss to a big win: the test was inconclusive. To detect +2 points with 80% power we would need about 3 800 users per group. I would also not compute "observed power" afterwards; it is just the p-value in disguise. Chapter 5.7 · Chapter 5.6

19. A result is "highly significant". Does that mean the feature has a big impact? How does your framework handle this?

No. Statistical significance only says the effect is probably not zero; it depends on the effect, the noise and $n$ together, so huge samples make tiny effects significant. Practical significance compares the effect, with its uncertainty, to the smallest effect worth acting on, agreed before the test. Example: a 95% interval of +0.06 to +0.44 points with a +0.5-point threshold is real but not worth shipping on its own. In my framework I can put the threshold into the decision: $P(\theta_B - \theta_A \gt \delta\mid D)$ instead of only $P(\theta_B \gt \theta_A\mid D)$. Chapter 5.6 · Chapter 5.10

20. Explain Type I and Type II errors. What is the false-positive rate of your Bayesian decision rule?

A Type I error is a false positive: declaring an effect that is not there, at rate $\alpha$, which we choose. A Type II error is a miss of a real effect, at rate $\beta$, which depends on the true effect, $n$ and the noise; power is $1 - \beta$. A rule like "ship B when $P(\theta_B \gt \theta_A\mid D) \gt 0.95$" makes both kinds of mistakes too, and its threshold is not $\alpha$. I get its false-positive rate by simulation: generate many A/A datasets with $\theta_A = \theta_B$, run the full model on each, and count how often the rule fires. With flat priors and one look it behaves like a one-sided 5% test; with more looks or a mismatched prior it fires more often. A pile-up of false alarms there also catches pipeline bugs. Chapter 5.7 · A/A tests

21. How many users do you need to detect 10% → 12% with 80% power? Walk me through it.

The effect must sit $z_{1-\alpha/2} + z_{1-\beta} = 1.96 + 0.84 \approx 2.8$ standard errors from zero. For two rates (pooled SE under $H_0$): $n = [1.96\sqrt{2\bar p\bar q} + 0.84\sqrt{p_Aq_A + p_Bq_B}]^2/(p_B - p_A)^2$ with $\bar p = 0.11$, which gives about 3 841 per group (tools give 3 835–3 841). A quick check is Lehr's rule $n \approx 16\sigma^2/\Delta^2$. Because the SE shrinks like $1/\sqrt n$, halving the lift you want to detect needs about four times the users: 10% → 11% needs 14 751 per arm. Then I turn $n$ into days from eligible traffic and round up to whole weeks. Chapter 5.7 · Chapter 5.10

22. How do you size a Bayesian A/B test? Doesn't Bayes make sample size unnecessary?

No. Bayesian inference gives direct probability statements, but how often a decision rule makes wrong calls still depends on $n$, the prior and the stopping rule. I size it by simulation (pre-posterior analysis): fix the prior, the decision rule and a candidate $n$; simulate many experiments where the true effect is the MDE and many where it is zero; record $P(\text{ship}\mid\delta)$ (Bayesian power) and $P(\text{ship}\mid 0)$ (false-ship rate); adjust $n$ or the threshold. Example: flat priors, one look, 14 750 per arm, 10% → 11%: the rule "$P(B \gt A\mid D) \gt 0.95$" ships in about 88% of worlds and falsely ships in about 5%, close to a one-sided z-test. Averaging over a prior for the effect gives assurance. Chapter 5.10

23. What does a 95% confidence interval mean?

In frequentist statistics the parameter is fixed and the interval is random. The 95% is a property of the method: if we repeated the experiment many times, about 95% of the intervals built this way would contain the true value. Once I have computed one interval, it either contains the truth or not; I cannot attach a 95% probability to that particular interval without a prior. A confidence interval is also about the parameter, not a range holding 95% of the data. Chapter 5.8

24. What is the difference between a confidence interval and a credible interval? Why does your framework's interval look like the classical one?

A confidence interval treats $\theta$ as fixed and the data as random: 95% is the hit rate of the recipe over repeated samples. A credible interval treats the observed data as fixed and $\theta$ as uncertain: $P(a \le \theta \le b\mid D) = 0.95$ given the model and the prior, so I may say "θ is in this range with probability 0.95", but only as good as the prior. With lots of data and weak priors the posterior is close to Normal with sd ≈ SE (Bernstein–von Mises), so the numbers nearly coincide: 60/500 gives a CI of [9.44%, 15.14%] and a flat-prior credible interval of [9.44%, 15.15%]. With small data or strong priors they differ: 3/20 with a Beta(10, 90) prior gives [5.2%, 36.0%] vs [5.9%, 17.0%]. That is why I would also run a prior-sensitivity check for small segments. Chapter 5.8

25. For a low-conversion segment with few users, which interval for the rate would you use?

Not Wald. $\hat p \pm z\sqrt{\hat p(1-\hat p)/n}$ breaks for small $n$ or rates near 0: with $n = 20$ and a true rate of 5% it covers only about 64% of the time, and 0 of 20 gives the absurd interval [0, 0]. Wilson inverts the score test, stays inside [0, 1] and covers well: 1 of 20 gives [0.9%, 23.6%]. Careful: statsmodels' proportion_confint defaults to Wald, so I pass method='wilson'. In my Bayesian framework the Beta posterior's quantiles are naturally inside [0, 1], and hierarchical pooling borrows strength from other segments. Chapter 5.8

26. What classical test would you compare each of your Bayesian likelihoods with?

The outcome type picks the likelihood the same way it picks the test. Conversions (Beta-Binomial) ↔ the two-proportion z-test, which equals the 2×2 χ² test. A categorical metric (Dirichlet-Multinomial) ↔ the χ² homogeneity test. A continuous metric (Normal or Student-t) ↔ Welch's t-test, or a rank test when outliers dominate. Counts (Poisson) ↔ a rate comparison or Poisson regression. With flat priors and large samples they agree numerically: for the checkout $P(\theta_B \gt \theta_A\mid D) \approx 0.843$ and $\Phi(z) = \Phi(1.01) \approx 0.844$. The Bayesian version gives the whole distribution of the lift instead of one p-value. Chapter 5.9 · z-tests

27. "I used a t-test." Which one, and why? What is the Bayesian analogue?

Welch's two-sample t-test. Student's version pools the two variances; if the smaller group is the noisier one, the pooled SE is too small and the false-positive rate can be several times $\alpha$ (about 30% when a group of 8 has 4× the sd of a group of 32). Welch keeps separate variances and adjusts the degrees of freedom; it costs almost nothing when variances are equal. SciPy's ttest_ind defaults to Student, so I pass equal_var=False. The Bayesian analogue is giving each variant its own scale parameter, $\sigma_A$ and $\sigma_B$, instead of one shared $\sigma$ (in my framework I would check which of the two the Normal or Student-t likelihood uses). Chapter 5.9

28. Revenue per user is very skewed. Would you compare the arms with Mann–Whitney?

Only if the question is "does a typical user spend more?". Mann–Whitney tests whether one group's values tend to be larger, $P(X \gt Y)$, not whether the means differ, and it can disagree with a comparison of means. If the decision is about total revenue, I compare means: Welch's t or a bootstrap interval, which the CLT supports at A/B sample sizes (heavy tails need more data). A Student-t likelihood has the same caveat: with skewed data its location follows the bulk of the data, not the mean. I choose the test from the business question first, then check its assumptions. Chapter 5.9

Designing and reading A/B tests, and causality (5.10–5.12)

29. What is an MDE? Is it "the minimum effect the test will detect"?

The minimum detectable effect is the smallest true effect the design detects with the planned power (say 80%) at the chosen $\alpha$: about $(z_{1-\alpha/2} + z_{1-\beta})\,SE$. Smaller true effects can still come out significant, just less often, and effects exactly at the MDE are missed 20% of the time. It is not the effect we expect. I agree the MDE with the business as (at most) the smallest effect worth shipping, take the baseline and variance from history, and compute $n$; if the $n$ is unaffordable I change the metric, use CUPED or triggering, or don't run the test. Chapter 5.10

30. What is a sample ratio mismatch, and why check it before running your Bayesian model?

An SRM is a split of users between arms that differs from the plan by more than chance allows. I test it with a χ² goodness-of-fit test on the unit counts only, before looking at any metric: planned 50/50 with 5 200 vs 4 800 users gives $\chi^2 = 16$, $p \approx 6\times10^{-5}$. It means users leaked from one arm for a reason, and leaked users are not random, so the arms are no longer comparable; the bias has unknown size and does not shrink with more data. The Beta-Binomial model takes $k$ and $n$ as given and would return a perfectly computed posterior about the wrong groups: no model can detect or repair a broken assignment. So a failed check means "results invalid: find the leak, fix it, rerun". Chapter 5.11

31. "We're Bayesian, so peeking is not a problem for us." How do you respond?

Half true. The posterior is valid whenever I compute it, because the stopping rule does not enter the likelihood. But a decision rule like "ship at the first day $P(B \gt A\mid D) \gt 0.95$" is a sequential procedure with frequentist error rates: in A/A simulations with flat priors it ships B about 5% of the time with one look and about 21% with 20 daily looks. Its stated probability is only guaranteed to be calibrated if the prior matches the real distribution of effects. (For comparison, a classical test checked at 1.96 has 8.3% false positives with 2 looks, 19% with 10, 28% with 30.) So peeking is still a design question: fix the rule and the look schedule in advance, set a minimum duration in whole weeks, use priors informed by past experiments, decide on $P(\theta_B - \theta_A \gt \delta\mid D)$, and simulate the whole procedure. Chapter 5.11 · peeking

32. You track 20 metrics and 10 segments. How do you keep false discoveries under control?

With 20 independent tests at 5% and no real effects, the chance of at least one false positive is $1 - 0.95^{20} \approx 64\%$. I name one primary metric in advance; guardrails are watched for harm; secondary metrics and segments are labelled exploratory. For a few decision metrics or variant comparisons I control the family-wise error rate with Holm (never weaker than Bonferroni); for a scorecard of many exploratory metrics I control the false discovery rate with Benjamini–Hochberg and confirm hits in a new test. For segments, hierarchical partial pooling shrinks noisy segment effects toward the overall effect, so one lucky segment needs real evidence to stand out. Pooling does not protect against scanning many metrics or re-analysing until something crosses 0.95. Chapter 5.11 · corrections

33. What is SUTVA, and where could it fail in your work?

SUTVA (the Stable Unit Treatment Value Assumption) says each unit has exactly one potential outcome per treatment value: (1) no interference, my outcome depends only on my own assignment, and (2) no hidden versions of the treatment. It is not "the groups are similar" (that is what randomization gives) and not "everyone has the same effect". My Beta-Binomial likelihood assumes it: each user's outcome depends only on that user's arm. It fails in marketplaces (shared stock, drivers) and social features; then $P(\theta_B \gt \theta_A\mid D)$ describes "B in a half-treated world", not the launch. The fix is in the design: randomize clusters or time slots (switchbacks), and model at that level. Chapter 5.11

34. Why can you read your A/B posterior causally but not a promotion coefficient in your forecasting model?

Because of how the treatment was assigned. In the A/B test a coin decided, so treatment is independent of the potential outcomes, $T \perp (Y(0), Y(1))$, and $\theta_B - \theta_A$ is the effect of showing B; the causal reading comes from the design, the Bayesian model adds a careful statement of uncertainty. Without randomization, a difference equals the effect on the treated plus selection bias, $E[Y(0)\mid T{=}1] - E[Y(0)\mid T{=}0]$. If promotions were always run in high-demand weeks, the promotion coefficient mixes "promotions raise demand" with "promotions happen when demand is already high": fine for forecasting under the same habits, not an answer to "what if we run an extra one?". Chapter 5.12 · correlation vs causation

35. "Feature adopters retain 30% better, so the feature raises retention by 30%." What do you say?

Adopters are self-selected and probably more engaged to begin with, a common cause of both adopting and retaining. Correlation can come from causation, reverse causation, a common cause (confounding) or selection; only a design that rules out the other three lets me read the gap causally. The true effect could be anywhere from 0 to 30%, or even negative. I would randomize access or an encouragement to use the feature (an instrument), or at least adjust for pre-adoption engagement and say which confounders could remain. Chapter 5.12 · confounding

36. What is CUPED? Is it a bias correction? Could you use it with your Beta-Binomial model?

CUPED is variance reduction, not bias correction: the plain difference in means is already unbiased in a randomized test. It uses a covariate $X$ measured before the experiment (usually the same metric in a pre-period): $Y^{cv} = Y - \theta(X - \bar X)$ with $\theta = Cov(Y, X)/Var(X)$, the value that minimizes the variance. The expected difference is unchanged because $X$ cannot be affected by the treatment, and the variance shrinks by $1 - \rho^2$: with $\rho = 0.7$ you need about 51% of the users. It is regression adjustment with one covariate. A conjugate Beta-Binomial has no slot for it, since adjusted values are no longer 0/1, so for conversions I would use a Bayesian logistic regression with the centred pre-period covariate; for Normal or Student-t metrics I add the covariate to the mean. Chapter 5.12

Regression, GLMs, covariance, PCA and t-SNE (5.13–5.17)

37. Explain your forecasting model as a regression. Why is it "linear" if it contains sines and hinge functions?

The mean is a linear model $\mu = X\beta$ with structured columns: a constant and a ramp $t$ for the trend, one hinge column $(t - s_j)_+$ per candidate changepoint (chosen by a Prophet-like grid plus PELT), $\sin$ and $\cos(2\pi nt/P)$ columns for each seasonality, 0/1 holiday columns, and one column per external regressor. Prophet's continuous piecewise trend, with $\gamma_j = -s_j\delta_j$, is exactly $kt + m + \sum_j\delta_j(t - s_j)_+$. "Linear" means linear in the coefficients: the columns can be any functions of the data, as long as the locations $s_j$, periods and orders are fixed beforehand. The Bayesian version adds priors (Laplace on the $\delta_j$, and whatever priors your code uses on the rest), a likelihood (Normal, Student-t or NB) and fits everything with SVI; the number of columns is what drives the choice between a full-rank and a low-rank guide. Chapter 5.13 · the model

38. A holiday always coincides with a promotion in your data. What happens to their coefficients and to the forecast?

That is multicollinearity: the data cannot split the credit between the two columns. Their coefficients get large standard errors (inflated by $\sqrt{VIF}$, $VIF = 1/(1 - R_j^2)$), can flip sign when a few rows change, and their posteriors are strongly correlated along a long ridge. Predictions stay fine as long as the two keep moving together; the danger is extrapolation, a future promotion without the holiday. Priors act like ridge and keep the coefficients in a sensible part of the ridge, and the posterior correlation reveals the problem. I would check a correlation matrix before adding regressors, and combine, drop or regularize them if I need to interpret them. Chapter 5.13

39. Your demand residuals fan out as the level grows. What does that mean, and what would you do?

Heteroscedasticity: the error variance grows with the level. The mean estimates are still unbiased, but uncertainty is wrong: a constant-σ Normal likelihood overstates uncertainty early in the series and understates it late, so the bands are too wide, then too narrow. Options: model log-demand, use the Negative Binomial likelihood (its variance $\mu + \mu^2/\alpha$ grows with the mean by construction), or let σ depend on the level. In a classical regression I would at least report robust (HC) standard errors or use weighted least squares. Chapter 5.13 · residual plots

40. Why does your Negative Binomial likelihood need a log or softplus link? Is that the same as log-transforming the data? Which parameterization do you use?

A count likelihood needs a positive mean, but trend + seasonality + holidays + regressors can be negative, so the linear predictor goes through a positive link. A link transforms the mean, not the data: $\log E[y] \ne E[\log y]$, zeros are fine, and predictions are means on the original scale. With a log link the components act multiplicatively (a holiday adds a percentage); with softplus they act almost additively at high levels. NB2 has $Var = \mu + \mu^2/\alpha$: in NumPyro NegativeBinomial2(mean, concentration=α), in statsmodels alpha $= 1/\alpha$, in SciPy nbinom(n=α, p=α/(α+μ)). I would check which class my code calls and whether the prior sits on $\alpha$ or $1/\alpha$. Chapter 5.14 · NB regression

41. Why use a Negative Binomial instead of a Poisson for daily demand?

A Poisson forces variance = mean. Real daily counts vary more (days differ in ways the model does not capture), which is overdispersion; check it with the Pearson dispersion $\hat\phi = \frac{1}{n-p}\sum(y - \hat\mu)^2/\hat\mu$, which is about 1 for a Poisson. Under overdispersion the Poisson mean estimates are still reasonable, but its standard errors are too small by about $\sqrt{\hat\phi}$ and its intervals too narrow. The NB adds a Gamma-distributed day-level rate (a Gamma–Poisson mixture); a large fitted $\alpha$ says the series is nearly Poisson. The mean model, a log link on a linear predictor, is unchanged. Chapter 5.14 · NB regression

42. A logistic model gives the treatment an odds ratio of 1.5. Is that a 50% lift in conversion?

No. An odds ratio of 1.5 means the odds $p/(1-p)$ are 1.5 times higher. The relative lift in conversion depends on the baseline: from 50%, odds 1 → 1.5 means 50% → 60%, a 20% lift. Since $OR = RR\cdot\frac{1-p_A}{1-p_B}$, the odds ratio is close to the risk ratio only for rare outcomes. I report the odds ratio with its interval as the model output and translate it into predicted conversion rates for control and treatment, with the absolute and relative difference. Chapter 5.14

43. Why does a full-rank Gaussian guide learn a Cholesky factor, and how do you choose between full-rank and low-rank?

A covariance matrix must be symmetric positive semi-definite, and a free gradient step on its entries can break that. Learning a lower-triangular $L$ with a positive diagonal and using $\Sigma = LL^\top$ guarantees a valid, invertible covariance, and $\log|\Sigma| = 2\sum\log L_{ii}$. The three Gaussian guides are three covariance structures: mean-field diagonal ($2d$ numbers in NumPyro), low-rank $WW^\top + \text{diag}$ ($d(r+2)$) and full-rank $LL^\top$ ($d + d(d+1)/2$). All are Gaussian approximations; none is exact. Full-rank captures every correlation but its memory and work grow like $d^2$, so I use it when $d$ is small and low-rank when $d$ is large, with $r$ set by how many large eigenvalues the posterior correlation has, validated against full-rank or NUTS on a smaller version. The size threshold in my code is a design choice I can defend this way. Chapter 5.15 · guides

44. PCA vs t-SNE: what does each keep? Would you feed t-SNE coordinates into a model, or read cluster sizes from a t-SNE plot?

PCA is a linear map that keeps as much variance as possible: deterministic up to sign, fast, with a transform that applies to new rows. It is unsupervised, so variance is not importance, and I centre with the training mean, standardize mixed units, and fit the scaler and PCA on the training window only. t-SNE chooses 2-D positions so that neighbours stay neighbours: Gaussian neighbour probabilities in the data (perplexity sets each width), a heavy-tailed Student-t in the map, and a KL objective weighted by $p_{ij}$, so far-apart pairs barely matter. Cluster sizes and gaps between far clusters are therefore not meaningful; I read only neighbourhoods and groups that are stable across seeds and perplexities. And no, its coordinates are a poor feature: there is no transform for new points, a re-fit changes the layout, and the geometry is distorted. I use PCA (or the raw features) for models and t-SNE or UMAP only for pictures. Chapter 5.17 · reading a map · PCA limits

Appendix D

Where to go next

This guide taught you to go backwards from data to the process that made it, with an honest statement of uncertainty: estimators, likelihood and priors, tests and intervals, experiments, regression and the geometry of many variables. Here is what it used from Guide 1, where Guides 3 and 4 pick it up, and a short plan for the rest of the series.

How this guide connects to the other three

Each row starts from an idea and points to where it is used next. Links to other guides open them at the right chapter.

Coming from Guide 1 · Probability & Data

From Guide 1Used in this guide for
Random variables, expectation and variance, the $n - 1$ divisor (4.4, 4.5)Estimators as random variables, bias, variance and MSE (5.1); $SE = \sigma/\sqrt n$ (5.5)
Bayes' theorem, $P(A\mid B) \ne P(B\mid A)$ (4.3)MAP (5.2); what a p-value is not (5.6)
Markov's and Chebyshev's inequalities, LLN and CLT (4.12, 4.13)Consistency (5.1); sampling distributions and z-tests (5.5, 5.6)
Bernoulli, Binomial, Normal, Student-t, Laplace (4.7, 4.9)Proportions and their SEs (5.5); t intervals (5.8); Laplace priors (5.3)
Poisson, Negative Binomial, overdispersion (4.8)Poisson and NB regression (5.14)
"At least one" $= 1 - (1-p)^m$ (4.2)Multiple testing and peeking (5.11)
Law of total variance, within vs between (4.6)Cluster sampling and the design effect (5.4); ANOVA (5.9)
Covariance and correlation, $Var(aX + bY)$ (4.15)CUPED (5.12); multicollinearity (5.13); covariance matrices (5.15)
Q-Q plots, standardization and the global scaler (4.17, 4.18)Residual diagnostics (5.13); correlation PCA (5.16); t-SNE inputs (5.17)

Going to Guide 3 · Bayesian Modeling & Computation

From this guideUsed next for
Likelihood, MAP vs the full posterior (5.2, 5.3)Bayesian inference: prior, likelihood, posterior, evidence, predictive (6.1); choosing priors (6.2)
Beta-Binomial MAP and pseudo-counts (5.2)Conjugate Beta-Binomial and Dirichlet-Multinomial models (6.3)
Confidence vs credible intervals; practical thresholds $\delta$ (5.8, 5.10)Credible intervals and posterior decisions such as $P(\theta_B - \theta_A \gt \delta\mid D)$ (6.4)
Bias–variance and shrinkage, strata, ANOVA's between vs within, CATEs (5.1, 5.4, 5.9, 5.12)Hierarchical models (6.5); pooling and shrinkage (6.6)
Simulating a decision rule's error rates; overdispersion checks (5.7, 5.11, 5.14)Posterior predictive checks, prior sensitivity, identifiability (6.8)
Curvature and Fisher information (5.2)The Laplace approximation and MCMC (6.9)
KL divergence (5.17)Variational inference and KL (6.11); the ELBO and SVI (6.12)
Covariance structure, Cholesky factors, the multivariate Normal (5.15)Mean-field, full-rank and low-rank guides and their parameter counts (6.13); SVI vs NUTS (6.15)

Going to Guide 4 · Time Series & Bayesian Forecasting

From this guideUsed next for
IID failures, correlated residuals (5.4, 5.13)Autocorrelation (7.3); residual diagnostics for forecasts (7.17)
The forecasting model as a regression with a design matrix (5.13)The additive model $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$ (7.7)
Hinge columns and Laplace priors on $\delta_j$ (5.3, 5.13)Changepoints (7.8); PELT (7.9); Laplace priors on trend changes (7.10)
Fourier columns; multicollinearity; scaler and PCA fitted on training data only (5.13, 5.16)Fourier seasonality (7.11); holidays, regressors and leakage (7.12)
Losses as likelihoods, links, NB2 and overdispersion (5.2, 5.14)Forecast likelihoods: Normal, Student-t, Negative Binomial (7.13)
Intervals, coverage, a forecast as a counterfactual (5.8, 5.12)Predictive distributions (7.14); coverage and calibration (7.16)
Cross-validation for λ, bias–variance (5.3, 5.1)Time-series cross-validation (7.15); model complexity across both projects (7.18)
Everything aboveThe capstone: one set of ideas, two projects (7.20)

A short study plan

  1. Close this guide properly. Reread the notebook boxes of the P0 chapters: estimators (5.1), MAP vs the posterior (5.2, 5.3), standard errors (5.5), tests, power and intervals (5.6–5.8), A/B testing (5.10, 5.11) and PCA (5.16). Then answer the question bank out loud.
  2. Do the checkout example both ways from memory. Pooled rate, $SE_0$, $z$, p-value, the Wald interval of the lift, the power for +2 points and the $n$ for 80% power; then the Beta posteriors and $P(\theta_B \gt \theta_A\mid D)$. Explain why the numbers agree and why the sentences differ.
  3. Simulate your own decision rule. Take the rule your A/B framework uses, simulate A/A and A/B datasets at your real traffic and look schedule, and write down its false-ship rate and its power (5.7, 5.10, 5.11). Then write a one-page experiment plan for a real test, with the planning sheet next to you.
  4. Guide 3, your A/B framework end to end. Chapters 6.1–6.8 cover the models (conjugacy, posterior decisions, hierarchical pooling, model checking); 6.9–6.15 explain how the posterior is computed (MCMC, NUTS, VI, the ELBO, guides, your training loop, SVI vs NUTS); 6.16–6.17 cover JAX and JIT. Reread 5.15 just before 6.13.
  5. Guide 4, your forecasting model component by component. Chapters 7.1–7.6 give the time-series basics, 7.7–7.13 go through your model piece by piece, 7.14–7.17 cover evaluation and diagnostics, and 7.18–7.20 tie both projects together.
  6. Every week. Run one "Code it" block, change one number and predict the result before you run it. Write one paragraph explaining an idea from your projects to an interviewer, in plain words.

How to make it stick

  • Simulate before you trust a formula. Twenty lines of NumPy check almost every claim in this guide: the bias of the ÷$n$ variance, $\sigma/\sqrt n$, the 5% of A/A tests, the peeking table, the coverage of Wald vs Wilson, the CUPED factor $1 - \rho^2$.
  • Always say "of what". SE of what, power to detect what, lift in points or percent, which posterior summary, which NB parameterization.
  • Check the design before the numbers. SRM, the analysis unit, triggering and the look schedule decide whether any p-value or posterior means what it says.
  • Separate the estimate from the decision. Report the effect with its interval or posterior, then compare it with a threshold that was agreed before the data.
  • Keep your notebook. The "Write this in your notebook" boxes, copied by hand, are the fastest revision sheet you can have.

Well-known resources

  • Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments: the practitioner's reference for A/B testing (SRM, triggering, metrics, pitfalls).
  • Deng, Xu, Kohavi & Walker (2013), "Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data": the CUPED paper.
  • James, Witten, Hastie & Tibshirani, An Introduction to Statistical Learning (free online) and Hastie, Tibshirani & Friedman, The Elements of Statistical Learning (free PDF): regression, ridge and lasso, PCA.
  • Gelman, Hill & Vehtari, Regression and Other Stories: regression, GLMs and causal inference with simulation; and Gelman & Carlin (2014), "Beyond Power Calculations" on Type S and Type M errors.
  • Hernán & Robins, Causal Inference: What If (free online) and Cunningham, Causal Inference: The Mixtape (free online): potential outcomes, DAGs, difference-in-differences, instruments.
  • Wattenberg, Viégas & Johnson, "How to Use t-SNE Effectively" (Distill, free): the reading rules for t-SNE maps, with interactive examples.
  • The SciPy scipy.stats and statsmodels documentation: the final word on defaults (equal_var, correction, method, alpha).

Companion guides

This guide stands on Guide 1 · Probability & Data and on three earlier guides: the Linear Algebra guide (projections, eigenvalues, SVD), the Calculus guide (derivatives and gradients for MLE) and the Optimization guide (gradient descent, regularization, proximal methods). Next stop: Guide 3 · Bayesian Modeling & Computation.