Stats 1 · Probability & Data
Formulas are drawn by KaTeX, which loads from the internet. You seem to be offline, so formulas show as plain text for now.
Statistics for ML · Guide 1 of 4

Probability & Data

Probability, random variables, distributions, averages and spreads, correlation, density plots, Q-Q plots and transformations, from zero. Every idea starts with a plain-English picture, then a worked example with real numbers, then the formal definition, and then a plot you can drag, resample and play with.

What is this guide about, in one sentence? It teaches the language for talking about randomness (probability) and the first tools for looking at data (descriptive statistics and plots).

Why care? Both of your projects are built on these ideas. A conversion in your A/B framework is a random yes/no event. A day of demand in your forecasting model is a random number drawn around a trend. Choosing a Beta, a Student-t or a Negative Binomial, reading a Q-Q plot of residuals, or deciding to log-transform a metric: all of it rests on this guide.

Three ways to say it:

  • Picture: probability runs from a model to the data it could produce; statistics runs back from the data to the model.
  • Numbers: "a 10% conversion rate" is a model; "53 of 500 users converted" is data.
  • Slogan: first learn how randomness behaves, then learn to read it in data.

The four guides

Your syllabus (probability_statistics_syllabus.md) is split into four guides, read in this order. Each guide links to the others whenever an idea is taught elsewhere.

1 · Probability & Data (you are here)

Statistical thinking, probability rules, Bayes' theorem, random variables, expectation and variance, all the key distributions, LLN and CLT, descriptive and robust statistics, correlation, KDE, Q-Q plots, transformations. Syllabus Parts 0–4.

2 · Estimation, Inference & Experiments

Estimators, MLE and MAP, regularization as priors, sampling, standard errors, hypothesis tests, confidence intervals, A/B testing, causal thinking, regression, GLMs, PCA and t-SNE. Parts 5–9, 26, 27.

3 · Bayesian Modeling & Computation

Priors and posteriors, conjugate models, credible intervals, hierarchical models and shrinkage, model checking, MCMC, HMC and NUTS, variational inference, the ELBO, SVI, guides, your training loop, JAX and JIT. Parts 10, 11, 19–22, 24, 25.

4 · Time Series & Bayesian Forecasting

Autocorrelation, stationarity, classical forecasting, the Prophet-style model (changepoints, PELT, Laplace priors, Fourier seasonality, holidays, regressors, likelihoods), forecast evaluation, diagnostics, production, and a capstone that ties both projects together. Parts 12–18, 23, 28–30.

Built around your two projects

Throughout the guides, boxes marked In your projects show exactly where an idea lives in the work you have done.

The Bayesian A/B testing framework

Conversions (Bernoulli, Binomial, Beta), categorical metrics (Categorical, Multinomial, Dirichlet), continuous and count metrics (Normal, Student-t, Poisson), segments and partial pooling, a global scaler, decisions like $P(\theta_B \gt \theta_A \mid D)$. In this guide: chapters 4.3, 4.6, 4.7, 4.11, 4.14 and 4.18.

The Prophet-style forecasting model

Demand as trend + seasonality + holidays + regressors + noise; Normal, Student-t and Negative Binomial likelihoods; Laplace priors on trend changes; residual checks with KDE and Q-Q plots. In this guide: chapters 4.1, 4.8, 4.9, 4.15, 4.16, 4.17 and 4.18.

How every topic is taught

Each concept follows the same steps, in this order. The reason always comes before the formula.

1 · Intuition

An everyday picture, no symbols yet, and the key idea said three different ways.

2 · Example

Real numbers, every arithmetic step shown, so you can redo it on paper.

3 · Definition

The precise statement. Every symbol is explained in words.

4 · Why · Where · How

Why do we need it? Where is it used? How is it used in practice?

5 · Picture and play

Diagrams and interactive plots: drag points, move sliders, draw new samples and watch what changes.

6 · Careful

The usual confusions, written as ✗ wrong idea and ✓ right idea.

7 · In your projects · Say it right

Where the idea sits in your A/B framework or forecasting model, and how to phrase subtle points precisely in an interview.

8 · Notebook and check

A short "write this in your notebook" summary, a quick self-test, and at the end of each chapter a recap, Python code, quizzes and practice problems.

Try your first interactive

The boxes marked Interactive react to you. This one shows the single most important idea in statistics: the same process gives different data every time.

The true conversion rate is set by the slider (this is the model, which in real life you never see). Press Run an experiment: 500 users arrive, each converts with that probability, and the observed rate is added to the histogram. Press Run 100 experiments and watch the observed rates pile up around the true rate. Notice: no single experiment gives exactly the true rate, but together they form a bell shape whose width you can predict. Try 50 users instead of 500: the bell gets wider.

The roadmap

Eighteen chapters, each building on the last. The small grey text names the syllabus modules each chapter covers.

What matters most (your P0 list). Probability and random variables (4.2–4.4), expectation, variance, covariance and correlation (4.5, 4.15), Pearson vs Spearman (4.15), LLN and CLT (4.13), the nine key distributions (4.7–4.11), Q-Q plots (4.17), KDE (4.16) and Box-Cox (4.18).

How to read the symbols

Do not memorise this table now. Come back whenever a symbol looks strange.

SymbolSay it asMeaning
$P(A)$"probability of A"How likely the event $A$ is, a number from 0 to 1.
$P(A\mid B)$"probability of A given B"How likely $A$ is once we know $B$ happened.
$X$ vs $x$"capital X", "small x"$X$ is a random variable (the process that produces a number); $x$ is one value it took.
$X \sim N(\mu, \sigma^2)$"X is distributed as Normal with mean mu and variance sigma squared"$X$ follows that distribution. In maths we write the variance; NumPyro and SciPy take the standard deviation.
$p(x)$, $f(x)$"p of x"A probability mass function (for counts) or a density (for continuous values).
$F(x)$"F of x"The cumulative distribution function, $P(X \le x)$.
$E[X]$, $\mu$"expected value of X", "mu"The long-run average of $X$.
$Var(X)$, $\sigma^2$, $\sigma$"variance", "sigma squared", "sigma"How spread out $X$ is; $\sigma$ is the standard deviation.
$\bar x$, $s$"x bar", "s"The sample mean and sample standard deviation, computed from data.
$\theta$"theta"A generic unknown parameter (a rate, a mean, a slope…).
$\sum_{i=1}^{n}$"sum from i equals 1 to n"Add up the terms for every data point.
iid"i-i-d"Independent and identically distributed: each value is a fresh draw from the same distribution.

Tips for studying

  • Play first, read second. If a paragraph confuses you, move the sliders in the plot below it, then reread.
  • Press "New sample" a lot. Watching numbers change from sample to sample is how statistical intuition is built.
  • Copy the notebook boxes. Writing the short summaries by hand is the fastest way to remember them.
  • Run the code. Each chapter ends with Python you can paste into a notebook (NumPy, SciPy, pandas, statsmodels).
  • Use the tools. Press / to search, use the sidebar to jump, switch dark mode at the top, and mark chapters complete as you go.

If a section feels too hard, it is almost always because a word from an earlier section is fuzzy. Find that word in the sidebar or the glossary at the end and reread it. Statistics takes two passes for everyone.

Chapter 4.1 · Syllabus Module 0

Statistical thinking

Before any formula, statistics is a way of looking at the world: the numbers in front of you are one random draw from a hidden process, and your job is to learn about that process. This chapter gives you the words (observation, feature, target, population, sample, parameter, statistic), the two directions (probability and statistics), the idea of a data-generating process, and the real difference between frequentist and Bayesian thinking.

  • Read any dataset as observations (rows) and variables (columns), and name each variable's type: continuous, discrete, nominal, ordinal or binary
  • Tell a population from a sample, and a parameter (fixed, usually unknown) from a statistic (computed from the sample, changes every time)
  • Name the six jobs of statistics: describe, estimate, infer, model, predict, decide
  • Explain the two directions: probability goes from a model to data; statistics goes from data back to the model
  • Describe a data-generating process (real process → random variables → distribution → observed data) for both of your projects
  • Explain frequentist vs Bayesian thinking in depth: what is random, what probability means, and how each one reasons

Data and variables: rows, columns, features and the target core

Picture a spreadsheet. Each row is one thing you looked at: one user, one day, one order. Each column is one thing you measured about it: the device, the minutes on the site, whether the user bought. That is all "data" means here: a table of measurements.

Usually one column is special. It is the thing you want to explain or predict, such as "did the user buy?" or "how many orders tomorrow?". That column is the target. The other columns are the clues you use, the features.

Columns also come in different kinds. "Minutes" is a number you can average. "Device" is a label: phone, tablet, desktop. You cannot average labels. Knowing the kind of each column tells you which summaries make sense and, later, which probability distribution fits it.

Three ways to say it:

  • Picture: a dataset is a table; a row is one observed thing, a column is one measured property.
  • Numbers: 5 users × 6 columns; the last column (0 or 1: bought or not) is the target; its average, 0.4, is the conversion rate.
  • Slogan: rows are things, columns are questions, and the target is the question you care about.

Here is a tiny dataset of five users who reached a checkout page (the same table is drawn in the figure below).

userdeviceplanpagesminutesconverted
u1phonebasic32.50
u2desktoppro79.11
u3phonebasic10.80
u4tabletenterprise1215.31
u5phonepro44.00
  1. Count the rows: 5. So there are $n = 5$ observations.
  2. Count the columns: 6. "user" is only a name tag (an ID), so there are 5 real variables.
  3. The question we care about is "did they buy?", so converted is the target. The other four (device, plan, pages, minutes) are features.
  4. Name the kinds: device = labels with no order (nominal); plan = labels with an order basic < pro < enterprise (ordinal); pages = whole-number counts (discrete); minutes = any value in a range (continuous); converted = only 0 or 1 (binary).
  5. Average the target: $(0+1+0+1+0)/5 = 2/5 = 0.4$. The mean of a 0/1 column is the proportion of 1s: here a 40% conversion rate.

A dataset is a table with $n$ rows and some columns.

  • An observation (also called a unit, record, example or row) is one thing measured: one user, one day.
  • A variable is one measured property, a column. It "varies" from observation to observation, hence the name.
  • The target (also response, outcome, label, dependent variable; usually written $y$) is the variable we want to explain or predict.
  • A feature (also predictor, input, explanatory or independent variable; written $x$) is a variable used to explain or predict the target. Covariate means the same thing; people often use it for a background variable they want to "adjust for", such as a user's spending before an experiment (you will meet this in Chapter 5.12).

Variable types:

  • Numeric (quantitative): arithmetic makes sense.
    • Continuous: any value in a range (minutes, revenue, temperature).
    • Discrete: separate, countable values, usually counts $0, 1, 2, \dots$ (orders per day, pages viewed).
  • Categorical (qualitative): the values are labels for groups.
    • Nominal: no natural order (device, country).
    • Ordinal: a natural order, but the gaps are not necessarily equal (satisfaction 1 to 5, plan tiers).
    • Binary: exactly two values (converted yes/no, group A/B). We usually code it as 0 and 1.
Why do we need it?

Every later method needs to know which column is the target and what kind of values each column holds. Averaging a label, or modelling a yes/no column with a bell curve, gives answers that look precise but mean nothing.

Where is it used?

Every pandas DataFrame and every ML training set (X = features, y = target); choosing a likelihood (binary → Bernoulli, counts → Poisson or Negative Binomial, categories → Multinomial); one-hot encoding in scikit-learn; the design matrix of a regression or a forecasting model.

How is it used?

Before any analysis, list the columns, mark the target, and write each column's type. In pandas, check df.dtypes, convert label columns to category (with ordered=True for ordinal ones), one-hot encode nominal features, and keep 0/1 targets as integers.

features = predictors = covariates (x) target (y) userdeviceplanpagesminutesconverted u1phonebasic32.50 u2desktoppro79.11 u3phonebasic10.80 u4tabletenterprise1215.31 u5phonepro44.00 IDnominalordinaldiscretecontinuousbinary name tagcategoricalcategoricalnumericnumericcategorical ← one observation (a row) the purple box is one variable (a column) n = 5 observations
The example table, annotated. Each row (blue band) is one observation; each column (purple box) is one variable. The bottom lines give each column's type. The ID column is only a name tag: it is not a feature.

Pick a variable from the list. The purple path through the tree shows the two questions that decide its type: can I do arithmetic with the values? and then how many values, and is there an order? Read the readout: it tells you which summaries make sense, how to feed the column to a model, and which distribution family usually describes it (taught in Chapters 4.7–4.10). Try "A/B group" and "converted?": both are binary, even though one is written with letters and one with digits.

Ten users used these devices: 5 phones, 3 desktops, 2 tablets. A computer needs numbers, so someone gave each device a code. Switch between Coding 1, 2 and 3. Watch the purple "mean code" move to a different device each time, although the data never changed. The mode (most common device) and the shares (50%, 30%, 20%) never move. Then switch on one-hot columns: the mean of each 0/1 column is exactly that device's share, which is meaningful.

"If a column contains numbers, it is a numeric variable."

Postal codes, user IDs, store numbers and A/B group codes are labels written with digits. Ask: does adding or averaging two values mean anything? If not, it is categorical.

"Ordinal levels are equally spaced, so I can average satisfaction scores like any number."

The step from "1 = very unhappy" to "2" need not equal the step from "4" to "5". The median and the share in each level are always safe. A mean of ordinal codes is a common shortcut, but it adds an assumption (equal steps), so say so when you use it.

"Feature, predictor and covariate are three different things."

They are near-synonyms for "a column used to explain the target". Different fields prefer different words (ML says feature; regression says predictor; experiments say covariate). The same holds for target = response = outcome = label.

"The mean of a yes/no column is meaningless."

Coded as 0/1, its mean is exactly the proportion of yeses. That is why a conversion rate is just an average.

In an A/B framework like yours, each row is usually one randomized unit (a user), the variant is a binary or nominal column, segments are nominal columns, and each metric is a target whose type picks the likelihood: a binary conversion → Bernoulli/Binomial with a Beta prior; a categorical choice → Multinomial with a Dirichlet prior; a count → Poisson; a continuous metric → Normal or Student-t. In your forecasting model each row is a day: the target is demand $y_t$; the features are time (for the trend), Fourier terms (for seasonality), 0/1 holiday indicators and the exogenous regressors $X_t$.

Row = observation; column = variable. Target = response = $y$ (what we predict); features = predictors = covariates = $x$ (the clues).

Types: numeric (continuous / discrete counts) vs categorical (nominal / ordinal / binary).

Traps: digits do not make a column numeric; never average nominal codes; the mean of a 0/1 column is a proportion.

Quick check: what type is "number of support tickets per day"? And "country of the user"?

Tickets per day are counts $0, 1, 2, \dots$, so numeric, discrete (a Poisson or Negative Binomial candidate). Country is a label with no natural order, so categorical, nominal: summarize it with shares per country, and one-hot encode it for a model.

Population and sample core

To check the salt in a big pot of soup, you do not drink the whole pot. You stir it, then taste one spoonful. The pot is the population: everything you want to know about. The spoonful is the sample: the part you actually look at.

Stirring matters. If you skim only the top, where the oil floats, the spoonful lies about the pot. Stirring is what statisticians call random sampling: every part of the pot gets a fair chance to end up in your spoon.

And one spoonful is never exactly like the pot. A second spoonful tastes a tiny bit different. That small, unavoidable wobble is sampling variation.

Three ways to say it:

  • Picture: the population is the pot; the sample is a spoonful taken after stirring.
  • Numbers: 10 000 visitors with a 12% conversion rate; a random 500 of them showed 11.4%.
  • Slogan: we look at the sample, but we want to know about the population.

Last month 10 000 visitors reached your checkout page and 1 200 of them bought. You cannot inspect every session by hand, so you pick 500 sessions at random.

  1. Population: all 10 000 visitors. Its conversion rate is $1\,200 / 10\,000 = 0.12$.
  2. Sample: the 500 randomly chosen visitors. The sample size is $n = 500$.
  3. In the sample, 57 bought. The sample conversion rate is $57 / 500 = 0.114$.
  4. The gap: $0.114 - 0.12 = -0.006$. Nobody made a mistake: this is sampling variation. Another random 500 would give a slightly different number, perhaps $0.126$.

In real life you usually do not know the population's 0.12. That is the whole point: you use the 0.114 to learn about it.

  • The population is the complete set of units we want to learn about. It can be a real, finite list (all 10 000 visitors last month) or a conceptual one: all future visitors who will see this page, or all future days of demand. A conceptual population is really a process that keeps producing data.
  • A sample is the part of the population we actually observe. Its size is written $n$.
  • A random sample is chosen by chance, so that every unit has a known (often equal) chance of being picked. A simple random sample gives every set of $n$ units the same chance.
  • When data come from a process, we often assume they are iid (independent and identically distributed): each observation is a fresh draw from the same distribution, not influenced by the others. You will check when this fails in Chapter 5.4.
  • Sampling variation: different samples from the same population give different results.
Why do we need it?

We almost never see the whole population: it is too big, too expensive, or it lies in the future. Every conclusion is therefore built on a sample, and we must know whether that sample can speak for the population.

Where is it used?

A/B tests (the users in the test stand for all future users), opinion polls, quality checks on a production line, train/test splits in ML, and forecasting, where the history you have is one sample of what the demand process can produce.

How is it used?

Name the population first ("all users who will see checkout next quarter"). Then make sure the sample is chosen at random from it (randomize, do not hand-pick), record $n$, and treat every number from the sample as an estimate that comes with a wobble.

Population (the pot) parameter: true rate p (usually unknown) random sample (stir, then scoop) statistic p̂ close to p only the top rows (biased) p̂ far from p, and more data won't fix it
A random sample (blue) gives a rate close to the population's. A sample taken from a convenient corner (red) can be badly off, and collecting more of the same kind of data does not remove that error.

Each dot is one user of a population of 1 000; teal dots converted (120 of them, so the true rate is 0.12). Press Draw a sample a few times: the blue rings mark the sampled users, and the readout gives the sample rate. Notice it changes every time but stays near 0.12. Raise $n$ to 300: the rates wobble less. Now pick First n rows only: the rows are in sign-up order and early users bought more often, so this "sample" is far off, and drawing again changes nothing.

"The population is 100 million people, so a sample of 2 000 (0.002%) is far too small."

For a random sample, how precise the answer is depends mainly on $n$, not on how big the population is (as long as the population is much bigger than the sample). 2 000 random people tell you about 100 million almost as well as about 1 million.

"With enough data, a biased sample becomes fine."

Bias does not shrink with $n$. A million users from the wrong group still answer the wrong question. More data only shrinks the random wobble.

"The population is always a list of real people or things."

Often it is a process: all future users, all future days. Then "the population mean" means the long-run average the process would produce.

In an A/B framework like yours, the users in the experiment are the sample; the population is the users who will see the winning variant after launch. The conclusion only transfers if the sample looks like them: randomization inside the experiment, and an experiment window that covers the normal mix of weekdays, weekends and user types. In your forecasting model, the history you fit on is one sample of what the demand process can produce, and the population you care about is the future days.

Population = everything we want to know about (often a process). Sample = what we observe, size $n$.

Random sampling makes the sample stand in for the population; sampling variation = different samples give different numbers.

Trap: precision depends on $n$, not on population size; bias does not go away with more data.

Quick check: you test a new checkout only on users who arrive from email campaigns, 50 000 of them. Can you say how all users will react?

Not safely. Email users are not a random sample of all users (they are often more loyal). 50 000 of them gives a precise answer to "how do email users react?", but the population you care about is all users. The size of a sample does not fix who is in it.

Parameter vs statistic core

Two kinds of numbers appear in every analysis, and mixing them up causes most beginner confusion.

  • A parameter describes the population, such as the true conversion rate of all visitors. It is one fixed number, but usually unknown.
  • A statistic is any number you compute from the sample, such as 57/500 = 0.114. You always know it, but it changes every time you take a new sample.

Watch the word. "Statistic" here does not mean the school subject "statistics". It means one number calculated from data: a mean, a median, a maximum, a count. A "test statistic" (Chapter 5.6) is just a statistic that is used for a test.

Three ways to say it:

  • Picture: the parameter is the bullseye that never moves; each sample's statistic is a dart that lands near it.
  • Numbers: true rate 0.12 (parameter); three samples give 0.114, 0.126, 0.108 (statistics).
  • Slogan: Population → Parameter, Sample → Statistic (P with P, S with S).

Five sampled users spent these minutes on the site: 2.5, 9.1, 0.8, 15.3, 4.0. Each line below computes a different statistic from the same sample.

  1. Sample mean: $(2.5 + 9.1 + 0.8 + 15.3 + 4.0)/5 = 31.7/5 = 6.34$ minutes.
  2. Sample median: sort them, $0.8, 2.5, 4.0, 9.1, 15.3$; the middle one is $4.0$ minutes.
  3. Sample maximum: $15.3$ minutes. Sample range: $15.3 - 0.8 = 14.5$ minutes.
  4. The population mean $\mu$ (the average over all users) is a parameter. We do not know it. If we did, it might be about 5.9 minutes, and our 6.34 would be a reasonable guess for it.
  • A parameter is a number that describes the population or the process, for example the mean $\mu$, the standard deviation $\sigma$, a proportion $p$, or in general $\theta$ ("theta"). Parameters are written with Greek letters.
  • A statistic is any number computed from the sample alone: $T = t(x_1, \dots, x_n)$. It must not need any unknown parameter to compute. Statistics are written with Latin letters or hats: $\bar x$ (sample mean), $s$ (sample standard deviation), $\hat p$ (sample proportion), $\hat\theta$.
  • Matching pairs: $\mu \leftrightarrow \bar x$, $\;\sigma \leftrightarrow s$, $\;p \leftrightarrow \hat p$, $\;\theta \leftrightarrow \hat\theta$.
  • Because the sample is random, a statistic is random too: it has its own distribution across repeated samples, called the sampling distribution (Chapter 5.5).
  • A statistic used to guess a parameter is an estimator (the recipe), and the number it gives on your data is an estimate (Chapter 5.1).
Why do we need it?

We want the parameter (the truth about the population) but can only compute statistics. Keeping the two apart is what lets us ask the central question of statistics: how far from the parameter is my statistic likely to be?

Where is it used?

Every estimate: a conversion rate $\hat p$ for the true rate $p$, a sample mean for $\mu$, a fitted regression slope for the true slope, a test statistic in a z-test or t-test, and the summary numbers in every dashboard.

How is it used?

Name the parameter you care about in words ("the true conversion rate of variant B"). Choose a statistic that estimates it ($k/n$). Report the statistic together with how much it would wobble across samples (a standard error or an interval, taught in Guide 2).

Population parameter μ = 5.9 fixed, unknown sample 1: x̄ = 5.4 sample 2: x̄ = 6.3 sample 3: x̄ = 5.8 5.0 6.0 7.0 μ statistics scatter around the fixed parameter
One fixed parameter (green), many possible statistics (blue). Each new sample gives a new value of x̄, scattered around μ.

The top plot is a population of 2 000 users' minutes on the site (sessions time out at 30 minutes). The green line is the parameter. Press Draw 1 sample: the blue ticks are your $n$ sampled users and the purple line is your statistic. Press Draw 100 samples: the bottom plot collects the statistic from every sample, and it spreads around the green line. Change $n$ and see the spread shrink. Then switch to maximum: the sample maximum is always at or below the population maximum, so it falls short on average.

"A statistic is the subject called statistics."

A statistic is one number computed from a sample: a mean, a count, a maximum. "Test statistic" just means a statistic used in a test.

"My sample mean is the population mean."

It is an estimate of the population mean. It is usually close, almost never exactly equal, and it changes with the sample.

"If I collect more data, the parameter changes."

The parameter is a property of the population; your data do not move it. More data changes your statistic (and, in Bayesian language, your uncertainty about the parameter), not the parameter itself.

"The conversion rate of variant B is 11.4%."

"The observed conversion rate of B in this experiment is 11.4%; it estimates B's true conversion rate."

Model answer: "The true rate is a parameter: fixed but unknown. 11.4% is a statistic from our sample. Another sample would give a slightly different number, so I always report it with its uncertainty."

In an A/B framework like yours, $\theta_A$ and $\theta_B$, the true conversion rates, are parameters; $k_A/n_A$ and $k_B/n_B$ are statistics. Your decision quantity $P(\theta_B \gt \theta_A \mid D)$ is a statement about the parameters, computed from the data $D$. In your forecasting model, the trend slope, the changepoint adjustments $\delta_j$, the seasonality coefficients and the noise scale $\sigma$ are parameters; the error of a backtest, or the standard deviation of the residuals, is a statistic.

Parameter = describes the population, Greek ($\mu, \sigma, p, \theta$), fixed and usually unknown.

Statistic = computed from the sample, Latin or hat ($\bar x, s, \hat p, \hat\theta$), known but changes from sample to sample.

Trap: "statistic" = one number from data, not the subject. Never call an observed rate "the true rate".

Quick check: which are statistics? (a) the median of your 200 sampled order values; (b) the true average order value of all customers; (c) the largest value in your sample.

(a) and (c) are statistics: you can compute them from the sample. (b) is a parameter: it describes the whole population and you usually do not know it.

What statistics is: six jobs

Statistics is learning from data when there is noise. "Noise" means the random wobble that makes two days, two users or two samples differ even when nothing important changed.

Think of a shop owner looking at a week of sales. First they describe it ("weekends were busy"). Then they guess how sales work ("a normal level plus a weekend boost, plus luck"). They put a number on the boost (estimate), wonder whether it is real or just this week's luck (infer), guess next Saturday (predict), and finally choose how much stock to buy (decide). Those are the six jobs of statistics, and every project uses most of them.

Three ways to say it:

  • Picture: describe what you see, model how it was made, then use the model to estimate, infer, predict and decide.
  • Numbers: "130 orders a day" (describe) → "weekend adds about 63" (estimate) → "next Saturday about 175" (predict) → "stock 190" (decide).
  • Slogan: statistics turns noisy data into reliable conclusions and good decisions.

One week of daily orders, Monday to Sunday: 100, 110, 110, 120, 120, 170, 180.

  1. Describe. Total $= 910$, so the mean is $910/7 = 130$ orders a day. Weekdays: $(100+110+110+120+120)/5 = 560/5 = 112$. Weekend: $(170+180)/2 = 175$.
  2. Model. Write down how the numbers might have been made: orders $=$ weekday level $+$ weekend bump (on Saturday and Sunday) $+$ random noise.
  3. Estimate. Put a number on the unknown bump: $175 - 112 = 63$ extra orders per weekend day.
  4. Infer. Ask whether the bump is real for all weeks or just luck of this one. With one week we cannot say much; with many weeks we can measure how sure we are (Guide 2).
  5. Predict. Next Saturday: about $112 + 63 = 175$ orders, give or take the usual noise.
  6. Decide. Stock a bit more than 175, perhaps about 190, so that a slightly busier Saturday does not run out. The exact amount depends on the noise and on what running out costs (see the widget).
  • Descriptive statistics: summarize the data you have (means, medians, spreads, counts, plots). No claim beyond these data.
  • Inferential statistics: draw conclusions about the population or process beyond the data, with a statement of uncertainty (intervals, tests, posterior probabilities).
  • Statistical modeling: write down a probability model of how the data were produced, with unknown parameters (the data-generating process, later in this chapter).
  • Estimation: give a value (a point estimate) or a range (an interval) for an unknown parameter. It is one kind of inference.
  • Prediction: say what a new or future observation will be, with uncertainty. It targets something you will later observe, not a parameter.
  • Decision-making: choose an action using the predictions or estimates, their uncertainty, and the costs of each kind of mistake.
Why do we need it?

Naming the job tells you which tool and which kind of answer you need. "What was our conversion rate?" (describe) needs no uncertainty; "Is B better?" (infer) and "How much stock?" (decide) need a model and honest uncertainty.

Where is it used?

Dashboards (describe), A/B tests (estimate and infer), demand forecasting (model and predict), inventory and capacity planning (decide), ML model training (model and estimate), and credit or medical risk scoring (predict and decide).

How is it used?

Start each analysis by writing the question and its job. Describe the data first (always plot it). Build a model only if you need to go beyond the data. Report estimates and predictions with ranges, and make decisions with the costs written down.

1 · DescribeWhat do the data look like?"130 orders a day" 2 · ModelWhat process made them?level + weekend bump + noise 3 · EstimateWhat is the unknown number?"bump ≈ 63 orders" 4 · InferDoes it hold beyond these data?real bump, or luck? 5 · PredictWhat happens next?"next Saturday ≈ 175" 6 · DecideWhat should we do?"stock about 190"
The six jobs in the order they usually happen, with the one-week demand example. Describing stays inside the data; the other five reach beyond it and therefore need a model and a statement of uncertainty.

The blue dots are 56 days of simulated orders (day 0 is a Monday). Click through the six jobs from left to right and read the readout each time. In Estimate, compare the estimates with the true values used to simulate. In Predict, notice the band: a single future day is uncertain even if the model is right. In Decide, move the risk slider: a lower risk of running out needs more stock. Press New data to see every number change a little.

"Statistics is calculating averages."

Describing is only the first of six jobs. Most of the value (and most of the difficulty) is in going beyond the data: inferring, predicting and deciding under uncertainty.

"Estimating a parameter and predicting a value are the same thing."

The weekend bump (a parameter) can be pinned down very precisely with lots of data. Next Saturday's orders (a future observation) always keep the day-to-day noise. So a prediction range is always wider than the matching estimate range.

"Once I have a forecast, the decision follows."

A decision also needs the costs: is running out worse than wasting stock? The same forecast gives different best actions for different costs.

In an A/B framework like yours: you describe the observed conversion rates per variant, model them (for example Beta-Binomial), estimate each variant's rate (its posterior), infer whether B beats A ($P(\theta_B \gt \theta_A \mid D)$), and decide whether to ship. In your forecasting model you model demand as trend + seasonality + holidays + regressors + noise, estimate the parameters with SVI, predict future days as a predictive distribution, and someone decides capacity or stock from it.

Six jobs: describe → model → estimate → infer → predict → decide.

Describe = about the data in hand. Estimate = a parameter. Predict = a future observation (always noisier). Decide = needs costs too.

Trap: an estimate's range is narrower than a prediction's range.

Quick check: "Last month our average order value was 42 dollars." "Next month's average order value will be between 40 and 46 dollars." Which jobs are these?

The first only summarizes data you already have: describe. The second is about a value you will observe in the future: predict (with an uncertainty range, as a prediction should have).

Probability vs statistics: two directions core

Imagine a baker and a food critic. The baker knows the recipe and asks: "what kinds of cakes will come out of my oven?" The critic tastes a cake and asks: "what recipe could have made this?"

Probability is the baker's question: start from a known model (a coin with a known chance of heads) and work out which data are likely. Statistics is the critic's question: start from the data you saw and work back to the unknown model. Same kitchen, opposite directions.

Three ways to say it:

  • Picture: probability walks from the model to the data; statistics walks back from the data to the model.
  • Numbers: "if the coin is fair, 7 heads in 10 flips has probability 0.117" (probability); "we saw 7 heads in 10, so the chance of heads is probably near 0.7, but 0.5 is not ruled out" (statistics).
  • Slogan: probability: model → data. Statistics: data → model.

Both directions with the same coin and the same formula. The probability of exactly $k$ heads in $n$ flips when the chance of heads is $\theta$ is $\binom{n}{k}\theta^k(1-\theta)^{n-k}$ (you will meet this Binomial formula properly in Chapter 4.7). Here $\binom{10}{7} = 120$ is the number of ways to place 7 heads among 10 flips.

  1. Probability direction. The model is known: a fair coin, $\theta = 0.5$. $P(7 \text{ heads}) = 120 \times 0.5^{10} = 120/1024 \approx 0.117$.
  2. Statistics direction. Now we only know the data: 7 heads in 10 flips. Try several values of $\theta$ and ask how well each explains the data.
  3. $\theta = 0.7$: $120 \times 0.7^7 \times 0.3^3 = 120 \times 0.0824 \times 0.027 \approx 0.267$. The best fit, because $7/10 = 0.7$.
  4. $\theta = 0.5$: $0.117$, which is $0.117/0.267 \approx 0.44$ times as good. Still quite plausible.
  5. $\theta = 0.3$: $120 \times 0.3^7 \times 0.7^3 \approx 0.009$, only $0.034$ times as good. Hard to believe.
  6. Conclusion: the data point to $\theta$ near 0.7, but 10 flips cannot rule out a fair coin.
  • Probability (probability theory): given a model with known parameters $\theta$, compute the chance of possible data, $P(\text{data} \mid \theta)$. The answer is a set of numbers that add up to 1 over all possible datasets.
  • Statistics (statistical inference): given observed data $D$, learn about the unknown $\theta$ (and check the model itself).
  • The bridge is one formula read two ways. With $\theta$ fixed and the data varying, $p(D \mid \theta)$ is a probability distribution over datasets. With the data fixed and $\theta$ varying, the same expression is the likelihood $L(\theta) = p(D \mid \theta)$: a score of how well each $\theta$ explains the data you saw (Chapter 5.2). The likelihood is not a probability distribution over $\theta$.
Why do we need it?

You can only reason backwards from data if you can first reason forwards: to judge whether 7 heads is surprising for a fair coin, you must know what a fair coin usually produces. Every statistical method uses probability as its engine.

Where is it used?

Forwards: simulating A/B tests to plan sample sizes, prior predictive checks, generating forecast paths. Backwards: maximum likelihood, Bayesian posteriors, hypothesis tests, fitting any ML model to data.

How is it used?

Write the model as "data given parameters", $p(D \mid \theta)$. Use it forwards to simulate fake data and check that it looks realistic; use it backwards (as a likelihood, plus a prior if Bayesian) to learn $\theta$ from the real data.

Model coin with chance θ (the recipe) Data 7 heads in 10 flips (the cake) probability: "what data will this model give?" statistics: "what model made these data?" both directions use the same formula p(data | θ)
Probability (orange) reasons from a known model to the data it tends to produce. Statistics (blue) reasons from observed data back to the unknown model.

Start in Probability mode: set the model's conversion rate $p$ and look at the orange bars, the chance of each possible result $k$ out of 20 users. Press Simulate data from the model a few times: the blue dot (your dataset) lands mostly where the bars are tall. Then switch to Statistics mode: the curve now runs over possible values of $p$, for the data $k$ you have, and peaks at $\hat p = k/20$. Drag the teal handle to try other values of $p$ and read how much worse they explain the data. The green line marks the $p$ that really produced the simulated data.

"The likelihood $L(\theta)$ is the probability that $\theta$ is the true value."

It is the probability of the data if $\theta$ were true. It need not add up to 1 over $\theta$. To get real probabilities about $\theta$ you need a prior and Bayes' rule (Chapter 6.1).

"7 heads in 10 proves the coin is biased towards heads."

A fair coin gives 7 or more heads in 17% of runs of 10 flips. The data lean towards $\theta \approx 0.7$, but they are weak evidence. Statistics measures how strong the evidence is, it does not just pick the best-fitting value.

Both projects use both directions. Forwards (probability): in an A/B framework like yours you can simulate conversions from assumed rates to plan how many users you need, or simulate from the prior to check it gives believable rates (a prior predictive check, Chapter 6.2); your forecasting model, once fitted, is run forwards to produce future demand paths. Backwards (statistics): inferring $\theta_A$ and $\theta_B$ from observed conversions, and learning trend, seasonality and noise from the demand history.

Probability: model → data, $P(\text{data} \mid \theta)$ with $\theta$ known.

Statistics: data → model, learn the unknown $\theta$ from observed $D$.

Same formula, read as a function of $\theta$ = the likelihood $L(\theta) = p(D \mid \theta)$. Trap: the likelihood is not a probability distribution over $\theta$.

Quick check: "If the true conversion rate is 10%, how many of 1 000 users will convert?" and "53 of 500 users converted; what is the conversion rate?" Which is probability and which is statistics?

The first starts from a known model (10%) and asks about data: probability. (The expected number is $1\,000 \times 0.1 = 100$, give or take some.) The second starts from data and asks about the unknown rate: statistics. (The natural estimate is $53/500 = 0.106$.)

The data-generating process core

Behind every dataset there is a machine that made it: real people deciding whether to buy, real customers ordering food on a rainy Saturday. You never see the machine, only what comes out of it.

A statistician describes the machine with a probability model, for example: "each user flips a hidden, slightly unfair coin; heads means buy". If you could run the same machine again on a fresh day, you would get different data, because part of the machine is random. But the machine itself would be the same.

That is the most important idea in this guide: your dataset is one random draw from a process. Statistics is about the process, not about the one draw.

Three ways to say it:

  • Picture: a machine with a hidden setting; every run prints a different dataset; we study the prints to learn the setting.
  • Numbers: the same process (each user buys with chance 0.10) gave 9, 13 and 11 buyers out of 100 on three days.
  • Slogan: same process, different data, every time.

Follow the four steps for each of your two kinds of data.

  1. Real process. (a) Visitors reach checkout and decide whether to buy. (b) Customers place orders every day; demand grows slowly and is higher at weekends.
  2. Random variables (numbers we do not know until they happen; Chapter 4.4). (a) $Y_i = 1$ if visitor $i$ buys, $0$ if not. (b) $Y_t$ = orders on day $t$.
  3. Distribution (the rule for how likely each value is). (a) $Y_i \sim \text{Bernoulli}(\theta)$: "buy" with probability $\theta$, independently for each visitor. (b) $Y_t = 100 + 0.5\,t + 40 \cdot \text{weekend}_t + \varepsilon_t$ with noise $\varepsilon_t \sim N(0, 10^2)$, where weekend$_t$ is 1 on Saturdays and Sundays and 0 otherwise.
  4. Observed data (one run of the machine). (a) $0, 0, 1, 0, 0, 1, \dots$: 57 buyers among 500 visitors. (b) Day 5 is a Saturday: its average is $100 + 0.5 \times 5 + 40 = 142.5$ orders; this time the noise was $+8.7$, so we observed $151.2$.

If we ran the same process again, step 4 would change; steps 1 to 3 would not.

  • The data-generating process (DGP) is the true, unknown mechanism that produced the data: real process → random variables → their distribution → the observed values.
  • A statistical model is our written-down guess of the DGP: a family of distributions $p(y \mid \theta)$, one for each value of the parameters $\theta$. For example $\{\text{Bernoulli}(\theta) : 0 \le \theta \le 1\}$.
  • The observed data $D = \{y_1, \dots, y_n\}$ are one realization (one run) of the process.
  • A model usually splits each observation into a systematic part (what we can explain: trend, weekend effect, the conversion rate) and a random part (noise), and it states assumptions such as independence and the shape of the noise. A famous warning from the statistician George Box: "all models are wrong, but some are useful".
Why do we need it?

Without a picture of how the data were made, you cannot say which differences are real and which are noise, you cannot choose a likelihood, and you cannot simulate what the future might look like.

Where is it used?

Every probabilistic model: NumPyro and Stan models are literally written as a DGP; A/B testing (conversions as Bernoulli draws); forecasting (trend + seasonality + noise); simulation studies; prior and posterior predictive checks; synthetic-data testing of a pipeline.

How is it used?

Before fitting anything, write the DGP in words and then in symbols: what is random, what distribution, which parameters. Simulate fake datasets from it and check that they look like real data. Then fit the same model to the real data.

Real processRandom variablesDistributionObserved data A/Bconversions Forecastdaily demand people visit checkoutand decide to buy Yᵢ = 1 if user i buys,0 if not Yᵢ ~ Bernoulli(θ)independent users 0, 0, 1, 0, 0, 1, …57 buyers of 500 customers order;growth, weekends, events Yₜ = orderson day t Yₜ ~ N(μₜ, σ²)μₜ = trend + season + … 131, 118, 176, …one history the model (orange) is our guess of the machine
The data-generating process, step by step, for both projects. Statistics works backwards along these arrows: from the observed data (blue) to the model (orange).

The green line is the process: the average demand for each day (growing slowly, higher at weekends). Each press of New dataset runs the process again and draws a new blue history; older histories stay as faint lines. Notice that the green line never moves while the blue data always do. Set the noise $\sigma$ to 0: every dataset is identical (real data never look like this). Switch to Student-t noise: same scale, but now and then a day is far from the line. Turn off Show the process to see what you really get in practice: only the blue data.

Each square is a user; teal means the user converted. The process is simple: every user converts with probability $p$, independently of the others. Press New dataset several times and watch the count change: 9, 13, 11… although $p$ never changes. Press Run 50 datasets: the counts pile up around $100 \times p$ (green line). Move $p$ and start again.

"The model is the data" or "the data are the truth".

The data are one noisy run of the process. The model is our guess of the process. The truth is the process itself, which we never see directly.

"My model fits this dataset well, so it is the true process."

Many different processes can produce similar-looking data, and a flexible model can fit noise. Check a model on new data, or by simulating from it and comparing with the real data (posterior predictive checks, Chapter 6.8).

"Noise is a nuisance; the model should get rid of it."

Noise is part of the model. Its shape (Normal, Student-t, Negative Binomial…) is a modelling choice, and it decides how much the model trusts unusual points (Chapter 4.9).

A/B framework. In an A/B framework like yours, the DGP of a conversion metric is: each user in variant A converts with probability $\theta_A$, independently, so the number of converters is $\text{Binomial}(n_A, \theta_A)$; the Beta-Binomial model adds a Beta prior on $\theta_A$. A categorical metric is Categorical/Multinomial with a Dirichlet prior. Count and continuous metrics use Poisson, Normal or Student-t likelihoods.

Forecasting model. Your DGP is $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$: trend, seasonality, holidays and regressors form the systematic part, and the random part comes from the likelihood (Normal or Student-t noise around the mean, or, for counts, a Negative Binomial draw around it). Writing the DGP first is how you choose likelihoods and priors; simulating from it before seeing data is the prior predictive check of Chapter 6.2.

DGP: real process → random variables → distribution → observed data. A model is our guess of it: $p(y \mid \theta)$.

Data = one realization. Same process, different data every time.

Conversions: $Y_i \sim \text{Bernoulli}(\theta)$. Demand: $y_t$ = systematic part + noise. Trap: a good fit does not prove the model is the true process.

Quick check: in $Y_t = 100 + 0.5\,t + 40 \cdot \text{weekend}_t + \varepsilon_t$, which parts change if you "rerun" the same month?

Only the noise $\varepsilon_t$, and therefore the observed $Y_t$. The systematic part $100 + 0.5t + 40\cdot\text{weekend}_t$ (and the parameters 100, 0.5, 40 and $\sigma$) belong to the process and stay the same.

Frequentist vs Bayesian thinking core

A friend flips a coin and covers it with a cup. Is it heads?

  • A frequentist says: "the coin already landed; it is heads or it is tails. There is no probability left in it. Probability describes the flipping: if we flipped many times, about half would be heads."
  • A Bayesian says: "I don't know which side is up, so for me heads has probability 0.5. Probability describes my uncertainty. If I peek, my probability jumps to 0 or 1."

Now replace the coin under the cup with an unknown conversion rate $\theta$. The frequentist treats $\theta$ as a fixed number and puts the randomness in the data ("what would happen if we repeated the experiment?"). The Bayesian puts a probability distribution on $\theta$ itself and reasons from the data actually observed ("given what I saw, what do I believe?"). That is the real difference. It is much deeper than "p-values vs posteriors".

Three ways to say it:

  • Picture: frequentist: one fixed target, many imaginary repeats of the experiment. Bayesian: one real experiment, a cloud of belief over the target.
  • Numbers: 7 heads in 10. Frequentist: "if $\theta$ were 0.5, 17% of repeats would give 7 or more heads". Bayesian: "given these flips, $P(\theta \gt 0.5) = 0.89$".
  • Slogan: frequentist: $\theta$ is fixed, data are random. Bayesian: data are fixed (once seen), $\theta$ is uncertain.

We flip a coin 10 times and see 7 heads. Both schools analyse the same data.

  1. Frequentist estimate. $\hat\theta = 7/10 = 0.7$.
  2. Frequentist check. "If the coin were fair ($\theta = 0.5$) and we repeated the 10 flips many times, how often would we see 7 or more heads?" Count the ways: $\binom{10}{7} + \binom{10}{8} + \binom{10}{9} + \binom{10}{10} = 120 + 45 + 10 + 1 = 176$ out of $2^{10} = 1024$, so $176/1024 \approx 0.172$. Not unusual: no strong evidence of bias.
  3. Frequentist interval. A 95% confidence interval (Wilson method, Chapter 5.8) is $[0.40, 0.89]$. Its promise is about the method: in 95% of repeated experiments, intervals built this way contain the true $\theta$.
  4. Bayesian prior. Before the flips, take a flat prior: every $\theta$ between 0 and 1 equally believable (a $\text{Beta}(1, 1)$ distribution, Chapter 4.11).
  5. Bayesian posterior. After 7 heads and 3 tails the belief becomes $\text{Beta}(1+7,\; 1+3) = \text{Beta}(8, 4)$ (why the counts just add: Chapter 6.3). Its mean is $8/12 \approx 0.667$.
  6. Bayesian answers. $P(\theta \gt 0.5 \mid \text{data}) \approx 0.887$, and a 95% credible interval is $[0.39, 0.89]$: given these data and this prior, $\theta$ lies there with probability 0.95.

The intervals almost match, but they mean different things: one is a promise about a procedure over repeats; the other is a probability about $\theta$ given this one dataset.

FrequentistBayesian
What is random?The data, and anything computed from them, across imagined repetitions of the experiment.Anything we are unsure about, including $\theta$. The observed data are fixed once seen.
What does probability mean?Long-run relative frequency over many repetitions.Degree of belief (plausibility) given the information we have.
The parameter $\theta$A fixed, unknown constant. "$P(\theta \gt 0.5)$" has no meaning: $\theta$ either is or is not above 0.5.Unknown, so it gets a distribution: the prior $p(\theta)$, updated to the posterior $p(\theta \mid D)$.
How it reasons"If $\theta$ had this value, how would my procedure behave over many repeats?""Given the data I actually saw, what should I believe?" (it conditions on $D$)
Typical toolsEstimators and their sampling distributions, confidence intervals, p-values.Prior, likelihood, posterior, posterior predictive distribution.
Extra input it needsThe sampling plan: how the data were collected and when you would have stopped.A prior distribution for $\theta$.
How a method is judgedBy its long-run error rates (coverage, false-positive rate).By coherence with the model and prior (and, in practice, by calibration checks too).

Both use the same likelihood $p(D \mid \theta)$. With a lot of data and a weak prior, their numbers often come out close; the interpretations stay different.

Why do we need it?

The two views answer different questions. Knowing which one you are using tells you how to state results correctly ("95% of such intervals cover θ" vs "θ is in here with probability 0.95") and what each method needs from you (a sampling plan, or a prior).

Where is it used?

Frequentist: classical A/B testing with p-values, confidence intervals, most of scikit-learn and statsmodels. Bayesian: NumPyro, Stan and PyMC models, Bayesian A/B testing with $P(\theta_B \gt \theta_A \mid D)$, Bayesian forecasting with predictive distributions, hierarchical models.

How is it used?

Frequentist: fix the design and the stopping rule in advance, compute an estimate, an interval and perhaps a p-value, and interpret them as properties of the procedure. Bayesian: choose a prior, compute the posterior from the data, and read direct probabilities about the parameters and future data from it.

Frequentist Bayesian θ: one fixed, unknown number D₁: 7 of 10 D₂: 5 of 10 D₃: 8 of 10 the data vary over imagined repeats probability = long-run frequency D: 7 of 10 observed, fixed p(θ | D) θ is uncertain: a distribution probability = degree of belief, given D
Left: the frequentist keeps θ fixed and asks how the data would vary across imagined repetitions (dashed boxes never happened). Right: the Bayesian keeps the one observed dataset fixed and describes θ with a probability distribution.

Left (frequentist): $\theta$ is held fixed at the value on the second slider, and the bars show how the number of heads would vary if the 10 flips were repeated; red bars are results at least as high as yours. Right (Bayesian): your flips are fixed, and the orange curve is the belief about $\theta$ after seeing them (dashed: the belief before). Move $k$ and watch both sides respond. Try the strong prior: the Bayesian answer changes, the frequentist one does not. Then move the tested $\theta$: only the left side changes.

A coin with true $\theta = 0.7$ (known here only because we simulate it) is flipped up to 400 times. Move the slider to reveal more flips. With the flat prior, the blue confidence interval and the orange credible interval are already close after 10 flips (this sequence starts with 7 heads in 10, our example), and both shrink around the truth as $n$ grows. Switch to the strong, wrong prior (centred on 0.5): with few flips the Bayesian interval is pulled towards 0.5; with hundreds of flips the data win. Press New sample for a different sequence of flips.

Going deeper: the stopping rule. You see 9 heads and 3 tails. Person A had planned "flip exactly 12 times". Person B had planned "flip until the 3rd tail". Same data, different plans.

  • Frequentist: "how often, with a fair coin, would my plan give a result at least this extreme?" For A it is $P(\ge 9 \text{ heads in } 12) = 299/4096 \approx 0.073$. For B it is $P(\ge 9 \text{ heads before the 3rd tail}) \approx 0.033$. The evidence depends on the plan, because the imagined repeats differ.
  • Bayesian: both likelihoods are proportional to $\theta^9(1-\theta)^3$, so with a flat prior both people get the posterior $\text{Beta}(10, 4)$ and $P(\theta \gt 0.5 \mid D) \approx 0.954$. Only the data seen matter, not what you would have done with other data.

This is why frequentists insist that you fix the design in advance (and why peeking at an A/B test is a frequentist problem, Chapter 5.11), while Bayesians insist that you state your prior.

"Frequentist means p-values; Bayesian means posteriors."

Those are tools that follow from the two views. The real difference is what is treated as random (the data vs the parameter), what probability means (long-run frequency vs degree of belief), and how you reason (over imagined repeats vs conditional on the observed data).

"Bayesian is subjective and frequentist is objective."

Both make choices: the model and likelihood (both), the sampling plan and the test (frequentist), the prior (Bayesian). The Bayesian choices are written down openly as a prior.

"A 95% confidence interval contains θ with probability 95%."

For the frequentist, θ is fixed, so a computed interval either contains it or not. The 95% describes the method over repeated experiments. "Probability 95% that θ is inside" is the Bayesian credible interval's statement (Chapter 5.8, Chapter 6.4).

"In the Bayesian view, θ is a random number that keeps changing."

θ is still one unknown value in the world. The distribution describes our knowledge of it, which sharpens as data arrive.

"The difference is that frequentists use p-values and Bayesians use priors."

"The difference is what is random and what probability means. Frequentists treat the parameter as a fixed unknown and the data as random, and they judge methods by how they behave over repeated experiments. Bayesians treat the unknown parameter as uncertain, describe it with a probability distribution, and condition on the data actually observed."

Model answer: "In my A/B framework I report $P(\theta_B \gt \theta_A \mid D)$: a direct probability about the parameters given the data, which only the Bayesian view allows. A frequentist test would instead tell me how surprising the data would be if there were no difference, over imagined repeats. With lots of data and weak priors the numbers often agree, but the statements mean different things, and the Bayesian one needs a prior while the frequentist one needs a fixed design."

Your A/B framework is Bayesian: $\theta_A$ and $\theta_B$ get priors, the posterior conditions on the observed conversions, and decisions use direct statements like $P(\theta_B \gt \theta_A \mid D)$. Your forecasting model is Bayesian too: it gives a predictive distribution $p(y_{\text{future}} \mid D)$ conditioned on the observed history. Frequentist ideas still matter for both: you can (and should) check how a Bayesian decision rule behaves across many simulated experiments, for example how often it would declare a winner in A/A tests where nothing differs (Chapter 5.11).

Frequentist: θ fixed, data random; probability = long-run frequency; reason over imagined repeats; needs a fixed design.

Bayesian: θ uncertain → $p(\theta \mid D) \propto p(D \mid \theta)\,p(\theta)$; probability = degree of belief; condition on the observed $D$; needs a prior.

Trap: not "p-values vs posteriors". With much data and weak priors the numbers agree; the meanings do not.

Quick check: a colleague says "there is a 95% chance the true lift is between 1% and 3%". Which view does that sentence belong to, and what would the other view say instead?

It is a probability statement about the parameter, so it is Bayesian (a credible interval). A frequentist would say: "the interval from 1% to 3% was produced by a method that captures the true lift in 95% of repeated experiments"; for this particular interval, the lift is either inside or not.

Recap, cheat sheet and practice

  • A dataset is a table: rows = observations, columns = variables. The target (response, $y$) is what we explain; features (predictors, covariates, $x$) are the clues.
  • Variable types: numeric (continuous, discrete) and categorical (nominal, ordinal, binary). The type decides the sensible summaries, the encoding and later the likelihood.
  • The population is what we want to know about (often a process); the sample is what we see. Random sampling lets the sample stand in for the population; bias does not shrink with $n$.
  • A parameter (Greek, fixed, unknown) describes the population; a statistic (Latin or hat) is computed from the sample and changes from sample to sample.
  • Six jobs: describe, model, estimate, infer, predict, decide.
  • Probability: model → data. Statistics: data → model. One formula $p(D \mid \theta)$, read as the likelihood when $D$ is fixed.
  • The data-generating process: real process → random variables → distribution → observed data. Your data are one run of it.
  • Frequentist: θ fixed, data random, probability = long-run frequency. Bayesian: θ uncertain, condition on the observed data, probability = degree of belief.

Cheat sheet

TermPlain meaningExample
Observationone row, one thing measuredone user, one day
Target / response ($y$)what we want to explain or predictconverted, orders on day $t$
Feature / predictor / covariate ($x$)a clue used to explain $y$device, holiday flag
Continuous / discreteany value in a range / countable valuesminutes / orders per day
Nominal / ordinal / binarylabels without order / with order / two valuescountry / plan tier / yes-no
Population / sampleeverything we care about / what we observeall future users / 500 test users
Parameter / statisticdescribes the population / computed from the sample$p = 0.12$ / $\hat p = 57/500$
Probability vs statisticsmodel → data vs data → model$P(7 \text{ heads} \mid \theta = 0.5) = 0.117$ vs $\hat\theta = 0.7$
Likelihood $L(\theta)$$p(D \mid \theta)$ as a function of $\theta$not a distribution over $\theta$
DGPprocess → random variables → distribution → data$Y_i \sim \text{Bernoulli}(\theta)$
Frequentist / Bayesianθ fixed, data random / θ uncertain, data fixedconfidence interval / $P(\theta_B \gt \theta_A \mid D)$
Code it · Python
import numpy as np
import pandas as pd
from scipy import stats
from statsmodels.stats.proportion import proportion_confint

# 1) A dataset: rows = observations, columns = variables
df = pd.DataFrame({
    "device":    ["phone", "desktop", "phone", "tablet", "phone"],        # nominal
    "plan":      pd.Categorical(["basic", "pro", "basic", "enterprise", "pro"],
                                categories=["basic", "pro", "enterprise"],
                                ordered=True),                             # ordinal
    "pages":     [3, 7, 1, 12, 4],                                         # discrete
    "minutes":   [2.5, 9.1, 0.8, 15.3, 4.0],                               # continuous
    "converted": [0, 1, 0, 1, 0],                                          # binary target
})
X, y = df.drop(columns="converted"), df["converted"]   # features vs target
print(df.dtypes)                       # text, category (ordered), int64, float64, int64
print(y.mean())                        # 0.4 -> the mean of a 0/1 column is a proportion
X_model = pd.get_dummies(X, columns=["device"])        # one-hot encode the nominal feature
print(list(X_model.columns))  # ['plan', 'pages', 'minutes', 'device_desktop', 'device_phone', 'device_tablet']

# 2) Population -> sample -> statistic
rng = np.random.default_rng(0)
population = rng.lognormal(mean=1.6, sigma=0.6, size=2000)   # minutes on site
mu = population.mean()                                        # parameter: one fixed number
xbars = [rng.choice(population, size=25, replace=False).mean() for _ in range(1000)]
print(round(mu, 2), np.round(xbars[:3], 2))   # 5.82 [5.13 6.03 5.74]: the statistic changes
print(round(np.mean(xbars), 2), round(np.std(xbars), 2))     # 5.83 0.75: centred near mu

# 3) Data-generating process: same process, different datasets
t = np.arange(28)
weekend = (t % 7 >= 5).astype(float)
process_mean = 100 + 0.5 * t + 40 * weekend
for seed in [1, 2, 3]:
    y_obs = process_mean + np.random.default_rng(seed).normal(0, 10, size=28)  # normal(mean, SD)
    print(seed, round(y_obs.mean(), 1))         # 117.3, 118.5, 118.5: new data every run
print(round(process_mean.mean(), 2))            # 118.18, the process average over these days

# 4) Two directions and two schools with 7 heads in 10 flips
print(stats.binom.pmf(7, 10, 0.5))              # probability: model -> data, 0.1171875
print(stats.binom.sf(6, 10, 0.5))               # frequentist: P(K >= 7 | theta = 0.5) = 0.171875
print(proportion_confint(7, 10, method="wilson"))   # 95% confidence interval (0.397, 0.892)
post = stats.beta(1 + 7, 1 + 3)                 # Bayesian: flat prior Beta(1,1) -> Beta(8, 4)
print(post.mean(), post.sf(0.5))                # 0.667 and P(theta > 0.5 | data) = 0.887
print(post.ppf([0.025, 0.975]))                 # 95% credible interval [0.390 0.891]
Test yourself

1. Which of these is a statistic?

A statistic is computed from the sample. The experiment's observed rate is computed from the 2 000 users you saw; the other three describe the population and are parameters.

2. A column "plan tier" takes the values basic, pro and enterprise. Its type is…

The values are labels (categorical) with a natural order basic < pro < enterprise (ordinal). The gaps between tiers need not be equal, so it is not numeric.

3. "If the true conversion rate is 10%, how many of 200 users will convert?" is a question of…

The model is known (10%) and we ask about possible data. That is the probability direction. (About $200 \times 0.1 = 20$, give or take.)

4. In frequentist thinking, what is treated as random?

For a frequentist, θ is a fixed unknown constant; the randomness lives in the data (and in statistics computed from them) over repeated experiments. Priors belong to the Bayesian view.

5. Two random samples of 1 000 people: one from a country of 50 million, one from a city of 5 million. Which is true?

For random samples much smaller than the population, precision depends on $n$, not on the population's size. Both have $n = 1\,000$.

6. Which sentence fits the idea of a data-generating process?

The data are one random run of the process. The parameters belong to the process and do not change between runs; the noise is part of the model, and a good fit does not prove the model is the truth.

Practice problems

A. A forecasting table has the columns date, orders, is_holiday, temperature, promo_type (none / discount / bundle) and store_size (S / M / L). Give the role and type of each.

orders: the target, numeric discrete (a count). is_holiday: feature, binary. temperature: feature, numeric continuous. promo_type: feature, categorical nominal (one-hot encode it). store_size: feature, categorical ordinal (S < M < L). date: the time index; it is not used as a raw number but to build features such as the trend $t$, the day of the week and Fourier terms.

B. A population of 8 000 users has 640 converters. A random sample of 400 contains 28 converters. Name the parameter and the statistic, give both values, and explain the gap.

Parameter: the population conversion rate $p = 640/8\,000 = 0.08$. Statistic: the sample rate $\hat p = 28/400 = 0.07$. The gap $0.07 - 0.08 = -0.01$ is sampling variation: a different random 400 would give a different $\hat p$. In practice you would only know the 0.07.

C. Name the job of each sentence: (i) "Median basket last week: 3 items." (ii) "Demand = trend + weekly pattern + noise." (iii) "B's true lift over A is about 2 points." (iv) "B's lift is unlikely to be just noise." (v) "Next Monday: 140 to 170 orders." (vi) "Ship B."

(i) describe; (ii) model; (iii) estimate; (iv) infer; (v) predict; (vi) decide.

D. A coin is flipped 4 times. (a) If it is fair, what is P(exactly 3 heads)? (b) You saw 3 heads in 4. Which of θ = 0.25, 0.5, 0.75 explains this best, and by how much?

(a) There are $\binom{4}{3} = 4$ ways to place 3 heads, each with probability $0.5^4 = 1/16$, so $4/16 = 0.25$. (b) $P(3 \text{ heads} \mid \theta) = 4\theta^3(1-\theta)$. For 0.25: $4 \times 0.015625 \times 0.75 = 0.047$. For 0.5: $0.25$. For 0.75: $4 \times 0.421875 \times 0.25 = 0.422$. So $\theta = 0.75$ explains the data best: $0.422/0.25 \approx 1.7$ times better than 0.5 and about 9 times better than 0.25. (a) is probability; (b) is statistics.

E. Write a data-generating process, in the four steps, for "number of support tickets per day" with a weekly pattern.

Real process: customers run into problems and write in; fewer write at weekends. Random variable: $Y_t$ = tickets on day $t$, a count. Distribution: for example $Y_t \sim \text{Poisson}(\lambda_t)$ with $\lambda_t = a \cdot (1 + c \cdot \text{weekend}_t)$, where $a$ is the weekday level and $c$ the weekend change (negative); if the days vary more than a Poisson allows, a Negative Binomial (Chapter 4.8). Observed data: one history such as 41, 37, 52, 44, 39, 18, 15, … Rerunning the process would give different counts with the same $a$ and $c$.

F. Interview: "Explain frequentist vs Bayesian thinking without using the words p-value or posterior."

"They disagree about what probability means and what is random. A frequentist reads probability as a long-run frequency: the unknown quantity is a fixed number, the data are random, and a method is judged by how often it would be right if the experiment were repeated many times. A Bayesian reads probability as a degree of belief: the unknown quantity is uncertain, so it gets a probability distribution, which is updated by conditioning on the data actually observed. So a frequentist needs a fixed design and talks about procedures; a Bayesian needs a stated prior belief and can talk directly about the unknown quantity, for example 'the chance that B beats A'."

Chapter 4.2 · Syllabus Modules 1.1–1.2

Probability: outcomes, events and rules

Probability is the language of uncertainty, and it has a tiny grammar: list what can happen, group the outcomes into events, give each event a number between 0 and 1, and combine those numbers with three or four rules. With two dice, a few users and a checkout page, this chapter builds every rule from pictures you can click.

  • Describe any random situation as an experiment with outcomes and a sample space
  • Write events as sets of outcomes, and combine them with "or" (union), "and" (intersection) and "not" (complement)
  • State the three axioms of probability and compute probabilities by counting equally likely outcomes
  • Explain the two readings of probability: long-run frequency and degree of belief
  • Use the complement rule, the addition rule (inclusion–exclusion) and the multiplication rule for independent events
  • Compute "at least one" probabilities, such as the chance that at least one of 20 metrics shows a false alarm

Experiments, outcomes and the sample space core

Probability always starts with a situation whose result you do not know yet: flip a coin, roll a die, show the checkout page to a visitor, wait to see how many orders arrive tomorrow. Statisticians call such a situation an experiment, even when nobody wears a lab coat.

Each possible result is an outcome. The complete list of all possible outcomes is the sample space. Think of it as a menu that lists everything that could happen; each time the experiment runs, exactly one item on the menu is served.

Three ways to say it:

  • Picture: the sample space is the menu; an outcome is the one dish you actually get.
  • Numbers: one die has 6 outcomes; two dice have $6 \times 6 = 36$; one visitor has 2 (buy, no buy).
  • Slogan: before you ask "how likely?", list "what can happen?".

Roll a red die and a blue die.

  1. An outcome is the pair (red result, blue result), for example $(2, 5)$.
  2. The red die has 6 possible results. For each of them, the blue die has 6. So there are $6 \times 6 = 36$ outcomes.
  3. $(2, 5)$ and $(5, 2)$ are different outcomes: in the first the red die shows 2, in the second it shows 5.
  4. Other experiments: one visitor, $\Omega = \{\text{buy}, \text{no buy}\}$ (2 outcomes); orders tomorrow, $\Omega = \{0, 1, 2, \dots\}$ (no largest value); time on site, $\Omega = [0, \infty)$ (every number from 0 up).
  • An experiment (random experiment) is any process with a well-defined set of possible results, where we do not know in advance which result will occur.
  • An outcome $\omega$ ("omega") is one possible result.
  • The sample space $\Omega$ ("capital omega"; some books write $S$) is the set of all outcomes. It must be complete (some outcome always happens) and its outcomes mutually exclusive (exactly one happens each time).
  • Sample spaces can be finite (a die), countably infinite (counts $0, 1, 2, \dots$) or continuous (an interval of real numbers).
  • Counting principle: if one step has $m$ possible results and a second step has $n$ for each of them, the pair has $m \times n$ outcomes. $k$ coin flips give $2^k$ outcomes.
Why do we need it?

You cannot give chances to things you have not listed. Writing the sample space first prevents the classic mistakes: forgetting outcomes, counting one outcome twice, or treating outcomes as equally likely when they are not.

Where is it used?

Choosing a likelihood (its support, the set of values it allows, must match the sample space: {0, 1} for a conversion, {0, 1, 2, …} for counts, real numbers for revenue changes); defining the classes of a classifier; designing the possible results of an experiment.

How is it used?

For each random quantity, write down: what exactly is one outcome, and what values are possible. Then check that your model gives probability to all of them and to nothing outside them (for example, a Normal model for counts also allows negative counts).

flip a coinroll a dieone visitororders tomorrowtime on site H T 1 2 3 4 5 6 buy no buy 0 1 2 3 … 0∞ Ω = {H, T}Ω = {1, …, 6}Ω = {buy, no buy}Ω = {0, 1, 2, …}Ω = [0, ∞) finite: 2finite: 6finite: 2countably infinitecontinuous
Five experiments and their sample spaces. The first three are finite lists; orders can be any whole number (no largest value); time on site can be any number from 0 up.

Pick an experiment and look at its sample space. Watch how the counting principle multiplies: one coin has 2 outcomes, two coins $2 \times 2 = 4$, three coins $2 \times 2 \times 2 = 8$. Press Run the experiment several times: exactly one outcome lights up each time. For orders tomorrow the list never ends; for time on site the outcomes are not a list at all but every point of a line.

"With two dice, (2, 5) and (5, 2) are the same outcome, so there are 21 outcomes."

There are 21 unordered pairs, but they are not equally likely: "a 2 and a 5" can happen two ways, "two 6s" only one way. Use the 36 ordered pairs; they are equally likely for fair dice.

"The sample space is the outcomes I observed."

It is every outcome that could happen, including ones you have never seen (a day with 0 orders, a user who stays 3 hours).

"Two outcomes means a 50/50 chance."

A visitor either buys or not, yet the chance of buying may be 0.1. Listing outcomes says what can happen, not how likely each one is.

In an A/B framework like yours, each metric has its own sample space, and the likelihood must match it: a conversion lives in {0, 1} (Bernoulli/Binomial); a choice among $K$ options lives in {1, …, $K$} (Categorical/Multinomial); a count lives in {0, 1, 2, …} (Poisson). In your forecasting model, daily demand is usually a count (orders, units), so its natural sample space is {0, 1, 2, …}: a Negative Binomial likelihood respects that, while a Normal likelihood also gives some probability to negative and fractional demand. That can be an acceptable approximation for large counts, but you should know you are making it (Chapter 7.13).

Experiment = uncertain process; outcome $\omega$ = one result; sample space $\Omega$ = all outcomes (complete, mutually exclusive).

Counting principle: $m$ then $n$ choices → $m \times n$ outcomes; $k$ coins → $2^k$; two dice → 36.

Trap: ordered pairs (2,5) ≠ (5,2); "two outcomes" ≠ 50/50.

Quick check: how many outcomes does "flip a coin and roll a die" have? Is it finite?

For each of the 2 coin results there are 6 die results, so $2 \times 6 = 12$ outcomes, from (H, 1) to (T, 6). Yes, it is finite.

Events, and how to combine them: or, and, not core

An event is a yes/no question about the outcome: "is the sum 7?", "did the user buy?", "were there more than 200 orders?". Each question picks out a group of outcomes: the ones whose answer is "yes". We say the event happened if the actual outcome landed inside that group.

Questions can be combined with three little words. "A or B" (at least one of them happened), "A and B" (both happened), "not A" (A did not happen). In pictures: join the two groups, take their overlap, or take everything outside.

Three ways to say it:

  • Picture: an event is a region you colour on the map of all outcomes.
  • Numbers: "sum is 7" colours 6 of the 36 dice cells; "at least one 6" colours 11; both at once: 2.
  • Slogan: an event is a set of outcomes; or = union, and = intersection, not = complement.

Two dice, outcomes (red, blue). Let A = "the sum is 7" and B = "at least one die shows 6".

  1. A $= \{(1,6), (2,5), (3,4), (4,3), (5,2), (6,1)\}$: 6 outcomes.
  2. B: 6 outcomes have red = 6, 6 have blue = 6, and $(6,6)$ is in both lists, so $6 + 6 - 1 = 11$ outcomes.
  3. A and B (both): $\{(1,6), (6,1)\}$: 2 outcomes.
  4. A or B (at least one): $6 + 11 - 2 = 15$ outcomes (the 2 shared ones must not be counted twice).
  5. Not A: $36 - 6 = 30$ outcomes.
  6. Suppose you roll $(3, 4)$. Then A happened (sum 7), B did not (no 6), "A or B" happened, "A and B" did not.
  • An event is a subset of the sample space, $A \subseteq \Omega$. "A occurs" means the outcome $\omega$ is in $A$: $\omega \in A$.
  • Special events: a single outcome $\{\omega\}$ (an elementary event), the sure event $\Omega$, the impossible event $\emptyset$ (the empty set).
  • Union $A \cup B$ ("A or B"): outcomes in $A$, in $B$, or in both. In probability "or" is always inclusive.
  • Intersection $A \cap B$ ("A and B"): outcomes in both. Some books write $AB$ or $P(A, B)$.
  • Complement $A^c$ ("not A"; also written $A'$ or $\bar A$): outcomes not in $A$.
  • Difference $A \setminus B$ ("A but not B") $= A \cap B^c$.
  • $A$ and $B$ are mutually exclusive (disjoint) if they cannot happen together: $A \cap B = \emptyset$.
Why do we need it?

Real questions are combinations: "converted and on mobile", "late or cancelled", "not a bot". Writing them as sets turns fuzzy English into something you can count and compute, without double counting.

Where is it used?

Filtering a DataFrame (&, |, ~ are and, or, not), defining segments and metrics in A/B tests, SQL WHERE clauses, the events a forecast should price, such as "demand exceeds capacity", and the rules of this chapter.

How is it used?

Translate the question word by word: "or" → $\cup$, "and" → $\cap$, "not" → complement. Then count (equally likely outcomes) or add up probabilities. In pandas, df[(df.device == "phone") & (df.converted == 1)] is the event "phone and converted".

A ∪ B (or)A ∩ B (and)Aᶜ (not A)A ∖ B (A only)disjoint ABABABABAB in A, in B, or bothin botheverything outside Ain A, not in BA ∩ B = ∅ the box is Ω, all outcomes; the purple region is the event
Venn diagrams of the event operations. The rectangle is the whole sample space Ω; each circle is an event; the purple region is the combined event named above it.

Each cell is one of the 36 outcomes (red die across, blue die up; the number is the sum). Choose events A (blue tint) and B (orange tint), then pick what to show: A, B, A ∪ B, A ∩ B or Aᶜ; the purple cells are that event, and its probability is just "purple cells / 36". Check the readout for A ∪ B: the overlap is subtracted once. Choose custom for A and click cells to build your own event. Press Roll the dice to see which events happened.

"'A or B' means one of them but not both."

In probability "or" is inclusive: $A \cup B$ contains the outcomes where both happen. "Exactly one of them" is a different event: $(A \setminus B) \cup (B \setminus A)$.

"The opposite of 'both users convert' is 'neither converts'."

The complement of "both" is "not both": at least one of them does not convert. "Neither converts" is only one part of it. With two dice, "not both 6" has 35 outcomes, "neither is 6" only 25.

"An event is one outcome."

An event is a set of outcomes: it can hold one, many, all (the sure event) or none (the impossible event).

In an A/B framework like yours, metrics and segments are events. "Converted" is an event; a segment such as "mobile users in India" is an intersection; a combined success such as "converted or upgraded" is a union; a guardrail such as "no checkout error" is a complement. In your forecasting model, "demand exceeds capacity on day $t$" is the event $\{y_t \gt C\}$; its probability comes from the predictive distribution (Chapter 7.14).

Event = set of outcomes, $A \subseteq \Omega$; "A happened" = $\omega \in A$.

or = $\cup$ (inclusive), and = $\cap$, not = $A^c$, A only = $A \setminus B$; disjoint: $A \cap B = \emptyset$.

Trap: not(both) = at least one fails, not "neither".

Quick check: with two dice, how many outcomes are in "a double, and the sum is 4 or less"? And in "a double, or the sum is 4 or less"?

Doubles: (1,1), (2,2), …, (6,6): 6 outcomes. Sum ≤ 4: (1,1), (1,2), (2,1), (1,3), (2,2), (3,1): 6 outcomes. Both: (1,1) and (2,2): 2. Either: $6 + 6 - 2 = $ 10.

What a probability is: three rules (the axioms) core

Imagine you have exactly one bucket of sand, and you spread all of it over the outcomes on the menu. Some outcomes get a big pile, some a small pile, some none, but no outcome can get a negative amount, and all the sand must be used.

The probability of an event is then the amount of sand sitting on its outcomes. "Even number" on a die collects the sand on 2, 4 and 6. That picture already contains all three rules of probability: no negative sand, one bucket in total, and to get the sand on a group of separate outcomes, you add their piles.

Three ways to say it:

  • Picture: one bucket of sand spread over the outcomes; an event's probability is the sand on it.
  • Numbers: loaded die 0.1, 0.1, 0.1, 0.1, 0.1, 0.5 (total 1), so $P(\text{even}) = 0.1 + 0.1 + 0.5 = 0.7$.
  • Slogan: never negative, adds up to 1, and separate pieces add.

A loaded die gives the 6 half of all the probability: $P(1) = P(2) = P(3) = P(4) = P(5) = 0.1$ and $P(6) = 0.5$.

  1. No negative numbers: all six are at least 0. ✓
  2. Total: $5 \times 0.1 + 0.5 = 0.5 + 0.5 = 1$. ✓
  3. $P(\text{even}) = P(2) + P(4) + P(6) = 0.1 + 0.1 + 0.5 = 0.7$ (the three outcomes cannot happen together, so their probabilities add).
  4. $P(\text{not } 6) = 1 - 0.5 = 0.5$.
  5. For a fair die every face gets $1/6$, and $P(\text{even}) = 3/6 = 0.5$.
  6. Two fair dice: all 36 ordered pairs are equally likely, so $P(\text{sum is } 7) = 6/36 = 1/6$. When all outcomes are equally likely, a probability is just "count the outcomes in the event, divide by the total".

A probability $P$ gives every event $A$ a number $P(A)$ so that the three axioms (basic rules, stated by Kolmogorov in 1933) hold:

  1. Non-negative: $P(A) \ge 0$ for every event $A$.
  2. Normalized: $P(\Omega) = 1$ (something in the sample space always happens).
  3. Additive: if $A$ and $B$ are mutually exclusive ($A \cap B = \emptyset$), then $P(A \cup B) = P(A) + P(B)$. The same holds for any finite or countably infinite list of mutually exclusive events.

Everything else follows from these three. For example $P(\emptyset) = 0$, $0 \le P(A) \le 1$, $P(A^c) = 1 - P(A)$ (derived below), and if $A \subseteq B$ then $P(A) \le P(B)$.

For a finite or countable sample space, $P(A) = \sum_{\omega \in A} P(\{\omega\})$: add the probabilities of the outcomes in $A$.

Equally likely outcomes (the "classical" case: fair coins, fair dice, a random draw from a list): $P(A) = \dfrac{|A|}{|\Omega|}$, where $|A|$ is the number of outcomes in $A$. This shortcut is only valid when every outcome really is equally likely.

Why do we need it?

The axioms are the contract every probability must honour. Any model that breaks them (probabilities that are negative or do not add up to 1) gives impossible answers, and every rule in this chapter is derived from these three.

Where is it used?

Softmax outputs of a classifier (non-negative, sum to 1), the class probabilities of a Categorical distribution, a Dirichlet draw, any PMF table, normalizing a posterior so it integrates to 1, and checking that a hand-built probability table is valid.

How is it used?

Whenever you build probabilities by hand, check: all ≥ 0? total = 1? If the total is not 1, divide each number by the total (normalize). To get the probability of an event, add the probabilities of its outcomes.

Drag the top of each bar to set the probability of each face. The readout checks the axioms: no bar can go below 0, and the total must be exactly 1. Push one bar up and watch the total pass 1: then the numbers are no longer probabilities; press Make it add up to 1 to divide each bar by the total. Pick an event above the plot: its faces turn purple and its probability is the sum of the purple bars (axiom 3).

"$P(A) = |A| / |\Omega|$ always."

Only when every outcome is equally likely. A visitor's sample space {buy, no buy} has 2 outcomes, but $P(\text{buy})$ is not $1/2$. A loaded die has 6 faces, but $P(6)$ may be 0.5.

"A probability can be bigger than 1 if the event is very likely."

Never: $P(A) \le P(\Omega) = 1$. (A probability density can be bigger than 1, which is a different thing: Chapter 4.4.)

"Probability 0 means impossible."

In a finite model, yes: an outcome with probability 0 is one the model rules out. For a continuous one it does not: the chance that a visit lasts exactly 5.0000… minutes is 0, yet some visit lasts some exact time. Only ranges of values get positive probability (Chapter 4.4).

Every likelihood in your projects is a probability model that obeys these axioms. A Bernoulli puts $\theta$ on "convert" and $1 - \theta$ on "no convert". A Categorical metric with $K$ options needs probabilities $\pi_1, \dots, \pi_K \ge 0$ with $\sum_k \pi_k = 1$: exactly the kind of vector a Dirichlet distribution produces, which is why it is the natural prior in your Dirichlet-Multinomial model (Chapter 4.11). In neural networks, softmax enforces the same two rules.

Axioms: (1) $P(A) \ge 0$; (2) $P(\Omega) = 1$; (3) disjoint events add: $P(A \cup B) = P(A) + P(B)$.

$P(A) = \sum_{\omega \in A} P(\omega)$; equally likely outcomes: $P(A) = |A|/|\Omega|$.

Trap: counting rule only for equally likely outcomes; probability 0 ≠ impossible for continuous outcomes.

Quick check: someone proposes P(convert) = 0.12, P(add to cart only) = 0.30, P(leave) = 0.65 for three separate outcomes of a visit. Is this a valid probability model?

No. All are non-negative, but $0.12 + 0.30 + 0.65 = 1.07 \ne 1$, which breaks axiom 2. Either an outcome is double counted (perhaps some cart users also converted) or the numbers need fixing.

What does "probability 0.5" mean? Long-run frequency and degree of belief

"The chance of heads is 0.5." There are two common ways to read that sentence, and both are useful.

  • Frequency reading: if you flip the coin very many times, about half of the flips come up heads, and the share gets closer to 0.5 the longer you go.
  • Belief reading: 0.5 is how strongly you expect heads, a number you could use to bet fairly. This reading also works for one-off events that can never be repeated: "the chance this product launch succeeds", "the chance that variant B's true rate beats A's".

The two readings follow the same three axioms, so all the rules in this chapter work for both. You already met the schools of thought built on them in frequentist vs Bayesian thinking.

Three ways to say it:

  • Picture: frequency is a running share that settles down; belief is the fair price of a bet.
  • Numbers: 507 heads in 1 000 flips (frequency); paying at most 30 cents for a ticket that pays 1 dollar if it rains means $P(\text{rain}) = 0.3$ for you (belief).
  • Slogan: same rules, two meanings: "how often" or "how sure".
  1. Frequency. In one simulated run of a fair coin (the code at the end of this chapter), the share of heads was $5/10 = 0.5$ after 10 flips, $47/100 = 0.47$ after 100, $507/1\,000 = 0.507$ after 1 000 and $4\,953/10\,000 = 0.4953$ after 10 000. It wobbles, but less and less.
  2. Belief. A ticket pays 100 dollars if it rains tomorrow and nothing otherwise. If your probability of rain is $p$, the ticket pays $100 \times p$ dollars on average, so its fair price for you is $100p$.
  3. If the most you would pay is 30 dollars, then $100p = 30$, so your degree of belief is $p = 30/100 = 0.3$.
  4. A belief like "$P(\theta_B \gt \theta_A) = 0.93$" can only be read this way: the true rates are fixed numbers in the world, so there is no long run of "repeating the universe" to count over.
  • Frequency (frequentist) interpretation: $P(A)$ is the long-run relative frequency of $A$: if $n_A$ is the number of times $A$ happens in $N$ independent repetitions, then $n_A/N$ gets as close to $P(A)$ as you like as $N$ grows. (That this really happens is the Law of Large Numbers, Chapter 4.13.) It needs a repeatable experiment.
  • Belief (Bayesian, "subjective" or "epistemic") interpretation: $P(A)$ is a degree of belief in $A$ given the information you have, for example measured by the fair price of a bet that pays 1 if $A$ happens. Beliefs must still obey the axioms; if they do not, someone could sell you a set of bets that loses money whatever happens (the "Dutch book" argument).
  • Classical interpretation: with equally likely outcomes, $P(A) = |A|/|\Omega|$. It is a special case that both readings agree with.
Why do we need it?

To say what a reported number means. "90% of days fall inside the forecast band" is a frequency claim you can check; "93% probability that B beats A" is a belief given the data and the prior. Mixing them up leads to wrong interpretations in interviews and reports.

Where is it used?

Frequency: simulation (Monte Carlo), coverage and calibration checks of forecast intervals, false-positive rates of A/B tests. Belief: Bayesian priors and posteriors, $P(\theta_B \gt \theta_A \mid D)$, the probability outputs of a classifier read as confidence, weather forecasts.

How is it used?

For a repeatable event, estimate or check a probability by its frequency: simulate or count many trials. For a one-off or a parameter, state it as a belief and say what it is conditioned on (data, prior). Either way, apply the same rules.

frequency: "how often" belief: "how sure" 0.5 1number of flips → share of heads TICKET pays 100 dollars if it rains pays nothing otherwise fair price 30 dollars → P(rain) = 0.3
Two readings of the same number. Left: the share of heads wobbles early and settles near 0.5 as flips accumulate. Right: a degree of belief measured as the fair price of a bet.

Five independent runs of the same experiment (blue lines; the dark one is run 1). Each line is the share of trials so far in which the event happened. The trial axis is on a log scale so you can see the wild start. Press +10, +100 and +1000: the lines crowd towards the green true probability, inside the orange band that narrows like $1/\sqrt{N}$. Switch to a die shows 6 or a visitor converts and repeat. Press New sample for fresh runs.

"After five tails in a row, heads is due."

The coin has no memory: the next flip is still 50/50. The share settles near 0.5 not because streaks get "corrected", but because a few early flips are swamped by thousands of later ones. (This mistake is called the gambler's fallacy.)

"After 100 flips the share must be within 0.01 of 0.5."

The share still wobbles: after $N$ flips its typical distance from $p$ is about $\sqrt{p(1-p)/N}$, which is 0.05 at $N = 100$. Settling is slow; the wobble halves only when you quadruple $N$.

"A belief probability can be any number I like."

It must obey the axioms (otherwise your bets can be exploited), and it should change sensibly when data arrive. Bayes' theorem (Chapter 4.3) is the rule for that update.

Both readings appear in your work. In an A/B framework like yours, "$P(\theta_B \gt \theta_A \mid D) = 0.95$" is a belief about two fixed but unknown rates, given the data and the priors. In your forecasting model, a 90% prediction interval can be checked with the frequency reading: over many days, about 90% of the actual values should fall inside it. That check is called coverage or calibration (Chapter 7.16), and it is how you test whether the model's beliefs deserve to be trusted.

Frequency: $P(A) = $ the long-run share $n_A/N$ as $N \to \infty$ (repeatable experiments).

Belief: $P(A) = $ degree of belief given information (fair bet price); works for one-off events and parameters.

Same axioms, same rules. Traps: gambler's fallacy; the wobble shrinks only like $1/\sqrt{N}$.

Quick check: which reading fits each sentence? (a) "Our model's 80% intervals contained the actual demand on 79% of last year's days." (b) "There is a 70% chance the new pricing page beats the old one."

(a) is a frequency statement: it counts how often something happened over many repetitions (days). (b) is a belief statement about a one-off unknown (which page is truly better), as in a Bayesian A/B analysis.

The complement rule: P(not A) = 1 − P(A) core

Every time the experiment runs, either A happens or it does not, never both and never neither. So the probability of A and the probability of "not A" share the whole bucket of sand between them: together they make exactly 1.

This is more useful than it looks, because often "not A" is much easier to count. "The sum of two dice is at most 10" has 33 outcomes, but its opposite, "the sum is 11 or 12", has only 3.

Three ways to say it:

  • Picture: the sample space is split into A and everything else; the two pieces fill it exactly.
  • Numbers: P(convert) = 0.12, so P(not convert) = 0.88.
  • Slogan: if A is hard to count, count "not A" and subtract from 1.
  1. A visitor converts with probability 0.12. Then $P(\text{does not convert}) = 1 - 0.12 = 0.88$.
  2. Two dice, A = "the sum is at least 3". The only outcome not in A is $(1, 1)$, with sum 2.
  3. So $P(A^c) = 1/36$ and $P(A) = 1 - 1/36 = 35/36 \approx 0.972$. Counting A directly would mean listing 35 outcomes.
  4. A = "the sum is at most 10". Its complement "11 or 12" is $\{(5,6), (6,5), (6,6)\}$: 3 outcomes. So $P(A) = 1 - 3/36 = 33/36 \approx 0.917$.

For any event $A$:

$$P(A^c) = 1 - P(A).$$

Why it is true (from the axioms): $A$ and $A^c$ are mutually exclusive, and together they make up everything, $A \cup A^c = \Omega$. By axiom 3, $P(A) + P(A^c) = P(\Omega)$, and by axiom 2, $P(\Omega) = 1$. Subtract $P(A)$ from both sides.

Two consequences: $P(\emptyset) = 1 - P(\Omega) = 0$, and $P(A) = 1 - P(A^c) \le 1$ because $P(A^c) \ge 0$.

Why do we need it?

Many events are long lists of cases ("at least one", "anything except …") while their opposite is a single short case. The complement rule turns a long count into a short one, and it is the key to every "at least one" calculation.

Where is it used?

"At least one false positive" in multiple testing, "at least one failure" in reliability (servers, pipelines), survival functions $P(X \gt x) = 1 - P(X \le x)$ (SciPy's sf), one-sided p-values, and "probability of not running out of stock".

How is it used?

Ask: is "not A" simpler than A? If so, compute $P(\text{not } A)$ and subtract it from 1. With SciPy, dist.sf(x) gives $1 - $ dist.cdf(x) directly, and more accurately when the answer is tiny.

A Aᶜ = not A everything else P(A) = 0.12 P(Aᶜ) = 1 − 0.12 = 0.88 Ω: the whole bar has probability 1
A and "not A" never overlap and together cover all outcomes, so their probabilities add to 1 (drawn to scale for P(A) = 0.12).

Event A is "the sum of two dice is at most $t$" (blue cells); its complement is the orange cells. Move $t$ from 2 to 12. The purple outline marks whichever side is smaller: that is the side to count. Notice that for $t = 10$ you only need to count 3 orange cells to know that A has probability $1 - 3/36$. At $t = 12$, A is certain and its complement is empty.

"The complement of 'at least one' is 'at most one'."

The complement of "at least one" is "none". The complement of "all of them" is "at least one is not".

"$P(\text{not } A) = 1 - P(A)$ only works for equally likely outcomes."

It follows straight from the axioms, so it works for every probability model: loaded dice, conversion rates, continuous distributions.

In your projects the complement rule appears constantly. A Bernoulli likelihood gives $P(\text{no conversion}) = 1 - \theta$. If your forecasting model gives $P(\text{demand} \le C) = 0.97$ for capacity $C$, then the risk of exceeding capacity is $1 - 0.97 = 0.03$. And an A/B decision such as "probability that B is not better" is $1 - P(\theta_B \gt \theta_A \mid D)$ (ignoring the zero-probability case of an exact tie for continuous rates).

$P(A^c) = 1 - P(A)$, because $A$ and $A^c$ are disjoint and fill $\Omega$.

Use it when "not A" is shorter to count; $P(\emptyset) = 0$ and $P(A) \le 1$ follow.

Trap: not(at least one) = none; not(all) = at least one fails.

Quick check: three fair coins are flipped. What is P(at least one head)?

The complement "no heads" is only TTT: $P = (1/2)^3 = 1/8$. So $P(\text{at least one head}) = 1 - 1/8 = 7/8 = 0.875$. (Direct counting: 7 of the 8 equally likely outcomes.)

The addition rule: P(A or B) core

30% of users open the newsletter and 20% click an ad. What share did at least one of the two? Adding gives 50%, but that is too much: the users who did both were counted once in the 30% and once again in the 20%. Count them twice and you overshoot by exactly their share. So subtract the overlap once.

If the two events can never happen together (a die cannot show 1 and 2 at once), there is no overlap, and plain adding is right.

Three ways to say it:

  • Picture: two overlapping circles; adding their areas paints the overlap twice, so remove one coat.
  • Numbers: $0.30 + 0.20 - 0.08 = 0.42$.
  • Slogan: "or" = add, then subtract the double-counted "and".
  1. A = "opens the newsletter", $P(A) = 0.30$. B = "clicks an ad", $P(B) = 0.20$. Both: $P(A \cap B) = 0.08$.
  2. Only A: $0.30 - 0.08 = 0.22$. Only B: $0.20 - 0.08 = 0.12$.
  3. At least one: only A + only B + both $= 0.22 + 0.12 + 0.08 = 0.42$. Same as $0.30 + 0.20 - 0.08 = 0.42$.
  4. Neither (complement rule): $1 - 0.42 = 0.58$.
  5. Dice check: $P(\text{sum } 7 \text{ or a } 6) = 6/36 + 11/36 - 2/36 = 15/36 \approx 0.417$, matching the 15 cells you counted in the dice widget.
  6. Mutually exclusive: $P(\text{die shows 1 or 2}) = 1/6 + 1/6 = 2/6$, nothing to subtract.

Addition rule (inclusion–exclusion for two events): for any events $A$ and $B$,

$$P(A \cup B) = P(A) + P(B) - P(A \cap B).$$

Derivation from the axioms. Split $A \cup B$ into three pieces that cannot overlap: "A only" ($A \setminus B$), "both" ($A \cap B$), "B only" ($B \setminus A$). By axiom 3, $P(A \cup B) = P(A \setminus B) + P(A \cap B) + P(B \setminus A)$. Also $P(A) = P(A \setminus B) + P(A \cap B)$ and $P(B) = P(B \setminus A) + P(A \cap B)$. Adding these two counts $P(A \cap B)$ twice, so $P(A) + P(B) - P(A \cap B)$ is exactly the three pieces.

  • Mutually exclusive events ($A \cap B = \emptyset$): $P(A \cup B) = P(A) + P(B)$ (axiom 3).
  • Three events: $P(A \cup B \cup C) = P(A) + P(B) + P(C) - P(A \cap B) - P(A \cap C) - P(B \cap C) + P(A \cap B \cap C)$: add singles, subtract pairs, add back the triple.
  • Union bound (Boole's inequality): $P(A \cup B) \le P(A) + P(B)$, and in general $P(A_1 \cup \dots \cup A_m) \le \sum_i P(A_i)$. It is always true, and close to exact only when overlaps are tiny.
Why do we need it?

"Or" questions are everywhere (did the user do X or Y? did any alarm fire?), and naive adding double-counts the cases where several things happen together, sometimes giving "probabilities" above 1.

Where is it used?

Funnel and engagement metrics ("reached checkout by search or by email"), deduplicating audiences in marketing, the union bound behind the Bonferroni correction for multiple tests, and reliability ("server A or server B fails").

How is it used?

Find the three numbers $P(A)$, $P(B)$, $P(A \cap B)$ (from counts, or from a table), then add and subtract. In data, just count users who satisfy A | B directly; use the formula when you only have the separate summaries.

Ω: all users A only0.22 both0.08 B only0.12 P(A) = 0.30P(B) = 0.20 0.30 + 0.20counts "both"twice → − 0.08= 0.42
Adding P(A) and P(B) paints the purple overlap twice. Subtracting it once gives P(A or B) = 0.30 + 0.20 − 0.08 = 0.42. (Areas not to scale.)

The bars are a Venn diagram flattened onto a line, drawn to scale: the top bar is all of Ω (probability 1). Set P(A), P(B) and their overlap P(A ∩ B) with the sliders. The purple part of A and B is the overlap. The bottom row lays P(A) and P(B) end to end: it overshoots the true union (purple row above it) by exactly the overlap, marked in red. Try the presets: Disjoint needs no correction; B inside A gives P(A ∪ B) = P(A).

"$P(A \text{ or } B) = P(A) + P(B)$."

Only for mutually exclusive events. Otherwise subtract $P(A \cap B)$. A quick alarm: if your sum comes out above 1, you double counted.

"Segments in a report always add up to 100%."

Only if every user is in exactly one segment (a partition). Overlapping tags such as "mobile" and "returning" can add up to more than 100%.

In an A/B framework like yours, a combined success metric such as "converted or upgraded" has rate $P(C) + P(U) - P(C \cap U)$, not the sum. Segment shares only add to 1 when the segments form a partition, which is also what a hierarchical model's groups should be: each unit belongs to exactly one group (Chapter 6.5). And the union bound $P(\text{any false alarm}) \le \sum_i P(\text{alarm}_i)$ is the idea behind the Bonferroni correction (Chapter 5.11).

$P(A \cup B) = P(A) + P(B) - P(A \cap B)$; disjoint: just add.

Three events: add singles, subtract pairs, add back the triple. Union bound: $P(\cup A_i) \le \sum P(A_i)$.

Trap: a sum above 1 means double counting.

Quick check: 60% of users use the app, 50% use the website, and 25% use both. What share uses neither?

$P(\text{app or web}) = 0.60 + 0.50 - 0.25 = 0.85$. Neither: $1 - 0.85 = 0.15$, so 15% of users.

The multiplication rule: P(A and B) for independent events core

Two strangers visit your site. Each converts with probability 0.1, and what one does has no effect on the other. What is the chance that both convert? Among all pairs of visits, the first converts in 1 out of 10; and among those, the second converts in 1 out of 10 again. A tenth of a tenth: $0.1 \times 0.1 = 0.01$.

That works because the events are independent: knowing that one happened does not change the chance of the other. If they were linked (two visits from the same person, or picking users from a small group without putting them back), the second chance would change, and you would multiply by that changed chance instead.

Three ways to say it:

  • Picture: in a unit square, A takes a strip of width P(A), B a strip of height P(B); "both" is the corner rectangle of area P(A) × P(B).
  • Numbers: $0.1 \times 0.1 = 0.01$ (both), $0.9 \times 0.9 = 0.81$ (neither), $2 \times 0.1 \times 0.9 = 0.18$ (exactly one).
  • Slogan: "and" for independent events = multiply.
  1. Two independent visitors, each converting with 0.1. Both: $0.1 \times 0.1 = 0.01$.
  2. Neither: $(1 - 0.1) \times (1 - 0.1) = 0.9 \times 0.9 = 0.81$.
  3. Exactly one: "first yes, second no" or "first no, second yes": $0.1 \times 0.9 + 0.9 \times 0.1 = 0.09 + 0.09 = 0.18$.
  4. Check: $0.01 + 0.81 + 0.18 = 1$. The four branches of the tree cover everything.
  5. Dice: $P(\text{red } 6 \text{ and blue } 6) = 1/6 \times 1/6 = 1/36$, the same as counting the single cell $(6, 6)$.
  6. Not independent: 10 users, 3 of whom converted. Pick 2 different users at random. The first is a converter with probability $3/10$. If so, only 2 converters are left among 9 users, so the second is a converter with probability $2/9$. $P(\text{both}) = 3/10 \times 2/9 = 6/90 \approx 0.067$, not $0.3 \times 0.3 = 0.09$.
  • Events $A$ and $B$ are independent if knowing that one happened does not change the probability of the other. The formal definition is the multiplication rule itself: $$P(A \cap B) = P(A)\,P(B).$$
  • For $m$ independent events (each one unaffected by any combination of the others): $P(A_1 \cap A_2 \cap \dots \cap A_m) = P(A_1)\,P(A_2) \cdots P(A_m)$.
  • General multiplication rule (always true): $P(A \cap B) = P(A)\,P(B \mid A)$, where $P(B \mid A)$, "the probability of B given A", is the chance of B once we know A happened. Independence is exactly the case $P(B \mid A) = P(B)$. Conditional probability is the topic of Chapter 4.3.
  • A probability tree draws a sequence of steps as branches; the probability of a path is the product of the probabilities along it, and the paths' probabilities add to 1.
Why do we need it?

Most data are many observations together, and the probability of the whole dataset is built by multiplying the probabilities of the pieces. Without the product rule there is no likelihood, and without spotting dependence you get wildly overconfident answers.

Where is it used?

Every likelihood of independent observations, $p(D \mid \theta) = \prod_i p(y_i \mid \theta)$ (Binomial, Poisson, Normal regressions); naive Bayes classifiers; reliability of systems in series; the "at least one" formula in the next section.

How is it used?

First ask whether the events are really independent (different users? separate days? sampling with replacement?). If yes, multiply. If not, multiply by the conditional probability instead. In code, multiply many small probabilities by adding their logarithms to avoid underflow.

converts 0.1does not 0.9 visitor 1 0.10.90.10.9 visitor 2 both: 0.1 × 0.1 = 0.01 first only: 0.1 × 0.9 = 0.09 second only: 0.9 × 0.1 = 0.09 neither: 0.9 × 0.9 = 0.81 total = 1
A probability tree for two independent visitors. Multiply along each path; the four paths cover every possibility, so their probabilities add to 1.

The square is the whole sample space (area 1). Event A is the blue vertical strip, of width P(A); event B is the orange horizontal strip, of height P(B). When they are independent, "A and B" is the purple corner rectangle, with area P(A) × P(B). Move the sliders. Then press Throw 1 000 random points: the share landing in both is close to the product, and among the points inside A, the share that is also in B is close to P(B). That second fact is what "independent" means.

Each level of the tree is one visitor: C = converts, N = does not. The number on a branch is its probability; a path's probability is the product along it (shown at the leaf). Pick an event: the matching leaves turn purple and the readout adds them. Compare at least one with none: they always add up to 1. Switch to 3 visitors and watch the tree double.

"Multiply whenever you see the word 'and'."

Multiply the plain probabilities only for independent events. Otherwise use $P(A)\,P(B \mid A)$: picking 2 converters out of 10 users is $3/10 \times 2/9$, not $3/10 \times 3/10$.

"Independent means the circles in the Venn diagram do not overlap."

Non-overlapping circles are mutually exclusive events, which are strongly dependent (see below). Independence is about sizes: the overlap's area equals the product of the two areas, as in the corner rectangle.

"Two visits are independent because they are two rows in the table."

Two visits by the same user, or by users in the same household, are often linked. Independence is a property of the process, not of the table layout.

"Mutually exclusive events are independent."

If $A$ and $B$ are mutually exclusive and both have positive probability, then $P(A \cap B) = 0$, but $P(A)P(B) \gt 0$. So they are dependent: learning that A happened tells you for certain that B did not.

Model answer: "Mutually exclusive means they cannot happen together. Independent means one tells you nothing about the other. A user choosing the basic plan and the same user choosing the pro plan are mutually exclusive, and therefore highly dependent: if I know they chose basic, I know they did not choose pro."

Your likelihoods are products. In an A/B framework like yours, the Binomial likelihood for variant A multiplies one Bernoulli term per user: $\prod_i \theta_A^{y_i}(1-\theta_A)^{1-y_i}$. That product assumes users are independent once the true rate is fixed (a kind of independence you will meet in Chapter 4.3). If one user can appear many times, or users influence each other, the product treats the data as more informative than it is and the posterior becomes too narrow (Chapter 5.10, 5.11). Your forecasting likelihood also multiplies days, assuming the noise on different days is independent given the mean; autocorrelated residuals are the warning sign (Chapter 7.17). In code, these products are computed as sums of log-probabilities.

Independent: $P(A \cap B) = P(A)P(B)$ (knowing A does not change B's chance).

Always: $P(A \cap B) = P(A)\,P(B \mid A)$. Tree: multiply along a path, add paths.

Traps: mutually exclusive ≠ independent (they are dependent); sampling without replacement is dependent.

Quick check: a fair coin is flipped and a fair die is rolled. What is P(heads and a number greater than 4)?

The coin and die do not affect each other, so multiply: $P(\text{heads}) \times P(5 \text{ or } 6) = 1/2 \times 2/6 = 2/12 = 1/6$. Counting check: 2 of the 12 outcomes, (H, 5) and (H, 6).

"At least one": the complement trick core

"At least one" hides many cases: exactly one, exactly two, exactly three, and so on. Its opposite has only one case: none at all. So compute the chance of "none" (for independent tries, multiply the chances of failing each time) and subtract it from 1.

The result surprises almost everyone: small chances pile up fast when you try many times. One metric has a 5% chance of a false alarm; look at 20 metrics and a false alarm somewhere becomes more likely than not.

Three ways to say it:

  • Picture: in the tree, "at least one" is every leaf except the single all-fail leaf.
  • Numbers: $1 - 0.95^{20} = 1 - 0.358 = 0.642$.
  • Slogan: at least one $= 1 -$ none.
  1. At least one 6 in 4 rolls. No 6 on one roll: $5/6$. No 6 on four independent rolls: $(5/6)^4 = 625/1296 \approx 0.482$. So $P(\text{at least one 6}) = 1 - 0.482 = 0.518$.
  2. 20 metrics. Nothing really changed, but each metric independently has a 5% chance to look "significant" by accident. No false alarm on one metric: $0.95$. On all 20: $0.95^{20} \approx 0.358$. At least one false alarm: $1 - 0.358 = 0.642$.
  3. A converter among 10 visitors at $p = 0.1$: $1 - 0.9^{10} = 1 - 0.349 = 0.651$.
  4. The wrong shortcut: $20 \times 0.05 = 1.00$ would claim a false alarm is certain. Adding the chances counts the overlaps (two or more alarms at once) several times; it is only an upper bound.

For any events $A_1, \dots, A_m$:

$$P(\text{at least one } A_i) = 1 - P(\text{none}) = 1 - P(A_1^c \cap A_2^c \cap \dots \cap A_m^c).$$

If the events are independent and each has probability $p$:

$$P(\text{at least one}) = 1 - (1 - p)^m.$$
  • Upper bound (always true, by the union bound): $P(\text{at least one}) \le m\,p$. It is a good approximation only when $m\,p$ is small (say below 0.1).
  • How many tries for a 50% chance? Solve $(1-p)^m = 0.5$: $m = \ln 0.5 / \ln(1 - p) \approx 0.69/p$ for small $p$.
  • Without independence, $P(\text{at least one})$ still lies between the largest single $P(A_i)$ and $\min(1, \sum_i P(A_i))$.
Why do we need it?

Whenever you look many times (many metrics, segments, variants, days, servers), rare events stop being rare. This formula tells you how quickly "unlikely" becomes "expected", which is the root of the multiple-testing problem.

Where is it used?

Family-wise error rates in A/B testing and the Bonferroni/Šidák corrections; the chance a forecast is breached on at least one day; reliability ("at least one of 100 servers fails"); the birthday problem; data-quality checks across many columns.

How is it used?

Write the event as "at least one", compute "none" as a product (if the tries are independent), and subtract from 1. To keep the overall false-alarm rate at 5% across $m$ independent checks, each check needs $1 - (1 - 0.05)^{1/m}$ (Šidák), roughly $0.05/m$ (Bonferroni).

Set the chance per try $p$ and the number of tries $m$. The purple curve is the correct $1 - (1-p)^m$ for every $m$ from 1 to 100; the red dashed line is the tempting but wrong $m \times p$, which shoots past 1. Try the presets. Notice how close the two are for small $m$, and how far apart they are by the time the curve passes 50%.

Each row is one A/B experiment in which nothing really changed; each square is one metric. A metric gives a false alarm (red) with probability α, independently. A row with at least one red square (marked "alarm") is an experiment where someone could claim "a significant effect". Press Run 25 more experiments several times and compare the running share of red rows with the formula. Then lower the number of metrics to 1: now the share drops to about α.

"20 metrics at 5% each means $20 \times 5\% = 100\%$: a false alarm is certain."

Adding counts the overlaps many times. The right answer for independent metrics is $1 - 0.95^{20} \approx 64\%$. $m \times p$ is only an upper bound.

"The 64% is exact for my 20 metrics."

It assumes the metrics are independent. Real metrics are related (revenue and conversions move together), so the true chance differs; for metrics that tend to move together it is usually lower. It is never below 5% (one metric alone) and never above $20 \times 5\%$.

"Each metric is tested at 5%, so my experiment's error rate is 5%."

5% is the rate per metric. The chance of at least one false alarm in the whole experiment is much larger. That is why multiple-testing corrections exist (Chapter 5.11).

"We tested 20 metrics at α = 0.05, and one came out significant, so it is a real effect."

"With 20 independent metrics and no real effect, the chance of at least one false positive is $1 - 0.95^{20} \approx 0.64$. One significant metric out of 20 is what pure noise usually produces."

Model answer: "I would decide the primary metric in advance, treat the rest as secondary, and control the family-wise error rate (the chance of any false alarm at all; Bonferroni or Holm) or the false discovery rate (the share of the alarms that are false; Benjamini–Hochberg) across them."

In an A/B framework like yours, many metrics, variants and segments create many chances for a lucky-looking result. A Bayesian rule such as "ship if $P(\theta_B \gt \theta_A \mid D) \gt 0.95$" is not immune: applied separately to 20 segments with no real difference, it will flag some segment far more often than it flags any single segment. Partial pooling in a hierarchical model reduces this by shrinking noisy segment estimates towards each other (Chapter 6.6). For your forecasts, if each day independently had a 2% chance of exceeding capacity, the chance of at least one breach in a 30-day month would be $1 - 0.98^{30} \approx 0.45$; real days are correlated, so check it by simulating whole future paths.

$P(\text{at least one}) = 1 - P(\text{none})$; independent with chance $p$ each: $1 - (1-p)^m$.

20 metrics at 5%: $1 - 0.95^{20} \approx 0.64$. 50% needs $m \approx 0.69/p$ tries.

Trap: $m \times p$ is only an upper bound; real metrics are not independent.

Quick check: a data pipeline has 12 independent steps, each failing on a given night with probability 0.01. What is the chance that at least one step fails tonight?

No failures: $0.99^{12} \approx 0.886$. At least one: $1 - 0.886 = 0.114$, about 11%. The shortcut $12 \times 0.01 = 0.12$ is close here because $m \times p$ is small.

Recap, cheat sheet and practice

  • An experiment has outcomes; the set of all outcomes is the sample space $\Omega$. Two dice: $6 \times 6 = 36$ ordered outcomes.
  • An event is a set of outcomes. "Or" = union $\cup$ (inclusive), "and" = intersection $\cap$, "not" = complement $A^c$. Mutually exclusive: $A \cap B = \emptyset$.
  • The axioms: $P(A) \ge 0$, $P(\Omega) = 1$, and probabilities of mutually exclusive events add. With equally likely outcomes, $P(A) = |A|/|\Omega|$.
  • Two readings of the same rules: long-run frequency ("how often") and degree of belief ("how sure").
  • Complement rule: $P(A^c) = 1 - P(A)$. Addition rule: $P(A \cup B) = P(A) + P(B) - P(A \cap B)$.
  • Multiplication rule: for independent events $P(A \cap B) = P(A)P(B)$; always $P(A \cap B) = P(A)P(B \mid A)$. Mutually exclusive events are dependent.
  • At least one $= 1 - $ none $= 1 - (1-p)^m$ for independent tries: 20 metrics at 5% give about 64%.

Cheat sheet

RuleFormulaWhen
Counting (classical)$P(A) = |A| / |\Omega|$only for equally likely outcomes
Sum over outcomes$P(A) = \sum_{\omega \in A} P(\omega)$finite or countable $\Omega$
Complement$P(A^c) = 1 - P(A)$always
Addition$P(A \cup B) = P(A) + P(B) - P(A \cap B)$always (drop the last term if disjoint)
Union bound$P(\cup_i A_i) \le \sum_i P(A_i)$always; close when overlaps are small
Multiplication$P(A \cap B) = P(A)\,P(B)$independent events only
General multiplication$P(A \cap B) = P(A)\,P(B \mid A)$always (Chapter 4.3)
At least one$1 - (1-p)^m$$m$ independent tries, chance $p$ each
Code it · Python
import itertools
import numpy as np
from fractions import Fraction

# 1) Sample space of two dice: all ordered pairs (red, blue)
omega = list(itertools.product(range(1, 7), repeat=2))
print(len(omega))                                  # 36

# 2) Events are subsets; with equally likely outcomes P(A) = |A| / |Omega|
A = {w for w in omega if sum(w) == 7}              # sum is 7
B = {w for w in omega if 6 in w}                   # at least one 6
P = lambda E: Fraction(len(E), len(omega))
print(P(A), P(B), P(A & B), P(A | B))              # 1/6 11/36 1/18 5/12
print(P(A | B) == P(A) + P(B) - P(A & B))          # True: the addition rule
print(1 - P(A), P(set(omega) - A))                 # 5/6 5/6: the complement rule

# 3) Long-run frequency: the running share of heads settles near 0.5
rng = np.random.default_rng(1)
flips = rng.random(10_000) < 0.5
running = np.cumsum(flips) / np.arange(1, 10_001)
print(running[[9, 99, 999, 9999]])                 # [0.5 0.47 0.507 0.4953]: wobbly, then close

# 4) Multiplication rule for independent events
print(Fraction(1, 6) * Fraction(1, 6))             # 1/36: red 6 and blue 6
print(0.1 * 0.1, 0.9 * 0.9, 2 * 0.1 * 0.9)         # 0.01 0.81 0.18: both, neither, exactly one
print(3 / 10 * 2 / 9)                              # 0.0667: two converters drawn WITHOUT replacement

# 5) "At least one" through the complement
print(1 - (5 / 6) ** 4)                            # 0.518: at least one 6 in 4 rolls
print(1 - 0.95 ** 20)                              # 0.642: at least one false alarm among 20 metrics
sims = rng.random((100_000, 20)) < 0.05            # 100 000 simulated experiments x 20 metrics
print(sims.any(axis=1).mean())                     # 0.646 in this run (about 0.64)
Test yourself

1. Two fair dice are rolled. What is P(the sum is 8)?

Sum 8 comes from (2,6), (3,5), (4,4), (5,3), (6,2): 5 of the 36 equally likely ordered outcomes. (1/11 comes from wrongly treating the 11 possible sums as equally likely.)

2. $P(A) = 0.5$, $P(B) = 0.4$ and $P(A \cap B) = 0.2$. What is $P(A \cup B)$?

Addition rule: $0.5 + 0.4 - 0.2 = 0.7$. Plain adding (0.9) counts the overlap twice.

3. A and B are mutually exclusive, and both have positive probability. Which is true?

$P(A \cap B) = 0$ but $P(A)P(B) \gt 0$, so the independence equation fails. Knowing A happened changes the chance of B to 0.

4. A fair coin is flipped 5 times. What is P(at least one head)?

The complement "no heads" is only TTTTT, with probability $(1/2)^5 = 1/32$. So $1 - 1/32 = 31/32$. (5/2 is the shortcut $m \times p$, which is not even a probability.)

5. Which of these is not a valid probability model for a coin?

$0.7 + 0.4 = 1.1 \ne 1$, which breaks the axiom $P(\Omega) = 1$. A coin that always lands heads (the second option) is unusual but perfectly valid.

6. "The share of heads settles near 0.5 because, after a run of tails, heads becomes more likely." This sentence is…

Independent flips stay 50/50 whatever happened before (believing otherwise is the gambler's fallacy). The share settles because a fixed early excess becomes a smaller and smaller fraction of a growing total.

Practice problems

A. Two fair dice. What is the probability that the two numbers differ by at most 1?

Difference 0 (doubles): 6 outcomes. Difference 1: (1,2), (2,1), (2,3), (3,2), (3,4), (4,3), (4,5), (5,4), (5,6), (6,5): 10 outcomes. Total $6 + 10 = 16$, so $P = 16/36 = 4/9 \approx 0.444$.

B. 40% of email recipients opened the email and 15% clicked a link; you can only click after opening. Find P(opened or clicked) and P(opened but did not click).

Every clicker opened, so "clicked" is inside "opened" and $P(\text{opened} \cap \text{clicked}) = 0.15$. Addition rule: $0.40 + 0.15 - 0.15 = 0.40$: the union is just "opened". Opened but did not click: $0.40 - 0.15 = 0.25$.

C. Three independent servers are each down on a given day with probability 0.02. Find P(all up), P(at least one down) and P(all down).

All up: $0.98^3 = 0.941192 \approx 0.941$. At least one down (complement): $1 - 0.941 = 0.059$. All down: $0.02^3 = 0.000008$. Note how much more likely "at least one" is than "all".

D. A test set has 5 accounts, 2 of which are bots. You pick 2 different accounts at random. Find P(both bots), P(no bots) and P(exactly one bot).

Picks are without replacement, so the second chance depends on the first. Both bots: $2/5 \times 1/4 = 2/20 = 0.1$. No bots: $3/5 \times 2/4 = 6/20 = 0.3$. Exactly one: the rest, $1 - 0.1 - 0.3 = 0.6$ (check: $2/5 \times 3/4 + 3/5 \times 2/4 = 6/20 + 6/20 = 0.6$).

E. Interview: "Your PM checked an A/B test separately in 10 user segments and found one significant at 5%. Should we ship for that segment?"

"Not on that evidence alone. If nothing really differs and the 10 segment checks were independent, the chance that at least one shows a false alarm is $1 - 0.95^{10} \approx 0.40$, so finding one is close to a coin flip even with no real effect. I would treat it as a hypothesis: correct for the 10 checks (Holm or Bonferroni), use a hierarchical model that shrinks segment estimates, or confirm it in a new experiment aimed at that segment."

F. Use the axioms to show: if every outcome of A is also in B ($A \subseteq B$), then $P(A) \le P(B)$.

Split $B$ into two pieces that cannot overlap: $A$ and "B but not A", $B \setminus A$. By axiom 3, $P(B) = P(A) + P(B \setminus A)$. By axiom 1, $P(B \setminus A) \ge 0$. So $P(B) \ge P(A)$. For example, "sum is 7" is inside "sum is at least 7", so its probability (6/36) cannot be larger (21/36).

Chapter 4.3 · Syllabus Modules 1.3–1.6

Conditional probability, independence, total probability and Bayes' theorem

New information changes probabilities. "What is the chance this user converts?" has one answer; "what is the chance this user converts, given they are on mobile?" has another. This chapter teaches the four rules for updating probabilities when you learn something, and the one rule (Bayes) that turns "how likely is the evidence if X is true?" into "how likely is X now that I have seen the evidence?". It is the probability engine under your whole A/B framework.

  • Compute a conditional probability $P(A\mid B)$ by "shrinking the world" to $B$, and see why $P(A\mid B) \ne P(B\mid A)$
  • Use the multiplication rule $P(A\cap B) = P(A\mid B)\,P(B)$ and draw probability trees and funnels
  • Know exactly what independence means, and tell apart independence, correlation and conditional independence
  • Split a hard probability into easy cases with the law of total probability
  • Use Bayes' theorem with natural frequencies, avoid the base-rate fallacy, and apply it to "which variant did this converter come from?"

Conditional probability: shrinking the world core

You roll two dice behind a screen. Before you look, the chance that the sum is at least 10 is small: 6 of the 36 equally likely outcomes. Now a friend peeks and says: "the first die is a 6." Suddenly a sum of 10 or more looks much more likely. Nothing about the dice changed. What changed is what you know.

The trick is simple. Throw away every outcome that cannot have happened any more (all the outcomes where the first die is not 6). Only 6 outcomes are left. That smaller set is your new world. Inside it, count how often your event happens: 3 of the 6 (6+4, 6+5, 6+6). So the chance is now $3/6 = 1/2$.

That is a conditional probability: the probability of one event given that another event is known to have happened. (An event is a set of outcomes, like "the sum is at least 10"; see Chapter 4.2.)

Three ways to say it:

  • Picture: cross out everything that did not happen, then measure your event inside what is left.
  • Numbers: P(sum ≥ 10) = 6/36 ≈ 0.17, but P(sum ≥ 10 given first die = 6) = 3/6 = 0.5.
  • Slogan: "given B" means "B is the new 100%".

Dice. Let $A$ = "sum ≥ 10" and $B$ = "first die = 6".

  1. All outcomes: $6 \times 6 = 36$ equally likely pairs.
  2. $A$ contains (4,6), (5,5), (6,4), (5,6), (6,5), (6,6): 6 outcomes, so $P(A) = 6/36 = 1/6 \approx 0.167$.
  3. $B$ contains (6,1), (6,2), …, (6,6): 6 outcomes. This is the new world.
  4. Outcomes in both $A$ and $B$ (written $A \cap B$, "A and B"): (6,4), (6,5), (6,6): 3 outcomes.
  5. $P(A \mid B) = \dfrac{3}{6} = 0.5$. Knowing $B$ tripled the chance of $A$.

Your world: 1,000 visitors. 400 are on mobile and 600 on desktop. 20 mobile visitors convert and 60 desktop visitors convert (80 conversions in total).

  1. Overall: $P(\text{convert}) = 80/1000 = 0.08$.
  2. Given mobile, the world is the 400 mobile visitors: $P(\text{convert} \mid \text{mobile}) = 20/400 = 0.05$.
  3. Given converted, the world is the 80 converters: $P(\text{mobile} \mid \text{convert}) = 20/80 = 0.25$.
  4. Same 20 people on top, different worlds underneath, so $0.05 \ne 0.25$. The order inside $P(\cdot \mid \cdot)$ matters.

For two events $A$ and $B$ with $P(B) \gt 0$, the conditional probability of $A$ given $B$ is

$$P(A \mid B) = \frac{P(A \cap B)}{P(B)}.$$
  • The bar "$\mid$" is read "given". $B$ is the conditioning event: the thing we know happened.
  • $A \cap B$ ("A intersect B", "A and B") is the set of outcomes in both events.
  • Dividing by $P(B)$ rescales the new world so that it has total probability 1. With equally likely outcomes this is just counting: $P(A\mid B) = \dfrac{\#(A\cap B)}{\#B}$.
  • $P(\cdot \mid B)$ is a full probability rule on the new world: for example $P(A \mid B) + P(A^c \mid B) = 1$, where $A^c$ ("A complement") means "not A".
  • If $P(B) = 0$ the formula divides by zero and $P(A \mid B)$ is not defined by it. (Conditioning on an exact value of a continuous variable needs densities, Chapter 4.4, and conditional distributions, Chapter 4.6.)
Why do we need it?

Almost every useful question carries a "given": the conversion rate given the variant, the demand given it is a holiday, the chance of fraud given a flag. Without conditional probability we could only talk about everyone at once.

Where is it used?

Conversion rate per variant in A/B tests ($\theta_A = P(\text{convert}\mid\text{saw A})$), segment metrics, classifiers that output $P(y=1\mid x)$, every likelihood $p(\text{data}\mid\theta)$, and every posterior $p(\theta\mid\text{data})$ in your NumPyro models.

How is it used?

Filter the data to the rows where the condition holds, then compute the share you care about. In pandas: df[df.device == "mobile"].converted.mean(), or pd.crosstab(df.device, df.converted, normalize="index") for every group at once.

All outcomes (the whole world) A B A∩B given B B is the new 100% P(A | B) = purple share of B everything outside B is crossed out
Conditioning on B throws away every outcome outside B and treats B as the whole world. P(A | B) is the share of B that is also in A.

Each square is one of the 36 equally likely outcomes (row = first die, column = second die). Pick an event $A$ and an event $B$. With Shrink the world to B on, every square outside $B$ fades: you are counting only inside $B$. Try $A$ = "sum ≥ 10", $B$ = "first die = 6", then swap to $B$ = "at least one 6" and compare $P(A\mid B)$ with $P(B\mid A)$: they are different numbers.

The square is all visitors. Column widths are the shares of mobile (blue) and desktop (orange) traffic; the dark part at the bottom of each column is the share of that column that converts. So area = probability. Choose Given mobile: the world is the blue column. Choose Given converted: the world is only the dark areas. Make mobile traffic large and its conversion rate small, and watch $P(\text{mobile}\mid\text{convert})$ and $P(\text{convert}\mid\text{mobile})$ move in different directions.

"$P(A \mid B)$ and $P(B \mid A)$ are basically the same thing."

They share the same top number $P(A\cap B)$ but divide by different worlds: $P(B)$ versus $P(A)$. $P(\text{mobile}\mid\text{convert}) = 0.25$ while $P(\text{convert}\mid\text{mobile}) = 0.05$. Swapping them is the most common probability mistake there is (it has a name: confusion of the inverse).

"$P(A \mid B)$ means $B$ causes $A$."

"Given" only means "restricted to the cases where $B$ happened". Mobile visitors converting less does not prove that the phone causes it; mobile users may just be different people (causal questions come in Chapter 5.12).

"$P(A\mid B)$ is the share of A that is in B."

It is the share of B that is in A. The thing after the bar is always the world you divide by.

In your A/B framework, the conversion rates are conditional probabilities: $\theta_A = P(\text{convert}\mid\text{assigned to A})$ and $\theta_B = P(\text{convert}\mid\text{assigned to B})$. Each Beta-Binomial model estimates one of them from the users in that arm only, which is exactly "shrinking the world" to that arm. Segment-level metrics are conditioned one step further: $P(\text{convert}\mid\text{variant B},\ \text{mobile})$.

$P(A\mid B) = \dfrac{P(A\cap B)}{P(B)}$, needs $P(B)\gt 0$. With equal outcomes: $\#(A\cap B)/\#B$.

Meaning: B becomes the new 100%; measure A inside it.

Trap: $P(A\mid B) \ne P(B\mid A)$ in general (0.05 vs 0.25 in the visitor example).

Quick check: on the dice grid, what is $P(\text{doubles} \mid \text{sum} = 8)$?

Sum = 8 has 5 outcomes: (2,6), (3,5), (4,4), (5,3), (6,2). Only (4,4) is a double. So $P = 1/5 = 0.2$, compared with $P(\text{doubles}) = 6/36 \approx 0.167$ before you knew the sum.

The multiplication rule, probability trees and funnels core

Think of a checkout funnel. Out of everyone who visits, some add an item to the cart. Out of those, some reach checkout. Out of those, some pay. To get the share who pay, you do not need a new formula: you just keep taking "a fraction of a fraction". 40% add to cart, half of those reach checkout, 60% of those pay: $0.4 \times 0.5 \times 0.6 = 0.12$.

Each step's fraction is a conditional probability ("pays, given reached checkout"). Multiplying them along the path gives the chance of the whole path. That is the multiplication rule. You met its simplest form in Chapter 4.2; here is the general version that works even when the steps depend on each other.

Three ways to say it:

  • Picture: walk down a tree; multiply the numbers written on the branches you walk along.
  • Numbers: P(mobile and converts) = P(mobile) × P(converts | mobile) = 0.4 × 0.05 = 0.02.
  • Slogan: "and" = first thing × (second thing given the first).

Funnel. 1,000 visitors. $P(\text{cart}) = 0.4$, $P(\text{checkout}\mid\text{cart}) = 0.5$, $P(\text{pay}\mid\text{checkout}) = 0.6$.

  1. Cart: $1000 \times 0.4 = 400$ visitors.
  2. Checkout: $400 \times 0.5 = 200$ visitors.
  3. Pay: $200 \times 0.6 = 120$ visitors.
  4. So $P(\text{cart} \cap \text{checkout} \cap \text{pay}) = 0.4 \times 0.5 \times 0.6 = 0.12$, and indeed $120/1000 = 0.12$.

Cards (steps that depend on each other). Draw two cards without putting the first back. What is the chance both are aces?

  1. First card is an ace: $4/52$.
  2. Given the first was an ace, 3 aces remain among 51 cards: $P(\text{2nd ace}\mid\text{1st ace}) = 3/51$.
  3. Both: $\dfrac{4}{52}\times\dfrac{3}{51} = \dfrac{12}{2652} = \dfrac{1}{221} \approx 0.0045$.

Rearranging the definition of conditional probability gives the multiplication rule:

$$P(A \cap B) = P(A \mid B)\,P(B) = P(B \mid A)\,P(A).$$

For many events, apply it again and again (the chain rule of probability):

$$P(A_1 \cap A_2 \cap \dots \cap A_n) = P(A_1)\,P(A_2\mid A_1)\,P(A_3 \mid A_1\cap A_2)\cdots P(A_n\mid A_1\cap\dots\cap A_{n-1}).$$
  • A probability tree draws this: each branch carries a conditional probability, the probability of a leaf is the product along its path, and all leaves add up to 1.
  • The rule always holds. It only simplifies to $P(A)\,P(B)$ when the events are independent (next section).
  • The same rule for random variables, $p(x, y) = p(y\mid x)\,p(x)$, is how every generative model is written: first draw one thing, then the next given the first.
Why do we need it?

Joint events ("visited AND bought AND returned") are hard to count directly, but the step-by-step conditional rates are easy to measure. The rule glues the steps together.

Where is it used?

Conversion funnels, survival and churn curves, language models ($P(w_1)P(w_2\mid w_1)\cdots$), naive Bayes, Markov chains, and every NumPyro model, whose sample statements multiply conditional densities into one joint density.

How is it used?

Write the process as a sequence of steps. Estimate each step's rate among the units that reached it. Multiply along the path. To find a total, add the paths that end in the outcome you want (that is the next-but-two section).

visit1000 cart400 checkout200 pay 120 × 0.4 × 0.5 (given cart) × 0.6 (given checkout) P(pay) = 0.4 × 0.5 × 0.6 = 0.12 = 120 / 1000
A funnel is the chain rule: every step's rate is a conditional probability among those who reached it, and the overall rate is their product.

The first split is the device; the second is whether the visitor converts given that device. Branch thickness shows how much probability flows through it. Read a leaf: it is the product of its two branches. Switch to All paths to "converts": the two highlighted leaves add up to $P(\text{converts})$. Set the two conversion rates equal and notice that $P(\text{converts})$ no longer depends on the device split.

"P(A and B) = P(A) × P(B), always."

Only when A and B are independent. In general you must multiply by the conditional probability: $P(A)\,P(B\mid A)$. For two aces, $\tfrac{4}{52}\times\tfrac{4}{52} \approx 0.0059$ is wrong; $\tfrac{4}{52}\times\tfrac{3}{51} \approx 0.0045$ is right.

"Each funnel step's rate is out of all visitors."

Each step's rate must be measured among the units that reached that step. "60% pay" means 60% of those at checkout, not 60% of all visitors.

Every NumPyro model is a multiplication rule written as code. In your forecasting model the joint density is $p(\theta, y) = p(\theta)\,p(y\mid\theta)$: first the priors (for example the trend slopes and the noise scale), then the data given those parameters. Each numpyro.sample(...) line adds one conditional factor (in log space: one term added to the joint log-density).

$P(A\cap B) = P(A\mid B)P(B) = P(B\mid A)P(A)$. Chain: $P(A_1\cap\dots\cap A_n) = P(A_1)P(A_2\mid A_1)\cdots$

Tree: multiply along a path; leaves sum to 1. Funnel = chain rule.

Trap: $P(A)P(B)$ only if independent.

Quick check: 30% of sessions see a promo banner; 8% of those who see it click it. What share of all sessions click the banner?

$P(\text{see}\cap\text{click}) = P(\text{see})\,P(\text{click}\mid\text{see}) = 0.3 \times 0.08 = 0.024$, i.e. 2.4% of all sessions.

Independence: when knowing B tells you nothing about A core

Flip a coin in Delhi and another in London. Learning that the Delhi coin landed heads does not change your 50% for London. The two events are independent: knowing one tells you nothing about the other.

There is a neat picture. Event $A$ takes up some share of the whole world, say half. $B$ is independent of $A$ when $A$ also takes up exactly half of $B$. In other words, $A$ cuts $B$ in the same proportion as it cuts everything. Then "shrinking the world to $B$" changes nothing.

Three ways to say it:

  • Picture: A occupies the same fraction of B as it does of the whole world.
  • Numbers: P(first die even) = 1/2, and P(first die even | sum = 7) is also 1/2.
  • Slogan: independent = "the news about B is useless for A".

Two dice. $A$ = "first die is even", $B$ = "sum = 7", $C$ = "sum = 8".

  1. $P(A) = 18/36 = 1/2$. $P(B) = 6/36 = 1/6$. $P(C) = 5/36$.
  2. $A\cap B$: (2,5), (4,3), (6,1), so $P(A\cap B) = 3/36 = 1/12$.
  3. $P(A)\,P(B) = \tfrac12\times\tfrac16 = \tfrac1{12}$. Equal, so $A$ and $B$ are independent. Check: $P(A\mid B) = 3/6 = 1/2 = P(A)$.
  4. $A\cap C$: (2,6), (4,4), (6,2), so $P(A\cap C) = 3/36 = 6/72$.
  5. $P(A)\,P(C) = \tfrac12\times\tfrac5{36} = \tfrac5{72}$. Not equal, so $A$ and $C$ are dependent: $P(A\mid C) = 3/5 = 0.6 \ne 0.5$.

Surprise: both events are about the same die, yet "even" and "sum = 7" are independent. Independence is a fact about probabilities, not about whether two things are physically connected.

Events $A$ and $B$ are independent when

$$P(A\cap B) = P(A)\,P(B).$$
  • If $P(B)\gt 0$ this is the same as $P(A\mid B) = P(A)$; if $P(A) \gt 0$, the same as $P(B\mid A) = P(B)$. It is symmetric: if A tells you nothing about B, then B tells you nothing about A.
  • If $A$ and $B$ are independent, so are $A$ and $B^c$, $A^c$ and $B$, and $A^c$ and $B^c$.
  • Several events $A_1,\dots,A_n$ are mutually independent when the product rule holds for every sub-group of them, for example $P(A_1\cap A_2\cap A_3) = P(A_1)P(A_2)P(A_3)$ as well as every pair. Pairs alone are not enough (see the quick check).
  • Two random variables $X$ and $Y$ (numbers produced by a random process; Chapter 4.4) are independent when every event about $X$ is independent of every event about $Y$. Equivalently the joint probability (or density) factorizes: $p(x, y) = p(x)\,p(y)$ for all $x, y$.
  • iid ("independent and identically distributed") means: mutually independent, and every one follows the same distribution.
Why do we need it?

Independence lets us multiply. Without it, the probability of 1,000 users' outcomes would need a gigantic joint table; with it, the joint probability is just a product of 1,000 small terms, which is what makes a likelihood computable.

Where is it used?

Every iid likelihood $\prod_i p(y_i\mid\theta)$, randomization in A/B tests (assignment independent of user traits), the "at least one" rule $1-(1-p)^n$, naive Bayes classifiers, and chi-square tests of independence (Chapter 5.9).

How is it used?

Ask "if I learned B, would my probability for A move?". If not, multiply. With data, compare $P(A\mid B)$ estimated in the data with $P(A)$; they will never match exactly, so use a test (chi-square) or a model to judge whether the gap is more than noise.

The square is the whole world (area 1). The blue strip is event $A$; set its width with the slider. The orange rectangle is event $B$; drag its two corners. Purple is $A\cap B$. Press Horizontal band: $A$ now cuts $B$ in exactly the same proportion as it cuts the square, so they are independent. Then press Disjoint from A: $P(A\cap B) = 0$ while $P(A)P(B)\gt 0$, so disjoint events are strongly dependent.

"Independent means the events cannot happen together."

That is disjoint (mutually exclusive): $P(A\cap B) = 0$. Disjoint events with positive probabilities are as dependent as it gets: if one happens, the other is impossible. Independent events overlap, in exactly the proportion $P(A)P(B)$.

"These two events are about the same thing, so they must be dependent."

"First die even" and "sum = 7" share a die and are still independent. Check the numbers, not the story.

"In my data $P(A\mid B) = 0.081$ and $P(A) = 0.080$, so they are dependent."

Estimates from finite data never match exactly. Independence is a statement about the true probabilities; whether a small gap is real or noise is a question for a test or a model (Chapter 5.9).

Randomization is a promise of independence. In an A/B framework like yours, assignment is random, so the variant a user gets is independent of their device, country or habits: $P(\text{variant B}\mid\text{mobile}) = P(\text{variant B})$. That is why a difference in conversion between arms can be blamed on the variant. A bug that assigns variants by device breaks this independence; comparing the device mix across arms (a balance check) or a sample-ratio check within each device can reveal it (Chapter 5.11).

Independent ⇔ $P(A\cap B) = P(A)P(B)$ ⇔ $P(A\mid B) = P(A)$ (if $P(B)\gt 0$). Variables: $p(x,y) = p(x)p(y)$.

Picture: A cuts B in the same proportion as it cuts the whole world.

Traps: independent ≠ disjoint; pairwise ≠ mutual; check numbers, not stories.

Quick check: flip two fair coins. A = "first is heads", B = "second is heads", C = "the two coins match". Are A, B, C pairwise independent? Mutually independent?

Each has probability 1/2. $P(A\cap B) = P(HH) = 1/4$, $P(A\cap C) = P(HH) = 1/4$, $P(B\cap C) = P(HH) = 1/4$: every pair multiplies, so they are pairwise independent. But $P(A\cap B\cap C) = P(HH) = 1/4 \ne 1/8$: once you know A and B, C is certain. So they are not mutually independent.

Independence, correlation and conditional independence: three different things core

People mix up three questions that sound alike:

  • Independence: does knowing X tell me anything at all about Y (its centre, its spread, its shape)?
  • Correlation: do X and Y tend to move together along a straight-line trend?
  • Conditional independence: once I already know a third thing Z, does X still tell me anything about Y?

Everyday picture: ice-cream sales and sunburns rise and fall together, so they are correlated and dependent. But compare only days with the same temperature: now ice-cream sales tell you nothing extra about sunburns. They are conditionally independent given temperature. The hot weather (a common cause) created the whole link.

Three ways to say it:

  • Picture: independence = every slice of the scatter plot looks the same; zero correlation = no tilted straight-line trend; conditional independence = no link inside each Z-group.
  • Numbers: $Y = X^2$ with $X\in\{-1,0,1\}$ has correlation exactly 0, yet $X$ fixes $Y$ completely.
  • Slogan: zero correlation is not independence, and independence overall is not independence given Z (in either direction).

(1) Uncorrelated but dependent. $X$ is $-1$, $0$ or $1$, each with probability $1/3$, and $Y = X^2$. Covariance (how much X and Y move together on average; defined below) is $Cov(X,Y) = E[XY] - E[X]E[Y]$, where $E[\cdot]$ means "the average value" (Chapter 4.5).

  1. $E[X] = (-1 + 0 + 1)/3 = 0$.
  2. $XY = X^3$, which is $-1, 0, 1$, so $E[XY] = 0$.
  3. $Cov(X,Y) = 0 - 0\cdot E[Y] = 0$, so the correlation is 0.
  4. But $P(Y = 1) = 2/3$ while $P(Y = 1\mid X = 0) = 0$. Knowing X changes Y completely: dependent.

(2) Dependent, but conditionally independent. 20% of sessions are bots. Two bot detectors use different signals; each flags a bot with probability 0.8 and a human with probability 0.1, independently given whether the session is a bot.

  1. $P(\text{D2 flags}) = 0.2\times0.8 + 0.8\times0.1 = 0.16 + 0.08 = 0.24$.
  2. $P(\text{both flag}) = 0.2\times0.8^2 + 0.8\times0.1^2 = 0.128 + 0.008 = 0.136$.
  3. $P(\text{D2 flags}\mid\text{D1 flagged}) = 0.136/0.24 \approx 0.567$, far above $0.24$: overall the detectors are dependent.
  4. Among bots only: $P(\text{D2}\mid\text{D1, bot}) = 0.8 = P(\text{D2}\mid\text{bot})$. Given the truth, they are independent.

(3) Independent, but dependent given Z. A slow page (S, probability 0.2) and a bad recommendation (R, probability 0.3) happen independently. A user leaves (L) if either happens.

  1. $P(L) = 1 - (1-0.2)(1-0.3) = 1 - 0.56 = 0.44$.
  2. Given the user left: $P(S\mid L) = P(S)/P(L) = 0.2/0.44 \approx 0.455$ (every slow page leads to leaving, so $S\cap L = S$).
  3. Given the user left and the recommendation was bad: $P(S\mid L, R) = 0.2$. R already "explains" the leaving, so S drops back to its usual 0.2.
  4. Inside the world "users who left", learning R changes the chance of S: S and R are dependent given L. This is called explaining away.
  • Independence: $p(x, y) = p(x)\,p(y)$ for all values. Written $X \perp Y$.
  • Covariance and correlation (full treatment in Chapter 4.15): $Cov(X,Y) = E[(X - E[X])(Y - E[Y])]$ and $\rho = \dfrac{Cov(X,Y)}{\sigma_X\sigma_Y}$, a number from $-1$ to $1$ ($\sigma$ = standard deviation, the typical spread). It measures only the straight-line part of the link.
  • Conditional independence: $X \perp Y \mid Z$ when $p(x, y\mid z) = p(x\mid z)\,p(y\mid z)$ for every $z$; equivalently $p(x\mid y, z) = p(x\mid z)$: once $Z$ is known, $Y$ adds nothing about $X$.
ClaimTrue?Why / counterexample
Independent ⇒ uncorrelatedYes (if the variances are finite)If $p(x,y) = p(x)p(y)$ then $E[XY] = E[X]E[Y]$, so $Cov = 0$.
Uncorrelated ⇒ independentNo$Y = X^2$: correlation 0, fully dependent. (Exception: for jointly Normal variables, zero correlation does mean independence.)
Conditionally independent ⇒ independentNoTwo detectors given the truth: independent within each group, dependent overall (common cause).
Independent ⇒ conditionally independentNoSlow page and bad recommendation: independent overall, dependent among users who left (common effect).
Why do we need it?

Models are built from independence assumptions, and most of them are conditional ones. Knowing which kind you are assuming tells you what can break it, and stops you from reading "r ≈ 0" as "no relationship".

Where is it used?

Bayesian likelihoods ("data iid given θ"), hierarchical models (groups independent given the hyperparameters), naive Bayes, graphical models and causal diagrams (confounders and colliders, Chapter 5.12), and feature screening with correlation matrices.

How is it used?

Plot the data, not just r: look for curved or fan shapes. Split by the third variable (device, segment, truth label) and check the link inside each group. Ask what common causes or common effects (like "only users who left") sit in your data.

Common cause bot? D1 D2 dependent overall · independent given "bot?" Common effect S R left independent overall · dependent given "left"
Arrows mean "influences". A common cause makes two things dependent until you condition on the cause. A common effect does the opposite: two independent causes become dependent once you look only at cases where the effect happened.

Each picture has 300 points. The purple line joins the average $y$ in each vertical slice of $x$; the purple band is ±1 standard deviation within the slice. Compare Line (large r) with Parabola, Ring and V shape: r is close to 0, yet the slices plainly change with $x$. Only Independent has slices that all look alike. Press New sample to see how much r wobbles by chance.

Blue bars: the chance that detector 2 flags a session. Orange bars: the same chance given detector 1 flagged it. For everyone, orange is much taller: the detectors are dependent, because both react to the same hidden truth. Inside bots only and humans only, blue and orange are equal: conditionally independent. Now raise the shared blind spot (the share of cases where both detectors use the same signal and give the same answer): the bars split even inside each group.

S = slow page, R = bad recommendation, independent of each other. A user leaves (L) if S or R happens. The first two bars (blue) look at all users: $P(S\mid R) = P(S)$, so S and R are independent. The last three bars (orange) look only at users who left. Inside that world, learning that R happened pushes S back down, and learning that R did not happen makes S certain. Move the sliders: the pattern never goes away.

"The correlation is 0.02, so the two variables are unrelated."

Correlation only sees straight-line trends. Curved (U, V), ring-shaped or spread-changing relationships can have r ≈ 0. Always plot the data and look at slices.

"Two signals that both react to the truth are independent evidence, so I can multiply their likelihood ratios."

You may multiply only if they are independent given the truth. If they share a data source or a blind spot, multiplying double-counts the evidence and makes you overconfident.

"Adding more 'control' variables or filters can only remove fake links."

Conditioning on a common effect (looking only at users who left, only at shipped features, only at survivors) can create a link between independent causes. This is one form of selection bias (Chapter 5.4).

Your Bayesian models lean on conditional independence. In the A/B framework, users' conversions are modelled as independent given the conversion rate: $y_i \perp y_j \mid \theta$, so the likelihood is $\prod_i p(y_i\mid\theta)$. Before you know θ they are not independent: many conversions so far raise your estimate of θ and therefore the chance that the next user converts. In the hierarchical model, segments are treated as conditionally independent given the global parameters (Chapter 6.5). In the forecasting model, the noise terms $\epsilon_t$ are modelled as independent given the trend, seasonality, holiday and regressor terms; autocorrelated residuals mean this assumption has failed, and the predictive intervals can come out too narrow (Chapter 7.17).

"Independent and uncorrelated mean the same thing."

Independence implies zero correlation (when the variances exist), but zero correlation does not imply independence: $Y = X^2$ is uncorrelated with a symmetric $X$ yet completely determined by it.

"If two things are independent, they stay independent whatever I condition on."

Conditional independence and independence do not imply each other. A common cause gives dependence that disappears given the cause; a common effect gives independence that disappears given the effect.

Model answer: "Independence means the joint distribution factorizes, so knowing one variable tells you nothing about the other. Correlation only measures the linear part of the relationship, so it can be zero for strongly dependent variables. Conditional independence is independence inside each level of a third variable. My Bayesian models assume observations are independent given the parameters, which is weaker than assuming they are independent outright."

Independent: $p(x,y) = p(x)p(y)$. Uncorrelated: $Cov(X,Y) = 0$ (linear only). Conditionally independent: $p(x,y\mid z) = p(x\mid z)p(y\mid z)$.

Independent ⇒ uncorrelated; not the reverse ($Y = X^2$).

Common cause: dependent, but independent given the cause. Common effect: independent, but dependent given the effect (explaining away).

Quick check: shoe size and reading score are strongly correlated in a sample of children aged 5 to 12. Within each age, they are unrelated. Which kind of (in)dependence is this?

Age is a common cause: older children have bigger feet and read better. Shoe size and reading score are dependent overall but conditionally independent given age.

The law of total probability: split into cases, then add core

Your manager asks: "what is our overall conversion rate?" You do not know it directly, but you know it for each device: desktop converts 10%, mobile converts 5%, and 60% of traffic is desktop. The overall rate is not the plain average of 10% and 5% (that would be 7.5%). Desktop has more visitors, so it should count more. The answer is a weighted average: each case's rate, weighted by how common that case is.

That is the law of total probability: split the world into cases that do not overlap and together cover everything, work out the probability inside each case, then add them up with the case sizes as weights.

Three ways to say it:

  • Picture: the event is cut into pieces by the cases; add up the pieces.
  • Numbers: P(convert) = 0.6 × 0.10 + 0.4 × 0.05 = 0.06 + 0.02 = 0.08.
  • Slogan: the overall rate is the size-weighted average of the case rates.

Devices. $P(\text{desktop}) = 0.6$, $P(\text{mobile}) = 0.4$, $P(\text{convert}\mid\text{desktop}) = 0.10$, $P(\text{convert}\mid\text{mobile}) = 0.05$.

  1. Desktop piece: $P(\text{convert}\cap\text{desktop}) = 0.10 \times 0.6 = 0.06$ (multiplication rule).
  2. Mobile piece: $P(\text{convert}\cap\text{mobile}) = 0.05\times 0.4 = 0.02$.
  3. The two cases do not overlap and cover every visitor, so add: $P(\text{convert}) = 0.06 + 0.02 = 0.08$.
  4. This matches the 1,000-visitor table from the first section: $80/1000 = 0.08$.

Three segments. New visitors are 50% of traffic and convert 4%; returning are 30% and convert 10%; loyal are 20% and convert 20%.

  1. $0.5 \times 0.04 = 0.02$, $\;0.3\times 0.10 = 0.03$, $\;0.2\times0.20 = 0.04$.
  2. $P(\text{convert}) = 0.02 + 0.03 + 0.04 = 0.09$.
  3. The plain average of the three rates is $(0.04 + 0.10 + 0.20)/3 \approx 0.113$: wrong, because it pretends the segments are equally large.

A partition of the sample space is a set of events $B_1, \dots, B_k$ that do not overlap (at most one of them happens) and together cover everything (at least one happens). For any event $A$:

$$P(A) = \sum_{i=1}^{k} P(A\mid B_i)\,P(B_i).$$
  • Two cases: $P(A) = P(A\mid B)P(B) + P(A\mid B^c)P(B^c)$.
  • Each term $P(A\mid B_i)P(B_i) = P(A\cap B_i)$ is one "piece" of $A$; the pieces do not overlap, so they add.
  • $P(A)$ computed this way is often called the marginal probability of $A$: "marginal" means we have summed out (averaged over) the other variable, here the case $B$.
  • For a continuous case variable $\theta$ the sum becomes an integral: $p(y) = \int p(y\mid\theta)\,p(\theta)\,d\theta$. You will meet this as the evidence or marginal likelihood in Chapter 6.1.
Why do we need it?

Many probabilities are easy to get within cases but hard to get overall. Splitting into cases turns one hard question into several easy ones. It also supplies the bottom line of Bayes' theorem.

Where is it used?

Overall metrics from segment metrics, mixture models (each data point comes from one of several components), the denominator of Bayes' theorem, the marginal likelihood $p(D)$ in Bayesian model comparison, and predictive distributions that average over parameter uncertainty.

How is it used?

List cases that cannot both happen and cover everything (devices, segments, hypotheses). Get each case's share and its conditional rate. Multiply and add. In pandas: (df.groupby("seg").converted.mean() * df.seg.value_counts(normalize=True)).sum(), which equals df.converted.mean().

B₁ new · 50% B₂ returning · 30% B₃ loyal · 20% A∩B₁ = 0.5 × 0.04 = 0.02 A∩B₂ = 0.03 A∩B₃ = 0.04 P(A) = 0.02 + 0.03 + 0.04 = 0.09
The cases B₁, B₂, B₃ split the world into non-overlapping columns (width = share of traffic). The event A = "converts" is cut into one piece per column (height = conversion rate in that case). The total is the sum of the pieces.

Each column is a segment: its width is its share of traffic and its height is its conversion rate, so its area is $P(\text{segment} \cap \text{convert})$. The purple line is the true overall rate (total area spread over the full width); the dashed grey line is the naive average of the three rates. Make the loyal segment tiny and very good: the naive average jumps, the true rate barely moves. Make all sizes equal: the two lines meet.

"The overall rate is the average of the segment rates."

Only if every segment is the same size. In general you must weight each rate by its segment's share: $\sum_i P(B_i)\,P(A\mid B_i)$.

"Any list of groups will do as cases."

The cases must not overlap and must cover everyone. "Mobile users" and "new users" overlap, so adding their pieces double-counts; "mobile" and "desktop" without "tablet" misses people.

"If B converts better than A in every segment, B must convert better overall."

Not if the arms have very different segment mixes; this reversal is Simpson's paradox (Chapter 4.6). Randomization protects A/B tests from it because both arms get the same mix on average.

In your A/B framework, an arm's overall conversion rate is the segment-mix-weighted average of its segment rates; that is the law of total probability, and it is why a shift in traffic mix can move a metric with no change in any segment. In both projects the Bayesian evidence $p(D) = \int p(D\mid\theta)\,p(\theta)\,d\theta$ is this law with infinitely many cases (one per value of θ), and the forecast distribution $p(\tilde y\mid D) = \int p(\tilde y\mid\theta)\,p(\theta\mid D)\,d\theta$ averages the predictions over parameter uncertainty in the same way (Chapter 6.1, Chapter 7.14).

$P(A) = \sum_i P(A\mid B_i)P(B_i)$ for a partition $B_1..B_k$ (no overlaps, covers everything).

Meaning: overall rate = size-weighted average of the case rates. Continuous: $p(y) = \int p(y\mid\theta)p(\theta)d\theta$.

Trap: do not take a plain average of segment rates.

Quick check: 70% of orders ship from warehouse 1 (late 5% of the time) and 30% from warehouse 2 (late 20% of the time). What share of all orders is late?

$0.7\times0.05 + 0.3\times0.20 = 0.035 + 0.06 = 0.095$, so 9.5% of orders are late.

Bayes' theorem: turning P(evidence | cause) into P(cause | evidence) core

A screening test catches 90% of people who have a rare condition and wrongly flags 9% of healthy people. Your test comes back positive. Is there a 90% chance you have it? Most people (including many doctors in surveys) say yes. The truth is closer to 9%.

Imagine 10,000 people. Only 1% (100 people) have the condition; 90 of them test positive. Of the 9,900 healthy people, 9% (891) also test positive. So 981 people get a positive result, and only 90 of them are sick: $90/981 \approx 9\%$. The healthy group is so much bigger that its small error rate produces far more positives than the sick group does.

Bayes' theorem does exactly this calculation. It starts from what you know (how likely the evidence is under each explanation, and how common each explanation is) and returns what you want (how likely each explanation is now that you have seen the evidence).

Three ways to say it:

  • Picture: shrink the world to "everyone who tested positive", then ask what share of that world is sick.
  • Numbers: 90 true positives out of 90 + 891 = 981 positives ≈ 9.2%.
  • Slogan: a positive test is a strong clue only if the thing it points to was not very rare to begin with.

Prevalence (how common the condition is, also called the base rate) $= 1\%$. Sensitivity $P(+\mid\text{sick}) = 90\%$. Specificity $P(-\mid\text{healthy}) = 91\%$, so the false-positive rate is $P(+\mid\text{healthy}) = 9\%$.

  1. Top: $P(+\mid\text{sick})\,P(\text{sick}) = 0.90\times0.01 = 0.009$.
  2. Bottom, by total probability: $P(+) = 0.90\times0.01 + 0.09\times0.99 = 0.009 + 0.0891 = 0.0981$.
  3. $P(\text{sick}\mid +) = 0.009 / 0.0981 \approx 0.0917$, about 9.2%.
  4. Odds version. Odds = probability "for" divided by probability "against". Before the test the odds of being sick are $0.01/0.99 = 1:99$. The likelihood ratio of a positive is $0.90/0.09 = 10$ (a positive is 10 times more likely if sick).
  5. After the test: odds $= 10 \times 1:99 = 10:99$, so $P = 10/(10+99) = 10/109 \approx 0.0917$. Same answer, easier arithmetic.

Write the multiplication rule both ways, $P(A\cap B) = P(A\mid B)P(B) = P(B\mid A)P(A)$, and divide by $P(B)$:

$$P(A\mid B) = \frac{P(B\mid A)\,P(A)}{P(B)}, \qquad P(B) = P(B\mid A)P(A) + P(B\mid A^c)P(A^c).$$
  • $P(A)$: the prior probability of the explanation, before seeing the evidence (here the base rate).
  • $P(B\mid A)$: the likelihood, how probable the evidence is if the explanation is true.
  • $P(B)$: the evidence (total probability of seeing B), which makes the answers add to 1.
  • $P(A\mid B)$: the posterior, the updated probability after seeing the evidence.
  • Many explanations $H_1,\dots,H_k$ (a partition): $P(H_i\mid E) = \dfrac{P(E\mid H_i)P(H_i)}{\sum_j P(E\mid H_j)P(H_j)}$.
  • Odds form: $\dfrac{P(A\mid B)}{P(A^c\mid B)} = \dfrac{P(A)}{P(A^c)}\times\dfrac{P(B\mid A)}{P(B\mid A^c)}$, i.e. posterior odds = prior odds × likelihood ratio.

These words are used loosely here; Chapter 6.1 gives them their full meaning for unknown parameters.

Why do we need it?

We usually know how causes produce evidence ($P(\text{evidence}\mid\text{cause})$), but decisions need the reverse ($P(\text{cause}\mid\text{evidence})$). Bayes is the only correct way to flip a conditional, and it forces you to include the base rate.

Where is it used?

Medical and fraud screening, spam filters and naive Bayes, alert systems (how many alarms are real?), and all of Bayesian statistics: $p(\theta\mid D) \propto p(D\mid\theta)\,p(\theta)$, including the posterior probability that variant B beats A.

How is it used?

Write down the base rates and the evidence rates. Use natural frequencies (imagine 10,000 cases) or odds × likelihood ratio. Always compute the bottom line with the law of total probability, so that false positives from the large group are counted.

10,000 people 1%99% 100 have it 9,900 do not 90%10%9%91% 90 positive 10 negative 891 positive 9,009 negative P(has it | positive) = 90 / (90 + 891) = 90 / 981 ≈ 9.2%
Natural frequencies make Bayes' theorem a counting exercise. The positives come from two places: 90 true positives (blue) and 891 false positives (red). The rare condition is swamped by the small error rate of the big healthy group.

Each dot is a person. Filled blue: has the condition and tests positive. Blue ring: has it but the test misses it. Red: healthy but tests positive (false alarm). Grey: healthy and negative. Turn on Show only positive tests to shrink the world. Then raise the prevalence from 1% to 20% and watch the share of blue among the positives climb, even though the test itself did not change. Finally set specificity to 99%: false alarms nearly vanish.

Start with the 1% base rate. Press Positive: the probability jumps to about 9%. Press Positive again: about 50%. A third time: about 91%. Each bar is the posterior after one more result, and it becomes the prior for the next test. Try a Negative after two positives. The readout shows the odds arithmetic. This assumes repeat results are independent given the true state (conditional independence, from the section above); a test that errs the same way twice would not deserve this much trust.

"The test is 90% accurate, so a positive means a 90% chance of having the condition."

90% is $P(+\mid\text{sick})$. You need $P(\text{sick}\mid +)$, which also depends on the base rate and the false-positive rate. Ignoring the base rate is called the base-rate fallacy.

"The chance of this evidence if he were innocent is 1 in 1,000, so he is guilty with probability 99.9%." (the prosecutor's fallacy)

That is $P(\text{evidence}\mid\text{innocent})$. To get $P(\text{innocent}\mid\text{evidence})$ you also need how many innocent people could have produced the same evidence, i.e. the prior.

"In Bayes' theorem, the bottom $P(B)$ is just $P(B\mid A)$."

$P(B)$ is the total probability of the evidence from every explanation: $P(B\mid A)P(A) + P(B\mid A^c)P(A^c)$. Forgetting the second term forgets the false positives.

"A p-value of 0.03 means there is a 3% chance the null hypothesis is true."

A p-value is computed assuming the null is true: it is about P(data this extreme | H0). The probability that H0 is true given the data is the reverse conditional and needs a prior, through Bayes' theorem.

Model answer: "P(A | B) and P(B | A) are different quantities; Bayes' theorem links them through the base rate P(A). A test's sensitivity is P(positive | condition), but what a patient needs is P(condition | positive), which can be small when the condition is rare. The same distinction separates a p-value from the posterior probability of a hypothesis (Chapter 5.6)."

Your A/B framework reports posterior decisions such as $P(\theta_B \gt \theta_A\mid D)$. That number comes from Bayes' theorem applied to the unknown conversion rates: $p(\theta\mid D) \propto p(D\mid\theta)\,p(\theta)$, the same "likelihood × prior, then normalize" as the screening test, with θ in place of "has the condition" (Chapter 6.1, Chapter 6.4). It answers the business question directly ("how likely is B better, given the data?"), which is a different quantity from a p-value.

$P(A\mid B) = \dfrac{P(B\mid A)P(A)}{P(B\mid A)P(A) + P(B\mid A^c)P(A^c)}$. Odds: posterior odds = prior odds × $\dfrac{P(B\mid A)}{P(B\mid A^c)}$.

Screening: prevalence 1%, sensitivity 90%, false positives 9% → $P(\text{sick}\mid +) \approx 9\%$.

Trap: base-rate fallacy; $P(\text{evidence}\mid H) \ne P(H\mid\text{evidence})$.

Quick check: 2% of transactions are fraud. A rule flags 95% of fraud and 5% of normal transactions. What share of flagged transactions are fraud?

$P(\text{fraud}\mid\text{flag}) = \dfrac{0.95\times0.02}{0.95\times0.02 + 0.05\times0.98} = \dfrac{0.019}{0.019 + 0.049} = \dfrac{0.019}{0.068} \approx 0.279$. Only about 28% of flags are real fraud. Odds check: $1:49 \times 19 = 19:49$, and $19/68 \approx 0.279$.

Bayes in your A/B framework: which variant did this converter come from?

A purchase shows up in the logs, but the variant label is missing. Which arm did this buyer most likely see? Two things matter. First, how much traffic each arm got: if 80% of users saw A, most buyers come from A even if B is a bit better. Second, how well each arm converts: the better arm produces more buyers per visitor. Bayes' theorem combines exactly these two: the traffic split is the prior, the conversion rates are the likelihoods.

Now turn the question around. Instead of "which arm?", ask "which conversion rate?". Write down a few candidate rates, give each a prior weight, and score each by how well it predicts the conversions you saw. The same rule gives a probability for every candidate. That is Bayesian inference in miniature, and your Beta-Binomial models do it with every possible rate at once.

Three ways to say it:

  • Picture: shrink the world to "all converters" and see what share of it came from B.
  • Numbers: 50/50 split, A converts 10%, B 12%: P(B | converted) = 0.06/(0.05 + 0.06) ≈ 0.545.
  • Slogan: posterior ∝ likelihood × prior, then rescale so the answers add to 1.

Which arm? Split 50/50; $P(\text{convert}\mid A) = 0.10$, $P(\text{convert}\mid B) = 0.12$.

  1. Pieces: $P(A\cap\text{conv}) = 0.5\times0.10 = 0.05$, $\;P(B\cap\text{conv}) = 0.5\times0.12 = 0.06$.
  2. Total: $P(\text{conv}) = 0.05 + 0.06 = 0.11$.
  3. $P(B\mid\text{conv}) = 0.06/0.11 \approx 0.545$.
  4. With an 80/20 split (20% to B): $\dfrac{0.2\times0.12}{0.8\times0.10 + 0.2\times0.12} = \dfrac{0.024}{0.104} \approx 0.231$. Same rates, very different answer: the prior matters.

Which rate? Three candidate rates 5%, 10%, 15%, equally likely before data. You observe $k = 3$ conversions among $n = 20$ users.

  1. Likelihood of the data under each rate (Binomial, Chapter 4.7): $P(k = 3\mid\theta) = \binom{20}{3}\theta^3(1-\theta)^{17}$, with $\binom{20}{3} = 1140$.
  2. $\theta = 0.05$: $0.0596$; $\;\theta = 0.10$: $0.1901$; $\;\theta = 0.15$: $0.2428$.
  3. Prior × likelihood: each times $1/3$: $0.0199,\ 0.0634,\ 0.0809$; their sum (the evidence) is $0.1642$.
  4. Posterior: $0.0199/0.1642 \approx 0.121$, $\;0.0634/0.1642\approx 0.386$, $\;0.0809/0.1642 \approx 0.493$.

For hypotheses $H_1,\dots,H_k$ that do not overlap and cover all possibilities, and observed evidence $E$:

$$P(H_i\mid E) = \frac{P(E\mid H_i)\,P(H_i)}{\sum_{j} P(E\mid H_j)\,P(H_j)} \quad\text{i.e.}\quad \text{posterior} \propto \text{likelihood}\times\text{prior}.$$
  • "$\propto$" means "proportional to": equal up to one constant factor that is the same for every $H_i$. Here that factor is $1/P(E)$, the bottom line, found by the law of total probability.
  • The likelihoods $P(E\mid H_i)$ need not add to 1 across hypotheses. Only after multiplying by the prior and dividing by the total do we get a probability distribution over the hypotheses.
  • When the hypothesis is a continuous number θ (any rate from 0 to 1), the sum becomes an integral and $P(\cdot)$ becomes a density: $p(\theta\mid D) = p(D\mid\theta)\,p(\theta)\,/\,p(D)$ (Chapter 6.1).
Why do we need it?

It turns "how well does each explanation predict what I saw?" into "how much should I believe each explanation now?", while respecting how plausible each explanation was to begin with. That is the core move of every Bayesian decision.

Where is it used?

Attribution questions (which campaign, arm or segment produced this event?), mixture models (which component produced this point?), classifiers like naive Bayes, and grid approximations of a posterior over an unknown rate, the warm-up for Beta-Binomial models.

How is it used?

List the hypotheses, attach prior weights, compute each likelihood for the observed data, multiply, then divide by the sum. In NumPy: post = prior * like / (prior * like).sum(). Work in logs (logsumexp) when likelihoods are tiny.

users prior: split arm A arm B likelihood θ_Alikelihood θ_B a converter posterior:P(B | converted)
Read forwards, the experiment is a probability tree: the split, then conversion given the arm. Bayes' theorem reads it backwards: given a conversion, which arm was it?

Column widths are the traffic shares of A (blue) and B (orange); heights are their conversion rates (the non-converters above are not drawn). The coloured areas are all the converters. Move traffic to B down to 20% and see $P(B\mid\text{converted})$ fall even though B converts better. Press Simulate 10,000 visitors to check the formula against random data, and New sample to see the noise.

Five candidate rates are the hypotheses. Blue bars: the prior. Orange bars: the likelihood $P(k\mid n,\theta)$ of your data under each rate (rescaled to add to 1 only so you can compare shapes). Green bars: the posterior = prior × likelihood, renormalized. Start with $k = 3$ of $n = 20$. Raise $n$ to 100 and $k$ to 15: the posterior concentrates. Switch the prior to "favours low rates" with small $n$ and then with large $n$: the data overrule the prior as $n$ grows.

"B converts 20% better than A, so a converter is 20% more likely to have come from B."

Only with an even split. The traffic split is the prior: with 20% of traffic on B, most converters still come from A.

"The likelihoods 0.06, 0.19, 0.24 are the probabilities of the three rates."

A likelihood is the probability of the data under each rate. Across rates they need not add to 1. To get probabilities of the rates you must multiply by the prior and normalize (Chapter 5.2 studies likelihoods on their own).

"The posterior says the true rate is 15%."

It says 15% is the most probable of the candidates you allowed, with probability about 0.49. Twenty users leave a lot of uncertainty, and a real model lets θ be any number in [0, 1] (Beta prior, Chapter 6.3).

A Beta-Binomial model in your framework is the "which rate?" widget with a continuous grid: every θ between 0 and 1 is a candidate, the Beta distribution is the prior, the Binomial is the likelihood, and the posterior is again a Beta (Chapter 6.3). From the two posteriors you then compute $P(\theta_B \gt \theta_A\mid D)$ by drawing pairs $(\theta_A, \theta_B)$ and counting how often B is larger (Chapter 6.4). If an interviewer asks how Bayes' theorem shows up in your framework, this is the answer: each arm's rate is updated by likelihood × prior, normalized.

$P(H_i\mid E) = \dfrac{P(E\mid H_i)P(H_i)}{\sum_j P(E\mid H_j)P(H_j)}$: posterior ∝ likelihood × prior.

Which arm? prior = traffic split, likelihood = arm conversion rate.

Trap: likelihoods are not probabilities of the hypotheses; the split (prior) changes the answer.

Quick check: in the "which rate?" example, what happens to the posterior if you multiply every prior weight by 10?

Nothing. Every prior × likelihood term is multiplied by 10, and so is their sum, so the ratios (the posterior) are unchanged. Only the relative prior weights matter.

Recap, cheat sheet and practice

  • Conditional probability $P(A\mid B) = P(A\cap B)/P(B)$: make B the new world and measure A inside it. $P(A\mid B) \ne P(B\mid A)$.
  • Multiplication rule $P(A\cap B) = P(A\mid B)P(B)$; chained, it describes funnels, trees and every generative model.
  • Independence: $P(A\cap B) = P(A)P(B)$, i.e. B is useless news about A. Not the same as disjoint; pairwise is not mutual.
  • Independence ≠ zero correlation ≠ conditional independence. Independence ⇒ uncorrelated, not the reverse. A common cause creates dependence that vanishes given the cause; a common effect creates dependence given the effect.
  • Total probability: $P(A) = \sum_i P(A\mid B_i)P(B_i)$, a size-weighted average over non-overlapping cases.
  • Bayes: $P(A\mid B) = P(B\mid A)P(A)/P(B)$; posterior odds = prior odds × likelihood ratio. Base rates matter; natural frequencies make it easy.

Cheat sheet

IdeaFormulaPicture / meaning
Conditional probability$P(A\mid B) = \dfrac{P(A\cap B)}{P(B)}$B is the new 100%
Multiplication (chain) rule$P(A\cap B) = P(A\mid B)P(B)$multiply along a tree path / funnel
Independence$P(A\cap B) = P(A)P(B)$A cuts B in the same proportion as the world
Conditional independence$p(x,y\mid z) = p(x\mid z)\,p(y\mid z)$no link inside each Z-group
Correlation$\rho = Cov(X,Y)/(\sigma_X\sigma_Y)$straight-line link only
Total probability$P(A) = \sum_i P(A\mid B_i)P(B_i)$weighted average over cases
Bayes' theorem$P(A\mid B) = \dfrac{P(B\mid A)P(A)}{P(B)}$flip the conditional, keep the base rate
Odds formpost. odds = prior odds × $\dfrac{P(B\mid A)}{P(B\mid A^c)}$each clue multiplies the odds
Code it · Python

import itertools
from fractions import Fraction
import numpy as np
import pandas as pd
from scipy import stats

# 1) Conditional probability by counting: shrink the world, then count
outcomes = list(itertools.product(range(1, 7), repeat=2))   # 36 equally likely pairs

def P(event, given=lambda o: True):
    world = [o for o in outcomes if given(o)]                # keep only outcomes where "given" holds
    return Fraction(sum(event(o) for o in world), len(world))

A = lambda o: o[0] + o[1] >= 10        # sum at least 10
B = lambda o: o[0] == 6                # first die shows 6
print(P(A), P(A, given=B), P(B, given=A))          # 1/6 1/2 1/2

# 2) Independence check: P(A and B) == P(A) * P(B)?
even = lambda o: o[0] % 2 == 0
sum7 = lambda o: sum(o) == 7
sum8 = lambda o: sum(o) == 8
print(P(lambda o: even(o) and sum7(o)) == P(even) * P(sum7))   # True  -> independent
print(P(lambda o: even(o) and sum8(o)) == P(even) * P(sum8))   # False -> dependent

# 3) Conditional probabilities from data with pandas
df = pd.DataFrame({
    "device":    ["mobile"] * 400 + ["desktop"] * 600,
    "converted": [1] * 20 + [0] * 380 + [1] * 60 + [0] * 540,
})
print(pd.crosstab(df.device, df.converted, normalize="index"))   # rows: P(converted | device): 0.10 desktop, 0.05 mobile
print(df[df.converted == 1].device.value_counts(normalize=True))  # P(device | converted): desktop 0.75, mobile 0.25

# 4) Law of total probability = size-weighted average of segment rates
share = df.device.value_counts(normalize=True)          # P(device)
rate = df.groupby("device").converted.mean()            # P(convert | device)
print((share * rate).sum(), df.converted.mean())        # 0.08 0.08

# 5) Bayes for a screening test, exact and by simulation
prev, sens, spec = 0.01, 0.90, 0.91
ppv = sens * prev / (sens * prev + (1 - spec) * (1 - prev))
print(round(ppv, 4))                                     # 0.0917

rng = np.random.default_rng(0)
n = 1_000_000
sick = rng.random(n) < prev
pos = np.where(sick, rng.random(n) < sens, rng.random(n) < 1 - spec)
print(round(sick[pos].mean(), 3))                        # about 0.092 (simulation noise)

# 6) Conditional independence: two bot detectors
is_bot = rng.random(n) < 0.2
p_flag = np.where(is_bot, 0.8, 0.1)                      # P(flag | truth), same for both detectors
d1 = rng.random(n) < p_flag                              # independent draws GIVEN the truth
d2 = rng.random(n) < p_flag
print(round(d2.mean(), 3), round(d2[d1].mean(), 3))      # about 0.24 vs 0.567: dependent overall
print(round(d2[is_bot].mean(), 3), round(d2[is_bot & d1].mean(), 3))   # both about 0.8: independent given bot

# 7) Bayes over several hypotheses: which conversion rate? (3 conversions in 20 users)
theta = np.array([0.05, 0.10, 0.15])
prior = np.array([1, 1, 1]) / 3
like = stats.binom.pmf(3, 20, theta)                     # P(k = 3 | theta) for each candidate
post = prior * like / np.sum(prior * like)               # Bayes: normalize by the total probability
print(like.round(4), post.round(3))                      # [0.0596 0.1901 0.2428] [0.121 0.386 0.493]
Test yourself

1. Two fair dice. What is $P(\text{sum} = 7 \mid \text{first die} = 3)$?

Given the first die is 3, the world has 6 outcomes (3,1)…(3,6). Only (3,4) sums to 7, so the answer is 1/6.

2. $P(A) = 0.4$, $P(B) = 0.5$ and $P(A\cap B) = 0.2$. The events are…

$P(A)P(B) = 0.4\times0.5 = 0.2 = P(A\cap B)$, which is the definition of independence. Overlapping is expected for independent events; disjoint would need $P(A\cap B) = 0$.

3. The correlation between X and Y in a large dataset is 0.00. Which statement is correct?

Correlation only measures straight-line association. $Y = X^2$ with symmetric X has zero correlation but is completely determined by X.

4. 70% of traffic is desktop (converts 10%) and 30% is mobile (converts 5%). What is the overall conversion rate?

Law of total probability: $0.7\times0.10 + 0.3\times0.05 = 0.07 + 0.015 = 0.085$. The plain average 7.5% ignores that desktop is bigger.

5. 10% of sessions are bots. A detector flags 80% of bots and 20% of humans. A session is flagged. What is the chance it is a bot?

$\dfrac{0.8\times0.1}{0.8\times0.1 + 0.2\times0.9} = \dfrac{0.08}{0.08 + 0.18} = \dfrac{0.08}{0.26} \approx 0.31$. The many humans produce more false flags than the few bots produce true ones.

6. In a Bayesian A/B model, conversions are iid given the rate θ. Before θ is known, are two users' conversions independent?

They are conditionally independent given θ, but θ is a shared unknown (a common cause), so marginally they are dependent: each observation carries information about θ.

Practice problems

A. Two dice. Find $P(\text{at least one 6}\mid\text{sum} = 8)$ and compare it with $P(\text{at least one 6})$.

Sum = 8: (2,6), (3,5), (4,4), (5,3), (6,2), 5 outcomes. Two contain a 6, so $P = 2/5 = 0.4$. Without the condition, $P(\text{at least one 6}) = 11/36 \approx 0.306$. Knowing the sum is 8 raised it.

B. An email campaign: 25% of recipients open it, 20% of openers click, 10% of clickers buy. How many buyers per 10,000 emails?

Chain rule: $P(\text{buy}) = 0.25\times0.20\times0.10 = 0.005$. Per 10,000 emails: $10000\times0.25 = 2500$ open, $2500\times0.2 = 500$ click, $500\times0.1 = 50$ buy.

C. 70% of traffic goes to A (converts 4%) and 30% to B (converts 6%). A buyer's variant label is lost. What is $P(B\mid\text{bought})$?

Pieces: $0.7\times0.04 = 0.028$ and $0.3\times0.06 = 0.018$. Total $0.046$. $P(B\mid\text{bought}) = 0.018/0.046 \approx 0.391$. Even though B converts 50% better, most buyers came from A because A got more traffic.

D. Explain to an interviewer why a test that is "99% accurate" (sensitivity and specificity both 99%) can be wrong for most people who test positive.

Suppose the condition affects 1 in 1,000 people. Of 100,000 people, 100 have it and 99 test positive; of the 99,900 without it, 1% = 999 test positive. So $P(\text{has it}\mid +) = 99/(99 + 999) \approx 0.090$. "99% accurate" describes $P(+\mid\text{has it})$ and $P(-\mid\text{healthy})$; the positive predictive value $P(\text{has it}\mid +)$ also depends on the base rate, and with a rare condition the false positives from the huge healthy group dominate.

E. Show that if A and B are independent, then A and $B^c$ are independent too.

$A$ splits into the non-overlapping pieces $A\cap B$ and $A\cap B^c$, so $P(A\cap B^c) = P(A) - P(A\cap B) = P(A) - P(A)P(B) = P(A)\,(1 - P(B)) = P(A)\,P(B^c)$. That is the definition of independence for $A$ and $B^c$.

F. In the screening example (1%, 90%, 91%), a person tests positive and then negative on an independent repeat test. What is the probability they have the condition?

After the positive: odds $= 1/99 \times 10 = 10/99$ (probability ≈ 0.092). A negative has likelihood ratio $P(-\mid\text{sick})/P(-\mid\text{healthy}) = 0.10/0.91 \approx 0.110$. New odds $= 10/99 \times 0.110 \approx 0.0111$, so $P = 0.0111/1.0111 \approx 0.011$: almost back to the base rate. This assumes the two results are independent given the true state.

Chapter 4.4 · Syllabus Modules 2.1–2.6

Random variables: PMF, PDF, CDF and quantiles

So far we have talked about events: "the user converts", "the sum is at least 10". Most of the time, though, we care about numbers: how many orders tomorrow, how long a delivery takes, how much a user spends. A random variable is the tool that turns random outcomes into numbers. This chapter gives you the four standard ways to describe one (PMF, PDF, CDF and quantiles) and shows exactly which question each one answers.

  • Say precisely what a random variable is, and why a random variable $X$ is not the same as an observed value $x$
  • Tell discrete from continuous variables, and know what "support" means
  • Read a PMF (probabilities of single values) and a PDF (a density: probability per unit, where area is probability and the height can exceed 1)
  • Use the CDF $F(x) = P(X\le x)$ to get the probability of any interval, for steps and for smooth curves
  • Find quantiles and percentiles, for a distribution and for a dataset, and use them for prediction intervals

What is a random variable? (and why $X$ is not $x$) core

Toss two coins. The outcome is something like "heads, tails". That is not a number, so you cannot average it or plot it. Now fix a rule: "count the heads". The rule turns every possible outcome into a number: HH → 2, HT → 1, TH → 1, TT → 0. That rule is a random variable.

Two things are easy to mix up. The rule exists before you toss: it knows every number it could give and how likely each one is. The number you actually get today (say 1, because you tossed TH) is a single observed value. We write the rule with a capital letter, $X$, and an observed value with a small letter, $x$.

The name is a little misleading: a random variable is neither "random" by itself nor a "variable" like in algebra. It is a fixed function. All the randomness comes from which outcome happens.

Three ways to say it:

  • Picture: X is a machine that stamps a number on each possible outcome; x is the number stamped on the outcome that actually happened.
  • Numbers: before the toss, X is 0, 1 or 2 with probabilities 1/4, 1/2, 1/4; after the toss, x = 1.
  • Slogan: capital X is the question "how many heads will I get?"; small x is today's answer.

Two coins. The sample space (all possible outcomes; Chapter 4.2) is $\{HH, HT, TH, TT\}$, each with probability $1/4$. Let $X$ = number of heads.

  1. Apply the rule to each outcome: $X(HH) = 2$, $X(HT) = 1$, $X(TH) = 1$, $X(TT) = 0$.
  2. The event "$X = 1$" is the set of outcomes that give 1: $\{HT, TH\}$. So $P(X = 1) = 1/4 + 1/4 = 1/2$.
  3. Likewise $P(X = 0) = 1/4$ and $P(X = 2) = 1/4$. These three numbers are the distribution of $X$.
  4. You toss and see TH. The observed value is $x = 1$. It is one number, not a distribution.
  5. The same outcomes can carry other rules: $Y$ = "1 if the coins match, else 0" gives $Y(HH) = Y(TT) = 1$ and $Y(HT) = Y(TH) = 0$.

Your world. One visitor session is an outcome. $C$ = "1 if the session converts, else 0" and $R$ = "revenue of the session in dollars" are two random variables on the same session. Your logged table holds their observed values $c_i$ and $r_i$.

A random variable is a function that assigns a real number to every outcome $\omega$ (omega, one outcome) in the sample space $\Omega$:

$$X:\ \Omega \to \mathbb{R}, \qquad \omega \mapsto X(\omega).$$
  • Capital $X$ = the random variable (the rule, before the experiment). Small $x$ = a particular number, such as a value $X$ could take or a value that was observed (a realization).
  • "$X = x$" and "$X \le x$" are events: $\{X = x\} = \{\omega : X(\omega) = x\}$, so $P(X = x)$ makes sense.
  • The distribution of $X$ is the collection of all probabilities $P(X \in \text{some set})$. The PMF, PDF, CDF and quantile function in this chapter are four ways to write it down.
  • A function of a random variable is again a random variable: $2X + 1$, $X^2$, $X_1 + X_2$. In particular the sample mean $\bar X = (X_1 + \dots + X_n)/n$ of data not yet collected is a random variable; the number $\bar x$ you compute afterwards is one of its values (Chapter 4.1, Chapter 4.13).
  • Several random variables measured together, $(X_1, \dots, X_d)$, form a random vector.
Why do we need it?

Outcomes like "heads, tails" or "this user's session" cannot be added, averaged or plotted. Random variables turn them into numbers, so all of algebra and calculus can be used on randomness: means, variances, likelihoods, gradients.

Where is it used?

Every model: a conversion $Y_i\sim$ Bernoulli($\theta$), tomorrow's demand $Y_t$, a residual $\epsilon_t$, a model parameter θ treated as random in Bayesian statistics, an estimator $\hat\theta$ (a function of random data), and every numpyro.sample site.

How is it used?

Name the quantity you care about, decide which numbers it can take, then describe how likely each is (PMF/PDF/CDF). Keep two columns in your head: the random variable you model, and the observed values in your table.

outcomes ω HH HT TH TT rule X = number of heads 2 1 0 P(X = 2) = 1/4 P(X = 1) = 2/4 P(X = 0) = 1/4 observed: x = 1
The random variable X is the whole set of arrows (a rule defined before anything happens), together with the probabilities it inherits. Today's outcome TH follows one arrow and gives the observed value x = 1.

The grid shows all 36 outcomes of two dice; each square shows the number the chosen rule $X$ gives that outcome. The bars on the right are the resulting probabilities $P(X = x)$. Change the rule and watch the same outcomes produce a different distribution. Press Roll: one outcome happens (purple), and the rule gives one observed value $x$. Roll a few times: $X$ stays the same rule, while $x$ changes every time.

"X is 3."

X is a rule with a whole distribution. "The observed value of X is 3" or "$x = 3$" is correct, and "$X = 3$" is an event that may or may not happen, with probability $P(X = 3)$.

"The sample mean is just a number, so it has no distribution."

The number you computed, $\bar x$, is one value. The procedure $\bar X$ applied to a fresh random sample gives a different number each time, so it is a random variable with its own distribution (the sampling distribution, Chapter 5.5). That is why estimates have standard errors.

"A random variable is a variable whose value is random."

It is a fixed function of the outcome. Its value looks random only because the outcome is random.

In NumPyro the line numpyro.sample("y", dist.Bernoulli(probs=theta), obs=y) declares a random variable named "y" with a Bernoulli distribution; the argument obs=y plugs in its observed values (your 0/1 conversion column). Leave out obs (for example when simulating from the prior) and NumPyro draws fresh values of the same random variable instead. In your forecasting model, $Y_t$ = demand on day $t$ is the random variable, the historical demand $y_t$ are its observed values, and the forecast for a future day is a distribution over the values $Y_t$ could take.

"The random variable is the data."

The data are realizations (observed values) of random variables. The random variable is the mechanism that could have produced many different datasets.

Model answer: "A random variable is a function from outcomes to numbers; it carries a distribution. An observed value is one realization of it. I write capital $Y_i$ for the conversion of user $i$ as a Bernoulli($\theta$) random variable and small $y_i$ for the 0 or 1 in my table. The same distinction is why an estimator like $\bar X$ has a sampling distribution while my estimate $\bar x$ is a single number."

Random variable $X:\Omega\to\mathbb{R}$, a rule that puts a number on every outcome.

Capital $X$ = the random quantity (has a distribution). Small $x$ = one value / one observation.

"$X = x$" is an event; functions of random variables ($\bar X$, $X^2$) are random variables too.

Quick check: you roll two dice and record the larger number. Write $P(X = 6)$ for this random variable.

$X = 6$ when at least one die is 6: 11 of the 36 outcomes ((6,1)…(6,6) and (1,6)…(5,6)). So $P(X = 6) = 11/36 \approx 0.306$, the most likely value of the maximum. Select "X = larger die" in the widget to see it.

Discrete and continuous random variables core

Some numbers come from counting: orders today, conversions in a session, items in a cart. They jump from one value to the next (0, 1, 2, …) with gaps in between, like the rungs of a ladder. These are discrete.

Other numbers come from measuring: delivery time, temperature, the exact revenue of a large order. Between any two values there are infinitely many more (20 minutes, 20.1, 20.01, …), like a smooth ramp. These are continuous.

The difference matters because of one strange fact: for a continuous variable, the chance of hitting one exact value is zero. The chance that a delivery takes exactly 20.000000… minutes is 0, even though deliveries around 20 minutes are common. So for continuous variables we talk about intervals ("between 19 and 21 minutes") and use a density, not a list of probabilities.

Three ways to say it:

  • Picture: discrete = a ladder of separate values; continuous = a smooth ramp.
  • Numbers: P(exactly 3 orders) can be 0.15; P(delivery exactly 20 min) is 0, but P(19.5 to 20.5 min) ≈ 0.039.
  • Slogan: discrete variables have probabilities at points; continuous variables have probability only over stretches.

Delivery time $T$ in minutes follows a smooth curve with mean 20 (a Gamma distribution, Chapter 4.10). Shrink the window around 20 minutes and watch the probability:

  1. $P(19.5 \le T \le 20.5) \approx 0.039$ (window 1 minute).
  2. $P(19.95 \le T \le 20.05) \approx 0.0039$ (window 0.1 minute): ten times narrower, ten times smaller.
  3. $P(19.995 \le T \le 20.005) \approx 0.00039$ (window 0.01 minute).
  4. Keep shrinking: the probability goes to 0. So $P(T = 20) = 0$ exactly.
  5. But the ratio probability ÷ window width stays near $0.039$ per minute. That steady ratio is the density at 20 (next-but-one section).

Sorting a few of your variables: converted (0/1): discrete. Orders per day (0, 1, 2, … with no upper limit): discrete. Delivery time: continuous. Revenue per visitor: mixed, because most visitors spend exactly 0 (a lump of probability at one point) and buyers spend a continuous amount.

  • A random variable is discrete if its possible values can be listed, $x_1, x_2, x_3, \dots$ (finitely many, or a countable list like 0, 1, 2, …). Each value can have a positive probability $P(X = x_k)$. It is described by a PMF.
  • A random variable is continuous if $P(X = x) = 0$ for every single number $x$ and probabilities of intervals come from a density (a PDF): $P(a \le X \le b)$ = area under the density curve.
  • The support of $X$ is the set of values it can actually take: where the PMF or the density is positive. Bernoulli: $\{0, 1\}$. Poisson: $\{0, 1, 2, \dots\}$. Exponential: $[0, \infty)$. Normal: all real numbers. Beta: $[0, 1]$.
  • A mixed random variable has both: lumps of probability at some points plus a density elsewhere (revenue: a lump at 0 plus a curve for buyers).
Why do we need it?

The type decides which tools work: sums of probabilities for discrete variables, areas under densities for continuous ones. It also decides which distributions are even allowed: a model whose support does not match the data puts probability on impossible values.

Where is it used?

Choosing a likelihood: Bernoulli/Binomial for yes/no, Poisson or Negative Binomial for counts, Normal or Student-t for real-valued noise, Gamma or Log-Normal for positive amounts, Beta for proportions (Chapter 4.11). SciPy even names methods by type: pmf for discrete, pdf for continuous.

How is it used?

Ask two questions about a column: "is it counted or measured?" and "what values are possible (negative? zero? an upper limit?)". The answers give the type and the support, and the support narrows down the family of distributions.

Discrete: orders per hour 01234 probability sits on points Continuous: delivery time area = P(18 to 22 min) probability is area over stretches Mixed: revenue per visitor P(0) = 0.92 0 a lump at 0 plus a curve
Three kinds of random variable. Left: probability sits on separate values (stems). Middle: probability is spread along a range, and only areas are probabilities. Right: a mix, with a lump of probability at exactly zero (non-buyers) and a density for buyers' spend (numbers illustrative).

2,000 simulated delivery times, grouped into bins of the chosen width. Left: the probability (share of deliveries) in each bin. Right: the same bars divided by the bin width (probability per minute) with the true density curve in orange. Shrink the bin width from 10 to 0.5 minutes: on the left every bar sinks towards 0 (the chance of any single value vanishes), while on the right the bars settle onto the curve. Press New sample to see the noise.

"P(X = 20) = 0, so a delivery time of 20 minutes can never happen."

Some exact value always happens. Zero probability for each single point just means probability lives on stretches, not on points. Ask about intervals, such as $P(19.5 \le T \le 20.5)$.

"The column holds whole numbers, so I must use a discrete model; it is stored as floats, so it is continuous."

Think about what the number is, not how it is stored. Star ratings stored as 4.0 are discrete (and ordered categories). Counts in the thousands are discrete but are sometimes modelled with continuous approximations. Small counts with many zeros need a count model.

"Every variable is either discrete or continuous."

Mixed variables are common in product data: revenue, time spent (many exact zeros), discounts. A pure continuous model ignores the lump at zero; zero-inflated and hurdle models handle it (Chapter 4.8).

Support drives likelihood choice in both projects. Conversions have support $\{0, 1\}$, so the A/B framework uses Bernoulli/Binomial (with a Beta prior on the rate, whose support is $[0, 1]$). Categorical metrics take one of $K$ labels, hence Categorical/Multinomial with a Dirichlet. Daily demand counts have support $\{0, 1, 2, \dots\}$, which is what makes a Negative Binomial a natural option for the forecasting model: a Normal likelihood would give probability to negative and fractional demand, which matters most when counts are small.

Discrete: values can be listed; $P(X = x)$ can be positive; described by a PMF.

Continuous: $P(X = x) = 0$ for every $x$; probabilities are areas under a density. Support = values X can take.

Trap: mixed variables (revenue: lump at 0 + continuous part) are neither.

Quick check: discrete, continuous or mixed? (a) number of support tickets per day, (b) a user's session length in seconds, (c) the discount a user received, where 80% received none.

(a) Discrete (a count, support 0, 1, 2, …). (b) Continuous (measured; in practice also often has a lump at 0 for bounced sessions). (c) Mixed: a lump of 0.8 at exactly 0 plus a spread of positive discounts.

The probability mass function (PMF) core

For a discrete variable, the full story fits in a small table: each possible value and its probability. Think of 1 kilogram of sand (all the probability) split into piles, one pile on each possible value. The PMF says how heavy each pile is. Because all the sand is used, the piles always weigh 1 in total.

Draw it as a bar (or stem) chart: the height of each bar is a probability, so it is between 0 and 1. To get the probability of a group of values ("at least 2 orders"), add up their bars.

Three ways to say it:

  • Picture: a bar chart whose bars add up to 1.
  • Numbers: sum of two dice: P(7) = 6/36, P(2) = 1/36, and all 11 bars add to 36/36.
  • Slogan: for discrete variables, bar height = probability; add bars to get any event.

Sum of two dice, $S$. Count the outcomes for each sum: 2 has 1, 3 has 2, …, 7 has 6, …, 12 has 1.

  1. $p(s) = P(S = s) = \dfrac{6 - |s - 7|}{36}$ for $s = 2, \dots, 12$. Example: $p(7) = 6/36$, $p(4) = 3/36$.
  2. $P(S \le 4) = p(2) + p(3) + p(4) = (1 + 2 + 3)/36 = 6/36 = 1/6$.
  3. $P(S \text{ even}) = (1 + 3 + 5 + 5 + 3 + 1)/36 = 18/36 = 1/2$.

Orders per hour, $N$, at a small shop:

$n$01234
$p(n)$0.200.350.250.150.05
  1. Check: $0.20 + 0.35 + 0.25 + 0.15 + 0.05 = 1$ ✓.
  2. $P(N \ge 2) = 0.25 + 0.15 + 0.05 = 0.45$.
  3. $P(1 \le N \le 3) = 0.35 + 0.25 + 0.15 = 0.75$.
  4. The most likely value (the mode) is 1, with probability 0.35.

The probability mass function of a discrete random variable $X$ is

$$p(x) = P(X = x).$$
  • $0 \le p(x) \le 1$ for every $x$, and $p(x) = 0$ outside the support.
  • $\sum_{x} p(x) = 1$ (sum over the support).
  • For any set $A$ of values: $P(X\in A) = \sum_{x\in A} p(x)$.
  • From data, the empirical PMF is the share of observations at each value: pd.Series(x).value_counts(normalize=True). It wobbles around the true PMF and settles as $n$ grows (Chapter 4.13).
  • In code: scipy.stats.binom(n, p).pmf(k); NumPyro's dist.Poisson(rate).log_prob(k) returns $\log p(k)$.
Why do we need it?

It is the complete description of a discrete random variable: every probability you could ask about it is a sum of PMF values. It is also what a count model's likelihood multiplies together.

Where is it used?

Binomial PMF for "k conversions out of n", Poisson and Negative Binomial PMFs for daily orders, the Categorical PMF in a softmax classifier, and the log-likelihood of any discrete model, $\sum_i \log p(y_i\mid\theta)$.

How is it used?

Tabulate it (or use a formula), check that it sums to 1, then add the bars you need. To check a model, overlay the model's PMF on the empirical PMF of your data and look for systematic gaps (too few zeros, too thin a tail).

Pick how many dice are summed. Orange stems: the exact PMF. Blue bars: the share of your rolls that gave each sum. Press Roll 100 a few times, then Roll 1,000: the bars settle onto the stems. With 1 die the PMF is flat; with 2 it is a triangle; with 3 it already looks like a bell (a first glimpse of the CLT, Chapter 4.13).

Drag the tops of the bars to set $P(N = n)$ for $n = 0, \dots, 5$ orders per hour. The readout adds the bars; it only counts as a PMF when the total is exactly 1, so press Normalize to divide every bar by the total. Then use the two sliders to pick a range: the purple bars are added to give $P(a \le N \le b)$.

"A PMF value can be 1.3 if the value is very likely."

PMF values are probabilities, so each is between 0 and 1, and together they add to exactly 1. (Densities are different, as the next section shows.)

"The most likely value is the average."

The most likely value is the mode (tallest bar). The average (the balance point, Chapter 4.5) can differ: for orders per hour the mode is 1 but the mean is $0(0.2) + 1(0.35) + 2(0.25) + 3(0.15) + 4(0.05) = 1.5$.

"My data's value counts are the PMF."

They are an estimate of it. With 100 rolls the bars miss the true PMF by a few percentage points; the gap shrinks as the sample grows.

Discrete likelihoods in your projects are PMFs: the Binomial PMF for "k of n users converted" in the A/B framework, and the Negative Binomial PMF for a day's demand count in the forecasting model. The log-likelihood of a count model is a sum of $\log p(y_t\mid\theta)$ terms, and since every PMF value is at most 1, each term is at most 0. Check which parameterization your code passes to the PMF (for the Negative Binomial there are several, Chapter 4.8).

PMF: $p(x) = P(X = x)$, with $0\le p(x)\le 1$ and $\sum_x p(x) = 1$.

$P(X\in A) = \sum_{x\in A} p(x)$: add the bars.

Trap: mode (tallest bar) ≠ mean; empirical frequencies only estimate the PMF.

Quick check: using the orders table, what is $P(N \ne 1)$?

Complement rule: $P(N\ne 1) = 1 - p(1) = 1 - 0.35 = 0.65$. (Or add the other bars: $0.20 + 0.25 + 0.15 + 0.05 = 0.65$.)

The probability density function (PDF): area is probability core

For a continuous variable, picture the 1 kilogram of probability-sand spread smoothly along the number line instead of piled on separate points. The density at a point is how thick the sand is there. The probability of a stretch, like "between 18 and 22 minutes", is how much sand lies on that stretch: the area under the density curve.

Thickness is not an amount. If all the sand is squeezed onto a short stretch, it must be piled high: a density can be 2, 5 or 100 without anything being wrong, as long as the total area is 1. A density is "probability per unit of $x$", like population density is "people per square kilometre".

Three ways to say it:

  • Picture: probability is the area under the curve; the height is only how crowded that region is.
  • Numbers: a delivery uniformly between 0 and 0.5 hours has density 2 per hour everywhere, and P(between 0.1 and 0.3 h) = 2 × 0.2 = 0.4.
  • Slogan: density ≠ probability; density × width ≈ probability.

Uniform delivery window. A courier promises delivery within half an hour, equally likely at any moment: $X \sim$ Uniform$(0, 0.5)$ hours.

  1. The density is flat. The total area must be 1, and the width is 0.5, so the height is $1/0.5 = 2$. So $f(x) = 2$ for $0 \le x \le 0.5$ (and 0 elsewhere). A density of 2 is fine.
  2. $P(0.1 \le X \le 0.3)$ = area of a rectangle = width × height $= 0.2 \times 2 = 0.4$.
  3. Same question in minutes: $X \sim$ Uniform$(0, 30)$, height $1/30 \approx 0.033$ per minute, and $P(6 \le X \le 18) = 12 \times \tfrac{1}{30} = 0.4$. The probability is the same; the density number changed because the unit changed.

Waiting for the next order. The waiting time has density $f(x) = 0.1\,e^{-0.1x}$ for $x \ge 0$ minutes (an Exponential distribution with mean 10 minutes, Chapter 4.10).

  1. $f(0) = 0.1$ per minute: the curve is tallest at 0 and decays.
  2. $P(X \le 5)$ is the area from 0 to 5. The CDF in the next section gives it without any hand calculation: $1 - e^{-0.1\times5} = 1 - e^{-0.5} \approx 1 - 0.607 = 0.393$.

A continuous random variable $X$ has probability density function $f$ when, for all $a \le b$,

$$P(a \le X \le b) = \int_a^b f(x)\,dx \quad(\text{the area under } f \text{ between } a \text{ and } b).$$
  • $f(x) \ge 0$ everywhere, and $\int_{-\infty}^{\infty} f(x)\,dx = 1$ (total area 1). There is no upper limit of 1 on $f(x)$ itself.
  • The integral sign $\int$ means "add up the areas of very many very thin strips"; for this chapter you only need that picture. For narrow strips, $P(x \le X \le x + \Delta x) \approx f(x)\,\Delta x$.
  • $P(X = a) = 0$, so for continuous variables $P(a \le X \le b) = P(a \lt X \lt b)$: including or excluding the ends changes nothing.
  • Units: $f$ is measured in "probability per unit of $x$" (per minute, per hour). Changing units rescales the density.
  • We write $f(x)$ or $p(x)$ for densities. In code: scipy.stats.norm(loc, scale).pdf(x) and NumPyro's dist.Normal(loc, scale).log_prob(x) (the log of the density).
Why do we need it?

Continuous variables have zero probability at every point, so a list of probabilities is useless. The density is what remains: it tells you where values are crowded, and its areas give every probability you need.

Where is it used?

Normal and Student-t likelihoods for residuals, Gamma or Log-Normal for positive amounts, Beta priors on conversion rates, kernel density estimates of data (Chapter 4.16), and histograms drawn with density=True.

How is it used?

Plot the density to see the shape (peak, skew, tails). For a probability, never read the height: compute an area, in practice with the CDF, F(b) - F(a). For a likelihood, evaluate the density at each observed value and sum the logs.

height f(x) Δx strip area ≈ f(x)·Δx = P(x ≤ X ≤ x + Δx) total area under the curve = 1 density (probability per unit of x)
Probability is area. A thin strip of width Δx under the curve holds about f(x)·Δx of the probability. The height f(x) is a rate (per unit of x), which is why it can be larger than 1 when the strip is narrow.

Pick a distribution, then drag the two handles $a$ and $b$ along the axis. The purple area is $P(a \le X \le b)$. Turn on Show 10 strips to see the area built from thin rectangles: their total is close to the exact area. Notice the y-axis numbers: they are densities, not probabilities (for the Uniform and Beta examples they go above 1).

Make the distribution narrower with the slider. The curve gets taller: with a Normal of $\sigma = 0.1$ the peak is about 4, far above the dashed line at height 1. Yet the readout's total area stays 1.000, and the probability of the small purple window around the peak never exceeds 1. Switch to Uniform: a box of width 0.25 must have height 4.

"$f(2) = 0.8$, so $P(X = 2) = 0.8$."

For a continuous variable $P(X = 2) = 0$. The value 0.8 is a density: about 0.8 × Δx of probability lies in a tiny window of width Δx around 2.

"A density above 1 means a bug."

Only the total area must be 1. A Normal with σ = 0.1 peaks near 4; a Beta(50, 50) peaks near 8.

"In a density=True histogram the bar heights add up to 1."

The bar areas (height × bin width) add up to 1. Heights are densities and depend on the bin width and the units.

"My log-likelihood is +350, so something is wrong: log-probabilities cannot be positive."

With a continuous likelihood each term is a log density, which is positive whenever the density exceeds 1 (for example a Normal with a small σ evaluated near its mean). Positive log-likelihoods are normal for continuous models; for discrete models (log PMFs) they cannot happen.

Model answer: "A PDF value is a density, probability per unit of x, not a probability. Probabilities are areas under it, $P(a\le X\le b) = F(b) - F(a)$, and only the total area has to be 1, so the height can exceed 1. That is also why a continuous log-likelihood can be positive and why its value depends on the units of the data."

In your forecasting model, a Normal or Student-t likelihood evaluates a density at each observed $y_t$; the Negative Binomial evaluates a PMF. So with the continuous likelihoods, the log-likelihood (and therefore the ELBO) can be positive, and the SVI loss (which NumPyro reports as the negative ELBO) can be negative: not a bug. Units matter too: dividing all $y_t$ by a scale $c$ (and rescaling the model to match) multiplies every density value by $c$, which shifts the total log-likelihood by $n\log c$. Log-likelihoods or ELBOs computed on differently scaled data are therefore not comparable; only compare them on the same scaling (for example after the same global scaler).

$P(a\le X\le b) = \int_a^b f(x)\,dx$ = area. $f\ge 0$, total area 1, but $f(x)$ itself can be $\gt 1$.

$P(X = a) = 0$; small window: $P \approx f(x)\,\Delta x$. Density units = probability per unit of x.

Trap: never read a density height as a probability; continuous log-likelihoods can be positive.

Quick check: $f(x) = 3x^2$ for $0 \le x \le 1$ (and 0 elsewhere). Is it a valid density? What is $P(X \le 0.5)$?

It is never negative, and its area from 0 to 1 is $[x^3]_0^1 = 1$, so yes (even though $f(1) = 3 \gt 1$). $P(X\le 0.5) = (0.5)^3 = 0.125$: little probability on the left half, because the density is low there.

The cumulative distribution function (CDF) core

Walk along the number line from the far left, carrying a bucket. Every time you pass some probability, scoop it into the bucket. The amount in your bucket when you reach the point $x$ is the CDF at $x$: the total probability of all values up to and including $x$.

The bucket only fills up, never empties, so the CDF only goes up, from 0 on the far left to 1 on the far right. For a discrete variable, probability sits on points, so the bucket jumps at each value: the CDF is a staircase. For a continuous variable, probability is spread smoothly, so the CDF is a smooth ramp, steep where the density is high.

Three ways to say it:

  • Picture: a running total of probability from the left: stairs for discrete, a ramp for continuous.
  • Numbers: for the sum of two dice, F(4) = P(S ≤ 4) = 6/36 and F(8) = 26/36, so P(5 ≤ S ≤ 8) = 20/36.
  • Slogan: F(x) = "probability of at most x"; any interval is a difference of two F values.

Sum of two dice (a staircase).

  1. $F(4) = P(S\le4) = (1+2+3)/36 = 6/36$.
  2. $F(4.5) = 6/36$ too: no probability lives between 4 and 5, so the staircase is flat there.
  3. $F(8) = (1+2+3+4+5+6+5)/36 = 26/36$.
  4. $P(5\le S\le 8) = P(4 \lt S \le 8) = F(8) - F(4) = 26/36 - 6/36 = 20/36 \approx 0.556$.
  5. Careful with $\lt$: $P(S \lt 8) = F(7) = 21/36$, but $P(S \le 8) = 26/36$. The difference is $p(8) = 5/36$, the height of the step at 8.

Waiting time for the next order (a ramp), mean 10 minutes: $F(x) = 1 - e^{-x/10}$ for $x \ge 0$.

  1. $F(5) = 1 - e^{-0.5} = 1 - 0.6065 = 0.3935$.
  2. $F(15) = 1 - e^{-1.5} = 1 - 0.2231 = 0.7769$.
  3. $P(5 \lt X \le 15) = F(15) - F(5) = 0.7769 - 0.3935 = 0.3834$.
  4. Wait longer than 30 minutes: $P(X \gt 30) = 1 - F(30) = e^{-3} \approx 0.0498$.

The cumulative distribution function of any random variable $X$ is

$$F(x) = P(X \le x) \quad\text{for every real number } x.$$
  • It never goes down, it starts at 0 on the far left ($F(x)\to 0$ as $x\to-\infty$) and ends at 1 on the far right ($F(x)\to1$ as $x\to\infty$).
  • Intervals: $P(a \lt X \le b) = F(b) - F(a)$. Tails: $P(X \gt x) = 1 - F(x)$, called the survival function $S(x)$ (sf in SciPy).
  • Discrete: $F(x) = \sum_{k\le x} p(k)$; a staircase that jumps by $p(k)$ at each value $k$ and is flat in between. At a jump, $F(k)$ takes the upper value (the filled dot): the CDF is right-continuous. Here $\lt$ and $\le$ differ: $P(X \lt k) = F(k) - p(k)$.
  • Continuous: $F(x) = \int_{-\infty}^{x} f(t)\,dt$, a smooth ramp with no jumps. Its slope is the density: $F'(x) = f(x)$ (slopes are derivatives, Calculus Chapter 2.3). Steep ramp = crowded values.
  • Empirical CDF of data $x_1,\dots,x_n$: $\hat F_n(t) = \dfrac{\#\{i : x_i \le t\}}{n}$, the share of observations at or below $t$. A staircase with a step of $1/n$ at each data point.
Why do we need it?

One function answers every "at most", "more than" and "between" question, for discrete and continuous variables alike. You never have to add bars or compute areas by hand: two CDF values and a subtraction are enough.

Where is it used?

Tail probabilities (p-values, risk of exceeding capacity), service levels ("95% of deliveries within 40 minutes"), quantiles (the next section inverts it), Kolmogorov–Smirnov tests, PIT histograms for checking forecasts (Chapter 7.16), and empirical CDF plots for comparing groups.

How is it used?

stats.expon(scale=10).cdf(15) - stats.expon(scale=10).cdf(5) for an interval; .sf(x) for an upper tail (more accurate than 1 - cdf far in the tail); from samples, np.mean(draws <= x) is the empirical CDF at $x$.

01 01234 jump = p(1) = 0.35 flat between values filled dot = value of F at the jump orders per hour n
The CDF of "orders per hour" (PMF 0.20, 0.35, 0.25, 0.15, 0.05). It jumps by p(n) at each possible value and is flat in between. At a jump, F takes the upper (filled) value, because F(n) = P(N ≤ n) includes n itself.

Top: the PMF or density. Bottom: the CDF. Drag the blue handle $b$ and read $F(b) = P(X \le b)$ where the dashed line meets the curve; the top plot shades the same probability. Then drag the orange handle $a$ to the right: the shaded probability is now $P(a \lt X \le b) = F(b) - F(a)$, the purple rise on the CDF. For a discrete distribution, move $b$ between two values (e.g. 4.5): the staircase is flat, nothing changes.

Blue staircase: the empirical CDF of your sample (each data point adds a step of $1/n$; the ticks at the bottom are the data). Orange: the true CDF. The red segment marks the largest vertical gap between them. Start with $n = 20$ and press New sample several times: the gap is often 0.1 or more. Slide $n$ up to 500: the staircase hugs the curve and the gap shrinks.

"For counts, $P(X \lt 8)$ and $P(X \le 8)$ are the same."

They differ by $p(8)$, the height of the step at 8. Only for continuous variables does including or excluding an endpoint not matter.

"$F(x)$ is the probability of $x$."

$F(x)$ is the probability of at most $x$, everything to the left included. It is always between 0 and 1 (unlike a density) and never decreases.

"To get an upper tail, always compute 1 - cdf(x)."

Mathematically fine, but far in the tail cdf(x) rounds to 1.0 and the subtraction returns 0. Use sf(x) (survival function) for tiny tail probabilities, and logsf for even smaller ones.

CDFs answer the questions stakeholders actually ask. In your forecasting model, "what is the chance demand exceeds capacity $c$ next Monday?" is $1 - F(c)$ for the predictive distribution; with posterior predictive draws you estimate it as the share of draws above $c$, i.e. one minus the empirical CDF at $c$ (Chapter 7.14). In the A/B framework, $P(\theta_B - \theta_A \gt 0\mid D)$ is one minus the CDF of the lift at 0, again computed as the share of posterior draws where B beats A (Chapter 6.4).

$F(x) = P(X\le x)$; rises from 0 to 1, never down. $P(a\lt X\le b) = F(b) - F(a)$; $P(X\gt x) = 1 - F(x)$ (sf).

Discrete: staircase, jump $= p(x)$, $P(X\lt k) = F(k) - p(k)$. Continuous: smooth ramp, slope $F' = f$.

Empirical CDF: share of data ≤ t. Trap: $\lt$ vs $\le$ for discrete variables.

Quick check: using the orders-per-hour PMF (0.20, 0.35, 0.25, 0.15, 0.05), find $F(2)$, $F(2.7)$ and $P(N \gt 2)$.

$F(2) = 0.20 + 0.35 + 0.25 = 0.80$. $F(2.7) = 0.80$ as well (no values between 2 and 3: flat stair). $P(N\gt 2) = 1 - F(2) = 0.20$ (check: $0.15 + 0.05$).

Quantiles and percentiles: the CDF read backwards core

The CDF answers: "given a value, what fraction of outcomes lies at or below it?" Very often we need the reverse: "given a fraction, which value has that fraction below it?" A delivery company promises: "90% of orders arrive within 33 minutes." The number 33 is the 90th percentile (the 0.9 quantile) of delivery time.

On a CDF plot, you find it by reading sideways: start at 0.9 on the vertical axis, walk right until you hit the curve, then drop straight down to the horizontal axis. That is why the quantile function is called the inverse CDF. The median is the 0.5 quantile: half below, half above.

Three ways to say it:

  • Picture: read the CDF from the vertical axis to the horizontal axis.
  • Numbers: for waiting times with mean 10 minutes, the median is 6.9 minutes and the 90th percentile is 23 minutes.
  • Slogan: CDF: value → fraction. Quantile: fraction → value.

From a formula. Waiting time with $F(x) = 1 - e^{-x/10}$. Solve $F(q) = p$ for $q$:

  1. $1 - e^{-q/10} = p \;\Rightarrow\; e^{-q/10} = 1 - p \;\Rightarrow\; q = -10\ln(1-p)$.
  2. Median ($p = 0.5$): $q = -10\ln 0.5 = 10\ln 2 \approx 6.93$ minutes. Smaller than the mean of 10: a few very long waits pull the mean up, but not the median.
  3. 90th percentile ($p = 0.9$): $q = -10\ln 0.1 \approx 23.03$ minutes.
  4. For a Normal: $q(0.975) = \mu + 1.96\,\sigma$ and $q(0.025) = \mu - 1.96\,\sigma$, the familiar "±1.96σ".

From a discrete distribution (orders per hour, CDF 0.20, 0.55, 0.80, 0.95, 1). The quantile is the smallest value whose CDF reaches p: the median is 1 (since $F(0) = 0.20 \lt 0.5 \le F(1) = 0.55$) and the 90th percentile is 3 ($F(2) = 0.80 \lt 0.9 \le F(3) = 0.95$).

From data. Ten delivery times, sorted: 12, 15, 17, 18, 20, 22, 25, 28, 35, 60 minutes. NumPy's default method counts positions from 0 to $n-1 = 9$ and interpolates:

  1. Median: position $h = 9\times0.5 = 4.5$, halfway between the values at positions 4 and 5: $(20 + 22)/2 = 21$.
  2. 90th percentile: $h = 9\times0.9 = 8.1$, so take the value at position 8 (35) plus 0.1 of the way to position 9 (60): $35 + 0.1\times(60 - 35) = 37.5$.
  3. Quartiles: $h = 2.25$ gives $17 + 0.25\times1 = 17.25$ (Q1); $h = 6.75$ gives $25 + 0.75\times 3 = 27.25$ (Q3). Interquartile range IQR $= 27.25 - 17.25 = 10$.
  4. The mean is 25.2, above the median 21, pulled up by the 60-minute delivery.

For $0 \lt p \lt 1$, the $p$-quantile of $X$ is

$$q(p) = F^{-1}(p) = \min\{x : F(x) \ge p\},$$

the smallest value whose CDF reaches $p$. For a continuous, strictly increasing CDF this is just the solution of $F(q) = p$, so $P(X \le q(p)) = p$.

  • The 100p-th percentile is the same thing in percent: the 0.9 quantile = the 90th percentile.
  • Median $= q(0.5)$. Quartiles: $Q_1 = q(0.25)$, $Q_3 = q(0.75)$. IQR $= Q_3 - Q_1$, a spread measure that ignores the extreme 25% on each side.
  • Sample quantiles estimate $q(p)$ from data. There are several conventions; NumPy's default (method="linear", also pandas' and R's default) sorts the data, computes the position $h = (n-1)p$ counting from 0, and interpolates between the two nearest sorted values. np.percentile(x, 90) = np.quantile(x, 0.9).
  • In SciPy the quantile function is ppf ("percent point function"): stats.norm.ppf(0.975) = 1.96. isf(p) is the inverse survival function, $q(1-p)$.
  • Inverse-transform sampling: if $U$ is uniform on $(0, 1)$, then $q(U)$ has CDF $F$. Uniform random numbers pushed through the quantile function become draws from any distribution you like.
Why do we need it?

Many decisions are about thresholds, not averages: "how much stock covers 95% of days?", "what latency do 99% of requests beat?". Quantiles give those thresholds directly, and they are robust to a few extreme values, unlike the mean.

Where is it used?

Prediction intervals (5% and 95% quantiles of a forecast), equal-tailed credible intervals, critical values of tests (1.96), box plots and the IQR outlier rule (Chapter 4.14), Q-Q plots (Chapter 4.17), quantile (pinball) loss, SLAs such as p95/p99 latency.

How is it used?

For a named distribution, call dist.ppf(p). For data or for posterior or predictive draws, call np.quantile(draws, [0.05, 0.5, 0.95]). Report the median with a quantile interval when data are skewed.

01 0.9 q(0.9) ≈ 23 min 1. start at p = 0.9 on the vertical axis 2. go right until you hit the CDF 3. drop down: that x is the quantile F(x), waiting time with mean 10 min
The CDF maps a value to a probability; the quantile function maps a probability back to a value. 90% of waits are at most about 23 minutes.

Drag the purple handle up and down the left edge to choose $p$. The arrow runs right to the CDF and down to the quantile $q(p)$. Compare the Exponential (median well below the mean) with the Normal (median = mean), and try $p = 0.975$ on the Normal: $q \approx 1.96$. On a discrete distribution, slide $p$ slowly: $q(p)$ jumps from one value to the next, because it is the smallest value whose CDF reaches $p$.

Ten delivery times (one per row). Choose $p$ with the slider; the purple line is the sample $p$-quantile (NumPy's default method) and the readout shows the arithmetic. The box at the bottom spans Q1 to Q3 with the median inside. Now drag the slowest delivery (60 min) far to the right, or press Make an outlier: the mean (green) moves, while the median and the box do not move at all. Set $p = 0.9$ and repeat: the 90th percentile moves too, because it sits next to the largest value. Quantiles in the middle are robust; quantiles near the edges of a small sample are not.

Each draw picks a uniform random number $u$ between 0 and 1 (on the vertical axis), runs right to the CDF and drops down to $x = q(u)$. Press Draw 1 a few times to watch single draws, then Draw 200 several times: the histogram below fills in the density. Where the CDF is steep, many $u$'s land on a short stretch of $x$, which is exactly where the density is high. This is one standard way to turn uniform random numbers into draws from any distribution (libraries often use faster special-purpose algorithms, but the resulting distribution is the same).

"The 90th percentile of delivery time is 90% of the longest delivery."

It is the time that 90% of deliveries do not exceed. It has nothing to do with the maximum.

"Every tool gives the same percentile for the same data."

With small samples, conventions differ. For the ten deliveries above, the 90th percentile is 37.5 with NumPy's default, 47.5 with method="hazen" and 57.5 with method="weibull". State the method when it matters; with large samples the differences fade.

"The 90th percentile of weekly demand is the sum of the seven daily 90th percentiles."

Quantiles do not add up. Summing daily P90s gives a number that is usually too high, because it assumes every day is unusually high at once. For a weekly quantile, sum each simulated sample path over the week first, then take the quantile of those weekly totals.

In a Bayesian forecasting model like yours, the prediction interval for each day can come from quantiles of the posterior predictive draws, for example jnp.quantile(draws, jnp.array([0.05, 0.5, 0.95]), axis=0): the 5% and 95% quantiles are the band, the 0.5 quantile is the median forecast. Equal-tailed credible intervals for a lift in the A/B framework are the 2.5% and 97.5% quantiles of its posterior draws (Chapter 6.4). When you report weekly or monthly totals, aggregate the draws first and take quantiles last. Quantile forecasts are scored with the pinball loss (Chapter 7.16).

$q(p) = F^{-1}(p) = \min\{x : F(x)\ge p\}$; percentile = $100p$. Median $= q(0.5)$; IQR $= q(0.75) - q(0.25)$.

Sample quantile (NumPy default): sort, $h = (n-1)p$, interpolate. SciPy: ppf.

Trap: quantiles do not add; methods differ for small n; the median ignores extremes, the mean does not.

Quick check: for a Normal with mean 100 and standard deviation 15, what is the 97.5th percentile? What fraction of values lies between the 2.5th and 97.5th percentiles?

$q(0.975) = 100 + 1.96\times15 = 129.4$. By definition 97.5% − 2.5% = 95% of values lie between the two percentiles ($100 - 29.4 = 70.6$ and $129.4$).

Recap, cheat sheet and practice

  • A random variable $X$ is a rule that puts a number on every outcome; an observed value $x$ is one realization. Estimators like $\bar X$ are random variables; estimates like $\bar x$ are numbers.
  • Discrete variables (counts) have probabilities at points; continuous variables (measurements) have $P(X = x) = 0$ and probabilities over intervals. The support is the set of possible values.
  • PMF $p(x) = P(X = x)$: bars between 0 and 1 that add to 1.
  • PDF $f(x)$: a density (probability per unit). Area = probability; the height can exceed 1; log densities can be positive.
  • CDF $F(x) = P(X\le x)$: a staircase or a ramp from 0 to 1; $P(a\lt X\le b) = F(b) - F(a)$.
  • Quantile $q(p) = F^{-1}(p)$: the value with a fraction $p$ below it. Median, quartiles, percentiles; sample quantiles depend on a method.

Cheat sheet

ObjectFormulaAnswers the questionSciPy
PMF (discrete)$p(x) = P(X = x)$how likely is exactly $x$?.pmf(x)
PDF (continuous)$P(a\le X\le b) = \int_a^b f$where are values crowded? (area = probability).pdf(x)
CDF$F(x) = P(X\le x)$how likely is at most $x$?.cdf(x)
Survival function$1 - F(x) = P(X\gt x)$how likely is more than $x$?.sf(x)
Interval$F(b) - F(a) = P(a\lt X\le b)$how likely is between $a$ and $b$?.cdf(b) - .cdf(a)
Quantile$q(p) = \min\{x: F(x)\ge p\}$which value has a fraction $p$ below it?.ppf(p)
Sample quantilesort, $h = (n-1)p$, interpolatethe same, from datanp.quantile(x, p)
Code it · Python

import numpy as np
from scipy import stats
from scipy.integrate import quad

rng = np.random.default_rng(0)

# 1) A random variable is a rule; the numbers you record are its observed values
coins = rng.integers(0, 2, size=(5, 2))          # 5 experiments, 2 coins each (1 = heads)
x = coins.sum(axis=1)                             # X = number of heads, applied to each outcome
print(x)                                          # five observed values of X (they depend on the seed)

# 2) PMF: sum of two dice, exactly (convolution) and by simulation
die = np.ones(6) / 6
pmf = np.convolve(die, die)                       # P(S = 2), ..., P(S = 12)
print((pmf * 36).round().astype(int))             # [1 2 3 4 5 6 5 4 3 2 1] outcomes out of 36
rolls = rng.integers(1, 7, size=(10_000, 2)).sum(axis=1)
print(np.mean(rolls == 7).round(3), pmf[5].round(3))   # about 0.17 (simulated) vs 0.167 (exact, 6/36)

orders = stats.rv_discrete(values=([0, 1, 2, 3, 4], [0.2, 0.35, 0.25, 0.15, 0.05]))   # orders per hour
print(orders.pmf(1), orders.cdf(2), orders.sf(2).round(2), orders.ppf(0.5), orders.ppf(0.9))   # 0.35 0.8 0.2 1.0 3.0

# 3) PDF: a density is not a probability; areas are
U = stats.uniform(loc=0, scale=0.5)               # Uniform(0, 0.5) hours; SciPy: loc = start, scale = width
print(U.pdf(0.2), round(U.cdf(0.3) - U.cdf(0.1), 4))   # 2.0 (a density above 1) and 0.4 (a probability)
narrow = stats.norm(loc=0, scale=0.1)             # scale is the standard deviation, not the variance
print(narrow.pdf(0).round(3), round(quad(narrow.pdf, -1, 1)[0], 6))   # 3.989 but total area 1.0
print(narrow.logpdf(0).round(4))                  # 1.3836: a log density can be positive

# 4) CDF, survival function, intervals: waiting time with mean 10 minutes
W = stats.expon(scale=10)                         # SciPy: scale = mean = 1 / rate
print(W.cdf(5).round(4), (W.cdf(15) - W.cdf(5)).round(4), W.sf(30).round(4))   # 0.3935 0.3834 0.0498

# 5) Quantiles: ppf is the inverse CDF
print(W.ppf([0.5, 0.9]).round(2))                 # [ 6.93 23.03]
print(stats.norm.ppf(0.975).round(4))             # 1.96

times = np.array([12, 15, 17, 18, 20, 22, 25, 28, 35, 60])
print(np.quantile(times, [0.25, 0.5, 0.75, 0.9])) # [17.25 21.   27.25 37.5 ]  (NumPy default: method="linear")
print(np.percentile(times, 90, method="hazen"))   # 47.5: another convention, another answer for small n

# 6) Inverse-transform sampling and the empirical CDF
draws = W.ppf(rng.random(100_000))                # uniform numbers pushed through the quantile function
print(np.mean(draws <= 5).round(3), draws.mean().round(2))   # about 0.393 and about 10
Test yourself

1. Which of these is a random variable?

A random variable is a function from outcomes to numbers. The 3 in your table is an observed value, the sample space is the set of outcomes, and $P(X = 1)$ is a probability.

2. A continuous model gives $f(2) = 1.7$. What follows?

A density is probability per unit of x. Only the total area must be 1. Probabilities of single points are 0 for continuous variables; a small window of width Δx around 2 has probability about 1.7·Δx.

3. $S$ is the sum of two dice with CDF $F$. Which expression equals $P(5 \le S \le 8)$?

$F(8) - F(4) = P(4 \lt S \le 8)$, and for whole-number values "more than 4" is "at least 5". $F(8) - F(5)$ would drop $S = 5$; $F(7) - F(4)$ would drop $S = 8$.

4. Which list is a valid PMF on the values 0, 1, 2?

PMF values must each be between 0 and 1 and add to 1. 0.5 + 0.6 − 0.1 = 1 but has a negative value; 0.2 + 0.3 + 0.4 = 0.9; 1.2 exceeds 1.

5. Waiting times are Exponential with mean 10 minutes. The median wait is…

Solve $1 - e^{-q/10} = 0.5$: $q = 10\ln 2 \approx 6.93$. The right tail of long waits pulls the mean (10) above the median. 23 minutes is the 90th percentile.

6. What does np.quantile([1, 2, 3, 4, 5, 6, 7, 8, 9, 10], 0.9) return?

Default method: position $h = (n-1)p = 9\times0.9 = 8.1$ (counting from 0). The value at position 8 is 9 and at position 9 is 10, so $9 + 0.1\times(10 - 9) = 9.1$.

Practice problems

A. Two dice; $X$ = the larger of the two numbers. Find the PMF of $X$ and check that it adds to 1.

$X \le k$ means both dice are at most $k$: $P(X\le k) = (k/6)^2 = k^2/36$. So $P(X = k) = P(X\le k) - P(X\le k-1) = \dfrac{k^2 - (k-1)^2}{36} = \dfrac{2k-1}{36}$: that is 1, 3, 5, 7, 9, 11 out of 36 for $k = 1,\dots,6$. Sum: $1+3+5+7+9+11 = 36$, so the total is $36/36 = 1$ ✓. Notice we found the PMF through the CDF.

B. A conversion rate has density $f(x) = 2x$ on $[0, 1]$. Find $F(x)$, $P(X \le 0.5)$ and the median. Is $f(1) = 2$ a problem?

Area from 0 to $x$ under the line $2t$ is a triangle: $F(x) = \tfrac12\cdot x\cdot 2x = x^2$. So $P(X\le0.5) = 0.25$. Median: $x^2 = 0.5$, so $x = \sqrt{0.5}\approx 0.707$. $f(1) = 2$ is fine: it is a density, and the total area $F(1) = 1$.

C. Orders per hour have PMF 0.20, 0.35, 0.25, 0.15, 0.05 on 0–4. Use the CDF to find $P(1\le N\le3)$, $P(N \lt 3)$ and the median.

CDF: $F(0) = 0.20$, $F(1) = 0.55$, $F(2) = 0.80$, $F(3) = 0.95$, $F(4) = 1$. $P(1\le N\le 3) = F(3) - F(0) = 0.95 - 0.20 = 0.75$. $P(N\lt3) = F(2) = 0.80$. Median: the smallest $n$ with $F(n)\ge0.5$ is $n = 1$.

D. Explain to an interviewer the difference between a PMF value and a PDF value, and why a continuous model's log-likelihood can be positive.

"A PMF value is a probability, $P(X = x)$, so it lies in $[0, 1]$ and its log is at most 0. A PDF value is a density, probability per unit of $x$; probabilities are areas $\int_a^b f$, and only the total area must be 1, so the density can exceed 1 where values are tightly packed. The log-likelihood of a continuous model sums log densities, so whenever densities exceed 1 (small noise scale, data measured in large units) the sum can be positive. Rescaling the data changes the log-likelihood by a constant, so I only compare log-likelihoods computed on the same data scale."

E. Waiting time to the next order is Exponential with mean 10 minutes. Which waiting time is exceeded only 5% of the time?

We need the 0.95 quantile: $q = -10\ln(1 - 0.95) = -10\ln 0.05 \approx 29.96$, about 30 minutes. Check with the survival function: $P(X\gt30) = e^{-3}\approx 0.0498 \approx 5\%$ ✓.

F. Revenue for five visitors: 0, 0, 0, 40, 100 dollars. Find the mean, the median and the 80th percentile (NumPy's default). What kind of random variable is revenue?

Mean $= 140/5 = 28$. Median: position $4\times0.5 = 2$, value 0. 80th percentile: position $4\times0.8 = 3.2$, so $40 + 0.2\times(100 - 40) = 52$. Revenue is a mixed variable: a lump of probability at exactly 0 (non-buyers) plus a continuous spread for buyers, which is why the median (0) and the mean (28) tell such different stories.

Chapter 4.5 · Syllabus Modules 2.7–2.9

Expected value, variance and standard deviation

A whole distribution is a lot to carry around. Most of the time two numbers do the job: where the values sit on average (the expected value) and how far they usually stray from that average (the variance and the standard deviation). These two numbers are inside every model you have built: the mean of a likelihood, the noise level σ, the uncertainty of a difference between two variants, and the scaler you apply before fitting.

  • See the expected value $E[X]$ as the balance point of a distribution and as the long-run average of many repeats
  • Use linearity of expectation: averages of sums are sums of averages, even when the parts depend on each other
  • Know that $E[g(X)] \ne g(E[X])$ for a curved $g$ (a first look at Jensen's inequality)
  • Read the variance as the average squared distance from the mean, and the standard deviation as the same thing back in the original units
  • Use the rules $Var(aX+b) = a^2Var(X)$ and "variances of independent pieces add"
  • Explain, with a simulation and a short proof, why the sample variance divides by $n-1$

Expected value: the balance point of a distribution core

What you need first: a random variable $X$ is a rule that turns each random outcome into a number, and its PMF $p(x) = P(X = x)$ lists how likely each value is (Chapter 4.4). The sum sign $\sum$ means "add up" (Linear Algebra guide, 1.1).

Take a long ruler and put small stacks of coins on it: a tall stack at mark 1, a smaller one at mark 2, smaller still at 3 and 4. Now try to balance the ruler on one finger. There is exactly one spot where it stays level. That spot is the expected value.

A probability distribution is just like that ruler. The values are the marks, and the probabilities are the weights of the stacks. Heavy stacks (likely values) pull the balance point toward them. Far-away stacks pull harder, because they sit on a longer lever.

Three ways to say it:

  • Picture: the expected value is where the distribution balances, like a seesaw with weights on it.
  • Numbers: if a basket holds 1, 2, 3 or 4 items with probabilities 0.4, 0.3, 0.2, 0.1, the balance point is 2.0 items.
  • Slogan: the expected value is the centre of mass of the probabilities.

Items in a basket. Let $X$ = the number of items in the next basket in an online shop. From past data: $P(X{=}1)=0.4$, $P(X{=}2)=0.3$, $P(X{=}3)=0.2$, $P(X{=}4)=0.1$. (They add to 1, as they must.)

  1. Multiply each value by its probability: $1\times0.4 = 0.4$, $\;2\times0.3 = 0.6$, $\;3\times0.2 = 0.6$, $\;4\times0.1 = 0.4$.
  2. Add them: $0.4 + 0.6 + 0.6 + 0.4 = 2.0$. So $E[X] = 2.0$ items.
  3. Check the balance. Distances from 2.0 are $-1, 0, +1, +2$. Weight each by its probability: $-1\times0.4 = -0.4$, $0$, $+1\times0.2 = +0.2$, $+2\times0.1 = +0.2$.
  4. Left pull $0.4$, right pull $0.2 + 0.2 = 0.4$. They cancel: the seesaw is level at 2.0.

A fair die. Each face 1–6 has probability $1/6$: $E[X] = (1+2+3+4+5+6)/6 = 21/6 = 3.5$. Notice that a die can never show 3.5. The expected value does not have to be a possible value.

For a discrete random variable $X$ with PMF $p(x)$, the expected value (also called the expectation or the mean of $X$) is

$$E[X] \;=\; \sum_{x} x\,p(x),$$

where the sum runs over every value $x$ that $X$ can take (this set of possible values is called the support of $X$). We often write $\mu$ ("mu") or $\mu_X$ for it.

  • Balance property: $\sum_x (x - \mu)\,p(x) = 0$. The probability-weighted distances to the left and right of $\mu$ cancel exactly.
  • $E[X]$ lies between the smallest and the largest possible value. It need not be a possible value itself (the die's 3.5).
  • For a constant $c$, $E[c] = c$.
  • $E[X]$ is a fixed property of the distribution. It is not computed from data. The data version is the sample mean $\bar x$, which changes from sample to sample.
  • For a yes/no variable ($X = 1$ with probability $p$, else $0$): $E[X] = 0\times(1-p) + 1\times p = p$. The expected value of a 0/1 variable is the probability of a 1.
Why do we need it?

A distribution lists many numbers; we need one number for "the typical size" to compare options, set prices, plan stock and make decisions. The expected value is that number, and it is the one that adds up correctly over many repeats.

Where is it used?

The mean of every likelihood (Normal $\mu$, Poisson $\lambda$, Bernoulli $p$), posterior means in your A/B framework, point forecasts, expected revenue per visitor, expected loss in decision rules, and every loss function in machine learning (the training loss is an average).

How is it used?

For a PMF: multiply each value by its probability and add (np.sum(x * p)). For data: take the sample mean x.mean() as an estimate of it. For a model: draw many samples and average them (Monte Carlo).

0.4 0.3 0.2 0.1 1 item 3 items 4 items balance point E[X] = 2 left pull: 0.4 × 1 = 0.4 right pull: 0.2 × 1 + 0.1 × 2 = 0.4
The probabilities are weights on a ruler. The ruler balances at the expected value, where the weight-times-distance on the left equals the weight-times-distance on the right.

The blue bars are the weights (how likely each value is). Drag the purple fulcrum under the beam until the beam is level: you have found $E[X]$. Then drag a bar up or down and watch the beam tip. Try Long right tail: a few far-right values drag the balance point to the right even though most of the weight is on the left. Try Two extremes: the balance point sits in the middle, where there is no weight at all.

"The expected value is the most likely value."

The most likely value is the mode. In the basket example the mode is 1 item, but the expected value is 2 items. They agree only for symmetric, single-peaked distributions.

"The expected value is a value we expect to see."

It is a balance point, not a forecast of one outcome. A die never shows 3.5, and a conversion (0 or 1) never equals 0.05.

"The mean and the median are the same thing."

The median splits the probability in half; the mean balances probability × distance. A long tail pulls the mean a lot and the median only a little (try "Long right tail").

In your A/B framework the posterior mean of a conversion rate is $E[\theta \mid D]$: the balance point of the posterior distribution of $\theta$ after seeing the data $D$ (Bayesian updating is taught in Chapter 6.1). In your forecasting model, a common point forecast is the mean of the predictive draws for a day, $E[y_t \mid D]$; some code reports the median instead, so check which one yours returns. Both are "one number from a distribution", and they differ when the distribution is skewed.

"$E[X]$ and $\bar x$ are the same thing."

$E[X]$ is a fixed number that belongs to the distribution (the process). $\bar x$ is computed from one sample and changes from sample to sample. $\bar x$ is an estimate of $E[X]$.

Model answer: "The expected value is the probability-weighted average of all possible values, a property of the distribution. The sample mean is what I compute from data; by the law of large numbers it gets close to the expected value as the sample grows."

$E[X] = \sum_x x\,p(x)$: the balance point of the distribution ($\sum (x-\mu)p(x) = 0$).

For a 0/1 variable, $E[X] = P(X=1)$.

Trap: the mean is not the most likely value and need not be a possible value (die: 3.5).

Quick check: a user converts with probability 0.08. Let $X = 1$ if they convert and $0$ if not. What is $E[X]$?

$E[X] = 0\times0.92 + 1\times0.08 = 0.08$. The expected value of a yes/no variable is the probability of "yes". This is why the conversion rate is literally an expected value.

The long-run average, and expected values of continuous variables core

A casino has no idea whether you will win the next spin. But over a million spins, its average profit per spin is almost exactly what the expected value says. That is the second meaning of the expected value: if you could repeat the random experiment again and again and average the results, the average would settle at $E[X]$. (This is the Law of Large Numbers, Chapter 4.13.)

For a continuous variable (a waiting time, an order value) the values are not separate marks on a ruler but a smooth shape: the density. The idea does not change. The expected value is where that shape would balance if you cut it out of cardboard.

Three ways to say it:

  • Picture: run the experiment many times; the running average wobbles, then flattens out at $E[X]$.
  • Numbers: if 5% of visitors buy and the average order is 40 dollars, the shop earns about 2 dollars per visitor in the long run.
  • Slogan: one draw is a surprise; the average of many draws is predictable.

Revenue per visitor. Let $R$ = the revenue from one visitor. With probability $0.95$ the visitor buys nothing ($R = 0$). With probability $0.05$ they buy, and an order is worth 40 dollars on average.

  1. Non-buyers contribute $0.95 \times 0 = 0$.
  2. Buyers contribute $0.05 \times 40 = 2$.
  3. $E[R] = 0 + 2 = 2$ dollars per visitor. For 10 000 visitors, expect about $10\,000 \times 2 = 20\,000$ dollars.

(We used "the average order given that the visitor buys". Averages inside a group like this are conditional expectations, the topic of Chapter 4.6.)

A lottery ticket. It costs 2 and pays 100 with probability $0.01$. Profit $X = 98$ with probability $0.01$, and $X = -2$ with probability $0.99$. So $E[X] = 98\times0.01 + (-2)\times0.99 = 0.98 - 1.98 = -1$. On average you lose 1 per ticket; over 1000 tickets, about 1000.

A continuous waiting time. A delivery arrives at a uniformly random time between 0 and 10 minutes, so the density is $f(x) = 1/10$ on $[0, 10]$.

  1. Replace "sum of value × probability" by "integral of value × density": $E[X] = \int_0^{10} x\cdot\tfrac{1}{10}\,dx$.
  2. The integral of $x$ is $x^2/2$: $\;\tfrac{1}{10}\left[\tfrac{x^2}{2}\right]_0^{10} = \tfrac{1}{10}\cdot\tfrac{100}{2}$.
  3. $= \tfrac{1}{10}\times 50 = 5$ minutes, the middle of the interval, as the balance picture predicts.

For a continuous random variable with density $f(x)$:

$$E[X] = \int_{-\infty}^{\infty} x\,f(x)\,dx .$$

Read the integral sign $\int$ as "a sum over very thin slices": cut the number line into tiny pieces of width $dx$; the piece near $x$ has probability about $f(x)\,dx$ (probability = area, Chapter 4.4); multiply each $x$ by that probability and add everything up. It is the continuous version of $\sum x\,p(x)$.

The long-run meaning. If $X_1, X_2, \dots$ are independent draws of the same $X$ ("iid"), then the running average $\bar X_n = (X_1 + \dots + X_n)/n$ gets closer and closer to $E[X]$ as $n$ grows. This needs $E[X]$ to exist.

It can fail to exist. If the tails of a distribution are too heavy, the sum or integral does not converge and there is no expected value. The Cauchy distribution is the classic case: its running averages never settle (Chapter 4.13).

Why do we need it?

Business questions are about many repeats: thousands of visitors, hundreds of days. The expected value turns "one uncertain outcome" into "what happens on average over many", which is what budgets, prices and capacity plans need.

Where is it used?

Revenue per visitor and average order value in A/B tests, expected demand in forecasting, the price of insurance and bets, Monte Carlo estimates (average many simulated draws), and stochastic gradient descent, whose noisy gradients are right on average.

How is it used?

If you know the distribution, compute the sum or integral. If you can only simulate it, draw many samples and average them: samples.mean(). If you only have data, the sample mean estimates it, and more data makes the estimate steadier.

Each step is one visitor: with probability $p$ they buy, and the order value is random with average $v$. The blue line is the running average revenue per visitor. Early on it jumps wildly (one big order moves it a lot). Later it settles near the green line $E[R] = p\times v$. Press New sample a few times, then choose 5 runs: all runs end up near the same value. Lower $p$ to 0.01 and notice it takes much longer to settle: rare buyers make the average noisy.

Pick a shape. Drag the purple fulcrum under the beam until it is level, or press Snap to the mean. Then press Snap to the median: for the symmetric shapes the beam stays level, but for Waiting time and Order value it tips, because the long right tail pulls the mean to the right of the median. In Two segments the mean sits in the valley where almost no values live.

"The mean is where most of the data is."

In a mix of two segments (60% around 3, 40% around 8) the mean is 5, in a valley with almost no values. The mean is a balance point, not a "typical" value.

"Every distribution has a mean."

Very heavy tails can make $\int x f(x)\,dx$ infinite or undefined (Cauchy). Then sample averages never settle, no matter how much data you collect.

"After 100 visitors the average revenue per visitor is reliable."

With rare conversions the running average is very noisy for a long time. How noisy is measured by the variance, later in this chapter.

A metric like revenue per visitor in an A/B test is a mix of many exact zeros and a skewed positive order value. Its expected value has two levers: $E[R] = P(\text{buy}) \times E[\text{order} \mid \text{buy}]$. A variant can raise revenue by converting more people, by raising order size, or both, so it is worth looking at the two parts separately. In your forecasting model, "expected demand on a day" is the mean of the predictive distribution, which you estimate by averaging the model's simulated draws.

Continuous: $E[X] = \int x\,f(x)\,dx$ (balance point of the density).

Long run: the average of many iid draws settles at $E[X]$ (if it exists).

Trap: the mean can sit where no data live (two segments) and can be far from the median (long tails).

Quick check: a waiting time has an Exponential distribution with rate 0.5 per minute. Which is larger, the mean or the median?

The mean is $1/0.5 = 2$ minutes, the median is $2\ln 2 \approx 1.39$ minutes. The long right tail pulls the mean above the median (the widget's "Waiting time" shape).

Linearity of expectation: averages of sums are sums of averages core

Three friends are coming to a party. On average each one brings 2 snacks. On average, how many snacks arrive? Six. You did not need to know whether the friends talk to each other first, or whether one brings more when another brings less. Averages simply add.

Same with scaling: if each snack costs 3 and there is a fixed delivery fee of 5, the average bill is $3 \times (\text{average snacks}) + 5$. Multiply and shift the values, and the average is multiplied and shifted in the same way.

Three ways to say it:

  • Picture: stacking two piles of random size gives a pile whose average height is the sum of the two average heights.
  • Numbers: 20 metrics each have a 5% chance of a false alarm, so on average $20\times0.05 = 1$ false alarm per experiment, correlated or not.
  • Slogan: the expected value of a sum is the sum of the expected values. Always.

(a) Scale and shift. Shipping costs $C = 3 + 2X$ dollars, where $X$ is the number of items (the basket example, $E[X] = 2$).

  1. Rule: $E[3 + 2X] = 3 + 2E[X] = 3 + 2\times2 = 7$.
  2. Check directly: costs are $5, 7, 9, 11$ with probabilities $0.4, 0.3, 0.2, 0.1$, so $E[C] = 2.0 + 2.1 + 1.8 + 1.1 = 7.0$. ✓

(b) Counting with indicators. 1000 visitors arrive: 600 on mobile, each buying with probability $0.03$, and 400 on desktop, each buying with probability $0.06$.

  1. Give each visitor $i$ a 0/1 variable $I_i$ ("did they buy?"). Its expected value is its buying probability.
  2. The number of buyers is $I_1 + I_2 + \dots + I_{1000}$, so the expected number is the sum of the probabilities.
  3. $600\times0.03 + 400\times0.06 = 18 + 24 = 42$ buyers on average. No independence was needed.

(c) Many metrics. An experiment with no real effect checks 20 metrics, each tested at level $\alpha = 0.05$ (each has a 5% chance of a false "significant" result). Expected number of false alarms $= 20 \times 0.05 = 1$, whether the metrics are independent or strongly correlated. (The chance of at least one false alarm is a different question, and it does depend on the correlation: $1 - 0.95^{20} \approx 0.64$ only if they are independent. See Chapter 4.2.)

For any random variables $X, Y$ (with finite means) and constants $a, b$:

$$E[aX + b] = a\,E[X] + b, \qquad E[X + Y] = E[X] + E[Y].$$

More generally $E\left[\sum_i a_i X_i\right] = \sum_i a_i E[X_i]$. This is called linearity of expectation.

  • No independence needed. The rule holds for dependent variables too. (Why, for two discrete variables with joint PMF $p(x,y)$: $E[X+Y] = \sum_{x,y}(x+y)p(x,y) = \sum_{x,y}x\,p(x,y) + \sum_{x,y}y\,p(x,y) = \sum_x x\,p_X(x) + \sum_y y\,p_Y(y)$, where $p_X(x) = \sum_y p(x,y)$ is the marginal PMF of $X$: the distribution of $X$ alone, ignoring $Y$.)
  • Indicator variable: for an event $A$, the variable $\mathbf{1}_A$ is 1 if $A$ happens and 0 if not. $E[\mathbf{1}_A] = P(A)$. A count of events is a sum of indicators, so $E[\text{count}] = \sum P(A_i)$.
  • Not for products: $E[XY] = E[X]\,E[Y]$ is true for independent variables, but not in general.
  • Not for curved functions: $E[X^2] \ne (E[X])^2$ in general (next section).
Why do we need it?

Most quantities we care about are totals or averages of many pieces that depend on each other in messy ways. Linearity lets us get their expected value from the pieces alone, without ever working out the joint distribution.

Where is it used?

Expected number of conversions or false positives, expected weekly demand from daily forecasts, the mean of an additive forecast $g(t) + s(t) + h(t) + X_t\beta$, unbiased stochastic gradients in SGD, and Monte Carlo averages in SVI (the ELBO estimate).

How is it used?

Write the quantity as a sum (often of indicators), take the expected value of each piece, and add. For scaled or shifted variables, apply the scale and shift to the mean directly: E[a*X + b] = a*E[X] + b.

Each visitor is a 0/1 indicator. Its expected value is its buying probability. p = 0.03 p = 0.03 … p = 0.06 p = 0.06 … E[count]= Σ pᵢ 600 mobile 400 desktop 600 × 0.03 + 400 × 0.06 = 18 + 24 = 42 expected buyers, even if visitors influence each other
The indicator trick: write a count as a sum of 0/1 variables, then add up their expected values (their probabilities).

Each run is one experiment with no real effect, checking $m$ metrics at level $\alpha$. The bars show how often a run had 0, 1, 2, … false alarms. The green line is the linearity answer $m\alpha$; the purple line is the simulated average. Slide the correlation between metrics up to 0.8: the purple line stays on the green one (linearity does not care), but the bars change shape: more runs with zero alarms and a few runs with many. Watch "P(at least one)" fall.

"$E[X + Y] = E[X] + E[Y]$ needs $X$ and $Y$ to be independent."

It never needs independence. Two metrics in the same experiment, two consecutive days of demand, two users from the same household: their expected values still add.

"$E[XY] = E[X]\,E[Y]$, by the same logic."

Products are different. $E[XY] = E[X]E[Y] + Cov(X, Y)$, so it holds only when the covariance is zero (for example, when $X$ and $Y$ are independent). Covariance is taught in Chapter 4.15.

"If the expected number of false alarms is 1, then the chance of at least one is the same whatever the correlation."

The mean count is fixed by linearity, but the probability of "at least one" depends on how the alarms cluster. Correlated metrics fail together.

With a Normal or Student-t likelihood, your forecasting model is a sum: $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$ with noise of mean zero. For fixed parameter values, linearity gives $E[y_t] = g(t) + s(t) + h(t) + X_t\beta$: each component adds its own expected contribution, and the expected demand for a week is the sum of the seven daily expectations even though neighbouring days are correlated. In your A/B framework, $P(\theta_B \gt \theta_A \mid D)$ is the expected value of the indicator $\mathbf{1}\{\theta_B \gt \theta_A\}$ under the posterior. That is why it can be computed as the fraction of posterior draws in which $\theta_B$ beats $\theta_A$.

$E[aX+b] = aE[X]+b$ and $E[X+Y] = E[X]+E[Y]$: always, dependent or not.

Indicator trick: $E[\mathbf{1}_A] = P(A)$, so $E[\text{count}] = \sum P(A_i)$.

Trap: not for products ($E[XY]$) and not for curved functions ($E[X^2]$).

Quick check: daily orders have expected values 120 (Mon–Fri) and 180 (Sat, Sun). What is the expected weekly total? Do you need independence between days?

$5\times120 + 2\times180 = 600 + 360 = 960$ orders. No independence is needed: linearity of expectation holds for any dependence between days. (The variance of the weekly total does depend on the dependence, as you will see below.)

The average of a curve is not the curve of the average: $E[g(X)] \ne g(E[X])$

Should you square the average, or average the squares? Take two equally likely values, 1 and 3. Their average is 2, and $2^2 = 4$. But their squares are 1 and 9, and the average of the squares is 5. Not the same!

The reason is the bend. Squaring stretches big numbers much more than small ones, so the big value gains more than the small value loses. Linearity (the last section) only works for straight-line functions like $3 + 2x$. As soon as the function curves, you cannot "push the average through it".

Three ways to say it:

  • Picture: on a bowl-shaped curve, the straight line (chord) between two points of the curve lies above the curve.
  • Numbers: for the basket, $E[X^2] = 5$ but $(E[X])^2 = 4$.
  • Slogan: you cannot move an average through a curve.

(a) Squares (basket example).

  1. $E[X^2] = \sum x^2 p(x) = 1^2(0.4) + 2^2(0.3) + 3^2(0.2) + 4^2(0.1) = 0.4 + 1.2 + 1.8 + 1.6 = 5.0$.
  2. $(E[X])^2 = 2^2 = 4$.
  3. The gap is $5 - 4 = 1$. Remember this number: it will turn out to be the variance (next section).

(b) Logs. A metric is 10 or 1000, each with probability 1/2.

  1. $E[X] = (10 + 1000)/2 = 505$.
  2. Average on the log scale: $E[\log_{10} X] = (1 + 3)/2 = 2$.
  3. Transform back: $10^2 = 100$. That is far below 505. Averaging on a log scale and transforming back does not give the mean (it gives the geometric mean). This is why back-transformed forecasts need care (Chapter 4.18).

(c) A ratio. Variant A's conversion rate $\theta_A$ is uncertain: 0.04 or 0.06, equally likely. Variant B's rate is $\theta_B = 0.06$.

  1. Ratio for each case: $0.06/0.04 = 1.5$ and $0.06/0.06 = 1.0$. So $E[\theta_B/\theta_A] = (1.5 + 1.0)/2 = 1.25$.
  2. Ratio of the averages: $E[\theta_B]/E[\theta_A] = 0.06/0.05 = 1.2$.
  3. $1.25 \ne 1.2$. The expected relative lift is not the ratio of expected rates.

For any function $g$, the expected value of $g(X)$ weights each $g(x)$ by the probability of $x$:

$$E[g(X)] = \sum_x g(x)\,p(x) \qquad\text{or}\qquad E[g(X)] = \int g(x)\,f(x)\,dx .$$

(You do not need the distribution of $g(X)$ itself; some books call this the "law of the unconscious statistician".)

  • In general $E[g(X)] \ne g(E[X])$. They are equal for every $X$ only when $g$ is a straight line, $g(x) = ax + b$ (that is linearity).
  • Jensen's inequality. If $g$ is convex (curves upward, like $x^2$, $e^x$, or $1/x$ for $x \gt 0$), then $E[g(X)] \ge g(E[X])$. If $g$ is concave (curves downward, like $\ln x$ or $\sqrt{x}$), then $E[g(X)] \le g(E[X])$. The two sides are equal only when $X$ does not vary (or $g$ is straight over the values $X$ takes). The full story is in the Optimization guide, Jensen's inequality.
Why do we need it?

Many numbers we report are curved functions of uncertain quantities: ratios, logs, squares, exponentials. Plugging the average into the function gives a biased answer, and Jensen tells you which way it is off.

Where is it used?

Relative lift $\theta_B/\theta_A$ in A/B tests, back-transforming log-scale forecasts (4.18), the exp or softplus link that keeps a count mean positive, the variance formula $E[X^2] - (E[X])^2 \ge 0$, and the derivation of the ELBO in variational inference (6.12).

How is it used?

Apply the function to every draw first, then average: np.mean(g(draws)), not g(np.mean(draws)). To guess the direction of the error, ask: does $g$ curve up (average of $g$ is bigger) or down (smaller)?

$X$ takes the two blue values $a$ and $b$ (drag them along the curve); the slider sets $P(X = a)$. The orange dot sits on the straight chord at height $E[g(X)]$; the purple dot sits on the curve at height $g(E[X])$. The red gap between them is the Jensen gap. Drag $a$ and $b$ far apart: the gap grows. Switch to ln x: the gap flips sign (concave). Switch to the straight line: the gap vanishes.

"The expected relative lift is $E[\theta_B]/E[\theta_A]$."

Compute the ratio for each posterior draw, then average (or summarise) those ratios. $E[\theta_B/\theta_A]$ is usually different, because $1/\theta_A$ is convex.

"If I model $\log y$ and transform the mean back with $e^{(\cdot)}$, I get the mean of $y$."

$e^{E[\log y]} \le E[y]$ by Jensen (log is concave). For log-normal data the back-transformed value is exactly the median of $y$, which sits below the mean.

"Jensen says the average of $g$ is always the bigger one."

Only for convex (upward-curving) $g$. For concave $g$ such as $\ln$ or $\sqrt{\ }$ the inequality reverses.

In your A/B framework, report relative lift by computing $\theta_B^{(s)}/\theta_A^{(s)} - 1$ for every posterior draw $s$ and summarising those numbers; dividing the two posterior means gives a different (and slightly wrong) answer. In your forecasting model, if a count mean passes through an exp (log link) or softplus to stay positive, then averaging the linear predictor first and transforming afterwards underestimates the mean demand, because both exp and softplus curve upward.

$E[g(X)] = \sum g(x)\,p(x)$; in general $E[g(X)] \ne g(E[X])$ (equal only for straight-line $g$).

Jensen: convex $g$ ⇒ $E[g(X)] \ge g(E[X])$; concave ⇒ $\le$.

Trap: ratio of means ≠ mean of ratios; $e^{\text{mean of logs}}$ ≠ mean.

Quick check: $X$ is 1 or 9, each with probability 1/2. Compare $E[\sqrt{X}]$ with $\sqrt{E[X]}$.

$E[\sqrt X] = (1 + 3)/2 = 2$. $\sqrt{E[X]} = \sqrt{5} \approx 2.24$. The square root is concave, so the average of the roots is smaller, as Jensen predicts.

Variance: the average squared distance from the mean core

Two delivery services both arrive in 30 minutes on average. The first always arrives between 28 and 32 minutes. The second arrives anywhere between 10 and 50. Same mean, very different experience. We need a number for how far from the mean the values usually land.

The natural idea: measure each value's distance from the mean and average those distances. One problem: distances to the left are negative and distances to the right are positive, and they cancel exactly (that is the balance property). The fix used by statistics: square each distance before averaging. Squares are never negative, and big misses count extra.

Three ways to say it:

  • Picture: build a square on each distance from the mean; the variance is the average area of those squares.
  • Numbers: basket sizes 1, 2, 3, 4 are at distances −1, 0, 1, 2 from the mean 2; the probability-weighted average of 1, 0, 1, 4 is 1.0.
  • Slogan: variance = average squared miss from the mean.

Basket sizes ($\mu = E[X] = 2$).

  1. Distances from the mean: $1-2 = -1$, $\;2-2 = 0$, $\;3-2 = 1$, $\;4-2 = 2$.
  2. Squares: $1, 0, 1, 4$.
  3. Weight by the probabilities and add: $1(0.4) + 0(0.3) + 1(0.2) + 4(0.1) = 0.4 + 0 + 0.2 + 0.4 = 1.0$. So $Var(X) = 1.0$ items².
  4. Shortcut check: $E[X^2] - \mu^2 = 5.0 - 4 = 1.0$. ✓ (The Jensen gap from the last section!)

A fair die ($\mu = 3.5$): $E[X^2] = (1+4+9+16+25+36)/6 = 91/6 \approx 15.17$, so $Var(X) = 91/6 - 3.5^2 = 15.17 - 12.25 = 35/12 \approx 2.92$.

A conversion (1 with probability $p$, else 0): $E[X^2] = p$ (because $0^2 = 0$ and $1^2 = 1$), so $Var(X) = p - p^2 = p(1-p)$. For $p = 0.05$: $0.05\times0.95 = 0.0475$.

The variance of a random variable $X$ with mean $\mu$ is

$$Var(X) = E\big[(X - \mu)^2\big] = \sum_x (x-\mu)^2\,p(x) \quad\Big(\text{or } \int (x-\mu)^2 f(x)\,dx\Big).$$

We write it $\sigma^2$ ("sigma squared"). Shortcut formula, derived with linearity:

$$E[(X-\mu)^2] = E[X^2 - 2\mu X + \mu^2] = E[X^2] - 2\mu E[X] + \mu^2 = E[X^2] - \mu^2 .$$
  • $Var(X) \ge 0$, and $Var(X) = 0$ only when $X$ is a constant (never varies).
  • It is in squared units (items², minutes², dollars²). Its square root, the standard deviation, comes next.
  • Bernoulli: $Var = p(1-p)$, largest ($0.25$) at $p = 0.5$.
  • Why squares and not absolute distances? Squares never cancel, they are smooth (easy to differentiate), and they give clean rules: variances of independent pieces add (later in this chapter), and the mean is exactly the point that makes the average squared distance smallest. The average absolute distance is also a valid measure of spread (Chapter 4.14) but has none of these neat rules.
Why do we need it?

The mean alone hides risk. Variance measures how much outcomes wobble, and it is the version of spread with the cleanest algebra, so every later formula for uncertainty (standard errors, intervals, sample sizes) is built from it.

Where is it used?

The noise term $\sigma^2$ of a Normal likelihood, Poisson (variance = mean) versus Negative Binomial (variance $\mu + \mu^2/\alpha$) in 4.8, $p(1-p)$ in A/B sample-size formulas, the bias–variance trade-off, PCA (directions of largest variance), and the law of total variance (4.6).

How is it used?

For a PMF: np.sum((x - mu)**2 * p). For a sample: np.var(x, ddof=1) (the $n-1$ version, explained at the end of the chapter). For counts, compare the variance with the mean to see whether a Poisson model is too narrow.

Five values sit on the number line (drag them). For each one, a red square is built on its distance to the purple mean, so its area is the squared distance. The variance is the average of the five areas, and the dashed purple square has exactly that area: its side is the standard deviation. Press One far value: a single value far away owns most of the total area. That is how strongly squaring punishes big distances.

"The variance is in the same units as the data."

It is in squared units: minutes², dollars², items². That is why we usually report its square root, the standard deviation.

"$E[X^2] - (E[X])^2$ is a good way to compute a variance in code."

With a large mean and a small spread it can lose all its digits. For the values $10^9 + 1,\ 10^9 + 2,\ 10^9 + 3$, NumPy gives np.mean(x**2) - np.mean(x)**2 = 0.0, while the true variance is $0.667$. Use np.var, which subtracts the mean first.

"A variance of 4 means values are typically 4 away from the mean."

Typically about $\sqrt{4} = 2$ away. The variance is an area, the standard deviation is a length.

For a conversion metric, each user is a 0/1 variable with variance $p(1-p)$: at $p = 0.05$ that is $0.0475$ per user. This one number drives how many users your A/B tests need (Chapter 5.7). For count metrics, the variance is what separates the likelihoods in your projects: a Poisson model forces variance = mean, while the Negative Binomial in your forecasting model allows variance $\mu + \mu^2/\alpha$, bigger than the mean (Chapter 4.8).

$Var(X) = E[(X-\mu)^2] = E[X^2] - \mu^2$ (squared units, always $\ge 0$).

Bernoulli: $p(1-p)$. Die: $35/12$.

Trap: squaring makes far values dominate; compute with np.var, not the shortcut.

Quick check: a conversion happens with probability 0.1. What are the variance and standard deviation of the 0/1 conversion variable?

$Var = 0.1\times0.9 = 0.09$, so $SD = \sqrt{0.09} = 0.3$.

Standard deviation: the typical distance, in the original units core

Two coffee shops both sell 100 cups a day on average. Shop A sells 98, 101, 100, 99, 102. Shop B sells 60, 140, 85, 115, 100. Same average, very different days. If you are ordering milk, Shop B's owner needs much more safety stock.

The variance captures this, but it speaks in "cups squared", which nobody can picture. Take its square root and you are back in cups: the standard deviation says how far a typical day is from the average, in the same units as the data.

Three ways to say it:

  • Picture: the standard deviation is the usual length of the "stick" between a value and the mean (the side of the average square from the last widget).
  • Numbers: Shop A's days are typically about 1.6 cups from the mean; Shop B's about 30 cups.
  • Slogan: the mean says where the data are; the standard deviation says how spread out they are.

A random variable. Basket sizes have $Var(X) = 1$ item², so $\sigma = \sqrt{1} = 1$ item.

Data. Shop A, five days: 98, 101, 100, 99, 102 cups.

  1. Mean: $(98+101+100+99+102)/5 = 500/5 = 100$.
  2. Distances from the mean: $-2, +1, 0, -1, +2$.
  3. Square them: $4, 1, 0, 1, 4$. Sum $= 10$.
  4. Average the squares. For a sample we divide by $n-1 = 4$ (the reason is the last section of this chapter): $10/4 = 2.5$. This is the sample variance $s^2$.
  5. Square root, back to cups: $s = \sqrt{2.5} \approx 1.58$ cups. (Dividing by $n = 5$ instead would give $\sqrt{2} \approx 1.41$.)

Shop B gives $s \approx 30.2$ cups (squared distances $1600, 1600, 225, 225, 0$; sum $3650$; $3650/4 = 912.5$; $\sqrt{912.5} \approx 30.2$).

The standard deviation (SD) is the square root of the variance.

  • Of a random variable: $\sigma = \sqrt{Var(X)} = \sqrt{E[(X-\mu)^2]}$ (a fixed property of the distribution).
  • Of a sample $x_1, \dots, x_n$ with sample mean $\bar x$: $\;s = \sqrt{\dfrac{1}{n-1}\displaystyle\sum_{i=1}^{n}(x_i - \bar x)^2}$ (computed from data; it changes from sample to sample).

Properties: same units as the data; $\sigma \ge 0$, and $0$ only when there is no spread at all; $SD(aX + b) = |a|\,\sigma$ (next section).

Reading it (rule of thumb): for bell-shaped data, about two thirds of the values lie within one SD of the mean (about 68% for a Normal, Chapter 4.9). For other shapes this can be quite different; the only guarantee for every distribution is Chebyshev's: at least 75% within two SDs (Chapter 4.12).

Why do we need it?

An average without a spread is half the story. Planning stock, judging whether a change is "big", and building error bars all need to know how much values normally wobble around the mean, in units people understand.

Where is it used?

Standard errors and confidence intervals, z-scores and standardization, the $\sigma$ of a Normal likelihood in a NumPyro model, risk in finance, control charts in factories, and the "noise level" of a forecast.

How is it used?

Compute it with np.std(x, ddof=1) or pandas.Series.std(). Report it next to the mean ("100 ± 1.6 cups"). Divide a difference by it to see whether the difference is large compared with normal day-to-day noise.

Shop A: s ≈ 1.6 cups Shop B: s ≈ 30 cups mean 100 a typical distance ≈ s
Same mean (purple line), very different spread. The standard deviation measures the typical length of the gap between a value and the mean, in cups.

Drag any day left or right. The purple line is the mean, the red sticks are the distances to the mean, and the shaded band is mean ± s. Pull one day far away and see how fast $s$ grows: distances are squared, so one far day counts a lot. Compare $s$ with the mean absolute distance in the readout: $s$ is never smaller. Use the buttons to compare the two shops.

"The standard deviation is the average distance from the mean."

It is the square root of the average squared distance. That is close to, but at least as large as, the average absolute distance, and it reacts much more strongly to far-away values.

"np.std(x) gives the sample standard deviation."

NumPy (and jax.numpy) divide by $n$ by default (ddof=0). Use np.std(x, ddof=1) to divide by $n-1$. pandas' .std() already uses $n-1$.

"The standard deviation and the standard error are the same."

The SD measures the spread of the data. The standard error is the SD of an estimate (such as $\bar x$) across repeated samples; it shrinks as $n$ grows (Chapter 5.5).

In your forecasting model the Normal likelihood $y_t \sim N(\mu_t, \sigma^2)$ has a noise standard deviation $\sigma$: the typical size of a day's surprise after trend, seasonality, holidays and regressors are explained. In NumPyro, dist.Normal(mu, sigma) takes this standard deviation, not the variance. For dist.StudentT(df, loc, scale) the scale is not the SD: for $\nu \gt 2$ the SD is $\text{scale}\times\sqrt{\nu/(\nu-2)}$ (for $\nu = 4$, about $1.41\times$ the scale), and for $\nu \le 2$ there is no finite variance at all (Chapter 4.9). In the A/B framework, the SD of a revenue metric sets how many users you need: noisier metrics need bigger experiments.

"σ and s are the same thing."

σ is a fixed property of the population (or of a random variable). s is computed from a sample, so it changes from sample to sample. s estimates σ.

Model answer: "σ describes the process; s is my estimate of it from the data I have. With more data, s gets closer to σ."

$\sigma = \sqrt{Var(X)}$; sample: $s = \sqrt{\tfrac{1}{n-1}\sum (x_i - \bar x)^2}$, in the units of the data.

Mean = where the data sit; SD = how far a typical value is from it.

Trap: np.std needs ddof=1 for the sample SD; Student-t scale ≠ SD.

Quick check: the data 5, 5, 5, 5. What is s, and why?

Every value equals the mean 5, so every distance is 0, the sum of squares is 0, and $s = 0$. A standard deviation of zero means "no spread at all".

Shifting and scaling: $Var(aX + b) = a^2\,Var(X)$ core

Add 10 to every value and the whole cloud of values slides 10 to the right. Every value moves, and so does the mean, by the same amount. The distances to the mean do not change, so the spread does not change.

Multiply every value by 2 and every distance to the mean doubles. Squared distances become $2^2 = 4$ times bigger, so the variance is multiplied by 4 and the standard deviation by 2. Multiply by $-1$ and the cloud is mirrored, but the distances (and so the spread) stay the same.

Three ways to say it:

  • Picture: sliding a histogram moves it but does not widen it; stretching it widens it.
  • Numbers: temperatures with SD 5 °C have SD $1.8 \times 5 = 9$ °F; the "+32" in the conversion does nothing to the spread.
  • Slogan: shifts move the centre; scales stretch the spread (and square the variance).

Shipping cost $C = 3 + 2X$ for the basket size $X$ ($E[X] = 2$, $Var(X) = 1$).

  1. Mean: $E[C] = 3 + 2\times2 = 7$.
  2. Variance: $Var(C) = 2^2 \times Var(X) = 4\times1 = 4$. The "+3" does not appear.
  3. SD: $SD(C) = |2| \times 1 = 2$ dollars.
  4. Check directly: costs $5, 7, 9, 11$ are at distances $-2, 0, 2, 4$ from 7; squares $4, 0, 4, 16$; weighted: $4(0.4) + 0 + 4(0.2) + 16(0.1) = 1.6 + 0.8 + 1.6 = 4.0$. ✓

Units. Daily temperature has mean 20 °C and SD 5 °C. In Fahrenheit, $F = 1.8C + 32$: mean $1.8\times20 + 32 = 68$ °F, SD $1.8\times5 = 9$ °F, variance $1.8^2\times25 = 3.24\times25 = 81$ °F² $= 9^2$. ✓

Standardizing. $Z = (X - \mu)/\sigma$ is $aX + b$ with $a = 1/\sigma$ and $b = -\mu/\sigma$. So $E[Z] = \mu/\sigma - \mu/\sigma = 0$ and $Var(Z) = Var(X)/\sigma^2 = 1$. Every z-score has mean 0 and SD 1 (Chapter 4.18).

For constants $a$ and $b$:

$$E[aX + b] = a\,E[X] + b,\qquad Var(aX + b) = a^2\,Var(X),\qquad SD(aX + b) = |a|\,SD(X).$$

Why (two lines): the new value minus the new mean is $(aX + b) - (a\mu + b) = a(X - \mu)$. Squaring gives $a^2(X-\mu)^2$, and averaging gives $a^2\,Var(X)$.

  • The shift $b$ never changes the spread.
  • The sign of $a$ never changes the spread ($Var(-X) = Var(X)$).
  • The SD uses $|a|$ (a length cannot be negative).
Why do we need it?

Data are constantly re-expressed: new units, currencies, per-thousand rates, standardized features. This rule tells you, without recomputing anything, what happens to the mean and the spread.

Where is it used?

z-scores and the global scaler in your projects, unit and currency conversion, converting a model's scaled outputs back to business units, and the derivation of the standard error $\sigma/\sqrt{n}$ (the average is $\tfrac1n\times$ a sum).

How is it used?

Mean: apply the same scale and shift. Variance: multiply by the square of the scale and ignore the shift. SD: multiply by the absolute scale. To undo standardization: multiply SDs by $\sigma$, add the mean back only to levels, never to differences.

The top row is the basket distribution $X$ (blue); the bottom row is $Y = aX + b$ (orange). Grey lines show where each value moves. The purple marks are the means and the bars underneath show mean ± SD. Move only $b$: everything slides, the bar keeps its length. Move $a$ to 2, then 3: the bar stretches by $|a|$. Make $a$ negative: the lines cross (the order of the values flips) but the spread is unchanged.

"$Var(2X) = 2\,Var(X)$."

$Var(2X) = 4\,Var(X)$. Variance is in squared units, so the scale gets squared. (It is the SD that doubles.)

"$Var(X + 10) = Var(X) + 10$."

Adding a constant moves the values, not their spread: $Var(X + 10) = Var(X)$.

"$Var(-X) = -Var(X)$."

A variance is never negative. Flipping the sign mirrors the values and leaves the spread alone: $Var(-X) = Var(X)$.

"$X + X$ has variance $Var(X) + Var(X) = 2\,Var(X)$."

$X + X = 2X$, so its variance is $4\,Var(X)$. "Variances add" only works for independent pieces, and $X$ is as dependent on itself as anything can be (next section).

Your A/B framework uses a global scaler: one mean $m$ and one SD $s$ for all groups, $z = (y - m)/s$. By this rule every group keeps its position and spread relative to the others: differences between groups are just divided by $s$. To report a posterior on the original scale, multiply SDs and differences by $s$ (never add $m$ to a difference). Scaling each group with its own mean and SD would be a different transformation per group, and it wipes out exactly the differences you want to estimate (Chapter 4.6, Chapter 4.18).

$E[aX+b] = aE[X]+b$; $\;Var(aX+b) = a^2Var(X)$; $\;SD(aX+b) = |a|\,SD(X)$.

Shifts never change spread; z-scores have mean 0 and SD 1.

Trap: $Var(2X) = 4Var(X)$, not $2Var(X)$.

Quick check: daily revenue has SD 300 dollars. What are the SD and variance in thousands of dollars?

Divide by 1000 ($a = 0.001$): SD $= 0.3$ thousand dollars, variance $= 0.3^2 = 0.09$ (thousand dollars)², which is $300^2/1000^2 = 90\,000/1\,000\,000$. ✓

Adding random quantities: variances add (for independent pieces) core

Add up two noisy things, for example demand on Monday and on Tuesday. Sometimes both are high, sometimes both are low, but often one is high and the other low, and their surprises partly cancel. So the total is noisier than each day, but less noisy than "SD of Monday + SD of Tuesday".

For independent pieces the exact rule is beautifully simple: the variances add. Standard deviations then combine like the sides of a right triangle (Pythagoras), not like lengths on a line.

Three ways to say it:

  • Picture: two independent SDs are the two short sides of a right triangle; the SD of the sum is the long side.
  • Numbers: SD 3 plus SD 4 (independent) gives SD $\sqrt{9 + 16} = 5$, not 7.
  • Slogan: for independent pieces, variances add, standard deviations do not.

(a) Two days of demand. Monday has SD 30 orders, Tuesday SD 40, independent.

  1. Variances: $30^2 = 900$ and $40^2 = 1600$.
  2. Add: $900 + 1600 = 2500$.
  3. SD of the two-day total: $\sqrt{2500} = 50$ orders (not $30 + 40 = 70$).

(b) Two dice. One die has variance $35/12 \approx 2.92$. The sum of two independent dice has variance $35/12 + 35/12 = 35/6 \approx 5.83$, so its SD is $\sqrt{5.83} \approx 2.42$, not $2\times1.71 = 3.42$.

(c) A difference. An A/B comparison looks at $\hat p_B - \hat p_A$, the difference of two independent estimates. Subtracting does not subtract variances: $Var(\hat p_B - \hat p_A) = Var(\hat p_B) + Var(\hat p_A)$. A difference is noisier than either part.

(d) An average. For $n$ independent values each with variance $\sigma^2$, the sum has variance $n\sigma^2$, and the average (sum × $\tfrac1n$) has variance $\tfrac{1}{n^2}\times n\sigma^2 = \sigma^2/n$. Its SD is $\sigma/\sqrt n$: the standard error (taught in Chapter 5.5).

For any two random variables (with finite variances):

$$Var(X + Y) = Var(X) + Var(Y) + 2\,Cov(X, Y),\qquad Var(X - Y) = Var(X) + Var(Y) - 2\,Cov(X, Y),$$

where the covariance $Cov(X,Y) = E[(X - \mu_X)(Y - \mu_Y)]$ measures whether $X$ and $Y$ tend to be above their means together (positive) or on opposite sides (negative). Covariance and correlation get their own chapter (Chapter 4.15); here you only need: independent variables have covariance 0.

Derivation. $(X + Y) - (\mu_X + \mu_Y) = (X - \mu_X) + (Y - \mu_Y)$. Square it: $(X-\mu_X)^2 + (Y-\mu_Y)^2 + 2(X-\mu_X)(Y-\mu_Y)$. Take expected values term by term (linearity): $Var(X) + Var(Y) + 2Cov(X,Y)$.

  • Independent (or just uncorrelated): $Var(X \pm Y) = Var(X) + Var(Y)$. Note the plus for a difference too.
  • Many independent pieces: $Var\left(\sum_i a_iX_i\right) = \sum_i a_i^2\,Var(X_i)$.
  • Positively correlated pieces add more variance; negatively correlated pieces partly cancel. $X + X$ (correlation 1) has variance $4Var(X)$.
Why do we need it?

Totals, differences and averages are everywhere, and we need their uncertainty. This rule gives it from the uncertainty of the parts, and it explains why averaging many observations makes an estimate steadier.

Where is it used?

The standard error $\sigma/\sqrt n$, the uncertainty of an A/B difference, error budgets in engineering, portfolio risk, forecast intervals for weekly or monthly totals, and the noise of a Monte Carlo average (the ELBO estimate in SVI shrinks its variance like $1/S$ with $S$ samples).

How is it used?

Square each SD, add the variances (plus twice any covariance), then take the square root at the end. Never add SDs directly unless the pieces are perfectly correlated.

SD(X) = 3 SD(Y) = 4 SD(X+Y) = √(3²+4²) = 5 Independent pieces: variances add: 9 + 16 = 25 SDs do not: 3 + 4 = 7 ✗ Correlated pieces: the angle is not 90°, and 2·Cov appears.
Standard deviations of independent pieces combine like the sides of a right triangle. Only variances (the squares) add.

Left: 1500 simulated pairs $(X, Y)$. Right: the histogram of $X + Y$ (or $X - Y$), with the orange curve predicted by the variance rule. Start with SDs 3 and 4 and correlation 0: the sum has SD 5. Push the correlation to +1: now the SDs do add (7). Push it to −1: the pieces cancel and the spread collapses to $|3 - 4| = 1$. Switch to $X - Y$ and notice the sign of the covariance term flips.

"Standard deviations add."

Variances add (for independent pieces); SDs combine as $\sqrt{\sigma_X^2 + \sigma_Y^2}$. SDs add only when the pieces are perfectly positively correlated.

"$Var(X - Y) = Var(X) - Var(Y)$."

For independent $X, Y$ it is $Var(X) + Var(Y)$. Subtracting a random quantity adds its noise; a difference can even have a larger SD than both parts.

"Variances always add."

Only when the covariance is zero. Days of demand that rise and fall together (positive covariance) make a weekly total noisier than the simple sum of variances says.

In your A/B framework, when the posteriors of the two variants are independent (separate groups of users, no shared parameters), $Var(\theta_B - \theta_A \mid D) = Var(\theta_A \mid D) + Var(\theta_B \mid D)$: the uncertainty of the lift is bigger than the uncertainty of either rate. With hierarchical sharing the posteriors can be correlated, and the covariance term matters; computing the difference draw by draw handles that automatically. In your forecasting model, the uncertainty of a weekly total depends on how the daily errors co-move, which is one reason to simulate whole paths rather than add daily SDs (Chapter 7.14).

"If the SD of each variant's conversion-rate estimate is 0.003, the SD of the difference is 0.006."

For independent estimates it is $\sqrt{0.003^2 + 0.003^2} = 0.003\sqrt2 \approx 0.0042$.

Model answer: "Variances of independent quantities add, so the variance of the difference is the sum of the two variances, and the SD is the square root of that sum. SDs add only under perfect positive correlation."

$Var(X \pm Y) = Var(X) + Var(Y) \pm 2Cov(X,Y)$; independent ⇒ $Var(X) + Var(Y)$ (plus, even for a difference).

Average of $n$ iid values: variance $\sigma^2/n$, SD $\sigma/\sqrt n$.

Trap: add variances, not SDs; $X + X$ is not "two independent copies".

Quick check: daily demand has SD 20 and the 7 days are independent. What is the SD of the weekly total?

Variance $7\times20^2 = 2800$, SD $\sqrt{2800} \approx 52.9$. Not $7\times20 = 140$. (If the days were positively correlated, the true SD would be larger than 52.9.)

Population variance vs sample variance: why we divide by $n - 1$ core

We want the spread around the true mean $\mu$, but we do not know $\mu$. So we measure distances to the sample mean $\bar x$ instead. Here is the catch: $\bar x$ is computed from these very points, so it always sits right in the middle of them. In fact $\bar x$ is the single point that makes the sum of squared distances as small as possible. The true mean is somewhere else, a little off, and distances to it are a little bigger.

So "average squared distance to $\bar x$" is too small on average. Dividing by $n - 1$ instead of $n$ makes the number slightly bigger, and it turns out to be exactly the right correction: on average over many samples it equals the true variance.

Three ways to say it:

  • Picture: the sample mean is the centre of its own points, so the points look closer to their centre than to the true centre.
  • Numbers: from the population $\{1, 3\}$ (variance 1), samples of size 2 give an average "÷n variance" of 0.5 but an average "÷(n−1) variance" of exactly 1.
  • Slogan: we spent one piece of information estimating the mean, so we divide by $n - 1$.

All possible samples. A population has two equally likely values, 1 and 3. Its mean is $\mu = 2$ and its variance is $\sigma^2 = \tfrac{(1-2)^2 + (3-2)^2}{2} = 1$. Draw a sample of size $n = 2$ (with replacement). There are four equally likely samples:

Sample$\bar x$$\sum (x_i - \bar x)^2$÷ $n$ = 2÷ $(n-1)$ = 1
1, 11000
1, 321 + 1 = 212
3, 12212
3, 33000
Average over the 4 samples210.51
  1. Samples (1, 1) and (3, 3) have zero spread around their own mean, although the population does vary.
  2. Dividing by $n$: average $= (0 + 1 + 1 + 0)/4 = 0.5$, only half the true variance.
  3. Dividing by $n - 1$: average $= (0 + 2 + 2 + 0)/4 = 1$, exactly the true variance.

For data $x_1, \dots, x_n$ drawn independently from a distribution with mean $\mu$ and variance $\sigma^2$:

  • Population variance $\sigma^2 = E[(X - \mu)^2]$: a fixed property of the distribution. (For a complete, finite population of $N$ values, $\sigma^2 = \tfrac1N\sum(x - \mu)^2$.)
  • Sample variance $s^2 = \dfrac{1}{n-1}\displaystyle\sum_{i=1}^n (x_i - \bar x)^2$. The divisor $n - 1$ is called Bessel's correction.
  • The "÷ n" version $\hat\sigma^2 = \tfrac1n\sum (x_i - \bar x)^2$ is also used (it is the maximum-likelihood estimate under a Normal model, Chapter 5.2).

An estimator is a recipe that turns data into a guess of an unknown quantity; it is unbiased if its average over many repeated samples equals the true value (Chapter 5.1). The facts: $E[s^2] = \sigma^2$ (unbiased) and $E[\hat\sigma^2] = \tfrac{n-1}{n}\sigma^2$ (too small).

Proof in three steps.

  1. Write $x_i - \mu = (x_i - \bar x) + (\bar x - \mu)$, square and add over $i$. The cross term is $2(\bar x - \mu)\sum(x_i - \bar x) = 0$, because deviations from $\bar x$ sum to zero. So $\sum(x_i - \bar x)^2 = \sum(x_i - \mu)^2 - n(\bar x - \mu)^2$.
  2. Take expected values: $E\left[\sum(x_i - \mu)^2\right] = n\sigma^2$, and $E[(\bar x - \mu)^2] = Var(\bar x) = \sigma^2/n$ (from the previous section).
  3. So $E\left[\sum(x_i - \bar x)^2\right] = n\sigma^2 - n\cdot\tfrac{\sigma^2}{n} = (n-1)\sigma^2$. Divide by $n - 1$ to get exactly $\sigma^2$.

Degrees of freedom. The $n$ deviations $x_i - \bar x$ always add to zero, so once you know $n - 1$ of them, the last one is fixed. Only $n - 1$ of them are free to vary: we say the sum of squares has $n - 1$ degrees of freedom.

Why do we need it?

Without the correction, every spread we measure from a small sample is too small on average, so error bars are too narrow and tests are too confident. The $n - 1$ divisor fixes the average error caused by measuring around $\bar x$ instead of $\mu$.

Where is it used?

The sample SD in every t-test and confidence interval, pandas .var()/.std(), R's var, ANOVA (sums of squares divided by degrees of freedom), and the $n - p$ divisor for residual variance in regression with $p$ fitted coefficients.

How is it used?

Use np.var(x, ddof=1) or np.std(x, ddof=1) ("delta degrees of freedom" = 1 means divide by $n-1$). Check the default of each library. For large $n$ the choice barely matters; for $n$ below about 30 it does.

sample mean x̄ true mean μ distances to x̄: short distances to μ: longer x̄ is the centre of its own points, so Σ(xᵢ − x̄)² ≤ Σ(xᵢ − μ)² for every sample.
The sample mean is fitted to the sample, so squared distances to it are never larger, and usually smaller, than squared distances to the unknown true mean. Dividing by n − 1 compensates on average.

Five values (blue, along the bottom) come from a population with true mean μ = 10 (green). The curve shows $SS(c) = \sum (x_i - c)^2$ for every possible centre $c$. Drag the orange point along the curve: the lowest point is always at the sample mean $\bar x$ (purple), never at μ. Press New sample several times. The gap $SS(\mu) - SS(\bar x)$ is always $n(\bar x - \mu)^2 \ge 0$, so measuring around $\bar x$ under-measures the spread.

Each click draws 1000 new samples of size $n$ from the chosen population and computes both spread estimates for each sample. The lines show their running averages divided by the true $\sigma^2$, so the target is the green line at 1. The orange "÷ n" line settles at $(n-1)/n$ (the dashed purple line), too low; the blue "÷ (n − 1)" line settles at 1. Set $n = 2$: "÷ n" is off by half! Set $n = 30$: the two nearly agree. Try every population: the correction works for skewed and yes/no data too.

"Dividing by $n - 1$ makes the sample SD $s$ unbiased."

It makes $s^2$ unbiased for $\sigma^2$. The square root is concave, so by Jensen $E[s] \le \sigma$: $s$ is still a little too small on average (for Normal data with $n = 5$, $E[s] \approx 0.94\,\sigma$). The bias fades as $n$ grows.

"$n - 1$ is always the best divisor."

It is the unbiased one. If you care about mean squared error instead, other divisors can win (for Normal data, dividing by $n + 1$ gives the smallest MSE). "Best" depends on the goal; this trade-off is the topic of Chapter 5.1.

"If I have the whole population, I should still divide by $n - 1$."

With every member of a finite population in hand, you know $\mu$ exactly and the variance is $\tfrac1N\sum(x - \mu)^2$. The correction exists because $\bar x$ is only an estimate of $\mu$.

jax.numpy.var and np.var divide by $n$ by default, while pandas divides by $n - 1$. If a scaler is computed in JAX on one machine and in pandas on another, the two SDs differ by the factor $\sqrt{n/(n-1)}$: invisible for large datasets, noticeable for small segments. Note also that inside a Bayesian NumPyro model like yours, a noise scale $\sigma$ is a parameter with a prior and a posterior, so there is no ddof choice there; the $n$ vs $n - 1$ question appears in preprocessing (scalers) and in quick summary statistics.

"We divide by $n - 1$ because one data point is lost."

No data point is lost. We divide by $n - 1$ because the deviations are measured from $\bar x$, which was fitted to the same data; this uses up one degree of freedom (the deviations must sum to zero) and makes the raw sum of squares too small by exactly one $\sigma^2$ on average.

Model answer: "$E[\sum(x_i - \bar x)^2] = (n-1)\sigma^2$, because $\bar x$ is the point that minimises the sum of squares, so it is always at least as close to the data as the true mean. Dividing by $n - 1$ makes $s^2$ an unbiased estimator of $\sigma^2$. $s$ itself is still slightly biased."

$s^2 = \tfrac{1}{n-1}\sum(x_i - \bar x)^2$ is unbiased: $E[s^2] = \sigma^2$; the ÷n version averages $\tfrac{n-1}{n}\sigma^2$.

Reason: $\sum(x_i - \bar x)^2 = \sum(x_i - \mu)^2 - n(\bar x - \mu)^2$, and $\bar x$ uses up one degree of freedom.

Trap: $s$ is still biased low; NumPy/JAX default ddof=0, pandas ddof=1.

Quick check: a population has values 0 and 6, equally likely. For samples of size 2, what are the average ÷n and ÷(n−1) variances?

$\sigma^2 = 9$. Samples (0,0), (0,6), (6,0), (6,6) have sums of squares 0, 18, 18, 0. ÷2: average $(0+9+9+0)/4 = 4.5$ (half of 9). ÷1: average $(0+18+18+0)/4 = 9$ (exactly $\sigma^2$).

Recap, cheat sheet and practice

  • The expected value $E[X] = \sum x\,p(x)$ (or $\int x f(x)\,dx$) is the balance point of a distribution and the long-run average of many repeats. It need not be a possible value or the most likely value.
  • Linearity: $E[aX + b] = aE[X] + b$ and $E[X+Y] = E[X]+E[Y]$, with no independence needed. Counts are sums of indicators, so $E[\text{count}] = \sum P(A_i)$.
  • For curved $g$, $E[g(X)] \ne g(E[X])$; Jensen: convex ⇒ $\ge$, concave ⇒ $\le$.
  • The variance $Var(X) = E[(X-\mu)^2] = E[X^2] - \mu^2$ is the average squared distance from the mean (squared units). The SD is its square root (original units).
  • $Var(aX+b) = a^2 Var(X)$: shifts do nothing to spread; scales are squared.
  • $Var(X \pm Y) = Var X + Var Y \pm 2Cov(X,Y)$: for independent pieces variances add (even for a difference), SDs do not. The average of $n$ iid values has variance $\sigma^2/n$.
  • The sample variance divides by $n - 1$ because deviations are measured from $\bar x$; this makes $s^2$ unbiased for $\sigma^2$ (but $s$ is still slightly low).

Cheat sheet

QuantityFormulaIn words
Expected value$E[X] = \sum x\,p(x)$, $\;\int x f(x)\,dx$balance point; long-run average
Bernoulli mean / variance$p$, $\;p(1-p)$0/1 variable: mean = probability of 1
Function of $X$$E[g(X)] = \sum g(x)\,p(x)$$\ne g(E[X])$ unless $g$ is linear
Linearity$E[aX+bY+c] = aE[X]+bE[Y]+c$always, dependent or not
Variance$E[(X-\mu)^2] = E[X^2]-\mu^2$average squared distance
Standard deviation$\sigma = \sqrt{Var(X)}$typical distance, original units
Shift and scale$Var(aX+b) = a^2Var(X)$, $SD = |a|\sigma$shift: no effect; scale: squared
Sum / difference$Var(X\pm Y) = Var X + Var Y \pm 2Cov$independent: variances add
Average of $n$ iid$Var(\bar X) = \sigma^2/n$SD shrinks like $1/\sqrt n$
Sample variance$s^2 = \frac{1}{n-1}\sum(x_i-\bar x)^2$unbiased; ddof=1
Code it · Python

import numpy as np
import pandas as pd

rng = np.random.default_rng(0)

# 1. Expected value and variance of a PMF: items in a basket
x = np.array([1, 2, 3, 4])
p = np.array([0.4, 0.3, 0.2, 0.1])
mu = np.sum(x * p)                      # balance point
var = np.sum((x - mu) ** 2 * p)         # average squared distance
print(mu, var, np.sum(x**2 * p) - mu**2)   # 2.0 1.0 1.0  (the shortcut agrees)

# 2. Long-run average: revenue per visitor (5% buy, average order 40)
n = 200_000
buys = rng.random(n) < 0.05
order = rng.gamma(shape=2.0, scale=20.0, size=n)   # mean = 2 * 20 = 40
revenue = np.where(buys, order, 0.0)
print(revenue.mean())                   # 1.979... close to 0.05 * 40 = 2.0

# 3. Linearity: expected false alarms across 20 correlated metrics
m, rho, R = 20, 0.6, 20_000
z0 = rng.standard_normal((R, 1))
z = np.sqrt(rho) * z0 + np.sqrt(1 - rho) * rng.standard_normal((R, m))
alarms = (np.abs(z) > 1.96).sum(axis=1)
print(alarms.mean(), (alarms > 0).mean())   # 0.99 0.355: mean still 20*0.05 = 1, but P(at least one) is far below 0.64

# 4. Jensen: E[g(X)] is not g(E[X])
print(np.sum(x**2 * p), mu**2)          # 5.0 4.0  (the gap is the variance)

# 5. Shift and scale: shipping cost C = 3 + 2X
c = 3 + 2 * x
mc = np.sum(c * p)
print(mc, np.sum((c - mc) ** 2 * p))     # 7.0 4.0  = 3 + 2*2 and 2**2 * 1

# 6. Variances of independent pieces add; SDs do not
a = rng.normal(0, 3, 1_000_000)
b = rng.normal(0, 4, 1_000_000)
print(np.var(a + b), np.std(a + b), np.var(a - b))   # about 25, 5, 25 (not 7 for the SD)

# 7. ddof: NumPy divides by n, pandas by n - 1
cups = np.array([98, 101, 100, 99, 102])
print(np.var(cups), np.var(cups, ddof=1), pd.Series(cups).var())   # 2.0 2.5 2.5

# 8. Why n - 1: average many sample variances (Normal, sigma^2 = 4, n = 5)
S = rng.normal(0, 2, size=(200_000, 5))
print(S.var(axis=1, ddof=0).mean())     # 3.199... about (n-1)/n * 4 = 3.2 (too small)
print(S.var(axis=1, ddof=1).mean())     # 3.999... about 4.0 (unbiased)
print(S.std(axis=1, ddof=1).mean())     # 1.879... below 2: s itself is biased low
Test yourself

1. A promotion pays out 10 with probability 0.3 and 0 otherwise. What is the expected payout?

$E[X] = 10\times0.3 + 0\times0.7 = 3$. The payout is never actually 3; the expected value is the long-run average per customer.

2. $Var(X) = 4$. What is $Var(3X + 5)$?

The shift 5 does nothing; the scale 3 is squared: $3^2\times4 = 36$. (The SD goes from 2 to 6.)

3. $X$ and $Y$ are independent with SDs 6 and 8. What is the SD of $X - Y$?

Variances add, even for a difference: $36 + 64 = 100$, so the SD is $\sqrt{100} = 10$. (100 is the variance, not the SD.)

4. Which statement is true for any two random variables with finite means?

Linearity of expectation holds with or without independence. The product rule and the variance rule need zero covariance, and $E[X^2] - (E[X])^2 = Var(X) \ge 0$.

5. Why does the sample variance divide by $n - 1$?

$E[\sum(x_i - \bar x)^2] = (n-1)\sigma^2$, so dividing by $n-1$ makes $s^2$ unbiased. $s$ itself stays slightly biased, and other divisors can have a smaller MSE.

6. Posterior draws of $\theta_A$ and $\theta_B$ are available. How should you estimate the expected relative lift $E[\theta_B/\theta_A]$?

$E[g(X)] \ne g(E[X])$ for a curved $g$; the ratio is curved in $\theta_A$. Apply the function per draw, then average.

Practice problems

A. A game costs 1 to play. You roll a die and win 6 if it shows a six (nothing otherwise). What is your expected profit per game?

Profit is $6 - 1 = 5$ with probability $1/6$ and $-1$ with probability $5/6$. $E = 5/6 - 5/6 = 0$: a fair game. (Check with linearity: expected winnings $6\times1/6 = 1$, minus the cost 1.)

B. $X$ takes the values 0, 2, 4 with probabilities 0.25, 0.5, 0.25. Find $E[X]$, $Var(X)$ and $SD(X)$ in two ways.

$E[X] = 0 + 1 + 1 = 2$. Definition: $(0-2)^2(0.25) + 0 + (4-2)^2(0.25) = 1 + 1 = 2$. Shortcut: $E[X^2] = 0 + 4(0.5) + 16(0.25) = 2 + 4 = 6$, and $6 - 2^2 = 2$. ✓ So $Var = 2$ and $SD = \sqrt2 \approx 1.41$.

C. A shop has 300 visitors from email (each buys with probability 0.1) and 700 from search (0.02 each). Expected number of buyers? Does it matter whether visitors influence each other?

Sum of indicators: $300\times0.1 + 700\times0.02 = 30 + 14 = 44$. Linearity does not need independence, so it does not matter for the expected number. It does matter for the variance of the number of buyers.

D. Interview: "Daily demand has SD 30 and days are independent. My colleague says the SD of two-day demand is 60. Is that right?"

No. Variances add: $30^2 + 30^2 = 1800$, so the SD is $\sqrt{1800} \approx 42.4$. SDs add only if the two days move together perfectly (correlation +1). If days are positively correlated in reality, the true SD lies between 42.4 and 60.

E. Show that $Var(\bar X) = \sigma^2/n$ for $n$ independent draws with variance $\sigma^2$, using only this chapter's rules.

$\bar X = \tfrac1n(X_1 + \dots + X_n)$. Independent pieces: $Var(X_1 + \dots + X_n) = n\sigma^2$. Scaling by $\tfrac1n$ multiplies the variance by $\tfrac{1}{n^2}$: $Var(\bar X) = n\sigma^2/n^2 = \sigma^2/n$. The SD is $\sigma/\sqrt n$.

F. Interview: "Why is the expected value of a conversion indicator equal to the conversion rate, and what is its variance at p = 0.5 and at p = 0.02?"

$E[X] = 0\cdot(1-p) + 1\cdot p = p$. $Var = p(1-p)$: $0.25$ at $p = 0.5$ (the maximum) and $0.02\times0.98 = 0.0196$ at $p = 0.02$. Relative to the mean, though, low rates are noisier: the SD $\sqrt{0.0196} = 0.14$ is 7 times the mean 0.02, which is why tests on rare conversions need many users.

Chapter 4.6 · Syllabus Modules 1.7–1.8

Conditioning on another variable: total expectation and total variance

Real data come in groups: devices, countries, user segments, weekdays and weekends. This chapter teaches two short laws that connect the groups to the whole. The first gets the overall average from the group averages. The second splits the total spread into spread inside the groups and spread between the groups. That split is the main idea behind the hierarchical models in your A/B framework.

  • Read a conditional distribution: the distribution of $X$ inside one group $Y = y$
  • Understand the conditional expectation $E[X \mid Y]$ as a random variable: one mean per group
  • Use the law of total expectation $E[X] = E\big[E[X \mid Y]\big]$: the overall mean is a weighted average of group means
  • Recognise Simpson's paradox: a change in the group mix can reverse a comparison
  • Derive and use the law of total variance: total = within-group + between-group
  • See mixtures of segments and overdispersed counts as applications of the two laws
  • Connect between-group variance $\tau^2$ and within-group variance $\sigma^2$ to hierarchical models and to the global scaler

The conditional distribution: the distribution inside one group

What you need first: conditional probability $P(A \mid B) = P(A \cap B)/P(B)$ as "shrinking the world to $B$" (Chapter 4.3); random variables and PMFs (Chapter 4.4); expected value and variance (Chapter 4.5).

Ask "how many items are in a basket?" and the answer is a distribution over all shoppers. Now ask "how many items are in a basket for shoppers on mobile?" You have zoomed into one group. The answer is a new distribution: same possible values, different probabilities.

That zoomed-in distribution is the conditional distribution of $X$ (items) given $Y$ (device) equals "mobile". To get it from a table of counts, keep only the mobile row and rescale it so it adds up to 1. Nothing else changes.

Three ways to say it:

  • Picture: cut one row out of the table and stretch it until it sums to 1.
  • Numbers: mobile baskets have 1, 2, 3 items with probabilities 0.6, 0.3, 0.1; desktop baskets 0.2, 0.4, 0.4.
  • Slogan: a conditional distribution is the distribution inside one group.

1000 orders, counted by device $Y$ and number of items $X$ (the running example of this chapter):

1 item2 items3 itemsrow total
mobile36018060600
desktop80160160400
column total4403402201000
  1. Joint probabilities: divide every cell by 1000. For example $P(X{=}1, Y{=}\text{mobile}) = 360/1000 = 0.36$.
  2. Marginal of $X$ (ignore the device): the column totals ÷ 1000 give $0.44, 0.34, 0.22$. The marginal of $Y$: $P(\text{mobile}) = 0.6$, $P(\text{desktop}) = 0.4$.
  3. Conditional of $X$ given mobile: divide the mobile row by its total. $360/600 = 0.6$, $\;180/600 = 0.3$, $\;60/600 = 0.1$.
  4. Conditional of $X$ given desktop: $80/400 = 0.2$, $\;160/400 = 0.4$, $\;160/400 = 0.4$.
  5. Each conditional distribution adds to 1: $0.6 + 0.3 + 0.1 = 1$ and $0.2 + 0.4 + 0.4 = 1$.

Let $X$ and $Y$ be two random variables observed together. (Here $Y$ is the "group label", but it can be any variable.)

  • The joint distribution $p(x, y) = P(X = x, Y = y)$ gives the probability of each pair.
  • The marginal distribution of $X$ is $p_X(x) = \sum_y p(x, y)$: the distribution of $X$ alone, ignoring $Y$. (The name comes from writing these sums in the margin of the table.)
  • The conditional distribution of $X$ given $Y = y$ is
$$p(x \mid y) = \frac{p(x, y)}{p_Y(y)}, \qquad \text{for every } y \text{ with } p_Y(y) \gt 0 .$$

For continuous variables the same formula holds with densities: $f(x \mid y) = f(x, y)/f_Y(y)$.

  • For each fixed $y$, $p(x \mid y)$ is a proper distribution in $x$: non-negative and summing to 1.
  • Rearranged: $p(x, y) = p(x \mid y)\,p_Y(y)$ (the multiplication rule of Chapter 4.3).
  • $X$ and $Y$ are independent exactly when $p(x \mid y) = p_X(x)$ for every $y$: knowing the group changes nothing.
Why do we need it?

Averages over everybody can hide the fact that groups behave very differently. The conditional distribution describes each group on its own, which is the first step for comparing segments, building group-level models, and computing overall numbers correctly.

Where is it used?

Segment analysis in A/B tests (mobile vs desktop), the likelihood $p(\text{data} \mid \theta)$ in every Bayesian model, $p(y_t \mid \text{past})$ in forecasting, class-conditional densities in naive Bayes and discriminant analysis, and contingency tables in chi-square tests.

How is it used?

In pandas: pd.crosstab(df.device, df.items, normalize="index") divides each row by its total. Or df[df.device == "mobile"].items.value_counts(normalize=True). Then compare the rows with each other and with the marginal.

joint counts 1 item2 items3 itemstotal mobile36018060600 desktop80160160400 margin4403402201000 ÷ 600 P(X | mobile) 123 0.60.30.1 Conditioning = keep one row and divide by its total. The margin ignores the device.
From a joint table to a conditional distribution: the mobile row (blue) divided by its total 600 gives P(X = x | mobile).

The table holds order counts (rows: mobile, desktop; columns: 1, 2, 3 items). Choose Mobile or Desktop: the coloured bars show that group's distribution, and the dashed outline shows the distribution of everybody (the marginal). Edit the counts: make the desktop row proportional to the mobile row (for example 120, 60, 20) and notice the bars stop changing between groups. That is independence: the device tells you nothing about basket size.

"$P(X = 1 \mid \text{mobile})$ is the share of all orders that are mobile one-item orders."

That is the joint probability $0.36$. The conditional probability divides by the mobile total only: $360/600 = 0.6$.

"$P(X \mid Y)$ and $P(Y \mid X)$ are the same table."

One divides rows by row totals, the other divides columns by column totals. $P(\text{1 item} \mid \text{mobile}) = 0.6$ but $P(\text{mobile} \mid \text{1 item}) = 360/440 \approx 0.82$ (the same trap as in Bayes' theorem, Chapter 4.3).

In your A/B framework, "the conversion distribution for users in segment $g$" is a conditional distribution $p(y \mid \text{segment} = g)$, and the model's likelihood $p(y \mid \theta_g)$ is a conditional distribution given the segment's parameter. In your forecasting model, the likelihood $p(y_t \mid g(t), s(t), h(t), X_t\beta, \sigma)$ is the distribution of demand given all the components on day $t$.

Joint $p(x,y)$; marginal $p_X(x) = \sum_y p(x,y)$; conditional $p(x \mid y) = p(x,y)/p_Y(y)$.

Conditioning = keep one row, divide by its total. Each row then sums to 1.

Trap: joint ≠ conditional, and $p(x \mid y) \ne p(y \mid x)$.

Quick check: using the table, what is $P(X = 3 \mid \text{desktop})$, and what is $P(\text{desktop} \mid X = 3)$?

$P(X=3 \mid \text{desktop}) = 160/400 = 0.4$ (divide by the desktop row total). $P(\text{desktop} \mid X=3) = 160/220 \approx 0.73$ (divide by the 3-item column total).

Conditional expectation: $E[X \mid Y]$ is one mean per group core

Once you have the distribution inside a group, you can take its mean. The average basket for mobile shoppers is 1.5 items; for desktop shoppers it is 2.2 items. These are conditional expectations: $E[X \mid Y = \text{mobile}] = 1.5$ and $E[X \mid Y = \text{desktop}] = 2.2$. Each is an ordinary number.

Now a subtle but powerful step. Before the next shopper arrives, you do not know their device. So you do not know which group mean will apply: it will be 1.5 (if they are on mobile, probability 0.6) or 2.2 (desktop, probability 0.4). That uncertain group mean is itself a random variable. We call it $E[X \mid Y]$, with no "$= y$". Think of it as replacing every shopper's basket by the average basket of their group.

Three ways to say it:

  • Picture: squash every data point onto its group's mean; what is left is $E[X \mid Y]$.
  • Numbers: $E[X \mid Y]$ equals 1.5 with probability 0.6 and 2.2 with probability 0.4.
  • Slogan: $E[X \mid Y = y]$ is a number; $E[X \mid Y]$ is a random variable, one mean per group.

From the conditional distributions of the last section:

  1. $E[X \mid \text{mobile}] = 1(0.6) + 2(0.3) + 3(0.1) = 0.6 + 0.6 + 0.3 = 1.5$.
  2. $E[X \mid \text{desktop}] = 1(0.2) + 2(0.4) + 3(0.4) = 0.2 + 0.8 + 1.2 = 2.2$.
  3. So $E[X \mid Y]$ is the random variable that equals $1.5$ when $Y$ = mobile (probability $0.6$) and $2.2$ when $Y$ = desktop (probability $0.4$).
  4. Like any random variable it has a mean: $0.6(1.5) + 0.4(2.2) = 0.9 + 0.88 = 1.78$. Hold on to this number: it is the overall mean (next section).
  5. It also has a variance: $0.6(1.5 - 1.78)^2 + 0.4(2.2 - 1.78)^2 = 0.6(0.0784) + 0.4(0.1764) = 0.04704 + 0.07056 = 0.1176$. This will be the "between-group" variance.

The spread inside each group is a conditional variance too: $Var(X \mid \text{mobile}) = E[X^2 \mid \text{mobile}] - 1.5^2 = (0.6 + 1.2 + 0.9) - 2.25 = 0.45$, and $Var(X \mid \text{desktop}) = (0.2 + 1.6 + 3.6) - 2.2^2 = 5.4 - 4.84 = 0.56$.

The conditional expectation of $X$ given $Y = y$ is the mean of the conditional distribution:

$$E[X \mid Y = y] = \sum_x x\,p(x \mid y) \qquad\Big(\text{or } \int x\,f(x \mid y)\,dx\Big).$$

Call this function $h(y)$: it gives one number for each group $y$. The conditional expectation of $X$ given $Y$ is the random variable

$$E[X \mid Y] = h(Y),$$

which takes the value $h(y)$ with probability $P(Y = y)$. In the same way, $Var(X \mid Y = y) = E\big[(X - h(y))^2 \mid Y = y\big]$ is the spread inside group $y$, and $Var(X \mid Y)$ is the random variable "the spread of whatever group you land in".

  • $E[X \mid Y]$ is a function of $Y$: once you know the group, you know its value exactly.
  • It is the best prediction of $X$ from $Y$ in the squared-error sense: no other function of $Y$ has a smaller average squared error. This is why regression models aim at $E[\text{target} \mid \text{features}]$ (Chapter 5.13).
  • If $X$ and $Y$ are independent, $E[X \mid Y] = E[X]$: every group has the same mean, and the random variable does not vary.
Why do we need it?

"The mean" is rarely one number in practice: it depends on the segment, the day, the features. $E[X \mid Y]$ is how we write "the mean for each group" as one object we can compute with, and it is exactly what prediction models try to learn.

Where is it used?

Segment means in A/B tests, the regression function $E[y \mid x]$ in linear and generalized linear models, the point forecast $E[y_t \mid \text{information so far}]$, the group means $\theta_g$ of hierarchical models, and the derivations of both laws in this chapter.

How is it used?

In pandas: df.groupby("device").value.mean() gives $E[X \mid Y = y]$ for each group; df.groupby("device").value.transform("mean") writes each row's group mean next to it, which is the random variable $E[X \mid Y]$ evaluated on your data.

the group Y (random) the value of E[X | Y] mobile (prob 0.6) desktop (prob 0.4) E[X | Y = mobile] E[X | Y = desktop] 1.5 items 2.2 items E[X | Y] = 1.5 (prob 0.6) or 2.2 (prob 0.4): a random variable with mean 1.78.
E[X | Y] is a function of the group: it sends each group to that group's mean. Because the group is random, E[X | Y] is random too.

Each dot is one order value, in three device groups (rows). The bottom row shows all orders together. Press Squash to group means (or move the slider): every dot slides to its group's mean. In the bottom row only three values remain, with sizes 50%, 33% and 17%: that is the random variable $E[X \mid Y]$. Its mean is still the overall mean, but its spread is smaller: only the between-group spread is left.

"$E[X \mid Y]$ is a number, like $E[X]$."

$E[X \mid Y = y]$ (a specific group) is a number. $E[X \mid Y]$ (no specific group) is a random variable: it is uncertain because $Y$ is uncertain.

"$E[X \mid Y]$ is a random variable, so it has its own extra randomness."

All its randomness comes from $Y$. Once the group is known, $E[X \mid Y]$ is known exactly. It is a function of $Y$.

In the A/B framework, the segment-level conversion rate $\theta_g$ is $E[\text{conversion} \mid \text{segment} = g]$: one mean per group. Across segments, these $\theta_g$ vary, and that variation is what a hierarchical model describes. In your forecasting model, for fixed parameters, the model's mean for day $t$ is $E[y_t \mid t, \text{holidays}, X_t] = g(t) + s(t) + h(t) + X_t\beta$: a conditional expectation given everything known about that day.

"$E[X \mid Y]$ and $E[X \mid Y = y]$ mean the same thing."

$E[X \mid Y = y]$ is the mean of $X$ in the group where $Y$ equals the specific value $y$, a number. $E[X \mid Y]$ is the random variable $h(Y)$ that reports the mean of whatever group $Y$ falls into.

Model answer: "Conditioning on an event gives a number; conditioning on a random variable gives a random variable, a function of that variable. That is why we can take its expectation and its variance, which is exactly what the laws of total expectation and total variance do."

$E[X \mid Y = y] = \sum_x x\,p(x \mid y)$: a number (the mean of group $y$).

$E[X \mid Y] = h(Y)$: a random variable, "replace each point by its group mean".

Trap: it is random only through $Y$; it is the best squared-error predictor of $X$ from $Y$.

Quick check: if $X$ and $Y$ are independent, what does $E[X \mid Y]$ look like?

Every group has the same distribution of $X$, so every group mean equals $E[X]$. $E[X \mid Y]$ is then the constant $E[X]$: a "random variable" that does not vary, with variance 0 (no between-group spread).

The law of total expectation: the overall mean is a weighted average of group means core

A school has two classes. Class A has 10 students and averages 70 points; class B has 30 students and averages 90. What is the school average? Not $(70 + 90)/2 = 80$, because class B has three times as many students. The right answer weights each class by its size: $(10\times70 + 30\times90)/40 = (700 + 2700)/40 = 85$.

That is the whole law. To get the overall mean, take each group's mean and weight it by how likely (how big) the group is. In symbols, the overall mean is the mean of the random variable $E[X \mid Y]$ from the last section.

Three ways to say it:

  • Picture: walk down a probability tree: each branch carries its probability and its group mean; multiply along the branches and add.
  • Numbers: $0.6\times1.5 + 0.4\times2.2 = 1.78$ items per basket.
  • Slogan: the average of the group averages, weighted by group size, is the overall average.

(a) Baskets.

  1. Group means: $E[X \mid \text{mobile}] = 1.5$, $E[X \mid \text{desktop}] = 2.2$. Group probabilities: $0.6$, $0.4$.
  2. Weighted average: $0.6\times1.5 + 0.4\times2.2 = 0.9 + 0.88 = 1.78$.
  3. Check with the marginal distribution $(0.44, 0.34, 0.22)$: $1(0.44) + 2(0.34) + 3(0.22) = 0.44 + 0.68 + 0.66 = 1.78$. ✓

(b) Revenue per visitor (from Chapter 4.5): split visitors into buyers (5%, average order 40) and non-buyers (95%, revenue 0). $E[R] = 0.05\times40 + 0.95\times0 = 2$ dollars per visitor.

(c) Order value by device. Mobile orders (60%) average 40 dollars; desktop orders (40%) average 70. Overall: $0.6\times40 + 0.4\times70 = 24 + 28 = 52$ dollars.

Law of total expectation (also called the law of iterated expectations, or the tower rule):

$$E[X] = E\big[\,E[X \mid Y]\,\big] = \sum_y P(Y = y)\;E[X \mid Y = y].$$

For a continuous $Y$: $E[X] = \int E[X \mid Y = y]\,f_Y(y)\,dy$.

Why it is true (discrete case, three steps):

  1. $\sum_y p_Y(y)\,E[X \mid Y = y] = \sum_y p_Y(y)\sum_x x\,p(x \mid y)$.
  2. $p_Y(y)\,p(x \mid y) = p(x, y)$ (multiplication rule), so this is $\sum_x x \sum_y p(x, y)$.
  3. $\sum_y p(x, y) = p_X(x)$ (the marginal), so we get $\sum_x x\,p_X(x) = E[X]$. ∎

The groups must cover every case exactly once (a partition: no overlaps, nothing left out). It is the "expected value version" of the law of total probability (Chapter 4.3).

Why do we need it?

We often know or estimate the mean of each group separately (each segment, each device, each day type) and need the overall mean, or we know the overall mean must be built from parts. Weighting by group size is the only correct way to combine them.

Where is it used?

Overall conversion and revenue from segment-level rates, revenue per visitor = conversion rate × order value given purchase, yearly averages from weekday and weekend averages, mixture models, and the derivation of the law of total variance.

How is it used?

Make a table of groups with their share and their mean; multiply share × mean and add. In pandas: (df.groupby("g").size() / len(df) * df.groupby("g").x.mean()).sum(), which equals df.x.mean().

P(mobile) = 0.6 P(desktop) = 0.4 E[X | mobile] = 1.5 E[X | desktop] = 2.2 0.6 × 1.5 = 0.90 0.4 × 2.2 = 0.88 E[X] = 0.90 + 0.88 = 1.78
The law of total expectation as a probability tree: multiply each branch probability by that group's mean, then add over branches.

Each coloured bar is a group: its position is the group mean, its height is the number of users. Drag the tops. The purple fulcrum sits at the size-weighted mean, and the beam is level there; the grey dashed line is the plain average of the three group means. Make one group tiny: the weighted mean ignores it, the plain average does not. Make all groups the same size: the two lines meet.

"The overall mean is the simple average of the group means."

Only when all groups are the same size. Otherwise weight each group mean by its share: big groups pull harder.

"If every group's mean goes up, the overall mean must go up."

Not if the shares change at the same time: moving people into a low-mean group can pull the overall mean down while every group improves. That is Simpson's paradox, next.

In your A/B framework with segments, a variant's overall conversion rate is $\sum_s P(\text{segment } s)\,\theta_{s}$: a weighted average of segment rates. When you report an overall effect from a segment-level (hierarchical) model, you must choose these weights (the actual traffic mix, or a fixed target mix) and use the same weights for both variants. In demand forecasting, average daily demand over whole weeks is the weighted average of the weekday mean (weight 5/7) and the weekend mean (weight 2/7).

$E[X] = E\big[E[X \mid Y]\big] = \sum_y P(Y=y)\,E[X \mid Y=y]$.

Overall mean = group means weighted by group shares (groups must form a partition).

Trap: a plain average of group means is wrong when group sizes differ.

Quick check: 80% of days are normal days with mean demand 100; 20% are promotion days with mean demand 250. What is the mean daily demand?

$0.8\times100 + 0.2\times250 = 80 + 50 = 130$. (The plain average $(100 + 250)/2 = 175$ is far off.)

Simpson's paradox: when a change in the mix beats a change in every group

Variant B converts better than A among desktop users, and also better among mobile users. Surely B converts better overall? Not necessarily. If most of B's traffic happened to come from mobile (which converts poorly for everyone) while most of A's came from desktop, then B's overall rate can be lower than A's.

The law of total expectation explains it: each overall rate is a weighted average of segment rates, and here the two variants use different weights. A comparison of weighted averages with different weights can point the opposite way from every group-level comparison.

Three ways to say it:

  • Picture: two seesaws with the same kinds of weights, placed in different amounts, balance at different points.
  • Numbers: B wins 22.5% vs 20% on desktop and 6% vs 5% on mobile, yet loses 9.3% vs 17% overall.
  • Slogan: before trusting an overall comparison, check the mix.
SegmentVariant AVariant BBetter
desktop160 / 800 = 20%45 / 200 = 22.5%B
mobile10 / 200 = 5%48 / 800 = 6%B
overall170 / 1000 = 17%93 / 1000 = 9.3%A
  1. A's users: 80% desktop, 20% mobile. Law of total expectation: $0.8\times20\% + 0.2\times5\% = 16\% + 1\% = 17\%$.
  2. B's users: 20% desktop, 80% mobile: $0.2\times22.5\% + 0.8\times6\% = 4.5\% + 4.8\% = 9.3\%$.
  3. B is better inside each segment, but B's weights put most of its traffic in the low-converting segment.
  4. Compare on a common mix, say 50/50: A $= 0.5\times20\% + 0.5\times5\% = 12.5\%$; B $= 0.5\times22.5\% + 0.5\times6\% = 14.25\%$. With equal weights, B wins again.

Simpson's paradox: a comparison (a difference, a trend, a correlation) that holds within every group can weaken, vanish or reverse when the groups are combined. In terms of this chapter:

$$\bar r_A = \sum_s w_{A,s}\,r_{A,s}, \qquad \bar r_B = \sum_s w_{B,s}\,r_{B,s},$$

where $r_{v,s}$ is variant $v$'s rate in segment $s$ and $w_{v,s}$ is the share of variant $v$'s users in segment $s$. If $w_A \ne w_B$, then $\bar r_B - \bar r_A$ mixes two effects: the within-segment difference and the difference in mix.

  • It happens when the grouping variable (here the segment) is related both to which variant a user saw and to the outcome. Such a variable is called a confounder (Chapter 5.12).
  • Standardization fixes the mix: compute both overall rates with the same weights $w_s$.
Why do we need it?

Aggregated dashboards are full of averages over changing mixes. Without this idea you can ship the worse variant, or conclude that a metric fell when only the traffic mix moved.

Where is it used?

A/B tests with traffic ramp-ups or allocations that change over time (days act as segments), before/after comparisons, marketing channel reports, medical studies (treatment success by severity), and university admissions (the famous Berkeley example).

How is it used?

Always break the comparison down by the main segments and look at each variant's mix. If the mixes differ, ask why. Then compare within segments, or standardize both variants to one common mix.

The segment rates are fixed (B is better in both segments). Move the two sliders, which set how much of each variant's traffic comes from mobile. With 20% vs 80% the overall bars flip: A looks better. Press Same mix: the paradox disappears. Turn on Compare on a 50/50 mix to see the standardized comparison, which is not fooled by the mix.

"B lost overall, so B is the worse variant."

First check each variant's segment mix. If the mixes differ, the overall comparison mixes the variant effect with a mix effect. Compare within segments, or standardize to a common mix.

"The segment-level comparison is always the right one."

It depends on why the mix differs. If the assignment process caused it (a ramp-up, a bug, a targeting rule), compare within segments. If the variant itself changes who arrives (B attracts more mobile users), the overall number may be the business-relevant one. Think about the cause before choosing.

"Simpson's paradox cannot happen in a randomized A/B test."

With a fixed random split, both variants have the same segment mix on average, so large reversals are unlikely. But it does happen when the split changes over time (ramp-ups: days become segments), when assignment is not random, or when you compare periods or self-selected groups.

In a segment-level model of your A/B framework, the overall rate for a variant is $\sum_s w_s\,\theta_{v,s}$. Computing $P(\text{B better overall} \mid D)$ from posterior draws needs explicit weights $w_s$; using the same $w_s$ for both variants (for example, the expected future traffic mix) avoids reporting a mix effect as a variant effect. If your experiments ramp traffic up over several days, treat the days as a segment when you sanity-check the overall numbers.

Overall rate = $\sum_s w_s\,r_s$ (law of total expectation). Different weights ⇒ comparisons can flip.

Fix: compare within segments or standardize to one common mix; ask why the mix differs.

Trap: "better in every segment" does not imply "better overall" when the mixes differ.

Quick check: in the example, what share of B's traffic on mobile would make B's overall rate equal to A's 17%?

Solve $(1 - w)\,22.5\% + w\,6\% = 17\%$: $22.5 - 16.5w = 17$, so $w = 5.5/16.5 \approx 0.33$. With more than about 33% of B's traffic on mobile (and A's mix fixed at 20% mobile), B looks worse overall.

The law of total variance: total = within groups + between groups core

Why do order values vary? For two separate reasons. First, even among mobile users alone, some orders are small and some are big: that is spread within a group. Second, mobile orders are typically smaller than desktop orders: the group averages themselves differ, which is spread between groups. The total spread is made of exactly these two parts, and the law of total variance says they simply add.

Picture a single order. Its distance from the overall average can be walked in two steps: from the overall average to its group's average (a between-group step), then from the group's average to the order itself (a within-group step). When we square and average, the two steps add up cleanly, with nothing left over.

Three ways to say it:

  • Picture: squash every point onto its group mean (between part); what you squashed away is the within part.
  • Numbers: order values: within 150 + between 216 = total 366.
  • Slogan: total spread = average spread inside the groups + spread of the group averages.

(a) Order value by device. Mobile: 60% of orders, mean 40, SD 10. Desktop: 40%, mean 70, SD 15. Overall mean (last section) = 52.

  1. Within (average of the group variances, weighted by share): $0.6\times10^2 + 0.4\times15^2 = 0.6\times100 + 0.4\times225 = 60 + 90 = 150$.
  2. Between (variance of the group means around 52): $0.6\times(40 - 52)^2 + 0.4\times(70 - 52)^2 = 0.6\times144 + 0.4\times324 = 86.4 + 129.6 = 216$.
  3. Total $= 150 + 216 = 366$, so the overall SD is $\sqrt{366} \approx 19.1$ dollars.
  4. Check the hard way with $E[X^2] - (E[X])^2$: inside each group $E[X^2 \mid Y] = Var + \text{mean}^2$, so $1700$ (mobile) and $5125$ (desktop). $E[X^2] = 0.6\times1700 + 0.4\times5125 = 1020 + 2050 = 3070$, and $3070 - 52^2 = 3070 - 2704 = 366$. ✓

(b) Baskets (the table example). Within: $0.6\times0.45 + 0.4\times0.56 = 0.27 + 0.224 = 0.494$. Between: $0.1176$ (computed in the conditional expectation section). Total $= 0.6116$. Check: $E[X^2] = 1(0.44) + 4(0.34) + 9(0.22) = 3.78$, and $3.78 - 1.78^2 = 3.78 - 3.1684 = 0.6116$. ✓ Here most of the spread (81%) is within the devices.

Law of total variance. For random variables $X$ and $Y$ (with $Var(X)$ finite):

$$Var(X) \;=\; \underbrace{E\big[Var(X \mid Y)\big]}_{\text{within-group}} \;+\; \underbrace{Var\big(E[X \mid Y]\big)}_{\text{between-group}} .$$
  • $E[Var(X \mid Y)] = \sum_y P(Y=y)\,Var(X \mid Y=y)$: the average spread inside the groups (sometimes called the "unexplained" part).
  • $Var(E[X \mid Y]) = \sum_y P(Y=y)\,\big(E[X \mid Y=y] - E[X]\big)^2$: the spread of the group means (the part "explained" by knowing the group).
  • The share $Var(E[X \mid Y])/Var(X)$ is the fraction of the variance explained by the grouping. Both parts are $\ge 0$, so $E[Var(X \mid Y)] \le Var(X)$: on average, knowing $Y$ can only reduce the spread (one particular group can still be more spread out than the whole population).

Derivation (six short steps, using only earlier rules):

  1. Shortcut formula: $Var(X) = E[X^2] - (E[X])^2$.
  2. Total expectation applied to $X^2$: $E[X^2] = E\big[E[X^2 \mid Y]\big]$.
  3. The shortcut formula inside one group: $E[X^2 \mid Y] = Var(X \mid Y) + \big(E[X \mid Y]\big)^2$.
  4. So $E[X^2] = E\big[Var(X \mid Y)\big] + E\big[(E[X \mid Y])^2\big]$.
  5. Total expectation applied to $X$: $(E[X])^2 = \big(E\big[E[X \mid Y]\big]\big)^2$.
  6. Subtract step 5 from step 4: $Var(X) = E\big[Var(X \mid Y)\big] + \Big\{E\big[(E[X \mid Y])^2\big] - \big(E\big[E[X \mid Y]\big]\big)^2\Big\}$, and the curly bracket is the shortcut formula for $Var\big(E[X \mid Y]\big)$. ∎

Data version. For data in groups $g$ with $n_g$ points, group means $\bar x_g$ and overall mean $\bar x$, the same split holds exactly for sums of squares:

$$\underbrace{\sum_{g}\sum_{i}(x_{gi} - \bar x)^2}_{SST\ (\text{total})} = \underbrace{\sum_{g}\sum_{i}(x_{gi} - \bar x_g)^2}_{SSW\ (\text{within})} + \underbrace{\sum_g n_g(\bar x_g - \bar x)^2}_{SSB\ (\text{between})}.$$

(Write $x_{gi} - \bar x = (x_{gi} - \bar x_g) + (\bar x_g - \bar x)$ and square; the cross terms add to zero because deviations from $\bar x_g$ sum to zero inside each group.) This identity is the heart of ANOVA (Chapter 5.9).

Why do we need it?

To know where variation comes from. If most of it is between groups, modelling the groups pays off; if most is within, the groups barely matter. It also gives the variance of a mixture or a two-level model without any integration.

Where is it used?

Hierarchical models ($\tau^2$ between groups, $\sigma^2$ within), ANOVA and $R^2$, overdispersion (Poisson counts with varying rates give the Negative Binomial), random-effects models, mixture models, and splitting forecast uncertainty into parameter and observation parts.

How is it used?

Per group compute share, mean and variance (df.groupby("g").x.agg(["size", "mean", "var"])). Within = share-weighted average of the variances; between = share-weighted variance of the means. Use ddof=0 throughout for the split to match the total exactly.

overall mean x̄ = 52 desktop mean x̄_g = 70 one order x = 90 between: 70 − 52 = 18 within: 90 − 70 = 20 total: 90 − 52 = 38 = 18 + 20
Each deviation from the overall mean is a between step plus a within step. After squaring and averaging, the cross terms cancel, so total variance = between + within.

Pick a group to edit, then move its mean, SD and size. The coloured curves are the groups (each scaled by its share); the black curve is everything together. The bar in the readout splits the total variance into within (grey) and between (purple). Press Same means: the between part drops to 0. Press Tight groups far apart: almost everything is between. Make one group's size 0: it vanishes from both parts.

Eight orders in two groups (drag them along their rows). Thin coloured lines join each order to its group mean (within parts); the purple dashed line is the overall mean. The bottom row shows the last order you dragged: its total deviation (black) is its between step (purple) plus its within step (coloured). The readout checks that the sums of squares add up exactly, whatever you do. Drag the two groups on top of each other: SSB goes to almost 0.

"Total variance = variance of group 1 + variance of group 2."

The within part is a weighted average of the group variances (not their sum), and you must add the between part: the spread of the group means.

"The between-group variance is the variance of the group averages I observed."

In a finite sample each observed group average also carries noise from the few points behind it, so observed averages vary more than the true group means: $Var(\bar x_g) = \tau^2 + \sigma^2/n$ in the two-level model below. The law itself is about the true conditional means.

"The split only works for Normal data."

It holds for any distribution with a finite variance: counts, yes/no data, skewed amounts. No Normality is needed anywhere in the derivation.

When you analyse a metric across segments in the A/B framework, the law tells you how much of its variation is between segments (real differences in segment means) and how much is within (user-to-user noise). A large between share means segments matter and a segment-level model is worth it; a tiny one means pooling the segments loses little. In a hierarchical model these two parts become the parameters $\tau^2$ and $\sigma^2$ (last section of this chapter).

"The law of total variance says the variance of $X$ is the average of the conditional variances."

That is only the within part. You must add the variance of the conditional means.

Model answer: "$Var(X) = E[Var(X \mid Y)] + Var(E[X \mid Y])$: the expected within-group variance plus the variance of the group means. In a hierarchical model with $y \sim N(\theta_g, \sigma^2)$ and $\theta_g \sim N(\mu, \tau^2)$, it gives $Var(y) = \sigma^2 + \tau^2$, and $\tau^2/(\tau^2 + \sigma^2)$ is the share of variance due to the groups."

$Var(X) = E[Var(X \mid Y)] + Var(E[X \mid Y])$ = within + between.

Data: $SST = SSW + SSB$, with $SSB = \sum n_g(\bar x_g - \bar x)^2$ (exact, cross terms vanish).

Trap: within is a weighted average of group variances; don't forget the between part.

Quick check: two equally large groups both have SD 5; their means are 10 and 20. What is the total variance?

Overall mean 15. Within $= 0.5\times25 + 0.5\times25 = 25$. Between $= 0.5\times(10-15)^2 + 0.5\times(20-15)^2 = 25$. Total $= 50$, SD $\approx 7.07$.

Mixtures of segments, and why mixing makes counts overdispersed

Pour the orders from all devices into one histogram and you get a mixture: a blend of the group distributions, each in proportion to its share. Even if every group is a neat bell curve, the blend can have two humps, a long tail, or heavy tails. And you can get its mean and variance from the two laws, without integrating anything.

The same thing happens with counts. Suppose orders on any given day are Poisson (spread equal to the mean), but the day's underlying rate changes from day to day (weather, promotions, paydays). Then the counts over many days are a mixture of Poissons, and the between-day part of the variance comes on top: the variance becomes larger than the mean. That is called overdispersion.

Three ways to say it:

  • Picture: several bell curves, each shrunk by its share, stacked into one outline.
  • Numbers: a Poisson rate that is 20 on half the days and 40 on the others gives counts with mean 30 but variance 130.
  • Slogan: mixing groups adds the between-group spread on top of the within-group spread.

(a) Order values: 60% from $N(40, 10^2)$, 40% from $N(70, 15^2)$.

  1. Mean (total expectation) $= 52$; variance (total variance) $= 150 + 216 = 366$; SD $\approx 19.1$.
  2. Probability of an order above 80 dollars, by total probability: $0.6\,P(N(40,10^2) \gt 80) + 0.4\,P(N(70,15^2) \gt 80) \approx 0.6(0.00003) + 0.4(0.2525) \approx 0.101$.
  3. A single Normal with the same mean and SD, $N(52, 19.1^2)$, gives only $P(\gt 80) \approx 0.072$. Same mean and variance, different shape, different tail answer.

(b) Overdispersed counts. Daily orders are Poisson with rate $\lambda$; $\lambda = 20$ on half the days, $40$ on the other half.

  1. Mean: $E[X] = E\big[E[X \mid \lambda]\big] = E[\lambda] = 0.5(20) + 0.5(40) = 30$.
  2. Within: a Poisson's variance equals its rate, so $E[Var(X \mid \lambda)] = E[\lambda] = 30$.
  3. Between: $Var(E[X \mid \lambda]) = Var(\lambda) = 0.5(20-30)^2 + 0.5(40-30)^2 = 100$.
  4. Total: $Var(X) = 30 + 100 = 130$, more than four times the mean. A plain Poisson model would claim 30.

A mixture distribution: first pick a group $Y = g$ with probability $w_g$ (the mixture weights, $\sum w_g = 1$), then draw $X$ from that group's distribution $f_g$. Its density (or PMF) is $f(x) = \sum_g w_g\,f_g(x)$, and the two laws give

$$E[X] = \sum_g w_g\,\mu_g, \qquad Var(X) = \underbrace{\sum_g w_g\,\sigma_g^2}_{\text{within}} + \underbrace{\sum_g w_g\,(\mu_g - E[X])^2}_{\text{between}} .$$

Overdispersion. If $X \mid \lambda \sim Poisson(\lambda)$ and the rate $\lambda$ itself varies, then

$$E[X] = E[\lambda], \qquad Var(X) = E[\lambda] + Var(\lambda) \;\ge\; E[X],$$

with equality only when $\lambda$ never varies. If $\lambda$ follows a Gamma distribution with mean $\mu$ and variance $\mu^2/\alpha$, the counts follow exactly the Negative Binomial with $Var(X) = \mu + \mu^2/\alpha$ (the NB2 form; all its parameterizations are in Chapter 4.8).

Why do we need it?

Real populations are blends of segments, days and users. Mixtures explain shapes that a single standard distribution cannot (two humps, extra-heavy tails), and the overdispersion formula explains why count data are almost always more variable than a Poisson model allows.

Where is it used?

Gaussian mixture models for clustering, the Negative Binomial as a gamma–Poisson mixture, the Student-t as a scale mixture of Normals (4.9), zero-inflated models (4.8), and posterior predictive distributions (a mixture over parameter draws, 6.1).

How is it used?

Check counts with y.var() / y.mean(): a ratio well above 1 signals overdispersion and points to a Negative Binomial or a model with group or day effects. For bimodal continuous data, look for a hidden grouping variable and model it, instead of forcing one bell curve.

2000 orders come from two segments (blue and orange) with the same inside-segment SD. The black curve is the true mixture; the red dashed curve is a single Normal with the same mean and SD. Pull the segments apart: two humps appear, and the single Normal puts its peak exactly in the valley. Shrink the distance to 0: the mixture is one bell again. Change the share of B and watch the overall mean slide (total expectation).

Each of 365 days gets its own rate $\lambda$ (drawn from a Gamma distribution with the chosen mean and SD), and then a Poisson count with that rate. Blue bars: the simulated counts. Orange stems: a plain Poisson with the same mean. Green stems: the Negative Binomial that the mixture produces exactly. Set the day-to-day SD to 0: blue, orange and green agree. Raise it: the counts spread far beyond the Poisson, and the variance/mean ratio climbs.

"Two humps in a histogram mean the data are weird or broken."

Usually it means two groups are mixed. Find the grouping variable (device, country, user type) and look at each group's conditional distribution.

"If daily counts have variance much bigger than the mean, the Poisson model is fine because the mean is right."

The mean can be right while the spread is badly wrong. Too-narrow predictive intervals and overconfident decisions follow. Use a Negative Binomial or model the source of the varying rate.

"A Normal with the right mean and variance describes a mixture well enough."

It can get tail probabilities and the location of the peak quite wrong (0.072 vs 0.101 above 80 dollars in the example).

This is the usual reason to offer a Negative Binomial likelihood, as your forecasting model does: day-to-day changes in the underlying demand rate that the trend, seasonality, holidays and regressors do not capture behave like a varying $\lambda$, which makes the counts overdispersed, $Var = \mu + \mu^2/\alpha$. In the A/B framework, a Poisson likelihood for a count metric assumes variance = mean for each user; if users have different rates, the pooled counts are overdispersed, so check var/mean before trusting Poisson intervals (Chapter 4.8).

Mixture: $f = \sum w_g f_g$; mean $\sum w_g\mu_g$; variance $\sum w_g\sigma_g^2 + \sum w_g(\mu_g - \mu)^2$.

Poisson with varying rate: $Var(X) = E[\lambda] + Var(\lambda) \ge E[X]$ (overdispersion); Gamma rates ⇒ NB2.

Trap: matching mean and variance does not match the shape (humps, tails).

Quick check: daily counts have mean 30 and variance 130. If you fit NB2, what dispersion α do you get?

$130 = 30 + 30^2/\alpha$, so $30^2/\alpha = 100$ and $\alpha = 900/100 = 9$. (Equivalently, the varying rate has variance 100, SD 10.)

Between-group $\tau^2$ and within-group $\sigma^2$: the law inside hierarchical models core

A hierarchical model generates data in two steps, exactly like the two laws. Step one: each group (segment, store, country) gets its own true mean $\theta_g$, drawn from a population of group means with centre $\mu$ and spread $\tau$. Step two: each observation in group $g$ scatters around $\theta_g$ with spread $\sigma$. So $\tau$ says how different the groups really are, and $\sigma$ says how noisy a single observation is.

The law of total variance then gives the spread of one random observation: $\sigma^2 + \tau^2$. And it explains a trap: the averages you observe per group vary more than the true $\theta_g$, because each average still carries some within-group noise, $\sigma^2/n$. Small groups' averages are the noisiest, which is why hierarchical models pull them toward the overall mean (shrinkage, Chapter 6.6).

Three ways to say it:

  • Picture: a two-level dart game: first pick where a group's target sits, then throw darts around that target.
  • Numbers: $\tau = 3$, $\sigma = 4$ gives one observation an SD of $\sqrt{9 + 16} = 5$, and 36% of its variance comes from the group.
  • Slogan: $\tau$ is real group difference; $\sigma$ is noise; observed group averages mix both.

Groups have true means $\theta_g \sim N(50, 3^2)$ (so $\mu = 50$, $\tau = 3$), and observations $y_{gi} \mid \theta_g \sim N(\theta_g, 4^2)$ (so $\sigma = 4$).

  1. Within: $E[Var(y \mid \theta_g)] = \sigma^2 = 16$.
  2. Between: $Var(E[y \mid \theta_g]) = Var(\theta_g) = \tau^2 = 9$.
  3. Total: $Var(y) = 16 + 9 = 25$, SD $= 5$.
  4. Share due to the groups: $\tau^2/(\tau^2 + \sigma^2) = 9/25 = 0.36$. This share is also the correlation between two observations from the same group (they share $\theta_g$), called the intraclass correlation.
  5. Average of $n = 4$ observations in a group: $\bar y_g = \theta_g + (\text{average of 4 noises})$, so $Var(\bar y_g) = \tau^2 + \sigma^2/4 = 9 + 4 = 13$. With $n = 100$: $9 + 0.16 = 9.16$, close to $\tau^2$.

A two-level (hierarchical) Normal model:

$$\theta_g \sim N(\mu, \tau^2) \quad (g = 1, \dots, G), \qquad y_{gi} \mid \theta_g \sim N(\theta_g, \sigma^2) \quad (i = 1, \dots, n_g),$$

with all draws independent. The words:

  • Group parameters $\theta_g$: one true mean per group.
  • Hyperparameters $\mu$ and $\tau$: the parameters of the distribution that the group parameters come from. In a Bayesian model they get priors of their own (hyperpriors), taught in Chapter 6.5.
  • $\tau^2$ = between-group variance; $\sigma^2$ = within-group variance.

By the law of total variance (conditioning on $\theta_g$):

$$Var(y_{gi}) = \sigma^2 + \tau^2, \qquad Corr(y_{gi}, y_{gj}) = \frac{\tau^2}{\tau^2 + \sigma^2}\ (i \ne j), \qquad Var(\bar y_g) = \tau^2 + \frac{\sigma^2}{n_g}.$$

(The syllabus writes $\theta_g \sim N(\mu, \tau)$ with $\tau$ as a standard deviation, which is how NumPyro's dist.Normal(mu, tau) reads; in maths we write the variance $\tau^2$ inside $N(\cdot,\cdot)$.)

Why do we need it?

To reason about grouped data honestly: how much do segments truly differ ($\tau$), how much is noise ($\sigma$), and how much should we trust a small group's average. Without the split we mistake noise for real segment differences.

Where is it used?

Hierarchical (multilevel) models with partial pooling, random-effects meta-analysis, the eight-schools example, intraclass correlation and design effects for clustered A/B tests (5.10), and empirical-Bayes shrinkage of segment estimates.

How is it used?

Simulate from the model to see what $\tau$ and $\sigma$ imply; estimate them by fitting the hierarchical model (or roughly: pooled within-group variance for $\sigma^2$, and variance of group averages minus $\sigma^2/n$ for $\tau^2$). Compare $\tau^2/(\tau^2 + \sigma^2)$ across metrics to see where segments matter.

μ (centre), τ (between SD) θ₁ (segment 1) θ₂ (segment 2) θ_G (segment G) θ_g ~ N(μ, τ²): between y₁₁ … y₁ₙ y₂₁ … y₂ₙ y_G1 … y_Gn y ~ N(θ_g, σ²): within … Var(y) = σ² + τ² (law of total variance, conditioning on θ_g)
The two-level model: hyperparameters μ, τ generate one true mean per group (between-group variation τ²); each group's observations scatter around it (within-group variation σ²).

Eight segments (rows). Green ticks are the true segment means $\theta_g$, orange ticks are the observed averages, blue dots are observations, purple is the overall centre. Set $n = 2$: the orange ticks scatter much more than the green ones (observed averages carry noise). Set $n = 50$: orange sits on green. Set $\tau = 0$: the segments are truly identical, yet the orange ticks still differ. Then switch the view to Global scaler (the between share is unchanged) and to Per-group z-scores (every segment collapses to mean 0: the between-group signal is destroyed).

"$\tau$ is the SD of the segment averages I observed."

The observed averages have variance $\tau^2 + \sigma^2/n$: they include noise. With small $n$, segments can look very different even when $\tau = 0$. Estimating $\tau$ needs the within-group noise subtracted, which is what the hierarchical model does.

"Standardizing each segment separately is harmless preprocessing."

Per-group z-scores force every segment mean to 0 and every segment SD to 1, so the between-group variance in the data becomes exactly 0. A model fit afterwards can only conclude "no segment differences". Use one global scaler instead.

"If 36% of the variance is between groups, the groups explain 36% of every user's behaviour."

It is a statement about variance across the whole population, not about individuals: knowing the group removes 36% of the variance of a random observation, on average.

Your A/B framework's hierarchical partial pooling is this model: segment-level parameters $\theta_g$ drawn around a global mean with between-segment spread $\tau$, and observations around each $\theta_g$ with within-segment noise. The law of total variance tells you why small segments' raw averages are unreliable ($\sigma^2/n$ is big) and why the model shrinks them toward the global mean (Chapter 6.5, 6.6). It also explains the syllabus warning about the global scaler: one shared mean and SD keeps the between/within split intact, while normalizing each group by its own mean and SD would set the between-group variance to zero and destroy the group effect you want to measure (Chapter 4.18).

"Partial pooling shrinks every segment by the same amount."

Shrinkage depends on how noisy each segment's average is compared with the real spread between segments: segments with small $n$ (big $\sigma^2/n$ relative to $\tau^2$) are pulled more, large segments hardly at all.

Model answer: "By the law of total variance, an observed segment average varies by $\tau^2 + \sigma^2/n$. Only $\tau^2$ is real segment difference; $\sigma^2/n$ is noise. The hierarchical model weighs the two, so small segments borrow strength from the global mean while large segments mostly keep their own average."

$\theta_g \sim N(\mu, \tau^2)$, $y \mid \theta_g \sim N(\theta_g, \sigma^2)$ ⇒ $Var(y) = \sigma^2 + \tau^2$; ICC $= \tau^2/(\tau^2+\sigma^2)$.

Observed group averages: $Var(\bar y_g) = \tau^2 + \sigma^2/n_g$ (noise inflates them, most for small $n_g$).

Trap: per-group z-scoring erases between-group variance; use a global scaler.

Quick check: $\tau = 2$, $\sigma = 6$, and a segment has $n = 9$ users. What is the variance of its observed average, and how much of it is noise?

$Var(\bar y_g) = 2^2 + 6^2/9 = 4 + 4 = 8$. Half of it ($\sigma^2/n = 4$) is noise, so this segment's average deserves substantial shrinkage toward the global mean.

Recap, cheat sheet and practice

  • A conditional distribution $p(x \mid y) = p(x,y)/p_Y(y)$ is the distribution inside one group: keep the row, divide by its total.
  • $E[X \mid Y = y]$ is a number (a group's mean); $E[X \mid Y]$ is a random variable, one mean per group ("squash each point onto its group mean").
  • Total expectation: $E[X] = E[E[X \mid Y]] = \sum_y P(y)\,E[X \mid y]$, a weighted average of group means.
  • Simpson's paradox: comparing weighted averages with different weights can reverse every within-group comparison. Check the mix; compare within segments or on a common mix.
  • Total variance: $Var(X) = E[Var(X \mid Y)] + Var(E[X \mid Y])$, within + between; in data, $SST = SSW + SSB$.
  • Mixtures get their mean and variance from the two laws; a Poisson with a varying rate is overdispersed: $Var = E[\lambda] + Var(\lambda)$, and Gamma rates give the Negative Binomial.
  • In a hierarchical model, $Var(y) = \sigma^2 + \tau^2$, observed group averages vary by $\tau^2 + \sigma^2/n$, and per-group scaling destroys $\tau^2$.

Cheat sheet

IdeaFormulaIn words
Conditional distribution$p(x \mid y) = p(x,y)/p_Y(y)$one row, rescaled to sum to 1
Marginal$p_X(x) = \sum_y p(x,y)$ignore $Y$ (row/column totals)
Conditional mean$E[X \mid Y=y] = \sum_x x\,p(x \mid y)$mean inside group $y$
$E[X \mid Y]$$h(Y)$, $h(y) = E[X \mid Y=y]$random: one mean per group
Total expectation$E[X] = \sum_y P(y)\,E[X \mid y]$weighted average of group means
Total variance$E[Var(X \mid Y)] + Var(E[X \mid Y])$within + between
Sums of squares$SST = SSW + \sum n_g(\bar x_g - \bar x)^2$data version (ANOVA)
Mixture variance$\sum w_g\sigma_g^2 + \sum w_g(\mu_g - \mu)^2$no integration needed
Poisson with varying rate$Var = E[\lambda] + Var(\lambda)$overdispersion; Gamma ⇒ NB2
Hierarchical Normal$Var(y) = \sigma^2 + \tau^2$, $Var(\bar y_g) = \tau^2 + \sigma^2/n$real differences + noise
Code it · Python

import numpy as np
import pandas as pd
from scipy import stats

rng = np.random.default_rng(1)

# 1. Order values from two devices (60% mobile, 40% desktop)
n_mob, n_desk = 6000, 4000
df = pd.DataFrame({
    "device": ["mobile"] * n_mob + ["desktop"] * n_desk,
    "value": np.concatenate([rng.normal(40, 10, n_mob), rng.normal(70, 15, n_desk)]),
})
g = df.groupby("device")["value"]
share = g.size() / len(df)            # P(Y = y)
means = g.mean()                      # E[X | Y = y]: one mean per group
vars_ = g.var(ddof=0)                 # Var(X | Y = y)
overall = df["value"].mean()

# Law of total expectation: overall mean = weighted average of group means
print((share * means).sum(), overall)          # 51.86 51.86 (true value 52)

# Law of total variance: within + between = total (exact with ddof=0)
within = (share * vars_).sum()
between = (share * (means - overall) ** 2).sum()
print(within, between, within + between, df["value"].var(ddof=0))
# 149.3 213.8 363.0 363.0   (true values 150 + 216 = 366)

# E[X | Y] as a random variable: replace every value by its group mean
df["E_X_given_Y"] = g.transform("mean")
print(df["E_X_given_Y"].mean(), df["E_X_given_Y"].var(ddof=0))   # 51.86 213.8 (= overall mean, = between)

# 2. Simpson's paradox
t = pd.DataFrame({"variant": ["A", "A", "B", "B"],
                  "segment": ["desktop", "mobile", "desktop", "mobile"],
                  "users": [800, 200, 200, 800], "conv": [160, 10, 45, 48]})
t["rate"] = t.conv / t.users
print(t.pivot(index="segment", columns="variant", values="rate"))   # B higher in both rows
tot = t.groupby("variant")[["conv", "users"]].sum()
print(tot.conv / tot.users)                    # A 0.170, B 0.093: A higher overall
print(t.groupby("variant").rate.mean())        # common 50/50 mix: A 0.125, B 0.1425

# 3. Overdispersion from day-to-day rate changes (gamma-Poisson = NB2)
mu, alpha, days = 30, 9, 200_000
lam = rng.gamma(shape=alpha, scale=mu / alpha, size=days)   # E = 30, Var = mu**2/alpha = 100
y = rng.poisson(lam)
print(y.mean(), y.var())                       # 29.97 129.2  (theory 30 and 30 + 100 = 130)
nb = stats.nbinom(n=alpha, p=alpha / (alpha + mu))   # SciPy's (n, p) form of NB2(mu=30, alpha=9)
print(nb.mean(), nb.var())                     # 30.0 130.0 (up to rounding)

# 4. Hierarchical model: Var(y) = sigma^2 + tau^2
mu0, tau, sigma, G, n = 50, 3, 4, 2000, 20
theta = rng.normal(mu0, tau, G)                # one true mean per group
yy = rng.normal(theta[:, None], sigma, (G, n)) # n observations per group
print(yy.var(), tau**2 + sigma**2)             # 25.02 25
print(yy.mean(axis=1).var(), tau**2 + sigma**2 / n)   # 9.82 9.8: observed averages vary more than tau^2 = 9
z = (yy - yy.mean(axis=1, keepdims=True)) / yy.std(axis=1, keepdims=True)
print(z.mean(axis=1).var())                    # about 1e-30: per-group z-scores erase the group effect
zg = (yy - yy.mean()) / yy.std()
print(zg.mean(axis=1).var(), (tau**2 + sigma**2 / n) / (tau**2 + sigma**2))   # 0.393 0.392: global scaler keeps it
Test yourself

1. Group 1 has 30 users with mean 10; group 2 has 10 users with mean 20. What is the overall mean?

Weight by group share: $0.75\times10 + 0.25\times20 = 7.5 + 5 = 12.5$. The plain average 15 ignores the sizes.

2. Which statement about $E[X \mid Y]$ is correct?

$E[X \mid Y] = h(Y)$ with $h(y) = E[X \mid Y = y]$. It is random because $Y$ is, and its mean is $E[X]$.

3. Within-group variances are 4 (share 0.5) and 16 (share 0.5); group means are 0 and 6. What is $Var(X)$?

Within $= 0.5\times4 + 0.5\times16 = 10$. Overall mean 3; between $= 0.5\times9 + 0.5\times9 = 9$. Total $= 19$.

4. Variant B converts better than A in every segment but worse overall. The most likely explanation is…

Simpson's paradox. Each overall rate is a weighted average of segment rates; different weights can flip the comparison. Compare within segments or on a common mix.

5. In a hierarchical model with $\tau = 3$ and $\sigma = 4$, what is the SD of a single observation, and what share of its variance comes from the groups?

$Var = \sigma^2 + \tau^2 = 16 + 9 = 25$, SD $= 5$. Between share $= 9/25 = 36\%$.

6. Daily counts are Poisson given the day's rate, and the rate varies from day to day. Compared with the mean, the variance of the counts is…

Total variance: $E[Var(X \mid \lambda)] + Var(E[X \mid \lambda]) = E[\lambda] + Var(\lambda)$. This is overdispersion; Gamma rates give the Negative Binomial.

Practice problems

A. Weekdays (5 of 7 days) have mean demand 100 and SD 10; weekend days (2 of 7) have mean 170 and SD 20. Find the overall mean, within, between and total variance.

Mean: $\tfrac57(100) + \tfrac27(170) = (500 + 340)/7 = 120$. Within: $\tfrac57(100) + \tfrac27(400) = (500 + 800)/7 \approx 185.7$. Between: $\tfrac57(100-120)^2 + \tfrac27(170-120)^2 = \tfrac57(400) + \tfrac27(2500) = (2000 + 5000)/7 = 1000$. Total $\approx 1185.7$, SD $\approx 34.4$. Most of the variance (84%) is the weekday/weekend difference, so a model without weekly seasonality would leave most of it unexplained.

B. Two equally large segments convert at 10% and 30%. Split the variance of a single user's 0/1 conversion into within and between, and check the total.

Overall $p = 0.2$, so $Var = 0.2\times0.8 = 0.16$. Within: $0.5(0.1\times0.9) + 0.5(0.3\times0.7) = 0.5(0.09 + 0.21) = 0.15$. Between: $0.5(0.1-0.2)^2 + 0.5(0.3-0.2)^2 = 0.01$. Total $0.15 + 0.01 = 0.16$. ✓ For a single 0/1 outcome the total variance is fixed by the overall $p$ (a mix of Bernoullis is again a Bernoulli), so segment differences add no extra variance; they only move some of it from the “within” part to the “between” part.

C. Revenue per visitor: mobile (70% of visitors) converts at 2% with average order 50; desktop (30%) converts at 5% with average order 80. What is the expected revenue per visitor?

Per segment (total expectation inside each): mobile $0.02\times50 = 1.0$; desktop $0.05\times80 = 4.0$. Overall: $0.7\times1.0 + 0.3\times4.0 = 0.7 + 1.2 = 1.9$ dollars per visitor.

D. Interview: "Some of our segments have only 20 users and show huge lifts. Why shouldn't we trust them?"

By the law of total variance, an observed segment average varies by $\tau^2 + \sigma^2/n$. With $n = 20$ the noise term $\sigma^2/20$ can dominate the real between-segment variation $\tau^2$, so extreme segment averages are mostly noise (and the most extreme ones are the ones you notice). A hierarchical model shrinks small segments toward the overall mean in proportion to how noisy they are.

E. Daily counts have mean 30 and variance 130. What does this tell you, and what NB2 dispersion matches it?

Variance/mean $\approx 4.3$, far above the Poisson value 1: the counts are overdispersed, consistent with a rate that varies across days by $Var(\lambda) = 130 - 30 = 100$. NB2: $130 = 30 + 30^2/\alpha$ gives $\alpha = 9$.

F. Interview: "Why is it a mistake to z-score each segment separately before fitting a model of segment effects?"

Per-group z-scoring subtracts each segment's own mean, so every segment ends with mean 0: $Var(E[X \mid \text{segment}])$ in the transformed data is exactly 0. Since total = within + between, the model now sees only within-segment variation and must conclude the segments are identical. A single global scaler is one linear transformation for everyone; it rescales both parts by the same factor and keeps their ratio.

Chapter 4.7 · Syllabus Module 9

Yes/no and category distributions: Bernoulli, Binomial, Categorical, Multinomial

A visitor converts or does not. A user picks the Free, Basic or Pro plan. Out of 1 000 visitors, some number converts. These are the simplest kinds of random data, and they are the heart of your A/B framework. This chapter teaches the four distributions that describe them, each with the same eight-question recipe, so that you can say exactly what each one assumes and when it breaks.

  • Use one recipe for every distribution: support → parameters → mean → variance → shape → assumptions → when used → relationships
  • Know the Bernoulli (one yes/no trial) and the Binomial (how many yeses in $n$ trials), including where $\binom{n}{k}$ comes from
  • State the Binomial's assumptions (fixed $n$, independent trials, the same $p$) and see what happens to the spread when they fail
  • Use the Normal approximation to the Binomial, and recognise when it fails (small $np$)
  • Know the Categorical (one pick among $K$ categories) and the Multinomial (counts of each category after $n$ picks), including why the counts are negatively correlated
  • See the four as one family, and place each one in your A/B framework (conversions, categorical metrics)

How to learn any distribution: the eight-question recipe core

What we need from earlier chapters: a random variable and its PMF (Chapter 4.4: the table of "value → probability" for a count), and the mean and variance (Chapter 4.5).

When you meet a new person you learn a few standard things: their name, where they live, what they do, who their family is. After that you can place them. Distributions are the same. There are dozens of them, but every one can be "met" with the same eight questions.

If you can answer all eight for a distribution, you understand it well enough to choose it for a model, defend the choice in an interview, and notice when it is the wrong choice. That is exactly what your syllabus asks for, so every distribution in this chapter and the next gets a recipe card with these eight answers.

Three ways to say it:

  • Picture: each distribution has an ID card with eight lines on it.
  • Numbers: "values 0 or 1, one knob $p$, mean $p$, variance $p(1-p)$" already tells you most of what a Bernoulli is.
  • Slogan: what can happen, what sets it, where is it, how wide, what shape, what story, what data, what relatives.

Warm-up: one roll of a fair die, run through the recipe.

  1. Support (the values that can happen): $\{1, 2, 3, 4, 5, 6\}$.
  2. Parameters (the knobs): none. A fair die is fully known. (A general "uniform on $a, \dots, b$" has knobs $a$ and $b$.)
  3. Mean: $(1+2+3+4+5+6)/6 = 21/6 = 3.5$.
  4. Variance: squared distances from 3.5 are $6.25, 2.25, 0.25, 0.25, 2.25, 6.25$; their average is $17.5/6 \approx 2.917$.
  5. Shape: flat. Every value has probability $1/6$.
  6. Assumptions (the story that makes it true): the die is balanced, so no face is preferred.
  7. When used: games, and "assign each user to one of 6 buckets at random".
  8. Relationships: it is a Categorical distribution (later in this chapter) with all six probabilities equal.

A distribution describes which values a random variable can take and how likely each one is. For a count (a discrete variable) it is given by a PMF $p(k) = P(X = k)$.

Most named distributions are really families: one formula with one or more parameters. Fix the parameters and you get one member of the family. We write $X \sim \text{Name}(\text{parameters})$, read "$X$ is distributed as…". The eight questions:

#QuestionMeaning in plain words
1SupportThe set of values with probability above zero. Example: $\{0, 1\}$, or $\{0, 1, \dots, n\}$, or all whole numbers $\ge 0$.
2ParametersThe knobs that pick one member of the family, and their allowed ranges (e.g. $0 \le p \le 1$).
3Mean$E[X]$, the long-run average (the balance point).
4Variance$Var(X)$, the average squared distance from the mean; its square root is the standard deviation.
5ShapeSymmetric or skewed (one long tail)? One peak or several? How does the shape change with the parameters?
6AssumptionsThe data-generating story under which this distribution is exactly right. When the story fails, the distribution is only an approximation.
7When usedReal data and models where it appears.
8RelationshipsSpecial cases, sums, limits and mixtures that connect it to other distributions.
Why do we need it?

Choosing a likelihood is one of the most important modelling decisions in both of your projects. A fixed checklist stops you from picking a distribution because "everyone uses it", and makes you check the support and the variance first.

Where is it used?

Choosing Binomial vs Multinomial metrics in the A/B framework, Normal vs Student-t vs Negative Binomial likelihoods in the forecasting model, GLM families (Chapter 5.14), and every "why did you choose this distribution?" interview question.

How is it used?

Look at your data: what values can it take (support)? Does its spread grow with its mean (variance)? Which story produced it (assumptions)? Then pick the distribution whose card matches, and say which assumption you are least sure about.

1 · Supportwhich values can happen? 2 · Parameterswhich knobs set it? 3 · Meanwhere is the centre? 4 · Variancehow wide? 5 · Shapeskewed? one peak? 6 · Assumptionswhat story makes it true? 7 · When usedwhich real data? 8 · Relationshipswhich relatives?
The recipe card. The top row is "the numbers" (what and where); the bottom row is "the story" (why, when, and how it connects). Most mistakes in practice come from the bottom row.

"A distribution is just a formula for probabilities."

A distribution is a model of a process. The formula is only correct when the story behind it is true (independent trials, the same probability…). The assumptions are part of the distribution, not a footnote.

"Parameters and statistics are the same kind of number."

A parameter (like the true conversion rate $p$) belongs to the distribution and is usually unknown. A statistic (like the observed rate $k/n$) is computed from a sample and changes from sample to sample (Chapter 4.1).

Every distribution: support → parameters → mean → variance → shape → assumptions → when used → relationships.

Choose a likelihood by matching the support and the variance first, then check the story.

Trap: the formula is only as good as its assumptions.

Quick check: what is the support of "the number of orders a shop gets tomorrow"?

All whole numbers $0, 1, 2, \dots$ with no fixed upper limit. That already rules out the Binomial (whose support stops at $n$) and points to the count distributions of Chapter 4.8.

Bernoulli: one yes/no trial core

A visitor lands on your checkout page. Either they buy (yes) or they do not (no). Before it happens, you do not know which. All you can say is "they buy with probability $p$".

We turn the answer into a number: 1 for yes, 0 for no. That 0/1 number is a Bernoulli random variable. The event we count as 1 is traditionally called a success, even when it is something bad (a churned user, a failed payment). "Success" just means "the outcome we code as 1".

A single attempt like this is called a trial.

Three ways to say it:

  • Picture: a bent coin that lands "1" with probability $p$ and "0" otherwise.
  • Numbers: with $p = 0.1$, about 10 of every 100 visitors convert, and the average of their 0/1 values is about 0.1.
  • Slogan: one trial, two outcomes, one knob $p$.

A visitor converts with probability $p = 0.1$. Let $X = 1$ if they convert and $X = 0$ if not.

  1. PMF: $P(X = 1) = 0.1$ and $P(X = 0) = 1 - 0.1 = 0.9$. They add up to 1.
  2. Mean: $E[X] = 0 \times 0.9 + 1 \times 0.1 = 0.1$. The mean of a 0/1 variable is the probability of a 1.
  3. $E[X^2]$: since $0^2 = 0$ and $1^2 = 1$, $X^2 = X$, so $E[X^2] = 0.1$.
  4. Variance: $Var(X) = E[X^2] - (E[X])^2 = 0.1 - 0.01 = 0.09$. This equals $p(1-p) = 0.1 \times 0.9$.
  5. Standard deviation: $\sqrt{0.09} = 0.3$.
  6. Compare a coin with $p = 0.5$: variance $0.5 \times 0.5 = 0.25$, the largest possible. A 50/50 outcome is the hardest to predict.

$X \sim \text{Bernoulli}(p)$ if $X$ takes the value 1 with probability $p$ and 0 with probability $1-p$. In one line:

$$P(X = x) = p^{x}(1-p)^{1-x}, \qquad x \in \{0, 1\}.$$

(Put in $x = 1$: you get $p$. Put in $x = 0$: you get $1 - p$.)

Recipe card · Bernoulli($p$)
Support$\{0, 1\}$
Parameters$p$ = probability of a 1, with $0 \le p \le 1$
Mean$p$
Variance$p(1-p)$; largest (0.25) at $p = 0.5$, zero at $p = 0$ or $1$
Shapetwo bars, heights $1-p$ and $p$
Assumptionsone trial with exactly two possible outcomes, coded 1 and 0
When useddid a user convert / click / churn / pay? Any single yes/no event; the label of a binary classifier
Relationships= Binomial with $n = 1$; = Categorical with $K = 2$ categories; a sum of independent Bernoullis with the same $p$ is Binomial; the Beta distribution is the usual prior for $p$ (Chapter 4.11)
Why do we need it?

It is the smallest possible random event, and bigger models are built from it. Writing a conversion as "$X \sim$ Bernoulli($p$)" separates the thing we never see (the rate $p$) from the thing we do see (0 or 1).

Where is it used?

Per-user conversion in A/B tests, logistic regression (each label is Bernoulli with $p$ from a sigmoid), binary cross-entropy loss (it is the Bernoulli negative log-likelihood, Chapter 5.2), and click-through modelling.

How is it used?

Code the outcome as 0/1. The sample mean of the 0/1 column is your estimate of $p$. In NumPyro: numpyro.sample("y", dist.Bernoulli(probs=p), obs=y) (or logits= for the log-odds).

visitor p = 0.1 1 − p = 0.9 buys → X = 1 leaves → X = 0 01 0.90.1 PMF of X
A Bernoulli turns a yes/no event into a 0/1 number. Its whole distribution is two bars whose heights add up to 1.

Set $p$ and press 100 visitors a few times. The orange bars are the true PMF; the blue bars are the fractions you actually observed, and they settle near the orange ones. On the right, the curve is the variance $p(1-p)$. Slide $p$ from 0.01 to 0.99 and watch the purple dot: the variance is largest at $p = 0.5$ and tiny near 0 or 1, because outcomes that almost always go one way are easy to predict.

"The variance of a conversion is bigger when the conversion rate is bigger."

The variance $p(1-p)$ grows only up to $p = 0.5$ and then shrinks again. A metric with $p = 0.9$ is exactly as noisy per user as one with $p = 0.1$.

"A success must be something good."

"Success" is only the outcome coded as 1. For a churn model, success = the user churned. Always say which outcome is the 1.

$X \sim \text{Bernoulli}(p)$: $P(X=1) = p$, $P(X=0) = 1-p$; $E[X] = p$, $Var(X) = p(1-p) \le 0.25$.

The mean of a 0/1 column is the observed rate.

Trap: variance peaks at $p = 0.5$; "success" is just the label for 1.

Quick check: a payment fails with probability 0.02. What are the mean and standard deviation of the 0/1 "failed" indicator?

Mean $= p = 0.02$. Variance $= 0.02 \times 0.98 = 0.0196$, so the standard deviation is $\sqrt{0.0196} = 0.14$.

Binomial: how many yeses in $n$ tries? core

Now five visitors arrive, and each one converts with probability $p = 0.2$, without affecting the others. How many of the five convert? It could be 0, 1, 2, 3, 4 or 5. The Binomial distribution gives the probability of each of these counts.

Think of it as flipping the same bent coin $n$ times and counting the 1s. Each flip is a Bernoulli trial; the Binomial is their total.

There is one subtle step. "Exactly 2 conversions out of 5" can happen in many orders (visitors 1 and 2, or 1 and 3, or 4 and 5…). Every order has the same probability, so we count the orders and multiply. That count of orders is the famous $\binom{n}{k}$.

Three ways to say it:

  • Picture: a ball falls through $n$ rows of pegs and bounces right with probability $p$ at each one; the bin it lands in is the number of right-bounces.
  • Numbers: with $n = 5$ and $p = 0.2$, one conversion is the most likely result (41%), zero happens 33% of the time, and all five only 0.03% of the time.
  • Slogan: count the yeses in $n$ independent tries that share the same $p$.

$n = 5$ visitors, each converts with $p = 0.2$. What is $P(\text{exactly } 2 \text{ convert})$? Write Y for "converts" and N for "does not".

  1. One particular order, say YYNNN: $0.2 \times 0.2 \times 0.8 \times 0.8 \times 0.8 = 0.2^2 \times 0.8^3 = 0.04 \times 0.512 = 0.02048$. (We may multiply because the visitors are independent.)
  2. Every order with two Y and three N has the same probability $0.02048$: the same numbers are multiplied, only in a different order.
  3. How many such orders? List them: YYNNN, YNYNN, YNNYN, YNNNY, NYYNN, NYNYN, NYNNY, NNYYN, NNYNY, NNNYY. That is 10. The shortcut is $\binom{5}{2} = \frac{5!}{2!\,3!} = \frac{120}{2 \times 6} = 10$.
  4. The orders cannot happen at the same time, so we add their probabilities: $P(X = 2) = 10 \times 0.02048 = 0.2048$.
  5. The same steps for every $k$ give the whole PMF: $P(0) = 0.3277$, $P(1) = 0.4096$, $P(2) = 0.2048$, $P(3) = 0.0512$, $P(4) = 0.0064$, $P(5) = 0.0003$. They add up to 1.
  6. Mean $= np = 5 \times 0.2 = 1$ conversion. Variance $= np(1-p) = 5 \times 0.2 \times 0.8 = 0.8$, so sd $= \sqrt{0.8} \approx 0.894$.

Let $X$ be the number of successes in $n$ independent Bernoulli($p$) trials. Then $X \sim \text{Binomial}(n, p)$ and

$$P(X = k) = \binom{n}{k} p^{k} (1-p)^{n-k}, \qquad k = 0, 1, \dots, n, \qquad \binom{n}{k} = \frac{n!}{k!\,(n-k)!}.$$
  • $\binom{n}{k}$ ("$n$ choose $k$", the binomial coefficient) counts the orders with exactly $k$ successes. $n! = n \times (n-1) \times \dots \times 1$ is "$n$ factorial", and $0! = 1$.
  • $p^k(1-p)^{n-k}$ is the probability of any one of those orders.

Why the mean and variance are what they are. Write $X = X_1 + \dots + X_n$ with each $X_i \sim$ Bernoulli($p$). By linearity of expectation (Chapter 4.5), $E[X] = p + \dots + p = np$. Because the trials are independent, their variances add: $Var(X) = p(1-p) + \dots + p(1-p) = np(1-p)$.

Recipe card · Binomial($n$, $p$)
Support$\{0, 1, \dots, n\}$: a bounded count (it can never exceed $n$)
Parameters$n$ = number of trials (a known whole number); $p$ = success probability per trial, $0 \le p \le 1$
Mean$np$ (and the observed proportion $X/n$ has mean $p$)
Variance$np(1-p)$ (and $X/n$ has variance $p(1-p)/n$)
Shapeone peak near $np$; symmetric when $p = 0.5$; a long right tail when $p$ is small; a long left tail when $p$ is near 1; more bell-shaped as $n$ grows (skewness $\frac{1-2p}{\sqrt{np(1-p)}}$ shrinks)
Assumptionsa fixed number $n$ of trials; each trial has two outcomes; the trials are independent; every trial has the same $p$
When usedconversions out of $n$ visitors, clicks out of $n$ impressions, defective items in a batch, correct answers out of $n$ questions
Relationshipssum of $n$ Bernoullis; Binomial($1, p$) = Bernoulli; ≈ Normal for large $np(1-p)$; ≈ Poisson($np$) for large $n$ and small $p$ (Chapter 4.8); Multinomial with $K = 2$; Beta prior for $p$ gives the Beta-Binomial model (Chapter 6.3)
Why do we need it?

We almost never look at one visitor; we look at totals. The Binomial tells us how much a total should wobble when the rate is fixed, which is the yardstick for deciding whether a difference between variants is real or just luck.

Where is it used?

Conversion counts in A/B tests (the likelihood of the Beta-Binomial model), two-proportion z-tests and their standard errors (Guide 2), logistic regression on grouped data, and quality control (defects per batch).

How is it used?

For each variant record $n$ (exposed users) and $k$ (converters). Then $k \sim$ Binomial($n$, $\theta$) is the likelihood: scipy.stats.binom(n, p).pmf(k), or in NumPyro dist.Binomial(total_count=n, probs=theta) with obs=k.

Y: pN: 1 − p visitor 1 → 2 → 3 YYY · k = 3 YYN · k = 2 YNY · k = 2 YNN · k = 1 NYY · k = 2 NYN · k = 1 NNY · k = 1 NNN · k = 0 3 paths with k = 2 each: p²(1 − p) P(X = 2) = 3 · p²(1 − p)
Three trials give $2^3 = 8$ paths. Three of them (orange) have exactly two successes, and each has probability $p^2(1-p)$. So $P(X = 2) = \binom{3}{2}p^2(1-p)$: count the paths, then multiply by the probability of one path.

Every small row is one possible sequence of $n$ visitors (blue square = converts, empty square = does not). The sequences are grouped by how many converted. Use the $k$ slider to pick a group: the purple box shows its $\binom{n}{k}$ sequences. The readout multiplies "number of sequences" by "probability of one sequence". Notice the groups in the middle are the tallest (most orders), but with a small $p$ the most probable count (lower plot) sits on the left.

Each ball meets $n$ rows of pegs and bounces right with probability $p$ (a success) or left otherwise. The bin it lands in = the number of right-bounces. Press Drop 1 ball to follow one path (purple), then Drop 100 balls several times: the blue bars (balls per bin) grow into the orange stems (the Binomial's expected counts). Try $p = 0.5$ (symmetric) and $p = 0.15$ (piled on the left, long right tail).

Move $n$ and $p$. The purple line is the mean $np$ and the shaded band is mean ± one standard deviation. Press the presets: Fair coin is symmetric; Rare event piles up near 0 with a long right tail; Almost sure is the mirror image. Then slide $n$ up with $p$ fixed: the shape becomes more symmetric and bell-like.

"$\binom{n}{k}p^k(1-p)^{n-k}$: the $\binom{n}{k}$ is a fudge factor."

It counts the orders. $p^k(1-p)^{n-k}$ is the probability of one particular order; there are $\binom{n}{k}$ orders with $k$ successes, and they add up.

"The most likely count is always exactly $np$."

$np$ is the mean and need not be a whole number ($np = 2.5$ is common). The most likely count (the mode) is a whole number near $np$, at most one step away.

"The Binomial variance is $p(1-p)$."

That is the variance of one trial. The count $X$ has variance $np(1-p)$; the proportion $X/n$ has variance $p(1-p)/n$. Mixing these up is the most common standard-error bug.

"Orders per day are Binomial, because each order is a success."

A Binomial needs a fixed, known number of trials $n$, so its count can never exceed $n$. Orders per day have no natural $n$ and no upper limit: that is a job for the Poisson or Negative Binomial (Chapter 4.8). "Converters out of 1 000 exposed users" is Binomial.

Model answer: "I use a Binomial when I have $n$ independent yes/no trials with a common probability: conversions out of exposed users. For counts without a fixed number of trials I use a Poisson, or a Negative Binomial if the variance is larger than the mean."

$X \sim \text{Binomial}(n, p)$: $P(X=k) = \binom{n}{k}p^k(1-p)^{n-k}$, $k = 0..n$.

$E[X] = np$, $Var(X) = np(1-p)$; the proportion $X/n$ has mean $p$ and variance $p(1-p)/n$.

$\binom{n}{k}$ = number of orders; $p^k(1-p)^{n-k}$ = probability of one order.

Trap: needs a fixed $n$; do not confuse the variance of one trial with the variance of the count.

Quick check: 10 visitors, $p = 0.3$. What are the mean, the variance and $P(X = 0)$?

Mean $= 10 \times 0.3 = 3$. Variance $= 10 \times 0.3 \times 0.7 = 2.1$. $P(X = 0) = 0.7^{10} \approx 0.0282$: about a 3% chance that nobody converts.

When is a count really Binomial? The four assumptions core

The Binomial formula is a promise: "if every trial is separate and has the same chance, the total will wobble by about $\sqrt{np(1-p)}$". In real data that promise is often broken, and when it breaks the total usually wobbles more than the Binomial says. Error bars built on the Binomial then look more certain than they should.

A handy way to remember the four conditions is BINS: Binary outcome, Independent trials, a fixed Number of trials, the Same probability for every trial.

Three ways to say it:

  • Picture: 1 000 separate coins versus 200 people who each flip one coin and then repeat the same answer 5 times. Same average, much bigger swings.
  • Numbers: in the example below, clustering multiplies the variance by 5 and a changing daily rate multiplies it by 6.5.
  • Slogan: the Binomial's spread is only honest when the trials are truly separate and truly alike.

Case 1: repeat users. A day has 1 000 sessions, but they come from 200 users with 5 sessions each. Take the extreme case: a user either converts in all 5 of their sessions or in none, with probability 0.1.

  1. If the 1 000 sessions were independent: $X \sim$ Binomial(1000, 0.1), mean $= 100$, variance $= 1000 \times 0.1 \times 0.9 = 90$, sd $\approx 9.5$.
  2. In reality the number of converting users is $U \sim$ Binomial(200, 0.1), and the session count is $X = 5U$.
  3. Mean: $E[X] = 5 \times 200 \times 0.1 = 100$. Same as before.
  4. Variance: $Var(5U) = 5^2 \, Var(U) = 25 \times (200 \times 0.1 \times 0.9) = 25 \times 18 = 450$, sd $\approx 21.2$. Five times the variance: the Binomial on sessions is badly overconfident.

Case 2: the rate changes from day to day. Each day has $n = 200$ visitors. On half the days $p = 0.05$, on the other half $p = 0.15$ (average 0.10). Using the law of total variance (Chapter 4.6):

  1. Within days (average Binomial variance): $\tfrac12(200 \times 0.05 \times 0.95) + \tfrac12(200 \times 0.15 \times 0.85) = \tfrac12(9.5 + 25.5) = 17.5$.
  2. Between days (variance of the daily mean $200p$): the daily mean is 10 or 30, each half the time, so its variance is $10^2 = 100$.
  3. Total: $17.5 + 100 = 117.5$, compared with $200 \times 0.1 \times 0.9 = 18$ for a fixed $p = 0.1$: about 6.5 times more.

A count is Binomial($n$, $p$) exactly when all four hold:

AssumptionPlain meaningTypical way it breaksWhat happens / what to use
Binaryeach trial has two outcomesa user picks one of 3 plansuse Categorical / Multinomial (below)
Independentone trial's result tells you nothing about another'srepeat sessions of the same user, users in the same store or city, word of mouthpositive dependence adds covariance terms: $Var(X) = np(1-p) + 2\sum_{i\lt j} Cov(X_i, X_j)$, so the variance grows
Number fixed$n$ is set in advance, not by the results"stop when we reach 10 conversions"; counts with no natural $n$ (orders per day)stopping rules change the distribution (counting failures until the $r$-th success gives the Negative Binomial, Chapter 4.8); unbounded counts → Poisson / NB
Same $p$every trial has the same success probabilitythe rate drifts across days, campaigns, batchesthe count is overdispersed (variance bigger than $np(1-p)$); a Beta-Binomial (the Binomial's $p$ itself varies from batch to batch, following a Beta curve) or a hierarchical model handles this (Chapter 6.3)

Overdispersion means "more spread than the model allows". It is the single most common way a textbook count model fails on real data, and it returns in full force in Chapter 4.8.

Why do we need it?

Every standard error, z-test and posterior built on a Binomial inherits its assumptions. If the trials are clustered or the rate drifts, the model is overconfident: intervals are too narrow and too many "significant" results appear.

Where is it used?

Choosing the unit of analysis in A/B tests (users, not sessions; Chapter 5.10), design effects in cluster sampling (Chapter 5.4), Beta-Binomial and hierarchical models for drifting rates, and checking a model with posterior predictive checks (Chapter 6.8).

How is it used?

Before using a Binomial, ask the BINS questions. Then check the data: compare the observed variance of daily or per-group counts with $np(1-p)$. A ratio far above 1 means overdispersion: change the unit, add structure (segments, days), or use a Beta-Binomial or hierarchical model.

independent: 10 separate trials clustered: 2 users × 5 sessions user 1: all yesuser 2: all no each circle is a fresh coin flip count wobbles like Binomial(10, p) only 2 real decisions, copied 5 times count = 5 × Binomial(2, p): jumps of 5 same number of sessions, same average, but the clustered count has 5 times the variance
Independence is about information. Ten separate sessions are ten pieces of evidence; ten sessions from two users who repeat their choice are really only two.

Each run simulates 300 days with 200 visits per day and an average conversion rate of 0.1. The blue histogram is the observed daily conversions; the orange stems are what Binomial(200, 0.1) predicts. In Textbook they match. Switch to p changes daily and raise the slider: the histogram spreads out far beyond the orange stems. Switch to Repeat users: with 5 sessions per user the counts jump in steps of 5 and the variance is about 5 times too big. Press New sample to see that this is not a fluke.

"If the average rate is right, the Binomial is right."

Both broken cases above have exactly the right mean (20 or 100). Only the spread is wrong, and the spread is what decides standard errors, intervals and posterior widths.

"Any difference between users makes the count overdispersed."

Careful. If each of the $n$ trials has its own fixed probability $p_i$ and they are independent, the variance is $\sum p_i(1-p_i)$, which is actually a little smaller than Binomial with the average $p$. The big inflation comes from dependence between trials or from a rate that changes between batches (days, campaigns) while being shared within a batch.

In an A/B framework like yours, the trials of the Beta-Binomial model must be the randomization units, usually users. If you fed sessions or page views into $k \sim \text{Binomial}(n, \theta)$, repeat visits by the same user would act like Case 1 above and the posterior for $\theta$ would be far too narrow, making $P(\theta_B \gt \theta_A \mid D)$ look more decisive than the data justify. If the conversion rate drifts strongly across days or segments, that structure belongs in the model (segments and hierarchical pooling, Chapter 6.5).

Binomial needs BINS: Binary, Independent, fixed Number, Same $p$.

Dependence or a drifting rate keeps the mean but inflates the variance (overdispersion).

Trap: right mean ≠ right model; check observed variance vs $np(1-p)$, and count users, not sessions.

Quick check: users visit 4 times each and always make the same decision on every visit. By how much does treating visits as independent understate the variance of the visit-level conversion count?

The visit count is $4U$ with $U$ Binomial over users, so its variance is $4^2 \times (\text{users} \times p(1-p))$, while the naive Binomial over visits gives $(4 \times \text{users}) \times p(1-p)$. The ratio is $16/4 = 4$: the naive variance is 4 times too small (the standard deviation 2 times too small).

The Normal approximation to the Binomial, and when it fails core

Look back at the shape explorer with $n = 100$, $p = 0.3$: the bars trace a smooth bell. That is no accident. A Binomial is a sum of many small independent pieces (the Bernoulli trials), and sums like that become bell-shaped. This is the Central Limit Theorem, which you will meet properly in Chapter 4.13.

The bell is useful: instead of adding hundreds of Binomial terms, we can use the Normal curve and a z-score. Most textbook formulas for proportions (z-tests, the simple "Wald" interval) quietly rely on this.

But the bell needs room. When $np$ is small (rare events), the bars are squeezed against zero and lean to the right. A symmetric bell then spills into negative counts, which cannot happen, and it gets the tail probabilities wrong.

Three ways to say it:

  • Picture: lots of expected successes → a bell; only a couple → a lopsided pile next to a wall at zero.
  • Numbers: for $n = 500$, $p = 0.1$ the Normal is off by about 0.003 in a tail; for $n = 100$, $p = 0.02$ it misses a tail probability by about a quarter.
  • Slogan: the bell needs about 10 expected successes and 10 expected failures.

Where it works. 500 visitors, $p = 0.1$. What is $P(X \le 40)$?

  1. Mean $np = 50$; variance $np(1-p) = 45$; sd $= \sqrt{45} \approx 6.708$.
  2. Exact (adding Binomial terms with software): $P(X \le 40) = 0.0751$.
  3. Normal with the continuity correction (the bar for 40 reaches up to 40.5): $z = (40.5 - 50)/6.708 \approx -1.416$, and $\Phi(-1.416) \approx 0.0784$. Close.
  4. Without the correction: $z = (40 - 50)/6.708 \approx -1.491$, $\Phi(-1.491) \approx 0.0680$. Worse, which is why the $+0.5$ is used.

Where it fails. 100 visitors, $p = 0.02$ (a rare event).

  1. Mean $np = 2$; sd $= \sqrt{100 \times 0.02 \times 0.98} = 1.4$.
  2. Exact: $P(X \ge 5) = 0.0508$.
  3. Normal: $P(X \ge 5) \approx P(Z \ge (4.5 - 2)/1.4) = P(Z \ge 1.786) \approx 0.0371$, about 27% too small.
  4. The same Normal also puts $\Phi((-0.5 - 2)/1.4) \approx 0.037$ of its probability below $-0.5$: on negative counts, which are impossible.

When $np(1-p)$ is large, the Binomial is close to a Normal with the same mean and variance:

$$X \approx N\big(np,\; np(1-p)\big), \qquad P(X \le k) \approx \Phi\!\left(\frac{k + 0.5 - np}{\sqrt{np(1-p)}}\right).$$
  • $\Phi$ is the standard Normal CDF (Chapter 4.4). The $+0.5$ is the continuity correction: a whole-number bar at $k$ covers the interval from $k - 0.5$ to $k + 0.5$.
  • For the proportion $\hat p = X/n$: $\hat p \approx N\big(p,\; p(1-p)/n\big)$. This is the basis of the two-proportion z-test and the Wald interval (Guide 2).
  • Rule of thumb (not a theorem): the approximation is usually good when $np \ge 10$ and $n(1-p) \ge 10$ (some books use 5). Below that, use the exact Binomial, or the Poisson when $n$ is large and $p$ small (Chapter 4.8).
Why do we need it?

Normal probabilities are quick to compute and easy to reason with (z-scores, "about 2 standard errors"). For large experiments they are accurate enough, and they explain where the classic formulas for proportions come from.

Where is it used?

The two-proportion z-test, Wald confidence intervals for a conversion rate, sample-size formulas for A/B tests (Guide 2), and quick sanity checks ("is 40 conversions surprising if I expect 50?").

How is it used?

Check $np \ge 10$ and $n(1-p) \ge 10$. If yes, standardize: $z = (k + 0.5 - np)/\sqrt{np(1-p)}$ and read $\Phi(z)$. If no, compute exactly with scipy.stats.binom.cdf: with a computer, there is rarely a reason not to.

Drag the purple handle to choose $k$. The dark blue bars add up to the exact $P(X \le k)$; the orange area under the bell up to $k + 0.5$ is the Normal approximation. Press Works ($n = 500$, $p = 0.1$): the two numbers agree well. Press Fails ($n = 100$, $p = 0.02$): the bars lean right, the bell spills past zero (red area = probability on impossible negative counts) and the tail numbers disagree. Then raise $n$ with the slider and watch the rule-of-thumb message change once $np$ passes 10.

"With $n = 1000$ the Normal approximation is always fine."

What matters is the expected number of successes $np$ (and failures), not $n$ alone. With $n = 1000$ and $p = 0.002$, $np = 2$: the Binomial is strongly skewed and the Poisson is the right approximation.

"The CLT says the conversions are Normally distributed."

Each conversion is 0 or 1, never Normal. The CLT is about the total (or the average) of many of them, and only approximately, for large $np(1-p)$.

For large $np(1-p)$: $X \approx N(np, np(1-p))$; $P(X\le k) \approx \Phi\big((k + 0.5 - np)/\sqrt{np(1-p)}\big)$.

Rule of thumb: $np \ge 10$ and $n(1-p) \ge 10$.

Trap: rare events (small $np$) are skewed and the bell puts mass on negative counts; compute exactly.

Quick check: 2 000 users, $p = 0.004$. Is the Normal approximation safe?

$np = 8$, which is below the rule-of-thumb 10, so expect some skew. The exact Binomial (or the Poisson with $\lambda = 8$) is the safer choice.

Categorical: one pick among $K$ options core

A new user signs up and picks one plan: Free, Basic or Pro. This is no longer yes/no; there are three possible answers. The Categorical distribution is the Bernoulli's big brother: one trial, $K$ possible outcomes, one probability for each.

Picture a spinner divided into slices of different sizes. The size of each slice is the probability of that category, and the slices fill the whole circle, so the probabilities add up to 1.

The categories are labels, not numbers. "Pro" is not three times "Free". That is why we usually store a categorical outcome as a one-hot vector: a list of $K$ zeros with a single 1 in the position of the chosen category.

Three ways to say it:

  • Picture: a spinner with unequal slices; one spin, one slice.
  • Numbers: with probabilities 0.5 / 0.3 / 0.2, out of 100 sign-ups about 50 pick Free, 30 Basic and 20 Pro.
  • Slogan: one pick among $K$ labels; the $K$ probabilities add to 1.

Plan choice with $\pi = (\pi_{\text{Free}}, \pi_{\text{Basic}}, \pi_{\text{Pro}}) = (0.5, 0.3, 0.2)$.

  1. Check: $0.5 + 0.3 + 0.2 = 1$. Only two numbers are free to choose; the third is whatever is left.
  2. $P(\text{Pro}) = 0.2$. $P(\text{not Free}) = 0.3 + 0.2 = 0.5$ (the categories cannot happen together, so we add).
  3. One-hot: a user who picks Basic is stored as $x = [0, 1, 0]$.
  4. Mean of the one-hot vector: the average of many such vectors is $[0.5, 0.3, 0.2] = \pi$. Each position is a Bernoulli indicator ("did they pick this plan?").
  5. Variance of the Pro indicator: $0.2 \times 0.8 = 0.16$. Covariance of the Free and Basic indicators: $E[x_F x_B] - E[x_F]E[x_B] = 0 - 0.5 \times 0.3 = -0.15$, negative because they can never both be 1.
  6. A trap: coding Free = 1, Basic = 2, Pro = 3 and taking the mean gives $1(0.5) + 2(0.3) + 3(0.2) = 1.7$. "Plan number 1.7" means nothing.

$X \sim \text{Categorical}(\pi)$ with $\pi = (\pi_1, \dots, \pi_K)$, $\pi_k \ge 0$, $\sum_k \pi_k = 1$, means $P(X = k) = \pi_k$ for $k = 1, \dots, K$. With the one-hot vector $x$:

$$p(x) = \prod_{k=1}^{K} \pi_k^{\,x_k}.$$

(Only the chosen category has $x_k = 1$, so the product is just its $\pi_k$.) It is also called the "multinoulli" distribution.

Recipe card · Categorical($\pi$)
Support$K$ labels $\{1, \dots, K\}$, or equivalently the $K$ one-hot vectors
Parametersthe probability vector $\pi$: $K$ numbers, but only $K - 1$ free (they must sum to 1)
Meannot meaningful for labels; for the one-hot vector, $E[x] = \pi$
Variancefor the one-hot vector: $Var(x_k) = \pi_k(1-\pi_k)$, $Cov(x_j, x_k) = -\pi_j\pi_k$ for $j \ne k$
Shapea bar chart; for nominal categories (no natural order) the order of the bars is arbitrary; the mode is the most probable category
Assumptionsone trial; exactly one of $K$ categories happens; the categories are mutually exclusive (no two at once) and exhaustive (one always happens)
When usedplan choice, device type, which button was clicked, the next word of a language model, the output of a multi-class classifier (softmax probabilities)
Relationships$K = 2$ gives the Bernoulli; the sum of $n$ independent one-hot draws is Multinomial; the Dirichlet is the usual prior for $\pi$ (Chapter 4.11); a softmax turns any $K$ scores into a valid $\pi$
Why do we need it?

Many outcomes are not yes/no: plan, channel, device, page, reason for cancelling. Squeezing them into several separate yes/no metrics loses the fact that they compete for the same user. The Categorical keeps them together and forces the probabilities to add up to 1.

Where is it used?

Multi-class classification (softmax + cross-entropy is the Categorical negative log-likelihood), language models (next token), mixture models (which component generated a point), and categorical A/B metrics.

How is it used?

Store the outcome as a label or one-hot vector; estimate $\pi_k$ by the fraction of users in category $k$. In NumPy: rng.choice(K, p=pi); in NumPyro: dist.Categorical(probs=pi) (or logits= scores).

Free 0.5 Basic 0.3 Pro 0.2 one spin = one user picks Free → one-hot [1, 0, 0] picks Basic → one-hot [0, 1, 0] picks Pro → one-hot [0, 0, 1] average of many one-hot vectors → [0.5, 0.3, 0.2] = π
A Categorical draw is one spin of a spinner whose slice sizes are the probabilities. Storing the result as a one-hot vector makes its average equal to the probability vector.

Drag the top of any orange bar. The other bars shrink or grow so that the total stays exactly 1: probabilities of exclusive categories compete. Press 1 user to see one draw and its one-hot vector, then 100 users a few times: the blue bars (observed fractions) approach the orange ones. Switch to 2 categories: a Categorical with $K = 2$ is just a Bernoulli.

"Code the plans 1, 2, 3 and use the average plan as a metric."

Nominal labels have no distances, so their average means nothing. Compare the probability vector (share of each plan), or a meaningful number attached to each plan, such as its revenue.

"Three separate yes/no metrics (picked Free? picked Basic? picked Pro?) carry the same information as one Categorical."

The three indicators are linked: they always sum to 1 and are negatively correlated. Treating them as unrelated metrics double-counts evidence and ignores that a gain in one plan is a loss in another.

$X \sim \text{Categorical}(\pi)$: $P(X = k) = \pi_k$, $\sum \pi_k = 1$ ($K - 1$ free numbers).

One-hot $x$: $E[x] = \pi$, $Var(x_k) = \pi_k(1-\pi_k)$, $Cov(x_j,x_k) = -\pi_j\pi_k$.

$K = 2$ → Bernoulli. Trap: labels are not numbers; never average the codes.

Quick check: a classifier outputs scores that softmax turns into (0.7, 0.2, 0.1). What distribution is that, and what is the variance of the indicator "class 1"?

A Categorical with $\pi = (0.7, 0.2, 0.1)$. The indicator for class 1 is Bernoulli(0.7), so its variance is $0.7 \times 0.3 = 0.21$.

Multinomial: counting how many picked each option core

Now ten users each pick a plan, independently, with the same probabilities. At the end we have a tally: maybe 5 Free, 3 Basic, 2 Pro. The Multinomial gives the probability of every possible tally. It is to the Categorical exactly what the Binomial is to the Bernoulli: $n$ trials, counted.

One new feature appears. The counts must add up to $n$. So they compete: if more people than usual pick Free, fewer people are left for Basic and Pro. The counts are therefore negatively correlated, even though every user chooses independently.

Three ways to say it:

  • Picture: drop $n$ balls into $K$ buckets of different widths and count the balls in each bucket.
  • Numbers: with $n = 10$ and $\pi = (0.5, 0.3, 0.2)$, the single most likely tally is $(5, 3, 2)$, and even it happens only 8.5% of the time.
  • Slogan: $n$ picks among $K$ options, counted; the counts add to $n$ and push against each other.

$n = 10$ users, $\pi = (0.5, 0.3, 0.2)$ for (Free, Basic, Pro). What is $P(\text{tally} = (5, 3, 2))$?

  1. One particular order with 5 Free, 3 Basic and 2 Pro: $0.5^5 \times 0.3^3 \times 0.2^2 = 0.03125 \times 0.027 \times 0.04 = 0.00003375$.
  2. Number of such orders (the multinomial coefficient): $\frac{10!}{5!\,3!\,2!} = \frac{3\,628\,800}{120 \times 6 \times 2} = 2520$.
  3. $P = 2520 \times 0.00003375 = 0.08505$.
  4. Expected counts: $n\pi = (5, 3, 2)$. Variances: $n\pi_k(1-\pi_k) = (2.5,\; 2.1,\; 1.6)$.
  5. Covariance of the Free and Basic counts: $-n\pi_F\pi_B = -10 \times 0.5 \times 0.3 = -1.5$. Correlation: $-1.5/\sqrt{2.5 \times 2.1} \approx -0.655$.
  6. Look at Pro alone: "Pro or not Pro" is yes/no, so the Pro count is Binomial(10, 0.2): mean 2, variance 1.6. The same numbers as in step 4.

Let $X = (X_1, \dots, X_K)$ count how many of $n$ independent Categorical($\pi$) trials fell in each category. Then $X \sim \text{Multinomial}(n, \pi)$ and

$$P(X_1 = x_1, \dots, X_K = x_K) = \frac{n!}{x_1!\,x_2!\cdots x_K!}\;\pi_1^{x_1}\pi_2^{x_2}\cdots\pi_K^{x_K}, \qquad x_k \ge 0,\;\; \sum_k x_k = n.$$
Recipe card · Multinomial($n$, $\pi$)
Supportall vectors of $K$ whole numbers $\ge 0$ that add up to $n$
Parameters$n$ (known number of trials) and $\pi$ ($K - 1$ free probabilities)
Mean$E[X_k] = n\pi_k$
Variance$Var(X_k) = n\pi_k(1-\pi_k)$; $Cov(X_j, X_k) = -n\pi_j\pi_k$ (always negative); $Corr(X_j,X_k) = -\sqrt{\frac{\pi_j\pi_k}{(1-\pi_j)(1-\pi_k)}}$, which does not depend on $n$
Shapea cloud of tallies centred at $n\pi$, squeezed onto the plane "counts add to $n$"; each single count has a Binomial shape
Assumptionsfixed $n$; independent trials; the same $\pi$ for every trial; exclusive and exhaustive categories (the BINS conditions, with "Binary" replaced by "$K$ categories")
When usedplan counts per A/B variant, survey answers, word counts in a document (bag of words), chi-square tests of category tables (Chapter 5.9)
Relationshipssum of $n$ one-hot Categoricals; $n = 1$ gives the Categorical; $K = 2$ gives the Binomial; each $X_k \sim$ Binomial($n, \pi_k$); merging categories adds their probabilities; independent Poisson counts, conditioned on their total, are Multinomial (Chapter 4.8); the Dirichlet prior gives the Dirichlet-Multinomial model (Chapter 6.3)
Why do we need it?

When an experiment's outcome is a category, the data per variant is a tally. The Multinomial is the one likelihood that treats the whole tally at once, including the fact that the counts must add to the number of users.

Where is it used?

The Dirichlet-Multinomial model for categorical A/B metrics, chi-square goodness-of-fit and independence tests, topic models and naive Bayes on word counts, and sample ratio checks across several variants.

How is it used?

Count users per category for each variant, giving a vector that sums to $n$. Use it as the observation: scipy.stats.multinomial(n, pi).pmf(x), or in NumPyro dist.Multinomial(total_count=n, probs=theta) with obs=counts, where theta has a Dirichlet prior.

Each dot on the left is one experiment with $n$ users: its Free count (across) against its Basic count (up). 400 experiments are drawn each time. The cloud slopes downward: tallies with many Free users have fewer Basic users, because all counts share the same $n$. The purple dashed line is "Free + Basic = $n$" (no Pro users at all); no dot can cross it. On the right, the Pro count alone (blue) matches a Binomial($n$, $\pi_{\text{Pro}}$) (orange). Change $n$: the cloud grows, but the correlation in the readout stays the same.

"The counts of a Multinomial are independent Binomials."

Each count on its own is Binomial($n, \pi_k$), but together they are tied by $\sum_k X_k = n$, so they are negatively correlated. Treating them as independent overstates the information in a tally.

"The correlation between the counts disappears with large $n$."

The covariance $-n\pi_j\pi_k$ grows with $n$ and the correlation $-\sqrt{\pi_j\pi_k/((1-\pi_j)(1-\pi_k))}$ does not change with $n$ at all.

"Categorical and Multinomial are the same thing."

A Categorical describes one trial: which category happened. A Multinomial describes the counts after $n$ trials. The Categorical is the Multinomial with $n = 1$, exactly as the Bernoulli is the Binomial with $n = 1$.

Model answer: "Per user, the plan choice is Categorical($\pi$). Per variant, I model the tally of plans as Multinomial($n$, $\pi$). With a Dirichlet prior on $\pi$, the posterior is again a Dirichlet, so the update is just prior pseudo-counts plus observed counts."

$X \sim \text{Multinomial}(n, \pi)$: $P(x) = \frac{n!}{\prod x_k!}\prod \pi_k^{x_k}$ with $\sum x_k = n$.

$E[X_k] = n\pi_k$, $Var(X_k) = n\pi_k(1-\pi_k)$, $Cov(X_j,X_k) = -n\pi_j\pi_k$. Each $X_k \sim$ Binomial($n, \pi_k$).

Trap: the counts are not independent; they share the total $n$.

Quick check: 200 users, $\pi = (0.6, 0.3, 0.1)$. What are the mean and variance of the third count, and the covariance of the first two?

Mean $200 \times 0.1 = 20$; variance $200 \times 0.1 \times 0.9 = 18$; covariance $-200 \times 0.6 \times 0.3 = -36$.

One family, and where it lives in your A/B framework core

The four distributions of this chapter are really one idea with two dials:

  • How many trials? One, or $n$ of them counted together.
  • How many possible outcomes per trial? Two (yes/no), or $K$ categories.

Each corner of that 2 × 2 table is one distribution. Once you see the table, every relationship ("a Bernoulli is a Binomial with $n = 1$", "a Multinomial with $K = 2$ is a Binomial") stops being a fact to memorise and becomes obvious.

Three ways to say it:

  • Picture: a 2 × 2 grid: (1 trial, 2 outcomes) = Bernoulli, ($n$ trials, 2 outcomes) = Binomial, (1 trial, $K$ outcomes) = Categorical, ($n$ trials, $K$ outcomes) = Multinomial.
  • Numbers: Bernoulli(0.3) = Binomial(1, 0.3) = Categorical((0.7, 0.3)).
  • Slogan: count the trials, count the outcomes, and you know the distribution.

An A/B test in this language. Variant A: 1 000 users, 100 convert. Variant B: 1 000 users, 120 convert. Each variant also records which plan each converter picked.

  1. Per user: converts or not, $y_i \sim$ Bernoulli($\theta_A$) for users in A. ($\theta$, "theta", is the usual letter for an unknown parameter.)
  2. Per variant: the total $k_A \sim$ Binomial(1000, $\theta_A$) and $k_B \sim$ Binomial(1000, $\theta_B$). Observed rates: 0.10 and 0.12.
  3. Plan choice per converter: Categorical($\pi_A$); the plan tally of variant A's 100 converters: Multinomial(100, $\pi_A$).
  4. Bernoulli per user or Binomial on the total? Take 4 users with outcomes 1, 0, 0, 1. The per-user model gives $\theta(1-\theta)(1-\theta)\theta = \theta^2(1-\theta)^2$. The Binomial on the total $k = 2$ gives $\binom{4}{2}\theta^2(1-\theta)^2 = 6\,\theta^2(1-\theta)^2$.
  5. At $\theta = 0.3$: $0.0441$ vs $0.2646$. At $\theta = 0.5$: $0.0625$ vs $0.375$. The ratio is always 6. A constant factor does not change which $\theta$ is more plausible, so both models lead to the same conclusions about $\theta$ (as long as the users really are independent with a shared $\theta$).
2 outcomes (yes/no)$K$ outcomes (categories)
1 trialBernoulli($p$): value 0 or 1Categorical($\pi$): one label / one-hot vector
$n$ trials, countedBinomial($n, p$): count $0..n$Multinomial($n, \pi$): tally vector summing to $n$

Relationships (each one is a line on the recipe cards):

  • Special cases: Bernoulli = Binomial($1, p$) = Categorical with $K = 2$; Categorical = Multinomial($1, \pi$); Binomial = Multinomial with $K = 2$.
  • Sums: the sum of $n$ independent Bernoulli($p$) is Binomial($n, p$); the sum of $n$ independent one-hot Categorical($\pi$) vectors is Multinomial($n, \pi$); Binomial($n_1, p$) + Binomial($n_2, p$) = Binomial($n_1 + n_2, p$) when independent with the same $p$.
  • Margins: each count of a Multinomial is Binomial; merging categories gives a smaller Multinomial.
  • Limits: Binomial → Normal when $np(1-p)$ is large; Binomial → Poisson when $n$ is large and $p$ small with $np = \lambda$ (Chapter 4.8).
  • Priors (conjugate pairs): Beta for $p$ (Beta-Binomial), Dirichlet for $\pi$ (Dirichlet-Multinomial) (Chapter 4.11, Chapter 6.3). "Conjugate" means the posterior has the same form as the prior.
Why do we need it?

Seeing the four as one family lets you move between the per-user view (Bernoulli, Categorical) and the per-variant view (Binomial, Multinomial) without changing the model, and tells you which prior goes with which likelihood.

Where is it used?

The conversion metric (Beta-Binomial) and the categorical metric (Dirichlet-Multinomial) of an A/B framework, logistic vs multinomial (softmax) regression, and Binomial vs Multinomial chi-square tests.

How is it used?

Decide the unit (user), the outcome type (yes/no or $K$ categories) and whether you work per user or with totals. Aggregating to totals is faster and gives the same posterior when users are independent with a shared rate.

Bernoulli(p)1 trial · 2 outcomes Binomial(n, p)n trials · 2 outcomes Categorical(π)1 trial · K outcomes Multinomial(n, π)n trials · K outcomes sum n sum n 2 → K 2 → K Normallarge np(1−p) Poisson (4.8)n large, p small Beta prior on p Dirichlet prior on π Read backwards: n = 1 turns a count into a single trial; K = 2 turns categories into yes/no.
The family map. Solid arrows: sums of trials and more categories. Dashed: the two limits of the Binomial, and the conjugate priors used in the A/B framework (purple).

Pick the number of trials and the number of outcomes; the widget names the distribution you get, draws one sample, and shows what you keep from it. Notice that with $n$ trials we throw away the order and keep only the count (Binomial) or the tally (Multinomial). Press Draw again several times and watch the kept number change.

Variant B is truly better (its true rate is higher by the lift you choose). The blue and orange curves show where the observed conversion rates of A and B land across many repeats of the same experiment. With 1 000 users per variant and a 10% lift they overlap a lot: the readout says how often B's observed rate is NOT above A's, even though B is better. Press Run one experiment to see single results jump around. Then raise the number of users and watch the curves pull apart.

"B converted more users in this experiment, so B has the higher true rate."

Observed counts are Binomial draws around the true rates. With 1 000 users per variant and a true 10% lift (0.10 vs 0.11), observed B fails to beat observed A in about 24% of experiments. Deciding needs a model of this noise: a test (Guide 2) or a posterior probability (Guide 3).

"Adding two Binomial counts always gives a Binomial."

Only if they are independent and share the same $p$. Adding the conversions of two variants with different rates gives something that is not Binomial (its variance is smaller than Binomial with the pooled rate), which is one reason we model each variant separately.

In an A/B framework like yours, these distributions are the likelihoods. Conversions: $k_v \sim \text{Binomial}(n_v, \theta_v)$ for each variant $v$, with a Beta prior on $\theta_v$: the Beta-Binomial model. Categorical metrics (such as which plan is chosen): the tally $x_v \sim \text{Multinomial}(n_v, \pi_v)$ with a Dirichlet prior on $\pi_v$: the Dirichlet-Multinomial model. Both update by adding observed counts to prior pseudo-counts (Chapter 6.3), and the decision quantities such as $P(\theta_B \gt \theta_A \mid D)$ are computed from the resulting posteriors (Chapter 6.4). Their validity rests on the assumptions of this chapter: independent randomization units and a shared rate within each variant (or segment).

"Modelling each user as a Bernoulli gives a different posterior from modelling the total as a Binomial."

If users are independent with a common $\theta$, the two likelihoods differ only by the constant $\binom{n}{k}$, so the posterior for $\theta$ is identical. The aggregated Binomial is simply the faster way to compute it.

Model answer: "The Bernoulli product is $\theta^k(1-\theta)^{n-k}$ and the Binomial is $\binom{n}{k}\theta^k(1-\theta)^{n-k}$. The binomial coefficient does not depend on $\theta$, so it cancels when the posterior is normalized. I aggregate to $(n, k)$ per variant unless I need user-level covariates or the users are not exchangeable."

2 × 2 table: (1, 2) Bernoulli · ($n$, 2) Binomial · (1, $K$) Categorical · ($n$, $K$) Multinomial.

Priors: Beta ↔ Binomial, Dirichlet ↔ Multinomial. Limits: Binomial → Normal (large $np(1-p)$), → Poisson (large $n$, small $p$).

Per-user Bernoulli and per-variant Binomial give the same posterior for $\theta$. Trap: observed winner ≠ true winner.

Quick check: a variant shows a page with 4 possible actions and you count what 500 users did. Which distribution, and which prior would you pair it with?

The tally of 500 actions over 4 categories is Multinomial(500, $\pi$). The conjugate prior for $\pi$ is a Dirichlet, giving the Dirichlet-Multinomial model.

Recap, cheat sheet and practice

  • Meet every distribution with the same eight questions: support, parameters, mean, variance, shape, assumptions, when used, relationships.
  • Bernoulli($p$): one yes/no trial; mean $p$, variance $p(1-p)$ (largest at $p = 0.5$).
  • Binomial($n, p$): number of yeses in $n$ independent trials with the same $p$; $\binom{n}{k}$ counts the orders; mean $np$, variance $np(1-p)$.
  • The Binomial needs BINS (Binary, Independent, fixed Number, Same $p$). Clustered trials or a drifting rate keep the mean but inflate the variance: overdispersion.
  • For large $np(1-p)$ the Binomial is close to a Normal (use the $+0.5$ continuity correction); for small $np$ it is skewed and the Normal fails.
  • Categorical($\pi$): one pick among $K$ labels; store it one-hot; never average the codes. Multinomial($n, \pi$): the tally of $n$ picks; each count is Binomial, and the counts are negatively correlated because they share $n$.
  • In the A/B framework: conversions are Beta-Binomial, categorical metrics are Dirichlet-Multinomial; per-user Bernoullis and per-variant totals give the same posterior when users are independent.

Cheat sheet

DistributionSupportParametersMeanVarianceTypical useSciPy / NumPyro
Bernoulli$\{0,1\}$$p$$p$$p(1-p)$one conversionbernoulli(p) / Bernoulli(probs=p)
Binomial$\{0..n\}$$n, p$$np$$np(1-p)$conversions out of $n$ usersbinom(n, p) / Binomial(total_count=n, probs=p)
Categorical$K$ labels$\pi$ (sums to 1)one-hot: $\pi$one-hot: $\pi_k(1-\pi_k)$, cov $-\pi_j\pi_k$which plan one user picksrng.choice(K, p=pi) / Categorical(probs=pi)
Multinomialtallies summing to $n$$n, \pi$$n\pi_k$$n\pi_k(1-\pi_k)$, cov $-n\pi_j\pi_k$plan tally per variantmultinomial(n, pi) / Multinomial(total_count=n, probs=pi)
RuleFormula / check
Binomial PMF$\binom{n}{k}p^k(1-p)^{n-k}$, $\binom{n}{k} = \frac{n!}{k!(n-k)!}$
Proportion $\hat p = X/n$mean $p$, variance $p(1-p)/n$
Normal approximation$P(X \le k) \approx \Phi\big((k + 0.5 - np)/\sqrt{np(1-p)}\big)$ if $np \ge 10$ and $n(1-p) \ge 10$ (rule of thumb)
Overdispersion checkobserved variance of counts ÷ $np(1-p)$ well above 1
Multinomial correlation$-\sqrt{\pi_j\pi_k/((1-\pi_j)(1-\pi_k))}$, independent of $n$
Code it · Python

import numpy as np
from scipy import stats

# --- Bernoulli: one visitor converts with p = 0.1
bern = stats.bernoulli(0.1)
print(bern.mean(), round(bern.var(), 4))            # 0.1 0.09

# --- Binomial: 5 visitors, p = 0.2
binom = stats.binom(n=5, p=0.2)
print(round(binom.pmf(2), 4))                       # 0.2048 = 10 * 0.2**2 * 0.8**3
print(binom.pmf(np.arange(6)).round(4).tolist())    # [0.3277, 0.4096, 0.2048, 0.0512, 0.0064, 0.0003]
print(binom.mean(), round(binom.var(), 4))          # 1.0 0.8

# --- Binomial = sum of Bernoullis (simulation)
rng = np.random.default_rng(0)
users = rng.random((100_000, 5)) < 0.2              # 100 000 experiments x 5 users (True = converts)
k = users.sum(axis=1)
print(k.mean().round(3), k.var().round(3))          # 1.001 0.797  (theory 1.0 and 0.8)

# --- Normal approximation: works for n = 500, p = 0.1 ...
big = stats.binom(500, 0.1)
print(big.cdf(40).round(4), stats.norm(50, np.sqrt(45)).cdf(40.5).round(4))   # 0.0751 0.0784
# ... and fails for a rare event (n = 100, p = 0.02, so np = 2)
small, approx = stats.binom(100, 0.02), stats.norm(2, np.sqrt(1.96))
print(small.sf(4).round(4), approx.sf(4.5).round(4))                           # 0.0508 0.0371
print(approx.cdf(-0.5).round(4))                    # 0.0371 of the Normal sits on impossible negative counts

# --- A broken assumption: repeat users (5 sessions each, same decision every time)
clustered = 5 * rng.binomial(200, 0.1, size=100_000)
print(clustered.mean().round(1), clustered.var().round(0), 1000 * 0.1 * 0.9)  # 100.0 451.0 90.0

# --- Categorical and Multinomial: plan choice (Free, Basic, Pro)
probs = np.array([0.5, 0.3, 0.2])
plan = rng.choice(3, p=probs)                       # one user: 0, 1 or 2
one_hot = np.eye(3)[plan]                           # e.g. [0. 1. 0.]
multi = stats.multinomial(n=10, p=probs)
print(multi.pmf([5, 3, 2]).round(5))                # 0.08505
X = rng.multinomial(10, probs, size=100_000)        # 100 000 tallies of 10 users
print(X.mean(axis=0).round(2))                      # [5.   2.99 2.01]  (theory 5, 3, 2)
print(np.cov(X.T)[0, 1].round(2))                   # -1.5  = -10 * 0.5 * 0.3

# --- Same evidence about theta: users as Bernoullis vs the total as a Binomial
y = np.array([1, 0, 0, 1])
for theta in [0.3, 0.5]:
    per_user = np.prod(stats.bernoulli(theta).pmf(y))
    total = stats.binom(4, theta).pmf(y.sum())
    print(theta, per_user.round(4), total.round(4), (total / per_user).round(4))
# 0.3 0.0441 0.2646 6.0
# 0.5 0.0625 0.375 6.0      -> the ratio is always C(4, 2) = 6

# --- The same distributions in NumPyro (float32, so tiny rounding differences)
import jax.numpy as jnp
import numpyro.distributions as dist
print(jnp.exp(dist.Binomial(total_count=5, probs=0.2).log_prob(2)))            # 0.2047999
print(jnp.exp(dist.Multinomial(total_count=10, probs=jnp.array(probs)).log_prob(jnp.array([5, 3, 2]))))  # 0.085050024
Test yourself

1. "Number of users, out of the 2 000 shown the new banner, who clicked it." Which distribution fits best?

There is a fixed number of yes/no trials ($n = 2000$ users), so the count of yeses is Binomial, assuming users act independently with the same click probability. A Bernoulli would describe a single user.

2. What is the variance of $X \sim$ Binomial(50, 0.2)?

$np(1-p) = 50 \times 0.2 \times 0.8 = 8$. (10 is the mean $np$; 0.16 is the variance of a single trial.)

3. For which success probability is a single yes/no outcome hardest to predict (largest variance)?

$p(1-p)$ is largest, 0.25, at $p = 0.5$. At $p = 1$ the outcome is certain and the variance is 0; 0.1 and 0.9 give the same variance 0.09.

4. In a Multinomial tally of plan choices, the counts for Free and Basic are…

$Cov(X_j, X_k) = -n\pi_j\pi_k \lt 0$. More Free users leaves fewer users for the other plans. The correlation does not depend on $n$.

5. $n = 300$ users, conversion probability 0.01. Is the Normal approximation to the Binomial safe?

What matters is the expected number of successes, $np = 3$. The distribution is squeezed against 0 and right-skewed; the bell would put probability on negative counts.

6. 1 000 sessions come from 250 users (4 sessions each), and each user makes the same decision in all of their sessions. If you treat the sessions as independent Binomial trials, the standard error of the conversion rate is…

The true variance of the session count is $4^2 \times 250\,p(1-p) = 4000\,p(1-p)$, while the naive Binomial gives $1000\,p(1-p)$: 4 times too small. The standard error is the square root, so it is 2 times too small.

Practice problems

A. Each of 10 visitors converts with probability 0.05, independently. What is the probability that at least one converts, and what is the expected number of conversions?

Use the complement: $P(\text{none}) = 0.95^{10} \approx 0.5987$, so $P(\text{at least one}) = 1 - 0.5987 = 0.4013$. The expected number is $np = 10 \times 0.05 = 0.5$.

B. Derive the mean and variance of the Binomial from the Bernoulli, and say which assumption each step needs.

Write $X = X_1 + \dots + X_n$ with $X_i \sim$ Bernoulli($p$). Linearity of expectation gives $E[X] = \sum E[X_i] = np$; this needs no independence, only the same $p$ for each trial. For the variance, $Var(X) = \sum Var(X_i) + 2\sum_{i\lt j}Cov(X_i,X_j)$; independence makes every covariance zero, leaving $np(1-p)$. Without independence the covariance terms appear, which is exactly the clustering problem.

C. Six users choose among (Free, Basic, Pro) with $\pi = (0.5, 0.25, 0.25)$. Find $P(\text{tally} = (3, 2, 1))$ and the expected tally.

Multinomial coefficient: $\frac{6!}{3!\,2!\,1!} = \frac{720}{6 \times 2 \times 1} = 60$. One order: $0.5^3 \times 0.25^2 \times 0.25 = 0.125 \times 0.0625 \times 0.25 = 0.001953$. So $P = 60 \times 0.001953 \approx 0.1172$. Expected tally $n\pi = (3, 1.5, 1.5)$.

D. A fair coin is flipped 400 times. Approximate $P(\text{at least } 220 \text{ heads})$ with the Normal, and say whether the approximation is trustworthy.

Mean $np = 200$, sd $\sqrt{400 \times 0.25} = 10$. With the continuity correction, $P(X \ge 220) \approx P(Z \ge (219.5 - 200)/10) = P(Z \ge 1.95) \approx 0.0256$. Here $np = n(1-p) = 200 \ge 10$ and $p = 0.5$ makes the Binomial symmetric, so the approximation is excellent (the exact value is $0.0255$).

E. Interview: "Your daily conversion counts vary much more than a Binomial predicts. Why might that be, and what would you do?"

Possible reasons: the rate changes from day to day (weekday effects, campaigns, traffic mix), so $p$ is not the same for every trial; or trials are not independent (repeat visits by the same user, sessions counted instead of users, clustered traffic from one source). Actions: check the unit of analysis (use users, the randomization unit); model the structure that changes the rate (day or segment effects, a hierarchical model); or use an overdispersed likelihood such as the Beta-Binomial. Then check the fit by comparing the observed variance of daily counts with what the model predicts.

F. Interview: "A colleague models the three plan shares of each variant as three independent Binomial metrics. What is wrong, and what would you use?"

The three shares are tied together: the counts add up to the number of users, and the probabilities add up to 1, so the counts are negatively correlated. Independent Binomials ignore this, can produce shares that do not add to 1, and treat one gain-and-loss as two separate pieces of evidence. The right likelihood is the Multinomial for the tally, with a Dirichlet prior on the share vector (the Dirichlet-Multinomial model). Each single share is still marginally Binomial, so simple per-plan summaries remain valid.

Chapter 4.8 · Syllabus Module 9

Count distributions: Poisson and Negative Binomial

Orders per day, support tickets per hour, clicks per session: counts with no fixed upper limit. The Poisson is the simplest model for them and has one striking property: its variance equals its mean. Real counts almost always vary more than that. The Negative Binomial fixes this, and it is the count likelihood of your forecasting model, so this chapter also untangles its many confusing parameterizations.

  • Know the Poisson: events in a window at a steady rate; $E[Y] = Var(Y) = \lambda$; its assumptions and its shape
  • See the Poisson as the limit of the Binomial (many tiny chances), and use sums and splits of Poisson counts
  • Recognise overdispersion (variance bigger than the mean) and explain where it comes from (rates that vary across days, users or stores)
  • Build the Negative Binomial as a gamma-Poisson mixture; know NB2: $Var = \mu + \mu^2/\alpha$, and that $\alpha \to \infty$ gives back the Poisson
  • Translate between all common NB parameterizations (SciPy, NumPy, NumPyro's three classes, statsmodels) and know which one your code uses
  • Tell zero inflation apart from overdispersion, and choose between Poisson, NB and zero-inflated models from the data

Poisson: counting events in a window core

What we need from earlier chapters: PMFs (Chapter 4.4), mean and variance (Chapter 4.5), and the eight-question recipe and the Binomial (Chapter 4.7).

A small online shop gets orders at random moments through the day, on average 2 per hour. How many orders will arrive between 3 pm and 4 pm? Maybe 0, maybe 2, now and then 5. Unlike the Binomial, there is no fixed number of trials: there is no list of "the $n$ people who might order this hour", and no upper limit on the count.

What we do have is a rate: the average number of events per unit of time (or per page, per square metre…). The window we count in (one hour) is called the interval or exposure. If events arrive one at a time, independently of each other, at a steady rate, the count in a window follows a Poisson distribution with one parameter, $\lambda$ ("lambda") = rate × window length = the expected count.

Three ways to say it:

  • Picture: raindrops falling on one paving stone during a minute of steady rain; you count the drops.
  • Numbers: with $\lambda = 2$ orders per hour, about 14% of hours have no orders, 27% have one, 27% have two, and 14% have four or more.
  • Slogan: independent events at a steady rate → Poisson, and its variance equals its mean.

Orders arrive at $\lambda = 2$ per hour. Let $Y$ be the number of orders in one hour. The Poisson formula is $P(Y = k) = \lambda^k e^{-\lambda}/k!$, with $e^{-2} \approx 0.1353$.

  1. $P(Y = 0) = 2^0 e^{-2}/0! = 1 \times 0.1353 / 1 = 0.1353$.
  2. $P(Y = 1) = 2^1 e^{-2}/1! = 2 \times 0.1353 = 0.2707$.
  3. $P(Y = 2) = 2^2 e^{-2}/2! = 4 \times 0.1353 / 2 = 0.2707$.
  4. $P(Y = 3) = 2^3 e^{-2}/3! = 8 \times 0.1353 / 6 = 0.1804$.
  5. $P(Y \ge 4) = 1 - (0.1353 + 0.2707 + 0.2707 + 0.1804) = 1 - 0.8571 = 0.1429$.
  6. Mean $= 2$ orders, variance $= 2$, standard deviation $= \sqrt 2 \approx 1.41$ orders.

For a whole 8-hour day the expected count is $8 \times 2 = 16$, so the daily count is Poisson(16): the parameter always means "rate × length of the window".

$Y \sim \text{Poisson}(\lambda)$, $\lambda \gt 0$, if

$$P(Y = k) = \frac{\lambda^{k} e^{-\lambda}}{k!}, \qquad k = 0, 1, 2, \dots$$

Why the mean and the variance are both $\lambda$. $E[Y] = \sum_k k\,\frac{\lambda^k e^{-\lambda}}{k!} = \lambda \sum_{k\ge1} \frac{\lambda^{k-1}e^{-\lambda}}{(k-1)!} = \lambda \times 1$ (the last sum adds up a whole Poisson PMF). The same trick gives $E[Y(Y-1)] = \lambda^2$, so $E[Y^2] = \lambda^2 + \lambda$ and $Var(Y) = \lambda^2 + \lambda - \lambda^2 = \lambda$.

Recipe card · Poisson($\lambda$)
Support$\{0, 1, 2, \dots\}$: all whole numbers, with no upper limit
Parameters$\lambda \gt 0$, the expected count = rate × window length (exposure)
Mean$\lambda$
Variance$\lambda$, the same as the mean (so the sd is $\sqrt\lambda$)
Shapefor small $\lambda$: piled near 0 with a long right tail ($P(0) = e^{-\lambda}$); for large $\lambda$: close to a bell $N(\lambda, \lambda)$; skewness $1/\sqrt\lambda$; the peak is at $\lfloor\lambda\rfloor$ (and also at $\lambda - 1$ when $\lambda$ is a whole number)
Assumptionsevents happen one at a time; counts in non-overlapping windows are independent; the rate is constant within the window
When usedorders per hour, tickets per day, page views per minute, defects per metre, goals in a match, count metrics per user in experiments
Relationshipslimit of Binomial($n, \lambda/n$); sums of independent Poissons are Poisson; thinning keeps it Poisson; the gaps between events are Exponential (Chapter 4.10); a Gamma-distributed $\lambda$ gives the Negative Binomial; Gamma is the conjugate prior (Chapter 6.3); Poisson regression uses a log link (Chapter 5.14)
Why do we need it?

Counts without a fixed number of trials are everywhere in business data. The Poisson is the baseline model for them: one parameter, easy to fit, and a clear prediction (variance = mean) that tells you immediately when the data need something richer.

Where is it used?

Poisson likelihoods for count metrics in experiments, Poisson regression (GLMs with a log link), queueing and staffing (calls per hour), anomaly alerts ("is 9 errors in a minute surprising if we expect 2?"), and as the starting point of the Negative Binomial.

How is it used?

Estimate $\lambda$ by the average count per window (it is the maximum-likelihood estimate). Check that the variance of the counts is close to the mean. In code: scipy.stats.poisson(lam), rng.poisson(lam), NumPyro dist.Poisson(rate=lam).

2 orders 0 orders 3 orders 1 order 1 pm2 pm3 pm4 pm5 pm gap = waiting time
A Poisson process: events (blue ticks) land at random moments at a steady rate. The count in each one-hour window is Poisson($\lambda$); the gaps between events follow an Exponential distribution (Chapter 4.10).

The top strip shows the first 12 hours of a simulated day-and-night of orders at rate $\lambda$ per hour (blue ticks), with each hour's count above it. Below, the counts of all 200 simulated hours (blue) are compared with the Poisson($\lambda$) PMF (purple). Press New sample a few times: the individual hours change a lot, but the histogram stays close to the purple stems and the readout's mean and variance stay close to each other. Try $\lambda = 0.5$ (mostly empty hours) and $\lambda = 12$ (a bell).

Slide $\lambda$ from 0.3 to 40. For small $\lambda$ the stems pile up at 0 with a long right tail; for large $\lambda$ they form a symmetric bell. The shaded band is mean ± one standard deviation $\sqrt\lambda$: notice that the band gets wider as $\lambda$ grows (spread grows with the mean). Switch on the Normal overlay to see when $N(\lambda, \lambda)$ is a fair stand-in.

"Variance = mean means the Poisson is a narrow distribution."

It means the spread is tied to the level: a Poisson with mean 100 has sd 10, a Poisson with mean 10 000 has sd 100. Bigger counts wobble more in absolute terms but less relative to their size ($\text{sd}/\text{mean} = 1/\sqrt\lambda$).

"$\lambda$ is a rate per hour, so a day's count is Poisson($\lambda$)."

The Poisson parameter is the expected count in the window you count over: rate × length. A rate of 2 per hour over an 8-hour day gives Poisson(16). In models this "length" is called the exposure or offset.

"The Poisson is the distribution of rare events."

The Poisson needs events that arrive independently, one at a time, at a steady rate. The count itself can be huge (visits per hour to a big site can be Poisson(5000)). "Rare" comes from one derivation, the limit of a Binomial with many trials that each rarely succeed (next section): each tiny slice of time rarely contains an event.

Model answer: "A Poisson count arises when independent events occur at a constant rate over a window. Its single parameter is the expected count, and its variance equals its mean, which is the first thing I check in real data."

$Y \sim \text{Poisson}(\lambda)$: $P(Y = k) = \lambda^k e^{-\lambda}/k!$, $k = 0, 1, \dots$

$E[Y] = Var(Y) = \lambda$ = rate × window; $P(0) = e^{-\lambda}$.

Assumptions: independent events, one at a time, constant rate. Trap: $\lambda$ includes the window length; "rare" is not the requirement.

Quick check: a site gets 0.5 errors per minute on average. What is the probability of a minute with no errors, and of a 10-minute stretch with no errors?

One minute: Poisson(0.5), $P(0) = e^{-0.5} \approx 0.607$. Ten minutes: Poisson($10 \times 0.5 = 5$), $P(0) = e^{-5} \approx 0.0067$.

Where the Poisson comes from: the limit of the Binomial core

Cut the hour into many tiny slots, say 3 600 seconds. In any one second, an order arrives with a tiny probability, and two orders in the very same second practically never happen. So each second is a yes/no trial, and the number of orders in the hour is the number of "yes" seconds: a Binomial with a huge $n$ and a tiny $p$.

Make the slots finer and finer while keeping the expected count $np = \lambda$ fixed, and the Binomial settles into a fixed shape. That limiting shape is the Poisson. This is why the Poisson has no $n$: it is what is left when $n$ is "infinitely large" and only the expected count $\lambda$ matters.

The same story fits users: 1 000 visitors, each buying with probability 0.002, give a count that is almost exactly Poisson(2).

Three ways to say it:

  • Picture: an hour chopped into thousands of slots, each one almost always empty.
  • Numbers: Binomial(1000, 0.002) gives $P(0) = 0.13506$; Poisson(2) gives $0.13534$.
  • Slogan: many chances, each tiny → Poisson with $\lambda = np$.

1 000 visitors, each buys with $p = 0.002$, so $\lambda = np = 2$.

  1. Binomial: $P(X = 0) = 0.998^{1000} \approx 0.13506$. Poisson: $e^{-2} \approx 0.13534$.
  2. $P(X = 1)$: Binomial $1000 \times 0.002 \times 0.998^{999} \approx 0.27067$; Poisson $2e^{-2} \approx 0.27067$.
  3. $P(X = 3)$: Binomial $0.18063$; Poisson $0.18045$.
  4. Variances: Binomial $np(1-p) = 2 \times 0.998 = 1.996$; Poisson $2$. When $p$ is tiny, $1 - p \approx 1$, which is why the variance ends up equal to the mean.
  5. With only $n = 10$ trials of $p = 0.2$ (same $\lambda = 2$), $P(0) = 0.8^{10} \approx 0.107$: far from 0.135. The limit needs many small chances.

Poisson limit theorem (the "law of rare events"). If $n \to \infty$ and $p \to 0$ with $np = \lambda$ held fixed, then for every $k$

$$\binom{n}{k} p^{k}(1-p)^{n-k} \;\longrightarrow\; \frac{\lambda^{k}e^{-\lambda}}{k!}.$$

Why (put $p = \lambda/n$ and regroup):

$$\binom{n}{k}\Big(\frac{\lambda}{n}\Big)^{k}\Big(1-\frac{\lambda}{n}\Big)^{n-k} = \frac{\lambda^k}{k!}\cdot\underbrace{\frac{n(n-1)\cdots(n-k+1)}{n^k}}_{\to\,1}\cdot\underbrace{\Big(1-\frac{\lambda}{n}\Big)^{n}}_{\to\,e^{-\lambda}}\cdot\underbrace{\Big(1-\frac{\lambda}{n}\Big)^{-k}}_{\to\,1}.$$

The middle limit $(1 - \lambda/n)^n \to e^{-\lambda}$ is the famous definition of $e$ from calculus (Calculus guide).

  • How close? A known bound (Le Cam's inequality) says that the largest difference between the two distributions' probabilities for any event (the "total variation distance") is less than $np^2 = \lambda p$. Small $p$ is what matters.
  • Rule of thumb: the Poisson is a good stand-in for the Binomial when $n \ge 20$ and $p \le 0.05$ (a convention, not a theorem).
Why do we need it?

It explains why the Poisson fits so many counts: whenever there are many independent opportunities, each with a small chance, the count is close to Poisson, whatever the exact number of opportunities.

Where is it used?

Modelling rare conversions or errors in large traffic, approximating Binomial tail probabilities for small rates, insurance claims, and justifying Poisson likelihoods for counts of users who perform a rare action.

How is it used?

If $n$ is large and $p$ small, replace Binomial($n, p$) by Poisson($np$). You do not even need to know $n$ exactly: only the expected count matters. Check that $p$ is small; with $p$ near 0.2 or above, keep the Binomial.

one hour, cut into n tiny slots 1 1 each slot: an event with tiny probability p = λ / n (at most one per slot) events in the hour = number of busy slots ~ Binomial(n, λ/n) → Poisson(λ) as the slots get finer
The Poisson as a Binomial with very many, very unlikely trials. Here 2 of the 20 slots are busy; with 3 600 one-second slots the count is practically Poisson.

Keep $\lambda$ fixed and slide the number of trials $n$ up from 10. The blue bars are Binomial($n$, $\lambda/n$); the purple stems are Poisson($\lambda$). With $n = 10$ they visibly disagree; by $n = 200$ you can hardly tell them apart, and the readout's distance shrinks roughly like $1/n$. Then try a larger $\lambda$: you need a larger $n$ for the same closeness, because $p = \lambda/n$ must be small.

"The Poisson approximation works whenever $n$ is large."

It needs $p$ small as well. Binomial(1000, 0.5) has mean 500 and variance 250, while Poisson(500) has variance 500: badly wrong. For large $n$ with moderate $p$, use the Normal approximation instead.

"The Binomial and the Poisson have the same variance."

A Binomial has variance $np(1-p)$, a little less than its mean $np$; the Poisson's variance equals its mean. In the limit $p \to 0$ the gap vanishes. Variance larger than the mean is a third situation (overdispersion, two sections below).

Binomial($n, \lambda/n$) → Poisson($\lambda$) as $n \to \infty$: many chances, each tiny.

Key step: $(1 - \lambda/n)^n \to e^{-\lambda}$. Error (total variation) $\lt np^2$.

Trap: needs small $p$, not just large $n$.

Quick check: 50 000 users each hit a rare bug with probability 0.0001. Approximately how likely is it that nobody hits it?

$\lambda = 50000 \times 0.0001 = 5$, so $P(0) \approx e^{-5} \approx 0.0067$. Here $p$ is tiny, so the Poisson is excellent.

Adding and splitting Poisson counts

Orders reach your shop through two channels: the app and the website. If each channel's orders arrive as an independent Poisson stream, then all orders together are again a Poisson stream, and the rates simply add.

It also works backwards. Visitors arrive as a Poisson stream; each visitor buys with probability 3%, independently. The buyers then form a Poisson stream of their own, with the rate multiplied by 3%. Throwing away events at random like this is called thinning.

Three ways to say it:

  • Picture: two rivers of raindrops merge into one river; a sieve that lets through 3% of the drops gives a thinner river.
  • Numbers: Poisson(2) + Poisson(3) = Poisson(5); 3% of Poisson(100) = Poisson(3).
  • Slogan: independent Poisson streams add; random sieving scales the rate.
  1. Adding. App orders: Poisson(2) per hour. Web orders: Poisson(3) per hour, independent. Total: Poisson(5). Mean 5, variance 5.
  2. Check with "no orders at all": the total is 0 only if both are 0, so $P(\text{total} = 0) = e^{-2}\times e^{-3} = 0.1353 \times 0.0498 = 0.0067 = e^{-5}$. ✓
  3. Scaling the window. A rate of 5 per hour over a 24-hour day is a sum of 24 independent hours: Poisson(120).
  4. Thinning. Visitors: Poisson(100) per hour. Each buys with probability 0.03. Buyers per hour: Poisson($100 \times 0.03 = 3$). $P(\text{no buyer in an hour}) = e^{-3} \approx 0.0498$.
  5. The non-buyers are Poisson(97), and, surprisingly, independent of the number of buyers.
  • Sums. If $Y_1 \sim \text{Poisson}(\lambda_1)$ and $Y_2 \sim \text{Poisson}(\lambda_2)$ are independent, then $Y_1 + Y_2 \sim \text{Poisson}(\lambda_1 + \lambda_2)$. The same holds for any number of independent Poisson counts.
  • Thinning. If $Y \sim \text{Poisson}(\lambda)$ and each event is kept independently with probability $q$, the kept count is Poisson($q\lambda$), the dropped count is Poisson($(1-q)\lambda$), and the two are independent.
  • Conditioning on the total. Given $Y_1 + Y_2 = n$, the split is Binomial: $Y_1 \mid (Y_1 + Y_2 = n) \sim \text{Binomial}\big(n, \tfrac{\lambda_1}{\lambda_1+\lambda_2}\big)$. With $K$ streams the split is Multinomial (Chapter 4.7).
Why do we need it?

Real counts are built from pieces: channels, hours, regions, user segments. These rules let us move between the pieces and the total without changing the model family, and tell us exactly how the rate scales.

Where is it used?

Aggregating hourly to daily counts, splitting total demand by channel, funnels (visits → buyers is thinning), exposure offsets in Poisson regression, and the Poisson trick behind the Multinomial likelihood.

How is it used?

Add the rates of independent streams; multiply a rate by the keep probability; multiply a per-hour rate by the window length. Before adding, ask whether the pieces are really independent: a shared driver (a promotion, the weather) breaks it.

Blue ticks are app orders, orange ticks are web orders (first 10 hours shown). In All orders mode the histogram of 300 hourly totals matches Poisson($\lambda_1 + \lambda_2$). Switch to Kept orders: each order is kept with probability $q$ (kept ones are drawn green, the rest faded). The kept counts follow Poisson($q(\lambda_1+\lambda_2)$): thinning only rescales the rate. In both modes, the readout's variance stays close to its mean.

"The average of two Poisson counts is Poisson."

$(Y_1 + Y_2)/2$ can be 1.5, which no count can be, and its variance is $\lambda/2$, not equal to its mean $\lambda$. Only sums of independent Poisson counts stay Poisson.

"If every day is Poisson, a month of days pooled together is Poisson."

Only if every day has the same rate. If quiet days have $\lambda = 10$ and busy days $\lambda = 30$, the pooled daily counts are a mixture, and a mixture of Poissons has variance larger than its mean. That is overdispersion, the subject of the next section.

Independent: Poisson($\lambda_1$) + Poisson($\lambda_2$) = Poisson($\lambda_1 + \lambda_2$). Window of length $t$: Poisson($\text{rate}\times t$).

Thinning with keep probability $q$: Poisson($q\lambda$). Given the total, the split is Binomial / Multinomial.

Trap: averages and mixtures of Poissons are not Poisson.

Quick check: 40 visitors per hour (Poisson), 5% of them buy. What is the distribution of buyers over a 10-hour day, and its standard deviation?

Buyers per hour: Poisson($40 \times 0.05 = 2$). Over 10 independent hours: Poisson(20). Standard deviation $\sqrt{20} \approx 4.47$.

Overdispersion: when the variance is bigger than the mean core

The Poisson makes a bold promise: variance = mean. Real counts rarely keep it. Your shop's daily orders are not driven by one steady rate: there are sale days and quiet Sundays, rainy weeks, a post that went viral, one business customer who orders 15 items at once. Each of these makes the counts spread out more than a single Poisson would.

"More spread than the model allows" has a name: overdispersion. The word dispersion just means spread. The usual measure is the dispersion index, variance ÷ mean: 1 for a Poisson, above 1 for overdispersed counts.

Three ways to say it:

  • Picture: pour the histograms of quiet days and busy days into one pile: the pile is much wider than a Poisson with the average rate.
  • Numbers: ten days with mean 20 orders and variance 60: three times what the Poisson allows.
  • Slogan: when the rate itself moves around, the counts spread more than Poisson.

Ten days of orders: 17, 24, 8, 30, 20, 15, 33, 12, 23, 18.

  1. Mean: $200/10 = 20$.
  2. Deviations from 20: $-3, 4, -12, 10, 0, -5, 13, -8, 3, -2$. Squares: $9, 16, 144, 100, 0, 25, 169, 64, 9, 4$, sum $540$.
  3. Sample variance: $540/9 = 60$. Dispersion index: $60/20 = 3$.
  4. If the days were Poisson(20), the sd would be $\sqrt{20} \approx 4.47$. The days with 8 and 33 orders would be $2.7$ and $2.9$ sd from the mean: $P(Y \le 8) = 0.0021$ and $P(Y \ge 33) = 0.0047$. Two such days in ten is very hard to believe.
  5. With sd $\sqrt{60} \approx 7.75$ (a model that allows variance 60), the same days are only $1.5$ and $1.7$ sd away: ordinary.

Where the extra variance comes from. Suppose half the days are quiet (rate 10) and half busy (rate 30). By the law of total variance (Chapter 4.6): $Var(Y) = E[Var(Y\mid\lambda)] + Var(E[Y\mid\lambda]) = E[\lambda] + Var(\lambda) = 20 + 10^2 = 120$. The mean is still 20, but the variance is 6 times the Poisson value.

A count $Y$ with mean $\mu$ is overdispersed relative to the Poisson if $Var(Y) \gt \mu$, and underdispersed if $Var(Y) \lt \mu$. The dispersion index is $D = Var(Y)/E[Y]$ (estimated by $s^2/\bar y$).

If each count is Poisson given its own rate $\lambda$, but $\lambda$ varies, then

$$E[Y] = E[\lambda], \qquad Var(Y) = \underbrace{E[\lambda]}_{\text{Poisson noise}} + \underbrace{Var(\lambda)}_{\text{rate variation}} \;\ge\; E[Y].$$

So any variation in the rate can only add variance. The common causes:

  • Heterogeneity (things differ): rates differ across days, users, stores, or because of factors the model does not include.
  • Clustering / contagion (events are not independent): one order brings another, items come in bulk, one user produces many events.
  • Excess zeros can also raise the variance, but they are a separate problem with a separate fix (later in this chapter).

Underdispersion ($D \lt 1$) is rarer: very regular arrivals (a machine that ticks every minute), or counts with a hard ceiling such as Binomial counts with a large $p$.

Why do we need it?

A Poisson model fitted to overdispersed data gets the mean right but the uncertainty wrong: its intervals are too narrow, its tests too eager, its forecasts overconfident. Spotting overdispersion tells you the likelihood must change.

Where is it used?

Choosing Poisson vs Negative Binomial likelihoods (demand forecasting, count metrics in experiments), quasi-Poisson and NB regression in GLMs, posterior predictive checks of variance (Chapter 6.8), and insurance and epidemiology models.

How is it used?

Compute $s^2/\bar y$ for groups that should share a rate (same weekday, same segment), or compare the residual variance with the fitted mean after modelling. Ratios well above 1 point to the Negative Binomial (or to missing structure in the model).

quiet days, λ = 10variance ≈ mean busy days, λ = 30variance ≈ mean all days pooledmean 20, variance 120 += each group alone is Poisson; the mixture is overdispersed (variance 6× the mean)
Overdispersion from heterogeneity: two groups that are each Poisson (variance ≈ mean) pool into counts whose variance is far above their mean.

Each blue dot is one of 12 stores: its average daily orders (across) and the variance of its daily orders (up), from 60 simulated days. The purple dashed line is the Poisson promise, variance = mean; the orange curve is the Negative Binomial's $\mu + \mu^2/\alpha$. With Poisson data the dots scatter around the line. With Negative Binomial data they bend upward along the orange curve, and the bend gets stronger as $\alpha$ gets smaller. This picture is a quick, powerful check on real data.

"My daily demand has variance 5 times its mean, so I need a Negative Binomial."

Maybe, but first ask whether the extra variance is explained by structure. A strong weekly cycle makes raw daily counts overdispersed even if every day is Poisson given its weekday. In a model with trend, seasonality and holidays, judge dispersion on what is left: compare each day's count with the model's mean for that day.

"Overdispersion only changes the error bars, so it is harmless for point forecasts."

The point forecast (the mean) may be fine, but intervals, tail probabilities ("P(demand > capacity)"), tests and posterior widths all come out too narrow. And in a model fitted by likelihood, a too-confident likelihood also lets a few noisy days pull the parameters around too much.

Dispersion index $D = Var/mean$: Poisson 1, overdispersed $\gt 1$.

Varying rate: $Var(Y) = E[\lambda] + Var(\lambda)$ (law of total variance) → always at least the mean.

Causes: heterogeneity, clustering. Trap: check dispersion after the model's structure, not on raw pooled counts.

Quick check: a store's daily orders have mean 50 and variance 50 within each weekday, but variance 300 over all days pooled. Is the store overdispersed?

Not in the sense that matters for a model with a weekday effect: given the weekday, variance ≈ mean, so a Poisson with weekday-specific rates fits. The pooled variance is large because the weekday means differ (between-group variance), which the model already explains.

Building the Negative Binomial: a Poisson whose rate varies (gamma-Poisson mixture) core

Here is a recipe that produces overdispersion on purpose. Each day, first nature picks how busy the day will be: a rate $\lambda_{\text{day}}$, drawn from a smooth distribution of positive numbers. Then orders arrive as Poisson($\lambda_{\text{day}}$). Repeat for many days and look at the counts.

The distribution used for the rate is the Gamma distribution: a flexible family for positive numbers with a shape parameter (here called $\alpha$) and a rate parameter (you will meet it properly in Chapter 4.10). A large $\alpha$ means the daily rates are all close to the average; a small $\alpha$ means wild swings between quiet and busy days.

Mixing many Poissons with Gamma-distributed rates gives a count distribution with a neat formula: the Negative Binomial. A mixture just means "first pick a parameter at random, then draw from the distribution with that parameter".

Three ways to say it:

  • Picture: two dice in a row: a "how busy is today" die, then a "how many orders" die whose odds depend on the first.
  • Numbers: average rate 4 with Gamma variance 8 gives counts with mean 4 and variance $4 + 8 = 12$.
  • Slogan: Negative Binomial = Poisson with a Gamma-distributed rate.

Target mean $\mu = 4$ orders and shape $\alpha = 2$. The rate is $\lambda \sim$ Gamma(shape $\alpha = 2$, rate $\beta = \alpha/\mu = 0.5$).

  1. Mean of the daily rate: $\alpha/\beta = 2/0.5 = 4 = \mu$. Variance of the rate: $\alpha/\beta^2 = 2/0.25 = 8 = \mu^2/\alpha$.
  2. Given the rate, the count is Poisson: $E[Y \mid \lambda] = \lambda$ and $Var(Y \mid \lambda) = \lambda$.
  3. Mean of the count: $E[Y] = E[\lambda] = 4$.
  4. Variance of the count (law of total variance): $E[\lambda] + Var(\lambda) = 4 + 8 = 12$. In general $\mu + \mu^2/\alpha = 4 + 16/2 = 12$. A Poisson(4) would have variance 4.
  5. Probability of a zero day: $P(Y = 0) = E[e^{-\lambda}] = \big(\tfrac{\beta}{\beta + 1}\big)^{\alpha} = \big(\tfrac{0.5}{1.5}\big)^2 = \tfrac{1}{9} \approx 0.111$. A Poisson(4) gives $e^{-4} \approx 0.018$: the mixture has six times more empty days, all from quiet days.

If $\lambda \sim \text{Gamma}(\text{shape} = \alpha,\ \text{rate} = \alpha/\mu)$ and $Y \mid \lambda \sim \text{Poisson}(\lambda)$, then $Y$ has the Negative Binomial distribution with mean $\mu$ and concentration $\alpha$:

$$P(Y = k) = \frac{\Gamma(k + \alpha)}{\Gamma(\alpha)\,k!}\left(\frac{\alpha}{\alpha + \mu}\right)^{\alpha}\left(\frac{\mu}{\alpha + \mu}\right)^{k}, \qquad k = 0, 1, 2, \dots$$ $$E[Y] = \mu, \qquad Var(Y) = \mu + \frac{\mu^2}{\alpha}.$$
  • $\Gamma(\cdot)$ is the gamma function, a smooth version of the factorial: $\Gamma(m) = (m-1)!$ for whole numbers $m$. It lets $\alpha$ be any positive number.
  • The first variance term $\mu$ is the Poisson noise; the second, $\mu^2/\alpha$, is the extra variance from the varying rate. It grows with the square of the mean.
  • Large $\alpha$: the Gamma is narrow, the rates barely vary, $\mu^2/\alpha \to 0$, and the NB becomes Poisson($\mu$). Small $\alpha$: strong rate variation, many zeros and a long right tail.
Why do we need it?

It turns the vague statement "the rate varies" into a concrete model with one extra parameter, and explains why the NB variance has the form $\mu + \mu^2/\alpha$. It also gives a way to simulate NB data: draw a Gamma, then a Poisson.

Where is it used?

NumPyro's GammaPoisson (and NegativeBinomial2, which is built on it), Bayesian Gamma-Poisson models for count metrics (Chapter 6.3), customer-purchase models, and the hierarchical view of overdispersion (each day has its own rate).

How is it used?

To simulate: lam = rng.gamma(shape=alpha, scale=mu/alpha) (NumPy uses the scale $= 1/\text{rate}$), then rng.poisson(lam). To model: use an NB likelihood with mean $\mu$ and a learned $\alpha$; the Gamma step is integrated out for you.

1 · pick today's rate λ ~ Gamma(α, α/μ) λ = 5.3 2 · draw the count Y | λ ~ Poisson(λ) Y = 6 over many days Y ~ NB(μ, α) Var = μ + μ²/α the Gamma step adds the variance μ²/α on top of the Poisson's μ
The gamma-Poisson mixture. Integrating out the hidden daily rate turns "a Poisson with a random rate" into the Negative Binomial.

Press Next day. Top: nature draws today's rate (purple dot) from the Gamma curve (teal). Bottom: the day's order count is drawn from Poisson(that rate) and added to the blue histogram. After a few days press Add 100 days several times. The blue bars approach the orange NB($\mu$, $\alpha$) stems, not the purple Poisson($\mu$) stems: more zeros, a longer tail. Raise $\alpha$ to 30: the Gamma narrows, every day has almost the same rate, and the NB collapses onto the Poisson.

"In NumPy, rng.gamma(alpha, alpha/mu) draws the Gamma rate with mean $\mu$."

NumPy's (and SciPy's) second Gamma argument is the scale $= 1/\text{rate}$. rng.gamma(alpha, alpha/mu) has mean $\alpha^2/\mu$. Use rng.gamma(shape=alpha, scale=mu/alpha). NumPyro's dist.Gamma(concentration, rate) uses the rate.

"The NB is just a Poisson with its variance multiplied by a constant."

That model is the quasi-Poisson (variance $= \phi\mu$). The NB2 variance is $\mu + \mu^2/\alpha$: the extra part grows with the square of the mean, so big counts are relatively much noisier than under a quasi-Poisson.

$\lambda \sim \text{Gamma}(\alpha, \text{rate } \alpha/\mu)$, $Y\mid\lambda \sim \text{Poisson}(\lambda)$ ⇒ $Y \sim \text{NB}(\mu, \alpha)$.

$E[Y] = \mu$, $Var(Y) = \mu + \mu^2/\alpha$ ($= E[\lambda] + Var(\lambda)$). $P(0) = (\alpha/(\alpha+\mu))^\alpha$.

$\alpha \to \infty$: Poisson. Trap: NumPy/SciPy Gamma takes the scale, not the rate.

Quick check: $\mu = 10$, $\alpha = 5$. What are the variance and the dispersion index of the NB?

$Var = 10 + 100/5 = 30$; dispersion index $30/10 = 3$.

The Negative Binomial: recipe card, two stories and its shape core

The Negative Binomial (NB) has two origin stories that lead to exactly the same distribution.

  • Story 1 (the modern one, from the last section): a Poisson count whose rate varies from unit to unit according to a Gamma distribution. This is the story to think with when you model demand.
  • Story 2 (the historical one, which gave the name): a salesperson keeps making calls until they get their $r$-th sale. Each call succeeds with probability $p$, independently. The NB counts the failed calls before the $r$-th success. (It is "negative binomial" because its formula uses a binomial coefficient with a negative upper number when written in a certain way; the name says nothing useful about the data.)

Matching the two: $r = \alpha$ and $p = \alpha/(\alpha + \mu)$. Story 2 is why SciPy describes the NB with $(n, p)$, while story 1 is why NumPyro's NegativeBinomial2 uses (mean, concentration).

Three ways to say it:

  • Picture: a Poisson whose rate wobbles, or the failures before the $r$-th success.
  • Numbers: NB with mean 4 and $\alpha = 2$ = failures before the 2nd success when each try succeeds with probability 1/3: mean 4, variance 12 either way.
  • Slogan: one distribution, two stories; think "Poisson with extra variance $\mu^2/\alpha$".

Story 2 with $r = 2$ sales needed and success probability $p = 1/3$ per call. Let $K$ = number of failed calls before the 2nd sale.

  1. $P(K = 0)$: the first two calls are both sales: $p^2 = 1/9 \approx 0.111$.
  2. $P(K = 1)$: exactly one failure among the calls before the last sale, and the last call is a sale. The failure can be in position 1 or 2 of the first two calls, so $\binom{2}{1}\,p^2(1-p) = 2 \times \tfrac19 \times \tfrac23 = \tfrac{4}{27} \approx 0.148$.
  3. In general $P(K = k) = \binom{k + r - 1}{k}p^r(1-p)^k$: choose where the $k$ failures go among the first $k + r - 1$ calls; the last call is the $r$-th sale.
  4. Mean $r(1-p)/p = 2 \times \tfrac23 / \tfrac13 = 4$. Variance $r(1-p)/p^2 = 2 \times \tfrac23 / \tfrac19 = 12$.
  5. Exactly the numbers of the gamma-Poisson example ($\mu = 4$, $\alpha = 2$: mean 4, variance 12, $P(0) = 1/9$). With $r = \alpha = 2$ and $p = \alpha/(\alpha + \mu) = 2/6 = 1/3$, the two stories give the same PMF.
  6. With $r = 1$ ("failures before the first success"), the NB is the Geometric distribution.

NB2, the mean–dispersion form used for modelling: $Y \sim \text{NB}(\mu, \alpha)$ with

$$P(Y = k) = \frac{\Gamma(k + \alpha)}{\Gamma(\alpha)\,k!}\left(\frac{\alpha}{\alpha + \mu}\right)^{\alpha}\left(\frac{\mu}{\alpha + \mu}\right)^{k}, \qquad E[Y] = \mu, \qquad Var(Y) = \mu + \frac{\mu^2}{\alpha}.$$

("NB2" because the variance has a $\mu^2$ term; an "NB1" variant with $Var = \mu(1 + \delta)$ also exists but is rarely used.)

Recipe card · Negative Binomial (NB2)
Support$\{0, 1, 2, \dots\}$: an unbounded count, like the Poisson
Parametersmean $\mu \gt 0$; concentration $\alpha \gt 0$ (also called shape, size, $r$, total_count, $\phi$ or $\theta$; some authors call $1/\alpha$ the "dispersion")
Mean$\mu$
Variance$\mu + \mu^2/\alpha$, always above the mean; the dispersion index is $1 + \mu/\alpha$
Shaperight-skewed with a longer right tail than the Poisson; more zeros than the Poisson with the same mean; when $\alpha \le 1$ the most likely value is 0; as $\alpha \to \infty$ it becomes Poisson($\mu$)
Assumptionscounts arising from Poisson events whose rate varies by a Gamma distribution across units or time (story 1), or failures before the $r$-th success in independent trials with constant $p$ (story 2)
When usedoverdispersed counts: daily demand, orders per customer, claims, website events per user, gene-expression read counts; NB regression; the count likelihood of a forecasting model
Relationshipsgamma-Poisson mixture; $\alpha \to \infty$ gives Poisson; $r = 1$ gives the Geometric; independent NBs with the same $p$ (same $\mu/\alpha$) add up to an NB; it is to the Poisson what the Beta-Binomial is to the Binomial (a mixture that adds variance); NB regression (Chapter 5.14); NB forecast likelihood (Chapter 7.13)
Why do we need it?

It is the simplest count distribution that lets the variance exceed the mean, with just one extra parameter, and it still contains the Poisson as a special case. When the data turn out to be Poisson, the fitted $\alpha$ simply becomes large.

Where is it used?

Demand forecasting with count data (the NB likelihood option of a Prophet-style model), NB regression in statsmodels and R (glm.nb), Stan's neg_binomial_2, NumPyro's NegativeBinomial2, and RNA-seq tools such as DESeq2.

How is it used?

Put the model's mean (trend, seasonality, regressors) into $\mu$ through a positive link such as exp, and learn $\alpha$ from the data. Report $\alpha$ together with the parameterization, and compare the fitted variance $\mu + \mu^2/\alpha$ with the observed spread.

story 2: count the failed calls before the r-th sale (r = 2) ✗✗sale 1 ✗✗✗sale 2 stop here. failures K = 5 K ~ NB(r = 2, p = P(sale)); with r = α and p = α / (α + μ) it is the same as NB2(μ, α)
The historical "failures before the $r$-th success" story. It explains SciPy's $(n, p)$ parameters, but for demand data the gamma-Poisson story is the useful way to think.

The orange stems are NB($\mu$, $\alpha$); the purple stems are Poisson($\mu$) with the same mean. Press Almost Poisson ($\alpha = 100$): the two nearly coincide. Press Typical demand ($\alpha = 5$) and Very overdispersed ($\alpha = 0.5$): the NB moves probability to 0 and to the far right tail, while the middle shrinks. Watch the readout's tail probability $P(Y \ge 2\mu)$: it is the number that matters for capacity planning.

"A larger α means more overdispersion."

In NB2 with concentration α, larger α means less overdispersion: the extra variance is $\mu^2/\alpha$, and $\alpha \to \infty$ is the Poisson. Libraries that use $1/\alpha$ (statsmodels' alpha) reverse the direction. Always check which way your parameter points.

"Adding two NB counts always gives an NB."

Only when they share the same $p = \alpha/(\alpha + \mu)$, i.e. the same ratio $\mu/\alpha$. Two NB counts with different ratios add up to something that is not exactly NB (unlike Poisson counts, which always add).

NB2: $E = \mu$, $Var = \mu + \mu^2/\alpha$; $\alpha \to \infty$ → Poisson; dispersion index $1 + \mu/\alpha$.

Same distribution as failures before the $r$-th success: $r = \alpha$, $p = \alpha/(\alpha+\mu)$; $r = 1$ → Geometric.

Trap: big α = less overdispersion (in NB2); know which parameter your library means.

Quick check: an NB with mean 20 has variance 100. What is α, and what would a Poisson with the same mean give for the sd?

$100 = 20 + 400/\alpha$, so $400/\alpha = 80$ and $\alpha = 5$. A Poisson(20) has sd $\sqrt{20} \approx 4.47$, against $\sqrt{100} = 10$ for this NB.

One distribution, many names: Negative Binomial parameterizations core

Twenty degrees Celsius and sixty-eight degrees Fahrenheit describe the same weather with different numbers. The Negative Binomial has the same problem, only worse: every library describes it with different knobs, sometimes with the same names meaning different things. SciPy's p and NumPyro's probs are complements of each other; statsmodels' alpha is the reciprocal of NumPyro's concentration.

The danger is that a wrong translation never crashes. The code runs, the model fits, and it silently describes a different distribution (often with the wrong mean). That is why your syllabus says: know the exact parameterization you used.

Three ways to say it:

  • Picture: one distribution in the middle, many name tags around it.
  • Numbers: mean 4 and variance 12 is nbinom(2, 1/3) in SciPy and NegativeBinomialProbs(2, 2/3) in NumPyro.
  • Slogan: translate through (mean, concentration), and check the mean afterwards.

Target: mean $\mu = 4$ and concentration $\alpha = 2$ (variance $4 + 16/2 = 12$, $P(0) = 1/9$). Each line below was checked to give exactly this distribution.

  1. SciPy stats.nbinom(n, p) counts failures before the $n$-th success: $n = \alpha = 2$, $p = \alpha/(\alpha + \mu) = 2/6 = 1/3$. Mean $n(1-p)/p = 2 \times \tfrac23 \times 3 = 4$. ✓
  2. NumPy rng.negative_binomial(n, p): same convention as SciPy, $n = 2$, $p = 1/3$.
  3. NumPyro NegativeBinomial2(mean=4, concentration=2): the mean–dispersion form, written directly.
  4. NumPyro GammaPoisson(concentration=2, rate=0.5): the gamma-Poisson story, rate $= \alpha/\mu = 0.5$.
  5. NumPyro NegativeBinomialProbs(total_count=2, probs=2/3): here probs $= \mu/(\alpha + \mu) = 4/6$, which is $1 -$ SciPy's $p$. Its mean is $\text{total\_count}\times\frac{\text{probs}}{1-\text{probs}} = 2 \times 2 = 4$. ✓
  6. NumPyro NegativeBinomialLogits(total_count=2, logits=log 2): logits $= \log\frac{\text{probs}}{1-\text{probs}} = \log(\mu/\alpha) = \log 2 \approx 0.693$.
  7. statsmodels NB2 (sm.families.NegativeBinomial(alpha=...) or the NegativeBinomial count model): its alpha $= 1/\alpha = 0.5$, with variance $\mu + \text{alpha}\cdot\mu^2 = 4 + 0.5 \times 16 = 12$. ✓

Translation table, starting from the mean $\mu$ and the concentration $\alpha$ (NB2: $Var = \mu + \mu^2/\alpha$):

WhereCall / formParameters in terms of $\mu, \alpha$Back to the mean
maths (NB2)$\text{NB}(\mu, \alpha)$$\mu$, $\alpha$$\mu$
classicNB($r$, $p$), failures before the $r$-th success$r = \alpha$, $p = \frac{\alpha}{\alpha+\mu}$$r(1-p)/p$
SciPystats.nbinom(n, p)$n = \alpha$, $p = \frac{\alpha}{\alpha+\mu}$$n(1-p)/p$
NumPyrng.negative_binomial(n, p)same as SciPy$n(1-p)/p$
NumPyroNegativeBinomial2(mean, concentration)mean $= \mu$, concentration $= \alpha$mean
NumPyroGammaPoisson(concentration, rate)concentration $= \alpha$, rate $= \alpha/\mu$concentration / rate
NumPyroNegativeBinomialProbs(total_count, probs)total_count $= \alpha$, probs $= \frac{\mu}{\alpha+\mu}$ ($= 1 -$ SciPy's $p$)total_count · probs / (1 − probs)
NumPyroNegativeBinomialLogits(total_count, logits)total_count $= \alpha$, logits $= \log(\mu/\alpha)$total_count · $e^{\text{logits}}$
statsmodelsNB2 alphaalpha $= 1/\alpha$(mean comes from the regression)
Stan / Rneg_binomial_2(mu, phi) / glm.nb's theta$\phi = \theta = \alpha$$\mu$

"Dispersion parameter" is ambiguous. Some sources mean $\alpha$ (bigger = closer to Poisson), others mean $1/\alpha$ (bigger = more overdispersed). Always say which, or better, say the variance formula.

Why do we need it?

Models get moved between libraries: a prototype in statsmodels, simulation in NumPy, the production model in NumPyro. A wrong translation changes the mean or the variance silently, and nothing in the training loop will warn you.

Where is it used?

Every NB likelihood in code: NumPyro forecasting and experiment models, SciPy simulations and checks, statsmodels NB regression, and any comparison of a fitted "dispersion" between tools or papers.

How is it used?

Keep $(\mu, \alpha)$ as the "master" values. Convert with the table, then check: compute the library's own mean and variance (d.mean, d.variance in NumPyro; d.mean(), d.var() in SciPy) and confirm they equal $\mu$ and $\mu + \mu^2/\alpha$.

one NB mean 4 · variance 12 NB2 / NegativeBinomial2mean 4, concentration 2 SciPy nbinom / NumPyn = 2, p = 1/3 GammaPoissonconcentration 2, rate 0.5 NegativeBinomialProbstotal_count 2, probs 2/3 NegativeBinomialLogitstotal_count 2, logits log 2 statsmodels NB2alpha = 1/2 (reciprocal!)
Six name tags, one distribution. Note the two traps: NumPyro's probs is $1 -$ SciPy's $p$, and statsmodels' alpha is $1/$concentration.

Set the mean $\mu$ and concentration $\alpha$ you want. The table gives the arguments for every library, and the last column recomputes each one's mean (for statsmodels, the variance) from its own formula: they all agree. Now choose a mistake. "SciPy p → NumPyro probs" passes SciPy's $p$ into NegativeBinomialProbs: the red stems show the distribution you would really get, with mean $\alpha^2/\mu$ instead of $\mu$. "statsmodels alpha → concentration" uses $1/\alpha$ as the concentration: the mean stays right but the variance is wrong.

"probs in NumPyro means the same as p in SciPy."

For the NB they are complements: NumPyro's probs $= \mu/(\alpha+\mu)$, SciPy's p $= \alpha/(\alpha+\mu)$. Passing one as the other changes the mean from $\mu$ to $\alpha^2/\mu$ (for $\mu = 4, \alpha = 2$: from 4 to 1), with no error message.

"I fitted α = 0.2 in statsmodels, so I can use concentration = 0.2 in NumPyro."

statsmodels' alpha is $1/\alpha$. The NumPyro concentration is $1/0.2 = 5$.

"I used a Negative Binomial likelihood with dispersion 2."

Name the parameterization: "an NB2 likelihood, parameterized by its mean $\mu_t$ and a concentration $\alpha$, so $Var(y_t) = \mu_t + \mu_t^2/\alpha$; in NumPyro that is NegativeBinomial2(mean, concentration)." Then a listener knows whether a larger value means more or less overdispersion.

Model answer: "Libraries disagree: SciPy uses (n, p) counting failures, NumPyro offers mean–concentration, total_count–probs and total_count–logits, and statsmodels' alpha is the reciprocal of the concentration. I keep the model in mean–concentration form and verify the implied mean and variance in code."

In your forecasting model the Negative Binomial is one of the likelihood options. With a mean–concentration class such as NegativeBinomial2(mean=mu_t, concentration=alpha), $\mu_t$ is built from trend, seasonality, holidays and regressors (kept positive by a link, Chapter 7.13) and $\alpha$ is learned: a large fitted $\alpha$ says the data are close to Poisson, a small one says the days are much noisier than Poisson. Check which NB class your code calls and how its arguments are computed; if it uses NegativeBinomialProbs or Logits, confirm the probability direction with a quick numeric test of the mean. Also note what your prior is placed on ($\alpha$ or $1/\alpha$): the same "weak" prior means different things on the two scales (Chapter 6.8).

Master form: mean $\mu$, concentration $\alpha$, $Var = \mu + \mu^2/\alpha$.

SciPy/NumPy: $n = \alpha$, $p = \alpha/(\alpha+\mu)$. NumPyro Probs: probs $= \mu/(\alpha+\mu)$; Logits: $\log(\mu/\alpha)$; GammaPoisson rate $= \alpha/\mu$. statsmodels alpha $= 1/\alpha$.

Trap: wrong translations run silently; always check the implied mean and variance.

Quick check: SciPy nbinom(n=5, p=0.25). What are $\mu$ and $\alpha$, and the NumPyro NegativeBinomial2 call?

$\alpha = n = 5$. Mean $\mu = n(1-p)/p = 5 \times 0.75/0.25 = 15$. So NegativeBinomial2(mean=15, concentration=5), with variance $15 + 225/5 = 60$.

Too many zeros: zero inflation is a different problem from overdispersion core

Count data often contain two kinds of zero. A shop that was open all day and simply got no orders is an ordinary zero: bad luck from the count distribution. A shop that was closed for renovation, a product that was out of stock, a user who never uses the feature: these zeros come from a separate switch that was off. They are called structural zeros.

A zero-inflated model writes this down: first a coin decides whether the unit is "switched off" (a structural zero, with probability $\pi$); if it is on, the count comes from an ordinary Poisson (or NB), which can also produce a zero now and then.

The Negative Binomial also has more zeros than the Poisson, but for a different reason: it spreads the rate out, so some units have tiny rates. Overdispersion and zero inflation can look alike in a summary (both raise the variance), yet they describe different mechanisms and give very different probabilities of zero.

Three ways to say it:

  • Picture: "the shop was closed" versus "the shop was open but nobody came".
  • Numbers: with the same mean 4 and the same variance 8, a zero-inflated Poisson has 20.5% zeros and the NB only 6.25%.
  • Slogan: too many zeros is a different illness from too much spread.

Zero-inflated Poisson with structural-zero probability $\pi = 0.2$ and Poisson rate $\lambda = 5$.

  1. $P(Y = 0) = \pi + (1 - \pi)e^{-\lambda} = 0.2 + 0.8 \times 0.00674 = 0.2054$: almost all of it from the switch.
  2. Mean: $(1 - \pi)\lambda = 0.8 \times 5 = 4$.
  3. $E[Y^2] = (1-\pi)(\lambda + \lambda^2) = 0.8 \times 30 = 24$, so $Var(Y) = 24 - 16 = 8$. (Formula: $(1-\pi)\lambda(1 + \pi\lambda) = 4 \times 2 = 8$.)
  4. An NB with the same mean 4 and variance 8 needs $8 = 4 + 16/\alpha$, so $\alpha = 4$. Its $P(0) = (4/8)^4 = 0.0625$.
  5. A Poisson with mean 4 (variance 4) gives $P(0) = e^{-4} = 0.0183$.
  6. Look at small counts too: $P(Y = 1)$ is 0.027 under the ZIP but 0.125 under the NB. The ZIP has a spike at 0, then a gap, then a Poisson hump around 5; the NB slopes smoothly from 0.

Zero-inflated Poisson, ZIP($\lambda$, $\pi$): with probability $\pi$, $Y = 0$ (structural); otherwise $Y \sim$ Poisson($\lambda$).

$$P(Y = 0) = \pi + (1-\pi)e^{-\lambda}, \qquad P(Y = k) = (1-\pi)\frac{\lambda^k e^{-\lambda}}{k!} \;\; (k \ge 1),$$ $$E[Y] = (1-\pi)\lambda, \qquad Var(Y) = (1-\pi)\lambda(1 + \pi\lambda).$$
  • The zero-inflated NB (ZINB) uses an NB instead of the Poisson, for data with both extra zeros and extra spread.
  • A hurdle model is a cousin: a yes/no part decides "zero or positive", and positive counts come from a count distribution with the zero removed (a "zero-truncated" Poisson or NB). In a hurdle model all zeros come from the yes/no part; in a zero-inflated model zeros come from both parts. More in Chapter 7.13.
  • In NumPyro: ZeroInflatedPoisson(gate=π, rate=λ) and ZeroInflatedNegativeBinomial2(mean, concentration, gate), where mean is the mean of the NB part (the overall mean is $(1 - \text{gate}) \times$ mean).
Why do we need it?

An NB can match the mean and variance of zero-heavy data and still put the zeros in the wrong place. Getting $P(0)$ right matters when the question is about zeros: "how many days will this product sell nothing?", "what share of users never convert?"

Where is it used?

Intermittent demand (spare parts, slow-moving products), stock-outs, feature usage per user, insurance claims, ecology counts (sites where a species is absent), and hurdle models for purchase counts.

How is it used?

Compare the observed share of zeros with the fitted model's $P(0)$. If zeros are far more common than even the NB predicts, and you can name the mechanism (closed, out of stock, inactive user), model it: ZIP / ZINB / hurdle, or better, add the cause (an "open" or "in stock" indicator) as data.

zero-inflated Poisson Negative Binomial hurdle model switch on? off: πon: 1 − π 0 (structural) Poisson(λ) can still give 0 rate λ ~ Gamma Poisson(λ) zeros only from small rates cross the hurdle? 0 (all zeros) count ≥ 1 zero-truncated same data can be summarised by the same mean and variance, yet the zeros arise in different ways
Three ways to get zeros. Zero inflation adds a separate "off" switch; the NB gets its zeros from units with low rates; a hurdle model sends every zero through its yes/no part.

Three distributions with the same mean $\mu$: Poisson (purple), NB (orange) and zero-inflated Poisson (teal). In Same mean and variance mode, the NB's $\alpha$ is chosen so its variance equals the ZIP's: two distributions with identical mean and variance. Compare the bars at 0, then at 1–3: the ZIP has a tall spike at zero and a gap after it. Raise the structural-zero probability $\pi$ and watch the gap widen. In Same mean only mode, set the NB's $\alpha$ yourself.

"Lots of zeros means I need a zero-inflated model."

A count with a small mean has lots of zeros under any model (Poisson(0.5) is 61% zeros), and an NB with small $\alpha$ has even more. "Zero-inflated" means more zeros than the count model predicts. Compare the observed share with the fitted $P(0)$ before adding a zero component.

"If the NB matches the mean and variance, it matches the data."

Two numbers do not pin down a distribution. Check the shape where your question lives: the share of zeros, the small counts, the upper tail.

"Zero inflation and overdispersion are the same thing: both mean more variance than Poisson."

Overdispersion is about the spread (variance above the mean), usually from rates that vary smoothly across units. Zero inflation is about an extra mechanism that produces zeros on top of an ordinary count. Zero inflation raises the variance too, so a variance check alone cannot tell them apart; the share of zeros and the shape near zero can.

Model answer: "I first compare variance with mean (NB vs Poisson), then compare the observed zero share with the model's $P(0)$. If there are still far too many zeros and I can name the cause, I add a zero-inflation or hurdle component, or model the cause directly."

ZIP: $P(0) = \pi + (1-\pi)e^{-\lambda}$; mean $(1-\pi)\lambda$; variance $(1-\pi)\lambda(1+\pi\lambda)$.

Hurdle: all zeros from a yes/no part; positives from a zero-truncated count.

Trap: overdispersion ≠ zero inflation; judge zeros against the fitted model's $P(0)$.

Quick check: a product's daily sales have mean 3. The fitted NB predicts 12% zero days, but 35% of days had zero sales. What would you look into?

Far more zeros than the NB predicts suggests a separate zero mechanism: stock-outs, the product not being listed, store closures. If you can record that cause (an "in stock" flag), add it to the model; otherwise consider a zero-inflated NB or a hurdle model.

Poisson, Negative Binomial or zero-inflated? Checking real counts core

You now have a small toolbox: Binomial (bounded counts), Poisson (variance = mean), Negative Binomial (variance above the mean) and zero-inflated versions (extra zeros). Choosing among them is a matter of asking the data three questions, in order: Is the count bounded? How does its variance compare with its mean? Are there more zeros than the model expects?

One subtlety decides most real cases: these questions must be asked after the model's structure is taken into account. Daily demand with a weekly pattern is "overdispersed" if you pool all days, but may be perfectly Poisson once each weekday has its own rate.

Three ways to say it:

  • Picture: a short flowchart: bounded? → variance vs mean? → zeros vs model?
  • Numbers: the ten days with mean 20 and variance 60 fit an NB ($\alpha \approx 10$) much better than a Poisson.
  • Slogan: support first, then spread, then zeros, always relative to the model's mean.

The ten days from the overdispersion section: 17, 24, 8, 30, 20, 15, 33, 12, 23, 18 (mean 20, variance 60).

  1. Support: orders per day, no fixed $n$ → the Poisson family, not the Binomial.
  2. Spread: dispersion index $60/20 = 3$, far above 1.
  3. Method-of-moments $\alpha$: solve $60 = 20 + 20^2/\alpha$, so $\alpha = 400/40 = 10$. (Maximum likelihood gives $\alpha \approx 11$.)
  4. Log-likelihood (how probable the data are under each fitted model): Poisson(20): $-37.68$; NB(20, 10): $-34.04$. The NB is higher by 3.6, more than the roughly 1 unit that one extra parameter is "worth" by AIC's rule, so the NB is preferred.
  5. Zeros: there are none, and both models predict almost none: no sign of zero inflation.
  6. Before trusting this, ask about structure: were the 8-order and 33-order days a Sunday and a sale day? If a model explains them, the leftover dispersion could be much smaller.

A practical procedure for choosing a count likelihood:

  1. Support. A fixed number of trials $n$ (bounded count) → Binomial (Beta-Binomial if overdispersed). No upper limit → Poisson family.
  2. Spread, given the model. Compare variance with mean within groups that share a rate (same weekday, segment, store), or compare residual spread with the fitted mean. Dispersion index ≈ 1 → Poisson; clearly above 1 → Negative Binomial.
  3. Zeros. Compare the observed share of zeros with the fitted $P(0)$. Far more zeros, with a nameable cause → model the cause, or use ZIP / ZINB / hurdle.
  4. Compare fits. Log-likelihood with a penalty for parameters (AIC), predictive checks of the variance, the zero share and the tails (Chapter 6.8), and holdout performance.

The NB contains the Poisson as a limit, so it never fits worse in-sample; the question is whether the improvement is real. Method-of-moments estimate: $\hat\alpha = \bar y^2/(s^2 - \bar y)$ when $s^2 \gt \bar y$; if $s^2 \le \bar y$ there is no evidence of overdispersion.

Why do we need it?

The likelihood decides how wide every interval, posterior and forecast band will be. A defensible, data-based choice (with a check you can show) is far stronger than "we used a Poisson because it is a count".

Where is it used?

Likelihood selection in both of your projects, GLM family choice (Poisson vs NB regression), demand-planning models, and as a standard interview question about count data.

How is it used?

Fit the simplest candidate, look at dispersion and zeros relative to its fitted means, switch to NB (or add zero inflation) only when the data demand it, and keep the check (a dispersion number, a PPC plot) as evidence for the choice.

fixed numberof trials n? yes Binomial(Beta-Binomial if spread) no variance ≈ mean(given the model)? yes Poisson bigger NegativeBinomial zeros ≫ P(0)?→ ZIP / ZINB /hurdle, or model the cause after any choice: check spread, zeros and tails of the fitted model against the data
Choosing a count likelihood: support first, then the spread relative to the model's mean, then the zeros. The zero check applies after either the Poisson or the NB.

Pick how the data were really generated, then judge them as if you did not know. The readout gives the dispersion index, the method-of-moments $\hat\alpha$, the log-likelihood of the fitted Poisson and NB, and the zero shares. Poisson data: dispersion ≈ 1 and the NB gains almost nothing. NB data: dispersion well above 1 and a clear log-likelihood gain. Zero-inflated: even the NB misses the zero share. Weekly cycle: pooled dispersion looks high, but within each weekday it is about 1, so the "overdispersion" is the weekly pattern. Press New sample and shrink the number of days to see how noisy these checks are with little data.

"The NB has a higher likelihood, so the data are overdispersed."

The NB contains the Poisson as a limit, so its maximized likelihood is never lower. Ask whether the gain is larger than the price of the extra parameter (AIC, a likelihood-ratio test, a holdout check), and whether it survives once the model's structure (weekdays, trend, holidays) is included.

"With 4 weeks of data I can reliably tell Poisson from NB."

Variance estimates are noisy with few points: for 28 days of truly Poisson counts, the dispersion index lands anywhere between about 0.54 and 1.6 in 95% of samples. With little data, prefer the NB with a sensible prior on $\alpha$ (it can still learn "close to Poisson") over a confident Poisson.

A/B framework. When a count metric (for example, orders per user during the test) gets a Poisson likelihood, the model assumes every user in a variant shares one rate. Heavy and light users break this: per-user counts are overdispersed, the posterior for each variant's rate is too narrow, and $P(\lambda_B \gt \lambda_A \mid D)$ looks more decisive than it should. Check the dispersion index per variant; if it is well above 1, consider an NB likelihood or modelling user heterogeneity.

Forecasting model. For count demand the NB likelihood is the natural choice: the support is right (whole numbers ≥ 0, unlike the Normal), and $\alpha$ lets the noise grow faster than the mean. Judge its fit on the counts relative to the fitted mean $\mu_t$, not on the raw series, and add a zero check if some days have stock-outs or closures (Chapter 7.13, Chapter 7.14).

Order of questions: bounded? → variance vs mean (given the model) → zeros vs fitted $P(0)$.

Method of moments (match the mean and variance): $\hat\alpha = \bar y^2/(s^2 - \bar y)$ if $s^2 \gt \bar y$. Compare fits with log-likelihood + parameter penalty and predictive checks.

Trap: pooled "overdispersion" can be unmodelled structure (weekdays); NB never fits worse in-sample.

Quick check: 60 days of tickets have mean 9 and variance 8.1. What do you conclude?

Dispersion index $8.1/9 = 0.9$: no sign of overdispersion (slightly below 1, well within noise for 60 days). A Poisson is reasonable; the method-of-moments $\hat\alpha$ does not exist because $s^2 \lt \bar y$.

Recap, cheat sheet and practice

  • Poisson($\lambda$): counts of independent events at a steady rate in a window; $\lambda$ = rate × window; $E = Var = \lambda$; $P(0) = e^{-\lambda}$.
  • It is the limit of Binomial($n, \lambda/n$): many chances, each tiny. Independent Poisson counts add (rates add); random thinning scales the rate.
  • Overdispersion: variance above the mean, usually because the rate varies across days, users or stores ($Var = E[\lambda] + Var(\lambda)$), or because events cluster. Judge it relative to the model's mean, not on pooled raw counts.
  • Negative Binomial = Poisson with a Gamma-distributed rate. NB2: mean $\mu$, $Var = \mu + \mu^2/\alpha$; large $\alpha$ → Poisson. Same distribution as "failures before the $r$-th success" with $r = \alpha$, $p = \alpha/(\alpha+\mu)$.
  • Parameterizations differ: SciPy/NumPy $(n, p)$; NumPyro NegativeBinomial2(mean, concentration), GammaPoisson, NegativeBinomialProbs (probs $= 1 -$ SciPy's $p$), NegativeBinomialLogits; statsmodels alpha $= 1/\alpha$. Always check the implied mean and variance.
  • Zero inflation (an extra "off" switch) is not overdispersion: same mean and variance can give very different $P(0)$. Compare observed zeros with the fitted model's $P(0)$.

Cheat sheet

DistributionSupportParametersMeanVariance$P(0)$Use when
Poisson$0, 1, 2, \dots$$\lambda$$\lambda$$\lambda$$e^{-\lambda}$variance ≈ mean (given the model)
Negative Binomial (NB2)$0, 1, 2, \dots$$\mu, \alpha$$\mu$$\mu + \mu^2/\alpha$$\big(\frac{\alpha}{\alpha+\mu}\big)^{\alpha}$variance > mean
Zero-inflated Poisson$0, 1, 2, \dots$$\lambda, \pi$$(1-\pi)\lambda$$(1-\pi)\lambda(1+\pi\lambda)$$\pi + (1-\pi)e^{-\lambda}$extra zeros from a separate mechanism
Binomial (for contrast)$0..n$$n, p$$np$$np(1-p)$$(1-p)^n$a fixed number of trials
NB in code, for mean $\mu$ and concentration $\alpha$Arguments
SciPy stats.nbinom(n, p) / NumPy negative_binomial(n, p)$n = \alpha$, $p = \alpha/(\alpha+\mu)$
NumPyro NegativeBinomial2(mean, concentration)$\mu$, $\alpha$
NumPyro GammaPoisson(concentration, rate)$\alpha$, $\alpha/\mu$
NumPyro NegativeBinomialProbs(total_count, probs)$\alpha$, $\mu/(\alpha+\mu)$
NumPyro NegativeBinomialLogits(total_count, logits)$\alpha$, $\log(\mu/\alpha)$
statsmodels NB2 alpha$1/\alpha$
Code it · Python

import numpy as np
from scipy import stats

# --- Poisson: orders per hour with rate 2
pois = stats.poisson(2)
print(pois.pmf([0, 1, 2, 3]).round(4).tolist())     # [0.1353, 0.2707, 0.2707, 0.1804]
print(pois.mean(), pois.var(), round(pois.sf(3), 4)) # 2.0 2.0 0.1429  (P(Y >= 4))

# --- Binomial with many tiny chances -> Poisson
print(round(stats.binom(1000, 0.002).pmf(0), 5), round(pois.pmf(0), 5))   # 0.13506 0.13534

# --- Sums and thinning (simulation)
rng = np.random.default_rng(1)
total = rng.poisson(2, 200_000) + rng.poisson(3, 200_000)                  # app + web
print(total.mean().round(2), total.var().round(2), (total == 0).mean().round(4))   # 5.0 5.01 0.0067  (Poisson(5): 5, 5, e**-5)
buyers = rng.binomial(rng.poisson(100, 200_000), 0.03)                     # keep each visitor with prob 0.03
print(buyers.mean().round(2), buyers.var().round(2))                       # 3.0 3.01 -> Poisson(3)

# --- Overdispersion in ten days of orders
days = np.array([17, 24, 8, 30, 20, 15, 33, 12, 23, 18])
m, v = days.mean(), days.var(ddof=1)
print(m, v, v / m)                                   # 20.0 60.0 3.0
alpha_mom = m**2 / (v - m)                           # method of moments: 10.0
print(stats.poisson(m).logpmf(days).sum().round(2),
      stats.nbinom(alpha_mom, alpha_mom / (alpha_mom + m)).logpmf(days).sum().round(2))   # -37.68 -34.04

# --- Gamma-Poisson mixture = NB2(mu=4, alpha=2)
mu, alpha = 4.0, 2.0
lam = rng.gamma(shape=alpha, scale=mu / alpha, size=400_000)    # NumPy uses SCALE = 1/rate
y = rng.poisson(lam)
print(y.mean().round(2), y.var().round(2), (y == 0).mean().round(3))   # 4.0 12.03 0.111  (theory 4, 12, 1/9)

# --- One NB, many names: all of these are NB2(mu=4, alpha=2)
import jax.numpy as jnp
import numpyro.distributions as dist
p_scipy = alpha / (alpha + mu)                      # 1/3
nb_scipy = stats.nbinom(n=alpha, p=p_scipy)
print(round(nb_scipy.mean(), 4), round(nb_scipy.var(), 4), round(nb_scipy.pmf(0), 4))   # 4.0 12.0 0.1111
for d in [dist.NegativeBinomial2(mean=mu, concentration=alpha),
          dist.GammaPoisson(concentration=alpha, rate=alpha / mu),
          dist.NegativeBinomialProbs(total_count=alpha, probs=mu / (alpha + mu)),
          dist.NegativeBinomialLogits(total_count=alpha, logits=np.log(mu / alpha))]:
    print(type(d).__name__, float(d.mean), float(d.variance), round(float(jnp.exp(d.log_prob(0))), 4))
# NegativeBinomial2 4.0 12.0 0.1111   (and the same line for the other three)

# --- The classic mistake: SciPy's p passed as NumPyro's probs
wrong = dist.NegativeBinomialProbs(total_count=alpha, probs=p_scipy)
print(float(wrong.mean))                             # 1.0, not 4.0  (= alpha**2 / mu)

# --- statsmodels' alpha is 1 / concentration
import statsmodels.api as sm
print(sm.families.NegativeBinomial(alpha=1 / alpha).variance(np.array([mu])))   # [12.]

# --- Zero inflation vs overdispersion: same mean 4 and variance 8
zip_ = dist.ZeroInflatedPoisson(gate=0.2, rate=5.0)
print(round(float(jnp.exp(zip_.log_prob(0))), 4), float(zip_.mean), float(zip_.variance))   # 0.2054 4.0 8.0
print(round(stats.nbinom(4, 0.5).pmf(0), 4), stats.nbinom(4, 0.5).mean(), stats.nbinom(4, 0.5).var())  # 0.0625 4.0 8.0
Test yourself

1. Support tickets arrive as Poisson with an average of 9 per day. What are the variance and the standard deviation of a day's count?

For a Poisson, variance = mean = 9, so the sd is $\sqrt 9 = 3$. No "number of trials" is needed.

2. Which data most clearly call for a Negative Binomial rather than a Poisson?

Variance five times the mean, on days that should share a rate, is strong overdispersion. Variance ≈ mean fits the Poisson; a fixed number of trials is Binomial; 0/1 is Bernoulli.

3. NumPyro NegativeBinomial2(mean=10, concentration=5) has variance…

$\mu + \mu^2/\alpha = 10 + 100/5 = 30$.

4. SciPy stats.nbinom(n=3, p=0.5) is the same distribution as which NumPyro call?

Concentration $\alpha = n = 3$; mean $n(1-p)/p = 3 \times 0.5/0.5 = 3$. (In NegativeBinomialProbs it would be total_count=3, probs=0.5, since probs $= 1 - p$.)

5. Your fitted NB matches the mean and variance of daily sales, but 30% of days have zero sales where the NB predicts 8%. The most likely explanation is…

Matching mean and variance does not pin down $P(0)$. Far more zeros than the fitted NB predicts points to an extra zero-producing process; model its cause, or use a zero-inflated or hurdle model.

6. Errors occur as Poisson with rate 1.5 per hour. What is the probability of a 4-hour shift with no errors?

Four independent hours add up: Poisson($4 \times 1.5 = 6$), so $P(0) = e^{-6}$. Equivalently, all four hours must be empty: $(e^{-1.5})^4 = e^{-6}$.

Practice problems

A. Tickets arrive at 3 per hour (Poisson). What is the probability of at least 2 tickets in an hour?

$P(0) = e^{-3} \approx 0.0498$ and $P(1) = 3e^{-3} \approx 0.1494$. So $P(Y \ge 2) = 1 - 0.0498 - 0.1494 = 1 - 4e^{-3} \approx 0.8009$.

B. Derive the mean and variance of the gamma-Poisson mixture with $\lambda \sim$ Gamma(shape $\alpha$, rate $\alpha/\mu$).

The Gamma has mean $\alpha/(\alpha/\mu) = \mu$ and variance $\alpha/(\alpha/\mu)^2 = \mu^2/\alpha$. Given $\lambda$, $Y$ is Poisson, so $E[Y\mid\lambda] = Var(Y\mid\lambda) = \lambda$. Law of total expectation: $E[Y] = E[\lambda] = \mu$. Law of total variance: $Var(Y) = E[Var(Y\mid\lambda)] + Var(E[Y\mid\lambda]) = E[\lambda] + Var(\lambda) = \mu + \mu^2/\alpha$.

C. A statsmodels NB2 fit reports mean 8 and alpha = 0.25. Write the NumPyro NegativeBinomial2 and SciPy nbinom versions, and the variance.

statsmodels' alpha is $1/\alpha$, so the concentration is $\alpha = 4$. NumPyro: NegativeBinomial2(mean=8, concentration=4). SciPy: nbinom(n=4, p=4/(4+8)=1/3). Variance $8 + 64/4 = 24$ (statsmodels: $8 + 0.25 \times 64 = 24$ ✓).

D. A zero-inflated Poisson has $\lambda = 10$ and $\pi = 0.3$. Find $P(0)$, the mean and the variance, and compare $P(0)$ with a Poisson of the same mean.

$P(0) = 0.3 + 0.7e^{-10} \approx 0.3000$. Mean $(1 - 0.3) \times 10 = 7$. Variance $7 \times (1 + 0.3 \times 10) = 28$. A Poisson(7) has $P(0) = e^{-7} \approx 0.0009$: the ZIP has over 300 times more zeros, and four times the variance.

E. Interview: "Why might a Poisson likelihood be too restrictive for demand forecasting, and what does the Negative Binomial add?"

A Poisson forces the variance to equal the mean at every time point. Real demand varies more because of things the model does not capture (local events, weather, bulk orders, customers who differ), so a Poisson model's prediction intervals would be too narrow and it would over-react to noisy days. The NB is a Poisson whose rate varies by a Gamma distribution; it adds one parameter, the concentration $\alpha$, giving $Var = \mu + \mu^2/\alpha$, so the noise can grow faster than the mean. It still has the right support (whole numbers ≥ 0) and becomes the Poisson when $\alpha$ is large, so the data can tell you how much extra variance there is.

F. Interview: "In a daily-sales dataset, how would you tell overdispersion apart from zero inflation?"

First fit the mean structure (trend, weekday, holidays), so that "dispersion" is measured around the right mean. Then (1) compare the variance of the counts with their fitted means: a ratio well above 1 is overdispersion, handled by an NB. (2) Compare the observed share of zero days with the zero probability of the fitted NB: if there are far more zeros than even the NB predicts, there is an extra zero mechanism. (3) Look for the cause of those zeros (stock-outs, closures, items not listed); if it can be recorded, use it as data, otherwise use a zero-inflated NB or a hurdle model. Finally, check the chosen model with predictive checks of the zero share, the variance and the upper tail.

Chapter 4.9 · Syllabus Module 9

Noise distributions: Normal, Student-t and Laplace

No model predicts perfectly. The part it misses is called noise, and the shape of that noise decides how your model reacts to surprises. This chapter teaches the three noise shapes your projects use: the Normal (the classic bell), the Student-t (a bell that expects the occasional big surprise) and the Laplace (a sharp tent that also makes a famous prior).

  • Describe any distribution with the eight-step recipe: support → parameters → mean → variance → shape → assumptions → when used → relationships
  • Use the Normal: its formula, the 68–95–99.7 rule, z-scores, and why sums of independent Normals stay Normal (variances add, not standard deviations)
  • Understand the Student-t: degrees of freedom ν, heavy tails, why it becomes Normal as ν grows, when its mean or variance does not exist, and why its scale σ is not its standard deviation
  • Understand the Laplace: $p(x) \propto e^{-|x-\mu|/b}$, its sharp peak and exponential tails, variance $2b^2$, and why a Laplace prior pulls values toward zero
  • Compare tails with numbers ($P(|X| \gt 3)$) and with the log-density picture, and explain why the mean, the median and the Student-t centre react so differently to one outlier

Colours in every plot of this chapter: blue = Normal, orange = Student-t, teal = Laplace, purple = a special point or line, red = residuals and tail areas. Labels always say the same thing in words.

What is a noise distribution? Location, scale and the recipe core

Your forecast says 200 orders tomorrow. Tomorrow comes and 212 orders arrive. The miss, $+12$, is the noise of that day. (People also call it the error or the residual: the actual value minus the predicted value.)

Every day misses by a different amount. You cannot predict each miss. But you can describe the pattern of the misses: "usually within about 10 orders, as often above as below, a miss of 40 almost never happens". A noise distribution is exactly that description, written as a probability distribution.

All three noise distributions in this chapter live on the whole number line (a miss can be any positive or negative number) and are symmetric around a centre. They differ in one important way: how often very big misses happen. That far-away part of a distribution is called its tail.

Three ways to say it:

  • Picture: a bell of possible misses sitting on top of the prediction line.
  • Numbers: "misses of ±10 are common, ±40 almost never" is a noise distribution in words.
  • Slogan: the model predicts the centre; the noise distribution predicts the size and shape of the surprises.

Five days of forecasts and actual orders:

Day12345
Forecast200210190205195
Actual212204193230191
Residual = actual − forecast+12−6+3+25−4

A noise model says how such residuals behave. Most noise models are built from one standard shape $Z$ (centred at 0, width 1), then shifted and stretched:

  1. Draw a standard value, say $z = 1.2$.
  2. Stretch it by the scale $s = 10$: $10 \times 1.2 = 12$.
  3. Shift it by the location $\mu = 200$ (the forecast): $200 + 12 = 212$. That is day 1's actual value.
  4. The density must stretch too. The standard Normal's peak height is $0.399$. Stretched 10 times wider, the same total area 1 must spread over 10 times more room, so the peak drops to $0.399/10 = 0.0399$.

A location–scale family is a set of distributions made from one standard shape $Z$ by

$$X = \mu + s\,Z, \qquad p_X(x) = \frac{1}{s}\, p_Z\!\left(\frac{x-\mu}{s}\right).$$
  • $\mu$, the location, says where the centre sits. $s \gt 0$, the scale, says how stretched the shape is. Both are parameters: fixed numbers (knobs) that pick one member of the family.
  • The factor $1/s$ keeps the total area equal to 1 (a density is a height whose area is probability; see Chapter 4.4).
  • Normal: $s$ is called $\sigma$ (and equals the standard deviation). Laplace: $s$ is called $b$. Student-t: $s$ is called $\sigma$ too, and there is one more parameter, $\nu$, that changes the shape itself.
  • The support of a distribution is the set of values that can occur at all. For all three noise families it is every real number, written $(-\infty, \infty)$ or $\mathbb{R}$.
  • The tails are the far-left and far-right parts of the distribution: how much probability sits far from the centre.

The syllabus asks you to learn every distribution with the same recipe: support → parameters → mean → variance → shape → assumptions → when used → relationships. Each distribution below ends with a summary card in exactly that order.

Why do we need it?

A model that only predicts a single number cannot say how surprised to be, cannot draw prediction intervals and cannot be fitted by likelihood. The noise distribution is the part of the model that turns "prediction" into "probability of what we saw".

Where is it used?

The likelihood of every regression and forecasting model (your Prophet-style model's $\epsilon_t$), continuous metrics in an A/B framework, Kalman filters, the error term of linear regression (Chapter 5.13) and the noise level in Gaussian processes.

How is it used?

Write the model as "observation = prediction + noise", choose a noise family (Normal, Student-t, Laplace…), and let the fitting procedure learn its location and scale. In NumPyro: numpyro.sample("y", dist.Normal(mu, sigma), obs=y).

model prediction actual values red stick = residual = actual − predicted the noisedistribution 0 = no miss
The model predicts a centre (orange line). Each actual value (blue) misses it by a residual (red stick). Collect all the red sticks and look at their pattern: that pattern is the noise distribution.
1 · Support which values can occur? 2 · Parameters which knobs can I turn? 3 · Mean where is the centre? 4 · Variance how spread out? 5 · Shape peak, skew, tails 6 · Assumptions what must be true? 7 · When used which jobs is it for? 8 · Relationships family links
The syllabus recipe for mastering any distribution. Every distribution in Chapters 4.7 to 4.11 gets a summary card that answers these eight questions in this order.

Drag the purple handle under the axis to move the location $\mu$, and the pink handle to change the scale $s$ (its distance from $\mu$). The dashed grey curve is the standard shape ($\mu = 0$, $s = 1$). Notice: doubling $s$ halves the peak height, because the area must stay 1. Switch between the three families and read the last line of the readout: only for the Normal does the scale equal the standard deviation.

"The noise distribution is the distribution of the data $y$."

It is the distribution of what is left after the model's prediction: $y_t - \hat y_t$. Daily demand with a trend and a weekly cycle is not bell-shaped as a whole, yet its residuals can be.

"A bigger scale makes the curve taller."

A bigger scale makes it wider and lower. The area under a density is always 1, so stretching it by $s$ divides its height by $s$.

Your forecasting model is $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$. Everything except $\epsilon_t$ builds the centre (the location). $\epsilon_t$ is the noise, and choosing a Normal, Student-t or (for counts) Negative Binomial likelihood is choosing its distribution. In an A/B framework like yours, a continuous metric per user is modelled the same way: group mean + noise, with a Normal or Student-t likelihood.

Residual = actual − predicted. A noise distribution describes the pattern of residuals.

Location–scale family: $X = \mu + sZ$, density $\tfrac{1}{s}p_Z\big(\tfrac{x-\mu}{s}\big)$: shift by $\mu$, stretch by $s$, peak drops by $s$.

Recipe for every distribution: support → parameters → mean → variance → shape → assumptions → when used → relationships.

Quick check: the standard Laplace has peak height 0.5. What is the peak height when the scale is $b = 4$?

Stretching by 4 divides the height by 4: $0.5/4 = 0.125$. (Directly from the formula: $1/(2b) = 1/8 = 0.125$.)

The Normal distribution core

Drop a ball through a board of pins. At every pin it bounces a little left or a little right. Where it lands is the sum of many small, independent nudges. Drop thousands of balls and they pile up in a bell shape: many near the middle, fewer and fewer further out.

Daily demand surprises work the same way: a bit of weather, a few extra visitors, one slow delivery, a small promotion elsewhere… Many small causes add up. Whenever a quantity is the sum of many small independent pieces, it tends to look like this bell, called the Normal (or Gaussian) distribution. (The precise reason is the Central Limit Theorem, Chapter 4.13.)

Three ways to say it:

  • Picture: a symmetric bell; its middle is the mean, its width is the standard deviation.
  • Numbers: with residual sd 10 orders, a miss bigger than 20 happens on about 4.6% of days, bigger than 30 on 0.27%.
  • Slogan: many small independent pushes, added up, make a bell.

Suppose forecast residuals follow a Normal with mean $\mu = 0$ and standard deviation $\sigma = 10$ orders.

  1. Height of the density at 0: $\dfrac{1}{\sigma\sqrt{2\pi}} = \dfrac{1}{10 \times 2.5066} = 0.0399$.
  2. Height at $x = 10$ (one sd away): multiply by $e^{-(10-0)^2/(2 \cdot 10^2)} = e^{-1/2} = 0.6065$, giving $0.0399 \times 0.6065 = 0.0242$.
  3. Probability of a miss within ±10 (one sd): $0.683$, about 68% of days.
  4. Probability of a miss bigger than 20 in size (two sd): $P(|e| \gt 20) = 0.0455$, about 1 day in 22.
  5. Bigger than 30 (three sd): $P(|e| \gt 30) = 0.0027$, about 1 day in 370.

Steps 3–5 come from the Normal's cumulative distribution function (Chapter 4.4), which has no simple formula; software or a table gives the numbers. The next concept shows how one table serves every Normal.

$X$ follows a Normal distribution with mean $\mu$ and variance $\sigma^2$, written $X \sim N(\mu, \sigma^2)$, when its density is

$$p(x) = \frac{1}{\sigma\sqrt{2\pi}}\exp\!\left(-\frac{(x-\mu)^2}{2\sigma^2}\right).$$
  • $(x-\mu)^2/\sigma^2$ is the squared distance from the centre, measured in standard deviations. The $\exp(-\ldots)$ makes the height fall as that distance grows. $1/(\sigma\sqrt{2\pi})$ is just the number that makes the area 1.
  • Take the log: $\log p(x) = -\frac{(x-\mu)^2}{2\sigma^2} + \text{constant}$. The log-density is an upside-down parabola. Keep this in mind for the tails section.
  • Notation trap: in maths the second number is the variance $\sigma^2$. NumPyro's dist.Normal(loc, scale), SciPy's norm(loc, scale) and NumPy's rng.normal(loc, scale) take the standard deviation $\sigma$.

Summary card: Normal

Supportall real numbers $(-\infty, \infty)$
Parameters$\mu$ (location) and $\sigma \gt 0$ (scale = standard deviation)
Mean$\mu$ (also the median and the mode)
Variance$\sigma^2$
Shapesymmetric bell; tails shrink like $e^{-x^2/2}$, extremely fast (skewness 0, excess kurtosis 0)
Assumptions (as noise)errors symmetric around the prediction, independent, with the same spread everywhere (unless modelled), and very large errors essentially never happen
When usedresidual noise in regression and forecasting, continuous A/B metrics, measurement error, priors on unbounded parameters such as coefficients, the approximate distribution of averages
Relationshipslimit of averages (CLT, 4.13); Student-t with $\nu \to \infty$; sums and $aX+b$ stay Normal; $e^X$ is Log-Normal (4.10); $Z^2$ of a standard Normal is $\chi^2_1$
Why do we need it?

It is the simplest sensible model for "many small causes add up". Its maths is easy: sums stay Normal, the log-likelihood is a sum of squares, and two numbers (mean and sd) describe it completely.

Where is it used?

Least-squares regression (squared error = Normal likelihood, Chapter 5.2), the Normal likelihood in your forecasting model, Normal priors on regression weights and seasonality coefficients, Kalman filters, Gaussian processes, z-tests and confidence intervals.

How is it used?

Estimate $\mu$ and $\sigma$ (by sample mean and sd, or inside a model), then read probabilities from the CDF: scipy.stats.norm(loc=mu, scale=sigma).cdf(x). Check the assumption with a residual histogram or a Q-Q plot (Chapter 4.17).

Move $\mu$ and watch the whole bell slide; move $\sigma$ and watch it widen and drop. The blue bars are 500 real random draws. Read the readout: the fractions of draws within 1, 2 and 3 standard deviations stay close to 68%, 95% and 99.7% whatever $\mu$ and $\sigma$ are. Press New sample several times to see how much a sample of 500 wobbles.

"$N(0, 4)$ has standard deviation 4."

In maths notation the second number is the variance, so the sd is $\sqrt4 = 2$. But dist.Normal(0, 4) in NumPyro has sd 4. Always check whether a library wants the variance or the standard deviation (NumPyro, SciPy and NumPy all want the sd).

"To use a Normal likelihood, my data must look Normal."

The assumption is about the noise around the model's mean, not about the raw data. Demand with trend and seasonality looks nothing like a bell overall; its residuals might.

"The Normal is the default because most real data is Normal."

It is popular because sums of small effects tend toward it and its maths is easy. Many real quantities are not Normal: counts (4.8), waiting times and revenues (4.10), and residuals with occasional big shocks (the Student-t, below).

$X \sim N(\mu, \sigma^2)$: $p(x) = \frac{1}{\sigma\sqrt{2\pi}}e^{-(x-\mu)^2/(2\sigma^2)}$. Mean $\mu$, variance $\sigma^2$.

Log-density = upside-down parabola. Comes from many small independent pushes added up.

Trap: maths writes the variance; NumPyro/SciPy/NumPy take the standard deviation.

Quick check: residuals are $N(0, 25)$. What is the standard deviation, and what NumPyro call describes them?

The variance is 25, so $\sigma = 5$. In NumPyro: dist.Normal(0.0, 5.0), because NumPyro's second argument is the standard deviation.

The 68–95–99.7 rule and z-scores (standardization) core

Every Normal curve has the same shape. A Normal with mean 200 and sd 20 is just the standard bell with its axis relabelled: "0" becomes 200 and every step of 1 becomes a step of 20.

So instead of asking "is 250 orders a lot?", ask "how many standard deviations above the mean is 250?". That number is the z-score. Once you have it, one single table (or one rule) answers the question for every Normal in the world.

Three ways to say it:

  • Picture: slide the bell to 0 and squeeze it to width 1; the value moves with it and lands at its z-score.
  • Numbers: 250 orders when the mean is 200 and the sd is 20 is $z = 2.5$: higher than about 99.4% of days.
  • Slogan: a z-score measures distance from the mean in units of standard deviations.

Daily demand is $N(200, 20^2)$. Today 250 orders arrived. How unusual is that?

  1. Distance from the mean: $250 - 200 = 50$ orders.
  2. In standard deviations: $z = 50/20 = 2.5$.
  3. From the standard Normal table: $P(Z \gt 2.5) = 1 - \Phi(2.5) = 1 - 0.99379 = 0.0062$.
  4. So a day this high or higher happens about once every $1/0.0062 \approx 161$ days.

A low day of 170 orders: $z = (170-200)/20 = -1.5$ and $P(Z \lt -1.5) = 0.0668$, about 1 day in 15. Not rare at all.

The middle 95% of days: $200 \pm 1.96 \times 20 = [160.8,\ 239.2]$ orders (1.96 is the z that leaves 2.5% in each tail).

The z-score of a value $x$ is

$$z = \frac{x - \mu}{\sigma}.$$

If $X \sim N(\mu, \sigma^2)$, then $Z = \dfrac{X-\mu}{\sigma} \sim N(0, 1)$, the standard Normal. Its cumulative distribution function (CDF) is written $\Phi(z) = P(Z \le z)$. Therefore

$$P(X \le x) = \Phi\!\left(\frac{x-\mu}{\sigma}\right).$$

The 68–95–99.7 rule (for Normal distributions only): $P(|Z| \le 1) = 0.6827$, $P(|Z| \le 2) = 0.9545$, $P(|Z| \le 3) = 0.9973$. For exactly 95%, use $\pm 1.96$.

Standardizing data means doing the same subtraction and division with the sample's mean and sd: $z_i = (x_i - \bar x)/s$. The result always has mean 0 and sd 1, but it keeps the data's shape (skew, tails). Chapter 4.18 covers standardization of data in depth.

Why do we need it?

Values in different units (orders, euros, seconds) cannot be compared directly. A z-score turns any value into "how unusual is this?", and the 68–95–99.7 rule turns that into a probability in your head.

Where is it used?

z-tests and p-values (Chapter 5.6), anomaly alerts on dashboards ("more than 3 sd from normal"), confidence intervals (±1.96 SE, Chapter 5.8), feature scaling before training, and the global scaler of your A/B framework.

How is it used?

Compute $z = (x-\mu)/\sigma$; then scipy.stats.norm.sf(z) gives $P(Z \gt z)$ and norm.cdf(z) gives $P(Z \le z)$. For data: (x - x.mean()) / x.std(ddof=1) or scikit-learn's StandardScaler.

μ−3σ μ−2σ μ−σ μ μ+σ μ+2σ μ+3σ 68.3% within 1 σ 95.4% within 2 σ 99.7% within 3 σ N(μ, σ²)
For every Normal distribution, whatever its μ and σ: about 68% of values lie within one standard deviation of the mean, 95% within two and 99.7% within three. Outside ±3σ there is only 0.3%: about 1 value in 370.

Drag the purple handle to a demand level. The axis has two rows: z-scores (grey) and orders (black). Change the mean and sd sliders: the bell never changes shape, only the black labels change. That is standardization. Try 250 with mean 200 and sd 20 (z = 2.5), then turn on Two-sided to count surprises in both directions.

"Standardizing (z-scoring) makes data Normal."

It only shifts and rescales. A right-skewed metric stays right-skewed after z-scoring, and the 68–95–99.7 percentages do not apply to it.

"68–95–99.7 holds for any data."

It is a Normal fact. For heavy-tailed noise, values beyond 3 sd are several times more common (you will see 5× for a Student-t with ν = 3 later in this chapter). The only rule that holds for every distribution is Chebyshev's much weaker bound, at least 75% within 2 sd (Chapter 4.12).

Your A/B framework's global scaler is this subtraction and division applied to data: one mean and one sd for all groups, so that standardized values are comparable and priors such as $N(0, 1)$ on a standardized effect mean the same thing everywhere. Why it must be one global pair and not one per group is the topic of Chapter 4.18.

$z = (x-\mu)/\sigma$; if $X \sim N(\mu, \sigma^2)$ then $P(X \le x) = \Phi(z)$.

68.3% within 1σ, 95.4% within 2σ, 99.7% within 3σ; 95% exactly within 1.96σ.

Trap: z-scoring rescales but does not change the shape; the rule is for Normals only.

Quick check: residuals are $N(0, 10^2)$. Roughly how often is a residual bigger than +20?

$z = 20/10 = 2$. About 95% lie within ±2σ, so 5% lie outside, half of it above: about 2.3% of days ($P(Z \gt 2) = 0.0228$), roughly 1 day in 44.

Adding Normals: variances add, standard deviations do not

Two stores each have their own daily surprise. What about the surprise in the total? Some days store A is up while store B is down, and the two partly cancel. So the total wobbles less than "A's wobble plus B's wobble".

The exact rule: for independent pieces, variances add. Standard deviations combine like the sides of a right-angled triangle: $\sqrt{15^2 + 20^2} = 25$, not $15 + 20 = 35$. And there is a bonus special to the Normal: the total is still exactly Normal.

Three ways to say it:

  • Picture: two bells added give a wider bell, but less wide than placing them end to end.
  • Numbers: sd 15 and sd 20 give a total with sd 25.
  • Slogan: independent noises add in variance; Normal plus Normal is Normal.

Store A's daily demand is $N(100, 15^2)$, store B's is $N(80, 20^2)$, and they are independent.

  1. Mean of the total: $100 + 80 = 180$.
  2. Variance of the total: $15^2 + 20^2 = 225 + 400 = 625$.
  3. Standard deviation: $\sqrt{625} = 25$.
  4. So the total is $N(180, 25^2)$. Chance the total exceeds 230: $z = (230-180)/25 = 2$, $P = 0.0228$.
  5. If you wrongly added standard deviations ($15 + 20 = 35$), you would get $z = 50/35 = 1.43$ and $P = 0.0766$: more than three times too high.

A week of one store. Seven independent days of $N(100, 15^2)$: the weekly total is $N(700,\ 7 \times 225) = N(700, 1575)$, sd $\sqrt{1575} = 39.7$ (not $7 \times 15 = 105$). The daily average over the week is $N(100,\ 225/7)$, sd $15/\sqrt7 = 5.67$.

If $X \sim N(\mu_1, \sigma_1^2)$ and $Y \sim N(\mu_2, \sigma_2^2)$ are independent, then

$$X + Y \sim N(\mu_1 + \mu_2,\ \sigma_1^2 + \sigma_2^2), \qquad aX + b \sim N(a\mu_1 + b,\ a^2\sigma_1^2).$$

For $n$ independent copies of $N(\mu, \sigma^2)$: the sum is $N(n\mu, n\sigma^2)$ and the average is $N(\mu, \sigma^2/n)$.

  • The mean and variance rules hold for any independent variables (Chapter 4.5). What is special about the Normal is that the shape stays Normal. Most families lose their shape when added: the sum of two independent Laplace noises is not Laplace, and the sum of two uniforms is a triangle.
  • If $X$ and $Y$ are correlated, add the covariance term: $Var(X+Y) = \sigma_1^2 + \sigma_2^2 + 2\,Cov(X, Y)$ (Chapter 4.15).
Why do we need it?

Totals and averages are everywhere: weekly demand from daily forecasts, total revenue from many users, the difference between two groups' means. This rule gives their distribution without any simulation.

Where is it used?

Aggregating forecasts across days or stores, the standard error of a mean $\sigma/\sqrt n$ and of a difference $\sqrt{SE_1^2 + SE_2^2}$ (Chapter 5.5), A/B test statistics, and the error budget of measurement chains.

How is it used?

Add the means, add the variances (only if independent), take the square root at the very end. If the pieces might be correlated (neighbouring days, nearby stores), include the covariance or you will be overconfident.

The blue bars are 3000 simulated days of the total demand of two stores. The orange curve uses "variances add"; the red dashed curve uses the wrong "sds add". Move both sliders and check that orange always fits. Then switch on correlated stores: the bars get wider than orange, because the independence assumption is broken and the covariance term is missing.

"Standard deviations add: sd 15 plus sd 20 gives sd 35."

Variances add (for independent pieces): $\sqrt{15^2 + 20^2} = 25$. Adding sds is only right in the extreme case where the two move in perfect lockstep (correlation 1).

"$2X$ and $X_1 + X_2$ are the same thing."

$2X$ doubles one draw: variance $4\sigma^2$. $X_1 + X_2$ adds two independent draws: variance $2\sigma^2$, because they partly cancel.

If your forecasting model's daily noise terms were independent $N(0, \sigma^2)$, the noise in a 7-day total would have sd $\sigma\sqrt7$, not $7\sigma$. If the residuals are correlated from day to day (autocorrelation, Chapter 7.3), the independent formula is too narrow, which is one reason to check residual autocorrelation before trusting weekly intervals.

Independent: $X+Y \sim N(\mu_1+\mu_2,\ \sigma_1^2+\sigma_2^2)$; $aX+b \sim N(a\mu+b,\ a^2\sigma^2)$.

Sum of $n$ iid: $N(n\mu, n\sigma^2)$; average: $N(\mu, \sigma^2/n)$.

Trap: variances add, sds do not. Correlated pieces need $+2Cov$.

Quick check: revenue = 2.5 × orders, and orders $\sim N(100, 15^2)$. What is the distribution of revenue?

$aX + b$ with $a = 2.5$, $b = 0$: mean $2.5 \times 100 = 250$, variance $2.5^2 \times 225 = 1406.25$, sd $2.5 \times 15 = 37.5$. Revenue $\sim N(250, 37.5^2)$.

The Student-t distribution: a bell that expects big surprises core

Most days your forecast misses by a little. But now and then something odd happens: a stock-out, a bot attack, a data glitch, a viral post. On those days the miss is huge. A Normal says a miss of 5 standard deviations happens about once in 1.7 million days. Real residuals often show such misses far more often.

The Student-t distribution is a bell with heavier tails: it keeps more probability far from the centre, so it "expects" the occasional big surprise. One extra knob, the degrees of freedom $\nu$ ("nu"), sets how heavy the tails are: small $\nu$ = very heavy tails; as $\nu$ grows the Student-t turns into the Normal.

Where do heavy tails come from? Imagine the noise level itself changes from day to day: calm days with small misses, wild days with big ones. Pool all days together and you get a mix of narrow and wide bells. That mix is taller in the middle and fatter in the tails than a single bell. A Student-t is exactly such a mix.

Three ways to say it:

  • Picture: a bell with a lower, longer skirt; $\nu$ is the dial from "very heavy skirt" to "plain Normal".
  • Numbers: with $\nu = 3$, a miss beyond 3 scale units has probability 0.058, about 21 times the Normal's 0.0027.
  • Slogan: a Student-t is a Normal that is not sure about its own width.

Compare the standard Student-t with $\nu = 3$ ($\mu = 0$, $\sigma = 1$) with the standard Normal.

  1. The t density at 0 is the constant $\dfrac{\Gamma(2)}{\Gamma(1.5)\sqrt{3\pi}}$. Here $\Gamma$ is the gamma function, a smooth version of the factorial: $\Gamma(2) = 1! = 1$ and $\Gamma(1.5) = \sqrt\pi/2 = 0.8862$. Also $\sqrt{3\pi} = 3.0700$. So the height is $1/(0.8862 \times 3.0700) = 0.3676$. The Normal's peak is $0.3989$: the t is a little lower in the middle.
  2. At $x = 3$: the t height is $0.3676 \times \left(1 + \tfrac{3^2}{3}\right)^{-2} = 0.3676/16 = 0.0230$. The Normal's height is $0.0044$. The t is about 5 times higher out there.
  3. Tail probability $P(|X| \gt 3)$: t with $\nu=3$ gives $0.0577$; Normal gives $0.0027$. Ratio about 21.
  4. Further out, $P(|X| \gt 5)$: t gives $0.0154$ (1 in 65); Normal gives $5.7 \times 10^{-7}$ (1 in 1.7 million).

$X$ follows a Student-t distribution with $\nu$ degrees of freedom, location $\mu$ and scale $\sigma$, written $X \sim t_\nu(\mu, \sigma)$, when

$$p(x) = \frac{\Gamma\!\left(\frac{\nu+1}{2}\right)}{\Gamma\!\left(\frac{\nu}{2}\right)\sqrt{\nu\pi}\;\sigma}\left(1 + \frac{1}{\nu}\left(\frac{x-\mu}{\sigma}\right)^2\right)^{-\frac{\nu+1}{2}}.$$
  • The fraction in front only makes the area 1. The shape is in the bracket: instead of $e^{-z^2/2}$ (Normal), the height falls like a power of $(1 + z^2/\nu)$, which is much slower. Far out it falls like $|x|^{-(\nu+1)}$.
  • Log-density: $-\frac{\nu+1}{2}\log\!\left(1 + \frac{z^2}{\nu}\right) + \text{constant}$: it grows only like a logarithm, not like a parabola.
  • Mixture form: draw a precision $\tau \sim \text{Gamma}(\tfrac{\nu}{2}, \text{rate } \tfrac{\nu}{2})$ (mean 1; Gamma is taught in Chapter 4.10), then draw $X \mid \tau \sim N(\mu, \sigma^2/\tau)$. The result is exactly $t_\nu(\mu, \sigma)$. "Precision" means $1/\text{variance}$: small $\tau$ = a wide Normal.
  • The name: W. S. Gosset published it in 1908 under the pen name "Student". It first appeared as the distribution of the t-statistic $(\bar x - \mu)/(s/\sqrt n)$, where $\nu = n - 1$ (Chapter 5.9). As a noise model, $\nu$ is just a tail-heaviness knob.

Summary card: Student-t

Supportall real numbers $(-\infty, \infty)$
Parameters$\nu \gt 0$ (degrees of freedom: tail heaviness), $\mu$ (location), $\sigma \gt 0$ (scale). NumPyro dist.StudentT(df, loc, scale); SciPy t(df, loc, scale)
Mean$\mu$ if $\nu \gt 1$; does not exist if $\nu \le 1$
Variance$\sigma^2\,\dfrac{\nu}{\nu-2}$ if $\nu \gt 2$; infinite if $1 \lt \nu \le 2$; undefined if $\nu \le 1$
Shapesymmetric bell; heavier (power-law) tails; slightly lower peak than a Normal with the same scale
Assumptions (as noise)errors symmetric, independent, constant scale; mostly ordinary with occasional large surprises
When usedrobust likelihood for residuals with outliers (both of your projects), metrics with occasional extreme users, priors that allow rare large effects, t-tests and t-intervals
Relationships$\nu \to \infty$: Normal; $\nu = 1$: Cauchy; Normal with a Gamma-distributed precision; $Z/\sqrt{V/\nu}$ with $Z \sim N(0,1)$ and $V \sim \chi^2_\nu$ independent
Why do we need it?

Real residuals often contain a few big shocks. Under a Normal likelihood those few points dominate the fit (their squared error is huge). A Student-t treats them as believable rare events, so the fit stays with the bulk of the data.

Where is it used?

The Student-t likelihood in your forecasting model and A/B framework, robust regression, financial returns, Kalman filters with outliers, heavy-tailed priors (for example on coefficients that are usually small but sometimes large), t-tests and t confidence intervals.

How is it used?

Replace dist.Normal(mu, sigma) by dist.StudentT(nu, mu, sigma). Either fix $\nu$ (3 to 5 is a common robust choice) or give it a prior and learn it. Compare against the Normal with residual Q-Q plots (Chapter 4.17) and predictive checks.

Start at ν = 3 and compare the orange t with the blue Normal: lower in the middle, much fatter in the tails. Switch to Zoom into the right tail: the shaded areas beyond 3 are the tail probabilities. Now slide ν up to 30 and 100: the t melts into the Normal. Slide ν down to 1 (the Cauchy) and read what happens to the mean and the variance.

calm, normal and wild days:three Normals, sd 0.5, 1 and 2 averagethem the mixture: tall middle,heavier tails dashed: one Normal with the same sd
Where heavy tails come from. If the noise level itself changes from day to day, the pooled residuals are a mix of narrow and wide Normals. The mix is taller in the middle and fatter in the tails than any single Normal with the same standard deviation. A Student-t is exactly such a mix (with the widths chosen by a Gamma distribution).

Each draw is made in two steps: first pick a precision $\tau$ from a Gamma distribution (mean 1), then draw from a Normal with variance $1/\tau$. The teal curves show 8 of these hidden Normals (drawn at 30% height). Press New sample: the histogram of all 2000 draws (blue) matches the orange Student-t, not the dashed Normal. Lower ν and the hidden widths become more varied, so the tails get heavier.

"ν is the sample size."

In a t-test, $\nu = n - 1$ comes from the sample size. In a noise model, $\nu$ is just a shape knob for tail heaviness, unrelated to how much data you have.

"With ν = 30 the t is basically Normal, so the tails do not matter."

Near the centre they are almost identical, but far out the difference remains: $P(|X| \gt 4)$ is 0.00038 for $\nu = 30$ and 0.000063 for the Normal, 6 times more.

"Heavier tails just means more spread everywhere."

Compared with a Normal of the same standard deviation, a Student-t is taller in the middle, thinner in the "shoulders" (around 1–2 sd) and fatter in the far tails. The probability moves from the shoulders to both the centre and the extremes.

Both of your projects offer a Student-t likelihood: dist.StudentT(df=nu, loc=mu, scale=sigma) in NumPyro. In the forecasting model it absorbs the occasional extreme day (a promotion that was not in the regressors, a data glitch) without bending the trend and seasonality toward it. In an A/B framework like yours it does the same for a few extreme users of a continuous metric. Check in your code whether $\nu$ is fixed or learned with a prior; if learned, its posterior tells you how heavy the tails of your residuals really are.

"The Student-t likelihood removes the outliers."

Nothing is removed. The Student-t assigns more probability to extreme residuals, so they exert less influence on the fit than under a Normal likelihood. Every point still enters the likelihood.

Model answer: "With a Normal likelihood, a residual of 10 standard deviations is nearly impossible, so the fit bends toward that point to make it less extreme. A Student-t gives that residual a believable probability, so the point keeps a small, smooth weight and the fit stays with the bulk of the data."

$t_\nu(\mu, \sigma)$: density $\propto \big(1 + \tfrac{1}{\nu}(\tfrac{x-\mu}{\sigma})^2\big)^{-(\nu+1)/2}$; tails fall like a power, not like $e^{-x^2}$.

Small $\nu$ = heavy tails; $\nu \to \infty$ = Normal; $\nu = 1$ = Cauchy. $P(|X| \gt 3)$: 0.058 ($\nu = 3$) vs 0.0027 (Normal).

Mixture: precision $\tau \sim$ Gamma($\nu/2$, rate $\nu/2$), then $N(\mu, \sigma^2/\tau)$. Say "less influence", never "removes outliers".

Quick check: as ν grows from 3 to 100, what happens to $P(|X| \gt 3)$ for a standard t?

It falls from 0.058 (ν = 3) through 0.013 (ν = 10) and 0.0054 (ν = 30) to 0.0034 (ν = 100), approaching the Normal's 0.0027. The tails thin out as the t turns into the Normal.

Degrees of freedom: when the mean or variance does not exist, and why the scale is not the sd

With very heavy tails, giant values arrive so often that averages never settle down. You average 1 000 draws and get something reasonable; then one enormous draw arrives and drags the average far away; later another one does it again. If this never stops, no matter how much data you collect, we say the mean (or the variance) does not exist or is infinite.

For the Student-t this depends only on $\nu$: for $\nu \le 2$ the running sample variance keeps jumping up forever (infinite variance); for $\nu \le 1$ even the running mean keeps jumping (no mean). The Cauchy ($\nu = 1$) is the famous example.

A second, very practical point: the number $\sigma$ inside StudentT(nu, mu, sigma) is a scale, not the standard deviation. Because the tails hold extra probability, the standard deviation is bigger: $\sigma\sqrt{\nu/(\nu-2)}$.

Three ways to say it:

  • Picture: a running average that keeps getting knocked off course by giants never finds a resting place.
  • Numbers: StudentT(df=3, scale=10) has standard deviation $10\sqrt3 = 17.3$, not 10.
  • Slogan: a moment exists only below ν: the mean needs ν > 1, the variance needs ν > 2.
  1. $\nu = 4$, $\sigma = 2$: variance $= \sigma^2\,\nu/(\nu-2) = 4 \times 4/2 = 8$, so sd $= \sqrt8 = 2.83$. The scale (2) and the sd (2.83) differ by 41%.
  2. $\nu = 3$, $\sigma = 10$: variance $= 100 \times 3/1 = 300$, sd $= 17.3$.
  3. $\nu = 2.5$: the factor is $\sqrt{2.5/0.5} = \sqrt5 = 2.24$: the sd is more than twice the scale.
  4. $\nu = 2$: the formula divides by $\nu - 2 = 0$: the variance is infinite. No finite sd exists.
  5. How much probability lies within ± one scale unit? For $\nu = 3$: $P(|X - \mu| \lt \sigma) = 0.609$. For a Normal (where scale = sd): 0.683.

For $X \sim t_\nu(\mu, \sigma)$, the average $E[|X|^k]$ is finite only when $k \lt \nu$. Hence:

  • Mean $E[X] = \mu$ exists only if $\nu \gt 1$. For $\nu \le 1$ the mean is undefined: the left and right tails are both infinitely heavy, so "the average" has no value at all.
  • Variance $Var(X) = \sigma^2\,\dfrac{\nu}{\nu - 2}$ for $\nu \gt 2$. For $1 \lt \nu \le 2$ it is infinite: the mean exists, but the average squared distance from it grows without limit.
  • Standard deviation $= \sigma\sqrt{\nu/(\nu-2)}$, always larger than the scale $\sigma$, and close to it only for large $\nu$ ($\nu = 30$: factor 1.035).

The sample mean and variance of any dataset are always finite numbers. When the true moment does not exist, these sample numbers never converge (the Law of Large Numbers needs a finite mean, Chapter 4.13).

Why do we need it?

If you fit a Student-t and then make intervals with "mean ± 2 × sigma", you use the wrong number. And if your residuals have no finite variance, any method that relies on the sample variance (standard errors, the CLT) quietly breaks.

Where is it used?

Reading fitted NumPyro StudentT parameters, comparing the noise level of a Normal model with a Student-t model, choosing a prior for $\nu$ that keeps it above 2 when you need a finite variance, and understanding why the Cauchy breaks the CLT (Chapter 4.13).

How is it used?

Convert scale to sd with $\sigma\sqrt{\nu/(\nu-2)}$ before comparing models. Make intervals from t quantiles: scipy.stats.t(df, loc, scale).ppf([0.025, 0.975]). If a learned $\nu$ ends up near or below 2, report quantiles, not standard deviations.

Each panel follows 2000 draws from a standard Student-t. Top: the running mean after n draws (the horizontal axis counts the draws). Bottom: the running sample variance (values above 10 run off the top). With ν = 30 and ν = 3 both settle (the purple dashed line is the true variance). With ν = 1.5 the mean settles but the variance keeps leaping up whenever a giant draw arrives. With ν = 1 (Cauchy) even the mean never settles. Press New sample a few times for each ν.

The orange curve is a Student-t with the chosen ν and scale σ. The purple bar spans ± σ (the scale); the pink bar spans ± one standard deviation. Lower ν and watch the pink bar grow away from the purple one; at ν = 2 there is no finite sd at all. The dashed blue curve is the Normal with the same sd as the t (shown when it exists): notice the t is taller in the middle.

"NumPyro's StudentT(df=3, loc=0, scale=10) has standard deviation 10."

Its sd is $10\sqrt{3/1} = 17.3$. The scale only equals the sd for a Normal.

"The Student-t model has a smaller sigma than the Normal model, so it found less noise."

The two sigmas measure different things. Fit both to the same residuals and the t's scale comes out smaller because its tails carry the big misses. Compare standard deviations or predictive intervals, not raw sigmas.

"If ν = 1.5 I can still compute np.var(residuals), so the variance exists."

A sample variance is always a finite number, but with infinite true variance it never settles: more data just brings bigger giants. Use quantiles or the scale instead.

If your forecasting model learns $\nu$ and the posterior puts weight on $\nu \le 2$, the implied residual variance is infinite and "± 2 sigma" bands are meaningless; use predictive quantiles from posterior samples instead (Chapter 7.14). A prior that keeps $\nu$ away from very small values (for example a Gamma prior with most mass above 2, see Chapter 4.10) is a common way to avoid this; check what your code does.

Student-t moments exist only below ν: mean needs ν > 1, variance needs ν > 2.

$Var = \sigma^2\nu/(\nu-2)$, so sd $= \sigma\sqrt{\nu/(\nu-2)} \gt \sigma$ (ν = 3: 1.73σ; ν = 4: 1.41σ).

Trap: the scale argument is not the sd. Infinite variance = sample variance never settles.

Quick check: you fit StudentT(df=5, loc=0, scale=4) to residuals. What is their standard deviation?

$4\sqrt{5/3} = 4 \times 1.291 = 5.16$. (Mean exists since ν > 1 and the variance is finite since ν > 2.)

The Laplace distribution: a sharp peak and exponential tails core

Take an exponential decay, "the height halves every so many steps", and put it on the right side of a centre point. Mirror it onto the left side. Glue the two halves together at the centre. You get a tent with a sharp tip: the Laplace distribution (also called the double exponential).

Compared with a Normal of the same spread, the Laplace puts more probability very close to the centre (the tall sharp tip), less in the "shoulders", and more far away (its tails shrink like $e^{-|x|}$, much more slowly than the Normal's $e^{-x^2}$). Its tails are still lighter than a Student-t's, which shrink only like a power.

It has a famous partner: the absolute value. A Laplace density is "exp of minus an absolute distance", so fitting a Laplace means minimizing absolute errors, and its best centre is the median.

Three ways to say it:

  • Picture: two exponential slides back to back, meeting at a sharp tip.
  • Numbers: with $b = 1$, $P(|X - \mu| \gt 3) = e^{-3} = 0.050$, about 18 times the standard Normal's 0.0027.
  • Slogan: Laplace = an absolute value inside the exponent.

Residuals follow a Laplace with centre $\mu = 0$ and scale $b = 5$ orders.

  1. Peak height: $1/(2b) = 1/10 = 0.1$.
  2. Right tail: $P(e \gt t) = \tfrac12 e^{-t/b}$. Both tails together: $P(|e| \gt t) = e^{-t/b}$.
  3. So $P(|e| \gt 10) = e^{-10/5} = e^{-2} = 0.135$.
  4. Variance $= 2b^2 = 2 \times 25 = 50$, so the sd is $\sqrt{50} = 7.07$ orders (not 5).
  5. The size of a typical miss: $|e|$ is exponential with mean $b = 5$; the median miss size is $b \ln 2 = 5 \times 0.693 = 3.47$.

$X \sim \text{Laplace}(\mu, b)$ when

$$p(x) = \frac{1}{2b}\exp\!\left(-\frac{|x-\mu|}{b}\right), \qquad \log p(x) = -\frac{|x-\mu|}{b} - \log(2b).$$
  • $|x - \mu|$ is the absolute distance from the centre; $b$ sets how fast the height decays. The log-density is a tent (a V upside down) with straight sides of slope $\pm 1/b$.
  • At $x = \mu$ the density has a corner (it is not smooth there), but its value is finite, $1/(2b)$. There is no lump of probability sitting exactly at $\mu$: $P(X = \mu) = 0$.
  • Fitting $\mu$ by maximum likelihood minimizes $\sum_i |x_i - \mu|$, whose answer is the sample median.

Summary card: Laplace

Supportall real numbers $(-\infty, \infty)$
Parameters$\mu$ (location) and $b \gt 0$ (scale). NumPyro dist.Laplace(loc, scale) with scale $= b$; SciPy laplace(loc, scale=b); NumPy rng.laplace(loc, scale)
Mean$\mu$ (also the median and the mode)
Variance$2b^2$, so sd $= \sqrt2\,b \approx 1.414\,b$
Shapesymmetric, sharp tip at $\mu$; exponential tails (heavier than Normal, lighter than any Student-t); excess kurtosis 3
Assumptions (as noise)symmetric, independent errors with a constant scale; errors cluster tightly near zero with occasional larger ones
When usedabsolute-error (L1) and median regression, robust noise models, priors that prefer values near zero (Bayesian lasso, your changepoint slope changes), the Laplace mechanism in differential privacy
Relationships$|X - \mu| \sim$ Exponential with mean $b$; the difference of two independent Exponentials with mean $b$ is Laplace(0, $b$); a Normal whose variance is exponentially distributed (mean $2b^2$) is Laplace
Why do we need it?

It is the noise model that matches absolute-error loss, and the prior that says "probably very close to zero, but a few large values are allowed". The Normal cannot express either idea.

Where is it used?

Median and quantile-style regression (L1 loss), the Laplace prior $\delta_j \sim \text{Laplace}(0, b)$ on trend changes in your forecasting model, the Bayesian lasso, sparse coding, and adding privacy noise to published statistics.

How is it used?

As a likelihood: dist.Laplace(mu, b) for the observations. As a prior: numpyro.sample("delta", dist.Laplace(0.0, b)) for coefficients you expect to be mostly near zero. Remember that the sd is $1.414\,b$ when comparing with Normal scales.

μ sharp tip (a corner) at μ left half: e^(−(μ−x)/b) right half: e^(−(x−μ)/b) dashed: Normal, same sd
A Laplace density is an exponential decay on each side of μ, mirrored and glued together. The glue point is a sharp corner. Compared with a Normal of the same standard deviation (dashed), it is taller at the centre, thinner in the shoulders and thicker far out.

Move $b$ and watch the tip rise and fall (its height is $1/(2b)$). Keep Normal with the same sd on: the Laplace is taller at the centre and has more mass far out. Switch to log density: the Laplace becomes a tent with straight sides, the Normal a parabola that dives faster and faster. Turn on the two halves to see the mirrored exponentials.

"$b$ is the standard deviation of the Laplace."

$b$ is the scale. The sd is $\sqrt2\,b \approx 1.41\,b$. NumPyro's Laplace(0, 0.5) has sd 0.71.

"The Laplace has heavier tails than the Student-t."

Laplace tails shrink exponentially ($e^{-|x|/b}$); Student-t tails shrink like a power ($|x|^{-(\nu+1)}$), which is eventually far slower. The order of tail weight is Normal < Laplace < Student-t.

"The sharp peak means there is a chunk of probability exactly at μ."

The density is finite at μ, so $P(X = \mu) = 0$. The tip only means that values near μ are more likely than under a Normal of the same sd.

$\text{Laplace}(\mu, b)$: $p(x) = \frac{1}{2b}e^{-|x-\mu|/b}$; mean $\mu$, variance $2b^2$ (sd $1.414b$).

Log-density = tent with slopes $\pm 1/b$. MLE of $\mu$ = sample median (absolute-error loss).

Tails: Normal < Laplace < Student-t. $P(|X-\mu| \gt t) = e^{-t/b}$.

Quick check: a Laplace has sd 2. What is $b$, and what is $P(|X - \mu| \gt 2)$?

$\sqrt2\,b = 2$ gives $b = \sqrt2 = 1.414$. Then $P(|X - \mu| \gt 2) = e^{-2/1.414} = e^{-1.414} = 0.243$. (A Normal with sd 2 gives $P(|Z| \gt 1) = 0.317$: the Laplace is more concentrated near the centre.)

Why a Laplace prior pulls values toward zero (a preview)

A prior is a distribution that describes what you believe about an unknown parameter before seeing the data (the full story is in Chapter 6.1). In your forecasting model, each candidate changepoint $j$ has a slope change $\delta_j$. You believe most of them are about zero (the trend did not really change there) and a few are large (real changes). You want a prior that says: "very likely tiny, but a big one is allowed".

The Laplace says exactly that: a tall sharp tip at 0 (lots of belief in tiny values) and tails that are not too thin (big values stay possible). Think of $-\log(\text{prior})$ as the price of a value. The Normal's price $\delta^2/(2\tau^2)$ is a U: almost flat near zero, so moving a little away from zero costs almost nothing. The Laplace's price $|\delta|/b$ is a V: it charges the same steady price per step all the way down to zero. That steady pull can push small estimates all the way to exactly zero.

Three ways to say it:

  • Picture: a ball in a V-shaped valley rolls right into the tip and can stop there; in a U-shaped valley it slows down and settles near, not at, the bottom.
  • Numbers: with the same sd of 1, the Laplace puts 13.2% of its probability within ±0.1 of zero; the Normal puts 8.0%.
  • Slogan: Laplace prior = "most effects about zero, a few big ones".

One noisy measurement $y$ of an effect $\delta$, with noise sd $s = 1$. Two priors with the same sd 1: Normal $N(0, 1^2)$ ($\tau = 1$) and Laplace$(0, b)$ with $b = 1/\sqrt2 = 0.707$. The MAP estimate is the single most probable value of $\delta$ after seeing $y$ (taught properly in Chapter 5.2).

  1. Normal prior: the MAP shrinks $y$ by a fixed fraction, $\hat\delta = y\,\dfrac{\tau^2}{\tau^2 + s^2} = y \times \dfrac12$.
  2. Laplace prior: minimize $\tfrac12(y-\delta)^2 + |\delta|/b$. For $\delta \gt 0$ the slope is $-(y-\delta) + 1/b$; setting it to 0 gives $\delta = y - 1/b = y - 1.414$, which is only allowed when $y \gt 1.414$. Otherwise the best value is the tip, $\delta = 0$. In one line: $\hat\delta = \text{sign}(y)\,\max(|y| - s^2/b,\ 0)$ ("soft thresholding").
  3. $y = 1$: Normal MAP $= 0.5$; Laplace MAP $= 0$ exactly (since $1 \le 1.414$).
  4. $y = 3$: Normal MAP $= 1.5$ (halved); Laplace MAP $= 3 - 1.414 = 1.586$ (only shifted).

The Laplace kills small signals and leaves big ones nearly intact; the Normal shrinks everything by the same fraction.

A Laplace prior $\delta \sim \text{Laplace}(0, b)$ has $-\log p(\delta) = |\delta|/b + \text{constant}$, the L1 penalty. A Normal prior $N(0, \tau^2)$ has $-\log p(\delta) = \delta^2/(2\tau^2) + \text{constant}$, the L2 penalty.

  • The "pull" toward zero is the slope of $-\log p$: Normal $\delta/\tau^2$ (fades to nothing near 0), Laplace $1/b$ for every $\delta \ne 0$ (never fades).
  • With a Normal likelihood, the MAP under a Laplace prior is the soft-threshold $\text{sign}(y)\max(|y| - s^2/b, 0)$: exactly zero inside the "dead zone" $|y| \le s^2/b$.
  • Precise statement: the L1 / lasso equivalence holds for the MAP point estimate. The full posterior under a Laplace prior has no lump at zero: its mean and median are small but generally not exactly zero. The posterior is "sparse-ish", not sparse. Chapter 5.3 derives this; Chapter 7.10 applies it to changepoints.
  • Smaller $b$ = narrower prior = stronger pull (bigger dead zone $s^2/b$).
Why do we need it?

With 25 candidate changepoints and limited data, letting every slope change be free overfits: the trend wiggles at every bump. A Laplace prior lets the data "buy" a few real changes while keeping the rest near zero.

Where is it used?

$\delta_j \sim \text{Laplace}(0, b)$ on trend changes in your forecasting model (Prophet does the same), the Bayesian lasso, sparse regression with many candidate regressors, and compressed sensing.

How is it used?

Put dist.Laplace(0.0, b) on coefficients you expect to be mostly near zero. Tune $b$: smaller means fewer, smaller changes. Expect exact zeros only from a MAP fit; posterior draws will be small but non-zero.

Top: the two priors, with the same sd. The shaded strip is ±0.1 around zero: the Laplace holds more there. Bottom: the pull toward zero (slope of −log prior). Drag the purple handle toward 0: the Normal's pull (blue line) fades to nothing, the Laplace's pull (teal) stays at full strength until the very tip. Change the prior sd: a narrower prior pulls harder.

The horizontal position of the purple handle is the observed value $y$. The dashed grey line is "no prior" (estimate = $y$). Blue: MAP with a Normal prior (a straight line through 0 with a smaller slope: everything shrinks by the same fraction). Teal: MAP with a Laplace prior: flat at exactly 0 inside the shaded dead zone, then parallel to the dashed line. Drag $y$ in and out of the dead zone; make the prior narrower or the noise bigger and watch the dead zone grow.

"A Laplace prior makes most $\delta_j$ exactly zero."

Only the MAP point estimate can land exactly on zero. Posterior draws, the posterior mean and the posterior median are pulled close to zero but are generally not exactly zero. The result is shrinkage, "sparse-ish", not true sparsity.

"A smaller $b$ is a weaker prior."

A smaller $b$ is a narrower prior that believes more strongly in zero: the dead zone $s^2/b$ gets wider and more changes are suppressed.

In your forecasting model, $\delta_j \sim \text{Laplace}(0, b)$ on the slope changes is exactly this V-shaped pull: most candidate changepoints get $\delta_j \approx 0$ and only changes the data strongly supports survive. Because you fit with SVI, each $\delta_j$ gets an approximate posterior (a smooth bump from the guide), not an exact zero. The choice of $b$ trades underfitting (too small) against overfitting (too large); Chapter 7.10 covers it in depth.

"Putting a Laplace prior on the coefficients is the same as Lasso regression."

The Lasso's L1 solution equals the MAP estimate under a Laplace prior (with a Normal likelihood). The full Bayesian posterior is different: it is continuous with no point mass at zero, so its mean is not sparse.

Model answer: "A Laplace prior adds an L1 penalty to the negative log posterior, so its MAP is the Lasso solution and can be exactly zero. But the posterior itself only shrinks the coefficients toward zero; in my forecasting model that gives sparse-ish changepoints, not exact zeros."

$-\log$ Laplace prior $= |\delta|/b$ (L1, a V); $-\log$ Normal prior $= \delta^2/(2\tau^2)$ (L2, a U).

Pull toward 0: Normal $\delta/\tau^2$ (fades), Laplace $1/b$ (constant). MAP = soft-threshold: $\text{sign}(y)\max(|y| - s^2/b, 0)$.

Trap: L1 = Laplace only at the MAP; the posterior is sparse-ish, not sparse. Smaller $b$ = stronger pull.

Quick check: noise sd $s = 1$ and a Laplace prior with $b = 0.5$. Which observations $y$ give a MAP of exactly zero?

The threshold is $s^2/b = 1/0.5 = 2$. Every observation with $|y| \le 2$ gives a MAP of exactly 0; $y = 3$ gives $3 - 2 = 1$.

Comparing the tails: a table and the log-density view core

Near the centre the three distributions look alike. They differ most in the tails, and that is exactly where the important behaviour lives: how surprised the model is by an extreme day, and how hard it fights to explain it.

On an ordinary plot the tails all look like "almost zero", so we cannot compare them. Two better views: a table of tail probabilities, and the log of the density. Taking the log stretches the tiny numbers apart: the Normal becomes a parabola that dives faster and faster, the Laplace becomes straight lines (a tent), and the Student-t bends and flattens out like a logarithm.

Three ways to say it:

  • Picture: on a log scale, Normal = falling parabola, Laplace = straight ramp, Student-t = a slope that keeps flattening.
  • Numbers: with the same sd of 1, a value beyond 5 has probability $5.7 \times 10^{-7}$ (Normal), $0.00085$ (Laplace) and $0.0032$ (Student-t, ν = 3).
  • Slogan: the tails decide how a model reacts to extremes.

Tail probabilities $P(|X - \mu| \gt x)$, computed with SciPy. First with the scale parameter equal to 1, then with all three rescaled to the same standard deviation 1 (t scale $= \sqrt{1/3} = 0.577$, Laplace $b = 1/\sqrt2 = 0.707$):

$x$NormalStudent-t, ν = 3Laplace
scale = 130.00270.05770.0498
same sd = 120.04550.04050.0591
same sd = 130.0027 (1 in 370)0.0138 (1 in 72)0.0144 (1 in 70)
same sd = 140.0000630.00620.0035
same sd = 150.000000570.00320.00085
  1. Laplace row, $x = 3$, same sd: $e^{-x/b} = e^{-3 \times 1.414} = e^{-4.243} = 0.0144$.
  2. At 2 sd the three are similar (the t even has a little less than the Normal: its probability sits in the centre and the far tails, not the shoulders).
  3. At 3 sd the t and the Laplace tie at about 5× the Normal. By 5 sd the t has 4× the Laplace and about 5 600× the Normal.

The tail probability at distance $x$ is $P(|X-\mu| \gt x)$. How fast it shrinks as $x$ grows defines the tail type:

  • Light (Gaussian) tails, Normal: shrinks like $e^{-x^2/(2\sigma^2)}$. Log-density: $-\frac{(x-\mu)^2}{2\sigma^2}$, a parabola.
  • Exponential tails, Laplace: exactly $e^{-x/b}$. Log-density: $-\frac{|x-\mu|}{b}$, straight lines.
  • Heavy (power-law) tails, Student-t: shrinks roughly like $x^{-\nu}$. Log-density: $-\frac{\nu+1}{2}\log\!\left(1 + \frac{(x-\mu)^2}{\nu\sigma^2}\right)$, growing only like $\log x$.

Minus the log-density is the penalty (or loss) the model charges for a residual. Fitting by maximum likelihood minimizes the total penalty. So the tail type is the loss function: squared error (Normal), absolute error (Laplace), log-type loss (Student-t).

Why do we need it?

Choosing a likelihood is choosing a tail. If your data has more extreme values than the likelihood allows, the fit is dragged around by them and the uncertainty bands are too narrow where it matters most.

Where is it used?

Choosing between Normal and Student-t likelihoods in your two projects, risk models (value at risk), alert thresholds ("3 sd" means very different things under different tails), loss-function choice in ML (MSE vs MAE vs robust losses such as Huber).

How is it used?

Look at residuals: count how many lie beyond 3 or 4 sd and compare with the table, draw a Q-Q plot (Chapter 4.17), or plot a log-histogram. If extremes are far more common than the Normal predicts, try a Student-t.

−4 −3 −2 −1 0 1 2 3 4 residual r (in scale units) penalty = −log density (+ constant) Normal: r²/2a U: steeper and steeper Laplace: |r|a V: the same slope everywhere Student-t (ν = 3): 2·log(1 + r²/3)flattens out far away
The "price" each model charges for a residual of size r is minus the log of its density. The Normal's price grows like r², so a far point is very expensive and drags the fit toward itself. The Laplace price grows like |r|. The Student-t price grows only like log r: a far point costs little extra, so the fit does not chase it.

The plot shows the right half (everything is symmetric). In log density view, compare the shapes: blue parabola, teal straight line, orange flattening curve. Drag the purple handle along the axis and read the tail table below. Turn off Same sd to use scale = 1 instead. Raise ν and watch the orange curve bend down toward the Normal.

"At 3 sd the Student-t (ν = 3) and the Laplace agree, so they are interchangeable."

They agree only around 3 sd. Further out the t's power-law tail wins by a growing factor (4× at 5 sd, more beyond). For rare extreme days the choice matters.

"Tails are tiny probabilities, so they do not matter in practice."

Three years of daily data is about 1 100 days. A "1 in 370" event happens about 3 times; a "1 in 1.7 million" event should never happen. If it does, a Normal model is badly surprised and bends its fit to explain it.

Tail type = how fast $P(|X - \mu| \gt x)$ shrinks: Normal $e^{-x^2/2}$ (light) < Laplace $e^{-x/b}$ (exponential) < Student-t $\approx x^{-\nu}$ (heavy).

Log-density: parabola / straight lines / log curve. Penalty = −log density: squared / absolute / log loss.

Same sd 1, beyond 3: 0.0027 / 0.0144 / 0.0138 (Normal / Laplace / t3).

Quick check: on a log-density plot, which family looks like two straight lines meeting at a point?

The Laplace: $\log p(x) = -|x-\mu|/b - \log(2b)$ is linear in $|x - \mu|$, so each side is a straight line with slope $\mp 1/b$, meeting at the corner $x = \mu$.

Robustness: how each noise model reacts to an outlier core

Fitting a model means finding the centre that makes the data most probable, which is the same as making the total penalty (minus the log density, summed over points) as small as possible. Now add one far-away point, an outlier:

  • Normal (penalty $r^2$): the outlier's penalty is enormous, and it shrinks a lot if the centre moves toward it. So the centre moves toward it. The answer is the mean.
  • Laplace (penalty $|r|$): every point pulls with the same strength, near or far. The answer is the median: the outlier counts as "one point on the right", not as "a point 20 units away".
  • Student-t (penalty $\log(1 + r^2/\nu)$): far points pull less and less. The model explains the outlier as a rare tail event and mostly ignores it. The answer is a weighted mean in which the outlier gets a tiny weight.

Three ways to say it:

  • Picture: each point pulls the centre with a rubber band. Normal bands get stronger the more they stretch; Laplace bands pull with a constant force; Student-t bands go slack when stretched too far.
  • Numbers: for 9, 10, 10, 11, 30 the centres are 14 (Normal), 10 (Laplace) and 10.10 (Student-t, ν = 3).
  • Slogan: robust = one bad point cannot drag the answer far.

Five days of orders: 9, 10, 10, 11 and one odd day with 30.

  1. Normal fit minimizes $\sum (x_i - \mu)^2$: the answer is the mean, $(9+10+10+11+30)/5 = 70/5 = 14$. Above four of the five days!
  2. Laplace fit minimizes $\sum |x_i - \mu|$: the answer is the median. Sorted: 9, 10, 10, 11, 30, so $\hat\mu = 10$.
  3. Student-t fit (ν = 3, centre and scale fitted together, by computer): $\hat\mu = 10.10$, $\hat\sigma = 1.49$. The fit is a weighted mean $\hat\mu = \sum w_i x_i / \sum w_i$ with weights $w_i = \dfrac{\nu+1}{\nu + (x_i - \hat\mu)^2/\hat\sigma^2}$.
  4. The outlier's weight: $r = 30 - 10.10 = 19.90$, $r/\hat\sigma = 13.39$, $r^2/\hat\sigma^2 = 179.4$, so $w = 4/(3 + 179.4) = 0.022$. The four ordinary days get weights 1.13, 1.33, 1.33 and 1.19.
  5. Move the odd day from 30 to 100: mean 28, median 10, Student-t 10.02. The t answer moves back toward 10 as the outlier gets more extreme.

For a location $\mu$ with penalty $\rho(r)$ on each residual $r_i = x_i - \mu$, the fit solves $\min_\mu \sum_i \rho(r_i)$, i.e. $\sum_i \psi(r_i) = 0$, where $\psi = \rho'$ is the influence (the pull of one point):

Noise modelPenalty $\rho(r)$ (scale 1)Pull $\psi(r)$Best centrePull of a far point
Normal$r^2/2$$r$meangrows without limit
Laplace$|r|$$\text{sign}(r)$medianconstant (bounded)
Student-t$\frac{\nu+1}{2}\log(1 + r^2/\nu)$$\dfrac{(\nu+1)\,r}{\nu + r^2}$weighted meanfalls back to 0 ("redescending")

An estimator is robust when a few extreme points cannot move it far. The t's pull is largest at $r = \sqrt\nu$ scale units and then fades. Chapter 4.14 covers robust statistics (median, MAD, trimmed means, breakdown point) in general.

Why do we need it?

Real data has glitches, bots, promotions and stock-outs. If one bad day can move your trend or your treatment effect a lot, your conclusions depend on luck. A robust likelihood makes the fit depend on the bulk of the data.

Where is it used?

Student-t likelihoods in your forecasting model and A/B framework, median and quantile regression, Huber loss in robust ML training, RANSAC in computer vision, and robust Kalman filters.

How is it used?

Swap the likelihood (dist.Normal → dist.StudentT), refit, and compare: if estimates change a lot, a few points were driving the Normal fit. Inspect those points (are they errors or real events?) rather than deleting them blindly.

Five days of orders; the bottom one is the odd day. Drag it right from 30 toward 60: the blue mean follows it, the teal median does not move, and the orange Student-t centre barely moves (and even drifts back). The small numbers are the weights the Student-t gives each day. Then drag the other points, or raise ν to 100 and watch the t centre become the mean.

Drag the purple handle to set the size of one residual $r$ (in scale units). The vertical axis is the pull $\psi(r)$, or in the other view the penalty $\rho(r)$ = −log density. In pull view: the Normal's pull (blue) keeps growing, the Laplace's (teal) is flat at 1, and the Student-t's (orange) rises, peaks at $r = \sqrt\nu$ and then falls back toward 0. Switch to penalty view to see the U, the V and the flattening log curve that cause these pulls.

"The Student-t fit drops the points beyond 3σ."

There is no cut-off. Every point keeps a smooth weight $(\nu+1)/(\nu + r^2/\sigma^2)$; far points just get small weights, decided by the model itself.

"The median is always better than the mean."

For truly Normal noise the mean is more efficient: the median needs about 57% more data for the same precision (relative efficiency $2/\pi \approx 0.64$, Chapter 5.1). Robustness has a price when there are no outliers.

"If the Student-t gives a different answer, delete the outliers and use the Normal."

First ask what the extreme points are. Data errors should be fixed; real events (a promotion, an outage) may need a regressor or a holiday term. The Student-t is a safeguard, not a replacement for looking.

This is the usual reason to offer both Normal and Student-t likelihoods, as both projects do. In the forecasting model, a single extreme day under a Normal likelihood can bend the trend, shift a changepoint slope $\delta_j$ or inflate a holiday effect; under a Student-t it keeps a small weight. In an A/B framework like yours, a few very heavy users of a continuous metric can move a Normal-likelihood group mean, and therefore $P(\theta_B \gt \theta_A \mid D)$; a Student-t likelihood reduces that sensitivity. Residual histograms, KDEs and Q-Q plots (4.16, 4.17) tell you which one your data needs.

Fit = minimize $\sum \rho(r_i)$ with $\rho = -\log$ density. Normal → mean; Laplace → median; Student-t → weighted mean, $w_i = \frac{\nu+1}{\nu + r_i^2/\sigma^2}$.

Pull $\psi$: Normal grows ($r$), Laplace constant ($\pm1$), Student-t redescends ($\frac{(\nu+1)r}{\nu + r^2}$, max at $r = \sqrt\nu$).

9, 10, 10, 11, 30 → 14 / 10 / 10.10. Robustness costs efficiency when the data really is Normal.

Quick check: with ν = 3 and scale 1, what weight does the Student-t give a point at residual 1, and one at residual 10?

$w = 4/(3 + r^2)$. At $r = 1$: $4/4 = 1$. At $r = 10$: $4/103 = 0.039$. The far point counts about 26 times less.

Recap, cheat sheet and practice

  • A noise distribution describes the residuals (actual − predicted). All three families here are location–scale families on the whole real line: $X = \mu + sZ$.
  • Normal $N(\mu, \sigma^2)$: the bell from many small independent pushes. 68–95–99.7 within 1, 2, 3 sd; z-scores $z = (x-\mu)/\sigma$; independent Normals add with variances adding and the shape staying Normal.
  • Student-t $t_\nu(\mu, \sigma)$: a Normal with a randomly varying width, hence heavy (power-law) tails. Becomes Normal as $\nu \to \infty$, Cauchy at $\nu = 1$. Mean needs $\nu \gt 1$, variance needs $\nu \gt 2$, and sd $= \sigma\sqrt{\nu/(\nu-2)}$, not $\sigma$.
  • Laplace$(\mu, b)$: $p \propto e^{-|x-\mu|/b}$, a sharp tip with exponential tails, variance $2b^2$. As a prior it pulls toward zero with constant force; its MAP is soft thresholding (L1), but the posterior is only sparse-ish.
  • Tails and robustness: Normal < Laplace < Student-t. Minus the log-density is the loss: squared → mean, absolute → median, log-type → a weighted mean that gives outliers small weights. Say "less influence", never "removes outliers".
NormalN(μ, σ²) Student-tt(ν, μ, σ) Cauchy= t with ν = 1 LaplaceLaplace(μ, b) Gamma(Chapter 4.10) Exponential(Chapter 4.10) ν → ∞ ν = 1 gives the randomprecision of a Normal difference of two Normal with an exponentiallydistributed random variance sums and aX + b stay Normal
How the three noise families relate. The Student-t becomes the Normal as ν grows and the Cauchy at ν = 1. Both the Student-t and the Laplace are "a Normal whose variance is itself random" (Gamma-distributed precision for the t, exponentially distributed variance for the Laplace). The difference of two independent Exponentials is Laplace, and |Laplace − μ| is Exponential. Arrows point from the ingredient to the result.

Cheat sheet

NormalStudent-tLaplace
Support$\mathbb{R}$$\mathbb{R}$$\mathbb{R}$
Parameters$\mu$, $\sigma$$\nu$, $\mu$, $\sigma$$\mu$, $b$
Density $\propto$$e^{-(x-\mu)^2/(2\sigma^2)}$$\left(1 + \frac{(x-\mu)^2}{\nu\sigma^2}\right)^{-(\nu+1)/2}$$e^{-|x-\mu|/b}$
Mean$\mu$$\mu$ if $\nu \gt 1$$\mu$
Variance$\sigma^2$$\sigma^2\nu/(\nu-2)$ if $\nu \gt 2$$2b^2$
Tailslight, $e^{-x^2/2}$heavy, $\approx x^{-\nu}$exponential, $e^{-x/b}$
Penalty / best centre$r^2$ / mean$\log(1 + r^2/\nu)$ / weighted mean$|r|$ / median
$P(|X| \gt 3)$, sd = 10.00270.0138 ($\nu = 3$)0.0144
NumPyroNormal(loc, scale)StudentT(df, loc, scale)Laplace(loc, scale)
SciPynorm(loc, scale)t(df, loc, scale)laplace(loc, scale)
Code it · Python

import numpy as np
from scipy import stats

# 1) The 68-95-99.7 rule
for k in (1, 2, 3):
    print(k, round(2 * stats.norm.cdf(k) - 1, 4))   # 0.6827, 0.9545, 0.9973

# 2) z-score: demand ~ N(200, 20^2), today 250 orders
z = (250 - 200) / 20
print(z, round(stats.norm.sf(z), 4))                 # 2.5 0.0062

# 3) Independent Normals: variances add (not sds)
rng = np.random.default_rng(0)
a = rng.normal(100, 15, 200_000)        # NumPy takes the sd, not the variance
b = rng.normal(80, 20, 200_000)
print(round((a + b).std(), 2))                       # 25.0  (not 35)

# 4) Tail probabilities P(|X| > 3) with scale 1
print(round(2 * stats.norm.sf(3), 4),                # 0.0027  Normal
      round(2 * stats.t(df=3).sf(3), 4),             # 0.0577  Student-t, nu = 3
      round(2 * stats.laplace(scale=1).sf(3), 4))    # 0.0498  Laplace, b = 1

# 5) Student-t: the scale is not the sd; nu <= 2 has infinite variance
print(round(stats.t(df=4, scale=2).std(), 3), stats.t(df=2).var())   # 2.828 inf

# 6) Laplace: variance 2 b^2
print(stats.laplace(scale=5).var())                  # 50.0

# 7) Robust centres for 9, 10, 10, 11, 30
x = np.array([9, 10, 10, 11, 30.0])
df, loc, scale = stats.t.fit(x, fdf=3)  # Student-t MLE with nu fixed at 3
print(x.mean(), np.median(x), round(loc, 2), round(scale, 2))   # 14.0 10.0 10.1 1.49

# 8) Laplace-prior MAP = soft thresholding (noise sd s, prior scale b)
def soft(y, s=1.0, b=1 / np.sqrt(2)):
    return np.sign(y) * np.maximum(np.abs(y) - s**2 / b, 0)
print(soft(np.array([1.0, 3.0])))                    # [0.  1.5858]

# 9) The same families in NumPyro: all take the SCALE, never the variance
import numpyro.distributions as dist
for d in (dist.Normal(0.0, 10.0), dist.StudentT(3.0, 0.0, 10.0), dist.Laplace(0.0, 0.5)):
    print(type(d).__name__, round(float(d.variance) ** 0.5, 3))
# Normal 10.0 / StudentT 17.321 (= 10 * sqrt(3)) / Laplace 0.707 (= 0.5 * sqrt(2))
Test yourself

1. In NumPyro you write dist.Normal(0.0, 4.0). What is the variance of this distribution?

NumPyro's second argument is the standard deviation (scale), so $\sigma = 4$ and the variance is $4^2 = 16$. In maths notation the same distribution is written $N(0, 16)$.

2. Two independent Normal noises have standard deviations 6 and 8. What is the standard deviation of their sum?

Variances add: $36 + 64 = 100$, so the sd is $\sqrt{100} = 10$. Adding sds (14) would be right only if the two moved in perfect lockstep.

3. A Student-t with $\nu = 2$ has…

The mean exists when $\nu \gt 1$ and the variance is finite only when $\nu \gt 2$. At $\nu = 2$ the mean is $\mu$ but the variance is infinite. "No mean" happens at $\nu \le 1$ (the Cauchy).

4. Which sentence describes a Student-t likelihood correctly?

Nothing is removed: every point stays in the likelihood with a smooth weight. The heavy tails make extreme residuals believable, so the fit does not chase them. ν is a shape knob here, and t tails are the heaviest of the three.

5. What is the standard deviation of dist.Laplace(0.0, 2.0)?

The scale is $b = 2$; the variance is $2b^2 = 8$, so the sd is $\sqrt8 = 2\sqrt2 \approx 2.83$.

6. Noise sd $s = 1$, Laplace prior with $b = 0.5$, and you observe $y = 1.5$. What is the MAP estimate?

Soft thresholding: the threshold is $s^2/b = 1/0.5 = 2$. Since $|y| = 1.5 \le 2$, the MAP sits exactly at the tip, 0. (The posterior mean would still be a small non-zero number.)

Practice problems

A. Daily demand is $N(500, 40^2)$. How often does demand exceed 600? Give the range that holds the middle 95% of days.

$z = (600-500)/40 = 2.5$, so $P = 1 - \Phi(2.5) = 0.0062$: about 1 day in 161. The middle 95%: $500 \pm 1.96 \times 40 = 500 \pm 78.4 = [421.6,\ 578.4]$.

B. Seven independent days of $N(500, 40^2)$. What is the distribution of the weekly total, and how likely is a week above 3 700?

Mean $7 \times 500 = 3500$; variance $7 \times 1600 = 11200$; sd $40\sqrt7 = 105.8$. So the total is $N(3500, 105.8^2)$. $z = 200/105.8 = 1.89$, $P = 0.029$, about 1 week in 34. (Wrongly using $7 \times 40 = 280$ as the sd would give $z = 0.71$ and a much larger probability.)

C. Residuals follow a Student-t with $\nu = 4$ and scale 3. What is their standard deviation? How often is a residual beyond ±3 scale units (±9), compared with a Normal?

sd $= 3\sqrt{4/2} = 3\sqrt2 = 4.24$. $P(|T| \gt 3)$ for $\nu = 4$ is 0.0399, about 1 in 25. For a Normal, ±3 of its own sd would give 0.0027: the t produces such misses about 15 times as often.

D. Laplace residuals with $b = 2$. Compute $P(|e| \gt 6)$ and compare with a Normal of the same sd.

$P(|e| \gt 6) = e^{-6/2} = e^{-3} = 0.0498$. The sd is $\sqrt2 \times 2 = 2.83$. A Normal with sd 2.83 gives $P(|Z| \gt 6/2.83) = P(|Z| \gt 2.12) = 0.034$. The Laplace has about 1.5 times more probability out there (and more near the centre too).

E. Data 4, 5, 5, 6, 20. Which centre does each noise model choose? What happens if 20 becomes 50?

Normal → mean $= 40/5 = 8$. Laplace → median $= 5$. Student-t → close to 5 (the 20 gets a tiny weight). If 20 becomes 50: the mean jumps to $70/5 = 14$, the median stays 5, and the Student-t centre stays near 5 (it even moves slightly closer to 5, because the pull of a far point fades).

F. Interview: "Your forecasting model has a Student-t option. Why, and what does ν do?"

Model answer: "Residuals from real demand data sometimes have rare large shocks that the regressors do not explain. Under a Normal likelihood such a day has a tiny probability, so the posterior bends the trend or seasonality to explain it. The Student-t has heavier tails, so it assigns those residuals more probability and they exert less influence; nothing is removed. ν controls tail heaviness: small ν means heavy tails, and as ν grows the likelihood approaches the Normal. The scale parameter is not the residual sd; the sd is $\sigma\sqrt{\nu/(\nu-2)}$ for ν > 2. I would choose between Normal and t with residual Q-Q plots and predictive checks."

Chapter 4.10 · Syllabus Module 9

Positive-quantity distributions: Exponential, Gamma, Log-Normal (+ Uniform)

Waiting times, order values, session lengths, prices and the noise scale σ of a model can never be negative, and they usually have a long right tail. The Normal ignores both facts. This chapter teaches the families built for positive numbers: the Exponential (waiting for the next event), the Gamma (adding waits, and the favourite prior for positive parameters), the Log-Normal (multiplying many factors), plus the Uniform as the "I only know the range" baseline.

  • Explain why positive, right-skewed data needs its own distributions, and use the coefficient of variation (sd/mean) as a first clue
  • Use the Exponential for waiting times: rate versus scale, the survival formula $e^{-\lambda t}$, and the memoryless property
  • Use the Gamma: as a sum of exponential waits, with shape α and rate β (and avoid the rate-versus-scale trap in SciPy and NumPy), and as a prior for positive rates and scales, including its link to the Negative Binomial
  • Use the Log-Normal: why products of many positive factors give it, why its μ and σ belong to log X, and why its mean is bigger than its median
  • Use the Uniform as a baseline and as the raw material for sampling every other distribution (inverse-CDF sampling)
  • Choose between these families from the data's story, its CV, its log-histogram and its tails

Colours in this chapter: blue bars = simulated data, orange curve = the distribution that explains them. When families are compared: Exponential blue, Gamma orange, Log-Normal teal, Uniform pink, and a Normal for comparison in grey dashes. Purple marks a special value (a threshold, a median); red marks impossible or tail regions.

Why positive data needs its own distributions core

How long until the next order? How much did this customer spend? How long did the session last? What is the noise level σ of my model? All of these are positive numbers. A wait of −3 minutes or a spend of −20 euros is impossible.

They usually share a second feature: a hard wall at zero on the left and a long tail on the right. Most orders are small, a few are huge. Most waits are short, a few are long. This lopsided shape is called right skew (the long tail points right).

A Normal distribution ignores both facts: it is symmetric, and it always puts some probability below zero. When the spread is large compared with the mean, that "impossible" probability becomes big and the model becomes nonsense.

Three ways to say it:

  • Picture: a wall at 0, a pile of small values near it, a long tail stretching right.
  • Numbers: orders with mean 40 euros and sd 30 euros: a Normal says 9% of orders are negative.
  • Slogan: if it cannot go below zero, do not model it with a bell that can.

Order values have mean 40 euros and standard deviation 30 euros.

  1. A Normal with these numbers: $P(X \lt 0) = \Phi\!\left(\frac{0 - 40}{30}\right) = \Phi(-1.33) = 0.091$. About 9% of orders would be negative: impossible.
  2. The coefficient of variation is $CV = \text{sd}/\text{mean} = 30/40 = 0.75$. For a Normal, $P(X \lt 0) = \Phi(-1/CV)$: the bigger the CV, the bigger the leak.
  3. A Gamma distribution (taught below) with the same mean and sd has shape $\alpha = (40/30)^2 = 1.78$ and rate $\beta = 40/30^2 = 0.0444$. Its probability below 0 is exactly 0, and its median is 32.8 euros, below the mean of 40, as the long right tail demands.

A distribution has positive support when every value it can produce is in $[0, \infty)$ (or $(0, \infty)$ if exactly 0 is impossible too). The support is the set of possible values (Chapter 4.9 uses the same word).

  • Right skew: a long tail on the right; then typically mode < median < mean.
  • Coefficient of variation: $CV = \sigma/\mu$, a unit-free measure of spread relative to the mean. It is a fingerprint: Exponential has $CV = 1$, Gamma $1/\sqrt\alpha$, Log-Normal $\sqrt{e^{\sigma^2} - 1}$.
  • Rule of thumb (not a law): a Normal model for a positive quantity is harmless when the CV is small, because it leaks $\Phi(-1/CV)$ below zero: CV 0.3 → 0.04%, CV 0.5 → 2.3%, CV 0.75 → 9.1%, CV 1 → 15.9%.
Why do we need it?

A model that puts probability on impossible values gives impossible predictions (negative revenue, negative waiting times), wrong intervals, and a symmetric shape that misses the long tail where the big money or the long waits are.

Where is it used?

Revenue per user and order values in A/B tests, delivery and waiting times, insurance claim sizes, server latencies, rainfall, and every positive parameter in a Bayesian model: noise scales σ, rates λ, Student-t ν, Negative Binomial dispersion α, Laplace scale b.

How is it used?

Check the support and the CV first. If the CV is small, a Normal may be fine. Otherwise pick a positive family (Exponential, Gamma, Log-Normal), or model log(y) with a Normal (Chapter 4.18). For parameters, use priors that live on $(0, \infty)$.

Whole line (−∞, ∞)Normal, Student-t, Laplace (Chapter 4.9) 0 Positive half-line [0, ∞)Exponential, Gamma, Log-Normal (this chapter) wall at 0 0waiting times, prices, revenue, σ Bounded [a, b]Uniform (this chapter); Beta on [0, 1] (Chapter 4.11) ab
The support is the set of values a quantity can take. Noise lives on the whole line. Waiting times, money amounts and scale parameters live on the positive half-line, with a hard wall at zero. Some quantities are bounded on both sides.

Set the mean and sd of a positive quantity (order value in euros). The grey dashed Normal puts the red shaded probability below 0, which is impossible. The orange Gamma has the same mean and sd but stops at the wall. Try mean 40, sd 30 (9% leak), then a small CV such as mean 100, sd 20 (the leak almost vanishes and both curves look alike), then a huge CV such as mean 20, sd 50.

"Just use a Normal and clip the negative predictions to zero."

Clipping changes the mean and piles fake probability at exactly 0, and the symmetric shape still misses the long right tail. Use a positive distribution, or model the log of the data.

"Positive data can never be modelled with a Normal."

When the CV is small (daily demand of 200 with sd 20 has CV 0.1), a Normal is a fine approximation: it leaks essentially nothing below zero. The problem appears when the spread is large compared with the mean.

Both of your models contain positive parameters: the noise scale σ of the Normal or Student-t likelihood, the Student-t degrees of freedom ν, the Negative Binomial concentration, the Laplace scale $b$ of the changepoint prior, the group-level spread in the hierarchical A/B model. Their priors must live on $(0, \infty)$; Half-Normal, Exponential, Gamma and Log-Normal are the usual choices. Check in your code which prior each positive parameter gets, and how wide it is.

Positive quantities: support $[0, \infty)$, usually right-skewed (mode < median < mean).

$CV = $ sd/mean. A Normal leaks $\Phi(-1/CV)$ below zero: fine for CV ≈ 0.1, bad for CV ≥ 0.5.

Trap: clipping a Normal at zero is not a fix; choose a positive family or log-transform.

Quick check: session lengths have mean 5 minutes and sd 5 minutes. How much probability would a Normal put below zero?

CV = 1, so $P(X \lt 0) = \Phi(-1) = 0.159$: almost 16% impossible sessions. A positive family is needed (CV = 1 is exactly the Exponential's fingerprint).

The Exponential distribution: waiting for the next event core

Orders arrive at random, but at a steady average rate: about 6 per hour. You have just received one. How long until the next?

Often the next one comes quickly; sometimes you wait a long time. Short waits are the most common, and the chance of a longer wait shrinks steadily: every extra minute of waiting multiplies the remaining chance by the same factor. That steady shrinking is an exponential decay, and the distribution of the wait is the Exponential distribution.

It is the twin of the Poisson distribution (Chapter 4.8): if the number of orders per hour is Poisson with rate λ, the gaps between orders are Exponential with the same rate λ.

Three ways to say it:

  • Picture: random dots on a timeline; the gaps between neighbouring dots are Exponential.
  • Numbers: 6 orders per hour means a mean gap of 10 minutes, and a wait longer than 20 minutes happens 13.5% of the time.
  • Slogan: Poisson counts the events; Exponential times the gaps.

Orders arrive at $\lambda = 6$ per hour $= 0.1$ per minute.

  1. Mean wait $= 1/\lambda = 1/0.1 = 10$ minutes. (This $1/\lambda$ is called the scale.)
  2. Chance the wait is longer than 20 minutes: $P(T \gt 20) = e^{-\lambda t} = e^{-0.1 \times 20} = e^{-2} = 0.135$.
  3. Chance the next order comes within 5 minutes: $P(T \le 5) = 1 - e^{-0.5} = 1 - 0.607 = 0.393$.
  4. Median wait: solve $e^{-0.1t} = 0.5$, so $t = \ln 2/0.1 = 6.93$ minutes, less than the mean of 10 (the long right tail pulls the mean up).
  5. Standard deviation $= 1/\lambda = 10$ minutes, equal to the mean, so $CV = 1$.

$T \sim \text{Exponential}(\lambda)$ with rate $\lambda \gt 0$ (events per unit time) has

$$p(t) = \lambda e^{-\lambda t}, \qquad P(T \le t) = 1 - e^{-\lambda t}, \qquad P(T \gt t) = e^{-\lambda t}, \qquad t \ge 0.$$

$P(T \gt t)$ is called the survival function: the chance that you are still waiting at time $t$. The scale $\theta = 1/\lambda$ is the mean waiting time; some libraries ask for $\lambda$, others for $\theta$.

Summary card: Exponential

Support$[0, \infty)$
Parametersrate $\lambda \gt 0$, or scale $\theta = 1/\lambda$. NumPyro dist.Exponential(rate=λ); SciPy expon(scale=1/λ); NumPy rng.exponential(scale=1/λ)
Mean$1/\lambda$ (median $\ln 2/\lambda \approx 0.69/\lambda$, mode 0)
Variance$1/\lambda^2$ (so sd = mean, CV = 1)
Shapehighest at 0, decays exponentially; right-skewed (skewness 2)
Assumptionsevents happen one at a time, independently, at a constant average rate (a Poisson process); no memory (next concept)
When usedtime between arrivals or failures, survival baselines, queueing, a weakly informative prior for a positive scale ("small values likely, large allowed")
RelationshipsGamma(1, λ); sum of $k$ → Gamma(k, λ); gaps of a Poisson process; $-\ln(U)/\lambda$ for Uniform $U$; minimum of independent Exponentials is Exponential with the rates added; $|$Laplace$(0, b) |$ is Exponential with mean $b$
Why do we need it?

To answer timing questions: how long until the next sale, the next failure, the next customer at the counter? And as the simplest "small is likely, big is possible" distribution for any positive quantity.

Where is it used?

Queueing models in call centres and servers, reliability (time to failure), survival analysis (as the constant-hazard baseline), simulation of arrival streams, and Exponential priors on scale parameters such as a noise σ in Bayesian models.

How is it used?

Estimate the rate as $\hat\lambda = 1/\bar t$ (one over the average gap). Read probabilities from $e^{-\lambda t}$. In code, be explicit: expon(scale=1/lam) in SciPy, dist.Exponential(rate=lam) in NumPyro.

0 min 15 min 30 min 45 min 60 min T₁=4 T₂=11 T₃=3 T₄=13 T₅=16 T₆=5 count in this hour = 6, a draw from Poisson(λ × 1 hour) gaps T₁, T₂, … between orders: each ~ Exponential(rate λ), independent blue dots: order arrival times
One random process, two descriptions. If orders arrive independently at a steady average rate λ, the number of orders in a window is Poisson (Chapter 4.8) and the gaps between orders are Exponential with the same rate. This is called a Poisson process.

Top: orders arriving at random over 3 hours at the chosen rate. Bottom: a histogram of all the gaps between orders (blue) and the Exponential density with the same rate (orange). Press Simulate 100 more hours a few times: the histogram settles onto the curve, the mean gap approaches 60/λ minutes and the sd of the gaps approaches the mean (CV ≈ 1).

Set the rate λ (per minute) and drag the purple handle to a waiting time $t$. In density view the red area is $P(T \gt t)$; in survival view the curve itself is $P(T \gt t) = e^{-\lambda t}$. Check the example: λ = 0.1 and t = 20 give $e^{-2} = 0.135$. Notice the median (purple dashed) is always below the mean (orange dashed).

"Exponential(2) has mean 2."

It depends on the convention. With rate 2 the mean is 0.5; with scale 2 the mean is 2. NumPyro takes the rate; SciPy and NumPy take the scale. Always write the keyword: rate= or scale=.

"The most likely waiting time is the mean."

The density is highest at 0: very short waits are the most likely. The median is only 69% of the mean, because a few long waits pull the mean up.

$T \sim \text{Exp}(\lambda)$: $p(t) = \lambda e^{-\lambda t}$, $P(T \gt t) = e^{-\lambda t}$; mean $1/\lambda$, sd $1/\lambda$, median $\ln2/\lambda$.

Gaps of a Poisson process with rate λ. Rate = events per time; scale = mean time = $1/\lambda$.

Trap: NumPyro uses the rate, SciPy/NumPy the scale.

Quick check: support tickets arrive at 3 per hour. What is the chance of a quiet hour with no ticket at all, computed in two ways?

Waiting time: $P(T \gt 1 \text{ hour}) = e^{-3 \times 1} = e^{-3} = 0.0498$. Counting: the number of tickets in an hour is Poisson(3), and $P(N = 0) = e^{-3} = 0.0498$. Same event, same answer.

The memoryless property: no event is ever "due"

You have already waited 20 minutes for the next order. Is it now "due", so that it will surely come soon? Under the Exponential model: no. The chance of waiting at least 10 more minutes is exactly the same as for someone who just started waiting. The process does not remember how long you have waited.

It is like tossing a fair coin: five tails in a row do not make heads more likely on the next toss. A constant rate means that in every minute the event has the same small chance of happening, whatever happened before.

Three ways to say it:

  • Picture: cut off the part of the waiting-time curve you have already lived through, rescale what is left, and you get the original curve back.
  • Numbers: with a mean wait of 10 minutes, $P(\text{10 more minutes}) = 0.368$, whether you have waited 0, 20 or 60 minutes already.
  • Slogan: an Exponential clock always starts fresh.

Orders at $\lambda = 0.1$ per minute. You have waited 20 minutes. What is the chance of waiting more than 10 further minutes?

  1. "Still waiting at 30, given still waiting at 20" is a conditional probability (Chapter 4.3): $P(T \gt 30 \mid T \gt 20) = \dfrac{P(T \gt 30)}{P(T \gt 20)}$ (being past 30 already includes being past 20).
  2. $= \dfrac{e^{-0.1 \times 30}}{e^{-0.1 \times 20}} = \dfrac{e^{-3}}{e^{-2}} = e^{-1} = 0.368$.
  3. Starting fresh: $P(T \gt 10) = e^{-0.1 \times 10} = e^{-1} = 0.368$. Identical.

Compare a Gamma waiting time with shape 3 and the same mean of 10 minutes: fresh, $P(T \gt 10) = 0.423$; after 20 minutes of waiting, $P(T \gt 30 \mid T \gt 20) = 0.101$. For the Gamma, having waited does make the end nearer.

A waiting time $T$ is memoryless when, for all $s, t \ge 0$,

$$P(T \gt s + t \mid T \gt s) = P(T \gt t).$$

For the Exponential: $\dfrac{e^{-\lambda(s+t)}}{e^{-\lambda s}} = e^{-\lambda t}$. ✓ The Exponential is the only continuous distribution with this property (the Geometric is its discrete cousin).

Equivalent idea: the hazard rate $h(t) = p(t)/P(T \gt t)$, the chance per unit time that the event happens right now given that it has not happened yet, is constant: $h(t) = \lambda e^{-\lambda t}/e^{-\lambda t} = \lambda$. Increasing hazard (wear-out, "due") or decreasing hazard (the longer a user has been away, the less likely they return) both break memorylessness.

Why do we need it?

It tells you when the Exponential is the right model and when it is not. If "time already waited" changes the outlook, an Exponential model will give wrong predictions, however well its mean fits.

Where is it used?

Queueing theory (call centres, server requests), radioactive decay, Markov models in continuous time, survival analysis (checking whether the hazard is constant), and the Poisson process behind count data.

How is it used?

Ask the domain question: does the chance per minute change with time already waited? Check the data: plot the empirical hazard or compare $P(T \gt s+t \mid T \gt s)$ for several $s$. If it changes, use a Gamma, Weibull or Log-Normal waiting time instead.

The blue curve is $P(T \gt x)$ from a fresh start. Drag the purple handle to the time $s$ you have already waited (grey). Orange: your remaining-wait curve, $P(T \gt x \mid T \gt s)$. Dashed blue: a fresh curve started at $s$. For the Exponential they lie exactly on top of each other wherever you put $s$. Switch to the Gamma: the orange curve drops below, because the end really is getting closer.

"We have had no order for 30 minutes, so one is due any second."

Under a constant-rate (Exponential) model, the wait from now on has the same distribution as always. Thinking otherwise is the gambler's fallacy.

"Real waiting times are always memoryless."

Many are not. Machines wear out (rising hazard), a delivery that has not left the warehouse yet is not "fresh", and users who have been inactive for months are less likely to return (falling hazard). Memorylessness is an assumption to check.

Memoryless: $P(T \gt s+t \mid T \gt s) = P(T \gt t)$. Only the Exponential (continuous) has it.

Same thing: constant hazard $h(t) = p(t)/P(T \gt t) = \lambda$.

Trap: if waiting longer changes the outlook (wear-out, churn), the Exponential is the wrong model.

Quick check: a light bulb's lifetime is Exponential with mean 1 000 hours. It has already burned for 1 500 hours. What is its expected remaining life?

Memoryless: the remaining life is again Exponential with mean 1 000 hours, so the expected remaining life is 1 000 hours. (Real bulbs wear out, which is exactly why real lifetimes are often not Exponential.)

The Gamma distribution: adding up waits, shape and rate core

Now wait not for the next order but for the third one. That wait is three exponential gaps added together. Adding waits gives the Gamma distribution: with α gaps, each with rate λ, the total wait is Gamma(α, λ).

Two knobs control it. The shape α decides the form: α = 1 is the Exponential (highest at 0); α = 2 to 5 is a lopsided hump; large α looks almost Normal (many gaps added, the Central Limit Theorem again). α need not be a whole number: it is a smooth dial between these shapes. The rate β (= λ) stretches or squeezes the time axis: a faster rate means shorter waits.

Because it is positive, flexible and mathematically convenient, the Gamma is also the most common prior for positive parameters (next concept).

Three ways to say it:

  • Picture: stack α exponential gaps end to end; the total length is one Gamma draw.
  • Numbers: three gaps of mean 10 minutes: total mean 30, sd $10\sqrt3 = 17.3$.
  • Slogan: Gamma = Exponential waits added up.

Orders at $\lambda = 0.1$ per minute. How long until the third order? $T_3 \sim \text{Gamma}(\alpha = 3, \beta = 0.1)$.

  1. Mean: three gaps of mean 10: $\alpha/\beta = 3/0.1 = 30$ minutes.
  2. Variance: independent gaps, so variances add: $3 \times 10^2 = 300$ $(= \alpha/\beta^2 = 3/0.01)$. sd $= \sqrt{300} = 17.3$ minutes.
  3. Most likely value (mode): $(\alpha - 1)/\beta = 2/0.1 = 20$ minutes. Median 26.7 minutes. Again mode < median < mean.
  4. Chance of waiting more than an hour for the third order: that is the same as "at most 2 orders in 60 minutes", and the count in 60 minutes is Poisson($0.1 \times 60 = 6$): $P = e^{-6}\left(1 + 6 + \tfrac{6^2}{2}\right) = 25e^{-6} = 0.062$.

$X \sim \text{Gamma}(\alpha, \beta)$ with shape $\alpha \gt 0$ and rate $\beta \gt 0$ has density

$$p(x) = \frac{\beta^\alpha}{\Gamma(\alpha)}\, x^{\alpha-1} e^{-\beta x}, \qquad x \gt 0.$$
  • $x^{\alpha-1}$ shapes the rise from zero; $e^{-\beta x}$ makes the long right tail decay exponentially; $\Gamma(\alpha)$, the gamma function ($\Gamma(n) = (n-1)!$ for whole numbers), makes the area 1.
  • The scale is $\theta = 1/\beta$. Some libraries take the rate, some the scale: this is the most common Gamma bug.

Summary card: Gamma

Support$(0, \infty)$
Parametersshape $\alpha \gt 0$ and rate $\beta \gt 0$ (or scale $\theta = 1/\beta$). NumPyro dist.Gamma(concentration=α, rate=β); SciPy gamma(a=α, scale=1/β); NumPy rng.gamma(shape=α, scale=1/β)
Mean$\alpha/\beta$
Variance$\alpha/\beta^2$; so $CV = 1/\sqrt\alpha$
Shaperight-skewed (skewness $2/\sqrt\alpha$); $\alpha \lt 1$: density shoots up at 0; $\alpha = 1$: Exponential; $\alpha \gt 1$: hump with mode $(\alpha-1)/\beta$; large $\alpha$: nearly Normal
Assumptionsas a waiting time: α independent exponential stages with the same rate; in general: positive, right-skewed, with an exponentially decaying right tail
When usedtotal waiting times, claim sizes and rainfall, priors for positive rates, scales and precisions, the mixing distribution behind the Negative Binomial and the Student-t
RelationshipsExponential = Gamma(1, β); sum of independent Gamma($\alpha_i$, β) = Gamma($\sum\alpha_i$, β); $\chi^2_k$ = Gamma($k/2$, rate $\tfrac12$); $cX \sim$ Gamma(α, β/c); Poisson with a Gamma rate = Negative Binomial (4.8); Normal with a Gamma precision = Student-t (4.9); $X_1/(X_1+X_2)$ of two independent Gammas with the same rate = Beta (4.11)
Why do we need it?

It is the flexible workhorse for positive quantities: one knob for shape, one for size. It covers spiky-at-zero, humped and nearly bell-shaped positive data, and it combines neatly with the Poisson and the Normal.

Where is it used?

Waiting for the k-th event, insurance and hydrology, Gamma priors on Poisson rates and Normal precisions, Gamma priors on the Student-t ν, Gamma GLMs for positive skewed outcomes, and the gamma–Poisson mixture behind Negative Binomial counts in forecasting.

How is it used?

Pick α and β from a mean and a CV ($\alpha = 1/CV^2$, $\beta = \alpha/\text{mean}$), or fit them by maximum likelihood (scipy.stats.gamma.fit(x, floc=0)). Always write the keyword (rate= or scale=) and print the mean to check.

gap 1: 8 min gap 2: 14 min gap 3: 6 min time until the 3rd order = 8 + 14 + 6 = 28 min: one draw from Gamma(α = 3, rate λ) each gap ~ Exponential(rate λ), independent 3rd order
Add up α independent exponential waits with the same rate and you get a Gamma(α, rate λ) wait. Its mean is α/λ: three gaps of 10 minutes on average give 30 minutes on average.

Choose how many gaps α to add (each gap is Exponential with mean 10 minutes). Top: six example totals, each made of α coloured gaps. Bottom: 2000 totals (blue) and the Gamma(α, rate 0.1) density (orange). With α = 1 you get the Exponential; with α = 3 the example's hump around 20–30 minutes; with α = 10 the histogram already looks almost Normal.

Move the shape α: below 1 the density shoots up at zero, at 1 it is the Exponential, above 1 it becomes a hump that grows more symmetric. Move the rate β: a bigger rate squeezes everything toward zero (mean α/β). Then switch on the trap: the red dashed curve is what you get if you pass β to SciPy or NumPy as the scale. With α = 2, β = 0.5 the correct mean is 4 but the trap gives 1.

"scipy.stats.gamma(2, 4) is the Gamma with shape 2 and rate 4."

SciPy's second positional argument is loc (a shift!), and its scale is $1/\text{rate}$. Shape 2, rate 4 is gamma(a=2, scale=0.25), mean 0.5. NumPy's rng.gamma(2, 4) means shape 2, scale 4, mean 8. NumPyro's Gamma(2, 4) means shape 2, rate 4, mean 0.5.

"The Gamma's β is always the scale."

Textbooks use both conventions. In this guide (and in NumPyro, Stan and PyMC) β is the rate and the mean is α/β. Whenever you see Gamma(α, β), find out which one is meant before using it.

"A Gamma(2, 4) prior."

"A Gamma prior with shape 2 and rate 4, so mean 0.5 and sd about 0.35." Naming the convention and the implied mean removes all ambiguity.

Model answer: "I parameterize the Gamma by shape α and rate β, mean α/β and variance α/β². SciPy and NumPy use scale = 1/β, so in SciPy that is gamma(a=α, scale=1/β); NumPyro's Gamma(concentration, rate) matches my convention."

$\text{Gamma}(\alpha, \beta)$: $p(x) \propto x^{\alpha-1}e^{-\beta x}$; mean $\alpha/\beta$, variance $\alpha/\beta^2$, CV $1/\sqrt\alpha$, mode $(\alpha-1)/\beta$.

Sum of α Exponential(β) waits; α = 1 is the Exponential; large α ≈ Normal.

Trap: rate vs scale. NumPyro Gamma(α, rate); SciPy gamma(a=α, scale=1/β); NumPy gamma(shape, scale).

Quick check: a Gamma has mean 8 and sd 4. What are α and β?

$CV = 4/8 = 0.5$, so $\alpha = 1/CV^2 = 4$, and $\beta = \alpha/\text{mean} = 4/8 = 0.5$. Check: mean $4/0.5 = 8$ ✓, variance $4/0.25 = 16$, sd 4 ✓.

The Gamma as a prior for positive parameters, and the gamma–Poisson link

Bayesian models are full of parameters that must be positive: a Poisson rate λ (orders per day), a noise scale σ, a precision $1/\sigma^2$, the Negative Binomial concentration, the Student-t ν. Before seeing data you describe your belief about each one with a prior distribution, and it must live on $(0, \infty)$.

The Gamma is the classic choice because its two knobs match two natural questions: "Where do I think the value is?" (the prior mean) and "How sure am I?" (the shape α: bigger α = narrower prior, since CV $= 1/\sqrt\alpha$).

It also has two famous partnerships. A Gamma prior on a Poisson rate stays a Gamma after seeing count data (it is conjugate, see below). And if the daily rate itself varies from day to day like a Gamma, the counts you observe are Negative Binomial: this is the gamma–Poisson mixture of Chapter 4.8.

Three ways to say it:

  • Picture: a positive hump you can slide (mean) and squeeze (α).
  • Numbers: "about 20 orders a day, give or take 10" is Gamma(α = 4, β = 0.2).
  • Slogan: mean says where, α says how sure.

You believe a store gets about 20 orders per day, and you would not be surprised by 10 or 35.

  1. Choose the prior mean $m = 20$ and a CV of 0.5 ("give or take half"): $\alpha = 1/0.5^2 = 4$.
  2. Rate: $\beta = \alpha/m = 4/20 = 0.2$. Prior: $\lambda \sim \text{Gamma}(4, 0.2)$, sd $= \sqrt4/0.2 = 10$, mode $(4-1)/0.2 = 15$.
  3. Its central 90% interval (from software) is 6.8 to 38.8 orders per day. With $\alpha = 16$ (CV 0.25) it shrinks to 12.5 to 28.9.
  4. What daily counts does this prior predict? Draw λ from the Gamma, then a count from Poisson(λ). The counts are Negative Binomial with mean 20 and variance $20 + 20^2/4 = 120$, far more spread than a plain Poisson(20) (variance 20).
  5. Conjugacy, as a preview: after 7 days with 175 orders in total, the posterior is again a Gamma, $\text{Gamma}(4 + 175,\ 0.2 + 7) = \text{Gamma}(179, 7.2)$, mean $179/7.2 = 24.9$ (Chapter 6.3 derives this).

To build a Gamma prior from a prior mean $m$ and a coefficient of variation $c$: $\alpha = 1/c^2$ and $\beta = \alpha/m$.

  • $\alpha \le 1$: highest density at 0; says "small values are quite possible". $\alpha \gt 1$: the density is 0 at 0 and peaks at $(\alpha-1)/\beta$; says "the value is not near zero". Useful when you know a scale cannot be tiny.
  • A prior is conjugate to a likelihood when the posterior is in the same family as the prior. Gamma(α, β) prior + Poisson counts $y_1, \dots, y_n$ → posterior Gamma$(\alpha + \sum y_i,\ \beta + n)$.
  • Gamma–Poisson mixture: if $\lambda \sim \text{Gamma}(\alpha, \alpha/\mu)$ and $Y \mid \lambda \sim \text{Poisson}(\lambda)$, then $Y \sim \text{NB}(\mu, \alpha)$ with $Var(Y) = \mu + \mu^2/\alpha$ (the NB2 form of Chapter 4.8; NumPyro NegativeBinomial2(mean, concentration)).
  • A common default prior for the Student-t ν is Gamma(2, rate 0.1): mean 20, mode 10, and only 1.8% of its mass below ν = 2. It is a widely used recommendation, not a law.
Why do we need it?

Every positive parameter needs a prior that cannot go negative and that you can tune with intuitive numbers. Without it, a sampler or an SVI optimizer can wander into impossible values or into absurdly large ones.

Where is it used?

Gamma priors on Poisson rates (count metrics in A/B tests), on precisions in Normal models, on the Student-t ν, on Negative Binomial concentrations; the gamma–Poisson mixture behind NB likelihoods in forecasting; Gamma GLMs.

How is it used?

Decide a prior mean and how uncertain you are, convert to α and β, then check the implied 90% interval (stats.gamma(a, scale=1/b).ppf([0.05, 0.95])) and simulate the data it predicts (a prior predictive check, Chapter 6.2).

Set what you believe the daily rate is (the prior mean) and how sure you are (α). The purple band is the central 90% of the prior. Raise α and watch the band narrow (CV = 1/√α). Then switch to daily counts it predicts: blue bars are counts simulated by drawing a rate from the prior and then a Poisson count. They are much wider than a Poisson with a fixed rate (grey) and match the Negative Binomial (orange). Large α brings them close to the Poisson.

"A Gamma(0.001, 0.001) prior is uninformative."

It looks vague, but it piles enormous density near zero and its tail is extremely long; for scale parameters it can strongly change the answer when data is scarce. "Flat-looking" is not "no information" (Chapter 6.2).

"A Gamma prior on σ and a Gamma prior on the precision $1/\sigma^2$ say the same thing."

They are different beliefs: a Gamma on the precision implies a very different shape for σ. Check which quantity your code puts the prior on.

In your models the positive parameters (noise scale σ, Student-t ν, Negative Binomial concentration, the Laplace scale $b$ of the changepoint prior, group-level spreads) each need a positive prior. A Gamma with α > 1 keeps a parameter away from zero (a common choice for ν, which should not sit near 0, the extreme heavy-tail end); an Exponential or Half-Normal allows values near zero (common for spreads). In the forecasting model, a Negative Binomial likelihood for daily counts is exactly the gamma–Poisson story: the day's underlying rate wobbles around the model's mean, and the concentration says how much. Check which priors and which NB parameterization your code uses.

Gamma prior from a mean $m$ and a CV $c$: $\alpha = 1/c^2$, $\beta = \alpha/m$. α ≤ 1 allows values near 0; α > 1 pushes away from 0.

Conjugate: Gamma(α, β) + Poisson counts → Gamma(α + Σy, β + n). Gamma-mixed Poisson = NB, Var $= \mu + \mu^2/\alpha$.

Trap: tiny-α "vague" Gammas are not uninformative; know whether the prior is on σ or on $1/\sigma^2$.

Quick check: you want a prior for a rate with mean 5 and sd 5. Which Gamma?

CV = 1, so α = 1 and β = 1/5 = 0.2: Gamma(1, 0.2), which is the Exponential with mean 5. Its highest density is at 0.

The Log-Normal distribution: products of many positive factors core

What a customer spends is a product of many positive factors: a base basket × a seasonal factor × a promotion factor × a device factor × a mood factor… Each factor nudges the total up or down by some percentage.

Logs turn products into sums: $\log(a \times b \times c) = \log a + \log b + \log c$. A sum of many small independent pieces is close to Normal (the Central Limit Theorem, Chapter 4.13). So the log of the spend is close to Normal, and the spend itself follows a Log-Normal distribution: positive, right-skewed, with a long right tail.

Three ways to say it:

  • Picture: a lopsided hump with a long right tail; take logs and it becomes a symmetric bell.
  • Numbers: if log-spend is $N(\ln 30, 0.8^2)$, the median spend is 30 euros but the mean is 41.3 euros.
  • Slogan: Normal adds, Log-Normal multiplies.

Products become sums. Start at 100 and multiply by 1.10, 0.95 and 1.20: $100 \times 1.10 \times 0.95 \times 1.20 = 125.4$. With logs: $4.605 + 0.095 - 0.051 + 0.182 = 4.832$, and $e^{4.832} = 125.4$.

Order values. Suppose $\log X \sim N(\mu, \sigma^2)$ with $\mu = \ln 30 = 3.401$ and $\sigma = 0.8$.

  1. Median: half the logs are below $\mu$, so half the values are below $e^\mu = 30$ euros.
  2. Mean: $e^{\mu + \sigma^2/2} = 30 \times e^{0.32} = 30 \times 1.377 = 41.31$ euros, 38% above the median.
  3. Mode (the peak): $e^{\mu - \sigma^2} = 30 \times e^{-0.64} = 15.82$ euros.
  4. Chance of an order above 100 euros: $P\!\left(Z \gt \frac{\ln 100 - \ln 30}{0.8}\right) = P(Z \gt 1.505) = 0.066$.
  5. Standard deviation: $\text{mean} \times \sqrt{e^{\sigma^2} - 1} = 41.31 \times 0.947 = 39.1$ euros, so CV $= 0.95$.

$X$ is Log-Normal, $X \sim \text{LogNormal}(\mu, \sigma^2)$, when $\log X \sim N(\mu, \sigma^2)$, i.e. $X = e^Y$ with $Y$ Normal. Its density is

$$p(x) = \frac{1}{x\,\sigma\sqrt{2\pi}}\exp\!\left(-\frac{(\ln x - \mu)^2}{2\sigma^2}\right), \qquad x \gt 0.$$

Important: $\mu$ and $\sigma$ are the mean and sd of $\log X$, not of $X$.

Summary card: Log-Normal

Support$(0, \infty)$
Parameters$\mu$ and $\sigma \gt 0$ of $\log X$. NumPyro dist.LogNormal(loc=μ, scale=σ); NumPy rng.lognormal(mean=μ, sigma=σ); SciPy lognorm(s=σ, scale=exp(μ))
Mean$e^{\mu + \sigma^2/2}$ (median $e^\mu$, mode $e^{\mu - \sigma^2}$)
Variance$(e^{\sigma^2} - 1)\,e^{2\mu + \sigma^2}$; CV $= \sqrt{e^{\sigma^2} - 1}$ (depends on σ only)
Shaperight-skewed, mode < median < mean; long right tail, heavier than a Gamma's far out; symmetric on a log scale
Assumptionspositive, built from many independent multiplicative effects, so that log X is roughly Normal
When usedrevenue per user or order, incomes, prices, file sizes, session durations, latencies, multiplicative errors, priors for positive scales
Relationships$\log X$ is Normal; products of independent Log-Normals are Log-Normal (μ's add, σ²'s add); $X^a$ is Log-Normal; the log is Box-Cox with λ = 0 (4.18)
Why do we need it?

Many business quantities grow by percentages, not by fixed amounts. Their spread is proportional to their size and their right tail is long. The Log-Normal captures that, and taking logs turns it into the familiar Normal world.

Where is it used?

Revenue and basket-size metrics in A/B tests, latency monitoring, income and price data, multiplicative noise in forecasting ("±10%" rather than "±10 units"), log-transformed regression targets, and LogNormal priors on positive scales.

How is it used?

Take logs of the data and check that they look Normal (histogram, Q-Q plot). Fit with mu = np.log(x).mean(), sigma = np.log(x).std(ddof=1). Report the median $e^\mu$ and remember that the mean is $e^{\mu+\sigma^2/2}$, not $e^\mu$.

100 4.605 × 1.10 + 0.095 110 4.700 × 0.95 − 0.051 104.5 4.649 × 1.20 + 0.182 125.4 4.832 take log ↓ value: multiply by random growth factors log of value: add random pieces → a sum → approximately Normal
Logs turn products into sums: $\log(100 \times 1.10 \times 0.95 \times 1.20) = 4.605 + 0.095 - 0.051 + 0.182 = 4.832$, and $e^{4.832} = 125.4$. A sum of many small independent pieces is close to Normal, so the log of a product of many positive factors is close to Normal: the product itself is Log-Normal.

1000 values each start at 100 and are multiplied, step by step, by random factors between (1 − g) and (1 + g). Top: the final values (blue) with the fitted Log-Normal (orange). Bottom: the logs of the same values with a fitted Normal. With 1 step the shape is flat; increase the number of steps and the top becomes right-skewed while the bottom becomes a symmetric bell. Watch the mean pull away from the median as the spread grows.

Set the median $e^\mu$ and the log-scale spread σ. Three markers: pink = mode $e^{\mu-\sigma^2}$, purple = median $e^\mu$, orange = mean $e^{\mu+\sigma^2/2}$. With small σ they nearly coincide; with large σ the mean runs far to the right of the median, and the shaded share of values above the mean drops well below one half.

"In LogNormal(mu, sigma), mu is the mean of the data."

μ is the mean of log X. $e^\mu$ is the median of X, and the mean is larger, $e^{\mu + \sigma^2/2}$. SciPy hides this even more: lognorm(s=σ, scale=e^μ).

"Average the logs and exponentiate: that is the average value."

$\exp(\text{mean of logs})$ is the geometric mean, which estimates the median, not the mean. Back-transforming a forecast made on the log scale gives a median forecast (Chapter 4.18).

"The very large values in my revenue data are outliers to delete."

In Log-Normal-like data, a long right tail is expected. Deleting the big values throws away a large share of the total revenue and biases the mean downwards.

"The Log-Normal has mean μ and variance σ²."

"μ and σ² are the mean and variance of log X. X has median $e^\mu$ and mean $e^{\mu + \sigma^2/2}$, which is larger."

Model answer: "Revenue per order is roughly Log-Normal, because it is built multiplicatively. On the log scale it is Normal with mean μ and sd σ. The typical order (the median) is $e^\mu$, while the average order is $e^{\mu+\sigma^2/2}$, pulled up by the long right tail. If I compare groups on the log scale, I am comparing medians (geometric means), not means."

Revenue-type metrics in an A/B framework like yours are positive and strongly right-skewed, often roughly Log-Normal. If such a metric is modelled with a Normal or Student-t likelihood, the long right tail is part of what the Student-t's heavy tails absorb; modelling log(revenue) is the other common route, but then the effect you estimate is a ratio of medians, not a difference of means. Be explicit about which one your decision $P(\theta_B \gt \theta_A \mid D)$ refers to.

$\log X \sim N(\mu, \sigma^2)$. Median $e^\mu$, mean $e^{\mu+\sigma^2/2}$, mode $e^{\mu-\sigma^2}$: mode < median < mean.

Comes from multiplying many positive factors (logs add → CLT). CV $= \sqrt{e^{\sigma^2}-1}$.

Trap: μ, σ belong to log X. SciPy: lognorm(s=σ, scale=exp(μ)); NumPy: lognormal(mean=μ, sigma=σ).

Quick check: log-latency is $N(\ln 200, 0.5^2)$ (milliseconds). What are the median and the mean latency?

Median $= e^{\ln 200} = 200$ ms. Mean $= 200 \times e^{0.5^2/2} = 200 \times e^{0.125} = 200 \times 1.133 = 226.6$ ms.

The Uniform distribution: "I only know the range"

A spinner with a pointer, numbered from 0 to 10, gives every position the same chance. Nothing below 0, nothing above 10, and no value inside the range is preferred. That is the Uniform distribution: a flat line between two walls.

It plays two roles. As a model, it is the baseline for "I only know the range": rounding errors, a random start time inside an hour, a hyperparameter searched at random between two limits. And it is the raw material of all randomness on a computer: random number generators produce Uniform(0, 1) numbers, and every other distribution is made from them. The standard trick is to push a uniform number backwards through a CDF (inverse-CDF sampling).

Three ways to say it:

  • Picture: a flat table-top between two walls.
  • Numbers: on [0, 10] the mean is 5 and the chance of landing between 2 and 5 is 3/10.
  • Slogan: uniform numbers in, any distribution out (through the inverse CDF).

$U \sim \text{Uniform}(0, 10)$.

  1. Density: $1/(b-a) = 1/10 = 0.1$ everywhere on [0, 10].
  2. $P(2 \lt U \lt 5) = (5-2)/10 = 0.3$: probability is length divided by total length.
  3. Mean: $(0+10)/2 = 5$. Variance: $(b-a)^2/12 = 100/12 = 8.33$, sd $= 2.89$.

Inverse-CDF sampling. The Exponential with rate 0.1 has CDF $F(t) = 1 - e^{-0.1t}$. To turn a uniform number $u = 0.75$ into an Exponential wait:

  1. Solve $F(t) = u$: $1 - e^{-0.1t} = 0.75$, so $e^{-0.1t} = 0.25$.
  2. $t = -\ln(0.25)/0.1 = 1.386/0.1 = 13.86$ minutes.
  3. In general $t = F^{-1}(u) = -\ln(1-u)/\lambda$. Feed in many uniform $u$'s and the $t$'s follow the Exponential exactly.

$U \sim \text{Uniform}(a, b)$ for $a \lt b$ has density $p(x) = \dfrac{1}{b-a}$ for $a \le x \le b$ (and 0 outside), and CDF $F(x) = \dfrac{x-a}{b-a}$ on $[a, b]$.

Inverse-CDF sampling: if $U \sim \text{Uniform}(0, 1)$ and $F$ is any CDF, then $X = F^{-1}(U)$ has CDF $F$. (Reason: $P(F^{-1}(U) \le x) = P(U \le F(x)) = F(x)$.)

Summary card: Uniform

Support$[a, b]$
Parameterslower end $a$, upper end $b$. NumPyro dist.Uniform(low, high); NumPy rng.uniform(low, high); SciPy uniform(loc=a, scale=b−a) (not (a, b)!)
Mean$(a+b)/2$ (also the median)
Variance$(b-a)^2/12$
Shapeflat, with hard edges; no tails at all; no single mode
Assumptionsevery value in the range equally likely; values outside impossible
When usedrandom number generation and inverse-CDF sampling, random hyperparameter search, rounding error, flat priors on bounded parameters (Uniform(0, 1) on a conversion rate), p-values when the null hypothesis is true
RelationshipsUniform(0, 1) = Beta(1, 1) (4.11); $F^{-1}(U) \sim F$; $-\ln(U)/\lambda \sim$ Exponential(λ); the sum of two independent Uniforms is triangular
Why do we need it?

It is the honest model when only the range is known, the simplest distribution to reason about, and the starting point from which computers generate every other distribution.

Where is it used?

Every simulation and Monte Carlo method (NumPy, JAX and NumPyro samplers start from uniform bits), random search for hyperparameters, a Beta(1, 1) prior on a conversion rate (a common default in A/B analysis; check what your code uses), checking p-values (Uniform under H₀, Chapter 5.6) and PIT calibration plots (Chapter 7.16).

How is it used?

Draw with rng.uniform(a, b). To sample a distribution with a known inverse CDF, apply it to uniforms: -np.log(1 - rng.uniform(size=n)) / lam. Before using a Uniform prior, ask whether the hard edges are really impossible boundaries.

Top: the CDF of the target distribution. Each uniform number $u$ (pink dot on the left axis) goes right to the curve and down to $x = F^{-1}(u)$ (purple path). The first draw is the example $u = 0.75$, giving 13.86 minutes. Press Draw 1 a few times, then Draw 500: the histogram of the $x$'s (bottom, blue) fills in the target density (orange), even though every input was uniform. Steep parts of the CDF catch many $u$'s, which is why values pile up there.

"A Uniform prior contains no information."

Flat on one scale is not flat on another: a Uniform(0, 1) prior on a conversion rate $p$ is far from flat on the log-odds $\log(p/(1-p))$ (Chapter 6.2). And its hard edges are strong information: if the truth lies outside $[a, b]$, no amount of data can find it.

"scipy.stats.uniform(2, 5) is Uniform(2, 5)."

SciPy uses loc and scale: uniform(2, 5) is Uniform(2, 7). Write uniform(loc=a, scale=b - a). NumPy's rng.uniform(2, 5) and NumPyro's Uniform(2, 5) do mean [2, 5].

Uniform(a, b): density $1/(b-a)$, mean $(a+b)/2$, variance $(b-a)^2/12$.

Inverse CDF: $F^{-1}(U)$ has CDF $F$; e.g. $-\ln(1-U)/\lambda$ is Exponential(λ).

Trap: SciPy is uniform(loc, scale); "uniform" is not "uninformative".

Quick check: what are the mean and sd of Uniform(0, 1)?

Mean $1/2$; variance $1/12 = 0.0833$; sd $\sqrt{1/12} = 0.289$.

Choosing among Exponential, Gamma and Log-Normal

You have positive, right-skewed data. Exponential, Gamma or Log-Normal? Matching the mean and the sd is not enough: two families with the same mean and sd can still differ a lot in where the peak sits and, above all, in the far right tail, which is where the big orders and the long waits live.

Use three clues. The story: waiting at a constant rate (Exponential), adding stages or a positive parameter (Gamma), multiplying factors (Log-Normal). The CV: an Exponential always has CV = 1. The log view: take logs of the data; a Log-Normal becomes a symmetric bell, while Exponential and Gamma data become skewed to the left.

Three ways to say it:

  • Picture: same centre and spread, different skirts.
  • Numbers: mean 10 and sd 10 for both, yet values above 100 are 16 times more common under the Log-Normal than under the Exponential.
  • Slogan: the story, the CV and the log-histogram choose the family; the tail is where the choice matters.

Positive data with mean 10 and sd 10 (CV = 1). Two candidates with exactly these moments:

  1. Exponential with rate 0.1 (mean 10, sd 10). Equivalently Gamma(1, 0.1).
  2. Log-Normal: from CV, $\sigma^2 = \ln(1 + CV^2) = \ln 2 = 0.693$, $\sigma = 0.833$; from the mean, $\mu = \ln 10 - \sigma^2/2 = 2.303 - 0.347 = 1.956$. Its median is $e^{1.956} = 7.07$.
  3. $P(X \gt 30)$: Exponential $e^{-3} = 0.050$; Log-Normal 0.041. Similar.
  4. $P(X \gt 100)$: Exponential $e^{-10} = 0.000045$; Log-Normal 0.00073, about 16 times more.
  5. A Normal with the same mean and sd would put $\Phi(-1) = 16\%$ of its mass below zero: useless here.

Matching a mean $m$ and an sd $s$ (method of moments):

  • Gamma: $\alpha = (m/s)^2$, $\beta = m/s^2$.
  • Log-Normal: $\sigma^2 = \ln(1 + s^2/m^2)$, $\mu = \ln m - \sigma^2/2$.
  • Exponential: only possible when $s = m$; then $\lambda = 1/m$.

Clues (rules of thumb):

ExponentialGammaLog-Normal
Storywait for one event at a constant ratesum of stages; positive parameterproduct of many factors
CVexactly 1$1/\sqrt\alpha$, any value$\sqrt{e^{\sigma^2}-1}$, any value
Peakat 0at 0 if α ≤ 1, else $(\alpha-1)/\beta$at $e^{\mu-\sigma^2} \gt 0$
Histogram of log(data)skewed leftskewed left (less as α grows)symmetric bell
Far right tailexponentialexponentialheavier (decays slower)

Then confirm with a Q-Q plot (Chapter 4.17) or by comparing fitted log-likelihoods (Chapter 5.2).

Why do we need it?

The tail decides the answers to the questions that matter: the chance of a very large order, the 99th-percentile latency, the cost of a rare long delay. Two models that agree on mean and sd can disagree on these by a factor of 10 or more.

Where is it used?

Choosing a likelihood for revenue or duration metrics in A/B tests, capacity and SLA planning (tail latencies), insurance pricing, and choosing priors for positive parameters.

How is it used?

Write down the story, compute the CV, plot a histogram of log(data), fit the candidates (stats.gamma.fit(x, floc=0), stats.lognorm.fit(x, floc=0)), then compare Q-Q plots and tail probabilities that matter for your decision.

Positive, continuous data Time until an event,constant rate, no memory? yes Exponential A sum of several waits, ora prior for a positive rate/scale? yes Gamma Made by multiplying factors?Is log(data) roughly Normal? yes Log-Normal Only a range is known,everything inside equally likely? yes Uniform no no no Counts?→ 4.8 Proportions?→ 4.11
A first guide for positive continuous data. The story of how the data was made is the strongest clue; then check with the CV (sd/mean), a histogram of log(data) and a Q-Q plot (Chapter 4.17). Counts and proportions have their own families.

All curves have mean 10 and the CV you choose. Orange: Gamma. Teal: Log-Normal. Grey dashed: Normal (with the impossible part below zero in red). At CV = 1 the Gamma is exactly the Exponential. Switch to right tail and compare the curves beyond 20; then to density of log X, where the Log-Normal turns into a symmetric bell and the Gamma does not. Read the tail table below.

A hidden machine draws 300 values with mean 10 from an Exponential, a Gamma (α = 4) or a Log-Normal (σ = 0.9). Look at the histogram, switch to log of the data, read the clues (CV and the skewness of the logs), then press a guess button. Press New mystery sample and play again. Rules of thumb: CV near 1 with strongly left-skewed logs → Exponential; CV near 0.5 → Gamma(4); symmetric logs → Log-Normal.

"Same mean and same sd, so the two models are practically the same."

Their tails can differ by orders of magnitude. If the decision depends on rare large values (capacity, risk, revenue concentration), check the tail directly.

"Pick whichever histogram looks closest."

Histograms of skewed data hide the tail. Look at the log-histogram and a Q-Q plot, and let the data-generating story guide you.

The syllabus's first connecting theme is choosing a likelihood from the support and variance of the data (the full chooser is in Chapter 4.11). For positive continuous metrics (revenue per user, time on site) the candidates are this chapter's families, or a log transform with a Normal or Student-t. For the forecasting model's daily demand counts, the right families are Poisson and Negative Binomial (Chapter 4.8), not these: counts are whole numbers, and the NB is the gamma–Poisson mixture you met above.

Match mean $m$ and sd $s$: Gamma $\alpha = (m/s)^2$, $\beta = m/s^2$; Log-Normal $\sigma^2 = \ln(1 + s^2/m^2)$, $\mu = \ln m - \sigma^2/2$.

Clues: Exponential CV = 1; log(data) symmetric → Log-Normal; Log-Normal's far tail is heavier than Gamma's.

Trap: same mean and sd ≠ same tails. Confirm with log-histograms and Q-Q plots.

Quick check: data has mean 50 and sd 25. Which Gamma and which Log-Normal match it?

CV = 0.5. Gamma: $\alpha = 4$, $\beta = 50/625 = 0.08$. Log-Normal: $\sigma^2 = \ln 1.25 = 0.223$ ($\sigma = 0.472$), $\mu = \ln 50 - 0.112 = 3.800$; its median is $e^{3.800} = 44.7$.

Recap, cheat sheet and practice

  • Positive data lives on $[0, \infty)$ and is usually right-skewed. A Normal leaks $\Phi(-1/CV)$ below zero; use it only when the CV is small.
  • Exponential(λ): the wait for the next event at a constant rate; $P(T \gt t) = e^{-\lambda t}$, mean $1/\lambda$, CV 1, memoryless (constant hazard). Twin of the Poisson count.
  • Gamma(α, β): the sum of α exponential waits; mean α/β, variance α/β², CV $1/\sqrt\alpha$. The standard prior for positive rates and scales, conjugate to the Poisson rate, and the mixing ingredient of the Negative Binomial and the Student-t. Watch the rate versus scale convention.
  • Log-Normal: log X is Normal; it arises from multiplying many factors. μ and σ belong to log X; median $e^\mu$ < mean $e^{\mu+\sigma^2/2}$; the right tail is heavier than a Gamma's.
  • Uniform(a, b): "only the range is known"; mean $(a+b)/2$, variance $(b-a)^2/12$; the raw material of sampling via $F^{-1}(U)$. Not the same as "no information".
  • Choose with the story, the CV, the log-histogram and the tails; equal means and sds do not mean equal tails.
Uniform(0, 1) Exponential(λ) Gamma(α, λ) χ²ₖGamma(k/2, rate ½) LaplaceChapter 4.9 Negative BinomialGamma-mixed Poisson, 4.8 Student-tGamma precision, 4.9 Normal Log-Normal −ln(U)/λ add up α of them(one alone is Gamma(1, λ)) special case difference of two mix rate mix precision exp( · )
How the positive families connect to each other and to the earlier chapters. Uniform numbers become Exponential waits; Exponential waits add up to a Gamma; the Gamma is the mixing ingredient behind the Negative Binomial and the Student-t; exponentiating a Normal gives a Log-Normal. Arrows point from the ingredient to the result.

Cheat sheet

ExponentialGammaLog-NormalUniform
Support$[0, \infty)$$(0, \infty)$$(0, \infty)$$[a, b]$
Parametersrate λshape α, rate βμ, σ of log Xa, b
Density$\lambda e^{-\lambda x}$$\frac{\beta^\alpha}{\Gamma(\alpha)}x^{\alpha-1}e^{-\beta x}$$\frac{1}{x\sigma\sqrt{2\pi}}e^{-(\ln x-\mu)^2/(2\sigma^2)}$$\frac{1}{b-a}$
Mean$1/\lambda$$\alpha/\beta$$e^{\mu+\sigma^2/2}$$(a+b)/2$
Variance$1/\lambda^2$$\alpha/\beta^2$$(e^{\sigma^2}-1)e^{2\mu+\sigma^2}$$(b-a)^2/12$
CV1$1/\sqrt\alpha$$\sqrt{e^{\sigma^2}-1}$$\frac{b-a}{\sqrt3\,(a+b)}$
NumPyroExponential(rate)Gamma(concentration, rate)LogNormal(loc=μ, scale=σ)Uniform(low, high)
SciPyexpon(scale=1/λ)gamma(a=α, scale=1/β)lognorm(s=σ, scale=e^μ)uniform(loc=a, scale=b−a)
NumPyexponential(scale=1/λ)gamma(shape=α, scale=1/β)lognormal(mean=μ, sigma=σ)uniform(low, high)
Code it · Python

import numpy as np
from scipy import stats
rng = np.random.default_rng(1)

# 1) Exponential waiting time: 6 orders per hour = rate 0.1 per minute
lam = 0.1
E = stats.expon(scale=1 / lam)          # SciPy and NumPy take the SCALE = 1/rate
print(E.mean(), round(E.sf(20), 4), round(E.median(), 2))     # 10.0 0.1353 6.93
print(round(E.sf(30) / E.sf(20), 4), round(E.sf(10), 4))      # 0.3679 0.3679  (memoryless)

# 2) Gamma = sum of 3 exponential waits
waits = rng.exponential(scale=1 / lam, size=(100_000, 3)).sum(axis=1)
G = stats.gamma(a=3, scale=1 / lam)     # shape 3, rate 0.1  ->  scale 10
print(round(waits.mean(), 2), G.mean(), round(G.std(), 2))      # ~30 (29.85) 30.0 17.32
print(round(G.sf(60), 4), round(stats.poisson(6).cdf(2), 4))   # 0.062 0.062: 3rd order after 60 min = at most 2 orders in 60 min

# 3) The rate-versus-scale trap (shape 2, rate 4)
print(stats.gamma(a=2, scale=1 / 4).mean(), stats.gamma(a=2, scale=4).mean())   # 0.5 (right) vs 8.0 (wrong)

# 4) A Gamma prior from a mean (20) and a CV (0.5): alpha = 1/CV^2, beta = alpha/mean
a, b = 1 / 0.5**2, 4 / 20
print(a, b, np.round(stats.gamma(a=a, scale=1 / b).ppf([0.05, 0.95]), 1))   # 4.0 0.2 [ 6.8 38.8]
lam_draws = rng.gamma(shape=a, scale=1 / b, size=200_000)      # gamma-Poisson mixture
counts = rng.poisson(lam_draws)
print(round(counts.mean(), 1), round(counts.var(), 0))            # ~20 and ~120 (NB: 20 + 20^2/4)

# 5) Log-Normal: SciPy uses s = sigma and scale = exp(mu)
mu, sigma = np.log(30), 0.8
LN = stats.lognorm(s=sigma, scale=np.exp(mu))
print(round(LN.median(), 2), round(LN.mean(), 2), round(LN.sf(100), 4))   # 30.0 41.31 0.0662
y = rng.lognormal(mean=mu, sigma=sigma, size=200_000)            # NumPy: mean and sigma OF THE LOG
print(round(np.exp(np.log(y).mean()), 1), round(y.mean(), 1))    # ~30 (geometric mean = median), ~41.3 (mean)

# 6) Multiplicative growth -> log-normal (1000 paths, 20 steps of up to +-20%)
v = 100 * np.prod(rng.uniform(0.8, 1.2, size=(1000, 20)), axis=1)
print(round(stats.skew(v), 2), round(stats.skew(np.log(v)), 2))  # about 2 (skewed values) vs about 0 (symmetric logs)

# 7) Uniform: SciPy's uniform(loc, scale) is U(loc, loc + scale)
U = stats.uniform(loc=0, scale=10)
print(U.mean(), round(U.var(), 3))                               # 5.0 8.333
u = rng.uniform(size=200_000)
x = -np.log(1 - u) / lam                 # inverse-CDF sampling of the Exponential
print(round(x.mean(), 2), round(-np.log(1 - 0.75) / lam, 2))     # ~10, and u = 0.75 gives 13.86

# 8) NumPyro uses RATES and log-scale parameters
import numpyro.distributions as dist
print(float(dist.Exponential(rate=0.1).mean), float(dist.Gamma(3.0, rate=0.1).mean),
      round(float(dist.LogNormal(mu, sigma).mean), 2), float(dist.Uniform(0.0, 10.0).mean))
# 10.0 30.0 41.31 5.0
Test yourself

1. Calls arrive at a rate of 0.2 per minute. What is the mean waiting time until the next call?

For an Exponential with rate λ, the mean is $1/\lambda = 1/0.2 = 5$ minutes. (The median is shorter: $\ln 2/0.2 = 3.47$ minutes.)

2. Waiting times are Exponential with mean 10 minutes. You have waited 20 minutes. What is the chance you wait more than 10 further minutes?

Memoryless: $P(T \gt 30 \mid T \gt 20) = e^{-3}/e^{-2} = e^{-1} = P(T \gt 10)$. The time already waited does not matter.

3. What is the mean of scipy.stats.gamma(a=3, scale=2)?

SciPy's scale is $1/\text{rate}$, so the rate is $\beta = 0.5$ and the mean is $\alpha/\beta = 3/0.5 = 6$ (equivalently shape × scale = 3 × 2).

4. $\log X \sim N(\ln 20, 1)$. Which pair is (median of X, mean of X)?

Median $= e^\mu = 20$. Mean $= e^{\mu + \sigma^2/2} = 20 \times e^{0.5} = 32.97$. The mean is above the median because of the long right tail.

5. A Gamma distribution has shape α = 9. What is its coefficient of variation?

CV $= \text{sd}/\text{mean} = (\sqrt\alpha/\beta)/(\alpha/\beta) = 1/\sqrt\alpha = 1/3$. The rate cancels: it only changes the units.

6. Which waiting-time model is not memoryless?

Only the Exponential (a Gamma with shape 1) has a constant hazard. A Gamma with shape 3 has a rising hazard: the longer you have waited, the sooner the end.

Practice problems

A. Support calls arrive at 12 per hour. What is the chance the next call comes within 2 minutes? What are the mean and median waits?

λ = 12/60 = 0.2 per minute. $P(T \le 2) = 1 - e^{-0.2 \times 2} = 1 - e^{-0.4} = 1 - 0.670 = 0.330$. Mean $1/0.2 = 5$ minutes; median $\ln 2/0.2 = 3.47$ minutes.

B. Same calls. What is the distribution of the time until the 5th call, its mean and sd, and the chance it takes more than 30 minutes?

Gamma(α = 5, rate 0.2): mean $5/0.2 = 25$ minutes, sd $\sqrt5/0.2 = 11.2$ minutes. "5th call after 30 minutes" = "at most 4 calls in 30 minutes", and the count is Poisson($0.2 \times 30 = 6$): $P = e^{-6}(1 + 6 + 18 + 36 + 54) = 115e^{-6} = 0.285$.

C. Build a Gamma prior for a noise scale σ with prior mean 2 and CV 0.5. What are α and β, and the central 90% interval?

$\alpha = 1/0.5^2 = 4$, $\beta = \alpha/\text{mean} = 4/2 = 2$. sd $= \sqrt4/2 = 1$. Central 90% (software): 0.68 to 3.88. Since α > 1 the density is 0 at σ = 0: this prior says σ is not tiny (only 1.9% of its mass is below 0.5).

D. Revenue per user is Log-Normal with median 25 euros and mean 40 euros. Find σ, and the share of users who spend more than 100 euros.

mean/median $= e^{\sigma^2/2} = 40/25 = 1.6$, so $\sigma^2 = 2\ln 1.6 = 0.940$ and $\sigma = 0.970$. $P(X \gt 100) = P\!\left(Z \gt \frac{\ln(100/25)}{0.970}\right) = P(Z \gt 1.429) = 0.076$, about 7.6% of users.

E. Using only uniform numbers $u_1 = 0.5$ and $u_2 = 0.9$, construct one draw from Gamma(2, rate 0.1).

Turn each into an Exponential(0.1) wait with $-\ln(1-u)/0.1$: $-\ln(0.5)/0.1 = 6.93$ and $-\ln(0.1)/0.1 = 23.03$. Add the two waits: $6.93 + 23.03 = 29.96$, one draw from Gamma(2, 0.1).

F. Interview: "Revenue per user is very skewed. Would you compare groups on the raw scale or on the log scale, and what does each comparison mean?"

Model answer: "Revenue is positive and roughly Log-Normal, because it is built multiplicatively. On the raw scale I compare means: that is what the business usually cares about (total revenue), but a few huge spenders make it noisy, which is why a heavy-tailed likelihood such as a Student-t can help. On the log scale the data is roughly Normal and better behaved, but a difference of log-means is a ratio of geometric means, roughly a ratio of medians, not a difference of means. Back-transforming with exp gives the median, and the mean needs the extra factor $e^{\sigma^2/2}$. So I choose the scale that matches the decision and I say which quantity my posterior probability refers to."

Chapter 4.11 · Syllabus Module 9

Distributions over probabilities: Beta and Dirichlet, and the distribution map

Your A/B framework does not know the true conversion rate of a variant. It holds a whole curve of beliefs about it. That curve is a Beta distribution, and its many-category cousin is the Dirichlet. This chapter teaches both, then steps back to draw one map that connects every distribution of Chapters 4.7 to 4.11, and ends with a practical tool: how to choose a likelihood from the kind of data you have.

  • Describe an unknown probability with a Beta distribution, and read its two numbers α and β as pseudo-counts (pretend successes and failures)
  • Compute the mean, mode, variance and concentration of a Beta, and recognise its shapes (U, flat, skewed, peaked, J)
  • Describe an unknown probability vector (shares that add up to 1) with a Dirichlet distribution, picture it on a triangle, and know that each share on its own follows a Beta
  • See why "prior pseudo-counts + real counts = posterior" makes Beta and Dirichlet the natural priors for the Binomial and the Multinomial (a preview of Chapter 6.3)
  • Read the distribution map: special cases, sums, limits, mixtures and conjugate pairs
  • Choose a likelihood from the support and the variance structure of your data, for both of your projects

The Beta distribution: a curve of belief about an unknown rate core

You launch a new checkout page. Its true conversion rate $p$ is some number between 0 and 1, but nobody tells you which. Still, you are not completely clueless. A rate near 10% is believable. A rate of 90% is absurd. A rate of exactly 0 is impossible, because someone has already bought.

So instead of one number, you hold a curve over the interval from 0 to 1. Where the curve is tall, that value of $p$ is believable. Where it is low, it is not. The Beta distribution is the standard family of such curves.

Compare it with the Binomial from Chapter 4.7. There, $p$ was known and the number of conversions was random. Here it is the other way round: the rate $p$ itself is the uncertain thing. That is why we call the Beta a distribution over probabilities: the values it describes are themselves probabilities.

Three ways to say it:

  • Picture: a hill drawn over a 1-metre plank; the hill is tall over the rates you find believable.
  • Numbers: Beta(3, 27) puts about 95% of its belief between a 2% and a 23% conversion rate, centred at 10%.
  • Slogan: the Binomial counts successes for a known rate; the Beta describes an unknown rate.

The simplest hump: Beta(2, 2). Its formula is $f(p) = 6\,p\,(1-p)$ for $p$ between 0 and 1. Let us check it is a proper distribution and use it.

  1. Height in the middle: $f(0.5) = 6 \times 0.5 \times 0.5 = 1.5$. A height above 1 is fine: it is a density, not a probability (Chapter 4.4).
  2. Heights at the ends: $f(0) = 6 \times 0 \times 1 = 0$ and $f(1) = 0$. So rates of exactly 0 or 1 are the least believable.
  3. Total area: $\int_0^1 6p - 6p^2\,dp = \big[3p^2 - 2p^3\big]_0^1 = 3 - 2 = 1$. Good: a density must have total area 1.
  4. Probability that the rate is at most 20%: $\big[3p^2 - 2p^3\big]_0^{0.2} = 3(0.04) - 2(0.008) = 0.12 - 0.016 = 0.104$.
  5. So under Beta(2, 2) belief, there is a 10.4% chance that $p \le 0.2$. By symmetry, the mean is exactly 0.5.

A conversion-rate example. Beta(3, 9) has formula $f(p) = 495\,p^2(1-p)^8$. At $p = 0.2$: $495 \times 0.04 \times 0.8^8 = 495 \times 0.04 \times 0.1678 \approx 3.32$. That is its peak. The area between 0.1 and 0.3 is about 0.60 (see the figure below).

A random variable $p$ follows a Beta distribution with parameters $\alpha \gt 0$ and $\beta \gt 0$, written $p \sim \text{Beta}(\alpha, \beta)$, if its density is

$$f(p) = \frac{p^{\alpha-1}(1-p)^{\beta-1}}{B(\alpha, \beta)}, \qquad 0 \lt p \lt 1, \qquad B(\alpha,\beta) = \frac{\Gamma(\alpha)\Gamma(\beta)}{\Gamma(\alpha+\beta)}.$$
  • Support (the set of values it can take): the interval from 0 to 1. Nothing outside it, so it fits probabilities, rates and proportions.
  • Parameters (the knobs that pick one curve from the family): $\alpha$ and $\beta$, two positive numbers called shape parameters. The power $p^{\alpha-1}$ lifts the curve near 1; the power $(1-p)^{\beta-1}$ lifts it near 0.
  • $B(\alpha,\beta)$ is the normalizing constant: the number that makes the total area exactly 1. $\Gamma$ (the Gamma function) is a smooth version of the factorial: $\Gamma(n) = (n-1)!$ for whole numbers, so $B(3, 9) = \frac{2!\,8!}{11!} = \frac{1}{495}$.
  • Mean $\dfrac{\alpha}{\alpha+\beta}$, variance $\dfrac{\alpha\beta}{(\alpha+\beta)^2(\alpha+\beta+1)}$ (the next section explains both).
Beta at a glance
Support$0 \lt p \lt 1$ (a probability)
Parameters$\alpha \gt 0$, $\beta \gt 0$ (pseudo-counts of successes and failures)
Mean · mode · variance$\frac{\alpha}{\alpha+\beta}$ · $\frac{\alpha-1}{\alpha+\beta-2}$ (when $\alpha, \beta \gt 1$) · $\frac{\alpha\beta}{(\alpha+\beta)^2(\alpha+\beta+1)}$
ShapeU, flat, one hump, skewed or J-shaped, depending on $\alpha$ and $\beta$
Used foran unknown conversion rate, click rate or any single probability; a prior for the Binomial; a likelihood for data that are proportions
RelationshipsBeta(1, 1) = Uniform(0, 1); the 2-category Dirichlet; conjugate prior of the Bernoulli and Binomial; $X/(X+Y)$ for independent Gammas
CodeSciPy beta(a, b) · NumPyro dist.Beta(concentration1=α, concentration0=β)
Why do we need it?

Many unknowns in ML and experiments are probabilities: a conversion rate, a click rate, a defect rate. We need a way to say "I am fairly sure it is near 10%, but it could be 5% or 20%", and only a distribution that lives between 0 and 1 can say that honestly.

Where is it used?

The Beta-Binomial model for conversion metrics in a Bayesian A/B test, Thompson sampling in multi-armed bandits, click-through-rate smoothing in ad systems, reliability (failure rates), and Beta regression for data that are proportions.

How is it used?

Pick $\alpha$ and $\beta$ to describe your belief before data (a prior), update them with data (Chapter 6.3), then read the result: its mean, an interval that holds 95% of the area, or the area above a threshold such as $P(p \gt 0.12)$ with beta(a, b).sf(0.12).

0 0.2 0.4 0.6 0.8 1 1 2 3 density (how plausible) mean 0.25 peak (mode) at 0.2, height ≈ 3.3 area ≈ 0.60 = P(0.1 ≤ p ≤ 0.3) The whole area under the curve is exactly 1. A height is a density, not a probability: it can be bigger than 1. The curve lives only on [0, 1], because p is a probability. p = the unknown conversion rate
A Beta(3, 9) curve describing belief about an unknown conversion rate $p$. Tall means plausible. Probabilities are areas: the shaded strip says there is about a 60% chance that $p$ lies between 0.1 and 0.3.

Move the sliders for $\alpha$ and $\beta$ and watch the curve. Drag the two purple handles on the axis: the shaded area is the probability that $p$ lies between them. Press 3 of 10 converted and then Belief near 10% and compare how wide they are. Switch on Show 1 000 random draws: a histogram of random rates drawn from the Beta piles up exactly under the curve.

"The Beta gives the probability of a conversion."

The Beta describes how uncertain we are about the conversion rate. The chance that the next user converts is the Beta's mean (under Beta(3, 9) that is 0.25), but the curve carries much more: how sure we are about that number.

"The peak height 3.3 means $p = 0.2$ has probability 3.3."

For a continuous variable every exact value has probability 0. Heights are densities; only areas are probabilities. A narrow curve must be tall so that its area is still 1.

"Beta and Binomial are the same thing written differently."

The Binomial is a distribution over counts (0, 1, …, n) for a fixed $p$. The Beta is a distribution over the rate $p$ itself. They are partners (one is the prior of the other), not twins.

In an A/B framework like yours, each variant's conversion rate $\theta_A$, $\theta_B$ is not one number but a Beta curve. Every decision quantity is read from these curves: the posterior mean, an interval, and $P(\theta_B \gt \theta_A \mid D)$, which compares two Beta curves (Chapter 6.4 shows how). Check which parameter names your code uses: NumPyro calls them Beta(concentration1, concentration0), where concentration1 is $\alpha$ (the success side) and concentration0 is $\beta$.

$p \sim \text{Beta}(\alpha,\beta)$: density $\propto p^{\alpha-1}(1-p)^{\beta-1}$ on $(0, 1)$; mean $\alpha/(\alpha+\beta)$.

It describes an unknown probability. Heights are densities (can exceed 1); probabilities are areas.

Trap: Binomial = counts for a known $p$; Beta = belief about $p$ itself. NumPyro: Beta(concentration1=α, concentration0=β).

Quick check: under Beta(2, 2), what is $P(p \ge 0.8)$?

Beta(2, 2) is symmetric around 0.5, so $P(p \ge 0.8) = P(p \le 0.2) = 0.104$. You can also compute it: $1 - [3p^2 - 2p^3]_0^{0.8} = 1 - (1.92 - 1.024) = 1 - 0.896 = 0.104$.

Reading α and β: pseudo-counts, mean, mode, variance and concentration core

Think of a restaurant rated 4.5 stars. With 2 reviews, you shrug. With 2 000 reviews, you trust it. The average is the same; the amount of evidence behind it is not.

The Beta works the same way. $\alpha$ behaves like a count of successes and $\beta$ like a count of failures that you pretend you have already seen. They are called pseudo-counts ("pseudo" means pretend). Their split sets where the curve is centred. Their total $\alpha + \beta$ sets how confident, and so how narrow, the curve is. This total is called the concentration.

Three ways to say it:

  • Picture: α green marbles and β red marbles in a bag; the share of green sets the centre, the number of marbles sets the confidence.
  • Numbers: Beta(3, 27) and Beta(30, 270) are both centred at 10%, but the second is about 3 times narrower.
  • Slogan: the ratio picks the centre; the total picks the certainty.

Past checkout pages converted at about 10%. Compare a weak belief, Beta(3, 27), with a strong one, Beta(30, 270).

  1. Concentration: $3 + 27 = 30$ and $30 + 270 = 300$ "pretend users".
  2. Means: $3/30 = 0.1$ and $30/300 = 0.1$. Same centre.
  3. Modes (peaks): $\frac{3-1}{30-2} = \frac{2}{28} \approx 0.071$ and $\frac{30-1}{300-2} = \frac{29}{298} \approx 0.097$. The weak curve is lopsided, so its peak sits left of its mean.
  4. Variances with the short form $\frac{m(1-m)}{\kappa+1}$, where $m$ is the mean and $\kappa = \alpha+\beta$: $\frac{0.1 \times 0.9}{31} = 0.00290$ and $\frac{0.09}{301} = 0.000299$.
  5. Standard deviations: $\sqrt{0.00290} \approx 0.054$ and $\sqrt{0.000299} \approx 0.017$. Ten times the pseudo-counts gives about $\sqrt{301/31} \approx 3.1$ times less spread.

For $p \sim \text{Beta}(\alpha, \beta)$ with concentration $\kappa = \alpha + \beta$:

$$E[p] = \frac{\alpha}{\alpha+\beta}, \qquad \text{mode} = \frac{\alpha-1}{\alpha+\beta-2}\ \ (\alpha, \beta \gt 1), \qquad Var(p) = \frac{\alpha\beta}{(\alpha+\beta)^2(\alpha+\beta+1)} = \frac{m(1-m)}{\kappa+1}.$$
  • Mean–concentration form. Instead of $(\alpha, \beta)$ you can give the mean $m$ and the concentration $\kappa$: $\alpha = m\kappa$ and $\beta = (1-m)\kappa$. Many people find this easier to set.
  • Where the mean comes from. $E[p] = \int_0^1 p\,f(p)\,dp = \frac{B(\alpha+1, \beta)}{B(\alpha, \beta)} = \frac{\Gamma(\alpha+1)\Gamma(\alpha+\beta)}{\Gamma(\alpha)\Gamma(\alpha+\beta+1)} = \frac{\alpha}{\alpha+\beta}$, using $\Gamma(x+1) = x\,\Gamma(x)$.
  • Pseudo-counts, precisely. In the mean and in updating (Section 6 below) $\alpha$ and $\beta$ act like counts that are added to real data. The mode $\frac{\alpha-1}{\alpha+\beta-2}$ is the success share among $\alpha-1$ successes and $\beta-1$ failures, which is why Beta(1, 1) behaves like "no data at all".
  • For a fixed mean $m$, the variance $\frac{m(1-m)}{\kappa+1}$ shrinks like $1/\kappa$: more pseudo-counts, narrower curve. It is always smaller than $m(1-m)$, the variance of a single yes/no outcome with that rate.
Why do we need it?

To turn a vague sentence like "about 10%, but I am not very sure" into two exact numbers a model can use, and to read a fitted Beta back into plain words ("centred at 11.7%, worth about 120 users of evidence").

Where is it used?

Setting Beta priors in Bayesian A/B tests, empirical-Bayes smoothing of click-through rates (add pretend clicks and views to every ad), Thompson sampling, and checking whether a prior is accidentally strong compared with the real sample size.

How is it used?

Choose a centre $m$ and how many users' worth of trust $\kappa$ you want; set $\alpha = m\kappa$, $\beta = (1-m)\kappa$. Compare $\kappa$ with your real number of users: if $\kappa$ is much smaller, the data will dominate.

Beta(3, 27) 3 successes + 27 failures = 30 pretend users, 10% success Beta(30, 270) 30 successes + 270 failures = 300 pretend users, 10% success 0 0.1 0.2 0.3 same mean 0.1 wide: sd ≈ 0.054 narrow: sd ≈ 0.017 p
α and β behave like counts of successes and failures. Both curves have the same split (10% successes), so the same centre. Ten times more pretend users makes the curve about three times narrower ($\sqrt{301/31}\approx 3.1$).

Keep the mean at 0.10 and slide the concentration from about 2 up to 1 000. The curve narrows, and the shaded band that holds the middle 95% of the belief shrinks. Read $\alpha = m\kappa$ and $\beta = (1-m)\kappa$ in the readout. Then try a very small concentration (below 2) and notice the curve stops being a hump: it piles up at 0.

"$\alpha$ is the number of successes I have seen."

$\alpha$ acts like a success count in the formulas. In a prior nobody has seen anything; you are encoding a belief. After updating with data (Section 6), $\alpha$ becomes "prior pseudo-successes + real successes".

"Beta(1, 1) means I saw one success and one failure."

Beta(1, 1) is the flat curve: every rate equally believable, a belief and not a record of anything you saw. For the peak (mode) it counts as no data, which is why the mode formula uses $\alpha - 1$ and $\beta - 1$. For the mean it still adds 2 to the pretend total, so a small sample is pulled a little toward 0.5.

"Small $\alpha$ and $\beta$, like Beta(0.3, 0.3), is the weakest, flattest belief."

Below 1 the curve becomes U-shaped: it says "the rate is either near 0 or near 1". That is a strong and odd statement, not a vague one.

"The mean is the most likely value."

The most likely value is the mode. For a lopsided curve (like Beta(3, 27)) mean, median and mode differ.

When you set a Beta prior for a variant's conversion rate in an A/B framework like yours, read it as "centre $m$, worth $\kappa$ users". If your experiment has 20 000 users per variant, a prior with $\kappa = 30$ hardly matters; a prior with $\kappa = 20\,000$ would be as strong as the whole experiment. Look up what your code actually uses (a flat Beta(1, 1) or something informative) so you can state its strength in users.

"I used a Beta(1, 1) prior, so my prior is uninformative."

Beta(1, 1) is flat on the rate p. Flat is still a choice: the same belief is not flat on the log-odds scale (Chapter 6.2), and with very little data it still pulls the estimate towards 0.5.

Model answer: "I used Beta(1, 1), which is flat on the conversion rate and worth about two pseudo-observations, so with thousands of users per variant the data dominate the posterior."

Mean $\frac{\alpha}{\alpha+\beta}$ · mode $\frac{\alpha-1}{\alpha+\beta-2}$ · Var $= \frac{m(1-m)}{\kappa+1}$ with $\kappa = \alpha+\beta$.

Mean–concentration form: $\alpha = m\kappa$, $\beta = (1-m)\kappa$. Ratio sets the centre, total sets the certainty.

Trap: Beta(1, 1) is flat, a belief and not a record of "1 success + 1 failure" (for the mean it still adds 2 to the pretend total). Parameters below 1 give U-shapes, not vagueness.

Quick check: you want a prior centred at 5% that is worth 40 users. Which Beta?

$\alpha = m\kappa = 0.05 \times 40 = 2$ and $\beta = 0.95 \times 40 = 38$, so Beta(2, 38). Its sd is $\sqrt{0.05 \times 0.95 / 41} \approx 0.034$.

The shapes a Beta can take

Picture two people pulling on a rope tied to a pile of sand on a plank from 0 to 1. $\alpha$ pulls the sand toward the right end (toward $p = 1$). $\beta$ pulls it toward the left end (toward $p = 0$). If both pull hard and equally, the sand forms a tall, narrow pile in the middle. If one pulls harder, the pile slides toward that end.

A value below 1 is special: instead of pulling the pile away from its end, it lets the sand pile up right at that end. Two values below 1 give a "U": sand at both ends, little in the middle.

Three ways to say it:

  • Picture: α tugs the hump to the right, β tugs it to the left; a value below 1 makes its end shoot up.
  • Numbers: Beta(2, 5) peaks at 0.2 with a long tail to the right; Beta(5, 2) is its mirror image, peaking at 0.8.
  • Slogan: both above 1 gives a hump; both below 1 gives a U; exactly 1 and 1 gives flat.

Why Beta(0.5, 0.5) is a U. Its density is $f(p) = \frac{1}{\pi\sqrt{p(1-p)}}$.

  1. In the middle: $f(0.5) = \frac{1}{\pi\sqrt{0.25}} = \frac{1}{\pi \times 0.5} \approx 0.637$.
  2. Near the end: $f(0.1) = \frac{1}{\pi\sqrt{0.09}} = \frac{1}{\pi \times 0.3} \approx 1.061$. Higher than the middle.
  3. As $p \to 0$ the square root goes to 0, so $f(p)$ grows without limit. The curve shoots up at both ends: a U.

Why Beta(2, 5) is skewed right.

  1. Mode: $\frac{2-1}{2+5-2} = \frac{1}{5} = 0.2$.
  2. Mean: $\frac{2}{7} \approx 0.286$.
  3. The mean is to the right of the peak, because a long thin tail stretches toward 1 and pulls the balance point along. A long tail on the right is called skewed right.

The shape of Beta($\alpha$, $\beta$) follows from the two powers in $p^{\alpha-1}(1-p)^{\beta-1}$:

ParametersShapeExample
$\alpha = \beta = 1$flat (the Uniform(0, 1))Beta(1, 1)
$\alpha \lt 1$ and $\beta \lt 1$U-shaped: density goes to infinity at both endsBeta(0.5, 0.5)
$\alpha \le 1 \lt \beta$highest at 0 and falling (J-shaped); infinite at 0 when $\alpha \lt 1$Beta(1, 3), Beta(0.5, 2)
$\beta \le 1 \lt \alpha$highest at 1 and rising toward it (mirror image)Beta(3, 1)
$\alpha \gt 1$ and $\beta \gt 1$one hump, at the mode $\frac{\alpha-1}{\alpha+\beta-2}$Beta(2, 5), Beta(20, 20)
  • Symmetry: $\alpha = \beta$ gives a curve symmetric around 0.5. In general, if $p \sim \text{Beta}(\alpha, \beta)$ then $1 - p \sim \text{Beta}(\beta, \alpha)$ (swap the roles of success and failure).
  • Skew: $\beta \gt \alpha$ leans toward 0 with a long tail toward 1 (skewed right); $\alpha \gt \beta$ is the mirror (skewed left).
  • Width: a bigger total $\alpha + \beta$ gives a narrower hump.
Why do we need it?

To recognise at a glance what a fitted or chosen Beta is saying, and to avoid priors that say something strange (a U-shaped prior claims the rate is near 0 or near 1, which is rarely what anyone means).

Where is it used?

Reading posteriors for rare-event rates (early in an experiment they are skewed right), choosing priors in A/B tests and bandits, and modelling proportions such as "share of a session spent on the page", which can pile up near 0 or 1.

How is it used?

Plot the curve before you use it (beta(a, b).pdf(x)). Check its shape against what you believe. When it is skewed, report the mean or median and an interval, not only the peak.

Beta(0.5, 0.5) · U-shaped 0 0.5 1 both ends pile up Beta(1, 1) · flat 0 0.5 1 every p equally plausible Beta(2, 5) · skewed right 0 0.5 1 hump near 0.2, tail to the right Beta(5, 2) · skewed left 0 0.5 1 mirror image of Beta(2, 5) Beta(20, 20) · peaked 0 0.5 1 symmetric and narrow Beta(1, 3) · J-shaped 0 0.5 1 highest at 0, falls to 0 at 1
The same two-number family makes all of these. Below 1, a parameter makes its end of the curve shoot up (the U-shaped curve really goes to infinity at 0 and 1; it is cut off here). Above 1, it pulls the hump away from that end. Equal parameters give a symmetric curve.

The left square is a map of all Betas: across is $\alpha$, up is $\beta$ (both on a doubling scale, so 1 sits in the middle). Drag the blue dot through the four coloured regions and watch the curve on the right: U-shaped (both below 1), piled up at 0, piled up at 1, one hump (both above 1). Slide along the purple diagonal: the curve stays symmetric and gets narrower as you go up-right.

"A U-shaped Beta is a mistake."

It is a real belief: "for this item the rate is either very low or very high". For example, a feature that most users never touch and a few use every time. It is just rarely a sensible prior for a conversion rate.

"The curve shoots up at 0, so $p = 0$ has a big probability."

The density is infinite at 0 but the area near 0 stays finite, and the probability of exactly 0 is still 0. Only areas are probabilities.

"Skewed right means the hump is on the right."

Skew names the side of the long tail. Beta(2, 5) has its hump on the left (at 0.2) and its long tail on the right, so it is skewed right.

Early in an experiment, a variant with only a handful of conversions has a posterior like Beta(4, 150): a hump close to 0 with a long right tail. Its mean, median and mode differ, so say which one you report. In the A/B framework, summaries such as the posterior mean and an interval from the draws are safer than "the peak".

Both above 1: one hump at $\frac{\alpha-1}{\alpha+\beta-2}$ · both below 1: U · (1, 1): flat · one below/at 1: piles up at that end.

$\alpha = \beta$: symmetric; $\beta \gt \alpha$: skewed right (long tail toward 1). $1-p \sim \text{Beta}(\beta, \alpha)$.

Trap: skew names the tail side, not the hump side.

Quick check: describe Beta(3, 1) without computing anything.

$\beta = 1$ and $\alpha = 3 \gt 1$, so the curve is highest at $p = 1$ and rises toward it. In fact $f(p) = 3p^2$: it starts at 0 at $p = 0$ and reaches 3 at $p = 1$. It is the mirror image of Beta(1, 3).

The Dirichlet distribution: a curve of belief about several shares at once core

Your product sells three plans: Basic, Pro and Enterprise. Each new customer picks one. The unknown thing is now a list of three shares, for example 50% Basic, 30% Pro, 20% Enterprise. The shares are tied together: they must add up to 1, so if one goes up, another must go down.

One Beta cannot describe three tied numbers. The Dirichlet (say "dee-ree-KLAY") can. It is a hill of belief over every possible list of shares. All those lists fit on a triangle: each corner means "everyone picks that plan", and each point inside is one possible mix. The Dirichlet paints that triangle: bright where a mix is believable, dark where it is not.

Three ways to say it:

  • Picture: a heat map painted on a triangle; the hot spot is the most believable mix of shares.
  • Numbers: Dirichlet(5, 3, 2) is centred at the mix (0.5, 0.3, 0.2).
  • Slogan: Beta is for one probability; Dirichlet is for a whole list of probabilities that add up to 1.

Take $\boldsymbol{\alpha} = (5, 3, 2)$ for (Basic, Pro, Enterprise).

  1. Total: $\alpha_0 = 5 + 3 + 2 = 10$.
  2. Mean shares: $(5/10,\ 3/10,\ 2/10) = (0.5,\ 0.3,\ 0.2)$. They add to 1, as they must.
  3. Normalizing constant: $\frac{\Gamma(10)}{\Gamma(5)\Gamma(3)\Gamma(2)} = \frac{9!}{4!\,2!\,1!} = \frac{362\,880}{48} = 7\,560$.
  4. Density at the mix (0.5, 0.3, 0.2): $7\,560 \times 0.5^{4} \times 0.3^{2} \times 0.2^{1} = 7\,560 \times 0.0625 \times 0.09 \times 0.2 = 7\,560 \times 0.001125 \approx 8.5$.
  5. Where it sits on the triangle (Basic at the bottom-left, Pro at the bottom-right, Enterprise at the top): across $= p_{\text{Pro}} + \tfrac12 p_{\text{Ent}} = 0.3 + 0.1 = 0.4$, up $= 0.866 \times p_{\text{Ent}} = 0.173$.

A probability vector is a list $\mathbf{p} = (p_1, \dots, p_K)$ with every $p_k \ge 0$ and $p_1 + \dots + p_K = 1$. The set of all of them is the simplex (for $K = 3$, a triangle; for $K = 2$, a line segment). It has only $K - 1$ free numbers, because the last share is 1 minus the others.

$\mathbf{p} \sim \text{Dirichlet}(\alpha_1, \dots, \alpha_K)$ with every $\alpha_k \gt 0$ and $\alpha_0 = \sum_k \alpha_k$ has density on the simplex

$$f(\mathbf{p}) = \frac{\Gamma(\alpha_0)}{\prod_{k=1}^{K}\Gamma(\alpha_k)}\ \prod_{k=1}^{K} p_k^{\alpha_k - 1}.$$
  • It is the Beta's formula with one power per category. With $K = 2$, Dirichlet($\alpha$, $\beta$) for $(p, 1-p)$ is exactly Beta($\alpha$, $\beta$) for $p$.
  • Mean share of category $k$: $\alpha_k / \alpha_0$.
  • A standard way to draw from it: draw independent $G_k \sim \text{Gamma}(\alpha_k, 1)$ and divide each by their sum: $p_k = G_k / \sum_j G_j$.
Dirichlet at a glance
Supportprobability vectors: $p_k \ge 0$, $\sum p_k = 1$ (the simplex)
Parameters$\alpha_1, \dots, \alpha_K \gt 0$ (one pseudo-count per category); total $\alpha_0$
Mean · variance$\alpha_k/\alpha_0$ · $\frac{\alpha_k(\alpha_0 - \alpha_k)}{\alpha_0^2(\alpha_0+1)}$
Shapea hump inside the triangle when all $\alpha_k \gt 1$; flat when all equal 1; pushed into corners when all are below 1
Used forunknown category shares; prior for the Categorical and Multinomial; topic mixtures
Relationships$K = 2$ gives the Beta; each single share is a Beta; normalized Gammas
CodeSciPy dirichlet([5, 3, 2]) · NumPyro dist.Dirichlet(concentration=jnp.array([5., 3., 2.]))
Why do we need it?

Categorical outcomes (which plan, which star rating, which device) come with several probabilities that must add to 1. We need to describe uncertainty about all of them together without ever breaking the "adds to 1" rule.

Where is it used?

Dirichlet-Multinomial models for categorical metrics in Bayesian A/B tests, the topic proportions of each document in LDA topic models, smoothing category frequencies in language and recommendation models, and Bayesian market-share models.

How is it used?

Give one pseudo-count per category as a prior; after counting customers in each category the posterior is Dirichlet(α + counts) (Section 6). Read posterior mean shares, an interval for each share, or draw many vectors and count how often one variant's share beats another's.

K = 2: every (p₁, p₂) with p₁ + p₂ = 1 p₁p₂ (1, 0)(0, 1) (0.7, 0.3) a line segment: the Beta lives here Basic share = 0.5 Basic (1, 0, 0) Pro (0, 1, 0) Enterprise (0, 0, 1) (0.5, 0.3, 0.2) centre (⅓, ⅓, ⅓) on this edge Enterprise = 0 a triangle:the Dirichletlives here
A simplex is the set of all probability vectors. With 2 categories it is a line segment (home of the Beta). With 3 categories it is a triangle (home of the Dirichlet): each corner is "everyone picks that plan", and a point is closer to a corner when that plan has a bigger share.

The colour shows the density (darker orange = more believable mix); blue dots are 300 random share-vectors drawn from the Dirichlet; the purple dot is the mean. Raise $\alpha_1$ (Basic) and watch the cloud slide toward the Basic corner. Press Confident (20, 12, 8): the same mean as (5, 3, 2) but four times the pseudo-counts, so a much tighter cloud. Press Flat (1, 1, 1): every mix equally likely. Press Sparse (0.3, 0.3, 0.3): the dots run to the corners and edges, so most draws give nearly everything to one or two plans.

"A Dirichlet is just three independent Betas, one per plan."

Independent Betas would not add up to 1. In a Dirichlet the shares are tied: when Basic's share is high, the others must be low, so the shares are negatively correlated.

"The triangle is a 3D picture."

It is flat (2D) because only two of the three shares are free; the third is 1 minus the other two. In general a $K$-category simplex has $K - 1$ dimensions.

"Dirichlet(1, 1, 1) says each plan has a share of one third."

Its mean is (⅓, ⅓, ⅓), but it says every mix is equally plausible, including (0.9, 0.05, 0.05). It is a statement of ignorance, not of equal shares.

Your A/B framework uses a Dirichlet-Multinomial model for categorical metrics. For each variant, the unknown vector of category probabilities gets a Dirichlet prior, the observed category counts follow a Multinomial (Chapter 4.7), and the posterior is again a Dirichlet. A question like "did the mix of plans shift in B?" is answered from draws of these vectors.

$\mathbf{p} \sim \text{Dirichlet}(\boldsymbol{\alpha})$: density $\propto \prod_k p_k^{\alpha_k - 1}$ on the simplex ($p_k \ge 0$, $\sum p_k = 1$); mean $\alpha_k/\alpha_0$.

$K = 2$ is the Beta. Draw by normalizing independent Gammas.

Trap: shares are tied (negatively correlated), not independent Betas; Dirichlet(1, …, 1) is flat, not "equal shares".

Quick check: what is the mean of Dirichlet(4, 4, 2)?

$\alpha_0 = 10$, so the mean is $(4/10, 4/10, 2/10) = (0.4, 0.4, 0.2)$.

Reading a Dirichlet: mean shares, concentration and Beta marginals

Everything you learned about Beta pseudo-counts carries over. Each $\alpha_k$ acts like a pretend count of customers who chose plan $k$. Their split sets the centre of the cloud on the triangle. Their total $\alpha_0$ (the concentration) sets how tight the cloud is.

And there is a lovely shortcut. If you only care about one plan, say "Pro or not Pro?", you have turned the question into yes/no. The share of Pro then follows a plain Beta: Pro's pseudo-count against everybody else's. A single share taken on its own is called a marginal (it is what you see in the "margin" when you ignore the other shares).

Three ways to say it:

  • Picture: the mean sets where the cloud of dots sits; $\alpha_0$ sets how tightly it clusters.
  • Numbers: in Dirichlet(5, 3, 2) the Basic share alone is Beta(5, 5) and the Enterprise share alone is Beta(2, 8).
  • Slogan: look at one category at a time and a Dirichlet is just a Beta.

Dirichlet(5, 3, 2) again, $\alpha_0 = 10$.

  1. Basic alone: Beta($\alpha_1$, $\alpha_0 - \alpha_1$) = Beta(5, 10 − 5) = Beta(5, 5). Mean 0.5, variance $\frac{5 \times 5}{10^2 \times 11} = 0.0227$, sd $\approx 0.151$.
  2. Enterprise alone: Beta(2, 8). Mean 0.2, variance $\frac{2 \times 8}{100 \times 11} = 0.0145$, sd $\approx 0.121$.
  3. Merge Pro and Enterprise into "paid extras": their pseudo-counts add, $3 + 2 = 5$, so (Basic, extras) ~ Dirichlet(5, 5).
  4. How Basic and Pro move together: $Cov = -\frac{5 \times 3}{10^2 \times 11} = -0.0136$, a correlation of about $-0.65$. Negative, because the shares compete for the same total.
  5. Ten times more confidence, Dirichlet(50, 30, 20): same means, Basic's sd drops to $\sqrt{0.5 \times 0.5 / 101} \approx 0.050$, about 3 times smaller.

For $\mathbf{p} \sim \text{Dirichlet}(\boldsymbol{\alpha})$ with $\alpha_0 = \sum_k \alpha_k$ and mean shares $m_k = \alpha_k / \alpha_0$:

$$E[p_k] = m_k, \qquad Var(p_k) = \frac{m_k(1-m_k)}{\alpha_0 + 1}, \qquad Cov(p_i, p_j) = -\frac{m_i m_j}{\alpha_0 + 1}\ \ (i \ne j).$$
  • Marginals are Beta: $p_k \sim \text{Beta}(\alpha_k,\ \alpha_0 - \alpha_k)$.
  • Merging categories adds pseudo-counts: $(p_1 + p_2, p_3, \dots) \sim \text{Dirichlet}(\alpha_1 + \alpha_2, \alpha_3, \dots)$.
  • Mean–concentration form: $\boldsymbol{\alpha} = \alpha_0 \cdot \mathbf{m}$. Bigger $\alpha_0$, tighter cloud (variances shrink like $1/\alpha_0$).
  • Symmetric Dirichlet (all $\alpha_k = a$): $a = 1$ flat; $a \gt 1$ pulls toward the centre (equal shares); $a \lt 1$ pushes toward corners and edges ("sparse" mixes where a few categories take almost everything).
Why do we need it?

To set a Dirichlet prior from plain beliefs ("about half choose Basic, worth 50 customers"), and to answer questions about one category with simple Beta tools instead of the whole triangle.

Where is it used?

Prior design for categorical metrics in A/B tests, reporting an interval for each category share, sparse priors ($a \lt 1$) in LDA topic models so that each document uses few topics, and merging rare categories into "other".

How is it used?

Set $\boldsymbol{\alpha} = \alpha_0 \cdot \mathbf{m}$. For one share use beta(α_k, α₀ − α_k).ppf([0.05, 0.95]). For questions about two shares at once (Pro minus Basic), draw whole vectors with rng.dirichlet(α, size=4000) and compute from the draws.

Drag the purple dot (the mean shares) anywhere in the triangle and set the concentration $\alpha_0$. The blue dots are 400 random share-vectors. On the right, the blue histogram shows one share from those draws, and the orange curve is the Beta($\alpha_k$, $\alpha_0 - \alpha_k$) that theory predicts: they match. Push $\alpha_0$ up to 500 (tight cloud, narrow Beta), then down below 3 with the mean at the centre (dots fly to the corners: sparse).

"A small $\alpha_0$ means a flat, vague Dirichlet."

It is flat only when every $\alpha_k = 1$ (so $\alpha_0 = K$). With equal $\alpha_k$ below 1 the belief runs into the corners: it expects one plan to dominate, which is a strong claim.

"If I know each share's Beta marginal, I know everything."

The marginals miss how the shares move together. For a question about two shares at once (the Pro share minus the Basic share), use whole vectors drawn from the Dirichlet.

In a Dirichlet-Multinomial metric, the posterior for one category's share in one variant is a Beta, so "did the Pro share go up in B?" can be answered with two Beta curves. "Did the whole mix change?" involves all shares together, so it needs joint draws of the vectors (for example from your SVI posterior or from the conjugate Dirichlet).

$E[p_k] = \alpha_k/\alpha_0$, $Var(p_k) = \frac{m_k(1-m_k)}{\alpha_0+1}$, $Cov(p_i, p_j) = -\frac{m_i m_j}{\alpha_0+1}$.

One share alone: $p_k \sim \text{Beta}(\alpha_k, \alpha_0 - \alpha_k)$. Merge categories: add their $\alpha$'s.

Trap: all $\alpha_k = 1$ is flat; all below 1 is sparse (corners), not vague.

Quick check: in Dirichlet(2, 6, 2), how is the Pro share distributed on its own, and what is its mean?

$\alpha_0 = 10$, so Pro ~ Beta(6, 10 − 6) = Beta(6, 4), with mean $6/10 = 0.6$.

A first look at updating: prior pseudo-counts + real counts

Here is why Beta and Dirichlet are everyone's favourite priors. A prior is your belief before seeing data. The posterior is your belief after seeing it (Bayes' theorem from Chapter 4.3 turns one into the other; Chapter 6.1 teaches the whole recipe).

Usually that update needs heavy maths or a computer. But when the data arrive as counts of the same kind as the pseudo-counts (successes and failures; customers per plan), the update is just addition: pour the real marbles into the bag that already holds the pretend ones. A prior that stays in the same family after updating is called conjugate to that kind of data.

Three ways to say it:

  • Picture: real green and red marbles join the pretend ones in the same bag.
  • Numbers: Beta(2, 18) plus 12 conversions out of 100 users gives Beta(14, 106).
  • Slogan: prior pseudo-counts + real counts = posterior pseudo-counts.

Beta and conversions. Prior Beta(2, 18): centre 10%, worth 20 users. Data: 12 of 100 users convert.

  1. Successes onto $\alpha$: $2 + 12 = 14$. Failures onto $\beta$: $18 + (100 - 12) = 18 + 88 = 106$.
  2. Posterior: Beta(14, 106), worth $14 + 106 = 120$ users.
  3. Posterior mean: $14 / 120 \approx 0.1167$.
  4. The same number as a weighted average: the prior mean 0.10 gets weight $20/120 = 1/6$; the data rate $12/100 = 0.12$ gets weight $100/120 = 5/6$. $\tfrac16 \times 0.10 + \tfrac56 \times 0.12 = 0.0167 + 0.1 = 0.1167$. ✓

Dirichlet and plan choices. Prior Dirichlet(1, 1, 1) (flat). Data: 50 Basic, 30 Pro, 20 Enterprise.

  1. Add counts to pseudo-counts: $(1 + 50,\ 1 + 30,\ 1 + 20) = (51, 31, 21)$, total 103.
  2. Posterior mean shares: $(51/103,\ 31/103,\ 21/103) \approx (0.495,\ 0.301,\ 0.204)$. Slightly pulled from (0.5, 0.3, 0.2) toward (⅓, ⅓, ⅓) by the prior.

Beta–Binomial. If $p \sim \text{Beta}(\alpha, \beta)$ and, given $p$, the number of successes in $n$ trials is $k \sim \text{Binomial}(n, p)$, then

$$p \mid k \ \sim\ \text{Beta}(\alpha + k,\ \beta + n - k).$$

Why (in one line). Posterior $\propto$ likelihood $\times$ prior $\propto p^{k}(1-p)^{n-k} \cdot p^{\alpha-1}(1-p)^{\beta-1} = p^{\alpha+k-1}(1-p)^{\beta+n-k-1}$, which is again the Beta shape. (The full derivation, and the predictive distribution, are in Chapter 6.3.)

Dirichlet–Multinomial. If $\mathbf{p} \sim \text{Dirichlet}(\boldsymbol{\alpha})$ and the category counts are $\mathbf{c} \sim \text{Multinomial}(n, \mathbf{p})$, then $\mathbf{p} \mid \mathbf{c} \sim \text{Dirichlet}(\boldsymbol{\alpha} + \mathbf{c})$.

  • Posterior mean as a weighted average: $\frac{\alpha + k}{\alpha + \beta + n} = w \cdot \frac{\alpha}{\alpha+\beta} + (1 - w)\cdot\frac{k}{n}$ with $w = \frac{\alpha + \beta}{\alpha + \beta + n}$. The prior's pull fades as $n$ grows.
  • Conjugate pair (prior family, likelihood): the posterior stays in the prior's family. Others you will meet: Gamma–Poisson (count rates) and Normal–Normal with known variance (Chapter 6.3).
Why do we need it?

Bayesian updating normally needs integrals, MCMC or variational inference. A conjugate pair gives the exact posterior with two additions: instant, transparent, and easy to explain to a product manager.

Where is it used?

Beta-Binomial conversion tests, Thompson sampling bandits (add 1 to α on a click, 1 to β on no click), Dirichlet-Multinomial tests for categorical metrics, streaming dashboards that update counts in real time, and click-rate smoothing.

How is it used?

Store two numbers per variant (α, β), or one per category. After each batch of data, add the counts. Read the posterior with scipy.stats.beta(α, β): its mean, an interval with .ppf([0.025, 0.975]), or a tail area with .sf(threshold).

0 0.2 0.4 Prior Beta(2, 18) centre 10%, worth 20 users 0 0.2 0.4 Posterior Beta(14, 106) centre 14/120 ≈ 0.117 data: 12 of 100 convert α: 2 + 12 = 14 β: 18 + 88 = 106
Updating a Beta with yes/no data is just addition: successes go onto α, failures onto β. The posterior is narrower because it now holds 120 users' worth of information instead of 20. (Both curves use the same vertical scale.)

Set a prior with the sliders (teal, dashed). Press +100 users a few times: each batch is simulated with a true rate of 12%. The blue curve shows what the data alone say (scaled to fit); the orange curve is the posterior, which always sits between the prior and the data and gets narrower with every batch. Try a very strong prior (β = 60, α = 6): it takes many users to move it.

Each group of bars is one plan: teal = prior mean share, blue = share in the data so far, orange = posterior mean share with a 90% interval (from its Beta marginal). Press +50 customers a few times (true shares 50% / 30% / 20%) and watch orange move from teal toward blue while its intervals shrink. Raise the prior pseudo-count to 20 per plan: the posterior clings to equal shares much longer.

"The pseudo-counts are real data, so I can report them as observations."

They encode a belief. Report the prior separately, and compare its total ($\alpha + \beta$) with the real sample size: if the prior is "worth" as many users as the experiment, it is driving the answer.

"Bayesian inference needs conjugate priors."

Conjugacy is a convenience, not a requirement. NumPyro fits non-conjugate models with MCMC or SVI (Chapters 6.9–6.12).

"The posterior mean is just the observed rate $k/n$."

It is a weighted average of the prior mean and $k/n$. With 100 users and a prior worth 20, the data get weight 5/6, not 1.

Both the Beta-Binomial and the Dirichlet-Multinomial models of your A/B framework are conjugate pairs, so their posteriors have exact closed forms. If your framework fits them with SVI, the closed form is a free test: for a single variant without pooling, the SVI posterior mean and spread should come out close to Beta($\alpha + k$, $\beta + n - k$). (With hierarchical pooling across segments the model is no longer this simple pair, which is one reason to use SVI at all.)

"We use a Beta prior because that is what everybody uses."

Give the reasons: the support matches (a rate lives in [0, 1]); it is conjugate to the Binomial, so the posterior is Beta($\alpha + k$, $\beta + n - k$); its parameters read as pseudo-counts, so its strength can be stated in users; and its shapes are flexible.

Model answer: "A conversion rate lives between 0 and 1, so a Beta is the natural prior. It is conjugate to the Binomial likelihood: I add conversions to α and non-conversions to β. Its strength α + β tells me how many users' worth of prior information I am adding."

Beta($\alpha$, $\beta$) + ($k$ of $n$) → Beta($\alpha + k$, $\beta + n - k$). Dirichlet($\boldsymbol\alpha$) + counts $\mathbf{c}$ → Dirichlet($\boldsymbol\alpha + \mathbf{c}$).

Posterior mean $= w \cdot$ prior mean $+ (1-w) \cdot k/n$, with $w = \frac{\alpha+\beta}{\alpha+\beta+n}$.

Trap: conjugacy is convenience, not a requirement; pseudo-counts are belief, not data.

Quick check: flat prior Beta(1, 1), then 3 conversions out of 10 users. What is the posterior and its mean?

Beta(1 + 3, 1 + 7) = Beta(4, 8), with mean $4/12 \approx 0.333$, between the prior mean 0.5 and the data rate 0.3.

The distribution map: how the distributions are related core

Chapters 4.7 to 4.11 introduced fifteen names. Learned one by one they feel like a zoo. They are really one family, and a few kinds of links connect them:

  • Special case: one distribution is another with a particular setting (Beta(1, 1) is the Uniform).
  • Sum: add independent copies (n Bernoullis make a Binomial).
  • Limit: under some condition one becomes almost the other (a Binomial with huge $n$ and tiny $p$ is almost a Poisson).
  • Mixture: a parameter is itself random (a Poisson whose rate changes from day to day becomes a Negative Binomial).
  • Transformation: apply a function (the exponential of a Normal is a Log-Normal).
  • Conjugate pair: a prior and a likelihood that update by simple arithmetic (Beta and Binomial).

Three ways to say it:

  • Picture: a metro map; each line is a rule for turning one distribution into another.
  • Numbers: Binomial(1000, 0.003) gives P(0) = 0.0496 and Poisson(3) gives 0.0498: nearly the same.
  • Slogan: learn the links, and the list of names becomes one family.

A mixture you will use: the Gamma–Poisson. Daily orders are Poisson with rate $\lambda$, but $\lambda$ itself changes from day to day: $\lambda \sim \text{Gamma}(\alpha = 2, \text{rate} = 0.4)$, which has mean $2/0.4 = 5$ and variance $2/0.4^2 = 12.5$ (Chapter 4.10).

  1. Mean (law of total expectation, Chapter 4.6): $E[Y] = E\big[E[Y\mid\lambda]\big] = E[\lambda] = 5$.
  2. Variance (law of total variance): $Var(Y) = E\big[Var(Y\mid\lambda)\big] + Var\big(E[Y\mid\lambda]\big) = E[\lambda] + Var(\lambda) = 5 + 12.5 = 17.5$.
  3. The Negative Binomial NB2 with mean $\mu = 5$ and dispersion $\alpha = 2$ has $Var = \mu + \mu^2/\alpha = 5 + 25/2 = 17.5$. Same mean, same variance; in fact the whole distribution is the same (Chapter 4.8).
  4. Zero days: NB gives $P(0) = \big(\tfrac{\alpha}{\alpha + \mu}\big)^{\alpha} = (2/7)^2 \approx 0.082$, while a plain Poisson(5) gives $e^{-5} \approx 0.0067$: about 12 times more zero days, purely because the rate wobbles.

A limit. Binomial(1000, 0.003): $P(0) = 0.997^{1000} \approx 0.0496$. Poisson(3): $P(0) = e^{-3} \approx 0.0498$.

The map in words (every arrow of the figure, plus a few more). "iid" means independent and identically distributed.

FromBecomesHowKind
Bernoulli($p$)Binomial($n$, $p$)sum of $n$ iid copiessum
Categorical($\mathbf{p}$)Multinomial($n$, $\mathbf{p}$)category counts of $n$ iid drawssum
Categorical / MultinomialBernoulli / Binomial$K = 2$ categoriesspecial case
Binomial($n$, $\lambda/n$)Poisson($\lambda$)$n \to \infty$ with $np = \lambda$ fixedlimit
Binomial($n$, $p$)Normal($np$, $np(1-p)$)$n$ large, $np$ and $n(1-p)$ not smalllimit
Poisson($\lambda$)Normal($\lambda$, $\lambda$)$\lambda$ largelimit
Poisson($\lambda_1$) + Poisson($\lambda_2$)Poisson($\lambda_1 + \lambda_2$)independent sumsum
Poisson($\lambda$), $\lambda \sim$ Gamma($\alpha$, rate $\alpha/\mu$)NB2($\mu$, $\alpha$)the rate is randommixture
NB2($\mu$, $\alpha$)Poisson($\mu$)$\alpha \to \infty$limit
Exponential($\lambda$)Gamma($k$, $\lambda$)sum of $k$ iid copies; Exponential = Gamma(1, $\lambda$)sum
$\chi^2_k$Gamma($k/2$, rate $1/2$)sum of $k$ squared standard Normalsspecial case
$G_k \sim$ Gamma($\alpha_k$, 1) independentDirichlet($\boldsymbol\alpha$)$p_k = G_k / \sum_j G_j$ (for $K = 2$: $X/(X+Y)$ is Beta)transformation
Dirichlet($\boldsymbol\alpha$)Beta($\alpha_k$, $\alpha_0 - \alpha_k$)one share alone; $K = 2$ is the Betaspecial case
Beta(1, 1)Uniform(0, 1)flat Betaspecial case
Normal($\mu$, $\sigma^2$)Log-Normal($\mu$, $\sigma$)$e^{X}$transformation
Normal + Normal (independent)Normalmeans add, variances addsum
Normal(0, $1/\tau$), $\tau \sim$ Gamma($\nu/2$, rate $\nu/2$)Student-t($\nu$)the precision is randommixture
Student-t($\nu$)Normal$\nu \to \infty$; with $\nu = 1$ it is the Cauchylimit
Normal(0, $V$), $V \sim$ Exponential(mean $2b^2$)Laplace(0, $b$)the variance is randommixture
$E_1 - E_2$, $E_i$ iid Exponential(mean $b$)Laplace(0, $b$)differencetransformation

Conjugate pairs (prior ↔ likelihood): Beta ↔ Bernoulli/Binomial; Dirichlet ↔ Categorical/Multinomial; Gamma ↔ Poisson (and Exponential) rates; Normal ↔ Normal mean when the variance is known.

Why do we need it?

The links let you derive properties instead of memorising them, explain why a model behaves as it does (why NB has extra zeros, why Student-t tolerates outliers), and know where an approximation stops working.

Where is it used?

Negative Binomial count models in demand forecasting (Gamma–Poisson), robust regression with Student-t errors (Normal scale mixture), Beta and Dirichlet samplers built from Gammas, Poisson approximations for rare events, and conjugate priors in A/B tests and bandits.

How is it used?

When a model fails a check, walk one link on the map: Poisson too narrow → Negative Binomial; Normal tails too light → Student-t; need a prior for a rate → its conjugate partner. Verify a link by simulation, exactly as the explorer below does.

sum of n sum of n K = 2 K = 2 K = 2 marginals Beta(1, 1) n→∞, p→0, np = λ λ large α → ∞ Poisson with a Gamma rate sum of k normalize K Gammas ν → ∞ Normal with Gamma precision difference of two exp( Normal ) n large (Normal approximation) conjugate conjugate conjugate Bernoulli Binomial Poisson Normal Categorical Multinomial Neg. Binomial Student-t Beta Dirichlet Gamma Laplace Uniform Exponential Log-Normal sum, special case or transform (exact) limit (approximately equal when …) mixture (a parameter is itself random) conjugate pair (prior ↔ likelihood) counts and categories real or positive numbers distributions over probabilities
The distribution map. Read an arrow as "becomes": a Bernoulli summed over $n$ trials becomes a Binomial; a Binomial with many trials and a tiny success chance becomes (approximately) a Poisson; a Poisson whose rate is itself Gamma-distributed becomes a Negative Binomial. Purple links join a prior to the likelihood it is conjugate to. The table below lists every link in words.

Pick a relationship. Blue is the starting construction (simulated draws or the exact distribution); orange is what the map says it becomes. Move the slider. For a limit, the two only meet when the slider goes far enough (for example $n$ large for Binomial → Poisson). For a sum, mixture or transformation, they match at every setting, up to the wobble of random sampling: press New sample to see it.

All the distributions of Chapters 4.7–4.11 at a glance

DistributionSupportParametersMeanVarianceTypical useNumPyro
Bernoulli{0, 1}$p$$p$$p(1-p)$one conversionBernoulli(probs=p)
Binomial0, …, $n$$n$, $p$$np$$np(1-p)$conversions out of $n$Binomial(total_count=n, probs=p)
Categorical{1, …, $K$}$\mathbf{p}$(category probabilities $p_k$)one choice among $K$Categorical(probs=p)
Multinomialcounts adding to $n$$n$, $\mathbf{p}$$np_k$$np_k(1-p_k)$counts per categoryMultinomial(total_count=n, probs=p)
Poisson0, 1, 2, …$\lambda$$\lambda$$\lambda$events per intervalPoisson(rate=λ)
Negative Binomial (NB2)0, 1, 2, …$\mu$, $\alpha$$\mu$$\mu + \mu^2/\alpha$overdispersed countsNegativeBinomial2(mean=μ, concentration=α)
Normalall reals$\mu$, $\sigma$$\mu$$\sigma^2$symmetric noiseNormal(loc=μ, scale=σ)
Student-tall reals$\nu$, $\mu$, $\sigma$$\mu$ ($\nu \gt 1$)$\sigma^2\frac{\nu}{\nu-2}$ ($\nu \gt 2$)noise with outliersStudentT(df=ν, loc=μ, scale=σ)
Laplaceall reals$\mu$, $b$$\mu$$2b^2$sharp-peaked noise; sparsity priorLaplace(loc=μ, scale=b)
Exponential$x \ge 0$rate $\lambda$$1/\lambda$$1/\lambda^2$waiting timesExponential(rate=λ)
Gamma$x \gt 0$shape $\alpha$, rate $\beta$$\alpha/\beta$$\alpha/\beta^2$positive amounts; prior for rates and scalesGamma(concentration=α, rate=β)
Log-Normal$x \gt 0$$\mu$, $\sigma$ (of $\log x$)$e^{\mu + \sigma^2/2}$$(e^{\sigma^2}-1)e^{2\mu+\sigma^2}$multiplicative, right-skewed amountsLogNormal(loc=μ, scale=σ)
Uniform$[a, b]$$a$, $b$$\frac{a+b}{2}$$\frac{(b-a)^2}{12}$"anything in this range"Uniform(low=a, high=b)
Beta$(0, 1)$$\alpha$, $\beta$$\frac{\alpha}{\alpha+\beta}$$\frac{\alpha\beta}{(\alpha+\beta)^2(\alpha+\beta+1)}$an unknown probabilityBeta(concentration1=α, concentration0=β)
Dirichletprobability vectors$\boldsymbol{\alpha}$$\frac{\alpha_k}{\alpha_0}$$\frac{\alpha_k(\alpha_0-\alpha_k)}{\alpha_0^2(\alpha_0+1)}$unknown category sharesDirichlet(concentration=α)

SciPy differs in places: gamma(a, scale=1/β), nbinom(n=α, p=α/(α+μ)), lognorm(s=σ, scale=e^μ), expon(scale=1/λ). NumPyro and SciPy both take the standard deviation (scale), not the variance, for the Normal.

"Binomial → Poisson means they are the same distribution."

A limit is an approximation that gets better under a condition ($n$ large, $p$ small). With $n = 5$ and $p = 0.6$ the Poisson is a poor stand-in. Check the condition before you use a limit.

"A Normal whose variance is random is still a Normal."

Mixing Normals with different spreads gives heavier tails and a sharper peak: Student-t or Laplace. That is exactly why these likelihoods tolerate outliers.

"Normal approximations are always fine for counts if the mean is large enough."

Often fine for large counts far from zero, but a Normal can put probability on negative counts, and it keeps the variance free of the mean, unlike Poisson or NB.

Your likelihoods sit on this map. In the forecasting model, a Negative Binomial for daily demand is a Poisson whose rate wobbles from day to day (Gamma mixture), which is why it fits overdispersed counts and extra zero days. A Student-t likelihood is a Normal whose scale wobbles from point to point, which is why one extreme day pulls the fit less. In the A/B framework, Beta-Binomial and Dirichlet-Multinomial are the two conjugate pairs of the left side of the map.

Six links: special case, sum, limit, mixture, transformation, conjugate pair.

Gamma–Poisson mixture = NB2 ($Var = \mu + \mu^2/\alpha$); Normal with Gamma precision = Student-t; Normal with Exponential variance = Laplace.

Trap: limits are approximations with conditions; mixtures of Normals are not Normal.

Quick check: daily rate $\lambda \sim$ Gamma(shape 4, rate 0.8) and orders $Y \mid \lambda \sim$ Poisson($\lambda$). What are $E[Y]$ and $Var(Y)$?

$E[\lambda] = 4/0.8 = 5$ and $Var(\lambda) = 4/0.64 = 6.25$. So $E[Y] = 5$ and $Var(Y) = E[\lambda] + Var(\lambda) = 5 + 6.25 = 11.25$. Check with NB2: $\mu = 5$, $\alpha = 4$: $5 + 25/4 = 11.25$. ✓

Choosing a likelihood from support and variance core

A likelihood is the distribution you assume produced each observation, once the model's parameters (and inputs) are fixed. It is the part of a model that says "this is what the data look like around the prediction" (Chapter 5.2 and Chapter 6.1 use it in full).

Choosing one is like choosing a container. You would not carry soup in a paper bag or nails in a sieve. First ask what values one observation can take: only 0 or 1? Whole counts with no top? Any real number? Only positive amounts? This set of possible values is the support, and the likelihood must have the same support. Then ask how the spread behaves: does the variance grow with the mean? Are there occasional huge values? That is the variance structure. Two questions, and most of the choice is made.

Three ways to say it:

  • Picture: match the shape of the container (the support) first, then its stretchiness (the variance).
  • Numbers: daily orders with mean 20 and variance 70 are counts with variance 3.5 times the mean: a Negative Binomial, not a Poisson.
  • Slogan: support first, variance second, check third.

Daily orders at one store. Over 60 days the counts have mean 20 and variance 70.

  1. Support: whole numbers 0, 1, 2, … with no fixed upper limit. Candidates: Poisson or Negative Binomial (not Normal, which allows negative and fractional values; not Binomial, which needs a known maximum $n$).
  2. Variance structure: a Poisson forces variance = mean = 20. The data show 70. The dispersion ratio is $70 / 20 = 3.5$, far above 1: the counts are overdispersed (Chapter 4.8).
  3. Fit an NB2 by matching mean and variance (the "method of moments"): $\mu + \mu^2/\alpha = 70$ gives $\alpha = \frac{\mu^2}{Var - \mu} = \frac{400}{50} = 8$.
  4. Why it matters: the chance of a busy day with at least 35 orders is $0.0015$ under Poisson(20) but $0.057$ under NB2(20, 8), about 38 times more. Staffing based on the Poisson would be badly surprised.
  5. Check: after fitting, compare predicted and observed variance, zero days and big days (posterior predictive checks, Chapter 6.8).

Three quick ones. Did this user convert? Support {0, 1}: Bernoulli. Seconds spent on a page? Positive and right-skewed: Log-Normal or Gamma. Forecast errors with a few huge misses? Real numbers with heavy tails: Student-t.

The support of a distribution is the set of values it gives positive probability (or density) to. The variance structure is how the variance relates to the mean (and how heavy the tails are). A recipe:

  1. Support. Match the possible values of one observation: {0, 1}; 0…$n$; 0, 1, 2, …; one of $K$ labels; all reals; positive reals; (0, 1); probability vectors.
  2. Variance structure. Compare the candidates' variance rules with the data (looked at within groups or around the model's prediction, not across everything at once):
LikelihoodSupportVariance as a function of the mean
Bernoulli / Binomial{0, 1} / 0…$n$$p(1-p)$ / $np(1-p)$: fixed by the mean, below it
Poisson0, 1, 2, …Var = mean
Negative Binomial (NB2)0, 1, 2, …Var = $\mu + \mu^2/\alpha$, above the mean
Normalall realsVar = $\sigma^2$, not tied to the mean; light tails
Student-tall reals$\sigma^2\nu/(\nu-2)$; heavy tails (outliers expected)
Laplaceall reals$2b^2$; sharp peak, exponential tails
Gamma / Log-Normalpositive realsVar $\propto$ mean² (the spread grows with the level)
Beta(0, 1)$m(1-m)/(\kappa+1)$: shrinks near 0 and 1
Categorical / Multinomial / Dirichletlabels / counts / shares$np_k(1-p_k)$ per count / $\frac{m_k(1-m_k)}{\alpha_0+1}$ per share
  1. Tails and zeros. Outliers → Student-t. Far more zeros than even an NB predicts → a zero-inflated model (an extra group that is always zero, mixed with the count distribution) or a hurdle model (first decide zero or not, then draw a positive count); see Chapter 7.13.
  2. Check. Fit, then compare model and data: residual Q-Q plot (Chapter 4.17), variance vs mean, predicted vs observed zeros and extremes.
Why do we need it?

A likelihood with the wrong support predicts impossible values (negative orders, a 120% rate). One with the wrong variance gives intervals that are far too narrow or too wide, and decisions (stock, launch or not) that are wrong with confidence.

Where is it used?

Every probabilistic model: the metric models of a Bayesian A/B framework, the observation model of a forecasting system, GLMs (Chapter 5.14), the loss of a neural network (squared error ↔ Normal, cross-entropy ↔ Bernoulli/Categorical).

How is it used?

Write down what one observation can be, compute the dispersion ratio (variance / mean) for counts or look at a residual Q-Q plot for real values, pick the candidate, fit it, then run predictive checks. In NumPyro it is the dist.… inside numpyro.sample("y", …, obs=y).

1. What can ONE observation be? 2. How does the spread behave? → candidates 0 or 1 (yes / no) Bernoulli · n users together: Binomial k successes out of a known n Binomial · extra spread between groups: Beta-Binomial a count 0, 1, 2, … (no top) Var ≈ mean: Poisson · Var > mean: Neg. Binomial many extra zeros: zero-inflated / hurdle one of K labels / K counts Categorical / Multinomial any real number light tails: Normal · outliers: Student-t · sharp peak: Laplace a positive amount Gamma · big right skew: Log-Normal · waits: Exponential a share in (0, 1) / K shares Beta / Dirichlet 3. Fit it, then CHECK: residual Q-Q plot, variance vs mean, posterior predictive checks
Choosing a likelihood in three questions. First the support (the values one observation can take), then the variance structure (how the spread behaves), then a check against the data. Orange names are the usual candidates.

Pick what one observation can be. For counts, real numbers and positive amounts, a second row of buttons asks how the spread behaves. The plot shows the candidates with typical parameters (orange = the suggested one, grey dashed = the alternatives), and the readout gives the NumPyro call, the reason, and what to check. For real numbers, tick log scale: the tails, invisible on the normal scale, suddenly show which likelihood expects outliers.

Set the mean and the variance of your counts (as a ratio Var / mean). At a ratio near 1 the Poisson (blue) is enough. Raise the ratio: the NB2 fitted by matching mean and variance (orange) spreads out, puts more mass on slow days and busy days, and the readout shows how much more often a busy day (75% above the mean, the red line) happens. Press Daily orders for the worked example (mean 20, variance 70).

"Use a Normal for everything: the Central Limit Theorem makes everything Normal."

The CLT is about averages of many observations (Chapter 4.13), not about each observation. A single day's orders near zero, or a single conversion, is not Normal.

"I measured the variance of my raw daily demand, so I know the likelihood's dispersion."

The likelihood describes $y$ given the model's prediction. Raw demand also varies because of trend, weekly seasonality and holidays. Judge the dispersion within comparable days or on residuals, otherwise you will overstate it.

"More zeros than a Poisson predicts, so I need a zero-inflated model."

An NB already predicts many more zeros than a Poisson with the same mean. Try it first; add zero inflation only if NB still falls short.

"Student-t removes the outliers."

It keeps them, but expects them: extreme residuals are given more probability, so they pull the fit less (Chapter 7.13).

This is item 1 of the topics that connect both of your projects. A/B framework: conversions (support {0, 1}, or $k$ of $n$) → Bernoulli/Binomial with a Beta prior; categorical metrics → Multinomial with a Dirichlet prior; count metrics → Poisson; continuous metrics → Normal, or Student-t when outliers appear. Forecasting model: continuous, well-behaved demand → Normal; demand residuals with occasional extreme days → Student-t; count demand whose variance exceeds its mean → Negative Binomial. In every case the reason is the same pair: support, then variance structure.

"I used a Negative Binomial because it is the standard for counts."

Give the reasoning chain: the support (non-negative integers with no upper bound), the variance structure (variance well above the mean, so Poisson is too narrow), and the check you ran afterwards.

Model answer: "Demand is a count, so I needed support on 0, 1, 2, …. Within comparable days the variance was several times the mean, which a Poisson cannot produce because it forces them equal. NB2 adds a dispersion parameter, Var = μ + μ²/α, and reduces to the Poisson as α grows. Posterior predictive checks of variance and zero days then looked right."

Choose a likelihood: 1. support (what values can one observation take?) → 2. variance structure (Var vs mean, tails, zeros) → 3. check (Q-Q, dispersion, predictive checks).

Counts: dispersion ratio Var/mean ≈ 1 → Poisson; > 1 → NB2 with $\alpha \approx \mu^2/(Var - \mu)$.

Trap: judge the variance given the model (within groups or on residuals), not on raw data.

Quick check: support tickets per day have mean 8 and variance 30. Which likelihood, and what α would you start from?

Counts with no upper limit and variance far above the mean (ratio 3.75): Negative Binomial NB2. Matching moments: $\alpha = 8^2 / (30 - 8) = 64/22 \approx 2.9$.

Recap, cheat sheet and practice

  • The Beta($\alpha$, $\beta$) describes an unknown probability. Mean $\frac{\alpha}{\alpha+\beta}$; $\alpha, \beta$ act as pseudo-counts; their total $\kappa = \alpha + \beta$ (the concentration) sets how narrow it is.
  • Its shapes: both above 1 → one hump; both below 1 → U; (1, 1) → flat; one at or below 1 → piles up at that end. Skew names the tail side.
  • The Dirichlet($\boldsymbol\alpha$) describes an unknown probability vector on the simplex (a triangle for 3 categories). Mean $\alpha_k/\alpha_0$; each share alone is Beta($\alpha_k$, $\alpha_0 - \alpha_k$); shares are negatively correlated.
  • Updating is addition: Beta($\alpha + k$, $\beta + n - k$) and Dirichlet($\boldsymbol\alpha + \mathbf{c}$). The posterior mean is a weighted average of prior mean and data rate (full story in Chapter 6.3).
  • The distribution map links everything through special cases, sums, limits, mixtures, transformations and conjugate pairs. Gamma–Poisson = Negative Binomial; Normal with random scale = Student-t or Laplace.
  • To choose a likelihood: support first, variance structure second, check third.

Cheat sheet

IdeaFormulaIn words
Beta density$\frac{p^{\alpha-1}(1-p)^{\beta-1}}{B(\alpha,\beta)}$, $0 \lt p \lt 1$a hill of belief over a rate
Beta mean · mode · variance$\frac{\alpha}{\alpha+\beta}$ · $\frac{\alpha-1}{\alpha+\beta-2}$ · $\frac{m(1-m)}{\kappa+1}$ratio = centre, total = certainty
Mean–concentration$\alpha = m\kappa$, $\beta = (1-m)\kappa$"centre $m$, worth $\kappa$ users"
Beta–Binomial updateBeta($\alpha + k$, $\beta + n - k$)add successes and failures
Dirichlet$\propto \prod p_k^{\alpha_k-1}$; mean $\alpha_k/\alpha_0$a hill over share-vectors
Dirichlet marginal$p_k \sim$ Beta($\alpha_k$, $\alpha_0 - \alpha_k$)one share alone is a Beta
Dirichlet–Multinomial updateDirichlet($\boldsymbol\alpha + \mathbf{c}$)add customers per category
Gamma–Poisson mixtureNB2($\mu$, $\alpha$): $Var = \mu + \mu^2/\alpha$a wobbling rate makes extra spread
Count dispersion checkratio $= Var/\text{mean}$; $\hat\alpha = \mu^2/(Var - \mu)$≈ 1 Poisson, > 1 NB
Choosing a likelihoodsupport → variance structure → checkcontainer shape, then stretch, then test
Code it · Python
import numpy as np
from scipy import stats

# --- Beta: mean, mode, sd, areas and intervals ---
a, b = 3, 9
B = stats.beta(a, b)
print(B.mean(), (a - 1) / (a + b - 2), B.std())    # 0.25 0.2 0.1201   (mean, mode, sd)
print(B.cdf(0.3) - B.cdf(0.1))                     # 0.5977              P(0.1 <= p <= 0.3)
print(B.ppf([0.025, 0.975]))                       # [0.0602 0.5178]     the middle 95% of the belief

# --- mean-concentration form: centre 10%, worth 30 users ---
m, kappa = 0.10, 30
print(stats.beta(m * kappa, (1 - m) * kappa).std())   # 0.0539  (sd of Beta(3, 27))

# --- conjugate update: prior Beta(2, 18) + 12 conversions out of 100 ---
post = stats.beta(2 + 12, 18 + 88)
print(post.mean(), post.ppf([0.025, 0.975]))       # 0.1167 [0.0658 0.1796]

# --- Dirichlet: random share-vectors, one share alone is a Beta ---
alpha = np.array([5, 3, 2])
rng = np.random.default_rng(0)
P = rng.dirichlet(alpha, size=100_000)             # every row adds up to 1
print(P.mean(axis=0))                              # about [0.50 0.30 0.20]  = alpha / alpha.sum()
print(P[:, 0].std(), stats.beta(5, 5).std())       # both about 0.1508: Basic alone is Beta(5, 5)
print(np.corrcoef(P[:, 0], P[:, 1])[0, 1])         # about -0.65: the shares compete for the same total

# --- Dirichlet update: flat prior + counts (50, 30, 20) ---
counts = np.array([50, 30, 20])
print((1 + counts) / (3 + counts.sum()))           # [0.4951 0.301  0.2039]

# --- a link of the map by simulation: Gamma-Poisson = NB2(mu=5, alpha=2) ---
lam = rng.gamma(shape=2, scale=5 / 2, size=200_000)   # NumPy's gamma takes scale = 1/rate
y = rng.poisson(lam)
print(y.mean(), y.var())                           # about 5.0 and 17.5  (= 5 + 25/2)
print((y == 0).mean(), (2 / 7) ** 2)               # about 0.081 vs 0.0816: NB2 zero probability (alpha/(alpha+mu))^alpha

# --- choosing a count likelihood: dispersion ratio and NB2 by moments ---
mean, var = 20, 70
alpha_hat = mean**2 / (var - mean)
nb = stats.nbinom(alpha_hat, alpha_hat / (alpha_hat + mean))   # SciPy's NB(n, p) form
print(var / mean, alpha_hat, nb.var())             # 3.5 8.0 70.0
print(stats.poisson(mean).sf(34), nb.sf(34))       # 0.0015 0.057   P(Y >= 35): about 38 times more likely under NB2

# --- the same objects in NumPyro: note the parameter names ---
import jax.numpy as jnp
import numpyro.distributions as dist
print(dist.Beta(concentration1=3.0, concentration0=9.0).mean)          # 0.25
print(dist.Dirichlet(concentration=jnp.array([5.0, 3.0, 2.0])).mean)   # [0.5 0.3 0.2]
print(dist.NegativeBinomial2(mean=20.0, concentration=8.0).variance)   # 70.0
Test yourself

1. What is the mean of Beta(6, 14)?

Mean $= \alpha/(\alpha+\beta) = 6/20 = 0.30$. (The mode is $5/18 \approx 0.28$. The value 0.43 comes from $6/14$, a common slip: the denominator must be $\alpha + \beta$.)

2. Which Beta is U-shaped (piles up at both 0 and 1)?

Both parameters below 1 make both ends shoot up. Beta(1, 1) is flat, Beta(2, 2) is a hump, Beta(5, 1) piles up at 1 only.

3. Prior Beta(3, 7). Then 4 of 20 users convert. What is the posterior?

Add successes to α and failures to β: $3 + 4 = 7$ and $7 + (20 - 4) = 23$.

4. In Dirichlet(6, 3, 1), the first share on its own follows…

A single share is Beta($\alpha_k$, $\alpha_0 - \alpha_k$) $=$ Beta(6, 10 − 6) $=$ Beta(6, 4), with mean 0.6.

5. Daily counts have mean 12 and variance 40 (within comparable days). The best first likelihood is…

Counts with no upper limit and variance well above the mean (ratio 3.3): overdispersed, so NB2. A Poisson forces variance = mean; a Binomial needs a known maximum and has variance below the mean.

6. A Poisson count whose rate $\lambda$ is itself Gamma-distributed is…

This Gamma–Poisson mixture is exactly the NB2. The random rate adds its own variance: $Var = \mu + \mu^2/\alpha$.

Practice problems

A. You believe a refund rate is "about 4%, trusted like 50 orders". Which Beta prior, and what is its sd?

$\alpha = m\kappa = 0.04 \times 50 = 2$, $\beta = 0.96 \times 50 = 48$: Beta(2, 48). Its variance is $\frac{m(1-m)}{\kappa+1} = \frac{0.04 \times 0.96}{51} = 0.000753$, so sd $\approx 0.027$.

B. For the posterior Beta(14, 106), find the mean, the mode and the sd.

Mean $= 14/120 \approx 0.1167$. Mode $= \frac{14-1}{120-2} = \frac{13}{118} \approx 0.1102$. Variance $= \frac{14 \times 106}{120^2 \times 121} = \frac{1484}{1\,742\,400} \approx 0.000852$, so sd $\approx 0.0292$. The mode sits a little left of the mean: a slight right skew.

C. Prior Dirichlet(4, 4, 2) for (Basic, Pro, Enterprise). You then see 16, 26 and 8 customers. Give the posterior, its mean shares, and the sd of the Pro share.

Posterior Dirichlet(4 + 16, 4 + 26, 2 + 8) = Dirichlet(20, 30, 10), $\alpha_0 = 60$. Mean shares $(20/60, 30/60, 10/60) = (0.333, 0.5, 0.167)$. Pro alone ~ Beta(30, 30): variance $\frac{0.5 \times 0.5}{61} = 0.0041$, sd $\approx 0.064$.

D. Show that if $\lambda \sim$ Gamma(shape $\alpha$, rate $\alpha/\mu$) and $Y \mid \lambda \sim$ Poisson($\lambda$), then $E[Y] = \mu$ and $Var(Y) = \mu + \mu^2/\alpha$.

A Gamma(shape $a$, rate $r$) has mean $a/r$ and variance $a/r^2$. Here: $E[\lambda] = \alpha / (\alpha/\mu) = \mu$ and $Var(\lambda) = \alpha / (\alpha/\mu)^2 = \mu^2/\alpha$. For a Poisson, mean and variance both equal $\lambda$. Law of total expectation: $E[Y] = E[E[Y\mid\lambda]] = E[\lambda] = \mu$. Law of total variance: $Var(Y) = E[Var(Y\mid\lambda)] + Var(E[Y\mid\lambda]) = E[\lambda] + Var(\lambda) = \mu + \mu^2/\alpha$. The first term is the Poisson noise; the second is the extra spread from the wobbling rate.

E. Interview: choose a likelihood for (a) whether a user clicked, (b) seconds spent on a page, (c) support tickets per day (mean 8, variance 30), (d) forecast residuals with a few huge misses. Explain each in one line.

(a) Bernoulli: support {0, 1}. (b) Log-Normal or Gamma: positive, right-skewed, spread grows with the level. (c) Negative Binomial: counts with variance far above the mean (ratio 3.75), start from $\alpha \approx 64/22 \approx 2.9$. (d) Student-t: real-valued with heavy tails, so the few huge misses pull the fit less (it does not delete them).

F. Which Beta has mean 0.25 and standard deviation 0.05?

Use $Var = \frac{m(1-m)}{\kappa+1}$: $0.05^2 = \frac{0.25 \times 0.75}{\kappa + 1}$, so $\kappa + 1 = \frac{0.1875}{0.0025} = 75$ and $\kappa = 74$. Then $\alpha = 0.25 \times 74 = 18.5$ and $\beta = 0.75 \times 74 = 55.5$: Beta(18.5, 55.5). This "match the mean and sd" trick is a quick way to turn a stated belief into a prior.

Chapter 4.12 · Syllabus Modules 3.1–3.2

How far can data spread? Markov's and Chebyshev's inequalities

Usually you do not know the exact shape of your data's distribution. You might only know its average, or its average and its standard deviation. Surprisingly, that is already enough to promise something: "at most 20% of sessions can be this long", "at least 75% of days fall within two standard deviations". These promises hold for every possible shape. They are often loose, but they are never wrong, and one of them is the engine behind the Law of Large Numbers.

  • Say what a tail probability, a bound and a distribution-free statement are, and why the mean and standard deviation alone do not fix a tail
  • State, use and prove (with a picture) Markov's inequality $P(X \ge a) \le E[X]/a$ for non-negative $X$
  • Derive Chebyshev's inequality $P(|X-\mu| \ge k\sigma) \le 1/k^2$ from Markov, and know its conditions
  • See why Chebyshev is usually loose, yet exact for a special three-point distribution; know the one-sided version
  • Compare "at least $1 - 1/k^2$ within $k$ standard deviations" with the 68–95–99.7 rule
  • Use Chebyshev on averages: the link between variance and concentration, and the road to the Law of Large Numbers

Tail probabilities: how likely is an extreme value? core

A café owner knows two facts about daily orders: the average is 200 and the typical wobble, the standard deviation, is 20 (Chapter 4.5). She asks: "How often could I get a crazy day of 260 orders or more?"

If she knew the exact shape of the distribution (say, a bell curve), she could compute it. Usually she does not. Different shapes with the same average and the same standard deviation can give very different answers. What she can get without knowing the shape is a guarantee: a ceiling that the true answer can never exceed, whatever the shape.

The far end of a distribution is called its tail, and the chance of landing there is a tail probability. This chapter is about guarantees for tail probabilities.

Three ways to say it:

  • Picture: the two thin ends of a distribution, far from the middle, coloured red.
  • Numbers: with mean 200 and sd 20, the chance of a 260-order day is 0.13% for a bell shape but 5.6% for a lumpy shape with the very same mean and sd.
  • Slogan: the mean and the spread do not fix the tails, but they do put a ceiling on them.

Three distributions, all with mean 200 and standard deviation 20. How likely is a day with at least 260 orders (3 standard deviations above the mean)?

  1. Bell-shaped (Normal): $P(X \ge 260) = P(Z \ge 3) \approx 0.00135$, about 0.13%.
  2. Right-skewed (a Gamma with the same mean and sd): about 0.28%: twice as often, because its right tail is longer.
  3. Lumpy: 140 orders with probability $1/18$, 200 with probability $16/18$, 260 with probability $1/18$. Mean: $\frac{140 + 16 \times 200 + 260}{18} = \frac{3600}{18} = 200$. ✓ Variance: $\frac{1}{18}(60^2) + \frac{1}{18}(60^2) = \frac{7200}{18} = 400$, so sd $= 20$. ✓ And $P(X \ge 260) = 1/18 \approx 5.6\%$.
  4. Same mean, same sd, yet the answer ranges from 0.13% to 5.6%: more than 40 times apart.
  5. The good news (proved later in this chapter): no distribution with sd 20 can put more than $1/9 \approx 11.1\%$ of its days 60 or more orders away from the mean, on both sides together. That ceiling is Chebyshev's inequality.
  • A right-tail probability is $P(X \ge a)$; a left-tail probability is $P(X \le a)$; a two-sided tail is $P(|X - \mu| \ge t) = P(X \ge \mu + t) + P(X \le \mu - t)$, the chance of landing at least $t$ away from the mean $\mu$ on either side.
  • Distances are often measured in standard deviations: $t = k\sigma$ means "$k$ standard deviations away".
  • A bound (more precisely an upper bound) on a probability is a number $B$ with $P(\text{event}) \le B$ guaranteed.
  • A statement is distribution-free when it holds for every distribution that meets a few basic conditions (like "never negative" or "finite variance"), whatever its shape.
  • A bound is tight when some distribution reaches it exactly, so no better distribution-free bound exists. It is loose for a particular distribution when the true probability is much smaller.
Why do we need it?

Risk questions ("how often will demand exceed capacity?", "how often will an error be this big?") are tail questions, and the true shape is usually unknown. A guarantee lets you reason safely before you have a trustworthy model of the shape.

Where is it used?

Capacity planning, "3-sigma" alert thresholds on monitoring dashboards, sanity checks of anomaly detectors, and the theory behind the Law of Large Numbers, generalization bounds in learning theory and randomized algorithms.

How is it used?

Compute the mean (and sd) from data, plug them into Markov's or Chebyshev's inequality to get a worst case. If the worst case is already acceptable, you are done. If not, you need a model of the shape (Normal, Student-t, …) for a sharper number.

μ (the mean) μ − kσ μ + kσ kσ most values are close to μ left tail P(X ≤ μ − kσ) right tail P(X ≥ μ + kσ) two-sided tail: P(|X − μ| ≥ kσ) = the two red areas together
A tail is the far end of a distribution. Here $k = 2$: the red areas are the chance of landing at least $k$ standard deviations from the mean. For this bell shape they are small (about 4.6% together), but the size of the tails depends on the shape, which we usually do not know.

All six shapes have mean 0 and standard deviation 1 (they are "standardized"). Pick one to highlight; its tails beyond $\pm k$ are shaded red. Move $k$ to 2 and read the table: the two-sided tail ranges from 0 (Uniform) to about 6% (Laplace). Move $k$ to 3: now the skewed and heavy-tailed shapes keep far more mass out there than the Normal. The last line shows the worst any shape could ever do: $1/k^2$.

"The mean and the standard deviation tell me the chance of a 3-sigma day."

Only together with a shape. With mean and sd alone, the chance could be anything from 0 up to $1/9$. The shape is extra information.

"The 68–95–99.7 rule works for any data."

It is a property of the Normal (bell) shape. Skewed or heavy-tailed data can have far more beyond 3 sd (Section 5 compares the two rules).

"A bound is an estimate of the probability."

A bound is a ceiling. The true probability is at most the bound, and is often much smaller.

In your forecasting model, a natural monitoring rule is "flag a day whose residual is more than 3 standard deviations from 0". How many days get flagged depends on the residuals' shape: about 0.27% if they are Normal, more if they are heavy-tailed (one reason the model offers a Student-t likelihood). Without assuming a shape, the only promise is "at most 11.1%" (Chebyshev).

Tail probability: $P(X \ge a)$, or two-sided $P(|X - \mu| \ge k\sigma)$.

Mean and sd do not fix the tails; a shape does. Distribution-free bounds give a ceiling valid for every shape.

Trap: a bound is "at most", not "about".

Quick check: why can the Uniform shape have a two-sided tail of exactly 0 at $k = 2$?

A Uniform with sd 1 lives on $[-\sqrt3, \sqrt3] \approx [-1.73, 1.73]$. Nothing lies 2 or more sd from the mean, so the tail is 0. Bounded shapes can have empty tails.

Markov's inequality: the mean alone limits big values

The average session on your site lasts 4 minutes. Could half of all sessions last 20 minutes or more? No. Those sessions alone would add at least $0.5 \times 20 = 10$ minutes to the average, which is already more than 4. And because a session cannot last a negative time, nothing can pull the average back down.

So big values must be rare. How rare? If a fraction $q$ of sessions lasts at least 20 minutes, they contribute at least $20q$ to the average, so $20q \le 4$, which means $q \le 4/20 = 20\%$. That is Markov's inequality.

Three ways to say it:

  • Picture: on a seesaw that cannot go below zero, heavy weights far out must be few, or the balance point (the mean) would move out with them.
  • Numbers: mean 4, threshold 20: at most $4/20 = 20\%$ of values can reach it.
  • Slogan: for non-negative quantities, the mean puts a ceiling on how often a big value can happen.

Session length $X$ in minutes: $X \ge 0$, $E[X] = 4$.

  1. Bound for $a = 20$: $P(X \ge 20) \le E[X]/a = 4/20 = 0.2$.
  2. Can 20% really happen? Yes: let 80% of sessions last 0 minutes and 20% last exactly 20. Mean $= 0.8 \times 0 + 0.2 \times 20 = 4$ ✓, and $P(X \ge 20) = 0.2$. So the bound cannot be improved: it is tight.
  3. A more realistic shape: if session lengths were Exponential with mean 4, then $P(X \ge 20) = e^{-20/4} = e^{-5} \approx 0.0067$, about 30 times smaller than the bound.
  4. A useless case: $a = 3$, below the mean. The bound $4/3$ is bigger than 1, so it says nothing.

Markov's inequality. If $X \ge 0$ (never negative) has a finite mean, then for every $a \gt 0$

$$P(X \ge a) \le \frac{E[X]}{a}.$$

Proof by picture. Write $\mathbf{1}\{X \ge a\}$ for the indicator: 1 when the event happens, 0 otherwise (its average is the probability: $E[\mathbf{1}\{X \ge a\}] = P(X \ge a)$). For every outcome,

$$a \cdot \mathbf{1}\{X \ge a\} \ \le\ X,$$

because when $X \ge a$ the left side is $a \le X$, and when $X \lt a$ the left side is $0 \le X$ (here we need $X \ge 0$). Taking the average of both sides keeps the order: $a \cdot P(X \ge a) \le E[X]$. Divide by $a$.

  • Same statement with $a = c \cdot E[X]$: $P(X \ge c\,E[X]) \le 1/c$. "At most a third of the values can be 3 times the mean or more."
  • Data version: for any list of non-negative numbers, (how many are $\ge a$) / $n \le \bar x / a$.
  • Equality holds only for distributions that take just the two values 0 and $a$.
Why do we need it?

It gives a quick, almost assumption-free ceiling on how often a non-negative quantity can be large, from nothing but its average. And it is the building block of Chebyshev's inequality and of many sharper bounds.

Where is it used?

Sanity checks on orders, queue lengths, latencies and costs; bounding the chance that a randomized algorithm runs long; and as the first step in the proofs of Chebyshev's inequality, Chernoff bounds and other concentration results in ML theory.

How is it used?

Check that the quantity can never be negative, compute or bound its mean, divide by the threshold, and read the result as "at most". If the result is 1 or more, the threshold is at or below the mean and the bound says nothing.

2 4 6 8 10 value of X → a = 4 X itself a · 1{X ≥ a}: 0 below a, then a 0 gap ≥ 0 For X ≥ 0 the orange line is never above the blue one, so its average is never above the blue one's average: a · P(X ≥ a) ≤ E[X], so P(X ≥ a) ≤ E[X] / a
Markov's inequality, proved by a picture. For every non-negative value $x$, the orange step $a \cdot \mathbf{1}\{x \ge a\}$ sits below the blue line $x$. Average both over the distribution of $X$: the average of the orange step is $a \cdot P(X \ge a)$ and the average of the blue line is $E[X]$.

Top: the blue line is the value $X$ itself; the orange step is $a \cdot \mathbf{1}\{X \ge a\}$. The step never rises above the line (the green gap is never negative). Bottom: the distribution of $X$, every option with mean 3, with $P(X \ge a)$ in red. Move $a$ and compare $a \cdot P(X \ge a)$ (the average of the orange step) with $E[X] = 3$ (the average of the blue line). Choose two-point worst case: all mass at 0 or at $a$, and the inequality becomes an equality.

Ten sessions (blue, in minutes) and a threshold $a$ (purple, on the bottom row). Try to make the fraction of sessions at or above $a$ bigger than $\bar x / a$. You cannot: dragging a session above $a$ raises the mean too. The closest you can get is Worst case: every session at 0 or exactly at $a$. Then press Add a huge outlier and see the bound jump: Markov only knows the mean.

"Markov's inequality works for any random variable."

Only for non-negative ones. Counter-example: $X = -9$ or $+11$, each with probability 1/2. The mean is 1, yet $P(X \ge 11) = 0.5$, far above $1/11 \approx 0.09$. The negative value pulls the mean down and hides the big one.

"Markov tells me the probability of a big value."

It tells you the most it can be. For realistic shapes the truth is often many times smaller (30 times smaller in the session example).

"Markov needs the variance."

It needs only the mean, which is exactly why it is weak. Add the variance and you get Chebyshev, which is stronger.

In your forecasting model, daily demand is never negative (and is a count under the Negative Binomial likelihood), so Markov applies to it. With a predicted mean of 120 orders and a capacity of 300, the chance of exceeding capacity is at most $120/300 = 40\%$, whatever the distribution. That ceiling is crude; the model's predictive distribution (with its Negative Binomial likelihood) gives the real number (Chapter 7.14). Residuals can be negative, so for them use Chebyshev instead.

$X \ge 0$: $P(X \ge a) \le E[X]/a$. Equivalently, at most $1/c$ of the mass sits at $c$ times the mean or beyond.

Proof: $a \cdot \mathbf{1}\{X \ge a\} \le X$, then take averages. Tight for a two-point distribution on $\{0, a\}$.

Trap: needs non-negative $X$; says nothing when $a \le E[X]$; "at most", not "about".

Quick check: a store averages 2 refunds a day. What can you say about days with 10 or more refunds?

Refund counts are never negative, so Markov gives $P(X \ge 10) \le 2/10 = 0.2$: at most 20% of days. The truth is probably much smaller, but without the shape that is all you can promise.

Chebyshev's inequality: the variance limits how far values stray core

The variance is the average squared distance from the mean. Suppose many values were far from the mean. Their squared distances would be huge, so the variance would be huge. Turn that around: if the variance is small, there cannot be many values far away.

Chebyshev's inequality (say "CHEB-ih-shev") turns this into a number. Measure distance in standard deviations. Then the fraction of values at least $k$ standard deviations from the mean is at most $1/k^2$. Two standard deviations: at most a quarter. Three: at most a ninth. Whatever the shape.

Three ways to say it:

  • Picture: the variance is a fixed budget of squared distance; far-away values are expensive, so only a few can afford to be far.
  • Numbers: with mean 200 and sd 20, at most 25% of days fall outside 160–240, and at most 11.1% outside 140–260.
  • Slogan: small variance forces the values to huddle near the mean, whatever the shape.

Daily orders: mean $\mu = 200$, standard deviation $\sigma = 20$, shape unknown.

  1. $k = 2$: two standard deviations is $2 \times 20 = 40$, so the band is 160 to 240. Chebyshev: $P(|X - 200| \ge 40) \le 1/2^2 = 1/4$. At most 25% of days fall outside; at least 75% fall inside.
  2. $k = 3$: band 140 to 260. At most $1/9 \approx 11.1\%$ outside; at least 88.9% inside.
  3. $k = 1$: the bound is $1/1 = 1$. It says nothing (any probability is at most 1).
  4. Any distance, not just whole $k$: $P(|X - 200| \ge 50) \le \sigma^2 / 50^2 = 400/2500 = 0.16$.
  5. For comparison, if the shape were Normal: 4.6% outside 160–240 and 0.27% outside 140–260. The guarantee is much more cautious, because it must also cover strange shapes.

Chebyshev's inequality. If $X$ has mean $\mu$ and a finite variance $\sigma^2 \gt 0$, then for every $k \gt 0$

$$P\big(|X - \mu| \ge k\sigma\big) \le \frac{1}{k^2}, \qquad \text{equivalently} \qquad P\big(|X-\mu| \ge t\big) \le \frac{\sigma^2}{t^2}\ \ \text{for every } t \gt 0.$$

So at least $1 - 1/k^2$ of the probability lies strictly within $k$ standard deviations of the mean.

Derivation from Markov, step by step.

  1. Let $Y = (X - \mu)^2$. It is never negative, and its average is the variance: $E[Y] = \sigma^2$.
  2. The events match: $|X - \mu| \ge k\sigma$ exactly when $(X - \mu)^2 \ge k^2\sigma^2$ (both sides are non-negative, so squaring keeps the order).
  3. Markov on $Y$ with $a = k^2\sigma^2$: $P(Y \ge k^2\sigma^2) \le \dfrac{E[Y]}{k^2\sigma^2} = \dfrac{\sigma^2}{k^2\sigma^2} = \dfrac{1}{k^2}$.
  • Conditions: only a finite variance. Any shape: discrete or continuous, skewed or symmetric, bounded or not.
  • Data version: for any dataset, with its own mean and its own sd computed with divisor $n$, the fraction of points at least $k$ sds from the mean is at most $1/k^2$. (The usual sd with divisor $n-1$ is a little larger, so the band is wider and the statement still holds.)
  • One-sided version (Cantelli's inequality, optional): $P(X - \mu \ge k\sigma) \le \dfrac{1}{1 + k^2}$ for $k \gt 0$. At $k = 2$ that is 20% (plain Chebyshev would give 25%); at $k = 3$, 10%.
Why do we need it?

The variance is often the only spread information we have. Chebyshev turns it into a promise about how much of the data can sit far out, valid for every shape. It is also the key step in proving the Law of Large Numbers.

Where is it used?

Shape-free control limits and outlier rules, sanity checks of "3-sigma" alerts, the proof of the weak Law of Large Numbers (Chapter 4.13), sample-size guarantees that assume nothing about the shape (Section 6), and concentration arguments in learning theory.

How is it used?

Compute $\mu$ and $\sigma$ (or $\bar x$ and $s$ from data), pick $k$, and read "at most $1/k^2$ outside $\mu \pm k\sigma$". For a one-sided question use $1/(1 + k^2)$. If the number is too cautious for your purpose, you need a model of the shape.

1. Start X has mean μ and variance σ² 2. Square the distance Y = (X − μ)² ≥ 0 E[Y] = σ² 3. Markov on Y a = k²σ²: P(Y ≥ k²σ²) ≤ σ²/(k²σ²) 4. Same event Y ≥ k²σ² ⇔ |X − μ| ≥ kσ P(|X − μ| ≥ kσ) ≤ 1/k² Chebyshev = Markov applied to the squared distance from the mean.
The whole proof of Chebyshev's inequality in four steps. Squaring makes the distance non-negative (so Markov applies), and its average is exactly the variance.

1 000 session lengths drawn from a skewed shape with true mean $\mu = 4$ minutes and sd $\sigma = 2$. Press Next to walk through the proof: (1) the values, with those at least $k\sigma$ from the mean in red; (2) square each distance, $Y = (X - \mu)^2$: the red ones are the same sessions; (3) apply Markov to $Y$; (4) translate back. Change $k$ at any step and watch the red set and the numbers update.

Left: the red curve is Chebyshev's bound as $k$ grows; each grey curve is the true tail of one standardized shape; your selected shape is orange. Every curve stays under the red one, always. Tick log scale to see how far under: at $k = 3$ the Normal's tail is about 40 times smaller than the bound. Switch to one-sided: the red curve becomes Cantelli's $1/(1+k^2)$ (the dashed red is plain $1/k^2$). Right: the selected shape with its tail at the current $k$ shaded.

Page response times: never negative, mean $\mu$, standard deviation $\sigma$ (in milliseconds). For a slow-response threshold $a$, compare three guarantees for $P(T \ge a)$: Markov (uses only the mean), Chebyshev (uses the variance too), and Cantelli (the one-sided version). The vertical axis is on a log scale: each step down is 10 times smaller. Shrink $\sigma$ and watch the variance-based bounds plunge while Markov does not move. The orange curve is the truth for one example shape (a Gamma with this mean and sd): even the best bound is far above it.

"Chebyshev says 25% of the values are beyond two standard deviations."

At most 25%. For most real shapes far fewer are (4.6% for a Normal). It is a ceiling, not a prediction.

"Chebyshev is useful for any $k$."

For $k \le 1$ the bound is 1 or more, which says nothing. The inequality only bites beyond one standard deviation.

"It holds for every distribution, even with infinite variance."

It needs a finite variance. A Student-t with $\nu \le 2$, or a Cauchy (Chapter 4.9), has no finite variance, so there is no Chebyshev guarantee for them.

"For a one-sided question, just halve it: $P(X - \mu \ge k\sigma) \le \frac{1}{2k^2}$."

Halving is only valid for symmetric distributions. In general the best one-sided bound is $\frac{1}{1+k^2}$, which is larger (20% versus 12.5% at $k = 2$). A two-point distribution with $\mu + 2\sigma$ (probability 1/5) and $\mu - \sigma/2$ (probability 4/5) reaches 20% exactly.

Both of your projects can use a Student-t likelihood. Remember that its scale $\sigma$ is not its standard deviation: the sd is $\sigma\sqrt{\nu/(\nu - 2)}$, and only when $\nu \gt 2$. If a fitted $\nu$ is 2 or below, the residuals have no finite variance, so "within a few standard deviations" statements, and Chebyshev with them, simply do not apply. With $\nu \gt 2$, Chebyshev still promises at most 11.1% of residuals beyond 3 sds, whatever the exact shape.

"By Chebyshev, 75% of the data lies within two standard deviations."

"At least 75%, for any distribution with finite variance. It is distribution-free, so it is usually loose, but it cannot be improved without more assumptions (it is tight for a three-point distribution)."

Model answer: "Chebyshev says $P(|X-\mu| \ge k\sigma) \le 1/k^2$ whenever the variance is finite. It comes from applying Markov to $(X-\mu)^2$. It is distribution-free and usually loose: for a Normal the 2σ tail is 4.6%, not 25%. Its real value is theoretical: it shows that small variance forces concentration, and applied to a sample mean it proves the weak law of large numbers."

$P(|X-\mu| \ge k\sigma) \le 1/k^2$ (any shape, finite variance) ⇔ at least $1 - 1/k^2$ within $k$ sds. Also $P(|X - \mu| \ge t) \le \sigma^2/t^2$.

Proof: Markov on $Y = (X-\mu)^2$ with $a = k^2\sigma^2$. One-sided (Cantelli): $P(X - \mu \ge k\sigma) \le 1/(1+k^2)$.

Trap: "at most", not "exactly"; useless for $k \le 1$; needs finite variance; do not halve for one-sided questions.

Quick check: for which $k$ does Chebyshev guarantee at least 90% of values within $k$ standard deviations?

Need $1 - 1/k^2 \ge 0.9$, so $1/k^2 \le 0.1$, $k^2 \ge 10$, $k \ge \sqrt{10} \approx 3.16$. (For a Normal, $k = 1.645$ would already give 90%.)

Usually loose, sometimes exact: the three-point worst case

Think of the variance as a budget of squared distance. You want to spend it so that as much probability as possible lands at least $k\sigma$ from the mean. Two tricks:

  • Put the far values exactly at distance $k\sigma$. Putting them further out costs more budget (the cost grows with the square of the distance) without counting any more.
  • Put every other value exactly at the mean, where it costs nothing.

The result is a strange, spiky distribution with three spikes. It reaches Chebyshev's bound exactly, so the bound cannot be improved. But real data are rarely three spikes, which is why the bound is usually far from the truth.

Three ways to say it:

  • Picture: one tall spike at the mean and two small spikes exactly $k\sigma$ away on each side.
  • Numbers: for $k = 2$: 1/8 at $\mu - 2\sigma$, 3/4 at $\mu$, 1/8 at $\mu + 2\sigma$; exactly 25% lies 2σ away.
  • Slogan: tight in the worst case, loose in the usual case.

Take $\mu = 0$, $\sigma = 1$, $k = 2$.

  1. Values $-2, 0, 2$ with probabilities $\frac18, \frac34, \frac18$. They add to 1. ✓
  2. Mean: $\frac18(-2) + \frac34(0) + \frac18(2) = 0$. ✓
  3. Variance: $\frac18(4) + \frac34(0) + \frac18(4) = \frac12 + \frac12 = 1$. ✓
  4. Tail: $P(|X| \ge 2) = \frac18 + \frac18 = \frac14 = \frac{1}{k^2}$. Exactly the bound.

Look back at the lumpy café distribution of Section 1: 140, 200 and 260 orders with probabilities $\frac1{18}, \frac{16}{18}, \frac1{18}$. With $\sigma = 20$, the outer spikes sit exactly $3\sigma$ away, and $P(|X - 200| \ge 60) = \frac{2}{18} = \frac19$. It was the worst case for $k = 3$.

For $k \ge 1$, the distribution

$$X = \begin{cases} \mu - k\sigma & \text{with probability } \frac{1}{2k^2} \\ \mu & \text{with probability } 1 - \frac{1}{k^2} \\ \mu + k\sigma & \text{with probability } \frac{1}{2k^2} \end{cases}$$

has mean $\mu$, variance $2 \cdot \frac{1}{2k^2} \cdot k^2\sigma^2 = \sigma^2$, and $P(|X - \mu| \ge k\sigma) = \frac{1}{k^2}$ exactly.

  • So Chebyshev's bound is tight: no better bound can hold for every distribution with this mean and variance.
  • Why nothing beats it: $\sigma^2 = E[(X-\mu)^2] \ge E\big[(X-\mu)^2 \text{ over the far values only}\big] \ge k^2\sigma^2 \cdot P(|X - \mu| \ge k\sigma)$. Equality needs every near value exactly at $\mu$ and every far value exactly at distance $k\sigma$: the three-point shape.
  • Each $k$ has its own worst case: there is no single distribution that is worst for all $k$ at once.
Why do we need it?

Knowing the bound is tight stops you hoping for a better shape-free promise. Seeing how strange the worst case is explains why real tails are usually far below it, and why adding shape information pays off so much.

Where is it used?

Interview questions such as "can Chebyshev be improved?", worst-case reasoning in robust statistics and stress testing, and understanding why distribution-free sample-size guarantees are so expensive (Section 6).

How is it used?

When Chebyshev gives a disappointingly large number, ask: "does my data look anything like three spikes?" If not, a shape assumption checked with a Q-Q plot, or an inequality that uses more information (for example that the data are bounded), gives a sharper answer.

probability 1/8 probability 3/4 probability 1/8 μ − 2σ μ μ + 2σ k = 2: variance = 1/8·(2σ)² + 3/4·0 + 1/8·(2σ)² = σ² ✓ P(|X − μ| ≥ 2σ) = 1/8 + 1/8 = 1/4 = 1/k²: exactly the bound
The worst case for Chebyshev with $k = 2$. Every bit of the variance "budget" is spent putting mass exactly at distance $2\sigma$; the rest sits at the mean, where it costs nothing. No distribution with this mean and variance can put more than 1/4 of its mass at distance $2\sigma$ or more.

Three spikes with mean 0 and variance 1. The outer spikes sit at $\pm c$ and each gets probability $1/(2c^2)$, so the variance stays exactly 1 whatever $c$ is. The question is $P(|X| \ge k)$. Slide $c$ beyond $k$: the spikes count, but they must shrink, so the tail $1/c^2$ falls below $1/k^2$. Slide $c$ below $k$: the spikes no longer count at all. Press Put the spikes at k: exactly the bound.

"Chebyshev is tight, so it is accurate for my data."

"Tight" means it cannot be improved for all distributions at once. For a particular shape it can be very loose: for a Normal at $k = 3$ it says 11.1% while the truth is 0.27%, about 40 times smaller.

"There is one worst-case distribution."

The worst case depends on $k$: its outer spikes sit exactly at $\pm k\sigma$. The worst case for $k = 2$ has no mass at all beyond $2\sigma$, so for $k = 3$ it is not the worst.

Worst case for $k \ge 1$: $\mu \pm k\sigma$ with probability $\frac{1}{2k^2}$ each, $\mu$ with $1 - \frac{1}{k^2}$. Variance $\sigma^2$, tail exactly $1/k^2$.

Tight (cannot be improved shape-free) but usually loose for real shapes.

Trap: "tight" ≠ "accurate for my data"; each $k$ has its own worst case.

Quick check: write down the worst case for $k = 4$, with $\mu = 0$, $\sigma = 1$.

$\pm 4$ with probability $\frac{1}{2 \times 16} = \frac{1}{32}$ each, and 0 with probability $1 - \frac{1}{16} = \frac{15}{16}$. Variance: $2 \times \frac{1}{32} \times 16 = 1$ ✓. Tail: $\frac{2}{32} = \frac{1}{16} = \frac{1}{4^2}$ ✓.

What fraction lies within $k$ standard deviations? Chebyshev versus the 68–95–99.7 rule

There are two different promises about "how much of the data sits near the mean", and people mix them up all the time.

  • The 68–95–99.7 rule: about 68% within 1 sd, 95% within 2, 99.7% within 3. It is true only for the Normal (bell) shape (Chapter 4.9).
  • Chebyshev: at least $1 - 1/k^2$ within $k$ sds: 0% within 1, 75% within 2, 88.9% within 3. True for every shape with a finite variance.

The first is a forecast for bell curves. The second is an insurance policy that covers everything, which is why it promises less.

Three ways to say it:

  • Picture: for each $k$, a tall bar (Normal) next to a shorter bar (the guarantee for any shape).
  • Numbers: within 2 sds: 95.4% if Normal, at least 75% for anything.
  • Slogan: 68–95–99.7 is a fact about bell curves; "at least $1 - 1/k^2$" is a fact about all distributions.

Ten session lengths in minutes, including one power user: 2, 3, 3, 4, 4, 4, 5, 5, 6, 24.

  1. Mean: $(2+3+3+4+4+4+5+5+6+24)/10 = 60/10 = 6$.
  2. Distances from the mean: $-4, -3, -3, -2, -2, -2, -1, -1, 0, 18$. Squares: $16, 9, 9, 4, 4, 4, 1, 1, 0, 324$, sum 372.
  3. Sample sd: $s = \sqrt{372/9} = \sqrt{41.3} \approx 6.43$.
  4. Within 2 sds: $6 \pm 12.86$, the range $-6.86$ to $18.86$. Nine of ten values (90%) are inside; only the 24 is out. Chebyshev's "at least 75%" holds. ✓
  5. Within 1 sd: $6 \pm 6.43$. Again 9 of 10 (90%), far more than the Normal rule's 68%: the one big value inflated $s$, so everything else looks close.
  6. The 24 sits $(24 - 6)/6.43 \approx 2.8$ sds out. Under a Normal shape that happens about 0.5% of the time; here it is 1 value in 10. The data are simply not bell-shaped.

For any distribution with finite variance and $k \gt 1$: $P\big(|X - \mu| \lt k\sigma\big) \ge 1 - \dfrac{1}{k^2}$. For a Normal: $P\big(|X - \mu| \lt k\sigma\big) = 2\Phi(k) - 1$, where $\Phi$ is the standard Normal CDF.

$k$11.522.534
Chebyshev: at least (any shape)0%55.6%75%84%88.9%93.75%
Normal shape: exactly68.3%86.6%95.4%98.8%99.73%99.994%

Both apply to data too, with $\bar x$ and $s$ in place of $\mu$ and $\sigma$ (Chebyshev holds exactly with the divisor-$n$ sd and therefore also with the slightly larger usual $s$).

Why do we need it?

"95% within 2 standard deviations" gets quoted for every dataset. Knowing which promise applies protects you from false alarms on skewed data and from false comfort on heavy-tailed data.

Where is it used?

"Beyond 3 sd" outlier rules, anomaly thresholds, quick summaries of business metrics, and sanity checks of prediction bands built as "mean ± 2σ" from a Normal likelihood.

How is it used?

Compute the fraction of your data within $\bar x \pm k s$. It must be at least $1 - 1/k^2$. If it is far from the Normal rule, the shape is not Normal: look at a histogram or Q-Q plot (Chapter 4.17) before using Normal-based thresholds.

0% 25% 50% 75% 100% 0.0% 68.3% k = 1 55.6% 86.6% k = 1.5 75.0% 95.4% k = 2 88.9% 99.7% k = 3 Chebyshev: guaranteed for ANY shape (at least) if the shape is Normal
Fraction of values within $k$ standard deviations of the mean. Purple: the most Chebyshev can promise without knowing the shape. Orange: what a Normal shape gives (the 68–95–99.7 rule). The gap is the price of knowing nothing about the shape.

Pick a dataset (500 values, except the last). The green band is $\bar x \pm k s$; red bars are outside it. Compare the measured fraction inside with Chebyshev's guarantee (it always holds) and with the Normal rule (it only fits the bell-shaped data). Try the heavy-tailed data at $k = 3$, and the one big outlier data at $k = 1$ and $k = 2$.

"10% of my data is beyond 2 standard deviations, so something is broken."

Up to 25% is allowed for some shapes. It only tells you the data are not bell-shaped. Check the shape before calling anything an error.

"Anything beyond 3 sd is an outlier, by definition."

For skewed or heavy-tailed data, many genuine values lie beyond 3 sd. And a big outlier inflates $s$ itself, which can hide it (in the example, the 24 is only 2.8 sds out). Robust rules based on the median and MAD or the IQR are safer (Chapter 4.14).

"Chebyshev says at least 75% are within 2 sd, so about 75% are."

At least. For most real shapes the fraction is far higher: between 94% and 100% for the six standardized shapes of Section 1.

If forecast intervals are built as "prediction ± 2σ" from a Normal likelihood, they cover about 95% of days only when the residuals really are close to Normal. Chebyshev tells you the worst they could do (75%). Measuring the actual coverage on held-out days tells you the truth (Chapter 7.16). In the A/B framework, the same caution applies to "mean ± 2 sd" summaries of metrics with a few very large values.

Any shape: at least $1 - 1/k^2$ within $k$ sds (0%, 75%, 88.9% for $k = 1, 2, 3$).

Normal shape only: 68.3%, 95.4%, 99.7%.

Trap: never apply 68–95–99.7 to data you have not checked; big outliers inflate $s$ and can hide themselves.

Quick check: someone reports that 30% of their values lie more than 2.5 sds from the mean. Possible?

No. Chebyshev allows at most $1/2.5^2 = 16\%$ beyond 2.5 sds (using the data's own mean and sd), whatever the shape. The report must contain an error, for example an sd computed on a different subset of the data.

Variance controls concentration: Chebyshev for averages and the road to the Law of Large Numbers core

One user's behaviour is noisy. The average over many users is much calmer: their ups and downs cancel. In numbers, the variance of an average of $n$ independent values is the single-value variance divided by $n$ (Chapter 4.5).

Now apply Chebyshev to the average. Its variance $\sigma^2/n$ shrinks as $n$ grows, so the chance that the average misses the true mean by more than any fixed amount shrinks to zero. That is the weak Law of Large Numbers, proved in two lines (Chapter 4.13 explores it fully). It is also the clearest picture of how variance controls concentration: less variance, tighter packing.

Three ways to say it:

  • Picture: a funnel: the band where the average can wander narrows as more data arrive.
  • Numbers: conversion rate 10%; with 10 000 users the observed rate misses by a full percentage point or more with probability at most 9%.
  • Slogan: more data, smaller variance of the average, less room to be wrong.

Each user converts with probability $p = 0.1$ (a Bernoulli, Chapter 4.7), so one user's variance is $\sigma^2 = p(1-p) = 0.09$. The observed rate $\hat p$ is the average over $n$ users.

  1. Variance of the average: $Var(\hat p) = 0.09/n$.
  2. Chebyshev with $\varepsilon = 0.01$ (one percentage point): $P(|\hat p - 0.1| \ge 0.01) \le \dfrac{0.09}{n \times 0.01^2} = \dfrac{900}{n}$.
  3. $n = 10\,000$: at most $0.09$. $n = 100\,000$: at most $0.009$. The bound falls like $1/n$.
  4. Guaranteed sample size: to push the bound to 5%, need $900/n \le 0.05$, so $n \ge 18\,000$.
  5. With the Normal approximation instead (the CLT, Chapter 4.13): at $n = 10\,000$ the miss chance is about 0.0009, and 5% needs only $n \approx (1.96 \times 0.3 / 0.01)^2 \approx 3\,458$. The shape-free guarantee costs about 5 times more users.

Let $X_1, \dots, X_n$ be independent, each with mean $\mu$ and finite variance $\sigma^2$, and let $\bar X_n = \frac1n\sum_{i=1}^n X_i$.

  1. $E[\bar X_n] = \frac1n \sum E[X_i] = \frac1n \cdot n\mu = \mu$.
  2. $Var\big(\sum X_i\big) = \sum Var(X_i) = n\sigma^2$ (variances add for independent variables).
  3. $Var(\bar X_n) = \frac{1}{n^2} \cdot n\sigma^2 = \frac{\sigma^2}{n}$ (using $Var(cY) = c^2 Var(Y)$ with $c = 1/n$).
  4. Chebyshev with $t = \varepsilon$: $\displaystyle P\big(|\bar X_n - \mu| \ge \varepsilon\big) \le \frac{\sigma^2}{n\varepsilon^2}$ for every $\varepsilon \gt 0$.
  • Weak Law of Large Numbers: the right side goes to 0 as $n \to \infty$, so $P(|\bar X_n - \mu| \ge \varepsilon) \to 0$. The average converges to $\mu$ "in probability".
  • Shape-free sample size: $n \ge \dfrac{\sigma^2}{\delta\,\varepsilon^2}$ guarantees $P(|\bar X_n - \mu| \ge \varepsilon) \le \delta$. For a rate, $\sigma^2 = p(1-p) \le 1/4$ always.
  • Assumptions: independence (step 2 uses it) and finite variance (the LLN itself needs only a finite mean, but this short proof uses the variance).
Why do we need it?

It is the simplest rigorous reason why averages from big samples can be trusted, and it shows exactly how the variance, the sample size and the precision trade off.

Where is it used?

The proof of the weak LLN, reasoning about Monte Carlo error (how many simulation draws you need), sample-complexity bounds in learning theory, and rough, assumption-free sanity checks of sample sizes before an experiment.

How is it used?

With $\sigma^2$ (or an upper bound such as 1/4 for a rate), a tolerance $\varepsilon$ and an allowed miss chance $\delta$, compute $n \ge \sigma^2/(\delta\varepsilon^2)$. For real experiment planning use CLT-based power calculations (Chapter 5.7), which need far fewer users.

For each sample size $n$ (across, on a log scale), 400 experiments are simulated and the blue dot shows how often the average missed the true mean by at least $\varepsilon$. Red: Chebyshev's guarantee $\sigma^2/(n\varepsilon^2)$. Orange dashed: the Normal (CLT) approximation. Shrink $\varepsilon$ and watch every curve move right: precision costs data. Every blue dot stays under the red curve, usually far under it.

Choose the baseline conversion rate, the margin $\varepsilon$ you want the observed rate to be within, and the miss chance $\delta$ you can live with. The red bar is the number of users Chebyshev needs to guarantee it; the orange bar is what the Normal approximation says (the usual planning tool). The ratio between them depends only on $\delta$: about 5 at $\delta = 5\%$, about 15 at $\delta = 1\%$.

"Chebyshev gives the sample size I need for my A/B test."

It gives a safe but wasteful guarantee. Standard experiment planning uses the Normal approximation for the difference of two rates (power analysis, Chapter 5.7), which needs several times fewer users.

"Averages always settle down, whatever the data."

The $\sigma^2/n$ step needs independence, and the bound needs a finite variance. Correlated data (consecutive days of a time series) settle more slowly, and averages of Cauchy draws never settle (Chapter 4.13).

"The average has the same spread as one value."

Its variance is $\sigma^2/n$ and its standard deviation $\sigma/\sqrt n$, the standard error (Chapter 5.5). Four times the data halves the spread.

In the A/B framework, the observed conversion rate of a variant has variance $p(1-p)/n$: this is why more users per variant make every estimate (and every posterior) tighter. In the forecasting model, daily residuals are often correlated from one day to the next, so the "divide by $n$" rule overstates how fast averages over days settle (Chapter 7.3). And in your custom SVI loop the ELBO is a Monte Carlo estimate built from random draws (NumPyro's num_particles); averaging over $S$ draws makes its noise variance fall like $1/S$, the same $\sigma^2/n$ law (Chapter 6.12).

$Var(\bar X_n) = \sigma^2/n$ (independence) ⇒ $P(|\bar X_n - \mu| \ge \varepsilon) \le \dfrac{\sigma^2}{n\varepsilon^2} \to 0$: the weak LLN.

Shape-free sample size: $n \ge \sigma^2/(\delta\varepsilon^2)$; for a rate $\sigma^2 \le 1/4$. About $1/(\delta z^2)$ times the Normal-approximation $n$ (≈ 5× at δ = 5%).

Trap: needs independence and finite variance; Chebyshev is a guarantee, not a planning tool.

Quick check: the conversion rate is unknown. How many users does Chebyshev need to guarantee that the observed rate is within 2 percentage points with probability at least 90%?

Use the worst case $\sigma^2 = p(1-p) \le 0.25$, $\varepsilon = 0.02$, $\delta = 0.1$: $n \ge \frac{0.25}{0.1 \times 0.0004} = 6\,250$. (The Normal approximation would say $(1.645 \times 0.5 / 0.02)^2 \approx 1\,691$.)

Recap, cheat sheet and practice

  • A tail probability is the chance of landing far from the middle. The mean and sd do not fix it; a shape does. Distribution-free bounds give ceilings valid for every shape.
  • Markov (non-negative $X$): $P(X \ge a) \le E[X]/a$. Proof: $a\cdot\mathbf{1}\{X \ge a\} \le X$. Tight for mass on $\{0, a\}$.
  • Chebyshev (finite variance): $P(|X - \mu| \ge k\sigma) \le 1/k^2$, i.e. Markov applied to $(X - \mu)^2$. At least 75% within 2 sds, 88.9% within 3, for any shape.
  • It is tight (a three-point distribution reaches it) but usually loose (40 times too big for a Normal at $k = 3$). One-sided: $1/(1 + k^2)$; do not halve.
  • The 68–95–99.7 rule is for Normal shapes only; "at least $1 - 1/k^2$" is for all shapes.
  • For averages: $Var(\bar X) = \sigma^2/n$, so $P(|\bar X - \mu| \ge \varepsilon) \le \sigma^2/(n\varepsilon^2) \to 0$: the weak Law of Large Numbers. Variance controls concentration.

Cheat sheet

IdeaFormulaNeeds
Markov$P(X \ge a) \le E[X]/a$; $P(X \ge c\,E[X]) \le 1/c$$X \ge 0$, finite mean, $a \gt 0$
Chebyshev$P(|X-\mu| \ge k\sigma) \le 1/k^2$; $P(|X - \mu| \ge t) \le \sigma^2/t^2$finite variance
Within form$P(|X - \mu| \lt k\sigma) \ge 1 - 1/k^2$finite variance, $k \gt 1$
One-sided (Cantelli)$P(X - \mu \ge k\sigma) \le 1/(1 + k^2)$finite variance
Worst case for $k$$\mu \pm k\sigma$ w.p. $\frac{1}{2k^2}$ each, $\mu$ w.p. $1 - \frac{1}{k^2}$$k \ge 1$
Normal rule68.3% / 95.4% / 99.7% within 1 / 2 / 3 sdsNormal shape
Averages$P(|\bar X_n - \mu| \ge \varepsilon) \le \sigma^2/(n\varepsilon^2)$independence, finite variance
Shape-free sample size$n \ge \sigma^2/(\delta\varepsilon^2)$ ($\sigma^2 \le 1/4$ for a rate)same
Code it · Python
import numpy as np
from scipy import stats

# --- Markov: P(X >= a) <= E[X] / a for X >= 0 ---
print(4 / 20)                                # 0.2      the Markov bound for sessions of 20+ minutes (mean 4)
print(stats.expon(scale=4).sf(20))           # 0.0067   the truth if sessions were Exponential with mean 4

x = np.array([1, 2, 2, 3, 3, 4, 4, 5, 6, 10])   # ten session lengths (minutes)
a = 8
print((x >= a).mean(), x.mean() / a)         # 0.1 0.5  fraction at or above a, never more than mean / a

# --- Chebyshev: P(|X - mu| >= k*sd) <= 1/k^2, against true tails ---
ks = np.array([1.5, 2.0, 3.0])
print(np.round(1 / ks**2, 4))                # [0.4444 0.25   0.1111]   Chebyshev bounds for k = 1.5, 2, 3
shapes = {
    "normal":  stats.norm(),
    "laplace": stats.laplace(scale=1 / np.sqrt(2)),   # sd 1
    "t3":      stats.t(3, scale=1 / np.sqrt(3)),       # sd 1
    "expon":   stats.expon(loc=-1),                     # mean 0, sd 1, skewed
}
for name, d in shapes.items():
    m, s = d.mean(), d.std()
    tail = d.sf(m + ks * s) + d.cdf(m - ks * s)
    print(name, np.round(tail, 4))           # normal  [0.1336 0.0455 0.0027]
                                             # laplace [0.1199 0.0591 0.0144]
                                             # t3      [0.0805 0.0405 0.0138]
                                             # expon   [0.0821 0.0498 0.0183]   all far below the bounds

# --- the three-point worst case reaches the bound exactly ---
k = 2
vals = np.array([-k, 0, k])
probs = np.array([1 / (2 * k**2), 1 - 1 / k**2, 1 / (2 * k**2)])
print(probs @ vals, probs @ vals**2, probs[0] + probs[2])   # 0.0 1.0 0.25   mean 0, variance 1, tail exactly 1/k^2

# --- Chebyshev holds on any dataset (its own mean and sd, ddof=0) ---
rng = np.random.default_rng(0)
data = rng.lognormal(0, 1, 100_000)          # very skewed
z = np.abs(data - data.mean()) / data.std()  # np.std divides by n
print([round(float((z >= k).mean()), 4) for k in (2, 3, 4)])   # [0.036, 0.0176, 0.0097]  below 0.25, 0.111, 0.0625

# --- Chebyshev for averages: guaranteed vs Normal-approximation sample size ---
p, eps, delta = 0.1, 0.01, 0.05
var = p * (1 - p)
n_cheb = var / (delta * eps**2)
zq = stats.norm.ppf(1 - delta / 2)
n_clt = (zq * np.sqrt(var) / eps) ** 2
print(round(n_cheb), int(np.ceil(n_clt)))    # 18000 3458     about 5 times more users for the shape-free guarantee

# --- simulate 2 000 experiments with 10 000 users each ---
n = 10_000
phat = rng.binomial(n, p, size=2000) / n
print((np.abs(phat - p) >= eps).mean(), var / (n * eps**2))   # about 0.0005 vs the bound 0.09
Test yourself

1. $X \ge 0$ has mean 5. The most you can say about $P(X \ge 25)$ is…

Markov: $P(X \ge 25) \le 5/25 = 0.2$. It needs only the mean and non-negativity. It is a ceiling, not an exact value. (0.04 squares the ratio, as Chebyshev would, but Chebyshev needs a variance, which we were not given.)

2. Daily orders have mean 50 and sd 5, shape unknown. At least what fraction of days fall strictly between 40 and 60?

40 to 60 is $\mu \pm 2\sigma$. Chebyshev: at least $1 - 1/2^2 = 75\%$. The 95% figure needs a Normal shape.

3. What does Chebyshev's inequality require?

Only a finite variance (and so a finite mean). Non-negativity is Markov's condition; Chebyshev gets around it by applying Markov to $(X-\mu)^2$, which is always non-negative.

4. With $\mu = 0$ and $\sigma = 1$, which distribution makes Chebyshev exact at $k = 2$?

Its variance is $\frac18 \cdot 4 + \frac18 \cdot 4 = 1$ and $P(|X| \ge 2) = \frac14 = 1/k^2$. The ±1 coin has variance 1 but nothing at distance 2.

5. For any distribution with finite variance, the best general bound on the one-sided $P(X - \mu \ge 2\sigma)$ is…

Cantelli: $1/(1 + k^2) = 1/5$. Halving to 1/8 is only valid for symmetric distributions; 0.023 is the Normal value; 1/4 holds but is not the best.

6. Independent values with $\sigma^2 = 4$; the average of $n = 400$ of them. Chebyshev bounds $P(|\bar X - \mu| \ge 0.5)$ by…

$Var(\bar X) = 4/400 = 0.01$, so the bound is $\frac{\sigma^2}{n\varepsilon^2} = \frac{4}{400 \times 0.25} = 0.04$.

Practice problems

A. Page response times have mean 200 ms. Bound the chance of a response of 1 second or more. Then use the extra fact that the sd is 100 ms.

Markov (times are non-negative): $P(T \ge 1000) \le 200/1000 = 0.2$. With the sd: 1000 ms is $800/100 = 8$ sds above the mean. Chebyshev: $P(|T - 200| \ge 800) \le 1/64 \approx 0.0156$. One-sided Cantelli: $P(T - 200 \ge 800) \le 1/(1 + 64) = 1/65 \approx 0.0154$. Knowing the variance improved the guarantee from 20% to about 1.5%.

B. Derive Chebyshev's inequality from Markov's, step by step.

Let $Y = (X - \mu)^2 \ge 0$, so $E[Y] = \sigma^2$. Because both sides are non-negative, $|X - \mu| \ge k\sigma$ holds exactly when $Y \ge k^2\sigma^2$. Markov on $Y$ with $a = k^2\sigma^2$: $P(Y \ge k^2\sigma^2) \le E[Y]/(k^2\sigma^2) = 1/k^2$. Hence $P(|X - \mu| \ge k\sigma) \le 1/k^2$.

C. A dataset has mean 50 and sd 10. Could 30% of its values lie below 25 or above 75?

25 and 75 are $2.5$ sds from the mean. Chebyshev allows at most $1/2.5^2 = 0.16$ = 16% beyond that, for any shape (using the data's own mean and divisor-$n$ sd; the usual sd is slightly larger, which only makes the band wider). So 30% is impossible.

D. Check that the café distribution of Section 1 (140, 200, 260 orders with probabilities 1/18, 16/18, 1/18) is the Chebyshev worst case for $k = 3$.

Mean: $(140 + 16 \times 200 + 260)/18 = 3600/18 = 200$. Variance: $(60^2 + 60^2)/18 = 400$, so $\sigma = 20$. The outer values are $60 = 3\sigma$ away, each with probability $1/18 = 1/(2 \cdot 3^2)$, and the middle one has $16/18 = 1 - 1/9$. So $P(|X - 200| \ge 60) = 2/18 = 1/9 = 1/k^2$: exactly the bound.

E. Interview: "Is Chebyshev's inequality useful in practice?" Answer in a few sentences.

"As a number, rarely: it is distribution-free, so it must cover strange shapes and is usually loose; for a Normal at 3σ it says 11% when the truth is 0.27%. It is tight, though: a three-point distribution reaches it, so no better bound exists without more assumptions. Its real value is conceptual and theoretical: it shows that a small variance forces values to concentrate near the mean, and applied to a sample mean with variance $\sigma^2/n$ it proves the weak law of large numbers. In practice I use it as a sanity check and use a shape model, such as a Normal or Student-t, for sharp numbers."

F. The conversion rate is unknown. With Chebyshev, how many users guarantee that the observed rate is within 2 percentage points of the truth with probability at least 90%? Compare with the Normal approximation.

Worst-case variance $p(1-p) \le 1/4$. Chebyshev: $n \ge \frac{0.25}{0.1 \times 0.02^2} = \frac{0.25}{0.00004} = 6\,250$. Normal approximation: $z = 1.645$ for 90% two-sided, $n \approx (1.645 \times 0.5 / 0.02)^2 \approx 1\,691$. Ratio $\approx 3.7 = 1/(0.1 \times 1.645^2)$.

Chapter 4.13 · Syllabus Modules 3.3–3.4

Law of Large Numbers and Central Limit Theorem

Two theorems explain why averages are so useful. The Law of Large Numbers says an average of many draws settles down near the true mean. The Central Limit Theorem says the small leftover wobble of that average has a bell shape, even when the data do not. Almost every error bar, A/B test and Monte Carlo estimate you will ever compute leans on these two facts.

  • See the sample mean $\bar X_n$ as a random variable with its own centre $\mu$ and its own spread $\sigma/\sqrt n$
  • State the Law of Large Numbers (weak and strong), prove the weak one with Chebyshev's inequality, and avoid the "law of averages" trap
  • See the LLN fail when there is no mean (the Cauchy distribution)
  • State the Central Limit Theorem precisely: what converges (the standardized mean), what becomes Normal (its sampling distribution, not the data), and what the theorem does not say
  • Explain why skewed data need a bigger $n$, and why dependence or infinite variance breaks the $\sigma/\sqrt n$ rule
  • Use both theorems to read Monte Carlo estimates such as $P(\theta_B \gt \theta_A \mid D)$ computed from posterior draws

The sample mean is a random variable core

You show a new checkout page to 100 users and 10 of them buy: an observed conversion rate of 0.10. A colleague runs the very same experiment with 100 other users and gets 0.13. Nobody made a mistake. The average just depends on which users happened to arrive.

So before you run an experiment, its average is not a fixed number yet. It is a random variable: a number produced by a chance process (Chapter 4.4). It has its own centre and its own spread, just like a single draw does. Both theorems in this chapter are about this one random variable.

Three ways to say it:

  • Picture: every experiment scoops a different handful of users, so every average lands in a slightly different place; pile up many such averages and they form a cloud around the truth.
  • Numbers: with a true rate of 0.1, an average over 100 users typically lands about 0.03 away from 0.1; over 400 users, about 0.015 away.
  • Slogan: the average is itself random, with centre $\mu$ and spread $\sigma/\sqrt n$.

Each user converts (value 1) with probability $p = 0.1$, or not (value 0). Users are independent. Find the centre and the spread of the average over $n = 100$ users.

  1. One user: $E[X] = 1\cdot 0.1 + 0\cdot 0.9 = 0.1$, and $Var(X) = p(1-p) = 0.1 \times 0.9 = 0.09$, so $\sigma = \sqrt{0.09} = 0.3$ (Chapter 4.7).
  2. The average is $\bar X = (X_1 + \dots + X_{100})/100$.
  3. Centre. Expectation is linear (Chapter 4.5): $E[\bar X] = \tfrac{1}{100}(0.1 + \dots + 0.1) = \tfrac{1}{100}\cdot 100 \cdot 0.1 = 0.1$.
  4. Spread of the sum. The users are independent, so their variances add: $Var(X_1 + \dots + X_{100}) = 100 \times 0.09 = 9$.
  5. Dividing a random variable by 100 divides its variance by $100^2$: $Var(\bar X) = 9/10\,000 = 0.0009$.
  6. Standard deviation of the average: $\sqrt{0.0009} = 0.03$. That is the "typical miss" of one experiment.
  7. With $n = 400$: $Var(\bar X) = 0.09/400 = 0.000225$ and the spread is $\sqrt{0.000225} = 0.015$. Four times the users gives half the spread.

Let $X_1, \dots, X_n$ be iid (independent and identically distributed: each is a fresh draw from the same distribution, and no draw tells you anything about another) with mean $\mu$ and variance $\sigma^2$. The sample mean and the sum are

$$\bar X_n = \frac{1}{n}\sum_{i=1}^{n} X_i, \qquad S_n = \sum_{i=1}^{n} X_i = n\bar X_n .$$

Then

$$E[\bar X_n] = \mu, \qquad Var(\bar X_n) = \frac{\sigma^2}{n}, \qquad SD(\bar X_n) = \frac{\sigma}{\sqrt n}; \qquad E[S_n] = n\mu, \quad Var(S_n) = n\sigma^2 .$$
  • Capital $\bar X_n$ is the random variable (the recipe, before the data exist). Small $\bar x$ is the one number you got.
  • The standard deviation of an estimate across repeated samples has a special name, the standard error (SE). Here $SE(\bar X_n) = \sigma/\sqrt n$. It is taught fully in Chapter 5.5.
  • Assumptions: $E[\bar X_n] = \mu$ needs only that every $X_i$ has mean $\mu$. The variance formula also needs a finite $\sigma^2$ and no correlation between the draws (independence is more than enough). Section 8 shows what happens without it.
  • The distribution of $\bar X_n$ over imagined repeated samples is called its sampling distribution.
Why do we need it?

We almost never know the true mean of anything. We only see an average of some data. To know how much to trust that average, we must know how much it would change if we repeated the data collection. That is exactly the spread $\sigma/\sqrt n$.

Where is it used?

Error bars and confidence intervals, A/B test sample-size planning (the $\sqrt n$ is why halving the noise costs four times the users), mini-batch gradients in SGD (an average over a batch), Monte Carlo estimates from posterior draws, and averaging forecasts over many days.

How is it used?

Compute $\bar x$ and $s$ from the data, then report $\bar x \pm$ about $2\,s/\sqrt n$. To shrink the error by a factor $k$, plan for $k^2$ times more data. In code: x.mean() and x.std(ddof=1) / np.sqrt(len(x)).

Population mean μ, sd σ sample 1 → x̄₁ = 0.11 sample 2 → x̄₂ = 0.08 sample 3 → x̄₃ = 0.13 … and so on, thousands of times μ the pile of all the x̄'s centre μ · spread σ/√n
The same process (left) gives a different sample every time, so the sample mean changes too. Pile up the means of many repeated experiments and you get the sampling distribution of $\bar X$: centred on $\mu$, with spread $\sigma/\sqrt n$.

The true conversion rate is 0.1. Each click of Run 1 experiment of each runs two experiments: one with the smaller number of users (blue, top) and one with four times as many users (orange, bottom). Press Run 200 of each a few times. Notice: both piles are centred on the green line (the truth), but the orange pile is about half as wide. Switch between the pairs of sizes: the ratio of the widths stays near 2, because $\sqrt 4 = 2$.

"The sample mean $\bar x$ is the true mean."

$\bar x$ is an estimate of $\mu$. $\mu$ is fixed; $\bar x$ changes from sample to sample. They are close when $n$ is large, which is what the rest of this chapter is about.

"The average of 100 values has spread $\sigma$, like one value."

One value has spread $\sigma$. The average of $n$ independent values has spread $\sigma/\sqrt n$: averaging cancels out part of the noise.

"Doubling the data halves the error."

The error shrinks like $1/\sqrt n$. Doubling $n$ only divides it by $\sqrt 2 \approx 1.41$. To halve it you need four times the data.

In your A/B framework, each variant's observed conversion rate is exactly this $\bar X$: an average of Bernoulli outcomes. Its spread $\sqrt{p(1-p)/n}$ is why a small segment with 80 users can show a rate far from the truth while the full variant with 20 000 users barely moves. That jumpiness of small groups is the reason partial pooling exists (Chapter 6.6). In your forecasting model, a weekly average of demand is less noisy than a single day, but only if the days' surprises are independent, which section 8 questions.

$E[\bar X_n] = \mu$, $\;Var(\bar X_n) = \sigma^2/n$, $\;SD(\bar X_n) = \sigma/\sqrt n$ (the standard error). For the sum: $E = n\mu$, $Var = n\sigma^2$.

The sample mean is a random variable; its distribution over repeated samples is its sampling distribution.

Trap: error shrinks like $1/\sqrt n$, so halving it costs 4× the data. The formula assumes independent draws.

Quick check: a fair die has $\sigma^2 = 35/12$. What is the spread (sd) of the average of 16 rolls?

$Var(\bar X) = (35/12)/16 = 35/192 \approx 0.182$, so $SD(\bar X) = \sqrt{0.182} \approx 0.427$. Equivalently $\sigma/\sqrt{16} = 1.708/4 \approx 0.427$.

The Law of Large Numbers: averages settle down core

Flip a fair coin 10 times and get 7 heads (70%). Nobody is surprised. Flip it 10 000 times and get 70% heads, and you would check the coin. With many flips you expect something like 49.6% or 50.3%: very close to one half.

That is the Law of Large Numbers (LLN): as you average more and more independent draws, the average gets closer and closer to the true mean, and stays there. Casinos, insurance companies and polling all rely on it. So does every estimate you compute from data.

Three ways to say it:

  • Picture: a line that tracks the running average jumps around wildly at first, then calms down and hugs the true mean.
  • Numbers: after 10 die rolls the average might easily be 4.2; after 10 000 rolls it lands between 3.45 and 3.55 about 99.7% of the time.
  • Slogan: averages settle down.

Roll a fair die ($\mu = 3.5$) eight times: 6, 2, 5, 1, 4, 3, 6, 2. Track the running average (the average of all rolls so far).

  1. Running sums: $6,\ 8,\ 13,\ 14,\ 18,\ 21,\ 27,\ 29$.
  2. Divide each by how many rolls so far ($1, 2, \dots, 8$): $6/1 = 6$, $8/2 = 4$, $13/3 \approx 4.33$, $14/4 = 3.5$, $18/5 = 3.6$, $21/6 = 3.5$, $27/7 \approx 3.86$, $29/8 = 3.625$.
  3. The first roll alone missed by $6 - 3.5 = 2.5$. After eight rolls the miss is only $3.625 - 3.5 = 0.125$.
  4. Why the moves get smaller: the 8th roll can change the running average by at most $(6 - 1)/8 \approx 0.6$, while the 2nd roll could change it by up to $(6-1)/2 = 2.5$. Each new draw is a smaller and smaller share of the total.
  5. How close after 10 000 rolls? The spread is $\sigma/\sqrt n = 1.708/100 \approx 0.017$, so a miss of $0.05$ is about $2.9$ spreads: rare (about 0.3% of the time, using the bell shape from section 5).

Let $X_1, X_2, \dots$ be iid with a finite mean $\mu = E[X]$ (precisely: $E|X| \lt \infty$). Let $\bar X_n$ be the average of the first $n$.

  • Weak LLN (convergence in probability): for every tolerance $\varepsilon \gt 0$, $$P\big(|\bar X_n - \mu| \gt \varepsilon\big) \to 0 \quad \text{as } n \to \infty .$$ In words: pick any small band around $\mu$; the chance that the average is outside the band shrinks to zero.
  • Strong LLN (almost sure convergence): $P\big(\bar X_n \to \mu\big) = 1$. In words: if you keep one running-average line going forever, it will (with probability 1) eventually enter any band around $\mu$ and never leave it again.

"Converges" means "gets arbitrarily close as $n$ grows". The LLN says where the average goes. It does not say how fast; the Central Limit Theorem (section 5) answers that.

Why do we need it?

It is the reason "estimate a mean by averaging data" works at all, and the reason probability can be read as a long-run frequency (Chapter 4.2). Without it, more data would not guarantee better answers.

Where is it used?

Every Monte Carlo estimate (posterior means from NUTS or SVI draws, $P(\theta_B \gt \theta_A \mid D)$ as a fraction of draws, ELBO estimates), training loss as an average that approximates the expected loss, conversion rates in A/B tests, insurance premiums, and polling.

How is it used?

Collect more independent draws and average them. In practice, plot the running average against $n$: once it has flattened inside the accuracy you need, you have enough draws. If it keeps jumping, suspect heavy tails (section 4) or dependence (section 8).

Five independent people each draw values from the same population and track their running average (x-axis: number of draws, on a log scale). The shaded band is $\mu \pm \varepsilon$. Notice that every line wobbles a lot at first and then settles inside the band. Make the band narrower: the lines need more draws to stay inside. Try Bernoulli p = 0.05 (rare events take longer) and then Cauchy, which has no mean: its lines never settle (section 4).

"After 10 heads in a row, tails is due: the law of averages will even things out."

Coins have no memory. The next flip is still 50/50. The proportion of heads still goes to one half, but by dilution, not by compensation: the early surplus of heads is not cancelled, it just becomes a tiny share of a huge total. The widget below shows this.

"The LLN says that with large $n$, $\bar x = \mu$."

It says $\bar x$ gets arbitrarily close to $\mu$ with high probability. For any finite $n$ there is still a miss, typically of size $\sigma/\sqrt n$.

"The LLN works for any data."

It needs independent (or weakly dependent) draws from a distribution with a finite mean. Section 4 shows a distribution where it fails completely.

Each line starts with 10 heads in a row, then flips a fair coin up to 2000 times. Top: the proportion of heads drifts to 0.5. Bottom: the surplus of heads (heads minus half the flips) does not go back to 0. Its expected value stays at +5 forever (green), and the lines just wander around it, more and more widely (the shaded band is ±2 standard deviations, $\pm\sqrt{n-10}$). Press New sample several times. The proportion still goes to 0.5 because a surplus of about 5 divided by $n$ goes to zero.

In your A/B framework the conversion rate of each variant settles as users accumulate: that is the LLN. The same law runs inside your Bayesian tools. A posterior mean reported by NumPyro is an average over posterior draws, and $P(\theta_B \gt \theta_A \mid D)$ is the fraction of draws where $\theta_B$ wins. Both are trustworthy only because averages over many draws settle (section 9).

"The Law of Large Numbers means that past bad luck gets corrected."

Nothing gets corrected. Future draws are independent of past ones. The past surplus is diluted by the growing number of draws.

Model answer: "The LLN says the sample average converges to the expected value as $n$ grows, for independent draws with a finite mean. It is about averages, not individual draws, and it works by dilution: the early deviations don't disappear, they become negligible relative to $n$."

LLN: iid draws with finite mean $\mu$ ⇒ $\bar X_n \to \mu$. Weak: $P(|\bar X_n - \mu| \gt \varepsilon) \to 0$. Strong: the running-average path itself converges, with probability 1.

It works by dilution, not compensation (no "tails is due").

Needs: independence (or weak dependence) and a finite mean. Says nothing about speed.

Quick check: after 20 flips you have 15 heads. What do you expect the proportion of heads to be after 1020 flips (fair coin from now on)?

The next 1000 flips give about 500 heads on average, so you expect $(15 + 500)/1020 \approx 0.505$. The surplus of 5 heads stays; it is just diluted.

Why the LLN holds: a two-line proof with Chebyshev

In Chapter 4.12 you met Chebyshev's inequality: if a quantity has a small variance, it cannot often be far from its mean, whatever its distribution. Now combine it with section 1: the average $\bar X_n$ has variance $\sigma^2/n$, which shrinks to zero as $n$ grows. A quantity whose variance goes to zero has less and less room to be far from its mean. That is the whole proof.

Three ways to say it:

  • Picture: the variance is a "budget" for spreading out; the average's budget shrinks like $1/n$, so big misses become impossible to afford.
  • Numbers: with $p = 0.1$, $n = 1000$ users and a tolerance of $0.03$, Chebyshev guarantees at most a 10% chance of missing by that much. The true chance is about 0.19%.
  • Slogan: variance going to zero forces the chance of a miss to go to zero.

Conversion rate $p = 0.1$ (so $\sigma^2 = 0.09$), $n = 1000$ users. How likely is the observed rate to miss $0.1$ by $0.03$ or more?

  1. Variance of the average: $Var(\bar X) = \sigma^2/n = 0.09/1000 = 0.00009$.
  2. Chebyshev in the form $P(|Y - E[Y]| \ge \varepsilon) \le Var(Y)/\varepsilon^2$, with $Y = \bar X$ and $\varepsilon = 0.03$: $\;0.00009/0.03^2 = 0.00009/0.0009 = 0.1$.
  3. So at most 10% of experiments can miss by $0.03$ or more. This is a guarantee that needs no bell-curve assumption.
  4. The exact answer (from the Binomial distribution, a miss means 70 or fewer, or 130 or more, conversions out of 1000) is $0.00057 + 0.00134 \approx 0.0019$, about 0.19%. The bound is true, but about 50 times too pessimistic.
  5. Turned around: Chebyshev says $n \ge \sigma^2/(\delta\,\varepsilon^2)$ users make the miss chance at most $\delta$. For $\delta = 0.05$: $n \ge 0.09/(0.05 \times 0.0009) = 2000$. The bell-curve answer of section 5 is $n \approx (1.96 \times 0.3/0.03)^2 \approx 385$. The guarantee is safe but costly.

Weak LLN with finite variance. If $X_1, \dots, X_n$ are iid (or just uncorrelated) with mean $\mu$ and finite variance $\sigma^2$, then for every $\varepsilon \gt 0$

$$P\big(|\bar X_n - \mu| \ge \varepsilon\big) \;\le\; \frac{Var(\bar X_n)}{\varepsilon^2} \;=\; \frac{\sigma^2}{n\,\varepsilon^2} \;\longrightarrow\; 0 \quad (n \to \infty).$$
  • Step 1 uses Chebyshev's inequality (Chapter 4.12) for the random variable $\bar X_n$, whose mean is $\mu$.
  • Step 2 uses $Var(\bar X_n) = \sigma^2/n$ from section 1.
  • Step 3: $\varepsilon$ and $\sigma^2$ are fixed numbers, so dividing by a growing $n$ sends the bound to 0.

This proof needs a finite variance. The LLN itself is stronger: it only needs a finite mean (Khinchin's weak LLN; Kolmogorov's strong LLN), but those proofs need more advanced tools.

Why do we need it?

It turns the LLN from a hopeful slogan into a proven fact, and it gives a sample size that is safe for any distribution with a known variance bound, with no Normal approximation needed.

Where is it used?

Proofs of consistency of estimators (Chapter 5.1), worst-case sample-size guarantees in randomized algorithms and A/B tooling, concentration arguments in learning theory (Hoeffding and Bernstein bounds are sharper cousins), and Monte Carlo error budgets.

How is it used?

Plug in a variance (or an upper bound: for any 0/1 metric $\sigma^2 \le 0.25$) and your tolerance into $\sigma^2/(n\varepsilon^2)$. Use it when you must be safe whatever the distribution; use the CLT when you want a realistic estimate.

1 · spread of x̄ Var(X̄ₙ) = σ²/n 2 · Chebyshev P(|Y − E Y| ≥ ε) ≤ Var(Y)/ε² 3 · put Y = X̄ₙ P(|X̄ₙ − μ| ≥ ε) ≤ σ²/(n ε²) 4 · let n grow σ²/(n ε²) → 0 miss chance → 0 needs: draws uncorrelated, variance σ² finite
The weak Law of Large Numbers in four steps: the average's variance shrinks like $1/n$, and Chebyshev turns a small variance into a small chance of a big miss.

The plot shows, for each number of users $n$, the chance that the observed conversion rate misses the true rate $p$ by at least $\varepsilon$ (vertical axis on a log scale: each grid line is 10 times smaller). Purple: Chebyshev's bound. Green: the exact chance (from the Binomial distribution). Orange dashed: the bell-curve (CLT) approximation. Move $n$ and notice the purple line is always above the green one (a guarantee can never be beaten), but often thousands of times higher. The little zig-zags in green come from counts being whole numbers.

"Chebyshev says the chance of a miss is 10%."

It says the chance is at most 10%. It is a ceiling that holds for every distribution with that variance, so for any particular distribution the truth is usually far lower.

"The Law of Large Numbers needs a finite variance."

This short proof needs it. The LLN itself only needs a finite mean. Section 4 shows what happens without a finite mean.

$P(|\bar X_n - \mu| \ge \varepsilon) \le \sigma^2/(n\varepsilon^2) \to 0$: Chebyshev + $Var(\bar X_n) = \sigma^2/n$.

Distribution-free guarantee: safe sample size $n \ge \sigma^2/(\delta\varepsilon^2)$; usually much more than needed.

Trap: a bound is a ceiling, not the probability.

Quick check: a 0/1 metric (so $\sigma^2 \le 0.25$). How many users does Chebyshev say guarantee at most a 1% chance of missing by 0.05 or more?

$n \ge \sigma^2/(\delta\varepsilon^2) = 0.25/(0.01 \times 0.0025) = 0.25/0.000025 = 10\,000$ users. This works for any true rate, because $p(1-p) \le 0.25$ always.

When averages never settle: the Cauchy distribution

A lighthouse stands 1 km from a long straight beach. It flashes once at a completely random angle (any direction toward the beach is equally likely), and you record where the beam hits the beach. Usually it lands nearby. But when the angle is almost parallel to the beach, the beam lands 50, 500 or 5 000 km away. These monster values are rare, yet not rare enough: one of them can wreck an average of thousands of ordinary values.

Where the beam lands follows the Cauchy distribution (it is also the Student-t with $\nu = 1$, Chapter 4.9). It has no mean. With no mean, there is nothing for the average to settle on, and the Law of Large Numbers fails completely.

Three ways to say it:

  • Picture: the running average looks calm for a while, then a single huge draw kicks it somewhere else, again and again, forever.
  • Numbers: the average of 1 000 Cauchy draws has exactly the same distribution as one draw: in both cases there is about a 6.3% chance of landing beyond ±10.
  • Slogan: no mean, nothing to settle on.

Why is there no mean? The mean exists only if the average size $E|X|$ is finite. For the standard Cauchy density $f(x) = \dfrac{1}{\pi(1+x^2)}$, compute the part of $E|X|$ that comes from $-M \le x \le M$:

  1. $\displaystyle\int_{-M}^{M} \frac{|x|}{\pi(1+x^2)}\,dx = \frac{2}{\pi}\int_0^M \frac{x}{1+x^2}\,dx = \frac{1}{\pi}\ln(1+M^2)$.
  2. $M = 10$: $\ln(101)/\pi \approx 1.47$.
  3. $M = 1000$: $\ln(1\,000\,001)/\pi \approx 4.40$.
  4. $M = 10^6$: $\ln(10^{12}+1)/\pi \approx 8.80$.
  5. It keeps growing (slowly, like a logarithm) and never levels off, so $E|X| = \infty$: the far tails carry "infinite average size". The mean is undefined.
  6. Tail check: $P(|X| \gt 10) = 1 - \tfrac{2}{\pi}\arctan(10) \approx 0.063$. For a standard Normal it is about $10^{-23}$.

The LLN needs $E|X| \lt \infty$. Three cases:

  • Finite variance (Normal, Poisson, Bernoulli, Exponential, Student-t with $\nu \gt 2$): the average settles, with error about $\sigma/\sqrt n$, and the CLT applies.
  • Finite mean, infinite variance (Student-t with $1 \lt \nu \le 2$, Pareto with tail index between 1 and 2): the LLN still holds (the average does settle), but slowly and with occasional jumps; the usual CLT with $\sigma/\sqrt n$ does not apply.
  • No mean (Cauchy, Student-t with $\nu \le 1$): the LLN fails. For iid standard Cauchy draws, $\bar X_n$ is again standard Cauchy for every $n$: averaging does not help at all.

The median still works for the Cauchy: the sample median converges to the centre 0. This is one reason robust statistics exist (Chapter 4.14).

Why do we need it?

The LLN is so familiar that people forget it has a condition. Knowing the failure case tells you when "just collect more data and average" is dangerous: when single values can be astronomically large.

Where is it used?

Heavy-tailed metrics (revenue per user with rare huge orders, latency spikes), Student-t likelihoods with small degrees of freedom $\nu$, ratio metrics whose denominator can be near zero (the ratio of two independent Normals centred at zero is exactly Cauchy), and importance-sampling weights with infinite variance.

How is it used?

Before trusting an average, look at a running-average plot and at the largest values. If the plot keeps jumping, use a median or a trimmed mean, model the tails explicitly (Student-t, log transform), or check that your prior keeps $\nu$ in a range where the mean exists.

2 000 draws, tracked as they arrive. Orange: the running mean. Green: the running median. Red ticks at the top mark where the three largest values (in size) arrived. With Cauchy, press New sample several times: the orange line jumps exactly at the red ticks and never settles, while the green line calmly approaches 0. With Student-t ν = 1.5 the mean exists, so the orange line does settle, but slowly and with jumps. With Normal both settle smoothly.

"Every distribution has a mean; I just need enough data to find it."

Some distributions have no mean at all (Cauchy, Student-t with $\nu \le 1$). Their sample average never converges, no matter how much data you collect.

"If the average of my data looks stable for a while, the mean exists."

A Cauchy running mean can look calm for hundreds of draws and then jump. Stability over a short stretch proves nothing; look at the tails (a Q-Q plot, Chapter 4.17, or the largest values).

Both of your projects use a Student-t likelihood. Its degrees of freedom $\nu$ decide which case you are in: $\nu \le 1$ means the noise has no mean, $1 \lt \nu \le 2$ means infinite variance, $\nu \gt 2$ means a finite variance $\sigma^2\nu/(\nu-2)$. If $\nu$ is learned from data, check what range your prior allows. Predictive means and variances computed from draws only make sense when they exist. Note that the Student-t scale parameter is never the standard deviation (Chapter 4.9).

LLN needs $E|X| \lt \infty$. Cauchy: no mean, $\bar X_n \sim$ Cauchy(0, 1) for every $n$; averaging never helps.

Finite mean but infinite variance (t with $1 \lt \nu \le 2$): the average settles, slowly; no usual CLT.

Trap: a calm-looking running mean is not proof that the mean exists; the median still works.

Quick check: you average 10 000 independent standard Cauchy draws. What is the chance the average is beyond ±10?

The average is itself standard Cauchy, so the chance is the same as for one draw: $1 - \tfrac{2}{\pi}\arctan(10) \approx 0.063$. Ten thousand draws bought nothing.

The Central Limit Theorem: why the bell curve is everywhere core

A day of demand is the sum of many small, separate decisions by many customers. A measurement error is the sum of many tiny disturbances. When many small independent pieces are added up, the total piles up in a bell shape, and this happens whatever the shape of each piece. One die is flat. Two dice make a triangle. Three dice already look like a bell.

The LLN told us the average goes to $\mu$. The Central Limit Theorem (CLT) describes the leftover wobble around $\mu$: for large $n$ it follows a Normal curve with spread $\sigma/\sqrt n$.

Three ways to say it:

  • Picture: add up enough independent pieces and a bell appears, even if each piece is flat, lopsided or two-humped.
  • Numbers: the conversion rate over 400 users with $p = 0.1$ is roughly $N(0.1,\ 0.015^2)$, so about 95% of such experiments land within $0.1 \pm 0.03$.
  • Slogan: averages are approximately Normal, even when the data are not.

Add up dice and watch the shape change (exact counting, no simulation).

  1. One die: each of 1–6 has probability $1/6$. Flat.
  2. Two dice: there are $6 \times 6 = 36$ equally likely pairs. Sum 2 happens 1 way, sum 3 two ways, …, sum 7 six ways, …, sum 12 one way: counts $1, 2, 3, 4, 5, 6, 5, 4, 3, 2, 1$. A triangle.
  3. Three dice: $216$ triples. Counting the ways for sums $3$ to $18$ gives $1, 3, 6, 10, 15, 21, 25, 27, 27, 25, 21, 15, 10, 6, 3, 1$. Already bell-shaped.
  4. Check with the Normal curve. For three dice: mean $3 \times 3.5 = 10.5$, variance $3 \times 35/12 = 8.75$, sd $\sqrt{8.75} \approx 2.958$.
  5. Exact $P(\text{sum} \ge 15) = (10 + 6 + 3 + 1)/216 = 20/216 \approx 0.0926$.
  6. Normal approximation (using 14.5 as the cut, because the sum jumps in whole steps): $P\!\left(Z \ge \frac{14.5 - 10.5}{2.958}\right) = P(Z \ge 1.35) \approx 0.088$. Close, with only three dice.

Central Limit Theorem (the classical iid version). Let $X_1, X_2, \dots$ be iid with mean $\mu$ and finite variance $\sigma^2 \gt 0$. Define the standardized mean

$$Z_n \;=\; \frac{\bar X_n - \mu}{\sigma/\sqrt n} \;=\; \frac{S_n - n\mu}{\sigma\sqrt n}.$$

Then for every number $z$,

$$P(Z_n \le z) \;\longrightarrow\; \Phi(z) \quad \text{as } n \to \infty,$$

where $\Phi$ is the standard Normal CDF (Chapter 4.9). This is called convergence in distribution: the CDF of $Z_n$ approaches the CDF of $N(0, 1)$.

  • Practical reading for large $n$: $\bar X_n \approx N(\mu,\ \sigma^2/n)$ and $S_n \approx N(n\mu,\ n\sigma^2)$.
  • For a proportion $\hat p$ (an average of 0/1 values): $\hat p \approx N\big(p,\ p(1-p)/n\big)$.
  • Replacing the unknown $\sigma$ by the sample sd $s$ keeps the same limit (this is what z- and t-tests in Chapter 5.6 rely on).
  • "Standardized" means: subtract the mean, divide by the standard deviation, so the result has mean 0 and sd 1 (the z-score idea, Chapter 4.18).
Why do we need it?

It lets us make probability statements about an average without knowing the shape of the data. Without it, every confidence interval and every test would need the exact distribution of the data, which we almost never know.

Where is it used?

z-tests and t-tests, confidence intervals for means and proportions, A/B test power and sample-size formulas, Monte Carlo error bars, the Normal approximation to the Binomial and Poisson, and the "many small causes" argument behind Normal noise in regression and forecasting.

How is it used?

Treat $\bar x$ as Normal with mean $\mu$ and sd $s/\sqrt n$. Then $\bar x \pm 1.96\, s/\sqrt n$ covers $\mu$ about 95% of the time, and $(\bar x - \mu_0)/(s/\sqrt n)$ can be compared with the standard Normal. First check that $n$ is large enough for your data's skewness (section 7).

1 die: flat 2 dice: triangle 3 dice: a bell already sum 1 … 6 sum 2 … 12 sum 3 … 18 (orange: Normal curve)
Exact distributions of the sum of one, two and three fair dice (bar heights proportional to the number of ways). Adding independent pieces turns a flat shape into a bell very quickly.

Pick a population, then raise $n$ (the number of draws in each average). Left: a histogram of 2 000 averages, with the Normal curve $N(\mu, \sigma^2/n)$ the CLT predicts (orange). Right: a Q-Q plot of the standardized averages against the Normal (Chapter 4.17): points on the diagonal mean "Normal". Notice: Fair die looks Normal by $n \approx 5$; Exponential needs more; Bernoulli p = 0.05 needs far more (watch the bent Q-Q line). Cauchy never becomes Normal. Turn off Zoom to the means to see the pile also get narrower as $n$ grows.

The bars are the exact distribution of the sum of $n$ dice (computed by counting, no randomness). The orange curve is the Normal curve with the same mean and variance. Raise $n$ for the fair die: by $n = 5$ the bars and the curve are almost identical. Now pick the loaded die (mostly 1s and 2s, so it is lopsided): the bars lean to one side and need many more dice. The readout gives the largest gap between the exact and Normal CDFs: at $n = 30$ the loaded die is still further from Normal than the fair die at $n = 3$.

"The CLT says sums are exactly Normal once $n$ is big."

It says the distribution gets closer and closer to Normal. At any finite $n$ there is an error, and it is largest for skewed or heavy-tailed data (sections 7 and 8).

"The CLT works for any data as long as $n$ is large."

It needs (roughly) independent pieces and a finite variance. The Cauchy option in the machine above never becomes Normal, however large $n$ is.

In your A/B framework, any Normal-approximation shortcut for a conversion rate or a difference of rates ($\hat p \approx N(p, p(1-p)/n)$) is the CLT. Your Beta-Binomial model does not need it: its posterior is exact for any $n$, which matters most for small samples and rare conversions. A cousin of the CLT (the Bernstein–von Mises theorem) says that, under regularity conditions, posteriors also become approximately Normal as data grow; this is one reason a Gaussian guide in SVI tends to fit well with lots of data. In your forecasting model, a Normal likelihood for daily noise is often justified by the CLT story "many small independent influences add up", and it breaks when one big shock (a promotion, an outage) dominates a day, which is when Student-t helps.

CLT: iid, mean $\mu$, finite $\sigma^2 \gt 0$ ⇒ $Z_n = \dfrac{\bar X_n - \mu}{\sigma/\sqrt n} \to N(0, 1)$ in distribution.

Use: $\bar X_n \approx N(\mu, \sigma^2/n)$, $S_n \approx N(n\mu, n\sigma^2)$, $\hat p \approx N(p, p(1-p)/n)$.

Trap: approximate, not exact; needs independence and finite variance.

Quick check: daily orders have mean 20 and sd 8, independent across days. Roughly, what is the chance that the total over 64 days exceeds 1 344?

Sum: mean $64 \times 20 = 1280$, sd $8\sqrt{64} = 64$. $1344 = 1280 + 64$, one sd above the mean, so $P \approx P(Z \gt 1) \approx 0.16$.

What exactly becomes Normal, and what the CLT does not say core

Three different things are easy to mix up:

  • the population: the shape of single values (for example, waiting times: mostly short, a few long);
  • one sample: the $n$ values you actually collected. Its histogram looks like the population, and it looks more like it as $n$ grows;
  • the sampling distribution of $\bar X$: how the average would vary over many imagined repeats of the whole data collection.

The CLT is only about the third one. Collecting more data makes your sample a sharper picture of the (possibly skewed) population. It never turns the data into a bell.

Three ways to say it:

  • Picture: the histogram of your data keeps the population's shape; only the histogram of many averages becomes a bell.
  • Numbers: 1000 exponential waiting times still have skewness about 2, but the average of those 1000 values has skewness $2/\sqrt{1000} \approx 0.06$.
  • Slogan: the CLT normalizes averages, not data.

Exponential waiting times have skewness $\gamma = 2$ (a long right tail). For an average of $n$ iid values, the skewness is $\gamma/\sqrt n$ (section 7 explains why).

  1. The data themselves: skewness 2, whether you collect 4 values or 4 million. A bigger sample estimates this 2 more precisely; it does not shrink it.
  2. Average of $n = 4$: skewness $2/\sqrt 4 = 1$. Still visibly lopsided.
  3. Average of $n = 30$: $2/\sqrt{30} \approx 0.37$. Mildly lopsided.
  4. Average of $n = 1000$: $2/\sqrt{1000} \approx 0.063$. Practically symmetric, practically Normal.
  5. And the average itself, unscaled, collapses toward $\mu$ (spread $\sigma/\sqrt n \to 0$). The bell only stays visible if you zoom in by $\sqrt n$: that is why the CLT is stated for the standardized mean $Z_n$.

What converges: the distribution (the CDF) of the standardized mean $Z_n = (\bar X_n - \mu)/(\sigma/\sqrt n)$ converges to the standard Normal CDF. This is a statement about the sampling distribution of one number computed from the data.

What the CLT does not say:

  • It does not say the data, the population or the residuals become Normal.
  • It does not say how large $n$ must be. That depends on skewness and tails (section 7); "$n \ge 30$" is a rule of thumb, not part of the theorem.
  • It does not promise accuracy far out in the tails: the approximation is best near the centre, and relative errors for tiny tail probabilities can be large.
  • It does not cover every statistic: the maximum of a sample never becomes Normal; the median has its own CLT with a different variance.
  • It does not apply without its assumptions: roughly independent pieces and a finite variance (section 8).

(The fact that the sample histogram approaches the population shape is a different result, the LLN applied to the CDF, known as the Glivenko–Cantelli theorem.)

Why do we need it?

Mixing up "the average is approximately Normal" with "the data are approximately Normal" leads to wrong models: a Normal likelihood for skewed counts, or Normal prediction intervals for heavy-tailed demand. Knowing what the CLT covers keeps both uses honest.

Where is it used?

Choosing a likelihood (Chapter 4.11): the CLT justifies Normal intervals for averages (A/B metric means), not a Normal model for single observations (a day's demand, one user's revenue). Residual checks with Q-Q plots and KDE exist exactly because the data need not be Normal.

How is it used?

Ask: "is the quantity I'm treating as Normal an average of many independent pieces?" If yes (a mean metric, a Monte Carlo average), the CLT helps. If it is a single observation or a residual, check its distribution directly (Chapter 4.17).

1 · the population 2 · one sample, n = 30 3 · many averages of n = 30 μμμ shape of single values still skewed, like the population bell-shaped, narrow (σ/√30) the CLT is only about panel 3
All three panels share one horizontal scale. A sample (middle) is a rough copy of the population (left) and stays skewed. Only the distribution of the average over many repeats (right) is bell-shaped, and it is much narrower.

Left: a histogram of one sample of $n$ values. Right: a histogram of 1 000 averages, each over $n$ values, with the CLT's Normal curve. Drag $n$ from 1 up to 500. Notice: the left histogram becomes a cleaner and cleaner picture of the skewed population (its skewness stays near 2 for the exponential), while the right one becomes a symmetric bell. More data never made the left side Normal.

"I have 50 000 rows, so by the CLT my data are Normal and a Normal likelihood is fine."

The CLT says nothing about individual rows. 50 000 skewed values are 50 000 skewed values. Check the data (or the residuals) with a histogram, KDE or Q-Q plot.

"The CLT says the sample mean converges to $\mu$."

That is the LLN. The CLT describes the fluctuations of the mean around $\mu$, magnified by $\sqrt n$: their shape is Normal and their size is $\sigma/\sqrt n$.

"The Central Limit Theorem says that with enough data, everything is Normal."

It says the sampling distribution of a standardized sum or average of many independent, finite-variance pieces approaches a Normal.

"$n \ge 30$ is enough for the CLT."

Thirty is a rule of thumb for mildly skewed data. A metric with rare events or heavy tails can need thousands (section 7), and data with infinite variance never get there.

Model answer: "The CLT is about the sampling distribution, not the data. If $X_1, \dots, X_n$ are iid with finite variance $\sigma^2$, then $\sqrt n(\bar X - \mu)/\sigma$ converges in distribution to $N(0,1)$, so $\bar X$ is approximately $N(\mu, \sigma^2/n)$. It doesn't make the data Normal, it doesn't say how big $n$ must be (that depends on skewness and tails), and it fails with strong dependence or infinite variance."

Three objects: population · one sample · sampling distribution of $\bar X$. The CLT is about the third only.

Data keep their skewness $\gamma$; the average's skewness is $\gamma/\sqrt n$.

Not in the CLT: Normal data, a safe $n$, accurate far tails, the maximum, dependent or infinite-variance data.

Quick check: your forecasting residuals are skewed, and you have 3 years of daily data. Does the CLT make a Normal likelihood for the residuals appropriate?

No. Each residual is a single observation, not an average, so the CLT says nothing about its shape. With more days you just see the skewness more sharply. Look at a Q-Q plot and consider a transformation or a different likelihood.

How large must $n$ be? Skewness slows the CLT

The bell forms quickly when each piece is already symmetric: a sum of five fair dice already looks Normal. It forms slowly when the pieces are lopsided. A conversion with $p = 1\%$ is a lopsided piece: almost always 0, rarely 1. You need very many of them before their average looks like a symmetric bell.

The reason is simple arithmetic: the skewness of an average shrinks only like $1/\sqrt n$. Start very skewed and you need a big $n$ to get close to zero.

Three ways to say it:

  • Picture: lopsided pieces need many more partners before they blend into a symmetric bell.
  • Numbers: to bring the average's skewness down to 0.1, exponential data need $n = 400$ but Bernoulli(0.05) data need about $n = 1700$.
  • Slogan: the more skewed the data, the bigger the $n$.

Conversion rate $p = 0.1$, $n = 400$ users. How good is the Normal approximation for "the observed rate is at least 2 standard errors above the truth"?

  1. Standard error: $\sqrt{0.1 \times 0.9/400} = 0.3/20 = 0.015$. Two SEs above: $0.1 + 0.03 = 0.13$, i.e. 52 or more conversions.
  2. Normal approximation: $P(Z \ge 2) \approx 0.0228$.
  3. Exact Binomial: $P(K \ge 52) \approx 0.0311$. The Normal curve underestimates the right tail by about a third.
  4. The left tail, $P(K \le 28)$ (2 SEs below), is $\approx 0.0235$: close to $0.0228$. The error is lopsided, because the data are.
  5. Why: one user has skewness $(1-2p)/\sqrt{p(1-p)} = 0.8/0.3 \approx 2.67$; the average of 400 has skewness $2.67/\sqrt{400} \approx 0.13$, small but not zero.

For an average of $n$ iid values with skewness $\gamma$ and excess kurtosis $\kappa$ (Chapter 4.14):

$$\text{skew}(\bar X_n) = \frac{\gamma}{\sqrt n}, \qquad \text{excess kurtosis}(\bar X_n) = \frac{\kappa}{n}.$$

So skewness is the slow part. To reach a target skewness $t$, you need roughly $n \approx (\gamma/t)^2$.

Populationskewness $\gamma$$n$ for skew$(\bar X) = 0.1$
Fair die, Normal, bimodal symmetric0any $n$ (already symmetric)
Poisson, $\lambda = 2$$1/\sqrt 2 \approx 0.71$50
Exponential2400
Bernoulli, $p = 0.05$$0.9/\sqrt{0.0475} \approx 4.13$≈ 1 705
Bernoulli, $p = 0.01$$0.98/\sqrt{0.0099} \approx 9.85$≈ 9 700

A sharper statement is the Berry–Esseen bound: the largest gap between the CDF of $Z_n$ and $\Phi$ is at most $C\,\rho/(\sigma^3\sqrt n)$, where $\rho = E|X-\mu|^3$ and $C$ is a constant below 0.48. Again: error $\propto 1/\sqrt n$, larger for lopsided, heavy-tailed data.

Rules of thumb (only rules of thumb): $n \ge 30$ for mildly skewed data; for a proportion, $np \ge 10$ and $n(1-p) \ge 10$ (some books say 5).

Why do we need it?

"Use the Normal approximation" is only safe if $n$ is big enough for this data. Rare-event and heavy-tailed metrics can make a textbook test give wrong p-values and intervals with wrong coverage even with thousands of users.

Where is it used?

A/B tests on low conversion rates, revenue per user (mostly zeros plus a few large orders), click-through rates on ads, error rates in monitoring, and any sample-size calculation that relies on the Normal approximation (Chapter 5.7).

How is it used?

Estimate the skewness of a single observation, compute $\gamma/\sqrt n$, and if it is not small, use an exact method (Binomial, Beta posterior), a Wilson interval (Chapter 5.8), a bootstrap, a transformation, or just more data.

Blue bars: the exact distribution of the observed rate $\hat p = K/n$ (from the Binomial). Orange: the CLT's Normal curve. The red region is "more than 2 SEs above the truth", which the Normal curve says has probability 0.0228. Set $p = 0.02$ and $n = 100$: the bars are visibly lopsided and the Normal curve even spills below 0. Raise $n$ and watch the exact tail probabilities approach 0.0228. Then try $p = 0.5$: symmetric pieces, fast convergence.

"I have 5 000 users, so the Normal approximation is fine."

With $p = 0.001$ you expect only 5 conversions: $np = 5$, skewness of the rate about $1/\sqrt 5 \approx 0.45$. What matters is the number of events and the skewness, not the headline $n$.

"If the middle of the histogram matches the bell, the tails match too."

The centre converges first. Tail probabilities (the ones that decide p-values and risk) can still be off by a large relative amount.

In an A/B framework like yours, conversion metrics with small rates and revenue-like metrics with heavy right tails are exactly the slow cases. This is one practical advantage of your Beta-Binomial model for conversions: the posterior is exact for any $n$ and needs no Normal approximation. For continuous, skewed metrics, the Student-t (or a transformed scale) is safer than trusting the CLT at moderate $n$.

skew$(\bar X_n) = \gamma/\sqrt n$; kurtosis$(\bar X_n) = \kappa/n$. Needed $n \approx (\gamma/\text{target})^2$.

Berry–Esseen: CDF error $\le C\rho/(\sigma^3\sqrt n)$, $C \lt 0.48$.

Rules of thumb only: $n \ge 30$; $np \ge 10$ and $n(1-p) \ge 10$. Rare events and heavy tails need far more.

Quick check: daily order counts are Poisson with $\lambda = 2$ (skewness $1/\sqrt 2$). What is the skewness of a 50-day average? And of a single day's count if you collect 50 days?

The 50-day average: $0.707/\sqrt{50} = 0.1$. A single day's count is still Poisson(2) with skewness $0.707$, however many days you collect.

The assumptions: independence and finite variance core

Averaging works because independent errors cancel: one day is too high, another too low, and in the average they partly wash out. If the errors move together (a busy week is busy every day; one heavy user creates many sessions), they do not cancel, and the average stays much noisier than $\sigma/\sqrt n$ suggests.

And the bell shape needs every piece to be "small" compared with the total. If single pieces can be enormous (infinite variance), one piece can dominate the whole sum, and no bell appears.

Three ways to say it:

  • Picture: a crowd of independent people walking randomly ends up near the start; a crowd marching in step does not.
  • Numbers: 50 days whose surprises carry over strongly from one day to the next ($\varphi = 0.8$) give an average about $2.9$ times noisier than $\sigma/\sqrt{50}$: they are worth only about 6 independent days.
  • Slogan: no independence, no $\sqrt n$; no finite variance, no bell.

The average of just two values, each with variance $\sigma^2$, whose correlation is $\rho$ (covariance $\rho\sigma^2$; covariance and correlation are taught in Chapter 4.15):

  1. $Var\!\left(\tfrac{X_1 + X_2}{2}\right) = \tfrac14\big(Var(X_1) + Var(X_2) + 2\,Cov(X_1, X_2)\big) = \tfrac14(\sigma^2 + \sigma^2 + 2\rho\sigma^2) = \tfrac{\sigma^2(1+\rho)}{2}$.
  2. Independent ($\rho = 0$): $\sigma^2/2$, the usual $\sigma^2/n$.
  3. $\rho = 0.8$: $0.9\,\sigma^2$. Averaging barely helped.
  4. $\rho = 1$ (a copy of the same value): $\sigma^2$. Two copies are worth exactly one value.
  5. The "worth" in independent values (the effective sample size) is $n_{\text{eff}} = \sigma^2/Var(\bar X) = 2/(1+\rho)$: $2$, $1.11$, and $1$ in the three cases.

Dependence. For a series whose values all have variance $\sigma^2$ and whose correlation at lag $k$ (between values $k$ steps apart) is $\rho_k$,

$$Var(\bar X_n) = \frac{\sigma^2}{n}\Big[1 + 2\sum_{k=1}^{n-1}\Big(1 - \frac{k}{n}\Big)\rho_k\Big], \qquad n_{\text{eff}} = \frac{n}{1 + 2\sum_{k=1}^{n-1}(1 - k/n)\rho_k}.$$

For an AR(1) series ($\rho_k = \varphi^k$; each day carries over a fraction $\varphi$ of yesterday's surprise; Chapter 7.3) the bracket is about $(1+\varphi)/(1-\varphi)$ for large $n$: $\varphi = 0.5$ gives 3, $\varphi = 0.8$ gives 9. Versions of the CLT exist for such weakly dependent series, but with this larger variance: the bell may survive, the naive $\sigma/\sqrt n$ does not.

Finite variance. If $Var(X) = \infty$ (Cauchy, Student-t with $\nu \le 2$), the classical CLT fails. Suitably rescaled sums can converge to other, heavy-tailed "stable" distributions instead of the Normal.

Identically distributed can be relaxed: the CLT still holds for independent pieces with different distributions, as long as no single piece dominates the total variance (the Lindeberg condition).

Why do we need it?

Ignoring dependence gives standard errors that are too small, confidence intervals that are too narrow and far too many "significant" results. It is one of the most common ways real analyses go wrong.

Where is it used?

A/B tests where the randomization unit (user) differs from the analysis unit (session, page view) (Chapter 5.10), clustered and hierarchical data, autocorrelated time-series residuals, and MCMC chains, whose effective sample size (ESS, reported by NumPyro's diagnostics) is built on the same idea, $n/(1 + 2\sum_k \rho_k)$ (Chapter 6.10).

How is it used?

Ask "what is the independent unit?" and compute averages and SEs at that level. For series, estimate the autocorrelations (ACF), compute $n_{\text{eff}}$, and use it in place of $n$. For heavy tails, check the largest values before trusting any average.

assumptionif it fails…example in your workwhat to do independentpieces σ/√n too small,intervals too narrow sessions of one user;autocorrelated days right unit, n_eff,model dependence finitevariance no bell; averagesjump at huge values ratio metrics; Student-tnoise with ν ≤ 2 median, trimming,model the tails samedistribution usually fine, unlessone piece dominates segments of verydifferent sizes check no singlegiant contributor the first two are the ones that break real analyses
The CLT's assumptions, what goes wrong when each fails, where you might meet the failure, and the usual fix.

Each run is a series of 50 days whose surprises carry over with strength $\varphi$ (an AR(1) series with variance 1). Left: one example series and its average (purple). Right: the averages of 400 such series, with the naive curve $N(0, 1/50)$ that assumes independence (orange) and the correct one (green). Start at $\varphi = 0$: the curves agree. Raise $\varphi$ to 0.8: the real averages spread about 3 times wider than the naive curve. Try a negative $\varphi$: alternating days cancel even better than independent ones.

"We logged 1 000 000 sessions, so $n = 1\,000\,000$."

If sessions come from 50 000 users and a user's sessions resemble each other, the independent unit is the user. Treating sessions as independent makes the standard error too small (Chapter 5.10).

"My residuals look Normal, so my uncertainty is right."

Normal-looking residuals can still be autocorrelated. Then any uncertainty that assumes independent days is too narrow. Check the residual ACF (Chapter 7.17).

In your A/B framework, the randomization unit must also be the unit of independence: if users are randomized but metrics are computed per session, sessions from one user are correlated and the effective sample size is closer to the number of users. Segments in a hierarchical model are another form of dependence: users in the same segment share a segment effect, which is exactly what the hierarchical structure models (Chapter 6.5). In your forecasting model, the noise $\epsilon_t$ is, in a Prophet-style design, treated as independent from day to day unless you added a residual-correlation term (check your code). If the real residuals are autocorrelated, the days are not independent pieces, and uncertainty computed as if they were can be too narrow. Posterior draws from MCMC are autocorrelated too: the ESS that diagnostics report is this same $n_{\text{eff}}$ idea.

$Var(\bar X_n) = \frac{\sigma^2}{n}[1 + 2\sum(1 - k/n)\rho_k]$; AR(1): factor $\approx (1+\varphi)/(1-\varphi)$. $n_{\text{eff}} = n/\text{factor}$.

Infinite variance: no classical CLT (sums can go to heavy-tailed stable laws).

Trap: count independent units, not rows. Correlation inflates the true SE.

Quick check: 100 days with AR(1) carry-over $\varphi = 0.5$. Roughly how many independent days are they worth?

The factor is about $(1+0.5)/(1-0.5) = 3$, so $n_{\text{eff}} \approx 100/3 \approx 33$ days. (The exact finite-$n$ factor is a little below 3.)

Monte Carlo: the LLN and the CLT inside your Bayesian tools

Some quantities are too hard to compute with a formula, but easy to simulate. "What is the probability that variant B's true rate beats A's, given the data?" Draw a plausible pair of rates from the posterior, check who wins, repeat thousands of times, and report the fraction of wins. This recipe is called Monte Carlo.

The LLN guarantees the fraction approaches the true probability as you draw more. The CLT tells you how far off you probably still are: about $\text{sd}/\sqrt S$ for $S$ independent draws.

Three ways to say it:

  • Picture: throw many random darts and count the hits; the hit rate settles (LLN), and its wobble shrinks like $1/\sqrt S$ (CLT).
  • Numbers: with 4 000 draws, a probability near 0.84 is estimated to within about ±0.011 (95%).
  • Slogan: simulate, average, and report the ± too.

Variant A: 50 conversions out of 500. Variant B: 60 out of 500. With uniform Beta(1, 1) priors the posteriors are $\theta_A \sim Beta(51, 451)$ and $\theta_B \sim Beta(61, 441)$ (conjugate updating is taught in Chapter 6.3).

  1. Draw $S = 4000$ pairs $(\theta_A^{(s)}, \theta_B^{(s)})$ from the two posteriors.
  2. For each pair, write $I_s = 1$ if $\theta_B^{(s)} \gt \theta_A^{(s)}$, else $0$.
  3. Estimate: $\hat P = \frac{1}{S}\sum_s I_s$. A typical run gives about $0.84$. (Numerical integration gives the exact value $0.8429$.)
  4. Each $I_s$ is a 0/1 value with variance $P(1-P) \approx 0.843 \times 0.157 \approx 0.132$.
  5. Monte Carlo standard error: $\sqrt{0.132/4000} \approx 0.0058$. So report about $0.84 \pm 0.011$ (two SEs).
  6. To get $\pm 0.001$ (95%) you would need $S \approx 1.96^2 \times 0.132/0.001^2 \approx 509\,000$ draws: the $\sqrt S$ law makes extra decimals expensive.

To estimate $E[g(\theta)]$ where $\theta \sim p$, draw $\theta^{(1)}, \dots, \theta^{(S)}$ iid from $p$ and compute

$$\hat I_S = \frac{1}{S}\sum_{s=1}^{S} g\big(\theta^{(s)}\big).$$
  • LLN: $\hat I_S \to E[g(\theta)]$ as $S \to \infty$ (if $E|g(\theta)| \lt \infty$).
  • CLT: $\hat I_S \approx N\big(E[g],\ Var(g)/S\big)$, so the Monte Carlo standard error is $\text{MCSE} = \text{sd}(g)/\sqrt S$, estimated by the sd of the $g(\theta^{(s)})$ values.
  • A probability is the special case $g = $ an indicator (1 if the event happens, 0 otherwise): $\text{MCSE} = \sqrt{\hat P(1-\hat P)/S}$.
  • With autocorrelated draws (MCMC), replace $S$ by the effective sample size: $\text{MCSE} = \text{sd}/\sqrt{\text{ESS}}$.

Monte Carlo error is only about the simulation. It does not include the uncertainty of the model itself or of its assumptions.

Why do we need it?

Most posterior quantities (probabilities of beating a rival, expected lifts, forecast quantiles) have no closed formula once models get realistic. Simulation plus averaging gets them, and the CLT tells you how many draws are enough.

Where is it used?

$P(\theta_B \gt \theta_A \mid D)$ and expected-loss decisions in A/B tools (Chapter 6.4), posterior means and intervals from NUTS or SVI draws, the Monte Carlo ELBO estimate in SVI (Chapter 6.12), forecast fan charts built from predictive draws, and the bootstrap.

How is it used?

Draw, compute the quantity per draw, average. Report the MCSE ($\text{sd}/\sqrt S$, or $\text{sd}/\sqrt{\text{ESS}}$ for MCMC). If the MCSE is too big for the decision, draw more; to halve it, draw four times as many.

A converted 50 of 500 users; set how many of B's 500 users converted. Three independent simulation runs (coloured lines) each draw up to 2 000 posterior pairs and track the running fraction of draws where $\theta_B \gt \theta_A$. The green line is the exact answer (by numerical integration); the purple band is the CLT's ±1.96 MCSE. Notice: all lines enter the band and the band narrows like $1/\sqrt S$. Press New sample: the early part changes a lot, the end barely. Set B to 50: the answer is 0.5 and the band is widest (a 50/50 indicator has the largest variance).

"The tool printed $P(B \gt A) = 0.843$, so that is the probability."

It is a Monte Carlo estimate with its own error (a 95% range of about ±0.011 at 4 000 draws), on top of the fact that it is only as good as the model and prior.

"4 000 MCMC draws give a Monte Carlo error of $\text{sd}/\sqrt{4000}$."

MCMC draws are autocorrelated. Use the effective sample size: $\text{sd}/\sqrt{\text{ESS}}$, which can be much larger (Chapter 6.10).

If your A/B framework computes its posterior decision $P(\theta_A \gt \theta_B \mid D)$ from posterior draws (the usual way with SVI), it works exactly like this widget: a fraction of draws. The LLN makes it converge; the CLT gives it an error bar $\sqrt{\hat P(1-\hat P)/S}$. Decisions near a threshold (say 0.95) need enough draws that the MCSE is small compared with the distance to the threshold. In both projects, SVI optimizes a Monte Carlo estimate of the ELBO, so the ELBO trace is noisy, with noise shrinking like $1/\sqrt{\text{particles}}$. This noise is the reason a training loop that compares against the best ELBO so far uses a patience window rather than stopping at the first non-improvement (Chapter 6.14). Forecast intervals summarized from predictive draws carry Monte Carlo error too.

"We ran 1 000 posterior samples, so our probability is accurate to 0.001."

The Monte Carlo error scales like $1/\sqrt S$, not $1/S$: with $S = 1000$ and $P$ near 0.9 it is about $\sqrt{0.09/1000} \approx 0.0095$.

Model answer: "A Monte Carlo estimate is an average, so the LLN makes it consistent and the CLT gives its standard error, $\text{sd}/\sqrt S$, or $\text{sd}/\sqrt{\text{ESS}}$ for correlated MCMC draws. I report it, and I make sure it's small relative to how close the estimate is to my decision threshold."

$\hat I_S = \frac1S\sum g(\theta^{(s)}) \to E[g]$ (LLN); $\text{MCSE} = \text{sd}(g)/\sqrt S$ (CLT).

For a probability: $\sqrt{\hat P(1-\hat P)/S}$. For MCMC: use ESS instead of $S$.

Trap: 10× accuracy costs 100× draws; MCSE ignores model error.

Quick check: from 2 500 independent draws you estimate $P = 0.5$. What is the Monte Carlo standard error?

$\sqrt{0.5 \times 0.5/2500} = \sqrt{0.0001} = 0.01$. A 95% Monte Carlo interval is about $0.5 \pm 0.02$.

Recap, cheat sheet and practice

  • The sample mean $\bar X_n$ is a random variable: centre $\mu$, spread $\sigma/\sqrt n$ (the standard error). Four times the data, half the error.
  • LLN: with iid draws and a finite mean, $\bar X_n \to \mu$. It works by dilution, not compensation. The weak LLN follows from Chebyshev: $P(|\bar X_n - \mu| \ge \varepsilon) \le \sigma^2/(n\varepsilon^2)$.
  • No mean, no LLN: the average of $n$ Cauchy draws is again Cauchy. The median still works.
  • CLT: with iid draws and a finite variance, the standardized mean $(\bar X_n - \mu)/(\sigma/\sqrt n)$ converges in distribution to $N(0, 1)$. So $\bar X_n \approx N(\mu, \sigma^2/n)$.
  • The CLT is about the sampling distribution of the average, not the data. It does not give a safe $n$: skewness of the average is $\gamma/\sqrt n$, so rare events and heavy tails need much larger $n$.
  • Dependence inflates the true variance of the mean ($n_{\text{eff}} \lt n$); infinite variance kills the bell.
  • Monte Carlo estimates (posterior probabilities, means, ELBOs) are averages: LLN for correctness, CLT for the error bar $\text{sd}/\sqrt S$ (or $\text{sd}/\sqrt{\text{ESS}}$).

Cheat sheet

IdeaFormulaNeeds
Mean of the average$E[\bar X_n] = \mu$same mean for every draw
Spread of the average (SE)$SD(\bar X_n) = \sigma/\sqrt n$uncorrelated draws, finite $\sigma$
Weak LLN$P(|\bar X_n - \mu| \gt \varepsilon) \to 0$iid, $E|X| \lt \infty$
Chebyshev version$P(|\bar X_n - \mu| \ge \varepsilon) \le \sigma^2/(n\varepsilon^2)$finite $\sigma^2$
CLT$\dfrac{\bar X_n - \mu}{\sigma/\sqrt n} \to N(0,1)$; $\bar X_n \approx N(\mu, \sigma^2/n)$iid (or weakly dependent), finite $\sigma^2$
Proportion$\hat p \approx N(p,\ p(1-p)/n)$rule of thumb: $np, n(1-p) \ge 10$
Skewness of the average$\gamma/\sqrt n$ (kurtosis $\kappa/n$)iid
Dependent data$Var(\bar X_n) = \frac{\sigma^2}{n}[1 + 2\sum(1 - \frac kn)\rho_k]$; AR(1) ≈ $\frac{\sigma^2}{n}\frac{1+\varphi}{1-\varphi}$stationary series
Monte Carlo error$\text{MCSE} = \text{sd}/\sqrt S$; probability: $\sqrt{\hat P(1-\hat P)/S}$independent draws (else use ESS)
Code it · Python
import numpy as np
from scipy import stats

rng = np.random.default_rng(0)

# 1) The sample mean is random: centre p, spread sigma/sqrt(n)
p, n = 0.1, 400
rates = rng.binomial(n, p, size=10_000) / n     # 10 000 repeated experiments of 400 users
print(rates.mean(), rates.std(ddof=1))           # 0.0997  0.0152   (theory: 0.1 and 0.3/sqrt(400) = 0.015)

# 2) Law of Large Numbers: running average of die rolls
rolls = rng.integers(1, 7, size=100_000)
running = np.cumsum(rolls) / np.arange(1, rolls.size + 1)
print(running[[9, 999, 99_999]])                 # [3.9  3.46  3.50135]  closer and closer to 3.5

# 3) Weak LLN: Chebyshev's bound vs the exact chance of missing by eps
n, eps = 1000, 0.03
bound = p * (1 - p) / (n * eps**2)
exact = stats.binom.cdf(70, n, p) + stats.binom.sf(129, n, p)   # K <= 70 or K >= 130
print(bound, exact)                              # 0.1  0.0019   (the bound is ~50x too pessimistic)

# 4) CLT with skewed data: averages of 30 exponential values
means = rng.exponential(1.0, size=(20_000, 30)).mean(axis=1)
print(stats.skew(means), 2 / np.sqrt(30))        # 0.39  0.365   skewness shrinks like 1/sqrt(n)
print(np.mean(means > 1 + 2 / np.sqrt(30)))      # 0.0326        the Normal approximation says 0.0228

# 5) Cauchy: averaging does not help
c = rng.standard_cauchy(size=(2_000, 1_000)).mean(axis=1)
print(np.mean(np.abs(c) > 10))                   # 0.0625   same as ONE draw: 1 - (2/pi)*arctan(10) = 0.063

# 6) Dependence: AR(1) with phi = 0.8 inflates the spread of a 50-day average
phi, n, reps = 0.8, 50, 20_000
x = np.empty((reps, n))
x[:, 0] = rng.standard_normal(reps)
for t in range(1, n):
    x[:, t] = phi * x[:, t - 1] + np.sqrt(1 - phi**2) * rng.standard_normal(reps)
print(x.mean(axis=1).std(), 1 / np.sqrt(n))      # 0.407  0.141   (theory: sqrt(8.2/50) = 0.405)

# 7) Monte Carlo: P(theta_B > theta_A | D) from posterior draws, with its standard error
S = 4000
theta_a = rng.beta(51, 451, S)                   # A: 50/500 with a Beta(1, 1) prior
theta_b = rng.beta(61, 441, S)                   # B: 60/500
est = np.mean(theta_b > theta_a)
mcse = np.sqrt(est * (1 - est) / S)
print(est, mcse)                                 # 0.848  0.0057   (exact answer 0.8429, well within 2 MCSE)
Test yourself

1. According to the Central Limit Theorem, what becomes approximately Normal as $n$ grows?

The CLT is about how the average would vary over repeated samples. The data keep the population's shape, however much you collect.

2. You want to halve the standard error of a mean. You need…

SE $= \sigma/\sqrt n$. Replacing $n$ by $4n$ divides the SE by $\sqrt 4 = 2$.

3. You average 1 000 independent standard Cauchy draws. The average…

The Cauchy has no mean (and no variance), so neither the LLN nor the CLT applies. The average of $n$ standard Cauchy draws is again standard Cauchy.

4. A fair coin has landed tails 8 times in a row. The probability that the next flip is heads is…

Flips are independent. The LLN works by dilution: past surpluses become a negligible share of a growing total; they are never actively corrected.

5. A 0/1 metric with $\sigma^2 = 0.25$, $n = 2500$ users, tolerance $\varepsilon = 0.02$. What does Chebyshev guarantee about $P(|\bar X - p| \ge 0.02)$?

$\sigma^2/(n\varepsilon^2) = 0.25/(2500 \times 0.0004) = 0.25/1 = 0.25$. It is an upper bound, not the probability (for $p = 0.5$ the true value is about 0.048).

6. Daily residuals follow an AR(1) with $\varphi = 0.8$. For a 100-day average, compared with the naive $\sigma/\sqrt{100}$, the true standard error is…

The variance is inflated by about $(1+\varphi)/(1-\varphi) = 9$ (8.6 exactly at $n = 100$), so the standard error grows by about $\sqrt 9 = 3$. Nine is the factor for the variance, not the SE.

Practice problems

A. Revenue per visitor has mean 2.5 and sd 10 (in some currency unit). For 10 000 independent visitors, what is the standard error of the average revenue, and a rough 95% range for it?

SE $= 10/\sqrt{10\,000} = 10/100 = 0.1$. By the CLT, the average is roughly $N(2.5, 0.1^2)$, so about 95% of the time it lands within $2.5 \pm 1.96 \times 0.1 = 2.5 \pm 0.196$. Caution: revenue is very skewed (mostly zeros, a few big orders), so check that $n$ is large enough for its skewness before trusting the Normal shape (section 7).

B. How many independent draws does Chebyshev say guarantee $P(|\bar X - \mu| \ge 0.1\sigma) \le 0.01$? What does the CLT suggest?

Chebyshev: $n \ge \sigma^2/(\delta\varepsilon^2) = \sigma^2/(0.01 \times 0.01\sigma^2) = 10\,000$. CLT: we need $0.1\sigma \ge 2.576\,\sigma/\sqrt n$ (2.576 is the Normal value leaving 1% in the two tails), so $\sqrt n \ge 25.76$, $n \ge 664$. The guarantee costs about 15 times more data, but holds for any distribution.

C. Explain to an interviewer why "with $n \ge 30$ the CLT makes everything Normal" is wrong.

Three problems. (1) The CLT is about the sampling distribution of an average or sum, not the data; individual values and residuals keep their own shape. (2) "30" is only a rule of thumb: the skewness of the average is $\gamma/\sqrt n$, so a metric like Bernoulli(0.01) (skewness ≈ 9.85) needs thousands of observations before the average looks Normal. (3) The theorem needs (roughly) independent draws and a finite variance; with autocorrelation the $\sigma/\sqrt n$ is wrong, and with infinite variance (Cauchy, t with $\nu \le 2$) there is no Normal limit at all.

D. A conversion rate of $p = 0.02$ measured on $n = 500$ users. What is the skewness of $\hat p$? Does the $np \ge 10$ rule pass? How many users would bring the skewness down to 0.1?

One user: $\gamma = (1 - 2p)/\sqrt{p(1-p)} = 0.96/\sqrt{0.0196} = 0.96/0.14 \approx 6.86$. Average of 500: $6.86/\sqrt{500} \approx 0.31$. $np = 10$: the rule just passes, yet the skewness is still sizeable. For skewness 0.1: $n \approx (6.86/0.1)^2 \approx 4\,700$.

E. Two measurements of the same thing have correlation $\rho = 0.5$ and the same variance $\sigma^2$. What is the variance of their average, and how many independent measurements is the pair worth?

$Var = \frac14(\sigma^2 + \sigma^2 + 2 \times 0.5\sigma^2) = \frac14 \times 3\sigma^2 = 0.75\sigma^2$. Two independent measurements would give $0.5\sigma^2$. $n_{\text{eff}} = \sigma^2/(0.75\sigma^2) = 2/(1 + 0.5) \approx 1.33$.

F. From $S = 1000$ independent posterior draws you estimate $P(\theta_B \gt \theta_A \mid D) = 0.97$. Give its Monte Carlo standard error. How many draws would make the MCSE 0.001?

MCSE $= \sqrt{0.97 \times 0.03/1000} = \sqrt{0.0000291} \approx 0.0054$; report about $0.97 \pm 0.011$. For MCSE $= 0.001$: $S = 0.97 \times 0.03/0.001^2 = 29\,100$ draws. If the draws come from MCMC, these are effective draws (ESS), so you may need more actual iterations.

Chapter 4.14 · Syllabus Module 4

Descriptive and robust statistics

Before any model, you summarize data with a few numbers: where it sits, how spread out it is, and what shape it has. Some of these numbers are thrown around by a single bad day; others barely notice it. Knowing which is which is the whole story of robust statistics, and it is exactly why your forecasting model offers a Student-t likelihood next to the Normal one.

  • Compute and interpret the centre of data: mean, median, mode, trimmed mean and winsorized mean
  • Compute and interpret spread: variance, standard deviation, range, IQR and MAD (raw and ×1.4826)
  • Describe shape: skewness, kurtosis, heavy tails, and outliers (the 1.5×IQR rule as a convention)
  • Explain robustness with two ideas: the breakdown point and the influence of one point
  • Connect it all to your projects: why a Normal likelihood is fragile under spikes and how a Student-t likelihood limits their influence

The centre of data: mean, median and mode core

Your shop had six ordinary days with about 11 to 15 orders each, and one promotion day with 90 orders. "How many orders on a typical day?" has more than one honest answer:

  • The mean (average) shares the total equally among the days: every order counts, including all 90 of the promotion day. It comes out about 24, more than any ordinary day.
  • The median is the middle day once you line the days up from smallest to largest: 14. The promotion day is just "the biggest one"; how big it is does not matter.
  • The mode is the most common value: 15 (it happened twice).

Three ways to say it:

  • Picture: the mean is the balance point of a seesaw loaded with the data; the median is the point that splits the data into two equal halves; the mode is the highest peak.
  • Numbers: one promotion day moves the mean from 13.3 to 24.3 but the median only from 13.5 to 14.
  • Slogan: the mean listens to every value; the median listens only to the order.

Seven days of orders: 12, 15, 11, 14, 13, 15, 90.

  1. Mean: add them up, $12 + 15 + 11 + 14 + 13 + 15 + 90 = 170$, and divide by the number of days: $170/7 \approx 24.3$.
  2. Median: sort them, $11, 12, 13, 14, 15, 15, 90$. With 7 values the middle one is the 4th: 14.
  3. Mode: 15 appears twice, every other value once, so the mode is 15.
  4. Without the promotion day (6 values: 11, 12, 13, 14, 15, 15): mean $80/6 \approx 13.3$; with an even count the median is the average of the two middle values, $(13 + 14)/2 = 13.5$.
  5. So the single extreme day raised the mean by $24.3 - 13.3 = 11$ orders and the median by only $0.5$.

For data $x_1, \dots, x_n$, write the sorted values as $x_{(1)} \le x_{(2)} \le \dots \le x_{(n)}$ (the order statistics).

  • Mean: $\bar x = \frac{1}{n}\sum_{i=1}^{n} x_i$. It is the value $c$ that makes the total squared distance $\sum (x_i - c)^2$ smallest.
  • Median: the middle order statistic, $x_{((n+1)/2)}$ for odd $n$, and the average of the two middle ones for even $n$. It makes the total absolute distance $\sum |x_i - c|$ smallest. For a distribution, the median is the 0.5 quantile (Chapter 4.4).
  • Mode: the most frequent value. For continuous data (no exact repeats), the mode means the highest peak of a histogram or density estimate (Chapter 4.16). Data with two clear peaks are called bimodal.

For a right-skewed distribution (a long tail to the right) the usual order is mode $\lt$ median $\lt$ mean, because the tail pulls the balance point to the right. This is a common pattern, not a law: it can fail, especially for discrete data.

Why do we need it?

One number for "where the data is" is the first summary anyone asks for. Choosing the wrong one tells the wrong story: an average salary or an average order value can describe nobody, because a few huge values pull it up.

Where is it used?

Mean: totals and revenue (total = $n \times$ mean), A/B metrics, MSE-trained predictions. Median: typical latency, house prices, robust baselines, MAE-trained predictions, the point forecast of a back-transformed log model (Chapter 4.18). Mode: the MAP estimate in Bayesian models (the posterior's peak).

How is it used?

Compute both mean and median. If they are close, the data are roughly symmetric and either works. If they differ a lot, look for skew or extreme values and decide which question you are answering: "total per unit" (mean) or "typical unit" (median). In code: np.mean, np.median, scipy.stats.mode.

mode = peak median: half the area on each side mean = balance point (the seesaw pivot) 50% long right tail pulls the mean right → 1 1.68 2
A right-skewed distribution (a Gamma with shape 2): the mode is the peak (1), the median splits the area into two halves (1.68), and the mean is where the shape would balance on a pivot (2). The long right tail drags the mean furthest.

Each dot is one day's orders (one day per row). Drag any day left or right. Watch the mean (purple) follow every move, the median (green) move only when the order of the days changes, and the mode (teal) jump between repeated values. Press Promotion day: one day jumps to 55 orders. The mean leaps; the median barely moves. Press Two peaks to see data with two modes, where a single "centre" describes neither group.

"The median is a more accurate mean."

The median answers a different question ("the typical day") than the mean ("the total shared equally"). For a symmetric distribution they estimate the same number; for a skewed one they do not, and neither is "wrong".

"In a right-skewed distribution the mean is always bigger than the median."

Usually, but not always. It is a rule of thumb that can fail, especially for discrete data and multimodal shapes. Compute both instead of guessing.

"Every dataset has one mode."

Data can have no repeated value (no mode in the counting sense), several tied modes, or two separate peaks (bimodal), which often means two different groups are mixed together.

In an A/B framework like yours, a revenue-per-user metric is usually mostly zeros with a few large values. Its median is often 0 for both variants, which says nothing. The business cares about total revenue $= n \times$ mean, so the mean is the right target, and you need tools that handle its heavy tail (a Student-t or a transformed scale) rather than switching to the median. In your forecasting model, the posterior predictive distribution of tomorrow's demand has a mean and a median; for a skewed predictive (counts, or a log-scale model) they differ, so say which one your point forecast reports.

Mean $\bar x = \frac1n\sum x_i$ (balance point, minimizes $\sum(x_i - c)^2$). Median = middle sorted value (minimizes $\sum|x_i - c|$). Mode = most frequent value / highest peak.

Right skew: usually mode $\lt$ median $\lt$ mean.

Trap: the median is not a "better mean", it is a different question. One extreme value can move the mean a lot and the median hardly at all.

Quick check: the data 3, 8, 8, 9, 100. Give the mean, the median and the mode.

Mean $= 128/5 = 25.6$. Median $=$ the 3rd sorted value $= 8$. Mode $= 8$ (twice). The mean is far from every "typical" value because of the 100.

In between: the trimmed mean and the winsorized mean

The mean uses every value fully; the median throws away almost all the information except the order. There is a middle road. In judged sports such as diving, the highest and the lowest scores are dropped before averaging, so one very generous or very harsh judge cannot decide the result. That is a trimmed mean.

A winsorized mean is the gentler version: instead of throwing the extreme values away, it pulls them in to the nearest value that is kept, then averages everything. The extreme day still counts as "a high day", just not as an enormous one.

Three ways to say it:

  • Picture: trimming cuts off both ends of the sorted row; winsorizing folds the ends back onto the last kept values.
  • Numbers: for 2, 5, 6, 7, 7, 8, 8, 9, 11, 37 the mean is 10.0, the 10% trimmed mean 7.63, the 10% winsorized mean 7.7 and the median 7.5.
  • Slogan: trim a little from each end and the extremes lose their power.

Ten values, sorted: 2, 5, 6, 7, 7, 8, 8, 9, 11, 37. Use 10% at each end, so cut $k = 0.1 \times 10 = 1$ value at each end.

  1. Ordinary mean: $(2 + 5 + 6 + 7 + 7 + 8 + 8 + 9 + 11 + 37)/10 = 100/10 = 10$.
  2. Trimmed mean: drop the 2 and the 37, average the other 8: $(5 + 6 + 7 + 7 + 8 + 8 + 9 + 11)/8 = 61/8 \approx 7.63$.
  3. Winsorized mean: replace the 2 by the next value, 5, and the 37 by the previous value, 11: $5, 5, 6, 7, 7, 8, 8, 9, 11, 11$. Average: $77/10 = 7.7$.
  4. Median (even count): $(7 + 8)/2 = 7.5$.
  5. The single 37 pulled the ordinary mean up by about 2.4 compared with the robust versions.

Let $k = \lfloor \alpha n \rfloor$ (round $\alpha n$ down) for a trimming proportion $\alpha$ between 0 and 0.5.

  • $\alpha$-trimmed mean: drop the $k$ smallest and the $k$ largest values and average the remaining $n - 2k$: $\;\bar x_{\alpha} = \frac{1}{n - 2k}\sum_{i=k+1}^{n-k} x_{(i)}$.
  • $\alpha$-winsorized mean: replace the $k$ smallest values by $x_{(k+1)}$ and the $k$ largest by $x_{(n-k)}$, then take the ordinary mean of all $n$ values.
  • $\alpha = 0$ gives the ordinary mean. As $\alpha$ approaches 0.5, the trimmed mean approaches the median.

Libraries can count the cut slightly differently when $\alpha n$ is not a whole number; scipy.stats.trim_mean(x, 0.1) cuts $\lfloor 0.1n \rfloor$ from each end. Check the documentation of the function you use.

Why do we need it?

The median wastes information when the data are mostly well behaved, and the mean is wrecked by a few wild values. Trimmed and winsorized means keep most of the information while capping the damage that extremes can do.

Where is it used?

Judged sports scores, inflation measures (trimmed-mean CPI), capping heavy revenue metrics before A/B analysis, robust baselines in monitoring, gradient clipping and loss clipping in ML (the same "cap the extremes" idea), and Yuen's trimmed-mean t-test.

How is it used?

Pick a small proportion in advance (5–20% is common) and apply it the same way to every group you compare: scipy.stats.trim_mean(x, 0.1), or scipy.stats.mstats.winsorize(x, limits=(0.01, 0.01)) followed by .mean(). Report that you did it, because it changes the quantity being estimated.

Twelve values, one per row. Choose how much to cut at each end. Trimmed values become hollow rings (they are ignored by the trimmed mean); winsorized values show a pink arrow to where they are pulled in. Drag the top value (45) further right: the mean keeps moving, while the trimmed and winsorized means stop caring once that value is cut. Set the cut to 0 and both become the ordinary mean; set it to 25% and the trimmed mean gets close to the median.

"Winsorizing just removes noise, so the winsorized mean estimates the same thing as the mean."

For skewed data it estimates a different, smaller number. Capping big orders lowers the average on purpose. That is fine for comparing variants, but it is no longer the true average revenue.

"I can choose the trimming proportion after looking at which result is significant."

Choose it before looking at the outcome, and apply the same rule to every group, or you are tuning the analysis to get the answer you want.

If your A/B framework (or the data pipeline feeding it) ever caps or trims a heavy metric, that changes what $\theta$ means in the model: the posterior is about the capped metric. Your framework's Student-t likelihood is an alternative that keeps the original scale: instead of editing the data, it lets the model expect occasional extreme values (section 10).

"We winsorized revenue at the 99th percentile, so our lift estimate is the lift in average revenue."

It is the lift in capped average revenue. It has less variance (smaller error bars) but is biased relative to the true average whenever the extremes differ between variants.

Model answer: "Winsorizing trades bias for variance. It changes the estimand (the exact quantity you are estimating), from mean revenue to mean capped revenue, so I report it as such, choose the cap before the test, and, if the decision is about total revenue, check that the uncapped result points the same way."

$k = \lfloor\alpha n\rfloor$. Trimmed mean: drop $k$ at each end, average the rest. Winsorized mean: replace the $k$ extremes at each end by the nearest kept value, average all $n$.

$\alpha = 0$: the mean; $\alpha \to 0.5$: the median.

Trap: both change the estimand (the quantity being estimated) for skewed data; fix $\alpha$ in advance.

Quick check: 1, 3, 4, 5, 6, 7, 8, 9, 10, 60. Compute the 10% trimmed mean and the 10% winsorized mean.

$k = 1$. Trimmed: drop 1 and 60, $(3+4+5+6+7+8+9+10)/8 = 52/8 = 6.5$. Winsorized: 1 → 3 and 60 → 10, giving $3,3,4,5,6,7,8,9,10,10$, sum 65, mean 6.5. (Here they happen to agree; the ordinary mean is $113/10 = 11.3$.)

Spread: variance, standard deviation and range

The centre says where the data sit; the spread (also called dispersion) says how far the values wander from it. You already know the two main tools from Chapter 4.5: the variance (the average squared distance from the mean) and the standard deviation (its square root, back in the data's units). The simplest spread of all is the range: largest minus smallest.

The range has a hidden flaw: it depends on how much data you have. Keep collecting days and sooner or later you see a new record high or low, so the range keeps growing even though the process has not changed. The standard deviation settles down instead.

Three ways to say it:

  • Picture: the SD is the typical distance from the centre; the range is the distance between the two most extreme points ever seen.
  • Numbers: for Normal data with $\sigma = 1$, the expected range is about 3.1 for 10 values, 5.0 for 100 and 6.5 for 1 000, while the SD stays near 1.
  • Slogan: the SD describes the process; the range describes your two most extreme days.

Six ordinary days: 12, 15, 11, 14, 13, 15. Then add the promotion day, 90.

  1. Mean of the six: $80/6 \approx 13.33$. Squared distances: $1.78, 2.78, 5.44, 0.44, 0.11, 2.78$ (for example $(12 - 13.33)^2 \approx 1.78$). Sum $\approx 13.33$.
  2. Sample variance: $13.33/(6-1) \approx 2.67$; standard deviation $\sqrt{2.67} \approx 1.63$ orders. Range: $15 - 11 = 4$.
  3. With the promotion day: mean $170/7 \approx 24.29$; the 90 alone contributes $(90 - 24.29)^2 \approx 4318$ to the sum of squares.
  4. The SD jumps to about $29.0$ and the range to $90 - 11 = 79$. One day multiplied the SD by about 18.
  5. Squaring is the reason: a distance 10 times larger counts 100 times more.
  • Sample variance: $s^2 = \frac{1}{n-1}\sum_{i=1}^{n}(x_i - \bar x)^2$; sample standard deviation: $s = \sqrt{s^2}$, in the data's units. (Why $n - 1$: Chapter 4.5.)
  • Range: $x_{(n)} - x_{(1)}$, the largest minus the smallest value.
  • Library defaults differ: np.var and np.std divide by $n$ (ddof=0); pandas .var() and .std() divide by $n - 1$ (ddof=1).
  • $s$ estimates a fixed population $\sigma$ and settles as $n$ grows (if $\sigma$ is finite). The range estimates no fixed quantity for unbounded distributions: its expected value keeps growing with $n$ (for Normal data, roughly like $2\sigma\sqrt{2\ln n}$ for very large $n$).
Why do we need it?

Two processes with the same average can behave completely differently. Stock, staffing, risk and error bars all depend on how much values wander, not only on where they sit.

Where is it used?

The SD: z-scores and standardization, standard errors, the noise scale $\sigma$ of a Normal likelihood, prediction intervals, sample-size planning for A/B tests. The range: quick data-quality checks (min and max reveal impossible values), and the range-based control charts of factory quality control.

How is it used?

Report the SD next to the mean (x.std(ddof=1)), and look at min and max for sanity (negative counts, impossible dates). Do not compare ranges of datasets with different sizes, and remember that both the SD and the range react strongly to a single extreme value.

One growing sample from a population with true $\sigma = 1$. Blue: the sample SD after $n$ values. Orange: the range after $n$ values (x-axis on a log scale). For Normal data the dashed orange curve is the expected range. Notice that the SD settles near the green line at 1, while the range climbs steadily. With Student-t (heavy tails) the range jumps up in big steps. With Uniform the range cannot exceed $2\sqrt 3 \approx 3.46$, so it levels off.

"Dataset A has a bigger range than dataset B, so A is more variable."

If A just has more rows, its range will tend to be bigger even when both come from the same process. Compare standard deviations (or IQRs), not ranges, across datasets of different sizes.

"The SD is robust because it averages over all the data."

Averaging squared distances makes it fragile: one value far away can dominate the sum. One promotion day took the SD from 1.6 to 29.

$s^2 = \frac{1}{n-1}\sum(x_i - \bar x)^2$, $s = \sqrt{s^2}$; range $= \max - \min$.

np.std uses ddof=0, pandas .std() uses ddof=1.

Trap: the range grows with $n$; the SD (and the range) react strongly to one extreme value.

Quick check: two weeks of data have range 30 and two years of data from the same process have range 75. Has the process become more variable?

Not necessarily. Two years contain about 50 times more days, so more chances to see an extreme high and low. Compare the standard deviations (or IQRs) of the two periods instead.

Quartiles, the IQR and the box plot core

Line all the days up from smallest to largest and cut the line into four groups with the same number of days. The three cut points are the quartiles: $Q_1$ (a quarter of the days are below it), the median $Q_2$ (half below) and $Q_3$ (three quarters below). The distance from $Q_1$ to $Q_3$ is the interquartile range (IQR): the width of the middle half of the data.

Because the IQR ignores the top quarter and the bottom quarter, a few extreme days cannot stretch it. A box plot draws exactly these numbers: a box from $Q_1$ to $Q_3$, a line at the median, "whiskers" out to the last ordinary values, and dots for anything unusually far out.

Three ways to say it:

  • Picture: the box holds the middle half of the data; its width is the IQR.
  • Numbers: for 11, 12, 13, 14, 15, 15, 90 the IQR is 2.5 orders, while the SD is 29.
  • Slogan: the IQR measures spread using only the middle half, so the extremes cannot bully it.

Sorted orders: 11, 12, 13, 14, 15, 15, 90 ($n = 7$). We use NumPy's default rule: the $p$ quantile sits at position $1 + (n-1)p$ in the sorted list, interpolating between neighbours when the position is not a whole number.

  1. $Q_1$: position $1 + 6 \times 0.25 = 2.5$, halfway between the 2nd value (12) and the 3rd (13): $Q_1 = 12.5$.
  2. Median: position $1 + 6 \times 0.5 = 4$: the 4th value, 14.
  3. $Q_3$: position $1 + 6 \times 0.75 = 5.5$, halfway between the 5th (15) and the 6th (15): $Q_3 = 15$.
  4. IQR $= Q_3 - Q_1 = 15 - 12.5 = 2.5$.
  5. Fences (section 8): $Q_1 - 1.5 \times 2.5 = 8.75$ and $Q_3 + 1.5 \times 2.5 = 18.75$. The 90 is outside, so the box plot draws it as a separate dot; the whiskers stop at 11 and 15.

Other quantile rules exist (for example "the median of each half", which gives 12 and 15 here). They differ a little for small $n$ and agree for large $n$.

  • Quartiles: $Q_1$, $Q_2$, $Q_3$ are the 0.25, 0.5 and 0.75 quantiles of the data (Chapter 4.4). np.quantile(x, [0.25, 0.5, 0.75]) uses the "linear" method by default (also pandas' and R's default).
  • Interquartile range: $\text{IQR} = Q_3 - Q_1$, in the data's units.
  • Five-number summary: minimum, $Q_1$, median, $Q_3$, maximum.
  • Box plot (Tukey's version): a box from $Q_1$ to $Q_3$ with a line at the median; whiskers from the box to the most extreme data values that lie within the fences $Q_1 - 1.5\,\text{IQR}$ and $Q_3 + 1.5\,\text{IQR}$; every value beyond a fence drawn as its own dot.
  • For Normal data, $\text{IQR} = 2 \times 0.6745\,\sigma \approx 1.349\,\sigma$, so $\text{IQR}/1.349$ is a robust estimate of $\sigma$.
Why do we need it?

We need a spread that a few extreme values cannot inflate, and a picture that shows centre, spread, skew and unusual values at a glance for many groups side by side.

Where is it used?

Box plots comparing segments or variants in an A/B dashboard, residuals by day of week in forecasting, latency percentiles in monitoring, scikit-learn's RobustScaler (subtract the median, divide by the IQR), and Tukey's outlier fences.

How is it used?

q1, q3 = np.quantile(x, [0.25, 0.75]), iqr = q3 - q1 (or scipy.stats.iqr(x)). Draw box plots with plt.boxplot or df.boxplot(by='segment'). A box much closer to one whisker than the other signals skew.

IQR = Q3 − Q1 (middle 50%) Q1Q3 median lower fence upper fence outliers whisker end the 19 data values · fences = Q1 − 1.5·IQR and Q3 + 1.5·IQR
Anatomy of a box plot, drawn above the data it summarizes. The box spans the middle half (Q1 to Q3), the green line is the median, the whiskers reach the most extreme values inside the fences, and values beyond a fence are drawn one by one.

The dots at the bottom are 11 daily values; the box plot above is built from them live. Drag the far-right value (40) toward the others: when it crosses the upper fence (dashed purple) it stops being an outlier and the whisker grows to reach it. Drag two values far to the left: the box hardly changes, because the IQR only looks at the middle half. Watch the readout for the quartile arithmetic.

"The whiskers go from the minimum to the maximum."

In Tukey's box plot (the default in matplotlib, seaborn and pandas) the whiskers stop at the last values inside the fences; anything beyond is drawn as a dot. Some tools offer min-to-max whiskers, so check the setting.

"Two tools gave different quartiles, so one of them has a bug."

There are several accepted quantile rules. For small samples they differ slightly. NumPy, pandas and R's default ("linear") agree with each other.

$Q_1, Q_2, Q_3$ = 0.25, 0.5, 0.75 quantiles; IQR $= Q_3 - Q_1$ (middle half). Normal: IQR $\approx 1.349\sigma$.

Box plot: box $Q_1$–$Q_3$, median line, whiskers to the last values inside $Q_1 - 1.5\,$IQR and $Q_3 + 1.5\,$IQR, dots beyond.

Trap: whiskers are not min/max; quantile rules differ slightly for small $n$.

Quick check: $Q_1 = 20$ and $Q_3 = 30$. Which of the values 3, 10, 44, 46 are drawn as separate dots in a box plot?

IQR $= 10$, fences $20 - 15 = 5$ and $30 + 15 = 45$. Beyond the fences: 3 (below 5) and 46 (above 45). The 10 and the 44 are inside and can be whisker ends.

The MAD: a spread that ignores extremes

The standard deviation measures distances from the mean, squares them and averages them. All three steps are sensitive to extreme values. The median absolute deviation (MAD) replaces each fragile step with a robust one: measure distances from the median, do not square them, and take the median of the distances.

In words: "how far is a typical value from the typical value?" Half of the data lie within one MAD of the median.

Three ways to say it:

  • Picture: draw a band around the median that holds exactly half the points; the MAD is the half-width of that band.
  • Numbers: adding the 90-order promotion day takes the SD from 1.6 to 29, but leaves the MAD at about 1.
  • Slogan: the MAD is the median's version of the standard deviation.

Orders 11, 12, 13, 14, 15, 15, 90.

  1. Median: 14.
  2. Absolute distances from 14: $3, 2, 1, 0, 1, 1, 76$.
  3. Sort them: $0, 1, 1, 1, 2, 3, 76$. Their median (the 4th) is 1. So MAD $= 1$.
  4. Scaled version: $1.4826 \times 1 \approx 1.48$. This is the MAD's estimate of $\sigma$ if the bulk of the data were Normal.
  5. Compare: the SD is about 29.0 and the range 79. Without the promotion day the MAD is 1.5 (scaled 2.22) and the SD 1.63: the MAD stayed small, the SD exploded.

$$\text{MAD} = \text{median}_i\,\big|x_i - \text{median}(x)\big|, \qquad \hat\sigma_{\text{MAD}} = 1.4826 \times \text{MAD}.$$

  • Why 1.4826? For Normal data, $|X - \mu|$ has median $0.6745\,\sigma$, because $P(|Z| \le 0.6745) = 0.5$ for a standard Normal $Z$. So MAD $\approx 0.6745\sigma$, and $1/0.6745 = 1/\Phi^{-1}(0.75) \approx 1.4826$ turns it into an estimate of $\sigma$.
  • The scaling is only "correct" for Normal-shaped data. For other shapes, the scaled MAD is still a sensible robust spread, but not exactly $\sigma$.
  • Library defaults differ: scipy.stats.median_abs_deviation(x) returns the raw MAD; add scale='normal' for the ×1.4826 version. statsmodels.robust.mad(x) returns the scaled version by default. Older pandas had .mad() meaning the mean absolute deviation around the mean, a different (non-robust) quantity; it was removed in pandas 2.0.
Why do we need it?

A spread estimate that the very outliers you are looking for cannot inflate. With the SD, one huge value makes itself look normal by stretching the yardstick it is measured with (this is called masking).

Where is it used?

Robust z-scores for anomaly detection in monitoring and data cleaning, the scale step of robust M-estimators (Huber regression), robust standardization of features, and quick checks of residual noise levels in forecasting when a few days are spikes.

How is it used?

Compute median_abs_deviation(x, scale='normal') and use it where you would use the SD, for example a robust z-score $(x - \text{median})/(1.4826\,\text{MAD})$; values beyond about 3.5 are worth a look. If MAD is 0 (over half the values identical), fall back to the IQR or report it.

Fifteen ordinary values (blue) come from a Normal distribution with $\sigma = 1$, plus one value you can drag (red). The bars show three estimates of $\sigma$: the SD, IQR/1.349 and 1.4826·MAD, plus the range for comparison. The green line marks the true $\sigma = 1$. Drag the red value from the crowd out to 30: the SD bar grows steadily and the range explodes, while IQR/1.349 and 1.4826·MAD barely move.

"The MAD and the SD measure the same thing, so they should be equal."

The raw MAD is about $0.67\sigma$ for Normal data. Only the scaled MAD ($\times 1.4826$) is comparable with the SD, and only for Normal-shaped data.

"MAD in pandas and MAD in SciPy are the same function."

Check what the function computes: raw or scaled, median absolute deviation or mean absolute deviation. These differ by large factors.

Scaling matters in both projects (the global scaler is taught in Chapter 4.18). If a scaler uses the mean and SD, one extreme day both shifts and stretches the whole scaled series. A robust alternative is to centre with the median and scale with the MAD or the IQR (that is what scikit-learn's RobustScaler does with the IQR). Check which statistics your scaler uses, and remember that the prior scales you set (for example on the noise $\sigma$) are expressed in those scaled units.

MAD $= \text{median}|x_i - \text{median}(x)|$; $\hat\sigma = 1.4826\,$MAD (Normal data; $1.4826 = 1/\Phi^{-1}(0.75)$).

Robust z-score: $(x - \text{median})/(1.4826\,\text{MAD})$, flag beyond ~3.5.

Trap: raw vs scaled; SciPy raw by default, statsmodels scaled; old pandas .mad() was the mean absolute deviation.

Quick check: data 2, 4, 4, 5, 7, 9, 40. Compute the MAD and the scaled MAD.

Median 5. Absolute distances: 3, 1, 1, 0, 2, 4, 35; sorted 0, 1, 1, 2, 3, 4, 35; median 2. MAD $= 2$; scaled $\approx 2.97$.

Shape, part 1: skewness

Daily order counts, waiting times, revenue and house prices pile up on the left and trail off in a long tail to the right. Exam scores on an easy test do the opposite: most are high, a few are very low. This lopsidedness is called skewness. A long right tail is positive (right) skew; a long left tail is negative (left) skew; a mirror-symmetric shape has skewness 0.

The number works by cubing the distances from the mean. Cubes keep the sign (left is negative, right is positive) and make far values count enormously, so the side with the longer tail wins.

Three ways to say it:

  • Picture: skewness says which side the long tail is on and how long it is.
  • Numbers: for 1, 2, 3, 4, 10 the one far value on the right gives a skewness of about $+1.14$.
  • Slogan: the tail tells the sign.

Data 1, 2, 3, 4, 10.

  1. Mean: $20/5 = 4$. Distances from the mean: $-3, -2, -1, 0, 6$.
  2. Average squared distance: $(9 + 4 + 1 + 0 + 36)/5 = 50/5 = 10$. Call it $m_2$.
  3. Average cubed distance: $(-27 - 8 - 1 + 0 + 216)/5 = 180/5 = 36$. Call it $m_3$. The single $+6$ contributes $216$, more than all the negatives together.
  4. Skewness: $g_1 = m_3/m_2^{3/2} = 36/10^{1.5} = 36/31.62 \approx 1.14$. Positive: the long tail is on the right.
  5. pandas reports $1.70$ for the same data, because it applies a small-sample correction (next box). With only 5 values, both numbers are very uncertain.

With $m_k = \frac{1}{n}\sum_{i}(x_i - \bar x)^k$ (the $k$-th central moment of the data):

$$g_1 = \frac{m_3}{m_2^{3/2}} \quad\text{(SciPy's default)}, \qquad G_1 = g_1\,\frac{\sqrt{n(n-1)}}{n-2} \quad\text{(pandas, Excel; SciPy with } \texttt{bias=False}\text{)}.$$

For a random variable: $\gamma = E[(X-\mu)^3]/\sigma^3$. Known values: Normal 0, Exponential 2, Gamma(shape $k$) $2/\sqrt k$, Poisson($\lambda$) $1/\sqrt\lambda$, Bernoulli($p$) $(1-2p)/\sqrt{p(1-p)}$.

  • Skewness has no units: the cube of distances is divided by the cube of the SD.
  • It is noisy: for Normal data its standard error is about $\sqrt{6/n}$ (about 0.24 at $n = 100$).
  • It is very sensitive to outliers (cubes!). One extreme value can create a large skewness by itself.
Why do we need it?

Skew changes which summaries make sense (mean vs median), how fast the CLT kicks in (Chapter 4.13), and whether a symmetric likelihood such as the Normal fits the data at all.

Where is it used?

Deciding to log-transform revenue or demand (Chapter 4.18), choosing a likelihood (Gamma, Log-Normal or Negative Binomial for right-skewed data), checking residuals of a forecast, and judging whether a Normal approximation for an A/B metric is safe.

How is it used?

Compute scipy.stats.skew(x) and compare the mean with the median, but always look at a histogram or Q-Q plot too (Chapter 4.17). As a rough guide (a rule of thumb, not a test): $|g_1| \lt 0.5$ fairly symmetric, $0.5$–$1$ moderate, above $1$ strong skew.

Choose a shape and draw a sample. The histogram shows the data, with the mean (purple) and median (green). Go from Normal to Gamma to Exponential to Log-Normal: the right tail grows, the skewness grows, and the mean moves further right of the median. Exam scores is skewed the other way. Use a small $n$ and press New sample repeatedly: the sample skewness jumps around a lot.

"Skewness 0 means the distribution is symmetric."

Symmetric implies skewness 0, but not the other way round: some lopsided shapes have positive and negative cubed parts that cancel. Look at the plot.

"My sample of 30 has skewness 0.6, so the population is skewed."

With $n = 30$ the noise in the sample skewness is about $\sqrt{6/30} \approx 0.45$ for Normal data. A value of 0.6 is weak evidence. And a single outlier can produce it.

Your forecasting target is often right-skewed (counts of orders, demand). That shows up three ways: the residuals of a Normal-likelihood fit are skewed, the predictive mean sits above the predictive median, and a Negative Binomial likelihood (or a log transform) usually fits better than a Normal (Chapter 7.13). In an A/B framework, strongly skewed metrics are the ones where Normal approximations need large samples (Chapter 4.13).

$g_1 = m_3/m_2^{3/2}$ (SciPy), adjusted $G_1 = g_1\sqrt{n(n-1)}/(n-2)$ (pandas). Positive = long right tail.

Exponential 2, Gamma $2/\sqrt k$, Poisson $1/\sqrt\lambda$. Noise ≈ $\sqrt{6/n}$.

Trap: very outlier-sensitive and noisy; skewness 0 does not prove symmetry; always plot.

Quick check: data 0, 0, 0, 4. Compute $g_1$.

Mean 1; distances $-1, -1, -1, 3$. $m_2 = (1+1+1+9)/4 = 3$; $m_3 = (-1-1-1+27)/4 = 6$. $g_1 = 6/3^{1.5} = 6/5.196 \approx 1.15$: right-skewed.

Shape, part 2: kurtosis and heavy tails core

Two kinds of daily noise can have the same mean and the same standard deviation and still behave very differently. One gives steady, moderate surprises. The other is quiet most days but now and then produces a huge surprise. The second has heavy tails: extreme values happen far more often than a Normal distribution with the same SD would allow.

Kurtosis puts a number on this. It averages the fourth power of the distances from the mean (in SD units). Fourth powers make far values dominate completely, so kurtosis is mostly a measure of how much the tails weigh. It is scaled so that a Normal gives 3; the excess kurtosis subtracts 3 so that the Normal gives 0.

Three ways to say it:

  • Picture: the same width in the middle, but fatter tails reaching further out.
  • Numbers: with the same SD, a day beyond 4 SDs happens about once in 16 000 days for Normal noise, once in 290 for Laplace noise and once in 160 for Student-t noise with $\nu = 3$.
  • Slogan: kurtosis is about the tails, not the peak.

Two tiny datasets, both with mean 0.

  1. Data $-1, -1, 1, 1$: $m_2 = (1+1+1+1)/4 = 1$ and $m_4 = (1+1+1+1)/4 = 1$. Kurtosis $m_4/m_2^2 = 1$; excess $1 - 3 = -2$. No tails at all: every value sits exactly one SD away. ($-2$ is the smallest excess kurtosis possible.)
  2. Data $0, 0, 0, 0, 0, 0, 0, 0, 1, -1$: $m_2 = 2/10 = 0.2$ and $m_4 = 2/10 = 0.2$.
  3. Kurtosis $m_4/m_2^2 = 0.2/0.04 = 5$; excess $5 - 3 = +2$. Mostly nothing, with two values far out compared with the SD ($\sqrt{0.2} \approx 0.45$): heavy tails.
  4. So "positive excess kurtosis" means: compared with a Normal of the same SD, more of the variance comes from rare, far-out values.

With central moments $m_k = \frac1n\sum(x_i - \bar x)^k$:

$$g_2 = \frac{m_4}{m_2^2} - 3 \quad\text{(excess kurtosis; SciPy's default, } \texttt{fisher=True}\text{)}.$$

For a random variable: $\kappa = E[(X-\mu)^4]/\sigma^4 - 3$. scipy.stats.kurtosis(x, fisher=False) returns $m_4/m_2^2$ (Normal = 3); pandas .kurt() returns a small-sample-adjusted excess kurtosis.

Distribution (SD 1)excess kurtosis$P(|X| \gt 3)$$P(|X| \gt 4)$
Uniform$-1.2$00
Normal00.00270.000063
Laplace30.01440.0035
Student-t, $\nu = 5$$6/(\nu - 4) = 6$0.01170.0036
Student-t, $\nu = 3$infinite ($\nu \le 4$)0.01380.0062
  • Heavy-tailed (informally): tails that shrink more slowly than the Normal's $e^{-x^2/2}$, for example like $e^{-|x|}$ (Laplace) or like a power $|x|^{-(\nu+1)}$ (Student-t). Light-tailed: shrinking faster, or bounded (Uniform).
  • The sample kurtosis is very noisy for heavy-tailed data (for $\nu \le 4$ it does not settle at all), and one outlier can dominate it.
Why do we need it?

Risk lives in the tails. Two noise models with the same SD can give the same 50% intervals but wildly different chances of a 4-SD day. A model with tails that are too light is overconfident exactly when it matters.

Where is it used?

Checking residuals of regression and forecasting models before trusting Normal intervals, choosing Student-t over Normal likelihoods (and the size of $\nu$), finance (fat-tailed returns), anomaly detection thresholds, and understanding why Laplace priors leave most changepoint slopes near zero but allow a few big ones.

How is it used?

Compute scipy.stats.kurtosis(resid) as a first hint, then confirm with a Q-Q plot (Chapter 4.17), whose ends bend away in an S-shape for heavy tails. Count how many residuals fall beyond 3 and 4 SDs and compare with the Normal's 0.27% and 0.006%.

Every distribution here has mean 0 and standard deviation 1. The dashed grey curve is the Normal for reference. Left: the ordinary density, with the region beyond ±3 SDs shaded red. Right: the same curves on a log scale, where tails become visible: the Normal falls like a parabola, the Laplace in straight lines, the Student-t much more slowly. Switch between shapes, then press New sample several times with Student-t ν = 3: the sample kurtosis of 1 000 values jumps all over the place.

"High kurtosis means a tall, sharp peak."

Kurtosis is dominated by the tails. A distribution can have a flat top and huge kurtosis, or a sharp peak and modest kurtosis. Read it as "how heavy are the tails compared with a Normal of the same SD".

"SciPy says my residuals have kurtosis 3, so they look Normal."

SciPy reports excess kurtosis by default, so 3 means heavy tails (Laplace-like). The Normal gives 0 by default, or 3 only with fisher=False.

"The sample kurtosis is small, so the tails are light."

For heavy-tailed data the sample kurtosis is unreliable: it can be modest in one sample and huge in the next. Use a Q-Q plot and tail counts.

In both projects the Student-t likelihood is the "heavy tails" choice, and its $\nu$ is a tail dial: $\nu \le 4$ means infinite kurtosis, $\nu = 5$ gives 6, and large $\nu$ approaches the Normal's 0. In your forecasting model the Laplace prior on the slope changes $\delta_j$ has excess kurtosis 3 compared with a Normal prior of the same SD: more mass very close to 0 and more mass far out. That shape is what allows most changepoints to be nearly switched off while a few take big values (Chapter 7.10).

Excess kurtosis $g_2 = m_4/m_2^2 - 3$: Normal 0, Uniform −1.2, Laplace 3, t$_\nu$ $6/(\nu-4)$ ($\nu \gt 4$), infinite for $\nu \le 4$.

Heavy tails = extreme values far more often than a Normal with the same SD.

Trap: not "peakedness"; SciPy reports excess by default; sample kurtosis is unreliable under heavy tails.

Quick check: residuals with SD 10 contain 12 values beyond 40 (4 SDs) out of 2 000. Is that plausible for Normal noise?

A Normal gives $P(|X| \gt 4\sigma) \approx 0.000063$, so about $2000 \times 0.000063 \approx 0.13$ such values expected. Twelve is far more: the tails are heavy (12/2000 = 0.006 is about what a Student-t with $\nu = 3$ would give).

Outliers and the 1.5 × IQR rule

An outlier is a value that sits far away from the bulk of the data. It can be three very different things: a mistake (a logging bug, a test account, a typo), a real but special event (a promotion, a holiday, an outage), or an ordinary draw from a heavy-tailed process. A rule can only flag values as "far"; deciding which of the three it is takes thinking.

The most common rule is Tukey's: flag anything more than 1.5 IQRs below $Q_1$ or above $Q_3$. It is a convention proposed by John Tukey, not a law of nature. On clean Normal data it flags only a small fraction of values, which makes it a handy first screen.

Three ways to say it:

  • Picture: put a fence 1.5 box-widths outside each end of the box; whatever lands beyond a fence gets a closer look.
  • Numbers: even perfectly clean Normal data has about 0.7% of values beyond the fences: 7 per 1 000.
  • Slogan: flagged means "look at it", not "delete it".

How often does clean Normal data cross Tukey's fences?

  1. For a standard Normal, $Q_1 = -0.6745$ and $Q_3 = +0.6745$, so IQR $= 1.349$.
  2. Upper fence: $Q_3 + 1.5 \times 1.349 = 0.6745 + 2.0235 = 2.698$. The lower fence is $-2.698$.
  3. $P(Z \gt 2.698) \approx 0.0035$, and the same below: in total about $0.0070$, i.e. 0.7%.
  4. With 1 000 clean values you should expect about 7 flagged points. With 100 000, about 700.
  5. For heavier tails the rate is higher: about 5.5% for a Student-t with $\nu = 3$ and about 7.8% (all on the high side) for a Log-Normal with $\sigma = 1$. These are not errors; they are what those distributions do.
  • Tukey's fences: $L = Q_1 - 1.5\,\text{IQR}$, $U = Q_3 + 1.5\,\text{IQR}$; values outside $[L, U]$ are "outliers" in the box-plot sense. Values beyond $Q_1 - 3\,\text{IQR}$ or $Q_3 + 3\,\text{IQR}$ are sometimes called "far out".
  • z-score rule: flag $|x - \bar x|/s \gt 3$. Weakness (masking): big outliers inflate $\bar x$ and $s$ and so hide themselves.
  • Robust z-score: flag $|x - \text{median}|/(1.4826\,\text{MAD}) \gt 3.5$ (a common rule of thumb). It does not suffer from masking.
  • All of these thresholds are conventions. For skewed data, symmetric rules flag many legitimate values on the long side; transform first (Chapter 4.18) or use a model-based check.
Why do we need it?

Extreme values can dominate means, SDs, regression fits and the fit of a likelihood. You need a quick, consistent way to find them so you can check whether they are errors, special events, or just the tails of the process.

Where is it used?

Data-quality checks before an A/B analysis (bots, test accounts, duplicated events), cleaning training data for forecasting, monitoring dashboards (alerts on robust z-scores), and box plots in every exploratory analysis.

How is it used?

Flag with fences or robust z-scores, then investigate each flag: fix or drop genuine errors (and document it), model real events (holiday and promotion regressors), and keep the heavy-tailed ones with a likelihood that expects them (Student-t, Negative Binomial). Never delete points just to make a result look cleaner.

box: Q1 to Q3 (IQR = 1.349σ) Q1 − 1.5·IQR Q3 + 1.5·IQR −0.67σ+0.67σ −2.70σ+2.70σ 0.35%0.35% clean Normal data: about 0.7% of values land beyond the fences anyway
Tukey's fences on a Normal distribution. The fences sit at about ±2.7 standard deviations, so roughly 0.35% of perfectly ordinary values fall beyond each one.

Draw $n$ values from a distribution and apply Tukey's fences (dashed purple). Red dots are flagged. For Normal data, move $n$ up to 5 000 and press New sample a few times: about 0.7% get flagged every time, although nothing is wrong with them. Then try Student-t ν = 3 (heavy tails) and Log-Normal (skewed): many more flags, all of them legitimate draws.

"The 1.5×IQR rule found 70 outliers in my 10 000 rows, so I'll delete them."

About 70 flags in 10 000 is exactly what clean Normal data produces. Flags are prompts to investigate. Delete only values you can show are errors, and say so in the report.

"An outlier is an error."

Often it is the most important data point: a holiday, a promotion, an outage. In forecasting such days should be explained by the model (holiday and event regressors), not removed.

"|z| > 3 is a safe outlier test."

Large outliers inflate the mean and SD and can hide themselves (masking). Use the median and MAD (robust z-score) instead.

In your forecasting model, many "outliers" are exactly what the holiday term $h(t)$ and the regressors $X_t\beta$ are for: a promotion day is not noise if you know it was a promotion. What remains unexplained is handled by the likelihood: a Student-t lets rare big residuals happen without dragging the trend. In an A/B framework like yours, clear data errors (bot traffic, duplicated events, test users) should be removed by rules fixed before the experiment, applied equally to both variants.

Fences: $Q_1 - 1.5\,$IQR and $Q_3 + 1.5\,$IQR (≈ ±2.7σ for Normal data; 0.7% of clean data flagged).

Outlier = error, special event, or a heavy-tail draw. Robust z: $|x - \text{med}|/(1.4826\,\text{MAD}) \gt 3.5$.

Trap: flagged ≠ wrong; never delete to make results look nicer; |z| > 3 suffers from masking.

Quick check: for 50 000 clean Normal values, about how many does the 1.5×IQR rule flag?

About $0.007 \times 50\,000 \approx 350$ values, give or take random variation. None of them is an error.

Robust statistics: breakdown point and influence core

A statistic is robust if a few bad values cannot push it around much. Two simple questions make this precise:

  • How many values can I corrupt before it breaks? Replace some data points by absurd numbers. The mean breaks with just one. The median survives until half the data are corrupted. This fraction is the breakdown point.
  • How hard does one value pull? Move a single point further and further out. Its pull on the mean keeps growing without limit; its pull on the median stays fixed. This is the influence of one point.

Three ways to say it:

  • Picture: the mean is a seesaw that one heavy child can tip over; the median is a vote, where one extreme voter is still only one vote.
  • Numbers: with 20 values, one corrupted value can send the mean anywhere; the median holds until 10 are corrupted.
  • Slogan: robust estimators cap the say of any single point.

Twenty clean values with mean 10. Replace some of them by 10 000 (a logging bug).

  1. One corrupted value: the mean becomes about $(19 \times 10 + 10\,000)/20 \approx 510$. Broken, after 1 bad value out of 20 (5%).
  2. 10% trimmed mean (drops 2 values at each end): survives 1 and 2 corrupted values (they are trimmed away); breaks at 3.
  3. Median of 20 values: the average of the 10th and 11th sorted values. With 9 corrupted values (all huge), the 10th and 11th are still clean; with 10 corrupted, the 11th is 10 000 and the median jumps to about 5 000.
  4. So the breakdown points are roughly: mean $1/n$ (essentially 0%), $\alpha$-trimmed mean $\alpha$, median 50%.
  5. For spread: SD and range break with one value (0%), IQR at 25%, MAD at 50%.
  • Breakdown point: the smallest fraction of the data that, when replaced by arbitrary values, can carry the estimate arbitrarily far away (to infinity, or to 0 for a spread). Large samples: mean 0, SD 0, range 0; $\alpha$-trimmed mean $\alpha$; IQR 25%; median 50%; MAD 50%. 50% is the maximum possible.
  • Influence function (idea): how much a single extra point at position $x$ changes the estimate, as a function of $x$. Mean: proportional to $x - \mu$, so unbounded. Median: a fixed step, bounded. A robust estimator has a bounded influence function.
  • M-estimators: estimates defined by minimizing $\sum \rho\big((x_i - \mu)/\sigma\big)$ for some loss $\rho$. $\rho(r) = r^2$ gives the mean; $\rho(r) = |r|$ the median; Huber's loss (quadratic for small $r$, linear for large) is a compromise. The Student-t maximum-likelihood estimate is also an M-estimator, with an influence that even falls back towards zero for very far points (it "redescends").
  • The price of robustness: if the data really are Normal and clean, robust estimators are a bit noisier. The median's variance is about $\pi/2 \approx 1.57$ times the mean's (its efficiency is $2/\pi \approx 64\%$; Chapter 5.1). Trimmed means and Huber estimators lose much less.
Why do we need it?

Real data contain errors and heavy tails. An estimate that can be ruined by one row is dangerous in automated pipelines, where nobody looks at each number before it feeds a decision or a model.

Where is it used?

Huber loss in regression and gradient boosting (HuberRegressor, loss='huber'), Student-t likelihoods in Bayesian models, median-based monitoring baselines, robust scalers, RANSAC in computer vision, and trimmed or winsorized A/B metrics.

How is it used?

Decide what the estimate is for. If you need a total (revenue), keep the mean and model the tails. If you need a typical value or a stable baseline, use a robust estimator. In models, swap a squared loss / Normal likelihood for Huber or Student-t to bound each point's influence.

mean: grows forever median: a fixed step Huber: grows, then capped Student-t (ν = 3): falls back r (distance from the centre) → pull on the estimate
The "influence" of one data point at distance $r$ from the centre. The mean's pull grows without limit (one point can drag it anywhere). The median's pull is a fixed step, Huber's grows and then stops, and the Student-t's ($\nu = 3$) even shrinks again for very far points: far-away values are treated as surprises rather than evidence about the centre.

Twenty clean values have centre 10 and spread 2 (green lines). Use the first slider to replace $k$ of them by a bad value, and the second to choose how bad. Each dot is one estimator, on a log scale; it turns red when it has moved more than 3 times away from the truth. Raise $k$ one step at a time: the mean and SD break at $k = 1$, the 10% trimmed mean at 3, the IQR and the 20% trimmed mean at 5, and the median holds until $k = 10$ (half the data). The MAD also cannot be carried away before $k = 10$, but it is pushed upward first (around $k = 9$ it crosses the 3× line, at about 9 instead of 2; at $k = 10$ it jumps to thousands).

"Robust estimators are always better, so always use the median."

On clean Normal data the median is noisier than the mean (about 64% efficiency), and it answers a different question for skewed data. Robustness is insurance: worth it when bad values or heavy tails are likely.

"The 20% trimmed mean has breakdown 20%, so it is safe if fewer than 20% of the values are bad."

It cannot be carried arbitrarily far, but moderately bad values below the breakdown fraction can still bias it. Breakdown is about the worst case, not about being unaffected.

"We use the median because it is more accurate than the mean."

The median is more robust: it has a 50% breakdown point and bounded influence. On clean Normal data it is actually less precise. And for skewed data it estimates a different quantity.

Model answer: "Robustness is about how much bad or extreme data can move an estimate. The mean has breakdown point zero and unbounded influence; the median and MAD have 50% breakdown; trimmed means sit in between. I pick based on the question (total vs typical) and the data: if extremes are errors or nuisance, a robust estimator or a heavy-tailed likelihood; if they are the signal, I model them."

Breakdown: mean, SD, range 0 · $\alpha$-trimmed mean $\alpha$ · IQR 25% · median, MAD 50% (the maximum).

Influence: mean unbounded; median bounded; Huber capped; Student-t redescending.

Trap: robustness costs efficiency on clean Normal data (median ≈ 64%); it is insurance, not free accuracy.

Quick check: 100 values. What is the smallest number of corrupted values that can ruin the 5% trimmed mean? And the IQR?

The 5% trimmed mean drops 5 values at each end, so 6 corrupted values (all at one end) get through: about 6%. The IQR uses the 25th and 75th percentiles, so about 25 or more extreme values on one side are needed to move them arbitrarily: roughly 25%.

Why your forecasting model offers both a Normal and a Student-t likelihood core

Fitting a model with a Normal likelihood means, in effect, estimating the centre with a mean-like rule and the noise level with an SD-like rule. You have just seen that both have breakdown point zero: a handful of spike days (an unmodelled promotion, an outage, a data glitch) drag the fitted level towards them and inflate the noise $\sigma$. A bigger $\sigma$ then widens every day's interval, even the calm days.

A Student-t likelihood (Chapter 4.9) has heavy tails, so it already "expects" an occasional big surprise. Fitting it automatically gives far-away points a small weight. The spikes stay in the data and are still explained (as rare tail events), but they no longer steer the centre or blow up the scale.

Three ways to say it:

  • Picture: under a Normal, a spike is "impossible", so the model bends everything to accommodate it; under a Student-t, a spike is "rare but allowed", so the model shrugs.
  • Numbers: three spikes among 63 days push the Normal's $\hat\sigma$ from 1 to about 1.9, while the Student-t ($\nu = 3$) scale only moves from about 0.8 to 0.9 and the spikes get weights between 0.04 and 0.09 instead of 1.
  • Slogan: Normal for well-behaved noise, Student-t when surprises happen.

How the Student-t down-weights a point. With $\nu = 3$, the fit gives each point the weight $w = (\nu + 1)/(\nu + r^2)$, where $r$ is its residual divided by the scale (this comes from the t's log-density; the Normal gives every point weight 1).

  1. An ordinary day with $r = 0.5$: $w = 4/(3 + 0.25) \approx 1.23$.
  2. A day at $r = 2$: $w = 4/(3 + 4) \approx 0.57$.
  3. A spike at $r = 6$: $w = 4/(3 + 36) \approx 0.10$.
  4. A spike at $r = 10$: $w = 4/(3 + 100) \approx 0.04$. The further out, the less it counts.
  5. The fitted centre is the weighted mean $\sum w_i x_i / \sum w_i$ and the scale uses the same weights, so spikes barely move either (the weights and the fit are updated together until they agree).

For noise $\epsilon_t = y_t - \mu_t$:

  • Normal likelihood $y_t \sim N(\mu_t, \sigma^2)$: $-\log p = \frac{(y_t - \mu_t)^2}{2\sigma^2} + \text{const}$. Squared loss: a residual's pull grows linearly with its size, without limit.
  • Student-t likelihood $y_t \sim \text{StudentT}(\nu, \mu_t, \sigma)$: $-\log p = \frac{\nu+1}{2}\log\!\Big(1 + \frac{(y_t - \mu_t)^2}{\nu\sigma^2}\Big) + \text{const}$. Log-shaped loss: a far residual's pull shrinks (redescending influence). Small $\nu$ = heavy tails; $\nu \to \infty$ gives the Normal back.
  • $\sigma$ in the Student-t is a scale, not the standard deviation: SD $= \sigma\sqrt{\nu/(\nu-2)}$ for $\nu \gt 2$. NumPyro: dist.StudentT(df, loc, scale); dist.Normal(loc, scale) takes the SD.
Why do we need it?

Real demand and metric data have occasional days the model cannot explain. With a Normal likelihood those days distort the trend, seasonality and noise level; with a Student-t they are absorbed as tail events and the rest of the fit stays honest.

Where is it used?

Your Prophet-style forecasting model (Normal or Student-t likelihood), the continuous metrics of your A/B framework (Normal or Student-t), robust Bayesian regression in general, Kalman filters with outliers, and any model whose residual Q-Q plot shows heavy tails.

How is it used?

Fit with a Normal first; check residual kurtosis, tail counts and the Q-Q plot. If the tails are heavy, switch to dist.StudentT(nu, mu, sigma) with $\nu$ fixed (3–5 is common) or learned with a prior that keeps it above 1 or 2. Compare forecasts and intervals on held-out data (Chapter 7.13).

Sixty ordinary residuals (noise with SD 1, rug at the bottom) plus three spike days you can drag (red handles; they are part of the data). Orange: the fitted Normal (centre = mean, $\sigma$ = SD). Teal: the fitted Student-t with $\nu$ from the slider. Drag the spikes far right: the orange curve shifts and flattens, so it no longer fits the 60 ordinary days (its "middle half" band now holds far more than half of them), while the teal curve stays put. Drag the spikes into the crowd: both fits agree. Raise $\nu$ to 30: the Student-t behaves like the Normal again.

"With a Student-t likelihood I don't need to look at outliers any more."

It limits their influence on the fit; it does not tell you why they happened. Recurring spikes (promotions, holidays) should become regressors or holiday terms, so the model can predict them.

"The Student-t's σ is the noise standard deviation."

It is a scale. The SD is $\sigma\sqrt{\nu/(\nu-2)}$ for $\nu \gt 2$ (and infinite for $\nu \le 2$). Compare scales across likelihoods with care.

This is the syllabus's "connects directly to why the forecasting model supports Normal and Student-t likelihoods". In your forecasting model, choosing Student-t protects the piecewise trend and the Fourier seasonality from being bent by a few unexplained days, and keeps the noise scale (and therefore the prediction intervals) honest on ordinary days. Your A/B framework offers the same choice for continuous metrics: a Student-t likelihood keeps a few extreme users from dominating the estimated difference between variants, without capping or deleting their data. In both, compare the fits with posterior predictive checks of tail behaviour (Chapter 6.8).

"The Student-t likelihood removes outliers."

"The Student-t assigns more probability to extreme residuals, so they exert less influence than under a Normal likelihood." Nothing is removed; every point is still in the likelihood.

Model answer: "A Normal likelihood is equivalent to squared loss, so a single large residual has unbounded influence on the fitted trend and inflates σ, widening all intervals. A Student-t has heavier tails: its log-density grows only logarithmically in the residual, so far points get small weight, roughly $(\nu+1)/(\nu + r^2)$. That's why the model offers both: Normal when the residuals look Normal, Student-t when there are occasional unexplained shocks. $\nu$ controls how heavy the tails are."

Normal likelihood ≙ squared loss: mean-like centre, SD-like scale, breakdown 0, unbounded influence.

Student-t: $-\log p \propto \frac{\nu+1}{2}\log(1 + r^2/\nu)$, weights $(\nu+1)/(\nu + r^2)$: far points count little. SD $= \sigma\sqrt{\nu/(\nu-2)}$.

Say it right: Student-t does not remove outliers; it gives extreme residuals more probability so they have less influence.

Quick check: with $\nu = 4$, what weight does a residual of 4 scales get, compared with one of 1 scale?

$r = 4$: $w = 5/(4 + 16) = 0.25$. $r = 1$: $w = 5/(4 + 1) = 1$. The far point counts a quarter as much as the ordinary one; under a Normal both would count fully.

Recap, cheat sheet and practice

  • Centre: the mean uses every value (balance point), the median only the order (middle value), the mode the most frequent value. Trimmed and winsorized means sit in between.
  • Spread: variance and SD (squared distances, fragile), range (grows with $n$), IQR (middle half), MAD (median distance from the median; ×1.4826 to estimate $\sigma$ for Normal data).
  • Shape: skewness (cubed distances: which side has the long tail) and excess kurtosis (fourth powers: how heavy the tails are; Normal 0). Both are noisy and outlier-sensitive.
  • Outliers: errors, special events or heavy-tail draws. Tukey's 1.5×IQR fences are a convention that flags about 0.7% of clean Normal data. Flag, investigate, then fix, model or keep.
  • Robustness: breakdown point (mean 0, trimmed mean $\alpha$, IQR 25%, median and MAD 50%) and influence (mean unbounded, median bounded, Student-t redescending). Robustness costs some efficiency on clean data.
  • In your models: a Normal likelihood behaves like the mean and SD (fragile under spikes); a Student-t likelihood gives extreme residuals more probability so they have less influence.

Cheat sheet

StatisticFormula / ruleBreakdownCode
Mean$\frac1n\sum x_i$0np.mean
Medianmiddle sorted value50%np.median
$\alpha$-trimmed meandrop $\lfloor\alpha n\rfloor$ each end, average$\alpha$stats.trim_mean(x, a)
Winsorized meanclip $\lfloor\alpha n\rfloor$ each end, average$\alpha$mstats.winsorize(x, (a, a)).mean()
SD$\sqrt{\frac{1}{n-1}\sum(x_i - \bar x)^2}$0x.std(ddof=1)
Rangemax − min (grows with $n$)0np.ptp
IQR$Q_3 - Q_1$; $\sigma \approx$ IQR/1.34925%stats.iqr
MADmedian$|x_i - $median$|$; $\sigma \approx 1.4826\,$MAD50%median_abs_deviation(x, scale='normal')
Skewness$m_3/m_2^{3/2}$0stats.skew (pandas: adjusted)
Excess kurtosis$m_4/m_2^2 - 3$ (Normal 0)0stats.kurtosis
Tukey fences$Q_1 - 1.5\,$IQR, $Q_3 + 1.5\,$IQR—~0.7% of Normal data flagged
Code it · Python
import numpy as np
import pandas as pd
from scipy import stats
from scipy.stats import mstats

orders = np.array([12, 15, 11, 14, 13, 15, 90])          # 7 days, one promotion day

# Centre
print(orders.mean(), np.median(orders))                   # 24.29  14.0
print(stats.mode(orders, keepdims=False).mode)            # 15
x = np.array([2, 5, 6, 7, 7, 8, 8, 9, 11, 37])
print(x.mean(), stats.trim_mean(x, 0.1))                  # 10.0  7.625  (drops 1 value at each end)
print(mstats.winsorize(x, limits=(0.1, 0.1)).mean())      # 7.7          (2 -> 5, 37 -> 11)

# Spread (watch the ddof defaults!)
print(orders.std(ddof=1), np.std(orders), pd.Series(orders).std())   # 29.02  26.86 (ddof=0!)  29.02
print(np.ptp(orders), stats.iqr(orders))                  # range 79, IQR 2.5
print(stats.median_abs_deviation(orders))                 # 1.0    raw MAD
print(stats.median_abs_deviation(orders, scale='normal')) # 1.4826 x 1.4826: estimates sigma for Normal data

# Shape
z = np.array([1, 2, 3, 4, 10])
print(stats.skew(z), pd.Series(z).skew())                 # 1.138 (SciPy g1)   1.697 (pandas adjusted G1)
print(stats.kurtosis(z), stats.kurtosis(z, fisher=False)) # -0.212 (excess)    2.788 (plain m4/m2^2)

# Tukey fences (1.5 x IQR rule)
q1, q3 = np.quantile(orders, [0.25, 0.75])               # NumPy's default 'linear' method
lo, hi = q1 - 1.5 * (q3 - q1), q3 + 1.5 * (q3 - q1)
print(q1, q3, lo, hi, orders[(orders < lo) | (orders > hi)])   # 12.5 15.0 8.75 18.75 [90]

# Clean Normal data still gets flagged about 0.7% of the time
rng = np.random.default_rng(0)
y = rng.standard_normal(100_000)
a, b = np.quantile(y, [0.25, 0.75])
print(np.mean((y < a - 1.5 * (b - a)) | (y > b + 1.5 * (b - a))))   # 0.00724

# Normal fit vs Student-t fit when three spike days are present
resid = np.concatenate([rng.standard_normal(60), [6.0, 7.5, 9.0]])
print(resid.mean(), resid.std(ddof=1))                    # 0.279  1.835   Normal: centre pulled up, sigma inflated
df, loc, scale = stats.t.fit(resid, fix_df=3)            # Student-t MLE with nu fixed at 3
print(loc, scale)                                         # -0.005  0.685  centre and scale barely affected
Test yourself

1. Which statistic has a breakdown point of 50%?

The median only moves arbitrarily far once half of the data are corrupted. The mean, SD and range can be ruined by a single value (breakdown point 0).

2. For the data 3, 4, 5, 6, 100…

Mean $= 118/5 = 23.6$; median = 3rd sorted value $= 5$. The 100 pulls the mean far above every typical value.

3. scipy.stats.kurtosis returns about 0 for a large Normal sample because…

The plain kurtosis of a Normal is 3. SciPy subtracts 3 by default (fisher=True), so 0 means "Normal-like tails".

4. You apply the 1.5×IQR rule to 2 000 clean values from a Normal distribution. About how many are flagged?

About 0.7% of Normal values lie beyond the fences: $0.007 \times 2000 \approx 14$. Flagged does not mean wrong.

5. What is the (raw) MAD of 1, 2, 3, 4, 100?

Median 3; absolute distances 2, 1, 0, 1, 97; their median is 1. The 100 has no effect beyond being "the largest distance".

6. Which sentence about a Student-t likelihood is correct?

Every point stays in the likelihood; far points just get small weight. σ is a scale (SD $= \sigma\sqrt{\nu/(\nu-2)}$), and intervals can be narrower on ordinary days but have heavier tails.

Practice problems

A. For 4, 7, 8, 8, 9, 10, 11, 12, 13, 60 compute the mean, the median, the 20% trimmed mean, the MAD (raw and scaled), the IQR (NumPy rule) and the Tukey fences.

Mean $= 142/10 = 14.2$. Median $= (9 + 10)/2 = 9.5$. 20% trimmed: cut $\lfloor 0.2 \times 10\rfloor = 2$ at each end, average $8, 8, 9, 10, 11, 12$: $58/6 \approx 9.67$. MAD: distances from 9.5 are $5.5, 2.5, 1.5, 1.5, 0.5, 0.5, 1.5, 2.5, 3.5, 50.5$; sorted, the middle two are 1.5 and 2.5, so MAD $= 2.0$, scaled $\approx 2.97$. Quartiles: $Q_1$ at position $1 + 9 \times 0.25 = 3.25$, between 8 and 8: $8$; $Q_3$ at position $7.75$, between 11 and 12: $11.75$; IQR $= 3.75$. Fences: $8 - 5.625 = 2.375$ and $11.75 + 5.625 = 17.375$, so 60 is flagged.

B. Explain to an interviewer why a forecasting model would offer both a Normal and a Student-t likelihood.

"A Normal likelihood corresponds to squared loss: the fitted level and the noise σ behave like a mean and an SD, which have breakdown point zero. A few unexplained spike days pull the trend toward them and inflate σ, which widens the intervals on every ordinary day. A Student-t has heavier tails; its negative log-density grows only like $\log(1 + r^2/\nu)$, so far residuals get small weight $(\nu+1)/(\nu+r^2)$. It doesn't remove outliers; it gives them more probability so they have less influence. I'd use the Normal when residual diagnostics look Normal, and the Student-t when the Q-Q plot or tail counts show heavy tails, and I'd still model recurring spikes such as holidays explicitly."

C. Compute the skewness $g_1$ of 0, 0, 0, 0, 10.

Mean 2; distances $-2, -2, -2, -2, 8$. $m_2 = (4 \times 4 + 64)/5 = 16$. $m_3 = (4 \times (-8) + 512)/5 = 96$. $g_1 = 96/16^{1.5} = 96/64 = 1.5$: strongly right-skewed.

D. Revenue for 1 000 users: 950 users spend 0, 45 spend 20, and 5 spend 2 000. Compute the mean, the median and the 1% winsorized mean. Which one answers "average revenue per user"?

Mean $= (45 \times 20 + 5 \times 2000)/1000 = (900 + 10\,000)/1000 = 10.9$. Median $= 0$. 1% winsorizing clips the top $\lfloor 0.01 \times 1000 \rfloor = 10$ values to the 990th sorted value, which is 20; so the five 2 000s become 20, and the mean becomes $50 \times 20/1000 = 1.0$. Only the plain mean (10.9) is the average revenue per user; the winsorized mean is a different, much smaller quantity, and the median says nothing here.

E. A metric has median 50 and MAD 4. Is the value 75 unusual by the robust z-score rule? What can happen with the ordinary z-score?

Robust z $= (75 - 50)/(1.4826 \times 4) = 25/5.93 \approx 4.2 \gt 3.5$: flagged. If a few large values have inflated the ordinary SD to, say, 12, then $z = 25/12 \approx 2.1$ and 75 is not flagged: the outliers have masked themselves.

F. With a Student-t likelihood, $\nu = 5$, how much weight does a residual of 5 scales get relative to a residual of 0? What does that mean for the fit?

$w(5) = 6/(5 + 25) = 0.2$ and $w(0) = 6/5 = 1.2$, so the far point counts $0.2/1.2 \approx 1/6$ as much as a point at the centre. It still enters the likelihood, but it can barely move the fitted level or inflate the scale. Under a Normal likelihood both would have equal weight, and the far point's squared residual (25) would dominate.

Chapter 4.15 · Syllabus Modules 2.10–2.11, 7

Covariance and correlation

Until now we looked at one variable at a time: its average, its spread, its shape. Real questions are about two things at once. Do warmer days bring more orders? Do users who see more emails convert more often? Does a candidate regressor move with demand, or with another regressor you already have? Covariance and correlation are the first tools for "do these two move together, and how tightly?"

  • Read the sign of covariance from a picture: the "rectangles from the mean point" view
  • See why covariance depends on units, and how dividing by both standard deviations gives correlation, a unit-free number between −1 and 1
  • Compute and interpret Pearson's r: what $r = -1, 0, +1$ mean, why it only sees straight lines, and why one outlier can fool it
  • Say precisely why zero correlation is not independence and why correlation is not causation
  • Use Spearman (ranks) and Kendall (agreeing pairs) for curved-but-monotone, ordinal or outlier-prone data
  • Remove a third variable with partial correlation (the residual method)
  • Read a correlation matrix heatmap before you add many regressors to your forecasting model

Covariance: do two things move together? core

Think about a small shop that sells cold drinks. On a warmer-than-usual day it usually sells more than usual. On a cooler-than-usual day it usually sells less. Temperature and orders move together.

How do we turn "move together" into a number? Look at one day. Ask two questions: was the temperature above or below its average? Were the orders above or below their average? If both answers agree (both above, or both below), that day is a vote for "together". If they disagree, it is a vote for "opposite". Covariance adds up these votes, and big surprises count more than small ones.

Three ways to say it:

  • Picture: draw a cross through the average point. Points in the top-right and bottom-left corners push covariance up; points in the top-left and bottom-right corners push it down.
  • Numbers: a day 4 °C warmer than average with 20 more orders than average contributes $(+4)\times(+20) = +80$, a strong vote for "together".
  • Slogan: covariance is the average of "how far above its mean x is" times "how far above its mean y is".

Five days of data. $x$ = temperature (°C), $y$ = orders.

day$x$ (°C)$y$ (orders)$x - \bar x$$y - \bar y$product
114110−4−20+80
216130−200
3181200−100
420150+2+20+40
522140+4+10+40
  1. Means: $\bar x = (14+16+18+20+22)/5 = 90/5 = 18$ °C and $\bar y = (110+130+120+150+140)/5 = 650/5 = 130$ orders.
  2. Distances from the means: the columns $x-\bar x$ and $y-\bar y$ above.
  3. Multiply each pair: $80, 0, 0, 40, 40$. No product is negative: no day "disagrees".
  4. Add them: $80 + 0 + 0 + 40 + 40 = 160$.
  5. For a sample, divide by $n - 1 = 4$ (the same reason as for the sample variance in Chapter 4.5): $s_{xy} = 160/4 = 40$.
  6. Units: $x$ is in °C and $y$ in orders, so the covariance is 40 °C·orders. Positive: they move together.

A random-variable version. Roll two fair dice. Let $X$ = the first die and $Y$ = the total. Then $E[X] = 3.5$, $E[Y] = 7$, and $E[XY] = E[X^2] + E[X]\,E[\text{second die}] = \tfrac{91}{6} + 3.5 \times 3.5 = 15.1\overline{6} + 12.25 = 27.41\overline{6}$. So $Cov(X, Y) = E[XY] - E[X]E[Y] = 27.41\overline{6} - 24.5 = 2.91\overline{6} = \tfrac{35}{12}$. Positive: a big first die tends to give a big total.

For two random variables $X$ and $Y$ with means $\mu_X = E[X]$ and $\mu_Y = E[Y]$, the covariance is

$$Cov(X, Y) = E\big[(X - \mu_X)(Y - \mu_Y)\big] = E[XY] - E[X]\,E[Y].$$

For paired data $(x_1, y_1), \dots, (x_n, y_n)$ the sample covariance is

$$s_{xy} = \frac{1}{n-1}\sum_{i=1}^{n} (x_i - \bar x)(y_i - \bar y).$$
  • Sign: positive = they tend to be above (or below) their means together; negative = when one is above, the other tends to be below; zero = no straight-line tendency either way.
  • Units: units of $X$ times units of $Y$ (°C·orders). Its size has no fixed scale.
  • Scaling rule: $Cov(aX + b,\; cY + d) = ac\,Cov(X, Y)$. Shifting does nothing; stretching multiplies. This is the scale dependence.
  • $Cov(X, X) = Var(X)$, and $Cov(X, Y) = Cov(Y, X)$.
  • Variance of a sum (the missing piece from Chapter 4.5): $Var(X + Y) = Var(X) + Var(Y) + 2\,Cov(X, Y)$ and $Var(X - Y) = Var(X) + Var(Y) - 2\,Cov(X, Y)$. Only when $Cov = 0$ do variances simply add.
  • If $X$ and $Y$ are independent, $Cov(X, Y) = 0$. The reverse is not true (see zero correlation is not independence).
Why do we need it?

Means and variances describe one variable at a time. To combine variables (add them, subtract them, predict one from the other) we must know how they move together. Without covariance, the variance of a sum or a difference cannot be computed.

Where is it used?

The variance of a difference of two metrics measured on the same users, portfolio risk, the covariance matrix of a multivariate Normal (Chapter 5.15), PCA, the slope of a regression line ($b = s_{xy}/s_x^2$), and CUPED variance reduction (Chapter 5.12).

How is it used?

Call np.cov(x, y) (it returns the 2×2 matrix; the covariance is entry [0, 1]) or df.cov() in pandas. Read the sign, not the size. To judge strength, divide by both standard deviations: that is correlation, next section.

x̄ ȳ x above, y above (+) × (+) = + x below, y below (−) × (−) = + x below, y above (−) × (+) = − x above, y below (+) × (−) = − a "together" day: its area counts + an "opposite" day: its area counts −
Draw a cross at the mean point $(\bar x, \bar y)$. Each data point makes a rectangle with the cross's centre. Rectangles in the green corners count as positive area, in the red corners as negative area. The covariance is (roughly) the average signed area.

Each blue dot is a day. The purple cross marks the average temperature and the average orders. Every dot draws a rectangle to the cross: green rectangles (top-right, bottom-left) add to the covariance, red ones (top-left, bottom-right) subtract. (1) Start with Worked example and check the readout against the table above. (2) Drag the hottest day down to 100 orders: its rectangle turns red and the covariance falls. (3) Press Move opposite and No pattern, and watch the balance of green and red.

"A bigger covariance means a stronger relationship."

The size of a covariance depends on the units. The same five days give 40 in °C·orders but 72 in °F·orders and 0.4 in °C·(hundreds of orders). Only the sign is meaningful on its own. Use correlation to judge strength.

"np.cov(x, y) gives me the covariance."

It gives the 2×2 covariance matrix $\begin{bmatrix} s_x^2 & s_{xy} \\ s_{xy} & s_y^2 \end{bmatrix}$; the covariance is np.cov(x, y)[0, 1]. Also, np.cov divides by $n-1$ by default, while np.var divides by $n$, so np.cov(x, y)[0, 0] and np.var(x) disagree unless you pass ddof=1 to np.var.

"Variances always add: $Var(X + Y) = Var(X) + Var(Y)$."

Only when $Cov(X, Y) = 0$ (for example when they are independent). Otherwise add $2\,Cov(X, Y)$. Two dice: $Var(X + Y) = \tfrac{35}{12} + \tfrac{35}{12} + 0 = \tfrac{70}{12}$, because two separate dice have covariance 0.

In an A/B framework like yours, two metrics measured on the same users (for example orders and revenue) are not independent, so the variance of a combined or derived metric needs the covariance term: $Var(X - Y) = Var(X) + Var(Y) - 2\,Cov(X, Y)$. The same idea powers CUPED (Chapter 5.12): a pre-experiment metric that covaries with the outcome can be used to remove noise. In your forecasting model, two regressors with a large covariance (temperature and "feels-like" temperature) carry almost the same information and compete for the same credit (see the correlation matrix).

$Cov(X,Y) = E[(X-\mu_X)(Y-\mu_Y)] = E[XY] - E[X]E[Y]$; sample: $s_{xy} = \tfrac{1}{n-1}\sum (x_i-\bar x)(y_i-\bar y)$.

Sign = direction (together / opposite). Size depends on units: $Cov(aX+b, cY+d) = ac\,Cov(X,Y)$.

$Var(X \pm Y) = Var X + Var Y \pm 2Cov(X,Y)$. Trap: np.cov returns a matrix and uses $n-1$.

Quick check: every day had exactly 130 orders (no change at all). What is the covariance with temperature?

Every $y_i - \bar y = 130 - 130 = 0$, so every product is 0 and $s_{xy} = 0$. A variable that never moves cannot move with anything. (Its correlation, as we will see, is not even defined, because its standard deviation is 0.)

Correlation: covariance with the units taken out core

Covariance has a measuring-stick problem. Switch temperature from Celsius to Fahrenheit and the covariance jumps from 40 to 72, even though nothing about the days changed. We cannot say whether 40 is "a lot".

The fix: measure each variable in its own standard deviations. A day that is "1.26 standard deviations warmer than usual" means the same thing in Celsius, Fahrenheit or kelvin. These rescaled values are called z-scores (a z-score says how many standard deviations a value is above its mean). The covariance of the two z-scores is a pure number, always between −1 and +1. That number is the correlation.

Three ways to say it:

  • Picture: stretch or squash each axis until both clouds have the same spread. Now only the shape of the cloud is left, and correlation describes that shape.
  • Numbers: $40 / (3.16 \times 15.81) = 40/50 = 0.8$, and you get 0.8 again in Fahrenheit: $72 / (5.69 \times 15.81) = 0.8$.
  • Slogan: correlation = covariance divided by both standard deviations = the covariance of z-scores.

Same five days. From the last section: $s_{xy} = 40$. Now the standard deviations:

  1. $s_x^2 = \big((-4)^2 + (-2)^2 + 0^2 + 2^2 + 4^2\big)/4 = 40/4 = 10$, so $s_x = \sqrt{10} \approx 3.162$ °C.
  2. $s_y^2 = \big((-20)^2 + 0^2 + (-10)^2 + 20^2 + 10^2\big)/4 = 1000/4 = 250$, so $s_y = \sqrt{250} \approx 15.81$ orders.
  3. Correlation: $r = \dfrac{s_{xy}}{s_x\, s_y} = \dfrac{40}{3.162 \times 15.81} = \dfrac{40}{50} = 0.8$. The units (°C·orders on top, °C × orders below) cancel.
  4. The z-score route gives the same answer. z-scores of $x$: $-4/3.162 = -1.265$, then $-0.632, 0, 0.632, 1.265$. z-scores of $y$: $-20/15.81 = -1.265$, then $0, -0.632, 1.265, 0.632$.
  5. Products of matching z-scores: $1.6, 0, 0, 0.8, 0.8$. Sum $= 3.2$. Divide by $n - 1 = 4$: $3.2/4 = 0.8$. ✓
  6. In Fahrenheit, $x' = 1.8x + 32$: covariance $1.8 \times 40 = 72$, $s_{x'} = 1.8 \times 3.162 = 5.692$, so $r = 72/(5.692 \times 15.81) = 0.8$ again.

The (population) correlation of $X$ and $Y$ is

$$\rho_{XY} = \frac{Cov(X, Y)}{\sigma_X\, \sigma_Y}, \qquad \text{and its sample version is } r = \frac{s_{xy}}{s_x\, s_y}.$$

Here $\rho$ ("rho") is a fixed property of the process and $r$ is computed from data, so $r$ changes from sample to sample and estimates $\rho$.

Why standardizing gives correlation (derivation). Let $Z_X = (X - \mu_X)/\sigma_X$ and $Z_Y = (Y - \mu_Y)/\sigma_Y$. Using the scaling rule $Cov(aX + b, cY + d) = ac\,Cov(X, Y)$ with $a = 1/\sigma_X$ and $c = 1/\sigma_Y$:

$$Cov(Z_X, Z_Y) = \frac{1}{\sigma_X}\cdot\frac{1}{\sigma_Y}\,Cov(X, Y) = \rho_{XY}.$$

Why it lies between −1 and +1. z-scores have variance 1. A variance can never be negative, so

$$0 \le Var(Z_X + Z_Y) = 1 + 1 + 2\rho \;\Rightarrow\; \rho \ge -1, \qquad 0 \le Var(Z_X - Z_Y) = 2 - 2\rho \;\Rightarrow\; \rho \le 1.$$

$\rho = 1$ exactly when $Var(Z_X - Z_Y) = 0$, that is when $Z_Y = Z_X$: every point lies on one rising straight line. Likewise $\rho = -1$ means one falling straight line.

  • Unit-free: unchanged by $X \to aX + b$ with $a \gt 0$. If $a \lt 0$ only the sign flips.
  • Not defined when $\sigma_X = 0$ or $\sigma_Y = 0$ (you would divide by zero). Software returns nan.
  • Geometry: $r$ is the cosine of the angle between the centred data vectors $(x_i - \bar x)$ and $(y_i - \bar y)$, the cosine similarity from the Linear Algebra guide.
Why do we need it?

We want to compare strengths across very different pairs (temperature and orders, price and conversions, two regressors) on one fixed scale. Correlation gives every pair the same ruler, from −1 to +1, whatever the units.

Where is it used?

Feature screening before modelling, correlation matrices of regressors, the ρ parameter of a bivariate Normal, posterior correlation between parameters (why full-rank guides exist, Chapter 6.13), and the autocorrelation of a time series (Chapter 7.3).

How is it used?

np.corrcoef(x, y)[0, 1], scipy.stats.pearsonr(x, y) or df.corr(). Always look at the scatter plot as well, and check that neither variable is constant (the result would be nan).

raw pair (22 °C, 140) units: °C, orders − mean centred (+4, +10) still in units ÷ sd z-scores (+1.26, +0.63) no units left × and average correlation r = 0.8 between −1 and +1 this day's product of z-scores: 1.26 × 0.63 = 0.8; average all five products (÷ 4) to get r
Correlation is covariance computed after each variable has been centred and divided by its own standard deviation. The last step averages the products of z-scores (dividing by $n-1$).

The left plot is fixed: the five example days in °C and orders. The right plot shows the same days in the units you pick. (1) Choose °F: the cloud looks identical, only the axis numbers change; the covariance becomes 72 but $r$ stays 0.8. (2) Choose hundreds of orders: the covariance shrinks to 0.4, $r$ is still 0.8. (3) Choose degrees below 30 °C (a flipped scale): the cloud mirrors and only the sign of $r$ changes.

"Correlation 0.8 means 80% of the points are on the line" or "x explains 80% of y".

$r$ is not a percentage. Its square, $r^2 = 0.64$, is the fraction of the variance of $y$ explained by the best straight line through the points (you will meet this as $R^2$ in Chapter 5.13). So $r = 0.8$ leaves 36% of the variance unexplained.

"Dividing by $n$ instead of $n-1$ changes the correlation."

It does not. The same factor appears in the covariance and in both standard deviations, and it cancels: $r = \dfrac{\sum (x_i - \bar x)(y_i - \bar y)}{\sqrt{\sum (x_i - \bar x)^2}\sqrt{\sum (y_i - \bar y)^2}}$. (The covariance itself does change.)

"If one variable is constant, the correlation is 0."

It is not defined: you would divide by $s = 0$. NumPy prints a warning and returns nan. Check for constant columns (a regressor that is always 0 in the training window) before computing a correlation matrix.

"Correlation is just covariance on a nicer scale, so they tell you the same thing."

They share the sign, but only correlation has a fixed scale. Covariance mixes the strength of the link with the units and spreads of the two variables; correlation removes the units and spreads, so only the tightness of the straight-line pattern is left.

Model answer: "Correlation is the covariance of the standardized variables, $\rho = Cov(X, Y)/(\sigma_X\sigma_Y)$. Standardizing removes units, which forces it into $[-1, 1]$, with $\pm 1$ meaning an exact straight line."

$\rho = \dfrac{Cov(X,Y)}{\sigma_X\sigma_Y} = Cov(Z_X, Z_Y)$; sample $r = \dfrac{s_{xy}}{s_x s_y}$.

Always in $[-1, 1]$ (because $Var(Z_X \pm Z_Y) \ge 0$); $\pm1$ = exact straight line. Unit-free; a negative rescaling flips the sign.

Traps: $r$ is not a percentage ($r^2$ is the variance explained); a constant variable makes $r$ undefined.

Quick check: what is the correlation between $x$ and $y = 3x - 2$? And between $x$ and $y = -0.5x + 7$?

Both are exact straight lines. The first rises, so $r = +1$. The second falls, so $r = -1$. The steepness (3 or −0.5) does not matter, only the direction and the fact that every point is exactly on the line.

Pearson's r: how tightly do the points hug a straight line? core

Draw the scatter plot. If the points form a thin cigar that rises from left to right, the two variables have a strong positive straight-line link. A thin cigar that falls: strong negative. A round blob: no straight-line link at all. Pearson's $r$ (the correlation of the last section, computed from data) turns that picture into one number.

Two things decide $r$: the direction of the cigar (rising or falling gives the sign) and how thin it is (thinner = closer to ±1). How steep the cigar looks does not matter: stretch the whole picture upward and it looks steeper, but $r$ does not change.

Three ways to say it:

  • Picture: thin rising cigar ≈ +0.9, round blob ≈ 0, thin falling cigar ≈ −0.9.
  • Numbers: $r = +1$: every point exactly on a rising line; $r = -1$: exactly on a falling line; $r = 0$: no straight-line trend.
  • Slogan: $r$ measures straight-line togetherness, and nothing else.

The five days again, with the three sums written out (this is the form most textbooks use):

  1. Cross-products: $\sum (x_i - \bar x)(y_i - \bar y) = 80 + 0 + 0 + 40 + 40 = 160$.
  2. Squares for $x$: $\sum (x_i - \bar x)^2 = 16 + 4 + 0 + 4 + 16 = 40$.
  3. Squares for $y$: $\sum (y_i - \bar y)^2 = 400 + 0 + 100 + 400 + 100 = 1000$.
  4. $r = \dfrac{160}{\sqrt{40}\,\sqrt{1000}} = \dfrac{160}{\sqrt{40\,000}} = \dfrac{160}{200} = 0.8$.
  5. Compare with the slope of the best straight line: $b = 160/40 = 4$ orders per °C. If we counted orders in hundreds, the slope would be $0.04$, but $r$ would still be $0.8$. Slope and $r$ answer different questions.
  6. They are linked: $b = r \cdot s_y / s_x = 0.8 \times 15.81 / 3.162 = 0.8 \times 5 = 4$. ✓

Pearson's correlation coefficient for paired data $(x_i, y_i)$, $i = 1, \dots, n$:

$$r = \frac{\sum_{i=1}^{n}(x_i - \bar x)(y_i - \bar y)}{\sqrt{\sum_{i=1}^{n}(x_i - \bar x)^2}\;\sqrt{\sum_{i=1}^{n}(y_i - \bar y)^2}}.$$
  • $r = +1$ or $-1$: all points on one straight line (rising or falling). $r = 0$: the best straight line is flat.
  • It measures linear (straight-line) association only. A perfect curve can have $r$ well below 1, or even $r = 0$.
  • It is sensitive to outliers: one extreme point can create or destroy a large $r$, because every term uses distances from the mean, and the sums of squares let far points dominate.
  • Slope link: the least-squares slope of $y$ on $x$ is $b = r\, s_y/s_x$. $r^2$ is the share of the variance of $y$ explained by that line.
  • Words often attached to $|r|$ (a rule of thumb; what counts as "strong" depends on the field): below 0.1 negligible, 0.1–0.3 weak, 0.3–0.7 moderate, above 0.7 strong.
Why do we need it?

A scatter plot is hard to compare across hundreds of variable pairs. One number on a fixed −1 to +1 scale lets us screen, rank and report straight-line relationships quickly.

Where is it used?

Feature screening, checking regressors against demand in a forecasting model, the correlation between two metrics in an experiment, model evaluation (correlation between predicted and actual values), and the "r" reported next to almost every scatter plot in papers.

How is it used?

scipy.stats.pearsonr(x, y) returns $r$ and a p-value; np.corrcoef and df.corr() return $r$ only. Plot first, then compute. If the cloud is curved or has outliers, also compute Spearman.

r = −1.0 r = −0.8 r = −0.4 r = 0.0 r = 0.4 r = 0.8 r = 1.0
Seven clouds of 36 points with $r = -1, -0.8, -0.4, 0, 0.4, 0.8, 1$. The sign is the direction of the tilt; the size is how thin the cloud is around a straight line.

Ten draggable days. Green rectangles push the covariance up, red ones push it down (purple cross = the mean point). Watch the four readouts change as you drag. (1) Press Negative, Near zero and Positive and compare the shapes. (2) In Positive, drag one point far to the bottom-right: Pearson drops much faster than Spearman and Kendall. (3) Press One outlier: nine points with no pattern plus one far point give Pearson ≈ 0.88 while Spearman stays near 0.16. (4) Press U-shape (all four are 0, yet y clearly depends on x) and Rising curve (Spearman and Kendall are exactly 1, Pearson is not).

These are Anscombe's four datasets (1973). Switch between I, II, III and IV. The means, variances, correlation (0.82) and best-fit line ($\hat y = 3 + 0.5x$) are the same in all four, but the stories are completely different. Read the story under the plot for each one.

A new cloud of 60 points appears with a hidden correlation. Set the slider to your guess, then press Reveal. Play five or six rounds: most people overestimate small correlations and underestimate how "fat" a cloud with r = 0.5 looks. Notice that r = 0.5 still looks quite messy.

"A steeper line means a larger correlation."

Steepness is the slope, which has units ("4 orders per °C"). Correlation is about how closely the points follow a line, whatever its steepness. A nearly flat line with points tightly on it has $|r|$ close to 1. (A perfectly flat line is the exception: then $y$ has no spread and $r$ is not defined.)

"r = 0.6 is twice as strong as r = 0.3."

$r$ is not on a ratio scale. In explained variance, $0.6^2 = 0.36$ versus $0.3^2 = 0.09$: four times as much.

"One number tells me everything about the relationship."

Anscombe's four datasets all have $r = 0.82$ and the same best-fit line, yet one is a curve, one is a line with an outlier and one is a single influential point. Always plot the data.

"r tells me how much y changes when x increases by one."

That is the slope $b = r\,s_y/s_x$, which has units. $r$ is unit-free: it tells you how tightly the points follow a straight line and in which direction.

Model answer: "Pearson's r measures the strength and direction of the linear association. It is the covariance of the standardized variables, it is sensitive to outliers, and it can be near zero even when there is a strong non-linear relationship, so I always look at the scatter plot and often report Spearman as well."

$r = \dfrac{\sum (x_i-\bar x)(y_i-\bar y)}{\sqrt{\sum (x_i-\bar x)^2}\sqrt{\sum (y_i-\bar y)^2}}$ — sign = direction, size = how thin the cloud is around a line.

$\pm1$ = exact line; 0 = no straight-line trend. Slope $b = r\,s_y/s_x$; $r^2$ = variance explained by the line.

Traps: linear only, outlier-sensitive, not a slope, not a percentage. Plot first (Anscombe).

Quick check: in the lab, why does dragging one point far to the bottom-right hurt Pearson so much more than Spearman?

Pearson uses the actual distances from the mean, so a point that is very far away makes a huge red rectangle that outweighs all the others. Spearman only uses the point's rank: however far you drag it, it can only become "the largest x and the smallest y", so its influence is capped.

Zero correlation does not mean independence core

Think of heating and cooling demand for electricity. On very cold days people heat; on very hot days they cool; on mild days they do neither. Plot energy use against temperature and you get a U-shape. Temperature clearly controls energy use. But the left half of the U falls and the right half rises, so the straight-line trends cancel, and the correlation is close to 0.

Correlation only looks for a straight-line trend. "No straight-line trend" is a much weaker statement than "no relationship at all". Independent means something much stronger: knowing $X$ tells you nothing about $Y$ (not its average, not its spread, not anything).

Three ways to say it:

  • Picture: a perfect U, a ring, an X: all have $r \approx 0$, and in all of them $X$ tells you a lot about $Y$.
  • Numbers: $X \in \{-1, 0, 1\}$ and $Y = X^2$ give $Cov(X, Y) = 0$, yet $Y$ is completely determined by $X$.
  • Slogan: independent ⇒ uncorrelated, but uncorrelated ⇏ independent.

Let $X$ be $-1$, $0$ or $1$, each with probability $\tfrac13$, and let $Y = X^2$.

  1. Possible pairs $(x, y)$: $(-1, 1)$, $(0, 0)$, $(1, 1)$, each with probability $\tfrac13$.
  2. $E[X] = (-1 + 0 + 1)/3 = 0$ and $E[Y] = (1 + 0 + 1)/3 = \tfrac23$.
  3. $E[XY]$: the products $xy$ are $-1, 0, 1$, so $E[XY] = (-1 + 0 + 1)/3 = 0$.
  4. $Cov(X, Y) = E[XY] - E[X]E[Y] = 0 - 0 \times \tfrac23 = 0$. So the correlation is 0.
  5. Independence test: $P(Y = 0) = \tfrac13$, but $P(Y = 0 \mid X = 0) = 1$. Knowing $X$ changed the probability, so $X$ and $Y$ are not independent.

$X$ and $Y$ are independent when $P(X \in A,\; Y \in B) = P(X \in A)\,P(Y \in B)$ for all sets $A$ and $B$ (see Chapter 4.3). Equivalently: the whole distribution of $Y$ is the same whatever value $X$ takes.

$X$ and $Y$ are uncorrelated when $Cov(X, Y) = 0$, that is $E[XY] = E[X]E[Y]$: only the straight-line trend is absent.

  • Independent (and finite variances) ⇒ uncorrelated. The reverse fails in general.
  • One special case: if $(X, Y)$ is jointly (bivariate) Normal, then uncorrelated ⇔ independent. Two variables that are each Normal on their own are not enough.
  • A useful check: if $Y$ were independent of $X$, every vertical slice of the scatter plot would show the same mean and the same spread of $Y$ (up to noise).
Why do we need it?

Screening variables by correlation alone throws away real but curved relationships, and "the residuals are uncorrelated with x" can hide a pattern the model missed. We need to know exactly what a zero correlation does and does not rule out.

Where is it used?

Residual-versus-regressor plots in regression and forecasting, feature selection (a U-shaped effect of temperature on demand), mutual information and other dependence measures used in feature selection, and the "uncorrelated but dependent" residuals of volatile time series (Chapter 7.17).

How is it used?

Never conclude "no relationship" from $r \approx 0$. Plot $y$ against $x$, look at slice means and slice spreads, and if needed try a curve (a squared term, a spline) or a dependence measure that is not limited to straight lines.

Pick a shape. Every one has Pearson (and usually Spearman) close to 0. The purple line joins the average y in each vertical slice, and the purple band shows the spread in each slice. If y were independent of x, the line would be flat and the band would have the same width everywhere. In U, V and Wave the slice means move; in Fan the means stay flat but the spread grows; in Ring and X the spread changes from slice to slice. Press New sample a few times.

"r ≈ 0, so x has no effect on y and I can drop it."

$r \approx 0$ only rules out a straight-line trend. A U-shaped effect (energy use against temperature, demand against price around a sweet spot) has $r \approx 0$ and can be the most important variable you have. Plot it.

"X and Y are each Normally distributed and uncorrelated, so they are independent."

That shortcut needs them to be jointly Normal. Counter-example: $X \sim N(0, 1)$ and $Y = SX$ where $S = \pm1$ with a fair coin. $Y$ is also $N(0, 1)$, and $Cov(X, Y) = E[S]E[X^2] = 0$, yet $|Y| = |X|$ always.

"Spearman is 0, so they are independent."

Spearman only detects monotone trends. The U, the ring and the X all have Spearman ≈ 0 too.

In your forecasting model, suppose demand really rises on both very cold and very hot days, but temperature enters as one straight-line regressor. The fitted model cannot capture the U, and the leftover U sits in the residuals. In least-squares regression with an intercept, residuals are exactly uncorrelated with every regressor by construction, so "residuals are uncorrelated with temperature" proves nothing. Plot residuals against each regressor (and against time) and look for slice means that move or spreads that change. The same warning applies to residual autocorrelation: zero autocorrelation does not mean independent residuals (Chapter 7.17).

"The correlation is zero, so the variables are unrelated."

Zero correlation means no linear association. Independence means the whole conditional distribution of $Y$ does not depend on $X$. Independence implies zero correlation, not the other way round.

Model answer: "Take $X$ uniform on $\{-1, 0, 1\}$ and $Y = X^2$: the covariance is 0 but $Y$ is a function of $X$. Only for jointly Normal variables does zero correlation imply independence."

Independent ⇒ $Cov = 0$. $Cov = 0$ ⇏ independent (U, ring, X shapes; $Y = X^2$).

Independence = every slice of the scatter has the same distribution of $Y$. Exception: jointly Normal ⇒ (uncorrelated ⇔ independent).

Trap: residuals uncorrelated with a regressor can still hide a curved pattern. Plot them.

Quick check: give a pair of variables in your own work that could have $r \approx 0$ but a strong relationship.

Demand against temperature when both cold and hot weather raise demand (a U). Conversion rate against price when a middle price converts best (an upside-down U). Residual size against day of week when only weekends are noisy (a change in spread, not in mean). In each case slice means or spreads move, while the straight-line trend is flat.

Correlation is not causation core

Across the summer, a shop's ice-cream orders and its sunscreen orders rise and fall together: their correlation is high. Does buying ice cream make people buy sunscreen? No. A third thing, hot sunny weather, pushes both up. If you hid all the ice cream, sunscreen sales would not drop.

A correlation tells you that two things move together in the data you have. A cause is something that, if you changed it yourself, would change the other thing. These are different questions, and the data alone usually cannot tell them apart.

Three ways to say it:

  • Picture: a hidden puppeteer (the weather) pulls two strings at once; the two puppets dance together without touching.
  • Numbers: in the widget below, ice-cream and sunscreen orders have $r \approx 0.8$, but among days with the same temperature their correlation is about 0.
  • Slogan: correlation says "they move together"; causation says "moving one moves the other".

Four different stories can produce the same correlation between $X$ and $Y$:

  1. X causes Y. A price cut raises orders.
  2. Y causes X (reverse causation). Stores with more staff have more sales, partly because busy stores hire more staff.
  3. Z causes both (confounding). Temperature drives ice-cream orders and sunscreen orders. $Z$ is called a confounder: a variable that affects both and creates a correlation between them.
  4. Chance, often helped by shared trends. Two quantities that both grow over time look correlated even if they have nothing to do with each other. In a simulation of 20 000 pairs of independent random walks of length 60, about 40% had $|r| \gt 0.5$; for pairs of independent white-noise series of the same length, only about 0.004% did.

Only story 1 says that changing $X$ will change $Y$. The correlation is the same number in all four.

Correlation is a property of the joint distribution of the data we observed. Causation is about what would happen under an intervention: $X$ causes $Y$ if setting $X$ to a different value (and changing nothing else) changes the distribution of $Y$.

  • A confounder $Z$ influences both $X$ and $Y$. It can create a correlation with no causal link, or hide a real one.
  • Randomized experiments (A/B tests) break confounding: a coin decides who gets the treatment, so the treatment is independent of every confounder. That is why an experiment can support causal claims that observational correlations cannot (Chapter 5.12 builds the causal language).
  • Correlation can still be useful for prediction without being causal: sunscreen orders can help forecast ice-cream orders, as long as the hidden cause keeps working the same way.
Why do we need it?

Acting on a non-causal correlation wastes money: pushing sunscreen will not sell ice cream. Business decisions ("should we launch this feature?") need causal answers, while forecasts only need stable associations. Knowing which one you have prevents costly mistakes.

Where is it used?

A/B testing (the whole reason it exists), observational studies, marketing attribution, choosing regressors for a forecasting model, and any dashboard where someone says "users who do X retain better, so let's make everyone do X".

How is it used?

For a correlation, ask: could $Y$ cause $X$? Is there a common cause $Z$? Are both just trending over time? Could it be chance among many comparisons? Then either run an experiment, or adjust for the confounders you can measure (partial correlation, regression) and say clearly what remains an assumption.

1 · common cause temperature ice cream sunscreen correlated, no arrow 2 · reverse causation staff sales busy stores hire the arrow may point the "wrong" way 3 · shared trend / chance time → (two unrelated growing series)
Three ways to get a strong correlation without "X causes Y". Arrows mean "causes". Only a randomized experiment, or careful adjustment for the common causes, can separate these stories from a true causal effect.

Each dot is a day: ice-cream orders (x) and sunscreen orders (y), both shown relative to their average. Days are cool, mild or hot. Neither product affects the other: temperature drives both. (1) With the default strength, the overall correlation is about 0.8. (2) Switch on Look inside each temperature group: within one colour the cloud is round and the correlations are near 0. (3) Slide the strength to 0: the overall correlation disappears too.

Top: two independent series (neither knows the other exists), each shown standardized. Bottom: the correlation $r$ for 150 such pairs, with the top pair marked in purple. (1) With Random walks (series that wander and trend, like many business metrics) the histogram is wide: |r| above 0.5 is common. (2) Switch to White noise: the histogram squeezes around 0. (3) Day-to-day changes of the same walks behave like white noise again. Press New sample.

"Users who use feature X retain better, so pushing everyone to X will raise retention."

Engaged users both use more features and retain better (a confounder: engagement). Forcing X on everyone may change nothing. Only a randomized test of pushing X answers the causal question.

"r = 0.9 between two monthly series is overwhelming evidence of a link."

If both series trend over time, a large $r$ is easy to get by chance. Check the correlation of the changes (differences) or of the series after removing trend and seasonality (Chapter 7.4).

"Correlation is useless without causation."

For prediction a stable non-causal correlation can be very useful (sunscreen orders can help forecast ice-cream orders). It becomes dangerous when you act on it, or when the hidden cause changes.

Your A/B framework exists because of this section: random assignment makes the variant independent of every confounder (engagement, device, season), so a difference in conversion between A and B can be read causally. In your forecasting model, a candidate regressor that grows over time (installed base, marketing budget) can correlate strongly with growing demand just because both trend. Before trusting it, look at the correlation of its changes with demand's changes, or its correlation with the residuals after trend and seasonality, and check that its values are actually known at forecast time (leakage, Chapter 7.12).

"X and Y are strongly correlated, so X drives Y."

A correlation is compatible with X → Y, Y → X, a common cause, or chance helped by shared trends. Correlation describes association in the observed data; causation is a claim about what happens when you intervene.

Model answer: "Correlation alone cannot separate causation from confounding or reverse causation. To claim causation I would randomize, as in an A/B test, or, with observational data, adjust for the confounders I can measure and state the remaining assumptions explicitly."

Correlation = moving together in observed data. Causation = changing X would change Y (intervention).

Same r from: X→Y, Y→X, common cause Z (confounder), chance / shared trends.

Randomization (A/B tests) breaks confounding. Trap: trending series correlate by accident; correlate the changes.

Quick check: cities with more ice-cream shops have more crime. Name a likely confounder.

Population size (and also warm climate). Bigger cities have more of everything: more shops and more crime. Compare cities of similar size (or use per-person rates) and the link largely disappears.

Spearman's rank correlation: Pearson on the ranks core

Sometimes we do not care whether the relationship is a straight line. We only ask: when x goes up, does y (almost) always go up too? Think of advertising: the first extra emails bring many orders, later ones bring fewer, but more emails never mean fewer orders. The curve bends, yet it only ever rises. That is called a monotone relationship (always rising, or always falling).

Spearman's trick is simple: throw away the actual values and keep only their ranks (1st smallest, 2nd smallest, …). A curve that only rises turns into a perfect straight line of ranks. Then compute ordinary Pearson on the ranks.

Three ways to say it:

  • Picture: replace each axis by "place in the queue"; any rising curve becomes a straight rising staircase.
  • Numbers: for $y = x^3$ with $x = 1, \dots, 5$, Pearson is 0.943 but Spearman is exactly 1.
  • Slogan: Pearson uses values, Spearman uses ranks: Pearson asks "straight line?", Spearman asks "always the same direction?".

A rising curve. $x = 1, 2, 3, 4, 5$ and $y = x^3 = 1, 8, 27, 64, 125$.

  1. Pearson on the values: the points bend upward, so they are not on a straight line: $r \approx 0.943$.
  2. Ranks of $x$: $1, 2, 3, 4, 5$. Ranks of $y$: 1 is the smallest (rank 1), 8 is rank 2, …, 125 is rank 5: $1, 2, 3, 4, 5$.
  3. The rank pairs $(1,1), (2,2), \dots, (5,5)$ lie on a straight line, so Pearson on the ranks is exactly 1. Spearman $\rho_s = 1$.

The shortcut formula. Five product categories ranked by page views and by orders:

  1. Ranks by views: $1, 2, 3, 4, 5$. Ranks by orders: $2, 1, 3, 5, 4$.
  2. Differences $d_i$: $-1, 1, 0, -1, 1$. Squares: $1, 1, 0, 1, 1$, so $\sum d_i^2 = 4$.
  3. $\rho_s = 1 - \dfrac{6\sum d_i^2}{n(n^2 - 1)} = 1 - \dfrac{6 \times 4}{5 \times 24} = 1 - \dfrac{24}{120} = 0.8$.

Ties get the average of the ranks they share: the values $10, 20, 20, 30$ get ranks $1, 2.5, 2.5, 4$.

Spearman's rank correlation $\rho_s$ is Pearson's correlation computed on the ranks:

$$\rho_s = r\big(\text{rank}(x),\; \text{rank}(y)\big).$$

When there are no ties this equals $\rho_s = 1 - \dfrac{6\sum_{i} d_i^2}{n(n^2-1)}$, where $d_i$ is the difference between the two ranks of observation $i$. With ties, use average ranks and the Pearson-on-ranks form (the shortcut is then only approximate).

  • $\rho_s = +1$: perfectly increasing (every larger $x$ has a larger $y$); $-1$: perfectly decreasing. It measures monotone association, straight or curved.
  • Unchanged by any increasing transformation of either variable ($\log$, $\sqrt{\ }$, $e^x$), because those keep the order.
  • Works for ordinal data: values that have an order but no meaningful distances (satisfaction 1–5, plan tier free < basic < pro).
  • Limits the damage of outliers: a point can be at most "the largest" or "the smallest"; pushing it further away does not change its rank.
Why do we need it?

Many real relationships bend (diminishing returns, saturation, exponential growth) or come with outliers and skewed values. Pearson underrates bent monotone links and overreacts to outliers. Ranks fix both, and they also handle ordinal scales.

Where is it used?

Screening regressors whose effect saturates, skewed business metrics (revenue, session length), survey and rating data, rank agreement between two models' scores, and the method='spearman' option of pandas' df.corr.

How is it used?

scipy.stats.spearmanr(x, y) or df.corr(method='spearman'). Compare it with Pearson: Spearman much larger than |Pearson| suggests a bent monotone curve or an outlier pulling Pearson down; Pearson much larger than Spearman suggests a few extreme points are making the straight line.

values: y = x³ Pearson 0.943 replace by ranks ranks: 1–5 against 1–5 Spearman = 1
The curve $y = x^3$ is perfectly increasing but not straight, so Pearson is below 1. Replacing every value by its rank straightens any increasing curve, so Spearman (Pearson on the ranks) is exactly 1.

Left: 25 points on a rising curve $y = 10\,(x/10)^k$ plus noise. Right: the same points with every value replaced by its rank. (1) Set the noise to 0 and raise the curvature $k$: the left cloud bends and Pearson falls, but the right plot stays a perfect straight line, so Spearman and Kendall stay at exactly 1. (2) Add noise: now all three fall below 1, and Spearman stays above Pearson.

Twelve grey-blue points have a clear rising trend (Pearson 0.82). The red point is one bad record (a data-entry error, a refund day). Drag it. (1) Put it inside the cloud: little changes. (2) Drag it to the bottom-right corner: Pearson falls towards 0 or below. (3) Once it is beyond all other points, keep dragging it further: Pearson keeps changing, but Spearman and Kendall do not move at all, because the point's rank no longer changes.

"Spearman is immune to outliers."

It is more resistant, not immune. In a small sample one extreme point still moves it (from 0.81 to 0.42 in the widget), but its influence is capped: once the point is the most extreme, moving it further changes nothing. Pearson's damage has no cap.

"Spearman = 1 means y is a straight-line function of x."

It means $y$ always increases with $x$: any increasing curve, straight or bent ($e^x$, $\log x$, $x^3$).

"A log transform changes the Spearman correlation."

Ranks do not change under any increasing transformation, so Spearman stays exactly the same. Pearson does change (often a log makes a bent link straighter and raises Pearson, see Chapter 4.18).

"Pearson and Spearman are two formulas for the same thing; just pick one."

They measure different things. Pearson measures linear association on the original values; Spearman measures monotone association on the ranks.

Model answer: "Pearson uses the values and asks how close the points are to a straight line; Spearman uses ranks and asks whether y consistently increases or decreases with x. I prefer Spearman for monotone non-linear relationships, ordinal data, or when outliers worry me; if the two disagree a lot, I plot the data to see whether it is curvature or a few extreme points."

$\rho_s$ = Pearson on ranks; no ties: $1 - \dfrac{6\sum d_i^2}{n(n^2-1)}$; ties get average ranks.

Measures monotone association; = ±1 for any perfectly increasing/decreasing curve; unchanged by log or any increasing transform.

Use for curved-monotone links, ordinal data, outliers (resistant, not immune).

Quick check: x = 1, 2, 3, 4 and y = 10, 100, 1000, 10000. What are Spearman and (roughly) Pearson?

The ranks are $1, 2, 3, 4$ for both, so Spearman = 1. The values explode upward, so the points are far from a straight line and Pearson is clearly below 1 (about 0.82). After a $\log_{10}$ of $y$ the values become $1, 2, 3, 4$ and Pearson would also be 1.

Kendall's tau: let every pair of points vote

Take any two days. If the warmer day also had more orders, the pair agrees about the direction (it is concordant). If the warmer day had fewer orders, the pair disagrees (it is discordant). Do this for every possible pair and count. Kendall's tau is the share of agreeing pairs minus the share of disagreeing pairs.

Three ways to say it:

  • Picture: join every pair of points with a segment. Rising segments are "agree" votes, falling ones are "disagree" votes.
  • Numbers: 5 of 6 pairs agree and 1 disagrees: $\tau = (5 - 1)/6 = 0.67$.
  • Slogan: τ = P(a random pair agrees) − P(it disagrees).

Four points: $(1, 1), (2, 3), (3, 2), (4, 4)$. There are $\binom{4}{2} = 6$ pairs.

  1. $(1,1)$ & $(2,3)$: x up, y up → agree (C).
  2. $(1,1)$ & $(3,2)$: x up, y up → C.   $(1,1)$ & $(4,4)$: C.
  3. $(2,3)$ & $(3,2)$: x up (2 → 3), y down (3 → 2) → disagree (D).
  4. $(2,3)$ & $(4,4)$: C.   $(3,2)$ & $(4,4)$: C.
  5. Totals: $C = 5$, $D = 1$. $\tau = \dfrac{C - D}{\text{number of pairs}} = \dfrac{5 - 1}{6} \approx 0.667$.
  6. For comparison, Spearman for the same points is 0.8 (rank differences $0, 1, -1, 0$; $1 - 6 \times 2/(4 \times 15) = 0.8$), and Pearson is also 0.8. Kendall is usually the smallest of the three in size.

The five temperature days give $C = 8$, $D = 2$ out of 10 pairs, so $\tau = 0.6$ (the two disagreeing pairs are days 2–3 and days 4–5).

For $n$ paired observations there are $n(n-1)/2$ pairs. A pair $i, j$ is concordant if $(x_j - x_i)(y_j - y_i) \gt 0$ and discordant if it is $\lt 0$. With $C$ concordant and $D$ discordant pairs and no ties,

$$\tau = \frac{C - D}{n(n-1)/2}.$$
  • With ties, the common version τ-b (SciPy's default) divides by $\sqrt{(n_0 - n_1)(n_0 - n_2)}$, where $n_0 = n(n-1)/2$ and $n_1$, $n_2$ are the numbers of pairs tied in $x$ and in $y$.
  • Range $[-1, 1]$; $\pm1$ for perfectly monotone data, like Spearman. Measures monotone association.
  • Usually smaller in size than Spearman. For jointly Normal data with correlation $\rho$: $\tau = \tfrac{2}{\pi}\arcsin\rho$, so $\rho = 0.8$ gives $\tau \approx 0.59$ (and Spearman $\approx 0.79$).
  • Simple meaning (no ties): $\tau = P(\text{concordant}) - P(\text{discordant})$ for a randomly chosen pair.
Why do we need it?

It has the clearest meaning of all the rank measures ("how much more likely is a pair to agree than disagree?"), it behaves well in small samples and with many ties, and it is robust to outliers.

Where is it used?

Comparing two rankings (a model's ranking of products against the true ranking, two judges, two time periods), small-sample studies, data with many ties (ratings on a 1–5 scale), and copula models in finance.

How is it used?

scipy.stats.kendalltau(x, y) (τ-b by default) or df.corr(method='kendall'). Do not compare its size directly with Pearson or Spearman: τ = 0.6 is a strong association.

Six draggable points A–F make 15 pairs. Every pair is joined by a segment: green = agree (concordant: rising), red = disagree (discordant: falling), grey dashed = tied. Press Next pair to walk through the pairs one at a time. Then drag a point down past its neighbours and watch green segments turn red and τ drop.

"τ = 0.6 is weaker than Pearson r = 0.8, so the link is weaker."

Kendall's scale is different: for Normal data with $\rho = 0.8$, τ is only about 0.59. Compare τ with other τ values, not with $r$.

"Kendall and Spearman will disagree about the direction."

They almost always have the same sign; both measure monotone association. They differ mainly in scale and in how they weigh big rank swaps (Spearman squares rank differences, Kendall just counts swapped pairs).

Pair is concordant if $(x_j-x_i)(y_j-y_i) \gt 0$, discordant if $\lt 0$. $\tau = \dfrac{C-D}{n(n-1)/2}$ (τ-b adjusts for ties).

Meaning: P(pair agrees) − P(pair disagrees). Monotone, robust, small-sample friendly.

Trap: smaller scale than Spearman/Pearson (Normal: $\tau = \tfrac{2}{\pi}\arcsin\rho$).

Quick check: five points are perfectly decreasing. What are C, D and τ?

All $5 \times 4/2 = 10$ pairs disagree: $C = 0$, $D = 10$, so $\tau = (0 - 10)/10 = -1$.

Partial correlation: what is left after removing a third variable?

Ice-cream orders and sunscreen orders are correlated, and we suspect temperature is behind it. Partial correlation asks a precise question: after we take out everything temperature explains, do the two still move together?

The recipe is a "leftovers" recipe. Predict ice-cream orders from temperature with a straight line and keep what the line could not predict (the residuals: actual minus predicted). Do the same for sunscreen orders. Then correlate the two sets of leftovers. If the leftovers are uncorrelated, temperature explained the whole link.

Three ways to say it:

  • Picture: subtract the temperature trend from both variables, then look at the scatter of what remains.
  • Numbers: $r_{XY} = 0.6$ can drop to a partial correlation of 0 once $Z$ is removed (example below).
  • Slogan: partial correlation = correlation of the leftovers after the third variable has had its say.

For one control variable $Z$ there is a formula that uses only the three ordinary correlations:

$$r_{XY\cdot Z} = \frac{r_{XY} - r_{XZ}\,r_{YZ}}{\sqrt{(1 - r_{XZ}^2)(1 - r_{YZ}^2)}}.$$

Case 1: the link is all temperature. Ice cream–sunscreen $r_{XY} = 0.6$; ice cream–temperature $r_{XZ} = 0.8$; sunscreen–temperature $r_{YZ} = 0.75$.

  1. The part of the link that travels through $Z$: $r_{XZ}\,r_{YZ} = 0.8 \times 0.75 = 0.6$.
  2. Numerator: $0.6 - 0.6 = 0$. So $r_{XY\cdot Z} = 0$: nothing is left after removing temperature.

Case 2: a real link remains. Marketing emails sent ($X$) and orders ($Y$) have $r_{XY} = 0.7$. Both are higher in the holiday season ($Z$): $r_{XZ} = 0.5$, $r_{YZ} = 0.6$.

  1. Numerator: $0.7 - 0.5 \times 0.6 = 0.7 - 0.3 = 0.4$.
  2. Denominator: $\sqrt{(1 - 0.25)(1 - 0.36)} = \sqrt{0.75 \times 0.64} = \sqrt{0.48} \approx 0.693$.
  3. $r_{XY\cdot Z} = 0.4 / 0.693 \approx 0.577$. The season explains part of the link (0.7 shrinks to 0.58), but a sizeable association remains.

The partial correlation of $X$ and $Y$ given $Z$, written $r_{XY\cdot Z}$, is the correlation between

  • $e_X$ = the residuals of $X$ after a least-squares straight-line fit on $Z$, and
  • $e_Y$ = the residuals of $Y$ after a least-squares straight-line fit on $Z$.

This is the residual method. With several control variables $Z_1, \dots, Z_k$, fit $X$ and $Y$ on all of them (multiple regression, Chapter 5.13) and correlate the residuals. For one $Z$ it equals the formula above.

  • It removes only the straight-line part of $Z$'s influence. A curved effect of $Z$ leaves traces in the residuals.
  • It is a statement about association after adjustment, not automatically about causation. Adjusting for the wrong variable (one that is caused by both $X$ and $Y$) can even create a fake correlation (Chapter 5.12).
Why do we need it?

Raw correlations mix direct links with links that travel through other variables. To ask "does this regressor carry information beyond the season?" we must take the season out first.

Where is it used?

Checking whether a candidate regressor adds anything beyond trend and seasonality, the partial autocorrelation function PACF of time series (Chapter 7.3), Gaussian graphical models, and adjustment for confounders in observational analysis.

How is it used?

Fit $X \sim Z$ and $Y \sim Z$ with np.polyfit, statsmodels.OLS or sklearn.LinearRegression, take residuals, then np.corrcoef(eX, eY). (The pingouin package also has partial_corr.)

X (ice cream) Y (sunscreen) Z (temperature) fit X on Z,keep leftovers e_X fit Y on Z,keep leftovers e_Y corr(e_X, e_Y)= r_XY·Z
The residual method. Remove from $X$ and from $Y$ everything a straight line in $Z$ can explain, then correlate what is left. That correlation is the partial correlation.

A hidden season variable $Z$ pushes both $X$ and $Y$ (dot colour = $Z$, teal low → pink high). Left: the raw scatter of $X$ and $Y$. Right: the leftovers after removing a straight line in $Z$ from each. (1) Keep the direct effect at 0: the raw correlation is large, the right-hand cloud is round and the partial correlation is near 0. (2) Raise the direct effect: now a link survives on the right. (3) Set the season's pull to 0: raw and partial correlation become the same.

"Partial correlation near 0 proves there is no causal effect."

It only says that, after removing the straight-line effect of the variables you controlled for, no straight-line association is left. Unmeasured confounders, curved effects and controlling for the wrong variable can all mislead.

"Control for every variable you have, just in case."

Controlling for a variable that is itself caused by both $X$ and $Y$ (a "collider", Chapter 5.12) can create a correlation that is not there. Choose controls by reasoning about causes, not by habit.

In your forecasting model, trend $g(t)$ and seasonality $s(t)$ play the role of $Z$. A candidate regressor may correlate with demand only because both have the same seasonal shape. The useful question is whether it helps after trend and seasonality: correlate the regressor's residuals (after removing trend and season) with the model's residuals. A strong raw correlation with a near-zero partial correlation means the regressor mostly duplicates what the Fourier terms already capture.

$r_{XY\cdot Z}$ = corr of residuals of $X$ on $Z$ and of $Y$ on $Z$ (residual method).

One control: $r_{XY\cdot Z} = \dfrac{r_{XY} - r_{XZ}r_{YZ}}{\sqrt{(1-r_{XZ}^2)(1-r_{YZ}^2)}}$.

Removes only the linear effect of $Z$; adjustment ≠ causation; never control for a collider.

Quick check: $r_{XY} = 0.5$, $r_{XZ} = 0$, $r_{YZ} = 0.6$. Is the partial correlation smaller or larger than 0.5?

Numerator $0.5 - 0 \times 0.6 = 0.5$; denominator $\sqrt{(1 - 0)(1 - 0.36)} = \sqrt{0.64} = 0.8$. So $r_{XY\cdot Z} = 0.5/0.8 = 0.625$: larger than 0.5. $Z$ has nothing to do with $X$, but it adds noise to $Y$; removing that noise makes the X–Y link clearer. (Careful: the three numbers must fit together. With $r_{YZ} = 0.9$ instead, the "partial correlation" would come out as 1.15, which is impossible: no real data has those three correlations, because that correlation matrix would have a negative eigenvalue.)

The correlation matrix: every pair at once core

Before you add ten candidate regressors to a forecasting model, you want to know which ones are near-duplicates of each other and which ones move with demand. With 10 variables there are 45 pairs. Instead of 45 scatter plots, put all the correlations in one table: row $i$, column $j$ holds the correlation of variable $i$ with variable $j$. Colour the cells and you have a heatmap you can read in seconds.

Three ways to say it:

  • Picture: a coloured square grid; strong colours off the diagonal are pairs that move together.
  • Numbers: the diagonal is all 1 (every variable with itself) and the table is a mirror image across the diagonal.
  • Slogan: one look at the matrix before you build the model saves hours of puzzling over unstable coefficients.

A simulated year of daily data (the widget below uses the same data). Three of the variables:

$$R = \begin{bmatrix} 1 & 0.98 & -0.48 \\ 0.98 & 1 & -0.46 \\ -0.48 & -0.46 & 1 \end{bmatrix} \quad \begin{matrix} \text{temperature} \\ \text{feels-like} \\ \text{rain} \end{matrix}$$
  1. Diagonal: 1, 1, 1. Every variable is perfectly correlated with itself.
  2. Symmetry: the (1, 2) entry and the (2, 1) entry are both 0.98, because $r_{XY} = r_{YX}$.
  3. Temperature and feels-like: 0.98 (0.984 before rounding). They carry almost the same information: a near-duplicate pair.
  4. Rain: about −0.48 with both: rainy days tend to be cooler. A moderate negative link, not a duplicate.
  5. Warning size for a model containing both temperatures: the variance inflation factor $VIF = 1/(1 - r^2) = 1/(1 - 0.984^2) \approx 31$. It says the variance of each temperature coefficient is about 31 times larger than it would be if the two were uncorrelated (VIF is taught with regression in Chapter 5.13). Note how touchy this is: with $r = 0.98$ it would be 25.

For variables $X_1, \dots, X_p$, the correlation matrix $R$ is the $p \times p$ matrix with entries $R_{ij} = \text{corr}(X_i, X_j)$.

  • Ones on the diagonal; symmetric ($R_{ij} = R_{ji}$); every entry in $[-1, 1]$.
  • It is the covariance matrix of the standardized variables: $R = D^{-1/2}\,\Sigma\,D^{-1/2}$, where $\Sigma$ is the covariance matrix and $D$ the diagonal matrix of variances (the covariance matrix itself is taught in Chapter 5.15).
  • It is positive semi-definite (no negative eigenvalues; see Chapter 1.12). So not every symmetric table of numbers in $[-1,1]$ is a valid correlation matrix: $r_{12} = 0.9$, $r_{13} = 0.9$, $r_{23} = -0.9$ is impossible (one eigenvalue is −0.8). If X is close to Y and to Z, Y and Z cannot be strongly opposite.
  • It can be Pearson, Spearman or Kendall: same layout, different measure in each cell.
Why do we need it?

Regressors that move together compete for the same credit: the model cannot tell which one deserves it, so their coefficients become unstable and hard to interpret (multicollinearity). The matrix finds such pairs before they cause trouble.

Where is it used?

Exploratory data analysis, choosing external regressors for a forecasting model, feature selection, PCA (an eigen-decomposition of $R$, Chapter 5.16), and posterior correlation matrices that show which parameters a variational guide must capture together.

How is it used?

df.corr() (or method='spearman'), then a heatmap (for example seaborn.heatmap(df.corr(), vmin=-1, vmax=1, cmap='RdBu')). Look for off-diagonal cells with large |r| between regressors, and check each regressor's correlation with the target.

tempfeelsrain tempfeelsrain 1 0.98 −0.48 0.98 1 −0.46 −0.48 −0.46 1 near-duplicate pair (|r| ≥ 0.9):expect unstable coefficients mirror image across the diagonal:R is symmetric blue = positive, orange = negative,stronger colour = larger |r|
Reading a correlation heatmap: the diagonal is always 1, the two triangles mirror each other, and the off-diagonal cells with strong colour flag pairs that move together.

A simulated year of daily data: six candidate regressors and the target (orders). Cells with |r| at or above the threshold between two regressors get a red frame (cells with the target are signal, not a problem). (1) Find the two near-duplicate pairs. (2) Switch to Spearman: promo day and discount % become almost 1, because discount is only positive on promo days. (3) Use the menu to look at the scatter behind any cell (both axes standardized).

"No pair has |r| above 0.8, so there is no multicollinearity."

Multicollinearity can involve three or more variables at once (for example $x_3 \approx x_1 + x_2$) without any single large pairwise correlation. The VIF from a regression of each regressor on all the others catches this; the pairwise matrix does not.

"Keep the regressors with high correlation to the target and drop the rest."

A correlation with the target is a marginal (one-at-a-time) view. A regressor with a small raw correlation can matter once seasonality is removed, and a large raw correlation can be just shared trend. Judge regressors after adjustment (partial correlation, holdout performance).

"Compute the matrix on all the data I have."

For forecasting, choose regressors using the training window only. Screening on data that includes the test period leaks future information into your model choice (Chapter 7.12).

Before adding many external regressors to your forecasting model, compute the correlation matrix of the candidates (training window only, after the same scaling you use in the model; see Chapter 4.18). Near-duplicates such as temperature and feels-like temperature split the effect between their two $\beta$s arbitrarily. In a Bayesian fit this shows up as a long, thin ridge in the posterior of $(\beta_1, \beta_2)$: the sum $\beta_1 + \beta_2$ is well determined but each one alone is not, so the two coefficients are strongly negatively correlated in the posterior. A mean-field guide cannot represent that correlation and understates each coefficient's uncertainty; a full-rank or low-rank guide can (Chapter 6.13). Usually it is simpler to keep one of the pair or combine them.

"Multicollinearity makes the model's predictions bad."

It mainly makes the individual coefficients unstable and hard to interpret (large standard errors, signs that flip between samples). Predictions can stay fine as long as new data keeps the same pattern between the correlated regressors; they suffer when that pattern breaks.

Model answer: "I check the correlation matrix and VIFs before adding regressors. Highly correlated regressors inflate the variance of their coefficients because the data cannot separate their effects. If I care about interpretation I drop or combine them, or use a shrinkage prior; if I only care about prediction, it matters less, provided the correlation structure is stable."

$R_{ij} = \text{corr}(X_i, X_j)$: ones on the diagonal, symmetric, positive semi-definite; $R = D^{-1/2}\Sigma D^{-1/2}$.

Large off-diagonal |r| between regressors → multicollinearity: unstable coefficients; $VIF = 1/(1-R_j^2)$ (two regressors: $1/(1-r^2)$).

Traps: pairwise view misses 3-way collinearity; screen on the training window only.

Quick check: in a correlation matrix of 12 regressors, how many distinct off-diagonal correlations are there?

$12 \times 11 / 2 = 66$. The matrix has $144$ cells: 12 on the diagonal (all 1) and 132 off the diagonal, which come in mirror pairs, so 66 distinct values.

Recap, cheat sheet and practice

  • Covariance $Cov(X,Y) = E[(X-\mu_X)(Y-\mu_Y)]$ says whether two variables move together (sign), but its size depends on units: $Cov(aX+b, cY+d) = ac\,Cov(X,Y)$. It completes the variance of a sum: $Var(X \pm Y) = Var X + Var Y \pm 2Cov(X,Y)$.
  • Correlation $\rho = Cov/(\sigma_X\sigma_Y)$ is the covariance of z-scores: unit-free, always in $[-1, 1]$, $\pm1$ only for an exact straight line.
  • Pearson's r measures straight-line association. It is not a slope, not a percentage, and is sensitive to outliers. Always plot (Anscombe).
  • Zero correlation ≠ independence (U, ring, $Y = X^2$), except for jointly Normal variables. Correlation ≠ causation: confounders, reverse causation and shared trends all create correlation; randomization breaks confounding.
  • Spearman = Pearson on ranks: monotone association, ordinal data, resistant to outliers. Kendall = (agreeing − disagreeing pairs)/pairs: same idea, smaller scale.
  • Partial correlation = correlation of residuals after removing a third variable. The correlation matrix shows every pair; near-duplicate regressors mean unstable coefficients.

Cheat sheet

MeasureFormulaMeasuresPython
Covariance$\frac{1}{n-1}\sum (x_i-\bar x)(y_i-\bar y)$direction; size has unitsnp.cov(x, y)[0, 1]
Pearson r$\dfrac{s_{xy}}{s_x s_y}$straight-line associationstats.pearsonr, df.corr()
Spearman $\rho_s$Pearson on ranks; $1-\frac{6\sum d_i^2}{n(n^2-1)}$monotone associationstats.spearmanr
Kendall τ$\dfrac{C-D}{n(n-1)/2}$ (τ-b with ties)monotone; pair agreementstats.kendalltau
Partial $r_{XY\cdot Z}$$\dfrac{r_{XY}-r_{XZ}r_{YZ}}{\sqrt{(1-r_{XZ}^2)(1-r_{YZ}^2)}}$link left after removing Zresiduals + np.corrcoef
Correlation matrix$R_{ij} = \text{corr}(X_i, X_j)$all pairs; multicollinearitydf.corr()
Variance of a sum$Var X + Var Y + 2Cov(X,Y)$combining variables—
Code it · Python
import numpy as np
import pandas as pd
from scipy import stats

temp   = np.array([14, 16, 18, 20, 22])        # °C
orders = np.array([110, 130, 120, 150, 140])   # orders per day

# --- covariance: np.cov returns the 2x2 matrix and divides by n-1 ---
C = np.cov(temp, orders)
print(C)                      # [[ 10.  40.]  [ 40. 250.]]
print(C[0, 1])                # 40.0   (°C·orders)
print(np.var(temp), np.var(temp, ddof=1))   # 8.0 10.0  <- np.var divides by n unless ddof=1

# --- correlation is unit-free ---
temp_f = 1.8 * temp + 32
print(round(np.cov(temp_f, orders)[0, 1], 6))          # 72.0  covariance changed (x 1.8)
print(round(np.corrcoef(temp, orders)[0, 1], 6),
      round(np.corrcoef(temp_f, orders)[0, 1], 6))     # 0.8 0.8  correlation did not

# --- Pearson, Spearman, Kendall ---
print(round(stats.pearsonr(temp, orders).statistic, 6))    # 0.8
print(round(stats.spearmanr(temp, orders).statistic, 6))   # 0.8
print(round(stats.kendalltau(temp, orders).statistic, 6))  # 0.6  (8 concordant, 2 discordant of 10 pairs)

x = np.arange(1, 6); y = x**3                               # monotone but curved
print(round(stats.pearsonr(x, y).statistic, 3),
      round(stats.spearmanr(x, y).statistic, 6))            # 0.943 1.0

# --- zero correlation, total dependence ---
x = np.linspace(-1, 1, 201)
print(round(np.corrcoef(x, x**2)[0, 1], 10))                # 0.0 (up to rounding), yet y = x^2

# --- partial correlation by the residual method ---
rng = np.random.default_rng(0)
z  = rng.normal(size=1000)                       # season / temperature
xs = 0.8 * z + 0.6 * rng.normal(size=1000)       # ice-cream orders
ys = 0.8 * z + 0.6 * rng.normal(size=1000)       # sunscreen orders (no direct link)

def resid(a, b):
    """residuals of a after a straight-line least-squares fit on b"""
    slope, intercept = np.polyfit(b, a, 1)
    return a - (intercept + slope * b)

print(round(np.corrcoef(xs, ys)[0, 1], 2))                       # 0.63  raw correlation
print(round(np.corrcoef(resid(xs, z), resid(ys, z))[0, 1], 2))   # 0.01  partial: about 0

# --- correlation matrix with pandas ---
df = pd.DataFrame({"temp": z, "ice_cream": xs, "sunscreen": ys})
print(df.corr().round(2))                        # Pearson (default): temp-ice 0.80, temp-sun 0.79, ice-sun 0.63
print(df.corr(method="spearman").round(2))       # Spearman; method="kendall" also works
Test yourself

1. Temperature (°C) and orders have covariance 40 and correlation 0.8. You convert temperature to °F ($1.8x + 32$). What are the new covariance and correlation?

Shifting by 32 changes nothing; multiplying by 1.8 multiplies the covariance: $1.8 \times 40 = 72$. Correlation is unit-free, so it stays 0.8 (it can never exceed 1).

2. Which statement is true?

Independence implies zero covariance, but not the other way round ($Y = X^2$). A high r does not show causation, and the change of Y per unit X is the slope $b = r\,s_y/s_x$, not r.

3. $y = e^x$ for $x = 1, 2, \dots, 10$, with no noise. What do Pearson and Spearman give?

$e^x$ always increases, so the ranks match perfectly: Spearman = 1. The curve explodes upward and is far from straight, so Pearson is below 1 (about 0.72 here).

4. Four points form 6 pairs: 4 pairs are concordant and 2 are discordant (no ties). Kendall's τ is…

$\tau = (C - D)/\text{pairs} = (4 - 2)/6 = 1/3$: agreeing pairs outnumber disagreeing ones by one third of all pairs.

5. $r_{XY} = 0.6$, $r_{XZ} = 0.8$, $r_{YZ} = 0.75$. What is the partial correlation $r_{XY\cdot Z}$?

Numerator $0.6 - 0.8 \times 0.75 = 0.6 - 0.6 = 0$. Once the straight-line effect of Z is removed from both, nothing is left.

6. Two candidate regressors for your forecasting model have $r = 0.97$. What is the main risk of putting both in?

Near-duplicate regressors cause multicollinearity: the sum of their effects is well determined, each one alone is not (VIF ≈ $1/(1-0.97^2) \approx 17$). In a Bayesian fit their coefficients become strongly negatively correlated in the posterior.

Practice problems

A. For $x = 1, 2, 3, 4$ and $y = 2, 4, 5, 9$, compute the covariance, Pearson's r, Spearman and Kendall.

$\bar x = 2.5$, $\bar y = 5$. Deviations: $x$: $-1.5, -0.5, 0.5, 1.5$; $y$: $-3, -1, 0, 4$. Products: $4.5, 0.5, 0, 6$, sum $11$. So $s_{xy} = 11/3 \approx 3.67$.

$\sum (x_i-\bar x)^2 = 2.25 + 0.25 + 0.25 + 2.25 = 5$; $\sum (y_i-\bar y)^2 = 9 + 1 + 0 + 16 = 26$. $r = 11/\sqrt{5 \times 26} = 11/\sqrt{130} \approx 11/11.40 \approx 0.965$.

$y$ increases every time $x$ does: ranks are $1,2,3,4$ for both, so Spearman = 1, and all 6 pairs are concordant, so Kendall = 1. Pearson is below 1 because the jump from 5 to 9 bends the pattern.

B. Two metrics measured on the same users have $Var(X) = 4$, $Var(Y) = 9$ and correlation $0.5$. Find $Var(X - Y)$. What would it be if you wrongly assumed independence?

$Cov(X, Y) = \rho\,\sigma_X\sigma_Y = 0.5 \times 2 \times 3 = 3$. Then $Var(X - Y) = 4 + 9 - 2 \times 3 = 7$. Assuming independence gives $4 + 9 = 13$, almost twice as large: the positive correlation makes the difference less noisy. This is exactly why paired designs and CUPED reduce variance.

C. Show that replacing $x$ by $3x + 5$ does not change Pearson's r, but replacing it by $-3x + 5$ flips its sign.

Let $x' = ax + b$. Then $x'_i - \bar x' = a(x_i - \bar x)$. The numerator of $r$ becomes $a\sum (x_i - \bar x)(y_i - \bar y)$ and the denominator becomes $\sqrt{a^2\sum (x_i-\bar x)^2}\sqrt{\sum (y_i - \bar y)^2} = |a|\sqrt{\cdots}\sqrt{\cdots}$. So $r' = (a/|a|)\,r$: unchanged for $a = 3$, sign flipped for $a = -3$. The shift $b$ disappears when we subtract the mean.

D. Interview: "Explain why zero correlation does not imply independence. When does it?"

"Correlation only measures the linear part of a relationship: it is zero when $E[XY] = E[X]E[Y]$. Independence requires the entire distribution of Y to be the same for every value of X. Example: X uniform on $\{-1, 0, 1\}$ and $Y = X^2$ have covariance 0, but Y is a function of X. The implication holds in one important case: if X and Y are jointly Normal, zero correlation implies independence. Marginal Normality alone is not enough."

E. Two models rank the same 6 products. Model A: 1, 2, 3, 4, 5, 6. Model B: 2, 1, 4, 3, 6, 5. Compute Spearman and Kendall.

Rank differences $d$: $-1, 1, -1, 1, -1, 1$, so $\sum d^2 = 6$ and $\rho_s = 1 - \dfrac{6 \times 6}{6 \times 35} = 1 - \dfrac{36}{210} \approx 0.829$.

Pairs: $6 \times 5/2 = 15$. Only the three swapped neighbours (1–2, 3–4, 5–6) disagree: $D = 3$, $C = 12$. $\tau = (12 - 3)/15 = 0.6$.

F. Interview: a dashboard shows that days with more push notifications have more orders ($r = 0.7$). Your manager wants to double the notifications. What do you say?

"The correlation does not tell us that notifications cause orders. Notifications are probably sent more on promo days and holidays, when orders are high anyway (a confounder), and some notifications are triggered by user activity (reverse causation). Both series may also share a trend. To know the causal effect of more notifications, I would run an A/B test: randomly assign users to the current and the doubled notification rate and compare orders, with guardrail metrics like unsubscribes. Randomization makes the groups comparable on every confounder."

Chapter 4.16 · Syllabus Module 5

Histograms and kernel density estimation

You have a pile of numbers: 365 forecast residuals, 10 000 order values, 4 000 posterior draws. What shape do they have? One peak or two? Symmetric or skewed? Heavy tails? A histogram and its smooth cousin, the kernel density estimate (KDE), are the two standard pictures of a distribution. Both have a hidden knob that can change the story completely. This chapter is about seeing the shape honestly.

  • Build a histogram by hand, and know the difference between counts, relative frequencies and density
  • See how the bin width and the bin origin change the picture, and use the common rules of thumb
  • Build a KDE as "one bump per data point": $\hat f_h(x) = \frac{1}{nh}\sum K\big(\frac{x - x_i}{h}\big)$
  • Know that the bandwidth $h$ matters far more than the kernel: small $h$ = noisy (overfitting), large $h$ = over-smoothed (underfitting)
  • Use Silverman's rule and SciPy's default as rules of thumb, and know how the libraries differ
  • Spot the classic failures: boundary bias for positive data (and the reflection fix), hidden modes, and the curse of dimensionality
  • Use a KDE to look at your model's residual distribution against a Normal fit

Histograms: counting in bins core

Your forecasting model made 200 daily forecasts, and each day has a residual: actual orders minus forecast orders. A list of 200 numbers tells you nothing at a glance. So cut the number line into equal pieces called bins (say −10 to −5, −5 to 0, 0 to 5, …), drop every residual into its bin, and draw a bar for each bin as tall as the number of residuals inside. The outline of the bars is the shape of your data.

Three ways to say it:

  • Picture: buckets in a row on the number line; each value is a marble dropped in its bucket; the piles show where the data is crowded.
  • Numbers: ten residuals in four bins of width 5 give the counts 1, 3, 4, 2.
  • Slogan: a histogram is a bar chart of "how many values fell in each interval".

Ten residuals (orders): $-7, -4, -3, -1, 0, 1, 2, 2, 5, 9$. Bins of width $w = 5$: $[-10, -5)$, $[-5, 0)$, $[0, 5)$, $[5, 10]$. (The round bracket means "this end is not included", so 0 goes into $[0, 5)$.)

  1. Counts: $[-10,-5)$ holds $-7$ → 1. $[-5, 0)$ holds $-4, -3, -1$ → 3. $[0, 5)$ holds $0, 1, 2, 2$ → 4. $[5, 10]$ holds $5, 9$ → 2. Check: $1 + 3 + 4 + 2 = 10$.
  2. Relative frequencies (count ÷ $n$): $0.1, 0.3, 0.4, 0.2$. They add to 1.
  3. Density heights (count ÷ ($n \times w$)): $1/50 = 0.02$, $3/50 = 0.06$, $4/50 = 0.08$, $2/50 = 0.04$.
  4. Areas of the density bars (height × width): $0.1, 0.3, 0.4, 0.2$. Total area $= 1$, just like a probability density (Chapter 4.4).

Choose bin edges $b_0 \lt b_1 \lt \dots \lt b_k$. Bin $j$ is the interval $[b_{j-1}, b_j)$ with width $w_j = b_j - b_{j-1}$ (NumPy also closes the last bin on the right). Let $c_j$ be the number of the $n$ observations in bin $j$. A histogram draws a bar over each bin with one of three heights:

  • count $c_j$ (bar heights add to $n$);
  • relative frequency $c_j/n$ (heights add to 1);
  • density $\hat f_j = \dfrac{c_j}{n\,w_j}$ (bar areas add to 1).

The density version is an estimate of the probability density: for $x$ in bin $j$, $\hat f(x) = c_j/(n w_j)$. It is the only version you may overlay on a pdf, and the only one that stays honest when bins have different widths.

Why do we need it?

Numbers in a list hide their shape. A histogram shows at once where values are crowded, whether they are symmetric or skewed, whether there are two groups, and whether there are extreme values. It is the first plot to draw for any new variable.

Where is it used?

Exploratory analysis of every metric, residual checks of regression and forecasting models, posterior draws from MCMC or SVI, the sampling distribution in a bootstrap, PIT histograms for forecast calibration (Chapter 7.16).

How is it used?

np.histogram(x, bins=...) returns counts and edges; plt.hist(x, bins=..., density=True) draws density bars; df['x'].hist() in pandas. Use density=True whenever you overlay a curve or compare samples of different sizes.

1 3 4 2 −10 −5 0 5 10 each dot = one residual width w = 5 bar height = count (here 1, 3, 4, 2) density = count ÷ (n·w) = 4 ÷ (10 × 5) = 0.08 bin edges: −10, −5, 0, 5, 10 0 goes in [0, 5): left end included, right end not
A histogram of ten residuals with bins of width 5. Each bar counts the dots below it. Dividing each count by $n \times w$ turns the bars into a density whose total area is 1.

200 simulated forecast residuals. (1) Switch between count, relative frequency and density: the shape stays the same, only the y-axis changes. Only in density mode does the true curve (green, dashed) line up with the bars. (2) Move the bin width: counts grow with wider bins, density heights do not. (3) Turn on unequal bins (wide bins in the tails) and switch back to count: the wide tail bins look important. Density mode fixes this.

"The bar height of a histogram is always a count."

It may be a count, a relative frequency or a density. Read the y-axis. When you overlay a pdf or compare two samples of different sizes, use density (density=True).

"With unequal bins, the tallest bar is where the data is most crowded."

Only in density mode. A wide bin collects more values simply because it is wide. Dividing by the width fixes this.

"Density heights can never be above 1."

A density is not a probability. If the values live in $[0, 0.1]$, the heights must be around 10 so that the area is 1 (Chapter 4.4).

Histogram = counts per bin. Three heights: count $c_j$, relative frequency $c_j/n$, density $c_j/(n w_j)$ (areas sum to 1).

Use density to overlay a pdf, to compare samples of different sizes, and whenever bins have different widths.

NumPy bins are $[a, b)$ except the last, which is closed.

Quick check: 400 values, a bin of width 0.5 contains 60 of them. What is the density height of that bar?

$60 / (400 \times 0.5) = 60/200 = 0.3$. Its area is $0.3 \times 0.5 = 0.15$, which is the fraction of the data in the bin ($60/400$).

Bin width and bin origin: the histogram's two hidden choices core

A histogram looks objective, but you made two choices: how wide the bins are, and where the first bin starts (the bin origin). Very wide bins blur everything into one or two blocks. Very narrow bins give a comb of spikes, most of them noise. And sliding all the bins sideways by half a width can make a bump appear or disappear.

Three ways to say it:

  • Picture: a histogram is a photo of the data taken through a grid; move or resize the grid and the photo changes.
  • Numbers: the same ten residuals give counts 1, 3, 4, 2 with bins starting at −10, but 3, 5, 1, 1 with bins starting at −7.5.
  • Slogan: never trust a single histogram; try a few widths and starting points.

The ten residuals again: $-7, -4, -3, -1, 0, 1, 2, 2, 5, 9$. Keep the width 5, but start the bins at $-7.5$ instead of $-10$.

  1. Bins: $[-7.5, -2.5)$, $[-2.5, 2.5)$, $[2.5, 7.5)$, $[7.5, 12.5)$.
  2. Counts: $\{-7, -4, -3\}$ → 3; $\{-1, 0, 1, 2, 2\}$ → 5; $\{5\}$ → 1; $\{9\}$ → 1.
  3. Before: 1, 3, 4, 2 (a hump slightly right of 0). Now: 3, 5, 1, 1 (a peak at 0 and a long right tail). Same data, a different story.
  4. Rules of thumb for the width (all from this sample: $n = 10$, range $= 9 - (-7) = 16$, IQR $= 2 - (-2.5) = 4.5$): Sturges $= 16/(\log_2 10 + 1) = 16/4.32 \approx 3.70$; Freedman–Diaconis $= 2 \times 4.5 \times 10^{-1/3} \approx 4.18$. NumPy's bins='auto' takes the smaller one, 3.70, then rounds up to a whole number of bins: $\lceil 16/3.70 \rceil = 5$ bins of width $3.2$.

A histogram depends on the bin width $w$ and the bin origin (the position of the first edge). Common rules of thumb for $w$ (each is a convenient default, not a law):

  • Sturges: about $\log_2 n + 1$ bins, so $w = \text{range}/(\log_2 n + 1)$. Built for roughly Normal data; gives too few bins for large $n$.
  • Freedman–Diaconis: $w = 2\,\text{IQR}\; n^{-1/3}$. Uses the IQR, so outliers do not inflate it.
  • Scott: $w = 3.49\, s\; n^{-1/3}$ (assumes roughly Normal data).
  • NumPy bins='auto': the smaller of the Sturges and Freedman–Diaconis widths (recent versions also stop the FD width from becoming tiny). plt.hist uses 10 bins by default, whatever $n$ is.

Why $n^{-1/3}$? With more data each bin can be narrower and still hold enough points to give a stable height. Narrow bins = less blur (bias) but noisier heights (variance): the same trade-off as the KDE bandwidth below.

Why do we need it?

A bin width that is too large can hide a second group (weekend days); one that is too small turns noise into fake peaks. Knowing the two hidden choices stops you from reading stories into binning accidents.

Where is it used?

Every histogram you draw: residual plots, metric distributions in experiment reports, PIT histograms, posterior histograms. Also in data pipelines that bucket a continuous feature (age bands, price bands) for a model.

How is it used?

Start with bins='auto' or 'fd', then try half and double the width and shift the origin by half a bin. Report a feature (two peaks, a gap) only if it survives these changes, or confirm it with a KDE.

150 days of orders: most are weekdays (around 120) and some are weekends (around 165). (1) Make the bins very wide (40): the weekend group disappears into one block. (2) Make them very narrow (2): a comb of random spikes. (3) Pick a medium width like 14 and slide the origin: the two peaks merge and separate as the grid moves. (4) Press the rule buttons to see what Sturges, Freedman–Diaconis and NumPy's auto choose.

"plt.hist(x) picks a good number of bins for me."

Its default is 10 bins for 50 values or for 5 million. Pass bins='auto' (or 'fd') and still try a few widths.

"There is a bump in my histogram, so there is a real second group."

A bump can be a binning accident. Shift the origin by half a bin and halve the width; if the bump survives, and a KDE shows it too, take it seriously.

Two hidden choices: bin width $w$ and bin origin. Wide = blurred (bias), narrow = spiky (variance).

Rules of thumb: Sturges $\approx \log_2 n + 1$ bins; Freedman–Diaconis $w = 2\,\text{IQR}\,n^{-1/3}$; NumPy 'auto' = the smaller width.

Trap: plt.hist default = 10 bins. Check features with other widths/origins.

Quick check: you have 8 000 values with IQR 20. What bin width does Freedman–Diaconis suggest?

$w = 2 \times 20 \times 8000^{-1/3} = 40/20 = 2$, because $8000^{1/3} = 20$.

Kernel density estimation: a bump on every point core

The histogram's trouble comes from hard bin edges: a value either is in a bin or is not. A kinder idea: put a small, smooth hill of sand on top of every data point, each hill holding the same amount of sand ($1/n$ of the total). Where points are crowded, the hills pile up into a high ridge. Where points are rare, the ground stays low. The outline of all the sand is the kernel density estimate (KDE).

No bins, so no bin origin. The only real choice is how wide each hill is. That width is the bandwidth $h$.

Three ways to say it:

  • Picture: one small bell over every data point; add the bells; the sum is the curve.
  • Numbers: for the points 1, 2, 4 and $h = 1$, the curve at $x = 2$ is $\tfrac13(0.242 + 0.399 + 0.054) = 0.232$.
  • Slogan: a KDE is the average of $n$ little bumps, one per data point.

Data $x_1 = 1$, $x_2 = 2$, $x_3 = 4$; Gaussian bumps with bandwidth $h = 1$. The Gaussian bump is the standard Normal density $\varphi(u) = e^{-u^2/2}/\sqrt{2\pi}$, with $\varphi(0) = 0.399$, $\varphi(\pm1) = 0.242$, $\varphi(\pm2) = 0.054$.

  1. Height at $x = 2$. Distances in units of $h$: $(2-1)/1 = 1$, $(2-2)/1 = 0$, $(2-4)/1 = -2$.
  2. Bump heights: $\varphi(1) + \varphi(0) + \varphi(-2) = 0.242 + 0.399 + 0.054 = 0.695$.
  3. Divide by $n h = 3 \times 1$: $\hat f(2) = 0.695/3 \approx 0.232$.
  4. Height at $x = 3$: distances $2, 1, -1$: $(0.054 + 0.242 + 0.242)/3 = 0.538/3 \approx 0.179$. Lower: $x = 3$ sits in the gap between 2 and 4.
  5. Each bump has area $1/3$, so the total area is exactly 1: the KDE is a proper density.

Given data $x_1, \dots, x_n$, a kernel $K$ (a symmetric probability density, usually the standard Normal) and a bandwidth $h \gt 0$, the kernel density estimate is

$$\hat f_h(x) = \frac{1}{n h}\sum_{i=1}^{n} K\!\left(\frac{x - x_i}{h}\right).$$
  • $\frac{1}{h}K\big(\frac{x - x_i}{h}\big)$ is one bump: the kernel's shape, centred at $x_i$, stretched to width $h$ (for a Gaussian kernel, $h$ is the bump's standard deviation). The factor $\frac1n$ gives each bump area $\frac1n$.
  • $\hat f_h(x) \ge 0$ and $\int \hat f_h(x)\,dx = 1$. It is smooth when $K$ is smooth.
  • It is nonparametric: it does not assume a family (Normal, Gamma…); the data decide the shape. The price: it needs more data, and its look depends heavily on $h$.
  • A KDE is itself a mixture of $n$ small densities, so you can sample from it: pick a data point at random, then add kernel noise of width $h$.
Why do we need it?

Histograms are blocky and depend on bin edges. A KDE gives a smooth picture without edges, which makes it easier to compare several distributions on one plot and to see peaks, skew and tails.

Where is it used?

Density plots of residuals, posterior density plots drawn from MCMC or SVI samples (for example in ArviZ), violin plots, comparing a metric between variants, anomaly scores (low density = unusual), and smoothing in seaborn's kdeplot.

How is it used?

kde = scipy.stats.gaussian_kde(x), then kde(grid) evaluates the curve; seaborn.kdeplot(x) draws it; sklearn.neighbors.KernelDensity and statsmodels' KDEUnivariate are alternatives. Always check which bandwidth was used.

x = 1 x = 2 x = 4 f̂(2) = 0.232 dashed: one bump per point (area 1/3 each) solid: their sum = the KDE
The KDE of the points 1, 2 and 4 with a Gaussian kernel and $h = 1$. Each orange bump has area $1/3$; the black curve is their sum. The purple dot is the value computed in the example, $\hat f(2) \approx 0.232$.

Six data points sit on the axis (drag them left and right). Each one carries an orange bump; the blue curve is their sum divided by $n$. (1) Push two points together: their bumps pile up into a high peak. (2) Make $h$ small: every point gets its own spike. Make $h$ large: one wide hill. (3) Move the purple probe and read the arithmetic of $\hat f(x_0)$ in the readout. The area under the blue curve is always 1.

"The KDE shows the true distribution of my data."

It is an estimate that depends on $h$. It smooths sharp features, puts some density where there are no observations (beyond the largest point, or below 0), and can show bumps that are just noise when $h$ is small.

"The KDE curve at 0.23 means a 23% chance of that value."

It is a density. Probabilities are areas under the curve over an interval (Chapter 4.4).

When you plot the posterior of $\theta_B - \theta_A$ in your A/B framework, or of a changepoint slope $\delta_j$ in your forecasting model, you usually have a few thousand draws and a smooth curve drawn through them. That curve is a KDE of the draws. Its smoothness is a display choice: with a large bandwidth the posterior can look wider and more Normal than the draws really are. For decisions use the draws directly (for example the fraction with $\theta_B \gt \theta_A$), not the plotted curve.

$\hat f_h(x) = \dfrac{1}{nh}\sum_{i=1}^n K\!\left(\dfrac{x - x_i}{h}\right)$: one bump (area $1/n$) per point; total area 1.

$K$ = kernel (shape of the bump, usually Normal); $h$ = bandwidth (its width).

Trap: a KDE is an estimate shaped by $h$; heights are densities, not probabilities.

Quick check: one single data point at 5, Gaussian kernel, $h = 2$. What does the KDE look like?

A single Normal curve centred at 5 with standard deviation 2: $\hat f(x) = \frac{1}{2}\varphi\big(\frac{x-5}{2}\big)$, which is the $N(5, 2^2)$ density. Its peak height is $0.399/2 \approx 0.2$.

The kernel: the bump's shape matters little

The bump does not have to be a bell. It can be a dome (Epanechnikov), a box (top-hat) or a tent (triangular). Once the bumps have the same spread, the final curves look almost identical. The shape of each grain of sand hardly matters; how far the sand spreads does.

Three ways to say it:

  • Picture: bell, dome, box or tent of the same width give nearly the same landscape when you add many of them.
  • Numbers: the Gaussian kernel is about 95% as efficient as the theoretically best (Epanechnikov) kernel; the box is about 93%.
  • Slogan: choose the kernel for convenience; spend your attention on the bandwidth.

Matching the spread. A top-hat (box) kernel that is flat on $[-1, 1]$ has standard deviation $1/\sqrt3 \approx 0.577$. To make box bumps as wide as Gaussian bumps with $h = 1$ (standard deviation 1), the box must stretch over $[-\sqrt3, \sqrt3] \approx [-1.73, 1.73]$.

  1. Gaussian, $h = 1$: each bump has standard deviation $1 \times 1 = 1$.
  2. Box with "bandwidth" $h_{box}$ (half-width): standard deviation $h_{box}/\sqrt3$. Setting $h_{box}/\sqrt3 = 1$ gives $h_{box} = \sqrt3 \approx 1.73$.
  3. With $h_{box} = 1$ instead, the box bumps would be only $0.58$ as wide, and that KDE would look much rougher. The difference comes from the width, not the shape.

Efficiencies relative to the Epanechnikov kernel, for large samples: biweight 99.4%, triangular 98.6%, Gaussian 95.1%, box 93.0%. An efficiency of 95% means you need about $1/0.95 \approx 1.05$ times as much data to reach the same accuracy: a tiny price.

A kernel is a function $K(u) \ge 0$ with $\int K(u)\,du = 1$, usually symmetric ($K(-u) = K(u)$). Common choices, written on their standard scale:

  • Gaussian: $K(u) = \tfrac{1}{\sqrt{2\pi}}e^{-u^2/2}$ (never exactly zero; standard deviation 1).
  • Epanechnikov: $K(u) = \tfrac34(1 - u^2)$ for $|u| \le 1$ (minimizes the large-sample error; sd $1/\sqrt5$).
  • Top-hat (box): $K(u) = \tfrac12$ for $|u| \le 1$ (sd $1/\sqrt3$). Triangular: $K(u) = 1 - |u|$ for $|u| \le 1$ (sd $1/\sqrt6$).

Because these standard scales differ, "$h = 1$" means a different bump width for different kernels and in different libraries. Compare kernels only after matching their spreads.

Why do we need it?

To stop worrying about the wrong knob. Beginners often try many kernels; the real lever is the bandwidth. Knowing this saves time and avoids fake "kernel effects" that are really width effects.

Where is it used?

sklearn.neighbors.KernelDensity(kernel='epanechnikov' | 'tophat' | 'gaussian' | ...); SciPy's gaussian_kde is Gaussian only; compact kernels (Epanechnikov, box) are used when speed matters, because far-away points contribute exactly zero.

How is it used?

Use the Gaussian kernel by default. If you switch kernels, convert the bandwidth so the spread stays the same (divide the Gaussian $h$ by the kernel's own standard deviation), and check how your library defines its bandwidth.

Gaussian Epanechnikov Top-hat (box) Triangular
Four kernels scaled to the same standard deviation (the dashed grey curve is the Gaussian, for reference). Their shapes differ, but bumps of the same spread give almost the same KDE.

The same 40 data points, smoothed with four kernels. (1) With match the spreads on, the four curves almost lie on top of each other at every bandwidth. (2) Turn it off: now each kernel uses the same number $h$ as its own width parameter, and the box and tent look much rougher. That difference is about width, not shape.

"I tried the Epanechnikov kernel and my KDE changed a lot, so the kernel matters."

Most likely the effective width changed. In scikit-learn, bandwidth is the Gaussian's standard deviation but the half-width of the box and Epanechnikov kernels, so the same number gives narrower compact bumps. Match spreads before comparing.

"Compact kernels are always better because they are optimal."

The Epanechnikov kernel is optimal only by a few percent in large samples. The Gaussian kernel's smoothness and never-zero tails are often more convenient.

Kernel = bump shape (Gaussian, Epanechnikov, box, tent). With matched spreads, KDEs look almost the same.

Efficiency vs Epanechnikov: Gaussian 95%, box 93%. The bandwidth matters far more than the kernel.

Trap: "h" means different widths for different kernels and libraries.

Quick check: a triangular kernel on $[-1, 1]$ has sd $1/\sqrt6$. What half-width matches a Gaussian kernel with $h = 0.5$?

Half-width $\times\, 1/\sqrt6 = 0.5$, so half-width $= 0.5\sqrt6 \approx 1.22$.

Bandwidth: the knob that really matters core

The bandwidth $h$ is how wide each bump is. With a tiny $h$, every data point gets its own sharp spike: the curve follows every accident of this particular sample. That is overfitting (also called under-smoothing): draw a new sample and the spikes move. With a huge $h$, all bumps melt into one wide hill: real features such as a second peak are smoothed away. That is underfitting (over-smoothing). A good $h$ sits in between, and it should shrink slowly as you collect more data.

Three ways to say it:

  • Picture: $h$ is the blur of a camera: too little blur shows every speck of dust, too much blur loses the faces.
  • Numbers: for 100 residuals with $s = 2$ and IQR $= 2.5$, Silverman's rule gives $h \approx 0.66$ and SciPy's default gives $h \approx 0.80$.
  • Slogan: bandwidth matters far more than the kernel.

$n = 100$ residuals, sample standard deviation $s = 2.0$, IQR $= 2.5$.

  1. Common factor: $n^{-1/5} = 100^{-0.2} \approx 0.398$.
  2. SciPy's default (gaussian_kde, "Scott's factor"): kernel sd $h = s \cdot n^{-1/5} = 2.0 \times 0.398 \approx 0.80$.
  3. Silverman's rule of thumb: $h = 0.9 \cdot \min\!\big(s, \tfrac{\text{IQR}}{1.349}\big)\, n^{-1/5}$. Here $\text{IQR}/1.349 = 2.5/1.349 \approx 1.85$, which is smaller than $s = 2$, so $h = 0.9 \times 1.85 \times 0.398 \approx 0.66$.
  4. Normal-reference rule: $h = 1.06\, s\, n^{-1/5} = 1.06 \times 2.0 \times 0.398 \approx 0.84$.
  5. Why $n^{-1/5}$: with 32 times more data ($n = 3200$), all three shrink by a factor of $32^{1/5} = 2$. More data allows narrower bumps.

The bandwidth trades two kinds of error (the same bias–variance trade-off you will meet for estimators in Chapter 5.1):

  • Small $h$: low bias (the curve can follow sharp features), high variance (it changes a lot from sample to sample).
  • Large $h$: low variance, high bias (peaks are flattened, valleys filled, tails pushed out).

For smooth densities the best $h$ shrinks like $n^{-1/5}$. Rules of thumb for a Gaussian kernel (they are exactly right only for Normal data):

  • Silverman: $h = 0.9\,\min(s,\ \text{IQR}/1.349)\,n^{-1/5}$ (R's default bw.nrd0; statsmodels' bw='silverman').
  • Normal reference / Scott: $h \approx 1.06\, s\, n^{-1/5}$ (statsmodels uses $1.059\min(s, \text{IQR}/1.349)$ for 'scott' and its default 'normal_reference').
  • SciPy gaussian_kde: default bw_method='scott' gives kernel sd $= s\,n^{-1/5}$; 'silverman' gives $s\,(3n/4)^{-1/5} \approx 1.06\,s\,n^{-1/5}$ (the normal-reference rule, not the 0.9 rule). A number passed as bw_method is a factor that multiplies $s$.
  • Data-driven choices (cross-validation, the Sheather–Jones plug-in) adapt better to multimodal or skewed data.
Why do we need it?

The same data can look unimodal or bimodal, smooth or spiky, depending only on $h$. Without understanding $h$ you cannot tell whether a feature of a density plot is in the data or in the smoothing.

Where is it used?

Every KDE: seaborn's kdeplot (with bw_adjust to scale SciPy-style rules), posterior density plots, density-based anomaly detection, mean-shift clustering, and kernel smoothing in general (the same idea gives Nadaraya–Watson regression).

How is it used?

Start from the default rule, then plot with half and double the bandwidth (bw_adjust=0.5 and 2 in seaborn). Trust features that survive. If the data may have several peaks, prefer a smaller $h$ or a data-driven rule.

h = 0.1: too small (noisy) h = 0.4: about right h = 1.5: too large (smooths away a peak)
The same 60 points (tick marks) with three bandwidths. Blue: the KDE; green dashed: the true two-peaked density. Too small: spiky noise. About right: both peaks. Too large: the two peaks melt into one wide hill.

Top: 150 points from a two-peaked distribution (green dashed = truth, which in real life you never see). Bottom: the error between the KDE and the truth (integrated squared error) for every bandwidth, on a log scale of $h$. (1) Slide $h$ to the far left: spikes, large error. Far right: one blob, large error again. The best $h$ is at the bottom of the U. (2) Press Silverman and SciPy default: both land to the right of the bottom of the U. These rules are built for one bell, so on two-peaked data they over-smooth (here the error is roughly double the best). (3) Press New sample: the U and its minimum move a little.

Ten different samples of 100 points from the same two-peaked distribution (green dashed). Thin blue lines: the ten KDEs. Thick blue: their average. (1) With a small $h$ the thin curves disagree wildly (high variance), yet their average follows the truth (low bias). (2) With a large $h$ the thin curves all agree (low variance), but they agree on the wrong shape: the peaks are too low and the valley is filled in (high bias). (3) The bars split the average error into these two parts; find the $h$ where their sum is smallest.

"gaussian_kde(x, bw_method='silverman') uses Silverman's rule of thumb $0.9\min(s, \text{IQR}/1.349)n^{-1/5}$."

SciPy's 'silverman' is $(3n/4)^{-1/5}$ times $s$, about $1.06\,s\,n^{-1/5}$: the normal-reference rule. The 0.9 version is R's bw.nrd0 and statsmodels' bw='silverman'. Names differ between libraries; check the formula.

"KernelDensity(bandwidth='scott') in scikit-learn adapts to my data's spread."

In scikit-learn, 'scott' and 'silverman' give $n^{-1/5}$ and $(3n/4)^{-1/5}$ without multiplying by the data's standard deviation, so they assume standardized data. And the default bandwidth=1.0 is in the data's own units. Standardize first, or set the bandwidth yourself.

"The smoother curve is the better one."

Smoother means more bias. Too much smoothing hides peaks and fattens tails. Look at several bandwidths.

"The key decision in a KDE is which kernel to use."

The bandwidth dominates. Kernels with matched spread give nearly identical estimates; the bandwidth decides between a noisy and an over-smoothed picture.

Model answer: "A KDE puts a kernel of width $h$ on every point and averages. A small $h$ gives low bias and high variance (spiky, overfit); a large $h$ gives low variance but high bias (over-smoothed, can hide modes). Rules of thumb like Silverman's assume roughly Normal data, so for multimodal or skewed data I check smaller bandwidths or use cross-validation."

Small $h$: spiky, low bias, high variance (overfit). Large $h$: smooth, high bias (underfit). Best $h \propto n^{-1/5}$.

Silverman: $0.9\min(s, \text{IQR}/1.349)n^{-1/5}$. SciPy default: $s\,n^{-1/5}$. Both are rules of thumb for one-bell data.

Traps: library names differ (SciPy 'silverman' ≈ 1.06 rule; sklearn 'scott' ignores the data's sd).

Quick check: with $n = 100$ the rule gives $h = 0.8$. Roughly what does it give with $n = 3200$ from the same distribution?

The rule scales like $n^{-1/5}$, and $3200/100 = 32 = 2^5$, so $h$ halves: about $0.4$.

KDE versus histogram: two views of the same idea

A histogram is secretly a KDE with box-shaped bumps whose positions are locked to a grid. That locked grid is why the bin origin matters. Now imagine drawing the histogram many times, each time sliding the grid a little, and averaging all of them. The jumps caused by any single grid position cancel out, and the average becomes a smooth curve: exactly a KDE with tent-shaped (triangular) bumps.

Three ways to say it:

  • Picture: stack many slightly shifted histograms on top of each other; the staircase melts into a smooth hill.
  • Numbers: averaging 32 shifted histograms of width $w$ gives almost exactly the triangular-kernel KDE with half-width $h = w$.
  • Slogan: histogram = boxes on a grid; KDE = bumps centred on the data.

One data point at $x = 0.4$, bin width 1 (so $n = 1$ and every non-empty bar has density $1/(1 \times 1) = 1$).

  1. Grid starting at 0: the point is in $[0, 1)$, so the histogram is 1 on $[0, 1)$ and 0 elsewhere.
  2. Grid starting at 0.5: the point is in $[-0.5, 0.5)$, so the histogram is 1 on $[-0.5, 0.5)$.
  3. Average of the two: $0.5$ on $[-0.5, 0)$, $1$ on $[0, 0.5)$, $0.5$ on $[0.5, 1)$: a small staircase peaked near the point.
  4. With more and more shifts the staircase becomes the tent $1 - |x - 0.4|$ on $[-0.6, 1.4]$: a triangular kernel of half-width 1 centred on the point.

The average shifted histogram (ASH) with bin width $w$ and $m$ equally spaced origins converges, as $m \to \infty$, to the KDE with the triangular kernel and $h = w$:

$$\frac{1}{m}\sum_{j=1}^{m} \hat f_{\text{hist}, j}(x) \;\longrightarrow\; \frac{1}{n w}\sum_{i=1}^{n}\Big(1 - \frac{|x - x_i|}{w}\Big)_+ ,$$

where $(z)_+ = \max(z, 0)$. The reason: for a randomly placed grid, the chance that $x$ and $x_i$ share a bin is $1 - |x - x_i|/w$ (when this is positive).

HistogramKDE
Knobsbin width and originbandwidth (and kernel)
Lookstep functionsmooth curve
Strengthsexact counts, fast for huge $n$, honest for discrete data and point masses (many exact zeros)no origin effect, easy to overlay several groups, smooth peaks
Weaknessesblocky, origin-dependentleaks past boundaries, smooths away sharp edges and spikes, can suggest impossible values (2.5 orders)
Why do we need it?

Choosing the right picture avoids false conclusions: a KDE of integer counts invents in-between values, a histogram with a bad origin invents bumps. Seeing them as one family tells you what each one smooths and why.

Where is it used?

Overlaying the residual distributions of several models (KDE), reporting how many users had exactly zero revenue (histogram or bar chart), comparing variants in an experiment report, and histogram-plus-KDE plots such as seaborn.histplot(x, kde=True).

How is it used?

For continuous data, draw both with matched smoothing (bin width about 2 to 3 bandwidths is a reasonable rough match). For counts or data with spikes at special values, use bars per value and do not smooth.

60 data points (ticks). Blue: the average of $m$ histograms whose grids are shifted by $w/m$ each. Orange dashed: the KDE with a tent kernel of half-width $w$. (1) With $m = 1$ you have an ordinary, blocky histogram. (2) Raise $m$ to 4, 8, 32: the blue staircase converges onto the orange curve. (3) Change $w$: the limit changes with it, because $w$ plays the role of the bandwidth.

"A KDE is always better than a histogram because it is smooth."

Smooth is not always honest. Daily order counts are whole numbers, revenue per user has a spike at exactly 0, and a metric can have a hard cap. A KDE smears all of these; a histogram (or a bar per value) shows them.

"KDE removes all arbitrary choices."

It removes the bin origin, but the bandwidth is still a choice, and it matters just as much as the bin width did.

Histogram = box bumps on a fixed grid; KDE = bumps centred on the data. Averaging shifted histograms → triangular-kernel KDE with $h = w$.

Histogram: exact counts, discrete data, spikes. KDE: smooth, overlays, no origin effect.

Trap: a KDE of counts or zero-inflated data invents values that cannot happen.

Quick check: you want to show the distribution of the number of orders per user (0, 1, 2, … with 60% zeros). Histogram/bar chart or KDE?

A bar chart with one bar per count. The data are whole numbers with a big spike at 0; a KDE would spread that spike into negative values and between the integers.

Boundary bias: when the data cannot go below zero

Waiting times, revenue, session lengths and daily counts can never be negative. But a KDE does not know that. It puts a symmetric bump on every point, so the bumps of points close to 0 spill over to the negative side. Two things go wrong at once: the curve shows density where no data can exist, and just above 0 the curve is too low, because part of the mass leaked across the wall.

Three ways to say it:

  • Picture: piles of sand next to a wall spill over the top; the sand on the other side is lost to your side.
  • Numbers: a point at 0.2 with $h = 0.5$ sends $\Phi(-0.4) \approx 34\%$ of its bump below zero.
  • Slogan: a KDE does not know about walls unless you tell it.
  1. One observation at $x_i = 0.2$, Gaussian kernel, $h = 0.5$. The part of its bump below 0 is $P(N(0.2, 0.5^2) \lt 0) = \Phi\big(\tfrac{0 - 0.2}{0.5}\big) = \Phi(-0.4) \approx 0.345$.
  2. For the whole dataset the leaked mass is the average over points: $\frac1n\sum_i \Phi(-x_i/h)$.
  3. Exponential data with rate 1 (true density $e^{-x}$, height 1 at 0). With $h = 0.3$ the plain KDE at 0 is, on average, only about $0.40$: less than half the truth.
  4. Reflection fix: give every point a mirror copy $-x_i$, add the mirrored bumps to the original ones (still dividing by $n$, not $2n$), and keep only $x \ge 0$. On the positive side the mirrored bumps return exactly the mass that leaked. At 0 the reflected KDE is exactly twice the plain one, about $0.80$ on average here.
  5. Still not 1: reflection forces the curve to be flat at the wall, while $e^{-x}$ is steep there. A smaller $h$ helps ($h = 0.1$ gives about 0.92).

Boundary bias: near the edge $b$ of the data's possible range (the support: the set of values the variable can take), a standard KDE underestimates the density and puts mass outside the support. For data with $x \ge b$, the reflection estimator is

$$\hat f_R(x) = \frac{1}{nh}\sum_{i=1}^n \left[K\!\left(\frac{x - x_i}{h}\right) + K\!\left(\frac{x - (2b - x_i)}{h}\right)\right] \quad \text{for } x \ge b, \qquad \hat f_R(x) = 0 \text{ for } x \lt b.$$
  • It integrates to 1 over $[b, \infty)$: the leaked mass is folded back.
  • Its slope at $b$ is zero, so it is still biased if the true density is steep at the wall.
  • Alternatives: estimate the KDE of $\log x$ and transform back (multiply by the Jacobian $1/x$), or use special boundary kernels. Simply cutting the plot at 0 (for example cut=0 or clip in seaborn) hides the leak but does not put the lost mass back.
Why do we need it?

Most business metrics are positive and many pile up near zero (waiting times, small orders, sparse counts). A plain KDE draws them with a fake dip at zero and a tail into impossible negative values, which can mislead a likelihood choice.

Where is it used?

Density plots of revenue, latencies, durations, prices, posterior draws of scale parameters (σ, τ, the Negative Binomial concentration α) that live on $(0, \infty)$, and probabilities on $[0, 1]$ (two walls).

How is it used?

Reflect at the wall (by hand: gaussian_kde(np.r_[x, -x]), evaluate on $x \ge 0$ and multiply by 2), or KDE on the log scale. For data with an exact spike at 0, report the share of zeros separately and draw a KDE of the positive values only.

boundary: values cannot be below 0 leaked mass bump of one point at 0.3 mirror copy at −0.3 (dotted) true density e^(−x) (dashed)
One data point at 0.3 gets a bump (orange) whose left part (red) falls below 0, where no data can exist. Reflection adds a mirror point at −0.3; on the positive side its bump (dotted) gives back exactly the mass that leaked.

300 waiting times between orders (minutes, exponential, true density in green). Blue: the plain KDE. The red area is density that leaked below 0, where no waiting time can be. (1) Look at the height at 0: the truth is 1, the blue curve is much lower. (2) Switch on reflect at 0: the orange curve folds the leak back and roughly doubles the height at 0. (3) Make $h$ smaller: both biases shrink, but the curve gets noisier.

"The KDE shows some negative revenue, so a few users must have negative revenue (refunds)."

Check the raw data first. A KDE puts density below 0 for any data that sits near 0, refunds or not. The leak is a property of the estimator.

"I cut the plot at 0, so the boundary problem is solved."

Cutting hides the negative part but the curve near 0 is still too low and the area is less than 1. Reflect or transform instead.

In an A/B framework like yours, a revenue-per-user metric is never negative and usually has a big spike at exactly 0 (users who did not buy). A KDE of all users smears that spike into a bump around 0 with a fake negative tail. Report the zero share (a conversion rate, Bernoulli/Beta) and draw a KDE of the positive amounts on a log scale: that matches how such metrics are often modelled (a yes/no part and a positive-amount part). In your forecasting model, daily demand counts are non-negative whole numbers: show them as a bar chart of counts; KDEs of positive posterior draws (the noise scale σ, the Negative Binomial concentration) need reflection or a log scale too.

Near a wall (e.g. $x \ge 0$) a KDE leaks mass outside and underestimates the density at the wall (about half, for small $h$).

Fix: reflect, $\hat f_R(x) = \frac{1}{nh}\sum[K(\frac{x-x_i}{h}) + K(\frac{x+x_i}{h})]$ for $x \ge 0$; or KDE on $\log x$.

Trap: cutting the plot at 0 hides the leak but does not fix it.

Quick check: a data point sits exactly at 0 and you use a Gaussian kernel. What fraction of its bump leaks below 0?

$\Phi(0) = 0.5$: exactly half, whatever the bandwidth. Reflection puts the mirror point also at 0, which doubles the positive half back to a full unit of area.

Multimodal data: a big bandwidth can hide a second peak

Order values from two kinds of customers (small top-ups and big monthly baskets), daily demand on weekdays versus weekends, conversion rates of two user segments: such data have two (or more) peaks, called modes. A mode is a local high point of the density, a most-typical value. A distribution with two of them is bimodal; with several, multimodal. If the bandwidth is too large, the KDE melts the peaks into one and you would wrongly conclude there is one homogeneous group.

Three ways to say it:

  • Picture: two hills close together look like one wide hill through a blurry lens.
  • Numbers: two equal Normal groups with sd 1 and centres 3 apart are bimodal, but a Gaussian KDE with $h \gt 1.12$ will on average show one peak.
  • Slogan: rules of thumb assume one bell; when you suspect groups, try a smaller bandwidth.

Two equally large groups: $\tfrac12 N(-d/2, 1) + \tfrac12 N(d/2, 1)$, centres $d$ apart.

  1. The true mixture is bimodal only when $d \gt 2$ (two standard deviations apart). With $d = 1.5$ the two groups exist, but the density still has one peak.
  2. On average, a Gaussian KDE with bandwidth $h$ equals the true density blurred by an $N(0, h^2)$: each group's variance grows from $1$ to $1 + h^2$.
  3. So the KDE (on average) shows two peaks only when $d \gt 2\sqrt{1 + h^2}$.
  4. For $d = 3$: two peaks need $\sqrt{1 + h^2} \lt 1.5$, that is $h \lt \sqrt{1.25} \approx 1.12$.
  5. For $d = 2.5$: two peaks need $h \lt \sqrt{1.5625 - 1} = 0.75$. Groups that are only a little apart need a small bandwidth to be seen at all, and therefore a lot of data.

A mode of a density is a local maximum. A density is unimodal with one mode, bimodal with two, multimodal with several. Mixtures of groups are the usual cause.

  • The number of modes of a Gaussian KDE can only go down as $h$ grows (a known property of the Gaussian kernel). The smallest $h$ at which the KDE becomes unimodal (the "critical bandwidth") is the basis of Silverman's test for multimodality.
  • Rules of thumb (Silverman, Scott) are tuned for one Normal bell; on multimodal data the between-group spread inflates $s$, so they tend to over-smooth.
  • Two groups do not always make two peaks: if they overlap a lot, the mixture is unimodal but wider or skewed.
Why do we need it?

A hidden second group changes everything downstream: a single mean is meaningless, a Normal likelihood is wrong, and a segment that behaves differently goes unnoticed. Being able to reveal (or rule out) a second mode is a core exploratory skill.

Where is it used?

Segment discovery (two customer types), mixture models and clustering, residual checks (a bimodal residual distribution signals a missing regressor such as a weekend or holiday indicator), and MCMC diagnostics (a bimodal posterior can trap samplers and variational guides).

How is it used?

Plot the KDE with the default rule and with half of it (bw_adjust=0.5). If a second peak appears, check that it survives resampling, then look for a variable that separates the groups (day of week, device, country).

200 values from two equal groups whose centres are $d$ apart (green dashed: the true density). Purple dots mark the peaks the KDE finds. (1) With $d = 3$ and a small $h$, two peaks. Raise $h$ past about 1.1: one peak. (2) Press Silverman h: does the rule keep both peaks? (3) Set $d = 1.5$: the truth itself has only one peak, even though there are two groups.

"The KDE has one peak, so the data come from one homogeneous group."

Groups closer than about two standard deviations make a single peak, and a large bandwidth merges peaks that are further apart. One peak is compatible with several groups. Look for a variable that might split them.

"The KDE has three small peaks, so there are three groups."

With a small bandwidth, noise makes bumps. A real mode should survive a moderately larger bandwidth and a new sample (or a bootstrap resample).

If the KDE of your forecasting model's residuals shows two bumps, the model is probably missing a variable that splits days into two kinds: weekends against weekdays, holiday periods, promo days, a changepoint the trend missed. In the A/B framework, a two-humped distribution of a per-user metric suggests two segments; that is exactly where per-segment parameters with partial pooling (Chapter 6.5) earn their keep.

Mode = local peak. Two equal Normal groups (sd 1) are bimodal only if their centres are more than 2 apart.

Gaussian KDE blurs each group to variance $1 + h^2$: two peaks visible (on average) only if $d \gt 2\sqrt{1 + h^2}$.

Rules of thumb over-smooth multimodal data: also look at half the bandwidth. Bumps must survive resampling.

Quick check: two equal groups with sd 1 and centres 4 apart. Up to which bandwidth does a Gaussian KDE (on average) keep both peaks?

Need $4 \gt 2\sqrt{1 + h^2}$, so $\sqrt{1 + h^2} \lt 2$, $1 + h^2 \lt 4$, $h \lt \sqrt3 \approx 1.73$.

The curse of dimensionality: why KDE stays in one or two dimensions

A KDE estimates the density at a point by looking at the data nearby. On a line, a small window around a point contains plenty of data. But as you add more variables (dimensions), "nearby" empties out. To catch even 10% of the data in a box around a point, the box must stretch across most of the range of every variable, and then it is no longer "nearby" at all. This is the curse of dimensionality: space grows so fast with the number of dimensions that any fixed amount of data becomes thin.

Three ways to say it:

  • Picture: 1000 marbles on a 1-metre stick are crowded; the same 1000 marbles spread through a 10-dimensional cube are lonely.
  • Numbers: to hold 10% of uniform data, a box must cover 10% of each axis in 1D, 32% in 2D, 46% in 3D and 79% in 10D.
  • Slogan: in high dimensions, everything is far from everything.

Data spread uniformly over the unit cube $[0, 1]^d$. A cube-shaped neighbourhood with side $\ell$ holds the fraction $\ell^d$ of the data.

  1. To hold the fraction $p$, we need $\ell^d = p$, so $\ell = p^{1/d}$.
  2. $p = 0.1$: $d = 1$: $\ell = 0.10$; $d = 2$: $\ell = \sqrt{0.1} \approx 0.32$; $d = 3$: $\ell \approx 0.46$; $d = 10$: $\ell = 0.1^{0.1} \approx 0.79$; $d = 100$: $\ell \approx 0.98$.
  3. The other way round: a fixed box of side 0.2 holds the fraction $0.2^d$. With $n = 1000$ points: $1000 \times 0.2 = 200$ points in 1D, $40$ in 2D, $8$ in 3D, $0.32$ in 5D, and $0.0001$ in 10D.
  4. KDE accuracy improves with $n$ like $n^{-4/(d+4)}$ (for its mean squared error). To halve the error you need $2^{(d+4)/4}$ times more data: about 2.4 times in 1D, but about 11 times in 10D.

The curse of dimensionality is the collection of problems that appear when the number of dimensions $d$ grows: the volume of space grows exponentially, so a fixed sample covers it ever more thinly; local neighbourhoods become nearly empty or nearly global; distances between points become similar to each other.

  • For kernel density estimation, the best achievable mean squared error falls like $n^{-4/(d+4)}$: $n^{-0.8}$ in 1D, $n^{-0.29}$ in 10D. The sample size needed for a fixed accuracy grows exponentially in $d$.
  • This is why KDE is used for 1D and 2D pictures (and occasionally up to a handful of dimensions), and why high-dimensional problems need structure: parametric models, independence assumptions, or dimension reduction first (PCA, Chapter 5.16).
Why do we need it?

It explains why "just estimate the joint density of all my features" does not work, why nearest-neighbour methods degrade with many features, and why statistical models impose structure instead of learning everything from scratch.

Where is it used?

Choosing between nonparametric and parametric methods, k-nearest-neighbours and kernel methods in many dimensions, anomaly detection on many features, and the reason MCMC and VI in high-dimensional posteriors rely on gradients and structure rather than grids (Chapter 6.9).

How is it used?

Use KDE for one or two variables at a time (marginals, pairwise plots). With more variables, reduce dimension first, assume a parametric form (a multivariate Normal), or model the variables conditionally (a regression) rather than jointly.

d = 1 side 0.10 d = 2 side 0.32 d = 3 side 0.46 each blue region holds 10% of uniformly spread data
To hold 10% of uniformly spread data, a neighbourhood must cover 10% of a line, 32% of each side of a square and 46% of each edge of a cube. In 10 dimensions it needs 79% of every axis: it is no longer a local neighbourhood.

Bars show the side length a cube neighbourhood needs to hold the share $p$ of uniformly spread data, for 1 to 30 dimensions (1 = the whole range of each variable). (1) Move $d$ from 1 to 10 and watch the purple bar climb towards 1. (2) Lower $p$ to 1%: even then, by $d = 20$ the "neighbourhood" covers about 80% of every axis. (3) Press Simulate 500 points to count how many random points really land in a small central box.

"KDE is nonparametric, so it works for any number of features if I have enough data."

"Enough data" grows exponentially with the dimension. Beyond a few dimensions no realistic dataset is enough. Use KDE for 1D/2D views; model higher dimensions with structure.

"My 2D plots look fine, so the 20-dimensional data is well covered."

Every pairwise plot can look dense while the full 20-dimensional space is almost empty. Low-dimensional views cannot show joint emptiness.

Uniform data: a cube holding share $p$ needs side $p^{1/d}$ (10%: 0.10, 0.32, 0.46, …, 0.79 in 10D).

KDE error $\propto n^{-4/(d+4)}$: data needs explode with $d$. Use KDE in 1–2 dimensions.

High dimensions need structure: parametric models, conditioning, or dimension reduction.

Quick check: how large must a cube be to hold 1% of uniform data in 2 dimensions? In 20?

$0.01^{1/2} = 0.1$ of each axis in 2D (small and local). $0.01^{1/20} \approx 0.79$ of each axis in 20D: almost the whole range, even though it holds only 1% of the data.

Looking at your model's residuals with a KDE core

A forecasting model with a Normal likelihood says: "after trend, seasonality, holidays and regressors, what is left over is Normal noise." The leftovers are the residuals $e_t = y_t - \hat y_t$ (actual minus fitted). Their histogram and KDE are the quickest picture of whether that promise holds. Put the KDE next to a Normal curve with the same mean and standard deviation, and look for differences in the centre, the symmetry, the peak and the tails.

Three ways to say it:

  • Picture: lay the residuals' smooth outline over the bell curve the likelihood assumes, and look where they disagree.
  • Numbers: a Normal puts only 0.27% of residuals beyond 3 standard deviations: about 1 day in a year. Seven such days means heavy tails.
  • Slogan: the residuals are the model's confession; the KDE lets you read it.

A year of daily residuals ($n = 365$), mean $0.3$ orders, standard deviation $10$ orders.

  1. Normal expectation beyond $\pm 3$ sd ($\pm 30$ orders): $2 \times (1 - \Phi(3)) \approx 0.0027$, so $365 \times 0.0027 \approx 1$ day.
  2. You count 7 such days. That is about 7 times more extreme days than the Normal predicts.
  3. The KDE also shows a narrower, taller peak than the Normal curve: most days are quieter than the sd suggests, and a few are wild. A narrow peak plus fat tails is the classic heavy-tail signature.
  4. Conclusion: a Normal likelihood will let those 7 days pull the fit; a Student-t likelihood would treat them as less surprising and give them less influence (Chapter 7.13).

For a fitted model, the residual of observation $t$ is $e_t = y_t - \hat y_t$. A residual-distribution check compares the KDE (and histogram) of the $e_t$ with the noise distribution the likelihood assumes, fitted with the same location and scale. Things to read:

  • Centre: clearly away from 0 → systematic bias.
  • Symmetry: a long tail on one side → skew (transformation, Chapter 4.18, or a different likelihood).
  • Peak and tails: taller peak with fatter tails than the Normal → heavy tails (Student-t). Excess kurtosis (Chapter 4.14) gives a number for this.
  • Several bumps: a missing variable that splits the observations into kinds.

The KDE is good for the bulk but weak in the tails, where there are few points. On a log-density scale a Normal is a downward parabola and heavy tails bend away from it; the Q-Q plot (Chapter 4.17) is the sharper tool for tails. A density check also says nothing about whether residuals are independent over time or have constant spread (Chapter 7.17).

Why do we need it?

The likelihood is an assumption about the noise. If the residuals contradict it (heavy tails, skew, two bumps), the uncertainty intervals and the influence of unusual days are wrong. A quick residual KDE catches the big mismatches before you trust the forecasts.

Where is it used?

Regression diagnostics, choosing between Normal and Student-t likelihoods in forecasting, checking a Gaussian assumption before a t-test or a confidence interval, and monitoring production models (a residual distribution that changes over time signals drift).

How is it used?

Compute residuals on the training data (and on a holdout), plot sns.histplot(e, stat='density', kde=True), overlay stats.norm(e.mean(), e.std()).pdf, count residuals beyond 3 sd, look at the log-density, then confirm with a Q-Q plot.

fit the model residualseₜ = yₜ − ŷₜ histogram + KDE(several bandwidths) overlay a Normalsame mean and sd tall peak + fat tails→ try Student-t (7.13) long tail on one side→ transform (4.18) two bumps→ missing variable
A residual-distribution check in four steps and the three most common findings. Tails are confirmed with a Q-Q plot (Chapter 4.17); independence and constant spread need plots against time and fitted values (Chapter 7.17).

300 residuals (orders). Blue: their KDE (SciPy's default bandwidth). Orange: the Normal with the same mean and sd. Teal: a Student-t fit with ν = 4. (1) Normal noise: blue and orange agree. (2) Heavy tails: blue has a taller, narrower peak and the count beyond 3 sd jumps. (3) Switch on log scale: the Normal becomes a parabola that drops fast, while heavy tails stay up. (4) Skewed and 3 outliers show other signatures.

"The residual KDE looks like a bell, so the Normal likelihood is fine."

The bulk can look Normal while the tails are much heavier; the KDE is weakest exactly where it matters (few points in the tails), and a large bandwidth flattens the tall narrow peak that signals heavy tails, so the residuals look more Normal than they are. Count extreme residuals, use the log scale, and confirm with a Q-Q plot.

"The residuals are Normal, so the model is correct."

A good-looking distribution says nothing about time structure. Residuals can be perfectly Normal and still autocorrelated (a missed seasonality) or have a spread that grows with the level. Plot them against time and against fitted values too (Chapter 7.17).

Your forecasting model offers Normal and Student-t likelihoods, and this plot is the everyday tool for choosing between them: fit with the Normal, compute residuals $e_t = y_t - \hat y_t$ (using the posterior mean or median prediction), draw the KDE with the Normal overlay, and count days beyond 3 sd. A tall narrow peak with several extreme days is the signature that the Student-t will fit better. For the Negative Binomial (count) likelihood, raw residuals are not supposed to look Normal or even symmetric, because the noise of counts is skewed and grows with the mean; check those models with probability integral transform (PIT) values instead (Chapter 7.16).

"I checked normality of the residuals with a KDE and it looked fine."

A KDE shows the bulk and is smoothed by a bandwidth choice; it is a first look, not a test of the tails.

Model answer: "I look at a histogram and KDE of the residuals with a fitted Normal overlay for the overall shape, count residuals beyond 3 sd against the 0.27% the Normal predicts, and confirm the tails with a Q-Q plot. If the tails are heavy I switch to a Student-t likelihood. I also check residuals over time and against fitted values, because the distribution plot cannot see autocorrelation or changing variance."

Residuals $e_t = y_t - \hat y_t$. Overlay a KDE on a Normal with the same mean and sd.

Heavy tails: tall narrow peak + more than 0.27% beyond 3 sd → Student-t. Skew → transform. Two bumps → missing variable.

Traps: KDE tails are unreliable (use log scale + Q-Q plot); a Normal shape says nothing about time structure.

Quick check: 730 daily residuals, 9 of them beyond 3 sd. Is that unusual for Normal noise?

Expected: $730 \times 0.0027 \approx 2$. Nine is more than four times that. For a Poisson count with mean 2, nine or more has probability about 0.0002, so this is strong evidence of heavier-than-Normal tails.

Recap, cheat sheet and practice

  • A histogram counts values in bins. Heights can be counts, relative frequencies or densities $c_j/(n w_j)$ (areas sum to 1; required for overlays and unequal bins).
  • It has two hidden choices: bin width and bin origin. Rules of thumb (Sturges, Freedman–Diaconis, NumPy 'auto') are starting points; always try a few.
  • A KDE puts one bump of area $1/n$ on every point: $\hat f_h(x) = \frac{1}{nh}\sum K\big(\frac{x-x_i}{h}\big)$. The kernel shape matters little; the bandwidth $h$ matters a lot.
  • Small $h$: spiky, low bias, high variance (overfit). Large $h$: smooth, high bias (underfit), can hide modes. Rules of thumb (Silverman, SciPy's Scott default) assume one bell; library names and conventions differ.
  • Averaging shifted histograms gives a triangular-kernel KDE: the two are one family.
  • Failures: boundary bias for positive data (fix: reflection or log scale), hidden modes with large $h$, and the curse of dimensionality (KDE belongs in 1–2 dimensions).
  • For residuals: KDE + Normal overlay + count beyond 3 sd + log scale, then a Q-Q plot.

Cheat sheet

IdeaFormula / rulePython
Histogram density$c_j/(n w_j)$np.histogram(x, bins, density=True)
Bin width rulesSturges: $\text{range}/(\log_2 n + 1)$; FD: $2\,\text{IQR}\,n^{-1/3}$bins='sturges' | 'fd' | 'auto'
KDE$\frac{1}{nh}\sum K\big(\frac{x - x_i}{h}\big)$stats.gaussian_kde(x)(grid)
SciPy default bandwidthkernel sd $= s\,n^{-1/5}$bw_method='scott' (a number = factor × $s$)
Silverman's rule$0.9\min(s, \text{IQR}/1.349)\,n^{-1/5}$statsmodels bw='silverman'
Reflection at 0$\frac{1}{nh}\sum[K(\frac{x-x_i}{h}) + K(\frac{x+x_i}{h})]$, $x \ge 0$2 * gaussian_kde(np.r_[x, -x])(grid)
Two-group bimodalitytrue: $d \gt 2\sigma$; KDE: $d \gt 2\sqrt{\sigma^2 + h^2}$sns.kdeplot(x, bw_adjust=0.5)
Curse of dimensionalityside for share $p$: $p^{1/d}$; error $\propto n^{-4/(d+4)}$—
Code it · Python
import numpy as np
from scipy import stats
from sklearn.neighbors import KernelDensity

# --- histogram: counts vs density ---
r = np.array([-7, -4, -3, -1, 0, 1, 2, 2, 5, 9])
edges = [-10, -5, 0, 5, 10]
counts, _ = np.histogram(r, bins=edges)
dens, _ = np.histogram(r, bins=edges, density=True)
print(counts)                                   # [1 3 4 2]
print(dens, (dens * np.diff(edges)).sum())      # [0.02 0.06 0.08 0.04] 1.0  (areas sum to 1)
print(np.histogram_bin_edges(r, bins="auto"))   # [-7.  -3.8 -0.6  2.6  5.8  9. ]  5 bins of width 3.2

# --- KDE by hand, then with SciPy ---
d, h = np.array([1.0, 2.0, 4.0]), 1.0
f2 = stats.norm.pdf((2 - d) / h).sum() / (len(d) * h)
print(round(f2, 4))                             # 0.2316
kde = stats.gaussian_kde(d, bw_method=h / d.std(ddof=1))   # bw_method = FACTOR that multiplies the data's sd
print(kde(2.0).round(4))                        # [0.2316]

# --- default bandwidths (rules of thumb) ---
rng = np.random.default_rng(1)
x = rng.normal(0, 2, size=100)
n, s = len(x), x.std(ddof=1)
iqr = np.subtract(*np.percentile(x, [75, 25]))
k = stats.gaussian_kde(x)                       # default bw_method='scott'
print(round(k.factor, 4), round(n ** -0.2, 4))  # 0.3981 0.3981
print(round(np.sqrt(k.covariance[0, 0]), 4), round(s * n ** -0.2, 4))   # 0.6814 0.6814  kernel sd h = s * n^(-1/5)
print(round(0.9 * min(s, iqr / 1.349) * n ** -0.2, 4))                  # 0.5292  Silverman's 0.9 rule (R's bw.nrd0)
print(round(stats.gaussian_kde(x, bw_method="silverman").factor, 4),
      round((3 * n / 4) ** -0.2, 4))            # 0.4217 0.4217  SciPy's 'silverman' = the ~1.06 rule

# --- boundary bias and reflection for positive data ---
w = rng.exponential(1.0, size=300)              # true density e^(-x), equal to 1 at x = 0
h = 0.3
plain = stats.gaussian_kde(w, bw_method=h / w.std(ddof=1))
both = np.r_[w, -w]                             # data plus mirror copies
refl = stats.gaussian_kde(both, bw_method=h / both.std(ddof=1))
print(plain(0.0).round(3), (2 * refl(0.0)).round(3))   # [0.432] [0.864]  truth is 1; reflected = 2 x plain at 0
print(round(stats.norm.cdf(-w / h).mean(), 3))          # 0.108  share of the plain KDE's area below 0

# --- scikit-learn: bandwidth is in data units and defaults to 1.0 ---
kd = KernelDensity(kernel="gaussian", bandwidth=h).fit(w[:, None])
print(np.exp(kd.score_samples([[0.0]])).round(3))       # [0.432]  same as plain(0.0)
Test yourself

1. 500 values; the bin $[10, 12)$ contains 50 of them. What is the density height of its bar?

Density $= c/(n w) = 50/(500 \times 2) = 0.05$. (0.1 is the relative frequency $50/500$; the area of the bar, $0.05 \times 2$, equals it.)

2. Which choice changes the look of a KDE the most?

With matched spreads, kernels give nearly identical curves (the Gaussian is about 95% as efficient as the best kernel). The bandwidth decides between spiky and over-smoothed.

3. You draw a Gaussian KDE of 50 points with a very small bandwidth. What do you see?

Each bump is very narrow, so they do not overlap: the curve chases every individual point. That is overfitting (high variance).

4. What does scipy.stats.gaussian_kde(x) use as its bandwidth by default?

SciPy's default 'scott' factor is $n^{-1/5}$, multiplied by the data's standard deviation. The fixed 1.0 is scikit-learn's default; the 0.9 rule is Silverman's (R, statsmodels); SciPy does not cross-validate.

5. Waiting times (all positive, many close to 0). What does a plain Gaussian KDE do near 0?

Bumps of points near 0 spill below 0 (boundary bias). The curve shows impossible negative values and is too low just above 0; for small $h$ it is about half the true height. Reflection or a log scale fixes most of this.

6. Two equal groups, each Normal with sd 1, centres 2.5 apart. Which statement is right?

Bimodal needs $d \gt 2$, true here. The KDE blurs each group to variance $1 + h^2$, so it needs $2.5 \gt 2\sqrt{1 + h^2}$, i.e. $h^2 \lt 0.5625$, $h \lt 0.75$.

Practice problems

A. Data: 2, 3, 3, 7. Bins $[0,2)$, $[2,4)$, $[4,6)$, $[6,8]$. Give the counts, relative frequencies and densities.

Counts: $[0,2)$: 0 (the value 2 belongs to $[2,4)$); $[2,4)$: 3 (2, 3, 3); $[4,6)$: 0; $[6,8]$: 1 (7). Relative frequencies: $0, 0.75, 0, 0.25$. Densities $c/(n w) = c/(4 \times 2)$: $0, 0.375, 0, 0.125$. Check: areas $0.375 \times 2 + 0.125 \times 2 = 0.75 + 0.25 = 1$.

B. Data $-1, 0, 3$; Gaussian kernel, $h = 1$. Compute $\hat f(0)$.

Distances in bandwidths: $(0-(-1))/1 = 1$, $0$, $(0-3)/1 = -3$. Bump heights: $\varphi(1) = 0.2420$, $\varphi(0) = 0.3989$, $\varphi(-3) = 0.0044$. Sum $= 0.6453$; divide by $nh = 3$: $\hat f(0) \approx 0.215$. The point at 3 is three bandwidths away and contributes almost nothing.

C. $n = 243$, $s = 6$, IQR $= 6.745$. Compute Silverman's rule, SciPy's default and the 1.06 normal-reference rule.

$243 = 3^5$, so $n^{-1/5} = 1/3$. $\text{IQR}/1.349 = 5.0 \lt s$. Silverman: $0.9 \times 5.0 \times \tfrac13 = 1.5$. SciPy default: $6 \times \tfrac13 = 2.0$. Normal reference: $1.06 \times 6 \times \tfrac13 = 2.12$. Silverman's rule is smallest because the IQR suggests the bulk is narrower than the sd (a sign of heavy tails or outliers).

D. Positive data. A point at 0.5 with $h = 0.25$: what share of its bump leaks below 0? And a point at 0?

At 0.5: $\Phi(-0.5/0.25) = \Phi(-2) \approx 0.023$, about 2%. At 0: $\Phi(0) = 0.5$, half the bump. Points within about one bandwidth of the wall cause most of the leak.

E. How large must a cube neighbourhood be to contain 1% of uniformly spread data in 2 dimensions? In 20? What does this mean for KDE?

$0.01^{1/2} = 0.10$ of each axis in 2D; $0.01^{1/20} \approx 0.79$ in 20D. In 20 dimensions, a box that holds just 1% of the data spans almost the whole range of every variable, so "local" averaging is impossible. KDE needs exponentially more data as $d$ grows; use it for one or two variables at a time.

F. Interview: "Your residual KDE looks bell-shaped. Is the Normal likelihood justified?"

"Not from that alone. The KDE describes the bulk and depends on the bandwidth; a large bandwidth flattens the tall narrow peak of heavy-tailed residuals, and the tails themselves have few points. I would overlay a Normal with the same mean and sd, count residuals beyond 3 sd against the 0.27% a Normal predicts, look at the log-density, and confirm with a Q-Q plot. If the tails are heavy, I would try a Student-t likelihood. I would also plot residuals against time and fitted values, since the distribution plot cannot reveal autocorrelation or changing variance."

Chapter 4.17 · Syllabus Module 6

Q-Q plots

A histogram shows you roughly what your data looks like. A Q-Q plot answers a sharper question: does my data have the same shape as a particular distribution, especially at the extremes? It is the plot you will use to decide whether a Normal likelihood is good enough for your residuals, or whether a Student-t would be better.

  • Say in plain words what a Q-Q plot compares: the sorted data against the quantiles a model predicts
  • Build one by hand: sort, give each point a probability label $(i-0.5)/n$, look up the model's quantile, plot, add a reference line
  • Know that shifting and stretching the data never bends the line: a Q-Q plot checks shape, and its slope and intercept give the spread and the centre
  • Read the shapes: straight, curved (skew), S-shaped (heavy or light tails), lonely points (outliers), steps (two groups, rounded or count data)
  • Judge how much wiggle is normal for your sample size with a simulation envelope
  • Use other references, especially a Student-t, and know what a P-P plot is and why it is weaker in the tails
  • Run the full residual check for your models: "is a Normal likelihood reasonable here, or would a Student-t be better?"

The idea: line up your data against what a model predicts core

Imagine a class of 20 children standing in a queue from shortest to tallest. Next to them stands a second queue of 20 "model children": the heights a Normal distribution says a typical class would have, also from shortest to tallest. Now pair them up: shortest with shortest, second-shortest with second-shortest, and so on, up to tallest with tallest.

If the real class has the same shape as the model, every real child is "the model child, a bit shifted and a bit stretched". Plot each pair as a dot (model height across, real height up) and the dots fall on a straight line. If the real class has, say, a few unusually tall children, the dots at the tall end leave the line. The place where the dots bend away tells you where the shapes differ.

That is the whole idea of a Q-Q plot ("quantile–quantile plot"). A quantile is a cut point: the 0.9-quantile (also called the 90th percentile) is the value with 90% of the data below it (you met quantiles in Chapter 4.4). A Q-Q plot puts the data's quantiles against the model's quantiles.

Why not just look at a histogram? A histogram is good at the middle, where most of the data lives. But the important question for choosing a likelihood is usually about the extremes (the "tails"), and in a histogram the extremes are a few tiny bars that change when you change the bins. In a Q-Q plot every single point, including the most extreme one, gets its own dot.

Three ways to say it:

  • Picture: pair the smallest with the smallest and the largest with the largest, then check whether the pairs fall on a straight line.
  • Numbers: if your data's 0.99-quantile is 28 and a standard Normal's is 2.33, the dot goes at (2.33, 28); if every dot has the same ratio of about 12, the dots lie on one straight line.
  • Slogan: straight line = same shape; a bend = a different shape, and the bend shows you where.

A year of daily forecast errors ("residuals", in orders) has these quantiles. We compare them with the quantiles of the standard Normal $N(0,1)$, the bell curve with mean 0 and standard deviation 1.

level $p$0.010.250.500.750.99
standard Normal quantile $z_p$−2.33−0.6700.672.33
Data A quantile−28−80828
Data B quantile−45−70745
  1. Data A, the ratio "data quantile ÷ Normal quantile" at the outer levels: $28 / 2.33 \approx 12.0$.
  2. Data A at the quartiles: $8 / 0.674 \approx 11.9$. Same ratio, so all five dots lie on one straight line through 0 with slope about 12. Data A has a Normal shape with a standard deviation of about 12 orders.
  3. Data B at the outer levels: $45 / 2.33 \approx 19.3$.
  4. Data B at the quartiles: $7 / 0.674 \approx 10.4$. The ratio is different: the middle dots follow a line with slope about 10, but the end dots sit far beyond it.
  5. Conclusion: Data B has a narrower middle and much longer tails than any Normal. Its plot bends away from the line at both ends. That is the classic sign of heavy tails.

A Q-Q plot of data $x_1, \dots, x_n$ against a reference distribution with quantile function $Q$ (the inverse of its CDF) is the scatter plot of the points

$$\bigl(\,Q(p_i),\; x_{(i)}\,\bigr), \qquad i = 1, \dots, n,$$
  • $x_{(i)}$ ("x sub i in brackets") is the $i$-th smallest data value. The sorted values $x_{(1)} \le x_{(2)} \le \dots \le x_{(n)}$ are called the order statistics.
  • $p_i$ is a probability level ("plotting position") that says which quantile $x_{(i)}$ represents, for example $p_i = (i - 0.5)/n$. The next section explains it.
  • $Q(p_i)$ is the theoretical quantile: where the reference distribution puts its $p_i$ point. It goes on the horizontal axis; the data ("sample quantiles") go on the vertical axis. (This is the SciPy and statsmodels convention; some software swaps the axes, so always read the axis labels.)
  • If the data have the same shape as the reference, the points lie close to a straight line.
Why do we need it?

Models assume a shape for the noise (Normal, Student-t, …). If the assumed shape is wrong in the tails, the model is overconfident or is pulled around by extreme points. A histogram is too blurry at the extremes to tell; a Q-Q plot shows every extreme point clearly.

Where is it used?

Checking residuals of linear regression and of your Prophet-style forecasting model, choosing Normal vs Student-t likelihoods, checking a metric before an A/B test, checking that standardized data or z-scores look Normal, and comparing two samples' shapes.

How is it used?

Call scipy.stats.probplot(x, dist="norm", plot=ax) or statsmodels.api.qqplot(x, line="s"). Look at the middle first (straight?), then the two ends (bending away?), then lonely points. Decide whether the departure matters for your purpose.

your data, sorted (blue) what a Normal model predicts, sorted (orange) pair them model quantile → ↑ data
Left: the sorted data (blue) and the sorted values a Normal model predicts (orange), paired smallest with smallest. Right: each pair becomes one dot. These seven dots lie close to the purple straight line, so the data has a Normal-like shape.

Both samples have 300 values, the same mean and the same standard deviation. Switch between Normal noise and Heavy-tailed noise and compare the two histograms: they look almost the same bell. Now compare the Q-Q plots: the heavy-tailed one bends away from the purple line at both ends. Press New sample a few times. Points outside the window are pinned to the edge as rings.

"A Q-Q plot shows the data over time" or "against the fitted values".

It shows the sorted data against a model's quantiles. The time order is thrown away when you sort. (Residuals over time and against fitted values are separate plots; you will meet them in Chapter 7.17.)

"If the histogram looks like a bell, the data is Normal enough."

Histograms hide the tails, which is exactly where Normal and heavy-tailed noise differ. Use a Q-Q plot to judge the extremes.

"The Q-Q plot tests whether the data is exactly Normal."

Real data is never exactly Normal. A Q-Q plot shows how and where the shape differs, so you can decide whether the difference matters for your model.

In your forecasting model with a Normal likelihood, the noise $\epsilon_t$ is assumed Normal. A Normal says a day 4 standard deviations from the forecast happens about once in 15 800 days. If your residuals show such days every few months, the Normal is badly wrong in the tails, and a Student-t likelihood (which your model also offers) may fit better. The Q-Q plot of the residuals is the quickest way to see this. The same check applies to a revenue-per-user metric in your A/B framework before you choose a Normal or Student-t likelihood for it.

Q-Q plot = points $(Q(p_i),\, x_{(i)})$: model quantiles across, sorted data up.

Straight line → same shape. Bend → different shape; where it bends tells you where (middle or tails).

Trap: a bell-shaped histogram says nothing reliable about the tails.

Quick check: the standard Normal's 0.9-quantile is 1.28 and your data's 0.9-quantile is 6.4. Where does this pair go on the Q-Q plot?

At the point $(1.28,\ 6.4)$: the model's quantile goes across, the data's quantile goes up. If the data were Normal with mean 0, this point would be on a line of slope $6.4/1.28 = 5$, which would mean a standard deviation of about 5.

Building a Q-Q plot step by step core

The recipe has five steps: sort the data, give each sorted value a probability label ("this one is the 10% point"), ask the model where its own 10% point is, plot the pairs, and draw a reference line to compare against.

The only new idea is the probability label. With 5 values, think of the probability line from 0 to 1 cut into 5 equal slices, one slice per value. Each value sits in the middle of its slice: 0.1, 0.3, 0.5, 0.7, 0.9. Why the middle and not the end of the slice? If the largest value got the label 1.0, we would ask the Normal for its 100% point, and that is infinity.

Three ways to say it:

  • Picture: cut the line from 0 to 1 into $n$ equal slices and put each sorted value in the middle of its slice.
  • Numbers: for $n = 5$ the labels are 0.1, 0.3, 0.5, 0.7, 0.9, and the Normal's quantiles there are −1.28, −0.52, 0, 0.52, 1.28.
  • Slogan: sort, label, look up, plot, compare.

Five daily forecast errors (in orders): $1, -3, 0, 3, -1$.

  1. Sort: $x_{(1)}, \dots, x_{(5)} = -3, -1, 0, 1, 3$.
  2. Label: $p_i = (i - 0.5)/5$ gives $0.5/5 = 0.1$, $1.5/5 = 0.3$, $0.5$, $0.7$, $0.9$.
  3. Look up the standard Normal quantiles $z_i = \Phi^{-1}(p_i)$ (the inverse of the Normal CDF, scipy.stats.norm.ppf): $-1.28, -0.52, 0, 0.52, 1.28$.
  4. Plot the pairs $(z_i, x_{(i)})$: $(-1.28, -3)$, $(-0.52, -1)$, $(0, 0)$, $(0.52, 1)$, $(1.28, 3)$.
  5. Reference line $x = \bar x + s\,z$. Here $\bar x = 0/5 = 0$ and $s = \sqrt{(9 + 1 + 0 + 1 + 9)/4} = \sqrt{5} \approx 2.24$.
  6. Line values at the five $z_i$: $2.24 \times (-1.28) \approx -2.87$, then $-1.17$, $0$, $1.17$, $2.87$. The data $-3, -1, 0, 1, 3$ are within about 0.2 of the line everywhere. With only five points, that is as straight as it gets.

Plotting positions. The $i$-th smallest value is matched with the level

$$p_i = \frac{i - a}{n + 1 - 2a}, \qquad i = 1, \dots, n,$$

for a constant $a$ between 0 and 1. Different tools pick different $a$:

  • $a = 0.5$ gives $p_i = (i - 0.5)/n$, the "middle of the slice" rule used in this guide (and by LA.stats.qq).
  • $a = 0$ gives $p_i = i/(n+1)$: the default of statsmodels qqplot / ProbPlot.
  • scipy.stats.probplot uses Filliben's approximation to the medians of the order statistics, and R's qqnorm uses $a = 3/8$ when $n \le 10$ and $a = 0.5$ otherwise.

The choice only visibly moves the two or three most extreme points when $n$ is small. It never changes the overall shape.

Reference lines. Common choices (with their statsmodels names):

  • line="s": the "standardized" line $\bar x + s\,z$ (mean and standard deviation of the data).
  • line="q": the line through the first and third quartiles (R's qqline). It ignores the tails, so it is not dragged around by extreme points. Most widgets in this chapter use it.
  • line="r": a least-squares line through the points (SciPy's probplot fits this one and reports its slope, intercept and correlation $r$).
  • line="45": the line $y = z$. Only meaningful if the data are already standardized (or you compare with a fully specified distribution).
Why do we need it?

If you know the five steps, a Q-Q plot stops being a black box: you can read any point ("this residual is my 99% point, and the Normal puts its 99% point here"), and you understand why different libraries give slightly different pictures.

Where is it used?

scipy.stats.probplot, statsmodels qqplot and ProbPlot, R's qqnorm/qqline, residual panels of regression and forecasting libraries, and probability plots in quality control.

How is it used?

Usually one call does all five steps. By hand: xs = np.sort(x), p = (np.arange(1, n+1) - 0.5) / n, z = norm.ppf(p), plt.scatter(z, xs), then add xs.mean() + xs.std(ddof=1) * z as the line.

theoretical quantile Q(pᵢ) (the model) → sorted data x₍ᵢ₎ (sample quantile) ↑ above the line: this value is larger than the model expects reference line: intercept ≈ centre, slope ≈ spread each dot = one data value, paired with the model's quantile
Anatomy of a Q-Q plot. Across: where the model puts its quantiles. Up: the sorted data. A dot above the line (red) is larger than the model predicts at that level; a dot below is smaller.

Press Next to walk through the recipe with eight daily forecast errors. Watch the table fill in one column at a time: sorted value, probability label $p_i = (i-0.5)/8$, Normal quantile $z_i$, and the reference-line value. In step 4 notice that the shaded slices under the bell all have the same area, 1/8. Press New data to repeat with other numbers.

"Use $p_i = i/n$."

Then the largest value gets $p_n = 1$, and every distribution with an unbounded right tail (Normal, Student-t) puts its 100% point at infinity. Use $(i - 0.5)/n$ or another rule that keeps every $p_i$ strictly between 0 and 1.

"sm.qqplot(x, line="45") on my raw residuals shows they are not Normal: the points are far from the line."

The 45° line $y = z$ assumes mean 0 and standard deviation 1. Raw residuals with a standard deviation of 12 will lie on a line with slope 12. Use line="s" or line="q", or standardize first (or pass fit=True).

Recipe: sort → $p_i = (i-0.5)/n$ → $z_i = \Phi^{-1}(p_i)$ → plot $(z_i, x_{(i)})$ → line $\bar x + s z$ (or through the quartiles).

Libraries use slightly different $p_i$ (statsmodels $i/(n+1)$, SciPy Filliben); only the extreme points move a little.

Trap: the 45° line only makes sense for standardized data.

Quick check: with $n = 4$ values, what are the probability labels $p_i = (i-0.5)/n$, and the Normal quantiles?

$p = 0.125, 0.375, 0.625, 0.875$. The standard Normal quantiles there are $\Phi^{-1}(0.125) \approx -1.15$, $\Phi^{-1}(0.375) \approx -0.32$, $+0.32$ and $+1.15$ (symmetric, because the Normal is symmetric).

Reading the straight line: shifting and stretching never bend it core

Convert temperatures from Celsius to Fahrenheit, or measure residuals in hundreds of orders instead of orders. The numbers change, but the shape of the data (one hump, symmetric, how long the tails are) does not. A Q-Q plot is built to see shape only, so these changes must not bend it. They don't: shifting the data (adding a number) moves the line up or down, and stretching it (multiplying by a positive number) tilts the line. It stays straight.

This is great news: you never need to know the mean $\mu$ or the standard deviation $\sigma$ before drawing the plot. You compare your data with the standard Normal, and the line itself tells you the centre (where it crosses $z = 0$) and the spread (its slope).

Three ways to say it:

  • Picture: shift moves the line up and down, stretch tilts it, and neither bends it.
  • Numbers: for $N(100, 15^2)$ the 10%, 50% and 90% points are 80.8, 100 and 119.2, plotted against $z = -1.28, 0, 1.28$: a line with intercept 100 and slope 15.
  • Slogan: straightness = shape, intercept = centre, slope = spread.

Daily orders that follow a Normal with mean $\mu = 100$ and standard deviation $\sigma = 15$, so $X = 100 + 15Z$ with $Z \sim N(0, 1)$.

  1. The 10% point of $Z$ is $z_{0.1} = -1.2816$. Shifting and stretching keeps the order of values, so the 10% point of $X$ is $100 + 15 \times (-1.2816) = 100 - 19.22 = 80.78$.
  2. The 50% point: $100 + 15 \times 0 = 100$.
  3. The 90% point: $100 + 15 \times 1.2816 = 119.22$.
  4. Plot $(-1.2816, 80.78)$, $(0, 100)$, $(1.2816, 119.22)$. Slope between the outer points: $(119.22 - 80.78)/(1.2816 - (-1.2816)) = 38.44 / 2.563 = 15$. Intercept (value at $z = 0$): 100.
  5. So the line is $x = 100 + 15z$: the intercept is the mean and the slope is the standard deviation. Every level $p$ gives a point on this same line.

If $X = \mu + \sigma Z$ with $\sigma \gt 0$, then for every level $p$

$$Q_X(p) = \mu + \sigma\, Q_Z(p).$$
  • So the theoretical Q-Q plot of $X$ against $Z$ is the straight line with intercept $\mu$ and slope $\sigma$. For a large sample from $N(\mu, \sigma^2)$, the dots scatter closely around $x = \mu + \sigma z$.
  • This works for any location–scale family: a family where every member is a shifted and stretched copy of one standard shape. Normal, Student-t with a fixed $\nu$, Laplace, logistic and uniform are such families.
  • The slope is the family's scale parameter. For the Normal the scale is the standard deviation. For a Student-t it is not: the standard deviation is $\sigma\sqrt{\nu/(\nu-2)}$ (for $\nu \gt 2$).
  • The correlation $r$ between the theoretical and the sample quantiles (the "probability plot correlation coefficient", PPCC) measures straightness. It does not change when you shift or stretch the data.
Why do we need it?

It explains why one plot against the standard Normal works for data of any size and units, and it lets you read the centre and the spread straight off the plot. It also tells you what the plot cannot see: the mean and the sd are not part of the shape check.

Where is it used?

Every Normal Q-Q plot of residuals, scipy.stats.probplot (which reports the fitted slope, intercept and $r$), the PPCC tests for normality, and quick estimates of a noise scale from a plot.

How is it used?

Plot against $N(0,1)$. If the dots are straight, read the intercept as the centre and the slope as the spread. In SciPy: (osm, osr), (slope, intercept, r) = probplot(x).

Move the centre μ slider: the whole cloud of dots moves up and down. Move the spread σ slider: the cloud tilts. Watch the straightness score $r$ in the readout: it does not change at all. Then pick Right-skewed sample: the dots bend, and no setting of μ or σ can straighten them, because a shift and a stretch cannot change shape.

"The dots are far from the 45° line, so my data is not Normal."

Far from the 45° line can just mean the mean is not 0 or the spread is not 1. Shape problems show up as a bend, not as a straight line with a different slope or intercept.

"The slope of a Q-Q plot against a Student-t is the standard deviation."

It is the t's scale parameter $\sigma$. The standard deviation is larger: $\sigma\sqrt{\nu/(\nu-2)}$. For $\nu = 3$ it is $\sigma\sqrt{3} \approx 1.73\sigma$.

In your forecasting model the noise is $\epsilon_t \sim N(0, \sigma^2)$, or a Student-t with scale $\sigma$. If the Normal Q-Q plot of the residuals is straight, its slope is a rough estimate of the noise sd $\sigma$, and its intercept should be near 0 (a clearly non-zero intercept means the model is biased: it under- or over-forecasts on average). For a Student-t likelihood, NumPyro's dist.StudentT(df, loc, scale) takes this scale, not the standard deviation.

$Q_X(p) = \mu + \sigma Q_Z(p)$: Normal data gives a straight line, intercept ≈ mean, slope ≈ sd.

A Q-Q plot checks shape only; shift and stretch never bend it.

Trap: against a Student-t, the slope is the scale, not the sd.

Quick check: the Normal Q-Q plot of your daily residuals is a straight line that crosses $z = 0$ at 2 and has slope 12. What does it tell you?

The residuals look Normal in shape, with mean about 2 and standard deviation about 12 orders. A mean of 2 means the model under-forecasts by about 2 orders on average (residual = actual − forecast), which is worth fixing.

Curved plots: skewed data core

Think of order values in a shop: most orders are small, a few are very large, and none are below zero. The data is bunched on the left with a long tail on the right. This is right skew (also called positive skew).

Compare it with a Normal of the same mean and spread. The Normal has a long left tail too, so its small quantiles go far below the mean; the real data cannot, so its smallest values sit above the line. In the middle the real data is bunched up, so it sits a little below the line. At the far right the real data has the longer tail, so it climbs above the line again. Above, below, above: the dots trace a curve that bends upward, like a smile.

Left skew (a long tail to the left, for example "days of stock remaining" when most shops are well stocked) is the mirror image: the curve bends downward, like a frown.

Three ways to say it:

  • Picture: right skew = a smile (curve bends up); left skew = a frown (curve bends down).
  • Numbers: Exponential data vs a Normal with the same mean 1 and sd 1: at 1% the data is at 0.01 but the Normal at −1.33; at 50%, 0.69 vs 1.00; at 99%, 4.61 vs 3.33.
  • Slogan: both ends on the same side of the line means skew.

Waiting times between orders often follow an Exponential distribution. Take Exponential with rate 1: its mean is 1 and its standard deviation is 1. The matching Normal is $N(1, 1)$, whose $p$-quantile is $1 + z_p$.

level $p$0.010.100.500.900.99
Exponential quantile $-\ln(1-p)$0.010.110.692.304.61
Normal quantile $1 + z_p$−1.33−0.281.002.283.33
data vs lineaboveabovebelowonabove
  1. Left end ($p = 0.01$): $0.01 \gt -1.33$. The data cannot go below 0, so it is far above the Normal's line.
  2. Middle ($p = 0.5$): $0.69 \lt 1.00$. The median of skewed data is below its mean, so the dot is below the line.
  3. Right end ($p = 0.99$): $4.61 \gt 3.33$. The long right tail puts the dot above the line again.
  4. Above–below–above is a curve that bends upward: a smile. That is the signature of right skew.

Skewness measures asymmetry: positive (right) skew has the longer tail on the right and usually mean > median; negative (left) skew is the mirror image. (The skewness number itself is defined in Chapter 4.14.)

  • Right skew on a Normal Q-Q plot: a convex curve (its slope keeps increasing from left to right). Both ends lie above a line fitted to the middle.
  • Left skew: a concave curve (slope decreasing). Both ends lie below the line.
  • Why the slope tells you this. The local slope of the Q-Q curve at level $p$ equals $\dfrac{f_{\text{ref}}(Q_{\text{ref}}(p))}{f_{\text{data}}(Q_{\text{data}}(p))}$, the reference density divided by the data density at the matching points. Where the data is thin (a long tail), its quantiles are far apart, so the slope is steep.
Why do we need it?

A Normal likelihood is symmetric. If the noise is skewed, the model puts too much probability on impossible values (negative demand) and too little on big positive surprises. Spotting the smile early tells you to transform or change the likelihood.

Where is it used?

Revenue per user, order values, session lengths and waiting times (right skew), low-volume count series in forecasting, and residuals of a model that should have used a log scale.

How is it used?

If you see a smile, try a log or Box-Cox transform and redraw the Q-Q plot (Chapter 4.18), or switch to a skewed likelihood (log-normal, Gamma, or Negative Binomial for counts).

Drag the skew slider to the right: the density (left) grows a long right tail and the Q-Q curve (right) bends into a smile. Drag it to the left for a frown. At 0 the curve is the straight purple line. Turn on Show a sample of 100 and press New sample: real data scatters around the curve, so a mild skew is hard to see with few points.

"The plot curves, so the data has outliers."

A smooth bend running through the whole plot is a shape (skew). Outliers are a few isolated dots that break away from an otherwise straight pattern.

"Standardize the data (z-scores) to fix the skew."

A z-score is a shift and a stretch, and those never change shape (previous section). To reduce skew you need a non-linear transform such as a log (Chapter 4.18).

In your A/B framework, a revenue-per-user metric is typically strongly right-skewed (many zeros and small values, a few big spenders). Its Q-Q plot against a Normal is a steep smile, which warns you that a plain Normal likelihood on the raw values is a poor description, and that a log transform, a heavy-tailed likelihood or a model for zeros may be needed. In your forecasting model, low-volume demand counts are right-skewed too, which is one reason the model offers a Negative Binomial likelihood.

Right skew → convex "smile": both ends above the line. Left skew → concave "frown": both ends below.

Local slope of the Q-Q curve = reference density ÷ data density: thin tail → steep.

Trap: z-scores do not remove skew; a log or Box-Cox transform can.

Quick check: on a Normal Q-Q plot, the lowest 10 dots sit below the line and the highest 10 dots also sit below the line. What shape is the data?

Both ends on the same side (below) with the middle above is a concave "frown": the data is left-skewed, with a long tail toward small values.

S-shapes: heavy tails, light tails and outliers core

Some processes are calm most of the time and then, now and then, produce a big surprise: a viral day, a system outage, a bulk order. Their data has heavy tails (also called fat tails): extreme values happen far more often than a Normal would allow. To keep the same overall spread, the middle of such data is usually tighter than a Normal's.

On a Normal Q-Q plot that gives a stretched S: the middle dots are flatter than the line, the lowest dots dive below it, and the highest dots shoot above it. The two ends bend away from the line in opposite directions (unlike skew, where both ends go the same way).

Light tails are the opposite: values are boxed in and extremes are rarer than Normal (for example a percentage that can only move within a narrow band). The ends tuck back toward the middle: a reversed S.

An outlier is different again: one or a few lonely points far from an otherwise straight pattern. Heavy tails bend the ends gradually; an outlier jumps.

Three ways to say it:

  • Picture: heavy tails = the ends flare away from the line (S); light tails = the ends tuck in (reversed S); outlier = one stray dot.
  • Numbers: Student-t with 3 degrees of freedom, rescaled to sd 1: its 75% point is 0.44 (Normal: 0.67) but its 99.9% point is 5.90 (Normal: 3.09).
  • Slogan: ends flaring out → heavy tails → think Student-t.

Compare the quantiles of a Student-t with $\nu = 3$, rescaled to standard deviation 1 (divide by its sd $\sqrt{3} \approx 1.732$), with the standard Normal.

level $p$0.750.950.990.999
Student-t(3) quantile0.7652.3534.54110.215
÷ 1.732 (sd 1)0.441.362.625.90
standard Normal0.671.642.333.09
  1. At $p = 0.75$ and $0.95$ the t's quantiles are smaller ($0.44 \lt 0.67$, $1.36 \lt 1.64$): the middle is tighter, so those dots sit below the line on the right half.
  2. The two curves cross near $p \approx 0.98$.
  3. Beyond that the t's quantiles are much larger ($2.62 \gt 2.33$, and $5.90 \gt 3.09$): the far dots shoot above the line.
  4. By symmetry the left half is the mirror image (above the line in the middle, far below at the end). Flat middle, flaring ends: an S.
  5. In counts: a Normal puts 0.27% of values beyond 3 sd; this t puts 1.4% there, about 5 times more. For 365 days that is about 5 days instead of 1.

The tails of a distribution are the regions far from its centre. "Tail weight" describes how much probability lives there.

  • Heavy (fat) tails: $P(|X - \mu| \gt k\sigma)$ shrinks more slowly than for a Normal as $k$ grows (Student-t, Laplace). Excess kurtosis (Chapter 4.14) is usually positive. On a Normal Q-Q plot: an S-shape, low end below the line and high end above it.
  • Light (thin) tails: extremes rarer than Normal, often because values are bounded (uniform: excess kurtosis −1.2). Q-Q plot: a reversed S, low end above the line and high end below it.
  • Outlier: an observation far from the bulk that the rest of the pattern does not lead you to expect. Q-Q plot: one or a few isolated dots at an end, with the rest straight. Outliers can be data errors, or genuine rare events (a holiday you did not model).
Why do we need it?

The tails decide how surprised a model is by extreme days. A Normal likelihood treats a 5-sd residual as almost impossible, so a few such days drag the fitted trend toward them and make the uncertainty too narrow. Knowing the tails are heavy tells you to use a heavier-tailed likelihood.

Where is it used?

Choosing Normal vs Student-t likelihoods in your two projects, robust regression, risk models in finance (fat-tailed returns), anomaly detection, and deciding whether a few extreme points are data errors or real behaviour.

How is it used?

Look at the two ends of a Normal Q-Q plot. Opposite flares (S): consider Student-t. A few stray dots: inspect those rows (dates, IDs) before deciding anything. Tucked-in ends: usually harmless, but check bounded data.

The slider changes the shape $\beta$ of one family of symmetric distributions (always rescaled to mean 0 and sd 1). At $\beta = 2$ it is exactly the Normal and the Q-Q curve is the straight line. At $\beta = 1$ it is the Laplace: notice the S, with the ends flaring away. Push $\beta$ below 1 for even heavier tails. Push it up to 8: the distribution becomes box-like (light tails) and the S reverses. Compare $P(|X| \gt 3)$ with the Normal's 0.0027 in the readout.

Ten forecast errors sit on the strip at the top; drag any of them left or right. The Q-Q plot below re-sorts and redraws at once. Press One outlier: only one dot leaves the purple line. Press Heavy tails: the ends bend away gradually on both sides (an S). Try Right skew and Light tails too. The purple line goes through the quartiles, so a single extreme dot cannot drag it around.

"The ends bend away, so those points are outliers: delete them."

A gradual S-shape means the process itself has heavy tails. Deleting the extremes makes the noise look calmer than it is, and the model becomes overconfident. Keep them and use a likelihood that expects them (Student-t).

"An S-shape means the data is skewed."

Skew bends both ends to the same side of the line. Heavy tails bend them to opposite sides: the low end below, the high end above.

"Light tails are as dangerous as heavy tails for a Normal likelihood."

Usually much less: the Normal just slightly overstates how often extremes happen, which makes intervals a little too wide. Bounded data near its limits (rates near 0 or 1) is the case to watch; a Beta or a logit scale fits it better.

In your forecasting model, daily residuals with an S-shaped Normal Q-Q plot are the classic reason to switch the likelihood from Normal to Student-t. Isolated stray dots are a different story: look up their dates first. If they are holidays or promotions, the better fix is a holiday or event regressor ($h(t)$ in your model), not a heavier tail.

Heavy tails: S-shape (low end below, high end above the line). Light tails: reversed S. Outlier: a lone dot.

Skew = both ends on the same side; tails = ends on opposite sides.

Trap: never delete the ends of an S; model them (Student-t).

Quick check: on a Normal Q-Q plot of 200 residuals, the middle is straight, the lowest 5 dots are far below the line and the highest 5 are far above it. What do you conclude?

Opposite flares at both ends with a straight middle: heavy tails. The residuals produce more extreme values in both directions than a Normal allows, so a Student-t likelihood is worth trying.

How much wiggle is normal? Sample size and simulation envelopes core

Even data that really comes from a Normal never gives a perfectly straight Q-Q plot. The dots wobble, and the wobble is biggest at the two ends (where there are few points) and when the sample is small. So before you announce "heavy tails!", ask: would truly Normal data of this size wiggle this much?

The honest way to answer is to simulate. Generate many fake datasets of the same size from a Normal, draw all their Q-Q curves in grey, and shade the band that contains 95% of them at each position. That band is a simulation envelope. If your dots stay inside it, their wiggles are the ordinary wiggles of Normal data.

Three ways to say it:

  • Picture: a grey cloud of fake-Normal Q-Q curves; your real dots should sit inside it.
  • Numbers: with 20 truly Normal points, the largest one (plotted at $z = 1.96$) lands anywhere between 0.96 and 3.02 in 95% of samples.
  • Slogan: before you trust a bend, compare it with the wobble.

How far can the largest of $n = 20$ standard Normal values wander? It is plotted at $z = \Phi^{-1}(19.5/20) = \Phi^{-1}(0.975) = 1.96$.

  1. The largest value is $\le x$ only if all 20 values are $\le x$. Independent values multiply: $P(\max \le x) = \Phi(x)^{20}$.
  2. Lower 2.5% point: solve $\Phi(x)^{20} = 0.025$, so $\Phi(x) = 0.025^{1/20} = e^{\ln(0.025)/20} = e^{-0.1844} = 0.8316$, giving $x = \Phi^{-1}(0.8316) \approx 0.96$.
  3. Upper 97.5% point: $\Phi(x) = 0.975^{1/20} = 0.99873$, giving $x \approx 3.02$.
  4. So in 95% of Normal samples of size 20, the top dot lies between 0.96 and 3.02, while the line says 1.96. A top dot at 2.9 is not evidence of a heavy tail when $n = 20$.
  5. With $n = 1000$ the same calculation gives a band of 2.68 to 4.05 around $z = 3.29$: still wide at the very end, but now the rest of the plot is tight, so a systematic S becomes easy to see.

A pointwise simulation envelope for a Normal Q-Q plot of $n$ points:

  1. Simulate $B$ datasets of size $n$ from the reference (in practice from $N(\bar x, s^2)$, using your data's mean and sd; this is a "parametric bootstrap").
  2. Sort each simulated dataset. For each rank $i$, collect the $B$ values of the $i$-th smallest.
  3. The envelope at rank $i$ runs from the 2.5% to the 97.5% point of those $B$ values.

When the reference is fully known, no simulation is needed: the $i$-th smallest of $n$ uniform values follows a $\text{Beta}(i,\, n + 1 - i)$ distribution, so the band is $Q\bigl(\text{Beta}(i, n+1-i)\text{ quantiles}\bigr)$. The widget below uses this exact version.

  • It is pointwise: each rank has its own 95% band. Even for truly Normal data, about 5% of the dots fall outside on average, and with $n = 50$ there is roughly a 47% chance that at least one dot pokes out (checked by simulation).
  • Neighbouring sorted values move together, so wiggles come in smooth waves, not as independent jitter.
Why do we need it?

Without a sense of normal wobble you will over-read small samples (seeing heavy tails in noise) and under-read huge samples (where everything looks slightly non-Normal). The envelope gives you a yardstick tied to your exact sample size.

Where is it used?

Residual diagnostics in regression and forecasting, the "half-normal plot with envelope" for GLMs, the qqPlot function of R's car package, and posterior predictive checks (which use the same "simulate fake data, compare" idea, Chapter 6.8).

How is it used?

Simulate 200–1000 Normal datasets with your data's $n$, mean and sd; sort each; take np.percentile(sims, [2.5, 97.5], axis=0); shade it behind your Q-Q dots. Look for systematic runs outside the band at the ends, not single dots.

Start with Truly Normal data and n = 20, and press New sample many times: the dots wander a lot, yet they stay inside the green band. Turn on show 19 simulated Normal samples to see where the band comes from. Now switch to Student-t (ν = 5): with n = 20 you cannot tell it apart from Normal; with n = 1000 the ends leave the band clearly.

"Two dots are outside the envelope, so the data is not Normal."

The band is pointwise 95%: about 5% of dots fall outside even for perfect Normal data. What matters is a systematic run of dots outside the band at one or both ends.

"With a million residuals the Q-Q plot is not perfectly straight, so the Normal likelihood is useless."

Huge samples reveal tiny departures that may not matter. Ask how big the departure is in practical terms: how much more often do 3-sd days happen than the Normal predicts, and does that change your intervals or decisions?

"Use a normality test (Shapiro–Wilk, Anderson–Darling) instead of looking."

Use both, carefully. With small $n$ the tests have little power (they miss real heavy tails); with large $n$ they reject for trivial departures. The plot shows what kind of departure you have, which a single p-value cannot.

"The Q-Q plot shows that my residuals are Normal."

A Q-Q plot can only show that the residuals are consistent with a Normal at this sample size, or that they depart from it in a specific way. It cannot prove Normality.

Model answer: "The Q-Q plot of the residuals stays inside a simulation envelope for Normal data of the same size, so I see no evidence against a Normal likelihood. With more data or a different period that could change, so I would recheck it."

Normal data wobbles most at the ends and for small $n$. Max of 20 Normals: 95% between 0.96 and 3.02.

Envelope: simulate Normal datasets of the same $n$, sort, take 2.5%/97.5% per rank (exact: Beta$(i, n+1-i)$ order statistics).

Trap: about 5% of dots poke out by chance; look for runs at the ends.

Quick check: with 15 residuals, the top dot is at 2.6 while the line says 1.83. Is that evidence of a heavy right tail?

No. With $n = 15$ the largest of 15 Normal values varies a lot: its 95% range is roughly 0.78 to 2.93 (from $\Phi(x)^{15} = 0.025$ and $0.975$). A value of 2.6 is well inside that range. You would need more data, or several dots in a run, to claim a heavy tail.

Q-Q against other references: Student-t and friends core

The reference does not have to be the Normal. A Q-Q plot can compare your data with any distribution whose quantiles you can compute. It is like trying on shoes: if the Normal pinches at the toes (the tails), try a Student-t with a smaller $\nu$ (degrees of freedom, the knob that sets how heavy its tails are). When the S-shape disappears and the dots line up, that reference has tails like yours.

Three ways to say it:

  • Picture: the same residuals bend into an S against the Normal and straighten against a Student-t with the right $\nu$.
  • Numbers: if your residuals' 99% point is 4.54 (scale 1), the Normal says 2.33 (far off the line) but a Student-t with $\nu = 3$ says 4.54 (right on it).
  • Slogan: change the reference until the line is straight; that reference is your candidate likelihood.

Residuals with these quantiles (they come from a Student-t with $\nu = 3$ and scale 1):

level $p$0.750.950.99
residual quantile0.772.354.54
Normal quantile0.6741.6452.326
Student-t(3) quantile0.7652.3534.541
  1. Against the Normal, the ratio "residual ÷ reference" is $0.77/0.674 = 1.14$, then $2.35/1.645 = 1.43$, then $4.54/2.326 = 1.95$. The ratio (the local slope) keeps growing: the plot bends upward at the end.
  2. Against the Student-t(3), the ratios are $0.77/0.765 = 1.01$, $2.35/2.353 = 1.00$, $4.54/4.541 = 1.00$. Constant: a straight line with slope 1 (scale 1).
  3. So a Student-t with about 3 degrees of freedom describes these residuals; a Normal does not.

A Q-Q plot against any reference $F$ plots $\bigl(Q_F(p_i), x_{(i)}\bigr)$. For a location–scale family with a fixed shape (such as Student-t with a fixed $\nu$) the plot is straight when the data have that shape, with slope = scale and intercept = location.

  • The shape parameter must be chosen. Try a few values ($\nu$ = 3, 5, 10, 30), or pick the one that makes the plot straightest: the $\nu$ that maximizes the probability plot correlation coefficient (PPCC), the correlation between the theoretical and the sorted sample quantiles (scipy.stats.ppcc_max, ppcc_plot).
  • Other useful references: Exponential (waiting times), a Normal Q-Q of $\log x$ (that is, a log-normal check), Gamma, Uniform (for p-values or PIT values, forward-link Chapter 7.16).
  • For count data (Poisson, Negative Binomial), ordinary Q-Q plots show stairs. A common fix is randomized quantile residuals (Dunn and Smyth): turn each count into a continuous value that should be $N(0,1)$ if the count model is right, then use a Normal Q-Q plot.
Why do we need it?

Seeing an S against the Normal tells you the tails are too heavy, but not how heavy. Comparing with Student-t references of different $\nu$ tells you roughly what tail weight would fit, which is exactly the knob of a Student-t likelihood.

Where is it used?

Choosing the likelihood in your forecasting model and A/B framework (Normal vs Student-t), checking waiting times against an Exponential, checking p-values against a Uniform, and scipy.stats.probplot(x, dist=stats.t, sparams=(nu,)).

How is it used?

Draw the Normal Q-Q first. If it is S-shaped, draw Student-t Q-Q plots for a few $\nu$ (or run ppcc_max), note the $\nu$ that straightens the plot, then fit a Student-t likelihood and compare predictive performance.

The residuals are generated from the noise you pick (default: Student-t with ν = 3). With the Normal reference you see an S. Switch to the Student-t reference and drag ν: the dots straighten near the true ν. The right panel shows the straightness score (PPCC) for many ν; the purple mark is the best one. Try Normal noise too: the best ν is large, which means "use a Normal".

"The Student-t with the straightest Q-Q plot proves the noise is Student-t with that $\nu$."

It gives you a candidate. The best $\nu$ depends on a handful of extreme points, so it is noisy; confirm by fitting the model (possibly learning $\nu$ with a prior) and checking predictions.

"Fit the reference's $\nu$, location and scale to the data and the Q-Q plot is a fair test."

If you tune the reference to the same data, the plot will look better than it should (you fitted the test). Treat it as a description, not as a proof.

This is the syllabus question for both projects: is a Normal likelihood reasonable for my residuals, or would Student-t be better? In your forecasting model: fit with a Normal likelihood, compute residuals, draw the Normal Q-Q plot; if it is S-shaped, compare Student-t references (or refit with dist.StudentT(df, loc, scale)) and check whether forecast intervals improve. You can fix $\nu$ (for example 3–5) or give it a prior and learn it. In the A/B framework, the same check on a continuous metric within each variant tells you whether the Normal or the Student-t likelihood is the safer choice.

"We use a Student-t likelihood because it removes the outliers."

A Student-t does not remove anything. It assigns more probability to extreme residuals, so they exert less influence on the fit than under a Normal likelihood.

Model answer: "The Normal Q-Q plot of the residuals was S-shaped: more extreme days than a Normal allows. Under a Normal likelihood those days pull the trend and inflate σ. A Student-t treats them as plausible, so they have less influence, and the predictive intervals have more realistic tails. All the data stays in the model."

Q-Q against any reference: $(Q_F(p_i), x_{(i)})$. Against Student-t($\nu$): straight → tails like a t with that $\nu$; slope = scale.

PPCC = correlation of the Q-Q points; the $\nu$ with the highest PPCC is a candidate (scipy.stats.ppcc_max).

Say: "Student-t gives extreme residuals less influence", never "Student-t removes outliers".

Quick check: your residuals are S-shaped against the Normal, but almost straight against a Student-t with $\nu = 4$, with slope 6. What is the residual standard deviation implied by this t?

The slope is the t's scale, $\sigma = 6$. For $\nu = 4$ the standard deviation is $\sigma\sqrt{\nu/(\nu-2)} = 6\sqrt{4/2} = 6\sqrt{2} \approx 8.5$.

P-P plots, for contrast

A Q-Q plot compares values: "your 99% point is at 4.1, the model's is at 2.3". A P-P plot ("probability–probability plot") compares probabilities instead: for each sorted data point it asks "what fraction of the data lies below this point?" and "what fraction does the model say lies below it?". Both answers are between 0 and 1.

That squashing into [0, 1] is the catch. Far out in the tails, every probability is close to 0 or close to 1, so all the extreme points get crammed into the two corners where you cannot see them. A P-P plot is sensitive in the middle of the distribution and nearly blind in the tails, which is exactly the opposite of what you need for choosing a likelihood.

Three ways to say it:

  • Picture: a Q-Q plot stretches the tails out; a P-P plot squeezes them into the corners.
  • Numbers: a residual at 4 sd: on a Q-Q plot it sits at height 4 where the line says about 2.3 (obvious); on a P-P plot it sits at 0.99997 where the line says about 0.995 (invisible).
  • Slogan: Q-Q for the tails, P-P for the middle.

The five residuals from before, $-3, -1, 0, 1, 3$, with mean 0 and sd $\sqrt5 \approx 2.236$. The fitted model is $N(0, 5)$.

  1. Data probabilities ("fraction below"), as before: $p_i = 0.1, 0.3, 0.5, 0.7, 0.9$.
  2. Model probabilities $\Phi(x_{(i)}/2.236)$: $\Phi(-1.342) = 0.090$, $\Phi(-0.447) = 0.327$, $\Phi(0) = 0.5$, $0.673$, $0.910$.
  3. P-P points: $(0.090, 0.1)$, $(0.327, 0.3)$, $(0.5, 0.5)$, $(0.673, 0.7)$, $(0.910, 0.9)$: all within 0.03 of the diagonal $y = x$.
  4. Now imagine adding a residual at 4 sd. Its model probability is $\Phi(4) = 0.99997$, squeezed into the corner next to the other high points. On the Q-Q plot the same point would stand out at height 4.

A P-P plot of sorted data $x_{(1)}, \dots, x_{(n)}$ against a reference CDF $F$ plots

$$\bigl(\,F(x_{(i)}),\; p_i\,\bigr), \qquad p_i = \tfrac{i - 0.5}{n},$$

and compares them with the diagonal $y = x$. (Some tools put the two coordinates the other way round.) It is the Q-Q plot with the reference CDF applied to both axes. The parameters of $F$ (for example $\bar x$ and $s$) must be chosen first, because unlike the Q-Q plot it does not absorb shifts and stretches into the line.

The largest vertical gap between the empirical CDF and $F$ is the Kolmogorov–Smirnov statistic. Like the P-P plot, it is most sensitive in the middle and weak in the tails.

Why do we need it?

Mostly to know why you should prefer the Q-Q plot for tail questions, and to recognise P-P-style pictures elsewhere. When the middle of the distribution is what matters, a P-P plot shows it clearly.

Where is it used?

Some statistics packages offer it next to the Q-Q plot; the Kolmogorov–Smirnov test is built on the same comparison; and the probability-integral-transform (PIT) checks of forecast calibration (Chapter 7.16) use the same "apply the CDF" idea.

How is it used?

Fit the reference, compute F = stats.norm.cdf(np.sort(x), loc, scale) and plot it against (np.arange(1, n+1) - 0.5)/n with the diagonal. Use it for the middle; switch to a Q-Q plot for the tails.

The same 100 standardized values are shown as a Q-Q plot (left) and a P-P plot (right). The 5 smallest and 5 largest values are red. With Heavy tails the red dots shoot away from the line on the Q-Q plot but hide in the corners of the P-P plot. Try Right skew: the P-P plot does show the bend, because skew also changes the middle.

"P-P and Q-Q plots show the same thing, so either will do."

They contain the same points but stretch the axes differently. Tail problems (the ones that matter for Normal vs Student-t) are clear on a Q-Q plot and nearly invisible on a P-P plot.

"A Kolmogorov–Smirnov test passed, so the tails are fine."

KS measures the biggest gap between CDFs, which is usually in the middle. It is weak at detecting heavy tails. (Also, when the parameters are estimated from the same data, the standard KS p-value is wrong; the Lilliefors version corrects this.)

P-P plot: $(F(x_{(i)}),\ (i-0.5)/n)$ against the diagonal; probabilities on both axes.

Q-Q is better for tails (likelihood choice); P-P is better for the middle.

Trap: a passed KS test says little about the tails.

Quick check: which plot would you use to see whether your residuals have more 4-sd days than a Normal predicts, and why?

The Q-Q plot. It plots values, so a 4-sd residual appears at height 4, far above the line. On a P-P plot it becomes the probability 0.99997, packed into the corner with the other large values.

Q-Q plots of model residuals: is a Normal likelihood reasonable? core

Your raw daily demand is not Normal and is not supposed to be: it has a trend, a weekly pattern and holidays. What your model assumes is that the leftovers are Normal: after the model has explained trend, season, holidays and regressors, the part it could not explain (the residual, $e_t = y_t - \hat y_t$, actual minus forecast) should look like random Normal noise. So the Q-Q plot you care about is the Q-Q plot of the residuals.

One more subtlety. If the noise is Normal on every day but its size changes over time (calm months and wild months), then all the residuals thrown together form a mixture of narrow and wide bells, and a mixture like that has heavy tails. The Q-Q plot shows an S, yet the cure is to model the changing variance, not to switch to a Student-t. That is why the Q-Q plot is always read together with a plot of the residuals over time.

Three ways to say it:

  • Picture: peel off trend and season; what is left should give a straight Q-Q line.
  • Numbers: with 365 days of Normal noise you expect about $365 \times 0.0027 \approx 1$ day beyond ±3σ; if you see 9, the tails are heavy (or the variance changes).
  • Slogan: Q-Q the leftovers, not the raw data, and look at them over time too.

You fit your forecasting model to a year of data and standardize the residuals: $r_t = e_t / \hat\sigma$.

  1. Under a Normal, $P(|r| \gt 3) = 2 \times (1 - \Phi(3)) = 2 \times 0.00135 = 0.0027$.
  2. Expected count in 365 days: $365 \times 0.0027 \approx 0.99$, about one day.
  3. You count 9 days beyond ±3. That is about 9 times the Normal's prediction, and it matches the S-shape on the Q-Q plot.
  4. For comparison, a Student-t with $\nu = 4$ rescaled to sd 1 puts $P(|r| \gt 3) \approx 0.013$, so about $365 \times 0.013 \approx 4.8$ days: much closer.
  5. Before switching likelihoods, plot $r_t$ over time. If the 9 days are spread out and the spread looks constant, heavy tails are the story: try Student-t. If they cluster in a period with visibly wider spread, the variance is changing: model that instead. If they are all holidays, add holiday regressors.

The residual Q-Q check:

  1. Fit the model and compute residuals $e_t = y_t - \hat y_t$ (for a Bayesian model, use the posterior mean forecast; the fuller version of this check is a posterior predictive check, Chapter 7.14).
  2. Standardize: $r_t = e_t / \hat\sigma$, where $\hat\sigma$ is the estimated noise sd.
  3. Draw a Normal Q-Q plot of the $r_t$ with the line $y = z$ and a simulation envelope.
  4. Read it: straight → a Normal likelihood is reasonable; S → heavy tails; smile or frown → skew; lone dots → specific days to investigate.
  5. Cross-check with residuals vs time and vs fitted values (changing variance), and with their autocorrelation (the Q-Q plot cannot see time order; Chapter 7.17).

Why changing variance looks like heavy tails: a mixture of Normals with different standard deviations is heavier-tailed than any single Normal. In fact the Student-t is such a mixture: a Normal whose variance is itself random (drawn from an inverse-Gamma distribution).

Why do we need it?

The likelihood is an assumption about the residuals. Checking it is how you justify "Normal" or "Student-t" in your model, and how you catch missing structure (holidays, regime changes, changing variance) that shows up as odd residuals.

Where is it used?

Regression diagnostics (the second panel of R's plot(lm)), your forecasting model's residual check, A/B metric modelling, and any NumPyro model with a Normal or Student-t observation likelihood.

How is it used?

r = (y - yhat) / sigma_hat; sm.qqplot(r, line="45") (fine here because $r$ is standardized); plot r against time; count np.sum(np.abs(r) > 3) and compare with 0.0027 * len(r); then decide.

Fit the model Residuals eₜ = yₜ − ŷₜ, divided by σ̂ (standardized) Normal Q-Q plot + simulation envelope Straight, inside band S-shape: heavy tails Smile or frown: skew A few lone dots A Normal likelihood is reasonable Spread changing over time? Model it or transform. If not: try Student-t Log or Box-Cox (Chapter 4.18), or a skewed or count likelihood Look up the dates: holidays? errors? Add regressors or fix the data Q-Q plots ignore time order: also plot residuals over time and their ACF (Chapter 7.17).
The residual Q-Q check as a decision flow. Each pattern points to a different fix; for the S-shape, always rule out changing variance before switching to a Student-t likelihood.

A 30-week daily demand series (trend + weekly pattern + noise) is fitted with a linear trend and weekday effects. The top panel shows the residuals over time, the bottom panel their Normal Q-Q plot with a 95% envelope. Switch the true noise: compare Heavy-tailed noise with Noise grows over time. Both give an S on the Q-Q plot, but only the top panel tells them apart. With 4 unmodelled holidays, look for the lone dots.

"For a Normal likelihood, the target $y$ itself must be Normally distributed."

Only the noise around the model's prediction needs to be (approximately) Normal. Demand with a trend and weekly pattern is far from Normal as a whole, and that is fine.

"The residual Q-Q plot looks straight, so the residuals are fine."

A Q-Q plot ignores time order. Residuals with strong autocorrelation (today's error predicts tomorrow's) can look perfectly Normal on a Q-Q plot. Check residuals over time and their ACF too.

"An S-shape always means: switch to Student-t."

First check whether the spread changes over time or with the level of the series. If it does, a transform (Chapter 4.18) or an explicit variance model is the real fix; a Student-t would only hide it.

Both projects use this exact check. In your forecasting model: residuals from the posterior mean forecast, standardized by the estimated noise scale; a straight Q-Q plot supports the Normal likelihood, an S supports Student-t, and stairs plus a smile for low-volume series support the Negative Binomial (for counts, use randomized quantile residuals to get a clean picture). Holidays show up as lone dots until $h(t)$ includes them. In your A/B framework: a Q-Q plot of a continuous metric within each variant (or of residuals after a regression adjustment) tells you whether the Normal or the Student-t likelihood is the better description before you compute $P(\theta_B \gt \theta_A \mid D)$.

"I checked that demand is Normally distributed, so I used a Normal likelihood."

"I checked that the residuals (actual minus forecast) are approximately Normal, using a Q-Q plot with an envelope together with residuals over time."

Model answer: "The likelihood describes the noise after trend, seasonality, holidays and regressors. So I look at a Normal Q-Q plot of the standardized residuals. If the tails flare out and the spread is stable over time, I switch to a Student-t; if the spread grows with the level, I transform or model the variance; if a few dates stand out, I add regressors."

Q-Q the standardized residuals $r_t = (y_t - \hat y_t)/\hat\sigma$, not the raw $y$.

Normal: about $0.27\%$ beyond ±3 (≈ 1 day per year). Many more → heavy tails or changing variance.

Trap: always pair the Q-Q plot with residuals over time (variance changes, holidays, autocorrelation).

Quick check: 730 daily residuals, 14 of them beyond ±3σ̂, scattered evenly through time with a steady spread. What would you do?

A Normal expects $730 \times 0.0027 \approx 2$ such days; 14 is about 7 times more. With a steady spread and no clustering, this points to heavy tails, so try a Student-t likelihood (and check that its predictive intervals now cover about the right fraction of days).

Recap, cheat sheet and practice

  • A Q-Q plot pairs the sorted data $x_{(i)}$ with a reference distribution's quantiles $Q(p_i)$, $p_i = (i - 0.5)/n$. Straight line = same shape.
  • Shifting and stretching never bend the line: intercept ≈ centre, slope ≈ spread (the scale, which for a Student-t is not the sd). The plot checks shape only.
  • Read it in order: middle → ends → loners → jumps. Smile/frown = skew; S = heavy tails; reversed S = light tails; lone dots = outliers; step = two groups; stairs = counts or rounding.
  • Normal data wobbles, most at the ends and for small $n$. Judge bends against a simulation envelope (about 5% of dots poke out by chance).
  • Any distribution can be the reference. Against a Student-t, the $\nu$ that straightens the plot (highest PPCC) is a candidate tail weight.
  • P-P plots compare probabilities; they squash the tails into the corners, so use Q-Q plots for tail questions.
  • In your models, Q-Q the standardized residuals and always look at them over time too: an S can be heavy tails (→ Student-t) or changing variance (→ model it or transform).

Cheat sheet

IdeaFormula or ruleRemember
Q-Q point$(Q(p_i),\ x_{(i)})$, $p_i = (i - 0.5)/n$model across, sorted data up
Normal line$x \approx \mu + \sigma z$intercept = centre, slope = spread
Library positionsstatsmodels $i/(n+1)$; SciPy Filliben; R $a = 3/8$ or $1/2$only the extreme dots move
Right / left skewconvex smile / concave frownboth ends on the same side
Heavy / light tailsS / reversed Sends on opposite sides
Wobble of the top dot$P(\max \le x) = \Phi(x)^n$$n = 20$: 0.96 to 3.02 around 1.96
Exact envelope$i$-th of $n$ uniforms $\sim \text{Beta}(i, n+1-i)$pointwise; ~5% outside by chance
StraightnessPPCC $= \text{corr}(Q(p_i), x_{(i)})$ppcc_max picks a shape like $\nu$
Student-t slopescale $\sigma$; sd $= \sigma\sqrt{\nu/(\nu-2)}$slope is not the sd
Residual check$r_t = (y_t - \hat y_t)/\hat\sigma$; $P(|r| \gt 3) = 0.0027$≈ 1 day per year under a Normal
Code it · Python
import numpy as np
from scipy import stats
import matplotlib
matplotlib.use("Agg")                     # no window needed; remove this line in a notebook
import statsmodels.api as sm

# 1) A Q-Q plot by hand: sort, label, look up, compare
x = np.array([1, -3, 0, 3, -1.0])         # five daily forecast errors
n = len(x)
xs = np.sort(x)                            # order statistics
p = (np.arange(1, n + 1) - 0.5) / n        # plotting positions (i - 0.5)/n
z = stats.norm.ppf(p)                      # standard Normal quantiles
line = xs.mean() + xs.std(ddof=1) * z      # reference line  x-bar + s*z
print(p)                                   # [0.1 0.3 0.5 0.7 0.9]
print(z.round(2))                          # [-1.28 -0.52  0.    0.52  1.28]
print(line.round(2))                       # [-2.87 -1.17  0.    1.17  2.87]

# 2) SciPy does it in one call (Filliben positions + least-squares line)
(osm, osr), (slope, intercept, r) = stats.probplot(x, dist="norm")
print(osm.round(2))                        # [-1.13 -0.49  0.    0.49  1.13]  (other positions)
print(round(slope, 2), round(intercept, 2), round(r, 4))   # 2.56 0.0 0.9964

# 3) Normal vs heavy-tailed residuals (365 days each)
rng = np.random.default_rng(0)
e_norm = rng.normal(0, 6, 365)
e_t = 6 / np.sqrt(3) * rng.standard_t(3, 365)          # Student-t(3) scaled to sd 6
for name, e in [("normal", e_norm), ("t3", e_t)]:
    rstd = (e - e.mean()) / e.std(ddof=1)               # standardized residuals
    (_, _), (_, _, ppcc) = stats.probplot(rstd, dist="norm")
    print(name, "PPCC:", round(ppcc, 4), "| beyond 3 sd:", int(np.sum(np.abs(rstd) > 3)),
          "| expected under Normal:", round(365 * 2 * stats.norm.sf(3), 2))
# normal PPCC: 0.9974 | beyond 3 sd: 3 | expected under Normal: 0.99
# t3 PPCC: 0.9672 | beyond 3 sd: 6 | expected under Normal: 0.99

# 4) Which Student-t straightens the heavy-tailed residuals?
for nu in [2, 3, 5, 10, 30]:
    (_, _), (_, _, ppcc) = stats.probplot(e_t, dist=stats.t, sparams=(nu,))
    print("nu =", nu, "PPCC:", round(ppcc, 4))
# nu = 2: 0.9681 | nu = 3: 0.9965 | nu = 5: 0.9915 | nu = 10: 0.9802 | nu = 30: 0.9715
print("best nu:", round(stats.ppcc_max(e_t, brack=(2, 10), dist="t"), 2))   # best nu: 3.28

# 5) A simulation envelope for n = 365 (parametric bootstrap from N(mean, sd))
sims = np.sort(rng.normal(e_t.mean(), e_t.std(ddof=1), size=(2000, 365)), axis=1)
lo, hi = np.percentile(sims, [2.5, 97.5], axis=0)
xs_t = np.sort(e_t)
print("t3 dots outside the envelope:", int(np.sum((xs_t < lo) | (xs_t > hi))), "of 365")
# 203 of 365: with this many points both the flat middle and the flaring ends of the S leave the band

# 6) The usual plotting calls
fig1 = sm.qqplot(e_t, line="q")                          # quartile line; reference N(0, 1)
fig2 = sm.qqplot((e_t - e_t.mean()) / e_t.std(ddof=1), line="45")   # 45-degree line needs standardized data
fig3 = sm.qqplot(e_t, dist=stats.t, distargs=(3,), line="q")         # Student-t(3) reference
Test yourself

1. On the Normal Q-Q plots in this chapter (the SciPy and statsmodels convention), what is on the horizontal axis?

The theoretical quantiles $Q(p_i)$ go across and the sorted data $x_{(i)}$ go up. Time order is thrown away when you sort.

2. Both ends of a Normal Q-Q plot lie above the line and the middle lies below it. The data is…

Above–below–above is a convex "smile": the small values are less small than a Normal allows and the big values are bigger. That is a long right tail.

3. Residuals give a straight middle, the lowest dots far below the line and the highest dots far above it. What is the best next step?

Opposite flares form an S: heavy tails. Rule out changing variance first (it also makes an S); with a steady spread, a Student-t likelihood is the natural fix. Deleting real extremes makes the model overconfident.

4. You multiply every residual by 10 (you change units). What happens to the Normal Q-Q plot?

A stretch multiplies every sample quantile by 10, so the slope (the spread) is multiplied by 10. Shape, and so straightness, is unchanged.

5. With $n = 20$ residuals, the largest one sits at 2.8, while the line says 1.96 there. What should you conclude?

$P(\max \le x) = \Phi(x)^{20}$ gives a 95% range of 0.96 to 3.02. A single top dot at 2.8 is ordinary wobble for $n = 20$.

6. Which sentence about a Student-t likelihood is correct?

Nothing is removed: the t simply expects extremes, so they pull the fit less. Its Q-Q slope is the scale $\sigma$; the sd is $\sigma\sqrt{\nu/(\nu-2)}$.

Practice problems

A. Build the Normal Q-Q points and the line $\bar x + s z$ for the data 12, 7, 9, 15, 10. Is it roughly straight?

Sorted: 7, 9, 10, 12, 15. Positions $0.1, 0.3, 0.5, 0.7, 0.9$, so $z = -1.28, -0.52, 0, 0.52, 1.28$. Mean $\bar x = 53/5 = 10.6$. Deviations $-3.6, -1.6, -0.6, 1.4, 4.4$; squares $12.96 + 2.56 + 0.36 + 1.96 + 19.36 = 37.2$; $s^2 = 37.2/4 = 9.3$, $s \approx 3.05$. Line values: $10.6 + 3.05z = 6.69, 9.00, 10.6, 12.20, 14.51$. Data minus line: $0.31, 0.00, -0.60, -0.20, 0.49$. All within about 0.6 of the line: roughly straight. (Ends slightly above and middle slightly below hints at a mild right skew, but five points cannot tell.)

B. Waiting times follow an Exponential with rate 1. Compute the 0.1 and 0.9 quantiles and compare them with $N(1, 1)$. Which shape do you expect on a Normal Q-Q plot?

Exponential quantile $-\ln(1-p)$: $-\ln 0.9 = 0.105$ and $-\ln 0.1 = 2.303$. The Normal $1 + z_p$: $1 - 1.28 = -0.28$ and $1 + 1.28 = 2.28$. At the low end the data (0.105) is above the Normal (−0.28); in the middle the median 0.693 is below 1; at the far right (0.99: 4.61 vs 3.33) the data is above again. Above–below–above: a convex smile, i.e. right skew.

C. Residuals follow a Student-t with $\nu = 5$ and scale 2. What slope do you expect on a Q-Q plot against Student-t(5), and what is their standard deviation?

Against the matching Student-t(5), the plot is straight with slope = scale = 2. The standard deviation is $2\sqrt{5/3} \approx 2.58$. (On a Normal Q-Q plot you would see a mild S instead.)

D. For $n = 10$ truly Normal values, find the 95% range of the largest value and the $z$ at which it is plotted.

$P(\max \le x) = \Phi(x)^{10}$. Lower: $\Phi(x) = 0.025^{1/10} = 0.6915$, so $x \approx 0.50$. Upper: $\Phi(x) = 0.975^{1/10} = 0.99747$, so $x \approx 2.80$. It is plotted at $z = \Phi^{-1}(9.5/10) = \Phi^{-1}(0.95) = 1.645$. So with 10 points the top dot can be anywhere from 0.50 to 2.80: never over-read the ends of a small Q-Q plot.

E. Every day's noise is Normal, but half the days have sd 1 and half have sd 3. Show that the pooled residuals are heavy-tailed.

Pooled variance: $0.5 \times 1 + 0.5 \times 9 = 5$. Fourth moment (a Normal has $E[X^4] = 3\sigma^4$): $3(0.5 \times 1 + 0.5 \times 81) = 123$. Kurtosis $= 123/5^2 = 4.92$, excess kurtosis $= 1.92 \gt 0$: heavier tails than a Normal. Beyond 3 pooled sd ($3\sqrt5 \approx 6.71$) the probability is about 0.0127, almost 5 times the Normal's 0.0027. So a Q-Q plot of the pooled residuals shows an S, even though no single day is heavy-tailed. This is why you check the spread over time before choosing a Student-t.

F. Explain to an interviewer how you would decide between a Normal and a Student-t likelihood for your forecasting model.

"I fit the model with a Normal likelihood and compute standardized residuals: actual minus posterior mean forecast, divided by the noise scale. I draw a Normal Q-Q plot with a simulation envelope and count how many residuals are beyond ±3; a Normal expects about 0.27%. If the plot is straight, the Normal is fine. If it is S-shaped, I first plot the residuals over time: if the spread changes, I model the variance or transform; if the extremes are holidays, I add regressors. If the spread is steady and the extremes are scattered, the noise is heavy-tailed, so I refit with a Student-t, which gives those days less influence, and check that its prediction intervals cover the right fraction of held-out days."

Chapter 4.18 · Syllabus Module 8

Transformations: log, square root, Box-Cox, Yeo-Johnson, standardization

Sometimes data is hard to model on its natural scale: it is skewed, its spread grows with its level, or its effects multiply instead of adding. A transformation changes the ruler you measure with, so that the data becomes easier to describe. This chapter shows when each transformation helps, how to choose one, how to come back to the original scale without fooling yourself, and why your A/B framework uses one global scaler instead of scaling each group separately.

  • Say why we transform: reduce skew, stabilize variance, turn multiplication into addition, straighten relationships
  • Use the log transform (and know what to do with zeros: log1p and $\log(y + c)$)
  • Use the square root for counts, and the general rule for picking a variance-stabilizing transform
  • Understand the Box-Cox family, why $\lambda = 0$ means the log, and how $\lambda$ is chosen by maximum likelihood
  • Use Yeo-Johnson when the data has zeros or negative values
  • Avoid the back-transformation trap: the mean of the logs is not the log of the mean, and back-transformed forecasts estimate the median
  • Standardize correctly, and explain why per-group scaling can wipe out the group effect you want to estimate

Why transform? Changing the ruler core

Picture a map of a country's cities by population: a few cities have millions of people, most have a few thousand. On an ordinary (linear) ruler, all the small towns are squashed into one corner and a couple of giants fill the rest of the page. Now use a ruler where every step means "ten times bigger": 1 000, 10 000, 100 000, 1 000 000 are equally spaced. Suddenly every city is visible. You did not change the cities; you changed the ruler.

A transformation applies one fixed function to every value, like $\log y$ or $\sqrt y$. The good ones are monotone (they keep the order: bigger stays bigger), so nothing is lost and you can always go back. What changes is the spacing: a log or square root pulls big values in and spreads small values out.

Why would we want that? Four everyday reasons: the data is skewed (a long right tail), its spread grows with its level (big stores vary more than small ones), its effects are multiplicative ("+20%" rather than "+20 orders"), or a relationship is curved and becomes straight on another scale. Many models (linear regression, Normal likelihoods) work best when the noise is symmetric, has constant spread and effects add up.

Three ways to say it:

  • Picture: stretch the ruler at the small end and squeeze it at the big end until the data looks evenly spread.
  • Numbers: 1, 10, 100, 1000 have gaps 9, 90, 900; after $\log_{10}$ they become 0, 1, 2, 3 with gaps 1, 1, 1.
  • Slogan: a transformation changes the ruler, not the data's order.

Daily orders for four stores: 1, 10, 100 and 1000.

  1. On the raw ruler the gaps are $10 - 1 = 9$, $100 - 10 = 90$, $1000 - 100 = 900$. The first three stores are crammed into the first 10% of the axis.
  2. Take $\log_{10}$: $\log_{10} 1 = 0$, $\log_{10} 10 = 1$, $\log_{10} 100 = 2$, $\log_{10} 1000 = 3$. Equal gaps of 1: each step is "×10".
  3. Take the square root instead: $1, 3.16, 10, 31.6$, with gaps $2.16, 6.84, 21.6$. Still growing, but each gap is only about 3 times the previous one instead of 10 times: the square root is a milder squeeze than the log.
  4. The order 1 < 10 < 100 < 1000 is the same on every ruler, so the median store is still the same store.

A transformation is a function $g$ applied to every value: $y \mapsto g(y)$. Useful ones are strictly increasing (monotone), so they keep the order and have an inverse $g^{-1}$ to get back.

  • Quantiles pass straight through a monotone transform: the median of $g(Y)$ is $g(\text{median of } Y)$, and the same holds for every percentile.
  • Means do not: in general $E[g(Y)] \ne g(E[Y])$ (more in the back-transformation section below).
  • The ladder of powers (Tukey): $y^2,\ y,\ \sqrt y,\ \log y,\ 1/\sqrt y,\ 1/y$. Going down the ladder squeezes the big values more and more strongly, so it fixes stronger and stronger right skew. Box-Cox (later in this chapter) turns this ladder into one continuous knob $\lambda$.
  • A linear transform $a + by$ (like a z-score or a change of units) changes location and scale but never the shape; it cannot remove skew.
Why do we need it?

Many tools assume symmetric noise with constant spread and additive effects. Real data (money, counts, sizes) breaks all three. A transformation can make the data match the assumptions, so simple models work and their uncertainty statements become trustworthy.

Where is it used?

Log-revenue in A/B tests, log-demand in forecasting, square-root counts, Box-Cox in classical time-series modelling, sklearn.preprocessing.PowerTransformer and StandardScaler in ML pipelines, and log axes in every plot of skewed data.

How is it used?

Look at a histogram and a Q-Q plot (Chapter 4.17). If they show right skew or spread growing with level, try a log or square root, redraw, and keep the transform that gives the most symmetric, even-spread picture. Remember to transform back at the end.

linear ruler 1, 10 100 1000 log ruler 1 10 100 1000 each step on the log ruler = ×10
The same four values on a linear ruler (top: three of them are squashed together on the left) and on a log ruler (bottom: equally spaced, because each is ten times the previous one). The order never changes.

The top ruler is the raw scale from 1 to 100, with 60 skewed data points as a rug (blue). The bottom ruler is the same values after the transform chosen by the λ slider (λ = 1: no change; 0.5: square root; 0: log; −1: reciprocal). Slide λ down from 1 and watch the small values spread out and the crowd on the left move toward the middle. At λ = 0 the ticks 1, 10, 100 are equally spaced.

"Transforming the data changes the information in it."

A monotone transform keeps every value's rank and can be undone exactly. It changes distances and therefore averages, variances and how a model weighs each point, but not the order.

"Any transform that makes the histogram look nicer is fine."

A transform also changes what your model's parameters mean (a difference of logs is a ratio) and what a back-transformed prediction estimates (often the median). Choose it for a reason you can explain, not only for looks.

Both projects meet the reasons above. In your A/B framework, revenue or order-value metrics are right-skewed and their treatment effects are often multiplicative ("+3%"), so a log scale is natural. In your forecasting model, demand often has a spread that grows with its level and seasonal swings that scale with the trend; a log transform turns these into constant spread and additive seasonality, while the Negative Binomial likelihood is the alternative that models counts on their own scale.

Transform = one function applied to every value; monotone transforms keep the order and can be undone.

Reasons: skew, spread growing with level, multiplicative effects, curved relationships.

Ladder: $y^2, y, \sqrt y, \log y, 1/y$; lower = stronger squeeze. Linear transforms (z-scores) never change shape.

Quick check: the median of 5 shop sizes is 40 m². What is the median of their square roots? Can you say the same about the mean?

The square root is increasing, so it keeps the order: the median of the square roots is $\sqrt{40} \approx 6.32$. The mean does not pass through like this: the mean of the square roots is in general smaller than the square root of the mean (because the square root bends downward).

The log transform core

Many quantities grow by multiplying: a price rises 10%, a user base doubles, a promotion lifts sales by 20% in every store. The logarithm is the tool that turns multiplying into adding: $\log(a \times b) = \log a + \log b$. On the log scale, "+20% for every store" becomes "add the same 0.18 to every store", whether the store sells 10 or 1000.

The same property fixes two other problems at once. Data that is a product of many positive factors (customer size × visit frequency × basket size …) is right-skewed, and its log is roughly symmetric (this is the log-normal distribution of Chapter 4.10). And when the spread grows in proportion to the level (big stores wobble by ±10% just like small ones), the log turns that into a constant spread.

Three ways to say it:

  • Picture: the log ruler is a "times" ruler: equal steps mean equal ratios.
  • Numbers: +20% takes 10 to 12 and 100 to 120; the logs move by the same $\ln 1.2 \approx 0.18$ in both cases.
  • Slogan: logs turn ratios into differences, skew into symmetry, and proportional spread into constant spread.

A promotion lifts sales by 20% everywhere. Store A sells 10 a day, store B sells 100.

  1. Raw effects: A goes $10 \to 12$ (+2), B goes $100 \to 120$ (+20). Same promotion, effects 10 times apart: a model with one additive "promotion effect" cannot describe both.
  2. Natural logs: $\ln 10 = 2.303 \to \ln 12 = 2.485$, a change of $0.182$. $\ln 100 = 4.605 \to \ln 120 = 4.787$, also $0.182$.
  3. That common change is $\ln 1.2 = 0.182$. One additive effect on the log scale describes both stores.
  4. To read it back: $e^{0.182} = 1.2$, i.e. +20%. For small effects the log change is close to the percentage itself: $\ln 1.05 = 0.049 \approx 5\%$.
  5. Spread: small stores vary by about ±10 on a level of 100, big ones by ±100 on 1000. On the log scale both vary by about $\pm 0.1$, because $\text{sd}(\ln Y) \approx \text{sd}(Y)/\text{mean}(Y)$, the coefficient of variation.

The log transform is $y \mapsto \ln y$ (natural log; the base only rescales the result: $\log_{10} y = \ln y / \ln 10$). It needs $y \gt 0$.

  • Multiplicative → additive: $\ln(ab) = \ln a + \ln b$, so a model $y = a \cdot b \cdot \varepsilon$ becomes $\ln y = \ln a + \ln b + \ln \varepsilon$.
  • Differences are ratios: $\ln y_2 - \ln y_1 = \ln(y_2/y_1)$. A coefficient $\beta$ on the log scale means "multiply by $e^\beta$"; for small $\beta$, about $100\beta\%$.
  • Variance stabilizing when $\text{sd}(Y) \propto \text{mean}(Y)$: by a first-order (delta-method) approximation, $\text{sd}(\ln Y) \approx \text{sd}(Y)/\mu$, the coefficient of variation.
  • Reduces right skew: if $Y$ is log-normal, $\ln Y$ is exactly Normal.
Why do we need it?

Money, sizes, durations and demand are positive, right-skewed and change by percentages. Without the log, a few huge values dominate averages and fits, effects differ wildly between big and small units, and Normal-noise models are badly wrong.

Where is it used?

Log-revenue and log-order-value metrics in A/B tests, log-demand and multiplicative seasonality in forecasting (Prophet's multiplicative mode is related), log-price models in economics, log axes in plots, and log links in Poisson and Negative Binomial GLMs.

How is it used?

np.log(y) for strictly positive data; fit the model on the log scale; read coefficients as percentage changes ($e^\beta - 1$); transform predictions back with care (see the back-transformation section).

400 simulated users' revenue (median 20 dollars). Left: the raw histogram, piled up on the left with a long right tail. Right: the histogram of $\ln(\text{revenue})$, a symmetric bell. Raise σ (spread on the log scale): the raw data becomes far more skewed (the mean runs away from the median), while the log histogram just gets wider and stays symmetric. Watch the skewness numbers.

Weekly demand grows by a fixed percentage per week, with noise that is a percentage too. On the raw scale (left) the curve bends upward and the noise band fans out. On the log scale (right) the same data is a straight line with a constant-width band. Change the growth rate and the noise and compare the week-to-week changes in the readout.

"Log-transforming makes any data Normal."

It makes log-normal data Normal and helps with right skew in general. Left-skewed or already-symmetric data gets worse, and data with many zeros cannot be logged at all without extra choices (next section).

"A coefficient of 0.4 on the log scale means +40%."

It means "multiply by $e^{0.4} = 1.49$", i.e. +49%. The "log change ≈ percentage" shortcut only works for small values (below about 0.1).

In your forecasting model, if weekly seasonality and holiday lifts are percentages of the current level, modelling $\ln(\text{demand})$ with the additive structure $g(t) + s(t) + h(t) + X_t\beta$ is the same as a multiplicative model on the raw scale, and the regressor coefficients become percentage effects. In your A/B framework, analysing log-revenue turns a "+3% lift" into an additive shift, which is often more stable across segments of very different size than an additive lift in dollars.

$\ln(ab) = \ln a + \ln b$: ratios become differences; $\beta$ on the log scale = ×$e^\beta$ (≈ $100\beta\%$ when small).

sd$(\ln Y) \approx$ sd$(Y)/\mu$: the log stabilizes spread that is proportional to the level. Log-normal → Normal.

Trap: needs $y \gt 0$; does not fix left skew.

Quick check: on the log scale a treatment coefficient is $\beta = 0.10$. What is the effect on the original scale?

Multiply by $e^{0.10} = 1.105$: about +10.5%. (For such a small $\beta$, "about +10%" is a fine shortcut.)

Zeros: log1p and $\log(y + c)$

Counts and money often contain zeros: days with no sales of a slow item, users who bought nothing. The log of zero is $-\infty$, so a plain log fails. The common fix is to add a small constant first: $\log(y + c)$. With $c = 1$ this is $\log(1 + y)$, available as np.log1p ("log of one plus").

But the constant is a real choice, not a technicality. With a tiny $c$ the zeros land far away from everything else (log of a tiny number is very negative), creating an artificial gap. With a huge $c$ the transform barely bends anything. The data has not changed; only your ruler has.

Three ways to say it:

  • Picture: $c$ sets where the zeros sit on the log ruler: far out on the left (small $c$) or right next to the ones (large $c$).
  • Numbers: with $c = 0.01$ the step from 0 to 1 is 6.7 times the step from 1 to 2; with $c = 1$ it is 1.7 times; with $c = 100$ the two steps are almost equal.
  • Slogan: the "+c" is a modelling decision; choose it on purpose, or use a model that handles zeros.

Daily units sold: 0, 1, 3, 9, 99.

  1. Plain log: $\ln 0 = -\infty$. Not usable.
  2. log1p: $\ln 1 = 0$, $\ln 2 = 0.693$, $\ln 4 = 1.386$, $\ln 10 = 2.303$, $\ln 100 = 4.605$. Zeros stay at 0 and big values are pulled in.
  3. With $c = 0.01$: $\ln 0.01 = -4.61$ for the zero but $\ln 1.01 = 0.01$ for the one: the zero sits 4.6 units below the one, farther than the distance from 1 to 99 ($\ln 99.01 - \ln 1.01 = 4.59$).
  4. Ratio of the steps "0 → 1" and "1 → 2": $\ln\frac{1+c}{c} \big/ \ln\frac{2+c}{1+c}$. For $c = 0.01$: $4.615/0.688 = 6.71$; for $c = 1$: $0.693/0.405 = 1.71$; for $c = 100$: about 1.01.
  5. Units matter too: log1p of revenue in cents is not a rescaled log1p of revenue in dollars, because the "+1" means 1 cent in one case and 1 dollar in the other.
  • $\text{log1p}(y) = \ln(1 + y)$, defined for $y \gt -1$; numerically accurate for tiny $y$ (np.log1p). Its inverse is np.expm1: $e^z - 1$.
  • $\log(y + c)$ for a chosen $c \gt 0$ (sometimes called a "started log"). The result depends on $c$ and on the units of $y$.
  • Alternatives that avoid the choice: the square root (fine at 0, next section), Yeo-Johnson (later), the inverse hyperbolic sine $\text{asinh}(y) = \ln(y + \sqrt{y^2+1})$, or, often best, a model for the original data: a Poisson or Negative Binomial likelihood with a log link, or a two-part model (probability of zero × size when non-zero).
Why do we need it?

Real metrics have zeros, and dropping those rows biases everything (the zero days are real days). You need a transform that keeps them, and you need to know that the choice of $c$ can change your conclusions.

Where is it used?

np.log1p for count and revenue features in ML pipelines, log1p targets in demand forecasting competitions, "log(y + 1)" outcomes in A/B analyses of revenue per user, and asinh in economics.

How is it used?

Use np.log1p(y) and np.expm1(z) to go back. Check that results do not swing when you try another $c$ (for example 0.1 and 10). If the zeros are a large share of the data, model them directly instead.

300 days of sales of a slow item (about a third of the days are zero). The orange ticks show where the values 0, 1, 2, 5, 10, 20 land after $\log(y + c)$, rescaled so that 0 is at the left and 20 at the right. Drag c down to 0.01: the zeros fly away from the rest and form their own island. Push it up to 100: the transform becomes almost a straight line and does nothing. Compare the skewness numbers.

"Add 1 and take the log; the +1 does not matter."

For small values it matters a lot: it decides how far the zeros sit from the ones, and it depends on the units (cents vs dollars). Check that conclusions survive a different $c$.

"Drop the zeros so the log works."

Zero days are real outcomes. Dropping them answers a different question ("sales on days with at least one sale") and biases the average upward.

Low-volume series in your forecasting model have many zero days. Instead of log1p plus a Normal likelihood, your model's Negative Binomial likelihood handles the zeros and the skew on the original count scale, with a log link keeping the mean positive (Chapter 4.8). In the A/B framework, revenue per user with many zeros is the classic case where a two-part view (conversion × spend given conversion) is clearer than log1p.

np.log1p(y) $= \ln(1+y)$, inverse np.expm1. $\log(y+c)$: the result depends on $c$ and on units.

Small $c$ pushes zeros far away (artificial gap); large $c$ barely transforms.

Trap: never drop zeros to make a log work; consider a count or two-part model.

Quick check: why is np.log1p(1e-15) better than np.log(1 + 1e-15)?

$1 + 10^{-15}$ is rounded in floating point before the log is taken, and most of the tiny number's digits are lost: np.log(1 + 1e-15) returns $1.11 \times 10^{-15}$, 11% too big. np.log1p(1e-15) computes $\ln(1+y)$ directly and returns $1.0 \times 10^{-15}$. Its partner expm1 does the same for $e^z - 1$.

The square-root transform for counts

Counts of independent events (orders per hour, clicks per page, arrivals per day) behave roughly like a Poisson: the variance equals the mean (Chapter 4.8). A quiet shop with 4 orders an hour wobbles by about ±2; a busy one with 100 an hour wobbles by about ±10. If you want to compare or model them together, the busy shop's noise drowns everything.

Take square roots and the wobble evens out: $\sqrt 4 = 2$ wobbles by about ±0.5, and $\sqrt{100} = 10$ also wobbles by about ±0.5. The square root is gentler than the log, and it works at zero ($\sqrt 0 = 0$), which makes it a natural first choice for counts.

Three ways to say it:

  • Picture: the square root squeezes busy counts just enough that every count gets the same wobble.
  • Numbers: Poisson with mean 4: sd 2; with mean 100: sd 10. After $\sqrt{\ }$: both about 0.5.
  • Slogan: for Poisson counts, the square root makes the noise level the same everywhere.

Why 0.5? Use the first-order "delta method" approximation: near the mean $\lambda$, $\sqrt y$ behaves like a straight line with slope $\frac{d}{dy}\sqrt y = \frac{1}{2\sqrt\lambda}$, and a straight line multiplies the sd by its slope.

  1. Quiet shop, $\lambda = 4$: raw sd $= \sqrt 4 = 2$. Slope $= 1/(2\sqrt 4) = 1/4$. sd of $\sqrt Y \approx 2 \times 1/4 = 0.5$.
  2. Busy shop, $\lambda = 100$: raw sd $= 10$. Slope $= 1/(2 \times 10) = 1/20$. sd of $\sqrt Y \approx 10/20 = 0.5$.
  3. In general: $\sqrt\lambda \times \frac{1}{2\sqrt\lambda} = \frac12$, whatever $\lambda$ is.
  4. Exact values (summing over the Poisson probabilities): 0.553 for $\lambda = 4$ and 0.501 for $\lambda = 100$. The approximation is good once $\lambda$ is about 10 or more.
  5. Anscombe's version $\sqrt{y + 3/8}$ is even flatter: 0.500 already at $\lambda = 4$.
  • Square-root transform: $y \mapsto \sqrt y$, defined for $y \ge 0$ (zeros are fine).
  • If $Y \sim \text{Poisson}(\lambda)$, then $\text{Var}(\sqrt Y) \approx \tfrac14$ (sd $\approx 0.5$) for moderate and large $\lambda$.
  • Anscombe transform: $\sqrt{y + 3/8}$, whose variance is closer to $\tfrac14$ for smaller $\lambda$. (Freeman–Tukey, $\sqrt y + \sqrt{y+1}$, is a similar variant with variance about 1.)
  • For very small means ($\lambda \lesssim 2$) no simple transform stabilizes the variance well; the data is too discrete.
  • It is a milder squeeze than the log: for Poisson data the log over-corrects (sd of $\ln Y \approx 1/\sqrt\lambda$, which keeps shrinking).
Why do we need it?

Comparing or regressing counts of very different sizes with a constant-variance model gives the busy units all the weight and makes the quiet ones look noiseless. The square root puts all counts on an equal-noise footing while keeping zeros.

Where is it used?

Count features in ML pipelines, rootograms (square-root histograms) for checking count models, square-root counts in classical ANOVA on counts, image photon counts (Anscombe), and quick plots of hourly or daily order counts.

How is it used?

np.sqrt(y) or np.sqrt(y + 3/8); fit or plot on that scale; square back for predictions (with the same back-transformation care as for the log). For serious count modelling, prefer a Poisson or Negative Binomial likelihood.

The curves show the exact standard deviation of a Poisson count after each transform, for means $\lambda$ from 0.5 to 100 (log scale across). The dashed line is the target 0.5. Drag λ and read the values. Notice: $\sqrt{y}$ (blue) and Anscombe (green) settle at 0.5, while $\ln(1+y)$ (orange) keeps falling: the log over-corrects counts. Raw sd is $\sqrt\lambda$, which keeps growing (see the readout).

"Counts should always be log-transformed."

For Poisson-like counts the square root is the variance-stabilizing choice; the log over-corrects and cannot handle zeros without a "+c". The log is right when the spread grows in proportion to the mean (next section), which happens for overdispersed counts at high volumes.

"After the square root, small counts are fine too."

Below a mean of about 2, counts are too lumpy (mostly 0, 1, 2) for any transform to make them look continuous and Normal. Use a count likelihood.

In your A/B framework, count metrics (for example orders per user) can use a Poisson likelihood rather than a transform, which keeps effects on the natural scale. The square-root idea is still useful there for plots and quick checks, for example a "rootogram" comparing observed and predicted count frequencies on a square-root axis, so that small frequencies stay visible.

Poisson: Var = mean. $\text{sd}(\sqrt Y) \approx \sqrt\lambda \cdot \frac{1}{2\sqrt\lambda} = 0.5$ (good for $\lambda \gtrsim 10$; Anscombe $\sqrt{y+3/8}$ from about 4).

Works at zero; milder than the log.

Trap: log over-corrects Poisson counts; nothing fixes very small means.

Quick check: a Poisson count has mean 25. What are the approximate sds of $Y$ and of $\sqrt Y$?

sd$(Y) = \sqrt{25} = 5$. Slope of $\sqrt y$ at 25 is $1/(2 \times 5) = 0.1$, so sd$(\sqrt Y) \approx 5 \times 0.1 = 0.5$.

Variance stabilization: match the transform to how the spread grows

The log and the square root are two answers to one question: how fast does the spread grow with the level? If the spread grows like the square root of the level (Poisson counts), the square root evens it out. If the spread grows in proportion to the level (±10% everywhere), the log evens it out. If it grows even faster, you need an even stronger squeeze, like $1/y$.

You can measure the growth from your data: split it into groups (stores, weeks, segments), compute each group's mean and sd, and see how the sd scales with the mean.

Three ways to say it:

  • Picture: the faster the spread grows with the level, the further down the ladder of powers you go.
  • Numbers: means 9, 36, 144 with sds 3, 6, 12: the sd doubles each time the mean quadruples, so sd ∝ √mean, so use $\sqrt y$.
  • Slogan: sd ∝ mean$^k$ → use $y^{1-k}$ (and the log when $k = 1$).

Three store groups have mean daily orders 9, 36, 144 and standard deviations 3, 6, 12.

  1. How does sd grow with the mean? From 9 to 36 the mean is ×4 and the sd is ×2. So sd $\propto$ mean$^k$ with $4^k = 2$, i.e. $k = \ln 2 / \ln 4 = 0.5$.
  2. The rule says: use $y^{1-k} = y^{0.5}$, the square root.
  3. Check with the delta method, sd$(\sqrt Y) \approx$ sd$(Y)/(2\sqrt\mu)$: $3/(2 \times 3) = 0.5$, $6/(2 \times 6) = 0.5$, $12/(2 \times 12) = 0.5$. Equal: stabilized.
  4. Try the log instead, sd$(\ln Y) \approx$ sd$(Y)/\mu$: $3/9 = 0.33$, $6/36 = 0.17$, $12/144 = 0.083$. Now the spread shrinks with the level: the log over-corrects here.

Delta method (first-order approximation): for a smooth $g$ and $Y$ with mean $\mu$ and small relative spread,

$$\text{Var}\bigl(g(Y)\bigr) \approx g'(\mu)^2\, \text{Var}(Y).$$

If $\text{sd}(Y) = c\,\mu^k$, we want $|g'(\mu)|\, c\,\mu^k$ to be constant, so $g'(\mu) \propto \mu^{-k}$. Integrating:

$$g(y) = \begin{cases} y^{\,1-k} & k \ne 1 \\ \ln y & k = 1. \end{cases}$$
spread grows like$k$stabilizing transformtypical data
constant0noneNormal noise
$\sqrt{\text{mean}}$½$\sqrt y$Poisson counts
mean1$\ln y$constant coefficient of variation; multiplicative noise
mean$^{3/2}$3/2$1/\sqrt y$(inverse Gaussian-like data)
mean²2$1/y$very strongly growing spread

To estimate $k$, plot $\ln(\text{sd})$ against $\ln(\text{mean})$ across groups (a "spread–level plot"); the slope is about $k$, and the matching Box-Cox $\lambda$ (next section) is about $1 - k$.

Why do we need it?

Models with one noise level (a single σ) assume the spread is the same everywhere. When it is not (heteroscedasticity), standard errors and prediction intervals are wrong: too wide for quiet units, too narrow for busy ones.

Where is it used?

Choosing between square root and log for counts and sales, the spread–level plot of exploratory data analysis, classical ANOVA on transformed data, and deciding whether a forecasting model needs a log scale or a likelihood whose variance grows with the mean (Poisson, Negative Binomial).

How is it used?

Group the data, compute means and sds, fit the slope of log-sd on log-mean, pick $y^{1-k}$ (or the log), and check that the group sds are now similar. Or skip the transform and use a likelihood with the right mean–variance relation.

Four groups with means 5, 20, 80 and 320 (40 values each), shown raw, after $\sqrt{\ }$ and after $\ln$. The bars are group mean ± sd. Choose the data type. With Poisson counts the square-root panel has equal bars, while the log panel's bars shrink. With constant CV (sd = 30% of the mean) the log panel wins. The readout gives each panel's largest sd ÷ smallest sd (1 = perfectly even). (The rare zero counts are set to 0.5 before taking the log.)

"The log is the universal variance stabilizer."

Only when the spread is proportional to the level. For Poisson-like counts it over-corrects; for spread that grows faster than the level it under-corrects. Measure how the spread grows first.

"The delta method gives the exact variance."

It is a first-order approximation that is good when the spread is small relative to the mean. For small counts or very wide data, check with a simulation or exact calculation (as in the previous widget).

This is the reasoning behind the likelihoods in your forecasting model. A Normal likelihood assumes constant spread ($k = 0$). Poisson assumes Var = mean ($k = ½$). The Negative Binomial, Var $= \mu + \mu^2/\alpha$, behaves like Poisson at low volume and like "spread ∝ mean" ($k = 1$) at high volume. So instead of choosing a transform, you can choose the likelihood whose mean–variance relation matches what the spread–level plot shows.

Delta method: $\text{Var}(g(Y)) \approx g'(\mu)^2 \text{Var}(Y)$.

sd ∝ mean$^k$ → use $y^{1-k}$; $k = ½$ → $\sqrt y$; $k = 1$ → $\ln y$; $k = 2$ → $1/y$. Estimate $k$ as the slope of ln sd vs ln mean.

Trap: the wrong transform can over-correct (log on Poisson counts).

Quick check: across products, a mean of 10 goes with sd 2 and a mean of 1000 goes with sd 200. Which transform?

The mean is ×100 and the sd is ×100, so $k = 1$: the spread is proportional to the level (constant CV of 20%). Use the log; on the log scale both have sd about $0.2$.

Box-Cox: one family of power transforms core

Instead of choosing between "square root" and "log" and "reciprocal" by hand, Box and Cox put the whole ladder of powers on one dial, $\lambda$ (lambda). Turn the dial to 1 and nothing changes shape; to 0.5 and you get a square-root shape; to 0 and you get exactly the log; to −1 and you get a reciprocal shape. Every value in between is allowed too.

The formula has a small twist, "subtract 1 and divide by $\lambda$", which looks odd but does two good things: it keeps the order of the data even for negative $\lambda$ (so bigger stays bigger), and it makes the dial smooth, so that as $\lambda$ approaches 0 the curve slides exactly into the log.

Three ways to say it:

  • Picture: one dial that runs smoothly down the ladder of powers, with the log sitting at 0.
  • Numbers: $y = 1, 4, 9, 16$ become $0, 2, 4, 6$ at $\lambda = 0.5$ (evenly spaced) and $0, 1.39, 2.20, 2.77$ at $\lambda = 0$.
  • Slogan: Box-Cox = power transform with a smooth dial, log at zero, positive data only.

Apply Box-Cox to $y = 1, 4, 9, 16$ at four settings of the dial.

$\lambda$formula$y=1$$y=4$$y=9$$y=16$
1$y - 1$03815
0.5$2(\sqrt y - 1)$0246
0$\ln y$01.3862.1972.773
−1$1 - 1/y$00.750.8890.938
  1. $\lambda = 1$: $(y^1 - 1)/1 = y - 1$. Only a shift: the shape is unchanged.
  2. $\lambda = 0.5$: $(\sqrt y - 1)/0.5$. Since these $y$ are squares, the results are evenly spaced.
  3. $\lambda \to 0$: try $\lambda = 0.01$ and $y = 4$: $4^{0.01} = e^{0.01 \ln 4} = e^{0.01386} = 1.01396$, so $(1.01396 - 1)/0.01 = 1.396$, already close to $\ln 4 = 1.386$.
  4. $\lambda = -1$: $(y^{-1} - 1)/(-1) = 1 - 1/y$. The plain reciprocal $1/y$ would reverse the order (1, 0.25, 0.11, 0.06); dividing by the negative $\lambda$ flips it back.

For strictly positive data $y \gt 0$, the Box-Cox transform with parameter $\lambda$ is

$$y^{(\lambda)} = \begin{cases} \dfrac{y^\lambda - 1}{\lambda} & \lambda \ne 0, \\[2mm] \ln y & \lambda = 0. \end{cases}$$
  • Why $\lambda = 0$ gives the log: $y^\lambda = e^{\lambda \ln y} = 1 + \lambda \ln y + \tfrac{(\lambda \ln y)^2}{2} + \dots$, so $\dfrac{y^\lambda - 1}{\lambda} = \ln y + \tfrac{\lambda (\ln y)^2}{2} + \dots \to \ln y$ as $\lambda \to 0$.
  • Every member is increasing in $y$ and passes through $(1, 0)$ with slope 1, so the curves are easy to compare.
  • $\lambda \lt 1$ squeezes large values (reduces right skew); $\lambda \gt 1$ stretches them (reduces left skew).
  • Needs $y \gt 0$: $\ln 0$ is undefined and negative numbers have no real fractional powers. A shift $y + c$ makes it work but adds an arbitrary choice; Yeo-Johnson (below) is the cleaner fix.
Why do we need it?

The right amount of squeezing is rarely exactly "square root" or exactly "log". One continuous dial lets the data choose the strength (next section), and the named transforms are just special points on it.

Where is it used?

scipy.stats.boxcox, sklearn.preprocessing.PowerTransformer(method="box-cox"), Box-Cox regression (R's MASS::boxcox), classical time-series modelling (forecasting libraries offer a Box-Cox option), and your syllabus' P0 list.

How is it used?

Make sure all values are positive, call yt, lam = scipy.stats.boxcox(y), model yt, and invert predictions with scipy.special.inv_boxcox(z, lam). Prefer a rounded, explainable $\lambda$ (0, 0.5, 1) when the data allows it.

−4 −2 0 2 4 0 1 2 3 all pass through (1, 0) with slope 1 λ = 2 (square-like) λ = 1 (no change in shape) λ = 0.5 (square-root-like) λ = 0 (log) λ = −1 (reciprocal-like) original value y (must be > 0) transformed value y(λ)
The Box-Cox family. All curves pass through (1, 0) with slope 1. Lower λ bends the curve more: large values are squeezed harder and small values are spread out more.

Drag λ. The blue curve is the Box-Cox transform, the dashed purple curve is $\ln y$. Bring λ slowly toward 0: the blue curve slides onto the log, with no jump. The readout follows one value, $y = 4$, and compares it with $\ln 4 = 1.386$. Go below 0 and notice the curve is still increasing (order kept).

"Box-Cox with $\lambda = -1$ is $1/y$."

It is $1 - 1/y$: the same shape as $1/y$ but increasing, so the order is kept. With a raw $1/y$, your largest value becomes your smallest, and every coefficient changes sign.

"Shift the data by +1 and Box-Cox works fine with zeros."

It runs, but the shift is arbitrary and changes the chosen $\lambda$ and the result (just like $c$ in $\log(y + c)$). Use Yeo-Johnson, or a model that handles zeros.

$y^{(\lambda)} = (y^\lambda - 1)/\lambda$, and $\ln y$ at $\lambda = 0$ (the limit). Needs $y \gt 0$.

$\lambda = 1$ shift only; 0.5 √-shape; 0 log; −1 reciprocal shape (order kept).

Trap: "−1 then ÷λ" is what keeps the order and makes the dial smooth; zeros need Yeo-Johnson.

Quick check: compute the Box-Cox transform of $y = 9$ for $\lambda = 0.5$ and for $\lambda = 2$.

$\lambda = 0.5$: $(9^{0.5} - 1)/0.5 = (3 - 1)/0.5 = 4$. $\lambda = 2$: $(81 - 1)/2 = 40$.

Choosing $\lambda$ by maximum likelihood core

Which $\lambda$ should you use? Box and Cox asked: for which $\lambda$ does the transformed data look most like a Normal with one constant spread? They score every $\lambda$ with a log-likelihood (a measure of how well a Normal model fits the transformed data) and pick the $\lambda$ with the highest score: the maximum likelihood estimate $\hat\lambda$ (MLE; you will meet MLE in general in Chapter 5.2).

One trap is built in. You might think "just pick the $\lambda$ that makes the transformed values least spread out". But each $\lambda$ changes the units: $\lambda = -1$ squeezes everything into the interval from 0 to 1, so its variance is tiny no matter what. The Box-Cox score adds a correction term, the Jacobian (a measure of how much the transform stretched or squeezed the data), that charges for squeezing, so every $\lambda$ is judged fairly.

Three ways to say it:

  • Picture: a curve of "how Normal does it look" over the dial; pick its peak.
  • Numbers: for $y = 1, 2, 4, 8, 16$ the score is −8.48 at $\lambda = 1$, −7.27 at 0.5, −6.83 at 0 (the peak) and −7.27 at −0.5.
  • Slogan: best $\lambda$ = highest profile log-likelihood, with a fair correction for squeezing.

Data $y = 1, 2, 4, 8, 16$ (each value doubles: a geometric series), $n = 5$, $\sum \ln y = \ln(1 \cdot 2 \cdot 4 \cdot 8 \cdot 16) = \ln 1024 = 6.931$. The score is $\ell(\lambda) = -\tfrac n2 \ln \hat\sigma^2(\lambda) + (\lambda - 1)\sum \ln y_i$, where $\hat\sigma^2(\lambda)$ is the variance (dividing by $n$) of the transformed values.

$\lambda$transformed values$\hat\sigma^2(\lambda)$$-\tfrac52 \ln \hat\sigma^2$$(\lambda - 1) \times 6.931$$\ell(\lambda)$
10, 1, 3, 7, 1529.76−8.4830−8.483
0.50, 0.83, 2, 3.66, 64.577−3.802−3.466−7.268
00, 0.69, 1.39, 2.08, 2.770.961+0.100−6.931−6.832
−0.50, 0.59, 1, 1.29, 1.50.286+3.129−10.397−7.268
  1. Look at the variance column alone: it keeps falling as $\lambda$ goes down (29.76 → 0.286). Picking the smallest variance would always push $\lambda$ as low as possible.
  2. The Jacobian column $(\lambda - 1)\sum\ln y$ falls too, and it exactly charges for that squeezing.
  3. Adding the two columns gives a curve with a peak at $\lambda = 0$: $\hat\lambda = 0$, the log. That makes sense: on the log scale this data is $0, 0.69, 1.39, 2.08, 2.77$, evenly spaced and symmetric.
  4. How sure are we? Values of $\lambda$ whose score is within $1.92$ of the peak form an approximate 95% interval. Here SciPy gives $(-1.08, 1.08)$: with only 5 points almost any $\lambda$ is plausible. More data narrows it.

Assume the transformed data are Normal: $y_i^{(\lambda)} \sim N(\mu, \sigma^2)$, independent. Writing the likelihood in terms of the original $y$ needs the Jacobian $\prod_i \frac{d y_i^{(\lambda)}}{d y_i} = \prod_i y_i^{\lambda - 1}$. Plugging in the best $\mu$ and $\sigma^2$ for each $\lambda$ gives the profile log-likelihood

$$\ell(\lambda) = -\frac{n}{2} \ln \hat\sigma^2(\lambda) + (\lambda - 1) \sum_{i=1}^{n} \ln y_i \quad (+\text{ a constant}),$$

where $\hat\sigma^2(\lambda) = \frac1n \sum_i \bigl(y_i^{(\lambda)} - \overline{y^{(\lambda)}}\bigr)^2$. This is exactly scipy.stats.boxcox_llf.

  • $\hat\lambda = \arg\max_\lambda \ell(\lambda)$ (scipy.stats.boxcox(y) returns it, boxcox_normmax offers other criteria).
  • Approximate 95% interval: $\{\lambda : \ell(\lambda) \ge \ell(\hat\lambda) - \tfrac12\chi^2_{1,\,0.95}\} = \{\lambda : \ell(\lambda) \ge \ell(\hat\lambda) - 1.92\}$ (boxcox(y, alpha=0.05)).
  • Common practice: if a simple value (1, 0.5, 0, −1) lies inside the interval, use it; it is easier to explain and to invert.
  • In a regression, the assumption is about the residuals, so $\lambda$ should be chosen with the model in place (as R's MASS::boxcox(lm(...)) does), not from the raw $y$ alone.
Why do we need it?

Guessing $\lambda$ by eye is slow and subjective. The likelihood gives a principled best value and an interval that tells you how strongly the data prefers it, so you can defend "we used the log" with evidence.

Where is it used?

scipy.stats.boxcox, boxcox_normmax, boxcox_llf; sklearn's PowerTransformer (which fits $\lambda$ by maximum likelihood); Box-Cox options in forecasting libraries; R's MASS::boxcox profile plot.

How is it used?

yt, lam, ci = stats.boxcox(y, alpha=0.05). Look at the interval; round $\lambda$ to a simple value inside it if possible; check the Q-Q plot of the transformed data (or of the model residuals); store $\lambda$ to invert predictions later.

Pick a data set (200 positive values). Drag λ and watch three panels: the histogram of the transformed data, its Normal Q-Q plot, and the profile log-likelihood $\ell(\lambda)$ with your λ (green line), the MLE (purple dot) and the 95% interval (shaded). Press Jump to the MLE: the histogram becomes the most symmetric and the Q-Q plot the straightest. For Normal already notice how wide the interval is.

"Pick the $\lambda$ that makes the transformed data's variance smallest."

Lower $\lambda$ squeezes everything, so the variance always shrinks. The Jacobian term $(\lambda - 1)\sum \ln y$ corrects for this; without it the comparison is meaningless.

"$\hat\lambda = 0.13$, so I must use exactly 0.13."

$\hat\lambda$ is an estimate with uncertainty. If 0 (the log) is inside the 95% interval, the log is usually the better choice: easier to explain (percent effects) and to invert.

"After Box-Cox the data is Normal."

Box-Cox picks the most Normal-looking member of its family. If no power transform can make the data Normal (two humps, many zeros, heavy tails on the log scale), the best member is still not Normal. Always look at the Q-Q plot afterwards.

For a forecasting model, a Box-Cox $\lambda$ chosen on the whole history would be a mild form of leakage if you then evaluate on part of that history (Chapter 7.12): choose it on the training window only. And since your model's Normal or Student-t likelihood is about the residuals, judge $\lambda$ by the residual Q-Q plot after fitting the trend and seasonality, not by the histogram of raw demand.

$\ell(\lambda) = -\frac n2 \ln \hat\sigma^2(\lambda) + (\lambda - 1)\sum \ln y_i$ (= boxcox_llf); $\hat\lambda$ = its maximum.

95% interval: $\ell(\lambda) \ge \ell(\hat\lambda) - 1.92$. Round to 1, 0.5, 0 or −1 if inside.

Trap: never pick λ by smallest variance; check the Q-Q plot after transforming.

Quick check: the profile log-likelihood peaks at $\hat\lambda = 0.21$ with $\ell = -812.4$, and $\ell(0) = -813.6$, $\ell(0.5) = -815.9$. Which simple $\lambda$ would you use?

The 95% interval keeps $\lambda$ values with $\ell \ge -812.4 - 1.92 = -814.32$. $\ell(0) = -813.6$ is inside, $\ell(0.5) = -815.9$ is outside. So the log ($\lambda = 0$) is consistent with the data and is the natural choice.

Yeo-Johnson: Box-Cox for zeros and negative values

Box-Cox only works for positive numbers. But plenty of skewed data crosses zero: daily profit (some days lose money), a change in sales versus last year, residuals of a model. Yeo and Johnson built a version of the dial that works for every real number.

The trick: treat the positive side and the negative side separately. Positive values (and zero) get a Box-Cox transform of $y + 1$, so zero maps to zero smoothly. Negative values get the mirror image treatment with the power $2 - \lambda$. The two halves join at zero without a kink. With $\lambda = 1$ nothing changes at all. With $\lambda \lt 1$ large positive values are pulled in and negative values are pushed out, which reduces right skew; $\lambda \gt 1$ does the opposite.

Three ways to say it:

  • Picture: one dial again, but the ruler now extends smoothly through zero into the negatives.
  • Numbers: with $\lambda = 0.5$: $3 \mapsto 2$, $0 \mapsto 0$, $-3 \mapsto -4.67$.
  • Slogan: Yeo-Johnson = Box-Cox on $y + 1$ for $y \ge 0$, mirrored with $2 - \lambda$ for $y \lt 0$.

Transform $y = 3$, $0$ and $-3$ with Yeo-Johnson at $\lambda = 0.5$.

  1. $y = 3 \ge 0$: $\dfrac{(y+1)^\lambda - 1}{\lambda} = \dfrac{4^{0.5} - 1}{0.5} = \dfrac{2 - 1}{0.5} = 2$.
  2. $y = 0$: $\dfrac{1^{0.5} - 1}{0.5} = 0$. Zero always stays at zero.
  3. $y = -3 \lt 0$: $-\dfrac{(1-y)^{2-\lambda} - 1}{2 - \lambda} = -\dfrac{4^{1.5} - 1}{1.5} = -\dfrac{8 - 1}{1.5} = -4.67$.
  4. So the positive side was squeezed (3 → 2) and the negative side stretched (−3 → −4.67): a long right tail is pulled in and the left side pushed out, making the data more symmetric.
  5. With $\lambda = 1$: $3 \mapsto (4 - 1)/1 = 3$ and $-3 \mapsto -(4^1 - 1)/1 = -3$. The identity.

The Yeo-Johnson transform with parameter $\lambda$ is defined for all real $y$:

$$\psi(y, \lambda) = \begin{cases} \dfrac{(y+1)^\lambda - 1}{\lambda} & y \ge 0,\ \lambda \ne 0 \\[2mm] \ln(y + 1) & y \ge 0,\ \lambda = 0 \\[2mm] -\dfrac{(1-y)^{2-\lambda} - 1}{2 - \lambda} & y \lt 0,\ \lambda \ne 2 \\[2mm] -\ln(1 - y) & y \lt 0,\ \lambda = 2 \end{cases}$$
  • For $y \ge 0$ it is Box-Cox applied to $y + 1$. For $y \lt 0$ it is minus Box-Cox of $1 - y = 1 + |y|$ with power $2 - \lambda$.
  • It is increasing, continuous and smooth at $y = 0$ (value 0, slope 1 from both sides). $\lambda = 1$ is the identity.
  • $\lambda$ is chosen by maximum likelihood exactly as for Box-Cox, with the Jacobian term $(\lambda - 1)\sum_i \text{sign}(y_i)\ln(1 + |y_i|)$ (scipy.stats.yeojohnson, yeojohnson_llf).
  • sklearn.preprocessing.PowerTransformer() uses Yeo-Johnson by default and also standardizes the result to mean 0, sd 1 (standardize=True).
Why do we need it?

Shifting data to make Box-Cox work adds an arbitrary constant. Yeo-Johnson handles zeros and negatives directly, so skewed quantities that cross zero (profits, changes, residuals) can still be made more symmetric with one fitted dial.

Where is it used?

sklearn's PowerTransformer (default method) in feature pipelines, scipy.stats.yeojohnson, tabular ML preprocessing for gradient-free linear models and neural nets, and AutoML systems that normalize skewed features.

How is it used?

pt = PowerTransformer(); Xt = pt.fit_transform(X_train), then pt.transform(X_test) with the same fitted $\lambda$s; pt.lambdas_ shows them and pt.inverse_transform goes back.

300 days of profit (in hundreds of dollars), strongly right-skewed and roughly one day in eight negative, so Box-Cox cannot be used. Left: the Yeo-Johnson curve for your λ (blue) and the identity λ = 1 (dashed). Right: the histogram of the transformed values (standardized) with a Normal bell. Lower λ from 1 and watch the right tail shrink; press Jump to the MLE; then try λ = 2 to see skew get worse.

"Yeo-Johnson with the same $\lambda$ is the same as Box-Cox."

For positive data it is Box-Cox of $y + 1$, not of $y$, so the best $\lambda$ and the result differ (especially for values near 0). Do not mix up the two $\lambda$s.

"Fit PowerTransformer on the whole dataset, then split into train and test."

The fitted $\lambda$ (and the standardization) then contains information from the test set. Fit on the training data only and apply the fitted transformer to the test data.

Skewed exogenous regressors in your forecasting model that can be negative (a price change versus last week, a temperature anomaly) are typical Yeo-Johnson candidates before standardizing them. Fit the transform on the training window only, so that a backtest does not secretly use the future (Chapter 7.12).

Yeo-Johnson: $y \ge 0$: $((y+1)^\lambda - 1)/\lambda$; $y \lt 0$: $-((1-y)^{2-\lambda} - 1)/(2-\lambda)$; $\lambda = 1$ = identity.

Works for zeros and negatives; $\lambda$ by maximum likelihood; sklearn PowerTransformer default.

Trap: its $\lambda$ is not Box-Cox's $\lambda$; fit on training data only.

Quick check: compute $\psi(3, 1.5)$ and $\psi(-3, 1.5)$. What does $\lambda \gt 1$ do?

$\psi(3, 1.5) = (4^{1.5} - 1)/1.5 = (8 - 1)/1.5 = 4.67$. $\psi(-3, 1.5) = -(4^{0.5} - 1)/0.5 = -(2 - 1)/0.5 = -2$. Positive values are stretched and negative values squeezed: $\lambda \gt 1$ reduces left skew (and would add right skew).

Coming back: the mean of the log is not the log of the mean core

You fitted a model on the log scale and it predicts $\ln(\text{demand}) = 4.61$ for tomorrow. Natural move: $e^{4.61} = 100$, so "we expect 100 orders". But what is that 100? On the log scale the prediction is the centre of a symmetric bell, so it is both the mean and the median of the logs. Exponentiating keeps the order of values, so it carries the median across: 100 is the median demand. The mean demand is larger, because the exponential stretches the high side of the bell much more than the low side.

This matters whenever you add things up: total stock to order for a week, total revenue across users. Medians do not add up to the median of the sum, and using medians as means systematically under-forecasts totals.

Three ways to say it:

  • Picture: exponentiating a symmetric bell gives a right-skewed shape whose mean sits to the right of its median.
  • Numbers: for 1, 10, 100 the mean is 37, but the mean of the $\log_{10}$s is 1, and $10^1 = 10$ (the median and the geometric mean).
  • Slogan: transforms carry quantiles across, not means.

Two calculations.

  1. Data 1, 10, 100. Arithmetic mean: $(1 + 10 + 100)/3 = 111/3 = 37$.
  2. Logs (base 10): 0, 1, 2. Their mean is 1. Back-transform: $10^1 = 10$. This is the geometric mean $\sqrt[3]{1 \cdot 10 \cdot 100} = \sqrt[3]{1000} = 10$, here equal to the median, and far below 37.
  3. Forecast: the log-scale model says $\ln Y \sim N(4.605, 1^2)$ (centre $\ln 100$, residual sd 1). Naive back-transform: $e^{4.605} = 100$, the median.
  4. The mean of a log-normal is $e^{\mu + \sigma^2/2} = e^{4.605 + 0.5} = 100 \times e^{0.5} = 100 \times 1.649 = 164.9$.
  5. So "100" understates the expected demand by $1 - 100/164.9 = 39\%$. With a smaller residual sd, say $\sigma = 0.3$, the gap is only $1 - e^{-0.045} = 4.4\%$.
  • Jensen's inequality (for the log, which bends downward): $E[\ln Y] \le \ln E[Y]$, with equality only when $Y$ does not vary. Equivalently $e^{E[\ln Y]}$ (the geometric mean) $\le E[Y]$.
  • Quantiles survive any increasing transform $g$: the $p$-quantile of $g^{-1}(Z)$ is $g^{-1}$ of the $p$-quantile of $Z$. So back-transformed medians and prediction-interval endpoints are correct.
  • Means do not survive: if $\ln Y \sim N(\mu, \sigma^2)$, then median$(Y) = e^\mu$ but $E[Y] = e^{\mu + \sigma^2/2}$.
  • Without the Normal assumption, Duan's smearing estimator estimates the mean as $e^{\hat z} \times \frac1n \sum_i e^{\hat e_i}$, where $\hat e_i$ are the log-scale residuals.
  • The same issue appears for every non-linear transform (square root, Box-Cox): the back-transformed mean prediction estimates something close to the median, not the mean.
Why do we need it?

Inventory, revenue and capacity decisions are about totals and averages. Back-transforming a log-scale forecast naively gives the median, which can be far below the mean when the noise is large, and the error adds up over many items and days.

Where is it used?

Log-scale demand forecasts, log-revenue A/B analyses ("the geometric mean went up 3%" is not "the average revenue went up 3%"), Box-Cox forecasting with bias adjustment, and any model fitted to np.log(y).

How is it used?

Report back-transformed predictions as medians, or correct them: multiply by $e^{\hat\sigma^2/2}$ (log-normal) or by the smearing factor. Back-transform interval endpoints directly. Better still, simulate: draw from the predictive distribution on the log scale, exponentiate the draws, then take means or sums.

mean = median = ln 100 ≈ 4.61 ln 25 ln 100 ln 400 log scale: symmetric bell median = e^4.61 = 100 mean ≈ 120 0 100 200 300 400 original scale: skewed, mean > median exp
Left: on the log scale the forecast is a symmetric bell, so its centre is both mean and median. Right: after exponentiating, the median moves across exactly (e^4.61 = 100), but the long right tail pulls the mean up to about 120 (here σ = 0.6).

The model on the log scale is $\ln Y \sim N(\ln 100, \sigma^2)$. Drag σ (the residual sd on the log scale). The purple line is the naive back-transform $e^{\mu} = 100$ (the median); the green line is the true mean $e^{\mu + \sigma^2/2}$. Watch the gap grow with σ. The readout also checks this on 60 simulated days: the plain average of the days, $e^{\text{mean of logs}}$, and the smearing estimate. Press New sample.

"exp(prediction on the log scale) is the expected value."

It is (approximately) the median. The expected value needs a correction such as $e^{\hat\sigma^2/2}$, the smearing factor, or simulation.

"Back-transformed prediction intervals are also wrong."

Interval endpoints are quantiles, and quantiles pass through increasing transforms exactly. A 90% interval for $\ln Y$ back-transforms to a 90% interval for $Y$ (it just becomes lopsided).

"Daily median forecasts add up to the weekly median forecast."

Medians are not additive. To forecast a weekly total, simulate daily paths, add them up per path, and then take the median or mean of the totals.

"We modelled log-revenue and the treatment increased average revenue by 3%."

An additive 0.03 on the log scale is a 3% increase in the geometric mean (or the median, if the log-scale noise is symmetric). The arithmetic mean can change by a different amount if the treatment also changes the spread.

Model answer: "The log-scale effect is a multiplicative effect on the median. If the business question is about average revenue per user, I back-transform with a variance correction, or better, I simulate from the posterior predictive and compare the means directly."

In your forecasting model, if you ever fit $\ln(\text{demand})$, the exponentiated forecast is a median forecast; for stock planning based on expected totals, use the predictive draws: exponentiate each draw and then aggregate. Your Bayesian setup makes this natural, because the posterior predictive already gives samples. With the Negative Binomial likelihood on the original scale, the model's mean $\mu_t$ is already the expected count, so no back-transformation correction is needed.

$E[\ln Y] \le \ln E[Y]$ (Jensen). $e^{\text{mean of logs}}$ = geometric mean ≈ median, not the mean.

Log-normal: median $e^\mu$, mean $e^{\mu + \sigma^2/2}$. Smearing: $e^{\hat z}\cdot \overline{e^{\hat e}}$.

Quantiles and interval endpoints back-transform exactly; means do not. Aggregate by simulation.

Quick check: a log-scale model predicts $\ln Y = 3$ with residual sd $0.8$. Give the back-transformed median and mean.

Median $e^3 = 20.1$. Mean $e^{3 + 0.32} = e^{3.32} = 27.7$ (the correction factor is $e^{0.32} = 1.38$, i.e. +38%).

Standardization: z-scores core

A temperature of 30 and a price of 30 mean completely different things. To compare or combine variables measured in different units, we re-express each value as "how many standard deviations above or below its own average is it?". That number is a z-score. A z-score of +2 means "two typical distances above average", whatever the original units were.

Standardization is a linear transform: subtract the mean (shift) and divide by the sd (stretch). As you saw with Q-Q plots, a shift and a stretch never change shape. A skewed variable stays exactly as skewed; an outlier stays an outlier. Standardization fixes units and scale, not shape.

Three ways to say it:

  • Picture: slide the data so the average sits at 0, then rescale the ruler so one sd is one unit.
  • Numbers: heights 160, 170, 180 cm (mean 170, sd 10) become −1, 0, +1.
  • Slogan: z-scores change units, not shape.

Heights 160, 170, 180 cm.

  1. Mean $\bar x = (160 + 170 + 180)/3 = 170$.
  2. Sample sd ($n - 1$): deviations $-10, 0, 10$; squares $100 + 0 + 100 = 200$; $200/2 = 100$; $s = 10$.
  3. z-scores: $(160 - 170)/10 = -1$, $(170 - 170)/10 = 0$, $(180 - 170)/10 = 1$.
  4. sklearn's StandardScaler divides by $n$ instead: $\sigma = \sqrt{200/3} = 8.165$, giving $\mp 1.225$ and 0. Same shape, slightly different numbers.
  5. To go back: $x = \bar x + s\,z$, e.g. $170 + 10 \times 1 = 180$. Keep $\bar x$ and $s$!

The z-score (standardized value) of $x_i$ is

$$z_i = \frac{x_i - \bar x}{s} \quad\text{(or } \frac{x - \mu}{\sigma} \text{ with population values).}$$
  • The $z_i$ have mean 0 and sd 1 (with the same ddof used to compute $s$). Order, relative gaps, skewness and kurtosis are unchanged.
  • Undefined if $s = 0$ (all values equal).
  • Library conventions: scipy.stats.zscore and StandardScaler use ddof = 0 by default; pandas .std() uses ddof = 1. The difference only matters for small $n$.
  • Robust variant: $(x - \text{median}) / \text{IQR}$ (RobustScaler), which an outlier cannot inflate.
  • Fit the scaler on the training data only, store $\bar x$ and $s$, and apply the same numbers to new data and to invert predictions.
Why do we need it?

Optimizers (gradient descent, SVI) converge badly when features have wildly different scales (Optimization guide). Priors like $N(0, 1)$ on coefficients only make sense when the inputs are on a standard scale. Distances (k-NN, clustering, PCA) are dominated by whichever feature has the biggest units.

Where is it used?

StandardScaler in ML pipelines, PCA on the correlation matrix (Chapter 5.16), regularized regression, neural network inputs, and the global scaler in your A/B framework.

How is it used?

mu, sd = x_train.mean(), x_train.std(ddof=1); z = (x - mu) / sd for every split; model on z; convert effects back by multiplying by sd (and adding mu for levels).

Top ruler: 8 days of orders (drag them). Bottom ruler: their z-scores. The dashed lines connect each day to its z-score. Notice that the pattern of gaps is identical on both rulers: only the units changed. Drag one day far to the right: it stays the far-right point, and the other z-scores shrink a little because the sd grew. Toggle the divisor to compare n − 1 with StandardScaler's n.

"Standardizing makes the data Normal."

It only shifts and rescales. Skewed data stays skewed and heavy tails stay heavy. To change shape you need a non-linear transform (log, Box-Cox, Yeo-Johnson).

"Compute the mean and sd on all the data, then split into train and test."

That leaks information from the test period into training. Fit the scaler on the training data, then apply it unchanged to validation, test and future data.

"A coefficient of 0.5 on a standardized input means +0.5 orders per unit of input."

It means +0.5 (in the output's units) per one standard deviation of the input. Divide by the input's sd to get the effect per original unit.

In your forecasting model, exogenous regressors are usually standardized so that one prior scale (for example $\beta \sim N(0, 1)$ or a Laplace prior) is sensible for all of them and SVI's optimizer sees well-scaled gradients. Store each regressor's training mean and sd: you need them to standardize future regressor values at forecast time and to read the coefficients back in original units.

$z = (x - \bar x)/s$: mean 0, sd 1; shape, order and outliers unchanged.

StandardScaler / zscore: ddof = 0; pandas: ddof = 1. Robust: (x − median)/IQR.

Trap: fit on training data only; keep $\bar x$, $s$ to invert; coefficients are "per sd".

Quick check: a regressor has training mean 20 and sd 4. A future day has value 30. What is its z-score, and what does a coefficient of 3 (on the z scale) mean per original unit?

$z = (30 - 20)/4 = 2.5$. A coefficient of 3 per standard deviation is $3/4 = 0.75$ per original unit of the regressor.

One global scaler, not one per group core

Suppose you want to know whether variant B beats variant A. You decide to standardize first, and it seems tidy to standardize each group on its own: subtract A's mean from A's values, B's mean from B's values. But look at what that does: every group now has mean exactly 0. The difference between the groups, the very thing you wanted to measure, has been subtracted away. Per-group scaling also divides each group by its own sd, so any difference in spread is erased too.

A global scaler uses one mean and one sd for all the data. Every value is shifted and stretched by the same amounts, so the groups keep their positions relative to each other. The effect is still there, just expressed in "global sd" units, and you can convert it back exactly.

Three ways to say it:

  • Picture: per-group scaling slides every group's centre to 0; a global scaler moves all groups together like one rigid block.
  • Numbers: A = 10, 12, 14 and B = 14, 16, 18: per-group z-scores are −1, 0, 1 for both (difference 0); global z-scores give means −0.71 and +0.71 (difference 1.41 sd = 4 original units).
  • Slogan: scale everything with one ruler, or you measure the groups against themselves.

Daily orders per user: variant A gives 10, 12, 14 and variant B gives 14, 16, 18. The true difference in means is $16 - 12 = 4$.

  1. Per-group scaling. A: mean 12, sd 2 → z = −1, 0, 1. B: mean 16, sd 2 → z = −1, 0, 1. Mean difference $0 - 0 = 0$. The effect is gone, and no back-transform can recover it from the z-scores alone.
  2. Global scaler. All six values: mean $84/6 = 14$. Deviations −4, −2, 0, 0, 2, 4; squares sum $16 + 4 + 0 + 0 + 4 + 16 = 40$; $s = \sqrt{40/5} = \sqrt 8 = 2.828$.
  3. Global z for A: $(10-14)/2.828 = -1.414$, $-0.707$, $0$, mean $-0.707$. For B: $0$, $0.707$, $1.414$, mean $+0.707$.
  4. Difference on the scaled data: $0.707 - (-0.707) = 1.414$ global sds.
  5. Back to orders: $1.414 \times 2.828 = 4.0$. Exactly the true difference.
  • Global scaler: $z_{gi} = (y_{gi} - \bar y)/s$ with one overall $\bar y$ and $s$ (from the training/pre-period data). A difference $\Delta_z$ on this scale corresponds to $\Delta_y = s\,\Delta_z$ in original units.
  • Per-group scaling: $z_{gi} = (y_{gi} - \bar y_g)/s_g$ forces every group mean to 0 and every group sd to 1. Between-group differences in means (and in spreads) are removed by construction.
  • In a hierarchical model, per-group scaling makes the between-group variation $\tau$ look like zero, so partial pooling has nothing left to learn (Chapter 4.6: total variance = within + between; per-group scaling deletes the "between" part).
  • Per-group scaling is fine only when the group differences are nuisance you deliberately want to remove (for example, comparing the shape of daily patterns across stores of different sizes), never when they are what you want to estimate.
Why do we need it?

Standardizing helps priors and optimizers, but it must not change the question. A single global scaler gives the numerical benefits while keeping every between-group difference intact and exactly convertible back.

Where is it used?

Your A/B framework's global scaler, hierarchical models across segments or stores, multi-series forecasting (scaling across series vs within series), and any pipeline where a groupby(...).transform(zscore) could quietly delete an effect.

How is it used?

Compute one mean and sd on the pooled (training or pre-experiment) data, apply them to every group, fit the model, and multiply effect estimates by $s$ to report them in original units.

raw orders global scaler per-group scaling AB AB AB B − A = 16 − 12 = 4 B − A = 1.41 sd → × 2.83 = 4 B − A = 0: effect wiped out
The same two groups three ways. A global scaler (middle) shifts and stretches both groups together, so the gap between their means (purple bars) survives and converts back exactly. Per-group scaling (right) puts both means at 0 and the effect disappears.

Three segments with 30 users each. Set the true B − A difference and the spread within groups. Left: raw values with group means (orange). Right: the same data after scaling; switch between Global scaler and Per-group scaling. With the global scaler the gaps survive and the readout converts them back to orders exactly; with per-group scaling every group mean is 0 and the between-group variation vanishes, whatever difference you set.

"Normalize each segment separately so that big and small segments are comparable."

That makes every segment's mean 0 and sd 1, so segment differences, including treatment effects estimated across segments, become invisible. Use one global scaler; let the model (for example hierarchical pooling) handle differences between segments.

"The global scaler changes the effect size."

It changes only the units. An effect of 1.41 global sds is the same as 4 orders when $s = 2.83$; multiply by $s$ to report it.

This is the reason to use one global scaler for all groups, as your A/B framework does. The hierarchical model estimates group parameters around a global mean with between-group spread $\tau$, and the decision quantity $P(\theta_B \gt \theta_A \mid D)$ compares group parameters. Scaling each group (or each variant) by its own mean and sd would set every group's mean to 0: $\tau$ would collapse toward 0, partial pooling would have nothing to pool, and the treatment difference you are trying to estimate would be removed before the model ever saw it. With one global mean and sd, priors such as $N(0, 1)$ are sensible for every group, and posterior effects convert back to original units by multiplying by the global sd.

"We standardized each group to mean 0 and sd 1 before fitting, to make the groups comparable."

Per-group standardization removes exactly the between-group differences we want to estimate. We use one global scaler, so groups stay comparable and effects are recoverable.

Model answer: "We standardize with a single global mean and sd. It keeps the numerics and priors well behaved, but every group is shifted and stretched by the same amounts, so differences between groups survive and convert back by multiplying by the global sd. Per-group normalization would force every group mean to zero and erase the treatment effect and the between-group variance that the hierarchical model needs."

Global: $z = (y - \bar y)/s$ for everyone → effects survive; $\Delta_y = s\,\Delta_z$.

Per-group: every group mean 0, sd 1 → the group effect and $\tau$ are erased.

Example: A = 10, 12, 14, B = 14, 16, 18 → global diff 1.414 sd × 2.828 = 4; per-group diff 0.

Quick check: a hierarchical model across 50 stores estimates the between-store sd $\tau$. Someone z-scores each store's sales separately before fitting. What will $\tau$ come out as, and why?

Close to zero. After per-store z-scoring, every store's mean sales is exactly 0, so there is no between-store variation left in the data for $\tau$ to describe. The model would wrongly conclude that the stores are all alike and pool them completely.

Choosing a transform, or changing the likelihood instead

There are two roads to the same goal. Road one: change the data's ruler (log, square root, Box-Cox) until a simple model with Normal, constant-spread noise fits. Road two: keep the data as it is and change the model's assumption: a Poisson or Negative Binomial for counts, a Gamma or log-normal for positive skewed values, a Student-t for heavy tails, often with a log "link" so the mean stays positive.

The two roads answer slightly different questions. A Normal model of $\ln y$ describes the typical (median, geometric-mean) value; a model with a log link describes the mean of $y$ directly. With modern tools (GLMs, NumPyro) the second road is often cleaner, because nothing has to be transformed back. The first road is still the right one for tools that need Normal-ish noise, for features, and for plots.

Three ways to say it:

  • Picture: either bend the ruler to fit the model, or pick a model that fits the ruler you have.
  • Numbers: for 1, 10, 100 a model of the logs has centre $e^{\text{mean log}} = 10$; a log-link model of the mean has centre 37.
  • Slogan: transform the data, or choose the likelihood; know which question each one answers.

Three users spent 1, 10 and 100 dollars. Fit the simplest possible model (just a centre) two ways.

  1. Transform road: Normal model for $\ln y$. Its centre is the mean of the logs: $(\ln 1 + \ln 10 + \ln 100)/3 = (0 + 2.303 + 4.605)/3 = 2.303$. Back-transformed: $e^{2.303} = 10$. That is the geometric mean (and here the median).
  2. Likelihood road: a Gamma or Poisson-type model with a log link, $\ln E[y] = \beta_0$. With only an intercept its estimate matches the sample mean: $E[y] = 37$, so $\beta_0 = \ln 37 = 3.61$.
  3. Both are "log models", but one estimates $E[\ln y]$ and the other $\ln E[y]$. By Jensen's inequality, $E[\ln y] \le \ln E[y]$: $2.303 \lt 3.61$.
  4. If the business cares about total revenue, it is the mean (37 per user) that matters; the transform road needs a back-transformation correction to get there.
situationtransform roadlikelihood road
counts, spread ≈ √mean$\sqrt y$ (Anscombe)Poisson, log link
counts, extra spreadlog1p / Box-CoxNegative Binomial, log link
positive, spread ∝ level, right skew$\ln y$ (or Box-Cox $\lambda$)Gamma or log-normal, log link
skewed, crosses zeroYeo-Johnsona skewed or heavy-tailed likelihood
symmetric, heavy tailsnone helps muchStudent-t
only the units differz-score (global scaler)(same; for priors and optimizers)
  • Tree models (gradient boosting, random forests) split on the order of a feature, so monotone feature transforms do not change them. Transforming the target still matters, because it changes the loss (e.g. squared error on $\ln y$ cares about ratios).
  • Whatever you choose: check the residual Q-Q plot afterwards (Chapter 4.17), fit any transform on training data only, and back-transform with care.
Why do we need it?

Transforming by habit can answer the wrong question (median instead of mean), break at zeros, or hide a variance problem. Knowing both roads lets you pick the one that matches your data and the decision you have to make.

Where is it used?

Choosing between log-revenue and a Gamma GLM in A/B analysis, between log-demand and a Negative Binomial likelihood in forecasting, feature engineering for linear models vs tree models, and the likelihood-selection question of your syllabus (Part 30).

How is it used?

Ask: what values can the data take? How does the spread grow? Can I choose the likelihood? Then follow the table above, fit, and compare the residual diagnostics (and predictive performance) of the candidates.

What do you need? Comparable units, a good optimizer, priors on one scale A Normal-noise tool needs symmetric, even-spread data You can choose the likelihood (Bayes, GLM) Standardize: z = (x − x̄) / s one global scaler shape stays the same fit on training data counts → √y spread ∝ level → log zeros → log1p, Yeo-Johnson negatives → Yeo-Johnson unsure → Box-Cox λ by MLE counts → Poisson / NB positive, skewed → Gamma or log-normal (log link) heavy tails → Student-t mean stays on original scale Always: check the residual Q-Q plot afterwards, and back-transform with care (median vs mean).
Three goals, three answers. Standardize for units; transform when a Normal-noise tool needs symmetric, even-spread data; change the likelihood when you can, so the model works on the original scale.

Answer the three questions about a variable. The ladder shows which power (Box-Cox λ) the transform road suggests, and the readout gives the likelihood road too, with the reason. Try: counts + spread grows like the mean + "I can choose the likelihood"; then the same data with "linear model / Normal noise".

"A log-transformed Normal model and a log-link model are the same thing."

The first models $E[\ln y]$ (a median-type centre); the second models $\ln E[y]$ (the mean). They give different centres, different effect interpretations and need different back-transformations.

"Always transform skewed features before gradient boosting."

Trees only use the order of each feature, so monotone transforms of features change nothing. Spend the effort on the target and on the loss instead.

Both projects mostly take the likelihood road: Beta-Binomial for conversions, Poisson for count metrics, Normal or Student-t for continuous metrics, and Normal, Student-t or Negative Binomial for demand in the forecasting model. Transforms still appear around them: the global scaler for numerical stability and sensible priors, standardized regressors, and possibly a log scale for strongly multiplicative series. Being able to say why you chose a likelihood instead of a transform (the mean stays on the original scale, zeros are handled, no back-transformation bias) is a strong interview answer.

Two roads: transform the data (√, log, Box-Cox, Yeo-Johnson) or choose the likelihood (Poisson, NB, Gamma, Student-t) with a log link.

Log-transform model → $E[\ln y]$ (median-type); log-link model → $\ln E[y]$ (mean).

Trees ignore monotone feature transforms. Always check residuals afterwards.

Quick check: daily demand counts with many zeros, spread growing faster than √mean, and you are writing a NumPyro model. Which road and which choice?

The likelihood road: a Negative Binomial likelihood with a log link (for example NumPyro's NegativeBinomial2(mean, concentration)). It handles zeros, has Var $= \mu + \mu^2/\alpha$ for the extra spread, and keeps the mean on the original scale, so no back-transformation is needed.

Recap, cheat sheet and practice

  • A transformation changes the ruler, not the order. Reasons: skew, spread growing with level, multiplicative effects, curved relationships.
  • Log: ratios become differences, log-normal becomes Normal, spread ∝ level becomes constant. Needs $y \gt 0$; zeros need log1p or $\log(y + c)$, and $c$ is a real choice.
  • Square root: stabilizes Poisson counts (sd ≈ 0.5). General rule: sd ∝ mean$^k$ → $y^{1-k}$ (log at $k = 1$).
  • Box-Cox $(y^\lambda - 1)/\lambda$, log at $\lambda = 0$, $y \gt 0$; choose $\lambda$ by the profile log-likelihood (with the Jacobian term) and round within the 95% interval. Yeo-Johnson extends it to zeros and negatives.
  • Back-transformation: quantiles pass through, means do not. $e^{\text{mean log}}$ is a median-type value; the log-normal mean is $e^{\mu + \sigma^2/2}$. Aggregate by simulation.
  • Standardization changes units, never shape. Fit on training data; keep $\bar x$, $s$.
  • One global scaler: per-group scaling erases the group effect and the between-group variance; this is why a single global scaler, as in your A/B framework, is the safe choice.
  • Often the better road is to choose the likelihood (Poisson, NB, Gamma, Student-t) instead of transforming.

Cheat sheet

ToolFormulaUse it when / remember
Log$\ln y$; $\beta$ → ×$e^\beta$positive, right-skewed, spread ∝ level, % effects
log1p$\ln(1 + y)$, inverse expm1zeros present; the "+1" depends on units
Square root$\sqrt y$, $\sqrt{y + 3/8}$Poisson counts: sd ≈ 0.5
Variance-stabilizingsd ∝ $\mu^k$ → $y^{1-k}$delta method $\text{Var}(g(Y)) \approx g'(\mu)^2\text{Var}(Y)$
Box-Cox$(y^\lambda - 1)/\lambda$; $\ln y$ at 0$y \gt 0$; $\ell(\lambda) = -\frac n2\ln\hat\sigma^2 + (\lambda-1)\sum\ln y$
λ interval$\ell(\lambda) \ge \ell(\hat\lambda) - 1.92$round to 1, 0.5, 0, −1 if inside
Yeo-Johnson$y \ge 0$: $((y+1)^\lambda - 1)/\lambda$; $y \lt 0$: $-((1-y)^{2-\lambda} - 1)/(2-\lambda)$zeros and negatives; sklearn default
Back-transformmedian $e^\mu$, mean $e^{\mu + \sigma^2/2}$quantiles exact, means need correction
z-score$(x - \bar x)/s$units only; StandardScaler ddof = 0
Global scalerone $\bar y$, $s$ for all groupseffects survive: $\Delta_y = s\,\Delta_z$
Code it · Python
import numpy as np
from scipy import stats, special
from sklearn.preprocessing import PowerTransformer, StandardScaler

rng = np.random.default_rng(0)

# 1) log, log1p and the "mean of logs" trap
y = np.array([1, 10, 100.0])
print(y.mean(), np.exp(np.log(y).mean()).round(6))       # 37.0 10.0   (mean vs geometric mean)
print(np.log1p([0, 1, 3, 9, 99]).round(3))               # [0.    0.693 1.386 2.303 4.605]

# 2) the square root stabilizes Poisson counts
for lam in [4, 100]:
    c = rng.poisson(lam, 200_000)
    print(lam, c.std().round(2), np.sqrt(c).std().round(3), np.sqrt(c + 3 / 8).std().round(3))
# 4 2.0 0.552 0.499      raw sd grows like sqrt(lambda) ...
# 100 10.01 0.501 0.5    ... but sd of sqrt(y) stays near 0.5

# 3) Box-Cox: the profile log-likelihood and the MLE
y = np.array([1, 2, 4, 8, 16.0])
for lam in [1, 0.5, 0, -0.5]:
    print(lam, round(stats.boxcox_llf(lam, y), 3))       # -8.483, -7.268, -6.832, -7.268
yt, lam_hat, ci = stats.boxcox(y, alpha=0.05)
print(round(lam_hat, 4), np.round(ci, 2))                # 0.0 [-1.08  1.08]   (5 points: very wide)

big = rng.lognormal(np.log(20), 0.6, 500)                 # log-normal data: expect lambda near 0
yt, lam_hat, ci = stats.boxcox(big, alpha=0.05)
print(round(lam_hat, 3), np.round(ci, 3))                # 0.048 [-0.079  0.176]  -> use the log
print(np.allclose(special.inv_boxcox(yt, lam_hat), big))  # True  (exact inverse)

# 4) Yeo-Johnson for data with negatives (sklearn default; it also standardizes)
profit = np.exp(np.log(20) + 0.8 * rng.standard_normal(300)) - 8
pt = PowerTransformer()                                   # method="yeo-johnson", standardize=True
z = pt.fit_transform(profit.reshape(-1, 1)).ravel()
print((profit <= 0).sum(), pt.lambdas_.round(3), round(stats.skew(profit), 2), round(stats.skew(z), 2))
# 37 [0.532] 2.21 0.17     (37 values <= 0, so Box-Cox could not be used)

# 5) back-transforming a log-scale forecast: median vs mean
mu, sigma = np.log(100), 1.0
print(round(np.exp(mu), 2), round(np.exp(mu + sigma**2 / 2), 2))   # 100.0 164.87
draws = np.exp(rng.normal(mu, sigma, 200_000))             # simulate, then summarize
print(round(np.median(draws), 1), round(draws.mean(), 1))  # 99.6 164.8

# 6) global scaler vs per-group scaling
A = np.array([10, 12, 14.0]); B = np.array([14, 16, 18.0])
allv = np.r_[A, B]; m, s = allv.mean(), allv.std(ddof=1)
zA, zB = (A - m) / s, (B - m) / s
print(round(zB.mean() - zA.mean(), 3), round((zB.mean() - zA.mean()) * s, 3))   # 1.414 4.0
pA = (A - A.mean()) / A.std(ddof=1); pB = (B - B.mean()) / B.std(ddof=1)
print(pB.mean() - pA.mean())                              # 0.0  -> the effect is wiped out

# StandardScaler divides by n (ddof = 0)
print(StandardScaler().fit_transform(np.array([[160.0], [170.0], [180.0]])).ravel().round(4))  # [-1.2247  0.  1.2247]
Test yourself

1. The Box-Cox transform with $\lambda = 0$ is…

$y^\lambda = e^{\lambda\ln y} \approx 1 + \lambda \ln y$ for small $\lambda$, so $(y^\lambda - 1)/\lambda \to \ln y$. The family is defined to be $\ln y$ at exactly 0, which makes it smooth.

2. A model of $\ln(\text{revenue})$ gives the treatment a coefficient of 0.4. On the original scale this means…

A difference of logs is a log ratio: $e^{0.4} = 1.49$. The "≈ 100β%" shortcut is only good for small β (below about 0.1).

3. Poisson counts from stores with means 4 and 100. Which transform makes their noise level about equal?

Poisson sd = √mean, so $k = ½$ and $y^{1-k} = \sqrt y$; the sd of $\sqrt Y$ is about 0.5 for both. The log over-corrects counts.

4. You forecast $\ln(\text{demand})$ and report $e^{\text{forecast}}$. What does that number estimate?

Increasing transforms carry quantiles across, so the centre of a symmetric log-scale distribution maps to the median. The mean is larger: $e^{\mu + \sigma^2/2}$ for a log-normal.

5. You standardize a strongly right-skewed variable to z-scores. Its skewness is now…

A z-score is a shift and a stretch. Those never change shape, so skewness, kurtosis and outliers are unchanged.

6. In an A/B framework with segments, why is per-group z-scoring dangerous?

Each group is centred on its own mean, so all group means become 0 and their spreads become 1. A single global scaler keeps the differences and converts back exactly.

Practice problems

A. Data 2, 8, 32. Take natural logs, check the spacing, and compare the arithmetic mean with the back-transformed mean of the logs.

$\ln 2 = 0.693$, $\ln 8 = 2.079$, $\ln 32 = 3.466$: equal gaps of $1.386 = \ln 4$, because each value is 4 times the previous one. Mean of logs $= 2.079 = \ln 8$, so the back-transform gives 8 (the geometric mean $\sqrt[3]{2 \cdot 8 \cdot 32} = \sqrt[3]{512} = 8$). The arithmetic mean is $42/3 = 14$. So $e^{\text{mean of logs}} = 8 \lt 14$, as Jensen's inequality says.

B. Counts have mean 50. Use the delta method to find the sd of $\sqrt Y$ and of $\ln Y$ if (i) $Y$ is Poisson, (ii) $Y$ is Negative Binomial with Var $= \mu + \mu^2/5$.

(i) sd$(Y) = \sqrt{50} = 7.07$. sd$(\sqrt Y) \approx 7.07/(2\sqrt{50}) = 0.5$; sd$(\ln Y) \approx 7.07/50 = 0.14$. (ii) Var $= 50 + 2500/5 = 550$, sd $= 23.45$. sd$(\sqrt Y) \approx 23.45/(2 \times 7.07) = 1.66$; sd$(\ln Y) \approx 23.45/50 = 0.47$. For overdispersed counts at high volume the spread grows nearly like the mean, so the log is closer to the right stabilizer (a simulation gives 1.64 and 0.50, close to these).

C. Three product groups have means 2, 8, 32 and standard deviations 0.4, 1.6, 6.4. Which transform stabilizes the spread? Check it.

Mean ×4 goes with sd ×4, so $4^k = 4$, $k = 1$: spread ∝ level (CV = 20%). Use the log. Delta method: sd$(\ln Y) \approx$ sd/mean $= 0.4/2 = 1.6/8 = 6.4/32 = 0.2$ for all three groups.

D. Compute Yeo-Johnson $\psi(8, 0)$ and $\psi(-8, 0)$. What does $\lambda = 0$ do to the negative side?

$y = 8 \ge 0$, $\lambda = 0$: $\ln(8 + 1) = \ln 9 = 2.197$. $y = -8 \lt 0$, $\lambda \ne 2$: $-\dfrac{(1 + 8)^{2} - 1}{2} = -\dfrac{81 - 1}{2} = -40$. The positive side is squeezed hard (like a log) while the negative side is stretched hard (power $2 - 0 = 2$). That is what removes a strong right skew; if the data had a long left tail instead, a small $\lambda$ would make it worse.

E. Each day's demand has $\ln Y \sim N(\ln 100, 0.5^2)$, independent over 7 days. A planner adds up the seven back-transformed forecasts $e^{\ln 100} = 100$. What is wrong, and what is the expected weekly total?

Each $e^{\ln 100} = 100$ is a daily median. The daily mean is $100\,e^{0.5^2/2} = 100\,e^{0.125} = 113.3$. Means add up, so the expected weekly total is $7 \times 113.3 = 793$, not 700: summing medians under-forecasts by about 12%. For the weekly median or intervals, simulate seven-day paths and add them per path.

F. Explain to an interviewer why your A/B framework uses one global scaler rather than normalizing each group separately.

"We standardize for numerical stability and so that priors like $N(0, 1)$ make sense, but we do it with one mean and one standard deviation computed on all the data. That shifts and stretches every group identically, so the differences between variants and segments survive, and any posterior effect converts back to original units by multiplying by the global sd. If we normalized each group by its own mean and sd, every group mean would become exactly 0: the treatment effect would be removed before the model saw the data, the between-group variance in the hierarchical model would collapse to zero, and partial pooling would have nothing left to estimate."

Appendix A

Glossary

Every important word of this guide in one place, explained in plain English (178 terms). Type in the box to filter: it searches the terms and their explanations. The small numbers after each entry link to the section that teaches it.

All terms, A to Z

No term matches. Try a shorter word, or press / to search the whole guide.

68–95–99.7 rule
For Normal data, about 68.3%, 95.4% and 99.7% of values lie within 1, 2 and 3 sds of the mean (95% within 1.96 sd). It holds for Normal shapes only. 4.9 4.12
Addition rule (inclusion–exclusion)
$P(A\cup B)=P(A)+P(B)-P(A\cap B)$: add the two chances, then subtract the overlap you counted twice. Union bound: $P(A_1\cup\dots\cup A_m)\le\sum_i P(A_i)$. 4.2
Anscombe's quartet
Four small datasets with the same means, variances and Pearson $r$ but completely different scatter plots. The lesson: plot before you trust $r$. 4.15
At least one (complement trick)
$P(\text{at least one}) = 1-P(\text{none})$; for $m$ independent tries with chance $p$ each, $1-(1-p)^m$. Twenty metrics tested at 5% give about 64%. 4.2
Axioms of probability
The three rules every probability obeys: it is never negative, the whole sample space has probability 1, and chances of events that cannot happen together add up. 4.2
Back-transformation
Converting results from a transformed scale (such as log) back to the original units. Quantiles come back exactly; means do not, because the mean of the log is not the log of the mean. Duan's smearing estimator corrects the mean. 4.18
Bandwidth ($h$)
The width of each bump in a kernel density estimate. Small $h$ gives a spiky, noisy curve; large $h$ gives an over-smoothed one that can hide peaks. It matters far more than the kernel shape. 4.16
Base-rate fallacy
Ignoring the base rate and reading $P(\text{evidence}\mid H)$ as if it were $P(H\mid\text{evidence})$. A 90%-sensitive test for a 1% condition gives only about a 9% chance of the condition after a positive. Natural frequencies ("imagine 10 000 people") make the base rate hard to ignore. 4.3
Base rate (prevalence)
How common something is before you see any evidence, for example 1% of people have the condition. Bayes' theorem cannot work without it. 4.3
Bayes' theorem
$P(A\mid B)=P(B\mid A)\,P(A)/P(B)$. It flips a conditional probability around, using the base rate $P(A)$: posterior ∝ likelihood × prior. 4.3 · Bayes' theorem 4.3 · Bayes in your A/B framework
Bayesian view
The unknown parameter is uncertain, so it gets a probability distribution (prior, then posterior); probability means degree of belief; you condition on the data you actually saw. 4.1
Bernoulli distribution
One yes/no trial: 1 with probability $p$, 0 otherwise. Mean $p$, variance $p(1-p)$. One user's conversion. 4.7
Bessel's correction ($n-1$)
Dividing the sum of squared deviations by $n-1$ instead of $n$. The deviations are measured from $\bar x$, which was fitted to the same data, so the raw sum is too small by one $\sigma^2$ on average. 4.5
Beta distribution
A distribution for an unknown probability between 0 and 1, with density $\propto p^{\alpha-1}(1-p)^{\beta-1}$ and mean $\alpha/(\alpha+\beta)$. The usual prior for a conversion rate. 4.11 · The Beta distribution 4.11 · Shapes of the Beta
Bias
Being off in the same direction on average. An estimator is biased if its average over repeated samples misses the true value (the ÷$n$ variance is biased low). Bias from a skewed way of collecting the data does not shrink as $n$ grows. 4.5 4.1
Bimodal / multimodal
Having two (or more) peaks. Often a sign of two mixed groups; a large KDE bandwidth can hide the second peak. 4.16
Binomial distribution
The number of yeses in $n$ independent trials with the same $p$: $P(k)=\binom nk p^k(1-p)^{n-k}$, where $\binom nk$ counts the possible orders. Mean $np$, variance $np(1-p)$; it needs a fixed, known $n$. 4.7
BINS conditions
Memory aid for when a count is Binomial: Binary outcomes, Independent trials, a fixed Number of trials, the Same $p$. Breaking them usually inflates the variance. 4.7
Boundary bias and reflection (KDE)
Near a hard edge (like 0 for positive data), a KDE leaks mass past the edge and shows too little density just inside it. Fix: reflection (add a mirror-image bump at $-x_i$ for every point) or work on a log scale. 4.16
Box-Cox transform
The family of power transforms $(y^\lambda-1)/\lambda$, with $\ln y$ at $\lambda=0$; it needs $y\gt0$. $\lambda$ is chosen by maximum likelihood, whose Jacobian term $(\lambda-1)\sum\ln y_i$ stops it from simply picking the smallest variance. 4.18 · Box-Cox: the family 4.18 · Choosing λ by maximum likelihood
Box plot
A box from $Q_1$ to $Q_3$ with a line at the median, whiskers to the most extreme points inside the 1.5 × IQR fences, and dots for points beyond them. 4.14
Breakdown point
The largest share of the data you can corrupt before an estimate can be pushed arbitrarily far. Mean 0, IQR 25%, median and MAD 50% (the most possible). 4.14
Categorical distribution
One pick among $K$ labels with probabilities $\pi_1,\dots,\pi_K$ that add to 1. With $K=2$ it is the Bernoulli. 4.7
Cauchy distribution
A bell-shaped but extremely heavy-tailed distribution (the Student-t with $\nu=1$). It has no mean, so averages of Cauchy draws never settle. 4.13 4.9
CDF (cumulative distribution function)
$F(x)=P(X\le x)$, rising from 0 to 1: a staircase for a discrete variable, a smooth ramp for a continuous one. $P(a\lt X\le b)=F(b)-F(a)$. 4.4
Central Limit Theorem (CLT)
For iid draws with finite variance, the CDF of the standardized sample mean approaches the $N(0,1)$ CDF (convergence in distribution), so $\bar X\approx N(\mu,\sigma^2/n)$. It is about the average, not the data, and gives no safe $n$. 4.13 · Central Limit Theorem 4.13 · What the CLT does not say
Chain rule of probability
$P(A_1\cap\dots\cap A_n)=P(A_1)\,P(A_2\mid A_1)\,P(A_3\mid A_1\cap A_2)\cdots$: multiply step by step along a funnel, or along a path of a probability tree (then add the paths you want). 4.3
Chebyshev's inequality
For any distribution with finite variance, $P(|X-\mu|\ge k\sigma)\le 1/k^2$: at least 75% within 2 sds and 88.9% within 3. Always true, usually loose. One-sided version (Cantelli): $P(X-\mu\ge k\sigma)\le 1/(1+k^2)$. 4.12 · Chebyshev's inequality 4.12 · Within k sds vs 68–95–99.7
Coefficient of variation (CV)
sd divided by mean: a unit-free measure of relative spread. A Normal model of positive data puts $\Phi(-1/CV)$ of its mass below zero. 4.10
Collider (common effect)
A variable caused by two others. Looking only at one value of it (say, users who left) can create a link between two independent causes; this is called explaining away. Never control for a collider. 4.3 4.15
Complement
$A^c$, "not A": every outcome that is not in $A$. $P(A^c)=1-P(A)$. 4.2
Concentration
How tightly a Beta or Dirichlet sits around its mean: $\kappa=\alpha+\beta$ or $\alpha_0=\sum\alpha_k$, read as "worth this many observations". In the NB2, the concentration $\alpha$ sets the extra spread: large $\alpha$ is close to Poisson. 4.11 4.8
Conditional distribution
The distribution of $X$ inside the group where $Y=y$: $p(x\mid y)=p(x,y)/p_Y(y)$. Keep one row of the table and divide by its total. 4.6
Conditional expectation
$E[X\mid Y=y]$ is the mean of one group (a number). $E[X\mid Y]$ is a random variable: the mean of whichever group $Y$ falls into. 4.6
Conditional independence
$X$ and $Y$ are independent inside each level of $Z$: $p(x,y\mid z)=p(x\mid z)\,p(y\mid z)$. It neither implies nor follows from plain independence. 4.3
Conditional probability
$P(A\mid B)=P(A\cap B)/P(B)$: the chance of $A$ once you know $B$ happened. $B$ becomes the new 100%. 4.3
Confounder (common cause)
A third variable that affects both $X$ and $Y$ and so creates a correlation between them, like temperature driving both ice-cream and sunscreen orders. Randomization breaks confounding. 4.15 4.3
Conjugate prior
A prior that gives a posterior from the same family, so updating is just arithmetic: Beta + Binomial counts → Beta; Dirichlet + Multinomial counts → Dirichlet; Gamma + Poisson counts → Gamma. 4.11 4.10
Continuous random variable
A random variable that can take any value in a range. $P(X=x)=0$ for every single value; probabilities are areas under a density. A mixed variable (revenue: a lump at 0 plus a spread of positive amounts) is neither purely discrete nor continuous. 4.4
Correlation
Covariance with the units taken out: $\rho=Cov(X,Y)/(\sigma_X\sigma_Y)$, always between −1 and 1, and ±1 only for an exact straight line. 4.15
Correlation matrix
The table of correlations between every pair of variables: ones on the diagonal, symmetric. Use it to spot near-duplicate regressors before fitting. 4.15
Covariance
$Cov(X,Y)=E[(X-\mu_X)(Y-\mu_Y)]$: positive when the two tend to be above (or below) their means together. Its size depends on the units. 4.15
Curse of dimensionality
In many dimensions a "neighbourhood" must be huge to hold a fair share of the data, so methods like KDE need enormous samples. Keep KDE to one or two dimensions. 4.16
Data-generating process (DGP)
The real mechanism that produced your data: real process → random variables → distribution → observed values. Your dataset is one run of it. 4.1
Dataset, observation, variable
A dataset is a table: each row is an observation (one user, one day), each column a variable (one thing measured about it). 4.1
Degrees of freedom
(1) In $s^2$: the $n$ deviations from $\bar x$ must add to zero, so only $n-1$ are free. (2) In the Student-t: the parameter $\nu$ that sets how heavy the tails are. 4.5 4.9
Delta method
An approximation for the spread of a transformed variable: $Var(g(Y))\approx g'(\mu)^2\,Var(Y)$. It tells you which transform makes the spread constant. 4.18
Density (PDF)
A curve $f(x)$ whose areas are probabilities: $P(a\le X\le b)=\int_a^b f(x)\,dx$. Its height is probability per unit of $x$ and can be above 1. 4.4
Dirichlet distribution
A distribution over probability vectors (shares that are ≥ 0 and add to 1), with mean $\alpha_k/\alpha_0$. The usual prior for category shares; with $K=2$ it is the Beta. 4.11 · The Dirichlet distribution 4.11 · Dirichlet: means, α₀, marginals
Discrete random variable
A random variable whose values can be listed (counts, labels). Single values can have positive probability, given by a PMF. 4.4
Disjoint (mutually exclusive) events
Events that cannot happen together: $A\cap B=\emptyset$. If both have positive probability they are dependent, not independent. 4.2 · Events: or, and, not 4.2 · Multiplication rule (and)
Dispersion index
Variance divided by mean for counts: 1 for a Poisson, well above 1 when the counts are overdispersed. 4.8
Distribution map
The web of links between distributions: special cases, sums, limits, mixtures, transformations and conjugate pairs (for example Gamma–Poisson = Negative Binomial). 4.11
Effective sample size ($n_{\text{eff}}$)
How many independent values your correlated data are worth. Positive autocorrelation makes $n_{\text{eff}}$ smaller than $n$, so the true standard error is larger than $\sigma/\sqrt n$. 4.13 · Independence & finite variance 4.13 · Monte Carlo estimates
Efficiency
How precise an estimator is for a given amount of data. On clean Normal data the median has about 64% of the mean's efficiency; robustness is insurance, not free accuracy. 4.14
Estimand
The exact quantity you are trying to estimate. Winsorizing revenue changes it from "mean revenue" to "mean capped revenue". 4.14
Estimator vs estimate
An estimator is a recipe (like $\bar X$) and is itself a random variable with a sampling distribution; an estimate is the number the recipe gave on your data (like $\bar x=4.2$). 4.4 4.1
Event
A set of outcomes, such as "the two dice sum to 7". "A or B" is the union $A\cup B$ (at least one happens); "A and B" is the intersection $A\cap B$ (both happen). 4.2
Excess kurtosis
$m_4/m_2^2-3$: how heavy the tails are compared with a Normal (Normal 0, Laplace 3, Uniform −1.2). It is not "peakedness". 4.14
Expected value $E[X]$
The probability-weighted average of all possible values, $\sum x\,p(x)$ or $\int x f(x)\,dx$: the balance point of the distribution and the long-run average. 4.5 · Expected value: the balance point 4.5 · Long-run average & continuous E[X]
Experiment, outcome, sample space
An experiment is any process with an uncertain result; an outcome $\omega$ is one possible result; the sample space $\Omega$ is the set of all outcomes. Counting principle: $m$ options then $n$ options give $m\times n$ outcomes (two dice: 36). 4.2
Exponential distribution
The waiting time until the next event when events arrive at a constant rate $\lambda$: $P(T\gt t)=e^{-\lambda t}$, mean $1/\lambda$. It is memoryless. 4.10
Feature (predictor, covariate)
A variable used as a clue to explain or predict the target, usually written $x$. 4.1
Frequentist view
Probability means long-run frequency; the parameter is a fixed unknown and the data are random; a method is judged by how it behaves over many imagined repeats. 4.1 4.2
Gamma distribution
A positive, right-skewed distribution with shape $\alpha$ and rate $\beta$: mean $\alpha/\beta$, variance $\alpha/\beta^2$. The sum of $\alpha$ exponential waits, and a standard prior for positive rates and scales. 4.10 · Gamma: shape and rate 4.10 · Gamma priors and gamma–Poisson
Gamma–Poisson mixture
A Poisson count whose rate varies by a Gamma distribution. The result is the Negative Binomial, with $Var=\mu+\mu^2/\alpha$. 4.8
Geometric distribution
The number of failures before the first success; the discrete cousin of the Exponential and the Negative Binomial with $r=1$. SciPy and NumPy count trials instead (one more). 4.8 4.10
Geometric mean
$e^{\text{mean of the logs}}$. For log-normal data it estimates the median, not the mean. 4.18 4.10
Global scaler
One mean and one sd used to standardize every group. Group differences survive and convert back; per-group scaling would erase them. 4.18
Hazard rate
$h(t)=p(t)/P(T\gt t)$: the chance per unit time that the event happens now, given it has not happened yet. Constant for the Exponential. 4.10
Heavy and light tails
Heavy tails: extreme values happen far more often than under a Normal with the same sd (Student-t, Laplace). Light tails: fewer extremes (Uniform). On a Q-Q plot: an S-shape and a reversed S. 4.9 4.14 4.17
Hierarchical model
A model where group parameters come from a shared distribution, $\theta_g\sim N(\mu,\tau^2)$, and each group's data from its own parameter. Then $Var(y)=\sigma^2+\tau^2$. 4.6
Histogram (bins and density)
Counts of values in bins, drawn as bars. Two hidden choices, the bin width and the bin origin, can change the picture. Heights can be counts, shares, or densities $c_j/(n\,w_j)$ whose bar areas add to 1. 4.16 · Histograms: counts vs density 4.16 · Bin width and bin origin
iid
Independent and identically distributed: each value is a fresh, unrelated draw from the same distribution. 4.11 4.13
Independence (events)
Knowing $B$ does not change the chance of $A$: $P(A\cap B)=P(A)P(B)$, or $P(A\mid B)=P(A)$. Several events are mutually independent only if the product rule holds for every sub-group. 4.3 4.2
Independence (random variables)
The joint distribution factorizes, $p(x,y)=p(x)\,p(y)$, so one tells you nothing about the other. It implies zero correlation, but not the reverse. 4.3 4.15
Influence
How much one data point can move an estimate. Mean: unbounded. Median: bounded. A Student-t fit: the pull falls back toward zero for very far points ("redescending"). 4.14 4.9
Interquartile range (IQR)
$Q_3-Q_1$, where the quartiles $Q_1, Q_2, Q_3$ are the 0.25, 0.5 and 0.75 quantiles: the width of the middle half of the data. For Normal data, IQR ≈ $1.349\sigma$. 4.14
Intraclass correlation (ICC)
$\tau^2/(\tau^2+\sigma^2)$: the share of the total variance that comes from real differences between groups (also the correlation of two observations from the same group). 4.6
Jensen's inequality
For a convex $g$, $E[g(X)]\ge g(E[X])$; for a concave $g$ it flips. The average of a curve is not the curve of the average. 4.5
Joint and marginal distributions
The joint distribution $p(x,y)$ gives the probabilities of pairs of values (the full two-way table). A marginal is one variable alone, found by adding over the other: $p_X(x)=\sum_y p(x,y)$. 4.6
Kendall's tau
Correlation from pairs of points: (concordant − discordant pairs) / all pairs. Concordant pairs move the same way, discordant pairs opposite ways; τ-b adjusts for ties. 4.15
Kernel
The shape of each bump in a KDE (Gaussian, Epanechnikov, box, triangle). Once the spreads are matched, the choice matters little. 4.16
Kernel density estimate (KDE)
A smooth density estimate that puts a small bump (area $1/n$) on every data point and adds them up: $\hat f_h(x)=\frac{1}{nh}\sum K\big(\frac{x-x_i}{h}\big)$. 4.16 · KDE: a bump on every point 4.16 · KDE vs histogram
Laplace distribution
$p(x)=\frac{1}{2b}e^{-|x-\mu|/b}$: a sharp peak with exponential tails and variance $2b^2$. Its maximum-likelihood centre is the median. 4.9
Laplace prior
A Laplace(0, $b$) prior on an effect, such as a changepoint slope change $\delta_j$. Its MAP is soft thresholding (like the Lasso), but the full posterior is only sparse-ish: no exact zeros. 4.9
Law of Large Numbers (LLN)
For iid draws with a finite mean, the sample average settles at $E[X]$ as $n$ grows. It works by dilution, not compensation: no outcome is ever "due" (that belief is the gambler's fallacy). Weak law: the chance of being more than $\varepsilon$ away goes to 0; strong law: the running-average path itself converges. 4.13 · Law of Large Numbers 4.13 · Proof of the weak LLN 4.2
Law of total expectation
$E[X]=E\big[E[X\mid Y]\big]$: the overall mean is the group means weighted by the group shares. 4.6
Law of total probability
$P(A)=\sum_i P(A\mid B_i)P(B_i)$ over cases $B_i$ that do not overlap and together cover everything. 4.3
Law of total variance
$Var(X)=E[Var(X\mid Y)]+Var(E[X\mid Y])$: total spread = within-group spread + between-group spread. 4.6
Likelihood
$L(\theta)=p(D\mid\theta)$, read as a function of $\theta$ with the data fixed: how well each parameter value explains the data. It is not a probability distribution over $\theta$. To choose one: support first, variance structure second, then check. 4.1 4.3 4.11
Linearity of expectation
$E[aX+bY+c]=aE[X]+bE[Y]+c$, always, even when $X$ and $Y$ are dependent. With indicators ($\mathbf 1_A$ is 1 if $A$ happens, else 0), an expected count is a sum of probabilities. 4.5
Location–scale family
Distributions of the form $X=\mu+sZ$: shifting by $\mu$ and stretching by $s$ never change the shape. The Normal, Student-t and Laplace are examples. 4.9 4.17
Log-Normal distribution
$X$ is Log-Normal if $\ln X$ is Normal. It comes from multiplying many positive factors; the median $e^\mu$ is below the mean $e^{\mu+\sigma^2/2}$. 4.10
Log transform
Replacing $y$ by $\ln y$: ratios become differences, right skew shrinks, and spread that grows with the level becomes constant. Needs $y\gt0$; with zeros use log1p, $\ln(1+y)$, where the "+1" is a real choice that depends on the units. 4.18 · The log transform 4.18 · Zeros: log1p and log(y + c)
M-estimator (and Huber loss)
An estimate found by minimizing $\sum\rho(\text{residual})$ for some loss $\rho$: squared loss gives the mean, absolute loss the median, and Huber's loss (quadratic for small residuals, linear for large ones) a compromise. 4.14
MAD (median absolute deviation)
The median of $|x_i-\text{median}|$: a robust spread. Multiply by 1.4826 to estimate $\sigma$ for Normal data. 4.14
MAP estimate
The single most probable parameter value after seeing the data: the peak of the posterior (full story in Chapter 5.2). 4.9
Markov's inequality
For a non-negative $X$, $P(X\ge a)\le E[X]/a$: the mean alone limits how often big values can occur. 4.12
Mean (sample mean $\bar x$)
The sum divided by $n$: the balance point of the data and the value that minimizes the sum of squared distances. One extreme value can move it a lot. 4.14
Median
The middle sorted value (the 0.5 quantile). It minimizes the sum of absolute distances and does not care how extreme the extremes are. 4.14 4.4
Memoryless property
$P(T\gt s+t\mid T\gt s)=P(T\gt t)$: having waited does not change the outlook. Among continuous distributions only the Exponential has it. 4.10
Method of moments
Estimating parameters by matching the sample mean and variance to the model's formulas, for example $\hat\alpha=\bar y^2/(s^2-\bar y)$ for the Negative Binomial. 4.8
Mixture
A population made of groups, $f=\sum_g w_g f_g$. Its mean and variance come from the laws of total expectation and total variance. 4.6
Mode
The most frequent value, or the highest point of a PMF or density. 4.14 4.4
Monte Carlo estimate
Estimating an expectation by averaging over random draws, $\frac1S\sum g(\theta^{(s)})$. The LLN makes it converge; the CLT gives its error, the MCSE $=\text{sd}/\sqrt S$. 4.13
Multicollinearity
Regressors that are strongly correlated with each other. Their individual coefficients become unstable (large SEs, flipping signs); the VIF measures it. 4.15
Multinomial distribution
The counts in each of $K$ categories after $n$ independent picks: $E[X_k]=n\pi_k$. The counts are negatively correlated because they share $n$. 4.7
Multiplication rule
$P(A\cap B)=P(A\mid B)\,P(B)$ always; it becomes $P(A)P(B)$ only for independent events. 4.2 4.3
Negative Binomial (NB2)
A count distribution with mean $\mu$ and variance $\mu+\mu^2/\alpha$, for overdispersed counts; it becomes Poisson as $\alpha\to\infty$. Libraries parameterize it in several different ways. 4.8 · Negative Binomial 4.8 · NB parameterizations
Nominal, ordinal, binary
Kinds of categorical variable: labels with no order (country), labels with an order (plan tier), and exactly two values (yes/no). 4.1
Normal approximation (to the Binomial)
For large $np(1-p)$, a Binomial count is close to $N(np, np(1-p))$; add 0.5 (the continuity correction) when you approximate $P(X\le k)$. For rare events (small $np$) it fails. 4.7
Normal distribution
The symmetric bell $N(\mu,\sigma^2)$ with very light tails; it appears when many small independent pushes add up. Maths writes the variance; NumPyro and SciPy take the sd. 4.9
Numeric variable (continuous, discrete)
A variable that holds real quantities: continuous (any value in a range, like minutes) or discrete (countable, like orders per day). 4.1
Odds and likelihood ratio
Odds = probability "for" / probability "against" (0.01 is 1:99). The likelihood ratio $P(B\mid A)/P(B\mid A^c)$ says how many times more likely the evidence is if $A$ is true. Posterior odds = prior odds × likelihood ratio. 4.3
One-hot encoding
Storing a category as a vector of zeros with a single 1 in its position. Then the mean of the vector is $\pi$, the category probabilities. 4.7
Outlier
A value far from the rest: a data error, a special event, or a genuine draw from a heavy tail. Flag and investigate; never delete just to make results look nicer. 4.14
Overdispersion
Counts whose variance is larger than a Poisson (or Binomial) allows, usually because the rate varies across days or users, or because events cluster. 4.8 4.7
P-P plot
Plots model probabilities $F(x_{(i)})$ against empirical ones. Like the Kolmogorov–Smirnov statistic (the largest gap between the empirical and the model CDF), it is good in the middle and weak in the tails; use Q-Q plots for tail questions. 4.17
Parameter
A number that describes the population or process (Greek letters: $\mu,\sigma,p,\theta$). It is fixed and usually unknown. 4.1
Partial correlation
The correlation between $X$ and $Y$ after removing the straight-line effect of $Z$ from both (correlate the two sets of residuals). 4.15
Partial pooling (shrinkage)
Letting each group's estimate borrow strength from the others: small, noisy groups are pulled toward the overall mean much more than large ones (full story in Chapter 6.6). 4.6
Partition
A set of cases that do not overlap and together cover everything (every user in exactly one segment). Total probability and total expectation need one. 4.3 4.6
Pearson's r
The sample correlation $r=s_{xy}/(s_x s_y)$: the direction and tightness of a straight-line pattern. Not a slope, not a percentage, and sensitive to outliers. 4.15
Plotting positions
The probability levels $p_i$ where the $i$-th sorted value is placed in a Q-Q plot, commonly $(i-0.5)/n$. Libraries differ slightly; only the extreme dots move. 4.17
PMF (probability mass function)
$p(x)=P(X=x)$ for a discrete variable: bars between 0 and 1 that add to 1. 4.4
Poisson distribution
Counts of independent events at a steady rate in a window: $P(k)=\lambda^k e^{-\lambda}/k!$, with mean = variance = $\lambda$ = rate × exposure (the window length). Independent Poisson counts add, and randomly keeping events (thinning) keeps it Poisson. 4.8 · Poisson 4.8 · Adding and splitting counts
Population
Everything you want to learn about, often an ongoing process (all future users) rather than a finite list. 4.1
Posterior
The updated distribution of a parameter after seeing the data: $p(\theta\mid D)\propto p(D\mid\theta)\,p(\theta)$. 4.1 4.3
PPCC (probability plot correlation coefficient)
The correlation of the points of a Q-Q plot: a straightness score. The reference (for example the Student-t $\nu$) with the highest PPCC is a candidate. 4.17
Prediction
A guess of a future or unseen observation. Its range is always wider than the range of a parameter estimate, because it adds the observation's own randomness. 4.1
Prior
The distribution that describes what you believe about a parameter before seeing the data, for example Beta(1, 1) on a conversion rate. 4.1 4.11
Probability
A number between 0 and 1 that obeys the three axioms. It can be read as a long-run frequency (frequentist) or a degree of belief (Bayesian); the rules are the same. 4.2 · What a probability is: axioms 4.2 · Frequency and belief
Probability vs statistics
Probability goes from a known model to the data it could produce; statistics goes from observed data back to the unknown model. 4.1
Pseudo-counts
Reading the Beta's $\alpha$ and $\beta$ (or the Dirichlet's $\alpha_k$) as imaginary prior successes and failures (or category counts). Updating adds the real counts to them. 4.11 · α, β as pseudo-counts 4.11 · Updating: pseudo + real counts
Q-Q plot
A plot of the sorted data (the order statistics) against the quantiles of a reference distribution. A straight line means the same shape; bends show skew, heavy tails, outliers or mixed groups. 4.17 · The idea of a Q-Q plot 4.17 · Build a Q-Q plot by hand 4.17 · The shape gallery
Quantile
The value with a fraction $p$ of the distribution at or below it: $q(p)=F^{-1}(p)$. The median is the 0.5 quantile; the 90th percentile is the 0.9 quantile. 4.4
Random variable
A rule that puts a number on every outcome of a random process, written with a capital letter ($X$). It carries a distribution. 4.4
Rank
A value's position in the sorted data (1 = smallest; tied values share the average rank). Spearman's correlation is Pearson's $r$ on ranks. 4.15
Realization
One observed value of a random variable (small $x$). Your dataset is one realization of the data-generating process. 4.4 4.1
Residual
Actual minus predicted, $e_t=y_t-\hat y_t$. The likelihood of a model describes its residuals, not the raw data. 4.9 4.17
Robust statistic
An estimate that a few bad or extreme values cannot push around much: median, MAD, trimmed mean, IQR. 4.14
Robust z-score
$(x-\text{median})/(1.4826\,\text{MAD})$, with values beyond about 3.5 flagged. Unlike the usual z-score, big outliers cannot hide themselves by inflating the mean and sd (masking). 4.14 · MAD (robust spread) 4.14 · Outliers & 1.5×IQR rule
Sample
The part of the population you actually observe, of size $n$. Random sampling (choosing it by chance) lets it stand in for the population. 4.1
Sample variance $s^2$ and sample sd $s$
$s^2=\frac{1}{n-1}\sum(x_i-\bar x)^2$ and $s=\sqrt{s^2}$, in the units of the data; $s$ estimates $\sigma$. NumPy, JAX and StandardScaler divide by $n$ by default (ddof=0); pandas divides by $n-1$ (ddof=1). 4.5 4.18
Sampling distribution
How a statistic (like $\bar X$) would vary over many imagined repeats of the whole data collection. The CLT describes it for averages. 4.13 · The sample mean is random 4.13 · What the CLT does not say
Sampling variation
Different samples from the same population give different numbers. Precision depends on $n$, not on the population size. 4.1
Sensitivity and specificity
Sensitivity $P(+\mid\text{condition})$: how often a test catches true cases. Specificity $P(-\mid\text{no condition})$; one minus it is the false-positive rate. 4.3
Silverman's rule
A rule-of-thumb KDE bandwidth for one-bell data: $h=0.9\min(s,\text{IQR}/1.349)\,n^{-1/5}$. SciPy's default (Scott) uses $s\,n^{-1/5}$. 4.16
Simplex
The set of probability vectors (all shares ≥ 0, adding to 1): a line segment for 2 categories, a triangle for 3. The Dirichlet lives on it. 4.11
Simpson's paradox
A comparison that holds inside every segment can reverse overall when the segment mixes differ. Compare within segments or on a common mix. 4.6
Simulation envelope
The band in which the Q-Q points of truly Normal data of the same size usually fall. Judge bends against it; about 5% of dots poke out by chance. 4.17
Skewness
Which side has the long tail: positive means a long right tail (usually mode < median < mean). The sample skewness $m_3/m_2^{3/2}$ is noisy and outlier-sensitive. 4.14
Soft thresholding
$\text{sign}(y)\max(|y|-t,0)$: shrink toward zero and set small values exactly to zero. It is the MAP under a Laplace prior with a Normal likelihood. 4.9
Spearman's rank correlation
Pearson's $r$ computed on the ranks. It measures monotone association (always up or always down), suits ordinal data and resists outliers. 4.15
Square-root transform
$\sqrt y$ for counts: it makes Poisson spread roughly constant (sd ≈ 0.5) and works at zero. Anscombe's version $\sqrt{y+3/8}$ already works from a mean of about 4. 4.18
Standard deviation ($\sigma$, $s$)
The square root of the variance: the typical distance from the mean, in the original units. $\sigma$ belongs to the process, $s$ to the sample. 4.5
Standard error (SE)
The standard deviation of an estimate across repeated samples, for example $SE(\bar X)=\sigma/\sqrt n$ (full story in Chapter 5.5). 4.13
Standardization (z-score)
$z=(x-\bar x)/s$: mean 0, sd 1. It changes the units, never the shape; fit $\bar x$ and $s$ on training data only. For Normal data $P(X\le x)=\Phi(z)$, with $\Phi$ the CDF of the standard Normal $N(0,1)$. 4.18 4.9
Statistic
Any number computed from the sample ($\bar x$, $s$, $\hat p$). It changes from sample to sample. Not the school subject. 4.1
Statistics: the six jobs
Describe (summarize the data in hand: descriptive statistics), model, estimate, infer (conclude beyond the sample, with uncertainty: inferential statistics), predict, decide (which also needs the costs of each mistake). 4.1
Student-t distribution
A bell with heavy (power-law) tails, controlled by the degrees of freedom $\nu$: Normal as $\nu\to\infty$, Cauchy at $\nu=1$. It is a Normal whose precision (one over the variance) varies by a Gamma distribution. Its scale $\sigma$ is not its sd. 4.9 · Student-t: heavy tails 4.9 · ν, missing moments, scale ≠ sd
Support
The set of values a random variable can take: $\{0,1\}$, $\{0,1,2,\dots\}$, $(0,\infty)$, $(0,1)$… Matching the support is the first step in choosing a likelihood. 4.4 4.11
Survival function
$P(X\gt x)=1-F(x)$; in SciPy, .sf(x). 4.4
Tail probability
The chance of landing far from the middle, $P(X\ge a)$ or $P(|X-\mu|\ge k\sigma)$. The mean and sd do not fix it; the shape does. 4.12
Target (response)
The variable you want to explain or predict, usually written $y$. 4.1
Trimmed mean
Drop the $\lfloor\alpha n\rfloor$ smallest and largest values and average the rest. It sits between the mean ($\alpha=0$) and the median. 4.14
Tukey fences (1.5 × IQR rule)
Flag values below $Q_1-1.5\,\text{IQR}$ or above $Q_3+1.5\,\text{IQR}$. A convention that flags about 0.7% of clean Normal data. 4.14
Unbiased estimator
An estimator whose average over many repeated samples equals the true value. $s^2$ is unbiased for $\sigma^2$, but $s$ is still slightly too low. 4.5
Uniform distribution
Every value in $[a,b]$ equally likely: mean $(a+b)/2$, variance $(b-a)^2/12$. "Uniform" is not "uninformative". If $U$ is Uniform(0, 1), then $F^{-1}(U)$ has CDF $F$ (inverse-CDF sampling). 4.10
Variance
The average squared distance from the mean, $Var(X)=E[(X-\mu)^2]=E[X^2]-\mu^2$, in squared units. 4.5
Variance of a sum
$Var(X\pm Y)=Var X+Var Y\pm2Cov(X,Y)$. For independent pieces the variances add (even for a difference); the SDs do not. 4.5
Variance-stabilizing transform
A transform that makes the spread roughly constant across levels: if sd ∝ mean$^k$, use $y^{1-k}$ (square root for Poisson counts, log for a constant CV). 4.18
VIF (variance inflation factor)
$1/(1-R_j^2)$: how much a regression coefficient's variance is inflated by the regressor's correlation with the others. Two regressors: $1/(1-r^2)$. 4.15
Winsorized mean
Replace the $\lfloor\alpha n\rfloor$ most extreme values at each end by the nearest kept value, then average all $n$. The estimand becomes the mean of the capped values. 4.14
Yeo-Johnson transform
An extension of Box-Cox that also works for zeros and negative values; its $\lambda$ is chosen by maximum likelihood (the default of sklearn's PowerTransformer). 4.18
Zero-inflated Poisson (ZIP)
A Poisson with an extra "off switch" that produces additional zeros: $P(0)=\pi+(1-\pi)e^{-\lambda}$, mean $(1-\pi)\lambda$. 4.8
Zero inflation (and hurdle models)
More zeros than the fitted count model predicts, coming from a separate mechanism (closed store, item out of stock); a different problem from overdispersion. A hurdle model is a cousin: a yes/no part decides "zero or positive", then a zero-truncated count gives the positive values. 4.8
Appendix B

Formula and notebook sheet

The key formulas of every chapter on one page, each with what it means in words and the trap that usually goes with it. Then one big table of every distribution in Chapters 4.7–4.11, and the library conventions that silently change your answer. Notation: $X$ a random variable, $x$ a value, $\mu$ and $\sigma$ population mean and sd, $\bar x$ and $s$ sample mean and sd, $N(\mu, \sigma^2)$ written with the variance.

4.1–4.3 · Statistical thinking, probability rules, Bayes

4.1 · Statistical thinking

FormulaIn wordsMain trapSee
$\theta$ vs $\hat\theta$; e.g. $p$ vs $\hat p = k/n$A parameter (Greek) describes the population: fixed, usually unknown. A statistic (Latin letter or hat) is computed from the sample and changes from sample to sample.Calling an observed rate "the true rate". "Statistic" means a number from data, not the school subject.4.1
$P(\text{data}\mid\theta)$ vs learning $\theta$ from $D$Probability goes model → data; statistics goes data → model.Mixing up the two directions.4.1
$L(\theta) = p(D\mid\theta)$The likelihood: the same formula, read as a function of $\theta$ with the data fixed.It is not a probability distribution over $\theta$.4.1
$Y_i \sim \text{Bernoulli}(\theta)$; $\;y_t = $ systematic part $+$ noiseThe data-generating process: real process → random variables → distribution → observed data. Your data are one run of it.A good fit does not prove the model is the true process.4.1
$p(\theta\mid D) \propto p(D\mid\theta)\,p(\theta)$Bayesian: $\theta$ is uncertain and gets a distribution; condition on the data you saw. Frequentist: $\theta$ fixed, data random, judge methods over repeats.The difference is not "p-values vs posteriors"; it is what is random and what probability means.4.1

4.2 · Probability: outcomes, events and rules

FormulaIn wordsMain trapSee
$P(A) = |A|/|\Omega|$Count the favourable outcomes; two dice have $6\times6 = 36$ ordered outcomes.Only for equally likely outcomes; "two outcomes" is not "50/50".4.2
$P(A)\ge0$, $\;P(\Omega)=1$, $\;P(A\cup B)=P(A)+P(B)$ if disjointThe three axioms; every other rule follows from them.For continuous outcomes, probability 0 does not mean impossible.4.2
$P(A^c) = 1 - P(A)$Complement rule: count "not A" when it is shorter.not(all) = at least one fails, not "none".4.2
$P(A\cup B) = P(A)+P(B)-P(A\cap B)$Addition rule ("or"): subtract the overlap you counted twice.A total above 1 means double counting.4.2
$P(\cup_i A_i) \le \sum_i P(A_i)$Union bound: adding chances can only overcount.Close only when overlaps are small.4.2
$P(A\cap B) = P(A)\,P(B)$Multiplication rule for independent events ("and").Mutually exclusive events are dependent, not independent.4.2
$1-(1-p)^m$Chance of at least one hit in $m$ independent tries; 20 metrics at 5% give about 0.64.$m\times p$ is only an upper bound; real metrics are correlated.4.2

4.3 · Conditional probability, independence and Bayes' theorem

FormulaIn wordsMain trapSee
$P(A\mid B) = \dfrac{P(A\cap B)}{P(B)}$B becomes the new 100%; measure A inside it.$P(A\mid B) \ne P(B\mid A)$ in general.4.3
$P(A\cap B) = P(A\mid B)P(B)$; $\;P(A_1\cap\dots\cap A_n) = P(A_1)P(A_2\mid A_1)\cdots$Chain rule: multiply along a funnel or a tree path; leaves add to 1.$P(A)P(B)$ only if independent.4.3
$P(A\cap B) = P(A)P(B)$ ⇔ $P(A\mid B) = P(A)$Independence: B is useless news about A.Independent ≠ disjoint; pairwise ≠ mutual.4.3
$p(x,y\mid z) = p(x\mid z)\,p(y\mid z)$Conditional independence: no link inside each $z$-group.It neither implies nor follows from plain independence (common cause, common effect).4.3
$P(A) = \sum_i P(A\mid B_i)P(B_i)$Total probability: a size-weighted average over non-overlapping cases.Not a plain average of segment rates.4.3
$P(A\mid B) = \dfrac{P(B\mid A)P(A)}{P(B\mid A)P(A)+P(B\mid A^c)P(A^c)}$Bayes: flip the conditional, keep the base rate. Prevalence 1%, sensitivity 90%, false positives 9% → $P(\text{sick}\mid +)\approx 9\%$.Base-rate fallacy.4.3
posterior odds = prior odds × $\dfrac{P(B\mid A)}{P(B\mid A^c)}$Odds form: each clue multiplies the odds by its likelihood ratio.A likelihood is not the probability of the hypothesis.4.3

4.4–4.6 · Random variables, expectation, variance, conditioning

4.4 · Random variables: PMF, PDF, CDF, quantiles

FormulaIn wordsMain trapSee
$X:\Omega\to\mathbb{R}$A random variable puts a number on every outcome. Capital $X$ = the random quantity; small $x$ = one value.The data are realizations, not the random variable.4.4
$p(x) = P(X=x)$, $\;\sum_x p(x) = 1$PMF: bars between 0 and 1 that add to 1.The tallest bar (mode) is not the mean.4.4
$P(a\le X\le b) = \int_a^b f(x)\,dx$PDF: area is probability; height is probability per unit of $x$.$f(x)$ can exceed 1, so continuous log-likelihoods can be positive.4.4
$F(x) = P(X\le x)$; $\;P(a\lt X\le b) = F(b)-F(a)$; $\;F' = f$CDF: a staircase (discrete) or a ramp (continuous) from 0 to 1. Survival function $P(X\gt x) = 1-F(x)$.$\lt$ vs $\le$ matters for discrete variables.4.4
$\hat F_n(t) = \#\{i: x_i\le t\}/n$Empirical CDF: the share of data at or below $t$.It only estimates $F$.4.4
$q(p) = F^{-1}(p) = \min\{x: F(x)\ge p\}$Quantile: the value with a fraction $p$ below it. Median $= q(0.5)$; IQR $= q(0.75)-q(0.25)$.Quantiles do not add; sample-quantile methods differ for small $n$.4.4

4.5 · Expected value, variance and standard deviation

FormulaIn wordsMain trapSee
$E[X] = \sum_x x\,p(x)$, $\;\int x f(x)\,dx$Balance point; long-run average. For a 0/1 variable, $E[X] = P(X=1)$.Not the most likely value; need not be a possible value (a die: 3.5).4.5
$E[aX+bY+c] = aE[X]+bE[Y]+c$Linearity: always, dependent or not. $E[\mathbf 1_A] = P(A)$, so $E[\text{count}] = \sum P(A_i)$.Not for products ($E[XY]$) or curves ($E[X^2]$).4.5
convex $g$: $E[g(X)]\ge g(E[X])$Jensen: the average of a curve is not the curve of the average (concave: $\le$).$e^{\text{mean of logs}} \ne$ mean; ratio of means ≠ mean of ratios.4.5
$Var(X) = E[(X-\mu)^2] = E[X^2]-\mu^2$Average squared distance from the mean (squared units). Bernoulli: $p(1-p)$.Far values dominate; compute with np.var, not the shortcut.4.5
$\sigma = \sqrt{Var(X)}$Typical distance from the mean, in the original units.$\sigma$ (process) ≠ $s$ (sample); Student-t scale ≠ sd.4.5
$Var(aX+b) = a^2Var(X)$, $\;SD(aX+b) = |a|\,SD(X)$Shifts never change spread; scales are squared.$Var(2X) = 4Var(X)$, not $2Var(X)$.4.5
$Var(X\pm Y) = Var X + Var Y \pm 2Cov(X,Y)$Independent pieces: variances add, even for a difference.Add variances, not SDs: SD of a difference of two 0.003s is 0.0042.4.5
$Var(\bar X) = \sigma^2/n$The average of $n$ iid values: SD shrinks like $1/\sqrt n$.Needs independence.4.5
$s^2 = \frac{1}{n-1}\sum(x_i-\bar x)^2$Unbiased: $E[s^2] = \sigma^2$; the ÷$n$ version averages $\frac{n-1}{n}\sigma^2$.$s$ is still slightly low; NumPy/JAX default ddof=0, pandas ddof=1.4.5

4.6 · Total expectation and total variance

FormulaIn wordsMain trapSee
$p(x\mid y) = p(x,y)/p_Y(y)$; $\;p_X(x) = \sum_y p(x,y)$Conditional = keep one row, divide by its total. Marginal = row or column totals.Joint ≠ conditional; $p(x\mid y)\ne p(y\mid x)$.4.6
$E[X\mid Y=y]$ (a number); $\;E[X\mid Y] = h(Y)$ (random)One mean per group; "replace each point by its group mean".It is random only through $Y$.4.6
$E[X] = E\big[E[X\mid Y]\big] = \sum_y P(Y=y)\,E[X\mid Y=y]$Total expectation: overall mean = group means weighted by group shares.A plain average of group means is wrong when sizes differ.4.6
overall rate $= \sum_s w_s\,r_s$Simpson's paradox: different mixes $w_s$ can reverse every within-segment comparison."Better in every segment" does not imply "better overall".4.6
$Var(X) = E[Var(X\mid Y)] + Var(E[X\mid Y])$Total variance = within groups + between groups. Data version: $SST = SSW + \sum n_g(\bar x_g-\bar x)^2$.Do not forget the between part.4.6
$\sum w_g\sigma_g^2 + \sum w_g(\mu_g-\mu)^2$Variance of a mixture of segments.Same mean and variance ≠ same shape.4.6
$Var(Y) = E[\lambda] + Var(\lambda)$A Poisson with a varying rate is overdispersed; Gamma rates give the NB2.Variance is always at least the mean.4.6
$Var(y) = \sigma^2+\tau^2$; $\;Var(\bar y_g) = \tau^2+\sigma^2/n_g$; $\;\text{ICC} = \frac{\tau^2}{\tau^2+\sigma^2}$Hierarchical Normal: real group differences ($\tau^2$) plus noise; small groups' averages are noisiest.Per-group z-scoring erases $\tau^2$; use a global scaler.4.6

4.7–4.11 · Distribution formulas

4.7 · Bernoulli, Binomial, Categorical, Multinomial

FormulaIn wordsMain trapSee
support → parameters → mean → variance → shape → assumptions → use → relationshipsThe eight questions to ask of any distribution.A formula is only as good as its assumptions.4.7
$P(X=1) = p$; $\;E = p$, $\;Var = p(1-p)\le 0.25$Bernoulli: one yes/no trial; the mean of a 0/1 column is the observed rate.Variance peaks at $p = 0.5$.4.7
$P(X=k) = \binom nk p^k(1-p)^{n-k}$; $\;E = np$, $\;Var = np(1-p)$Binomial: yeses in $n$ trials; $\binom nk$ counts the orders. $\hat p = X/n$ has variance $p(1-p)/n$.Needs a fixed $n$; orders per day are not Binomial.4.7
$Var(X) = np(1-p) + 2\sum_{i\lt j}Cov(X_i,X_j)$BINS: Binary, Independent, fixed Number, Same $p$. Dependence or a drifting rate inflates the variance.Right mean ≠ right model; count users, not sessions.4.7
$P(X\le k)\approx\Phi\!\left(\frac{k+0.5-np}{\sqrt{np(1-p)}}\right)$Normal approximation with continuity correction, if $np\ge10$ and $n(1-p)\ge10$ (rule of thumb).Fails for rare events (small $np$).4.7
$P(X=k) = \pi_k$; one-hot: $E[x] = \pi$, $Cov(x_j,x_k) = -\pi_j\pi_k$Categorical: one pick among $K$ labels ($K-1$ free numbers).Never average the label codes.4.7
$\frac{n!}{\prod x_k!}\prod\pi_k^{x_k}$; $\;E[X_k] = n\pi_k$, $Cov(X_j,X_k) = -n\pi_j\pi_k$Multinomial: counts per category after $n$ picks; each $X_k\sim$ Binomial$(n,\pi_k)$.The counts are not independent; they share $n$.4.7

4.8 · Poisson and Negative Binomial

FormulaIn wordsMain trapSee
$P(Y=k) = \lambda^k e^{-\lambda}/k!$; $\;E = Var = \lambda$; $\;P(0) = e^{-\lambda}$Poisson: independent events at a steady rate; $\lambda$ = rate × window."Rare" is not the requirement; $\lambda$ includes the window length.4.8
Binomial$(n, \lambda/n)\to$ Poisson$(\lambda)$; $\;(1-\lambda/n)^n\to e^{-\lambda}$Many chances, each tiny.Needs small $p$, not just large $n$.4.8
Poisson$(\lambda_1)$ + Poisson$(\lambda_2)$ = Poisson$(\lambda_1+\lambda_2)$; thinning → Poisson$(q\lambda)$Independent counts add; keeping each event with chance $q$ scales the rate.Averages and mixtures of Poissons are not Poisson.4.8
$D = Var/\text{mean}$Dispersion index: 1 for a Poisson, above 1 = overdispersed.Judge it given the model's structure, not on raw pooled counts.4.8
$\lambda\sim$ Gamma$(\alpha, \text{rate }\alpha/\mu)$, $Y\mid\lambda\sim$ Poisson$(\lambda)$ ⇒ NB2$(\mu,\alpha)$Gamma–Poisson mixture: $E = \mu$, $Var = \mu+\mu^2/\alpha$, $P(0) = \big(\frac{\alpha}{\alpha+\mu}\big)^\alpha$; $\alpha\to\infty$ gives Poisson.Large $\alpha$ = less overdispersion (in NB2).4.8
SciPy: $n = \alpha$, $p = \frac{\alpha}{\alpha+\mu}$; NumPyro probs $= \frac{\mu}{\alpha+\mu}$, logits $= \log(\mu/\alpha)$Translating NB2 into each library (full table below).Wrong translations run silently; check the implied mean and variance.4.8
$\hat\alpha = \bar y^2/(s^2-\bar y)$ (if $s^2\gt\bar y$)Method-of-moments guess of the NB concentration.Pooled "overdispersion" can be unmodelled structure (weekdays).4.8
ZIP: $P(0) = \pi+(1-\pi)e^{-\lambda}$; mean $(1-\pi)\lambda$; var $(1-\pi)\lambda(1+\pi\lambda)$Zero inflation: an extra "off switch" makes additional zeros.Overdispersion ≠ zero inflation; compare observed zeros with the fitted $P(0)$.4.8

4.9 · Normal, Student-t, Laplace

FormulaIn wordsMain trapSee
$X = \mu + sZ$, density $\frac1s p_Z\!\left(\frac{x-\mu}{s}\right)$Location–scale family: shift by $\mu$, stretch by $s$; the shape never changes.Residual = actual − predicted; the noise model describes residuals.4.9
$p(x) = \frac{1}{\sigma\sqrt{2\pi}}e^{-(x-\mu)^2/(2\sigma^2)}$Normal $N(\mu,\sigma^2)$: the bell from many small independent pushes; log-density is a parabola.Maths writes the variance; NumPyro/SciPy/NumPy take the sd.4.9
$z = (x-\mu)/\sigma$, $\;P(X\le x) = \Phi(z)$68.3% / 95.4% / 99.7% within 1 / 2 / 3 sd; 95% within 1.96 sd.The rule is for Normal shapes only.4.9
$X+Y\sim N(\mu_1+\mu_2,\ \sigma_1^2+\sigma_2^2)$; $\;\bar X\sim N(\mu, \sigma^2/n)$Independent Normals add: variances add, the shape stays Normal.SDs do not add; correlated pieces need $+2Cov$.4.9
$p(x)\propto\big(1+\frac{1}{\nu}(\frac{x-\mu}{\sigma})^2\big)^{-(\nu+1)/2}$Student-t: power-law tails; Normal as $\nu\to\infty$, Cauchy at $\nu = 1$. $P(|X|\gt3)$ = 0.058 ($\nu=3$, scale 1) vs 0.0027 (Normal).Say "less influence", never "removes outliers".4.9
$Var = \sigma^2\frac{\nu}{\nu-2}$ ($\nu\gt2$); mean needs $\nu\gt1$sd $= \sigma\sqrt{\nu/(\nu-2)}$: 1.73σ at $\nu = 3$, 1.41σ at $\nu = 4$.The scale argument is not the sd.4.9
$p(x) = \frac{1}{2b}e^{-|x-\mu|/b}$; $\;Var = 2b^2$Laplace: sharp tip, exponential tails; the maximum-likelihood centre is the median.sd $= \sqrt2\,b$, not $b$.4.9
MAP $= \text{sign}(y)\max(|y| - s^2/b,\,0)$One noisy measurement $y$ (noise sd $s$) with a Laplace$(0,b)$ prior: soft thresholding (L1). A Normal prior instead shrinks by a fixed fraction.L1 = Laplace only at the MAP; the posterior is sparse-ish, not sparse.4.9
$P(|X-\mu|\gt3\,\text{sd})$: 0.0027 / 0.0144 / 0.0138Tails at the same sd: Normal / Laplace / Student-t(3). Penalty $-\log p$: squared / absolute / logarithmic.Same sd does not mean same tails.4.9
$w_i = \frac{\nu+1}{\nu+r_i^2/\sigma^2}$Student-t fit = weighted mean; far points get small weights. Normal → mean; Laplace → median.Robustness costs efficiency when the data really are Normal.4.9

4.10 · Exponential, Gamma, Log-Normal, Uniform

FormulaIn wordsMain trapSee
$CV = \text{sd}/\text{mean}$; $\;P(X\lt0) = \Phi(-1/CV)$How much a Normal model of positive data leaks below zero.Clipping a Normal at zero is not a fix.4.10
$p(t) = \lambda e^{-\lambda t}$, $\;P(T\gt t) = e^{-\lambda t}$Exponential: wait for the next event; mean $1/\lambda$ = sd, median $\ln2/\lambda$.NumPyro uses the rate, SciPy/NumPy the scale $1/\lambda$.4.10
$P(T\gt s+t\mid T\gt s) = P(T\gt t)$; $\;h(t) = \lambda$Memoryless; constant hazard.Wear-out or churn that depends on age breaks it.4.10
$p(x)\propto x^{\alpha-1}e^{-\beta x}$; mean $\alpha/\beta$, var $\alpha/\beta^2$, CV $1/\sqrt\alpha$Gamma(shape, rate): a sum of $\alpha$ exponential waits; mode $(\alpha-1)/\beta$ when $\alpha\gt1$.Rate vs scale: SciPy gamma(a=α, scale=1/β).4.10
$\alpha = 1/c^2$, $\;\beta = \alpha/m$; $\;$Gamma$(\alpha+\sum y,\ \beta+n)$A Gamma prior from a mean $m$ and CV $c$; conjugate update for a Poisson rate.Tiny-α "vague" Gammas are not uninformative.4.10
$\log X\sim N(\mu,\sigma^2)$: median $e^\mu$, mean $e^{\mu+\sigma^2/2}$, mode $e^{\mu-\sigma^2}$Log-Normal: products of many positive factors; mode < median < mean.$\mu,\sigma$ belong to $\log X$.4.10
mean $(a+b)/2$, var $(b-a)^2/12$; $\;F^{-1}(U)\sim F$Uniform: only the range is known; inverse-CDF sampling.SciPy is uniform(loc, scale); uniform ≠ uninformative.4.10
Gamma: $\alpha = (m/s)^2$, $\beta = m/s^2$; LogN: $\sigma^2 = \ln(1+s^2/m^2)$, $\mu = \ln m - \sigma^2/2$Match a given mean $m$ and sd $s$.Same mean and sd ≠ same tails; check log-histograms and Q-Q plots.4.10

4.11 · Beta, Dirichlet, the map, choosing a likelihood

FormulaIn wordsMain trapSee
$p\sim$ Beta$(\alpha,\beta)$: $\propto p^{\alpha-1}(1-p)^{\beta-1}$; mean $\frac{\alpha}{\alpha+\beta}$A curve of belief about an unknown rate.Binomial = counts for a known $p$; Beta = belief about $p$.4.11
mode $\frac{\alpha-1}{\alpha+\beta-2}$; $\;Var = \frac{m(1-m)}{\kappa+1}$, $\kappa = \alpha+\beta$; $\;\alpha = m\kappa$, $\beta = (1-m)\kappa$Ratio sets the centre, total sets the certainty (pseudo-counts).Beta(1, 1) = flat, not "no information"; parameters below 1 give U-shapes.4.11
$\mathbf p\sim$ Dirichlet$(\boldsymbol\alpha)$: $\propto\prod p_k^{\alpha_k-1}$; mean $\alpha_k/\alpha_0$Belief about several shares on the simplex; $K = 2$ is the Beta.Shares are tied (negatively correlated), not independent Betas.4.11
$p_k\sim$ Beta$(\alpha_k, \alpha_0-\alpha_k)$; $\;Cov(p_i,p_j) = -\frac{m_im_j}{\alpha_0+1}$One share alone is a Beta; merging categories adds their $\alpha$'s.All $\alpha_k\lt1$ is sparse (corners), not vague.4.11
Beta$(\alpha+k,\ \beta+n-k)$; $\;$Dirichlet$(\boldsymbol\alpha+\mathbf c)$Updating = adding real counts to pseudo-counts. Posterior mean $= w\cdot$prior mean $+ (1-w)\,k/n$, $w = \frac{\alpha+\beta}{\alpha+\beta+n}$.Conjugacy is a convenience, not a requirement.4.11
special case, sum, limit, mixture, transformation, conjugate pairThe six kinds of links in the distribution map (e.g. Gamma–Poisson = NB2; Normal with Gamma precision = Student-t).Limits are approximations with conditions.4.11
support → variance structure → check; $\;\hat\alpha\approx\mu^2/(Var-\mu)$Choosing a likelihood: what values can one observation take, how does the spread behave, then test it.Judge the variance given the model, not on raw data.4.11

The distribution table (Chapters 4.7–4.11)

Fifteen main distributions plus three side characters (zero-inflated Poisson, Geometric, Cauchy). The first table is the maths; the second is how to call each one in code. "Use it for" names the situation where it usually appears.

DistributionSupportParametersMeanVarianceUse it for
Bernoulli$(p)$$\{0,1\}$$0\le p\le1$$p$$p(1-p)$one conversion, one click
Binomial$(n,p)$$\{0,\dots,n\}$$n$ (known), $p$$np$$np(1-p)$conversions out of $n$ exposed users
Categorical$(\pi)$$K$ labels$\pi$ ($K-1$ free)one-hot: $\pi$one-hot: $\pi_k(1-\pi_k)$, cov $-\pi_j\pi_k$which plan one user picks
Multinomial$(n,\pi)$count vectors adding to $n$$n$, $\pi$$n\pi_k$$n\pi_k(1-\pi_k)$, cov $-n\pi_j\pi_k$plan tally per variant
Poisson$(\lambda)$$\{0,1,2,\dots\}$$\lambda\gt0$ (rate × window)$\lambda$$\lambda$events at a steady rate; count metrics
Negative Binomial (NB2)$(\mu,\alpha)$$\{0,1,2,\dots\}$mean $\mu\gt0$, concentration $\alpha\gt0$$\mu$$\mu+\mu^2/\alpha$overdispersed counts: daily demand
Zero-inflated Poisson$(\lambda,\pi)$$\{0,1,2,\dots\}$$\lambda$, extra-zero share $\pi$$(1-\pi)\lambda$$(1-\pi)\lambda(1+\pi\lambda)$extra zeros from a separate cause
Geometric$(p)$$\{0,1,2,\dots\}$ (failures before the first success)$p$$(1-p)/p$$(1-p)/p^2$NB with $r = 1$; discrete cousin of the Exponential
Normal$(\mu,\sigma^2)$$\mathbb R$$\mu$, $\sigma\gt0$$\mu$$\sigma^2$symmetric, light-tailed noise; continuous metrics
Student-t$(\nu,\mu,\sigma)$$\mathbb R$$\nu\gt0$, $\mu$, scale $\sigma\gt0$$\mu$ if $\nu\gt1$$\sigma^2\frac{\nu}{\nu-2}$ if $\nu\gt2$noise with occasional shocks (robust likelihood)
Cauchy$(x_0,\gamma)$$\mathbb R$$x_0$, $\gamma\gt0$does not existdoes not existStudent-t with $\nu = 1$; averages never settle
Laplace$(\mu,b)$$\mathbb R$$\mu$, $b\gt0$$\mu$$2b^2$sharp-peaked noise; prior on changepoint slope changes $\delta_j$
Exponential$(\lambda)$$[0,\infty)$rate $\lambda\gt0$$1/\lambda$$1/\lambda^2$waiting times; weak prior for a positive scale
Gamma$(\alpha,\beta)$$(0,\infty)$shape $\alpha$, rate $\beta$$\alpha/\beta$$\alpha/\beta^2$positive amounts; priors for rates, scales, precisions
Log-Normal$(\mu,\sigma)$$(0,\infty)$$\mu$, $\sigma$ of $\log X$$e^{\mu+\sigma^2/2}$$(e^{\sigma^2}-1)e^{2\mu+\sigma^2}$revenue, prices, durations (multiplicative)
Uniform$(a,b)$$[a,b]$$a\lt b$$(a+b)/2$$(b-a)^2/12$"only the range is known"; random number generation
Beta$(\alpha,\beta)$$(0,1)$$\alpha,\beta\gt0$ (pseudo-counts)$\frac{\alpha}{\alpha+\beta}$$\frac{\alpha\beta}{(\alpha+\beta)^2(\alpha+\beta+1)}$an unknown conversion rate (prior and posterior)
Dirichlet$(\boldsymbol\alpha)$probability vectors (simplex)$\alpha_1,\dots,\alpha_K\gt0$, $\alpha_0 = \sum\alpha_k$$\alpha_k/\alpha_0$$\frac{\alpha_k(\alpha_0-\alpha_k)}{\alpha_0^2(\alpha_0+1)}$unknown category shares (prior for a categorical metric)

The same distributions in code

NumPy means rng = np.random.default_rng(seed); SciPy means scipy.stats; NumPyro means numpyro.distributions as dist. The right-hand column is the trap that changes your answer without any error message. Whatever you remember, check which call your own code uses.

DistributionNumPySciPyNumPyroWatch out
Bernoullirng.binomial(1, p)bernoulli(p)Bernoulli(probs=p) (or logits=)probs vs logits argument
Binomialrng.binomial(n, p)binom(n, p)Binomial(total_count=n, probs=p)$n$ must be a known number of trials
Categoricalrng.choice(K, p=pi)(use multinomial(1, pi))Categorical(probs=pi)labels are $0,\dots,K-1$ in code
Multinomialrng.multinomial(n, pvals)multinomial(n, p)Multinomial(total_count=n, probs=p)counts share $n$ (negatively correlated)
Poissonrng.poisson(lam)poisson(mu)Poisson(rate)"rate" here is the expected count of the window
NB2 $(\mu,\alpha)$rng.negative_binomial(α, α/(α+μ))nbinom(n=α, p=α/(α+μ))NegativeBinomial2(mean=μ, concentration=α); GammaPoisson(α, rate=α/μ); NegativeBinomialProbs(α, probs=μ/(α+μ)); NegativeBinomialLogits(α, logits=log(μ/α))SciPy's $p$ is $1-$ NumPyro's probs; statsmodels' NB2 alpha $= 1/\alpha$
Zero-inflated Poissonsimulate: zero with prob. $\pi$, else Poissonnone (statsmodels has a ZeroInflatedPoisson model)ZeroInflatedPoisson(gate=π, rate=λ)gate = extra-zero probability, not $P(0)$
Geometricrng.geometric(p) counts trials: $1,2,\dots$geom(p) counts trials: $1,2,\dots$Geometric(probs=p) counts failures: $0,1,\dots$means differ by one: $1/p$ vs $(1-p)/p$
Normalrng.normal(loc=μ, scale=σ)norm(loc=μ, scale=σ)Normal(loc=μ, scale=σ)scale is the sd, not the variance
Student-tμ + σ * rng.standard_t(ν)t(df=ν, loc=μ, scale=σ)StudentT(df=ν, loc=μ, scale=σ)scale ≠ sd: sd $= \sigma\sqrt{\nu/(\nu-2)}$
Cauchyx0 + γ * rng.standard_cauchy()cauchy(loc, scale)Cauchy(loc, scale)sample means never settle
Laplacerng.laplace(loc=μ, scale=b)laplace(loc=μ, scale=b)Laplace(loc=μ, scale=b)sd $= \sqrt2\,b$
Exponentialrng.exponential(scale=1/λ)expon(scale=1/λ)Exponential(rate=λ)scale in NumPy/SciPy, rate in NumPyro
Gammarng.gamma(shape=α, scale=1/β)gamma(a=α, scale=1/β)Gamma(concentration=α, rate=β)rate vs scale: Gamma(2, 4) means 0.5 or 8
Log-Normalrng.lognormal(mean=μ, sigma=σ)lognorm(s=σ, scale=np.exp(μ))LogNormal(loc=μ, scale=σ)$\mu,\sigma$ are for $\log X$; SciPy wants $e^\mu$
Uniformrng.uniform(low=a, high=b)uniform(loc=a, scale=b-a)Uniform(low=a, high=b)SciPy's second argument is the width
Betarng.beta(α, β)beta(α, β)Beta(concentration1=α, concentration0=β)concentration1 is the "success" pseudo-count
Dirichletrng.dirichlet(alpha)dirichlet(alpha)Dirichlet(concentration=alpha)one share alone is Beta$(\alpha_k, \alpha_0-\alpha_k)$

4.12–4.13 · Bounds, the LLN and the CLT

4.12 · Markov's and Chebyshev's inequalities

FormulaIn wordsMain trapSee
$P(X\ge a)\le E[X]/a$ ($X\ge0$)Markov: the mean alone limits big values; at most $1/c$ of the mass sits at $c$ times the mean or beyond.Needs non-negative $X$; "at most", not "about".4.12
$P(|X-\mu|\ge k\sigma)\le 1/k^2$Chebyshev: at least 75% within 2 sd, 88.9% within 3, for any shape with finite variance. Proof: Markov on $(X-\mu)^2$.Usually loose (a Normal has 4.6% beyond 2 sd, not 25%); useless for $k\le1$.4.12
$P(X-\mu\ge k\sigma)\le\frac{1}{1+k^2}$One-sided version (Cantelli).Do not just halve the two-sided bound.4.12
$\mu\pm k\sigma$ w.p. $\frac{1}{2k^2}$ each, $\mu$ w.p. $1-\frac{1}{k^2}$The three-point distribution that makes Chebyshev exact."Tight" ≠ "accurate for my data".4.12
$P(|\bar X_n-\mu|\ge\varepsilon)\le\frac{\sigma^2}{n\varepsilon^2}$; $\;n\ge\frac{\sigma^2}{\delta\varepsilon^2}$Chebyshev for averages: proves the weak LLN and gives a shape-free sample size ($\sigma^2\le 1/4$ for a rate).A guarantee, not a planning tool (about 5× the Normal-based $n$ at δ = 5%).4.12

4.13 · Law of Large Numbers and Central Limit Theorem

FormulaIn wordsMain trapSee
$E[\bar X_n] = \mu$, $\;SD(\bar X_n) = \sigma/\sqrt n$The sample mean is a random variable; its sd is the standard error.Halving the error costs 4× the data; assumes independent draws.4.13
$P(|\bar X_n-\mu|\gt\varepsilon)\to0$Weak LLN (iid, finite mean). Strong LLN: the running-average path itself converges with probability 1.Works by dilution, not compensation; says nothing about speed.4.13
Cauchy: $\bar X_n\sim$ Cauchy$(0,1)$ for every $n$No mean, no LLN; the median still works.A calm running mean is not proof that the mean exists.4.13
$Z_n = \dfrac{\bar X_n-\mu}{\sigma/\sqrt n}\to N(0,1)$CLT (iid, finite $\sigma^2$): $\bar X_n\approx N(\mu,\sigma^2/n)$, $\hat p\approx N(p, p(1-p)/n)$.About the average, not the data; approximate, not exact.4.13
skew$(\bar X_n) = \gamma/\sqrt n$; kurtosis $\kappa/n$Skewness slows the CLT: Bernoulli(0.05) needs about 1 700 for skew 0.1.$n\ge30$ is a rule of thumb, not a guarantee.4.13
$Var(\bar X_n) = \frac{\sigma^2}{n}\big[1+2\sum(1-\frac kn)\rho_k\big]$; AR(1) ≈ $\frac{\sigma^2}{n}\frac{1+\varphi}{1-\varphi}$Dependence inflates the true SE; $n_{\text{eff}} = n/\text{factor}$.Count independent units, not rows.4.13
$\text{MCSE} = \text{sd}(g)/\sqrt S$; probability: $\sqrt{\hat P(1-\hat P)/S}$Monte Carlo error of a posterior summary from $S$ draws (use ESS for MCMC).10× accuracy costs 100× draws; MCSE ignores model error.4.13

4.14–4.15 · Describing data, robustness and correlation

4.14 · Descriptive and robust statistics

FormulaIn wordsMain trapSee
$\bar x$ minimizes $\sum(x_i-c)^2$; median minimizes $\sum|x_i-c|$Mean = balance point; median = middle value; mode = most frequent. Right skew: usually mode < median < mean.The median is not a "better mean"; it answers a different question.4.14
$k = \lfloor\alpha n\rfloor$ at each endTrimmed mean drops them; winsorized mean caps them at the nearest kept value.Both change the estimand for skewed data; fix $\alpha$ in advance.4.14
$s = \sqrt{\frac{1}{n-1}\sum(x_i-\bar x)^2}$; range = max − minClassical spread.np.std uses ddof=0; the range grows with $n$.4.14
IQR $= Q_3-Q_1$; Normal: IQR ≈ $1.349\sigma$Spread of the middle half; the box of a box plot.Whiskers are not min and max.4.14
MAD $= \text{median}|x_i-\text{median}|$; $\;\hat\sigma = 1.4826\,$MADRobust spread; robust z $= (x-\text{median})/(1.4826\,\text{MAD})$, flag beyond about 3.5.Raw vs scaled: SciPy raw by default, statsmodels scaled.4.14
$g_1 = m_3/m_2^{3/2}$Skewness: positive = long right tail (Exponential 2, Poisson $1/\sqrt\lambda$).Noisy and outlier-sensitive; 0 does not prove symmetry.4.14
$g_2 = m_4/m_2^2-3$Excess kurtosis: tail heaviness (Normal 0, Uniform −1.2, Laplace 3, $t_\nu$: $6/(\nu-4)$).Not "peakedness"; infinite for a $t$ with $\nu\le4$.4.14
$Q_1-1.5\,\text{IQR}$, $\;Q_3+1.5\,\text{IQR}$Tukey fences (≈ ±2.7σ for Normal data; flags about 0.7% of clean data).Flagged ≠ wrong; never delete to make results look nicer.4.14
breakdown: mean 0, trimmed $\alpha$, IQR 25%, median and MAD 50%How much bad data an estimate survives; influence: mean unbounded, median bounded, Student-t redescending.Robustness costs efficiency on clean data (median ≈ 64%).4.14
$-\log p\propto\frac{\nu+1}{2}\log(1+r^2/\nu)$Student-t likelihood: far residuals get weight about $(\nu+1)/(\nu+r^2)$.It gives extreme residuals less influence; it does not remove them.4.14

4.15 · Covariance and correlation

FormulaIn wordsMain trapSee
$Cov(X,Y) = E[(X-\mu_X)(Y-\mu_Y)] = E[XY]-E[X]E[Y]$Sign = do they move together; sample version divides by $n-1$.Size depends on units: $Cov(aX+b, cY+d) = ac\,Cov(X,Y)$.4.15
$\rho = \dfrac{Cov(X,Y)}{\sigma_X\sigma_Y} = Cov(Z_X,Z_Y)$Correlation = covariance of z-scores; always in $[-1,1]$.$r$ is not a percentage ($r^2$ is the variance explained).4.15
$r = \dfrac{\sum(x_i-\bar x)(y_i-\bar y)}{\sqrt{\sum(x_i-\bar x)^2}\sqrt{\sum(y_i-\bar y)^2}}$; slope $b = r\,s_y/s_x$Pearson: how tightly points hug a straight line.Linear only, outlier-sensitive, not a slope; plot first (Anscombe).4.15
independent ⇒ $Cov = 0$; not the reverse$Y = X^2$ is uncorrelated with a symmetric $X$ but fully dependent.Only jointly Normal variables: uncorrelated ⇔ independent.4.15
$\rho_s = 1-\dfrac{6\sum d_i^2}{n(n^2-1)}$ (no ties)Spearman = Pearson on ranks: monotone association.Resistant to outliers, not immune.4.15
$\tau = \dfrac{C-D}{n(n-1)/2}$Kendall: concordant minus discordant pairs (τ-b adjusts for ties).Smaller scale: for Normals $\tau = \frac2\pi\arcsin\rho$.4.15
$r_{XY\cdot Z} = \dfrac{r_{XY}-r_{XZ}r_{YZ}}{\sqrt{(1-r_{XZ}^2)(1-r_{YZ}^2)}}$Partial correlation: correlate the residuals after removing $Z$.Removes only linear effects; never control for a collider.4.15
$R_{ij} = \text{corr}(X_i,X_j)$; $\;VIF = 1/(1-R_j^2)$Correlation matrix and variance inflation: spot near-duplicate regressors.Pairwise views miss three-way collinearity; screen on the training window only.4.15

4.16–4.18 · Densities, Q-Q plots, transformations

4.16 · Histograms and kernel density estimation

FormulaIn wordsMain trapSee
height $= c_j/(n\,w_j)$Density histogram: bar areas add to 1; needed for overlays and unequal bins.NumPy bins are $[a,b)$ except the last.4.16
Sturges ≈ $\log_2n+1$ bins; FD width $= 2\,\text{IQR}\,n^{-1/3}$Rules of thumb for bin width (NumPy 'auto' takes the smaller width).plt.hist defaults to 10 bins; try other widths and origins.4.16
$\hat f_h(x) = \dfrac{1}{nh}\sum_{i=1}^n K\!\left(\dfrac{x-x_i}{h}\right)$KDE: one bump of area $1/n$ on every point.An estimate shaped by $h$; heights are densities.4.16
Silverman $0.9\min(s, \text{IQR}/1.349)\,n^{-1/5}$; SciPy (Scott) $s\,n^{-1/5}$Rule-of-thumb bandwidths for one-bell data; the best $h$ shrinks like $n^{-1/5}$.Bandwidth matters far more than the kernel; library names differ.4.16
$\frac{1}{nh}\sum\big[K(\frac{x-x_i}{h})+K(\frac{x+x_i}{h})\big]$, $x\ge0$Reflection fixes boundary bias for positive data (or use a log scale).Cutting the plot at 0 hides the leak but does not fix it.4.16
two peaks visible if $d\gt2\sqrt{\sigma^2+h^2}$A big bandwidth can hide a second mode.Also look at half the rule-of-thumb bandwidth.4.16
side for share $p$: $p^{1/d}$; error $\propto n^{-4/(d+4)}$Curse of dimensionality: keep KDE to 1–2 dimensions.High dimensions need structure, not more bumps.4.16

4.17 · Q-Q plots

FormulaIn wordsMain trapSee
points $(Q(p_i),\ x_{(i)})$, $\;p_i = (i-0.5)/n$Model quantiles across, sorted data up; straight line = same shape.A bell-shaped histogram says little about the tails.4.17
$Q_X(p) = \mu+\sigma\,Q_Z(p)$Shifts and stretches never bend the line: intercept ≈ centre, slope ≈ spread.Against a Student-t, the slope is the scale, not the sd.4.17
smile / frown / S / reversed SRight skew / left skew / heavy tails / light tails; lone dots = outliers; flat–step–flat = two groups.Never delete the ends of an S; model them.4.17
$P(\max\le x) = \Phi(x)^n$; $i$-th uniform order statistic ∼ Beta$(i, n+1-i)$How much Normal data wobbles; the basis of a simulation envelope.About 5% of dots poke out by chance; look for runs at the ends.4.17
PPCC $= \text{corr}(Q(p_i), x_{(i)})$Straightness score; the $\nu$ with the highest PPCC is a candidate tail weight.A candidate, not a proof.4.17
$\big(F(x_{(i)}),\ (i-0.5)/n\big)$P-P plot: probabilities on both axes; best for the middle.Use Q-Q, not P-P or a KS test, for tail questions.4.17
$r_t = (y_t-\hat y_t)/\hat\sigma$; $\;P(|r|\gt3) = 0.0027$Q-Q the standardized residuals; a Normal expects about one day per year beyond ±3.Pair it with residuals over time (changing variance also makes an S).4.17

4.18 · Transformations

FormulaIn wordsMain trapSee
$\ln(ab) = \ln a+\ln b$; $\;\beta$ on log scale = ×$e^\beta$Log: ratios become differences; sd$(\ln Y)\approx$ sd$(Y)/\mu$ stabilizes spread ∝ level.Needs $y\gt0$; does not fix left skew.4.18
$\text{log1p}(y) = \ln(1+y)$, inverse expm1A log that works at zero.The "+1" (or $c$) is a real choice that depends on units; never drop zeros.4.18
sd$(\sqrt Y)\approx0.5$; Anscombe $\sqrt{y+3/8}$Square root stabilizes Poisson counts (good for $\lambda\gtrsim10$; Anscombe from about 4).The log over-corrects Poisson counts.4.18
$Var(g(Y))\approx g'(\mu)^2\,Var(Y)$; sd ∝ mean$^k$ → $y^{1-k}$Delta method; $k=\frac12$ → $\sqrt y$, $k=1$ → $\ln y$, $k=2$ → $1/y$.Estimate $k$ as the slope of ln sd vs ln mean.4.18
$y^{(\lambda)} = \dfrac{y^\lambda-1}{\lambda}$, $\;\ln y$ at $\lambda = 0$Box-Cox: one dial of power transforms; needs $y\gt0$.Zeros or negatives need Yeo-Johnson.4.18
$\ell(\lambda) = -\frac n2\ln\hat\sigma^2(\lambda)+(\lambda-1)\sum\ln y_i$Choose $\lambda$ by maximum likelihood; 95% interval: $\ell(\lambda)\ge\ell(\hat\lambda)-1.92$.Never pick $\lambda$ by the smallest variance (the Jacobian term matters).4.18
$y\ge0$: $\frac{(y+1)^\lambda-1}{\lambda}$; $\;y\lt0$: $-\frac{(1-y)^{2-\lambda}-1}{2-\lambda}$Yeo-Johnson: Box-Cox for zeros and negatives; $\lambda = 1$ is the identity.Its $\lambda$ is not Box-Cox's $\lambda$; fit on training data only.4.18
$E[\ln Y]\le\ln E[Y]$; median $e^\mu$, mean $e^{\mu+\sigma^2/2}$; smearing $e^{\hat z}\cdot\overline{e^{\hat e}}$Back-transformation: quantiles pass through, means need a correction.$e^{\text{mean of logs}}$ is a median-type value, not the mean.4.18
$z = (x-\bar x)/s$Standardization: mean 0, sd 1; shape, order and outliers unchanged.Fit $\bar x$, $s$ on training data; StandardScaler uses ddof=0.4.18
one $\bar y$, $s$ for all groups; $\;\Delta_y = s\,\Delta_z$Global scaler: group effects survive and convert back.Per-group z-scoring erases the group effect and $\tau$.4.18

Library conventions that silently change your answer

None of these raise an error. Each one gives a number that looks fine and is wrong for what you meant.

WhatConventionSee
Variance and sd divisornp.var, np.std, jnp.var and sklearn's StandardScaler divide by $n$ (ddof=0); pandas .var(), .std() divide by $n-1$ (ddof=1).4.5 4.18
Sample quantilesnp.quantile uses the "linear" method by default ($h = (n-1)p$, interpolate); other methods differ for small $n$.4.4
Normal, Student-t, LaplaceAll libraries take the sd (Normal) or the scale (t, Laplace), never the variance. For the t, scale ≠ sd; for the Laplace, sd $= \sqrt2\,b$.4.9 4.9
Rate vs scaleExponential and Gamma: NumPyro takes the rate; SciPy and NumPy take scale $= 1/\text{rate}$.4.10
Log-Normal, UniformSciPy lognorm(s=σ, scale=np.exp(μ)) and uniform(loc=a, scale=b-a).4.10 4.10
Negative BinomialSciPy/NumPy $(n,p)$ count failures with $p = \alpha/(\alpha+\mu)$; NumPyro has four forms; statsmodels' alpha $= 1/\alpha$. Always print the implied mean and variance.4.8
GeometricSciPy and NumPy count trials ($1,2,\dots$); NumPyro counts failures ($0,1,\dots$).4.8
BetaNumPyro Beta(concentration1=α, concentration0=β): concentration1 goes with successes.4.11
Skewness, kurtosisSciPy skew is the plain $g_1$ and kurtosis is excess (Normal = 0) by default; pandas .skew() uses the adjusted formula.4.14 4.14
MADSciPy median_abs_deviation is raw unless scale='normal' (×1.4826); statsmodels' mad is already scaled; old pandas .mad() was the mean absolute deviation.4.14
KDE bandwidthSciPy gaussian_kde default (Scott) uses kernel sd $s\,n^{-1/5}$; its 'silverman' option is about $1.06\,s\,n^{-1/5}$, not the $0.9\min(\cdot)$ rule; other libraries mean different things by the same names.4.16
Histogramsplt.hist defaults to 10 bins; NumPy bins are half-open $[a,b)$ except the last.4.16 4.16
Q-Q plotting positionsstatsmodels $i/(n+1)$, SciPy probplot Filliben's positions, the textbook $(i-0.5)/n$: only the extreme dots move.4.17
Appendix C

Interview question bank

Forty questions an interviewer could ask about this guide's material while you walk them through your Bayesian A/B framework or your forecasting model. Most come straight from the chapters' "Say it right" boxes.

How to use this bank. Read the question, answer it out loud in about a minute, and only then open the model answer. A strong answer usually has four parts: (1) a one-sentence definition in plain words, (2) a picture or a small number, (3) where it lives in your project, (4) the trap you avoid. If an answer feels shaky, follow its link back to the chapter.

Statistical thinking, probability and Bayes' rule

1. What is the difference between frequentist and Bayesian statistics?

It is about what is random and what probability means, not "p-values vs posteriors". A frequentist treats the parameter (say B's true conversion rate) as a fixed unknown number and the data as random. Probability is a long-run frequency, and a method is judged by how it behaves over many imagined repeats of the experiment. A Bayesian treats the unknown parameter as uncertain, gives it a prior, and conditions on the data actually observed to get a posterior. That is why my A/B framework can report $P(\theta_B \gt \theta_A \mid D)$ directly. With lots of data and weak priors the numbers often agree, but the statements mean different things: the Bayesian one needs a prior, the frequentist one needs a fixed design. Chapter 4.1

2. A dashboard says variant B converts at 11.4%. Is that B's conversion rate?

Not exactly. B's true conversion rate is a parameter: fixed and unknown. 11.4% is a statistic: the rate observed in this sample of users. Another sample would give a slightly different number. So I say "the observed rate of B is 11.4%; it estimates B's true rate", and I always report it with its uncertainty (in my framework, the posterior of $\theta_B$). Chapter 4.1

3. Walk me through the data-generating process behind each of your projects.

A/B framework: every exposed user is a random yes/no draw, $Y_i \sim \text{Bernoulli}(\theta_v)$, where $\theta_v$ is the conversion rate of their variant. The count per variant is then Binomial, and with a Beta prior, $\theta_v$ gets a Beta posterior. Forecasting model: demand on day $t$ is a systematic part (trend + seasonality + holidays + regressors) plus noise drawn from a likelihood (Normal, Student-t or Negative Binomial). The data we see are one run of that process. A model is my written-down guess of it, and a good fit does not prove it is the true process. Chapter 4.1

4. Are mutually exclusive events independent?

No. When both have positive probability they are strongly dependent. Mutually exclusive means they cannot happen together, so $P(A\cap B) = 0$; independence would need $P(A\cap B) = P(A)P(B) \gt 0$. Knowing one happened tells you for certain that the other did not: a user who picked the basic plan did not pick the pro plan. Chapter 4.2

5. Suppose you track 20 metrics at α = 0.05 and one comes out significant. Is it a real effect?

Not necessarily. If none of the 20 metrics really changed and the tests were independent, the chance of at least one false positive is $1 - 0.95^{20} \approx 0.64$. One hit out of 20 is what pure noise usually produces. I would name one primary metric in advance and control the family-wise error rate (Bonferroni, Holm) or the false discovery rate (Benjamini–Hochberg) across the rest (Chapter 5.11). Real metrics are correlated, so 0.64 is a guide, not an exact number. Chapter 4.2

6. A p-value is 0.03. Is there a 3% chance that the null hypothesis is true?

No. A p-value is computed assuming the null is true: it is about the chance of data at least this extreme given H0. The chance that H0 is true given the data is the reverse conditional, and getting it needs a prior through Bayes' theorem. It is the same trap as a screening test: with prevalence 1%, sensitivity 90% and a 9% false-positive rate, $P(\text{condition}\mid +) = 0.009/0.0981 \approx 9\%$, not 90%. $P(A\mid B) \ne P(B\mid A)$. Chapter 4.3

7. Explain independence, zero correlation and conditional independence. Which one does your model assume?

Independence: the joint distribution factorizes, $p(x,y) = p(x)p(y)$, so one variable tells you nothing about the other. Zero correlation: only "no straight-line link"; $Y = X^2$ with a symmetric $X$ is uncorrelated but completely determined by $X$. Conditional independence: independent inside each level of a third variable, $p(x,y\mid z) = p(x\mid z)p(y\mid z)$. A common cause makes two things dependent overall but independent given the cause. My Bayesian models assume observations are conditionally independent given the parameters, which is weaker than independent outright: users of one segment share a rate, so they are correlated overall. Chapter 4.3

Random variables, means and variances

8. What is a random variable, and how is it different from your data?

A random variable is a rule that puts a number on every outcome of a random process; it carries a distribution. The data are realizations of it. I write capital $Y_i \sim \text{Bernoulli}(\theta)$ for the conversion of user $i$ and small $y_i$ for the 0 or 1 in my table. The same distinction separates an estimator like $\bar X$, which has a sampling distribution, from my estimate $\bar x$, which is a single number. Chapter 4.4

9. Suppose your model's log-likelihood comes out at +350. Is something broken?

Not necessarily. For a continuous likelihood, each term is a log density, not a log probability. A density can be larger than 1 (a Normal with a small σ, near its mean), so its log is positive. Probabilities are areas under the density. Positive log-likelihoods are normal for continuous models, and their value depends on the units of $y$. They cannot happen for discrete PMFs, such as a Poisson or Negative Binomial count likelihood. Chapter 4.4

10. Each variant's conversion-rate estimate has an SD of 0.003. What is the SD of the difference B − A?

For independent estimates, the variances add, even for a difference: $\sqrt{0.003^2 + 0.003^2} = 0.003\sqrt2 \approx 0.0042$, not 0.006. SDs only add under perfect positive correlation. In general $Var(X - Y) = Var X + Var Y - 2Cov(X,Y)$. Chapter 4.5

11. Why does the sample variance divide by n − 1?

Because the deviations are measured from $\bar x$, which was fitted to the same data, so they come out a little too small: $E\big[\sum(x_i - \bar x)^2\big] = (n-1)\sigma^2$. Dividing by $n - 1$ makes $s^2$ unbiased for $\sigma^2$. No data point is "lost": one degree of freedom is used up because the deviations must sum to zero. $s$ itself is still slightly too low on average. In code, NumPy and JAX default to ddof=0 and pandas to ddof=1. Chapter 4.5

12. Variant B wins in every segment but loses overall. How can that happen?

Simpson's paradox. The overall rate is a weighted average of segment rates, $\sum_s w_s r_s$, and the weights (the segment mix) differ between the variants. In the chapter's example B is better on desktop (22.5% vs 20%) and on mobile (6% vs 5%), but 800 of B's 1 000 users are on mobile, so overall A shows 17% and B 9.3%. Compare within segments or on a common mix, and ask why the mix differs: in a properly randomized test the mix should be similar in both arms, so a big difference is itself a warning sign (Chapter 5.11). Chapter 4.6

13. State the law of total variance and show where it appears in your hierarchical A/B model.

$Var(X) = E[Var(X\mid Y)] + Var(E[X\mid Y])$: total variance = the average variance within groups + the variance between the group means. In a hierarchical model with $\theta_g \sim N(\mu, \tau^2)$ and $y \mid \theta_g \sim N(\theta_g, \sigma^2)$, it gives $Var(y) = \sigma^2 + \tau^2$, and an observed segment average varies by $\tau^2 + \sigma^2/n_g$. Only $\tau^2$ is real segment difference; $\sigma^2/n_g$ is noise. That is why partial pooling does not shrink every segment by the same amount: small segments are pulled toward the global mean much more than large ones (Chapter 6.6). Chapter 4.6 · τ² vs σ²

Distributions and likelihood choice

14. Why are conversions Binomial, but orders per day are not?

A Binomial needs a fixed, known number of independent yes/no trials with the same $p$ (Binary, Independent, fixed Number, Same $p$), so it can never exceed $n$. "Converters out of 1 000 exposed users" fits. Orders per day have no natural $n$ and no upper limit, so I use a Poisson, or a Negative Binomial when the variance is well above the mean. Even conversions become overdispersed if the same user appears many times or the rate drifts, so I count users, not sessions. Chapter 4.7 · BINS

15. Does modelling each user as a Bernoulli give a different posterior from modelling each variant's total as a Binomial?

No, if users are independent with a common $\theta$. The Bernoulli product is $\theta^k(1-\theta)^{n-k}$; the Binomial only adds the constant $\binom nk$, which does not depend on $\theta$ and cancels when the posterior is normalized. Aggregating to $(n, k)$ per variant is simply faster. I keep user-level rows only when I need user-level covariates or the users are clearly not interchangeable. Chapter 4.7

16. How do you model a categorical metric, such as which plan a user picks?

Per user, the choice is Categorical($\pi$): one pick among $K$ options. Per variant, the tally is Multinomial($n$, $\pi$): each count is Binomial($n$, $\pi_k$), and the counts are negatively correlated because they share $n$. With a Dirichlet($\boldsymbol\alpha$) prior on $\pi$, the posterior is Dirichlet($\boldsymbol\alpha$ + counts): updating is adding the observed counts to the prior pseudo-counts. That is the Dirichlet-Multinomial model of my framework (Chapter 6.3). Chapter 4.7 · Dirichlet

17. Is the Poisson "the distribution of rare events"?

Not really. It needs events that arrive one at a time, independently, at a steady rate within a window; the count itself can be huge (visits per hour to a big site can be Poisson(5000)). "Rare" comes from one derivation: Binomial($n$, $\lambda/n$) → Poisson($\lambda$), where each tiny slice of time rarely holds an event. Its fingerprint is mean = variance = $\lambda$, the first thing I check in real counts. Chapter 4.8

18. Why did you use a Negative Binomial likelihood for demand?

Support first: demand is a count (0, 1, 2, … with no upper limit), so the Poisson family. Variance second: I compare variance and mean within comparable days (not on raw pooled counts); when the variance is well above the mean, a Poisson cannot produce it, because it forces the two to be equal. The NB2 adds a concentration $\alpha$: $Var = \mu + \mu^2/\alpha$, and it becomes the Poisson as $\alpha \to \infty$. It has a natural story: a Poisson whose rate wobbles from day to day according to a Gamma distribution. Then I check it: posterior predictive checks of the variance and of the number of zero days. Chapter 4.8 · choosing a likelihood

19. Suppose you say "an NB likelihood with dispersion 2". What exactly is the 2?

That phrase is ambiguous, because libraries disagree. I should say: "an NB2 likelihood with mean $\mu_t$ and concentration $\alpha$, so $Var(y_t) = \mu_t + \mu_t^2/\alpha$; in NumPyro, NegativeBinomial2(mean, concentration)." Then everyone knows that a larger $\alpha$ means less overdispersion. SciPy's nbinom(n, p) uses $n = \alpha$, $p = \alpha/(\alpha+\mu)$; NumPyro's NegativeBinomialProbs uses probs $= \mu/(\alpha+\mu)$; statsmodels' alpha is $1/\alpha$. I always print the implied mean and variance in code. Chapter 4.8

20. What is the difference between zero inflation and overdispersion?

Overdispersion is about spread: variance above the mean, usually because rates vary across days or users. Zero inflation is about an extra mechanism that produces zeros on top of an ordinary count (a closed store, an item out of stock). Zero inflation also raises the variance, so a variance check alone cannot tell them apart. I compare the observed share of zeros with the fitted model's $P(0)$ (for NB2, $(\alpha/(\alpha+\mu))^\alpha$). Only if there are still far too many zeros and I can name the cause do I add a zero-inflated or hurdle part. Chapter 4.8

21. Does a Student-t likelihood remove outliers?

No. Nothing is removed; every point stays in the likelihood. The Student-t assigns more probability to extreme residuals, so they have less influence on the fit than under a Normal likelihood. Under a Normal, a 10-sd residual is nearly impossible, so the trend bends toward it and σ inflates; under a $t$ it is plausible, and its weight is roughly $(\nu+1)/(\nu+r^2)$. One more detail: the $t$'s scale σ is not its sd; the sd is $\sigma\sqrt{\nu/(\nu-2)}$ for $\nu \gt 2$ (1.73σ at $\nu = 3$). Chapter 4.9 · 4.14

22. Is your Laplace prior on the changepoint slope changes the same as the Lasso?

Only at the MAP. With a Normal likelihood, the negative log posterior gains an L1 penalty $|\delta|/b$, so the posterior mode is a soft-thresholded value that can be exactly zero, like the Lasso. But the full posterior is continuous with no point mass at zero, so its mean and its draws are almost never exactly zero. In my forecasting model the Laplace prior gives sparse-ish changepoints: most $\delta_j$ shrunk close to zero, a few large. A smaller $b$ means a stronger pull toward zero (Chapter 5.3, Chapter 7.10). Chapter 4.9

23. A colleague writes "Gamma(2, 4) prior". What do you ask?

"Shape and rate, or shape and scale?" With shape 2 and rate 4 the mean is $\alpha/\beta = 0.5$ and the sd about 0.35; with shape 2 and scale 4 the mean is 8. NumPyro's Gamma(concentration, rate) takes the rate; SciPy uses gamma(a=α, scale=1/β) and NumPy gamma(shape, scale). I always state the convention and the implied mean. Chapter 4.10

24. Revenue per order is Log-Normal with parameters μ and σ. What are its median and mean?

$\mu$ and $\sigma$ belong to $\log X$, not to $X$. The median is $e^\mu$; the mean is $e^{\mu+\sigma^2/2}$, which is larger because of the long right tail (mode < median < mean). So if I compare groups on the log scale, I am comparing medians (geometric means), not means. Chapter 4.10

25. Is a Beta(1, 1) prior on a conversion rate uninformative?

It is flat on the rate, which is a choice, not "no information": the same belief is not flat on the log-odds scale (Chapter 6.2), and with very little data it still pulls the estimate toward 0.5. Its strength is about two pseudo-observations, so with thousands of users per variant the data dominate. I like the Beta because the support matches (a rate lives in [0, 1]), it is conjugate to the Binomial, and updating is counting: Beta($1+k$, $1+n-k$). Chapter 4.11 · updating

26. How do you choose a likelihood for a new metric?

Three steps. 1. Support: what values can one observation take? Binary → Bernoulli; bounded count → Binomial; unbounded count → Poisson or NB; real and symmetric → Normal or Student-t; positive and skewed → Gamma or Log-Normal; a proportion → Beta; categories → Categorical/Multinomial. 2. Variance structure: how the variance relates to the mean, how heavy the tails are, how many zeros; judge it given the model (within groups or on residuals), not on raw pooled data. 3. Check: Q-Q plots, dispersion, posterior predictive checks. Chapter 4.11

Bounds, the LLN, the CLT and Monte Carlo

27. What does Chebyshev's inequality say, and why is it useful if it is so loose?

For any distribution with finite variance, $P(|X - \mu| \ge k\sigma) \le 1/k^2$: at least 75% within 2 sd and 88.9% within 3. It comes from Markov's inequality applied to $(X-\mu)^2$. It is usually loose (a Normal has 4.6% beyond 2 sd, not 25%), but it cannot be improved without assumptions: a three-point distribution hits it exactly. Its real value is theoretical: small variance forces concentration, and applied to the sample mean it proves the weak law of large numbers. Chapter 4.12

28. What does the Central Limit Theorem say? Is n ≥ 30 enough?

For iid draws with finite variance, the standardized sample mean $(\bar X - \mu)/(\sigma/\sqrt n)$ approaches $N(0, 1)$, so $\bar X \approx N(\mu, \sigma^2/n)$. It is about the sampling distribution of the average, not the data: skewed data stay skewed. It gives no safe $n$: the skewness of the average is $\gamma/\sqrt n$, so a conversion metric with $p = 0.05$ needs about 1 700 users before the average's skewness drops to 0.1. Heavy tails need more; strong dependence or infinite variance break it. "30" is a rule of thumb for mildly skewed data. Chapter 4.13 · how large n?

29. Suppose P(B > A | D) = 0.9 was estimated from 1 000 posterior draws. How accurate is that number?

A Monte Carlo estimate is an average, so the LLN makes it converge and the CLT gives its error: $\text{MCSE} = \sqrt{\hat P(1-\hat P)/S} = \sqrt{0.09/1000} \approx 0.0095$. So it is 0.90 ± about 0.02, not accurate to 0.001. The error falls like $1/\sqrt S$: ten times more accuracy costs a hundred times more draws. For correlated MCMC draws, use the effective sample size instead of $S$ (Chapter 6.10). And MCSE says nothing about whether the model itself is right. Chapter 4.13

30. Daily demand is usually autocorrelated. What happens to the standard error of the mean?

The formula $\sigma/\sqrt n$ assumes independent draws. With positive autocorrelation, $Var(\bar X) = \frac{\sigma^2}{n}\big[1 + 2\sum(1 - k/n)\rho_k\big]$, which is larger than $\sigma^2/n$; for an AR(1) series with coefficient $\varphi$ the inflation factor is about $(1+\varphi)/(1-\varphi)$ (3 for $\varphi = 0.5$). The data are worth fewer independent values ($n_{\text{eff}} \lt n$), so naive intervals are too narrow. Count independent units, not rows (Chapter 7.3). Chapter 4.13

Describing data, robustness and relationships

31. Is the median more accurate than the mean?

It is more robust, not more accurate. The median has a 50% breakdown point and bounded influence; the mean has breakdown point 0 and unbounded influence. On clean Normal data the median is actually less precise (about 64% efficiency), and for skewed data it estimates a different quantity (the typical value, not the average). I choose by the question (total revenue needs the mean) and by the data: if extremes are errors or nuisance, a robust estimator or a heavy-tailed likelihood; if they are the signal, I model them. Chapter 4.14

32. Suppose revenue is winsorized at the 99th percentile before the analysis. What does the lift estimate mean now?

It is the lift in mean capped revenue, not in mean revenue: winsorizing changes the estimand. It trades bias for variance: smaller error bars, but biased relative to the true average whenever the extremes differ between variants. I fix the cap before the test, report the result as capped, and check that the uncapped result points the same way if the decision is about total revenue. Chapter 4.14

33. The correlation between two variables is zero. Are they unrelated?

Not necessarily. Zero correlation means no straight-line link. Independence means the whole distribution of $Y$ does not change with $X$. Independence implies zero correlation, not the other way round: take $X$ uniform on $\{-1, 0, 1\}$ and $Y = X^2$; the covariance is 0, yet $Y$ is a function of $X$. Only for jointly Normal variables does zero correlation mean independence. For model residuals: being uncorrelated with a regressor can still hide a curved pattern, so plot them. Chapter 4.15

34. Pearson or Spearman: which one, and why?

They measure different things. Pearson measures linear association on the raw values and is sensitive to outliers. Spearman is Pearson on the ranks, so it measures monotone association (y consistently goes up, or down, with x); it suits ordinal data, does not change under a log transform, and resists outliers. If they disagree a lot, I plot the data to see whether the cause is curvature or a few extreme points. Neither one says anything about causation. Chapter 4.15

35. Suppose you add many correlated regressors to your forecasting model. Will multicollinearity hurt the forecasts?

Mainly it makes the individual coefficients unstable: large standard errors, signs that flip between samples, and an arbitrary split of the effect between near-duplicates ($VIF = 1/(1-r^2)$ for two regressors; $r = 0.984$ gives about 31). Forecasts can stay fine as long as the correlation pattern between the regressors holds, and suffer when it breaks. I check the correlation matrix and VIFs on the training window only, then drop, combine, or put a shrinkage prior on correlated regressors (Chapter 7.12). Chapter 4.15

36. Users who use a feature convert much more. Does the feature drive conversion?

Correlation alone cannot tell. The same correlation appears if the feature causes conversion, if converting users explore more features (reverse causation), if a common cause such as engagement drives both (confounding), or by chance helped by shared trends. Causation is a claim about what happens when you intervene. That is exactly why we run an A/B test: randomization breaks confounding. With observational data I would adjust for the confounders I can measure and state the remaining assumptions (Chapter 5.12). Chapter 4.15

Density plots, Q-Q plots and transformations

37. In a kernel density estimate, which matters more: the kernel or the bandwidth?

The bandwidth, by far. Kernels with matched spread give nearly identical curves. A small $h$ gives a spiky curve with low bias and high variance (overfit); a large $h$ gives a smooth curve with high bias (underfit) that can hide a second peak. Silverman's rule and SciPy's default (Scott) assume roughly one-bell data, so for skewed or multimodal data I also look at smaller bandwidths. Positive data also needs a boundary fix (reflection or a log scale). Chapter 4.16

38. How do you check whether a Normal likelihood is reasonable for your forecasting model?

I look at the residuals (actual minus forecast), not the raw demand: the likelihood describes the noise left after trend, seasonality, holidays and regressors. I make a Normal Q-Q plot of the standardized residuals with a simulation envelope, count residuals beyond ±3 (a Normal expects about 0.27%, roughly one day a year), and plot the residuals over time. An S-shape with stable spread points to a Student-t; spread growing with the level points to a transform or a variance model; a few odd dates point to missing regressors. A straight Q-Q plot only shows the residuals are consistent with a Normal at this sample size; it never proves Normality. Chapter 4.17 · 4.16 · envelopes

39. A model of log revenue gives a treatment coefficient of 0.03. Did average revenue rise by 3%?

Not necessarily. 0.03 on the log scale is a multiplicative effect of about $e^{0.03} \approx 1.03$ on the geometric mean (the median, if the log-scale noise is symmetric). The arithmetic mean can change by a different amount if the treatment also changes the spread: for log-normal data the mean is $e^{\mu + \sigma^2/2}$. The mean of the log is not the log of the mean (Jensen). If the question is average revenue per user, I back-transform with a variance correction or, better, simulate from the posterior predictive and compare the means directly. Chapter 4.18

40. Why does your A/B framework use one global scaler instead of standardizing each group separately?

Per-group standardization forces every group to mean 0 and sd 1, which erases exactly the between-group differences (the treatment effect and $\tau$) the model is supposed to estimate. One global mean and sd shift and stretch every group by the same amounts, so the differences survive and convert back by multiplying by the global sd: $\Delta_y = s\,\Delta_z$. Example: A = 10, 12, 14 and B = 14, 16, 18. Global scaling keeps a difference of 1.414 sd, which is $1.414 \times 2.828 = 4$ units; per-group scaling gives 0. Chapter 4.18

Appendix D

Where to go next

This guide gave you the language: probability, random variables, distributions, and the first tools for looking at data. The other three guides use that language every day. Here is how they connect, and a short plan for the rest of the series.

How this guide feeds the next three

Each row starts from something you learned here and points to where it is used next. All links open the other guides at the right chapter.

Guide 2 · Estimation, Inference & Experiments

From this guideUsed next for
Parameter vs statistic, estimator vs estimate, $n-1$ (4.1, 4.5)Bias, variance and MSE of estimators (5.1)
Likelihood; the Laplace prior's MAP is soft thresholding (4.1, 4.9)Maximum likelihood and MAP (5.2); ridge, lasso and priors (5.3)
Population and sample, iid, $\sigma/\sqrt n$, the CLT (4.1, 4.13)Sampling biases (5.4); standard errors and the bootstrap (5.5)
$P(A\mid B) \ne P(B\mid A)$; Normal tails and $\Phi$ (4.3, 4.9)Hypothesis tests and p-values (5.6); errors and power (5.7); confidence intervals (5.8)
"At least one" $= 1-(1-p)^m$ (4.2)Multiple metrics and repeated peeking in A/B tests (5.11)
Simpson's paradox, confounders, correlation ≠ causation (4.6, 4.15)Designing A/B experiments (5.10); causal thinking and CUPED (5.12)
Correlation matrix, multicollinearity, residuals (4.15, 4.17)Linear regression (5.13); covariance matrices and the multivariate Normal (5.15); PCA (5.16)
Poisson, NB2, log link vs log transform (4.8, 4.18)Generalized linear models (5.14)

Guide 3 · Bayesian Modeling & Computation

From this guideUsed next for
Bayes' theorem; prior, likelihood, posterior (4.1, 4.3)Bayesian inference (6.1); choosing priors (6.2)
Beta, Dirichlet, pseudo-counts, conjugate updating (4.11)Beta-Binomial and Dirichlet-Multinomial models (6.3); credible intervals and $P(\theta_B \gt \theta_A \mid D)$ (6.4)
Law of total variance, $\tau^2$ vs $\sigma^2$, the global scaler (4.6, 4.18)Hierarchical models (6.5); pooling and shrinkage (6.6); centered vs non-centered (6.7)
Dispersion checks, zero counts, Q-Q plots (4.8, 4.17)Posterior predictive checks, prior sensitivity, identifiability (6.8)
LLN, CLT, Monte Carlo error, effective sample size (4.13)MCMC (6.9); HMC, NUTS, R̂ and ESS (6.10)
Densities and log densities, Jensen's inequality (4.4, 4.5)Variational inference and KL (6.11); the ELBO and SVI (6.12)
Correlation between quantities (4.15)Mean-field, full-rank and low-rank guides (6.13); your training loop and SVI vs NUTS (6.14, 6.15)

Guide 4 · Time Series & Bayesian Forecasting

From this guideUsed next for
The data-generating process of demand (4.1)Components of a series (7.2); the additive model $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$ (7.7)
Dependence breaks $\sigma/\sqrt n$ (4.13)Autocorrelation (7.3); stationarity (7.4)
The Laplace distribution and prior (4.9)Changepoints (7.8); PELT (7.9); Laplace priors on $\delta_j$ (7.10)
Correlated regressors, standardization fitted on training data (4.15, 4.18)Holidays, exogenous regressors and leakage (7.12)
Normal, Student-t, NB2, zero inflation (4.8, 4.9)Forecast likelihoods (7.13)
Quantiles and CDFs (4.4)Predictive distributions and intervals (7.14); coverage, calibration, CRPS and pinball loss (7.16)
KDE and Q-Q plots of residuals (4.16, 4.17)Residual diagnostics (7.17)
Everything aboveThe capstone: one set of ideas, two projects (7.20)

A short study plan

  1. Close this guide properly. Reread the notebook boxes of the P0 chapters (4.2–4.5, 4.7–4.11, 4.13, 4.15–4.18). Then answer the question bank out loud, and try to rebuild the distribution table from memory: support, mean, variance and the NumPyro call for each row.
  2. Guide 2, in order. Chapters 5.1–5.8 are the core of classical inference. Then read 5.10–5.12 with your A/B framework open next to you, and 5.13–5.14 as "the regression view" of your forecasting model. Read 5.15 before Guide 3's chapter on variational guides.
  3. Guide 3, your A/B framework end to end. Chapters 6.1–6.8 cover the models (conjugacy, posterior decisions, hierarchical pooling, model checking). Chapters 6.9–6.15 explain how the posterior is computed (MCMC, NUTS, VI, the ELBO, guides, your training loop, SVI vs NUTS), and 6.16–6.17 cover JAX and JIT.
  4. Guide 4, your forecasting model component by component. Chapters 7.1–7.6 give the time-series basics, 7.7–7.13 go through your model piece by piece, 7.14–7.17 cover how to evaluate and diagnose a forecast, and 7.18–7.20 tie both projects together.
  5. Every week. Run one "Code it" block from any chapter and change one number to predict the result before you run it. Write one paragraph that explains an idea from your projects to an interviewer, in plain words.

How to make it stick

  • Simulate before you trust a formula. Ten lines of NumPy (draw, compute, repeat) check almost every claim in this guide: $n-1$, $\sigma/\sqrt n$, the CLT, the NB variance, Chebyshev's bound.
  • Always name the parameterization. Write "shape and rate", "mean and concentration", "scale, not sd" next to every distribution in your notes and code, and print the implied mean and variance.
  • Plot before you summarize. A histogram, a KDE and a Q-Q plot take seconds and catch skew, heavy tails, mixed groups and data errors that a mean and an sd hide.
  • Look at residuals, not raw data, when you judge a likelihood.
  • Keep your notebook. The "Write this in your notebook" boxes, copied by hand, are the fastest revision sheet you can have.

Free and well-known resources

  • Blitzstein & Hwang, Introduction to Probability (free online) and the Harvard Stat 110 lectures: probability, random variables and distributions in depth.
  • Allen Downey, Think Stats and Think Bayes (free online): the same ideas, taught through Python simulation.
  • Richard McElreath, Statistical Rethinking (lectures free online): the best gentle path into Bayesian modeling, with many NumPyro ports of the code.
  • Gelman et al., Bayesian Data Analysis (free PDF from the authors): the reference for hierarchical models and model checking.
  • Hyndman & Athanasopoulos, Forecasting: Principles and Practice (free online): time series, evaluation and forecasting practice.
  • The SciPy scipy.stats and NumPyro distribution documentation: the final word on what each argument means.

Companion guides

This series stands on three earlier guides: the Linear Algebra guide (vectors, matrices, eigenvalues, least squares), the Calculus guide (derivatives, integrals, gradients) and the Optimization guide (gradient descent, Adam, regularization). Revisit them whenever a step feels shaky. Next stop: Guide 2 · Estimation, Inference & Experiments.