Time Series & Bayesian Forecasting
Data that arrives in order, and how to forecast it honestly. Autocorrelation and stationarity, the classical forecasting toolbox, then your Prophet-style Bayesian model piece by piece (trend and changepoints, PELT, Laplace priors, Fourier seasonality, holidays, regressors, likelihoods), how to evaluate point and probabilistic forecasts, how to diagnose residuals, how to run models in production, and a capstone that ties both of your projects together.
What is this guide about, in one sentence? Learning from the past of a series to say, with honest uncertainty, what its future might look like.
Why care? This is your forecasting project. Every term in its resume bullet (piecewise-linear trend, changepoints, PELT, Laplace priors, Fourier seasonality, holidays, exogenous regressors, Normal, Student-t and Negative Binomial likelihoods) is a chapter here, together with the questions an interviewer will ask next: "Why not ARIMA?", "How did you validate it?", "Are the intervals calibrated?", "What leaks?".
Three ways to say it:
- Picture: a series is a recipe of layers (a slow trend, a repeating weekly pattern, spikes on holidays, and noise); forecasting means continuing each layer into the future.
- Numbers: "tomorrow ≈ 120 orders" is a point forecast; "80% chance between 95 and 150, 4% chance above our capacity of 170" is a forecast you can act on.
- Slogan: the order of the rows is information. Never shuffle it, and never let the future leak into the past.
Before you start. You need correlation, the Normal, Student-t, Poisson and Negative Binomial distributions and Q-Q plots from Guide 1 (Probability & Data), regression, GLMs and Laplace priors as L1 from Guide 2 (Inference & Experiments), and posterior predictive distributions and SVI from Guide 3 (Bayesian Modeling & Computation). Each time one appears you get a one-line reminder and a link.
The four guides
1 · Probability & Data
Probability, distributions, LLN/CLT, descriptive statistics, correlation, KDE, Q-Q plots, transformations. Parts 0–4.
2 · Estimation, Inference & Experiments
Estimators, MLE/MAP, tests, intervals, A/B testing, causal thinking, regression, GLMs, PCA, t-SNE. Parts 5–9, 26, 27.
3 · Bayesian Modeling & Computation
Bayesian inference, hierarchical models, checking, MCMC, NUTS, VI, ELBO, SVI, guides, your training loop, JAX. Parts 10, 11, 19–22, 24, 25.
4 · Time Series & Bayesian Forecasting (you are here)
Time-series foundations, classical forecasting, the Prophet-style model, evaluation, diagnostics, production, capstone. Parts 12–18, 23, 28–30.
Your model, piece by piece
Try your first interactive
The roadmap
- 7.1What makes time series differentOrder, lag, horizon, origin · Module 30
- 7.2Components of a time seriesTrend, seasonality, events, noise · Module 31
- 7.3AutocorrelationACF, PACF, white noise, random walks · Module 32
- 7.4Stationarity and differencingWhat stays the same over time · Module 33
- 7.5Baselines and exponential smoothingNaive, seasonal naive, SES, Holt, Holt-Winters · Module 34
- 7.6The ARIMA familyAR, MA, ARIMA, SARIMA, time features · Module 34
- 7.7The additive forecasting modelg + s + h + Xβ + ε · Module 35
- 7.8Piecewise-linear trends and changepointsSlopes, continuity, three approaches · Modules 36–37
- 7.9PELT change-point detectionCosts, penalties, pruning · Module 38
- 7.10Laplace priors on trend changesShrinkage and sparsity · Module 39
- 7.11Fourier seasonalityHarmonics, order, aliasing · Modules 40–41
- 7.12Holidays, regressors and leakageIndicators, future availability · Modules 42–43
- 7.13Forecast likelihoodsNormal, Student-t, Negative Binomial · Modules 44–46
- 7.14Bayesian forecastingPredictive distributions, uncertainty, PPCs · Modules 47–49
- 7.15Cross-validation and accuracy metricsRolling origin, MAE, RMSE, MAPE, MASE · Modules 50–51
- 7.16Probabilistic forecast evaluationCoverage, calibration, sharpness, CRPS · Module 52
- 7.17Residual diagnosticsHeteroscedasticity, autocorrelated errors · Modules 69–71
- 7.18Complexity and alternativesBias–variance knobs, which model when · Modules 81–82
- 7.19Production statistical modelingReproducibility, monitoring, stability · Modules 83–85
- 7.20Capstone: one set of ideas, two projectsThe connecting themes and an interview drill · Part 30
What matters most (your P0 list). Decomposition (7.2), autocorrelation and stationarity (7.3–7.4), piecewise trends, changepoints, PELT and Laplace priors (7.8–7.10), Fourier seasonality (7.11), holidays, regressors and leakage (7.12), Normal vs Student-t vs Negative Binomial (7.13), rolling validation (7.15), calibration (7.16), residual diagnostics (7.17), alternatives (7.18).
How to read the symbols
| Symbol | Say it as | Meaning |
|---|---|---|
| $y_t$ | "y at time t" | The value of the series at time $t$. |
| $y_{t-k}$ | "y lagged by k" | The value $k$ steps earlier. |
| $\hat y_{T+h\mid T}$ | "forecast of y at T plus h made at T" | A forecast $h$ steps ahead from the forecast origin $T$. |
| $\rho_k$ | "rho k" | Autocorrelation at lag $k$. |
| $g(t), s(t), h(t)$ | "g, s, h of t" | Trend, seasonality and holiday components. |
| $k, m, \delta_j, s_j$ | "k, m, delta j, s j" | Base slope, intercept, slope change at changepoint $j$, and the changepoint's time. |
| $P$, $N$ | "period", "order" | Length of a seasonal cycle, and how many Fourier harmonics are used. |
| $\nu$ | "nu" | Degrees of freedom of a Student-t (small = heavy tails). |
| $\alpha$ (in NB) | "alpha", concentration | Negative Binomial dispersion: $Var = \mu + \mu^2/\alpha$. |
| $e_t$ | "residual at t" | $y_t - \hat y_t$, what the model did not explain. |
If a chapter feels too hard, the missing piece is usually in an earlier guide: autocorrelation needs correlation (Guide 1, 4.15), the trend model needs regression (Guide 2, 5.13), and Laplace priors need MAP vs posterior (Guide 2, 5.3). Follow the links and come back.
What makes time-series data different
In every earlier guide, the rows of a dataset could be shuffled without losing anything: 500 users are 500 users in any order. A time series is different. Its rows are days (or hours, or weeks), and the order of the rows is part of the information. Today depends on yesterday. This one fact changes how we plot, how we split data for testing, how we compute error bars, and what a "forecast" even means.
- Say what a time series is, read the notation $y_t$, $y_{t-1}$, $y_{t-7}$, and tell it apart from ordinary (cross-sectional) data
- See with your own eyes why shuffling a time series keeps its histogram but destroys its information
- Build lags (shifted copies) and read a lag plot
- Explain temporal dependence: why it makes forecasting possible, and why it makes the usual standard error too small
- Choose and change the frequency of a series (daily → weekly) without the classic aggregation mistakes
- Use the words forecast origin, horizon and information set precisely: $\hat y_{T+h\mid T}$
- Explain why a random train/test split leaks the future, and what to do instead (a preview of Chapter 7.15)
A time series: one value per time step, in order core
Picture two piles of paper. The first pile is 500 survey forms, one from each user of your app. If the wind scatters them and you pick them up in a different order, you have lost nothing: it is still the same 500 answers.
The second pile is a shop's diary: one page per day, saying how many orders came in. If the wind scatters these pages, you have lost something. You no longer know which day came after which, so you cannot see the growth, the busy weekends, or the slow week after a holiday. The page numbers were part of the data.
A time series is the diary: numbers that come with a time stamp, kept in time order.
Three ways to say it:
- Picture: ordinary data is a bag of marbles; a time series is a string of beads, and the string matters.
- Numbers: "orders on Monday 10 June = 126" means more when you also know Sunday had 157 and the Monday before had 120.
- Slogan: in a time series, when is part of what.
Two weeks of daily orders for a small online shop, starting on Monday 3 June 2024.
| $t$ | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| day | Mon | Tue | Wed | Thu | Fri | Sat | Sun | Mon | Tue | Wed | Thu | Fri | Sat | Sun |
| $y_t$ | 120 | 115 | 118 | 125 | 140 | 170 | 157 | 126 | 119 | 121 | 130 | 146 | 178 | 160 |
- Each column is one time step. We number them $t = 1, 2, \dots, 14$. The first value is $y_1 = 120$ (Mon 3 Jun). The last is $y_{14} = 160$ (Sun 16 Jun), so the length is $T = 14$.
- The steps are equally spaced (one day) and no day is missing: this is a regular daily series.
- Pick $t = 8$ (Mon 10 Jun): $y_8 = 126$. "Yesterday" is $y_{t-1} = y_7 = 157$. "The same weekday last week" is $y_{t-7} = y_1 = 120$.
- Change from yesterday: $y_8 - y_7 = 126 - 157 = -31$. That looks like a collapse, but it is only the usual Monday dip after a busy weekend.
- Change from last Monday: $y_8 - y_1 = 126 - 120 = +6$. This comparison is fairer: it says the shop is slowly growing.
Notice that steps 3–5 are only possible because we know the order of the days.
A time series is a sequence of observations of one quantity, each tied to a time, kept in time order:
$$y_1,\; y_2,\; \dots,\; y_T \qquad\text{(written } (y_t)_{t=1}^{T}\text{)}.$$- $t$ is the time index (day 1, day 2, …). $y_t$ ("y at time t") is the value observed at time $t$. $T$ is the last observed time.
- Regular series: the gaps between time steps are equal (every hour, every day). Most tools in this guide assume this. An irregular log of events (one row per order, at random moments) is usually aggregated into a regular series first (orders per day).
- Univariate: one quantity over time. Multivariate: several quantities on the same time grid, for example orders $y_t$ together with price and temperature $x_t$.
- Cross-sectional data (the opposite case): many units observed once, like users in an A/B test. Its rows have no natural order.
Notation from the guide's symbol table: $y_{t-k}$ is the value $k$ steps earlier than time $t$.
Why do we need it?
Many questions are about change over time: is demand growing, is Saturday busier, what happens next week? None of these can be asked if we treat the data as an unordered bag. Naming the time index makes "before", "after" and "the future" precise.
Where is it used?
Demand and sales forecasting, web traffic, server load, stock prices, sensor readings, daily conversion rates during an A/B test, energy use, and every model in this guide: ARIMA, exponential smoothing, Prophet, and your own Prophet-style model.
How is it used?
In pandas, put the time stamps in the index: pd.Series(values, index=pd.date_range("2024-06-03", periods=14, freq="D")). Sort by time, check that the spacing is regular, and decide what a missing step means before you model anything.
"A time series is just a column of numbers."
It is a column of numbers plus their time stamps, in order. Keep the dates in the index; many mistakes come from losing them (a merge that re-sorts rows, a reset_index before a shuffle).
"A missing day is a zero."
Missing (not recorded) and zero (recorded, nothing happened) are different facts. Filling a logging outage with zeros invents a crash in demand. Decide case by case: zero, missing (leave it out of the likelihood), or interpolated.
"Irregular time stamps are fine, the model will cope."
Lags ("yesterday"), seasonal periods ("7 days") and most models assume a regular grid. Aggregate an event log to a regular frequency first (you will see how in the frequency section).
In your forecasting model, $y_t$ is the value on day (or time step) $t$, and the time index is used directly: the trend $g(t)$ and the seasonality $s(t)$ are functions of $t$, the holiday term $h(t)$ is read from the calendar date, and the regressors $X_t$ must sit on the same time grid as $y_t$. In your A/B framework the data are mostly cross-sectional (users), but a daily conversion rate during an experiment is a time series, which is why novelty effects and peeking (Chapter 5.11) are time questions.
Time series: $y_1, \dots, y_T$, one value per equally spaced time step, in order. $t$ = time index, $T$ = last observed time, $y_{t-k}$ = value $k$ steps earlier.
Cross-sectional rows can be reordered; time-series rows cannot.
Trap: missing ≠ zero; keep the dates in the index; make the grid regular first.
Quick check: in the two-week table, what are $y_{13}$, $y_{12}$ and $y_6$, and which comparison tells you more about growth?
$y_{13} = 178$ (Sat 15 Jun), $y_{12} = 146$ (the day before) and $y_6 = 170$ (the Saturday before). $y_{13} - y_{12} = +32$ is mostly the usual Saturday jump; $y_{13} - y_6 = +8$ compares like with like (Saturday with Saturday), so it says more about growth.
Time order is information: why shuffling destroys it core
Take a film and cut it into single frames. Put the frames in a bag and shake it. Every frame is still there: the same faces, the same colours, the same number of frames. But the story is gone. You cannot tell who walked in, or what happens next.
A histogram of a time series is the "bag of frames". It shows which values happened and how often. A time plot is the film. Shuffling the days keeps the histogram perfectly, and destroys everything that lives in the order: the trend, the weekly rhythm, and the link between one day and the next.
Three ways to say it:
- Picture: same frames, shuffled film: nothing missing, nothing makes sense.
- Numbers: the mean and the standard deviation of a shuffled series are exactly the same; the error of "tomorrow = today" jumps (from 1.8 to 4.8 in the example below).
- Slogan: a histogram forgets time; a forecast needs it.
Six days of sign-ups in time order: 10, 12, 11, 14, 15, 17. The same six numbers shuffled: 14, 10, 17, 11, 15, 12.
- Mean of both: $(10+12+11+14+15+17)/6 = 79/6 \approx 13.17$. Shuffling cannot change a sum.
- Standard deviation of both: $s \approx 2.64$. The distances from the mean are the same numbers, only in a different order.
- Forecast each day with "tomorrow = yesterday's value" (the naive forecast). In time order the errors are $|12-10|, |11-12|, |14-11|, |15-14|, |17-15| = 2, 1, 3, 1, 2$. Their mean (the MAE, mean absolute error) is $9/5 = 1.8$.
- Shuffled order: $|10-14|, |17-10|, |11-17|, |15-11|, |12-15| = 4, 7, 6, 4, 3$. MAE $= 24/5 = 4.8$.
- Same numbers, same histogram, but yesterday predicts today 2.7 times worse after shuffling ($4.8/1.8 \approx 2.7$). The information was in the order.
Shuffling (randomly reordering) the values $y_1, \dots, y_T$:
- keeps everything about the marginal distribution, the distribution of a single value ignoring when it happened: the histogram, mean, variance, quantiles, min and max;
- destroys everything that depends on the order: trend, seasonality, the correlation between $y_t$ and $y_{t-k}$ (autocorrelation, Chapter 7.3), and therefore our ability to forecast.
A dataset is called exchangeable when every reordering of it is equally plausible, so the order carries no information. IID data (independent and identically distributed, Chapter 5.4) is exchangeable. A time series usually is not: move Saturday's 178 into a Tuesday slot and you get a week the shop never has. Some orders of the values are far more plausible than others.
Why do we need it?
Most of statistics (means, t-tests, bootstraps, K-fold cross-validation) silently assumes the rows can be reordered. Knowing that a time series cannot be reordered tells you which of those tools will quietly give wrong answers.
Where is it used?
It is the reason for time-based splits and rolling-origin evaluation (7.15), block bootstraps (resampling whole chunks of days), "shuffle=False" in sequence models, and the permutation test for "is there any time structure at all?" (compare a statistic with its value on shuffled copies).
How is it used?
Plot the series against time before anything else, never only its histogram. Keep data sorted by time. When a tool shuffles rows (a DataLoader, train_test_split, KFold(shuffle=True)), switch it off or shuffle whole windows instead of single rows.
"The histogram of daily orders tells me the main facts about the data."
It tells you which values occur, not when. Two series with identical histograms can be a steady climb or random noise. Always look at the time plot first.
"Sorting by date is cosmetic; I can sort later."
Lags, differences, rolling windows and resampling all read the rows in their current order. Compute them on unsorted data and they are silently wrong. Sort (and check for duplicate dates) first.
"So shuffling must never happen when training on time series."
Shuffling single rows destroys the lags. Shuffling whole windows (each window keeps its own internal order, for example 28 days of history and the next 7 days to predict) is fine and common when training neural forecasters, as long as every window lies before the test period.
"Each day is one observation, so the 365 days are a sample of 365 IID points."
The days are neither independent (today depends on yesterday) nor identically distributed (there is a trend and a weekly rhythm). The order carries information, so the data are not exchangeable.
Model answer: "A time series is one realisation of a process that unfolds in time. Its values are dependent and usually not identically distributed, so tools that assume IID rows, like random K-fold or a plain standard error, need time-aware versions."
Shuffling keeps the marginal distribution (histogram, mean, sd) and destroys trend, seasonality and autocorrelation.
Exchangeable = order carries no information. IID data are exchangeable; time series are not.
Trap: never judge a time series from its histogram alone; never shuffle single rows.
Quick check: you shuffle a year of daily temperatures. Which of these change: the mean, the 90th percentile, the correlation between today and yesterday, the number of days above 30 °C?
Only the correlation between today and yesterday changes (it drops to about 0). The mean, the 90th percentile and the count of hot days depend only on which values occurred, not on their order.
Lags: looking $k$ steps back core
Your phone shows "one year ago today" photos. Shops compare "this Saturday" with "last Saturday". Weather apps say "warmer than yesterday". Each of these lines up today with a moment a fixed distance back in time.
That fixed distance is called a lag. "Lag 1" means one step back (yesterday, for daily data). "Lag 7" means seven steps back (the same weekday last week). A lagged copy of a series is the series slid along the time axis, so that each row sees the value from $k$ rows earlier.
Once you have the lagged copy next to the original, you can ask: when yesterday was high, is today high too? That question is the start of everything in Chapter 7.3 (autocorrelation).
Three ways to say it:
- Picture: a second copy of the series, slid $k$ steps to the right, laid on top of the first.
- Numbers: for Mon 10 Jun, lag 1 is Sunday's 157 and lag 7 is last Monday's 120.
- Slogan: a lag is "the same series, $k$ steps ago".
Six days of sign-ups: $y_1, \dots, y_6 = 10, 12, 11, 14, 15, 17$. Build the lag columns and measure how strongly yesterday goes with today.
| $t$ | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| $y_t$ | 10 | 12 | 11 | 14 | 15 | 17 |
| $y_{t-1}$ (lag 1) | — | 10 | 12 | 11 | 14 | 15 |
| $y_{t-2}$ (lag 2) | — | — | 10 | 12 | 11 | 14 |
- The lag-1 row is the $y_t$ row pushed one place to the right. The first entry is empty: there is no day 0.
- With lag $k$ we lose the first $k$ entries, so lag 1 gives $6 - 1 = 5$ complete pairs $(y_{t-1}, y_t)$: $(10,12), (12,11), (11,14), (14,15), (15,17)$.
- Means of the two columns of pairs: yesterday $(10+12+11+14+15)/5 = 62/5 = 12.4$; today $(12+11+14+15+17)/5 = 69/5 = 13.8$.
- Distances from the means. Yesterday: $-2.4, -0.4, -1.4, 1.6, 2.6$. Today: $-1.8, -2.8, 0.2, 1.2, 3.2$.
- Products: $4.32, 1.12, -0.28, 1.92, 8.32$; sum $= 15.4$. Sums of squares: $17.2$ (yesterday) and $22.8$ (today).
- Pearson correlation of the pairs (Chapter 4.15): $r_1 = 15.4/\sqrt{17.2 \times 22.8} = 15.4/19.80 \approx 0.78$. High days follow high days.
This is exactly what pandas y.autocorr(1) returns. (Chapter 7.3 uses the standard ACF formula, which uses one overall mean and divides by $n$; on a tiny trending series like this one it gives a much smaller number, 0.37. On long series the two agree closely.)
The lag-$k$ value of a series at time $t$ is $y_{t-k}$, the value $k$ steps earlier ($k = 1, 2, \dots$). The lagged series is $y_{t-k}$ viewed as a new column next to $y_t$; its first $k$ entries are missing.
- Lag pairs: $(y_{t-k}, y_t)$ for $t = k+1, \dots, T$; there are $T - k$ of them.
- Lag plot: a scatter plot of these pairs, $y_{t-k}$ on the horizontal axis and $y_t$ on the vertical axis.
- Lag-$k$ correlation: the correlation of the lag pairs. It measures how strongly the series is linearly linked with its own past $k$ steps back (autocorrelation = correlation with itself).
- A lead is the opposite direction: $y_{t+k}$, the value $k$ steps later. As an input feature, a lead is the future and is never allowed.
- Books on ARIMA write the lag (backshift) operator $L y_t = y_{t-1}$, so $L^k y_t = y_{t-k}$ (you will meet it in Chapter 7.6).
Why do we need it?
"Today depends on the past" is only useful if we can line today up with the past. Lags turn that idea into columns we can plot, correlate and feed to a model. Without them there is no autocorrelation, no AR model and no seasonal naive forecast.
Where is it used?
The ACF and PACF (7.3), AR and ARIMA models (7.6), the seasonal naive forecast "same day last week" (7.5), lag features in gradient-boosting and neural forecasters, the Ljung–Box test on residuals (7.17), and differencing $y_t - y_{t-1}$ (7.4).
How is it used?
In pandas: y.shift(1) is lag 1, y.shift(7) is lag 7, and y.autocorr(k) is the lag-$k$ correlation. Drop the first $k$ rows (they are NaN) before fitting. Draw a lag plot with pd.plotting.lag_plot(y, lag=k) and look at its shape.
"y.shift(-1) is yesterday."
shift(1) moves values down one row, so each row sees the previous day: that is the lag. shift(-1) brings tomorrow's value into today's row: a lead. Used as a feature, a lead is the answer key, and your backtest will look perfect for the wrong reason.
"pandas autocorr and statsmodels acf are the same number."
autocorr(k) is Pearson's r of the lag pairs (each column with its own mean). acf uses one overall mean and divides by $n$. On the six-day example they give 0.78 and 0.37. Know which one a plot or a report used.
"A lag plot is only interesting when it is a straight line."
Separate blobs (a weekly pattern seen at the wrong lag), a fan (spread that grows with the level) or a few far points (outliers, holidays) are all messages. Read the shape, not only the correlation.
Your forecasting model is a regression on time, not on past values: its inputs are $t$ (for $g$ and $s$), the calendar (for $h$) and $X_t$. There is no $y_{t-1}$ column. That is a design choice with a consequence: whatever "memory" the data have beyond trend, seasonality, holidays and regressors stays in the residuals as lag correlation. Checking the residuals' lag-1 and lag-7 correlations is therefore one of your key diagnostics (7.3, 7.17).
Lag $k$: $y_{t-k}$ = y.shift(k); first $k$ rows empty; $T-k$ lag pairs.
Lag plot = scatter of $(y_{t-k}, y_t)$; lag-$k$ correlation = how strongly the series follows its own past.
Trap: shift(-1) is the future (a lead = leakage); autocorr ≠ acf on short series.
Quick check: daily data with a strong weekly pattern. At which lags do you expect the lag plot to hug the diagonal?
At lags 7, 14, 21, …: whole weeks, where every day is compared with the same weekday. At lags like 3 or 4 a busy Saturday is paired with a quiet Tuesday or Wednesday, so the cloud splits into blobs and the correlation drops (it can even turn negative).
Temporal dependence: today depends on yesterday core
If today is unusually hot, tomorrow is probably warm too. Weather has memory. So do most business series: a busy day is often followed by another busy day, because the same campaign, the same weather and the same news are still around.
This memory is a gift and a trap at the same time. A gift, because knowing yesterday narrows down today: that is the whole reason forecasting is possible. A trap, because 100 days with memory do not carry 100 independent pieces of evidence. Neighbouring days repeat each other's story, so the usual error bars come out far too narrow.
Three ways to say it:
- Picture: in a lag plot of a series with memory, a thin vertical slice ("yesterday was about 116") contains a much narrower spread of "today" values than the whole cloud.
- Numbers: with memory $\phi = 0.8$, knowing yesterday cuts the spread of today by 40%, and 100 days are worth only about 11 independent days for estimating the mean.
- Slogan: memory makes forecasts possible and naive standard errors wrong.
A simple model with memory: each day's surprise (its distance from the long-run mean 100) is 0.8 times yesterday's surprise plus a fresh random shock with standard deviation 6. In symbols, $y_t - 100 = 0.8\,(y_{t-1} - 100) + e_t$ with $e_t \sim N(0, 6^2)$.
- Yesterday was $y_{t-1} = 120$, a surprise of $+20$. Today's expected surprise is $0.8 \times 20 = 16$, so the best guess for today is $100 + 16 = 116$.
- Given yesterday, the only uncertainty left is the fresh shock: today's spread is $\sigma = 6$.
- Without knowing yesterday, the spread of a random day is larger. For this model the long-run variance is $\sigma^2/(1-\phi^2) = 36/(1 - 0.64) = 36/0.36 = 100$, so the standard deviation is $10$.
- Knowing yesterday shrinks the spread from 10 to 6: a 40% cut. That is the gift.
- The trap: for estimating the long-run mean from $n = 100$ days, a standard approximation for this model says the variance of the sample mean is about $\frac{1+\phi}{1-\phi} = \frac{1.8}{0.2} = 9$ times what the IID formula assumes. So the true standard error is about $\sqrt{9} = 3$ times the naive $s/\sqrt{n}$.
- Equivalently, the 100 days are worth about $n_{\text{eff}} = 100/9 \approx 11$ independent days. (A simulation in the Code-it block gives a ratio of about 3.1 for $n = 100$.)
A series has temporal (serial) dependence when the distribution of $y_t$ changes once you know its past:
$$p(y_t \mid y_{t-1}, y_{t-2}, \dots) \;\ne\; p(y_t).$$Read it as "the conditional distribution of today, given the past, is not the same as the distribution of a random day" (conditional distributions: Chapter 4.6).
The simplest example is the AR(1) model (autoregressive of order 1; full story in Chapter 7.6):
$$y_t - \mu = \phi\,(y_{t-1} - \mu) + e_t, \qquad e_t \overset{iid}{\sim} N(0, \sigma^2),\quad |\phi| \lt 1.$$- $\mu$ = long-run mean; $\phi$ ("phi") = memory: 0 means no memory, close to 1 means long memory; $e_t$ = the fresh shock ("innovation") of day $t$.
- Long-run variance: $Var(y_t) = \sigma^2/(1-\phi^2)$. Given yesterday: $Var(y_t \mid y_{t-1}) = \sigma^2$.
- Lag-$k$ correlation: $\phi^k$ (it fades geometrically).
- Consequence for averages (large $n$, this model): $SE(\bar y) \approx \dfrac{\sigma_y}{\sqrt n}\sqrt{\dfrac{1+\phi}{1-\phi}}$ and the effective sample size is $n_{\text{eff}} \approx n\,\dfrac{1-\phi}{1+\phi}$. With $\phi \gt 0$ the IID formula $\sigma_y/\sqrt n$ (Chapter 5.5) is too small.
Why do we need it?
Dependence is what a forecaster exploits and what an inference procedure must respect. Without naming it, we cannot explain why yesterday helps predict today, or why a confidence interval built from 100 correlated days is far too confident.
Where is it used?
AR, ARIMA and state-space models; lag features in ML forecasters; HAC (Newey–West) standard errors and block bootstraps for correlated data; the effective sample size of an MCMC chain (Chapter 6.10), which is exactly this idea; and checks of residual autocorrelation (7.17).
How is it used?
Measure it with the lag-1 correlation or the ACF. When you average or test dependent data, either model the dependence or use a time-aware error bar (block bootstrap, HAC). When you forecast, use it: recent values (or residuals) carry information about the next steps.
"Dependence is a problem to be removed."
Dependence is the signal a forecaster lives on. It only becomes a problem when a method assumes independence (an IID standard error, a random split, a plain bootstrap).
"100 daily points give the same precision as 100 users."
With positive memory, neighbouring days repeat each other. With $\phi = 0.8$, 100 days carry about as much information about the mean as 11 independent days, and a naive interval is about 3 times too narrow.
"The lag-1 correlation is near 0, so the days are independent."
A near-zero lag-1 correlation only rules out a linear link with yesterday. There can still be a lag-7 link (weekly pattern), a non-linear link, or dependence in the spread (calm and wild periods). Check several lags and the residual plots (7.17).
In your forecasting model the noise $\epsilon_t$ is assumed independent from day to day once $g$, $s$, $h$ and $X\beta$ are accounted for. If the residuals still have memory, the model treats 365 correlated days as 365 independent facts: the posterior becomes over-confident and the prediction intervals too narrow for multi-day totals. In your A/B framework the same issue appears when one user contributes many sessions or when a metric is tracked day by day: the effective sample size is smaller than the row count (Chapter 5.10).
Dependence: $p(y_t \mid \text{past}) \ne p(y_t)$. AR(1): $y_t - \mu = \phi(y_{t-1} - \mu) + e_t$; lag-$k$ correlation $\phi^k$.
Gift: knowing the past narrows today ($\sigma$ instead of $\sigma/\sqrt{1-\phi^2}$). Trap: $SE(\bar y)$ is about $\sqrt{(1+\phi)/(1-\phi)}$ times the IID value; $n_{\text{eff}} \approx n(1-\phi)/(1+\phi)$.
Quick check: an AR(1) series has $\phi = 0.5$. How many independent days are 300 days worth, and how much wider is the true standard error of the mean?
$n_{\text{eff}} \approx 300 \times (1 - 0.5)/(1 + 0.5) = 300 \times 0.5/1.5 = 100$ days. The standard error is $\sqrt{1.5/0.5} = \sqrt 3 \approx 1.73$ times the IID value. (Both are large-$n$ approximations for this model.)
Frequency: how often we record, and how to change it
Photograph a river once a second, once a day and once a month. The three photo albums tell different stories: ripples, a flood that came and went, the slow change of the seasons. Nothing is wrong with any of them; they answer different questions.
The frequency of a series is how often it is recorded: hourly, daily, weekly, monthly. Resampling changes it. Going to a coarser frequency (daily → weekly) means combining several values into one, and you must choose how: add them up, average them, or keep the last one. Patterns faster than the new step (the weekend peak, for weekly data) disappear inside each combined value.
Three ways to say it:
- Picture: zooming out on a map: streets vanish, motorways remain.
- Numbers: two weeks of daily orders become two weekly totals, 945 and 980; the Saturday peak of 170 is hidden inside the 945.
- Slogan: pick the frequency of the decision, and aggregate each quantity the way it adds up.
The two weeks from the first section: week 1 = 120, 115, 118, 125, 140, 170, 157; week 2 = 126, 119, 121, 130, 146, 178, 160.
- Orders are a flow (an amount per day), so a weekly value is a sum. Week 1: $120+115+118+125+140+170+157 = 945$. Week 2: $126+119+121+130+146+178+160 = 980$.
- As an average per day: $945/7 = 135$ and $980/7 = 140$. Same information, different units ("orders per week" vs "orders per day").
- At weekly frequency the Saturday peak and the Monday dip are gone: they live inside each weekly number. What is left is the growth from 945 to 980 ($+3.7\%$).
- A rate needs care. Suppose Saturday had 1 000 visitors converting at 6% and Sunday 3 000 visitors converting at 2%. Conversions: $0.06 \times 1000 = 60$ and $0.02 \times 3000 = 60$.
- The weekend conversion rate is total conversions / total visitors $= 120/4000 = 3\%$, not the average of the two daily rates $(6\% + 2\%)/2 = 4\%$. The plain average gives Saturday's 1 000 visitors the same weight as Sunday's 3 000.
The frequency (or sampling interval) of a regular series is the time between consecutive steps. Resampling converts a series to another frequency:
- Downsampling (to a coarser step, e.g. daily → weekly) combines the values inside each new period. The right combination depends on the kind of quantity:
- flow (orders, revenue, visits per period) → sum;
- level measured at moments (temperature, price) → mean;
- stock at the end of a period (inventory, account balance) → last value;
- rate or ratio (conversion rate) → sum the numerators and the denominators separately, then divide.
- Upsampling (to a finer step, e.g. monthly → daily) has to fill new slots by repeating or interpolating. It creates no new information.
- The seasonal period is counted in steps, so it depends on the frequency: daily data has a weekly period of 7 and a yearly period of about 365.25; weekly data a yearly period of about 52.18; monthly data 12; hourly data 24 (daily) and 168 (weekly).
- Periods are anchored to the calendar. In pandas,
"W"means weeks ending on Sunday ("W-SUN"), and month-end totals use"ME"(plain"M"raises an error in pandas 3).
Why do we need it?
Raw data rarely arrives at the frequency of the decision. Orders are logged one by one, but staff are rostered per day and stock is ordered per week. Choosing and changing the frequency correctly decides which patterns you can see and which you have averaged away.
Where is it used?
pandas resample and groupby pipelines, weekly business dashboards, choosing the seasonal periods of a Fourier seasonality (7.11) and of a seasonal naive baseline (7.5), hierarchical forecasting (daily forecasts that must add up to weekly targets), and daily conversion-rate series in experiments.
How is it used?
y.resample("W").sum() for flows, .mean() for levels, .last() for stocks, and for rates resample numerator and denominator separately and divide. Check the anchor day, drop or flag incomplete first and last periods, and set the seasonal periods in the units of the new step.
"To get a weekly conversion rate, average the seven daily rates."
Sum the conversions and the visitors separately, then divide. Averaging rates gives a quiet day the same weight as a busy day.
"The last bar of my weekly dashboard dropped 60%: demand is collapsing."
Check whether the last period is complete. A week that has only reached Wednesday has fewer (and quieter) days in it. Drop partial periods or compare per-day averages of matching weekdays.
"Upsampling monthly data to daily gives me 30 times more data."
It gives 30 times more rows and zero new information: the new values are copies or interpolations. Any error bar computed on them is fake.
"resample("W") groups Monday to Sunday by default, and "M" means month."
"W" is weeks ending Sunday, labelled with the Sunday date; use "W-MON" etc. for other anchors. In pandas 3 use "ME" for month end; "M" raises an error.
In your forecasting model the seasonal period $P$ of a Fourier seasonality is measured in the units of $t$. If $t$ counts days, the weekly term uses $P = 7$ and the yearly term $P = 365.25$ (7.11). If you ever resample to weekly data, the weekly term must go (there is nothing left to model) and the yearly period becomes about 52.18. For count likelihoods (Negative Binomial), resample counts by summing, so the values stay whole counts.
Frequency = spacing of the steps. Downsample flows by sum, levels by mean, stocks by last, rates by ratio of sums.
Seasonal period is in steps: daily → 7 and 365.25; weekly → 52.18; monthly → 12; hourly → 24 and 168.
Traps: incomplete edge periods; pandas "W" = weeks ending Sunday, "ME" not "M"; upsampling adds no information.
Quick check: you have hourly server requests and hourly average CPU load. How do you get daily values for each?
Requests are a flow: sum the 24 hourly counts. CPU load is a level measured over each hour: take the mean of the 24 values (weighted by duration if hours can be partial). Do not sum CPU percentages.
Forecast origin, horizon and the information set core
You stand on a road in the fog. Behind you, you can see every step you have taken. Ahead, you can guess the next few metres well, the next hundred metres roughly, and the next kilometre hardly at all.
The spot where you stand is the forecast origin: the last moment whose data you are allowed to use. How far ahead you are guessing is the horizon. Everything you know at the origin (all past values, plus things known in advance, like the holiday calendar or a planned price) is the information set. A forecast is always "made at the origin, for some horizon".
Three ways to say it:
- Picture: a vertical line at the origin; known road on the left, fog on the right that thickens with distance.
- Numbers: standing on Sunday 16 June ($T = 14$), the forecast for Wednesday 19 June is the 3-step-ahead forecast $\hat y_{17\mid 14}$.
- Slogan: every forecast has a "made at" date and a "for" date.
Origin $T = 14$ (Sun 16 Jun), the last day of the two-week table. We forecast the next week ($h = 1, \dots, 7$) with the seasonal naive rule "same weekday last week" (Chapter 7.5).
- Horizon $h = 1$ is day $T + 1 = 15$ (Mon 17 Jun). Last week's Monday is day $15 - 7 = 8$, so $\hat y_{15\mid 14} = y_8 = 126$.
- $h = 3$ is day 17 (Wed 19 Jun): $\hat y_{17\mid 14} = y_{10} = 121$.
- $h = 7$ is day 21 (Sun 23 Jun): $\hat y_{21\mid 14} = y_{14} = 160$.
- $h = 8$ is day 22 (Mon 24 Jun). Last week's Monday (day 15) is not known at the origin, so the rule reuses the last known week: $\hat y_{22\mid 14} = y_8 = 126$ again. In general $\hat y_{T+h\mid T} = y_{T+h-7\lceil h/7\rceil}$, where $\lceil \cdot \rceil$ rounds up.
- On Wednesday 19 June a new origin $T = 17$ is available, and the forecast for Sunday 23 June can be remade with three more days of data. Same target day, different origin, different forecast.
- Forecast origin $T$: the last time step whose observation may be used.
- Information set $\mathcal{I}_T$: everything known at the origin: $y_1, \dots, y_T$, past regressors, and future inputs that are genuinely known in advance (calendar, holidays, planned promotions).
- Horizon $h \ge 1$: how many steps ahead. The $h$-step-ahead forecast is written $\hat y_{T+h\mid T}$: "forecast of $y$ at $T + h$, made at $T$".
- Forecast window: all horizons you need, $h = 1, \dots, H$ (for example $H = 14$ days for staff rostering).
- One-step-ahead forecasts ($h = 1$) vs multi-step forecasts ($h \gt 1$): multi-step forecasts are harder because errors accumulate.
- Fitted values $\hat y_t$ for $t \le T$ are the model's values inside the training period. They are not forecasts: the model has already seen $y_t$ when it produced them.
The uncertainty of $y_{T+h}$ given $\mathcal{I}_T$ usually grows with $h$. For the AR(1) model, the forecast standard deviation is $\sigma\sqrt{1 + \phi^2 + \dots + \phi^{2(h-1)}}$; for a random walk ($\phi = 1$) it is $\sigma\sqrt{h}$.
Why do we need it?
Without an explicit origin, it is easy to use information that would not have existed at the time ("leakage"), and to report a one-step error when the business needs a 14-day-ahead forecast. The notation $\hat y_{T+h\mid T}$ forces both questions: what was known, and how far ahead.
Where is it used?
Backtesting and rolling-origin evaluation (7.15), accuracy reported per horizon, prediction intervals that widen with $h$ (7.14), staffing and inventory planning (lead times), and the "future dataframe" that Prophet-style models are evaluated on.
How is it used?
Fix the origin first, cut the data at it, and build every feature only from the information set. Produce forecasts for $h = 1..H$, store them with both dates (made-at, for), and evaluate errors separately for each horizon. Never call fitted values "forecasts".
"My model fits the training data closely, so its forecasts will be accurate."
Fitted values were produced with $y_t$ already in the training data; they measure fit, not foresight. Accuracy must be measured on forecasts made from an origin, on data after it.
"The horizon is the length of the training window."
The horizon $h$ counts steps after the origin. The training window is how much history before the origin you use. They are separate choices.
"Anything in the database at forecast time is in the information set."
Only what would truly have been known at $T$. Late-arriving or revised data (yesterday's orders still being counted, refunds booked next week) must be excluded or used in its "as-of" version.
"The model's predictions for last year are its forecasts for last year."
Those are in-sample fitted values. A forecast for a day is made from an origin before that day, using only what was known then.
Model answer: "I always state a forecast as made-at $T$, for $T + h$. Fitted values show how well the model describes the past; forecast accuracy comes from re-fitting at past origins and predicting the days after each origin, reported per horizon."
Your forecasting model produces $\hat y_{T+h\mid T}$ by extending $g(t)$, $s(t)$, $h(t)$ and $X_t\beta$ past the last training day. Seasonality and holidays are known functions of the calendar, so they extend easily. Regressors are in the information set only if they are known in advance or forecast themselves (7.12). The trend's uncertainty is what grows most with $h$; check that your predictive intervals actually widen with the horizon (7.14).
$\hat y_{T+h\mid T}$ = forecast for time $T + h$ made at origin $T$, using only the information set $\mathcal{I}_T$.
Uncertainty grows with $h$: AR(1) sd $= \sigma\sqrt{\sum_{j=0}^{h-1}\phi^{2j}}$; random walk sd $= \sigma\sqrt h$.
Trap: fitted values are not forecasts; horizon ≠ training window; use as-of data.
Quick check: on Friday evening you forecast next Tuesday's orders with data up to Friday. What are $T$, $h$, and the notation?
The origin is Friday ($T$ = Friday). Saturday is $h = 1$, Sunday $h = 2$, Monday $h = 3$, Tuesday $h = 4$. The forecast is $\hat y_{T+4\mid T}$, a 4-step-ahead forecast.
Why a random train/test split leaks the future core
Cover one day in the middle of a known week and ask a friend to guess it. Easy: they look at the days on both sides and pick something in between. Now cover next week and ask the same. Much harder: there is no "other side".
The first task is interpolation (filling a gap between known points). The second is extrapolation (going beyond the last known point). Forecasting is always extrapolation. A random train/test split hides random days, so it tests interpolation: the model gets to see days after each test day. It scores beautifully and then disappoints in real life. That is leakage: information from the future slipped into training.
Three ways to say it:
- Picture: random split = test dots scattered among training dots; time split = all test dots to the right of a wall.
- Numbers: in the widget below, a random split reports an error of about 9 orders, the honest time split about 27: three times more.
- Slogan: train on the past, test on the future, never the other way round.
Six days that grow steadily: 100, 104, 108, 112, 116, 120. The "model" predicts a hidden day from the nearest known days.
- Random split: day 4 is chosen as the test day. Its neighbours, days 3 and 5, are in training. Prediction: $(108 + 116)/2 = 112$. Actual: 112. Error: $0$.
- Time split: train on days 1–4, test on days 5 and 6. The nearest known day is day 4, so both predictions are 112.
- Errors: $|116 - 112| = 4$ and $|120 - 112| = 8$. MAE $= (4 + 8)/2 = 6$.
- The random split said "error 0", the time split "error 6". Only the second one resembles real use, where tomorrow is never in the training data.
Data leakage in forecasting: any information that would not be available at the forecast origin is used for training, feature building or model selection.
- A random split (or shuffled K-fold) puts days from after each test day into training: temporal leakage. It measures interpolation error and is too optimistic for forecasting.
- A temporal (time-based) split trains on $t \le T$ and tests on $t \gt T$.
- Rolling-origin evaluation repeats the temporal split at many origins and averages the errors per horizon (Chapter 7.15). A gap (embargo) between training and test is added when features use windows that would otherwise overlap.
- Other common leaks: scaling or normalising with statistics of the full series; selecting features, tuning hyperparameters or detecting changepoints on the full series; features built from the target's future (a lead, a centred moving average) (7.12).
Why do we need it?
A validation number is only useful if it predicts how the model will do after deployment. Random splits answer "how well can you fill gaps?", a question nobody asked. Time-based splits answer "how well can you see ahead?", the actual job.
Where is it used?
scikit-learn's TimeSeriesSplit, Prophet's cross_validation (cutoffs = origins, with an initial window and a horizon), backtesting in finance, demand-forecasting competitions (M5 used the last 28 days as the test window), and any honest comparison of your model against baselines.
How is it used?
Sort by time, choose an origin, fit everything (scaling, feature selection, changepoint detection, hyperparameters) on data up to the origin only, forecast the horizon after it, record the errors, move the origin forward, repeat. Report the average error per horizon.
"Random K-fold cross-validation is fine as long as I shuffle."
Shuffling is the problem: each fold trains on days after its test days. Use time-ordered folds (expanding or sliding windows, 7.15).
"My model has no lag features, so a random split cannot leak."
It still leaks: the training days right after a test day tell the model the level, the trend and any jump around that day. In the widget the "model" has no lag features either.
"I standardised the whole series first and then split it; that is harmless."
The mean and standard deviation of the full series contain the test period (for example its higher level after a jump). Fit any scaler, detector or selector on the training part only, then apply it to the test part.
"I validated my forecasting model with 5-fold cross-validation and got an RMSE of 4."
"I validated it with rolling-origin backtests: at several origins I refit on the past only and forecast the next $H$ days, and I report the error per horizon."
Model answer: "Random folds leak the future: the model sees days on both sides of each test day, which turns forecasting into interpolation and makes the error look too small. Time-ordered splits are the only honest test. I also make sure everything fitted from data (scalers, changepoints, hyperparameters) is fitted inside each training window."
For your forecasting model, every backtest should refit (or at least re-scale and re-detect) using only data up to each origin. Two places where leakage hides in a model like yours: PELT changepoint detection run once on the full history (it will place changepoints using the test period, 7.9), and any scaling of $y$ or $X$ computed on the full series. Both make the backtest look better than live performance.
Forecasting = extrapolation. Random splits test interpolation → too optimistic (temporal leakage).
Train on $t \le T$, test on $t \gt T$; repeat over origins (rolling origin, 7.15).
Trap: fit scalers, feature selection, changepoint detection and tuning inside the training window only.
Quick check: why can a random split look good even for a model that cannot forecast at all?
Because the test days have training days just before and just after them. A model that only averages nearby days (no understanding of the future) fills those gaps well. That skill disappears as soon as all test days lie after the last training day.
Recap, cheat sheet and practice
- A time series $y_1, \dots, y_T$ is one value per equally spaced time step, in order. Unlike cross-sectional data, its rows cannot be reordered: shuffling keeps the histogram, mean and sd, and destroys trend, seasonality and autocorrelation.
- A lag $y_{t-k}$ (
shift(k)) lines today up with $k$ steps ago. The lag plot and the lag-$k$ correlation show how strongly a series follows its own past. A lead (shift(-k)) is the future. - Temporal dependence: $p(y_t \mid \text{past}) \ne p(y_t)$. It makes forecasting possible (knowing yesterday narrows today) and makes IID standard errors too small ($n_{\text{eff}} \approx n(1-\phi)/(1+\phi)$ for AR(1)).
- Frequency is the step size. Downsample flows by sum, levels by mean, stocks by last, rates by ratio of sums; watch incomplete edge periods; seasonal periods are counted in steps.
- A forecast is $\hat y_{T+h\mid T}$: made at the origin $T$ from the information set, for horizon $h$. Uncertainty grows with $h$. Fitted values are not forecasts.
- Random splits leak: they test interpolation. Train on the past, test on the future, and fit every data-driven step inside the training window (rolling origin, 7.15).
Cheat sheet
| Idea | Symbol / formula | Meaning | Python |
|---|---|---|---|
| Time series | $y_t,\ t = 1..T$ | values in time order, regular steps | pd.Series(v, index=pd.date_range(...)) |
| Lag $k$ | $y_{t-k}$ | value $k$ steps back; $T-k$ pairs | y.shift(k) |
| Lead (never a feature) | $y_{t+k}$ | value $k$ steps ahead | y.shift(-k) |
| Lag-$k$ correlation | $r_k = \text{corr}(y_{t-k}, y_t)$ | link with the own past | y.autocorr(k) (ACF in 7.3) |
| AR(1) memory | $y_t-\mu = \phi(y_{t-1}-\mu)+e_t$ | lag-$k$ correlation $\phi^k$ | statsmodels ARIMA(order=(1,0,0)) |
| Effective sample size | $n\,\frac{1-\phi}{1+\phi}$ | AR(1), large $n$ | — |
| Resample | sum / mean / last / ratio of sums | flow / level / stock / rate | y.resample("W").sum(), "ME" |
| Forecast | $\hat y_{T+h\mid T}$ | for $T+h$, made at origin $T$ | — |
| Forecast sd, AR(1) | $\sigma\sqrt{\sum_{j=0}^{h-1}\phi^{2j}}$ | grows with $h$ ($\sigma\sqrt h$ if $\phi = 1$) | — |
| Honest split | train $t \le T$, test $t \gt T$ | no future in training | sklearn TimeSeriesSplit |
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestRegressor
# --- a daily series = values + a DatetimeIndex (the order is part of the data) ---
w1 = [120, 115, 118, 125, 140, 170, 157]
w2 = [126, 119, 121, 130, 146, 178, 160]
y = pd.Series(w1 + w2, index=pd.date_range("2024-06-03", periods=14, freq="D"), name="orders")
print(y.index[0].day_name(), "->", y.index[-1].day_name()) # Monday -> Sunday
# --- lags: shift(k) puts the value from k steps back on each row ---
lags = pd.DataFrame({"y": y, "lag1": y.shift(1), "lag7": y.shift(7)})
print(lags.iloc[7].to_dict()) # Mon 10 Jun: {'y': 126.0, 'lag1': 157.0, 'lag7': 120.0}
print(round(y.autocorr(1), 3), round(y.autocorr(7), 3)) # 0.632 0.997 (Pearson of the lag pairs)
# --- shuffling keeps the histogram but destroys the order ---
rng = np.random.default_rng(0)
shuffled = pd.Series(rng.permutation(y.to_numpy()))
print(y.mean(), shuffled.mean(), round(y.std(), 2), round(shuffled.std(), 2)) # 137.5 137.5 21.15 21.15
print(round(shuffled.autocorr(7), 2)) # -0.4: the weekly link is gone (only 7 pairs, so noisy)
naive_mae = lambda s: np.mean(np.abs(np.diff(np.asarray(s, float)))) # "tomorrow = today"
print(round(naive_mae(y), 2), round(naive_mae(shuffled), 2)) # 14.46 26.92
# --- resampling: "W" means weeks ENDING on Sunday (W-SUN) ---
print(y.resample("W").sum().tolist(), y.resample("W").mean().tolist()) # [945, 980] [135.0, 140.0]
# month-end totals use "ME"; plain "M" raises an error in pandas 3
# --- a rate is a ratio of sums, not a mean of daily rates ---
visitors, conversions = np.array([1000, 3000]), np.array([60, 60])
print(np.mean(conversions / visitors), conversions.sum() / visitors.sum()) # 0.04 0.03
# --- origin and horizon: seasonal naive forecast made at the origin T ---
T = y.index[-1] # origin: Sun 16 Jun, the last known day
fc = pd.Series(y.iloc[-7:].to_numpy(), # "same weekday last week"
index=pd.date_range(T + pd.Timedelta(days=1), periods=7, freq="D"))
print(fc.index[0].date(), fc.iloc[0], "...", fc.index[-1].date(), fc.iloc[-1]) # 2024-06-17 126 ... 2024-06-23 160
# --- temporal dependence: AR(1) with phi = 0.8 makes s/sqrt(n) about 3x too small ---
phi, sig, n, reps = 0.8, 6.0, 100, 4000
x = np.empty((reps, n))
x[:, 0] = rng.normal(0, sig / np.sqrt(1 - phi**2), reps) # start in the long-run state
for t in range(1, n):
x[:, t] = phi * x[:, t - 1] + rng.normal(0, sig, reps)
true_sd_of_mean = x.mean(axis=1).std()
naive_se = (x.std(axis=1, ddof=1) / np.sqrt(n)).mean()
print(round(true_sd_of_mean, 2), round(naive_se, 2), round(true_sd_of_mean / naive_se, 1)) # 2.96 0.96 3.1
# --- random split vs time split: the random split leaks the future ---
days = np.arange(300)
dow = days % 7
level = 100 + 0.2 * days + np.where(days >= 240, 25, 0) # a jump in the last fifth
demand = level + np.array([-10, -14, -12, -6, 4, 22, 16])[dow] + rng.normal(0, 5, days.size)
X = np.column_stack([days, dow])
model = RandomForestRegressor(n_estimators=200, random_state=0)
test_random = rng.choice(days, size=60, replace=False) # random 20%
train_random = np.setdiff1d(days, test_random)
model.fit(X[train_random], demand[train_random])
mae_random = np.mean(np.abs(model.predict(X[test_random]) - demand[test_random]))
train_time, test_time = days[:240], days[240:] # last 20%
model.fit(X[train_time], demand[train_time])
mae_time = np.mean(np.abs(model.predict(X[test_time]) - demand[test_time]))
print(round(mae_random, 1), round(mae_time, 1)) # 5.0 34.4: the random split flatters the model
1. You shuffle the rows of a year of daily sales. Which of these changes?
2. A colleague adds df["y"].shift(-1) as a feature and the backtest error drops to almost zero. What happened?
shift(1) is the lag (yesterday). shift(-1) moves tomorrow's value into today's row. That information does not exist at forecast time: classic leakage.3. Daily values follow an AR(1) process with $\phi = 0.8$. You average 100 days and report $s/\sqrt{100}$ as the standard error. Roughly how does the true standard error compare?
4. You have daily visitors and conversions for one week. What is the correct weekly conversion rate?
5. On Friday evening (data up to Friday) you forecast the following Tuesday. What is the horizon $h$?
6. Why does a random 80/20 split usually overstate forecast accuracy?
Practice problems
A. For the series 5, 8, 6, 9, 12, 10, write the lag-1 and lag-2 columns, count the pairs, and compute the lag-1 correlation of the pairs.
Lag 1: —, 5, 8, 6, 9, 12 (5 pairs). Lag 2: —, —, 5, 8, 6, 9 (4 pairs).
Lag-1 pairs $(y_{t-1}, y_t)$: $(5,8), (8,6), (6,9), (9,12), (12,10)$. Means: yesterday $40/5 = 8$, today $45/5 = 9$. Deviations: yesterday $-3, 0, -2, 1, 4$; today $-1, -3, 0, 3, 1$. Products: $3, 0, 0, 3, 4$, sum $10$. Sums of squares: $9+0+4+1+16 = 30$ and $1+9+0+9+1 = 20$.
$r_1 = 10/\sqrt{30 \times 20} = 10/\sqrt{600} \approx 10/24.49 \approx 0.41$: a moderate positive link with yesterday.
B. You have daily revenue, daily average basket price and end-of-day inventory. How do you turn each into a weekly series? And what goes wrong if the data end on a Thursday?
Revenue is a flow: sum the 7 days. Average basket price is a ratio (revenue ÷ baskets): sum revenue and baskets over the week and divide (or at least weight by baskets). Inventory is a stock: take the last value of the week.
If the data end on a Thursday, the last weekly revenue total covers only 4 days and will look like a sharp drop. Drop that partial week, flag it, or compare per-day averages of the same weekdays.
C. An AR(1) series has $\phi = 0.6$ and shock sd $\sigma = 4$. Find the long-run sd, the sd of today given yesterday, the effective sample size of 200 days, and the standard-error inflation factor.
Long-run sd $= \sigma/\sqrt{1 - \phi^2} = 4/\sqrt{1 - 0.36} = 4/0.8 = 5$. Given yesterday, only the shock is left: sd $= 4$ (20% narrower).
$n_{\text{eff}} \approx 200 \times (1 - 0.6)/(1 + 0.6) = 200 \times 0.4/1.6 = 50$. Inflation factor $\sqrt{(1+\phi)/(1-\phi)} = \sqrt{1.6/0.4} = \sqrt 4 = 2$: the true standard error of the mean is about twice the IID value.
D. With the two-week table and origin $T = 14$ (Sun 16 Jun), give the seasonal naive forecasts for $h = 2$ and $h = 9$.
$h = 2$ is day 16 (Tue 18 Jun): $\hat y_{16\mid 14} = y_{16-7} = y_9 = 119$.
$h = 9$ is day 23 (Tue 25 Jun). $\lceil 9/7 \rceil = 2$, so the rule looks back $14$ days: $\hat y_{23\mid 14} = y_{23-14} = y_9 = 119$. Day 16 would have been "last Tuesday", but it is not known at the origin, so the last known Tuesday is used.
E. Interview: "Your teammate validated a demand model with a random 80/20 split and reports MAE = 5. Do you trust it?"
"Not as a forecast accuracy. A random split leaves training days before and after each test day, so the model is interpolating; any level shift, trend change or holiday effect near a test day is visible in training. I would re-run the evaluation with time-ordered splits: choose several origins, fit on data up to each origin only (including any scaling, feature selection and changepoint detection), forecast the next $H$ days, and report the error per horizon. The honest number is often noticeably larger, and that is the number the business will experience."
F. Interview: "Explain forecast origin, horizon and information set. Is tomorrow's weather in the information set? Is next week's planned promotion?"
"The origin $T$ is the last time whose data I may use. The horizon $h$ is how many steps after $T$ I am forecasting, written $\hat y_{T+h\mid T}$. The information set is everything known at $T$: the history up to $T$ plus inputs that are genuinely known in advance."
"A planned promotion is decided before $T$, so it is in the information set and can be used as a future regressor. Tomorrow's actual weather is not known at $T$; I can only use a weather forecast available at $T$ (with its own error), or leave weather out. Using the actual weather in a backtest would leak the future and overstate accuracy."
Components of a time series
A daily demand series looks messy, but it is usually a handful of simple layers added together: a slow trend, a pattern that repeats every week, spikes on special days, the pull of outside drivers like price or weather, and noise. Learning to see these layers, and to pull them apart, is the first step of every forecasting model, including yours: $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$.
- See a series as a stack of layers and map each layer to a term of your model: trend $g(t)$, seasonality $s(t)$, holidays $h(t)$, regressors $X_t\beta$, noise $\epsilon_t$
- Describe a trend and see why its shape decides long-horizon forecasts
- Define seasonality (fixed, known period) and why seasonal effects are centred to sum to zero
- Tell cycles from seasonality: the interview distinction
- Explain why holidays/events and regressors need their own terms, and what good noise looks like
- Choose between additive and multiplicative structure and use the log to switch
- Run a classical moving-average decomposition by hand, and know its limits (missing ends, no forecasts)
A series is a stack of layers core
Listen to a song. A slow bass line moves under everything. A drum loop repeats every bar. Now and then a cymbal crashes. The singer reacts to the crowd. And there is a little hiss on the recording. You hear one sound, but a sound engineer hears five tracks that were mixed together.
A demand series is mixed the same way. A slow trend (the business growing), a season that repeats (busy weekends), events (a holiday spike), outside drivers (a hot day, a discount) and noise (random customer choices). Decomposing a series means un-mixing it: estimating each track so you can understand it, check it, and continue it into the future.
Three ways to say it:
- Picture: five thin strips stacked on top of each other add up to the one wiggly line you see.
- Numbers: a Black Friday's 288 orders = trend 200 + Friday effect 30 + holiday 45 + discount effect 20 + noise −7.
- Slogan: one line on the screen, several stories underneath.
One day in your shop's history: a Friday that is also Black Friday, with a 10% discount running. Each layer, in orders:
- Trend $g(t)$: the business's underlying level that week is 200 orders a day.
- Seasonality $s(t)$: Fridays are on average 30 above the weekly level: $+30$.
- Holiday $h(t)$: Black Friday adds $+45$ on top of a normal Friday.
- Regressor $X_t\beta$: each percentage point of discount adds about 2 orders ($\beta = 2$), and the discount is $x_t = 10$: $2 \times 10 = +20$.
- Noise $\epsilon_t$: the part nobody could predict that day: $-7$.
- Add the layers: $y_t = 200 + 30 + 45 + 20 - 7 = 288$ orders.
Only the 288 is observed. The five numbers on the left are what a model has to estimate.
An additive decomposition writes the series as a sum of components, each with its own job:
$$y_t \;=\; \underbrace{g(t)}_{\text{trend}} + \underbrace{s(t)}_{\text{seasonality}} + \underbrace{h(t)}_{\text{holidays/events}} + \underbrace{X_t\beta}_{\text{regressors}} + \underbrace{\epsilon_t}_{\text{noise}}.$$- Trend $g(t)$: the slow, long-run level and direction.
- Seasonality $s(t)$: a pattern that repeats with a fixed, known period (7 days, 365.25 days).
- Holidays/events $h(t)$: effects tied to specific dates that do not repeat on a fixed period.
- Regressors $X_t\beta$: effects of measured outside variables $X_t$ (price, discount, temperature) with coefficients $\beta$.
- Noise $\epsilon_t$: what is left; ideally unpredictable, with no pattern.
Classical textbooks use a shorter form, $y_t = T_t + S_t + R_t$ (trend-cycle, seasonal, remainder); holidays and regressors then hide inside the remainder. Cycles (slow rises and falls without a fixed period) are a sixth idea, discussed in their own section. Only $y_t$ is observed: how the total is split into components is a modelling choice.
Why do we need it?
A raw series mixes causes that behave very differently in the future: a trend keeps going, a weekly pattern repeats, a holiday comes back on a known date, noise does not repeat. Separating them is what lets us continue each one sensibly and explain the forecast to a person.
Where is it used?
Prophet and your Prophet-style model (the sum above is the model), classical and STL decomposition in statsmodels, the ETS family (Holt–Winters keeps level, trend and season states, 7.5), structural time-series models, and every "trend + seasonality" chart in a business review.
How is it used?
First plot the series and say out loud which layers you see. Then estimate them: quickly with seasonal_decompose or STL for exploration, or jointly in a model with one term per layer. Finally look at the remainder: any pattern left there is a missing layer.
"The decomposition tells me the true trend and the true seasonality."
Only $y_t$ is observed. The components are estimates, and the same series can be split in several ways (a level shift can be "trend" or "regressor"; a weekly mean can sit in the trend or in the season). Conventions and priors decide the split.
"Noise is just the part we were too lazy to model."
Some noise is irreducible: individual customer decisions. A good model leaves noise with no pattern; it does not try to explain every wiggle (that is overfitting).
This sum is your forecasting model: $g(t)$ is the piecewise-linear trend with changepoints (7.8–7.10), $s(t)$ the Fourier seasonality (7.11), $h(t)$ the holiday effects and $X_t\beta$ the exogenous regressors (7.12), and $\epsilon_t$ is described by the likelihood: Normal, Student-t or Negative Binomial (7.13). Chapter 7.7 puts the whole model together; this chapter is about what each layer means.
$y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$: trend + seasonality + holidays + regressors + noise.
Each layer continues differently into the future: extend, repeat, look up the calendar, needs future X, only its spread.
Trap: components are estimated, not observed; the split is a modelling choice.
Quick check: on a normal Tuesday (no holiday, no discount) the trend is 180, the Tuesday effect is −15 and the noise is +4. What is $y_t$?
$h(t) = 0$ and $X_t\beta = 0$ (no discount), so $y_t = 180 - 15 + 0 + 0 + 4 = 169$ orders.
Trend: the slow, long-run direction core
At the beach, waves come and go every few seconds, but the tide slowly rises over hours. If you only watch one wave you cannot tell whether the tide is coming in. If you step back and blur out the waves, the tide is obvious.
The trend of a series is its tide: the slow level and direction once the fast wiggles (weekdays, noise) are blurred away. It is often the most important layer for forecasting far ahead, because every other layer either repeats (seasonality) or averages out (noise), while the trend keeps moving the whole series up or down.
Three ways to say it:
- Picture: the line you would draw through the middle of the wiggles with a thick pen.
- Numbers: "+2 orders per day" before day 20 and "+0.5 per day" after it is a trend with one change of slope.
- Slogan: seasons repeat, noise cancels, the trend carries the series away.
A straight trend $g(t) = k\,t + m$ with slope $k = 2$ orders per day and intercept $m = 100$ (the level at $t = 0$).
- At $t = 30$: $g(30) = 2 \times 30 + 100 = 160$.
- Now the growth slows at a changepoint on day $s_1 = 20$: the slope changes by $\delta_1 = -1.5$, to $2 - 1.5 = 0.5$ per day.
- The trend must not jump at day 20. Up to day 20 it climbs $2 \times 20 = 40$ to reach 140; after that it climbs $0.5$ per day: $g(30) = 140 + 0.5 \times 10 = 145$.
- The Prophet-style formula gives the same: $g(t) = (k + \delta_1)\,t + (m - s_1\delta_1) = 0.5 \times 30 + (100 + 20 \times 1.5) = 15 + 130 = 145$. The term $-s_1\delta_1$ keeps the line connected (derived in Chapter 7.8).
- Forecast for day 60: the old slope says $2 \times 60 + 100 = 220$; the changed slope says $140 + 0.5 \times 40 = 160$. A 60-order difference, from one slope change. That is why the trend dominates long-horizon forecasts.
The trend $g(t)$ is the slowly varying part of the series: its level and long-run direction, after seasonality, events and noise are removed. Common shapes:
- Level only: $g(t) = m$ (no growth).
- Linear: $g(t) = k\,t + m$; $k$ = slope (change per step), $m$ = intercept.
- Piecewise linear: the slope changes by $\delta_j$ at changepoints $s_j$ and the line stays connected; this is your model's trend (7.8).
- Saturating (logistic): growth slows toward a capacity, an option in Prophet.
- Smooth, estimated: a moving average or LOESS curve; used for description (classical decomposition, STL), not for extrapolation.
The trend-cycle of classical decomposition is the trend together with any slow cycles, since a smoother cannot tell them apart.
Why do we need it?
Growth, decline and changes in direction are what planners care most about ("are we growing, and how fast?"). For forecasts beyond a few weeks, the assumed trend shape decides almost everything, so it must be chosen and checked deliberately.
Where is it used?
Prophet's piecewise-linear and logistic growth, the level and slope states of Holt's method (7.5), differencing in ARIMA, which removes a trend instead of modelling it (7.4), trend lines in dashboards, and capacity planning.
How is it used?
Look at a smoothed version (a 7-day or 28-day moving average) to see the trend. Pick a shape that can be extrapolated sensibly (linear pieces, saturation), let changepoints capture changes of direction, and never extend a high-degree polynomial into the future.
"Any slow up-and-down movement is the trend."
The trend is the long-run level and direction. Repeating ups and downs on a fixed period are seasonality; irregular multi-year swings are cycles. A flexible smoother will happily mix them, so name what you are estimating.
"A polynomial trend that fits the history well is a good forecast."
Polynomials explode outside the fitted range (the cubic in the widget). Use shapes that extrapolate calmly: linear pieces, damped or saturating growth.
"The latest slope will continue."
It is the best guess only if nothing changes, and slopes do change (that is what changepoints are). Long-horizon intervals must include uncertainty about future slope changes.
Your $g(t)$ is piecewise linear: a base slope $k$, an intercept $m$, and slope changes $\delta_j$ at changepoints $s_j$ chosen from a Prophet-like grid plus PELT detection, with Laplace priors $\delta_j \sim Laplace(0, b)$ that keep most changes near zero (7.8–7.10). For forecasts, the model continues the last slope; how much the trend's uncertainty should grow past the last changepoint is one of the most important questions about your intervals.
Trend = slow level and direction. Linear $g(t) = kt + m$; piecewise linear: slope $k \to k + \delta_1 \to \dots$, offsets $-s_j\delta_j$ keep it connected.
The trend dominates long horizons. Trap: never extrapolate a polynomial; slopes can change in the future too.
Quick check: $g(t) = 3t + 50$ with a changepoint at $s_1 = 10$ and $\delta_1 = -2$. What is $g(20)$?
Up to day 10: $3 \times 10 + 50 = 80$. After it the slope is $3 - 2 = 1$: $g(20) = 80 + 1 \times 10 = 90$. Check with the formula: $(3 - 2)\times 20 + (50 + 10 \times 2) = 20 + 70 = 90$.
Seasonality: a pattern that repeats with a fixed period core
Every week has the same rhythm: quiet Monday to Wednesday, busier Friday, a big Saturday. It comes back every 7 days, without fail, because the calendar repeats. The same is true of the year: December is busy every December.
That is seasonality: a pattern tied to the calendar, with a fixed and known period. Because we know the period in advance, seasonality is the easiest layer to forecast: next Saturday will look like past Saturdays.
We describe a season by its effects: how far each position (each weekday) sits above or below the level. By convention the effects average to zero, so the level stays in the trend and the season only says "relatively busier or quieter".
Three ways to say it:
- Picture: cut the series into weeks and stack them: the weekly shapes lie on top of each other.
- Numbers: Sunday $+42$, Tuesday $-25$, and the seven effects add up to 0.
- Slogan: same time on the calendar, same push.
Average orders by weekday over a quiet stretch with no trend: Mon 130, Tue 125, Wed 128, Thu 135, Fri 150, Sat 190, Sun 192.
- Weekly level = average of the seven: $(130+125+128+135+150+190+192)/7 = 1050/7 = 150$.
- Effect = weekday average minus level: Mon $130-150 = -20$, Tue $-25$, Wed $-22$, Thu $-15$, Fri $0$, Sat $+40$, Sun $+42$.
- Check: $-20-25-22-15+0+40+42 = 0$. The effects sum to zero, so they carry only the shape of the week, not its level.
- Forecast for a future Saturday when the trend is at 160: $160 + 40 = 200$.
- Why insist on zero? If the effects summed to $+70$ (average $+10$), the series would look identical with level $150 + 10$ and effects each $10$ lower. Without a convention the level and the season cannot be told apart.
A seasonal component with period $P$ is a function that repeats every $P$ steps:
$$s(t + P) = s(t) \quad \text{for all } t.$$- $P$ is fixed and known in advance from the calendar: 7 for weekly seasonality in daily data, 365.25 for yearly seasonality in daily data, 12 for yearly in monthly data, 24 for daily in hourly data (Chapter 7.1, frequency).
- The values over one period are the seasonal effects. Additive convention: they sum to zero over a period, $\sum_{j=1}^{P} s(j) = 0$ (multiplicative: the factors average to one).
- Multiple seasonalities can be present at once: weekly and yearly in daily demand; daily and weekly in hourly traffic.
- Two common ways to build $s(t)$ in a regression: seasonal dummies ($P - 1$ free numbers, one per position) or Fourier terms (sines and cosines; $2N$ numbers for order $N$, smooth shapes, works for $P = 365.25$; Chapter 7.11).
Why do we need it?
Without a seasonal term, a busy Saturday looks like a surprise every week, the trend gets pulled up and down by weekdays, and forecasts miss the shape of next week. With it, the model expects the rhythm and only has to explain what is unusual.
Where is it used?
Weekly and yearly Fourier terms in Prophet and your model, seasonal dummies in regression, the seasonal states of Holt–Winters, SARIMA's seasonal lags, the seasonal naive baseline ("same day last week", 7.5), and retail, energy and traffic planning.
How is it used?
Find the periods from the calendar and from a plot that folds the series by the period (widget below). Model each one, keep the effects centred, and check the remainder for leftover rhythm (a lag-7 correlation in daily residuals means the weekly term is too weak).
"Seasonality means the seasons of the year."
Any pattern with a fixed calendar period: daily (in hourly data), weekly, monthly, yearly. Daily demand usually has at least weekly and yearly seasonality.
"The seasonal effects do not need to sum to zero."
Without a centring convention (or a prior that does the same job), the level can move between the trend and the season freely: the model is not identifiable, and the split you report is arbitrary.
"The yearly period is 365 days."
On average it is 365.25 (leap years). With a 365-day period the yearly pattern slowly drifts by a day every four years. Prophet uses 365.25.
Your $s(t)$ is a Fourier seasonality: a sum of sines and cosines with period $P$ in the units of $t$, so $P = 7$ (weekly) and $P = 365.25$ (yearly) for daily data (7.11). Sines and cosines each average to zero over a full period, so a Fourier seasonality is centred automatically and the level stays in $g(t)$. The Fourier order $N$ controls how wiggly the seasonal shape may be.
Seasonality: $s(t + P) = s(t)$ with a fixed, known period $P$ (7, 365.25, 12, 24 …). Effects are centred: sum to 0 (additive) or average 1 (multiplicative).
Find $P$ from the calendar; check by folding. Trap: uncentred effects make level and season unidentifiable; yearly $P = 365.25$.
Quick check: weekday averages are 100, 100, 100, 100, 110, 150, 140. What are the centred effects, and what is the level?
Level $= 800/7 \approx 114.3$. Effects: four days at $100 - 114.3 = -14.3$, Fri $-4.3$, Sat $+35.7$, Sun $+25.7$. Check: $4 \times (-14.3) - 4.3 + 35.7 + 25.7 \approx 0$ (up to rounding).
Cycles are not seasonality core
Saturdays come back every 7 days. Christmas comes back every year. You can write their dates in your diary years ahead. That is seasonality.
Economic booms and slowdowns also come back: sales rise for a few years, fall, recover. But nobody can write the next downturn into a diary. It might come in 3 years or in 6. The rises and falls are real, but they have no fixed period. That is a cycle.
The difference matters for forecasting: a seasonal pattern can be copied into the future at the right dates; a cycle cannot, because you do not know when the next turn will happen.
Three ways to say it:
- Picture: seasonal peaks sit on an evenly spaced ruler; cyclic peaks sit at uneven gaps.
- Numbers: seasonal gaps 12, 12, 12 months; cyclic gaps 37, 67, 71 months.
- Slogan: seasons follow the calendar; cycles follow the economy.
Monthly sales of a furniture company over 30 years.
- Every December is a peak: the peaks are 12 months apart, year after year. Gaps: $12, 12, 12, \dots$ This is yearly seasonality ($P = 12$ months).
- On top of that, the yearly totals rise and fall with the housing market. Peaks of this slow swing come in month 30, 67, 134, 205 and 267. Gaps: $67 - 30 = 37$, $134 - 67 = 67$, $205 - 134 = 71$, $267 - 205 = 62$ months.
- The gaps are between 3 and 6 years and never repeat exactly. There is no single period to put into a seasonal term. This is a cycle.
- Forecasting December next year: add the December effect (known). Forecasting the next housing downturn: the series alone cannot tell you when; you need outside information (regressors) or wide uncertainty.
- Seasonality: a pattern that repeats with a fixed, known period tied to the calendar (or the clock). The timing of peaks is predictable for any future date.
- Cycle (cyclic behaviour): rises and falls that are not of fixed length; in economic data typically longer than a year (2–10 years). Both the length and the size of each swing vary.
Cycles come from the dynamics of the system (the economy, a market, a population), not from the calendar. Statistically, a series with memory can show cycle-like swings on its own: an AR(2) process whose coefficients produce damped oscillations has "pseudo-cycles" around an average length (Chapter 7.6). Classical decomposition does not separate cycles from the trend: it estimates a trend-cycle.
Why do we need it?
Calling a cycle "seasonality" leads to a fixed-period term that repeats the last swing forever, at the wrong time. Recognising a cycle tells you to model it differently (trend changes, regressors, AR dynamics) or to widen the uncertainty.
Where is it used?
Business-cycle analysis in economics, commodity and housing markets, ARIMA and state-space models (cycles from AR dynamics), the "trend-cycle" of classical and X-13 decomposition, and capacity planning that must survive a downturn.
How is it used?
Measure the gaps between peaks: constant gaps tied to the calendar mean seasonality; varying gaps mean a cycle. Model seasons with seasonal terms; handle cycles with trend changepoints, leading indicators as regressors, AR terms, or honest wide intervals.
"The series goes up and down, so it has seasonality."
Only if the ups and downs repeat on a fixed calendar period. Uneven swings are cycles (or just a wandering series with memory).
"Add a Fourier term with a 4-year period to capture the business cycle."
That term assumes the next peak comes exactly 4 years after the last, at the same phase, forever. Real cycles drift, so the term will eventually forecast a boom during a bust. Use changepoints, regressors, AR dynamics or wider intervals.
"Seasonality and cycles are the same thing at different speeds."
The difference is not speed but regularity: seasonality has a fixed, known period from the calendar; a cycle has a variable, unknown length.
Model answer: "Seasonality repeats on a known calendar period, like weekly or yearly patterns, so I can model it with seasonal dummies or Fourier terms and forecast it at the right dates. Cycles are rises and falls without a fixed period, often multi-year and driven by the economy, so their timing cannot be extrapolated from the calendar; I handle them with trend changes, regressors or AR-type dynamics, and wider uncertainty."
Your model has no separate cycle term. With a few years of daily data, a multi-year cycle shows up in one of two places: the piecewise trend bends at changepoints to follow it, or it stays in the residuals as slow, autocorrelated wandering (7.17). Do not try to catch it with a long-period Fourier term; if a known driver exists (a market index, marketing spend), add it as a regressor instead.
Seasonality = fixed, known calendar period. Cycle = rises and falls of varying length (often multi-year).
Test: are the gaps between peaks constant and calendar-tied? Classical decomposition lumps cycles into the "trend-cycle".
Trap: never model a cycle with a fixed-period seasonal term.
Quick check: hourly website traffic peaks every evening around 20:00, and also has slow swings that last 2–5 months after big product launches. Which is which?
The evening peak repeats every 24 hours: daily seasonality ($P = 24$ for hourly data). The launch-driven swings have no fixed period: they behave like cycles (or events), and are better explained by launch indicators or trend changes than by a seasonal term.
Holidays and events: spikes on known dates
Black Friday is always a Friday at the end of November, but its date moves: 29 November in 2024, 28 November in 2025, 27 November in 2026. Easter moves by up to five weeks from year to year. A product launch or a TV ad happens once. A weekly or yearly seasonal pattern cannot catch these, because they do not sit at a fixed position of a fixed period.
So events get their own layer, $h(t)$: "on this specific date (and maybe the days around it), add this much". The dates come from a calendar we know in advance, so this layer can be continued into the future, as long as we keep the calendar up to date.
Three ways to say it:
- Picture: a flat line with a few tall, narrow spikes on marked dates.
- Numbers: three Black Fridays left $+44$, $+52$, $+48$ after trend and season, so the effect is about $+48$.
- Slogan: seasons repeat on a ruler; holidays repeat on a calendar.
After removing trend and weekly seasonality, the remainders on Black Friday in three years were $+44$, $+52$ and $+48$ orders, and on the Thursday before it $+10$, $+14$ and $+12$.
- Black Friday effect: average the three remainders: $(44 + 52 + 48)/3 = 144/3 = 48$.
- Day-before effect (a "window" of one day before): $(10 + 14 + 12)/3 = 36/3 = 12$.
- So $h(t) = 48$ on Black Friday, $12$ on the day before, and $0$ on all other days.
- Forecast for next Black Friday when trend + Friday effect = 230: $230 + 48 = 278$, and $12$ extra on the Thursday.
- Caution: each estimate rests on only three days. With so few observations, a model should shrink holiday effects toward zero with a prior instead of trusting the raw average (7.12).
A holiday/event component adds an effect on listed dates:
$$h(t) = \sum_{j} \kappa_j \, \mathbf{1}[t \in D_j],$$- $D_j$ is the set of dates of event $j$ (for example every Black Friday), possibly widened by a window of days before and after; $\mathbf{1}[\cdot]$ is an indicator (1 if true, 0 if not), so each event is one 0/1 column.
- $\kappa_j$ ("kappa j") is the effect of event $j$ in units of $y$ (or one effect per day of the window).
- Recurring events (Christmas, Black Friday) share one $\kappa_j$ across years; one-off events (an outage, a launch) appear once and are mainly modelled to stop them distorting the other layers.
- Events differ from seasonality because their dates are not at a fixed position of a fixed period. They are still known in advance from a calendar, so $h(t)$ is available at forecast time.
Why do we need it?
Unmodelled spikes do double damage: the forecast misses the next holiday, and the spike leaks into the trend and the seasonal estimates (a Black Friday makes "Fridays" and "November" look bigger than they are). A separate term keeps the other layers clean.
Where is it used?
Prophet's holidays dataframe with lower_window / upper_window, the holidays Python package, intervention dummies in ARIMA, retail and travel demand models, and event studies (the effect of a launch or an outage).
How is it used?
Make a table of events and dates, past and future. Build one 0/1 column per event (and per window day), fit their coefficients jointly with the other layers, put shrinkage priors on rare events, and keep the future calendar current so the forecast knows the next dates.
"The yearly seasonality will take care of Christmas and Black Friday."
A yearly term ties effects to a fixed day of the year and is smooth; it cannot follow moving dates (Easter, Black Friday) or produce a one-day spike without bending the whole season around it. Give events their own columns.
"A one-off outage in the history does not matter for the forecast."
Left unmodelled, it pulls down the trend and the seasonal estimates around it. Mark it with an indicator (or treat those days as missing) so it does not distort the other layers.
"An event seen once gives a reliable effect."
One or two observations give a very noisy estimate. Use a shrinkage prior on rare event effects (7.12).
In your model, $h(t)$ is built from holiday indicator columns with their own coefficients and priors (7.12). Because holidays are fitted jointly with the trend and seasonality (not one after the other, as in the classical decomposition above), a holiday spike is much less likely to leak into $g(t)$ or $s(t)$. The holiday calendar is part of the information set at the forecast origin: it must list future dates too.
$h(t) = \sum_j \kappa_j\,\mathbf{1}[t \in D_j]$: one 0/1 column per event (plus window days), effect $\kappa_j$.
Events sit on calendar dates, not on a fixed period; they are known in advance.
Trap: unmodelled spikes leak into trend and season; rare events need shrinkage.
Quick check: why can't a weekly seasonal term absorb Black Friday, even though Black Friday is always a Friday?
The weekly term must give the same effect to every Friday. Black Friday is one Friday a year with a much bigger effect. Spreading its spike over all 52 Fridays would raise every Friday a little and still miss Black Friday badly.
Regressors: outside drivers $X_t\beta$
Some ups and downs have nothing to do with the calendar. Ice-cream orders jump on a hot Tuesday in May. Orders rise when a discount is running. A competitor's sale pulls orders away. The time index cannot know any of this; only an outside measurement can.
A regressor (also called an exogenous variable, a covariate or a feature) is such a measurement, recorded on the same time grid as the series. The model learns a coefficient for it: "each extra degree adds about 3 orders". The catch: to forecast, you need the regressor's future values too.
Three ways to say it:
- Picture: the remainder wiggles in step with the temperature curve, until temperature becomes part of the model.
- Numbers: a promo Tuesday had 165 orders and a similar Tuesday without the promo had 125: about $+40$ per promo day.
- Slogan: the calendar explains "when"; regressors explain "because".
Two Tuesdays with the same trend level and the same weekday effect: one had a promotion (165 orders), one did not (125 orders). On another day the temperature was 5 °C above normal.
- Promotion effect, by comparing like with like: $165 - 125 = 40$ orders. A regression does this comparison across all days at once and gives $\beta_{\text{promo}} \approx 40$ per promo day.
- Temperature: suppose the regression finds $\beta_{\text{temp}} = 3$ orders per °C. A day 5 °C above normal gets $3 \times 5 = +15$.
- For a day with both: $X_t = (x_{\text{promo}}, x_{\text{temp}}) = (1, 5)$ and $\beta = (40, 3)$, so $X_t\beta = 1 \times 40 + 5 \times 3 = 55$ orders on top of trend, season and holidays.
- To forecast next Tuesday, you need next Tuesday's $X$: the promo is planned (known), the temperature must come from a weather forecast (uncertain).
With $p$ regressors, $X_t = (x_{t,1}, \dots, x_{t,p})$ is the row of their values at time $t$ and $\beta = (\beta_1, \dots, \beta_p)^\top$ the column of coefficients:
$$X_t\beta = \sum_{j=1}^{p} x_{t,j}\,\beta_j.$$- $\beta_j$ = expected change in $y_t$ when $x_{t,j}$ rises by one unit, with the other terms held fixed. Its units are "units of $y$ per unit of $x_j$".
- Regressors can be continuous (temperature, price), binary (promo on/off), categorical (as dummies), or lagged (yesterday's ad spend).
- Standardizing a regressor (subtract its mean, divide by its sd) changes the units of $\beta_j$ to "per standard deviation"; fit the scaling on training data only (7.1).
- Future availability: $X_{T+h}$ must be known at the origin (planned prices, calendars) or forecast itself (weather), which adds uncertainty (7.12).
- $\beta$ describes association in the data; it is a causal effect only under extra conditions, for example a randomized promotion (Chapter 5.12).
Why do we need it?
Trend, seasonality and holidays only know the date. Demand also responds to prices, promotions, weather and marketing. Without regressors those effects end up in the noise, the intervals get wider, and the model cannot answer "what if we run a promotion?".
Where is it used?
Prophet's add_regressor, regression with ARIMA errors (ARIMAX/SARIMAX in statsmodels), price-elasticity models, energy load forecasting with temperature, the feature columns of gradient-boosting forecasters, and the design matrix of linear regression (Chapter 5.13).
How is it used?
Pick drivers you will know (or can forecast) at the origin. Check their correlations with each other (4.15), add them as columns, fit their $\beta$ jointly with the other layers, and look at whether the remainder stops following them. For future periods, supply planned or forecast values.
"The regressor improved the backtest, so it belongs in the model."
Only if its future values are available at the origin. Actual weather, actual competitor prices or same-day traffic look great in a backtest that uses the true values and are unavailable live. Backtest with the values you would really have had (7.12).
"$\hat\beta = 3$ means raising the temperature (or the price) by one unit causes 3 more orders."
It is an association learned from the data. It is causal only if nothing else moves with the regressor, which is guaranteed only under randomization.
"Add every available signal as a regressor."
Strongly correlated regressors split the credit unstably, and each extra column can overfit. Check the correlation matrix and use regularizing priors.
In your model, $X_t\beta$ holds the exogenous regressors with priors on $\beta$. Two practical rules follow from this section: every regressor needs future values at forecast time (known in advance or forecast, and then the forecast's own uncertainty should be acknowledged), and any scaling of $X$ must be fitted on the training window only. If a regressor and the trend can both explain the same level shift, the split between them is weakly identified (Chapter 6.8).
$X_t\beta = \sum_j x_{t,j}\beta_j$; $\beta_j$ = change in $y$ per unit of $x_j$, other terms fixed.
Regressors explain variation the calendar cannot (price, promo, weather).
Trap: you need $X$ in the future; association ≠ causation; correlated regressors share credit unstably.
Quick check: $\beta_{\text{price}} = -8$ orders per euro. The price will be cut from 20 to 18 euros next week. What does the regressor term predict, and what assumption does it rely on?
$\Delta x = 18 - 20 = -2$ euros, so the term changes by $(-8) \times (-2) = +16$ orders per day. It assumes the historical association between price and orders holds for this change, for example that past price cuts were not always paired with other promotions.
Noise: what is left when the layers are removed
After you account for the trend, the weekday, the holiday and the weather, today's orders are still not exactly predictable. Some customers just happened to buy, others happened not to. That leftover is noise.
Good noise is boring. It hovers around zero, has a steady spread, and has no memory: knowing yesterday's leftover tells you nothing about today's. If the leftover still has a pattern (a weekly rhythm, spikes on certain dates, long runs above zero), then it is not noise: a layer is missing.
Three ways to say it:
- Picture: static on a radio: no tune left in it.
- Numbers: leftovers 3, −5, 2, 0, −4, 6, −2: mean 0, spread about 4, no pattern.
- Slogan: if you can still see a pattern, it is not noise yet.
A week of remainders $e_t = y_t - (\hat g + \hat s + \hat h + X_t\hat\beta)$: 3, −5, 2, 0, −4, 6, −2.
- Mean: $(3 - 5 + 2 + 0 - 4 + 6 - 2)/7 = 0/7 = 0$. No leftover level.
- Squares: $9, 25, 4, 0, 16, 36, 4$; sum $= 94$.
- Standard deviation: $\sqrt{94/(7-1)} = \sqrt{15.67} \approx 3.96$ orders. This estimates the noise scale $\sigma$.
- Signs: $+, -, +, 0, -, +, -$: they flip back and forth, no long runs. Consecutive products $3 \times (-5) = -15$, $(-5) \times 2 = -10$, … are mostly negative, so there is no positive day-to-day memory here.
- Compare a bad remainder: $8, 9, 7, -6, -8, -7, 9$ (three highs, three lows, a high): long runs of the same sign, the signature of a missing slow layer.
The noise (error, remainder, residual when estimated) is what the other components leave:
$$\epsilon_t = y_t - \big(g(t) + s(t) + h(t) + X_t\beta\big).$$- Standard assumptions of the basic model: mean zero, constant spread, independent over time, and a chosen distribution: Normal, Student-t (heavy tails) or Negative Binomial (counts) (7.13).
- Noise with mean zero, constant variance and no autocorrelation is called white noise (7.3).
- The estimated noise $e_t = y_t - \hat y_t$ (the residual) is our window on these assumptions: structure in $e_t$ (lag correlation, spikes, a spread that grows with the level) means the model is missing something (7.17).
Why do we need it?
The noise layer is what makes forecasts uncertain. Its distribution becomes the likelihood, which sets the width of prediction intervals, and checking it is how we find missing layers.
Where is it used?
The likelihood of every probabilistic forecaster (your Normal, Student-t and Negative Binomial options), residual diagnostics such as the ACF and the Ljung–Box test (7.17), prediction intervals (7.14), and anomaly detection (flag days whose remainder is far outside the usual spread).
How is it used?
After fitting, plot the residuals over time, check their mean and spread, look at their lag-1 and lag-7 correlations and at a histogram or Q-Q plot. A pattern points to the missing layer; a heavy tail points to a Student-t likelihood; spread growing with the level points to a log or multiplicative form.
"The smaller the training remainder, the better the model."
Noise cannot be forecast. A model flexible enough to make the training remainder tiny has fitted the noise (overfitting) and will forecast worse. Judge models on future data.
"A large remainder sd means the model is wrong."
Some series are simply noisy. What signals a wrong model is structure in the remainder: lag correlation, spikes on dates, a spread that changes with the level.
In your model, $\epsilon_t$ is described by the likelihood: Normal (symmetric, light tails), Student-t (heavy tails, so extreme days have less influence) or Negative Binomial (counts whose variance grows with the mean) (7.13). The model assumes the noise is independent from day to day; a lag-1 or lag-7 bar outside the band in your residual ACF means that assumption, and therefore your interval widths, need a second look (7.17).
$\epsilon_t = y_t - (g + s + h + X\beta)$; assumed mean 0, steady spread, independent; its distribution = the likelihood.
Residual structure = missing layer: lag-7 bar → seasonality; date spikes → holidays; long runs → trend or cycle.
Trap: a tiny training remainder is overfitting, not success.
Quick check: the residual ACF of your daily model has a large bar at lag 7 and small bars elsewhere. What is missing, and what would you change?
A weekly rhythm is left over, so the weekly seasonality is missing or too weak (for example, too low a Fourier order for the weekly term, or a weekly pattern that changes over time). Add or strengthen the weekly term, or let it vary (for example multiplicatively with the level), then re-check the ACF.
Additive or multiplicative? core
A small shop sells 100 orders on a normal weekday and 120 on Saturday. Five years later it has grown and sells 200 on a normal weekday. Does Saturday now bring 220 (still 20 more) or 240 (still 20% more)?
If the Saturday boost stays a fixed number of orders, the layers add: additive structure. If it stays a fixed percentage, the layers multiply: multiplicative structure. For businesses that grow, percentages are usually the more natural story: a bigger shop has bigger Saturdays. You can see it in a plot: the weekly swings get taller as the series rises, like a funnel.
Three ways to say it:
- Picture: additive = a band of constant width around the trend; multiplicative = a funnel that widens as the trend rises.
- Numbers: additive: Saturday $+20$ at level 100 and at level 200; multiplicative: $+20\%$, i.e. $+20$ then $+40$.
- Slogan: taking logs turns "times" into "plus".
Trend 100 in January and 200 in December. The Saturday effect is either $+20$ orders (additive) or $\times 1.2$ (multiplicative).
- Additive: January Saturday $= 100 + 20 = 120$; December Saturday $= 200 + 20 = 220$. The swing above the trend stays 20.
- Multiplicative: January Saturday $= 100 \times 1.2 = 120$; December Saturday $= 200 \times 1.2 = 240$. The swing doubles to 40, because the trend doubled.
- Take logs of the multiplicative version: $\log 120 - \log 100 = \log 1.2 \approx 0.182$ and $\log 240 - \log 200 = \log 1.2 \approx 0.182$. On the log scale the swing is constant again.
- So multiplicative data can be modelled additively after a log: $\log y_t = \log g(t) + \log s(t) + \log r_t$.
- Additive: $y_t = T_t + S_t + R_t$. Seasonal effects are in units of $y$ and sum to 0 over a period; their size does not depend on the level.
- Multiplicative: $y_t = T_t \times S_t \times R_t$. Seasonal factors are unit-free, average 1 over a period (1.2 means "20% above"), and the seasonal swing in units of $y$ grows with the trend. The remainder factor $R_t$ is around 1, so the noise also grows with the level.
- Log link: for positive data, $\log y_t = \log T_t + \log S_t + \log R_t$ is additive. Requires $y_t \gt 0$; zeros need
log1p, a different transformation, or a count likelihood (Chapter 4.18). - Prophet's multiplicative mode scales the chosen terms by the trend: $y_t = g(t)\,\big(1 + s(t)\big) + \dots$, so $s(t) = 0.2$ means 20% above the trend.
How to tell from data: the seasonal swing (or the spread of the noise) grows with the level → multiplicative or log; it stays constant → additive.
Why do we need it?
With the wrong structure the model systematically under-predicts peaks when the level is high and over-predicts them when it is low, and its residual spread changes with the level. Choosing the right form (or a log) fixes both with no extra parameters.
Where is it used?
seasonal_decompose(model="multiplicative"), Holt–Winters with multiplicative seasonality, Prophet's seasonality_mode="multiplicative", log transforms before ARIMA, and count models with a log link, where every term multiplies the mean.
How is it used?
Plot the series and compare the size of the seasonal swing early and late (widget). If it grows with the level, fit on the log scale or switch the seasonal term to multiplicative mode. Remember that back-transformed log forecasts estimate the median, not the mean.
"A log transform fixes every growing pattern."
It needs $y \gt 0$. With zeros (quiet days, small stores) the log breaks; log1p distorts small values. For counts, a count likelihood with a log link (or a multiplicative mode) is often cleaner.
"Back-transforming the log forecast gives the expected value."
$e^{\text{mean of } \log y}$ is the median (when the errors on the log scale are symmetric), which is below the mean for right-skewed data (Chapter 4.18). Totals built from back-transformed medians are biased low.
"Multiplicative seasonality means the model has more parameters."
Same parameters, different structure: the seasonal effect is a percentage of the trend instead of a fixed number of units.
Model answer: "I check whether the seasonal swing grows with the level. If it does, the effects are proportional, so I either model log demand additively or use multiplicative seasonality, where $s(t)$ scales the trend. For counts I prefer a log-link likelihood, which makes every component act multiplicatively on the mean."
Your model, as written, is additive: $g + s + h + X\beta$. If the weekly swing in your demand grows with the trend, there are three standard responses: fit on $\log y$, give the seasonal (and holiday) terms a multiplicative form $g(t)\,(1 + s(t))$, or, for counts, put the components through a log link so they multiply on the original scale. Which link your Negative Binomial likelihood uses (log, softplus or identity on the mean) decides whether your components add or multiply: check it in your code (7.13).
Additive $y = T + S + R$ (constant swing) vs multiplicative $y = T \times S \times R$ (swing ∝ level); $\log$ turns the second into the first.
Diagnose: does the seasonal swing (or noise spread) grow with the level?
Traps: log needs $y \gt 0$; back-transformed log forecasts are medians.
Quick check: the December peak is 30 above the trend when the trend is 150, and 60 above when the trend is 300. Additive or multiplicative? What is the seasonal factor?
Multiplicative: the peak is $30/150 = 20\%$ and $60/300 = 20\%$ above the trend both times. The factor is $1.2$; on the log scale the effect is $\log 1.2 \approx 0.18$ in both years.
Classical decomposition by moving averages core
How do you see the trend through the weekly wiggles? Average over windows exactly one week wide. Every such window contains one Monday, one Tuesday, …, one Sunday, so the busy and quiet days cancel and only the level of that week is left. Slide the window along the series and you get a smooth trend.
Then subtract the trend. What remains is "season + noise". Average it separately for all Mondays, all Tuesdays, and so on: the noise averages out and the weekday effects remain. Subtract those too, and what is left is the remainder. That is the whole classical decomposition: three subtractions and two kinds of averaging.
Three ways to say it:
- Picture: blur with a one-period window to get the trend; stack the periods to get the season; whatever is left is the remainder.
- Numbers: for quarterly data, average four quarters at a time; for daily data with weekly seasonality, seven days at a time.
- Slogan: window = period, average by position, subtract, repeat.
Eight quarters of sales (two years), built as trend $100 + 4t$ plus quarterly effects, with no noise so every step can be checked: $y_1, \dots, y_8 = 116, 104, 92, 128, 132, 120, 108, 144$. The period is $m = 4$.
- Moving averages of 4 consecutive quarters (each sits between two quarters): $(116+104+92+128)/4 = 440/4 = 110$; then $456/4 = 114$; $472/4 = 118$; $488/4 = 122$; $504/4 = 126$.
- Because 4 is even, average neighbouring pairs to centre them on a quarter (the "2×4" moving average): $\hat T_3 = (110 + 114)/2 = 112$, $\hat T_4 = 116$, $\hat T_5 = 120$, $\hat T_6 = 124$. Quarters 1, 2, 7 and 8 get no trend value: the window would stick out of the data.
- Detrend: $y_t - \hat T_t$: Q3: $92 - 112 = -20$; Q4: $128 - 116 = +12$; Q1 (year 2): $132 - 120 = +12$; Q2 (year 2): $120 - 124 = -4$.
- Average by quarter (here one value each) and centre: the mean of $(+12, -4, -20, +12)$ is $0/4 = 0$, so the seasonal effects are Q1 $+12$, Q2 $-4$, Q3 $-20$, Q4 $+12$.
- Remainder $y_t - \hat T_t - \hat S_t$: e.g. Q3: $92 - 112 - (-20) = 0$. All remainders are 0 because the example had no noise. With real data they would be the noise.
Classical additive decomposition with seasonal period $m$:
- Trend: $\hat T_t$ = centred moving average of order $m$. For odd $m$: $\hat T_t = \frac1m\sum_{j=-(m-1)/2}^{(m-1)/2} y_{t+j}$. For even $m$: the "$2\times m$" average, weights $\frac{1}{2m}$ on the two end values and $\frac1m$ on the $m-1$ values between.
- Detrend: $d_t = y_t - \hat T_t$.
- Season: for each position $j$ in the period (each weekday), average all $d_t$ at that position; subtract the mean of these $m$ averages so the effects sum to 0. Repeat them over the whole series: $\hat S_t$.
- Remainder: $\hat R_t = y_t - \hat T_t - \hat S_t$.
Multiplicative version: $d_t = y_t/\hat T_t$, factors normalised to average 1, $\hat R_t = y_t/(\hat T_t \hat S_t)$.
Limits: no trend (and no remainder) for the first and last $\lfloor m/2 \rfloor$ points; the seasonal pattern is assumed identical every period; outliers and holidays leak into the trend and the season; sudden trend changes are smoothed over; there is no uncertainty and no forecast. STL (seasonal-trend decomposition using LOESS) fixes several of these: it covers the ends, lets the season change slowly, and has a robust option against outliers.
Why do we need it?
It is the quickest honest look at a series' layers: how strong the trend is, what the weekly shape looks like, whether the swing grows with the level, and what is left over. It is exploration before modelling, and a sanity check after.
Where is it used?
statsmodels.tsa.seasonal.seasonal_decompose and STL, R's decompose and stl, official statistics (X-13ARIMA-SEATS for seasonal adjustment of economic data), and the "trend / weekly / yearly" component plots that Prophet-style models produce.
How is it used?
Call seasonal_decompose(y, model="additive", period=7) (the period is inferred from a daily index), plot the four panels, compare additive vs multiplicative remainders, and use STL with robust=True when there are outliers. Do not use it to forecast.
"I ran seasonal_decompose, so I can forecast with its components."
It only describes the past: it has no formula to extend the trend, no uncertainty, and its trend stops $\lfloor m/2 \rfloor$ steps before the end, exactly where a forecast starts. (extrapolate_trend fills the ends with a straight-line guess; it is a guess.)
"Any smoothing window will do for the trend."
The window must be the period (or a multiple of it), otherwise the seasonal pattern leaks into the trend as ripples. For an even period use the centred $2\times m$ average.
"The seasonal effects from the decomposition are the same every year because the business is the same."
The classical method forces them to be the same. If the weekly shape changes over time (more weekend shopping lately), use STL, which lets the season evolve slowly.
"Classical decomposition and my forecasting model do the same thing."
Both split a series into layers, but decomposition is a descriptive smoother applied after the fact, while the model writes each layer as a parametric function of time, fits all layers jointly with a likelihood and priors, and can extend them into the future with uncertainty.
Model answer: "I use classical or STL decomposition for exploration: to see the trend, the seasonal shape, whether the effects look additive or multiplicative, and what is left. My forecasting model is different: $g$, $s$, $h$ and $X\beta$ are parametric, they are estimated jointly so that a holiday does not leak into the trend, and the posterior gives forecast distributions, not just a smoothed past."
Before fitting your forecasting model, a quick seasonal_decompose or STL on the training window answers useful questions: is the weekly shape stable, does its size grow with the level (additive vs multiplicative), are there spikes on known dates (holidays to add), and does the remainder still show slow waves (trend changes or cycles)? Run it only on data before the forecast origin, so it does not become a source of leakage.
Classical: $\hat T$ = centred MA with window $m$ (2×$m$ if even) → $d = y - \hat T$ → average $d$ by position, centre → $\hat S$ → $\hat R = y - \hat T - \hat S$ (divide for multiplicative).
Limits: ends missing, fixed season, outliers leak, no forecast. STL: robust, evolving season.
Trap: window ≠ period leaves ripples; decomposition is description, not a forecasting model.
Quick check: monthly data with yearly seasonality. Which moving average gives the trend, and how many months have no trend value at each end?
$m = 12$ is even, so use the centred $2\times12$ moving average: 13 months, weights $1/24$ on the two end months and $1/12$ on the 11 months between. It needs 6 months on each side, so the first 6 and last 6 months get no trend value.
Recap, cheat sheet and practice
- A series is a stack of layers: $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$. Only $y_t$ is observed; the split is a modelling choice, made unique by conventions (centred seasons) and priors.
- Trend $g(t)$: slow level and direction; decides long-horizon forecasts; piecewise-linear in your model. Never extrapolate polynomials.
- Seasonality $s(t)$: fixed, known period ($s(t+P) = s(t)$); effects centred (sum 0, or average 1); several periods can coexist (7 and 365.25 days).
- Cycles: rises and falls with no fixed period; not seasonality; they end up in the trend or in autocorrelated residuals.
- Holidays/events $h(t)$: indicator columns on known, moving dates. Regressors $X_t\beta$: outside drivers that need future values. Noise $\epsilon_t$: no pattern; its distribution is the likelihood; structure in residuals = a missing layer.
- Additive vs multiplicative: constant swing vs swing ∝ level; the log turns the second into the first.
- Classical decomposition: centred MA with window = period → detrend → average by position and centre → remainder. Descriptive only: missing ends, fixed season, leaks from outliers; STL is the robust alternative.
Cheat sheet
| Layer | Symbol | What it captures | How to see it | In your model |
|---|---|---|---|---|
| Trend | $g(t)$ | slow level and direction | 7- or 28-day centred MA | piecewise linear, changepoints, Laplace priors (7.8–7.10) |
| Seasonality | $s(t)$, $s(t+P)=s(t)$ | fixed calendar period | fold by $P$; weekday averages | Fourier terms, $P = 7$, $365.25$ (7.11) |
| Cycle | — | swings of varying length | gaps between peaks vary | no term: trend or residuals |
| Holidays | $h(t)=\sum_j\kappa_j\mathbf 1[t\in D_j]$ | effects on known dates | spikes in the remainder | indicator columns + priors (7.12) |
| Regressors | $X_t\beta$ | outside drivers | remainder follows a driver | exogenous columns; need future $X$ (7.12) |
| Noise | $\epsilon_t$ | unpredictable rest | residual plot, ACF, Q-Q | Normal / Student-t / NB likelihood (7.13) |
| Multiplicative | $y = T\times S\times R$ | swing ∝ level | swing grows with the trend | log $y$, multiplicative mode, or a log link |
| Classical decomposition | $\hat T$, $\hat S$, $\hat R$ | descriptive split | seasonal_decompose, STL | exploration before fitting |
import numpy as np
import pandas as pd
from statsmodels.tsa.seasonal import seasonal_decompose, STL
# --- 1) the worked example: 8 quarters, classical additive decomposition ---
q = np.array([116, 104, 92, 128, 132, 120, 108, 144], dtype=float)
res = seasonal_decompose(q, model="additive", period=4)
print(res.trend) # [ nan nan 112. 116. 120. 124. nan nan]
print(res.seasonal[:4]) # [ 12. -4. -20. 12.] Q1..Q4 effects, sum 0
print(res.resid[2:6]) # [0. 0. 0. 0.] (the example has no noise)
w = np.array([1, 2, 2, 2, 1]) / 8 # the 2x4 moving average by hand
print(np.convolve(q, w, mode="valid")) # [112. 116. 120. 124.]
# --- 2) a daily series built from known layers ---
rng = np.random.default_rng(1)
t = np.arange(16 * 7)
idx = pd.date_range("2024-01-01", periods=t.size, freq="D") # 1 Jan 2024 is a Monday
week = np.array([-10, -14, -12, -6, 4, 22, 16]) # Mon..Sun, sums to 0
trend = 100 + 0.5 * t
y_add = pd.Series(trend + week[t % 7] + rng.normal(0, 3, t.size), index=idx)
dec = seasonal_decompose(y_add, model="additive") # period 7 inferred from freq "D"
print(dec.seasonal.iloc[:7].round(1).tolist()) # [-9.6, -13.6, -12.8, -5.7, 3.6, 22.1, 15.9] close to the truth
print(round(dec.seasonal.iloc[:7].sum(), 10)) # 0.0: effects sum to zero
print(int(dec.trend.isna().sum())) # 6: no trend for 3 days at each end
print(round(dec.resid.std(), 2)) # 2.52 (noise sd was 3; the MA and averages soak up a little)
# --- 3) multiplicative data: the weekly swing grows with the level ---
y_mul = pd.Series(trend * (1 + 0.25 * week[t % 7] / 22) * (1 + rng.normal(0, 0.02, t.size)), index=idx)
swing = lambda s, k: s.iloc[7 * k: 7 * k + 7].max() - s.iloc[7 * k: 7 * k + 7].min()
print(round(swing(y_mul, 0), 1), round(swing(y_mul, 15), 1)) # 43.4 64.7: grows
print(round(swing(np.log(y_mul), 0), 3), round(swing(np.log(y_mul), 15), 3)) # 0.409 0.41: constant on the log scale
dm = seasonal_decompose(y_mul, model="multiplicative")
print(dm.seasonal.iloc[:7].round(3).tolist()) # [0.889, 0.849, 0.86, 0.928, 1.045, 1.261, 1.169]
print(round(dm.seasonal.iloc[:7].mean(), 10)) # 1.0: factors average to one
# --- 4) STL: LOESS-based, robust option, no missing ends ---
stl = STL(y_add, period=7, robust=True).fit()
print(int(stl.trend.isna().sum()), round(stl.resid.std(), 2)) # 0 2.29
1. Which of these is seasonality?
2. Weekday averages (Mon to Sun) are 90, 95, 100, 105, 110, 130, 140. What is Saturday's centred seasonal effect?
3. Over three years the trend doubled, and the weekly swing doubled too. What is the most natural fix?
4. Daily data with a weekly pattern. Which centred moving average gives a classical trend estimate without weekly ripples?
5. Why can't you forecast with seasonal_decompose's output directly?
6. The residual ACF of your daily model shows a large bar at lag 7 and nothing else. Which layer is most likely missing or too weak?
Practice problems
A. Quarterly data: 50, 30, 40, 60, 58, 38, 48, 68. Run the classical additive decomposition (period 4) by hand.
4-quarter averages: $(50+30+40+60)/4 = 45$, $(30+40+60+58)/4 = 47$, $(40+60+58+38)/4 = 49$, $(60+58+38+48)/4 = 51$, $(58+38+48+68)/4 = 53$.
Centre (2×4): $\hat T_3 = (45+47)/2 = 46$, $\hat T_4 = 48$, $\hat T_5 = 50$, $\hat T_6 = 52$; quarters 1, 2, 7, 8 have no trend value.
Detrended: Q3 $40 - 46 = -6$, Q4 $60 - 48 = 12$, Q1 $58 - 50 = 8$, Q2 $38 - 52 = -14$. Their mean is $0$, so the seasonal effects are Q1 $+8$, Q2 $-14$, Q3 $-6$, Q4 $+12$. Remainders are 0: these data were built as trend $40 + 2t$ plus exactly these effects.
B. In year 1 the trend is 400 and December is 80 above it; in year 3 the trend is 800 and December is 160 above it. Additive or multiplicative? Predict December in year 4 if the trend will be 1 000.
$80/400 = 160/800 = 0.2$: the effect is a constant 20% of the trend, so the structure is multiplicative with factor $1.2$. Year 4: $1000 \times 1.2 = 1200$ (an additive model would have said $1000 + 160 = 1160$ or less). On the log scale the December effect is $\log 1.2 \approx 0.18$ every year.
C. A trend has base slope $k = 1.5$, intercept $m = 80$, and changepoints at $s_1 = 10$ ($\delta_1 = +1$) and $s_2 = 30$ ($\delta_2 = -2$). Find $g(40)$ two ways.
Step by step: up to day 10 slope 1.5: $80 + 15 = 95$. Days 10–30 slope $2.5$: $95 + 50 = 145$. Days 30–40 slope $0.5$: $145 + 5 = 150$.
Formula: slope $= 1.5 + 1 - 2 = 0.5$; offset $= 80 - 10 \times 1 - 30 \times (-2) = 80 - 10 + 60 = 130$; $g(40) = 0.5 \times 40 + 130 = 150$.
D. A one-day spike of $+70$ sits in the middle of a daily series. What does a classical additive decomposition (period 7) do with it?
The 7-day moving average includes the spike in the seven windows centred within 3 days of it, so $\hat T$ rises by $70/7 = 10$ on those seven days. The remainder on the spike day is about $70 - 10 = 60$, and the six neighbouring days get remainders about $-10$ (small dips). The spike also nudges its weekday's seasonal average up by $60/(\text{number of weeks})$. A model that fits a holiday indicator jointly with the other layers avoids this smearing.
E. Interview: "What is the difference between seasonality and a cycle, and how does your model treat each?"
"Seasonality repeats on a fixed calendar period, like weekly or yearly patterns, so its timing is known for any future date; my model represents it with Fourier terms with periods 7 and 365.25 days. A cycle is a rise and fall without a fixed length, like a business cycle; its timing cannot be read from the calendar. My model has no cycle term: a cycle would be absorbed by trend changepoints or remain as slow autocorrelation in the residuals, which I check. If a known driver exists, I would add it as a regressor rather than force a fixed-period term."
F. Interview: "Walk me through the components of your forecasting model and how each one is extended into the future."
"$y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$. The trend $g$ is piecewise linear with changepoints; beyond the last data point it continues with the latest slope, and its uncertainty is the main reason intervals widen with the horizon. The seasonality $s$ is a Fourier sum with known periods, so it simply repeats. The holiday term $h$ uses indicator columns from a calendar that includes future dates. The regressor term $X\beta$ needs future values of $X$: known in advance (planned prices, promotions) or forecast, which adds uncertainty. The noise $\epsilon$ is not extended; its likelihood (Normal, Student-t or Negative Binomial) sets the spread of the predictive distribution around the mean."
Autocorrelation: ACF, PACF, white noise, random walks
A busy day is often followed by another busy day. A time series remembers its own past, and that memory is both a gift (it makes forecasting possible) and a trap (it makes ordinary standard errors lie). This chapter gives you the tools to see the memory: the ACF and PACF plots, the two reference series every analyst must know (white noise and the random walk), and the question an interviewer will ask about your model: "your Prophet-style model treats the noise as independent; what if it is not?"
- Compute an autocorrelation $r_k$ by hand, with the exact convention statsmodels uses (one overall mean, divide by $n$)
- Read an ACF plot (correlogram): decay, waves, seasonal spikes, and the ±1.96/√n band
- Understand the PACF as "the direct link after removing the in-between days", and compute it from the ACF
- Recognise the ACF/PACF signatures of white noise, AR(1), AR(2), MA(1), random walks, trends and weekly patterns
- Define white noise precisely (and see why it is not the same as "independent")
- Know the random walk: wandering paths, spread growing like $\sqrt t$, an ACF that refuses to die
- See how serial dependence makes naive standard errors and intervals too narrow
- Answer the interview question: a Prophet-style model does not model autocorrelated residuals explicitly; what follows, and what can you do?
Autocorrelation at lag k: a series correlated with its own past core
Busy days come in runs. A heat wave, a promotion or a payday week pushes orders up for several days in a row; a holiday lull pulls them down for several days in a row. So if today was above average, tomorrow will probably be above average too. The series remembers.
In Chapter 4.15 correlation measured whether two different variables move together. Autocorrelation ("auto" = self) measures whether a series moves together with its own past. The lag $k$ is how many steps back we look: lag 1 = yesterday, lag 7 = the same weekday last week (lags were introduced in Chapter 7.1).
Three ways to say it:
- Picture: write the series on a strip of paper, make a copy, slide the copy $k$ days to the right, and ask: do the two strips go up and down together?
- Numbers: for the 8 days below, $r_1 = 0.5$ (above-average days tend to be followed by above-average days) and $r_3 = -0.5$ (days three apart tend to sit on opposite sides of the average).
- Slogan: autocorrelation is correlation with yourself, $k$ steps ago.
Eight days of orders: $y = 7, 8, 10, 12, 13, 11, 9, 10$.
- Mean: $\bar y = (7+8+10+12+13+11+9+10)/8 = 80/8 = 10$.
- Deviations $d_t = y_t - \bar y$: $-3, -2, 0, 2, 3, 1, -1, 0$.
- Squares: $9, 4, 0, 4, 9, 1, 1, 0$. Their sum is $28$. (This is the "total spread" we will divide by.)
- Lag 1: pair each day with the day before, $d_t \times d_{t-1}$ for $t = 2, \dots, 8$: $(-2)(-3) = 6$, $(0)(-2) = 0$, $(2)(0) = 0$, $(3)(2) = 6$, $(1)(3) = 3$, $(-1)(1) = -1$, $(0)(-1) = 0$. Sum $= 14$.
- $r_1 = 14/28 = 0.5$. Positive: neighbours tend to be on the same side of the mean.
- Lag 3: $d_t \times d_{t-3}$ for $t = 4, \dots, 8$: $(2)(-3) = -6$, $(3)(-2) = -6$, $(1)(0) = 0$, $(-1)(2) = -2$, $(0)(3) = 0$. Sum $= -14$, so $r_3 = -14/28 = -0.5$. The series rises for about three days and then falls, so days three apart disagree.
- Lag 2 the same way: sum $= -5$, so $r_2 = -5/28 \approx -0.179$.
Notice two conventions. We used one overall mean (10) for both members of each pair, and we divided by the sum of all 8 squared deviations even though lag 3 has only 5 pairs. This is exactly what statsmodels.tsa.stattools.acf does.
Let $y_1, \dots, y_n$ be a series with mean $\bar y$. The sample autocovariance and sample autocorrelation at lag $k$ are
$$\hat\gamma_k = \frac{1}{n}\sum_{t=k+1}^{n} (y_t - \bar y)(y_{t-k} - \bar y), \qquad r_k = \frac{\hat\gamma_k}{\hat\gamma_0}.$$For the random process behind the data, the autocovariance is $\gamma(k) = Cov(y_t, y_{t-k})$ and the autocorrelation is $\rho(k) = Corr(y_t, y_{t-k}) = \gamma(k)/\gamma(0)$. These only make sense as "one number per lag" when the process is stationary: its mean, its spread and its links to the past do not change over time, so $\rho(k)$ is the same for every $t$ (made precise in Chapter 7.4).
- $r_0 = 1$ always (a series is perfectly correlated with itself), and $-1 \le r_k \le 1$.
- The lag is symmetric: the link between today and $k$ days ago is the same as between today and $k$ days ahead, so we only plot $k \ge 0$.
- Serial dependence means "values at different times are not independent". Autocorrelation measures the straight-line part of it, exactly as Pearson correlation does for two variables.
- Dividing by $n$ (not $n-k$) pulls $r_k$ slightly toward 0 at big lags, but it guarantees a well-behaved set of autocorrelations (they always form a valid correlation structure). It is the standard choice.
Why do we need it?
Ordinary statistics assume independent rows. Time series break that assumption, and autocorrelation is the number that says by how much. Without it we cannot tell whether yesterday helps predict today, or whether our standard errors can be trusted.
Where is it used?
Choosing ARIMA orders (Chapter 7.6), checking forecast residuals (Chapter 7.17), finding the weekly cycle in daily orders (a big $r_7$), MCMC effective sample size (Chapter 6.10), and correcting standard errors for dependent data.
How is it used?
Call acf(y, nlags=30) from statsmodels.tsa.stattools (or plot_acf) on the series, or better on the residuals of your model. Look at lag 1 (short memory), lag 7 (weekly pattern) and how fast the values fade.
"$r_k$ is just np.corrcoef(y[k:], y[:-k])" (or pandas.Series.autocorr(k)).
Those compute a Pearson correlation of the overlapping pairs, with each piece's own mean and spread. The ACF uses one overall mean and divides by the whole series' spread. For long series the two are close; for our 8 days they give 0.63 instead of 0.5 at lag 1, and −0.87 instead of −0.5 at lag 3. Use statsmodels' acf when you want the ACF.
"A large $r_1$ in my raw sales data means I need an autoregressive model."
A trend or a seasonal pattern alone makes neighbouring values similar and gives a large $r_1$. The raw ACF mixes trend, seasonality and short-term memory. The ACF that tells you about leftover memory is the one of the residuals after your model (see the last section).
"Negative autocorrelation is impossible for demand."
It happens: a stock-up day followed by a quiet day, alternating shifts, or a series that you differenced once too often (Chapter 7.4). A negative $r_1$ means a zigzag.
$r_k = \dfrac{\sum_{t=k+1}^{n}(y_t-\bar y)(y_{t-k}-\bar y)}{\sum_{t=1}^{n}(y_t-\bar y)^2}$: one overall mean, divide by the full spread; $r_0 = 1$.
Autocorrelation = correlation of a series with itself $k$ steps back. Positive = runs, negative = zigzag.
Trap: np.corrcoef of shifted pieces and pandas .autocorr() are not the ACF; raw-data ACF mixes trend, season and memory.
Quick check: a series has 8 values. How many pairs are used for $r_3$, and what do we divide by?
$8 - 3 = 5$ pairs (days 4–8 paired with days 1–5). We still divide by the sum of all 8 squared deviations, $\sum_{t=1}^{8}(y_t - \bar y)^2$, not by a sum over 5 days.
The ACF plot: the series' memory profile core
One lag is one number. Compute $r_k$ for every lag $k = 1, 2, \dots, 30$ and draw each one as a vertical stick: that picture is the ACF (autocorrelation function), also called a correlogram.
Think of shouting in a canyon. The echo after one second is loud, after two seconds quieter, and soon it is gone: that is a short memory. In a canyon with a wall exactly seven seconds away you hear a strong echo at 7, 14 and 21 seconds: that is a weekly pattern in daily data. And some series echo almost forever: a trend or a random walk.
Three ways to say it:
- Picture: a row of sticks, one per lag. Tall sticks = strong memory at that distance; the shape of the row tells you what kind of memory it is.
- Numbers: if "today = 0.8 × yesterday + a fresh surprise", the sticks are $0.8, 0.64, 0.51, 0.41, \dots$: each one 0.8 times the one before.
- Slogan: the ACF is the memory profile of a series: how strongly, and for how long, it remembers.
A simple series with memory: today = 0.8 × yesterday + a fresh surprise, written $y_t = 0.8\,y_{t-1} + \epsilon_t$, where the surprises $\epsilon_t$ are independent with mean 0. This is called an AR(1) series ("autoregressive of order 1": a regression of the series on its own previous value; the full family comes in Chapter 7.6). The number 0.8 is the carry-over $\varphi$ ("phi").
- Lag 1: today is built from yesterday, so $\rho(1) = \varphi = 0.8$.
- Lag 2: today carries 0.8 of yesterday, which carried 0.8 of the day before: $\rho(2) = 0.8 \times 0.8 = 0.64$.
- In general $\rho(k) = 0.8^k$: $\rho(3) = 0.512$, $\rho(5) \approx 0.328$, $\rho(7) \approx 0.210$, $\rho(10) \approx 0.107$.
- Half-life of the memory: $0.8^k = 0.5$ gives $k = \ln 0.5 / \ln 0.8 \approx 3.1$ days.
- Why the rule holds: $Cov(y_t, y_{t-k}) = Cov(0.8\,y_{t-1} + \epsilon_t,\; y_{t-k}) = 0.8\,Cov(y_{t-1}, y_{t-k}) + 0$, because today's fresh surprise is independent of the past. So each extra lag multiplies the correlation by 0.8.
With $\varphi = -0.6$ the sticks alternate in sign: $-0.6, 0.36, -0.216, \dots$ (a zigzag series). With $\varphi = 0$ every stick is 0: no memory at all.
The autocorrelation function of a stationary process is $k \mapsto \rho(k) = Corr(y_t, y_{t-k})$ for $k = 0, 1, 2, \dots$. Its estimate from data is $k \mapsto r_k$, and the correlogram (ACF plot) draws $r_k$ as sticks against $k$.
How to read it:
- Lag 0 is always 1 (statsmodels draws it;
plot_acf(..., zero=False)hides it). - Fast, smooth decay (like $\varphi^k$): short memory, typical of a stationary series with carry-over.
- Very slow, almost straight-line decay from near 1: a trend or a random walk. The series is not stationary; difference or detrend it first (Chapter 7.4).
- Peaks at 7, 14, 21 (daily data): a weekly pattern. Peaks at 12, 24 in monthly data: a yearly pattern.
- Alternating signs: zigzag behaviour (negative carry-over, or over-differencing).
- All sticks inside the band (next section): no linear memory that this much data can detect.
For the AR(1) series $y_t = \varphi y_{t-1} + \epsilon_t$ with $|\varphi| \lt 1$: $\rho(k) = \varphi^k$.
Why do we need it?
One picture shows how far back the past matters, whether there is a weekly or yearly cycle, and whether the series is wandering (non-stationary). Reading numbers lag by lag would hide the pattern.
Where is it used?
The first plot in every time-series analysis; ARIMA/SARIMA order selection (Box–Jenkins method, Chapter 7.6); choosing seasonal periods for Fourier terms (Chapter 7.11); residual diagnostics of your forecasting model (Chapter 7.17); MCMC chain autocorrelation.
How is it used?
from statsmodels.graphics.tsaplots import plot_acf, then plot_acf(y, lags=40). Look at three things: how fast the sticks fade, where the peaks are, and which sticks cross the shaded band. Do it on the raw series, then again on the residuals.
"The ACF of my series decays slowly, so the series has a long, strong memory I should model with many AR lags."
A slowly fading ACF is the classic sign of a trend or a random walk, not of a stationary process with many lags. The ACF formula subtracts one overall mean; when the level drifts, every value is far from that mean in the same direction as its neighbours, so all $r_k$ are large. Remove the trend (or difference) first, then read the ACF again.
"Each stick is a separate, independent piece of evidence."
For an AR(1) series all sticks come from one carry-over number $\varphi$: lag 2 is large only because lag 1 is. The PACF (below) separates direct from indirect links.
ACF = $r_k$ for $k = 0, 1, \dots$, drawn as sticks. AR(1): $\rho(k) = \varphi^k$ (geometric decay; alternating if $\varphi \lt 0$).
Slow straight-line decay → trend / random walk (non-stationary). Peaks at 7, 14, 21 → weekly pattern.
Trap: read the ACF of the residuals, not only of the raw series.
Quick check: daily orders have $r_1 = 0.45$, $r_2 = 0.2$, and peaks $r_7 = 0.6$, $r_{14} = 0.5$. What do you conclude?
A short-term memory (lag 1 and 2) plus a strong weekly pattern (lags 7 and 14). In your forecasting model the weekly Fourier terms are meant to capture the weekly part; whether the short-term memory remains should be checked in the residual ACF.
Is this spike real? The ±1.96/√n band and the Ljung–Box test core
Even a series with no memory at all (every day an independent coin toss) gives sample autocorrelations that are not exactly zero. By luck there are always a few short runs. So we need a ruler: how big can $r_k$ get by chance alone?
For a memoryless series of length $n$, each $r_k$ behaves roughly like a Normal number with mean 0 and standard deviation $1/\sqrt n$. About 95% of such chance values land within $\pm 1.96/\sqrt n$. That is the shaded band in every ACF plot: the noise floor.
Three ways to say it:
- Picture: the band is the water level; only sticks that rise above the water are worth a second look.
- Numbers: with 100 days the band is $\pm 1.96/\sqrt{100} = \pm 0.196$; with 365 days it shrinks to about $\pm 0.103$.
- Slogan: a spike must clear the noise floor, and with many lags a few will clear it by luck.
You have $n = 100$ days of residuals and look at 20 lags.
- Standard deviation of $r_k$ under "no memory": $1/\sqrt{100} = 0.1$.
- Band: $\pm 1.96 \times 0.1 = \pm 0.196$.
- $r_1 = 0.31$ is outside the band ($0.31 \gt 0.196$): evidence of lag-1 memory.
- $r_9 = -0.21$ is just outside, and it is the only other spike among 20 lags. Each lag has about a 5% chance of a false alarm, so among 20 lags up to about $20 \times 0.05 = 1$ false spike is expected. One lone, barely-outside spike at an odd lag is weak evidence.
- To halve the band to $\pm 0.098$ you need four times the data: $n = 400$ (precisely $n = (1.96/0.098)^2 = 400$).
If $y_1, \dots, y_n$ are independent and identically distributed with finite variance (the null hypothesis "no memory"; hypothesis tests are explained in Chapter 5.6), then for large $n$
$$r_k \approx N\!\left(0, \tfrac{1}{n}\right) \text{ for each lag } k \ge 1, \qquad \text{95\% band: } \pm\frac{1.96}{\sqrt n}.$$- Different lags are approximately independent under this null, so among $m$ lags about $0.05\,m$ fall outside by chance (a little fewer in short series, because dividing by $n$ shrinks $r_k$ at large lags).
- Bartlett band (what
plot_acfdraws by default): at lag $k$ it uses $\pm 1.96\sqrt{(1 + 2\sum_{j=1}^{k-1} r_j^2)/n}$. It asks "is lag $k$ zero, given the earlier lags?", so it widens after big early sticks.pandas.plotting.autocorrelation_plotdraws flat bands. - Ljung–Box test: one test for all lags $1, \dots, h$ together, $$Q = n(n+2)\sum_{k=1}^{h}\frac{r_k^2}{n-k},$$ which is approximately chi-square with $h$ degrees of freedom when the series has no memory. A small p-value says "there is memory somewhere in the first $h$ lags". (When testing the residuals of a fitted ARMA model, the degrees of freedom are reduced by the number of fitted coefficients; see Chapter 7.17.)
Why do we need it?
Without a reference for "how big is big", every wiggle in an ACF looks meaningful. The band and the Ljung–Box test separate real memory from the chance correlations that any finite series shows.
Where is it used?
Every plot_acf/plot_pacf figure, residual checks after ARIMA, ETS or your Prophet-style model, acorr_ljungbox in statsmodels' diagnostics, and the "residuals are white noise" check in forecasting courses and production monitoring.
How is it used?
Look for sticks clearly outside the band at meaningful lags (1, 2, 7, 14). Ignore a lone barely-outside stick at an odd lag. Then run acorr_ljungbox(resid, lags=[10, 20]); a p-value below about 0.05 says the residuals still contain memory.
"One stick out of 40 crosses the band, so the residuals are not white noise."
With 40 lags, one or two sticks typically cross the 95% band by chance even for perfect white noise. Judge the overall pattern (or Ljung–Box), and give most weight to lags that make sense (1, 2, 7, 14).
"All sticks are inside the band, so there is no autocorrelation."
Inside the band means "not detectable with this $n$". With 60 days the band is $\pm 0.25$, so a real $\rho(1) = 0.2$ can hide inside it. Absence of evidence is not evidence of absence (Chapter 5.6).
"The band is a confidence interval for the true autocorrelation."
The flat band is centred on 0: it shows where $r_k$ would fall if there were no memory. It is a test region, not an interval around your estimate.
No memory (iid): $r_k \approx N(0, 1/n)$ → band $\pm 1.96/\sqrt n$ (n = 100 → ±0.196).
Roughly 1 in 20 lags crosses by chance (a bit fewer in short series). plot_acf uses widening Bartlett bands.
Ljung–Box $Q = n(n+2)\sum_{k\le h} r_k^2/(n-k) \sim \chi^2_h$: tests many lags at once; small p = memory somewhere.
Quick check: how many days do you need for the white-noise band to be ±0.05?
$1.96/\sqrt n = 0.05$ gives $\sqrt n = 39.2$, so $n \approx 1537$ days (about 4.2 years of daily data).
The PACF: the direct link after removing the in-between days core
Grandma tells Mum a story, and Mum tells you. Your version and Grandma's version are similar, but only because of Mum. If you already know exactly what Mum said, Grandma's version adds nothing new: there is no direct line from Grandma to you.
Time series work the same way. In "today = 0.8 × yesterday + surprise", today is correlated with the day before yesterday ($\rho(2) = 0.64$), but only through yesterday. The ACF counts this indirect link. The partial autocorrelation (PACF) at lag $k$ asks a sharper question: after I already know the days in between, does the day $k$ steps back still tell me anything?
Three ways to say it:
- Picture: a chain $y_{t-2} \to y_{t-1} \to y_t$. The ACF measures everything that flows from $y_{t-2}$ to $y_t$; the PACF measures only a bypass arrow that skips $y_{t-1}$.
- Numbers: for "0.8 × yesterday", ACF at lag 2 = 0.64 but PACF at lag 2 = 0.
- Slogan: the PACF is what lag $k$ adds once the closer lags are known.
For lag 2 there is a short formula (derived in the definition): $\alpha(2) = \dfrac{\rho(2) - \rho(1)^2}{1 - \rho(1)^2}$.
- AR(1) with $\varphi = 0.8$: $\rho(1) = 0.8$, $\rho(2) = 0.64$. Then $\alpha(2) = (0.64 - 0.8^2)/(1 - 0.8^2) = (0.64 - 0.64)/0.36 = 0$. The lag-2 link is entirely "through yesterday".
- A series with $\rho(1) = 0.5$ and $\rho(2) = 0.5$: $\alpha(2) = (0.5 - 0.25)/(1 - 0.25) = 0.25/0.75 \approx 0.333$. If the link went only through yesterday we would expect $\rho(2) = 0.5^2 = 0.25$; the extra $0.25$ is a direct lag-2 effect.
- Our 8 days of orders: $r_1 = 0.5$, $r_2 \approx -0.179$. So $\alpha(2) \approx (-0.179 - 0.25)/0.75 \approx -0.571$. This is exactly what
pacf(y, method="ywm")returns. - At lag 1 nothing lies in between, so $\alpha(1) = \rho(1)$ always (0.5 for the 8 days).
The partial autocorrelation at lag $k$, written $\alpha(k)$ or $\varphi_{kk}$, has two equivalent descriptions:
- Residual view: predict $y_t$ from the in-between values $y_{t-1}, \dots, y_{t-k+1}$ (best straight-line prediction), and do the same for $y_{t-k}$. $\alpha(k)$ is the correlation of the two leftovers (residuals). This is the partial correlation of Chapter 4.15, applied to lags.
- Regression view: regress $y_t$ on $y_{t-1}, \dots, y_{t-k}$; $\alpha(k)$ is the coefficient of the last lag $y_{t-k}$.
Lag 2, step by step (standardize so every $y$ has variance 1; then the best prediction of $y_t$ from $y_{t-1}$ is $\rho(1) y_{t-1}$, and the same for $y_{t-2}$ by symmetry):
$$\begin{aligned} Cov\big(y_t - \rho_1 y_{t-1},\; y_{t-2} - \rho_1 y_{t-1}\big) &= \rho_2 - \rho_1\rho_1 - \rho_1\rho_1 + \rho_1^2 = \rho_2 - \rho_1^2,\\ Var\big(y_t - \rho_1 y_{t-1}\big) &= 1 - 2\rho_1^2 + \rho_1^2 = 1 - \rho_1^2 \;\;(\text{same for the other residual}),\\ \alpha(2) &= \frac{\rho_2 - \rho_1^2}{1 - \rho_1^2}. \end{aligned}$$Durbin–Levinson recursion (how software gets every lag from the ACF): $\varphi_{11} = \rho_1$ and, for $k \ge 2$,
$$\varphi_{kk} = \frac{\rho_k - \sum_{j=1}^{k-1}\varphi_{k-1,j}\,\rho_{k-j}}{1 - \sum_{j=1}^{k-1}\varphi_{k-1,j}\,\rho_j}, \qquad \varphi_{k,j} = \varphi_{k-1,j} - \varphi_{kk}\,\varphi_{k-1,k-j}\;\;(j \lt k).$$Key property: for an AR($p$) series, $\alpha(k) = 0$ for every $k \gt p$ (the PACF "cuts off"). The white-noise band for the PACF is the same $\pm 1.96/\sqrt n$.
Why do we need it?
The ACF of a sticky series is large at many lags even when only yesterday matters directly. The PACF removes this "echo through the middle", so we can see how many past days genuinely carry their own information.
Where is it used?
Choosing the order $p$ of an AR or ARIMA model (Box–Jenkins, Chapter 7.6), deciding how many lag features a regression or gradient-boosting forecaster needs, and diagnosing what structure is left in residuals.
How is it used?
plot_pacf(y, lags=30, method="ywm") or pacf(y, nlags=30, method="ywm"). Count the significant spikes before the PACF drops inside the band: that count is a first guess for $p$. Always look at ACF and PACF side by side.
"pacf(y) and plot_pacf(y) give the same numbers."
Not by default. pacf uses method="ywadjusted" (it divides lag-$k$ sums by $n - k$), while plot_pacf uses method="ywm" (the usual ACF with denominator $n$, as above). For long series they nearly agree; for our 8 days, lag 1 is 0.571 versus 0.5. Pass method= explicitly and say which one you used.
"PACF at lag 2 = the correlation between today and two days ago."
That is the ACF. The PACF is the correlation that is left after the in-between day has been accounted for; it can be 0 even when the ACF is large, and it can even have the opposite sign.
PACF $\alpha(k)$ = correlation of $y_t$ and $y_{t-k}$ after removing the straight-line effect of $y_{t-1}, \dots, y_{t-k+1}$ = last coefficient of an AR($k$) regression.
$\alpha(1) = \rho_1$; $\alpha(2) = (\rho_2 - \rho_1^2)/(1 - \rho_1^2)$; Durbin–Levinson for higher lags. AR($p$): PACF = 0 after lag $p$.
Trap: statsmodels pacf default (ywadjusted) ≠ plot_pacf default (ywm).
Quick check: $\rho_1 = 0.6$ and $\rho_2 = 0.36$. What is the PACF at lag 2, and what kind of series could this be?
$\alpha(2) = (0.36 - 0.36)/(1 - 0.36) = 0$. The lag-2 link is fully explained by lag 1: consistent with an AR(1) with $\varphi = 0.6$ (since $0.6^2 = 0.36$).
Reading ACF and PACF together: the signatures
Doctors read two charts together (pulse and blood pressure) because each alone is ambiguous. Forecasters read the ACF and PACF together for the same reason. Each common kind of series leaves a recognizable signature: which plot decays slowly and which one stops suddenly.
Two new building blocks appear here. AR (autoregressive) series remember their past values: "today = some of yesterday + surprise". MA (moving-average) series remember past surprises: "today = today's surprise + part of yesterday's surprise", like a shock whose effect lasts exactly two days. (The full ARIMA family is Chapter 7.6; here we only read their fingerprints.)
Three ways to say it:
- Picture: AR = ACF fades, PACF stops. MA = ACF stops, PACF fades. They are mirror images.
- Numbers: AR(1) with $\varphi = 0.7$: ACF 0.7, 0.49, 0.34…, PACF 0.7, 0, 0… MA(1) with $\theta = 0.6$: ACF 0.44, 0, 0…, PACF 0.44, −0.24, 0.14…
- Slogan: the plot that stops tells you the order; the plot that fades tells you the type.
Three small calculations.
- MA(1), $y_t = \epsilon_t + 0.6\,\epsilon_{t-1}$ (surprises with variance 1). Today and yesterday share one surprise, $\epsilon_{t-1}$: $Cov(y_t, y_{t-1}) = 0.6 \times 1 = 0.6$. $Var(y_t) = 1 + 0.6^2 = 1.36$. So $\rho_1 = 0.6/1.36 \approx 0.441$. Today and two days ago share no surprise, so $\rho_2 = 0$: the ACF stops after lag 1.
- AR(2), $y_t = 0.5\,y_{t-1} + 0.3\,y_{t-2} + \epsilon_t$: $\rho_1 = \varphi_1/(1 - \varphi_2) = 0.5/0.7 \approx 0.714$ and $\rho_2 = \varphi_1\rho_1 + \varphi_2 = 0.5 \times 0.714 + 0.3 \approx 0.657$. The PACF is $0.714$, then exactly $\varphi_2 = 0.3$, then 0.
- AR(2) with a wave, $\varphi_1 = 1.0$, $\varphi_2 = -0.5$: $\rho_1 = 1/1.5 \approx 0.667$, $\rho_2 = 0.667 - 0.5 \approx 0.167$, $\rho_3 = 0.167 - 0.5 \times 0.667 \approx -0.167$, $\rho_4 = -0.167 - 0.5 \times 0.167 = -0.25$. The ACF swings below zero and back: a damped wave. The series looks cyclic, yet it is stationary.
| Series | ACF | PACF |
|---|---|---|
| White noise | all ≈ 0 | all ≈ 0 |
| AR($p$): $y_t = \sum_{i=1}^{p}\varphi_i y_{t-i} + \epsilon_t$ | decays (geometric, or a damped wave) | cuts off after lag $p$ |
| MA($q$): $y_t = \epsilon_t + \sum_{j=1}^{q}\theta_j\epsilon_{t-j}$ | cuts off after lag $q$ | decays |
| Random walk or trend | near 1, fades very slowly | one big spike (≈ 1) at lag 1 |
| Weekly pattern (daily data) | peaks at 7, 14, 21, … | spike(s) at or near lag 7 |
Useful facts: MA(1) has $\rho_1 = \theta/(1+\theta^2)$, which is never larger than 0.5 in size. AR(2) is stationary only inside the triangle $\varphi_1 + \varphi_2 \lt 1$, $\varphi_2 - \varphi_1 \lt 1$, $|\varphi_2| \lt 1$; it shows a damped wave when $\varphi_1^2 + 4\varphi_2 \lt 0$.
Why do we need it?
The signature turns two plots into a first model guess (AR or MA, which order, seasonal or not) and tells you when a series must be differenced before anything else, instead of trying models blindly.
Where is it used?
The Box–Jenkins workflow for ARIMA/SARIMA (Chapter 7.6), choosing lag features for gradient-boosting forecasters, and residual diagnostics: an AR(1)-shaped residual ACF says "add one lag of memory" (Chapter 7.17).
How is it used?
Plot plot_acf and plot_pacf side by side. Slow ACF decay → difference first. Otherwise: PACF stops at $p$ → try AR($p$); ACF stops at $q$ → try MA($q$); both fade → ARMA. Then check the residuals of the chosen model.
"The PACF stops after lag 2, so the true process is exactly AR(2)."
The signature is a starting guess. Sample ACF/PACF values are noisy (look at $n = 60$ in the widget), real series mix AR and MA parts, and several models can fit almost equally well. Compare candidate models on held-out data and check their residuals.
"A wavy series must have a seasonal pattern."
An AR(2) with $\varphi_1^2 + 4\varphi_2 \lt 0$ produces irregular waves (pseudo-cycles) with no fixed period. A real seasonal pattern repeats at a fixed lag (7, 12, 365) and shows peaks at its multiples. Cycles of varying length are different from seasonality (Chapter 7.2).
AR($p$): ACF fades, PACF stops after $p$. MA($q$): ACF stops after $q$, PACF fades. Both fade → ARMA.
Random walk / trend: ACF ≈ 1 fading slowly, PACF one spike → difference first. Seasonal: peaks at multiples of the period.
MA(1): $\rho_1 = \theta/(1+\theta^2)$, $|\rho_1| \le 0.5$. Trap: a signature is a first guess, not a proof.
Quick check: the ACF has one clear spike at lag 1 ($r_1 = -0.45$) and nothing else; the PACF fades with negative values. What is your first guess?
An MA(1) with a negative $\theta$: the ACF stops after lag 1 and the PACF fades. Solving $\theta/(1+\theta^2) = -0.45$ gives $\theta \approx -0.63$ (the other solution, $\theta \approx -1.59$, gives the same ACF but is usually excluded). This shape often appears after differencing a series that did not need it (Chapter 7.4).
White noise: the series with nothing left to learn core
Remember the hiss of an untuned radio or the "snow" on an old TV: no tune, no pattern, nothing you could predict. A white-noise series is the time-series version: every value has the same average and the same spread, and the past says nothing (in a straight-line sense) about the next value. (The name comes from white light, which mixes all colours equally; white noise mixes all frequencies equally.)
White noise is the benchmark of "done": a good forecasting model should leave residuals that look like white noise. If they do not, there is structure the model has not used.
Three ways to say it:
- Picture: a fuzzy horizontal band: no trend, no waves, no runs, no bursts of calm or chaos.
- Numbers: mean 0, constant standard deviation $\sigma$, and every autocorrelation $\rho(k) = 0$ for $k \ge 1$ (sample values inside $\pm 1.96/\sqrt n$).
- Slogan: white noise is the part nobody can predict; it is what should be left after a perfect model.
Each day toss a fair coin: $\epsilon_t = +1$ for heads, $-1$ for tails.
- Mean: $E[\epsilon_t] = 0.5(+1) + 0.5(-1) = 0$, the same every day.
- Variance: $E[\epsilon_t^2] - 0^2 = 1$, the same every day.
- Different days are independent tosses, so $Cov(\epsilon_t, \epsilon_s) = E[\epsilon_t\epsilon_s] = E[\epsilon_t]E[\epsilon_s] = 0$ for $t \ne s$.
- So this is white noise. The best forecast of tomorrow is the mean, 0; knowing the last 100 tosses does not help.
A trickier case: "big days follow big days, but the direction is a coin toss" (calm weeks and turbulent weeks). The values are uncorrelated (the direction is unpredictable), but their sizes are not independent. This is white noise that is not independent (see the widget).
A series $\epsilon_t$ is white noise, written $\epsilon_t \sim WN(0, \sigma^2)$, if for all times $t$ and $s \ne t$:
$$E[\epsilon_t] = 0, \qquad Var(\epsilon_t) = \sigma^2 \lt \infty, \qquad Cov(\epsilon_t, \epsilon_s) = 0.$$- Its ACF and PACF are 0 at every lag $k \ge 1$.
- iid noise (independent and identically distributed, with finite variance) is a stronger condition: independence rules out every kind of link, not only straight-line ones. Every iid series with finite variance and mean 0 is white noise; not every white noise is iid.
- Gaussian white noise, $\epsilon_t \overset{iid}{\sim} N(0, \sigma^2)$, is the strongest: for jointly Normal values, uncorrelated already means independent.
- Some books allow a constant non-zero mean; "white noise" is about the absence of memory and the constant spread.
Why do we need it?
It is the yardstick for "no exploitable structure". Every test of residuals (ACF band, Ljung–Box) compares your leftovers against white noise. It is also the raw ingredient from which AR, MA and random-walk series are built.
Where is it used?
The noise term $\epsilon_t$ of ARIMA, ETS and your Prophet-style model; residual diagnostics (Chapter 7.17); the shocks of a random walk (next section); simulation of synthetic series for testing forecasting code.
How is it used?
After fitting a model, plot the residuals over time and their ACF, run acorr_ljungbox, and also look at the ACF of the squared residuals to catch changing volatility. If all look like white noise, the remaining error is (linearly) unpredictable.
"White noise means the values are independent."
White noise only requires zero correlation (no straight-line memory) and constant variance. Volatility clustering (calm and turbulent periods) is uncorrelated yet dependent. Check the ACF of squared or absolute residuals to see it.
"White noise means Normally distributed."
Coin tosses (±1) and Student-t surprises are white noise too. "White" is about memory, not about the bell shape. Whether the residuals look Normal is a separate check (Q-Q plot, Chapter 4.17).
"My residuals are white noise, so my model is right."
It means there is no linear, time-ordered structure left. The model can still have the wrong level for future regimes, miss non-linear effects, or be badly calibrated. White residuals are necessary, not sufficient.
In your forecasting model, writing $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$ with $\epsilon_t \sim N(0, \sigma^2)$ or Student-t, independently for each day, is exactly the assumption "the leftover is iid noise" (the inner boxes of the figure). With the Negative Binomial likelihood the same idea reads "counts are independent given their means". So after fitting, the residuals (for NB, use standardized residuals $(y_t - \mu_t)/\sqrt{\mu_t + \mu_t^2/\alpha}$, where $\alpha$ is the NB concentration, so the variance is $\mu_t + \mu_t^2/\alpha$; check which parameterization your code uses) should look like white noise: ACF inside the band, Ljung–Box not significant, and no clustering in the squared residuals. Choosing Student-t makes the model robust to occasional huge days; it does not model memory or clustering.
$WN(0, \sigma^2)$: $E = 0$, $Var = \sigma^2$ constant, $Cov(\epsilon_t, \epsilon_s) = 0$ for $t \ne s$. ACF = PACF = 0 at all lags ≥ 1.
Gaussian iid ⊂ iid ⊂ white noise. White = no linear memory; not necessarily independent or Normal.
Good residuals look like white noise; also check the ACF of squared residuals.
Quick check: is $\epsilon_t = Z_t Z_{t-1}$, with $Z_t$ iid $N(0, 1)$, white noise? Is it independent?
$E[\epsilon_t] = E[Z_t]E[Z_{t-1}] = 0$; $Var(\epsilon_t) = E[Z_t^2]E[Z_{t-1}^2] = 1$; $Cov(\epsilon_t, \epsilon_{t-1}) = E[Z_t Z_{t-1}^2 Z_{t-2}] = 0 \cdot 1 \cdot 0 = 0$, and larger lags are 0 too. So it is white noise. But it is not independent: $\epsilon_t$ and $\epsilon_{t-1}$ share $Z_{t-1}$, so a huge $|\epsilon_{t-1}|$ makes a huge $|\epsilon_t|$ more likely (their squares are correlated).
The random walk: white noise, added up core
Picture someone walking along a line who, at every step, flips a coin: heads one step right, tails one step left. Where is the walker after 100 steps? Each step is pure noise, but the position is the sum of all steps so far. Nothing pulls the walker back to the start. Sometimes they drift far away and stay there for a long time.
That is a random walk: today = yesterday + a fresh random step. Stock prices, the level of a cumulative balance, and many "level" series behave roughly like this. It is the opposite extreme to white noise: white noise forgets everything; a random walk forgets nothing, because every past step is still inside today's value.
Three ways to say it:
- Picture: many walkers starting together fan out into a widening cone; none of them is pulled back to the middle.
- Numbers: with daily steps of standard deviation 10 orders, the spread after 100 days is $10\sqrt{100} = 100$ orders, after 400 days $10\sqrt{400} = 200$.
- Slogan: a random walk is white noise added up; it has no home to return to.
Start at $y_0 = 100$ with steps $\epsilon = +2, -1, +3, +1, -2$.
- $y_1 = 100 + 2 = 102$, $y_2 = 102 - 1 = 101$, $y_3 = 101 + 3 = 104$, $y_4 = 104 + 1 = 105$, $y_5 = 105 - 2 = 103$.
- So $y_5 = y_0 + (2 - 1 + 3 + 1 - 2) = 100 + 3$: the value is the start plus the sum of all steps.
- Differences give the steps back: $102 - 100 = 2$, $101 - 102 = -1$, … Differencing turns a random walk into white noise (Chapter 7.4).
- Spread: the steps are independent with variance $\sigma^2$, and variances of independent pieces add (Chapter 4.5), so $Var(y_t) = t\sigma^2$ and $sd(y_t) = \sigma\sqrt t$. With $\sigma = 10$: $sd(y_{100}) = 10 \times 10 = 100$.
- Memory: $y_{110}$ contains all of $y_{100}$ plus 10 new steps, so $Corr(y_{100}, y_{110}) = \sqrt{100/110} \approx 0.953$. But $Corr(y_{10}, y_{20}) = \sqrt{10/20} \approx 0.707$. The correlation depends on when you look, not only on the lag.
A random walk is $y_t = y_{t-1} + \epsilon_t$ with $\epsilon_t \sim WN(0, \sigma^2)$. A random walk with drift adds a constant step $\mu$: $y_t = \mu + y_{t-1} + \epsilon_t$. Unrolling,
$$y_t = y_0 + \mu t + \sum_{i=1}^{t}\epsilon_i, \qquad E[y_t] = y_0 + \mu t, \qquad Var(y_t \mid y_0) = t\sigma^2, \qquad Corr(y_t, y_{t+k}) = \sqrt{\tfrac{t}{t+k}}.$$- Not stationary: the variance grows with $t$ and the correlation depends on $t$ (Chapter 7.4). It is an AR(1) with $\varphi = 1$, which is called a unit root.
- Sample ACF: starts near 1 and fades very slowly, almost in a straight line; PACF: one spike near 1.
- Its first difference $\nabla y_t = y_t - y_{t-1} = \mu + \epsilon_t$ is white noise (plus the drift). A series that becomes stationary after one difference is called integrated of order 1, I(1).
- Best forecast: $\hat y_{T+h} = y_T + h\mu$ (the last value, plus drift: the "naive" forecast of Chapter 7.5), with forecast standard deviation $\sigma\sqrt h$, which grows without limit.
Why do we need it?
It is the reference model for "wandering" series. Many business levels behave like it, and treating a random walk as if it were stable around a mean gives terrible forecasts, fake correlations between unrelated series, and much too narrow intervals.
Where is it used?
The naive and drift forecasts (Chapter 7.5), the "I" in ARIMA, unit-root tests (ADF, KPSS, Chapter 7.4), random-walk priors for slowly changing parameters in Bayesian state-space models, and random-walk Metropolis proposals in MCMC (Chapter 6.9).
How is it used?
Simulate with np.cumsum(rng.normal(size=n)). Recognize it by a wandering path and an ACF that refuses to die; confirm with a unit-root test; difference it before ARIMA modelling; always compare your forecasts with the naive "tomorrow = today" baseline.
"The walk has been above its starting value for 150 days, so it is due to come back."
A random walk has no level to return to. Its best forecast is simply its last value (plus drift). Long excursions are normal, not a sign that a reversal is coming.
"Two of my business metrics are strongly correlated over the last year, so they are related."
Two independent random walks often show $|r| \gt 0.5$ just because both wander (spurious correlation; try the "Two unrelated trends look correlated" widget in Chapter 4.15). Correlate their day-to-day changes instead.
"A slowly fading ACF tells me the true autocorrelation is about 0.9 at lag 5."
For a random walk there is no single true $\rho(k)$ (it depends on $t$). The slowly fading sample ACF is a symptom of non-stationarity, not an estimate of a stable memory.
$y_t = \mu + y_{t-1} + \epsilon_t$ ⇒ $y_t = y_0 + \mu t + \sum \epsilon_i$; $Var(y_t) = t\sigma^2$, sd $= \sigma\sqrt t$; $Corr(y_t, y_{t+k}) = \sqrt{t/(t+k)}$.
Random walk = AR(1) with $\varphi = 1$ (unit root); not stationary; differencing gives white noise. Forecast: last value + drift, interval width grows like $\sqrt h$.
Traps: "it will come back", and spurious correlation between two wandering series.
Quick check: a random walk has steps with $\sigma = 5$. What is the 95% range for the value 36 days ahead, given today's value 200 and no drift?
Forecast $200$; sd $= 5\sqrt{36} = 30$; 95% range $200 \pm 1.96 \times 30 = 200 \pm 58.8$, i.e. about 141 to 259. For 9 days ahead it would be only $200 \pm 1.96 \times 15 = 200 \pm 29.4$.
Serial dependence and standard errors: correlated days are worth less core
Suppose you ask 50 people for their opinion, but each person mostly repeats what the previous person said. You have 50 answers, yet much less than 50 people's worth of information. Positively autocorrelated days are like that: each day partly repeats yesterday.
Formulas such as $SE = s/\sqrt n$ assume that every row brings new, independent information. With serial dependence they are too small: intervals too narrow, p-values too small, and "significant" results that are not. (The full derivation is in Chapter 4.13; here we see what it does to time-series work.)
Three ways to say it:
- Picture: 50 sticky days are like 17 truly separate days (when the carry-over is 0.5).
- Numbers: carry-over $\varphi = 0.8$ makes the variance of an average about 9 times the naive value, so the honest standard error is about 3 times the naive one; a "95%" interval for the mean then covers the truth only about 45% of the time (50 days, simulated).
- Slogan: count information, not rows.
100 days of residuals with standard deviation $s = 6$ and lag-1 autocorrelation about 0.5 (AR(1)-like). Is their average really close to 0?
- Naive standard error: $6/\sqrt{100} = 0.6$. Naive 95% half-width: $1.96 \times 0.6 \approx 1.18$.
- Inflation factor for an AR(1), large $n$: $\dfrac{1+\varphi}{1-\varphi} = \dfrac{1.5}{0.5} = 3$.
- Honest standard error: $0.6 \times \sqrt 3 \approx 1.04$. Honest half-width: $1.96 \times 1.04 \approx 2.04$.
- Effective sample size: $n_{\text{eff}} = 100/3 \approx 33$ independent days' worth of information.
- An average residual of 1.5 looks "significant" with the naive SE ($1.5 \gt 1.18$) but not with the honest one ($1.5 \lt 2.04$).
For a stationary series with variance $\sigma^2$ and autocorrelations $\rho_k$ (from Chapter 4.13):
$$Var(\bar y) = \frac{\sigma^2}{n}\Big[1 + 2\sum_{k=1}^{n-1}\Big(1 - \frac kn\Big)\rho_k\Big], \qquad n_{\text{eff}} = \frac{n}{1 + 2\sum_{k=1}^{n-1}(1 - k/n)\rho_k}.$$- AR(1): the bracket tends to $\dfrac{1+\varphi}{1-\varphi}$ as $n$ grows ($\varphi = 0.5 \to 3$, $\varphi = 0.8 \to 9$). Negative $\varphi$ makes it smaller than 1.
- Regression with autocorrelated errors: the least-squares coefficients are still unbiased when the mean model is right, but the usual standard errors are wrong, typically too small when both the errors and the regressor (time itself, a smooth seasonal term) are positively autocorrelated. Fixes: HAC (heteroscedasticity- and autocorrelation-consistent, Newey–West) standard errors, modelling the errors (AR errors, GLS), or a block bootstrap (resampling whole blocks of days).
- Bayesian version: a likelihood that multiplies $n$ independent terms "believes" it has $n$ independent days, so the posterior is too narrow in the same way.
Why do we need it?
Every interval, test and posterior built on independent-row formulas is overconfident when days are correlated. Knowing the inflation factor tells you how much to trust a time-series result, or how much more data you really need.
Where is it used?
Confidence intervals for average forecast errors, comparing two forecasting models on the same days (the Diebold–Mariano test uses HAC variances), regression on time series (HAC/Newey–West), MCMC effective sample size (Chapter 6.10), and A/B metrics computed on repeated days per user.
How is it used?
Estimate $r_1$ (or the residual ACF), compute $n_{\text{eff}}$, and report it. In statsmodels use .fit(cov_type="HAC", cov_kwds={"maxlags": L}). In a Bayesian model, model the dependence explicitly (an AR error term) instead of assuming independent days.
"I have 730 days of data, so my standard errors are tiny."
If the residuals have $r_1 = 0.8$, those 730 days carry roughly the information of $730/9 \approx 81$ independent days. Report $n_{\text{eff}}$ or use HAC standard errors.
"Autocorrelation biases my regression coefficients."
With a correct mean model and no lagged target among the regressors, the coefficients stay unbiased; it is their uncertainty that is misstated (and the estimates are less efficient). Bias appears if you include a lagged target while the errors are autocorrelated.
In your forecasting model, the likelihood multiplies one independent term per day. If the true residuals carry over from day to day, the posterior for the trend slope, the changepoint adjustments $\delta_j$ and the seasonal coefficients is narrower than it should be, exactly like the naive intervals in the widget, and runs of positive residuals can look like trend changes. In an A/B framework like yours, the same arithmetic appears whenever a unit contributes several correlated observations (several days or sessions of one user): the randomization unit is the unit of independence (Chapter 5.10).
$Var(\bar y) = \frac{\sigma^2}{n}[1 + 2\sum(1 - k/n)\rho_k]$; AR(1) factor $\approx (1+\varphi)/(1-\varphi)$; $n_{\text{eff}} = n/\text{factor}$.
Positive autocorrelation → naive SEs too small, intervals too narrow, posteriors overconfident.
Fixes: HAC (Newey–West) SEs, model the AR errors, block bootstrap. Trap: count information, not rows.
Quick check: 365 days of residuals with AR(1) carry-over 0.6. Roughly how many independent days are they worth, and by what factor is the naive SE too small?
Factor $(1 + 0.6)/(1 - 0.6) = 1.6/0.4 = 4$, so $n_{\text{eff}} \approx 365/4 \approx 91$ days, and the naive SE is too small by $\sqrt 4 = 2$.
What a Prophet-style model does (and does not do) with autocorrelation core
Your forecasting model draws the expected path from the calendar: a trend, a weekly and yearly rhythm, holiday bumps, regressor effects. Around that path it assumes independent noise, a fresh surprise every day. It never asks "was yesterday unusually high?".
Real demand often has short-term momentum that the calendar cannot see: a warm spell, a viral post, a competitor's stock-out. If Monday was 10 orders above the model, Tuesday is likely to be above it too. A Prophet-style model has no term for this, so the momentum stays in the residuals as autocorrelation.
Three ways to say it:
- Picture: the model's line follows the calendar; the real series wiggles around it in runs, and the model never looks at the last run.
- Numbers: if the residuals have $r_1 = 0.6$, using yesterday's residual would shrink the typical one-step-ahead error by a factor $\sqrt{1 - 0.6^2} = 0.8$ (20% smaller).
- Slogan: Prophet-style = calendar + independent noise; ARIMA = memory. Leftover memory shows up in the residual ACF.
Residuals of a fitted Prophet-style model have $r_1 = 0.6$ and standard deviation 6.25 orders. Yesterday (the forecast origin $T$) ended $e_T = +10$ orders above the model.
- An AR(1) view of the residuals: $e_{t} \approx 0.6\,e_{t-1} + \eta_t$, where $\eta_t$ is the truly new surprise.
- One day ahead the expected leftover is $0.6 \times 10 = 6$: add 6 orders to the model's forecast for tomorrow.
- Two days ahead: $0.6^2 \times 10 = 3.6$. Seven days ahead: $0.6^7 \times 10 \approx 0.28$. The correction fades geometrically.
- Typical one-step error: from $6.25$ (model alone) to $6.25 \times \sqrt{1 - 0.36} = 6.25 \times 0.8 = 5.0$ with the correction.
- Conclusion: the missing memory costs accuracy at short horizons (1–3 days here) and matters little for horizons of a few weeks.
The Prophet-style model is a structural regression on time:
$$y_t = \mu_t + \epsilon_t, \qquad \mu_t = g(t) + s(t) + h(t) + X_t\beta, \qquad \epsilon_t \text{ independent across days}$$(Normal or Student-t noise; for the Negative Binomial likelihood, counts independent given $\mu_t$). Prophet itself has no autoregressive term. Two kinds of autocorrelation must be separated:
- Autocorrelation of $y_t$ caused by the structure (trend, seasonality, holidays, slowly moving regressors): the model does capture this, through $\mu_t$.
- Autocorrelation of the residuals $e_t = y_t - \hat\mu_t$ (short-term memory): the model does not model it.
If residuals follow $e_t = \varphi e_{t-1} + \eta_t$ with $\varphi \ne 0$:
- Point forecasts ignore recent residuals, so short horizons are less accurate than they could be; long horizons are nearly unaffected ($\varphi^h \to 0$).
- Parameter uncertainty is understated: the likelihood counts $n$ independent days (previous section), and runs of residuals can be mistaken for changepoints.
- Uncertainty of sums is understated: the variance of a weekly total includes covariance terms that independent noise sets to 0.
Responses: add an AR residual term $\epsilon_t = \varphi\epsilon_{t-1} + \eta_t$ (in NumPyro with jax.lax.scan, Chapter 6.16); add lagged-target features (only lags available at forecast time, Chapter 7.12); add missing structure (richer seasonality, holidays, regressors), because residual memory often means a missed component; or use a state-space model or a Gaussian process for the residual (Chapter 7.18). Always check with the residual ACF and Ljung–Box (Chapter 7.17).
Why do we need it?
It is the honest answer to "what does your model assume?". Knowing that the noise is assumed independent tells you what to check, where the model can lose accuracy, and why its intervals or posteriors may be too confident.
Where is it used?
Prophet and Prophet-style Bayesian models (your project), structural time-series models, regression with ARMA errors (SARIMAX with exogenous columns), and the residual-check step of every forecasting pipeline and model review.
How is it used?
Fit the model, compute residuals on the training period (and on rolling-origin holdouts), plot their ACF/PACF, run Ljung–Box. If lag 1 or 2 is clearly outside the band, try an AR(1) residual term and compare short-horizon accuracy and calibration on holdouts.
"My raw series has a huge ACF, so a Prophet-style model is the wrong tool; I need ARIMA."
Most of a raw series' ACF usually comes from trend and seasonality, which the structural model handles. The question is what is left: look at the ACF of the residuals.
"Switching from a Normal to a Student-t likelihood fixes autocorrelated residuals."
Student-t makes the model robust to occasional extreme days; Negative Binomial handles overdispersed counts. Neither links one day's noise to the next. Autocorrelation needs a memory term or a missing component.
"I will add yesterday's sales as a regressor; problem solved."
For a 14-day-ahead forecast, yesterday's sales are not known for days 2–14. Either use lags of at least the horizon, forecast recursively (feeding predictions back in, which compounds errors), or put the memory in the noise model. Otherwise the backtest leaks future information (Chapter 7.12).
Your forecasting model is exactly the left graph of the figure: $\mu_t = g(t) + s(t) + h(t) + X_t\beta$ with independent Normal, Student-t or Negative Binomial noise. Its changepoint trend, Fourier seasonality and holiday terms explain the autocorrelation that comes from structure; they do not explain short-term momentum. So the honest workflow is: fit (with your SVI loop), compute residuals (standardized ones for NB), plot their ACF/PACF with the $\pm 1.96/\sqrt n$ band and run Ljung–Box, on the training period and on rolling-origin holdouts. If lag-1 memory remains, the options are an AR(1) residual term written with lax.scan, horizon-safe lag features, or a richer component (often a missed holiday window or seasonality). Remember the second effect too: with independent noise, the posterior over $\delta_j$ and $\beta$ can be too narrow.
"Prophet handles autocorrelation."
"Prophet-style models are useless if the residuals are autocorrelated."
"It models the autocorrelation that comes from trend, seasonality, holidays and regressors. It assumes the remaining noise is independent, so it does not model short-term memory in the residuals. I check the residual ACF and Ljung–Box; if memory remains, short-horizon forecasts lose accuracy and the uncertainty is understated, and I add an AR residual term or horizon-safe lag features, or look for a missed component."
Model answer: "My model is a structural regression on time with independent noise. That is a deliberate trade-off: interpretable components and robust long-horizon behaviour, at the cost of ignoring day-to-day momentum. I verify the trade-off with residual diagnostics and fix it with an AR(1) error term when the residual ACF shows it matters."
Prophet-style: $y_t = \mu_t + \epsilon_t$, $\epsilon_t$ independent. Structure-driven autocorrelation is modelled; residual memory is not.
If $e_t = \varphi e_{t-1} + \eta_t$: worse short-horizon forecasts (correction $\varphi^h e_T$ fades), overconfident posteriors, understated variance of totals.
Fixes: AR(1) residual term (scan), horizon-safe lags, missing components. Traps: Student-t/NB do not fix memory; raw-series ACF ≠ residual ACF.
Quick check: residual $r_1 = 0.5$ and $e_T = -8$ at the forecast origin. What correction does an AR(1) residual term add at horizons 1, 2 and 3?
$0.5 \times (-8) = -4$, then $0.25 \times (-8) = -2$, then $0.125 \times (-8) = -1$. The forecast starts 4 orders below the calendar line and returns to it within a few days.
Recap, cheat sheet and practice
- Autocorrelation $r_k$ is the correlation of a series with itself $k$ steps back: one overall mean, divide by the full spread (statsmodels' convention; not
np.corrcoefof shifted pieces). - The ACF (correlogram) is the memory profile: fast decay = short memory, slow straight-line decay = trend or random walk, peaks at 7, 14, 21 = weekly pattern. AR(1): $\rho(k) = \varphi^k$.
- Under "no memory", $r_k \approx N(0, 1/n)$: the band ±1.96/√n. About 1 lag in 20 crosses by chance; Ljung–Box tests many lags at once.
- The PACF is the direct link after removing the in-between lags; $\alpha(2) = (\rho_2 - \rho_1^2)/(1 - \rho_1^2)$; Durbin–Levinson for all lags. AR($p$): PACF stops after $p$; MA($q$): ACF stops after $q$.
- White noise: constant mean and variance, zero autocorrelation; not necessarily independent or Normal. Good residuals look like it.
- Random walk = white noise added up: variance $t\sigma^2$, no level to return to, ACF that refuses to die, unit root; differencing gives white noise.
- Serial dependence inflates the true SE: AR(1) factor $(1+\varphi)/(1-\varphi)$, $n_{\text{eff}} = n/\text{factor}$; use HAC SEs or model the dependence.
- A Prophet-style model captures structure-driven autocorrelation but assumes independent residuals; check the residual ACF and add an AR term if memory remains.
Cheat sheet
| Idea | Formula / rule | Python |
|---|---|---|
| Sample ACF | $r_k = \sum_{t \gt k}(y_t-\bar y)(y_{t-k}-\bar y)\,/\,\sum_t (y_t-\bar y)^2$ | acf(y, nlags=30), plot_acf(y) |
| White-noise band | $\pm 1.96/\sqrt n$ (Bartlett band widens with lag) | plot_acf(y, alpha=0.05) |
| Ljung–Box | $Q = n(n+2)\sum_{k=1}^{h} r_k^2/(n-k) \sim \chi^2_h$ | acorr_ljungbox(y, lags=[10]) |
| PACF | $\alpha(1) = \rho_1$, $\alpha(2) = \frac{\rho_2-\rho_1^2}{1-\rho_1^2}$, Durbin–Levinson | pacf(y, nlags=30, method="ywm") |
| AR(1) | $\rho(k) = \varphi^k$; PACF stops after 1 | ArmaProcess(ar=[1, -phi]) |
| MA(1) | $\rho_1 = \theta/(1+\theta^2)$, then 0; PACF decays | ArmaProcess(ma=[1, theta]) |
| Random walk | $y_t = y_{t-1} + \epsilon_t$; $Var = t\sigma^2$; forecast sd $\sigma\sqrt h$ | np.cumsum(rng.normal(size=n)) |
| SE inflation | AR(1): $\times\sqrt{(1+\varphi)/(1-\varphi)}$; $n_{\text{eff}} = n/\text{factor}$ | .fit(cov_type="HAC", cov_kwds={"maxlags": L}) |
import numpy as np
import statsmodels.api as sm
from statsmodels.tsa.stattools import acf, pacf
from statsmodels.tsa.arima_process import ArmaProcess
from statsmodels.stats.diagnostic import acorr_ljungbox
# --- 1) lag-k autocorrelation by hand vs statsmodels (denominator n) ---
y = np.array([7, 8, 10, 12, 13, 11, 9, 10], dtype=float) # 8 days of orders
d = y - y.mean() # deviations from the mean 10
r1 = np.sum(d[1:] * d[:-1]) / np.sum(d**2) # 14 / 28
print(r1) # 0.5
print(acf(y, nlags=3).round(3)) # [ 1. 0.5 -0.179 -0.5 ]
print(round(np.corrcoef(y[1:], y[:-1])[0, 1], 3)) # 0.629 <- Pearson of the pairs: NOT the ACF
# --- 2) PACF: the method matters for short series ---
print(pacf(y, nlags=2, method="ywm").round(3)) # [ 1. 0.5 -0.571] (= plot_pacf default)
print(pacf(y, nlags=2).round(3)) # [ 1. 0.571 -0.838] (default "ywadjusted")
# --- 3) theory: AR(1) with phi = 0.8 and MA(1) with theta = 0.6 ---
ar1 = ArmaProcess(ar=[1, -0.8], ma=[1]) # note the sign: y_t - 0.8 y_{t-1} = e_t
print(ar1.acf(5).round(4)) # [1. 0.8 0.64 0.512 0.4096]
print(ar1.pacf(4).round(4)) # [1. 0.8 0. 0. ] (0.0 may print as -0.)
ma1 = ArmaProcess(ar=[1], ma=[1, 0.6])
print(ma1.acf(4).round(4)) # [1. 0.4412 0. 0. ]
# --- 4) simulate, then read the sample ACF/PACF with the +-1.96/sqrt(n) band ---
rng = np.random.default_rng(1)
x = ar1.generate_sample(nsample=500, distrvs=rng.standard_normal, burnin=200)
print(acf(x, nlags=3).round(2), pacf(x, nlags=3, method="ywm").round(2))
# [1. 0.81 0.66 0.55] [1. 0.81 0.03 0.03] ACF fades, PACF stops after lag 1
print(round(1.96 / np.sqrt(len(x)), 3)) # 0.088 white-noise band for n = 500
# --- 5) white noise and Ljung-Box ---
wn = rng.normal(size=200)
lb = acorr_ljungbox(wn, lags=[10])
print(lb.round(3)) # lb_stat 6.234, lb_pvalue 0.795: no memory found
print(acorr_ljungbox(x, lags=[10]).round(3)) # AR(1): lb_stat 1042.573, lb_pvalue 0.0
# --- 6) random walk: slow ACF decay; its differences are white noise ---
rw = np.cumsum(rng.normal(size=500))
print(acf(rw, nlags=5).round(3)) # [1. 0.986 0.972 0.959 0.946 0.932] fades very slowly
print(acf(np.diff(rw), nlags=3).round(2)) # [ 1. -0.05 0.02 -0. ] differences: white noise
# --- 7) serial dependence inflates the standard error of a mean ---
X = np.ones((len(x), 1))
naive = sm.OLS(x, X).fit() # assumes independent days
hac = sm.OLS(x, X).fit(cov_type="HAC", cov_kwds={"maxlags": 20}) # Newey-West
print(round(naive.bse[0], 3), round(hac.bse[0], 3)) # 0.075 0.216 (HAC is about 2.9x larger)
print(round(np.sqrt((1 + 0.8) / (1 - 0.8)), 2)) # 3.0 theory: true SE ~ 3x naive for phi = 0.8
# --- 8) a Prophet-style regression (trend + weekly Fourier terms) can leave AR(1) residuals ---
t = np.arange(280)
noise = ArmaProcess(ar=[1, -0.6]).generate_sample(280, scale=5, distrvs=rng.standard_normal, burnin=200)
y2 = 100 + 0.3 * t + 10 * np.sin(2 * np.pi * t / 7) + noise
Xd = np.column_stack([np.ones(280), t, np.sin(2 * np.pi * t / 7), np.cos(2 * np.pi * t / 7)])
resid = sm.OLS(y2, Xd).fit().resid
print(acf(y2, nlags=2).round(2)) # [1. 0.94 0.85] raw series: big ACF (mostly trend + season)
print(acf(resid, nlags=3).round(2)) # [1. 0.64 0.36 0.18] residual memory the model ignored
phi_hat = acf(resid, nlags=1)[1]
corrected = resid[1:] - phi_hat * resid[:-1] # one-step errors if we also use yesterday's residual
print(round(resid.std(), 2), round(corrected.std(), 2)) # 6.51 5.02 one-step errors shrink (theory 6.25 -> 5)
1. A series has $\sum_t (y_t - \bar y)^2 = 40$ and $\sum_{t} (y_t - \bar y)(y_{t-1} - \bar y) = 12$. What is $r_1$ (statsmodels convention)?
2. For an AR(1) with $\varphi = 0.6$, what is $\rho(3)$?
3. You have 400 days. Roughly how wide is the 95% white-noise band on the ACF plot?
4. The ACF fades geometrically (0.7, 0.5, 0.35, …) and the PACF has one spike at lag 1, then nothing. Best first guess?
5. Which statement about white noise is correct?
6. The residuals of your Prophet-style model have $r_1 = 0.5$ and a Ljung–Box p-value below 0.001. What is the main consequence?
Practice problems
A. For $y = 4, 6, 5, 7, 8$, compute $r_1$, $r_2$ and the PACF at lag 2.
Mean $= 30/5 = 6$; deviations $-2, 0, -1, 1, 2$; $\sum d^2 = 4 + 0 + 1 + 1 + 4 = 10$.
Lag 1: $(0)(-2) + (-1)(0) + (1)(-1) + (2)(1) = 0 + 0 - 1 + 2 = 1$, so $r_1 = 1/10 = 0.1$. Lag 2: $(-1)(-2) + (1)(0) + (2)(-1) = 2 + 0 - 2 = 0$, so $r_2 = 0$.
$\alpha(2) = (0 - 0.1^2)/(1 - 0.1^2) = -0.01/0.99 \approx -0.0101$. (statsmodels: acf → 0.1, 0; pacf(..., method="ywm") → 0.1, −0.0101.)
B. An MA(1) series $y_t = \epsilon_t + 0.5\,\epsilon_{t-1}$. Find $\rho_1$, $\rho_2$ and the PACF at lag 2.
$Cov(y_t, y_{t-1}) = 0.5\sigma^2$, $Var(y_t) = (1 + 0.25)\sigma^2 = 1.25\sigma^2$, so $\rho_1 = 0.5/1.25 = 0.4$. Two days apart they share no surprise: $\rho_2 = 0$.
$\alpha(2) = (0 - 0.16)/(1 - 0.16) = -0.16/0.84 \approx -0.190$. The ACF stops after lag 1 while the PACF keeps going (decaying, alternating): the MA(1) signature.
C. Daily inventory levels behave like a random walk with step sd 4 units and no drift. Today's level is 250. Give the forecast and 95% range for 30 days ahead.
Forecast: 250 (the last value). sd $= 4\sqrt{30} \approx 21.9$. 95% range: $250 \pm 1.96 \times 21.9 = 250 \pm 42.9$, about 207 to 293. The range keeps growing with the horizon, unlike a stationary series.
D. You have 200 days of residuals that behave like AR(1) with $\varphi = 0.7$. What is $n_{\text{eff}}$, and how much too small is a naive standard error of their mean?
Factor $(1 + 0.7)/(1 - 0.7) = 1.7/0.3 \approx 5.67$. $n_{\text{eff}} \approx 200/5.67 \approx 35$ days. The naive SE is too small by $\sqrt{5.67} \approx 2.38$: a naive 95% interval should be about 2.4 times wider.
E. Interview: "Your raw daily sales have $r_1 = 0.9$ and $r_7 = 0.85$. Does that mean you should use ARIMA instead of your Prophet-style model?"
"Not by itself. A trend makes all neighbouring values similar and a weekly pattern makes $r_7$ large; my model captures both through $g(t)$ and the weekly Fourier terms. What matters is the ACF of the residuals after the model. If the residuals look like white noise, the structural model has used the autocorrelation. If they still show lag-1 memory, I would add an AR(1) residual term or compare against SARIMA on rolling-origin backtests."
F. Interview: "What assumption does your forecasting model make about the noise over time, and how do you check it?"
"The noise is independent from day to day given the components: Normal or Student-t errors, or Negative Binomial counts given the mean. In other words the residuals should be white noise. I check the residual ACF and PACF against the $\pm 1.96/\sqrt n$ band, run Ljung–Box at lags 10 and 20, and look at the ACF of squared residuals for volatility clustering. If memory remains, I add an AR(1) error term (with scan in NumPyro) or the missing component, and I treat the uncertainty from the independent-noise fit as optimistic."
Stationarity and differencing
To learn from the past you need something that stays the same over time. "Stationary" is the precise name for that: the rules that generate the series do not drift. ARIMA insists on it and gets there by differencing; your Prophet-style model gets around it by modelling the drifting parts (trend, seasonality, holidays) directly. To compare the two, and to answer "is your series stationary?" in an interview, you need the idea clearly: what it means, how to see it, how to fix it, and how the ADF and KPSS tests check it.
- Say what stationarity means in plain words, and check it by eye with rolling mean and rolling spread
- Know strict vs weak stationarity (constant mean, constant variance, autocovariance depending only on the lag) and how they relate
- Sort common series into stationary or not, including the traps (slow AR(1), cycles, level shifts, weekly patterns)
- Use differencing (first, second, seasonal) and undo it; recognise over-differencing
- Tell trend-stationary from difference-stationary series, and choose between detrending and differencing
- Know what ADF and KPSS test (opposite null hypotheses), how to read them together, and their weaknesses
- Explain why a Prophet-style decomposition does not need a stationary series the way ARIMA does, and why you still need the idea
What stationarity means: the rules do not change over time core
Imagine cutting a 30-day window out of a long series, hiding the dates, and handing it to a friend. Can they tell whether it came from the beginning or the end? If the series is stationary, they cannot: every window looks like it was made by the same machine, with the same typical level, the same amount of wobble, and the same kind of memory. If the series trends upward, or gets noisier over time, they can tell at once.
"Stationary" (it means "not moving") does not mean "flat" or "constant". A stationary series wiggles all the time. What stays still are the rules behind the wiggles.
Three ways to say it:
- Picture: slide a window along the series; its average, its spread and its "stickiness" stay about the same everywhere.
- Numbers: days 1–5 average 10 with sd 1.6, and days 96–100 also average about 10 with sd about 1.6. If days 96–100 averaged 30, or had sd 9.5, the series would not be stationary.
- Slogan: stationary = the past is a fair sample of the future, because the machine making the data never changes.
Three windows of five days each, compared with a first window $A = 9, 11, 10, 12, 8$.
- Window A: mean $= 50/5 = 10$; deviations $-1, 1, 0, 2, -2$; squares sum $1 + 1 + 0 + 4 + 4 = 10$; $s = \sqrt{10/4} \approx 1.58$.
- Window B (later) $= 11, 9, 10, 8, 12$: mean 10, the same deviations in another order, so $s \approx 1.58$. Same level, same spread: consistent with stationarity.
- Window B′ $= 29, 31, 30, 32, 28$: mean 30, $s \approx 1.58$. The spread is the same but the level moved: mean not stable.
- Window C $= 4, 16, 10, 22, -2$: mean 10, deviations $-6, 6, 0, 12, -12$, squares sum $360$, $s = \sqrt{90} \approx 9.49$. Same level, but the spread grew six-fold: variance not stable.
- A third property is about memory: the correlation between days 1 and 2 should be the same as between days 96 and 97. Only the distance (lag) should matter, not the date.
A series is (weakly) stationary when three things do not depend on time $t$:
- Mean stability: $E[y_t] = \mu$ for every $t$ (no trend, no level shifts, no seasonal pattern in the mean).
- Variance stability: $Var(y_t) = \sigma^2 \lt \infty$ for every $t$ (no funnel, no calm-then-wild regimes in the variance).
- Autocovariance depends only on the lag: $Cov(y_t, y_{t+k}) = \gamma(k)$, the same at every $t$ (the memory works the same way throughout).
The precise strict and weak versions are on the next card. A rolling mean and rolling standard deviation (computed over a sliding window of $w$ days) are the quickest visual check: for a stationary series both stay roughly flat.
Why do we need it?
Estimating anything from the past (a mean, an autocorrelation, an AR coefficient) only makes sense if the past and the future follow the same rules. Stationarity is the formal version of "the history is representative".
Where is it used?
ARIMA and ARMA models (they model a stationary series, after differencing), the ACF/PACF (meaningful only for stationary series, Chapter 7.3), unit-root tests, residual checks of your forecasting model, and drift monitoring in production (Chapter 7.19).
How is it used?
Plot the series with y.rolling(30).mean() and y.rolling(30).std() in pandas. A drifting rolling mean → trend or level shifts; a drifting rolling sd → changing variance (try a log transform). Then confirm with ADF/KPSS (later in this chapter).
"Stationary means the series is flat and boring."
A stationary series can wiggle a lot and can have strong memory (an AR(1) with $\varphi = 0.9$ wanders for weeks). What must stay fixed are the rules: the long-run mean, the spread and the lag structure.
"The rolling mean moved a bit, so the series is not stationary."
Rolling statistics of a stationary series also wobble, especially with short windows and strong memory. Look for systematic drift (a steady climb, a jump that stays, a widening band), and press "New sample" to see how much wobble is normal.
Stationary: $E[y_t] = \mu$, $Var(y_t) = \sigma^2$, $Cov(y_t, y_{t+k}) = \gamma(k)$ — none depends on $t$.
Quick check: rolling mean and rolling sd (y.rolling(w).mean()/.std()) stay flat.
Trap: stationary ≠ flat; it means "the machine making the data does not change".
Quick check: daily orders double every year, and the day-to-day noise doubles with them. Which stationarity conditions fail? What simple transformation helps?
Both the mean (it grows) and the variance (it grows) fail. Taking logs turns "doubling" into "adding a constant" and turns noise that grows with the level into noise of constant size (Chapter 4.18); a trend in the log would then still need a difference or a trend term.
Strict vs weak stationarity
There are two levels of "the rules do not change". The strict version asks that everything about the series looks the same after a time shift: the whole shape of the distribution of any day, any pair of days, any week. The weak version only asks that the first two "summaries" stay fixed: the mean, the variance and the covariances.
Weak stationarity is what practitioners mean almost always, because the ACF, ARMA models and most tests only use means and covariances.
Three ways to say it:
- Picture: strict = the whole photograph is the same after shifting; weak = only the average brightness and contrast are the same.
- Numbers: odd days drawn from a Normal(0, 1) and even days from a flat (uniform) distribution with mean 0 and variance 1: means, variances and covariances never change (weakly stationary), but the shape alternates (not strictly stationary).
- Slogan: strict = same distribution after any shift; weak = same mean, variance and lag-covariances after any shift.
- Odd days: $y_t \sim N(0, 1)$: mean 0, variance 1.
- Even days: $y_t \sim \text{Uniform}(-\sqrt 3, \sqrt 3)$: mean 0, variance $(2\sqrt 3)^2/12 = 12/12 = 1$.
- All days independent, so $Cov(y_t, y_{t+k}) = 0$ for every $k \ne 0$, whatever $t$ is.
- Mean, variance and autocovariances do not depend on $t$: weakly stationary. But $P(y_t \gt 2)$ is about 0.023 on odd days and exactly 0 on even days ($\sqrt 3 \approx 1.73 \lt 2$): the distribution changes with $t$, so it is not strictly stationary.
- The other way round: independent Cauchy draws are strictly stationary (every day has the identical distribution) but have no finite mean or variance, so they are not weakly stationary.
Strict stationarity. For every number of time points $m$, every set of times $t_1, \dots, t_m$ and every shift $h$, the joint distribution of $(y_{t_1}, \dots, y_{t_m})$ equals that of $(y_{t_1+h}, \dots, y_{t_m+h})$.
Weak (covariance, second-order) stationarity. For all $t$ and $k$:
$$E[y_t] = \mu, \qquad Var(y_t) = \gamma(0) \lt \infty, \qquad Cov(y_t, y_{t+k}) = \gamma(k).$$- Strict + finite variance ⇒ weak.
- Weak ⇏ strict in general (example above). Weak + jointly Gaussian ⇒ strict, because a Gaussian process is fully described by its means and covariances.
- Strict ⇏ weak when moments are infinite (iid Cauchy).
- Examples of weakly stationary series: white noise, AR(1) with $|\varphi| \lt 1$, any MA($q$). Not stationary: random walks, trends, deterministic seasonal means, growing variance, level shifts.
Why do we need it?
"Stationary" can mean two different things, and interviewers sometimes ask which one. Knowing that ARMA, the ACF and unit-root tests only need the weak version tells you which properties you actually have to check.
Where is it used?
The theory of ARMA models (Wold's decomposition is about weakly stationary series), Gaussian processes (stationary kernels depend only on the distance $|t - s|$), spectral analysis, and the stationary distribution of a Markov chain in MCMC (Chapter 6.9, a different but related use of the word).
How is it used?
In practice you check the weak conditions: rolling mean, rolling variance, and whether the ACF looks the same in different parts of the series. Strict stationarity matters when you rely on the whole distribution (quantile forecasts, tail risk): then also compare histograms of early and late windows.
"Strictly stationary is just a stronger version of weakly stationary, so it always implies it."
Only when the variance is finite. iid Cauchy draws are strictly stationary but not weakly, because the mean and variance do not exist.
"A deterministic weekly pattern plus white noise is stationary because the pattern repeats."
The mean $E[y_t]$ is different on Saturdays and Mondays, so it depends on $t$: not stationary. It becomes stationary after you remove the pattern (regression on weekday terms) or take a seasonal difference.
Strict: every joint distribution is shift-invariant. Weak: constant mean, finite constant variance, $Cov(y_t, y_{t+k}) = \gamma(k)$.
Strict + finite variance ⇒ weak; weak + Gaussian ⇒ strict. Counter-examples: alternating Normal/Uniform (weak only), iid Cauchy (strict only).
In practice "stationary" = weakly stationary.
Quick check: $y_t = \epsilon_t$ for $t \le 100$ and $y_t = 2\epsilon_t$ for $t \gt 100$, with $\epsilon_t$ iid $N(0, 1)$. Stationary?
No. The mean is 0 throughout, but the variance jumps from 1 to 4 at $t = 100$, so it depends on $t$. Neither weakly nor strictly stationary.
Stationary or not? Train your eye
Most of the time you decide by looking. The obvious cases are easy: a climbing trend is not stationary, pure static is. The useful skill is the traps: series that look wild but are stationary (strong but temporary memory, irregular cycles, calm-and-turbulent spells), and series that look tame but are not (a single level shift, a weekly pattern in the mean).
The question to ask each time: is something pulling the series back to a fixed level with a fixed spread, or is the level or spread itself changing with time?
Three ways to say it:
- Picture: a dog on an elastic leash tied to a post (stationary: it roams but always comes back) versus a dog with no leash (random walk) or a post that is being moved (trend, level shift).
- Numbers: AR(1) with $\varphi = 0.95$ has memory half-life $\ln 0.5/\ln 0.95 \approx 13.5$ days: long swings, yet its variance is fixed at $\sigma^2/(1 - 0.95^2) \approx 10.3\,\sigma^2$.
- Slogan: wild is allowed; drifting rules are not.
Classify each with one reason.
- $y_t = 0.95\,y_{t-1} + \epsilon_t$: stationary ($|\varphi| \lt 1$: the pull-back exists, even if weak).
- $y_t = y_{t-1} + \epsilon_t$: not ($Var(y_t) = t\sigma^2$ grows).
- $y_t = 5 + 0.2t + \epsilon_t$: not ($E[y_t] = 5 + 0.2t$ depends on $t$).
- $y_t = \epsilon_t + 0.8\,\epsilon_{t-1}$: stationary (any MA is).
- $y_t = 10 + 4\cdot\mathbb 1[t \ge 125] + \epsilon_t$ (a level shift at day 125): not (the mean jumps).
- $y_t = s_{t \bmod 7} + \epsilon_t$ (a fixed weekly pattern): not in the strict textbook sense (the mean depends on the weekday).
Quick rules for the most common building blocks:
- White noise, any MA($q$), and AR($p$) with all roots outside the unit circle (AR(1): $|\varphi| \lt 1$; AR(2): inside the triangle of Chapter 7.3) are stationary.
- A unit root (random walk, AR(1) with $\varphi = 1$), a deterministic trend, a deterministic seasonal mean, a level shift or a trend break (changepoint), and a variance that changes with $t$ make a series non-stationary.
- Volatility clustering of the ARCH type (calm and turbulent spells arriving at random) is stationary: the conditional variance changes, the unconditional variance does not.
- Sums: stationary + stationary (jointly stationary) is stationary; stationary + trend is not; stationary + random walk is not.
Why do we need it?
Formal tests have low power and assumptions of their own; a trained eye catches the cause (a trend, a break, a funnel, a weekly mean) and therefore the fix, which a single p-value does not tell you.
Where is it used?
The first look at any new series in exploratory analysis, deciding between differencing, detrending, a log transform or explicit components, explaining residual plots in model reviews, and spotting regime changes in production monitoring.
How is it used?
Plot the series with its rolling mean and rolling sd, look at its ACF, and name the violation: drifting level → trend or unit root; jump → changepoint; funnel → transform or a variance model; weekday-dependent mean → seasonal terms or a seasonal difference.
"It has long up-and-down swings, so it cannot be stationary."
A strongly persistent AR(1) (φ = 0.95) or an AR(2) with pseudo-cycles swings for a long time and still returns to a fixed mean with a fixed spread. Look for a drift in the level or the spread, not for swings.
"Volatility clustering means non-stationary."
If the bursts arrive at random and the long-run variance is fixed (ARCH/GARCH), the series is stationary. A variance that grows with time, or jumps at a fixed date, is the non-stationary case.
Stationary: white noise, MA($q$), AR with $|\varphi| \lt 1$, ARCH-type clustering, differences of a random walk.
Not: random walk, trend, level shift / trend break, weekday-dependent mean, variance growing with $t$.
Ask: is there a fixed level and spread the series is pulled back to? Wild is allowed; drifting rules are not.
Quick check: is $y_t = \epsilon_t + 0.5\,t$ made stationary by subtracting $0.5\,t$? By differencing?
Subtracting the known trend leaves $\epsilon_t$: white noise, stationary. Differencing gives $0.5 + \epsilon_t - \epsilon_{t-1}$: also stationary, but now an MA(1) with $\theta = -1$ (lag-1 autocorrelation $-0.5$). Both work here; detrending is the more natural fix for a trend-stationary series (below).
Differencing: model the changes instead of the levels core
If you cannot predict where a wandering walker will be, you can still describe their steps. Differencing does exactly that: replace each value by its change from the day before. A random walk becomes its random steps (white noise). A straight upward trend becomes a constant step size.
This is the "I" (integrated) in ARIMA: difference the series until it is stationary, model the changes, then add the changes back up (integrate) to get forecasts of the levels.
Three ways to say it:
- Picture: instead of the height of a staircase, record the height of each step.
- Numbers: sales 100, 104, 107, 112, 115, 120 become changes 4, 3, 5, 3, 5: the trend has become an average step of 4.
- Slogan: differencing turns "where am I?" into "how much did I move?".
- Weekly sales $y = 100, 104, 107, 112, 115, 120$.
- First difference $\nabla y_t = y_t - y_{t-1}$: $104 - 100 = 4$, $107 - 104 = 3$, $112 - 107 = 5$, $115 - 112 = 3$, $120 - 115 = 5$. One value is lost (the first week has no previous week).
- Average change: $(4 + 3 + 5 + 3 + 5)/5 = 20/5 = 4 = (120 - 100)/5$: the slope of the trend (the drift).
- Undo it (integrate): start at 100 and add the changes: $100, 104, 107, 112, 115, 120$. A forecast of the next change, say 4, gives the next level $120 + 4 = 124$.
- A curved (quadratic) trend needs two differences: $y = 0, 1, 4, 9, 16$ gives $\nabla y = 1, 3, 5, 7$ and $\nabla^2 y = 2, 2, 2$.
- Over-differencing: differencing white noise $\epsilon_t$ gives $\epsilon_t - \epsilon_{t-1}$, with $Var = 2\sigma^2$ and $Cov(\nabla\epsilon_t, \nabla\epsilon_{t-1}) = -\sigma^2$, so $\rho_1 = -\sigma^2/2\sigma^2 = -0.5$. A strong negative lag-1 autocorrelation after differencing means you differenced once too often.
The first difference is $\nabla y_t = y_t - y_{t-1}$ (also written $(1 - B)y_t$ with the backshift operator $B y_t = y_{t-1}$). The $d$-th difference applies it $d$ times: $\nabla^2 y_t = y_t - 2y_{t-1} + y_{t-2}$.
- A series is integrated of order $d$, I($d$), if it needs $d$ differences to become stationary. Random walk: I(1). Stationary series: I(0).
- A linear trend $a + bt$ becomes the constant $b$; a quadratic trend needs $d = 2$.
- Each difference loses one observation and adds noise (variance of $\nabla\epsilon$ is $2\sigma^2$).
- Over-differencing a stationary series creates negative lag-1 autocorrelation: an AR(1) with carry-over $\varphi$ becomes a series with $\rho_1 = -(1 - \varphi)/2$ (−0.5 for white noise).
- Undo by cumulative summation: $y_t = y_0 + \sum_{s=1}^{t}\nabla y_s$. Forecasts of the changes become forecasts of the levels this way, and their uncertainties add up, so the intervals widen.
Why do we need it?
Methods like ARMA need a stationary series. Differencing is the simplest way to remove a unit root or a trend without having to specify what the trend looks like.
Where is it used?
The $d$ in ARIMA($p, d, q$) (Chapter 7.6), returns in finance (differences of log prices), the naive and drift forecasts (Chapter 7.5), unit-root testing, and correlating the changes of two series instead of their levels to avoid spurious correlation.
How is it used?
y.diff() in pandas or np.diff(y). Difference once, check the plot, ACF and an ADF/KPSS test; difference again only if still needed (rarely more than $d = 2$). If $r_1$ of the differenced series is near −0.5, you over-differenced.
"Differencing is harmless, so difference to be safe."
Differencing a stationary series injects negative autocorrelation (an MA part with a unit root), inflates the noise variance and makes long-run forecasts drift-like with ever-growing intervals. Difference only when the series needs it.
"After differencing, my forecast is the forecast of the changes."
You must add the forecast changes back onto the last observed level (cumulative sum). Libraries like statsmodels' ARIMA(order=(p, 1, q)) do this for you; if you difference by hand, you undo it by hand.
$\nabla y_t = y_t - y_{t-1}$; $\nabla^2 y_t = y_t - 2y_{t-1} + y_{t-2}$. I($d$): $d$ differences to stationarity.
Linear trend → constant (the drift); random walk → white noise; undo with a cumulative sum from the last level.
Over-differencing: $r_1 \approx -0.5$ (white noise) or $-(1-\varphi)/2$ (AR(1)). Rarely need $d \gt 2$.
Quick check: $y = 3, 5, 9, 15, 23$. Difference until the result is constant. What is $d$?
$\nabla y = 2, 4, 6, 8$; $\nabla^2 y = 2, 2, 2$. Constant after two differences, so $d = 2$ (the levels follow a quadratic, $y_t = t^2 + t + 3$ for $t = 0, \dots, 4$).
Seasonal differencing: compare with the same day last week
Saturdays are always busy and Mondays always quiet. Comparing a Saturday with the Friday before mixes "the weekly pattern" with "real change". Comparing this Saturday with last Saturday cancels the weekly pattern and leaves only the change. That is a seasonal difference: subtract the value one full season ago.
Three ways to say it:
- Picture: lay this week on top of last week and look only at the gaps.
- Numbers: week 1: 20, 22, 21, 23, 30, 45, 40; week 2: 22, 24, 23, 25, 32, 47, 42. Seasonal differences: 2, 2, 2, 2, 2, 2, 2. The weekly shape is gone; only the growth of 2 per week is left.
- Slogan: to remove a season, subtract the same point one season ago.
- Seasonal difference with period $m = 7$: $\nabla_7 y_t = y_t - y_{t-7}$.
- Monday: $22 - 20 = 2$; Tuesday: $24 - 22 = 2$; … Saturday: $47 - 45 = 2$; Sunday: $42 - 40 = 2$.
- The weekly pattern (from 20 on Mondays up to 45 on Saturdays) cancels completely because it is identical in both weeks.
- We lose the first 7 values (one full season), compared with 1 value for an ordinary difference.
- If a trend is also present, $\nabla_7$ leaves a constant (here 2 per week); if the series also wanders, combine both: $\nabla\nabla_7 y_t = (y_t - y_{t-7}) - (y_{t-1} - y_{t-8})$.
For a season of length $m$ (7 for daily data with a weekly pattern, 12 for monthly data with a yearly pattern), the seasonal difference is
$$\nabla_m y_t = y_t - y_{t-m} = (1 - B^m)\,y_t.$$- It removes any seasonal pattern that repeats exactly, and also a seasonal pattern that evolves slowly (a "seasonal random walk").
- It is the $D$ in SARIMA$(p, d, q)(P, D, Q)_m$ (Chapter 7.6). Usually $D \le 1$, and $d + D \le 2$.
- Over-seasonal-differencing: if the pattern is perfectly fixed, $\nabla_m$ applied to the noise creates a negative autocorrelation of about $-0.5$ at lag $m$ (the same effect as over-differencing, one season apart). Then explicit seasonal terms (weekday dummies or Fourier terms) are the better tool.
Why do we need it?
A seasonal pattern makes the mean depend on the time of year or week, which breaks stationarity. Seasonal differencing removes it without having to estimate the pattern.
Where is it used?
SARIMA models, "same day last week" and "same month last year" comparisons in business reporting, year-over-year growth rates, and the seasonal naive forecast (Chapter 7.5).
How is it used?
y.diff(7) for daily data with a weekly cycle (y.diff(12) for monthly data). Check the ACF: the spikes at 7, 14, 21 should disappear. A strong negative spike at lag 7 afterwards suggests the pattern was fixed and Fourier or dummy terms would be cleaner.
"An ordinary first difference removes seasonality too."
$\nabla$ removes trends and unit roots, not weekly patterns: Saturday minus Friday is still large every week. Use $\nabla_7$ (or seasonal terms) for the season.
"Seasonal differencing and Fourier terms are interchangeable."
Seasonal differencing assumes the pattern may evolve and throws away one season of data; Fourier or dummy terms estimate a fixed (or slowly varying, if you let them) pattern explicitly and keep the levels interpretable. Your model uses the second approach (Chapter 7.11).
$\nabla_m y_t = y_t - y_{t-m}$ ($m = 7$ daily/weekly, $m = 12$ monthly/yearly). Loses $m$ values.
Removes repeating patterns; combine with $\nabla$ for trend + season ($d + D \le 2$).
Trap: a fixed pattern + $\nabla_m$ → negative spike at lag $m$; then prefer Fourier/dummy terms.
Quick check: monthly sales with a yearly pattern and a steady growth of 5 units per month. What does $\nabla_{12}$ leave?
The yearly pattern cancels, and the growth over 12 months adds up: $\nabla_{12} y_t \approx 12 \times 5 = 60$ plus noise. A constant: the year-over-year change.
Detrend or difference? Trend-stationary vs difference-stationary series core
Two series can both climb steadily and look almost identical, yet react to a shock in opposite ways. Suppose a big promotion adds 10 extra orders one day.
- In a trend-stationary world the series lives on an invisible line (the trend) and wobbles around it. After the shock it drifts back to the line within days. The shock is temporary.
- In a difference-stationary world (a random walk with drift) there is no line to return to. After the shock the whole future path is 10 higher. The shock is permanent.
The right fix follows the world: subtract the line (detrend) in the first, take differences in the second.
Three ways to say it:
- Picture: a dog on a leash walking beside a moving owner (trend-stationary) versus a dog with no leash that keeps a steady pace on average (difference-stationary).
- Numbers: with leash strength $\varphi = 0.6$, a shock of 10 is $10 \times 0.6^5 \approx 0.78$ after 5 days; in the random walk it is still 10 after 500 days.
- Slogan: temporary shocks → detrend; permanent shocks → difference.
- Trend-stationary: $y_t = 10 + 0.2t + u_t$ with $u_t = 0.6\,u_{t-1} + \epsilon_t$. A shock of $+10$ in $u$ at day 50 adds $10 \times 0.6^h$ to day $50 + h$: $6$ after 1 day, $3.6$ after 2, $0.78$ after 5, $0.0004$ after 20.
- Difference-stationary: $y_t = 0.2 + y_{t-1} + \epsilon_t$. The same shock adds $10$ to day 50 and, since each day starts from the previous one, $10$ to every later day.
- Detrending the first (subtract a fitted line $\hat a + \hat b t$) leaves the stationary AR(1) $u_t$.
- Differencing the second leaves $0.2 + \epsilon_t$: white noise around the drift.
- The wrong fixes: differencing the first gives $0.2 + u_t - u_{t-1}$, stationary but over-differenced (negative $r_1$). Detrending the second leaves a series that still wanders, and the fitted slope looks far more certain than it is (a spurious trend).
- Forecasts: trend-stationary → back toward the line, with an interval whose width stops growing (at most $\pm 1.96\,\sigma/\sqrt{1 - \varphi^2}$ around the line); difference-stationary → last value plus drift, with an interval $\pm 1.96\,\sigma\sqrt h$ that grows forever.
A series is trend-stationary if $y_t = f(t) + u_t$ where $f$ is a deterministic function of time (a line, a curve, a piecewise line) and $u_t$ is stationary. Removing $f$ (detrending, usually by regression on $t$) makes it stationary. Shocks fade.
A series is difference-stationary (has a unit root, is I(1)) if $\nabla y_t$ is stationary but $y_t$ is not: $y_t = \mu + y_{t-1} + u_t$. Differencing makes it stationary. Shocks persist forever.
- Detrending a difference-stationary series leaves a non-stationary remainder, and a regression of $y_t$ on $t$ reports a "significant" slope even for a pure random walk with zero drift (spurious regression).
- Differencing a trend-stationary series works but over-differences: $\nabla u_t$ has $\rho_1 = -(1 - \varphi)/2$.
- Forecast error variance at horizon $h$: bounded for trend-stationary (it tends to $Var(u_t)$, plus parameter uncertainty); $h\sigma^2$ for a random walk.
- Unit-root tests (next section) are designed to tell the two apart, with limited power.
Why do we need it?
The two worlds need different fixes and give very different long-range forecasts and interval widths. Picking the wrong one either throws away information (over-differencing) or invents certainty (spurious trends).
Where is it used?
Choosing between ARIMA with $d = 1$ and a regression on time with stationary errors; the long-standing economics debate about whether GDP has a unit root; and comparing a Prophet-style deterministic trend with an ARIMA forecast (last section of this chapter).
How is it used?
Ask what a shock does: does the series return to a line? Look at plots of residuals after detrending and of differences, check their ACFs (wandering vs $r_1 \approx -0.5$), and run ADF with a trend (regression="ct") plus KPSS.
"A straight line fits my series with t = 25, so the trend is real."
If the series has a unit root, the residuals are not independent and the regression's t-statistic is meaningless: pure random walks routinely give |t| above 10. Check whether the residuals are stationary before trusting a trend.
"Detrending and differencing are two equivalent ways to remove a trend."
They encode different beliefs about shocks (temporary vs permanent) and give different forecasts: back to the line with bounded intervals, versus from the last value with intervals growing like $\sqrt h$.
Trend-stationary: $y_t = f(t) + u_t$, $u_t$ stationary → detrend; shocks fade ($S\varphi^h$); bounded forecast variance.
Difference-stationary (unit root): $y_t = \mu + y_{t-1} + u_t$ → difference; shocks are permanent; forecast variance $h\sigma^2$.
Wrong fixes: detrending a random walk → spurious trend; differencing a trend-stationary series → over-differencing.
Quick check: after an outage, a series drops by 30 and is still 30 lower three months later, with no sign of returning. Which world does this suggest?
Difference-stationary-like behaviour (a permanent shock), or a level shift (a changepoint) in an otherwise trend-stationary series. The two are hard to tell apart from one event; your model would describe it as a changepoint in $g(t)$.
Unit-root tests: ADF and KPSS ask opposite questions core
Two detectives investigate the same series with opposite starting assumptions.
- The ADF detective assumes "it is a random walk" (non-stationary) and looks for evidence of a pull back toward a mean. Only strong evidence makes it say "stationary".
- The KPSS detective assumes "it is stationary" and looks for evidence of wandering. Only strong evidence makes it say "not stationary".
A hypothesis test only ever rejects its starting assumption (the null hypothesis) or fails to; it never proves it (Chapter 5.6). Using both detectives covers both directions.
Three ways to say it:
- Picture: ADF looks at the scatter of "today's change" against "yesterday's level": a downward slope means the series is pulled back. KPSS adds up the deviations from the mean: a running total that drifts far away means the series wanders.
- Numbers: AR(1) with $\varphi = 0.5$: ADF statistic about −9 (far below −2.87: reject the unit root), KPSS about 0.05 (below 0.463: keep stationarity). Random walk: ADF about −1.5 (keep the unit root), KPSS about 2.5 (reject stationarity).
- Slogan: ADF: null = unit root; KPSS: null = stationary. Agreement is evidence; disagreement is a warning.
Real output from the Code-it block at the end (300 days each; statsmodels; ADF lags chosen by AIC):
| series | ADF stat (p) | KPSS stat (p) | reading |
|---|---|---|---|
| AR(1), $\varphi = 0.5$ | −8.96 (< 0.001) | 0.051 (≥ 0.10) | both say stationary |
| random walk | −1.54 (0.51) | 2.456 (≤ 0.01) | both say unit root → difference |
| differences of the walk | −13.36 (< 0.001) | 0.152 (≥ 0.10) | stationary after one difference: I(1) |
trend + AR(1), tested with "c" | −1.20 (0.67) | 2.743 (≤ 0.01) | looks like a unit root (wrong question!) |
trend + AR(1), tested with "ct" | −8.94 (< 0.001) | 0.051 (≥ 0.10) | trend-stationary → detrend |
- Random walk, ADF: $-1.54 \gt -2.87$ (the 5% critical value for $n = 300$ with a constant), so we cannot reject "unit root".
- Random walk, KPSS: $2.456 \gt 0.463$ (5% critical value, level version), so we reject "stationary".
- Both point the same way: a unit root. Difference once and test again: the differences pass both tests.
- The trend series shows why the regression option matters: asked "stationary around a constant?" (
"c"), both tests say no; asked "stationary around a line?" ("ct"), both say yes.
ADF (augmented Dickey–Fuller). Regress today's change on yesterday's level:
$$\Delta y_t = \alpha + \beta t + \gamma\,y_{t-1} + \sum_{i=1}^{p}\delta_i\,\Delta y_{t-i} + e_t$$($\beta t$ only with regression="ct"; the lagged changes absorb short-term autocorrelation, and statsmodels chooses $p$ by AIC). For an AR(1), $\gamma = \varphi - 1$.
- $H_0$: $\gamma = 0$, a unit root (non-stationary). $H_1$: $\gamma \lt 0$ (stationary around a constant, or around a trend for
"ct"). - Statistic: $\tau = \hat\gamma/SE(\hat\gamma)$. Under $H_0$ it does not follow a t-distribution; it follows the Dickey–Fuller distribution, shifted to the left. 5% critical values: about $-2.86$ (constant; $-2.89$ for $n = 100$) and $-3.41$ (constant + trend), versus $-1.645$ for an ordinary one-sided test.
- Reject (small p) → evidence of stationarity. Fail to reject → "a unit root cannot be ruled out", not "proved non-stationary".
KPSS (Kwiatkowski–Phillips–Schmidt–Shin). Picture the series as stationary noise plus a random-walk part with variance $\sigma_u^2$.
- $H_0$: $\sigma_u^2 = 0$, the series is stationary around a level (
"c") or a trend ("ct"). $H_1$: a random-walk part exists. - Statistic: with residuals $e_t$ from regressing $y$ on a constant (or constant + trend) and their running sums $S_t = e_1 + \dots + e_t$, $$\eta = \frac{\frac{1}{n^2}\sum_{t=1}^{n} S_t^2}{\hat\sigma^2_{LR}},$$ where $\hat\sigma^2_{LR}$ is a long-run variance (it accounts for autocorrelation). A stationary series keeps $S_t$ near 0; a wandering one lets it drift far.
- 5% critical values: 0.463 (level), 0.146 (trend). Large $\eta$ → reject stationarity. statsmodels looks the p-value up in a small table and clips it to [0.01, 0.10] (with an
InterpolationWarning).
Reading them together: ADF rejects + KPSS does not → stationary. ADF does not + KPSS rejects → unit root, difference. Both reject → conflicting (often a trend tested with "c", a structural break, or a near-unit root). Neither rejects → inconclusive (too little data).
Why do we need it?
Eyeballing can be fooled by persistent but stationary series and by slow wandering. A formal test, read in both directions, gives a documented reason for choosing $d$ in ARIMA or for detrending instead of differencing.
Where is it used?
The $d$ step of ARIMA (pmdarima.auto_arima uses KPSS by default to choose $d$), econometrics (unit roots in prices and GDP), cointegration analysis, and checking that model residuals are stationary.
How is it used?
adfuller(y, regression="c", autolag="AIC") and kpss(y, regression="c", nlags="auto") from statsmodels.tsa.stattools; use "ct" when the series trends. Read both, difference if they point to a unit root, test again, and always plot.
"ADF p = 0.3, so the series has a unit root."
Failing to reject is not proof. With 100 days and $\varphi = 0.95$, a stationary series passes as "unit root" most of the time (in our simulation of 1 000 such series, ADF rejected only about 15%). Say "a unit root cannot be ruled out".
"KPSS p = 0.10, so the series is stationary with probability 0.9."
A p-value is not the probability of a hypothesis, and statsmodels clips KPSS p-values to the range 0.01–0.10, so 0.10 means "0.10 or more".
"The tests say non-stationary, so the series must have a unit root."
Structural breaks fool both tests. In a simulation of 300 series (300 days, AR(1) noise with $\varphi = 0.5$) around a trend whose slope flips from +0.1 to −0.1 halfway, ADF with a trend never rejected and KPSS rejected every time, although every series was stationary around its bent line. A changepoint can look exactly like a unit root.
Your forecasting model detects changepoints (grid + PELT) precisely because demand trends bend. That same bending makes ADF and KPSS on the raw series say "unit root" quite often, so a failed test is not a reason to switch to differencing or to doubt the trend component. The useful places for these tests in your workflow are the residuals of the fitted model (they should pass as stationary) and an ARIMA baseline you compare against, where the tests help choose $d$.
ADF: $\Delta y_t = \alpha (+\beta t) + \gamma y_{t-1} + \sum\delta_i\Delta y_{t-i} + e_t$; $H_0$: $\gamma = 0$ (unit root). Reject → stationary. 5% crit ≈ −2.86 ("c"), −3.41 ("ct").
KPSS: $H_0$: stationary (level or trend); $\eta = \sum S_t^2/(n^2\hat\sigma^2_{LR})$; crit 0.463 ("c"), 0.146 ("ct"). Reject → not stationary.
Use both; choose "c" vs "ct" deliberately. Traps: low power near φ = 1; breaks/changepoints mimic unit roots; fail to reject ≠ accept.
Quick check: ADF p = 0.02 and KPSS p = 0.01 (both "c") on a series that clearly trends upward. What do you do?
Conflict: ADF says "no unit root", KPSS says "not stationary around a constant". The obvious cause is the trend: rerun both with regression="ct". If both then agree on trend-stationarity, detrend (or model the trend explicitly, as your model does); if not, look for breaks.
Stationarity, ARIMA and your Prophet-style model core
There are two strategies for a series whose level and pattern change over time.
- ARIMA's strategy: transform the problem away. Difference (and seasonally difference) until what is left is stationary, model that stationary remainder with AR and MA terms, then add the changes back up.
- Your model's strategy: describe the changes directly. Write the non-stationary parts as explicit functions of time (a piecewise-linear trend with changepoints, Fourier seasonality, holiday effects, regressors) and assume that only the leftover noise is stationary.
Think of a passenger on a moving train. ARIMA studies how the passenger moves relative to their last position. A Prophet-style model draws the train's route on a map and describes the passenger's small wobble around it.
Three ways to say it:
- Picture: ARIMA flattens the landscape first; a Prophet-style model draws the landscape.
- Numbers: for a random walk with step sd 5, a 30-day forecast interval is $\pm 1.96 \times 5\sqrt{30} \approx \pm 54$; for a fixed trend line with noise sd 5 it is about $\pm 9.8$ (before parameter uncertainty). Same data, very different promises.
- Slogan: ARIMA needs a stationary series; a decomposition needs stationary residuals.
Daily demand with a slope of about 0.5 orders per day, a weekly pattern, and noise. History ends at day $T$.
- ARIMA route: ADF/KPSS say "unit root" → take $\nabla$ (and $\nabla_7$ for the week). Fit an ARMA to the changes, forecast the changes, add them to $y_T$. Each forecast change adds its own uncertainty, so with an I(1) model the interval keeps widening (like $\sqrt h$ for a pure random walk).
- Prophet-style route: no differencing. Fit $g(t)$ (slope 0.5 with changepoints), $s(t)$ (weekly Fourier terms) and $h(t)$ on the levels. Forecast by extending each component. If the noise is iid with sd $\sigma$, the noise part of the interval has constant width $\pm 1.96\,\sigma$; the widening must come from trend uncertainty.
- Prophet itself, according to its documentation, adds trend uncertainty by assuming the future has changepoints with the same average frequency and size as the history and simulating them. Whether your own model does something similar decides how its long-horizon intervals behave: check it.
- Stationarity still appears in the Prophet-style route: the residuals $e_t = y_t - \hat\mu_t$ must be stationary. A residual mean that drifts means a missed changepoint or level shift; a residual spread that grows means a variance problem.
ARIMA($p, d, q$) assumes that $\nabla^d y_t$ is a stationary ARMA($p, q$) process. Its coefficients are only meaningful (and its forecasts only revert sensibly) because the differenced series has a constant mean, variance and autocovariance. Non-stationarity is removed before modelling.
A Prophet-style decomposition assumes
$$y_t = \underbrace{g(t) + s(t) + h(t) + X_t\beta}_{\text{non-stationary parts, modelled explicitly}} + \underbrace{\epsilon_t}_{\text{iid, so stationary}}.$$Non-stationarity is modelled, not removed: no differencing, forecasts in levels, interpretable components. It is a trend-stationary view of the world (shocks to $\epsilon_t$ are temporary), made flexible by changepoints that let the trend bend.
Where you still need stationarity with your model:
- Residual checks: rolling mean and rolling sd of residuals should be flat, and their ACF should look like white noise (Chapter 7.3, Chapter 7.17).
- Comparing approaches: "unit root or trend-stationary?" is exactly "should shocks persist (ARIMA-like intervals growing like $\sqrt h$) or fade (decomposition with bounded noise intervals)?". Rolling-origin backtests and interval coverage at long horizons decide in practice (Chapter 7.15, 7.16).
- Regressors: a trending regressor and a trending target can look strongly related by accident (spurious regression), and a trending regressor can steal credit from $g(t)$ (identifiability, Chapter 6.8).
- Unit-root tests on your data: changepoints make a trend-stationary series look like a unit root to ADF/KPSS (previous section), so "the tests say non-stationary" is not a reason to difference before a changepoint model.
Why do we need it?
"Why didn't you make the series stationary?" and "Why not ARIMA?" are standard interview questions about a Prophet-style model. The honest answer needs the distinction between removing non-stationarity and modelling it, and the residual checks that remain.
Where is it used?
Prophet, NeuralProphet and structural time-series models (decomposition route); ARIMA, SARIMA and auto_arima (differencing route); hybrid models such as regression with ARIMA errors (SARIMAX with exogenous columns), which combine both.
How is it used?
For your model: fit on levels, then test the residuals (rolling mean/sd, ACF, Ljung–Box, ADF/KPSS on the residuals). To compare with ARIMA: run both in the same rolling-origin backtest and compare point accuracy and interval coverage by horizon.
"Prophet-style models require stationary data, so I differenced the series before fitting."
Differencing removes exactly the trend and level information the trend component is supposed to model, and the forecasts would then be of changes, not levels. Fit the decomposition on the levels (a log transform for multiplicative growth is a different matter).
"My model has no stationarity assumption at all."
It assumes the residual noise is iid, which includes being stationary: constant mean (zero), constant variance (given the likelihood) and no memory. And its deterministic trend encodes a trend-stationary view of shocks; changepoints soften this, but it is still an assumption to check against the data.
In your forecasting model the piecewise-linear trend with changepoints (grid + PELT) is what replaces differencing: each slope change $\delta_j$ lets the deterministic trend follow a level that drifts, so much of what an ADF test would call a "unit root" is absorbed as changepoints. Three practical consequences. (1) Do not difference before fitting; check the residuals instead: rolling mean and sd of residuals over time, their ACF, and ADF/KPSS on the residuals. (2) With the Negative Binomial likelihood, a variance that grows with the level is part of the model ($Var = \mu + \mu^2/\alpha$, with $\alpha$ the NB concentration; check which parameterization your code uses), so judge variance stability on standardized residuals $(y_t - \mu_t)/\sqrt{\mu_t + \mu_t^2/\alpha}$, not on raw ones. (3) Long-horizon intervals depend on how your model treats future trend changes; if it keeps the last slope fixed and adds only iid noise, its fan has the constant-width shape of the orange band above, and backtested coverage at long horizons is the test of whether that is acceptable.
"Prophet needs the series to be stationary."
"Stationarity doesn't matter for Prophet-style models."
"ARIMA needs a stationary (differenced) series because it models a stationary process. A Prophet-style decomposition models the non-stationary parts (trend with changepoints, seasonality, holidays, regressors) explicitly as functions of time, so it does not require making the series stationary. It does assume the residuals are stationary iid noise, and its deterministic trend implies shocks are temporary. I use stationarity to check the residuals and to compare the two approaches, especially how their uncertainty grows with the horizon."
Model answer: "My model doesn't difference; it puts the non-stationarity into $g(t)$, $s(t)$ and $h(t)$. So the stationarity question moves to the residuals, which I check with rolling statistics, the ACF and Ljung–Box. Compared with ARIMA, my model assumes trend-stationarity with occasional slope changes, while ARIMA with $d = 1$ assumes every shock is permanent. Which is right shows up in rolling-origin backtests, especially in long-horizon coverage."
ARIMA: make $\nabla^d y_t$ stationary, model it, integrate back (removes non-stationarity). Prophet-style: $y_t = g + s + h + X\beta + \epsilon_t$ (models non-stationarity); requires stationary iid residuals.
Trend-stationary worldview → bounded noise bands; unit-root worldview → bands ∝ $\sqrt h$.
Traps: differencing before a decomposition; "no stationarity assumption"; unit-root tests fooled by changepoints.
Quick check: the residuals of your fitted model have a rolling mean that drops by about 20 after a certain date and stays there. What does this tell you?
The residuals are not stationary: a level shift happened that the trend did not capture (a missed changepoint, an unmodelled event, or a change in a regressor). Fix the mean model (a changepoint near that date, a regressor, or an event indicator) rather than differencing.
Recap, cheat sheet and practice
- Stationary = the rules do not change over time: constant mean, constant finite variance, and autocovariance depending only on the lag (weak stationarity). Strict stationarity asks for every joint distribution to be shift-invariant.
- Check by eye with rolling mean and rolling sd; wild swings are allowed, drifting rules are not. Traps: persistent AR(1), pseudo-cycles and ARCH bursts are stationary; level shifts and fixed weekly means are not.
- Differencing $\nabla y_t = y_t - y_{t-1}$ removes unit roots and linear trends; seasonal differencing $\nabla_m$ removes repeating patterns; undo with a cumulative sum. Over-differencing gives $r_1 \approx -0.5$.
- Trend-stationary (shocks fade, detrend, bounded intervals) vs difference-stationary (shocks persist, difference, intervals ∝ $\sqrt h$). Detrending a random walk gives spurious trends.
- ADF: $H_0$ unit root; KPSS: $H_0$ stationary. Use both, choose "c" or "ct" deliberately, remember low power near $\varphi = 1$ and that changepoints mimic unit roots.
- ARIMA removes non-stationarity; a Prophet-style model describes it with explicit components and needs stationary residuals. Stationarity is still how you check the residuals and compare the two.
Cheat sheet
| Idea | Formula / rule | Python |
|---|---|---|
| Weak stationarity | $E[y_t] = \mu$, $Var(y_t) = \gamma(0) \lt \infty$, $Cov(y_t, y_{t+k}) = \gamma(k)$ | y.rolling(30).mean(), .std() |
| First / second difference | $\nabla y_t = y_t - y_{t-1}$; $\nabla^2 y_t = y_t - 2y_{t-1} + y_{t-2}$ | y.diff(), np.diff(y, n=2) |
| Seasonal difference | $\nabla_m y_t = y_t - y_{t-m}$ | y.diff(7) |
| Over-differencing | white noise → $\rho_1 = -0.5$; AR(1) → $-(1-\varphi)/2$ | acf(np.diff(y)) |
| Trend- vs difference-stationary | $f(t) + u_t$ (detrend) vs $\mu + y_{t-1} + u_t$ (difference) | np.polyfit residuals vs np.diff |
| ADF | $H_0$: $\gamma = 0$ in $\Delta y_t = \alpha + \gamma y_{t-1} + \dots$; 5% ≈ −2.86 ("c"), −3.41 ("ct") | adfuller(y, regression="c", autolag="AIC") |
| KPSS | $H_0$: stationary; $\eta = \sum S_t^2/(n^2\hat\sigma^2_{LR})$; 5%: 0.463 ("c"), 0.146 ("ct") | kpss(y, regression="c", nlags="auto") |
import warnings
import numpy as np
import pandas as pd
from statsmodels.tsa.stattools import adfuller, kpss, acf
from statsmodels.tsa.arima_process import ArmaProcess
warnings.simplefilter("ignore") # after the imports: hides statsmodels' FutureWarning / InterpolationWarning
rng = np.random.default_rng(42)
n = 300
t = np.arange(n)
ar = ArmaProcess(ar=[1, -0.5]).generate_sample(n, distrvs=rng.standard_normal, burnin=200) # stationary AR(1)
rw = np.cumsum(rng.normal(size=n)) # random walk (unit root)
ts = 10 + 0.05 * t + ar # trend-stationary: line + stationary noise
# KPSS p-values are looked up in a table and clipped to [0.01, 0.10]: 0.100 means ">= 0.10"
def tests(name, x, reg="c"):
adf_stat, adf_p = adfuller(x, regression=reg, autolag="AIC")[:2] # H0: unit root
kpss_stat, kpss_p = kpss(x, regression=reg, nlags="auto")[:2] # H0: stationary
print(f"{name:22s} ADF {adf_stat:6.2f} (p={adf_p:.3f}) KPSS {kpss_stat:5.3f} (p={kpss_p:.3f})")
tests("AR(1), phi=0.5", ar) # ADF -8.96 (p=0.000) KPSS 0.051 (p=0.100): both say stationary
tests("random walk", rw) # ADF -1.54 (p=0.512) KPSS 2.456 (p=0.010): unit root
tests("diff(random walk)", np.diff(rw)) # ADF -13.36 (p=0.000) KPSS 0.152 (p=0.100): I(1)
tests("trend + AR(1), 'c'", ts) # ADF -1.20 (p=0.672) KPSS 2.743 (p=0.010): wrong question
tests("trend + AR(1), 'ct'", ts, "ct") # ADF -8.94 (p=0.000) KPSS 0.051 (p=0.100): trend-stationary
tests("random walk, 'ct'", rw, "ct") # ADF -1.43 (p=0.853) KPSS 0.235 (p=0.010): a trend does not fix a unit root
# --- differencing by hand and with pandas ---
y = pd.Series([100, 104, 107, 112, 115, 120])
print(y.diff().tolist()) # [nan, 4.0, 3.0, 5.0, 3.0, 5.0] first value is lost
print(y.diff().mean()) # 4.0 = (120 - 100) / 5, the average step
print((y.iloc[0] + y.diff().fillna(0).cumsum()).tolist()) # undo: cumulative sum gives y back
week = [20, 22, 21, 23, 30, 45, 40]
s = pd.Series(week + [v + 2 for v in week]) # two weeks, same weekly shape, +2 growth
print(s.diff(7).dropna().tolist()) # [2.0, 2.0, 2.0, 2.0, 2.0, 2.0, 2.0] pattern gone
# --- over-differencing white noise creates lag-1 autocorrelation -0.5 ---
e = rng.normal(size=5000)
print(acf(np.diff(e), nlags=2).round(2)) # [ 1. -0.48 -0.02] about -0.5 at lag 1
# --- rolling statistics: a quick visual check ---
z = pd.Series(rw)
print(z.rolling(50).mean().iloc[[49, 149, 249, 299]].round(2).tolist()) # [-0.09, -11.18, -24.57, -19.52] drifts
print(pd.Series(ar).rolling(50).mean().iloc[[49, 149, 249, 299]].round(2).tolist()) # [-0.26, 0.24, -0.28, 0.15] near 0
1. Which of these series is (weakly) stationary?
2. What is the null hypothesis of the ADF test?
3. On a series (regression "c"), ADF gives p = 0.45 and KPSS gives p ≤ 0.01. The best reading is…
4. After differencing, the ACF shows $r_1 = -0.49$ and nothing else. What happened?
5. Daily orders have a fixed weekly pattern plus a steady growth of 0.5 orders per day. What does $\nabla_7 y_t$ look like?
6. Which statement about your Prophet-style model and stationarity is right?
Practice problems
A. $y = 12, 15, 19, 24, 30$. How many differences make it constant? Use that to forecast the next value.
$\nabla y = 3, 4, 5, 6$; $\nabla^2 y = 1, 1, 1$: constant after $d = 2$. Forecast: the next second difference is 1, so the next first difference is $6 + 1 = 7$, and the next value is $30 + 7 = 37$ (undoing the differences step by step).
B. Show that differencing an AR(1) with $\varphi = 0.8$ gives lag-1 autocorrelation $-0.1$.
Let $\gamma_k = \gamma_0\varphi^k$. For $z_t = y_t - y_{t-1}$: $Var(z_t) = 2\gamma_0 - 2\gamma_1 = 2\gamma_0(1 - \varphi)$ and $Cov(z_t, z_{t-1}) = \gamma_1 - \gamma_2 - \gamma_0 + \gamma_1 = \gamma_0(2\varphi - \varphi^2 - 1) = -\gamma_0(1 - \varphi)^2$. So $\rho_1 = -(1 - \varphi)^2/(2(1 - \varphi)) = -(1 - \varphi)/2 = -0.2/2 = -0.1$. The closer $\varphi$ is to 1, the less harm differencing does (at $\varphi = 1$ it is exactly right).
C. A shock of +20 hits a trend-stationary series with AR(1) noise, $\varphi = 0.7$. How much is left after 3 days, and what is the half-life? What if the series were a random walk?
$20 \times 0.7^3 = 20 \times 0.343 = 6.86$. Half-life: $\ln 0.5/\ln 0.7 \approx 1.94$ days. In a random walk the full $+20$ remains forever.
D. Give an example of a series that is weakly but not strictly stationary, and one that is strictly but not weakly stationary.
Weak only: independent draws with odd days $N(0, 1)$ and even days Uniform$(-\sqrt 3, \sqrt 3)$: same mean 0, variance 1 and zero covariances, but different shapes. Strict only: iid Cauchy draws: identical distribution every day, but no finite mean or variance, so the weak conditions cannot hold.
E. Interview: "Is your demand data stationary? Did you make it stationary before modelling?"
"No, the raw demand is not stationary: it has a trend with changepoints, weekly and yearly seasonality and holiday spikes. I did not difference it, because my model is a decomposition: it puts those parts into $g(t)$, $s(t)$ and $h(t)$ and fits on the levels. What must be stationary is the residual noise, so I check the residuals with rolling mean and standard deviation, their ACF and Ljung–Box. ARIMA would need the differenced series to be stationary instead."
F. Interview: "ADF and KPSS both say your daily series is non-stationary. Shouldn't you difference it first?"
"Not necessarily. Both tests treat a trend that changes slope as evidence of a unit root; a stationary series around a piecewise-linear trend fails them routinely. My model includes that piecewise trend, so I test the residuals, not the raw series. If the residuals are stationary, the decomposition has handled the non-stationarity. If I wanted to compare with ARIMA, I would let a rolling-origin backtest and long-horizon interval coverage decide between the trend-stationary and unit-root views."
Baselines and exponential smoothing
Before you trust a clever forecasting model, you must know what "doing almost nothing" would have scored. This chapter builds the simplest honest forecasts (copy the last value, copy last week, the average, a straight line, the average of the last few days) and then the exponential-smoothing family (SES, Holt, Holt–Winters, ETS), which many forecasting teams use as their strong default. All the way through, we ask the interview question: why would someone choose this instead of your Prophet-style Bayesian model, and when should your model win?
- Make the five baseline forecasts by hand: naive, seasonal naive, mean, drift and moving average, and score them on a holdout with the MAE
- Know which baseline is the best possible forecast for which kind of series, and why a model that cannot beat the right baseline adds nothing
- Understand simple exponential smoothing as a weighted average whose weights shrink geometrically, and as "move the level a fraction α of each surprise"
- Add a trend (Holt, and the damped version) and a season (Holt–Winters, additive and multiplicative), and read every update equation in words
- See exponential smoothing as a statistical model (ETS) with a likelihood and prediction intervals that widen with the horizon
- Explain clearly when exponential smoothing beats a Prophet-style model and when it loses, and run all of it in
statsmodels
What we need from earlier chapters: forecast origin, horizon and "never shuffle" (Chapter 7.1); trend and seasonality (Chapter 7.2); random walks and white noise (Chapter 7.3); stationarity (Chapter 7.4); the mean and the Normal distribution (Chapters 4.5, 4.9); maximum likelihood (Chapter 5.2); regression (Chapter 5.13). Notation: $y_t$ is the value at time $t$ (for us, usually orders on day $t$). $T$ is the forecast origin, the last time we have seen. $h$ is the horizon, how many steps ahead we forecast. $\hat y_{T+h\mid T}$ ("y hat at T plus h given T") is the forecast for time $T+h$ made with data up to $T$. $m$ is the season length (the number of steps in one repeating cycle: $m = 7$ for daily data with a weekly pattern). $\ell_t$ ("ell") is a level, $b_t$ a trend (slope per step), $s_t$ a seasonal effect; $\alpha, \beta, \gamma$ ("alpha, beta, gamma") are smoothing parameters between 0 and 1; $\phi$ ("phi") is a damping factor. An error (or forecast error) is $e = y - \hat y$, actual minus forecast.
The naive and seasonal naive forecasts core
A weather forecaster's oldest trick: "tomorrow will be like today." It sounds lazy, yet for many places it is right surprisingly often, and every new weather model must prove it beats this rule. A restaurant manager uses a cleverer version: "this Saturday will be like last Saturday." Saturdays are busy, Mondays are quiet, so copying the same day of last week keeps the weekly rhythm.
These two rules are the naive and the seasonal naive forecasts. They need no fitting and no parameters. That is exactly why they are useful: they are the baseline, the yardstick. A baseline is a simple forecast that any serious method must beat; if your model cannot beat it, the model adds nothing.
To compare forecasts fairly we hide the most recent days (the holdout), forecast them from the older days only, and measure the misses. The simplest score is the MAE (mean absolute error): the average size of the misses, in the same units as the data.
Three ways to say it:
- Picture: naive draws a flat line at the last value; seasonal naive photocopies the last week and pastes it into the future.
- Numbers: if last Sunday had 37 orders, naive says 37 for every future day; if last Monday had 22, seasonal naive says 22 for next Monday.
- Slogan: "same as last time" and "same as this time last cycle" are the forecasts to beat.
A shop's daily orders for two weeks (Monday to Sunday). We forecast the next Monday, Tuesday and Wednesday, then compare with what really happened.
| Mon | Tue | Wed | Thu | Fri | Sat | Sun | |
|---|---|---|---|---|---|---|---|
| week 1 | 20 | 22 | 21 | 23 | 30 | 40 | 35 |
| week 2 | 22 | 24 | 23 | 25 | 32 | 42 | 37 |
| week 3 (holdout) | 24 | 25 | 25 | not used | |||
- Naive: the last value seen is Sunday of week 2: 37. So $\hat y = 37$ for Monday, Tuesday and Wednesday.
- Naive errors (actual − forecast): $24 - 37 = -13$, $25 - 37 = -12$, $25 - 37 = -12$. MAE $= (13 + 12 + 12)/3 = 37/3 \approx 12.33$ orders.
- Seasonal naive ($m = 7$): copy the same weekday from the last week seen: Monday 22, Tuesday 24, Wednesday 23.
- Seasonal naive errors: $24 - 22 = 2$, $25 - 24 = 1$, $25 - 23 = 2$. MAE $= (2 + 1 + 2)/3 = 5/3 \approx 1.67$ orders.
- Conclusion: on this weekly data, seasonal naive misses by about 1.7 orders a day, naive by about 12.3. Any model for this shop must beat 1.67, not 12.33.
Given data $y_1, \dots, y_T$ (the origin is $T$), for every horizon $h = 1, 2, \dots$:
$$\text{naive: } \hat y_{T+h\mid T} = y_T, \qquad \text{seasonal naive: } \hat y_{T+h\mid T} = y_{T+h-m(k+1)}, \quad k = \left\lfloor \tfrac{h-1}{m} \right\rfloor .$$- $\lfloor x \rfloor$ means "round down". The index $T + h - m(k+1)$ is simply "the same position in the last full cycle you have seen". For $h \le m$ it is $y_{T+h-m}$.
- Neither method has parameters to fit; seasonal naive needs at least one full cycle of history ($T \ge m$).
- When each is optimal: if the series is a random walk ($y_t = y_{t-1} + \varepsilon_t$ with independent mean-zero shocks, Chapter 7.3), the best forecast in squared error is exactly $y_T$, so naive is optimal. If it is a seasonal random walk ($y_t = y_{t-m} + \varepsilon_t$), seasonal naive is optimal.
- Holdout scoring: keep the last $H$ points aside, forecast them from the rest, and compute $\text{MAE} = \frac{1}{H}\sum_{h=1}^{H} \lvert y_{T+h} - \hat y_{T+h\mid T}\rvert$. (Other scores such as RMSE, MAPE and MASE are compared in Chapter 7.15.)
Why do we need it?
"Our model has an MAE of 6" means nothing on its own. Is 6 good? Only a baseline answers that. Naive and seasonal naive cost nothing to compute, so they tell you the accuracy you get for free.
Where is it used?
As the reference in every forecasting benchmark (the M-competitions, the scaled error MASE divides by the naive or seasonal naive error), in retail and staffing dashboards ("same day last week"), as a fallback when a fancy model fails in production, and as the yardstick for your Prophet-style model.
How is it used?
Split by time: train on everything before the origin, hold out the last weeks. Compute y_train[-1] (naive) and np.tile(y_train[-m:], ...) (seasonal naive), score them with the MAE, and write these numbers next to your model's score in every report.
"A baseline is a weak straw man I include so that my model looks good."
A baseline should be the strongest simple method that suits the data. For daily data with a weekly rhythm that is seasonal naive, not naive. Beating a weak baseline proves nothing.
"I'll score the forecasts on a random 20% of the days."
Hold out the last part of the series. A random holdout lets the method see the days just before and after each test day, which is information you never have when you forecast for real (Chapter 7.15).
"Seasonal naive averages all past Mondays."
It copies exactly one value: the most recent Monday. Averaging several past Mondays is a different baseline (a "seasonal mean"); it is less noisy but slower to react.
When you describe your forecasting model's accuracy in an interview, put its holdout error next to seasonal naive with $m = 7$ (for daily data with a weekly cycle) and next to naive. The first follow-up to "our MAE was 6.2" is always "compared to what?". A sentence like "seasonal naive scored 9.0 on the same rolling holdout, so the model cuts the error by about 30%" answers it before it is asked.
Naive: $\hat y_{T+h\mid T} = y_T$ (optimal for a random walk). Seasonal naive: copy the same position from the last cycle (optimal for $y_t = y_{t-m} + \varepsilon_t$).
Score on a holdout taken from the end of the series: MAE $= \frac1H\sum\lvert y - \hat y\rvert$.
Trap: always compare against the right baseline (seasonal naive for seasonal data).
Quick check: monthly data with a yearly pattern, last observed month is March 2025. What does seasonal naive forecast for May 2025 and for May 2026?
Here $m = 12$. For May 2025 ($h = 2$) it copies May 2024. For May 2026 ($h = 14$), $k = \lfloor 13/12 \rfloor = 1$, so it goes back $m(k+1) = 24$ months from May 2026: again May 2024, the most recent May it has seen. Seasonal naive repeats the last observed cycle forever.
The mean and drift forecasts
Two more yardsticks. The mean forecast says: "the future will look like the long-run average of everything I have seen." It is the right idea for a series that just jitters around a fixed level, like the daily number of typos in a long book.
The drift forecast says: "draw a straight line from the first point to the last point, and keep going." It is naive plus the average growth per step. It suits a series that wanders but tends to climb, like a company's user count.
Three ways to say it:
- Picture: mean = a flat line through the middle of the cloud; drift = a ruler laid from the first dot to the last dot and extended.
- Numbers: weekly sign-ups 50, 52, 55, 56, 58, 65: mean forecast 56, drift forecast 65 + 3 = 68 next week.
- Slogan: mean remembers everything equally; drift remembers only where you started, where you are, and how long it took.
Weekly sign-ups for six weeks: $y = 50, 52, 55, 56, 58, 65$ (so $T = 6$). Forecast the next three weeks.
- Mean: $(50 + 52 + 55 + 56 + 58 + 65)/6 = 336/6 = 56$. Forecast 56 for weeks 7, 8 and 9.
- Naive for comparison: 65 for all three weeks.
- Drift slope: $(y_T - y_1)/(T - 1) = (65 - 50)/(6 - 1) = 15/5 = 3$ sign-ups per week.
- Same slope from the week-to-week changes: $2, 3, 1, 2, 7$, sum $15$, average $15/5 = 3$. The middle values cancel out ("telescoping"), so only the first and last values matter.
- Drift forecasts: $65 + 1\times3 = 68$, $65 + 2\times3 = 71$, $65 + 3\times3 = 74$.
Notice how different the three answers are for week 9: 56 (mean), 65 (naive), 74 (drift). Which one is right depends on what kind of process made the data.
For data $y_1, \dots, y_T$:
$$\text{mean: } \hat y_{T+h\mid T} = \bar y = \frac{1}{T}\sum_{t=1}^{T} y_t, \qquad \text{drift: } \hat y_{T+h\mid T} = y_T + h\,\frac{y_T - y_1}{T - 1}.$$- The drift slope equals the average one-step change: $\frac{1}{T-1}\sum_{t=2}^{T}(y_t - y_{t-1}) = \frac{y_T - y_1}{T-1}$.
- When each is optimal: the mean is the best forecast (in squared error) when $y_t = \mu + \varepsilon_t$ with independent noise: a constant level plus noise. Drift is the natural forecast for a random walk with drift, $y_t = y_{t-1} + c + \varepsilon_t$, because $\frac{y_T - y_1}{T-1}$ is the average of the observed steps, an estimate of $c$.
- Both have no parameters to choose. Drift needs $T \ge 2$.
Why do we need it?
Naive ignores everything except the last value. The mean uses all values equally, which is best when the level never moves; drift adds the typical growth per step, which is the cheapest way to say "this series has been climbing".
Where is it used?
Mean: stable metrics, error rates, the "climatology" forecast in weather. Drift: growth metrics (users, revenue) and as the default "random walk with drift" benchmark in finance and forecasting libraries (R's rwf(drift=TRUE), sktime's NaiveForecaster(strategy="drift")).
How is it used?
Mean: np.repeat(y.mean(), h). Drift: y[-1] + (y[-1] - y[0]) / (len(y) - 1) * np.arange(1, h + 1). Score both with the other baselines on the same holdout and keep the best as the bar to clear.
"The drift line is a regression line through all the points."
It only uses the first and last values (the middle steps cancel). One unusual value at the start or at the end can swing it a lot. A regression trend uses every point (you will build one in Chapter 7.6).
"The mean forecast is useless; nobody forecasts with an average."
For a stable, noisy series (no trend, no season) the mean beats naive, because naive copies the last day's random noise into every future day while the mean averages the noise away.
"Drift is the same as Holt's trend method."
Drift uses one fixed slope from the whole history. Holt (later in this chapter) keeps updating a local slope that follows recent changes.
Mean: $\hat y = \bar y$ (optimal for constant level + noise). Drift: $\hat y_{T+h\mid T} = y_T + h\,\frac{y_T - y_1}{T-1}$ (random walk with drift).
Drift slope = average one-step change = depends only on the first and last value.
Trap: an odd first or last value distorts drift; the mean ignores trends entirely.
Quick check: a series of 11 daily values starts at 200 and ends at 250. What does the drift method forecast 4 days ahead?
Slope $= (250 - 200)/(11 - 1) = 5$ per day. Forecast $= 250 + 4\times5 = 270$. The nine values in between do not matter.
Moving-average forecasts: average the last k values
Naive trusts only yesterday, so it copies yesterday's random noise. The mean trusts all of history, so it cannot follow a level that moves. The moving average sits in between: average the last $k$ days and use that as the forecast. A few days of averaging cancels much of the noise, but the average is always a little "behind the times", because it includes days that are already a while ago.
So there is a dial, $k$. Small $k$: quick to react, but noisy. Large $k$: smooth, but slow; on a rising series it is always too low. This tug-of-war between noise and lag appears in every smoothing method in this chapter.
Three ways to say it:
- Picture: a window of width $k$ slides along the series; the forecast is the average height inside the window.
- Numbers: last three days 16, 14, 18 → forecast $(16 + 14 + 18)/3 = 16$ for every future day.
- Slogan: a bigger window means less noise but more lag.
Six days of orders: 10, 14, 12, 16, 14, 18. Use a 3-day trailing moving average ($k = 3$).
- After day 3: $(10 + 14 + 12)/3 = 36/3 = 12$. This is the forecast for day 4. Actual 16: error $16 - 12 = +4$.
- After day 4: $(14 + 12 + 16)/3 = 42/3 = 14$. Forecast for day 5; actual 14: error $0$.
- After day 5: $(12 + 16 + 14)/3 = 42/3 = 14$. Forecast for day 6; actual 18: error $+4$.
- After day 6: $(16 + 14 + 18)/3 = 48/3 = 16$. This is the forecast for day 7, day 8, and every later day.
- The errors $+4, 0, +4$ are never negative: the series is climbing and the average is always behind it. On a perfectly straight line with slope $b$ per day, the $k$-day average is always $b(k+1)/2$ too low for the next day. With $b = 2$ and $k = 3$ that is $2\times4/2 = 4$.
The (trailing) moving-average forecast with window $k$ is
$$\hat y_{T+h\mid T} = \frac{1}{k}\left(y_T + y_{T-1} + \dots + y_{T-k+1}\right) \quad\text{for every } h \ge 1 .$$- It is a weighted average with weight $1/k$ on each of the last $k$ values and weight 0 on everything older.
- $k = 1$ gives naive; $k = T$ gives the mean. The forecast is flat (no trend, no season).
- On a straight line $y_t = a + bt$ the lag error is exactly $b(k+1)/2$ for the next step: the window's average sits at its middle time, $(k-1)/2$ steps back, and the target is one more step ahead.
- With $k = m$ (one full season) the seasonal pattern averages out completely, so the forecast is the de-seasonalized level. That is useful for smoothing, but it is a poor forecast for a seasonal series.
- "Trailing" means it uses only past values. A centred moving average (half the window before, half after) is a smoother for decomposition (Chapter 7.2); it needs future values, so it cannot be a forecast.
Why do we need it?
Real series are noisy. Averaging a few recent values removes part of the noise while still following slow changes in the level. It is the simplest "smoothing" idea, and exponential smoothing is a refinement of it.
Where is it used?
7-day rolling averages in dashboards (cases, sales, traffic), inventory reorder rules, technical indicators in finance, the trend step of classical decomposition, and as a smoothed baseline for a noisy metric, such as a noisy training loss curve.
How is it used?
In pandas: y.rolling(k).mean() gives the trailing average at each time; the last value is the forecast. Choose $k$ by holdout error, or set $k = m$ when you only want a de-seasonalized level to look at.
"A bigger window is always safer."
A bigger window removes more noise but adds more lag. On a trending series the lag $b(k+1)/2$ grows with $k$, so past some point a bigger window makes things worse. Choose $k$ by holdout error.
"A 7-day moving average is a good forecast for data with a weekly pattern."
It is a good summary of the level, because it removes the weekly pattern. As a forecast it predicts the same value for Saturday and Monday, which is exactly what you do not want.
"A moving-average forecast is the same as an MA model in ARIMA."
They share a name and nothing else. The moving-average forecast averages past observed values $y$. An MA($q$) model (Chapter 7.6) says each value is a mean plus a weighted sum of the last $q$ unobserved random shocks $\varepsilon$.
Model answer: "A rolling mean is a smoothing rule on the data. MA($q$) is a probability model for the noise: $y_t = \mu + \varepsilon_t + \theta_1\varepsilon_{t-1} + \dots + \theta_q\varepsilon_{t-q}$. Its signature is an autocorrelation function that cuts off after lag $q$."
MA forecast: $\hat y_{T+h\mid T} = \frac1k\sum_{j=0}^{k-1} y_{T-j}$, flat; $k=1$ → naive, $k = T$ → mean.
Noise vs lag: on a line with slope $b$, it is $b(k+1)/2$ too low.
Trap: a rolling mean of $y$ ≠ the MA($q$) model of shocks.
Quick check: daily sales grow by exactly 3 per day with no noise. How far behind is a 9-day trailing average forecast for tomorrow?
Lag $= b(k+1)/2 = 3\times10/2 = 15$ units too low. The average of the last 9 days sits at the middle of the window, 4 days back ($4\times3 = 12$ below today), and tomorrow is 3 above today: $12 + 3 = 15$.
Why strong baselines matter, and when they win core
An investment fund that earned 8% last year sounds great, until you learn the stock-market index (which you could buy with no skill at all) earned 12%. Forecasting works the same way. A model is only as good as its margin over the cheapest sensible alternative.
There is a deeper reason simple baselines are hard to beat: each one is the best possible forecast for some simple kind of series. Naive is optimal for a random walk, the mean for noise around a fixed level, seasonal naive for a seasonal random walk, drift for a random walk with drift. Many real series are close to one of these. When they are, a complex model can only lose: it has more knobs, and every knob it tunes on noise makes it worse.
Three ways to say it:
- Picture: five simple rulers; for every kind of series, one of them already fits almost perfectly.
- Numbers: model MAE 8 vs seasonal naive MAE 10 is a 20% gain; model MAE 12 vs 10 means the model is worse than copying last week.
- Slogan: no baseline, no claim.
Skill scores. On the same holdout, seasonal naive has MAE 10 orders per day.
- Model A has MAE 8. Its skill is $1 - 8/10 = 0.2$: it removes 20% of the baseline's error.
- Model B has MAE 12. Skill $= 1 - 12/10 = -0.2$: it is 20% worse than copying last week.
- Model C has MAE 9.8. Skill $= 1 - 9.8/10 = 0.02$, a 2% gain. With one holdout window this small difference could easily be luck; you would need many forecast origins (Chapter 7.15) before believing it.
- Scaled errors do the same thing in one number: MASE divides the model's MAE by the in-sample MAE of the (seasonal) naive forecast, so MASE $\lt 1$ means "better than naive" (owned by Chapter 7.15).
For an error measure $E$ (MAE, RMSE, …) computed on the same holdout,
$$\text{Skill} = 1 - \frac{E_{\text{model}}}{E_{\text{baseline}}}\quad (\gt 0:\ \text{better than the baseline},\ \lt 0:\ \text{worse}).$$Which simple forecast is optimal (minimum expected squared error, with the parameters known) for which process:
| Process (how the data were made) | Best simple forecast |
|---|---|
| $y_t = \mu + \varepsilon_t$ (constant level + noise) | mean |
| $y_t = y_{t-1} + \varepsilon_t$ (random walk) | naive |
| $y_t = y_{t-1} + c + \varepsilon_t$ (random walk with drift) | drift |
| $y_t = y_{t-m} + \varepsilon_t$ (seasonal random walk) | seasonal naive |
| slowly wandering level + noise ("local level") | simple exponential smoothing (next section); a moving average is a rough version |
Here $\varepsilon_t$ are independent, mean-zero shocks. A strong baseline means: the best of these simple methods for your series, scored on the same holdout as your model.
Why do we need it?
Without a strong baseline you cannot tell real skill from complexity. Complex models can look impressive on a chart and still lose to "same as last week" on fresh data.
Where is it used?
The M3 and M4 forecasting competitions (where simple methods and exponential smoothing were hard to beat and many sophisticated entries did not beat them), MASE and skill scores in papers and dashboards, model-release checklists in industry forecasting teams.
How is it used?
Before building anything, compute naive, seasonal naive, mean, drift and a moving average on rolling holdouts. Keep the best as the bar. Report every model as "error, and skill versus the best baseline", over many origins rather than one.
"My model beat the baseline on the holdout, so it is better."
One holdout is one draw. Repeat the comparison over many forecast origins (rolling-origin evaluation, Chapter 7.15) and look at the average and the spread of the difference.
"Baselines are for beginners; real teams use machine learning."
Real teams keep baselines forever: as the bar in every evaluation, as a sanity check in monitoring (if the model suddenly loses to seasonal naive, something broke), and as a fallback in production.
"If a baseline wins, forecasting this series is hopeless."
It means the series behaves like a simple process (often a random walk). That is useful knowledge: you now know the honest uncertainty, and you can stop over-engineering.
Why choose a baseline instead of your Prophet-style model, and vice versa?
- Choose (or at least keep) the baseline when: the series behaves like a random walk (no stable trend or calendar pattern to learn); the history is very short (a Prophet-style model cannot learn weekly and yearly patterns from a few weeks); you forecast thousands of series and need something cheap, fast and impossible to mis-fit; or your model's rolling-origin error is not clearly lower.
- Choose the Prophet-style model when: there is real structure the baselines cannot express: a trend that bends at changepoints, several seasonalities (weekly and yearly), holidays that move around the calendar, regressors such as price or promotions, a count likelihood (Negative Binomial) and honest predictive intervals. Then it should beat seasonal naive clearly, and you should show that it does.
"Our forecasting model has 8% MAPE, so it is very accurate."
"Our model's MAPE is 8% over 20 rolling origins; seasonal naive scores 11% on the same origins, so the model removes about a quarter of the baseline's error."
Model answer: "An error number only means something next to a baseline evaluated in the same way. I compare against seasonal naive and drift over many forecast origins, and I report skill or MASE. If the model cannot beat them, its complexity is not justified."
Skill $= 1 - E_{\text{model}}/E_{\text{baseline}}$; MASE $\lt 1$ = beats naive (7.15).
Each baseline is optimal for a simple process: mean ↔ level + noise, naive ↔ random walk, drift ↔ random walk with drift, seasonal naive ↔ seasonal random walk.
Trap: comparing against a weak baseline, or on a single holdout window.
Quick check: on a holdout, naive has RMSE 15, seasonal naive RMSE 9, your model RMSE 8. What is your model's skill, and versus which baseline should you report it?
Report it against the strongest baseline, seasonal naive: skill $= 1 - 8/9 \approx 0.11$ (11% better). Against naive it would look like $1 - 8/15 \approx 0.47$, which flatters the model because naive is a weak baseline for seasonal data.
Simple exponential smoothing (SES): fading memory core
You guess your commute will take 30 minutes. Today it took 40. You do not now expect 40 tomorrow (that overreacts to one bad day), and you do not stay at 30 (that ignores the news). You move part of the way: "30% of the surprise", so your new guess is $30 + 0.3\times10 = 33$ minutes. Do this every day and you have simple exponential smoothing.
The running guess is called the level: our current estimate of "where the series is now". The fraction you move is $\alpha$ (alpha), between 0 and 1. Unrolled, the level turns out to be a weighted average of all past values in which the newest value gets weight $\alpha$, the one before $\alpha(1-\alpha)$, then $\alpha(1-\alpha)^2$, and so on: the weights shrink geometrically (by the same factor each step back). That geometric fade is why it is called "exponential". It sits between naive ($\alpha = 1$: only the last value counts) and the mean (all values count equally).
Three ways to say it:
- Picture: a memory that fades: yesterday is vivid, last week is blurry, last year is almost gone.
- Numbers: with $\alpha = 0.5$ the last four days get weights 0.5, 0.25, 0.125, 0.0625.
- Slogan: new level = old level + α × surprise.
Data 20, 24, 22, 26 with $\alpha = 0.5$ and a starting level $\ell_0 = 20$. Each forecast is the current level; after seeing $y_t$ we update the level.
- Day 1: forecast $\ell_0 = 20$; actual 20; surprise $0$; new level $\ell_1 = 20 + 0.5\times0 = 20$.
- Day 2: forecast 20; actual 24; surprise $+4$; $\ell_2 = 20 + 0.5\times4 = 22$.
- Day 3: forecast 22; actual 22; surprise $0$; $\ell_3 = 22$.
- Day 4: forecast 22; actual 26; surprise $+4$; $\ell_4 = 22 + 0.5\times4 = 24$. The forecast for day 5, 6, 7, … is 24 (flat).
- Check with the weights: $0.5\times26 + 0.25\times22 + 0.125\times24 + 0.0625\times20 + 0.0625\times\ell_0 = 13 + 5.5 + 3 + 1.25 + 1.25 = 24$ ✓. The weights $0.5 + 0.25 + 0.125 + 0.0625 + 0.0625$ add up to 1; the last one is the weight left on the starting level.
The Code-it block reproduces these numbers with statsmodels.
Simple exponential smoothing with smoothing parameter $0 \lt \alpha \le 1$:
$$\ell_t = \alpha\,y_t + (1-\alpha)\,\ell_{t-1}, \qquad \hat y_{T+h\mid T} = \ell_T \ \text{ for all } h \ge 1 .$$- Error-correction form: $\ell_t = \ell_{t-1} + \alpha\,e_t$ with $e_t = y_t - \ell_{t-1}$, the one-step forecast error ("surprise").
- Weighted-average form: $\ell_T = \sum_{j=0}^{T-1}\alpha(1-\alpha)^j\,y_{T-j} + (1-\alpha)^T\ell_0$. The weights add to 1.
- $\alpha = 1$ → naive. Small $\alpha$ → long memory, heavy smoothing. The average age of the weights is $(1-\alpha)/\alpha$ steps; a $k$-day moving average with the same average age has $k = 2/\alpha - 1$ (this is pandas'
span: $\alpha = 2/(\text{span}+1)$). - Choosing α and ℓ₀: minimize the sum of squared one-step errors $\text{SSE}(\alpha) = \sum_t (y_t - \ell_{t-1})^2$ over the history (what
statsmodels'ExponentialSmoothing.fit()does). For the additive-error model of the ETS section, maximizing the Gaussian likelihood picks the same values. - SES has no trend and no season: its forecast is a flat line at the last level. It is the optimal forecast for a "local level" process (a random-walk level observed with noise).
Why do we need it?
We want a forecast that follows a level which slowly moves, without copying every bit of noise. A moving average does this with a hard cut-off; SES does it with a smooth fade and only one number to remember (the level).
Where is it used?
Demand forecasting for slow movers, inventory systems, "EWMA" control charts in factories, pandas ewm, the moving averages of gradients inside the Adam optimizer ($\beta_1 = 0.9$ is SES with $\alpha = 0.1$), and smoothing a noisy loss or ELBO curve before deciding whether training has stopped improving.
How is it used?
SimpleExpSmoothing(y).fit() estimates α by minimizing the one-step SSE; read params["smoothing_level"] and forecast(h). An α close to 1 says "this behaves like a random walk"; a small α says "mostly noise around a slowly moving level".
"Exponential smoothing fits an exponential curve to the data."
The weights decay exponentially (geometrically). The forecast itself is a flat line at the last level.
"A larger α means more smoothing."
The opposite. Large α means "move a lot towards each new value": little smoothing, close to naive. Small α means heavy smoothing and long memory.
"pandas y.ewm(alpha=0.3).mean() is exactly SES."
Only with adjust=False, which uses the recursion $\ell_t = \alpha y_t + (1-\alpha)\ell_{t-1}$ starting from $y_0$. The default adjust=True re-normalizes the weights over the finite history, which differs at the start. Also, ewm(...).mean() at time $t$ is the level after seeing $y_t$, i.e. the forecast for $t+1$.
If your SVI training loop smooths the noisy ELBO before comparing it with the best value so far (a common way to make early stopping less jumpy, Chapter 6.14), an exponentially weighted moving average is exactly SES: the smoothed value moves a fraction α of each new noisy ELBO. A small α hides noise but reacts late to a real plateau, the same noise-versus-lag trade-off as in this section.
$\ell_t = \alpha y_t + (1-\alpha)\ell_{t-1} = \ell_{t-1} + \alpha e_t$; forecast $\hat y_{T+h\mid T} = \ell_T$ (flat).
Weights $\alpha(1-\alpha)^j$ fade geometrically; α = 1 → naive; small α → heavy smoothing. Choose α by minimizing one-step SSE.
Trap: large α = less smoothing; SES has no trend and no season.
Quick check: current level 50, α = 0.2, today's value is 60. What is the new level, and what does SES forecast for next week?
Surprise $= 60 - 50 = 10$; new level $= 50 + 0.2\times10 = 52$. SES forecasts 52 for every day of next week (a flat line).
Holt's linear trend method, and damping
SES forecasts a flat line. On a series that keeps growing, it is always behind, just like the moving average. Holt's method adds a second running estimate: the trend, the current slope (growth per step). Think of tracking a car: you keep an estimate of its position (the level) and of its speed (the trend). To predict where it will be in 5 minutes you take position + 5 × speed. Every new sighting nudges both estimates.
A straight line forever is often too bold. Sales rarely grow at the same rate for years. The damped version lets the slope fade out over the horizon, so long-range forecasts flatten. In large forecasting comparisons, damped trend methods have been among the most reliable simple methods.
Three ways to say it:
- Picture: SES is a flat ruler; Holt is a tilted ruler; damped Holt is a ruler that bends down to flat.
- Numbers: level 19.5 and trend 3.25 → forecasts 22.75, 26, 29.25 for the next three steps.
- Slogan: forecast = where you are + how fast you are moving × how far ahead.
Data 10, 13, 16, 20 with $\alpha = 0.5$ and $\beta = 0.5$. Start from the first two values: level $\ell_1 = 10$, trend $b_1 = 13 - 10 = 3$.
- Forecast for day 2: $\ell_1 + b_1 = 13$; actual 13. New level $\ell_2 = 0.5\times13 + 0.5\times(10 + 3) = 13$; new trend $b_2 = 0.5\times(13 - 10) + 0.5\times3 = 3$.
- Forecast for day 3: $13 + 3 = 16$; actual 16. $\ell_3 = 0.5\times16 + 0.5\times16 = 16$; $b_3 = 0.5\times(16 - 13) + 0.5\times3 = 3$.
- Forecast for day 4: $16 + 3 = 19$; actual 20 (surprise $+1$). $\ell_4 = 0.5\times20 + 0.5\times19 = 19.5$; $b_4 = 0.5\times(19.5 - 16) + 0.5\times3 = 1.75 + 1.5 = 3.25$.
- Forecasts: $\hat y_5 = 19.5 + 1\times3.25 = 22.75$, $\hat y_6 = 19.5 + 2\times3.25 = 26$, $\hat y_7 = 19.5 + 3\times3.25 = 29.25$.
- Damped with $\phi = 0.8$ (same level and trend): $\hat y_5 = 19.5 + 0.8\times3.25 = 22.1$, $\hat y_6 = 19.5 + (0.8 + 0.64)\times3.25 = 24.18$, and far ahead the forecasts level off at $19.5 + \frac{0.8}{1 - 0.8}\times3.25 = 19.5 + 13 = 32.5$.
Holt's linear trend method with smoothing parameters $0 \lt \alpha, \beta \le 1$:
$$\begin{aligned} \ell_t &= \alpha\,y_t + (1-\alpha)(\ell_{t-1} + b_{t-1}) && \text{(level: blend the new value with the old prediction)}\\ b_t &= \beta(\ell_t - \ell_{t-1}) + (1-\beta)\,b_{t-1} && \text{(trend: blend the latest change in level with the old slope)}\\ \hat y_{T+h\mid T} &= \ell_T + h\,b_T . \end{aligned}$$Damped trend ($0 \lt \phi \lt 1$; $\phi = 1$ gives Holt):
$$\ell_t = \alpha y_t + (1-\alpha)(\ell_{t-1} + \phi b_{t-1}),\quad b_t = \beta(\ell_t - \ell_{t-1}) + (1-\beta)\phi\,b_{t-1},\quad \hat y_{T+h\mid T} = \ell_T + (\phi + \phi^2 + \dots + \phi^h)\,b_T .$$- As $h \to \infty$ the damped forecast approaches $\ell_T + \frac{\phi}{1-\phi}b_T$: a horizontal line.
- $b_t$ is a local slope that follows recent changes; it is not the slope of a line through all the data.
- Notation differs between sources:
statsmodelscalls βsmoothing_trendand φdamping_trend; some textbooks write the trend parameter as β* and use β for α·β*. Check which one a number refers to before comparing. - Typical φ values are 0.8 to 0.98 (a rule of thumb; it is usually estimated).
Why do we need it?
Growing or shrinking series make SES lag forever. Holt follows a slope that itself can change slowly, and the damped version stops that slope from running away over long horizons.
Where is it used?
Sales and capacity planning for products in growth, the "Comb" benchmark of the M4 competition (the average of SES, Holt and damped Holt forecasts), automatic forecasting tools (statsmodels Holt, R forecast::holt, the ETS family) and as a strong baseline for trending metrics.
How is it used?
Holt(y, damped_trend=True).fit() estimates α, β (and φ) by minimizing one-step SSE. Look at the fitted trend $b_T$, plot the forecast, and prefer the damped version for long horizons unless you have a strong reason to extrapolate growth.
"Holt's forecast is a safe straight-line extrapolation."
It extrapolates the latest slope forever. For long horizons that is often too optimistic (or too pessimistic). The damped trend is a safer default unless the growth is known to continue.
"β is the slope of the trend."
β is a smoothing parameter: how quickly the slope estimate $b_t$ reacts to changes. The slope itself is $b_t$, in units of "data units per step".
"Holt can learn a trend that bends at a changepoint, just like the Prophet-style trend."
It adapts to a new slope gradually, after the fact (how fast depends on β), and it has no notion of a changepoint at a specific date. A piecewise-linear trend puts the bend at a chosen time and estimates the slope change there (Chapter 7.8).
Holt: $\ell_t = \alpha y_t + (1-\alpha)(\ell_{t-1} + b_{t-1})$, $b_t = \beta(\ell_t - \ell_{t-1}) + (1-\beta)b_{t-1}$, forecast $\ell_T + h b_T$.
Damped: forecast $\ell_T + (\phi + \dots + \phi^h)b_T \to \ell_T + \frac{\phi}{1-\phi}b_T$.
Trap: undamped trends run away at long horizons; β is a reaction speed, not a slope.
Quick check: level 100, trend 4 per week, φ = 0.5. What are the damped forecasts 1 and 2 weeks ahead, and where do they level off?
$h = 1$: $100 + 0.5\times4 = 102$. $h = 2$: $100 + (0.5 + 0.25)\times4 = 103$. Limit: $100 + \frac{0.5}{0.5}\times4 = 104$. With a strong damping the forecast flattens almost immediately.
Holt–Winters: adding a season (additive and multiplicative) core
Holt keeps two memories: where the series is (level) and how fast it moves (trend). Holt–Winters adds a third: one seasonal effect for every position in the cycle. For daily data with a weekly pattern that is 7 numbers: "Mondays are 5 below the level, Saturdays 20 above", and so on. Each day, only that day's seasonal number gets updated, using what we just saw.
There are two ways to combine the season with the level. Additive: "Saturday adds 20 orders", the same amount whether the shop is small or big. Multiplicative: "Saturday is 40% above the level", so the seasonal swing grows as the business grows.
Three ways to say it:
- Picture: a level line, a slope, and a row of seven little knobs (one per weekday) that are each re-tuned once a week.
- Numbers: level 104, trend 3, Tuesday effect +5 → forecast for Tuesday $= 104 + 3 + 5 = 112$.
- Slogan: forecast = level + trend × steps + the effect of that position in the cycle.
Quarterly data ($m = 4$), $\alpha = \beta = \gamma = 0.5$. Before the new quarter: level $\ell = 100$, trend $b = 2$, seasonal effects Q1 $= -10$, Q2 $= +5$, Q3 $= +15$, Q4 $= -10$. A new Q1 arrives with value $y = 96$.
- Forecast for this Q1: $\ell + b + s_{Q1} = 100 + 2 - 10 = 92$. Surprise $96 - 92 = +4$.
- Level: remove the season from the new value ($96 - (-10) = 106$) and blend it with the old prediction of the level ($100 + 2 = 102$): $\ell_{new} = 0.5\times106 + 0.5\times102 = 104$.
- Trend: the level rose by $104 - 100 = 4$; blend with the old slope: $b_{new} = 0.5\times4 + 0.5\times2 = 3$.
- Season Q1: what this quarter showed beyond the old level-plus-trend ($96 - 102 = -6$), blended with the old Q1 effect: $s_{Q1,new} = 0.5\times(-6) + 0.5\times(-10) = -8$. The other three effects are untouched.
- Forecasts: next quarter (Q2, $h = 1$): $104 + 3 + 5 = 112$. Next Q1 ($h = 4$): $104 + 4\times3 + (-8) = 108$.
Some textbooks update the season with $y - \ell_{new}$ instead of $y - (\ell + b)$; here that gives $0.5\times(96 - 104) + 0.5\times(-10) = -9$. The two versions are the same model with γ defined differently ($\gamma_{\text{here}} = \gamma_{\text{other}}(1-\alpha)$), so seasonal parameters from different sources are not directly comparable.
Additive Holt–Winters with season length $m$ and smoothing parameters $\alpha, \beta, \gamma \in (0, 1]$ (the form used by statsmodels' ExponentialSmoothing):
Multiplicative season (the swing is a percentage of the level; needs $y \gt 0$):
$$\ell_t = \alpha\,\frac{y_t}{s_{t-m}} + (1-\alpha)(\ell_{t-1} + b_{t-1}),\qquad s_t = \gamma\,\frac{y_t}{\ell_{t-1} + b_{t-1}} + (1-\gamma)\,s_{t-m},\qquad \hat y_{T+h\mid T} = (\ell_T + h b_T)\, s_{T+h-m(k+1)} .$$- Additive seasonal effects are in data units and add to about 0 over a cycle; multiplicative ones are factors around 1 (1.4 = "40% above the level").
- You need at least two full cycles of history to start the season sensibly (a rule of thumb; more is better).
- Classic Holt–Winters has one season length $m$. A trend can be dropped (no β) or damped (add φ as in Holt).
Why do we need it?
Most business series have a calendar rhythm. SES and Holt cannot represent it, so they miss every busy Saturday. Holt–Winters keeps a running estimate of each part of the cycle that adapts if the pattern slowly changes.
Where is it used?
Retail and call-centre staffing, energy load, web traffic, hotel bookings, supply-chain planning systems, monitoring tools that forecast "expected traffic now" to raise alerts, and as the strong seasonal benchmark next to a Prophet-style model.
How is it used?
ExponentialSmoothing(y, trend="add", seasonal="add", seasonal_periods=7).fit(); use seasonal="mul" when the swings grow with the level (or log-transform and stay additive). Compare it with seasonal naive on a holdout and inspect the fitted seasonal effects.
"Holt–Winters handles weekly and yearly seasonality at the same time."
Classic Holt–Winters has one season length $m$. For daily data with weekly and yearly patterns you need Fourier terms (Chapter 7.11), models built for several seasonalities (such as TBATS), or a Prophet-style model.
"If the series grows, I must use the multiplicative season."
Use multiplicative when the seasonal swing grows in proportion to the level. A growing level with a constant swing is additive. A common alternative is to log-transform (Chapter 4.18) and use an additive model.
"A γ of 0.1 in one library means the same as 0.1 in another."
Two equivalent seasonal update forms exist, with $\gamma$ defined differently. Compare forecasts, not parameter values, across tools.
Additive HW: level from the de-seasonalized value, trend from the change in level, season $s_t = \gamma(y_t - \ell_{t-1} - b_{t-1}) + (1-\gamma)s_{t-m}$.
Forecast $\ell_T + hb_T + s_{\text{same position, last cycle}}$; multiplicative: $(\ell_T + hb_T)\times s$.
Trap: one season length only; multiplicative only when the swing scales with the level.
Quick check: multiplicative weekly model, level 200, trend 0, Saturday factor 1.3. What is the Saturday forecast? What would an additive model with Saturday effect +60 give if the level doubled to 400?
Multiplicative: $200\times1.3 = 260$. If the level doubles to 400, multiplicative gives $520$ (the swing doubles too), while additive gives $400 + 60 = 460$ (the swing stays 60 orders).
ETS: exponential smoothing as a statistical model, with prediction intervals
So far exponential smoothing has been a recipe: update the level, update the trend, forecast. A recipe gives a single line, not an answer to "how sure are you?". The ETS view writes the same recipe as a statistical model: every day the value is the level plus a random shock, and that same shock also nudges the level. Now we can write down a likelihood, estimate the parameters by maximum likelihood, compare variants with AIC, and simulate the future to get prediction intervals.
Why do the intervals widen? One step ahead, you only face one new shock. Ten steps ahead, nine earlier shocks will each have moved the level a bit (by $\alpha$ times the shock) before the tenth arrives. Uncertainty piles up.
Three ways to say it:
- Picture: a fan that opens to the right: many possible futures that start together and spread apart.
- Numbers: with α = 0.5 and shock sd 4, the forecast sd is 4 one step ahead and 8 thirteen steps ahead.
- Slogan: E for Error, T for Trend, S for Season: the recipe plus an honest noise model.
SES as a model with $\alpha = 0.5$ and shock standard deviation $\sigma = 4$. The variance of the error $h$ steps ahead is $\sigma^2[1 + \alpha^2(h-1)]$.
- $h = 1$: $16\times[1 + 0.25\times0] = 16$, so sd $= 4$. An 80% interval is forecast $\pm 1.2816\times4 = \pm5.13$.
- $h = 5$: $16\times[1 + 0.25\times4] = 16\times2 = 32$, sd $= \sqrt{32} \approx 5.66$.
- $h = 13$: $16\times[1 + 0.25\times12] = 16\times4 = 64$, sd $= 8$. The 95% interval is $\pm1.96\times8 = \pm15.68$, about twice as wide as at $h = 1$.
- With $\alpha = 0.1$ the same $h = 13$ gives $16\times[1 + 0.01\times12] = 17.92$, sd $\approx 4.23$: a slowly moving level means the far future is not much less certain than tomorrow.
ETS(A,N,N) (additive error, no trend, no season), the model behind SES:
$$y_t = \ell_{t-1} + \varepsilon_t, \qquad \ell_t = \ell_{t-1} + \alpha\,\varepsilon_t, \qquad \varepsilon_t \overset{iid}{\sim} N(0, \sigma^2).$$- $\varepsilon_t = y_t - \ell_{t-1}$ is the one-step error, so the level equation is exactly the SES update. Point forecast $\hat y_{T+h\mid T} = \ell_T$; forecast variance $\sigma_h^2 = \sigma^2[1 + \alpha^2(h-1)]$. (Here $N(0,\sigma^2)$ is written with the variance;
numpyandscipytake the sd.) - The name ETS(Error, Trend, Season): Error ∈ {A, M} (additive or multiplicative noise), Trend ∈ {N, A, Ad} (none, additive, damped), Season ∈ {N, A, M}. SES = ETS(A,N,N), Holt = ETS(A,A,N), damped Holt = ETS(A,Ad,N), additive Holt–Winters = ETS(A,A,A).
- With the same parameter values, additive- and multiplicative-error versions give the same point forecasts but different intervals.
- Parameters (α, β, γ, φ, initial states, σ) are fitted by maximum likelihood (
statsmodelsETSModel); variants are compared with AIC/AICc. - These intervals assume the model is right and treat the fitted parameters as known: they include future shocks, not uncertainty about α or σ.
Why do we need it?
Decisions need ranges, not single numbers: how much stock covers 95% of plausible demand? The ETS model turns the smoothing recipe into a probability model, so intervals, likelihoods and model selection all become possible.
Where is it used?
Automatic forecasting software (statsmodels ETSModel, R forecast::ets and fable::ETS, Nixtla's statsforecast AutoETS), inventory safety-stock calculations, and as the classical probabilistic benchmark against which Bayesian forecasts are compared.
How is it used?
ETSModel(y, error="add", trend="add", seasonal="add", seasonal_periods=7).fit(), then get_prediction(...).summary_frame(alpha=0.2) for an 80% interval. Check its coverage on a holdout (Chapter 7.16) before trusting it.
"An ETS 95% interval accounts for all the uncertainty."
It accounts for future shocks if the model is correct, with α and σ treated as known. It ignores uncertainty in the estimated parameters and the chance that the model is the wrong one, so real coverage is often a bit lower. Check it on a holdout.
"Wider intervals mean a worse model."
Intervals should be as narrow as possible while still covering the truth at the stated rate (calibration first, then sharpness, Chapter 7.16). A model with narrow intervals that miss half the time is worse.
ETS(A,N,N): $y_t = \ell_{t-1} + \varepsilon_t$, $\ell_t = \ell_{t-1} + \alpha\varepsilon_t$; forecast variance $\sigma^2[1 + \alpha^2(h-1)]$.
ETS(Error, Trend, Season): SES = (A,N,N), Holt = (A,A,N), additive HW = (A,A,A); fit by maximum likelihood, compare by AIC.
Trap: intervals ignore parameter and model uncertainty; check coverage.
Quick check: ETS(A,N,N) with α = 1. What model is this, and how does the forecast sd grow?
With α = 1 the level jumps all the way to each new value: $\ell_t = y_t$, so the forecast is naive and the model is a random walk. The variance is $\sigma^2[1 + (h-1)] = h\sigma^2$, so the sd grows like $\sigma\sqrt h$: 4 steps ahead is twice as uncertain as 1 step ahead.
Exponential smoothing versus your Prophet-style model core
The two approaches have different philosophies. Exponential smoothing is local: "the recent past is the best guide; keep re-estimating where we are and how fast we move." A Prophet-style model is global: "one recipe (trend + seasonality + holidays + regressors) explains the whole history; changepoints let the trend bend." Think of a car's GPS that keeps recalculating from your current position (local) versus a route planned in advance on a map that knows where every bridge and roadwork is (global).
Neither is better everywhere. Local methods react quickly to things nobody planned for (a sudden level shift, a slow drift). Global methods use known structure that local methods cannot see (a holiday next Tuesday, a price change, two seasonalities at once) and give a decomposition you can explain.
Three ways to say it:
- Picture: ETS hugs the recent data; the Prophet-style model draws one structured curve through all of it.
- Numbers: after a +15 level shift, ETS with α = 0.3 has closed all but $15\times0.7^7 \approx 1.2$ of the gap a week later; a global line without a changepoint there stays off.
- Slogan: local adapts to surprises; global uses what you already know.
Two situations, same shop.
- Level shift. Two weeks ago a competitor closed and daily orders jumped by 15. Exponential smoothing with $\alpha = 0.3$ moves its level by 30% of each surprise. After one day the remaining gap is $15\times0.7 = 10.5$, after two days $15\times0.7^2 = 7.35$, after a week $15\times0.7^7 = 15\times0.082 \approx 1.24$. By the forecast origin it has fully adapted.
- A global linear trend fitted to 14 weeks, without a changepoint near the jump, can only tilt: it ends up too low at the end of the history and too low in the forecast. (A Prophet-style model fixes this only if a changepoint is allowed near the jump, Chapter 7.8, and even then a slope change is not the same as a step.)
- Known holiday. A public holiday next Friday has added about +30 orders on each of the last three occurrences. Holt–Winters knows only weekday effects, so its Friday forecast misses by about 30 that day.
- A regression with a 0/1 holiday column estimates the holiday effect from the three past holidays ($\approx 30$) and adds it on the right future date.
| Exponential smoothing (ETS, Holt–Winters) | Prophet-style Bayesian model | |
|---|---|---|
| Core idea | local: states (level, trend, season) updated after every observation | global: $y_t = g(t) + s(t) + h(t) + X_t\beta + \varepsilon_t$ fitted to all history |
| Reacts to unplanned shifts | quickly, automatically (speed set by α, β, γ) | only through changepoints (where allowed) or not at all |
| Seasonality | one season length (classic form) | several at once (weekly + yearly Fourier terms) |
| Holidays, regressors | none in standard ETS | indicator and regressor columns with priors |
| Data requirements | regular, gap-free grid; ≥ 2 seasons for HW | regression on time: gaps and irregular dates are natural |
| Noise model | Gaussian (additive or multiplicative error) | any likelihood: Normal, Student-t, Negative Binomial |
| Uncertainty | future shocks, with parameters fixed at their estimates | posterior predictive: future noise and parameter uncertainty (as well as the inference method captures it) |
| Cost and robustness | milliseconds per series, few parameters, hard to break | more parameters, priors and inference to check (SVI/NUTS) |
| Explanation | states are hard to tell as a business story | components (trend, weekly, holiday, price effect) are easy to plot and explain |
Why do we need it?
"Why not just use Holt–Winters?" is one of the first questions about a custom forecasting model. You need to know what each approach assumes, so you can say when yours is the right tool and prove it with a fair comparison.
Where is it used?
Model-selection discussions in demand-planning teams, design reviews of forecasting systems, academic comparisons (ETS is a standard benchmark in the M-competitions and in papers on neural and Bayesian forecasters), and in interviews about your forecasting project.
How is it used?
Fit ETS (or Holt–Winters) and your model on the same rolling origins; compare point error, interval coverage and CRPS; look at where each one wins (around holidays? after shifts? at long horizons?). Often the answer is "keep both" or "combine them".
"ETS assumes the series is stationary, like ARIMA."
ETS models with a level or trend state are non-stationary by design (the level itself wanders like a random walk). They need no differencing; the smoothing handles the moving level (compare Chapter 7.4).
"The Bayesian model must be better: it gives full posterior uncertainty."
Uncertainty is only useful if it is calibrated and the point forecasts are good. A misspecified Bayesian model can be confidently wrong. Compare coverage and CRPS on rolling origins (Chapter 7.16), not philosophies.
"Holt–Winters can use our promotion calendar."
The standard ETS implementations (such as statsmodels' ExponentialSmoothing and ETSModel) take no regressors. If known future events matter, you need a regression-type model (a Prophet-style model, or ARIMA with regressors, Chapter 7.6).
Why choose exponential smoothing instead of your Prophet-style model, and vice versa?
- Choose ETS / Holt–Winters when: the series has one dominant seasonality and no important known events; levels shift in ways nobody can schedule; you need thousands of forecasts quickly and robustly (few parameters, no priors or inference diagnostics to babysit); or a fast, strong benchmark for your model.
- Choose the Prophet-style model when: there are several seasonalities, holidays that move around the calendar, or regressors (price, promotions) whose future values are known; you need a count or heavy-tailed likelihood (Negative Binomial, Student-t); the history has gaps or irregular dates; stakeholders need an explainable decomposition; or you want predictive distributions that include parameter uncertainty.
- Honest middle ground: in your model the trend adapts only through its changepoints (grid + PELT, with Laplace priors on the slope changes), so a sudden level shift after the last allowed changepoint is exactly where Holt–Winters can beat it. Showing that you know this is a strong interview answer.
"We used a Bayesian Prophet-style model because it is more advanced than Holt–Winters."
"We needed weekly and yearly seasonality, moving holidays, external regressors and a Negative Binomial likelihood for counts, which standard ETS cannot express. We would still benchmark against seasonal naive and ETS on rolling origins."
Model answer: "Exponential smoothing is local: it re-estimates level, trend and season after every observation, so it adapts quickly to unplanned shifts and is cheap and robust. A Prophet-style model is a global regression on time with structured components, priors and a flexible likelihood, so it can use known calendar information and give an explainable decomposition with full predictive distributions. I would pick ETS for many simple series or when shifts are unpredictable, and the structured model when calendar effects, regressors or count likelihoods matter, and I would show the choice with a rolling-origin comparison."
ETS = local (states updated every step; adapts to surprises; one season; no regressors; Gaussian errors; plug-in intervals).
Prophet-style = global regression (several seasonalities, holidays, regressors, any likelihood, posterior predictive; adapts to shifts only via changepoints).
Trap: "more advanced" is not an argument; a rolling-origin comparison is.
Quick check: you must forecast daily demand for 20 000 products, most with 6 months of history, no promotion data and frequent unexplained jumps. Which family, and why?
Exponential smoothing (ETS or Holt–Winters with a weekly season), or even seasonal naive. With no known events or regressors, short histories and unplanned jumps, a local method adapts automatically and is cheap and robust at that scale. A Prophet-style model would have little extra information to use and many more things to check per series.
Recap, cheat sheet and practice
- Baselines: naive ($y_T$), seasonal naive (last cycle), mean ($\bar y$), drift ($y_T + h\frac{y_T - y_1}{T-1}$), moving average (last $k$). No parameters (except $k$); each is optimal for a simple process.
- Strong baseline: the best simple method for your series, scored on the same time-ordered holdout; report skill $= 1 - E_{\text{model}}/E_{\text{baseline}}$ over many origins.
- SES: $\ell_t = \ell_{t-1} + \alpha(y_t - \ell_{t-1})$; geometric weights $\alpha(1-\alpha)^j$; flat forecast; α = 1 is naive; α chosen by one-step SSE.
- Holt: adds a local slope; forecast $\ell_T + hb_T$. Damped: $\ell_T + (\phi + \dots + \phi^h)b_T$, levels off at $\ell_T + \frac{\phi}{1-\phi}b_T$.
- Holt–Winters: adds one seasonal effect per position in the cycle (additive: units; multiplicative: factors). One season length.
- ETS: the smoothing recipes as state-space models with a noise term: likelihood, AIC, intervals that widen with $h$ (SES: $\sigma^2[1 + \alpha^2(h-1)]$), conditional on fitted parameters.
- vs Prophet-style: local and adaptive vs global and structured; choose by the data's structure and prove it with rolling-origin comparisons.
Cheat sheet
| Method | Forecast $\hat y_{T+h\mid T}$ | Best when / remember |
|---|---|---|
| Naive | $y_T$ | random walk; = SES with α = 1 |
| Seasonal naive | $y_{T+h-m(k+1)}$, $k = \lfloor (h-1)/m\rfloor$ | strong seasonal pattern; the default bar for daily data ($m = 7$) |
| Mean | $\bar y$ | constant level + noise |
| Drift | $y_T + h\,(y_T - y_1)/(T-1)$ | random walk with drift; uses only first and last values |
| Moving average | $\frac1k\sum_{j=0}^{k-1}y_{T-j}$ | lag $b(k+1)/2$ on a trend; ≠ MA($q$) model |
| SES | $\ell_T$; $\ell_t = \alpha y_t + (1-\alpha)\ell_{t-1}$ | wandering level; large α = little smoothing |
| Holt (damped) | $\ell_T + (\phi + \dots + \phi^h)b_T$ | local trend; damp for long horizons |
| Holt–Winters | $\ell_T + hb_T + s$ (or $\times s$) | one seasonality; multiplicative if swing ∝ level |
| ETS(A,N,N) interval | $\ell_T \pm z\,\sigma\sqrt{1 + \alpha^2(h-1)}$ | widens with $h$; ignores parameter uncertainty |
| Skill | $1 - E_{\text{model}}/E_{\text{baseline}}$ | > 0 beats the baseline; use many origins |
import numpy as np
import pandas as pd
from statsmodels.tsa.holtwinters import SimpleExpSmoothing, Holt, ExponentialSmoothing
from statsmodels.tsa.exponential_smoothing.ets import ETSModel
# 0) The SES worked example of this chapter, by library (alpha = 0.5, starting level 20)
y4 = np.array([20, 24, 22, 26.0])
ses_hand = SimpleExpSmoothing(y4, initialization_method="known", initial_level=20).fit(
smoothing_level=0.5, optimized=False)
print(ses_hand.fittedvalues, ses_hand.forecast(3)) # [20. 20. 22. 22.] [24. 24. 24.]
# 1) Simulated daily demand: wandering level + small upward drift + weekly pattern + noise (20 weeks)
rng = np.random.default_rng(0)
n = 140
t = np.arange(n)
weekly = np.array([0, 2, 1, 3, 10, 20, 15]) # Mon..Sun bumps
level = 50 + 0.15 * t + np.cumsum(rng.normal(0, 1.0, n))
y = pd.Series(level + weekly[t % 7] + rng.normal(0, 2, n),
index=pd.date_range("2024-01-01", periods=n, freq="D"))
train, test = y[:-28], y[-28:] # hold out the LAST 4 weeks (never shuffle)
h, T = len(test), len(train)
# 2) The five baselines
fc = {
"mean": np.repeat(train.mean(), h),
"naive": np.repeat(train.iloc[-1], h),
"seasonal naive": np.tile(train.iloc[-7:].to_numpy(), h // 7),
"drift": train.iloc[-1] + (train.iloc[-1] - train.iloc[0]) / (T - 1) * np.arange(1, h + 1),
"moving avg k=7": np.repeat(train.iloc[-7:].mean(), h),
}
mae = lambda f: np.mean(np.abs(test.to_numpy() - np.asarray(f)))
for name, f in fc.items():
print(f"{name:15s} MAE = {mae(f):5.2f}")
# mean 15.27 · naive 8.01 · seasonal naive 4.61 · drift 9.44 · moving avg k=7 7.12
# 3) Exponential smoothing (parameters chosen by minimizing the in-sample one-step squared errors)
ses = SimpleExpSmoothing(train, initialization_method="estimated").fit()
holt = Holt(train, initialization_method="estimated").fit()
hw = ExponentialSmoothing(train, trend="add", seasonal="add", seasonal_periods=7,
initialization_method="estimated").fit()
for name, m in [("SES", ses), ("Holt", holt), ("Holt-Winters", hw)]:
p = m.params
print(f"{name:12s} alpha={p['smoothing_level']:.2f} beta={p['smoothing_trend']:.2f} "
f"gamma={p['smoothing_seasonal']:.2f} MAE = {mae(m.forecast(h)):5.2f}")
# SES alpha=1.00 beta=nan gamma=nan MAE = 8.01 (alpha = 1: SES became naive)
# Holt alpha=0.00 beta=0.00 gamma=nan MAE = 8.21 (a fixed line: it cannot see the week)
# Holt-Winters alpha=0.40 beta=0.00 gamma=0.00 MAE = 2.35 (the only one that beats seasonal naive)
# 4) Holt-Winters as a statistical model, ETS(A,A,A): prediction intervals that widen with h
ets = ETSModel(train, error="add", trend="add", seasonal="add", seasonal_periods=7).fit(disp=False)
pi = ets.get_prediction(start=test.index[0], end=test.index[-1]).summary_frame(alpha=0.2) # 80%
inside = (test >= pi["pi_lower"]) & (test <= pi["pi_upper"])
print("80% interval coverage on the holdout:", round(inside.mean(), 3)) # 0.964 (27 of 28 days)
width = pi["pi_upper"] - pi["pi_lower"]
print("interval width, day 1 vs day 28:", round(width.iloc[0], 1), round(width.iloc[-1], 1)) # 6.6 15.2
Output checked with statsmodels 0.15.0. The skill of Holt–Winters versus seasonal naive here is $1 - 2.35/4.61 \approx 0.49$. Coverage on one 28-day holdout is a noisy number; judge calibration over many origins (Chapter 7.16).
1. A series behaves like a random walk (no trend, no season). Which baseline is the best possible forecast?
2. SES with α = 0.2 has level 40. Today's value is 50. What is the forecast for next week?
3. Which statement about the SES parameter α is correct?
4. Weekly sales have a Saturday peak that was +20 when the level was 100 and is +60 now that the level is 300. Which method fits this best?
5. ETS(A,N,N) with α = 0.5 and σ = 2. What is the forecast standard deviation 5 steps ahead?
6. In which situation does exponential smoothing have the clearest advantage over a Prophet-style model?
Practice problems
A. Quarterly data $y = 10, 20, 30, 15, 12, 22, 33, 17$ ($T = 8$, $m = 4$). Give the naive, seasonal naive, mean and drift forecasts for the next four quarters.
- Naive: $17, 17, 17, 17$.
- Seasonal naive: copy the last year: $12, 22, 33, 17$.
- Mean: $(10 + 20 + 30 + 15 + 12 + 22 + 33 + 17)/8 = 159/8 = 19.875$ for all four.
- Drift: slope $(17 - 10)/(8 - 1) = 1$; forecasts $18, 19, 20, 21$.
- The data clearly repeat a yearly shape (low, higher, peak, drop), so seasonal naive is the baseline to beat.
B. SES with α = 0.3 and starting level 100. The data are 110, 95, 105. Compute the levels and the forecast.
- $\ell_1 = 100 + 0.3\times(110 - 100) = 103$.
- $\ell_2 = 103 + 0.3\times(95 - 103) = 103 - 2.4 = 100.6$.
- $\ell_3 = 100.6 + 0.3\times(105 - 100.6) = 100.6 + 1.32 = 101.92$.
- Forecast for every future step: 101.92. The one-step errors were $+10, -8, +4.4$.
C. Damped Holt with level 200, trend 10 and φ = 0.9. Compute the forecasts for h = 1, 2 and the long-run limit. What would undamped Holt say for h = 20?
- $h = 1$: $200 + 0.9\times10 = 209$. $h = 2$: $200 + (0.9 + 0.81)\times10 = 217.1$.
- Limit: $200 + \frac{0.9}{0.1}\times10 = 290$.
- Undamped, $h = 20$: $200 + 20\times10 = 400$, and it keeps rising forever. The damped forecast at $h = 20$ is $200 + 10\times\sum_{j=1}^{20}0.9^j = 200 + 10\times9(1 - 0.9^{20}) \approx 200 + 10\times7.91 = 279.1$.
D. (Interview) "Your Prophet-style model loses to seasonal naive on the holdout. What do you do?"
"First I check the evaluation: is it one window or many rolling origins, is the holdout really after the training data, and did any preprocessing (scaling, PELT changepoints, feature selection) see the holdout? If the loss is real, I look at where it loses: right after a level shift (changepoints not allowed late enough, or Laplace scale too small), around holidays (missing or mis-dated holiday columns), at long horizons (trend extrapolation), or everywhere (wrong seasonality or likelihood). Residual diagnostics tell me which component is wrong. If the series simply behaves like a seasonal random walk, I accept that and ship the baseline or an ETS model; complexity has to earn its place."
E. (Interview) "Explain ETS(A,A,A). Why do its prediction intervals widen, and what do they leave out?"
"It's additive Holt–Winters written as a state-space model: level, trend and seasonal states, each updated by a fraction of the one-step error, with additive Gaussian errors. Because every future shock also moves the states, errors accumulate with the horizon, so the forecast variance grows with $h$; for SES it is $\sigma^2[1 + \alpha^2(h-1)]$. The standard intervals treat the estimated parameters as exact and assume the model is correct, so they miss parameter uncertainty and model misspecification; I'd check their coverage on rolling origins."
F. For each case pick a starting method and say why: (1) hourly server load with a daily cycle, 2 years of clean data; (2) weekly sales of a new product, 10 weeks of history; (3) daily store orders with weekly and yearly patterns, public holidays and a price promotion calendar.
- (1) Holt–Winters / ETS with $m = 24$ (one strong seasonality, regular grid, lots of data), benchmarked against seasonal naive. If a weekly cycle also matters, a model with two seasonalities (Fourier terms) is needed.
- (2) Naive, drift or damped Holt: 10 weeks cannot support a seasonal model, and a damped trend avoids extrapolating early growth forever.
- (3) A Prophet-style model: two seasonalities, holidays and a known future regressor (promotions) are exactly what it represents and ETS cannot. Still compare it with seasonal naive and Holt–Winters on rolling origins.
The ARIMA family and regression with time features
Exponential smoothing describes a series by its level, trend and season. The ARIMA family describes it by its memory: how today depends on the last few values (AR) and on the last few random shocks (MA), after differencing away trends (I) and seasons (S). Then we turn to the third classical option, a plain regression on time features (trend, weekday dummies, Fourier terms), which is the skeleton of your Prophet-style model. You do not need to build all of these from scratch, but you do need to know what each assumes, how it forecasts, and why someone might choose it instead of your Bayesian model.
- Read and simulate AR($p$) and MA($q$) models; know when an AR model is stationary and why MA effects have a finite memory
- Combine them into ARMA, and recognise the classic ACF/PACF signatures (AR: PACF cuts off; MA: ACF cuts off; ARMA: both tail off)
- Add differencing to get ARIMA($p,d,q$) and seasonal terms to get SARIMA; see naive, drift, SES and seasonal naive as special cases
- Predict how ARIMA forecasts behave: mean reversion and levelling intervals for stationary models, flat forecasts and $\sqrt h$-growing intervals for a random walk
- Fit and choose orders with the Box–Jenkins loop, AIC and residual checks (Ljung–Box)
- Build a regression with time features and fix its autocorrelated errors with ARIMA errors
- Explain why someone might choose ARIMA/ETS or a time-feature regression instead of your Prophet-style model, and vice versa
What we need from earlier chapters: autocorrelation, the ACF and PACF, white noise and random walks (Chapter 7.3); stationarity, differencing and unit roots (Chapter 7.4); baselines and exponential smoothing (Chapter 7.5); linear regression and design matrices (Chapter 5.13); maximum likelihood (Chapter 5.2). Notation: $\varepsilon_t$ ("epsilon t") is a shock (also called an innovation): an unpredictable random kick at time $t$, independent across time with mean 0 and variance $\sigma^2$ (white noise). $\phi_i$ ("phi") are autoregressive coefficients, $\theta_j$ ("theta") moving-average coefficients, $c$ a constant, $\mu$ the mean of a stationary series. $\Delta y_t = y_t - y_{t-1}$ is the first difference; $\Delta_m y_t = y_t - y_{t-m}$ the seasonal difference. $\rho_k$ is the autocorrelation at lag $k$.
Autoregressive models AR($p$): today leans on yesterday core
Think of a bathtub's water temperature. If it is hot now, it will still be warm in a minute, but it drifts back towards room temperature. Today's value is a fraction of yesterday's distance from normal, plus a new random kick. That is an autoregressive model: "auto" (self) + "regressive" (regression): the series is regressed on its own past values.
The key number is $\phi$ (phi), how much of yesterday's deviation survives until today. With $\phi = 0.5$, half of any surprise is left tomorrow, a quarter the day after, and so on: the series is pulled back to its mean (mean reversion). With $\phi = 1$ nothing is ever forgotten: that is a random walk, which never returns. With $\phi \gt 1$ deviations grow and the series explodes.
Three ways to say it:
- Picture: a ball on a spring tied to the mean: kicked away by shocks, pulled back each step.
- Numbers: mean 20, $\phi = 0.5$, today 28 → forecasts 24, 22, 21, 20.5: the gap of 8 halves every step.
- Slogan: AR = a regression of the series on its own lags.
AR(1) $y_t = 10 + 0.5\,y_{t-1} + \varepsilon_t$. The last value is $y_T = 28$.
- The long-run mean solves $\mu = 10 + 0.5\mu$, so $\mu = 10/(1 - 0.5) = 20$.
- Forecast 1 step ahead (future shocks have mean 0): $10 + 0.5\times28 = 24$.
- 2 steps: $10 + 0.5\times24 = 22$. 3 steps: $10 + 0.5\times22 = 21$. 4 steps: $20.5$.
- Same numbers from the gap to the mean: $\hat y_{T+h} = \mu + \phi^h(y_T - \mu) = 20 + 0.5^h\times8$: $24, 22, 21, 20.5, \dots \to 20$.
- Autocorrelation: $\rho_k = \phi^k$, so $\rho_1 = 0.5$, $\rho_2 = 0.25$, $\rho_3 = 0.125$: it decays geometrically.
An autoregressive model of order $p$, AR($p$):
$$y_t = c + \phi_1 y_{t-1} + \phi_2 y_{t-2} + \dots + \phi_p y_{t-p} + \varepsilon_t,\qquad \varepsilon_t \sim \text{white noise}(0, \sigma^2).$$- AR(1) is stationary (mean, variance and autocorrelations do not change over time, Chapter 7.4) exactly when $\lvert\phi_1\rvert \lt 1$. Then $\mu = \frac{c}{1-\phi_1}$, $Var(y_t) = \frac{\sigma^2}{1-\phi_1^2}$ and $\rho_k = \phi_1^k$. A negative $\phi_1$ makes the series zig-zag (the ACF alternates in sign).
- AR(2) is stationary exactly when $\phi_1 + \phi_2 \lt 1$, $\phi_2 - \phi_1 \lt 1$ and $\lvert\phi_2\rvert \lt 1$ (a triangle in the $(\phi_1, \phi_2)$ plane). If also $\phi_1^2 + 4\phi_2 \lt 0$, the series shows damped pseudo-cycles with period $2\pi/\arccos\!\big(\phi_1/(2\sqrt{-\phi_2})\big)$.
- In general, AR($p$) is stationary when all roots of $1 - \phi_1 z - \dots - \phi_p z^p = 0$ lie outside the unit circle ($\lvert z\rvert \gt 1$). $\phi = 1$ in AR(1) is a unit root: a random walk.
- Signature: the ACF tails off (geometric decay or damped waves); the PACF cuts off after lag $p$.
Why do we need it?
Many series have momentum: a busy day is usually followed by another busy day. An AR model turns "recent values carry information" into a forecast that fades back to normal at the right speed.
Where is it used?
Short-term demand and load forecasting, macroeconomic series (inflation, unemployment), sensor data, the error term of regressions with autocorrelated residuals, and as the "AR" part of ARIMA, SARIMA and state-space models.
How is it used?
Look for a PACF that cuts off after lag $p$; fit with ARIMA(y, order=(p, 0, 0)).fit(); check that the estimated $\phi$'s are inside the stationary region, the residuals look like white noise, and use forecast(h).
"An AR(1) with $\phi = 0.9$ is almost a random walk, so it forecasts like one."
Close in the short run, very different in the long run. The AR(1) forecast always returns to the mean and its interval stops growing; the random walk forecast stays flat at the last value and its interval grows forever (below).
"The constant $c$ is the mean of the series."
The mean is $\mu = c/(1 - \phi_1 - \dots - \phi_p)$. With $c = 10$ and $\phi = 0.5$ the mean is 20. (Libraries differ: statsmodels' ARIMA reports the mean μ as const (≈ 20 here), while SARIMAX(..., trend="c") reports the intercept $c$ (≈ 10). Check which one your tool prints.)
AR($p$): $y_t = c + \sum_{i=1}^p\phi_i y_{t-i} + \varepsilon_t$. AR(1): stationary iff $\lvert\phi\rvert \lt 1$; $\mu = c/(1-\phi)$; $\rho_k = \phi^k$; forecast $\mu + \phi^h(y_T - \mu)$.
AR(2) stationary triangle: $\phi_1 + \phi_2 \lt 1$, $\phi_2 - \phi_1 \lt 1$, $\lvert\phi_2\rvert \lt 1$; cycles if $\phi_1^2 + 4\phi_2 \lt 0$.
Trap: $c$ is not the mean; $\phi = 1$ is a random walk, not "a strong AR".
Quick check: $y_t = 6 + 0.8\,y_{t-1} + \varepsilon_t$ with $\sigma = 3$. What are the mean, the standard deviation of $y_t$, and the 2-step forecast from $y_T = 40$?
Mean $= 6/(1 - 0.8) = 30$. Variance $= 9/(1 - 0.64) = 25$, so sd $= 5$. Forecast: $30 + 0.8^2\times(40 - 30) = 30 + 6.4 = 36.4$ (check: $6 + 0.8\times40 = 38$, then $6 + 0.8\times38 = 36.4$).
Moving-average models MA($q$): shocks with a short memory
A newspaper misprint causes a flood of complaint calls today and some follow-up calls tomorrow; after that it is forgotten. The effect of the shock lasts exactly two days, then stops completely. A moving-average model describes this: today's value is the mean plus today's shock plus a fraction of the last few shocks.
The contrast with AR is about memory. In an AR model a shock feeds into tomorrow's value, which feeds into the day after, so its effect fades but never ends. In an MA($q$) model a shock affects only the next $q$ values and then disappears completely. That is why the ACF of an MA($q$) is exactly zero after lag $q$.
Three ways to say it:
- Picture: AR memory is an echo that fades away; MA memory is a footprint that is wiped clean after $q$ steps.
- Numbers: MA(1) with $\theta = 0.5$: a shock of $+6$ today adds $+3$ tomorrow and $0$ after that.
- Slogan: MA = a weighted sum of recent shocks, not of recent values.
MA(1) $y_t = 20 + \varepsilon_t + 0.5\,\varepsilon_{t-1}$ with $\sigma = 1$. Today's shock was $\varepsilon_T = 6$.
- Forecast for tomorrow: $20 + E[\varepsilon_{T+1}] + 0.5\times6 = 20 + 0 + 3 = 23$.
- Two days ahead: $20 + E[\varepsilon_{T+2}] + 0.5\,E[\varepsilon_{T+1}] = 20$. From $h = 2$ on, the forecast is just the mean.
- Variance: $Var(y_t) = \sigma^2(1 + \theta^2) = 1 + 0.25 = 1.25$.
- Lag-1 autocorrelation: $\rho_1 = \frac{\theta}{1 + \theta^2} = \frac{0.5}{1.25} = 0.4$. Lag 2 and beyond: $\rho_k = 0$, because $y_t$ and $y_{t-2}$ share no shocks.
A moving-average model of order $q$, MA($q$):
$$y_t = \mu + \varepsilon_t + \theta_1\varepsilon_{t-1} + \dots + \theta_q\varepsilon_{t-q}.$$- Always stationary (a finite sum of white-noise terms). $E[y_t] = \mu$, $Var(y_t) = \sigma^2(1 + \theta_1^2 + \dots + \theta_q^2)$.
- MA(1): $\rho_1 = \theta/(1+\theta^2)$ (never larger than 0.5 in size), $\rho_k = 0$ for $k \ge 2$. MA(2): $\rho_1 = \frac{\theta_1 + \theta_1\theta_2}{1 + \theta_1^2 + \theta_2^2}$, $\rho_2 = \frac{\theta_2}{1 + \theta_1^2 + \theta_2^2}$, zero after.
- Invertibility: $\theta$ and $1/\theta$ give the same ACF, so we pick the version with the roots of $1 + \theta_1 z + \dots + \theta_q z^q$ outside the unit circle (MA(1): $\lvert\theta\rvert \lt 1$). Then the shocks can be recovered from past data, which is what forecasting needs.
- Signature: the ACF cuts off after lag $q$; the PACF tails off.
- Sign conventions differ:
statsmodelsand this guide write $+\theta\varepsilon_{t-1}$; Box and Jenkins' original books write $-\theta\varepsilon_{t-1}$. Always check the sign before interpreting a fitted θ.
Why do we need it?
Some effects last a fixed, short time: a one-day promotion with a next-day hangover, a measurement error that is corrected the next day. AR models describe these poorly; a short MA describes them with one or two numbers.
Where is it used?
The MA part of ARIMA and SARIMA (the "airline model" SARIMA(0,1,1)(0,1,1)12 is pure MA after differencing), over-differenced series (differencing white noise gives MA(1) with θ = −1), and SES, which is ARIMA(0,1,1).
How is it used?
Look for an ACF that drops to (near) zero after lag $q$; fit ARIMA(y, order=(0, 0, q)); read θ with the sign convention in mind; forecasts return to the mean after $q$ steps.
"MA($q$) means averaging the last $q$ observations."
It is a weighted sum of the last $q$ unobserved shocks. The rolling mean of observations is a smoothing rule from Chapter 7.5; the two are unrelated apart from the name.
"A fitted MA coefficient of −0.98 is a strong, healthy effect."
An MA coefficient near −1 (a near-unit root in the MA part) is a classic sign of over-differencing: you differenced a series that did not need it. Try one difference less.
MA($q$): $y_t = \mu + \varepsilon_t + \sum_{j=1}^q\theta_j\varepsilon_{t-j}$; always stationary; $Var = \sigma^2(1 + \sum\theta_j^2)$.
MA(1): $\rho_1 = \theta/(1+\theta^2)$, $\rho_k = 0$ for $k\ge2$; forecasts equal the mean after $q$ steps.
Trap: shocks, not observations; check the sign convention; θ ≈ −1 means over-differenced.
Quick check: an MA(1) has $\rho_1 = 0.4$. What are the two possible values of θ, and which one do we report?
Solve $\theta/(1 + \theta^2) = 0.4$: $0.4\theta^2 - \theta + 0.4 = 0$, so $\theta = (1 \pm \sqrt{1 - 0.64})/0.8 = (1 \pm 0.6)/0.8$, i.e. $\theta = 0.5$ or $\theta = 2$. They give the same ACF; we report the invertible one, $\theta = 0.5$ ($\lvert\theta\rvert \lt 1$).
ARMA($p,q$) and the ACF/PACF signatures core
Real series often have both kinds of memory: a slow fade (AR) and a short footprint (MA). ARMA($p,q$) simply adds them: $p$ lags of the series and $q$ lags of the shocks. Mixing the two lets a model with very few numbers describe a rich autocorrelation pattern.
How do you guess $p$ and $q$ from data? With two diagnostic plots you met in Chapter 7.3. The ACF shows the total correlation at each lag. The PACF shows the direct correlation at lag $k$ after removing what the lags in between already explain. Each model leaves a fingerprint: an AR($p$) has exactly $p$ direct links, so its PACF stops after $p$; an MA($q$) forgets shocks after $q$ steps, so its ACF stops after $q$; ARMA has both and neither plot stops cleanly.
Three ways to say it:
- Picture: two bar charts, and you look for the one that drops off a cliff.
- Numbers: PACF 0.71, 0.30, then ≈ 0 → AR(2). ACF 0.55, 0.26, then ≈ 0 → MA(2).
- Slogan: PACF cuts → AR; ACF cuts → MA; both tail → ARMA.
ARMA(1,1) $y_t = 0.7\,y_{t-1} + \varepsilon_t + 0.4\,\varepsilon_{t-1}$.
- Response to one unit shock (ψ-weights): $\psi_0 = 1$, $\psi_1 = \phi + \theta = 0.7 + 0.4 = 1.1$, then each next one is $0.7\times$ the previous: $\psi_2 = 0.77$, $\psi_3 = 0.539$, …
- Lag-1 autocorrelation: $\rho_1 = \frac{(1 + \phi\theta)(\phi + \theta)}{1 + 2\phi\theta + \theta^2} = \frac{1.28\times1.1}{1 + 0.56 + 0.16} = \frac{1.408}{1.72} \approx 0.819$.
- After that the AR part takes over: $\rho_2 = 0.7\times0.819 \approx 0.573$, $\rho_3 = 0.7\times0.573 \approx 0.401$: a geometric tail off.
- The PACF also tails off (0.819, −0.294, 0.116, −0.046, …), so neither plot cuts off: the signature of a mixed model.
ARMA($p,q$):
$$y_t = c + \sum_{i=1}^{p}\phi_i y_{t-i} + \varepsilon_t + \sum_{j=1}^{q}\theta_j\varepsilon_{t-j}.$$Stationary when the AR part is (roots of $1 - \phi_1 z - \dots - \phi_p z^p$ outside the unit circle); invertible when the MA part is. Signatures (theoretical; sample plots are noisy, so judge against the $\pm1.96/\sqrt n$ band):
| Model | ACF | PACF |
|---|---|---|
| white noise | all ≈ 0 | all ≈ 0 |
| AR($p$) | tails off (geometric or damped wave) | cuts off after lag $p$ |
| MA($q$) | cuts off after lag $q$ | tails off |
| ARMA($p,q$) | tails off (after lag $q$) | tails off (after lag $p$) |
| random walk / unit root | stays near 1, decays very slowly | one spike near 1 at lag 1 |
"Tails off" means it shrinks gradually; "cuts off" means it falls inside the band and stays there.
Why do we need it?
To choose a sensible starting model without trying every combination, and to check afterwards that the residuals have no signature left (a clean ACF means the memory has been captured).
Where is it used?
The identification step of the Box–Jenkins method, residual checks of any forecasting model (including your Prophet-style model), and textbook interview questions ("this PACF cuts off at 2, what model?").
How is it used?
Make the series stationary first (difference if needed). Plot plot_acf and plot_pacf from statsmodels.graphics.tsaplots. Read off candidate $(p, q)$, then let AIC and residual diagnostics decide between a few nearby candidates.
"The PACF has a spike at lag 2, so it is an AR(2), done."
Signatures suggest candidates. Sample plots are noisy (one bar in 20 crosses the band by chance), and mixed models are hard to read. Fit two or three nearby candidates and let AIC and residual checks decide.
"I can read the signatures from the raw series."
Only for a stationary series. A trend or unit root makes the ACF decay very slowly and hides everything else; difference (or detrend) first (Chapter 7.4).
ARMA($p,q$) = AR($p$) + MA($q$): $y_t = c + \sum\phi_i y_{t-i} + \varepsilon_t + \sum\theta_j\varepsilon_{t-j}$.
PACF cuts after $p$ → AR($p$); ACF cuts after $q$ → MA($q$); both tail → ARMA; ACF stuck near 1 → difference.
Trap: signatures give candidates, not answers; read them on a stationary series.
Quick check: after one difference, the ACF has one big negative bar at lag 1 (about −0.5) and nothing else; the PACF decays slowly. What model, and what does it suggest?
ACF cutting off after lag 1 with tailing PACF is an MA(1). A lag-1 value near −0.5 means θ near −1, the fingerprint of over-differencing: the original series was probably already stationary (differencing white noise gives exactly $\rho_1 = -0.5$). Try $d = 0$.
ARIMA($p,d,q$): difference, model, add back up core
ARMA models need a stationary series: one that keeps returning to a fixed mean. Sales that grow for years do not. The trick, from Chapter 7.4: look at the changes instead of the levels. "How much did sales change since yesterday?" is often stationary even when sales are not. Model the changes with ARMA, forecast the future changes, then add them back up onto the last value to get forecasts of the level.
That "add back up" step is called integration (the opposite of differencing), and it is the "I" in ARIMA. $d$ counts how many times we differenced: usually 0 or 1, occasionally 2.
Three ways to say it:
- Picture: flatten the hill by looking at its slopes, model the slopes, then rebuild the hill from the predicted slopes.
- Numbers: values 100, 103, 105, 110, 112 → changes 3, 2, 5, 2 → average change 3 → forecast 115, 118, 121.
- Slogan: AR + I + MA: differencing handles the trend, ARMA handles the memory.
Data $y = 100, 103, 105, 110, 112$. Differences $\Delta y = 3, 2, 5, 2$.
- ARIMA(0,1,0) (random walk): future changes have mean 0, so the forecast is $112$ for every step. This is the naive forecast.
- ARIMA(0,1,0) with a constant (random walk with drift): future changes average $\bar{\Delta y} = (3 + 2 + 5 + 2)/4 = 3$, so the forecasts are $115, 118, 121$. This is the drift forecast, since $(112 - 100)/4 = 3$.
- ARIMA(1,1,0) with $\phi = 0.5$, no constant: the changes follow AR(1). The last change was 2, so the forecast changes are $0.5\times2 = 1$, then $0.5$, then $0.25$.
- Add them back up: $112 + 1 = 113$, $113 + 0.5 = 113.5$, $113.5 + 0.25 = 113.75$. The forecast bends and levels off near $112 + 2 = 114$.
Write $B$ for the backshift operator: $By_t = y_{t-1}$, so $(1 - B)y_t = \Delta y_t$. A series follows ARIMA($p,d,q$) if its $d$-th difference follows a stationary, invertible ARMA($p,q$):
$$\underbrace{(1 - \phi_1 B - \dots - \phi_p B^p)}_{\text{AR}}\ \underbrace{(1 - B)^d}_{\text{I}}\ y_t = c + \underbrace{(1 + \theta_1 B + \dots + \theta_q B^q)}_{\text{MA}}\ \varepsilon_t .$$- $d = 1$ removes a stochastic trend (a random-walk-like level); a constant with $d = 1$ means a linear drift in the forecasts; with $d = 2$ a constant would mean a quadratic trend, which is rarely sensible.
- Choose $d$ from plots and unit-root tests (ADF, KPSS, Chapter 7.4); prefer the smallest $d$ that makes the series look stationary. Over-differencing leaves an MA coefficient near −1.
- Special cases: ARIMA(0,0,0) with a constant = mean forecast; ARIMA(0,1,0) = random walk (naive); ARIMA(0,1,0) + constant = drift; ARIMA(0,1,1) = SES with $\theta = \alpha - 1$; ARIMA(0,2,2) corresponds to Holt's linear method.
- In
statsmodels'ARIMA, a constant cannot be combined with $d \ge 1$; ask for drift withtrend="t"(a linear trend in the levels is a constant in the differences).
Why do we need it?
Most business series trend or wander, so plain ARMA does not apply. Differencing inside the model handles that, and the forecasts are automatically turned back into levels with correct, growing intervals.
Where is it used?
Economic and financial forecasting, demand planning, automatic forecasting tools (R's auto.arima, pmdarima.auto_arima, Nixtla's AutoARIMA), and as a classical benchmark next to ETS in forecasting papers.
How is it used?
Plot the series, pick $d$ (usually 0 or 1), read candidate $p, q$ from the ACF/PACF of the differenced series, fit ARIMA(y, order=(p, d, q)).fit(), compare AICs, check residuals, then get_forecast(h) for means and intervals.
"When in doubt, difference once more; it can't hurt."
Over-differencing makes the series noisier (white noise has variance $\sigma^2$, its difference $2\sigma^2$), creates a fake MA(1) with θ ≈ −1, and widens forecast intervals for no reason.
"I'll add a constant to my ARIMA(1,1,1) to give it a mean."
With $d = 1$ a constant is a drift: the forecasts will rise (or fall) by that amount every step, forever. Include it only if the series really has a persistent average growth.
"ARIMA can only be used on stationary data."
"The differenced series must be stationary. ARIMA itself, with $d \ge 1$, is a non-stationary model: differencing is built into it."
Model answer: "ARIMA models the $d$-th difference as a stationary ARMA and integrates back for forecasts. That is different from a Prophet-style model, which never differences: it writes the trend as an explicit function of time (piecewise linear with changepoints) and assumes the remaining noise is independent. Both handle trends, in different ways."
ARIMA($p,d,q$): ARMA($p,q$) on $\Delta^d y$, then integrate back. $d$ = 0 or 1 usually (2 rarely).
Special cases: (0,1,0) naive, (0,1,0)+c drift, (0,1,1) SES with θ = α − 1.
Trap: over-differencing (θ ≈ −1, higher variance); a constant with $d = 1$ is a drift.
Quick check: an ARIMA(0,1,1) fit reports θ = −0.7. Which SES α is this, and what does that say about the series?
$\alpha = 1 + \theta = 0.3$. The forecast is an exponentially weighted average with a fairly long memory: the series behaves like a slowly wandering level plus a lot of short-term noise, not like a pure random walk (which would be θ = 0, α = 1).
SARIMA: seasonal differences and seasonal terms
With a weekly pattern, this Monday is best compared with last Monday, not with yesterday (Sunday). SARIMA adds this idea to ARIMA in two ways. First, a seasonal difference: look at $y_t - y_{t-7}$, "how much higher than the same day last week", which removes a repeating pattern the way an ordinary difference removes a trend. Second, seasonal AR and MA terms that connect a value to the values and shocks one or more whole seasons back (lags 7, 14, …).
Three ways to say it:
- Picture: stack the weeks on top of each other and model how each weekday changes from week to week.
- Numbers: week 2 minus week 1 is +2 for every weekday → next Monday ≈ last Monday + 2.
- Slogan: ARIMA for the short memory, the seasonal part for the "same time last cycle" memory.
The shop data of Chapter 7.5: week 1 = 20, 22, 21, 23, 30, 40, 35; week 2 = 22, 24, 23, 25, 32, 42, 37.
- Seasonal differences $\Delta_7 y_t = y_t - y_{t-7}$ for week 2: $22 - 20 = 2$, $24 - 22 = 2$, $23 - 21 = 2$, … all equal to $2$.
- So the seasonally differenced series is a constant 2 (plus noise in real data): the model SARIMA(0,0,0)(0,1,0)7 with a constant, "same day last week + 2".
- Forecasts for week 3: Monday $22 + 2 = 24$, Tuesday $24 + 2 = 26$, Wednesday $23 + 2 = 25$.
- Actual values were 24, 25, 25: errors $0, -1, 0$, MAE $= 1/3 \approx 0.33$, better than seasonal naive's 1.67, because the model also learned the weekly growth of +2.
SARIMA($p,d,q$)($P,D,Q$)$_m$ adds seasonal polynomials in $B^m$:
$$\Phi(B^m)\,\phi(B)\,(1 - B)^d(1 - B^m)^D\,y_t = c + \Theta(B^m)\,\theta(B)\,\varepsilon_t ,$$- $\phi(B), \theta(B)$: the ordinary AR and MA polynomials (orders $p, q$). $\Phi(B^m) = 1 - \Phi_1B^m - \dots - \Phi_PB^{Pm}$ and $\Theta(B^m) = 1 + \Theta_1B^m + \dots + \Theta_QB^{Qm}$: seasonal AR and MA at lags $m, 2m, \dots$
- $(1 - B^m)^D$: $D$ seasonal differences, $\Delta_m y_t = y_t - y_{t-m}$ (usually $D \le 1$).
- Seasonal naive = SARIMA(0,0,0)(0,1,0)$_m$. The classic "airline model" is SARIMA(0,1,1)(0,1,1)$_{12}$ for monthly data.
- Seasonal signatures appear at multiples of $m$: e.g. a single ACF spike at lag $m$ suggests $Q = 1$; PACF spikes at $m, 2m$ suggest seasonal AR.
- One season length $m$, which must be an integer and not too large: $m = 7$ (daily, weekly) or $12$ (monthly) is fine; $m = 365$ (daily, yearly) is slow and fragile, and $52.18$ (weekly, yearly) is not an integer. Then Fourier terms as regressors are the usual fix (below, Chapter 7.11).
Why do we need it?
Seasonal series have strong correlations at lags $m, 2m, \dots$ that a short ARMA cannot reach without dozens of coefficients. Seasonal terms capture them with one or two numbers.
Where is it used?
Monthly airline passengers and retail sales (the airline model), daily traffic and orders with a weekly cycle, official statistics (X-13ARIMA-SEATS seasonal adjustment uses seasonal ARIMA models), and automatic seasonal ARIMA in forecasting libraries.
How is it used?
ARIMA(y, order=(p, d, q), seasonal_order=(P, D, Q, m)) or SARIMAX(...). Look at the ACF at lags $m, 2m$ after seasonal differencing; compare a few candidates by AIC; check the residual ACF at seasonal lags too.
"Every seasonal series should be seasonally differenced."
If the weekly pattern is fixed (deterministic: the same shape every week), seasonal differencing over-differences it, and the fit shows a seasonal MA coefficient near −1. In the Code-it block below, SARIMA(1,0,0)(0,1,1)7 fitted to a fixed weekly pattern returns ma.S.L7 = −0.999. Seasonal dummies or Fourier terms in a regression are the natural model for a fixed pattern; seasonal differencing suits a pattern that slowly changes.
"SARIMA handles weekly and yearly seasonality in daily data."
Only one $m$. Daily data with a yearly cycle would need $m = 365$ (impractical); the usual approach is SARIMA for the weekly cycle plus Fourier regressors for the yearly one, or a regression-based model such as your Prophet-style model.
SARIMA($p,d,q$)($P,D,Q$)$_m$: ordinary ARIMA × seasonal ARIMA in $B^m$; $\Delta_m y_t = y_t - y_{t-m}$.
Seasonal naive = (0,0,0)(0,1,0)$_m$; airline = (0,1,1)(0,1,1)$_{12}$.
Trap: one integer $m$ only; a fixed pattern + seasonal differencing → seasonal MA ≈ −1 (use dummies/Fourier instead).
Quick check: monthly data. After one ordinary and one seasonal difference, the ACF has a single significant negative spike at lag 12 and one at lag 1. Which model does this suggest?
An ACF cut-off at lag 1 suggests $q = 1$; a single spike at lag 12 suggests $Q = 1$. With $d = D = 1$ and $m = 12$ this is SARIMA(0,1,1)(0,1,1)$_{12}$, the airline model.
How ARIMA forecasts behave: mean reversion versus random walk core
Ask two models about next month. A stationary model (say AR(1)) says: "right now we are above normal, but we will drift back to the mean; and since the mean is stable, my uncertainty stops growing after a while." A random walk says: "my best guess is where we are now, but every day adds a new, permanent shock, so the further ahead you ask, the less I know, without limit."
This difference in the shape of the forecast and its interval is the most practical consequence of choosing $d = 0$ or $d = 1$. Long-range ARIMA forecasts are deliberately boring: they converge to the mean (stationary), to a flat line (random walk) or to a straight line (random walk with drift). All the interesting detail lives in the first few steps.
Three ways to say it:
- Picture: AR(1): a fan that opens and then stops opening; random walk: a fan that keeps opening like $\sqrt h$.
- Numbers: σ = 2: the AR(1) (φ = 0.5) forecast sd goes 2, 2.24, … and never exceeds 2.31; the random walk's goes 2, 2.83, 4 (h = 4), 8 (h = 16).
- Slogan: stationary forgets, integrated remembers.
Shock size σ = 2. The forecast error $h$ steps ahead is a sum of the future shocks, each multiplied by its ψ-weight: $Var = \sigma^2(\psi_0^2 + \psi_1^2 + \dots + \psi_{h-1}^2)$.
- AR(1), φ = 0.5: $\psi_j = 0.5^j$. $h = 1$: $4\times1 = 4$, sd $2$. $h = 2$: $4\times(1 + 0.25) = 5$, sd $\approx 2.24$.
- As $h \to \infty$: $4\times(1 + 0.25 + 0.0625 + \dots) = 4/(1 - 0.25) = 5.33$, sd $\approx 2.31$: the unconditional sd of the series. A 95% interval never gets wider than $\pm1.96\times2.31 = \pm4.53$.
- Random walk: $\psi_j = 1$ for all $j$. $Var = 4h$, sd $= 2\sqrt h$: $h = 1$: 2, $h = 4$: 4, $h = 16$: 8. The 95% interval at $h = 16$ is $\pm15.7$, and still growing.
- ARIMA(0,1,1) (= SES): $\psi_0 = 1$, $\psi_j = 1 + \theta = \alpha$ for $j \ge 1$, so $Var = \sigma^2[1 + (h-1)\alpha^2]$, the same formula as ETS(A,N,N) in Chapter 7.5.
Write the (possibly integrated) model in ψ-weight form $y_t = \mu_t + \sum_{j\ge0}\psi_j\varepsilon_{t-j}$ with $\psi_0 = 1$. Then the $h$-step forecast error and its variance are
$$y_{T+h} - \hat y_{T+h\mid T} = \sum_{j=0}^{h-1}\psi_j\,\varepsilon_{T+h-j},\qquad \sigma_h^2 = \sigma^2\sum_{j=0}^{h-1}\psi_j^2 ,$$and with Normal shocks a $100(1-a)\%$ prediction interval is $\hat y_{T+h\mid T} \pm z_{1-a/2}\,\sigma_h$.
- Stationary ($d = 0$): forecasts revert to the mean; $\sigma_h^2 \to Var(y_t)$, so the interval width levels off.
- $d = 1$: forecasts approach a flat line (no constant) or a straight line (with drift); $\sigma_h^2$ grows roughly linearly in $h$, so the width grows like $\sqrt h$.
- $d = 2$: forecasts follow a straight line set by the last local slope; the width grows even faster.
- These intervals treat the estimated parameters (and the chosen orders) as exactly right. They include future shocks only.
Why do we need it?
Choosing $d = 0$ versus $d = 1$ changes the long-range forecast completely: a return to normal or a permanent new level, a bounded or an ever-widening interval. You must know which story the model is telling before acting on it.
Where is it used?
Inventory and capacity planning (how wide must the safety margin be 8 weeks out?), finance (prices behave like random walks; spreads and rates often mean-revert), and every get_forecast(h).conf_int() call in statsmodels.
How is it used?
After fitting, plot the forecast with its interval for a long horizon. Ask: does reverting to the mean (or extrapolating the drift) make business sense? Check interval coverage on rolling origins before using the bands for decisions.
"The ARIMA forecast is flat (or a straight line) after a few steps, so the model is broken."
That is what these models are built to say: beyond their memory, the best guess is the mean, the last level, or the drift line. If you expect seasonality or known events in the long run, they must be in the model (seasonal terms, regressors).
"95% ARIMA intervals contain 95% of future values."
Only if the model and its estimated parameters are right. Parameter uncertainty, order selection and changing behaviour are not included, so real coverage is often lower, especially far ahead. Measure it on a holdout (Chapter 7.16).
In your Prophet-style model the noise $\varepsilon_t$ is independent from day to day, so the predictive band does not widen through accumulated shocks the way a random walk's does. Apart from the noise, the band widens with the horizon through parameter uncertainty from the posterior, mainly the trend (the slope and the slope changes), and through simulated future slope changes only if your forecast code does that, as Prophet does (check your code). If your residuals turn out to be strongly autocorrelated (Chapter 7.17), the short-horizon bands can be miscalibrated: the model treats today as unrelated to yesterday when it is not.
$h$-step error variance: $\sigma^2\sum_{j=0}^{h-1}\psi_j^2$. AR(1): $\sigma^2\frac{1-\phi^{2h}}{1-\phi^2} \to \frac{\sigma^2}{1-\phi^2}$; random walk: $h\sigma^2$.
Stationary → forecast reverts to the mean, width levels off; $d = 1$ → flat/drift line, width $\propto \sqrt h$.
Trap: intervals ignore parameter and model uncertainty.
Quick check: a random walk has σ = 3. How many steps ahead is the forecast sd twice as large as one step ahead? And for an AR(1) with φ = 0.6?
Random walk: sd $= 3\sqrt h$, which doubles at $h = 4$. AR(1) with φ = 0.6: the sd can never exceed $3/\sqrt{1 - 0.36} = 3/0.8 = 3.75$, only 1.25 times the 1-step sd, so it never doubles.
Fitting and choosing orders: Box–Jenkins, AIC and residual checks
Fitting an ARIMA model is a loop, a bit like a doctor's visit: look at the patient (plots), make a guess (candidate orders), run tests (fit and compare), check the result (are the residuals clean?), and if not, go round again. This is the Box–Jenkins method, named after the two statisticians who made it popular in the 1970s.
To compare candidates we need a score that rewards fit but punishes extra parameters, otherwise the biggest model always "wins" on the training data. The AIC (Akaike information criterion) does exactly this: lower is better, and every extra parameter must improve the log-likelihood by more than 1 to pay for itself.
Three ways to say it:
- Picture: identify → estimate → check → (repeat) → forecast.
- Numbers: AIC 1840.8, 1462.0, 1408.8, 1409.5, 1410.3 for AR(0) … AR(4): pick AR(2).
- Slogan: simplest model whose residuals look like white noise.
Output of the Code-it block for 500 values simulated from AR(2) with $\phi = (0.5, 0.3)$:
- Identify: the sample PACF is 0.73, 0.32, 0.05, −0.04 (band ±0.088): two big bars, then nothing → candidate AR(2).
- Estimate AR(0) … AR(4) by maximum likelihood. AIC: 1840.8, 1462.0, 1408.8, 1409.5, 1410.3. The minimum is at $p = 2$; AR(3) adds a parameter but improves $-2\log L$ by only $1408.8 + 2 - 1409.5 = 1.3 \lt 2$.
- Fitted AR(2): $\hat\phi_1 = 0.495$, $\hat\phi_2 = 0.324$, $\hat\sigma^2 = 0.963$ (truth 0.5, 0.3, 1).
- Check: Ljung–Box on the residuals with 10 lags (df $= 10 - 2 = 8$): $Q = 9.60$, p-value $= 0.29$. No evidence of leftover autocorrelation → accept and forecast.
- Estimation: ARIMA parameters are usually fitted by (Gaussian) maximum likelihood;
statsmodelswrites the model in state-space form and computes the likelihood with the Kalman filter. A quick approximation for pure AR is least squares on the lagged values ("conditional least squares"). - AIC $= -2\log L + 2k$ and BIC $= -2\log L + k\log n$, with $k$ the number of estimated parameters and $n$ the number of observations. Lower is better. BIC punishes size more and picks smaller models; AICc is a small-sample correction of AIC. Compare criteria only between models fitted to the same data (same $d$ and $D$, since differencing changes the data).
- Ljung–Box test on residuals $e_t$: $Q = n(n+2)\sum_{k=1}^{h}\frac{r_k^2}{n-k}$, compared with a $\chi^2$ distribution with $h - (p + q)$ degrees of freedom. A small p-value means autocorrelation is left over. (Residual diagnostics in depth: Chapter 7.17.)
- Automatic selection (e.g. the Hyndman–Khandakar stepwise search in R's
auto.arimaandpmdarima.auto_arima): choose $d$ by unit-root tests, then search nearby $(p, q, P, Q)$ by AICc.
Why do we need it?
Many $(p, d, q)$ combinations fit similarly. A disciplined loop with a complexity penalty and a residual check stops you from both under-fitting (memory left in the residuals) and over-fitting (coefficients that model noise).
Where is it used?
Every ARIMA workflow; AIC/BIC also choose the number of components in mixtures, lags in VAR models and features in regressions; Ljung–Box checks the residuals of any forecasting model, including Prophet-style ones.
How is it used?
Fit a handful of candidates with ARIMA(...).fit(), read .aic, keep the lowest few, run acorr_ljungbox(res.resid, lags=[10], model_df=p+q), and finally compare the survivors on rolling-origin forecast error, which is what you actually care about.
"Model A has a lower AIC than model B, which used $d = 1$ instead of $d = 0$, so A is better."
Differencing changes the data the likelihood is computed on, so AICs with different $d$ (or $D$) are not comparable. Choose $d$ first (plots, unit-root tests), then compare AICs within that $d$, and finally compare everything by out-of-sample forecast error.
"Ljung–Box p = 0.3, so the residuals are white noise."
It means there is no strong evidence of autocorrelation at the lags tested. Absence of evidence is not proof (Chapter 5.6), and the test says nothing about changing variance or outliers.
Box–Jenkins: identify (d, then p, q from ACF/PACF) → estimate (MLE) → check (residual ACF, Ljung–Box) → forecast.
AIC $= -2\log L + 2k$, BIC $= -2\log L + k\log n$; lower is better; same $d$ only.
Trap: AIC compares fit on training data; the final judge is rolling-origin forecast error.
Quick check: model 1 has $\log L = -700$ with 3 parameters; model 2 has $\log L = -698.5$ with 5 parameters. Which does AIC prefer?
AIC₁ $= 1400 + 6 = 1406$; AIC₂ $= 1397 + 10 = 1407$. Model 1 wins: the two extra parameters improved $\log L$ by 1.5, less than the 2 they cost.
Regression with time features, and regression with ARIMA errors core
The third classical option forgets about memory altogether and treats forecasting as ordinary regression (Chapter 5.13). For each day we write down features that are known in advance: the day number $t$ (for the trend), which weekday it is, sine and cosine waves for smooth seasonal shapes, a 0/1 flag for holidays. Then least squares finds one weight per feature, and forecasting is just plugging in the features of future days.
This is exactly the skeleton of your Prophet-style model: $g(t) + s(t) + h(t) + X_t\beta$ is a regression whose columns are trend pieces, Fourier terms, holiday flags and regressors. What the plain regression lacks is (a) priors, changepoints and a flexible likelihood, and (b) any memory: it assumes the errors are independent. When they are not (a busy day is followed by a busy day), a fix is to give the errors an ARIMA model: regression with ARIMA errors.
Three ways to say it:
- Picture: a table with one row per day and columns "trend, Tuesday?, …, Sunday?, sin, cos, holiday?"; the model learns a price tag for each column.
- Numbers: $50 + 0.5\times12 + 8 + 20 = 84$ orders on day 12, a Saturday holiday.
- Slogan: structure from the calendar, memory from the errors.
A fitted regression (day 0 is a Monday): $\hat y_t = 50 + 0.5\,t + 8\,\text{Sat}_t + 5\,\text{Sun}_t + 20\,\text{Hol}_t$ (other weekday effects 0 for simplicity).
- Day 12: $12 \bmod 7 = 5$, a Saturday, and it is a holiday: $50 + 0.5\times12 + 8 + 20 = 50 + 6 + 8 + 20 = 84$.
- Day 13: a Sunday, no holiday: $50 + 6.5 + 5 = 61.5$.
- Day 14: a Monday: $50 + 7 = 57$ (Monday is the baseline day, so it has no column).
- Why no Monday column? The 7 weekday flags always add up to 1, the same as the intercept column, so with all 7 the design matrix would be singular (the "dummy variable trap"). Dropping one day makes it the baseline.
- Now the errors. Suppose yesterday's residual was $u_T = +6$ (actual 6 above the regression line) and the errors follow AR(1) with $\phi = 0.6$. Regression with AR(1) errors forecasts tomorrow as the regression value $+ 0.6\times6 = +3.6$, the day after $+0.36\times6 = +2.16$, then the correction fades. The plain regression ignores the $+6$.
Regression with time features: $y_t = \mathbf{x}_t^\top\beta + u_t$, where $\mathbf{x}_t$ holds deterministic, known-in-advance features:
- trend: $t$ (and possibly hinge columns $(t - s_j)_+$ for slope changes, Chapter 7.8);
- seasonal dummies: $m - 1$ indicators (e.g. Tue…Sun), or Fourier terms $\sin(2\pi kt/P), \cos(2\pi kt/P)$, $k = 1..K$ (Chapter 7.11). For $m = 7$, $K = 3$ spans exactly the same shapes as the 6 dummies; for long periods ($P = 365.25$) Fourier needs far fewer columns ($2K$ instead of 364) and allows non-integer periods;
- event indicators: holidays, promotions; and regressors whose future values are known (Chapter 7.12).
OLS assumes the errors $u_t$ are independent. Regression with ARIMA errors (also called dynamic regression) keeps the same mean but lets $u_t$ follow an ARIMA model, e.g. $u_t = \phi u_{t-1} + \varepsilon_t$. Its forecast is $\mathbf{x}_{T+h}^\top\hat\beta + \hat\phi^h\hat u_T$: the regression line plus a fading correction from the latest error. In statsmodels: ARIMA(y, exog=X, order=(1, 0, 0)) or SARIMAX(y, exog=X, order=...).
Why do we need it?
Calendar structure (weekdays, holidays, smooth yearly shapes) and known future drivers are often the biggest predictable part of a series. A regression uses them directly and explains each effect with a coefficient; ARIMA errors add short-term memory on top.
Where is it used?
Electricity load and call-centre models (temperature and calendar regressors), Prophet and your Prophet-style model (the same regression with priors), "dynamic harmonic regression" (Fourier terms + ARIMA errors) for long seasonal periods, and feature-based gradient boosting, which uses the same calendar features.
How is it used?
Build the design matrix from the date index only (never from future values of $y$), fit OLS, check the residual ACF; if it shows memory, refit with ARIMA(y, exog=X, order=(p, 0, q)) and pass the future rows of $X$ to forecast(h, exog=X_future).
"OLS on time features gave tiny standard errors, so every coefficient is very precise."
With autocorrelated errors the usual OLS standard errors are too small (the days are not independent pieces of evidence, Chapter 7.3). The coefficients are still fine as estimates; their uncertainty needs ARIMA errors, HAC standard errors, or a model of the dependence.
"SARIMAX(y, exog=X) fits a regression with independent errors."
Its default order is $(1, 0, 0)$, so it silently adds AR(1) errors. Pass order=(0, 0, 0) for a plain regression. Also, in statsmodels the "X" part is a regression with ARIMA errors, not a regression on lagged $y$; some textbooks use "ARIMAX" for the other meaning.
"I'll use yesterday's sales as a feature in my time-feature regression."
Then it is no longer a "known in advance" feature: for a 14-day forecast you do not know the sales of day 13. Lagged targets need a recursive forecast (or ARIMA errors), and careless use leaks future information into evaluation (Chapter 7.12).
Why choose a plain time-feature regression (OLS, or regression with ARIMA errors) instead of your Prophet-style model, and vice versa?
- Choose the plain regression when: the structure is simple and stable (no changepoints needed), you want a fast, transparent baseline with closed-form estimates, or you want ARIMA errors for short-horizon accuracy that a model with independent noise cannot give.
- Choose the Prophet-style model when: the trend changes at unknown times (changepoint grid + PELT with Laplace priors on the slope changes $\delta_j$), there are many holiday or Fourier columns that need shrinkage priors to avoid over-fitting, the data are counts or heavy-tailed (Negative Binomial, Student-t likelihoods), or you need full predictive distributions. Its mean is the same regression; the priors, changepoints and likelihood are what you add.
$y_t = \mathbf{x}_t^\top\beta + u_t$ with known-in-advance columns: $t$, $m-1$ dummies or Fourier pairs, holiday flags, known regressors.
Autocorrelated $u_t$ → regression with ARIMA errors: forecast $\mathbf{x}_{T+h}^\top\hat\beta + \hat\phi^h\hat u_T$.
Trap: dummy trap (use $m-1$); SARIMAX default order is (1,0,0); OLS SEs too small with autocorrelated errors.
Quick check: daily data with a yearly pattern. How many columns do you need for yearly seasonal dummies, and for Fourier terms of order 10?
Day-of-year dummies need 364 columns (365 days minus the baseline; leap years complicate it further). Fourier order 10 needs $2\times10 = 20$ columns, works with the non-integer period 365.25, and gives a smooth yearly shape. This is why Prophet-style models use Fourier terms for yearly seasonality.
The ARIMA family versus your Prophet-style model core
ARIMA and a Prophet-style model get their information from different places. ARIMA asks: "what did the last few values and shocks do?" It is excellent at short-term memory: if today was unusually high, tomorrow probably will be too. A Prophet-style model asks: "what does the calendar and the trend say about this date?" It is excellent at structure: weekday and yearly shapes, holidays, regressors, a trend that bends.
That suggests where each wins. Right after the forecast origin, memory matters, and ARIMA (or regression with ARIMA errors) has an edge. Far from the origin, memory has faded (an AR(1) correction shrinks like $\phi^h$), and structure is all that is left, so the structured model's calendar knowledge dominates. A Prophet-style model with independent noise leaves that short-term memory on the table, which is one of the first critiques an interviewer raises.
Three ways to say it:
- Picture: ARIMA looks in the rear-view mirror; the Prophet-style model reads the calendar and the map.
- Numbers: with AR(1) errors and φ = 0.6, the memory correction is 60% of yesterday's surprise at $h = 1$ but $0.6^{14} \approx 0.08\%$ at $h = 14$.
- Slogan: memory wins short horizons; structure wins long ones.
From the Code-it block (daily series: trend + weekly pattern + AR(1) errors with $\phi = 0.7$, $\sigma = 2$; 71 rolling origins, 14-day forecasts):
- Regression with independent errors ("Prophet-like" mean): MAE 2.39 at $h = 1$ and 2.24 at $h = 14$. Its error at every horizon is the whole AR(1) error, sd $2/\sqrt{1 - 0.49} \approx 2.80$; for a Normal error, MAE ≈ $0.8\times$ sd $\approx 2.24$.
- Same regression with AR(1) errors (estimated $\hat\phi = 0.69$): MAE 1.50 at $h = 1$. One step ahead only the new shock is unknown: sd 2, MAE ≈ $0.8\times2 = 1.6$.
- At $h = 14$: MAE 2.29, the same as the plain regression (2.24) within noise: $0.69^{14} \approx 0.006$, so the memory correction has vanished.
- Conclusion: modelling the error memory cut the 1-day-ahead error by about 37% and did nothing for 2-week-ahead forecasts.
| ARIMA / SARIMA | Regression with time features (+ ARIMA errors) | Prophet-style Bayesian model | |
|---|---|---|---|
| Main information | recent values and shocks | calendar columns (+ recent errors) | trend with changepoints, Fourier seasonality, holidays, regressors |
| Trend handling | differencing ($d$) | $t$ column (fixed slope) | piecewise-linear $g(t)$, Laplace priors on slope changes |
| Seasonality | one integer $m$ | dummies or Fourier, several periods | Fourier, several periods, with priors |
| Short-term memory | yes (its core) | only with ARIMA errors | no (independent noise) |
| Data needs | regular grid, enough history to estimate dependence | known future features | known future features; tolerates gaps |
| Likelihood | Gaussian | Gaussian | Normal, Student-t, Negative Binomial |
| Uncertainty | plug-in intervals | plug-in intervals | posterior predictive (incl. parameter uncertainty, as well as SVI captures it) |
| Interpretation | coefficients on lags: hard to explain | coefficients per feature: easy | component plots: easy |
Why do we need it?
Interviewers will ask why you did not use ARIMA. A good answer names the assumption each model makes, the situations where each wins, and the evidence (rolling-origin comparisons by horizon) that justified your choice.
Where is it used?
Model reviews in forecasting teams, papers comparing forecasting methods (ARIMA and ETS are the standard classical benchmarks), and hybrid systems: a structured model for the calendar plus an ARIMA model on its residuals for short horizons.
How is it used?
Evaluate both on the same rolling origins and report error by horizon. If ARIMA wins at short horizons, consider adding an AR error term or an ARIMA residual correction to the structured model; if the structured model wins everywhere, say so with numbers.
"ARIMA is outdated; the Bayesian model replaces it."
They encode different information. ARIMA (or ETS) remains a strong, cheap benchmark, and its short-term memory is exactly what a Prophet-style model with independent noise lacks. Many production systems combine them.
"ARIMA can use holidays if I add a seasonal term."
Seasonal terms repeat every $m$ steps; most holidays do not fall on the same position of a fixed cycle (Easter moves; Black Friday moves within November). Holidays need indicator regressors, i.e. regression with ARIMA errors or a structured model.
Why choose ARIMA/SARIMA instead of your Prophet-style model, and vice versa?
- Choose ARIMA/SARIMA when: short horizons matter most and the series has strong short-term dependence (yesterday predicts today); there is one regular seasonality and no important calendar events; you want a well-understood classical model with cheap automatic order selection (
auto_arima) for many series; or you need a benchmark that captures dependence your model ignores. - Choose the Prophet-style model when: forecasts are needed weeks or months ahead (where calendar structure dominates and memory has faded); there are multiple seasonalities, moving holidays and known regressors; the trend changes at changepoints; the data are counts or have outliers (Negative Binomial, Student-t likelihoods); data have gaps; or stakeholders need component plots and full predictive distributions.
- Both: fit the structured model, check the residual ACF (Chapter 7.17); if lag-1 autocorrelation is large and short horizons matter, add an AR term on the residuals (a regression-with-ARIMA-errors idea) and show the gain by horizon.
"My model doesn't need to handle autocorrelation because the trend and seasonality explain everything."
"The model assumes the remaining noise is independent. I checked the residual ACF and Ljung–Box: if there is memory left, short-horizon point forecasts and interval coverage suffer, and I'd add an AR error term or compare with ARIMA errors by horizon."
Model answer: "ARIMA models dependence on recent values; a Prophet-style model models the mean as a function of time with independent noise. So ARIMA tends to win one or two steps ahead when errors are autocorrelated, while the structured model wins at longer horizons and whenever calendar effects, regressors or non-Gaussian likelihoods matter. I compare them by horizon on rolling origins, and the residual ACF tells me whether a hybrid is worth it."
ARIMA = memory (recent values/shocks); Prophet-style = structure (trend + calendar + regressors) with independent noise.
AR error correction fades like $\phi^h$: memory wins short horizons, structure wins long ones.
Trap: Prophet-style ignores residual autocorrelation; ARIMA ignores calendar events. Compare by horizon.
Quick check: your residuals have lag-1 autocorrelation 0.5. Roughly how much of yesterday's residual should a 1-day-ahead forecast add back, and how much a 5-day-ahead one?
With AR(1) errors and φ ≈ 0.5: $0.5$ of yesterday's residual at $h = 1$, and $0.5^5 \approx 0.03$ at $h = 5$. The correction matters for tomorrow and is negligible within a week.
Recap, cheat sheet and practice
- AR($p$): regress on own lags; AR(1) stationary iff $\lvert\phi\rvert \lt 1$, mean $c/(1-\phi)$, $\rho_k = \phi^k$; AR(2) stationary inside the triangle; complex roots give pseudo-cycles.
- MA($q$): weighted sum of the last $q$ shocks; always stationary; ACF cuts off after $q$; invertible if $\lvert\theta\rvert \lt 1$ (MA(1)); not a rolling mean of observations.
- Signatures: PACF cuts → AR; ACF cuts → MA; both tail → ARMA; ACF stuck near 1 → difference first.
- ARIMA($p,d,q$): ARMA on the $d$-th difference, integrated back. Naive = (0,1,0), drift = (0,1,0) + constant, SES = (0,1,1) with θ = α − 1. Over-differencing → θ ≈ −1.
- SARIMA($p,d,q$)($P,D,Q$)$_m$: seasonal differences and seasonal AR/MA at lags $m, 2m$; one integer $m$; a fixed pattern is better as dummies/Fourier.
- Forecast behaviour: variance $\sigma^2\sum_{j\lt h}\psi_j^2$; stationary → mean reversion and bounded width; $d = 1$ → flat or drift line and width ∝ $\sqrt h$; intervals ignore parameter uncertainty.
- Fitting: Box–Jenkins (identify, estimate by MLE, check residuals with Ljung–Box, forecast); AIC/BIC within the same $d$; final judge = rolling-origin error.
- Time-feature regression: $t$, $m-1$ dummies or Fourier pairs, holiday flags, known regressors; autocorrelated errors → regression with ARIMA errors. This is the skeleton of the Prophet-style model.
- vs Prophet-style: memory (ARIMA) wins short horizons; structure (calendar, changepoints, likelihoods, priors) wins long horizons and explains itself. Compare by horizon.
Cheat sheet
| Model | Equation | Remember |
|---|---|---|
| AR(1) | $y_t = c + \phi y_{t-1} + \varepsilon_t$ | $\mu = c/(1-\phi)$; $\rho_k = \phi^k$; forecast $\mu + \phi^h(y_T - \mu)$ |
| AR(2) stationarity | $\phi_1 + \phi_2 \lt 1$, $\phi_2 - \phi_1 \lt 1$, $\lvert\phi_2\rvert \lt 1$ | cycles if $\phi_1^2 + 4\phi_2 \lt 0$ |
| MA(1) | $y_t = \mu + \varepsilon_t + \theta\varepsilon_{t-1}$ | $\rho_1 = \theta/(1+\theta^2)$, then 0; forecast = μ after 1 step |
| ARIMA($p,d,q$) | $\phi(B)(1-B)^d y_t = c + \theta(B)\varepsilon_t$ | $d$ = 0/1 usually; constant with $d = 1$ = drift |
| SARIMA | $\Phi(B^m)\phi(B)(1-B)^d(1-B^m)^D y_t = c + \Theta(B^m)\theta(B)\varepsilon_t$ | seasonal naive = (0,0,0)(0,1,0)$_m$ |
| Forecast variance | $\sigma^2\sum_{j=0}^{h-1}\psi_j^2$ | AR(1) → $\sigma^2/(1-\phi^2)$; RW → $h\sigma^2$ |
| AIC / BIC | $-2\log L + 2k$ / $-2\log L + k\log n$ | lower is better; same data (same $d$, $D$) |
| Ljung–Box | $Q = n(n+2)\sum_{k\le h} r_k^2/(n-k) \sim \chi^2_{h-p-q}$ | small p → memory left in residuals |
| Dynamic regression | $y_t = \mathbf{x}_t^\top\beta + u_t$, $u_t$ ARIMA | forecast $\mathbf{x}^\top\hat\beta + \hat\phi^h\hat u_T$ |
statsmodels | ARIMA(y, exog=X, order=(p,d,q), seasonal_order=(P,D,Q,m), trend=...) | SARIMAX default order is (1,0,0) |
import numpy as np
import warnings
from statsmodels.tsa.arima_process import ArmaProcess
from statsmodels.tsa.arima.model import ARIMA
from statsmodels.tsa.statespace.sarimax import SARIMAX
from statsmodels.tsa.holtwinters import SimpleExpSmoothing
from statsmodels.tsa.stattools import acf, pacf
from statsmodels.stats.diagnostic import acorr_ljungbox
warnings.filterwarnings("ignore") # silence convergence chatter in this demo
rng = np.random.default_rng(1)
# 1) The AR(1) worked example: c = 10, phi = 0.5 (mean 20), last value 28
c, phi, y = 10, 0.5, 28.0
fc = []
for h in range(4):
y = c + phi * y
fc.append(y)
print(fc) # [24.0, 22.0, 21.0, 20.5]: the gap to 20 halves each step
# 2) Signatures: AR(2) has a PACF that cuts off after lag 2; MA(1) has an ACF that cuts off after lag 1
# ArmaProcess wants the lag polynomials: AR "1 - 0.5L - 0.3L^2" -> [1, -0.5, -0.3]; MA "1 + 0.6L" -> [1, 0.6]
ar2 = ArmaProcess([1, -0.5, -0.3], [1]).generate_sample(500, distrvs=rng.standard_normal, burnin=200)
ma1 = ArmaProcess([1], [1, 0.6]).generate_sample(500, distrvs=rng.standard_normal, burnin=200)
print("AR(2) acf ", acf(ar2, nlags=4)[1:].round(2), " pacf", pacf(ar2, nlags=4)[1:].round(2))
# AR(2) acf [0.73 0.68 0.59 0.5 ] pacf [ 0.73 0.32 0.05 -0.04] (PACF cuts off after lag 2)
print("MA(1) acf ", acf(ma1, nlags=4)[1:].round(2), " pacf", pacf(ma1, nlags=4)[1:].round(2))
# MA(1) acf [ 0.42 0. -0.01 -0.07] pacf [ 0.42 -0.21 0.1 -0.14] (ACF cuts off after lag 1)
print("bands +-", round(1.96 / np.sqrt(500), 3)) # +- 0.088
# 3) Choose the AR order by AIC, then check the residuals with Ljung-Box
for p in range(5):
print(p, round(ARIMA(ar2, order=(p, 0, 0)).fit().aic, 1)) # 1840.8, 1462.0, 1408.8, 1409.5, 1410.3 -> p = 2
fit = ARIMA(ar2, order=(2, 0, 0)).fit()
print(fit.params.round(3)) # const, ar.L1, ar.L2, sigma2 (truth 0, 0.5, 0.3, 1)
print(acorr_ljungbox(fit.resid, lags=[10], model_df=2)) # lb_stat 9.60, lb_pvalue 0.29: no structure left
# 4) Forecast behaviour: AR(1) intervals level off, random-walk intervals keep growing
ar1 = ArmaProcess([1, -0.5], [1]).generate_sample(300, distrvs=rng.standard_normal, burnin=200)
rw = np.cumsum(rng.standard_normal(300))
for name, series, order in [("AR(1)", ar1, (1, 0, 0)), ("random walk", rw, (0, 1, 0))]:
f = ARIMA(series, order=order).fit().get_forecast(20)
width = np.diff(f.conf_int(alpha=0.05), axis=1).ravel()
print(name, "95% width at h=1, 5, 20:", width[[0, 4, 19]].round(2))
# AR(1) 95% width at h=1, 5, 20: [3.93 4.76 4.77] levels off
# random walk 95% width at h=1, 5, 20: [ 3.78 8.44 16.88] grows like sqrt(h)
# 5) SES is ARIMA(0,1,1) with theta = alpha - 1
lvl = np.cumsum(rng.normal(0, 1, 400)) + rng.normal(0, 2, 400) # random-walk level + noise
theta = ARIMA(lvl, order=(0, 1, 1)).fit().params[0]
alpha = SimpleExpSmoothing(lvl, initialization_method="estimated").fit().params["smoothing_level"]
print("1 + theta =", round(1 + theta, 3), " SES alpha =", round(alpha, 3)) # 0.405 and 0.404
# 6) Daily data with a weekly season: SARIMA(1,0,0)(0,1,1)_7 with a linear trend term
n = 420
t = np.arange(n)
weekly = np.array([0, 2, 1, 3, 10, 20, 15])
u = ArmaProcess([1, -0.7], [1]).generate_sample(n, distrvs=rng.standard_normal, burnin=200) * 2
y = 40 + 0.05 * t + weekly[t % 7] + u # trend + weekly pattern + AR(1) errors (phi = 0.7, sd 2)
sar = ARIMA(y[:-14], order=(1, 0, 0), seasonal_order=(0, 1, 1, 7), trend="t").fit()
print(sar.params.round(3)) # [ 0.044 0.704 -0.999 4.352] = drift, ar.L1, ma.S.L7, sigma2
print("SARIMA holdout MAE:", round(np.mean(np.abs(y[-14:] - sar.forecast(14))), 2)) # 1.94
# ma.S.L7 near -1: the weekly pattern is fixed, so the seasonal difference over-differenced it
# 7) Regression with time features (trend + weekday dummies): independent errors vs AR(1) errors
X = np.column_stack([t] + [(t % 7 == d).astype(float) for d in range(1, 7)])
reg = ARIMA(y[:336], exog=X[:336], order=(0, 0, 0)).fit() # "Prophet-like": independent errors
arx = ARIMA(y[:336], exog=X[:336], order=(1, 0, 0)).fit() # same regression + AR(1) errors
print("AR coefficient of the errors:", round(dict(zip(arx.param_names, arx.params))["ar.L1"], 2)) # 0.69 (truth 0.7)
err = {"independent errors": [], "AR(1) errors": []}
for o in range(336, 407): # 71 rolling origins, parameters kept fixed
for name, res in [("independent errors", reg), ("AR(1) errors", arx)]:
f = res.apply(y[:o], exog=X[:o]).forecast(14, exog=X[o:o + 14])
err[name].append(np.abs(y[o:o + 14] - f))
for name, e in err.items():
e = np.array(e)
print(f"{name:18s} MAE at h=1: {e[:, 0].mean():.2f} at h=14: {e[:, 13].mean():.2f}")
# independent errors MAE at h=1: 2.39 at h=14: 2.24
# AR(1) errors MAE at h=1: 1.50 at h=14: 2.29
Output checked with statsmodels 0.15.0 (runs in about 2 seconds). Results depend on the random seed; the patterns (cut-offs, AIC minimum at the true order, bounded vs growing widths, α = 1 + θ, the short-horizon gain from AR errors) do not.
1. A stationary series has a PACF with significant bars at lags 1 and 2 only, and an ACF that decays slowly. Which model is the natural first candidate?
2. $y_t = 6 + 0.7\,y_{t-1} + \varepsilon_t$. What value do long-range forecasts approach?
3. Which forecasting method is the same as an ARIMA(0,1,1) model?
4. A random walk has shock sd σ = 2. What is the forecast standard deviation 9 steps ahead?
5. You call SARIMAX(y, exog=X).fit() to fit "a plain regression on calendar features". What did you actually fit?
SARIMAX's default order is $(1,0,0)$, and its exogenous part is a regression with ARIMA errors. Pass order=(0, 0, 0) for a plain regression.6. The residuals of your Prophet-style model have lag-1 autocorrelation 0.7. Where would adding an AR(1) error term help most?
Practice problems
A. $y_t = 4 + 0.8\,y_{t-1} + \varepsilon_t$, σ = 1, last value 30. Give the mean, the forecasts for h = 1, 2, 3, the forecast sd for h = 1, 2, and its limit.
- Mean: $4/(1 - 0.8) = 20$.
- Forecasts: $20 + 0.8\times10 = 28$; $20 + 0.64\times10 = 26.4$; $20 + 0.512\times10 = 25.12$ (check: $4 + 0.8\times28 = 26.4$, $4 + 0.8\times26.4 = 25.12$).
- sd: $h = 1$: 1; $h = 2$: $\sqrt{1 + 0.64} \approx 1.28$; limit $1/\sqrt{1 - 0.64} = 1/0.6 \approx 1.67$.
B. MA(1) $y_t = 10 + \varepsilon_t - 0.5\,\varepsilon_{t-1}$. Compute $\rho_1$, $\rho_2$, and the forecasts for h = 1, 2 if the last shock was $\varepsilon_T = 2$.
- θ = −0.5: $\rho_1 = -0.5/(1 + 0.25) = -0.4$; $\rho_2 = 0$.
- $h = 1$: $10 + (-0.5)\times2 = 9$. $h = 2$: $10$ (the shock is forgotten after one step).
- Note the sign convention: in Box–Jenkins notation ($y_t = \mu + \varepsilon_t - \theta\varepsilon_{t-1}$) the same model has $\theta = +0.5$.
C. Data 50, 52, 55, 57, 60. Forecast three steps with (i) ARIMA(0,1,0) with drift, (ii) ARIMA(1,1,0) with φ = 0.4 and no constant.
- Differences: 2, 3, 2, 3; mean 2.5.
- (i) Drift: $60 + 2.5 = 62.5$, $65$, $67.5$.
- (ii) Forecast differences: $0.4\times3 = 1.2$, $0.48$, $0.192$. Levels: $61.2$, $61.68$, $61.872$: the forecast flattens because, without a constant, future changes revert to 0.
D. (Interview) "Why didn't you just use ARIMA for your forecasting project?"
"ARIMA models short-term dependence through lags and differencing, with Gaussian errors and one seasonal period. My data had weekly and yearly seasonality, holidays that move around the calendar, external regressors with known future values, a trend whose slope changes, and counts that need a Negative Binomial likelihood. A Prophet-style model represents all of that directly, with priors (Laplace on the changepoint slope changes) and a posterior predictive distribution. ARIMA's strength is short-horizon memory, which my model's independent noise does not capture, so I would benchmark against ARIMA/ETS by horizon and check the residual ACF; if one-day-ahead accuracy mattered and residuals were autocorrelated, I'd add an AR error term."
E. (Interview) "Here is an ACF with one large spike at lag 1 (about −0.48) and nothing else, computed after first-differencing. What do you conclude?"
"ACF cutting off after lag 1 means an MA(1) for the differenced series, and a lag-1 value near −0.5 means θ ≈ −1. That is the fingerprint of over-differencing: the original series was probably already stationary (differencing white noise gives exactly −0.5). I'd go back to $d = 0$, check the ACF of the original series, and compare AIC within $d = 0$; the variance of the differenced series being larger than the original is another hint."
F. Daily orders with a weekly pattern, a yearly pattern and 10 public holidays per year, 3 years of history. Sketch a classical model and say how many seasonal parameters it needs, compared with SARIMA with m = 365.
- A dynamic harmonic regression: columns $t$ (trend), 6 weekday dummies, yearly Fourier terms of order $K$ (say $K = 10$: 20 columns), and holiday indicators (10 columns, or fewer if grouped), with ARIMA errors (e.g. AR(1) or chosen by AIC).
- Seasonal parameters: $6 + 20 = 26$ coefficients, all estimated from 3 years (≈ 1 095 days).
- SARIMA with $m = 365$: the seasonal difference uses up the whole first year of data, the state-space form needs a state vector of more than 365 numbers (slow and fragile to fit), and moving holidays are still not handled. This is why both the classical answer (Fourier + ARIMA errors) and the Prophet-style answer use Fourier terms for long seasonal periods.
The additive forecasting model
Your forecasting model writes every day's value as a sum: a slow trend, plus a repeating seasonal pattern, plus holiday bumps, plus the effect of outside drivers, plus noise. This chapter is about that sum itself: what each layer is for, what it quietly assumes, why "adding things up" makes a model easy to read and to debug, how the whole thing is one Bayesian regression with structured columns and priors, and when adding is the wrong way to combine the layers.
- Read the model $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$ out loud, symbol by symbol, and compute one day's value as a sum of five numbers
- Say each layer's job and its assumptions, and recognise the pattern a missing layer leaves in the residuals
- Reason about the components separately (what-if questions, component plots) and know when the data cannot split the credit between them
- See the whole model as one regression with column groups, priors on each group and a chosen likelihood, and count its parameters
- Use the model generatively: simulate series from the priors before fitting (a prior predictive check)
- Choose between additive and multiplicative seasonality from a plot, and know what a log link does
- Explain how a forecast is made: continue every layer into the future, and know which layers need outside information
One equation, five layers core
Think of a shop receipt. The total at the bottom is not a mystery number: it is a list of items added up. If the total looks strange, you read the lines and find the item that caused it.
Your forecasting model treats each day's value the same way. Tomorrow's orders are "the usual level for this time of the year's growth" plus "the extra you always get on a Saturday" plus "the bump because it is a holiday" plus "the effect of the marketing e-mails we sent" plus "a bit of luck we cannot predict". Each line on the receipt is one component (one layer) of the model.
Because the lines are just added, every line is in the same units as the data (orders), and you can read, check and change each one on its own.
Three ways to say it:
- Picture: five transparent sheets laid on top of each other; each sheet draws one cause; looking through all of them you see the series.
- Numbers: Saturday, day 40: trend 140 + Saturday +22 + holiday +35 + e-mails +12 = mean 209 orders; the luck of the day (−5) makes the observed 204.
- Slogan: one day's value is a receipt with five lines.
A shop whose trend is $g(t) = 120 + 0.5\,t$ orders, where $t$ counts days. Day 0 is a Monday, so day 40 is a Saturday ($40 = 5\times7 + 5$, and the 6th day of the week is Saturday). Day 40 is also a public holiday, and the shop sent 3 thousand marketing e-mails that morning; each thousand e-mails adds about $\beta = 4$ orders.
- Trend (the slow level): $g(40) = 120 + 0.5\times40 = 120 + 20 = 140$ orders.
- Seasonality (the weekly pattern): Saturdays are usually 22 orders above the weekly average, so $s(40) = +22$.
- Holiday: this holiday usually adds 35 orders, so $h(40) = +35$.
- Regressor (an outside driver): $X_{40}\beta = 3 \times 4 = +12$.
- Mean (what the model expects): $\mu_{40} = 140 + 22 + 35 + 12 = 209$ orders.
- Noise (the luck of the day): the shop actually received 204 orders, so $\epsilon_{40} = 204 - 209 = -5$.
- Next day, Sunday 41, no holiday, 1 thousand e-mails: $g(41) = 140.5$, $s = +16$, $h = 0$, $X\beta = 4$, so $\mu_{41} = 140.5 + 16 + 0 + 4 = 160.5$. The holiday line just drops to zero; the other lines do not care.
The additive forecasting model (the form used by Prophet and by your model) says that the value at time $t$ is
$$y_t \;=\; \underbrace{g(t)}_{\text{trend}} \;+\; \underbrace{s(t)}_{\text{seasonality}} \;+\; \underbrace{h(t)}_{\text{holidays}} \;+\; \underbrace{X_t\beta}_{\text{regressors}} \;+\; \underbrace{\epsilon_t}_{\text{noise}}.$$- $y_t$: the observed value at time $t$ (orders on day $t$).
- $g(t)$: the trend, the slow, long-run level. In your model it is piecewise-linear with changepoints (Chapter 7.8).
- $s(t)$: the seasonality, patterns that repeat with a fixed period (a week, a year), built from sines and cosines (Fourier terms, Chapter 7.11). With several periods, $s(t)$ is their sum.
- $h(t)$: the holiday (event) effects, bumps on known dates (Chapter 7.12).
- $X_t$: a row of $R$ numbers, the exogenous regressors on day $t$ ("exogenous" = coming from outside the series: e-mails sent, price, temperature). $\beta$: a column of $R$ weights, one per regressor. $X_t\beta = \sum_{r=1}^{R} x_{t,r}\beta_r$.
- $\epsilon_t$: the observation noise, the part no layer explains, with mean 0. Its distribution is the model's likelihood (Normal, Student-t, or Negative Binomial for counts, Chapter 7.13).
Everything except the noise is the mean $\mu_t = g(t) + s(t) + h(t) + X_t\beta$, so the model can also be written as "$y_t$ is drawn around $\mu_t$", for example $y_t \sim N(\mu_t, \sigma^2)$. Two conventions keep the layers tidy: the seasonal pattern averages to about zero over one period (so the level lives in the trend), and the noise has mean zero (so the average lives in $\mu_t$).
Why do we need it?
A real series is a mix of several causes at once. Writing it as a sum gives each cause its own small, interpretable piece, so you can forecast each piece in the way that suits it (continue the trend, repeat the season, read the calendar) and explain any day's forecast line by line.
Where is it used?
Prophet and Prophet-style NumPyro or Stan models (yours is one), classical decomposition and STL (Chapter 7.2), structural time-series models (local level + season + regression), marketing-mix models, and regression with time features as a baseline (Chapter 7.6).
How is it used?
Write the mean as a sum of component functions, give each its own parameters and priors, fit them all together (in your case by SVI), then report a forecast together with its component breakdown: "the trend says 140, Saturday adds 22, the holiday adds 35…".
"Additive means the five layers are independent random things."
"Additive" only says how the layers are combined: by adding. Their values can be related (more e-mails are sent before holidays, for example). Such overlaps matter when you try to split the credit between layers (see reasoning about components separately).
"The seasonal layer tells you the level of a Saturday."
The seasonal layer is a deviation that averages to about zero over a week: "+22 compared with an average day". The level lives in the trend. A seasonal value of +22 means nothing without the trend underneath it.
"$\epsilon_t$ is the model's error on day $t$."
$\epsilon_t$ is the part of the process that no layer explains. What you can compute after fitting is the residual $e_t = y_t - \hat\mu_t$, an estimate of $\epsilon_t$ that also contains the model's mistakes (Chapter 7.17).
This equation is your forecasting model: $g(t)$ is the piecewise-linear trend whose changepoints come from a Prophet-like grid plus PELT detection, with Laplace priors on the slope changes $\delta_j$; $s(t)$ is Fourier seasonality; $h(t)$ is the holiday effects; $X_t\beta$ is the exogenous regressors; and the distribution of $\epsilon_t$ is the likelihood you choose (Normal, Student-t or Negative Binomial). In an interview, start every explanation from this line and point at the term you are talking about.
$y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$; mean $\mu_t$ = everything except $\epsilon_t$; every term is in the units of $y$.
Seasonality is a zero-average deviation; the level lives in the trend; noise has mean 0.
Trap: "additive" is about how layers combine, not about them being unrelated; $\epsilon_t$ (unobservable) ≠ residual $e_t$.
Quick check: trend 150, it is a Tuesday (weekly effect −14), no holiday, 2 thousand e-mails at 4 orders each, and 139 orders were observed. What are $\mu_t$ and $\epsilon_t$?
$\mu_t = 150 - 14 + 0 + 2\times4 = 144$ orders. $\epsilon_t = y_t - \mu_t = 139 - 144 = -5$: a slightly unlucky Tuesday.
Each layer's job, and what it assumes core
Picture a small team cleaning a house. One person does the floors, one the windows, one the kitchen. If the window-cleaner is off sick, nobody else does the windows: the dirt stays where it was, and anyone walking in can see exactly which job was not done.
The layers of the model are that team. Each one has one job: explain one kind of pattern. Whatever no layer explains is left behind in the residuals $e_t = y_t - \hat\mu_t$ (the "leftovers"). So if a layer is missing or wrong, its pattern shows up in the leftovers: a weekly wave, spikes on holidays, a slow bend, or a wiggle that follows an outside driver.
Each layer also makes assumptions (the trend is straight between changepoints, the weekly pattern has the same shape every week, a holiday has the same effect every year…). When reality breaks an assumption, the leftovers show that too.
Three ways to say it:
- Picture: a team where each person has one job; a missing person leaves their mess visible.
- Numbers: drop the weekly layer and the Saturday residuals sit around +22 while the Tuesday residuals sit around −14.
- Slogan: what a layer does not explain, the residuals show.
The shop's true weekly pattern is Mon −10, Tue −14, Wed −12, Thu −6, Fri +4, Sat +22, Sun +16 orders (these add up to $-10-14-12-6+4+22+16 = 0$). Fit the model without a seasonal layer.
- Without $s(t)$, the model's mean for every day of a week is about the same (trend + holidays + e-mails). The weekly pattern has nowhere to go.
- So on a Saturday the residual is about $y_t - \hat\mu_t \approx +22 + \epsilon_t$, and on a Tuesday about $-14 + \epsilon_t$.
- Average the residuals by weekday over many weeks: the noise averages out and you read back roughly $-10, -14, -12, -6, +4, +22, +16$: the missing layer, printed in the leftovers.
- The residual standard deviation grows too. The weekly values have mean 0 and mean square $(100+196+144+36+16+484+256)/7 = 1232/7 = 176$, so they add about 176 to the residual variance. With noise sd 5 (variance 25): $\sqrt{25 + 176} = \sqrt{201} \approx 14.2$ instead of about 5.
Each layer's job, usual form and main assumptions:
| Layer | Job (what it explains) | Usual form | Assumes | If it is missing or wrong, residuals show… |
|---|---|---|---|---|
| $g(t)$ trend | the slow, long-run level and growth | piecewise-linear: $kt + m$ plus slope changes $\delta_j$ at changepoints $s_j$ | straight between changepoints; changes are occasional; the last slope continues | a slow bend or drift (residuals positive for weeks, then negative) |
| $s(t)$ seasonality | patterns that repeat every period $P$ (7 days, 365.25 days) | Fourier terms $\sum_n a_n\cos\frac{2\pi n t}{P} + b_n\sin\frac{2\pi n t}{P}$ | known period; the same shape every cycle; in additive mode the same size at every level | a regular wave with the period; or a growing wave if its size should scale with the level |
| $h(t)$ holidays | bumps on known dates (holidays, events) | 0/1 indicator columns, often with windows of days around the date | dates known in advance; the same effect each time the event recurs | spikes on the event dates |
| $X_t\beta$ regressors | effects of measured outside drivers | one column per driver, one weight each | a linear, constant effect; the value is known (or forecast) for future days; no future information leaks in | a wiggle that follows the driver |
| $\epsilon_t$ noise | whatever is left | the likelihood: Normal, Student-t, Negative Binomial | independent over time given $\mu_t$; the chosen shape and spread (constant $\sigma$ for a Normal) | — (it is the leftovers: they should look like plain noise) |
A model is well specified when every systematic pattern has a layer to go to, so the residuals have no structure left: no trend, no wave, no spikes, no link to drivers, no autocorrelation (Chapter 7.17).
Why do we need it?
Knowing each layer's job tells you where to look when something is off, and knowing its assumptions tells you when the model will fail before it does: a seasonal pattern whose size grows, a holiday that moved, an e-mail effect that saturates.
Where is it used?
Model building and debugging for Prophet-style models, residual diagnostics (Chapter 7.17), feature reviews in demand forecasting, and interview questions like "what does each component of your model assume?".
How is it used?
After fitting, plot the residuals against time, group them by weekday and month, mark holidays on the plot, and scatter them against each regressor. Each pattern you find points to one layer to add, enlarge (more Fourier terms, a holiday window) or change (multiplicative mode, another likelihood).
"The fit line looks close to the data, so every layer is right."
A missing weekly layer can hide inside a wiggly fit line, and a wrong one can still look close. Judge each layer by the residuals: group them by weekday, look at them on holidays, scatter them against each regressor and plot them over time.
"The noise term will absorb anything the layers miss, so nothing breaks."
It does absorb it, and that is exactly the problem: the leftovers become patterned, the estimated $\sigma$ is too big on normal days and too small on the special ones, and the forecast intervals are wrong in a systematic way.
"A layer with a good shape is enough."
Each layer also has assumptions about stability: the same weekly shape every week, the same holiday effect every year, the same e-mail effect at any volume. A pattern that drifts over time breaks the layer even when its shape was right at the start.
When your forecasting model misbehaves, this table is the debugging plan: a 7-day wave in the residuals points at the weekly Fourier order (7.11), spikes at the holiday list and windows (7.12), a slow bend at the changepoints and their Laplace scale (7.8–7.10), a funnel at the likelihood or the additive form (7.13, multiplicative seasonality), and day-to-day memory at the missing autoregressive part (7.17).
Trend = slow level; seasonality = fixed-period repeats; holidays = known dates; regressors = outside drivers; noise = the rest.
A missing or wrong layer leaves its pattern in the residuals $e_t = y_t - \hat\mu_t$.
Trap: a close-looking fit proves nothing; check residuals by weekday, on holidays, against each regressor, and over time.
Quick check: after fitting, the residuals on the three Black Fridays in your data are +80, +95 and +70, and all other residuals look like noise. Which layer is missing or wrong?
The holiday layer: a known, dated event leaves positive spikes on exactly its dates. Add a Black Friday indicator (perhaps with a window for the days around it, Chapter 7.12). The average residual, about $(80 + 95 + 70)/3 \approx 82$ orders, is a first guess of its effect.
Reasoning about each component separately core
Because the layers are added, each one is a separate number in the same units, and changing one does not change the others. That gives you three superpowers:
- Read each layer on its own: "Saturdays add about 22 orders", "the holiday adds about 35", "each thousand e-mails adds about 4".
- Ask what-if questions: "if we cancel tomorrow's e-mails, the forecast drops by exactly the e-mail line, whatever the day or the trend".
- Debug one layer at a time: a strange weekly plot is a seasonality problem, not a trend problem.
There is one condition. The data must be able to tell the layers apart. If the shop sends its e-mails every Saturday, the data only ever sees "Saturday-with-e-mails". It can measure the total lift, but not how much of it belongs to Saturday and how much to the e-mails.
Three ways to say it:
- Picture: five separate dials; turning one moves the total by exactly that dial's amount.
- Numbers: cancelling 3 thousand e-mails lowers the forecast by $3 \times 4 = 12$ orders, on a Tuesday or on a Saturday, in January or in June.
- Slogan: additive layers can be read one by one, as long as the data can tell them apart.
Part 1: a what-if. The model's forecast for next Saturday is trend 150, Saturday +22, no holiday, 3 thousand e-mails at 4 orders each.
- Forecast with the e-mails: $150 + 22 + 0 + 3\times4 = 184$ orders.
- Forecast without them ($x = 0$): $150 + 22 + 0 + 0 = 172$ orders.
- Difference: $184 - 172 = 12 = 3 \times 4$. In an additive model the effect of a change in one layer is that layer's change, nothing more. It would also be 12 on a Tuesday.
Part 2: when the split is impossible. Suppose every Saturday had exactly 3 thousand e-mails and no other day had any. On Saturdays the data show a lift of about 34 orders over an average day.
- The model's Saturday lift is $s(\text{Sat}) + 3\beta$. The data pin down only this sum: $s(\text{Sat}) + 3\beta = 34$.
- "Saturday +22, e-mails 4 each" gives $22 + 12 = 34$. "Saturday +34, e-mails do nothing" gives $34 + 0 = 34$. "Saturday +10, e-mails 8 each" gives $10 + 24 = 34$.
- All three fit the data equally well. The total forecast for a usual Saturday is the same (34 above average), but the credit is not determined, and the what-if "cancel the e-mails" gives 12, 0 or 24. Only data where e-mails and Saturdays sometimes come apart (or a prior) can split it.
After fitting, the forecast splits into component contributions:
$$\hat\mu_t = \hat g(t) + \hat s(t) + \hat h(t) + X_t\hat\beta,$$and each term can be plotted on its own (a "components plot": the trend over time, the weekly pattern over Monday…Sunday, the holiday effects, the regressor effects).
- Separability: in an additive model, a change in one layer changes $\mu_t$ by exactly that change ($\partial\mu_t/\partial h(t) = 1$, and the same for every layer). There are no interactions: the Saturday effect does not depend on the trend level, the e-mail effect does not depend on the weekday. (If you need such an interaction, you must add it as its own column, or use multiplicative mode, below.)
- Identifiability (in plain words: whether the data can pin a quantity down; taught in depth in Chapter 6.8): the sum $\mu_t$ is usually well determined, but the individual layers are determined only when their columns are not (nearly) copies or combinations of each other. When two layers can make the same shape, their estimates trade off: one goes up, the other goes down, and their uncertainties become large and strongly (negatively) correlated.
Why do we need it?
Stakeholders do not ask "what is the forecast?" only; they ask "why is it high?" and "what if we cancel the promotion?". Separate, additive layers answer both directly, and the identifiability caveat tells you when such answers are not supported by the data.
Where is it used?
Prophet's plot_components, explaining demand forecasts to planners, what-if scenarios in marketing-mix models, holiday-effect reports, and debugging a Prophet-style NumPyro model by plotting each latent component's posterior.
How is it used?
Compute each layer's contribution from the fitted (or posterior) weights, plot them in separate panels, and check that each looks sensible. For a what-if, change one input (a regressor, a holiday flag) and recompute. Before trusting a split, check the correlation between the layers' columns and the posterior correlation of their weights.
"The components plot shows the true causes of the series."
It shows how this model splits the fitted mean under its assumptions. If two layers can make the same shape (a rising regressor and the trend, e-mails that are sent every Saturday and the weekly pattern), the split is only as good as the data that separates them, plus the priors.
"Additive layers can capture any interaction automatically."
They cannot. "E-mails work better on weekends" or "the weekly swing is bigger when the shop is bigger" are interactions. You must add them explicitly (an extra column such as e-mails × weekend) or change the model form (multiplicative seasonality).
"A large uncertainty on the trend slope means the forecast is very uncertain."
Not necessarily: if the slope and a regressor weight trade off, their errors cancel in the sum. Look at the uncertainty of the quantity you care about (the forecast), not only at individual weights.
In your forecasting model this is why a regressor that trends with time (cumulative users, a growing marketing budget) is risky: it competes with the piecewise trend for the same slow rise, and the Laplace prior on the slope changes $\delta_j$ and the prior on $\beta$ decide the split as much as the data do. When you present a component breakdown, check the posterior correlation between the trend weights and the regressor weights first, and say "the total is well determined; the split is uncertain" when that is the case (Chapter 6.8).
"The trend component is the true underlying growth of the business."
The trend component is the model's estimate of the slow part of the mean, after the other layers took their share, under its assumptions (piecewise-linear, chosen changepoints, chosen priors). It is a useful summary, not a measurement.
Model answer: "Each component is a contribution to the fitted mean. The sum is usually well determined. Individual components are interpretable when their columns are distinct; when, say, a regressor rises with time, the trend and that regressor trade off, so I check the posterior correlation before I interpret either one."
$\hat\mu_t = \hat g(t) + \hat s(t) + \hat h(t) + X_t\hat\beta$: each part is in units of $y$; a change in one layer moves $\mu_t$ by exactly that change (no interactions).
What-ifs and component plots come for free. The sum is usually well determined; the split needs distinct columns (identifiability).
Trap: overlapping layers (rising regressor vs trend, e-mails sent only on Saturdays vs the weekly pattern) make the split arbitrary.
Quick check: a holiday falls on a Sunday in every year of your data. Can the model estimate the holiday effect separately from the Sunday effect?
Only partly. The weekly layer learns "Sunday" from all the other Sundays (dozens per year), so the holiday column is still distinct from the Sunday columns: the holiday effect is the extra lift on top of a usual Sunday. What cannot be learned is whether the holiday would behave differently on a weekday: the model assumes the same additive effect on any weekday (no interaction), and the data never test that.
The whole model is one regression with structured columns core
The five layers look like five different machines. Look closer and each one is just a few columns of numbers, one row per day, each column with one weight. The trend is the columns "1" and "$t$" plus one ramp per changepoint. Weekly seasonality is six sine/cosine columns. Each holiday is a 0/1 column. Each regressor is its own column.
Put all those columns side by side and you have one big table, the design matrix $X$. The mean is that table times one long list of weights. That is a linear regression (Chapter 5.13). What makes it Bayesian is that every block of weights gets a prior (a belief about its size before seeing data), the noise gets a likelihood, and we learn a whole distribution of weights instead of one best set.
Three ways to say it:
- Picture: a wide table with colour-coded column groups (trend, changepoints, seasonality, holidays, regressors) times a long weight vector.
- Numbers: 2 + 25 + 20 + 6 + 12 + 3 = 68 columns, so 68 weights, plus one noise scale.
- Slogan: components are column groups; priors say how big each group's weights may be.
Count the columns of a typical daily model with two years of history.
- Base trend: an intercept column of 1s (weight $m$) and the time column $t$ (weight $k$): 2 columns.
- Changepoints: 25 candidate locations, one ramp column each (weights $\delta_1, \dots, \delta_{25}$): 25.
- Yearly seasonality with Fourier order 10: a sine and a cosine per harmonic, $2\times10 =$ 20.
- Weekly seasonality with order 3: $2\times3 =$ 6.
- Holidays: 4 holidays, each with a window of 3 days (day before, the day, day after), one column per day of the window: $4\times3 =$ 12.
- Regressors: e-mails, price discount, temperature: 3.
- Total: $2 + 25 + 20 + 6 + 12 + 3 = 68$ weights for the mean. With a Normal likelihood add one noise scale $\sigma$: $d = 69$ unknown parameters.
- What this means for inference (Chapter 6.13): a mean-field Gaussian guide has $2d = 138$ numbers to learn; a full-rank Gaussian guide has $d + d(d+1)/2 = 69 + 2415 = 2484$; a low-rank guide with rank 10 has $d(10 + 2) = 828$.
Stack the column groups into one $n\times p$ design matrix and the weights into one vector:
$$X = \big[\,\underbrace{1,\ t}_{\text{base trend}} \;\big|\; \underbrace{X_{\text{cp}}}_{\text{changepoint ramps}} \;\big|\; \underbrace{F_{\text{year}},\ F_{\text{week}}}_{\text{Fourier}} \;\big|\; \underbrace{X_{\text{hol}}}_{\text{0/1 flags}} \;\big|\; \underbrace{X_{\text{reg}}}_{\text{regressors}}\,\big], \qquad \theta = (m, k, \delta, \beta_{\text{season}}, \beta_{\text{hol}}, \beta_{\text{reg}}),$$ $$\mu = X\theta, \qquad y_t \sim \text{Likelihood}(\mu_t, \text{noise parameters}).$$- Each row $X_t$ is one day; each column is one feature; $\mu_t = X_t\theta$ is the mean of day $t$. The changepoint ramps make the trend continuous (why, and how they relate to Prophet's $a(t)$ and $\gamma$ form: Chapter 7.8).
- Fixed before fitting (not estimated): the changepoint locations, the seasonal periods and Fourier orders, the holiday dates and windows, and which regressors to use. That is what keeps the mean linear in the weights.
- Bayesian part: each block gets a prior, for example $\delta_j \sim \text{Laplace}(0, b)$ for slope changes (Chapter 7.10) and Normal priors for seasonal, holiday and regressor weights. The posterior is $p(\theta \mid y) \propto p(y \mid \theta)\,p(\theta)$. With a Normal likelihood, its MAP (most probable point) is a penalized least-squares fit: Normal priors act like a ridge (L2) penalty and Laplace priors like a lasso (L1) penalty (with another likelihood it is a penalized fit of that likelihood) (Chapter 5.3).
- With a Normal likelihood and Normal priors the posterior is Gaussian and could be written down exactly; with Laplace priors, a Student-t or a Negative Binomial likelihood it cannot, which is one reason to fit with SVI or MCMC (Chapter 6.12).
Why do we need it?
Seeing the model as one regression lets you reuse everything you know: collinearity checks, parameter counting, ridge/lasso intuition for priors, residual diagnostics. It also tells you the size of the inference problem, which decides how hard SVI is and which guide is affordable.
Where is it used?
Prophet (in its default linear-growth, additive mode its Stan model is exactly a linear mean with these blocks), Prophet-style NumPyro models like yours, dynamic regression baselines, marketing-mix models (adstock columns + priors), and Bayesian structural time-series models with a static regression part.
How is it used?
Build each block as an array of shape (n_days, n_columns) with NumPy, np.column_stack them, and in the model compute mu = X @ theta (jnp.dot in JAX), with one numpyro.sample per block giving its prior. Print X.shape and the rank before fitting.
"A Bayesian forecasting model with priors is something different from a regression."
Its mean is a regression with structured columns. The Bayesian parts are the priors on the weight blocks, the choice of likelihood, and the fact that you learn a posterior distribution instead of one point. All the regression intuitions (collinearity, column scaling, counting parameters) still apply.
"The model learns where the changepoints are and which periods to use."
In this design the changepoint locations, periods, Fourier orders, holiday windows and regressors are chosen before fitting. Only the weights (and noise parameters) are learned. Choices made from the same data (like PELT on the training series) carry uncertainty the posterior does not see (Chapter 7.9).
"With priors, extra columns are free."
Priors keep extra weights small, but every column is one more latent parameter: the guide grows (quadratically for a full-rank guide), SVI becomes slower and noisier, and overlapping columns make the split between components uncertain.
This is the cleanest way to describe your model's structure: "the mean is $X\theta$ with blocks for the base trend, one ramp per changepoint candidate (from the grid plus PELT), Fourier columns per seasonality, holiday indicators and exogenous regressors; Laplace priors on the slope changes, other priors on the other blocks; a Normal, Student-t or Negative Binomial likelihood; fitted with SVI." The column count is also the latent dimension that your automatic choice between a full-rank and a low-rank Gaussian guide depends on: count it for your own configuration with the formula in the example.
$\mu = X\theta$, $X = [1, t \mid \text{ramps} \mid \text{Fourier} \mid \text{flags} \mid \text{regressors}]$; $y_t \sim$ Lik$(\mu_t, \cdot)$; priors per block.
Columns: $2 + S + 2N_{\text{year}} + 2N_{\text{week}} + (\text{holiday columns}) + R$. Guide sizes: $2d$, $d(r+2)$, $d + d(d+1)/2$.
Trap: locations, periods, orders, windows are fixed in advance; only weights and noise parameters are learned.
Quick check: 15 changepoints, yearly order 6, weekly order 3, 5 one-day holidays, 2 regressors, Student-t likelihood (σ and ν). How many parameters?
Mean weights: $2 + 15 + 12 + 6 + 5 + 2 = 42$. Noise parameters: σ and ν, 2 more. Total $d = 44$. A full-rank Gaussian guide would learn $44 + 44\times45/2 = 44 + 990 = 1034$ numbers.
Priors on every block, and simulating the model before fitting it core
A Bayesian model is a recipe for making fake data. Read it top to bottom: pick a starting level, pick a slope, pick a few slope changes, pick a weekly pattern, pick holiday effects, add them up, add noise. Every "pick" is a draw from a prior (a distribution that says which values you find believable before seeing any data).
So before fitting, you can ask the model to dream: run the recipe many times and look at the series it makes. If the dreams look like plausible histories of your shop, the priors are sensible. If half of them crash to minus 2 000 orders or explode to a million, the priors are telling the model that such worlds are believable, and that will leak into your forecasts. This is a prior predictive check (taught in general in Chapter 6.2); here we apply it to the whole additive model.
Three ways to say it:
- Picture: the model draws imaginary sales charts before it has seen yours; you check whether they look like sales charts.
- Numbers: 25 slope changes with Laplace scale 2 orders/day each give a final slope with standard deviation about 14 orders per day: hundreds of orders of drift per month for a shop selling 150 a day.
- Slogan: simulate before you fit; if the dreams are absurd, fix the priors.
How many "big" slope changes does the prior $\delta_j \sim \text{Laplace}(0, b)$ expect among $S = 25$ candidates? Call a change big if $|\delta_j| > 0.5$ orders per day. For a Laplace distribution, $P(|\delta| > c) = e^{-c/b}$ (Chapter 4.9).
- Scale $b = 0.2$: $P(|\delta_j| > 0.5) = e^{-0.5/0.2} = e^{-2.5} \approx 0.082$.
- Expected number of big changes: $25 \times 0.082 \approx 2.05$. "About two real bends, the rest small": a sensible belief.
- The slope at the end of the history is $k + \sum_j \delta_j$. Each $\delta_j$ has variance $2b^2$, so the 25 changes add variance $25\times2b^2 = 50b^2$, a standard deviation of $b\sqrt{50} \approx 7.07\,b = 1.41$ orders per day.
- Scale $b = 2$: $P(|\delta_j| > 0.5) = e^{-0.25} \approx 0.78$, so about $25\times0.78 \approx 19.5$ big changes, and the final slope has standard deviation $7.07\times2 \approx 14.1$ orders per day.
- Over a 30-day horizon that slope alone moves the forecast by about $14.1\times30 \approx 424$ orders (one standard deviation), for a shop selling about 150 a day. The prior believes in absurd futures.
The generative story of the additive model (illustrative priors in original units; your code has its own):
$$\begin{aligned} m &\sim N(150, 20^2), \quad k \sim N(0, \sigma_k^2), \quad \delta_j \sim \text{Laplace}(0, b),\ j = 1..S,\\ \beta_{\text{season}} &\sim N(0, \sigma_s^2), \quad \beta_{\text{hol}} \sim N(0, \sigma_h^2), \quad \beta_{\text{reg}} \sim N(0, \sigma_r^2), \quad \sigma \sim \text{HalfNormal}(\sigma_0),\\ \mu_t &= X_t\theta, \qquad y_t \sim N(\mu_t, \sigma^2). \end{aligned}$$- $N(\mu, \sigma^2)$ is written with the variance; NumPyro's
dist.Normal(loc, scale)takes the standard deviation.dist.Laplace(loc, scale)takes $b$. HalfNormal is a Normal folded to positive values (for scales). - Prior predictive simulation: draw $\theta$ from the priors, compute $\mu$, draw $y$; repeat. The resulting series are what the model believes before seeing data.
- For comparison, Prophet uses (on its scaled data: $y$ divided by its maximum absolute value, time rescaled to $[0, 1]$ over the history) $k, m \sim N(0, 5^2)$, $\delta_j \sim \text{Laplace}(0, 0.05)$ (the default
changepoint_prior_scale), seasonal, holiday and regressor weights $\sim N(0, 10^2)$ (default prior scales 10), and $\sigma \sim \text{HalfNormal}(0.5)$. Because of the scaling, the same number means different things for different series (Chapter 7.10).
Why do we need it?
Priors on 25 slope changes, 26 Fourier weights and many holiday weights interact: each looks harmless on its own, but together they can describe absurd worlds. Simulating the whole model is the only easy way to see what the priors jointly believe.
Where is it used?
The Bayesian workflow before any fit (prior predictive checks), choosing the Laplace scale for changepoints, choosing holiday prior scales, checking that a Negative Binomial model with a log link does not produce billions of orders, and in Prophet-style NumPyro models with Predictive.
How is it used?
In NumPyro run Predictive(model, num_samples=500)(key, X) without passing y; plot 20 of the simulated series and compute a few numbers (share of negative values, range after one year, size of the weekly swing). Adjust the prior scales until the simulations cover plausible worlds and not much more.
"Wide priors are the safe, uninformative choice."
For one parameter, perhaps. For 25 slope changes, 26 Fourier weights and many holidays together, wide priors describe wild worlds, and with a short or noisy history the posterior inherits that wildness, especially in the forecast. Weakly informative priors that keep simulations plausible are the safer default (Chapter 6.2).
"With enough data the priors do not matter."
Some blocks never get much data: a holiday seen twice, a changepoint near the end of the history, a regressor that barely varied. For those the prior still decides a lot. Check prior sensitivity for exactly these blocks (Chapter 6.8).
"Prophet's default $b = 0.05$ is a universal good value."
It is set on Prophet's scaled data (y divided by its maximum, time from 0 to 1). The same number means a different slope change in orders per day for every series and every history length. Check what your own model's scale means in your own units.
In your forecasting model the slope changes have Laplace priors $\delta_j \sim \text{Laplace}(0, b)$: most candidates should stay near zero while a few make real bends. Simulating from the model is the quickest way to see what your $b$ (and your other prior scales) believe, in your units: count the big bends per simulated history, look at the spread of the final slope, and check the share of negative values (with a Normal likelihood) or absurdly large counts (with a Negative Binomial and a log link). Do the same check whenever you change how the data or time are scaled.
The model is a recipe for fake data: draw weights from priors → $\mu = X\theta$ → draw $y$. Running it without data = prior predictive check.
Laplace: $P(|\delta| > c) = e^{-c/b}$, $Var = 2b^2$; $S$ changes give final-slope sd $b\sqrt{2S}$.
Trap: many "harmless" wide priors together make absurd worlds; priors matter most for blocks with little data.
Quick check: with $S = 32$ candidates and $b = 0.25$ orders/day, what is the prior standard deviation of the total slope change $\sum_j\delta_j$?
Each $\delta_j$ has variance $2b^2 = 2\times0.0625 = 0.125$. The sum of 32 independent ones has variance $32\times0.125 = 4$, so the standard deviation is $\sqrt4 = 2$ orders per day ($= b\sqrt{2S} = 0.25\times8$).
Additive or multiplicative seasonality? core
A small shop sells 100 orders on an average day and 120 on Saturdays. A year later it has grown to 300 orders on an average day. What happens on Saturdays now?
- If Saturdays add a fixed number (+20 orders), Saturdays now sell 320. That is additive seasonality: the size of the weekly swing does not depend on the level.
- If Saturdays add a fixed percentage (+20%), Saturdays now sell 360. That is multiplicative seasonality: the swing grows with the level.
For a growing business the second is often more realistic: the busy day is busy in proportion to how big the business is. You can see which one you have in a plot: if the seasonal waves get taller as the series rises, the seasonality is multiplicative.
Three ways to say it:
- Picture: additive = waves of constant height riding the trend; multiplicative = waves that grow with the trend, like a funnel.
- Numbers: +20 orders at level 100 and +20 at level 300 (additive) versus +20 at 100 and +60 at 300 (multiplicative +20%).
- Slogan: additive adds orders; multiplicative adds percent.
Saturday effect +20%, trend 100 in week 1 and 300 in week 52, no noise.
- Multiplicative: $y = g(t)\,(1 + s(t))$. Week 1 Saturday: $100\times1.2 = 120$, a lift of 20. Week 52 Saturday: $300\times1.2 = 360$, a lift of 60.
- An additive model must use one lift for all Saturdays. Least squares picks something in between (say about +40): too big in week 1 (residual about $20 - 40 = -20$), too small in week 52 (residual about $60 - 40 = +20$). The residuals form a weekly wave that grows over time.
- Take logs of the multiplicative model: $\log y = \log g(t) + \log(1 + s(t))$. The Saturday term is now a constant $\log 1.2 \approx 0.182$ every week: on the log scale the seasonality is additive.
- So "multiplicative" can be handled either by a model of the form $g\,(1 + s)$, or by modelling $\log y$ additively (then remember that exponentiating a mean of logs gives a median, not a mean: Chapter 4.18).
Three ways to combine trend and seasonality:
$$\text{additive: } y_t = g(t) + s(t) + \epsilon_t, \qquad \text{multiplicative: } y_t = g(t)\,\big(1 + s(t)\big) + \epsilon_t, \qquad \text{log-additive: } \log y_t = \tilde g(t) + \tilde s(t) + \tilde\epsilon_t.$$- In additive mode $s(t)$ is in units of $y$ (orders); in multiplicative mode it is a fraction of the trend ($s = 0.2$ means +20%).
- Prophet offers both:
Prophet(seasonality_mode='multiplicative')computes $y = g(t)\,(1 + \text{multiplicative terms}) + \text{additive terms}$. Added seasonalities and extra regressors follow that mode by default and can be switched one by one withmode='additive'ormode='multiplicative'inadd_seasonality/add_regressor. - A log link does the same job automatically: if a count model's mean is $\mu_t = \exp(\eta_t)$ with $\eta_t$ = trend + seasonality + …, then every additive term in $\eta_t$ multiplies the mean: $\exp(a + b) = e^a e^b$ (Chapter 5.14). A softplus link, $\log(1 + e^{\eta})$, behaves almost like the identity when $\eta$ is large, so its terms stay roughly additive there.
Why do we need it?
Growing (or shrinking) series usually have seasonal swings that scale with the level. An additive model forces one swing size for the whole history, so it overshoots early, undershoots late, and its forecast swings are wrong exactly where the series is largest.
Where is it used?
Retail and e-commerce demand, web traffic, airline passengers (the classic multiplicative series), Prophet's seasonality_mode, Holt–Winters multiplicative (Chapter 7.5), log transforms before ARIMA, and count models with a log link (Poisson or Negative Binomial GLMs).
How is it used?
Plot the series: do the seasonal waves grow with the level? Fit both forms and compare the residuals over time (a growing wave or funnel means the form is wrong) and the holdout error. Or model $\log y$, or use a log link, and check what your back-transformation returns (median or mean).
"A growing series always needs multiplicative seasonality."
Only if the seasonal swing grows with the level. Some businesses grow while the weekend lift stays a fixed number of orders. Look at the plot and the residuals; do not decide from the trend alone.
"Model log y, then exponentiate the forecast and you have the mean."
$\exp(\text{mean of } \log y)$ is the median on the original scale (for a symmetric log-scale error), which is below the mean. For a Normal error with variance $\tilde\sigma^2$ on the log scale, the mean is $\exp(\tilde\mu + \tilde\sigma^2/2)$.
"Multiplicative mode makes the holidays and regressors multiplicative too, and that is always fine."
In Prophet they follow seasonality_mode unless you set mode= on each. A temperature effect may well be additive while the weekly swing is multiplicative. Decide per component.
Your model is written additively: $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$. Two things can still make it behave multiplicatively, so check which applies to your code: (1) a multiplicative option for seasonality like Prophet's, if you implemented one; (2) the link used for the Negative Binomial mean. If the NB mean is $\exp(\eta_t)$, every additive term in $\eta_t$ multiplies the mean (a +20% Saturday, a +35% holiday); if it is softplus, terms are close to additive in orders when the mean is large. With a Normal or Student-t likelihood on raw values the effects are additive in orders.
Additive: $y = g + s$ (swing in orders, constant); multiplicative: $y = g(1 + s)$ (swing in percent, grows with the level); $\log y$ additive ⇔ multiplicative on the original scale.
A log link $\mu = e^{\eta}$ turns additive terms in $\eta$ into multiplicative effects on $\mu$.
Trap: decide from the plot and residuals (growing wave = wrong form); exp of a log-scale mean is a median.
Quick check: in a Negative Binomial model with mean $\mu_t = \exp(\eta_t)$, the holiday weight in $\eta_t$ is 0.3. What does the holiday do to the expected orders?
It multiplies the mean by $e^{0.3} \approx 1.35$: about +35%, whatever the level. At a mean of 100 that is about +35 orders; at a mean of 300, about +105.
Forecasting with the additive model: continue every layer core
To forecast, the model does not "look at the last few days and guess the next one". It takes each layer and continues it into the future in the way that suits that layer:
- the trend line is extended with its last slope;
- the seasonal waves keep repeating;
- the holidays are read from next year's calendar;
- the regressors must be given to the model: planned e-mails are known, tomorrow's temperature must itself be forecast;
- the noise cannot be continued, only described: it becomes the width of the forecast band.
Three ways to say it:
- Picture: slide each transparent sheet to the right and keep drawing in its own style.
- Numbers: trend 190 + 9 days × 0.8 = 197.2, Tuesday −14, 2 thousand planned e-mails +8 → 191.2 orders, ± about 8 for the noise.
- Slogan: forecast = every layer continued, plus honest noise.
History up to day $T = 111$. The fitted trend is $\hat g(111) = 190$ orders with last slope 0.8 orders per day; the weekly effect for Tuesday is −14; the residual sd is $\hat\sigma = 6$. Forecast day 120.
- Horizon: $h = 120 - 111 = 9$ days. Day 120 is a Tuesday ($120 = 17\times7 + 1$, and day 0 was a Monday).
- Trend: continue the last slope: $\hat g(120) = 190 + 0.8\times9 = 190 + 7.2 = 197.2$.
- Seasonality: the same weekly pattern, Tuesday: $-14$ → $197.2 - 14 = 183.2$.
- Holidays: the calendar says day 120 is not a holiday: $+0$.
- Regressor: the marketing team plans 2 thousand e-mails that day: $2\times4 = +8$ → $\hat y_{120\mid111} = 191.2$ orders.
- Noise only, 80% band: $\pm z_{0.9}\hat\sigma = \pm1.2816\times6 \approx \pm7.7$, so about $[183.5,\ 198.9]$. This band ignores the uncertainty in the weights, in future trend changes and in the regressor, so a full Bayesian band is wider (Chapter 7.14).
The point forecast made at origin $T$ for horizon $h$ is the fitted mean evaluated at the future time:
$$\hat y_{T+h\mid T} = \hat g(T+h) + \hat s(T+h) + \hat h(T+h) + X_{T+h}\hat\beta.$$- $\hat g(T+h)$: the piecewise-linear trend beyond the last changepoint is a straight line with the final slope $k + \sum_j\delta_j$. Prophet also simulates possible future slope changes to widen the trend uncertainty (Chapter 7.8).
- $\hat s(T+h)$: the same Fourier columns, evaluated at the future $t$; $\hat h(T+h)$: the holiday columns from the future calendar.
- $X_{T+h}$: the regressor values on the future day. They are not produced by the model: they must be known in advance (planned promotions, prices, calendar features) or forecast separately (weather), and their error passes straight into the forecast (Chapter 7.12).
- The predictive distribution combines parameter uncertainty (the posterior over the weights) with observation noise from the likelihood: for each posterior draw, compute $\mu_{T+h}$, then draw $y_{T+h}$ from the likelihood (Chapter 7.14).
Why do we need it?
Knowing how each layer is continued tells you what the forecast can and cannot know: it will follow the calendar and the planned drivers perfectly, it will extend the trend in a straight line, and it will not react to yesterday's surprise or to a driver nobody supplied.
Where is it used?
model.make_future_dataframe + predict in Prophet, NumPyro's Predictive on a future design matrix, staffing and inventory plans built from demand forecasts, and scenario planning ("what if we send twice the e-mails?").
How is it used?
Build the design matrix for the future dates with exactly the same column definitions as in training (same changepoints, periods, orders, holiday windows, regressor scaling), fill in the future regressor values, then compute the mean (and posterior predictive draws for bands). Check which regressors are truly known at the forecast origin.
"The model forecasts the regressors too."
It does not. Future regressor values are an input. If they are not known at the forecast origin you must forecast them separately, use scenarios, or drop the regressor, and the error of that guess is added to the forecast error. A regressor that is only available after the fact is a leak (Chapter 7.12).
"The noise band is the forecast uncertainty."
It is only one part. The weights are uncertain (posterior), the trend may change again in the future, and the regressors may be guessed. A predictive distribution from posterior draws, plus future-trend and regressor uncertainty, is wider and widens with the horizon (Chapter 7.14).
"A good fit on the history means a good forecast."
The forecast depends most on the parts that are extrapolated: the final slope and the future regressor values. Those are exactly the parts the history checks least. Judge forecasts by rolling-origin backtests (Chapter 7.15).
In your forecasting model the future design matrix must be built exactly like the training one: the same changepoint locations (the grid plus PELT ones, all inside the history), the same Fourier periods and orders, the same holiday windows, and the same regressor scaling if you scale them (the scaler fitted on the training period only). For every exogenous regressor, be ready to say whether its future values are known at the origin, forecast, or scenario inputs, and how that uncertainty is (or is not) included in your intervals.
"The model learns from yesterday's value, so it adapts quickly to a sudden change."
An additive Prophet-style model is a regression on time and known inputs. It does not use $y_{t-1}$ unless you add lagged values as regressors. Yesterday's surprise does not move tomorrow's forecast; the model adapts only when refitted, and then only through its layers (a new changepoint, a new holiday weight).
Model answer: "My model's mean is $g(t) + s(t) + h(t) + X_t\beta$: a function of time, the calendar and known drivers. Forecasting continues each layer. It does not have an autoregressive part, so short-term persistence stays in the residuals; I check their autocorrelation, and if it matters I'd add an AR term or compare with a model that has one (Chapters 7.3, 7.17)."
$\hat y_{T+h\mid T} = \hat g(T+h) + \hat s(T+h) + \hat h(T+h) + X_{T+h}\hat\beta$: trend extended with its last slope, seasonality repeated, holidays from the calendar, regressors supplied from outside.
Build the future design exactly like the training one.
Trap: the model does not forecast regressors and does not use yesterday's value; the noise band is not the whole uncertainty.
Quick check: at origin $T$ the trend is 200 with final slope −0.5 per day; a Saturday 14 days later has weekly effect +22 and is a holiday (+35); no regressors. What is the point forecast?
Trend: $200 - 0.5\times14 = 193$. Add Saturday and the holiday: $193 + 22 + 35 = 250$ orders.
Recap, cheat sheet and practice
- The model: $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$; the mean $\mu_t$ is everything except the noise; every term is in the units of $y$; seasonality averages to about zero over a period.
- Jobs: trend = slow level (piecewise-linear); seasonality = fixed-period repeats (Fourier); holidays = known dates (indicators); regressors = outside drivers (one weight each); noise = the rest (the likelihood). A missing or wrong layer leaves its pattern in the residuals.
- Separately: additive layers can be read, plotted and changed one at a time (what-ifs), with no interactions; the sum is usually well determined, the split only when the layers' columns are distinct (identifiability).
- One regression: $\mu = X\theta$ with column groups $[1, t \mid \text{ramps} \mid \text{Fourier} \mid \text{flags} \mid \text{regressors}]$; locations, periods, orders and windows are fixed in advance; priors per block (Laplace on slope changes), a chosen likelihood, fitted by SVI.
- Generative: draw weights from the priors, compute $\mu$, draw $y$: a prior predictive check shows what the priors jointly believe.
- Additive vs multiplicative: $g + s$ (fixed swing) vs $g(1 + s)$ (swing in percent); log scale or a log link makes effects multiplicative.
- Forecasting: continue each layer (last slope, repeated seasons, calendar, supplied regressors); the model does not use yesterday's value; the noise band is only part of the uncertainty.
Cheat sheet
| Idea | Formula | Remember |
|---|---|---|
| Additive model | $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$ | a receipt with five lines, all in units of $y$ |
| Mean | $\mu_t = g(t) + s(t) + h(t) + X_t\beta$; $y_t \sim \text{Lik}(\mu_t, \cdot)$ | $\epsilon_t$ unobservable; residual $e_t = y_t - \hat\mu_t$ |
| Design matrix | $\mu = X\theta$, $X = [1, t \mid (t - s_j)_+ \mid \sin, \cos \mid \text{flags} \mid X_{\text{reg}}]$ | linear in θ; columns fixed in advance |
| Column count | $2 + S + 2N_{\text{year}} + 2N_{\text{week}} + \#\text{holiday cols} + R$ | plus noise parameters (σ, ν or α) |
| Guide sizes | mean-field $2d$; low-rank $d(r + 2)$; full-rank $d + d(d+1)/2$ | Chapter 6.13 |
| Laplace prior | $P(|\delta| > c) = e^{-c/b}$, $Var = 2b^2$ | $S$ changes: final-slope sd $b\sqrt{2S}$ |
| Multiplicative | $y = g(1 + s)$; $\log y = \log g + \log(1 + s)$ | swing grows with the level |
| Log link | $\mu = e^{\eta}$: weight $w$ multiplies $\mu$ by $e^{w}$ | 0.3 → ×1.35 |
| Forecast | $\hat y_{T+h\mid T} = \hat g(T+h) + \hat s(T+h) + \hat h(T+h) + X_{T+h}\hat\beta$ | regressors must be supplied |
import numpy as np
import jax, jax.numpy as jnp
import numpyro, numpyro.distributions as dist
from numpyro.infer import SVI, Trace_ELBO, Predictive, init_to_median
from numpyro.infer.autoguide import AutoNormal
# 1) Simulate 16 weeks of a shop (day 0 is a Monday): trend + weekly + holidays + e-mails + noise
rng = np.random.default_rng(0)
n = 140; t = np.arange(n, dtype=float) # days 0..111 = history, 112..139 = future
week = np.array([-10, -14, -12, -6, 4, 22, 16.0]) # Mon..Sun, sums to 0
hol_days = [19, 47, 96, 124]
emails = np.zeros(n); a = 0.0
for i in range(n): # thousands of e-mails: smooth and positive
a = 0.8 * a + rng.normal(0, 0.6); emails[i] = max(0.0, 2 + a)
trend = 110 + 0.6 * t - 0.45 * np.maximum(0, t - 70)
y_all = trend + week[t.astype(int) % 7] + 35 * np.isin(t, hol_days) + 4 * emails + rng.normal(0, 5, n)
T = 112; y = y_all[:T]
# 2) Build the design matrix block by block (same function for history and future)
def fourier(t, period, order):
return np.column_stack([f(2 * np.pi * k * t / period) for k in range(1, order + 1) for f in (np.sin, np.cos)])
def blocks(t, x):
return {"base": np.column_stack([np.ones(len(t)), t]),
"changept": np.maximum(0, t - 70)[:, None],
"weekly": fourier(t, 7, 3),
"holidays": np.isin(t, hol_days).astype(float)[:, None],
"e-mails": x[:, None]}
B = blocks(t[:T], emails[:T])
X = np.column_stack(list(B.values()))
print(X.shape, np.linalg.matrix_rank(X)) # (112, 11) 11
# 3) Least squares, then split the fitted mean into components (one receipt per day)
theta, *_ = np.linalg.lstsq(X, y, rcond=None)
parts, j = {}, 0
for name, Bk in B.items():
parts[name] = Bk @ theta[j:j + Bk.shape[1]]; j += Bk.shape[1]
d = 47 # a Saturday (47 % 7 = 5) and a holiday
print({k: round(float(v[d]), 1) for k, v in parts.items()}, round(float(sum(v[d] for v in parts.values())), 1), round(float(y[d]), 1))
# {'base': 138.1, 'changept': 0.0, 'weekly': 20.9, 'holidays': 32.0, 'e-mails': 17.7} 208.8 220.4
print(np.round(parts["weekly"][:7], 1)) # [-10.9 -14.3 -8.6 -5.4 2.3 20.9 16.] (truth -10 -14 -12 -6 4 22 16)
# 4) Remove one layer at a time: the leftovers grow (true noise sd 5)
for drop in ["changept", "weekly", "holidays", "e-mails"]:
Xd = np.column_stack([Bk for k, Bk in B.items() if k != drop])
r = y - Xd @ np.linalg.lstsq(Xd, y, rcond=None)[0]
print(f"without {drop:8s} residual sd = {r.std(ddof=Xd.shape[1]):.2f}")
# changept 6.17 · weekly 13.46 · holidays 7.14 · e-mails 6.42 (the full model: about 5)
# 5) The Bayesian version: a prior per block, a Normal likelihood, fitted by SVI
def model(t, F, hol, x, y=None):
m = numpyro.sample("m", dist.Normal(150, 50)) # level at t = 0 (orders)
k = numpyro.sample("k", dist.Normal(0, 1)) # base slope (orders/day)
delta = numpyro.sample("delta", dist.Laplace(0, 0.5)) # slope change at day 70
beta_s = numpyro.sample("beta_s", dist.Normal(0, 20).expand([F.shape[1]]).to_event(1))
beta_h = numpyro.sample("beta_h", dist.Normal(0, 50)) # holiday effect
beta_x = numpyro.sample("beta_x", dist.Normal(0, 10)) # orders per thousand e-mails
sigma = numpyro.sample("sigma", dist.HalfNormal(20)) # noise sd (NumPyro takes the sd)
mu = m + k * t + delta * jnp.maximum(0, t - 70) + F @ beta_s + beta_h * hol + beta_x * x
numpyro.sample("y", dist.Normal(mu, sigma), obs=y)
args = (jnp.array(t[:T]), jnp.array(B["weekly"]), jnp.array(B["holidays"][:, 0]), jnp.array(emails[:T]))
guide = AutoNormal(model, init_loc_fn=init_to_median)
svi = SVI(model, guide, numpyro.optim.Adam(0.05), Trace_ELBO())
res = svi.run(jax.random.PRNGKey(0), 6000, *args, y=jnp.array(y), progress_bar=False)
post = guide.sample_posterior(jax.random.PRNGKey(1), res.params, sample_shape=(2000,))
for name in ["k", "delta", "beta_h", "beta_x", "sigma"]:
print(name, round(float(post[name].mean()), 3), "+/-", round(float(post[name].std()), 3))
# k 0.62 +/- 0.008 · delta -0.455 +/- 0.035 · beta_h 31.598 +/- 3.308 · beta_x 3.948 +/- 0.158 · sigma 5.56 +/- 0.349
print("least squares:", np.round(theta[[1, 2, 9, 10]], 3)) # [0.611 -0.479 31.965 3.898] (truth 0.6, -0.45, 35, 4)
# 6) Prior predictive check: let the model dream before seeing data
prior = Predictive(model, num_samples=500)(jax.random.PRNGKey(2), *args)["y"]
print("share of dreams outside [0, 400] on some day:", round(float(((prior < 0) | (prior > 400)).any(axis=1).mean()), 3))
# 0.424: k ~ N(0, 1) allows +-100 orders of drift over 112 days; tighten it if that is implausible
# 7) Forecast days 112..139: same design, planned e-mails supplied, posterior predictive draws
Bf = blocks(t[T:], emails[T:])
fargs = (jnp.array(t[T:]), jnp.array(Bf["weekly"]), jnp.array(Bf["holidays"][:, 0]), jnp.array(emails[T:]))
pred = Predictive(model, guide=guide, params=res.params, num_samples=2000)(jax.random.PRNGKey(3), *fargs)["y"]
lo, hi = np.percentile(np.asarray(pred), [10, 90], axis=0)
print("MAE:", round(float(np.mean(np.abs(np.asarray(pred).mean(0) - y_all[T:]))), 2),
"80% coverage:", round(float(np.mean((y_all[T:] >= lo) & (y_all[T:] <= hi))), 2))
# MAE: 5.32 80% coverage: 0.89 (28 days only, so coverage is noisy)
1. After fitting, the residuals are plain noise except on four dates where they jump by about +60. Which layer should you look at first?
2. Your series triples over two years and the weekend peaks look three times taller at the end than at the start. Which form fits best?
3. In your forecasting model, which of these is not learned when the model is fitted?
4. A model uses daily temperature as a regressor. What is needed to forecast next week?
5. A Negative Binomial model has mean $\mu_t = \exp(\eta_t)$ and the holiday weight in $\eta_t$ is 0.2. On a holiday the expected orders are…
6. What does a prior predictive check of the full forecasting model tell you?
Practice problems
A. Trend $g(t) = 80 + 1.5t$. Day 30 is a Sunday (weekly effect +16), a holiday (+30), with a price-discount regressor $x = 0.5$ and weight $\beta = 20$. The shop received 195 orders. Write the receipt and the noise.
- Trend: $80 + 1.5\times30 = 80 + 45 = 125$.
- Weekly +16, holiday +30, regressor $0.5\times20 = 10$.
- Mean: $125 + 16 + 30 + 10 = 181$ orders.
- Noise: $\epsilon_{30} = 195 - 181 = +14$ (a lucky day, or a sign that one layer is too small: check other holidays and Sundays).
B. A model has 20 changepoint candidates, yearly order 8, weekly order 3, 6 holiday columns, 4 regressors and a Negative Binomial likelihood (concentration α). Count the parameters and the size of a full-rank and a rank-5 Gaussian guide.
- Mean weights: $2 + 20 + 16 + 6 + 6 + 4 = 54$.
- Noise: the NB concentration α: 1. So $d = 55$.
- Full-rank: $d + d(d+1)/2 = 55 + 55\times56/2 = 55 + 1540 = 1595$ numbers.
- Low-rank, $r = 5$: $d(r + 2) = 55\times7 = 385$ numbers. Mean-field: $2d = 110$.
C. (Interview) "Walk me through the components of your forecasting model and what each one assumes."
"The mean is a sum of four layers and the noise is a likelihood. The trend is piecewise-linear: a base slope plus slope changes at candidate changepoints from a Prophet-like grid and PELT, with Laplace priors so most changes stay near zero; it assumes straight segments and extrapolates the last slope. Seasonality is Fourier terms per period; it assumes a known period and a stable shape. Holidays are indicator columns; they assume known dates and the same effect each time. Regressors are linear terms; they assume a constant linear effect and that the values are available when I forecast, without leakage. The noise is Normal, Student-t or Negative Binomial depending on the data, assuming independence given the mean. Because everything adds up, I can explain any forecast line by line, and I check the residuals for whatever a layer missed."
D. With 25 candidates and $\delta_j \sim \text{Laplace}(0, 0.1)$, how many slope changes larger than 0.3 in size does the prior expect, and what is the prior sd of the final slope change $\sum\delta_j$?
- $P(|\delta_j| > 0.3) = e^{-0.3/0.1} = e^{-3} \approx 0.0498$.
- Expected count: $25\times0.0498 \approx 1.24$.
- Variance of the sum: $25\times2\times0.1^2 = 0.5$; sd $= \sqrt{0.5} \approx 0.71$ orders per day ($= 0.1\sqrt{50}$).
E. Multiplicative data: trend 50 in week 1 and 250 in week 40, Saturday +30%. You fit an additive model whose single Saturday lift comes out at about 45. What do the Saturday residuals look like?
- True Saturday lift: $50\times0.3 = 15$ in week 1, $250\times0.3 = 75$ in week 40.
- Residuals: week 1 about $15 - 45 = -30$; week 40 about $75 - 45 = +30$.
- So Saturday residuals climb from negative to positive over time (and other weekdays show the mirror pattern): a weekly wave that grows. Switch to $g(1 + s)$ or model $\log y$.
F. (Interview) "You added 'cumulative registered users' as a regressor and the trend became almost flat. What happened, and is the forecast wrong?"
"Cumulative users rises steadily, just like the trend, so the two layers compete for the same slow rise: the regressor took the credit and the trend's slopes shrank (helped by the Laplace prior pulling the $\delta_j$ to zero). The fitted total can be just as good, so the in-sample forecast is not necessarily wrong, but the component breakdown is no longer interpretable and the posterior of the two will be strongly correlated. The forecast now depends on future values of cumulative users, which I would have to forecast myself. I would check the posterior correlation, compare holdout errors with and without the regressor, and prefer the simpler model unless the regressor adds information beyond time."
Piecewise-linear trends and changepoints
A business does not grow in one straight line forever. It grows, then a competitor arrives and growth slows, then a new product makes it speed up again. The trend layer of your model handles this with a line that is allowed to bend at a few moments, the changepoints. This chapter builds that bending line from scratch: slopes and intercepts, slope changes $\delta_j$, the offsets $\gamma_j = -s_j\delta_j$ that keep it connected, the changepoint matrix $A(t)$, how many changepoints to allow, how the trend is extended into the future, and the three ways to decide where the bends may happen.
- Read $g(t) = kt + m$: the slope $k$, the intercept $m$, their units, and what extrapolation means
- Add changepoints $s_j$ with slope changes $\delta_j$, so the slope goes $k \to k + \delta_1 \to k + \delta_1 + \delta_2 \to \cdots$
- Derive the continuity offsets $\gamma_j = -s_j\delta_j$ and the equivalent "hinge" form $kt + m + \sum_j \delta_j (t - s_j)_+$
- Build the changepoint matrix $A(t)$ and the trend as a matrix formula, the way Prophet does and the way a JAX model can
- Say why trends change, and what a continuous piecewise-linear trend can and cannot represent (slope changes yes, sudden level jumps no)
- Explain the risk of too many or too few changepoints, and why the end of the history matters most for the forecast
- Compare the three approaches: a fixed grid (Prophet's defaults), data-driven detection (PELT), and Bayesian latent changepoints
The straight trend $g(t) = kt + m$: slope and intercept core
Imagine walking up a long, even ramp. Two numbers describe it completely: where you start (your height at the bottom) and how steep it is (how much you rise per step). Knowing those two, you can say how high you will be after any number of steps, even steps you have not taken yet.
A straight trend is that ramp. The intercept $m$ is the level at time zero; the slope $k$ is how many orders the level gains per day. The trend ignores the weekly ups and downs and the holidays (those are other layers, Chapter 7.7): it describes only the slow, underlying level.
Forecasting with a straight trend means extrapolating: walking further up the same ramp, past the last day you have seen.
Three ways to say it:
- Picture: a ramp: a starting height and a steepness.
- Numbers: start at 100 orders, gain 2 per day: day 30 is $100 + 2\times30 = 160$, day 45 (a forecast) is 190.
- Slogan: intercept = where you start; slope = how fast you climb.
A shop's underlying level (weekly pattern removed) is $g(t) = 2t + 100$ orders, with $t$ in days since 1 March.
- Intercept: $m = 100$ orders, the level on 1 March ($t = 0$).
- Slope: $k = 2$ orders per day. Units matter: $k$ is "orders per day", $m$ is "orders".
- Day 30: $g(30) = 2\times30 + 100 = 60 + 100 = 160$ orders.
- Forecast for day 45 (15 days after the last data on day 30): $g(45) = 90 + 100 = 190$ orders. Extrapolation continues the line.
- If you count time from day 30 instead ($t' = t - 30$), the same line is $g = 2t' + 160$: the slope stays 2, but the intercept becomes 160. The intercept depends on where time zero is; the slope does not.
A linear trend is
$$g(t) = k\,t + m,$$- $t$: time (days, or time rescaled to run from 0 to 1 over the history, as Prophet does);
- $k$: the slope or growth rate, the change in $g$ per unit of time ($g(t+1) - g(t) = k$), in units of $y$ per unit of time;
- $m$: the intercept or offset, the value at $t = 0$, in units of $y$.
As columns of a design matrix (Chapter 7.7): a column of 1s with weight $m$ and the column $t$ with weight $k$. Fitting by least squares gives $\hat k = S_{ty}/S_{tt}$ and $\hat m = \bar y - \hat k\,\bar t$ (Chapter 5.13). Extrapolation means evaluating $g$ beyond the last observed time; it assumes the slope stays the same.
Why do we need it?
The slow level of a series moves, and a forecast must say where it is heading. A line is the simplest honest summary of "where it is" and "how fast it moves", and it is the building block that changepoints bend.
Where is it used?
Prophet's growth='linear' trend (the base of your model's trend), Holt's linear trend (Chapter 7.5), drift forecasts, regression with a time column, and capacity planning ("at 2 more orders a day, when do we hit 300?").
How is it used?
Put a column of 1s and a column of $t$ in the design matrix and give their weights priors; read $k$ as the daily growth. Check units and the time origin before interpreting $m$; if time is rescaled to 0–1, convert $k$ back to "per day" by dividing by the history length.
"The intercept $m$ is the shop's starting level, so it has a business meaning."
It is the value at $t = 0$, wherever you put time zero. Rescale or shift time (Prophet rescales the history to run from 0 to 1) and $m$ changes while the line does not. Interpret the slope; treat $m$ as bookkeeping.
"A slope fitted on the history will hold in the future."
Extrapolation assumes it will. The further ahead you forecast, the more a small slope error (or a future change of slope) matters: an error of 0.5 orders per day becomes 45 orders after 90 days.
Your model's trend starts from exactly this line: a base slope $k$ and an offset $m$, each with a prior. If your code rescales time (for example to 0–1 over the history, as Prophet does) or standardizes $y$, the fitted $k$ is "change per whole history in scaled units": multiply by the scale of $y$ and divide by the history length to report it in orders per day.
$g(t) = kt + m$: $k$ = change per unit time (orders/day), $m$ = value at $t = 0$.
Extrapolation continues the last line; slope errors grow linearly with the horizon.
Trap: $m$ depends on where time zero is (and on rescaling); $k$ depends on the time units.
Quick check: Prophet-style scaled time runs from 0 (first day) to 1 (day 364 of a one-year history), and $y$ was divided by 500. The fitted scaled slope is 0.4. What is the slope in orders per day?
Over the whole history (364 days) the scaled level rises by 0.4, i.e. $0.4\times500 = 200$ orders. Per day: $200/364 \approx 0.55$ orders per day.
Changepoints and slope changes: $k \to k + \delta_1 \to k + \delta_1 + \delta_2$ core
Now the ramp has bends. You walk up at one steepness, then at a certain moment the ramp gets steeper, later it flattens. The moments where the steepness changes are the changepoints $s_1, s_2, \dots$ The amount by which the steepness changes at each one is the slope change (or slope adjustment) $\delta_j$: positive means "steeper from now on", negative means "flatter".
The slopes add up as you pass the changepoints: start with $k$; after the first changepoint the slope is $k + \delta_1$; after the second, $k + \delta_1 + \delta_2$; and so on. The line itself stays connected: the ramp bends, it does not break.
Three ways to say it:
- Picture: a ramp with hinges; at each hinge the steepness changes, but there is no step.
- Numbers: 2 orders/day, then $+3$ at day 20 (now 5/day), then $-4$ at day 50 (now 1/day).
- Slogan: a changepoint changes the speed, not the position.
Base slope $k = 2$ orders/day, intercept $m = 100$. Changepoint $s_1 = 20$ with $\delta_1 = +3$ (a new product launches), changepoint $s_2 = 50$ with $\delta_2 = -4$ (a competitor opens).
- Slopes: before day 20: $k = 2$; from day 20: $k + \delta_1 = 2 + 3 = 5$; from day 50: $k + \delta_1 + \delta_2 = 5 - 4 = 1$ orders per day.
- Level on day 20: $100 + 2\times20 = 140$.
- Level on day 50: start from 140 and climb 5 per day for 30 days: $140 + 5\times30 = 290$.
- Level on day 60: start from 290 and climb 1 per day for 10 days: $290 + 10 = 300$.
- Forecast for day 90 (the last slope continues): $300 + 1\times30 = 330$.
A piecewise-linear trend with changepoints $s_1 \lt s_2 \lt \dots \lt s_S$ and slope changes $\delta_1, \dots, \delta_S$:
- A changepoint $s_j$ is a time at which the slope is allowed to change. A segment is the stretch between two neighbouring changepoints.
- The slope at time $t$ is the base slope plus every change already passed: $$\text{slope}(t) = k + \sum_{j:\ s_j \le t} \delta_j.$$
- On each segment the trend is a straight line with that slope, and consecutive segments meet at the changepoints (the trend is continuous; how this is guaranteed is the next section).
- $\delta_j \gt 0$: growth speeds up; $\delta_j \lt 0$: growth slows (or turns into decline); $\delta_j = 0$: the changepoint does nothing.
The final slope, which the forecast uses, is $k + \sum_{j=1}^{S}\delta_j$.
Why do we need it?
Real growth rates change: launches, price changes, saturation, competitors, a pandemic. One straight line either misses those turns or is dragged by old history. Changepoints let the slope adapt while every segment stays a simple line.
Where is it used?
Prophet's linear and logistic trends, your model's trend, segmented ("broken-stick") regression in economics and epidemiology, piecewise-linear features in gradient-boosted models, and regression splines of degree 1 (linear splines).
How is it used?
Choose candidate locations $s_j$ (a grid, a detector such as PELT, or both), give each a slope change $\delta_j$ with a shrinkage prior (Laplace in your model), and let the fit decide which changes are real. Read the fitted slope plot (a step function) to explain when growth changed.
"$\delta_j$ is the new slope after changepoint $j$."
$\delta_j$ is the change in slope. The new slope is $k + \delta_1 + \dots + \delta_j$. A large positive $\delta_2$ after a large negative $\delta_1$ may just bring the slope back to where it was.
"A changepoint is a jump in the level."
In this model a changepoint changes the speed (slope) and the trend stays connected. A sudden step in the level (a tracking change that doubles the counts overnight) is something else (see why trends change).
In your forecasting model, the changepoint slopes $\delta_j$ are exactly these step heights, and each one has a Laplace prior $\delta_j \sim \text{Laplace}(0, b)$ so that most candidates keep $\delta_j \approx 0$ and only a few make a real bend (Chapter 7.10). When you explain a fitted trend, plot the slope staircase $k + \sum_{s_j \le t}\delta_j$: it shows when and by how much growth changed.
slope$(t) = k + \sum_{s_j \le t}\delta_j$; final slope $k + \sum_j\delta_j$ (used for the forecast).
$\delta_j$ = change of slope at $s_j$ (not the new slope); the line stays connected.
Trap: changepoints change speed, not position: level jumps need something else.
Quick check: $k = 1.5$, $\delta_1 = -2$ at day 30, $\delta_2 = +1$ at day 70, $m = 200$. What are the three slopes and the level on day 80?
Slopes: 1.5, then $1.5 - 2 = -0.5$, then $-0.5 + 1 = 0.5$. Level: day 30: $200 + 1.5\times30 = 245$; day 70: $245 - 0.5\times40 = 225$; day 80: $225 + 0.5\times10 = 230$ orders.
Keeping the trend connected: deriving $\gamma_j = -s_j\delta_j$ core
Here is a trap. Suppose you write the trend after a changepoint as "the new slope times $t$, plus the old intercept": $(k + \delta_1)\,t + m$. Before the changepoint it was $kt + m$. At the changepoint $t = s_1$ the two disagree by $\delta_1 s_1$: the line jumps. And the jump is bigger the later the changepoint happens, which makes no sense: why should a slope change on day 300 cause a bigger jump than the same change on day 10? It only happens because the new slope is measured from time zero.
The fix: each time the slope changes, also shift the intercept by exactly the amount that cancels the jump. That shift is the offset adjustment $\gamma_j$. One line of algebra shows it must be $\gamma_j = -s_j\delta_j$.
Three ways to say it:
- Picture: when you tilt a stick at a hinge, you must hold the hinge in place; $\gamma_j$ is the hand that holds it.
- Numbers: slope change $+3$ at day 20 would jump by $3\times20 = 60$; adding $\gamma_1 = -60$ to the intercept removes it.
- Slogan: new slope, same position: pay for the tilt with an offset of $-s_j\delta_j$.
$k = 2$, $m = 100$, one changepoint $s_1 = 20$ with $\delta_1 = +3$.
- Just before day 20: $g = 2\times20 + 100 = 140$.
- Naive formula just after: $(2 + 3)\times20 + 100 = 100 + 100 = 200$. A jump of $200 - 140 = 60 = s_1\delta_1 = 20\times3$.
- Add the offset $\gamma_1 = -s_1\delta_1 = -60$ to the intercept after day 20: $(2 + 3)\times20 + (100 - 60) = 100 + 40 = 140$. Connected.
- Day 30: $5\times30 + 40 = 190$, the same as "140 plus 10 days at 5 per day" $= 140 + 50 = 190$.
- Two changepoints ($\delta_2 = -4$ at $s_2 = 50$): $\gamma_2 = -50\times(-4) = +200$. Day 60: slope $2 + 3 - 4 = 1$, intercept $100 - 60 + 200 = 240$, so $g(60) = 60 + 240 = 300$, matching the previous section.
Derivation. Write the trend with a slope and an intercept that both change at the changepoints:
$$g(t) = \Big(k + \sum_{j:\ s_j \le t}\delta_j\Big)\,t + \Big(m + \sum_{j:\ s_j \le t}\gamma_j\Big).$$- Just before $s_j$ the line is $(k + K_{j-1})\,t + (m + G_{j-1})$, where $K_{j-1} = \sum_{i \lt j}\delta_i$ and $G_{j-1} = \sum_{i \lt j}\gamma_i$ collect the earlier changes.
- At $s_j$ the slope gains $\delta_j$ and the intercept gains $\gamma_j$: $(k + K_{j-1} + \delta_j)\,t + (m + G_{j-1} + \gamma_j)$.
- Continuity at $t = s_j$: the two lines must give the same value there: $(k + K_{j-1})s_j + m + G_{j-1} = (k + K_{j-1} + \delta_j)s_j + m + G_{j-1} + \gamma_j$.
- Everything cancels except $0 = \delta_j s_j + \gamma_j$, so $\boxed{\gamma_j = -s_j\,\delta_j}$.
- Substitute back: after $s_j$ the two new terms are $\delta_j t - s_j\delta_j = \delta_j(t - s_j)$. Summing over passed changepoints gives the hinge form $$g(t) = kt + m + \sum_{j=1}^{S}\delta_j\,(t - s_j)_+, \qquad (u)_+ = \max(u, 0).$$
This is the trend of Prophet (Taylor & Letham's paper writes it with an indicator vector $a(t)$, next section) and of your model. Each $(t - s_j)_+$ is a ramp ("hinge") that is 0 before $s_j$ and grows by 1 per day after it.
Why do we need it?
Without the offsets, the trend would jump at every changepoint by an amount that depends on where time zero happens to be, an artefact with no business meaning. The offsets make a slope change mean only "the speed changed here".
Where is it used?
Prophet's piecewise_linear trend and its Stan model, your model's trend, linear regression splines (the truncated power basis $(t - s_j)_+$), and segmented regression with continuity constraints.
How is it used?
Either compute $\gamma = -s \odot \delta$ and use the $a(t)$ form, or build the ramp columns $(t - s_j)_+$ directly and use them as regression columns with weights $\delta_j$. Both give identical trends; check it numerically once in your code.
"The offsets $\gamma_j$ are extra parameters the model must learn."
They are not free: each $\gamma_j$ is fixed by $\delta_j$ and $s_j$ as $-s_j\delta_j$. The only learned quantities in the trend are $k$, $m$ and the $\delta_j$.
"The jump without $\gamma$ is a real effect, so leaving it in adds flexibility."
The jump's size depends on where you put time zero, so it has no meaning. If you want real level shifts, model them on purpose with a step column (next sections).
In your forecasting model the trend is continuous, so either your code computes the offsets $\gamma_j = -s_j\delta_j$ (Prophet's form) or it uses the ramp columns $(t - s_j)_+$ (the hinge form); they are the same function. A one-line numerical check (both forms on the same $t$, $k$, $m$, $\delta$, $s$) is a good unit test, and the derivation above is a classic interview question: "why is $\gamma_j = -s_j\delta_j$?".
"γ is there to make the model fit better."
γ is there to make the trend continuous: it cancels the jump $s_j\delta_j$ that appears when a slope change is applied to a line measured from time zero.
Model answer: "At changepoint $s_j$ the slope gains $\delta_j$. If the intercept stayed the same, the line would jump by $s_j\delta_j$. Requiring the values just before and after $s_j$ to match gives $\delta_j s_j + \gamma_j = 0$, so $\gamma_j = -s_j\delta_j$, and the trend becomes $kt + m + \sum_j\delta_j(t - s_j)_+$."
Continuity at $s_j$: $\delta_j s_j + \gamma_j = 0$ ⇒ $\gamma_j = -s_j\delta_j$.
$g(t) = (k + \sum_{s_j \le t}\delta_j)\,t + (m + \sum_{s_j \le t}\gamma_j) = kt + m + \sum_j\delta_j(t - s_j)_+$.
Trap: without γ the trend jumps by $s_j\delta_j$, a size that depends on where time zero is; γ is not a free parameter.
Quick check: $s = (10, 40)$, $\delta = (0.5, -1.2)$. Compute $\gamma$, and check that $kt + m + \sum\delta_j(t - s_j)_+$ and the $\gamma$ form agree at $t = 50$ for $k = 1$, $m = 20$.
$\gamma_1 = -10\times0.5 = -5$, $\gamma_2 = -40\times(-1.2) = +48$. γ form: slope $1 + 0.5 - 1.2 = 0.3$, intercept $20 - 5 + 48 = 63$, $g(50) = 15 + 63 = 78$. Hinge form: $50 + 20 + 0.5\times40 - 1.2\times10 = 70 + 20 - 12 = 78$. They agree.
The changepoint matrix $A(t)$: the whole trend in one matrix formula core
To compute the trend for every day at once, a computer needs a simple bookkeeping table: for each day, which changepoints have already happened? That table is the changepoint matrix $A$. It has one row per day and one column per changepoint, with a 1 if that changepoint is in the past (or today) and a 0 if it is still in the future.
Each row of $A$ looks like a staircase of 1s that gets longer as time goes on: early days have all 0s, late days have all 1s. With this table, "add up the slope changes that already happened" is just "row of $A$ times $\delta$", and the whole trend becomes one line of matrix code.
Three ways to say it:
- Picture: a calendar where you tick off each changepoint once it has passed.
- Numbers: changepoints on days 2 and 4: day 3's row is $(1, 0)$, day 5's row is $(1, 1)$.
- Slogan: $A$ says which bends are behind you; $A\delta$ adds up their slope changes.
Days $t = 0, \dots, 6$; changepoints $s = (2, 4)$; $k = 1$, $m = 10$; slope changes $\delta = (2, -2)$, so $\gamma = -s\odot\delta = (-2\times2,\ -4\times(-2)) = (-4, 8)$.
$$A = \begin{bmatrix} 0 & 0 \\ 0 & 0 \\ 1 & 0 \\ 1 & 0 \\ 1 & 1 \\ 1 & 1 \\ 1 & 1 \end{bmatrix}, \qquad A\delta = \begin{bmatrix} 0 \\ 0 \\ 2 \\ 2 \\ 0 \\ 0 \\ 0 \end{bmatrix}, \qquad A\gamma = \begin{bmatrix} 0 \\ 0 \\ -4 \\ -4 \\ 4 \\ 4 \\ 4 \end{bmatrix}.$$- Row $t = 3$: $a(3) = (1, 0)$ (day 2 has passed, day 4 has not). Slope $k + a\cdot\delta = 1 + 2 = 3$; intercept $m + a\cdot\gamma = 10 - 4 = 6$; $g(3) = 3\times3 + 6 = 15$.
- Row $t = 4$: $a(4) = (1, 1)$. Slope $1 + 2 - 2 = 1$; intercept $10 - 4 + 8 = 14$; $g(4) = 4 + 14 = 18$.
- Row $t = 6$: slope 1, intercept 14, $g(6) = 6 + 14 = 20$.
- All rows: $g = (10, 11, 12, 15, 18, 19, 20)$. The slope is 1, then 3 (from day 2), then 1 again (from day 4), and the values connect.
- Hinge check for $t = 6$: $kt + m + \delta_1(6 - 2)_+ + \delta_2(6 - 4)_+ = 6 + 10 + 2\times4 - 2\times2 = 20$. Same answer.
For times $t_1, \dots, t_n$ and changepoints $s_1, \dots, s_S$, the changepoint matrix is the $n\times S$ matrix of 0s and 1s
$$A_{ij} = a_j(t_i) = \begin{cases} 1 & \text{if } t_i \ge s_j,\\ 0 & \text{otherwise.}\end{cases}$$With $\gamma = -s\odot\delta$ ($\odot$ = element-by-element product), the trend for all days at once is
$$g = (k\,\mathbf{1} + A\delta)\odot t + (m\,\mathbf{1} + A\gamma) \;=\; k\,t + m\,\mathbf{1} + R\,\delta, \qquad R_{ij} = (t_i - s_j)_+ = A_{ij}\,(t_i - s_j).$$- $A\delta$ is, for each day, the sum of the slope changes already passed; $A\gamma$ the sum of the offsets.
- $R$ is the ramp (hinge) matrix: the changepoint block of the design matrix (Chapter 7.7). Since $g$ is linear in $(m, k, \delta)$, the trend is a linear regression on the columns $\mathbf{1}, t, R$.
- Prophet's Stan model computes exactly $(k + A\delta)\odot t + (m + A\gamma)$ with $\gamma = -s\odot\delta$, on its rescaled time.
- NumPy / JAX:
A = (t[:, None] >= s[None, :]).astype(float), theng = (k + A @ delta) * t + (m + A @ (-s * delta)). Static shapes ($n\times S$), no Python loops, so it JIT-compiles well.
Why do we need it?
A loop over changepoints for every day is slow and awkward under JAX's JIT. One matrix product computes the whole trend for all days and all changepoints at once, with fixed shapes, and makes the linear (regression) structure explicit.
Where is it used?
Prophet's Stan code and its Python helper that builds the changepoint matrix, Prophet-style NumPyro models (yours), PyMC and Stan re-implementations of Prophet, and regression splines (the truncated power basis).
How is it used?
Build A once from the training times and the candidate locations (and again for the future times, with the same s). In the model compute the trend with two matrix-vector products, or use the ramp matrix R = A * (t[:, None] - s) as the changepoint columns of the design matrix.
"$A$ has to be learned or updated during training."
$A$ depends only on the times and the changepoint locations, which are fixed before fitting. Build it once (for training days and, separately, for future days) and treat it as data.
"Use t > s or t >= s, it matters a lot."
At $t = s_j$ the ramp $(t - s_j)$ is 0 either way, so the trend values are identical. Pick one convention and use it for both the training and the future matrix.
"Boolean masks like t[t >= s_j] are a fine way to build the trend in JAX."
Masked indexing gives arrays whose length depends on the data, which fails under jit. The 0/1 matrix $A$ (a comparison cast to float) keeps every shape fixed (Chapter 6.17).
In your forecasting model, the changepoint locations (from the grid plus PELT) define $A$ for the training days, and the same locations define it for the forecast days; the slope changes $\delta_j$ with their Laplace priors are the only trend weights besides $k$ and $m$. Building $A$ with a comparison and a cast, instead of boolean masking, keeps every shape fixed under JIT: the same lesson as the "boolean masking before tracing" problem you met in your A/B framework.
$A_{ij} = 1[t_i \ge s_j]$; $g = (k + A\delta)\odot t + (m + A\gamma)$, $\gamma = -s\odot\delta$.
Equivalent: $g = kt + m + R\delta$ with ramps $R_{ij} = (t_i - s_j)_+$: a regression on $\mathbf{1}, t, R$.
Trap: $A$ is fixed data (built from fixed locations), the same for training and future days; avoid boolean masks under JIT.
Quick check: times $t = 0..9$, changepoints $s = (3, 7)$. Write the rows of $A$ for $t = 2, 3, 8$, and the ramp values $(t - s_j)_+$ for $t = 8$.
$a(2) = (0, 0)$, $a(3) = (1, 0)$, $a(8) = (1, 1)$. Ramps at $t = 8$: $(8 - 3)_+ = 5$ and $(8 - 7)_+ = 1$.
Why trends change, and what a bending line cannot do core
Growth rates change for real reasons: a new product launches, a price changes, a marketing channel opens or closes, a market fills up (saturation), a competitor arrives, the economy turns, a pandemic changes how people shop. Each of these changes the speed at which the level moves, and a changepoint with a slope change $\delta_j$ captures it.
But not every break in a series is a change of speed. Some are jumps in the level (a tracking change suddenly counts twice as many orders; a big new customer is onboarded on one day), some are one-day spikes (an outage, a viral post), some are temporary dips that come back (a two-week closure). A continuous piecewise-linear trend can only bend. It handles a jump badly (it needs two very close changepoints to fake a steep ramp) and it should not chase a one-day spike at all: those belong to other layers (a step regressor, a holiday/event column, a robust likelihood).
Three ways to say it:
- Picture: a changepoint bends the road; it cannot make a cliff or a pothole.
- Numbers: to fake a +40 level jump with bends you need $+13.3$ per day for 3 days and then $-13.3$: two changepoints doing the job of one step column.
- Slogan: changepoints are for changes in speed; jumps and spikes need their own columns.
Underlying level $50 + 0.5t$ orders; on day 60 a tracking fix adds 40 orders to every later day (a level shift).
- A continuous trend cannot jump: just before and just after day 60 it has (almost) the same value. One changepoint at day 60 can only change the slope there, so it leaves a residual pattern shaped like a "Z" around day 60.
- Fake the jump with two changepoints, at days 60 and 63: an extra slope of $40/3 \approx 13.3$ orders per day for 3 days, then cancel it. So $\delta_1 \approx +13.3$, $\delta_2 \approx -13.3$.
- That works, but it uses two parameters, makes the jump look like a 3-day ramp, and with a Laplace prior those two large $|\delta|$ values are heavily penalized (the prior expects few large slope changes), so the fit will smooth the jump.
- The direct fix is a step column $1[t \ge 60]$ (0 before, 1 from day 60) with one weight, estimated near 40. It is an exogenous "event" regressor, like a holiday column that never switches off (Chapter 7.12).
Four kinds of change in a series and the layer that should handle each:
| Kind of change | What happens | Typical causes | Right tool |
|---|---|---|---|
| Slope change (trend break) | the growth rate changes; the level stays connected | launch, pricing, saturation, competitor, new channel | changepoint with $\delta_j$ |
| Level shift | the level jumps once and stays | tracking or definition change, acquisition, a big contract | step column $1[t \ge s]$ (or fix the data) |
| Spike / outlier | one or a few days far off, then back | outage, bot traffic, viral post, data error | event column, data cleaning, or a heavy-tailed likelihood (Student-t) |
| Temporary regime | a block of days shifted, then back to normal | store closure, lockdown, a two-week campaign | a window indicator (event with a duration) or a regressor |
A structural break is the general term (econometrics) for any lasting change in the process that generates the data; slope changes and level shifts are the two most common kinds.
Why do we need it?
Using changepoints for the wrong kind of break bends the trend in strange ways: spikes pull nearby slopes, level shifts become steep ramps, and the forecast inherits the distortion. Matching each kind of change to the right layer keeps the trend honest.
Where is it used?
Data reviews before training any Prophet-style model, intervention analysis (step and pulse dummies in regression and ARIMA), outlier handling in demand forecasting, and annotating known business events (launches, tracking changes) in a forecasting pipeline.
How is it used?
Plot the series and the residuals; for each break ask "did the speed change, the level jump, or did a few days misbehave?". Keep slope changes for changepoints, add step or window columns for known level shifts and regimes, and clean or down-weight one-off spikes.
"Any break in the series is a changepoint."
In a Prophet-style trend, a changepoint is specifically a change in slope. Level shifts, spikes and temporary regimes are different kinds of break and need other columns (or data fixes).
"More changepoints will eventually handle the jump."
They approximate it with steep ramps that cost two large slope changes each, which a sparse Laplace prior resists, and which can leave the trend with a wrong final slope. A single step column is cheaper and more honest when you know the date.
"Outliers should be absorbed by the trend."
The trend should describe the slow level. A one-day spike pulls a flexible trend towards itself on both sides; handle it with an event column, data cleaning, or a heavy-tailed likelihood (Chapter 7.13).
Your forecasting model has the tools for each kind of break: changepoints with Laplace-shrunk $\delta_j$ for slope changes, holiday/event columns for spikes and windows, exogenous regressors for level shifts you know about (a step column), and a Student-t likelihood for outliers you cannot explain. A detector such as PELT reacts to any change in the data's behaviour, so check what kind of change each detected point really is before you feed it in as a trend changepoint (Chapter 7.9).
Changepoint = change in slope (speed), trend stays connected. Level shift = step column $1[t \ge s]$. Spike = event column / robust likelihood. Temporary regime = window indicator.
Faking a jump of size $J$ over $w$ days needs $\delta = \pm J/w$ at two close changepoints.
Trap: forcing every break into the trend distorts the slopes and the forecast.
Quick check: on 1 April your analytics tool switched from counting sessions to counting users, and the series dropped by 30% overnight and stayed there. Changepoint or something else?
Something else: it is a level shift caused by a definition change, not a change in growth speed. Best: fix the data (rescale one side so both use the same definition). Otherwise add a step column $1[t \ge \text{1 April}]$ (multiplicative, or on the log scale, since the drop is a percentage). A changepoint would only bend the trend.
Too many or too few changepoints core
With too few changepoints the trend is stiff. It cannot follow a real turn, so it cuts through the middle: too high in one period, too low in the next, and its final slope mixes the old growth rate with the new one. The forecast goes in the wrong direction. This is underfitting (high bias).
With too many changepoints, and nothing to hold them back, the trend becomes a nervous zig-zag that follows the noise. It looks great on the history, but its last segment is fitted on a handful of noisy days, so its final slope, the one the forecast extends, can point anywhere. This is overfitting (high variance).
Two standard defences: put the candidates only in the earlier part of the history (so the last slope is estimated from many days), and shrink the slope changes towards zero with a sparse prior (Laplace in your model, Chapter 7.10), so that most candidates do nothing unless the data insist.
Three ways to say it:
- Picture: a stiff ruler misses the bends; a wet noodle follows every bump and points anywhere at the end.
- Numbers: on the demo series, 0 changepoints give a future trend error near 30 orders; 40 free changepoints over the whole history give about 25; a few well-placed ones, about 2–5.
- Slogan: too few = bias; too many (unshrunk) = variance; the forecast pays for both.
A trend with slope 0.8, then $+1.2$ from day 40 (slope 2.0), then $-1.5$ from day 85 (slope 0.5). History: days 0–119, noise sd 4. Forecast: days 120–149. Average errors over 30 simulated histories (least squares, evenly spaced candidates):
- 0 changepoints: one straight line through a curve that rises steeply in the middle and flattens at the end. Training error (RMSE) about 8.9; the fitted slope is far from the final 0.5, so the 30-day trend forecast misses by about 29.6 on average (RMSE).
- 10 candidates in the first 80%: training RMSE about 3.8, future trend RMSE about 2.7. The last 20% of the history (24 days) has no candidates, so the final slope is estimated from many days.
- 40 candidates over the whole history (including the last days), no shrinkage: training RMSE about 3.2 (the best), but future trend RMSE about 24.8: the last segments are fitted on a few noisy days and their slope is extrapolated.
- Training error always falls as you add changepoints; forecast error falls and then rises. Only out-of-sample checks (rolling origin, Chapter 7.15) show the turning point.
The number of changepoints (and how strongly their $\delta_j$ are shrunk) is a bias–variance knob (Chapter 5.1; across the whole model: Chapter 7.18):
- Underfitting (too few, or too much shrinkage): systematic residual patterns (bows, long runs of one sign), biased final slope, forecasts that keep the old growth rate.
- Overfitting (too many, too little shrinkage): residuals look like white noise or smaller, slopes swing from segment to segment, the final slope is unstable from one sample to the next, and forecast intervals (if they reflect slope uncertainty) explode.
- Two practical knobs: the number and placement of candidates (for example Prophet's 25 candidates in the first 80% of the history) and the prior scale $b$ of $\delta_j \sim \text{Laplace}(0, b)$. With a sparse prior you can afford many candidates, because most get $\delta_j \approx 0$.
Why do we need it?
The trend is the layer that dominates long-horizon forecasts, and its flexibility is the easiest thing to get wrong. Too stiff and the forecast keeps an outdated growth rate; too loose and it extrapolates noise.
Where is it used?
Tuning Prophet's n_changepoints, changepoint_range and changepoint_prior_scale; choosing PELT's penalty (Chapter 7.9); choosing the Laplace scale in your model (Chapter 7.10); knot selection in regression splines.
How is it used?
Fix a generous set of candidates away from the very end, then tune the shrinkage (or the penalty) by rolling-origin backtests on the horizon you care about. Look at the fitted slope staircase: many tiny zig-zags mean too loose; residual bows mean too stiff.
"The model with the lowest training error has the right number of changepoints."
Training error always drops when you add flexibility. Choose the number (or the shrinkage) by forecast error on held-out future periods.
"Changepoints near the end of the history help the forecast react to recent changes."
They can, but the slope of a segment fitted on a few days is very noisy, and that is the slope that gets extrapolated. This is why Prophet keeps its default candidates in the first 80% of the history, and why you need strong shrinkage (or many days) to trust a late changepoint.
"With a Laplace prior the number of candidates no longer matters."
It matters less, because most $\delta_j$ are pulled to near zero, but candidate spacing, the prior scale $b$ and the late candidates still shape the fit and the forecast. Check sensitivity to both (Chapter 7.10).
In your forecasting model, the grid plus PELT decides where slope changes may happen and the Laplace prior decides how much each one may do. Together they are your trend's bias–variance knob. In an interview, describe it in exactly those terms: "too few changepoints biases the final slope; too many unshrunk changepoints overfit and make the extrapolated slope unstable; I placed candidates with a grid plus PELT and shrank their slope changes with a Laplace prior, and I validated the setting on rolling-origin forecasts." (Say the last part only if you really did it; otherwise say how you would.)
Too few changepoints → stiff trend, biased final slope (underfit). Too many unshrunk → zig-zag trend, unstable final slope (overfit).
Defences: candidates away from the end (Prophet: first 80%), sparse prior on $\delta_j$, choose by out-of-sample forecast error.
Trap: training error always prefers more changepoints.
Quick check: a changepoint 5 days before the end of the history is fitted by least squares with noise sd 10. Roughly how uncertain is the slope of that last 5-day segment?
The slope of a line fitted to 5 equally spaced points has standard error $\sigma/\sqrt{S_{tt}}$ with $S_{tt} = \sum(t - \bar t)^2 = 4 + 1 + 0 + 1 + 4 = 10$, so about $10/\sqrt{10} \approx 3.2$ orders per day (a bit less here because the continuity constraint ties the segment to the earlier line). Extrapolated 30 days, that is an uncertainty of the order of $\pm95$ orders from the slope alone. A late changepoint makes the forecast's slope very noisy.
Extrapolating the trend: the last slope rules the forecast core
Once the history ends, there are no more data to bend the line. The forecast continues the last segment with the final slope $k + \sum_j\delta_j$. Everything earlier in the history matters only through how it pins down that last segment.
Two kinds of uncertainty follow. First, the final slope is an estimate: if the last segment is short, it is a noisy estimate, and its error grows linearly with the horizon. Second, the future may contain new changepoints nobody has seen yet. Prophet's answer to the second is to assume the future will see slope changes about as often and about as large as the history did, and to simulate them.
Three ways to say it:
- Picture: a car with its steering wheel locked at the last angle; a short last road segment means you barely know that angle.
- Numbers: a slope uncertainty of 0.3 orders/day becomes about ±15 orders after 30 days and ±44 after 90 days (90% bands).
- Slogan: the trend forecast is the final slope, extended; its uncertainty grows with the horizon.
At the forecast origin $T$, the trend is $\hat g(T) = 300$ with final slope $\hat k_{\text{final}} = 1.0$ orders/day and standard error 0.3.
- Point forecast 30 days ahead: $300 + 1.0\times30 = 330$.
- Slope uncertainty alone, 90% band: $\pm1.645\times0.3\times30 \approx \pm14.8$ orders after 30 days, and $\pm1.645\times0.3\times90 \approx \pm44.4$ after 90 days. It grows linearly with the horizon.
- Future changepoints (Prophet's idea): with $S = 25$ changepoints spread over a 365-day history, the rate is about $25/364$ per day, so a 90-day horizon expects $25\times90/364 \approx 6.2$ new slope changes.
- If they are drawn from $\text{Laplace}(0, \lambda)$ with $\lambda = 0.06$ orders/day (the average size of the fitted $|\delta_j|$), six of them change the slope by a total with standard deviation $\sqrt{6\times2\times0.06^2} = 0.06\sqrt{12} \approx 0.21$ orders/day by the end, widening the band further.
For $t \gt T$ (after the last observation), with all changepoints inside the history,
$$\hat g(t) = \hat g(T) + \Big(\hat k + \sum_{j=1}^{S}\hat\delta_j\Big)(t - T).$$- Slope uncertainty: $Var(\hat g(T + h)) = Var(\hat g(T)) + h^2\,Var(\hat k_{\text{final}}) + 2h\,Cov(\hat g(T), \hat k_{\text{final}})$, dominated by the $h^2$ term for long horizons. In a Bayesian fit you get it automatically from posterior draws of $(m, k, \delta)$.
- Future trend changes (Prophet): Prophet's documentation explains that it assumes future slope changes will be as frequent and, on average, as large as those seen in the history. In its implementation, future changepoints arrive at the historical rate (about $S$ per history length) at random times, and their sizes are drawn from $\text{Laplace}(0, \lambda)$ with $\lambda$ = the average of the fitted $|\delta_j|$. Many simulated trend paths give the trend part of the forecast interval.
- Why candidates stop before the end: Prophet's default
changepoint_range = 0.8leaves the last 20% of the history free of candidates; its documentation gives two reasons: to leave room ("runway") for projecting the trend forward, and to avoid overfitting fluctuations at the very end of the series.
Why do we need it?
For long horizons the trend is the largest source of forecast error. Knowing that only the final slope is extrapolated, and how uncertain it is, tells you how far ahead a forecast can be trusted and why intervals must widen with the horizon.
Where is it used?
Prophet's trend uncertainty (simulated future changepoints), Bayesian posterior predictive forecasts of a piecewise trend, capacity planning months ahead, and every discussion of "why is the interval so wide in six months?".
How is it used?
Report the final slope and its uncertainty; keep candidates (or strong shrinkage) away from the last days; check the forecast fan's width at your planning horizon; and if your model does not simulate future slope changes, say that its trend intervals cover only the uncertainty of the current slope.
"The forecast uses the average slope of the history."
It uses the final slope $k + \sum_j\delta_j$. Early history matters only through how well it pins down the last segment.
"A narrow trend interval means the future trend is well known."
It may only mean the model assumes no future changes. Intervals from the current slope's uncertainty alone ignore new changepoints; Prophet adds them by simulation, and with MAP fitting it does not include the uncertainty of the fitted slopes themselves unless you ask for full sampling (mcmc_samples).
"Simulating future changepoints from least-squares slope changes is fine."
The simulated size $\lambda$ is the average fitted $|\delta_j|$. An overfitted zig-zag history inflates it and makes the future fan enormous. Shrinkage on the $\delta_j$ matters for the interval width too.
In your forecasting model, posterior draws of $k$ and the $\delta_j$ (from your SVI guide) give the uncertainty of the final slope automatically, and it widens linearly with the horizon. Whether your forecasts also include future slope changes (as Prophet's simulation does) is a design choice of your code: know which, and say it when you present intervals. Note also that a full-rank guide can represent the strong negative correlation between neighbouring $\delta_j$ (and between $k$ and the $\delta_j$), which a mean-field guide ignores; that correlation affects the final-slope uncertainty.
Forecast trend: $\hat g(T) + (\hat k + \sum_j\hat\delta_j)(t - T)$; slope error grows like $h$ (variance like $h^2$).
Prophet: future changepoints at the historical rate ($S$ per history length), sizes $\text{Laplace}(0, \text{mean}|\hat\delta|)$; candidates in the first 80% only.
Trap: a short last segment means a noisy final slope; overfitted $\delta$'s inflate simulated future changes.
Quick check: 30 changepoints were spread over a 600-day history. How many new changes does Prophet's simulation expect in the next 120 days, on average?
The rate is 30 per (about) 600 days, i.e. 0.05 per day, so $0.05\times120 = 6$ new slope changes on average (the actual number is random, Poisson with mean about 6).
Where may the trend bend? Approach 1: a fixed grid of candidates core
We know the trend may bend, but not where. There are three ways to decide:
- A fixed grid: place many candidate changepoints at regular intervals and let a sparse prior switch most of them off.
- Data-driven detection: run a detector (such as PELT) that finds the moments where the series' behaviour changes, and use those.
- Bayesian latent changepoints: treat the locations themselves as unknown parameters and learn a posterior over where the bends are.
The grid is the simplest. It is like putting hinges every few metres along a fence: most stay straight, and only the hinges where the fence really needs to turn are used. The work of choosing is moved from "where" to "how much": the slope changes $\delta_j$ and their prior.
Three ways to say it:
- Picture: hinges every few metres; a sparse prior keeps most of them straight.
- Numbers: Prophet's default on one year of daily data: 25 candidates, about every 11.6 days, all in the first 292 days.
- Slogan: offer many bends, pay for each one, keep the few the data really want.
Prophet's default grid for 365 daily rows (n_changepoints=25, changepoint_range=0.8).
- Rows eligible for candidates: the first 80%: $\lfloor 365\times0.8 \rfloor = 292$ rows (indices 0–291).
- Prophet takes 26 evenly spaced positions from row 0 to row 291 (step $291/25 = 11.64$), rounds them, and drops the first one (row 0, where a slope change would just be the base slope).
- Candidates: rows 12, 23, 35, 47, 58, …, 268, 279, 291: about every 11.6 days.
- The last $365 - 292 = 73$ days have no candidates: the final slope is estimated from at least 73 days.
- Each candidate gets $\delta_j \sim \text{Laplace}(0, 0.05)$ on Prophet's scaled data (
changepoint_prior_scale=0.05). Fitted by MAP, most $\delta_j$ come out tiny and a few are far from zero: the bends the data support.
Fixed-grid approach. Choose candidate locations $s_1 \lt \dots \lt s_S$ in advance, without looking at the shape of the data (only at its time range), then fit all slope changes $\delta_1, \dots, \delta_S$ with a sparsity-inducing prior (or penalty) so that most are near zero.
- Prophet's defaults (from its documentation):
n_changepoints=25potential changepoints placed uniformly in the first 80% of the history (changepoint_range=0.8), andchangepoint_prior_scale=0.05, the scale of the Laplace prior on the slope changes. You can also pass your own list of dates withchangepoints=[...]. - Pros: simple, no detection step, the model stays linear with fixed shapes (fast, JIT-friendly), and real bends anywhere in the eligible range can be picked up.
- Cons: a bend between two candidates is approximated by the nearest one or two; results depend on the prior scale; no candidates in the last part of the history means a late real change is missed until the next refit.
Why do we need it?
Searching for the exact locations of bends is hard and uncertain. A dense grid plus a sparse prior turns "where?" into an easy continuous problem ("how much at each candidate?") that a regression or an SVI loop can solve directly.
Where is it used?
Prophet's default trend, Prophet-style models in NumPyro, Stan and PyMC, your model's grid part, and lasso-type trend filtering (L1 penalties on slope changes, related to "ℓ1 trend filtering").
How is it used?
Compute candidate positions from the training time range only (first 80% by default), build the ramp columns, put a Laplace (or similar) prior on their weights, fit, then inspect which $\delta_j$ are far from zero. Tune the prior scale by rolling-origin backtests.
"The grid says the trend changes at those 25 dates."
The grid only says where the trend may change. The fitted $\delta_j$ decide which candidates matter; most are near zero.
"Prophet's 25 candidates and $b = 0.05$ are tuned for my data."
They are generic defaults on scaled data. Too stiff or too wiggly a trend is the first thing Prophet's documentation tells you to fix by changing changepoint_prior_scale; check the effect by backtesting.
"A grid can place a bend exactly where it happened."
Only at candidate locations. A real bend between two candidates is shared by the two neighbours (two smaller $\delta$'s), which the sparse prior may smooth a little. A finer grid or a detector (next section) helps.
Your model's candidate set starts from a Prophet-like grid: evenly spaced candidates over the eligible part of the history, each with a slope change $\delta_j \sim \text{Laplace}(0, b)$. Know your own settings (how many candidates, which share of the history, the value of $b$ and the scaling it applies to) and be ready to explain why the last part of the history has no candidates.
Grid: many fixed candidates + sparse prior on $\delta_j$; "where?" becomes "how much at each candidate?".
Prophet defaults: n_changepoints=25 in the first 80% (changepoint_range=0.8), $\delta_j \sim \text{Laplace}(0, 0.05)$ on scaled data (changepoint_prior_scale).
Trap: candidates are possibilities, not detected changes; defaults are not tuned for your series.
Quick check: 200 daily rows, n_changepoints=10, changepoint_range=0.9. Where do the candidates go and how many final days are free of them?
Eligible rows: $\lfloor200\times0.9\rfloor = 180$ (indices 0–179). Ten candidates evenly spread from row 0 to 179 with step $179/10 = 17.9$, after dropping row 0: about rows 18, 36, 54, 72, 90, 107, 125, 143, 161, 179. The last $200 - 180 = 20$ days have no candidates.
Approach 2: data-driven detection (PELT) core
Instead of offering bends everywhere, we can look at the data first and ask: where does the series' behaviour change? A change-point detector splits the series into segments so that each segment is well described by something simple (here: a straight line), while paying a fixed penalty for every extra segment. Without the penalty it would cut the series into tiny pieces; with it, it keeps a cut only if the cut makes the fit much better.
PELT (Pruned Exact Linear Time) is a fast, exact way to find the best such segmentation. Its details (cost functions, the penalty, the pruning trick) are the whole of Chapter 7.9; here we only need what it hands to the trend model: a short list of locations.
Three ways to say it:
- Picture: a tailor who cuts the cloth only where the pattern really changes, and charges for every cut.
- Numbers: one straight line through "0, 1, 2, 3, 4, 4, 4, 4, 4, 4" leaves squared errors of 5.15; two lines leave 0. With a penalty of 3 per cut, cutting wins; with a penalty of 8, it does not.
- Slogan: detection = let the data propose the bends, at a price per bend.
Ten days: $y = 0, 1, 2, 3, 4, 4, 4, 4, 4, 4$ (growth of 1 per day, then flat). Segment cost = the squared error of the best straight line on that segment.
- One segment (no changepoint): the best line through all ten points is $\hat y = 1.09 + 0.424\,t$, with squared error 5.15. Total: $5.15 + 0\times\beta$.
- Two segments, split at day 5: days 0–4 lie exactly on a line (error 0), days 5–9 too (error 0). Total: $0 + 1\times\beta$.
- Penalty $\beta = 3$: $3 \lt 5.15$, so the detector reports a changepoint at day 5. Penalty $\beta = 8$: $8 \gt 5.15$, no changepoint.
- The detected location (day 5, or day 4 for the continuous model) then becomes a fixed changepoint in the trend: here the continuous fit with a ramp at day 4 has $k = 1$, $\delta = -1$ and zero error.
Data-driven detection chooses changepoint locations $\tau_1 \lt \dots \lt \tau_M$ from the data by solving
$$\min_{M,\ \tau_1, \dots, \tau_M}\ \sum_{i=0}^{M} C\big(y_{\tau_i : \tau_{i+1}}\big) + \beta M,$$where $C$ is a segment cost (for a trend: the squared error of a straight-line fit; for mean shifts: the squared error around the segment mean), $\beta$ is the penalty per changepoint, $\tau_0$ and $\tau_{M+1}$ are the series' start and end. PELT finds the exact minimizer quickly (Chapter 7.9).
- The detected locations are then used as the $s_j$ of the trend (alone, or added to a grid), and the model fits their $\delta_j$.
- Choices that matter: what series PELT sees (the raw series, a deseasonalized one, the day-to-day changes), the cost, the penalty $\beta$, and the minimum segment length. With Prophet you can feed such locations in with
changepoints=[...]. - The big caveat: the locations are chosen from the same data and then treated as known. The posterior does not know they could have been elsewhere, so its uncertainty is too small (selection uncertainty is not propagated, Chapter 7.9), and running the detector on data that includes the test period leaks the future (Chapter 7.12).
Why do we need it?
A grid cannot put a bend exactly where it happened, and with a short history or a strong prior it may smooth real changes away. A detector proposes locations where the data really change, so few candidates can do the work of many.
Where is it used?
The ruptures library (PELT, binary segmentation, window methods), the R changepoint package, monitoring and anomaly systems, segmenting sensor or traffic data, and as the data-driven half of your model's grid + PELT changepoint selection.
How is it used?
Run PELT on the training period only, with a cost that matches the change you care about and a penalty tuned (for example BIC-like, a constant times $\log n$); inspect each detected point; pass the accepted ones as changepoint locations; still keep a shrinkage prior on their $\delta_j$.
"PELT finds the true changepoints."
It finds the best segmentation for the chosen cost and penalty on this sample. Change the penalty and the answer changes; resample the noise and the locations move. The output is an estimate.
"Once PELT has found the changepoints, the model's uncertainty is complete."
The model treats detected locations as known, so the posterior ignores "the bend could have been on day 35 instead of 40". Intervals are somewhat too narrow; see Chapter 7.9 for a simulation.
"Run PELT on the whole series, then backtest."
If the detector saw the test period, its locations carry information from the future: a leak that makes backtests look better than live forecasts. Detect inside each training window only.
Your model combines a Prophet-like grid with PELT detection. Be precise about the details of your own pipeline, because an interviewer will ask: which series PELT runs on (raw, deseasonalized, differenced), which cost and penalty it uses, the minimum segment length, how detected points are merged with the grid (added, used instead, snapped to it), whether detection runs only on each training window, and that every candidate, grid or detected, still gets a Laplace-shrunk $\delta_j$. And state the limitation plainly: the detected locations are fixed inputs, so their uncertainty is not in your posterior.
Detection: minimize $\sum C(\text{segment}) + \beta M$; PELT solves it exactly and fast (Chapter 7.9). Detected $\tau$'s become fixed $s_j$'s.
Small β → over-detection; large β → under-detection.
Trap: locations are estimates from the same data: their uncertainty is lost, and detecting on the test period leaks.
Quick check: one line through a segment costs 120; the best split into two lines costs $40 + 50 = 90$. For which penalties does the detector split?
Splitting changes the total from $120$ to $90 + \beta$. It splits when $90 + \beta \lt 120$, i.e. $\beta \lt 30$.
Approach 3: Bayesian latent changepoints (the location is uncertain too) core
In the first two approaches the possible bend locations are fixed before the model is fitted. The fully Bayesian alternative says: the location of a bend is just another unknown. Give it a prior ("any day is equally likely"), and let the data produce a posterior over the location: perhaps "day 44 with 13%, day 42 with 13%, day 45 with 11%, …". It is called latent because it is never observed directly.
The reward: the forecast averages over all plausible locations, so the uncertainty about where the trend bent flows into the uncertainty about where it is going. The price: locations are discrete and the posterior over them is lumpy, which gradient-based tools (NUTS, SVI) cannot handle directly, and the computation grows quickly with the number of bends.
Three ways to say it:
- Picture: instead of one pin on the calendar, a smudge showing where the bend probably is.
- Numbers: locations 3, 5, 7 with posterior 0.1, 0.6, 0.3 and final slopes 0.8, 0.5, 0.2 give an averaged slope $0.08 + 0.30 + 0.06 = 0.44$.
- Slogan: do not pick the bend; average over it.
A tiny model with one bend, at one of three candidate days $s \in \{3, 5, 7\}$, with equal prior probability $1/3$ each. After fitting each option, the data are 1, 6 and 3 times as likely under $s = 3, 5, 7$ (their marginal likelihoods are in the ratio $1 : 6 : 3$).
- Posterior $\propto$ prior × likelihood: $\tfrac13\times1 : \tfrac13\times6 : \tfrac13\times3 = 1 : 6 : 3$.
- Normalize (divide by $1 + 6 + 3 = 10$): $P(s = 3\mid y) = 0.1$, $P(s = 5\mid y) = 0.6$, $P(s = 7\mid y) = 0.3$.
- Each location implies a different final slope: 0.8, 0.5, 0.2 orders per day.
- Posterior-averaged final slope: $0.1\times0.8 + 0.6\times0.5 + 0.3\times0.2 = 0.08 + 0.30 + 0.06 = 0.44$.
- A detector would pick $s = 5$ and report slope 0.5 as if certain. The Bayesian answer is 0.44, with extra spread because 40% of the probability sits on other locations.
A Bayesian latent changepoint model treats the locations (and possibly their number) as parameters:
$$s \sim p(s), \quad \delta \sim p(\delta), \quad (k, m) \sim p(k, m), \quad y_t \mid s, \delta, k, m \sim \text{Lik}\Big(kt + m + \sum_j\delta_j(t - s_j)_+,\ \sigma\Big),$$ $$p(s \mid y) \propto p(s)\,p(y \mid s), \qquad p(y \mid s) = \int p(y \mid s, \theta)\,p(\theta)\,d\theta \ \ (\text{the marginal likelihood of location } s).$$- Forecasts average over locations: $p(\tilde y\mid y) = \sum_s p(\tilde y \mid s, y)\,p(s\mid y)$.
- How it is computed: for one or two bends, enumerate every candidate location (sum it out); for an unknown number of bends, reversible-jump MCMC or product-partition models; online versions such as Bayesian online changepoint detection (Adams and MacKay, 2007); packages such as BEAST for trend + season breaks.
- Gradients: a discrete location has no gradient, so NUTS and reparameterized SVI cannot move it directly. Workarounds: enumerate (sum over) a small discrete set, which NumPyro supports for discrete sites in MCMC, or replace the hard step $1[t \ge s]$ with a smooth sigmoid so that $s$ becomes continuous (the posterior can then be multimodal).
Why do we need it?
It is the honest answer to "where did the trend change?" when the data cannot say exactly. Its intervals include location uncertainty, which grid-plus-detection pipelines leave out, and it tells you when a "detected" change is weakly supported.
Where is it used?
Bayesian structural-break analysis in econometrics and climate science, BEAST for remote-sensing time series, Bayesian online changepoint detection in monitoring, epidemiological growth-rate change analysis, and as a validation tool for detector-based pipelines.
How is it used?
For a few changes: compute the marginal likelihood of the model for each candidate location (or enumerate them inside NumPyro), normalize to a posterior, and average forecasts over it. For many: use a dedicated sampler or package. Compare the posterior's spread with the detector's single answer.
The three approaches side by side.
| 1 · Fixed grid | 2 · Detection (PELT) | 3 · Bayesian latent | |
|---|---|---|---|
| Where may bends be? | many preset candidates | a few locations chosen from the data | anywhere (a prior over locations) |
| What is learned in the model? | $\delta_j$ at every candidate (sparse prior) | $\delta_j$ at the detected locations | locations and $\delta_j$ together |
| Location uncertainty in the forecast? | partly (spread over neighbouring candidates) | no (locations fixed after detection) | yes |
| Computation | easy, linear, fixed shapes (SVI/NUTS friendly) | a fast detection step, then the same easy model | hard: discrete, multimodal; enumeration or special samplers |
| Main risks | prior scale too loose/tight; no late candidates | penalty choice; leakage; overconfidence | cost; tuning; harder to explain and to scale |
"The best location is the answer; the rest of the posterior is noise."
When the posterior spreads over 10–20 days, the single best day is only slightly more likely than its neighbours. Reporting one day as the changepoint hides real uncertainty, and so does a pipeline that fixes it.
"Bayesian latent changepoints are always better."
They are more honest about location, but much harder to compute and scale, and they still depend on priors (on the number and size of changes). A grid with a sparse prior gets much of the benefit at a fraction of the cost.
Your model uses approaches 1 and 2 together (a Prophet-like grid plus PELT detection), with a Laplace prior on every $\delta_j$. That is a sound, fast design, and approach 3 is the right thing to name when an interviewer asks about its weakness: "the PELT locations are treated as fixed, so the posterior does not include uncertainty about where the trend changed; the grid with shrinkage partly compensates by letting neighbouring candidates share a bend; a fully Bayesian alternative would put a prior on the locations and sum over them, at a much higher computational cost".
"Our model estimates the changepoints."
The model estimates the slope changes at given candidate locations. The locations come from a grid and a detector, chosen before fitting, so the posterior is conditional on them.
Model answer: "There are three ways to handle changepoints: a fixed grid with a sparse prior, data-driven detection like PELT, or treating locations as latent variables. I combined a Prophet-like grid with PELT: the grid gives broad coverage, PELT proposes extra locations where the data strongly suggest a change, and Laplace priors shrink every slope change. The trade-off is that location uncertainty from detection is not propagated into the posterior; a latent-changepoint model would fix that but is much harder to fit with SVI."
Latent: $p(s\mid y) \propto p(s)\,p(y\mid s)$; forecasts average over $s$. Captures location uncertainty; grid and detection do not (fully).
Computation: enumerate a few locations, RJ-MCMC for unknown counts, online BOCPD; discrete $s$ has no gradient (enumerate or smooth it).
Trap: one "best" day can hide a wide posterior.
Quick check: the posterior over a bend's location is 0.25 on day 30, 0.5 on day 31 and 0.25 on day 32, and the forecast for next month is 200, 210 and 230 under the three. What is the posterior-averaged forecast, and what would a detector that picks day 31 report?
Average: $0.25\times200 + 0.5\times210 + 0.25\times230 = 50 + 105 + 57.5 = 212.5$. The detector reports 210, with no extra spread from the 50% chance that the bend was on another day.
Recap, cheat sheet and practice
- Straight trend: $g(t) = kt + m$; $k$ = change per unit time, $m$ = value at $t = 0$ (depends on where time zero is and on rescaling). Extrapolation continues the line.
- Changepoints: at $s_j$ the slope changes by $\delta_j$: slope$(t) = k + \sum_{s_j \le t}\delta_j$; the final slope $k + \sum_j\delta_j$ drives the forecast.
- Continuity: without an offset the line jumps by $s_j\delta_j$; requiring equal values at $s_j$ gives $\gamma_j = -s_j\delta_j$, and $g(t) = kt + m + \sum_j\delta_j(t - s_j)_+$.
- Matrix form: $A_{ij} = 1[t_i \ge s_j]$, $g = (k + A\delta)\odot t + (m + A\gamma)$ = $kt + m + R\delta$ with ramps $R_{ij} = (t_i - s_j)_+$: a regression on fixed columns.
- Why trends change: launches, prices, saturation, competitors, shocks. A continuous bend handles slope changes only; level shifts, spikes and temporary regimes need step, event or window columns (or a robust likelihood).
- How many: too few → stiff, biased final slope; too many unshrunk → zig-zag, unstable final slope. Keep candidates away from the end and shrink the $\delta_j$; choose by forecast error.
- Extrapolation: slope uncertainty grows linearly with the horizon; Prophet also simulates future changepoints at the historical rate with Laplace sizes of scale mean$|\hat\delta|$.
- Three approaches: fixed grid + sparse prior (Prophet: 25 candidates in the first 80%, $b = 0.05$ on scaled data); data-driven detection (PELT, Chapter 7.9; locations then fixed); Bayesian latent changepoints (posterior over locations; honest but expensive). Your model: grid + PELT with Laplace priors.
Cheat sheet
| Idea | Formula | Remember |
|---|---|---|
| Linear trend | $g(t) = kt + m$ | $k$ in units of $y$ per unit time |
| Slope after changepoints | $k + \sum_{s_j \le t}\delta_j$ | $\delta_j$ is a change, not the new slope |
| Continuity offset | $\gamma_j = -s_j\delta_j$ | from $\delta_j s_j + \gamma_j = 0$ |
| Prophet form | $(k + a(t)^\top\delta)\,t + (m + a(t)^\top\gamma)$ | $a_j(t) = 1[t \ge s_j]$ |
| Hinge form | $kt + m + \sum_j\delta_j(t - s_j)_+$ | identical function; linear in $(m, k, \delta)$ |
| Matrix form | $g = (k + A\delta)\odot t + (m + A\gamma)$ | A = (t[:,None] >= s[None,:]) |
| Jump without γ | $s_j\delta_j$ | depends on where time zero is |
| Forecast trend | $\hat g(T) + (\hat k + \sum\hat\delta_j)(t - T)$ | SE grows like the horizon $h$ |
| Prophet defaults | 25 candidates, first 80%, $\delta_j \sim \text{Laplace}(0, 0.05)$ | on scaled data; tune by backtest |
| Future changes (Prophet) | rate $\approx S$ per history length; size $\sim \text{Laplace}(0, \overline{|\hat\delta|})$ | overfit δ's inflate the fan |
| Detection objective | $\min \sum C(\text{segment}) + \beta M$ | locations then fixed (no location uncertainty) |
| Latent location | $p(s\mid y) \propto p(s)\,p(y\mid s)$ | average forecasts over $s$ |
import numpy as np
import jax, jax.numpy as jnp
import numpyro, numpyro.distributions as dist
from numpyro.infer import SVI, Trace_ELBO
from numpyro.infer.autoguide import AutoDelta, AutoNormal
np.set_printoptions(precision=2, suppress=True)
# 1) The changepoint matrix A(t), the offsets gamma = -s*delta, and the trend (Prophet's form)
t = np.arange(7.0); s = np.array([2.0, 4.0]); k, m = 1.0, 10.0; delta = np.array([2.0, -2.0])
A = (t[:, None] >= s[None, :]).astype(float) # A[i, j] = 1 if t_i >= s_j
gamma = -s * delta
g = (k + A @ delta) * t + (m + A @ gamma)
print(A.T) # [[0. 0. 1. 1. 1. 1. 1.] [0. 0. 0. 0. 1. 1. 1.]]
print(gamma, g) # [-4. 8.] [10. 11. 12. 15. 18. 19. 20.]
# 2) The hinge (ramp) form is the same function; without gamma the line jumps
R = np.maximum(0.0, t[:, None] - s[None, :]) # R[i, j] = (t_i - s_j)_+
print(np.allclose(g, k * t + m + R @ delta)) # True
print((k + A @ delta) * t + m - g) # [ 0. 0. 4. 4. -4. -4. -4.] = A @ (s * delta)
# 3) Prophet's default grid: 25 candidates evenly spread over the first 80% of the rows
def prophet_grid(n, n_changepoints=25, changepoint_range=0.8):
hist = int(np.floor(n * changepoint_range))
return np.linspace(0, hist - 1, n_changepoints + 1).round().astype(int)[1:]
print(prophet_grid(365)[:5], prophet_grid(365)[-2:]) # [12 23 35 47 58] [279 291]
# 4) A trend with two real bends (day 40: +1.2, day 85: -1.5), 200 days, noise sd 4
rng = np.random.default_rng(1)
n = 200; t = np.arange(n, dtype=float)
g_true = 60 + 0.8 * t + 1.2 * np.maximum(0, t - 40) - 1.5 * np.maximum(0, t - 85)
y = g_true + rng.normal(0, 4, n)
cp = prophet_grid(n).astype(float) # 25 candidates in the first 160 days
X = np.column_stack([np.ones(n), t, np.maximum(0, t[:, None] - cp[None, :])])
b_ls = np.linalg.lstsq(X, y, rcond=None)[0]
print("LS: mean |delta| =", round(np.abs(b_ls[2:]).mean(), 3), " final slope =", round(b_ls[1] + b_ls[2:].sum(), 3))
# LS: mean |delta| = 0.797 final slope = 0.465 (every candidate gets a sizeable delta; true final slope 0.5)
# 5) The same trend in NumPyro, Prophet-style scaling (t in [0, 1], y / max|y|), Laplace prior on delta
T, ys = t[-1], np.abs(y).max()
ts, yy, cs = jnp.array(t / T), jnp.array(y / ys), jnp.array(cp / T)
As = (ts[:, None] >= cs[None, :]).astype(jnp.float32)
def model(ts, A, s, y=None, b=0.05):
k = numpyro.sample("k", dist.Normal(0, 5))
m = numpyro.sample("m", dist.Normal(0, 5))
delta = numpyro.sample("delta", dist.Laplace(0, b).expand([A.shape[1]]).to_event(1))
sigma = numpyro.sample("sigma", dist.HalfNormal(0.5))
g = (k + A @ delta) * ts + (m + A @ (-s * delta)) # continuous piecewise-linear trend
numpyro.sample("y", dist.Normal(g, sigma), obs=y)
def fit(guide, steps=4000, lr=0.01):
svi = SVI(model, guide, numpyro.optim.Adam(lr), Trace_ELBO())
return svi.run(jax.random.PRNGKey(0), steps, ts, As, cs, y=yy, progress_bar=False).params
map_guide = AutoDelta(model) # MAP, like Prophet's default fit
p = fit(map_guide)
d_map = np.asarray(p["delta_auto_loc"]) * ys / T # back to orders per day
k_map = float(p["k_auto_loc"]) * ys / T
print("MAP: |delta| > 0.05 at days", cp[np.abs(d_map) > 0.05].astype(int), np.round(d_map[np.abs(d_map) > 0.05], 2))
# MAP: |delta| > 0.05 at days [32 38 45 83 89] [ 0.16 0.61 0.38 -0.66 -0.78] (sums 1.15 and -1.44: the two real bends)
print("MAP: mean |delta| =", round(np.abs(d_map).mean(), 3), " final slope =", round(k_map + d_map.sum(), 3))
# MAP: mean |delta| = 0.105 final slope = 0.52 (Adam does not give exact zeros; small deltas are just tiny)
normal_guide = AutoNormal(model) # a posterior (mean-field Gaussian) instead of a point
p2 = fit(normal_guide, steps=6000)
post = normal_guide.sample_posterior(jax.random.PRNGKey(1), p2, sample_shape=(2000,))
final = (post["k"] + post["delta"].sum(-1)) * ys / T
print("posterior final slope:", round(float(final.mean()), 3), "+/-", round(float(final.std()), 3))
# posterior final slope: 0.539 +/- 0.065 (the slope the forecast extends, with its uncertainty)
1. $k = 3$, $\delta_1 = -1$ at day 10, $\delta_2 = +2$ at day 30. What is the slope on day 45?
2. Which offset keeps the trend continuous at changepoint $s_j$ with slope change $\delta_j$?
3. If the offsets are left out, the trend jumps at each changepoint. The size of the jump depends on…
4. A change in how orders are counted adds 25% to every day from 1 May onwards. How should the trend model handle it?
5. Why does Prophet place its default candidate changepoints only in the first 80% of the history?
6. Which approach carries the uncertainty about where the trend changed into the forecast interval?
Practice problems
A. $k = 0.5$, $m = 40$, changepoints $s = (20, 60)$, $\delta = (1.5, -1.0)$. Compute $g(10)$, $g(40)$ and $g(80)$ with the $\gamma$ form and check $g(80)$ with the hinge form.
- $\gamma = (-20\times1.5,\ -60\times(-1.0)) = (-30, +60)$.
- $g(10)$: no changepoint passed: $0.5\times10 + 40 = 45$.
- $g(40)$: first passed: slope $0.5 + 1.5 = 2$, intercept $40 - 30 = 10$: $2\times40 + 10 = 90$.
- $g(80)$: both passed: slope $0.5 + 1.5 - 1 = 1$, intercept $40 - 30 + 60 = 70$: $80 + 70 = 150$.
- Hinge: $0.5\times80 + 40 + 1.5\times(80 - 20) - 1.0\times(80 - 60) = 40 + 40 + 90 - 20 = 150$. ✓
B. (Interview) "Derive the continuity condition for Prophet's piecewise-linear trend."
"Write the trend as $(k + \sum_{s_j \le t}\delta_j)t + (m + \sum_{s_j \le t}\gamma_j)$. Just before $s_j$ the line is $(k + K)t + (m + G)$, where $K$ and $G$ hold the earlier changes; just after, it gains $\delta_j t + \gamma_j$. For the values to match at $t = s_j$ we need $\delta_j s_j + \gamma_j = 0$, so $\gamma_j = -s_j\delta_j$. Plugging in, each changepoint adds $\delta_j(t - s_j)$ after $s_j$, so $g(t) = kt + m + \sum_j\delta_j(t - s_j)_+$, a sum of ramps. Without γ the line would jump by $s_j\delta_j$, which depends on where time zero is."
C. Times $t = 0, 1, \dots, 5$ and changepoints $s = (1, 3)$. Write $A$ and the ramp matrix $R$, and compute $g$ for $k = 2$, $m = 0$, $\delta = (-1, -1)$.
- Rows of $A$: $t = 0$: (0, 0); $t = 1, 2$: (1, 0); $t = 3, 4, 5$: (1, 1).
- Ramps $R = ((t - 1)_+, (t - 3)_+)$: (0, 0), (0, 0), (1, 0), (2, 0), (3, 1), (4, 2).
- $g = 2t + R\delta$: $t = 0$: 0; 1: 2; 2: $4 - 1 = 3$; 3: $6 - 2 = 4$; 4: $8 - 3 - 1 = 4$; 5: $10 - 4 - 2 = 4$.
- Slopes: 2, then 1 (from $t = 1$), then 0 (from $t = 3$): the trend rises and then flattens, with no jumps.
D. Prophet's default grid for 500 daily rows: how many rows are eligible, roughly how far apart are the 25 candidates, and where are the first and last ones?
- Eligible: $\lfloor500\times0.8\rfloor = 400$ rows (indices 0–399).
- 26 evenly spaced positions from 0 to 399, step $399/25 = 15.96$, rounded, first one dropped.
- Candidates: 16, 32, 48, 64, 80, …, 383, 399 (about every 16 days).
- The last 100 days (rows 400–499) have no candidates.
E. (Interview) "Your backtests got worse when you doubled the number of changepoints. Why could that happen, and what would you do?"
"More changepoints make the trend more flexible. If the slope changes are not shrunk enough, the trend starts following noise; the last segments are fitted on few days, so the final slope, which is what gets extrapolated, becomes unstable. In-sample error goes down while forecast error goes up: classic overfitting. I would keep candidates away from the end of each training window, strengthen the Laplace shrinkage (smaller $b$) or go back to fewer candidates, and pick the setting by rolling-origin forecast error at the horizon that matters, not by in-sample fit."
F. A one-bend model's posterior over the location is 0.2 (day 50), 0.5 (day 52), 0.3 (day 55); the implied final slopes are 1.0, 0.8 and 0.4. Compute the posterior-averaged final slope and its spread over locations, and say what a detector picking day 52 misses.
- Mean: $0.2\times1.0 + 0.5\times0.8 + 0.3\times0.4 = 0.2 + 0.4 + 0.12 = 0.72$.
- Spread from the location alone: $E[\text{slope}^2] = 0.2\times1 + 0.5\times0.64 + 0.3\times0.16 = 0.2 + 0.32 + 0.048 = 0.568$; variance $0.568 - 0.72^2 = 0.568 - 0.5184 = 0.0496$; sd $\approx 0.22$ orders per day.
- The detector reports 0.8 with no location spread. It misses both the shift of the mean (to 0.72) and the extra 0.22 sd, which after 60 days is about $\pm13$ orders (one sd) of trend uncertainty.
PELT change-point detection
Your forecasting model lets the trend bend at changepoints, and part of its candidate list comes from a detector called PELT. This chapter opens that detector up: what it is trying to minimise, why it needs a penalty, how dynamic programming finds the best answer, why PELT's shortcut is safe, how it fails (too many or too few changepoints), and the two things an interviewer will ask about most: the changepoints it picks are uncertain, but the Bayesian model treats them as certain; and if PELT ever sees the test period, your backtest is lying.
- Say what segmentation and a segment cost are, and compute the L2 cost of a cut by hand
- Know which change each cost can see: L2 (mean shifts) vs Normal (mean and variance), and what to do for slope changes
- Explain why the objective $\sum C + \beta m$ needs a penalty, and choose a BIC-like β (and know it depends on the scale of y)
- Run optimal partitioning (dynamic programming) by hand, and explain PELT's pruning rule and why it never throws away the winner
- Recognise over- and under-detection, penalty sensitivity, and the effect of min_size and ruptures' jump
- Explain (P0) why selection uncertainty is not propagated into the Bayesian posterior, and why PELT must never see the test period
What we need from earlier chapters: piecewise-linear trends, changepoints and the three ways of choosing them (Chapter 7.8); the sum of squared deviations and variance (Chapter 4.5); the likelihood and the Normal log-likelihood (Chapter 5.2); time-series splits and leakage (Chapter 7.1). Words used throughout: a changepoint is a time where the behaviour of the series changes (its level, its spread, or its slope); detection means finding those times from the data alone; an algorithm is a step-by-step recipe a computer follows; a penalty is a fixed price we add to the score for each changepoint we use.
Segmentation: cutting a series into calm pieces core
Look at a shop's daily orders. For two weeks they hover around 10. Then a campaign starts and they hover around 30. Nobody needs statistics to see that "something changed on day 3". Changepoint detection is the computer's version of that glance: find the days where the story of the series changes.
The computer does it by cutting the timeline into pieces called segments, so that inside each piece the series behaves the same way (for now: it wobbles around one level). The days where we cut are the changepoints. Choosing the cuts is called segmentation.
To compare two ways of cutting we need a score. Inside each segment, draw a flat line at the segment's average and measure how badly the points miss it. That "badness of fit" is the segment's cost. Good cuts make every piece calm, so the total cost is small.
Three ways to say it:
- Picture: lay a few flat rulers along the series; put the joins where the series jumps, and every ruler fits snugly.
- Numbers: for 10, 12, 11, 30, 31, 29, one flat line costs 545.5; cutting between 11 and 30 leaves two calm pieces that cost 2 + 2 = 4.
- Slogan: a good segmentation makes every piece boring.
Six days of orders: $y = 10, 12, 11, 30, 31, 29$. We number days from 0 (as Python does), so day 3 is the value 30. The cost of a segment is the sum of squared distances from the segment's own mean.
- No cut (one segment): mean $= (10+12+11+30+31+29)/6 = 123/6 = 20.5$. Distances: $-10.5, -8.5, -9.5, 9.5, 10.5, 8.5$. Squares: $110.25, 72.25, 90.25, 90.25, 110.25, 72.25$. Cost $= 545.5$.
- Cut before day 3 (segments $10, 12, 11$ and $30, 31, 29$): first mean 11, squares $1, 1, 0$, cost 2. Second mean 30, squares $0, 1, 1$, cost 2. Total cost $= 2 + 2 = 4$.
- Cut before day 2 instead (segments $10, 12$ and $11, 30, 31, 29$): first cost $1 + 1 = 2$; second mean $101/4 = 25.25$, squares $203.06, 22.56, 33.06, 14.06$, cost $272.75$. Total $274.75$: a bad cut, because the 11 is forced into the "busy" piece.
- So the cut before day 3 is the best single cut: the cost dropped from 545.5 to 4. We say "one changepoint at $\tau = 3$".
Let $y_0, y_1, \dots, y_{n-1}$ be a series. A segmentation with $m$ changepoints is a list of cut positions $0 \lt \tau_1 \lt \tau_2 \lt \dots \lt \tau_m \lt n$. Add $\tau_0 = 0$ and $\tau_{m+1} = n$. Segment $j$ is $y_{\tau_j:\tau_{j+1}} = (y_{\tau_j}, \dots, y_{\tau_{j+1}-1})$: it starts at day $\tau_j$ and stops just before day $\tau_{j+1}$ (Python slicing).
A segment cost $C(\cdot)$ is a number that is small when one simple model fits the segment well. The L2 cost (mean-shift cost) of a segment of length $L$ with mean $\bar y$ is
$$C_{L2}(y_{a:b}) = \sum_{t=a}^{b-1} (y_t - \bar y_{a:b})^2 = L \cdot \hat\sigma^2_{a:b}, \qquad \hat\sigma^2_{a:b} = \text{the segment's variance with divisor } L.$$- $\tau_j$ is the first day of a new segment. The library
rupturesinstead returns the ends of the segments, including $n$: for the example it returns[3, 6]. - Costs add up: the cost of a segmentation is $\sum_{j=0}^{m} C(y_{\tau_j:\tau_{j+1}})$.
- Cutting never increases the L2 cost: fitting two means can only fit better than one. (Remember this: it is the reason a penalty is needed, and the reason PELT's shortcut works.)
- The L2 cost assumes the noise has the same size everywhere and only the level changes.
Why do we need it?
"The trend changed around here" must become a precise question before a computer can answer it. Segmentation plus a cost turns it into one: find the cuts that make the pieces fit best. Every detector (PELT, binary segmentation, window methods) is built on this score.
Where is it used?
Changepoint candidates for Prophet-style trends (your model), finding regime changes in demand or prices, detecting when a metric's level shifted after a release, monitoring sensors, and splitting genomic copy-number data.
How is it used?
Pick a cost that matches the change you care about (level, spread, slope), compute it for each candidate piece, and let an algorithm search for the cuts. In Python: ruptures.Pelt(model="l2") uses exactly the L2 cost above.
"A changepoint is a weird day, like an outlier."
An outlier is one odd value; afterwards the series goes back to normal. A changepoint is a lasting change: the days after it all follow a new rule (new level, new spread, new slope).
"ruptures says [3, 6], so there are two changepoints."
ruptures lists the end of every segment, and the last one is always $n$. [3, 6] means one changepoint: a new segment starts at index 3.
"The cost of a segment is its variance, so long and short segments are comparable."
The L2 cost is the sum of squares, which is $L$ times the variance. A segment twice as long with the same wobble costs about twice as much. That is what we want: every day's miss counts once.
In your forecasting model the trend $g(t)$ bends at changepoints $s_j$ with slope changes $\delta_j$ (Chapter 7.8). Part of the list of $s_j$ comes from PELT. PELT itself knows nothing about your model: it only sees a series and a cost, and returns cut positions. Keep this separation clear in an interview: "PELT proposes where; the Bayesian model estimates how much (the δⱼ)".
Segmentation = cut positions $0 \lt \tau_1 \lt \dots \lt \tau_m \lt n$; segment $j$ = $y_{\tau_j:\tau_{j+1}}$.
L2 cost of a segment = $\sum (y_t - \bar y_{\text{seg}})^2$; total = sum over segments. More cuts never raise it.
Trap: ruptures returns segment ends including $n$.
Quick check: what is the L2 cost of the segment 4, 6, 8?
Mean $= 18/3 = 6$; squared distances $4, 0, 4$; cost $= 8$ (= $L \times \hat\sigma^2 = 3 \times 8/3$).
Which change can the cost see? L2 (mean) vs Normal (mean and variance)
A cost is a pair of glasses. It only sees the kind of change it was built for. The L2 cost measures distances from a flat line, so it sees changes in level. It is blind to a change in spread: if orders keep the same average but become much more erratic (say, after a new delivery partner), every segment still has the same mean, and cutting does not lower the L2 cost.
The Normal cost fits a Normal distribution to each segment, with its own mean and its own standard deviation, and scores how well it fits. A calm piece fitted with a small σ scores very well, so cutting between "calm" and "wild" pays off even when the means are equal.
Three ways to say it:
- Picture: L2 glasses see the height of the cloud; Normal glasses see its height and its thickness.
- Numbers: for 10, 10.5, 9.5, 10, 4, 16, 7, 13 a cut in the middle saves 0 with L2 but 15.3 with the Normal cost.
- Slogan: choose the cost for the change you are hunting.
Eight days: $10, 10.5, 9.5, 10$ (calm) then $4, 16, 7, 13$ (wild). Both halves have mean 10.
- L2, no cut: squares from 10 are $0, 0.25, 0.25, 0, 36, 36, 9, 9$; cost $90.5$.
- L2, cut in the middle: the halves cost $0.5$ and $90$; total $90.5$. The cut saves nothing: L2 cannot see this change.
- Normal cost of a segment $= L \log \hat\sigma^2$ (with $\hat\sigma^2$ the variance, divisor $L$). No cut: $\hat\sigma^2 = 90.5/8 = 11.31$, cost $8 \log 11.31 = 19.41$.
- Cut in the middle: calm half $\hat\sigma^2 = 0.5/4 = 0.125$, cost $4 \log 0.125 = -8.32$; wild half $\hat\sigma^2 = 90/4 = 22.5$, cost $4\log 22.5 = 12.45$. Total $4.14$.
- The cut saves $19.41 - 4.14 = 15.27$. (Negative costs are fine: only differences matter.)
The Normal cost (mean-and-variance cost) of a segment of length $L$ is twice its negative maximised Normal log-likelihood, with constants dropped:
$$C_{N}(y_{a:b}) = L \log \hat\sigma^2_{a:b}, \qquad -2\max_{\mu,\sigma}\log p(y_{a:b}\mid\mu,\sigma) = L\log(2\pi\hat\sigma^2_{a:b}) + L.$$- The constants $L\log 2\pi + L$ add up to $n\log 2\pi + n$ over any segmentation, so they never change which cuts win.
- The L2 cost is the same idea with the variance held fixed at one value σ² for all segments: $-2\log p = \frac{1}{\sigma^2}\sum (y_t - \bar y)^2 + \text{const}$. So L2 = "Normal noise, one σ, only the mean may change".
- Other costs exist for other changes:
"l1"(median shifts, robust to outliers),"rbf"(any change in distribution),"linear"(a regression inside each segment; slope changes, concept 8),"ar"(changes in autocorrelation). - A Normal cost needs segments long enough to estimate a variance; a 2-point segment with two almost equal values has $\hat\sigma^2 \approx 0$ and a hugely negative cost. Use a larger minimum segment length with it (here 5).
Why do we need it?
A detector can only find the changes its cost can see. Use L2 on a series whose noise suddenly grows and it will "explain" the extra noise with many fake level changes; use the Normal cost and it finds the single real change in spread.
Where is it used?
ruptures models "l2", "normal", "l1", "rbf", "linear"; the R package changepoint (cpt.mean, cpt.var, cpt.meanvar); volatility regimes in finance; demand series whose noise changes with a new fulfilment process.
How is it used?
Ask "what kind of change matters for my model?" For trend changepoints the level/slope matters, so L2 or a linear cost; if the noise level shifts too, try "normal" with a larger min_size and look at the residuals per segment.
"L2 found no change at day 80, so nothing changed there."
L2 cannot see changes in spread at all. "No changepoint found" always means "no change of the kind my cost measures, big enough to pay the penalty".
"The Normal cost is better, so always use it."
It has more freedom (a variance per segment), so it needs longer segments and more data per change, and it is fooled by outliers (one huge value inflates a segment's σ̂). Use it when spread changes matter.
For your trend, what matters is a change in level or slope of $g(t)$, so a mean-type or linear cost is the natural fit. But if your demand series also changes its noise level (common with counts: busier periods are noisier, the reason you have a Negative Binomial likelihood, Chapter 7.13), an L2 detector may report extra "changepoints" in the busy periods. Check which model= your PELT call uses and look at where its cuts fall.
L2 cost $=\sum(y-\bar y)^2$: sees level shifts; assumes one noise level.
Normal cost $= L\log\hat\sigma^2$: sees level and spread shifts; needs longer segments.
Trap: "no changepoint found" only means "none of the kind my cost can see".
Quick check: you run L2 PELT on daily orders whose noise doubles in December. What do you expect to see?
Extra changepoints inside December: L2 assumes one noise level, so the larger wobbles look like small level changes that are worth paying the penalty for. A Normal cost (or modelling the variance) would show one change in spread instead.
The penalty β: why "more cuts" must cost something core
If the only goal were a small total cost, the computer would cheat: put every day in its own segment. Each one-day segment fits its own mean perfectly, the cost is 0, and we have learned nothing (every day is a "changepoint").
So we charge a fixed price per cut, the penalty β. Now a cut is worth making only if it lowers the cost by more than β. A real jump lowers the cost a lot and easily pays; a cut that only tidies up noise lowers it a little and does not pay. β is the bar a change must clear to be believed.
Three ways to say it:
- Picture: each cut is a toll gate; a cut gets built only if it saves more than the toll.
- Numbers: on 10, 12, 11, 30, 31, 29 the first cut saves 541.5 (pays any sensible toll); a second cut saves only 1.5 (pays only if β is below 1.5).
- Slogan: fit + price × number of cuts; the price stops you from fitting noise.
Same six days, $10, 12, 11, 30, 31, 29$, with known noise sd $\sigma = 1$ and the BIC-like penalty $\beta = 2\sigma^2\log n = 2 \times 1 \times \log 6 = 3.58$ (defined below).
- Best cost with $m = 0, 1, 2, 3, 4, 5$ cuts (found by trying every placement): $545.5,\ 4,\ 2.5,\ 1,\ 0.5,\ 0$. The cost keeps falling as cuts are added; with 5 cuts every day is alone and the cost is 0.
- Add $\beta m$: $545.5 + 0 = 545.5$; $\;4 + 3.58 = 7.58$; $\;2.5 + 7.17 = 9.67$; $\;1 + 10.75 = 11.75$; $\;0.5 + 14.33 = 14.83$; $\;0 + 17.92 = 17.92$.
- The smallest total is $7.58$ at $m = 1$: one changepoint, at day 3. The second cut would save $4 - 2.5 = 1.5 \lt 3.58$: not worth its price.
- With $\beta = 0$ the winner would be $m = 5$ (cost 0): every day a changepoint. With $\beta = 600$ the winner would be $m = 0$: even the obvious jump (saving 541.5) cannot pay.
Penalised changepoint objective. Over all numbers of changepoints $m$ and all positions $\tau_1 \lt \dots \lt \tau_m$, minimise
$$\sum_{j=0}^{m} C\big(y_{\tau_j:\tau_{j+1}}\big) \;+\; \beta\, m .$$- $\beta \gt 0$ is the penalty: the price of one changepoint, in the same units as the cost (for L2: units of $y^2$).
- BIC-like choice for L2. With Normal noise of known sd σ, $-2\log(\text{likelihood}) = \frac{1}{\sigma^2}\sum (y_t - \text{segment mean})^2 + \text{const}$. The BIC adds $\log n$ for each parameter. A changepoint adds a new mean and a new location: 2 parameters. In cost units that is $\beta = 2\sigma^2 \log n$ (some people count only the mean: $\sigma^2\log n$). For the Normal cost, which is already in $-2\log$ units, a change in mean and variance adds 3 parameters: $\beta = 3\log n$.
- σ is unknown in practice. A robust estimate uses the day-to-day differences (a level shift makes only one big difference): $\hat\sigma = 1.4826 \cdot \text{MAD}(y_t - y_{t-1})/\sqrt2$, where MAD is the median absolute deviation and the $\sqrt2$ is there because a difference of two noise values has variance $2\sigma^2$.
- These are rules of thumb, not laws. They assume independent Normal noise; with autocorrelated noise (common in daily data, Chapter 7.3) they detect too many changes.
Why do we need it?
Without a price per cut, the best "segmentation" puts every day in its own segment. The penalty turns "fit as well as possible" into "fit well with as few changes as the data can justify", a bias–variance trade-off.
Where is it used?
Every penalised detector: ruptures (predict(pen=β)), R's changepoint (penalty="BIC", "MBIC", "Manual"), the BIC/AIC model-selection criteria, and lasso-style penalties in general (Chapter 7.10 uses the same idea for slope changes).
How is it used?
Estimate σ̂ (for example from the differences), set $\beta = 2\hat\sigma^2\log n$ as a starting point, then look at how the number of changepoints changes as you move β up and down (concept 6), and check that the cuts make sense on a plot.
"β = 10 is a sensible default."
β is in the units of the cost. For L2 that is units of $y^2$: measure orders in hundreds instead of units and every cost shrinks 10 000 times, so the same β suddenly allows almost nothing. Always set β relative to $\hat\sigma^2$ (or standardise y first and say so).
"The BIC penalty is the statistically correct one."
It is a rule of thumb derived for independent Normal noise and large n. With autocorrelated or heavy-tailed noise it over-detects; with many small changes it under-detects. Treat it as a starting point, then look at sensitivity.
"A smaller total cost always means a better segmentation."
Only once the penalty is included. Without it, more cuts always win.
In your pipeline, find the line that calls PELT and note three things: the penalty value (or how it is computed), whether it is tied to an estimate of the noise level, and whether y was scaled before the call. If the series is standardised first (for example by a global scaler), $\hat\sigma^2 \approx$ the noise variance in standard units, and $2\log n$ (about 13 for two years of daily data) is the BIC-like scale; if it runs on raw orders, β must be in orders².
Objective: $\min \sum_{j=0}^m C(y_{\tau_j:\tau_{j+1}}) + \beta m$. A cut is kept only if it saves more than β.
BIC-like for L2: $\beta = 2\hat\sigma^2\log n$ (Normal cost: $3\log n$). σ̂ from the differences: $1.4826\,\text{MAD}(\Delta y)/\sqrt2$.
Trap: β lives in cost units ($y^2$ for L2): rescaling y changes its meaning.
Quick check: daily orders have noise sd 4 and you have 365 days. What BIC-like β would you start with for an L2 cost?
$\beta = 2 \times 4^2 \times \log 365 = 32 \times 5.90 \approx 189$ (in orders²). A level shift over a segment boundary is kept only if it lowers the sum of squares by more than about 189.
Optimal partitioning: finding the best cuts by dynamic programming core
How many ways can you cut 365 days? Each of the 364 gaps is "cut" or "no cut": $2^{364}$ ways, more than the atoms in the universe. Trying them all is impossible. Yet a simple trick finds the best one exactly.
The trick: think about the last segment. Whatever the best segmentation of days $0$ to $t-1$ is, its last segment starts at some day $s$. Everything before $s$ must then be the best segmentation of days $0$ to $s-1$ (if it were not, swapping in the better one would improve the whole). So if we already know the best score for every shorter stretch, the best score for $0..t-1$ is "best score up to $s$ + cost of the last piece + β", minimised over the start $s$. We fill in these best scores left to right, like filling a table. That is dynamic programming: solve small problems once, reuse them.
Three ways to say it:
- Picture: to find the best route to town t, look at every town s you could have come from: best route to s + one last hop.
- Numbers: for 1, 3, 10, 12 the table is $F = 0, 2, 7, 9$; the last value, 9, is the best total and its last segment starts at day 2.
- Slogan: the best answer ends with a last segment, and everything before it is a smaller best answer.
Data $y = 1, 3, 10, 12$ (days 0–3), L2 cost, $\beta = 5$, segments may have length 1. Let $F(t)$ be the best penalised score of the first $t$ values, with the bookkeeping start $F(0) = -\beta = -5$ (the first segment has no changepoint in front of it, so it gets its β back).
- Costs we need: $C(1) = C(3) = C(10) = C(12) = 0$; $C(1,3) = 2$; $C(10,12) = 2$; $C(3,10) = 24.5$; $C(1,3,10) = 44.67$; $C(3,10,12) = 44.67$; $C(1,3,10,12) = 85$.
- $F(1)$: only start $s = 0$: $F(0) + C(1) + 5 = -5 + 0 + 5 = 0$.
- $F(2)$: $s = 0$: $-5 + C(1,3) + 5 = 2$; $\;s = 1$: $F(1) + C(3) + 5 = 5$. Best $F(2) = 2$ (start 0: no cut yet).
- $F(3)$: $s = 0$: $-5 + 44.67 + 5 = 44.67$; $\;s = 1$: $0 + 24.5 + 5 = 29.5$; $\;s = 2$: $2 + 0 + 5 = 7$. Best $F(3) = 7$, last segment starts at 2.
- $F(4)$: $s = 0$: $85$; $\;s = 1$: $0 + 44.67 + 5 = 49.67$; $\;s = 2$: $2 + 2 + 5 = 9$; $\;s = 3$: $7 + 0 + 5 = 12$. Best $F(4) = 9$, last segment starts at 2.
- Walk back: the last segment of $0..3$ starts at 2; the best for $0..1$ starts at 0. Segments $(1, 3)$ and $(10, 12)$: one changepoint at $\tau = 2$, total $2 + 2 + 5 = 9$. We evaluated $1 + 2 + 3 + 4 = 10$ candidate costs.
Optimal partitioning (OP) (Jackson et al., 2005) computes, for $t = 1, \dots, n$,
$$F(t) = \min_{0 \le s \lt t}\Big[\,F(s) + C\big(y_{s:t}\big) + \beta\,\Big], \qquad F(0) = -\beta,$$and stores the minimising $s$ as $\text{last}(t)$. Then $F(n)$ equals the minimum of $\sum_j C(y_{\tau_j:\tau_{j+1}}) + \beta m$, and following $\text{last}(n), \text{last}(\text{last}(n)), \dots$ back to 0 gives the changepoints.
- It is exact: it returns the true minimiser of the objective (not an approximation like binary segmentation).
- Cost of the search: at time $t$ it evaluates $t$ candidates, so about $n^2/2$ segment costs in total: $O(n^2)$. With cumulative sums each L2 cost takes constant time.
- With a minimum segment length $L_{\min}$, only starts with $t - s \ge L_{\min}$ are allowed.
Why do we need it?
The number of possible segmentations explodes ($2^{n-1}$). Dynamic programming finds the exact best one in about $n^2/2$ steps, which is fine for a few thousand points; PELT (next concept) then makes it much faster.
Where is it used?
The engine inside PELT and inside ruptures.Dynp (a fixed number of changepoints), the same "best last step" idea as the Viterbi algorithm for hidden Markov models and as shortest-path algorithms.
How is it used?
You rarely write it yourself: ruptures.Pelt runs it with pruning. Knowing it lets you explain why PELT's answer is exact for the chosen cost and penalty, and why runtime grows when there are few changes.
"Dynamic programming is a heuristic: it finds a good segmentation."
It finds the best segmentation for the chosen cost and β, exactly. Heuristics such as binary segmentation (cut once, then cut each half…) can miss the optimum.
"Exact means correct."
Exact means "the true minimiser of this objective". If the cost is wrong for the data (L2 on a trend, L2 with changing noise) or β is badly chosen, the exact answer is exactly wrong.
$F(t) = \min_{s \lt t}[F(s) + C(y_{s:t}) + \beta]$, $F(0) = -\beta$; walk back through the stored argmins.
Exact optimum of the penalised objective; about $n^2/2$ cost evaluations.
Trap: exact ≠ true changepoints; it is exact for the cost and β you chose.
Quick check: in the 4-day example, what is $F(4)$ with $\beta = 50$, and how many changepoints?
Now $F(0) = -50$, $F(1) = 0$, $F(2) = 2$, $F(3) = \min(44.67, 74.5, 52) = 44.67$, and $F(4) = \min(85,\ 94.67,\ 2 + 2 + 50 = 54,\ 94.67) = 54$. Still one changepoint at day 2, total $54 = 2 + 2 + 50$. The cut disappears only when β exceeds the saving $85 - 4 = 81$.
PELT: drop starts that can never win again core
Optimal partitioning keeps every old day as a possible start of the last segment, forever. But some old starts are hopeless. If a start $s$ is already losing to today's best by more than one penalty, it can never catch up: whatever happens later, the winner could simply add one cut at today and stay ahead of $s$.
PELT ("Pruned Exact Linear Time", Killick, Fearnhead and Eckley, 2012) throws such starts away. The list of candidates stays short, so each step is quick, and the answer is exactly the same as optimal partitioning.
Three ways to say it:
- Picture: a race where a runner more than one "penalty" behind at any checkpoint is sent home, because the leader can always pay that penalty and still win.
- Numbers: for 1, 3, 10, 12 at $t = 3$, starts 0 and 1 trail the best (7) by far more than β = 5 (39.67 and 24.5 before the β), so they are dropped: 8 evaluations instead of 10.
- Slogan: if you are behind by more than one cut, you are out for good.
Same data $1, 3, 10, 12$, $\beta = 5$. PELT runs the same recursion, then after computing $F(t)$ removes every start $s$ with $F(s) + C(y_{s:t}) \gt F(t)$.
- $t = 1$: candidates $\{0\}$; $F(1) = 0$. Check $s = 0$: $F(0) + C(1) = -5 + 0 = -5 \le 0$: keep. Add 1: candidates $\{0, 1\}$.
- $t = 2$: values $2$ (s=0) and $5$ (s=1); $F(2) = 2$. Checks: $-5 + 2 = -3 \le 2$ keep; $0 + 0 = 0 \le 2$ keep. Add 2: $\{0, 1, 2\}$.
- $t = 3$: values $44.67, 29.5, 7$; $F(3) = 7$. Checks: $s = 0$: $-5 + 44.67 = 39.67 \gt 7$: prune. $s = 1$: $0 + 24.5 = 24.5 \gt 7$: prune. $s = 2$: $2 + 0 \le 7$: keep. Add 3: $\{2, 3\}$.
- $t = 4$: only 2 evaluations: $9$ (s=2) and $12$ (s=3); $F(4) = 9$, same as before.
- Work: $1 + 2 + 3 + 2 = 8$ cost evaluations instead of 10. On long series with frequent changes the saving is huge.
PELT pruning rule. After computing $F(t)$, remove from the candidate set every $s$ with
$$F(s) + C\big(y_{s:t}\big) \;\gt\; F(t) \qquad (\text{"$s$ trails by more than one } \beta\text{"}).$$Why it is safe. For costs where splitting never hurts, $C(y_{s:T}) \ge C(y_{s:t}) + C(y_{t:T})$ for every later $T \gt t$ (true for L2 and the Normal cost: two fitted pieces fit at least as well as one). Then for every future $T$:
$$F(s) + C(y_{s:T}) + \beta \;\ge\; F(s) + C(y_{s:t}) + C(y_{t:T}) + \beta \;\gt\; F(t) + C(y_{t:T}) + \beta,$$and the right side is the score of starting the last segment at $t$. So start $t$ always beats start $s$ from now on: $s$ can never be the winner, and dropping it changes nothing.
- Exact: PELT returns the same segmentation as optimal partitioning (with segments allowed to have length 1, the setting of the proof).
- Speed: worst case still $O(n^2)$ (for example when there are no changes: nothing can be pruned). When the number of changepoints grows in proportion to $n$ (a change every few weeks, say), the expected cost is $O(n)$ under the conditions in Killick et al. Few changes in a long series → little pruning.
- The general rule has a constant $K$ ($F(s) + C + K \gt F(t)$) for costs where splitting can add up to $K$; for L2 and Normal costs $K = 0$.
Why do we need it?
Optimal partitioning on 10 000 points means 50 million cost evaluations, and you may want to rerun it in every backtest fold and for many penalties. PELT gives the same exact answer with a small fraction of the work when changes are frequent.
Where is it used?
ruptures.Pelt, R's changepoint::cpt.mean(method="PELT"), the candidate-changepoint step of your forecasting pipeline, and changepoint tools in monitoring, finance and genomics.
How is it used?
rpt.Pelt(model="l2", min_size=2, jump=1).fit(y).predict(pen=beta) returns segment ends. Choose the cost (model), the penalty (pen), the minimum segment length and the grid (jump) on purpose; the defaults are not neutral (concept 7).
"PELT is an approximation that trades accuracy for speed."
Pruning only removes starts that provably cannot win. PELT's answer equals optimal partitioning's answer. The speed comes for free.
"PELT is always linear time."
Only when changes are frequent (their number grows with n). With no changes, nothing gets pruned and PELT does about $n^2/2$ evaluations, like optimal partitioning.
"PELT finds the true changepoints."
"PELT finds the exact minimiser of a penalised cost. Whether those are the true changepoints depends on the cost, the penalty and the noise."
Model answer: "PELT solves the penalised segmentation problem $\min \sum C + \beta m$ exactly, by dynamic programming over the start of the last segment, and it prunes any start that already trails the current best by more than β: such a start can never win again, because splitting never increases the cost. That makes it about linear time when there are many changes. Its answer is only as good as the cost and β: the wrong cost or a badly scaled β gives an exact but wrong segmentation."
Prune $s$ when $F(s) + C(y_{s:t}) \gt F(t)$: it trails by more than one β and can never win (because $C(y_{s:T}) \ge C(y_{s:t}) + C(y_{t:T})$).
Same answer as optimal partitioning; ~$O(n)$ when changes are frequent, $O(n^2)$ worst case.
Trap: "exact" refers to the objective, not to the truth.
Quick check: at time $t$, $F(t) = 50$, and a start $s$ has $F(s) + C(y_{s:t}) = 53$ with $\beta = 5$. Is $s$ pruned?
Yes: $53 \gt 50$. Its candidate score $53 + 5 = 58$ trails the best (50) by 8, more than one β. It could only win later if adding a cut at $t$ cost more than β, which is impossible.
Too many or too few: over-detection, under-detection and penalty sensitivity core
A detector can be wrong in two ways. It can over-detect: report changepoints where nothing changed (false alarms), because it is explaining noise. Or it can under-detect: miss a real change (a miss), because the change was too small or too short to pay the penalty. The penalty β is the dial between the two: turn it down and false alarms pile up; turn it up and real changes disappear.
How sensitive is the answer to the dial? On clean data there is usually a wide range of β that gives the same answer (a "plateau"). If the number of changepoints changes every time you nudge β, the data do not speak clearly, and you should not trust any single answer.
Three ways to say it:
- Picture: a smoke alarm: too sensitive and it rings for toast; too dull and it sleeps through a fire.
- Numbers: on the 200-day series, β = 2 gives 48 changepoints, every β between 9.7 and 192 gives the true 3, and β = 300 gives 1.
- Slogan: look at the whole β dial, not one setting.
The 200-day series of the widgets (true changes at days 50, 90, 150; level jumps of 4, −3 and 5; noise sd 1.5), L2 PELT, a detection counts as a hit if it is within 5 days of a true change.
- $\beta = 2$: 48 changepoints. 3 hits, 45 false alarms: severe over-detection.
- $\beta = 5$: 26 changepoints (3 hits, 23 false alarms).
- $\beta = 25.75$ (BIC-like, $2\hat\sigma^2\log 200$ with $\hat\sigma = 1.56$): exactly 50, 89, 150: 3 hits, 0 false alarms.
- Every β from about 9.7 up to about 192 gives the same 3 changepoints: a wide plateau (a factor of 20), so the answer is robust.
- $\beta = 300$: only day 150 (the biggest jump): 2 misses. Above about 741: no changepoints at all.
- Over-detection (false alarms, false positives): reported changepoints with no true change nearby. Cause: β too small, the wrong cost (L2 with changing noise or on a trend), outliers, autocorrelated noise.
- Under-detection (misses, false negatives): true changes with no reported changepoint nearby. Cause: β too large, small changes, short segments, changes near the end of the series.
- Penalty sensitivity: how the solution changes with β. The number of changepoints $m(\beta)$ is a step function that can only go down as β goes up. A long flat step = a stable answer.
- Scale: for L2, multiplying y by $c$ multiplies every cost by $c^2$, so the same plateau moves to β values $c^2$ times larger. A β is meaningful only together with the scale of y.
- Tools: run PELT over a grid of β (the R package
changepointautomates this with CROPS, "changepoints for a range of penalties"), plot $m(\beta)$, and pick the middle of a plateau or the elbow.
Why do we need it?
The penalty is the least principled number in the pipeline, and it decides how many trend bends your model is allowed to consider. Knowing whether the answer is stable across β tells you whether to trust it.
Where is it used?
Choosing pen in ruptures, CROPS in R's changepoint, elbow plots of cost vs number of changepoints, and the bias–variance discussion of trend flexibility (Chapters 7.10 and 7.18).
How is it used?
Run PELT for a grid of β (say 0.25× to 4× the BIC-like value), plot the number of changepoints, mark where your chosen β sits, and report it. If there is no plateau, prefer more candidates and let the Laplace prior (7.10) shrink the useless ones.
"PELT found 3 changepoints, so the series has 3 changes."
It found 3 at this β. Report the range of β that gives the same answer; if the count keeps moving, the number of changes is genuinely uncertain.
"False alarms are harmless; the model can ignore them."
Each false alarm gives the trend one more place to bend. With a shrinkage prior on the slope change (Chapter 7.10) a useless candidate costs little; with no shrinkage it lets the trend chase noise and makes the extrapolated slope unstable.
"Misses are rare if β is the BIC value."
Small changes, short segments and changes near the end of the series are often missed at any reasonable β: a change in the last few days simply has too few points after it to pay the penalty.
In a "grid + PELT" design like yours, the two kinds of error cost different things. A false alarm adds one more candidate whose $\delta_j$ the Laplace prior can shrink toward 0. A miss removes a place where the trend could bend, and no prior can bring it back (unless the grid happens to have a nearby candidate). That is a reasonable argument for erring toward a lower β when PELT only proposes candidates, a judgement call to state and to check by backtest, not a rule.
Small β → over-detection (false alarms); large β → under-detection (misses).
Plot $m(\beta)$ over a β grid; trust a long plateau. For L2, rescaling y by $c$ moves everything by $c^2$.
Trap: one β, one answer, no sensitivity check.
Quick check: you change the unit of your series from orders to thousands of orders but keep β. What happens, and why?
Every L2 cost shrinks by $1000^2 = 10^6$, so a fixed β is now enormous relative to the savings: PELT finds far fewer changepoints (probably none). The penalty must be rescaled by $10^{-6}$ too, which happens automatically if β is computed as $2\hat\sigma^2\log n$ from the data.
Minimum segment length, outliers, and ruptures' jump
A single crazy day (a data glitch, a one-off bulk order) is not a change in the trend. But a detector that may use segments of length 1 can "explain" it perfectly: cut just before it and just after it, and the crazy day sits in its own segment with cost 0. If the spike is big enough, those two cuts pay for themselves, and you get two fake changepoints.
A minimum segment length forbids pieces shorter than $L_{\min}$ days. A lone spike can no longer have its own segment, so it stops producing changepoints. The price: two real changes closer than $L_{\min}$ cannot both be found, and a change in the last $L_{\min} - 1$ days cannot be found at all.
ruptures has a second, easy-to-miss setting: jump (default 5). It only considers changepoints at every 5th index, for speed. A real change on day 72 is then reported at day 70 or 75.
Three ways to say it:
- Picture: a minimum length is a rule "no piece shorter than a week";
jumpis a ruler that only has marks every 5 days. - Numbers: one spike of +8 (noise sd 1) creates changepoints at 30 and 31 when $L_{\min} = 1$, and none when $L_{\min} = 2$; with
jump=5a change at 72 comes back as 70. - Slogan: set
min_sizefor robustness andjump=1for precise dates.
100 days, level 10 then 13 from day 72, noise sd 1, plus one spike of +8 on day 30. σ̂ (from the differences) $= 1.17$, β $= 2\hat\sigma^2\log 100 = 12.6$.
- $L_{\min} = 1$: PELT returns 30, 31, 72. The spike gets its own one-day segment: removing it from its neighbours saves about $8^2 = 64$ in cost, more than the price of two cuts ($2 \times 12.6 = 25.2$).
- $L_{\min} = 2$ (ruptures' default): PELT returns only 72. A spike can only sit in a segment of at least 2 days, and that saves too little to pay for two cuts.
- $L_{\min} = 2$ with
jump=5: PELT returns 70. The real change is on day 72, but only multiples of 5 are allowed.
- Minimum segment length $L_{\min}$ (
min_sizein ruptures; default 2 for"l2"): only segmentations whose segments all have at least $L_{\min}$ points are allowed. In the recursion: $F(t) = \min_{s:\ t - s \ge L_{\min}}[\dots]$. jump(ruptures; default 5): candidate changepoints are restricted to multiples ofjump. It divides the work by aboutjump² but rounds locations. Usejump=1whenever the exact day matters (it does for trend changepoints).- Robustness to outliers comes from $L_{\min}$, from a robust cost (
"l1"), or from cleaning outliers first; the L2 cost itself is very sensitive to them (it squares the misses). - A fine point: Killick's proof that pruning is exact assumes segments of any length are allowed. With $L_{\min} \ge 2$, ruptures (and the PELT used in this guide, which matches it) can very occasionally return a slightly worse segmentation than full optimal partitioning; in our tests this happened only with tiny penalties that were already over-segmenting.
Why do we need it?
Daily demand has glitches and one-off spikes. Without a minimum length they become pairs of fake changepoints; with the wrong jump the real ones land on the wrong dates. Both settings change the candidate list your trend gets.
Where is it used?
rpt.Pelt(model=..., min_size=..., jump=...), minseglen in R's changepoint, and Prophet-style pipelines that require a minimum gap between candidate changepoints.
How is it used?
Set jump=1; choose min_size as the shortest stretch you would believe is a real regime (for daily data often a week or more); handle known outliers (holidays, data errors) before detection, or model them separately (Chapter 7.12).
"I used the defaults, so the changepoint dates are exact."
rpt.Pelt(model="l2") uses jump=5: every reported changepoint is a multiple of 5. Pass jump=1.
"A bigger min_size is always safer."
It removes outlier cuts but also hides changes that are close together and, importantly for forecasting, any change in the last min_size − 1 days: exactly where a new trend matters most.
Check your PELT call for min_size and jump. If jump is left at the default, your PELT changepoints sit on multiples of 5 days; that is usually harmless when they are only candidates with Laplace-shrunk slopes, but it is worth knowing (and saying). If holidays produce spikes, either remove the holiday effect before detection or make sure min_size stops them from becoming trend candidates; otherwise the trend and the holiday terms compete for the same bump (identifiability, Chapter 6.8).
min_size = shortest allowed segment: stops one-day spikes from becoming two fake changepoints; hides changes in the last min_size − 1 days.
jump (default 5 in ruptures) rounds changepoints to multiples of 5: use jump=1.
Trap: trusting the defaults' dates.
Quick check: with min_size = 7 and 365 days of data, what is the latest day on which PELT could report a new segment starting?
The last segment needs at least 7 points, so it can start at day $365 - 7 = 358$ at the latest (days 358–364). A change on day 360 cannot be reported yet.
Trend changes are slope changes: what should PELT look at?
Your trend is piecewise-linear: at a changepoint the slope changes, while the line stays connected. The L2 cost, however, looks for flat pieces. Feed it a steadily growing series and it sees a staircase: a sloped line is badly fitted by one flat mean, and cutting it into short flat steps lowers the cost a lot. You get many changepoints on a trend that never changed.
There are three common fixes. (1) Run L2 PELT on the day-to-day differences $y_t - y_{t-1}$: a slope change in y is a level change in the differences. But differencing doubles the noise variance, so gentle slope changes drown. (2) Use a linear cost: fit a straight line (not a flat mean) inside each segment. (3) Remove the trend or seasonality first and look for changes in what is left.
Three ways to say it:
- Picture: flat rulers on a ramp make a staircase; tilted rulers lie flat on it.
- Numbers: the perfectly straight series 0, 1, …, 9 has no change at all, yet one L2 cut in the middle "saves" 82.5 − 20 = 62.5.
- Slogan: match the cost to the shape of your trend.
Straight line $y = 0, 1, 2, \dots, 9$ (slope 1, no noise, no change).
- L2, no cut: mean 4.5; squares $20.25, 12.25, 6.25, 2.25, 0.25, 0.25, 2.25, 6.25, 12.25, 20.25$; cost $82.5$.
- L2, cut in the middle: each half (0–4 and 5–9) has squares $4, 1, 0, 1, 4$, cost 10; total $20$. "Saving" $62.5$: any β below that reports a fake changepoint.
- Differences: $1, 1, \dots, 1$, all equal: L2 cost 0, no cut can help: no changepoint (correct).
- Linear cost: a straight line fits perfectly: cost 0, no changepoint (correct).
- Differences: if $y_t = g(t) + \epsilon_t$ with a slope change δ at $s$, then $\Delta y_t = y_t - y_{t-1}$ has mean $k$ before $s$ and $k + \delta$ after: a mean shift of size δ. Its noise $\epsilon_t - \epsilon_{t-1}$ has variance $2\sigma^2$ and is negatively correlated from day to day, so detection power drops and the BIC-like β must use $2\hat\sigma^2$.
- Linear (regression) cost: $C_{lin}(y_{a:b}) = \min_{\alpha,\gamma}\sum_{t=a}^{b-1}(y_t - \alpha - \gamma t)^2$. A change adds a slope, an intercept and a location: BIC-like $\beta = 3\hat\sigma^2\log n$. In ruptures:
rpt.Pelt(model="linear")on the columns[y, t, 1]. This cost lets the line jump at a changepoint, unlike Prophet's continuous trend; ruptures'"clinear"cost forces continuity in a simple way. - Detrend / deseasonalise first: subtract a fitted trend or seasonal pattern and run a level or variance cost on what remains. Weekly seasonality left in the series also produces fake changepoints with any of these costs.
Why do we need it?
The changepoints your model uses are places where the slope of g(t) may change. A detector that hunts for level steps on a trending series hands the model a staircase of meaningless candidates and may miss the real bends.
Where is it used?
Candidate generation for Prophet-style trends, ruptures' "linear" and "clinear" costs, trend-break detection in economics (structural breaks), and the R packages for piecewise-linear changepoints.
How is it used?
Decide what PELT sees: differences (simple, noisy), a linear cost (direct, needs larger min_size), or residuals after removing seasonality. Plot the detected cuts on the raw series and ask: are these bends in the trend, or steps of a staircase?
"PELT found 9 changepoints in my trend, so the trend changed 9 times."
If PELT ran with an L2 (level) cost on a trending series, most of those are staircase steps of one straight line. Check which cost and which signal it saw.
"Differencing is the clean way to turn slope changes into level changes."
It is correct on paper, but it doubles the noise variance and makes neighbouring differences negatively correlated; small slope changes become very hard to detect.
This is a question to answer precisely about your own code: what series does PELT see, and with which model=? The raw (scaled) series with "l2", its differences, a deseasonalised series, or a linear cost all give different candidates. None is "the" right one; what matters is that you know which one you use and why. If it is an L2 cost on the raw trending series, expect staircase candidates, and expect the Laplace prior on δⱼ (7.10) to be doing much of the real selection.
Trend changes = slope changes. L2 on a trend → staircase of fake cuts.
Options: L2 on $\Delta y$ (noise var $2\sigma^2$), a linear cost (β ≈ $3\hat\sigma^2\log n$), or detrend/deseasonalise first.
Trap: not knowing which signal and cost your pipeline feeds PELT.
Quick check: a slope changes from 2 to 3 units per day, and daily noise has sd 4. In the differences, how big is the level shift compared with the noise?
The differences jump from mean 2 to mean 3: a shift of 1, while their noise sd is $\sqrt2 \times 4 \approx 5.7$. The shift is less than a fifth of the noise sd per day, so it needs many days on each side to be detected. A linear cost uses the information much better.
The key limitation: PELT's uncertainty never reaches the posterior core
PELT gives one answer: "the trend changed on day 97". But the data are noisy. With a different few weeks of noise from the very same process, PELT might have said day 92, or day 104, or "no change at all". The location of a changepoint is itself an uncertain estimate.
Now the Bayesian model takes "day 97" as a fixed, known input and estimates the slope change there. Its posterior is honest about everything given day 97, and completely silent about "was it really day 97?". The uncertainty from the selection step is not propagated (not passed on) into the posterior. Forecast intervals come out too narrow, most of all when the latest changepoint is uncertain, because the latest slope is what the forecast extends.
There is a second, quieter problem: the same data were used twice, once to choose where the change is and once to measure it. PELT picks places where the noise happened to make the change look big, so measured changes come out too large on average (the "winner's curse").
Three ways to say it:
- Picture: a weather map drawn as if the storm's position were certain, when the forecaster was only sure it was "somewhere in this region".
- Numbers: in 300 simulated copies of one series, a small real change was found in only 67% of copies, anywhere between day 88 and day 109, and measured 21% too big on average.
- Slogan: conditioning on an estimate is not the same as knowing it.
One truth, 300 noise replicates: 150 days, level 10, then 13 from day 50 (a big change: 2σ), then 14 from day 100 (a small change: 0.67σ), noise sd 1.5. Each replicate: L2 PELT with β $= 2\hat\sigma^2\log 150$, σ̂ estimated from that replicate's differences (code in the recap).
- The big change was found in every replicate; 88% of the time within ±2 days of day 50, and 90% of the time between days 48 and 53.
- The small change was found in only 67% of replicates; when found, 90% of the locations were spread between day 88 and day 109.
- The number of changepoints varied: 1 in 77 replicates, 2 in 197, 3 or more in 26.
- When the small change was found, its measured size (mean after − mean before) averaged 1.21, although the truth is 1.0: about 21% too large, because PELT kept it only when the noise made it look big.
- In each replicate the downstream model would treat its own list as certain. None of this spread would appear in its posterior.
Let $\tau$ be the changepoint locations and $\theta$ the other model parameters (k, m, δ, seasonality, σ…). The full Bayesian posterior averages over the uncertainty in τ:
$$p(\theta\mid y) = \sum_{\tau} p(\theta\mid y, \tau)\, p(\tau\mid y).$$A "detect, then fit" pipeline instead reports the plug-in posterior $p(\theta\mid y, \hat\tau(y))$, where $\hat\tau(y)$ is the PELT output.
- It ignores the spread of $p(\tau\mid y)$: credible and prediction intervals are too narrow, especially for quantities near uncertain changepoints (the latest slope, the forecast).
- It uses $y$ twice (selection, then estimation): estimates at selected locations are biased away from zero (selection bias, the winner's curse).
- It is a hard choice: a missed change cannot come back, a false one is always there (unless a shrinkage prior on its δ makes it harmless).
- Ways to reduce the problem: keep PELT points only as candidates alongside a dense grid, with $\delta_j \sim Laplace(0, b)$ so the model can shrink wrong ones (Chapter 7.10); check stability (rerun PELT on resampled or perturbed series, or over a range of β, and see which points survive); refit with alternative changepoint sets and compare forecasts (a sensitivity analysis, Chapter 6.8); or make the locations latent parameters (Bayesian changepoint models, the third approach of Chapter 7.8), which is more expensive and needs marginalisation or MCMC because locations are discrete.
Why do we need it?
It is the most important honest limitation of a "grid + PELT, then Bayesian fit" design. Without seeing it, you would read the posterior intervals as complete, and promise forecast coverage that the model cannot deliver after a recent, uncertain trend change.
Where is it used?
Any two-stage pipeline: feature or variable selection followed by a fit (post-selection inference), changepoints chosen then treated as known, hyperparameters tuned on the same data, and in interviews about your forecasting model (it is on your P0 list).
How is it used?
State it, measure it (replicates or a bootstrap of the detection step), and soften it: candidates plus shrinkage, sensitivity refits, wider intervals checked by rolling-origin coverage (Chapter 7.16), or a latent-changepoint model when the stakes justify it.
"The model is Bayesian, so its intervals include all the uncertainty."
They include the uncertainty of the parameters inside the model, given its inputs. The changepoint locations chosen by PELT are inputs, so their uncertainty is not included.
"PELT found the change on day 97, so the slope change happened on day 97."
Day 97 is a noisy estimate. A small change can easily be located a week or two off, or missed. The posterior over its location (when you can compute it) is often wide.
"The estimated slope change at a PELT point is unbiased."
PELT keeps locations where the data made the change look big; re-estimating it on the same data overstates it on average (in the example, 1.21 instead of 1.0). Shrinkage priors partly correct this.
In your forecasting model the PELT changepoints are selected on the observed series and then passed to the Bayesian model as fixed candidate locations. This means selection uncertainty is not propagated into the posterior: the SVI posterior over $k$, $m$, $\delta_j$ and the forecast is conditional on those locations. Two things in your design help: the Laplace prior on every $\delta_j$ can shrink a poor candidate toward zero, and a grid of candidates gives the trend other places to bend. What it cannot do is widen the forecast for "we are not sure where the last change was". Ways to show you have thought about it: a stability check of the PELT points across β or resampled series, refits with alternative changepoint sets, and rolling-origin coverage checks of the intervals.
"My Bayesian forecast accounts for all the uncertainty, including the changepoints."
"The posterior is conditional on the changepoint candidates chosen by PELT. The uncertainty in where the changes are is not propagated."
Model answer: "PELT chooses the changepoint locations on the observed series, and the Bayesian model treats them as known. So the posterior is $p(\theta\mid y, \hat\tau)$, not $p(\theta\mid y)$: it ignores the spread of $p(\tau\mid y)$ and it uses the data twice, so intervals near uncertain changepoints are too narrow and changes at selected points look bigger than they are. I reduce the damage by using PELT points only as candidates next to a grid, with Laplace priors that shrink unnecessary slope changes, and by checking stability and interval coverage in rolling-origin backtests. A fully Bayesian alternative treats locations as latent variables, at a higher computational cost."
Full: $p(\theta\mid y) = \sum_\tau p(\theta\mid y,\tau)p(\tau\mid y)$. Pipeline: $p(\theta\mid y, \hat\tau(y))$.
Consequences: intervals too narrow near uncertain changepoints; data used twice → selected changes look too big.
Mitigate: candidates + Laplace shrinkage, stability checks, sensitivity refits, coverage backtests, or latent changepoints.
Quick check: why does this limitation matter more for the latest changepoint than for one two years ago?
The forecast extends the latest slope. If the last change is uncertain (few points after it), both its location and the new slope are uncertain, and both drive the forecast. An old changepoint with lots of data on both sides is located precisely and barely affects the future.
Leakage: if PELT sees the test period, the backtest lies core
A backtest pretends to stand at a past date (the forecast origin), forecast the next weeks, and compare with what really happened. It is only honest if every step uses only what was known at the origin (Chapter 7.1).
Suppose a real change happened three days before the origin. Standing at the origin, those three days are all the evidence you have; PELT will often miss the change or misplace it. But if you ran PELT once on the whole series (training + test) and reused its changepoints in every fold, the test weeks confirm the change loudly, PELT places it exactly, and the backtest forecast "knows" the level moved. The backtest looks better than real life will ever be. This is leakage: future information sneaking into the past.
Three ways to say it:
- Picture: marking last week's exam answers with a pen that was dipped in next week's answer key.
- Numbers: with a change 3 days before the origin, honest PELT gives a test MAE of about 1.86; PELT run on the full series gives about 1.41, a 24% "improvement" that does not exist live.
- Slogan: everything that looks at the data, including PELT, goes inside the fold.
Training: days 0–99, test: days 100–129, level 10 then 13 from day 97, noise sd 1.5. The "model" forecasts the mean of the last training segment. 400 simulated replicates (Python code in the recap).
- Honest: PELT on days 0–99 only (β from that data). Average test MAE $\approx 1.86$. With 3 points after the change, PELT often misses or misplaces it, so the forecast mixes the old level in.
- Leaky: PELT on days 0–129, keep the changepoints before day 100, then forecast from the training data. Average test MAE $\approx 1.41$.
- The difference is pure leakage: the forecast still uses only training values, but the location of the cut was chosen with test data.
- For comparison, a perfect forecast (the true level) would have MAE $\approx 1.5\sqrt{2/\pi} \approx 1.20$. If the change is 15 days before the origin, the gap almost vanishes (1.28 vs 1.24 in the widget's simulation): leakage matters most exactly when a change is recent.
Leakage through changepoint detection happens when any input to the model of a backtest fold (changepoint locations, their number, the penalty β, the noise estimate σ̂, the scaling of y) was computed using data after that fold's forecast origin.
- Rule: in rolling-origin evaluation (Chapter 7.15), rerun the whole pipeline for each origin: scaling, PELT with its penalty, candidate grid, fit, forecast, using only data up to that origin.
- The honest backtest also shows a real property of the live system: changes right before the origin are detected late. That delay is part of your forecast error, and it must appear in the backtest.
- Related leaks: choosing β or the number of changepoints by looking at the full-history plot; tuning the Laplace scale b on the test period; holiday or regressor features built from the full series (Chapter 7.12).
Why do we need it?
A leaky backtest overstates accuracy and coverage, exactly in the situations (recent trend changes) where the live model will do worst. Decisions made on it (which model to ship, how much safety stock to hold) are made on fantasy numbers.
Where is it used?
Rolling-origin backtests of Prophet-style models, any pipeline with a data-driven preprocessing step (detectors, scalers, feature selection), and interview questions about how you validated the forecasting model.
How is it used?
Wrap PELT inside the fold function: for origin in origins: train = y[:origin]; cps = pelt(train); model = fit(train, cps); score(forecast, y[origin:origin+h]). Never compute changepoints once, outside the loop.
"The model was only trained on the training data, so there is no leakage."
Leakage can enter through any preprocessing step. If the changepoint locations, the penalty or the scaling were computed with test data, information from the test period is in the model.
"Running PELT once on the full history is just a convenience."
It is a convenience that inflates backtest accuracy most where it matters most: right after real changes. The honest version also measures the detection delay your live system really has.
When you describe your backtests, be ready for: "Did PELT run inside each fold?" The honest answer for a correct pipeline is: "Yes: for each forecast origin I rescale, run PELT with its penalty, build the candidate list and fit SVI using only data up to that origin." If it ran once on the full series, say so and treat the backtest numbers as optimistic, especially for origins right after trend changes (Chapter 7.12 covers the other leakage routes, 7.15 the rolling-origin set-up).
Leakage = any input to a fold (changepoints, β, σ̂, scaling) computed with data after the origin.
Rule: run PELT (and everything else data-driven) inside every backtest fold.
Trap: "the model only saw training data" while its changepoints were placed using the test period.
Quick check: your backtest MAE is much better for origins just after trend changes than live performance. What is the first thing you check?
Whether the changepoints (or β, σ̂, the scaler) were computed on the full series. A detector that sees the test period places recent changes perfectly in the backtest, while live it needs several days of evidence after a change.
Recap, cheat sheet and practice
- Segmentation cuts a series into segments; a cost scores how well one simple model fits each piece. L2 sees level shifts; the Normal cost sees level and spread; a linear cost sees slope changes.
- The objective is $\sum_j C(y_{\tau_j:\tau_{j+1}}) + \beta m$. Without the penalty β, every day becomes a changepoint. BIC-like for L2: $\beta = 2\hat\sigma^2\log n$, in the units of $y^2$.
- Optimal partitioning solves it exactly by dynamic programming over the start of the last segment: $F(t) = \min_s[F(s) + C(y_{s:t}) + \beta]$, about $n^2/2$ evaluations.
- PELT prunes starts with $F(s) + C(y_{s:t}) \gt F(t)$: they trail by more than one β and can never win. Same answer, roughly linear time when changes are frequent.
- Small β → over-detection; large β → under-detection. Check the whole penalty path; trust plateaus.
min_sizestops spike cuts;jump=5(ruptures default) rounds dates. - L2 on a trending series gives a staircase: for trend changes use differences, a linear cost, or detrend first, and know which one your pipeline uses.
- P0: PELT's choices enter the Bayesian model as fixed inputs, so selection uncertainty is not propagated: intervals near uncertain changepoints are too narrow and selected changes look too big. And PELT must run inside each backtest fold, or the test period leaks.
Cheat sheet
| Idea | Formula / setting | Remember |
|---|---|---|
| Segment, changepoint | $y_{\tau_j:\tau_{j+1}}$, $\tau_j$ = first day of a new segment | ruptures returns segment ends, last = n |
| L2 cost | $\sum (y_t - \bar y_{seg})^2$ | level shifts only; one noise level |
| Normal cost | $L\log\hat\sigma^2_{seg}$ | level + spread; needs longer segments |
| Objective | $\sum_{j=0}^{m} C(y_{\tau_j:\tau_{j+1}}) + \beta m$ | a cut must save more than β |
| BIC-like β | L2: $2\hat\sigma^2\log n$; Normal: $3\log n$; linear: $3\hat\sigma^2\log n$ | rules of thumb; β scales with $y^2$ |
| Noise estimate | $\hat\sigma = 1.4826\,\text{MAD}(\Delta y)/\sqrt2$ | robust to level shifts |
| Optimal partitioning | $F(t) = \min_{s}[F(s) + C(y_{s:t}) + \beta]$, $F(0) = -\beta$ | exact, $O(n^2)$ |
| PELT pruning | drop $s$ if $F(s) + C(y_{s:t}) \gt F(t)$ | safe because splitting never raises the cost |
| Work | ~$O(n)$ with frequent changes; $O(n^2)$ with none | same answer as OP |
| Settings | rpt.Pelt(model="l2", min_size=2, jump=1) | default jump=5 rounds dates |
| Slope changes | L2 on $\Delta y$ (noise $2\sigma^2$) or model="linear" on [y, t, 1] | L2 on levels → staircase |
| Selection uncertainty | $p(\theta\mid y) = \sum_\tau p(\theta\mid y,\tau)p(\tau\mid y)$ vs $p(\theta\mid y,\hat\tau)$ | too narrow; winner's curse |
| Leakage | PELT, β, σ̂, scaling inside every fold | recent changes look easy only in leaky backtests |
import warnings
import numpy as np
import ruptures as rpt
warnings.filterwarnings("ignore", category=UserWarning) # ruptures' "normal" cost prints a notice
def sigma_hat(y):
"""Robust noise sd from day-to-day differences: 1.4826 * MAD(diff) / sqrt(2)."""
d = np.diff(y)
return 1.4826 * np.median(np.abs(d - np.median(d))) / np.sqrt(2)
# 1) A series with three level changes; BIC-like penalty; PELT with ruptures
rng = np.random.default_rng(5)
mu = np.repeat([10.0, 14.0, 11.0, 16.0], [52, 39, 57, 52]) # true changes at 52, 91, 148
y = mu + 1.5 * rng.standard_normal(mu.size)
n = y.size
s = sigma_hat(y)
pen = 2 * s**2 * np.log(n) # beta = 2 sigma^2 log n (L2 cost units)
print(round(s, 2), round(pen, 1)) # 1.41 21.1
algo = rpt.Pelt(model="l2", min_size=2, jump=1).fit(y)
print(algo.predict(pen=pen)) # [52, 92, 148, 200] segment ENDS; the last one is n
print(rpt.Pelt(model="l2").fit(y).predict(pen=pen)) # [50, 90, 145, 150, 200] default jump=5: multiples of 5 only
# 2) Penalty sensitivity: number of changepoints for a grid of penalties
for b in [1, 5, 10, 25, 100, 300, 1000]:
print(b, len(algo.predict(pen=b)) - 1)
# 1 60 | 5 18 | 10 7 | 25 3 | 100 3 | 300 1 | 1000 0 (over-detection ... plateau at 3 ... under-detection)
# 3) Optimal partitioning vs PELT, written out (L2 cost, segments of length >= 1), counting the work
def segment_dp(y, beta, prune=True):
n = len(y); c1 = np.r_[0, np.cumsum(y)]; c2 = np.r_[0, np.cumsum(y * y)]
cost = lambda a, b: c2[b] - c2[a] - (c1[b] - c1[a]) ** 2 / (b - a)
F = np.full(n + 1, np.inf); F[0] = -beta; last = np.zeros(n + 1, int); R = [0]; evals = 0
for t in range(1, n + 1):
vals = {s_: F[s_] + cost(s_, t) + beta for s_ in R}; evals += len(vals)
last[t] = min(vals, key=vals.get); F[t] = vals[last[t]]
if prune:
R = [s_ for s_ in R if vals[s_] - beta <= F[t]] # drop starts that trail by more than one beta
R.append(t)
cps, t = [], n
while t > 0:
t = last[t]
if t > 0: cps.append(int(t))
return sorted(cps), evals
print(segment_dp(np.array([1.0, 3, 10, 12]), 5.0, prune=False)) # ([2], 10)
print(segment_dp(np.array([1.0, 3, 10, 12]), 5.0, prune=True)) # ([2], 8)
op, pe = segment_dp(y, pen, prune=False), segment_dp(y, pen, prune=True)
print(op[0] == pe[0], op[1], pe[1]) # True 20100 5607 same answer, about 28% of the work
# 4) Mean-only vs mean-and-variance cost on a change in spread
r = np.random.default_rng(1)
z = np.r_[10 + r.standard_normal(40), 13 + r.standard_normal(40), 13 + 3 * r.standard_normal(40)]
sz = sigma_hat(z)
print(rpt.Pelt(model="l2", min_size=2, jump=1).fit(z).predict(pen=2 * sz**2 * np.log(120))[:-1])
print(rpt.Pelt(model="normal", min_size=5, jump=1).fit(z).predict(pen=3 * np.log(120))[:-1])
# [40, 82, 84, 95, 101, 106, 108] L2: extra cuts inside the noisy part
# [40, 80] normal cost: the level change and the variance change
# 5) Slope changes: L2 on levels makes a staircase; a linear (regression) cost finds the bends
t = np.arange(150.0)
g = 100 + 0.2 * t + 0.6 * np.clip(t - 50, 0, None) - 0.9 * np.clip(t - 100, 0, None)
yt = g + 2.0 * np.random.default_rng(0).standard_normal(150)
st = sigma_hat(yt)
print(len(rpt.Pelt(model="l2", min_size=2, jump=1).fit(yt).predict(pen=2 * st**2 * np.log(150))) - 1) # 11: a staircase
X = np.column_stack([yt, t, np.ones_like(t)]) # model="linear": first column y, then regressors
print(rpt.Pelt(model="linear", min_size=5, jump=1).fit(X).predict(pen=3 * st**2 * np.log(150))[:-1]) # [51, 102]
# 6) Selection uncertainty: same truth, 300 noise replicates
truth = np.repeat([10.0, 13.0, 14.0], 50) # big change at 50, small change (0.67 sd) at 100
found_small, size_small, first_big, n_cps = [], [], [], []
for rep in range(300):
yr = truth + 1.5 * np.random.default_rng(rep).standard_normal(150)
sr = sigma_hat(yr)
b = rpt.Pelt(model="l2", min_size=2, jump=1).fit(yr).predict(pen=2 * sr**2 * np.log(150))
n_cps.append(len(b) - 1); e = [0] + b
near_big = [c for c in b[:-1] if abs(c - 50) <= 15]
if near_big: first_big.append(near_big[0])
for i in range(1, len(e) - 1):
if abs(e[i] - 100) <= 15:
found_small.append(e[i]); size_small.append(yr[e[i]:e[i + 1]].mean() - yr[e[i - 1]:e[i]].mean())
first_big = np.array(first_big)
print(len(first_big) / 300, np.percentile(first_big, [5, 95]), np.mean(np.abs(first_big - 50) <= 2)) # 1.0 [48. 53.] 0.883
print(round(len(found_small) / 300, 2), np.percentile(found_small, [5, 95]), round(np.mean(size_small), 2)) # 0.67 [ 88. 109.] 1.21 (true size 1.0)
print(np.bincount(n_cps)) # [ 0 77 197 22 2 1 1] replicates with 0, 1, 2, ... changepoints
# 7) Leakage: PELT on the full series vs on the training part only
T, Hh = 100, 30
mae_honest, mae_leaky = [], []
for rep in range(400):
yl = np.r_[np.full(97, 10.0), np.full(T + Hh - 97, 13.0)] + 1.5 * np.random.default_rng(rep).standard_normal(T + Hh)
train = yl[:T]
c_h = rpt.Pelt(model="l2", min_size=2, jump=1).fit(train).predict(pen=2 * sigma_hat(train)**2 * np.log(T))[:-1]
c_l = [c for c in rpt.Pelt(model="l2", min_size=2, jump=1).fit(yl).predict(pen=2 * sigma_hat(yl)**2 * np.log(T + Hh))[:-1] if c < T]
for cps, out in [(c_h, mae_honest), (c_l, mae_leaky)]:
start = cps[-1] if cps else 0
out.append(np.mean(np.abs(yl[T:] - train[start:].mean()))) # forecast = mean of the last training segment
print(round(np.mean(mae_honest), 2), round(np.mean(mae_leaky), 2)) # 1.86 1.41 the leaky backtest looks better than reality
1. For a series of 6 values, ruptures returns [3, 6]. How many changepoints were found?
2. Why does the changepoint objective need the penalty $\beta m$?
3. At time $t$, PELT removes a candidate start $s$ when…
4. You switch your series from units to tenths (multiply y by 10) and keep the same β for an L2 PELT. What happens?
5. Changepoints are chosen by PELT on the observed series and then used by a Bayesian forecasting model. Which statement is correct?
6. A colleague runs PELT once on the full history, then does a rolling-origin backtest with those changepoints. What is the most likely effect?
Practice problems
A. Daily orders 4, 6, 5, 11, 9, 10 with noise sd 1. Using the L2 cost and $\beta = 2\sigma^2\log n$, how many changepoints, and where?
- $\beta = 2 \times 1 \times \log 6 = 3.58$.
- No cut: mean 7.5; squares $12.25, 2.25, 6.25, 12.25, 2.25, 6.25$; cost 41.5.
- Cut at day 3: $(4, 6, 5)$ mean 5, cost $1 + 1 + 0 = 2$; $(11, 9, 10)$ mean 10, cost 2. Total $4 + 3.58 = 7.58 \lt 41.5$.
- A second cut: the best is $(4), (6, 5), (11, 9, 10)$ or $(4, 6, 5), (11), (9, 10)$, cost 2.5: saves 1.5, less than 3.58. Not worth it.
- Answer: one changepoint at $\tau = 3$ (ruptures would return
[3, 6]).
B. Run optimal partitioning by hand on $y = 2, 2, 8$ with $\beta = 3$ and segments of length ≥ 1. Which starts would PELT prune at $t = 3$?
- Costs: single points 0; $C(2,2) = 0$; $C(2,8) = 18$; $C(2,2,8) = 4 + 4 + 16 = 24$ (mean 4).
- $F(0) = -3$. $F(1) = -3 + 0 + 3 = 0$. $F(2) = \min(-3 + 0 + 3,\ 0 + 0 + 3) = 0$ (start 0).
- $F(3) = \min(-3 + 24 + 3,\ 0 + 18 + 3,\ 0 + 0 + 3) = \min(24, 21, 3) = 3$ (start 2): one changepoint at 2, total $0 + 0 + 3$.
- Pruning at $t = 3$: $s = 0$: $F(0) + C = -3 + 24 = 21 \gt 3$, prune. $s = 1$: $0 + 18 = 18 \gt 3$, prune. $s = 2$: $0 + 0 = 0 \le 3$, keep.
C. Your pipeline standardises the series (sd 1) before PELT; the noise sd in standard units is about 0.3 and n = 730. What BIC-like β fits, and what goes wrong if someone hard-codes $\beta = 2\log n$?
$\beta = 2 \times 0.3^2 \times \log 730 = 0.18 \times 6.59 \approx 1.19$. Hard-coding $2\log n \approx 13.2$ silently assumes noise sd 1, so it is about 11 times too large ($1/0.09$): PELT will under-detect and only the largest changes survive. β must be tied to the noise level of the series PELT actually sees.
D. (Interview) "PELT is exact and linear time. Isn't that too good to be true?"
"It is exact because it solves the same dynamic program as optimal partitioning and only discards starts that provably cannot win: if a start trails the current best by more than one penalty, then adding a cut now always beats it later, since splitting a segment never increases an L2 or Normal cost. The linear time is not guaranteed: it holds in expectation when the number of changepoints grows with the length of the series. With no changes, nothing gets pruned and it is quadratic. And 'exact' only means the exact minimiser of the chosen cost plus penalty, not the true changepoints."
E. (Interview, your forecasting model) "Your 80% forecast intervals cover only 65% in the weeks after a trend change. Could the changepoint step be involved?"
"Yes, in three ways. First, selection uncertainty: the PELT locations are fixed inputs, so the posterior ignores uncertainty about where the last change was, and the forecast extends a slope that is less certain than the model thinks. Second, detection delay: right after a real change there are few points after it, so PELT may miss it and the Laplace prior will shrink the new slope change, so the forecast reacts late. Third, I would check that PELT ran inside each backtest fold; if not, the backtest itself is optimistic. Remedies: denser candidates with Laplace shrinkage, sensitivity refits with alternative changepoint sets, and calibrating the intervals on rolling-origin coverage."
F. With default settings, rpt.Pelt(model="l2").fit(y).predict(pen=pen) returns [50, 90, 145, 150, 200], while a plot shows only three clear jumps. What do you suspect, and how do you check?
The default jump=5 only allows changepoints at multiples of 5. A real change between them (here at 148) cannot be placed, so PELT approximates it with two cuts (145 and 150) and a short segment. Rerun with jump=1: in the Code-it data it returns [52, 92, 148, 200], three changepoints at the right days.
Laplace priors on trend changes
Your trend may bend at many candidate changepoints, but you believe it really bends at only a few. The Laplace prior $\delta_j \sim Laplace(0, b)$ is how the model says this: most slope changes are pulled to almost nothing, a few large ones survive. This chapter shows how that pull works, why the MAP fit is a lasso with exact zeros while the full posterior is not, what Prophet's famous 0.05 really measures, and how the one number b moves the model from a stiff straight line to a trend that chases noise.
- Explain why a trend model with many candidate changepoints needs a shrinkage prior, and what "most δⱼ ≈ 0, a few large" means
- Compare Laplace and Normal priors: the sharp peak, the heavier tails, and the constant "pull" toward zero
- See how much evidence a real change needs before the prior lets it through, and why forecasts react late
- Derive the L1 relationship: the MAP is a lasso, solved by soft-thresholding (proximal gradient), with $\lambda = \sigma^2/b$
- Know that the full posterior is shrunk but not exactly sparse, and how to report changes honestly
- Read the scale b in real units, including Prophet's
changepoint_prior_scale = 0.05on its scaled data - Diagnose underfitting (b too small) and overfitting (b too large), and choose b by backtest and sensitivity checks
What we need from earlier chapters: the piecewise-linear trend with changepoints $s_j$ and slope changes $\delta_j$ (Chapter 7.8) and where candidates come from (grid + PELT, Chapter 7.9); the Laplace distribution (Chapter 4.9); MAP, lasso and "the Laplace posterior is not sparse" (Chapter 5.3); shrinkage priors and prior sensitivity (Chapters 6.2 and 6.8). Words used throughout: shrinkage = pulling estimates toward a reference value (here 0); sparse = most entries exactly zero; MAP = the single most probable parameter value (the peak of the posterior); the scale b of a Laplace distribution = its typical size: $E|\delta| = b$.
Many candidate changepoints, few real changes core
We do not know where the trend bends, so we offer it many places where it may bend: 25 candidates on a grid, plus any that PELT proposes. At each candidate $s_j$ the slope may change by $\delta_j$. Our real belief is "at most candidates nothing happened, at a few something did".
If every $\delta_j$ is a free number, the fit does not share that belief. It uses all 25 bends to follow the noise: the trend wiggles, neighbouring slope changes come out huge with opposite signs, and the last slope (the one the forecast extends) is unstable. We need to tell the model "keep most $\delta_j$ near 0 unless the data insist". That message is a prior: $\delta_j \sim Laplace(0, b)$.
Three ways to say it:
- Picture: 25 hinges on a ruler; without a rule, every hinge flaps; with the Laplace prior, most hinges are stiff and only a few bend.
- Numbers: with free slope changes, the 25 fitted δⱼ on one series range from −4.8 to +5.7; with the Laplace prior (b = 0.05), 21 of them are exactly 0 at the MAP.
- Slogan: offer many bends, believe in few.
100 training days and 30 test days, time scaled so the history is $[0, 1]$, y divided by its largest value (as Prophet does). The true trend bends twice. 25 candidates. Averages over 20 simulated series (the widget below shows one at a time); the noise sd is about 0.045 in these units.
- Count the unknowns: $m$, $k$ and 25 slope changes: 27 numbers from 100 points.
- No prior (least squares): training RMSE 0.037, below the noise level 0.045. The fit is explaining noise.
- Its test RMSE (the next 30 days) is 0.057: the wiggles and the unstable last slope cost accuracy.
- With $\delta_j \sim Laplace(0, 0.05)$ (MAP fit): training RMSE 0.047 (about the noise level, as it should be) and test RMSE 0.050, better than without the prior.
With candidates $s_1 \lt \dots \lt s_J$, the piecewise-linear trend of Chapter 7.8 can be written as a sum of "hinge" columns:
$$g(t) = \Big(k + \sum_{j:\,s_j \le t}\delta_j\Big)t + \Big(m - \sum_{j:\,s_j \le t} s_j\delta_j\Big) = m + k\,t + \sum_{j=1}^{J} \delta_j\,(t - s_j)_+,$$where $(t - s_j)_+ = \max(0, t - s_j)$ is 0 before the candidate and grows by 1 per unit of time after it. The prior on each slope change is
$$\delta_j \overset{iid}{\sim} Laplace(0, b), \qquad p(\delta_j) = \frac{1}{2b}\,e^{-|\delta_j|/b}.$$- $b \gt 0$ is the scale: $E|\delta_j| = b$, $Var(\delta_j) = 2b^2$, sd $= \sqrt2\,b$.
- In code: NumPyro
dist.Laplace(0.0, b)(loc, scale = b); SciPylaplace(loc=0, scale=b); Standouble_exponential(0, b)(what Prophet uses). - "Most δⱼ ≈ 0, a few large" is a statement about the prior; whether the fitted changes are exactly 0 depends on how you summarise the posterior (concepts 4 and 5).
Why do we need it?
Without shrinkage, offering many candidates means the trend bends at all of them and chases noise. With it, we can offer many candidates cheaply (we do not need to know where the changes are) and let the data switch on only the ones they support.
Where is it used?
Prophet's trend (changepoint_prior_scale is b), Prophet-style NumPyro models such as your forecasting model, the Bayesian lasso, trend filtering with L1 penalties, and sparse regression in general.
How is it used?
Build one hinge column per candidate, give every δⱼ the same Laplace(0, b) prior on scaled data, fit (MAP or posterior), then look at the fitted trend and slope over time rather than at single δⱼ values.
"More candidates make the trend more flexible, so more is always better."
Without shrinkage, more candidates mean more noise-chasing. With shrinkage, extra candidates are cheap, but the overall flexibility depends on both the number of candidates and b.
"Just delete the candidates where nothing happened."
We do not know which those are; that is the whole problem. The prior lets the data decide, softly, without a hard yes/no choice (compare the hard choice of PELT in Chapter 7.9).
In your forecasting model the candidates come from a Prophet-like grid plus PELT detection, and every slope change gets $\delta_j \sim Laplace(0, b)$. This is the step that turns "many places where the trend may bend" into "a few places where it does". It is also your answer to "won't so many changepoints overfit?": the Laplace prior keeps most adjustments small, giving a sparse-ish changepoint representation.
$g(t) = m + kt + \sum_j \delta_j (t - s_j)_+$, with $\delta_j \sim Laplace(0, b)$: $E|\delta_j| = b$, sd $= \sqrt2 b$.
Many candidates + shrinkage = flexible where the data insist, stiff elsewhere.
Trap: free δⱼ (no prior) chase noise; training error below the noise level is the tell-tale sign.
Quick check: what is the slope of $g(t)$ after the third candidate, if $k = 0.3$, $\delta_1 = 0$, $\delta_2 = 0.5$, $\delta_3 = -0.2$?
$k + \delta_1 + \delta_2 + \delta_3 = 0.3 + 0 + 0.5 - 0.2 = 0.6$. The slope only depends on the sum of the changes so far; the offsets $-s_j\delta_j$ keep the line connected.
Laplace versus Normal: a sharp peak, heavier tails, a constant pull
Both priors are centred at 0, but they have different shapes. The Normal is a round hill: it expects medium changes everywhere. The Laplace is a tent: a sharp tip at 0 and long straight sides. It expects most changes tiny and a few big.
The clearest way to see the difference is the pull: how hard the prior drags an estimate back toward 0. A prior acts like a penalty $-\log p(\delta)$ added to the fit, and the pull is that penalty's slope. The Normal's pull grows with the size of δ: weak for small changes, very strong for big ones. The Laplace's pull is the same for every δ: relatively strong for small changes (it can drag them all the way to 0) and relatively weak for big ones (they survive).
Three ways to say it:
- Picture: Normal = a spring (the further you stretch it, the harder it pulls); Laplace = a constant rope tension.
- Numbers: with b = 0.05 and a Normal of the same variance, the pull at δ = 0.01 is 20 (Laplace) vs 2 (Normal), but at δ = 0.2 it is 20 vs 40.
- Slogan: Laplace shrinks small changes hard and big changes gently.
Laplace(0, b = 0.05) versus a Normal with the same variance $2b^2 = 0.005$ (sd $\sqrt{0.005} = 0.0707$).
- Tiny changes, $|\delta| \lt b/2 = 0.025$: Laplace $1 - e^{-0.5} = 0.393$; Normal $2\Phi(0.025/0.0707) - 1 = 2\Phi(0.354) - 1 = 0.276$. The Laplace has more mass near 0.
- Large changes, $|\delta| \gt 3b = 0.15$: Laplace $e^{-3} = 0.050$; Normal $2(1 - \Phi(2.12)) = 0.034$. The Laplace also has more mass far out.
- Penalties: Laplace $-\log p = |\delta|/b + c = 20|\delta| + c$; Normal $-\log p = \delta^2/(2 \cdot 0.005) + c = 100\,\delta^2 + c$.
- Pulls (slopes of the penalties): Laplace $1/b = 20$ for every $\delta \ne 0$; Normal $\delta/0.005 = 200\,\delta$: 2 at $\delta = 0.01$, 40 at $\delta = 0.2$.
For $\delta \sim Laplace(0, b)$ and $\delta \sim N(0, \tau^2)$:
$$-\log p_{Lap}(\delta) = \frac{|\delta|}{b} + \log(2b), \qquad -\log p_{N}(\delta) = \frac{\delta^2}{2\tau^2} + \tfrac12\log(2\pi\tau^2).$$- Pull (derivative of $-\log p$): Laplace $\text{sign}(\delta)/b$ (constant size, undefined exactly at 0: the "corner"); Normal $\delta/\tau^2$ (proportional to δ).
- Same variance means $\tau^2 = 2b^2$. Then the Laplace has more mass both near 0 and in the tails; the Normal has more in between.
- Tails: $P(|\delta| \gt x) = e^{-x/b}$ (exponential) for Laplace versus $\approx e^{-x^2/2\tau^2}$ for Normal: big changes are much more plausible under Laplace.
- The corner at 0 is what produces exact zeros at the MAP (concept 4). It does not put any probability on exactly 0 (concept 5).
Why do we need it?
Trend changes are rare and sometimes large. A Normal prior either crushes the large real changes (if tight) or lets every candidate wobble (if wide). The Laplace's shape fits "rare but possibly large" with one number.
Where is it used?
Changepoint slope changes in Prophet and Prophet-style models; the lasso (L1) penalty in regression; robust (absolute-error) regression uses the same shape as a likelihood (Chapter 4.9).
How is it used?
Choose Laplace when you expect sparse changes, Normal when you expect many small ones (as for Fourier coefficients). Compare them by their penalties and pulls, or by prior predictive draws of the trend (Chapter 6.2).
"The Laplace prior is just a narrower Normal."
With the same variance it is both more peaked and heavier-tailed. No Normal can say "mostly tiny, sometimes big" the way a Laplace does.
"b is the standard deviation of δⱼ."
b is the scale: $E|\delta_j| = b$, and the sd is $\sqrt2\,b$. If you match a Normal to a Laplace, use $\tau = \sqrt2\,b$.
Penalties: Laplace $|\delta|/b$ (constant pull $1/b$), Normal $\delta^2/2\tau^2$ (pull $\delta/\tau^2$).
Equal variance ($\tau = \sqrt2 b$): Laplace has more mass near 0 and in the tails; pulls harder below $|\delta| = 2b$, softer above.
Trap: b is not the sd ($\text{sd} = \sqrt2 b$).
Quick check: with b = 0.05, what fraction of prior slope changes exceed 0.1 in size?
$P(|\delta| \gt 0.1) = e^{-0.1/0.05} = e^{-2} \approx 0.135$: about 13.5%, so about 3 of 25 candidates in a typical prior draw.
Shrinkage: how much evidence a real change needs core
Think of a tug-of-war over one slope change $\delta_j$. The prior pulls toward 0 with constant force $1/b$. The data pull toward the value they suggest, and their strength grows with the evidence: the number of days after the candidate, and how far those days drift from the old slope. A slope change shows up slowly at first (one day after a bend the line has barely moved) and then quickly, because the gap between old and new line grows every day.
So right after a real trend change, the prior wins: the model keeps the old slope, and its forecast reacts late. After a couple of weeks the data win, and the estimate moves close to the truth. With the Laplace prior, once the evidence is strong the change is let through almost fully; small or uncertain changes stay near 0.
Three ways to say it:
- Picture: a gate that opens only when enough data push on it; small pushes leave it shut.
- Numbers: a true slope change of 0.5 orders/day per day (noise sd 5) is held at exactly 0 by the MAP for 10 days, reaches 0.30 after 15 days and 0.47 after 30.
- Slogan: shrinkage buys stability with delay.
One candidate. Daily noise sd $\sigma = 5$ orders/day; true slope change $\delta^* = 0.5$ orders/day per day; we assume the old slope and level are known from a long history, and that the data's estimate $x$ happens to equal the truth (0.5). Prior $Laplace(0, b = 0.1)$; the Normal for comparison has the same variance ($\tau^2 = 2b^2 = 0.02$).
- After $n$ days the hinge column is $1, 2, \dots, n$, so the standard error of $x$ is $se = \sigma/\sqrt{1^2 + \dots + n^2}$. For $n = 5$: $\sqrt{55} = 7.42$, $se = 0.674$. For $n = 10$: $\sqrt{385} = 19.6$, $se = 0.255$. For $n = 15$: $\sqrt{1240} = 35.2$, $se = 0.142$. For $n = 30$: $\sqrt{9455} = 97.2$, $se = 0.051$.
- Laplace MAP $= \text{sign}(x)\max(0, |x| - se^2/b)$. Thresholds $se^2/b$: $4.55$ ($n = 5$), $0.65$ ($n = 10$), $0.20$ ($n = 15$), $0.026$ ($n = 30$). MAP: $0,\ 0,\ 0.30,\ 0.47$.
- Laplace posterior mean (numerical): $0.02,\ 0.12,\ 0.30,\ 0.47$. Small but not 0 even when the MAP is 0.
- Normal posterior mean $x\tau^2/(\tau^2 + se^2)$: $0.02,\ 0.12,\ 0.25,\ 0.44$. Early on both priors hold the change back about equally; once the evidence is strong, the Laplace lets it through more fully (0.47 vs 0.44 of the true 0.5).
For one slope change with the rest of the model known, the data give an estimate $x \sim N(\delta, se^2)$ with $se^2 = \sigma^2/\sum_t (t - s_j)_+^2$. With prior $Laplace(0, b)$:
$$p(\delta\mid x) \propto \exp\Big(-\frac{(x - \delta)^2}{2\,se^2} - \frac{|\delta|}{b}\Big), \qquad \hat\delta_{MAP} = \text{sign}(x)\,\max\big(0,\ |x| - se^2/b\big).$$- Shrinkage = the gap between the estimate the data alone would give ($x$) and the Bayesian estimate. For large $|x|$ the Laplace MAP shrinks by the constant amount $se^2/b$; a Normal prior shrinks by the constant fraction $se^2/(\tau^2 + se^2)$.
- Evidence for a slope change grows very fast after the candidate: $\sum_{i=1}^n i^2 = n(n+1)(2n+1)/6 \approx n^3/3$.
- A candidate near the end of the history has few days after it, so its δⱼ is always heavily shrunk. Prophet places candidates only in the first 80% of the history partly for this reason (Chapter 7.8).
Why do we need it?
It explains two behaviours you will see and be asked about: fitted slope changes are smaller than the truth (shrinkage bias), and after a real change the forecast keeps the old slope for a while (detection delay).
Where is it used?
Every shrinkage prior: changepoint slopes, holiday effects with few occurrences (Chapter 7.12), partial pooling of segments in the A/B framework (Chapter 6.6), ridge and lasso.
How is it used?
Ask "how many days after a change until the model believes it?" for your b and noise level; if the delay is too long for the business, raise b (more flexible) or add an explicit candidate where you know a change happened (a launch, a price change).
"The fitted slope change is the size of the real change."
It is shrunk toward 0, by roughly $se^2/b$ when the evidence is strong and by much more when it is weak. Report it as a shrunk estimate, ideally with its posterior interval.
"The model missed the new trend because it is wrong."
A few days after a real change, holding the old slope is the designed behaviour: the prior demands evidence before believing a bend. The delay is the price of not reacting to every noisy week.
In your forecasting model this explains a pattern you may have seen in backtests: after a real shift in growth, forecasts lag for a while and then catch up. The Laplace prior on δⱼ (and the end-of-history effect: candidates near the end have little data after them) causes it. If an interviewer asks how to react faster, the honest options are a larger b (more flexible, more noise-chasing), explicit candidates at known events, or a model with a local level that adapts quickly; each trades stability for speed.
One change: $x \sim N(\delta, se^2)$, $se^2 = \sigma^2/\sum(t - s_j)_+^2$; Laplace MAP $= \text{sign}(x)\max(0, |x| - se^2/b)$.
Evidence for a slope change grows like $n^3/3$; until it beats $se^2/b$ the MAP stays at 0 (forecasts react late).
Trap: reading the shrunk δⱼ as the true size of the change.
Quick check: σ = 5, b = 0.1, data estimate x = 0.5 after 20 days. Is the MAP still 0?
$\sum_{i=1}^{20} i^2 = 2870$, $se^2 = 25/2870 = 0.0087$, threshold $se^2/b = 0.087$. MAP $= 0.5 - 0.087 = 0.413$: the gate is open.
The L1 relationship: the MAP is a lasso, solved by soft-thresholding core
The MAP is the single most probable set of slope changes: the peak of the posterior. To find a peak we minimise "minus log posterior", which is "squared misfit + penalty". With a Laplace prior the penalty is $\sum_j |\delta_j|/b$: an L1 penalty. So the MAP fit is exactly a lasso regression on the hinge columns (Chapter 5.3).
The absolute value has a corner at 0. At the corner, the penalty pulls with force $1/b$ from both sides, so a small δⱼ whose data pull is weaker than that sits exactly at 0. A computer finds this with soft-thresholding: move each δⱼ as the data suggest, then shrink it toward 0 by a fixed amount, and set it to 0 if it would cross. Repeating "gradient step, then soft-threshold" is the proximal gradient method (ISTA; its accelerated version is FISTA).
Three ways to say it:
- Picture: walk downhill on the squared error, then let a magnet at 0 grab every δⱼ that is close enough.
- Numbers: data estimate 0.3, threshold 0.2 → MAP 0.1; data estimate 0.15 → MAP exactly 0.
- Slogan: Laplace prior + MAP = lasso; the corner makes the zeros.
- One slope change with $se = 0.1$ and $b = 0.05$: the threshold is $se^2/b = 0.01/0.05 = 0.2$.
- Data estimate $x = 0.3$: MAP $= \text{sign}(0.3)\max(0, 0.3 - 0.2) = 0.1$. Data estimate $x = 0.15$: $\max(0, 0.15 - 0.2) = 0$, exactly.
- Many slope changes, noise sd $\sigma = 0.04$ (scaled units), $b = 0.05$. Minus log posterior $= \frac{1}{2\sigma^2}\text{RSS} + \frac1b\sum|\delta_j|$. Multiply by $\sigma^2$: $\tfrac12\text{RSS} + \lambda\sum|\delta_j|$ with $\lambda = \sigma^2/b = 0.0016/0.05 = 0.032$.
- scikit-learn's
Lassominimises $\frac{1}{2n}\text{RSS} + \alpha\sum|w_j|$, so $\alpha = \lambda/n = 0.032/100 = 0.00032$ for 100 training days. - On the widgets' series (noise sd 0.044 after scaling, same b) the MAP sets 21 of the 25 slope changes exactly to 0.
With Normal noise of sd σ, flat priors on $m, k$, and $\delta_j \sim Laplace(0, b)$, the MAP solves
$$\hat\delta_{MAP} = \arg\min_{m,k,\delta}\ \frac{1}{2\sigma^2}\sum_{i}\Big(y_i - m - k t_i - \sum_j \delta_j (t_i - s_j)_+\Big)^2 + \frac1b\sum_j|\delta_j| ,$$a lasso with $\lambda = \sigma^2/b$ in the "½RSS + λ·L1" convention (only δ is penalised). Soft-thresholding is $S_\theta(u) = \text{sign}(u)\max(0, |u| - \theta)$. Proximal gradient (ISTA) repeats, with step size $\eta = 1/L$ ($L$ = largest eigenvalue of $X^\top X/\sigma^2$):
$$\delta \leftarrow S_{\eta/b}\Big(\delta + \frac{\eta}{\sigma^2} X^\top (y - X\delta)\Big).$$- FISTA adds a momentum step and converges much faster; both converge to the exact lasso solution.
- Exact zeros need a method that handles the corner (soft-thresholding, coordinate descent, LARS). Plain gradient methods on the non-smooth objective (Adam in an SVI fit with an
AutoDeltaguide, quasi-Newton optimisers) typically end up near 0, not exactly at 0 (the Code-it shows this). - Prophet's default fit (
mcmc_samples=0) is a MAP of this kind, computed with Stan's optimiser.
Why do we need it?
It connects the Bayesian prior to the familiar L1/lasso picture and explains the "exact zeros" people see in MAP fits. It also gives you a fast, exact way to fit the trend for exploration and for choosing b.
Where is it used?
Prophet's default fit, lasso regression (sklearn.linear_model.Lasso), trend filtering, compressed sensing, and any MAP fit of a model with Laplace priors.
How is it used?
To get a quick MAP trend: build the hinge columns, project out $m$ and $k$, run a lasso with $\alpha = \sigma^2/(b\,n)$ (or FISTA), then read which candidates are non-zero. Remember it is one point, not a posterior.
"My SVI fit with an AutoDelta guide is the MAP, so it has exact zeros."
It approximates the MAP with gradient steps, which jitter around the corner at 0: values come out tiny (like $10^{-6}$), not exactly 0. Exact zeros need soft-thresholding or a lasso solver.
"λ in the lasso and b in the prior are the same knob, so the numbers carry over."
They are linked by $\lambda = \sigma^2/b$ (and sklearn's α = λ/n). A bigger b means a smaller λ (weaker penalty), and the link involves the noise level σ.
MAP with $\delta_j \sim Laplace(0,b)$ = lasso: $\tfrac12\text{RSS} + \lambda\sum|\delta_j|$, $\lambda = \sigma^2/b$ (sklearn α = λ/n).
Soft-threshold $S_\theta(u) = \text{sign}(u)\max(0, |u| - \theta)$; ISTA: gradient step then $S_{\eta/b}$.
Trap: exact zeros only from corner-aware solvers, and only at the MAP.
Quick check: σ = 0.05 (scaled), b = 0.1, 200 training days. What λ and what sklearn α correspond to the MAP?
$\lambda = \sigma^2/b = 0.0025/0.1 = 0.025$; $\alpha = \lambda/n = 0.025/200 = 0.000125$.
The full posterior: shrunk, but not exactly sparse core
The MAP is the top of the mountain; the posterior is the whole mountain. Your model is fitted with SVI (or could be with NUTS), which describes the whole mountain. And on the whole mountain, "δⱼ is exactly 0" has probability zero: δⱼ is a continuous quantity, and a single point has no area under a density. Posterior draws are never exactly 0; posterior means and medians of "switched-off" candidates are small but not 0.
There is a second surprise. When the data say "the trend bent somewhere around here" but cannot say exactly where, the posterior spreads the change over several neighbouring candidates, each with a modest δⱼ whose interval includes 0. Look at a single δⱼ and you might say "nothing happened"; look at the slope over that stretch and the change is obvious.
Three ways to say it:
- Picture: the MAP is a photo of the tallest point; the posterior is a fog that is thick near 0 but never collapses onto it.
- Numbers: on the widgets' series the MAP has 21 of 25 slope changes exactly 0; 75 000 posterior draws contain none, and 24 of 25 individual 90% intervals include 0, even though two real changes are there.
- Slogan: sparse-ish, not sparse; read slopes, not single deltas.
The widgets' series (seed 3), $b = 0.05$, noise sd known. The posterior was sampled with a Gibbs sampler (3 000 draws; it agrees with NumPyro's NUTS to within about 0.02 on every posterior mean).
- MAP: 4 non-zero slope changes: $\delta_7 = 0.30$, $\delta_8 = 0.27$ (near the true change of +1.0 at $t = 0.25$) and $\delta_{17} = -0.78$, $\delta_{19} = -0.52$ (near the true −1.55 at $t = 0.55$). The other 21 are exactly 0.
- Posterior draws: 0 of the $3000 \times 25 = 75\,000$ values are exactly 0. The smallest posterior mean in size is 0.009, not 0. About 6.5% of draws have $|\delta_j| \lt 0.005$: concentrated near 0, not on it.
- The +1.0 change is shared: posterior means 0.12, 0.15, 0.14, 0.10 at four neighbouring candidates. Each 90% interval includes 0 (24 of the 25 do).
- The slope after the last candidate: MAP −0.19; posterior 90% interval $[-0.43, -0.19]$; truth −0.22. The slope is well determined even when single δⱼ are not.
For a continuous posterior $p(\delta_j\mid y)$, $P(\delta_j = 0\mid y) = 0$. Shrinkage shows up as posterior mass concentrated near 0, not as a point mass at 0.
- MAP (mode): can be exactly 0 (the corner). Posterior mean and median: shrunk toward 0, almost never exactly 0. Draws (SVI guide or NUTS): never exactly 0.
- Hence the precise statement (Chapter 5.3): the Laplace prior is equivalent to L1 regularisation at the MAP; the full posterior is shrunk but not sparse.
- Truly sparse Bayesian models put a point mass at 0 (spike-and-slab priors) or use stronger shrinkage (horseshoe priors); a Laplace prior is neither.
- Honest summaries: the posterior of the slope $k + \sum_{j: s_j \le t}\delta_j$ over time, the cumulative change over a window, or $P(|\text{slope change over a window}| \gt c\mid y)$ for a meaningful $c$.
Why do we need it?
"Which changepoints are active?" has a clean answer only for the MAP. For the SVI posterior you need a different, honest summary, or you will either call every candidate "inactive" (all intervals include 0) or claim a sparsity the model does not have.
Where is it used?
Reporting trend changes from SVI or NUTS fits of Prophet-style models, comparing your model with Prophet's MAP output, and every "Bayesian lasso" analysis.
How is it used?
Plot the posterior slope over time with a band; report changes over windows; if you need a yes/no list, threshold on a practically meaningful effect size with posterior probability, and say that it is a decision rule, not a property of the prior.
"The posterior says most changepoints are exactly zero."
Only the MAP has exact zeros. The posterior concentrates near zero; its draws, means and medians are not exactly zero.
"The 90% interval of δⱼ includes 0, so the trend did not change there."
Neighbouring candidates often share one real change, so each δⱼ alone is uncertain. Check the posterior of the slope across the stretch (or the sum of the neighbours' δⱼ) before concluding anything.
Your model is fitted with SVI (a full-rank or low-rank Gaussian guide). Its approximate posterior over δⱼ is a continuous Gaussian: no δⱼ is ever exactly 0 in it, and with correlated hinge columns the guide's covariance matters (a low-rank guide may not capture all the negative correlation between neighbouring δⱼ that "share" a change). So "which changepoints are active?" is answered best by the posterior slope over time with a band, not by a list of non-zero δⱼ. If you compare with Prophet's default output, remember that Prophet reports a MAP.
"A Laplace prior is L1 regularisation, so my Bayesian changepoint model is sparse."
"A Laplace prior gives L1 regularisation at the MAP. The full posterior is shrunk toward zero but not sparse."
Model answer: "Minus the log posterior with δⱼ ~ Laplace(0, b) is squared error over 2σ² plus Σ|δⱼ|/b, so the MAP is a lasso with λ = σ²/b and can set some slope changes exactly to zero. But I fit the full posterior with SVI, and a continuous posterior gives zero probability to δⱼ = 0: most δⱼ are concentrated near zero, a few are large, so it is sparse-ish. I report the posterior trend slope over time rather than a list of active changepoints."
$P(\delta_j = 0\mid y) = 0$: posterior draws/means/medians are shrunk, not zero. Only the MAP is sparse.
Real changes are often shared among neighbouring candidates: read the slope over time, not single δⱼ.
Trap: "Laplace prior ⇒ sparse posterior" or "interval includes 0 ⇒ no change".
Quick check: a colleague reports "the posterior mean of δ₁₂ is 0.004, so changepoint 12 was switched off". What would you say?
That the posterior mean is small but not 0 (no posterior summary is exactly 0 under a Laplace prior), and that a single δⱼ can look small when a real change is shared with its neighbours. Check the posterior of the slope around candidate 12, or the sum of nearby δⱼ, before calling it switched off.
The scale b, and what Prophet's 0.05 really means core
b is the typical size of a slope change: $E|\delta_j| = b$. But "size" has units. A slope change is "extra y per unit of time", so the same number b means completely different things if y is measured in orders or in thousands of orders, and if time is measured in days or in years.
Prophet fixes the units before applying its prior. It divides y by the largest absolute value in the history (so the scaled series has maximum size 1) and measures time so that the whole history runs from 0 to 1. Its default changepoint_prior_scale = 0.05 is b in those units: a typical slope change adds "5% of the series' maximum per history length". Translated back to orders per day, that depends on how big the series is and how long the history is.
Three ways to say it:
- Picture: b is a ruler length, and Prophet first shrinks the drawing to a 1 × 1 box before measuring with it.
- Numbers: for a 2-year history with maximum 1 000 orders/day, b = 0.05 is about 25 extra orders/day per year (0.068 per day per day); with a 1-year history the same 0.05 is 50 per year.
- Slogan: a prior scale is meaningless without the scaling of y and t.
Daily orders with largest value 1 000 orders/day over a 730-day history; Prophet-style scaling; $b = 0.05$.
- Scaled units: y in "fractions of 1 000", t in "fractions of 730 days". A slope change δ = 0.05 means +0.05 scaled units per history length.
- Back to orders: $0.05 \times 1000 = 50$ orders/day per 730 days, that is $50/730 = 0.0685$ orders/day per day, or $0.0685 \times 365 \approx 25$ orders/day per year.
- Effect on the level: 90 days after a changepoint with δ = b, the trend is $0.0685 \times 90 \approx 6.2$ orders/day higher than without it.
- Rare changes: $P(|\delta| \gt 3b) = e^{-3} \approx 5\%$, so changes bigger than $3 \times 25 = 75$ orders/day per year are a priori rare.
- Same b, same series, but only 365 days of history: $50/365 = 0.137$ orders/day per day (50 per year): twice as flexible per calendar day.
Prophet's scaling and trend prior (as in recent Prophet 1.1.x releases; confirm against the source of the version you have installed):
- $\tilde y = y/\max|y|$ over the history (
scaling="absmax", the default;"minmax"uses $(y - \min y)/(\max y - \min y)$). - $\tilde t = (t - t_{first})/(t_{last} - t_{first})$, so the history is $[0, 1]$; candidates:
n_changepoints = 25evenly spaced rows in the firstchangepoint_range = 0.8of the history. - Priors (Stan): $k \sim N(0, 5^2)$, $m \sim N(0, 5^2)$, $\delta_j \sim \text{double\_exponential}(0, \tau)$ with $\tau = $
changepoint_prior_scale$= 0.05$ by default. "double_exponential" is the Laplace, so $b = \tau$. - The default fit is the MAP (
mcmc_samples = 0). For forecast intervals Prophet simulates new future changepoints at the historical rate, with sizes drawn from a Laplace whose scale is the average fitted $|\delta_j|$: so b also shapes the width of Prophet's trend uncertainty. - Conversion to original units: $\delta_{\text{per day}} = \tilde\delta \times \max|y| / (\text{history length in days})$.
Why do we need it?
"We used b = 0.05 like Prophet" is only meaningful if your model scales y and t the way Prophet does. Translating b into orders per day per year lets you judge whether the prior allows plausible trend changes.
Where is it used?
Prophet's changepoint_prior_scale, any NumPyro reimplementation of a Prophet-style trend (your model), and prior predictive checks of trend flexibility (Chapter 6.2).
How is it used?
Write down how your code scales y and t, convert b into business units with the formula above, and draw a few prior trends with that b (the widget below). If they look absurd (or rigid), the b, the scaling, or both need changing.
"0.05 is the standard value for the changepoint prior."
0.05 is Prophet's default on Prophet's scaling. In a model that standardises y (z-scores), measures time in days, or does not scale at all, the same 0.05 is a different prior, possibly absurdly tight or loose.
"b only affects the fit, not the forecast intervals."
In a full Bayesian fit, b shapes the posterior of the recent slope, which drives the forecast spread. In Prophet, future trend changes are simulated with a Laplace whose scale comes from the fitted |δⱼ|, which depend on b.
"We used Prophet's default changepoint prior scale of 0.05, so our trend prior is standard."
"0.05 is the Laplace scale b in Prophet's scaled units (y divided by its max, time over the history scaled to [0, 1]); I checked what it means in our units."
Model answer: "b is the expected absolute slope change, $E|\delta_j| = b$. Prophet applies it after dividing y by its maximum and mapping the history to [0, 1], so 0.05 means a typical change of 5% of the max per history length; for example, with two years of data and a max around 1 000 orders/day that is about 25 orders/day per year. In our NumPyro model I check how y and t are scaled before reusing any number, and I tune b by rolling-origin error and a sensitivity check." (Adapt to what you actually did.)
Before quoting a value of b for your forecasting model, check two lines of your code: how y is scaled before the trend (divided by its max, standardised, log-transformed, or raw) and how time is encoded (days, or scaled to [0, 1]). Only with Prophet-like scaling does 0.05 carry Prophet's meaning. Then translate your b into orders/day per year as above; it is the clearest way to defend the choice in an interview.
Prophet: $\tilde y = y/\max|y|$, $\tilde t \in [0, 1]$ over the history, $\delta_j \sim Laplace(0, 0.05)$ (default), MAP fit by default.
Real units: $\delta_{\text{per day}} = \tilde\delta \cdot \max|y| / \text{days}$; e.g. 0.05 · 1000 / 730 ≈ 0.068/day per day ≈ 25/day per year.
Trap: reusing 0.05 under a different scaling of y or t.
Quick check: a series with max 400 orders/day and 365 days of history, b = 0.05 in Prophet units. What is a typical slope change in orders/day per year?
$0.05 \times 400 = 20$ orders/day per history length (one year): about 20 orders/day per year, or $20/365 \approx 0.055$ per day per day.
Sensitivity to b: too much shrinkage underfits, too little overfits core
b is the trend's flexibility knob. Turn it down and the pull toward 0 is overwhelming: every slope change is crushed, the trend becomes one straight line, and real bends are missed. That is underfitting (too much bias): the forecast keeps an old slope that is no longer true. Turn it up and the pull nearly vanishes: the trend bends at many candidates to follow noise, and the last slope, which the forecast extends, jumps around from sample to sample. That is overfitting (too much variance).
In between there is a sweet spot, and the way to find it is to forecast data the model has not seen: holdout or rolling-origin error over a range of b. Then check that your conclusions do not hinge on the exact value (a sensitivity analysis).
Three ways to say it:
- Picture: a stiff steel ruler (tiny b), a bendy strip of rubber (huge b), and a flexible ruler that bends only where pushed (good b).
- Numbers: averaged over 20 series, test error is 0.29 at b = 0.001, 0.047 at b = 0.1, and 0.053 at b = 1, while training error keeps falling from 0.095 to 0.041.
- Slogan: training error always likes a bigger b; test error does not.
The widgets' setting: 100 training days, 30 test days, 25 candidates, noise sd ≈ 0.045 (scaled units). MAP fits for a grid of b, averaged over 20 simulated series (the U-curve widget below reproduces this).
- $b = 0.001$: no slope change survives (0 non-zero on average). The trend is a straight line; training RMSE 0.095, test RMSE 0.288. Severe underfitting.
- $b = 0.01$: about 1.5 changes survive; test RMSE 0.122. Still underfitting (the big bends are only partly followed).
- $b = 0.1$: about 4 non-zero; training 0.043 (about the noise level), test 0.047: the best region.
- $b = 1$: about 8 non-zero; training 0.041 (lower), test 0.053 (higher). Without any prior: training 0.037, test 0.057.
- The curve is steep on the "too small b" side and gentle on the "too large b" side in this example: crushing real changes is very expensive, chasing noise is moderately expensive.
- Underfitting (b too small): systematic error; residuals show structure (runs of positive then negative residuals around the missed bends, Chapter 7.17); forecasts keep an outdated slope.
- Overfitting (b too large): training error below the noise level, many active changes, a noisy last slope, forecasts that vary a lot between neighbouring origins.
- Choosing b: evaluate a log grid (Prophet's documentation suggests searching roughly 0.001 to 0.5 on its scale) by rolling-origin error (Chapter 7.15), ideally also by interval coverage and CRPS (Chapter 7.16); pick a value near the minimum; prefer a flat region.
- Sensitivity analysis (Chapter 6.8): refit with, say, b/3 and 3b; report whether forecasts and decisions change.
- Learning b: a hyperprior such as $b \sim \text{HalfNormal}$ lets the data choose (a hierarchical prior), but b is weakly identified with few real changes; check the result with prior predictive draws.
- The effective flexibility depends on b and on the number and placement of candidates (grid + PELT): changing the candidate list changes the best b.
Why do we need it?
b is the single most influential setting of the trend: it decides whether the forecast follows a new growth rate or ignores it. An interviewer will ask how you chose it and how sensitive the forecasts are to it.
Where is it used?
Tuning changepoint_prior_scale in Prophet, the Laplace scale in your NumPyro trend, the λ of any lasso, and the bias–variance discussion of model complexity (Chapter 7.18).
How is it used?
Loop over a log grid of b; for each, run the rolling-origin backtest (refitting everything inside each fold); plot error vs b; choose near the minimum; then report forecasts at b/3 and 3b to show robustness.
"Pick the b with the lowest training error."
Training error always prefers the largest b (least shrinkage). Only out-of-sample error (holdout, rolling origin) can show overfitting.
"A wide prior (large b) is the safe, uninformative choice."
With 25 candidates, a wide prior lets the trend chase noise and makes the extrapolated slope unstable. "Letting the data speak" with many free bends means letting the noise speak too.
"Tune b on the same test period you report."
Then the reported error is optimistic (the test period was used for a choice). Tune within rolling-origin folds, and report on data not used for tuning.
In your forecasting model, b controls the bias–variance balance of the trend, together with the candidate list from the grid and PELT. Underfitting (b too small) shows up as forecasts that keep an outdated growth rate and as residuals with long runs; overfitting (b too large) as wiggly fitted trends, training error below the noise level and unstable last slopes. Because the candidates partly come from PELT, the best b and the PELT penalty interact: more PELT candidates means each must be shrunk harder. A short sensitivity table (forecast and interval width at b/3, b, 3b) is a strong answer to "how did you choose your priors?".
"We set b large so the prior would not influence the results."
"b always influences the trend: small b underfits, large b overfits. We chose it by out-of-sample error and checked sensitivity."
Model answer: "The Laplace scale b is the flexibility of the trend. Too small and every slope change is shrunk away, so the model misses real changes in growth; too large and it bends at many candidates to fit noise, which makes the extrapolated slope unstable. I translated b into business units, evaluated a log grid with rolling-origin backtests, picked a value near the minimum in a flat region, and showed that forecasts at b/3 and 3b tell the same story." (Adapt to what you actually did.)
Small b → stiff straight trend (underfit, high bias); large b → wiggly trend, noisy last slope (overfit, high variance).
Choose b by rolling-origin error over a log grid; check b/3 and 3b; the best b depends on the candidate list.
Trap: choosing b by training error or on the reported test period.
Quick check: training RMSE 0.030, noise level about 0.045, test RMSE much worse than at a smaller b. Which way should b move?
Down. A training error clearly below the noise level means the trend is fitting noise (overfitting); more shrinkage (smaller b) should lower the test error.
Recap, cheat sheet and practice
- Many candidate changepoints + free slope changes = a trend that chases noise. The prior $\delta_j \sim Laplace(0, b)$ says "most changes tiny, a few large" ($E|\delta_j| = b$, sd $\sqrt2 b$).
- Compared with a Normal of the same variance, the Laplace has a sharper peak and heavier tails, and pulls with a constant force $1/b$: it squashes small changes and spares big ones.
- Shrinkage needs evidence to overcome: for one slope change the MAP stays at 0 until $|x| \gt se^2/b$, and the evidence grows like $n^3/3$ days after the change. Forecasts react late by design.
- L1 relationship: the MAP is a lasso, $\tfrac12\text{RSS} + \lambda\sum|\delta_j|$ with $\lambda = \sigma^2/b$, solved by soft-thresholding (ISTA/FISTA); only corner-aware solvers give exact zeros.
- The full posterior is shrunk, not sparse: no draw is exactly 0, real changes are often shared by neighbouring candidates; report the slope over time.
- Scale: Prophet's 0.05 is b after $y/\max|y|$ and $t \in [0, 1]$; in real units $\delta_{\text{per day}} = \tilde\delta\,\max|y|/\text{days}$.
- Sensitivity: small b underfits (straight line, outdated slope), large b overfits (wiggles, unstable last slope). Choose b by rolling-origin error over a log grid and show b/3 and 3b.
Cheat sheet
| Idea | Formula / setting | Remember |
|---|---|---|
| Trend with candidates | $g(t) = m + kt + \sum_j \delta_j (t - s_j)_+$ | slope after $s_j$: $k + \sum_{i \le j}\delta_i$ |
| Laplace prior | $p(\delta) = \frac{1}{2b}e^{-|\delta|/b}$; $E|\delta| = b$; Var $= 2b^2$ | NumPyro Laplace(0., b), SciPy laplace(0, scale=b), Stan double_exponential(0, b) |
| Penalty and pull | Laplace $|\delta|/b$, pull $1/b$; Normal $\delta^2/2\tau^2$, pull $\delta/\tau^2$ | equal variance: $\tau = \sqrt2 b$; pulls cross at $|\delta| = 2b$ |
| Tails | $P(|\delta| \gt c) = e^{-c/b}$ | $P(\lt b/2) = 0.39$, $P(\gt 3b) = 0.05$ |
| One change, MAP | $\text{sign}(x)\max(0, |x| - se^2/b)$ | exact 0 until the evidence beats $se^2/b$ |
| Evidence after a bend | $se^2 = \sigma^2/\sum_{i=1}^n i^2$, $\sum i^2 \approx n^3/3$ | late reaction after real changes |
| MAP = lasso | $\lambda = \sigma^2/b$ (½RSS convention); sklearn $\alpha = \lambda/n$ | bigger b = weaker penalty |
| ISTA step | $\delta \leftarrow S_{\eta/b}(\delta + \frac{\eta}{\sigma^2}X^\top(y - X\delta))$, $\eta = 1/L$ | FISTA = with momentum |
| Posterior | $P(\delta_j = 0\mid y) = 0$ | sparse-ish; read slopes, not single δⱼ |
| Prophet scaling | $y/\max|y|$, $t \in [0, 1]$, 25 candidates in first 80%, $\tau = 0.05$, MAP fit | 0.05 only means this under this scaling |
| Real units | $\delta_{\text{day}} = \tilde\delta\,\max|y|/\text{days}$ | 0.05 · 1000/730 ≈ 25 orders/day per year |
| Choosing b | log grid + rolling origin; sensitivity at b/3, 3b | never by training error |
import numpy as np
import jax
import jax.numpy as jnp
import numpyro
import numpyro.distributions as dist
from numpyro.infer import MCMC, NUTS, SVI, Trace_ELBO
from numpyro.infer.autoguide import AutoDelta
from sklearn.linear_model import Lasso
# 1) Data like the widgets: 100 training days + 30 test days, scaled the way Prophet scales
rng = np.random.default_rng(3)
T, Hh = 100, 30
t = np.arange(T + Hh) / (T - 1) # time: the history runs from 0 to 1
g_true = 0.4 + 0.3 * t + 0.9 * np.clip(t - 0.25, 0, None) - 1.4 * np.clip(t - 0.55, 0, None)
y_raw = g_true + 0.04 * rng.standard_normal(T + Hh)
y = y_raw / np.abs(y_raw[:T]).max() # 'absmax' scaling: divide by the largest |y| of the history
hist = int(np.floor(0.8 * T))
s = t[np.linspace(0, hist - 1, 26).round().astype(int)[1:]] # Prophet's rule: 25 candidates in the first 80%
X = np.clip(t[:, None] - s[None, :], 0, None) # hinge columns (t - s_j)_+
# 2) The model: delta_j ~ Laplace(0, b) (NumPyro: Laplace(loc, scale) with scale = b)
def model(t, X, b, y=None):
m = numpyro.sample("m", dist.Normal(0.0, 5.0))
k = numpyro.sample("k", dist.Normal(0.0, 5.0))
delta = numpyro.sample("delta", dist.Laplace(0.0, b).expand([X.shape[1]]))
sigma = numpyro.sample("sigma", dist.HalfNormal(0.5))
g = m + k * t + X @ delta
numpyro.sample("y", dist.Normal(g, sigma), obs=y)
# 3) Full posterior (NUTS) for a small and a large b
for b in [0.005, 0.05, 0.5]:
mcmc = MCMC(NUTS(model), num_warmup=500, num_samples=1000, progress_bar=False)
mcmc.run(jax.random.PRNGKey(0), t[:T], X[:T], b, y=y[:T])
p = {k_: np.asarray(v) for k_, v in mcmc.get_samples().items()}
d = p["delta"]
trend = p["m"][:, None] + p["k"][:, None] * t + d @ X.T # 1000 trend draws over all 130 days
test_rmse = np.sqrt(np.mean((trend.mean(0)[T:] - y[T:]) ** 2))
train_rmse = np.sqrt(np.mean((trend.mean(0)[:T] - y[:T]) ** 2))
last_slope = p["k"] + d.sum(1) # slope after the last candidate
print(b, int((d == 0).sum()), np.round(np.percentile(last_slope, [5, 95]), 2),
round(float(train_rmse), 3), round(float(test_rmse), 3))
# b exact zeros 90% interval of the last slope train RMSE test RMSE
# 0.005 0 [ 0.44 0.57] 0.096 0.275 too much shrinkage: a straight line
# 0.05 0 [-0.36 -0.07] 0.053 0.051
# 0.5 0 [-0.49 -0.03] 0.048 0.055 lower training error, worse test error
print(round(float((g_true[-1] - g_true[-2]) * (T - 1) / np.abs(y_raw[:T]).max()), 2)) # -0.23 (true last slope)
# 4) The MAP with sigma fixed is a lasso: lambda = sigma^2 / b; sklearn uses alpha = lambda / n
sigma, b = 0.04 / np.abs(y_raw[:T]).max(), 0.05
A = np.column_stack([np.ones(T), t[:T]]) # m and k are not penalised: project them out
P = A @ np.linalg.solve(A.T @ A, A.T)
lasso = Lasso(alpha=sigma**2 / (b * T), fit_intercept=False, tol=1e-12, max_iter=10**6)
d_map = lasso.fit(X[:T] - P @ X[:T], y[:T] - P @ y[:T]).coef_
print(int((d_map == 0).sum()), np.round(d_map[d_map != 0], 3)) # 22 [ 0.227 -0.41 -0.737] 22 of 25 exactly 0
# 5) A MAP found by gradient steps (SVI with an AutoDelta guide) lands NEAR zero, not ON zero
def model_fixed_sigma(t, X, b, y=None):
m = numpyro.sample("m", dist.Normal(0.0, 5.0))
k = numpyro.sample("k", dist.Normal(0.0, 5.0))
delta = numpyro.sample("delta", dist.Laplace(0.0, b).expand([X.shape[1]]))
numpyro.sample("y", dist.Normal(m + k * t + X @ delta, sigma), obs=y)
svi = SVI(model_fixed_sigma, AutoDelta(model_fixed_sigma), numpyro.optim.Adam(0.002), Trace_ELBO())
res = svi.run(jax.random.PRNGKey(1), 20000, t[:T], X[:T], b, y=y[:T], progress_bar=False)
d_svi = np.asarray(res.params["delta_auto_loc"])
print(int((d_svi == 0).sum()), int((np.abs(d_svi) < 0.01).sum()), f"{np.abs(d_svi).min():.1e}") # 0 21 1.9e-06: tiny, never exactly 0
1. For $\delta \sim Laplace(0, b)$, what are $E|\delta|$ and the standard deviation?
2. With Normal noise of sd σ and $\delta_j \sim Laplace(0, b)$, the MAP of the slope changes is the same as…
3. You fit the model with SVI (a Gaussian guide) and draw 4 000 samples of the 25 slope changes. How many draws are exactly 0?
4. What does Prophet's default changepoint_prior_scale = 0.05 mean?
5. As b grows from 0.001 to 2 on the same data, what happens to training and test error?
6. One candidate with known rest of the model: $se^2 = 0.04$, $b = 0.1$, data estimate $x = 0.3$. What is the Laplace MAP of δ?
Practice problems
A. With b = 0.05, compute $P(|\delta| \lt 0.01)$ and $P(|\delta| \gt 0.2)$, and the expected number of the 25 candidates with $|\delta| \gt 0.15$ in a prior draw.
- $P(|\delta| \lt 0.01) = 1 - e^{-0.01/0.05} = 1 - e^{-0.2} = 0.181$.
- $P(|\delta| \gt 0.2) = e^{-0.2/0.05} = e^{-4} = 0.018$.
- $P(|\delta| \gt 0.15) = e^{-3} = 0.0498$; expected count $25 \times 0.0498 \approx 1.2$ candidates.
B. One candidate, $se = 0.2$, $b = 0.05$. Find the Laplace MAP for $x = 1.0$ and $x = -0.5$, and compare with the posterior mean under a Normal prior of the same variance.
- Threshold $se^2/b = 0.04/0.05 = 0.8$. $x = 1.0$: MAP $= 1.0 - 0.8 = 0.2$. $x = -0.5$: $|x| \lt 0.8$, MAP $= 0$.
- Normal prior with $\tau^2 = 2b^2 = 0.005$: factor $\tau^2/(\tau^2 + se^2) = 0.005/0.045 = 0.111$. Means: $0.111$ and $-0.056$.
- The Laplace keeps more of the large estimate (0.2 vs 0.11) and kills the small one (0 vs −0.056): squash small, spare big.
C. A series has max 250 orders/day and 3 years (1 095 days) of history. With Prophet-style scaling and b = 0.05, what is a typical slope change in orders/day per year? The business expects growth changes of about 20 orders/day per year. Is b too small?
$0.05 \times 250/1095 = 0.0114$ orders/day per day $\approx 4.2$ per year. A change of 20 per year is about $4.8b$, and $P(|\delta| \gt 4.8b) = e^{-4.8} \approx 0.008$: the prior finds such changes very unlikely, so it will shrink them hard. b looks too small for this business; try larger values and confirm with a rolling-origin backtest.
D. (Interview) "In your posterior, every δⱼ has a 90% interval that includes 0. So the trend never changed?"
"Not necessarily. When the data cannot pin down exactly where a bend happens, the posterior spreads it over several neighbouring candidates; each δⱼ alone is uncertain and its interval includes 0, but their sum, and the slope over that stretch, can be clearly different from before. So I look at the posterior of the slope over time, or the total change over a window, not at single δⱼ. Also, with a Laplace prior no δⱼ is ever exactly 0 in the posterior, so 'includes 0' is the normal state for most candidates."
E. (Interview) "Why a Laplace prior on the slope changes and not a Normal one?"
"Because we believe most candidate changepoints do nothing and a few matter. With the same variance, a Laplace puts more mass near zero and more in the tails than a Normal, and its penalty |δ|/b pulls with constant force, so it shrinks small, noise-driven changes strongly while letting large, well-supported changes through. A Normal shrinks every change by the same fraction, so it either flattens real changes or lets every candidate wiggle. At the MAP the Laplace gives a lasso, which can switch changes off exactly; in the full posterior it gives sparse-ish, not sparse, changes."
F. Noise sd 0.03 (scaled), b = 0.02, 365 training days. What λ and what scikit-learn α give the MAP of the slope changes?
$\lambda = \sigma^2/b = 0.0009/0.02 = 0.045$; $\alpha = \lambda/n = 0.045/365 \approx 1.23 \times 10^{-4}$ (and remember to leave $m$ and $k$ unpenalised, for example by projecting them out as in the Code-it).
Fourier seasonality and Fourier order
Your forecasting model draws the weekly and yearly rhythm of demand with sines and cosines. This chapter builds that idea from a point going round a circle: period and frequency, harmonics, amplitude and phase, how a few waves add up to any repeating shape, and how those waves become plain columns of a regression. Then the practical questions an interviewer will ask: how do you choose the Fourier order, why can a large order hurt, what goes wrong with a short history, how do several seasonalities live together, and why can daily data never show a pattern faster than two days?
- Read a sine and a cosine wave: period $P$, frequency $1/P$, angular frequency $2\pi/P$, and why the waves average to zero
- Know what a harmonic is and why every harmonic of $P$ also repeats every $P$
- Derive $a\cos(\omega t) + b\sin(\omega t) = R\cos(\omega t - \varphi)$ and turn a sine/cosine pair into an amplitude and a phase (the peak day)
- Write the Fourier series $s(t) = \sum_{n=1}^{N}\left[a_n\cos\frac{2\pi n t}{P} + b_n\sin\frac{2\pi n t}{P}\right]$ and see how the order $N$ controls the detail of the seasonal shape
- Build Fourier terms as regression columns (2 columns per harmonic, $2N$ per seasonality) and fit them by least squares
- Treat the order as a bias–variance knob: watch training error fall and holdout error turn back up
- Spot identifiability problems: a short history cannot tell a yearly wave from a trend
- Combine weekly, yearly and other seasonalities, and state Prophet's documented default orders (yearly 10, weekly 3, daily 4)
- Explain aliasing and the Nyquist frequency: why weekly order 4 adds nothing for daily data, and why daily data cannot fit a daily cycle
What we need from earlier chapters: seasonality as a pattern with a fixed, known period, centred so its effects sum to zero (Chapter 7.2); frequency and the units of $t$ (Chapter 7.1); regression as "columns times weights", least squares, collinearity and the VIF (Chapter 5.13); ridge regression as a Normal prior (Chapter 5.3); the bias–variance trade-off for estimators (Chapter 5.1); identifiability, "which component gets the credit?" (Chapter 6.8); the additive model $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$ (Chapter 7.7). The Calculus guide already met the unit circle, radians, and $A\sin(\omega x + \varphi)$ (Chapter 2.1, trigonometric functions); here we use those waves as regression columns. Notation: $t$ is time (in days unless we say otherwise), $P$ the period in the same units as $t$, $n$ the harmonic number, $N$ the Fourier order, $a_n, b_n$ the cosine and sine weights, $\omega = 2\pi/P$ ("omega").
Sine and cosine: the simplest repeating shapes core
Sit on a Ferris wheel that turns at a steady speed. Your height goes up, then down, then up again, and the pattern repeats exactly once per lap. Draw your height against time and you get a smooth wave. That wave is a sine wave. Your left-right position draws the same wave, just a quarter of a lap earlier: a cosine wave.
Why do we care in forecasting? A seasonal pattern is something that repeats every $P$ days (every 7 days for a week). Sine and cosine are the simplest smooth shapes that repeat. If we stretch them so that one lap takes exactly $P$ days, they become the building bricks of any seasonal pattern. The rest of this chapter is about stacking these bricks.
Three ways to say it:
- Picture: a point going round a circle at a steady speed; its height over time is a sine wave, its sideways position a cosine wave.
- Numbers: with $P = 7$ days, $\sin(2\pi t/7)$ takes the values 0, 0.78, 0.97, 0.43, −0.43, −0.97, −0.78 on days 0 to 6, and is back at 0 on day 7.
- Slogan: one lap of the circle = one period of the season.
A weekly wave, day by day. Period $P = 7$ days. One full lap is $2\pi \approx 6.283$ radians (a radian is the angle unit maths and NumPy use; a full turn is $2\pi$ radians $= 360°$).
- Angle gained per day: $2\pi/7 \approx 0.898$ radians ($360°/7 \approx 51.4°$). This is the angular frequency $\omega$.
- Day 1: angle $0.898$. Height $\sin(0.898) \approx 0.782$; sideways position $\cos(0.898) \approx 0.623$.
- Day 2: angle $1.795$. $\sin \approx 0.975$, $\cos \approx -0.223$. Day 3: $\sin \approx 0.434$, $\cos \approx -0.901$.
- Days 4, 5, 6 mirror days 3, 2, 1 with the opposite sign for the sine: $-0.434, -0.975, -0.782$.
- Day 7: angle $7 \times 0.898 = 6.283 = 2\pi$: one full lap, back to $\sin = 0$, $\cos = 1$. The wave repeats.
- Add the seven sine values: $0 + 0.782 + 0.975 + 0.434 - 0.434 - 0.975 - 0.782 = 0$. Over a whole period the wave averages to zero: it pushes some days up and others down by the same total.
- Frequency $= 1/P = 1/7 \approx 0.143$ cycles per day: about one seventh of a lap per day.
For a period $P \gt 0$ (measured in the same units as the time $t$), the two basic waves are
$$\sin\!\left(\frac{2\pi t}{P}\right), \qquad \cos\!\left(\frac{2\pi t}{P}\right).$$- Period $P$: the time for one full cycle. The waves repeat: $f(t + P) = f(t)$.
- Frequency $f = 1/P$: cycles per unit of time (per day if $t$ is in days).
- Angular frequency $\omega = 2\pi/P = 2\pi f$: radians per unit of time. The waves can be written $\sin(\omega t)$ and $\cos(\omega t)$.
- Both waves stay between $-1$ and $1$, and both average to zero over one full period.
- A cosine is a sine shifted a quarter period: $\cos(x) = \sin(x + \pi/2)$. Their peaks are a quarter of a period apart.
- The input $2\pi t/P$ is an angle in radians.
np.sin,jnp.sinandMath.sinall expect radians, never degrees.
Why do we need it?
Seasonality is a pattern that repeats every $P$ days. Sine and cosine are the simplest smooth functions that repeat, they average to zero (so they never steal the level from the trend), and they can be sized and shifted. They are the bricks of every Fourier seasonality.
Where is it used?
The seasonal terms of Prophet and your Prophet-style NumPyro model, "Fourier terms" in regression with ARIMA errors (statsmodels and R's fourier()), spectral analysis, audio signals, electricity-load models with daily and weekly cycles, and positional encodings in Transformers.
How is it used?
Pick the period from the calendar ($P = 7$ days for a week, $365.25$ for a year), make sure $t$ and $P$ use the same unit, compute np.sin(2*np.pi*t/P) and np.cos(2*np.pi*t/P), and use them as columns in the model.
"Period and frequency are two names for the same thing."
They are inverses. The weekly wave has period 7 days and frequency $1/7$ cycle per day. A bigger period means a slower wave and a smaller frequency.
"np.sin(t) with $t$ in days gives a weekly wave."
$\sin(t)$ repeats every $2\pi \approx 6.28$ days, which matches no calendar cycle. You must write $\sin(2\pi t/P)$ so that one full lap takes exactly $P$ days. And the argument is in radians, not degrees.
"A sine wave can describe any weekly pattern."
One sine wave has one smooth peak and one smooth trough, exactly half a period apart. Real weeks (a big Saturday, a quiet Tuesday) need several waves added together: the harmonics of the next section.
In your forecasting model, the seasonal columns are sines and cosines of $2\pi n t/P$. The period must be in the units of the $t$ that goes into those columns: $P = 7$ and $P = 365.25$ if $t$ counts days, $P = 168$ and $P = 24$ if $t$ counts hours. Prophet, for example, builds its seasonal columns from $t$ measured in days since 1970-01-01, even though its trend uses a rescaled time. If your code rescales time to $[0, 1]$ anywhere, check which time the Fourier columns receive: a weekly period of 7 applied to a rescaled $t$ is a silent bug.
$\sin(2\pi t/P)$, $\cos(2\pi t/P)$: period $P$ (time per cycle), frequency $f = 1/P$, angular frequency $\omega = 2\pi/P$.
Both lie in $[-1, 1]$, average 0 over a period; cosine = sine shifted by a quarter period.
Trap: the argument is in radians and $t$, $P$ must share a unit.
Quick check: yearly seasonality in daily data uses $P = 365.25$. What are its frequency and angular frequency?
Frequency $f = 1/365.25 \approx 0.00274$ cycles per day; angular frequency $\omega = 2\pi/365.25 \approx 0.0172$ radians per day. The wave moves about one degree around the circle per day ($360°/365.25 \approx 0.99°$).
Harmonics: faster waves that still fit the period core
Pluck a guitar string. You hear the main note, but the string also vibrates in two halves, in three thirds, and so on. Those extra vibrations are called harmonics, and they all fit exactly into the length of the string.
A weekly pattern works the same way. The slowest wave fits once into the week (period 7 days). A wave that fits twice into the week (period 3.5 days) can describe "busy at the start of the week and again near the end". A wave that fits three times (period 2.33 days) adds even finer detail. The key fact: because each of these waves fits a whole number of times into 7 days, every one of them also repeats every 7 days. So any sum of them is still a weekly pattern.
Three ways to say it:
- Picture: harmonic $n$ squeezes $n$ full waves into one period.
- Numbers: for $P = 7$: harmonic 1 has period 7 days, harmonic 2 has period 3.5 days, harmonic 3 has period $7/3 \approx 2.33$ days.
- Slogan: faster waves add detail, but they all come back in step after one period.
Two harmonics added together. $s(t) = 10\cos(2\pi t/7) + 4\cos(4\pi t/7)$: harmonic 1 with weight 10 plus harmonic 2 with weight 4.
- Day 0: both cosines are $\cos 0 = 1$, so $s(0) = 10 + 4 = 14$.
- Day 1: $\cos(2\pi/7) \approx 0.623$ and $\cos(4\pi/7) \approx -0.223$, so $s(1) = 6.235 - 0.890 \approx 5.35$.
- Harmonic 2 has period $7/2 = 3.5$ days: at day 3.5 it has finished one full cycle, at day 7 it has finished two.
- Harmonic 1 finishes its single cycle at day 7. So at day 7 both are back at the start: $s(7) = 14 = s(0)$.
- The same holds for every day: $s(8) = s(1) \approx 5.35$, $s(9) = s(2)$, and so on. The sum repeats every 7 days.
- Values on days 0–6: $14, 5.35, -5.83, -6.52, -6.52, -5.83, 5.35$. One high peak and a broad flat trough: a shape one wave alone could not make (a single cosine would give a symmetric, rounder trough).
For a period $P$, the $n$-th harmonic ($n = 1, 2, 3, \dots$) is the pair of waves
$$\sin\!\left(\frac{2\pi n t}{P}\right), \qquad \cos\!\left(\frac{2\pi n t}{P}\right).$$- Its period is $P/n$ and its frequency is $n/P$: harmonic $n$ is $n$ times faster than harmonic 1 (the fundamental).
- Because $n$ cycles fit exactly into $P$, every harmonic satisfies $f(t + P) = f(t)$. Any weighted sum of harmonics of $P$ is periodic with period $P$.
- Higher harmonics carry finer detail: harmonic $n$ can describe features about $P/(2n)$ wide (half of its own period).
- Each harmonic also averages to zero over a full period, so sums of harmonics are automatically centred.
Why do we need it?
One wave has only one smooth bump per period. Real seasonal shapes are lopsided: a sharp weekend peak, a short December rush. Higher harmonics add exactly the extra detail needed, while keeping the pattern periodic and centred.
Where is it used?
The "order" of every Fourier seasonality (Prophet's fourier_order, the $K$ in R's fourier(x, K)), MP3 and JPEG compression (keep the important frequencies), electricity and traffic models, and music synthesizers.
How is it used?
Include harmonics $n = 1, \dots, N$ of each period as columns. Read the fitted weights harmonic by harmonic: large weights on high $n$ mean a sharp, detailed seasonal shape; tiny weights mean you could use a lower order.
"Harmonic 2 has period $2P$."
Harmonic $n$ is faster: period $P/n$. Harmonic 2 of a week has period 3.5 days. A slower wave with period $2P$ would not repeat every $P$ and is not part of the seasonality of period $P$.
"A twice-a-week pattern needs its own seasonality with $P = 3.5$."
It is already harmonic 2 of the weekly term. Adding a separate $P = 3.5$ seasonality on top of weekly order $\ge 2$ duplicates the same columns and makes the fit unidentifiable.
Harmonic $n$ of period $P$: $\sin(2\pi n t/P)$, $\cos(2\pi n t/P)$; period $P/n$, frequency $n/P$.
All harmonics repeat every $P$ and average to 0, so any sum of them is a centred pattern with period $P$.
Trap: higher $n$ = faster and more detailed, not slower.
Quick check: yearly seasonality ($P = 365.25$ days) with order 10. What is the period of the fastest harmonic, and roughly how narrow a feature can it draw?
Harmonic 10 has period $365.25/10 \approx 36.5$ days, so it can describe bumps about half that wide, roughly 18 days. A spike that lasts 2 or 3 days (a single holiday) is far too narrow for yearly order 10; that is a job for a holiday indicator (Chapter 7.12).
Amplitude and phase: why each harmonic gets a sine and a cosine core
A weekly wave needs two facts: how big it is (the height of its peak) and when the peak happens (Tuesday? Saturday?). Those are its amplitude and its phase.
You might expect the model to learn "size" and "peak day" directly. It does not. It learns two weights: one for a cosine and one for a sine of the same speed. Mixing a cosine (peak on day 0) and a sine (peak a quarter period later) in the right amounts gives a wave of the same speed whose peak sits anywhere you like. It is like mixing a "north" step and an "east" step to walk in any direction.
Why not just learn size and peak day? Because then the model would be non-linear and awkward (peak day wraps around: day 6.9 is right next to day 0.1). With a cosine weight and a sine weight, the model stays linear: plain least squares and plain Normal priors work.
Three ways to say it:
- Picture: the pair $(a, b)$ is an arrow; its length is the amplitude and its angle is the phase.
- Numbers: $3\cos(\omega t) + 4\sin(\omega t) = 5\cos(\omega t - 0.927)$: amplitude 5, peak about 1 day after the start of the week.
- Slogan: a cosine weight and a sine weight = a size and a timing, in a form the model can fit linearly.
Weekly harmonic 1 with weights $a = 3$ (cosine) and $b = 4$ (sine); $\omega = 2\pi/7 \approx 0.898$ per day; $t = 0$ is Monday 00:00.
- Amplitude: $R = \sqrt{a^2 + b^2} = \sqrt{9 + 16} = \sqrt{25} = 5$.
- Phase: $\varphi = \operatorname{atan2}(b, a) = \operatorname{atan2}(4, 3) \approx 0.927$ radians $\approx 53.1°$.
- So $3\cos(\omega t) + 4\sin(\omega t) = 5\cos(\omega t - 0.927)$. The peak is where the angle inside the cosine is 0: $\omega t = 0.927$, so $t = 0.927/0.898 \approx 1.03$ days: just after midnight on Tuesday (about 00:48).
- Check at $t = 1$: $3 \times 0.623 + 4 \times 0.782 = 1.870 + 3.127 = 4.997 \approx 5$. Almost exactly the peak height, as expected less than an hour before the peak.
- The trough is half a period later, at $1.03 + 3.5 = 4.53$ days (Friday midday), with value $-5$. Peak-to-trough difference: $2R = 10$.
- Going back: amplitude 5 and peak at day 1.03 give $a = R\cos\varphi = 5 \times 0.6 = 3$ and $b = R\sin\varphi = 5 \times 0.8 = 4$.
For any weights $a$ and $b$ (not both zero) and any angular frequency $\omega$:
$$a\cos(\omega t) + b\sin(\omega t) = R\cos(\omega t - \varphi), \qquad R = \sqrt{a^2 + b^2}, \quad \varphi = \operatorname{atan2}(b, a).$$Derivation. The cosine difference rule says $\cos(x - \varphi) = \cos x\cos\varphi + \sin x\sin\varphi$. So
$$R\cos(\omega t - \varphi) = (R\cos\varphi)\cos(\omega t) + (R\sin\varphi)\sin(\omega t).$$This equals $a\cos(\omega t) + b\sin(\omega t)$ for every $t$ exactly when $a = R\cos\varphi$ and $b = R\sin\varphi$. Squaring and adding: $a^2 + b^2 = R^2(\cos^2\varphi + \sin^2\varphi) = R^2$. Dividing: $\tan\varphi = b/a$; atan2(b, a) picks the angle in the right quarter of the circle from the signs of $a$ and $b$.
- $R$ = amplitude: the height of the peak above the average (peak to trough $= 2R$).
- $\varphi$ = phase: the angle where the peak sits. The peak time is $t^* = \varphi/\omega = \varphi P/(2\pi)$, plus any multiple of $P$ (add $P$ if negative).
- If $a = b = 0$ there is no wave: $R = 0$ and the phase means nothing.
- The same holds for each harmonic $n$ with $\omega_n = 2\pi n/P$ (its peaks repeat every $P/n$).
Why do we need it?
It explains why every harmonic gets two columns: one column can only make waves that peak on a fixed day. It also turns fitted weights into numbers people understand: "the weekly swing is ±5 orders and peaks on Tuesday".
Where is it used?
Reporting seasonal effects from Prophet-style models, AC electricity (amplitude and phase of a current), tides, phasors in engineering, and the polar form of complex numbers ($a + ib = Re^{i\varphi}$), which is what the FFT computes.
How is it used?
After fitting, for each harmonic compute R = np.hypot(a, b) and phi = np.arctan2(b, a), then peak time phi * P / (2*np.pi) % P. With posterior draws, do this for every draw and summarise the draws of $R$, not the $R$ of the averaged weights.
"The sine coefficient is the weekly effect; the cosine coefficient is a small correction."
Neither weight means anything alone. Only the pair matters: $R = \sqrt{a^2 + b^2}$ is the size and $\operatorname{atan2}(b, a)$ the timing. Shift the start of time $t = 0$ by one day and the two weights rotate into each other while the fitted pattern stays the same.
"Average the posterior draws of $a$ and $b$, then compute $R$ from the averages."
That gives a smaller number than the posterior mean of $R$ when the phase is uncertain (draws pointing in different directions cancel). Compute $R$ for every draw, then summarise. $R$ is a non-linear function, and the average of a function is not the function of the average.
"A peak at $\varphi = -1$ radian is a peak before time zero, so it is meaningless."
The pattern repeats, so add one period: $t^* = -1/\omega + P$. Phases live on a circle.
In your forecasting model, each seasonal harmonic contributes a (sine weight, cosine weight) pair. When you explain the weekly pattern in an interview, convert each pair into an amplitude and a peak day, or better, simply plot the fitted seasonal component $s(t)$ over one period (that is what Prophet's component plots do). A useful prior fact: if both weights of a harmonic get independent $N(0, \tau^2)$ priors, the implied prior on the phase is uniform (no favourite peak day) and the amplitude follows a Rayleigh distribution with typical size $\tau\sqrt{\pi/2} \approx 1.25\tau$. So a Normal prior on the weights is a prior on "how big", with no opinion about "when".
"We model weekly seasonality with $\sin(2\pi t/7)$ times a coefficient."
A sine alone forces the peak to day $7/4$; you need the sine and the cosine of each harmonic so the phase can be learned linearly.
Model answer: "Each harmonic has a sine and a cosine column. Together their weights $a, b$ encode an amplitude $\sqrt{a^2 + b^2}$ and a phase $\operatorname{atan2}(b, a)$, so the model can place the weekly peak on any day while staying linear in its parameters."
$a\cos\omega t + b\sin\omega t = R\cos(\omega t - \varphi)$, $R = \sqrt{a^2 + b^2}$, $\varphi = \operatorname{atan2}(b, a)$, peak at $t^* = \varphi/\omega$ (mod $P$).
Two linear weights per harmonic = one size + one timing.
Trap: single weights mean nothing; compute $R$ per posterior draw.
Quick check: a weekly harmonic-1 fit gives $a = -4$, $b = 0$. What are the amplitude and the peak day ($t = 0$ is Monday 00:00)?
$R = \sqrt{16 + 0} = 4$. $\varphi = \operatorname{atan2}(0, -4) = \pi$ radians (180°). Peak at $t^* = \pi/\omega = \pi \times 7/(2\pi) = 3.5$ days: Thursday midday. A negative cosine weight simply means the peak is half a week away from $t = 0$, where the cosine peaks.
The Fourier series: any repeating shape from a few waves core
With a box of Lego bricks you can build almost anything; with more bricks you get finer detail. Harmonics are the bricks of repeating shapes. A Fourier series is the recipe "take harmonic 1 in this amount, harmonic 2 in that amount, and so on". The French mathematician Fourier showed that, with enough harmonics, this recipe can draw essentially any repeating shape.
The number of harmonics you allow is the Fourier order $N$. Order 1 gives one smooth hill and one valley per period. Order 3 can draw a week with a big weekend and a quiet midweek. Order 10 can draw a year with a summer bump and a pre-Christmas rush. Smooth shapes need few harmonics. Sharp jumps need many, and even then the curve slightly overshoots next to the jump.
There is also a ceiling. With daily data, a week has only 7 distinct days. A centred weekly pattern is just 6 free numbers (the 7th is fixed because the effects sum to zero), and order 3 has exactly 6 weights. So weekly order 3 already reproduces any weekday pattern exactly.
Three ways to say it:
- Picture: stack harmonic 1, 2, 3, … like layers of paint; each layer adds finer detail to the same repeating shape.
- Numbers: order $N$ means $2N$ weights; weekly order 3 = 6 weights = the 6 free weekday effects.
- Slogan: the order is the resolution of the seasonal shape.
Quarterly sales, period $P = 4$ quarters. Centred quarterly effects: Q1 $-10$, Q2 $+5$, Q3 $-20$, Q4 $+25$ (sum 0). Put $t = 0, 1, 2, 3$ for Q1 to Q4.
- Harmonic 1 columns at $t = 0..3$: $\cos(2\pi t/4) = 1, 0, -1, 0$ and $\sin(2\pi t/4) = 0, 1, 0, -1$. Harmonic 2: $\cos(2\pi \cdot 2t/4) = \cos(\pi t) = 1, -1, 1, -1$, and $\sin(\pi t) = 0, 0, 0, 0$ (a column of zeros: it carries no information).
- Over one full period these columns are orthogonal (their dot products are 0), so each least-squares weight is simply (data · column) / (column · column).
- $a_1 = \dfrac{(-10)(1) + 5(0) + (-20)(-1) + 25(0)}{1 + 0 + 1 + 0} = \dfrac{10}{2} = 5$.
- $b_1 = \dfrac{(-10)(0) + 5(1) + (-20)(0) + 25(-1)}{2} = \dfrac{-20}{2} = -10$.
- $a_2 = \dfrac{(-10)(1) + 5(-1) + (-20)(1) + 25(-1)}{4} = \dfrac{-60}{4} = -15$.
- Order 1 ($a_1, b_1$ only): fitted values $5, -10, -5, 10$. Errors: $-15, +15, -15, +15$. A smooth wave cannot follow the up-down-up-down zigzag.
- Order 2 (add $a_2$): $5 - 15 = -10$, $-10 + 15 = 5$, $-5 - 15 = -20$, $10 + 15 = 25$. Exact. Three weights for three free quarterly effects.
The same logic for weekly data: 7 days, 6 free effects, and order 3 gives 6 weights, so order 3 is exact (and order 4 is impossible to use, as we will see in the aliasing section).
A Fourier series of order $N$ with period $P$ is
$$s(t) = \sum_{n=1}^{N}\left[a_n\cos\!\left(\frac{2\pi n t}{P}\right) + b_n\sin\!\left(\frac{2\pi n t}{P}\right)\right].$$- $N$ = the Fourier order: how many harmonics are used. There are $2N$ weights $a_1, b_1, \dots, a_N, b_N$. (Strictly, a finite sum is a partial Fourier sum; the full series lets $N \to \infty$.)
- There is no constant term: $s(t)$ averages to zero over a period. The level lives in the trend (or the intercept).
- Approximation: for any reasonable periodic shape, the best order-$N$ sum gets closer as $N$ grows. Smooth shapes converge fast. At a jump the partial sums overshoot by about 9% of the jump size however large $N$ is (the Gibbs phenomenon), although the overshoot gets narrower.
- Exact ceiling for sampled data: if each period contains $P$ equally spaced observations ($P$ a whole number), there are only $P - 1$ free centred values, and order $N = \lfloor P/2 \rfloor$ reproduces any pattern exactly. For even $P$, the top harmonic $n = P/2$ has a cosine but its sine column is all zeros. Weekly in daily data: $N = 3$ (6 weights) equals the 6 day-of-week dummies.
Why do we need it?
It gives one flexible, smooth recipe for any seasonal shape, with a single knob ($N$) for detail. For long periods like a year (365.25 days) it is far cheaper than a dummy per day: order 10 uses 20 weights instead of 364.
Where is it used?
Prophet's yearly, weekly and daily seasonalities, your model's $s(t)$, TBATS and dynamic harmonic regression, climate and energy models with yearly cycles, and signal processing (the FFT computes the weights for all $N$ at once).
How is it used?
Choose $P$ from the calendar and $N$ from the detail you need (defaults first, then check on a holdout). For short periods with whole-number $P$, remember the ceiling $\lfloor P/2 \rfloor$. For a sharp, dated spike, do not raise $N$: add a holiday indicator instead.
"A higher Fourier order is always more accurate."
On the true shape, yes. On noisy data, extra harmonics also fit the noise and add parameters, which hurts forecasts (next sections). And beyond $\lfloor P/2 \rfloor$ for sampled data they add nothing at all.
"To catch the Christmas spike, raise the yearly order to 50."
A spike tied to a date is better modelled by a holiday indicator (one weight, exactly placed). A huge yearly order would fit last year's noise all year round and ring around every spike.
"Fourier terms and day-of-week dummies are different models."
For daily data, weekly order 3 and six weekday dummies span exactly the same columns, so least squares gives identical fitted values. Fourier wins for long periods (yearly) and non-integer periods (365.25), where dummies are impossible or wasteful.
$s(t) = \sum_{n=1}^{N}[a_n\cos(2\pi nt/P) + b_n\sin(2\pi nt/P)]$: order $N$, $2N$ weights, averages to 0.
Small $N$ = smooth; large $N$ = detailed. Ceiling for $P$ samples per period: $\lfloor P/2 \rfloor$ (weekly daily: 3, exact = weekday dummies).
Trap: jumps ring (Gibbs, ~9% overshoot); dated spikes belong to holidays, not to a huge $N$.
Quick check: monthly data with yearly seasonality ($P = 12$). What is the largest useful order, and how many non-zero columns does it give?
$\lfloor 12/2 \rfloor = 6$. Harmonics 1 to 5 give 2 columns each (10), and harmonic 6 gives only its cosine ($\cos(\pi t) = \pm1$; its sine is all zeros). Total 11 columns = the 11 free monthly effects, the same as 11 month dummies.
Fourier terms are just regression columns core
We do not know the weights $a_n, b_n$ in advance. We learn them from data, and the trick is that this is ordinary regression. For each day, compute the sines and cosines; put them in columns next to each other; the seasonal part of the model is "these columns times their weights". The model never has to learn a wave: it only learns how much of each fixed wave to use.
Two practical facts make this pleasant. Over whole periods the columns are orthogonal (they do not overlap at all), so each weight is learned almost independently of the others. And the period $P$ is not learned: you choose it. If you choose it wrong, no amount of fitting can repair the shape.
Three ways to say it:
- Picture: a table with one row per day and $2N$ wavy columns; the model mixes them.
- Numbers: weekly order 2 for 8 weeks of daily data = a 56 × 4 block of numbers; $X^\top X = 28\,I$ (perfectly orthogonal).
- Slogan: fixed waves in, learned weights out.
One row of the seasonal block. Weekly period $P = 7$, order $N = 2$, day $t = 1$. Columns in Prophet's order: $\sin_1, \cos_1, \sin_2, \cos_2$.
- $\sin(2\pi \cdot 1/7) \approx 0.782$, $\cos(2\pi/7) \approx 0.623$, $\sin(4\pi/7) \approx 0.975$, $\cos(4\pi/7) \approx -0.223$. Row: $(0.782, 0.623, 0.975, -0.223)$.
- Weights in the same order: $\beta = (6, 4, 3, 0)$.
- Seasonal effect on day 1: $6(0.782) + 4(0.623) + 3(0.975) + 0(-0.223) = 4.691 + 2.494 + 2.925 + 0 \approx 10.11$.
- Stack 56 such rows (8 weeks): a 56 × 4 matrix. Over whole weeks, each column's squares add up to $56/2 = 28$ and every pair of different columns has dot product 0, so $X^\top X = 28 I$.
- Consequence: each least-squares weight is (column · data)/28, and with noise sd $\sigma$ each weight has standard error $\sigma/\sqrt{28}$ (for $\sigma = 15$: about 2.8).
For a seasonality with period $P$ and order $N$, the seasonal design block is the $T \times 2N$ matrix whose row for time $t$ is
$$x_s(t) = \left[\sin\tfrac{2\pi t}{P},\ \cos\tfrac{2\pi t}{P},\ \sin\tfrac{4\pi t}{P},\ \cos\tfrac{4\pi t}{P},\ \dots,\ \sin\tfrac{2\pi N t}{P},\ \cos\tfrac{2\pi N t}{P}\right],$$and the seasonal component is $s(t) = x_s(t)^\top\beta_s$ with $2N$ weights $\beta_s$ (the $b_n$ and $a_n$ of the Fourier series, in sin, cos order).
- The model stays linear in the weights, so least squares, ridge, Normal priors and SVI all treat these columns like any other (Chapter 5.13).
- Over whole periods of equally spaced data the columns are orthogonal: $X_s^\top X_s = \tfrac{T}{2}I$ (for harmonics $n \lt P/2$). With partial periods, gaps or a trend column, they become slightly correlated.
- $P$ and $N$ are fixed before fitting; only $\beta_s$ is learned. A wrong $P$ makes the pattern drift out of step.
- Moving the time origin rotates each $(a_n, b_n)$ pair but leaves the fitted $s(t)$ unchanged.
- Prophet's priors: every seasonal weight gets $N(0, \tau^2)$ with $\tau$ =
seasonality_prior_scale(default 10, on its internally rescaled $y$), the same $\tau$ for all harmonics of a seasonality.
Why do we need it?
Turning seasonality into columns means one fitting machine handles everything (trend, season, holidays, regressors), standard errors and priors work as usual, and the number of seasonal parameters is just the number of columns.
Where is it used?
prophet.Prophet (fourier_series), your NumPyro model's seasonal block (jnp.dot(X_season, beta)), statsmodels' DeterministicProcess with CalendarFourier, R's forecast::fourier with auto.arima(xreg=…), and LA.stats.ts.fourier in this guide.
How is it used?
Build the block once for training and future dates with the same $t$ origin and the same $P$, $N$; concatenate with the other blocks; fit. To plot the seasonal shape, multiply one period of rows by the fitted weights.
"The model will learn the right period if I give it roughly the right one."
$P$ is fixed; only the weights are learned. With $P = 7.5$ instead of 7, after 8 weeks the waves are 4 days out of step. Periods come from the calendar, not from the data.
"I built the Fourier columns separately for training and forecast dates, each starting at $t = 0$."
The forecast rows must continue the same clock: day 400 must use $t = 400$, not $t = 0$. Restarting $t$ shifts the phase and moves the weekly peak to the wrong day. Prophet avoids this by using days since 1970-01-01 for every row.
In your forecasting model the seasonal part is a block of Fourier columns per seasonality, multiplied by a weight vector that has a prior (typically Normal; check the scale your code uses and on which scale of $y$). Three things worth confirming in your code: the column order matches whatever you use to plot or name the weights (Prophet's convention is sin, cos per harmonic); future rows use the same time origin as training rows; and the period is in the units of that time index. These are the classic silent bugs because the model still fits, just worse.
Seasonal block: row $= [\sin\frac{2\pi t}{P}, \cos\frac{2\pi t}{P}, \dots, \sin\frac{2\pi Nt}{P}, \cos\frac{2\pi Nt}{P}]$; $s(t) = x_s(t)^\top\beta_s$, $2N$ weights.
Whole periods: columns orthogonal, $X^\top X = \frac{T}{2}I$. $P$, $N$ fixed in advance; weights learned (least squares or with priors).
Trap: same time origin and units for training and future rows; a wrong $P$ cannot be fixed by fitting.
Quick check: you fit weekly order 3 to 70 days of daily data with noise sd 14. Roughly what is the standard error of each Fourier weight?
70 days = 10 whole weeks, so each column's squares sum to $70/2 = 35$ and the columns are orthogonal. SE $= \sigma/\sqrt{35} = 14/5.92 \approx 2.37$ for every weight (if the other columns, such as the trend, are close to orthogonal to them).
Fourier order is a bias–variance knob core
Think of drawing last year's sales pattern with a thick marker or a fine pen. The thick marker (small $N$) gives a smooth, simple shape. It cannot draw the narrow pre-Christmas rush: it is biased, always wrong in the same places. The fine pen (large $N$) can draw everything, including last year's random wiggles, which will not come back next year. It changes a lot from one year of data to another: it has high variance.
Training error always falls as you add harmonics, because a bigger model can always copy the training data at least as well. So training error cannot choose $N$. Only data the model has not seen, from later in time, tells you when extra detail stops being real.
Three ways to say it:
- Picture: holdout error against $N$ is a U: too smooth on the left, too wiggly on the right.
- Numbers: in the simulation below (2 years of weekly data), training RMSE falls from 25.8 to 8.3 as $N$ goes from 1 to 25, while holdout RMSE falls from 28.6 to about 15.8 and then climbs back to 17.9.
- Slogan: $N$ buys detail with variance; the holdout decides the price.
A simulation with a known answer (the same code is in the Code-it block). Weekly sales for 3 years: level 1000 rising 2 per week, a smooth yearly wave plus a 2-week Christmas rush, noise sd 15. Train on the first 104 weeks, hold out the third year. Yearly period in weeks: $P = 365.25/7 \approx 52.18$. Fit intercept + trend + yearly order $N$ by least squares.
- $N = 1$ (2 seasonal weights): train RMSE 25.8, holdout 28.6. Too smooth: the Christmas rush is missed every year (bias).
- $N = 4$ (8 weights): train 18.7, holdout 23.3. Better, still blunt.
- $N = 10$ (20 weights): train 11.5, holdout 16.1. $N = 12$: train 10.9, holdout 15.8 (the best of the orders tried; the noise alone gives 15).
- $N = 20$ (40 weights): train 9.1, holdout 17.3. $N = 25$ (50 weights for 104 points): train 8.3, holdout 17.9.
- Training error fell at every step (25.8 → 8.3), but the holdout error turned around after about $N = 12$: the extra weights were fitting noise.
Let $\hat s_N$ be the seasonal shape fitted with order $N$.
- Training error never increases with $N$: the order-$N$ columns are contained in the order-$(N + 1)$ columns (nested models), so least squares can always do at least as well.
- Expected holdout error $\approx$ noise $+$ bias$^2$ $+$ variance (Chapter 5.1): bias falls as $N$ grows (a richer shape), variance grows roughly with the number of weights $2N$ divided by the amount of data.
- So holdout error is typically U-shaped in $N$; the best $N$ grows with the length of the history and shrinks with the noise level.
- Choosing $N$: start from defaults, compare a few orders by time-based holdout or rolling-origin evaluation (Chapter 7.15), and check the residuals for leftover seasonal structure (Chapter 7.17). A prior on the weights is a second, softer knob (below).
Why do we need it?
The order is the single setting that decides whether the seasonal shape is too crude or too noisy. Getting it wrong either leaves a pattern in the residuals (too small) or makes forecasts repeat last year's accidents (too large).
Where is it used?
Prophet's fourier_order per seasonality (and yearly_seasonality=20-style overrides), the $K$ chosen by AICc in R's dynamic harmonic regression, and the same knob as polynomial degree, tree depth and number of changepoints (Chapter 7.18).
How is it used?
Fit a few orders (for example yearly 5, 10, 15), score each on later data with rolling origins, pick the smallest order close to the best score, and confirm with a component plot that the seasonal shape looks plausible (no wiggles that no business would believe).
An order is a hard cut: harmonics up to $N$ are in, the rest are out. A prior on the weights is a soft alternative. Prophet puts the same Normal prior $N(0, \tau^2)$ on every Fourier weight of a seasonality. The widget below shows what that kind of prior can and cannot do, and compares it with a "smoothness" prior whose scale shrinks for higher harmonics ($\tau_n = \tau/n$), which acts like a soft order.
"The order with the lowest training error is the best."
Training error always prefers the largest order. Choose on later data (holdout or rolling origin), never on the data used for fitting, and never on a random split (Chapter 7.1).
"Prophet's default yearly order of 10 is the correct order."
It is a sensible default for several years of daily data, not a law. Short or noisy histories may want less; long histories with sharp seasonal features may support more.
"seasonality_prior_scale and fourier_order are the same knob."
The order decides which harmonics exist (how wiggly the shape can be). An equal-scale prior mostly decides how big the seasonal weights may be. They interact, but a small prior scale shrinks the real pattern too.
Fourier order is one of the bias–variance knobs of your forecasting model, next to the number of candidate changepoints, the Laplace scale $b$, and the guide's rank (Chapter 7.18). If you were asked "how did you choose the Fourier orders?", a strong answer is: "I started from Prophet-like defaults, compared a few orders on rolling-origin backtests, and checked the residual ACF at the seasonal lags (7 for weekly) for leftover structure." Only claim the steps you actually did; if you used defaults, say so and explain how you would validate them.
"A higher Fourier order makes the model more accurate."
It makes the seasonal shape more flexible: lower bias, higher variance, more parameters, and more room to confuse seasonality with holidays or trend.
Model answer: "Fourier order trades bias for variance. Low order underfits sharp seasonal features; high order fits noise and repeats it every year. Training error always improves with order, so I choose it by time-based validation, and I keep dated spikes in holiday terms rather than raising the order."
Training error ↓ always as $N$ ↑ (nested models). Holdout error: U-shaped (bias ↓, variance ↑ with $2N$/data).
Choose $N$ by rolling-origin validation + residual checks; best $N$ grows with history length, falls with noise.
Prior $N(0, \tau^2)$ on weights = ridge $\sigma^2/\tau^2$; equal τ mainly limits size, τ/n acts like a soft order.
Quick check: you double the length of the history and keep the noise level. Should the best Fourier order go up, down, or stay the same, and why?
Usually up (or at least not down). With twice the data, each weight is estimated with about $1/\sqrt2$ of the standard error, so the variance cost of each extra harmonic falls, while the bias benefit of extra detail stays the same. The bottom of the U moves to the right. You can see this in the order widget by switching the history from 1.5 to 3 years.
Parameters and identifiability: you cannot learn a year from three months core
Look at a very big circle up close and its edge looks like a straight line. A yearly wave is the same: in a 3-month window, a quarter of a yearly sine looks just like a line going up. If your history is that short, the model cannot tell "demand is growing" (trend) from "we are in the rising part of the yearly season" (seasonality). Both stories fit the data you have almost perfectly, and they forecast completely different futures.
This is an identifiability problem: the data do not pin down how to split the credit between components (Chapter 6.8). It is made worse by a large order, because every extra harmonic is 2 more weights that the same short window must explain. In a Bayesian fit the symptom is a long, thin posterior ridge (trend up and season down, or trend down and season up) and very wide forecast intervals; in least squares it is huge standard errors and wild coefficients.
Three ways to say it:
- Picture: a short piece of a long wave is indistinguishable from a straight line.
- Numbers: over 90 days, a line rising 0.2 per day and a yearly wave (plus a level) agree to within 0.35 everywhere; at day 270 one forecasts 54 and the other 0.6.
- Slogan: to learn a cycle you must see it, ideally twice.
Two stories, one short history. Daily data for days 0 to 89 that look like a straight line $y = 0.2t$ (from 0 up to 17.8).
- Story A, trend: $y = 0.2t$. It fits exactly. Forecast for day 180: $0.2 \times 180 = 36$; day 270: 54.
- Story B, a yearly wave and no trend: least squares of $0.2t$ on $\{1, \sin\frac{2\pi t}{365.25}, \cos\frac{2\pi t}{365.25}\}$ over the 90 days gives $8.90 + 8.90\sin(\cdot) - 8.55\cos(\cdot)$, an amplitude of 12.3.
- Over days 0–89, story B differs from story A by at most 0.35 (typical difference 0.14). With any realistic noise you cannot tell them apart.
- Story B's forecasts: day 180: 17.8 (the wave has flattened), day 270: 0.6 (it has come back down). Story A says 36 and 54.
- In the model with both a trend column and yearly order 1, the columns $t$, $\sin$, $\cos$ are nearly copies of each other over 90 days: correlation of $t$ with the sine column 0.98, with the cosine column −0.98, and variance inflation factors (VIF) about 1 390, 400 and 340. The trend's standard error is $\sqrt{1390} \approx 37$ times larger than it would be without the yearly columns.
- With 365 days of data the VIFs drop to about 2.6, 2.6 and 1.0; with 730 days, about 1.2, 1.2 and 1.0. A full cycle (better two) makes the components separable.
- Parameter count. A seasonality of order $N$ adds $2N$ weights. Several seasonalities add up: yearly order 10 + weekly order 3 = 26 seasonal weights, before trend, changepoints, holidays and regressors.
- Identifiability (plain words): different parameter values must give noticeably different predictions on the observed data. If they do not, the data cannot choose between them. Weak identification = almost the same predictions: very wide SEs, a ridge-shaped posterior, and results driven by the prior (Chapter 6.8).
- Fourier columns of a period $P$ are weakly identified against the trend when the history is shorter than about one period; against changepoints near the start of a seasonal swing; against holidays when the order is high enough to draw a holiday-sized bump; and against regressors that follow the season (temperature follows the year).
- Remedies: drop or lower the order of a seasonality you have not seen at least once (Prophet only switches yearly seasonality on automatically with at least two years of history); use priors on the seasonal weights (and say they are doing the work); borrow the shape from a related series (a hierarchical or informative prior); fix the shape from outside knowledge; or collect more history.
Why do we need it?
A model can fit a short history beautifully and still forecast nonsense, because the data never forced it to choose between "trend" and "season". Knowing when this happens tells you which components to switch off or regularise before the forecast embarrasses you.
Where is it used?
Prophet's automatic seasonality rules (yearly needs about two years), new products and new stores with little history, cold-start forecasting, and the posterior correlations you inspect after fitting your NumPyro model with SVI or NUTS.
How is it used?
Before fitting, compare history length with each period. After fitting, look at posterior (or least-squares) correlations between trend and seasonal weights, and at the spread of forecasts across refits or posterior draws. If they are huge, simplify or add prior information, and report it.
"The model has 400 data points and only 8 parameters, so everything is well determined."
Counting is not enough. If 400 days cover only part of a cycle, the yearly columns and the trend are nearly the same column. What matters is whether the data separate the components, which the VIFs and posterior correlations reveal.
"Adding a prior fixed the identifiability problem."
A prior makes the fit stable, but then the prior, not the data, decides the split between trend and season. That is fine if the prior is defensible, and you should say so ("with under a year of data, the yearly shape comes from the prior"), but it is not evidence.
In your forecasting model, trend changepoints, yearly Fourier terms, holidays and regressors can all compete for the same bump; this is the syllabus's "which component gets the credit?" question (Chapter 6.8). Concretely: with less than about two years of history, consider switching yearly seasonality off or giving it a tight prior; check the posterior correlation between the trend slope (or the $\delta_j$ near the start of the series) and the yearly weights; and remember that every Fourier weight is one more latent dimension in your SVI guide (2 per harmonic), which also feeds the full-rank versus low-rank decision (Chapter 6.13).
Order $N$ = $2N$ weights per seasonality; parameters add across seasonalities.
Less than one period of history ⇒ season ≈ trend (huge VIF, posterior ridge, wild forecasts). Rule of thumb: see the cycle at least once, ideally twice (Prophet's yearly default needs two years).
Fixes: drop/lower the order, priors (say they decide), outside information, more data.
Quick check: a new store has 5 months of daily sales. Which seasonalities would you include, and why?
Weekly, yes: 5 months contain about 21 full weeks, so the weekly shape is well identified (order 3 is exact for daily data). Yearly, no (or only with a strong prior borrowed from similar stores): 5 months is less than half a cycle, so a yearly wave would be confused with the trend and its forecasts for the unseen months would be guesses.
Weekly, yearly and multiple seasonalities core
Daily demand has a fast weekly rhythm riding on top of a slow yearly wave. Hourly website traffic has a daily rhythm (quiet at night), a weekly rhythm (quiet on Sundays) and a yearly one. Each rhythm has its own period, so each gets its own Fourier block with its own period $P$ and its own order $N$. The seasonal component is simply their sum.
The units matter. If $t$ counts days, the periods are 1, 7 and 365.25. If $t$ counts hours, the same three seasonalities have periods 24, 168 and 8 766.
Three ways to say it:
- Picture: a small fast wave riding on a big slow wave.
- Numbers: yearly order 10 (20 columns) + weekly order 3 (6 columns) = 26 seasonal weights for daily data.
- Slogan: one block per rhythm, all added together.
Counting columns in two set-ups, using Prophet's default orders (yearly 10, weekly 3, daily 4).
- Daily data, 3 years. Yearly ($P = 365.25$ days): $2 \times 10 = 20$ columns. Weekly ($P = 7$): $2 \times 3 = 6$. Daily: not possible with one value per day (aliasing section). Total $20 + 6 = 26$ seasonal weights.
- Hourly data, 2 years ($t$ in hours). Yearly: $P = 365.25 \times 24 = 8\,766$ hours, 20 columns. Weekly: $P = 7 \times 24 = 168$ hours, 6 columns. Daily: $P = 24$ hours, $2 \times 4 = 8$ columns. Total $20 + 6 + 8 = 34$.
- One day of the seasonal component: $s(t) = s_{\text{yearly}}(t) + s_{\text{weekly}}(t)$. For example, a December Saturday might be $+60$ (yearly, pre-Christmas) $+ 25$ (weekly, Saturday) $= +85$ above the trend.
- Additive blocks cannot make the Saturday effect bigger in December than in July: the weekly block adds the same $+25$ all year. If the weekly swing grows with the yearly level, use multiplicative seasonality (Chapter 7.7) or a conditional seasonality.
With seasonalities $k = 1, \dots, K$ of periods $P_k$ and orders $N_k$,
$$s(t) = \sum_{k=1}^{K} \sum_{n=1}^{N_k}\left[a_{kn}\cos\frac{2\pi n t}{P_k} + b_{kn}\sin\frac{2\pi n t}{P_k}\right],$$i.e. the Fourier blocks are concatenated side by side, with $2\sum_k N_k$ weights in total.
- Prophet's documented defaults (for $t$ in days): yearly $P = 365.25$, order 10; weekly $P = 7$, order 3; daily $P = 1$, order 4. With the default setting
'auto', each is switched on only when the data can support it: yearly with at least two years of history, weekly with at least two weeks of data spaced less than a week apart, daily only for sub-daily data. Every one usesseasonality_prior_scale(default 10) unless you override it. - Custom seasonalities: any period you can justify, for example Prophet's documentation example of a monthly seasonality with period 30.5 days and order 5 (
add_seasonality(name='monthly', period=30.5, fourier_order=5)). - Conditional seasonalities: a block multiplied by a 0/1 condition column (a weekly shape that applies only in the sports season, say). Prophet supports this with
condition_name. - Keep the blocks' frequencies apart: a yearly harmonic $n$ has frequency $n/365.25$ per day, so yearly order 52 would sit almost on top of weekly harmonic 1 ($52/365.25 \approx 0.1424$ vs $1/7 \approx 0.1429$) and the two blocks would fight. Normal yearly orders (10–20) are far below that.
Why do we need it?
Real series have several calendar rhythms at once. Modelling only one leaves the others in the residuals (and in the forecast errors). Separate blocks also let you plot and explain each rhythm on its own.
Where is it used?
Prophet's yearly + weekly (+ daily) defaults, your model's seasonal terms, electricity load (daily, weekly, yearly), call centres, ride-hailing and web traffic, TBATS and MSTL in the classical toolbox.
How is it used?
List the calendar rhythms; express each period in the units of $t$; give each block an order (defaults first); check the history covers each period at least twice; plot each fitted block over one period; look for leftover rhythm in the residual ACF at each seasonal lag.
"Daily data: switch on Prophet's daily seasonality too, just in case."
With one value per day, the daily-period columns are constant (cosine 1, sine 0) and say nothing. Daily seasonality needs several observations per day (next section).
"Hourly data: keep the weekly period at 7."
The period is in the units of $t$: 168 for hours. A weekly period of 7 on an hourly clock is a 7-hour cycle.
"Weekly and yearly blocks can be fitted separately, one after the other."
Fit them jointly (one design matrix), so neither absorbs the other's pattern. Sequential fitting works only when the blocks are orthogonal over the data, which partial years break.
For daily demand, the natural set-up in your model is a weekly block and, if you have about two or more years of history, a yearly block, each a set of Fourier columns with its own prior. If your data were ever hourly, a daily block becomes possible and the periods change units. In an interview, being able to say "weekly order 3 is exact for daily data, so it is equivalent to day-of-week dummies, and yearly order 10 gives 20 columns that can draw features about two to three weeks wide" shows you understand the defaults rather than just using them.
$s(t) = \sum_k$ (Fourier block with period $P_k$, order $N_k$); total $2\sum N_k$ weights. Periods in the units of $t$.
Prophet defaults: yearly 365.25 d, order 10; weekly 7 d, order 3; daily 1 d, order 4 (auto-enabled only when the data allow).
Trap: additive blocks give the same weekly swing all year; use multiplicative or conditional seasonality if it changes.
Quick check: half-hourly electricity data. Write the periods of the daily, weekly and yearly seasonalities in the units of $t$ (one step = 30 minutes).
Daily: 48 steps. Weekly: $7 \times 48 = 336$ steps. Yearly: $365.25 \times 48 = 17\,532$ steps. With Prophet you pass timestamps, so it handles units for you; in a hand-built model you must convert yourself.
Aliasing and the Nyquist frequency: what daily data can never see
In films, a fast-spinning wagon wheel sometimes seems to turn slowly backwards. The camera takes a picture only 24 times a second, and between pictures the wheel turns almost a full circle; you see only the small leftover movement. A fast motion, looked at too rarely, disguises itself as a slow one. That disguise is called aliasing.
Daily data are pictures taken once a day. A pattern that repeats every 0.9 days moves 0.9 of a cycle per day, which looks exactly like moving 0.1 of a cycle the other way: a slow 10-day wave. The fastest pattern daily data can show honestly is one that flips every single day (period 2 days). That limit, half the sampling rate, is the Nyquist frequency.
For Fourier seasonality this explains two facts: weekly order 4 adds nothing for daily data (harmonic 4 is harmonic 3 in disguise), and a daily seasonality cannot be fitted from daily data at all (it disguises itself as a constant).
Three ways to say it:
- Picture: dots taken once a day can be joined by a slow wave or by a fast one; you cannot tell which.
- Numbers: daily sampling: Nyquist = 0.5 cycle per day (period 2 days). Weekly harmonic 4 has frequency $4/7 \approx 0.571 \gt 0.5$ and folds back to $3/7$.
- Slogan: to see a rhythm you must look at least twice per cycle.
- A fast wave looking slow. $\cos(2\pi \cdot 0.9\,t)$ at $t = 0, 1, \dots, 6$: $1, 0.809, 0.309, -0.309, -0.809, -1, -0.809$. The slow wave $\cos(2\pi \cdot 0.1\,t)$ gives exactly the same seven numbers, because $0.9 = 1 - 0.1$ and a whole cycle per day is invisible.
- Weekly harmonic 4 vs 3. $\cos(2\pi \cdot 4t/7)$ on days 0–6: $1, -0.901, 0.623, -0.223, -0.223, 0.623, -0.901$, identical to $\cos(2\pi \cdot 3t/7)$. And $\sin(2\pi\cdot 4t/7) = 0, -0.434, 0.782, -0.975, 0.975, -0.782, 0.434$ is exactly minus $\sin(2\pi\cdot 3t/7)$.
- So with weekly order 4 the design has 8 Fourier columns but only 6 different directions. Together with the intercept: 9 columns, rank 7. $X^\top X$ is singular; least squares has no unique answer, and in a Bayesian model the posterior for those weights is a ridge held up only by the prior.
- Daily seasonality on daily data. $P = 1$ day: $\cos(2\pi t) = 1$ and $\sin(2\pi t) = 0$ for every whole $t$. The cosine column is a copy of the intercept and the sine column is all zeros.
- Monthly data, yearly period 12. Nyquist allows harmonics up to $n = 6$; harmonic 6 has $\sin(\pi t) = 0$ for all whole $t$. Intercept + order 6 = 13 columns, rank 12.
Observe a series every $\Delta$ time units (sampling rate $f_s = 1/\Delta$). The Nyquist frequency is $f_N = f_s/2$.
- Aliasing: at the sample times $t = k\Delta$, a wave of frequency $f$ is indistinguishable from waves of frequency $|f - m f_s|$ for any whole number $m$. Every frequency therefore has a twin (its alias) in the range $[0, f_N]$. Cosines match exactly; sines match up to a sign flip.
- For a Fourier seasonality with period $P$ (in units of $\Delta$), harmonic $n$ has frequency $n/P$. The samples can tell it apart from slower waves only if $n/P \lt 1/2$ (at exactly $n/P = 1/2$ only its cosine survives). So the useful order is at most $\lfloor P/2 \rfloor$; for whole-number $P$, higher harmonics exactly duplicate lower ones.
- Daily data ($\Delta$ = 1 day, $f_N$ = 0.5 cycle per day): weekly order ≤ 3; daily seasonality impossible; yearly order ≤ 182 (never a constraint in practice). Hourly data: daily order ≤ 12, so Prophet's default 4 is fine.
- Snapshots vs totals: aliasing is a problem for point readings (a temperature read at 9:00 each day). Daily totals sum over the whole day, which cancels an exact daily cycle instead of disguising it: the intra-day shape is simply gone from the data.
Why do we need it?
It tells you the maximum seasonal detail your data frequency can support, explains singular design matrices when someone sets weekly order 4 on daily data, and warns you that point readings taken at a fixed time of day can show fake slow cycles.
Where is it used?
Choosing Fourier orders and seasonalities for daily, weekly or monthly data; Prophet's rule that daily seasonality needs sub-daily data; sensor and IoT sampling rates; audio (44.1 kHz sampling to capture sounds up to about 22 kHz); and anti-aliasing filters in cameras.
How is it used?
For each seasonality, check $N \le \lfloor P/2 \rfloor$ in sampling units. If a pattern faster than two samples per cycle matters, collect data more often. Before using point readings, ask what happens between the samples.
"Weekly order 4 on daily data gives a slightly more detailed weekly shape."
It gives exactly the same shape as order 3, plus two redundant columns that make $X^\top X$ singular (or leave the extra weights held up only by the prior in a Bayesian fit, which wastes guide dimensions and can slow SVI).
"Aliasing means daily totals contain fake slow cycles from the hourly pattern."
Totals over whole days cancel an exact daily cycle. Fake slow cycles come from point readings taken once per period, or from a cycle whose period is not an exact number of samples.
"We set daily seasonality on to be safe; it doesn't hurt."
On daily data it cannot be estimated: the once-a-day cycle aliases to a constant. Prophet itself turns it on automatically only for sub-daily data.
Model answer: "The highest frequency daily data can represent is the Nyquist frequency, half a cycle per day. Weekly harmonic $n$ has frequency $n/7$, so only orders up to 3 are identifiable; harmonic 4 aliases to harmonic 3. A daily cycle aliases to frequency zero, so daily seasonality needs sub-daily data."
Nyquist $f_N = f_s/2$ (daily data: 0.5 cycle/day, period 2 days). Frequencies above it fold: $f \to |f - m f_s|$.
Useful order ≤ $\lfloor P/2 \rfloor$ in sampling units: weekly on daily data ≤ 3; daily seasonality needs sub-daily data.
Trap: extra aliased columns are redundant (singular $X^\top X$), not extra detail.
Quick check: weekly data (one value per week) and a "monthly" pattern with period about 4.35 weeks. What is the largest useful order?
$\lfloor 4.35/2 \rfloor = 2$. Harmonic 2 has frequency $2/4.35 \approx 0.46$ cycles per week, just under the Nyquist 0.5. Harmonic 3 has frequency $3/4.35 \approx 0.69 \gt 0.5$, so in weekly samples it looks like a slower wave of frequency $|0.69 - 1| \approx 0.31$ cycles per week (period about 3.2 weeks). Because 4.35 is not a whole number its columns are not exact copies of the others, but they no longer describe finer monthly detail: weekly sampling cannot tell that harmonic from a 3.2-week wiggle. Keep the order at 2 or below.
Recap, cheat sheet and practice
- Waves: $\sin(2\pi t/P)$, $\cos(2\pi t/P)$ repeat every $P$ (period), have frequency $1/P$ and angular frequency $\omega = 2\pi/P$, stay in $[-1, 1]$ and average 0 over a period. Arguments in radians; $t$ and $P$ in the same units.
- Harmonics: harmonic $n$ has period $P/n$ and frequency $n/P$; all repeat every $P$, so any sum of them is a centred pattern with period $P$.
- Amplitude and phase: $a\cos\omega t + b\sin\omega t = R\cos(\omega t - \varphi)$ with $R = \sqrt{a^2 + b^2}$, $\varphi = \operatorname{atan2}(b, a)$, peak at $t^* = \varphi/\omega$. Two linear weights per harmonic encode a size and a timing.
- Fourier series of order $N$: $\sum_{n=1}^{N}[a_n\cos(2\pi nt/P) + b_n\sin(2\pi nt/P)]$, $2N$ weights, no constant. Smooth shapes need few harmonics; jumps ring (Gibbs, ~9%). With $P$ samples per period, order $\lfloor P/2\rfloor$ is exact (weekly on daily data: 3 = weekday dummies).
- Columns: the seasonal block has rows $[\sin, \cos]$ per harmonic (Prophet and statsmodels order); the model is linear in the weights; over whole periods $X^\top X = \frac{T}{2}I$. $P$ and $N$ are fixed, only weights are learned; keep the same time origin for future rows.
- Order = bias–variance knob: training error always falls with $N$; holdout error is U-shaped. Choose $N$ by time-based validation and residual checks. A prior $N(0, \tau^2)$ = ridge $\sigma^2/\tau^2$; equal τ mainly limits size; τ/n acts like a soft order.
- Identifiability: with less than one period of history, a seasonal wave and a trend are nearly the same column (VIF ≈ 1 400 for 90 days vs a yearly wave). See the cycle, ideally twice; otherwise drop it, shrink it (and say the prior decides), or borrow information.
- Several seasonalities add as separate blocks. Prophet's documented defaults: yearly (365.25 d) order 10, weekly (7 d) order 3, daily (1 d) order 4, each auto-enabled only when the data allow; prior scale 10.
- Nyquist and aliasing: sampling once per $\Delta$ shows frequencies only up to $1/(2\Delta)$; faster waves fold back. Weekly harmonic 4 = harmonic 3 on daily data; a daily cycle on daily data is a constant.
Cheat sheet
| Idea | Formula / fact | Remember |
|---|---|---|
| Basic waves | $\sin(2\pi t/P)$, $\cos(2\pi t/P)$; $f = 1/P$, $\omega = 2\pi/P$ | radians; same units for $t$ and $P$ |
| Harmonic $n$ | $\sin, \cos(2\pi n t/P)$: period $P/n$ | faster, still repeats every $P$ |
| Amplitude, phase | $R = \sqrt{a^2 + b^2}$, $\varphi = \operatorname{atan2}(b, a)$, $t^* = \varphi P/(2\pi)$ | compute $R$ per posterior draw |
| Fourier series | $\sum_{n=1}^{N}[a_n\cos + b_n\sin]$ | $2N$ weights; averages to 0 |
| Exact ceiling | $N \le \lfloor P/2 \rfloor$ ($P$ samples per period) | weekly daily: 3; monthly yearly: 6 (11 columns) |
| Seasonal block | row $= [\sin_1, \cos_1, \dots, \sin_N, \cos_N]$ | whole periods: $X^\top X = \tfrac{T}{2}I$, SE $= \sigma/\sqrt{T/2}$ |
| Order choice | train error ↓ in $N$; holdout U-shaped | rolling-origin validation, not training fit |
| Prior on weights | $N(0, \tau^2)$ ⇔ ridge $\lambda = \sigma^2/\tau^2$ | Prophet: same τ for all harmonics (default 10) |
| Short history | less than one period ⇒ wave ≈ trend | VIF, posterior correlations; prior decides |
| Prophet defaults | yearly 365.25 d / 10 · weekly 7 d / 3 · daily 1 d / 4 | auto: yearly ≥ 2 years, daily only sub-daily |
| Nyquist | $f_N = f_s/2$; $f \to |f - m f_s|$ | harmonic 4/7 → 3/7; daily on daily → constant |
import numpy as np
import pandas as pd
from statsmodels.tsa.deterministic import Fourier
from statsmodels.stats.outliers_influence import variance_inflation_factor
def fourier(t, period, order):
"""Prophet-style Fourier columns: sin, cos for each harmonic n = 1..order."""
t = np.asarray(t, dtype=float)
return np.column_stack([f(2 * np.pi * n * t / period)
for n in range(1, order + 1) for f in (np.sin, np.cos)])
# 1) One row of the weekly block (P = 7, order 2) and its seasonal effect
row = fourier([1], 7, 2)[0]
print(row.round(3), round(float(row @ [6, 4, 3, 0]), 2)) # [ 0.782 0.623 0.975 -0.223] 10.11
idx = pd.date_range("2024-01-01", periods=3, freq="D") # statsmodels uses the same column order
print(Fourier(period=7, order=2).in_sample(idx).columns.tolist())
# 2) Amplitude and phase of a*cos + b*sin
a, b, P = 3.0, 4.0, 7.0
R, phi = np.hypot(a, b), np.arctan2(b, a)
print(R, round(phi, 3), round(phi * P / (2 * np.pi) % P, 2)) # 5.0 0.927 1.03 (peak day)
# 3) Quarterly effects: order 2 is exact for P = 4 (the sin of harmonic 2 is all zeros)
s4, t4 = np.array([-10, 5, -20, 25.0]), np.arange(4)
X4 = np.column_stack([np.cos(np.pi * t4 / 2), np.sin(np.pi * t4 / 2), np.cos(np.pi * t4)])
print(np.linalg.lstsq(X4, s4, rcond=None)[0].round(6)) # [ 5. -10. -15.]
# 4) Fourier order as a bias-variance knob: weekly data, yearly period in weeks
P = 365.25 / 7
t = np.arange(156.0) # 3 years of weekly sales
d = ((t % P) - 50.5 + P / 2) % P - P / 2 # weeks away from the Christmas week
truth = 60 * np.cos(2 * np.pi * t / P - 3.3) + 120 * np.exp(-0.5 * (d / 1.3) ** 2)
rng = np.random.default_rng(1)
y = 1000 + 2 * t + truth + rng.normal(0, 15, t.size)
train, test = t < 104, t >= 104 # train 2 years, hold out year 3
for N in [1, 4, 10, 12, 20, 25]:
X = np.column_stack([np.ones_like(t), t, fourier(t, P, N)])
beta = np.linalg.lstsq(X[train], y[train], rcond=None)[0]
e = y - X @ beta
rm = lambda m: np.sqrt(np.mean(e[m] ** 2))
print(N, round(rm(train), 1), round(rm(test), 1))
# 1 25.8 28.6 | 4 18.7 23.3 | 10 11.5 16.1 | 12 10.9 15.8 | 20 9.1 17.3 | 25 8.3 17.9
# 5) Short history: trend vs yearly wave (VIFs of the columns t, sin, cos)
for L in [90, 365, 730]:
tt = np.arange(L, dtype=float)
Xs = np.column_stack([np.ones(L), tt, fourier(tt, 365.25, 1)])
print(L, [round(float(variance_inflation_factor(Xs, j)), 1) for j in (1, 2, 3)])
# 90 [1388.1, 395.8, 340.6] | 365 [2.6, 2.6, 1.0] | 730 [1.2, 1.2, 1.0]
# 6) Aliasing: weekly order 4 on daily data adds no new columns; daily seasonality is a constant
td = np.arange(70.0)
for N in [3, 4]:
Xw = np.column_stack([np.ones(70), fourier(td, 7, N)])
print(N, Xw.shape[1], np.linalg.matrix_rank(Xw)) # 3 7 7 | 4 9 7
F4, F3 = fourier(td, 7, 4)[:, 6:], fourier(td, 7, 3)[:, 4:]
print(np.allclose(F4[:, 1], F3[:, 1]), np.allclose(F4[:, 0], -F3[:, 0])) # True True
print(fourier(td, 1.0, 1)[:3].round(6)) # sin = 0, cos = 1 on every whole day
1. In daily data, what is the period of the second harmonic of the weekly seasonality?
2. A weekly harmonic-1 fit gives cosine weight $a = -3$ and sine weight $b = 4$ ($t = 0$ is Monday 00:00). What are the amplitude and the peak time?
3. A colleague sets the weekly Fourier order to 4 on daily data "for extra detail". What actually happens?
4. You fit yearly orders 1 to 25 and record training and holdout errors. Which pattern should you expect?
5. Which are Prophet's documented default Fourier orders?
6. A product launched 4 months ago. You include yearly seasonality (order 10) and a linear trend. What is the main risk?
Practice problems
A. Weekly order 2, Prophet column order, weights $\beta = (6, 4, 3, 0)$. Compute the row for day $t = 3$ and the seasonal effect $s(3)$.
- Angles: harmonic 1: $2\pi \cdot 3/7 = 6\pi/7$; harmonic 2: $12\pi/7$.
- $\sin(6\pi/7) \approx 0.434$, $\cos(6\pi/7) \approx -0.901$, $\sin(12\pi/7) \approx -0.782$, $\cos(12\pi/7) \approx 0.623$. Row: $(0.434, -0.901, -0.782, 0.623)$.
- $s(3) = 6(0.434) + 4(-0.901) + 3(-0.782) + 0(0.623) = 2.603 - 3.604 - 2.346 + 0 \approx -3.35$.
B. You want a weekly wave with amplitude 10 that peaks on Saturday at noon ($t^* = 5.5$, $P = 7$). Which cosine and sine weights produce it?
- Phase: $\varphi = \omega t^* = 2\pi \times 5.5/7 \approx 4.937$ radians (283°, equivalently $-1.346$).
- $a = R\cos\varphi = 10 \times 0.2225 \approx 2.23$; $b = R\sin\varphi = 10 \times (-0.975) \approx -9.75$.
- Check: $\sqrt{2.23^2 + 9.75^2} = \sqrt{4.97 + 95.05} \approx 10$, and $\operatorname{atan2}(-9.75, 2.23) \approx -1.346$, i.e. $t^* = -1.346/0.898 + 7 \approx 5.5$. The "Peak on Saturday" button of the dial widget does this with $R = 5$.
C. Daily data, 3 years. Yearly order 10, weekly order 3, and a custom monthly seasonality with period 30.5 days and order 5. How many seasonal weights, and is any order above its Nyquist limit?
Weights: $2(10 + 3 + 5) = 36$. Limits (daily sampling, so $N \le \lfloor P/2 \rfloor$): yearly $\lfloor 365.25/2 \rfloor = 182 \ge 10$; weekly $3 \le 3$ (at the limit, exact); monthly $\lfloor 30.5/2 \rfloor = 15 \ge 5$. None is above its limit. Separately, make sure 3 years is enough to identify everything: 36 seasonal weights from about 1 096 days is fine, and the yearly cycle is seen three times.
D. Data every 3 hours (8 per day). What is the largest useful daily order, and how many daily-seasonality columns are non-zero at that order?
A day is $P = 8$ samples, so $N \le \lfloor 8/2 \rfloor = 4$. Harmonic 4 has frequency $4/8 = 0.5$, exactly the Nyquist frequency: its cosine alternates $+1, -1$, but its sine is $\sin(\pi t) = 0$ at every sample. So order 4 gives $3 \times 2 + 1 = 7$ non-zero columns, matching the 7 free effects of an 8-slot daily pattern. Prophet's default daily order 4 is the maximum here.
E. (Interview) "How did you choose your Fourier orders, and why not just set them very high?"
"I started from Prophet-like defaults (yearly 10, weekly 3 for daily data). Weekly 3 is already exact for daily data, since harmonic 4 aliases to harmonic 3. For the yearly order I compared a few values with rolling-origin backtests and looked at the residual ACF and the component plot. A very high order lowers training error but fits noise that will not repeat, adds two parameters per harmonic, and can steal credit from holidays and the trend; dated spikes belong in holiday terms instead. With short histories I would lower or drop the yearly order rather than let the prior silently decide." (Adapt to what you actually did.)
F. Why is it a problem to include six weekday dummies and weekly Fourier order 3 in the same daily model?
For daily data both sets span exactly the same 6-dimensional space of centred weekly patterns (order 3 is exact). Together they are perfectly collinear: $X^\top X$ is singular and the split of the weekly effect between dummies and Fourier weights is arbitrary. In a Bayesian model the fit still runs, but the split is decided by the priors, the posterior for those weights is a ridge, and the guide wastes 6 dimensions. Use one or the other.
Holidays, exogenous regressors and leakage
Trend and seasonality only know the date. Real demand also jumps on holidays and responds to prices, promotions and the weather. This chapter adds those two layers of your forecasting model, $h(t)$ and $X_t\beta$: how holidays become 0/1 columns with windows and shrinkage priors, how outside regressors enter, what their coefficients mean, and how they fight each other when they move together. Then the most important practical question of the whole guide: will this information really be available when you forecast? Getting that wrong is called leakage, and it produces beautiful backtests and disappointing forecasts.
- Turn a holiday into an indicator column $h_t = \beta_{holiday}\,I(t \in \text{Holiday})$ and read its coefficient
- Use event windows (days before and after), and tell recurring holidays from one-off events
- Handle overlapping holidays (who gets the credit?) and sparse observations with shrinkage priors; specify a prior scale that means something
- Add exogenous regressors: continuous, categorical (dummies and the dummy trap), lagged, and interactions
- Recognise multicollinearity, standardize regressors (on training data only), interpret coefficients correctly and regularize them with priors
- Ask the key question for every regressor: is $X_{T+h}$ known in advance, or must it be forecast (adding uncertainty)?
- Find and prevent data leakage: future information in features, target leakage, and preprocessing (scaling, feature selection, changepoint detection) fitted on the full series
What we need from earlier chapters: holidays and regressors as layers of a series (Chapter 7.2); forecast origin, horizon and why random splits leak (Chapter 7.1); the additive model and its design matrix (Chapter 7.7, Chapter 5.13); collinearity and the VIF (Chapter 5.13); correlation and correlation matrices (Chapter 4.15); standardization (Chapter 4.18); ridge and lasso as Normal and Laplace priors (Chapter 5.3); shrinkage as a precision-weighted average (Chapter 6.6); identifiability (Chapter 6.8); PELT (Chapter 7.9); Fourier seasonality (Chapter 7.11). Words: an exogenous regressor ("exogenous" = coming from outside the series) is an outside variable used to explain $y$; it is also called a covariate or a feature. An indicator $I(\cdot)$ is a switch: 1 when the condition is true, 0 otherwise.
Holidays as indicator columns core
Think of a light switch that is off every day of the year except Christmas Eve. On Christmas Eve it is on. A holiday term is exactly that switch, plus a number that says how much extra demand the switch adds when it is on. The model learns that number from past Christmas Eves.
"Extra" means extra on top of everything else the model already explains: the trend level on that date, the weekday, the yearly season, the regressors. If Christmas Eve is a Tuesday in December, the holiday coefficient answers: "how much busier was it than a normal December Tuesday?"
Why bother? Without the switch, the holiday spike has nowhere to go. The forecast misses next Christmas Eve, the noise level looks bigger than it is (so every interval is too wide), and nearby trend or seasonal terms bend to chase the spike.
Three ways to say it:
- Picture: a column of zeros with a 1 on each holiday date, multiplied by one learned number.
- Numbers: three Christmas Eves were 60, 50 and 70 orders above what trend and season predicted, so the holiday effect is about 60.
- Slogan: a holiday is a switch with a learned price tag.
Three years of daily orders. On each Christmas Eve, the rest of the model (trend + weekday + yearly season) predicts 180, 186 and 192 orders. The actual orders were 240, 236 and 262.
- Remainders on the holiday: $240 - 180 = 60$, $236 - 186 = 50$, $262 - 192 = 70$.
- Indicator column: $I(t \in \text{Christmas Eve}) = 1$ on those three days, 0 on the other 1 092 days.
- If the other components are already right, least squares for the single coefficient is the average remainder on the "on" days: $\hat\beta_{holiday} = (60 + 50 + 70)/3 = 60$. (In a real fit everything is estimated together, which is better; this is the idea.)
- Holiday term: $h_t = 60 \times I(t \in \text{Christmas Eve})$: 60 on Christmas Eve, 0 on every other day.
- Forecast for next Christmas Eve, when trend and season predict 198: $198 + 60 = 258$. Without the holiday column the forecast would be 198, off by about 60.
For holidays $j = 1, \dots, J$, each with a list of dates $D_j$ (past and future), the holiday component is
$$h(t) = \sum_{j=1}^{J} \beta_j\, I(t \in D_j), \qquad \text{one holiday: } h_t = \beta_{holiday}\,I(t \in \text{Holiday}).$$- $I(t \in D_j)$ is the indicator (dummy) column: 1 on the dates of holiday $j$, 0 elsewhere.
- $\beta_j$ is the expected extra $y$ on holiday $j$ compared with a non-holiday day that has the same trend, seasonality and regressor values. It can be negative (a closed store, a quiet public holiday).
- The columns go into the same design matrix as trend, Fourier and regressor columns and are fitted jointly. In an additive model $\beta_j$ is in units of $y$; in a multiplicative one it is a percentage change (Chapter 7.7).
- The dates come from a calendar, so the column is known for future dates: holidays are always available at forecast time, as long as the future calendar is filled in.
- Prophet: a
holidaystable with columnsholiday,ds(and optionallylower_window,upper_window,prior_scale); every holiday weight gets a Normal prior with scaleholidays_prior_scale(default 10).
Why do we need it?
Holiday spikes are large, predictable and dated. Modelling them gives a good forecast for the next holiday, keeps the noise estimate honest (so intervals are not too wide all year), and stops the trend and seasonal terms from bending to chase spikes.
Where is it used?
Prophet's holidays argument and add_country_holidays, the holidays Python package for national calendars, intervention dummies in ARIMA/SARIMAX, retail, travel and delivery demand models, and the $h(t)$ block of your forecasting model.
How is it used?
Build a table of event names and dates covering the training period and the forecast horizon; turn each event into a 0/1 column; fit it jointly with the rest; check that residuals on holiday dates no longer stand out; keep the calendar updated before each forecast run.
"The holiday coefficient is the number of orders on the holiday."
It is the extra over what the rest of the model predicts for that date. A coefficient of 60 on a day whose baseline is 198 means a forecast of 258.
"Just delete the holidays from the training data."
That protects the other components, but then the model knows nothing about the next holiday and forecasts an ordinary day. Deleting (or masking) is reasonable for one-off events; recurring holidays deserve a column.
"A high yearly Fourier order will pick up Christmas anyway."
A one-day spike is far narrower than what yearly order 10 can draw (about 18 days, Chapter 7.11). The Fourier terms smear it over weeks instead, over-predicting the days around it. And moving holidays (Easter, Black Friday) do not sit at a fixed day of the year at all.
In your forecasting model, $h(t)$ is a block of holiday indicator columns with their own weights and priors. Three things worth knowing about your own set-up: where the holiday dates come from (a library calendar or a hand-made table), whether that table also covers the forecast horizon, and which prior scale the holiday weights get and on which scale of $y$. In an interview, "holidays are 0/1 columns fitted jointly with trend and seasonality, so their coefficients are effects over the baseline for that date" is the one-sentence answer.
$h_t = \sum_j \beta_j I(t \in D_j)$: one 0/1 column per holiday (dates known past and future), fitted jointly.
$\beta_j$ = extra $y$ vs. a normal day with the same trend, season and regressors (can be negative).
Trap: without the column, spikes inflate σ̂, bend nearby terms and the next holiday is missed.
Quick check: a store is closed on 25 December, so sales are 0 that day; the baseline predicts 300. What holiday coefficient should the model learn, and is a Normal likelihood a good idea for that day?
About $-300$ (extra = $0 - 300$). It is a perfectly predictable event, so a holiday column is right. But a Normal likelihood around a mean of exactly 0 can put probability on negative sales; for count data a Negative Binomial likelihood with a log link handles "closed" days more naturally, or you can simply mask closed days and set the forecast to 0 by rule (Chapter 7.13).
Event windows: the days before and after core
Holidays rarely affect just one day. People shop in the days before Christmas, deliveries stop on the day itself, and returns and sales come after it. Black Friday drags in the whole weekend until Cyber Monday. A single switch on one date misses all of that.
An event window gives the holiday a few extra switches: one for "two days before", one for "the day before", one for "the day", one for "the day after", and so on. Each switch gets its own coefficient, so the model can learn the whole shape of the event: a build-up, a peak, a dip.
Three ways to say it:
- Picture: instead of one spike, a little profile of bars around the date.
- Numbers: Christmas with a window from 2 days before to 1 day after: 4 columns (Dec 23, 24, 25, 26) and 4 coefficients.
- Slogan: one switch per day of the event.
An online shop. Average remainders (after trend, weekday and season) over three years around Christmas: Dec 22: +2, Dec 23: +20, Dec 24: +45, Dec 25: −30 (no deliveries), Dec 26: +25 (sales start), Dec 27: +3.
- Dec 22 and Dec 27 are close to 0: no clear effect, so the window starts 2 days before and ends 1 day after the holiday: lower window $-2$, upper window $+1$.
- Columns: one per offset: $I(t = \text{Dec } 23)$, $I(t = \text{Dec } 24)$, $I(t = \text{Dec } 25)$, $I(t = \text{Dec } 26)$, each 1 on that date in every year.
- Coefficients: $+20, +45, -30, +25$. One holiday, four numbers, a shape that a single coefficient could never capture (its best single number would be the average, $+15$, wrong on every day).
- Cost: 4 parameters instead of 1, and each is learned from only 3 days (one per year). That is why windows should be as short as the evidence supports, and why their weights need shrinkage priors (section 5).
A holiday with dates $D$ and a window from $L \le 0$ to $U \ge 0$ days gets one column per offset $k = L, \dots, U$:
$$h(t) = \sum_{k=L}^{U} \beta_k\, I(t - k \in D).$$- $I(t - k \in D) = 1$ when day $t$ is exactly $k$ days after a holiday date (before it if $k \lt 0$).
- $U - L + 1$ coefficients per holiday, each learned from as many days as there are occurrences.
- Shared vs separate: a single column that is 1 on all window days gives one coefficient (fewer parameters, assumes the same effect on every day); separate columns per offset allow a shape. A middle path: separate columns with a prior that pulls them toward each other.
- Prophet:
lower_window(≤ 0) andupper_window(≥ 0) per holiday; it creates one column per holiday per offset, all with that holiday's prior scale.
Why do we need it?
Many events have lead-up and after-effects that are as large as the day itself. Without a window, those days look like noise, and the trend or seasonal terms try to absorb them.
Where is it used?
Prophet's lower_window/upper_window, Black Friday to Cyber Monday windows in retail, pre-holiday travel peaks, post-holiday returns, the days around a product launch, and event studies in economics.
How is it used?
Plot the average remainder for each day offset around past occurrences; start the window where the effect starts and end it where it fades; give each offset a column; refit and check the residuals around the event are flat; prefer short windows when there are few occurrences.
"A wide window is safer: include a week on each side."
Every offset is another parameter estimated from a handful of days. Unneeded offsets fit noise (and can collide with other holidays' windows). Use the shortest window the evidence supports, plus shrinkage.
"The window effect is the same on every day of the window."
Only if you choose a shared column. Separate columns let the day before be +45 and the day itself −30.
If your holiday table includes windows, each window day is one more indicator column and one more latent weight in your SVI guide. With, say, 15 holidays and windows of 4 days, that is 60 holiday weights, more than the seasonal block. This is where shrinkage priors on holiday weights (section 5) and prior sensitivity checks (Chapter 6.8) earn their keep.
Window $[L, U]$: columns $I(t - k \in D)$ for $k = L..U$, one coefficient per offset ($U - L + 1$ per holiday).
Choose it from the average remainder by offset; shortest window that flattens it.
Trap: each offset is learned from only as many days as there are occurrences.
Quick check: Black Friday with lower window 0 and upper window 3 (to Cyber Monday). How many columns, and with 4 years of data how many observations per coefficient?
Offsets 0, 1, 2, 3: 4 columns. Each is 1 on one day per year, so each coefficient rests on 4 observations. Black Friday is always a Friday, so offset 3 is always a Monday; the weekday effect is still estimated from all other Mondays, so the Cyber Monday coefficient is the extra over a normal Monday.
Recurring holidays vs one-off events
Some events come back: Christmas every 25 December, Black Friday every late November, Easter on a different date each year but every year. Others happen once: a website outage, a warehouse strike, a viral video, a lockdown. They need different treatment.
A recurring holiday shares one coefficient (or one set of window coefficients) across all its occurrences. Each year teaches the model a little more, and the coefficient is used again for next year's date. The date list must include future occurrences, and for moving holidays it must come from the real calendar, not from "day 358 of the year".
A one-off event teaches nothing about the future. We still model it, but for a different reason: to protect the other components. An unexplained crash in the last two weeks of data can convince a flexible trend that demand is collapsing. An indicator for the event (or removing those days) says "this was special, ignore it when learning the trend".
Three ways to say it:
- Picture: recurring events have marks in every year and in the future; a one-off event has one mark and nothing ahead.
- Numbers: an outage day 180 orders below normal, with noise sd 10, weighs as much as 324 ordinary days in the sum of squares if left unexplained.
- Slogan: recurring events are for forecasting; one-off events are for protection.
A one-day website outage. The model expected 200 orders; the day had 20. The noise sd on normal days is $\sigma = 10$.
- Residual if unmodelled: $20 - 200 = -180$. In standard units: $-180/10 = -18$.
- Its contribution to the sum of squares (Normal likelihood): $18^2 = 324$, while a typical day contributes about $1^2 = 1$. One bad day pulls on the fit as hard as 324 ordinary days.
- Add a one-off indicator (1 on the outage day only). Its coefficient becomes $-180$, the day's residual becomes 0, and the day no longer influences the trend, seasonality or $\sigma$.
- The indicator is 0 on every future date, so the forecast is not affected: we do not expect another outage.
- Alternatives with the same protective effect: drop (mask) that day from the likelihood, or use a heavy-tailed Student-t likelihood that lets extreme residuals pull less (Chapter 7.13).
- Recurring event: a column that is 1 on every occurrence, past and future, with one shared coefficient (or one per window offset). Moving holidays (Easter, Thanksgiving, Black Friday, Diwali, Lunar New Year, Ramadan) need the actual date of each year from a calendar.
- One-off event: a column that is 1 only on the event's days in the history and 0 for all future dates. Its coefficient absorbs the event so it does not distort trend, seasonality, regressor coefficients or the noise scale. It does not change the forecast directly.
- Alternatives for one-offs: remove (mask) the affected days from the likelihood; a robust likelihood (Student-t); or, if the event changed the level permanently, a genuine changepoint (Chapter 7.8).
- One-off events near the end of the history are the dangerous ones: the trend's last segment and any detected changepoints are learned from very few days.
Why do we need it?
Treating a one-off like a recurring holiday predicts a fake repeat; treating it as normal data lets it bend the trend. Treating a moving holiday as a fixed date puts its effect on the wrong day every year.
Where is it used?
COVID-lockdown indicators in demand models, outage and stock-out flags, intervention analysis in ARIMA, Prophet's advice to treat unusual periods as holidays or remove them, and the holidays package for moving dates.
How is it used?
Keep an incident log next to the holiday calendar. Recurring events: list every past and future date. One-off events: flag their days, add an indicator or mask them, and never give them future 1s. Re-check changepoints near flagged periods.
"Model the 2020 lockdown as a holiday, like Christmas."
A holiday column with future 1s would predict a lockdown next year. A one-off event gets a column only on its own days (or the days are masked).
"Easter is a holiday on day 100 of the year."
Easter moves by up to five weeks. Use each year's actual date from a calendar library; otherwise the coefficient is spread over the wrong days.
This matters doubly in your model because the trend uses changepoints chosen by a grid plus PELT. A one-off dip or spike near the end of the history can be detected by PELT as a mean shift and then treated as a permanent change (Chapter 7.9), and the Laplace prior on $\delta_j$ (Chapter 7.10) only shrinks, it does not know the dip was special. Flagging known incidents before changepoint detection and fitting (indicator or mask) is the cheap fix; say in an interview how you handled such periods, if you did.
Recurring: one shared β, dates listed past and future (moving holidays from a calendar). One-off: column = 1 only on its days, 0 in the future; it protects the other components.
Alternatives for one-offs: mask the days, robust likelihood, or a real changepoint if the change is permanent.
Trap: an unflagged one-off near the end of the history becomes a fake trend change.
Quick check: a 3-day marketing campaign happened once last spring and will not be repeated. Should it get a column, and what does that column contain for future dates?
Yes: a one-off indicator that is 1 on the three campaign days in the history (or a "campaign" regressor), so the uplift does not leak into the trend or the spring seasonality. For future dates it is 0 (no campaign planned). If campaigns will recur, make it a regressor with planned future values instead (section 10).
Overlapping holidays: who gets the credit? core
Your shop's anniversary sale happens to fall on Black Friday. Sales jump by 80. Was that the anniversary or Black Friday? If the two always happen on the same day, the data cannot say. They only ever show the total. Any split (40 + 40, 80 + 0, 100 − 20) explains the data equally well.
If the two sometimes happen on different days, those separate days settle the question. One separate occurrence of each can be enough to pin down both effects.
When the data cannot decide, a Bayesian model still produces an answer, but that answer comes from the prior. With the same Normal prior on both weights, the prior prefers an even split. That is a choice, not a finding. Overlaps also happen inside one calendar: Christmas Eve's window and Christmas Day's window can cover the same dates, or a public holiday can fall on a weekend.
Three ways to say it:
- Picture: the data draw a line "$\beta_A + \beta_B = 78$"; every point on it fits; only separate days cut the line to a point.
- Numbers: two overlapping years (+80, +76) and one separate year (A alone +50, B alone +30) give $\hat\beta_A = 49.2$, $\hat\beta_B = 29.2$.
- Slogan: if two switches are always flipped together, you can only learn what they do together.
Holiday A (Black Friday) and event B (anniversary sale). Remainders after the rest of the model: years 1 and 2, same day: $+80$, $+76$. Year 3, different days: A alone $+50$, B alone $+30$.
- Only years 1–2. Each overlapping day says $\beta_A + \beta_B \approx 78$ (the average). Nothing separates $\beta_A$ from $\beta_B$: the least-squares solution is a whole line, and $X^\top X$ (built from rows $(1, 1)$) is singular.
- Add year 3. Rows: $(1,1)$ with 80, $(1,1)$ with 76, $(1,0)$ with 50, $(0,1)$ with 30. Least squares minimises $(80 - A - B)^2 + (76 - A - B)^2 + (50 - A)^2 + (30 - B)^2$.
- Setting the derivatives to zero: $3A + 2B = 206$ and $2A + 3B = 186$.
- Subtract to get $A - B = 20$; add to get $5A + 5B = 392$, so $A + B = 78.4$. Hence $\hat\beta_A = 49.2$, $\hat\beta_B = 29.2$.
- Bayesian, overlap only (3 overlapping days averaging 78, noise sd 8, prior $N(0, 30^2)$ on each): posterior means $38.5$ and $38.5$ (an even split chosen by the symmetric prior), each with sd $21.3$ and correlation $-0.98$ between them. The sum is well known; the parts are not.
- Two indicator columns that are 1 on exactly the same days are perfectly collinear: only $\beta_A + \beta_B$ is identified (Chapter 6.8). Columns that overlap on most of their days are nearly collinear: large standard errors and a strongly negative correlation between the two estimates.
- Separate occurrences (days where only one is on) identify the individual effects; the more of them, the better.
- With priors, the posterior is always proper, but in the overlapping direction it equals the prior: the split reflects the prior scales (equal scales ⇒ even split; a smaller scale on one ⇒ the other takes more credit).
- Remedies: merge the two into one combined event when they always coincide; add an explicit "A and B together" column if their joint effect differs from the sum; choose priors deliberately and report that the split is prior-driven; avoid overlapping windows of related holidays when possible.
Why do we need it?
Holiday effects are often reported ("Black Friday adds 18%") and reused (planning next year's anniversary sale on a different day). If the split was decided by the prior, those reports and plans are built on sand.
Where is it used?
Prophet holiday tables with windows that cover the same dates, national plus regional holidays on the same day, promotions scheduled on holidays, marketing-mix models where campaigns coincide with seasonal peaks.
How is it used?
Count, for each pair of holiday columns, the days where both are 1 and the days where only one is. If one of the "only" counts is zero, merge the columns or accept that the split is a prior choice. After fitting, check posterior correlations between holiday weights.
"The model estimated Black Friday at +38 and the anniversary at +38, so both matter equally."
If they always coincided, the data only know the sum (+77); the even split is the prior's symmetry. Report the combined effect, or say clearly that the split is an assumption.
"The posterior is proper and the sampler converged, so the effects are identified."
Priors make any posterior proper. Identification is about whether the data narrow the posterior; a posterior that matches the prior in one direction (correlation near −1 between the two weights) is the warning sign.
In your forecasting model, holiday columns can overlap with each other (windows), with promotion regressors (a sale planned on a holiday), and with high-order yearly Fourier terms (Chapter 7.11). After fitting with SVI, look at the guide's or the posterior's correlations between those weights: a full-rank guide can show such a negative correlation, while a mean-field guide cannot represent it and will report each weight as more certain than it really is (Chapter 6.13).
Always-together columns: only $\beta_A + \beta_B$ is identified; the split comes from the prior (equal priors ⇒ even split).
Days where only one is on identify each effect. Check overlap counts; merge or add a joint column; report prior-driven splits.
Trap: posterior correlation near −1 = not identified, even if everything "converged".
Quick check: with priors $N(0, 30^2)$ on $\beta_A$ and $N(0, 5^2)$ on $\beta_B$, and only overlapping days showing a sum of 80, roughly how will the posterior split the credit?
Most of it goes to A. The tight prior on B says "B is probably small", so the posterior keeps $\beta_B$ near 0 and lets $\beta_A$ take most of the 80. For a known sum the split is proportional to the prior variances: $\beta_A \approx 80 \times 900/(900 + 25) \approx 78$ and $\beta_B \approx 2$ (slightly less in total because the priors also shrink the sum a little). Same data, different priors, different "findings".
Sparse observations and shrinkage priors on holiday effects core
A holiday happens once a year. With three years of data, its coefficient rests on three days. If one of those days was unusually busy for unrelated reasons (good weather, a viral post), the raw estimate is badly off, and the forecast will repeat that accident next year.
A shrinkage prior says: "holiday effects are usually moderate; believe a big effect only if the data insist". It pulls each raw estimate toward zero, more strongly when there are few observations and the noise is large, and hardly at all when the evidence is strong. It is exactly the partial pooling you met for small groups (Chapter 6.6), applied to rare days.
Three ways to say it:
- Picture: an arrow from the raw estimate toward 0, long for rare noisy holidays, short for well-measured ones.
- Numbers: raw +60 from 3 days with noise sd 20 and prior sd 30 → posterior +52; from 1 day → +42.
- Slogan: little evidence, little trust.
A holiday with average remainder $\bar r = 60$ on its days; noise sd per day $\sigma = 20$; prior $\beta \sim N(0, 30^2)$.
- Precisions (precision = 1/variance): data $n/\sigma^2$, prior $1/\tau^2 = 1/900 \approx 0.00111$.
- $n = 3$: data precision $3/400 = 0.0075$. Weight on the data $w = 0.0075/(0.0075 + 0.00111) \approx 0.871$.
- Posterior mean $= w\,\bar r = 0.871 \times 60 \approx 52.3$; posterior sd $= 1/\sqrt{0.00861} \approx 10.8$.
- $n = 1$: data precision $0.0025$, $w = 0.0025/0.00361 \approx 0.692$, posterior mean $\approx 41.5$, sd $\approx 16.6$. One day gets less trust.
- $n = 30$ (a weekly event, say): $w = 0.075/0.0761 \approx 0.985$: almost no shrinkage.
With a shrinkage prior $\beta_j \sim N(0, \tau^2)$ and (for simplicity) the rest of the model known, the posterior of a holiday effect seen on $n_j$ days with average remainder $\bar r_j$ is Normal with
$$E[\beta_j \mid D] = w_j\,\bar r_j, \qquad w_j = \frac{n_j/\sigma^2}{n_j/\sigma^2 + 1/\tau^2}, \qquad sd = \left(\frac{n_j}{\sigma^2} + \frac{1}{\tau^2}\right)^{-1/2}.$$- The MAP estimate equals ridge regression with penalty $\lambda = \sigma^2/\tau^2$ (Chapter 5.3). Smaller τ = stronger shrinkage.
- Shrinkage adds a little bias but removes a lot of variance when $n_j$ is small, so the expected squared error is usually lower (Chapter 5.1). If τ is far smaller than the real effects, the bias wins and true holidays are underestimated.
- Alternatives: a Laplace prior (many holidays near zero, a few large; sparse-ish at the MAP); a hierarchical prior $\beta_j \sim N(\mu_h, \tau_h^2)$ that learns the typical holiday size from all holidays together (Chapter 6.5).
- Prior specification: τ is in the units of the $y$ the model sees. Prophet divides $y$ by its maximum absolute value by default, so
holidays_prior_scale = 10(the default) allows effects of about ten times the series' maximum: very little regularization. Prophet's documentation suggests lowering it to dampen holiday effects that overfit; a per-holidayprior_scaleis also possible.
Why do we need it?
Holiday and window coefficients are the most data-starved parameters of a forecasting model (one observation per year each). Without shrinkage they absorb one-time noise and replay it every year; with it, rare effects are cautious and well-supported ones stay large.
Where is it used?
Prophet's holidays_prior_scale and per-holiday prior_scale, Normal or Laplace priors on holiday weights in NumPyro models, hierarchical holiday effects across stores or products, empirical-Bayes shrinkage in demand planning.
How is it used?
Decide what a plausible holiday effect is in the units the model uses (after any scaling), set τ so that most prior mass is within that range, check with a prior predictive simulation, and run a sensitivity check with τ halved and doubled (Chapter 6.8).
"Shrinkage priors bias the estimates, so they make the model worse."
They add a little bias and remove a lot of variance; for effects seen on 1–3 days the total error is usually smaller. Only a prior that is far too tight for the real effects makes things worse, which a prior sensitivity check reveals.
"Prophet's default holidays_prior_scale = 10 is a strong prior."
On a series divided by its maximum, a scale of 10 permits effects ten times the largest value ever observed: it barely regularizes. Prophet's documentation says to reduce it if holidays overfit.
Your model has the same two problems as the A/B framework: small groups there, rare holidays here, and the same cure (shrinkage). In the A/B framework, hierarchical partial pooling shrinks small segments toward the overall mean; in the forecasting model, the prior on holiday weights shrinks rare holidays toward zero (or toward a learned typical holiday size, if you make it hierarchical). If your code standardizes or rescales $y$, the holiday prior scale lives on that scale: be able to say what a prior sd of, say, 0.1 means in orders. The syllabus lists sparse holiday effects explicitly among the places to run prior sensitivity checks.
"We estimated the Black Friday effect at +23% from our data."
"From three Black Fridays; with a shrinkage prior the posterior mean is +19% with a wide interval, and the estimate moves by a few points if I halve or double the prior scale."
Model answer: "Holiday effects are estimated from very few days, so I use shrinkage priors on them. The posterior is a precision-weighted compromise between the data average and the prior: rare, noisy holidays are pulled toward zero, well-observed ones barely move. I check the prior scale in the units the model sees and test sensitivity."
$\beta_j \sim N(0, \tau^2)$: posterior mean $= w\bar r$, $w = \frac{n/\sigma^2}{n/\sigma^2 + 1/\tau^2}$; MAP = ridge with $\lambda = \sigma^2/\tau^2$.
Few occurrences + noisy days ⇒ strong shrinkage (lower MSE); many ⇒ little. Options: Laplace, hierarchical.
Trap: τ is in the units of the (scaled) $y$; Prophet's default 10 on $y/\max|y|$ is very weak.
Quick check: two holidays both have raw average remainder +50 and noise sd 20. Holiday P was seen 4 times, holiday Q once. With prior sd 25, what are the posterior means?
Prior precision $1/625 = 0.0016$. P: data precision $4/400 = 0.01$, $w = 0.01/0.0116 \approx 0.862$, mean $\approx 43.1$. Q: data precision $0.0025$, $w = 0.0025/0.0041 \approx 0.610$, mean $\approx 30.5$. Same raw number, different trust.
Exogenous regressors: continuous and categorical core
A regressor brings outside information into the model: the temperature, the price, whether a promotion is running. It enters the same way as everything else: a column of numbers and a learned coefficient.
A continuous regressor (temperature, price, ad spend) can take many values; its coefficient is "change in $y$ per one unit of $x$". A categorical regressor (promotion type: none, e-mail, banner) has a few labels with no natural numbers. We give it one 0/1 column per label except one, the reference level; each coefficient is then "how different is this label from the reference". Including a column for every label as well as the intercept repeats the intercept (the columns add up to 1 on every day): that is the dummy-variable trap.
Three ways to say it:
- Picture: a continuous regressor tilts the prediction; a categorical one shifts it up or down by a label-specific step.
- Numbers: banner promo +40 vs no promo, temperature +5 orders per °C: a banner day at +4 °C gets $40 + 20 = 60$ extra.
- Slogan: categories need a reference; coefficients are differences from it.
Promotion type with levels none (reference), e-mail, banner, and temperature anomaly $x_{temp}$ (°C above normal). Fitted: $\beta_{email} = 25$, $\beta_{banner} = 40$, $\beta_{temp} = 5$.
- Columns: $I(\text{email})$, $I(\text{banner})$, $x_{temp}$. No column for none: a "none" day has both dummies 0.
- A banner day at $+4$ °C: row $(0, 1, 4)$; regressor effect $0 \times 25 + 1 \times 40 + 4 \times 5 = 60$ orders above an otherwise identical no-promo, normal-temperature day.
- An e-mail day at $-2$ °C: $25 + (-2)(5) = 15$.
- Change the reference to e-mail: the coefficients become $\beta_{none} = -25$, $\beta_{banner} = +15$. Different numbers, same fitted values: only the meaning ("compared with what?") changed.
- Dummy trap: columns for none, e-mail and banner add up to 1 every day, exactly the intercept column. $X^\top X$ is singular; the level can be moved freely between the intercept and the three dummies.
With regressor row $X_t = (x_{t,1}, \dots, x_{t,p})$ the regressor block is $X_t\beta = \sum_j x_{t,j}\beta_j$.
- Continuous $x_j$: $\beta_j$ = expected change in $y_t$ for a one-unit increase in $x_{t,j}$, holding the other columns fixed; units "units of $y$ per unit of $x_j$". Linear by default; transform $x$ (log, splines, bins) if the effect bends.
- Binary $x_j \in \{0, 1\}$: $\beta_j$ = expected difference when it is on.
- Categorical with $K$ levels: $K - 1$ dummy columns (one-hot encoding minus a reference). Each $\beta$ is the difference from the reference level. With an intercept, never include all $K$ (dummy trap); in a Bayesian model with priors it would "fit", but the level would be split by the prior.
- Coefficients describe associations in the data; they are causal effects only under extra conditions (randomized promotions, no confounding; Chapter 5.12).
- Prophet:
add_regressor(name, prior_scale=None, standardize='auto', mode=None); the regressor must be supplied for the history and for every future date.
Why do we need it?
Trend, seasonality and holidays only know the date. Prices, promotions, weather and campaigns move demand on specific days; without them those movements stay in the noise and the model cannot answer "what if we run a banner promotion?".
Where is it used?
Prophet's add_regressor, SARIMAX's exog, marketing-mix models, price-elasticity models, energy load with temperature, pandas.get_dummies(drop_first=True) and scikit-learn's OneHotEncoder(drop='first').
How is it used?
List candidate drivers; encode categories with a deliberate reference level; keep units in mind (or standardize, section 9); fit jointly; read each coefficient as "per unit, compared with the reference, others fixed"; confirm future values will exist (section 10).
"β_banner = 40 means a banner day sells 40 orders."
It means 40 more than a reference (no-promo) day with the same date effects and temperature.
"One-hot encode every level; the Bayesian model will cope."
It will run, but the level is then shared between the intercept and the dummies according to the priors, wasting a dimension and making coefficients hard to read. Drop a reference level.
$X_t\beta = \sum_j x_{t,j}\beta_j$. Continuous: per unit; binary: on vs off; categorical: $K - 1$ dummies, each vs the reference.
Changing the reference changes coefficients, not fitted values. All $K$ dummies + intercept = dummy trap (singular).
Trap: coefficients are associations, "others held fixed", not automatically causal.
Quick check: weather type has 4 levels (sun, cloud, rain, snow). How many columns with an intercept, and what does $\beta_{snow} = -35$ mean if "sun" is the reference?
3 columns (cloud, rain, snow). $\beta_{snow} = -35$: on snowy days the model expects 35 fewer orders than on sunny days with the same trend, season, holidays and other regressors.
Lagged regressors and interactions
Some causes act with a delay. Money spent on ads today brings orders in two or three days; a heat wave yesterday empties shelves today. A lagged regressor uses an earlier value of $x$: $x_{t-L}$, "the value $L$ days ago". A happy side effect: if the lag is at least the forecast horizon, the value is already known at forecast time.
Some effects depend on each other. Hot weather lifts ice-cream orders more on weekends than on weekdays; a promotion works better in December. An interaction is a column made by multiplying two others, so the effect of one can change with the other.
Three ways to say it:
- Picture: a lag slides the regressor's curve to the right until its bumps line up with the demand bumps; an interaction gives each group its own slope.
- Numbers: orders$_t$ = 100 + 2 × spend$_{t-2}$; a Saturday at +5 °C gets $3 \times 5 + 4 \times 5 = 35$ extra, a Wednesday only 15.
- Slogan: lags move effects in time; interactions let effects depend on context.
- Lag. Model: orders$_t = 100 + 2\,x_{t-2}$ with $x$ = ad spend (in thousands). Forecast origin $T$ (today's data are in).
- Tomorrow ($T + 1$) needs $x_{T-1}$: known. The day after ($T + 2$) needs $x_T$: known. $T + 3$ needs $x_{T+1}$: not known yet, unless tomorrow's spend is planned. With lag $L$, horizons $h \le L$ need no regressor forecast.
- Interaction. Columns: temperature $x$, weekend $w \in \{0, 1\}$, and their product $x \cdot w$. Fitted: $\beta_x = 3$, $\beta_{x \cdot w} = 4$ (and a weekend main effect inside the weekly seasonality).
- Weekday at +5 °C: $3 \times 5 + 4 \times 5 \times 0 = 15$. Saturday at +5 °C: $3 \times 5 + 4 \times 5 \times 1 = 35$. The temperature slope is 3 on weekdays and $3 + 4 = 7$ on weekends.
- Lagged regressor: column $x_{t-L}$ for a lag $L \ge 1$. Several lags ($x_{t-1}, \dots, x_{t-K}$) form a distributed lag; adjacent lags are strongly correlated, so their individual coefficients are wobbly and benefit from priors. Choose $L$ from domain knowledge and validation (a cross-correlation plot of $y$ against lagged $x$ helps).
- Availability rule: for horizon $h$, $x_{T+h-L}$ is known at the origin $T$ when $L \ge h$ (or when $x$ is planned in advance). Lags of the target itself ($y_{t-L}$) follow the same rule and turn the model into an autoregression.
- Interaction: a product column $x_{t,a}\,x_{t,b}$ with its own coefficient. With both main effects in the model, $\beta_{ab}$ is how much the slope of $x_a$ changes per unit of $x_b$. Keep the main effects when you add an interaction, and centre continuous variables first so the main effects stay interpretable.
- Interactions multiply the column count quickly (every pair!); add the few that domain knowledge suggests, with priors.
Why do we need it?
Delayed effects are invisible to a same-day regressor, and context-dependent effects are averaged away by a single slope. Lags and interactions capture both while keeping the model linear in its weights.
Where is it used?
Marketing-mix models (lagged and "adstock" spend), distributed-lag models in epidemiology and energy, lag features in gradient-boosting forecasters, Prophet's conditional seasonalities and multiplicative regressors, holiday × promotion columns.
How is it used?
Plot the correlation of $y$ with $x$ at several lags; pick lags with a reason; check availability at each horizon; add interaction columns only for effects you can explain; validate by rolling-origin backtests and look at the residuals by group.
"Lagged regressors are always available for forecasting."
Only for horizons up to the lag. A lag-1 feature is unknown two days ahead unless you forecast it or it is planned. Check every lag against every horizon you report.
"With the interaction in the model, the main effect of temperature is the temperature effect."
It is the effect when the other variable is 0 (weekdays here). On weekends the slope is the sum. Always read main effects and interactions together.
Lag: column $x_{t-L}$; known at the origin for horizons $h \le L$ (or if planned). Choose $L$ by domain knowledge + cross-correlation + validation.
Interaction: column $x_a x_b$; slope of $x_a$ becomes $\beta_a + \beta_{ab}x_b$. Keep main effects; centre continuous inputs.
Trap: lags beyond the horizon, and reading a main effect alone when an interaction is present.
Quick check: you forecast 14 days ahead every Monday. Which of these regressors can be used without forecasting it: (a) yesterday's web sessions, (b) spend lagged by 21 days, (c) the price you will set, published a month in advance?
(a) No: lag 1 covers only horizon 1. (b) Yes: lag 21 ≥ 14, so every needed value is already in the past at the origin. (c) Yes: it is planned and known in advance (assuming the plan is followed; if prices often change last-minute, treat it as uncertain).
Multicollinearity: regressors that move together core
Two people always push a stuck car together. The car moves. Who pushed harder? If they always push at the same time, you cannot tell. You only know their total. Regressors that rise and fall together are the same: temperature and the yearly season (summer is hot), price cuts and promotions (they are planned together), web sessions and clicks.
The model can still predict well, because the total effect is well determined. But the individual coefficients become wobbly: in one sample temperature gets most of the credit, in the next the season does. Their uncertainties are large and strongly negatively correlated. This is the regressor version of the overlapping-holiday problem (section 4).
Three ways to say it:
- Picture: estimates scattered along a line "$\beta_1 + \beta_2 \approx$ total" instead of a small round cloud.
- Numbers: correlation 0.9 between two regressors ⇒ VIF $= 1/(1 - 0.81) \approx 5.3$, each SE about 2.3 times larger; correlation 0.99 ⇒ VIF ≈ 50, SE about 7 times larger.
- Slogan: collinear regressors share credit unpredictably; their sum is stable.
- Two regressors with correlation $r = 0.9$ and no other inputs. The variance inflation factor (VIF) of each is $1/(1 - r^2) = 1/(1 - 0.81) = 1/0.19 \approx 5.26$ (Chapter 5.13).
- Standard errors grow by $\sqrt{VIF} \approx 2.29$: an interval that would be ±1 is ±2.3.
- With $r = 0.99$: VIF $= 1/0.0199 \approx 50.3$, SEs × 7.1.
- Temperature vs yearly seasonality: if daily temperature follows the yearly cycle with correlation about 0.95 to the yearly Fourier columns, its VIF is around $1/(1 - 0.9) = 10$. Its coefficient then mostly reflects day-to-day deviations from the seasonal norm, which is often what you want, but it is estimated from much less variation than it seems.
- Predictions for days that look like the training days (both regressors high together) stay stable; predictions for unusual combinations (hot but no promo, when they always came together) are unreliable.
- Multicollinearity: a regressor column is (nearly) a linear combination of other columns, including trend, Fourier and holiday columns. Perfect collinearity makes $X^\top X$ singular; near collinearity makes it ill-conditioned.
- $VIF_j = 1/(1 - R_j^2)$, where $R_j^2$ comes from regressing column $j$ on all the others; $SE_j$ is multiplied by $\sqrt{VIF_j}$. Rules of thumb (rules of thumb only): above 5–10 deserves attention.
- Effects: wobbly individual coefficients, wide intervals, strongly correlated estimates (posterior correlation near −1 for two positively correlated regressors), coefficient signs that flip between samples; stable fitted values within the range of the data.
- Remedies: drop or combine regressors; use the deviation from the seasonal norm (temperature anomaly) instead of the raw value; regularize with priors (ridge-like Normal priors pull the wobble toward a sensible region); report joint effects; use a full-rank guide in SVI so the correlation is represented.
Why do we need it?
Interviewers and stakeholders ask "what is the effect of temperature?". If temperature is collinear with seasonality, the honest answer is uncertain, and a confident number would be misleading. Knowing this also explains odd signs and unstable coefficients between retrains.
Where is it used?
VIF checks in statsmodels (variance_inflation_factor), correlation-matrix heatmaps before adding regressors (Chapter 4.15), marketing-mix models (channels launched together), and posterior correlation plots in Bayesian models.
How is it used?
Before fitting, compute the correlation matrix and VIFs of all regressor, holiday and seasonal columns. After fitting, look at coefficient correlations. If they are extreme, simplify, reparameterize (anomalies, ratios) or use priors, and interpret only the stable combinations.
"The coefficient of temperature is negative, so heat lowers demand."
With a collinear partner (yearly seasonality, a summer promotion), the sign of one coefficient can flip from sample to sample. Look at the joint effect and the coefficient correlations before interpreting a single sign.
"Multicollinearity makes the forecast bad."
Not by itself: predictions for situations like the training data are fine. It hurts interpretation, and forecasts in situations where the usual co-movement breaks (a promotion without the usual price cut).
In your forecasting model, regressors are collinear not only with each other but with the other blocks: temperature with the yearly Fourier columns, promotions with holidays, a marketing regressor with a trend changepoint at the campaign launch. These are the "which component gets the credit?" cases of Chapter 6.8. Practical checks: a correlation matrix of all design columns before fitting, and posterior correlations after. Note that the full-rank vs low-rank guide choice matters here: a mean-field guide would hide exactly these negative correlations.
$VIF_j = 1/(1 - R_j^2)$; SE × $\sqrt{VIF}$. Two regressors with correlation $r$: VIF $= 1/(1 - r^2)$ (0.9 → 5.3, 0.99 → 50).
Effects: wobbly, correlated, sign-flipping coefficients; stable sums and in-range predictions.
Fixes: drop/combine, anomalies instead of raw values, priors, report joint effects, full-rank guide.
Quick check: regressing your price regressor on all other columns gives $R^2 = 0.96$. What are its VIF and SE inflation, and what might you do?
VIF $= 1/(1 - 0.96) = 25$; SE × 5. The price effect is barely estimable separately from the other columns (probably promotions or seasonality). Options: use price relative to the usual price for that season, combine price and promotion into one "discount depth" regressor, put an informative prior on the price elasticity from past experiments, or report only the joint effect.
Standardization, coefficient interpretation and regularization core
Regressors arrive in all sorts of units: degrees, percentages written as fractions, prices in cents, visits in thousands. The coefficient's size depends on the unit: an effect of 150 orders "per 1.0 of discount fraction" is the same as 1.5 orders "per percentage point". So a single prior, say "coefficients are probably within ±10", means something completely different for each regressor: far too tight for one, meaninglessly wide for another.
Standardizing fixes this. Subtract the regressor's mean and divide by its standard deviation (both computed on the training window). Now "one unit" means "one typical change", every coefficient is "the effect of a typical change", and one prior scale is fair to all of them. To talk to the business, convert back: effect per original unit = standardized coefficient ÷ sd.
Three ways to say it:
- Picture: put every regressor on the same ruler before asking the prior to judge them.
- Numbers: temperature sd 4 °C, 5 orders per °C ⇒ 20 orders per sd; discount sd 0.08, 150 orders per unit ⇒ 12 orders per sd.
- Slogan: same ruler, same prior, fair shrinkage.
Two regressors: temperature anomaly (sd 4 °C, true effect 5 orders per °C) and discount as a fraction (sd 0.08, true effect 150 orders per 1.0 of discount, i.e. 1.5 per percentage point). Prior on every coefficient: $N(0, 10^2)$. Noise sd 8, 120 days.
- Raw units. Temperature: 5 is well inside ±10, fine. Discount: the true coefficient is 150, but the prior says "probably within ±20": it pulls hard. The discount column has little spread in its own units (sd 0.08), so the data precision is only $120 \times 0.08^2/8^2 = 0.012$ against a prior precision $1/100 = 0.01$: weight on the data $\approx 0.55$. The estimate is shrunk roughly halfway, to about 80.
- Standardized. $z = (x - \text{mean})/\text{sd}$. Effects per sd: temperature $5 \times 4 = 20$, discount $150 \times 0.08 = 12$. Data precision per coefficient: $120/64 \approx 1.9$, far above 0.01: the prior barely touches either.
- Back to business units: $\hat\beta_{\text{per unit}} = \hat\beta_z/\text{sd}$: $12/0.08 = 150$ orders per 1.0 of discount (1.5 per percentage point).
- The intercept changes meaning too: after standardizing, it is the prediction at average regressor values instead of at zero.
- Standardization: $z_t = (x_t - \bar x_{train})/s_{train}$, with the mean and sd computed on the training window only and then reused unchanged for test and future rows. (Using full-series statistics leaks future information, section 12.)
- Interpretation: raw coefficient = change in $y$ per unit of $x$, holding the other columns fixed; standardized coefficient = change per one training-sd of $x$; $\beta_{raw} = \beta_z/s$. Binary regressors are usually left as 0/1 so their coefficient stays "on vs off".
- Regularization by priors (Chapter 5.3): Normal $N(0, \tau^2)$ ≈ ridge (shrinks all coefficients smoothly); Laplace ≈ lasso at the MAP (pushes weak regressors toward 0, sparse-ish); hierarchical or horseshoe-type priors when there are many candidates. The prior scale τ is only meaningful relative to the units of $x$ and $y$, hence standardize first.
- Prophet:
add_regressor(..., standardize='auto')standardizes a regressor unless it is binary; its prior scale defaults toholidays_prior_scale. Your own model may or may not standardize: check.
Why do we need it?
Without a common scale, one prior shrinks some effects to nothing and ignores others, coefficient sizes cannot be compared, and optimizers (SVI with Adam) struggle with parameters of wildly different sizes.
Where is it used?
Prophet's regressor standardization, scikit-learn's StandardScaler inside a Pipeline, NumPyro models that standardize inputs before putting $N(0, 1)$-style priors on coefficients, and the global scaler of your A/B framework (Chapter 4.18).
How is it used?
Fit the scaler on the training window; transform training, test and future rows with the same numbers; put one sensible prior scale on standardized coefficients; convert estimates back to original units for reporting; store the scaler with the model for production.
"Standardize the regressors using the whole series before splitting into train and test."
That uses future values (test-period mean and sd) in training: leakage. Compute the mean and sd on the training window, store them, and reuse them for every later row.
"Its standardized coefficient is 20, so temperature matters more than discount (12) in every sense."
It is bigger per typical change in the training data. A discount you could raise to 50% (6 sds) would matter more. Compare effects for the changes that are realistic for the decision at hand.
"Standardizing changes the model's predictions."
Without priors it changes nothing but the units of the coefficients. With priors it changes how hard each coefficient is shrunk, which is the point.
Both of your projects face this. In the A/B framework the syllabus insists on one global scaler rather than per-group normalization, because per-group scaling would erase the group differences you want to measure (Chapter 4.18). In the forecasting model, the same logic applies over time: one scaler fitted on the training window, applied unchanged to the test and future periods; never refit it on data that includes the period you evaluate. Then a Normal or Laplace prior on standardized coefficients has a clear meaning: "a typical change in this regressor probably moves demand by less than τ (in the units of $y$ the model sees)".
"The coefficient on standardized temperature is 20, so each degree adds 20 orders."
"Each training standard deviation of temperature (4 °C) adds about 20 orders, so about 5 orders per degree, holding the other terms fixed."
Model answer: "I standardize continuous regressors with training-window statistics so one prior scale is fair to all of them and the optimizer sees similar scales. Coefficients are then per standard deviation; I divide by the sd to report per-unit effects. The scaler is fitted inside each training window to avoid leakage."
$z = (x - \bar x_{train})/s_{train}$; $\beta_{raw} = \beta_z/s$; the intercept becomes the prediction at average inputs.
Priors regularize: Normal ≈ ridge, Laplace ≈ lasso (MAP); τ only means something on a known scale.
Trap: scaling with full-series statistics = leakage; "per sd" is not "per unit".
Quick check: a standardized web-traffic regressor has coefficient 8 orders; the training sd of traffic is 2 000 visits. What is the effect per 1 000 extra visits?
Per visit: $8/2000 = 0.004$ orders; per 1 000 visits: 4 orders (holding the other terms fixed). Before using it, though, ask section 10's question: will tomorrow's web traffic be known when you forecast tomorrow? Usually not.
Will $X_{T+h}$ be available? Known-in-advance vs forecast regressors core
A regressor helps you forecast tomorrow only if you know its value for tomorrow today. Some regressors are known in advance: the calendar, holidays, prices and promotions you have already planned, scheduled TV campaigns. Others are not: tomorrow's actual temperature, a competitor's price, tomorrow's website traffic. For those you have three options: forecast the regressor too (and accept its error), use a lagged value that is already known, or leave it out.
The trap is in the backtest. When you test the model on last year, the actual temperatures of last year are sitting in your table, so it is tempting to use them. That is using information you would never have had at the forecast origin: the backtest becomes an oracle and looks much better than the model will ever be in real use.
Three ways to say it:
- Picture: at the forecast origin, draw a wall; for each regressor ask "do I really have its values to the right of the wall?".
- Numbers: with the actual temperature the forecast error sd is 6; with a 7-day-ahead weather forecast that is assumed (for illustration) to be no better than the seasonal normal, it is about 21, the same as having no temperature at all.
- Slogan: a regressor is only as good as its future values.
Daily orders: $y_t = \text{baseline}_t + 5 \times \text{temp}_t + \epsilon_t$, noise sd $\sigma = 6$; the temperature anomaly has sd 4 °C around its seasonal normal. Assume (for illustration) weather-forecast errors with sd 1 °C one day ahead, 2 °C three days ahead and 4 °C (no better than the seasonal normal) seven days ahead.
- Oracle (actual future temperature, only possible in a careless backtest): error sd $= \sigma = 6$.
- Forecast regressor, 1 day ahead: the temperature error adds $5 \times 1 = 5$ orders of sd. Variances add: $\sqrt{6^2 + 5^2} = \sqrt{61} \approx 7.8$.
- 3 days ahead: $\sqrt{36 + (5 \times 2)^2} = \sqrt{136} \approx 11.7$.
- 7 days ahead: $\sqrt{36 + (5 \times 4)^2} = \sqrt{436} \approx 20.9$.
- No temperature regressor (use the seasonal normal): $\sqrt{36 + (5 \times 4)^2} \approx 20.9$ at every horizon. At 7 days the regressor no longer helps; at 1 day it helps a lot.
- A backtest that used actual temperatures would report error sd 6 at every horizon: more than three times too optimistic at a week ahead.
At the forecast origin $T$ the information set $\mathcal F_T$ is everything known at that moment. A regressor value $X_{T+h}$ is
- known in advance if $X_{T+h} \in \mathcal F_T$: calendar features, holidays, planned prices and promotions (if the plan is reliable), lags with $L \ge h$;
- to be forecast otherwise: weather, competitor actions, macro indicators, traffic. Then the predictive distribution must average over the regressor's uncertainty: $$p(y_{T+h} \mid \mathcal F_T) = \int p(y_{T+h} \mid X_{T+h}, \theta)\; p(X_{T+h} \mid \mathcal F_T)\; dX_{T+h},$$ in practice: draw regressor paths (from a weather ensemble or a model for $X$), then draw $y$ given each path. For a linear term, the extra variance is $\beta^2\,Var(X_{T+h} - \hat X_{T+h})$.
- A backtest must feed each origin the regressor values that were available at that origin (archived forecasts, plans as of that date), not the values recorded later.
Prophet's documentation states the requirement plainly: an extra regressor must be known for the history and for future dates, so it must either have known future values or be forecast separately.
Why do we need it?
It decides whether a regressor belongs in a forecasting model at all, how wide the intervals must be, and whether the backtest is honest. Forgetting it is the most common reason a model that "validated great" fails in production.
Where is it used?
Energy load forecasting with weather forecasts (and ensembles), retail models with planned promotions and price calendars, Prophet's future dataframe (you must fill in every regressor), SARIMAX's exog for the forecast period, and the regressor-uncertainty source of forecast uncertainty (Chapter 7.14).
How is it used?
For each regressor write down: known in advance (by whom, how far ahead) or forecast (by what, with what error). Backtest with as-of-origin values. If it must be forecast, simulate its uncertainty into the predictive distribution, or compare against the honest alternative of leaving it out.
"Temperature improved the backtest RMSE by 40%, so we add it."
Only if the backtest used temperature forecasts as they were available at each origin. With actual temperatures it measured an oracle. Redo it with as-of-origin values; for long horizons the gain may vanish.
"Planned promotions are known in advance, so they are always safe."
Only as reliable as the plan. If promotions are often added or cancelled at short notice, the planned calendar is itself a forecast with errors, and the backtest should use the plan as it stood at each origin.
For every exogenous regressor in your forecasting model, be ready to answer: "where do its future values come from at forecast time?". Calendar-type columns (holidays, Fourier terms) are always available; planned business levers usually are; anything measured after the fact (weather, traffic, competitor data) must be forecast, lagged beyond the horizon, or dropped. If you propagate regressor uncertainty, it becomes one of the sources of forecast uncertainty in Chapter 7.14; if you do not, say that your intervals are conditional on the regressor values.
For each regressor: $X_{T+h}$ known at origin (calendar, plans, lags $L \ge h$) or must be forecast (weather, competitors, traffic)?
Forecast regressor ⇒ extra variance $\beta^2 Var(\text{error of } \hat X)$; simulate $X$ paths, then $y$.
Trap: backtesting with actual future regressor values = oracle = leakage.
Quick check: a model uses "number of delivery drivers on shift" as a regressor. The rota is published two weeks ahead. Can you use it for 7-day forecasts? For 21-day forecasts?
7 days: yes, the rota is known (if it is not changed at short notice), and the backtest should use the rota as published at each origin. 21 days: not directly; beyond two weeks you need a forecast or a default (for example the usual staffing level), and the intervals must reflect that uncertainty.
Data leakage: future information and target leakage core
Imagine a student who sees the exam answers while "practising". Their practice score is brilliant and tells you nothing about the real exam. Leakage is the forecasting version: during training or backtesting, the model gets information it will not have when it forecasts for real. The backtest score becomes a lie, usually a flattering one.
Leaks hide in innocent-looking features. A "7-day average" that is centred on the day uses the three days after it. A "same-day web sessions" column is recorded at the end of the day you are trying to forecast, and it is driven by the very orders you want to predict (that is target leakage: a feature that is partly the target in disguise). A weather column holds the actual temperature, which nobody knew in advance. Values that were corrected weeks later ("backfilled") look cleaner in the table than they were on the day.
The test for every feature is one question: at the forecast origin, would I have had this exact number?
Three ways to say it:
- Picture: a feature with a hidden arrow pointing from the future (or from the target) into the past.
- Numbers: in the widget below, a same-day-sessions feature gives a backtest MAE of about 4 orders and a live MAE of about 30, far worse than using no feature at all (about 12).
- Slogan: if you would not have had it at the origin, you cannot use it in the backtest.
Daily orders, forecast made at the end of day $T$ for days $T + 1, \dots, T + 7$. Three candidate features:
- Same day last week $y_{t-7}$. For day $T + 7$ it needs $y_T$: known. For every horizon up to 7 the value is in the past. Honest. Backtest and live behave the same.
- Centred 3-day average $\frac{1}{3}(y_{t-1} + y_t + y_{t+1})$. It contains the target $y_t$ itself and tomorrow's value. In the stored table it exists for every past day, so a backtest happily uses it. At the origin, the latest one you can compute is for day $T - 1$ (it needs $y_T$); for $T + 1$ onward it does not exist. Leak: future information.
- Same-day web sessions $s_t$: recorded at the end of day $t$, and high because orders are high. It "predicts" $y_t$ almost perfectly in the table. At the origin you only know $s_T$. Leak: target leakage (and future information).
- Typical result (widget below, averaged over 9 origins): lag-7 feature: backtest MAE ≈ live MAE; centred average and sessions: backtest MAE far below live MAE. The leaky features win the backtest and lose in real use.
Data leakage in forecasting: any use, during training, feature construction, model selection or evaluation, of information that would not be available at the forecast origin $T$ for the forecast being made. Common forms:
- Future information in features: centred or forward-looking windows, leads $x_{t+k}$, actual values of regressors that are only known later (weather, competitor prices), lags shorter than the horizon, data revised or backfilled after the fact, joins on tables updated later.
- Target leakage: features computed from the target or caused by it (same-day revenue when forecasting orders, items shipped, sessions driven by the orders, "returns" logged later).
- Evaluation leakage: random K-fold or shuffled splits (Chapter 7.1); tuning on the test period.
- Pipeline leakage: scaling, imputation, feature selection or changepoint detection fitted on the full series (next section).
Symptoms: suspiciously good backtests, a feature that "explains everything", a big gap between backtest and live accuracy.
Why do we need it?
Leakage makes model choices on false evidence: you ship the model with the leakiest feature. It is the most common reason a forecasting model that validated well fails in production, and interviewers ask about it specifically.
Where is it used?
Every backtest: rolling-origin evaluation (Chapter 7.15), Prophet's cross_validation, scikit-learn's TimeSeriesSplit, Kaggle-style competitions (where leaks are notorious), and feature stores that keep "as-of" timestamps.
How is it used?
For each feature, write down when its value becomes known and compare with the origin and horizon. Build features with an "as of" time; in backtests, rebuild them as they would have been at each origin. Treat any feature that makes the backtest dramatically better with suspicion until proven honest.
"The feature is in our data warehouse for every past day, so it is available."
Availability is about when the value became known, not whether it is stored. A table of past days contains everything, including values that were only known later.
"Rolling-origin evaluation protects me from leakage."
It protects against one leak (training on the future). If the features themselves contain future information, rolling origins still leak. Features must be rebuilt as of each origin.
"Our backtest MAPE is 2%, so the model is excellent."
"Our backtest MAPE is 2% using only inputs available at each origin: calendar features, planned promotions, and lags at least as long as the horizon; scalers and changepoints were fitted inside each training window."
Model answer: "Leakage means using information at training or evaluation time that would not exist at the forecast origin. In forecasting it comes from future-looking features, regressors whose future values are unknown, target-derived features, and preprocessing or model selection done on the full series. I check every feature's availability time against the horizon, rebuild features as of each origin in backtests, and treat a sudden large improvement as a leak until proven otherwise."
Leakage is a P0 topic of your syllabus for good reason: your forecasting model has several doors through which the future can sneak in, namely exogenous regressors (are their future values known?), holiday and event tables (were one-off events flagged only with hindsight?), scaling of $y$ and $X$, and changepoints detected by PELT. When you describe your validation, say explicitly which inputs were available at each origin and that every fitted step ran on the training window only. If any of them did not, the honest move is to say so and explain how you would fix it.
Leakage = using information not available at the origin (features, regressors, target-derived columns, tuning, preprocessing).
Test per feature: when does its value become known? Compare with origin + horizon. Rebuild features as of each origin.
Trap: a great backtest from a centred window, actual future regressors or a same-day proxy of the target.
Quick check: is "number of orders cancelled on day $t$" a safe feature for forecasting orders on day $t$?
No. Cancellations on day $t$ are only known after day $t$, and they are driven by day $t$'s orders (target leakage). Lagged cancellations (for example from $t - 7$, for horizons up to 7) are fine if they were recorded by the origin.
Leakage through the pipeline: scaling, feature selection, changepoints core
Even with honest features, the steps around the model can peek. Every step that learns something from data is part of the model: computing a mean and sd to standardize, filling gaps by interpolation, deciding which regressors to keep, choosing the Fourier order or prior scale, detecting changepoints with PELT. If any of these steps sees the test period, the test period has helped build the model that is then judged on it.
The effect can be small (a scaler fitted on all data shifts numbers a little) or large (choosing, among 30 candidate regressors, the 3 that happened to fit the test period best; or letting PELT see that the level jumped just before the origin). Either way the backtest is no longer a fair rehearsal of a real forecast.
Three ways to say it:
- Picture: draw the origin wall through the whole pipeline, not just through the model fit.
- Numbers: keeping the 3 best of 30 pure-noise regressors by backtest error improved the backtest MAE by about 6% in simulation while making fresh forecasts about 4% worse.
- Slogan: everything that is fitted is fitted on the training window.
- Scaling. Standardizing temperature with the mean of all 3 years (including the test year) uses a number that did not exist at the origin. Correct: compute the mean and sd on the training window, store them, reuse them for the test period. (Same for dividing $y$ by its maximum.)
- Imputation. Filling a missing day by interpolating between the day before and the day after uses the future if the gap is at the end of the training window. Use only past values there.
- Feature selection. 30 candidate regressors that are pure noise (unrelated to demand). For each, fit on training data and score on the 4-week backtest window; keep the 3 best. Averaged over 20 simulated datasets (the Code-it block): backtest MAE 7.66 without them, 7.16 with them; on the next 4 weeks: 8.47 without, 8.78 with. Selection on the test window rewarded luck.
- Changepoints. PELT run on the full series sees a level jump 3 days before the origin clearly, because 30 more days of the new level follow it. PELT on the training window alone sees only 3 days of the new level. In 100 simulations (jump +12, noise sd 8, penalty $2\sigma^2\log n$; the Code-it block uses
ruptures), the full-series run found the jump 85 times, the training-window run 40 times; with the jump 1 day before the origin, 90 times vs 2. The leaky backtest therefore forecasts the new level far more often than a real forecast at that origin could.
A forecasting pipeline is every step from raw data to forecast. For an honest backtest at origin $T$:
- Every step that estimates anything (scalers, imputers, encoders, feature or order selection, hyperparameters, prior scales, changepoint detection, outlier flags) is fitted on data up to $T$ only, then applied unchanged to later rows.
- Model and hyperparameter choices are made with an inner validation inside the training window; the outer test window is used once, for the final report.
- With rolling origins, the whole pipeline is refitted at each origin. A gap (embargo) between training and test protects against features built from windows that would otherwise overlap.
- Tools: scikit-learn
PipelinewithTimeSeriesSplit; Prophet'scross_validationrefits the model at each cutoff, but any preprocessing you do before calling it (scaling, PELT, feature selection) is your responsibility.
Why do we need it?
Pipeline leaks are invisible in the model code: the model is fitted on the training rows, so everything looks correct. The leak sits in a preprocessing script that ran once on the whole dataset.
Where is it used?
scikit-learn Pipeline/ColumnTransformer, feature stores with point-in-time joins, Prophet and statsmodels backtests, PELT/ruptures changepoint pre-processing, AutoML and hyperparameter searches that must be nested inside time splits.
How is it used?
Write the pipeline as one function fit(train) → predictor; call it at each origin; never compute anything on the full table outside it. Audit by asking, for every number the model uses, "which rows was this computed from?".
"I only fit the model on the training data, so there is no leakage."
Check every step before the model: scalers, imputers, outlier flags, selected regressors, Fourier orders, prior scales, changepoints. If any of them was computed on rows after the origin, the backtest leaks.
"A scaler fitted on the full series is a tiny leak, not worth fixing."
Often small, sometimes not (a test-period peak changes $\max|y|$; a trend changes the mean). It costs nothing to fit it inside the training window, and it removes a question from every review.
Your forecasting pipeline uses PELT to choose changepoints before the Bayesian fit (Chapter 7.9). In a backtest, PELT must run on each training window separately; running it once on the full history lets test-period behaviour place the changepoints, which is exactly the leak in the widget above. The same applies to any scaling of $y$ and $X$, to holiday and event flags added with hindsight, and to the choice of Fourier orders and prior scales. A clear statement you can make: "the whole pipeline, including PELT and scaling, is refitted at every origin of the rolling backtest".
Pipeline leakage: any learned step (scaling, imputation, selection, tuning, PELT) that saw data after the origin.
Rule: fit every step on the training window, freeze it, apply to test; refit everything at each rolling origin; nest tuning inside training.
Trap: "the model was fitted on train" is not enough; selecting by backtest score rewards luck.
Quick check: you tried 12 Fourier-order and prior-scale combinations and report the best one's backtest error. What is wrong, and how do you fix it?
The backtest window was used to choose among 12 options, so its error is optimistically biased (the winner is partly the luckiest). Fix: choose the combination with an inner validation inside each training window (or on an earlier period), then report the error on a later window that played no part in the choice.
Recap, cheat sheet and practice
- Holidays are 0/1 indicator columns: $h_t = \sum_j \beta_j I(t \in D_j)$; $\beta_j$ is the extra over a normal day with the same trend, season and regressors; dates must cover the forecast horizon.
- Windows $[L, U]$ give one column (one coefficient) per day offset; choose them from the average remainder by offset; each offset is learned from only as many days as there are occurrences.
- Recurring events share one coefficient and repeat in the future (moving dates from a calendar); one-off events get a column only on their own days, to protect the trend and seasonality (or are masked).
- Overlapping holidays: if always together, only the sum is identified and the prior decides the split; separate occurrences identify each.
- Sparse holidays need shrinkage: posterior mean $= w\bar r$ with $w = \frac{n/\sigma^2}{n/\sigma^2 + 1/\tau^2}$; MAP = ridge. τ lives on the scale of the (possibly rescaled) $y$; Prophet's default 10 on $y/\max|y|$ is weak.
- Regressors: continuous (per unit), binary (on vs off), categorical ($K - 1$ dummies vs a reference; all $K$ + intercept = dummy trap), lagged ($x_{t-L}$, known for $h \le L$), interactions (product columns; slopes that depend on context).
- Multicollinearity: VIF $= 1/(1 - R_j^2)$; wobbly, correlated coefficients but stable sums and in-range predictions. Fix by combining, anomalies, priors, full-rank guides.
- Standardize with training-window mean and sd; coefficients become "per sd" ($\beta_{raw} = \beta_z/s$); one prior scale then treats all regressors fairly.
- Future availability: for each regressor, is $X_{T+h}$ known at the origin (calendar, plans, long lags) or must it be forecast (weather, competitors, traffic)? Forecast regressors add variance $\beta^2 Var(\hat X\text{ error})$; backtests must use as-of-origin values.
- Leakage: future information in features, target leakage, evaluation on shuffled data, and pipeline steps (scaling, imputation, selection, tuning, PELT) fitted on the full series. Every learned step runs on the training window, at every origin.
Cheat sheet
| Idea | Formula / rule | Remember |
|---|---|---|
| Holiday term | $h_t = \sum_j \beta_j I(t \in D_j)$ | extra over the baseline; future dates needed |
| Window | columns $I(t - k \in D)$, $k = L..U$ | $U - L + 1$ coefficients per holiday |
| One-off event | column = 1 on its days only | protects other terms; 0 in the future |
| Overlap | always together ⇒ only $\beta_A + \beta_B$ | split = prior; check "only A / only B" counts |
| Shrinkage | $w = \frac{n/\sigma^2}{n/\sigma^2 + 1/\tau^2}$, mean $= w\bar r$ | MAP = ridge $\lambda = \sigma^2/\tau^2$ |
| Prophet priors | holidays_prior_scale = 10 (default) | on $y/\max|y|$: weak; lower to dampen |
| Categorical | $K - 1$ dummies + intercept | coefficients vs the reference level |
| Lag | $x_{t-L}$ known at origin for $h \le L$ | else plan or forecast it |
| Interaction | slope of $x_a$ = $\beta_a + \beta_{ab}x_b$ | keep main effects; centre inputs |
| VIF | $1/(1 - R_j^2)$; two regressors: $1/(1 - r^2)$ | SE × $\sqrt{VIF}$; 0.9 → 5.3 |
| Standardize | $z = (x - \bar x_{train})/s_{train}$, $\beta_{raw} = \beta_z/s$ | training stats only |
| Forecast regressor | $Var = \sigma^2 + \beta^2 Var(X - \hat X)$ | oracle backtests are leakage |
| Leakage test | "would I have had this number at the origin?" | features, regressors, preprocessing, tuning |
import numpy as np
import pandas as pd
import ruptures as rpt
from sklearn.preprocessing import StandardScaler
from statsmodels.stats.outliers_influence import variance_inflation_factor
# 1) Holiday indicator columns with a window (lower -2, upper +1), Prophet-style: one column per offset
days = pd.date_range("2022-01-01", "2024-12-31", freq="D")
xmas = pd.to_datetime(["2022-12-25", "2023-12-25", "2024-12-25"])
H = pd.DataFrame({f"xmas_{k:+d}": days.isin(xmas + pd.Timedelta(days=k)).astype(int)
for k in range(-2, 2)}, index=days)
print(H.sum().to_dict()) # {'xmas_-2': 3, 'xmas_-1': 3, 'xmas_+0': 3, 'xmas_+1': 3}
print(H.loc["2023-12-22":"2023-12-27"].values.tolist()) # a shifted diagonal of 1s
# 2) Shrinkage of a holiday effect: prior N(0, 30^2), noise sd 20, average remainder 60
for n in (1, 3):
w = (n / 20**2) / (n / 20**2 + 1 / 30**2)
print(n, round(w, 3), round(w * 60, 1), round((n / 20**2 + 1 / 30**2) ** -0.5, 1))
# 1 0.692 41.5 16.6 | 3 0.871 52.3 10.8
# 3) Overlapping holidays: rows (A on, B on); two overlaps, then one separate year each
X = np.array([[1, 1], [1, 1], [1, 0], [0, 1.0]]); r = np.array([80, 76, 50, 30.0])
print(np.linalg.lstsq(X, r, rcond=None)[0]) # [49.2 29.2]
print(np.linalg.matrix_rank(X[:2])) # 1: overlaps alone fix only A + B
# 4) Categorical regressor: the dummy trap
promo = pd.Series(["none", "email", "banner", "none", "email", "none"])
D_all = pd.get_dummies(promo, dtype=float) # all 3 levels
D_ref = pd.get_dummies(promo, drop_first=True, dtype=float) # reference level dropped
print(np.linalg.matrix_rank(np.column_stack([np.ones(6), D_all])), np.linalg.matrix_rank(np.column_stack([np.ones(6), D_ref])))
# 3 3: four columns but rank 3 (dummy trap) vs three columns, full rank
# 5) Multicollinearity: VIF for two regressors with correlation 0.9
print(round(1 / (1 - 0.9**2), 2)) # 5.26 (theory)
rng = np.random.default_rng(0)
a = rng.normal(size=20000); b = 0.9 * a + np.sqrt(1 - 0.81) * rng.normal(size=20000)
Xc = np.column_stack([np.ones_like(a), a, b])
print(round(float(variance_inflation_factor(Xc, 1)), 2)) # 5.18 here: close to the theory
# 6) A regressor that must be forecast: error sd at horizons with temperature errors 0, 1, 2, 4 degrees
for e in (0, 1, 2, 4):
print(e, round(np.sqrt(6**2 + (5 * e) ** 2), 1)) # 6.0, 7.8, 11.7, 20.9
# 7) Leakage: a centred average and same-day sessions vs an honest lag-7 feature
def fourier(t, P, N):
return np.column_stack([f(2 * np.pi * k * t / P) for k in range(1, N + 1) for f in (np.sin, np.cos)])
wk = np.array([-20, -25, -22, -15, 0, 40, 42.0])
n = 160; t = np.arange(n); u = np.zeros(n)
for i in range(1, n): u[i] = 0.7 * u[i - 1] + rng.normal(0, 10)
y = 200 + 0.8 * wk[t % 7] + u
feat = {"lag7": np.r_[np.full(7, np.nan), y[:-7]],
"ma3": np.r_[np.nan, (y[:-2] + y[1:-1] + y[2:]) / 3, np.nan], # uses y[t+1]: future
"sess": 20 * y + rng.normal(0, 100, n)} # caused by the target
for name, x in feat.items():
eb, el = [], []
for T in range(90, 147, 7): # 9 rolling origins, 7-day horizon
tr = np.arange(7, T); tr = tr[~np.isnan(x[tr])]
Xf = np.column_stack([np.ones(n), fourier(t, 7, 3), x])
beta = np.linalg.lstsq(Xf[tr], y[tr], rcond=None)[0]
te = np.arange(T, T + 7)
Xlive = Xf[te].copy()
if name == "ma3": Xlive[:, -1] = x[T - 2] # latest value computable at the origin
if name == "sess": Xlive[:, -1] = x[T - 1] # latest value known at the origin
eb += list(np.abs(y[te] - Xf[te] @ beta)); el += list(np.abs(y[te] - Xlive @ beta))
print(name, round(np.mean(eb), 1), round(np.mean(el), 1)) # backtest MAE vs live MAE
# lag7 11.5 11.5 | ma3 5.1 22.8 | sess 4.2 38.9
# 8) Pipeline leakage: choose 3 of 30 noise regressors by backtest error (20 simulated datasets)
res = []
for seed in range(20):
g = np.random.default_rng(seed); m = 196; tt = np.arange(m)
yy = 200 + 0.8 * wk[tt % 7] + g.normal(0, 10, m); C = g.normal(0, 1, (m, 30))
tr, bt, fr = tt < 140, (tt >= 140) & (tt < 168), tt >= 168
B = np.column_stack([np.ones(m), fourier(tt, 7, 3)])
def score(X):
b = np.linalg.lstsq(X[tr], yy[tr], rcond=None)[0]; e = np.abs(yy - X @ b)
return e[bt].mean(), e[fr].mean()
best = np.argsort([score(np.column_stack([B, C[:, k]]))[0] for k in range(30)])[:3]
res.append(score(B) + score(np.column_stack([B, C[:, best]])))
print(np.round(np.mean(res, axis=0), 2)) # [7.66 8.47 7.16 8.78]: base bt, base fresh, selected bt, selected fresh
# 9) Changepoints: PELT on the full series vs on the training window only (ruptures, l2 cost)
pen = lambda N: 2 * 8**2 * np.log(N) # BIC-like penalty, noise sd 8
for J in (119, 117, 110): # jump of +12 at day J; origin = day 120
found_full = found_train = 0
for seed in range(100):
z = 100 + 12 * (np.arange(150) >= J) + np.random.default_rng(seed).normal(0, 8, 150)
full = rpt.Pelt(model="l2", min_size=2, jump=1).fit(z).predict(pen=pen(150))[:-1]
train = rpt.Pelt(model="l2", min_size=2, jump=1).fit(z[:120]).predict(pen=pen(120))[:-1]
found_full += any(abs(b - J) <= 3 for b in full); found_train += any(abs(b - J) <= 3 for b in train)
print(120 - J, found_full, found_train) # 1: 90 vs 2 | 3: 85 vs 40 | 10: 95 vs 85 (of 100)
# 10) Scaling inside the training window only
temp = rng.normal(15, 6, 400)
sc = StandardScaler().fit(temp[:300, None]) # fit on training rows only
print(round(sc.mean_[0], 2), round(sc.scale_[0], 2), sc.transform(temp[300:305, None]).ravel().round(2))
# the test rows are transformed with the TRAINING mean and sd, never refitted
1. A fitted holiday coefficient for Christmas Eve is +60 orders. What does it mean?
2. Your anniversary sale has always fallen on Black Friday, and both have their own indicator column. What can the data tell you?
3. A holiday has been seen twice and the daily noise is large. What does a $N(0, \tau^2)$ prior on its coefficient do?
4. You forecast 7 days ahead every morning. Which regressor can you use without forecasting it?
5. A standardized discount regressor has coefficient 12 orders; the training sd of the discount is 0.08 (as a fraction). What is the effect of raising the discount from 0.10 to 0.20?
6. Which of these steps leaks future information into a rolling-origin backtest?
Practice problems
A. You add 12 holidays, each with a window from 1 day before to 1 day after, to 3 years of daily data. How many holiday coefficients are there, and how many observations inform each?
Each holiday has offsets −1, 0, +1: 3 columns. $12 \times 3 = 36$ coefficients. Each offset column is 1 on one day per year, so each coefficient rests on 3 observations (fewer if some holidays overlap). That is a strong argument for shrinkage priors and for shorter windows where the remainders show no effect.
B. A holiday seen twice has average remainder +40; daily noise sd 15; prior $N(0, 20^2)$. Find the posterior mean and sd.
- Data precision $2/15^2 = 2/225 \approx 0.00889$; prior precision $1/400 = 0.0025$.
- $w = 0.00889/(0.00889 + 0.0025) \approx 0.780$; posterior mean $= 0.780 \times 40 \approx 31.2$.
- Posterior sd $= (0.01139)^{-1/2} \approx 9.37$. A 95% credible interval is roughly $31.2 \pm 18.4$: the effect is clearly positive but its size is uncertain.
C. Events A and B coincided twice (remainders +70, +74); once they fell apart (A alone +45, B alone +28). Find the least-squares effects.
- Minimise $(70 - A - B)^2 + (74 - A - B)^2 + (45 - A)^2 + (28 - B)^2$.
- Derivatives: $3A + 2B = 70 + 74 + 45 = 189$ and $2A + 3B = 70 + 74 + 28 = 172$.
- $A - B = 17$ and $A + B = 361/5 = 72.2$, so $\hat\beta_A = 44.6$, $\hat\beta_B = 27.6$. The single separate year carries the identification; with only the two overlaps you would know just $A + B \approx 72$.
D. Orders rise 3 per °C; noise sd 10. Five days ahead, the temperature forecast has error sd 3 °C. What is the forecast error sd, and what would a backtest with actual temperatures report?
$\sqrt{10^2 + (3 \times 3)^2} = \sqrt{100 + 81} = \sqrt{181} \approx 13.5$ orders. A backtest with actual temperatures (an oracle) would report about 10: roughly 26% too optimistic at this horizon, and its intervals would be too narrow.
E. (Interview) "Your backtest looks excellent. How do you know it is not leakage?"
"I checked three things. First, features: every input's value is available at the forecast origin for every horizon I report: calendar and holiday columns, planned promotions as published at that origin, and lags at least as long as the horizon; nothing centred, nothing measured on the forecast day, nothing derived from the target. Second, the pipeline: scaling, imputation, changepoint detection with PELT, Fourier orders and prior scales are all fitted inside each training window of a rolling-origin backtest, and choices are made on an inner validation, not on the reported test window. Third, sanity: no single feature produces a dramatic jump in accuracy, and backtest errors are in line with live errors from a shadow period." (Say only what you actually did; offer the rest as how you would check.)
F. (Interview) "How would you add weather to your forecasting model?"
"Temperature is not known at the forecast origin, so I would use weather forecasts as the regressor, both when training on history (archived forecasts as of each origin, if available) and when forecasting. I would standardize it with training statistics, use the anomaly from the seasonal normal to reduce collinearity with the yearly Fourier terms, put a Normal prior on its coefficient, and propagate the forecast error into the predictive distribution by sampling weather scenarios. Then I would compare against the model without weather in a rolling-origin backtest by horizon: the benefit is usually large for a day or two ahead and fades by a week."
Forecast likelihoods: Normal, Student-t, Negative Binomial
The layers of your forecasting model (trend, seasonality, holidays, regressors) say what a day is expected to be. They do not say how far a real day can land from that expectation. That job belongs to the likelihood, the noise model. This chapter teaches the three likelihoods in your forecasting model: the Normal (the classic bell), the Student-t (a bell that expects the occasional huge surprise) and the Negative Binomial (for counts that wobble more than a Poisson allows). You will learn what each one assumes, how to see which one your data needs, and the exact wording an interviewer expects.
- Say what a likelihood does in a forecasting model: it turns the expected value $\mu_t$ from the layers into a whole distribution for $y_t$, and the fit adds up its log-scores
- State the Normal assumptions (continuous, symmetric, light tails, constant variance) and spot which ones your residuals break
- Explain the Student-t likelihood (heavy tails, $\nu$, relation to the Normal) and say it right: it assigns more probability to extreme residuals, so they exert less influence. It does not remove outliers. Prove it with a demo of influence
- Explain why a Normal fails for counts, why a Poisson is too restrictive, what overdispersion is, and use the Negative Binomial with $Var(y) = \mu + \mu^2/\alpha$
- Know the exact parameterization you used: NumPyro
NegativeBinomial2,Probs,Logits,GammaPoisson, SciPynbinom, statsmodels; check the implied mean and variance - Keep a mean positive with a log or softplus link, and know what each link does to trends and holiday effects
- Tell zero inflation, hurdle models and plain overdispersion apart, and compare likelihoods fairly (held-out log-likelihood, not apples against oranges)
What we need from earlier chapters: the additive model $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$ (Chapter 7.7); the Normal and Student-t distributions (Chapter 4.9); the Poisson and Negative Binomial, overdispersion and the parameterization table (Chapter 4.8); choosing a likelihood from support and variance (Chapter 4.11); Q-Q plots (Chapter 4.17); the likelihood and maximum likelihood (Chapter 5.2); GLMs, link functions and Negative Binomial regression (Chapter 5.14); the posterior predictive (Chapter 6.1). You will meet the tools for checking a likelihood later: posterior predictive checks (Chapter 7.14), probabilistic scores (Chapter 7.16) and residual diagnostics (Chapter 7.17); here you get a first look. Notation: $\mu_t$ is the expected value of $y_t$ given the layers (the "bullseye" of day $t$); $r_t = y_t - \mu_t$ is the residual; $\varphi$ stands for the extra noise parameters ($\sigma$, $\nu$ or $\alpha$). $N(\mu, \sigma^2)$ is written with the variance; NumPyro and SciPy take the standard deviation $\sigma$. $\sim$ reads "is distributed as".
Colours in every plot of this chapter: blue = Normal, orange = Student-t, teal = Negative Binomial, purple = a special line (a mean, a threshold), red = residuals, errors and impossible values. The data are drawn in the neutral ink colour, and labels always say the same thing in words.
The likelihood is the noise story: from an expected value to a whole distribution core
Think of archery. The layers of your model (trend, weekly pattern, holiday, regressors) decide where the bullseye is for each day: "Tuesday, no holiday, a promo running: aim at 120 orders". The real day is an arrow. It rarely hits the bullseye exactly. The likelihood describes how the arrows scatter around the bullseye: tightly or loosely, evenly on both sides or not, with or without the odd arrow that lands far away.
Choosing a likelihood is choosing that scatter story. A Normal story says "arrows land in a round cloud, and a wild one is nearly impossible". A Student-t story says "mostly a round cloud, but wild arrows happen". A Negative Binomial story says "arrows are whole numbers, and the cloud gets bigger when the bullseye is bigger". The story matters for three jobs: fitting the layers, widths of forecast intervals, and the probability of extremes such as "demand above capacity".
Three ways to say it:
- Picture: the layers place the bullseye; the likelihood draws the cloud of arrows around it.
- Numbers: bullseye 120, spread 10. Under a Normal story a day at 150 is 3 spreads away, about 1 day in 740. Under a heavy-tailed story it is not shocking.
- Slogan: the layers say where; the likelihood says how far.
Three days. The layers give expected orders $\mu = 120, 100, 140$. The real orders were $126, 91, 141$, so the residuals are $+6, -9, +1$. Use a Normal likelihood with $\sigma = 8$. A likelihood gives each day a score: the log of how probable that day was.
- Day 1: $z = 6/8 = 0.75$. Score $= -\tfrac12 z^2 - \log\sigma - \tfrac12\log(2\pi) = -0.281 - 2.079 - 0.919 = -3.280$.
- Day 2: $z = -9/8 = -1.125$. Score $= -0.633 - 2.079 - 0.919 = -3.631$.
- Day 3: $z = 1/8 = 0.125$. Score $= -0.008 - 2.079 - 0.919 = -3.006$.
- The days are treated as independent given their bullseyes, so the scores add: $\ell = -3.280 - 3.631 - 3.006 = -9.917$. This number is the log-likelihood. Higher (closer to 0) means the data were less surprising to the model.
- Try another $\sigma$. The best one is the root-mean-square residual: $\sqrt{(36 + 81 + 1)/3} = \sqrt{39.33} = 6.27$. With $\sigma = 6.27$ the log-likelihood rises to $-9.765$. With $\sigma = 4$ it falls to $-10.603$, and with $\sigma = 12$ to $-10.621$. Fitting a noise parameter is just "pick the value that makes the data least surprising".
A likelihood for a forecasting model says how each observation is distributed around its expected value:
$$y_t \mid \mu_t,\varphi \;\sim\; p(y \mid \mu_t, \varphi), \qquad \mu_t = g(t) + s(t) + h(t) + X_t\beta .$$$\mu_t$ comes from the layers (or from a function of them, see the link in a later section), and $\varphi$ holds the noise parameters. The log-likelihood of all the days is
$$\ell = \sum_{t=1}^{T} \log p(y_t \mid \mu_t, \varphi).$$- Maximizing $\ell$ over the layers' parameters and $\varphi$ gives the maximum-likelihood fit; multiplying by priors gives the Bayesian posterior (Chapter 5.2, Chapter 6.1).
- Assumption: given $\mu_t$, the days are independent. If the residuals are autocorrelated (a good Tuesday follows a good Monday) that assumption is broken; it shows up in the residual diagnostics of Chapter 7.17.
| Likelihood | Possible values | Noise parameters | Spread of $y_t$ around $\mu_t$ | Natural fit |
|---|---|---|---|---|
| Normal | any real number | $\sigma$ | sd $= \sigma$, the same every day | continuous, symmetric, light tails |
| Student-t | any real number | $\sigma,\ \nu$ | scale $\sigma$ (sd $= \sigma\sqrt{\nu/(\nu-2)}$ if $\nu \gt 2$), the same every day | continuous, with occasional extreme residuals |
| Negative Binomial | $0, 1, 2, \dots$ | $\alpha$ | sd $= \sqrt{\mu_t + \mu_t^2/\alpha}$, grows with $\mu_t$ | counts that wobble more than a Poisson |
Why do we need it?
The layers give only a centre. To get a fit, an interval, or the probability that demand beats capacity, you need a distribution around that centre. A wrong noise story gives intervals that are too wide, too narrow or in impossible places, even when the centre is perfect.
Where is it used?
The numpyro.sample("obs", dist.Normal(mu, sigma), obs=y) line of any NumPyro regression, the likelihood switch in your forecasting model, GLMs (Chapter 5.14), and neural forecasters such as DeepAR that output the parameters of a Normal or Negative Binomial for each day.
How is it used?
Ask three questions: is $y$ a count or a real number, are there extreme residuals, and does the spread grow with the level? Pick the likelihood, put $\mu_t$ in its location slot, fit, then check residuals and posterior predictive draws (Chapter 7.14).
"If my trend, seasonality and holidays are right, the likelihood does not matter."
The centre may be fine, but the widths, the tail probabilities and (with outliers) even the centre itself depend on the noise story. "Probability that demand exceeds capacity" is a statement about the tail, so it is a statement about the likelihood.
"Residuals from a good model should look Normal, so I should always use a Normal likelihood."
Residuals are only Normal if the noise is. Counts, spikes and a spread that grows with the level are typical reasons they are not. Check first; choose the likelihood second.
Your forecasting model offers Normal, Student-t and Negative Binomial likelihoods because daily business series come in different flavours: revenue-like continuous values, series with occasional glitches or one-off events, and order counts. In your A/B framework the same logic applies to the metric likelihood: Normal and Student-t for continuous metrics, Poisson for counts, and Beta-Binomial or Dirichlet-Multinomial for conversions and categories. In both projects the likelihood is the line where the data enter the model, and the layers (or group means) enter as its location argument.
$y_t \sim p(\,\cdot \mid \mu_t, \varphi)$ with $\mu_t = g + s + h + X\beta$. $\ell = \sum_t \log p(y_t \mid \mu_t, \varphi)$; the fit makes $\ell$ (or $\ell$ + log prior) large.
Layers say where (the centre); the likelihood says how far (scatter, tails, whether the spread grows with $\mu$).
Trap: a great centre with a wrong likelihood still gives wrong intervals and wrong tail probabilities. Days are assumed independent given $\mu_t$.
Quick check: $\mu = 120$, Normal likelihood with $\sigma = 10$. Roughly how often should a day reach 150 or more?
$z = (150 - 120)/10 = 3$. The one-sided tail beyond 3 is $0.00135$, about 1 day in 740 (a two-sided "3 sigma" statement would say 0.27%). A forecaster who sees such a day every few months should doubt the Normal story, not just call it bad luck.
The Normal likelihood: a round cloud of the same size every day core
The Normal likelihood is the default noise story, and it makes four quiet promises about your residuals. They are continuous (any decimal number is possible), symmetric (as likely to miss high as low), light-tailed (a miss three typical-sizes away is rare and one five typical-sizes away is practically impossible), and the same size every day (one spread number $\sigma$ for all days, quiet or busy).
When those promises are roughly true, the Normal is wonderful: simple, fast, and its fit is ordinary least squares. When one is false, the Normal does not crash. It quietly gives the wrong answers: an inflated $\sigma$, intervals that are too wide in calm times and too narrow in wild times, and a trend dragged around by a few strange days.
Three ways to say it:
- Picture: a round cloud of arrows around the bullseye, always the same size, never lopsided, almost never far away.
- Numbers: 68% of days within $\pm\sigma$, 95% within $\pm2\sigma$, 99.7% within $\pm3\sigma$. A day $4.5\sigma$ away is a once-in-hundreds-of-years event.
- Slogan: Normal likelihood = least squares = "big misses are very unlikely, so they are very expensive".
Ten residuals (actual minus expected orders): $2, -1, 3, -4, 0, 1, -2, 5, -3, -1$. Fit the one noise parameter $\sigma$ of a Normal likelihood.
- Square each residual: $4, 1, 9, 16, 0, 1, 4, 25, 9, 1$. Their sum is $70$.
- The maximum-likelihood $\hat\sigma$ is the root of the average squared residual: $\hat\sigma = \sqrt{70/10} = 2.65$. (The sample standard deviation divides by $n-1$ and gives $2.79$; with 10 points the two differ visibly, with 1 000 they do not.)
- Check the promise "68% within $\pm\sigma$": residuals with $|r| \le 2.65$ are $2, -1, 0, 1, -2, -1$, that is 6 of 10 = 60%. Close enough for 10 points.
- Now add one strange day with residual $+12$. It is $12/2.65 = 4.54$ sigmas away. Under the fitted Normal, a miss that big (in either direction) happens with probability $5.7\times10^{-6}$ per day, about once in 174 000 days (477 years). Its log-score is about $-\tfrac12 \times 4.54^2 = -10.3$, against $-0.5$ for a typical day.
- Refit with the strange day included: sum of squares $70 + 144 = 214$, $\hat\sigma = \sqrt{214/11} = 4.41$. One day made $\hat\sigma$ jump by a factor $4.41/2.65 = 1.67$, so every 90% interval became 67% wider: $\pm7.26$ instead of $\pm4.35$.
A Normal likelihood says $y_t \mid \mu_t \sim N(\mu_t, \sigma^2)$, that is
$$p(y_t \mid \mu_t,\sigma) = \frac{1}{\sqrt{2\pi\sigma^2}}\exp\!\Big(-\frac{(y_t-\mu_t)^2}{2\sigma^2}\Big),\qquad \ell = -\frac{T}{2}\log(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum_{t=1}^T (y_t-\mu_t)^2 .$$- For fixed $\sigma$, maximizing $\ell$ over the layers is least squares: minimize $\sum (y_t - \mu_t)^2$ (Chapter 5.2).
- For fixed layers, the best noise level is $\hat\sigma^2 = \frac1T\sum_t r_t^2$, the mean squared residual.
- Assumptions: (1) $y$ continuous; (2) residuals symmetric; (3) light tails; (4) constant variance $\sigma^2$ across days (unless you model it, for example $\sigma_t = \sigma\,\mu_t$ for a constant relative error); (5) independence given $\mu_t$.
- NumPyro:
dist.Normal(loc, scale). The second argument is the standard deviation $\sigma$, not the variance.
Why do we need it?
We need one simple, well-understood noise story that gives a fit by least squares, closed-form intervals ($\mu_t \pm 1.645\sigma$ for 90%) and easy maths. Many series are close enough to it, and every other likelihood is judged against it.
Where is it used?
Linear regression and its t-tests (Chapter 5.13), the default likelihood of Prophet-style and structural time-series models, Kalman filters, ARIMA's maximum-likelihood fit, and the Normal option of your forecasting model and your A/B framework.
How is it used?
Use it for continuous, roughly symmetric residuals whose spread does not depend on the level. Fit, then draw a residual histogram, a Q-Q plot and residual-vs-fitted plot (Chapter 7.17). If the tails are fat, the spread fans out, or values must be positive counts, change the likelihood.
"A Normal likelihood means my data must look Normal."
It says the noise around the expected value is Normal. The series itself has trend, seasons and holidays, so its histogram is usually far from a bell. Look at the residuals.
"Residuals with mean zero and a bell-ish histogram: the Normal is fine."
Check all four promises, not just the bell. Constant variance and tail weight are the two that fail most often, and a single histogram can hide the first (look at residual-vs-fitted).
"σ is the variance."
In dist.Normal(mu, sigma), sigma is the standard deviation, in the units of the data. The variance is $\sigma^2$. A prior such as HalfNormal(20) on sigma is a prior on a typical miss of around 20 orders, not 20 squared orders.
In a forecasting model where one global $\sigma$ describes the noise, every day gets the same interval half-width $\approx 1.645\sigma$ for 90%, in a quiet January as in a promotion-heavy December. A few extreme days inflate $\sigma$, and a larger $\sigma$ widens every future interval and loosens the fit of the trend and the changepoint slopes $\delta_j$ (they chase fewer residuals). In the A/B framework, a Normal likelihood for a revenue-like metric makes each group's mean sensitive to a few huge buyers; compare with the Student-t in the next sections. Always say which standard deviation you mean when you quote "noise level".
"We used a Normal likelihood, so we assumed the demand data are Normally distributed."
The Normal likelihood assumes the residuals (demand minus the model's expected demand) are independent Normal draws with a constant spread. Demand itself is a sum of trend, seasonality and events plus noise.
Model answer: "$y_t \sim N(\mu_t, \sigma^2)$ with $\mu_t$ from the trend, seasonality, holiday and regressor terms. Fitting it is least squares. I checked the residuals for symmetry, tails and constant variance, because those are the assumptions the Normal adds."
$y_t \sim N(\mu_t, \sigma^2)$; $\ell = -\tfrac T2\log(2\pi\sigma^2) - \tfrac1{2\sigma^2}\sum r_t^2$; $\hat\sigma^2 = \tfrac1T\sum r_t^2$. Fitting = least squares.
Promises: continuous · symmetric · light tails · constant variance · independent. NumPyro Normal(loc, scale) takes the sd.
Trap: it is the residuals, not the series, that must be Normal; one far day inflates $\hat\sigma$ (2.65 to 4.41 in the example) and widens every interval.
Quick check: residuals $3, -3, 0, 0$. What is the maximum-likelihood $\hat\sigma$ of a Normal likelihood?
Squares: $9, 9, 0, 0$, sum $18$. $\hat\sigma = \sqrt{18/4} = \sqrt{4.5} = 2.12$. (Dividing by $n-1 = 3$ would give $2.45$, the sample standard deviation.)
The Student-t likelihood: a bell that expects the occasional big surprise core
Real demand series have rare days that no layer predicts: a competitor's outage sends you extra customers, a payment glitch loses a day of orders, a post goes viral. These days land far from the bullseye. A Normal likelihood calls such a day nearly impossible and, to make it less impossible, bends the whole model toward it. The Student-t likelihood says instead: "mostly the usual small scatter, but now and then a big surprise happens, and that is part of the story".
A simple way to see it: imagine every day has its own noise width. Calm days have a small width and stormy days a large one. Mix many Normal bells of different widths together and the result has a normal-looking middle but fatter tails. That mixture is the Student-t. One dial, $\nu$ (the Greek letter "nu"), says how often the stormy days happen: small $\nu$ means stormy days are common (very heavy tails), large $\nu$ means they almost never happen (the Normal).
Three ways to say it:
- Picture: the same bell shape with fatter feet: the tails fade slowly instead of dying out.
- Numbers: a miss 4 typical sizes away has probability 0.000063 (1 in 15 800) under a Normal and 0.016 (1 in 62) under a Student-t with $\nu = 4$: about 250 times more likely.
- Slogan: a Normal that is allowed to be surprised.
Ten days of orders at a small shop, and a straight-line trend $y = a + b\cdot\text{day}$ to fit: $100, 103, 103, 107, 108, 111, 111, 115, 116, 160$ for days $0, \dots, 9$. The nine ordinary days follow $100 + 2\cdot\text{day}$ with small wiggles; day 9 should have been about 118 but a data glitch made it 160. Fit the line under each likelihood.
- Normal (least squares): $a = 94.15$, $b = 4.28$ orders per day, $\hat\sigma = 10.7$. The line is tilted up to chase the glitch. Its residuals on days 0 to 8 run $+5.9, +4.6, +0.3, 0.0, -3.3, -4.5, -8.8, -9.1, -12.4$: a clear pattern, and every one of them is a miss that exists only because of the glitch.
- Student-t with $\nu = 4$ (intercept, slope and scale fitted together by computer): $a = 100.23$, $b = 2.01$, scale $\hat\sigma = 1.06$. The residuals on the nine ordinary days are small (between $-1.3$ and $+0.8$), and day 9's residual is $41.7$, which is $41.7/1.06 = 39$ scale units.
- How does the fit treat each day? As a weighted least-squares fit: ordinary least squares in which each day $t$ gets a weight $w_t$, a number saying how much that day counts. Here $w_t = \dfrac{\nu+1}{\nu + r_t^2/\sigma^2}$. For an ordinary day with $r/\sigma = 1$: $w = 5/(4+1) = 1$. For day 9: $r^2/\sigma^2 = 1549$, so $w = 5/(4 + 1549) = 0.003$.
- So day 9 pulls the line about $0.003$ as hard as an ordinary day. Fit the nine ordinary days alone and you get $b = 2.00$, almost exactly the t answer ($2.01$). The Normal answer ($4.28$) is more than double.
- Cost under each model for day 9: the Normal pays a squared price $\tfrac12(r/\sigma)^2$ that would be 774 if it kept the line where the t put it and used the small scale the ordinary days call for, so it moves the line (and inflates $\sigma$) instead. The t pays a log price $\tfrac{\nu+1}{2}\log(1 + r^2/(\nu\sigma^2)) = 2.5\log(1+387) = 14.9$ and keeps the line where it is.
A Student-t likelihood says $y_t \mid \mu_t \sim \text{StudentT}(\nu, \mu_t, \sigma)$ with density
$$p(y_t \mid \mu_t,\nu,\sigma) = \frac{\Gamma\!\big(\tfrac{\nu+1}{2}\big)}{\Gamma\!\big(\tfrac\nu2\big)\sqrt{\nu\pi}\,\sigma}\Big(1 + \frac{(y_t-\mu_t)^2}{\nu\sigma^2}\Big)^{-\frac{\nu+1}{2}} .$$- $\mu_t$ is the centre (it comes from the layers), $\sigma$ is the scale (how wide the bell is), and $\nu \gt 0$ is the degrees of freedom, here simply a tail-heaviness dial. NumPyro:
dist.StudentT(df, loc, scale). - Relation to the Normal: as $\nu \to \infty$ it becomes $N(\mu_t, \sigma^2)$. At $\nu = 1$ it is the Cauchy (no mean). For $\nu \gt 2$ the standard deviation is $\sigma\sqrt{\nu/(\nu-2)}$ (for $\nu = 4$: $1.41\sigma$); for $\nu \le 2$ the variance is infinite.
- Fitting: the fit is a weighted least squares with weights $w_t = \dfrac{\nu+1}{\nu + r_t^2/\sigma^2}$ that depend on the residuals themselves, so the data decide how much each day counts. This is the same idea as robust centre estimation in Chapter 4.9, applied to a whole regression.
- $\nu$ is usually given a prior and learned, but the data can pin it down only loosely (see the demo).
| $\nu$ | $P(\lvert T\rvert \gt 4)$ in scale units | about 1 in | sd for scale 1 |
|---|---|---|---|
| 2 | 0.0572 | 17 | infinite |
| 4 | 0.0161 | 62 | 1.41 |
| 10 | 0.0025 | 397 | 1.12 |
| 30 | 0.00038 | 2 619 | 1.04 |
| $\infty$ (Normal) | 0.000063 | 15 787 | 1 |
Why do we need it?
A few unexplained shocks should not be able to bend your trend, inflate your noise level and widen every future interval. The Student-t keeps the middle of the data in charge and treats rare large misses as plausible tail events.
Where is it used?
Robust regression, the Student-t option of your forecasting model and A/B framework, financial return models (returns have fat tails), Student-t output heads in neural forecasters, and Kalman filters that must survive sensor glitches.
How is it used?
Replace dist.Normal(mu, sigma) by dist.StudentT(nu, mu, sigma), give nu a prior (or fix it), refit and compare. Look at the posterior of $\nu$: small means heavy tails are needed; large means the Normal was fine. Check that the trend and holiday effects stopped moving with single days.
"$\nu$ is the degrees of freedom, so it must be $n - 1$."
That is the role of $\nu$ in a t-test. In a t likelihood, $\nu$ is a free tail-heaviness dial, learned from the residuals or fixed. It has nothing to do with how many days you have.
"The t scale $\sigma$ is the noise standard deviation."
The standard deviation is $\sigma\sqrt{\nu/(\nu-2)}$ for $\nu \gt 2$ and infinite otherwise. Two models with the same $\sigma$ but different $\nu$ have different noise levels, and a prior on $\sigma$ means something different from a prior on a standard deviation.
"A small $\nu$ is always safer."
If the noise is really close to Normal, a small $\nu$ wastes statistical power and gives overly generous extreme-quantile intervals (99.9% bands are much wider). The data decide; and as the demo shows, they can rarely pin down large $\nu$.
Your forecasting model and your A/B framework both list a Student-t option next to the Normal. In the forecasting model, if you give $\nu$ a prior and learn it, its posterior is itself a diagnostic: mass on small values says "this series has heavy-tailed residuals", mass spread over large values says "the Normal was fine". Because $\nu$ and the scale $\sigma$ trade off against each other, expect them to be correlated in the posterior; this is one reason a full-rank or low-rank guide (Chapter 6.13) can describe the posterior better than a fully independent one. For a metric in an A/B test, a Student-t group-mean estimate is less swayed by a handful of very large users.
$y_t \sim \text{StudentT}(\nu, \mu_t, \sigma)$; $\nu \to \infty$ gives the Normal; sd $= \sigma\sqrt{\nu/(\nu-2)}$ for $\nu \gt 2$. NumPyro StudentT(df, loc, scale).
Fit = weighted least squares with $w_t = \frac{\nu+1}{\nu + r_t^2/\sigma^2}$: ordinary days $\approx 1$, an extreme day $\approx 0$ (0.003 in the example).
Trap: $\nu$ is a tail dial, not $n-1$; $\sigma$ is not the sd; the data can rarely tell $\nu = 30$ from $\nu = 200$.
Quick check: $\nu = 4$, scale $\sigma = 2$. What weight does a day with residual $r = 10$ get, and what is the sd of this noise?
$r/\sigma = 5$, so $w = (4+1)/(4 + 25) = 5/29 = 0.17$, about a sixth of an ordinary day. The sd is $\sigma\sqrt{\nu/(\nu-2)} = 2\sqrt{2} = 2.83$.
Outliers and influence: the Student-t does not remove them, it lets them pull less core
In a tug of war, each day pulls the fitted line toward itself with a rope. Under a Normal likelihood the rope gets stronger the farther the day is from the line. A day 40 orders away pulls 40 times as hard as a day 1 order away, so a single glitch wins the tug of war against nine honest days.
Under a Student-t likelihood the rope strengthens only up to a point and then goes slack. A day that is far enough away pulls hardly at all. But it is still in the tug of war: its distance is still counted, it still adds to the log-likelihood, it still teaches the model that big surprises exist ($\nu$ gets smaller, the predictive tails get fatter). Nothing was deleted. That is why the correct sentence is "the Student-t assigns more probability to extreme residuals, so they exert less influence", and the incorrect one is "the Student-t removes outliers".
Three ways to say it:
- Picture: a rope for each day: Normal ropes tighten with distance; t ropes go slack when stretched too far.
- Numbers: in the shop example the glitch day has weight 0.003 under the t, so it moves the slope from 2.00 to 2.01, while under the Normal it moves it from 2.00 to 4.28.
- Slogan: less influence, not zero presence.
Same ten days as before ($100, 103, 103, 107, 108, 111, 111, 115, 116, 160$). Turn the dial $\nu$ and watch the fitted slope (true value 2 orders per day) and the glitch day's weight. Every row is a fit of $a + b\cdot\text{day}$ with $t_\nu$ errors by maximum likelihood.
| Likelihood | slope $b$ | scale $\hat\sigma$ | weight of day 9 |
|---|---|---|---|
| Student-t, $\nu = 2$ | 2.01 | 0.83 | 0.001 |
| Student-t, $\nu = 4$ | 2.01 | 1.06 | 0.003 |
| Student-t, $\nu = 10$ | 2.61 | 5.81 | 0.209 |
| Student-t, $\nu = 30$ | 3.94 | 9.95 | 0.800 |
| Normal ($\nu = \infty$) | 4.28 | 10.7 | 1 |
- The weight of day 9 is $w = \dfrac{\nu+1}{\nu + r^2/\sigma^2}$, with $r$ its residual under that fit. At $\nu = 4$: $r = 41.7$, $\sigma = 1.06$, so $w = 5/(4 + 1549) = 0.003$.
- The glitch is still there. Under the $\nu = 4$ fit it is 39 scale units from the line, and its log-probability is $-15.9$, against about $-1.3$ for an ordinary day. The model finds it extremely surprising; it simply refuses to bend the line for it.
- Leave day 9 out and refit the other nine days: the t slope changes from 2.011 to 2.000 (a shift of 0.011), the Normal slope from 4.279 to 2.000 (a shift of 2.28). That shift is the influence of day 9.
- For moderate $\nu$ (10, 30) the t is only partly robust: it still gives day 9 a weight of 0.2 to 0.8. Robustness is not a switch; it comes from small $\nu$ and a good fit of the scale.
Fitting by maximum likelihood solves $\sum_t \psi(r_t) \cdot (\text{regressor of day } t) = 0$, where $\psi(r) = \dfrac{d}{dr}\big(-\log p(r)\big)$ is the influence (pull) of a residual $r$. For residuals measured in scale units:
| Likelihood | Pull $\psi(r)$ of a residual $r$ | What happens to a far day |
|---|---|---|
| Normal | $r$ | pull grows without limit |
| Student-t, $\nu$ | $\dfrac{(\nu+1)\,r}{\nu + r^2}$ | largest at $\lvert r\rvert = \sqrt\nu$, then falls back toward 0 |
- The t pull is $w(r)\cdot r$ with the weight $w(r) = \frac{\nu+1}{\nu+r^2}$ of the previous section, so "less influence" is the weight shrinking faster than the residual grows.
- An estimator is robust when no single observation can move it arbitrarily far. Robust statistics in general are in Chapter 4.14.
- What the t does not do: it does not detect errors, does not delete or cap the data, and does not know whether the far day is a data bug (to be fixed) or a real event (to be modelled with a holiday or regressor, Chapter 7.12).
Why do we need it?
Seeing the mechanism (weights and pulls) is what lets you answer follow-up questions honestly: why the trend stopped moving, why the noise scale shrank, why the intervals still have fat tails, and when the t would not help.
Where is it used?
Robust regression (Huber, Tukey, Student-t), iteratively reweighted least squares, outlier-resistant Kalman and trend filters, and every model review where someone asks "what did you do about outliers?".
How is it used?
Fit both likelihoods and compare the trend and holiday estimates. Print the t weights per day ((nu+1)/(nu + (r/sigma)**2) from posterior means) and list the days with weight below, say, 0.3. Investigate those days: bug, real event, or genuine tail?
"The Student-t removes the outliers."
Nothing is removed. Every day stays in the likelihood and keeps a weight; the extreme ones just get a very small weight. The extreme day also still affects the scale, $\nu$ and the width of the predictive tails.
"The Student-t ignores everything beyond $3\sigma$."
There is no cut-off. The weight $(\nu+1)/(\nu + r^2/\sigma^2)$ falls smoothly; a day at $2\sigma$ is already slightly down-weighted, and one at $40\sigma$ is almost, but never exactly, ignored.
"If the t gives a different answer, the Normal was wrong, so delete the strange days."
First ask what the strange days are. A data bug should be fixed at the source. A real, repeating event (a sale, an outage) belongs in the model as a holiday or regressor; a Student-t would quietly down-weight exactly the days you should be explaining.
"We used a Student-t likelihood to remove outliers."
"The Student-t ignores outliers."
"A Student-t likelihood assigns more probability to extreme residuals, so they exert less influence on the fit than under a Normal likelihood."
Model answer: "Under a Normal likelihood a residual's pull grows linearly with its size, so one glitch can tilt the trend. Under a Student-t the pull is $(\nu+1)r/(\nu+r^2)$ in scale units: it rises, peaks and falls back toward zero, so a far-away day gets a tiny weight, 0.003 in my toy example. The day is not removed. It is still in the likelihood and still shapes $\nu$ and the width of the forecast tails. I would still look at those days, because if they are real events they should be modelled, not down-weighted."
This is exactly why your forecasting model can switch from a Normal to a Student-t likelihood: with a Normal one corrupted or unmodelled day can bend the piecewise-linear trend, shift a slope change $\delta_j$, inflate a holiday coefficient and widen all intervals. In an A/B framework like yours, the same mechanism protects a group mean for a heavy-tailed continuous metric. To audit a fit, compute the t weights from the posterior means, list the lowest-weight days, and ask for each: bug, real event or genuine tail?
Pull of a residual: Normal $\psi = r$ (unbounded); Student-t $\psi = \frac{(\nu+1)r}{\nu+r^2}$ (rises, peaks at $\sqrt\nu$, falls). Weight $w = \frac{\nu+1}{\nu+r^2/\sigma^2}$.
Example: glitch day weight 0.003 (t, $\nu=4$): slope 2.01; Normal slope 4.28; true 2.
Say it: "assigns more probability to extreme residuals, so they exert less influence". Never "removes" or "ignores" outliers.
Quick check: with $\nu = 4$ and scale 1, which has the larger pull, a residual of 2 or a residual of 20?
$\psi(r) = 5r/(4 + r^2)$. At $r = 2$: $10/8 = 1.25$. At $r = 20$: $100/404 = 0.25$. The far day pulls five times less, even though it is ten times farther. Under a Normal the pulls would be 2 and 20.
Counts are not bell-shaped: why a Normal likelihood fails for small counts core
A count is "how many": orders, sign-ups, support tickets, items sold. It is a whole number and it can never be below zero. Now picture a shop that sells about 3 items a day. Some days it sells 0, 1 or 2, a good day sells 6 or 8, and it can never sell $-2$. The numbers pile up against a wall at zero, with a long tail stretching to the right. A symmetric bell cannot sit against a wall: put its middle at 3 and a good part of it hangs over the edge, onto days with negative sales.
There is a second, quieter problem. For counts, busy days are noisier than quiet days in absolute terms: a day with 100 expected orders swings by many more orders than a day with 3. The Normal gives every day the same $\sigma$.
The good news: when counts are large (hundreds), the wall is far away and the bell is a fine description of the shape. The trouble is with small counts, and with the noise that depends on the level.
Three ways to say it:
- Picture: a bell squeezed against a wall: the part past the wall is impossible, and the leftover piles up on a spike at zero.
- Numbers: mean 3, variance 6. A Normal puts 11% of its probability below zero and gives the 90% interval $[-1.03,\ 7.03]$. A Negative Binomial with the same mean and variance gives $P(0) = 12.5\%$ and the interval $[0,\ 8]$.
- Slogan: counts live on $0, 1, 2, \dots$; a Normal lives on the whole line.
A small shop averages $\mu = 3$ orders a day, with variance 6 (noisier than a Poisson, whose variance would be 3). Compare a Normal likelihood and a Negative Binomial with the same mean and variance ($\alpha = 3$, because $3 + 3^2/3 = 6$).
- Standard deviation: $\sqrt6 = 2.45$. The Normal is $N(3,\ 2.45^2)$.
- Probability of negative orders under the Normal: $z = (0 - 3)/2.45 = -1.22$, so $P(y \lt 0) = 0.110$. Counting only values that would round to $-1$ or below ($y \le -0.5$): $z = -1.43$, $P = 0.077$. Either way, between 8% and 11% of the Normal's probability sits on impossible days.
- 90% interval under the Normal: $3 \pm 1.645 \times 2.45 = [-1.03,\ 7.03]$. A forecast with a negative lower end.
- Negative Binomial: $P(0) = \big(\tfrac{\alpha}{\alpha+\mu}\big)^\alpha = (3/6)^3 = 0.125$. Its 5%, 50% and 95% quantiles are 0, 2 and 8: the interval $[0, 8]$ is a set of real, possible counts, and the median (2) is below the mean (3), the sign of a right-skewed distribution.
- Upper tail: $P(y \ge 8)$ is 0.055 under the Negative Binomial but only 0.033 under the Normal (using $y \gt 7.5$). The Normal also under-estimates the chance of a very busy day.
- Now a big shop: $\mu = 100$, $\alpha = 20$ (variance 600). Normal 90% interval: $[59.7,\ 140.3]$; Negative Binomial: $[63,\ 143]$. Close. The Normal is acceptable when counts are large and the spread is modest.
Four ways in which a Normal likelihood does not fit count data:
- Support. The Normal allows negative and fractional values; counts are in $\{0, 1, 2, \dots\}$.
- Shape. Small counts are right-skewed with a pile-up at 0 (or a spike at 0 and 1); the Normal is symmetric.
- Variance depends on the mean. For counts, $Var(y)$ grows with $\mu$ (Poisson: $\mu$; Negative Binomial: $\mu + \mu^2/\alpha$); the Normal's $\sigma$ is the same on every day.
- Zeros. The Normal gives $P(y = 0) = 0$ as a density; a count model gives zero a real probability.
Your options, in order of preference when counts are small: (a) a count likelihood (Poisson or Negative Binomial) with a positive mean from a link (a later section); (b) a Normal on a transformed count (log1p or square root, Chapter 4.18), remembering that back-transforming the mean gives something closer to a median; (c) a plain Normal when counts are large (a rule of thumb: expected counts above about 20 to 30, and the day-to-day spread roughly constant). A Poisson with mean $\mu$ is approximately $N(\mu, \mu)$ when $\mu$ is large.
Why do we need it?
For small counts a Normal forecast reports impossible values, hides the chance of zero and gets the width wrong. Knowing exactly where it breaks tells you when you may still use it (large counts, stable spread) and when you must not.
Where is it used?
Demand for slow-moving products, hourly or per-segment order counts, support tickets, sign-ups and clicks per day, rare-event series, and the count likelihoods of GLMs (Chapter 5.14) and of your forecasting model.
How is it used?
Look at the series: is it whole numbers, how small, how many zeros? If yes, choose Poisson or Negative Binomial with a positive-mean link. If counts are large and smooth, a Normal may be fine; check the lower end of the 90% interval and the residual-vs-fitted plot.
"If the mean is 3, a Normal likelihood with $\sigma = 2.5$ is a fine model; I can clip negative forecasts at zero."
Clipping repairs the printed number, not the model: the fit, the interval widths and the chance of exceeding capacity are still computed from a distribution that put 8–11% of its mass on impossible days and too little on busy ones.
"Counts are large in my business, so the count likelihood never matters."
With large counts the shape problem disappears, but the variance-depends-on-the-mean problem does not: a busy weekend is noisier in orders than a quiet Tuesday. The Negative Binomial and the log link handle both (next sections).
In your forecasting model, the Negative Binomial option exists for count-valued series. If you forecast something like orders per hour, per product or per small segment, a Normal likelihood would put probability on negative orders and use one $\sigma$ for every hour. In the A/B framework, the Poisson likelihood for count metrics has the same limitation in a different form: it forces variance equal to the mean (the next section).
Counts: whole numbers $\ge 0$, right-skewed near 0, variance grows with the mean. A Normal ignores all three.
Example: $\mu = 3$, $Var = 6$: Normal puts 11% below 0, interval $[-1.03, 7.03]$; NB gives $P(0) = 12.5\%$, interval $[0, 8]$.
OK to use a Normal when counts are large (rule of thumb: above about 20–30) and the spread is stable; otherwise Poisson or NB.
Quick check: a Normal with mean 2 and sd 2 is used for daily tickets. Roughly what share of its probability is below zero?
$z = (0 - 2)/2 = -1$, and $\Phi(-1) = 0.159$: about 16% of the probability is on negative ticket counts. That is a large share of the forecast sitting on impossible values.
Poisson is too restrictive: overdispersion, and why it happens core
A Poisson distribution has one dial, the mean $\lambda$, and its spread is glued to it: variance = mean. That is what you get when events happen independently at a perfectly steady rate, like raindrops on a roof in steady drizzle.
Real daily counts are rarely that calm, because the rate itself changes from day to day in ways your layers do not capture: the weather, a mention on social media, which customers happened to visit, a slow payment page. Every day secretly has its own rate, and given that rate the count is Poisson. But all you see is the counts, and across days they spread out more than any single steady rate would allow. That extra spread is called overdispersion ("over" = more than Poisson, "dispersion" = spread).
Three ways to say it:
- Picture: raindrops under a wobbling tap: the average flow is fixed, but some minutes it runs faster and some slower, so the drop counts vary more than steady rain.
- Numbers: eight Saturdays of orders: $52, 38, 61, 45, 30, 58, 47, 41$ have mean $46.5$ but variance $107$, about $2.3$ times the mean. A Poisson would say the variance is about $46.5$.
- Slogan: Poisson says variance equals mean; real counts say variance is bigger.
Orders on eight consecutive Saturdays (same weekday, so the expected value is about the same): $52, 38, 61, 45, 30, 58, 47, 41$.
- Mean: $(52+38+61+45+30+58+47+41)/8 = 372/8 = 46.5$.
- Squared distances from the mean: $5.5^2 = 30.25$, $8.5^2 = 72.25$, $14.5^2 = 210.25$, $1.5^2 = 2.25$, $16.5^2 = 272.25$, $11.5^2 = 132.25$, $0.5^2 = 0.25$, $5.5^2 = 30.25$. Sum $= 750$.
- Sample variance: $750/(8-1) = 107.1$. The Poisson promise is variance $\approx$ mean $= 46.5$. The dispersion index is $107.1/46.5 = 2.30$.
- A Poisson with mean 46.5 says 90% of Saturdays fall in $[36, 58]$. Two of the eight (30 and 61) are outside: 25% instead of the 10% expected.
- Fit a Negative Binomial by matching moments: $\alpha = \mu^2/(Var - \mu) = 46.5^2/(107.1 - 46.5) = 35.7$. Its 90% interval is $[31, 64]$, which contains all eight days.
- Why does this happen? If the rate varies from Saturday to Saturday with some variance $\tau^2$, then total variance = average Poisson variance + variance of the rates (the law of total variance, Chapter 4.6): $Var(y) = \mu + \tau^2$. Here $\tau^2 \approx 60.6$, so the rate itself wobbles with sd about 7.8 orders.
A count model is overdispersed when $Var(y_t \mid \mu_t) \gt \mu_t$. Two practical tools:
- Pearson dispersion index: $\hat\phi = \dfrac{1}{n-p}\sum_t \dfrac{(y_t - \hat\mu_t)^2}{\hat\mu_t}$, where $p$ is the number of fitted parameters in $\mu_t$. About 1 means "Poisson is fine"; clearly above 1 means overdispersed (below 1 is underdispersed, which is rarer).
- Variance-against-mean plot: group days with similar expected values, plot each group's sample variance against its sample mean. Poisson data lie on the line $y = x$; Negative Binomial data lie on the curve $y = x + x^2/\alpha$.
The gamma-Poisson mixture: if the daily rate is $\lambda_t \sim \text{Gamma}(\text{shape } \alpha,\ \text{rate } \alpha/\mu)$ (mean $\mu$, variance $\mu^2/\alpha$) and $y_t \mid \lambda_t \sim \text{Poisson}(\lambda_t)$, then $y_t$ is Negative Binomial with $E[y_t] = \mu$ and $Var(y_t) = \mu + \mu^2/\alpha$. (Poisson, Negative Binomial and this mixture are taught in Chapter 4.8.)
Why do we need it?
A Poisson on overdispersed counts gives intervals that are far too narrow and overconfident probabilities of hitting capacity, and its standard errors for the layers' coefficients are too small, so trends and holidays look more certain than they are.
Where is it used?
Count regressions (glm.nb in R, statsmodels NegativeBinomial, scikit-learn's PoissonRegressor is the Poisson-only version), demand and web-traffic forecasting, epidemic counts, insurance claims, and the A/B framework's Poisson likelihood for count metrics.
How is it used?
Fit the Poisson (or take the layers' fitted means), compute $\hat\phi$ and the variance-against-mean plot by weekday or level. If $\hat\phi \gg 1$, switch to a Negative Binomial. After fitting, a posterior predictive check on the variance (Chapter 7.14) confirms the fix.
"Overdispersion means there are outliers."
Not necessarily. It means the variance is bigger than the mean everywhere, even without a single outlier, because the rate wobbles. (Outliers are a heavy-tail issue, the Student-t's job for continuous data.)
"My fitted Poisson mean is right, so the Poisson is fine."
The mean can be perfect while the variance is wrong by a factor of 2 or 10. Compute the dispersion index or plot variance against mean before trusting the intervals.
"If the data are overdispersed, add more regressors until it goes away."
Real missing structure (a missed holiday, a missing seasonality) does inflate dispersion, and adding it helps. But some day-to-day wobble stays however good the layers are; a Negative Binomial absorbs it honestly.
In an A/B framework like yours, a Poisson likelihood for a count metric (for example support tickets per user) is the simplest choice. If the metric is overdispersed, the Poisson posterior for each group's rate is too narrow and $P(\theta_B \gt \theta_A \mid D)$ becomes overconfident. In your forecasting model, daily counts are almost always overdispersed after trend, seasonality and holidays, so the Negative Binomial's extra parameter is not a luxury: it sets how wide the forecast intervals are.
Poisson: $Var = \mu$. Overdispersion: $Var \gt \mu$. Index $\hat\phi = \frac1{n-p}\sum \frac{(y-\hat\mu)^2}{\hat\mu}$ (1 = Poisson).
Cause: the rate wobbles. Gamma rate + Poisson count = Negative Binomial; $Var = \mu + \mu^2/\alpha$ (total variance = Poisson part + rate part).
Example: Saturdays mean 46.5, variance 107.1, index 2.3, $\hat\alpha \approx 35.7$.
Quick check: four Mondays of orders: $6, 14, 5, 15$. Is a Poisson reasonable? Give the index and a moment estimate of $\alpha$.
Mean $= 40/4 = 10$. Squared distances: $16, 16, 25, 25$, sum $82$; variance $= 82/3 = 27.3$. Index $= 27.3/10 = 2.7$, well above 1 (with only four values this is a rough signal, but a warning). $\hat\alpha = 10^2/(27.3 - 10) = 5.8$.
The Negative Binomial for forecasts: mean $\mu$, concentration $\alpha$, $Var = \mu + \mu^2/\alpha$ core
The Negative Binomial is "a Poisson whose rate is shaken a little every day". It has two dials. The mean $\mu$ is where the day is expected to land (this is what your layers produce). The concentration $\alpha$ says how steady the daily rate is: a big $\alpha$ means the rate hardly moves and the counts look Poisson; a small $\alpha$ means the rate swings wildly and the counts are very spread out.
The important consequence for forecasting: the spread grows faster than the mean. The variance has a Poisson part ($\mu$) and a rate-wobble part ($\mu^2/\alpha$) that grows with the square of the level. So a busy day is noisier, in absolute terms, than a quiet day, and in relative terms the noise settles at a floor of about $1/\sqrt\alpha$ (a Poisson's relative noise would keep shrinking as the level grows).
Three ways to say it:
- Picture: a Poisson cloud whose size also wobbles; the busier the day, the bigger the wobble.
- Numbers: mean 40, $\alpha = 5$: variance $40 + 1600/5 = 360$, sd 19. A Poisson with mean 40 has sd 6.3.
- Slogan: Poisson part + wobble part: $Var = \mu + \mu^2/\alpha$.
A busy-day forecast: expected orders $\mu = 40$, concentration $\alpha = 5$.
- Variance: $\mu + \mu^2/\alpha = 40 + 1600/5 = 40 + 320 = 360$. Standard deviation $= 18.97$. A Poisson with the same mean has variance 40 and sd $6.32$: three times narrower.
- 90% interval: Negative Binomial $[14,\ 75]$; Poisson $[30,\ 51]$.
- How likely is a day with 90 or more orders? Poisson: $7.8\times10^{-12}$ (never). Negative Binomial: $0.0169$, about 1 day in 59.
- The mixture story behind the numbers: the day's rate is $\lambda \sim \text{Gamma}(\text{shape } 5,\ \text{rate } 5/40 = 0.125)$, which has mean $5/0.125 = 40$ and sd $\sqrt5/0.125 = 17.9$. Then the count is Poisson($\lambda$). Total variance $= 40 + 17.9^2 = 40 + 320 = 360$. ✓.
- Relative noise (standard deviation divided by the mean): $\sqrt{1/\mu + 1/\alpha}$. At $\mu = 40$: Poisson $0.158$, Negative Binomial $0.474$. At $\mu = 400$: Poisson $0.05$, but Negative Binomial $\sqrt{1/400 + 1/5} = 0.45$. The wobble term does not fade with level.
- The dial $\alpha$ at $\mu = 40$: $\alpha = 5$ gives sd 19.0; $\alpha = 40$ gives sd 8.9; $\alpha = 1000$ gives sd 6.4, practically the Poisson's 6.3.
The Negative Binomial in mean–concentration form (called NB2) has probability mass function
$$P(y \mid \mu,\alpha) = \frac{\Gamma(y+\alpha)}{y!\,\Gamma(\alpha)}\Big(\frac{\alpha}{\alpha+\mu}\Big)^{\alpha}\Big(\frac{\mu}{\alpha+\mu}\Big)^{y},\qquad y = 0, 1, 2, \dots$$ $$E[y] = \mu,\qquad Var(y) = \mu + \frac{\mu^2}{\alpha}.$$- $\alpha \to \infty$ gives the Poisson($\mu$). Small $\alpha$ gives a very wide, spiky distribution. $P(y=0) = (\alpha/(\alpha+\mu))^\alpha$.
- In a forecasting model: $y_t \sim \text{NB2}(\mu_t, \alpha)$, with $\mu_t$ built from the layers through a positive link (a later section) and one $\alpha$ shared by all days. NumPyro:
dist.NegativeBinomial2(mean=mu_t, concentration=alpha). - "NB2" is the name for variance quadratic in the mean. Some software also offers "NB1", with variance $\phi\mu$ (linear in the mean). They are different models; this guide always means NB2.
- $\alpha$ needs a prior on positive numbers. As with $\nu$ in the Student-t, the data pin down small values of $\alpha$ well and large values only loosely.
Why do we need it?
Daily business counts are overdispersed and their noise grows with the level. The Negative Binomial gives both with one extra parameter, so forecast bands are wider on busy days and the probability of an extreme day is realistic.
Where is it used?
The Negative Binomial option of your forecasting model, NB regression in statsmodels and R (glm.nb), RNA-seq analysis (DESeq2), epidemic and insurance counts, and DeepAR-style neural forecasters that output a Negative Binomial for each time step.
How is it used?
Put the layers' mean in the first slot and a learned positive $\alpha$ in the second. After fitting, read $\hat\alpha$ next to your typical $\mu$: $Var/\mu = 1 + \mu/\hat\alpha$, so $\hat\alpha$ much larger than $\mu$ means nearly Poisson, and the relative noise never falls below $1/\sqrt{\hat\alpha}$ (for example $\hat\alpha = 5$ means at least 45% day-to-day noise). Check the variance and zero-count statistics in a posterior predictive check.
"A bigger $\alpha$ means more overdispersion."
The opposite: smaller $\alpha$ means more overdispersion (the extra variance is $\mu^2/\alpha$). $\alpha \to \infty$ is the Poisson. Some libraries and papers use the reciprocal $1/\alpha$ and call it "dispersion", for which bigger does mean more noise. Always quote the variance formula.
"The Negative Binomial variance is $\alpha\mu$" or "$\mu(1+\alpha)$".
Those are other parameterizations (NB1-style or reciprocal forms). For the mean–concentration NB2 used here the variance is $\mu + \mu^2/\alpha$.
"One $\alpha$ for the whole series means the relative noise is the same every day."
Close, but not exactly. The relative noise is $\sqrt{1/\mu_t + 1/\alpha}$: nearly constant for busy days (about $1/\sqrt\alpha$), larger for very quiet days where the Poisson part $1/\mu_t$ matters.
In your forecasting model a learned $\alpha$ is the knob that sets how wide the count forecast intervals are. A fitted $\hat\alpha$ near 3 says daily counts are very noisy even after trend, seasonality and holidays; a very large $\hat\alpha$ (much bigger than your typical daily count $\mu$) says the Poisson would have done. Note that the model's other uncertainty (the posterior over the layers' weights) sits on top of this noise: the full predictive distribution adds both (Chapter 7.14). If you also put a prior on $\alpha$, remember that the "same" weak prior behaves differently on $\alpha$ and on $1/\alpha$.
NB2: $E[y] = \mu$, $Var(y) = \mu + \mu^2/\alpha$, $P(0) = (\frac{\alpha}{\alpha+\mu})^\alpha$; $\alpha \to \infty$ is Poisson; small $\alpha$ is noisy.
Gamma-Poisson: rate $\sim \text{Gamma}(\alpha, \alpha/\mu)$, count $\sim \text{Poisson}$. Relative noise $\sqrt{1/\mu + 1/\alpha}$ never goes below $1/\sqrt\alpha$.
Example: $\mu = 40$, $\alpha = 5$: sd 19.0, 90% interval $[14, 75]$ (Poisson $[30, 51]$).
Quick check: $\mu = 100$ and $\alpha = 25$. What are the variance and sd, and how do they compare with a Poisson?
$Var = 100 + 10000/25 = 100 + 400 = 500$, sd $= 22.4$. A Poisson with mean 100 has sd 10: the Negative Binomial is 2.2 times as wide.
Know the exact parameterization you used: one distribution, many argument lists core
Your syllabus says it in one line: libraries differ, know the exact NumPyro parameterization you used. Here is why it matters so much for forecasting. A Negative Binomial can be written with the mean and the concentration, or with a "number of successes" and a "success probability", or with a log-odds, or as a Gamma-Poisson with a rate, or (in statsmodels) with a dispersion that is the reciprocal of the concentration. They are the same family but the numbers you pass are different, and one library's "probability" is another's "one minus probability".
The danger is silent. If you pass a SciPy-style number into a NumPyro-style slot, the code runs, the model fits, the plots look reasonable, and the model has quietly changed to a different distribution, often with the wrong mean. This chapter's first rule of safe use: after you build the distribution, print its mean and variance and compare with $\mu$ and $\mu + \mu^2/\alpha$.
The same care applies to the Normal and Student-t: NumPyro's Normal(loc, scale) takes the standard deviation, and StudentT(df, loc, scale) takes a scale that is not the standard deviation.
Three ways to say it:
- Picture: one person with six name tags in six languages; do not call them by the wrong one.
- Numbers: mean 40, concentration 5 is
NegativeBinomial2(40, 5)in NumPyro andnbinom(5, 1/9)in SciPy. Pass SciPy's $1/9$ as NumPyro'sprobsand the mean becomes $0.625$. - Slogan: translate through $(\mu, \alpha)$, then check the mean and variance.
Target: mean $\mu = 40$ orders and concentration $\alpha = 5$ (variance $40 + 1600/5 = 360$). The arguments for each library, each checked to produce exactly this distribution:
- NumPyro
NegativeBinomial2(mean=40, concentration=5): written directly. - NumPyro
GammaPoisson(concentration=5, rate=0.125): rate $= \alpha/\mu = 5/40$. - NumPyro
NegativeBinomialProbs(total_count=5, probs=0.8889):probs$= \mu/(\alpha+\mu) = 40/45$. Its mean is $5 \times 0.8889/0.1111 = 40$. - NumPyro
NegativeBinomialLogits(total_count=5, logits=2.079): logits $= \log(\mu/\alpha) = \log 8$. - SciPy
stats.nbinom(n=5, p=0.1111)(and NumPy'srng.negative_binomial(5, 0.1111)): $p = \alpha/(\alpha+\mu) = 5/45$, the complement of NumPyro'sprobs. Mean $n(1-p)/p = 5 \times 0.8889/0.1111 = 40$. - statsmodels NB2:
alpha$= 1/5 = 0.2$, and the variance is $\mu + \text{alpha}\cdot\mu^2 = 40 + 0.2 \times 1600 = 360$. - Two silent mistakes. SciPy's $p = 0.1111$ passed to NumPyro's
probs: mean becomes $5 \times 0.1111/0.8889 = 0.625$. statsmodels'alpha = 0.2passed as NumPyro'sconcentration: the mean is still 40, but the variance is $40 + 1600/0.2 = 8040$, which is 22 times too big.
Translation table from the master pair $(\mu, \alpha)$ (NB2, $Var = \mu + \mu^2/\alpha$). The full derivations and a longer widget are in Chapter 4.8.
| Where | Call | Arguments for $\mu = 40,\ \alpha = 5$ | Mean from the arguments |
|---|---|---|---|
| NumPyro | NegativeBinomial2(mean, concentration) | $(\mu, \alpha) = (40, 5)$ | mean |
| NumPyro | GammaPoisson(concentration, rate) | $(\alpha, \alpha/\mu) = (5, 0.125)$ | concentration / rate |
| NumPyro | NegativeBinomialProbs(total_count, probs) | $(\alpha, \frac{\mu}{\alpha+\mu}) = (5, 0.8889)$ | total_count · probs / (1 − probs) |
| NumPyro | NegativeBinomialLogits(total_count, logits) | $(\alpha, \log\frac\mu\alpha) = (5, 2.079)$ | total_count · elogits |
| NumPyro | NegativeBinomial(total_count, probs=…, logits=…) | a factory: it returns the Probs or Logits class above | as above |
| SciPy / NumPy | stats.nbinom(n, p), rng.negative_binomial(n, p) | $(\alpha, \frac{\alpha}{\alpha+\mu}) = (5, 0.1111)$ | $n(1-p)/p$ |
| statsmodels | NB2 alpha | $1/\alpha = 0.2$ | comes from the regression |
| Stan / R | neg_binomial_2(mu, phi), glm.nb theta | $\phi = \theta = \alpha = 5$ | $\mu$ |
- Check recipe in NumPyro:
d = dist.NegativeBinomial2(mu, alpha); print(d.mean, d.variance). In SciPy:d.mean(), d.var(). Both must equal $\mu$ and $\mu + \mu^2/\alpha$. - "Dispersion parameter" is ambiguous: say whether you mean $\alpha$ (large = close to Poisson) or $1/\alpha$ (large = noisier).
- The same warning for the other likelihoods:
dist.Normal(loc, scale)takes the standard deviation;dist.StudentT(df, loc, scale)takes the scale (sd $= \text{scale}\times\sqrt{\nu/(\nu-2)}$).
Why do we need it?
Models get moved between tools: explored in statsmodels, simulated in NumPy or SciPy, deployed in NumPyro. A wrong translation changes the mean or the variance without an error message, and a forecast with the wrong variance has the wrong intervals and the wrong probability of exceeding capacity.
Where is it used?
Every NB likelihood in code: NumPyro forecasting and experiment models, SciPy tests of simulated data, statsmodels count regressions, Stan and R code in papers you may reproduce, and any discussion comparing "the dispersion" fitted by two tools.
How is it used?
Keep $(\mu, \alpha)$ as the master form, convert with the table, then verify the implied mean and variance numerically. When reporting a fitted value, say which parameter ("concentration $\hat\alpha = 4.2$, so $Var = \mu + \mu^2/4.2$").
"probs in NumPyro and p in SciPy are the same thing."
For the Negative Binomial they are complements: NumPyro's probs $= \mu/(\alpha+\mu)$, SciPy's p $= \alpha/(\alpha+\mu)$. Swapping them turns the mean 40 into 0.625 with no error.
"I fitted alpha = 0.2 in statsmodels; I'll set the NumPyro concentration to 0.2."
statsmodels' alpha is the reciprocal of the concentration: use $1/0.2 = 5$.
"The Student-t scale and the Normal scale both mean standard deviation."
Only the Normal's. The Student-t's standard deviation is scale * sqrt(df / (df - 2)) for df $\gt 2$.
"We used a Negative Binomial likelihood with dispersion 2."
"We used an NB2 likelihood: NegativeBinomial2(mean=mu_t, concentration=alpha), so $Var(y_t) = \mu_t + \mu_t^2/\alpha$. A larger $\alpha$ means closer to Poisson."
Model answer: "The Negative Binomial has several parameterizations: SciPy uses (n, p) counting failures, NumPyro offers mean–concentration, total_count–probs and total_count–logits, and statsmodels' alpha is the reciprocal of the concentration. I keep the model in mean–concentration form, and I verify the implied mean and variance in code after building the distribution."
Open your forecasting model's likelihood code and find the line that builds the Negative Binomial. Which class is it? What exactly is passed in each slot, and is the mean computed by a link (exp or softplus)? If it uses NegativeBinomialProbs or Logits, confirm the direction of the probability with a quick numeric test. Also find where the prior on the dispersion sits (on $\alpha$ or on $1/\alpha$) and what the posterior mean of that parameter means in words. These are the first follow-up questions an interviewer asks after "we used a Negative Binomial".
Master form: $(\mu, \alpha)$, $Var = \mu + \mu^2/\alpha$. NumPyro: NegativeBinomial2(mean, concentration); Probs: $(\alpha, \frac{\mu}{\alpha+\mu})$; Logits: $(\alpha, \log\frac\mu\alpha)$; GammaPoisson: $(\alpha, \alpha/\mu)$.
SciPy nbinom(n, p): $(\alpha, \frac\alpha{\alpha+\mu})$. statsmodels alpha $= 1/\alpha$.
Always print d.mean and d.variance after building a distribution; mistakes are silent (mean 40 became 0.625).
Quick check: SciPy nbinom(n=10, p=0.2). Write the NumPyro NegativeBinomial2 call and the variance.
$\alpha = n = 10$; mean $\mu = n(1-p)/p = 10 \times 0.8/0.2 = 40$. So NegativeBinomial2(mean=40, concentration=10), with variance $40 + 1600/10 = 200$.
Keeping the mean positive: the log link and the softplus link core
Your layers add up to a number we will call $\eta_t$ (the Greek letter "eta"), the linear predictor: trend + season + holidays + regressors, on a scale where any number is allowed. A steep downward trend line will eventually carry $\eta_t$ below zero. But the mean of a count, $\mu_t$, must stay above zero. A link function is the bridge: it turns any $\eta_t$ into a positive $\mu_t$.
Two bridges are common. The log link says $\mu_t = e^{\eta_t}$. Adding layers in $\eta$ then means multiplying effects in orders: "holiday: $\times 1.3$". The softplus link says $\mu_t = \log(1 + e^{\eta_t})$, a smooth version of "take the larger of $\eta_t$ and 0". When $\eta_t$ is large, $\mu_t \approx \eta_t$, so layers still add in orders: "holiday: $+30$ orders". Near and below zero it flattens to a small positive number.
Three ways to say it:
- Picture: the log link is an elevator that speeds up as it climbs; the softplus link is a straight ramp with a soft floor at zero.
- Numbers: baseline 100 orders and a holiday effect of "30": log link, $100 \times 1.3 = 130$; softplus, $100 + 30 = 130$. Baseline 10: log link gives 13, softplus gives 40.
- Slogan: the log link multiplies; the softplus link adds (when the level is high).
Daily orders with a level of 100, a holiday effect and a trend. Same story, two links.
- Log link. The holiday coefficient is $\beta = \log 1.3 = 0.262$ ("30% more"). Level 100: $\eta = \log 100 + 0.262 = 4.605 + 0.262 = 4.867$, $\mu = e^{4.867} = 130$. Level 10: $\eta = 2.303 + 0.262$, $\mu = 13$. The same $\beta$ means +30% at every level.
- Softplus link. The layers are in orders. Holiday coefficient $= 30$. Level 100: $\eta = 100 + 30 = 130$, $\mu = \text{softplus}(130) = 130.0$. Level 10: $\eta = 40$, $\mu = 40$. The same coefficient means +30 orders at every level, which is +300% on a quiet baseline of 10.
- Growing trend. Log link with a slope of $0.01$ in $\eta$ per day: growth of about 1% per day, compounding. After 90 days: $100 \times e^{0.9} = 246$. Softplus link with a slope of 1 order per day: $100 + 90 = 190$. A straight line in $\eta$ is an exponential in $\mu$ under the log link, which can run away in a long forecast.
- Falling trend. Level 50, falling 3% per day under the log link: day 90 gives $50 \times e^{-2.7} = 3.4$ (a smooth decay). Under softplus with the slope $-1.5$ orders per day, $\eta$ crosses zero on day 33 and by day 90 is $-85$; $\mu = \text{softplus}(-85)\approx 0$. The forecast sticks to the floor.
- Uncertainty. If $\eta_t$ is uncertain, say $\eta \sim N(4.6, 0.3^2)$, then $E[e^\eta] = e^{4.6 + 0.3^2/2} = 104.1$, while $e^{E[\eta]} = 99.5$. Average $\mu$ over posterior draws; do not plug the average $\eta$ into the link.
With a link, the model for a count series becomes
$$\eta_t = g(t) + s(t) + h(t) + X_t\beta,\qquad \mu_t = f(\eta_t),\qquad y_t \sim \text{NB2}(\mu_t, \alpha),$$where $f$ is the inverse link:
| Link | $\mu = f(\eta)$ | Effect of adding $b$ to $\eta$ | For large $\eta$ | For very negative $\eta$ | Code |
|---|---|---|---|---|---|
| identity (Normal, Student-t) | $\eta$ | $+b$ orders | $\eta$ | negative (not allowed for counts) | mu = eta |
| log | $e^{\eta}$ | $\times e^{b}$ | grows exponentially | $e^\eta \to 0$ smoothly | jnp.exp(eta) |
| softplus | $\log(1+e^{\eta})$ | about $+b$ orders if $\eta$ is large | $\approx \eta$ (linear) | $\approx e^\eta \to 0$ | jax.nn.softplus(eta) |
- The log link is the standard ("canonical") link for Poisson GLMs and the usual choice for Negative Binomial regression (Chapter 5.14). Seasonal terms in $\eta$ become multiplicative seasonality (Chapter 7.7).
- $\text{softplus}(3) = 3.049$ vs $e^3 = 20.1$; $\text{softplus}(-2) = 0.127$ vs $e^{-2} = 0.135$.
- The link changes the units of every coefficient and of every prior on the layers: a prior $N(0, 1)$ on a holiday coefficient means "between $\times 0.37$ and $\times 2.7$" under the log link, but "within about $\pm1$ order" under softplus.
Why do we need it?
Without a link, a falling trend or a negative seasonal term can push the expected count below zero, and the Negative Binomial or Poisson is undefined (or the sampler crashes with NaNs). The link guarantees a legal mean and decides how the layers combine.
Where is it used?
Poisson and NB regression (log link), the count likelihoods of Prophet-style models in NumPyro, positive-output heads in neural forecasters (softplus), exposure offsets in insurance and epidemiology, and any positive-valued parameter such as a scale $\sigma$ or concentration $\alpha$.
How is it used?
Build $\eta_t$ from the layers, apply jnp.exp or jax.nn.softplus, pass the result as the mean. Set priors on the layers in the units of $\eta$. Interpret a log-link coefficient as a percentage ($e^\beta - 1$), and average over posterior draws of $\mu$ rather than of $\eta$.
"With a log link, the trend slope is in orders per day."
It is in log units, roughly "percent per day". A straight line in $\eta$ is an exponential curve in the orders, and a long forecast can grow explosively. Check the far end of the forecast.
"Average the draws of $\eta$, then apply the link."
Apply the link to each draw, then average. The link is curved, so $E[e^\eta] \gt e^{E[\eta]}$ (a Jensen effect, Chapter 4.5): in the example, 104.1 against 99.5.
"The same prior on the holiday effect works with either link."
The prior scale is in different units. Moving a model from exp to softplus (or the reverse) means re-thinking every prior scale, including the Laplace scale on the trend changes $\delta_j$.
"Fitting a Normal to $\log y$ and exponentiating gives the mean forecast."
It gives (roughly) the median; the mean is larger by a factor $e^{\sigma^2/2}$ (Chapter 4.18). A count likelihood with a log link models the mean directly.
Find in your forecasting model the line that turns the layers into the mean of a count likelihood. Is it jnp.exp(eta), jax.nn.softplus(eta) or something else, and is the same layer sum used for the Normal and Student-t options? That one line decides whether your holiday coefficients are percentages or orders, whether Fourier seasonality is multiplicative, and what scale a Laplace prior on a slope change $\delta_j$ should have. It also decides how your long-horizon forecast behaves: compounding with exp, roughly linear with softplus.
$\eta = g + s + h + X\beta$; $\mu = e^\eta$ (log link: effects multiply, a line in $\eta$ is exponential) or $\mu = \log(1+e^\eta)$ (softplus: $\approx\eta$ when large, effects add).
Baseline 100, holiday: $\times1.3$ (log) or $+30$ (softplus) both give 130; at baseline 10: 13 vs 40.
Trap: coefficient and prior units change with the link; apply the link per draw before averaging.
Quick check: under the log link, a regressor coefficient is $\beta = 0.10$ per unit of the regressor. How does one more unit change the expected orders?
It multiplies the mean by $e^{0.10} = 1.105$, a 10.5% increase, whatever the current level. Under a softplus link (large level) the same coefficient would mean about +0.10 orders.
Too many zeros: excess zeros, zero inflation and hurdle models core
Some count series have a lot of zero days: a spare part that sells twice a month, orders per hour at night, a feature used by a few customers. First ask why the zeros are there, because there are two very different kinds.
- Sampling zeros: the process could have produced something, but by chance produced nothing. A Poisson with mean 0.5 gives a zero 61% of the time. A Poisson or Negative Binomial already describes these.
- Structural zeros: the process could not produce anything that day: the shop was closed, the product was out of stock, the sensor was off. These are extra zeros on top of what a Poisson or NB expects. They are called excess zeros.
Three modelling stories exist. Plain Negative Binomial: no extra mechanism, the zeros come from low rates and the wide spread. Zero-inflated model: first flip a coin with probability $\pi$: heads means a structural zero; tails means draw from the count model (which can still give a zero). Hurdle model: flip a coin for zero versus not-zero; if not-zero, draw a count that is at least 1 from a "zero-truncated" count model. The difference: in a zero-inflated model zeros have two sources, in a hurdle model only one.
Three ways to say it:
- Picture: zero-inflated: a shop door that is sometimes locked (structural zero) plus a till that may also ring up nothing. Hurdle: a gate that decides "any sale today, yes or no?"; only after passing it does the till count 1, 2, 3, ….
- Numbers: $\pi = 0.3$, $\lambda = 4$. Zero-inflated Poisson: $P(0) = 0.3 + 0.7e^{-4} = 0.313$. A Poisson with the same mean (2.8): $P(0) = 0.061$. A Negative Binomial with the same mean and variance: $P(0) = 0.159$.
- Slogan: ask first why the zeros are there.
A product with daily sales that follow a Poisson with rate $\lambda = 4$ on the days it is available, but on 30% of the days ($\pi = 0.3$) it is unavailable and sells nothing.
- Zero-inflated Poisson (ZIP): $P(0) = \pi + (1-\pi)e^{-\lambda} = 0.3 + 0.7 \times 0.0183 = 0.3128$. For $k \ge 1$: $P(k) = (1-\pi)\,e^{-\lambda}\lambda^k/k!$, for example $P(1) = 0.7 \times 0.0733 = 0.0513$.
- Mean $= (1-\pi)\lambda = 0.7 \times 4 = 2.8$. Variance $= (1-\pi)\lambda(1 + \pi\lambda) = 0.7 \times 4 \times (1 + 1.2) = 6.16$. The variance is more than twice the mean: zero inflation also produces overdispersion.
- A Poisson with the same mean 2.8 has $P(0) = e^{-2.8} = 0.061$: five times too few zeros.
- A Negative Binomial with the same mean and variance ($\alpha = \mu^2/(Var - \mu) = 7.84/3.36 = 2.33$) has $P(0) = (2.33/5.13)^{2.33} = 0.159$. Better, but still only half the 0.313: an NB cannot put a large spike at 0 and keep a smooth rest.
- Hurdle with the same $\pi = 0.3$ and $\lambda = 4$: $P(0) = 0.3$ exactly, and for $k \ge 1$: $P(k) = (1-\pi)\dfrac{e^{-\lambda}\lambda^k/k!}{1 - e^{-\lambda}}$, so $P(1) = 0.7 \times 0.0733/0.9817 = 0.0522$. Mean $= (1-\pi)\lambda/(1-e^{-\lambda}) = 2.85$, variance $6.13$.
- How to tell them apart in practice: both give a spike at zero. The hurdle's zero probability is a free number; the ZIP's is at least $e^{-\lambda}$ (a gate plus a Poisson). Choose by the story: if some zeros are accidents of low counts and others structural, ZIP; if "any sale today at all?" and "how many, given some?" are separate decisions, hurdle.
Zero-inflated count model (base distribution $f$ = Poisson or NB, gate probability $\pi$):
$$P(y=0) = \pi + (1-\pi)f(0),\qquad P(y=k) = (1-\pi)f(k)\ \ (k \ge 1).$$Hurdle count model:
$$P(y=0) = \pi,\qquad P(y=k) = (1-\pi)\,\frac{f(k)}{1-f(0)}\ \ (k \ge 1).$$- NumPyro has
dist.ZeroInflatedPoisson(gate, rate),dist.ZeroInflatedNegativeBinomial2(mean, concentration, gate=...)and the generaldist.ZeroInflatedDistribution(base, gate=...). For a hurdle model, recent NumPyro releases include a ready-made class (for exampledist.HurdleNegativeBinomial2; check your version); otherwise you write it yourself (a mixture, ornumpyro.factorwith the log-probabilities above). - The gate $\pi$ can itself depend on covariates (for example a logistic regression on a "store closed" indicator).
- Check for excess zeros with a posterior predictive check on the number of zero days (Chapter 7.14): if the observed number is far above what the fitted Poisson or NB replicates, zeros are in excess.
- Prefer knowledge over inference. If you know why a day is zero (closed, out of stock, no data collected), say so with a regressor or by masking those days (Chapter 7.12) instead of asking $\pi$ to discover it. In a forecast you also need to know the future values of that indicator.
Why do we need it?
If a large share of days are zero for a reason that is not "a low rate", a Poisson or NB tries to explain them by pulling the mean down and inflating the dispersion. Both the average forecast and the intervals then become wrong on the days when the item is available.
Where is it used?
Intermittent demand (spare parts, slow sellers), counts of rare events such as claims or defects, ecology (species counts), health-care visits, usage counts of a rarely used product feature, and the zero-inflated classes of statsmodels, R's pscl and NumPyro.
How is it used?
Plot the share of zero days next to what the fitted Poisson or NB expects. If there is a clear excess, find out why. If the reason is known, add a regressor or mask. If it is not, try a zero-inflated or hurdle likelihood and compare with held-out log-likelihood and a zero-count check.
"Lots of zeros means I need a zero-inflated model."
Not necessarily. Low rates give many zeros without any extra mechanism, and a Negative Binomial can add more through its spread. Check the observed zero count against a fitted NB's replicates first (Chapter 7.14).
"Zero inflation and overdispersion are the same thing."
They are different problems that overlap: a zero-inflated Poisson is overdispersed, but an overdispersed NB can have few zeros. Fixing one does not fix the other.
"The gate $\pi$ will find the closed days for me."
Only if the data contain enough information. If you already know the closed days, give them to the model (regressor or mask). You also need that information for the future dates you forecast.
Your forecasting model offers the Negative Binomial for counts, not a zero-inflated or hurdle option (as far as the project facts go), so the practical question for a zero-heavy series is: do the zeros come from low rates (the NB is fine), or from days when nothing can happen? In the second case the cleanest fix is usually a regressor or a mask for those days (see Chapter 7.12), or modelling the series at a coarser level (weekly) where zeros disappear. A posterior predictive check on the share of zero days tells you which case you are in.
ZIP: $P(0) = \pi + (1-\pi)f(0)$, $P(k) = (1-\pi)f(k)$. Hurdle: $P(0) = \pi$, $P(k) = (1-\pi)f(k)/(1-f(0))$.
Example $\pi=0.3$, $\lambda=4$: ZIP $P(0)=0.313$; Poisson (same mean) 0.061; NB (same mean and variance) 0.159.
Ask why the zeros are there; known structural zeros belong in a regressor or mask; check the zero count with a PPC.
Quick check: ZIP with $\pi = 0.5$ and $\lambda = 2$. What are $P(0)$ and the mean?
$P(0) = 0.5 + 0.5e^{-2} = 0.5 + 0.5 \times 0.1353 = 0.5677$. Mean $= (1-\pi)\lambda = 0.5 \times 2 = 1$. (A Poisson with mean 1 would give $P(0) = 0.368$.)
Choosing between likelihoods: held-out log-likelihood, fair comparisons and checks core
So far you have three noise stories and several ways each can be wrong. How do you choose? Think of three forecasters who each publish a probability distribution for every future day. When the real days arrive, the best forecaster is the one who put the most probability on what actually happened. The log of that probability, summed over days, is the log predictive density (or log-likelihood on new data). Higher is better.
Two rules keep the comparison honest. First, score on days the model did not see (the last stretch of the series), because a flexible model always looks better on the data it was fitted to. Second, compare scores only for the same kind of outcome: a density for a continuous number and a probability for a whole number are not the same currency, and a density even changes when you change units.
Three ways to say it:
- Picture: a spelling bee: every model names the probability it gave to the true day; add the logs.
- Numbers: on 30 held-out days, a Poisson scored far below a Negative Binomial on overdispersed counts (see the table below); on clean Normal data the extra Student-t parameter bought nothing.
- Slogan: score on new days, in the same units, and then still look at the plots.
120 days of counts with a known weekly pattern and a trend ($\mu_t$ known and shared by all models); the true noise is Negative Binomial with $\alpha = 6$. Fit the noise parameters on the first 90 days, then score the last 30. The table below is the widget's default state (seed 12).
- Fit on days 1 to 90: the Normal gets $\hat\sigma = 9.46$ (one constant spread); the Poisson has no noise parameter; the Negative Binomial gets $\hat\alpha = 10.4$ (the true value is 6; 90 days of data pin $\alpha$ down only roughly).
- Score on days 91 to 120, summing $\log$ of the probability each model gave to the actual count: Poisson $-5.54$ per day, Normal (discretised to integers) $-4.17$, Negative Binomial $-4.03$.
- The difference between the Negative Binomial and the Poisson is $1.51$ nats per day, about $45$ over the 30 days: the Poisson found the actual counts far less likely, because its intervals were far too narrow. (The Normal, with $-4.17$, is much closer: its one constant $\sigma$ is a fair compromise, but it is too wide on quiet days and too narrow on busy ones.)
- In-sample, the Negative Binomial's AIC ($2k - 2\ell$ with $k$ = number of noise parameters) is also the lowest, but AIC only approximates the held-out score, and with the layers also fitted from the data you would also count their parameters.
- The unit trap: a Normal density on raw counts scored $-4.17$ per day, which looks close to the others. Measure the same orders in hundreds and every density value is 100 times larger: the raw row jumps by $\log 100 = 4.6$ per day and it would "beat" the Negative Binomial. Nothing about the model changed. Probabilities of whole numbers (Poisson, NB, a Normal discretised to integers) do not change when you rescale; densities do. So compare a continuous model with a count model only after discretising it (probability of the rounding interval $[y - \tfrac12, y + \tfrac12]$).
For new days $y^{\text{new}}_1, \dots, y^{\text{new}}_m$ and a fitted model with parameter estimate $\hat\theta$ (or posterior draws $\theta^{(s)}$), the held-out log predictive density is
$$\text{lpd} = \sum_{i=1}^{m} \log p(y^{\text{new}}_i \mid \hat\theta)\qquad\text{or, in a Bayesian model,}\qquad \sum_{i=1}^{m}\log\Big(\frac1S\sum_{s=1}^{S} p\big(y^{\text{new}}_i \mid \theta^{(s)}\big)\Big).$$- Higher is better. Divide by $m$ to get a per-day score. For time series the "new days" must come after the training days (time-ordered split, Chapter 7.15).
- In-sample criteria approximate this without a split: AIC $= 2k - 2\ell$ (lower is better), and in Bayesian models WAIC and PSIS-LOO (named here; computed from posterior draws). They assume exchangeable data, so for time series a genuine rolling-origin score is more trustworthy.
- Same currency: compare densities with densities (same units, same transformation of $y$) and probabilities with probabilities. If $y$ is transformed (log, sqrt), add the Jacobian before comparing with a model on the raw scale.
- Beyond one number: the log score can favour a model for the wrong reason (a better centre). Complement it with residual plots (Chapter 7.17), posterior predictive checks (Chapter 7.14) and interval coverage and CRPS (Chapter 7.16).
Why do we need it?
"The Student-t looks better" or "the Negative Binomial fits" are opinions until a number backs them up. A held-out score turns the choice of likelihood into a measured comparison and protects against picking the most flexible model just because it hugs the training data.
Where is it used?
Model comparison in NumPyro, PyMC and Stan (ArviZ compare with LOO/WAIC), time-series competitions that score probabilistic forecasts by log score or CRPS, and the evaluation of neural forecasters with Normal, Student-t or Negative Binomial output heads.
How is it used?
Fit each likelihood on the early days, compute the mean log predictive density on the later days (with logsumexp over posterior draws), compare in the same units, then confirm with a posterior predictive check on the feature that matters (tails, variance, zeros).
"Higher training log-likelihood means the better likelihood."
Extra parameters always raise the training score. Use held-out days (or an information criterion that charges for parameters) and for time series let the held-out days come after the training days.
"I compared the Normal and Poisson log-likelihoods and the Normal won."
A density and a probability mass function are different quantities, and the density changes with the unit of $y$. Discretise the Normal to the same integer grid, or compare both on the probability of the same event.
"The log score picked the winner, so I can skip the residual checks."
A single number can reward a better centre while hiding a wrong tail or a wrong variance. Look at residuals, posterior predictive checks and interval coverage as well.
Both of your projects switch likelihood by configuration (Normal, Student-t, and in the forecasting model also Negative Binomial). The disciplined way to choose is a rolling-origin comparison on held-out periods (Chapter 7.15) using a probabilistic score (log score, CRPS and interval coverage, Chapter 7.16), then a posterior predictive check on the feature that motivated the choice: tails for the Student-t, variance and zeros for the Negative Binomial. If you compare a Normal-likelihood run against an NB-likelihood run, remember the density-versus-probability trap above.
Held-out log predictive density $=\sum_i \log p(y_i^{\text{new}}\mid\hat\theta)$ (or log of the posterior-average of $p$); higher is better; AIC $=2k-2\ell$, lower is better.
Rules: score on later days; compare in the same currency (density vs density, probability vs probability); dividing $y$ by $c$ (dollars to thousands, say) raises every density log-score by $\log c$ per day, and multiplying by $c$ lowers it by the same amount.
Then confirm with residuals, a posterior predictive check and interval coverage.
Quick check: $y$ is measured in dollars and a Normal model scores $-4.0$ per day. You re-express $y$ in thousands of dollars (divide by 1000). What does the same Normal model now score, and what does that tell you?
Densities scale: the log-density rises by $\log 1000 = 6.91$ per day, to $+2.9$. The model is exactly as good as before; the score of a continuous model only means something when compared with other continuous models in the same units.
Recap, cheat sheet and practice
- The likelihood is the noise story: layers give $\mu_t$; the likelihood spreads it into $p(y_t \mid \mu_t, \varphi)$. It controls the fit, the interval widths and the tail probabilities. Days are assumed independent given $\mu_t$.
- Normal: continuous, symmetric, light tails, constant variance; fit = least squares; $\hat\sigma^2$ = mean squared residual; NumPyro
Normal(loc, scale)takes the sd. A few extreme days inflate $\hat\sigma$ and bend the trend. - Student-t: fat tails with dial $\nu$ (not $n-1$); $\nu \to \infty$ is Normal; sd $= \sigma\sqrt{\nu/(\nu-2)}$. Weights $w = \frac{\nu+1}{\nu+r^2/\sigma^2}$. Say it right: it assigns more probability to extreme residuals, so they exert less influence. It does not remove outliers.
- Counts: Normal fails for small counts (negative values, skew, constant spread); Poisson forces $Var = \mu$; real counts are overdispersed because the rate wobbles. NB2: $Var = \mu + \mu^2/\alpha$, $\alpha \to \infty$ is Poisson, small $\alpha$ is noisy.
- Parameterization: SciPy
nbinom(n, p)has $p = \alpha/(\alpha+\mu)$; NumPyroprobs$= \mu/(\alpha+\mu)$;logits$= \log(\mu/\alpha)$; GammaPoisson rate $= \alpha/\mu$; statsmodels alpha $= 1/\alpha$. Mistakes are silent: print mean and variance. - Positive mean: log link (effects multiply, a line in $\eta$ is exponential) or softplus (about linear when large, effects add). Priors and coefficients change units with the link; apply the link per posterior draw.
- Zeros: ask why. Low rates: NB. Structural zeros: regressor or mask, else zero-inflated ($P(0) = \pi + (1-\pi)f(0)$) or hurdle ($P(0) = \pi$). Compare likelihoods on held-out days in the same currency (densities vs densities, probabilities vs probabilities), then check residuals and posterior predictive draws.
Cheat sheet
| Normal | Student-t | Negative Binomial (NB2) | Zero-inflated / hurdle | |
|---|---|---|---|---|
| Values | real | real | $0, 1, 2, \dots$ | $0, 1, 2, \dots$ with a zero spike |
| Noise parameters | $\sigma$ | $\sigma$, $\nu$ | $\alpha$ | gate $\pi$ + base parameters |
| Mean | $\mu_t$ | $\mu_t$ ($\nu \gt 1$) | $\mu_t$ | ZIP: $(1-\pi)\lambda$ |
| Variance | $\sigma^2$, constant | $\sigma^2\nu/(\nu-2)$ ($\nu \gt 2$) | $\mu_t + \mu_t^2/\alpha$ | ZIP: $(1-\pi)\lambda(1+\pi\lambda)$ |
| Tails | light | heavy | right-skewed, long right tail | spike at 0 |
| Layers enter through | identity: $\mu_t = \eta_t$ | identity | log or softplus link | link for the count part (and maybe for the gate) |
| Fit as | least squares | weighted least squares, $w = \frac{\nu+1}{\nu + r^2/\sigma^2}$ | GLM-type fit | mixture |
| NumPyro | Normal(loc, scale) | StudentT(df, loc, scale) | NegativeBinomial2(mean, concentration) | ZeroInflatedPoisson(gate, rate), ZeroInflatedNegativeBinomial2(mean, concentration, gate=…) |
| SciPy | norm(loc, scale) | t(df, loc, scale) | nbinom(n=α, p=α/(α+μ)) | (write the pmf) |
| Breaks when | counts, fat tails, fanning spread | many small counts; $\nu$ poorly identified | excess zeros; autocorrelated noise | zeros are explainable by a known cause |
import numpy as np, jax, jax.numpy as jnp
import numpyro, numpyro.distributions as dist
from numpyro.infer import MCMC, NUTS, Predictive
from scipy import stats
from scipy.special import logsumexp
# 1) Know the parameterization: one Negative Binomial (mean 40, concentration 5) written four ways
mu, alpha = 40.0, 5.0
p_scipy = alpha / (alpha + mu) # SciPy: n = alpha, p = alpha / (alpha + mu)
ds = {"SciPy nbinom(n, p)": stats.nbinom(alpha, p_scipy),
"NumPyro NegativeBinomial2(mean, concentration)": dist.NegativeBinomial2(mu, alpha),
"NumPyro NegativeBinomialProbs(total_count, probs)": dist.NegativeBinomialProbs(alpha, mu / (alpha + mu)),
"NumPyro NegativeBinomialLogits(total_count, logits)": dist.NegativeBinomialLogits(alpha, np.log(mu / alpha))}
for name, d in ds.items():
m, v = (d.mean(), d.var()) if name.startswith("SciPy") else (float(d.mean), float(d.variance))
print(f"{name:52s} mean {m:5.1f} variance {v:6.1f}") # every line: mean 40.0, variance 360.0
print("MISTAKE (SciPy p given to NumPyro probs): mean =", float(dist.NegativeBinomialProbs(alpha, p_scipy).mean)) # 0.625
# 2) Heavy tails: trend + t(3) noise + three bad days near the end; Normal vs Student-t likelihood
rng = np.random.default_rng(11)
n = 120; t = np.arange(n)
y = 100 + 0.5 * t + 3 * rng.standard_t(3, n) # true slope 0.5 per day
y[[88, 94, 98]] += 45 # three bad days in the training window
tt, yj = jnp.array(t / 100.0), jnp.array(y)
train, test = slice(0, 100), slice(100, 120)
def trend_model(tt, y=None, likelihood="normal"):
a = numpyro.sample("a", dist.Normal(100, 50))
b = numpyro.sample("b", dist.Normal(0, 100)) # slope per 100 days
sigma = numpyro.sample("sigma", dist.HalfNormal(20))
mu = a + b * tt
if likelihood == "normal":
numpyro.sample("obs", dist.Normal(mu, sigma), obs=y) # scale = standard deviation
else:
nu = numpyro.sample("nu", dist.Gamma(2.0, 0.1))
numpyro.sample("obs", dist.StudentT(nu, mu, sigma), obs=y) # df, loc, scale
for lik in ("normal", "studentt"):
mc = MCMC(NUTS(trend_model), num_warmup=300, num_samples=400, progress_bar=False)
mc.run(jax.random.PRNGKey(0), tt[train], yj[train], likelihood=lik)
s = mc.get_samples()
mu_te = s["a"][:, None] + s["b"][:, None] * tt[test][None, :]
d = dist.Normal(mu_te, s["sigma"][:, None]) if lik == "normal" else dist.StudentT(s["nu"][:, None], mu_te, s["sigma"][:, None])
lpd = (logsumexp(np.asarray(d.log_prob(yj[test])), axis=0) - np.log(mu_te.shape[0])).sum() # held-out log density
lo, hi = np.quantile(s["b"], [0.05, 0.95]) / 100
print(f"{lik:9s} slope/day {s['b'].mean()/100:.3f} (90% interval {lo:.3f} to {hi:.3f}), sigma {s['sigma'].mean():.2f}, held-out log density {lpd:.1f}")
# 3) Overdispersed counts: Poisson vs Negative Binomial with a log link, held-out score and a PPC
rng = np.random.default_rng(5)
n = 140; t = np.arange(n)
weekday = np.array([0.0, -0.1, -0.1, 0.0, 0.2, 0.5, 0.3])
cnt = rng.negative_binomial(4, 4 / (4 + np.exp(np.log(12.0) + 0.002 * t + weekday[t % 7]))) # true concentration 4
X, cj, tn = jnp.array(np.eye(7)[t % 7]), jnp.array(cnt), jnp.array(t / 100.0)
trn, tst = slice(0, 112), slice(112, 140)
def count_model(X, tt, y=None, likelihood="poisson"):
a = numpyro.sample("a", dist.Normal(2.5, 1.0))
w = numpyro.sample("w", dist.Normal(0, 0.5).expand([7]))
b = numpyro.sample("b", dist.Normal(0, 0.5))
mu = jnp.exp(a + X @ w + b * tt) # log link keeps the mean positive
if likelihood == "poisson":
numpyro.sample("obs", dist.Poisson(mu), obs=y)
else:
alpha = numpyro.sample("alpha", dist.Gamma(2.0, 0.2)) # concentration; Var = mu + mu^2/alpha
numpyro.sample("obs", dist.NegativeBinomial2(mu, alpha), obs=y)
for lik in ("poisson", "negbin"):
mc = MCMC(NUTS(count_model), num_warmup=300, num_samples=400, progress_bar=False)
mc.run(jax.random.PRNGKey(1), X[trn], tn[trn], y=cj[trn], likelihood=lik)
s = mc.get_samples()
mu_te = jnp.exp(s["a"][:, None] + (X[tst] @ s["w"].T).T + s["b"][:, None] * tn[tst][None, :])
lp = dist.Poisson(mu_te).log_prob(cj[tst]) if lik == "poisson" else dist.NegativeBinomial2(mu_te, s["alpha"][:, None]).log_prob(cj[tst])
lpd = (logsumexp(np.asarray(lp), axis=0) - np.log(lp.shape[0])).sum()
rep = np.asarray(Predictive(count_model, posterior_samples=s)(jax.random.PRNGKey(2), X[trn], tn[trn], likelihood=lik)["obs"])
print(f"{lik:8s} held-out log density {lpd:.1f} | variance of the daily counts: observed {cnt[trn].var():.1f}, "
f"replicated {rep.var(1).mean():.1f} | zero days: observed {(cnt[trn]==0).sum()}, replicated {(rep==0).sum(1).mean():.2f}")
# Output of this script (about 7 seconds on a laptop):
# SciPy nbinom(n, p) mean 40.0 variance 360.0 (all four lines)
# MISTAKE (SciPy p given to NumPyro probs): mean = 0.625
# normal slope/day 0.582 (90% interval 0.532 to 0.639), sigma 9.57, held-out log density -69.9
# studentt slope/day 0.516 (90% interval 0.493 to 0.539), sigma 3.22, held-out log density -57.8
# poisson held-out log density -120.0 | variance ... observed 118.4, replicated 49.5 | zero days: observed 1, replicated 0.00
# negbin held-out log density -96.0 | variance ... observed 118.4, replicated 127.9 | zero days: observed 1, replicated 0.38
# Reading it: the Normal slope interval misses the true 0.5 (the bad days bent the trend); the Student-t covers it, with a
# much smaller sigma and a higher held-out score. The Poisson replicates a variance of 50 against an observed 118.
1. In NumPyro you write dist.StudentT(4.0, 0.0, 2.0). What is the standard deviation of this noise?
2. Which sentence about a Student-t likelihood is the one to use in an interview?
3. A forecast uses a Negative Binomial with mean $\mu = 50$ and concentration $\alpha = 10$. What is the variance of that day's count?
4. SciPy stats.nbinom(n=4, p=0.25) describes the same distribution as which NumPyro call?
probs would have to be $1 - p = 0.75$, not 0.25 (passing 0.25 gives mean $4 \times 0.25/0.75 = 1.33$).5. A zero-inflated Poisson has gate probability $\pi = 0.4$ and rate $\lambda = 3$. What is $P(y = 0)$?
6. Your NB model uses a log link and the holiday coefficient has posterior mean 0.20. What does it say about holiday demand?
Practice problems
A. Residuals $4, -2, 0, 6, -3, -5$ under a Normal likelihood. Find $\hat\sigma$ and the 90% band half-width. Then add one more residual of $+30$ and refit. What changed?
Squares: $16, 4, 0, 36, 9, 25$, sum $90$. $\hat\sigma = \sqrt{90/6} = \sqrt{15} = 3.87$; the 90% half-width is $1.645 \times 3.87 = 6.37$. With the extra residual: sum of squares $= 90 + 900 = 990$, $n = 7$, $\hat\sigma = \sqrt{990/7} = 11.89$, half-width $19.6$. One day tripled $\hat\sigma$ (factor 3.07) and made every interval 3.07 times wider. A Student-t fit would keep its scale much closer to the first answer.
B. A Student-t likelihood with $\nu = 5$ and scale 2. Compute the weights $w = (\nu+1)/(\nu + r^2/\sigma^2)$ of residuals $r = 2$, $10$ and $40$, and the standard deviation of the noise.
$r = 2$: $r^2/\sigma^2 = 1$, $w = 6/(5+1) = 1$. $r = 10$: $r^2/\sigma^2 = 25$, $w = 6/30 = 0.2$. $r = 40$: $r^2/\sigma^2 = 400$, $w = 6/405 = 0.0148$. So the day at 40 counts about 1.5% as much as a normal day, but not zero. Standard deviation $= 2\sqrt{5/3} = 2.58$ (the scale 2 is not the sd).
C. Orders on six Fridays: $35, 52, 41, 66, 28, 58$. Compute the mean, variance and dispersion index; estimate $\alpha$ by moments; compare the 90% intervals of a Poisson and a Negative Binomial and count how many Fridays fall outside each.
Mean $= 280/6 = 46.67$. Squared distances sum to $1047.3$, variance $= 1047.3/5 = 209.5$, index $= 209.5/46.67 = 4.49$. Moment estimate $\hat\alpha = 46.67^2/(209.5 - 46.67) = 13.4$. Poisson 90% interval $[36, 58]$: three Fridays (35, 28, 66) are outside (50% instead of 10%). Negative Binomial interval $[25, 73]$ contains all six. (Six values are a small sample, but the index of 4.5 is a strong warning.)
D. You want a Negative Binomial with mean 25 and concentration 4. Give the arguments for NumPyro NegativeBinomial2, NegativeBinomialProbs, NegativeBinomialLogits, GammaPoisson, SciPy nbinom and the statsmodels alpha, and the variance.
NegativeBinomial2(mean=25, concentration=4). NegativeBinomialProbs(total_count=4, probs=25/29=0.862). NegativeBinomialLogits(total_count=4, logits=log(25/4)=1.833). GammaPoisson(concentration=4, rate=4/25=0.16). SciPy nbinom(n=4, p=4/29=0.138). statsmodels $\text{alpha} = 1/4 = 0.25$. Variance $= 25 + 625/4 = 181.25$ (sd 13.5). Check in code: d.mean must print 25 and d.variance 181.25.
E. Baseline demand is 80 orders. A holiday lifts it by 25%. Give the holiday coefficient under (i) a log link and (ii) a softplus link on the count scale. What happens to each if the baseline falls to 20?
(i) Log link: $\beta = \log 1.25 = 0.223$; demand $= 80 \times 1.25 = 100$. At baseline 20: $20 \times 1.25 = 25$ (+5 orders). (ii) Softplus: the coefficient is $+20$ orders ($\eta = 80 + 20 = 100$, softplus$(100) = 100$). At baseline 20: $\eta = 40$, demand 40 (+20 orders, which is +100%). So the log link keeps the holiday effect proportional to the level; the softplus link keeps it a fixed number of orders. Choose by looking at which pattern your data show.
F. Interview: "Your forecasting model has Normal, Student-t and Negative Binomial likelihoods. How do you choose, and what does each protect you from?"
Model answer: "First the data type: counts get a Negative Binomial (or Poisson if the variance equals the mean) with a log or softplus link, because a Normal puts mass on negative values and ignores that busy days are noisier. For real-valued series I start with a Normal; if residuals show heavy tails or isolated spikes I use a Student-t, which assigns more probability to extreme residuals so they exert less influence on the trend and the noise scale. It does not delete anything, and I still look at the low-weight days. For a Negative Binomial, $Var = \mu + \mu^2/\alpha$, and I check the exact NumPyro parameterization by printing the mean and variance. I compare likelihoods on held-out later days, in the same units, and confirm with posterior predictive checks for the variance, tails and the share of zero days."
Bayesian forecasting: predictive distributions, uncertainty and posterior predictive checks
"Tomorrow: 120 orders" is a number you cannot act on. Should you staff for 120? What if demand reaches 150 and the warehouse can only ship 140? A Bayesian forecasting model does not hand you one number; it hands you a whole cloud of possible futures with weights, and you read the answer to your question off that cloud. This chapter shows how to read it, where its width comes from (four different sources of uncertainty), why it opens up with the horizon, and how to check, before you trust it, that the model can produce data that look like yours.
- Tell a point forecast from a predictive distribution, and read the mean, the median, quantiles, a prediction interval and $P(\text{demand} \gt \text{capacity})$ from it
- Answer decision questions by counting posterior predictive draws, including questions about several days at once (use whole paths, not daily quantiles) and "which number should I stock?"
- Name and separate the four sources of forecast uncertainty: parameter, observation noise, model, and exogenous regressors; say which ones more data can shrink
- Explain why intervals widen with the horizon and which mechanism applies: parameter uncertainty, future trend changes, accumulating noise
- Run posterior predictive checks for forecasts: replicate whole series, compare the mean, variance, tails, seasonality, extremes, zero frequency and autocorrelation, and name the fix for each failure
- Check the forecast itself on held-out days (coverage and where the actuals landed), as a preview of Chapters 7.15 and 7.16
What we need from earlier chapters: the posterior predictive distribution $p(\tilde y\mid D) = \int p(\tilde y\mid\theta)\,p(\theta\mid D)\,d\theta$ and how to simulate it (draw $\theta$, then draw $\tilde y$), with its split into parameter uncertainty and observation noise (Chapter 6.1); posterior predictive checks in general and why they find misfit (Chapter 6.8); credible intervals (Chapter 6.4); the law of total variance (Chapter 4.6); the additive model and how it forecasts (Chapter 7.7); trend changes (Chapter 7.8); the likelihoods (Chapter 7.13); regressors and their future availability (Chapter 7.12). You will meet the full backtesting tools later: time-series cross-validation (Chapter 7.15), probabilistic scores such as coverage and CRPS (Chapter 7.16) and residual diagnostics (Chapter 7.17); this chapter only previews them. Notation: $T$ is the forecast origin (the last day you have data for) and $h$ the horizon (how many days ahead); $\tilde y_{T+h}$ ("y tilde") is a future, not-yet-seen observation; $\theta^{(s)}$ is the $s$-th posterior draw out of $S$; $y^{\text{rep}}$ is a replicated series simulated from the fitted model; $T(\cdot)$ is a summary statistic of a series.
Colours: the observed data are drawn in the neutral ink colour; the forecast is orange; a capacity or threshold is purple. In the uncertainty-source widgets: blue = observation noise, teal = parameter uncertainty, pink = regressor uncertainty, orange = model uncertainty; green marks what really happened; red marks misses and bad cases. Labels always say the same thing in words.
A point forecast versus a predictive distribution core
A weather app can say "tomorrow: 24 degrees" or "tomorrow: 24 degrees, 70% chance of rain, 5% chance of a storm". The first is a point forecast: one number. The second describes the whole range of tomorrows the forecaster considers possible, with weights. You can only plan for the rain with the second.
A Bayesian forecasting model produces the second kind. For tomorrow's orders it gives a predictive distribution: every demand level that could happen and how plausible each one is. The point forecast is just one summary of that cloud (its mean, or its median). Different questions need different summaries of the same cloud: "how many orders on average?" (the mean), "what is a typical day?" (the median), "how likely is demand above our capacity?" (a tail probability), "what range should I plan for?" (an interval).
Three ways to say it:
- Picture: one dot on a number line versus a whole histogram with a mean, a median, an interval and a capacity line on it.
- Numbers: point forecast 120. Predictive distribution: mean 120, median 117, 90% interval 67 to 185, and a 19% chance of demand above 150.
- Slogan: a forecast is a distribution; a point forecast is one summary of it.
Tomorrow's orders. To keep the arithmetic exact, let the predictive distribution be a Negative Binomial with mean 120 and concentration $\alpha = 12$ (in a real model it also carries parameter uncertainty, which the next sections add). The warehouse can ship 150 orders a day.
- Mean (expected demand): 120. Variance $= \mu + \mu^2/\alpha = 120 + 14400/12 = 1320$, standard deviation $36.3$.
- Median: 117. Half of the possible futures are below 117. It is below the mean because the distribution has a long right tail that pulls the average up.
- Quantiles: a quantile is the value below which a given share of the futures falls (the 5th percentile is the value that 5% of the futures stay under). Here the 5th percentile is 67 and the 95th is 185, so the 90% prediction interval is $[67,\ 185]$: tomorrow lands in it with probability 0.90.
- Exceeding capacity: $P(\tilde y \gt 150) = 0.191$, about 1 day in 5. The point forecast of 120 is 30 below capacity and looks safe; the distribution says it is not.
- Treating "demand = 120" as certain would give $P(\tilde y \gt 150) = 0$. That is the cost of reporting only a point forecast.
- With posterior draws you get the same numbers by counting: with 1 000 draws about 191 of them exceed 150, and that estimate has a Monte Carlo error (the wobble that comes from using only 1 000 random draws instead of infinitely many) of about $\sqrt{0.19 \times 0.81/1000} = 0.012$, i.e. $\pm1.2$ percentage points.
The predictive distribution for a future observation, given the data $D$ and the model, is $p(\tilde y_{T+h} \mid D) = \int p(\tilde y_{T+h} \mid \theta)\,p(\theta \mid D)\,d\theta$. You almost never write it down; you sample it: for $s = 1, \dots, S$, draw $\theta^{(s)} \sim p(\theta \mid D)$, then draw $\tilde y^{(s)} \sim p(\tilde y \mid \theta^{(s)})$. Every summary is then a simple count or sort:
| Question | Summary of the $S$ draws | Best single number for this loss |
|---|---|---|
| expected demand | mean $\frac1S\sum_s \tilde y^{(s)}$ | minimizes squared error |
| typical demand | median (50th percentile) | minimizes absolute error |
| demand level not exceeded with probability $q$ | $q$-quantile of the sorted draws | minimizes the pinball (asymmetric) loss for $\tau = q$ (Chapter 7.16) |
| chance of exceeding capacity $c$ | $\frac1S\,\#\{s : \tilde y^{(s)} \gt c\}$ | (a probability, not a point) |
| 90% prediction interval | 5th and 95th percentiles | (a range) |
- A prediction interval is about the future observation $\tilde y$. A credible interval for a parameter or for the mean demand $\mu$ is narrower, because it leaves out the noise of a single day (Chapter 6.4). A "95% credible interval for the mean" is not a "95% prediction interval for tomorrow".
- Monte Carlo error: a tail probability $p$ estimated from $S$ draws has standard error $\sqrt{p(1-p)/S}$. Rare events need many draws.
Why do we need it?
Decisions depend on risk, not only on the centre: staffing, stock, server capacity and cash buffers are set by how bad a plausible bad day is. A point forecast cannot say how bad, and it hides how much to trust it.
Where is it used?
Capacity and inventory planning, call-centre staffing, energy-demand forecasting, probabilistic forecasting competitions, the Prophet-style model's uncertainty intervals, and Bayesian forecasting libraries (NumPyro Predictive, Stan generated quantities, PyMC sample_posterior_predictive).
How is it used?
Draw posterior samples, push them through the model for the future dates to get an array of shape (draws, days), then take means, medians, np.quantile along the draws axis, and (draws > capacity).mean(). Report the question you answered, not only a number.
"The forecast is 120."
The point forecast is 120 under one loss (say, squared error). The forecast is the whole distribution; 120 is its mean, and its median, 117, is a different "point forecast" of the same cloud.
"A 90% prediction interval means the true expected demand is in it with 90% probability."
A prediction interval is about one future observation $\tilde y$ and includes day-to-day noise. A 90% credible interval for the expected demand $\mu$ is a different, much narrower interval.
"My model says P(exceed capacity) = 0.019, so it is certain to be fine on 98% of days."
That number is conditional on the model being right. It includes parameter uncertainty and noise but not the chance that the model is wrong; the next sections and the checks at the end of this chapter address that.
"Our 95% interval says we are 95% sure the forecast is 120."
"The 95% interval is the interval for the average demand."
"The forecast is a predictive distribution. Its 95% prediction interval contains tomorrow's orders with probability 0.95 under the model; the credible interval for the expected demand is much narrower."
Model answer: "A point forecast is one summary, the mean or the median, of the posterior predictive distribution. I report the median, an 80% and a 95% prediction interval and the probability of exceeding capacity. The prediction interval includes day-to-day noise and parameter uncertainty; an interval for the mean demand would leave the noise out and would be far too narrow to plan with."
In your forecasting model, "the forecast" is a (draws, days) array: for each posterior draw of the trend, changepoint slopes $\delta_j$, Fourier weights, holiday and regressor effects and noise parameters, one simulated future path from the likelihood. In NumPyro this is what Predictive returns when you do not pass the observed data. From the A/B-testing side, the analogue is the posterior of a metric difference: there the "future" is the next user rather than the next day, and the decision question ($P(\theta_B \gt \theta_A \mid D)$) is a statement about parameters, not about $\tilde y$ (the two uncertainty statements of the capstone, Chapter 7.20).
Predictive: $p(\tilde y\mid D)=\int p(\tilde y\mid\theta)p(\theta\mid D)d\theta$; sample it: draw $\theta^{(s)}$, then $\tilde y^{(s)}$.
Mean = expected demand; median = typical day; $P(\tilde y \gt c)$ = share of draws above $c$ (error $\sqrt{p(1-p)/S}$); interval = percentiles of the sorted draws.
Example NB(120, 12): median 117, 90% interval [67, 185], $P(\gt150)=0.19$. A prediction interval is wider than a credible interval for $\mu$.
Quick check: out of 2 000 predictive draws for next Monday, 130 exceed capacity. Estimate $P(\text{exceed})$ and its Monte Carlo error.
$\hat p = 130/2000 = 0.065$. Error $= \sqrt{0.065 \times 0.935/2000} = 0.0055$, so the estimate is $6.5\% \pm 0.55$ percentage points. To pin a 1-in-1000 event to ±10% of itself you would need about $0.001 \times 0.999/(0.0001)^2 \approx 100\,000$ draws.
Answering questions with draws: tail probabilities, intervals and why whole paths matter core
Picture a spreadsheet. Each row is one simulated future: a complete trajectory, "Monday 112, Tuesday 98, Wednesday 131, …", produced from one draw of the model's parameters. Each column is one future day. A question about one day ("is tomorrow above capacity?") uses one column: count the rows where it is true. A question about several days ("does any day next week exceed capacity?" or "what will next week's total be?") uses whole rows: for each row, ask the question about the row, then count rows.
Why can't you just use the daily intervals? Because days are not independent. Every day in a row shares the same parameter draw: the same trend, the same weekday weights, the same holiday effects. If this row has a trend that is a little too high, all its days are a little too high. Daily quantiles throw this connection away; rows keep it.
Three ways to say it:
- Picture: S rows (futures) by h columns (days); one day = one column, a week = whole rows.
- Numbers: each day's 95th percentile is 163 orders; add seven of them and you get 1 141 for the week. But 95% of simulated weeks total less than 936.
- Slogan: keep the rows together; answer by counting rows.
A shop expects about 100 orders a day. Each day's orders are Negative Binomial with concentration 12 around a level that is itself uncertain by about 15% and is the same for all seven days of the week (a stand-in for shared parameter uncertainty). Simulate 400 000 futures (rows) of 7 days each. Capacity is 150 orders a day.
- One column (Monday): the 90% interval is $[52,\ 163]$, and $P(\text{Monday} \gt 150) = 0.082$.
- "Does any of the seven days exceed 150?" For each row take the maximum of its seven numbers, and count the rows where it exceeds 150: $0.379$. If you wrongly treat the days as independent and combine the daily probabilities, $1 - (1 - 0.082)^7 = 0.451$. The shared level makes the days move together, so the "at least one" chance is smaller than independence suggests.
- "What is next week's total?" Add each row: the mean total is 700 (the sum of the seven daily means, which does work for means). The 90% interval of the totals is $[502,\ 936]$.
- Wrong shortcut 1: add the seven daily 95th percentiles, $7 \times 163 = 1\,141$. That asks for all seven days to be extremely high at once. Real weeks above 936 happen only 5% of the time.
- Wrong shortcut 2: shuffle each day's column independently (pretend days are unrelated) and add: the 90% interval of the total shrinks to $[558,\ 857]$, a spread (standard deviation) of 91 instead of 133. The model's shared uncertainty has been thrown away, and the interval is about 30% too narrow.
Let $\tilde Y$ be the $S \times H$ array of predictive draws, $\tilde y^{(s)}_j$ = day $j$ of draw $s$ (all days of draw $s$ come from the same $\theta^{(s)}$ and one simulated path). Then:
- one day: $P(\tilde y_j \gt c) \approx \frac1S\sum_s \mathbf 1[\tilde y^{(s)}_j \gt c]$, interval $=$ quantiles of column $j$;
- some day in a window: $P(\max_{j \le 7}\tilde y_j \gt c) \approx \frac1S\sum_s \mathbf 1[\max_{j\le7}\tilde y^{(s)}_j \gt c]$;
- total over a window: $\tilde Y_{\text{tot}}^{(s)} = \sum_{j \le 7} \tilde y^{(s)}_j$; take its mean, quantiles and tail probabilities;
- the mean of a sum is the sum of the means (linearity), but quantiles do not add: $q_p(\sum_j \tilde y_j) \ne \sum_j q_p(\tilde y_j)$, and the variance of a sum includes covariances: $Var(\sum_j \tilde y_j) = \sum_j Var(\tilde y_j) + 2\sum_{j \lt k} Cov(\tilde y_j, \tilde y_k)$ (Chapter 4.15).
The marginal distributions of the days (the columns) do not determine the joint distribution (the rows). Store the whole draws array, not just a table of daily quantiles.
Why do we need it?
Real decisions span several days: a week's staffing plan, a month of stock, "will we ever hit the limit this quarter". Daily intervals cannot answer them and, used naively, give badly wrong totals and risk numbers.
Where is it used?
Weekly and monthly demand totals, peak-load probabilities for servers and power grids, budget forecasts with a probability of overspend, inventory policies over a lead time, and any model output summarized with "probability of at least one…".
How is it used?
Keep the array ytil[draws, days]. For a window use ytil[:, a:b].sum(axis=1) or (ytil[:, a:b] > cap).any(axis=1).mean(), then quantiles of the result. Never add percentiles across days.
"I have daily 90% intervals, so the weekly total is the sum of the lower ends to the sum of the upper ends."
Adding endpoints gives a range that is far too wide (it needs every day to be at its extreme together). Sum each simulated path, then take quantiles of the sums.
"Days are independent, so I can multiply the daily probabilities."
Only if the days share no parameters and the noise is independent. They share the posterior draw, so they are positively linked; and an autocorrelated noise model would link them further. Independence made the total interval about 30% too narrow in the example.
"500 draws are plenty for any probability."
The error of a probability estimate is $\sqrt{p(1-p)/S}$: fine for $p \approx 0.2$, useless for a 1-in-1000 event. Use more draws for rare events, or an analytic tail.
When your forecasting model feeds a capacity or staffing decision, the output you want is the (draws, days) array of posterior predictive paths, not a table of daily means and bands. Bands per day are fine for plotting; totals and "ever exceeds" questions need rows. If you save forecasts for downstream users, save the draws (or enough draws), not only quantiles. In the A/B framework the equivalent habit is to keep joint posterior draws of $(\theta_A, \theta_B)$ together when you compute $P(\theta_B - \theta_A \gt \delta)$ (Chapter 6.4): comparing separate summaries of the two posteriors loses the same kind of information.
Draws array $\tilde Y$ ($S$ rows = futures, $H$ columns = days). One day: a column. A window: a row (max, sum), then count or take quantiles across rows.
Means add across days; quantiles do not; variances add only with the covariances (shared $\theta$ makes them positive).
Example: daily 95th percentile 163, weekly total 90% interval [502, 936], sum of daily percentiles 1141 (wrong).
Quick check: 10 000 paths of the next 3 days. 700 paths have day 1 above capacity, 800 have day 2 above, 900 have day 3 above, and 1 600 paths have at least one of the three above. What is $P(\text{at least one})$, and what would independence have predicted?
Over whole paths: $1600/10000 = 0.16$. Independence: $1 - (1-0.07)(1-0.08)(1-0.09) = 1 - 0.93 \times 0.92 \times 0.91 = 0.221$. The true probability is lower because the exceedances overlap (the same bad draws exceed on several days), which only the paths can show.
Which number do I act on? Costs, quantiles and the stock decision core
A baker must decide tonight how many loaves to bake for tomorrow. Bake too few and customers walk away: a lost sale. Bake too many and the leftovers are wasted: a smaller loss. If a lost sale costs five times as much as a leftover loaf, then baking the average demand is a poor plan: half the days you run short, and shortages are the expensive mistake. The better plan is to bake enough to cover demand on five days out of six.
That amount is not the mean or the median of tomorrow's demand. It is a quantile of the predictive distribution (the 83rd percentile), chosen by the costs. This is the practical meaning of "a forecast is a distribution": the number you act on depends on what a mistake in each direction costs, and the distribution lets you compute it.
Three ways to say it:
- Picture: a see-saw loaded with the two costs: the balance point of the predictive distribution moves to the right when under-stocking hurts more.
- Numbers: lost sale costs 5, unsold unit costs 1: stock the $5/6 = 83$rd percentile, 154 units instead of the mean 120; expected cost falls from 86.4 to 59.0 a day, 32% lower.
- Slogan: stock at the quantile $c_u/(c_u + c_o)$.
Tomorrow's demand has the predictive distribution Negative Binomial with mean 120 and concentration 12 (as in the first section). A lost sale costs $c_u = 5$ per unit short; an unsold unit costs $c_o = 1$ per unit left over. Decide the stock level $s$.
- Cost of a day with demand $y$ and stock $s$: $5 \cdot (y - s)_+ + 1 \cdot (s - y)_+$, where $(x)_+$ means "$x$ if positive, otherwise 0".
- Stock the mean, $s = 120$: expected lost sales $14.4$ units (cost 72.0), expected leftovers $14.4$ units (cost 14.4), total expected cost $\mathbf{86.4}$. Stock the median, $s = 117$: $92.0$.
- Add one more unit to the stock. With probability $1 - F(s)$ demand is above $s$ and you gain $c_u = 5$ (a sale saved); with probability $F(s)$ demand is at or below $s$ and you lose $c_o = 1$ (one more leftover). Keep adding while $5(1 - F(s)) \gt 1\cdot F(s)$, that is while $F(s) \lt 5/6 = 0.833$. ($F(s)$ is the service level: the chance that demand is covered by a stock of $s$.)
- $F(153) = 0.828 \lt 0.833 \le F(154) = 0.834$, so the best stock is $s^* = \mathbf{154}$.
- At $s^* = 154$: expected lost sales $4.2$ (cost 20.8), expected leftovers $38.2$ (cost 38.2), total $\mathbf{59.0}$, which is 32% below the cost of stocking the mean. You hold a lot more leftovers on purpose, because they are cheap.
- Flip the costs ($c_u = 1$, $c_o = 5$): the quantile is $1/6$, the best stock is only 85 units, and the cost at the mean would be 86.4 against 48.9 at 85.
Choose the action $s$ that minimizes the expected cost under the predictive distribution:
$$s^* = \arg\min_s\ E\big[c_u(\tilde y - s)_+ + c_o(s - \tilde y)_+\big] = F^{-1}\!\Big(\frac{c_u}{c_u + c_o}\Big),$$where $F$ is the predictive CDF and $\frac{c_u}{c_u+c_o}$ is the critical ratio (the "newsvendor" solution). Other losses give other summaries:
| Loss for a miss $e = \tilde y - s$ | Best single number |
|---|---|
| squared error $e^2$ | the predictive mean |
| absolute error $\lvert e\rvert$ | the predictive median |
| linear, with weights $\tau$ above and $1-\tau$ below (the pinball loss) | the predictive $\tau$-quantile (with $\tau = c_u/(c_u + c_o)$) |
- The result needs the right tail of the predictive distribution to be accurate. A Normal likelihood on counts, or intervals that are too narrow, will mislead exactly here (Chapter 7.13, Chapter 7.16).
- If the decision covers several days (a lead time of a week), apply the same rule to the distribution of the total demand over those days, computed from paths.
Why do we need it?
"Forecast accuracy" is not the goal; good decisions are. When mistakes in one direction cost more, the best action is not the centre of the forecast, and only a predictive distribution lets you find it.
Where is it used?
Inventory and retail replenishment, staffing and call-centre capacity, cloud autoscaling headroom, energy procurement, ad-budget pacing, and quantile (pinball-loss) forecasting in competitions and in neural forecasters.
How is it used?
Estimate the two costs, compute $\tau = c_u/(c_u+c_o)$, take np.quantile(draws, tau) (of the total over the lead time if needed). Check the realized service level later: demand should stay below your stock about $\tau$ of the time (calibration, Chapter 7.16).
"The best forecast is the most accurate one, so I should stock the median (or the mean)."
Accuracy is measured in a loss function, and so is the decision. With unequal costs the best number is a quantile away from the centre. The same model, the same forecast, a different action.
"If I stock the 83rd percentile, demand will stay below it exactly 83% of days."
Only if the predictive distribution is well calibrated. Check the realized fraction on later days (Chapter 7.16); an overconfident forecast gives a lower real service level.
"I can read the daily quantile for a week-long lead time."
Use the distribution of the total demand over the lead time (sum the paths); the quantile of a total is not the sum of daily quantiles.
Your forecasting model is probabilistic, so every downstream decision can use the quantile that matches its costs instead of a single "forecast column". This is also the practical reason to care about calibration and the tails of the likelihood: a Student-t or Negative Binomial predictive can move a high quantile by tens of percent compared with a Normal, even when the mean forecast is identical. In the A/B framework the same idea is the expected-loss decision rule: act on the posterior of the lift with the cost of each wrong launch decision in mind (Chapter 6.4).
$s^* = F^{-1}\big(\frac{c_u}{c_u+c_o}\big)$: stock the quantile given by the critical ratio. Mean for squared loss, median for absolute loss, $\tau$-quantile for the pinball loss.
Example NB(120, 12), $c_u=5$, $c_o=1$: $\tau=0.833$, $s^*=154$, cost 59.0 vs 86.4 at the mean.
Needs a calibrated right tail; for multi-day lead times use the total over the paths.
Quick check: a lost sale costs 3 and an unsold unit costs 1. What predictive quantile should you stock, and is it above or below the median?
$\tau = 3/(3+1) = 0.75$, the 75th percentile, which is above the median (50th). Under-stocking is three times as expensive as over-stocking, so you lean toward having extra.
Four sources of forecast uncertainty: noise, parameters, model and regressors core
Suppose tomorrow's orders come in far from your forecast. There are four quite different explanations, and each needs a different cure.
- Observation noise. Even if the model and all its numbers were perfect, real days scatter around the expected value: customers are random. This is a floor. More data does not remove it; only a better explanation of the day (more informative layers) shrinks it.
- Parameter uncertainty. The trend slope, the weekday weights and the holiday effect were learned from a limited history, so they are known only roughly. The posterior describes exactly this, and more data shrinks it.
- Model uncertainty. The structure may be wrong: maybe the trend is about to flatten, the seasonality is changing, a competitor appeared. The posterior is conditional on the model being right, so it says nothing about this. Cures: compare models, average them, run scenarios, check the model (end of this chapter).
- Exogenous-regressor uncertainty. To forecast with a regressor (planned e-mails, price, temperature) you need its future values. If they are known by plan, no problem. If they must be forecast themselves, their errors pass through the coefficient into your forecast (Chapter 7.12).
Three ways to say it:
- Picture: four nested bands: noise in the middle, parameter doubt around it, then regressor doubt, then model doubt on the outside.
- Numbers: noise sd 12, parameters 8, regressors 6, model 6: variances $144 + 64 + 36 + 36 = 280$, total sd 16.7. A band built from noise and parameters alone has sd 14.4.
- Slogan: noise you cannot shrink, parameters you can, model and regressors you must not forget.
One future day. The base model predicts 200 orders. Each source separately:
- Observation noise: sd 12, so variance $144$.
- Parameters: the posterior of the expected value on that day has sd 8, variance $64$.
- Regressor: the model says 4 orders per 1 000 marketing e-mails, and the e-mail volume for that day is not known better than $\pm1.5$ thousand: effect sd $4 \times 1.5 = 6$, variance $36$.
- Model: a second plausible model (for example one that lets the trend flatten) predicts 212 instead of 200. With equal belief in the two, the mixture mean is $206$ and the between-model variance is $\big(\tfrac{212-200}{2}\big)^2 = 36$.
- If the sources are independent, variances add: $144 + 64 + 36 + 36 = 280$, sd $\sqrt{280} = 16.7$. The 90% prediction interval is $206 \pm 1.645 \times 16.7 = [178.5,\ 233.5]$.
- A forecast that included only noise and parameters from the first model has variance $208$, sd $14.4$, interval $200 \pm 23.7 = [176.3,\ 223.7]$: 14% narrower and 10 orders too low at the top end.
- Shares of the total variance: noise 51%, parameters 23%, regressor 13%, model 13%. Collecting more history would shrink only the parameter share; the noise floor and the model and regressor shares stay.
By the law of total variance (Chapter 4.6), conditioning on everything that is unknown except the noise,
$$Var(\tilde y\mid D) = \underbrace{E\big[Var(\tilde y\mid \theta, M, x)\big]}_{\text{observation noise}} + \underbrace{Var\big(E[\tilde y\mid \theta, M, x]\big)}_{\text{everything else}},$$and the second term splits (approximately, when the sources are independent) into parameters $\theta$, regressors $x$ and model $M$:
| Source | What is uncertain | In the posterior predictive? | Shrinks with more data? | How to include it |
|---|---|---|---|---|
| Observation noise | day-to-day scatter around $\mu_t$ | yes (the likelihood) | no (a floor) | choose the likelihood well (7.13) |
| Parameters | trend, $\delta_j$, season, holiday and regressor weights, $\sigma$, $\alpha$ | yes (posterior draws) | yes | use the posterior, not a point estimate |
| Model | structure: trend shape, future changepoints, seasonality that drifts, wrong likelihood | no, unless built in (for example simulated future trend changes) | not by itself | model comparison, averaging, scenarios, PPC |
| Exogenous regressors | future values of $X_{T+h}$ (when not known in advance) | no: the regressor is treated as given | no | simulate future $X$ from its own forecast; or run named scenarios |
- "Additive variances" is an approximation: the sources can interact (for example, a poorly known slope matters more if the model is also unsure about the future trend).
- Keeping a regressor fixed at a planned value gives a conditional forecast ("given this plan"). Simulating it gives an unconditional one. Both are legitimate; say which you report.
Why do we need it?
Knowing which source dominates tells you what to do. Wide because of noise: accept it or explain more. Wide because of parameters: collect more history or use stronger priors. Honest bands need model and regressor doubt too, or the intervals will be too narrow when they matter most.
Where is it used?
Prophet-style interval settings (it simulates trend uncertainty), scenario planning with different future promotions, ensemble forecasts that average several models, weather and energy forecasting where input forecasts are themselves uncertain, and risk reports that split "known unknowns" from "model risk".
How is it used?
Simulate with sources switched on one at a time and compare band widths: only noise (fix parameters at their posterior mean), noise plus parameters (full posterior), then add draws of future regressors, then mix models. The differences show each source's contribution.
"The posterior predictive interval contains all the uncertainty."
It contains the uncertainty inside the model you fitted: parameters and observation noise. Model uncertainty and future-regressor uncertainty are outside it unless you build them in. This is why a well-fitted Bayesian model can still produce intervals that are too narrow.
"More data will make my forecast interval shrink to zero."
Only the parameter part shrinks. The noise floor stays, and model and regressor doubt do not vanish with more history of the same kind.
"The posterior of the coefficient on the regressor already includes the uncertainty of the regressor's future values."
It describes how strongly the regressor affected the past, not how well you will know the regressor in the future. Those are separate sources.
"Our Bayesian model gives full uncertainty quantification."
"Our intervals include parameter uncertainty and observation noise under the assumed model. We handle regressors by [planned values / simulated forecasts / scenarios] and we assess model risk with posterior predictive checks and a rolling-origin backtest."
Model answer: "I separate four sources: observation noise, parameter uncertainty (the posterior), model uncertainty and exogenous-regressor uncertainty. The first two come out of the posterior predictive. The last two I add by simulating future regressors and by comparing or averaging model variants, and I validate the resulting coverage on held-out periods."
In your forecasting model the posterior draws cover the parameter source: the base slope and changepoint adjustments $\delta_j$, Fourier weights, holiday coefficients and regressor weights. The likelihood (Normal, Student-t or Negative Binomial) covers observation noise. Two things are easy to miss: future regressor values (do you pass planned values, or simulate them?) and the possibility that the trend changes after the end of the history, which a model that keeps the last fitted slope treats as certain (next section). Also note that with SVI, a too-simple guide can under-state parameter correlations and make the parameter band too narrow (Chapter 6.11, 6.13).
$Var(\tilde y\mid D) = E[Var(\tilde y\mid\cdot)] + Var(E[\tilde y\mid\cdot])$: observation noise + (parameters + regressors + model), adding when independent.
Example: $144 + 64 + 36 + 36 = 280$, sd 16.7, 90% interval [178.5, 233.5]; noise + parameters only: sd 14.4.
Posterior predictive covers noise and parameters. Model and regressor uncertainty must be added on purpose.
Quick check: a forecast has noise variance 100 and parameter variance 25. A second model that disagrees adds a between-model variance of 25. What fraction of the total variance is parameter uncertainty, and what is the total sd?
Total variance $= 100 + 25 + 25 = 150$, sd $= 12.2$. Parameter share $= 25/150 = 16.7\%$. Without the model term the sd would have been $\sqrt{125} = 11.2$, so ignoring model disagreement understates the sd by about 9%.
Why forecast intervals widen with the horizon (and why sometimes they do not) core
Think of steering a ship by compass. A tiny error in the heading hardly matters after one kilometre, but after a hundred kilometres you are far from where you meant to be. A forecast is the same: small errors in direction grow with distance. That is why forecast fans open up like a cone as the horizon grows.
There are four different reasons a band can widen (or not), and they grow at different speeds:
- Independent daily noise adds a band of the same width on every day. It does not widen with the horizon at all.
- An uncertain slope (parameter uncertainty): an error $\varepsilon$ in the slope becomes $h\varepsilon$ after $h$ days. The standard deviation grows in proportion to $h$.
- Future trend changes: the trend may bend again after the data end. Every possible bend early in the horizon affects all later days, so the standard deviation grows like $h^{1.5}$, faster than linear.
- Noise that carries over (a random walk, or an autoregressive error): yesterday's shock is still partly here. A random walk's standard deviation grows like $\sqrt h$; a mean-reverting AR(1) grows and then levels off.
Three ways to say it:
- Picture: a cone that opens: a thin tube for noise, a straight cone for slope doubt, a trumpet that flares for future trend changes.
- Numbers: noise sd 10, level sd 2, slope sd 0.05 per day, three possible trend changes per 100 days of size 0.15: total sd is 10.2 tomorrow, 10.9 in 30 days and 21.3 in 90 days.
- Slogan: small errors in direction grow with distance.
A forecast model with: daily noise sd $\sigma = 10$ (independent); level uncertainty at the forecast origin sd 2; slope uncertainty sd $0.05$ orders per day; future slope changes arriving at a rate of 3 per 100 days ($\lambda = 0.03$ per day), each change Laplace distributed with scale $b = 0.15$ (variance $2b^2 = 0.045$). Compute the predictive sd at $h = 1, 30, 90$ days.
- Noise: variance $100$ at every horizon.
- Parameters (level and slope): variance $2^2 + (h \times 0.05)^2$. At $h = 1$: $4.0$. At $h = 30$: $4 + 2.25 = 6.25$. At $h = 90$: $4 + 20.25 = 24.25$.
- Future trend changes: a change at day $s$ of size $\delta$ shifts the trend at day $h$ by $\delta\,(h - s)$. With changes arriving at rate $\lambda$, the variance is $\lambda \cdot 2b^2 \int_0^h (h-s)^2\,ds = \lambda \cdot 2b^2 \cdot h^3/3$. At $h = 1$: $0.03 \times 0.045/3 = 0.0005$ (nothing). At $h = 30$: $0.03 \times 0.045 \times 27\,000/3 = 12.15$. At $h = 90$: $0.03 \times 0.045 \times 729\,000/3 = 328.05$.
- Add the variances: $h=1$: $100 + 4.0 + 0.0 = 104.0$, sd $10.2$. $h=30$: $100 + 6.25 + 12.15 = 118.4$, sd $10.9$. $h=90$: $100 + 24.25 + 328.05 = 452.3$, sd $21.3$.
- The 80% half-width is $1.28 \times$ sd: $\pm13.1$ tomorrow, $\pm13.9$ in 30 days, $\pm27.3$ in 90 days. For the first month the band is almost flat; the cone only opens later, and most of the opening at 90 days comes from the possibility of future trend changes (72% of the variance), not from the fitted slope's uncertainty.
- Compare with other noise models for the same $\sigma$-scale: a random walk with step sd 2.5 has sd $2.5\sqrt h$ = 2.5, 13.7, 23.7 at $h = 1, 30, 90$; an AR(1) with $\phi = 0.9$ and innovation sd 4 has 4.0, 9.2, 9.2: it widens fast, then settles at its long-run sd.
For independent sources the predictive variance at horizon $h$ is the sum
$$Var(\tilde y_{T+h}\mid D) = \underbrace{\sigma^2}_{\text{noise}} + \underbrace{s_a^2 + h^2 s_b^2}_{\text{level and slope}} + \underbrace{\lambda\,\tfrac{2b^2}{3}\,h^3}_{\text{future trend changes}},$$where $s_a$, $s_b$ are the posterior sds of the level and slope at the origin (assumed independent here; in a fitted model they are correlated through $Cov(a, b)$), $\lambda$ is the rate of future changes per day, and $2b^2$ is the variance of one Laplace$(0, b)$ slope change.
| Mechanism | Standard deviation at horizon $h$ | Shape |
|---|---|---|
| independent noise | $\sigma$ | flat |
| uncertain slope | $\sqrt{s_a^2 + h^2 s_b^2}$ | about linear in $h$ |
| future trend changes | $b\sqrt{2\lambda/3}\;h^{3/2}$ | flares (power 1.5) |
| random-walk noise, step sd $s$ | $s\sqrt h$ | square-root growth |
| AR(1) noise, innovation sd $\sigma_e$, coefficient $\phi$ | $\sigma_e\sqrt{\dfrac{1-\phi^{2h}}{1-\phi^2}}$ | grows, then levels off |
- Prophet-style models simulate future slope changes with the same average frequency and Laplace scale seen in the history, which is how they get the flaring term. A model that keeps the last fitted slope treats the future trend as certain and gives a narrower, more optimistic cone.
- Which model families give which shape (ARIMA with differencing, ETS, regression with time features) is discussed in Chapter 7.6.
Why do we need it?
Planning horizons are different (next week's staffing, next quarter's budget), and the honest width changes between them. Knowing the shape also lets you spot a broken model: bands that stay flat for 90 days are probably ignoring something.
Where is it used?
Fan charts in central-bank and weather forecasting, Prophet's uncertainty_samples and interval_width, ARIMA and ETS forecast intervals (random-walk-like growth), long-range capacity plans, and horizon-specific evaluation (Chapter 7.15).
How is it used?
Compute the predictive sd (or the band width) at several horizons from your draws and compare with the shape you expect. Check calibration separately at short and long horizons: a long-horizon interval that is too narrow is the usual failure.
"Intervals widen because the noise piles up day after day."
In a trend model with independent daily noise, nothing piles up: every day has the same noise. The widening comes from parameter uncertainty and possible future trend changes. Noise does pile up in random-walk and ARIMA-with-differencing models, where today's shock persists.
"A narrow long-horizon band means a confident, accurate model."
It may mean the model treats the future trend as known. Check calibration at long horizons on held-out data (Chapter 7.15, Chapter 7.16).
"The band at 90 days is just the 1-day band stretched."
The mechanisms have different growth laws (flat, linear, power 1.5, square root), so the shape of the cone is informative: it shows which sources dominate.
In your forecasting model the interval at a long horizon depends on three things you can inspect: the posterior spread of the base slope and of the changepoint slope changes $\delta_j$ (the parameter term), whether your forecast code simulates new changepoints after the end of the history (the flaring term) or holds the last slope fixed, and the likelihood (the flat term). The Laplace scale $b$ on the $\delta_j$ also matters: a larger $b$ fitted on the history means larger simulated future bends if your model reuses it (Chapter 7.10). If your long-horizon intervals fail a coverage check, the flare is the first suspect.
$Var(\tilde y_{T+h}) = \sigma^2 + s_a^2 + h^2 s_b^2 + \lambda\frac{2b^2}{3}h^3$. Growth: noise flat; slope $\propto h$; future changes $\propto h^{3/2}$; random walk $\propto\sqrt h$; AR(1) levels off.
Example: sd 10.2 at $h=1$, 10.9 at 30, 21.3 at 90; future changes are 72% of the variance at 90 days.
Trap: noise does not "pile up" in a trend model with independent noise.
Quick check: the slope has posterior sd 0.2 orders per day (and nothing else is uncertain). By how much does the parameter sd grow between 10 and 40 days ahead?
Parameter sd $= h \times 0.2$: 2 orders at $h = 10$ and 8 orders at $h = 40$. Four times the horizon gives four times the sd (linear growth). A random-walk noise term would grow only by a factor of $\sqrt4 = 2$, and a future-trend-change term by a factor of $4^{1.5} = 8$.
Posterior predictive checks for forecasts: can the fitted model fake your series? core
Remember the forger test of Chapter 6.8: hide the real painting among the forger's copies and see whether an expert can pick it out. For a forecasting model the "painting" is your history, and the forger is the fitted model. Ask the model to write fake histories: same dates, same weekdays, same regressors, same length, each from a different posterior draw. If the model has understood the series, the real one should look like just another fake.
"Look like" must be made concrete. For a time series we compare summaries that matter for forecasting: the average level, the spread, the extremes, the number of zero days, the weekly shape, the day-to-day memory (autocorrelation) and whether the level shifted. Each summary probes one assumption of the model, so when one fails you know which part to repair. A model can pass five checks and fail the sixth; the sixth is the one that tells you something.
Three ways to say it:
- Picture: a lineup of fake histories with the real one hidden among them; you look for features that make it stand out.
- Numbers: on a bursty series the fitted model's fake histories have a lag-1 autocorrelation between $-0.14$ and $0.22$; the real series has $0.57$. Its mean, variance and maximum all look fine.
- Slogan: a good model can fake your data in every way you care about.
12 weeks (84 days) of daily orders. The model: Negative Binomial with a log link, a linear trend and a weekly Fourier pattern, fitted with NUTS. The real series has "bursts": runs of busy days and runs of quiet days that the model does not know about (its noise is independent day to day). This is exactly the run in the Code-it block at the end of the chapter.
- Take 500 posterior draws $\theta^{(s)}$ (trend, weekly weights, concentration $\alpha$). For each, simulate a whole 84-day series $y^{\text{rep}(s)}$ from $\text{NB}(\mu_t^{(s)}, \alpha^{(s)})$ on the same calendar.
- For the real series and for every fake series compute five summaries: mean, variance, maximum, weekend premium (Sat+Sun average divided by Mon–Fri average) and lag-1 autocorrelation.
- Mean: real $128.5$; fakes (middle 90%) $[113.5,\ 149.5]$; $P(\text{fake} \ge \text{real}) = 0.53$. Fine. Variance: real $6\,228$; fakes $[3\,504,\ 10\,200]$; $0.39$. Fine. Maximum: real $465$; fakes $[287,\ 599]$; $0.22$. Fine. Weekend premium: real $1.07$; fakes $[0.79,\ 1.44]$; $0.47$. Fine.
- Lag-1 autocorrelation: real $0.57$; fakes $[-0.14,\ 0.22]$; $P(\text{fake} \ge \text{real}) = 0.00$. None of the 500 fake series is as "sticky" as the real one. A clear misfit.
- Reading: the model reproduces the level, the spread, the extremes and the weekly shape, but not the persistence. The cause is its independent-noise assumption (days are independent given $\mu_t$). Fixes to try: an autoregressive error term (today's noise partly carries over from yesterday's), a local-level (state-space) component (a hidden level that drifts from day to day), or, if the bursts are really events, regressors for them. Then run the checks again.
- A variance-only check would have passed this model. And on a series generated from the model's own world, the same five checks give $P$-values between $0.14$ and $0.57$: no flag.
Recipe (Chapter 6.8 gives the general version): for $s = 1, \dots, S$ take a posterior draw $\theta^{(s)}$ and simulate $y^{\text{rep}(s)} \sim p(y \mid \theta^{(s)})$ with exactly the structure of the real data (same time points, same regressors, same group sizes). Pick summaries $T(\cdot)$ and compare $T(y)$ with the spread of $T(y^{\text{rep}(s)})$. The posterior predictive $p$-value is
$$p_T = P\big(T(y^{\text{rep}}) \ge T(y)\mid D\big) \approx \frac1S\#\{s: T(y^{\text{rep}(s)}) \ge T(y)\}.$$- $p_T$ near 0 or 1 (say below 0.025 or above 0.975, a convention) means the real value is out in the tail of what the model produces: a misfit. It is a diagnostic, not a hypothesis test; mid-range values are what a good model gives.
- For a discrete summary (number of zero days) ties are common; count half of the ties.
| Summary $T$ | What it probes | If it fails, look at |
|---|---|---|
| mean / level | trend, regressors, holiday sizes | missing changepoint or regressor; wrong link |
| variance, sd | noise size, dispersion | Poisson instead of NB; heteroscedastic noise |
| maximum, upper quantiles (tails) | extreme days, capacity risk | Student-t or NB; event regressors for spikes |
| number of zero days | zero frequency | closed days (regressor), zero inflation, hurdle |
| weekend premium, seasonal means by weekday or month | seasonality shape | Fourier order too low; missing seasonality; changing seasonality |
| lag-1 and lag-7 autocorrelation | serial dependence left in the noise | AR/state-space error term, richer seasonality, event regressors |
| last weeks vs first weeks | level shifts, trend breaks | more or better-placed changepoints (7.8–7.10) |
- Choose summaries by the decision the forecast serves: upper quantiles for capacity, weekday shape for staffing, autocorrelation if you will sum days into weeks.
- An in-sample PPC cannot see extrapolation problems (a trend that will break after the data end); that is what the held-out check in the next section is for.
Why do we need it?
A posterior exists for any model, even a bad one, and its forecast intervals will look smooth either way. The check is the standard way to learn that the model is missing something and what it is missing, before you trust the bands.
Where is it used?
The Bayesian workflow in Stan, PyMC and NumPyro (Predictive, ArviZ plot_ppc and plot_bpv), time-series model criticism for demand and traffic, and model reviews that ask "what did you check besides RMSE?".
How is it used?
Call Predictive(model, posterior_samples)(key, t, F) without the observed values to get y_rep of shape (draws, days). Write the summaries as functions of one series, apply them to the data and to each row of y_rep, and plot the observed value on top of the histogram of the fake values.
"The model passed the PPC, so it is correct."
It passed these summaries. A PPC finds misfit; it cannot prove a model right. Choose summaries that match the decisions and the assumptions you are worried about, and add a held-out check (next section).
"A $p$-value of 0.03 means the model is rejected at the 5% level."
The posterior predictive $p$-value is a diagnostic. It is not uniformly distributed under the model and is often conservative for summaries the model fits directly (like the mean). Read the picture: how far out is the real value, and does it matter for your use?
"My residual plot looks fine, so I do not need the PPC."
Residual plots (Chapter 7.17) and PPCs overlap, but the PPC speaks the language of the likelihood: it checks zero counts, extremes and dispersion on the original scale, which residual plots can hide.
"One summary that fails means throw the model away."
It tells you which assumption to repair. A bursty series is often fixed with one extra component, not a new model class.
"The posterior predictive p-value is the probability that our model is true."
"We ran a posterior predictive check, and it passed, so the forecast is validated."
"A posterior predictive check compares the observed value of a chosen summary with its distribution in data simulated from the fitted model. It can reveal misfit; it cannot validate a model."
Model answer: "I simulate replicated series from posterior draws on the same calendar and regressors, and compare summaries that matter for the forecast: level, spread, tails, zero days, weekly shape, autocorrelation and level shift. Each failure points to a fix, such as an AR or state-space term or a zero-inflated likelihood. Passing these checks means the model can imitate the history in those respects; I then test the forecast itself with a rolling-origin backtest."
For your forecasting model, run these checks after every fit, using the same calendar and regressors as the data: the mean and variance (level and noise size), the maximum or upper quantiles (does the likelihood have the right tails: Normal versus Student-t?), the number of zero days (for a Negative Binomial on small counts), weekday and yearly shape (is the Fourier order enough?), lag-1 and lag-7 autocorrelation (the model has independent noise, so serial dependence will show up here, Chapter 7.3), and last-vs-first level (a missed changepoint). With SVI the replicates should come from draws of the guide: if the guide is too narrow, the fake series will be too alike each other and the checks will flag problems that are really inference problems. Prior sensitivity and identifiability (the other half of Chapter 6.8) are checks on the model's inputs rather than its outputs.
PPC: draw $\theta^{(s)}$, simulate $y^{\text{rep}(s)}$ on the real calendar, compare $T(y)$ with $T(y^{\text{rep}})$; $p_T = P(T(y^{\text{rep}}) \ge T(y)\mid D)$; near 0 or 1 = misfit.
Summaries: mean, variance, max/tails, zero days, weekend/seasonal shape, lag-1 and lag-7 autocorrelation, last-vs-first level; each failure points to a fix.
Example: lag-1 real 0.57 vs fakes [−0.14, 0.22] while mean, variance and max pass: independent noise assumption violated.
Quick check: the real series has 9 zero days; across 1 000 fake series from your Negative Binomial model, the number of zero days is 0 in 940 series, 1 in 50 series, 2 in 10 series. What is the mid-$p$ for this summary, and what does it suggest?
Fake series with more zero days than 9: none. Ties at 9: none. So $p = 0$. The model essentially never produces nine zero days, while the real data has them: an excess-zeros problem. Look for a cause (closed days, stock-outs) and use a regressor or mask; otherwise a zero-inflated or hurdle likelihood (Chapter 7.13).
Checking the forecast itself: where did reality land? core
A posterior predictive check on the history asks "can the model fake the past?" But you use the model for the future, and a model can fake its past perfectly and still extrapolate badly: a trend that bends after the data end is invisible to any in-sample check. So also ask the question the way you will use the model: pretend to be at an earlier forecast origin. Fit using only the days before it, forecast the days that followed, and then look at where the real values landed inside their predictive distributions.
If the forecast is honest, reality should land everywhere in the predictive distribution, about equally: one day in ten in the lowest tenth, one in ten in the highest tenth, and so on. An 80% interval should contain about 80% of the real values. If real values keep falling outside, in the tails, the forecast was too confident; if they cluster in the middle, it was too timid. This is a posterior predictive check for the future, and it is the first step toward the full backtests of Chapter 7.15 and the calibration tools of Chapter 7.16.
Three ways to say it:
- Picture: darts thrown at a target drawn by the forecast: they should scatter over the whole board, not pile up at the rim.
- Numbers: an 80% interval should hold about 80% of the real days; one that holds 20% is badly overconfident.
- Slogan: does reality land where the forecast said it might?
The model forecast the next five days as $N(150, 12^2)$ each, and the real orders were $171, 150, 138, 162, 141$. For each real value compute the PIT ("probability integral transform"): the share of the forecast distribution lying below it.
- Day 1: $z = (171-150)/12 = 1.75$, PIT $= \Phi(1.75) = 0.960$. Day 2: $z = 0$, PIT $0.500$. Day 3: $z = -1$, PIT $0.159$. Day 4: $z = 1$, PIT $0.841$. Day 5: $z = -0.75$, PIT $0.227$.
- The 80% interval is $150 \pm 1.28 \times 12 = [134.6,\ 165.4]$. Days 2, 3, 4 and 5 are inside; day 1 (171) is outside, above. Coverage $4/5 = 80\%$. The PITs are spread over $(0,1)$: no sign of trouble (five points can show little).
- Now suppose the model had been overconfident, forecasting $N(150, 6^2)$ for the same days. The PITs become $0.9998,\ 0.500,\ 0.023,\ 0.977,\ 0.067$: four of five in the outer tails. The 80% interval $[142.3,\ 157.7]$ contains only day 2: coverage $1/5 = 20\%$.
- So a U-shaped pile-up of PITs near 0 and 1, with coverage far below the nominal level, is the signature of intervals that are too narrow. Piling in the middle would mean intervals too wide.
For each held-out day $j$ with predictive CDF $F_j$ and real value $y_j$, the PIT is $u_j = F_j(y_j) = P(\tilde y_j \le y_j \mid D_{\text{before the origin}})$. In terms of draws, $u_j \approx \frac1S\#\{s: \tilde y^{(s)}_j \le y_j\}$ (for counts, split ties). If the forecasts are calibrated (the probabilities they state match how often things really happen), the $u_j$ look like draws from Uniform$(0,1)$. The coverage of the central 80% interval is the share of days with $0.1 \le u_j \le 0.9$.
- Backtest recipe: choose an origin $T_0$; fit on days up to $T_0$ only (including any preprocessing, scaling and changepoint detection, Chapter 7.12); forecast days $T_0 + 1, \dots, T_0 + h$; record $u_j$ and whether the real value fell in each interval; move the origin and repeat.
- Daily errors are autocorrelated and the days share parameters, so one 28-day window gives a noisy coverage estimate (it can easily be 60% or 95% for a calibrated model). Use many origins and report coverage by horizon.
- Typical patterns: too many PITs near 0 and 1 means overconfident (too narrow); a hump in the middle means timid (too wide); PITs mostly near 1 (or 0) means biased (forecasts too low or too high), for example when a trend bends.
Why do we need it?
The model's own bands are only statements; the held-out days are the test. Checking the forecast the way you will use it finds extrapolation problems, such as an ignored trend break or underestimated volatility, that no in-sample check can see.
Where is it used?
Rolling-origin backtests in demand-forecasting teams, forecast competitions that score coverage and calibration, central-bank fan-chart reviews, and monitoring dashboards that track "share of actuals inside the 80% band" over time (Chapter 7.19).
How is it used?
For each past origin refit (or reuse a rolling fit), forecast the next $h$ days, compute the PIT or the 10/90% interval hit for each real value, and aggregate over origins and horizons. Look at coverage by horizon and at the PIT histogram, not only at an overall score.
"The 80% band contained 17 of 28 days, so the model is miscalibrated."
One window of 28 correlated days is a weak test: even a calibrated model can show 60% or 95%. Collect many origins and horizons before you conclude (Chapter 7.15).
"My PPC on the history passed, so the forecast intervals are fine."
The history check cannot see the future. A trend that bends after the origin, a new holiday pattern or a volatility change are invisible until you compare forecasts with what happened.
"I tuned my model on the whole series and then checked coverage on the last 28 days."
If those 28 days influenced the fitting, the model's changepoints, scaling or regressor choice, the check is contaminated (Chapter 7.12, leakage). Fit on days before the origin only.
Your forecasting model produces draws for the next days, so every past origin gives you the PIT and interval hits for free. Compute them for several origins and report coverage by horizon: if the 80% band holds 80% of the next-day values but only 55% of the values 4 weeks ahead, the long-horizon flare of the previous section is too small (for example, future trend changes are not simulated). If coverage is low for one likelihood and good for another, you have a reason to choose between Normal, Student-t and Negative Binomial that is independent of the training loss. Remember that changepoint detection and scaling must also be done using days before each origin.
PIT $u_j = F_j(y_j) \approx$ share of predictive draws below the real value. Calibrated: $u_j \sim$ Uniform$(0,1)$; 80% band holds about 80% of the real days.
Pile-up near 0 and 1: too narrow. Hump in the middle: too wide. All near 0 or 1: biased.
Use many origins and horizons; fit only on days before the origin. Full tools: Chapters 7.15 and 7.16.
Quick check: for a forecast you observe 100 days over many origins and only 62 fall inside the 80% band, with 25 below and 13 above. What is the problem, and which direction does it lean?
Coverage 62% against 80%: the intervals are too narrow (overconfident). Misses are lopsided: 25 below versus 13 above, so the forecast also leans too high (real values fall below the band about twice as often as above). Both the width and the centre deserve attention.
Recap, cheat sheet and practice
- A point forecast is one summary of a predictive distribution $p(\tilde y\mid D)$. Sample it: draw $\theta^{(s)}$, then $\tilde y^{(s)}$. Mean = expected demand, median = typical day, quantiles = intervals and stock levels, share of draws above $c$ = $P(\text{demand} \gt c)$ (Monte Carlo error $\sqrt{p(1-p)/S}$).
- Keep whole paths (rows of the draws array) for multi-day questions: weekly totals, "any day above capacity". Means add across days; quantiles and variances do not (shared parameters link the days).
- The action depends on the costs: the best stock is the quantile $c_u/(c_u + c_o)$ of the predictive distribution (mean for squared loss, median for absolute loss).
- Four sources: observation noise (a floor), parameter uncertainty (the posterior; shrinks with data), model uncertainty and exogenous-regressor uncertainty (both outside the posterior unless added: model averaging, scenarios, simulated future regressors). Variances add when independent.
- Horizon: independent noise is flat; slope doubt grows $\propto h$; future trend changes $\propto h^{3/2}$; random-walk noise $\propto \sqrt h$; AR(1) noise levels off. Small errors in direction grow with distance.
- Posterior predictive checks replicate whole series on the real calendar and compare summaries: mean, variance, tails and maximum, zero days, weekly shape, lag-1/lag-7 autocorrelation, level shift. Each failure points to a fix. They find misfit; they never prove a model right.
- Check the forecast itself at past origins: coverage of the 80% band and PIT values (uniform if calibrated). One window is noisy; use many origins and horizons (Chapters 7.15, 7.16).
Cheat sheet
| Question | Answer from the draws ytil[draws, days] | Watch out for |
|---|---|---|
| expected demand day $j$ | ytil[:, j].mean() | a prediction mean, not a typical day if skewed |
| typical day | np.median(ytil[:, j]) | below the mean for right-skewed counts |
| 90% prediction interval | np.quantile(ytil[:, j], [0.05, 0.95]) | wider than a credible interval for $\mu$ |
| $P(\text{demand} \gt c)$ on day $j$ | (ytil[:, j] > c).mean() | error $\sqrt{p(1-p)/S}$ |
| some day in a window exceeds $c$ | (ytil[:, a:b] > c).any(axis=1).mean() | needs whole paths; not $1-\prod(1-p_j)$ |
| total over a window | ytil[:, a:b].sum(axis=1) then quantiles | quantiles do not add across days |
| best stock for costs $c_u, c_o$ | np.quantile(total, c_u / (c_u + c_o)) | needs a calibrated right tail |
| parameters only | draw mu (the mean), not obs | compare band widths to see each source |
| PPC statistic | f(y) vs [f(r) for r in y_rep]; $p = P(f(\text{rep}) \ge f(y))$ | same calendar and regressors; near 0 or 1 = misfit |
| forecast check on held-out days | PIT $=$ (ytil < y_actual).mean(axis=0); coverage of 10/90% band | one window is noisy; use many origins |
import numpy as np, jax, jax.numpy as jnp
import numpyro, numpyro.distributions as dist
from numpyro.infer import MCMC, NUTS, Predictive
# 1) Simulate 12 weeks of daily orders (trend + weekly pattern, Negative Binomial noise, log link) and keep 14 days as the "future"
rng = np.random.default_rng(3)
T, H = 84, 14
t = np.arange(T + H)
def fourier(t, period, K):
a = 2 * np.pi * np.outer(t, np.arange(1, K + 1)) / period
return np.concatenate([np.sin(a), np.cos(a)], axis=1)
F = fourier(t, 7, 2) # weekly Fourier features (order 2)
eta_true = 4.5 + 0.004 * t + F @ np.array([-0.10, 0.04, -0.15, 0.03])
y_all = rng.negative_binomial(12, 12 / (12 + np.exp(eta_true))) # true concentration 12
y, y_future = y_all[:T], y_all[T:]
# 2) Model: log link + NB2 likelihood; fit with NUTS (a few hundred draws is plenty here)
def model(tt, F, y=None):
a = numpyro.sample("a", dist.Normal(4.5, 1.0))
b = numpyro.sample("b", dist.Normal(0.0, 0.5)) # trend per 100 days
w = numpyro.sample("w", dist.Normal(0.0, 0.5).expand([F.shape[1]]))
alpha = numpyro.sample("alpha", dist.Gamma(2.0, 0.1)) # concentration
mu = numpyro.deterministic("mu", jnp.exp(a + b * tt / 100.0 + F @ w))
numpyro.sample("obs", dist.NegativeBinomial2(mu, alpha), obs=y)
mcmc = MCMC(NUTS(model), num_warmup=300, num_samples=500, progress_bar=False)
mcmc.run(jax.random.PRNGKey(0), jnp.array(t[:T]), jnp.array(F[:T]), y=jnp.array(y))
post = mcmc.get_samples()
keep = {k: post[k] for k in ("a", "b", "w", "alpha")}
# 3) Posterior predictive for the next 14 days: one row per draw = one whole future PATH
out = Predictive(model, posterior_samples=keep)(jax.random.PRNGKey(1), jnp.array(t[T:]), jnp.array(F[T:]))
ytil, mu_f = np.asarray(out["obs"]), np.asarray(out["mu"]) # (500, 14) each
capacity = 150
d = ytil[:, 0] # tomorrow
print(f"tomorrow: mean {d.mean():.1f}, median {np.median(d):.0f}, 90% interval {np.quantile(d, [0.05, 0.95])}, P(> {capacity}) = {(d > capacity).mean():.3f}")
print(f"P(some day in the next 7 exceeds {capacity}) = {(ytil[:, :7] > capacity).any(axis=1).mean():.3f} (needs whole paths)")
wk = ytil[:, :7].sum(axis=1)
print(f"next-7-day total: mean {wk.mean():.0f}, 90% interval {np.quantile(wk, [0.05, 0.95])}")
# 4) How wide? parameter uncertainty alone (the mean mu) versus the full predictive (adds observation noise)
for k in (0, 13):
w_mu = np.diff(np.quantile(mu_f[:, k], [0.05, 0.95]))[0]
w_y = np.diff(np.quantile(ytil[:, k], [0.05, 0.95]))[0]
print(f"day {k + 1:2d} ahead: 90% width of the MEAN {w_mu:5.1f} | 90% width of the full predictive {w_y:5.1f}")
# 5) Check the forecast itself on the 14 days we held back: PIT = share of predictive draws below the actual
pit = (ytil < y_future).mean(axis=0) + 0.5 * (ytil == y_future).mean(axis=0)
lo, hi = np.quantile(ytil, [0.1, 0.9], axis=0)
print("80% interval coverage on the held-out days:", ((y_future >= lo) & (y_future <= hi)).mean().round(2), "| PIT values:", pit.round(2))
# 6) Posterior predictive checks on the HISTORY: replicate 84-day series, compare statistics with the observed series
rep = np.asarray(Predictive(model, posterior_samples=keep)(jax.random.PRNGKey(2), jnp.array(t[:T]), jnp.array(F[:T]))["obs"])
def lag1(x):
x = x - x.mean(); return (x[1:] * x[:-1]).sum() / (x * x).sum()
def weekend_premium(x): # (Sat+Sun mean) / (Mon-Fri mean); day 0 is a Monday
wd = np.arange(len(x)) % 7; return x[wd >= 5].mean() / x[wd < 5].mean()
stats_ = {"mean": np.mean, "variance": np.var, "max": np.max, "weekend premium": weekend_premium, "lag-1 autocorrelation": lag1}
def ppc(series, title):
print(title)
for name, f in stats_.items():
obs, sim = f(series), np.array([f(r) for r in rep])
print(f" {name:22s} observed {obs:8.2f} replicated 90% range [{np.quantile(sim, 0.05):7.2f}, {np.quantile(sim, 0.95):7.2f}] P(replicated >= observed) = {(sim >= obs).mean():.2f}")
ppc(y, "PPC, data from the model's own world:")
# 7) Same fitted model, but the real series has persistent bursts (autocorrelated noise the model does not have)
u = np.zeros(T)
for i in range(1, T): u[i] = 0.85 * u[i - 1] + 0.25 * rng.standard_normal()
y_burst = rng.negative_binomial(12, 12 / (12 + np.exp(eta_true[:T] + u)))
mcmc.run(jax.random.PRNGKey(0), jnp.array(t[:T]), jnp.array(F[:T]), y=jnp.array(y_burst))
keep = {k: mcmc.get_samples()[k] for k in ("a", "b", "w", "alpha")}
rep = np.asarray(Predictive(model, posterior_samples=keep)(jax.random.PRNGKey(2), jnp.array(t[:T]), jnp.array(F[:T]))["obs"])
ppc(y_burst, "PPC, same model but the real series has bursts:")
# Output of this script (about 6 seconds on a laptop):
# tomorrow: mean 118.2, median 111, 90% interval [ 66.9 182.05], P(> 150) = 0.166
# P(some day in the next 7 exceeds 150) = 0.910 (needs whole paths)
# next-7-day total: mean 924, 90% interval [ 738.85 1138.15]
# day 1 ahead: 90% width of the MEAN 38.2 | 90% width of the full predictive 115.1
# day 14 ahead: 90% width of the MEAN 40.3 | 90% width of the full predictive 118.0
# 80% interval coverage on the held-out days: 1.0 | PIT values: [0.68 0.21 0.12 0.83 0.74 0.46 0.32 0.76 0.62 0.3 0.5 0.16 0.66 0.54]
# PPC, data from the model's own world:
# mean 109.80 [101.74, 120.08] p 0.51 | variance 1600.69 [1022.90, 2404.82] p 0.44 | max 284 [191, 319] p 0.14
# weekend premium 0.95 [0.82, 1.14] p 0.57 | lag-1 autocorrelation 0.12 [-0.06, 0.30] p 0.56 (nothing flagged)
# PPC, same model but the real series has bursts:
# mean 128.46 [113.48, 149.50] p 0.53 | variance 6228.25 [3503.67, 10199.96] p 0.39 | max 465 [286.95, 598.60] p 0.22
# weekend premium 1.07 [0.79, 1.44] p 0.47 | lag-1 autocorrelation 0.57 [-0.14, 0.22] p 0.00 (only the autocorrelation fails)
# Notes: the interval for the MEAN is about a third of the full predictive width (noise dominates here);
# 14/14 held-out days inside the 80% band is a little better than the nominal 80% (14 days is a small sample).
1. Out of 4 000 posterior predictive draws for tomorrow, 520 exceed capacity. What is the estimate of $P(\text{demand} \gt \text{capacity})$ and roughly how accurate is it?
2. A right-skewed predictive distribution of demand has mean 120. The median is most likely…
3. Which source of forecast uncertainty does NOT shrink when you collect more history of the same kind?
4. You need the 90% interval for next week's total orders. You have a (draws × 7 days) array of posterior predictive paths. What do you do?
5. In a trend model with independent daily noise, why does the forecast interval get wider as the horizon grows?
6. In a posterior predictive check, the real series has lag-1 autocorrelation 0.57 while the middle 90% of 500 replicated series gives $-0.14$ to $0.22$. The mean, variance and maximum all look fine. What does this tell you?
Practice problems
A. 1 000 draws for next Friday: 62 exceed capacity 140 and 9 of them exceed 180. Estimate $P(\gt 140)$, $P(\gt 180)$ and $P(140 \lt \text{demand} \le 180)$ with the Monte Carlo errors.
$P(\gt140) = 62/1000 = 0.062 \pm \sqrt{0.062 \times 0.938/1000} = \pm 0.0076$. $P(\gt 180) = 0.009 \pm 0.003$ (a rare event is estimated with a relatively large error: 33% of itself). $P(140 \lt \text{demand} \le 180) = 0.062 - 0.009 = 0.053$. To estimate the 0.9% tail to ±10% of itself you need about $0.009 \times 0.991/(0.0009)^2 \approx 11\,000$ draws.
B. For one future day the noise sd is 15, the parameter sd 9, the regressor effect sd 12, and the between-model sd 8 around a mixture mean of 300. Find the total sd, the 90% interval and the biggest source. What would you do about it?
Variances: $225 + 81 + 144 + 64 = 514$, sd $22.7$. The 90% interval is $300 \pm 1.645 \times 22.7 = [262.7,\ 337.3]$. Shares: noise 44%, regressor 28%, parameters 16%, model 12%. Noise is the floor; the largest source you can attack is the regressor (28%): improve the forecast of the input, or report scenarios for it. More history would only shrink the 16% parameter share.
C. Tomorrow's demand is $N(200, 30^2)$. A lost sale costs 4, an unsold unit costs 1. What stock minimizes expected cost, and how does it compare with stocking the mean?
Critical ratio $\tau = 4/5 = 0.8$, so $s^* = 200 + 0.8416 \times 30 = 225.2$. Expected cost: $E[(y-s)_+] = 3.35$ and $E[(s-y)_+] = 3.35 + 25.2 = 28.6$ (the identity $E[(s-y)_+] = E[(y-s)_+] + s - \mu$), so cost $= 4 \times 3.35 + 28.6 = 42.0$. Stocking the mean: $E[(y-200)_+] = 30\,\varphi(0) = 11.97$ both ways, cost $= 4 \times 11.97 + 11.97 = 59.8$. The quantile stock is 30% cheaper (42.0 against 59.8).
D. Noise sd 6 (independent); level sd 1 and slope sd 0.1 per day at the origin; future slope changes arrive at 2 per 100 days with Laplace scale $b = 0.1$. Compute the predictive sd at 20 and 60 days. Which mechanism dominates at 60 days?
$h = 20$: noise $36$; parameters $1 + (20 \times 0.1)^2 = 5$; future changes $0.02 \times 2(0.1)^2 \times 20^3/3 = 0.02 \times 0.02 \times 2667 = 1.07$; total $42.07$, sd $6.49$. $h = 60$: noise $36$; parameters $1 + 36 = 37$; future changes $0.02 \times 0.02 \times 216000/3 = 28.8$; total $101.8$, sd $10.09$. At 60 days the parameter term (37) is slightly larger than the future-change term (28.8), and both together are bigger than the noise (36); at 120 days the future-change term would dominate because it grows like $h^3$.
E. A PPC of a Negative Binomial forecasting model gives (real vs middle 90% of replicates): mean 52 vs [48, 56]; variance 410 vs [300, 520]; maximum 130 vs [90, 150]; zero days 7 vs [0, 1]; lag-1 autocorrelation 0.10 vs [−0.05, 0.25]. What is wrong, and what would you try?
Only the number of zero days is far outside (7 against at most 1); everything else is inside. The model produces too few zeros: the data have excess zeros, most likely days when nothing could be sold. First find out why (closed days, stock-outs, missing data coded as zero). If known, add a regressor or mask those days; otherwise try a zero-inflated Negative Binomial. Then re-run the check, including the variance (the extra zeros may have inflated the fitted dispersion).
F. Interview: "Your forecast says 90% interval 100 to 140. How do you know it is right, and what does it not cover?"
Model answer: "It is the 5th and 95th percentile of posterior predictive draws, so it combines observation noise from the likelihood with parameter uncertainty from the posterior. I check it in two ways: posterior predictive checks on the history for the mean, variance, tails, zero count, weekly shape and autocorrelation, and a rolling-origin backtest where I count how often the real values fall inside the 80% and 90% bands, by horizon, and look at PIT histograms. What it does not cover is model uncertainty (a trend that changes in a way the model has not seen) and uncertainty in future regressors, unless I simulated them; I report the regressor assumptions, and I treat long-horizon intervals with extra caution because their coverage is the first thing to degrade."
Time-series cross-validation and accuracy metrics
A forecast is only as trustworthy as the test you gave it. This chapter shows how to test a forecasting model honestly (hide the future, move the forecast origin forward again and again, look at each horizon separately, never shuffle) and which numbers to compute afterwards: MAE, MSE, RMSE, MAPE, sMAPE, WAPE and MASE, including exactly where each one misleads. It is the chapter behind the interview question "how did you validate your forecasting model?"
- Evaluate by time: train on the past, test on the future, and name the pieces (forecast origin $T$, horizon $h$, forecast error $e_{T+h\mid T}$)
- Run a rolling-origin evaluation (expanding and sliding windows) and explain why one holdout is only a single noisy draw
- Report horizon-specific error and know why one blended number hides the truth
- Explain, with numbers, why random K-fold leaks, and what to use instead (
TimeSeriesSplit, forward chaining, a gap) - Compute and interpret MAE, MSE, RMSE and the bias, and say which point forecast (mean or median) each one rewards
- Know the limits of MAPE (explodes near zero, undefined at zero, asymmetric) and use sMAPE, WAPE and MASE correctly; read "MASE below 1" precisely
- Choose metrics for a decision, compare models fairly on the same origins, and run all of it in
numpyandstatsmodels
What we need from earlier chapters: forecast origin, horizon and "never shuffle" (Chapter 7.1, especially why random splits leak); seasonality and trend (Chapter 7.2); random walks (Chapter 7.3); the baselines naive, seasonal naive and drift (Chapter 7.5); the mean and the median (Chapter 4.5, 4.14). Notation: $T$ is the forecast origin (the last time whose data we may use). $h = 1, 2, \dots, H$ is the horizon (how many steps ahead). $\hat y_{T+h\mid T}$ ("y hat at T plus h, given T") is the forecast for time $T+h$ made at time $T$. The forecast error is $e_{T+h\mid T} = y_{T+h} - \hat y_{T+h\mid T}$, actual minus forecast (some books use the opposite sign; only the sign of the bias changes). $m$ is the season length ($m = 7$ for daily data with a weekly pattern). $n$ is the number of forecasts being scored.
Holdout by time: hide the future, forecast it, then look core
A teacher cannot test you fairly with the exact questions you practised on; you would just remember the answers. She keeps some questions hidden until the exam. A forecaster has a natural hidden exam: the future. So we pretend we are standing at some day, hide everything after it, make the forecast, and only then uncover the hidden days and compare.
The day we stand on is the forecast origin $T$. Everything up to $T$ is the training data: the only thing the model (and every step that prepares the data) may look at. The hidden block after $T$ is the holdout (also called the test set). The gap between a forecast and what really happened is the forecast error.
Three ways to say it:
- Picture: a wall in time. Everything left of the wall may be used; everything right of it is an exam question.
- Numbers: train on days 1–14, test on days 15–21. The forecast for day 17 is a 3-steps-ahead forecast ($h = 3$) made at $T = 14$.
- Slogan: fit on the past, score on the future, touch nothing in between.
A shop's daily orders for three weeks (Monday to Sunday). We stand at the end of week 2 ($T = 14$) and forecast week 3 with the seasonal naive rule from Chapter 7.5 ("each day will be like the same weekday last week").
| Mon | Tue | Wed | Thu | Fri | Sat | Sun | |
|---|---|---|---|---|---|---|---|
| week 1 (train) | 20 | 22 | 21 | 23 | 30 | 40 | 35 |
| week 2 (train) | 22 | 24 | 23 | 25 | 32 | 42 | 37 |
| week 3 (holdout, actual) | 24 | 25 | 25 | 26 | 33 | 44 | 38 |
- Seasonal naive copies week 2: the forecasts for $h = 1, \dots, 7$ are $22, 24, 23, 25, 32, 42, 37$.
- Errors, actual minus forecast: $24-22 = 2$, $25-24 = 1$, $25-23 = 2$, $26-25 = 1$, $33-32 = 1$, $44-42 = 2$, $38-37 = 1$.
- Mean absolute error: $(2+1+2+1+1+2+1)/7 = 10/7 \approx 1.43$ orders.
- For comparison, the plain naive rule (copy Sunday's 37 for all seven days) has errors $-13, -12, -12, -11, -4, 7, 1$ and MAE $= 60/7 \approx 8.57$ orders.
- Notice that every error has its own horizon: $e_{T+3\mid T} = 2$ is a 3-steps-ahead error, $e_{T+7\mid T} = 1$ a 7-steps-ahead error. Later we look at the horizons one by one.
Holdout (test) evaluation by time. Choose an origin $T$ and a horizon $H$. Use only $y_1, \dots, y_T$ for everything that learns from data. Then compare the forecasts with the hidden values $y_{T+1}, \dots, y_{T+H}$:
$$e_{T+h\mid T} = y_{T+h} - \hat y_{T+h\mid T}, \qquad h = 1, \dots, H, \qquad \text{MAE} = \frac1H \sum_{h=1}^{H} \lvert e_{T+h\mid T}\rvert .$$- "Everything that learns from data" includes the model parameters, but also scalers (means and standard deviations), changepoint detection, feature selection, and tuning of settings such as the Fourier order or the prior scale. All of it uses training data only.
- When you also choose settings by looking at scores, use three blocks in time order: train (fit), validation (compare and choose settings) and test (look once, at the end). If you tune on the test block it is no longer a test.
- If your features use windows that reach forward in time (a 7-day average, a lag), leave a gap between training and test so no training row overlaps the test period.
Why do we need it?
A score on data the model has already seen only measures memory. Forecasting is always a prediction about days that have not happened, so the only honest exam is days the model could not have seen.
Where is it used?
Every forecasting competition (the M-competitions score a hidden final block), Prophet's cross_validation (each "cutoff" is an origin), backtests in finance and demand planning, scikit-learn's TimeSeriesSplit, and the comparison of your model against baselines.
How is it used?
Sort by time. Pick an origin. Fit everything on the data up to it. Forecast the next $H$ steps. Uncover the actual values, record the errors $e_{T+h\mid T}$ with their horizon $h$, and compute a metric. Then do it again at other origins (next section).
"I held out the last 20% of the data, so my evaluation is honest."
It is honest about leakage, but it is still a single draw. The score depends on where the cut falls (before or after a holiday, a promotion, a level jump). One holdout cannot tell luck from skill; repeat the exam at many origins (next section).
"I tuned my settings on the holdout and then reported the holdout score."
Then the holdout was part of training. Keep three blocks in time order: train to fit, validation to choose, a final test block you look at once.
"The sign of the error does not matter."
For MAE and RMSE it does not. For the bias (the mean error) it does. Say which convention you use: here, error = actual − forecast, so a positive mean error means the forecasts were too low.
In your forecasting model, the training window is the only data that may touch: the scaling of $y$ and of the regressors, the choice of changepoints (the grid and the PELT run, 7.9), the Fourier order, the Laplace scale of the slope changes, and the SVI fit itself. If PELT is run once on the whole history and then the model is "backtested", the test period has already influenced the changepoints and the backtest looks better than reality (7.12). In an A/B framework like yours the same rule holds for daily metric series you forecast, for example expected daily conversions used to plan test duration.
Origin $T$: fit on $y_1..y_T$ only. Holdout: $y_{T+1}..y_{T+H}$. Error $e_{T+h\mid T} = y_{T+h} - \hat y_{T+h\mid T}$.
Train → validation (choose) → test (look once), always in time order.
Trap: scalers, changepoint detection and tuning must also see training data only. One holdout is one draw.
Quick check: a team standardizes the whole series first (mean and sd of all 3 years) and then trains on year 1–2 and tests on year 3. What is wrong?
The mean and standard deviation contain year 3, so information about the test period (for example a higher level) leaked into the training inputs. The scaler must be fitted on the training years only and then applied, unchanged, to the test year.
Rolling-origin evaluation: expanding and sliding windows core
One exam question tells you little. A driving test is not one corner but a whole route. To judge a forecaster we want many exams at many moments: after a quiet week, after a holiday, after a jump. So we walk the forecast origin forward: train up to $T_1$ and forecast the next block; add the newly revealed days, move to $T_2$, train again, forecast the next block; and so on. Then we average all the errors.
This also copies real life. A deployed model is rerun every day or week with fresh data, so a backtest should rerun it the same way. It has two flavours. With an expanding window the training set keeps all history from day 1 (it grows). With a sliding window it keeps only the latest $W$ days (it moves, and old days are dropped).
Three ways to say it:
- Picture: a staircase of "blue training block, orange test block" pairs that steps to the right.
- Numbers: 20 origins with a 7-day horizon give 140 scored forecasts instead of 7.
- Slogan: test it the way you will use it, again and again.
The same three weeks of orders. Our model is the seasonal mean: forecast each weekday as the average of that weekday over the training window. We use two origins, $T = 7$ and $T = 14$, with a horizon of 7 days.
- Origin $T = 7$. Training = week 1 only. The seasonal mean is just week 1: $20, 22, 21, 23, 30, 40, 35$. Actual week 2: $22, 24, 23, 25, 32, 42, 37$. Every error is $+2$, so MAE $= 2.00$.
- Origin $T = 14$, expanding window (weeks 1 and 2). Seasonal means: $(20+22)/2 = 21$, $23$, $22$, $24$, $31$, $41$, $36$. Actual week 3: $24, 25, 25, 26, 33, 44, 38$. Errors: $3, 2, 3, 2, 2, 3, 2$; MAE $= 17/7 \approx 2.43$.
- Origin $T = 14$, sliding window of 7 days (week 2 only). Forecasts $22, 24, 23, 25, 32, 42, 37$; errors $2, 1, 2, 1, 1, 2, 1$; MAE $= 10/7 \approx 1.43$.
- Average over the two origins: expanding $(2.00 + 2.43)/2 = 2.21$; sliding $(2.00 + 1.43)/2 = 1.71$.
- Why did sliding win? Orders grow by about 2 per week, so week 1 drags the expanding average too low. If demand had been stable, the expanding window (more data, less noise) would have won.
Rolling-origin evaluation (also called time-series cross-validation, backtesting or walk-forward validation). Choose an initial training size $T_0$, a horizon $H$ and a step $s$. For each origin $T = T_0, T_0 + s, T_0 + 2s, \dots$ (while $T + H \le N$):
- Training window: expanding $= \{1, \dots, T\}$, or sliding $= \{T - W + 1, \dots, T\}$.
- Fit everything on that window only (scalers, features, model).
- Forecast $\hat y_{T+1\mid T}, \dots, \hat y_{T+H\mid T}$ and store the errors $e_{T+h\mid T}$ together with $T$ and $h$.
The number of origins is $K = \lfloor (N - T_0 - H)/s \rfloor + 1$. Metrics are then averaged over origins (and, separately, per horizon).
- Expanding: more data (smaller variance), best when the process is stable; cost grows with each origin; old regimes stay in.
- Sliding: constant cost; forgets old regimes and so adapts to change; less data (noisier); $W$ is a setting you must choose.
- In practice: Prophet's
cross_validation(model, initial, period, horizon)usesinitialas $T_0$,periodas the step $s$ andhorizonas $H$ (expanding). scikit-learn'sTimeSeriesSplit(n_splits, test_size, gap, max_train_size)is expanding by default; settingmax_train_sizemakes it sliding.
Why do we need it?
One holdout is one draw and may be unusually easy or hard. Many origins average out luck, show how stable the model is, and imitate the real "retrain every week" routine.
Where is it used?
Prophet's cross-validation and performance_metrics, statsmodels and sktime backtesting tools, TimeSeriesSplit, demand-planning backtests, and the comparison table of any forecasting paper.
How is it used?
Pick $T_0$ (enough history to learn the seasonality), $H$ (the horizon your business uses), a step $s$ (often one horizon, or smaller). Loop over origins, refit, forecast, store errors. Report the mean and the spread across origins, and the error per horizon.
"Rolling origin means I train once and test on many windows."
At every origin the model is refitted (or at least updated) using only the data available at that origin. Training once on the first window would test a model that gets staler at each step, which is not how it will be used.
"20 origins means 20 independent tests."
The windows overlap and the errors are correlated (neighbouring origins share most of their training data, and errors within one block share the same shocks). Treat the spread across origins as a useful guide, not as 20 independent samples.
"Rolling origins are too expensive for my slow model."
Use fewer, well-spaced origins (even 6 to 12 across different regimes are far better than one), and reuse work that only saw the past, such as starting the fit from the previous origin's parameters.
For your forecasting model, one origin means one complete fit: re-scale, re-choose the changepoints (grid and PELT), and run the SVI loop on that window. So $K$ origins cost $K$ fits. Starting each fit from the previous origin's learned parameters is a common speed-up and is allowed, because those parameters only saw earlier data. Choose the first window long enough to contain at least a couple of full yearly cycles if you model yearly seasonality, otherwise the early origins test a model that could not learn it. The sliding window is a design choice that matters when the trend changes: it makes the model forget the old slopes.
"I did cross-validation with 5 folds."
"I did a rolling-origin backtest: 12 origins, a 14-day horizon, an expanding window, refitting everything inside each window."
Model answer: "For time series I never use shuffled folds. I move the forecast origin forward, refit on the data available at each origin, forecast the next block, and report the error per horizon averaged over origins, with its spread. I use an expanding window if the process is stable and a sliding window if it drifts."
Rolling origin: for $T = T_0, T_0+s, \dots$ → refit on the window → forecast $h = 1..H$ → store $e_{T+h\mid T}$. Origins $K = \lfloor (N - T_0 - H)/s\rfloor + 1$.
Expanding = all history (stable processes). Sliding = last $W$ days (drifting processes).
Trap: refit inside each window; origins overlap, so errors are correlated.
Quick check: $N = 400$ days, first window 100, horizon 14, step 14. How many origins?
Origins are $T = 100, 114, 128, \dots$ while $T + 14 \le 400$, i.e. $T \le 386$. That is $T = 100 + 14j$ for $j = 0, \dots, 20$: $K = \lfloor (400 - 100 - 14)/14 \rfloor + 1 = 20 + 1 = 21$ origins.
Forecast horizon and horizon-specific error core
"What will the weather be tomorrow?" is a much easier question than "what will it be in ten days?". Forecasts get worse the further ahead we look, because more surprises can happen in between. A single score for "the next four weeks" mixes the easy first days with the hard last days, and hides how the model behaves at the horizon you actually care about.
So after a rolling-origin run we do not only average everything. We also group the errors by horizon $h$ (how many steps ahead): all the 1-step-ahead errors together, all the 2-step-ahead errors together, and so on. A kitchen planning tomorrow's staff cares about $h = 1$; a buyer ordering stock eight weeks ahead cares about $h \approx 56$.
Three ways to say it:
- Picture: a curve with the horizon on the bottom axis and the error on the vertical axis, one line per model.
- Numbers: for a random walk the naive error is 4 orders at $h = 1$, 8 at $h = 4$, 16 at $h = 16$: four times the horizon, twice the error.
- Slogan: "How good is the model?" has no answer until you say "how far ahead?".
A series behaves like a random walk (Chapter 7.3): each day it moves by an independent random step with standard deviation $\sigma = 4$ orders. We use the naive forecast "tomorrow = today".
- One step ahead ($h = 1$): the error is exactly one random step, so its standard deviation is $4$.
- Four steps ahead ($h = 4$): $y_{T+4} - y_T$ is the sum of 4 independent steps. Variances add: $4 \times 4^2 = 64$, so the standard deviation is $\sqrt{64} = 8$.
- Nine steps ahead: $9 \times 16 = 144$, standard deviation $12$. Sixteen steps ahead: $16 \times 16 = 256$, standard deviation $16$.
- The RMSE of the naive forecast is therefore $\sigma\sqrt h$: multiplying the horizon by 4 only doubles the error. The average of the RMSE over $h = 1, \dots, 16$ would be a single number near $11.1$ that describes neither the easy nor the hard end.
Suppose a rolling-origin evaluation used origins $T_1, \dots, T_K$. The horizon-specific errors are
$$\text{MAE}_h = \frac1K \sum_{j=1}^{K} \lvert e_{T_j + h\mid T_j}\rvert, \qquad \text{RMSE}_h = \sqrt{\frac1K \sum_{j=1}^{K} e_{T_j+h\mid T_j}^2}, \qquad h = 1, \dots, H.$$- A multi-horizon summary (the mean of $\text{MAE}_h$ over $h$) is fine as one headline number, as long as the per-horizon curve is shown next to it.
- For a random walk with step standard deviation $\sigma$, the naive forecast has $\text{RMSE}_h = \sigma\sqrt h$ (the variance of a sum of independent steps adds).
- The curve is not always increasing. With a weekly pattern, the naive error follows the week: it is smallest at $h = 7, 14, \dots$ (same weekday) and largest when the two compared days are a quiet and a busy weekday.
- Origins must cover the whole cycle. If every origin is the same weekday (step 7 on daily data), horizon $h$ always means the same target weekday, so the curve shows the weekly pattern instead of the true effect of distance. Use a step that is not a multiple of $m$, for example 3.
Why do we need it?
Decisions have a lead time. Staffing, ordering and capacity each need a forecast at a specific horizon, and models can be good at one distance and poor at another. One blended number cannot answer "good for what?".
Where is it used?
The "performance metrics by horizon" plot of Prophet's cross-validation, the M-competition tables (short, medium and long horizon columns), weather and energy forecasting, and any supply-chain forecast tied to a lead time.
How is it used?
During the rolling-origin loop store the error with its horizon. Afterwards group by $h$ (a groupby on the horizon column, or np.abs(A - F).mean(axis=0) on an origins × horizons array), plot the curve for each model, and read off the horizon that matches the decision.
"One MAE for the whole 28-day horizon is enough."
That number averages easy and hard horizons and lets a model that is great at $h = 1$ and poor at $h = 28$ look the same as one that is average everywhere. Report the curve, plus the single horizon your decision depends on.
"Forecast error always grows with the horizon."
Usually, but not always. A random walk grows like $\sqrt h$; a strong, regular seasonal pattern can keep the error nearly flat; naive on weekly data zigzags with period 7.
"I used origins every 7 days, so I have lots of data per horizon."
Every origin was the same weekday, so each horizon always meant the same target weekday. The curve then measures the weekly pattern, not distance. Use a step that is not a multiple of the season length.
In your forecasting model the main sources of error change with distance. At short horizons the error is mostly noise and how well the seasonality and holiday columns explain the next days. At long horizons the trend term dominates: how far the last slope keeps going, and whether a new changepoint will appear that the grid and the Laplace prior did not anticipate (7.8, 7.10). If a regressor is not known in advance, its own forecast error also enters at long horizons (7.12). So always plot your model's error by horizon, and compare it with seasonal naive at the same horizons.
$\text{MAE}_h = \frac1K\sum_j \lvert e_{T_j+h\mid T_j}\rvert$ (likewise RMSE$_h$). Random walk + naive: RMSE$_h = \sigma\sqrt h$.
Report error by horizon; match the horizon to the decision's lead time.
Trap: a step that is a multiple of the season length ties horizon to weekday; the error curve is not always increasing.
Quick check: a random walk has step sd 3. What is the naive RMSE at $h = 25$, and at $h = 100$?
$\text{RMSE}_{25} = 3\sqrt{25} = 15$ and $\text{RMSE}_{100} = 3\sqrt{100} = 30$. Four times the horizon doubles the error, because the variance grows in proportion to $h$ and the standard deviation is its square root.
Why random K-fold cross-validation leaks, and what to use instead core
Chapter 7.1 showed that a random split lets the model peek at days on both sides of a test day. Here we look closely at how much it peeks and at the whole family of splitting schemes. The usual K-fold recipe cuts the data into $K$ piles, trains on $K-1$ of them and tests on the remaining one, $K$ times. For customer records that are independent of each other this is a great recipe. For a series it fails in two ways.
- The future is in the training set. Most training days lie after a typical test day, so the model has seen the level, the trend and any jump that came next.
- Neighbours are almost copies. A test day sits between training days that look very much like it, because a series is smooth and repeats. Predicting it is easy.
The cure is to respect time: every training row comes before every test row (forward chaining, i.e. the rolling-origin scheme of this chapter).
Three ways to say it:
- Picture: shuffled K-fold = orange test days sprinkled among blue training days; forward chaining = a wall, blue on the left, orange on the right.
- Numbers: with 100 days, a shuffled test day has on average about 40 training days after it; a forward-chaining test day has 0.
- Slogan: if the training set contains tomorrow, the score is about interpolation, not forecasting.
Twelve days, 4 folds, and fold 1 of a shuffled split has the test days $\{3, 6, 9\}$. Count how many training days lie in the future of each test day.
- Training days (all the others): $\{1, 2, 4, 5, 7, 8, 10, 11, 12\}$, nine of them.
- Test day 3: the training days after it are $4, 5, 7, 8, 10, 11, 12$, which is 7.
- Test day 6: the training days after it are $7, 8, 10, 11, 12$, which is 5.
- Test day 9: the training days after it are $10, 11, 12$, which is 3.
- So $7 + 5 + 3 = 15$ of the $3 \times 9 = 27$ (test day, training day) pairs, about 56%, look into the future.
- A forward-chaining fold with test days $\{11, 12\}$ trains on days $1$ to $10$: every training day is in the past, and the count is 0.
Splitting schemes for a series of $N$ points (see the picture below):
- Shuffled K-fold (
KFold(shuffle=True)): test days are random. Leaks heavily. - Blocked K-fold (
KFold()without shuffling, which is whatcross_val_score(model, X, y, cv=5)uses for regressors): each test fold is a block of consecutive days. It leaks less, but folds before the last still train on later blocks. - Forward chaining / expanding window (
TimeSeriesSplit, rolling origin): train on $1..T_j$, test on the block after $T_j$. No leak of the future. The first folds have little training data. - Sliding window: the same with a fixed-length training window (
max_train_size). - Gap / embargo: leave $g$ days between the end of training and the start of the test block (
TimeSeriesSplit(gap=g)) whenever features or labels use windows that reach across the boundary (a 7-day rolling mean, a lag).
Whatever the scheme, anything fitted from data (scalers, changepoint detection, feature selection, hyperparameters) must be fitted inside the training part of each fold.
Why do we need it?
You need a validation score that predicts live performance. Shuffled folds give a score that is too good, so a model that cannot forecast can win a model-selection contest and then fail after deployment.
Where is it used?
Model selection and hyperparameter tuning for forecasting (TimeSeriesSplit in scikit-learn pipelines, Prophet's cross-validation, sktime and darts backtests), finance (purged and embargoed cross-validation) and machine-learning forecasting with lag features.
How is it used?
Replace KFold by TimeSeriesSplit(n_splits=5, gap=...), fit pipelines (scaler, selector, model) inside each fold, and compare the fold scores. The gap between a shuffled score and a time-ordered score is an estimate of how much the shuffled score was flattering you.
"My model has no lag features, so a shuffled split cannot leak."
It leaks through the level, the trend and any jump: training days after the test day carry them. In the widget the toy model has no lag features at all.
"cross_val_score(model, X, y, cv=5) is the standard, so it must be fine."
For regressors it uses unshuffled contiguous folds (blocked K-fold). That is not random, but it still trains on later blocks. Pass a time-ordered splitter such as TimeSeriesSplit explicitly.
"Time-ordered folds solve leakage completely."
Only the split. Scalers, changepoint detection, feature selection and tuning must still be fitted inside each training part, and a gap is needed if features overlap the boundary. Also the first folds train on very little data, so their errors are noisier.
In an A/B framework like yours, the units in a test (users) are roughly exchangeable, so shuffling them is not a time problem. But anything you forecast as a daily series (expected conversions per day, traffic for planning test duration) is a time series and needs time-ordered validation. In your forecasting model, tune the Fourier orders, the Laplace scale of the changepoint slopes and the holiday window with time-ordered folds only. A shuffled split would flatter flexible settings, since a flexible curve can fill gaps between neighbouring days.
"I tuned the model with 5-fold cross-validation and picked the best RMSE."
"I tuned it with time-ordered folds and a gap equal to the feature window, so no training row overlaps a test row."
Model answer: "In a shuffled or even a blocked K-fold the model trains on days after the ones it is tested on, so it can interpolate the level and the trend. That makes the validation error too small, and especially so for flexible models, so they win the tuning. I use forward chaining: every training row comes before every test row, with a gap when features overlap, and any scaler or changepoint detection is fitted inside the training part of each fold."
Shuffled K-fold: training contains days after the test day → interpolation, flattering error.
Use time-ordered folds (TimeSeriesSplit, rolling origin), a gap if features overlap, everything fitted inside each training part.
Trap: cv=5 on a regressor is blocked K-fold, still not time-ordered.
Quick check: why is a shuffled split's flattering effect larger when the series has a strong trend?
A trend means the future differs from the past. With shuffled folds the model has training days on both sides of a test day, so a centred average follows the trend with no lag. With time-ordered folds it can only look back and lags behind the trend. The more the future differs from the past, the bigger the difference.
Absolute and squared errors: MAE, MSE, RMSE and the bias core
A forecast misses by $+3$ orders on Monday and by $-3$ on Tuesday. If we average the errors as they are, we get $0$ and the forecast looks perfect, even though it was wrong on both days. The plus and minus signs cancelled. To measure "how wrong", we must get rid of the signs first.
There are two natural ways. Drop the sign (absolute value): a miss of 10 counts exactly twice a miss of 5. Or square it: a miss of 10 counts four times a miss of 5, so big misses start to dominate. Averaging the absolute errors gives the MAE. Averaging the squared errors gives the MSE, and its square root, which is back in the original units, is the RMSE. The plain average of the signed errors is not useless: it is the bias, which says whether forecasts run too high or too low.
Three ways to say it:
- Picture: draw the misses as red sticks. MAE is the average stick length. RMSE is the same average with a spotlight on the long sticks.
- Numbers: misses of $2, 4, 0, 3, 30$ orders give MAE $= 7.8$ but RMSE $\approx 13.6$. One big miss is the whole difference.
- Slogan: MAE is the typical miss; RMSE is the typical miss when big misses hurt extra.
Five days. Actual orders $100, 120, 90, 110, 130$; forecasts $98, 124, 90, 107, 100$. (Errors are actual minus forecast.)
- Errors: $2, -4, 0, 3, 30$.
- Absolute errors: $2, 4, 0, 3, 30$; sum $= 39$; MAE $= 39/5 = 7.8$ orders.
- Squared errors: $4, 16, 0, 9, 900$; sum $= 929$; MSE $= 929/5 = 185.8$ (orders squared); RMSE $= \sqrt{185.8} \approx 13.63$ orders.
- Mean error (bias): $(2 - 4 + 0 + 3 + 30)/5 = 31/5 = 6.2$. Positive, so the forecasts were too low on average (the one big miss did most of that).
- Now suppose the last forecast had been 127 instead of 100 (error $3$). Errors $2, -4, 0, 3, 3$: MAE $= 12/5 = 2.4$, MSE $= 38/5 = 7.6$, RMSE $\approx 2.76$, bias $= 4/5 = 0.8$.
- One big miss multiplied the MAE by $7.8/2.4 = 3.25$ but the RMSE by $13.63/2.76 \approx 4.9$.
For $n$ forecasts $\hat y_i$ of actual values $y_i$, with errors $e_i = y_i - \hat y_i$:
$$\text{MAE} = \frac1n \sum_{i=1}^{n} \lvert e_i\rvert, \qquad \text{MSE} = \frac1n \sum_{i=1}^{n} e_i^2, \qquad \text{RMSE} = \sqrt{\text{MSE}}, \qquad \text{ME (bias)} = \frac1n \sum_{i=1}^{n} e_i .$$- MAE and RMSE are in the same units as the data (orders). MSE is in squared units, which is why we usually report its square root.
- Always $\text{MAE} \le \text{RMSE}$, with equality exactly when all the $\lvert e_i\rvert$ are equal. Also $\text{RMSE} \le \sqrt n\,\text{MAE}$. A large ratio RMSE/MAE means the errors are uneven: a few big misses.
- MAE, MSE and RMSE are all scale-dependent: you cannot compare them across series measured in different units or sizes (that is the job of MASE, below).
- The bias $\text{ME} = \bar y - \bar{\hat y}$ can be $0$ for a poor forecaster (errors cancel), so never report it alone.
Why do we need it?
To compare forecasters we need one number per forecaster that grows when forecasts are worse. Absolute and squared errors are the simplest ones, and they differ in how they treat big misses, which is exactly the question a business has to answer.
Where is it used?
Every forecasting library: Prophet's performance_metrics reports mse, rmse, mae, mape, mdape, smape and coverage; sklearn.metrics has mean_absolute_error and mean_squared_error; competitions and dashboards use them as the first headline numbers.
How is it used?
Compute the errors on the holdout (per origin and per horizon), then np.mean(np.abs(e)) and np.sqrt(np.mean(e**2)). Look at MAE and RMSE together: if RMSE is much larger than MAE, find the days with the big misses.
"An RMSE of 13.6 means a typical forecast is off by 13.6."
The typical (average-sized) miss was 7.8, the MAE. The RMSE is pulled up by the one miss of 30. Read RMSE as "a typical miss, with big misses counted extra".
"Lower RMSE always means a better forecast."
Only if big misses really are more costly for you than their size suggests. If running out of stock by 30 units is just six times worse than by 5, MAE matches the cost better. Choose the metric from the decision, not from habit.
"The mean error is almost zero, so the forecast is accurate."
Positive and negative errors cancel. Zero bias only means no systematic over- or under-forecasting. Report it next to the MAE or RMSE, never instead of them.
These scores are cousins of the likelihoods in your forecasting model (7.13). A Normal likelihood is built from squared errors, so maximizing it is minimizing a sum of squares (Chapter 5.2); a Laplace likelihood is built from absolute errors; a Student-t likelihood gives big errors less weight than a Normal does. Still, the metric you report should come from the decision (what a miss costs), not from the likelihood you trained with. Also note that your model produces a whole predictive distribution (7.14); MAE and RMSE judge only one number taken from it. Judging the whole distribution is the job of Chapter 7.16.
"RMSE is just a more modern MAE."
"They answer different questions. MAE is the typical miss and is robust to a few huge errors. RMSE punishes large misses more, so I use it when big misses are disproportionately costly."
Model answer: "MAE averages the absolute errors, so every unit of error counts the same, and it is minimized by the median. RMSE averages squared errors, so large misses dominate, and it is minimized by the mean. I report both and look at the ratio: if RMSE is much bigger than MAE, a few days drive the error and I look at them."
$e = y - \hat y$; MAE $= \overline{\lvert e\rvert}$; RMSE $= \sqrt{\overline{e^2}}$; bias $= \bar e$. Always MAE ≤ RMSE.
MAE = typical miss; RMSE = typical miss with big misses counted extra. Same units as the data; not comparable across series.
Trap: zero bias ≠ accurate; a big RMSE/MAE ratio flags a few huge errors.
Quick check: errors are $+5, -5, +5, -5$. What are the bias, MAE and RMSE?
Bias $= 0$ (they cancel). MAE $= 5$. RMSE $= \sqrt{25} = 5$. All errors have the same size, so RMSE equals MAE; the zero bias says nothing about accuracy.
Which point forecast does each metric reward: the mean or the median? core
You must write down one number for how many orders arrive in the next hour. How you will be graded decides the best number. Grading by absolute miss rewards the "middle" number: half the time the truth is below it, half the time above (the median). Grading by squared miss rewards the "balance point" number (the mean): a single huge hour pulls the balance point toward it, because squares make big misses count a lot.
When demand is nicely symmetric the mean and the median are the same, so nothing changes. When demand is skewed (many ordinary hours and a few giant ones) they differ, and the choice of metric quietly chooses the forecast. A model fitted by minimizing squared error aims at the mean, a model fitted by minimizing absolute error aims at the median, and each will look better on its own metric.
Three ways to say it:
- Picture: a seesaw (mean) versus the midpoint of a queue of people sorted by height (median).
- Numbers: hours with $2, 3, 3, 4, 38$ orders: the median is 3 and the mean is 10. Forecasting 3 wins on MAE; forecasting 10 wins on MSE.
- Slogan: squared error asks for the mean, absolute error asks for the median.
Orders in five hours: $2, 3, 3, 4, 38$. Compare two constant forecasts, $c = 3$ (the median) and $c = 10$ (the mean).
- For $c = 3$: absolute errors $1, 0, 0, 1, 35$, so MAE $= 37/5 = 7.4$. Squared errors $1, 0, 0, 1, 1225$, so MSE $= 1227/5 = 245.4$.
- For $c = 10$: absolute errors $8, 7, 7, 6, 28$, so MAE $= 56/5 = 11.2$. Squared errors $64, 49, 49, 36, 784$, so MSE $= 982/5 = 196.4$.
- On MAE the median (3) wins: $7.4 \lt 11.2$. On MSE (and RMSE) the mean (10) wins: $196.4 \lt 245.4$.
- Nothing is "wrong" with either forecast. They answer two different questions, and the hour with 38 orders is the reason they differ.
For a sample $y_1, \dots, y_n$ and a constant forecast $c$:
$$\arg\min_c \sum_{i}(y_i - c)^2 = \bar y \ (\text{the mean}), \qquad \arg\min_c \sum_{i}\lvert y_i - c\rvert = \text{the median}.$$- Why the mean: the slope of $\sum (y_i - c)^2$ with respect to $c$ is $-2\sum(y_i - c)$, which is zero exactly when $c = \bar y$.
- Why the median: the slope of $\sum\lvert y_i - c\rvert$ is (number of $y_i$ below $c$) minus (number above $c$). It is zero when the same number of points lie on each side, which is the median.
- The same rule holds for a whole predictive distribution: the forecast that minimizes the expected squared error is its mean, the one that minimizes the expected absolute error is its median. For a skewed distribution (counts, Log-Normal) the two differ, with the mean above the median for a right-skewed one.
- General principle: a point forecast should be matched to the score that will judge it. (A third case, the quantile forecast and the pinball loss, comes in Chapter 7.16.)
Why do we need it?
"Which number do I report?" has no answer until the scoring rule is fixed. Without this link you may train for one thing (the mean), grade on another (MAE) and wonder why the leaderboard disagrees with the loss curve.
Where is it used?
Choosing the loss for gradient boosting or neural forecasters (l2 versus l1 objectives), deciding between the posterior predictive mean and median as a model's point forecast, quantile forecasts for inventory, and robust regression.
How is it used?
Decide the metric from the decision first. Then report the matching summary of the predictive distribution: the mean for RMSE, the median for MAE. With posterior predictive samples (7.14) these are samples.mean(0) and np.median(samples, 0).
"The best forecast is the average."
The average is best when you are scored by squared error. Scored by absolute error, the median is best. The scoring rule decides.
"A median forecast is biased, so it must be worse."
For skewed demand its mean error is not zero, but it has the smaller MAE. "Biased" and "worse" are different statements: bias is judged against the mean, quality against the metric you chose.
"I will compare models on whichever metric makes mine look best."
Fix the metric before you look at results, and fix it from the decision. Otherwise you are choosing the model and the score together.
Your forecasting model returns posterior predictive samples, one distribution for every future day (7.14). With a Normal or Student-t likelihood the distribution is symmetric, so mean and median nearly agree. With a Negative Binomial likelihood for counts, or a log link, the distribution is right-skewed and its mean is above its median. So "the forecast" is not one number: say whether you report the predictive mean (RMSE-optimal) or median (MAE-optimal), and use the metric that matches.
$\arg\min_c \sum (y_i - c)^2 = \bar y$ (RMSE) · $\arg\min_c \sum\lvert y_i - c\rvert = \text{median}$ (MAE).
Skewed demand → mean above median; the metric chooses the forecast.
Trap: reporting the mean but grading with MAE (or the reverse); choose the metric before the model.
Quick check: a right-skewed daily demand has mean 60 and median 45. You forecast 60 every day. Is that MAE-optimal?
No. 60 is the RMSE-optimal constant. The MAE-optimal constant is the median, 45. Forecasting 60 is not "wrong", it just optimizes the other metric.
Percentage errors: MAPE, sMAPE and WAPE, and where they break core
A miss of 5 orders is a disaster for a product that sells 10 a day and nothing for one that sells 1 000. Percentages feel natural: "we were 5% off" means the same for both. So we divide each error by the actual value. That is the MAPE, the mean absolute percentage error.
But dividing by the actual has a price. When an actual is tiny, even a small miss becomes a gigantic percentage (a miss of 5 on a day with 2 orders is 250%), and one such day can dominate the whole average. When an actual is exactly zero, the percentage cannot even be computed. And percentages are lopsided: an over-forecast can be wrong by 300%, but an under-forecast can never be wrong by more than 100%. Two repairs exist: sMAPE divides by the size of actual and forecast together, and WAPE adds up all the misses first and divides by the total of the actuals.
Three ways to say it:
- Picture: a fire alarm that goes off whenever the actual is near zero; one noisy day rings louder than all the other days together.
- Numbers: the same miss of 5 on actuals $100, 50, 10, 2$ is $5\%, 10\%, 50\%, 250\%$. MAPE $= 78.75\%$, yet the MAE is just 5.
- Slogan: MAPE shouts about small days and cannot see zero.
Four days: actual $100, 50, 10, 2$; forecast $105, 55, 15, 7$. Every miss is exactly 5 orders.
- Absolute percentage errors $\lvert e\rvert / y$: $5/100 = 5\%$, $5/50 = 10\%$, $5/10 = 50\%$, $5/2 = 250\%$.
- MAPE $= (5 + 10 + 50 + 250)/4 = 315/4 = 78.75\%$. The last day alone contributes $62.5$ of the $78.75$ points.
- sMAPE (version $200\,\lvert e\rvert/(\lvert y\rvert + \lvert\hat y\rvert)$): $4.88\%$, $9.52\%$, $40\%$, $111.11\%$, average $41.38\%$.
- WAPE $= \sum\lvert e\rvert / \sum y = 20/162 \approx 12.35\%$. The MAE is $20/4 = 5$ orders, and $5$ divided by the mean actual $40.5$ is the same $12.35\%$.
- The lopsidedness. Actual 100, forecast 50: $50/100 = 50\%$. Swap the numbers: actual 50, forecast 100: $50/50 = 100\%$. The same pair of numbers, two different percentages, because the denominator is the actual.
- If the last actual had been $0$ (same forecast $7$, so the miss is $7$): $7/0$ cannot be computed, so the MAPE does not exist. WAPE still does: $(5+5+5+7)/(100+50+10+0) = 22/160 = 13.75\%$.
With errors $e_i = y_i - \hat y_i$:
$$\text{MAPE} = \frac{100}{n}\sum_{i=1}^n \frac{\lvert e_i\rvert}{\lvert y_i\rvert}, \qquad \text{sMAPE} = \frac{100}{n}\sum_{i=1}^n \frac{2\,\lvert e_i\rvert}{\lvert y_i\rvert + \lvert \hat y_i\rvert}, \qquad \text{WAPE} = 100\cdot\frac{\sum_i \lvert e_i\rvert}{\sum_i \lvert y_i\rvert} = 100\cdot\frac{\text{MAE}}{\overline{\lvert y\rvert}}.$$- MAPE: not defined if any $y_i = 0$; explodes when some $y_i$ is small; unbounded above; favours low forecasts (the constant that minimizes it is a weighted median with weights $1/y_i$, which is never above the ordinary median). "100 minus MAPE" is not an accuracy: it can be negative.
- sMAPE: lies between 0% and 200%; several different formulas share this name (some divide by the average $(\lvert y\rvert + \lvert\hat y\rvert)/2$, some multiply by 100 instead of 200), so check your library. It is not symmetric despite the name: for a given actual, an under-forecast costs more than an over-forecast of the same size. Not defined when $y_i = \hat y_i = 0$.
- WAPE (a.k.a. weighted MAPE, or MAE divided by the mean actual): one division at the end, so zero days are fine and tiny days do not explode. Large days weigh more, which usually matches the business. Not defined only if all actuals sum to zero.
Why do we need it?
Managers think in percentages, and one number that works for a product selling 10 and one selling 1 000 is attractive. These metrics make errors relative. You need to know when each relative metric is safe, because the wrong choice can rank models in the wrong order.
Where is it used?
MAPE in retail, energy and finance reports and in Prophet's performance_metrics (mape, smape); sMAPE in the M3 and M4 forecasting competitions; WAPE (and weighted versions) in supply-chain planning for slow and intermittent products.
How is it used?
Check the smallest actual first. If actuals can be zero or tiny (counts, low-volume segments, weekends), use WAPE or MASE instead. If you do report MAPE, show also the MAE, drop or flag near-zero days, and never use it alone to choose between models.
"MAPE of 8% means our forecasts are 92% accurate."
"Accuracy = 100 − MAPE" is not a real quantity: MAPE can exceed 100%, so it can go negative. Say "on average the forecast misses by 8% of the actual".
"sMAPE fixes the problems of MAPE."
It bounds the score at 200% but keeps the problems in another form: it is not symmetric, it is undefined when actual and forecast are both zero, it behaves strangely for forecasts near zero, and its formula varies between libraries.
"My counts contain zeros, so I will add 1 to every actual and then use MAPE."
That silently changes the metric and the ranking of models. Use WAPE or MASE, which are built for data with zeros.
"A lower MAPE always means a better forecast."
MAPE rewards forecasting low (see the widget). A model that is deliberately biased low can win on MAPE while losing on MAE and RMSE.
Count and demand series like those in your forecasting model (modelled with a Negative Binomial likelihood, 7.13) contain small values and often zeros: quiet days, low-volume segments, holidays. There MAPE is undefined or dominated by the tiny days, so prefer WAPE (volume-weighted, zero-safe) and MASE. If you report the predictive median as your point forecast, remember that MAPE would reward an even lower number, which is gaming the metric, not improving the model.
"MAPE is the standard accuracy metric, so I use it for everything."
"I avoid MAPE when actuals can be near zero. I report MAE or RMSE in the original units, WAPE for a relative number, and MASE to compare against the naive forecast."
Model answer: "MAPE has four problems. It is undefined at zero and explodes near zero, so a few small days dominate it. It is asymmetric: it cannot punish an under-forecast by more than 100% but punishes over-forecasts without limit, so it favours low forecasts and the best constant forecast under MAPE is below the median. It depends on the denominators, so it is not comparable between series of different mix. And 'accuracy = 100 − MAPE' is meaningless. I use WAPE, which sums the errors first, and MASE, which scales by the naive error on the training data."
MAPE $= \frac{100}{n}\sum\lvert e\rvert/\lvert y\rvert$ · sMAPE $= \frac{100}{n}\sum 2\lvert e\rvert/(\lvert y\rvert + \lvert\hat y\rvert)$ · WAPE $= 100\sum\lvert e\rvert/\sum\lvert y\rvert$.
MAPE: undefined at 0, explodes near 0, favours low forecasts, unbounded for over-forecasts. WAPE = MAE / mean(y).
Trap: "100 − MAPE = accuracy"; sMAPE is not symmetric; names hide different formulas.
Quick check: actuals $4$ and $400$, forecasts $8$ and $404$ (each misses by 4). Compute MAPE and WAPE.
Percentage errors: $4/4 = 100\%$ and $4/400 = 1\%$, so MAPE $= 50.5\%$. WAPE $= (4 + 4)/(4 + 400) = 8/404 \approx 1.98\%$. The tiny day dominates MAPE; WAPE sees that nearly all the volume was forecast well.
MASE: the scale-free error that says "better than naive?" core
We want a yardstick with three properties. (1) It must not divide by the actual values, so zeros and tiny days are harmless. (2) It must be scale-free, so a shop with 30 orders a day and a shop with 3 000 orders a day can be compared and averaged. (3) It should say directly whether we beat a simple forecast.
The idea: measure the miss of the cheap naive forecast on the training data, in the same units as the data, and use it as a ruler. Your model's average miss on the test data, divided by that ruler, tells you how many "naive-sized misses" your typical miss is. That is the MASE, the mean absolute scaled error.
Three ways to say it:
- Picture: a ruler built from the naive forecast's misses on the history; your model's test miss is measured against it.
- Numbers: test MAE 1.43 orders, naive ruler 2.0 orders: MASE $= 0.71$, i.e. about 29% smaller than the ruler.
- Slogan: MASE 1 means "as good as the naive forecast on the history"; below 1 is better.
Back to the three weeks of orders. Train on weeks 1 and 2, test on week 3 with the seasonal naive forecast (MAE $= 10/7 \approx 1.43$, Section 1).
- Seasonal ruler ($m = 7$). In the training data compare each day with the same weekday a week earlier: $\lvert 22-20\rvert, \lvert 24-22\rvert, \dots, \lvert 37-35\rvert$ are all $2$. Their mean is $d_7 = 14/7 = 2.0$.
- MASE $= 1.43/2.0 \approx 0.71$. The test misses are about 29% smaller than the in-sample seasonal naive miss.
- Day-to-day ruler ($m = 1$). The 13 differences $\lvert y_t - y_{t-1}\rvert$ are $2,1,2,7,10,5,13,2,1,2,7,10,5$, sum $67$, so $d_1 = 67/13 \approx 5.15$. Now the seasonal naive MASE is $1.43/5.15 \approx 0.28$, and the plain naive forecast (MAE $= 60/7 \approx 8.57$) has MASE $8.57/5.15 \approx 1.66$.
- The same forecast gets different MASE values for different $m$: always say which $m$ you used (here $m = 7$ matches the seasonal naive baseline we want to beat).
- Scale-free. Multiply every order count by 100. Every error and the ruler are multiplied by 100, so MASE stays $0.71$. The MAE would have jumped from 1.43 to 143.
- Why does plain naive score 1.66 and not 1? The ruler is a one-day-ahead naive error, but we scored naive up to seven days ahead, where errors are larger. More on this in the box below.
With training data $y_1, \dots, y_T$, test forecasts $\hat y_{T+h\mid T}$ for $h = 1, \dots, H$, and a season length $m$ ($m = 1$ for non-seasonal data):
$$\text{MASE} = \frac{\text{MAE}_{\text{test}}}{d_m}, \qquad \text{MAE}_{\text{test}} = \frac1H\sum_{h=1}^{H}\lvert y_{T+h} - \hat y_{T+h\mid T}\rvert, \qquad d_m = \frac{1}{T-m}\sum_{t=m+1}^{T}\lvert y_t - y_{t-m}\rvert .$$- $d_m$ is the in-sample MAE of the (seasonal) naive forecast. It uses the training data only, never the test data.
- MASE $\lt 1$: the forecast errors are smaller than the average in-sample naive error. MASE $\gt 1$: larger. (Hyndman and Koehler proposed it in 2006.)
- Properties: scale-free; defined even if actuals are zero; symmetric between under- and over-forecasts; can be averaged across series. Not defined if the training series is constant ($d_m = 0$).
- With several origins, compute each origin's MASE with its own training window and average them (or average the scaled errors $\lvert e\rvert / d_m$ over all origins and horizons).
- Horizon caveat. $d_m$ is a one-step ($m$-step) error but test errors reach $H$ steps ahead. So even the naive forecast usually scores MASE $\gt 1$ on a long horizon. To decide "do I beat naive?", compare with the naive forecast's own MASE (or MAE) on the same test windows: relative MAE $= \text{MAE}_{\text{model}}/\text{MAE}_{\text{naive}}$, i.e. the skill score of Chapter 7.5.
Why do we need it?
MAE and RMSE cannot be compared between series of different size, MAPE breaks at zero, and none of them says whether the number is good. MASE removes the units, tolerates zeros and anchors the number to a baseline everyone understands.
Where is it used?
The M4 forecasting competition (together with sMAPE) and the sktime, darts and R forecast accuracy tools; retail and energy teams that forecast thousands of series of very different size and need one average.
How is it used?
Compute $d_m$ from the training window, e.g. np.mean(np.abs(train[m:] - train[:-m])); divide the test MAE by it; report the average over origins (and per horizon) next to the MASE of your strongest baseline.
"MASE below 1 means my model beats the naive forecast on the test set."
It means the test errors are smaller than the in-sample, one-step naive error. For long horizons even the naive forecast itself scores above 1. To claim "beats naive" compare with the naive forecast's score on the same origins and horizons (skill $= 1 - \text{MAE}_{\text{model}}/\text{MAE}_{\text{naive}}$).
"I computed the ruler from the test data to be safe."
The ruler is the naive error on the training data. Using test data would leak the future into the metric, and it changes with every holdout.
"MASE with $m = 1$ is the default, so I do not need to say which $m$."
For seasonal data $m = 1$ gives a much larger ruler than $m = 7$ (day-to-day changes include the weekly swing), which makes every model look better. State $m$, and match it to the baseline you care about.
MASE suits your forecasting model for two reasons. Count-like demand with a Negative Binomial likelihood has zeros, where MAPE fails but MASE does not. And an A/B framework or planning system with many segments (hierarchy of groups) forecasts series of very different size: MASE averages them fairly, while an average RMSE would be dominated by the biggest segment. Use $m = 7$ for daily data with a weekly pattern, report it per horizon, and put the MASE of seasonal naive next to your model's MASE.
"MASE is just a percentage error."
"MASE is the test MAE divided by the in-sample MAE of a naive forecast. It is a ratio of errors, not a percentage of the actuals."
Model answer: "I use MASE to compare across series and to anchor to a baseline. The denominator is the in-sample MAE of the naive (or seasonal naive) forecast on the training data, so it never divides by the actuals and works with zeros. Below 1 means my errors are smaller than the average naive error on the history. Because that ruler is a one-step error, I also check the model against the naive forecast's score at the same horizons before saying I beat the baseline."
$\text{MASE} = \text{MAE}_{\text{test}} / d_m$, $d_m = \frac{1}{T-m}\sum_{t=m+1}^{T}\lvert y_t - y_{t-m}\rvert$ (training data only).
Scale-free, zero-safe, below 1 = smaller than the in-sample naive error. State $m$.
Trap: the ruler is one-step, so long horizons give MASE above 1 even for naive; compare with naive on the same windows.
Quick check: the training data is $10, 12, 11, 14, 13$ and $m = 1$. The test MAE of your model is 1.5. What is the MASE?
Differences $\lvert 12-10\rvert, \lvert 11-12\rvert, \lvert 14-11\rvert, \lvert 13-14\rvert = 2, 1, 3, 1$, mean $d_1 = 7/4 = 1.75$. MASE $= 1.5/1.75 \approx 0.86$: a bit better than the average in-sample naive miss.
Choosing metrics and running a fair bake-off core
Imagine a baking contest. It is fair only if every cake is baked in the same oven, judged by the same people and compared with the same plain sponge cake. A forecasting bake-off works the same way: every model gets the same origins, the same training windows, the same horizons and the same metrics, and the baselines are in the contest too.
Then we read the whole table, not only the first column. Different metrics answer different questions, and they can disagree about the winner. The question to ask is never "which metric is best?" but "which metric matches the decision?".
Three ways to say it:
- Picture: a table with models as rows and metrics as columns, the best cell of each column starred.
- Numbers: Holt–Winters beats seasonal naive on MAE ($3.62$ against $3.99$), but wins at 16 of 20 origins with a gain whose spread is $1.79$: promising, not proven.
- Slogan: same oven, same judges, plain sponge in the contest, and read the whole table.
The Code-it block at the end of this chapter simulates 210 days of orders (weekly pattern, slow trend, and a level jump on day 140) and runs a rolling-origin evaluation: expanding window, first origin at day 70, horizon 7, 20 origins. Its printed result:
| model | MAE | RMSE | MAPE % | sMAPE % | WAPE % | MASE ($m = 7$) |
|---|---|---|---|---|---|---|
| naive | 11.52 | 13.36 | 16.9 | 15.3 | 15.1 | 3.39 |
| seasonal naive | 3.99 | 6.32 | 5.3 | 5.6 | 5.2 | 1.20 |
| drift | 11.93 | 13.83 | 17.6 | 15.8 | 15.6 | 3.51 |
| Holt–Winters | 3.62 | 6.09 | 4.7 | 4.8 | 4.7 | 1.08 |
- Strongest baseline first. On this weekly series that is seasonal naive, not naive. Skill of Holt–Winters $= 1 - 3.62/3.99 \approx 0.09$ (9% smaller MAE). Against naive it would look like $1 - 3.62/11.52 \approx 0.69$, which flatters it.
- Spiky errors? RMSE/MAE for Holt–Winters is $6.09/3.62 \approx 1.7$: a few big misses exist. They come from the two origins just after the level jump on day 140: Holt–Winters has an MAE of 21.8 at the origin $T = 140$ and 9.1 at $T = 147$, against 1.4 to 4.1 at most other origins.
- MASE against 1. $1.08$ is above 1, which is not alarming: the test horizon is up to 7 days and the data contain a jump.
- Do the metrics agree? Yes here: the order is Holt–Winters, seasonal naive, naive, drift in every column. (The widget below shows series where they disagree.)
- Is the gap real? The per-origin MAE difference (seasonal naive minus Holt–Winters) has mean $0.37$ and standard deviation $1.79$ over 20 origins, so its standard error is about $1.79/\sqrt{20} \approx 0.40$: the mean gain is about one standard error. Holt–Winters wins at 16 of 20 origins, but the two origins right after the jump (where it reacts badly, 9.1 against 2.6 at $T = 147$) pull its average back.
A fair evaluation protocol for comparing forecasting models:
- Fix the metric from the decision before looking at results (see the guide below), and write down the primary one.
- Same rolling origins, windows and horizons for every model, with everything fitted inside each window.
- Include baselines (naive, seasonal naive, drift) and report a skill score or MASE against the strongest.
- Report more than the mean: the metric per horizon, its spread across origins (standard deviation, quantiles or a paired difference between two models), and the bias.
- Compare two models by the paired per-origin loss difference. A formal test of equal forecast accuracy for such differences is the Diebold–Mariano test (named here so you can look it up; we do not need its details).
Aggregating over several series: do not average RMSE or MAE of series with different sizes (the biggest series dominates). Average a scaled score such as MASE, or a volume-weighted one such as WAPE (total error over total volume).
Why do we need it?
Model choice is where most forecasting mistakes happen: a favourable metric, a lucky window, a missing baseline, or an unfair training setup can make a worse model look better. A written protocol takes the luck and the temptation out.
Where is it used?
Forecasting competitions (M4, M5) and papers, internal model-review meetings, sktime benchmarking, MLOps "champion versus challenger" tests before a new forecasting model replaces an old one in production.
How is it used?
Loop over origins and models in one script that shares the data splits. Store forecasts with their origin and horizon. Build a table of metrics and a per-horizon plot. Pick the winner by the pre-chosen metric, and inspect the paired differences and the worst origins before deciding.
"I will choose the metric after I see which model wins."
That chooses the model and the yardstick together, which guarantees a flattering result. Write the primary metric down first.
"Model A has an MAE 0.2 lower, so it is better."
Look at how that difference varies across origins. If it is smaller than its spread, or comes from one or two origins, it may be luck.
"I averaged the RMSE of 500 products to get one score."
The biggest products dominate that average. Use a scaled score (MASE) or a volume-weighted one (WAPE) when series differ in size.
"My model beat the naive forecast, so it is good."
Beat the strongest simple baseline for your data (for weekly data, seasonal naive), at the horizons you care about, over many origins.
For your forecasting model, a good first bake-off compares: seasonal naive, a Holt–Winters or ETS model (7.5), and the Prophet-style model, on the same rolling origins and horizons, reporting MAE or RMSE, WAPE, MASE per horizon, and the paired differences. The Bayesian model also gives a full predictive distribution, so the fair comparison continues with coverage, CRPS and the log score (Chapter 7.16), which can show a gain that RMSE alone hides. For the A/B framework the same discipline applies to a forecast of baseline conversion rates used in planning: compare against "last week's rate" first.
"We compared three models on the test set and picked the one with the lowest RMSE."
"We compared them over 20 rolling origins at horizons 1 to 14, against seasonal naive, with MAE, WAPE and MASE chosen in advance, and checked that the improvement held across origins."
Model answer: "I fix the evaluation before running it: same rolling origins and horizons for every model, baselines included, everything refit inside each window. I report error by horizon, a scale-free metric such as MASE, and the spread of the per-origin differences between my model and the strongest baseline. Only then do I say that a model is better."
Fair bake-off: same origins, windows, horizons, metrics; baselines included; metric chosen from the decision beforehand.
Report: error by horizon · MASE or skill vs the strongest baseline · spread of per-origin differences · bias.
Trap: averaging RMSE/MAE across series of different size; picking the metric after the winner.
Quick check: a model's MAE is 10% lower than seasonal naive's, but the per-origin difference has mean 0.3 and standard deviation 2.0 over 12 origins. Is it convincing?
Not yet. The standard error of the mean difference is about $2.0/\sqrt{12} \approx 0.58$, so the average gain (0.3) is only about half a standard error. That is well within what luck of the origins can produce. Use more origins, look at which origins drive the gain, and check the horizons that matter.
Recap, cheat sheet and practice
- Holdout by time: fit everything (parameters, scalers, changepoints, tuning) on $y_1..y_T$; score on $y_{T+1}..y_{T+H}$; error $e_{T+h\mid T} = y_{T+h} - \hat y_{T+h\mid T}$. One holdout is one draw.
- Rolling origin: repeat at origins $T_0, T_0+s, \dots$; refit inside each window; expanding window (all history, stable processes) or sliding window (last $W$ days, drifting processes); $K = \lfloor (N-T_0-H)/s\rfloor + 1$ origins.
- Horizon-specific error: $\text{MAE}_h$, $\text{RMSE}_h$ per horizon; random walk + naive gives $\sigma\sqrt h$; seasonal series can zigzag; use a step that is not a multiple of the season length.
- Random K-fold leaks: training days lie after the test day (interpolation). Use forward chaining (
TimeSeriesSplit), a gap when features overlap, everything fitted inside each training part. - MAE / RMSE / bias: typical miss / typical miss with big misses counted extra / systematic direction. MAE ≤ RMSE. MAE is minimized by the median, RMSE by the mean.
- MAPE is undefined at 0, explodes near 0, is lopsided and favours low forecasts; sMAPE is bounded but not symmetric and has several formulas; WAPE $= \sum\lvert e\rvert/\sum\lvert y\rvert$ is zero-safe.
- MASE $= \text{MAE}_{\text{test}}/d_m$ with $d_m$ the in-sample naive MAE: scale-free, zero-safe; below 1 beats the in-sample naive error (compare with naive on the same windows for horizons above 1).
- Fair bake-off: same origins, windows, horizons; baselines included; metric chosen from the decision first; report by horizon, spread across origins, paired differences and bias.
Cheat sheet
| Quantity | Formula | Remember |
|---|---|---|
| Forecast error | $e_{T+h\mid T} = y_{T+h} - \hat y_{T+h\mid T}$ | actual − forecast; positive bias = forecasts too low |
| Origins | $K = \lfloor (N - T_0 - H)/s\rfloor + 1$ | refit at every origin; expanding or sliding |
| Horizon error | $\text{MAE}_h = \frac1K\sum_j\lvert e_{T_j+h\mid T_j}\rvert$ | random walk, naive: $\text{RMSE}_h = \sigma\sqrt h$ |
| MAE | $\frac1n\sum\lvert e_i\rvert$ | typical miss; best constant = median |
| RMSE | $\sqrt{\frac1n\sum e_i^2}$ | big misses count extra; best constant = mean; ≥ MAE |
| Bias (ME) | $\frac1n\sum e_i$ | can be 0 for a bad forecaster |
| MAPE | $\frac{100}{n}\sum\lvert e_i\rvert/\lvert y_i\rvert$ | undefined at 0; explodes near 0; favours low forecasts |
| sMAPE | $\frac{100}{n}\sum 2\lvert e_i\rvert/(\lvert y_i\rvert + \lvert\hat y_i\rvert)$ | 0–200%; not symmetric; check the library's formula |
| WAPE | $100\sum\lvert e_i\rvert/\sum\lvert y_i\rvert$ | = MAE / mean(y); big days weigh more; zero-safe |
| MASE | $\text{MAE}_{\text{test}}\big/\frac{1}{T-m}\sum_{t=m+1}^{T}\lvert y_t - y_{t-m}\rvert$ | scale-free; ruler from training data; state $m$ |
| Skill | $1 - E_{\text{model}}/E_{\text{baseline}}$ | use the strongest baseline, same windows |
import warnings; warnings.filterwarnings("ignore")
import numpy as np
from statsmodels.tsa.holtwinters import ExponentialSmoothing
from sklearn.model_selection import TimeSeriesSplit, KFold
# 1) Simulated daily orders: weekly pattern + slow trend + noise, and a level jump on day 140
rng = np.random.default_rng(7)
n = 210
t = np.arange(n)
weekly = np.array([0, 3, 2, 4, 12, 24, 18])
y = 40 + 0.12 * t + weekly[t % 7] + rng.normal(0, 3, n)
y[140:] += 22
# 2) The metrics, written out (a = actual, f = forecast)
mae = lambda a, f: np.mean(np.abs(a - f))
rmse = lambda a, f: np.sqrt(np.mean((a - f) ** 2))
mape = lambda a, f: 100 * np.mean(np.abs((a - f) / a)) # breaks if an actual is 0
smape = lambda a, f: 100 * np.mean(2 * np.abs(a - f) / (np.abs(a) + np.abs(f)))
wape = lambda a, f: 100 * np.sum(np.abs(a - f)) / np.sum(np.abs(a))
def mase(a, f, train, m=7): # ruler = in-sample MAE of the seasonal naive forecast (training data only)
return np.mean(np.abs(a - f)) / np.mean(np.abs(train[m:] - train[:-m]))
# The worked example of this chapter: 3 weeks of orders, forecast week 3 from weeks 1-2
w = np.array([20,22,21,23,30,40,35, 22,24,23,25,32,42,37, 24,25,25,26,33,44,38.])
tr, te = w[:14], w[14:]
print("naive MAE", round(mae(te, np.repeat(tr[-1], 7)), 2), "| seasonal naive MAE", round(mae(te, tr[-7:]), 2),
"| MASE (m=7)", round(mase(te, tr[-7:], tr), 3))
# naive MAE 8.57 | seasonal naive MAE 1.43 | MASE (m=7) 0.714
# 3) Four forecasters. Each one sees ONLY the training window it is given.
def f_naive(tr, h): return np.repeat(tr[-1], h)
def f_snaive(tr, h): return np.tile(tr[-7:], h // 7 + 1)[:h]
def f_drift(tr, h): return tr[-1] + np.arange(1, h + 1) * (tr[-1] - tr[0]) / (len(tr) - 1)
def f_hw(tr, h):
fit = ExponentialSmoothing(tr, trend="add", seasonal="add", seasonal_periods=7,
initialization_method="estimated").fit()
return fit.forecast(h)
models = {"naive": f_naive, "seasonal naive": f_snaive, "drift": f_drift, "Holt-Winters": f_hw}
# 4) Rolling-origin evaluation: first origin at day 70, horizon 7, the origin moves 7 days each time
def rolling(window, init=70, h=7, step=7):
out = {k: {"a": [], "f": [], "tr": []} for k in models}
for T in range(init, n - h + 1, step):
train = y[:T] if window == "expanding" else y[T - init:T] # sliding = the last 70 days only
for k, fn in models.items():
out[k]["a"].append(y[T:T + h]); out[k]["f"].append(fn(train, h)); out[k]["tr"].append(train)
return out
for window in ["expanding", "sliding"]:
res = rolling(window)
print(f"\n{window} window, {len(res['naive']['a'])} origins, horizon 7")
print(f"{'model':15s} {'MAE':>6s} {'RMSE':>6s} {'MAPE%':>6s} {'sMAPE%':>7s} {'WAPE%':>6s} {'MASE':>5s}")
for k, r in res.items():
a, f = np.concatenate(r["a"]), np.concatenate(r["f"])
ms = np.mean([mase(aa, ff, tt) for aa, ff, tt in zip(r["a"], r["f"], r["tr"])]) # average of per-origin MASE
print(f"{k:15s} {mae(a,f):6.2f} {rmse(a,f):6.2f} {mape(a,f):6.1f} {smape(a,f):7.1f} {wape(a,f):6.1f} {ms:5.2f}")
# expanding window, 20 origins, horizon 7
# naive 11.52 13.36 16.9 15.3 15.1 3.39
# seasonal naive 3.99 6.32 5.3 5.6 5.2 1.20
# drift 11.93 13.83 17.6 15.8 15.6 3.51
# Holt-Winters 3.62 6.09 4.7 4.8 4.7 1.08
# sliding window, 20 origins, horizon 7
# naive 11.52 13.36 16.9 15.3 15.1 3.11 (naive ignores old data, so only MASE's ruler changes)
# seasonal naive 3.99 6.32 5.3 5.6 5.2 1.13
# drift 12.38 14.28 18.1 16.2 16.2 3.31
# Holt-Winters 3.87 6.28 5.0 5.2 5.1 1.06
# 4b) Do the two best models really differ? Paired per-origin MAE difference (expanding window)
res = rolling("expanding")
d = [mae(a, fs) - mae(a, fh) for a, fs, fh in zip(res["seasonal naive"]["a"], res["seasonal naive"]["f"], res["Holt-Winters"]["f"])]
print("\nper-origin MAE(seasonal naive) - MAE(Holt-Winters): mean", round(np.mean(d), 2), "sd", round(np.std(d, ddof=1), 2),
"| Holt-Winters better at", sum(x > 0 for x in d), "of", len(d), "origins")
# per-origin MAE(seasonal naive) - MAE(Holt-Winters): mean 0.37 sd 1.79 | Holt-Winters better at 16 of 20 origins
# 5) Error by horizon (expanding window). The step is 3, not 7: with step 7 every origin falls on the same
# weekday, so "horizon h" would always mean the same weekday and the curve would show the weekly pattern.
res = rolling("expanding", h=14, step=3)
print("\nMAE by horizon h (expanding window,", len(res["naive"]["a"]), "origins, step 3)")
print(f"{'h =':15s}", " ".join(f"{h:5d}" for h in [1, 2, 4, 7, 10, 14]))
for k in ["naive", "seasonal naive", "Holt-Winters"]:
A = np.array(res[k]["a"]); F = np.array(res[k]["f"])
print(f"{k:15s}", " ".join(f"{v:5.1f}" for v in np.abs(A - F).mean(axis=0)[[0, 1, 3, 6, 9, 13]]))
# h = 1 2 4 7 10 14
# naive 8.1 11.5 13.6 3.4 14.2 6.2 (zigzags with the week: dips at h = 7 and 14, the same weekday)
# seasonal naive 3.5 4.8 3.4 3.4 5.2 6.2
# Holt-Winters 2.4 2.9 3.3 3.8 4.5 5.1 (best at h = 1, 2, 4, 10, 14; seasonal naive wins at h = 7)
# 6) Why shuffled K-fold leaks: (a) the future is inside the training set, (b) the score is too good
def future_share(splits):
return np.mean([np.sum(train > i) for train, test in splits for i in test])
def local_weekday_mae(splits): # toy model: mean of the same weekday +-1, 2, 3 weeks, from TRAINING days only
errs = []
for train, test in splits:
ok = np.zeros(n, bool); ok[train] = True
for i in test:
nb = [i + k for k in (-21, -14, -7, 7, 14, 21) if 0 <= i + k < n and ok[i + k]]
errs.append(abs(y[i] - (y[nb].mean() if nb else y[train].mean())))
return np.mean(errs)
shuf = list(KFold(5, shuffle=True, random_state=0).split(y))
tsp = list(TimeSeriesSplit(5).split(y))
print("\nmean number of training days AFTER the test day: shuffled K-fold =", round(future_share(shuf), 1),
"| TimeSeriesSplit =", round(future_share(tsp), 1))
print("MAE of the same toy model: shuffled K-fold =", round(local_weekday_mae(shuf), 2),
"| TimeSeriesSplit =", round(local_weekday_mae(tsp), 2))
# mean number of training days AFTER the test day: shuffled K-fold = 84.0 | TimeSeriesSplit = 0.0
# MAE of the same toy model: shuffled K-fold = 3.31 | TimeSeriesSplit = 12.23 (most of the gap is the day-140 jump)
Output checked with numpy 2.5, statsmodels 0.15 and scikit-learn 1.9. The Holt–Winters fits use initialization_method="estimated"; exact digits can differ slightly with other library versions. Note that the sliding-window MASE of the naive forecast changes only because its ruler is computed on a shorter window.
1. Which procedure gives an honest estimate of how a forecasting model will perform after deployment?
2. A model's forecast errors on four days are $+3, -3, +3, -3$. What are the bias, the MAE and the RMSE?
3. Daily demand is strongly right-skewed (mean 60, median 45) and you are judged by MAE. Which constant forecast is best?
4. The actual value is 100. Which statement about the percentage error behind MAPE is true?
5. In the training data the average of $\lvert y_t - y_{t-7}\rvert$ is 2.0. Your model's test MAE is 3.0. What is the MASE ($m = 7$) and what does it say?
6. Daily data. You run a rolling-origin evaluation with an origin every 7 days and plot the error against the horizon $h$. What is the danger?
Practice problems
A. A series is $y_1, \dots, y_8 = 10, 12, 14, 13, 15, 17, 16, 18$. Use the naive forecast, horizon $H = 2$, first origin $T_0 = 4$, step 2, expanding window. How many origins are there? Compute all errors, $\text{MAE}_1$, $\text{MAE}_2$, the overall MAE, the RMSE and the bias.
- Origins: $K = \lfloor (8 - 4 - 2)/2\rfloor + 1 = 2$, namely $T = 4$ and $T = 6$.
- Origin $T = 4$: naive forecast $= y_4 = 13$ for $y_5, y_6$. Errors: $15 - 13 = 2$ and $17 - 13 = 4$.
- Origin $T = 6$: forecast $= y_6 = 17$ for $y_7, y_8$. Errors: $16 - 17 = -1$ and $18 - 17 = 1$.
- By horizon: $\text{MAE}_1 = (\lvert 2\rvert + \lvert -1\rvert)/2 = 1.5$; $\text{MAE}_2 = (4 + 1)/2 = 2.5$. The two-steps-ahead error is larger.
- Overall: MAE $= (2 + 4 + 1 + 1)/4 = 2.0$. MSE $= (4 + 16 + 1 + 1)/4 = 5.5$, RMSE $= \sqrt{5.5} \approx 2.35$. Bias $= (2 + 4 - 1 + 1)/4 = 1.5$ (positive: forecasts mostly too low on this rising series).
B. Actuals $80, 40, 20, 10$; forecasts $90, 36, 25, 5$. Compute MAE, RMSE, bias, MAPE, sMAPE and WAPE, and say which day dominates MAPE.
- Errors $y - \hat y$: $-10, 4, -5, 5$; absolute: $10, 4, 5, 5$.
- MAE $= 24/4 = 6.0$. MSE $= (100 + 16 + 25 + 25)/4 = 41.5$, RMSE $\approx 6.44$. Bias $= (-10 + 4 - 5 + 5)/4 = -1.5$ (forecasts a little too high).
- Percentage errors: $10/80 = 12.5\%$, $4/40 = 10\%$, $5/20 = 25\%$, $5/10 = 50\%$. MAPE $= 97.5/4 = 24.375\%$. The smallest day (10 orders) alone gives $50/97.5 \approx 51\%$ of the sum.
- sMAPE terms $200\cdot\lvert e\rvert/(y + \hat y)$: $11.76, 10.53, 22.22, 66.67$; mean $\approx 27.79\%$.
- WAPE $= 24/150 = 16\%$. It is much lower than the MAPE because the big day (80) was forecast well and carries the most volume.
C. Training data $100, 104, 103, 108, 110$. A model forecasts $112, 114$ for actuals $113, 118$. Compute the MASE with $m = 1$. The naive forecast would say $110, 110$: compute its MASE too and explain why it is above 1.
- Ruler: $\lvert 104-100\rvert, \lvert 103-104\rvert, \lvert 108-103\rvert, \lvert 110-108\rvert = 4, 1, 5, 2$; $d_1 = 12/4 = 3$.
- Model: errors $1, 4$, MAE $= 2.5$, MASE $= 2.5/3 \approx 0.83$.
- Naive: errors $113 - 110 = 3$ and $118 - 110 = 8$, MAE $= 5.5$, MASE $= 5.5/3 \approx 1.83$.
- Naive scores above 1 even though it is "the naive forecast", because the ruler is a one-day-ahead error and the test errors are one and two days ahead on a rising series. So "MASE below 1" cannot be read as "beats naive at the same horizon". Here the honest statement is that the model's MAE (2.5) is less than half the naive forecast's (5.5): skill $= 1 - 2.5/5.5 \approx 0.55$.
D. Hourly orders are $1, 2, 2, 3, 12$. Find the best constant forecast for MAE and for MSE, and the two scores of each.
- Median $= 2$; mean $= 20/5 = 4$.
- At $c = 2$: absolute errors $1, 0, 0, 1, 10$, MAE $= 12/5 = 2.4$; squared errors $1, 0, 0, 1, 100$, MSE $= 102/5 = 20.4$.
- At $c = 4$: absolute errors $3, 2, 2, 1, 8$, MAE $= 16/5 = 3.2$; squared errors $9, 4, 4, 1, 64$, MSE $= 82/5 = 16.4$.
- So the median (2) is best for MAE ($2.4 \lt 3.2$) and the mean (4) is best for MSE ($16.4 \lt 20.4$). The single hour with 12 orders is what separates them.
E. (Interview) "Why can't I just use 5-fold cross-validation for my forecasting model?"
"In a shuffled K-fold the training set contains days after each test day, so the model sees the level, trend and any jump that came next. That turns forecasting into interpolation, and the error looks too small, most of all for flexible models and when the series changes over time. Even unshuffled blocked folds train on later blocks. I use forward chaining: every training row comes before every test row, a gap if my features use windows, and I refit everything (scalers, changepoint detection, hyperparameters) inside each training part. I repeat it at many origins and report the error by horizon."
F. (Design) You must compare three models on 2 000 products, with sizes from 1 to 10 000 units a day, many days with zero sales, and a 14-day planning horizon. Describe the evaluation.
- Split: rolling origins, e.g. 8 to 12 origins spaced across a year (including a holiday season), 14-day horizon, step not a multiple of 7, expanding or sliding window to mirror how the model will be retrained; everything refit inside each window.
- Baselines: naive, seasonal naive ($m = 7$) and drift in the same contest; report skill against the strongest.
- Metrics: not MAPE or sMAPE (zeros and tiny products). Use MASE ($m = 7$) per product, averaged over products, so a 1-unit and a 10 000-unit product count equally; and WAPE over the whole volume if big products matter more. Show bias too, since over- and under-stocking have different costs.
- Aggregation: never average raw MAE or RMSE across products (the biggest ones dominate); show metrics by horizon, by product size band, and the spread of the paired per-origin differences between models.
- Decision: choose the primary metric beforehand, and then look at the worst origins before declaring a winner.
Probabilistic forecast evaluation
Your Bayesian forecasting model does not return one number for tomorrow; it returns a whole distribution. This chapter shows how to grade that distribution. Do the intervals contain the outcomes as often as they claim (coverage and calibration)? Are they as narrow as they can honestly be (sharpness)? And which scores reward the whole distribution at once (log score, CRPS, pinball loss)? These are the tools that take you beyond "the RMSE improved".
- Explain why a point-forecast score such as RMSE cannot see whether a forecast's uncertainty is honest
- Measure prediction-interval coverage, and judge it against its own sampling noise ("a 90% interval should contain about 90% of outcomes")
- Read a PIT histogram: flat = calibrated, U = overconfident, hump = underconfident, slope = biased
- State the rule "calibrated but not unnecessarily wide" and measure sharpness, with the interval score as the combined check
- Compute and interpret the log score (log predictive density), the CRPS (area picture; reduces to the absolute error for a point forecast) and the pinball loss for quantiles
- Know what a proper scoring rule is, why these scores qualify and coverage alone does not
- Evaluate by horizon and origin, and compute everything from forecast samples in
numpy
What we need from earlier chapters: a predictive distribution, its mean, median, quantiles and intervals (Chapter 7.14); rolling-origin evaluation, MAE and RMSE (Chapter 7.15); the Normal CDF and quantiles (Chapter 4.4, 4.9); the Student-t and Negative Binomial as forecast likelihoods (Chapter 7.13); credible intervals (Chapter 6.4); KL divergence (Chapter 6.11). Notation: $F_t$ is the predictive CDF (cumulative distribution function: $F_t(x) = P(Y_t \le x)$) for day $t$ and $p_t$ its density; $y_t$ is the outcome that actually happens; $[l_t, u_t]$ is a central prediction interval with nominal level $1-\alpha$ (so $\alpha = 0.2$ gives an 80% interval); $\tau$ ("tau") is a quantile level between 0 and 1 and $q_\tau$ the forecast's $\tau$-quantile; $x_1, \dots, x_S$ are $S$ forecast samples, for example posterior predictive draws.
From one number to a distribution: what "RMSE improved" hides core
Two weather forecasters both say "tomorrow's high will be 20 degrees". One adds "give or take 1 degree", the other "give or take 8". If it turns out to be 25, the first was badly overconfident and the second was fine, yet a score that only looks at "20 versus 25" gives them exactly the same mark. The uncertainty is half of the forecast, and a point score throws it away.
A Bayesian forecasting model (Chapter 7.14) gives a predictive distribution for every future day, and decisions use it: "what is the chance demand exceeds capacity?", "how much stock covers 95% of days?". Those answers are only worth trusting if the distribution is honest. So we need scores that look at the whole distribution, not just its centre.
Three ways to say it:
- Picture: a forecast is a fan (a median line with bands around it); a point score looks at the middle line only, and ignores the width of the fan.
- Numbers: forecaster X has the better centre (RMSE 1.00 against 1.04) but an 80% interval that catches the outcome only 39% of the time; forecaster Y has a slightly worse centre and catches it 82% of the time.
- Slogan: a forecast has two jobs, to be centred and to be honest about its doubt; RMSE grades only the first.
The true outcomes are $Y \sim N(0, 1)$ every day (mean 0, standard deviation 1). Two forecasters each issue a Normal predictive distribution.
- X: mean $0$ (exactly right), standard deviation $0.4$ (very confident).
- Y: mean $0.3$ (a bit off), standard deviation $1.1$ (honest about the doubt).
- Point score. The RMSE of a forecaster's mean is $\sqrt{\text{Var} + \text{offset}^2}$. For X: $\sqrt{1 + 0} = 1.000$. For Y: $\sqrt{1 + 0.09} \approx 1.044$. RMSE prefers X.
- 80% interval for X: $0 \pm 1.2816 \times 0.4 = \pm 0.51$. The chance that an outcome lands inside is $P(\lvert Z\rvert \le 0.51) = 2\Phi(0.51) - 1 \approx 0.39$. It claims 80% and delivers 39%.
- 80% interval for Y: $0.3 \pm 1.2816 \times 1.1 = [-1.11, 1.71]$. The chance of landing inside is $\Phi(1.71) - \Phi(-1.11) \approx 0.956 - 0.134 = 0.82$. It claims 80% and delivers 82%.
- Whole-distribution scores (formulas in the next sections), averaged over many days: mean CRPS 0.634 for X and 0.590 for Y (lower is better); mean log score $-3.13$ for X and $-1.46$ for Y (higher is better).
- So RMSE says "X is better" and both whole-distribution scores say "Y is much better". For anyone acting on the intervals, Y is the forecaster to trust.
A probabilistic forecast for day $t$ is a full predictive distribution $F_t$ (often given as forecast samples, quantiles, or a few interval pairs). Evaluating it has three parts:
- Calibration (reliability): the forecast's probabilities match how often things happen. Events it calls "20% likely" happen about 20% of the time; its 90% intervals contain about 90% of outcomes.
- Sharpness: how concentrated the forecast is (how narrow its intervals are). A property of the forecasts alone, not of the outcomes.
- Proper scoring rules: single numbers (log score, CRPS, pinball, interval score) that reward calibration and sharpness together and are minimized in expectation by the true distribution (Section 8).
Guiding principle (Gneiting and colleagues): maximize sharpness subject to calibration. A forecast must first be honest; among honest forecasts, the sharper the better.
Why do we need it?
Decisions use probabilities and intervals (stock for 95% of days, the chance of exceeding capacity). If those are wrong, decisions are wrong, even when the centre is perfect. A point score is blind to this by construction.
Where is it used?
Weather and energy forecasting (probabilistic scores are standard there), the quantile track of forecasting competitions (the M5 uncertainty track used a pinball-based score), financial risk (Value at Risk is a quantile), and any Bayesian forecast you want to trust.
How is it used?
Keep the forecast samples (or quantiles) from every origin and horizon of your rolling-origin run (Chapter 7.15). Then compute coverage, the PIT histogram, and the mean CRPS and log score. Report them next to RMSE, per horizon.
"My model's RMSE went down, so the forecasts are better."
Only the centre improved. The intervals may have become more overconfident or wider than needed. Check coverage and a proper score such as CRPS before you say "better".
"If the point forecast is good, the intervals are fine."
The interval comes from the model's spread, which is a separate quantity (the noise scale, parameter uncertainty, the tails). A model can have a perfect centre and a spread that is 3 times too small.
"Probabilistic evaluation is only for fancy models."
Any forecast with an interval needs it, from ETS to a Bayesian model; and even naive forecasts can be turned into probabilistic ones (the naive forecast plus the spread of its past errors).
Your forecasting model produces a predictive distribution for every future day through its Normal, Student-t or Negative Binomial likelihood plus the uncertainty in the parameters (7.14). When someone says "the new version improved RMSE", the follow-up questions are: did the 80% and 95% intervals keep their coverage? Did the CRPS or the log score improve on rolling origins? In an A/B framework like yours the same distinction exists: a posterior probability such as $P(\theta_B \gt \theta_A \mid D)$ is only useful if the model behind it is calibrated, which is exactly what the checks of this chapter test for forecasts.
A forecast = a distribution. Grade it for calibration (probabilities match frequencies), sharpness (narrow) and with proper scores (CRPS, log score, pinball).
Principle: maximize sharpness subject to calibration.
Trap: RMSE of the mean is blind to the spread; "RMSE improved" says nothing about the intervals.
Quick check: two models have identical RMSE. Model A's 90% intervals contain 62% of outcomes, model B's contain 91%. Which would you ship, and why?
Model B. Its intervals do what they claim. Model A is overconfident: it says 90% but delivers 62%, so any decision based on its intervals (stock for 95% of days, chance of exceeding capacity) will be wrong far too often. With equal RMSE the centres are equally good, so honesty about uncertainty is the deciding difference.
Prediction-interval coverage: does the 90% interval contain 90%? core
A weather service says "we are 90% sure the temperature stays inside this band". If you collect a year of its bands and outcomes and find that the temperature was inside the band only 60% of days, you stop believing the "90%". That check is coverage: of all the intervals a forecaster issued, which share contained the outcome?
If the intervals are honest, the share should be close to the stated level. It will not be exactly equal, because a finite number of days is a noisy sample: out of 20 days an honest 80% interval may easily catch 14 or 18. So we compare the observed share with the amount of wobble that luck alone would produce.
Three ways to say it:
- Picture: outcomes as dots and the interval as a band; count the dots outside the band.
- Numbers: an 80% interval should catch about 16 of 20 days (and 160 of 200). Catching 13 of 20 is within luck; catching 130 of 200 is not.
- Slogan: a 90% interval should contain about 90% of outcomes, no more, no less.
Each day a model issues an 80% interval for the next day's orders. Look at the hits (outcome inside the interval).
- 20 days, 13 hits. Observed coverage $= 13/20 = 65\%$. Nominal coverage $= 80\%$.
- If the intervals were honest, the number of hits would be like a Binomial with $n = 20$ and $p = 0.8$: standard deviation of the share $= \sqrt{0.8 \times 0.2/20} \approx 0.089$.
- A rough 95% range for the observed share under honesty: $0.80 \pm 1.96 \times 0.089 = [0.625, 0.975]$. The observed $0.65$ is inside it, so 20 days are not enough to say the intervals are wrong.
- 200 days, 130 hits. Same observed share, $65\%$. Now the standard deviation is $\sqrt{0.16/200} \approx 0.028$ and the range is $0.80 \pm 0.055 = [0.745, 0.855]$. The observed $0.65$ is far outside: the intervals are clearly overconfident.
- So "coverage 65% against 80%" is a weak signal with 20 days and a strong one with 200. Always quote $n$.
For outcomes $y_1, \dots, y_n$ and central prediction intervals $[l_t, u_t]$ with nominal level $1-\alpha$:
$$\text{coverage} = \frac1n \sum_{t=1}^{n} \mathbf 1\{\, l_t \le y_t \le u_t \,\}, \qquad \text{standard error under honesty} \approx \sqrt{\frac{(1-\alpha)\,\alpha}{n}} .$$- The indicator $\mathbf 1\{\cdot\}$ is 1 when the outcome is inside the interval and 0 otherwise. Coverage is just the share of hits.
- Calibrated (for this level) means the long-run share equals $1-\alpha$. Under-coverage (observed below nominal) means the intervals are too narrow: overconfident. Over-coverage means too wide: underconfident (or cautious).
- The hits should also be roughly independent over time. Misses that come in clusters (all on holidays, all after a level jump) mean the model is wrong in specific situations even if the average looks fine.
- Check several levels (50%, 80%, 95%) and slice the hits by horizon, weekday and holiday versus normal day. A predictive interval comes from the predictive distribution, which includes observation noise; it is wider than a credible interval for the trend alone (Chapter 7.14).
Why do we need it?
Coverage is the simplest honest test of an interval: it asks only "did the outcome fall inside?". It needs no model internals, so it works for any forecaster, Bayesian or not, and anyone can understand the result.
Where is it used?
Prophet's cross-validation output includes a coverage column; ETS and ARIMA interval checks in statsmodels; weather and energy forecast verification; and risk models, where it is called a backtest of Value at Risk.
How is it used?
For every origin and horizon, take the interval from the forecast samples (np.quantile(samples, [0.1, 0.9]) for an 80% interval), compare it with the outcome, average the hits, and compare the result with the nominal level using the binomial wobble $\sqrt{p(1-p)/n}$.
"My 90% intervals contained 85% of the 20 test days, so they are miscalibrated."
With $n = 20$ the standard error is $\sqrt{0.9 \times 0.1/20} \approx 0.067$. 85% is less than one standard error below 90%. Judge coverage with its sampling noise, and with as many forecasts as you can get (many origins and horizons).
"Coverage of 90% on average means the intervals are good everywhere."
Average coverage can hide failures: perfect on quiet days and badly wrong on holidays, or good at $h = 1$ and poor at $h = 14$. Slice by horizon, weekday and event.
"Higher coverage is always better."
100% coverage is easy to get: make the interval from $-\infty$ to $+\infty$. Coverage must match the stated level, and intervals should be as narrow as that allows (next sections).
"A 90% credible interval from my Bayesian model is guaranteed to cover 90% of future outcomes."
It is a statement inside the model. If the model is wrong (wrong likelihood, missed changepoint, overconfident approximate posterior) the real coverage can be much lower. Coverage is how you find out.
For your forecasting model, build the interval from the posterior predictive samples of future days (the draws that include the observation noise of the Normal, Student-t or Negative Binomial likelihood, 7.14), not from the posterior of the trend alone. The trend-only band is far narrower and will show terrible coverage that says nothing about the real forecast. Compute coverage at 50%, 80% and 95% over your rolling origins, and per horizon. If you use a mean-field or low-rank SVI guide, parameter uncertainty can be understated (Chapter 6.13); under-coverage that grows with the horizon is the typical sign.
"A 90% prediction interval means there is a 90% chance the next value is inside it, so the model's job is done."
"A 90% prediction interval is calibrated if, over many forecasts, about 90% of outcomes fall inside. I verify that on a rolling-origin backtest, per horizon, and compare the observed share with its binomial wobble."
Model answer: "Coverage is the share of outcomes that land inside the stated interval. If it is lower than nominal the model is overconfident, if higher it is too cautious. I check it at several levels and by horizon, and I judge it with its standard error, $\sqrt{p(1-p)/n}$, because a short record is noisy. Coverage alone is not enough, since a very wide interval covers everything; I add a sharpness measure or a proper score such as CRPS."
coverage $= \frac1n\sum \mathbf 1\{l_t \le y_t \le u_t\}$; honest wobble $\approx \sqrt{(1-\alpha)\alpha/n}$.
Below nominal = overconfident (too narrow); above = too cautious. Check by level, horizon and event.
Trap: judging coverage without $n$; use predictive (not trend-only) intervals; very wide intervals cover everything.
Quick check: 50 days, an 80% interval, 44 hits. Is that evidence of miscalibration?
Observed $44/50 = 88\%$. Standard error $= \sqrt{0.16/50} \approx 0.057$; $1.96\times0.057 \approx 0.111$, so the luck range is $[69\%, 91\%]$. 88% is inside, so there is no clear evidence of miscalibration (the intervals may be a little cautious, but 50 days cannot tell).
Calibration and the PIT histogram core
Coverage checks one interval at a time. A more complete check asks: where in its own forecast distribution does each outcome land? Picture the forecast as a ladder of probability from 0 (the very bottom of the distribution) to 1 (the very top). If the forecast is honest, outcomes should land on every rung equally often: as often in the bottom tenth as in the middle tenth or the top tenth. Each outcome's position on the ladder is its PIT value (probability integral transform): the forecast's CDF evaluated at the outcome.
Pile all the PIT values into a 10-bar histogram. An honest forecaster gives a flat histogram. A forecaster that is too sure of itself is surprised too often, so outcomes land in the extreme bins and the histogram is a U. One that is too vague never gets surprised, and its histogram is a hump in the middle. A forecaster whose centre is shifted piles all outcomes to one side: a slope.
Three ways to say it:
- Picture: a 10-bar histogram of "how high up its own forecast did the outcome land?"; flat is good.
- Numbers: forecast $N(100, 10^2)$: an outcome of 85 has PIT $\Phi(-1.5) = 0.067$ (near the bottom), 100 gives $0.5$, 125 gives $\Phi(2.5) = 0.994$ (near the top).
- Slogan: U = overconfident, hump = underconfident, slope = biased, flat = calibrated.
A model forecasts each day's orders as a Normal $N(100, 10^2)$. Three outcomes arrive: 85, 100 and 125.
- Standardize: $z = (y - 100)/10$ gives $-1.5$, $0$ and $2.5$.
- PIT $= \Phi(z)$ (the Normal CDF): $\Phi(-1.5) \approx 0.067$, $\Phi(0) = 0.5$, $\Phi(2.5) \approx 0.994$.
- So the first outcome sat at the 7th percentile of the forecast, the second at the median, the third at the 99.4th percentile. The last is a very surprising value.
- Now 100 outcomes in total and their PIT values counted in 10 bins give, say, $30, 8, 6, 5, 4, 5, 6, 7, 9, 20$. An honest forecast would give about 10 per bin. Here 50 of the 100 outcomes fall in the two extreme bins (expected: 20): the histogram is U-shaped, and the forecast is overconfident. Its distribution is too narrow, so outcomes land in its tails too often.
- The lopsided counts (30 at the bottom against 20 at the top) also hint that the forecast is a bit too high on average.
For a forecast with CDF $F_t$ and outcome $y_t$, the PIT value is
$$u_t = F_t(y_t) \in [0, 1].$$- Theorem (plain words): if $F_t$ is the true distribution of $y_t$, then $u_t$ is uniformly distributed on $[0, 1]$ (every tenth of the probability is equally likely to contain the outcome), and the $u_t$ are independent over time if the forecasts use all the information in the past.
- From forecast samples $x_1, \dots, x_S$: $u_t \approx \frac1S \#\{s : x_s \le y_t\}$, the share of samples below the outcome.
- Reading the histogram: flat = calibrated · U-shape = overconfident (forecast too narrow) · hump = underconfident (too wide) · piled to the left = the forecast is too high · piled to the right = too low.
- Counts (Negative Binomial forecasts) have a CDF with jumps, so $F(y)$ is not exactly uniform. Use a randomized PIT (pick a random value inside the jump at the observed count) before drawing the histogram.
- A calibration curve plots, for each level $c$, the share of outcomes inside the central $c$ interval against $c$ (outcome inside the central $c$ interval means $u_t \in [\frac{1-c}{2}, \frac{1+c}{2}]$). Calibrated means it lies on the diagonal.
Why do we need it?
Coverage tests one level at a time. The PIT histogram tests all levels at once and also shows how a forecast is wrong (too narrow, too wide, shifted), which tells you what to fix.
Where is it used?
Weather and economic forecast verification (rank histograms are the same idea for ensembles), Bayesian model checking (it is a posterior predictive check on the probability scale, Chapter 6.8), and the diagnostic plots of probabilistic forecasting libraries.
How is it used?
For each forecast, compute (samples < y).mean() (or the Normal CDF at the outcome). Plot a 10-bin histogram of all PIT values over origins and horizons. Compare each bar with the flat line and its binomial luck range; fix the model according to the shape.
"A flat PIT histogram means the forecast is good."
It means it is calibrated, which is necessary but not sufficient. A lazy forecaster who always issues the long-run distribution of the series (the same wide distribution every day) is also calibrated, but useless. Sharpness (next section) and proper scores catch that.
"A U-shaped histogram means the outcomes are too extreme."
It means the forecast distribution is too narrow for the outcomes it faces: the outcomes keep landing in its tails. The fix is to widen the forecast (more noise, more parameter uncertainty, heavier tails), not to change the data.
"Bars wobbling around the flat line prove miscalibration."
Each bar is a count and counts wobble. Compare with the binomial luck range ($\pm 1.96\sqrt{0.1\times0.9/n}$ relative to 0.1), not with an exact flat line.
"I can use the same PIT for counts directly."
The CDF of a count has jumps, so $F(y)$ is not exactly uniform even for a perfect forecast; use a randomized PIT (or a non-randomized variant designed for counts).
Draw the PIT histogram of your model's one-day-ahead forecasts on rolling origins. A U-shape is the typical sign of an overconfident forecast: a Normal likelihood on heavy-tailed residuals, too little parameter uncertainty (for example from a mean-field SVI guide), or a trend that cannot follow a new changepoint. A hump suggests the forecast spread is larger than it needs to be, for example a noise scale that grew during fitting to absorb a pattern the mean function missed. If your forecasts are counts under a Negative Binomial likelihood, use the randomized PIT. Compare the PIT of a Normal, Student-t and Negative Binomial version of the model (7.13).
"The PIT is just another accuracy number."
"The PIT histogram is a calibration diagnostic: it shows whether the observed outcomes are spread over the forecast distribution the way an honest forecast predicts."
Model answer: "For each forecast I compute the CDF at the outcome. If the forecasts are honest, these values are uniform on 0 to 1. A U-shaped histogram means too many outcomes in the tails, so the forecasts are overconfident. A hump means the forecasts are too wide. A tilt means a bias. It is a necessary check, not a sufficient one, because a forecast that is always the long-run distribution is calibrated but not sharp."
PIT $u_t = F_t(y_t)$; honest forecast → $u_t \sim$ Uniform(0, 1). From samples: share of samples $\le y_t$.
flat = calibrated · U = overconfident · hump = underconfident · slope = biased. Counts: randomized PIT.
Trap: flat is necessary, not sufficient (always-the-same-wide forecast is flat); judge bars against their luck range.
Quick check: a forecaster's PIT values are mostly near 0.5, with very few below 0.1 or above 0.9. What is wrong, and what do you change?
A hump: outcomes land in the middle of the forecast far more often than they should, so the forecast distributions are too wide (underconfident). Reduce the forecast spread (tighter noise scale or less prior uncertainty) until the histogram flattens, then check that coverage stays at the nominal level.
Sharpness: calibrated, but not unnecessarily wide core
Suppose a forecaster says: "tomorrow's orders will be between 0 and 10 000". The statement is almost certainly true, so the forecaster is perfectly "calibrated" at any level. It is also useless: you cannot plan staff from it. At the other extreme, "tomorrow's orders will be 123 to 124" is very informative, but if it is wrong most days it is dangerous.
A good forecast must be honest first (calibrated) and then as narrow as honesty allows (sharp). Sharpness is how tightly the forecast concentrates: for intervals, the width. It is a property of the forecasts alone; it does not look at the outcomes. That is why it cannot be used by itself (narrower is always "sharper") and why coverage cannot be used by itself (wider always "covers more").
Three ways to say it:
- Picture: coverage rises as you widen the band; the width cost rises too; the best band sits where the two meet.
- Numbers: for outcomes $N(100, 10^2)$, an 80% band of width 25.6 covers 80%; a band of width 12.8 covers only 48%; a band of width 51.2 covers 99% but is twice as wide as needed.
- Slogan: be calibrated first, then be as sharp as you can.
The truth is $N(100, 10^2)$. Three forecasters give an 80% interval (so $\alpha = 0.2$ and $2/\alpha = 10$), each centred at 100: narrow $\pm 6.4$, honest $\pm 12.8$ (that is $1.2816 \times 10$), wide $\pm 25.6$. Score each with the interval score: the width, plus $10 \times$ the distance by which the outcome misses the interval.
- Widths: narrow $12.8$, honest $25.6$, wide $51.2$. Coverage of the truth: $P(\lvert Z\rvert \le 0.64) = 48\%$, $80\%$, $P(\lvert Z\rvert \le 2.56) = 99\%$.
- Outcome 100 (right in the middle): all three contain it, so the interval score is just the width: narrow $12.8$ (the winner), honest $25.6$, wide $51.2$.
- Outcome 120 (an upper surprise): narrow misses by $120 - 106.4 = 13.6$, score $12.8 + 10\times13.6 = 148.8$. Honest misses by $120 - 112.8 = 7.2$, score $25.6 + 72 = 97.6$. Wide contains it: score $51.2$.
- Averaged over all possible outcomes (expectation under the truth): narrow $\approx 44.4$, honest $\approx 35.1$, wide $\approx 51.6$. The honest interval has the lowest expected score. The narrow one is punished by its misses, the wide one by its width.
For a central $(1-\alpha)$ interval $[l, u]$ and outcome $y$, the interval score (also called the Winkler score) is
$$\text{IS}_\alpha(l, u; y) = (u - l) + \frac{2}{\alpha}(l - y)\,\mathbf 1\{y \lt l\} + \frac{2}{\alpha}(y - u)\,\mathbf 1\{y \gt u\}.$$- First term: width (sharpness). The other terms: a penalty for missing, growing with how far the outcome is outside, and larger for a smaller $\alpha$ (a 95% interval that misses is punished more than an 80% one).
- Lower is better. It is a proper score for the pair of quantiles $(\alpha/2,\, 1-\alpha/2)$: in expectation it is smallest for the true quantiles (it equals $\tfrac{2}{\alpha}$ times the sum of two pinball losses, Section 7).
- Calibration is what coverage and the PIT measure; sharpness is the average interval width; the interval score combines the two. A forecaster cannot improve it by widening (the width term grows) or by narrowing (the miss penalty grows) beyond the honest interval.
Why do we need it?
Coverage alone is cheap to satisfy (make the interval huge). Sharpness alone is cheap too (make it tiny). Reporting them together, or a score that combines them, prevents both cheats and tells you the cost of extra caution.
Where is it used?
Forecast verification in weather and epidemic forecasting (for example the COVID-19 forecast hubs used interval scores), economic forecasting, and the interval-width columns of any forecasting report; the width is also what a planner pays for in extra stock or staff.
How is it used?
Report three things per level: coverage, mean interval width (in the data's units, or relative to the level of the series), and the mean interval score. Among models with the same (good) coverage, prefer the narrower one.
"Narrower intervals mean a better model."
Only if the coverage stays at the nominal level. A narrow interval that misses too often is overconfident, and its interval score is worse than an honest one.
"My 95% intervals contain 99% of outcomes, so I am on the safe side."
That is over-coverage: the intervals are wider than the model's uncertainty justifies, and someone paid for it (more stock, more staff). Calibration works in both directions.
"Sharpness is a property of the model's accuracy."
Sharpness is a property of the forecasts, not of whether they were right. Two forecasters can be equally sharp and very different in quality; that is why it is paired with calibration.
Prior and likelihood choices in your forecasting model move the width of the predictive intervals: a heavier-tailed likelihood or an extra source of parameter uncertainty widens them, a tight prior or an over-confident guide narrows them. Sharpness tells you what these choices cost or save. When you compare model variants (Normal against Student-t, different Fourier orders, different Laplace scales for the changepoints), look at coverage and mean width, or at the interval score, rather than at width alone. If two variants have the same good coverage, the narrower one is the better forecast.
"We want the tightest possible intervals, so we trained for narrow bands."
"We want calibrated intervals that are as narrow as that allows: maximize sharpness subject to calibration."
Model answer: "Coverage and sharpness pull in opposite directions: widening the interval raises coverage and lowers sharpness. I first check that coverage matches the nominal level at every horizon, then compare mean widths among the calibrated models. The interval score or CRPS combines both, and an honest forecast minimizes them in expectation."
$\text{IS}_\alpha = (u - l) + \frac2\alpha(l-y)\mathbf 1\{y \lt l\} + \frac2\alpha(y-u)\mathbf 1\{y \gt u\}$ = width + miss penalty. Lower is better.
Principle: maximize sharpness (narrowness) subject to calibration.
Trap: coverage alone rewards wide intervals, width alone rewards narrow ones; use them together.
Quick check: 80% interval $[90, 110]$, outcome $120$. What is the interval score?
Width $= 20$. The outcome is above the interval by $120 - 110 = 10$. With $\alpha = 0.2$: $\text{IS} = 20 + (2/0.2)\times10 = 20 + 100 = 120$. Had the outcome been inside, the score would just be the width, 20.
The log score: how much probability did you give to what happened? core
A forecast is a bet spread over all possible outcomes. When the outcome arrives, ask: how much of my bet was sitting on it? A forecast that put a tall, narrow spike of density near the value that happened did well. One that put almost nothing there did badly. The height of the forecast curve above the outcome is the predictive density at the outcome; its logarithm is the log score (higher is better).
Why the logarithm? Over many days the probabilities multiply (independent days: $p(y_1)\,p(y_2)\cdots$), and a product of many small numbers is awkward. The log turns the product into a sum, exactly like the log-likelihood you know from Chapter 5.2. The log score of a model on held-out data is its out-of-sample log-likelihood.
The log score is merciless to overconfidence. A forecast that is very narrow gets a huge density when it is right, but a density near zero, whose log is hugely negative, when the outcome is only a few standard deviations away.
Three ways to say it:
- Picture: read the height of the forecast curve at the place where the outcome fell, then take its log.
- Numbers: outcome 112; forecast $N(100, 10^2)$ scores $-3.94$, a too-narrow $N(100, 5^2)$ scores $-5.41$, and for an outcome of 135 the narrow one collapses to $-27.0$ while the honest one scores $-9.35$.
- Slogan: log score = the log of the probability density that the forecast gave to what happened.
Three forecasters all centre on 100 with standard deviations 10 (honest), 5 (too confident) and 20 (too vague). The Normal log density is $\log p(y) = -\tfrac12 \log(2\pi\sigma^2) - \dfrac{(y-\mu)^2}{2\sigma^2}$.
- Outcome 112. $\sigma = 10$: $-\tfrac12\log(2\pi\cdot100) - \tfrac{144}{200} = -3.222 - 0.72 = -3.94$. $\sigma = 5$: $-\tfrac12\log(2\pi\cdot25) - \tfrac{144}{50} = -2.528 - 2.88 = -5.41$. $\sigma = 20$: $-\tfrac12\log(2\pi\cdot400) - \tfrac{144}{800} = -3.915 - 0.18 = -4.09$.
- Outcome 94 (typical): the three scores are $-3.40$, $-3.25$ and $-3.96$. Here the narrow forecaster is best, because it was lucky that the outcome was close.
- Average of the two outcomes: honest $-3.67$, too confident $-4.33$, too vague $-4.03$. The honest forecaster already leads.
- Outcome 135 (a 3.5-sigma surprise for the honest forecaster): $-9.35$ for $\sigma = 10$, but $-27.03$ for $\sigma = 5$ (a 7-sigma surprise for it) and $-5.45$ for $\sigma = 20$. One surprising day can wipe out the gains of many lucky days.
- Posterior samples. If the predictive density is an average over $S$ posterior draws, $p(y) \approx \frac1S\sum_s p(y\mid\theta_s)$, take the log of the average. With two draws giving densities $0.1$ and $0.3$: $\log\frac{0.1+0.3}{2} = \log 0.2 = -1.609$. The average of the logs, $(\log 0.1 + \log 0.3)/2 = -1.753$, is lower and wrong.
For a predictive density $p_t$ (or probability mass function for counts) and outcome $y_t$:
$$\text{LS}_t = \log p_t(y_t), \qquad \text{mean log score} = \frac1n\sum_{t=1}^{n}\log p_t(y_t) .$$- Higher is better. Some books use the negative log score (also called the negative log predictive density, NLPD) as a loss: lower is better. Check which one a library prints.
- The units are nats (natural log). A difference of 0.1 between two models' mean log scores means one assigned about $e^{0.1} \approx 1.105$ times more probability to the outcomes per forecast.
- It is a proper score: in expectation it is highest for the true distribution. The expected gap between the true density $p$ and a forecast $f$ is the KL divergence, $E_p[\log p - \log f] = \text{KL}(p\,\|\,f)$ (Chapter 6.11).
- From posterior draws $\theta_1, \dots, \theta_S$: $\log p(y\mid D) \approx \log\!\big(\tfrac1S\sum_s p(y\mid\theta_s)\big) = \text{logsumexp}_s \log p(y\mid\theta_s) - \log S$. Summed over held-out points this is the expected log predictive density (ELPD) used to compare Bayesian models.
- Limits: it needs a density, so forecast samples alone are not enough (you must smooth them or use the model's own likelihood); and it is very sensitive to outcomes in the far tail, so a single extreme day can dominate it.
Why do we need it?
We want one number that rewards putting probability where the truth lands and punishes false certainty. The log score does this with the same quantity (the likelihood) that statisticians already use to fit and compare models.
Where is it used?
Bayesian model comparison (ELPD, LOO and WAIC), training of probabilistic forecasters (the DeepAR network is trained by minimizing the negative log-likelihood), and probabilistic prediction competitions that score a density.
How is it used?
On a held-out window, evaluate log p(y_t | forecast) with the model's likelihood. With posterior draws use logsumexp. Average over forecasts, and compare models by the difference. Keep an eye on the worst few days, which usually dominate.
"The log score is a probability, so it lies between 0 and 1."
It is the log of a density (which can exceed 1, so the log can be positive) or of a probability (log between $-\infty$ and 0). Only differences between models are meaningful, and only on the same outcomes and the same scale of the data.
"I average the log-likelihood over the posterior draws to get the predictive log score."
Average the likelihood first, then take the log (logsumexp minus $\log S$). The average of the logs is always lower and answers a different question.
"A higher log score on 10 days proves model A is better."
The log score is dominated by the few worst days. Look at the per-day scores, and compare models over many origins before concluding.
"I can compare log scores of a count model and a continuous model."
A pmf (probability) and a density are on different scales. Compare models of the same kind of outcome, or discretize consistently.
For your forecasting model, the log predictive density of a held-out day is $\log\frac1S\sum_s p(y\mid\theta_s)$ over posterior (or SVI-approximate posterior) draws, using the same Normal, Student-t or Negative Binomial likelihood as in training (7.13); for the Negative Binomial it is the log of a probability mass. This is the held-out version of the term your training objective already contains (the expected log-likelihood part of the ELBO, Chapter 6.12). Because it punishes tail surprises so hard, it is the score that shows most directly why a Student-t likelihood can beat a Normal on data with outliers, even when RMSE is similar.
"Log score is the same as the likelihood on the training data."
"The log score is the log predictive density of new outcomes, so it is the out-of-sample log-likelihood; the training log-likelihood uses data the model has already seen."
Model answer: "The log score takes the log of the density the forecast assigned to the observed value, averaged over held-out forecasts. It rewards sharp forecasts that are right and heavily punishes confident forecasts that are wrong, so it is a proper score. The difference between two models' mean log scores is how many nats of probability per forecast one gains over the other, and in expectation the gap between the true distribution and a model is a KL divergence, so the difference between two models is a difference of two KL divergences."
$\text{LS}_t = \log p_t(y_t)$ (higher is better, in nats). Posterior draws: $\log\frac1S\sum_s p(y\mid\theta_s)$ — log of the mean, not mean of the logs.
Proper; expected gap to the truth = KL divergence. Very sensitive to tail surprises.
Trap: needs a density; one extreme day can dominate; do not mix count and continuous scales.
Quick check: forecast $N(50, 4^2)$, outcome 58. What is the log score? Is the outcome a surprise?
$z = (58-50)/4 = 2$. $\log p = -\tfrac12\log(2\pi\cdot16) - 2^2/2 = -2.305 - 2 = -4.305$. A 2-sigma outcome happens about 5% of the time, so it is a mild surprise; the honest score for a typical outcome ($z \approx 0$) would be $-2.305$, so the surprise cost 2 nats.
CRPS: the gap between the forecast CDF and the outcome core
Draw the forecast as a CDF: an S-shaped curve that climbs from 0 to 1 and says "the probability of ending up below this value". Now draw what a perfect, fully certain forecast would draw once the outcome $y$ is known: a step that is 0 before $y$ and jumps to 1 at $y$. The closer the S-curve is to the step, the better the forecast. The CRPS (continuous ranked probability score) measures that gap by squaring the vertical difference at every point and adding it up along the horizontal axis.
Two nice facts follow. (1) The CRPS has the units of the data (orders), like the MAE. (2) If the forecast is a single number $f$, its "S-curve" is itself a step at $f$, so the gap between two steps is just $\lvert y - f\rvert$: the CRPS of a point forecast is the absolute error, and its average is the MAE. So the CRPS is the MAE for whole distributions. A good spread can even make it smaller than the absolute error of the centre.
Three ways to say it:
- Picture: the forecast's S-curve against the outcome's step; the shaded gap (squared) is the score.
- Numbers: forecast $N(100, 10^2)$, outcome 112: CRPS $= 7.48$ orders, while the bare point forecast 100 has an absolute error of 12.
- Slogan: CRPS = MAE for distributions.
A small sample forecast. A forecaster gives three equally likely values $\{8, 10, 12\}$ (its forecast samples). The outcome is $y = 11$.
- The forecast CDF is a staircase: $0$ below 8, $1/3$ on $[8,10)$, $2/3$ on $[10,12)$, $1$ from 12 on. The outcome step is $0$ below 11 and $1$ from 11 on.
- Squared gaps: on $[8,10)$ the gap is $1/3$, squared $1/9$, over length 2: $2/9$. On $[10,11)$ the gap is $2/3$, squared $4/9$, length 1: $4/9$. On $[11,12)$ the gap is $2/3 - 1 = -1/3$, squared $1/9$, length 1: $1/9$.
- CRPS $= 2/9 + 4/9 + 1/9 = 7/9 \approx 0.778$.
- Sample formula check: $E\lvert X - y\rvert = (3 + 1 + 1)/3 = 5/3$. $E\lvert X - X'\rvert$ over all 9 ordered pairs $= (0+2+4+2+0+2+4+2+0)/9 = 16/9$. CRPS $= 5/3 - \tfrac12\cdot\tfrac{16}{9} = 15/9 - 8/9 = 7/9$. Same answer.
A Normal forecast. $N(100, 10^2)$, outcome 112, so $z = 1.2$. The closed form is $\sigma\big[z(2\Phi(z) - 1) + 2\varphi(z) - 1/\sqrt\pi\big]$ with $\Phi$ the Normal CDF and $\varphi$ its density.
- $\Phi(1.2) = 0.8849$, so $2\Phi(1.2) - 1 = 0.7699$ and $z \times 0.7699 = 0.9238$.
- $\varphi(1.2) = 0.1942$, so $2\varphi = 0.3884$. And $1/\sqrt\pi = 0.5642$.
- Sum: $0.9238 + 0.3884 - 0.5642 = 0.7480$; times $\sigma = 10$: CRPS $= 7.48$ orders.
- Compare: a point forecast of 100 has CRPS $= \lvert 112 - 100\rvert = 12$. The plain (unsquared) area between the CDF and the step is $E\lvert X - y\rvert = 13.12$, larger than the CRPS because squaring shrinks small gaps.
For a forecast CDF $F$ and outcome $y$:
$$\text{CRPS}(F, y) = \int_{-\infty}^{\infty} \big(F(x) - \mathbf 1\{x \ge y\}\big)^2\,dx = E\lvert X - y\rvert - \tfrac12 E\lvert X - X'\rvert ,$$where $X$ and $X'$ are independent draws from $F$.
- Same units as $y$; lower is better; $\text{CRPS} \ge 0$, and $0$ only for a perfect point forecast that equals $y$.
- Point forecast: if $F$ is a step at $f$, $\text{CRPS} = \lvert y - f\rvert$. Averaged over forecasts, CRPS reduces to the MAE.
- Normal forecast: $\text{CRPS} = \sigma\big[z(2\Phi(z)-1) + 2\varphi(z) - 1/\sqrt\pi\big]$, $z = (y-\mu)/\sigma$ (this is
LA.stats.metrics.crpsNormalin the widgets). - From samples $x_1, \dots, x_S$: $\text{CRPS} \approx \frac1S\sum_i \lvert x_i - y\rvert - \frac{1}{2S^2}\sum_{i,j}\lvert x_i - x_j\rvert$. No density is needed, which makes it ideal for posterior predictive samples.
- It is a proper score, and it relates to pinball loss (Section 7) as $\text{CRPS} = 2\int_0^1 \rho_\tau\big(y, F^{-1}(\tau)\big)\,d\tau$: the average quantile loss over all quantile levels.
- Compared with the log score: it grows only linearly with the distance of the outcome, so a single extreme day does not dominate, and it works directly with samples. It is less sensitive to the exact shape of the tails.
Why do we need it?
We want one number for a whole predictive distribution that is in the units of the data, easy to explain ("average miss, but for distributions"), robust to extreme days, and computable straight from forecast samples.
Where is it used?
Weather and ensemble forecast verification (the standard score there), energy and electricity-price forecasting, probabilistic forecasting libraries (properscoring, scoringrules, gluonts), and as the "probabilistic MAE" in forecasting papers.
How is it used?
For each forecast keep $S$ samples (for example posterior predictive draws); compute np.mean(abs(x - y)) - 0.5 * np.mean(abs(x[:, None] - x[None, :])) (or the sorted-sample shortcut), average over origins and horizons, and compare it with the MAE of the median or mean.
"CRPS is the area between the predictive CDF and the outcome step."
Almost: it is the area of the squared gap. The plain area between the curves is $E\lvert X - y\rvert$, which is larger. Say "the integrated squared difference between the forecast CDF and the outcome's step function" if you are asked to be precise.
"CRPS of a probabilistic forecast can never beat the absolute error of its own centre."
It often does. A well-chosen spread lowers the CRPS below $\lvert y - \text{centre}\rvert$ (7.48 against 12 in the example). Only when the spread shrinks to zero do the two coincide.
"CRPS is only for Normal forecasts."
The sample formula works for any forecast you can sample from, including Student-t and Negative Binomial posterior predictive draws. The Normal closed form is only a shortcut.
"I can compare CRPS values between series on different scales."
CRPS is in the units of the data, like MAE. Scale it (divide by a naive CRPS or by the mean level) before averaging across series of different size.
CRPS is the natural headline score for your forecasting model because you can compute it directly from the posterior predictive samples of each forecast day, whatever the likelihood (Normal, Student-t or Negative Binomial). Average it over rolling origins and horizons, and compare it with the CRPS of the same samples collapsed to their median (which is the absolute error, i.e. the MAE) to see how much the uncertainty adds. To compare against a baseline such as seasonal naive, give the baseline an honest spread too (for example its past errors) so both are probabilistic forecasts, and report the CRPS ratio (a "CRPS skill score").
"CRPS is a kind of RMSE for distributions."
"CRPS generalizes the absolute error: for a point forecast it equals $\lvert y - f\rvert$."
Model answer: "CRPS compares the forecast CDF with the step function of the realised outcome, integrating the squared difference. It has the units of the data, it is proper, it reduces to the absolute error for a deterministic forecast, and it can be computed from samples as $E\lvert X - y\rvert - \tfrac12 E\lvert X - X'\rvert$. It also equals twice the average pinball loss over all quantile levels. I use it because it rewards both calibration and sharpness and is less tail-sensitive than the log score."
$\text{CRPS}(F, y) = \int (F(x) - \mathbf 1\{x \ge y\})^2 dx = E\lvert X - y\rvert - \tfrac12 E\lvert X - X'\rvert$. Units of $y$; lower is better.
Point forecast → $\lvert y - f\rvert$ (MAE). Normal: $\sigma[z(2\Phi(z)-1) + 2\varphi(z) - 1/\sqrt\pi]$. Samples: no density needed.
Trap: the plain area between CDF and step is $E\lvert X-y\rvert$, not the CRPS; scale before comparing series of different size.
Quick check: a forecast is the three samples $\{4, 6, 8\}$ and the outcome is 6. What is the CRPS?
$E\lvert X - 6\rvert = (2 + 0 + 2)/3 = 4/3$. The pairwise differences over the 9 ordered pairs sum to $2(2 + 4 + 2) = 16$, so $E\lvert X - X'\rvert = 16/9$ and half of it is $8/9$. CRPS $= 4/3 - 8/9 = 12/9 - 8/9 = 4/9 \approx 0.444$. (The point forecast 6 would have scored 0, but the spread {4, 6, 8} is a statement of doubt, and it costs a little when the doubt was not needed.)
Pinball loss: grading a quantile forecast core
You order stock for tomorrow. If you run out, each missing unit costs you 9 (lost sale). If you have too much, each spare unit costs 1 (storage). The best order is not the average demand. You should order enough that you run out only about 1 day in 10 (cost ratio $1 : 9$ means a $90\%$ service level): the 90% quantile of demand, the value that demand stays below 90% of the time.
The pinball loss (also called the quantile loss) is the score whose cost has exactly this lopsided shape. For a quantile level $\tau$ (tau, between 0 and 1) it charges $\tau$ per unit when the outcome is above your forecast (you under-forecast) and $1-\tau$ per unit when it is below (you over-forecast). The name comes from the shape of the graph, which looks like the tilted V of a pinball flipper. Averaged over many days, the forecast that minimizes it is the true $\tau$-quantile.
Three ways to say it:
- Picture: a V-shaped cost curve tilted so one side is steep and the other shallow; $\tau$ sets the tilt.
- Numbers: $\tau = 0.9$, forecast $q = 120$: outcome 140 costs $0.9 \times 20 = 18$; outcome 100 costs $0.1 \times 20 = 2$.
- Slogan: pinball loss turns "be right 90% of the time" into a score.
Orders in five hours: $2, 3, 3, 4, 38$. You must announce a number $q$ that should be exceeded only 10% of the time ($\tau = 0.9$). Compare $q = 4$ and $q = 38$.
- The loss for each outcome $y$: if $y \ge q$ it is $0.9\,(y - q)$; if $y \lt q$ it is $0.1\,(q - y)$.
- For $q = 4$: outcomes $2, 3, 3$ are below, costing $0.1\times2, 0.1\times1, 0.1\times1 = 0.2, 0.1, 0.1$; outcome 4 costs $0$; outcome 38 is above: $0.9 \times 34 = 30.6$. Total $31.0$, average $6.2$.
- For $q = 38$: all four small outcomes are below: $0.1\times(36 + 35 + 35 + 34) = 14.0$; outcome 38 costs $0$. Average $14/5 = 2.8$.
- So the quantile forecast $q = 38$ is much better for $\tau = 0.9$: it avoids the big under-forecast penalty on the one giant hour. With only five points the 90% quantile is the largest value; with more data it settles.
- Newsvendor link: with under-stock cost $c_u = 9$ and over-stock cost $c_o = 1$, the best order is the quantile $\tau^\ast = c_u/(c_u + c_o) = 9/10 = 0.9$.
For a quantile forecast $q$ at level $\tau$ and outcome $y$:
$$\rho_\tau(y, q) = \begin{cases} \tau\,(y - q) & \text{if } y \ge q, \\ (1-\tau)\,(q - y) & \text{if } y \lt q, \end{cases} \qquad\text{equivalently}\qquad \rho_\tau(y, q) = \big(\tau - \mathbf 1\{y \lt q\}\big)(y - q).$$- $\tau = 0.5$ gives $\tfrac12\lvert y - q\rvert$: half the absolute error. The median is the 0.5-quantile.
- It is minimized in expectation by the true $\tau$-quantile (proper for quantile forecasts). So a low $\tau$ rewards low forecasts and a high $\tau$ rewards high ones, each at the right level.
- A central $(1-\alpha)$ interval is the pair of quantiles $\tau = \alpha/2$ and $1-\alpha/2$, and the interval score of Section 4 equals $\tfrac2\alpha$ times the sum of those two pinball losses.
- Averaging the pinball loss over a grid of levels (for example 0.1, 0.2, ..., 0.9) gives a weighted quantile loss; averaging over all levels gives half the CRPS: $\text{CRPS} = 2\int_0^1 \rho_\tau(y, F^{-1}(\tau))\,d\tau$.
- Quantile forecasts are checked in two ways: pinball loss (accuracy) and the share of outcomes at or below $q_\tau$ (should be about $\tau$; this is calibration at one level).
Why do we need it?
Many decisions are about one quantile, not the mean: stock for 95% of days, capacity so overload happens 1 day in 20, a safe delivery date. The pinball loss grades exactly that and also trains models that output quantiles directly.
Where is it used?
Inventory and capacity planning, quantile regression and gradient boosting with a quantile objective, deep forecasters that output quantiles (such as the Temporal Fusion Transformer), the M5 uncertainty track, and Value at Risk in finance.
How is it used?
Pick the levels your decisions need (0.1, 0.5, 0.9, or 0.95). Take the forecast quantile at each level (np.quantile(samples, tau)), compute the pinball loss, average over origins and horizons, and also report the share of outcomes below each quantile.
"For the 90% quantile I just take the mean plus 1.28 standard deviations."
That is right only for a Normal forecast. For skewed demand (Negative Binomial, Log-Normal) take the 90% quantile of the actual predictive distribution, for example np.quantile(samples, 0.9).
"The 90% quantile forecast should be exceeded 90% of the time."
The outcome should be at or below the 90% quantile 90% of the time, so it is exceeded 10% of the time.
"A low pinball loss at one level means the whole distribution is good."
Each level is graded separately. Average over a grid of levels, or use CRPS, to judge the whole distribution, and check that quantile forecasts do not cross (the 90% quantile must not fall below the 50% one).
Capacity questions about your forecast ("how much stock or compute covers 95% of days?", "what is the chance demand exceeds 170?") are quantile questions on the posterior predictive samples (7.14): the answer is np.quantile(samples, 0.95) for each future day. Grade those quantiles with the pinball loss at the same $\tau$ and check the share of outcomes below them; this is the right test for a decision based on a quantile. In the A/B framework, a decision threshold on a posterior ("ship if the lift is above $\delta$ with 95% probability") is also a quantile statement about the posterior.
$\rho_\tau(y, q) = \tau(y-q)$ if $y \ge q$, else $(1-\tau)(q-y)$. Best $q$ = the true $\tau$-quantile. $\tau = 0.5$: half the absolute error.
Newsvendor: $\tau^\ast = c_u/(c_u + c_o)$. Average over levels → CRPS/2. Interval score = $\frac2\alpha\,(\rho_{\alpha/2} + \rho_{1-\alpha/2})$.
Trap: use the predictive quantile, not mean + 1.28 sd, for skewed data; quantiles must not cross.
Quick check: $\tau = 0.8$, forecast $q = 50$. What is the loss if the outcome is 60, and if it is 30?
Outcome 60 is above: $0.8\times(60-50) = 8$. Outcome 30 is below: $(1 - 0.8)\times(50-30) = 0.2\times20 = 4$. Under-forecasting is four times as costly per unit as over-forecasting at $\tau = 0.8$.
Proper scoring rules: why these scores, and why not RMSE or coverage alone? core
A scoring rule is a game. The forecaster announces a distribution, the outcome arrives, and a score is paid. A good game must make honesty the best strategy: if you truly believe tomorrow is $N(0, 1)$, announcing exactly that should give you the best expected score. If instead the game rewarded announcing something wider "to be safe" or narrower "to look sharp", forecasters would be pushed to lie. A score where telling the truth is optimal is called proper.
Two popular summaries fail this test for distributions. The RMSE of the mean ignores the spread, so every announced spread gets the same mark. Coverage alone rewards width: announce an enormous interval and you "cover" everything. The log score, the CRPS, the pinball loss and the interval score are proper; they are the tools to compare Bayesian forecasts with.
Three ways to say it:
- Picture: plot the expected score against the spread you announce; for a proper score the curve has its bottom exactly at the true spread.
- Numbers: truth $N(0,1)$: announcing spread 0.5, 1, 2 gives expected CRPS $0.610, 0.564, 0.656$ (best at 1), while the RMSE of the mean is $1.00$ for all three.
- Slogan: a proper score cannot be gamed; telling the truth is the best policy.
The true outcomes are $N(0,1)$. A forecaster announces $N(0, s^2)$ with $s = 0.5$ (overconfident), $1$ (honest) or $2$ (too vague). Expected values over many days (closed forms, checked by simulation):
| announced $s$ | RMSE of the mean | 80% interval: coverage | 80% interval: width | expected CRPS | expected log score |
|---|---|---|---|---|---|
| 0.5 | 1.00 | 47.8% | 1.28 | 0.610 | $-2.226$ |
| 1 (honest) | 1.00 | 80.0% | 2.56 | 0.564 | $-1.419$ |
| 2 | 1.00 | 99.0% | 5.13 | 0.656 | $-1.737$ |
- RMSE of the mean is $1.00$ in all three rows. It is the same number for all spreads: it can neither detect the overconfident forecaster nor reward the honest one.
- Coverage alone says "the more the better": $99\%$ looks better than $80\%$ if you only ask for at least 80%. A forecaster can always reach full coverage by widening.
- CRPS and log score are both best at $s = 1$, the true spread, and worse on either side. The honest spread is the unique best announcement.
- The log score is lopsided: being too narrow by half costs $2.226 - 1.419 = 0.81$ nats, being too wide by a factor 2 costs only $1.737 - 1.419 = 0.32$. Overconfidence is punished more than caution.
- You can check the closed forms: for a Normal forecast of spread $s$ against truth spread 1, expected CRPS $= \sqrt{2/\pi}\sqrt{s^2 + 1} - s/\sqrt\pi$ (at $s = 1$ it is $1/\sqrt\pi = 0.564$), and the expected negative log score is $\tfrac12\log(2\pi s^2) + 1/(2s^2)$.
A scoring rule $\text{S}(F, y)$ assigns a number (a loss here: lower is better) to a forecast distribution $F$ and an outcome $y$. If the outcome really comes from a distribution $G$, the forecaster's expected score is $E_{Y\sim G}\,\text{S}(F, Y)$.
- The rule is proper if this expected score is smallest when $F = G$ (reporting your true belief cannot be beaten), and strictly proper if $F = G$ is the only best report.
- Strictly proper: the log score; the CRPS; the Brier score for a yes/no event, $(p - \mathbf 1\{\text{event}\})^2$ (for example "demand exceeds capacity"). Proper for one summary: the pinball loss at level $\tau$ (best at the $\tau$-quantile); the interval score (best at the true pair of quantiles); the squared error (best at the mean) and the absolute error (best at the median, Chapter 7.15).
- Not proper for a distribution: the RMSE or MAE of one number taken from it (blind to the spread); coverage alone; interval width alone; MAPE (it favours low forecasts).
- Relationships (figure below): pinball at $\tau = 0.5$ is half the absolute error; the interval score is $\frac2\alpha$ times two pinball losses; the CRPS is twice the average pinball loss over all $\tau$; the CRPS of a point forecast is the absolute error.
Why do we need it?
If you rank models by a score that can be gamed, you will pick the model that games it. Proper scores guarantee that the model whose forecasts are closest to the truth has the best expected score, so model choice and honesty point the same way.
Where is it used?
Weather forecasting (the field that developed this theory), probabilistic forecasting competitions, Bayesian model comparison (the log score is the ELPD), and the evaluation of any classifier that outputs probabilities (log loss and Brier score are the same idea).
How is it used?
Choose the score from the question: whole distribution → CRPS or log score; a quantile decision → pinball at that level; an interval → interval score with coverage; an event probability → Brier score. Report one proper score as the headline and calibration diagnostics (coverage, PIT) beside it.
"RMSE is a proper scoring rule, so it is enough for probabilistic forecasts."
The squared error is proper for the mean (a point forecast). Applied to the mean of a distribution it ignores everything else about the distribution, so it cannot rank the spread, the tails or the intervals.
"If my 90% interval covers at least 90%, the model passes."
An interval that is far too wide also "covers at least 90%". The requirement is coverage close to 90% with the narrowest width, which is what the interval score and CRPS measure.
"Proper means the forecast is good."
Proper is a property of the score, not of the forecast. It guarantees that honesty is optimal in expectation; a proper score still ranks a poor-quality honest forecast below a better one.
"Log score and CRPS always agree on the ranking."
Both are proper, but they weigh the tails differently. The log score punishes a single tail surprise very hard; CRPS punishes it linearly. Rankings can differ when forecasts differ mainly in the tails; look at both.
When you compare variants of your forecasting model (Normal against Student-t against Negative Binomial likelihood; a mean-field against a full-rank guide; different Laplace scales for the changepoint slopes), report a proper score as the headline: the mean CRPS, with the log score as a second opinion, over the same rolling origins (Chapter 7.15). A change that improves RMSE but worsens CRPS or the log score has made the centre a little better and the uncertainty worse, which usually matters more for decisions. For the A/B framework, a posterior probability of a binary event (such as $P(\theta_B \gt \theta_A \mid D)$ calibrated against actual future results) is scored by the Brier score or the log loss.
"We chose the model with the narrowest intervals, since sharper is better."
"We chose by a proper scoring rule, because it cannot be gamed by making intervals too narrow or too wide."
Model answer: "A scoring rule is proper if the forecaster's expected score is best when they report their true belief. The log score and the CRPS are strictly proper, the pinball loss is proper for quantiles and the interval score for interval endpoints. RMSE of the mean ignores the spread, and coverage alone can be bought with width. I report a proper score such as CRPS next to coverage and the PIT, so I measure both calibration and sharpness."
Proper score: expected score is best when you report the true distribution (strictly: only then). Strictly proper: log score, CRPS, Brier. Proper for quantiles: pinball. Intervals: interval score.
pinball(τ=½) = ½|e| · interval score = (2/α)(ρ$_{α/2}$ + ρ$_{1-α/2}$) · CRPS = 2 avg ρ$_\tau$ · CRPS(point) = |e|.
Trap: RMSE of the mean ignores the spread; coverage alone rewards width; proper ≠ good forecast.
Quick check: a forecaster discovers that announcing a wider spread than they believe lowers their RMSE of the mean. Is that possible, and what does it say about RMSE?
No: the RMSE of the mean does not depend on the announced spread at all. The forecaster cannot gain or lose by changing it, which is exactly why RMSE gives no incentive to be honest about uncertainty. A proper score (CRPS, log score) does depend on it, with the best value at the true spread.
Evaluating over horizons and origins: where intervals quietly fail core
The further ahead you forecast, the less you know, so a good forecast should widen as the horizon grows (Chapter 7.14). The most common failure of a home-made interval is a band of constant width: it is right for tomorrow and increasingly wrong for next week. An average coverage over all horizons can hide that: it may look acceptable while the 14-days-ahead intervals are catastrophically too narrow.
So probabilistic evaluation uses the same machinery as Chapter 7.15: many rolling origins, a forecast distribution stored for every origin and every horizon, and every score computed per horizon (coverage at $h = 1, 2, \dots$; CRPS at each $h$) and then averaged over origins. Slicing by weekday or by holiday against normal days finds the next layer of hidden failures.
Three ways to say it:
- Picture: coverage on the vertical axis, horizon on the horizontal axis; a calibrated forecast is a flat line at the nominal level.
- Numbers: an 80% band that does not widen covers 80% at $h = 1$, 48% at $h = 4$, 33% at $h = 9$ and 25% at $h = 16$ for a random walk.
- Slogan: check calibration at every horizon, not on average.
A random walk with step standard deviation $\sigma = 1$, forecast by "tomorrow = today". A forecaster draws an 80% interval of half-width $z\sigma = 1.2816$ at every horizon (it never widens). The true error at horizon $h$ is Normal with standard deviation $\sqrt h$.
- $h = 1$: error sd $1$; coverage $= P(\lvert N(0,1)\rvert \le 1.2816) = 80\%$.
- $h = 4$: error sd $2$; coverage $= P(\lvert Z\rvert \le 1.2816/2 = 0.64) = 2\Phi(0.64) - 1 \approx 47.8\%$.
- $h = 9$: error sd $3$; $2\Phi(0.427) - 1 \approx 33.1\%$. $h = 16$: error sd $4$; $2\Phi(0.320) - 1 \approx 25.1\%$.
- An honest interval has half-width $z\sigma\sqrt h$ and covers 80% at every $h$. The average coverage of the constant band over $h = 1, \dots, 16$ is about $39\%$: a single number that already signals trouble, but would hide where the trouble is.
- Python check of the same effect on 147 rolling origins of a simulated walk (Code-it block): constant band $0.77, 0.44, 0.34, 0.22$ at $h = 1, 4, 9, 14$; widening band $0.77, 0.74, 0.79, 0.79$.
A probabilistic rolling-origin evaluation: at each origin $T_j$ and horizon $h = 1, \dots, H$ store the forecast distribution $F_{T_j+h\mid T_j}$ (as samples or quantiles) and the outcome. Then compute, for each $h$ and averaging over origins $j = 1, \dots, K$:
$$\text{coverage}_h = \frac1K\sum_j \mathbf 1\{l \le y \le u\}, \qquad \text{CRPS}_h = \frac1K\sum_j \text{CRPS}\big(F_{T_j+h\mid T_j},\, y_{T_j+h}\big),$$and likewise the log score, the pinball loss at the levels you care about, and the PIT histogram per horizon group.
- Judge each $\text{coverage}_h$ against its luck range $\pm 1.96\sqrt{p(1-p)/K}$. With few origins the curve is noisy (and neighbouring origins overlap, so the noise is larger than the formula says).
- CRPS skill score against a baseline (for example seasonal naive with an honest spread): $1 - \text{CRPS}_{\text{model}}/\text{CRPS}_{\text{baseline}}$, per horizon. Positive means the model's whole distribution beats the baseline's.
- Slice also by weekday, by holiday and normal day, by level (low / high demand) and by regime (before / after a changepoint): calibration often fails in a slice while the overall average looks fine.
- Use a step that is not a multiple of the season length, for the same reason as in Chapter 7.15.
Why do we need it?
Uncertainty is not constant: it grows with the horizon and differs by day type. A forecast can be calibrated on average yet dangerously wrong at the horizons and days that matter for a decision, so the checks must be sliced the way decisions are.
Where is it used?
Prophet's cross-validation reports coverage per horizon; weather and energy verification reports scores by lead time; central banks and demand planners publish fan charts whose calibration is checked by horizon.
How is it used?
Store samples for every (origin, horizon). Compute coverage, CRPS and log score per horizon, plot them with luck ranges, compare with a baseline's CRPS, then repeat for slices (weekday, holiday, regime). Fix the model where a slice fails.
"Average coverage over all horizons is 78%, close enough to 80%."
An average can hide opposite failures at different horizons (too narrow far ahead, too wide near). Report coverage per horizon, with its luck range.
"Intervals that do not widen with the horizon are simpler and fine."
Forecast error grows with the horizon (for a random walk like $\sqrt h$). A constant-width band is calibrated only at the horizon where it was measured.
"If the model is calibrated overall it is calibrated on holidays too."
Check slices. A model without proper holiday columns can be perfectly calibrated on normal days and badly overconfident on holidays, and holidays are often the days that matter.
"I can use the one-step residual sd to build the whole interval."
Multi-step errors also include the accumulation of trend and level uncertainty (and parameter uncertainty). Use the model's own multi-step predictive distribution, or estimate the spread per horizon from rolling-origin errors.
Evaluate your forecasting model's posterior predictive samples per horizon over rolling origins: coverage at 50%, 80% and 95%, CRPS and log score, each with its luck range, and a CRPS skill score against seasonal naive with an empirical spread. Then slice by weekday and holiday (do the holiday columns of 7.12 keep the intervals calibrated?), by regime (just after a changepoint, when the trend is least certain), and by level (a Negative Binomial's variance grows with the mean, so check low and high demand days separately). Under-coverage that grows with the horizon points at missing trend uncertainty (for example from an over-confident approximate posterior, Chapter 6.13), over-coverage at short horizons at a noise scale that is too large.
Per horizon $h$, averaged over origins: $\text{coverage}_h$, $\text{CRPS}_h$, log score$_h$; luck range $\pm1.96\sqrt{p(1-p)/K}$.
Random walk + naive: honest interval $\pm z\sigma\sqrt h$; a constant band covers $2\Phi(z/\sqrt h) - 1$.
Trap: average coverage hides horizon and slice failures; use the model's multi-step distribution.
Quick check: an 80% interval of constant half-width $1.2816$ for a random walk with step sd 1. What coverage do you expect at $h = 25$?
The error sd is $\sqrt{25} = 5$, so coverage $= 2\Phi(1.2816/5) - 1 = 2\Phi(0.256) - 1 \approx 0.20$. Only about 20% of outcomes land inside an interval that claims 80%.
Recap, cheat sheet and practice
- Why: a probabilistic forecast has two jobs, a good centre and honest doubt. RMSE of the mean grades only the centre and is blind to the spread; "RMSE improved" says nothing about the intervals.
- Coverage: the share of outcomes inside the $(1-\alpha)$ interval should be close to $1-\alpha$; judge it against its luck range $\pm1.96\sqrt{p(1-p)/n}$. Below nominal = overconfident, above = too cautious. Use predictive (not trend-only) intervals; check by horizon and slice.
- PIT: $u_t = F_t(y_t)$ is uniform for an honest forecast. Flat = calibrated, U = overconfident, hump = underconfident, slope = biased. Randomized PIT for counts. Necessary, not sufficient.
- Sharpness: narrow intervals are good only if calibrated: maximize sharpness subject to calibration. Interval score $= $ width $+ \frac2\alpha\times$ miss distance combines the two.
- Log score $\log p_t(y_t)$ (in nats, higher is better): rewards probability on the outcome, punishes overconfidence and tail surprises; from posterior draws use the log of the average likelihood.
- CRPS $= \int (F - \mathbf 1\{x\ge y\})^2dx = E\lvert X-y\rvert - \frac12E\lvert X-X'\rvert$: units of the data, = absolute error for a point forecast, computable from samples.
- Pinball loss $\rho_\tau$ grades a $\tau$-quantile; minimized by the true quantile; $\tau = \frac12$ gives half the absolute error; CRPS $= 2\times$ the average over $\tau$.
- Proper scores (log score, CRPS, Brier, pinball for quantiles, interval score for interval ends) reward honesty; coverage alone and RMSE of the mean do not. Evaluate per horizon and origin, then slice.
Cheat sheet
| Tool | Formula or rule | Reading / trap |
|---|---|---|
| Coverage | $\frac1n\sum\mathbf 1\{l_t \le y_t \le u_t\}$; luck range $\pm1.96\sqrt{p(1-p)/n}$ | below nominal = overconfident; very wide intervals cover everything |
| PIT value | $u_t = F_t(y_t)\approx$ share of samples $\le y_t$ | flat / U / hump / slope; counts: randomized |
| Interval score | $(u-l) + \frac2\alpha\max(l-y, 0) + \frac2\alpha\max(y-u, 0)$ | width + miss penalty; minimized by the honest interval |
| Log score | $\log p_t(y_t)$; draws: $\text{logsumexp}_s \log p(y\mid\theta_s) - \log S$ | higher is better; tail-sensitive; not mean of logs |
| CRPS | $E\lvert X-y\rvert - \frac12E\lvert X-X'\rvert$; Normal: $\sigma[z(2\Phi(z)-1)+2\varphi(z)-1/\sqrt\pi]$ | units of $y$; point forecast → $\lvert y-f\rvert$ |
| Pinball loss | $\tau(y-q)$ if $y\ge q$, else $(1-\tau)(q-y)$ | best $q$ = $\tau$-quantile; newsvendor $\tau^\ast = c_u/(c_u+c_o)$ |
| Brier score | $(p - \mathbf 1\{\text{event}\})^2$ | for "demand exceeds capacity" probabilities |
| Skill score | $1 - \text{score}_{\text{model}}/\text{score}_{\text{baseline}}$ | same origins and horizons; honest baseline spread |
| Constant band, random walk | coverage at $h$ $= 2\Phi(z/\sqrt h)-1$ | 80% → 48% at $h=4$, 33% at $h=9$ |
import warnings; warnings.filterwarnings("ignore")
import numpy as np
import pandas as pd
from scipy import stats
from statsmodels.tsa.exponential_smoothing.ets import ETSModel
rng = np.random.default_rng(12)
# ---------------------------------------------------------------------------------------------
# Part 1. Four forecasters for the same outcomes. Each one gives S = 2000 forecast samples per day
# (just like posterior predictive draws). The true noise sd is 4.
# ---------------------------------------------------------------------------------------------
n, S, sd_true = 400, 2000, 4.0
m = 100 + 10 * np.sin(2 * np.pi * np.arange(n) / 7) # the predictable part of demand
y = m + rng.normal(0, sd_true, n) # the outcomes
forecasters = { # (centre shift, spread of the forecast)
"honest": (0.0, 4.0),
"overconfident": (0.0, 2.0),
"underconfident": (0.0, 8.0),
"biased low": (-3.0, 4.0),
}
def crps_samples(y, xs): # CRPS = E|X - y| - 0.5 E|X - X'|, exact for the sample
x = np.sort(xs); k = len(x)
return np.mean(np.abs(x - y)) - np.sum((2 * np.arange(k) - k + 1) * x) / k**2
def crps_normal(y, mu, s): # closed form for a Normal forecast
z = (y - mu) / s
return s * (z * (2 * stats.norm.cdf(z) - 1) + 2 * stats.norm.pdf(z) - 1 / np.sqrt(np.pi))
def pinball(y, q, tau):
return np.mean(np.where(y >= q, tau * (y - q), (1 - tau) * (q - y)))
print(f"{'forecaster':15s} {'RMSE':>5s} {'cov80':>6s} {'width':>6s} {'CRPS':>5s} {'logS':>6s} {'pin.9':>6s} PIT histogram (10 bins)")
for name, (shift, spread) in forecasters.items():
X = rng.normal(m[:, None] + shift, spread, (n, S)) # n x S forecast samples
mean = X.mean(1)
lo, hi = np.quantile(X, [0.10, 0.90], axis=1) # central 80% interval from the samples
cov = np.mean((y >= lo) & (y <= hi))
crps = np.mean([crps_samples(y[i], X[i]) for i in range(n)])
logs = np.mean(stats.norm.logpdf(y, mean, X.std(1))) # Gaussian fit to the samples
pit = (X < y[:, None]).mean(1) # PIT = predictive CDF at the outcome
hist = np.histogram(pit, bins=10, range=(0, 1))[0]
q90 = np.quantile(X, 0.9, axis=1)
print(f"{name:15s} {np.sqrt(np.mean((y - mean)**2)):5.2f} {cov:6.2f} {np.mean(hi - lo):6.2f} {crps:5.2f} {logs:6.2f} "
f"{pinball(y, q90, 0.9):6.2f} {hist}")
# forecaster RMSE cov80 width CRPS logS pin.9 PIT histogram (10 bins)
# honest 3.97 0.80 10.24 2.24 -2.80 0.66 [44 28 40 36 39 44 37 48 47 37] flat (wobble is normal)
# overconfident 3.97 0.50 5.12 2.41 -3.59 0.85 [ 88 32 27 21 23 16 27 23 31 112] U-shape: too narrow
# underconfident 3.99 0.99 20.48 2.62 -3.12 1.02 [ 2 22 30 54 79 82 76 39 15 1] hump: too wide
# biased low 5.05 0.65 10.24 2.95 -3.10 0.94 [ 12 16 16 15 19 38 38 44 73 129] outcomes pile up near 1
# (RMSE of the mean is IDENTICAL for the first three: it cannot see the spread at all.)
# Closed-form check: for the honest Normal forecaster the average CRPS should be sd * 1/sqrt(pi) = 4 * 0.5642
print("theory CRPS of the honest forecaster:", round(sd_true / np.sqrt(np.pi), 3))
# theory CRPS of the honest forecaster: 2.257 (the sample mean 2.24 agrees)
# ---------------------------------------------------------------------------------------------
# Part 2. Score a real model: ETS (Holt-Winters as a statistical model, Chapter 7.5) on rolling origins
# ---------------------------------------------------------------------------------------------
t = np.arange(210)
weekly = np.array([0, 3, 2, 4, 12, 24, 18])
series = pd.Series(40 + 0.12 * t + weekly[t % 7] + rng.normal(0, 3, 210), index=pd.date_range("2024-01-01", periods=210, freq="D"))
z80 = stats.norm.ppf(0.9)
cover, crps_l, logs_l, width_l = [], [], [], []
for T in range(70, 210 - 7 + 1, 7): # 20 origins, horizon 7, expanding window
fit = ETSModel(series.iloc[:T], error="add", trend="add", seasonal="add", seasonal_periods=7).fit(disp=False)
pf = fit.get_prediction(start=series.index[T], end=series.index[T + 6]).summary_frame(alpha=0.2) # 80% prediction interval
mu = pf["mean"].to_numpy(); s = (pf["pi_upper"] - pf["pi_lower"]).to_numpy() / (2 * z80)
a = series.iloc[T:T + 7].to_numpy()
cover += list((a >= pf["pi_lower"].to_numpy()) & (a <= pf["pi_upper"].to_numpy()))
crps_l += list(crps_normal(a, mu, s)); logs_l += list(stats.norm.logpdf(a, mu, s)); width_l += list(2 * z80 * s)
print(f"\nETS(A,A,A), 20 origins x 7 days: coverage of the 80% interval = {np.mean(cover):.2f}, mean width = {np.mean(width_l):.1f}, "
f"mean CRPS = {np.mean(crps_l):.2f}, mean log score = {np.mean(logs_l):.2f}")
# ETS(A,A,A), 20 origins x 7 days: coverage of the 80% interval = 0.77, mean width = 6.8, mean CRPS = 1.52, mean log score = -2.40
# (140 forecasts: the standard error of a coverage near 0.8 is sqrt(0.8*0.2/140) = 0.034, so 0.77 is consistent with 0.80)
# ---------------------------------------------------------------------------------------------
# Part 3. Coverage by horizon: a random walk, naive forecasts, intervals that widen (honest) or do not (constant)
# ---------------------------------------------------------------------------------------------
rw = np.cumsum(rng.normal(0, 1, 3000))
H, hs = 14, np.arange(1, 15)
ok_h, ok_c = [], []
for T in range(60, 3000 - H + 1, 20): # 147 origins
sh = np.std(np.diff(rw[:T]), ddof=1) # one-step sd estimated from the past only
err = np.abs(rw[T:T + H] - rw[T - 1]) # naive forecast = last value
ok_h.append(err <= z80 * sh * np.sqrt(hs)); ok_c.append(err <= z80 * sh)
print("\nhorizon h ", " ".join(f"{h:5d}" for h in (1, 4, 9, 14)))
print("honest (sd*sqrt h)", " ".join(f"{v:5.2f}" for v in np.mean(ok_h, 0)[[0, 3, 8, 13]]))
print("constant (sd) ", " ".join(f"{v:5.2f}" for v in np.mean(ok_c, 0)[[0, 3, 8, 13]]), " theory:",
" ".join(f"{2 * stats.norm.cdf(z80 / np.sqrt(h)) - 1:5.2f}" for h in (1, 4, 9, 14)))
# horizon h 1 4 9 14
# honest (sd*sqrt h) 0.77 0.74 0.79 0.79 <- widening intervals stay near 0.80 at every horizon
# constant (sd) 0.77 0.44 0.34 0.22 theory: 0.80 0.48 0.33 0.27 <- a constant-width band collapses with the horizon
Output checked with numpy 2.5, SciPy 1.18 and statsmodels 0.15. Part 1 builds forecast samples by hand so every score can be traced; Part 2 scores a real statsmodels model (the ETS intervals are Normal, so its CRPS and log score use the Normal formulas with the standard deviation recovered from the interval); Part 3 shows the horizon effect. Sampling noise: a coverage near 0.8 from 140 forecasts has a standard error of about 0.034.
1. When is a 90% prediction interval calibrated?
2. A PIT histogram has tall bars in the first and last bins and low bars in the middle. What does this say about the forecasts?
3. Which of these is a proper scoring rule for a full predictive distribution?
4. A deterministic forecast says 100 and the outcome is 112. What is its CRPS?
5. Pinball loss with $\tau = 0.9$, forecast $q = 120$, outcome $100$. What is the loss?
6. You have $S$ posterior draws $\theta_s$ and want the log predictive density of an outcome $y$. What do you compute?
Practice problems
A. A model issues 90% intervals. Case 1: 40 days, 33 hits. Case 2: 400 days, 330 hits. For each, find the observed coverage, the luck range if the intervals were honest, and your verdict.
- Both cases: observed coverage $= 33/40 = 330/400 = 82.5\%$.
- Case 1: standard error $= \sqrt{0.9\times0.1/40} \approx 0.0474$; $1.96\times0.0474 \approx 0.093$; range $[0.807, 0.993]$. The observed $0.825$ is inside: no clear evidence against the intervals (40 days are too few).
- Case 2: standard error $= \sqrt{0.09/400} = 0.015$; $1.96\times0.015 = 0.029$; range $[0.871, 0.929]$. The observed $0.825$ is far below: the 90% intervals are overconfident (too narrow).
- Same observed share, different verdicts: the amount of data decides how much a gap means.
B. Forecasts are $N(50, 5^2)$ and the outcomes are $44, 50, 61, 38, 52$. Compute the five PIT values. Then a histogram of 100 PIT values has counts $19, 8, 6, 5, 5, 4, 5, 6, 8, 34$ (10 bins). Diagnose.
- $z = (y - 50)/5 = -1.2, 0, 2.2, -2.4, 0.4$.
- PIT $= \Phi(z) \approx 0.115,\ 0.500,\ 0.986,\ 0.008,\ 0.655$. Two of the five are in the extreme tenths ($0.008$ and $0.986$), where an honest forecast expects about one in five.
- Histogram: the two extreme bins hold $19 + 34 = 53$ of 100 (expected about 20): a strong U shape, so the forecasts are overconfident.
- The last bin ($34$) is much taller than the first ($19$): more outcomes land in the top of the forecast than in the bottom, so the forecasts also tend to be a little too low. Remedy: widen the predictive distribution (more noise or parameter uncertainty, or a heavier-tailed likelihood) and correct the level.
C. Forecast $N(20, 4^2)$, outcome $26$. Compute the log score and the CRPS, and compare the CRPS with the absolute error of the centre.
- $z = (26 - 20)/4 = 1.5$.
- Log score $= -\tfrac12\log(2\pi\cdot16) - 1.5^2/2 = -2.305 - 1.125 = -3.430$.
- CRPS: $\Phi(1.5) = 0.9332$ so $2\Phi - 1 = 0.8664$ and $z\times0.8664 = 1.2996$; $\varphi(1.5) = 0.1295$ so $2\varphi = 0.2590$; the sum $1.2996 + 0.2590 - 0.5642 = 0.9944$; times $\sigma = 4$: CRPS $\approx 3.98$.
- The absolute error of the centre is $\lvert 26 - 20\rvert = 6$. The CRPS (3.98) is smaller because the forecast admitted a spread of 4 and the outcome was within 1.5 spreads of the centre.
D. Selling a product earns 6 per unit sold and unsold stock costs 2 per unit. Demand is forecast $N(100, 20^2)$. What should you order? What is the pinball loss at that order for demands 130 and 90?
- Under-stock cost $c_u = 6$ (lost profit per missing unit), over-stock cost $c_o = 2$. The best order is the quantile $\tau^\ast = c_u/(c_u + c_o) = 6/8 = 0.75$.
- The 0.75 quantile of $N(100, 20^2)$ is $100 + 0.6745\times20 \approx 113.5$ units.
- Pinball loss at $q = 113.5$ with $\tau = 0.75$: demand 130 (short by 16.5): $0.75\times16.5 \approx 12.4$; demand 90 (23.5 spare): $0.25\times23.5 \approx 5.9$.
- These are exactly the costs per unit rescaled: under-stocking costs 3 times more than over-stocking ($0.75/0.25 = 3 = 6/2$).
E. (Interview) "Your new model has a lower RMSE but a worse CRPS than the old one. What could have happened, and what do you do?"
"RMSE only grades the centre; CRPS grades the whole distribution. A lower RMSE with a worse CRPS means the centre improved while the uncertainty got worse: typically the new model is overconfident (intervals too narrow, for example from an approximate posterior that understates parameter uncertainty or a Normal likelihood on heavy-tailed data) or too wide. I would check the coverage at 50%, 80% and 95% by horizon, draw the PIT histogram (U-shape means overconfident, hump means too wide), and compare the interval widths. The fix is to the uncertainty (noise model, tails, parameter uncertainty), not to the centre. I would only ship the new model if the CRPS and calibration are at least as good, because decisions use the intervals."
F. (Design) Describe how you would evaluate the probabilistic forecasts of your Prophet-style Bayesian model before trusting its intervals.
- Protocol: rolling origins (Chapter 7.15), horizon up to the planning horizon, refit everything inside each window, step not a multiple of 7; store posterior predictive samples for every origin and horizon.
- Calibration: coverage at 50%, 80% and 95% per horizon with luck ranges; PIT histogram at $h=1$ and for horizon groups; randomized PIT if the likelihood is Negative Binomial.
- Sharpness and proper scores: mean interval width; mean CRPS from the samples and the log score from the likelihood (log of the average over draws); a CRPS skill score against seasonal naive with an empirical spread.
- Decisions: pinball loss and the share of outcomes below the quantiles used for capacity (for example 0.95), and the Brier score of "demand exceeds capacity" probabilities.
- Slices: weekday, holidays versus normal days, just after a changepoint, low versus high demand. Fix the model where a slice fails, and compare variants (Normal, Student-t, Negative Binomial; different guides) by a proper score, not by RMSE alone.
Residual diagnostics
Your forecasting model tells a story about the past: a trend, a weekly rhythm, some holidays, and then "the rest is random noise". The residuals are that "rest". If the story is complete, the leftovers really do look like random static. If it is not, the leftovers keep a shape, and the shape tells you what the model missed. This chapter teaches you to read it: the mean and spread, the histogram and Q-Q plot, the residual-vs-fitted and residual-vs-time plots, the ACF and the Ljung–Box test, what to do when the spread changes (heteroscedasticity), and what to do when the residuals remember yesterday.
- Define a residual $e_t = y_t - \hat y_t$ precisely, say which prediction $\hat y_t$ you used (in-sample or holdout; for a Bayesian model: posterior mean or posterior-predictive draws), and explain why a good model leaves no exploitable structure
- Run the five questions on any residual series: is it centred, is the spread constant, is the shape right, is there a pattern against fitted values and time, is there memory?
- Read the mean and variance, the histogram and KDE and the Q-Q plot of residuals, and turn what you see into a likelihood decision (Normal, Student-t, Negative Binomial)
- Recognise the classic patterns in residual vs fitted and residual vs time plots and say what each one means
- Use the residual ACF, the Ljung–Box test (and its lag and degrees-of-freedom choices) and the Durbin–Watson number
- Detect heteroscedasticity and respond in three ways: a transformation, a different likelihood (such as the Negative Binomial) or an explicit variance model
- Use Pearson and randomized quantile residuals for count models, and judge a residual plot against replicated residuals (the Bayesian way of not over-reading noise)
- Explain autocorrelated residuals: what the ACF shape says about the missing piece, and the four responses (an AR residual component, a state-space model, a Gaussian process, richer seasonal terms)
- Run a full diagnostic panel for five kinds of misspecification and say what to try next
What we need from earlier chapters: Q-Q plots and KDE (Chapter 4.17 and 4.16: we use them here and do not re-teach them); residuals of a regression (Chapter 5.13); the ACF, the white-noise band and Ljung–Box (Chapter 7.3); the likelihoods (Chapter 7.13); and posterior predictive checks (Chapter 6.8). Each time one returns you get a one-line reminder.
Residuals: what the model did not explain core
A tailor cuts a suit from a roll of cloth. The scraps on the floor are what is left over. If the suit fits well, the scraps are random little pieces. If you find a long strip shaped like a sleeve, the tailor forgot a sleeve. The scraps tell you what the suit is missing.
A forecasting model works the same way. It says: "demand = trend + weekly rhythm + holidays + regressors + random noise". After fitting, take what really happened and subtract what the model explains. What remains is the residual. If the model told the whole story, the residuals look like random static with no shape. If they have a shape (a wave, a staircase, a fan that opens up), the shape is a clue to what is missing.
Three ways to say it:
- Picture: the off-cuts on the tailor's floor. Random scraps mean a good fit; a recognisable strip means a missing piece.
- Numbers: the model said 120 orders, the shop got 130, so the residual is $+10$. A good model has residuals that are small and have no pattern.
- Slogan: a good model leaves nothing to learn from its leftovers.
One week of orders and the model's fitted values (day 1 is a Monday):
| Day | Mon | Tue | Wed | Thu | Fri | Sat | Sun |
|---|---|---|---|---|---|---|---|
| actual $y_t$ | 112 | 98 | 105 | 130 | 142 | 160 | 118 |
| model $\hat y_t$ | 108 | 100 | 110 | 128 | 146 | 150 | 120 |
- Residuals $e_t = y_t - \hat y_t$: $+4,\; -2,\; -5,\; +2,\; -4,\; +10,\; -2$.
- Mean residual: $(4 - 2 - 5 + 2 - 4 + 10 - 2)/7 = 3/7 \approx 0.43$ orders. That is close to 0, so the model does not lean high or low overall.
- Typical size: the squares are $16, 4, 25, 4, 16, 100, 4$ with sum $169$. The mean square is $169/7 \approx 24.1$ and its root is $\approx 4.9$ orders, the typical miss.
- Look at the signs: $+ - - + - + -$. They jump around; there is no long run of the same sign, so nothing obvious is left.
- The biggest miss is Saturday, $+10$. In units of the typical miss that is $10/4.9 \approx 2.0$: unusual, not shocking. If every Saturday were $+10$, the model would be missing a Saturday effect; that is what looking at more weeks is for.
Let $y_t$ be the actual value on day $t$ and $\hat y_t$ the value the model predicts for that day. The residual is
$$e_t = y_t - \hat y_t.$$The standardized residual is $z_t = e_t/\hat\sigma$, the residual measured in units of the typical miss $\hat\sigma$ (so $|z_t| \gt 3$ flags a surprising day).
- Residual is not the same as noise. The model says $y_t = \mu_t + \epsilon_t$, where $\epsilon_t$ is the true noise (never seen). Then $e_t = \epsilon_t + (\mu_t - \hat\mu_t)$: the noise plus whatever the model got wrong or missed. If the model is right and well estimated, $e_t \approx \epsilon_t$, so the residuals should look like the noise we assumed (independent, constant spread, the chosen shape). Any difference is a symptom.
- Which prediction $\hat y_t$? (a) In-sample: the fitted value on days the model was trained on (optimistic, because the model has seen those days). (b) Out-of-sample: a forecast error on days held back (the honest version, Chapter 7.15). (c) For a Bayesian model there are several ways to get $\hat y_t$: the posterior-mean prediction $\hat\mu_t = E[\mu_t \mid D]$ (average the mean function over posterior draws; this is what this chapter uses unless it says otherwise), or one residual per posterior draw $e_t^{(s)} = y_t - \mu_t^{(s)}$ (keeps parameter uncertainty), or posterior-predictive draws $y_t^{\text{rep},(s)}$ to compare the real residuals with residuals of simulated data (the lineup in the section on replicated residuals below).
- No exploitable structure. The goal is that nothing available when the forecast is made (the date, the fitted value, other columns, the previous residuals) can be used to predict $e_t$ better than "zero, give or take $\hat\sigma$".
That goal gives five questions to ask every residual series: (1) is it centred (mean about 0, in every group)? (2) is the spread constant? (3) is the shape what the likelihood assumes? (4) is there a pattern against the fitted values, time or other columns? (5) is there memory (autocorrelation)? The rest of the chapter takes them in turn.
Why do we need it?
The noise assumption (independent, constant spread, a chosen shape) is the part of a model that is checked least and breaks most quietly. Residuals are the only direct view of it. They show what the model missed, and a missed piece means wrong forecasts or intervals that are too narrow.
Where is it used?
Every regression and forecasting workflow: the residual panels of statsmodels (plot_diagnostics) and R's checkresiduals, model reviews before a forecast is trusted, and the checking step of a Bayesian workflow. In your forecasting model it is the step after fitting with SVI.
How is it used?
Fit the model, compute residuals on the training days and on rolling-origin holdouts, plot them against time, against the fitted values and as an ACF, look at their histogram and Q-Q plot, run Ljung–Box. If you see structure, fix the model (a missing component, another likelihood), not the plot.
"The residual is the noise term $\epsilon_t$ of the model."
The noise $\epsilon_t$ is an unseen part of the model. The residual $e_t$ is what you can compute, and it equals the noise plus every mistake the model made. If the model is right, the two look alike.
"The residuals average to 0, so the model is unbiased."
For least squares with an intercept the in-sample residuals sum to exactly 0 by construction, whatever the model. The mean is only informative on a holdout, or inside groups (Mondays, holidays, high-demand days).
"Small in-sample residuals mean good forecasts."
A flexible model can fit the training days almost perfectly and forecast badly. In-sample residuals are optimistic. Check residuals on rolling-origin holdouts too (Chapter 7.15).
In your forecasting model the mean is $\mu_t = g(t) + s(t) + h(t) + X_t\beta$, and each layer corresponds to something that can be missing from the residuals: too few changepoints (or too strong a Laplace shrinkage) leaves a slow wave or a staircase; too low a Fourier order leaves a seasonal wave; a holiday window that is too short leaves spikes on the holiday dates; a regressor you did not include leaves its own shape. After the SVI fit, take the mean function averaged over the guide's samples (or its posterior mean) as $\hat y_t$, compute the residuals, and use the five questions on them, on the training period and on rolling-origin holdouts.
"The residual is the error of the model."
"A residual is the observed value minus the model's prediction for that day. It is an estimate of the noise, but it also contains whatever the model missed. I state which prediction I used: fitted values on training days, or forecasts on held-out days."
Model answer: "I check residuals because the noise assumption is the weakest-checked part of the model. For a Bayesian model I compute them from the posterior-mean prediction for a quick look, and compare with residuals of data simulated from the model when I need to judge whether a pattern is real."
Residual $e_t = y_t - \hat y_t$; standardized $z_t = e_t/\hat\sigma$. Residual = true noise + what the model missed.
Goal: no exploitable structure. Five questions: centred? constant spread? right shape? pattern vs fitted/time? memory?
Say which $\hat y_t$: in-sample fit, holdout forecast, posterior mean or a posterior draw. Trap: with an intercept the in-sample mean is 0 by construction.
Quick check: a model's residuals on 5 days are $+3, -1, +4, -2, -4$. What is the mean, and what do you conclude about "leaning"?
Sum $= 3 - 1 + 4 - 2 - 4 = 0$, so the mean is 0. On this sample the model does not lean high or low overall. But the mean alone cannot reveal a pattern such as "all weekends are above the model and all weekdays below". That needs the residuals by group, which is the next section.
The first two numbers: residual mean and residual variance core
Think of a dart player. Over the whole evening the darts land, on average, right on the bull's-eye: some a bit left, some a bit right. Does that prove the player is accurate? No. Perhaps every dart in the morning went far left and every dart in the evening far right, and the two cancelled out. The overall average hides the problem; you must look at the morning and the evening separately.
The mean of the residuals says whether the model leans high or low. The variance (or standard deviation) of the residuals says how big the typical miss is. Both are quick, but both can hide trouble inside a group: Mondays, holidays, busy days.
Three ways to say it:
- Picture: a see-saw whose two ends are Monday and Sunday. It looks level overall (mean 0) while each end is far from the ground.
- Numbers: residuals $+20, +4, 0, -2, +6, -10, -18$ for Mon to Sun have mean $0$, yet Monday is 20 orders too low in the model and Sunday 18 too high.
- Slogan: the overall mean can be 0 while every group is wrong; always look at the mean by group.
A model that ignores the weekday predicts a flat 100 orders every day. The week really was Mon 120, Tue 104, Wed 100, Thu 98, Fri 106, Sat 90, Sun 82.
- Residuals $y - \hat y$: $+20,\; +4,\; 0,\; -2,\; +6,\; -10,\; -18$.
- Overall mean: $(20 + 4 + 0 - 2 + 6 - 10 - 18)/7 = 0/7 = 0$. It looks perfectly centred.
- Typical miss: squares $400, 16, 0, 4, 36, 100, 324$ have sum $880$; mean square $880/7 \approx 125.7$; root $\approx 11.2$ orders.
- By group: Monday $+20$, Sunday $-18$. Almost all of the 11.2 comes from the weekday pattern, not from random noise.
- How sure can we be that a group mean is real? With 10 such weeks and a daily spread of 8 orders, the standard error of one weekday's mean is about $8/\sqrt{10} \approx 2.5$. A Monday mean of $+20$ is $20/2.5 = 8$ standard errors away from 0: clearly real. A Wednesday mean of $+1$ would be only $0.4$ standard errors away: nothing to see.
For residuals $e_1, \dots, e_n$:
$$\bar e = \frac1n\sum_{t=1}^n e_t, \qquad s_e^2 = \frac{1}{n-p}\sum_{t=1}^n (e_t - \bar e)^2, \qquad \bar e_g = \frac{1}{n_g}\sum_{t \in g} e_t,\quad SE(\bar e_g) \approx \frac{s_e}{\sqrt{n_g}}.$$- $\bar e$ is the residual mean; $s_e$ the residual standard deviation; $p$ the number of fitted coefficients (for a regression with priors or shrinkage, $p$ is less clear and dividing by $n$ is a fine rough choice).
- $\bar e_g$ is the mean residual in a group $g$ (a weekday, a holiday, a month, a bin of fitted values) with $n_g$ days. The standard error formula assumes independent residuals, so treat it as a rough guide.
- For least squares with an intercept $\bar e = 0$ exactly. For a Bayesian fit with priors the posterior-mean residuals are close to 0 but not exactly. Either way the overall mean is nearly useless on the training days; the group means and the holdout mean are informative.
- Compare $s_e$ with the noise the model thinks it has. Normal: $\sigma$. Student-t: its standard deviation $\sigma\sqrt{\nu/(\nu - 2)}$ (the scale $\sigma$ is not the sd, Chapter 7.13). Negative Binomial: $\sqrt{\mu + \mu^2/\alpha}$ (changes from day to day). A big gap means the model's noise story is wrong.
- Heavy tails inflate $s_e$. A robust spread, the median absolute deviation scaled by 1.4826 (Chapter 4.14), ignores a few wild days. If $s_e \gg 1.4826\times\text{MAD}$, a few days dominate the variance.
Why do we need it?
A model that leans in a particular group (always too low on Mondays and holidays) wastes information that is available before the forecast is made, and an overall mean of 0 hides exactly that. The residual spread also tells you how wide an honest interval must be.
Where is it used?
Bias checks by segment in forecasting reviews (weekday, month, promotion on/off, store), calibration of the noise scale of Normal and Student-t likelihoods, and the first line of every residual report. In an A/B setting, it is the same check as "does the model fit all segments equally well?".
How is it used?
Group the residuals by every column you know in advance (weekday, month, holiday flag, regressor bins, fitted-value bins), compute the mean and standard error in each group, and look for bars outside about two standard errors. Compare the residual sd (and its robust version) with the model's own noise scale.
"Mean residual 0 means the forecasts have no bias."
It means the positives and negatives cancel over the training days (and for least squares with an intercept that is guaranteed). Bias hides inside groups and on the holdout. Check both.
"A small residual standard deviation proves a good model."
A model with too many parameters can drive the in-sample residuals down by memorising noise (Chapter 7.18). Residual size is judged on holdout days, and compared with what a baseline achieves (Chapter 7.5).
"The residual sd is the forecast error I should expect."
It ignores parameter uncertainty and any change after the training period, so real forecast errors are larger, and they grow with the horizon (Chapter 7.14).
In NumPyro, dist.Normal(mu, sigma) takes the standard deviation, so the residual standard deviation is directly comparable with the posterior of sigma. With a Student-t likelihood the scale parameter is smaller than the sd; with a Negative Binomial the noise level changes with the mean. Check the segments your model should handle: weekday, holiday windows, and bins of the fitted value. A lean in "all Mondays" means a missing weekly component or a Fourier order that is too low; a lean on "holiday days" means the holiday window or effect prior is too narrow or too tight.
$\bar e = \frac1n\sum e_t$ (0 by construction for OLS with an intercept); $s_e^2 = \frac{1}{n-p}\sum (e_t - \bar e)^2$; group mean $\bar e_g$ with $SE \approx s_e/\sqrt{n_g}$.
Look at means by group (weekday, holiday, fitted bin) and on a holdout. Compare $s_e$ with the model's own noise scale; use the robust $1.4826\times$MAD next to it.
Trap: overall mean 0 can hide a lean in every group; a small in-sample $s_e$ can be overfitting.
Quick check: a group of $n_g = 25$ holiday days has mean residual $+9$, and $s_e = 15$. Is the lean real?
$SE \approx 15/\sqrt{25} = 3$, so the mean is $9/3 = 3$ standard errors above 0 (beyond the usual cut-off of about 2). Yes, it looks real: the model under-predicts holidays by about 9 orders on average, so the holiday effect (or its window) is probably too small.
The shape of the residuals: histogram, KDE and Q-Q plot core
Your likelihood is a claim about the noise: "most days miss by a little, a few miss by more, and very large misses are very rare" (for a Normal), or "large misses happen more often" (for a Student-t). The residuals are your evidence. If the claim is right, the residuals have that shape; if not, the shape tells you which likelihood to try.
You already know the tools. A histogram and a KDE (a smooth version of the histogram) show the shape directly (Chapter 4.16). A Q-Q plot compares the sorted residuals with the values a Normal (or other) distribution would give; a straight line means "same shape" (Chapter 4.17). Here we use them, we do not re-teach them: the new part is what the shapes say about a forecasting model.
Three ways to say it:
- Picture: a bell, or a bell with fat tails, or a bell with a long arm on one side, or two bells.
- Numbers: in 200 days a Normal makes about $200\times0.0027 = 0.54$ days beyond 3 standard deviations; seeing 5 of them is a sign of heavy tails.
- Slogan: the tails of the residuals decide the likelihood; the middle of the plot decides little.
You fitted a model with a Normal likelihood and standardized the 200 residuals ($z_t = e_t/\hat\sigma$). You find 5 days with $|z_t| \gt 3$ and the largest is $|z| = 4.1$.
- A Normal puts $P(|Z| \gt 3) = 0.0027$ on one day. In 200 days the expected number of such days is $200 \times 0.0027 = 0.54$.
- Seeing 5 or more when 0.54 are expected has probability about $0.00024$ (a Poisson count with mean 0.54). That is roughly 1 in 4 000: not luck.
- The largest value: $P(|Z| \gt 4.1) = 0.000041$, so the expected number in 200 days is $0.008$. Seeing one is again a surprise for a Normal.
- A Student-t with $\nu = 4$, scaled to the same standard deviation, puts $P(|Z| \gt 3) \approx 0.013$ (expected $2.6$ days in 200) and $P(|Z| \gt 4.1) \approx 0.0044$ (expected $0.9$ days). Five days beyond 3 and one beyond 4.1 are much less surprising.
- Conclusion: the tails are heavier than the Normal says. First check whether the wild days share a pattern (the same date each year = a missing holiday). If not, a heavier-tailed likelihood is reasonable (Chapter 7.13).
For residuals $e_t$ with mean $\bar e$ and standard deviation $s_e$, the standardized residuals are $z_t = (e_t - \bar e)/s_e$. Two numbers summarize the shape (SciPy defaults, Chapter 4.14):
$$\text{skewness} = \frac{1}{n}\sum_{t=1}^n z_t^3, \qquad \text{excess kurtosis} = \frac{1}{n}\sum_{t=1}^n z_t^4 - 3 \qquad (\text{with } s_e \text{ computed by dividing by } n).$$Both are 0 for a Normal. Positive skewness: a long right tail. Positive excess kurtosis: heavier tails than the Normal. Reading a Q-Q plot of the standardized residuals against a reference distribution:
| What you see | What it means | What to try |
|---|---|---|
| Points on a straight line | The shape matches the reference | Keep the likelihood |
| S-shape: both ends bend away from the line | Heavy tails (more extreme days than Normal) | First check the dates of the wild days; then a Student-t likelihood |
| Only the upper end bends up | Right skew: a long arm of big positive misses | Missing positive spikes (holidays, promotions), a log or Box–Cox transform, or a skewed positive likelihood |
| A kink or step in the middle | Two groups mixed together | Find the hidden group (an indicator the model lacks) |
| Ends flatter than the line | Lighter tails than Normal | Usually harmless; check for bounded data or heavy smoothing |
What the shape matters for. Under non-Normal noise with finite variance, the fitted mean function of a least-squares model is still sensible; what goes wrong is the interval (its width and its tail probabilities) and the way the likelihood weighs extreme days. So shape decides the likelihood and the interval, not whether the trend is right.
Why do we need it?
Prediction intervals and "how likely is demand above capacity" come from the tails of the noise distribution. A Normal likelihood fitted to heavy-tailed residuals gives intervals that are too narrow for the extreme days and lets a few wild days pull the whole fit around.
Where is it used?
Choosing between Normal and Student-t errors in regression and forecasting models, checking a GLM (Poisson, Negative Binomial), the plot in statsmodels' plot_diagnostics, R's checkresiduals, and every model review of a Bayesian forecasting model.
How is it used?
Standardize the residuals, draw the histogram with a KDE and a Q-Q plot against the Normal, count the days beyond $|z| \gt 3$, and compare with what the Normal expects. If the tails are heavy, look at the dates first, then try Student-t and compare on a holdout.
"The Q-Q plot must be a perfect straight line, or the model is wrong."
With a few hundred days the ends wobble even for perfectly Normal noise. Look for a clear, repeatable bend (and use the lineup idea below to calibrate your eye).
"Heavy-tailed residuals mean the noise is heavy-tailed, so I need a Student-t."
Heavy tails in the residuals can come from a missing piece in the mean (an unmodelled holiday gives a big residual on the same date every year) or from changing spread (calm and busy periods mixed together look heavy-tailed when pooled). Check the dates and the residual-vs-fitted plot before changing the likelihood.
"The data must be Normal for a Normal likelihood."
The raw series is never Normal (it has a trend and a season). Only the residuals, after the mean function is removed, are compared with the likelihood.
"With 2 000 days I can use a normality test to decide."
With many days a normality test rejects tiny, harmless departures. Look at how big the departure is (the plot, the tail counts) instead of only at a p-value.
Your model offers Normal, Student-t and Negative Binomial likelihoods. The Q-Q plot of the standardized residuals is the practical way to choose between the first two: S-shape means ask for the Student-t, and the learned degrees of freedom $\nu$ should then be modest (a very large $\nu$ means the Student-t is not needed). For counts, the Normal Q-Q plot is not the right yardstick at all; use the count residuals in the section below. Before moving to a heavier tail, mark the biggest residuals on a calendar: the same date every year means a holiday column or a longer holiday window (Chapter 7.12), not a new likelihood.
Standardize: $z_t = (e_t - \bar e)/s_e$. Count the days beyond $|z| \gt 3$ (Normal expects 0.27%). Skewness and excess kurtosis are 0 for a Normal.
Q-Q: S-shape = heavy tails (Student-t?), one bent end = skew, kink = two groups. Straight line vs a Student-t with the right $\nu$.
Trap: heavy tails may come from a missing holiday or from changing spread. Check dates and residual-vs-fitted first.
Quick check: the 12 biggest residuals of a daily model fall on 25 December, 1 January, Thanksgiving and the same three days the year before. Student-t or something else?
Something else first: a holiday effect. Wild days that repeat on known dates are a missing part of the mean, not random heavy-tailed noise. Add holiday columns (and a window around them); then re-check the tails. A Student-t would hide the problem by treating predictable spikes as chance.
Residual vs fitted and residual vs time: the two workhorse plots core
After a storm you survey the damage. If the damaged houses are scattered randomly over the town, it was bad luck. If they are all on the east side, there is a cause: the wind came from the east. A residual plot is a damage survey. It asks where the misses are.
Two questions come first. Do the misses depend on what the model predicted? (Plot residual against the fitted value: are the misses bigger when the prediction is bigger? Do they curve?) Do the misses depend on when? (Plot residual against time: waves, steps, runs, spikes on certain dates?) Any column that is known in advance can be asked the same: weekday, month, a regressor, the holiday flag.
Three ways to say it:
- Picture: a shapeless cloud around zero is good; a bowl, a fan, a wave, a staircase or lonely spikes each say something specific.
- Numbers: residuals for fitted values below 130 have a typical size of 6, and above 130 a typical size of 16: the spread is 2.5 times larger on busy days.
- Slogan: plot the misses against everything that could explain them; any shape you can see is information the model could have used.
Eight days sorted by the model's fitted value, with their residuals:
| fitted $\hat y$ | 60 | 80 | 100 | 120 | 140 | 160 | 180 | 200 |
|---|---|---|---|---|---|---|---|---|
| residual $e$ | −9 | +6 | −3 | +5 | −11 | +14 | −16 | +20 |
- The signs are mixed, so there is no curve and no lean: the mean in each half is close to 0.
- Look at the sizes. Lower half (fitted 60 to 120): residuals $-9, +6, -3, +5$. Root mean square: $\sqrt{(81 + 36 + 9 + 25)/4} = \sqrt{37.75} \approx 6.1$.
- Upper half (fitted 140 to 200): $-11, +14, -16, +20$. Root mean square: $\sqrt{(121 + 196 + 256 + 400)/4} = \sqrt{243.25} \approx 15.6$.
- Ratio of spreads: $15.6/6.1 \approx 2.5$. The correlation of $|e|$ with the fitted value is about $0.82$. The plot is a funnel: heteroscedasticity (the section after the next ones).
- A rough rule: a ratio far above 1.5 to 2 (with enough days) is worth acting on. It is a rule of thumb, not a test.
A residual plot draws $e_t$ (vertical) against a variable (horizontal) that is known when the forecast is made: the fitted value $\hat y_t$, time $t$, or any column (weekday, regressor, holiday flag). To see a pattern through the scatter, add a smoother: the orange curve in our plots is the mean residual in equal-count bins (LOWESS does something similar). For a well-specified model the smoother is flat at 0 and the vertical spread is the same everywhere.
| Shape | What it says | What to try |
|---|---|---|
| 1 · Random cloud | No structure left in this view | Move on to the other views |
| 2 · Curve (bowl or hill) vs fitted | The model is too straight: a nonlinear effect or a bent trend is missing | More trend flexibility (changepoints), a transform, a nonlinear regressor effect |
| 3 · Funnel vs fitted | The spread grows (or shrinks) with the level: heteroscedasticity | Log or square-root transform, Negative Binomial, a variance model |
| 4 · Waves vs time | A cycle the model does not have | Another seasonal period, higher Fourier order, a driver |
| 5 · Step vs time | A level shift or slope change the trend missed | Changepoints near that date (grid, PELT), a regressor that explains the shift |
| 6 · Spikes at certain dates | Events the model does not know | Holiday / event columns with windows |
Why do we need it?
Numbers like the mean and the standard deviation squash the residuals into one value and lose the where. The plots keep the where, and the where says what is missing: a changepoint near day 170, a spread that grows with demand, a holiday nobody told the model about.
Where is it used?
Every regression course and every forecasting review: the "residuals vs fitted" panel of plot_regress_exog and R's plot(lm), residual-vs-time plots in backtest reports, and diagnostics of GLMs. It is also how a missing holiday or a slope change is found in practice.
How is it used?
Make the two plots with a smoother, then plot the residuals against each known column (weekday, month, holiday flag, regressors). Read the shape with the table, change the model in the way it suggests, refit, and look again. Repeat until the plots look shapeless.
"The residual-vs-fitted plot looks fine, so the model is fine."
That plot cannot see time. A step, a wave or a spike on certain dates is invisible there (the fitted values mix the days). Make the residual-vs-time plot too, and plot against each known column.
"I see a pattern in the residuals, so the model is wrong."
With few points the eye finds patterns in pure noise. Add the smoother, add the band, and compare with residual plots of simulated data from the model (the lineup, below) before changing anything.
"Plot the residuals against the actual $y$."
Residuals are always correlated with the actual values (big $y$ often means a big positive miss), so that plot shows a slope even for a perfect model. Plot against the fitted value or against a column known in advance.
These two plots find most of the practical problems in your forecasting model. A step or slow wave against time points at the trend: not enough changepoints in the grid, a changepoint range that stops before the recent history, or a Laplace scale $b$ so small that real slope changes are shrunk away. Waves with a fixed period point at the Fourier order or a missing second seasonality. Spikes on repeating dates point at the holiday table. A funnel against the fitted value points at the likelihood (next sections). The fix is always chosen by the shape.
Plot $e_t$ against: fitted value, time, and every known column. Add a smoother (bin means) and a ±2σ̂ band.
Bowl = too straight · funnel = changing spread · waves = missed cycle · step = missed level shift · spikes = missed events.
Trap: one plot can hide a problem another shows; and noise makes patterns, so compare with simulated residuals.
Quick check: the residual-vs-fitted plot is a clean cloud, but the residual-vs-time plot shows every residual after day 200 above zero. Which plot is right, and what is missing?
Both are right; they show different things. The fitted plot cannot show a jump at a date. The time plot says the model under-predicts from day 200 on: a level shift or slope change the trend missed. Look at the changepoint grid and PELT near day 200, and at whether the changepoint range actually reaches that late in the history.
Memory in the residuals: the ACF, the Ljung–Box test and Durbin–Watson core
Imagine you guess the weather for tomorrow every day, and each evening you write down how wrong you were. If your mistakes are pure luck, today's mistake tells you nothing about tomorrow's. If you notice "after a day where I guessed too low, I usually guess too low again", then your mistakes remember, and you could use yesterday's mistake to correct today's guess. A model whose residuals remember has left information on the table.
The ACF (autocorrelation function) measures that memory at each lag: how correlated is the residual with the residual $k$ days earlier? You met it in Chapter 7.3. A single lag can be a fluke (about 1 lag in 20 crosses the band by chance), so we also use the Ljung–Box test, which asks about many lags at once.
Three ways to say it:
- Picture: a bar chart of correlations at lags 1, 2, 3, …; for good residuals all bars stay inside a thin blue band.
- Numbers: with 100 days the band is $\pm 1.96/\sqrt{100} = \pm 0.196$; a residual correlation of $0.25$ at lag 1 is outside it.
- Slogan: if yesterday's miss predicts today's, the model is not finished.
Part 1, the numbers. Eight residuals with mean 0: $2, 3, 1, -2, -3, -2, 1, 0$.
- Sum of squares: $4 + 9 + 1 + 4 + 9 + 4 + 1 + 0 = 32$.
- Lag-1 products $e_t e_{t-1}$: $3\cdot2 = 6,\; 1\cdot3 = 3,\; (-2)\cdot1 = -2,\; (-3)(-2) = 6,\; (-2)(-3) = 6,\; 1\cdot(-2) = -2,\; 0\cdot1 = 0$. Sum $= 17$.
- $r_1 = 17/32 \approx 0.53$: neighbouring residuals have the same sign more often than not (runs of $+$ then $-$).
- Durbin–Watson: $DW = \sum_{t\ge2}(e_t - e_{t-1})^2/\sum e_t^2 = (1 + 4 + 9 + 1 + 1 + 9 + 1)/32 = 26/32 \approx 0.81$. Compare $2(1 - r_1) = 0.94$; the two agree up to end effects. $DW \approx 2$ means no lag-1 memory; below 2 means positive memory; above 2 negative memory.
Part 2, a test on many lags. With $n = 100$ days a regression's residuals have $r_1 = 0.25$, $r_2 = 0.10$, $r_3 = -0.05$.
- $Q(3) = n(n+2)\sum_{k=1}^{3} \dfrac{r_k^2}{n-k} = 100\cdot102\left(\dfrac{0.0625}{99} + \dfrac{0.01}{98} + \dfrac{0.0025}{97}\right) = 10\,200 \times 0.000759 \approx 7.74$.
- Compare with a $\chi^2$ distribution with 3 degrees of freedom: $p \approx 0.052$. Borderline.
- Test only lag 1: $Q(1) = 10\,200 \times 0.0625/99 \approx 6.44$, $p \approx 0.011$. The memory sits at lag 1 and testing more lags dilutes the evidence. The choice of $h$ matters; the widget below shows how.
For residuals $e_1, \dots, e_n$ with mean $\bar e$, the residual autocorrelation at lag $k$ is (statsmodels' convention, as in Chapter 7.3)
$$r_k = \frac{\sum_{t=k+1}^{n} (e_t - \bar e)(e_{t-k} - \bar e)}{\sum_{t=1}^{n} (e_t - \bar e)^2}.$$If the residuals are white noise, each $r_k$ is roughly Normal with sd $1/\sqrt n$, hence the band $\pm1.96/\sqrt n$. The Ljung–Box statistic over the first $h$ lags is
$$Q(h) = n(n+2)\sum_{k=1}^{h} \frac{r_k^2}{n-k}, \qquad Q(h) \approx \chi^2_{h - m}\ \text{ if there is no memory},$$where $m$ is the number of ARMA parameters that were fitted to produce the residuals (0 for the residuals of a regression on calendar columns; in statsmodels the argument is model_df). Small p-value: evidence of memory somewhere in the first $h$ lags.
- Choice of $h$ (rules of thumb, Hyndman): about 10 for non-seasonal data; about $2m$ for seasonal data of period $m$ (14 for daily data with a weekly pattern, which is what we use); never more than about $n/5$. $h$ must reach the lag where memory is expected, or the test misses it.
- Squared residuals. Running the same test on $e_t^2$ detects volatility clustering: big misses follow big misses (quiet and stormy periods), even when $e_t$ itself has no memory.
- Durbin–Watson $DW = \sum_{t=2}^{n}(e_t - e_{t-1})^2 / \sum_t e_t^2 \approx 2(1 - r_1)$ looks at lag 1 only. Use it as a quick number from a regression table; it is not valid when lagged values of $y$ are among the regressors.
- Bayesian models. With a fitted prior and a non-ARMA structure, the exact $\chi^2$ reference is only approximate. Treat $Q$ as a score and compare it with the same score for residuals of data simulated from your model (the lineup below).
Why do we need it?
Memory in the residuals means the model ignores a pattern it could use, and makes its own uncertainty too small, because the likelihood treats correlated days as independent evidence (Chapter 7.3). One test over many lags gives an honest verdict on "is there any memory?" that a single bar cannot.
Where is it used?
Every ARIMA and regression-with-ARMA-errors workflow (plot_diagnostics, acorr_ljungbox), econometrics (Durbin–Watson in regression tables), model reviews of forecasting systems, and the validation step before intervals are trusted.
How is it used?
Compute the residual ACF up to about 2–3 seasons, draw the band, and run Ljung–Box at a lag that covers the seasonal period (for daily data with a weekly pattern, 14). Also test the squared residuals. A small p-value means: look at which lags are outside the band to see what is missing.
"No ACF bar crosses the band, so there is no memory."
Each bar is a separate small test, and a pattern of several bars that are individually inside the band (all slightly positive) can still be real. Use Ljung–Box for the joint question, and a lag range that covers the seasonal period.
"Ljung–Box p = 0.40 proves the residuals are white noise."
It only says there is no evidence of memory in the lags you tested, with the data you have. Absence of evidence is not evidence of absence, especially with few days or the wrong $h$.
"I tested h = 10 on daily data with a weekly pattern and passed."
That test never looked at lag 14 or 21. Cover at least one or two full seasons.
"Durbin–Watson close to 2 means the residuals are independent."
It only checks lag 1. A weekly memory (lag 7) leaves $DW$ near 2 and the ACF at lag 7 far outside the band.
For your model, compute the residual ACF to at least 28 or 35 lags (daily data with a weekly pattern) and run Ljung–Box at lag 14 and lag 28, on the training days and on each rolling-origin fold. If the lag-1 and lag-2 bars are outside the band, think of short-term momentum and an AR residual term. If lag 7, 14, 21 stick out, think of the weekly Fourier order or a weekday-specific effect (a holiday falling on a particular weekday). If the ACF is positive and decays very slowly, think of a trend the changepoints did not capture. If the residuals look fine but squared residuals have memory, think of changing spread (next sections).
$r_k = \sum (e_t - \bar e)(e_{t-k} - \bar e) / \sum (e_t - \bar e)^2$; band $\pm1.96/\sqrt n$. $Q(h) = n(n+2)\sum_{k=1}^{h} r_k^2/(n-k) \sim \chi^2_{h-m}$. $DW \approx 2(1 - r_1)$.
Choose $h$ to reach the season: about 10 non-seasonal, $2m$ seasonal (14 for daily/weekly), at most $n/5$. Test $e_t^2$ for volatility clustering.
Trap: small p = memory somewhere (read the ACF to see where); large p is not a proof; DW sees lag 1 only.
Quick check: $n = 250$ daily residuals have $r_7 = 0.45$ and all other lags are near 0. Does Ljung–Box with $h = 5$ catch it? With $h = 14$?
With $h = 5$ it never looks at lag 7, so $Q$ is near its null value and the test passes: a miss. With $h = 14$ it includes $r_7^2/(n - 7) \times n(n+2) \approx 0.2025 \times 252 \times 250/243 \approx 52$ (plus a bit from lag 14), against a $\chi^2_{14}$ whose typical size is about 14, so $p$ is tiny. The test only finds memory at lags it includes.
Heteroscedasticity: when the spread of the misses changes core
A small bakery sells about 50 loaves a day and is usually off by 5. A big supermarket bakery sells about 200 and is usually off by 20. Same kind of product, same kind of forecaster, but a ±5 error would be a disaster for the supermarket's planner and a ±20 error would be wild for the small bakery. Misses scale with the size of the thing.
Most simple models assume the miss is the same size on every day (one number $\sigma$). When the spread really changes with the level (or the season, or the weekday), that assumption is false. The long word for it is heteroscedasticity ("different scatter"); the opposite, constant spread, is homoscedasticity.
Three ways to say it:
- Picture: in the residual-vs-fitted plot the cloud opens up like a funnel; a constant-width band cannot follow it.
- Numbers: sd 5 on 50-order days and sd 20 on 200-order days. A single pooled $\sigma = 14.6$ is 3 times too big for the first and 27% too small for the second.
- Slogan: constant-$\sigma$ intervals are too wide when it is quiet and too narrow when it matters.
Daily demand is about 50 on quiet days and 200 on busy days, with noise of 10% of the level: sd 5 and sd 20. The model uses one Normal with a pooled $\sigma$ and builds 80% intervals. Half the days are quiet, half busy.
- Pooled variance: $(5^2 + 20^2)/2 = (25 + 400)/2 = 212.5$, so $\hat\sigma = 14.6$.
- An 80% interval is $\pm1.2816\,\hat\sigma = \pm18.7$ orders on every day.
- Quiet days (true sd 5): the half-width is $18.7/5 = 3.7$ standard deviations. Coverage $= 2\Phi(3.74) - 1 \approx 99.98\%$. The interval is far too wide and says nothing useful.
- Busy days (true sd 20): the half-width is $18.7/20 = 0.93$ standard deviations. Coverage $= 2\Phi(0.93) - 1 \approx 65\%$.
- So the "80%" interval holds 65% of the busy days: the days where under-forecasting means empty shelves. A model with $\sigma_t = 0.10\,\mu_t$ (spread proportional to the level) would give 80% on both kinds of day.
The model $y_t = \mu_t + \epsilon_t$ is homoscedastic if $Var(\epsilon_t) = \sigma^2$ for every $t$, and heteroscedastic if $Var(\epsilon_t) = \sigma_t^2$ changes with $t$. In forecasting the spread usually depends on one of three things: the level ($\sigma_t$ grows with $\mu_t$), the calendar (busy season, weekends, holiday weeks), or recent shocks (volatility clustering: big misses follow big misses).
How to detect it (plots first, numbers second):
- Residual vs fitted: a funnel. Residual vs time: a band that widens or has calm and stormy stretches. The bar chart of the residual sd per bin of the fitted value (used in the panel at the end).
- The spread ratio: sd of residuals in the top third of the fitted values divided by the sd in the bottom third. Far above 1.5 to 2 is worth acting on (rule of thumb).
- Correlation of $|e_t|$ with $\hat y_t$, and the ACF of $e_t^2$ (volatility clustering).
- Named tests, for completeness: Breusch–Pagan and White regress $e_t^2$ on the regressors (or the fitted value) and test whether anything explains the squared residuals (statsmodels
het_breuschpagan). Use them as a second opinion after the plots, not instead of them.
Three responses (module 70):
- Transform the data so the spread becomes constant (Chapter 4.18). If the sd grows in proportion to the level (a constant percentage error), a log does it. If the variance grows in proportion to the level (counts), a square root does it. Box–Cox chooses the power by likelihood. Price: forecasts must be back-transformed, and back-transformed point forecasts estimate the median, not the mean.
- Use a likelihood whose variance depends on the mean. Poisson: $Var = \mu$. Negative Binomial: $Var = \mu + \mu^2/\alpha$ (Chapter 7.13). Gamma or log-Normal for positive amounts: $sd \propto \mu$. No extra work: the funnel is part of the model.
- Model the variance explicitly. Let the scale depend on something known: $\sigma_t = c\,\mu_t$, or $\log\sigma_t = a + b\log\mu_t$, or $\sigma_t = \exp(a + b\,x_t)$ with $x_t$ a weekday, holiday flag or month. In least-squares language this is weighted least squares (weight $1/\sigma_t^2$); in a Bayesian model you write the scale as a function of known columns and give $a$, $b$ priors.
What heteroscedasticity does not do: it does not bias the least-squares estimate of the mean function. It makes the usual standard errors, posterior widths and prediction intervals wrong, and wastes efficiency (days with small noise should count more).
Why do we need it?
Forecast intervals and "probability that demand exceeds capacity" depend on the spread on the particular day. With a constant spread, quiet days waste safety stock and busy days get too little. Heteroscedasticity is common wherever noise is proportional: sales, revenue, traffic, counts.
Where is it used?
Log-transformed regressions and Box–Cox forecasting, Poisson and Negative Binomial GLMs (variance built in), weighted least squares, heteroscedastic regression and "distributional" models in Bayesian work, and GARCH-type models in finance (for volatility clustering).
How is it used?
Make the residual-vs-fitted plot and the spread ratio. If the funnel is there, choose: transform (log / sqrt), switch the likelihood (Negative Binomial for counts), or give the scale a formula. Refit and look at the same plot and at the coverage of the intervals by level (low, middle, high days).
"Heteroscedasticity makes my trend and seasonal estimates biased."
The least-squares estimate of the mean function stays unbiased and consistent. What breaks is the uncertainty: standard errors, posterior widths and prediction intervals are wrong, and the fit is not the most efficient.
"Switching to a Student-t likelihood fixes a funnel."
A Student-t has heavier tails but still one constant scale. It does not make the scale grow with the level. The funnel needs a transform, a likelihood with a mean-dependent variance, or a variance model.
"Always take logs."
Logs fix a spread that grows in proportion to the level and need $y \gt 0$ (zeros are a problem; log1p is a patch). For counts, a Negative Binomial is more natural. And the exponential of the mean of the log is the median forecast, not the mean: it is lower by a factor of about $e^{\sigma^2/2}$.
"The funnel is only about the level. Time cannot matter."
Spread can also depend on the calendar (a stormy December) or on recent shocks. Plot the residuals against time and the ACF of the squared residuals as well.
The Normal and Student-t likelihoods of your forecasting model use one scale $\sigma$ for all days unless you give it a formula. If your series is proportional (revenue, orders that scale with traffic), expect a funnel in the residual-vs-fitted plot. Your options map onto what the model already offers: model on a log scale, use the Negative Binomial for counts (its variance $\mu + \mu^2/\alpha$ grows with the mean), or make the scale a function of known columns. In an A/B framework like yours, where Normal and Student-t likelihoods are used for other metrics, the same question appears: if segments with a higher mean also have a higher spread, a single $\sigma$ for all groups under-states the uncertainty of the big segments and over-states it for the small ones. Check the spread by group, and by bins of the fitted mean.
Heteroscedastic: $Var(\epsilon_t) = \sigma_t^2$ changes. Detect: funnel, spread ratio $\gg 1.5$–2, $|e|$ vs fitted, ACF of $e^2$, Breusch–Pagan (second opinion).
Responses: (1) transform (log: $sd \propto \mu$; sqrt: $Var \propto \mu$), (2) likelihood with mean-dependent variance (Poisson, NB: $\mu + \mu^2/\alpha$), (3) explicit scale model $\sigma_t = c\mu_t$ or $\exp(a + bx_t)$ (WLS).
Trap: the mean stays unbiased, the intervals break; Student-t does not fix it; back-transformed forecasts are medians.
Quick check: residual sd is 6 for fitted values below 100 and 15 for fitted values above 200. Which transform is the natural first try, and why?
A log. The sd grows roughly in proportion to the level (the level at least doubles from below 100 to above 200 and the sd rises by 2.5 times), the signature of a constant percentage error. After taking logs, a constant percentage error becomes a constant absolute error. For counts you would try the square root or a Negative Binomial.
Residuals for counts and Bayesian models: Pearson and randomized quantile residuals core
For counts there is a trap. A model for orders per day should have bigger misses on big days: if the model expects 20 orders, a typical miss may be 6; if it expects 100, a typical miss may be 30. So the raw residuals of a correct count model already form a funnel. Reading that funnel as "something is wrong" is a false alarm.
The fair way to compare days is to measure each miss in units of the model's own idea of a typical miss on that day. That is the Pearson residual: $(y - \hat\mu)/\text{(typical miss under the model)}$. If the model's idea of the spread is right, these residuals have spread 1 everywhere. For whole numbers there is a second trouble: the residuals are lumpy (0, 1, 2, …). The randomized quantile residual removes the lumps: it pushes each observation through the model's own CDF and turns the result into a standard Normal score.
Three ways to say it:
- Picture: raw residuals fan out (even for a good model); Pearson residuals form an even band; quantile residuals fall on the straight line of a Normal Q-Q plot.
- Numbers: expected 100, observed 130: under a Poisson that is a miss of $30/\sqrt{100} = 3$ standard deviations; under a Negative Binomial with $\alpha = 10$ it is $30/33.2 = 0.9$.
- Slogan: measure each miss with the model's own ruler; the model is right when every ruler gives spread 1.
Two days from a count model. For the Negative Binomial in the mean-dispersion form, $Var = \mu + \mu^2/\alpha$ (Chapter 4.8); suppose $\hat\alpha = 10$.
- Day A: $\hat\mu = 100$, observed $y = 130$. Raw residual $+30$.
- Negative Binomial sd: $\sqrt{100 + 100^2/10} = \sqrt{1100} \approx 33.2$. Pearson residual $= 30/33.2 \approx 0.90$: an ordinary day.
- Poisson sd: $\sqrt{100} = 10$. Pearson residual $= 30/10 = 3.0$: a "3-sigma" day. If the model is a Poisson, every busy day looks wild, a sign that the Poisson variance is too small (overdispersion).
- Day B: $\hat\mu = 20$, $y = 28$. Negative Binomial sd $= \sqrt{20 + 400/10} = \sqrt{60} \approx 7.75$, Pearson $= 8/7.75 \approx 1.03$. Poisson: $8/\sqrt{20} \approx 1.79$.
- Randomized quantile residual for day A: the model's CDF gives $F(129) = 0.8215$ and $F(130) = 0.8280$. Draw $u$ between these (say the midpoint $0.8248$) and take $\Phi^{-1}(u) \approx 0.93$. It is close to the Pearson residual here; for small counts (0, 1, 2) the two differ a lot and the quantile version is the trustworthy one.
Let $\hat\mu_t$ be the model's predicted mean and $V(\hat\mu_t)$ its predicted variance. The Pearson residual is
$$r_t = \frac{y_t - \hat\mu_t}{\sqrt{V(\hat\mu_t)}}, \qquad V = \sigma^2\ (\text{Normal}),\quad \mu\ (\text{Poisson}),\quad \mu + \frac{\mu^2}{\alpha}\ (\text{NB2}),\quad \sigma^2\frac{\nu}{\nu - 2}\ (\text{Student-t},\ \nu \gt 2).$$If the model is right: $E[r_t] \approx 0$ and $Var(r_t) \approx 1$ on every kind of day. The average of $r_t^2$ over the days is the Pearson dispersion: near 1 is fine; clearly above 1 means more spread than the model allows (overdispersion).
The randomized quantile residual (Dunn and Smyth) needs the model's CDF $F(\cdot \mid \hat\theta_t)$. For a discrete $y_t$:
$$u_t = F(y_t - 1) + v_t\,\big(F(y_t) - F(y_t - 1)\big),\quad v_t \sim \text{Uniform}(0,1),\qquad r_t^{q} = \Phi^{-1}(u_t).$$For a continuous $y_t$ use $u_t = F(y_t)$ directly. If the model is right, $u_t$ is Uniform(0,1) and $r_t^q$ is $N(0,1)$, for any likelihood (Normal, Student-t, Poisson, NB, zero-inflated). Then use the usual tools on them: histogram, Q-Q plot against the Normal, vs fitted, vs time, ACF.
- Which prediction for a Bayesian model? Quick and good enough: plug in the posterior means $\hat\mu_t$ and $\hat\alpha$. Better: use the whole posterior, $\bar u_t = \frac1S\sum_s F(y_t \mid \theta^{(s)})$ (averaging the CDF value over $S$ posterior draws). This is the posterior-predictive probability integral transform (PIT) that Chapter 7.16 uses to judge calibration; $\Phi^{-1}(\bar u_t)$ is a Bayesian quantile residual. A third option is one set of residuals per draw, which shows parameter uncertainty.
- The randomization is not cheating: it is the exact discrete analogue of the PIT. Without it, quantile residuals of low counts would sit on a few fixed values.
Why do we need it?
Raw residuals of any count or positive-skew model fan out by design, so they cannot tell a good model from a bad one. You need residuals that are comparable across days and have a known distribution when the model is right.
Where is it used?
GLM diagnostics (Poisson, Negative Binomial, logistic: resid_pearson in statsmodels, the DHARMa package in R for simulated quantile residuals), checking overdispersion, and Bayesian model checking of count likelihoods (posterior-predictive PITs).
How is it used?
Compute the predicted mean and variance for each day, form Pearson residuals, and plot them against the fitted value (flat band expected) and time. Check that their variance is near 1. For Q-Q plots, ACF or lumpy low counts, use randomized quantile residuals instead.
"The raw residuals of my count model form a funnel, so the model is wrong."
A correct Poisson or Negative Binomial model has a funnel in its raw residuals, because the spread depends on the mean. Judge Pearson or quantile residuals, which should be an even band.
"Pearson dispersion of 8 for a Poisson fit means my trend and seasonality are wrong."
It means the variance function is too small. The mean function can be fine. Switch to a Negative Binomial (or a quasi-Poisson / robust standard errors) and re-check; only then look at the mean function.
"Randomizing the residuals is fudging."
It is the standard exact trick for discrete data (the analogue of the probability integral transform). Different random draws give slightly different plots, so look at two or three draws if it matters.
"A Normal Q-Q plot of Pearson residuals of counts must be straight."
For small counts Pearson residuals are lumpy and skewed even for a perfect model. Use randomized quantile residuals for that plot.
Your forecasting model has a Negative Binomial likelihood for counts. After the fit, take the posterior mean of the mean function and of the concentration $\alpha$ (check which parameterization your code uses before computing $V$: for NumPyro's NegativeBinomial2(mean, concentration) the variance is $\mu + \mu^2/\text{concentration}$; other parameterizations have different formulas, Chapter 4.8). Compute Pearson residuals and look at their dispersion and the residual-vs-fitted band; then randomized quantile residuals for the Q-Q plot and the ACF. A dispersion well above 1 with the NB means $\alpha$ is too large for some days (or excess zeros / a missing component); a dispersion well below 1 means the NB is more spread than the data. In an A/B framework with Poisson metrics the same check is how you notice overdispersion before trusting a Poisson posterior.
Pearson: $r_t = (y_t - \hat\mu_t)/\sqrt{V(\hat\mu_t)}$ with $V = \mu$ (Poisson), $\mu + \mu^2/\alpha$ (NB2), $\sigma^2$ (Normal). Right model: mean 0, variance about 1, flat band.
Randomized quantile: $u_t = F(y_t - 1) + v_t[F(y_t) - F(y_t - 1)]$, $r_t^q = \Phi^{-1}(u_t) \sim N(0,1)$ for any likelihood.
Bayesian: plug in posterior means, or average $F$ over draws (the posterior-predictive PIT). Trap: raw residuals of a correct count model fan out.
Quick check: a Negative Binomial model with $\hat\alpha = 4$ predicts $\hat\mu = 40$ and you observe 70. What is the Pearson residual?
Variance $= 40 + 40^2/4 = 40 + 400 = 440$, sd $\approx 20.98$. Pearson $= (70 - 40)/20.98 \approx 1.43$. An unremarkable day. Under a Poisson (sd $= \sqrt{40} \approx 6.3$) it would be $4.7$ standard deviations, which is why Poisson fits to overdispersed counts report so many "outliers".
Is that pattern real? Compare with residuals of simulated data core
Look at a random scatter of stars and you will "see" shapes: a ladle, a hunter. The human eye is a pattern-finding machine, and it finds patterns in noise. Residual plots of a perfectly good model also contain waves, clumps and small funnels.
A police lineup solves a similar problem: put the suspect among five innocent people and ask a witness to pick. For residuals: draw the real residual plot and five plots of residuals from data simulated by the model itself (what the model believes the world is like), shuffle them, and see whether you can tell which one is real. If you cannot, there is no pattern worth chasing. If it jumps out, the model is missing something.
Three ways to say it:
- Picture: six small plots, one of them is the real one.
- Numbers: if there is nothing to find, you pick the real panel with chance $1/6 \approx 17\%$ by luck alone. Picking it right is weak evidence (p ≈ 0.17); with 20 panels it would be $1/20 = 0.05$.
- Slogan: judge your residuals against what your model itself would produce.
The same idea with a single number instead of a picture: a posterior predictive check (Chapter 6.8) on the Ljung–Box statistic.
- The real residuals give $Q(14) = 31.4$.
- Simulate 1 000 datasets from the fitted model (so the noise is independent, by assumption). For each, fit the same model again, compute its residuals, and compute $Q(14)$.
- If the model were right, $Q(14)$ would follow about a $\chi^2_{14}$ distribution: median about 13, and only 5% of values above 23.7.
- The share of simulated $Q$ values above 31.4 is about $0.005$ (the $\chi^2_{14}$ tail at 31.4). So 31.4 would be extremely unusual if the model's independence story were true.
- Conclusion: the pattern is real; the model's story about independence is wrong. The share is a posterior predictive p-value for this statistic: not a probability that the model is true, only a measure of how surprising the real residuals are under the model.
The lineup protocol. (1) Fit the model to the real data and keep its residuals. (2) Simulate $R$ replicated datasets from the fitted model, $y^{\text{rep}} = \hat\mu + \text{noise drawn from the model's likelihood}$ (for a Bayesian model: first draw the parameters from the posterior, then the data, i.e. posterior-predictive draws). (3) Process every replicate exactly like the real data (refit, same residual formula). (4) Show the real residual plot hidden among the $R$ replicated ones in random order. (5) If an observer picks the real one more often than chance, the model has missed something the plot shows.
The numeric version: choose a discrepancy statistic $T$ (the Ljung–Box $Q$, the lag-1 autocorrelation, the largest $|z|$, the spread ratio, the number of days beyond $3\sigma$), compute $T$ for the real residuals and for each replicate, and report the share of replicates at least as extreme as the real value.
- For a linear-Gaussian model the residuals of a refitted replicate depend only on the simulated noise, not on the parameter draw, so drawing the parameters does not change the picture. For nonlinear, regularised or SVI-fitted models, drawing from the (approximate) posterior matters.
- Compare like with like: if the real residuals use the posterior mean, so must the replicates.
Why do we need it?
People over-read noise and, just as often, miss real structure that is small but important. A lineup turns "does this look odd?" into a fair comparison with what chance alone produces under your own model.
Where is it used?
Graphical inference in statistics (Buja and colleagues' "lineup" protocol), posterior predictive checks in Bayesian workflows (replicated data and test statistics, as in the bayesplot and ArviZ libraries), and model reviews where a plot is contested.
How is it used?
Draw posterior-predictive datasets from the fitted model, run the residual pipeline on each, and put the real plot among them. Or pick a statistic ($Q$, $r_1$, spread ratio, max $|z|$) and report where the real value sits among the replicated ones.
"If I can find a pattern in the lineup, the pattern must be big."
It only means the real residuals differ from what the model would produce. The difference can be small but important, or it can be large but harmless for forecasting. Judge the size afterwards (does a fix improve rolling-origin accuracy and calibration?).
"If I cannot pick the real panel, the model is correct."
It means this view and this amount of data show no problem. Other views (the ACF, vs fitted, the Q-Q plot, a statistic you did not plot) may still show one, and 280 days cannot show small effects.
"The posterior predictive p-value is the probability that the model is right."
It is how often data simulated from the model show a statistic at least as extreme as the real data. It measures surprise under the model, nothing more.
For your SVI-fitted model, the lineup is a cheap and honest way to review a contested plot. Draw posterior-predictive series from the guide (parameters from the fitted guide, then data from the likelihood: Normal, Student-t or Negative Binomial), compute the same residual and ACF pictures for each, and place the real ones among them. If the real ACF stands out from the replicated ones, the independent-noise assumption is wrong for the data. Remember that SVI posteriors can be too narrow, so replicated data may be slightly too tidy; compare with a second guide or an NUTS run if the decision matters (Chapter 6.15). Posterior predictive checks for whole forecasts are in Chapter 7.14.
Lineup: real residual plot + $R$ plots of residuals from data simulated from the fitted model, processed identically; shuffle; can you tell?
Numeric version: statistic $T$ (Ljung–Box $Q$, $r_1$, max $|z|$, spread ratio); share of replicates with $T \ge T_{\text{real}}$ = posterior predictive p-value.
Trap: a PPC p-value measures surprise under the model, not the probability the model is right; not seeing a pattern is not a proof of none.
Quick check: in a lineup of 20 panels (19 simulated, 1 real) a colleague picks the real one at once. How strong is that evidence, and what does it say about the model?
If there were nothing to find, picking the real panel would have probability $1/20 = 0.05$ by luck, so it is moderate evidence (p ≈ 0.05) that the real residuals differ from what the model produces. It says the model is missing something visible in that plot; what exactly (memory, level shift, wild days) you read from the shape.
Autocorrelated residuals: what the model missed, and four responses core
After a flood a river stays high for days. If your forecast of today's level is too low, tomorrow's will probably be too low as well, because the water that caused today's surprise is still there. A forecast that ignores this keeps being wrong in the same direction.
A calendar model (trend, weekly rhythm, holidays, regressors) draws the expected path from the date alone. Its residuals have memory when something outside the calendar pushes the series for a while (a viral post, the weather, a competitor's stock-out) or when part of the calendar is missing (a slow cycle, a level shift, a weekday shape). Both leave runs of same-signed residuals. The ACF tells you that memory exists; its shape and the residual-vs-time plot hint at why.
Three ways to say it:
- Picture: in the residual-vs-time plot, long runs above zero, then long runs below, instead of a jittery cloud.
- Numbers: if $r_1 = 0.55$, yesterday's miss predicts $55\%$ of today's.
- Slogan: leftover memory is leftover information: either model it or find the missing piece.
Reading three residual ACFs (module 71 asks: "if $Corr(e_t, e_{t-1}) \ne 0$, what is missing?").
- Series A: $r_1 = 0.55,\ r_2 = 0.30,\ r_3 = 0.17,\ r_4 = 0.09$, lag 7 about 0. The ratios $0.30/0.55 = 0.55$ and $0.17/0.30 = 0.57$ are about equal: a geometric decay $r_k \approx 0.55^k$ ($0.55^2 = 0.3025$, $0.55^3 = 0.166$, $0.55^4 = 0.092$). That is the signature of an AR(1) with $\varphi \approx 0.55$: short-term momentum. Using it would shrink one-step errors by a factor $\sqrt{1 - 0.55^2} = 0.835$, about 16% (compare Chapter 7.3).
- Series B: $r_7 = 0.45,\ r_{14} = 0.20,\ r_{21} = 0.09$, lags 1 to 6 near 0. The same geometric pattern but at lag 7: the miss on a given weekday repeats a week later. The weekly shape is not captured, e.g. a Fourier order that is too low for a sharp weekday pattern.
- Series C: $r_1 = 0.86,\ r_2 = 0.76,\ r_3 = 0.68,\ r_4 = 0.60$, fading slowly, with the residuals wandering up and down over weeks in the time plot. A slow component is missing: a level shift or slope change the changepoints did not catch, or a regressor that moves slowly.
The residuals are autocorrelated if $Corr(e_t, e_{t-k}) \ne 0$ for some lag $k$ beyond what chance explains (Ljung–Box, previous section). Reading the shape:
| ACF of the residuals | Likely meaning | First thing to try |
|---|---|---|
| Fast geometric decay from lag 1 (AR(1)-like) | Short-term momentum from things outside the calendar | An AR residual term, or a regressor that carries the shock |
| Spikes at 7, 14, 21 (maybe decaying) | The weekly pattern is not captured | Higher weekly Fourier order, weekday-specific effects (holiday × weekday) |
| Slow, almost linear decay; the time plot wanders | A slow component is missing (level shift, slope change, drifting driver) | More or better-placed changepoints, a weaker shrinkage on $\delta_j$, a regressor |
| Waves with a long period (e.g. positive at lag 30, negative at 15) | A cycle of another length | Another seasonal period, a regressor with that cycle |
| Nothing in the ACF but isolated big residuals on repeating dates | Missing events (the ACF is blind to isolated spikes) | Holiday columns (Chapter 7.12) |
Consequences (see Chapter 7.3): short-horizon forecasts are worse than they could be, posteriors for trend and regression coefficients are too narrow because correlated days count as independent evidence, and intervals for sums (a week's total) are too narrow.
The four responses (module 71):
- Richer deterministic terms (seasonal, trend, holidays, regressors). Preferred whenever the shape points at a missing piece: it keeps the model explainable.
- An AR residual component. Write the noise as $\epsilon_t = \varphi\,\epsilon_{t-1} + \eta_t$ with $|\varphi| \lt 1$ and independent $\eta_t$ (AR(1); AR($p$) adds more lags). The forecast gets a correction $\varphi^h e_T$ that fades with the horizon. In NumPyro this is a recursion, written with
jax.lax.scan(Chapter 6.16). - A state-space model. Let the level, slope or seasonal pattern itself drift over time (a hidden state updated by small random shocks, filtered with a Kalman filter). Memory is then permanent (the level keeps its new value) rather than fading.
- A Gaussian process for the leftover: a random smooth function with a correlation length $\ell$ (a kernel). It handles irregular times and smooth wandering but costs $O(n^3)$.
The next section compares the AR, state-space and Gaussian-process ways of giving the noise a memory. Chapter 7.18 puts them next to the alternatives to a Prophet-style model.
Why do we need it?
Independent noise is the cheapest assumption in a calendar model, and also the one most often broken by real series. Naming what the residual memory means tells you whether to add a calendar term (cheap, explainable) or a memory term (better short-term forecasts), instead of guessing.
Where is it used?
Regression with ARMA errors (SARIMAX with exogenous columns), Prophet-style models extended with AR terms, structural time-series models, Gaussian-process forecasting, and backtest reviews that decompose errors by horizon.
How is it used?
Read the ACF shape and the residual-vs-time plot, try the mean fix that matches the shape, refit, and re-run Ljung–Box. Only if memory remains and short horizons matter, add an AR term, then compare short- and long-horizon accuracy and calibration on rolling origins.
"Autocorrelated residuals mean I should switch to ARIMA."
First ask what the memory is made of. A weekday shape, a slope change or a missing holiday can all show up as autocorrelation, and each is fixed in an easier (and more explainable) way by adding the missing term.
"After adding an AR(1) term the ACF is flat, so the problem is solved."
An AR(1) term can soak up the symptom of a missed slope change (lag 1 looks fine) while the real cause remains: the longer-horizon forecasts stay off. Check accuracy by horizon on rolling origins, not just the ACF.
"I will add yesterday's value $y_{t-1}$ as a regressor."
For a forecast several days ahead, yesterday's value is unknown on days 2, 3, … Use lags at least as old as the horizon, or put the memory in the noise model. A backtest that uses $y_{t-1}$ on the horizon days leaks the answer (Chapter 7.12).
"Autocorrelation only makes forecasts less accurate."
It also makes the stated uncertainty too small: posterior intervals for trend and regression coefficients and intervals for totals assume independent days.
A Prophet-style model with changepoints, Fourier terms, holidays and regressors treats the leftover as independent noise (Chapter 7.3). Where do its residual runs most often come from? (1) A slope change after the last allowed changepoint, or a Laplace scale $b$ so small that a real change is shrunk away: the residuals drift in one direction over weeks. (2) A Fourier order too low for the weekly or yearly shape: spikes at the seasonal lags. (3) A holiday window that is too short: runs of same-signed residuals around holidays. (4) True momentum from outside: lag-1 geometric decay. In that last case an AR(1) error term is the natural extension, written with scan in NumPyro; keep the horizon in mind, because the AR correction fades quickly.
"The residuals are autocorrelated, so I add an AR term."
"The residual ACF shows memory, so I first ask what is missing. I look at the shape: spikes at the weekly lags point at the weekly terms, a slow drift points at the trend or a regressor, a fast geometric decay points at true momentum. I fix the mean function first and add an AR term only if momentum remains. I confirm with Ljung–Box and with accuracy by horizon on rolling origins."
Model answer: "Residual autocorrelation is a symptom with several possible causes. A calendar model such as mine assumes independent noise, so I treat memory as information I may have left on the table. I match the response to the shape and verify that the fix helps the horizons that matter."
Memory in $e_t$ = leftover information. Shape of the ACF: geometric from lag 1 = momentum; spikes at 7, 14, 21 = weekly shape; slow decay = missing slow component; long waves = another cycle.
Responses: (1) richer seasonal / trend / holiday terms, (2) AR residual term $\epsilon_t = \varphi\epsilon_{t-1} + \eta_t$ (correction $\varphi^h e_T$), (3) state-space (drifting components), (4) Gaussian process.
Trap: a flat ACF after an AR term does not prove the cause was momentum; check horizons.
Quick check: the residual ACF has $r_1 = -0.45$, $r_2 \approx 0$, and the time plot is jagged (zig-zag). What does a negative lag-1 correlation suggest?
Over-correction: a big positive miss is followed by a negative one, as if the model chases its own tail. Typical causes are a feature or smoother that over-reacts to yesterday, a differencing step applied twice, or residuals taken from a series that was already differenced. It is not the usual "momentum" case (that is positive). Durbin–Watson would be well above 2.
Giving the noise a memory: AR errors, state-space models and Gaussian processes
Three pictures of a wobbling quantity.
A rubber band (AR error): pull a ball tied to a post with a rubber band. Each day a random gust pushes it, and the band pulls it back a little. Today's position is "a fraction $\varphi$ of yesterday's position plus a fresh push". The memory fades geometrically.
A drifting buoy (state-space, local level): a buoy moves with the water by small random pushes and never returns to a fixed point. You do not see the buoy directly, only a noisy photograph of it every day. The memory is permanent: wherever the buoy has drifted to, it stays.
A random smooth curve (Gaussian process): choose a curve at random from all curves that change smoothly over about $\ell$ days. Days that are close in time get similar values, days far apart do not.
Three ways to say it:
- Picture: a ball on a rubber band; a buoy and a blurry photo; a random smooth ribbon.
- Numbers: with $\varphi = 0.7$ the correlation at lags 1, 2, 3 is $0.70, 0.49, 0.34$. A smooth curve with correlation length $\ell = 2.8$ days has exactly the same numbers.
- Slogan: an AR(1) is a Gaussian process in disguise; the buoy is the version whose memory never fades.
Take $\varphi = 0.7$ and compare the first two descriptions, then add the buoy.
- AR(1): $\epsilon_t = 0.7\,\epsilon_{t-1} + \eta_t$. Correlation at lag $k$: $0.7^k = 0.70,\ 0.49,\ 0.343,\ 0.240$.
- Gaussian process with the exponential kernel $k(\Delta) = \exp(-|\Delta|/\ell)$. Matching lag 1: $\exp(-1/\ell) = 0.7$ gives $\ell = -1/\ln 0.7 = 2.80$ days.
- Check lag 3: $\exp(-3/2.80) = 0.343 = 0.7^3$. The two agree at every integer lag: $\varphi^k = e^{-k/\ell}$ when $\varphi = e^{-1/\ell}$.
- Local level: $y_t = \mu_t + \varepsilon_t$, $\mu_t = \mu_{t-1} + \eta_t$. Let $q = \sigma_\eta^2/\sigma_\varepsilon^2 = 0.25$. The differences $\Delta y_t$ have lag-1 correlation $-1/(2 + q) = -1/2.25 = -0.444$ and nothing after that. The best forecast is the filtered level; it is the same as exponential smoothing with $\alpha = (\sqrt{q^2 + 4q} - q)/2 = (\sqrt{1.0625} - 0.25)/2 \approx 0.39$ (Chapter 7.18 shows why).
| AR(1) residual term | State-space (local level) | Gaussian process | |
|---|---|---|---|
| Model | $\epsilon_t = \varphi\epsilon_{t-1} + \eta_t$, $\eta_t \sim N(0, \sigma_\eta^2)$ | $y_t = \mu_t + \varepsilon_t$, $\mu_t = \mu_{t-1} + \eta_t$ (hidden state $\mu_t$) | $\epsilon(t) \sim GP(0, k)$, e.g. $k(t, t') = \sigma_f^2 e^{-\lvert t - t'\rvert/\ell}$ |
| Does memory fade? | Yes: $\varphi^k$ (stationary if $\lvert\varphi\rvert \lt 1$) | No: the level keeps wherever it drifted (random-walk-like) | Yes, at a rate set by the kernel and $\ell$ |
| Forecast | $\varphi^h e_T$ added to the calendar forecast | The current filtered level, flat; interval variance grows about linearly in $h$ | Mean reverts to the calendar line; variance grows to the prior variance |
| Parameters | $\varphi$, $\sigma_\eta$ | $\sigma_\eta$, $\sigma_\varepsilon$ (their ratio $q$ is the "speed") | kernel type, $\ell$, $\sigma_f$, noise |
| Cost | $O(n)$ (a recursion; scan) | $O(n)$ (Kalman filter) | $O(n^3)$ for $n$ days (exact) |
| Best when | short-term momentum around the calendar curve | the level or season itself drifts | smooth wandering, irregular times, small $n$, flexible shapes |
For equally spaced times an AR(1) is the exponential-kernel Gaussian process with $\varphi = e^{-1/\ell}$ (the process is also known as the Ornstein–Uhlenbeck process). A smoother kernel (squared-exponential, Matérn) or a periodic kernel gives memory shapes an AR(1) cannot. Higher-order memory is covered by AR($p$) and ARMA terms, or by a sum of kernels. Richer seasonal terms (the fourth response of the previous section) are not a memory model at all: they fix the mean.
For a Bayesian model any of the three is just another component with priors: a prior for $\varphi$ on $(-1, 1)$ (or $(0, 1)$), for $\sigma_\eta/\sigma_\varepsilon$, or for $\ell$ and $\sigma_f$.
Why do we need it?
When the residual ACF shows true momentum, ignoring it makes short-term forecasts worse and posteriors too narrow. These three descriptions are the standard ways to add memory while keeping everything else in the model, and they differ in whether the memory fades, how fast the model runs and what kind of shapes it allows.
Where is it used?
ARIMA errors in regression (SARIMAX), structural time series and the Kalman filter (UnobservedComponents), Gaussian-process regression (scikit-learn, GPyTorch, NumPyro's GP examples), and hybrid models such as a Prophet-style trend with an AR term.
How is it used?
Add the term to the mean function's noise: a recursion over days (AR), a filtered hidden state (state-space), or a kernel matrix (GP). Give each parameter a prior, refit, and check that the residual ACF is flat and short-horizon accuracy improved on rolling origins.
"A Gaussian process is non-parametric, so it makes no assumptions."
The kernel is a strong assumption: how smooth the curve is, whether the pattern is stationary, the length scale $\ell$. A wrong kernel gives confident wrong answers; the length is usually learned by maximising the marginal likelihood or with a prior.
"A state-space model and an ARIMA model are different things."
They overlap: many ARIMA and exponential-smoothing models have a state-space form, and the Kalman filter fits them. The state-space view adds flexibility: time-varying parameters, missing days, extra components.
"An AR(1) with $\varphi = 0.99$ is just a strong AR(1)."
It is almost a random walk: shocks fade over about 100 days, so it competes with the trend and the changepoints for the same slow wiggles (identifiability, Chapter 6.8). Put a prior on $\varphi$ that respects the horizon you care about, and check trend and AR terms together.
If the residual ACF of your forecasting model shows lag-1 momentum and the weekly and trend checks are clean, the smallest extension is an AR(1) error: $\mu_t$ stays as is, and the likelihood sees $y_t$ given $y_{t-1}$'s residual. In NumPyro the recursion over days is written with lax.scan; give $\varphi$ a prior inside $(-1, 1)$. Two cautions from the earlier chapters: this adds a sequential loop to the model, which matters for speed under JIT (Chapter 6.17), and a $\varphi$ near 1 trades off with the trend and the changepoint slopes $\delta_j$. A Gaussian process for the residual is neat for small data but its $O(n^3)$ cost grows fast.
AR(1): $\epsilon_t = \varphi\epsilon_{t-1} + \eta_t$, ACF $\varphi^k$, forecast $\varphi^h e_T$, cost $O(n)$. GP with exponential kernel $e^{-|\Delta|/\ell}$ = AR(1) with $\varphi = e^{-1/\ell}$; $O(n^3)$.
Local level: $y_t = \mu_t + \varepsilon_t$, $\mu_t = \mu_{t-1} + \eta_t$; permanent memory; $\alpha = (\sqrt{q^2+4q} - q)/2$ with $q = \sigma_\eta^2/\sigma_\varepsilon^2$.
Trap: the kernel is an assumption; $\varphi$ near 1 competes with the trend; richer seasonal terms fix the mean, they are not memory.
Quick check: an AR(1) residual term has $\varphi = 0.5$. By how much has the effect of today's surprise faded after 3 days, and what $\ell$ does the equivalent Gaussian process have?
After 3 days the surprise is multiplied by $0.5^3 = 0.125$ (it has faded to 12.5%). The equivalent kernel length is $\ell = -1/\ln 0.5 = 1/0.693 \approx 1.44$ days. The exponential kernel at 3 days gives $e^{-3/1.44} = 0.125$, the same.
The full diagnostic panel: seven worlds, six plots, one reading core
A doctor does not decide from one number. She takes the temperature, listens to the heart, looks at the throat, and then reads the whole picture. A residual panel does the same for a model: six small plots drawn together, so that one problem shows up in several of them and a false alarm in only one.
The panel in this section has: residual vs time, residual vs fitted, histogram with KDE, Q-Q plot, residual ACF and the spread by fitted bin. You choose what is wrong with the model (or nothing), look at the six plots and the numbers, and then read the shape. Switch on the matching fix and look again: the loop diagnose, change one thing, refit, diagnose is the real workflow.
Three ways to say it:
- Picture: six plots, one per question; a real problem leaves fingerprints in more than one.
- Numbers: AR errors give $r_1 \approx 0.7$, Durbin–Watson $\approx 2(1 - 0.7) = 0.6$, a tiny Ljung–Box p-value, and a normal-looking histogram.
- Slogan: each misspecification has its own fingerprint; learn the fingerprints, not the plots.
Reading a panel from its numbers. A model has these residual numbers (numbers like the ones you will see for the AR-errors world): mean $0.00$, sd $8.0$, $1.4826\times$MAD $= 8.2$, skewness $-0.5$, excess kurtosis $0.1$, Ljung–Box(14) $p \lt 0.001$, $r_1 = 0.71$, $r_7 \approx 0.03$, Durbin–Watson $0.58$, spread ratio $1.2$, largest $|z| = 3.5$, mean error over the last 28 days $+2.4$.
- Centred? The mean is 0 (guaranteed in-sample). The holdout mean $+2.4$ is small next to $\text{sd}/\sqrt{28}\approx 1.5$ times a few: not alarming.
- Constant spread? Spread ratio $1.2$: close to 1. Fine.
- Right shape? The sd (8.0) and the robust sd (8.2) agree, skewness and kurtosis are near 0, and the largest $|z|$ is 3.5 in 280 days (a Normal expects about 0.76 days beyond 3). The shape is acceptable.
- Pattern? Nothing against the fitted value; the time plot shows long runs above and below zero.
- Memory? $p \lt 0.001$ and $r_1 = 0.71$ with $r_7 \approx 0$. Check the Durbin–Watson identity: $2(1 - 0.71) = 0.58$, matching the reported $0.58$. A fast geometric decay at lag 1 and nothing at lag 7: short-term momentum.
- Reading: the calendar part looks fine; the noise has memory. Try an AR(1) term. The forecast correction would be $0.71^h e_T$: $0.71^7 \approx 0.09$, so it matters for about a week.
A symptom-to-action map for the six plots (everything in this chapter in one table):
| Where it shows | Symptom | Likely cause | First response |
|---|---|---|---|
| Time plot, ACF | Runs; slow ACF decay; recent holdout bias | Missed slope or level change | Changepoints near it, a weaker shrinkage, a regressor |
| Time plot, ACF | Waves; positive/negative swings in the ACF | Missed cycle | Another seasonal period, a driver |
| ACF | Spikes at 7, 14, 21 | Weekly shape not captured | Higher weekly order, weekday terms |
| ACF, Durbin–Watson | Fast geometric decay | Short-term momentum | AR(1) residual term, state-space, GP |
| Fitted plot, spread bars | Funnel; spread ratio $\gg 1.5$ | Heteroscedasticity | Transform, NB, variance model |
| Time plot, Q-Q | Huge residuals at regular spacing | Missed events | Holiday / event columns |
| Q-Q, histogram | S-shape, high kurtosis, irregular wild days | Heavy-tailed noise | Student-t likelihood |
| Histogram, Q-Q | Two humps; one-sided tail | Hidden group; skew | Find the indicator; log or skewed likelihood |
Why do we need it?
No single plot finds every problem, and each plot raises false alarms of its own. Looking at the six together, with the numbers, separates a real fingerprint from a fluke and tells you which part of the model to change.
Where is it used?
Forecast model reviews, ARIMA diagnostics (plot_diagnostics), regression checking, and the validation step before a forecasting system goes to production, where the same panel is recomputed on new data to watch for drift (Chapter 7.19).
How is it used?
After each fit make the panel for the training days and for a rolling-origin holdout. Read the shape with the table, change one thing in the model, refit and compare. Stop when nothing is left that you could exploit; then move on to accuracy and calibration.
"All six plots look fine, so the model is good."
It means no visible structure is left in this sample and these views. An over-flexible model also has beautiful residuals (it memorised them). The final judge is accuracy and calibration on rolling-origin holdouts (Chapter 7.15, 7.16).
"One panel looks odd, so I must change the model."
With six plots and several numbers, something will look odd in a good model about as often as chance says. Compare with a lineup of simulated residuals and fix only what shows up in more than one place and matters for forecasts.
"The panel told me to add a term, and the panel is clean after, so it is the right term."
Different fixes can silence the same plot (an AR term and a changepoint both flatten a lag-1 ACF). Prefer the fix with a story you can explain, and confirm on a holdout.
"The residual checks use in-sample data only."
Do them on the training days and on forecasts made from rolling origins; problems that appear only out of sample (a level change near the end) are the ones that hurt.
A practical routine for your forecasting model: (1) fit with SVI and take the mean function averaged over the guide's samples; (2) build the six plots for the training days, with the standardized (Pearson or quantile) residuals if the likelihood is a Negative Binomial; (3) repeat on the held-out days of a few rolling origins, where the time plot and the holdout mean matter most; (4) read the shapes with the table: runs and steps point at the trend (changepoint grid, PELT range, Laplace scale $b$), waves at Fourier order or a missing period, spikes on repeating dates at the holiday table, a funnel at the likelihood or a variance model, S-shaped Q-Q at Student-t; (5) change one thing, refit, re-check. Keep the panel as a standing report for the monitoring in Chapter 7.19.
"We checked the model: the p-value was above 0.05."
"We checked that the residuals have no exploitable structure: they are centred overall and within groups, the spread is constant against the fitted value and time, the shape matches the likelihood, and the ACF and Ljung–Box show no memory at the weekly lags. We did it on the training days and on rolling-origin holdouts, and we compared the plots with residuals of simulated data."
Model answer: "Residual diagnostics tell me whether my noise assumptions hold: independence, constant spread, the right shape and no leftover pattern. Each failure has a fingerprint and a fix: memory points at a missing component or an AR term, a funnel at a transform or a Negative Binomial, heavy tails at a Student-t. The residual panel can show that something is missing, but forecast accuracy and calibration decide whether the model is good."
Six plots: time, fitted, histogram+KDE, Q-Q, ACF, spread by fitted bin. Numbers: mean (0 in-sample), sd vs robust sd, skewness, excess kurtosis, Ljung–Box, DW, spread ratio, max $|z|$, holdout mean.
Fingerprints: AR = geometric $r_k$; weekly shape = spikes at 7, 14, 21; missed change = slow ACF decay + holdout bias; hetero = funnel; heavy tails = S-shaped Q-Q; events = huge residuals at regular spacing.
Loop: fit → panel → read → change one thing → refit. Trap: a clean panel is not a proof; accuracy and calibration decide.
Quick check: Q-Q plot S-shaped, Ljung–Box fine, spread ratio 1.1, and the five biggest residuals are on days 20, 72, 124, 176, 228. What would you do?
The days are 52 apart: events at regular spacing. Add an event/holiday column for those dates (with a window if needed) and refit; the S-shape and heavy-tail numbers should shrink. A Student-t would make the model tolerate the spikes instead of explaining them, and could not predict them next time.
Recap, cheat sheet and practice
- A residual is $e_t = y_t - \hat y_t$: the true noise plus whatever the model got wrong. State which prediction you used: in-sample fit, holdout forecast, posterior mean or a posterior draw. The goal is no exploitable structure.
- Five questions: centred (also inside groups and on a holdout; in-sample mean is 0 by construction), constant spread, right shape, no pattern against fitted values, time and known columns, no memory.
- Shape: histogram, KDE and Q-Q plot of standardized residuals decide the likelihood. S-shape = heavy tails (check the dates first, then a Student-t); one bent end = skew; a kink = two groups.
- Plots: bowl = too straight, funnel = changing spread, waves = missed cycle, step = missed level shift, spikes at regular spacing = missed events. Always make both the fitted and the time plot.
- Memory: residual ACF with the $\pm1.96/\sqrt n$ band, Ljung–Box $Q(h)$ with $h$ reaching the season (14 for daily data with a weekly pattern), Durbin–Watson $\approx 2(1 - r_1)$ (lag 1 only), ACF of $e^2$ for volatility clustering.
- Heteroscedasticity: detect with the funnel, the spread ratio ($\gg 1.5$) and Breusch–Pagan as a second opinion; respond with a transform (log, sqrt), a likelihood with mean-dependent variance (Poisson, NB: $\mu + \mu^2/\alpha$) or an explicit variance model ($\sigma_t = c\mu_t$, WLS). It biases the intervals, not the mean.
- Counts: raw residuals of a correct count model fan out; use Pearson residuals (dispersion near 1) and randomized quantile residuals ($N(0,1)$ if the model is right; the posterior-predictive PIT in a Bayesian model).
- Autocorrelated residuals: read the ACF shape (geometric from lag 1 = momentum; spikes at 7, 14, 21 = weekly shape; slow decay = missing slow component). Responses: richer seasonal / trend / holiday terms first; then an AR residual term (correction $\varphi^h e_T$), a state-space model, a Gaussian process. AR(1) = exponential-kernel GP with $\varphi = e^{-1/\ell}$.
- Calibrate your eye with replicated residuals (the lineup, or a posterior predictive check of a statistic). A clean panel is not a proof; accuracy and calibration on rolling origins decide.
Cheat sheet
| Question | What to compute or plot | Good looks like | If not | Python |
|---|---|---|---|---|
| 1 · Centred? | group means (weekday, holiday, fitted bin), holdout mean | inside $\pm2\,SE$ | missing term for that group | e.groupby(...).mean() |
| 2 · Constant spread? | residual vs fitted, spread ratio, $\lvert e\rvert$ vs fitted | even band, ratio near 1 | log / sqrt, NB, variance model | het_breuschpagan(e, X) |
| 3 · Right shape? | skewness, excess kurtosis, Q-Q, days with $\lvert z\rvert \gt 3$ | straight Q-Q, about 0.27% beyond 3 | dates first; then Student-t, NB, transform | stats.skew, stats.kurtosis, sm.qqplot(e, line="q") |
| 4 · Any pattern? | $e$ vs time, vs fitted, vs each column | flat smoother at 0 | match the shape to the table | plt.scatter(fitted, e) |
| 5 · Any memory? | ACF to 2–3 seasons, $Q(h)$, DW, ACF of $e^2$ | inside the band, $p \gt 0.05$ | richer mean first; then AR / state-space / GP | acorr_ljungbox(e, lags=[14]), durbin_watson(e) |
| Counts | Pearson $(y-\mu)/\sqrt{V}$; randomized quantile $\Phi^{-1}(u)$ | variance about 1; $N(0,1)$ | NB / zero-inflation / missing term | res.resid_pearson |
| Real pattern or noise? | lineup; PPC on $Q$, $r_1$, max $\lvert z\rvert$ | real plot not special | pattern is real: fix the model | simulate from the fit, refit, compare |
import numpy as np
import statsmodels.api as sm
from scipy import stats
from statsmodels.stats.diagnostic import acorr_ljungbox, het_breuschpagan
from statsmodels.stats.stattools import durbin_watson
from statsmodels.tsa.stattools import acf
rng = np.random.default_rng(7)
n = 280
t = np.arange(n)
def fourier(t, period, order):
cols = []
for k in range(1, order + 1):
cols += [np.sin(2 * np.pi * k * t / period), np.cos(2 * np.pi * k * t / period)]
return np.column_stack(cols)
# --- 1) data: trend + weekly pattern + AR(1) noise (phi = 0.7); the model below will assume INDEPENDENT noise ---
noise = np.zeros(n)
for i in range(1, n):
noise[i] = 0.7 * noise[i - 1] + rng.normal(0, 5.7)
y = 60 + 0.5 * t + 14 * np.sin(2 * np.pi * t / 7) + 6 * np.cos(2 * np.pi * t / 7) + noise
# --- 2) the "calendar regression": trend + weekly Fourier terms (order 2) ---
X = sm.add_constant(np.column_stack([t, fourier(t, 7, 2)])) # 6 columns
fit = sm.OLS(y, X).fit()
e = fit.resid
print(round(e.mean(), 6), round(e.std(ddof=X.shape[1]), 2)) # 0.0 7.51 mean is 0 by construction; residual sd
# --- 3) the five questions ---
# (1) centred inside groups? mean residual by weekday (all small: the weekly pattern is captured)
print(np.round([e[t % 7 == d].mean() for d in range(7)], 1)) # [ 0.3 -0.3 0.3 -0.1 -0. 0.2 -0.3]
# (2) constant spread? sd of the top third vs the bottom third of the fitted values, and Breusch-Pagan p
idx = np.argsort(fit.fittedvalues)
print(round(e[idx[-n // 3:]].std() / e[idx[:n // 3]].std(), 2), round(het_breuschpagan(e, X)[1], 3)) # 1.14 0.503
# (3) shape: skewness, excess kurtosis, robust sd = 1.4826 * MAD
print(round(stats.skew(e), 2), round(stats.kurtosis(e), 2), round(1.4826 * stats.median_abs_deviation(e), 2)) # -0.34 0.52 7.06
# (5) memory: ACF at lags 1, 2, 7; Ljung-Box at lag 14; Durbin-Watson vs 2(1 - r1)
r = acf(e, nlags=7)
print(np.round(r[[1, 2, 7]], 2)) # [0.71 0.52 0.16] fast geometric decay: momentum
print(acorr_ljungbox(e, lags=[14]).round(4)) # lb_stat 339.7229 lb_pvalue 0.0 -> memory left
print(round(durbin_watson(e), 2), round(2 * (1 - r[1]), 2)) # 0.58 0.59
# --- 4) fix: regression with AR(1) errors (feasible GLS); its one-step errors should look like white noise ---
glsar = sm.GLSAR(y, X, rho=1)
res = glsar.iterative_fit(maxiter=10)
print(np.round(glsar.rho, 2)) # [0.71] estimated phi
print(acorr_ljungbox(res.wresid, lags=[14], model_df=1).round(3)) # lb_stat 12.312 lb_pvalue 0.502 (df reduced by 1 AR parameter)
# --- 5) lineup / posterior-predictive style check: is the real Q(14) unusual under independent noise? ---
Q = acorr_ljungbox(e, lags=[14])["lb_stat"].iloc[0]
sigma = np.sqrt(fit.scale)
Qrep = []
for _ in range(500):
yrep = fit.fittedvalues + rng.normal(0, sigma, n) # data simulated from the fitted model
erep = sm.OLS(yrep, X).fit().resid # processed exactly like the real data
Qrep.append(acorr_ljungbox(erep, lags=[14])["lb_stat"].iloc[0])
print(round(Q, 1), round(np.mean(np.array(Qrep) >= Q), 3)) # 339.7 0.0 no simulated Q comes near the real one
# --- 6) counts: Pearson and randomized quantile residuals (Negative Binomial, alpha = 8) ---
nn = 210
tc = np.arange(nn)
mu_true = np.exp(np.log(20) + 0.0085 * tc + 0.25 * np.sin(2 * np.pi * tc / 7))
alpha = 8.0
yc = rng.negative_binomial(alpha, alpha / (alpha + mu_true)) # NumPy: n = alpha, p = alpha/(alpha + mu) -> mean mu, Var = mu + mu^2/alpha
Xc = sm.add_constant(np.column_stack([tc / 100, fourier(tc, 7, 1)]))
pois = sm.GLM(yc, Xc, family=sm.families.Poisson()).fit()
mu = pois.fittedvalues
print(round(np.mean(pois.resid_pearson ** 2), 1)) # 7.2 Pearson dispersion of the Poisson fit (should be about 1)
a_hat = 1 / np.mean(((yc - mu) ** 2 - mu) / mu ** 2) # method of moments for alpha
pear_nb = (yc - mu) / np.sqrt(mu + mu ** 2 / a_hat)
print(round(a_hat, 1), round(np.mean(pear_nb ** 2), 2)) # 7.8 0.98 alpha-hat near 8; dispersion near 1
d = stats.nbinom(a_hat, a_hat / (a_hat + mu)) # SciPy: nbinom(n=alpha, p=alpha/(alpha + mu))
u = d.cdf(yc - 1) + rng.uniform(size=nn) * (d.cdf(yc) - d.cdf(yc - 1))
rq = stats.norm.ppf(u) # randomized quantile residuals
print(round(rq.mean(), 2), round(rq.std(), 2), round(stats.skew(rq), 2)) # 0.0 0.96 0.18 close to N(0, 1)
1. A least-squares model with an intercept has a residual mean of exactly 0.00 on its training days. What does that tell you?
2. The residual-vs-fitted plot of a daily revenue model fans out to the right. What is the most natural first response?
3. A residual ACF has spikes at lags 7, 14, 21 (decaying) and nothing at lags 1 to 6. The most likely reading?
4. A Negative Binomial model has $\hat\alpha = 4$ and predicts $\hat\mu = 40$ for a day with 70 orders. The Pearson residual is about
5. Daily residuals with a weekly pattern pass a Ljung–Box test at $h = 10$ with $p = 0.40$. What can you conclude?
6. In a lineup of six residual plots (five simulated from your fitted model, one real) you cannot tell which is the real one. The best conclusion?
Practice problems
A. Actuals $20, 22, 19, 25, 27, 24$ and model values $21, 21, 21, 24, 25, 25$. Compute the residuals, their mean and root mean square, $r_1$ and the Durbin–Watson number, and say what the sign of $r_1$ suggests.
Residuals $y - \hat y$: $-1, +1, -2, +1, +2, -1$. Sum $= 0$, mean $0$. Squares: $1, 1, 4, 1, 4, 1$, sum $12$, mean $2$, RMS $\sqrt2 \approx 1.41$.
Lag-1 products: $(1)(-1) + (-2)(1) + (1)(-2) + (2)(1) + (-1)(2) = -1 - 2 - 2 + 2 - 2 = -5$. $r_1 = -5/12 \approx -0.42$. Differences: $2, -3, 3, 1, -3$; squares $4, 9, 9, 1, 9 = 32$; $DW = 32/12 \approx 2.67$, close to $2(1 - r_1) = 2.83$ (end effects).
A negative lag-1 correlation (zig-zag) means a big positive miss is followed by a negative one: the model over-reacts to the previous day (or the residuals come from an over-differenced series). With six values this is only a hint.
B. Revenue has sd equal to 10% of its level. On quiet days the level is 100 and on busy days 400 (half each). A constant-$\sigma$ Normal builds 90% intervals. Find the coverage on each kind of day and propose a fix.
True sds: $10$ and $40$. Pooled $\hat\sigma = \sqrt{(100 + 1600)/2} = \sqrt{850} \approx 29.2$. The 90% half-width is $1.645 \times 29.2 \approx 48.0$.
Quiet days: $48.0/10 = 4.8$ sd, coverage $2\Phi(4.8) - 1 \approx 100\%$ (far too wide). Busy days: $48.0/40 = 1.20$ sd, coverage $2\Phi(1.20) - 1 \approx 77\%$ (far below 90%).
Fix: model $\log y$ (a constant 10% error becomes sd about $0.10$ on the log scale) and back-transform: interval $= \text{median} \times e^{\pm1.645\times0.10} = \text{median} \times [0.848,\, 1.179]$. Coverage is then about 90% on both kinds of day. Alternatives: a Gamma-type likelihood or $\sigma_t = 0.10\mu_t$.
C. A count model predicts $\hat\mu = 25$ with $\hat\alpha = 5$ (Negative Binomial) and you observe 40. Give the Pearson residual under the NB and under a Poisson, and the randomized-quantile residual if $u$ is taken at the midpoint of $[F(39), F(40)]$.
NB: variance $= 25 + 625/5 = 150$, sd $= 12.25$, Pearson $= 15/12.25 \approx 1.22$. Poisson: sd $= 5$, Pearson $= 3.0$.
The NB CDF gives $F(39) = 0.8783$ and $F(40) = 0.8902$; the midpoint is $0.8843$ and $\Phi^{-1}(0.8843) \approx 1.20$, close to the NB Pearson value. The Poisson assumption would flag this ordinary day as a 3-sigma event; if many days do that, the Poisson variance is too small.
D. $n = 200$ daily residuals have $r_1 = 0.12$, $r_7 = 0.38$, $r_{14} = 0.17$ and the other lags near 0. Which bars are outside the band, what is the Ljung–Box statistic at $h = 14$ (use only these three lags) and what do you do?
Band $= 1.96/\sqrt{200} = 0.139$. Outside: $r_7 = 0.38$ and $r_{14} = 0.17$. $r_1 = 0.12$ is inside.
$Q = 200\cdot202\,(0.12^2/199 + 0.38^2/193 + 0.17^2/186) = 40\,400\,(0.0000724 + 0.000748 + 0.000155) = 40\,400 \times 0.000976 \approx 39.4$ on 14 degrees of freedom: $p \approx 0.0003$.
The memory sits at the weekly lags with geometric decay ($0.38^2 = 0.14 \approx 0.17$): the weekly pattern is not captured. Raise the weekly Fourier order or add weekday-specific terms (and check holidays falling on specific weekdays); re-run the test with $h = 14$ afterwards.
E. Interview: "Our forecast residuals have a Ljung–Box p-value below 0.001, so the Prophet-style model is wrong." Respond.
"A tiny p-value says there is memory left in the residuals; it does not say what. The model assumes independent noise given its components, so I read the ACF: fast geometric decay at lag 1 is short-term momentum, spikes at 7, 14, 21 mean the weekly shape is not captured, and slow decay with a wandering time plot means a trend change was missed. I fix the mean function first (Fourier order, changepoints, holidays), and add an AR(1) residual term only if momentum remains. Then I re-check the residuals and, more importantly, the accuracy by horizon and the interval coverage on rolling origins. So the model is incomplete in a specific way, not wrong in general."
F. Interview: "How do you check the noise assumption of your forecasting model, and what do you do if it fails?"
"I compute residuals from the posterior-mean prediction on the training days and on rolling-origin holdouts and ask five questions: are they centred inside groups, is the spread constant against the fitted value and time, does the shape match the likelihood (Q-Q plot, tail counts), is there any pattern against time or other columns, and is there memory (ACF, Ljung–Box at the seasonal lag). For counts I use Pearson and randomized quantile residuals. If the Q-Q plot is S-shaped I first check whether the wild days share dates (missing holidays) before moving to a Student-t. A funnel gets a transform, a Negative Binomial or a variance model. Memory gets a richer seasonal or trend term first, and an AR term if momentum remains. I compare suspicious plots with residuals of data simulated from the model so I do not chase noise."
Model complexity and forecasting alternatives
Every forecasting model is a bet: a bet on how complicated the world is, and on what kind of pattern will repeat. This chapter does two jobs. First, the complexity knobs: how flexible should a model be, and where do the knobs of your own model (Fourier order, number of changepoints, Laplace scale, guide rank, pooling strength) sit between "too stiff" and "too wiggly"? Second, the alternatives: ARIMA, exponential smoothing, state-space models, Gaussian processes, gradient boosting and neural forecasters, each with the assumption it makes and the situations in which it beats a Prophet-style model, and the ones in which it does not. The skill interviewers test is one question: what assumption does my model make, and when would another model be better?
- Explain underfitting, overfitting, uncertainty and generalization with the bias–variance–complexity picture, and see why training error always falls while holdout error is U-shaped
- Place each knob of your model on that picture: Fourier order, number of changepoints, Laplace scale, guide covariance rank and pooling strength (and see that the guide rank is a different kind of knob)
- State the assumptions of a Prophet-style model and the symptom, diagnostic and alternative for each one that fails
- For ARIMA / SARIMA, exponential smoothing, state-space models, Gaussian processes, gradient boosting and neural forecasting (N-BEATS, TFT: awareness only): the core assumption, when it beats a Prophet-style model, and when it does not
- Use a small model chooser and a simulated bake-off to see that different worlds crown different winners, and that only a rolling-origin comparison on your data decides
- Answer, with a short model answer, "What assumption does my model make, and when would another model be better?"
What we need from earlier chapters: bias and variance of estimators (Chapter 5.1); pooling and shrinkage (Chapter 6.6); guide families and their parameter counts (Chapter 6.13); exponential smoothing and baselines (Chapter 7.5); the ARIMA family (Chapter 7.6); the additive model, changepoints, the Laplace prior and the Fourier order (7.7, 7.8, 7.10, 7.11); rolling-origin evaluation (Chapter 7.15); and residual diagnostics (Chapter 7.17).
Bias, variance and complexity: stiff versus wiggly core
A tailor has two ways to fail. With a stiff template (one standard size) every suit is wrong in the same way for the same kind of customer: that is bias. With a wet noodle for a measuring tape that follows every wrinkle in today's posture, the suit fits today perfectly and tomorrow badly, and a different day gives a very different suit: that is variance. Good tailoring sits in between.
A forecasting model is the same. A model that is too simple (a straight line through a seasonal series) misses patterns that are really there: it underfits. A model that is too flexible (a curve that passes through every noisy point) learns the noise as if it were pattern: it overfits. The forecast then depends heavily on which noise happened to be in the training window, and its stated uncertainty is overconfident. Generalization means doing well on data the model has not seen, which for forecasting means the future.
Three ways to say it:
- Picture: thirty curves fitted to thirty noisy copies of the same history. Stiff model: the curves lie on top of each other but away from the truth. Wiggly model: they surround the truth but fan out.
- Numbers: with 56 days and noise sd 6, a model with 7 parameters has an expected squared error of 40.5; with 25 parameters, 52.1; with 1 parameter, 321.6.
- Slogan: simple models are wrong the same way every time; flexible models are wrong a different way every time.
The true seasonal pattern repeats every $P = 28$ days (a four-week cycle) and is built from three harmonics (a harmonic of amplitude $A$ carries "energy" $A^2/2$: $232$, $40.5$ and $12.5$; see the widget). We have $n = 56$ days (two full periods) of data with noise sd $\sigma = 6$, so $\sigma^2 = 36$. We fit a Fourier model of order $N$ (Chapter 7.11): $p = 2N + 1$ parameters. The expected squared error for predicting a new noisy day is bias$^2$ + variance + $\sigma^2$.
- Bias$^2$: the energy of the harmonics the model cannot draw. Order 0 draws none: $232 + 40.5 + 12.5 = 285$. Order 1 misses harmonics 2 and 3: $53$. Order 2 misses harmonic 3: $12.5$. Order 3 or more can draw everything: $0$.
- Variance: for least squares it is $\sigma^2\,p/n = 36\,p/56$. Order 0 ($p = 1$): $0.64$. Order 3 ($p = 7$): $4.5$. Order 6 ($p = 13$): $8.36$. Order 12 ($p = 25$): $16.07$.
- Test error $=$ bias$^2$ + variance $+ 36$: order 0: $321.6$; order 1: $90.9$; order 2: $51.7$; order 3: $\mathbf{40.5}$; order 6: $44.4$; order 12: $52.1$. A U-shape with the minimum at the true order 3.
- Training error is about bias$^2 + \sigma^2(n - p)/n$: order 3: $31.5$; order 6: $27.6$; order 12: $19.9$. It keeps falling as the model gets more flexible (the model is fitting noise), which is why training error cannot choose the complexity.
Suppose the truth at a time $t$ is $f(t)$, we observe $y = f(t) + \epsilon$ with $Var(\epsilon) = \sigma^2$, and we fit a model many times on fresh noisy training data, giving predictions $\hat f(t)$. Then
$$E\big[(y - \hat f(t))^2\big] \;=\; \underbrace{\big(E[\hat f(t)] - f(t)\big)^2}_{\text{bias}^2} \;+\; \underbrace{Var(\hat f(t))}_{\text{variance}} \;+\; \underbrace{\sigma^2}_{\text{noise, cannot be removed}}.$$- Bias: the average prediction is off from the truth because the model cannot represent the truth (too stiff). Variance: the prediction jumps around from one training sample to the next (too flexible). The same split for estimators is in Chapter 5.1.
- Complexity is the model's freedom: roughly the number of free parameters (the effective number when priors or penalties shrink them). Increasing complexity lowers bias and raises variance.
- Underfitting: error is high on training and holdout data (bias dominates). Overfitting: error is very low on training data but high on holdout data (variance dominates). Generalization gap: holdout error minus training error.
- Uncertainty: a flexible model fitted to little data also has wide, unstable posteriors and predictive intervals; a model that is too stiff gives intervals that look confident but are centred in the wrong place.
- You cannot see bias and variance on real data, because the truth is unknown. You estimate the U-curve with a holdout in time: rolling-origin evaluation (Chapter 7.15).
Why do we need it?
Every modelling choice (how many seasonal terms, how many changepoints, how strong a prior) trades stiffness against wiggliness. Knowing the trade-off turns "tuning" from guesswork into a search for the lowest holdout error, and explains why a model that fits history beautifully can forecast badly.
Where is it used?
Regularized regression (ridge, lasso), tree depth in gradient boosting, the number of ARIMA lags, the smoothing parameters of exponential smoothing, bandwidths in KDE and Gaussian processes, early stopping in neural networks, and every comparison of model sizes in forecasting.
How is it used?
Pick the complexity knob, vary it over a grid, measure the error on rolling-origin holdouts (not on training data), and choose near the lowest holdout error, preferring the simpler model when two are close. With priors, tune the prior scale instead of the parameter count.
"More parameters always make a better model."
More parameters always make the training error smaller. Holdout error falls and then rises, and the lowest point is where the model's flexibility matches the structure that is really in the data.
"Overfitting is only about too many parameters."
What matters is the effective flexibility. A model with 50 changepoints and a tight prior can be stiffer than one with 5 changepoints and no prior. The prior scale is a complexity knob too.
"Pick the complexity that gives the smallest error on all my data."
That is the training error again. Choose on time-ordered holdouts, and never tune on the final test period (Chapter 7.15).
Both of your projects are full of this trade-off. In the forecasting model, a high Fourier order or many weakly shrunk changepoints give a flexible mean function: it can memorise noise and extrapolate a wild last slope. In the A/B framework, the hierarchical model's pooling strength decides how much each segment's estimate is pulled toward the overall mean (the same bias–variance trade-off, Chapter 6.6). The next section places each knob on the U-curve.
Expected squared error $=$ bias$^2$ + variance + $\sigma^2$. Complexity ↑ → bias ↓, variance ↑. Training error always falls; holdout error is U-shaped.
Least squares with $p$ parameters, $n$ points: variance $\approx \sigma^2 p/n$; bias$^2$ $=$ the part of the truth the model cannot draw.
Trap: choose complexity on a time-ordered holdout, not on training error. Priors change the effective complexity.
Quick check: model A has training RMSE 3 and holdout RMSE 9; model B has training RMSE 7 and holdout RMSE 8. Which is overfitting, and which would you pick?
Model A: the gap between 3 and 9 is large, so it is overfitting (variance dominates). Model B has a small gap and the lower holdout error, so pick B. The training error alone would have chosen A, the wrong model.
Where each knob of your model sits: Fourier order, changepoints, Laplace scale, guide rank, pooling core
Picture a mixing desk with five sliders. Moving a slider changes how "free" the model is, but the sliders are not all the same kind. Two of them (Fourier order and number of candidate changepoints) set how much capacity the model has: how many shapes it can draw. Two others (the Laplace scale on the changepoint slopes and the pooling strength across groups) set how much of that capacity is actually used: they are strength-of-shrinkage knobs. The fifth (the rank of the guide's covariance) is different again: it does not change how well the model fits the data at all, only how faithfully the fitted posterior is approximated.
That distinction is practical. Prophet-style models usually give generous capacity (many candidate changepoints) and control the wiggliness with the prior scale. Adding capacity is cheap when the prior is strict; it is dangerous when the prior is loose.
Three ways to say it:
- Picture: capacity knobs open the door; strength knobs decide how far through it the data may push the model.
- Numbers: for eight groups, no pooling has squared error 1, full pooling has 1, and partial pooling has 0.5625.
- Slogan: tune the strength knobs; give the capacity knobs room; treat the guide rank as a cost-versus-accuracy choice.
Two calculations with exact numbers.
(a) Pooling strength. Eight groups (segments). Each group's true effect $\mu_g$ is drawn from $N(0, \tau^2)$ with $\tau^2 = 1$. We see 4 observations per group with noise sd 2, so each raw group mean has variance $v = 4/4 = 1$. The estimate is $\hat\mu_g = B\,\bar y_g + (1 - B)\,\bar y$ ($B$ = weight on the group's own mean, $\bar y$ = the mean of all group means).
- No pooling, $B = 1$: error $= v = 1$.
- Complete pooling, $B = 0$: every group gets the grand mean. Error $= v/8 + \tau^2(1 - 1/8) = 0.125 + 0.875 = 1.0$.
- Partial pooling, $B = 0.5$: error $= B^2 v + (1 - B)^2\,(1.0) + 2B(1 - B)\,v/8 = 0.25 + 0.25 + 0.0625 = 0.5625$.
- The best weight is $B^* = \tau^2/(\tau^2 + v) = 1/2$ (Chapter 6.6). So the middle setting has $44\%$ less squared error than either extreme: a U-curve in the pooling knob, with minimum at the right amount of shrinkage.
(b) Guide rank. A model with $d = 40$ latent parameters (Chapter 6.13).
- Mean-field (diagonal) guide: $2d = 80$ numbers.
- Full-rank Normal guide: $d + d(d+1)/2 = 40 + 820 = 860$ free numbers (a mean and the lower triangle of a Cholesky factor). NumPyro's
AutoMultivariateNormalstores a full $d \times d$ array whose upper triangle is fixed at zero, so the stored arrays hold $40 + 1\,600 = 1\,640$ numbers, of which 860 are free. - Low-rank guide of rank $r$: $d(r + 2)$ numbers (a mean $d$, a $d\times r$ factor, a diagonal $d$). Rank 5: $40 \times 7 = 280$ (33% of full-rank). Rank 20: $40\times22 = 880$, more than full-rank. A low-rank guide only saves anything while $r \lt (d - 1)/2$, that is below about $d/2$.
| Knob | "More flexible" means | Kind | Too stiff looks like | Too loose looks like | How to tune |
|---|---|---|---|---|---|
| Fourier order $N$ (7.11) | higher $N$: $2N$ columns | capacity | seasonal shape too smooth, residual waves at the seasonal lags | seasonal curve chases noise; unstable with short history | holdout error vs $N$; residual ACF |
| Candidate changepoints $C$ (7.8) | more candidates | capacity (ceiling) | cannot bend where the data bends | harmless if $b$ is small; wiggly trend if $b$ is large | generous $C$, then tune $b$ |
| Laplace scale $b$ (7.10) | larger $b$: slope changes less shrunk | strength of shrinkage | trend ignores real slope changes (underfit) | trend chases noise, wild last slope | rolling-origin error vs $b$; prior predictive and sensitivity checks |
| Pooling strength ($\tau$ small = strong, 6.6) | larger $\tau$: groups less alike | strength of shrinkage | real group differences erased | small groups get noisy estimates | estimate $\tau$ from the data (hierarchical prior), check by group |
| Guide rank $r$ (6.13) | larger $r$: more correlations captured | posterior approximation quality, not data fit | posterior too narrow along correlated directions | no overfit, but memory and time grow like $d\,r$ | compare with a richer guide or NUTS; weigh cost |
The first four knobs sit on the bias–variance curve of the mean function: the more freedom, the lower the training error, and the holdout error is U-shaped. The guide rank does not have a U-curve: the family of rank-$r$ guides is nested (a rank-$r$ guide contains every rank-$(r-1)$ guide), so with perfect optimisation a bigger rank can only approximate the true posterior as well or better. The price is parameters, memory and a harder optimisation problem. This is why it is a cost-versus-accuracy choice rather than a generalization choice.
Why do we need it?
A model with five knobs has too many settings to guess. Knowing which knobs control the bias–variance trade-off, which are shrinkage strengths and which only affect cost lets you tune two or three of them on a holdout and leave the rest generous or fixed.
Where is it used?
Prophet's changepoint_prior_scale, n_changepoints and seasonality_prior_scale; fourier_order in every Fourier-based seasonality; the group-level scale of hierarchical models; and the choice between mean-field, low-rank and full-rank guides in NumPyro (AutoNormal, AutoLowRankMultivariateNormal, AutoMultivariateNormal).
How is it used?
Fix generous capacity (a moderate Fourier order, many candidate changepoints), tune the strength knobs on rolling-origin holdouts, and check the result with posterior predictive checks. Choose the guide rank by comparing a few ranks (or NUTS) on the posterior quantities you care about, and by what the model size makes affordable.
"More candidate changepoints always overfit."
With a strict prior (small $b$) the extra candidates are pulled to zero unless the data demand them, so a generous grid is safe. With a loose prior, the same grid chases noise. The knob that controls wiggliness is the prior scale.
"Guide rank is an overfitting knob: a smaller rank regularizes."
It does not regularize the model. A rank that is too small gives a posterior approximation that is too narrow along correlated directions (overconfident), and a larger rank only improves the approximation, at higher cost.
"Low-rank is always cheaper than full-rank."
$d(r + 2)$ parameters beat $d + d(d+1)/2$ only while $r \lt (d - 1)/2$ (for $d = 40$: ranks up to 19). Beyond that the low-rank guide has more parameters and no advantage.
"Tune all five knobs by grid search."
The knobs interact (a loose $b$ with many changepoints is a different model from a tight $b$ with few). Fix generous capacity, tune one or two strength knobs on rolling origins, and check the rest with posterior predictive and sensitivity checks.
In the forecasting model, $N$ (Fourier order), the grid of candidate changepoints (from the Prophet-like grid and PELT) and the Laplace scale $b$ on $\delta_j \sim Laplace(0, b)$ are the capacity and strength knobs of the mean function; in the A/B framework the hierarchical scale across groups plays the same role as pooling strength. The guide choice (a full-rank or a low-rank Gaussian guide based on model size) is the fifth kind of knob: it is about the cost of approximating the posterior, not about fitting the data. If your code's rule picks the low-rank guide for large models, that is exactly the "past a certain $d$ full-rank is too expensive" logic from the table; check the chosen rank against a larger rank on one model if the posterior correlations matter for your decision.
Capacity knobs: Fourier order $N$ ($2N$ columns), candidate changepoints $C$. Strength knobs: Laplace scale $b$ (small = stiff), pooling $\tau$ (small = strong pooling). Cost knob: guide rank $r$, parameters $d(r + 2)$ vs full-rank $d + d(d+1)/2$.
First four: training error falls, holdout error is U-shaped. Guide rank: nested family, no U-curve, diminishing returns.
Trap: a generous $C$ with a small $b$ is safe; a generous $C$ with a large $b$ is not. Tune on rolling origins.
Quick check: a model has $d = 100$ latent parameters. How many numbers does a rank-10 low-rank guide have, and how many a full-rank guide?
Low-rank: $d(r + 2) = 100 \times 12 = 1\,200$. Full-rank: $d + d(d+1)/2 = 100 + 5\,050 = 5\,150$. The low-rank guide stores about 23% as many numbers. (Mean-field: $2d = 200$.)
The key question: what does my model assume, and when would another be better? core
Every map leaves things out. A road map assumes roads do not move; a hiking map assumes you care about hills. Before you trust a map you ask three things: what does it assume, when does it fail, and what other map would I use then?
A forecasting model is a map of a time series, and a Prophet-style model is one particular map: it draws the series as a trend with a few bends, plus a repeating seasonal curve, plus holiday bumps, plus regressors, plus independent noise. That map is excellent when the world really looks like that. It is the wrong map for a series whose level jumps without warning, whose seasonal pattern slowly changes shape, or whose misses carry over from day to day.
Three ways to say it:
- Picture: a sticky note on every model with the words "I assume that…".
- Numbers: a Prophet-style model makes about eight assumptions; each one, when false, leaves a specific fingerprint in the residual panel of Chapter 7.17.
- Slogan: a model is its assumptions. Know them and you know when to switch.
Three real-looking situations, each solved by reading the residuals and naming the broken assumption.
- A competitor closes; orders jump by 20 and stay there. The jump happened after the last allowed changepoint (the changepoint range stops before the end of the history). Fingerprint: residual-vs-time shows a step in the last weeks and the holdout mean error is far from 0. Broken assumption: the trend changes only inside the allowed range. Alternatives: a local level (exponential smoothing, state-space) that follows the new level within days, or a wider changepoint range.
- The weekly pattern slowly flattens over three years. Fingerprint: seasonal waves in the residuals whose size changes with time; residual ACF spikes at lags 7, 14, 21 that shift in sign. Broken assumption: the seasonal shape is fixed. Alternatives: Holt–Winters (the seasonal indices are updated), a state-space model with an evolving seasonal state.
- Demand follows weather streaks. Fingerprint: residual ACF decays geometrically from lag 1 ($r_1 = 0.6$). Broken assumption: independent noise. Alternatives: an AR residual term, ARIMA errors, a state-space or GP component.
In each case the question "what does my model assume?" turned a vague "the forecast is off" into a diagnosis and a candidate.
The assumptions of a Prophet-style Bayesian model, with the evidence that tests each one and the family that relaxes it:
| Assumption | Where it lives in the model | If it fails, you see | Test it with | Alternative that relaxes it |
|---|---|---|---|---|
| 1 · Components add (or multiply after a log) | $y = g + s + h + X\beta + \epsilon$ | Seasonal swings grow with the level; funnel | residual vs fitted (7.17) | multiplicative form, log scale, NB likelihood |
| 2 · Trend is piecewise linear, few sparse slope changes in the allowed range, and the last slope continues | $g(t)$, $\delta_j \sim$ Laplace | Steps or drift in residuals; holdout bias; wide, wild forecast fans | residual vs time; rolling-origin bias | local level / trend (ETS, state-space), more or later changepoints, damped trend |
| 3 · Seasonality is periodic with a fixed shape and amplitude | Fourier terms of order $N$ | Waves or spikes at seasonal lags that change over time | residual ACF at 7, 14, 21 | Holt–Winters, state-space seasonal states, quasi-periodic GP |
| 4 · Event and regressor effects are the same each time, additive and known in the future | $h(t)$, $X\beta$ | Spikes on dates; errors that depend on a combination of drivers | residual vs each column | interactions, gradient boosting, forecasts of regressors |
| 5 · Noise is independent given the mean, with the chosen likelihood and a constant scale (unless modelled) | $\epsilon_t$, Normal / Student-t / NB | Memory, funnel, heavy tails | ACF, Ljung–Box, Q-Q, spread bars | AR errors, state-space, GP, variance model, another likelihood |
| 6 · The future resembles the past (parameters are constant between changepoints) | all parameters | Error and coverage drift over time | error by origin, monitoring (7.15) | retraining, local or adaptive models |
| 7 · The approximate posterior is close to the real one | SVI guide family | Intervals too narrow; poor calibration | coverage and PIT (7.16) | richer guide, NUTS on a subset |
| 8 · One series at a time | per-series fit | Many short noisy series; no sharing of strength | compare with a pooled model | hierarchical or global models (gradient boosting, neural) |
The same exercise works for any model. For each family below we will write its core assumption in one line, then ask "when does it beat a Prophet-style model, and when does it not?".
Why do we need it?
No model is best everywhere. When a forecast disappoints, you need to know which assumption failed in order to pick the cure (a missing term, another noise model, another family). Without the habit you end up changing models at random.
Where is it used?
Design reviews of forecasting systems, model-selection discussions ("why not ARIMA?"), post-mortems of bad forecasts, and every interview about a forecasting project. It is also the basis of the residual-diagnostics step in Chapter 7.17.
How is it used?
Write the assumptions down in one line each, attach to each a test you can run on residuals or on rolling-origin errors, and a named alternative. When a test fails, try the alternative on the same rolling origins and compare accuracy and calibration.
"Prophet-style models are the best because they handle seasonality, holidays and trend automatically."
They handle them as long as the assumptions above hold. "Automatically" hides eight assumptions. Where they fail, other families win (see the bake-off at the end of this chapter).
"If the residuals look fine I know the model's assumptions hold."
Residual checks catch some failures (memory, spread, shape, missed events) but not others: the future may change in a way that history does not show (assumption 6). Only rolling-origin accuracy and monitoring catch that.
"A more flexible model avoids having assumptions."
Every model assumes something. A flexible one assumes smoothness, enough data and stability of the learned pattern; those assumptions are less visible, not absent.
For your own forecasting model, write the list yourself in interview form: (1) additive components with a piecewise-linear trend whose $\delta_j$ have Laplace priors; (2) periodic seasonality with Fourier order $N$; (3) holiday and regressor effects that repeat and are known in advance; (4) independent noise with a Normal, Student-t or Negative Binomial likelihood; (5) an SVI guide that is close enough to the posterior. For each, name the check you would run (the residual panel, rolling origins, coverage) and the model you would try if it failed. Phrase links to your projects as "in my forecasting model, the changepoint slopes $\delta_j$…", not as facts about code you did not write.
"We used a Prophet-style model because it is state of the art."
"We should use deep learning because it learns everything from the data."
"The model assumes additive components: a piecewise-linear trend with sparse slope changes, periodic seasonality with a fixed shape, repeatable holiday effects and independent noise with a chosen likelihood. That fits long, stable, calendar-driven series. I would choose something else when the level shifts without warning (exponential smoothing or state-space), when momentum matters at short horizons (ARIMA errors), when there are many related series or rich nonlinear drivers (global gradient boosting or a neural model), or when the data are few and smooth and I need honest uncertainty (a Gaussian process). I decide with rolling-origin accuracy and calibration, not with fashion."
Model answer: "My model is a structured Bayesian regression on time. Its strengths are interpretability, calendar information and uncertainty; its assumptions are a bending-but-stable trend, a fixed seasonal shape, repeatable events and independent noise. Each assumption has a residual diagnostic and an alternative family, and I compare candidates on the same rolling origins."
Assumptions of a Prophet-style model: additive components; sparse bends in a piecewise-linear trend inside the allowed range; fixed seasonal shape; repeatable events and known regressors; independent noise with the chosen likelihood; stable future; good-enough posterior; one series at a time.
For each: a symptom, a residual test, an alternative family. Always: "what does it assume, when does it fail, what would I use then?"
Trap: "automatic" and "state of the art" are not arguments; rolling-origin accuracy and calibration are.
Quick check: your residual-vs-time plot shows an upward step in the last four weeks and the last-28-day mean error is $+18$ orders. Which assumption failed and which family is a natural alternative?
Assumption 2: the trend changes only inside the allowed changepoint range and then continues with the last slope. A level shift near the end is not covered. Natural alternatives: a local method that re-estimates the level (exponential smoothing or a state-space local level), or widening the changepoint range and re-checking the Laplace scale. Compare them on rolling origins that include the shift.
ARIMA, SARIMA and exponential smoothing: forecasters that read the recent past core
There are two kinds of weather forecaster. One reads the calendar: "it is December, so it will be cold." The other reads the last few days: "it was cold yesterday and the day before, so it will be cold today." A Prophet-style model is the calendar reader. ARIMA and exponential smoothing are the recent-past readers: tomorrow is a weighted mix of the latest values (and the latest surprises), and the weights are learned from the series itself.
Each is better at something. The recent-past reader is great when today is a good guide to tomorrow: momentum, a level that wanders, short horizons. The calendar reader is great when the date carries the information: a holiday next Friday, a yearly peak, a promotion.
Three ways to say it:
- Picture: a driver looking at the road just ahead (memory-based) versus a driver following a printed route (calendar-based).
- Numbers: with day-to-day momentum $\varphi = 0.8$, using the recent past cuts the one-day-ahead error by 40%, but by only about 2% a week ahead.
- Slogan: memory wins at short horizons; structure wins at long horizons and on known dates.
The size of the memory advantage. Suppose the calendar part of both models is the same and the leftover noise is an AR(1) with $\varphi = 0.8$ (unit variance). A calendar-only forecast ignores the last leftover $e_T$; a memory-aware forecast adds $\varphi^h e_T$.
- The calendar-only error at horizon $h$ is the leftover itself: variance $1$.
- The memory-aware error removes the part that is predictable from $e_T$. Its variance is $1 - \varphi^{2h}$ (the standard AR(1) result: the unexplained part of $e_{T+h}$ given $e_T$).
- So the ratio of typical errors (memory-aware ÷ calendar-only) is $\sqrt{1 - \varphi^{2h}}$: $h = 1$: $\sqrt{1 - 0.64} = 0.60$. $h = 2$: $\sqrt{1 - 0.4096} = 0.77$. $h = 3$: $0.86$. $h = 5$: $0.95$. $h = 7$: $0.98$. $h = 14$: $0.999$.
- Reading: 40% smaller errors tomorrow, 14% smaller in 3 days, 2% smaller in a week, nothing in two weeks. At $\varphi = 0.5$ the advantage is $13\%$ tomorrow and gone after 3 days.
- Conclusion: a memory model beats a calendar model on short-horizon forecasts of a series with real momentum, and the advantage fades fast. If you forecast 28 days ahead, the structure matters far more.
ARIMA($p$, $d$, $q$) (Chapter 7.6): after differencing the series $d$ times, each value is a linear combination of the previous $p$ values (AR part) and the previous $q$ shocks (MA part) plus a new shock. SARIMA adds the same idea at the seasonal lag $m$ (seasonal AR/MA terms and seasonal differencing). With regressors (SARIMAX) the regression part carries the outside drivers and the ARMA part carries the memory of the errors. Exponential smoothing / ETS (Chapter 7.5): level, trend and seasonal states, each updated by a fraction ($\alpha, \beta, \gamma$) of the latest surprise.
Core assumption. After differencing (or through the updating states), the series is generated by a linear process with fixed coefficients: the next value is a weighted sum of recent values and shocks; the seasonal period is one fixed number; shocks are roughly Gaussian (for the usual intervals). The trend and level are allowed to wander: that is what the differencing and the smoothing states are for.
| When ARIMA / ETS beats a Prophet-style model | When it does not |
|---|---|
| Strong day-to-day momentum and short horizons (days) | Several seasonalities at once (weekly + yearly on daily data): SARIMA with $m = 365$ is impractical; classic ETS has one season |
| The level wanders or jumps without warning: differencing or smoothing follows it | Holidays that move around the calendar, promotions, outside drivers with complex effects |
| Short, clean series with one seasonality: few parameters, little to tune | Gaps and irregular dates (the classic forms want a regular grid); count and heavy-tailed likelihoods |
Thousands of series where automatic selection (auto_arima, ETS) is the only affordable approach | Long horizons: intervals grow like a random walk's; and you need explainable components and a full predictive distribution with parameter uncertainty |
Why do we need it?
"Why not ARIMA?" is the most likely follow-up question about a forecasting project. You need to say what ARIMA and ETS assume, where they are strong (short horizons, momentum, wandering levels, little data) and why they are weaker for calendar-heavy series.
Where is it used?
Supply-chain and finance forecasting at scale, statsmodels (SARIMAX, ExponentialSmoothing, ETSModel), pmdarima.auto_arima, R's forecast package, and as the standard baselines in forecasting competitions.
How is it used?
Difference until the series looks stationary, read the ACF/PACF to pick $p$ and $q$ (or let an automatic search choose), fit, check the residuals, forecast. Compare with your structured model on the same rolling origins and look at accuracy by horizon.
"ARIMA is outdated; modern models always beat it."
ARIMA and ETS stay among the hardest baselines to beat on short horizons and on single series, and they are the standard benchmarks in forecasting papers. A model that cannot beat them has not earned its complexity.
"ARIMA needs a stationary series, so it cannot handle trend."
Differencing ($d \ge 1$) removes a trend or wandering level. The forecast then follows the latest level, which is exactly why ARIMA recovers quickly after a level shift (Chapter 7.4).
"SARIMAX with dummy variables is the same as a Prophet-style model."
It can include the same columns, but the trend and seasonality are handled by differencing and by one fixed seasonal period, with no priors, no changepoint shrinkage and no non-Gaussian likelihood; getting several seasonalities in is awkward.
As you described it, your model has no memory term: its noise is independent (Chapter 7.3; check your code). That is a trade-off. At horizons of a few days ARIMA-type memory could beat it if the residual ACF shows lag-1 momentum; at horizons of weeks the structure wins. The practical hybrid is regression with AR errors: keep $g(t) + s(t) + h(t) + X_t\beta$ and add an AR(1) error term. An honest interview answer: "I would first check the residual ACF; if it shows momentum I would add an AR term and compare accuracy by horizon against a SARIMAX baseline."
ARIMA/SARIMA: after differencing, a linear process of recent values and shocks, one fixed seasonal period. ETS: level/trend/season states updated by $\alpha, \beta, \gamma$.
Memory advantage with AR(1) noise: error ratio $\sqrt{1 - \varphi^{2h}}$ (0.60 at $h = 1$ for $\varphi = 0.8$, 0.98 at $h = 7$).
Beats Prophet-style: momentum, short horizons, wandering levels, little data, single season. Loses: multiple seasonalities, moving holidays, regressors with complex effects, counts, long horizons, explainability.
Quick check: leftover noise has $\varphi = 0.5$. At which horizon does a memory correction first become less than a 5% improvement?
The ratio is $\sqrt{1 - 0.5^{2h}}$: $h = 1$: $0.866$ (13% better); $h = 2$: $\sqrt{1 - 0.0625} = 0.968$ (3% better). So from the second day on the gain is below 5%. With $\varphi = 0.5$ a memory model only helps for tomorrow.
State-space models: a hidden level that moves, seen through noise core
Your car's GPS in a tunnel. The car has a true position you cannot see (the hidden state). Every second you get a noisy reading. The GPS keeps a belief: "I think the car is here, give or take this much". When a new reading arrives it moves its belief toward the reading, by a lot if it trusts the reading, by a little if it trusts its own prediction more. That "move toward the reading by a fraction" is the Kalman filter.
A state-space model describes a time series in exactly this way: a hidden state (the level, the slope, the seasonal pattern) that changes a little each day by random pushes, and the data are noisy views of it. Where a Prophet-style model says "the trend bends at a few chosen dates", a state-space model says "the level and slope can drift a little every day, and we learn how fast".
Three ways to say it:
- Picture: a hidden buoy drifting on the water, photographed through fog; the filter's estimate follows the buoy.
- Numbers: if the level moves $q = 0.25$ as much (in variance) as the observation noise, the filter moves 39% of the way to each new surprise.
- Slogan: exponential smoothing is a Kalman filter in disguise.
The local-level model with observation noise variance $\sigma_\varepsilon^2 = 1$ and level-shock variance $\sigma_\eta^2 = q = 0.25$ (so $q = 0.25$). Work in units of $\sigma_\varepsilon^2$. One filter step: the current belief about the level is $a = 100$ with variance $P = 0.5$, and the next observation is $y = 104$.
- Predict. The level may have drifted: the predicted level is still $100$ and its variance grows to $P + q = 0.5 + 0.25 = 0.75$.
- Gain. $K = \dfrac{0.75}{0.75 + 1} = 0.4286$: trust in the new reading relative to the prediction (the observation noise adds 1).
- Update. New level $= 100 + 0.4286 \times (104 - 100) = 101.71$. New variance $= (1 - K)\times0.75 = 0.4286$.
- Repeat with the next observations: the gain settles at $0.4286 \to 0.4043 \to 0.3955 \to \dots \to 0.3904$. In general the settled gain solves $\alpha = (\sqrt{q^2 + 4q} - q)/2$; for $q = 0.25$: $(\sqrt{1.0625} - 0.25)/2 = 0.3904$.
- That settled gain is the smoothing weight $\alpha$ of simple exponential smoothing (Chapter 7.5): the new level is $\alpha\,y + (1 - \alpha)\,\text{old level}$. A fast-moving level (large $q$) gives a large $\alpha$; a slow level (small $q$) gives a small $\alpha$.
- Forecasts: the best forecast of every future day is the current level (flat). The variance of $y_{T+h}$ is $P_T + h\,q + 1$: with $P_T = 0.39$ this is $1.64$ tomorrow, $3.14$ in a week and $4.89$ in two weeks (sds $1.28$, $1.77$, $2.21$ in units of $\sigma_\varepsilon$): the interval grows because the level may drift further.
A linear-Gaussian state-space model has two equations:
$$y_t = Z\,\alpha_t + \varepsilon_t, \qquad \alpha_t = T\,\alpha_{t-1} + \eta_t,\qquad \varepsilon_t \sim N(0, \sigma_\varepsilon^2),\ \ \eta_t \sim N(0, Q).$$- $\alpha_t$ is the state, a short vector of hidden things: level, slope, seasonal values. The observation equation says how the data come from the state; the state equation says how the state moves.
- Local level: $\alpha_t = \mu_t$, $\mu_t = \mu_{t-1} + \eta_t$. Local linear trend: the state also holds a slope $\beta_t = \beta_{t-1} + \zeta_t$. Structural time series add seasonal states and fixed-coefficient regressors; "Bayesian structural time series" give the variances and regression weights priors.
- The Kalman filter computes, in one pass over the data, the best estimate of the state and its uncertainty at every time (predict, then update by the gain). It also produces the likelihood of the data, so the variances can be estimated by maximum likelihood or with priors. Missing observations are skipped (no update).
- ETS and ARIMA models have state-space forms, so they are special cases;
statsmodelsfits them with the same machinery (UnobservedComponents,SARIMAX,ETSModel). - Core assumption: the components follow linear dynamics with Gaussian shocks, and the shock sizes are constant; the series is a noisy window on a smoothly evolving hidden state.
Compared with a Prophet-style model. A Prophet-style trend has a Laplace prior on slope changes at chosen knots: a few large, sudden bends. A local-linear-trend state-space model lets the slope change by small Gaussian pushes at every step: many small changes. Likewise a Prophet-style Fourier seasonality has fixed coefficients, while a state-space seasonal can evolve over the years.
| When a state-space model beats a Prophet-style model | When it does not |
|---|---|
| The level, slope or seasonal pattern drifts slowly or unpredictably, and you want it to follow | You have many holidays and regressors that need sparse shrinkage priors, or events at known dates |
| Missing observations, irregular gaps, real-time updating as each new point arrives (cheap, $O(n)$) | Counts and heavy tails: the exact filter needs Gaussian noise; other likelihoods need approximations or particle filters |
| You care about uncertainty about today's level and short-to-medium forecasts | You want sparse "a few clear changepoints" explanations; or long horizons, where the variance $h\,q$ makes the intervals widen steadily |
Why do we need it?
Many series do not follow a fixed recipe: the level drifts, a seasonal pattern fades, data points are missing. A state-space model lets those parts evolve while still giving a clean decomposition and honest uncertainty, and the filter does it in one cheap pass.
Where is it used?
Structural time-series forecasting (UnobservedComponents in statsmodels, R's KFAS, TensorFlow Probability's STS), nowcasting in economics, tracking and navigation, CausalImpact-style analysis built on Bayesian structural time series, and the engine inside ETS and ARIMA fits.
How is it used?
Choose the components (level, trend, seasonal, regression), let the library estimate the shock variances (or give them priors), run the filter and smoother, plot the estimated components, and forecast. The ratio of level variance to noise variance tells you how fast the level may move.
"A state-space model is a different kind of model from ARIMA and ETS."
It is a common framework containing them. The state-space view adds flexibility (time-varying components, missing data, regressors) and a single algorithm (the Kalman filter) for all of them.
"The hidden state is the true level, so the filtered estimate is the truth."
The state is a modelling device. The filter gives an estimate with uncertainty, and the quality depends on the assumed shock variances (the ratio $q$). A wrong $q$ gives a filter that lags (too small) or chases noise (too large).
"Local level and Prophet-style trends are the same: both let the trend change."
They differ in the prior: the state-space slope changes a little at every step (Gaussian increments), the Prophet-style trend bends rarely and sharply (Laplace $\delta_j$ at knots). They suit different series, and their long-horizon uncertainty differs.
A state-space model is the natural answer to "your trend only changes at the allowed changepoints; what if the level keeps drifting?". In your forecasting model the knots come from a Prophet-like grid and PELT, and the Laplace prior on $\delta_j$ expresses "most days no change, a few big changes". A local-linear-trend model expresses "small changes all the time". If your residual-vs-time plot shows slow drift between changepoints, or the last weeks are off, a state-space component (or an AR/GP term for the leftover, Chapter 7.17) is the principled extension. Exact Kalman filtering needs Gaussian noise; with the Negative Binomial likelihood you would need approximations.
State-space: $y_t = Z\alpha_t + \varepsilon_t$, $\alpha_t = T\alpha_{t-1} + \eta_t$. Kalman filter: predict ($P + q$), gain $K = (P + q)/(P + q + 1)$, update by $K\times$ surprise.
Local level settled gain: $\alpha = (\sqrt{q^2 + 4q} - q)/2$, $q = \sigma_\eta^2/\sigma_\varepsilon^2$ = exponential smoothing. Forecast variance $P_T + hq + \sigma_\varepsilon^2$.
Beats Prophet-style: drifting components, missing data, online updating. Loses: many events/regressors with sparse priors, non-Gaussian counts, sparse bends.
Quick check: for $q = 1$ what is the settled gain $\alpha$, and what is the memory of the level?
$\alpha = (\sqrt{1 + 4} - 1)/2 = (2.236 - 1)/2 = 0.618$. The level moves 62% of the way to each new surprise; the memory is about $1/\alpha = 1.6$ days. A level that changes as much per step as the noise is very hard to smooth.
Gaussian processes: a random smooth curve, pinned down by the data core
Imagine a long, springy ribbon lying across a graph. You tell it only how stiff it is: a very stiff ribbon bends slowly, a floppy one wiggles quickly. Now you pin it to your data points. Where there are data, the ribbon is held in place; between and beyond the data it follows its own stiffness, and your uncertainty about where it runs grows the farther you are from a pin.
That ribbon is a Gaussian process (GP): a probability distribution over whole curves. You do not choose a list of columns (trend, Fourier terms) as in your Prophet-style model; you choose a kernel, a rule saying how similar the values at two times should be. Times close together are similar; times far apart are not. Choose a periodic kernel and the ribbon repeats.
Three ways to say it:
- Picture: a springy ribbon pinned at the data; the band around it widens away from the pins.
- Numbers: one observation at $t = 0$ with the kernel length $\ell = 2$: the GP is $0.88$ of the way to the data one day later and only $0.14$ four days later.
- Slogan: a GP is Bayesian regression where the "columns" are chosen by a kernel.
The smallest possible GP. The prior mean is 0, the kernel is $k(t, t') = \exp\!\big(-\tfrac12 (t - t')^2/\ell^2\big)$ with $\ell = 2$ (and variance 1), and we observe one point $y = 1$ at $t = 0$ with no noise. What do we believe about $f$ at $t = 1$ and $t = 4$?
- Covariance between the value at $t^*$ and the observed value at 0: $k(t^*, 0)$. At $t^* = 1$: $\exp(-\tfrac12\cdot1/4) = e^{-0.125} = 0.8825$. At $t^* = 4$: $\exp(-\tfrac12\cdot16/4) = e^{-2} = 0.1353$.
- Posterior mean $= k(t^*, 0)\,\big/\,k(0,0) \times y$. At $t^* = 1$: $0.8825$. At $t^* = 4$: $0.1353$. The curve decays back to the prior mean (0) away from the data.
- Posterior variance $= k(t^*, t^*) - k(t^*, 0)^2/k(0,0) = 1 - k(t^*,0)^2$. At $t^* = 1$: $1 - 0.7788 = 0.2212$ (sd $0.47$). At $t^* = 4$: $1 - 0.0183 = 0.9817$ (sd $0.99$).
- So near the data we are fairly sure ($\text{sd}\ 0.47$), and far away we know almost nothing more than the prior (sd 0.99, prior sd 1). With noisy data and many points the same formulas run with matrices.
A Gaussian process $f \sim GP(m, k)$ says: for any set of times $t_1, \dots, t_n$ the vector $(f(t_1), \dots, f(t_n))$ is multivariate Normal with mean $m(t_i)$ and covariance matrix $K_{ij} = k(t_i, t_j)$. With data $y_i = f(t_i) + \varepsilon_i$, $\varepsilon_i \sim N(0, \sigma_n^2)$, the posterior at new times $t_*$ is Normal with
$$\text{mean} = m_* + k_*^\top (K + \sigma_n^2 I)^{-1}(y - m), \qquad \text{variance} = k(t_*, t_*) - k_*^\top (K + \sigma_n^2 I)^{-1} k_*.$$- Kernels encode assumptions: squared-exponential (RBF): very smooth; exponential / Matérn-1/2: rough, exactly the AR(1) memory of Chapter 7.17; periodic: repeats with period $P$; sums and products combine them (trend kernel + periodic kernel × slow decay ≈ "a seasonal pattern that may drift").
- Hyperparameters (length $\ell$, signal sd, noise) are chosen by the marginal likelihood (the evidence, Chapter 6.1): $\log p(y) = -\tfrac12 y^\top(K + \sigma_n^2 I)^{-1}y - \tfrac12\log|K + \sigma_n^2 I| - \tfrac n2\log 2\pi$. It rewards fitting the data and penalises needless flexibility, like an Occam's razor.
- Cost: exact inference needs a Cholesky factorisation of an $n\times n$ matrix: about $n^3/3$ operations and $n^2$ numbers of memory. $n = 30$: 9 000 operations. $n = 1\,000$: $3\times10^8$ and 8 MB. $n = 10\,000$ days: $3.3\times10^{11}$ and 800 MB. Approximations (sparse GPs, state-space forms of GPs) exist.
- Core assumption: the function is a draw from a Gaussian process with the chosen kernel: a given smoothness and (usually) the same behaviour everywhere (stationarity); Gaussian noise.
| When a GP beats a Prophet-style model | When it does not |
|---|---|
| Small data and a smooth, flexible shape you do not want to specify with columns | Long series: $O(n^3)$ cost (daily data over many years) |
| Irregular times; honest uncertainty that grows away from the data | Extrapolating a trend: the RBF GP falls back to the mean; a trend must be put in by hand (a linear mean or kernel) |
| You can write the structure as a kernel (trend + periodic) and want it learned from data | Several strong seasonalities, holidays and regressors: more kernel engineering than adding columns |
| Residual smooth wandering on top of a calendar model | Counts and heavy tails (non-Gaussian likelihoods need approximate inference) |
Why do we need it?
When you do not know the shape of a function but you believe it is smooth, a GP lets the data decide the shape and reports uncertainty everywhere, including how fast it grows beyond the data. It is also the cleanest way to see how a kernel is an assumption.
Where is it used?
Small-data forecasting and calibration curves, Bayesian optimization (surrogate models), spatial statistics (kriging), residual models on top of a parametric forecast, and libraries such as scikit-learn (GaussianProcessRegressor), GPyTorch and NumPyro.
How is it used?
Choose a kernel from the structure you believe in (smooth, rough, periodic, sums), fit the hyperparameters by maximising the marginal likelihood (or give them priors), predict with mean and band, and check by holding out the end of the series.
"A GP forecasts the future as well as it interpolates."
Away from the data an RBF GP reverts to its prior mean with a wide band. Extrapolation needs a kernel with structure that extends (periodic, linear) or an explicit mean function.
"A GP fits itself: no assumptions, no tuning."
The kernel and its hyperparameters are the assumptions. A wrong kernel gives confident wrong answers, and hyperparameters fitted on little data can be badly determined (use priors or compare evidence).
"A GP is too slow to be useful."
It is $O(n^3)$, which is fine for hundreds of points and bad for tens of thousands. Sparse and state-space approximations reduce the cost; and a GP is a good residual component on top of a fast calendar model.
Your Prophet-style model is a Bayesian linear regression with a fixed list of columns (trend pieces, Fourier terms, holiday indicators). A GP is the same kind of model with the columns replaced by a kernel: more flexible, but the assumptions hide in the kernel. Two natural uses next to your model: (1) a GP (exponential kernel) for the leftover when the residual ACF shows momentum, equivalent to an AR(1) term for daily data (Chapter 7.17); (2) a GP trend instead of changepoints when the trend is smooth rather than bending sharply. Mind the $O(n^3)$ cost with years of daily data.
$f \sim GP(m, k)$: any finite set of values is multivariate Normal with covariance $K_{ij} = k(t_i, t_j)$. Posterior mean $k_*^\top(K + \sigma_n^2 I)^{-1}y$, variance $k_{**} - k_*^\top (K + \sigma_n^2 I)^{-1}k_*$.
Kernel = assumption (smooth, rough, periodic, sums). Hyperparameters by marginal likelihood. Cost $O(n^3)$, memory $O(n^2)$.
Beats: small data, smooth shapes, honest uncertainty. Loses: long series, extrapolating trends, many seasonalities and regressors, counts.
Quick check: why does a GP with an RBF kernel forecast the long-run average once you are far beyond the data?
The RBF correlation $k(t_*, t_i)$ drops to nearly 0 when $t_*$ is many length-scales from every data point. Then the posterior mean $k_*^\top(\dots)^{-1}y$ goes to 0 (the prior mean, here the average of the data) and the variance to the prior variance. The kernel says "values far apart are unrelated", so the data can say nothing about the far future. A periodic kernel says the opposite for times a whole period apart.
Gradient boosting: a forest of small question trees on features core
A decision tree plays "twenty questions". Is it a weekend? Is a promotion running? Was last week's sales above 120? Each answer sends the day down a branch, and at the bottom of the branch is a number: the average of the training days that ended up there. Gradient boosting builds hundreds of small trees, each one fitted to the mistakes of the trees before it, and adds them up.
For forecasting this has a big consequence. The model knows nothing about time, trends or seasons; it only knows the features you give it: calendar flags, prices, promotions, lagged sales. That is its strength (it can learn complicated interactions such as "promotions only help at weekends" with no extra work) and its weakness (a tree can only predict an average it has seen, so it cannot extrapolate a trend).
Three ways to say it:
- Picture: a stack of small flow-charts of yes/no questions, each correcting the last; the answer is the sum of the numbers at the bottoms.
- Numbers: trained on sales that never exceeded 146, the forecast stays near or below the top of that range (about 143 in the widget), even when the true trend heads to 160.
- Slogan: trees interpolate; they do not extrapolate.
The smallest possible failure. Eight days, a perfect upward trend: $t = 1,\dots,8$ and $y = 10, 12, 14, 16, 18, 20, 22, 24$ (two more orders every day). A depth-1 tree on the feature $t$.
- Try the split "$t \le 4.5$". Left bucket: $10, 12, 14, 16$ with mean $13$. Right bucket: $18, 20, 22, 24$ with mean $21$.
- Squared error: left $(10-13)^2 + (12-13)^2 + (14-13)^2 + (16-13)^2 = 9 + 1 + 1 + 9 = 20$; right the same, $20$. Total $40$. This is the best single split (any other split leaves a larger error).
- The tree predicts $13$ for every $t \le 4.5$ and $21$ for every $t \gt 4.5$.
- Forecast for $t = 9, 10, 11, 12$: the tree says $21$ each time. The truth is $26, 28, 30, 32$. At $t = 12$ the error is $11$ and keeps growing with the horizon.
- Boosting more trees and splitting deeper makes the staircase finer inside the data, but beyond $t = 8$ it is still flat at about the last bucket (the final value $24$ or a little less). A linear regression through the same eight points would be exact at $t = 12$ ($y = 8 + 2t$ gives $32$).
- The standard fixes: detrend first (model the trend with a regression or differences and give the trees the remainder), predict ratios or changes instead of levels, or add a linear model at the leaves. The widget shows the first fix.
Gradient boosting builds the prediction as a sum of small regression trees $h_m$, each fitted to the current residuals (for squared error, the negative gradient of the loss), added with a small learning rate $\eta$:
$$F_M(x) = F_0 + \eta\sum_{m=1}^{M} h_m(x),\qquad h_m \text{ fitted to } y - F_{m-1}(x).$$- The inputs $x$ are features you engineer: calendar (weekday, month, holiday flags, days to the next holiday), known-in-advance drivers (price, promotion), lags $y_{t-k}$ and rolling averages computed only from the past, and, in a model for many series, an identifier of the series. Typical libraries: LightGBM, XGBoost, CatBoost.
- Loss can be squared error, Poisson or Tweedie (counts and non-negative amounts), or a quantile loss to predict quantiles. Intervals are not automatic: use quantile models, conformal methods or ensembles.
- Horizon. A lag feature $y_{t-1}$ is unknown for days 2, 3, … of the forecast. Either use lags at least as old as the horizon, train one model per horizon (direct), or feed forecasts back in (recursive, errors compound). Computing a rolling feature with the future in it is the leakage of Chapter 7.12.
- Core assumption: the target is a function of the supplied features, and that function is the same in the future as in the training data. The prediction is built from averages of training targets, so it cannot move meaningfully outside their range (a single tree exactly; a boosted sum only slightly).
| When boosting beats a Prophet-style model | When it does not |
|---|---|
| Many related series (a "global" model shares what it learns), lots of rows | One short series: too few rows, and the features must be invented |
| Rich drivers with nonlinear interactions (price × promotion × weekday, weather) | Trends that must be extrapolated; level shifts to new, unseen levels |
| Count or skewed targets handled by a Poisson / Tweedie loss, no assumptions about seasonal shape | You need an explainable decomposition (trend, season, holiday) and a posterior predictive distribution with parameter uncertainty |
| Reported strong results in large retail benchmarks (many of the leading entries of the M5 competition were LightGBM-based global models; check the published summaries for details) | Hierarchical uncertainty, calibrated intervals by construction, small data |
Why do we need it?
When many series share patterns, or when sales depend on complicated combinations of drivers, a flexible learner that finds interactions by itself can beat a hand-built additive model. You need to know its blind spots (extrapolation, horizon, intervals) to use it safely.
Where is it used?
Retail and demand forecasting at scale (store × product × day), click-through and conversion models, tabular machine learning in general (LightGBM, XGBoost, CatBoost), and as the "residual learner" on top of a statistical forecast.
How is it used?
Build features that are known at the forecast origin (calendar, lags at least as old as the horizon), detrend or difference the target if there is a trend, train with a time-ordered validation set, choose the loss to match the target, and get intervals from quantile models or conformal prediction.
"Gradient boosting is a more powerful version of regression, so it will also extrapolate."
Trees predict averages of training targets, so they flatten outside the range they saw. Detrend, difference or predict ratios, or use a model with a linear part.
"I will add rolling means of sales as features, and the backtest is great."
If the rolling window includes days after the forecast origin, or if a lag is younger than the horizon, the backtest leaks the answer. Compute every feature as of the origin and validate with rolling origins (Chapter 7.15).
"Boosting gives probabilities and intervals out of the box."
A squared-error booster gives a point forecast. Intervals need a quantile loss, conformal prediction or an ensemble, and their coverage must be checked (Chapter 7.16).
In your forecasting model the trend is explicit (piecewise-linear with changepoints), so extrapolation is built in; a boosted model would have to be told about it. Two sensible ways to use boosting next to your model: (1) as a residual learner on top of the structured forecast, to catch nonlinear interactions between drivers; (2) as a global model across many series, which your one-series-at-a-time model does not do. The caution to quote in an interview: it needs horizon-safe features, a detrended target, and a separate plan for uncertainty.
Boosting: $F_M(x) = F_0 + \eta\sum h_m(x)$, each tree fitted to the current residuals. Features: calendar, known drivers, lags at least as old as the horizon, series id.
Predictions are averages of training targets: no extrapolation. Fix: detrend / difference / ratios. Intervals need quantile loss or conformal.
Beats Prophet-style: many series, rich nonlinear drivers, big data. Loses: one short series, trends, explainable decomposition, calibrated uncertainty by design.
Quick check: a boosted model is trained on daily sales whose highest value was 146. The true trend reaches 160 in the forecast window. About what is the highest value the model can forecast?
About 146 or a little less (the widget shows about 143): every prediction is built from training-target averages, so it cannot climb meaningfully above the range of the targets. The forecast flattens and the error grows with the horizon unless the trend is removed from the target first.
Deep learning and neural forecasting: N-BEATS and TFT (awareness only)
A new restaurant manager has three weeks of data and must guess next Friday. A manager who has run a hundred restaurants already knows that Fridays are busy, even though this one has seen only three Fridays. She shares what she has learned across restaurants.
That is the main reason neural forecasters can win. Models such as N-BEATS and the Temporal Fusion Transformer (TFT) are neural networks trained on many series at once (a "global" model). They take a window of recent values, plus covariates (holidays, prices, store type), and output the next values. Nobody tells them about trend, weekly pattern or holidays; with enough data they learn these from the examples.
Three ways to say it:
- Picture: one apprentice learning from three weeks of one shop, versus an experienced manager who has seen a hundred.
- Numbers: in the widget, sharing the weekly shape across 16 similar series cuts the typical error from 9.2 to about 8.2 (noise floor 8); if the series are very different, full sharing makes it worse (about 11).
- Slogan: neural forecasters win by sharing strength across series, not by magic; on one short series they have nothing to share.
Why sharing helps, with numbers. Each series has a fixed weekly shape plus noise with sd $\sigma = 8$ (variance 64), and we have 3 weeks of history per series (so 3 values for each weekday). We forecast next week.
- Local model (each series alone): the forecast for Friday is the average of its 3 Fridays. That average has error variance $64/3 = 21.3$.
- Expected squared forecast error: new noise $64$ + estimation error $21.3$ = $85.3$, so a typical error of $\sqrt{85.3} \approx 9.2$.
- Global model (16 similar series share the weekly shape): the shared estimate averages $16 \times 3 = 48$ Fridays, error variance $64/48 = 1.3$. Each series still estimates its own overall level from 21 days (error variance $64/21 = 3.05$).
- Expected squared error: $64 + 3.05 + 1.3 = 68.4$, a typical error of $\sqrt{68.4} \approx 8.3$ (simulation in the widget: 8.2). That is 10% better, and no model can go below the noise sd $8$.
- But if the series really have different shapes (each weekday effect differs by a further sd of 8 between series), forcing one shared shape adds a bias of variance $64$: total $64 + 3.05 + 1.3 + 64 \approx 132$, a typical error of about $11.5$ (simulation: 11.0). The right answer is in between: share the common part, keep a series-specific part, i.e. partial pooling (Chapter 6.6).
Neural forecasting (awareness). A neural network (feed-forward, recurrent/LSTM, convolutional, attention-based) is trained by gradient descent to map a window of past values and covariates to the next $H$ values, usually across thousands of series. Two examples to recognise:
- N-BEATS (Oreshkin et al.): a stack of fully connected blocks. Each block reads the input window and produces a backcast (the part of the window it explains) and a forecast. The next block works on what is left (input minus backcast), and the final forecast is the sum of the blocks' forecasts. An interpretable variant constrains blocks to polynomial trend shapes and Fourier seasonal shapes: the same kinds of components as $g(t)$ and $s(t)$ in your model. The basic form is univariate.
- TFT (Temporal Fusion Transformer, Lim et al.): recurrent layers plus attention, with separate inputs for static information (store type), future-known inputs (holidays, prices) and past-observed inputs; gating and variable-selection layers; quantile outputs for several horizons; the attention weights give some interpretability.
- Others to recognise: DeepAR, an autoregressive recurrent model that outputs the parameters of a likelihood (Normal, Negative Binomial for counts) and is trained on many series.
Core assumption. With enough series and history, a flexible function can learn trend, seasonality, event effects and cross-series patterns from examples, and what it learned in the training period holds in the future. Inputs are scaled (per series) so one network can serve series of different sizes.
| When neural forecasting beats a Prophet-style model | When it does not |
|---|---|
| Thousands of related series with long histories, so there is a lot to share | One series or a handful, or short histories: nothing to learn from but noise |
| Rich covariates and cross-series effects that are hard to write as columns (promotions across products, cannibalisation) | You need an explainable decomposition and explicit priors (a regulated or high-stakes setting) |
| You can afford GPUs, tuning, ensembling (neural runs vary with the random seed) and monitoring | Limited compute or time; or the gain over exponential smoothing / boosting does not justify the complexity |
| Quantile or likelihood heads give the intervals you need and are validated | Calibrated uncertainty with parameter uncertainty is required: intervals come from the output heads, not from a posterior |
Evidence, kept modest. Results from forecasting competitions are mixed and depend on the data: the summary of the M4 competition reported that the winner was a hybrid of exponential smoothing and a recurrent network and that the pure machine-learning submissions did not beat a simple combination of statistical methods; N-BEATS later reported results on M4 that improved on the winning hybrid; and many top entries of the M5 retail competition used gradient boosting rather than deep networks. Treat these as pointers to read the papers, not as a verdict.
Why do we need it?
When a company forecasts hundreds of thousands of series, one model per series is costly and wastes shared information. A global neural model learns once and forecasts everything, using covariates, and can capture patterns nobody wrote down. Knowing when that pays off keeps you from using it when it cannot.
Where is it used?
Large-scale demand forecasting at retailers and marketplaces, energy-load and traffic forecasting, research benchmarks (the M4 data, electricity and traffic sets), and libraries such as GluonTS, PyTorch Forecasting, Darts and NeuralForecast.
How is it used?
Scale each series, build training windows from many series (history window in, horizon out), train with a time-ordered validation set, ensemble several seeds, and evaluate with rolling origins against strong statistical baselines (seasonal naive, ETS), checking both accuracy and calibration.
"Deep learning is more advanced, so it will beat classical models."
On a single series, or on a few short ones, a plain exponential-smoothing or seasonal-naive forecast is often as good or better. Neural models pay off when there are many related series to learn from, and only if they are validated against strong baselines on rolling origins.
"The network learns trend and seasonality by itself, so I don't need to think about them."
It learns what the training examples show. Series scaling, trend handling, holiday inputs and the training window design all matter, and an unseen level or regime breaks it like any other model.
"A neural model with quantile outputs gives calibrated intervals."
The intervals come from output heads trained with a quantile loss, not from a posterior. Check coverage and calibration (Chapter 7.16) before trusting them.
Your forecasting model fits one series at a time, so it cannot share strength across series. In your A/B framework, hierarchical partial pooling does exactly this across groups and segments (Chapter 6.6): groups with little data borrow from the whole. A hierarchical version of your forecasting model, with shared priors over the components of many related series, would be the Bayesian route to the benefit neural global models get from scale. The N-BEATS interpretable variant (trend + Fourier seasonality) is a neural cousin of your $g(t) + s(t)$ decomposition; mention it as awareness, not as something you built.
Global neural models (N-BEATS: stacked blocks with backcast and forecast; TFT: attention with static / known-future / observed inputs; DeepAR: autoregressive with a likelihood) learn from many series at once.
Gain comes from sharing: local MSE $64 + 64/3$, global $64 + 64/21 + 64/(3K)$ in the example; too much sharing biases (partial pooling in between).
Beats: thousands of related series, rich covariates, resources. Loses: one short series, explanation, small data. Evidence is mixed; validate on rolling origins.
Quick check: you must forecast one product's weekly demand with 40 weeks of history and a known promotion calendar. Is a global neural model the natural first choice?
No. One series with 40 points leaves nothing to share, and a structured regression (or ETS, or a Prophet-style model with a promotion column) uses the promotion calendar directly and can be explained. A neural model would be a candidate if you had thousands of similar products, where learning once and sharing pays.
A model chooser: from the situation to the candidates core
A doctor does not start with a favourite medicine. She asks about the symptoms (how long, how bad, other illnesses) and the answers point to a short list of treatments; the tests then decide. Choosing a forecasting model works the same way. A handful of questions about your situation (how many series, how much history, are there events, does the level jump, do you need to explain, do you need honest intervals) point to a short list of candidate families. Then rolling-origin accuracy and calibration choose among them.
Two habits keep the shortlist honest. Always include the cheap baselines (seasonal naive, exponential smoothing) in the race. And when two families are close, prefer the one you can explain and monitor.
Three ways to say it:
- Picture: a short questionnaire on the left and a ranked shortlist on the right.
- Numbers: eight questions and eight families; the chooser keeps three. The decision between those three is made by rolling origins, not by the chooser.
- Slogan: the situation picks the candidates; the data pick the winner.
Three situations and the shortlist you would draw.
- 20 000 products, 6 months of history each, no promotion data, frequent unexplained jumps. Little to learn per series, jumps nobody can schedule, huge scale. Shortlist: seasonal naive, exponential smoothing (cheap, adapts to jumps), maybe a global boosting model with calendar features. A per-series Bayesian model with priors and inference diagnostics is too heavy to babysit 20 000 times.
- One national demand series, five years of daily data, holidays, price changes, counts, components must be explainable, intervals must be honest. Shortlist: a Prophet-style Bayesian model with a Negative Binomial likelihood (explainable components, event and regressor columns, posterior predictive), with ETS and SARIMAX as baselines to beat.
- 4 000 products × 300 stores, prices and promotions, strong interactions. Shortlist: global gradient boosting with price and promotion features (and a neural global model if you have the resources), against seasonal naive and ETS baselines; add a statistical top-level reconciliation if totals matter.
In each case the winner is decided afterwards: same rolling origins, same metrics (accuracy and calibration, Chapters 7.15 and 7.16).
The chooser in the widget scores each family by adding or subtracting points for each answer. The points are teaching heuristics, not laws; they encode the assumptions of this chapter. In brief:
| Answer | Points toward | Points away from |
|---|---|---|
| One series | everything per-series (ETS, ARIMA, structured, state-space, GP) | global learners (nothing to share) |
| Hundreds or thousands of series | global boosting and neural; cheap automatic methods | per-series Bayesian fits, GP |
| Short history | baselines, models with priors | flexible learners |
| Long history | richer models (structured, state-space, neural) | exact GP (cost $n^3$) |
| Several seasonalities | structured regression, state-space, GP | classic ETS and ARIMA (one season) |
| Known events or regressors | structured regression, boosting | seasonal naive, ETS |
| Counts or many zeros | count likelihoods (NB), Poisson / Tweedie losses | Gaussian-only models |
| Unplanned level shifts | local models: ETS, state-space, differencing | global trends with limited changepoint ranges, boosting |
| Strong short-term momentum | ARIMA-type memory, state-space, ETS | models with independent noise |
| Must explain components | structured regression, state-space, ETS | boosting, neural |
| Need honest predictive distributions | Bayesian structured, state-space, GP | point-forecast boosters |
| Nonlinear interactions between drivers | boosting, neural, GP | additive models, baselines |
Why do we need it?
"Which model should I use?" has no answer without the situation. A short, explicit mapping from the situation to candidates stops you from reaching for a favourite, and it makes the choice easy to defend in a design review.
Where is it used?
Forecasting platform design (what to offer for which kind of series), project kick-offs, the "model zoo" of libraries such as Darts, sktime and statsforecast, and interviews that ask "how would you forecast X?".
How is it used?
Answer the questions for your series, read the shortlist and the reasons, put the baselines in the race, then compare the candidates on the same rolling origins. Record the reasons in the model card so the choice can be revisited when the situation changes.
"The chooser said Prophet-style, so Prophet-style is the answer."
It gives candidates and reasons. The answer comes from rolling-origin accuracy and calibration on your data, including the baselines.
"Start with the most advanced model and simplify if needed."
Climb a ladder instead: seasonal naive → exponential smoothing → ARIMA / structured regression → boosting / neural. Each step must beat the previous one on rolling origins to justify itself.
"The best model on average is the best for every series."
Averages hide structure: some series (with events, with jumps) favour different families. Look at accuracy by series group, horizon and season, and consider combining models.
Use the same questionnaire for your two projects. For the forecasting model: a calendar-driven series with events, regressors, several seasonalities, counts and a need for explainable components and honest intervals points at a structured Bayesian regression, with ETS and SARIMAX as the baselines it must beat and global boosting as the challenger if many related series appear. For the A/B framework the "model families" are different (conjugate and hierarchical Bayesian models), but the habit is the same: say what the situation implies, name the alternative, and test it.
Questions: how many series? how much history? several seasonalities? events / regressors? counts? level shifts? momentum? explain? intervals? interactions?
Answers shortlist families; rolling-origin accuracy and calibration choose. Ladder: seasonal naive → ETS → ARIMA / structured → boosting / neural.
Trap: a chooser gives candidates, not winners; scores are heuristics.
Quick check: one series, only 30 days of history, a clear weekly pattern, no outside drivers. Which families would you start with?
Seasonal naive and a small regression with weekday effects (few parameters, easy to estimate from four weeks). Holt–Winters can be run, but its smoothing parameters are poorly determined from so little data; a Prophet-style model would need informative priors. Flexible learners (boosting, neural networks, a GP with many hyperparameters) have too little to learn from.
A small simulated bake-off: five worlds, five forecasters core
Imagine a sports day with five different tracks: a flat sprint, a steep hill, a muddy field, a sand pit and an obstacle course. A sprinter wins the first, a climber the second, and nobody wins all five. A bake-off is a sports day for forecasting models: the same tracks (worlds), the same starting line (the forecast origin), the same stopwatch (an error metric). What you learn is not "who is best" but which assumption each track tests.
Our five worlds are small, made-up series, each built to break one assumption of a Prophet-style model while keeping the others true. The five contestants are: seasonal naive (copy last week), Holt–Winters (exponential smoothing with a weekly season), a Prophet-style regression (trend with 8 candidate changepoints in the first 80% of the history, shrunk slope changes, weekly Fourier terms, an event column when the world has known events), the same regression with an AR(1) error correction (the SARIMAX-like memory idea), and gradient boosting on weekday and event features after removing a linear trend.
Three ways to say it:
- Picture: five tracks, five different winners.
- Numbers: in the calendar world the regression's typical error over the first week is about 4, Holt–Winters' about 9; in the interaction world boosting is about 5.5 and the regression about 8.5; in the level-shift world the regression's 28-day error is about 11 against about 5 for the others.
- Slogan: no model wins everywhere; the winner is the one whose assumptions match the world.
How the scoreboard is built, with a tiny example. Three forecast days: actual $100, 104, 98$. Forecaster A says $102, 101, 103$; forecaster B says $99, 105, 97$.
- A's absolute errors: $|100 - 102| = 2$, $|104 - 101| = 3$, $|98 - 103| = 5$. Mean absolute error: $(2 + 3 + 5)/3 = 3.33$.
- B's absolute errors: $1, 1, 1$. MAE $= 1.0$. B wins this round.
- One round is luck. The widget repeats the whole experiment on 20 freshly simulated series for each world (new noise, same world) and averages the MAE.
- It scores two windows: days 1 to 7 (short horizon, where memory models shine) and days 1 to 28 (where structure matters).
- The winner of each column is marked with a star. Read the table row by row: each world has its own winner, and the reason is always one assumption.
A fair bake-off protocol (the habits behind the toy):
- Same data, same origins, same metric for every model; rolling origins rather than one split (Chapter 7.15), and report accuracy by horizon.
- Baselines first (seasonal naive, ETS). A model that cannot beat them has not earned its complexity.
- Equal effort: a model tuned by the author against a baseline left at defaults proves nothing. Give each a small, honest search.
- Judge more than accuracy: calibration of the intervals (Chapter 7.16), cost, explainability, maintenance.
- No free lunch: no method dominates all problems; performance depends on how well a model's assumptions match the process that generated the data.
Caveat for this toy. The worlds are stylised and the contestants simple (a coarse Holt–Winters grid, boosting with two features). The table shows which assumption matters in each world, not which library wins in practice. The same lesson holds on real data, but only your own rolling-origin comparison can tell you the numbers.
Why do we need it?
Arguments about models ("Prophet versus ARIMA") are stopped by a comparison on shared data. A small simulated bake-off also teaches the diagnosis: when a model loses, it is because one assumption was false in that world, and that is what you want to be able to say.
Where is it used?
Forecasting competitions (M4, M5), internal model reviews, hyper-parameter and model-family selection with rolling-origin cross-validation, and academic benchmark papers that compare new methods with seasonal naive, ETS and ARIMA.
How is it used?
Define the origins and metrics, run all candidates and the baselines, tabulate error by horizon and by series group, add calibration and cost, then pick (or combine) and write down why. Re-run when the data change.
"The winner of this table is the best forecasting model."
The winner is the best model in that made-up world, with these simple contestants. The point is the diagnosis: which assumption was true or false. On your data, run your own rolling-origin comparison.
"A model that loses in one world is a bad model."
Each model loses where its assumption is false and wins where it is true. The regression is the best in the calendar world and the worst in the late-shift world.
"If one model wins on average, use it for all series."
Check by group and horizon. A mix (a structured model for calendar-driven series, a local model for jumpy ones, a combination of both) often beats any single choice.
The level-shift world is the one to remember for your forecasting model. Its trend bends only at candidate changepoints (a Prophet-like grid plus PELT detections), usually in an early part of the history, so a shift near the end is where Holt–Winters and the naive forecast can beat it. The remedy is not to abandon the model but to know the failure and test for it: check the residual-vs-time plot and the last-28-day error (Chapter 7.17), widen the changepoint range or loosen the shrinkage if appropriate, or combine the model with a local one. The interaction world shows when a structured additive model would need an explicit interaction column, or a boosted residual model.
"Prophet is better than ARIMA."
"ARIMA is outdated."
"It depends on the assumptions. On series with several seasonalities, holidays and regressors, a structured model wins; on short horizons with strong momentum or wandering levels, ARIMA or exponential smoothing win. I compare them on the same rolling origins by horizon, with calibration and cost, and I keep the simple baselines in the race."
Model answer: "Neither is better in general. Each makes assumptions: my structured Bayesian model assumes additive components, a sparsely bending trend and independent noise; ARIMA assumes a linear stationary process after differencing. I test both on the same rolling origins and look at where each fails: at the start of the horizon, around holidays, after level shifts."
Bake-off protocol: same data, origins and metric; baselines first; equal tuning effort; by horizon and group; add calibration and cost.
Worlds and winners (toy): calendar → structured regression; late shift → local models (ETS, naive) and the regression fails at long horizons; momentum → memory (ETS, AR correction) early on; evolving season → Holt–Winters; interaction → boosting.
No free lunch: the winner is the model whose assumptions match the process. Trap: toy results are about assumptions, not products.
Quick check: in the late-level-shift world the regression has a 28-day MAE of about 11 while seasonal naive has about 5. Which assumption is violated, and what would you do first in a real project?
The assumption that the trend changes only at candidate changepoints in the early part of the history (and that the last slope continues). In a real project: look at the residual-vs-time plot and the last-28-day mean error, check the changepoint range and Laplace scale, add a local model (ETS / state-space) or a late changepoint, and re-compare on rolling origins that include the shift.
Recap, cheat sheet and practice
- Bias–variance–complexity: expected squared error $=$ bias$^2$ + variance + $\sigma^2$. More complexity lowers bias and raises variance; training error only falls, holdout error is U-shaped. For least squares with $p$ parameters and $n$ points the variance is $\sigma^2 p/n$. Choose complexity on time-ordered holdouts, not on training error.
- Where the knobs sit: Fourier order and candidate changepoints are capacity knobs; the Laplace scale $b$ and the pooling strength are strength-of-shrinkage knobs (a generous $C$ with a small $b$ is safe, with a large $b$ it is not); the guide rank controls posterior approximation quality versus cost, has no U-curve, and a low-rank guide ($d(r+2)$ numbers) only saves memory while $r \lt (d-1)/2$.
- The key question: what does my model assume, and when would another be better? A Prophet-style model assumes additive components, sparse bends of a piecewise-linear trend inside the allowed range, a fixed seasonal shape, repeatable events and known regressors, independent noise, a stable future, a good posterior and one series at a time. Each has a residual symptom and an alternative.
- ARIMA / SARIMA, ETS: read the recent past; win with momentum, wandering levels, short horizons, little data; lose with several seasonalities, moving holidays, complex regressors, counts, long horizons. Memory advantage under AR(1) noise: error ratio $\sqrt{1 - \varphi^{2h}}$.
- State-space: a hidden level / slope / season moves by small random pushes; the Kalman filter updates by a gain; the local level model is exponential smoothing with $\alpha = (\sqrt{q^2 + 4q} - q)/2$. Wins when components drift, data are missing, you need online updating.
- Gaussian processes: a random smooth curve fixed by a kernel; the kernel is the assumption; cost $O(n^3)$; wins on small data and smooth shapes, reverts to the mean beyond the data unless the kernel extends.
- Gradient boosting: trees on features you supply; cannot extrapolate (predictions stay essentially within the target range), needs horizon-safe features, detrending and a plan for intervals; wins with many series, rich drivers and interactions.
- Neural forecasting (N-BEATS, TFT; awareness): global models whose advantage comes mostly from sharing strength across many related series; evidence is mixed; validate against ETS and seasonal naive on rolling origins.
- Choosing: the situation selects candidates (how many series, history, events, shifts, momentum, explanation, intervals, interactions); a rolling-origin bake-off with accuracy and calibration selects the winner. No model wins every world.
Cheat sheet
| Family | Core assumption | Beats a Prophet-style model when… | Loses when… | Python |
|---|---|---|---|---|
| Seasonal naive | this week looks like last week | very short history; as a benchmark | trend, noise, events | np.tile(y[-7:], k) |
| ETS / Holt–Winters | level, trend, season evolve by smoothing | unplanned shifts, one season, scale | several seasons, events, counts | ExponentialSmoothing, ETSModel |
| ARIMA / SARIMA(X) | linear stationary process after differencing, one season | momentum, short horizons, wandering level | several seasons, moving holidays, long horizons | SARIMAX, pmdarima.auto_arima |
| State-space (structural) | linear-Gaussian hidden components that drift | drifting season or trend, missing data, online | many sparse events, non-Gaussian counts | UnobservedComponents |
| Gaussian process | smooth function from a chosen kernel | small data, smooth shape, irregular times | long series, trend extrapolation, counts | GaussianProcessRegressor |
| Gradient boosting | target is a function of supplied features | many series, nonlinear drivers, interactions | one short series, trends, explanation, intervals | HistGradientBoostingRegressor, LightGBM |
| Neural global (N-BEATS, TFT) | patterns learned from many series persist | thousands of related series, rich covariates | few or short series, need for priors and explanation | GluonTS, Darts, NeuralForecast |
import warnings
import numpy as np
import statsmodels.api as sm
from statsmodels.tsa.holtwinters import ExponentialSmoothing, SimpleExpSmoothing
from statsmodels.tsa.statespace.sarimax import SARIMAX
from statsmodels.tsa.statespace.structural import UnobservedComponents
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.gaussian_process import GaussianProcessRegressor
from sklearn.gaussian_process.kernels import ExpSineSquared, RBF, WhiteKernel, ConstantKernel as Ck
warnings.filterwarnings("ignore") # statsmodels re-enables its own warnings on import, so filter afterwards
def fourier(t, period, order):
cols = []
for k in range(1, order + 1):
cols += [np.sin(2 * np.pi * k * t / period), np.cos(2 * np.pi * k * t / period)]
return np.column_stack(cols) if cols else np.zeros((len(t), 0))
# --- 1) bias-variance by Fourier order, EXACT for least squares (period 28, 56 days, noise sd 6) ---
P, n, sig = 28, 56, 6.0
tt = np.arange(n)
truth = lambda t: 20*np.sin(2*np.pi*t/P) + 8*np.cos(2*np.pi*t/P) + 9*np.sin(4*np.pi*t/P + 0.5) + 5*np.cos(6*np.pi*t/P)
for N in [0, 1, 2, 3, 6, 12]:
X = np.column_stack([np.ones(n), fourier(tt, P, N)])
Hm = X @ np.linalg.pinv(X) # hat matrix
bias2 = np.mean((Hm @ truth(tt) - truth(tt)) ** 2) # squared bias = harmonics the model cannot draw
var = sig**2 * np.trace(Hm) / n # variance = sigma^2 * p / n
print(N, X.shape[1], round(bias2, 1), round(var, 2), round(bias2 + var + sig**2, 1))
# printed columns: order N, parameters p = 2N + 1, bias^2, variance, expected squared error of a new noisy day
# 0 1 285.0 0.64 321.6
# 1 3 53.0 1.93 90.9
# 2 5 12.5 3.21 51.7
# 3 7 0.0 4.5 40.5 <- lowest: the true order
# 6 13 0.0 8.36 44.4
# 12 25 0.0 16.07 52.1
# --- 2) one trending weekly series with AR(1) leftovers: seven forecasters, 28-day holdout ---
rng = np.random.default_rng(3)
n, H = 196, 28
t = np.arange(n + H)
e = np.zeros(n + H)
for i in range(1, n + H):
e[i] = 0.6 * e[i - 1] + rng.normal(0, 4)
y = 50 + 0.4*t + 12*np.sin(2*np.pi*t/7) + 5*np.cos(4*np.pi*t/7) + e
ytr, yte = y[:n], y[n:]
mae = lambda f: round(float(np.mean(np.abs(yte - f))), 2)
out = {}
out["seasonal naive"] = mae(np.tile(ytr[-7:], 4))
out["Holt-Winters"] = mae(ExponentialSmoothing(ytr, trend="add", seasonal="add", seasonal_periods=7).fit().forecast(H))
X = sm.add_constant(np.column_stack([t, fourier(t, 7, 2)]))
out["regression + Fourier"] = mae(sm.OLS(ytr, X[:n]).fit().predict(X[n:]))
glsar = sm.GLSAR(ytr, X[:n], rho=1)
res = glsar.iterative_fit(10) # regression with AR(1) errors
last_res = (ytr - X[:n] @ res.params)[-1]
out["regression + AR(1) errors"] = mae(X[n:] @ res.params + glsar.rho[0] ** np.arange(1, H + 1) * last_res)
sarimax = SARIMAX(ytr, exog=X[:n, 1:], order=(1, 0, 0), trend="c").fit(disp=False)
out["SARIMAX (AR(1) errors)"] = mae(sarimax.forecast(H, exog=X[n:, 1:]))
Xf = np.column_stack([t, t % 7]) # boosting features: day number, weekday
gb = HistGradientBoostingRegressor(max_depth=3, learning_rate=0.1, max_iter=100, random_state=0).fit(Xf[:n], ytr)
out["boosting on [t, weekday]"] = mae(gb.predict(Xf[n:]))
lin = np.polyfit(t[:n], ytr, 1) # detrend first, give the trees the remainder
gb2 = HistGradientBoostingRegressor(max_depth=3, learning_rate=0.1, max_iter=100, random_state=0).fit(Xf[:n, 1:], ytr - np.polyval(lin, t[:n]))
out["boosting on detrended y"] = mae(np.polyval(lin, t[n:]) + gb2.predict(Xf[n:, 1:]))
for k, v in out.items():
print(f"{k:28s}{v}")
# holdout MAE: seasonal naive 6.02 | Holt-Winters 4.43 | regression + Fourier 4.79 | + AR(1) errors 4.55 | SARIMAX 4.53 | boosting on [t, weekday] 5.85 | boosting on detrended y 4.57
print(round(float(gb.predict(Xf[n:]).max()), 1), round(float(ytr.max()), 1), round(float(yte.max()), 1)) # 135.0 142.2 144.7: trees stay below the training maximum
# --- 3) Gaussian process: the kernel decides what happens beyond the data ---
ts = np.arange(0, 60, 2.0)
f = lambda t: 10*np.sin(2*np.pi*t/28) + 4*np.sin(2*np.pi*t/14 + 1)
yg = f(ts) + np.random.default_rng(5).normal(0, 1.5, len(ts))
tg = np.arange(60, 91, dtype=float)
kernels = {"RBF": Ck(64.0) * RBF(5.0) + WhiteKernel(2.0),
"periodic x slow decay": Ck(64.0) * ExpSineSquared(1.0, 28.0, periodicity_bounds="fixed") * RBF(100.0, length_scale_bounds="fixed") + WhiteKernel(2.0)}
for name, k in kernels.items():
gp = GaussianProcessRegressor(kernel=k, normalize_y=True, n_restarts_optimizer=2, random_state=0).fit(ts[:, None], yg)
print(name, round(gp.log_marginal_likelihood_value_, 1), round(float(np.sqrt(np.mean((gp.predict(tg[:, None]) - f(tg)) ** 2))), 2))
# RBF -15.2 6.96 (evidence, forecast RMSE: falls back to the mean)
# periodic x slow decay -7.8 1.59 (higher evidence, continues the wave)
# --- 4) the local-level state-space model IS exponential smoothing: alpha = (sqrt(q^2 + 4q) - q) / 2 ---
r5 = np.random.default_rng(5)
ylev = np.cumsum(r5.normal(0, 0.5, 300)) + r5.normal(0, 1.0, 300) # level sd 0.5, noise sd 1 -> q = 0.25
uc = UnobservedComponents(ylev, "llevel").fit(disp=False)
q = uc.params[1] / uc.params[0] # sigma2.level / sigma2.irregular
alpha_formula = (np.sqrt(q**2 + 4*q) - q) / 2
alpha_ses = SimpleExpSmoothing(ylev, initialization_method="estimated").fit().params["smoothing_level"]
print(round(q, 3), round(alpha_formula, 3), round(float(alpha_ses), 3)) # 0.268 0.401 0.399 q-hat, alpha from the formula, alpha estimated by exponential smoothing (true q = 0.25 -> alpha 0.39)
# --- 5) guide parameter counts in NumPyro (d = 40 latent values) ---
import jax, jax.numpy as jnp, numpyro, numpyro.distributions as dist
from numpyro.infer import SVI, Trace_ELBO
from numpyro.infer.autoguide import AutoNormal, AutoMultivariateNormal, AutoLowRankMultivariateNormal
d = 40
def model():
numpyro.sample("theta", dist.Normal(jnp.zeros(d), 1.0).to_event(1))
def stored(guide):
svi = SVI(model, guide, numpyro.optim.Adam(0.01), Trace_ELBO())
return sum(int(v.size) for v in svi.get_params(svi.init(jax.random.PRNGKey(0))).values())
print(stored(AutoNormal(model)), stored(AutoMultivariateNormal(model)), stored(AutoLowRankMultivariateNormal(model, rank=5)))
# 80 1640 280 AutoMultivariateNormal stores a full d x d array (upper triangle fixed at 0); its FREE numbers are d + d(d+1)/2:
print(2 * d, d + d * (d + 1) // 2, d * (5 + 2)) # 80 860 280 mean-field, full-rank, low-rank of rank 5
1. As you increase the Fourier order of a seasonal model fitted to a fixed training window, which statement is correct?
2. A colleague says "let's lower the rank of the low-rank guide to avoid overfitting". What is the best reply?
3. A boosted-tree model is trained on daily sales whose maximum was 146. The true series keeps rising. What will its forecasts do?
4. Which situation most favours exponential smoothing or ARIMA over a Prophet-style model?
5. A local-level model has level-shock variance equal to the observation-noise variance ($q = 1$). The settled Kalman gain, i.e. the exponential-smoothing weight $\alpha$, is
6. Which statement about neural forecasters such as N-BEATS and TFT is the most accurate?
Practice problems
A. Noise variance $\sigma^2 = 25$, $n = 50$ days. A model with $p = 3$ parameters has squared bias 40, one with $p = 7$ has squared bias 2, one with $p = 17$ has squared bias 0. Compute the expected squared error of a new noisy day for each and choose.
Variance $= \sigma^2 p/n = 25p/50 = p/2$: $1.5$, $3.5$, $8.5$. Expected error $=$ bias$^2$ + variance + $\sigma^2$: $40 + 1.5 + 25 = 66.5$; $2 + 3.5 + 25 = 30.5$; $0 + 8.5 + 25 = 33.5$. Choose $p = 7$ (30.5): it has a little bias but much less variance than the 17-parameter model. The 17-parameter model has the lowest bias and the highest variance.
B. A model has $d = 500$ latent parameters. How many numbers do a mean-field, a full-rank and a rank-10 low-rank guide have? Up to which rank is the low-rank guide smaller than the full-rank one?
Mean-field: $2d = 1\,000$. Full-rank: $d + d(d+1)/2 = 500 + 125\,250 = 125\,750$ free numbers. Low-rank, rank 10: $d(r + 2) = 500 \times 12 = 6\,000$ (about 4.8% of full-rank). The low-rank guide is smaller than the full-rank one while $d(r+2) \lt d + d(d+1)/2$, i.e. $r \lt (d - 1)/2 = 249.5$: up to rank 249. Beyond that it has more numbers and no advantage.
C. Local-level model with $q = 0.04$ and with $q = 4$. Find the settled weight $\alpha$ and the level memory $1/\alpha$ for each.
$q = 0.04$: $\alpha = (\sqrt{0.0016 + 0.16} - 0.04)/2 = (0.402 - 0.04)/2 = 0.181$; memory $\approx 5.5$ days: a slow, smooth estimate. $q = 4$: $\alpha = (\sqrt{16 + 16} - 4)/2 = (5.657 - 4)/2 = 0.828$; memory $\approx 1.2$ days: the estimate follows the data almost step by step.
D. Leftover noise is AR(1) with $\varphi = 0.9$. By what factor does a memory correction reduce the typical error at horizons 1, 3 and 7 days? What does that say about when ARIMA beats your model?
Ratio $= \sqrt{1 - \varphi^{2h}}$: $h = 1$: $\sqrt{1 - 0.81} = 0.436$ (56% smaller); $h = 3$: $\sqrt{1 - 0.531} = 0.685$ (31% smaller); $h = 7$: $\sqrt{1 - 0.229} = 0.878$ (12% smaller). With strong momentum the memory advantage lasts a week or more, so a memory model (or an AR error term) can beat an independent-noise calendar model at short horizons. At 28 days $\varphi^{56} \approx 0.003$ and the advantage is gone.
E. Interview: "Why didn't you just use a neural network for your forecasts?"
"A neural forecaster gets its edge from learning across many related series with plenty of history. My problem is a calendar-driven series (or a few) with holidays, regressors and counts, where I need explainable components, honest uncertainty and priors for short stretches of data. A structured Bayesian regression gives me that. I would consider a global neural or boosting model if I had thousands of related series with rich covariates, and I would still benchmark it against seasonal naive and ETS on the same rolling origins, with calibration, before adopting it."
F. Interview: "What assumption does your forecasting model make, and when would another model be better?"
"It assumes additive components: a piecewise-linear trend that bends rarely inside the allowed changepoint range, periodic seasonality of a fixed shape (Fourier terms), repeatable holiday effects and known regressors, and independent noise with a Normal, Student-t or Negative Binomial likelihood. That fits long, stable, calendar-driven series. I would look at exponential smoothing or a state-space model if the level shifts without warning or the seasonal shape drifts; at ARIMA errors if the residual ACF shows momentum and the horizon is a few days; at global boosting or a neural model if I had many related series or strong nonlinear drivers; and at a Gaussian process for small, smooth data. I check each assumption with the residual panel and decide by rolling-origin accuracy and calibration."
Production statistical modeling
A model that works in your notebook has passed only the first test. In production it must be rebuilt next month by someone else, watched while the world changes under it, and kept from breaking on numbers that are too small, too big or not numbers at all. This chapter gives you three habits: write everything down (reproducibility), keep looking (monitoring), and keep the arithmetic safe (numerical stability).
- List the ten things a run manifest must record (data, preprocessing, model version, priors, seed, inference settings, library versions, guide type, optimizer, iterations) and say what goes wrong when one is missing
- Explain what a random seed does and does not guarantee, and use the seed-to-seed spread as the yardstick before claiming "model B is better"
- Monitor a live forecast: error over time, residual drift, interval coverage, data and feature drift, missing regressors, changepoint frequency, with alarm rules that balance speed against false alarms
- Work with log probabilities: see why products of probabilities underflow, and compute sums of probabilities safely with log-sum-exp
- Handle constrained parameters (positive, between 0 and 1, probability vectors) with transforms such as exp, softplus and sigmoid, and know what each does to gradients
- Prevent and debug bad initial values, exploding gradients and NaNs, and explain why scaling the inputs is the cheapest fix
Reproducibility: write the recipe down core
A cake recipe that says "flour, sugar, eggs" is useless. A good one says "200 g flour, oven at 180 °C, 35 minutes". A fitted model is a cake too. The result depends on the ingredients (the data), the recipe (code, priors, the form of the model) and the oven (settings, software versions, the random numbers used).
Three months from now somebody asks, "Why did the forecast say 4,000 orders for that Friday?" or "Which model produced this number?". If you cannot rebuild the result, you cannot explain it, debug it or improve it. A result nobody can rebuild is closer to a rumour than to a measurement.
Three ways to say it:
- Picture: a label glued to every model: "made from this data, by this code, with these settings, on this software, on this day".
- Numbers: with 10 things that can differ, two runs can differ in $2^{10} - 1 = 1023$ different ways. A saved record turns "search 1023 possibilities" into "read one list of differences".
- Slogan: if it is not written down, it did not happen.
On Monday your forecasting job reports a holdout error (MAE) of 4.1 orders. A week later, "the same code" reports 4.9. Nobody changed anything, they say.
- Open the two saved manifests (small text files written at the end of each run) and compare them line by line.
- Only one line differs: the data hash. A late backfill corrected three old days, so the training table is not the table of last week.
- Check the claim: run last week's code on last week's data hash. You get 4.1 again, so the code is fine and the run is reproducible.
- Run this week's data with last week's settings. You get 4.9. The data changed, not the model.
- Cost: one comparison and two re-runs, about 2 hours. Without the manifest you would test the 10 suspects one at a time, up to 10 hours, and still might not find it (if the cause is not on your list).
A run manifest is a small text file (JSON or YAML) saved next to every fitted model. It records everything needed to rebuild the same result and to compare two runs. Reproducible means: given the manifest, someone else gets the same numbers (up to tiny floating-point differences, see the next section).
| Group | Field (syllabus Module 83) | Example value (illustrative) |
|---|---|---|
| Data | data version | hash of the training table, history cut-off date, row count |
| preprocessing | missing-value rule, filters, scaler statistics learned from the history, transforms | |
| Model | model version | code commit id, list of components (trend, Fourier orders, holidays, regressors), likelihood |
| prior configuration | every prior scale, for example the Laplace scale $b$ of the changepoint slopes | |
| Inference | random seed | the seed (or PRNG key) of every random step |
| inference settings | SVI or NUTS, learning rate, tolerance, patience, minimum steps | |
| guide type, optimizer | full-rank or low-rank (and the rank), Adam and its settings | |
| number of iterations | steps actually run, the step of the best state, the stopping reason | |
| Environment | library versions | Python, NumPy, JAX, jaxlib, NumPyro; float32 or float64; the device |
Good manifests are written automatically (never by hand), are small and diff-able (text, not a binary), and travel with the saved parameters.
Why do we need it?
Results that cannot be rebuilt cannot be debugged, audited, compared fairly or trusted. When a forecast changes between two runs, a manifest turns "something changed" into one line you can read.
Where is it used?
Experiment trackers (MLflow, Weights & Biases), data versioning (DVC, dataset hashes), Git commit ids and lock files (pip freeze), model cards, and every team that retrains a forecast weekly or must explain a number to an auditor.
How is it used?
At the end of every run write one JSON file: data hash, config, seed, library versions (from importlib.metadata), steps run. To compare two runs, diff the two files. To check reproducibility, rebuild from the file alone and compare the numbers.
"I fixed the random seed, so my result is reproducible."
A seed is one of ten ingredients. The same seed with different data, a different library version or a different number of iterations gives a different answer. Reproducible means the whole manifest is fixed.
"I saved the model file, so the run is recorded."
A file of fitted parameters does not say which data, which priors or which software produced it. Save the manifest beside it, and keep the training data (or an immutable snapshot) addressable by its hash.
"Reproducible means identical to the last digit on every machine."
Aim for: same numbers on the same software and hardware, and the same conclusion elsewhere. Parallel sums, float32 and other chip types change the last digits (next section).
In your forecasting model several choices are made automatically, and an automatic choice is exactly what a later reader cannot guess: the full-rank or low-rank guide chosen from the model size, the changepoint candidates from the grid and from PELT, the scaler statistics. Record the guide that was picked and the model size that triggered it. If your loop uses relative-ELBO early stopping with patience and best-state checkpointing, record the three settings, the step where it stopped, and the step of the best state. In the A/B framework, record the data snapshot of each experiment, the global scaler's statistics and the prior parameters of the Beta or Dirichlet priors, so a "B beats A" result can be rebuilt.
"It is reproducible, I set the seed."
"Every run writes a manifest: data hash and cut-off, preprocessing, code commit, prior scales, seed, inference settings, guide and optimizer, iterations run, and library versions. From that file I can rebuild the same forecast, and I can diff two manifests to find out why two forecasts differ."
Manifest = data hash + preprocessing + code version + priors + seed + inference settings + guide + optimizer + iterations run + library versions.
Written automatically at the end of each run, saved beside the parameters, compared with a diff.
Trap: a seed alone is not reproducibility; record the iterations actually run, not only the maximum.
Quick check: two runs differ only in the number of SVI steps they ran (stopped at step 900 and at step 1500 by early stopping). Which manifest lines explain it?
The "iterations actually run" and the stopping settings (tolerance, patience, minimum steps), and often the seed too, because a noisy loss makes the stopping step depend on the random numbers. Recording only "maximum steps = 5000" would hide the difference, which is why the manifest stores what actually happened.
Seeds, randomness and "the same result" core
Training your model rolls dice. The starting values of the parameters are random. The noisy estimate of the objective (the ELBO of Chapter 6.12) uses random draws. Forecast draws are random too. A seed is a note of how the dice were set at the start: the same note gives the same rolls, a different note gives different rolls and a slightly different model.
Neither answer is "wrong". The gap between them tells you how much of what you are looking at is dice. If model B beats model A by less than the gap between two seeds of model A, you have not learned anything yet.
Three ways to say it:
- Picture: train the same model with five seeds and you get a thin bundle of five slightly different forecast lines, not one line.
- Numbers: model A scores 4.30, 4.10, 4.00, 4.20, 4.40 over five seeds (mean 4.20, spread 0.16). A gap of 0.15 between A and B is smaller than that spread.
- Slogan: a difference smaller than the seed-to-seed spread is not a difference.
Five seeds, two models, holdout MAE (orders per day):
- Model A: 4.30, 4.10, 4.00, 4.20, 4.40. Mean $= 21.0/5 = 4.20$.
- Deviations from 4.20: $+0.10, -0.10, -0.20, 0, +0.20$. Squares: $0.01, 0.01, 0.04, 0, 0.04$, sum $0.10$. Divide by $n - 1 = 4$: $0.025$. Seed spread $s_A = \sqrt{0.025} = 0.158$.
- Model B (one extra regressor): 4.15, 3.95, 3.85, 4.05, 4.25. Mean $= 20.25/5 = 4.05$, and $s_B = 0.158$ as well.
- Gap $= 4.20 - 4.05 = 0.15$, which is $0.15/0.158 \approx 0.95$ seed spreads.
- Each mean of 5 seeds has its own wobble $0.158/\sqrt 5 = 0.071$, so the gap between two means has $SE = 0.158\sqrt{2/5} = 0.10$. The gap is only $0.15/0.10 = 1.5$ standard errors: suggestive, not convincing. Run 20 seeds and several forecast origins before you decide.
A pseudo-random number generator (PRNG) produces numbers that look random from a starting value, the seed. The same seed always gives the same sequence. NumPy: np.random.default_rng(seed). JAX has no hidden global generator: you pass an explicit key, jax.random.PRNGKey(seed), and make new keys with jax.random.split (Chapter 6.16).
- Where randomness enters a Bayesian forecast pipeline: initial parameter values; Monte Carlo draws inside the ELBO (reparameterized samples); mini-batch choice if you use batches; MCMC initial points and proposals; drawing forecast samples from the posterior predictive.
- Not random: PELT and a grid of candidate changepoints (same data, same penalty, same answer).
- "Same seed" is not "same bits". Computers add numbers in a different order on a different chip, library version or number of threads; floating-point addition is not exactly associative. The last digits move, and an iterative optimizer can amplify a tiny difference. float32 and float64 runs also differ.
Why do we need it?
Without a fixed seed you cannot repeat a run to debug it. Without several seeds you cannot tell a real improvement from luck of the dice. Both matter, and they are two different jobs.
Where is it used?
JAX and NumPyro (rng_key in SVI.run and Predictive), scikit-learn's random_state, benchmark papers that report "mean ± sd over 5 seeds", and tracker logs such as MLflow parameters.
How is it used?
Fix one seed to debug. Then train with 5 to 20 seeds, report mean ± spread, and compare models by their gap relative to that spread. Split keys, never reuse one: k1, k2 = jax.random.split(key).
"Run it with a few seeds and pick the best one for the report."
That selects noise, and the number will not repeat. Report the mean and spread over seeds, and choose the final model by its average score (or train the final one with the seed fixed in advance).
"The same seed gives the same result on any machine and any library version."
It gives the same random numbers only if the generator is unchanged; the arithmetic that follows can differ in the last digits (different chips, thread counts, float32 vs float64, library updates). Treat tiny differences as expected, and large ones as a bug or an unrecorded change.
"Seeds only matter for neural networks."
Any stochastic fitting method depends on them: SVI's starting values and ELBO draws, MCMC starting points, and the draws used to build a forecast fan.
"If seed spread is small, the model is certain."
Seed spread measures only the optimizer's randomness. The uncertainty from having a finite history (the posterior width, Chapter 7.14) and from which days are in the test set (several forecast origins, Chapter 7.15) are different things.
Your custom SVI loop draws fresh random numbers at every step, so its ELBO is a noisy number. If it compares each ELBO with the best one seen so far, then the step where it stops, and the step of the saved "best state", both depend on the seed. When you report that a changepoint scale or a guide rank is better, run a few seeds and give the spread. In the A/B framework the same holds for a decision such as $P(\theta_B > \theta_A \mid D) = 0.97$: with SVI draws it has Monte Carlo noise, so check that two seeds agree to the precision you quote (for example, two decimals).
"How do you know model B is better than model A?" — "Its holdout MAE is 0.15 lower."
"Lower on average over several seeds and over several rolling forecast origins, by more than the seed-to-seed spread, with the same splits for both models."
Seed = starting note for the dice. Same seed, same code and data: same random numbers; the last digits can still move.
Compare models by (gap) ÷ (seed-to-seed spread), using 5 to 20 seeds; $SE_{\text{gap}} = s\sqrt{2/m}$ for $m$ seeds each.
Trap: reporting the best seed; trusting one run.
Quick check: you train with 10 seeds. Model A: mean 4.2, spread 0.3. Model B: mean 4.0, spread 0.3. What do you report?
The gap is 0.2, which is $0.2/0.3 \approx 0.67$ seed spreads, and its standard error is $0.3\sqrt{2/10} = 0.134$, so the gap is about 1.5 standard errors: not convincing. Add seeds, train longer or with a smaller learning rate so the spread shrinks, and compare over several forecast origins.
Monitoring: is the forecast still as good as at launch?
A car has a dashboard: speed, fuel, a warning light. You do not open the engine every morning. You glance at a few gauges, and a light tells you when to stop and look inside. A live forecast needs gauges too, because the world keeps changing (a new competitor, a price change, a broken data feed) and the model does not know.
On launch day your backtest said "typical error: 4 orders a day". The question for every later day is: is it still about 4? If the typical error creeps up to 8, or the forecasts start to be too low day after day, something changed and somebody should look.
Three ways to say it:
- Picture: a line of weekly errors that normally wobbles around 4 and then climbs and stays high.
- Numbers: launch MAE 4.0; raise an alarm when the weekly MAE is above 6.0 (1.5 times launch) two weeks in a row.
- Slogan: good at launch is not good forever; set the gauge before the trip.
Launch MAE (from a rolling-origin backtest, Chapter 7.15) is 4.0 orders. After launch, the weekly MAE is 4.3, 3.8, 4.6, 5.1, 6.4, 7.8, 8.2. The rule is: alarm when the weekly MAE is above 1.5 times the launch MAE for two weeks in a row.
- Alarm level $= 1.5 \times 4.0 = 6.0$ orders.
- Weeks 1 to 4 (4.3, 3.8, 4.6, 5.1) are below 6.0. No alarm, even though week 4 is already creeping up.
- Week 5 is 6.4, above 6.0: the first breach. One bad week might be a one-off promotion, so we wait.
- Week 6 is 7.8, above 6.0: the second breach in a row. Alarm at the end of week 6.
- The price of the rule: the error began to rise in week 4 and the alarm came two weeks later. A looser rule would be faster and would also fire on harmless bad weeks.
Bias check with the residuals $e_t = y_t - \hat y_t$ (actual minus forecast). A 7-day average residual has a wobble of $\sigma_e/\sqrt 7$. With $\sigma_e = 5$ the limit at three wobbles is $3 \times 5/\sqrt 7 = 5.67$ orders. An average residual of $+6$ is beyond the limit: actuals sit above forecasts, so the model under-forecasts.
Each forecast is scored when its actual value arrives. For a forecast made on day $T$ for day $T + h$, that is on day $T + h$ (or later, if data arrive late). Keep the forecast that was issued then, not a recomputed one. Track these statistics over a rolling window of $w$ days:
- Rolling error: rolling MAE (or WAPE, or MASE, Chapter 7.15), compared with the launch level from the backtest and with a naive baseline. A rolling MASE above 1 means the naive forecast would have been better.
- Residual drift: the rolling mean residual $\bar e_w = \tfrac1w\sum e_t$. For a healthy model with roughly independent residuals of standard deviation $\sigma_e$ it is near 0 with standard error $\sigma_e/\sqrt w$. Limits at $\pm 3\sigma_e/\sqrt w$ are a common rule of thumb; autocorrelated residuals (Chapter 7.17) need wider limits.
- Alarm rule: a statistic beyond its limit for $k$ windows or days in a row. Shorter windows and lower limits react sooner and raise more false alarms. (Another detector with a name worth knowing: CUSUM, the running sum of residuals, which is good at catching small steady bias.)
Why do we need it?
The validation score tells you how good the model was on launch day. Data pipelines break, customers change behaviour and trends turn. Without a gauge you find out from an angry stakeholder instead of from an alert.
Where is it used?
Demand-forecast dashboards at retailers, capacity planning, control charts in factories (Shewhart and CUSUM charts), and alert rules in MLOps tooling that page a person when a metric leaves its band.
How is it used?
Each day: join the stored forecast with the arrived actual, compute the residual, update the rolling MAE or MASE and the rolling mean residual, compare with the launch level and limits, and alert when a rule fires on consecutive windows. Slice by horizon and segment.
"The validation score was good, so the model is good."
Validation describes launch day. Monitoring describes today. The two can drift apart without any code change.
"An alarm means we must retrain."
An alarm is a question. Often the cause is a broken data feed or a missing regressor (fix the pipeline), a one-off event the model does not know (add it as a holiday), or a real change (refresh the changepoints and refit). Retraining on bad inputs makes things worse.
"Watch the average error only."
The average hides direction and pattern. Also track the signed residual (bias), the error by horizon and by segment, and the error against the naive baseline. In the widget, a stronger weekly pattern raises the error but leaves the 7-day average residual near zero.
"Score each day's forecast against a recomputed forecast with today's model."
Store the forecast as it was issued. Re-scoring with a newer model or with revised data hides the real live error.
Your forecasting model fixes its changepoint candidates when it is fitted. If the world turns after the last one, the model keeps extrapolating the last slope, and the residual monitor shows a steady drift like the "trend turns" case above: the signal to refresh the changepoints or refit. Score forecasts with the same rolling-origin logic you used to validate, so the launch level is comparable. In the A/B framework the same habit applies to the inputs of the decision: watch the baseline conversion rate behind your prior, the assignment ratio (a sample ratio mismatch alarm, Chapter 5.11) and the metric definition, because a changed metric silently changes what $P(\theta_B \gt \theta_A \mid D)$ means.
"How do you know your forecast is still good in production?" — "We checked accuracy when we launched it."
"Every forecast is stored as issued and scored when its actual arrives. I track rolling MAE or MASE against the launch backtest and the naive baseline, the rolling mean residual for bias, and the interval coverage. A rule fires after two bad windows in a row, and the alert links to the investigation steps: data first, then events, then the model."
Residual $e_t = y_t - \hat y_t$; rolling MAE $=\tfrac1w\sum|e_t|$; bias limits $\pm 3\sigma_e/\sqrt w$ (rule of thumb, independent residuals).
Alarm = a statistic beyond its limit for $k$ windows in a row; a short window is fast but noisy.
Trap: launch accuracy is not current accuracy; an alarm is a question, not a retrain order.
Quick check: the 7-day mean residual has $\sigma_e = 5$ and reads $+4$. Alarm at $\pm 3\sigma_e/\sqrt 7$?
The limit is $3 \times 5/\sqrt 7 = 15/2.646 = 5.67$. A reading of $+4$ is inside the limit, so no alarm yet. Still, it is $4/(5/\sqrt 7) = 2.1$ standard errors above zero, so it is worth watching: if it stays there for a few days, a lower limit (or a CUSUM) would catch it earlier.
Monitoring: are the intervals still honest? (coverage and calibration drift)
A weather forecaster says "80% chance of rain" on 100 days. On about 80 of those days it should rain. If it rains on only 55, she is not just a little off: her numbers are misleading. A forecast interval is the same kind of promise. An 80% interval says "about 8 days in 10, the actual value lands inside". Coverage checks whether the promise is kept.
This matters for decisions. Safety stock, staffing and capacity are set from the edges of the interval. If the interval quietly becomes too narrow, you run out of stock. If it becomes too wide, you pay for stock you do not need.
Three ways to say it:
- Picture: a row of 10 intervals and the 10 actual values: about 8 dots should sit inside their intervals.
- Numbers: over 28 days an honest 80% interval misses about 5.6 days. Twelve misses has a chance of only about 0.5% if the interval is honest.
- Slogan: an interval is a promise; coverage checks whether it is kept.
An 80% interval is checked over the last 28 days. It missed on 12 days.
- An honest 80% interval misses on 20% of days. Expected misses in 28 days: $28 \times 0.2 = 5.6$.
- The count of misses behaves like a Binomial with $n = 28$, $p = 0.2$: standard deviation $\sqrt{28 \times 0.8 \times 0.2} = 2.12$.
- Observed coverage: $(28 - 12)/28 = 16/28 = 0.571$, not 0.80.
- Twelve misses is $(12 - 5.6)/2.12 = 3.0$ standard deviations above the expected 5.6. The exact chance of 12 or more misses for an honest interval is $0.005$ (about 1 in 200).
- Which way? If 9 misses were above the interval and 3 below, the forecasts are too low (bias). If 6 and 6, the interval is too narrow on both sides: the noise grew.
- How big a change is that? If the true noise became 1.6 times larger than the model assumes, the coverage of a Normal 80% interval falls to $2\Phi(1.2816/1.6) - 1 = 0.577$, which matches what we saw.
- Coverage of a central $(1-\alpha)$ interval: the fraction of actuals that fall inside it. Nominal coverage is $1-\alpha$.
- Calibration: the observed frequencies match the stated probabilities at every level (50%, 80%, 95% intervals), not only at one level.
- PIT (probability integral transform): $u_t = F_t(y_t)$, the forecast's own cumulative distribution function evaluated at the actual value. For a calibrated forecast the $u_t$ are uniform on (0, 1). For a Normal forecast $N(\hat y_t, \sigma_m^2)$: $u_t = \Phi\big((y_t - \hat y_t)/\sigma_m\big)$. Flat histogram: calibrated. U-shape: too narrow (overconfident). Hump: too wide (underconfident). Lopsided: biased. (Details in Chapter 7.16.)
- Calibration drift: coverage or the PIT shape moves away from nominal over time. A rolling window of $w$ days has a misses count that is Binomial$(w, \alpha)$ for an honest interval, which gives exact alarm counts. (The sampling wobble of the coverage is about $\sqrt{(1-\alpha)\alpha/w}$ when days are roughly independent; forecasts that are correlated across days need wider limits.)
Why do we need it?
Point-forecast accuracy can look fine while the uncertainty statement is wrong. Decisions such as stock levels, staffing and capacity use the interval edges, so a wrong interval costs money even if the middle line is good.
Where is it used?
Inventory and capacity planning, energy and demand forecasting, risk limits (value-at-risk back-testing counts exceedances in exactly this way), and any probabilistic forecast whose quantiles drive a decision.
How is it used?
Each day record whether the actual fell inside the stored interval (and on which side). Track the rolling coverage with exact Binomial limits, and keep a PIT histogram of the recent weeks. Alert when coverage leaves the band, then check bias, noise level and missing inputs.
"A 95% interval from my model contains 95% of the actuals, by definition."
Only if the model is right. The 95% is what the model claims. Coverage is what you measure, and drifting away from the claim is exactly what you are monitoring.
"Coverage above nominal is good news."
Intervals that are too wide waste safety stock and hide useful information. A calibrated interval is neither too narrow nor too wide. Use two-sided limits.
"Only one level matters, so check 80% and stop."
Check several levels (50%, 80%, 95%) or the PIT histogram. A forecast can be right at 80% and wrong in the tails (heavy tails make 99% intervals too narrow even when 80% looks fine).
"One bad window proves the model broke."
With a 1% test per day and 140 monitored days, a few false alarms are expected. Use persistence rules or stricter limits, and look at the cause before acting.
Your forecast intervals come from posterior predictive draws, so they combine parameter uncertainty and observation noise (Chapter 7.14). Two things can make live coverage too low even if the model was fine at launch: the noise level grew (a likelihood with a fixed scale is now too tight) and a guide that is too simple shrinks the posterior (a mean-field guide under-states correlated uncertainty, Chapter 6.13). Student-t and Negative Binomial likelihoods keep tails and variance honest on spiky days (Chapter 7.13). In an A/B framework, run A/A comparisons from time to time: the share of times your 95% credible interval for a difference excludes 0 is a calibration check of the whole pipeline.
"Your model gives 90% intervals. Are they right?" — "Yes, that is what the model outputs."
"The model claims 90%. I measure it: on the rolling backtest and in production, about 9 of 10 actuals should fall inside, and the PIT histogram should be flat. If the live coverage drifts to 70%, I look at bias, noise level, missing inputs and the guide before I trust the intervals."
Coverage = share of actuals inside the interval. Misses in $w$ days $\sim$ Binomial$(w, \alpha)$ if honest; mean $w\alpha$, sd $\sqrt{w\alpha(1-\alpha)}$.
PIT $u_t = F_t(y_t)$ is uniform when calibrated: U = too narrow, hump = too wide, lopsided = bias.
Trap: the 80% is a claim, not a fact; check it at several levels and in both directions.
Quick check: 28 days, 80% interval, only 3 misses. Alarm?
Expected misses are 5.6 and 3 is a little low: coverage $25/28 = 0.89$. For an honest interval the chance of 3 or fewer misses is 0.16, so this is ordinary luck. Intervals that are really too wide would show many such windows in a row and a hump-shaped PIT histogram.
Monitoring: are the inputs still the same? (data drift, missing regressors, changepoint frequency)
A bakery's recipe is tuned for last year's flour. If the supplier changes the flour, the bread changes, although the recipe (the model) is untouched. Or the delivery truck does not arrive and somebody quietly uses "0 kg of flour" for the day. In forecasting, the inputs can shift (new kinds of customers, bigger marketing campaigns) or break (the feed for a regressor stops and the column fills with zeros).
If you watch only the final error, you learn about the damage after it happened and you do not know where it came from. Watching the inputs gives an earlier warning and points at the culprit.
Three ways to say it:
- Picture: the histogram of a regressor in the training period and the histogram of last month, no longer on top of each other.
- Numbers: a drift score (PSI) of 0.135 between them; or 31% of last week's regressor values are exactly 0, against 2% in training.
- Slogan: check the inputs before you blame the model.
Compare the e-mail volume regressor in training with the last 300 days, using 5 equal-sized bins cut at the training quintiles. In training each bin holds 20% of the days. Recently the bins hold 10%, 15%, 20%, 25%, 30%: the regressor moved to bigger values.
- The drift score PSI (population stability index) adds up $(a_i - e_i)\ln(a_i/e_i)$ over the bins, where $e_i$ is the training share and $a_i$ the recent share.
- Bin 1: $(0.10 - 0.20)\ln(0.10/0.20) = (-0.10)(-0.693) = 0.0693$.
- Bin 2: $(-0.05)\ln(0.75) = (-0.05)(-0.288) = 0.0144$.
- Bin 3: $0 \times \ln 1 = 0$. Bin 4: $0.05\ln 1.25 = 0.0112$. Bin 5: $0.10\ln 1.5 = 0.0405$.
- Total: $0.0693 + 0.0144 + 0 + 0.0112 + 0.0405 = 0.135$. A common rule of thumb calls PSI below 0.1 small, 0.1 to 0.25 moderate and above 0.25 large. So this is a moderate shift: look, do not panic.
A broken feed is easier to see. In training the regressor was 0 on 2% of days. This week 31% of values are exactly 0 (the feed is missing and the pipeline fills with 0). With a weight of 4 orders per thousand e-mails and a typical volume of 2 thousand, each affected day loses about $4 \times 2 = 8$ orders, around 7% of a 110-orders day, with no sign of trouble in the model code.
- Data drift / feature drift: the distribution of the inputs (all columns, or one column) changes. Concept drift: the relationship between inputs and the target changes while the inputs look the same (it shows up in the residuals, not in input histograms). Target drift: the distribution of $y$ itself moves (a level shift).
- PSI $=\sum_{i=1}^{k}(a_i - e_i)\ln(a_i/e_i)$ over $k$ bins fixed from the reference period. Even with no drift it is not 0: it is about $(k-1)(1/n_{ref} + 1/n_{new})$ (with $k = 10$, $n_{ref} = 1000$, $n_{new} = 300$ that is $0.039$). KS (Kolmogorov–Smirnov) distance: $\max_x|F_{ref}(x) - F_{new}(x)|$, the largest gap between the two cumulative distribution functions.
- Pipeline health checks that need no statistics: share of missing values, freshness (how old is the last value?), range (inside the training minimum and maximum?), schema (type, units).
- Changepoint frequency: how many changepoints PELT (Chapter 7.9) finds per window of the series (or of the residuals) compared with the history. A jump in the count says the world is changing more often than the model's fixed changepoint candidates allow for.
Why do we need it?
Bad inputs cause bad forecasts long before anyone can measure the error, and they show where to look. A zero-filled regressor and a real demand drop look the same in the error plot but need opposite fixes.
Where is it used?
PSI in credit-scoring model monitoring, data-quality checks (missing share, ranges, freshness) in pipelines and feature stores, covariate-shift tests such as the two-sample KS test, and changepoint monitors on sensor and demand data.
How is it used?
Pick the training period as the reference. Every day or week compare the recent window with it, column by column: missing share, range, PSI or KS. Alert from thresholds learned on past windows. When a feature alarm fires, check the upstream source first.
"The input histogram moved, so the model is broken."
Not necessarily. A shift inside the training range, with the same relationship to the target, may leave the model fine. It is a signal to look. Outside the training range a linear regressor extrapolates without any warning, which is more dangerous.
"The inputs look the same, so nothing changed."
Concept drift hides here: the same inputs now lead to different outcomes (a competitor, a pricing change). Only the residual and error monitors can see it.
"PSI should be 0 when nothing changed."
It is a noisy statistic: about $(k - 1)(1/n_{ref} + 1/n_{new})$ even with no drift. Learn the normal range from past windows before setting an alarm level, and treat 0.1 and 0.25 as rules of thumb, not laws.
"Filling missing values with 0 keeps the pipeline running, so it is harmless."
To the model 0 is a real value. A filled-in zero silently changes the forecast. Decide in advance what happens on missing inputs (skip the forecast, use the last good value, or fall back to a model without that regressor) and log it.
In your forecasting model the regressors and the holiday calendar must be available at forecast time (Chapter 7.12). In production that becomes a schedule: who supplies next week's values, how late can they arrive, and what does the model do if they are missing? Add freshness and missing-share checks for every regressor, and compare each regressor's recent distribution with the training one (PSI or KS). The changepoint candidates (grid and PELT) describe how often the trend changed in the past; re-running PELT on a window of recent residuals and counting is a cheap way to notice that the trend is now changing faster. In the A/B framework, watch the inputs of a decision in the same way: the baseline conversion rate behind your priors, the exposure counts and the share of missing metric values per variant.
"Data drift and concept drift are the same thing."
Data (feature) drift means the distribution of the inputs $P(x)$ changed. Concept drift means the relationship $P(y \mid x)$ changed. The first can be seen by comparing input histograms; the second only shows up in the errors.
Model answer: "I monitor both. Input checks (missing share, ranges, PSI) catch feed and population changes early. Residual and coverage monitors catch a changed relationship. An input alarm without an error alarm tells me to look; both together tell me to act."
PSI $=\sum (a_i - e_i)\ln(a_i/e_i)$; rule of thumb 0.1 / 0.25; noise level $\approx (k-1)(1/n_{ref} + 1/n_{new})$. KS $=\max|F_{ref} - F_{new}|$.
Data drift = inputs change; concept drift = relationship changes; also watch missing share, freshness, range and changepoint counts.
Trap: a zero-filled feed looks like real data; input drift is a prompt to look, not proof the model failed.
Quick check: a regressor feed breaks and 30% of its values become 0. Which alarm fires first, the rolling MAE or a missing-share check, and why?
The missing-share (or exact-zero) check fires on the first bad day, because it looks directly at the input. The rolling MAE needs several days of errors to rise above its level and gives no hint about the cause. This is why input checks are the early warning, and error monitors the proof.
Numerical stability 1: log probabilities, underflow and log-sum-exp
Take a calculator and multiply 0.04 by itself again and again. After 3 steps you have 0.000064. After 10 steps the display shows a tiny number with a long exponent. A computer has a floor, too: below about $10^{-324}$ it gives up and stores exactly 0. That is called underflow. The likelihood of 400 days of data is a product of 400 small numbers, so it falls through the floor, and "the likelihood is 0" ends the story, even though the model may be perfectly good.
The fix is a change of units. Instead of the number we store its logarithm, which turns a product into a sum: $\log(a \times b) = \log a + \log b$. A sum of 400 logs is a modest number such as $-1287.55$. Nothing falls off the floor.
Three ways to say it:
- Picture: a staircase of 400 steps down: in ordinary units the bottom is below the ground floor; in log units it is just a long number line.
- Numbers: $0.04^{400} = 10^{-559}$ is stored as 0 in double precision, while $400 \times \ln 0.04 = -1287.55$ is stored exactly enough.
- Slogan: work in logs; leave the logs only at the very end, and then with log-sum-exp.
Part 1: the product. 400 days, each with density about 0.04 under the model.
- Product of 400 factors: $0.04^{400}$. In base 10, $\log_{10} 0.04 = -1.398$, so $400 \times (-1.398) = -559.18$: the product is $10^{-559.18}$.
- Double precision (float64) cannot store anything below about $5 \times 10^{-324}$. The product reaches that floor after 232 factors, and from there on it is exactly 0. Single precision (float32, the JAX default) hits its floor, about $1.4 \times 10^{-45}$, after only 33 factors.
- Sum of logs: $400 \times \ln 0.04 = 400 \times (-3.2189) = -1287.55$. Easy to store, and exact enough.
Part 2: adding probabilities. A mixture, a sum over posterior draws or a softmax needs $\log(e^{a_1} + e^{a_2})$ with $a_1 = -1000$, $a_2 = -1001$. Computing $e^{-1000}$ gives 0, so the naive answer is $\log 0 = -\infty$. Instead pull out the biggest term:
- $\log(e^{-1000} + e^{-1001}) = \log\big(e^{-1000}(1 + e^{-1})\big) = -1000 + \log(1 + e^{-1})$.
- $e^{-1} = 0.3679$, so $\log(1.3679) = 0.3133$.
- Answer: $-1000 + 0.3133 = -999.6867$. This is log-sum-exp. All the exponentials inside are of numbers $\le 0$, so none can overflow, and the biggest one equals $e^0 = 1$, so the sum is never zero.
Part 3: a posterior-predictive score. For one test point, three posterior draws give log-likelihoods $-1000, -1001, -1002$. The log of the average likelihood is $\mathrm{LSE} - \log 3 = -1000 + \log\frac{1 + e^{-1} + e^{-2}}{3} = -1000 + \log 0.5011 = -1000.691$. The average of the logs is $-1001$, which is not the same (it is smaller; Jensen's inequality).
- Log-likelihood. For independent data, $p(D \mid \theta) = \prod_i p(y_i \mid \theta)$, so $\log p(D \mid \theta) = \sum_i \log p(y_i \mid \theta)$. For continuous data $p$ is a density, and a density can exceed 1 (a Normal with $\sigma = 0.01$ has peak density 39.9, log density $+3.69$), so log densities are not always negative.
- Underflow / overflow. Float64 stores positive numbers from about $2.2\times10^{-308}$ (normal) or $5\times10^{-324}$ (the smallest) up to $1.8\times10^{308}$. Float32: $1.2\times10^{-38}$ (normal), $1.4\times10^{-45}$ (smallest), up to $3.4\times10^{38}$. $e^{x}$ overflows to $\infty$ for $x \gt 709.8$ in float64 and $x \gt 88.7$ in float32.
- Log-sum-exp. $\mathrm{LSE}(a_1,\dots,a_K) = \log\sum_k e^{a_k} = m + \log\sum_k e^{a_k - m}$ with $m = \max_k a_k$. Then softmax is $e^{a_k - \mathrm{LSE}(a)}$, and the log of an average is $\mathrm{LSE}(a) - \log K$.
- Other stable primitives:
log1p(x)$=\log(1+x)$ andexpm1(x)$=e^x - 1$ (accurate for tiny $x$),logaddexp,log_softmax,log_sigmoid, and parameterizing a Bernoulli or Binomial by logits rather than probabilities.
Why do we need it?
Likelihoods are products of hundreds or thousands of tiny numbers, and exponentials overflow. Logs turn products into sums that computers store safely, and their gradients become sums as well, which is what an optimizer needs.
Where is it used?
Every log-likelihood and ELBO (the ELBO is a sum of log densities), mixture models, discrete latent variables summed out by enumeration, softmax and cross-entropy, log-score and log predictive density, and jax.scipy.special.logsumexp.
How is it used?
Call dist.log_prob(y) and sum. To combine probabilities, use logsumexp, never log(sum(exp(...))). Use logits for probabilities, log1p and expm1 for small values. Subtract log(S) when averaging S likelihoods.
"Logs are only there to make the maths tidy."
They are needed. After about 230 factors of 0.04 a float64 product is exactly 0, and after 33 factors in float32. A likelihood computed as a product is wrong long before the data set is large.
"If I get $\log 0$, add a tiny number like $10^{-10}$ and move on."
That hides the bug and distorts the result. First ask why a probability was 0 (a masked row, a wrong support, a probability of exactly 0 or 1 in float32). Prefer stable primitives (log_softmax, logits, log1p, logsumexp) that never create the zero.
"The mean of the log-likelihoods over the posterior draws is the log predictive density."
The log predictive density is the log of the mean likelihood: $\mathrm{LSE}(\log p_s) - \log S$. The mean of the logs is smaller (Jensen's inequality) and punishes draws that fit a point badly too hard.
"Log probabilities are always negative."
A log probability of a discrete outcome is $\le 0$. A log density can be positive when the density is above 1 (narrow distributions).
NumPyro works in log densities everywhere: the ELBO that your SVI loop maximizes is a sum of log_prob terms over the observed days and the latent parameters, so a long history does not underflow. A masked or padded observation should add exactly 0 to that sum (it contributes probability 1), which is why masks are done with jnp.where or a mask handler and not with boolean indexing (Chapter 6.17). If you ever write a custom score, such as the log predictive density of a test day from posterior draws, use logsumexp over the draws and subtract $\log S$. In the A/B framework the same holds for the Beta-Binomial or Student-t likelihood summed over thousands of users: sums of log-probabilities, never products.
"We use log-likelihoods because logs are easier to differentiate."
That is one reason, but the main one in practice is numerical.
Model answer: "Three reasons. Numerical: a product of hundreds of probabilities underflows to 0 in float32 or float64, while the sum of logs is fine. Mathematical: independent pieces multiply, so their logs add, and so do their gradients. Statistical: many families are exponential families, whose log densities are simple. Where I must add probabilities, I use log-sum-exp."
$\log\prod p_i = \sum \log p_i$. $\mathrm{LSE}(a) = m + \log\sum e^{a_k - m}$, $m = \max a$. Softmax $= e^{a_k - \mathrm{LSE}(a)}$.
Log of the average likelihood $=\mathrm{LSE}(\log p_s) - \log S$, not the average of the logs.
Trap: epsilon hacks hide bugs; float32 underflows after about 33 factors of 0.04 (float64 after 232).
Quick check: $a = (-500, -501)$. Compute $\log(e^{a_1} + e^{a_2})$ without a calculator that can handle $e^{-500}$.
Pull out the maximum: $-500 + \log(1 + e^{-1}) = -500 + 0.3133 = -499.687$. The exponentials inside are $e^0 = 1$ and $e^{-1} = 0.368$, which are ordinary numbers.
Numerical stability 2: positive and constrained parameters (exp, softplus, sigmoid)
Some numbers in a model are only allowed in a small region. A noise level $\sigma$ must be positive. The degrees of freedom $\nu$ of a Student-t, the dispersion $\alpha$ of a Negative Binomial and the scale $b$ of a Laplace prior must be positive. A probability must stay between 0 and 1. An optimizer does not know any of that. It takes steps along the whole number line and may happily step to $\sigma = -0.3$, where the likelihood does not exist.
The trick: let the optimizer move a free number $u$ that can be anything, from $-\infty$ to $+\infty$, and compute the real parameter from it with a function that can only produce allowed values. $\sigma = e^{u}$ is always positive. $p = 1/(1 + e^{-u})$ is always between 0 and 1. The optimizer works in an open field; the model lives on the road.
Three ways to say it:
- Picture: a rubber sheet that maps the whole number line onto the positive half-line, so any step lands somewhere legal.
- Numbers: $u = -1, 0, 2$ gives $\sigma = 0.37, 1, 7.39$ with exp, or $0.31, 0.69, 2.13$ with softplus.
- Slogan: optimize on the open field; live on the road.
Two ways to make $\sigma$ positive: exp, $\sigma = e^u$, and softplus, $\sigma = \ln(1 + e^u)$.
- At $u = -1, 0, 2$: exp gives $0.368,\ 1,\ 7.389$. Softplus gives $\ln 1.368 = 0.313,\ \ln 2 = 0.693,\ \ln 8.389 = 2.127$.
- How fast $\sigma$ reacts to $u$ (the slope $d\sigma/du$): for exp it is $\sigma$ itself ($0.368,\ 1,\ 7.39$). For softplus it is the sigmoid $1/(1 + e^{-u})$: $0.269,\ 0.5,\ 0.881$, never above 1.
- One optimizer step of $+1$ from $u = 2$: exp takes $\sigma$ from $7.39$ to $e^{3} = 20.09$ (times 2.72). Softplus takes it from $2.13$ to $\ln(1 + e^3) = 3.05$ (plus 0.92).
- At the extremes: for $u = 20$ exp gives $4.85 \times 10^{8}$ and softplus gives $20.0$. For $u = -20$ both give about $2.06 \times 10^{-9}$, practically zero. So exp moves in percent everywhere, softplus moves in percent only when $\sigma$ is small and in plain units when $\sigma$ is large.
- Transform (bijector). A smooth one-to-one map $g$ from the free line $\mathbb{R}$ onto the allowed set. Positive:
exporsoftplus. In $(0,1)$:sigmoid. In $(-1, 1)$:tanh. A probability vector (a simplex):softmaxor a stick-breaking transform. Ordered values $c_1 \lt c_2 \lt \dots$: first value plus cumulative sums of positive steps. - Change of variables. If $\theta = g(u)$, the density of $u$ is $p_u(u) = p_\theta(g(u))\,|g'(u)|$. The extra factor is the Jacobian (for exp: $|g'(u)| = e^u$, so $\log|g'| = u$). Samplers and SVI work in $u$-space, so this term must be included. NumPyro does it for you.
- In NumPyro: every
numpyro.samplesite has asupport, andbiject_to(support)picks the transform (positive supports useExpTransform, unit-interval supports useSigmoidTransform, aDirichletusesStickBreakingTransform). Guide parameters are constrained too: the scales ofAutoNormaluseconstraints.softplus_positiveby default. For your own parameters:numpyro.param("s", 1.0, constraint=constraints.positive). - Link functions do the same job for a mean: $\mu_t = e^{\eta_t}$ (log link) or $\mu_t = \mathrm{softplus}(\eta_t)$ keep a count's mean positive (Chapter 7.13, Chapter 5.14).
Why do we need it?
Gradient steps are unconstrained. Without a transform, one big step can produce a negative scale or a probability above 1, and the log-likelihood becomes NaN. A transform makes every step legal.
Where is it used?
NumPyro and Stan constrained parameters, TensorFlow Probability bijectors, softmax output layers, the log link in Poisson and Negative Binomial regression, and the standard-deviation parameters of every variational guide.
How is it used?
Declare the right distribution (its support carries the constraint) and let NumPyro transform. For your own parameter use constraint=. Choose exp for multiplicative effects, softplus when large values must not explode, and consider adding a small floor such as softplus(u) + 1e-3.
"Just clip the parameter: sigma = jnp.maximum(sigma, 1e-6)."
Below the clip the gradient is exactly 0, so the optimizer cannot find its way back (a "dead" parameter). Use a smooth transform, then add a floor on top if you need one: jax.nn.softplus(u) + 1e-3.
"sigma = jnp.abs(u) is a fine way to keep a scale positive."
It has a kink at 0, and both $u$ and $-u$ give the same $\sigma$, so the free number is not identified and gradients flip sign. Use exp or softplus, which are one-to-one.
"The most probable $\sigma$ is the same whichever coordinates I use."
The highest point of a density (the MAP) changes under a change of variables, because the density is multiplied by a Jacobian. Medians and other quantiles map across exactly; means do not (Chapter 5.2).
"Exp and softplus are interchangeable."
They agree for very negative $u$ and differ for large $u$. Exp is multiplicative (and can explode); softplus is gentle for large values. Pick one deliberately and record it in the manifest.
Almost every interesting parameter in your two projects lives behind a transform, as long as it is learned and not fixed: the Normal noise $\sigma$, the Student-t degrees of freedom $\nu$ and the Negative Binomial concentration $\alpha$ in the forecasting model; the Laplace scale $b$ and the hierarchical scale $\tau$; the Beta-Binomial rates and the Dirichlet category shares in the A/B framework. Your guide (mean-field, full-rank or low-rank) is a Gaussian in the free space, so for a positive parameter it is a log-normal shape in the original space: skewed to the right, never below 0. If the mean of a count is written as $e^{\eta_t}$, an additive effect inside $\eta_t$ multiplies the demand, which is usually what you want for holidays, and exp may explode when $\eta_t$ gets large during early training; a softplus link or good initial values keep it tame.
"Why does your variational posterior for $\sigma$ look skewed when the guide is Gaussian?"
The Gaussian lives in the unconstrained coordinate $u = \log\sigma$ (NumPyro transforms every site to its free space). Mapping it back through $\exp$ gives a log-normal shape for $\sigma$: positive and right-skewed. The Gaussian approximation is therefore in log-scale, which is often a better fit for a scale than a Gaussian in $\sigma$ itself.
Positive: $\sigma = e^u$ or $\ln(1 + e^u)$. Probability: $1/(1 + e^{-u})$. Simplex: softmax or stick-breaking. Ordered: $c_1 = u_1$, $c_2 = c_1 + e^{u_2}$.
Density in free space: $p_u(u) = p_\theta(g(u))|g'(u)|$ (exp: extra $e^u$). Quantiles map across; the mode and mean do not.
Trap: clipping kills gradients; abs is not one-to-one; exp can explode.
Quick check: an optimizer step changes the free number from $u = 2$ to $u = 3$. By what factor and by how much does $\sigma$ change under exp, and under softplus?
Exp: $\sigma$ goes from $7.39$ to $20.09$, a factor $e = 2.72$ (a rise of 12.7). Softplus: from $2.127$ to $3.049$, a rise of $0.92$ (a factor of 1.43). The same step is a big multiplicative jump with exp and a gentle step with softplus.
Numerical stability 3: initialization, exploding gradients, NaNs and scaling the inputs
Picture walking down a long, narrow valley with very steep walls and a floor that is almost flat. If your steps are big enough to make progress along the floor, you bounce from wall to wall, higher each time, and finally fly out of the valley. If your steps are small enough to be safe, you hardly move along the floor. The valley is that shape when inputs have very different sizes: a column of day numbers 0 to 999 next to a column of 0-or-1 flags.
The cure is to reshape the valley into a round bowl by scaling the inputs so that every direction is about equally steep. The other two cures are to start near the road (sensible initial values) and to cap runaway steps (gradient clipping, smaller learning rate). When all of that fails a NaN ("not a number") appears, and it contaminates everything it touches.
Three ways to say it:
- Picture: a long thin valley (hard) versus a round bowl (easy).
- Numbers: with time in days 0 to 999 the largest safe learning rate is $6.0 \times 10^{-6}$; with time rescaled to 0 to 1 it is $1.58$, about 260,000 times larger.
- Slogan: scale first, then train.
Fit $y \approx a + b\,t$ on $n = 1000$ days by gradient descent, with loss $L = \frac{1}{2n}\sum_i (a + b t_i - y_i)^2$ and the true line $y = 50 + 0.2t$ plus noise (sd 5). The best possible loss is about $\sigma^2/2 = 12.5$.
- The curvature of this loss is the matrix $\begin{bmatrix} 1 & \bar t \\ \bar t & \overline{t^2} \end{bmatrix}$. For $t = 0, \dots, 999$: $\bar t = 499.5$ and $\overline{t^2} = 332{,}833$. Its eigenvalues are $332{,}834$ (steep) and $0.25$ (flat).
- Gradient descent is stable only if the learning rate is below $2/\lambda_{\max} = 2/332{,}834 = 6.0 \times 10^{-6}$.
- With $\mathrm{lr} = 3 \times 10^{-6}$ it is stable, but along the flat direction each step shrinks the error by only $1 - 3\times10^{-6}\times 0.25 \approx 1 - 7.5 \times 10^{-7}$. After 60 steps the loss is still 322. With $\mathrm{lr} = 10^{-5}$ the loss exceeds $10^{15}$ after 15 steps.
- Rescale time to $t' = t/999 \in [0, 1]$. The eigenvalues become $1.268$ and $0.0659$, the safe learning rate is $2/1.268 = 1.58$, and with $\mathrm{lr} = 1$ the loss reaches 12.3 within 60 steps.
- The ratio of largest to smallest curvature (the condition number) fell from $1{,}329{,}345$ to $19.2$. Nothing about the model changed, only the units of the input.
- Initialization: the starting values of the parameters (and of the guide). In NumPyro the default
init_to_uniform(radius=2)draws each free (unconstrained) number uniformly between −2 and 2, so a positive scale starts anywhere between $e^{-2} = 0.14$ and $e^{2} = 7.4$.init_to_medianstarts at the prior medians;init_to_valuesets them by hand.AutoNormalstarts every guide standard deviation atinit_scale=0.1. - Exploding gradients: gradient sizes that grow from step to step (learning rate above the stable limit, badly scaled inputs, an
explink or a scale heading to 0). The loss jumps up instead of down. Remedies: smaller learning rate, scaled inputs, gradient clipping (numpyro.optim.ClippedAdamclips each gradient value into $[-c, c]$, with $c = 10$ by default), gentler links, priors that keep scales away from 0. - NaN: "not a number", produced by $0/0$, $\infty - \infty$, $0 \times \infty$, $\log$ or $\sqrt{}$ of a negative number. Once a parameter is NaN, every later loss and gradient is NaN. Debugging order: (1) are the inputs finite? (2) is the loss finite at initialization? (3) which step first goes non-finite? (4) are the gradients finite? JAX can stop at the first culprit:
jax.config.update("jax_debug_nans", True). In a custom SVI loop,svi.stable_updatekeeps the old state when the new loss or state is not finite (count the skipped steps; it hides the cause, it does not remove it). - Scaling: standardize the target and the regressors (subtract a mean, divide by a scale) and map time to $[0, 1]$. Fit the scaler on the history only, apply the same numbers to the future, and transform forecasts back (Chapter 4.18; leakage in Chapter 7.12).
Why do we need it?
Training with SVI or gradient descent only works when the numbers stay finite and the steps are well sized. One NaN ruins hours of training, and badly scaled inputs force a learning rate that is either unstable or far too slow.
Where is it used?
Every gradient-based fit: SVI with Adam in NumPyro, neural networks (normalization layers, gradient clipping), regression with raw dates, and priors set on standardized data (Prophet-style models scale time and the target for this reason).
How is it used?
Standardize inputs and the target, scale time to 0 to 1, start with init_to_median or sensible values, clip gradients, log the loss and the count of non-finite steps, and when a NaN appears find the first step and the first operation that produced it.
"My loss became NaN, so lower the learning rate and rerun."
Sometimes that works, but first find where the first non-finite value appeared (inputs, initial loss, which step, which operation). A NaN caused by a masked row or a zero probability comes back at any learning rate.
"jnp.where(mask, jnp.log(x), 0) is safe because masked values are replaced by 0."
The forward value is safe; the gradient is not (the discarded branch still sends 0 × ∞). Use the double-where pattern, with the safe value inside the function.
"Scale the whole data set (history and future) once, that is simpler."
That leaks future information into the scaler (Chapter 7.12). Fit the scaler on the history only, save its numbers in the manifest, apply them to new data and invert them on the forecasts. Keep one scaler for all groups: scaling each group separately removes the group differences you want to estimate (Chapter 4.18).
"Skipping bad steps with stable_update fixes the model."
It only keeps the run alive. If many steps are skipped the model is unhealthy: log how many, and fix the cause (scaling, priors, links, learning rate).
Your forecasting model has exactly the ingredients that make training fragile: time in days (large numbers), regressors on their own scales, scale parameters ($\sigma$, $\nu$, $\alpha$, $b$), possibly a log link on a count mean, and a Student-t or Negative Binomial likelihood. Check that time, target and regressors are scaled before the SVI loop starts, that the scaler comes from the history only (and is stored in the manifest), and that scale parameters sit behind exp or softplus. If your loop uses relative-ELBO early stopping, add a guard: a non-finite loss is never "an improvement", so it should neither update the best state nor count toward a patience reset. In the A/B framework, a global scaler fitted on all groups together, masked rows handled with jnp.where (not boolean indexing under JIT) and Beta-Binomial counts passed as counts rather than as long product chains are the same ideas in a different dress.
"Training diverged, so I made the learning rate smaller."
"First I check whether the loss is finite at initialization and when it first turns non-finite. Typical causes are unscaled inputs, an exp link or scale heading to zero, and masked-log gradients. I fix the scaling and the parameterization, clip gradients, and only then tune the learning rate."
Stable learning rate for a quadratic loss: $\mathrm{lr} < 2/\lambda_{\max}$ of the curvature. Scaling inputs shrinks the condition number $\lambda_{\max}/\lambda_{\min}$ (here 1.3 million → 19).
NaN sources: $0/0$, $\infty-\infty$, $0\times\infty$, $\log$ or $\sqrt{}$ of negatives. Double-where for masked logs. Clip, scale, constrain, then debug.
Trap: scaler fitted on all data leaks the future; per-group scaling erases group effects.
Quick check: why can the loss be finite and the gradient NaN for the same input?
The forward pass only evaluates the branch that was chosen. The backward pass multiplies the derivatives of both branches by 0 or 1 and adds them. If the discarded branch has an infinite derivative at that point (for example $1/x$ at $x = 0$), then $0 \times \infty = \mathrm{NaN}$ spoils the gradient even though the forward value looked fine.
Recap, cheat sheet and practice
- Reproducibility: a run manifest saved beside the parameters: data hash and cut-off, preprocessing (scaler numbers), code commit, prior scales, seed, inference settings, guide and optimizer, iterations actually run, library versions (plus precision and forecast origin). A seed alone is one ingredient out of ten.
- Seeds: they make a run repeatable, not "the same to the last digit everywhere". Compare models by the gap relative to the seed-to-seed spread, over 5 to 20 seeds and several forecast origins; never report the best seed.
- Monitoring error and bias: store forecasts as issued, score them when actuals arrive, track rolling MAE or MASE against the launch level, and the rolling mean residual against $\pm 3\sigma_e/\sqrt w$. Alarm after $k$ bad windows in a row; short windows are fast and noisy. An alarm is a question, not a retrain order.
- Monitoring coverage: misses in $w$ days are Binomial$(w, \alpha)$ if the interval is honest; PIT histograms show U (too narrow), hump (too wide), lopsided (bias). Check several levels and both directions.
- Monitoring inputs: missing share, freshness, range, PSI or KS (rule of thumb 0.1 / 0.25; noise level $\approx (k-1)(1/n_1 + 1/n_2)$), changepoint counts per window. Data drift changes $P(x)$; concept drift changes $P(y\mid x)$ and only shows in the errors.
- Log probabilities: sum logs, never multiply probabilities (float64 dies after 232 factors of 0.04, float32 after 33); combine probabilities with log-sum-exp $m + \log\sum e^{a_k - m}$; the log predictive density is $\mathrm{LSE} - \log S$.
- Constrained parameters: optimize free numbers and map them: exp or softplus for positives, sigmoid for $(0,1)$, softmax or stick-breaking for simplexes. The density in free space carries a Jacobian; the MAP moves with the coordinates, quantiles do not.
- Training stability: scale inputs (history-only scaler, time to $[0,1]$), start at sensible values, constrain scales, clip gradients, use the double-where for masked logs, and debug a NaN by finding the first non-finite step and operation.
Cheat sheet
| Idea | Formula or rule | Remember |
|---|---|---|
| Manifest | data hash + preprocessing + code + priors + seed + inference + guide + optimizer + iterations run + versions | written automatically, diff-able |
| Seed noise | gap ÷ spread; $SE_{gap} = s\sqrt{2/m}$ | gap below about 1 to 2 spreads: no claim |
| Rolling MAE alarm | rolling MAE $\gt c \times$ launch MAE, $k$ days in a row | e.g. $c = 1.5$, $k = 2$ |
| Bias limits | $\bar e_w \notin \pm 3\sigma_e/\sqrt w$ | independent residuals; wider if autocorrelated |
| Coverage | misses in $w$ days $\sim$ Bin$(w, \alpha)$: mean $w\alpha$, sd $\sqrt{w\alpha(1-\alpha)}$ | 28 days, 80%: 5.6 ± 2.1; alarm at 12 or more |
| PIT | $u_t = F_t(y_t)$ uniform if calibrated | U narrow · hump wide · tilt bias |
| PSI | $\sum (a_i - e_i)\ln(a_i/e_i)$ | 0.1 / 0.25 rules of thumb |
| Log-likelihood | $\sum_i \log p(y_i \mid \theta)$ | log density can be positive |
| Log-sum-exp | $m + \log\sum e^{a_k - m}$; log mean $= \mathrm{LSE} - \log S$ | never log(sum(exp())) |
| Float limits | float64 $\approx 5\times10^{-324}$ to $1.8\times10^{308}$; float32 $\approx 1.4\times10^{-45}$ to $3.4\times10^{38}$ | exp overflows above 709.8 / 88.7 |
| Transforms | $e^u$ · $\ln(1+e^u)$ · $1/(1+e^{-u})$ · softmax | density gets $|g'(u)|$ (exp: $+u$ in the log) |
| Stable learning rate | $\mathrm{lr} \lt 2/\lambda_{\max}$ of the curvature | scaling shrinks $\lambda_{\max}/\lambda_{\min}$ |
| Masked log | $\log$(where(m, x, 1)), then where(m, ·, 0) | double-where keeps gradients finite |
import hashlib, json, importlib.metadata as md
import numpy as np
import jax, jax.numpy as jnp
import numpyro, numpyro.distributions as dist
from numpyro.infer import SVI, Trace_ELBO
from numpyro.infer.autoguide import AutoNormal
from numpyro.optim import Adam
from numpyro.distributions import constraints
from numpyro.distributions.transforms import biject_to
# 1) A small forecasting problem; the scaler is learned on the history only
rng = np.random.default_rng(0)
n, T = 112, 98 # 98 days of history, 14 to forecast
t = np.arange(n, dtype=float)
week = np.array([-10, -14, -12, -6, 4, 22, 16.0])
y = 100 + 0.1 * t + week[(t % 7).astype(int)] + rng.normal(0, 5, n)
mu_y, sd_y = y[:T].mean(), y[:T].std(ddof=1)
ys = (y - mu_y) / sd_y # standardized target
ts = t / (T - 1) # time mapped to [0, 1] on the history
F = np.column_stack([f(2 * np.pi * k * t / 7) for k in (1, 2) for f in (np.sin, np.cos)])
def model(ts, F, y=None):
a = numpyro.sample("a", dist.Normal(0, 1))
b = numpyro.sample("b", dist.Normal(0, 2))
w = numpyro.sample("w", dist.Normal(0, 1).expand([F.shape[1]]).to_event(1))
sigma = numpyro.sample("sigma", dist.HalfNormal(1)) # positive: NumPyro uses exp behind the scenes
numpyro.sample("y", dist.Normal(a + b * ts + F @ w, sigma), obs=y)
def fit(seed, steps=1500):
guide = AutoNormal(model)
svi = SVI(model, guide, Adam(0.05), Trace_ELBO())
res = svi.run(jax.random.PRNGKey(seed), steps, jnp.array(ts[:T]), jnp.array(F[:T]),
y=jnp.array(ys[:T]), progress_bar=False)
p = res.params
pred = (p["a_auto_loc"] + p["b_auto_loc"] * ts[T:] + F[T:] @ p["w_auto_loc"]) * sd_y + mu_y
return float(np.mean(np.abs(pred - y[T:]))), float(p["b_auto_loc"]), steps
# 2) The run manifest: everything needed to rebuild this result
mae, b_hat, steps = fit(0)
manifest = {
"data_sha256": hashlib.sha256(y[:T].tobytes()).hexdigest()[:12], "rows": T,
"scaler": {"mean": round(mu_y, 3), "sd": round(sd_y, 3)},
"model": "Normal | a + b*t + weekly Fourier(2)", "priors": "a N(0,1), b N(0,2), w N(0,1), sigma HalfNormal(1)",
"seed": 0, "guide": "AutoNormal", "optimizer": "Adam(lr=0.05)", "steps_run": steps,
"libs": {p: md.version(p) for p in ["numpy", "jax", "jaxlib", "numpyro"]},
"x64": bool(jax.config.jax_enable_x64), "holdout_mae": round(mae, 3)}
print(json.dumps(manifest))
# {"data_sha256": "ef5688324529", "rows": 98, "scaler": {"mean": 105.404, "sd": 14.489}, ..., "seed": 0, "guide": "AutoNormal",
# "optimizer": "Adam(lr=0.05)", "steps_run": 1500, "libs": {"numpy": "2.5.3", "jax": "0.11.2", "jaxlib": "0.11.2", "numpyro": "0.22.0"},
# "x64": false, "holdout_mae": 4.583} (versions are those of my machine: recording them is the point)
# 3) Seed-to-seed spread: the yardstick before you compare two models
maes = np.array([fit(s)[0] for s in range(5)])
print("MAE over 5 seeds:", np.round(maes, 3), "mean", round(maes.mean(), 3), "sd", round(maes.std(ddof=1), 3))
# MAE over 5 seeds: [4.583 5.735 4.72 5.313 4.461] mean 4.962 sd 0.542 <- a 0.2 "improvement" would be inside this spread
# 4) Monitoring: rolling MAE and rolling 80% coverage after a level shift on day 120
r = np.random.default_rng(3)
N, D0, L0, w = 200, 120, 60, 7
e = 5 * r.standard_normal(N); e[D0:] += 10 # residuals: the world shifts up by 10 orders at D0
mae0 = np.abs(e[:L0]).mean(); thr = 1.5 * mae0
rmae = np.array([np.nan] * (w - 1) + [np.abs(e[i - w + 1:i + 1]).mean() for i in range(w - 1, N)])
bad = rmae > thr; two = bad[1:] & bad[:-1]; days = np.where(two)[0] + 1 # alarm = two bad days in a row
days = days[days >= L0]
print("launch MAE", round(mae0, 2), "alarm level", round(thr, 2), "| alarm days before the change:", int((days < D0).sum()),
"| first alarm after it: day", int(days[days >= D0][0]))
# launch MAE 3.95 alarm level 5.93 | alarm days before the change: 5 | first alarm after it: day 122 (a 7-day window gives some false alarms)
hit = np.abs(e / 5) <= 1.2816 # inside the 80% interval of a model that thinks sigma = 5
print("coverage before the change", round(hit[:D0].mean(), 2), "after", round(hit[D0:].mean(), 2))
# coverage before the change 0.78 after 0.16 (the 80% promise is broken once the world shifts by 10 = two sigma)
# 5) Log probabilities
p = 0.04
print("0.04**400 =", p ** 400, "| 400*log(0.04) =", round(400 * np.log(p), 2))
# 0.04**400 = 0.0 | 400*log(0.04) = -1287.55
with np.errstate(all="ignore"):
print("float32 0.04**40 =", np.float32(p) ** 40, "| float32 exp(89) =", np.exp(np.float32(89)))
# float32 0.04**40 = 0.0 | float32 exp(89) = inf
a = jnp.array([-1000.0, -1001.0, -1002.0])
print("naive:", jnp.log(jnp.sum(jnp.exp(a))), "| logsumexp:", jax.scipy.special.logsumexp(a),
"| log mean likelihood:", jax.scipy.special.logsumexp(a) - jnp.log(3), "| mean of logs:", a.mean())
# naive: -inf | logsumexp: -999.5924 | log mean likelihood: -1000.69104 | mean of logs: -1001.0
# 6) Transforms for constrained parameters
u = np.array([-1.0, 0.0, 2.0])
print("exp", np.round(np.exp(u), 3), "softplus", np.round(np.log1p(np.exp(u)), 3))
print(type(biject_to(constraints.positive)).__name__, type(biject_to(constraints.unit_interval)).__name__,
type(biject_to(dist.Dirichlet(jnp.ones(3)).support)).__name__)
# exp [0.368 1. 7.389] softplus [0.313 0.693 2.127]
# ExpTransform SigmoidTransform StickBreakingTransform
# 7) The masked-log gradient trap, and scaling
bad = lambda x: jnp.where(x > 0, jnp.log(x), 0.0)
good = lambda x: jnp.where(x > 0, jnp.log(jnp.where(x > 0, x, 1.0)), 0.0)
print("grad at 0: naive where =", jax.grad(bad)(0.0), "| double where =", jax.grad(good)(0.0))
# grad at 0: naive where = nan | double where = 0.0
for name, x in [("days 0..999", np.arange(1000.0)), ("scaled", np.arange(1000.0) / 999)]:
H = np.array([[1, x.mean()], [x.mean(), (x ** 2).mean()]]); ev = np.linalg.eigvalsh(H)
print(name, "eigenvalues", np.round(ev, 4), "condition", round(ev[1] / ev[0]), "max safe lr", float(f"{2 / ev[1]:.3g}"))
# days 0..999 eigenvalues [2.504e-01 3.328e+05] condition 1329345 max safe lr 6.01e-06
# scaled eigenvalues [0.0659 1.2676] condition 19 max safe lr 1.58
1. A teammate says "my forecast is reproducible, I fixed the seed". What is the best reply?
2. Over 5 seeds model A has mean MAE 4.20 and model B 4.05; both have a seed spread of 0.16. What do you conclude?
3. A 7-day mean residual is monitored with $\sigma_e = 4$ and limits at three standard errors. What are the limits?
4. An 80% interval missed on 12 of the last 28 days. What is the best reading?
5. Why is the sum of logs safe when the product of 400 densities of 0.04 is not?
6. jnp.where(x > 0, jnp.log(x), 0.0) returns 0 at $x = 0$ but its gradient is NaN. What is the standard fix?
Practice problems
A. Five seeds each. Model A: 10.2, 10.6, 9.8, 10.4, 10.0. Model B: 9.6, 10.0, 9.2, 9.8, 9.4. Compare them.
- Mean A $= 51.0/5 = 10.2$; mean B $= 48.0/5 = 9.6$.
- Deviations of A: $0, +0.4, -0.4, +0.2, -0.2$; squares sum to $0.4$; divide by 4: $0.1$; $s_A = 0.316$. B has the same deviations, so $s_B = 0.316$.
- Gap $= 0.6$, which is $0.6/0.316 = 1.9$ seed spreads.
- $SE_{gap} = 0.316\sqrt{2/5} = 0.20$, so the gap is $0.6/0.20 = 3.0$ standard errors: B is convincingly better (on this data and these splits; repeat over several forecast origins).
B. Launch MAE 6.0; alarm when the weekly MAE exceeds 1.4 times it for two weeks in a row. Weekly MAEs: 6.5, 7.9, 8.6, 8.0, 9.5, 9.9. When does the alarm fire?
- Alarm level $= 1.4 \times 6.0 = 8.4$.
- Weeks above 8.4: week 3 (8.6), week 5 (9.5), week 6 (9.9). Week 4 (8.0) is below.
- Week 3 is a single breach (week 4 resets it). Weeks 5 and 6 are two in a row, so the alarm fires at the end of week 6.
- Comment: the error was already rising from week 2, and a lower level or a shorter rule would have fired earlier at the price of more false alarms.
C. A 90% interval is checked over 56 days and missed 12 times. How surprising is that?
- Expected misses: $56 \times 0.1 = 5.6$. Standard deviation: $\sqrt{56 \times 0.9 \times 0.1} = \sqrt{5.04} = 2.24$.
- $z = (12 - 5.6)/2.24 = 2.85$.
- Exact Binomial: $P(\text{misses} \ge 12) = 0.0084$, about 1 in 120. Observed coverage is $44/56 = 0.786$ instead of 0.90.
- This is a real signal at the 1% level. Look at the sides of the misses, then at noise, bias and inputs.
D. Training shares of a regressor in four quartile bins are 25% each; this month they are 15%, 20%, 30%, 35%. Compute PSI and interpret it.
- Bin 1: $(0.15 - 0.25)\ln(0.15/0.25) = (-0.10)(-0.511) = 0.0511$.
- Bin 2: $(-0.05)\ln(0.8) = 0.0112$. Bin 3: $0.05\ln 1.2 = 0.0091$. Bin 4: $0.10\ln 1.4 = 0.0336$.
- PSI $= 0.0511 + 0.0112 + 0.0091 + 0.0336 = 0.105$.
- By the rule of thumb that is a moderate shift (just above 0.1). With only four bins and a few hundred recent days the noise level is $(4-1)(1/n_1 + 1/n_2)$, small, so it is a real but mild shift: look at the regressor's source and at the error monitors.
E. Three log-weights are $-800, -801, -803$. Compute the log-sum-exp and the softmax weights.
- Maximum $m = -800$. Shifted values: $0, -1, -3$.
- $e^0 + e^{-1} + e^{-3} = 1 + 0.3679 + 0.0498 = 1.4177$, $\ln 1.4177 = 0.3490$.
- LSE $= -800 + 0.3490 = -799.651$. (A naive calculation gives $\log 0 = -\infty$.)
- Softmax weights: $1/1.4177 = 0.705$, $0.3679/1.4177 = 0.259$, $0.0498/1.4177 = 0.035$; they sum to 1.
F. (Interview) "Your forecasting model goes to production. What do you put around it?"
"Four things. First, a run manifest saved with every fit: data hash and cut-off, the scaler numbers, code commit, prior scales, seed, inference settings, guide type, optimizer, iterations actually run and library versions, so I can rebuild or diff any forecast. Second, a monitor that scores each stored forecast when its actual arrives: rolling MAE or MASE against the launch backtest and the naive baseline, the mean residual for bias, and interval coverage with exact binomial limits and a PIT histogram. Third, input checks: freshness and missing share for every regressor, ranges, PSI against the training period, and the count of changepoints per window. Fourth, numerical safety: log-space likelihoods, scale parameters behind exp or softplus, inputs scaled with a scaler fitted on the history only, gradient clipping, a guard that never accepts a non-finite loss as the best state, and safe masks with the double-where. Alarms need a persistence rule, and an alarm triggers an investigation, not an automatic retrain."
Capstone: one set of ideas, two projects
Your Bayesian A/B framework and your forecasting model look like different machines, but they are built from the same eight ideas: choosing a likelihood, robustness, shrinkage, the bias–variance trade-off, approximate inference, honest uncertainty, distribution diagnostics, and computational scale. This is your final revision hub: a map, eight compact lessons with links into all four guides, an interview drill, and the 44-item core checklist.
- See the eight connecting themes and say, for each one, where it lives in the A/B framework and where it lives in the forecasting model
- Choose a likelihood from the support and the variance structure of the data, for either project
- Explain robustness (Student-t) and shrinkage (hierarchical partial pooling vs Laplace shrinkage of changepoint slopes) as two views of one idea
- Place each modelling knob on the bias–variance curve and say how you would set it
- Tell $P(\theta_B \gt \theta_A \mid D)$ apart from $p(y_{\text{future}} \mid D)$ and name what each is for
- Pick the right diagnostic for each distribution question, and reason about computational scale (JIT, SVI, low-rank guides, model dimension)
- Answer 40 interview questions out loud and tick off all 44 core topics, each with a link back to the chapter that teaches it
How to use this chapter as a revision plan. (1) Open the map and click through the matrix until you can say, for each row, what it does in each project. (2) Take the eight themes one per sitting: read the compact lesson, play with the widget, then say the "Three ways to say it" out loud. (3) Run the interview drill daily: five random questions, answer aloud before you reveal. (4) Keep the 44-item checklist open: every item you cannot explain in one minute is a link to the chapter to re-read. (5) The last section has a one-page cheat sheet, a runnable script that touches all eight themes, six quiz questions and six practice problems.
The map: two projects, one toolbox core
Imagine two houses built from the same box of tools. From the street they look different: one is for deciding "does the new checkout page convert better?", the other for saying "how many orders will we get next month?". Inside, the walls are made of the same bricks. Both choose a noise model for the data, both pull noisy estimates toward a sensible centre, both approximate a posterior with a Gaussian guide trained by SVI, and both must say how sure they are.
That is good news for an interview. If you understand an idea once, you can explain it through whichever project the interviewer asks about, and you can answer a question about one project with an example from the other.
Three ways to say it:
- Picture: two islands joined by eight bridges.
- Numbers: 8 connecting themes, 2 projects, 44 core topics, 72 chapters (18 + 17 + 17 + 20) to find them in.
- Slogan: one toolbox, two projects.
One sentence, told twice: "a noisy estimate is pulled toward a centre, and how far depends on how much evidence there is."
- A/B framework. A segment has only 4 users and an average of 14, while all segments average 10. With a within-segment sd of 10 and a between-segment sd of 2, the segment keeps a weight of $\frac{4/100}{4/100 + 1/4} = 0.138$ on its own data. Its pooled estimate is $10 + 0.138 \times 4 = 10.55$.
- Forecasting model. A candidate changepoint has a raw slope-change estimate of 0.02 with standard error 0.05, and the Laplace scale is $b = 0.1$. The MAP estimate is soft-thresholded by $s^2/b = 0.0025/0.1 = 0.025$, so $0.02 \to 0$. A raw 0.30 becomes $0.30 - 0.025 = 0.275$.
- Same move (shrink toward a centre), different centre (the grand mean, or 0) and different knob (the between-segment sd $\tau$, or the Laplace scale $b$). That is theme 3, worked out in full below.
A concept × project matrix lists topics in rows and the two projects in columns. Each cell says how the topic is used there:
- ● core: part of the model or the workflow as you described it;
- ◐ used: used, usually as a supporting tool or a checking step;
- ○ idea only: not in your code as far as the syllabus says, but the same idea applies (or is the natural next step);
- – not used.
Only facts from the project descriptions are marked core. Where the syllabus says "know the exact version you used" (for example the Negative Binomial parameterization or which guide your code selects), the matrix reminds you to check your own code.
Why do we need it?
Interviewers jump between topics and between your two projects. A single map of "what is used where, and where did I learn it" lets you answer from either project and find the chapter to revise in seconds.
Where is it used?
Interview preparation, project write-ups and design reviews: it is the same table you would draw to explain to a new teammate which statistical choices the experimentation framework and the forecasting model share.
How is it used?
Click a row to see the role in each project, a one-line summary and links into all four guides. Filter by theme to rehearse one idea at a time. For each row, practise saying the A/B sentence and the forecasting sentence out loud.
"The A/B framework is the Bayesian project and the forecasting model is the time-series project."
Both are Bayesian models fitted with SVI in NumPyro and JAX. What differs is the data structure (independent users in groups versus an ordered series) and the question (which arm is better versus what comes next).
"If a topic is marked ○ it does not matter for the interview."
Interviewers love "what would you do if...". The ○ rows are exactly those answers: the Negative Binomial for an overdispersed A/B metric, hierarchical pooling across forecast series, peeking as leakage.
"Memorize the matrix."
Use it to find the chapter, then rebuild the explanation from the reason, the picture and one number. A memorized table breaks at the first follow-up question.
This whole chapter is "in your projects". The A/B framework: Beta-Binomial and Dirichlet-Multinomial for yes/no and category metrics, Normal, Student-t and Poisson for others, hierarchical partial pooling across groups, a global scaler, posterior decisions $P(\theta_A \gt \theta_B \mid D)$, SVI, and the boolean-masking-under-JIT lesson. The forecasting model: $y_t = g(t) + s(t) + h(t) + X_t\beta + \epsilon_t$, a piecewise-linear trend with changepoints from a grid and PELT, Laplace priors on the slope changes, Fourier seasonality, holidays, regressors, Normal, Student-t and Negative Binomial likelihoods, and a custom JIT-compiled SVI loop with relative-ELBO early stopping, patience, best-state checkpointing and a full-rank or low-rank guide chosen by model size. Use only these facts when you describe the projects; for anything else, say "I would have to check the code".
"What do your two projects have in common?" — "Both use NumPyro."
"Four things at least. Both choose a likelihood from the support and variance of the data; both use Student-t where tails are heavy; both use shrinkage toward a centre, partial pooling in the experiments and Laplace priors on the changepoint slopes in the forecast; and both fit a Gaussian guide with SVI, so the same questions about ELBO, guide family and convergence apply. They differ in the output: a decision probability in one, a predictive distribution in the other."
Eight themes: likelihood · robustness · shrinkage · bias–variance · approximate inference · uncertainty · diagnostics · scale.
For every topic know: what it does in the A/B framework, what it does in the forecasting model, which chapter teaches it.
Trap: describing the projects with details you did not state; say "I would check the code" instead.
Quick check: which two rows of the matrix are "core" in both columns?
SVI with the ELBO, and JIT compilation. Everything else is core on one side only (for example Student-t is "used" in the A/B framework and "core" in the forecasting model), which is why you need to translate an idea between the projects.
Theme 1 · Choose the likelihood from the support and the variance core
A likelihood is your model's guess about the shape of surprise: given what the model expects, how likely is each value that could actually show up? Before you pick one, ask two plain questions about a single observation. First: what values can it take? (the support: yes or no; one of several categories; a count 0, 1, 2, …; a positive amount; any real number). Second: how does the scatter behave? (the variance structure: the same everywhere, growing with the average, or mostly calm with a few extreme values).
It is like choosing a container. You decide by what you carry (water, eggs, sand) before you worry about the size. A Normal likelihood for a yes/no metric is a bucket for eggs.
Three ways to say it:
- Picture: a decision tree: "what can the number be?" leads to a family; "how does it scatter?" picks the member.
- Numbers: orders per day have mean 100 and variance 400. Poisson says the variance equals the mean (100), four times too small, so use a Negative Binomial.
- Slogan: support first, variance second, habit never.
Four metrics, four choices.
- 50 of 500 users converted. Each user is yes or no, so the support is $\{0, 1\}$ and the count per arm is Binomial$(n = 500, p)$ with variance $np(1-p) = 500 \times 0.1 \times 0.9 = 45$. With a Beta prior the posterior is exactly Beta: the Beta-Binomial model (Chapter 6.3).
- Orders per day: mean 100, variance 400. Counts, so Poisson or Negative Binomial. Poisson forces variance $=$ mean $= 100$. The data are four times more spread out. The Negative Binomial has variance $\mu + \mu^2/\alpha$: $400 = 100 + 10000/\alpha$, so $\alpha = 10000/300 = 33.3$. Check: $100 + 10000/33.3 = 400$. ✓
- Daily revenue residuals, symmetric, with a few huge days. Any real number, calm centre, heavy tails: Student-t. The Normal would let the huge days drag the mean and the noise level.
- Minutes on a page: positive, long right tail. A Normal can go negative and misses the skew. Take logs (Chapter 4.18) and use a Normal or Student-t on the log scale, or use a log-normal or Gamma.
Support = the set of values an observation can take. Variance function = how the variance depends on the mean. Matching both is the first step of model building.
| One observation is… | Likelihood | Variance | A/B framework | Forecasting model |
|---|---|---|---|---|
| yes / no | Bernoulli$(p)$, counts: Binomial$(n, p)$ | $p(1-p)$, $np(1-p)$ | ● Beta-Binomial | – |
| one of $K$ categories | Categorical, counts: Multinomial | $np_k(1-p_k)$ per category | ● Dirichlet-Multinomial | – |
| count, spread $\approx$ mean | Poisson$(\lambda)$ | $\lambda$ | ◐ | ○ (limit of NB) |
| count, spread above mean | Negative Binomial$(\mu, \alpha)$ | $\mu + \mu^2/\alpha$ | ○ | ● |
| real, light tails | Normal$(\mu, \sigma)$ | $\sigma^2$ | ◐ | ● |
| real, heavy tails | Student-t$(\nu, \mu, \sigma)$ | $\sigma^2\nu/(\nu - 2)$, $\nu \gt 2$ | ◐ | ● |
| positive, right-skewed | log-normal, Gamma, or Normal on the log scale | grows like the mean squared | transform (4.18) | transform or log link |
| proportion in (0, 1) | Beta, or Normal on the logit scale | $\mu(1-\mu)/(1+\phi)$ | ○ | ○ |
● core in your description · ◐ used · ○ idea only. Libraries: NumPyro NegativeBinomial2(mean, concentration) has variance $\mu + \mu^2/\alpha$, but NegativeBinomialProbs(total_count, probs), SciPy nbinom(n, p) and "failures before $r$ successes" use different parameters. Know the one your code calls.
Why do we need it?
A likelihood that cannot produce the data (negative counts, probabilities above 1) or that gets the spread wrong gives wrong intervals and wrong decisions, even if the average is right. Variance decides how fast you learn and how wide intervals are.
Where is it used?
Every NumPyro model: dist.Binomial, dist.Multinomial, dist.Poisson, dist.Normal, dist.StudentT, dist.NegativeBinomial2; GLMs (logistic, Poisson, NB regression); and the A/B and forecasting likelihood lists in Part 30 of your syllabus.
How is it used?
Write down the support, compute mean, variance, skewness and the share of zeros, pick the family, then check it with a Q-Q plot, a variance-vs-mean plot or a posterior predictive check. If a check fails, move to the next family on the table.
"Everything is roughly Normal, so use a Normal likelihood."
A Normal can produce negative counts, probabilities outside (0, 1), and gets the variance of conversions wrong (it depends on $p$). With big samples the average is approximately Normal (CLT), but the individual observations are not.
"Poisson is the default for counts."
Poisson forces variance $=$ mean. Real counts (orders, sessions) usually have extra spread, and a Poisson model then gives intervals that are too narrow. Check variance ÷ mean first.
"Excess zeros and overdispersion are the same problem."
A Negative Binomial with the same mean has more zeros than a Poisson, but sometimes there are even more zeros than that: a separate "no activity" process (zero-inflated or hurdle model).
"The variance ÷ mean ratio decides alone."
It is a rule of thumb from one sample (the wizard uses 1.3 as a cut). Confirm with a posterior predictive check on the fitted model.
Part 30 of your syllabus lists the choices exactly: experiments use Normal, Student-t, Poisson, Binomial and Multinomial; forecasting uses Normal, Student-t and Negative Binomial. The A/B framework's Beta-Binomial and Dirichlet-Multinomial are the Binomial and Multinomial with their conjugate priors. In the forecasting model the count mean has to stay positive (a log link is the usual way; check how your code does it), and the same support-and-variance logic explains why the Negative Binomial replaces the Poisson. When you describe the NB in an interview, say which parameterization your code uses (mean–concentration NB2, or total count and probability).
"We used a Normal likelihood because it is standard."
"I start from the support and the variance. For the conversion metric the observation is yes/no, so Binomial with a Beta prior. For orders per day, counts with variance well above the mean, so a Negative Binomial. For noisy continuous metrics with spikes, a Student-t. Then I check the choice with a Q-Q plot or a posterior predictive check."
Support first (yes/no, categories, counts, positive, real), variance second (constant, grows with mean, heavy tails).
Binomial $np(1-p)$ · Poisson $\lambda$ · NB $\mu + \mu^2/\alpha$ · Normal $\sigma^2$ · Student-t $\sigma^2\nu/(\nu-2)$.
Trap: Normal for bounded or count data; Poisson when the variance is above the mean; mixing NB parameterizations.
Quick check: a metric is a count with mean 4 and variance 4.2 in 200 users. Poisson or Negative Binomial?
The ratio variance ÷ mean is 1.05, close to 1, so Poisson is reasonable (the Negative Binomial estimate $\hat\alpha = 16/0.2 = 80$ is huge, which means "almost Poisson"). The estimate is noisy with 200 users: confirm with a posterior predictive check of the variance.
Theme 2 · Robustness: Student-t in both systems core
Fifteen students score close to each other, and a sixteenth score was typed in wrongly as 100. A fit that assumes Normal noise thinks "every point is equally trustworthy and big surprises basically never happen", so it runs toward the typo and the average jumps. A Student-t fit has "once in a while a point lands far away" built in. When a point is very far, it simply listens less. It does not delete the point; it limits how hard the point can pull.
Both of your projects have this problem: a few huge spenders in an experiment metric, a few spike days in a demand series. The cure is the same likelihood.
Three ways to say it:
- Picture: the Normal's thin tails make a far point look almost impossible, so the fit lunges at it; the t's thick tails make it unremarkable.
- Numbers: a residual 10 scale-units out pulls a Normal fit with force 10 but a Student-t fit ($\nu = 4$) with force 0.48, about 21 times less.
- Slogan: heavy tails, less panic.
Compare a Normal with a Student-t with $\nu = 4$ degrees of freedom (both with scale 1).
- How surprising is a point 4 scale-units away? $P(|X| \gt 4)$ is $0.000063$ for the Normal and $0.0161$ for the $t_4$: the $t_4$ finds it 255 times more plausible.
- The pull a point exerts (the derivative of the negative log-density, called the influence): Normal $\psi(r) = r$. Student-t $\psi(r) = \frac{(\nu+1) r}{\nu + r^2} = \frac{5r}{4 + r^2}$.
- At $r = 1, 2, 5, 10$ the $t_4$ pull is $1.00,\ 1.25,\ 0.86,\ 0.48$. The Normal pull is $1, 2, 5, 10$.
- The pull of the Student-t rises to a maximum (here 1.25 at $r = \sqrt\nu = 2$) and then falls: the further out a point is, the less it counts.
- The same story in cost (negative log-density, up to constants): at $r = 10$ the Normal pays $r^2/2 = 50$, the $t_4$ pays $\frac{\nu+1}{2}\ln(1 + r^2/\nu) = 8.1$.
- Student-t with $\nu$ degrees of freedom, location $\mu$ and scale $\sigma$: heavy tails that thin out like a power. As $\nu \to \infty$ it becomes Normal. The variance is $\sigma^2\nu/(\nu - 2)$ for $\nu \gt 2$, infinite for $\nu \le 2$; $\nu \le 1$ has no mean. The scale is not the standard deviation. NumPyro:
StudentT(df, loc, scale). - Influence function $\psi(r) = -\frac{d}{dr}\log p(r)$: how hard a residual $r$ (in scale units) pulls the fit. Normal: $r$ (unbounded). Laplace: $\mathrm{sign}(r)$ (bounded; it gives the median). Student-t: $\frac{(\nu+1)r}{\nu + r^2}$ (bounded, and falling for $|r| \gt \sqrt\nu$).
- Robust here means: a few bad points cannot change the answer by much. Other robust tools: median, MAD, trimmed means (Chapter 4.14).
Why do we need it?
Real data have occasional extreme values. A Normal likelihood lets one of them move the trend, the seasonal shape or an A/B mean, and it inflates the noise level, which widens every interval.
Where is it used?
Robust regression, Student-t likelihoods in NumPyro, Bayesian A/B tests for revenue-like metrics, demand forecasting with spiky days, and heavy-tailed returns in finance.
How is it used?
Replace dist.Normal(mu, sigma) by dist.StudentT(nu, mu, sigma), give $\nu$ a prior or a fixed small value, and check that the posterior of $\nu$ is not stuck at its prior. Compare Q-Q plots of residuals under both likelihoods.
"A Student-t likelihood removes outliers."
It removes nothing. It assigns more probability to extreme residuals, so they exert less influence on the fit. The outlier stays in the data and in the residual plot.
"Use a small $\nu$ to be safe."
A very small $\nu$ (1 or 2) has infinite variance and tells the model that almost anything can happen, so intervals become wide and the other parameters get weakly determined. Let the data speak (a prior on $\nu$) and check that the result is not just the prior.
"The Student-t scale is the standard deviation."
The standard deviation is $\text{scale}\times\sqrt{\nu/(\nu - 2)}$ for $\nu \gt 2$. For $\nu = 4$ it is $1.41$ times the scale.
"A spike on a holiday is an outlier, so the Student-t will handle it."
If you know why the day is special, put the cause in the model (a holiday column). Use heavy tails for the surprises you cannot explain.
Both of your systems offer the Student-t. In the forecasting model it protects the trend, seasonal and holiday weights from spike days, and it keeps the noise level from being inflated by a few bad days, which is what makes forecast intervals honest on ordinary days. In the A/B framework it plays the same role for continuous metrics with a few extreme users, so the estimate of an arm's mean does not hinge on one whale. In both, check whether your code estimates $\nu$ or fixes it, and look at how much the posterior of $\nu$ moves from its prior.
"We used Student-t to remove outliers."
"We used Student-t because it assigns more probability to extreme residuals, so a few extreme days or users have bounded influence on the fit. The outliers are still in the data and show up in the residual plot. I looked at a Q-Q plot of Normal-likelihood residuals first: the ends bent away from the line, which is the heavy-tail signature."
Influence: Normal $\psi = r$ (unbounded); Laplace $\mathrm{sign}(r)$; Student-t $\frac{(\nu+1)r}{\nu + r^2}$ (bounded, falls after $\sqrt\nu$).
Student-t: $sd = \text{scale}\sqrt{\nu/(\nu-2)}$ for $\nu \gt 2$; $\nu \to \infty$ is Normal.
Trap: it down-weights, it does not delete; explain known spikes with regressors.
Quick check: for $\nu = 4$, at which residual is the Student-t pull largest, and how large is it?
At $|r| = \sqrt\nu = 2$ the pull is $(\nu + 1)/(2\sqrt\nu) = 5/4 = 1.25$. Beyond that it falls: 0.86 at $r = 5$ and 0.48 at $r = 10$.
Theme 3 · Shrinkage: partial pooling vs Laplace shrinkage core
You see two restaurants. One has 3 reviews, all five stars. The other has 3,000 reviews averaging 4.6. You do not believe the first "5.0" as much as the second "4.6": you mentally pull the first toward a typical rating, and the fewer the reviews, the harder you pull. That is partial pooling: a small segment's average is pulled toward the overall average, and a big segment is left mostly alone.
The same instinct fixes another problem. In the forecasting model, each candidate changepoint has an estimated slope change. Most candidates are probably "nothing happened here". A tiny estimated change is mostly noise, so pull it all the way to zero. A big change is believed. That is Laplace shrinkage of the changepoint slopes.
Three ways to say it:
- Picture: arrows from raw estimates toward a centre line: short arrows for well-measured estimates, long arrows for poorly measured ones.
- Numbers: a segment of 4 users keeps 14% of its own average (the rest comes from the overall mean); a segment of 400 keeps 94%. A slope change of 0.02 becomes 0; one of 0.30 becomes 0.275.
- Slogan: trust an estimate in proportion to the evidence behind it.
A/B framework. Overall mean $\mu = 10$. Within a segment, users vary with sd $\sigma = 10$. Segment means vary with sd $\tau = 2$. The weight on a segment's own average is $w = \dfrac{n/\sigma^2}{n/\sigma^2 + 1/\tau^2}$.
- $1/\tau^2 = 1/4 = 0.25$. For $n = 4$: $n/\sigma^2 = 4/100 = 0.04$, so $w = 0.04/(0.04 + 0.25) = 0.138$. If its average is 14, the pooled estimate is $0.138 \times 14 + 0.862 \times 10 = 10.55$.
- For $n = 100$: $n/\sigma^2 = 1$, $w = 1/1.25 = 0.8$, and an average of 14 becomes $0.8 \times 14 + 0.2 \times 10 = 13.2$.
- For $n = 400$: $w = 4/4.25 = 0.94$. Big segments keep their own averages.
Forecasting model. A candidate slope change has a raw estimate $z$ with standard error $s = 0.05$, and the prior is $\delta \sim \text{Laplace}(0, b)$ with $b = 0.1$. The most probable value (MAP) is $\hat\delta = \mathrm{sign}(z)\max(|z| - s^2/b,\ 0)$, where $s^2/b = 0.0025/0.1 = 0.025$.
- $z = 0.02$: $|z| - 0.025 \lt 0$, so $\hat\delta = 0$.
- $z = 0.04$: $0.04 - 0.025 = 0.015$.
- $z = 0.30$: $0.30 - 0.025 = 0.275$; $z = -0.25$: $-0.225$. Big changes lose only 0.025.
- Partial pooling (Normal-Normal): $\theta_g \sim N(\mu, \tau^2)$, group averages $\bar y_g \sim N(\theta_g, \sigma^2/n_g)$. The posterior mean is $\hat\theta_g = w_g\bar y_g + (1 - w_g)\mu$ with $w_g = \dfrac{n_g/\sigma^2}{n_g/\sigma^2 + 1/\tau^2}$. As $\tau \to 0$: complete pooling ($w \to 0$, everyone equals $\mu$). As $\tau \to \infty$: no pooling ($w \to 1$). In practice $\mu$ and $\tau$ are learned (a hierarchical model, Chapter 6.6).
- Laplace shrinkage: minimizing $\frac{(z - \delta)^2}{2s^2} + \frac{|\delta|}{b}$ gives the soft threshold $\hat\delta = \mathrm{sign}(z)\max(|z| - s^2/b, 0)$ (Chapter 5.3, Chapter 7.10). This is the MAP. The full posterior is a continuous distribution: shrunk, with $P(\delta = 0 \mid D) = 0$.
- Shared idea: both are priors that pull noisy estimates toward a centre (the overall mean, or 0), by an amount that depends on the evidence relative to the prior spread ($\tau$ or $b$). Both give up a little bias for a lot less variance (Chapter 5.1).
| Hierarchical partial pooling | Laplace on changepoint slopes | |
|---|---|---|
| Pulled toward | the overall mean $\mu$ (learned) | zero |
| Strength set by | $\tau$ and each group's $n_g/\sigma^2$ | scale $b$ and each estimate's standard error $s_j$ |
| Shape of the pull | proportional: every group moves a fraction $1 - w_g$ | thresholding: small estimates go to 0, large ones lose a fixed amount |
| Exact "no effect"? | no (except $\tau = 0$) | only in the MAP, not in the posterior |
Why do we need it?
Small groups and weak signals give noisy estimates that look extreme by chance. Believing them at face value gives bad decisions (ship the "best" tiny segment) and wiggly forecasts (a trend with 25 invented changes). Shrinkage lowers the average error.
Where is it used?
Hierarchical models for segments, schools, hospitals and sports averages; ridge and lasso regression; Prophet-style changepoint priors; empirical-Bayes methods such as the James–Stein estimator; sparse holiday effects with shrinkage priors.
How is it used?
Add a prior that centres the effects (a hierarchical Normal for groups, a Laplace for slope changes). Let the data set the amount where possible (a hyperprior on $\tau$), check sensitivity to the scale ($b$), and compare estimates before and after shrinking.
"A Laplace prior makes the changepoint slopes exactly zero, so the model is sparse."
Only the MAP is exactly sparse. The full posterior under a Laplace prior is a continuous distribution: its mean and median are shrunk toward 0 but are almost never exactly 0, and $P(\delta_j = 0 \mid D) = 0$. Say "sparse-ish" or "shrunk", and say whether you mean MAP or posterior.
"Partial pooling averages all groups together."
It is a weighted average of the group's own mean and the overall mean, with weights from the evidence: $w_g = (n_g/\sigma^2)/(n_g/\sigma^2 + 1/\tau^2)$. A big group is barely changed.
"Shrinkage helps every group."
It lowers the error on average. A group with a truly large effect is pulled too far (bias). That is the price, and the reason the amount should be learned or checked for sensitivity.
"To remove group differences, standardize each group separately."
Per-group standardization erases the group effect you wanted to estimate. Use one global scaler and let the hierarchical model share information (Chapter 4.18).
The syllabus's own words: "learn the idea once and both projects make sense". In the A/B framework, hierarchical partial pooling pulls segments with little data toward the overall effect: it is the answer to "what if a segment has 40 users?". In the forecasting model, the Laplace prior $\delta_j \sim \text{Laplace}(0, b)$ pulls the slope changes at the candidate changepoints (from your grid and from PELT) toward zero, so only changes the data really support survive. The scale $b$ is the knob: too small and the trend is too stiff, too large and it follows noise. Check whether your code fixes $b$ or learns a hyperprior, and run a prior-sensitivity check.
"Hierarchical models and regularization are different topics."
"They are the same idea. Partial pooling is a Normal prior centred on the overall mean, with a width $\tau$ learned from the data; ridge regression is a Normal prior centred on zero. The Laplace prior is the L1 version, with a kink at zero. In my A/B framework the pooling stops small segments from looking extreme, and in the forecasting model the Laplace prior keeps the candidate changepoints from fitting noise. In both, the MAP and the full posterior are different objects."
Pooling: $\hat\theta_g = w_g\bar y_g + (1-w_g)\mu$, $w_g = \frac{n_g/\sigma^2}{n_g/\sigma^2 + 1/\tau^2}$. Laplace MAP: $\mathrm{sign}(z)\max(|z| - s^2/b, 0)$.
Both pull toward a centre (overall mean, zero); the strength comes from the prior scale ($\tau$, $b$) against the evidence.
Trap: Laplace posterior is not exactly sparse; per-group scaling kills group effects.
Quick check: with $\sigma = 10$ and $\tau = 2$, a segment has 25 users. What weight does it put on its own average?
$n/\sigma^2 = 25/100 = 0.25$ and $1/\tau^2 = 0.25$, so $w = 0.25/0.5 = 0.5$: half its own average, half the overall mean.
Theme 4 · The bias–variance trade-off: one curve, many knobs core
A tailor can make a suit too loose or too tight. A suit that is too loose fits nobody well: every customer is wrong in the same way (that is bias, the model is too simple to follow the real shape). A suit stitched exactly around one customer's posture on one day fits that person that day and nobody else, and it changes if you re-measure (that is variance, the model follows the noise of this particular sample).
Almost every setting in your two projects slides along this one scale: how many Fourier harmonics, how many changepoints, how large the Laplace scale, how much pooling, how many regressors, how rich the guide. More flexibility lowers bias and raises variance. The best setting is in the middle, and you find it on held-out data, never on the training fit.
Three ways to say it:
- Picture: a U-shaped curve of holdout error, with the training error sliding down the whole way.
- Numbers: in the widget below, training RMSE falls from 6.7 (order 1) to 3.7 (order 14) while holdout RMSE is lowest, 5.1, at order 3 and climbs to 6.0 by order 14.
- Slogan: flexibility buys lower bias and costs higher variance; pick the middle on held-out data.
The cleanest version: estimate a true effect $\mu = 2$ from a noisy average $\bar x$ whose variance is $\sigma^2/n = 1$. Compare $\bar x$ with the shrunk estimate $c\,\bar x$.
- Bias of $c\bar x$: $(c - 1)\mu$, so bias$^2 = (1-c)^2\mu^2 = 4(1-c)^2$. Variance: $c^2\sigma^2/n = c^2$.
- $\mathrm{MSE}(c) = \text{bias}^2 + \text{variance} = 4(1-c)^2 + c^2$.
- $c = 1$ (no shrinkage, unbiased): $0 + 1 = 1.00$. $c = 0.8$: $4(0.04) + 0.64 = 0.80$. $c = 0.5$: $4(0.25) + 0.25 = 1.25$. $c = 0$: $4$.
- The best shrinkage is $c^* = \mu^2/(\mu^2 + \sigma^2/n) = 4/5 = 0.8$, with MSE $0.80$: a 20% error reduction from a small, deliberate bias. Pooling, ridge and Laplace shrinkage are this idea in different clothes.
For an estimate $\hat\theta$: $\mathrm{MSE} = \mathrm{Bias}^2 + \mathrm{Var}$. For a prediction at a new point: expected squared error $= \mathrm{Bias}^2 + \mathrm{Var} + \sigma^2_{\text{noise}}$; the last term cannot be reduced by any model (Chapter 5.1). A knob is any setting that changes how flexible the model is:
| Knob | Project | Turn it toward more flexible | Risk of too much | Risk of too little |
|---|---|---|---|---|
| Fourier order $N$ | forecast | higher $N$ | wiggly season that fits noise | misses sharp peaks |
| Changepoints and Laplace scale $b$ | forecast | more candidates, larger $b$ | trend chases noise | trend too stiff, late to turn |
| Regressors, holiday windows | forecast | more columns | overfit, collinearity | unexplained spikes |
| Pooling strength ($\tau$) | A/B | larger $\tau$ (less pooling) | small segments look extreme | real segment differences hidden |
| Prior strength | both | weaker prior | noisy posterior | prior overrules data |
| Guide rank / family | forecast (and A/B SVI) | full-rank, larger $r$ | cost and noisy fitting | correlations missed, intervals too narrow |
Choose by rolling-origin validation (Chapter 7.15) with point and probabilistic scores (Chapter 7.16). A common tie-breaker (a rule of thumb, the "one-standard-error rule") is to take the simplest setting whose holdout error is within one standard error of the best.
Why do we need it?
A model that fits training data perfectly can forecast badly, and a model that is too simple misses real structure. Naming the trade-off lets you set every flexibility knob with one question: what happens on data the model has not seen?
Where is it used?
Model selection and cross-validation everywhere: ridge and lasso penalties, tree depth, polynomial degree, Fourier order, the changepoint prior scale in Prophet-style models, hierarchical pooling, and early stopping in training.
How is it used?
Pick one knob, fit a range of settings, score each on held-out time periods (not on the training fit), look for the U shape, and choose the middle. Report the setting, the score and its spread over several forecast origins.
"Pick the setting with the lowest training error."
Training error always favours the most flexible model. Judge on held-out periods: rolling-origin evaluation for time series.
"An unbiased estimator is always the best one."
The best estimator minimizes total error, $\mathrm{Bias}^2 + \mathrm{Var}$. A small, deliberate bias (shrinkage) often removes much more variance than it adds.
"More data makes the trade-off disappear."
More data lowers the variance term, so the best setting moves toward more flexibility. The trade-off remains.
"The best setting is the one that gets the lowest error once."
Errors from one split are noisy. Compare settings over several forecast origins and seeds, and prefer the simpler setting when the scores are within the noise (Chapter 7.19).
Your forecasting model has many knobs on this one curve: the Fourier order for weekly and yearly seasonality, the number and placement of changepoints and the Laplace scale $b$, how many regressors and holiday windows, and the guide's rank. Tune them with rolling-origin validation, not training fit, and score with both a point metric and a probabilistic one such as CRPS (a model can improve MAE while its intervals get worse). In the A/B framework, the same trade-off is the pooling strength $\tau$ and the prior strength: pool too little and small segments swing wildly, pool too much and real segment differences vanish. Hierarchical models learn the amount from the data, which is why they are attractive.
"How did you choose the Fourier order and the number of changepoints?" — "We tried a few and kept the one with the best fit."
"Both are flexibility knobs on the bias–variance curve. I fitted a range of values and scored each on rolling-origin holdouts, looking at MAE and at a probabilistic score like CRPS. Training fit always improves with more harmonics or changepoints, so it cannot choose. I picked the simplest setting whose holdout score was within noise of the best, and checked it was stable over several forecast origins."
$\mathrm{MSE} = \mathrm{Bias}^2 + \mathrm{Var}$. For $c\bar x$: $\mathrm{MSE}(c) = (1-c)^2\mu^2 + c^2\sigma^2/n$, best $c^* = \mu^2/(\mu^2 + \sigma^2/n)$.
Knobs: Fourier order, changepoints and $b$, regressors, pooling $\tau$, prior strength, guide rank. Choose on held-out data over several origins.
Trap: training error rewards complexity; unbiased is not best.
Quick check: $\mu = 3$ and $\sigma^2/n = 1$. What shrinkage factor $c^*$ minimizes the MSE of $c\bar x$, and what is that MSE?
$c^* = 9/(9 + 1) = 0.9$. MSE $= (1 - 0.9)^2 \times 9 + 0.9^2 \times 1 = 0.09 + 0.81 = 0.90$, against $1.00$ without shrinkage.
Theme 5 · Approximate inference: SVI in NumPyro and JAX, in both projects core
The exact posterior is often out of reach: to normalize it you must add up the model's fit over every possible parameter value, and with dozens or hundreds of parameters that integral cannot be done. So we change the question. Instead of computing the posterior, we choose a simple shape that we can handle, a Gaussian bell, and tune its knobs (centre and spread) until it looks as much as possible like the posterior. It is like pressing a flexible mould onto a sculpture: you never get the sculpture exactly, but you get a shape you can work with.
That is variational inference. Stochastic variational inference (SVI) does the tuning by noisy gradient steps, which is why it is fast and why your forecasting loop has an early-stopping rule. Both of your projects run on this machinery.
Three ways to say it:
- Picture: a bell curve that slides and squeezes until it sits on top of a lopsided posterior.
- Numbers: for 50 conversions in 500 users the exact posterior has mean 0.1016 and sd 0.01347; a Gaussian guide (on the logit scale) fitted by maximizing the ELBO gives mean 0.1016 and sd 0.01352.
- Slogan: SVI turns integration into optimization; the answer is only as good as the guide.
One arm, 50 conversions out of 500, flat Beta(1, 1) prior. Here the exact answer is known, so we can grade the approximation.
- Exact: the posterior is Beta$(1 + 50, 1 + 450) = $ Beta$(51, 451)$: mean $51/502 = 0.1016$, sd $0.01347$, 95% interval $[0.0767, 0.1295]$, and $P(\theta \gt 0.12) = 0.0906$.
- Guide: $q$ says "the logit of $\theta$ is Normal$(m, s^2)$". We choose $m$ and $s$ to maximize the ELBO, which is the same as minimizing $KL(q \,\|\, \text{posterior})$ because $\log p(D) = \mathrm{ELBO} + KL(q \,\|\, p(\cdot \mid D))$.
- Result: $m = -2.188$, $s = 0.148$. Mean $0.1016$, sd $0.01352$, 95% interval $[0.0774, 0.1303]$, $P(\theta \gt 0.12) = 0.0926$.
- Grade: the mean is exactly right (for this model the best logit-Normal guide always matches the posterior mean). The spread and tail probability are close (0.0926 vs 0.0906). With only 5 users and none converting, the approximation is rougher: $P(\theta \gt 0.12)$ is 0.424 vs the exact 0.464.
- Approximate inference menu (Chapter 6.9): conjugate (exact, simple models), MCMC / NUTS (samples that target the posterior as the run length grows; costly; needs convergence checks), Laplace (a Gaussian at the mode), variational / SVI (optimize a guide).
- SVI maximizes the ELBO $= E_q[\log p(D, \theta)] - E_q[\log q(\theta)]$ over the guide's parameters $\phi$, using noisy gradient estimates from random draws (the reparameterization trick) and an optimizer such as Adam (Chapter 6.12).
- Guide families (Chapter 6.13): mean-field (independent, $2d$ numbers), full-rank (all correlations, $d + d(d+1)/2$), low-rank (main correlations, $d(r+2)$).
- Limits: the best guide in its family can still be wrong (too narrow, one mode). NUTS is asymptotically exact, but a finite run still has Monte Carlo error and convergence risks, so never say "NUTS gives the exact posterior" (Chapter 6.15).
Why do we need it?
Realistic models have too many parameters for exact integration, and sampling them with MCMC can be too slow to iterate on or to rerun every week. SVI trades some exactness for speed and scale.
Where is it used?
NumPyro and Pyro (SVI, AutoNormal, AutoLowRankMultivariateNormal), variational autoencoders, large topic models, and production Bayesian models that must refit often.
How is it used?
Write the model, pick a guide, run the optimizer on the ELBO with a stopping rule, then draw posterior samples from the guide. Validate: compare with an exact or NUTS answer on a small case, run several seeds, and check posterior predictive fits.
"NUTS gives the exact posterior."
NUTS produces samples that target the posterior as the run gets longer. A finite run still has Monte Carlo error, can fail to explore (divergences, poor R̂), and needs diagnostics. "Asymptotically exact" and "exact" are different claims.
"A converged ELBO means the posterior is right."
It means the optimizer found the best guide of that family. A mean-field guide on a correlated posterior converges to a too-narrow answer. Convergence of the optimization and quality of the approximation are two separate questions.
"SVI replaces conjugate formulas."
If the posterior is available in closed form (an arm with a Beta prior and a Binomial likelihood), use it, and use it to test SVI. Reach for approximation when the model has no closed form (hierarchies, Student-t, NB regression with many terms).
"A larger ELBO always means a better model."
ELBOs are comparable only for the same data and model, as lower bounds. A looser guide gives a lower ELBO for the same model, which says nothing about the model itself.
Both projects fit with SVI in NumPyro and JAX. In the forecasting model the loop is yours: JIT-compiled updates, relative-ELBO early stopping with patience, best-state checkpointing (a noisy ELBO makes "best so far" a moving target, so keep the best state, not the last), and an automatic choice between a full-rank and a low-rank Gaussian guide based on model size. In the A/B framework, an arm with a Beta prior needs no inference at all; the hierarchical models over segments do, and you can validate the SVI result on a single arm against the exact Beta answer as in the widget. For decisions like $P(\theta_B \gt \theta_A \mid D)$, remember that tails are what a too-simple guide gets wrong: check against a longer run, another seed, or NUTS on a subset.
"We use SVI because it is faster than MCMC."
"SVI turns inference into optimization: I choose a family of distributions, a guide, and maximize the ELBO, which minimizes the KL divergence from the guide to the posterior. It scales and refits quickly, at the price of guide-dependent error, typically too-narrow intervals if the guide ignores correlations. I validate it on a case with a known answer, compare seeds, and would check against NUTS on a subset when the decision is high-stakes. I do not claim NUTS is exact either: it needs convergence diagnostics."
$\log p(D) = \mathrm{ELBO} + KL(q\,\|\,p(\cdot\mid D))$; maximize the ELBO = minimize that KL. ELBO $= E_q[\log p(D,\theta)] - E_q[\log q]$.
Guide sizes: mean-field $2d$, low-rank $d(r+2)$, full-rank $d + d(d+1)/2$.
Trap: "converged" is about the optimizer, not the approximation; NUTS is asymptotically exact, not exact.
Quick check: why can the ELBO of a mean-field guide be lower than the ELBO of a full-rank guide for the same model?
The ELBO equals $\log p(D) - KL(q \,\|\, \text{posterior})$. The model (and so $\log p(D)$) is the same, but a mean-field guide cannot represent the posterior's correlations, so its best KL is larger and its ELBO lower. The gap is a measure of how much the simpler guide loses, not a property of the model.
Theme 6 · Uncertainty: $P(A \gt B \mid D)$ versus $p(y_{\text{future}} \mid D)$ core
Two different questions hide behind the word "uncertainty". The first: "Is B really better than A?" It is about the settings of the world (the true conversion rates, the true average gain), which we cannot see. The second: "What will the next customer, or the next day, do?" It is about a single future observation.
More data answers the first one more and more firmly, because our uncertainty about the settings shrinks. The second one never gets below a floor: even if we knew every setting exactly, people and days stay random. Your A/B framework is built to answer the first question; your forecasting model is built to answer the second.
Three ways to say it:
- Picture: as data grows, one bell curve (the posterior of the gain) narrows to a spike, while the two bell curves for individual outcomes stay wide and on top of each other.
- Numbers: with 10,000 users per arm and a true gain of 0.05 standard deviations, $P(\theta_B \gt \theta_A \mid D) = 0.9998$, yet a random B user beats a random A user only 51.4% of the time.
- Slogan: the posterior shrinks with data; the predictive has a noise floor.
A metric with known user-level standard deviation $\sigma = 1$. Arm B has a true average $0.05$ higher than arm A. Both arms have $n = 10{,}000$ users, a flat prior, and data that look exactly typical.
- Question 1: is B better? The posterior of the difference of means is $N(0.05,\ 2\sigma^2/n)$: standard deviation $\sqrt{2/10000} = 0.01414$.
- $P(\Delta \gt 0 \mid D) = \Phi(0.05/0.01414) = \Phi(3.54) = 0.9998$. We are 99.98% sure B has the higher average.
- Question 2: will a random B user beat a random A user? A new user's outcome has the whole noise on top of the uncertain mean: variance $\sigma^2(1 + 1/n)$. The difference of two new users has sd $\sqrt{2(1 + 1/10000)} = 1.4143$.
- $P(Y_B \gt Y_A) = \Phi(0.05/1.4143) = \Phi(0.0354) = 0.514$. A coin flip with a tiny tilt.
- Both statements are true. The first supports "ship it if a 0.05 gain pays for the change". The second says "individual users will not notice". Different decisions need different quantities.
- Forecast twin. After $n = 30$ days of history, the 80% interval for the next day is $\pm 1.2816\,\sigma\sqrt{1 + 1/30} = \pm 1.303\sigma$. After infinitely many days: $\pm 1.2816\sigma$, the noise floor.
- Posterior decision quantity: $P(\theta_B \gt \theta_A \mid D) = \int\!\!\int \mathbf{1}[\theta_B \gt \theta_A]\, p(\theta_A, \theta_B \mid D)\, d\theta_A\, d\theta_B$, in practice the share of posterior draws with $\theta_B \gt \theta_A$ (Chapter 6.4). More useful: $P(\theta_B - \theta_A \gt \delta \mid D)$ for a gain $\delta$ that matters.
- Posterior predictive: $p(y_{\text{future}} \mid D) = \int p(y_{\text{future}} \mid \theta)\,p(\theta \mid D)\,d\theta$: draw $\theta$ from the posterior, then draw $y$ from the likelihood (Chapter 6.1, Chapter 7.14).
- Sources of forecast uncertainty: parameters (shrinks with data), observation noise (the floor), model structure, and the future values of regressors. Intervals widen with the horizon because trend and changepoint uncertainty accumulate.
- Credible vs confidence interval: a credible interval is a probability statement about the parameter given the data and the prior; a 95% confidence interval is a procedure that covers the truth in 95% of repeated experiments (Chapter 5.8).
Why do we need it?
Decisions come in two kinds: "which option is better on average?" and "how much will we actually see tomorrow?". Mixing them up leads to over-claiming ("99.9% sure" mistaken for "almost every user improves") or under-planning (capacity set from the mean, not the spread).
Where is it used?
Bayesian A/B decisions (probability of being best, expected loss, a minimum worthwhile gain), demand forecast fans, capacity and inventory planning from quantiles, and any prediction interval or credible interval report.
How is it used?
For decisions about parameters: compute the share of posterior draws above a margin $\delta$. For predictions: draw $\theta$ then $y$, report quantiles and exceedance probabilities such as $P(\text{demand} \gt \text{capacity})$, and say which kind of uncertainty each number carries.
"$P(B \gt A \mid D) = 0.99$ means B is better for 99% of users."
It means: given the data and the prior, the probability that B's true average is higher is 0.99. Individual users overlap heavily. How big the gain is, and whether it matters, is a separate question: use $P(\theta_B - \theta_A \gt \delta \mid D)$ for a worthwhile $\delta$.
"With enough data the forecast interval shrinks to a point."
Parameter uncertainty shrinks, observation noise does not. The interval for one future value never gets narrower than the noise.
"A 95% credible interval and a 95% confidence interval are the same thing."
They coincide numerically in some simple cases but answer different questions: the credible interval is a probability about the parameter given the data and prior; the confidence interval describes the long-run behaviour of a procedure.
"A forecast interval from the noise term alone is the whole uncertainty."
It leaves out parameter uncertainty (which grows with the horizon through the trend), model uncertainty and the uncertainty of future regressors.
Your A/B framework reports posterior decisions such as $P(\theta_A \gt \theta_B \mid D)$: statements about parameters, which sharpen as users accumulate and which depend on the prior and the guide (tails!). Your forecasting model reports $p(y_{\text{future}} \mid D)$: draw the parameters from the guide, then the noise from the likelihood (Normal, Student-t or Negative Binomial), so the fan combines both kinds of uncertainty and widens with the horizon. Do not quote one as if it were the other: a "95% sure the new page wins" is not a promise about any single visitor, and an 80% forecast band is not a statement about the uncertain parameters alone.
"What is the difference between the posterior and the posterior predictive?"
"The posterior $p(\theta \mid D)$ describes what we believe about the parameters; its width shrinks as data accumulate. The posterior predictive $p(y_{\text{new}} \mid D)$ averages the likelihood over that posterior, so it carries the parameter uncertainty plus the observation noise, and it has a noise floor. My A/B decisions use the first, for example $P(\theta_B \gt \theta_A \mid D)$ or the probability the gain exceeds a margin. My forecasts are the second: a distribution for future days."
$P(\theta_B \gt \theta_A \mid D)$ = share of posterior draws with $\theta_B \gt \theta_A$ (parameter question). $p(y_{\text{new}} \mid D) = \int p(y \mid \theta)\,p(\theta \mid D)\,d\theta$ (observation question).
Normal, known $\sigma$: gain posterior sd $\sigma\sqrt{2/n}$; one new value sd $\sigma\sqrt{1 + 1/n}$ (floor $\sigma$).
Trap: "99% sure it is better" $\ne$ "better for 99% of users"; intervals never shrink below the noise.
Quick check: $n = 100$ users per arm, gain 0.05σ. What is $P(\theta_B \gt \theta_A \mid D)$ and what is the chance a new B user beats a new A user?
Posterior sd $=\sqrt{2/100} = 0.1414$, so $P = \Phi(0.05/0.1414) = \Phi(0.354) = 0.638$: not convincing yet. For new users: sd $\sqrt{2(1 + 0.01)} = 1.4213$, $\Phi(0.0352) = 0.514$, essentially unchanged by $n$.
Theme 7 · Distribution diagnostics: Q-Q plots, KDE, residual plots, predictive checks core
A doctor does not judge your health from one number. She takes an X-ray for bones, a blood test for chemistry, and listens to your heart: different tools, each good at catching one kind of problem. A fitted model deserves the same treatment. A histogram or density plot shows the overall shape. A Q-Q plot shows the tails and the skew precisely. A residual plot shows whether the spread or a pattern depends on the fitted value or on time. A posterior predictive check asks any question you like of the model: "if it were true, would it produce data like mine, in zeros, variance and extremes?"
All four look at the leftovers: what the model did not explain. If the leftovers have structure, the model missed something.
Three ways to say it:
- Picture: four small plots, each answering one question about the leftovers.
- Numbers: a count series with mean 3.0 and variance 9.0: a Poisson model produces replicated variances between 2.2 and 4.0, so 9.0 is far outside; a Negative Binomial with $\alpha = 1.5$ produces 5.5 to 14.0, which covers it.
- Slogan: look at the leftovers from four sides before you trust the likelihood.
A posterior predictive check on 100 days of orders: observed mean 3.0, variance 9.0. Which likelihood is credible?
- Poisson fit: mean 3, so variance 3. Simulate many replicated series of 100 days from the fitted model and compute each replicate's variance: 95% of them fall between 2.2 and 4.0, median 3.0.
- The observed 9.0 is far outside that range. The Poisson cannot reproduce the data's spread, however well it matches the mean.
- Negative Binomial fit: mean 3 and $\text{Var} = \mu + \mu^2/\alpha = 3 + 9/1.5 = 9$. Replicated variances: 95% between 5.5 and 14.0, median 8.7. The observed 9.0 sits in the middle. The model reproduces the spread.
- Second check, zeros: Poisson(3) puts $e^{-3} = 5.0\%$ of days at zero; the NB with $\alpha = 1.5$ puts $(1.5/4.5)^{1.5} = 19.2\%$ there. If the data have about 20 zero days in 100, the NB is the one that gets them.
| Tool | Question it answers | What "bad" looks like | Typical response |
|---|---|---|---|
| Histogram, KDE (4.16) | What is the overall shape? Several humps? | skew, two humps, a spike at 0 | transform, mixture, zero-inflated model |
| Q-Q plot (4.17) | Do the tails and skew match the assumed distribution? | S-shape (heavy tails), bow (skew), jumps (outliers) | Student-t, transform, explain the outliers |
| Residual vs fitted | Does the spread depend on the level? | a funnel, a curve | log link, NB, variance model, missing term |
| Residual vs time, ACF, Ljung–Box (7.17) | Is there structure left in time? | runs, waves, spikes in the ACF | richer seasonality, AR error term, state-space model |
| Posterior predictive check (6.8, 7.14) | Can the model reproduce this feature of the data (mean, variance, zeros, maximum)? | observed value outside the replicated range | change likelihood or structure |
Rules of thumb used by the widget (screens, not tests; about one clean sample in ten trips the Ljung–Box screen by chance): tail ratio $(q_{95} - q_{50})/(q_{50} - q_{05})$ outside 0.625 to 1.6 means skew; excess kurtosis $\gt 1$ means heavy tails; correlation of $|\text{residual}|$ with the fitted value $\gt 0.3$ means a funnel; Ljung–Box $p \lt 0.05$ on 10 lags means autocorrelation.
Why do we need it?
A likelihood is an assumption. Fit statistics such as the ELBO do not tell you whether the assumption fits. Diagnostics show the specific way a model is wrong, and each shape points to a specific remedy.
Where is it used?
Regression diagnostics, ARIMA residual checks (Ljung–Box), Bayesian workflows (prior and posterior predictive checks), A/B analysis of metric distributions before a test, and model monitoring in production.
How is it used?
Standardize the residuals, draw the four plots, compute the screening numbers, name the problem (tails, skew, variance, autocorrelation), change one thing in the model and look again. Keep the plots in the model report.
"The Q-Q plot looks fine, so the model is fine."
A Q-Q plot sees only the marginal distribution of the residuals. Autocorrelation, a funnel in the spread and a missing seasonal term can all hide behind a straight Q-Q line.
"Residuals must be exactly Normal."
Compare to what the model assumes, and to what matters for the decision. Small wiggles are fine; ends that bend away (heavy tails) or a funnel are not. With few points the plot is noisy (simulate a reference band).
"Heavy tails: switch to Student-t immediately."
First ask whether the extreme residuals have a cause you can model (a holiday, a promotion, a bad data day). Heavy-tailed noise is for the surprises you cannot explain.
"A posterior predictive check proves the model is right."
It can only show that the model can or cannot reproduce the features you checked. Pick features that matter for the decision (zeros, variance, maximum, autocorrelation).
Both projects need this before and after the fit. In the A/B framework, look at the metric's histogram and Q-Q plot to choose between Normal, Student-t and Poisson, and run a predictive check for the zero share and the variance. In the forecasting model, the residual panel is the main review: Q-Q for the Normal-vs-Student-t decision, residual-vs-fitted for variance (and the Negative Binomial), ACF and Ljung–Box for missed structure (your model as described has no explicit AR error term, so autocorrelated residuals mean the trend, seasonality, holidays or regressors missed something), and replicated series for extremes and seasonality (Chapter 7.17).
"How do you know your likelihood is appropriate?" — "It converged."
"Convergence only says the optimizer finished. To check the likelihood I look at standardized residuals: a Q-Q plot for tails and skew, residual-vs-fitted for variance, the ACF and Ljung–Box for leftover structure, and posterior predictive checks for the features I care about, such as variance and zero counts. Each pattern has a specific response: Student-t for tails, a log link or Negative Binomial for variance that grows, a richer seasonal term for autocorrelation."
Four views: shape (histogram, KDE), tails (Q-Q), pattern (residual vs fitted and vs time, ACF), reproduction (posterior predictive check).
Signature → remedy: S-shaped Q-Q → Student-t · funnel → log link / NB / variance model · waves in the ACF → missed structure · replicated variance too small → NB.
Trap: Q-Q alone misses structure in time and in the spread.
Quick check: the residual-vs-fitted plot is a funnel (small spread at low forecasts, large at high ones). Which of your likelihood choices could explain it, and which fix is natural?
A Normal likelihood with constant $\sigma$ assumes the spread does not depend on the level. For counts and amounts the spread usually grows with the mean, as for a Negative Binomial ($\mu + \mu^2/\alpha$). Natural fixes: a Negative Binomial (or another likelihood whose variance grows with the mean), a log link or log transform so effects are multiplicative, or an explicit variance model.
Theme 8 · Computational scalability: JIT, SVI, low-rank guides, model dimension core
Cooking for 4 people and cooking for 4,000 need different kitchens. A model has three kinds of cost. (1) How many numbers the guide must learn. A guide that tracks the correlation between every pair of parameters needs a table of all pairs, which grows like the square of the number of parameters. (2) Work per step. Handling that table (a matrix factorization) grows like the cube. (3) Overhead. Python is slow in a loop; JIT compilation turns one step into fast machine code, but compiles again whenever the shapes of its inputs change.
The remedies match the costs: a low-rank guide keeps only the main correlations, SVI avoids sampling thousands of draws, and static shapes with jnp.where keep the compiled step reusable.
Three ways to say it:
- Picture: a full square table of all parameter pairs (full-rank) versus a thin strip of the main directions (low-rank).
- Numbers: with $d = 1000$ parameters a full-rank guide has 501,500 numbers, a rank-5 guide 7,000 and a mean-field guide 2,000.
- Slogan: pay only for the correlations you need; compile once, run many; keep shapes fixed.
- Count the parameters $d$. A forecasting model with 25 changepoint candidates, yearly order 10 (20 columns), weekly order 3 (6 columns), 8 holiday columns, 4 regressors, base slope and intercept (2) and a Student-t likelihood (scale and degrees of freedom: 2): $d = 25 + 20 + 6 + 8 + 4 + 2 + 2 = 67$. (Illustrative; use your own counts.)
- Guide sizes (free numbers): mean-field $2d = 134$. Full-rank $d + d(d+1)/2 = 67 + 2278 = 2345$. Rank-5: $d(r + 2) = 67 \times 7 = 469$. For $d = 1000$: $2000$, $501{,}500$ and $7000$.
- Memory. A full-rank covariance factor stored as a $d \times d$ float32 matrix takes $4d^2$ bytes: 18 kB at $d = 67$, 4 MB at $d = 1000$, 400 MB at $d = 10{,}000$. Its Cholesky factorization costs about $d^3/3$ operations: $3 \times 10^{8}$ at $d = 1000$ but $3 \times 10^{11}$ at $d = 10{,}000$.
- JIT arithmetic (illustrative numbers). Compile once: 5 s. A compiled step: 2 ms; an un-compiled Python step: 40 ms. For 3,000 steps: $5 + 3000 \times 0.002 = 11$ s versus $3000 \times 0.04 = 120$ s. Break-even after $5/0.038 \approx 132$ steps. If five different input shapes each trigger a compile, add $4 \times 5 = 20$ s.
- Model dimension $d$: the number of free (unconstrained) latent numbers the guide must describe, counting every weight, scale and hyperparameter.
- Guide sizes:
AutoNormal(mean-field) $2d$;AutoMultivariateNormal(full-rank) $d + d(d+1)/2$ free numbers (NumPyro stores the lower-triangular factor as a $d \times d$ array);AutoLowRankMultivariateNormal$d(r + 2)$ (a location, a $d \times r$ factor and $d$ diagonal scales). I checked these counts in NumPyro on a 40-parameter model: 80, 860 free (1640 stored) and 280 for $r = 5$ (Chapter 6.13). - JIT: Python is traced once with abstract values, converted to XLA and compiled; later calls with the same shapes and dtypes reuse the compiled code. New shapes, dtypes or static arguments recompile. Data-dependent shapes (
x[mask]with a boolean mask) cannot be traced, so usejnp.whereor masked sums (Chapter 6.17). - SVI scales because each step is one gradient on one or a few Monte Carlo draws (and, for large data, on a mini-batch scaled by $N/B$) (Chapter 6.12).
Why do we need it?
A model that is right but too slow or too big is never used, and one that is refit weekly must finish quickly. Cost decides which guide, which inference method and how many components you can afford.
Where is it used?
JAX and NumPyro training loops, Gaussian-process and neural-network approximations (low-rank structure), mini-batch SGD, XLA-compiled code on CPU, GPU and TPU, and any production model that is refit on a schedule.
How is it used?
Count $d$, compare the guide sizes, start with the cheapest guide that captures the correlations you care about, jit the update step, keep shapes fixed (pad and mask), and measure: compile time, time per step, memory.
"A bigger guide is always better, so use full-rank."
Full-rank captures all correlations but costs $d^2$ numbers and $d^3$ work, and with limited data and a noisy ELBO the extra numbers are harder to fit. Low-rank often captures the important correlations at a fraction of the cost. Compare on a small case.
"JIT makes everything fast."
The first call pays compilation, every new shape pays again, and tiny functions may not be worth it. Fix the shapes (pad and mask) and keep the compiled step in the loop.
"Mask the bad rows with x[mask] before the model; it is the same as jnp.where."
Boolean indexing has a data-dependent shape, which cannot be traced under jit. Use jnp.where, masked sums or a mask argument so the shape stays fixed (Chapter 6.17).
"Low-rank means a low-rank model."
The low-rank structure belongs to the guide's covariance, an approximation of the posterior's correlations. The model keeps all its parameters.
Your forecasting model chooses between a full-rank and a low-rank Gaussian guide based on model size, and your SVI update is JIT-compiled. Both are scalability decisions on the curves in the widget: count $d$ from your components, record the choice (and what triggered it) in the run manifest (Chapter 7.19), and remember that a model that grows (more changepoint candidates, higher Fourier order, more regressors) can cross the threshold and change the guide, and with it the width of your intervals. Your A/B framework met the other half of the lesson: boolean masking before tracing gives data-dependent shapes under JIT; the fix is to keep the shapes fixed and mask with jnp.where or masked sums. I will not guess your threshold: read it from your code.
"Why did you use a low-rank guide?" — "It was faster."
"A full-rank Gaussian guide has $d + d(d+1)/2$ free numbers and $O(d^3)$ work, which is fine for small models and heavy for large ones. A low-rank guide uses $d(r+2)$ numbers and still captures the main posterior correlations, which a mean-field guide would ignore, understating uncertainty. My loop picks the guide from the model size; to be safe I would check on a small case that the low-rank result matches full-rank, and log which guide was used."
Guide sizes: mean-field $2d$ · low-rank $d(r+2)$ · full-rank $d + d(d+1)/2$ (memory $\approx 4d^2$ bytes in float32, work $\approx d^3/3$).
JIT: trace once, compile, reuse; new shapes recompile; no boolean masks (use jnp.where); break-even steps $=$ compile time ÷ time saved per step.
Trap: bigger is not better; quoting a size threshold you have not read in your code.
Quick check: $d = 200$ and $r = 10$. How many numbers do the full-rank and low-rank guides have?
Full-rank: $200 + 200 \times 201/2 = 200 + 20{,}100 = 20{,}300$. Low-rank: $200 \times 12 = 2400$, about 8.5 times fewer. Mean-field: 400.
The interview drill: 40 short questions with model answers core
An interview is a conversation, and good answers have a shape. If you only memorize facts, the first follow-up ("why?", "what if the data are skewed?") breaks the answer. If you practise the shape, you can build an answer on the spot: say the idea in plain words, show one number, point at your own project, and name the trap you avoid.
This section has 40 questions, five for each of the eight themes. Most come from the "Say it right" boxes of the whole series, now seen through both of your projects. Answer each one out loud in about a minute before you look.
Three ways to say it:
- Picture: a four-step staircase: plain words, one number, your project, the trap.
- Numbers: 40 questions, 8 themes, about 60 seconds each: a full drill is under 45 minutes.
- Slogan: say it simply, show a number, point at your project, name the trap.
Question: "Does a Student-t likelihood remove outliers?" The four steps:
- Plain words: "No. It has heavier tails, so extreme residuals are less shocking and pull the fit less."
- One number: "With 4 degrees of freedom a residual 10 scale-units away pulls with force 0.48, a Normal would pull with force 10."
- Your project: "In my forecasting model it keeps a few spike days from bending the trend and from inflating the noise level, and the same likelihood is available for continuous metrics in the A/B framework."
- The trap: "The outliers are still in the data and in the residual plot, and if a spike has a known cause, like a holiday, I model the cause instead."
- The four-part answer: (1) plain definition; (2) a number or a tiny picture; (3) where it lives in your project; (4) the classic trap.
- Self-grading: 2 = all four parts, no notes. 1 = one part missing. 0 = I had to look. Revisit the 0s and 1s the next day (the "review pile").
- Honesty rule: describe your projects only with the facts you know (listed in the matrix at the top of this chapter). For anything else, say "I would have to check the code".
Why do we need it?
Understanding that you cannot say aloud under pressure is not yet useful. Speaking an answer once finds the gaps that reading never shows, and repeating the shape makes the explanations fast and calm.
Where is it used?
Applied-scientist and ML-engineer interviews, project reviews, design discussions, and explaining a model to a non-technical stakeholder who asks "but how sure are we?".
How is it used?
Pick a theme, press "Next question", answer aloud, press "Show the model answer", grade yourself, and add misses to the review pile. Follow the links under a model answer to re-read the chapter.
The 40 questions in order, with model answers. Open one to read it; the links go to the chapters that teach the idea.
Theme 1 · Choosing the likelihood
1. How do you choose a likelihood for a new metric?
I start from two things: the support of one observation (yes/no, categories, counts, positive amounts, any real number) and how its variance behaves (constant, growing with the mean, heavy tails). Yes/no gives a Binomial with a Beta prior; categories a Multinomial with a Dirichlet; counts with variance equal to the mean a Poisson, with variance above the mean a Negative Binomial; real values with light tails a Normal, with heavy tails a Student-t; positive skewed amounts a log transform first. Then I check the choice with a Q-Q plot, a variance-versus-mean plot or a posterior predictive check. Revise: 4.7 · 4.8 · 4.9 · 4.11 · 7.13
2. Daily orders have mean 100 and variance 400. Poisson or Negative Binomial, and what is the dispersion?
Negative Binomial. A Poisson forces the variance to equal the mean (100), four times too small, so its intervals would be too narrow. With the mean–dispersion form $\text{Var} = \mu + \mu^2/\alpha$ we get $400 = 100 + 10000/\alpha$, so $\alpha = 33.3$; as $\alpha \to \infty$ the Negative Binomial becomes Poisson. I would confirm with a posterior predictive check of the variance. Revise: 4.8 · 7.13
3. Why not use a Normal likelihood for a conversion metric?
A conversion is yes/no, so the number of conversions out of $n$ users is Binomial with variance $np(1-p)$, which depends on $p$. A Normal can give impossible values (rates below 0 or above 1) and gets the variance wrong for small $n$ or rare events. The Beta-Binomial is exact and conjugate: Beta$(\alpha + k, \beta + n - k)$. For large $n$ the Normal approximation to the rate works (the CLT), but that is an approximation, not the model. Revise: 4.7 · 6.3
4. Which Negative Binomial parameterizations exist, and how do you avoid mixing them up?
NumPyro NegativeBinomial2(mean, concentration) has $\text{Var} = \mu + \mu^2/\alpha$. NegativeBinomialProbs and NegativeBinomialLogits take total_count and probabilities or logits. SciPy nbinom(n, p) uses $n = \alpha$ and $p = \alpha/(\alpha + \mu)$. I check which one my code calls, and I verify by simulating a few thousand draws: the sample mean and variance must match $\mu$ and $\mu + \mu^2/\alpha$. Revise: 4.8 · 7.13
5. A metric has many exact zeros and otherwise positive amounts. What do you do?
First I check what a zero means in the metric definition. If zeros come from a separate "did not act" process, I model two parts: whether the user acts (Bernoulli) and the amount given action (log-normal, Gamma, or a Normal on the log scale), which is a hurdle model; for counts a zero-inflated model. A single Normal on the raw values fits neither the spike at zero nor the skew. I then check the zero share and the tail with a posterior predictive check. Revise: 4.8 · 4.10 · 4.18
Theme 2 · Robustness
6. Does a Student-t likelihood remove outliers?
No. It assigns more probability to extreme residuals, so they exert less influence on the fit. The influence of a residual $r$ on a Normal fit grows like $r$; on a Student-t fit it is $(\nu+1)r/(\nu + r^2)$, which is bounded and falls for large $r$ (at $\nu = 4$ a residual of 10 pulls with 0.48 instead of 10). The outlier stays in the data and in the residual plot. Revise: 4.9 · 7.13
7. What happens to a Student-t as the degrees of freedom grow, and what if they are very small?
As $\nu \to \infty$ it becomes a Normal. For $\nu \le 2$ the variance is infinite and for $\nu \le 1$ even the mean is undefined (the Cauchy). A very small $\nu$ says "almost anything can happen", which widens intervals and weakens the other parameters. I give $\nu$ a prior or fix a moderate value, and check how far the posterior moves from its prior. Revise: 4.9 · 7.13
8. How can residuals tell you that you need a Student-t instead of a Normal?
In a Q-Q plot against a Normal the two ends bend away from the line (an S-shape) while the middle follows it, and the excess kurtosis is clearly positive. Before changing the likelihood I check that the extreme residuals are not explained by something I left out, such as holidays or promotions. Revise: 4.17 · 7.17
9. One huge day is pulling your trend. Student-t, or a new regressor?
If I know the cause (a holiday, a campaign), I model it with an indicator or a regressor: it explains the day and sharpens the other parameters. The Student-t is for surprises I cannot explain. If the day is a data error, I fix or flag the data. Using heavy tails to hide a known event wastes information. Revise: 7.12 · 7.13
10. Is the scale of a Student-t the standard deviation?
No. The standard deviation is $\text{scale}\times\sqrt{\nu/(\nu - 2)}$ for $\nu \gt 2$; for $\nu = 4$ it is 1.41 times the scale. NumPyro StudentT(df, loc, scale) takes the scale. The difference matters when you compare the noise level of a Student-t model with a Normal model. Revise: 4.9
Theme 3 · Shrinkage
11. What do hierarchical partial pooling and the Laplace prior on changepoints have in common?
Both are priors that pull noisy estimates toward a centre (the overall mean for groups, zero for slope changes) by an amount set by the evidence compared with the prior spread. Both give up a little bias for less variance, so the total error falls. They differ in shape: pooling shrinks every group proportionally, while the Laplace MAP soft-thresholds small values to exactly zero. Revise: 5.3 · 6.6 · 7.10
12. How much does a small segment shrink toward the overall mean?
In the Normal-Normal model the weight on its own average is $w = (n/\sigma^2)/(n/\sigma^2 + 1/\tau^2)$. With $\sigma = 10$ and $\tau = 2$, a segment of 4 users has $w = 0.14$ and a segment of 100 users has $w = 0.8$. In a hierarchical model $\tau$ is learned from the data. Revise: 6.5 · 6.6
13. Is the posterior under a Laplace prior sparse?
Only the MAP is. With a Gaussian likelihood the MAP is the soft threshold $\mathrm{sign}(z)\max(|z| - s^2/b, 0)$. The full posterior is a continuous distribution: its mean and median are shrunk toward zero but not exactly zero, and $P(\delta = 0 \mid D) = 0$. I say "shrunk" or "sparse-ish" and state whether I mean the MAP or the posterior. Revise: 5.3 · 7.10
14. What happens to the pooled estimates as $\tau \to 0$ and as $\tau \to \infty$?
As $\tau \to 0$ the groups are forced to be identical: complete pooling, every estimate equals the overall mean. As $\tau \to \infty$ the groups are unrelated: no pooling, each uses only its own data. Partial pooling is the learned middle, and groups with little data move the most. Revise: 6.5 · 6.6
15. Why not standardize each segment separately to remove segment differences?
Per-group standardization removes exactly the between-group differences I want to estimate, setting every group's mean to zero. I use one global scaler fitted on the training data and let the hierarchical model share information across groups. Revise: 4.18 · 6.6
Theme 4 · Bias–variance
16. Explain the bias–variance trade-off using your forecasting model.
Every flexibility knob moves along it: more Fourier harmonics, more changepoints or a larger Laplace scale lower the bias and raise the variance. Training error always falls with flexibility, while holdout error is U-shaped. I choose settings with rolling-origin validation and prefer the simplest setting within the noise of the best. Revise: 5.1 · 7.11 · 7.18
17. Name the knobs of your two projects and the direction of each.
Forecasting: Fourier order (higher is more flexible), the number of changepoints and the Laplace scale $b$ (more or larger is more flexible), regressors and holiday windows (more columns), guide richness. A/B: the between-segment scale $\tau$ (larger means less pooling) and the strength of the priors (weaker means more flexible). Too flexible gives variance, too stiff gives bias. Revise: 6.6 · 7.10 · 7.11 · 7.18
18. How do you choose the Fourier order or the changepoint scale?
Not by the training fit. I fit a range of values, score each on rolling-origin holdouts with a point metric and a probabilistic metric such as CRPS, check stability across origins and seeds, and take the simplest setting within the noise of the best. Prior predictive checks help set sensible ranges. Revise: 7.15 · 7.16
19. Can a biased estimator beat an unbiased one?
Yes, because $\text{MSE} = \text{bias}^2 + \text{variance}$. A shrunk estimate $c\bar x$ with $c = \mu^2/(\mu^2 + \sigma^2/n)$ has a lower MSE than $\bar x$: with $\mu = 2$ and $\sigma^2/n = 1$, $c = 0.8$ gives MSE 0.80 against 1.00. Pooling, ridge and Laplace shrinkage all rely on this. Revise: 5.1 · 5.3
20. Training error keeps dropping but holdout error rises. What is happening, and what do you do?
Overfitting: the model is following noise in the training window (variance). I reduce flexibility or add shrinkage (fewer harmonics, smaller $b$, stronger priors), make sure the validation is leakage-free, and add data or regressors that explain real structure. Revise: 7.15 · 7.18
Theme 5 · Approximate inference
21. Why do you need approximate inference at all?
The posterior needs the evidence integral over all parameters, which is intractable for realistic models. Conjugate pairs such as Beta-Binomial are exact. Otherwise I use MCMC (NUTS), variational inference (SVI) or a Laplace approximation. SVI is fast and scales, with guide-dependent error. Revise: 6.9 · 6.11
22. SVI or NUTS: how do you choose?
SVI is optimization-based: fast, scalable and guide-dependent, so its intervals can be too narrow. NUTS is sampling-based: costly, asymptotically targets the posterior, gives richer uncertainty, and needs diagnostics ($\hat R$, effective sample size, divergences). I use SVI for speed and refits, would use NUTS on a subset to validate, and I never say NUTS is exact: a finite run has Monte Carlo error. Revise: 6.10 · 6.15
23. What does the ELBO equal, and why does maximizing it fit the posterior?
$\log p(D) = \mathrm{ELBO} + KL(q \,\|\, p(\theta \mid D))$. Since $\log p(D)$ does not depend on $q$, maximizing the ELBO minimizes the KL divergence from the guide to the posterior. The ELBO is $E_q[\log p(D, \theta)] - E_q[\log q(\theta)]$: expected fit plus the entropy of $q$. Revise: 6.11 · 6.12
24. Mean-field, full-rank and low-rank guides: what is the difference?
Mean-field: independent Gaussians, $2d$ numbers, ignores correlations and usually understates joint uncertainty. Full-rank: a full covariance, $d + d(d+1)/2$ free numbers, $O(d^3)$ work. Low-rank: a diagonal plus a rank-$r$ part, $d(r+2)$ numbers, capturing the main correlations. I choose by model size and compare with full-rank on a small case. Revise: 6.13
25. How do you decide when to stop an SVI run?
The ELBO is noisy, so I compare with the best value so far using a relative tolerance (absolute thresholds depend on the data scale), add patience (many loops also set a minimum number of steps) and keep the best state, not the last. NumPyro's svi.update returns the loss, the negative ELBO, so "improvement" means the loss went down. I also check across seeds. Revise: 6.14
Theme 6 · Uncertainty
26. What is the difference between $P(\theta_B \gt \theta_A \mid D)$ and the posterior predictive distribution?
The first is a statement about parameters: given the data and prior, how likely is it that B's true average is higher. Its uncertainty shrinks as data accumulate. The posterior predictive $p(y_{\text{new}} \mid D)$ averages the likelihood over the posterior, so it carries parameter uncertainty plus observation noise and has a noise floor. My A/B decisions use the first; my forecasts are the second. Revise: 6.1 · 6.4 · 7.14
27. $P(B \gt A \mid D) = 0.99$. Does B win for 99% of users?
No. 0.99 is the probability that B's true average is higher. Individual users overlap heavily: with a true gain of 0.05 standard deviations and 10,000 users per arm the probability is 0.9998, yet a random B user beats a random A user only 51.4% of the time. I report the size of the gain and $P(\text{gain} \gt \delta)$ for a worthwhile $\delta$. Revise: 6.4
28. Which uncertainties does a forecast interval contain?
Parameter uncertainty (the posterior over level, trend, seasonal and holiday weights), observation noise (the likelihood), model uncertainty (structure, and changepoints chosen before fitting are not propagated) and the uncertainty of future regressor values. The posterior predictive includes the first two; I say what is missing, and check coverage on a backtest. Revise: 7.14 · 7.16
29. Why do forecast intervals widen with the horizon?
The trend is extrapolated: a small error in the last slope is multiplied by the number of days ahead, and future slope changes are possible, so trend uncertainty grows with the horizon. Seasonal and noise uncertainty stay roughly constant, so the fan opens up from a floor set by the noise. Revise: 7.8 · 7.14
30. Credible interval versus confidence interval, in one minute.
A 95% credible interval says: given the data and my prior, the parameter lies in this range with probability 0.95. A 95% confidence interval comes from a procedure that covers the true value in 95% of repeated experiments; any one interval either contains it or not. They can coincide numerically in simple cases, but only the credible interval is a probability statement about the parameter. Revise: 5.8 · 6.4
Theme 7 · Distribution diagnostics
31. How do you check whether a Normal likelihood is reasonable?
I look at standardized residuals from four sides: a histogram or KDE for shape, a Q-Q plot against a Normal for tails and skew, residual-versus-fitted for the spread, and the ACF (with Ljung–Box) for leftover structure in time. Then a posterior predictive check on the features that matter for the decision, such as the variance and the maximum. Revise: 4.16 · 4.17 · 7.17
32. What does an S-shaped Q-Q plot mean?
Heavy tails: the sample's quantiles are more extreme than the Normal's at both ends (low points below the line, high points above) while the middle follows the line. A bow in one direction means skew. I then consider a Student-t after ruling out outliers that I can explain. Revise: 4.17
33. Your residuals have significant autocorrelation. What does that tell you, and what can you do?
The model missed temporal structure: the trend, seasonality, holidays or regressors do not capture everything. Since the model as I built it has no explicit AR error term, options are richer seasonal terms or changepoints, an AR error component, a state-space model or a Gaussian process. Intervals that assume independent residuals are too narrow. Revise: 7.3 · 7.17
34. What is a posterior predictive check, and what would you check for a count series?
Simulate replicated datasets from the fitted model, compute a statistic on each, and compare with the observed value. For counts: the variance, the share of zeros, the maximum and the autocorrelation. If the observed variance lies outside the replicated range, as with a Poisson model for overdispersed counts, the likelihood is wrong. Revise: 6.8 · 7.14
35. The residual-versus-fitted plot is a funnel: the spread grows with the fitted value. What now?
Constant-variance Normal noise is wrong. Options: model on the log scale or use a log link (multiplicative effects), use a Negative Binomial for counts, or model the variance explicitly. Then re-check the residual-versus-fitted plot. Revise: 5.14 · 7.13 · 7.17
Theme 8 · Computational scalability
36. Why does a full-rank guide get expensive, and what is the alternative?
It has $d + d(d+1)/2$ free numbers and needs $O(d^3)$ work for the Cholesky factor: at $d = 1000$ that is 501,500 numbers. A low-rank guide uses $d(r+2)$: 7,000 for $r = 5$, capturing the main correlations. Mean-field uses $2d = 2000$ but ignores correlations. Revise: 6.13 · 6.17
37. What does JIT do, and what are its catches?
JIT traces the function once, compiles it with XLA and reuses the compiled code for inputs with the same shapes and dtypes. The first call is slow, every new shape recompiles, Python control flow that depends on traced values needs lax.cond or jnp.where, and data-dependent shapes such as x[mask] fail. Revise: 6.16 · 6.17
38. You hit an error with boolean masking under JIT. What happened, and how did you fix it?
x[mask] has a shape that depends on the data, but JIT needs static shapes at trace time. The fix is to keep the full array and mask with jnp.where or masked sums, so the shape is fixed and invalid entries contribute zero. In a log-likelihood a masked entry should add 0, and if a log or square root is involved I use the double-where pattern to keep gradients finite. Revise: 6.17 · 7.19
39. How would you speed up a slow SVI loop?
Check that the update step is JIT-compiled and its shapes are fixed; avoid Python work and device transfers inside the loop; evaluate and log the ELBO every few steps; use a cheaper guide (low-rank or mean-field) if correlations are not essential; scale the inputs so a larger learning rate is stable; use mini-batches for large data; stop early with a relative rule. I measure before and after. Revise: 6.14 · 6.17
40. What would you put around a forecasting model in production?
A run manifest saved with every fit (data hash, scaler numbers, code commit, priors, seed, inference settings, guide, optimizer, iterations run, library versions); monitoring of rolling error, bias and interval coverage plus input checks (missing share, ranges, drift) and changepoint counts; and numerical safeguards (log-space likelihoods, scale parameters behind exp or softplus, scaled inputs, gradient clipping, a guard that never accepts a non-finite loss). An alarm triggers an investigation, not an automatic retrain. Revise: 7.19
"I read the answer and it made sense, so I know it."
Recognizing an answer is far easier than producing it. Say it aloud first, then compare.
"Learn the model answers word for word."
Learn the skeleton (plain words, number, project, trap). A recited paragraph falls apart at the first follow-up; a skeleton can be rebuilt.
"Say that your model does what the textbook says."
Interviewers ask about your code. If you are not sure whether your model fixes $\nu$ or learns it, or which Negative Binomial form it calls, say so and say how you would check.
Use the questions as a bridge between your two projects. For each answer, practise two endings: one in the A/B framework ("in my experimentation framework, this is the Beta-Binomial arm comparison / the hierarchical segment model") and one in the forecasting model ("in my forecast, this is the Laplace scale on the changepoint slopes / the Student-t noise / the guide choice"). Being able to switch between them shows that you understand the idea and not only the code.
"I don't remember the formula for that." (and stop)
"I don't remember the exact formula, but here is the idea and how I would derive or look it up..." then give the picture and the reason. A clear idea with a missing formula is a good answer; a formula with no idea is not.
Answer shape: plain words → one number → your project → the trap. Grade 2 / 1 / 0; revisit the 0s and 1s the next day.
Five questions per theme: likelihood, robustness, shrinkage, bias–variance, inference, uncertainty, diagnostics, scale.
Trap: reading answers instead of speaking them; claiming code details you have not checked.
Quick check: which part of the four-part answer do beginners most often forget?
The trap. Naming the classic mistake ("it down-weights outliers, it does not delete them"; "only the MAP is sparse"; "NUTS is asymptotically exact, not exact") shows that you have used the idea and not only read about it.
The core (P0) checklist: 44 items, each with a link core
A checklist is a way to find holes. "I read the chapter" is not the same as "I can explain it in a minute". The syllabus names 44 core topics, called P0 (priority zero): the ones you must be able to explain well to be comfortable with both projects. Below, every item has a one-line prompt of what you should be able to say, and links to the chapters that teach it.
Tick an item only after you have said it aloud, without notes, in about a minute. Anything you cannot tick is a link to follow.
Three ways to say it:
- Picture: 44 boxes; a progress bar fills as you tick them.
- Numbers: 8 items start in Guide 1, 7 in Guide 2, 15 in Guide 3 and 14 in Guide 4 (44 in all).
- Slogan: if you cannot say it in a minute, follow the link.
Item 17, Beta-Binomial, as a one-minute check.
- Sentence: "Conversions out of $n$ users are Binomial; with a Beta prior the posterior is again Beta, so no inference machinery is needed."
- Update: prior Beta$(\alpha, \beta)$ and $k$ conversions out of $n$ give Beta$(\alpha + k,\ \beta + n - k)$; $\alpha$ and $\beta$ act like pseudo-counts.
- Number: flat Beta$(1,1)$ with 50 of 500: Beta$(51, 451)$, mean $51/502 = 0.1016$, sd $0.0135$.
- Project: "In my A/B framework this is the conversion-metric model; two arms give two Betas, and I compare them by $P(\theta_B \gt \theta_A \mid D)$ from posterior draws."
- P0: the syllabus's list of core topics, the ones marked as must-know in the learning plan.
- Ticking rule: tick when you have explained the item aloud, without notes, covering the idea, one formula or number, and where it appears in a project.
- Where the items sit: items most tied to the A/B framework: 7 to 10, 14 to 18, 41; to the forecasting model: 24 to 36, 44; shared by both: 11 to 13, 19 to 23, 37 to 40, 42, 43; foundations used everywhere: 1 to 6.
Why do we need it?
A syllabus has hundreds of bullets; a short list shows what matters most. A checklist with links turns "I should revise" into "item 28, PELT, follow this link now".
Where is it used?
The week before an interview, when planning what to revise, when onboarding a colleague to your projects, and as a self-assessment at the end of this series.
How is it used?
Go down the list. Say each item aloud. Tick the ones you can explain; follow the links for the rest, then come back. The progress bar and the "pick my next item" button choose a random unticked item so you do not only revise your favourites.
| # | Item and what you should be able to say | Where to revise | |
|---|---|---|---|
| 1 | Probability and random variables Events, the rules of probability, and how a random variable turns outcomes into numbers (PMF, PDF, CDF). | 4.2 · 4.3 · 4.4 | |
| 2 | Expectation, variance, covariance, correlation The average, the spread around it, and how two variables move together (and the scale-free version). | 4.5 · 4.6 · 4.15 | |
| 3 | Pearson vs Spearman Straight-line association (outlier-sensitive) versus rank association (any monotone curve). | 4.15 | |
| 4 | LLN and CLT Averages settle on the mean; the average, not the data, becomes Normal with standard deviation σ/√n. | 4.12 · 4.13 | |
| 5 | Sampling distributions and standard error How an estimate would vary over repeated samples; SE of a mean is σ/√n. | 5.4 · 5.5 | |
| 6 | Estimators, bias, variance and MSE MSE = bias² + variance, and why a biased estimator can win. | 5.1 | |
| 7 | Hypothesis testing Null, test statistic, null distribution, p-value, decision; the logic, not just the recipe. | 5.6 · 5.9 | |
| 8 | p-values and confidence intervals What a p-value and a 95% interval do, and do not, say. | 5.6 · 5.8 | |
| 9 | Type I/II error and power False alarms, misses, power = 1 − β, and what drives it. | 5.7 | |
| 10 | A/B testing Design (unit, metrics, sample size), pitfalls (SRM, peeking, interference) and the causal reading. | 5.10 · 5.11 · 5.12 | |
| 11 | Normal, Student-t, Binomial, Beta, Poisson, Negative Binomial, Multinomial, Dirichlet and Laplace For each: support, parameters, variance, when to use it, how it relates to the others. | 4.7 · 4.8 · 4.9 · 4.11 | |
| 12 | Bayesian inference Posterior ∝ likelihood × prior; the evidence; sequential updating. | 6.1 | |
| 13 | Priors, likelihood, posterior, posterior predictive Each term in words; how to choose and check a prior. | 6.1 · 6.2 | |
| 14 | Credible intervals and posterior probabilities Equal-tailed vs highest-density; P(B > A | D); P(gain > δ | D). | 6.4 | |
| 15 | Hierarchical Bayesian models Group parameters drawn from a shared distribution; hyperparameters; centered vs non-centered. | 6.5 · 6.7 | |
| 16 | Partial pooling and shrinkage The weight formula; τ → 0 and τ → ∞; Laplace shrinkage of changepoint slopes. | 6.6 · 7.10 | |
| 17 | Beta-Binomial Beta(α + k, β + n − k); pseudo-counts; the posterior mean as a weighted average. | 6.3 | |
| 18 | Dirichlet-Multinomial Add the category counts to α; marginals are Beta. | 6.3 | |
| 19 | MCMC / NUTS A Markov chain targeting the posterior; HMC and NUTS; R̂, effective sample size, divergences. | 6.9 · 6.10 | |
| 20 | Variational inference Turn inference into optimization; the guide family and its limits. | 6.11 | |
| 21 | KL divergence and ELBO log p(D) = ELBO + KL(q‖posterior); why reverse KL is mode-seeking. | 6.11 · 6.12 | |
| 22 | SVI Noisy gradients of the ELBO, the reparameterization trick, the optimizer, and your training loop. | 6.12 · 6.14 | |
| 23 | Full-rank vs low-rank guides Parameter counts, correlations captured, cost, and how to choose. | 6.13 | |
| 24 | Time-series decomposition Trend, season, holidays, regressors, noise; additive versus multiplicative. | 7.2 · 7.7 | |
| 25 | Autocorrelation and stationarity ACF and PACF, white noise versus random walk, differencing. | 7.3 · 7.4 | |
| 26 | Piecewise-linear trends Base slope, slope adjustments δ_j, and the offsets γ_j = −s_j δ_j that keep the line connected. | 7.8 | |
| 27 | Changepoints Why trends change; too many versus too few; grid, detection and latent approaches. | 7.8 | |
| 28 | PELT Segment cost plus penalty, pruning; selection uncertainty is not propagated into the posterior. | 7.9 | |
| 29 | Laplace changepoint priors δ_j ~ Laplace(0, b); the MAP is a soft threshold; the posterior is shrunk, not exactly sparse. | 7.10 · 5.3 | |
| 30 | Fourier seasonality Harmonics, order N (2N columns), amplitude and phase, aliasing. | 7.11 | |
| 31 | Holiday and exogenous regressors Indicator columns and windows; future availability of regressors. | 7.12 | |
| 32 | Forecasting leakage Future information in features, scalers, PELT or feature selection; random K-fold. | 7.12 · 7.15 | |
| 33 | Normal vs Student-t vs Negative Binomial forecasting What each assumes and when each is the right noise model. | 7.13 | |
| 34 | Rolling time-series validation Rolling origin, expanding versus sliding window, error by horizon. | 7.15 | |
| 35 | Probabilistic forecast evaluation and calibration Coverage, PIT, sharpness, CRPS, log score; monitoring calibration. | 7.16 · 7.19 | |
| 36 | Residual diagnostics Mean, variance, Q-Q, residual plots, ACF, Ljung–Box: what each pattern means. | 7.17 | |
| 37 | Q-Q plots How to build one and what the shapes mean. | 4.17 | |
| 38 | KDE A bump on every point; the bandwidth matters far more than the kernel. | 4.16 | |
| 39 | Box-Cox Power transform, λ by maximum likelihood, needs positive data; back-transform caveat. | 4.18 | |
| 40 | PCA Directions of maximum variance; eigenvectors of the covariance; explained variance. | 5.16 · 5.15 | |
| 41 | JAX / JIT fundamentals Pure functions, grad, vmap, scan, PRNG keys, tracing, static shapes, the masking trap. | 6.16 · 6.17 | |
| 42 | Identifiability and misspecification Which component gets the credit; what a wrong model looks like in the checks. | 6.8 · 7.7 · 7.17 | |
| 43 | Prior / posterior predictive checks Simulate before and after fitting and compare features that matter. | 6.2 · 6.8 · 7.14 | |
| 44 | Model comparison and alternatives Bias–variance–complexity; ARIMA, exponential smoothing, state-space, GP, boosting, neural models. | 7.18 · 7.5 · 7.6 |
"I recognize the title, so I can tick it."
Tick only after saying it aloud without notes. Recognition is the first step, not the last.
"All 44 items matter equally for my interview."
Start with the items tied to your two projects (see the grouping above), then the shared ones, then the foundations. But do not skip foundations: an interviewer's follow-up often goes there.
"A complete checklist means I am done."
It means you have covered the core. Depth comes from the drill, from the code blocks (change a number, predict the result, run it) and from explaining the ideas to someone else.
The list is the skeleton of your two projects. For the A/B framework: items 7 to 10 (testing and A/B), 11 and 14 to 18 (the likelihoods, credible intervals, hierarchical models, pooling, Beta-Binomial, Dirichlet-Multinomial), 19 to 23 (inference and guides) and 41 (JAX and JIT). For the forecasting model: items 24 to 36 (decomposition through calibration and residual checks), 44 (alternatives) and again 19 to 23 and 41. Everything else supports these. If an interviewer picks a project and asks "tell me about it", you should be able to touch most of these items in order.
"I have studied all of these."
"For each core topic I can give a one-minute explanation, a number, and where it appears in my two projects. For example, on PELT: it finds changepoints by minimizing segment cost plus a penalty with pruning; I used it with a grid for candidate changepoints; its selection uncertainty is not propagated into the Bayesian posterior; and I would guard against it seeing the test period."
P0 = the 44 core topics of the syllabus. Tick = said aloud, no notes, about a minute, with a number and a project link.
Groups: foundations (1 to 6), testing and A/B (7 to 10), distributions (11), Bayesian (12 to 18), inference (19 to 23), forecasting (24 to 35), residual checks, Q-Q, KDE and Box-Cox (36 to 39), PCA (40), JAX (41), model checking and alternatives (42 to 44).
Trap: ticking on recognition.
Quick check: which P0 items would you touch first if an interviewer asks "walk me through your forecasting model"?
24 (decomposition), 26 and 27 (the trend and its changepoints), 28 (PELT), 29 (Laplace priors), 30 (Fourier terms), 31 (holidays and regressors), 33 (the likelihoods), 22 and 23 (SVI and the guide), 34 and 35 (validation and calibration), 36 (residual checks) and 32 (leakage). That is roughly the order of the chapters in this guide.
Recap, cheat sheet and practice
- One toolbox, two projects. Eight themes appear in both: likelihood choice, robustness, shrinkage, bias–variance, approximate inference, uncertainty, distribution diagnostics, scale. Use the concept × project matrix to find, for any topic, its role in each project and the chapter that teaches it.
- 1 · Likelihood: support first (yes/no, categories, counts, positive, real), variance second (constant, growing with the mean, heavy tails). Binomial $np(1-p)$, Poisson $\lambda$, Negative Binomial $\mu + \mu^2/\alpha$, Normal $\sigma^2$, Student-t $\sigma^2\nu/(\nu-2)$.
- 2 · Robustness: Student-t gives extreme residuals bounded influence; it down-weights, it does not delete. Explain known spikes with regressors.
- 3 · Shrinkage: pooling $w = \frac{n/\sigma^2}{n/\sigma^2 + 1/\tau^2}$ (toward the overall mean); Laplace MAP $\mathrm{sign}(z)\max(|z| - s^2/b, 0)$ (toward zero); only the MAP is exactly sparse.
- 4 · Bias–variance: $\mathrm{MSE} = \mathrm{Bias}^2 + \mathrm{Var}$; every flexibility knob (Fourier order, changepoints, $b$, $\tau$, rank) sits on one curve; choose on held-out time periods.
- 5 · Approximate inference: SVI maximizes the ELBO ($\log p(D) = \mathrm{ELBO} + KL$); the answer is only as good as the guide; NUTS is asymptotically exact, not exact.
- 6 · Uncertainty: $P(\theta_B \gt \theta_A \mid D)$ is about parameters (shrinks with data); $p(y_{\text{future}} \mid D)$ is about observations (has a noise floor).
- 7 · Diagnostics: histogram/KDE (shape), Q-Q (tails, skew), residual vs fitted (spread), ACF and Ljung–Box (time structure), posterior predictive checks (anything you care about).
- 8 · Scale: guide sizes $2d$, $d(r+2)$, $d + d(d+1)/2$; Cholesky work $\sim d^3$; JIT compiles once per shape; no boolean masks under JIT.
Cheat sheet
| Theme | Key rule | A/B framework | Forecasting model | Revise |
|---|---|---|---|---|
| 1 Likelihood | support, then variance function | Beta-Binomial, Dirichlet-Multinomial, Normal, Student-t, Poisson | Normal, Student-t, Negative Binomial | 4.7–4.11, 5.14, 6.3, 7.13 |
| 2 Robustness | influence $\frac{(\nu+1)r}{\nu + r^2}$, bounded | metrics with a few huge users | spike days; noise level | 4.9, 4.14, 7.13 |
| 3 Shrinkage | pull toward a centre by evidence ÷ prior spread | hierarchical partial pooling | Laplace on $\delta_j$ | 5.3, 6.5, 6.6, 7.10 |
| 4 Bias–variance | MSE = bias² + variance; choose on holdouts | pooling strength, prior strength | Fourier order, changepoints, $b$, guide | 5.1, 7.11, 7.15, 7.18 |
| 5 Inference | ELBO; guide family; validate | SVI; exact Beta check | custom SVI loop, early stopping | 6.9–6.15 |
| 6 Uncertainty | parameter vs observation uncertainty | $P(\theta_A \gt \theta_B \mid D)$, $P(\text{gain} \gt \delta)$ | predictive fan, coverage | 6.1, 6.4, 7.14, 7.16 |
| 7 Diagnostics | Q-Q, KDE, residuals, ACF, PPC | metric distributions, PPC | residual panel, PPC | 4.16, 4.17, 6.8, 7.17 |
| 8 Scale | $d$, rank $r$, static shapes | boolean mask lesson, global scaler | JIT loop, guide by size | 6.13, 6.16, 6.17, 7.19 |
import numpy as np
from scipy import stats
import jax, jax.numpy as jnp
import numpyro, numpyro.distributions as dist
from numpyro.infer import SVI, Trace_ELBO
from numpyro.infer.autoguide import AutoNormal, AutoMultivariateNormal, AutoLowRankMultivariateNormal
from numpyro.optim import Adam
# 1) Likelihood from support and variance: orders per day with mean 100 and variance 400
mu, var = 100.0, 400.0
alpha = mu ** 2 / (var - mu) # NB2: Var = mu + mu^2 / alpha
nb = stats.nbinom(n=alpha, p=alpha / (alpha + mu)) # the same NB in SciPy's (n, p) form
print("alpha", round(alpha, 1), "| mean", round(nb.mean(), 1), "var", round(nb.var(), 1), "| Poisson var would be", mu)
# alpha 33.3 | mean 100.0 var 400.0 | Poisson var would be 100.0
# 2) Robustness: the pull (influence) of a residual r under Normal and Student-t (nu = 4)
r = np.array([1, 2, 5, 10.0]); nu = 4
print("Normal pull", r, "| t4 pull", np.round((nu + 1) * r / (nu + r ** 2), 2))
# Normal pull [ 1. 2. 5. 10.] | t4 pull [1. 1.25 0.86 0.48]
print("P(|X| > 4): Normal", round(2 * stats.norm.sf(4), 6), "t4", round(2 * stats.t(4).sf(4), 4))
# P(|X| > 4): Normal 6.3e-05 t4 0.0161 (a 255-fold difference in how plausible a 4-scale-unit point is)
# 3) Shrinkage: partial pooling weights and the Laplace soft threshold
sigma, tau = 10.0, 2.0
for n in (4, 100, 400):
w = (n / sigma ** 2) / (n / sigma ** 2 + 1 / tau ** 2); print("n", n, "weight on own mean", round(w, 3), "pooled", round(w * 14 + (1 - w) * 10, 2))
# n 4 weight on own mean 0.138 pooled 10.55 | n 100 weight on own mean 0.8 pooled 13.2 | n 400 weight on own mean 0.941 pooled 13.76
z, s, b = np.array([0.02, 0.04, 0.30, -0.25]), 0.05, 0.1
print("Laplace MAP", np.sign(z) * np.maximum(np.abs(z) - s ** 2 / b, 0))
# Laplace MAP [ 0. 0.015 0.275 -0.225]
# 4) Bias-variance: MSE of c * xbar for mu = 2, sigma^2/n = 1
c = np.array([1.0, 0.8, 0.5]); print("MSE", np.round((1 - c) ** 2 * 4 + c ** 2 * 1, 2), "best c", 4 / (4 + 1))
# MSE [1. 0.8 1.25] best c 0.8
# 5) Approximate inference: SVI with a Gaussian guide against the exact Beta-Binomial posterior (50 of 500)
def model(k=None):
theta = numpyro.sample("theta", dist.Beta(1.0, 1.0))
numpyro.sample("k", dist.Binomial(500, theta), obs=k)
ex = stats.beta(51, 451)
def fit(seed, particles, lr, steps=3000):
guide = AutoNormal(model) # Gaussian on the logit scale, mean-field
svi = SVI(model, guide, Adam(lr), Trace_ELBO(num_particles=particles))
res = svi.run(jax.random.PRNGKey(seed), steps, k=50, progress_bar=False)
d = np.asarray(guide.sample_posterior(jax.random.PRNGKey(99), res.params, sample_shape=(20000,))["theta"])
return round(float(d.mean()), 4), round(float(d.std()), 4), round(float((d > 0.12).mean()), 4) # mean, sd, P(theta > 0.12)
print("noisy (1 particle, lr 0.05):", [fit(s, 1, 0.05) for s in range(4)])
print("steady (32 particles, lr 0.01):", [fit(s, 32, 0.01) for s in range(4)])
# noisy (1 particle, lr 0.05): [(0.106, 0.0139, 0.1583), (0.1038, 0.0143, 0.1304), (0.1025, 0.0128, 0.0909), (0.108, 0.0147, 0.1998)]
# steady (32 particles, lr 0.01): [(0.1014, 0.0132, 0.0846), (0.1016, 0.0137, 0.0935), (0.1014, 0.0137, 0.0922), (0.1014, 0.0136, 0.0907)]
print("exact:", round(ex.mean(), 4), round(ex.std(), 4), round(ex.sf(0.12), 4))
# exact: 0.1016 0.0135 0.0906 (a noisy single-draw ELBO makes the tail probability swing from seed to seed: 0.09 to 0.20)
# 6) Uncertainty: the gain posterior and one new user (sigma = 1, gain 0.05, n = 10000 per arm)
n, d = 10000, 0.05
print("P(B>A | D)", round(stats.norm.cdf(d / np.sqrt(2 / n)), 4), "| P(new B user > new A user)", round(stats.norm.cdf(d / np.sqrt(2 * (1 + 1 / n))), 4))
# P(B>A | D) 0.9998 | P(new B user > new A user) 0.5141
# 7) Diagnostics: posterior predictive check of the variance of 100 counts (observed mean 3, variance 9)
rng = np.random.default_rng(0)
pois = rng.poisson(3, (4000, 100)).var(axis=1, ddof=1); nbr = rng.negative_binomial(1.5, 1.5 / 4.5, (4000, 100)).var(axis=1, ddof=1)
print("replicated variance 95% range: Poisson", np.round(np.percentile(pois, [2.5, 97.5]), 1), "NB", np.round(np.percentile(nbr, [2.5, 97.5]), 1))
# replicated variance 95% range: Poisson [2.2 4. ] NB [ 5.5 14. ] -> the observed 9.0 is only plausible under the NB
# 8) Scalability: guide sizes, counted from NumPyro (a model with 40 free numbers)
def big():
numpyro.sample("w", dist.Normal(0, 1).expand([39]).to_event(1)); numpyro.sample("s", dist.HalfNormal(1))
for name, make in [("mean-field", AutoNormal), ("full-rank", AutoMultivariateNormal), ("low-rank r=5", lambda m: AutoLowRankMultivariateNormal(m, rank=5))]:
sv = SVI(big, make(big), Adam(0.01), Trace_ELBO()); p = sv.get_params(sv.init(jax.random.PRNGKey(0)))
print(name, "stored numbers:", sum(v.size for v in p.values()))
print("formulas for d = 40:", 2 * 40, 40 + 40 * 41 // 2, 40 * (5 + 2))
# mean-field stored numbers: 80 | full-rank stored numbers: 1640 (860 free: NumPyro stores the Cholesky factor as a 40 x 40 array) | low-rank r=5 stored numbers: 280
# formulas for d = 40: 80 860 280
1. Daily orders have mean 100 and variance 400. Which likelihood is the natural first choice, and why?
2. Which statement about the Student-t likelihood is correct?
3. A segment has 25 users, within-segment sd $\sigma = 10$, between-segment sd $\tau = 2$. How much weight does partial pooling give to the segment's own average?
4. A model has $d = 200$ free parameters. How many numbers do a rank-5 low-rank guide and a full-rank guide have?
5. After an A/B test, $P(\theta_B \gt \theta_A \mid D) = 0.99$. What does it mean?
6. Which diagnostic can reveal that your forecasting residuals are autocorrelated, while a Q-Q plot cannot?
Practice problems
A. Clicks per user have mean 2 and variance 7 in a large sample. Poisson or Negative Binomial? Give the moment estimate of the dispersion and check it.
- Variance ÷ mean $= 7/2 = 3.5$, far above 1, so a Poisson (ratio 1) is wrong.
- NB2: $\text{Var} = \mu + \mu^2/\alpha$, so $\alpha = \mu^2/(\text{Var} - \mu) = 4/5 = 0.8$.
- Check: $2 + 4/0.8 = 2 + 5 = 7$. ✓ A small $\alpha$ (0.8) means strong overdispersion.
- Next step: confirm with a posterior predictive check of the variance and the share of zeros.
B. (Shrinkage) A segment has $n = 36$ users, $\bar y = 20$, $\sigma = 12$, $\tau = 3$, overall mean 15. Separately, a changepoint slope change has $z = 0.12$ with $s = 0.06$ under a Laplace prior with $b = 0.05$. Compute both shrunk estimates.
- Pooling: $n/\sigma^2 = 36/144 = 0.25$, $1/\tau^2 = 1/9 = 0.111$, $w = 0.25/0.361 = 0.692$.
- Pooled estimate $= 15 + 0.692 \times (20 - 15) = 18.46$.
- Laplace threshold: $s^2/b = 0.0036/0.05 = 0.072$.
- MAP $= 0.12 - 0.072 = 0.048$ (positive and above the threshold, so it survives, reduced by 0.072).
C. (Bias–variance) $\mu = 3$ and $\sigma^2/n = 4$. Find the best shrinkage factor $c^*$ for $c\bar x$ and the MSE with and without it.
- $c^* = \mu^2/(\mu^2 + \sigma^2/n) = 9/13 = 0.692$.
- Without shrinkage ($c = 1$): $\mathrm{MSE} = 0 + 4 = 4$.
- With $c^*$: bias$^2 = (1 - 0.692)^2 \times 9 = 0.852$, variance $= 0.692^2 \times 4 = 1.917$, total $= 2.769$ ($= \mu^2 \sigma^2/n/(\mu^2 + \sigma^2/n) = 36/13$).
- A deliberate bias cuts the error by 31%.
D. (Uncertainty) 400 users per arm, user-level sd 1, a true gain of 0.1. Compute $P(\theta_B \gt \theta_A \mid D)$ and the chance a new B user beats a new A user.
- Posterior sd of the gain $= \sqrt{2/400} = 0.0707$; $z = 0.1/0.0707 = 1.414$; $P = \Phi(1.414) = 0.921$.
- New users: sd of the difference $=\sqrt{2(1 + 1/400)} = 1.416$; $z = 0.0706$; $P = \Phi(0.0706) = 0.528$.
- So "92% sure B's average is higher" and "52.8% of random pairs favour B" are both true: they answer different questions.
E. (Scale) A model has $d = 300$ parameters. Count the numbers in a mean-field, a rank-8 and a full-rank guide, and the memory of the full-rank factor in float32.
- Mean-field: $2d = 600$. Rank-8: $d(r+2) = 300 \times 10 = 3000$. Full-rank: $d + d(d+1)/2 = 300 + 45{,}150 = 45{,}450$.
- Full-rank is $45{,}450/3000 = 15.2$ times larger than rank-8.
- Memory of a $300 \times 300$ float32 factor: $4 \times 300^2 = 360{,}000$ bytes $= 0.36$ MB. Not the problem at this size; the cost explodes as $d$ grows ($d^2$ memory, $d^3$ work).
F. (Interview) "What ideas do your two projects share? You have two minutes."
"Eight ideas, and I can show each in both. One: choose the likelihood from the support and variance; the experiments use Binomial, Multinomial, Normal, Student-t and Poisson, the forecast Normal, Student-t and Negative Binomial. Two: robustness, where Student-t gives extreme values bounded influence, in both. Three: shrinkage, hierarchical partial pooling for segments and Laplace priors on the changepoint slopes, the same idea with a different centre. Four: bias–variance, with pooling strength, Fourier order, changepoint scale and guide rank as knobs, chosen on held-out data. Five: approximate inference, SVI in NumPyro and JAX in both, with my own loop and checkpointing in the forecast. Six: uncertainty: the experiments report $P(\theta_B \gt \theta_A \mid D)$, a statement about parameters; the forecast reports a predictive distribution with a noise floor. Seven: diagnostics, Q-Q plots, residual plots and posterior predictive checks. Eight: scale, JIT and static shapes, and the full-rank or low-rank guide chosen by model size. I would check the code before quoting any threshold."
Glossary
Every important word of this guide in one place, explained in plain English (230 terms). Type in the box to filter: it searches the terms and their explanations. The small numbers after each entry link to the section that teaches it (hover for its title).
All terms, A to Z
No term matches. Try a shorter word, or press / to search the whole guide.
- ACF/PACF signature
- AR($p$): ACF tails off (fades gradually), PACF cuts off after $p$ (falls inside the band and stays). MA($q$): ACF cuts off after $q$, PACF tails off. Random walk: ACF near 1 fading slowly, one big PACF spike. A signature is a first guess, not a proof. 7.3 7.6
- ACF plot (correlogram)
- The sticks $r_k$ drawn against the lag $k$: a memory profile. Fast decay means short memory, slow straight-line decay means trend or random walk, peaks at 7, 14, 21 mean a weekly pattern. 7.3
- Additive forecasting model
- $y_t=g(t)+s(t)+h(t)+X_t\beta+\epsilon_t$: trend, seasonality, holidays, regressors and noise, all in the units of $y$; everything except the noise is the mean $\mu_t$. Every explanation of your model should start from this line. 7.7
- Additive vs multiplicative
- Additive: constant seasonal swing, $y=T+S+R$. Multiplicative: swing proportional to the level, $y=T\times S\times R$. Taking logs turns the second into the first. 7.2 7.7
- ADF test
- Augmented Dickey–Fuller: $H_0$ is a unit root, so rejecting is evidence of stationarity. The statistic follows the Dickey–Fuller distribution, with 5% critical values near $-2.86$ (constant) and $-3.41$ (constant plus trend). 7.4
- AIC, BIC, AICc
- AIC $=-2\log L+2k$ and BIC $=-2\log L+k\log n$ trade fit against parameter count; lower is better, only between models on the same data. AICc corrects AIC for small samples. 7.6 7.5
- Aliasing
- At the sample times a fast wave is indistinguishable from a slower twin of frequency $|f-mf_s|$. The Nyquist frequency $f_s/2$ is the fastest wave the samples can show: on daily data weekly harmonic 4 duplicates harmonic 3, and a once-a-day cycle aliases to a constant. 7.11
- Amplitude and phase
- $a\cos\omega t+b\sin\omega t=R\cos(\omega t-\varphi)$ with $R=\sqrt{a^2+b^2}$ and $\varphi=\operatorname{atan2}(b,a)$. Two linear weights per harmonic encode a size and a timing, so the peak day can be learned while the model stays linear. 7.11
- ARIMA($p,d,q$)
- An ARMA($p,q$) model on the $d$-th difference, integrated back for forecasts. ARIMA(0,1,0) is naive, plus a constant it is drift, and ARIMA(0,1,1) with $\theta=\alpha-1$ is SES. 7.6
- ARMA($p,q$)
- An AR($p$) part plus an MA($q$) part. Both its ACF and PACF tail off; it is stationary when the AR part is and invertible when the MA part is. 7.6
- Assumptions of a Prophet-style model
- Additive components; a piecewise-linear trend with sparse bends in the allowed range whose last slope continues; a fixed seasonal shape; repeatable, known-in-advance events and regressors; independent noise with the chosen likelihood; a stable future; a close-enough posterior; one series at a time. 7.18
- Autocorrelated residuals
- Residuals that still carry day-to-day memory. A Prophet-style model with independent noise ignores it, so short-horizon forecasts lose accuracy and uncertainty is understated. An AR residual term $\epsilon_t=\varphi\epsilon_{t-1}+\eta_t$ (written with
lax.scan) adds a correction $\varphi^he_T$ that fades with the horizon. 7.3 7.17 - Autocorrelation (ACF value $r_k$)
- The sample autocovariance (one overall mean, divided by $n$) divided by the variance, $r_k=\hat\gamma_k/\hat\gamma_0$, always between $-1$ and $1$ with $r_0=1$. It measures the straight-line part of a series' link with its own past;
np.corrcoefof shifted pieces is not the ACF. 7.3 - Autoregressive model AR($p$)
- Today is a constant plus a weighted sum of the last $p$ values plus a shock. AR(1) is stationary exactly when $|\phi|\lt1$, with $\rho_k=\phi^k$ and mean $c/(1-\phi)$ (the constant $c$ is not the mean). 7.6
- Backshift operator $B$
- The operator with $By_t=y_{t-1}$, so the first difference is $(1-B)y_t$ and the seasonal difference is $(1-B^m)y_t$. Used in ARIMA notation. 7.4 7.6
- Bayesian latent changepoints
- Treating the number and locations of changepoints as parameters, so forecasts average over them. It is honest about location uncertainty but needs enumeration, reversible-jump MCMC or smoothing the step. 7.8
- Bias (mean error)
- $\bar e=\frac1n\sum e_i$, the systematic direction of the error. It can be zero for a poor forecaster, so never report it alone. 7.15
- Bias–variance decomposition
- Expected squared error $=\text{bias}^2+\text{variance}+\sigma^2$; more complexity (roughly the effective number of parameters) lowers bias and raises variance, which is about $\sigma^2p/n$ for least squares. You see the U-curve only on a time-ordered holdout. 7.18
- Box–Jenkins procedure
- Identify (choose $d$, then $p,q$ from ACF/PACF), estimate by maximum likelihood, check the residuals with the ACF and Ljung–Box, then forecast. 7.6
- Brier score
- $(p-\mathbf 1\{\text{event}\})^2$, a strictly proper score for the probability of a yes/no event such as “demand exceeds capacity”. 7.16
- Calibration (reliability)
- The stated probabilities match how often things happen: events called “20% likely” happen about 20% of the time and 90% intervals contain about 90% of outcomes. 7.16 7.16
- Capacity knobs vs strength knobs
- Capacity: Fourier order $N$ and the number of candidate changepoints. Strength of shrinkage: the Laplace scale $b$ and the pooling strength $\tau$. A generous candidate set with a small $b$ is safe; with a large $b$ it is not. 7.18
- Categorical regressor and dummy trap
- A variable with $K$ levels becomes $K-1$ dummy columns, each a difference from the reference level. Using all $K$ with an intercept is the dummy trap (singular); the reference choice changes coefficients, not fitted values. 7.12
- Changepoint
- A time $s_j$ at which the trend's slope is allowed to change. The stretch between two neighbouring changepoints is a segment. 7.8
- Changepoint matrix $A$
- The $n\times S$ 0/1 matrix $A_{ij}=1[t_i\ge s_j]$. The trend for all days is $g=(k+A\delta)\odot t+(m+A\gamma)$; building it with a comparison and a cast keeps shapes fixed under JIT. 7.8
changepoint_prior_scale- Prophet's Laplace scale $b$ for the slope changes, default 0.05, which only carries that meaning on Prophet's scaling ($y/\max|y|$, time in $[0,1]$); in real units $\delta_{\text{per day}}=\tilde\delta\max|y|/\text{days}$. Prophet's default fit is a MAP, not a posterior. 7.10 7.8
- Classical decomposition
- A descriptive split: centred moving-average trend, season by averaging the detrended values by position, then the remainder. It has missing ends, a fixed season and no forecast, so it is for exploration only. 7.2
- Concept × project matrix
- A table of topics against the two projects, marking each cell core, used, idea only or not used. Only facts from the project descriptions are marked core; anything else is “I would have to check the code”. 7.20
- Constrained-parameter transform
- A smooth one-to-one map from the free line to the allowed set: exp or softplus for positives, sigmoid for $(0,1)$, tanh for $(-1,1)$, softmax or stick-breaking for simplexes. NumPyro's
biject_to(support)picks it and adds the Jacobian. 7.19 - Continuity offset $\gamma_j=-s_j\delta_j$
- The intercept change that cancels the jump $s_j\delta_j$ a slope change would cause in a line measured from time zero. It comes from requiring equal values at $s_j$, and is not a free parameter. 7.8
- Coverage drift
- Misses in a window of $w$ days are Binomial$(w,\alpha)$ for an honest interval (mean $w\alpha$, sd $\sqrt{w\alpha(1-\alpha)}$). Live coverage drifting from nominal means the noise grew, a fixed-scale likelihood is now too tight, or the guide understates uncertainty. 7.19
- Cross-sectional data
- Many units observed once, like users in an A/B test. Its rows have no natural order, which is why the usual IID tools work there and not for a time series. 7.1
- CRPS
- $\int(F(x)-\mathbf 1\{x\ge y\})^2dx=E|X-y|-\frac12E|X-X'|$, in the units of $y$; for a point forecast it is the absolute error. It is computed straight from posterior predictive samples, is less tail-sensitive than the log score, and a skill score compares it with a baseline given an honest spread. 7.16
- Cycle
- Rises and falls of varying, unpredictable length, often multi-year and driven by the economy. A cycle is not seasonality and must not be modelled with a fixed-period term. 7.2
- Damped trend
- Holt with a factor $\phi\in(0,1)$ that shrinks the slope each step, so the forecast levels off at $\ell_T+\frac{\phi}{1-\phi}b_T$. Undamped trends run away at long horizons. 7.5
- Data-driven changepoint detection
- Choosing locations from the data by minimising the sum of segment costs plus a penalty $\beta M$ (PELT does this exactly). The locations are then treated as fixed, so their uncertainty is lost. 7.8 7.9
- Data (feature) drift vs concept drift
- Data drift: the distribution of the inputs $P(x)$ changed, visible by comparing input histograms. Concept drift: the relationship $P(y\mid x)$ changed, which shows only in the errors. Target drift: $y$ itself moves. 7.19
- Data leakage
- Any information that would not be available at the forecast origin is used for training, feature building or model selection. Backtests with leakage look better than live performance. 7.1 7.12 7.12
- Design matrix
- The $n\times p$ matrix $X=[1,t\mid\text{ramps}\mid\text{Fourier}\mid\text{flags}\mid\text{regressors}]$ with $\mu=X\theta$. Locations, periods, orders and windows are fixed in advance; only weights and noise parameters are learned. 7.7
- Difference-stationary
- A unit-root series whose changes are stationary: $y_t=\mu+y_{t-1}+u_t$. Difference it; shocks are permanent and forecast variance grows like $h\sigma^2$. 7.4
- Differencing
- Replacing the series by its changes, $\nabla y_t=y_t-y_{t-1}$, to remove a unit root or a linear trend. Undo it with a cumulative sum from the last level; each difference loses one observation. 7.4
- Direct vs recursive multi-step forecasting
- Direct: one model per horizon. Recursive: feed forecasts back as lags, so errors compound. A lag $y_{t-1}$ is unknown for days 2, 3, … of the forecast. 7.18
- Double-where trick
- For a masked log, apply
where(m, x, 1)before the log andwhere(m, ·, 0)after, so the gradient of the masked branch stays finite. 7.19 7.19 - Draws array (forecast paths)
- An $S\times H$ array of predictive draws: rows are whole simulated futures from one $\theta^{(s)}$, columns are days. Keep the whole array, because windows (totals, “any day above capacity”) need rows, not per-day quantiles. 7.14
- Effective sample size
- How many independent observations a correlated sample is worth. For AR(1) and large $n$ it is about $n(1-\phi)/(1+\phi)$, so a positively correlated series is worth fewer days than its row count. 7.1 7.3
- Eight connecting themes
- Likelihood choice, robustness, shrinkage, bias–variance, approximate inference, uncertainty, distribution diagnostics and scale: the ideas that appear in both the A/B framework and the forecasting model. 7.20
- ELBO and SVI (reprise)
- SVI maximises the ELBO $=E_q[\log p(D,\theta)]-E_q[\log q]$ with noisy gradients and an optimizer such as Adam; $\log p(D)=\text{ELBO}+KL(q\|p(\cdot\mid D))$. The answer is only as good as the guide; NUTS is asymptotically exact, not exact. 7.20
- ETS model
- Exponential smoothing written as a statistical (state-space) model with a noise term: Error, Trend, Season. For ETS(A,N,N), the model behind SES, the forecast variance is $\sigma^2[1+\alpha^2(h-1)]$; parameters come from maximum likelihood and intervals treat them as known. 7.5
- Event window
- Extra columns for the days before and after a holiday, $I(t-k\in D)$ for $k=L,\dots,U$, giving $U-L+1$ coefficients per holiday. Each offset is learned from only as many days as there are occurrences. 7.12
- Exchangeable
- A dataset is exchangeable when every reordering is equally plausible, so the order carries no information. IID data are exchangeable; a time series is not. 7.1
- Expanding vs sliding window
- Expanding trains on all history up to $T$ (more data, best for stable processes, old regimes stay in). Sliding trains on the last $W$ days (constant cost, forgets old regimes, noisier, $W$ must be chosen). 7.15
- Exploding gradients
- Gradient sizes that grow step to step (learning rate above the stable limit, badly scaled inputs, an exp link or a scale heading to 0). Remedies: smaller learning rate, scaled inputs, clipping (
ClippedAdam), gentler links, priors that keep scales off 0. 7.19 - Fair bake-off
- Same rolling origins, windows and horizons for every model, baselines included, everything refit inside each window, the metric fixed from the decision beforehand. Report by horizon, spread across origins and paired differences. 7.15
- Fitted values
- The model's values $\hat y_t$ inside the training period. They are not forecasts, because the model has already seen $y_t$ when it produced them. 7.1
- Fixed grid of candidates
- Choose many candidate locations in advance and fit all $\delta_j$ with a sparsity-inducing prior. Prophet puts 25 in the first 80% of the history, leaving runway to project the trend; a bend between two candidates is approximated by the nearest ones. 7.8
- Forecast horizon
- How many steps ahead we forecast. The $h$-step forecast is $\hat y_{T+h\mid T}$, “forecast of $y$ at $T+h$ made at $T$”; uncertainty usually grows with $h$. 7.1 7.14
- Forecast origin
- The last time step $T$ whose observation may be used for a forecast. Everything about a forecast is stated as “made at $T$”. 7.1
- Fourier columns
- The seasonal block with row $[\sin,\cos]$ per harmonic ($T\times2N$). The model is linear in the weights; over whole periods $X_s^\top X_s=\frac T2 I$. Period and order are fixed, only weights are learned. 7.11
- Fourier order $N$
- How many harmonics a seasonality uses. Training error always falls as $N$ grows, but holdout error is U-shaped: it is a bias–variance knob. With $P$ samples per period the exact ceiling is $\lfloor P/2\rfloor$ (weekly on daily data: 3, equal to the weekday dummies). 7.11 7.11
- Fourier series (order $N$)
- $s(t)=\sum_{n=1}^N[a_n\cos(2\pi nt/P)+b_n\sin(2\pi nt/P)]$, with $2N$ weights and no constant, so the level lives in the trend. Smooth shapes need few harmonics. 7.11
- Four-part answer
- The interview answer shape: a plain definition, a number or tiny picture, where it lives in your project, and the classic trap. Self-grade 2, 1 or 0 and revisit the 0s and 1s the next day. 7.20
- Four sources of forecast uncertainty
- Observation noise (a floor), parameters (the posterior, shrinks with data), model structure and exogenous regressors. The posterior predictive covers the first two; the last two must be added on purpose. 7.14
- Frequency (sampling interval)
- The time between consecutive steps of a regular series. The seasonal period is counted in steps, so it depends on the frequency (daily: 7 and 365.25; weekly: about 52.18; monthly: 12). 7.1
- Future availability of $X$
- A regressor is known in advance (calendar, planned prices and promotions, lags with $L\ge h$) or must be forecast (weather, competitors, traffic). A forecast regressor adds variance $\beta^2Var(X-\hat X)$ to the forecast. 7.12
- Future trend changes (Prophet)
- Prophet simulates future slope changes at the historical rate, with sizes drawn from a Laplace whose scale is the average fitted $|\delta_j|$. Whether your forecasts also do this is a design choice of your code. 7.8
- Gamma-Poisson mixture
- If a daily rate is Gamma-distributed (mean $\mu$, variance $\mu^2/\alpha$) and the count is Poisson given the rate, the count is Negative Binomial with variance $\mu+\mu^2/\alpha$: Poisson part plus rate part. 7.13 7.13
- Gaussian process (GP)
- A random smooth function $f\sim GP(m,k)$ whose kernel $k(t,t')$ is the assumption (smooth, rough or AR(1)-like, periodic), with hyperparameters chosen by the marginal likelihood. Cost is $O(n^3)$ time and $O(n^2)$ memory, and an RBF GP reverts to the mean beyond the data. 7.18
- Generative story (prior predictive simulation)
- Draw the weights from their priors, compute $\mu=X\theta$, draw $y$ from the likelihood. Run without data it shows what the priors jointly believe, before any fitting. 7.7
- Global model
- One model trained across many related series, sharing what it learns; the gain of boosting and neural forecasters comes mostly from sharing. A one-series-at-a-time Prophet-style model cannot do this (a hierarchical version could). 7.18 7.18
- Global scaler
- One standardization fitted on the history and applied to all groups (A/B) or to all periods (forecasting). Per-group scaling would erase the group differences you want to measure, and fitting it on the full series leaks the future. 7.12 7.19
- Gradient boosting (for forecasting)
- A sum of small trees fitted to the current residuals on engineered features: calendar, known drivers, lags at least as old as the horizon. Predictions are averages of training targets, so it cannot extrapolate a trend; intervals need quantile loss, conformal methods or ensembles. 7.18
- Guide family and size
- Mean-field has $2d$ learned numbers, low-rank $d(r+2)$ and full-rank $d+d(d+1)/2$, where $d$ is the number of unconstrained latent numbers (every weight, scale and hyperparameter). The size is a cost choice; a growing model can cross a threshold and change the guide, so read your own threshold from your code. 7.20 7.18
- HAC (Newey–West) standard errors
- Standard errors that stay valid when errors are autocorrelated and have non-constant spread. Alternatives are modelling AR errors (GLS) or a block bootstrap. 7.3
- Harmonic
- The $n$-th harmonic of period $P$ is the pair $\sin(2\pi nt/P)$, $\cos(2\pi nt/P)$, with period $P/n$. It is $n$ times faster than the fundamental, repeats every $P$ and averages to zero. 7.11
- Heteroscedasticity
- $Var(\epsilon_t)=\sigma_t^2$ changes with the level, the calendar or recent shocks; it biases the intervals, not the least-squares mean. Detect with a funnel or a spread ratio above about 1.5–2; respond with a transform (log or square root), a likelihood with mean-dependent variance (NB), or an explicit scale model. Student-t does not fix it. 7.17
- Hinge (ramp) form of the trend
- $g(t)=kt+m+\sum_j\delta_j(t-s_j)_+$ with $(u)_+=\max(u,0)$. It is the same function as Prophet's offset form and is linear in $(m,k,\delta)$. 7.8 7.8
- Holdout by time
- Choose an origin $T$ and horizon $H$, fit everything that learns from data on $y_1..y_T$ only, and score the hidden $y_{T+1..T+H}$ with error $e_{T+h\mid T}=y_{T+h}-\hat y_{T+h\mid T}$ (positive bias means forecasts too low). Use train, validation, then test blocks in time order, looking at the test block once. 7.15
- Holidays and events
- Effects tied to specific dates that do not sit on a fixed period: $h(t)=\sum_j\beta_jI(t\in D_j)$, one 0/1 indicator column per event, known in advance from a calendar. $\beta_j$ is the extra over a normal day with the same trend, season and regressors. 7.2 7.12
- Holiday shrinkage prior
- A prior $\beta_j\sim N(0,\tau^2)$ gives posterior mean $w\bar r_j$ with $w=\frac{n_j/\sigma^2}{n_j/\sigma^2+1/\tau^2}$; its MAP is ridge with $\lambda=\sigma^2/\tau^2$. Rare, noisy holidays are pulled toward zero. A hierarchical or Laplace prior are alternatives. 7.12
- Holt's linear trend
- SES plus a local slope $b_t$ that follows recent changes in the level; forecast $\ell_T+hb_T$. The slope is a reaction speed to recent change, not the slope of a line through all the data. 7.5
- Holt–Winters
- Holt's method plus one seasonal effect per position in the cycle, additive (in units) or multiplicative (factors around 1). It has one season length and needs about two full cycles to start. 7.5
- Horizon-specific error
- $\text{MAE}_h$ and $\text{RMSE}_h$ averaged over origins for each horizon $h$. For a random walk with naive forecasts $\text{RMSE}_h=\sigma\sqrt h$; a step that is a multiple of the season length ties horizon to weekday. 7.15
- Hurdle model
- A zero versus non-zero decision plus a zero-truncated count: $P(0)=\pi$, $P(k)=(1-\pi)f(k)/(1-f(0))$. Known structural zeros (closed, out of stock) belong in a regressor or a mask rather than in a gate; recent NumPyro releases include a hurdle class (for example
dist.HurdleNegativeBinomial2; check your version), otherwise you write it yourself. 7.13 - Identifiability (in the forecasting model)
- Whether the data can pin a quantity down. The sum $\mu_t$ is usually well determined, but two layers that can make the same shape trade off, so say “the total is well determined; the split is uncertain”. 7.7
- Influence (pull $\psi$) of a residual
- The derivative of $-\log p$: unbounded ($\psi=r$) for a Normal, and $(\nu+1)r/(\nu+r^2)$ for a t, which rises, peaks at $\sqrt\nu$ and falls back. Equivalently fitting a t is weighted least squares with weights $(\nu+1)/(\nu+r^2/\sigma^2)$, near 0 for a glitch day. 7.13
- Information set
- Everything genuinely known at the origin: past values, past regressors and future inputs known in advance (calendar, holidays, planned promotions). Tomorrow's weather is not in it. 7.1 7.12
- Initialization
- Starting values of the parameters and guide. NumPyro's default
init_to_uniform(radius=2)draws each free number in $(-2,2)$ andAutoNormalstarts every guide sd atinit_scale=0.1;init_to_medianandinit_to_valueare alternatives. 7.19 - Integrated of order $d$, I($d$)
- A series that needs $d$ differences to become stationary. A random walk is I(1); a stationary series is I(0). 7.4 7.3
- Interaction
- A product column $x_ax_b$ that makes the slope of $x_a$ equal $\beta_a+\beta_{ab}x_b$. Keep the main effects and centre continuous inputs. 7.12
- Interval coverage
- The share of outcomes inside a central $(1-\alpha)$ interval, $\frac1n\sum\mathbf 1\{l_t\le y_t\le u_t\}$. Below nominal means overconfident (too narrow), above means too cautious; misses that cluster signal a wrong model in specific situations. 7.16 7.16
- Interval score (Winkler)
- $(u-l)+\frac2\alpha(l-y)\mathbf 1\{y\lt l\}+\frac2\alpha(y-u)\mathbf 1\{y\gt u\}$: width plus a miss penalty; lower is better. It equals $\frac2\alpha$ times the sum of two pinball losses and is minimised by the honest interval. 7.16 7.16
- Jacobian term
- If $\theta=g(u)$ then $p_u(u)=p_\theta(g(u))|g'(u)|$ (for exp, $+u$ in the log). Samplers and SVI work in $u$-space, so the MAP moves with the coordinates while quantiles do not. 7.19
- JIT compilation (reprise)
- Python is traced once with abstract values, compiled by XLA and reused for the same shapes and dtypes; new shapes recompile. Boolean masks such as
x[mask]give data-dependent shapes, so usejnp.whereor masked sums. 7.20 7.8 - Kalman filter
- One pass over the data that gives the best estimate of the hidden state and its uncertainty at every time (predict, then update by a gain) and the likelihood of the data. Missing observations are skipped. It needs Gaussian noise. 7.18
- KPSS test
- Kwiatkowski–Phillips–Schmidt–Shin: $H_0$ is stationarity (around a level or a trend), the opposite of ADF; large $\eta$ rejects (5% critical values 0.463 and 0.146). Both tests have low power near $\varphi=1$ and are fooled by changepoints, so failing to reject proves nothing. 7.4
- L2 (mean-shift) cost
- $\sum(y_t-\bar y_{\text{seg}})^2$: Normal noise with one fixed $\sigma$ where only the mean may change. It sees level shifts, is very sensitive to outliers, and on a trend gives a staircase of fake cuts. 7.9 7.9
- Lag
- The value $k$ steps earlier, $y_{t-k}$, written as a new column with
y.shift(k). Its first $k$ entries are missing and there are $T-k$ usable pairs. 7.1 - Lagged regressor (distributed lag)
- A column $x_{t-L}$; several lags form a distributed lag, with strongly correlated adjacent columns. For horizon $h$ it is known at the origin when $L\ge h$. 7.12
- Lag-$k$ correlation
- The correlation of the lag pairs $(y_{t-k},y_t)$: how strongly a series is linearly linked to its own past. The ACF (Chapter 7.3) is the standard way to compute it for every $k$. 7.1 7.3
- Lag plot
- A scatter plot of $y_{t-k}$ (horizontal) against $y_t$ (vertical). Points hugging the diagonal mean the series follows its own past $k$ steps back. 7.1
- Laplace prior $\delta_j\sim\text{Laplace}(0,b)$
- A sharply peaked prior with exponential tails: $p(\delta)=\frac1{2b}e^{-|\delta|/b}$, $E|\delta|=b$, sd $\sqrt2\,b$. It says “most slope changes are tiny, a few are large”, and pulls with a constant force $1/b$ toward 0, where a Normal pulls in proportion to $\delta$. 7.10 7.10
- Laplace scale $b$
- The expected absolute slope change and the flexibility knob of the trend (not the sd, which is $\sqrt2\,b$). Small $b$ underfits, large $b$ overfits; refit at $b/3$ and $3b$ to test sensitivity. NumPyro
Laplace(loc, scale)and SciPylaplace(scale=b)take $b$; Stan calls itdouble_exponential. 7.10 7.10 - Lasso and soft-thresholding
- With Normal noise the MAP under Laplace priors is a lasso with $\lambda=\sigma^2/b$, solved by $S_\theta(u)=\text{sign}(u)\max(0,|u|-\theta)$ inside proximal-gradient loops (ISTA, or FISTA with momentum). Only corner-aware solvers give exact zeros. 7.10
- Level shift
- The level jumps once and stays. A continuous bend cannot do it; use a step column $1[t\ge s]$ or fix the data. 7.8
- Likelihood (the noise story)
- How each observation is spread around its expected value, $y_t\mid\mu_t,\varphi\sim p(y\mid\mu_t,\varphi)$. The layers say where the centre is; the likelihood says how far, how heavy the tails and whether the spread grows with $\mu$. Days are assumed independent given $\mu_t$. 7.13
- Linear cost (regression cost)
- The squared error of a straight-line fit inside each segment, so slope changes can be seen. Unlike Prophet's continuous trend, it lets the line jump at a changepoint. 7.9
- Linear trend
- $g(t)=kt+m$ with slope $k$ (change in $y$ per unit time) and intercept $m$ (value at $t=0$). $m$ depends on where time zero is and $k$ on the time units and any rescaling. 7.8
- Lineup protocol
- Show the real residual plot hidden among plots of residuals from data simulated from the fitted model, processed identically. If you can pick the real one, the model missed something; the numeric version is a posterior predictive check of a statistic. 7.17
- Link function and inverse link
- The function $f$ that turns the layer sum $\eta_t$ into a positive mean $\mu_t=f(\eta_t)$ for counts. It changes the units of every coefficient and prior; apply it per posterior draw before averaging. 7.13
- Ljung–Box test
- One test for all lags $1,\dots,h$ together: $Q=n(n+2)\sum r_k^2/(n-k)\sim\chi^2_h$ if there is no memory. A small p-value says there is memory somewhere in the first $h$ lags. 7.3 7.17
- Local level model
- A level that wanders as a random walk and is observed with noise. SES is its optimal forecast, and ARIMA(0,1,1) with $\theta=\alpha-1$ is the same thing. 7.5 7.6
- Local vs global model
- Exponential smoothing is local: its states are updated after every observation, so it adapts quickly. A Prophet-style model is global: one regression fitted to the whole history, which adapts to shifts only through changepoints. 7.5
- Log-likelihood $\ell$
- $\sum_t\log p(y_t\mid\mu_t,\varphi)$. Maximising it gives the maximum-likelihood fit, adding the log prior gives the MAP, and its exponential times the prior gives the posterior. 7.13
- Log link
- $\mu=e^\eta$: adding $b$ to $\eta$ multiplies the mean by $e^b$, so seasonal and holiday terms act as percentages and a line in $\eta$ is exponential. 7.13 7.7
- Log score
- $\log p_t(y_t)$, the log density (or mass) the forecast gave to what happened, in nats; higher is better. Held out on later days it is the log predictive density (ELPD summed, NLPD as a loss). Strictly proper but needs a density and is dominated by tail surprises; from draws use the log of the average likelihood. 7.16 7.13
- Log-sum-exp
- $\text{LSE}(a)=m+\log\sum_ke^{a_k-m}$ with $m=\max a$. The log of an average likelihood is $\text{LSE}-\log S$, never the average of the logs. 7.19
- MAE (mean absolute error)
- $\frac1n\sum|e_i|$, the typical miss, in the data's units. The best constant forecast under MAE is the median. 7.15 7.15
- MAPE
- $\frac{100}{n}\sum|e_i|/|y_i|$: undefined at zero, exploding near zero, unbounded for over-forecasts and favouring low forecasts. “100 minus MAPE” is not an accuracy. 7.15
- MASE
- Test MAE divided by $d_m$, the in-sample MAE of the (seasonal) naive forecast on the training data. Scale-free, zero-safe, symmetric; below 1 beats the in-sample naive error. Because $d_m$ is a one-step ruler, compare with the naive forecast's own MASE on the same long-horizon windows, and state $m$. 7.15
- Mean and drift forecasts
- The mean forecast $\bar y$ is optimal for a constant level plus noise and ignores trends. The drift forecast $y_T+h(y_T-y_1)/(T-1)$ adds the average one-step change, so an odd first or last value distorts it. 7.5
- Mean reversion
- In a stationary AR model, forecasts decay toward the mean ($\mu+\phi^h(y_T-\mu)$) and the interval width levels off. With $d=1$ there is no level to return to and the width grows like $\sqrt h$. 7.6
- Mean vs median forecast
- A point forecast should match its score: squared error rewards the predictive mean, absolute error the median. They differ for skewed distributions (counts, log link), where the mean sits above the median. 7.15
- Minimum segment length
min_sizeinruptures: no segment may be shorter, which stops a one-day spike becoming two fake changepoints but hides changes in the lastmin_size−1 days. Itsjump(default 5) rounds dates to multiples of 5; usejump=1when the exact day matters. 7.9- Model chooser
- A questionnaire (how many series, how much history, several seasonalities, events, counts, level shifts, momentum, explanation, intervals, interactions) that shortlists families. The points are heuristics; rolling-origin accuracy and calibration pick the winner. 7.18
- Moving-average forecast vs MA($q$) model
- A trailing moving-average forecast averages the last $k$ observed values into a flat forecast and lags a trend by $b(k+1)/2$. An MA($q$) model is something else: a mean plus a weighted sum of unobserved shocks; the two share a name and nothing more. 7.5 7.6
- Moving-average model MA($q$)
- A mean plus a weighted sum of the last $q$ unobserved shocks. It is always stationary and its ACF cuts off after lag $q$; for MA(1), $\rho_1=\theta/(1+\theta^2)$, never above 0.5 in size. 7.6
- MSE and RMSE
- MSE $=\frac1n\sum e_i^2$ in squared units; RMSE is its square root, in the data's units, and counts big misses extra. The best constant under RMSE is the mean. Always MAE $\le$ RMSE; a large ratio flags a few huge errors. 7.15 7.15
- Multicollinearity
- A regressor column that is nearly a combination of other columns (including trend, Fourier and holiday columns). Individual coefficients wobble and flip sign while sums and in-range predictions stay stable. 7.12
- Multiple seasonalities
- Fourier blocks concatenated side by side, one per period, with $2\sum_kN_k$ weights. Prophet's documented defaults: yearly 365.25 d order 10, weekly 7 d order 3, daily 1 d order 4, each switched on only when the data can support it. 7.11
- Multiplicative seasonality
- $y_t=g(t)(1+s(t))+\epsilon_t$: the seasonal effect is a fraction of the trend. A log link, or fitting $\log y$, makes every additive term multiply the mean. 7.7
- Naive forecast
- Tomorrow equals today: $\hat y_{T+h\mid T}=y_T$. It is the optimal squared-error forecast for a random walk and the same as SES with $\alpha=1$. 7.5
- NaN
- “Not a number” from $0/0$, $\infty-\infty$, $0\times\infty$, or $\log$ and $\sqrt{\ }$ of negatives. Once a parameter is NaN every later loss is NaN. Find the first non-finite step;
jax_debug_nansandsvi.stable_updatehelp but do not remove the cause. 7.19 - N-BEATS
- A neural forecaster of stacked blocks, each producing a backcast (the part of the window it explains) and a forecast; its interpretable variant uses trend and Fourier shapes, a cousin of $g(t)+s(t)$. TFT (attention) and DeepAR (autoregressive likelihood heads) are the other names to know; awareness only, and M4/M5 evidence is mixed. 7.18
- Negative Binomial (NB2)
- The mean–concentration form with $E[y]=\mu$ and $Var(y)=\mu+\mu^2/\alpha$; $\alpha\to\infty$ is Poisson and small $\alpha$ is noisy. NB1 (variance linear in the mean) is a different model. 7.13
- Negative Binomial parameterizations
- The same NB2 law has many argument lists: NumPyro
NegativeBinomial2(mean, concentration),NegativeBinomialProbs(total_count=α, probs=μ/(α+μ)); SciPynbinom(n=α, p=α/(α+μ)); statsmodels'alphais $1/\alpha$. “Dispersion” is ambiguous, so say which; mistakes are silent, so print the mean and variance. 7.13 - Newsvendor rule (critical ratio)
- With cost $c_u$ per unit short and $c_o$ per unit over, the best stock is the predictive quantile $F^{-1}(c_u/(c_u+c_o))$. Squared loss gives the mean, absolute loss the median; it needs a calibrated right tail. 7.14
- No exploitable structure
- Nothing known at forecast time (date, fitted value, other columns, previous residuals) predicts $e_t$ better than zero give or take $\hat\sigma$. Five questions: centred, constant spread, right shape, no pattern, no memory. 7.17
- No free lunch (bake-off)
- No method dominates all worlds; the winner is the model whose assumptions match the process. A fair bake-off uses the same data, origins and metric, baselines first, equal tuning effort, and judges calibration, cost and explainability too. 7.18
- Noise (remainder)
- What is left after the other layers: $\epsilon_t=y_t-(g+s+h+X\beta)$. Ideally mean zero, steady spread and independent; its distribution is the likelihood. 7.2 7.13
- Normal likelihood
- $y_t\sim N(\mu_t,\sigma^2)$: continuous, symmetric, light-tailed, constant variance, independent. Fitting is least squares; NumPyro's
Normal(loc, scale)takes the sd, not the variance. 7.13 - Normal (mean-and-variance) cost
- $L\log\hat\sigma^2_{\text{seg}}$, twice the negative maximised Normal log-likelihood without constants. It sees level and spread shifts and needs longer segments. 7.9
- One-standard-error rule
- A rule of thumb for choosing a knob: take the simplest setting whose holdout error is within one standard error of the best. 7.20
- Optimal partitioning (OP)
- Exact dynamic programming, $F(t)=\min_s[F(s)+C(y_{s:t})+\beta]$ with $F(0)=-\beta$, costing about $n^2/2$ segment costs. Exact for the chosen cost and penalty, not necessarily for the truth. 7.9
- Over-detection and under-detection
- Over-detection: reported changepoints with no real change (penalty too small, wrong cost, outliers, autocorrelation). Under-detection: missed changes (penalty too large, small changes, changes near the end). Plot the number found against $\beta$ and trust a long plateau. 7.9
- Over-differencing
- Differencing a series that was already stationary. It creates negative lag-1 autocorrelation ($-0.5$ for white noise, $-(1-\varphi)/2$ for an AR(1)) and adds noise. 7.4 7.4
- Overdispersion
- $Var(y_t\mid\mu_t)\gt\mu_t$, the usual state of daily counts because the underlying rate wobbles, so a Poisson (variance equals mean) gives intervals that are too narrow. A Normal fits counts badly too: negative values, symmetric shape, constant spread. 7.13
- Overlapping holidays
- Indicator columns that are on together are collinear: if always together, only $\beta_A+\beta_B$ is identified and the split comes from the priors. Days where only one is on identify each effect. 7.12
- P0 checklist
- The syllabus's 44 core topics: tick an item only when you have explained it aloud, without notes, with the idea, one formula or number and where it appears in a project. 7.20
- Partial autocorrelation (PACF)
- The correlation of $y_t$ and $y_{t-k}$ after removing the straight-line effect of the in-between lags; equivalently the last coefficient of an AR($k$) regression. For an AR($p$) it is zero after lag $p$. 7.3
- Partial pooling weight
- In the Normal–Normal case $\hat\theta_g=w_g\bar y_g+(1-w_g)\mu$ with $w_g=\frac{n_g/\sigma^2}{n_g/\sigma^2+1/\tau^2}$. $\tau\to0$ is complete pooling and $\tau\to\infty$ is no pooling. It is the A/B framework's version of shrinkage. 7.20 7.18
- Pearson dispersion index
- $\hat\phi=\frac1{n-p}\sum(y_t-\hat\mu_t)^2/\hat\mu_t$. About 1 means Poisson is fine; clearly above 1 means overdispersed. A variance-against-mean plot (Poisson on $y=x$, NB on $y=x+x^2/\alpha$) shows the same thing. 7.13
- Pearson residual
- $(y_t-\hat\mu_t)/\sqrt{V(\hat\mu_t)}$ with $V=\mu$ for Poisson and $\mu+\mu^2/\alpha$ for NB2. If the model is right its mean is about 0 and the average of $r_t^2$ (the Pearson dispersion) is about 1. Raw residuals of a correct count model fan out. 7.17
- PELT (Pruned Exact Linear Time)
- Optimal partitioning that drops every start $s$ with $F(s)+C(y_{s:t})\gt F(t)$ (it trails by more than one $\beta$ and can never win again, because splitting never raises the cost). Same answer, roughly linear time when changes are frequent, $O(n^2)$ worst case. 7.9
- Penalty $\beta$
- The price of one changepoint, in the cost's units (for L2, $y^2$); a cut is kept only if it saves more than $\beta$. A BIC-like rule of thumb is $2\hat\sigma^2\log n$ for L2, $3\log n$ for the Normal cost, with $\hat\sigma=1.4826\,\text{MAD}(\Delta y)/\sqrt2$; it assumes independent Normal noise. 7.9
- Period, frequency and angular frequency
- Period $P$ is the time for one cycle, frequency $f=1/P$ cycles per unit, angular frequency $\omega=2\pi/P$ radians per unit. The argument of sine and cosine is in radians, and $t$ and $P$ must share a unit. 7.11
- Pinball loss
- $\rho_\tau(y,q)=\tau(y-q)$ if $y\ge q$, else $(1-\tau)(q-y)$; minimised by the true $\tau$-quantile, and $\tau=0.5$ gives half the absolute error. Use predictive quantiles, not mean plus 1.28 sd, for skewed data. 7.16
- Pipeline leakage
- Any learned step (scaling, imputation, feature or order selection, tuning, PELT) fitted on data after the origin. Fit every step on the training window, freeze it, and refit everything at each rolling origin. 7.12 7.9
- PIT histogram
- Histogram of $u_t=F_t(y_t)$: flat means calibrated, a U overconfident, a hump underconfident, a tilt biased. Flat is necessary, not sufficient. For counts the CDF jumps, so use a randomized PIT. 7.16
- PIT (probability integral transform)
- $u_j=F_j(y_j)$, the share of predictive draws below the real value. If forecasts are calibrated the $u_j$ look Uniform(0,1); pile-ups near 0 and 1 mean too narrow, a hump in the middle too wide, all near one end biased. 7.14 7.16
- Pooling vs Laplace shrinkage
- Both are priors that pull noisy estimates toward a centre (the overall mean, or zero) by an amount set by the evidence against the prior scale. Pooling pulls proportionally; the Laplace MAP thresholds, and only the MAP is exactly sparse. 7.20 7.10
- Posterior decision vs posterior predictive
- $P(\theta_B\gt\theta_A\mid D)$ is about parameters and sharpens as users accumulate; $p(y_{\text{future}}\mid D)$ is about observations and has a noise floor. “99% sure B is better” is not “better for 99% of users”. 7.20 7.14
- Posterior is shrunk, not sparse
- A continuous posterior has $P(\delta_j=0\mid y)=0$: draws, means and medians are never exactly zero, only the MAP can be. Read the posterior slope over time, not a list of “active” changepoints. 7.10
- Posterior predictive check (PPC)
- Simulate replicated series from posterior draws on the real calendar and regressors and compare summaries $T(y)$ with $T(y^{\text{rep}})$ (mean, variance, max, zero days, weekday shape, lag-1 and lag-7 autocorrelation, level shift). A posterior predictive $p$-value near 0 or 1 flags misfit; it cannot validate a model. 7.14
- Prediction interval vs credible interval
- A prediction interval is about the future observation, including day-to-day noise. A credible interval for the parameter or the mean demand $\mu$ is narrower because it leaves the noise out; a 95% credible interval for the mean is not a 95% prediction interval for tomorrow. 7.14 7.16
- Predictive distribution (forecast)
- $p(\tilde y_{T+h}\mid D)=\int p(\tilde y\mid\theta)p(\theta\mid D)d\theta$, in practice sampled: draw $\theta^{(s)}$ from the posterior, then $\tilde y^{(s)}$ from the likelihood. A point forecast is one summary of it. 7.14
- Probabilistic forecast
- A full predictive distribution $F_t$ for each day (as samples, quantiles or interval pairs), judged on calibration, sharpness and proper scores. “RMSE improved” says nothing about its intervals. 7.16
- Proper scoring rule
- A rule whose expected score is best when you report your true belief (strictly proper: only then). Log score, CRPS and Brier are strictly proper; pinball and interval score are proper for quantiles; RMSE of one number, coverage alone, width alone and MAPE are not. 7.16
- PSI and KS distance
- PSI $=\sum(a_i-e_i)\ln(a_i/e_i)$ over bins fixed from a reference period (rules of thumb 0.1 and 0.25; the no-drift noise level is about $(k-1)(1/n_{ref}+1/n_{new})$). KS is the largest gap between two CDFs. 7.19
- Q-Q plot of residuals
- A straight line matches the reference; an S-shape means heavy tails (check the dates first, then try a Student-t), one bent end means skew, a kink means two mixed groups. Skewness and excess kurtosis are 0 for a Normal. For counts the Normal Q-Q is the wrong yardstick. 7.17
- Randomized quantile residual
- $r^q_t=\Phi^{-1}(u_t)$ with $u_t=F(y_t-1)+v_t[F(y_t)-F(y_t-1)]$ and $v_t\sim U(0,1)$: $N(0,1)$ for any right likelihood, the discrete analogue of the PIT. In a Bayesian model average $F$ over posterior draws. 7.17
- Random (shuffled) K-fold leakage
- Test days are random, so training contains days after the test day and the model interpolates. Blocked K-fold leaks less but still trains on later blocks; use forward chaining (
TimeSeriesSplit), train on $1..T_j$ and test after $T_j$, with a gap when features overlap. 7.15 7.1 - Random walk
- White noise added up: $y_t=y_{t-1}+\epsilon_t$, with variance $t\sigma^2$ and no level to return to. Adding a constant step gives a random walk with drift, $y_t=\mu+y_{t-1}+\epsilon_t$; the best forecast is the last value plus drift, with sd $\sigma\sqrt h$. 7.3
- Recurring vs one-off event
- A recurring event is 1 on every occurrence, past and future, with one shared coefficient (moving dates come from a calendar). A one-off is 1 only on its own days in the history, so it protects trend and seasonality and does not change the forecast. 7.12
- Regression with ARIMA errors (dynamic regression)
- The same regression mean as OLS but the errors $u_t$ follow an ARIMA model. The forecast is $\mathbf x_{T+h}^\top\hat\beta+\hat\phi^h\hat u_T$: the regression line plus a fading correction from the latest error. 7.6 7.6
- Regressors (exogenous variables)
- Measured outside drivers $X_t$ (price, discount, temperature) with coefficients $\beta$. They help only if their future values are known or forecast. 7.2 7.12
- Resampling
- Converting a series to another frequency. Downsample flows by sum, levels by mean, stocks by last value and rates by a ratio of sums; upsampling creates no new information. 7.1
- Residual
- $e_t=y_t-\hat y_t$: the true noise plus whatever the model got wrong or missed; the standardized version $z_t=e_t/\hat\sigma$ flags $|z_t|\gt3$. Say which $\hat y_t$ you used (in-sample fit, holdout forecast, posterior mean or draw); the goal is no exploitable structure. 7.17
- Residual ACF and Durbin–Watson
- The residual ACF with the $\pm1.96/\sqrt n$ band shows memory; $DW\approx2(1-r_1)$ sees lag 1 only. Choose the Ljung–Box lag $h$ to reach the season (14 for daily data with a weekly pattern, at most $n/5$). 7.17
- Residual drift monitor
- The rolling mean residual $\bar e_w$ is near 0 with standard error $\sigma_e/\sqrt w$ for a healthy model; limits at $\pm3\sigma_e/\sqrt w$ are a rule of thumb for independent residuals and need widening if they are autocorrelated. 7.19
- Residual panel
- Six plots (time, fitted, histogram with KDE, Q-Q, ACF, spread by fitted bin) read through a symptom-to-action table. A clean panel is not proof: accuracy and calibration on rolling origins decide. 7.17 7.20
- Residual vs fitted and vs time
- Bowl: too straight. Funnel: changing spread. Waves: a missed cycle. Step: a missed level or slope change. Spikes at certain dates: missed events. Add a bin-mean smoother and compare with simulated residuals. 7.17
- Rolling error monitor
- Store each forecast as issued, score it when its actual arrives, and track rolling MAE, WAPE or MASE against the launch backtest and a naive baseline. A rolling MASE above 1 means naive would have been better. 7.19
- Rolling-origin evaluation
- Also called time-series cross-validation, backtesting or walk-forward validation: for each origin $T_0, T_0+s,\dots$ refit on the window, forecast $h=1..H$, store the errors. The number of origins is $K=\lfloor(N-T_0-H)/s\rfloor+1$; origins overlap, so the errors are correlated. 7.15
- Run manifest
- A small text file (JSON or YAML) saved beside every fitted model: data hash and cut-off, preprocessing, code commit, prior scales, seed, inference settings, guide and optimizer, iterations actually run, library versions. Written automatically and diff-able, so a later reader can rebuild the run. 7.19
- Same seed is not same bits
- Floating-point addition is not associative, so a different chip, library version, thread count or precision moves the last digits, and an iterative optimizer can amplify them. Compare models by the gap divided by the seed-to-seed spread over 5–20 seeds; never report the best seed. 7.19
- SARIMA
- ARIMA with seasonal terms: $(P,D,Q)_m$ adds seasonal differences and AR/MA terms at lags $m,2m,\dots$ Seasonal naive is SARIMA(0,0,0)(0,1,0)$_m$ and it allows only one integer $m$. 7.6
- Seasonal differencing
- $\nabla_m y_t=y_t-y_{t-m}$ compares each value with the same day last season (7 for weekly, 12 for monthly). It is the $D$ in SARIMA; usually $D\le1$. 7.4 7.6
- Seasonal dummies
- One 0/1 column per position in the period ($P-1$ free numbers). Fourier terms are the smooth alternative and are cheaper for long periods such as 365.25. 7.2 7.6
- Seasonality
- A pattern with a fixed, known period $P$ tied to the calendar: $s(t+P)=s(t)$. Its effects are centred (sum to 0 if additive, average 1 if multiplicative). 7.2 7.11
- Seasonal naive forecast
- Copy the same position from the last full cycle (for daily data with a weekly cycle, $m=7$, the same weekday last week). It is the default bar to beat for seasonal data. 7.5 7.5
- Seasonal period
- The length $P$ of one repeat, counted in time steps (7 for weekly seasonality in daily data, 365.25 for yearly). Read it off the calendar and check by folding the series. 7.2 7.1
- Seed and PRNG key
- The starting value of a pseudo-random generator; the same seed gives the same sequence. NumPy uses
default_rng(seed); JAX has no hidden global generator and takes an explicitPRNGKeythat yousplit. 7.19 7.19 - Segmentation
- Cutting a series into calm pieces at positions $0\lt\tau_1\lt\dots\lt\tau_m\lt n$, scoring each piece with a cost that is small when one simple model fits (the cost never rises when you cut). In this guide $\tau_j$ is the first day of a new segment;
rupturesreturns segment ends, the last being $n$. 7.9 - Selection uncertainty
- The spread of $p(\tau\mid y)$, which a “detect, then fit” pipeline throws away: the posterior is $p(\theta\mid y,\hat\tau)$, not $p(\theta\mid y)$. Intervals near uncertain changepoints are too narrow, and because the data are used twice, selected changes look too big (winner's curse). 7.9
- Separability (no interactions)
- In an additive model a change in one layer moves $\mu_t$ by exactly that change. A Saturday effect cannot depend on the trend level unless you add that interaction as its own column or use a multiplicative form. 7.7
- Sharpness
- How concentrated (narrow) the forecasts are. It is a property of the forecasts alone; the principle is to maximise sharpness subject to calibration. 7.16 7.16
- Shrinkage
- The gap between what the data alone say and the Bayesian estimate. With one change the Laplace MAP is $\text{sign}(x)\max(0,|x|-se^2/b)$: exactly zero until the evidence beats $se^2/b$. 7.10
- Simple exponential smoothing (SES)
- A level that moves a fraction $\alpha$ toward each new value, $\ell_t=\ell_{t-1}+\alpha e_t$, so past weights $\alpha(1-\alpha)^j$ fade geometrically; the forecast is the last level, flat for every horizon. Large $\alpha$ means little smoothing; $\alpha=1$ is naive. 7.5
- Skill score
- $1-E_{\text{model}}/E_{\text{baseline}}$ on the same holdout: positive means better than the baseline. Report it over many origins, not one window. 7.5 7.15
- Slope change $\delta_j$
- The amount added to the slope at changepoint $s_j$, not the new slope: $\text{slope}(t)=k+\sum_{s_j\le t}\delta_j$. Positive speeds growth up, negative slows it, and the final slope $k+\sum_j\delta_j$ drives the forecast. 7.8 7.10
- sMAPE
- $\frac{100}{n}\sum2|e_i|/(|y_i|+|\hat y_i|)$, bounded between 0% and 200% but not symmetric, and several formulas share the name, so check your library. 7.15
- Softplus link
- $\mu=\log(1+e^\eta)$: about $\eta$ when $\eta$ is large (effects add in orders) and about $e^\eta$ when very negative. A prior $N(0,1)$ on a holiday coefficient means ×0.37 to ×2.7 under log but about ±1 order under softplus. 7.13
- Spurious regression
- Regressing a random walk on time (or on another random walk) often shows a “significant” slope although nothing real is there. Detrending a difference-stationary series leaves this trap. 7.4 7.4
- Standardizing a regressor
- $z=(x-\bar x_{\text{train}})/s_{\text{train}}$ with training-window statistics only, reused unchanged for test and future rows. Coefficients become “per standard deviation”; $\beta_{\text{raw}}=\beta_z/s$. 7.12
- State-space model
- A model where the level, slope or season is a hidden state updated by small random shocks and filtered with a Kalman filter; the local level model is exponential smoothing. Unlike an AR term, memory is permanent: the level keeps wherever it drifted. 7.17 7.18
- Stationarity (weak)
- Mean, finite variance and autocovariance $\gamma(k)$ do not depend on time $t$: the machine making the data does not change. Stationary does not mean flat. Rolling mean and sd staying roughly flat is the quick check; strict stationarity (every joint distribution shift-invariant) is stronger, and weak plus jointly Gaussian implies strict. 7.4 7.4
- STL
- Seasonal-trend decomposition using LOESS. It covers the ends, lets the season change slowly and has a robust option against outliers. 7.2
- Strong baseline
- The best of the simple methods for your series, scored on the same time-ordered holdout as your model. An accuracy number only means something next to one. 7.5
- Student-t does not remove outliers
- It assigns more probability to extreme residuals, so they exert less influence. The day stays in the likelihood and still shapes $\nu$ and the forecast tails; real events should be modelled, not down-weighted. 7.13 7.20
- Student-t likelihood
- A bell with a tail-heaviness dial $\nu$ (not $n-1$): $y_t\sim\text{StudentT}(\nu,\mu_t,\sigma)$. As $\nu\to\infty$ it is Normal, $\nu=1$ is Cauchy, and the sd is $\sigma\sqrt{\nu/(\nu-2)}$ only for $\nu\gt2$ ($\sigma$ is the scale). The data can rarely tell $\nu=30$ from $\nu=200$. 7.13
- Support and variance function
- Support is the set of values an observation can take; the variance function says how the variance depends on the mean. Choosing a likelihood means matching both, support first. 7.20 7.13
- Target leakage
- Features computed from the target or caused by it (same-day revenue when forecasting orders, returns logged later). The model appears to explain everything and fails live. 7.12
- Temporal (serial) dependence
- The distribution of $y_t$ changes once you know the past: $p(y_t\mid y_{t-1},\dots)\ne p(y_t)$. It makes forecasting possible and makes IID standard errors too small. 7.1 7.3
- Temporal split
- Train on $t\le T$ and test on $t\gt T$. A random split puts days from after each test day into training and measures interpolation, not forecasting. 7.1 7.15
- Time-feature regression
- A regression $y_t=\mathbf x_t^\top\beta+u_t$ on known-in-advance columns: $t$, hinge columns, seasonal dummies or Fourier pairs, holiday flags, known regressors. It is the skeleton of the Prophet-style model. 7.6 7.7
- Time series
- A sequence of observations of one quantity, each tied to a time and kept in time order: $y_1,\dots,y_T$. Its rows cannot be reordered without destroying information. 7.1
- Too many vs too few changepoints
- Too few (or too much shrinkage) underfits: stiff trend, biased final slope. Too many unshrunk overfits: a zig-zag trend and an unstable final slope. Training error always prefers more. 7.8 7.10
- Trend
- The slow, long-run level and direction of a series after seasonality, events and noise are removed. It dominates long horizons; polynomials should never be extrapolated. 7.2 7.8
- Trend-stationary
- $y_t=f(t)+u_t$ with a deterministic $f$ and stationary $u_t$. Detrend it; shocks fade and forecast variance stays bounded. This is the view behind your model's trend plus iid noise. 7.4 7.4
- Underfitting and overfitting
- Underfitting: high error on training and holdout data (bias dominates). Overfitting: very low training error but high holdout error (variance dominates). The generalization gap is holdout minus training error. 7.18
- Underflow and overflow
- Products of hundreds of probabilities underflow to 0 (float32 after about 33 factors of 0.04, float64 after 232); $e^x$ overflows above 709.8 in float64 and 88.7 in float32. Work with log probabilities. 7.19
- Unit root
- AR(1) with $\varphi=1$, as in a random walk: shocks never fade and the variance grows with time. The series needs one difference to become stationary. 7.3 7.4
- VIF (variance inflation factor)
- $VIF_j=1/(1-R_j^2)$, where $R_j^2$ regresses column $j$ on all the others; the standard error is multiplied by $\sqrt{VIF_j}$. For two regressors with correlation $r$, $1/(1-r^2)$ (0.9 gives 5.3). Above 5–10 is a rule of thumb for attention. 7.12
- Volatility clustering (ACF of $e^2$)
- Running the Ljung–Box test on squared residuals detects big misses following big misses even when $e_t$ itself has no memory. 7.17 7.17
- WAPE
- $100\sum|e_i|/\sum|y_i|$, equal to MAE divided by the mean actual. One division at the end, so zero days are fine and large days weigh more, which usually matches the business. 7.15
- Weak identification (short history)
- With less than one period of history a seasonal wave is nearly the same column as the trend (huge VIF, a ridge-shaped posterior, results driven by the prior). See the cycle at least once, ideally twice. 7.11
- White noise
- A series with mean 0, constant variance $\sigma^2$ and zero autocorrelation at every lag: no linear memory. It is weaker than iid (independence rules out all links, not only straight-line ones), which is weaker than Gaussian iid; good residuals look like it. 7.3
- White-noise band ($\pm1.96/\sqrt n$)
- Under “no memory”, $r_k\approx N(0,1/n)$, so about 95% of sticks fall inside $\pm1.96/\sqrt n$ and roughly 1 lag in 20 crosses by chance. A single crossing is weak evidence. 7.3
- Why intervals widen with the horizon
- Variance at horizon $h$ is noise (flat) plus level and slope doubt ($s_a^2+h^2s_b^2$, about linear) plus future trend changes (sd $\propto h^{3/2}$). Random-walk noise grows like $\sqrt h$ and AR(1) noise levels off. 7.14 7.6
- Zero-inflated model
- A gate probability $\pi$ of a structural zero on top of a base count distribution: $P(0)=\pi+(1-\pi)f(0)$ and $P(k)=(1-\pi)f(k)$ for $k\ge1$. 7.13
Formula and code sheet
The key formulas of every chapter on one page, each with what it means in words and the trap that usually goes with it. Then the special tables: the one-page map of the forecasting model (each component with its parameters, priors and chapter), the metrics, the three forecast likelihoods with their library parameterizations, a diagnostics-to-fix table, a leakage checklist, Prophet's defaults as the chapters state them, and a short code sheet. Notation: $y_t$ the series, $\hat y_{T+h\mid T}$ a forecast made at origin $T$ for horizon $h$, $g,s,h,X\beta,\epsilon$ the five layers, $\delta_j$ the slope changes at changepoints $s_j$, $b$ the Laplace scale, $N(\mu,\sigma^2)$ written with the variance (NumPyro and SciPy take the sd). Every threshold is a rule of thumb unless a chapter proves it.
7.1–7.6 · Time-series basics, autocorrelation, stationarity, baselines and the ARIMA family
7.1 · What makes time-series data different
| Formula | In words | Main trap | See |
|---|---|---|---|
$y_1,\dots,y_T$ in time order; lag $y_{t-k}$ (y.shift(k)), $T-k$ lag pairs; lag-$k$ correlation = correlation of the pairs | A series is values plus an ordered time index. Shuffling keeps the histogram, mean and sd and destroys trend, season and memory. | shift(-k) is a lead, the future, so leakage. Missing is not zero. Keep dates in the index. | 7.1 7.1 7.1 |
| AR(1): $y_t-\mu=\phi(y_{t-1}-\mu)+e_t$; $Var(y_t)=\sigma^2/(1-\phi^2)$; lag-$k$ correlation $\phi^k$ | The simplest memory. Knowing yesterday narrows today from $\sigma/\sqrt{1-\phi^2}$ to $\sigma$. | $\phi$ near 1 is long memory; $\phi=1$ is a random walk, not a strong AR. | 7.1 |
| $SE(\bar y)\approx\dfrac{\sigma_y}{\sqrt n}\sqrt{\dfrac{1+\phi}{1-\phi}}$, $\;n_{\text{eff}}\approx n\dfrac{1-\phi}{1+\phi}$ | Correlated days are worth fewer than their row count: $\phi=0.5$ gives a standard error 1.7 times the IID one. | The IID formula $\sigma/\sqrt n$ is too small whenever $\phi\gt0$. | 7.1 7.3 |
| Resample: flows by sum, levels by mean, stocks by last, rates by ratio of sums; period in steps: daily 7 and 365.25, weekly 52.18, monthly 12, hourly 24 and 168 | The right combination depends on the kind of quantity; the seasonal period is counted in steps. | pandas "W" means weeks ending Sunday; use "ME", not "M"; upsampling creates no information; mean of daily rates is not the rate. | 7.1 |
| $\hat y_{T+h\mid T}$ from the information set $\mathcal I_T$; AR(1) forecast sd $\sigma\sqrt{\sum_{j=0}^{h-1}\phi^{2j}}$; random walk $\sigma\sqrt h$ | Forecast made at origin $T$ for horizon $h$. Uncertainty grows with $h$. | Fitted values are not forecasts. Horizon is not the training window. | 7.1 |
| Honest split: train $t\le T$, test $t\gt T$, every learned step fitted inside the training window | Forecasting is extrapolation; a random split tests interpolation. | Scalers, feature selection, PELT and tuning on the full series are leakage. | 7.1 |
7.2 · Components of a time series
| Formula | In words | Main trap | See |
|---|---|---|---|
| $y_t=g(t)+s(t)+h(t)+X_t\beta+\epsilon_t$ | Trend, seasonality, holidays/events, regressors and noise. Only $y_t$ is observed. | Components are estimated, not observed; the split is a modelling choice. | 7.2 |
| $s(t+P)=s(t)$, centred: $\sum_{j=1}^{P}s(j)=0$ (additive) or average 1 (multiplicative) | A fixed, known calendar period ($P=7$, $365.25$, $12$, $24$). Find $P$ from the calendar and check by folding. | Cycles have varying length and are not seasonality. Uncentred effects make level and season unidentifiable. | 7.2 7.2 |
| $h(t)=\sum_j\kappa_j\mathbf 1[t\in D_j]$; $X_t\beta=\sum_jx_{t,j}\beta_j$ | One 0/1 column per event; $\beta_j$ is the change in $y$ per unit of regressor $j$ with the rest fixed. | Regressors need future values; association is not causation; correlated regressors share credit unstably. | 7.2 7.2 |
| Additive $y=T+S+R$ vs multiplicative $y=T\times S\times R$; $\log y$ additive $\Leftrightarrow$ multiplicative on the original scale | Does the seasonal swing (or noise spread) grow with the level? Then use logs, multiplicative mode or a log link. | Log needs $y\gt0$; back-transformed log forecasts estimate medians. | 7.2 |
| Classical: $\hat T$ = centred moving average (window $m$, $2\times m$ if even) $\to$ $d=y-\hat T$ $\to$ average $d$ by position and centre $\to$ $\hat S$ $\to$ $\hat R=y-\hat T-\hat S$ | A descriptive split for exploration, run on the training window only. STL is the robust version. | No trend at the ends, a fixed season, outliers leak, no forecast and no uncertainty. | 7.2 |
7.3 · Autocorrelation: ACF, PACF, white noise, random walks
| Formula | In words | Main trap | See |
|---|---|---|---|
| $r_k=\dfrac{\sum_{t=k+1}^{n}(y_t-\bar y)(y_{t-k}-\bar y)}{\sum_{t=1}^{n}(y_t-\bar y)^2}$, $r_0=1$ | One overall mean, divided by the full spread (statsmodels' convention). Positive means runs, negative means zigzag. | np.corrcoef of shifted pieces and .autocorr() are not the ACF. The raw-series ACF mixes trend, season and memory; read the residual ACF. | 7.3 7.3 |
| No memory: $r_k\approx N(0,1/n)$, band $\pm1.96/\sqrt n$; Ljung–Box $Q=n(n+2)\sum_{k=1}^{h}\dfrac{r_k^2}{n-k}\sim\chi^2_h$ | About 1 lag in 20 crosses the band by chance. Ljung–Box tests many lags at once; a small p means memory somewhere. | plot_acf draws widening Bartlett bands. A large p is not proof of no memory. | 7.3 |
| PACF: $\alpha(1)=\rho_1$, $\alpha(2)=\dfrac{\rho_2-\rho_1^2}{1-\rho_1^2}$; AR($p$): PACF $=0$ after lag $p$ | The direct link after removing the in-between lags; the last coefficient of an AR($k$) regression. | statsmodels' pacf default differs from plot_pacf (ywm). | 7.3 |
| AR($p$): ACF fades, PACF cuts at $p$. MA($q$): ACF cuts at $q$, PACF fades. Random walk: ACF near 1, one PACF spike. MA(1): $\rho_1=\theta/(1+\theta^2)$, $|\rho_1|\le0.5$ | Signatures give candidate models. Peaks at 7, 14, 21 mean a weekly pattern. | A signature is a first guess, not a proof; read it on a stationary series. | 7.3 |
| White noise $WN(0,\sigma^2)$; random walk $y_t=\mu+y_{t-1}+\epsilon_t$: $Var(y_t)=t\sigma^2$, $Corr(y_t,y_{t+k})=\sqrt{t/(t+k)}$ | White noise has nothing left to learn. A random walk is white noise added up (AR(1) with $\varphi=1$, a unit root); one difference gives white noise back. | White is weaker than iid. “It will come back” is wrong for a random walk; two wandering series correlate spuriously. | 7.3 7.3 |
| $Var(\bar y)=\dfrac{\sigma^2}{n}\Big[1+2\sum_{k=1}^{n-1}\big(1-\tfrac kn\big)\rho_k\Big]$; fixes: HAC (Newey–West), AR errors, block bootstrap | Positive autocorrelation makes naive standard errors, intervals and posteriors too narrow. | Count information, not rows. A likelihood that multiplies $n$ terms believes it has $n$ independent days. | 7.3 7.3 |
7.4 · Stationarity and differencing
| Formula | In words | Main trap | See |
|---|---|---|---|
| Weak: $E[y_t]=\mu$, $Var(y_t)=\gamma(0)\lt\infty$, $Cov(y_t,y_{t+k})=\gamma(k)$ for all $t$ | The machine making the data does not change. Check rolling mean and sd. Strict + finite variance gives weak; weak + Gaussian gives strict. | Stationary is not flat. Persistent AR(1), pseudo-cycles and ARCH bursts are stationary; level shifts and weekday means are not. | 7.4 7.4 |
| $\nabla y_t=y_t-y_{t-1}$; $\nabla^2y_t=y_t-2y_{t-1}+y_{t-2}$; $\nabla_my_t=y_t-y_{t-m}$; undo with a cumulative sum | Differencing removes a unit root or a linear trend; seasonal differencing removes a repeating pattern. Each loses observations. | Over-differencing: $r_1\approx-0.5$ for white noise, $-(1-\varphi)/2$ for AR(1). A fixed pattern + $\nabla_m$ leaves a negative spike at lag $m$. | 7.4 7.4 |
| Trend-stationary: $y_t=f(t)+u_t$ (detrend). Difference-stationary: $y_t=\mu+y_{t-1}+u_t$ (difference) | Shocks fade vs shocks persist; forecast variance bounded vs $h\sigma^2$. | Detrending a random walk gives a spurious trend; differencing a trend-stationary series over-differences. | 7.4 |
| ADF: $\Delta y_t=\alpha(+\beta t)+\gamma y_{t-1}+\sum\delta_i\Delta y_{t-i}+e_t$, $H_0$: $\gamma=0$; 5%: $-2.86$ (c), $-3.41$ (ct). KPSS: $H_0$ stationary; 5%: 0.463 (c), 0.146 (ct) | Opposite null hypotheses: use both. ADF rejects + KPSS does not: stationary. The reverse: unit root. | Low power near $\varphi=1$; changepoints mimic unit roots; failing to reject is not accepting. | 7.4 |
| ARIMA: make $\nabla^dy_t$ stationary, model it, integrate back. Prophet-style: $y_t=g+s+h+X\beta+\epsilon_t$ with stationary iid residuals | ARIMA removes non-stationarity; a Prophet-style model describes it with explicit components. Stationarity still checks the residuals. | “Prophet needs stationary data” and “stationarity does not matter” are both wrong. | 7.4 |
7.5 · Baselines and exponential smoothing
| Formula | In words | Main trap | See |
|---|---|---|---|
| Naive $\hat y_{T+h\mid T}=y_T$; seasonal naive $y_{T+h-m(k+1)}$, $k=\lfloor(h-1)/m\rfloor$; mean $\bar y$; drift $y_T+h\frac{y_T-y_1}{T-1}$ | No parameters. Each is optimal for a simple process: random walk, seasonal random walk, constant level + noise, random walk with drift. | An odd first or last value distorts drift. Compare against the right baseline: seasonal naive for seasonal data. | 7.5 7.5 7.5 |
| Moving average $\frac1k\sum_{j=0}^{k-1}y_{T-j}$ (flat); lag on a line of slope $b$: $b(k+1)/2$ | $k=1$ is naive and $k=T$ is the mean. A trailing rule, not a model. | A rolling mean of $y$ is not the MA($q$) model of shocks. | 7.5 |
| Skill $=1-E_{\text{model}}/E_{\text{baseline}}$ on the same holdout | Positive beats the baseline. Report it over many origins. | A weak baseline or a single window proves nothing. | 7.5 |
| SES: $\ell_t=\alpha y_t+(1-\alpha)\ell_{t-1}=\ell_{t-1}+\alpha e_t$; forecast $\ell_T$; weights $\alpha(1-\alpha)^j$; pandas span: $\alpha=2/(\text{span}+1)$ | Fading memory: a flat forecast at the last level. Choose $\alpha$ by one-step SSE. | Large $\alpha$ is less smoothing. SES has no trend and no season. | 7.5 |
| Holt: $\ell_t=\alpha y_t+(1-\alpha)(\ell_{t-1}+b_{t-1})$, $b_t=\beta(\ell_t-\ell_{t-1})+(1-\beta)b_{t-1}$, forecast $\ell_T+hb_T$; damped: $\ell_T+(\phi+\dots+\phi^h)b_T\to\ell_T+\frac{\phi}{1-\phi}b_T$ | A local level and slope that follow recent change; damping levels off long horizons. | Undamped trends run away; $\beta$ is a reaction speed, not a slope; notation for $\beta$ and $\phi$ differs between sources. | 7.5 |
| Holt–Winters (additive): $s_t=\gamma(y_t-\ell_{t-1}-b_{t-1})+(1-\gamma)s_{t-m}$, forecast $\ell_T+hb_T+s$ (multiplicative: $(\ell_T+hb_T)\,s$) | Adds one seasonal effect per position of the cycle. | One season length only; needs about two cycles of history; multiplicative only if the swing scales with the level. | 7.5 |
| ETS(A,N,N): $y_t=\ell_{t-1}+\varepsilon_t$, $\ell_t=\ell_{t-1}+\alpha\varepsilon_t$; forecast variance $\sigma^2[1+\alpha^2(h-1)]$ | Smoothing as a statistical model: likelihood, AIC, intervals that widen with $h$. | Intervals take fitted parameters as known: check coverage. ETS is local; a Prophet-style model is global. | 7.5 7.5 |
7.6 · The ARIMA family and regression with time features
| Formula | In words | Main trap | See |
|---|---|---|---|
| AR(1): $y_t=c+\phi y_{t-1}+\varepsilon_t$, stationary iff $|\phi|\lt1$, $\mu=c/(1-\phi)$, $\rho_k=\phi^k$, forecast $\mu+\phi^h(y_T-\mu)$ | Today leans on yesterday; forecasts revert to the mean. | $c$ is not the mean; $\phi=1$ is a random walk. | 7.6 |
| AR(2) stationary: $\phi_1+\phi_2\lt1$, $\phi_2-\phi_1\lt1$, $|\phi_2|\lt1$; pseudo-cycles if $\phi_1^2+4\phi_2\lt0$ | A triangle of allowed coefficients; complex roots give damped waves. | Pseudo-cycles come from memory, not the calendar. | 7.6 |
| MA($q$): $y_t=\mu+\varepsilon_t+\sum_{j=1}^{q}\theta_j\varepsilon_{t-j}$; $Var=\sigma^2(1+\sum\theta_j^2)$; ACF cuts off after $q$ | Always stationary; forecasts equal the mean after $q$ steps; invertible if $|\theta|\lt1$ (MA(1)). | Shocks, not observations. Sign conventions differ (statsmodels $+\theta$). $\theta\approx-1$ means over-differenced. | 7.6 |
| ARIMA: $\phi(B)(1-B)^dy_t=c+\theta(B)\varepsilon_t$; (0,1,0) naive; (0,1,0)+$c$ drift; (0,1,1) is SES with $\theta=\alpha-1$ | ARMA on the $d$-th difference, integrated back. $d$ is usually 0 or 1. | ARIMA with $d\ge1$ is itself non-stationary. In statsmodels a constant cannot be combined with $d\ge1$: use trend="t". | 7.6 |
| SARIMA$(p,d,q)(P,D,Q)_m$: $\Phi(B^m)\phi(B)(1-B)^d(1-B^m)^Dy_t=c+\Theta(B^m)\theta(B)\varepsilon_t$; seasonal naive = (0,0,0)(0,1,0)$_m$ | Seasonal differences and seasonal AR/MA terms at lags $m,2m,\dots$ | One integer $m$ only; $m=365$ is slow and fragile; a fixed pattern + $D=1$ gives seasonal MA near $-1$: use dummies or Fourier. | 7.6 |
| $h$-step error variance $\sigma^2\sum_{j=0}^{h-1}\psi_j^2$; AR(1): $\sigma^2\frac{1-\phi^{2h}}{1-\phi^2}\to\frac{\sigma^2}{1-\phi^2}$; random walk $h\sigma^2$ | Stationary: the interval width levels off. $d=1$: width grows like $\sqrt h$. | These intervals include future shocks only, not parameter or model uncertainty. | 7.6 |
| AIC $=-2\log L+2k$, BIC $=-2\log L+k\log n$; Ljung–Box on residuals with $h-(p+q)$ degrees of freedom | Box–Jenkins: identify, estimate by maximum likelihood, check residuals, forecast. | Compare criteria only for the same $d$ and $D$. The final judge is rolling-origin error. | 7.6 |
| Dynamic regression: $y_t=\mathbf x_t^\top\beta+u_t$, $u_t$ ARIMA; forecast $\mathbf x_{T+h}^\top\hat\beta+\hat\phi^h\hat u_T$ | The Prophet-style mean plus a fading correction from the latest error. | Dummy trap: $m-1$ columns. SARIMAX default order is (1,0,0). OLS standard errors are too small with autocorrelated errors. | 7.6 7.6 |
7.7–7.10 · The additive model, piecewise-linear trends, PELT and Laplace priors
7.7 · The additive forecasting model
| Formula | In words | Main trap | See |
|---|---|---|---|
| $y_t=g(t)+s(t)+h(t)+X_t\beta+\epsilon_t$; mean $\mu_t=g+s+h+X\beta$; $y_t\sim\text{Lik}(\mu_t,\cdot)$ | Five layers, all in the units of $y$. Seasonality is a zero-average deviation, the level lives in the trend, noise has mean 0. | $\epsilon_t$ (unobservable) is not the residual $e_t=y_t-\hat\mu_t$. “Additive” is about how layers combine. | 7.7 |
| $\hat\mu_t=\hat g+\hat s+\hat h+X_t\hat\beta$: a change in one layer moves $\mu_t$ by exactly that change | No interactions, so what-ifs and component plots come free. The sum is usually well determined; the split needs distinct columns. | A regressor that rises with time competes with the trend: check the posterior correlation before interpreting either. | 7.7 |
| $\mu=X\theta$, $X=[1,t\mid\text{ramps}\mid\text{Fourier}\mid\text{flags}\mid\text{regressors}]$; columns $2+S+2N_{\text{year}}+2N_{\text{week}}+\#\text{hol}+R$ | One regression with structured columns and a prior per block. Locations, periods, orders and windows are fixed in advance; only weights and noise parameters are learned. | The column count (plus noise parameters) is the latent dimension $d$ behind the guide choice. | 7.7 |
| Prior predictive: draw $\theta$ from the priors, $\mu=X\theta$, draw $y$. Laplace: $P(|\delta|\gt c)=e^{-c/b}$, $Var=2b^2$; $S$ changes give final-slope sd $b\sqrt{2S}$ | The model is a recipe for fake data; run without data it shows what the priors believe. | Many “harmless” wide priors together make absurd worlds. NumPyro Normal takes the sd, Laplace takes $b$. | 7.7 |
| Additive $y=g+s$; multiplicative $y=g(1+s)$; log link $\mu=e^{\eta}$: weight $w$ multiplies $\mu$ by $e^w$ (0.3 gives ×1.35) | A log link turns additive terms in $\eta$ into multiplicative effects on the mean. | Decide from the plot and residuals: a growing wave means the wrong form. $\exp$ of a log-scale mean is a median. | 7.7 |
| $\hat y_{T+h\mid T}=\hat g(T+h)+\hat s(T+h)+\hat h(T+h)+X_{T+h}\hat\beta$ | Continue every layer: last slope, repeated seasons, future calendar, regressors supplied from outside. | The model does not forecast regressors and does not use yesterday's value; the noise band is not the whole uncertainty. | 7.7 |
7.8 · Piecewise-linear trends and changepoints
| Formula | In words | Main trap | See |
|---|---|---|---|
| $g(t)=kt+m$; slope after changepoints $k+\sum_{s_j\le t}\delta_j$; final slope $k+\sum_j\delta_j$ | $k$ is in units of $y$ per unit time, $m$ the value at $t=0$. $\delta_j$ is a change, not the new slope. | $m$ depends on where time zero is; $k$ on the time units and any rescaling. Changepoints change speed, not level. | 7.8 7.8 |
| Continuity: $\delta_js_j+\gamma_j=0\Rightarrow\gamma_j=-s_j\delta_j$; hinge form $g(t)=kt+m+\sum_j\delta_j(t-s_j)_+$ | Without the offset the line would jump by $s_j\delta_j$. The offset form and the hinge form are the same function. | $\gamma_j$ is not a free parameter and not there “to fit better”. | 7.8 |
| $A_{ij}=\mathbf 1[t_i\ge s_j]$, $g=(k+A\delta)\odot t+(m+A\gamma)=kt+m+R\delta$, $R_{ij}=(t_i-s_j)_+$ | The whole trend in one matrix formula; linear in $(m,k,\delta)$. | Build $A$ with a comparison and a cast, not boolean masks under JIT; the same locations must be used for future days. | 7.8 |
| Forecast trend $\hat g(T)+(\hat k+\sum_j\hat\delta_j)(t-T)$; slope uncertainty grows like $h$ (variance $\propto h^2$) | The last slope rules the forecast. Prophet also simulates future changes at the historical rate, sizes Laplace(0, mean $|\hat\delta|$). | A short last segment gives a noisy final slope. Whether your forecasts simulate future bends is a design choice: know it. | 7.8 |
| Too few: stiff trend, biased final slope. Too many unshrunk: zig-zag trend, unstable final slope | Candidates (grid + PELT) and the Laplace scale $b$ are the trend's bias–variance knob. | Training error always prefers more changepoints; choose by out-of-sample forecast error. | 7.8 |
| Three approaches: fixed grid + sparse prior; detection (PELT) then fixed locations; Bayesian latent $p(s\mid y)\propto p(s)p(y\mid s)$ | Where may the trend bend? Many candidates, a few chosen from the data, or an uncertain location. | Your model's posterior is conditional on the grid and PELT locations: it does not include location uncertainty. | 7.8 7.8 7.8 |
7.9 · PELT change-point detection
| Formula | In words | Main trap | See |
|---|---|---|---|
| L2 cost $\sum(y_t-\bar y_{\text{seg}})^2$; Normal cost $L\log\hat\sigma^2_{\text{seg}}$; linear cost: line fit per segment | How well one simple model fits a segment. L2 sees level shifts, Normal sees level and spread, linear sees slope changes. | “No changepoint found” only means none of the kind the cost can see. L2 on a trend gives a staircase. | 7.9 7.9 7.9 |
| Objective $\min\sum_{j=0}^{m}C(y_{\tau_j:\tau_{j+1}})+\beta m$; BIC-like $\beta$: L2 $2\hat\sigma^2\log n$, Normal $3\log n$, linear $3\hat\sigma^2\log n$; $\hat\sigma=1.4826\,\text{MAD}(\Delta y)/\sqrt2$ | A cut is kept only if it saves more than $\beta$. | $\beta$ lives in cost units ($y^2$ for L2): rescaling $y$ changes its meaning. Rules of thumb assume independent Normal noise. | 7.9 |
| Optimal partitioning $F(t)=\min_{s\lt t}[F(s)+C(y_{s:t})+\beta]$, $F(0)=-\beta$; about $n^2/2$ cost evaluations | Exact dynamic programming over the start of the last segment. | Exact for the cost and $\beta$ you chose, not for the truth. | 7.9 |
| PELT: drop start $s$ if $F(s)+C(y_{s:t})\gt F(t)$; safe because $C(y_{s:T})\ge C(y_{s:t})+C(y_{t:T})$ | Same answer as optimal partitioning; roughly linear time when changes are frequent, $O(n^2)$ with none. | “PELT finds the true changepoints” is wrong: it minimises a penalised cost. | 7.9 |
rpt.Pelt(model="l2", min_size=2, jump=1); plot $m(\beta)$ over a grid of $\beta$ | Small $\beta$ over-detects, large $\beta$ under-detects; trust a long plateau. | jump defaults to 5 and rounds dates; ruptures returns segment ends, the last being $n$. | 7.9 7.9 |
| $p(\theta\mid y)=\sum_\tau p(\theta\mid y,\tau)p(\tau\mid y)$ vs the pipeline's $p(\theta\mid y,\hat\tau(y))$ | Selection uncertainty is not propagated: intervals near uncertain changepoints are too narrow and selected changes look too big. | Mitigate with candidates + Laplace shrinkage, stability checks over $\beta$ and resamples, sensitivity refits. | 7.9 |
| Leakage: PELT, $\beta$, $\hat\sigma$ and scaling are computed inside every backtest fold | Any input to a fold computed with data after the origin lets the test period place the changepoints. | Recent changes are detected late in real life; an honest backtest shows that delay. | 7.9 |
7.10 · Laplace priors on trend changes
| Formula | In words | Main trap | See |
|---|---|---|---|
| $g(t)=m+kt+\sum_j\delta_j(t-s_j)_+$, $\delta_j\sim\text{Laplace}(0,b)$, $p(\delta)=\frac1{2b}e^{-|\delta|/b}$; $E|\delta|=b$, sd $=\sqrt2\,b$ | Many candidates, few real changes: flexible where the data insist, stiff elsewhere. | $b$ is not the sd. Free $\delta_j$ (no prior) chase noise; training error below the noise level is the tell-tale. | 7.10 |
| Penalty and pull: Laplace $|\delta|/b$ (constant pull $1/b$); Normal $\delta^2/2\tau^2$ (pull $\delta/\tau^2$); tails $P(|\delta|\gt c)=e^{-c/b}$ ($P(\lt b/2)=0.39$, $P(\gt3b)=0.05$) | At equal variance ($\tau=\sqrt2\,b$) the Laplace has more mass near 0 and in the tails; its pull is stronger below $|\delta|=2b$. | Big changes are far more plausible under a Laplace than a Normal. | 7.10 |
| One change: $x\sim N(\delta,se^2)$, $se^2=\sigma^2/\sum(t-s_j)_+^2$; MAP $=\text{sign}(x)\max(0,|x|-se^2/b)$; evidence $\sim n^3/3$ | The MAP stays at 0 until the evidence beats $se^2/b$, so forecasts react late to real shifts, especially near the end of the history. | A shrunk $\delta_j$ is not the true size of the change. | 7.10 |
| MAP is a lasso: $\frac1{2\sigma^2}\sum(\cdot)^2+\frac1b\sum|\delta_j|$, $\lambda=\sigma^2/b$; ISTA step $\delta\leftarrow S_{\eta/b}(\delta+\frac\eta{\sigma^2}X^\top(y-X\delta))$ | Soft-thresholding $S_\theta(u)=\text{sign}(u)\max(0,|u|-\theta)$; FISTA adds momentum. | Exact zeros need a corner-aware solver and only exist at the MAP. Adam on a non-smooth objective ends near 0, not at 0. | 7.10 |
| $P(\delta_j=0\mid y)=0$: draws, means and medians are shrunk, never zero | Sparse-ish, not sparse. Report the posterior slope over time with a band; real changes are often shared by neighbouring candidates. | “Laplace prior gives a sparse posterior” and “interval includes 0 means no change” are both wrong. | 7.10 |
| Prophet: $\tilde y=y/\max|y|$, $\tilde t\in[0,1]$, $\tau=0.05$; real units $\delta_{\text{per day}}=\tilde\delta\,\max|y|/\text{days}$ | 0.05 means a typical change of 5% of the max per history length; with max about 1 000 and two years it is about 25 orders/day per year. | Reusing 0.05 under a different scaling of $y$ or $t$. | 7.10 |
| Choose $b$ on a log grid by rolling-origin error; refit at $b/3$ and $3b$ | Small $b$ underfits (outdated slope), large $b$ overfits (wiggles, unstable last slope). | Never choose by training error or on the reported test period; the best $b$ depends on the candidate list. | 7.10 |
7.11–7.12 · Fourier seasonality, holidays, regressors and leakage
7.11 · Fourier seasonality and Fourier order
| Formula | In words | Main trap | See |
|---|---|---|---|
| $\sin(2\pi t/P)$, $\cos(2\pi t/P)$; $f=1/P$, $\omega=2\pi/P$; harmonic $n$: $\sin,\cos(2\pi nt/P)$, period $P/n$ | Waves in $[-1,1]$ that average to 0; every harmonic repeats every $P$. | Radians, and $t$ and $P$ in the same unit. A weekly period of 7 applied to a time rescaled to $[0,1]$ is a silent bug. | 7.11 7.11 |
| $a\cos\omega t+b\sin\omega t=R\cos(\omega t-\varphi)$, $R=\sqrt{a^2+b^2}$, $\varphi=\operatorname{atan2}(b,a)$, peak at $t^*=\varphi/\omega$ | Two linear weights encode a size and a timing, so the model stays linear. | A single sine forces the peak day; single weights mean nothing, compute $R$ per posterior draw. | 7.11 |
| $s(t)=\sum_{n=1}^{N}[a_n\cos(2\pi nt/P)+b_n\sin(2\pi nt/P)]$; block row $[\sin_1,\cos_1,\dots,\sin_N,\cos_N]$; $2N$ weights; whole periods $X^\top X=\frac T2I$ | Any repeating shape from a few waves, as regression columns. $P$ and $N$ are fixed, only weights are learned. | Jumps ring (Gibbs, about 9% overshoot); dated spikes belong in holiday terms. Same time origin and units for future rows. | 7.11 7.11 |
| Order $N$: training error falls, holdout is U-shaped; exact ceiling $N\le\lfloor P/2\rfloor$; prior $N(0,\tau^2)$ is ridge with $\lambda=\sigma^2/\tau^2$ | Order is a bias–variance knob; weekly order 3 on daily data equals the six weekday dummies. | Higher order is not more accurate. Choose by rolling-origin validation and residual ACF at the seasonal lags. | 7.11 |
| Parameters: $2N$ per seasonality (yearly 10 + weekly 3 = 26); less than one period of history makes the wave nearly the trend | See the cycle at least once, ideally twice (Prophet's yearly default needs two years). | Weak identification: ridge-shaped posterior, results driven by the prior. | 7.11 7.11 |
| Nyquist $f_N=f_s/2$; aliasing $f\to|f-mf_s|$; useful order $\le\lfloor P/2\rfloor$ | Daily data cannot show a daily cycle (it aliases to a constant) and weekly harmonic 4 duplicates harmonic 3. | Extra aliased columns are redundant (singular $X^\top X$), not extra detail. | 7.11 |
7.12 · Holidays, exogenous regressors and leakage
| Formula | In words | Main trap | See |
|---|---|---|---|
| $h_t=\sum_j\beta_jI(t\in D_j)$; window: columns $I(t-k\in D)$, $k=L..U$, $U-L+1$ coefficients | One 0/1 column per holiday, fitted jointly with trend and season; dates must cover the forecast horizon. | Without the column spikes inflate $\hat\sigma$ and bend nearby terms. Each offset is learned from as many days as there are occurrences. | 7.12 7.12 |
| Recurring: one shared $\beta$, past and future dates. One-off: 1 only on its own days, 0 in the future. Overlap: if always together only $\beta_A+\beta_B$ is identified | A one-off protects trend and season; overlapping columns split credit by the prior. | An unflagged one-off near the end of the history becomes a fake trend change. Posterior correlation near $-1$ means not identified. | 7.12 7.12 |
| $\beta_j\sim N(0,\tau^2)$: $E[\beta_j\mid D]=w\bar r_j$, $w=\dfrac{n/\sigma^2}{n/\sigma^2+1/\tau^2}$; MAP = ridge, $\lambda=\sigma^2/\tau^2$ | Few, noisy occurrences means strong shrinkage and lower mean squared error. | $\tau$ is in the units of the (scaled) $y$; Prophet's default 10 on $y/\max|y|$ is very weak. | 7.12 |
| Categorical: $K-1$ dummies vs a reference; lag $x_{t-L}$ known at the origin if $L\ge h$; interaction slope $\beta_a+\beta_{ab}x_b$ | Coefficients are associations, with the other columns held fixed. | All $K$ dummies + intercept is the dummy trap. Keep main effects with an interaction; centre inputs. | 7.12 7.12 |
| $VIF_j=1/(1-R_j^2)$, $SE\times\sqrt{VIF}$; two regressors: $1/(1-r^2)$ (0.9 gives 5.3, 0.99 gives 50) | Collinear columns give wobbly, correlated coefficients but stable sums and in-range predictions. | Rule of thumb only (above 5–10 deserves a look). A mean-field guide hides the negative correlations. | 7.12 |
| Standardize: $z=(x-\bar x_{\text{train}})/s_{\text{train}}$, $\beta_{\text{raw}}=\beta_z/s$ | One prior scale then treats all regressors fairly; coefficients are per standard deviation. | Full-series statistics leak the future. “Per sd” is not “per unit”. | 7.12 |
| Forecast regressor: $Var=\sigma^2+\beta^2Var(X-\hat X)$; $p(y\mid\mathcal F_T)=\int p(y\mid X,\theta)p(X\mid\mathcal F_T)dX$ | Known in advance (calendar, plans, long lags) or must be forecast (weather, competitors, traffic). | Backtesting with the actual future regressor values is an oracle, i.e. leakage. | 7.12 |
| Leakage test: “would I have had this number at the origin?” for features, regressors, preprocessing and tuning | Fit every learned step on the training window, freeze it, refit everything at each rolling origin. | A great backtest from a centred window, actual future regressors or a same-day proxy of the target is a leak until proven otherwise. | 7.12 7.12 |
7.13–7.14 · Forecast likelihoods and Bayesian forecasting
7.13 · Forecast likelihoods: Normal, Student-t, Negative Binomial
| Formula | In words | Main trap | See |
|---|---|---|---|
| $y_t\mid\mu_t,\varphi\sim p(\cdot\mid\mu_t,\varphi)$, $\ell=\sum_t\log p(y_t\mid\mu_t,\varphi)$ | The layers say where (the centre); the likelihood says how far, how heavy the tails, and whether the spread grows with $\mu$. | A great centre with the wrong likelihood still gives wrong intervals and tail probabilities. Days are assumed independent given $\mu_t$. | 7.13 |
| Normal: $\ell=-\frac T2\log(2\pi\sigma^2)-\frac1{2\sigma^2}\sum r_t^2$, $\hat\sigma^2=\frac1T\sum r_t^2$ | Fitting is least squares. Promises: continuous, symmetric, light tails, constant variance, independent. | It is the residuals, not the series, that must be Normal. One far day inflates $\hat\sigma$ and widens every interval. | 7.13 |
| Student-t: sd $=\sigma\sqrt{\nu/(\nu-2)}$ ($\nu\gt2$); weights $w_t=\dfrac{\nu+1}{\nu+r_t^2/\sigma^2}$; pull $\psi(r)=\dfrac{(\nu+1)r}{\nu+r^2}$ | $\nu\to\infty$ is the Normal, $\nu=1$ the Cauchy. A far day gets a tiny weight (0.003 in the chapter's glitch example), not zero. | It assigns more probability to extreme residuals so they exert less influence; it does not remove outliers. $\nu$ is a tail dial, $\sigma$ is not the sd. | 7.13 7.13 |
| Counts: Poisson $Var=\mu$; Pearson index $\hat\phi=\frac1{n-p}\sum\frac{(y-\hat\mu)^2}{\hat\mu}$; gamma–Poisson mixture $\Rightarrow$ NB | Counts are whole numbers, right-skewed near 0, with variance growing with the mean. Overdispersion comes from a wobbling rate. | A Normal puts probability on negative counts; it is fine only for large counts with steady spread (rule of thumb above about 20–30). | 7.13 7.13 |
| NB2: $P(y)=\frac{\Gamma(y+\alpha)}{y!\,\Gamma(\alpha)}\big(\frac\alpha{\alpha+\mu}\big)^\alpha\big(\frac\mu{\alpha+\mu}\big)^y$; $E=\mu$, $Var=\mu+\mu^2/\alpha$, $P(0)=(\frac\alpha{\alpha+\mu})^\alpha$ | $\alpha\to\infty$ is Poisson; small $\alpha$ is noisy. Example $\mu=40$, $\alpha=5$: sd 19.0, 90% interval [14, 75]. | “Dispersion” may mean $\alpha$ or $1/\alpha$. NB1 (variance linear in the mean) is a different model. | 7.13 7.13 |
| Link: $\eta_t=g+s+h+X\beta$, $\mu_t=e^{\eta_t}$ (log) or $\log(1+e^{\eta_t})$ (softplus) | Log: effects multiply (a +30% holiday). Softplus: about $\eta$ when large, effects add (+30 orders). Baseline 100 gives 130 either way; at baseline 10: 13 vs 40. | Coefficient and prior units change with the link; apply the link per posterior draw before averaging. | 7.13 |
| ZIP: $P(0)=\pi+(1-\pi)f(0)$, $P(k)=(1-\pi)f(k)$. Hurdle: $P(0)=\pi$, $P(k)=(1-\pi)f(k)/(1-f(0))$ | $\pi=0.3$, $\lambda=4$: ZIP $P(0)=0.313$; a Poisson with the same mean 0.061. | Ask why the zeros are there: low rates (NB fine) or structural (regressor or mask first). | 7.13 |
| Held-out $\text{lpd}=\sum_i\log\big(\frac1S\sum_sp(y_i^{\text{new}}\mid\theta^{(s)})\big)$; AIC $=2k-2\ell$ | Compare likelihoods on later days, in the same currency. | Density vs probability: dividing $y$ by $c$ raises every density log-score by $\log c$ per day (multiplying lowers it by the same amount). In-sample criteria assume exchangeable data. | 7.13 |
7.14 · Bayesian forecasting: predictive distributions, uncertainty and PPCs
| Formula | In words | Main trap | See |
|---|---|---|---|
| $p(\tilde y_{T+h}\mid D)=\int p(\tilde y\mid\theta)p(\theta\mid D)d\theta$; sample: $\theta^{(s)}$ then $\tilde y^{(s)}$ | Mean = expected demand, median = typical day, $q$-quantile = level not exceeded with probability $q$, share above $c$ = $P(\text{demand}\gt c)$. | A 95% credible interval for the mean is not a 95% prediction interval for tomorrow. Monte Carlo error of a tail probability: $\sqrt{p(1-p)/S}$. | 7.14 |
| Draws array $\tilde Y$ ($S$ futures × $H$ days): a window question uses rows (max, sum), then counts or quantiles across rows | Means add across days; quantiles do not; variances add only with the covariances. | Save the draws, not a table of daily quantiles: $q_p(\sum\tilde y_j)\ne\sum q_p(\tilde y_j)$. | 7.14 |
| $s^*=F^{-1}\!\big(\frac{c_u}{c_u+c_o}\big)$ (critical ratio); squared loss $\to$ mean, absolute loss $\to$ median | The best stock is a predictive quantile set by the costs of running short and running over. | Needs a calibrated right tail; for multi-day lead times use the total over the paths. | 7.14 |
| $Var(\tilde y\mid D)=E[Var(\tilde y\mid\cdot)]+Var(E[\tilde y\mid\cdot])$: noise + parameters + regressors + model | Posterior predictive covers noise and parameters. Model and regressor uncertainty must be added on purpose. | “Full uncertainty quantification” overclaims. Say whether regressors are held at planned values (conditional) or simulated. | 7.14 |
| $Var(\tilde y_{T+h})=\sigma^2+s_a^2+h^2s_b^2+\lambda\frac{2b^2}{3}h^3$; sd: noise flat, slope $\propto h$, future changes $\propto h^{3/2}$, random-walk noise $\sqrt h$, AR(1) levels off | Intervals widen because trend doubt grows with distance. Noise does not pile up when it is independent. | A model that holds the last slope fixed treats the future trend as certain: a narrower, optimistic cone. | 7.14 |
| PPC: $p_T=P(T(y^{\text{rep}})\ge T(y)\mid D)$; PIT $u_j=F_j(y_j)\approx$ share of draws below the actual | Summaries: mean, variance, max/tails, zero days, weekday shape, lag-1 and lag-7 autocorrelation, last-vs-first level. Calibrated forecasts have uniform PIT. | $p_T$ is a diagnostic, not the probability the model is true; passing a PPC does not validate the forecast, a backtest does. | 7.14 7.14 |
7.15–7.17 · Accuracy, probabilistic evaluation and residual diagnostics
7.15 · Time-series cross-validation and accuracy metrics
| Formula | In words | Main trap | See |
|---|---|---|---|
| Holdout: fit on $y_1..y_T$, error $e_{T+h\mid T}=y_{T+h}-\hat y_{T+h\mid T}$; train → validation → test in time order | Everything that learns from data (scalers, changepoints, tuning) sees training data only. | One holdout is one draw. Tuning on the test block turns it into a validation block. | 7.15 |
| Rolling origin: $K=\lfloor(N-T_0-H)/s\rfloor+1$ origins; expanding $\{1..T\}$ or sliding $\{T-W+1..T\}$; $\text{MAE}_h=\frac1K\sum_j|e_{T_j+h\mid T_j}|$ | Refit inside each window, store errors with $T$ and $h$, report by horizon. Random walk + naive: $\text{RMSE}_h=\sigma\sqrt h$. | Origins overlap, so errors are correlated. A step that is a multiple of the season length ties horizon to weekday. Shuffled K-fold leaks; use forward chaining with a gap. | 7.15 7.15 7.15 |
| MAE $=\frac1n\sum|e_i|$; RMSE $=\sqrt{\frac1n\sum e_i^2}$; bias $=\frac1n\sum e_i$; $\arg\min_c\sum(y_i-c)^2=\bar y$, $\arg\min_c\sum|y_i-c|=$ median | Typical miss; typical miss with big misses counted extra; systematic direction. Always MAE $\le$ RMSE. | Zero bias is not accuracy. Report the predictive mean for RMSE, the median for MAE. All three are scale-dependent. | 7.15 7.15 |
| MAPE $=\frac{100}n\sum\frac{|e_i|}{|y_i|}$; sMAPE $=\frac{100}n\sum\frac{2|e_i|}{|y_i|+|\hat y_i|}$; WAPE $=100\frac{\sum|e_i|}{\sum|y_i|}$ | WAPE is MAE over the mean actual: zero-safe, big days weigh more. | MAPE is undefined at 0, explodes near 0, favours low forecasts; “100 − MAPE” is not accuracy; sMAPE is not symmetric and has several formulas. | 7.15 |
| MASE $=\text{MAE}_{\text{test}}\Big/\frac1{T-m}\sum_{t=m+1}^{T}|y_t-y_{t-m}|$ | Scale-free, zero-safe, below 1 means smaller than the in-sample naive error. State $m$ (7 for daily data with a weekly pattern). | The ruler is a one-step error, so even naive scores above 1 on long horizons: compare with naive on the same windows. | 7.15 |
| Fair bake-off: same origins, windows, horizons; baselines included; metric chosen from the decision first | Report by horizon, the spread across origins, paired differences and bias; average a scaled score (MASE, WAPE) across series. | Averaging RMSE across series of different size lets the biggest dominate; picking the metric after the winner. | 7.15 |
7.16 · Probabilistic forecast evaluation
| Formula | In words | Main trap | See |
|---|---|---|---|
| Coverage $=\frac1n\sum\mathbf 1\{l_t\le y_t\le u_t\}$; luck range $\pm1.96\sqrt{p(1-p)/n}$ | Share of outcomes inside the interval. Below nominal means overconfident (too narrow); above means too cautious. | Judge with $n$; very wide intervals cover everything; use predictive, not trend-only, intervals; check by level, horizon and slice. | 7.16 |
| PIT $u_t=F_t(y_t)\approx$ share of samples $\le y_t$ | Flat = calibrated, U = overconfident, hump = underconfident, slope = biased. Randomized PIT for counts. | Flat is necessary, not sufficient (an always-wide forecast is flat). | 7.16 |
| Interval score $=(u-l)+\frac2\alpha(l-y)\mathbf 1\{y\lt l\}+\frac2\alpha(y-u)\mathbf 1\{y\gt u\}$ | Width plus a miss penalty; lower is better. Maximise sharpness subject to calibration. | Coverage alone rewards width, width alone rewards narrowness: use them together. | 7.16 |
| Log score $\log p_t(y_t)$ (nats, higher is better); from draws $\text{logsumexp}_s\log p(y\mid\theta_s)-\log S$ | Rewards probability on the outcome and punishes confident misses; expected gap to the truth is a KL divergence. | Needs a density; one extreme day can dominate; log of the mean, not mean of the logs. | 7.16 |
| CRPS $=\int(F(x)-\mathbf 1\{x\ge y\})^2dx=E|X-y|-\frac12E|X-X'|$; Normal: $\sigma[z(2\Phi(z)-1)+2\varphi(z)-1/\sqrt\pi]$ | In the units of $y$; for a point forecast it is $|y-f|$. Computable from samples, no density needed. | Scale before comparing series of different size. The plain area between CDF and step is $E|X-y|$, not the CRPS. | 7.16 |
| Pinball $\rho_\tau(y,q)=\tau(y-q)$ if $y\ge q$, else $(1-\tau)(q-y)$; CRPS $=2\int_0^1\rho_\tau(y,F^{-1}(\tau))d\tau$ | Best $q$ is the true $\tau$-quantile; $\tau=\frac12$ is half the absolute error; newsvendor $\tau^*=c_u/(c_u+c_o)$. | Use the predictive quantile, not mean + 1.28 sd, for skewed data; quantiles must not cross. | 7.16 |
| Constant band, random-walk forecasts: coverage at horizon $h$ is $2\Phi(z/\sqrt h)-1$ (80% gives 48% at $h=4$, 33% at $h=9$) | Average coverage hides horizon and slice failures. | Use the model's own multi-step distribution; slice by weekday, holiday, level and regime. | 7.16 |
7.17 · Residual diagnostics
| Formula | In words | Main trap | See |
|---|---|---|---|
| $e_t=y_t-\hat y_t$; $z_t=e_t/\hat\sigma$; five questions: centred? constant spread? right shape? pattern vs fitted/time? memory? | Residual = true noise + what the model missed. Goal: no exploitable structure. | With an intercept the in-sample mean is 0 by construction: look at group means and the holdout. | 7.17 7.17 |
| $r_k=\dfrac{\sum(e_t-\bar e)(e_{t-k}-\bar e)}{\sum(e_t-\bar e)^2}$, band $\pm1.96/\sqrt n$; $Q(h)\sim\chi^2_{h-m}$; $DW\approx2(1-r_1)$ | Choose $h$ to reach the season: about 10 non-seasonal, $2m$ seasonal (14 for daily data with a weekly pattern), at most $n/5$. Test $e_t^2$ for volatility clustering. | Small p means memory somewhere (read the ACF to see where); DW sees lag 1 only. | 7.17 |
| Q-Q of standardized residuals: S-shape = heavy tails; one bent end = skew; kink = two groups. Plots: bowl, funnel, waves, step, spikes | Each fingerprint points to a fix: Student-t, transform, hidden group, changepoints, holiday columns. | Heavy tails may be a missed holiday or changing spread: check dates and residual-vs-fitted first. | 7.17 7.17 |
| Heteroscedastic: $Var(\epsilon_t)=\sigma_t^2$; spread ratio far above 1.5–2; fixes: log ($sd\propto\mu$), sqrt ($Var\propto\mu$), NB, $\sigma_t=c\mu_t$ or $\exp(a+bx_t)$ | Biases the intervals, not the mean. | Student-t does not fix it; back-transformed forecasts are medians. | 7.17 |
| Pearson $r_t=(y_t-\hat\mu_t)/\sqrt{V(\hat\mu_t)}$ ($V=\mu+\mu^2/\alpha$ for NB2); randomized quantile $r^q_t=\Phi^{-1}(u_t)$, $u_t=F(y_t-1)+v_t[F(y_t)-F(y_t-1)]$ | A right model gives variance about 1 and $N(0,1)$ for any likelihood. | Raw residuals of a correct count model fan out. Check which NB parameterization your code uses before computing $V$. | 7.17 |
| Lineup / PPC on a statistic ($Q$, $r_1$, max $|z|$, spread ratio): share of replicates at least as extreme as the real value | Noise makes patterns: compare the real residual plot with ones from data simulated from the fitted model. | A clean panel is not a proof; accuracy and calibration decide. | 7.17 7.17 |
| AR(1) error $\epsilon_t=\varphi\epsilon_{t-1}+\eta_t$, forecast correction $\varphi^he_T$; exponential-kernel GP $=$ AR(1) with $\varphi=e^{-1/\ell}$; local level has permanent memory | Fix the mean function first; add an AR term, state-space model or GP only if momentum remains. | A flat ACF after an AR term does not prove the cause was momentum; check horizons. | 7.17 7.17 |
7.18–7.20 · Complexity and alternatives, production, the capstone
7.18 · Model complexity and forecasting alternatives
| Formula | In words | Main trap | See |
|---|---|---|---|
| Expected squared error $=\text{bias}^2+\text{variance}+\sigma^2$; least squares variance $\approx\sigma^2p/n$ | Complexity up: bias down, variance up. Training error only falls; holdout error is U-shaped. | Choose complexity on a time-ordered holdout, not on training error. | 7.18 |
| Knobs: capacity (Fourier $N$, candidates $C$), strength (Laplace $b$, pooling $\tau$), cost (guide rank: $2d$, $d(r+2)$, $d+d(d+1)/2$) | The first four sit on the U-curve; the guide rank has no U-curve (nested family, diminishing returns). | A generous $C$ with a small $b$ is safe; with a large $b$ it is not. | 7.18 |
| Assumptions: additive layers; sparse bends, last slope continues; fixed season; repeatable known events; independent noise; stable future; good posterior; one series | Each assumption has a residual symptom, a test and an alternative family. | “State of the art” is not an argument; rolling-origin accuracy and calibration are. | 7.18 |
| AR(1) noise: error ratio of a memory model vs independent noise $\sqrt{1-\varphi^{2h}}$ (0.60 at $h=1$, 0.98 at $h=7$ for $\varphi=0.8$) | ARIMA and ETS read the recent past: they win with momentum, wandering levels and short horizons. | They lose with several seasonalities, moving holidays, complex regressors, counts and long horizons. | 7.18 |
| State space: $y_t=Z\alpha_t+\varepsilon_t$, $\alpha_t=T\alpha_{t-1}+\eta_t$; local level gain $\alpha=(\sqrt{q^2+4q}-q)/2$, $q=\sigma_\eta^2/\sigma_\varepsilon^2$ | A hidden level/slope/season that drifts; Kalman filter in $O(n)$; equals exponential smoothing. | Exact filtering needs Gaussian noise; many sparse events and counts are awkward. | 7.18 |
| GP: posterior mean $k_*^\top(K+\sigma_n^2I)^{-1}y$, variance $k_{**}-k_*^\top(K+\sigma_n^2I)^{-1}k_*$; cost $O(n^3)$ | The kernel is the assumption; honest uncertainty on small, smooth data. | Long series, trend extrapolation (an RBF GP reverts to the mean), many seasonalities and counts. | 7.18 |
| Boosting $F_M(x)=F_0+\eta\sum_mh_m(x)$; global models share strength across series; N-BEATS, TFT, DeepAR (awareness) | Win with many series and rich nonlinear drivers. | Trees cannot extrapolate a trend; lag features must be at least as old as the horizon; intervals need quantile loss or conformal methods. | 7.18 7.18 |
7.19 · Production statistical modeling
| Formula | In words | Main trap | See |
|---|---|---|---|
| Run manifest = data hash + preprocessing + code version + priors + seed + inference settings + guide + optimizer + iterations run + library versions | Written automatically beside the parameters, compared with a diff; includes the automatic choices (guide picked, PELT candidates, scaler numbers, best-state step). | A seed alone is not reproducibility; record the iterations actually run, not only the maximum. | 7.19 |
| Compare models by gap $\div$ seed-to-seed spread over 5–20 seeds; $SE_{\text{gap}}=s\sqrt{2/m}$ | Same seed gives the same numbers up to floating-point differences, not the same bits. | Reporting the best seed; trusting one run; a gap below 1–2 spreads is no claim. | 7.19 |
| Rolling MAE/MASE vs launch; bias limits $\bar e_w\notin\pm3\sigma_e/\sqrt w$; alarm after $k$ bad windows | Store forecasts as issued, score them when actuals arrive. | Launch accuracy is not current accuracy; an alarm is a question, not a retrain order; autocorrelated residuals need wider limits. | 7.19 |
| Misses in $w$ days $\sim\text{Bin}(w,\alpha)$: mean $w\alpha$, sd $\sqrt{w\alpha(1-\alpha)}$; PIT shape; PSI $=\sum(a_i-e_i)\ln(a_i/e_i)$ (0.1 / 0.25); KS $=\max|F_{ref}-F_{new}|$ | 28 days at 80%: 5.6 ± 2.1 misses. Data drift changes $P(x)$; concept drift changes $P(y\mid x)$ and shows only in the errors. | The 80% is a claim, not a fact; check several levels and both directions. Input drift is a prompt to look, not proof the model failed. | 7.19 7.19 |
| $\text{LSE}(a)=m+\log\sum_ke^{a_k-m}$, $m=\max a$; log of an average $=\text{LSE}-\log S$; float64 range about $5\times10^{-324}$ to $1.8\times10^{308}$, $e^x$ overflows above 709.8 | Sum logs, never multiply probabilities: float32 underflows after about 33 factors of 0.04, float64 after 232. | Epsilon hacks hide bugs. A log density can be positive. | 7.19 |
| Free coordinate $u$ and constrained $\theta=g(u)$: $p_u(u)=p_\theta(g(u))|g'(u)|$ (exp: $+u$ in the log); positive: $e^u$ or $\ln(1+e^u)$; $(0,1)$: sigmoid; simplex: softmax or stick-breaking | Guides are Gaussian in the free space, so a positive parameter is log-normal-shaped and right-skewed. | Quantiles map across the transform; the mode and the mean do not. Clipping kills gradients; exp can explode. | 7.19 |
| Stable learning rate for a quadratic: $\text{lr}\lt2/\lambda_{\max}$; scaling inputs shrinks the condition number; double-where for masked logs | NaN debugging: finite inputs? finite loss at init? first non-finite step? finite gradients? | A scaler fitted on all data leaks the future; per-group scaling erases group effects. A non-finite loss is never an improvement in early stopping. | 7.19 |
7.20 · Capstone: one set of ideas, two projects
| Formula | In words | Main trap | See |
|---|---|---|---|
| Likelihood from support, then variance: Binomial $np(1-p)$, Poisson $\lambda$, NB $\mu+\mu^2/\alpha$, Normal $\sigma^2$, Student-t $\sigma^2\nu/(\nu-2)$ | Yes/no, categories, counts, positive, real; then constant, growing with the mean, heavy tails. | Normal for bounded or count data; Poisson when variance exceeds the mean; mixing NB parameterizations. | 7.20 |
| Partial pooling $\hat\theta_g=w_g\bar y_g+(1-w_g)\mu$ vs Laplace MAP $\text{sign}(z)\max(|z|-s^2/b,0)$ | Both pull toward a centre (the overall mean, or zero) by evidence against the prior scale ($\tau$ or $b$). | Laplace posterior is not exactly sparse; per-group scaling kills group effects. | 7.20 |
| $\text{MSE}(c\bar x)=(1-c)^2\mu^2+c^2\sigma^2/n$, best $c^*=\mu^2/(\mu^2+\sigma^2/n)$ | Every flexibility knob sits on one bias–variance curve; choose on held-out time periods. | Training error rewards complexity; unbiased is not best. | 7.20 |
| $\log p(D)=\text{ELBO}+KL(q\|p(\cdot\mid D))$; guide sizes $2d$, $d(r+2)$, $d+d(d+1)/2$ | SVI maximises the ELBO; the answer is only as good as the guide family. | “Converged” is about the optimizer, not the approximation; NUTS is asymptotically exact, not exact. | 7.20 7.20 |
| Normal, known $\sigma$: gain posterior sd $\sigma\sqrt{2/n}$; one new value sd $\sigma\sqrt{1+1/n}$ (floor $\sigma$) | $P(\theta_B\gt\theta_A\mid D)$ is about parameters and sharpens with data; $p(y_{\text{future}}\mid D)$ has a noise floor. | “99% sure it is better” is not “better for 99% of users”. | 7.20 |
The forecasting model on one page: $y=g+s+h+X\beta+\epsilon$
Each component with what is learned, what is fixed in advance, and its prior. The priors are illustrative; the Prophet values are the ones the chapters quote for Prophet's own scaled data, and your NumPyro model has its own settings, so check which scale yours uses.
| Component | Form | Learned parameters | Fixed in advance | Priors (illustrative; Prophet's where the chapters state them) | See |
|---|---|---|---|---|---|
| Trend $g(t)$ | $g(t)=kt+m+\sum_{j=1}^{S}\delta_j(t-s_j)_+$ (hinge form; the same as the offset form with $\gamma_j=-s_j\delta_j$) | base slope $k$, offset $m$, slope changes $\delta_1..\delta_S$ ($S$ weights) | locations $s_j$ (grid + PELT), their number $S$, time scaling | $\delta_j\sim\text{Laplace}(0,b)$; $k,m$ Normal (Prophet: $N(0,5^2)$ on scaled data); Prophet $b=0.05$ on $y/\max|y|$ | 7.8 7.8 7.9 7.10 |
| Seasonality $s(t)$ | $s(t)=\sum_k\sum_{n=1}^{N_k}[a_{kn}\cos\frac{2\pi nt}{P_k}+b_{kn}\sin\frac{2\pi nt}{P_k}]$ | $a_{kn},b_{kn}$: $2N_k$ weights per period (yearly 10 + weekly 3 = 26) | periods $P_k$ (7, 365.25 for daily data), orders $N_k$, time origin | Normal $N(0,\tau^2)$ on the weights (Prophet: scale 10 on scaled $y$); a ridge in disguise | 7.11 7.11 7.11 |
| Holidays $h(t)$ | $h(t)=\sum_j\beta_jI(t\in D_j)$ (plus window offsets $I(t-k\in D)$) | one weight $\beta$ per holiday and window offset | dates $D_j$ (past and future), windows $[L,U]$, which events | shrinkage prior $N(0,\tau^2)$ (Prophet: scale 10); Laplace or a hierarchical prior are alternatives | 7.12 7.12 7.12 |
| Regressors $X_t\beta$ | $X_t\beta=\sum_jx_{t,j}\beta_j$ (continuous, binary, $K-1$ dummies, lags, interactions) | $R$ weights $\beta_j$ | which columns, lags, the training-window scaler, future values of $X$ | Normal (ridge-like) or Laplace on standardized coefficients; hierarchical for many candidates | 7.12 7.12 7.12 |
| Noise $\epsilon_t$ (the likelihood) | $y_t\sim\text{Lik}(\mu_t,\varphi)$: Normal($\mu_t,\sigma^2$), StudentT($\nu,\mu_t,\sigma$), NB2($\mu_t,\alpha$) | $\sigma$; $\sigma,\nu$; $\alpha$ (one shared set of noise parameters) | the family, the link for counts (log or softplus; identity for Normal and Student-t) | positive priors on $\sigma$, $\nu$, $\alpha$ (Prophet: $\sigma\sim\text{HalfNormal}(0.5)$ on scaled data) | 7.13 7.13 7.13 7.13 7.13 |
| Inference | SVI: maximise the ELBO over a Gaussian guide in unconstrained space | guide parameters: mean-field $2d$, low-rank $d(r+2)$, full-rank $d+d(d+1)/2$ | guide family and rank, optimizer, stopping rule, seed | $d=$ (columns of $X$) + noise parameters + hyperparameters; read your own guide threshold from your code | 7.7 7.18 7.19 |
Reading it in an interview: point at one row, say its form, say which numbers are learned and which were fixed before fitting, and name the check that would catch a mistake there (the diagnostics table).
Metrics: formula, unit, strength, trap
$e_i=y_i-\hat y_i$. Point metrics (the first seven) grade one number taken from the forecast; the last four grade intervals, quantiles or the whole distribution. A point forecast should match its score: squared error rewards the predictive mean, absolute error the median.
| Metric | Formula | Unit | Strength | Trap | See |
|---|---|---|---|---|---|
| MAE | $\frac1n\sum|e_i|$ | units of $y$ | the typical miss; robust; best constant is the median | scale-dependent; says nothing about the spread of the forecast | 7.15 7.15 |
| MSE | $\frac1n\sum e_i^2$ | units of $y$ squared | smooth; matches a Normal likelihood | squared units, so quote RMSE; a few big errors dominate | 7.15 |
| RMSE | $\sqrt{\text{MSE}}$ | units of $y$ | big misses count extra; best constant is the mean; RMSE $\ge$ MAE | scale-dependent; large RMSE/MAE ratio means a few huge errors | 7.15 7.15 |
| MAPE | $\frac{100}n\sum|e_i|/|y_i|$ | percent | easy to say | undefined at 0, explodes near 0, favours low forecasts; “100 − MAPE” is not accuracy | 7.15 |
| sMAPE | $\frac{100}n\sum2|e_i|/(|y_i|+|\hat y_i|)$ | percent, 0–200 | bounded | not symmetric; several formulas share the name; undefined when $y_i=\hat y_i=0$ | 7.15 |
| WAPE | $100\sum|e_i|/\sum|y_i|$ | percent | zero-safe; volume-weighted; = MAE / mean(y) | big days dominate; undefined only if all actuals sum to 0 | 7.15 |
| MASE | $\text{MAE}_{\text{test}}/\frac1{T-m}\sum_{t\gt m}|y_t-y_{t-m}|$ | unitless ratio | scale-free; zero-safe; below 1 beats the in-sample naive error | one-step ruler (naive often $\gt1$ at long horizons); state $m$; undefined for a constant training series | 7.15 |
| Coverage | $\frac1n\sum\mathbf 1\{l_t\le y_t\le u_t\}$ | share (compare with $1-\alpha$) | directly tests the stated interval | very wide intervals cover everything; needs $n$ and the luck range $\pm1.96\sqrt{p(1-p)/n}$ | 7.16 |
| Pinball loss | $\tau(y-q)$ if $y\ge q$, else $(1-\tau)(q-y)$ | units of $y$ | proper for the $\tau$-quantile; matches a stock or capacity decision | grades one level at a time; quantiles must not cross | 7.16 |
| CRPS | $E|X-y|-\frac12E|X-X'|$ | units of $y$ | grades the whole distribution; from samples; = $|y-f|$ for a point forecast | scale-dependent: scale before comparing series; compare on the same origins | 7.16 |
| Log score | $\log p_t(y_t)$; draws: $\text{LSE}_s\log p(y\mid\theta_s)-\log S$ | nats (higher is better) | strictly proper; sharp-and-right is rewarded | needs a density or mass; one tail surprise dominates; do not compare densities with probabilities | 7.16 7.13 |
The three forecast likelihoods side by side
Support first, variance second. All three take the mean $\mu_t$ from the layers; the noise parameters are shared by all days. Mistakes in argument lists are silent, so print the mean and variance of any distribution you build.
| Normal | Student-t | Negative Binomial (NB2) | |
|---|---|---|---|
| Support | any real number | any real number | $0,1,2,\dots$ |
| Noise parameters | $\sigma$ | $\sigma,\ \nu$ | $\alpha$ (concentration) |
| Mean | $\mu_t$ | $\mu_t$ if $\nu\gt1$ | $\mu_t$ |
| Variance | $\sigma^2$, the same every day | $\sigma^2\nu/(\nu-2)$ if $\nu\gt2$ (infinite for $\nu\le2$); $\sigma$ is the scale, not the sd | $\mu_t+\mu_t^2/\alpha$, grows with the mean; $\alpha\to\infty$ is Poisson |
| Shape and tails | symmetric, light tails | symmetric, heavy tails; $\nu\to\infty$ is Normal, $\nu=1$ Cauchy | right-skewed, long right tail, $P(0)=(\frac{\alpha}{\alpha+\mu})^\alpha$ |
| Layers enter through | identity: $\mu_t=\eta_t$ | identity: $\mu_t=\eta_t$ | log ($e^\eta$: effects multiply) or softplus ($\approx\eta$ when large: effects add) |
| NumPyro | Normal(loc, scale): scale is the sd | StudentT(df, loc, scale) | NegativeBinomial2(mean, concentration); also GammaPoisson(concentration, rate), NegativeBinomialProbs(total_count, probs), NegativeBinomialLogits(total_count, logits) |
| SciPy | norm(loc, scale) | t(df, loc, scale) | nbinom(n=alpha, p=alpha/(alpha+mu)) (counts failures) |
| statsmodels / others | least squares / OLS | — | NB2 alpha $=1/\alpha$; Stan neg_binomial_2(mu, phi), R glm.nb theta $=\alpha$ |
| Fitted as | least squares | weighted least squares, $w=\frac{\nu+1}{\nu+r^2/\sigma^2}$ | GLM-type maximum likelihood / Bayesian |
| Use when | continuous, symmetric, steady spread (revenue-like) and light tails | continuous with occasional extreme residuals (glitches, unmodelled spikes); $\nu$ is a diagnostic | daily counts whose spread grows with the level; overdispersed relative to Poisson |
| Breaks when | counts, fat tails, a spread that fans with the level | many small counts; $\nu$ poorly identified (data cannot tell 30 from 200) | excess zeros with a known cause; autocorrelated noise |
| Check with | Q-Q plot, spread by fitted bin | Q-Q S-shape; posterior of $\nu$; weights of the lowest days | Pearson dispersion, variance-vs-mean plot, zero count in a PPC |
| Taught in | 7.13 | 7.13 7.13 | 7.13 7.13 7.13 7.13 |
One Negative Binomial, many argument lists
Translation table from the master pair $(\mu,\alpha)$ with $Var=\mu+\mu^2/\alpha$. For $\mu=40$, $\alpha=5$ all of these give mean 40 and variance 360 (checked in NumPyro 0.22 and SciPy 1.18). Check which class and which prior (on $\alpha$ or $1/\alpha$) your own code uses.
| Where | Arguments for the same law | Note |
|---|---|---|
| NB2 master pair | $(\mu,\alpha)$ | $Var=\mu+\mu^2/\alpha$; $\alpha$ large means close to Poisson |
NumPyro NegativeBinomial2 | (mean=mu, concentration=alpha) | the form the chapters use; d.mean, d.variance to check |
NumPyro GammaPoisson | (concentration=alpha, rate=alpha/mu) | gamma rate then Poisson count; the gamma rate, not the scale |
NumPyro NegativeBinomialProbs | (total_count=alpha, probs=mu/(alpha+mu)) | mean $=\text{total\_count}\cdot\text{probs}/(1-\text{probs})$; NegativeBinomial(total_count, probs=…) is a factory |
NumPyro NegativeBinomialLogits | (total_count=alpha, logits=log(mu/alpha)) | probs $=\text{sigmoid(logits)}$ |
SciPy nbinom(n, p), NumPy negative_binomial(n, p) | $n=\alpha$, $p=\alpha/(\alpha+\mu)$ | counts failures before $n$ successes; mean $=n(1-p)/p$ |
| statsmodels NB2, Stan, R | alpha $=1/\alpha$; neg_binomial_2(mu, phi=alpha); theta $=\alpha$ | the reciprocal convention is the usual surprise |
Diagnostics to fix: symptom, likely cause, remedy
The residual panel and the forecast checks in one table. Change one thing, refit, re-check. A clean panel is not proof; rolling-origin accuracy and calibration decide.
| Symptom | Likely cause | Remedy | See |
|---|---|---|---|
| Residuals run above then below zero for weeks; slow ACF decay; holdout bias after a recent change | A slope or level change the trend missed: too few or badly placed changepoints, $b$ too small, candidate range ends early | Changepoints near it (grid, PELT), a larger $b$, a regressor that explains the shift; check the last-28-day error | 7.17 7.17 7.10 |
| A step in residual vs time | A level shift (definition change, launch) | Step column $\mathbf 1[t\ge s]$ or a changepoint; fix the data if it is a tracking change | 7.8 7.17 |
| ACF spikes at lags 7, 14, 21 (daily data); waves with a fixed period | Weekly Fourier order too low, a missing second seasonality, or a weekday-specific holiday effect | Raise the order (validate it), add the period, add holiday × weekday terms | 7.17 7.11 |
| Waves of another length in residual vs time or the ACF | A cycle the model does not have | Another seasonal period, a regressor that carries the cycle | 7.17 7.2 |
| Fast geometric ACF decay from lag 1 | Short-term momentum from outside the calendar | Check the mean function first; then an AR(1) residual term, state-space or GP; confirm by horizon | 7.17 7.17 |
| Isolated huge residuals on repeating dates (the ACF stays quiet) | Missing holidays or a window that is too short | Holiday columns and windows; flag one-offs before PELT and fitting | 7.17 7.12 |
| Funnel in residual vs fitted; spread ratio well above 1.5–2; ACF of $e^2$ has memory | Heteroscedasticity: spread grows with the level or with recent shocks | Log or square-root transform, a Negative Binomial, or an explicit scale model; Student-t does not fix it | 7.17 |
| S-shaped Q-Q, high excess kurtosis, a few wild days | Heavy-tailed noise (or a missed event) | Look at the dates first; then a Student-t, with the learned $\nu$ as a diagnostic | 7.17 7.13 |
| Q-Q bent at one end, or two humps | Right skew (missing positive spikes) or a hidden group | Event columns, log/Box–Cox, or find the missing indicator | 7.17 |
| Pearson dispersion well above 1 for an NB; replicated variance too small in a PPC | $\alpha$ too large for some days, excess zeros or a missing component | Check the zero count in a PPC; regressor or mask for structural zeros; ZIP or hurdle; richer mean | 7.17 7.13 |
| Coverage below nominal, worse with the horizon; PIT is U-shaped | Overconfidence: missing trend uncertainty (no simulated future bends), a too-simple guide, a likelihood with light tails | Simulate future trend changes, a richer guide or an NUTS check, a Student-t or NB; recheck by horizon and slice | 7.16 7.16 7.14 |
| PIT hump; coverage above nominal at short horizons | Forecast spread larger than it needs to be (for example a noise scale inflated by a pattern the mean function missed) | Add the missing structure; check the noise scale and its prior | 7.16 |
| PIT piled to one side; bias in the rolling mean residual | A missed trend bend or level shift | Changepoints, regressor, shorter sliding window | 7.16 7.19 |
| Backtest much better than live accuracy; one feature “explains everything” | Leakage | Rebuild features as of each origin; refit scalers, PELT and tuning inside every fold; see the leakage checklist | 7.12 7.9 |
| Coefficients wobble or flip sign between fits; posterior correlation near $-1$ | Collinear columns or overlapping holidays; weak identification on a short history | Drop or combine, use anomalies, shrinkage priors, report the sum, a full-rank guide; lower the yearly order | 7.12 7.12 7.11 |
Leakage checklist
Leakage means using, in training, feature building, tuning or evaluation, information that would not exist at the forecast origin. Ask each question for every backtest you report.
| Area | Question to ask | See |
|---|---|---|
| Split | Is every training row earlier than every test row? No shuffled or blocked K-fold? A gap when features use windows? | 7.1 7.15 |
| Features | No leads, no centred windows, lags at least as old as the horizon, no data revised or backfilled after the origin? | 7.12 |
| Regressors | Is each future value known at the origin, or forecast? Do backtests use the values available at each origin, not the actuals? | 7.12 |
| Target leakage | Is any column computed from the target, or caused by it (same-day revenue for orders, returns logged later)? | 7.12 |
| Scaling and imputation | Mean, sd, fills and encoders fitted on the training window only, stored, and applied unchanged to later rows? | 7.12 7.12 |
| Changepoints | PELT, its penalty $\beta$, $\hat\sigma$ and the candidate list rerun inside every fold? | 7.9 |
| Tuning | Fourier order, Laplace scale $b$, windows and priors chosen on an inner validation block; the test block looked at once? | 7.12 7.15 |
| Events and flags | Were one-off events and holidays flagged without hindsight of the test period? | 7.12 |
| Exploration | Decomposition, unit-root tests and plots run on the training window only? | 7.2 7.4 |
| Refits | Is the whole pipeline refit at every rolling origin (warm starts only from earlier origins)? | 7.15 7.12 |
| Evaluation | Metric and baselines fixed before the results; the same origins for every model? | 7.15 |
| Smell tests | A suddenly large improvement, a feature that explains everything, or a gap between backtest and live? Treat as a leak until proven otherwise. | 7.12 |
Prophet's defaults, as the chapters state them
Prophet's defaults are fixed constants that every version of Prophet uses.
Prophet's defaults can change between versions, and they are tuned for Prophet's own scaling. The values below are the ones the chapters quote (Chapter 7.10 states the scaling as in recent Prophet 1.1.x releases, the rest is from its documentation); look up your installed version before you quote one, and remember your NumPyro model has settings of its own.
| Setting | Default as the chapters state it | What it means | See |
|---|---|---|---|
n_changepoints | 25 candidate changepoints | placed uniformly in the first part of the history (next row) | 7.8 7.8 |
changepoint_range | 0.8 | candidates only in the first 80% of the history: runway to project the trend and no overfitting of the very end | 7.8 7.8 |
changepoint_prior_scale | 0.05 | the Laplace scale $b$ of $\delta_j$ (Stan double_exponential) on Prophet's scaled data | 7.10 7.8 |
| Scaling of $y$ and $t$ | scaling="absmax": $y/\max|y|$; time mapped to $[0,1]$ over the history | so 0.05 is a fraction of the maximum per history length; in real units $\delta_{\text{day}}=\tilde\delta\max|y|/\text{days}$ | 7.10 |
| Trend priors | $k\sim N(0,5^2)$, $m\sim N(0,5^2)$ | on the scaled data | 7.7 7.10 |
| Noise prior | $\sigma\sim\text{HalfNormal}(0.5)$ | on the scaled data | 7.7 |
seasonality_prior_scale | 10 | Normal $N(0,\tau^2)$ on every seasonal weight, the same $\tau$ for all harmonics of a seasonality | 7.11 |
| Fourier orders | yearly 10 ($P=365.25$), weekly 3 ($P=7$), daily 4 ($P=1$) | columns: 20, 6, 8; weekly order 3 is exact for daily data | 7.11 |
| Automatic seasonalities | yearly with at least two years of history; weekly with at least two weeks of data spaced less than a week apart; daily only for sub-daily data | the data must be able to support each one | 7.11 7.11 |
holidays_prior_scale | 10 | Normal prior on holiday weights, on $y/\max|y|$: almost no regularization; lower_window, upper_window and a per-holiday prior_scale are optional | 7.12 7.12 |
add_regressor | standardize='auto' (not for binary columns); prior scale defaults to holidays_prior_scale | the regressor must be supplied for the history and every future date | 7.12 7.12 |
mcmc_samples | 0 | the default fit is a MAP computed with Stan's optimiser, not a posterior | 7.10 7.10 |
| Future trend changes | simulated at the historical rate; sizes Laplace with scale = mean of the fitted $|\delta_j|$ | this is where the trend part of Prophet's forecast interval comes from | 7.8 |
| Time for seasonal columns | days since 1970-01-01 (the trend uses rescaled time) | check which time your own Fourier columns receive | 7.11 |
Code sheet: diagnostics, trend and seasonality, likelihoods and metrics
Three short blocks that were run as written (Python with NumPy 2.5, SciPy 1.18, statsmodels 0.15, ruptures 1.1.9, NumPyro 0.22, JAX 0.11); the output is shown as comments. Random-number outputs depend on the seed and library versions, so expect the same shapes and orders of magnitude rather than identical digits on your machine.
Residual memory: ACF, PACF and Ljung–Box (7.3, 7.17)
import numpy as np
from statsmodels.tsa.stattools import acf, pacf
from statsmodels.stats.diagnostic import acorr_ljungbox
# a toy residual series with memory: AR(1) with phi = 0.6
rng = np.random.default_rng(8)
n, phi = 300, 0.6
e = np.zeros(n)
for t in range(1, n):
e[t] = phi * e[t - 1] + rng.normal()
r = acf(e, nlags=14) # statsmodels ACF: one overall mean, divide by n
p = pacf(e, nlags=3, method="ywm") # the method plot_pacf uses by default
band = 1.96 / np.sqrt(n) # white-noise band
print(np.round(r[1:4], 2), round(band, 3)) # lag 1-3 ACF and the band
print(np.round(p[1:4], 2)) # PACF: one big spike at lag 1, then small
lb = acorr_ljungbox(e, lags=[14]) # Ljung-Box over the first 14 lags
print(float(lb["lb_pvalue"].iloc[0]) < 0.001) # True: memory somewhere in the first 14 lags
# output: [0.59 0.35 0.24] 0.113
# output: [0.59 0. 0.06]
# output: True
Trend with changepoints, Fourier columns and PELT (7.8, 7.11, 7.9)
import numpy as np
# --- trend with changepoints: hinge form and Prophet (offset) form give the same curve ---
t = np.arange(0, 100.)
k, m = 0.5, 10.
s = np.array([30., 60.]) # changepoint times
delta = np.array([0.4, -0.7]) # slope changes
A = (t[:, None] >= s[None, :]).astype(float) # n x S matrix, A_ij = 1[t_i >= s_j]
g_prophet = (k + A @ delta) * t + (m + A @ (-s * delta))
g_hinge = k * t + m + np.maximum(t[:, None] - s[None, :], 0) @ delta
print(np.allclose(g_prophet, g_hinge), "final slope", round(k + delta.sum(), 2)) # True 0.2
# --- Fourier columns, Prophet order (sin, cos per harmonic) ---
def fourier(t, P, N):
cols = []
for n in range(1, N + 1):
cols += [np.sin(2 * np.pi * n * t / P), np.cos(2 * np.pi * n * t / P)]
return np.column_stack(cols)
td = np.arange(0, 28.)
print(fourier(td, 7, 3).shape, fourier(td, 365.25, 10).shape) # (28, 6) (28, 20)
# --- amplitude and phase of one harmonic: a cos + b sin = R cos(w t - phi) ---
a, b = 3., 4.
R, ph = np.hypot(a, b), np.arctan2(b, a)
w = 2 * np.pi / 7
tt = np.linspace(0, 7, 50)
print(R, np.allclose(a * np.cos(w * tt) + b * np.sin(w * tt), R * np.cos(w * tt - ph))) # 5.0 True
# --- PELT with ruptures (returns segment ENDS, the last one is n) ---
import ruptures as rpt
rng = np.random.default_rng(0)
y = np.r_[rng.normal(0, 1, 60), rng.normal(4, 1, 60)]
sigma = 1.4826 * np.median(abs(np.diff(y) - np.median(np.diff(y)))) / np.sqrt(2)
beta = 2 * sigma**2 * np.log(len(y)) # BIC-like penalty for the L2 cost
ends = rpt.Pelt(model="l2", min_size=2, jump=1).fit(y).predict(pen=beta)
print(ends) # [60, 120] -> one change, starting on day 60
# output: True final slope 0.2
# output: (28, 6) (28, 20)
# output: 5.0 True
# output: [60, 120]
Negative Binomial argument lists, metrics, PIT, coverage and CRPS (7.13, 7.15, 7.16)
import numpy as np
from scipy import stats
import numpyro.distributions as dist
# --- Negative Binomial: the same distribution, three argument lists (mu = 40, alpha = 5) ---
mu, alpha = 40.0, 5.0
d_np = dist.NegativeBinomial2(mean=mu, concentration=alpha) # NumPyro: mean, concentration
d_sp = stats.nbinom(n=alpha, p=alpha / (alpha + mu)) # SciPy: n = alpha, p = alpha/(alpha+mu)
d_pr = dist.NegativeBinomialProbs(total_count=alpha, probs=mu / (alpha + mu))
print(float(d_np.mean), float(d_np.variance)) # 40.0 360.0 (= mu + mu^2/alpha)
print(round(d_sp.mean(), 1), round(d_sp.var(), 1), float(d_pr.mean)) # 40.0 360.0 40.0
# --- metrics ---
rng = np.random.default_rng(0)
y_train = 100 + 20 * np.sin(2 * np.pi * np.arange(140) / 7) + rng.normal(0, 4, 140)
y = y_train[-14:]; f = y + rng.normal(0, 6, 14); ytr = y_train[:-14]
mae = np.mean(np.abs(y - f)); rmse = np.sqrt(np.mean((y - f) ** 2))
wape = np.abs(y - f).sum() / np.abs(y).sum()
m = 7
mase = mae / np.mean(np.abs(ytr[m:] - ytr[:-m])) # scale: in-sample seasonal-naive MAE, training data only
print(round(mae, 2), round(rmse, 2), round(100 * wape, 1), round(mase, 2))
def pinball(y, q, tau): # quantile loss at level tau
d = y - q
return np.mean(np.maximum(tau * d, (tau - 1) * d))
def crps_samples(x, y): # E|X - y| - 0.5 E|X - X'| from draws
return np.mean(np.abs(x - y)) - 0.5 * np.mean(np.abs(x[:, None] - x[None, :]))
# --- forecast draws ytil[S, H] -> PIT, coverage, CRPS ---
S, H = 2000, 14
ytil = rng.normal(f, 6, size=(S, H)) # pretend posterior-predictive draws
pit = (ytil < y).mean(axis=0) # share of draws below the actual value
lo, hi = np.quantile(ytil, [0.05, 0.95], axis=0)
print(round(((y >= lo) & (y <= hi)).mean(), 2), round(float(np.mean([crps_samples(ytil[:, j], y[j]) for j in range(H)])), 2))
med = np.median(ytil, axis=0)
print(round(pinball(y, med, 0.5), 2), round(0.5 * np.mean(np.abs(y - med)), 2)) # pinball at tau = 0.5 is half the absolute error
# output: 40.0 360.0
# output: 40.0 360.0 40.0
# output: 3.35 4.34 3.3 0.74
# output: 0.93 2.48
# output: 1.69 1.69
Interview question bank
50 questions an interviewer could ask about this guide's material while you walk them through your forecasting model (and, where it fits, your A/B framework). They complement the forty-question drill of Chapter 7.20 rather than repeat it, and most come from the chapters' “Say it right” boxes: shuffled splits, fitted values versus forecasts, stationarity and Prophet, the trend as “true growth”, the continuity offset, what PELT does and does not give you, sparse Laplace posteriors, Fourier order, leakage, Negative Binomial argument lists, prediction versus credible intervals, MAPE, coverage and the PIT, residual memory, alternatives, and seeds, drift and NaNs in production.
How to use this bank. Read the question, answer it out loud in about a minute, and only then open the model answer. A strong answer usually has four parts: (1) a one-sentence definition in plain words, (2) a small number or picture, (3) where it lives in your project, (4) the trap you avoid. Where an answer talks about your own code (which Negative Binomial class you call, how PELT is configured, your guide-size threshold, your stopping tolerance), the answers are written as “if my code …” or “I would check …”: look at what your code actually does before you quote it. If an answer feels shaky, follow its link back to the chapter.
Time-series basics, autocorrelation, stationarity and baselines (7.1–7.6)
1. A teammate shuffled a year of daily demand rows and then split them 80/20. Their MAE is 5. Do you trust it?
No. A time series is not a bag of IID rows: shuffling keeps the histogram, the mean and the sd but destroys the trend, the weekly pattern and the link between today and yesterday. A random split also puts days from after each test day into training, so the model only has to fill gaps between neighbours: it measures interpolation, and the more flexible the model, the more it is flattered. I would train on the days up to an origin $T$, test on the days after it, repeat over many origins (rolling origin), refit everything that learns from data inside each window (scalers, changepoint detection, tuning), and put seasonal naive with $m=7$ next to the model on the same windows. Chapter 7.1 · Chapter 7.1 · Chapter 7.15
2. Explain forecast origin, horizon and information set. For your forecasting model, are the holiday calendar, next week's promotion plan and tomorrow's weather in the information set?
The origin $T$ is the last day whose value I may use, the horizon $h$ says how many days ahead, and $\hat y_{T+h\mid T}$ is the forecast made at $T$ for $T+h$. The information set is everything known at $T$: past values, past regressors and future inputs that are genuinely known in advance. The holiday calendar is in it, because it is a function of the date (as long as my holiday table covers the horizon). A promotion plan is in it if it was fixed at the origin and is reliable. Tomorrow's weather is not: it has to be forecast (and its error enters my forecast), lagged beyond the horizon, or dropped. And fitted values on training days are not forecasts, because the model has already seen those days. Chapter 7.1 · Chapter 7.12
3. Why is the standard error of an average of daily values too small when the days are correlated? How big is the effect for $\phi=0.5$, and why does it matter for your forecasting model?
Neighbouring days repeat information, so $n$ correlated days are worth fewer independent ones. For an AR(1) with $\phi=0.5$ the effective sample size is $n(1-\phi)/(1+\phi)=n/3$, and the true standard error is $\sqrt3\approx1.7$ times the IID value $\sigma/\sqrt n$. In my forecasting model the likelihood multiplies one term per day, so if the residuals still have memory it believes it has 365 independent facts: the posterior for the trend slope, the $\delta_j$ and the seasonal weights is too narrow, runs of same-signed residuals can be mistaken for trend changes, and intervals for weekly totals are too tight. I look at the residual ACF and Ljung–Box; the cures are HAC standard errors or a block bootstrap for a quick estimate, or modelling the memory (an AR error term). Chapter 7.1 · Chapter 7.3 · Chapter 7.3
4. You plot the ACF of a daily series: a slow, almost straight-line decay from near 1, with bumps at lags 7, 14 and 21. What do you conclude, and what do you do before modelling?
The slow straight-line decay is the signature of a trend or a random walk (the series is not stationary), and the bumps at 7, 14, 21 are a weekly pattern. The ACF of a raw series mixes trend, season and memory, so I would not read an AR order from it. For an ARIMA-style model I would difference (once, and seasonally with lag 7 for SARIMA) and look at the ACF and PACF again. For my Prophet-style model I would not difference at all: $g(t)$ carries the trend, Fourier terms carry the week, and I read the ACF of the residuals. One caution: ADF and KPSS on a series that bends at changepoints often say “unit root”, which is not a reason to difference before a changepoint model. Chapter 7.3 · Chapter 7.3 · Chapter 7.4
5. “ARIMA is only for stationary data, and Prophet needs stationary data too.” Correct that.
Both halves are shaky. ARIMA with $d\ge1$ is itself a non-stationary model: it differences the series $d$ times, models the differences as a stationary ARMA and integrates back. A Prophet-style model never differences; it puts the non-stationary parts into $g(t)$, $s(t)$, $h(t)$ and $X\beta$ and assumes only the leftover noise is stationary and independent. So I do not need to make the series stationary, but I still use the idea: to check the residuals (flat rolling mean and sd, an ACF like white noise, Ljung–Box) and to compare the two worldviews. Trend-stationary means shocks fade and noise bands stay bounded; a unit root means shocks persist and bands grow like $\sqrt h$, and long-horizon coverage in a backtest settles which is closer to the truth. Chapter 7.4 · Chapter 7.6
6. Your model's holdout MAE is 6.2. Is that good?
Compared with what? An error number needs a baseline scored the same way. I would score seasonal naive ($m=7$, same weekday last week), naive, drift and an ETS or Holt–Winters model on the same rolling origins and horizons, and report the skill $1-E_{\text{model}}/E_{\text{baseline}}$ or MASE against the strongest. If seasonal naive scored 9.0 on the same windows (a made-up number), the model removes about $1-6.2/9.0\approx31\%$ of the baseline's error. If it cannot clearly beat the strongest simple baseline, its extra structure is not earning its keep. Chapter 7.5 · Chapter 7.15
Model structure: components, identifiability and the design matrix (7.2, 7.7)
7. Walk me through $y_t=g(t)+s(t)+h(t)+X_t\beta+\epsilon_t$: the job and the assumption of each layer, and what you would see in the residuals if one were missing.
$g(t)$ is the slow level and growth: piecewise linear with slope changes $\delta_j$, straight between bends, and the last slope continues. $s(t)$ repeats with a fixed known period (7 and 365.25 days) built from Fourier terms and has the same shape every cycle. $h(t)$ is bumps on known dates, indicator columns with windows, the same effect each time. $X_t\beta$ is measured outside drivers with a constant linear effect, and their future values must be known or forecast. $\epsilon_t$ is described by the likelihood (Normal, Student-t, Negative Binomial), independent given the mean. A missing layer leaves its fingerprint: a slow bend means changepoints, a wave with a fixed period means Fourier order, spikes on dates mean holidays, a funnel means the likelihood, day-to-day memory means an AR term. Chapter 7.7 · Chapter 7.7
8. Is the trend component the true underlying growth of the business? What if one of your regressors rises with time?
No. The trend is the model's estimate of the slow part of the mean after the other layers took their share, under its assumptions (piecewise linear, chosen changepoints, chosen priors): a useful summary, not a measurement. The total $\mu_t$ is usually well determined; the split between layers is only as good as their columns are distinct. A regressor that rises with time (cumulative users, a growing budget) competes with the trend for the same slow rise, and then the Laplace prior on $\delta_j$ and the prior on $\beta$ decide the split as much as the data do. Before I show a component breakdown I check the posterior correlation between the trend weights and that regressor's weight, and I say “the total is well determined, the split is uncertain”. Chapter 7.7 · Chapter 7.2
9. Does your forecasting model use yesterday's value? What follows for how it reacts to a sudden change?
No. It is a regression on time, the calendar and known drivers; there is no $y_{t-1}$ column unless I add lags as regressors. So yesterday's surprise does not move tomorrow's forecast: the model adapts only when it is refitted, and then only through its layers (a new changepoint, a new holiday weight). The price is that short-term momentum stays in the residuals, so I check the residual ACF and Ljung–Box. If lag-1 correlation is large and short horizons matter, I would add an AR(1) error term, whose correction $\varphi^he_T$ fades with the horizon, and compare with ARIMA errors by horizon. At long horizons memory has faded and the structure wins. Chapter 7.7 · Chapter 7.3 · Chapter 7.6
10. How do you decide between additive and multiplicative seasonality, and what does a log link have to do with it?
I ask whether the seasonal swing (or the noise spread) grows with the level. A constant swing in orders is additive, $y=g+s$; a swing that is a constant percentage is multiplicative, $y=g(1+s)$, or equivalently additive on $\log y$. The parameter count is the same, only the structure differs. A log link does the same job automatically: with $\mu_t=e^{\eta_t}$ every additive term in $\eta_t$ multiplies the mean, so a holiday coefficient of $\ln1.3$ is a 30% bump at any level, whereas with a softplus link terms stay roughly additive in orders when the mean is large. My model is written additively, so I would check the line in my code that turns the layers into the mean of the count likelihood, because it decides whether coefficients are percentages or orders and what scale a prior needs. Chapter 7.2 · Chapter 7.7 · Chapter 7.13
11. In what sense is the whole model “one regression with structured columns”? What is fixed in advance, what is learned, and how do you count the latent dimension?
The mean is $\mu=X\theta$ with $X=[1,t\mid\text{changepoint ramps}\mid\text{Fourier}\mid\text{holiday flags}\mid\text{regressors}]$. Fixed before fitting: the changepoint locations (grid plus PELT), the periods and Fourier orders, the holiday dates and windows, and which regressors to use. Learned: the weights $(m,k,\delta,\text{seasonal},\text{holiday},\text{regressor})$ and the noise parameters, each block with its own prior (Laplace on $\delta$, Normal on the rest). Because the mean is linear in the weights, any inference method can plug in. The column count is $2+S+2N_{\text{year}}+2N_{\text{week}}+\#\text{holiday columns}+R$: for example with $S=25$ candidates, orders 10 and 3, 10 holiday columns and 2 regressors that is $2+25+20+6+10+2=65$ weights plus $\sigma$ (or $\sigma,\nu$ or $\alpha$). That count is the latent dimension $d$ that the full-rank versus low-rank guide choice depends on. Chapter 7.7 · Chapter 7.18
Changepoints, PELT and Laplace priors on trend changes (7.8–7.10)
12. Derive $\gamma_j=-s_j\delta_j$ and say why the offset is needed.
Write the trend as $\big(k+\sum_{s_j\le t}\delta_j\big)t+\big(m+\sum_{s_j\le t}\gamma_j\big)$. At $s_j$ the slope gains $\delta_j$ and the intercept gains $\gamma_j$. Just before $s_j$ the line is $(k+K)s_j+m+G$ and just after it is $(k+K+\delta_j)s_j+m+G+\gamma_j$. Continuity needs the two to be equal, and everything cancels except $0=\delta_js_j+\gamma_j$, so $\gamma_j=-s_j\delta_j$. Without it the line would jump by $s_j\delta_j$, a size that depends on where time zero is. Substituting back gives the hinge form $g(t)=kt+m+\sum_j\delta_j(t-s_j)_+$, the same function, which is also how I would write a unit test: compute both forms on the same $t$, $k$, $m$, $\delta$, $s$ and compare. Chapter 7.8 · Chapter 7.8
13. What are the three ways to handle changepoint locations? Which does your model use, and what is its main weakness?
A fixed grid of many candidates with a sparse prior on the slope changes (simple, linear, JIT-friendly, but a bend between two candidates is approximated by the nearest ones); data-driven detection such as PELT (locations chosen from the data, then fixed); and Bayesian latent changepoints (locations are parameters and forecasts average over them: honest, but discrete, multimodal and expensive). My model combines the first two: a Prophet-like grid plus PELT-detected locations, with a Laplace prior on every $\delta_j$. The weakness is that the locations are inputs, so the posterior is conditional on them and does not contain uncertainty about where the trend changed. The grid with shrinkage partly compensates by letting neighbouring candidates share a bend. Chapter 7.8 · Chapter 7.8 · Chapter 7.8
14. Explain PELT in a minute. Does it find the true changepoints?
A segment cost $C$ says how well one simple model fits a stretch (L2: squared distance to the segment mean; Normal: also the variance), and a penalty $\beta$ is the price of one more cut. PELT minimises $\sum C(\text{segments})+\beta m$. Optimal partitioning does that exactly with dynamic programming, $F(t)=\min_s[F(s)+C(y_{s:t})+\beta]$, at $O(n^2)$ cost. PELT prunes every start $s$ with $F(s)+C(y_{s:t})\gt F(t)$: it trails by more than one $\beta$ and, because cutting never raises the cost, can never win later. The answer is the same, in roughly linear time when changes are frequent. But “exact” refers to the penalised objective, not the truth: the wrong cost or a badly scaled $\beta$ gives an exact but wrong segmentation, small $\beta$ over-detects and large $\beta$ under-detects. Chapter 7.9 · Chapter 7.9 · Chapter 7.9 · Chapter 7.9
15. A trend change is a change of slope. What series and cost should PELT see, and what goes wrong with L2 on the raw series? What would you check in your own code?
L2 on a trending series sees every day as shifted from the previous segment's mean, so it cuts the trend into a staircase of fake changepoints. Better inputs: L2 on the day-to-day differences (a slope change becomes a mean shift of size $\delta$, with noise variance $2\sigma^2$, so the penalty must use $2\hat\sigma^2$), a linear (regression) cost, or a series with trend and season removed first; weekly seasonality left in also creates fake cuts. In my own code I would look up which series PELT sees (raw scaled, differenced, deseasonalised), which model=, how the penalty is set and whether it is tied to the noise level, min_size, and jump (default 5, which rounds dates to multiples of 5). Whatever the choice, I should be able to say it plainly. Chapter 7.9 · Chapter 7.9 · Chapter 7.9 · Chapter 7.9
16. Does your Bayesian forecast include the uncertainty about where the changepoints are? How do you limit the damage?
No. The posterior is $p(\theta\mid y,\hat\tau)$, conditional on PELT's output, not $p(\theta\mid y)=\sum_\tau p(\theta\mid y,\tau)p(\tau\mid y)$. It ignores how spread out the plausible locations are, so intervals near uncertain changepoints (the latest slope, the forecast) are too narrow, and because the same data pick the locations and then estimate the changes there, selected changes look too big. To limit the damage I use PELT points only as candidates next to a dense grid, so the Laplace prior can shrink a wrong one; rerun PELT over a range of penalties and on perturbed series to see which points survive; refit with alternative candidate sets and compare forecasts; and check interval coverage in rolling-origin backtests. A fully Bayesian alternative treats the locations as latent, at a much higher computational cost. Chapter 7.9 · Chapter 7.8
17. Why a Laplace prior on the slope changes rather than a Normal? Is the posterior sparse?
I believe most candidate slope changes are tiny and a few are large. The Laplace has a sharp peak at 0 and exponential tails, $P(|\delta|\gt c)=e^{-c/b}$, and pulls with a constant force $1/b$: small changes are squashed while big ones are spared. A Normal pulls in proportion to $\delta$ and has light tails. At the MAP a Laplace prior is a lasso with $\lambda=\sigma^2/b$, which can set slope changes exactly to zero. But I fit a posterior, a continuous distribution with $P(\delta_j=0\mid y)=0$: draws, means and medians are shrunk, never exactly zero. So it is “sparse-ish”, not sparse, and I report the posterior slope over time with a band rather than a list of active changepoints. Chapter 7.10 · Chapter 7.10 · Chapter 7.10
18. After a real shift in growth your forecasts lag for a while and then catch up. Why, and what are your options?
Evidence for a slope change builds up slowly: $n$ days after a bend the standard error of $\delta$ is $\sigma/\sqrt{\sum i^2}\approx\sigma\sqrt3/n^{3/2}$, and the Laplace MAP stays at 0 until $|x|$ beats $se^2/b$. A candidate near the end of the history has few days after it and is always heavily shrunk. So the lag is a design consequence, not a bug. The options trade stability for speed: a larger $b$ (more flexible, more noise-chasing), explicit candidates or a step column at a known event, a changepoint range that reaches closer to the end, or a local model (ETS, state space) that adapts quickly or runs next to mine. I would compare them by rolling-origin error, especially the last-28-day error. Chapter 7.10 · Chapter 7.8 · Chapter 7.18
19. What does Prophet's changepoint_prior_scale of 0.05 mean in your own units, and how would you choose $b$?
In Prophet 0.05 is the Laplace scale $b$ of $\delta_j$ after $y$ is divided by its maximum absolute value and time is mapped to $[0,1]$ over the history. So $E|\delta_j|=0.05$ means a typical slope change of 5% of the maximum per history length; in real units $\delta_{\text{per day}}=\tilde\delta\,\max|y|/\text{days}$. For two years (730 days) and a maximum near 1 000 orders per day that is about $0.05\cdot1000/730\approx0.07$ orders per day per day, roughly 25 orders per day per year. In my NumPyro model I would first check how $y$ and $t$ are scaled (raw, standardised, log, days or $[0,1]$), because the number only carries Prophet's meaning under Prophet's scaling. Then I would choose $b$ on a log grid by rolling-origin error and show that $b/3$ and $3b$ tell the same story. Chapter 7.10 · Chapter 7.10
Seasonality, holidays, regressors and leakage (7.11–7.12)
20. Why does each Fourier harmonic get both a sine and a cosine column?
A sine alone has its peak at a fixed place, so it could not put the weekly peak on any chosen day. The identity $a\cos\omega t+b\sin\omega t=R\cos(\omega t-\varphi)$ with $R=\sqrt{a^2+b^2}$ and $\varphi=\operatorname{atan2}(b,a)$ says two linear weights encode a size and a timing, so the model learns the peak day and stays a linear regression. If $a$ and $b$ get independent $N(0,\tau^2)$ priors, the implied phase is uniform (no favourite day) and the amplitude has typical size about $1.25\tau$: a prior on “how big” with no opinion on “when”. When I explain a fitted season I plot $s(t)$ over one period or compute $R$ per posterior draw, because a single weight means nothing on its own. Chapter 7.11
21. How many parameters do weekly order 3 and yearly order 10 add? Why is weekly order 3 “exact” for daily data, and why can daily data never show a daily seasonality?
Each harmonic is a sine and a cosine, so order $N$ adds $2N$ weights: weekly order 3 gives 6 and yearly order 10 gives 20, 26 in total before trend, holidays and regressors. With $P$ samples per period only $P-1$ centred values are free, so the exact ceiling is $\lfloor P/2\rfloor$: for a week of daily data, $N=3$ reproduces any weekly shape and spans the same shapes as six weekday dummies. Higher harmonics alias: harmonic 4 of period 7 duplicates harmonic 3, so extra columns make $X^\top X$ singular and add no detail. The same Nyquist argument says daily data (half a cycle per day at most) cannot show a once-a-day cycle: it aliases to a constant, which is why Prophet only switches daily seasonality on for sub-daily data. Chapter 7.11 · Chapter 7.11 · Chapter 7.11
22. You have only eight months of daily history. Should the model have yearly seasonality?
Probably not as it stands. With less than one period of history a yearly wave and the trend are nearly the same column (the chapter's example has a variance inflation factor near 1 400 for 90 days against a yearly wave), the posterior becomes a ridge and the result is driven by the prior, with wild forecasts a year ahead. Prophet itself only switches yearly seasonality on automatically with at least two years of history. My options: drop yearly seasonality or lower its order, give its weights a tight prior and say openly that the prior is doing the work, borrow the shape from a related series or from outside knowledge, or wait for more data. I would also look at the posterior correlation between the trend slope (or the early $\delta_j$) and the yearly weights. Chapter 7.11
23. Black Friday has occurred three times in your data. How do you estimate and report its effect?
Three occurrences means very little information, so I use a shrinkage prior $\beta_j\sim N(0,\tau^2)$. The posterior mean is $w\bar r_j$ with $w=\frac{n_j/\sigma^2}{n_j/\sigma^2+1/\tau^2}$, a precision-weighted compromise between the data average and zero (the MAP is ridge with $\lambda=\sigma^2/\tau^2$): rare, noisy holidays are pulled in, well-observed ones barely move. I would report something like “about +19% with a wide interval, moving by a few points if I halve or double the prior scale” rather than a bare +23%. I also check the prior scale in the units the model sees (Prophet's default 10 on $y/\max|y|$ is very weak), and a hierarchical or Laplace prior over holidays is an alternative. The syllabus lists sparse holiday effects among the prior-sensitivity checks. Chapter 7.12 · Chapter 7.12
24. Two holidays always fall on the same days. How do you split the credit?
If two indicator columns are on exactly the same days they are perfectly collinear: only $\beta_A+\beta_B$ is identified. With priors the posterior is proper, but in the overlapping direction it equals the prior, so equal prior scales give an even split and a smaller scale on A gives B more credit. Days where only one is on identify each effect. I count the “only A” and “only B” days, merge the two when they always coincide, add an explicit joint column if the combined effect is not the sum, and report the split as prior-driven. A posterior correlation near $-1$ means not identified, even if the optimiser “converged”; a full-rank guide can show it, a mean-field guide cannot. Chapter 7.12
25. One of your regressors is not known in advance. What do you do at forecast time?
First I ask where its future values come from at the origin. Calendar columns, planned prices and promotions and lags at least as old as the horizon are known in advance. Otherwise I forecast the regressor and let the predictive distribution average over it: draw regressor paths, then $y$ given each path; for a linear term the extra variance is $\beta^2Var(X-\hat X)$. The other options are to lag it beyond the horizon or to drop it. In a backtest each origin must see the values that were available at that origin (archived forecasts), not later actuals: using actuals is an oracle, which is leakage. If I do not propagate the regressor's uncertainty I say my intervals are conditional on its values. Chapter 7.12 · Chapter 7.14
26. A standardized temperature coefficient is 20. What does it mean, and how should the scaling be done?
With $z=(x-\bar x_{\text{train}})/s_{\text{train}}$, a coefficient of 20 means about 20 more orders per one training standard deviation of temperature, other terms fixed; if the sd is 4 °C that is about 5 orders per degree ($\beta_{\text{raw}}=\beta_z/s$). Binary regressors are usually left as 0/1. I standardize so that one prior scale is fair to every regressor and the optimiser sees similar scales, and the mean and sd must come from the training window only, be stored in the run manifest and be applied unchanged to test and future rows; a scaler fitted on the full series leaks the future. The coefficient is an association, not a causal effect, unless something like a randomized promotion supports that reading. Chapter 7.12
27. Name the doors through which the future can leak into a forecasting backtest, and the ones that are specific to your pipeline.
Features: leads, centred windows, lags shorter than the horizon, data revised or backfilled later. Regressors: actual future values instead of what was known (an oracle). Target leakage: columns computed from $y$ or caused by it. Pipeline: scalers, imputation, feature or order selection, hyperparameters and prior scales, PELT with its penalty and $\hat\sigma$, event flags added with hindsight, all of which must be fitted on the training window and refitted at every origin. Evaluation: shuffled folds, tuning on the test block. In my forecasting model the specific doors are PELT (run once on the full history, it lets the test period place the changepoints), the scaling of $y$ and $X$, the holiday table, and the Fourier orders and prior scales. Smell tests: a suddenly great backtest, one feature that explains everything, a big gap between backtest and live. Chapter 7.12 · Chapter 7.12 · Chapter 7.9
Likelihoods: Student-t, Negative Binomial, links and comparison (7.13)
28. Derive the Negative Binomial from a Poisson with a wobbling rate, and say what $\alpha$ does.
A Poisson forces $Var=\mu$, but real daily counts wobble more because the underlying rate is not constant. Let the rate be Gamma with mean $\mu$ and variance $\mu^2/\alpha$ (shape $\alpha$, rate $\alpha/\mu$) and let the count be Poisson given the rate. By the law of total variance $Var(y)=E[\lambda]+Var(\lambda)=\mu+\mu^2/\alpha$: a Poisson part plus a rate part, and $E[y]=\mu$. As $\alpha\to\infty$ the rate stops wobbling and I get the Poisson; small $\alpha$ is very noisy, and the relative noise never goes below $1/\sqrt\alpha$. For $\mu=40$, $\alpha=5$ the sd is 19.0 and the 90% interval is [14, 75] against [30, 51] for a Poisson. In the forecasting model the one shared $\alpha$ sets how wide the count intervals are. Chapter 7.13 · Chapter 7.13
29. Your code builds NegativeBinomialProbs(total_count=5, probs=0.8889). What are the mean and variance, and how would you write the same law as NB2 and in SciPy?
For this class the mean is $\text{total\_count}\cdot\text{probs}/(1-\text{probs})=5\cdot0.8889/0.1111\approx40$ and the variance is $\text{total\_count}\cdot\text{probs}/(1-\text{probs})^2\approx360$. That is NB2 with $\mu=40$ and $\alpha=\text{total\_count}=5$, because $\text{probs}=\mu/(\alpha+\mu)=40/45$ and $Var=\mu+\mu^2/\alpha=40+1600/5=360$. As NB2: NegativeBinomial2(mean=40, concentration=5). In SciPy: nbinom(n=5, p=5/45), with $p=0.1111$, not 0.8889: passing the wrong direction silently gives a mean of 0.625 instead of 40. statsmodels' alpha would be $1/5=0.2$. After building any distribution I print its mean and variance, because these mistakes raise no error. Chapter 7.13 · Chapter 7.13
30. A holiday adds 30% at a baseline of 100 orders. How does the coefficient differ under a log link and a softplus link, and what happens at a baseline of 10?
With a log link $\mu=e^{\eta}$ the effect multiplies: the coefficient is $\ln1.3\approx0.26$ and a baseline of 100 becomes 130, a baseline of 10 becomes 13. With a softplus link $\mu=\log(1+e^{\eta})$, which is about $\eta$ when $\eta$ is large, the effect adds: to get 130 from 100 the coefficient is about 30, and then at a baseline of 10 the same coefficient gives about 40. So the link decides whether holiday coefficients are percentages or orders, whether a Fourier seasonality is multiplicative, what scale a prior needs ($N(0,1)$ means ×0.37 to ×2.7 under log but about $\pm1$ order under softplus), and how a long-horizon forecast behaves: compounding with exp, roughly linear with softplus. I would find the line in my code that turns the layers into the count mean, and apply the link per posterior draw before averaging. Chapter 7.13
31. If you give $\nu$ a prior and learn it, what does its posterior tell you, and why is it hard to learn?
$\nu$ is a tail-heaviness dial, not $n-1$. Mass on small values says the residuals have heavy tails; mass spread over large values (or a very large $\nu$) says the Normal was fine. But the data pin down small $\nu$ reasonably well and large $\nu$ only loosely (it can hardly tell 30 from 200), and $\nu$ trades off with the scale: the sd is $\sigma\sqrt{\nu/(\nu-2)}$ for $\nu\gt2$, so the two are correlated in the posterior, which is one reason a full-rank or low-rank guide can describe it better than mean-field. I also remember that the t does not delete outliers: a glitch day gets a tiny weight (0.003 in the chapter's example) but stays in the likelihood, so I list the lowest-weight days and ask of each whether it is a bug, a real event to model, or a genuine tail. Chapter 7.13 · Chapter 7.13
32. How would you compare a Normal, a Student-t and a Negative Binomial version of the forecasting model fairly?
On held-out later days, with a rolling-origin evaluation: the same origins and horizons for all three versions, scored with proper scores (CRPS and the log predictive density) plus interval coverage, followed by a posterior predictive check of the feature that motivated each choice (tails for the t, variance and zero days for the NB). The currency matters: a Normal or Student-t gives a density on the real line and an NB gives a probability mass on integers, and dividing $y$ by $c$ raises every density log-score by $\log c$ per day (multiplying lowers it by the same amount), so I do not set a Normal's density log-score against an NB's mass log-score directly. CRPS from posterior-predictive samples is the safe common ruler. In-sample criteria such as AIC, WAIC and PSIS-LOO assume exchangeable data, so for time series the genuine rolling-origin score is more trustworthy. Chapter 7.13 · Chapter 7.16
Forecast distributions, backtests and scoring (7.14–7.16)
33. Describe your rolling-origin evaluation. What is refitted at each origin, how many fits does it cost, and expanding or sliding window?
I choose a first window $T_0$, a horizon $H$ and a step $s$; for each origin $T=T_0,T_0+s,\dots$ I refit everything on that window (scaling, changepoints from the grid and PELT, the SVI fit), forecast the next $H$ days and store the errors with $T$ and $h$. That is $K=\lfloor(N-T_0-H)/s\rfloor+1$ origins, so $K$ full fits; starting each fit from the previous origin's parameters is allowed because they only saw earlier data. An expanding window uses all history (less variance, best for a stable process, old regimes stay in); a sliding window uses the last $W$ days (it forgets old slopes when the trend changes, but is noisier). The first window should contain a couple of yearly cycles if I model yearly seasonality, and the step should not be a multiple of 7, or horizon gets tied to weekday. I report error by horizon with its spread across origins. Chapter 7.15 · Chapter 7.15
34. Why do you avoid MAPE for demand series, and what do you report instead?
MAPE is $\frac{100}{n}\sum|e_i|/|y_i|$: undefined if any actual is 0, exploding when actuals are small, capped at 100% for under-forecasts but unbounded for over-forecasts, so it favours low forecasts (the best constant under MAPE is below the median). “100 minus MAPE” is not an accuracy. Count-like demand with small values and zero days is exactly where it fails. I report MAE or RMSE in the original units, WAPE $=\sum|e|/\sum|y|$ (zero-safe, weights big days more, which usually matches the business), and MASE to anchor to the naive forecast. If I report the predictive median, MAPE would reward an even lower number, which is gaming the metric, not improving the model. Chapter 7.15
35. What does a MASE of 0.8 mean, and when can it mislead?
MASE is the test MAE divided by $d_m=\frac1{T-m}\sum_{t\gt m}|y_t-y_{t-m}|$, the in-sample MAE of the (seasonal) naive forecast on the training data ($m=7$ for daily data with a weekly pattern). So 0.8 means my errors are 20% smaller than the average in-sample naive error. It is scale-free, safe with zeros and can be averaged across series. The trap is that $d_m$ is a one-step ruler while my test errors reach $H$ days ahead: even the naive forecast usually scores above 1 on a long horizon, so “below 1” is not the same as “beats naive”. To claim that, I compute the naive forecast's own MASE or MAE on the same test windows and report the relative MAE, per horizon. Chapter 7.15 · Chapter 7.5
36. How do you answer “what is the chance we exceed capacity next week?” from your model?
From the draws array $\tilde Y$ ($S$ simulated futures by $H$ days): for each draw (a row) take the maximum over the week, then count the share of rows where it exceeds the capacity $c$. That is not the product $1-\prod(1-p_j)$ of daily probabilities and not a sum of daily quantiles, because all days of one row share the same parameters, so the days are linked; quantiles do not add and variances need the covariances. The Monte Carlo error is $\sqrt{p(1-p)/S}$: about 0.006 for $p=0.19$ and $S=4\,000$. So I save the draws, not just a table of daily bands, and in the A/B framework I keep the joint draws of $(\theta_A,\theta_B)$ for the same reason. Chapter 7.14 · Chapter 7.14
37. Your 90% intervals contain only 70% of the actuals in the backtest. How do you diagnose it?
First I check the luck range, $\pm1.96\sqrt{p(1-p)/n}$, remembering that neighbouring origins overlap so the real noise is larger. Then I slice the hits by level (50%, 80%, 95%), horizon, weekday, holiday versus normal day and regime. Coverage that falls as the horizon grows points at missing trend uncertainty: the forecast holds the last slope instead of simulating future bends, or an over-simple guide (mean-field) understates parameter correlations; I would simulate future trend changes, try a richer guide or check against NUTS on a subset. Low coverage at every horizon points at noise that is too small or tails that are too light (Student-t or NB, a variance model). Clustered misses on holidays point at the holiday columns. I also confirm the intervals come from posterior-predictive samples with observation noise, not from a trend-only band. Chapter 7.16 · Chapter 7.16 · Chapter 7.14
38. What does a U-shaped PIT histogram tell you, and why is a flat one not enough?
The PIT $u_t=F_t(y_t)$ is, from samples, the share of predictive draws below the actual value; for honest forecasts it is uniform. A U-shape means too many outcomes in the tails, so the forecasts are overconfident (too narrow); a hump means too wide; a pile to the left or right means the forecasts are biased high or low. A flat histogram is necessary but not sufficient: always issuing the long-run distribution is calibrated yet not sharp. So I add sharpness through CRPS or the interval score (maximise sharpness subject to calibration), judge the bars against their luck range, and use a randomized PIT for counts. Chapter 7.16 · Chapter 7.16
39. Which score do you put in the headline for the forecasting model: RMSE, CRPS or the log score?
I would lead with CRPS averaged over rolling origins and horizons, with the log score as a second opinion and coverage and the PIT for diagnosis. CRPS is in the units of $y$, strictly proper, equals the absolute error for a point forecast, is computed straight from posterior-predictive samples whatever the likelihood, and grows only linearly with distance, so one extreme day does not dominate. The log score needs a density and punishes tail surprises hard, which is why it shows so directly that a Student-t beats a Normal on outlier-rich data. RMSE of the mean ignores the spread completely: a change that improves RMSE but worsens CRPS has made the centre a little better and the uncertainty worse, which usually matters more for decisions. Chapter 7.16 · Chapter 7.16 · Chapter 7.16
40. Stock-outs cost 5 per unit and overstock costs 1 per unit. Which number from the forecast do you act on?
The critical ratio is $\tau=c_u/(c_u+c_o)=5/6\approx0.83$, so I stock the 83rd percentile of the predictive distribution of demand over the lead time (for a multi-day lead time, the distribution of the total over the paths), not the mean. In the chapter's example, with demand $\sim$ NB(120, 12), that gives 154 units at an expected cost of 59.0 against 86.4 for stocking the mean. This needs a calibrated right tail: a Normal likelihood on counts or intervals that are too narrow mislead exactly here, so I grade such quantiles with the pinball loss at $\tau=0.83$ and check that about 83% of outcomes fall below them. Chapter 7.14 · Chapter 7.16
Residual diagnostics (7.17)
41. Your residual ACF is not flat. How do the different shapes change what you do?
I match the response to the shape. A geometric decay from lag 1 is short-term momentum: I check the mean function first, then an AR(1) residual term (correction $\varphi^he_T$), a state-space model or a GP. Spikes at 7, 14, 21 mean the weekly shape is not captured: a higher weekly Fourier order or weekday-specific effects (a holiday on a particular weekday). A slow, almost linear decay with a wandering time plot means a slow component is missing: more or better-placed changepoints, a weaker shrinkage (larger $b$), a regressor. Long waves point at another cycle. A quiet ACF with isolated huge residuals on repeating dates means missing holidays, because the ACF is blind to isolated spikes. I choose the Ljung–Box lag to reach the season (14 for daily data with a weekly pattern) and confirm that a fix helps the horizons that matter. Chapter 7.17 · Chapter 7.17
42. How do you compute and read residuals for the Negative Binomial model?
Raw residuals of a correct count model fan out with the mean, so I standardize. The Pearson residual is $r_t=(y_t-\hat\mu_t)/\sqrt{\hat\mu_t+\hat\mu_t^2/\hat\alpha}$: its average is about 0 and the average of $r_t^2$ (the Pearson dispersion) about 1; clearly above 1 means more spread than the model allows ($\alpha$ too large for some days, excess zeros, a missing component), below 1 means the NB is more spread than the data. For the Q-Q plot and the ACF I use randomized quantile residuals, $u_t=F(y_t-1)+v_t[F(y_t)-F(y_t-1)]$, $\Phi^{-1}(u_t)\sim N(0,1)$ for any right likelihood. In a Bayesian fit I plug in posterior means or, better, average $F$ over draws (the posterior-predictive PIT). First I check which NB parameterization my code uses before writing down $V$. Chapter 7.17 · Chapter 7.13
43. How do you tell whether a pattern in a residual plot is real or just noise?
Noise makes patterns, so I compare with data simulated from the fitted model. In a lineup I draw $R$ replicate datasets (parameters from the guide, data from the likelihood), process them exactly like the real data, hide the real residual plot among them and see whether I can pick it out. The numeric version picks a statistic (Ljung–Box $Q$, $r_1$, the largest $|z|$, the spread ratio, the zero count) and reports the share of replicates at least as extreme as the real value. With SVI the replicates can be too tidy if the guide is too narrow, so for a contested decision I compare with a second guide or NUTS. Not seeing a pattern is not a proof there is none, and a posterior predictive $p$-value measures surprise under the model, not the probability that the model is right. Chapter 7.17 · Chapter 7.14
Alternatives to a Prophet-style model (7.5, 7.6, 7.18)
44. Where could Holt–Winters beat your model, and what would you do about it?
Exponential smoothing is local: the level, trend and season are updated after every observation, so after an unplanned level shift it follows within a few steps. My trend bends only at candidate changepoints, from a grid and PELT, which usually stop before the end of the history, so a shift near the end is not followed until the next refit; in the chapter's simulated “late shift” world the local models win at long horizons. I would still choose mine when there are several seasonalities, moving holidays, regressors, counts, explainable components and full predictive distributions. The honest procedure is to run both on the same rolling origins, look at the residual-vs-time plot and the last-28-day error, and then widen the changepoint range, loosen the shrinkage or combine the model with a local one. Chapter 7.5 · Chapter 7.18
45. When would you add an AR error to your model, and how does it change the forecast by horizon?
I check the residual ACF first. If lag 1 (and 2) are outside the band while the weekly and trend checks are clean, there is momentum. The smallest extension keeps $\mu_t=g+s+h+X\beta$ and makes the noise AR(1), $\epsilon_t=\varphi\epsilon_{t-1}+\eta_t$ (regression with AR errors; in NumPyro a recursion with lax.scan, with a prior for $\varphi$ in $(-1,1)$). The forecast gets a correction $\varphi^he_T$ that fades: for $\varphi=0.8$ it is 0.21 of the last residual after 7 days and 0.04 after 14, and the error ratio against independent noise is $\sqrt{1-\varphi^{2h}}$, 0.60 at $h=1$ and 0.98 at $h=7$. So the gain is at the first days, not at long horizons. Cautions: the loop is sequential under JIT, and $\varphi$ near 1 trades off with the trend and the $\delta_j$. I would show the gain by horizon against a SARIMAX baseline. Chapter 7.6 · Chapter 7.17 · Chapter 7.18
46. What does your forecasting model assume, how would you test each assumption, and what would you use if it failed?
(1) Components add: swing grows with the level, a funnel in the residuals; fix with logs, a multiplicative form or a count likelihood. (2) A piecewise-linear trend with sparse bends in the allowed range, last slope continuing: steps or drift in the residuals and holdout bias; local level/ETS or state space, more or later changepoints. (3) A fixed seasonal shape: waves at lags 7, 14, 21 that change over time; Holt–Winters or state-space seasonal states. (4) Events and regressors that repeat and are known ahead: spikes on dates; interactions or boosting. (5) Independent noise with a chosen likelihood: ACF, Q-Q, spread bars; AR errors, state space, GP, a variance model. (6) The future resembles the past: drifting error and coverage; monitoring and retraining. (7) The SVI posterior is close to the real one: intervals too narrow; a richer guide, NUTS on a subset. (8) One series at a time: hierarchical or global models. I decide with rolling-origin accuracy and calibration, not fashion. Chapter 7.18 · Chapter 7.18
Production: reproducibility, monitoring and numerical care (7.19)
47. What do you record so that a forecasting run can be rebuilt, and which of your choices are automatic?
A small text file written automatically beside each fitted model: the data version (hash, cut-off date, row count); preprocessing (missing-value rules and the scaler numbers learned from the history); the code commit and the model components and likelihood; every prior scale, such as the Laplace $b$; the seed or PRNG key of each random step; the inference settings (SVI or NUTS, learning rate, tolerance, patience, minimum steps); the guide type and rank and the optimizer; the iterations actually run, the step of the best state and the stopping reason; and the library versions, precision and device. Several choices in my forecasting model are automatic and a later reader cannot guess them: the guide picked from the model size, the changepoint candidates from the grid and PELT, the scaler statistics. So I record what was picked and what triggered it, and I diff two manifests to explain why two forecasts differ. Chapter 7.19
48. Model B has a 0.15 lower holdout MAE than model A after one SVI run each. Do you believe it?
Not yet. SVI draws fresh random numbers at every step, so the ELBO is noisy and the stopping step, the best-state step and even the final parameters depend on the seed; and “same seed” does not even guarantee the same bits on another chip, library version or precision. I would run 5 to 20 seeds per model, over several rolling origins and with identical splits, and compare the mean gap with the seed-to-seed spread ($SE_{\text{gap}}=s\sqrt{2/m}$ for $m$ seeds each). A gap smaller than one or two spreads is not a claim, and I never report the best seed. The same goes for a decision quantity such as $P(\theta_B\gt\theta_A\mid D)=0.97$ from SVI draws: I check that two seeds agree to the precision I quote. Chapter 7.19
49. What would you monitor once the model is live, and what is the difference between data drift and concept drift?
I store every forecast as issued and score it when its actual arrives. Error: rolling MAE, WAPE or MASE against the launch backtest and the naive baseline (a rolling MASE above 1 means naive would have been better). Bias: the rolling mean residual against $\pm3\sigma_e/\sqrt w$, wider limits if the residuals are autocorrelated. Intervals: coverage by level (misses in $w$ days are Binomial$(w,\alpha)$ for an honest interval, so 28 days at 80% give $5.6\pm2.1$ misses) and the PIT shape. Inputs: missing share, freshness, ranges, PSI or KS against the training distribution, and how often PELT finds changepoints in recent residuals. Data drift means $P(x)$ changed, visible in the inputs; concept drift means $P(y\mid x)$ changed, visible only in the errors. An alarm after $k$ bad windows in a row is a question (data first, then events, then the model), not an order to retrain. Chapter 7.19 · Chapter 7.19 · Chapter 7.19
50. The training loss turns into NaN. How do you debug it?
In order: are the inputs finite? Is the loss finite at initialization? Which step first becomes non-finite? Are the gradients finite? jax_debug_nans can stop at the first culprit. The usual causes are unscaled inputs (time in days, regressors on their own scales) with a learning rate above the stable limit ($\text{lr}\lt2/\lambda_{\max}$ of the curvature, and scaling shrinks the condition number), an exp link or a scale heading to 0, $0/0$, $\log$ or $\sqrt{\ }$ of a negative, and masked logs (the double-where trick keeps the gradient finite). The fixes: a history-only scaler, positive parameters behind exp or softplus, sensible initial values, clipped gradients, then learning-rate tuning. If my loop does relative-ELBO early stopping, a non-finite loss must never count as an improvement: it should neither update the best state nor reset the patience counter. svi.stable_update hides the cause, so I count the skipped steps. Chapter 7.19 · Chapter 7.19
Where to go next
This guide taught you to forecast with honest uncertainty: what makes time series different, the classical toolbox, your Prophet-style model piece by piece (trend and changepoints, PELT, Laplace priors, Fourier seasonality, holidays, regressors, likelihoods), how to evaluate point and probabilistic forecasts, how to read residuals, what else you could have used, how to run it in production, and how it joins your A/B framework. Here is how it connects to the other guides, a study plan for the whole series, and some resources.
How this guide connects to the other five
Each row starts from an idea and points to where it is used. Links to other guides open them at the right chapter.
Coming from Guide 1 · Probability & Data
| From Guide 1 | Used in this guide for |
|---|---|
| Correlation, partial correlation (4.15) | Autocorrelation as correlation with your own past, and the PACF as a partial correlation of lags (7.1, 7.3) |
| The Normal, Student-t and Laplace distributions (4.9) | The Normal and Student-t forecast likelihoods and the Laplace prior on slope changes (7.10, 7.13) |
| Poisson, Negative Binomial and the gamma–Poisson mixture (4.8) | Overdispersed counts, NB2 $Var=\mu+\mu^2/\alpha$ and its parameterizations (7.13) |
| Descriptive and robust statistics, MAD, influence (4.14) | Why a Student-t gives extreme days less influence; robust residual spread (7.13, 7.17) |
| Histograms, KDE and Q-Q plots (4.16, 4.17) | The shape of the residuals and the Normal-vs-Student-t decision (7.17) |
| Transformations, standardization and the global scaler (4.18) | Log scales and multiplicative seasonality, heteroscedasticity fixes, and a history-only scaler (7.2, 7.12, 7.17) |
| Law of total variance (4.6), LLN and CLT, the $1/\sqrt n$ error (4.13) | Noise plus parameter uncertainty in a forecast, and standard errors of correlated averages (7.3, 7.14) |
| Data-generating process, IID, exchangeability (4.1) | Why a time series is not IID and why shuffling destroys it (7.1) |
Coming from Guide 2 · Estimation, Inference & Experiments
| From Guide 2 | Used in this guide for |
|---|---|
| Bias, variance, MSE (5.1) | Fourier order, changepoint count and the Laplace scale as bias–variance knobs (7.11, 7.18) |
| Maximum likelihood and MAP (5.2) | Least squares as a Normal likelihood; the Laplace MAP and the full posterior (7.10, 7.13) |
| Regularization as a prior: ridge, lasso, Laplace (5.3) | Soft-thresholding for the changepoint slopes and shrinkage of holiday and seasonal weights (7.10, 7.12) |
| Standard error and the bootstrap (5.5) | Standard errors of correlated data, HAC and the block bootstrap (7.3) |
| Hypothesis tests (5.6) and confidence intervals (5.8) | Ljung–Box, ADF and KPSS tests; prediction intervals versus credible and confidence intervals (7.3, 7.4, 7.14) |
| A/B design, SRM, peeking (5.10, 5.11) | A/A checks and input monitoring in production (7.19) |
| Causal thinking (5.12) | Regressor coefficients are associations unless the data support more (7.12) |
| Linear regression and GLMs (5.13, 5.14) | The design matrix, dummy columns, VIF, and the log link for counts (7.7, 7.12, 7.13) |
| The multivariate Normal and covariance matrices (5.15) | Gaussian processes and guide covariances (7.17, 7.18) |
Coming from Guide 3 · Bayesian Modeling & Computation
| From Guide 3 | Used in this guide for |
|---|---|
| Posterior predictive distributions (6.1), credible intervals and decisions (6.4) | The forecast as a draws array, tail probabilities, and the capacity and stock decisions (7.14) |
| Prior predictive checks (6.2) | Simulating the whole forecasting model before fitting it (7.7, 7.10) |
| Hierarchical models and partial pooling (6.5, 6.6) | Shrinking rare holidays; the pooling-vs-Laplace comparison in the capstone (7.12, 7.20) |
| Posterior predictive checks, prior sensitivity, identifiability (6.8) | Which component gets the credit, changepoint-scale sensitivity, PPCs for forecasts (7.7, 7.10, 7.14) |
| MCMC, HMC and NUTS (6.9, 6.10) | Latent changepoints and checking an SVI forecast against NUTS on a subset (7.8, 7.18) |
| KL divergence and the ELBO (6.11, 6.12) | The log score as an out-of-sample log-likelihood and its KL meaning (7.16, 7.20) |
| Guides: mean-field, full-rank, low-rank (6.13) | Latent dimension of the forecasting model, correlated holiday and trend weights, under-coverage (7.7, 7.12, 7.16) |
| The custom SVI loop (6.14) | Run manifest, seed noise, guards against non-finite losses (7.19) |
JAX, scan, PRNG keys, JIT and boolean masking (6.16, 6.17) | AR errors with lax.scan, fixed-shape changepoint matrices, explicit keys and numerical care (7.8, 7.17, 7.19) |
The foundation guides · Linear Algebra, Calculus, Optimization
| From the foundation guides | Used in this guide for |
|---|---|
| Linear Algebra: least squares (1.10), definiteness (1.12), decompositions and Cholesky (1.13), numerical linear algebra and conditioning (1.15) | Fitting the design matrix, ridge, the $O(n^3)$ cost of a Gaussian process, and why scaling inputs helps training (7.7, 7.18, 7.19) |
| Calculus: gradients (2.4), backpropagation and automatic differentiation (2.9) | The pull of a prior as a gradient, and the gradients behind SVI (7.10, 7.19) |
| Optimization: gradient descent and Adam (3.3, 3.4), regularization as optimization (3.11), coordinate and proximal methods (3.13), stochastic optimization (3.14), convergence (3.5) | ISTA/FISTA and soft-thresholding for the Laplace MAP, a noisy ELBO, learning rates that are too big, and what “converged” means (7.10, 7.19) |
A study plan for the whole four-guide series
Reading the series for the first time? Take the guides in order, 1, 2, 3, 4: each assumes the one before it, and the foundation guides are there when a chapter leans on a matrix, a gradient or an optimizer. The plan below is the final revision before an interview. It starts with Guide 4, the closest to your forecasting project and the most likely to be probed, and then goes back to the earlier guides for the bridge chapters. Do the steps in order; the box after the list says how to squeeze it if time is short.
- Draw the model on a blank page. Write $y_t=g(t)+s(t)+h(t)+X_t\beta+\epsilon_t$, then under each term its form, its parameters, its prior and the chapter that teaches it; check against the model map. Add the two sentences about how it is fitted (SVI, a Gaussian guide) and how it is judged (rolling origin, coverage, CRPS).
- Re-read the P0 chapters of this guide by their notebook boxes: components (7.2), autocorrelation and stationarity (7.3, 7.4), changepoints, PELT and Laplace priors (7.8–7.10), Fourier seasonality (7.11), holidays, regressors and leakage (7.12), likelihoods (7.13), rolling validation (7.15), calibration (7.16), residuals (7.17), alternatives (7.18). Then work through the 44-item P0 checklist and tick only what you said aloud without notes.
- Re-derive five things on paper: the continuity offset $\gamma_j=-s_j\delta_j$ (7.8), the Laplace soft-threshold $\text{sign}(x)\max(0,|x|-se^2/b)$ (7.10), $Var=\mu+\mu^2/\alpha$ from a gamma–Poisson mixture (7.13), MASE and why its ruler is one step (7.15), and CRPS as $E|X-y|-\tfrac12E|X-X'|$ (7.16).
- Answer the question bank and the chapter 7.20 drill out loud, one minute each, in the four-part shape (definition, number, your project, trap). Grade 2 / 1 / 0 and revisit the 0s and 1s the next day.
- Refresh the bridge chapters from the earlier guides, in this order: Guide 1: noise distributions and counts (4.9, 4.8), robust statistics and Q-Q plots (4.14, 4.17), transformations and the global scaler (4.18). Guide 2: estimators, MLE/MAP and regularization (5.1–5.3), regression and GLMs (5.13, 5.14). Guide 3: Bayesian inference and the predictive (6.1), priors and prior predictive checks (6.2), credible intervals and decisions (6.4), pooling (6.6), model checking (6.8), the ELBO, guides and the loop (6.12–6.14), JAX and JIT (6.16, 6.17).
- Prepare two project stories of three minutes each, using only facts you can stand behind: the Bayesian A/B framework (Beta-Binomial and Dirichlet-Multinomial, Normal/Student-t/Poisson, hierarchical partial pooling, the global scaler, posterior decisions, SVI, boolean masking under JIT) and the forecasting model (the additive model, grid plus PELT changepoints with Laplace priors, Fourier seasonality, holidays, regressors, three likelihoods, the custom SVI loop with relative-ELBO stopping, patience, checkpointing and a size-based guide choice). For any code detail not on that list, practise saying “I would check the code”.
- Run the code. Run each “Code it” block of the P0 chapters and the code sheet, change one number and predict the result before you run it.
- The night before: reread the diagnostics-to-fix table, the leakage checklist and the likelihood table, then sleep.
If you have less time. One day: steps 1, 2 (notebook boxes only), 4 and 8. Three days: add steps 3 and 6, and the Guide 3 items of step 5. One week: all eight steps, with step 5 for all three guides and step 7 every day for 20 minutes.
How to make it stick, and where to read more
- Always say “of what”. Which interval (prediction or credible), which mass, which score (CRPS, log score, MASE and with which $m$), which Negative Binomial parameterization, which link, which scaling of $y$ and $t$, which horizon.
- Never let a number stand alone. An MAE needs a baseline on the same windows; a coverage needs its $n$ and luck range; a $b$ needs its units.
- Check the time ordering of everything that learns. If you cannot say “this was fitted on the training window only”, assume it leaked.
- Read residuals before reaching for a bigger model. Each fingerprint has a specific, cheaper fix.
- Describe your projects only with facts you know. “I would check the code” is a good answer for a detail you did not write down.
- Keep your notebook. The “Write this in your notebook” boxes, copied by hand, are the fastest revision sheet you can have.
Well-known resources
- Hyndman & Athanasopoulos, Forecasting: Principles and Practice (3rd ed., free online at otexts.com/fpp3): the standard friendly textbook for decomposition, ETS, ARIMA, regression with time features, cross-validation and accuracy measures.
- Taylor & Letham (2018), “Forecasting at Scale” and the Prophet documentation: the model your forecasting project is modelled on, with the defaults of Appendix B's table (check the version you use).
- Box, Jenkins, Reinsel & Ljung, Time Series Analysis: Forecasting and Control, and Hyndman, Koehler, Ord & Snyder, Forecasting with Exponential Smoothing: The State Space Approach: ARIMA and ETS in depth.
- Killick, Fearnhead & Eckley (2012), “Optimal detection of changepoints with a linear computational cost” (PELT) and Truong, Oudre & Vayatis (2020), “Selective review of offline change point detection methods” (the
ruptureslibrary); Adams & MacKay (2007), “Bayesian online changepoint detection” for the latent view. - Hyndman & Koehler (2006), “Another look at measures of forecast accuracy” (MASE); Gneiting & Raftery (2007), “Strictly proper scoring rules, prediction, and estimation”; and Gneiting, Balabdaoui & Raftery (2007), “Probabilistic forecasts, calibration and sharpness”: the evaluation chapters' sources.
- Durbin & Koopman, Time Series Analysis by State Space Methods, and Rasmussen & Williams, Gaussian Processes for Machine Learning (free online): state-space models and Gaussian processes.
- Dunn & Smyth (1996), “Randomized quantile residuals”: the residuals of Chapter 7.17 for counts.
- Makridakis, Spiliotis & Assimakopoulos on the M4 and M5 competitions (International Journal of Forecasting), Oreshkin et al. (2020), N-BEATS, Lim et al. (2021), Temporal Fusion Transformers and Salinas et al. (2020), DeepAR: pointers for the awareness-only part of Chapter 7.18.
- The statsmodels, ruptures, NumPyro and JAX documentation: the final word on
acf,acorr_ljungbox,SARIMAX,ExponentialSmoothing,Pelt,NegativeBinomial2,Predictiveand PRNG keys.
Companion guides
This guide stands on Guide 1 · Probability & Data, Guide 2 · Estimation, Inference & Experiments and Guide 3 · Bayesian Modeling & Computation, and on three earlier guides: the Linear Algebra guide, the Calculus guide and the Optimization guide. Guide 4 is the last of the series; the next step is to explain your two projects to a person.