Optimization for Machine Learning
Learn how to find the best choice, from zero. Every idea starts with a plain-English picture, then a worked example, then the formal definition, and then something you can drag, slide, rotate and race against other methods. Everything points toward one goal: understanding how a machine-learning model trains.
What is optimization, in one sentence? It is the art of choosing the numbers (the parameters) that make some score (the objective) as good as possible: as small as possible for an error, as large as possible for a profit.
Why care? Training a model is optimization. A neural network has millions of numbers to choose, and an error score that says how bad the current choice is. Training is the search for numbers that make the error small. Gradient descent, Adam, learning-rate schedules, regularization: all of them are optimization ideas.
Companion guides. Optimization builds on the Calculus guide (derivatives, gradients, the Hessian, Taylor series) and the Linear Algebra guide (vectors, matrices, eigenvalues). Whenever one of those ideas appears we remind you in one plain sentence and link to the exact place that explains it.
How every topic is taught
Each concept follows the same six steps, in this order. The order matters: understanding comes before formulas.
1 · Intuition
A picture or everyday story. No symbols yet. If you only read this part, you will already get the idea.
2 · Example
A small problem with real numbers, solved slowly, step by step. You can redo it with pen and paper.
3 · Definition
Now the precise statement and the notation. Because you have seen the idea, the symbols just give it a name.
4 · Play
An interactive visual. Drag points, move sliders, race optimizers against each other. In 3D you can also rotate the whole scene. Each one tells you what to try.
5 · Why · Where · How
Every concept answers three questions in plain English: Why do we need it? Where is it used? How is it used? So you always know why the idea is worth your time.
6 · Check
Short questions to test yourself, a recap of the key points, a NumPy code block to type out, and practice problems with worked answers.
Try your first interactive
The boxes marked Interactive react to you. Drop a ball on a landscape and let it roll downhill.
The roadmap
Seventeen chapters, in an order where each one builds on the last. Click any card to jump there.
- 3.1Optimization FundamentalsObjectives, constraints, local vs global
- 3.2Unconstrained OptimizationCritical points and the second-order test
- 3.3Gradient DescentThe core algorithm, step by step
- 3.4Advanced Gradient MethodsMomentum, Adam, AdamW, schedules
- 3.5ConvergenceRates, smoothness, stopping rules
- 3.6Convex OptimizationWhy local = global
- 3.7Constrained OptimizationFeasible regions and the big picture
- 3.8Lagrange MultipliersTangent curves and the Lagrangian
- 3.9KKT ConditionsInequality constraints made precise
- 3.10DualityThe dual problem and what it tells you
- 3.11RegularizationLoss + λ·penalty, L1, L2, sparsity
- 3.12Second-Order MethodsNewton, BFGS, L-BFGS
- 3.13Coordinate & Proximal MethodsSoft thresholding and the Lasso
- 3.14Stochastic OptimizationNoisy gradients, mini-batches
- 3.15Optimization in Deep LearningLandscapes, vanishing gradients, clipping
- 3.16Numerical OptimizationRounding, gradient checking, autodiff
- 3.17The Final ToolboxWhich method for which problem
What matters most for ML. If you only have time for the essentials: 3.3 gradient descent, 3.4 momentum and Adam, 3.6 convexity, 3.11 regularization, 3.14 stochastic optimization and 3.15 deep-learning optimization. Chapters 3.7 to 3.10 (constraints, Lagrange, KKT, duality) are what you need to understand SVMs and many "classical" algorithms: learn what the conditions mean rather than how to prove them.
How to read the maths symbols
You will meet these symbols again and again. Do not memorise them now. Come back to this table whenever one looks strange.
| Symbol | Say it as | Meaning |
|---|---|---|
| $\min_x f(x)$ | "minimise f of x" | Find the $x$ that makes $f(x)$ as small as possible. |
| $\arg\min_x f(x)$ | "the argmin" | The input $x$ that gives the smallest value (not the smallest value itself). |
| $x^\star$ | "x star" | The optimal solution. |
| $\text{s.t.}$, $g(x)\le 0$ | "subject to" | The rules the solution must obey (constraints). |
| $\nabla f$ | "the gradient of f" | The vector of partial derivatives. It points uphill. |
| $\nabla^2 f$, $H$ | "the Hessian" | The matrix of second derivatives (curvature). |
| $\eta$, $\alpha$ | "eta", "alpha" | The learning rate or step size. |
| $\lambda$, $\mu$ | "lambda", "mu" | A multiplier (Lagrange, KKT) or a regularization strength. |
| $\|\mathbf{x}\|$ | "norm of x" | The length of $\mathbf{x}$. |
| $L$-smooth, $\mu$-strongly convex | Roughly: the curvature (the Hessian's eigenvalues) is at most $L$ and at least $\mu$ everywhere. | |
| $\mathbb{E}[\cdot]$ | "expected value" | The average over randomness (for example over random mini-batches). |
| $O(1/k)$ | "big O of one over k" | How fast an error shrinks after $k$ steps. |
Tips for studying
- Play first, read second. If a paragraph confuses you, go to the interactive below it and move things around. Then re-read.
- Race the optimizers. Many widgets let you run several methods on the same landscape. Watching them move teaches more than any formula.
- Always ask "what is the landscape?" Almost every optimization idea is a statement about a hilly surface: where you stand, which way is down, how far to step.
- Don't rush. One chapter a week is a good pace. Come back to earlier chapters whenever you need to.
- Use the tools. Press / to search, use the sidebar to jump around, switch dark mode with the button at the top, and mark chapters complete as you go.
If a section feels too hard, it is almost always because a word from an earlier section is fuzzy. Go back, find that word in the sidebar, and reread it. Nobody gets optimization in a single pass. Seeing it twice is normal.
Optimization Fundamentals
Every time you pick the cheapest route or tune a recipe, you are doing optimization. This chapter gives the everyday idea its exact vocabulary: objective, decision variables, constraints, feasible region, local and global optimum. By the end you will be able to read the sentence "training a model is minimising a loss over its parameters" and know what every word in it means.
- Say what optimization is, and name its three ingredients: an objective, decision variables and constraints
- Tell decision variables, problem parameters and hyperparameters apart
- Describe the feasible region, feasible solutions and the optimal solution
- Know that max $f = -\min(-f)$, and tell minimum, maximum, local, global and saddle points apart
- Write any problem in the general form $\min_{\mathbf{x}} f(\mathbf{x})$ subject to $g_i(\mathbf{x}) \le 0$, $h_j(\mathbf{x}) = 0$
- Solve a small problem by hand, see what adding a constraint does, and map all of it onto machine-learning training
What is optimization? core
You make small optimization problems every day. Which route to work is cheapest? How much sugar goes in the tea? How many hours should you study tonight? Each time, three things are going on:
- There is a list of choices you can make (this route or that route; 1 spoon of sugar or 2).
- There is a score that tells you how good or bad a choice is (the cost of the trip; how much you dislike the tea).
- Sometimes there are rules (you cannot spend more than 20 dollars; you cannot use a negative amount of sugar).
Optimization is the craft of finding the choice with the best score that obeys the rules. Machine learning is the same story with millions of choices: the "choices" are the numbers inside the model, the "score" says how wrong the model is, and training is the search for numbers that make the score small.
Cheapest route. Three ways to get to work. You value your time at 0.25 dollars per minute (15 dollars per hour). The total cost of a route is $0.25 \times \text{minutes} + \text{toll}$.
- Route A: 40 minutes, no toll. Cost $= 0.25\cdot 40 + 0 = 10$ dollars.
- Route B: 30 minutes, toll 2 dollars. Cost $= 0.25\cdot 30 + 2 = 7.5 + 2 = 9.5$ dollars.
- Route C: 25 minutes, toll 5 dollars. Cost $= 0.25\cdot 25 + 5 = 6.25 + 5 = 11.25$ dollars.
The best choice is Route B. With only three choices you can simply try them all. With infinitely many choices (any amount of sugar from 0 to 10 spoons, in any fraction) you cannot try them all, and you need ideas from calculus. That is what the rest of this guide is about.
Optimization means choosing values for some unknowns, the decision variables $\mathbf{x}$, so that a number called the objective function $f(\mathbf{x})$ is as small (or as large) as possible, while obeying any constraints. For a "make it small" problem we write
$$\min_{\mathbf{x}} \; f(\mathbf{x}).$$Read it as: "find the $\mathbf{x}$ that makes $f(\mathbf{x})$ smallest". The rest of this chapter names each part of that sentence.
Why do we need it?
Many questions have too many possible answers to try one by one. Optimization gives a systematic way to search for the best answer instead of guessing.
Where is it used?
Training every machine-learning model (linear regression, logistic regression, neural networks), delivery routing, portfolio selection, scheduling flights, and tuning engineering designs.
How is it used?
First turn the question into a score to minimise (or maximise) and a list of rules. Then pick a method (by hand with calculus, or an algorithm like gradient descent in Chapter 3.3) that searches for the best choice.
Quick check: with the route costs above, what if your time were worth 0.10 dollars per minute?
A: $4 + 0 = 4$. B: $3 + 2 = 5$. C: $2.5 + 5 = 7.5$. Now Route A wins. Same choices, same tolls, but a different score gives a different best answer. The objective decides what "best" means.
The objective function core
The objective function is the score sheet. You hand it one candidate choice and it hands back one number saying how good that choice is.
A great way to picture it: a landscape. Each choice is a spot on the ground, and the objective is the height of the ground at that spot. "Minimise" means "find the lowest point"; "maximise" means "find the highest point". Whenever you meet an optimization problem, ask: what is the landscape?
Different fields use different names for the score. In machine learning the thing we minimise is usually called the loss, cost or error. A thing we maximise is called a reward, profit, utility or likelihood.
Tea. Let $x$ be spoons of sugar, and let the "dislike score" be $f(x) = \tfrac12 (x-4)^2 + 2$.
- $f(1) = \tfrac12 \cdot 9 + 2 = 6.5$ (too bitter).
- $f(4) = \tfrac12 \cdot 0 + 2 = 2$ (just right: the lowest possible).
- $f(6) = \tfrac12 \cdot 4 + 2 = 4$ (too sweet).
A loss in machine learning. One training example: the input is $1$ and the true answer is $2$. A model predicts $w \cdot 1 = w$. Its squared-error loss is $L(w) = (w - 2)^2$. Then $L(0) = 4$, $L(1) = 1$, $L(2) = 0$. The loss is the "height of the landscape" over the choice of $w$, and the lowest point is $w = 2$.
An objective function is a function
$$f : \mathbb{R}^n \to \mathbb{R}, \qquad \mathbf{x} \mapsto f(\mathbf{x}),$$which takes the decision variables $\mathbf{x} = (x_1, \dots, x_n)$ and returns a single real number. We either minimise it or maximise it.
- It must output one number. If you care about two things at once (speed and cost), you must combine them into one score, for example $0.25 \cdot \text{minutes} + \text{toll}$ as in the routes above.
- Which function you choose defines the problem. Change $f$ and the best answer changes.
Why do we need it?
"Best" means nothing until you say how to measure it. The objective turns a vague wish ("a good model", "a cheap trip") into one number that a computer can compare and improve.
Where is it used?
Mean squared error in regression, cross-entropy loss in classification, negative log-likelihood in probabilistic models, reward in reinforcement learning, and total cost in logistics.
How is it used?
Write the objective as code that takes the current choice and returns one number. Optimization algorithms call it again and again (and also look at its slope, see Chapter 3.3) to find choices with lower values.
A model's accuracy is not usually the objective. Accuracy jumps in steps (a prediction is right or wrong), so it gives a landscape of flat plateaus with no slope to follow. Training minimises a smooth stand-in, the loss, and we then hope accuracy improves along the way.
Quick check: can the objective return a list of two numbers, like (cost, time)?
No. A single objective returns one number, because "smallest" only makes sense for numbers on a line. To handle two goals you either add them with weights (like $0.25\cdot\text{minutes} + \text{toll}$), or treat one as a constraint ("time at most 30 minutes").
Decision variables core
Decision variables are the dials you are allowed to turn. They are the unknowns of the problem: the thing you are trying to choose. Everything else in the problem is fixed.
Picture a mixing desk with a few sliders. The sound quality is the score; the sliders are the decision variables. One slider means a one-dimensional problem; ten sliders, a ten-dimensional one. A neural network is a mixing desk with millions or billions of sliders.
- Tea: one dial, the spoons of sugar $x$.
- Fencing a field: two dials, the width $x$ and the length $y$.
- Fitting a line $\hat y = w x + b$ to data: two dials, the slope $w$ and the intercept $b$. The data points are not dials: they are given.
Take three data points $(1, 2)$, $(2, 3)$, $(3, 5)$. Try $w = 1, b = 0$. The predictions are $1, 2, 3$ and the errors (prediction minus truth) are $-1, -1, -2$. Sum of squared errors: $1 + 1 + 4 = 6$. Now try $w = 1.5$, $b = 0.33$: predictions $1.83, 3.33, 4.83$, errors $-0.17, 0.33, -0.17$, sum of squares $\approx 0.167$. Much better. Turning the two dials changed the score from 6 to 0.17.
The decision variables are the unknowns we choose, collected into a vector
$$\mathbf{x} = \begin{bmatrix} x_1 \\ \vdots \\ x_n \end{bmatrix} \in \mathbb{R}^n .$$The number $n$ is the dimension of the problem. Choosing $\mathbf{x}$ means choosing a point in $\mathbb{R}^n$; the objective $f(\mathbf{x})$ is the height of the landscape over that point. (Vectors are explained in the Linear Algebra guide.)
Why do we need it?
To search for a best answer you must first say exactly what you are allowed to change. Naming the decision variables turns a vague problem into a point in a space that an algorithm can move around in.
Where is it used?
The weights and bias of a linear model, the width and length in an engineering design, the amount of money put into each stock in a portfolio, and the billions of weights in a language model.
How is it used?
Stack all the numbers you may change into one vector $\mathbf{x}$. Write the objective as a function of that vector. The optimizer then moves $\mathbf{x}$ and watches how $f$ responds.
Not every number is a dial. The three data points, the toll prices and the field's required area are given. If you also let the optimizer change them, you would "improve" the score by cheating (move the data onto the line!). Always ask: which numbers do I get to choose?
Quick check: a linear model with 5 input features and a bias. How many decision variables?
$5$ weights $+ 1$ bias $= 6$, so $n = 6$ and the weight vector lives in $\mathbb{R}^6$.
Parameters (and how they differ from decision variables and hyperparameters)
Think of a cooking recipe. Given numbers are things like "serves 4 people" and "the oven is 180 degrees": they define this problem, and you do not change them while solving it. The dials you turn are the decision variables (how much salt, how long to bake). And a third kind of number is a setting of your search method: how big a change you try each time you adjust a dial, or how many attempts you allow.
The word "parameter" is used for the given numbers (in optimization textbooks) and, in machine learning, for the dials you turn (the weights). That causes real confusion, so let us be careful.
Fit the line $y = w x$ (no intercept) to the data $(1, 2)$ and $(2, 3)$.
- Errors: $w\cdot 1 - 2$ and $w \cdot 2 - 3$.
- Total squared error: $L(w) = (w-2)^2 + (2w-3)^2 = w^2 - 4w + 4 + 4w^2 - 12 w + 9 = 5w^2 - 16 w + 13$.
- Smallest where the slope is zero: $L'(w) = 10 w - 16 = 0$, so $w^\star = 1.6$, and $L(1.6) = 5(2.56) - 25.6 + 13 = 0.2$.
Here $w$ is the decision variable. The numbers $2, 3, 1, 2$ (the data) are the parameters of the problem. If the data change to $(1,2)$ and $(2,5)$, the same recipe gives a different best $w$. The learning rate you would use if you solved it by gradient descent (Chapter 3.3) is a hyperparameter.
| Kind of number | Who sets it? | Changes during optimization? | Examples |
|---|---|---|---|
| Decision variables $\mathbf{x}$ | the optimizer | yes, that is the whole point | $w, b$; the width and length of the field; neural-network weights |
| Problem parameters (data, constants) | the world / the problem statement | no | the data points; prices; the field area 100; the bound in a constraint |
| Hyperparameters | you, before running the optimizer | no (tuned in an outer loop by trial) | learning rate, regularisation strength $\lambda$, number of layers, batch size |
Vocabulary warning. In optimization textbooks "parameters" are the problem constants. In machine learning, "the model's parameters" almost always means the weights, which are the decision variables. In this guide we say "decision variables" (or "weights") for the numbers being optimised, "data" or "problem parameters" for the fixed inputs, and "hyperparameters" for the knobs of the method.
Why do we need it?
Mixing up the three kinds leads to bugs: tuning the data, or treating a learning rate as something the optimizer will fix. Keeping them apart tells you what to change, what to freeze, and what to search over separately.
Where is it used?
Every ML pipeline: the training loop updates the weights (decision variables) using the training data (parameters of the problem), while a separate search (grid search, random search) picks hyperparameters such as the learning rate and $\lambda$.
How is it used?
Before you start, list your numbers and sort them: "does the optimizer change this? does the world fix this? do I set it by hand?" Pass only the first kind as the variable of the optimizer; everything else is a constant of the function.
Quick check: in training a neural network, which of these is a hyperparameter: the weights, the learning rate, the training images?
The learning rate. The weights are the decision variables; the training images are the problem's data (its parameters in the textbook sense).
Constraints core
Constraints are the rules of the game. They say which choices are allowed at all, however good their score might look.
Without rules, many problems are silly. "Cheapest fence around a field": use no fence. "Best exam score": give yourself 100. Rules make the problem real: "the field must have an area of 100 square metres"; "you can spend at most 50 dollars"; "the amounts of the ingredients cannot be negative"; "the probabilities must add up to 1".
Rules come in two flavours: "at most / at least" (a limit, an inequality) and "exactly" (a precise requirement, an equality).
Two decision variables $x$ and $y$ (say, kilograms of two ingredients). The rules:
- "the total is at most 4": $x + y \le 4$
- "no negative amounts": $x \ge 0$ and $y \ge 0$
Test the choice $(1, 2)$: $1 + 2 = 3 \le 4$ ✓, $1 \ge 0$ ✓, $2 \ge 0$ ✓. Allowed. Test $(3, 3)$: $3 + 3 = 6 \not\le 4$ ✗. Not allowed.
An equality example: "the field must have area exactly 100": $x \cdot y = 100$. The choice $(10, 10)$ obeys it; the choice $(10, 9)$ does not, because $90 \ne 100$.
Notation, stated once for the whole guide. We always move everything to one side so that a constraint compares a function with zero:
- An inequality constraint is written $g_i(\mathbf{x}) \le 0$, for $i = 1, \dots, m$. The letter $g$ means "inequality".
- An equality constraint is written $h_j(\mathbf{x}) = 0$, for $j = 1, \dots, p$. The letter $h$ means "equality".
Converting is just moving terms: $x + y \le 4$ becomes $g_1(x,y) = x + y - 4 \le 0$. A "$\ge$" rule is flipped by multiplying by $-1$: $x \ge 0$ becomes $g_2(x,y) = -x \le 0$. And $xy = 100$ becomes $h_1(x,y) = xy - 100 = 0$.
A constraint is active at a point if it holds with equality there ($g_i = 0$, the point sits on the edge of the rule) and inactive if it holds with room to spare ($g_i \lt 0$). Equality constraints are always active. (We study active constraints properly in Chapter 3.7.)
Why do we need it?
Real choices have limits: money, time, physics, non-negativity, probabilities that sum to 1. Constraints let us state those limits in the problem itself, instead of throwing away bad answers afterwards.
Where is it used?
The margin rules of a support vector machine, norm limits on weights (the ball $\|\mathbf{w}\| \le r$ is the constrained form of ridge regression), probability simplexes in mixture models and softmax outputs, and resource budgets in operations research.
How is it used?
Write each rule as $g_i(\mathbf{x}) \le 0$ or $h_j(\mathbf{x}) = 0$. A point obeys a rule if the number $g_i$ is zero or negative (for equalities: exactly zero). Methods in Chapters 3.7 to 3.10 and 3.17 handle them.
Equality constraints are much more restrictive than they look. One equality in two variables cuts the allowed set from a whole area down to a curve. That is why you almost never hit an equality by luck: the point must lie exactly on it.
Hard versus soft rules. Sometimes a rule is not absolute. Instead of forbidding large weights, you can charge a price for them by adding a penalty to the objective. That is how regularisation works (Chapter 3.11); the same trick for general constraints is the penalty method (Chapter 3.17).
Quick check: write "$x_1 + 2x_2 \ge 6$" in the form $g(\mathbf{x}) \le 0$.
Multiply by $-1$ (which flips the direction) and move everything to the left: $-x_1 - 2x_2 + 6 \le 0$. So $g(\mathbf{x}) = 6 - x_1 - 2x_2$. Check with $(4, 1)$: $4 + 2 = 6 \ge 6$ holds, and $g = 6 - 4 - 2 = 0 \le 0$ holds too.
Feasible region and feasible solutions core
Draw the landscape of the objective. Now put a fence around the part of the ground where you are allowed to stand. The fenced-in ground is the feasible region. ("Feasible" is just a grand word for "allowed".) A spot inside is a feasible solution; a spot outside is infeasible.
Two things are easy to mix up, so say them slowly: a feasible solution is just a choice that obeys the rules. It may still be a terrible choice. The optimal solution is the best of all the feasible ones.
Objective: $f(x, y) = (x-3)^2 + (y-2)^2$ (the squared distance from the point $(3, 2)$). Rules: $x + y \le 4$, $x \ge 0$, $y \ge 0$. The feasible region is the triangle with corners $(0,0)$, $(4,0)$, $(0,4)$.
| Choice | $x+y$ | Feasible? | $f$ |
|---|---|---|---|
| $(1, 1)$ | 2 | yes | $4 + 1 = 5$ |
| $(2, 1)$ | 3 | yes | $1 + 1 = 2$ |
| $(2.5, 1.5)$ | 4 | yes (on the edge) | $0.25 + 0.25 = 0.5$ |
| $(3, 2)$ | 5 | no | $0$ (lowest of all, but not allowed) |
All three feasible choices obey the rules, and their scores differ: 5, 2, 0.5. The point $(3,2)$ has the lowest score in the whole plane but is infeasible, so it does not count. Which feasible point is best? Section "Worked example 2" below finds it by hand.
The feasible region (or feasible set) is the set of all points that obey every constraint:
$$\mathcal{F} = \{\, \mathbf{x} \in \mathbb{R}^n \;:\; g_i(\mathbf{x}) \le 0 \text{ for all } i, \;\; h_j(\mathbf{x}) = 0 \text{ for all } j \,\}.$$- A point $\mathbf{x} \in \mathcal{F}$ is a feasible solution (or feasible point). A point outside is infeasible.
- No constraints at all means $\mathcal{F} = \mathbb{R}^n$: an unconstrained problem (Chapter 3.2).
- If the rules contradict each other (for example $x \ge 2$ and $x \le 1$), $\mathcal{F}$ is empty and the problem is infeasible: it has no solution at all.
- $\mathcal{F}$ can be bounded (the triangle) or unbounded (just $x \ge 0$: a whole half-line).
Why do we need it?
It separates two questions: "is this choice allowed?" and "is this choice good?". Solvers need both. The feasible region is the map of where they may search.
Where is it used?
Linear programming (the polygon of allowed plans), SVMs (weights that classify every training point correctly), projected gradient descent (stay inside the region), and checking that a model's output is a valid probability.
How is it used?
Test a candidate by plugging it into every constraint. If any $g_i \gt 0$ or any $h_j \ne 0$, it is infeasible. Algorithms either start inside and stay inside, or are allowed to wander out and are pulled back.
Feasible does not mean good, and good does not mean feasible. The point with the lowest objective value can be infeasible (like $(3,2)$ above), and a feasible point can have a terrible score. A solver must satisfy both.
Quick check: is $(1, 2)$ feasible for the rules $x + y \le 4$, $x \ge 0$, $y \ge 0$, and also $x - y = 1$?
It obeys the three inequalities ($3 \le 4$, $1 \ge 0$, $2 \ge 0$), but $x - y = 1 - 2 = -1 \ne 1$. So it breaks the equality: infeasible. A point must obey every constraint.
The optimal solution core
The optimal solution is the winner: the feasible choice that no other feasible choice beats. We write it $\mathbf{x}^\star$ and say "x star".
There are two different things you might want to know, so keep them separate:
- Where is the best spot? That is the optimal solution $\mathbf{x}^\star$ (also called the optimiser or minimiser). "Add 4 spoons."
- How good is it? That is the optimal value $f^\star = f(\mathbf{x}^\star)$. "The tea will score 2."
Beware: a winner does not always exist, and it is not always unique. Watch for this in the widget below.
Tea again, $f(x) = \tfrac12(x-4)^2 + 2$. With no rules the winner is $x^\star = 4$ and the best score is $f^\star = 2$.
Now add the rule "I only have 3 spoons: $x \le 3$", that is $g(x) = x - 3 \le 0$. On the allowed spoons $0 \le x \le 3$, the score keeps falling as $x$ grows toward 3 (because 3 is still below 4). So $x^\star = 3$ and $f^\star = \tfrac12 \cdot 1 + 2 = 2.5$. The rule moved the winner and made the best score worse (from 2 to 2.5).
A point $\mathbf{x}^\star \in \mathcal{F}$ is an optimal solution of the problem $\min_{\mathbf{x} \in \mathcal{F}} f(\mathbf{x})$ if
$$f(\mathbf{x}^\star) \le f(\mathbf{x}) \quad \text{for every feasible } \mathbf{x}.$$Then $f^\star = f(\mathbf{x}^\star)$ is the optimal value, written $\min_{\mathbf{x}} f(\mathbf{x})$. The set of all winners is written $\arg\min_{\mathbf{x}} f(\mathbf{x})$ ("the argument that gives the minimum"). The two answers are different objects: $\min$ returns a number $f^\star$; $\arg\min$ returns an input $\mathbf{x}^\star$.
- An optimum may be non-unique (a flat-bottomed valley has a whole interval of winners; then $\arg\min$ is a set).
- An optimum may not exist: the objective may keep improving forever, or approach a value it never reaches. In that case we speak of the infimum (the greatest lower bound).
- It is guaranteed to exist if $f$ is continuous and $\mathcal{F}$ is non-empty, closed and bounded (the extreme value theorem). Closed means the edge belongs to $\mathcal{F}$; bounded means $\mathcal{F}$ does not stretch out to infinity.
Why do we need it?
It is the thing we are searching for. Knowing exactly what counts as "optimal" (and when it may not exist) tells us what an algorithm can promise and when to stop.
Where is it used?
The trained weights $\mathbf{w}^\star$ of a model are the optimiser of its training loss; the optimal value $f^\star$ is the lowest loss achievable; a flat valley explains why two different weight vectors can give identical predictions.
How is it used?
Report both $\mathbf{x}^\star$ and $f^\star$. During training you rarely know $f^\star$ in advance, so you watch how much $f$ still falls. In theory you compare any candidate with $f^\star$ to measure how far from optimal it is.
"The minimum" and "the minimiser" are not the same thing. For $f(x) = (x-1)^2 + 1$ the minimum is $1$ and the minimiser is $x^\star = 1$. They happen to look alike here by coincidence; for tea, the minimum is 2 and the minimiser is 4.
In machine learning we rarely reach the exact optimum. We stop when the loss is "good enough", so keep the difference between "the optimal solution" and "the solution the algorithm returned" in mind.
Quick check: for $f(x) = (x-2)^2$ subject to $x \ge 5$, what are $x^\star$ and $f^\star$?
On $x \ge 5$ the parabola is rising, so the lowest allowed point is the edge $x^\star = 5$, and $f^\star = (5-2)^2 = 9$. The unconstrained minimiser $x = 2$ is infeasible.
Minimum and maximum: one problem, two names core
A minimum is the bottom of a valley; a maximum is the top of a hill. Cost, loss and error are things you want low. Profit, reward, accuracy and likelihood are things you want high.
Here is the key trick: turn the picture upside down. A hilltop becomes a valley bottom, and it is in the same spot, at the same left-to-right position. So finding the highest point of $f$ is exactly the same job as finding the lowest point of $-f$. That is why optimization books (and this one) mostly talk about minimising: one kind of problem is enough.
Profit. A shop's weekly profit when it sets a price level $x$ is $P(x) = 3 - (x-1)^2$ (in thousands of dollars).
- $(x-1)^2 \ge 0$ always, so $P(x) \le 3$, with equality when $x = 1$. The maximum of $P$ is $3$, at $x = 1$.
- Flip it: $-P(x) = (x-1)^2 - 3$. This is smallest when $(x-1)^2 = 0$, that is $x = 1$, and then $-P(1) = -3$. The minimum of $-P$ is $-3$, at the same $x = 1$.
- Check: $\max P = 3 = -(-3) = -\min(-P)$ ✓.
Machine-learning version. To estimate the chance $p$ that a coin shows heads, you could maximise the likelihood of the data you saw. By flipping the sign of the logarithm, you instead minimise the negative log-likelihood. Same $p$ comes out (try the second example in the widget).
For any objective $f$ and feasible set $\mathcal{F}$:
$$\max_{\mathbf{x}\in\mathcal{F}} f(\mathbf{x}) = -\min_{\mathbf{x}\in\mathcal{F}} \bigl(-f(\mathbf{x})\bigr), \qquad \arg\max_{\mathbf{x}} f(\mathbf{x}) = \arg\min_{\mathbf{x}} \bigl(-f(\mathbf{x})\bigr).$$- The maximiser is the same point; only the value changes sign.
- More generally, the best point does not change if you add a constant to $f$, multiply it by a positive number, or apply any strictly increasing function (such as $\log$): $\arg\min f = \arg\min (a f + c)$ for $a \gt 0$, and $\arg\max L = \arg\max \log L$ when $L \gt 0$.
- Maximum (global): $f(\mathbf{x}^\star) \ge f(\mathbf{x})$ for all feasible $\mathbf{x}$. Minimum (global): $f(\mathbf{x}^\star) \le f(\mathbf{x})$ for all feasible $\mathbf{x}$. Local versions come in the next section.
Why do we need it?
It lets us write one kind of algorithm (a minimiser) and still solve "maximise" problems, by negating the objective. Everyone in the field then speaks the same language.
Where is it used?
Maximum-likelihood estimation (maximise $\log L$, i.e. minimise the negative log-likelihood), reinforcement learning (maximise reward by minimising negative reward), and optimisers like Adam, which are written as minimisers.
How is it used?
If your goal is "maximise $f$", hand $-f$ to a minimiser, then flip the sign of the answer to get the maximum value. The maximiser $\mathbf{x}^\star$ comes out unchanged.
Do not flip the answer twice. If a library "maximises" your function by minimising $-f$, the number it reports as the optimal loss may be $-f^\star$. Check which sign you are looking at.
Quick check: $g(x) = -x^2 + 4x$. What is $\max g$, and what is $\min(-g)$?
$g(x) = 4 - (x-2)^2$, so $\max g = 4$ at $x = 2$. Then $-g(x) = (x-2)^2 - 4$ has minimum $-4$ at $x = 2$. Indeed $\max g = 4 = -(-4)$.
Local and global optima core
Imagine hiking in thick fog. You can only see the ground a few steps around you. You walk downhill until every direction goes up, and you stop in a pit. Are you at the lowest point in the whole country? You cannot tell. You are only sure that you are at the lowest point nearby.
- A local optimum is best in its neighbourhood: nothing close by beats it.
- A global optimum is best everywhere (among all feasible choices).
Every global minimum is also a local one. The reverse is false: a landscape with many valleys has many local minima, and only the deepest is global. A method that simply walks downhill (like gradient descent, Chapter 3.3) is the hiker in the fog: it finds the nearest valley, and where you start decides which valley that is.
Take $f(x) = \sin(1.5x) + x^2/10$ on the interval $-5.5 \le x \le 5.5$. Its slope is $f'(x) = 1.5\cos(1.5x) + x/5$. Setting the slope to zero and checking the shape (we will do this properly in Chapter 3.2) gives three valleys:
| where | $x \approx$ | $f(x) \approx$ | type |
|---|---|---|---|
| left valley | $-4.78$ | $1.51$ | local minimum only |
| middle valley | $-0.96$ | $-0.90$ | global minimum (the lowest of the three) |
| right valley | $2.88$ | $-0.09$ | local minimum only |
A hiker starting at $x = 4$ rolls into the right valley and stops at $-0.09$, never learning that $-0.90$ exists.
Let $\mathcal{F}$ be the feasible set. A point $\mathbf{x}^\star \in \mathcal{F}$ is
- a local minimum if there is some radius $\varepsilon \gt 0$ such that $f(\mathbf{x}^\star) \le f(\mathbf{x})$ for every feasible $\mathbf{x}$ with $\|\mathbf{x} - \mathbf{x}^\star\| \lt \varepsilon$;
- a strict local minimum if the inequality is strict ($\lt$) for every such $\mathbf{x} \ne \mathbf{x}^\star$;
- a global minimum if $f(\mathbf{x}^\star) \le f(\mathbf{x})$ for all $\mathbf{x} \in \mathcal{F}$.
Local and global maxima are defined the same way with $\ge$. Global $\Rightarrow$ local, never the other way round. A problem may have several global minima (all with the same value), as in the Himmelblau landscape below.
The big good news (Chapter 3.6). For a special family called convex problems, every local minimum is global. Linear and logistic regression are convex; deep neural networks are not.
Why do we need it?
It tells us what an algorithm that only looks nearby can promise. "I found a local minimum" is a much weaker statement than "I found the best answer", and we must know which one we are claiming.
Where is it used?
Neural-network training (landscapes with many local minima and flat regions), clustering with k-means (different starting centres reach different local optima), and mixture-model fitting with EM.
How is it used?
In practice you try several starting points ("random restarts"), keep the best result, and compare their scores. For convex problems a single run suffices.
"Local minimum" is not a failure. In deep learning, many local minima are often found to be nearly as good as the best one, and often a point that is "good enough" is all we need. Do not assume a local minimum is bad, or that the global one is always reachable.
You cannot see "local vs global" from the formula of one point. To tell them apart you must either know something about the whole function (like convexity) or compare several results.
Quick check: a landscape has three valleys with depths $-1.0$, $-0.4$ and $-1.0$. How many local minima, and how many global minima?
Three local minima. The two valleys with depth $-1.0$ both are global minima (they tie), so there are two global minima.
Saddle points
A mountain pass between two peaks is flat at the very top of the path, yet it is neither a valley nor a hilltop. Walk along the path through the pass and you are at the highest point of that walk. Walk across the pass (towards a peak, up either side) and you are at the lowest point of that walk. A horse saddle has the same shape. A marble placed there balances for a moment, then rolls away if you nudge it the wrong way.
That balance point is a saddle point: the ground is level (the slope is zero), but it curves up in some directions and down in others.
$f(x, y) = x^2 - y^2$ at the origin.
- Along the $x$-axis ($y = 0$): $f = x^2$, a smile. The origin looks like a minimum.
- Along the $y$-axis ($x = 0$): $f = -y^2$, a frown. The origin looks like a maximum.
- $f(0,0) = 0$. A step to $(0.1, 0)$ gives $+0.01$ (up); a step to $(0, 0.1)$ gives $-0.01$ (down). So it is neither a local min nor a local max.
A saddle point is a point where the slope is zero ($\nabla f = \mathbf{0}$) but which is neither a local minimum nor a local maximum: in some directions $f$ goes up, in others it goes down.
In one variable there is no real saddle, but $f(x) = x^3$ at $x = 0$ is a cousin: the slope is zero there, and yet the function keeps falling on the left and rising on the right. Chapter 3.2 gives the exact test using the Hessian (if the Hessian has both positive and negative eigenvalues at a zero-slope point, that point is a saddle); Chapter 3.15 explains why saddles matter in deep learning.
Why do we need it?
"The slope is zero" does not mean "we are at a minimum". Saddles are flat spots that look like solutions to a naive test but are not. Knowing about them stops us from stopping too early.
Where is it used?
The loss landscapes of neural networks are thought to have a great many saddle points; saddle shapes also appear in game theory (a Nash equilibrium of a two-player zero-sum game) and in the Lagrangian of constrained problems (Chapters 3.8 to 3.10).
How is it used?
When training stalls on a plateau, suspect a saddle. Methods with momentum, noise (SGD) or curvature information often help to escape them. At a candidate solution, compute the Hessian's eigenvalues to rule a saddle out.
Quick check: for $f(x,y) = x^2 + y^2$, is the origin a saddle point?
No. The slope is zero there, but $f$ goes up in every direction ($x^2 + y^2 \ge 0$). It is a (global) minimum. A saddle needs mixed directions: some up, some down.
The general form of an optimization problem core
Letters follow a template: greeting, body, sign-off. Optimization problems have a template too, and once you know it you can read any of them. It has just three lines:
- What do I minimise? (the objective $f$)
- Over what? (the decision variables $\mathbf{x}$)
- Subject to which rules? (the constraints $g_i$ and $h_j$)
The template always says "minimise". So a "maximise" problem is rewritten by flipping the sign, and a "$\ge$" rule is rewritten as a "$\le$" rule.
"Maximise the profit $3x + 2y$, where $x + y \ge 2$ and $x \le 5$."
- Maximise becomes minimise by flipping the sign: $f(x, y) = -(3x + 2y) = -3x - 2y$.
- $x + y \ge 2$ becomes (multiply by $-1$, move to the left) $g_1(x, y) = 2 - x - y \le 0$.
- $x \le 5$ becomes $g_2(x, y) = x - 5 \le 0$.
No "exactly" rule appeared, so there is no $h_j$. In the general form: $\min_{x,y} \; -3x - 2y$ subject to $2 - x - y \le 0$ and $x - 5 \le 0$.
The general form (also called the standard form) of a constrained optimization problem is
$$\begin{aligned} \min_{\mathbf{x} \in \mathbb{R}^n} \quad & f(\mathbf{x}) \\ \text{subject to} \quad & g_i(\mathbf{x}) \le 0, \quad i = 1, \dots, m, \\ & h_j(\mathbf{x}) = 0, \quad j = 1, \dots, p. \end{aligned}$$"Subject to" is often abbreviated s.t. Remember the convention from the Constraints section: $g_i$ are the inequality constraints, $h_j$ are the equality constraints. We keep this notation in every later chapter.
- If $m = p = 0$ there are no rules: an unconstrained problem, $\min_{\mathbf{x}} f(\mathbf{x})$ (Chapter 3.2).
- Rewriting rules: $\max f \to \min(-f)$; "$g \ge 0$" $\to$ "$-g \le 0$"; "$a \le b$" $\to$ "$a - b \le 0$".
- Problems get names from the shape of $f$, $g_i$, $h_j$: all linear $\Rightarrow$ linear program; quadratic $f$ with linear rules $\Rightarrow$ quadratic program; convex pieces $\Rightarrow$ convex problem (Chapter 3.6).
Why do we need it?
One standard shape means one theory and one set of solvers for every problem. Once your question is in this form, you can hand it to a general-purpose tool.
Where is it used?
Every optimization library (SciPy's minimize, CVXPY, solvers inside SVM packages) expects something like it, and every theorem in Chapters 3.7 to 3.10 (Lagrangian, KKT, duality) is stated for it.
How is it used?
Translate your story step by step: pick the decision variables, write one objective to minimise (flip signs if needed), write each limit as $g_i \le 0$ and each exact requirement as $h_j = 0$. Then check that every constraint is on the form "something compared with 0".
Equalities hide in disguise. "Probabilities add up to 1" is an equality ($\sum p_i - 1 = 0$). "Nothing negative" is an inequality ($-p_i \le 0$). Some people write an equality as two inequalities ($h \le 0$ and $-h \le 0$); mathematically fine, but solvers treat equalities much better when they are kept as equalities.
Library conventions differ. SciPy's minimize asks for inequality constraints as fun(x) >= 0, the opposite sign of our $g_i(\mathbf{x}) \le 0$. Always read the documentation.
Quick check: put "minimise $x^2 + y^2$, with $x + y = 3$ and $x \ge 1$" in general form.
$f(x,y) = x^2 + y^2$; equality $h_1(x,y) = x + y - 3 = 0$; inequality $g_1(x,y) = 1 - x \le 0$ (from $x \ge 1$, multiplying by $-1$ and moving everything left). No further rewriting is needed.
Worked example 1: the least fence for a field
A farmer wants a rectangular field of exactly 100 square metres, and wants to buy as little fence as possible. Long, thin fields need a lot of fence. A very wide, short field also needs a lot. Somewhere between "long and thin" and "wide and short" there must be a shape that uses the least fence. Which one?
- Decision variables. The width $x$ and the length $y$ (both positive).
- Objective. The fence goes round the whole rectangle: $P = 2x + 2y$. Minimise it.
- Constraint. Area must be exactly 100: $xy = 100$, that is $h(x, y) = xy - 100 = 0$. (Also $x \gt 0$, $y \gt 0$.)
- General form: $\min\; 2x + 2y$ s.t. $xy - 100 = 0$, $-x \le 0$, $-y \le 0$.
- Use the equality to remove a variable. From $xy = 100$ we get $y = 100/x$. Now the fence length depends on $x$ alone: $P(x) = 2x + \dfrac{200}{x}$ for $x \gt 0$. (The constraint is used up, and what remains is a one-variable problem.)
- Find where the slope is zero. $P'(x) = 2 - \dfrac{200}{x^2}$. Set it to 0: $2 = \dfrac{200}{x^2}$, so $x^2 = 100$ and $x = 10$ (we need $x \gt 0$). Then $y = 100/10 = 10$.
- Check it is a valley, not a hilltop. $P''(x) = \dfrac{400}{x^3}$, and $P''(10) = 0.4 \gt 0$ (curving up, like a smile). So $x = 10$ is a minimum. Compare with neighbours: $P(5) = 10 + 40 = 50$, $P(10) = 20 + 20 = 40$, $P(20) = 40 + 10 = 50$.
- It is the global minimum. $P(x) \to \infty$ both as $x \to 0^+$ ($200/x$ blows up) and as $x \to \infty$ ($2x$ blows up), and there is only one place where the slope is zero.
Answer: $\mathbf{x}^\star = (10, 10)$, a square, and $f^\star = 40$ metres of fence. Among all rectangles of a given area, the square needs the least fence.
What if there were no constraint? $\min 2x + 2y$ over positive $x, y$ has no answer: shrink both toward 0 and the fence disappears (the infimum 0 is never reached). The constraint is what makes the problem meaningful.
The recipe used here (substitute out an equality, then use calculus on what remains):
- Write the objective and the constraints.
- Solve the equality for one variable and substitute it into the objective.
- Set the derivative of the result to 0, and check the second derivative.
- Check the boundaries and compare with other candidates; then rebuild the other variables.
Substitution works when the equality is easy to solve. For harder constraints we need Lagrange multipliers (Chapter 3.8). The derivative test is explained fully in Chapter 3.2; derivatives themselves in the Calculus guide.
Why do we need it?
It shows the whole pipeline on a problem small enough to do on paper: model the situation, reduce it to one variable, find the flat spot, and verify. Every bigger method automates these same steps.
Where is it used?
Design problems (least material for a fixed volume), economics (cost minimisation), and as the "sanity check" for solvers: a good test of any code is that it reproduces the answer you derived by hand.
How is it used?
Model, substitute, differentiate, solve, check the second derivative, check the boundary. Then verify numerically by trying a few neighbouring values, as in step 7.
Quick check: redo the problem with area 64. What is the best rectangle and its fence length?
$y = 64/x$, $P(x) = 2x + 128/x$, $P'(x) = 2 - 128/x^2 = 0 \Rightarrow x^2 = 64 \Rightarrow x = 8$, $y = 8$. $P(8) = 16 + 16 = 32$. Again a square, with fence length $4\sqrt{64} = 32$.
Worked example 2: what happens when you add a constraint
Suppose the best spot on a map is a lovely lake at the centre of a park. Now a new rule says "you may not enter the north half". If the lake was in the south, nothing changes. If the lake was in the north, you can no longer stand there. The best you can do is stand as close as possible, on the border.
Adding a rule never helps: you were choosing from a bigger set before, so the best score can only stay the same or get worse. And when the rule does bite, the new best point usually lies on the edge of the allowed region.
Start with the unconstrained problem $\min_{x,y}\; f(x,y) = (x-3)^2 + (y-2)^2$.
- Both squares are $\ge 0$, so $f \ge 0$, with equality only at $(3, 2)$. So $\mathbf{x}^\star = (3,2)$ and $f^\star = 0$.
Now add the rule $x + y \le 4$, i.e. $g(x,y) = x + y - 4 \le 0$.
- Is $(3,2)$ feasible? $3 + 2 - 4 = 1 \gt 0$. No. So the old answer is gone.
- The new best point must lie on the border $x + y = 4$. (If it were strictly inside, it would be a local minimum of the unconstrained problem too, and the unconstrained problem has only one: $(3,2)$.)
- On the border, $y = 4 - x$, so $f = (x-3)^2 + (4 - x - 2)^2 = (x-3)^2 + (2-x)^2 = 2x^2 - 10x + 13$.
- Slope zero: $4x - 10 = 0$, so $x = 2.5$ and $y = 1.5$. Second derivative $4 \gt 0$: a minimum.
- Value: $f^\star = 0.5^2 + 0.5^2 = 0.5$.
Result: the optimum moved from $(3, 2)$ to $(2.5, 1.5)$ and the best score got worse, from $0$ to $0.5$. The constraint is now active ($g = 0$ at the answer). Geometrically: $(2.5, 1.5)$ is the closest point of the line to $(3,2)$; the distance is $1/\sqrt2$, and $f^\star$ is its square, $1/2$ ✓.
A loose constraint. With $x + y \le 6$ instead: $(3,2)$ has $3 + 2 = 5 \le 6$, so it is still feasible and still optimal. The constraint is inactive and changes nothing. In general, with $x + y \le c$ the best value is $f^\star(c) = \dfrac{(5 - c)^2}{2}$ when $c \lt 5$, and $0$ when $c \ge 5$.
Three facts to remember.
- Adding constraints can never improve the optimal value of a minimisation problem: $f^\star_{\text{constrained}} \ge f^\star_{\text{unconstrained}}$ (a smaller allowed set has a worse-or-equal best).
- If the unconstrained optimum is feasible, the constraint is inactive and the answer is unchanged.
- If it is not feasible, the constrained optimum usually lies on the boundary: an active constraint, $g_i(\mathbf{x}^\star) = 0$. Chapters 3.7 to 3.9 make this precise.
Why do we need it?
It shows why constrained problems need their own methods: the answer is no longer where the slope is zero, but where the objective "meets" the constraint. It also shows what a rule costs: the gap $f^\star(c) - 0$.
Where is it used?
Budget limits in planning (relaxing the budget lowers the cost), norm limits on weights (a tighter limit pushes ridge regression toward simpler models), and sensitivity analysis ("what is one extra unit of resource worth?", Chapter 3.10).
How is it used?
First solve the unconstrained problem. Check whether its answer is feasible. If yes, you are done. If not, look for the best point on the boundary of the feasible region.
Do not just "clip" the unconstrained answer. It is tempting to take $(3,2)$ and shrink it toward the origin until the rule holds, which gives $(2.4, 1.6)$ with $f = 0.36 + 0.16 = 0.52$. That is feasible, but slightly worse than the true constrained optimum $(2.5, 1.5)$ with $f = 0.5$. Constrained problems need real methods, not quick fixes.
Quick check: with the rule $x \le 1$ added to the unconstrained problem of this section, what are $\mathbf{x}^\star$ and $f^\star$?
The two variables are independent in $f = (x-3)^2 + (y-2)^2$. The unconstrained best $x = 3$ breaks $x \le 1$, so take the nearest allowed value $x = 1$ (the function falls toward 3, so the best allowed point is the border). $y = 2$ is unaffected. $\mathbf{x}^\star = (1, 2)$ and $f^\star = (1-3)^2 + 0 = 4$.
Training a model is an optimization problem core
Here is the whole story of machine-learning training in one sentence: the model has numbers (the weights); a loss function measures how wrong the model is on the data; training searches for the weights that make the loss as small as possible.
That is exactly the landscape picture: each possible weight vector is a spot on the ground, the loss is the height there, and training is a hike to a low spot. Every word of this chapter now has a job.
Fit the line $\hat y = w x + b$ to the data $(1,2)$, $(2,3)$, $(3,5)$ by minimising the total squared error
$$L(w, b) = (w + b - 2)^2 + (2w + b - 3)^2 + (3w + b - 5)^2 = 14w^2 + 12wb + 3b^2 - 46w - 20b + 38.$$- Slope in $w$ (partial derivative): $\partial L/\partial w = 28w + 12b - 46$. Slope in $b$: $\partial L/\partial b = 12 w + 6 b - 20$.
- Set both to 0. From the second equation, $b = (20 - 12w)/6 = \tfrac{10}{3} - 2w$. Put that in the first: $28w + 12(\tfrac{10}{3} - 2w) - 46 = 28w + 40 - 24w - 46 = 4w - 6 = 0$, so $w = 1.5$.
- Then $b = \tfrac{10}{3} - 3 = \tfrac13$. The best line is $\hat y = 1.5x + \tfrac13$, with $L^\star = \tfrac16 \approx 0.167$.
Now add a constraint: "the weight may not exceed 1" ($g(w,b) = w - 1 \le 0$). The unconstrained optimum $w = 1.5$ is infeasible, so the answer moves to the border $w = 1$. With $w = 1$, minimising over $b$ gives $b = \tfrac{10}{3} - 2 = \tfrac43$ and $L = \tfrac23 \approx 0.667$, worse than $\tfrac16$, exactly as the "constraints never help" fact predicts.
Almost every supervised-learning problem is empirical risk minimisation: given $N$ training examples $(\mathbf{x}_i, y_i)$ and a model $\hat y = m(\mathbf{x}; \boldsymbol{\theta})$ with weights $\boldsymbol{\theta}$,
$$\min_{\boldsymbol{\theta}} \;\; \frac{1}{N}\sum_{i=1}^{N} \ell\bigl(m(\mathbf{x}_i; \boldsymbol{\theta}),\, y_i\bigr) \;\;(+\; \text{optional penalty or constraint on } \boldsymbol{\theta}),$$where $\ell$ is a per-example loss (squared error, cross-entropy…). The dictionary from optimization words to ML words:
| Optimization word | What it is in ML training |
|---|---|
| objective function $f$ | the training loss (average of the per-example losses, plus optionally a penalty) |
| decision variables $\mathbf{x}$ | the model's parameters: weights and biases $\boldsymbol{\theta}$ |
| problem parameters | the training data $(\mathbf{x}_i, y_i)$ and the fixed model architecture |
| hyperparameters | learning rate, regularisation strength $\lambda$, batch size, number of layers (set by you, not by the optimizer) |
| constraints $g_i, h_j$ | norm limits $\|\mathbf{w}\| \le r$ (ridge, weight clipping), sparsity limits, probabilities summing to 1, SVM margin rules |
| feasible region | the set of allowed weights (for example a ball of radius $r$) |
| optimal solution $\mathbf{x}^\star$ | the best weights $\boldsymbol{\theta}^\star$ |
| local vs global optimum | deep networks: many local minima; linear and logistic regression: the single (global) minimum, thanks to convexity (Chapter 3.6) |
| saddle point | flat regions of a deep network's loss where training can stall |
| maximum | likelihood or reward; handled by minimising the negative |
Why do we need it?
Seeing training as optimization makes it a well-defined mathematical problem. Then every tool of optimization (gradients, convexity, constraints, second-order information) can be used to train better and to explain what goes wrong.
Where is it used?
Linear and logistic regression, support vector machines, neural networks and transformers, matrix factorisation for recommendation, k-means, and Gaussian-mixture fitting: all minimise a loss over parameters.
How is it used?
Choose a model and a loss (the objective), pick starting weights, then let an optimizer (Chapters 3.3 and 3.4) repeatedly improve them. Add penalties or constraints for regularisation (Chapter 3.11). Tune the hyperparameters in an outer loop.
The loss on training data is a stand-in. What we really want is a model that works on new data. A tiny training loss does not guarantee that. Constraints and penalties (regularisation, Chapter 3.11) are the main tool for trading a little training loss for better behaviour on new data.
Quick check: in training a logistic-regression classifier, name the objective, the decision variables and the data.
Objective: the average cross-entropy loss. Decision variables: the weight vector $\mathbf{w}$ and bias $b$. Data (problem parameters): the training examples and their labels. Hyperparameters: for example the learning rate and the strength of any regularisation.
Recap, cheat sheet and practice
- Optimization = choose decision variables $\mathbf{x}$ to make an objective $f(\mathbf{x})$ as small (or large) as possible, subject to constraints.
- Three kinds of numbers: decision variables (the optimizer changes them), problem parameters / data (fixed), hyperparameters (you set them before the search). In ML, "parameters" usually means the decision variables (weights).
- General form: $\min f(\mathbf{x})$ s.t. $g_i(\mathbf{x}) \le 0$ (inequalities) and $h_j(\mathbf{x}) = 0$ (equalities).
- The feasible region is every point that obeys all constraints. A feasible solution obeys them; the optimal solution $\mathbf{x}^\star$ is the best feasible one, with optimal value $f^\star$. It may not exist or may not be unique.
- $\max f = -\min(-f)$, and the best point is the same. Adding a constraint can only keep the optimal value the same or make it worse.
- A local optimum is best nearby; a global optimum is best everywhere. A saddle point has slope zero but is neither a local min nor a local max.
- Training a model is optimization: objective = loss, decision variables = weights, constraints = norm limits and the like.
Cheat sheet
| Term | Symbol | Plain meaning |
|---|---|---|
| Objective | $f(\mathbf{x})$ | the one number we minimise (the landscape height) |
| Decision variables | $\mathbf{x} \in \mathbb{R}^n$ | the dials we may turn |
| Inequality constraint | $g_i(\mathbf{x}) \le 0$ | a limit (at most / at least) |
| Equality constraint | $h_j(\mathbf{x}) = 0$ | an exact requirement |
| Feasible region | $\mathcal{F}$ | all points that obey every rule |
| Optimal solution / value | $\mathbf{x}^\star$, $f^\star$ | the best feasible point and its score |
| argmin vs min | $\arg\min f$, $\min f$ | the input that wins vs the number it achieves |
| Max to min | $\max f = -\min(-f)$ | flip the picture upside down |
| Local min | $f(\mathbf{x}^\star) \le f(\mathbf{x})$ nearby | best in the neighbourhood |
| Global min | $f(\mathbf{x}^\star) \le f(\mathbf{x})$ for all feasible $\mathbf{x}$ | best anywhere |
| Saddle point | $\nabla f = \mathbf{0}$, mixed curvature | flat, but up one way and down another |
| Active / inactive | $g_i = 0$ / $g_i \lt 0$ | on the edge / with room to spare |
import numpy as np
from scipy.optimize import minimize
# 1. objective f(x, y) = (x - 3)^2 + (y - 2)^2, no constraints
f = lambda v: (v[0] - 3) ** 2 + (v[1] - 2) ** 2
res = minimize(f, x0=[0.0, 0.0])
print(np.round(res.x, 3), round(res.fun, 6)) # [3. 2.] 0.0
# 2. add x + y <= 4. SciPy wants fun(v) >= 0 for 'ineq', so pass -g(v) = 4 - x - y
cons = [{'type': 'ineq', 'fun': lambda v: 4 - v[0] - v[1]}]
res = minimize(f, x0=[0.0, 0.0], constraints=cons)
print(np.round(res.x, 3), round(res.fun, 6)) # [2.5 1.5] 0.5 (the optimum moved to the border)
# 3. the fence: min 2x + 2y s.t. x*y = 100, x, y > 0
fence = lambda v: 2 * v[0] + 2 * v[1]
eq = [{'type': 'eq', 'fun': lambda v: v[0] * v[1] - 100}]
res = minimize(fence, x0=[5.0, 20.0], constraints=eq, bounds=[(1e-3, None), (1e-3, None)])
print(np.round(res.x, 3), round(res.fun, 3)) # [10. 10.] 40.0
# 4. max f = -min(-f): P(x) = 3 - (x - 1)^2
P = lambda x: 3 - (x[0] - 1) ** 2
res = minimize(lambda x: -P(x), x0=[3.0])
print(np.round(res.x, 3), -round(res.fun, 3)) # [1.] 3.0 (maximiser 1, maximum value 3)
# 5. local vs global: random restarts on a function with several valleys
g = lambda x: np.sin(1.5 * x[0]) + x[0] ** 2 / 10
rng = np.random.default_rng(0)
ends = sorted({round(float(minimize(g, x0=[s]).x[0]), 2) for s in rng.uniform(-5, 5, 12)})
print(ends) # [-4.78, -0.96, 2.88] three local minima
print(min(ends, key=lambda e: g([e]))) # -0.96 the global one has the lowest g
1. You fit $\hat y = w x + b$ to a dataset by minimising the squared error. Which of these are the decision variables?
2. Which is the correct general-form version of the rule "$x_1 + x_2 \ge 3$"?
3. Which statement is always true?
4. Which statement about local and global minima is correct?
5. What can you say about $\min_x e^x$ over all real $x$?
6. You add one more constraint to a minimisation problem. The optimal value will...
Practice problems
A. In the three-route example (A: 40 min, no toll; B: 30 min, toll 2; C: 25 min, toll 5), for which values $v$ (dollars per minute) is each route the cheapest?
Costs: A $= 40v$, B $= 30v + 2$, C $= 25v + 5$.
A beats B when $40v \lt 30v + 2$, i.e. $v \lt 0.2$. B beats C when $30v + 2 \lt 25v + 5$, i.e. $v \lt 0.6$. So: A is best for $v \lt 0.2$, B for $0.2 \lt v \lt 0.6$, C for $v \gt 0.6$ (ties exactly at 0.2 and 0.6). Check $v = 0.25$: A $= 10$, B $= 9.5$, C $= 11.25$, so B ✓.
B. A farmer fences a rectangle of area 100, but one side runs along a river and needs no fence. Find the best width and the fence length.
Let $x$ be the two sides perpendicular to the river and $y$ the side parallel to it. Fence $P = 2x + y$ with $xy = 100$, so $y = 100/x$ and $P(x) = 2x + 100/x$.
$P'(x) = 2 - 100/x^2 = 0 \Rightarrow x^2 = 50 \Rightarrow x = \sqrt{50} \approx 7.07$. Then $y = 100/\sqrt{50} \approx 14.14$. $P'' = 200/x^3 \gt 0$, so it is a minimum. $P^\star = 2\sqrt{50} + 100/\sqrt{50} = 4\sqrt{50} \approx 28.28$ metres. (Not a square any more: the long side is twice the short side.)
C. Minimise $(x-3)^2 + (y-2)^2$ subject to $x \le 1$. Is the constraint active?
The objective is a sum of a function of $x$ and a function of $y$, so we can treat them separately. For $x \le 1$ the best $x$ is the border $x = 1$ (the parabola $(x-3)^2$ is still falling there); $y = 2$ is free. So $\mathbf{x}^\star = (1, 2)$, $f^\star = 4 + 0 = 4$. The unconstrained optimum $(3,2)$ breaks $x \le 1$, and $g(\mathbf{x}^\star) = 1 - 1 = 0$: the constraint is active.
D. Write in general form: "maximise $5a + 4b$ subject to $6a + 4b \le 24$, $a + 2b \le 6$, $a \ge 0$, $b \ge 0$."
$\min_{a,b}\; -5a - 4b$ subject to $g_1 = 6a + 4b - 24 \le 0$, $g_2 = a + 2b - 6 \le 0$, $g_3 = -a \le 0$, $g_4 = -b \le 0$. (There are no equality constraints, so no $h_j$.)
E. $P(x) = -x^2 + 6x - 5$. Find $\max P$ and the minimiser of $-P$. Check $\max P = -\min(-P)$.
Complete the square: $P(x) = -(x-3)^2 + 4$. So $\max P = 4$ at $x = 3$. Then $-P(x) = (x-3)^2 - 4$ has minimum $-4$ at $x = 3$ (the same point). Indeed $4 = -(-4)$ ✓.
F. A small network has layers 784 → 100 → 10, each with weights and biases. How many decision variables does training have?
Layer 1: $784 \cdot 100 = 78{,}400$ weights $+ 100$ biases. Layer 2: $100 \cdot 10 = 1000$ weights $+ 10$ biases. Total $78{,}400 + 100 + 1000 + 10 = \mathbf{79{,}510}$. Training is a search in $\mathbb{R}^{79510}$ for the point of lowest loss.
Unconstrained Optimization
No rules, just a landscape. How do you find the bottom of a valley without trying every spot? You look for places where the ground is level (the gradient is zero), and then you check how the ground curves (the Hessian). Those two ideas are the whole story of this chapter, and they explain every gradient-based method that follows.
- Solve one-dimensional problems: find where $f'(x) = 0$ and check the sign of $f''(x)$
- Extend this to many variables with the gradient and the Hessian
- Find critical points ($\nabla f = \mathbf{0}$) and know why this first-order condition is necessary but not sufficient
- Use the second-order conditions to sort critical points into local minima, local maxima and saddle points, and know when the test is inconclusive
- Write down the exact minimiser of a quadratic $\tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x}$ by solving $A\mathbf{x} = \mathbf{b}$
What you should know first. We use derivatives, the gradient and the Hessian from the Calculus guide, and positive-definite matrices from the Linear Algebra guide. Each time one appears we remind you in a sentence, so you can read on even if those chapters are not fresh. Notation: $\mathbf{x} \in \mathbb{R}^n$, the gradient $\nabla f$ is a column vector, the Hessian $\nabla^2 f$ (also written $H$) is a symmetric $n \times n$ matrix, and a "stationary" or "critical" point is written $\hat{\mathbf{x}}$.
One-dimensional optimization core
Ride a bike along a hilly road. At the top of a hill, and at the bottom of a valley, the road is level for a moment: it has stopped going up and has not yet started going down (or the other way round). Everywhere else the road is sloping.
So to find the hilltops and valley bottoms, look for the places where the slope is zero. To tell a hilltop from a valley bottom, look at how the slope changes as you pass: on a valley bottom the slope goes from negative (downhill) to positive (uphill), so the road curves upward, like a smile. On a hilltop it curves downward, like a frown. The "curving" is measured by the second derivative.
Find the local minima and maxima of $f(x) = x^3 - 3x^2 - 9x + 5$.
- Slope: $f'(x) = 3x^2 - 6x - 9 = 3(x^2 - 2x - 3) = 3(x - 3)(x + 1)$.
- Level ground where $f'(x) = 0$: $x = 3$ or $x = -1$.
- Curving: $f''(x) = 6x - 6$.
- At $x = -1$: $f'' = -12 \lt 0$, a frown, so a local maximum, with $f(-1) = -1 - 3 + 9 + 5 = 10$.
- At $x = 3$: $f'' = 12 \gt 0$, a smile, so a local minimum, with $f(3) = 27 - 27 - 27 + 5 = -22$.
Global? On the whole real line, no: a cubic falls to $-\infty$ on the left and rises to $+\infty$ on the right. But on a limited interval, say $-4 \le x \le 4$, compare all candidates, including the two ends: $f(-4) = -64 - 48 + 36 + 5 = -71$, $f(-1) = 10$, $f(3) = -22$, $f(4) = 64 - 48 - 36 + 5 = -15$. The global minimum on this interval is $-71$, at the left end $x = -4$ (the local minimum at $x=3$ is not global), and the global maximum is $10$ at $x = -1$.
Recipe for a smooth function of one variable.
- First-order step. Solve $f'(x) = 0$. The solutions are the critical points (candidates).
- Second-order step. At each candidate $\hat x$ look at $f''(\hat x)$: $f''(\hat x) \gt 0 \Rightarrow$ local minimum; $f''(\hat x) \lt 0 \Rightarrow$ local maximum; $f''(\hat x) = 0 \Rightarrow$ the test is inconclusive (see the later section on semidefinite cases).
- Global step. To find the global optimum on an interval $[a, b]$, compare $f$ at every critical point and at the endpoints $a$ and $b$. On all of $\mathbb{R}$, also check what $f$ does as $x \to \pm\infty$.
The alternative to step 2 is the first-derivative test: if $f'$ changes from negative to positive at $\hat x$, it is a minimum; from positive to negative, a maximum; no change of sign, neither. (Derivatives are explained in the Calculus guide.)
Why do we need it?
It turns "search every point" into "solve one equation": the best points of a smooth function can only be where the slope is zero (or at the ends of the allowed interval). That shrinks infinitely many candidates to a handful.
Where is it used?
One-parameter fits (the best scale factor, a single learning rate), the line search inside many algorithms (find the best step size along a direction), and every one-variable model in economics and engineering (profit versus price, cost versus size).
How is it used?
Differentiate, solve $f' = 0$, test the sign of $f''$ at each solution, then compare the values at the candidates and the interval ends. Software does the same with numbers when the algebra is too hard.
Do not forget the ends. On a closed interval the best point is often not a critical point at all, but an endpoint (like $x = -4$ above). At an endpoint the slope need not be zero. This is our first look at how constraints change the answer (Chapter 3.1, and Chapters 3.7 to 3.9).
"Local" is only about a small neighbourhood. The local minimum at $x = 3$ above has value $-22$, yet the function reaches $-71$ elsewhere in the interval.
Quick check: $f(x) = x^2 - 4x + 7$. Where is the minimum, and what is its value?
$f'(x) = 2x - 4 = 0 \Rightarrow x = 2$; $f''(x) = 2 \gt 0$ (a smile), so it is a minimum. $f(2) = 4 - 8 + 7 = 3$. It is also global, because the parabola opens upward.
Multivariable optimization core
With one variable there are only two directions to walk: left or right. With two variables you stand on a hillside and can walk in every compass direction. With a million weights you have a million directions (and every mixture of them).
For a point to be the bottom of a valley, every direction must lead uphill. A single downhill direction is enough to rule it out. The good news is that we can reduce the multi-direction question to many one-direction questions: pick a direction, walk a straight line along it, and you get a plain one-variable curve (a "slice"). The one-variable tools from the previous section apply to each slice.
$f(x, y) = (x-1)^2 + 2(y+2)^2$. At the point $(1, -2)$ the value is $f = 0$. Walk away from it by a distance $d$ in a few directions:
- Right: $f(1+d, -2) = d^2$. Up: $f(1, -2+d) = 2d^2$. Diagonal: $f(1+d, -2+d) = d^2 + 2d^2 = 3d^2$.
- In all of them $f \ge 0 = f(1,-2)$, and in fact $f(1+a, -2+b) = a^2 + 2b^2 \ge 0$ for every direction $(a,b)$. So $(1,-2)$ is a minimum, and we checked all directions at once by writing $f$ as a sum of squares.
For a harder function we cannot expand like this. Instead we use the tools of the next sections: the gradient (first-order information) and the Hessian (second-order information).
A multivariable unconstrained optimization problem is
$$\min_{\mathbf{x} \in \mathbb{R}^n} f(\mathbf{x}), \qquad f: \mathbb{R}^n \to \mathbb{R}.$$A point $\mathbf{x}^\star$ is a local minimum if $f(\mathbf{x}^\star) \le f(\mathbf{x})$ for all $\mathbf{x}$ in some small ball around it, that is, in every direction at once (Chapter 3.1). The two one-variable tools generalise as follows:
| One variable | Many variables |
|---|---|
| slope $f'(x)$ (a number) | gradient $\nabla f(\mathbf{x})$ (a vector of $n$ slopes) |
| curvature $f''(x)$ (a number) | Hessian $\nabla^2 f(\mathbf{x})$ (an $n \times n$ matrix of curvatures) |
| $f'' \gt 0$: smile | Hessian positive definite: bowl in every direction |
Why do we need it?
Real problems have many unknowns. The key trick, "a multi-variable problem is many one-variable slices", gives us a way to reason about hundreds or millions of variables using pictures from one or two.
Where is it used?
Every training problem: a linear model with $n$ weights is a search in $\mathbb{R}^n$, and a neural network with millions of weights is a search in millions of dimensions. Line searches in optimizers literally walk along one slice.
How is it used?
Pick a point and a direction $\mathbf{u}$, then study the one-variable function $\phi(t) = f(\mathbf{x} + t\mathbf{u})$. Its slope at $t=0$ is $\nabla f(\mathbf{x})^\top \mathbf{u}$ and its curvature is $\mathbf{u}^\top \nabla^2 f(\mathbf{x})\, \mathbf{u}$ (the next sections derive both).
Quick check: for $f(x,y) = x^2 - y^2$, the slice along the $x$-axis is $f(t, 0) = t^2$, which smiles. Can we conclude that the origin is a minimum?
No. A minimum needs every direction to go up. Along the $y$-axis the slice is $f(0, t) = -t^2$, which frowns. One good direction proves nothing; one bad direction is enough to rule the point out. (The origin is a saddle point.)
The gradient: a compass for uphill core
Stand on a hillside in the fog and feel the ground with your feet. You can tell which way is steepest uphill, and how steep it is. The gradient is exactly that: an arrow that points straight uphill, with a length equal to the steepness. Walk the opposite way, and you go straight downhill, as fast as possible.
On a contour map (lines of equal height) the gradient is always at a right angle to the contour line through your spot, pointing toward higher ground. On level ground the gradient has length zero.
$f(x, y) = x^2 + 3y^2$ at the point $(2, 1)$. The slope in the $x$ direction (treating $y$ as fixed) is $2x$; in the $y$ direction it is $6y$.
- $\nabla f(x,y) = \begin{bmatrix} 2x \\ 6y \end{bmatrix}$, so $\nabla f(2,1) = \begin{bmatrix} 4 \\ 6 \end{bmatrix}$.
- Steepness: $\|\nabla f\| = \sqrt{16 + 36} = \sqrt{52} \approx 7.21$.
- Check "downhill": $f(2,1) = 4 + 3 = 7$. Take a small step against the gradient, $\mathbf{x} - 0.01\nabla f = (2 - 0.04,\; 1 - 0.06) = (1.96, 0.94)$: $f = 3.8416 + 2.6508 = 6.4924$. The value dropped by $0.5076$. The straight-line prediction: $0.01 \cdot \|\nabla f\|^2 = 0.01 \cdot 52 = 0.52$ ✓ (close; the small gap comes from the curve bending).
For $f: \mathbb{R}^n \to \mathbb{R}$, the gradient is the column vector of partial derivatives:
$$\nabla f(\mathbf{x}) = \begin{bmatrix} \partial f / \partial x_1 \\ \vdots \\ \partial f / \partial x_n \end{bmatrix}.$$- $\nabla f$ points in the direction of steepest ascent; $-\nabla f$ in the direction of steepest descent. Its length $\|\nabla f\|$ is the steepness.
- The slope in any unit direction $\mathbf{u}$ is the dot product $\nabla f(\mathbf{x})^\top \mathbf{u}$ (the directional derivative).
- $\nabla f$ is perpendicular to the contour line through $\mathbf{x}$.
All of this is derived in the Calculus guide (the gradient) and steepest ascent and descent.
Why do we need it?
It tells us, from information at one spot, which way is downhill: the thing every optimization algorithm wants to know. It also detects the flat spots (where it is zero), which are the only places a smooth minimum can hide.
Where is it used?
Gradient descent and all its relatives (momentum, Adam), backpropagation (which is an efficient way to compute the gradient of a network's loss), and the optimality conditions of this chapter.
How is it used?
Compute $\nabla f$ at the current point (by formula, by backpropagation, or numerically). Either move against it (a step of gradient descent, Chapter 3.3) or set it to zero to find candidate optima (this chapter).
The gradient is local. It describes the slope right here, not the shape far away. A steep arrow does not mean the minimum is far, and a short arrow does not mean it is near (near a saddle it is short, too).
Shape matters. $\nabla f$ is a column vector with the same number of entries as $\mathbf{x}$ (we use this convention in every chapter). It is not the Jacobian of a vector-valued function.
Quick check: $f(x,y) = x^2 y$. Find $\nabla f$ at $(1, 3)$.
$\partial f/\partial x = 2xy = 6$ and $\partial f/\partial y = x^2 = 1$, so $\nabla f(1,3) = [6, 1]^\top$.
Critical points: where the ground is level core
A critical point (also called a stationary point) is a spot where the ground is level in every direction: the gradient is the zero vector. A marble placed there would not roll.
Valley bottoms, hilltops and mountain passes are all level spots. So every candidate for a minimum or maximum is on the list of critical points. Finding the list is step one; deciding which is which is step two (the next sections).
Find the critical points of $f(x, y) = x^3 - 3x + y^2 - 2y$.
- Gradient: $\nabla f = \begin{bmatrix} 3x^2 - 3 \\ 2y - 2 \end{bmatrix}$.
- Set both entries to zero: $3x^2 - 3 = 0$ gives $x = 1$ or $x = -1$; $2y - 2 = 0$ gives $y = 1$.
- So there are two critical points: $(1, 1)$ and $(-1, 1)$.
- Values: $f(1,1) = 1 - 3 + 1 - 2 = -3$ and $f(-1, 1) = -1 + 3 + 1 - 2 = 1$.
We cannot yet say what kind each one is. (We will finish this example in the second-order section: one is a minimum, the other a saddle.)
$\hat{\mathbf{x}}$ is a critical point of a differentiable $f$ if
$$\nabla f(\hat{\mathbf{x}}) = \mathbf{0}, \qquad\text{that is,}\qquad \frac{\partial f}{\partial x_1}(\hat{\mathbf{x}}) = 0, \;\dots,\; \frac{\partial f}{\partial x_n}(\hat{\mathbf{x}}) = 0.$$This is a system of $n$ equations in $n$ unknowns. Critical points can be minima, maxima, saddle points, or flat plateaus. Every unconstrained local minimum or maximum of a differentiable function is a critical point (the next section proves it).
Why do we need it?
It cuts an infinite search to a short list. Instead of checking every point of the landscape, we only need to look at the places where all slopes vanish.
Where is it used?
The normal equations of least squares (set the gradient of the squared error to zero), maximum-likelihood estimates (set the gradient of the log-likelihood to zero), and the stopping rule of gradient descent ($\|\nabla f\|$ almost zero, Chapter 3.5).
How is it used?
Compute $\nabla f$, set it to $\mathbf{0}$, and solve for $\mathbf{x}$ (by hand for small problems, by a numerical solver otherwise). Make a list of all solutions, then classify each one.
A critical point is a candidate, not an answer. Some critical points are minima, some maxima, some saddles. Always classify them (second-order conditions) or compare their values.
Solving $\nabla f = \mathbf{0}$ can be hard. For most real models (a neural network, logistic regression) the equations cannot be solved by hand, which is why we need iterative algorithms such as gradient descent (Chapter 3.3).
Quick check: find the critical points of $f(x, y) = x^2 + xy + y^2 - 3x$.
$\nabla f = [2x + y - 3,\; x + 2y]^\top = \mathbf{0}$. From the second equation $x = -2y$. Put it in the first: $-4y + y - 3 = 0$, so $y = -1$ and $x = 2$. One critical point: $(2, -1)$.
First-order conditions: necessary, but not sufficient core
Why must the slope be zero at a valley bottom? Because if it were not zero, there would be a downhill direction, and you could simply walk a tiny bit that way and end up lower. So a point with a non-zero slope cannot be the lowest point nearby.
That makes "gradient equals zero" a necessary condition: every minimum must satisfy it. But it is not sufficient: a hilltop, a saddle and a plateau also satisfy it. It rules points out; it does not rule them in.
Three one-variable functions, all with $f'(0) = 0$:
- $f(x) = x^2$: $f'(0) = 0$ and $0$ is a minimum.
- $f(x) = -x^2$: $f'(0) = 0$ and $0$ is a maximum.
- $f(x) = x^3$: $f'(0) = 0$ and $0$ is neither: the function keeps rising right through the flat spot ($f(-0.1) = -0.001$, $f(0.1) = 0.001$).
Proof that a minimum must have zero gradient. Suppose $\mathbf{x}^\star$ is a local minimum and $\nabla f(\mathbf{x}^\star) \ne \mathbf{0}$. Walk a small step $t \gt 0$ in the downhill direction $\mathbf{d} = -\nabla f(\mathbf{x}^\star)$. By the first-order Taylor approximation,
$$f(\mathbf{x}^\star + t\mathbf{d}) \approx f(\mathbf{x}^\star) + t\,\nabla f(\mathbf{x}^\star)^\top \mathbf{d} = f(\mathbf{x}^\star) - t\,\|\nabla f(\mathbf{x}^\star)\|^2 \lt f(\mathbf{x}^\star)$$for small $t$, which contradicts "minimum". So $\nabla f(\mathbf{x}^\star) = \mathbf{0}$. ∎ (Numerical check: for $f = x^2 + 3y^2$ at $(2,1)$ a downhill step of 0.01 lowered $f$ from 7 to 6.4924, as computed in the gradient section.)
First-order necessary condition. If $\mathbf{x}^\star$ is a local minimum (or local maximum) of a differentiable function $f$ and $\mathbf{x}^\star$ lies in the interior of the allowed set (here: there are no constraints), then
$$\nabla f(\mathbf{x}^\star) = \mathbf{0}.$$- Necessary: every unconstrained local optimum satisfies it.
- Not sufficient: a point with $\nabla f = \mathbf{0}$ may be a minimum, a maximum, a saddle or a plateau. We need second-order information to decide.
- Only for unconstrained problems. If the optimum sits on the edge of the allowed set, its gradient need not be zero (see the last preset in the widget). Handling that case is the job of Chapters 3.7 to 3.9.
Why do we need it?
It is the rule that makes optimisation by calculus possible: if you want a smooth unconstrained minimum, you only have to look at the points where the gradient is zero.
Where is it used?
Deriving closed-form solutions (least squares, ridge regression, Gaussian maximum likelihood), the stopping test of every gradient-based optimizer ("stop when $\|\nabla f\|$ is tiny"), and the starting point of the theory of constrained optimization.
How is it used?
Use it to reject candidates (non-zero gradient: not an optimum) and to find candidates (solve $\nabla f = \mathbf{0}$). Never use it alone to accept a candidate.
"Gradient zero" is not "optimum". This is the most common mistake in the whole subject. If a method reports "converged: gradient is zero", the point may be a saddle or a maximum (rare in practice with gradient descent, but possible).
Non-smooth functions break it. $f(x) = |x|$ has its minimum at $0$, where the derivative does not even exist. The first-order condition assumes differentiability.
Quick check: $f'(3) = 0$ for some function. Can you conclude that $3$ is a local minimum?
No. $f'(3) = 0$ only says $3$ is a candidate. It could be a minimum ($x^2$ shifted), a maximum, or an inflection like $(x-3)^3$. Check $f''(3)$ or the sign change of $f'$.
The Hessian: curvature in every direction core
The gradient tells you how steep the ground is. The Hessian tells you how the ground bends: does it curve up like a bowl, down like a hill, or both ways like a saddle? It is the many-variable version of the second derivative.
Why it matters at a flat spot: if the ground is level, the first thing that decides whether you are at the bottom of a bowl or the top of a hill is the bending. And the bending can be different in different directions: a Pringle chip curves up one way and down the other.
For $f(x, y) = x^3 - 3x + y^2 - 2y$ the second derivatives are $\partial^2 f/\partial x^2 = 6x$, $\partial^2 f/\partial y^2 = 2$, and the mixed ones $\partial^2 f/\partial x \partial y = 0$. So
$$\nabla^2 f(x,y) = \begin{bmatrix} 6x & 0 \\ 0 & 2 \end{bmatrix}, \quad \nabla^2 f(1,1) = \begin{bmatrix} 6 & 0 \\ 0 & 2 \end{bmatrix}, \quad \nabla^2 f(-1,1) = \begin{bmatrix} -6 & 0 \\ 0 & 2 \end{bmatrix}.$$At $(1,1)$ the ground curves up by 6 in the $x$ direction and by 2 in the $y$ direction: a bowl. At $(-1,1)$ it curves down by 6 in $x$ and up by 2 in $y$: a saddle.
A tilted example. $H = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}$. Along $\mathbf{u} = [1, 0]^\top$: $\mathbf{u}^\top H \mathbf{u} = 2$. Along $\mathbf{u} = \tfrac{1}{\sqrt2}[1, 1]^\top$: $H\mathbf{u} = \tfrac{1}{\sqrt2}[3, 3]^\top$, so $\mathbf{u}^\top H \mathbf{u} = \tfrac12(3 + 3) = 3$. Along $\tfrac{1}{\sqrt2}[1, -1]^\top$: $H\mathbf{u} = \tfrac{1}{\sqrt2}[1, -1]^\top$ and $\mathbf{u}^\top H\mathbf{u} = 1$. The eigenvalues of $H$ are $3$ and $1$: the largest and smallest curvatures.
The Hessian of $f: \mathbb{R}^n \to \mathbb{R}$ at $\mathbf{x}$ is the symmetric $n \times n$ matrix of second partial derivatives:
$$\nabla^2 f(\mathbf{x}) = \begin{bmatrix} \dfrac{\partial^2 f}{\partial x_1^2} & \cdots & \dfrac{\partial^2 f}{\partial x_1 \partial x_n} \\ \vdots & \ddots & \vdots \\ \dfrac{\partial^2 f}{\partial x_n \partial x_1} & \cdots & \dfrac{\partial^2 f}{\partial x_n^2} \end{bmatrix}, \qquad H_{ij} = \frac{\partial^2 f}{\partial x_i\, \partial x_j}.$$- Symmetric ($H_{ij} = H_{ji}$) when the second derivatives are continuous.
- Curvature along a direction. Along a unit vector $\mathbf{u}$, the slice $\phi(t) = f(\mathbf{x} + t\mathbf{u})$ has $\phi'(0) = \nabla f^\top\mathbf{u}$ and $\phi''(0) = \mathbf{u}^\top H \mathbf{u}$.
- Eigenvalues. $\mathbf{u}^\top H \mathbf{u}$ always lies between the smallest and the largest eigenvalue of $H$; the extremes are reached along the eigenvectors (the principal directions of curvature).
- The second-order Taylor picture. Near $\mathbf{x}$, for a small step $\mathbf{d}$: $$f(\mathbf{x} + \mathbf{d}) \approx f(\mathbf{x}) + \nabla f(\mathbf{x})^\top \mathbf{d} + \tfrac12\, \mathbf{d}^\top H(\mathbf{x})\, \mathbf{d}.$$
See the Hessian in the Calculus guide, and for eigenvalues the Linear Algebra guide.
Why do we need it?
At a flat spot the gradient is zero and says nothing more. The Hessian is the next piece of information: it decides bowl, hill or saddle, and it measures how fast the slope changes (which limits the safe learning rate).
Where is it used?
The second-order test of this chapter, Newton's method and quasi-Newton methods (Chapter 3.12), the learning-rate limit $\eta \lt 2/\lambda_{\max}$ (Chapter 3.3), and the study of how flat or sharp a minimum of a neural network is.
How is it used?
Compute it (or approximate it), take its eigenvalues, and read the signs: all positive (bowl), all negative (hill), mixed (saddle). In big models we avoid forming it and use products $H\mathbf{v}$ instead.
Direction matters. There is no single "second derivative" in many dimensions: the curvature depends on the direction, and the Hessian packages all those curvatures in one matrix. For a unit direction, use $\mathbf{u}^\top H \mathbf{u}$; for a non-unit step $\mathbf{d}$, $\mathbf{d}^\top H \mathbf{d} = \|\mathbf{d}\|^2\, \mathbf{u}^\top H \mathbf{u}$.
The Hessian changes from point to point (unless $f$ is quadratic): evaluate it at the point you care about, for instance at a critical point.
Quick check: find the Hessian of $f(x, y) = x^2 y$ at $(1, 2)$.
$f_x = 2xy$, $f_y = x^2$. Then $f_{xx} = 2y$, $f_{yy} = 0$, $f_{xy} = f_{yx} = 2x$. At $(1,2)$: $H = \begin{bmatrix} 4 & 2 \\ 2 & 0 \end{bmatrix}$. Its determinant is $-4 \lt 0$, so one eigenvalue is positive and one negative.
Second-order conditions: bowl, hill or saddle? core
You are standing on level ground (the gradient is zero). Now look at how the ground bends, in every direction:
- It curves up in every direction: you are in a bowl, so a local minimum. (Hessian positive definite.)
- It curves down in every direction: you are on a hilltop, so a local maximum. (Hessian negative definite.)
- It curves up in some directions and down in others: a saddle point. (Hessian indefinite.)
- It is perfectly flat in some direction and does not bend the opposite way anywhere else: the test cannot tell. (Hessian semidefinite: all eigenvalues $\ge 0$, or all $\le 0$, and at least one of them equals 0.)
Finish the example $f(x, y) = x^3 - 3x + y^2 - 2y$ with critical points $(1,1)$ and $(-1,1)$ and Hessian $\mathrm{diag}(6x,\, 2)$.
- At $(1,1)$: $H = \mathrm{diag}(6, 2)$, eigenvalues $6, 2$, both positive: positive definite, so a strict local minimum, $f = -3$.
- At $(-1,1)$: $H = \mathrm{diag}(-6, 2)$, eigenvalues $-6, 2$: one negative, one positive: indefinite, so a saddle point, $f = 1$.
Numerical check with small steps of 0.1. Second-order prediction: $f(\hat{\mathbf{x}} + \mathbf{d}) \approx f(\hat{\mathbf{x}}) + \tfrac12 \mathbf{d}^\top H\mathbf{d}$.
- At $(1,1)$, $\mathbf{d} = (0.1, 0)$: predicted $-3 + \tfrac12 \cdot 6 \cdot 0.01 = -2.97$; actual $f(1.1, 1) = -2.969$ ✓. For $\mathbf{d} = (0, 0.1)$: predicted $-3 + 0.01 = -2.99$; actual $f(1, 1.1) = -2.99$ ✓. Both are above $-3$: a valley bottom.
- At $(-1,1)$, $\mathbf{d} = (0.1, 0)$: predicted $1 + \tfrac12(-6)(0.01) = 0.97$; actual $f(-0.9, 1) = 0.971$ ✓ (lower!). For $\mathbf{d} = (0, 0.1)$: predicted $1.01$; actual $f(-1, 1.1) = 1.01$ ✓ (higher). Down one way, up the other: a saddle.
The pure saddle $f(x,y) = x^2 - y^2$. $\nabla f = [2x, -2y]^\top = \mathbf{0}$ only at the origin. $H = \begin{bmatrix} 2 & 0 \\ 0 & -2 \end{bmatrix}$, eigenvalues $2$ and $-2$: indefinite, so a saddle. Check: $f(0.1, 0) = 0.01 \gt 0 = f(0,0)$ but $f(0, 0.1) = -0.01 \lt 0$.
2×2 shortcut. For $H = \begin{bmatrix} a & b \\ b & c \end{bmatrix}$ with $\det H = ac - b^2$: if $\det H \gt 0$ and $a \gt 0$, both eigenvalues are positive (min); if $\det H \gt 0$ and $a \lt 0$, both are negative (max); if $\det H \lt 0$ the eigenvalues have opposite signs (saddle); if $\det H = 0$, one eigenvalue is zero (inconclusive).
Let $\nabla f(\hat{\mathbf{x}}) = \mathbf{0}$ and let $H = \nabla^2 f(\hat{\mathbf{x}})$.
| Hessian $H$ at $\hat{\mathbf{x}}$ | Eigenvalues | Verdict |
|---|---|---|
| positive definite ($\mathbf{d}^\top H\mathbf{d} \gt 0$ for all $\mathbf{d} \ne \mathbf{0}$) | all $\gt 0$ | strict local minimum |
| negative definite | all $\lt 0$ | strict local maximum |
| indefinite | some $\gt 0$ and some $\lt 0$ | saddle point |
| positive or negative semidefinite, but singular | all $\ge 0$ (or all $\le 0$) with at least one $= 0$ | inconclusive |
Why it works. At a critical point the gradient term vanishes, so for small $\mathbf{d}$, by Taylor's theorem, $f(\hat{\mathbf{x}} + \mathbf{d}) = f(\hat{\mathbf{x}}) + \tfrac12 \mathbf{d}^\top H\mathbf{d} + (\text{smaller terms})$. If $H$ is positive definite, then $\mathbf{d}^\top H \mathbf{d} \ge \lambda_{\min}\|\mathbf{d}\|^2$ with $\lambda_{\min} \gt 0$, which beats the smaller terms: $f$ goes up in every direction. If $H$ has an eigenvector $\mathbf{v}$ with eigenvalue $\lambda \lt 0$, then $f(\hat{\mathbf{x}} + t\mathbf{v}) \approx f(\hat{\mathbf{x}}) + \tfrac12 \lambda t^2$ goes down.
Necessary versus sufficient. (Notation: $H \succ 0$ means "$H$ is positive definite", all eigenvalues $\gt 0$; $H \succeq 0$ means "positive semidefinite", all eigenvalues $\ge 0$; $\prec$ and $\preceq$ are the same for negative.) Sufficient: $\nabla f = \mathbf{0}$ and $H \succ 0$ $\Rightarrow$ strict local minimum. Necessary: a local minimum must have $\nabla f = \mathbf{0}$ and $H$ positive semidefinite (no negative eigenvalue). The gap between "$\succ 0$" and "$\succeq 0$" is exactly the inconclusive case.
Definiteness is explained in the Linear Algebra guide; practical tests (eigenvalues, Cholesky, leading minors) are in tests for definiteness; the Calculus guide has the second-derivative test.
Why do we need it?
The first-order condition produces candidates but cannot tell valley from hilltop from pass. The second-order condition makes the final decision for each candidate, without trying every nearby point.
Where is it used?
Checking that a solver's answer is truly a minimum, saddle detection in deep-learning research (eigenvalues of the loss Hessian), Newton's method (which only makes sense where the Hessian is positive definite), and proving that a quadratic loss has a unique minimum.
How is it used?
At each critical point compute $H$ and its eigenvalues (or $\det H$ and $H_{11}$ for $2\times2$). All positive: minimum. All negative: maximum. Mixed: saddle. A zero: look closer (next sections).
Check all the eigenvalues. One negative eigenvalue among a thousand positive ones still makes a saddle. A diagonal that is positive does not guarantee a bowl: $H = \begin{bmatrix} 1 & 3 \\ 3 & 1 \end{bmatrix}$ has positive diagonal but $\det = 1 - 9 \lt 0$ (eigenvalues $4$ and $-2$): a saddle.
The second-order test is local. "Strict local minimum" says nothing about other valleys far away. And a numerical Hessian may show tiny eigenvalues like $10^{-9}$ that are really zero: decide with a tolerance.
Quick check: at a critical point $H = \begin{bmatrix} 2 & 0 \\ 0 & -5 \end{bmatrix}$. What is it?
Eigenvalues $2$ and $-5$: mixed signs, so indefinite: a saddle point. (It curves up along $x$ and down along $y$.)
Local minima, local maxima and saddle points: sorting the candidates core
Think of a mountain range seen from above. There are valley bottoms (local minima), summits (local maxima) and passes (saddle points). All of them are flat spots. A walker who only goes downhill ends in a valley bottom, never on a summit, and (with a little luck or noise) slides off a pass.
Mathematically: find all critical points, then use the Hessian at each to sort them into these three bins. If you want the global minimum, compare the valley bottoms (and think about what happens far away).
$f(x, y) = x^3 + y^3 - 3xy$.
- Gradient: $\nabla f = [3x^2 - 3y,\; 3y^2 - 3x]^\top$. Set to zero: $y = x^2$ and $x = y^2$. Substituting, $x = x^4$, so $x(x^3 - 1) = 0$: $x = 0$ or $x = 1$. Critical points: $(0,0)$ and $(1,1)$.
- Hessian: $H = \begin{bmatrix} 6x & -3 \\ -3 & 6y \end{bmatrix}$.
- At $(0,0)$: $H = \begin{bmatrix} 0 & -3 \\ -3 & 0 \end{bmatrix}$, $\det = -9 \lt 0$ (eigenvalues $3$ and $-3$): a saddle, $f = 0$.
- At $(1,1)$: $H = \begin{bmatrix} 6 & -3 \\ -3 & 6 \end{bmatrix}$, $\det = 27 \gt 0$ and $6 \gt 0$ (eigenvalues $9$ and $3$): a strict local minimum, $f = -1$.
Is $(1,1)$ the global minimum? No! Along the diagonal $x = y = -t$ we get $f = -2t^3 - 3t^2 \to -\infty$ as $t \to \infty$. This $f$ has a local minimum but is unbounded below, so there is no global minimum at all. A local test can never prove that a point is global.
For a smooth $f$ and a critical point $\hat{\mathbf{x}}$:
- Local minimum: $f(\hat{\mathbf{x}}) \le f(\mathbf{x})$ for all nearby $\mathbf{x}$. Detected by $H \succ 0$ (strict), sometimes by $H \succeq 0$ plus extra checks.
- Local maximum: $f(\hat{\mathbf{x}}) \ge f(\mathbf{x})$ for all nearby $\mathbf{x}$. Detected by $H \prec 0$.
- Saddle point: $\nabla f = \mathbf{0}$, but $f$ goes up in some directions and down in others. Detected by $H$ indefinite.
From local to global. To get a global minimum you need more: (i) if $f \to \infty$ in every direction far away (e.g. $x^4 + y^4$), then a global minimum exists and it is the lowest of the local minima; (ii) if $f$ is convex (Chapter 3.6), every local minimum is global.
Why do we need it?
It organises the whole search: after solving $\nabla f = \mathbf{0}$ you have a pile of candidates, and this sorting tells you which ones are worth keeping (the minima) and which to discard (maxima and saddles).
Where is it used?
Analysing loss surfaces (how many minima and saddles does a small network have?), clustering (local optima of k-means), physics (equilibria are critical points of an energy: stable ones are minima), and economics (profit maxima).
How is it used?
List all critical points, compute the Hessian eigenvalues at each, label them, and compare the function values of the minima. If $f$ can run off to $-\infty$, no global minimum exists.
Local conclusions, global caution. A strict local minimum can be far from the best point (or the function may be unbounded below, as in the cubic example). A gradient-based optimizer started in the wrong valley will faithfully report a worse minimum.
Counting critical points is hard. A deep network can have astronomically many. We do not list them; we just try to reach a good one (Chapter 3.15).
Quick check: a function has critical points with $f$-values $-3$ (local min), $-1$ (local min) and $2$ (saddle). Is the first one the global minimum?
Not necessarily. Among the listed candidates it is the lowest, so it is the global minimum if a global minimum exists (for instance if $f \to \infty$ far away). But if $f$ can dip to $-\infty$ elsewhere (as with $x^3 + y^3 - 3xy$) there is no global minimum at all.
When the second-order test is inconclusive
The Hessian measures bending to second order only. If in some direction the ground is perfectly flat to second order (a zero eigenvalue), the Hessian has nothing to say about that direction. Whether the ground eventually curves up, curves down, or rises on one side and falls on the other is decided by the higher derivatives (third, fourth and beyond).
Think of a very, very flat-bottomed bowl, like $x^4$: its second derivative at the bottom is exactly zero, yet it is certainly a minimum.
One variable: all three functions have $f'(0) = 0$ and $f''(0) = 0$.
- $f(x) = x^4$: $f' = 4x^3$, $f'' = 12x^2 = 0$ at $0$. It is a minimum ($x^4 \ge 0$).
- $f(x) = -x^4$: same derivatives at 0, but a maximum.
- $f(x) = x^3$: $f'' = 6x = 0$ at $0$. Neither: it rises through $0$.
The first-derivative test still works: for $x^4$, $f' = 4x^3$ goes from negative to positive (min); for $-x^4$ from positive to negative (max); for $x^3$, $f' = 3x^2$ is positive on both sides (neither).
Two variables: four functions that all have $\nabla f(0,0) = \mathbf{0}$ and the same Hessian $H = \begin{bmatrix} 2 & 0 \\ 0 & 0 \end{bmatrix}$ (eigenvalues $2$ and $0$):
| $f(x,y)$ | What is really there |
|---|---|
| $x^2$ | a flat-bottomed trough: every point $(0,y)$ is a (non-strict) minimum |
| $x^2 + y^4$ | $f \ge 0$ with equality only at the origin: a strict local minimum |
| $x^2 - y^4$ | $f(0, y) = -y^4 \lt 0$ but $f(x, 0) = x^2 \gt 0$: a saddle |
| $x^2 + y^3$ | $f(0, y) = y^3$ takes both signs: a saddle |
The Hessian cannot distinguish them. What to do: look at the higher-order terms along the flat direction, or evaluate $f$ on a small circle around the point, or use another argument (like convexity).
At a critical point $\hat{\mathbf{x}}$, the second-order test is inconclusive when the Hessian is positive semidefinite but not positive definite (or negative semidefinite but not negative definite): all eigenvalues have the same sign or are zero, and at least one is zero.
- If $H$ has both a positive and a negative eigenvalue, the point is a saddle regardless of any zero eigenvalues.
- In one variable, $f'(\hat{x}) = 0$ and $f''(\hat{x}) = 0$ is inconclusive; use the sign pattern of $f'$, or the first non-zero higher derivative: if it is of even order, positive means minimum and negative means maximum; if odd, it is neither.
- Never call a point a minimum just because "$H \succeq 0$": that is only necessary.
Why do we need it?
It tells us when to stop trusting a quick test. A zero eigenvalue means "look closer", so we do not wrongly declare victory or defeat.
Where is it used?
Over-parameterised models, where whole valleys are flat in some directions (a zero eigenvalue is common), logistic regression on perfectly separable data, and degenerate saddles in physics. Redundant parameters (like two copies of the same feature) make the Hessian singular.
How is it used?
If a Hessian eigenvalue is (numerically) zero, probe the function around the point: sample a small circle, look at the sign of the first non-zero higher-order term, or use the first-derivative test in 1D.
"Inconclusive" does not mean "not a minimum". $x^4$ at 0 is a perfectly good minimum. It means this test is silent, and we need another tool.
Zero eigenvalues are common in practice. Redundant features, symmetries and over-parameterisation make a Hessian singular. Then a minimum is not a single point but a whole flat valley (infinitely many equally good answers). That is normal for deep nets.
Quick check: $f(x) = x^4 - 4x^3$ has $f'(x) = 4x^2(x - 3)$. At $x = 0$, $f'' = 0$. Min, max or neither?
Use the first-derivative test. For $x \lt 0$: $4x^2 \gt 0$ and $(x-3) \lt 0$, so $f' \lt 0$. For $0 \lt x \lt 3$: still $f' \lt 0$. The sign does not change, so $x = 0$ is neither (a flat spot on a downhill slope). At $x = 3$: $f'' = 12\cdot 9 - 24\cdot 3 = 36 \gt 0$, a minimum with $f(3) = -27$.
The minimiser of a quadratic: just solve $A\mathbf{x} = \mathbf{b}$ core
The simplest landscape in many dimensions is a perfect bowl, the many-variable version of the parabola $\tfrac12 a x^2 - bx$. Its slope is a straight-line (linear) function of $\mathbf{x}$. Setting a linear function equal to zero is a linear equation, and linear equations can be solved exactly, in one go, with no search at all.
That is why least squares (the most common ML objective in history) has an exact answer: its loss is a quadratic bowl.
$f(\mathbf{x}) = \tfrac12 \mathbf{x}^\top A \mathbf{x} - \mathbf{b}^\top \mathbf{x}$ with $A = \begin{bmatrix} 3 & 1 \\ 1 & 2 \end{bmatrix}$ and $\mathbf{b} = \begin{bmatrix} 1 \\ 1 \end{bmatrix}$. Written out: $f(x,y) = \tfrac32 x^2 + xy + y^2 - x - y$.
- Gradient: $\nabla f = [3x + y - 1,\; x + 2y - 1]^\top$. (This is exactly $A\mathbf{x} - \mathbf{b}$.)
- Set to zero: $3x + y = 1$ and $x + 2y = 1$. From the second, $x = 1 - 2y$. Put in the first: $3 - 6y + y = 1$, so $y = 0.4$ and $x = 0.2$.
- Hessian $= A$, with eigenvalues $\tfrac{5 \pm \sqrt5}{2} \approx 3.618$ and $1.382$, both positive: positive definite. So $\mathbf{x}^\star = [0.2, 0.4]^\top$ is a strict local minimum, and in fact the global one (see the definition).
- Best value: $f^\star = -\tfrac12 \mathbf{b}^\top\mathbf{x}^\star = -\tfrac12(0.2 + 0.4) = -0.3$ (the formula is derived below).
- Check the identity below at the point $(1, 0)$: $f(1,0) = 1.5 - 1 = 0.5$, so $f - f^\star = 0.8$. And $\mathbf{x} - \mathbf{x}^\star = (0.8, -0.4)$, $A(\mathbf{x} - \mathbf{x}^\star) = (2.4 - 0.4,\; 0.8 - 0.8) = (2.0, 0)$, so $\tfrac12 (\mathbf{x}-\mathbf{x}^\star)^\top A(\mathbf{x}-\mathbf{x}^\star) = \tfrac12 (0.8 \cdot 2.0 + 0) = 0.8$ ✓.
Least squares is this problem. For a model $X\mathbf{w}$ and targets $\mathbf{y}$: $\|X\mathbf{w} - \mathbf{y}\|^2 = \mathbf{w}^\top (X^\top X)\mathbf{w} - 2 (X^\top\mathbf{y})^\top \mathbf{w} + \mathbf{y}^\top\mathbf{y} = 2\left[\tfrac12 \mathbf{w}^\top A \mathbf{w} - \mathbf{b}^\top\mathbf{w}\right] + \text{const}$ with $A = X^\top X$ and $\mathbf{b} = X^\top \mathbf{y}$. Solving $A\mathbf{w} = \mathbf{b}$ gives the normal equations $X^\top X \mathbf{w} = X^\top \mathbf{y}$. With the three data points of Chapter 3.1 (columns "$x$" and "$1$"): $X^\top X = \begin{bmatrix} 14 & 6 \\ 6 & 3 \end{bmatrix}$, $X^\top \mathbf{y} = \begin{bmatrix} 23 \\ 10 \end{bmatrix}$. Solve: $14w + 6b = 23$, $6w + 3b = 10$. From the second, $b = (10 - 6w)/3$; into the first, $14w + 2(10 - 6w) = 23$, so $2w = 3$, $w = 1.5$, $b = \tfrac13$. This is exactly the line we found in Chapter 3.1, with loss $\tfrac16$ ✓. (Single feature, no bias: $X = [1, 2]^\top$, $\mathbf{y} = [2, 3]^\top$ gives $A = 5$, $b = 8$, $w^\star = 1.6$, as in the earlier example.)
Let $A$ be a symmetric $n \times n$ matrix and $f(\mathbf{x}) = \tfrac12 \mathbf{x}^\top A \mathbf{x} - \mathbf{b}^\top \mathbf{x}$. Then
$$\nabla f(\mathbf{x}) = A\mathbf{x} - \mathbf{b}, \qquad \nabla^2 f(\mathbf{x}) = A \quad(\text{the same everywhere}).$$(The gradient rule for $\tfrac12\mathbf{x}^\top A\mathbf{x}$ is derived in the Calculus guide.) Critical points solve the linear system $A\mathbf{x} = \mathbf{b}$.
- $A$ positive definite (all eigenvalues $\gt 0$; see here): there is exactly one critical point $\mathbf{x}^\star = A^{-1}\mathbf{b}$, and it is the unique global minimum, with $f^\star = -\tfrac12 \mathbf{b}^\top A^{-1}\mathbf{b}$.
- $A$ singular (some eigenvalue $0$) and positive semidefinite: either $A\mathbf{x} = \mathbf{b}$ has no solution (then $f$ falls to $-\infty$: no minimum), or it has infinitely many (a flat valley of equally good minimisers).
- $A$ with a negative eigenvalue: no minimum ($f$ falls to $-\infty$ along that eigenvector). A non-singular indefinite $A$ gives a saddle at $A^{-1}\mathbf{b}$.
Why the PD minimiser is global (completing the square). Let $A\mathbf{x}^\star = \mathbf{b}$. Expand $\tfrac12(\mathbf{x}-\mathbf{x}^\star)^\top A(\mathbf{x}-\mathbf{x}^\star) = \tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{x}^\top A\mathbf{x}^\star + \tfrac12 \mathbf{x}^{\star\top} A \mathbf{x}^\star = \tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x} + \tfrac12\mathbf{b}^\top\mathbf{x}^\star$. Therefore
$$f(\mathbf{x}) = f(\mathbf{x}^\star) + \tfrac12 (\mathbf{x} - \mathbf{x}^\star)^\top A\, (\mathbf{x} - \mathbf{x}^\star), \qquad f(\mathbf{x}^\star) = -\tfrac12 \mathbf{b}^\top \mathbf{x}^\star.$$If $A \succ 0$ the last term is $\gt 0$ for every $\mathbf{x} \ne \mathbf{x}^\star$, so $\mathbf{x}^\star$ beats every other point. ∎
Why do we need it?
It is the one case where optimization needs no iteration: the exact answer comes from one linear solve. It is also the model case that explains how every optimizer behaves near a minimum (any smooth function looks like a quadratic there).
Where is it used?
Linear regression (normal equations), ridge regression ($A = X^\top X + \lambda I$), Gaussian processes, Kalman filters, the inner step of Newton's method, and the analysis of gradient descent on a bowl (Chapter 3.3).
How is it used?
Form $A$ and $\mathbf{b}$, check that $A$ is positive definite (eigenvalues or a Cholesky factorisation), and call np.linalg.solve(A, b). Do not compute $A^{-1}$ explicitly: solving is faster and more accurate.
Solve, do not invert. Writing $\mathbf{x}^\star = A^{-1}\mathbf{b}$ is the maths; in code, call a solver. For a million variables, even storing $A$ is impossible, which is why gradient descent and conjugate gradient (Chapters 3.3 and 3.17) exist.
It needs the factor $\tfrac12$ and symmetric $A$. If $f(\mathbf{x}) = \mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x}$ (no $\tfrac12$), the gradient is $2A\mathbf{x} - \mathbf{b}$ and the solution is $\mathbf{x}^\star = \tfrac12 A^{-1}\mathbf{b}$. For non-symmetric $A$, replace $A$ by $(A + A^\top)/2$.
Quick check: $A = \begin{bmatrix} 4 & 0 \\ 0 & 1 \end{bmatrix}$, $\mathbf{b} = [8, 3]^\top$. Find the minimiser and the minimum value.
$A\mathbf{x} = \mathbf{b}$ gives $4x = 8$ and $y = 3$, so $\mathbf{x}^\star = [2, 3]^\top$. $A$ is diagonal with positive entries, so it is positive definite. $f^\star = -\tfrac12 \mathbf{b}^\top\mathbf{x}^\star = -\tfrac12(16 + 9) = -12.5$.
The whole recipe in one place
Everything in this chapter comes down to a five-step routine, the same for a one-variable cubic and a hundred-variable loss: find the flat spots, then ask each one how the ground bends.
When step 2 is too hard to solve by hand (as it is for almost all real models), we cannot do this by algebra. We then walk toward a flat spot using the gradient, which is the topic of Chapter 3.3.
Try it on $f(x,y) = x^4 + y^4 - 4xy$ (details in the widget below).
- Gradient: $\nabla f = [4x^3 - 4y,\; 4y^3 - 4x]^\top$.
- Solve $\nabla f = \mathbf{0}$: $y = x^3$ and $x = y^3$, so $x = x^9$, giving $x \in \{0, 1, -1\}$. Critical points: $(0,0)$, $(1,1)$, $(-1,-1)$.
- Hessian: $H = \begin{bmatrix} 12x^2 & -4 \\ -4 & 12y^2 \end{bmatrix}$.
- Classify: at $(0,0)$: $H = \begin{bmatrix} 0 & -4 \\ -4 & 0 \end{bmatrix}$, eigenvalues $\pm 4$: saddle. At $(\pm1, \pm1)$ (same signs): $H = \begin{bmatrix} 12 & -4 \\ -4 & 12 \end{bmatrix}$, eigenvalues $16$ and $8$: strict local minima.
- Compare values: $f(0,0) = 0$, $f(1,1) = f(-1,-1) = 1 + 1 - 4 = -2$. Since $4xy \le 2x^2 + 2y^2$, we have $f \ge (x^2-1)^2 + (y^2-1)^2 - 2 \ge -2$, so $-2$ is the global minimum, attained at both $(1,1)$ and $(-1,-1)$.
The unconstrained recipe.
- Compute the gradient $\nabla f$.
- Solve $\nabla f = \mathbf{0}$: these are the critical points (first-order condition).
- Compute the Hessian and evaluate it at each critical point.
- Read the eigenvalues: all $\gt 0$ min; all $\lt 0$ max; mixed saddle; some $0$ inconclusive (probe further).
- For a global answer, compare the values of the minima and ask whether $f$ can go lower far away (or use convexity, Chapter 3.6).
Why do we need it?
It is a checklist that always works for small smooth problems, and it is the template that numerical methods copy: find $\nabla f \approx 0$ using the gradient, then verify with curvature.
Where is it used?
Deriving closed-form estimators by hand (least squares, maximum likelihood for Gaussians), verifying a solver's output, and teaching what "converged" means in every optimizer.
How is it used?
Do the algebra for toy cases (exercises, sanity checks). For real models, let software compute $\nabla f$ and $H$ (autodiff) and apply the same logic numerically: small gradient norm, Hessian eigenvalues all positive.
Do not skip step 5. A local minimum is only a local minimum. (In the cubic example it is not even close to the best you can do.)
Do not skip the algebra check of step 2. Make sure you found all real solutions of $\nabla f = \mathbf{0}$, for example by checking every case when you divide by something that could be zero.
Quick check: $f(x,y) = x^2 + xy + y^2 - 3x$. Run the recipe.
$\nabla f = [2x + y - 3,\; x + 2y]^\top = \mathbf{0}$ gives $(2, -1)$. $H = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}$, eigenvalues $3$ and $1$ (positive definite). So $(2,-1)$ is a strict local minimum. Since $f$ is a quadratic with $A \succ 0$ it is the global minimum, $f(2,-1) = 4 - 2 + 1 - 6 = -3$.
Recap, cheat sheet and practice
- In one variable: solve $f'(x) = 0$, then look at $f''$: $\gt 0$ min, $\lt 0$ max, $= 0$ inconclusive. For a global answer on an interval also compare the endpoints.
- In many variables: the gradient $\nabla f$ (column vector) points uphill; the Hessian $\nabla^2 f$ (symmetric matrix) measures curvature in every direction: $\mathbf{u}^\top H\mathbf{u}$.
- A critical point has $\nabla f = \mathbf{0}$. This first-order condition is necessary for an unconstrained local optimum, not sufficient (maxima, saddles, plateaus also satisfy it).
- Second-order conditions at a critical point: $H \succ 0$ strict local min; $H \prec 0$ strict local max; $H$ indefinite saddle; $H$ semidefinite with a zero eigenvalue: inconclusive.
- $x^4$, $-x^4$ and $x^3$ share $f' = f'' = 0$ at 0 but behave differently; the first-derivative test or a higher-order look is needed.
- For $f = \tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x}$: $\nabla f = A\mathbf{x} - \mathbf{b}$, $\nabla^2 f = A$; if $A \succ 0$ the unique global minimiser solves $A\mathbf{x} = \mathbf{b}$, with $f^\star = -\tfrac12\mathbf{b}^\top\mathbf{x}^\star$. Least squares is exactly this.
- Local tests never prove global optimality. For that you need extra structure such as convexity (Chapter 3.6).
Cheat sheet
| Idea | Condition / formula | Meaning |
|---|---|---|
| Critical point | $\nabla f(\hat{\mathbf{x}}) = \mathbf{0}$ | level ground (necessary for an unconstrained local optimum) |
| Strict local minimum | $\nabla f = \mathbf{0}$ and $H \succ 0$ | bowl: all eigenvalues $\gt 0$ |
| Strict local maximum | $\nabla f = \mathbf{0}$ and $H \prec 0$ | hill: all eigenvalues $\lt 0$ |
| Saddle point | $\nabla f = \mathbf{0}$ and $H$ indefinite | up one way, down another |
| Inconclusive | $\nabla f = \mathbf{0}$, $H \succeq 0$ (or $\preceq 0$) singular | look at higher-order terms |
| 2×2 shortcut | $\det H \gt 0$, $H_{11} \gt 0$: min; $\det H \gt 0$, $H_{11} \lt 0$: max; $\det H \lt 0$: saddle | no eigenvalues needed |
| Slope / curvature along $\mathbf{u}$ | $\nabla f^\top\mathbf{u}$, $\mathbf{u}^\top H\mathbf{u}$ | one-variable slice $\phi(t) = f(\mathbf{x} + t\mathbf{u})$ |
| Second-order Taylor | $f(\mathbf{x} + \mathbf{d}) \approx f + \nabla f^\top\mathbf{d} + \tfrac12\mathbf{d}^\top H\mathbf{d}$ | why the Hessian decides |
| Quadratic minimiser | $A\mathbf{x}^\star = \mathbf{b}$, $f^\star = -\tfrac12\mathbf{b}^\top\mathbf{x}^\star$ | exact solution when $A \succ 0$ |
| Normal equations | $X^\top X\mathbf{w} = X^\top\mathbf{y}$ | least squares as a quadratic |
import numpy as np
import sympy as sp
# 1. critical points and classification, step by step, with SymPy
x, y = sp.symbols('x y', real=True)
f = x**3 - 3*x + y**2 - 2*y
grad = [sp.diff(f, v) for v in (x, y)]
H = sp.hessian(f, (x, y))
for c in sp.solve(grad, (x, y), dict=True):
ev = np.linalg.eigvalsh(np.array(H.subs(c), dtype=float))
kind = ('local min' if (ev > 0).all() else 'local max' if (ev < 0).all()
else 'saddle' if (ev > 0).any() and (ev < 0).any() else 'inconclusive')
print(c, float(f.subs(c)), ev, kind)
# {x: -1, y: 1} 1.0 [-6. 2.] saddle
# {x: 1, y: 1} -3.0 [2. 6.] local min
# 2. closed-form minimiser of 1/2 x^T A x - b^T x (solve A x = b, do not invert A)
A = np.array([[3.0, 1.0], [1.0, 2.0]])
b = np.array([1.0, 1.0])
xs = np.linalg.solve(A, b)
print(xs, -0.5 * b @ xs) # [0.2 0.4] -0.3
print(np.round(np.linalg.eigvalsh(A), 3)) # [1.382 3.618] both positive: A is positive definite
# 3. check the identity f(x) - f* = 1/2 (x - x*)^T A (x - x*) at a point
q = lambda v: 0.5 * v @ A @ v - b @ v
p = np.array([1.0, 0.0])
print(round(q(p) - q(xs), 6), round(0.5 * (p - xs) @ A @ (p - xs), 6)) # 0.8 0.8
# 4. least squares is a quadratic: A = X^T X, b = X^T y (the line fit of Chapter 3.1)
X = np.array([[1.0, 1.0], [2.0, 1.0], [3.0, 1.0]]) # columns: x, 1
yv = np.array([2.0, 3.0, 5.0])
w = np.linalg.solve(X.T @ X, X.T @ yv)
print(np.round(w, 4), round(np.sum((X @ w - yv) ** 2), 4)) # [1.5 0.3333] 0.1667
# 5. inconclusive Hessian: x^2 + y^4 and x^2 - y^4 at the origin both have H = diag(2, 0)
H0 = np.array([[2.0, 0.0], [0.0, 0.0]])
print(np.linalg.eigvalsh(H0)) # [0. 2.] a zero eigenvalue: the test says nothing
ring = lambda g, r=0.5: [g(r * np.cos(t), r * np.sin(t)) for t in np.linspace(0, 2 * np.pi, 73)]
print(min(ring(lambda a, c: a**2 + c**4)) > 0) # True -> every point of the ring is above 0: a minimum
print(min(ring(lambda a, c: a**2 - c**4)) < 0) # True -> the ring also dips below 0 (and above): a saddle
1. For some $f$, $f'(2) = 0$ and $f''(2) = -3$. What is $x = 2$?
2. Which statement about $\nabla f(\mathbf{x}^\star) = \mathbf{0}$ is true for an unconstrained problem?
3. At a critical point the Hessian has eigenvalues $4$ and $-1$. The point is...
4. At a critical point the Hessian has eigenvalues $3$ and $0$. What can you conclude?
5. For $f(\mathbf{x}) = \tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x}$ with $A$ symmetric positive definite, the global minimiser is...
6. For $f(x) = x^4$ at $x = 0$, the second-derivative test is inconclusive. In reality $x = 0$ is...
Practice problems
A. Find and classify all critical points of $f(x) = x^4 - 4x^3$, and its global minimum.
$f'(x) = 4x^3 - 12x^2 = 4x^2(x - 3)$, so $x = 0$ or $x = 3$. $f''(x) = 12x^2 - 24x$. At $x = 3$: $f'' = 108 - 72 = 36 \gt 0$: local minimum, $f(3) = 81 - 108 = -27$. At $x = 0$: $f'' = 0$, inconclusive. First-derivative test: $f' \lt 0$ on both sides of 0 (since $x^2 \ge 0$ and $x - 3 \lt 0$ near 0), no sign change: neither. As $x \to \pm\infty$, $f \to +\infty$, so the local minimum at 3 is the global minimum, $-27$.
B. Minimise $f(x,y) = x^2 + xy + y^2 - 3x$ using the quadratic formula $A\mathbf{x} = \mathbf{b}$.
Write $f = \tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x}$ with $A = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}$ and $\mathbf{b} = [3, 0]^\top$ (check: $\tfrac12(2x^2 + 2xy + 2y^2) = x^2 + xy + y^2$ ✓). Eigenvalues of $A$: $3, 1 \gt 0$. Solve $2x + y = 3$, $x + 2y = 0$: $x = -2y$, $-4y + y = 3$, $y = -1$, $x = 2$. $f^\star = -\tfrac12 \mathbf{b}^\top\mathbf{x}^\star = -\tfrac12(6) = -3$. This matches the recipe quick check.
C. What are the critical points of $f(x,y) = (x - y)^2$? Are they strict minima?
$\nabla f = [2(x-y),\, -2(x-y)]^\top = \mathbf{0}$ whenever $x = y$: a whole line of critical points. $H = \begin{bmatrix} 2 & -2 \\ -2 & 2 \end{bmatrix}$ has eigenvalues $4$ and $0$ (semidefinite). Since $f \ge 0$ everywhere and $f = 0$ on the line, every point on the line is a global minimum, but none is strict: the minimiser is not unique (a flat valley along the eigenvector $[1,1]^\top$ with eigenvalue 0).
D. Classify the critical point of $f(x,y) = x^2 - 3y^2 + 2x$.
$\nabla f = [2x + 2,\; -6y]^\top = \mathbf{0}$ gives $(-1, 0)$. $H = \mathrm{diag}(2, -6)$: eigenvalues $2$ and $-6$, mixed: a saddle point, with $f(-1, 0) = 1 - 2 = -1$.
E. Find the line $y = wx$ that best fits $(1, 2), (2, 4), (3, 5)$ in the least-squares sense, and the smallest total squared error.
Here $A = X^\top X = 1 + 4 + 9 = 14$ and $b = X^\top\mathbf{y} = 2 + 8 + 15 = 25$, so $w^\star = 25/14 \approx 1.786$. The loss is $L(w) = 14w^2 - 50w + 45$ (since $\mathbf{y}^\top\mathbf{y} = 4 + 16 + 25 = 45$). $L(w^\star) = 45 - 25^2/14 = 45 - 44.643 = 5/14 \approx 0.357$. Check: predictions $1.786, 3.571, 5.357$; errors $-0.214, -0.429, 0.357$; squares $0.046 + 0.184 + 0.128 \approx 0.357$ ✓.
F. At a critical point, $H = \begin{bmatrix} 1 & 3 \\ 3 & 1 \end{bmatrix}$. Both diagonal entries are positive. Is it a minimum?
No. $\det H = 1 - 9 = -8 \lt 0$, so the eigenvalues have opposite signs ($4$ and $-2$). It is a saddle. Positive diagonal entries do not make a matrix positive definite: check the eigenvalues (or the determinant for $2\times 2$).
Gradient Descent
This is the algorithm that trains almost every model in machine learning. The idea fits in one sentence: feel which way is downhill, take a step that way, and repeat. In this chapter we build it slowly, prove when it works, see exactly how it fails, and then meet its three practical versions: batch, stochastic and mini-batch.
- Explain gradient descent as "walking downhill in fog", and say what the objective, the gradient and the learning rate are
- Derive the update rule $x \leftarrow x - \eta\,\nabla f(x)$ from the first-order (linear) approximation
- Run gradient descent by hand on a small least-squares problem, and in code from scratch
- Prove what happens on a quadratic: each direction shrinks by $(1-\eta\lambda)$, so it converges exactly when $0 < \eta < 2/\lambda_{\max}$, and the best fixed step gives rate $\frac{\kappa-1}{\kappa+1}$
- Recognise overshooting, oscillation, divergence, slow convergence (ill-conditioning, plateaus) and know the remedies
- Choose the step size (fixed, decreasing, line search) and the starting point (initialization)
- Tell batch, stochastic and mini-batch gradient descent apart: they differ only in how the gradient is estimated
Gradient descent: walking downhill in fog core
You are standing on a hillside in thick fog. You want the bottom of the valley, but you cannot see it. All you can do is feel the slope under your feet: does the ground tilt up to your left, or down?
So you do the obvious thing. Feel the slope. Take one step in the direction that goes down. Stop, feel the slope again, take another step. Repeat until the ground feels flat. You never needed a map. You only needed to know the slope right where you stand.
That is gradient descent. The "slope under your feet" is the gradient. "How long a step" is the learning rate. This chapter is about those two things and what can go wrong with them.
Let the "hillside" be the height $f(x_1, x_2) = x_1^2 + 3x_2^2$. It is a bowl whose lowest point is $(0,0)$. Its slopes are $\partial f/\partial x_1 = 2x_1$ and $\partial f/\partial x_2 = 6x_2$, so the gradient is $\nabla f = [2x_1,\ 6x_2]$. Start at $(4, 2)$ and use a step size $\eta = 0.1$ ("eta", the learning rate).
- At $(4,2)$ the gradient is $[8, 12]$. The step is $-\eta\cdot[8,12] = [-0.8, -1.2]$. New point: $(3.2,\ 0.8)$. The height fell from $28$ to $3.2^2 + 3\cdot0.8^2 = 12.16$.
- At $(3.2, 0.8)$ the gradient is $[6.4, 4.8]$. Step $[-0.64, -0.48]$. New point $(2.56,\ 0.32)$, height $6.8608$.
- At $(2.56, 0.32)$ the gradient is $[5.12, 1.92]$. Step $[-0.512, -0.192]$. New point $(2.048,\ 0.128)$, height $4.2435$.
The height keeps falling: $28 \to 12.16 \to 6.86 \to 4.24$. Notice that $x_2$ shrank much faster than $x_1$ (the bowl is steeper in the $x_2$ direction, so its slope, and therefore the step, is bigger there). We will come back to that: it is the seed of "slow convergence".
Gradient descent finds a low point of a function $f$ (the objective) with this loop. Here $\mathbf{x}_k$ is the point after $k$ steps (the iterate), and one pass through the loop is one iteration:
- Choose a starting point $\mathbf{x}_0$ and a learning rate $\eta \gt 0$.
- Feel the slope: compute the gradient $\mathbf{g}_k = \nabla f(\mathbf{x}_k)$.
- Step downhill: $\ \mathbf{x}_{k+1} = \mathbf{x}_k - \eta\,\mathbf{g}_k$.
- If the stopping rule says "good enough" (for example, the gradient is almost zero, or we ran out of time), stop. Otherwise go back to step 2.
Everything in this chapter is a closer look at one of these four lines.
Why do we need it?
For almost every model there is no formula for the best parameters. But we can always compute the slope at the current parameters. Gradient descent turns "find the best" into "keep going downhill", which a computer can do for millions of parameters at once.
Where is it used?
Training linear and logistic regression, neural networks (as SGD, Adam and friends), matrix factorisation for recommenders, and almost every model that is fitted by minimising a loss.
How is it used?
Write the loss, get its gradient (by hand or by automatic differentiation), pick a starting point and a learning rate, loop "gradient, step" and watch the loss curve go down. If it does not, change the learning rate first.
Gradient descent only sees the ground under its feet. It never looks ahead. That is why it can end in a shallow valley (the double well) and why a bad step size makes it wander. The next sections fix each of these questions in turn: which direction, how far, how do we know we arrived, and where do we start.
You met the 1D version in the Calculus guide (gradient descent in one dimension). Here we go much deeper and in many dimensions at once.
Quick check: at $x = 3$ the slope of $f$ is $f'(3) = -2$ and $\eta = 0.5$. Where is the next point?
$x_{\text{new}} = 3 - 0.5\cdot(-2) = 3 + 1 = 4$. The slope was negative (the ground rises to the left), so the step goes right. The minus sign in the rule takes care of the direction automatically.
The objective function and the gradient: which way is steepest? core
The objective function is the thing we want to make small. It takes the numbers we can choose (the parameters) and returns one number that says how bad they are. In machine learning it is the loss. Think of it as the height of the ground above each point of a map: we are looking for the lowest ground. (We met this in Chapter 3.1.)
Draw the map from above with contour lines, the rings that connect points of equal height (like a hiking map). Close rings mean steep ground, wide rings mean gentle ground, and the very centre of the rings is the bottom.
The gradient is an arrow standing at your feet. It points straight uphill, along the steepest way up. Its length says how steep that is. It is always at a right angle to the contour ring you are standing on. To go downhill as fast as possible, walk the opposite way, along $-\nabla f$. At the very bottom the ground is flat, the gradient is the zero arrow, and we stop.
For $f(x_1,x_2) = x_1^2 + 3x_2^2$ at the point $(4,2)$ the gradient is $\nabla f = [8, 12]$, with length $\|\nabla f\| = \sqrt{64+144} = \sqrt{208} \approx 14.42$. How steep is the ground if you walk in a unit-length direction $\mathbf{u}$? The slope is the dot product $\nabla f\cdot\mathbf{u}$:
- Due east, $\mathbf{u} = [1,0]$: slope $8\cdot1 + 12\cdot0 = 8$.
- Due north, $\mathbf{u} = [0,1]$: slope $12$.
- Along the gradient, $\mathbf{u} = [8,12]/14.42 = [0.555, 0.832]$: slope $8\cdot0.555 + 12\cdot0.832 = 14.42$. The steepest possible.
- Opposite, $\mathbf{u} = -[0.555,0.832]$: slope $-14.42$. Steepest downhill.
- Sideways, $\mathbf{u} = [-12, 8]/14.42$: slope $\frac{8(-12)+12\cdot 8}{14.42} = 0$. Walking along a contour ring does not change the height.
An objective function is a function $f:\mathbb{R}^n\to\mathbb{R}$. We want a point $\mathbf{x}^\star$ where $f$ is as small as possible.
Its gradient is the column vector of partial derivatives (see the gradient in the Calculus guide):
$$\nabla f(\mathbf{x}) = \begin{bmatrix} \partial f/\partial x_1 \\ \vdots \\ \partial f/\partial x_n \end{bmatrix}.$$The slope of $f$ at $\mathbf{x}$ when you walk in the unit direction $\mathbf{u}$ is the directional derivative $\nabla f(\mathbf{x})\cdot\mathbf{u} = \|\nabla f\|\cos\theta$, where $\theta$ is the angle between $\mathbf{u}$ and $\nabla f$. Since $-1\le\cos\theta\le 1$:
- $\theta = 0$ (walk along $\nabla f$): the slope is $+\|\nabla f\|$, the largest possible. So the gradient is the direction of steepest ascent.
- $\theta = 180^\circ$ (walk along $-\nabla f$): slope $-\|\nabla f\|$, the most negative. Steepest descent.
- $\theta = 90^\circ$: slope $0$. The gradient is perpendicular to the contour line.
A point where $\nabla f = \mathbf{0}$ is a stationary point (flat ground: a minimum, a maximum or a saddle; Chapter 3.2).
Why do we need it?
To walk downhill we need to know which direction is downhill. The gradient answers that using only local information: one arrow at the point where we stand, no map of the whole landscape.
Where is it used?
The gradient of the loss with respect to the weights is what backpropagation computes in every neural network, and what loss.backward() fills in. It also drives logistic regression, linear regression and every optimizer in the next chapters.
How is it used?
Compute $\nabla f$ at the current point (by hand, or by autodiff), flip its sign and scale it: that is the next step. A gradient of length near $0$ tells you that you are near flat ground.
Sign errors are the classic bug. $\nabla f$ points up. Gradient descent uses $-\nabla f$. If your loss goes up instead of down, check the minus sign first.
"Steepest" is a local statement. $-\nabla f$ is the steepest way down right here, for an infinitely small step. In the stretched bowl above it does not aim at the bottom. That is exactly why gradient descent zig-zags in narrow valleys (see "Slow convergence" below).
Quick check: $f(x_1,x_2) = x_1^2 + x_1 x_2$ at $(1, 2)$. Find $\nabla f$ and the steepest-descent direction.
$\partial f/\partial x_1 = 2x_1 + x_2 = 4$, $\partial f/\partial x_2 = x_1 = 1$. So $\nabla f = [4, 1]$ and the steepest-descent direction is $-\nabla f = [-4,-1]$ (length $\sqrt{17}\approx 4.12$; as a unit vector $[-0.970,-0.243]$).
The update rule, derived from the first-order idea core
Why $\mathbf{x} - \eta\nabla f$? Because of one fact: if you zoom in far enough, any smooth landscape looks like a flat, tilted board. (This is "linearization", from the Calculus guide: the first-order Taylor idea.)
On a tilted board, the fastest way down is straight against the tilt. That fixes the direction. But a tilted board keeps going down forever, so it cannot tell us how far to walk. The board is only a good picture very close to us. So we walk a small step, then zoom in again at the new place and redo the picture. The length of the step is the learning rate (times the tilt).
Use $f(x_1,x_2) = x_1^2 + 3x_2^2$ at $\mathbf{x} = (4,2)$. Here $f = 28$ and $\mathbf{g} = \nabla f = [8,12]$, so $\|\mathbf{g}\|^2 = 208$. The tilted-board (first-order) prediction for the step $-\eta\mathbf{g}$ is: new height $\approx f - \eta\|\mathbf{g}\|^2 = 28 - 208\eta$.
| $\eta$ | predicted height $28-208\eta$ | actual height $f(\mathbf{x}-\eta\mathbf{g})$ | gap |
|---|---|---|---|
| $0.01$ | $25.92$ | $25.9696$ | $0.0496$ |
| $0.1$ | $7.2$ | $12.16$ | $4.96$ |
| $0.3$ | $-34.4$ (impossible: $f\ge0$) | $28 - 62.4 + 44.64 = 10.24$ | $44.64$ |
For a small step the board is almost right (gap $0.0496$). For a bigger step the gap grows like $\eta^2$ ($10\times$ the step, $100\times$ the gap). The curve bends away from the board. So: small step, trustworthy; big step, risky.
Derivation, in three steps.
- The tilted board. For a small step $\mathbf{d}$: $\ f(\mathbf{x}+\mathbf{d}) \approx f(\mathbf{x}) + \nabla f(\mathbf{x})^\top\mathbf{d}$.
- The best direction. Fix the step length $\|\mathbf{d}\| = r$. We want $\nabla f^\top\mathbf{d}$ as negative as possible. By the Cauchy–Schwarz inequality, $\nabla f^\top\mathbf{d} \ge -\|\nabla f\|\,\|\mathbf{d}\|$, with equality exactly when $\mathbf{d}$ points opposite to $\nabla f$. So $\mathbf{d} = -r\,\nabla f/\|\nabla f\|$.
- The length. Make the length proportional to the slope: $r = \eta\|\nabla f\|$. Then $\mathbf{d} = -\eta\,\nabla f(\mathbf{x})$.
A second way to see it (the "price for distance" view). Build a model that trusts the board but charges a price for walking far: $\ m(\mathbf{d}) = f + \mathbf{g}^\top\mathbf{d} + \frac{1}{2\eta}\|\mathbf{d}\|^2$. Its lowest point is where the slope of $m$ is zero: $\mathbf{g} + \mathbf{d}/\eta = \mathbf{0}$, so $\mathbf{d} = -\eta\mathbf{g}$, the same rule. Now $1/\eta$ reads as "how curved we assume the ground is": a big $\eta$ means we assume gentle curvature and walk far.
When is a step guaranteed to help? Suppose the ground never curves more than $L$ (all eigenvalues of the Hessian are at most $L$). Then $f(\mathbf{x}+\mathbf{d}) \le f + \mathbf{g}^\top\mathbf{d} + \frac L2\|\mathbf{d}\|^2$ (we prove this in Chapter 3.5). Put $\mathbf{d} = -\eta\mathbf{g}$:
$$f(\mathbf{x}_{k+1}) \le f(\mathbf{x}_k) - \eta\Big(1 - \frac{L\eta}{2}\Big)\|\mathbf{g}_k\|^2 .$$The bracket is positive exactly when $\eta \lt 2/L$, so every step goes down. With $\eta = 1/L$ the guaranteed drop is $\|\mathbf{g}\|^2/(2L)$. In our example $L = 6$ (the largest second derivative of $x_1^2+3x_2^2$), and for $\eta = 0.1$ the guaranteed drop is $0.1(1-0.3)\cdot208 = 14.56$; the actual drop was $28 - 12.16 = 15.84 \ge 14.56$. ✓
Why do we need it?
"Step downhill" is a feeling; the update rule makes it exact and tells us why it works and when it fails. The derivation shows that the rule is the best move under a flat-ground model, and that a step-size limit ($2/L$) appears naturally.
Where is it used?
The same idea (a local model, then minimise it) gives every optimizer in this guide: Newton's method uses a curved model, momentum and Adam change the direction, and proximal methods add a penalty term to the model.
How is it used?
You only ever type the one line x = x - eta * grad(x). The derivation tells you how to pick $\eta$: below $2/L$ to be safe, around $1/L$ for guaranteed progress, and smaller when you do not know $L$.
The update rule is a rule about one step. It does not promise that the next point is the lowest point. Each step only needs to go down a bit, and we repeat. Also, the guaranteed-drop formula needs a known curvature bound $L$. In practice we often do not know $L$, so we tune $\eta$ by watching the loss.
Quick check: for $f(x)=x^2$ ($L = 2$) from $x=3$, what is the guaranteed drop with $\eta = 1/L = 0.5$, and what is the real drop?
$f'(3) = 6$, so $\|g\|^2 = 36$. Guaranteed drop $= \eta(1-L\eta/2)\|g\|^2 = 0.5\cdot0.5\cdot36 = 9$. Real step: $x = 3 - 0.5\cdot6 = 0$, height $0$: real drop $9$. For a quadratic the bound is exact (the curvature really equals $L$).
A full hand-worked run: gradient descent on least squares core
Fitting a straight line to data means choosing two numbers: the intercept $b$ (where the line starts) and the slope $m$ (how steep it is). For every choice $(b,m)$ we can measure how badly the line misses the data points. That "badness" is a height over the $(b,m)$ plane, and it makes a bowl. The best line sits at the bottom of the bowl.
So fitting the line is an optimization problem, and gradient descent can solve it: start with any line, see which small change of $(b,m)$ lowers the badness fastest, change it, repeat.
Data: three points $(x,y)$: $(0,1),\ (1,3),\ (2,4)$. Model: $\hat y = b + m x$. The residual of point $i$ is $r_i = b + m x_i - y_i$ (prediction minus truth). Objective: $f(b,m) = \tfrac12\sum_i r_i^2$. (The $\tfrac12$ is only there so the derivatives have no stray factor $2$.)
Gradient. By the chain rule, $\partial(\tfrac12 r_i^2)/\partial b = r_i\cdot 1$ and $\partial(\tfrac12 r_i^2)/\partial m = r_i\cdot x_i$. Adding over the points: $\ \partial f/\partial b = \sum r_i,\quad \partial f/\partial m = \sum r_i x_i$.
Run with start $(b,m) = (0,0)$ and $\eta = 0.1$:
- Iteration 1. $\hat y = [0,0,0]$, so $\mathbf{r} = [-1,-3,-4]$ and $f = \tfrac12(1+9+16) = 13$. $\partial f/\partial b = -1-3-4 = -8$. $\partial f/\partial m = 0\cdot(-1) + 1\cdot(-3) + 2\cdot(-4) = -11$. Update: $(b,m) = (0,0) - 0.1\,(-8,-11) = (0.8,\ 1.1)$.
- Iteration 2. $\hat y = [0.8, 1.9, 3.0]$, $\mathbf{r} = [-0.2,-1.1,-1.0]$, $f = \tfrac12(0.04+1.21+1) = 1.125$. Gradient $= (-2.3,\ -1.1-2.0) = (-2.3, -3.1)$. Update: $(0.8+0.23,\ 1.1+0.31) = (1.03,\ 1.41)$.
- Iteration 3. $\hat y = [1.03, 2.44, 3.85]$, $\mathbf{r} = [0.03,-0.56,-0.15]$, $f = 0.1685$. Gradient $= (-0.68,\ -0.56-0.30) = (-0.68,-0.86)$. Update: $(1.098,\ 1.496)$.
- Iteration 4. $\hat y = [1.098, 2.594, 4.090]$, $\mathbf{r} = [0.098,-0.406,0.090]$, $f = 0.0913$. Gradient $=(-0.218,-0.226)$. Update: $(1.1198,\ 1.5186)$.
Where is it heading? The exact answer solves $\nabla f = \mathbf{0}$, the normal equations $X^\top X\mathbf{w} = X^\top\mathbf{y}$. Here $X^\top X = \begin{bmatrix}3&3\\3&5\end{bmatrix}$ and $X^\top\mathbf{y} = [8, 11]$, which gives $\mathbf{w}^\star = (7/6,\ 3/2) = (1.1667,\ 1.5)$ with $f^\star = 1/12 = 0.0833$. Check: $3\cdot\frac76+3\cdot\frac32 = 8$ ✓ and $3\cdot\frac76+5\cdot\frac32 = 11$ ✓. Our iterates $(0.8,1.1)\to(1.03,1.41)\to(1.098,1.496)\to(1.12,1.519)$ are walking there.
The loss went $13 \to 1.125 \to 0.168 \to 0.091 \to 0.085$: a huge drop at first, then a crawl. The first step removed most of the error and later steps only polish. Gradient descent stops when $\|\nabla f\| < 10^{-3}$ after 47 iterations. Why the crawl? That is explained in the section on quadratics below (the bowl is a long, thin valley).
Stack the inputs into the design matrix $X$ (row $i$ is $[1,\ x_i]$) and the targets into $\mathbf{y}$. Then, with $\mathbf{w} = [b, m]$:
$$f(\mathbf{w}) = \tfrac12\|X\mathbf{w}-\mathbf{y}\|^2,\qquad \nabla f(\mathbf{w}) = X^\top(X\mathbf{w}-\mathbf{y}),\qquad \nabla^2 f = X^\top X.$$Gradient descent on least squares is therefore $\ \mathbf{w}\leftarrow\mathbf{w} - \eta\,X^\top(X\mathbf{w}-\mathbf{y})$. The Hessian $X^\top X$ is constant, so $f$ is exactly a quadratic bowl (positive definite when the columns of $X$ are independent; see quadratic forms and least squares). If your loss is the mean squared error $\frac1n\|X\mathbf{w}-\mathbf{y}\|^2$ (as in the Calculus guide), the gradient is $\frac2nX^\top(X\mathbf{w}-\mathbf{y})$: the same direction, just a different scale, which only rescales $\eta$.
Why do we need it?
A full run by hand, with every number, is the best way to see that gradient descent is just arithmetic. Linear regression is also the one model where we can check the answer exactly, with the normal equations.
Where is it used?
The same loop trains linear regression on data too big for the normal equations, logistic regression, and (with a different gradient) neural networks. Least squares is also the standard test problem for every optimizer.
How is it used?
Compute the residuals $X\mathbf{w}-\mathbf{y}$, multiply by $X^\top$ to get the gradient, subtract $\eta$ times it. Compare with np.linalg.solve to confirm the result: the two must agree.
Why $\eta=0.1$ and not $1$? The largest curvature of this bowl is the biggest eigenvalue of $X^\top X$, which is $4+\sqrt{10}\approx7.162$. The safe range is $\eta < 2/7.162 = 0.279$. A step of $1$ would explode. The next sections explain where this threshold comes from.
Centring the inputs helps. If we use $x-1$ instead of $x$ (so the inputs are centred at zero), $X^\top X$ becomes $\mathrm{diag}(3,2)$. The valley becomes nearly round (condition number $1.5$ instead of $8.55$), and gradient descent converges much faster. This is why we standardise features before training.
Quick check: one step of gradient descent on the same data starting from $(b,m) = (1,1)$ with $\eta = 0.1$.
$\hat y = 1 + x = [1,2,3]$, so $\mathbf{r} = [0,-1,-1]$. Gradient $= (\sum r,\ \sum r x) = (-2,\ -1-2) = (-2,-3)$. Update: $(1,1) - 0.1(-2,-3) = (1.2,\ 1.3)$.
Putting it together: the three dashboards of a run core
A pilot reads three dials. We read three too, whenever we run gradient descent:
- The map: where is the iterate $\mathbf{x}_k$ right now, and what path did it take?
- The loss curve: the height $f(\mathbf{x}_k)$ against the step number $k$. It should go down.
- The gradient length: $\|\nabla f(\mathbf{x}_k)\|$ against $k$. It tells us how flat the ground is. It also sets the step size: the step has length $\eta\|\nabla f\|$, so steps shrink by themselves as the ground flattens.
Almost all debugging of training is reading these three dials and asking: what shape is the curve, and what does it mean?
Here is how to read the shapes (you will see every one of them in the playground below):
| What you see | What it usually means | First thing to try |
|---|---|---|
| Loss falls smoothly and flattens at the bottom | All is well | Nothing |
| Loss drops fast, then crawls for a long time | $\eta$ small for the flat direction, or the valley is long and thin (ill-conditioned) | More steps, a larger $\eta$ (while stable), rescale the inputs |
| Path zig-zags across a valley; loss still falls | $\eta$ is near its limit in the steep direction | Lower $\eta$ a little, or use momentum (next chapter) |
Loss goes up, or becomes NaN | $\eta$ too large (divergence) | Cut $\eta$ by a factor 3–10 |
| Loss stuck high, gradient length near $0$ | Flat region: local minimum, saddle or plateau | Different start, bigger $\eta$, a method with momentum |
For a run $\mathbf{x}_0,\mathbf{x}_1,\dots$ of gradient descent we track
$$f(\mathbf{x}_k),\qquad \|\nabla f(\mathbf{x}_k)\|,\qquad \|\mathbf{x}_{k+1}-\mathbf{x}_k\| = \eta\,\|\nabla f(\mathbf{x}_k)\|.$$The last identity follows directly from the update rule: the step is $-\eta\nabla f$, so its length is $\eta$ times the gradient length. The plots below use a log scale on the vertical axis, because a loss that shrinks by the same factor each step is a straight line on a log plot (we explain why in the quadratic section). "Loss gap" means $f(\mathbf{x}_k) - f^\star$, the height above the lowest point.
Why do we need it?
You cannot see a million-dimensional loss surface, but you can always plot a number against the step count. Reading these curves is how you diagnose a training run without seeing the landscape.
Where is it used?
Every training log: TensorBoard and Weights & Biases loss curves, the loss and grad_norm lines printed by training loops, and learning-rate "range tests" before a long run.
How is it used?
Log the loss (and ideally the gradient norm) every few steps. Look at the shape: smooth fall, crawl, wiggle or blow-up. Then change one thing, usually $\eta$, and compare.
A small gradient does not mean you found the best point. The gradient is small on any flat ground: a minimum, but also a plateau or a saddle. Always ask which flat spot you are on (try the bumpy bowl).
A falling loss does not prove the learning rate is good. A loss that falls slowly is falling, but you may be wasting time. Compare two or three learning rates on a log scale (for example $0.1$, $0.03$, $0.01$).
Quick check: at some iterate $\|\nabla f\| = 5$ and $\eta = 0.02$. How long is the next step?
Step $= -\eta\nabla f$, so its length is $0.02\times5 = 0.1$. As the iterate nears the bottom, $\|\nabla f\|\to0$ and the steps shrink on their own, even with a fixed $\eta$.
Convergence: what it means, and when to stop core
Gradient descent produces a list of points $\mathbf{x}_0,\mathbf{x}_1,\mathbf{x}_2,\dots$ Each point is a little nearer the bottom than the last. We say the method converges when the points settle down: they get closer and closer to one spot and stay there. It is like turning a radio dial until the station comes in clearly: the changes get smaller and smaller.
We never reach the bottom exactly (a computer would need infinitely many steps, and rounds numbers anyway). So we choose a moment to stop and say "close enough". But close in what sense? There are three natural meanings. We can ask for the position to be close to the answer, for the height to be close to the lowest height, or for the slope to be close to zero. Only the third can be checked without already knowing the answer.
Take the least-squares run from earlier ($\eta=0.1$, answer $\mathbf{w}^\star = (7/6, 3/2)$, $f^\star = 1/12$). Here are the three measurements at each step:
| $k$ | distance $\|\mathbf{w}_k-\mathbf{w}^\star\|$ | loss gap $f-f^\star$ | gradient length $\|\nabla f\|$ | step length $\eta\|\nabla f\|$ |
|---|---|---|---|---|
| 0 | 1.9003 | 12.9167 | 13.6015 | 1.3601 |
| 1 | 0.5426 | 1.0417 | 3.8601 | 0.3860 |
| 2 | 0.1636 | 0.0852 | 1.0964 | 0.1096 |
| 3 | 0.0688 | 0.0079 | 0.3140 | 0.0314 |
| 4 | 0.0504 | 0.0015 | 0.0972 | 0.0097 |
All four columns shrink, but we could only compute the last two without knowing $\mathbf{w}^\star$. In practice we watch the gradient length. For a quadratic like this one there is a safety net: $\nabla f = A(\mathbf{w}-\mathbf{w}^\star)$, so $\|\mathbf{w}-\mathbf{w}^\star\|\le\|\nabla f\|/\lambda_{\min}$. With $\lambda_{\min} = 0.8377$, at $k=3$ the bound says the distance is at most $0.314/0.8377 = 0.375$; the truth is $0.0688$. The bound is safe but loose.
A sequence $\mathbf{x}_k$ converges to $\mathbf{x}^\star$ if $\|\mathbf{x}_k-\mathbf{x}^\star\|\to0$ as $k\to\infty$. Several related statements are used:
- Iterates converge: $\|\mathbf{x}_k-\mathbf{x}^\star\|\to0$.
- Values converge: $f(\mathbf{x}_k)\to f^\star$.
- Stationarity: $\|\nabla f(\mathbf{x}_k)\|\to0$ (the only one we can always measure).
Stopping rules (stop at the first $k$ where a rule fires; use a small tolerance $\varepsilon$ you choose):
- Gradient small: $\|\nabla f(\mathbf{x}_k)\|\le\varepsilon$ (the usual one).
- Steps small: $\|\mathbf{x}_{k+1}-\mathbf{x}_k\|\le\varepsilon$.
- Loss stalls: $|f(\mathbf{x}_{k})-f(\mathbf{x}_{k+1})|\le\varepsilon$.
- Budget: $k$ reaches a maximum number of iterations (always keep this one as a safety net).
What does it converge to? For a smooth convex function with a suitable $\eta$: to a global minimum. For a non-convex function: to a stationary point ($\nabla f=\mathbf{0}$), which can be a local minimum, and in unlucky cases a saddle. How fast it converges is the topic of Chapter 3.5; in the next section we compute it exactly for quadratics.
Why do we need it?
A loop needs an exit. "Run until it looks right" is not an algorithm. A stopping rule says exactly when the answer is good enough, so we neither stop too early nor burn compute on tiny improvements.
Where is it used?
scipy.optimize.minimize options gtol, ftol and maxiter; scikit-learn's tol and max_iter; early stopping in deep learning (stop when the validation loss stops improving).
How is it used?
Pick a tolerance (for example $10^{-6}$) and a maximum number of iterations. Stop at the first rule that fires, and print which one it was. If it was the maximum, you probably did not converge.
The "steps are small" rule is easy to fool. The step length is $\eta\|\nabla f\|$. If $\eta$ is tiny, steps are tiny from the very start, and the rule says "converged" long before you are anywhere near the bottom. The gradient rule does not have this problem (it does not involve $\eta$). Scale also matters: a fixed $\varepsilon$ means different things for a loss around $10^{6}$ and a loss around $10^{-6}$, so some libraries use relative tolerances.
Small gradient is not "global minimum". It only says the ground is flat. On a non-convex function that can be a local minimum or a saddle.
Quick check: a run stops because it hit the maximum iteration count. Did it converge?
Not necessarily. The budget rule only says "we ran out of time". Check the gradient length (or the loss curve): if it is still falling, the answer is not final. Always report which rule stopped the run.
Gradient descent on a quadratic: every direction shrinks by $(1-\eta\lambda)$ core
Why study a quadratic bowl? Because every smooth landscape looks like one near its bottom. Zoom in on any minimum and the ground curves like a bowl (the second-order Taylor idea). So how gradient descent behaves on a quadratic tells us how it behaves at the end of almost every training run.
A quadratic bowl is usually stretched: steep in one direction, gentle in another. The key insight is that gradient descent treats the bowl's axes independently. Along each axis it is plain 1D gradient descent on a parabola, and on a parabola with curvature $\lambda$ each step multiplies the distance to the bottom by the same number, $1-\eta\lambda$. A steep axis (big $\lambda$) shrinks fast, or overshoots if $\eta$ is too big. A gentle axis (small $\lambda$) shrinks slowly. The slowest axis decides how long the whole run takes.
Take $f(x_1,x_2) = \tfrac12(x_1^2 + 10x_2^2)$. The gradient is $[x_1,\ 10x_2]$, so one step is
$$x_1\leftarrow x_1-\eta x_1 = (1-\eta)\,x_1,\qquad x_2\leftarrow x_2-10\eta\,x_2 = (1-10\eta)\,x_2.$$The two coordinates are completely separate. Each is multiplied by its own factor every step:
| $\eta$ | factor for $x_1$: $1-\eta$ | factor for $x_2$: $1-10\eta$ | overall rate (the larger size) | what happens |
|---|---|---|---|---|
| 0.05 | 0.95 | 0.5 | 0.95 | converges, slowly (the flat axis $x_1$ crawls) |
| 0.1 | 0.9 | 0 | 0.9 | $x_2$ is killed in one step; $x_1$ still slow |
| $2/11 = 0.1818$ | 0.8182 | $-0.8182$ | 0.8182 | best: both axes equally fast ($x_2$ flips sign each step) |
| 0.19 | 0.81 | $-0.9$ | 0.9 | converges, but worse: $x_2$ overshoots more |
| 0.2 | 0.8 | $-1$ | 1 | $x_2$ jumps between $+c$ and $-c$ forever: never converges |
| 0.25 | 0.75 | $-1.5$ | 1.5 | diverges |
So the best step balances the two axes. Smaller makes the flat axis slow; larger makes the steep axis overshoot.
Setting. Let $f(\mathbf{x}) = \tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x}$ with $A$ symmetric positive definite (positive definite: all eigenvalues $\lambda_i \gt 0$). Its gradient is $A\mathbf{x}-\mathbf{b}$ and its minimiser is $\mathbf{x}^\star = A^{-1}\mathbf{b}$, so $\nabla f(\mathbf{x}) = A(\mathbf{x}-\mathbf{x}^\star)$.
Derivation.
- The error. Let $\mathbf{e}_k = \mathbf{x}_k-\mathbf{x}^\star$. The update gives $\mathbf{e}_{k+1} = \mathbf{e}_k - \eta A\mathbf{e}_k = (I-\eta A)\,\mathbf{e}_k$. So $\mathbf{e}_k = (I-\eta A)^k\mathbf{e}_0$.
- Separate the axes. $A$ is symmetric, so $A = Q\Lambda Q^\top$ with $Q$ orthogonal and $\Lambda = \mathrm{diag}(\lambda_1,\dots,\lambda_n)$ (symmetric eigen-decomposition; the columns of $Q$ are the bowl's axes). Write the error in the axis coordinates: $\mathbf{c}_k = Q^\top\mathbf{e}_k$.
- Each axis on its own. Multiply step 1 by $Q^\top$: $\mathbf{c}_{k+1} = (I-\eta\Lambda)\mathbf{c}_k$. Since $I-\eta\Lambda$ is diagonal, coordinate $i$ evolves alone: $$c_i(k) = (1-\eta\lambda_i)^k\,c_i(0).$$
- When does everything shrink? We need $|1-\eta\lambda_i|\lt 1$ for every $i$, that is $0\lt \eta\lambda_i\lt 2$. The binding one is the largest eigenvalue: $$\boxed{\ 0\lt \eta\lt \frac{2}{\lambda_{\max}}\ }$$
- The rate. The slowest axis decides: $\|\mathbf{e}_k\|\le\rho^k\|\mathbf{e}_0\|$ with $\rho(\eta) = \max_i|1-\eta\lambda_i| = \max\big(|1-\eta\lambda_{\min}|,\ |1-\eta\lambda_{\max}|\big)$. (Because $Q$ is orthogonal it does not change lengths.)
- The best fixed step. $1-\eta\lambda_{\min}$ decreases as $\eta$ grows and $|1-\eta\lambda_{\max}|$ is first shrinking, then growing. The best $\eta$ makes the two equal in size, $1-\eta\lambda_{\min} = -(1-\eta\lambda_{\max})$: $$\boxed{\ \eta^\star = \frac{2}{\lambda_{\min}+\lambda_{\max}},\qquad \rho^\star = \frac{\lambda_{\max}-\lambda_{\min}}{\lambda_{\max}+\lambda_{\min}} = \frac{\kappa-1}{\kappa+1}\ }$$ where $\kappa = \lambda_{\max}/\lambda_{\min}$ is the condition number. (Check: $1-\eta^\star\lambda_{\min} = \frac{\lambda_{\max}+\lambda_{\min}-2\lambda_{\min}}{\lambda_{\max}+\lambda_{\min}} = \rho^\star$ ✓.)
Numbers. $\kappa = 10$: $\eta^\star = 2/11$, $\rho^\star = 9/11 = 0.818$ (we measured $0.8182$ by running the code). To shrink the error by $10^{-6}$ you need $k\ge \ln(10^{6})/\ln(1/\rho)$ steps: $68.8$ steps for $\kappa=10$, $690.8$ for $\kappa=100$. In general this is about $\frac\kappa2\ln\frac1\varepsilon$: the number of steps grows in proportion to $\kappa$.
The loss shrinks twice as fast on a log scale. $f-f^\star = \tfrac12\sum\lambda_ic_i^2$, so each term shrinks like $(1-\eta\lambda_i)^{2k}$: the loss gap decays like $\rho^{2k}$. On a log plot each $\log|c_i|$ is a straight line with slope $\log|1-\eta\lambda_i|$, which is why we plot loss curves on a log axis.
Beyond quadratics. Near a minimum $\mathbf{x}^\star$ of any smooth $f$, $f\approx f^\star+\tfrac12(\mathbf{x}-\mathbf{x}^\star)^\top H(\mathbf{x}-\mathbf{x}^\star)$ with $H=\nabla^2f(\mathbf{x}^\star)$. So close enough to the minimum, GD behaves like the quadratic with $A=H$ (we prove this properly in Chapter 3.5, where $\lambda_{\max}$ becomes the smoothness constant $L$ and $\lambda_{\min}$ the strong-convexity constant $\mu$).
Back to the least-squares run. There $A = X^\top X = \begin{bmatrix}3&3\\3&5\end{bmatrix}$ with eigenvalues $\lambda = 4\pm\sqrt{10} = 7.162$ and $0.838$ ($\kappa = 8.55$). With $\eta = 0.1$ the factors are $1-0.716 = 0.284$ (steep axis) and $1-0.0838 = 0.916$ (flat axis). The start error has size $1.899$ along the steep axis and only $0.069$ along the flat one. After $k$ steps: steep $1.899\cdot0.284^k$ ($0.54, 0.15, 0.043, 0.012$ for $k=1..4$), flat $0.069\cdot0.916^k$ ($0.064, 0.058, 0.053, 0.049$). The steep part is gone by $k=3$, and from then on only the flat part is left, shrinking by a mere $8\%$ per step. That is exactly the "drop, then crawl" we saw. With the best step $\eta^\star = 2/8 = 0.25$ the rate would be $\sqrt{10}/4 = 0.7906$ and the run needs $41$ steps instead of $47$ to reach $\|\nabla f\|<10^{-3}$.
Why do we need it?
It turns "gradient descent works if the learning rate is not too big" into exact numbers: the safe range, the best rate and the number of steps. It also explains why stretched valleys are slow, which is the root of most optimizer design.
Where is it used?
It is the standard analysis behind learning-rate advice for linear and logistic regression, behind preconditioning and feature scaling, and behind the motivation for momentum and Adam (next chapter), conjugate gradient and Newton's method.
How is it used?
Estimate the largest eigenvalue $\lambda_{\max}$ of the Hessian (for example with a few power iterations) and keep $\eta$ below $2/\lambda_{\max}$, ideally near $1/\lambda_{\max}$. The ratio $\lambda_{\max}/\lambda_{\min}$ predicts how many iterations you need.
"Best" depends on which meaning of best. $\eta^\star = 2/(\lambda_{\min}+\lambda_{\max})$ is the best constant step. It is only available if you know both eigenvalues. In practice you rarely know them, so people use $\eta$ somewhat below $1/\lambda_{\max}$ and live with the slow flat axis.
This is exact only for quadratics. For general functions the same numbers are the local behaviour near a minimum (where the Hessian plays the role of $A$). Far from the minimum the curvature can be much larger or smaller, which is why non-convex runs sometimes need a smaller step at the start.
Quick check: $f=\tfrac12(2x_1^2 + 8x_2^2)$. Find the converging range of $\eta$, the best $\eta$ and the best rate.
$A = \mathrm{diag}(2,8)$, so $\lambda_{\min}=2$, $\lambda_{\max}=8$, $\kappa=4$. Range: $0\lt \eta\lt 2/8 = 0.25$. Best: $\eta^\star = 2/(2+8) = 0.2$, rate $\rho^\star=(4-1)/(4+1) = 0.6$. Check: factors $1-0.2\cdot2 = 0.6$ and $1-0.2\cdot8 = -0.6$ ✓.
The learning rate: the one number you must choose core
The learning rate $\eta$ is how bold each step is. The step you take is "slope times $\eta$". Too timid and you crawl; too bold and you leap across the valley and may land higher than you started.
What counts as "too bold" depends on the ground. On gently curved ground you can stride. On sharply curved ground (a narrow, steep bowl) even a short step overshoots. So the safe learning rate is set by the sharpest curvature, which is the largest eigenvalue $\lambda_{\max}$ of the Hessian (called $L$ when we think of it as an upper bound on the curvature). We just derived the exact limit for a quadratic: $\eta\lt 2/L$.
The bowl $f = \tfrac12(x_1^2+10x_2^2)$ has $L=10$, $\mu = 1$ (the smallest curvature), so the limit is $2/L = 0.2$. Steps needed to shrink the error by $10^{-6}$ (from $k\ge\ln(10^6)/\ln(1/\rho)$):
| $\eta$ | rate $\rho=\max(\lvert1-\eta\rvert,\ \lvert1-10\eta\rvert)$ | steps to $10^{-6}$ | verdict |
|---|---|---|---|
| 0.01 | 0.99 | 1375 | too small: safe but 20 times slower than the best |
| 0.1 ($=1/L$) | 0.9 | 131 | safe and good |
| 0.1818 ($=2/(L+\mu)$) | 0.8182 | 69 | the best fixed value |
| 0.19 | 0.9 | 131 | converging but zig-zagging |
| 0.2 ($=2/L$) | 1 | never | edge: bounces forever |
| 0.21 | 1.1 | never | diverges: the error grows $10\%$ per step |
Notice the shape: a gentle curve going down as $\eta$ rises, then a cliff at $2/L$. It is always safer to be a bit too small than a bit too big.
The learning rate (or step size, when it is constant) $\eta\gt 0$ is the scalar in $\mathbf{x}_{k+1}=\mathbf{x}_k-\eta\nabla f(\mathbf{x}_k)$. It is a hyperparameter: a setting you choose, not something the method learns. For an $L$-smooth function:
- Stability: $0\lt \eta\lt 2/L$ (exact for quadratics; for general functions it is the threshold for guaranteed decrease, see the derivation above).
- Safe default: $\eta=1/L$ (guaranteed drop $\|\nabla f\|^2/(2L)$ per step).
- Best constant step on a quadratic: $2/(\lambda_{\min}+\lambda_{\max})$, which gives rate $(\kappa-1)/(\kappa+1)$.
Scale rules (they save debugging time). If you multiply the loss by $c$, then $L$ becomes $cL$, so $\eta$ must become $\eta/c$. In particular, switching from the mean loss over $n$ examples to the sum multiplies $L$ by $n$, so the same $\eta$ would explode.
If you do not know $L$: try $\eta$ on a log grid ($1, 0.3, 0.1, 0.03, 0.01,\dots$) and keep the largest value whose loss curve falls smoothly. Many people also run a short "learning-rate range test" (increase $\eta$ every step and watch where the loss starts to rise).
Why do we need it?
The gradient says which way to go but not how far. $\eta$ is the answer to "how far", and a wrong answer breaks everything: too small wastes days of compute, too large gives a loss of NaN.
Where is it used?
Every optimizer has one: lr in PyTorch's SGD and Adam, learning_rate in scikit-learn and XGBoost. It is usually the first setting tuned, and typically the one that matters most.
How is it used?
Start from a value known to work for similar models, try factors of 3 or 10 up and down, and judge by the loss curve (not the final accuracy). Lower it if the loss spikes; raise it if the curve is a slow, straight, shallow line.
The best learning rate for the loss is not always the best for the final model. In machine learning we care about performance on new data, and the largest stable $\eta$ is not always best for that. Treat the formulas here as a guide to stability and speed, and check validation results too.
Do not copy a learning rate between different problems. It depends on the curvature, the loss scale (mean vs sum), the input scaling, and the optimizer. The same number can be perfect for one model and explode for another.
Quick check: the loss is $f=5(x-1)^2$. Which $\eta$ is the limit, and which gives a one-step jump to the minimum?
$f''=10$, so $L=\lambda=10$. Limit $2/L=0.2$. One step jump: $\eta=1/\lambda=0.1$, since $x\leftarrow x-0.1\cdot10(x-1) = 1$ exactly.
Overshooting: oscillation and divergence core
Picture rolling a marble into a bowl by pushing it each time towards the bottom. If the push is too strong, the marble flies across the bottom and up the other wall. That is overshooting.
There are two possible outcomes. If the marble lands lower than where it started (it crossed the bottom but climbed a smaller hill), the next push is weaker, and it zig-zags in: oscillation that still converges. If it lands higher, the next push is even stronger, and each swing is bigger than the last: divergence.
Minimise $f(x)=2x^2$ (curvature $\lambda=4$, so $f'(x)=4x$). One step is $x\leftarrow x-4\eta x=(1-4\eta)x$. Start at $x_0=3$; the multiplier is $m=1-4\eta$:
| $\eta$ | $m$ | $x_0,x_1,x_2,x_3$ | name |
|---|---|---|---|
| 0.1 | 0.6 | 3, 1.8, 1.08, 0.648 | smooth approach (same side) |
| 0.25 | 0 | 3, 0, 0, 0 | one-step jump to the bottom |
| 0.4 | $-0.6$ | 3, $-1.8$, 1.08, $-0.648$ | overshoot, oscillation, converges |
| 0.5 | $-1$ | 3, $-3$, 3, $-3$ | endless bounce (never converges) |
| 0.55 | $-1.2$ | 3, $-3.6$, 4.32, $-5.184$ | overshoot and diverge |
The whole story is the sign and size of one number, $m=1-\eta\lambda$.
On a parabola with curvature $\lambda$, gradient descent is the straight-line map $x_{k+1}=m\,x_k$ with the multiplier $m=1-\eta\lambda$. (Measure $x$ from the minimum.) Then:
- $0\lt m\lt 1$ ($0\lt \eta\lt 1/\lambda$): the iterates approach from one side. No overshoot.
- $m=0$ ($\eta=1/\lambda$): lands on the minimum in one step.
- $-1\lt m\lt 0$ ($1/\lambda\lt \eta\lt 2/\lambda$): overshoot. $x$ changes sign every step but $|x|$ shrinks: oscillating convergence.
- $m=-1$ ($\eta=2/\lambda$): a two-cycle $+c,-c,+c,\dots$, no progress.
- $m\lt -1$ ($\eta\gt 2/\lambda$): oscillating divergence; $|x|$ grows like $|m|^k$.
In several dimensions each axis has its own $m_i=1-\eta\lambda_i$. The steepest axis (largest $\lambda_i$) overshoots first and diverges first, so $\eta\lt 2/\lambda_{\max}$ protects all axes. A zig-zag across a valley is exactly "the steep axis has $-1\lt m\lt 0$ while the flat axis crawls with $m$ close to $+1$".
Why do we need it?
"Loss went up" or "loss is jumping around" are the two most common training symptoms. Knowing the multiplier $1-\eta\lambda$ tells you which one you are seeing, and that the cure for both is a smaller $\eta$.
Where is it used?
Diagnosing training curves; choosing the step in numerical solvers; one common explanation of learning-rate warmup (start with a small $\eta$ because the early curvature may be large) and of gradient clipping (next chapters).
How is it used?
If the loss alternates up and down with a shrinking envelope, you are slightly over $1/\lambda$ (fine, but wasteful). If it grows, you are over $2/\lambda$: cut $\eta$ by at least a factor of 2 and restart from the last good point.
Overshoot is not always a bug. For $1/\lambda\lt \eta\lt 2/\lambda$ the iterates still converge, and for a stretched bowl the best fixed step ($2/(\lambda_{\min}+\lambda_{\max})$) deliberately overshoots on the steep axis, as the table in the quadratic section showed.
Real losses are not parabolas. The curvature changes from place to place, so a step that is fine in a flat region can overshoot in a steep one. Training runs sometimes show a sudden spike early on (steep region) and then settle. This is one reason for learning-rate warmup and gradient clipping (Chapter 3.4 and Chapter 3.15).
Quick check: $f=3x^2$ ($\lambda=6$), $x_0=2$, $\eta=0.25$. Write $x_1,x_2$ and say what happens.
$m = 1-0.25\cdot6 = -0.5$. So $x_1=-1$, $x_2=0.5$, $x_3 = -0.25$. Overshoot (the sign flips), but $|m|\lt 1$, so it converges.
Slow convergence: narrow valleys (ill-conditioning) and plateaus core
Gradient descent can be slow for two different reasons.
1. A long, narrow valley. The walls are steep, the floor is almost flat. The steepest-descent arrow points mostly at the opposite wall, not along the valley. To avoid flying over the walls you must take small steps, but small steps along the floor go nowhere. The path zig-zags from wall to wall, creeping down the valley. A round bowl has no such problem: every direction points at the bottom.
2. A plateau. A region where the ground is almost flat. The gradient is tiny, so the step $\eta\|\nabla f\|$ is tiny, even if we are still far from the bottom. It feels like walking through deep sand.
Valley. From the quadratic section: with the best fixed step, the number of steps to shrink the error by $10^{-6}$ is $\ln(10^6)/\ln\!\big(\tfrac{\kappa+1}{\kappa-1}\big)$:
| $\kappa=\lambda_{\max}/\lambda_{\min}$ | 1 (round) | 3 | 10 | 30 | 100 | 1000 |
|---|---|---|---|---|---|---|
| rate $\rho=\frac{\kappa-1}{\kappa+1}$ | 0 | 0.5 | 0.818 | 0.935 | 0.980 | 0.998 |
| steps for $10^{-6}$ | 1 | 20 | 69 | 207 | 691 | 6908 |
Ten times more stretched means about ten times more steps.
Plateau. Take $f(x)=x^4/4$, so $f'(x)=x^3$ and one step is $x\leftarrow x-\eta x^3 = x(1-\eta x^2)$. Start at $x_0=1$ with $\eta=0.1$: $x_1=0.9$, $x_2=0.8271$, $x_3=0.7705$, and after 10, 100 and 1000 steps: $x=0.5604,\ 0.2158,\ 0.0704$. Compare the parabola $x^2/2$ with the same step: $x_k=0.9^k$ gives $0.349$ after 10 steps and $0.00003$ after 100. The quartic bowl is flat at the bottom ($f''=3x^2\to0$), so as $x$ shrinks the slope $x^3$ shrinks even faster, and the steps almost stop. The decay is only about $1/\sqrt{2\eta k}$: very slow.
The condition number of the Hessian is $\kappa=\lambda_{\max}/\lambda_{\min}\ge1$: how much steeper the steepest direction is than the flattest. A problem with large $\kappa$ is ill-conditioned. On a quadratic, gradient descent with the best fixed step needs about $\frac\kappa2\ln\frac1\varepsilon$ iterations to reach accuracy $\varepsilon$: proportional to $\kappa$.
A plateau is a region where $\|\nabla f\|$ is small although we are far from a minimum (or where the curvature vanishes at the minimum, as in $x^4$). Near such places the convergence is sublinear (the error shrinks like $1/\sqrt{k}$ or $1/k$, not like $\rho^k$).
Remedies (each is a topic later in the guide):
- Rescale the problem (you can do this today): standardise and centre the features, which makes the valley rounder. Our least-squares example went from $\kappa=8.55$ to $1.5$ just by centring $x$. This idea, in general, is called preconditioning.
- Momentum (Chapter 3.4): averaging past gradients damps the zig-zag and speeds up the floor; on quadratics it turns $\kappa$ into about $\sqrt\kappa$ in the iteration count.
- Adaptive methods such as AdaGrad, RMSProp and Adam (Chapter 3.4): a separate step size for each coordinate.
- Second-order methods (Chapter 3.12): Newton's method and L-BFGS use curvature to point at the bottom, and on a quadratic Newton needs one step.
Why do we need it?
Most slow training runs are not "the optimizer is broken" but "the landscape is a stretched valley or a plateau". Knowing the cause tells you the cure: rescale, add momentum or switch method.
Where is it used?
Feature standardisation before regression and neural networks, batch and layer normalisation (one common explanation of why they help is that they make the landscape rounder), Adam as the default for Transformers, L-BFGS for logistic regression in scikit-learn, and plateau detection in learning-rate schedulers.
How is it used?
Standardise your inputs first. If the loss still creeps with a smooth, shallow slope, suspect a narrow valley: try momentum or Adam. If it is flat for many steps and then drops, you probably crossed a plateau: be patient or raise $\eta$.
Slow is not the same as stuck. A shallow but steady downward loss curve means you are still improving. A truly flat curve with a near-zero gradient means a flat spot, which might be a minimum or might not (plateaus and saddles are covered in Chapter 3.15).
Raising $\eta$ is not the cure for a narrow valley. The limit $2/\lambda_{\max}$ is set by the steep axis, so a bigger step only makes the zig-zag worse. The fix must change the geometry (rescaling, preconditioning) or the method (momentum, Adam, Newton).
Quick check: why does centring the inputs speed up gradient descent on least squares?
The Hessian is $X^\top X$. With an uncentred input column next to a column of ones, the two columns are strongly correlated, so $X^\top X$ has one big and one tiny eigenvalue (a narrow valley, large $\kappa$). Centring makes the columns uncorrelated, so the eigenvalues become comparable (a rounder bowl, small $\kappa$) and the step count, which grows like $\kappa$, drops.
Step-size strategies: fixed, decreasing and line search
So far the step size $\eta$ was one fixed number. There are three common ways to treat it.
- Fixed: one number for the whole run. Simple and, on smooth problems, hard to beat when it is well chosen.
- Decreasing: start bold, get more careful. $\eta_k$ shrinks as $k$ grows. It matters most when the gradient is noisy (stochastic gradient descent, below): then a fixed step keeps getting kicked around the bottom, and a smaller step later calms it. On a smooth deterministic problem it usually just makes the run slower.
- Line search: at every step, test a big step and shrink it until it really goes down. Like testing the ice with your foot before putting your weight on it. It adapts to the local curvature automatically, at the price of extra function evaluations.
Backtracking on our standard bowl. $f=x_1^2+3x_2^2$ at $(4,2)$, $\mathbf{g}=[8,12]$, $\|\mathbf{g}\|^2=208$, $f=28$. Try $\alpha=1$ first and halve it until the Armijo test $f(\mathbf{x}-\alpha\mathbf{g})\le f(\mathbf{x})-c\,\alpha\|\mathbf{g}\|^2$ passes (use $c=10^{-4}$):
- $\alpha=1$: new point $(-4,-10)$, $f=16+300=316$. Needed $\le 28-0.0208=27.98$. Reject.
- $\alpha=0.5$: new point $(0,-4)$, $f=48$. Needed $\le27.99$. Reject.
- $\alpha=0.25$: new point $(2,-1)$, $f=4+3=7$. Needed $\le27.995$. Accept. The height fell from $28$ to $7$.
For comparison, the exact line search (the best $\alpha$ along the ray) is $\alpha=208/992=0.2097$, giving $f=6.19$. Backtracking found a nearly-as-good step with three tests and no knowledge of the curvature.
Fixed: $\eta_k=\eta$. Decreasing (common schedules): $\eta_k=\dfrac{\eta_0}{1+d\,k}$ (inverse time), $\eta_k=\eta_0\gamma^k$ (exponential), $\eta_k=\dfrac{\eta_0}{\sqrt{k+1}}$. (More schedules, such as cosine and warmup, are in Chapter 3.4.) For noisy gradients the classical requirement is $\sum_k\eta_k=\infty$ (we can still travel any distance) and $\sum_k\eta_k^2\lt \infty$ (the noise eventually dies out) (Chapter 3.14).
Backtracking line search (awareness): choose a starting length $\alpha_0$ (often $1$), a shrink factor $\beta\in(0,1)$ (often $\tfrac12$) and a tiny $c$ (often $10^{-4}$). At each iteration, with direction $\mathbf{d}=-\nabla f(\mathbf{x}_k)$:
- Set $\alpha=\alpha_0$.
- While $f(\mathbf{x}_k+\alpha\mathbf{d})\gt f(\mathbf{x}_k)+c\,\alpha\,\nabla f(\mathbf{x}_k)^\top\mathbf{d}$: set $\alpha\leftarrow\beta\alpha$.
- Take the step $\mathbf{x}_{k+1}=\mathbf{x}_k+\alpha\mathbf{d}$.
The inequality is the Armijo (sufficient decrease) condition: "the function must fall by at least a fraction $c$ of what the tilted-board model promised". Since $\nabla f^\top\mathbf{d}=-\|\nabla f\|^2$ here, this is the test in the example. It always stops after finitely many halvings. An exact line search minimises $f$ along the ray; it is rarely worth its cost. Stronger conditions (Wolfe) are used with quasi-Newton methods (Chapter 3.12).
Why do we need it?
The right step size is unknown in advance and changes along the path: steep early, flat later, noisy sometimes. A strategy that adapts spares us a lot of trial and error.
Where is it used?
Fixed steps for most deep-learning runs (often with a schedule on top); decreasing steps for SGD in theory and in practice; backtracking and Wolfe line searches inside scipy.optimize (BFGS, CG, L-BFGS-B) and classical solvers.
How is it used?
Deterministic, smooth problem with cheap function values: use a line search and forget about $\eta$. Mini-batch training: a fixed (or scheduled) $\eta$, because the loss on one batch is too noisy for a line search to be meaningful.
Decreasing steps are a cure for noise, not for slowness. On a smooth, deterministic problem a constant step already converges linearly. A schedule that shrinks too fast (for example $1/k^2$) can stop long before the bottom, because $\sum\eta_k$ stays finite. We return to this in Chapter 3.14.
Line searches need a trustworthy function value. With mini-batches the loss changes from batch to batch, so "did the loss go down?" is not a meaningful test. That is why deep learning rarely uses line search and instead tunes $\eta$ and a schedule.
Quick check: $f=x^2$ at $x=2$, $\mathbf{g}=4$, $\alpha_0=1$, halving, $c=10^{-4}$. Which $\alpha$ is accepted?
$\alpha=1$: new $x=2-4=-2$, $f=4$. Needed $f\le 4-10^{-4}\cdot16=3.9984$. Not met: reject. $\alpha=0.5$: new $x=0$, $f=0\le3.9992$: accept. Here $\alpha=0.5=1/f''$, the perfect step.
Initialization: where you start decides where you end core
Gradient descent needs a starting point $\mathbf{x}_0$. Does the choice matter?
On a single bowl, no. Wherever you start, you roll to the same bottom. The start only changes how long the trip takes.
On a landscape with several valleys, yes. Think of rain falling on a mountain range: a drop that lands on the east side of a ridge ends up in the east river, a drop on the west side ends up in the west river. The set of starting points that flow to the same valley is that valley's basin (its watershed). The ridge between basins is a line you can cross by a tiny shove.
One dimension. $f(x)=0.1x^4-0.8x^2+0.4x$ has two valleys, at $x=-2.1149$ (height $-2.4236$, the deepest, the global minimum) and at $x=1.8608$ (height $-0.8268$, a local minimum). The hilltop between them is at $x=0.2541$. Gradient descent with $\eta=0.1$ gives:
| start $x_0$ | $-1$ | $0$ | $0.2$ | $0.3$ | $1$ |
|---|---|---|---|---|---|
| ends at | $-2.1149$ | $-2.1149$ | $-2.1149$ | $1.8608$ | $1.8608$ |
Starting just left of the hilltop ($0.2$) or just right of it ($0.3$) gives completely different answers.
Two dimensions. Himmelblau's function has four minima, all of height $0$ (equally good). With $\eta=0.01$: from $(0,0)$ the run ends at $(3,2)$; from $(-1,1)$ at $(-2.805,3.131)$; from $(1,-1)$ at $(3.584,-1.848)$; from $(-2,-2)$ at $(-3.779,-3.283)$. On a $30\times30$ grid of starting points the four valleys catch $231$, $233$, $213$ and $223$ of the $900$ starts: each takes about a quarter of the square.
Initialization is the choice of $\mathbf{x}_0$. The basin of attraction of a minimum $\mathbf{x}^\star$ is the set of starting points from which gradient descent (with a given $\eta$) converges to $\mathbf{x}^\star$.
- Convex problems (Chapter 3.6): there are no bad valleys; with $\eta\lt 2/L$ every start converges to a global minimum. A zero start $\mathbf{x}_0=\mathbf{0}$ is a perfectly good default (linear regression, logistic regression).
- Non-convex problems: the start decides which stationary point you reach. Practical answers: random restarts (run from several random starts and keep the lowest loss), warm starts (start from a solution of a similar problem, such as a pre-trained model), or a problem-specific initialization.
- Neural networks (a forward note, Chapter 3.15): weights are started at small random numbers, scaled by the width of the layer (the names to look up are Xavier/Glorot and He initialization). Two reasons: if all weights start equal (for example zero) all the hidden units in a layer get identical gradients and stay identical forever (the symmetry is never broken), and the scale controls whether signals and gradients shrink or blow up as they pass through many layers.
Symmetry traps. Starting exactly on a line of symmetry can lock you out of the good valleys: for $f=x^2-y^2+y^4/4$ a start with $y=0$ never leaves the saddle at the origin ($\partial f/\partial y = -2y+y^3=0$ stays $0$), while any $y\neq0$ slides away. This is the same reason zero-initialized networks do not learn.
Why do we need it?
Because there is no "no start". And because on non-convex losses the start can decide whether you get a great model or a mediocre one, so it is worth knowing which starts are safe.
Where is it used?
Zero or small starts for regression and logistic regression; random restarts for $k$-means and matrix factorisation; Glorot/He initialization in every deep-learning framework; fine-tuning (warm-starting from a pre-trained model).
How is it used?
Convex model: start at zero (or any point) and stop worrying. Non-convex model: use the framework's default initializer, set a random seed so runs are reproducible, and, if results vary a lot, compare several seeds.
"Random" does not mean "careless". Starting a deep network with too large or too small random weights makes the signals vanish or explode, so the scale of the random numbers matters as much as the randomness (Chapter 3.15).
A "good" start does not mean the global minimum. Gradient descent finds a valley near the start. Whether that valley is the deepest one is not guaranteed on non-convex problems.
Quick check: logistic regression has a convex loss. Is it safe to start all weights at zero? And a neural network with one hidden layer?
For logistic regression, yes: convex loss, so every start (with a sensible $\eta$) leads to the global minimum, and zero is fine. For a neural network, no: all hidden units would receive identical gradients and remain copies of each other. Networks need random starting weights to break the symmetry.
Batch, stochastic and mini-batch gradient descent: only the gradient estimate changes core
In machine learning the objective is an average over the training examples: $f(\mathbf{w})=\frac1n\sum_{i=1}^n\ell_i(\mathbf{w})$, where $\ell_i$ is the loss on example $i$ (see the Calculus guide). The exact gradient is the average of $n$ per-example gradients. With $n=10^7$ images, one single step would need ten million gradient computations. Far too slow.
Think of an election poll. To learn the opinion of ten million voters, you do not ask all of them. You ask a few hundred at random, and the answer is close. Gradient descent does the same: estimate the average gradient from a random few examples. The estimate is noisy, but it points roughly the right way, it is thousands of times cheaper, and we take many steps anyway, so the noise averages out.
The three versions differ in only one thing: how many examples we "poll" for each step. All of them (batch), one (stochastic), or a few (mini-batch). The update rule is the same: $\mathbf{w}\leftarrow\mathbf{w}-\eta\,\hat{\mathbf{g}}$, where $\hat{\mathbf{g}}$ is whatever gradient estimate we computed.
A tiny problem: one parameter, model $\hat y = w x$, four examples $(x,y) = (1,2),(2,3),(3,5),(4,9)$, loss $\ell_i=\tfrac12(wx_i-y_i)^2$, so $\nabla\ell_i=(wx_i-y_i)\,x_i$. At $w=1$:
- Residuals $wx_i-y_i = 1-2,\ 2-3,\ 3-5,\ 4-9 = -1,\,-1,\,-2,\,-5$.
- Per-example gradients (residual times $x_i$): $-1,\ -2,\ -6,\ -20$.
- Batch: the average is $(-1-2-6-20)/4 = \mathbf{-7.25}$. This is the true gradient.
- SGD: pick one example at random: the estimate is $-1$, $-2$, $-6$ or $-20$, each with probability $\tfrac14$. Very different from $-7.25$ in any single draw, but the average over the four equally likely draws is exactly $-7.25$. (Unbiased.) The typical spread (standard deviation) is $7.60$.
- Mini-batch, $B=2$: there are six possible pairs, with averages $-1.5,\ -3.5,\ -10.5,\ -4,\ -11,\ -13$. Their average is again $-7.25$, and their spread is $4.39$, smaller than $7.60$ (about $1/\sqrt3$ of it here: roughly $1/\sqrt B$ with a small bonus for sampling without replacement).
Let $f(\mathbf{w})=\frac1n\sum_i\ell_i(\mathbf{w})$. An epoch is one pass through all $n$ examples. In every variant the update is $\mathbf{w}\leftarrow\mathbf{w}-\eta\,\hat{\mathbf{g}}$ with
| Variant | Gradient estimate $\hat{\mathbf{g}}$ | Cost of one update | Updates per epoch | Noise |
|---|---|---|---|---|
| Batch gradient descent | $\frac1n\sum_{\text{all }i}\nabla\ell_i=\nabla f$ (exact) | $n$ gradients | $1$ | none |
| Stochastic (SGD) | $\nabla\ell_i$ for one random example $i$ | $1$ gradient | $n$ | largest: variance $\sigma^2$ |
| Mini-batch, size $B$ | $\frac1B\sum_{i\in\text{batch}}\nabla\ell_i$ | $B$ gradients | $n/B$ | variance about $\sigma^2/B$ |
Here $\sigma^2=\frac1n\sum_i\|\nabla\ell_i-\nabla f\|^2$ measures how much the examples disagree. Two facts make the whole idea work:
- Unbiased: $\mathbb{E}[\hat{\mathbf{g}}]=\nabla f(\mathbf{w})$. On average the estimate is the true gradient.
- Variance falls like $1/B$ (for samples drawn independently), so the noise (standard deviation) falls like $1/\sqrt B$: four times the batch, half the noise.
Cost per epoch is the same for all three (about $n$ gradient computations); what differs is how many updates you get out of it: $1$, $n$ or $n/B$. Typical loop: shuffle the data, cut it into batches of size $B$, and for each batch compute $\hat{\mathbf{g}}$ and update. In deep-learning libraries "SGD" almost always means mini-batch SGD. Common batch sizes are $32$ to a few hundred (a habit, not a law).
Behaviour near the bottom. With a fixed $\eta$ the noise never fully fades, so SGD does not settle at $\mathbf{w}^\star$: it hovers in a "noise ball" around it, whose size grows with $\eta$ and shrinks with $B$. Shrinking $\eta$ over time (or growing $B$) tightens the ball (details in Chapter 3.14).
Why do we need it?
Real data sets are too large to compute the exact gradient for every step. Using a small random sample gives many cheap, slightly noisy steps, and that wins over a few exact, expensive ones.
Where is it used?
All deep-learning training (mini-batch SGD, Adam and others) uses mini-batches; SGDRegressor and SGDClassifier in scikit-learn use single examples; full-batch gradient descent is used for small convex problems and with L-BFGS.
How is it used?
Choose a batch size $B$ that fits in memory, shuffle every epoch, and compute the gradient on each batch. Tune $\eta$ for that batch size. Bigger batches give smoother but fewer updates; smaller batches are noisier but update more often.
"Faster" depends on what you count. Per epoch (a pass over the data), SGD and mini-batch usually make much more progress than batch. Per update, batch is the best (its gradient is exact). Per second of wall-clock time on a GPU, mini-batches win because the hardware computes 64 or 256 gradients almost as fast as one.
Always reshuffle. If the examples are sorted (all the cats, then all the dogs), consecutive batches are very different from each other and training zig-zags badly. Shuffle once per epoch.
A learning rate tuned for batch GD can break SGD. A single example has much larger curvature than the average (compare the per-example Hessian $x_ix_i^\top$ with the mean), so the stable $\eta$ for SGD can be (and often is) smaller than $2/\lambda_{\max}$ of the full problem.
Quick check: 50,000 examples, batch size 100. How many updates per epoch, and how many for 10 epochs?
$50000/100=500$ updates per epoch, so $5000$ updates for 10 epochs. Each update used 100 gradients, so the total is $5000\times100=500{,}000=10\times50{,}000$ gradient evaluations, exactly what 10 epochs of any variant costs.
The whole picture in 3D: walking down the loss surface
Until now we looked at landscapes from above, as contour maps. Now stand at the side and look at the real surface: the two parameters run along the floor, and the loss is the height. Gradient descent is a walker on that surface who always steps in the steepest-downhill direction as seen from above (perpendicular to the contour line), then re-measures.
It is not a rolling ball. A ball has momentum and keeps moving; gradient descent has none. It takes a step, forgets everything, and measures the slope again. (A ball with memory is "momentum": Chapter 3.4.)
Our first example, $f=x_1^2+3x_2^2$ with $\eta=0.1$, visits the points $(4,2)\to(3.2,0.8)\to(2.56,0.32)\to(2.048,0.128)$. On the surface these are the heights $28\to12.16\to6.86\to4.24$: a staircase of points sliding down one wall of the bowl, falling quickly across the steep $x_2$ direction first and then crawling along the gentle $x_1$ direction towards the origin. Seen from above (the Top view of the widget below) that is the contour picture you already know.
The loss surface (or loss landscape) of an objective $f$ with two parameters is the graph $z=f(x_1,x_2)$. Its contour lines are the level sets $\{f=c\}$. Valleys are minima, hilltops are maxima, and a saddle is a point that is a minimum in one direction and a maximum in another (like a mountain pass; Chapter 3.2). Gradient descent produces a path that, seen from above, crosses each contour line at a right angle (the gradient is perpendicular to the contour).
Real models have millions of parameters, so their surface cannot be drawn. People draw two-dimensional slices through it and rely on the intuition from low dimensions, which is useful but not the whole truth (Chapter 3.15).
Why do we need it?
A picture of the landscape builds the mental model for every phenomenon in this chapter: valleys, ridges, flat plateaus, saddles, several minima. With it, "ill-conditioned" or "local minimum" are shapes you can see, not just words.
Where is it used?
Loss-landscape plots in papers on neural networks, teaching demos of optimizers, and sanity checks of a loss in two parameters (for example the weight and bias of a tiny model).
How is it used?
Fix all parameters except two, evaluate the loss on a grid, and plot the surface or its contours. Overlay the optimizer's path. Look for narrow valleys, flat regions and several basins.
3D pictures can fool the eye. A step that looks short on the screen can be long in the parameters (and the other way round), and the height is scaled. Use the Top view to judge distances and angles, and the Front view to judge heights.
Low-dimensional intuition has limits. In two dimensions a local minimum looks like a trap. In very high dimensions, saddle points are much more common than in two (whether they are the main obstacle to training is still debated; see Chapter 3.15), and real loss surfaces are far more complicated than any picture here.
Quick check: in the Top view, how does a gradient-descent step meet the contour line it starts on?
At a right angle. The step is $-\eta\nabla f$, and the gradient is perpendicular to the contour line through the point (the directional derivative along the contour is zero).
Recap, cheat sheet and practice
- Gradient descent walks downhill in fog: feel the gradient, step against it, repeat. $\mathbf{x}_{k+1}=\mathbf{x}_k-\eta\nabla f(\mathbf{x}_k)$. The direction $-\nabla f$ is steepest descent because the slope along a unit direction $\mathbf{u}$ is $\nabla f\cdot\mathbf{u}=\|\nabla f\|\cos\theta$.
- Derivation: the first-order model $f(\mathbf{x}+\mathbf{d})\approx f+\nabla f^\top\mathbf{d}$ is minimised by $\mathbf{d}\parallel-\nabla f$ (Cauchy–Schwarz); adding the price $\|\mathbf{d}\|^2/(2\eta)$ gives exactly $\mathbf{d}=-\eta\nabla f$. If the curvature is at most $L$: $f(\mathbf{x}_{k+1})\le f(\mathbf{x}_k)-\eta(1-L\eta/2)\|\nabla f\|^2$, a drop for every $\eta\lt 2/L$.
- On a quadratic $\tfrac12\mathbf{x}^\top A\mathbf{x}-\mathbf{b}^\top\mathbf{x}$: the error along eigen-direction $i$ is multiplied by $(1-\eta\lambda_i)$ each step. Converges iff $0\lt \eta\lt 2/\lambda_{\max}$. Rate $\rho(\eta)=\max_i|1-\eta\lambda_i|$; best fixed step $\eta^\star=2/(\lambda_{\min}+\lambda_{\max})$ with $\rho^\star=\frac{\kappa-1}{\kappa+1}$; steps for accuracy $\varepsilon$ are about $\frac\kappa2\ln\frac1\varepsilon$.
- Learning rate: too small crawls, $1/\lambda\lt \eta\lt 2/\lambda$ overshoots but converges (zig-zag), above $2/\lambda$ diverges. Multiplier $m=1-\eta\lambda$. Scale the loss by $c$, divide $\eta$ by $c$.
- Slow convergence: narrow valleys (large $\kappa$: zig-zag, steps $\propto\kappa$) and plateaus (vanishing gradient, sublinear). Remedies: rescale and centre features, momentum and Adam (Chapter 3.4), Newton and L-BFGS (Chapter 3.12).
- Stopping: gradient small (best), steps small (fooled by tiny $\eta$), loss stalls, and always a maximum number of iterations. Small gradient means flat ground, not necessarily the global minimum.
- Initialization: any start works on a convex problem (zero is fine); on non-convex problems the start picks the valley (basins), so use random restarts; neural networks need random starts to break symmetry.
- Step-size strategies: fixed; decreasing (for noisy gradients); backtracking line search with the Armijo test $f(\mathbf{x}-\alpha\mathbf{g})\le f-c\alpha\|\mathbf{g}\|^2$.
- Batch, SGD, mini-batch differ only in the gradient estimate $\hat{\mathbf{g}}$ (all $n$, one, or $B$ examples). It is unbiased with variance about $\sigma^2/B$; an epoch is $n$ gradient evaluations and gives $1$, $n$ or $n/B$ updates.
Cheat sheet
| Idea | Formula | Remember |
|---|---|---|
| Update rule | $\mathbf{x}_{k+1}=\mathbf{x}_k-\eta\,\nabla f(\mathbf{x}_k)$ | minus sign: gradient points uphill |
| Step length | $\|\mathbf{x}_{k+1}-\mathbf{x}_k\|=\eta\|\nabla f\|$ | steps shrink as the ground flattens |
| Guaranteed decrease | $f_{k+1}\le f_k-\eta(1-\tfrac{L\eta}{2})\|\mathbf{g}_k\|^2$ | needs curvature $\le L$; $\eta=1/L$ drops by $\|\mathbf{g}\|^2/2L$ |
| Quadratic: one axis | $c_i(k)=(1-\eta\lambda_i)^kc_i(0)$ | straight line on a log plot |
| Stability limit | $0\lt \eta\lt 2/\lambda_{\max}$ ($=2/L$) | edge: bounces forever |
| Best fixed step and rate | $\eta^\star=\frac{2}{\lambda_{\min}+\lambda_{\max}}$, $\rho^\star=\frac{\kappa-1}{\kappa+1}$ | $\kappa=\lambda_{\max}/\lambda_{\min}$ |
| Steps for accuracy $\varepsilon$ | $\ln(1/\varepsilon)/\ln(1/\rho)\approx\frac\kappa2\ln\frac1\varepsilon$ | 10 times more stretched, 10 times more steps |
| Multiplier (1D) | $m=1-\eta\lambda$ | $0\lt m\lt 1$ smooth; $-1\lt m\lt 0$ zig-zag; $m\lt -1$ diverges |
| Stopping rules | $\|\nabla f\|\le\varepsilon$; $\|\Delta\mathbf{x}\|\le\varepsilon$; $|\Delta f|\le\varepsilon$; $k\le k_{\max}$ | report which rule fired |
| Armijo test | $f(\mathbf{x}-\alpha\mathbf{g})\le f(\mathbf{x})-c\,\alpha\|\mathbf{g}\|^2$ | halve $\alpha$ until it holds; $c\approx10^{-4}$ |
| Decreasing step | $\eta_k=\eta_0/(1+dk)$ | for noisy gradients |
| Batch gradient | $\frac1n\sum_{i=1}^n\nabla\ell_i$ | $n$ gradients, $1$ update per epoch |
| SGD | $\nabla\ell_i$, random $i$ | $1$ gradient, $n$ updates per epoch |
| Mini-batch | $\frac1B\sum_{i\in\text{batch}}\nabla\ell_i$ | unbiased; noise $\propto1/\sqrt B$; $n/B$ updates per epoch |
| Initialization | convex: any start; non-convex: random restarts | networks: random, never all equal |
import numpy as np
# ---------- 0. the hand-worked run from the chapter: three points, eta = 0.1 ----------
X = np.array([[1., 0.], [1., 1.], [1., 2.]]) # each row is [1, x_i]
y = np.array([1., 3., 4.])
f_ls = lambda w: 0.5 * np.sum((X @ w - y) ** 2)
grad_ls = lambda w: X.T @ (X @ w - y) # gradient = X^T (Xw - y)
w = np.zeros(2)
for k in range(4):
w = w - 0.1 * grad_ls(w)
print(k + 1, w.round(4), round(f_ls(w), 4))
# 1 [0.8 1.1] 1.125
# 2 [1.03 1.41] 0.1685
# 3 [1.098 1.496] 0.0913
# 4 [1.1198 1.5186] 0.0849
# ---------- 1. the quadratic theory: Hessian eigenvalues, 2/L, best step, best rate ----------
lam = np.linalg.eigvalsh(X.T @ X)
print(lam.round(3), "2/L =", round(2 / lam[1], 4))
kappa = lam[1] / lam[0]
print("best eta", 2 / (lam[0] + lam[1]), " rate", round((kappa - 1) / (kappa + 1), 4))
# [0.838 7.162] 2/L = 0.2792
# best eta 0.25 rate 0.7906
# ---------- 2. gradient descent from scratch, with a stopping rule ----------
rng = np.random.default_rng(0)
n = 100
x = rng.uniform(-1, 1, n)
data_y = 1 + 2 * x + 0.3 * rng.standard_normal(n) # y = 1 + 2x + noise
Xd = np.column_stack([np.ones(n), x])
def loss(w): # f(w) = (1/2n) * sum (w.x_i - y_i)^2
return 0.5 * np.mean((Xd @ w - data_y) ** 2)
def grad(w, idx=None): # gradient on the chosen examples (all by default)
Xb, yb = (Xd, data_y) if idx is None else (Xd[idx], data_y[idx])
return Xb.T @ (Xb @ w - yb) / len(yb)
w_star = np.linalg.solve(Xd.T @ Xd, Xd.T @ data_y) # exact answer (normal equations), to compare with
print("exact answer ", w_star.round(4), round(loss(w_star), 5))
# exact answer [0.9782 1.9804] 0.0423
def gradient_descent(w0, eta, tol=1e-6, max_iter=10_000):
w = w0.copy()
for k in range(max_iter):
g = grad(w) # the gradient on ALL n examples
if np.linalg.norm(g) < tol: # stopping rule: gradient (almost) zero
return w, k
w = w - eta * g # the update rule
return w, max_iter # ran out of budget: did NOT converge
w, k = gradient_descent(np.zeros(2), eta=0.5)
print("batch GD ", w.round(4), "after", k, "steps")
# batch GD [0.9782 1.9804] after 68 steps
# ---------- 3. the learning rate: the stability threshold is 2 / (largest Hessian eigenvalue) ----------
H = Xd.T @ Xd / n # Hessian of f
L = np.linalg.eigvalsh(H).max()
print("largest eigenvalue", round(L, 4), " -> converges for eta <", round(2 / L, 4))
for eta in (0.5, 1.0, 2 / L + 0.05):
w, k = gradient_descent(np.zeros(2), eta, max_iter=200)
print(f" eta = {eta:.3f}: |w - w*| = {np.linalg.norm(w - w_star):.2e}")
# largest eigenvalue 1.0146 -> converges for eta < 1.9712
# eta = 0.500: |w - w*| = 2.28e-06
# eta = 1.000: |w - w*| = 2.51e-06
# eta = 2.021: |w - w*| = 2.51e+04 (just above 2/L: it blows up)
# ---------- 4. stochastic and mini-batch: the SAME update, a different gradient estimate ----------
def sgd(w0, eta, epochs, batch_size, seed=0):
r = np.random.default_rng(seed)
w = w0.copy(); updates = 0
for epoch in range(epochs):
order = r.permutation(n) # reshuffle once per epoch
for start in range(0, n, batch_size):
idx = order[start:start + batch_size]
w = w - eta * grad(w, idx) # gradient estimated from just these examples
updates += 1
return w, updates
for name, B, eta in (("batch (B = n)", n, 0.5), ("mini-batch B = 10", 10, 0.2), ("SGD (B = 1)", 1, 0.05)):
w, u = sgd(np.zeros(2), eta, epochs=20, batch_size=B)
print(f"{name:18s}", w.round(4), "updates:", u, " loss:", round(loss(w), 5))
# batch (B = n) [0.9831 1.9474] updates: 20 loss: 0.0425
# mini-batch B = 10 [0.9918 1.9826] updates: 200 loss: 0.0424
# SGD (B = 1) [1.009 1.9718] updates: 2000 loss: 0.04277 (close, but it never settles exactly)
1. Which line is gradient descent?
2. $f(x_1,x_2)=\tfrac12(x_1^2+10x_2^2)$. For which learning rate do the iterates blow up?
3. A quadratic has Hessian eigenvalues $1$ and $4$. What are the best fixed learning rate and the resulting rate?
4. Which stopping rule can be fooled into stopping far from the answer just by choosing a tiny learning rate?
5. A data set has $n=1000$ examples. Mini-batch gradient descent with $B=50$ runs for 5 epochs. How many updates are made?
6. Which statement about initialization is correct?
Practice problems
A. $f(x_1,x_2)=x_1^2+2x_2^2$. Do one gradient-descent step from $(3,-1)$ with $\eta=0.2$. What is the stability limit, and what is the best fixed $\eta$ with its rate?
$\nabla f=[2x_1,\ 4x_2]=[6,-4]$. New point $=(3,-1)-0.2[6,-4]=(1.8,\ -0.2)$. The height fell from $9+2=11$ to $3.24+0.08=3.32$. The Hessian is $\mathrm{diag}(2,4)$, so $\lambda_{\max}=4$ and the limit is $\eta\lt 2/4=0.5$. Best: $\eta^\star=2/(2+4)=1/3$, rate $\rho^\star=(4-2)/(4+2)=1/3$.
B. $f=\tfrac12(2x_1^2+5x_2^2)$. Find the range of $\eta$ that converges, the best $\eta$ and rate, and how many steps are needed to shrink the error by $10^{-6}$ at the best rate.
$\lambda_{\min}=2$, $\lambda_{\max}=5$, $\kappa=2.5$. Converges for $0\lt \eta\lt 2/5=0.4$. Best $\eta^\star=2/7=0.2857$, rate $\rho^\star=3/7=0.4286$. Steps: $\ln(10^6)/\ln(7/3)=13.816/0.8473\approx16.3$, so $17$ steps.
C. One-parameter least squares: model $\hat y=wx$, data $(1,1),(2,3),(3,4)$, $f(w)=\tfrac12\sum(wx_i-y_i)^2$. Find the exact minimiser, then do three gradient-descent steps from $w=0$ with $\eta=0.05$. What is the largest stable $\eta$?
$f'(w)=\sum(wx_i-y_i)x_i=w\sum x_i^2-\sum x_iy_i=14w-19$ (since $\sum x^2=14$, $\sum xy=1+6+12=19$). Minimiser $w^\star=19/14=1.357$. Steps: $w_1=0-0.05(-19)=0.95$; $w_2=0.95-0.05(14\cdot0.95-19)=0.95+0.285=1.235$; $w_3=1.235-0.05(17.29-19)=1.3205$. The error shrinks by $1-0.05\cdot14=0.3$ per step ($1.357,\ 0.407,\ 0.122,\ 0.037$). The curvature is $L=14$, so $\eta\lt 2/14=0.143$ is stable, and $\eta=1/14=0.0714$ lands on $w^\star$ in one step.
D. A data set has $10^6$ examples. Compare batch sizes $B=100$ and $B=10{,}000$: updates per epoch, cost per update, and noise of the gradient estimate.
Updates per epoch: $10^6/100=10{,}000$ versus $10^6/10{,}000=100$. Cost per update: $100$ versus $10{,}000$ gradient evaluations (100 times more). Noise (standard deviation) $\propto1/\sqrt B$: the larger batch has $\sqrt{100}=10$ times less noise. Cost per epoch is the same ($10^6$ gradient evaluations): the larger batch buys smoother updates but gets 100 times fewer of them.
E. $f=3x^2$ ($\lambda=6$), $x_0=2$. Give $x_1,x_2,x_3$ and the behaviour for $\eta=0.25$, $\eta=1/3$ and $\eta=0.4$.
The multiplier is $m=1-6\eta$. $\eta=0.25$: $m=-0.5$, so $2,\,-1,\,0.5,\,-0.25$: overshoot but converges. $\eta=1/3$: $m=-1$, so $2,-2,2,-2$: endless bounce (this is $2/\lambda$). $\eta=0.4$: $m=-1.4$, so $2,\,-2.8,\,3.92,\,-5.49$: diverges.
F. Backtracking on $f(x)=3x^2$ at $x=1$ with $\alpha_0=1$, halving, and a large test constant $c=0.1$ (to make the test strict). Which $\alpha$ is accepted, and what is the new point?
$g=f'(1)=6$, $\|g\|^2=36$, $f=3$. The test is $f(1-6\alpha)\le3-0.1\cdot36\,\alpha=3-3.6\alpha$. $\alpha=1$: $x=-5$, $f=75\gt -0.6$: reject. $\alpha=0.5$: $x=-2$, $f=12\gt 1.2$: reject. $\alpha=0.25$: $x=-0.5$, $f=0.75\le3-0.9=2.1$: accept. New point $-0.5$, height $0.75$. (The perfect step is $1/f''=1/6$; backtracking got within a factor of 2 without knowing the curvature.)
Advanced Gradient-Based Optimization
Plain gradient descent works, but it is slow in long narrow valleys, it uses one step size for every number it is tuning, and it never changes its mind about how big a step to take. This chapter fixes all three. Momentum remembers where you were going. Adaptive methods (AdaGrad, RMSProp, Adam, AdamW) give every parameter its own step size. Schedules and warmup change the step size over time. Together they are how almost every modern neural network is trained.
- See exactly why plain gradient descent struggles: narrow valleys, mixed scales, one fixed step size
- Use the exponentially weighted average, the single tool behind momentum, RMSProp and Adam
- Derive and run momentum and Nesterov momentum, and know why they damp the zig-zag
- Derive AdaGrad, RMSProp, Adam (including the bias correction $1-\beta^t$) and AdamW (decoupled weight decay), with exact update rules and worked numbers
- Choose a learning-rate schedule (step, exponential, inverse-time, cosine) and understand warmup
- Race all the optimizers on the same landscapes, and say honestly when each one is a good choice
What we assume from Chapter 3.3. You know plain gradient descent, $\mathbf{x}_{t+1} = \mathbf{x}_t - \eta\,\nabla f(\mathbf{x}_t)$, where $\eta > 0$ is the learning rate (step size); you know that it diverges when $\eta$ is bigger than about $2/L$ (where $L$ is the largest curvature); and you have seen the zig-zag it makes in a long, narrow valley (an ill-conditioned problem). If any of that is fuzzy, revisit Chapter 3.3 first. The gradient is the vector of slopes that points uphill; curvature says how quickly that slope changes.
Notation used in this whole chapter. $\mathbf{x}_t$ are the parameters after $t$ steps and $\mathbf{g}_t = \nabla f(\mathbf{x}_t)$ is the gradient there. Whenever we write $\mathbf{g}^2$, $\sqrt{\mathbf{s}}$ or a division of two vectors, we mean it entry by entry. The letters $\mathbf{v}$ (momentum's velocity) and $\mathbf{m}, \mathbf{v}$ (Adam's two averages) are used in separate sections; each section says which one it means.
Why plain gradient descent is not enough core
Plain gradient descent is a hiker in thick fog. At every step the hiker feels the slope under their feet, takes a step of the same fixed size straight downhill, and then forgets everything. Three weaknesses follow from that.
- Narrow valleys. If the valley is steep on the sides and nearly flat along the bottom, the downhill direction points mostly at the opposite wall. The hiker bounces from wall to wall and creeps along the bottom.
- One step size for everything. A step that is safe for a steep direction is tiny for a flat one. There is no way to say "go bold here, go careful there".
- The step size never changes. Big steps are great at the start, when you are far away. They are bad at the end, when you want to settle into the bottom.
Every method in this chapter fixes one of these weaknesses by remembering something: the past directions, the past gradient sizes, or how far along training is.
Take the bowl $f(x, y) = \tfrac12 x^2 + 5y^2$. Its gradient is $(x,\ 10y)$: the slope in $y$ is ten times steeper than in $x$. Start at $(4, 1)$ with $\eta = 0.19$. This is almost the biggest safe step, because the steep direction has curvature $10$ and plain gradient descent needs $\eta \lt 2/10 = 0.2$.
- Each step multiplies $x$ by $1 - \eta\cdot 1 = 0.81$, and multiplies $y$ by $1 - \eta\cdot 10 = -0.9$.
- So $y$ goes $1,\ -0.9,\ 0.81,\ -0.729,\ \dots$ It flips sign every step and shrinks by only $10\%$ each time. That is the zig-zag.
- And $x$ goes $4,\ 3.24,\ 2.62,\ 2.13,\ \dots$ It shrinks by only $19\%$ per step.
- After 8 steps we are at $(0.74,\ 0.43)$. Both numbers are still far from $0$.
The steep direction forbids a bigger $\eta$, and the gentle direction needs a bigger $\eta$. A single number cannot satisfy both.
Recall the numbers from Chapter 3.3. For a bowl whose curvature ranges from $\mu$ (flattest direction) to $L$ (steepest direction), the condition number is $\kappa = L/\mu$. With the best fixed step $\eta = 2/(L+\mu)$, the distance to the minimum shrinks, at best, by the factor
$$\frac{\kappa - 1}{\kappa + 1} \quad\text{per step.}$$For $\kappa = 100$ that is $99/101 \approx 0.980$, so you need about $115$ steps just to make the error ten times smaller. The family tree of fixes for this chapter:
| Weakness of plain GD | Idea | Methods |
|---|---|---|
| Zig-zag across a valley, crawl along it | Remember the recent direction | Momentum, Nesterov |
| One step size for all coordinates | Divide each coordinate by its recent gradient size | AdaGrad, RMSProp, Adam |
| Weight decay behaves oddly inside Adam | Apply the decay separately from the gradient | AdamW |
| Big steps early, small steps late | Change $\eta$ with time | Schedules, warmup |
Why do we need it?
The loss of a real model is almost never a round bowl. It has directions that are thousands of times steeper than others. Plain gradient descent on such a loss is painfully slow, or unstable if you raise $\eta$.
Where is it used?
Every deep-learning training run: image classifiers (ResNets), language models (Transformers), recommendation models, and reinforcement-learning agents. Almost none of them use plain gradient descent.
How is it used?
You keep the gradient-descent loop and replace the single line "$\mathbf{x} \leftarrow \mathbf{x} - \eta\,\mathbf{g}$" with a better update rule. In a framework this is a one-word change of optimizer, plus choosing $\eta$ and a schedule.
Slow is not the same as wrong. Plain gradient descent still gets there, and for well-conditioned problems it is perfectly fine. The methods in this chapter earn their keep on ill-conditioned or noisy problems, which is what neural-network losses usually are.
Quick check: a bowl has curvatures $1$ and $50$. What is the largest $\eta$ plain GD can use, and by what factor does the flat direction shrink per step at that $\eta$?
$\eta$ must stay below $2/L = 2/50 = 0.04$. At $\eta = 0.04$ the flat direction is multiplied by $1 - 0.04\cdot 1 = 0.96$, so it loses only $4\%$ of its error per step. That is very slow, and it is the reason for this whole chapter.
The exponentially weighted average: a memory with fading core
How warm is it "these days"? You do not average the whole year. You also do not look only at today. You let yesterday count a bit less than today, the day before a bit less again, and so on, until old days fade away. That is an exponentially weighted average (EWA, also called an exponential moving average).
It is the single tool behind momentum, RMSProp and Adam. One number $\beta$ (between $0$ and $1$) says how long the memory is: $\beta$ near $0$ means "almost no memory", $\beta$ near $1$ means "a long memory".
A nice property: you never store the past. You keep one running number and update it with every new value.
The gradient of one coordinate over four steps was $g = 2,\ 4,\ 0,\ 6$. Use $\beta = 0.5$ and start the running number at $m_0 = 0$. The rule is: new $m$ = $0.5\cdot$ (old $m$) $+\ 0.5\cdot$ (new gradient).
- $m_1 = 0.5\cdot 0 + 0.5\cdot 2 = 1$
- $m_2 = 0.5\cdot 1 + 0.5\cdot 4 = 2.5$
- $m_3 = 0.5\cdot 2.5 + 0.5\cdot 0 = 1.25$
- $m_4 = 0.5\cdot 1.25 + 0.5\cdot 6 = 3.625$
With a longer memory, $\beta = 0.9$, the same rule gives $m = 0.2,\ 0.58,\ 0.522,\ 1.0698$. It moves much more slowly, and it is much smoother. Also notice that it starts too low: the average of the four gradients is $3$, but $m_4 = 1.07$. This is because we started the memory at $0$. Adam will fix this with a correction, in the Adam section.
Given a stream $g_1, g_2, \dots$ and a number $0 \le \beta \lt 1$, the exponentially weighted average is
$$m_t = \beta\, m_{t-1} + (1 - \beta)\, g_t, \qquad m_0 = 0.$$Unroll it to see the weights. Substitute $m_{t-1} = \beta m_{t-2} + (1-\beta)g_{t-1}$, then $m_{t-2}$, and so on:
$$m_t = (1-\beta)\big(g_t + \beta\, g_{t-1} + \beta^2 g_{t-2} + \dots + \beta^{t-1} g_1\big).$$So the gradient that is $k$ steps old gets the weight $(1-\beta)\beta^k$: each step back multiplies the weight by $\beta$. Three facts follow.
- The weights add up to $1 - \beta^t$. The sum $(1-\beta)(1 + \beta + \dots + \beta^{t-1}) = (1-\beta)\dfrac{1-\beta^t}{1-\beta} = 1 - \beta^t$. After many steps this is almost $1$, so $m_t$ really is an average. In the first few steps it is less than $1$, which is why $m_t$ starts low.
- Average age. The weight on a $k$-step-old value is $(1-\beta)\beta^k$, so the average age is $\sum_{k\ge 0} k(1-\beta)\beta^k = \beta/(1-\beta)$.
- Effective window $\approx 1/(1-\beta)$. After $1/(1-\beta)$ steps a weight has shrunk by $\beta^{1/(1-\beta)}$, which tends to $e^{-1} \approx 0.37$ as $\beta \to 1$. So "the last $1/(1-\beta)$ values" is a good picture of what is being averaged. This is a rule of thumb, not an exact window.
| $\beta$ | window $\approx 1/(1-\beta)$ | weight left after one window, $\beta^{1/(1-\beta)}$ |
|---|---|---|
| 0.5 | 2 steps | 0.25 |
| 0.9 | 10 steps | 0.349 |
| 0.99 | 100 steps | 0.366 |
| 0.999 | 1000 steps | 0.368 |
Why do we need it?
Raw gradients are noisy and jumpy. We want a steady summary of "what has the gradient been doing lately" without storing every past gradient, which would cost a lot of memory.
Where is it used?
Momentum (an average of gradients), RMSProp and Adam (averages of gradients and of squared gradients), batch-norm running statistics, smoothed loss curves in training dashboards, and "EMA of weights" used for evaluation in many deep-learning recipes.
How is it used?
Keep one running number per parameter. At each step do m = beta*m + (1-beta)*g. Choose $\beta$ so that $1/(1-\beta)$ is the number of steps you want to remember.
Two meanings of "average". A plain average gives every value the weight $1/t$, so after a million steps one new gradient changes it by $0.0001\%$. An exponentially weighted average gives the newest value a fixed weight $1-\beta$ forever. It keeps reacting. That is exactly what we want in training, where the gradient keeps changing as the parameters move.
It starts low. Because $m_0 = 0$, the first few values of $m_t$ are biased toward $0$ (the weights only add up to $1-\beta^t$). We will fix this with bias correction in the Adam section.
Quick check: with $\beta = 0.99$, roughly how many past gradients does the average "remember"? And what is the weight of the newest one?
About $1/(1-0.99) = 100$ gradients. The newest gradient has weight $1-\beta = 0.01$, the one before it $0.01\cdot 0.99 = 0.0099$, and so on, fading slowly.
Momentum: the heavy ball core
Imagine a heavy ball rolling inside the valley instead of a hiker who forgets every step. Two things change.
- The slope is a force, not a velocity. The ground pushes the ball a little each moment. The ball keeps the speed it already has (inertia) and adds the new push to it.
- A little friction slowly takes speed away, so the ball does not run forever.
Now watch the valley. The pushes across the valley alternate: left wall, right wall, left wall. They cancel out in the ball's speed. The pushes along the valley always point the same way, so they pile up and the ball speeds up. That is exactly what we wanted: damp the zig-zag, speed up the crawl.
Use the bowl from before, $f = \tfrac12 x^2 + 5y^2$ with gradient $(x, 10y)$, start $(4, 1)$, $\eta = 0.19$. Momentum keeps a velocity $\mathbf{v}$ (starting at $\mathbf{0}$). Each step: new velocity = $\beta\cdot$ old velocity $-\ \eta\cdot$ gradient, then move by the velocity. Take $\beta = 0.3$.
- Step 1: gradient $(4, 10)$. $\mathbf{v}_1 = 0.3\cdot(0,0) - 0.19\cdot(4,10) = (-0.76,\ -1.9)$. New position $(4,1) + \mathbf{v}_1 = (3.24,\ -0.9)$. (Same as plain GD: there is no past yet.)
- Step 2: gradient $(3.24,\ -9)$. $\mathbf{v}_2 = 0.3\cdot(-0.76, -1.9) - 0.19\cdot(3.24, -9) = (-0.228 - 0.6156,\ -0.57 + 1.71) = (-0.8436,\ 1.14)$. New position $(2.3964,\ 0.24)$.
- Step 3: gradient $(2.3964,\ 2.4)$. $\mathbf{v}_3 = 0.3\cdot(-0.8436, 1.14) - 0.19\cdot(2.3964, 2.4) = (-0.2531 - 0.4553,\ 0.342 - 0.456) = (-0.7084,\ -0.114)$. New position $(1.688,\ 0.126)$.
Compare after 3 steps. Plain GD: $(2.13,\ -0.73)$, with $y$ still big and flipping sign. Momentum: $(1.69,\ 0.13)$. The old velocity in $y$ ($-1.9$) and the new push ($+1.71$) nearly cancel, and the $x$-velocity keeps adding up. After 8 steps momentum is at $(0.22,\ 0.01)$ while GD is at $(0.74,\ 0.43)$.
The momentum (heavy-ball) update, with learning rate $\eta$ and momentum coefficient $\beta \in [0, 1)$ (often $0.9$), is
$$\mathbf{v}_{t+1} = \beta\,\mathbf{v}_t - \eta\,\nabla f(\mathbf{x}_t), \qquad \mathbf{x}_{t+1} = \mathbf{x}_t + \mathbf{v}_{t+1}, \qquad \mathbf{v}_0 = \mathbf{0}.$$Since the step taken is $\mathbf{x}_{t+1} - \mathbf{x}_t = \mathbf{v}_{t+1}$, this is the same as: new step = plain gradient step + $\beta\,\times$ the previous step,
$$\mathbf{x}_{t+1} = \mathbf{x}_t - \eta\,\mathbf{g}_t + \beta\,(\mathbf{x}_t - \mathbf{x}_{t-1}).$$It is an average of past gradients. Unroll $\mathbf{v}_t$ exactly as in the previous section: $\mathbf{v}_t = -\eta\,(\mathbf{g}_{t-1} + \beta\,\mathbf{g}_{t-2} + \beta^2\mathbf{g}_{t-3} + \dots)$. The weights $\beta^k$ add up to $1/(1-\beta)$, so
$$\mathbf{v}_t = -\frac{\eta}{1-\beta}\ \times\ (\text{an exponentially weighted average of past gradients}).$$Why it damps zig-zag and speeds up valleys. Look at one coordinate.
- Steady gradient $g$ (along the valley): the velocity settles where $v = \beta v - \eta g$, so $v = -\eta g/(1-\beta)$. With $\beta = 0.9$ that is a step 10 times larger than plain GD's $-\eta g$. This is the "effective window $1/(1-\beta)$" from the previous section.
- Alternating gradient $+g, -g, +g, \dots$ (across the valley): the velocity settles into a see-saw of size only $\eta g/(1+\beta)$. With $\beta = 0.9$ that is $0.53\,\eta g$, smaller than plain GD's $\eta g$.
- So the ratio "useful steady step : useless see-saw" improves from $1$ to $\dfrac{1/(1-\beta)}{1/(1+\beta)} = \dfrac{1+\beta}{1-\beta}$, which is $19$ for $\beta = 0.9$.
How fast, exactly (quadratics). In one direction with curvature $h$ the gradient is $hx$ and the update becomes $x_{t+1} = (1 + \beta - \eta h)\,x_t - \beta\,x_{t-1}$. Its characteristic equation is $\lambda^2 - (1+\beta-\eta h)\lambda + \beta = 0$, and the error shrinks like $\rho^t$ where $\rho$ is the larger $|\lambda|$. Two facts follow.
- Stability: it converges if and only if $0 \lt \eta h \lt 2(1+\beta)$. Plain GD is the case $\beta=0$, limit $2$. So momentum even allows a larger $\eta h$.
- Exactly $\sqrt{\beta}$ in a whole band: the two roots are complex whenever $(1-\sqrt\beta)^2 \le \eta h \le (1+\sqrt\beta)^2$, and then $|\lambda| = \sqrt{\beta}$ (the product of the roots is $\beta$), no matter what $h$ is. Pick $\eta$ and $\beta$ so that every curvature between $\mu$ and $L$ falls inside that band: $\beta = \Big(\dfrac{\sqrt\kappa - 1}{\sqrt\kappa + 1}\Big)^2$ and $\eta = \dfrac{4}{(\sqrt L + \sqrt\mu)^2}$. Then the rate is $\sqrt\beta = \dfrac{\sqrt\kappa - 1}{\sqrt\kappa + 1}$ (Polyak's result for quadratics).
For $\kappa = 100$: plain GD at its best step shrinks the error by $0.980$ per step, tuned momentum by $0.818$. The number of steps needed grows like $\sqrt\kappa$ instead of $\kappa$. (Quick numerical check, error in $f$ down to $10^{-12}$ on a bowl with $\kappa = 100$: GD takes 824 steps, tuned momentum 109.) These formulas are exact for quadratics; for general losses they are a guide, not a guarantee.
Why do we need it?
In a long narrow valley, plain GD must use a small $\eta$ (to avoid bouncing out) and then crawls. Momentum lets the gentle direction build up speed while the bouncing cancels itself.
Where is it used?
SGD with momentum is the classic way to train image classifiers (ResNet-style networks) and is built into every framework (torch.optim.SGD(..., momentum=0.9)). It is also the first-moment half of Adam.
How is it used?
Keep one velocity vector the same size as the parameters. Each step update the velocity, then add it to the parameters. Usual choice: $\beta = 0.9$. Because the effective step is $\eta/(1-\beta)$, lower $\eta$ when you add momentum.
Momentum can overshoot. The ball does not stop at the bottom: it rolls through, climbs the other side, and comes back. With $\beta$ too high you get long, slow oscillations. A middle $\beta$ (between $0.5$ and $0.9$ is typical) trades speed against overshoot.
Two ways of writing it. Our form $\mathbf{v} \leftarrow \beta\mathbf{v} - \eta\mathbf{g}$ puts $\eta$ inside the velocity. Deep-learning libraries (for example PyTorch) store $\mathbf{b} \leftarrow \beta\mathbf{b} + \mathbf{g}$ and then do $\mathbf{x} \leftarrow \mathbf{x} - \eta\mathbf{b}$. With a constant $\eta$ the two give the same iterates. They differ only if you change $\eta$ while training (as schedules do), because the old velocity then carries the old or the new $\eta$.
Quick check: the gradient is the constant $2$ for a long time. With $\eta = 0.1$ and $\beta = 0.9$, how fast does the ball end up moving per step, and how does that compare with plain GD?
The terminal velocity solves $v = 0.9v - 0.1\cdot 2$, so $v = -0.2/0.1 = -2$. The ball moves $2$ per step. Plain GD moves $\eta g = 0.2$ per step. That is $1/(1-\beta) = 10$ times faster.
Nesterov momentum: look before you push core
Ordinary momentum measures the slope where the ball is, and then adds that push to the speed it already has. But we already know that, because of its speed, the ball is about to move by roughly $\beta\mathbf{v}$. So why measure the slope at the old spot?
Nesterov momentum jumps ahead first by $\beta\mathbf{v}$, and then measures the slope there. It is like a driver who looks at the road ahead instead of at the bumper. If the ball is about to shoot past the bottom, the slope at the look-ahead point already points back uphill, so the correction (the brake) comes one step earlier.
One-dimensional bowl $f(x) = \tfrac12 x^2$ (gradient $= x$). Start at $x_0 = 5$, $\eta = 0.2$, $\beta = 0.8$, velocity $0$. The minimum is at $x = 0$.
- Momentum (gradient at the current $x$). $v_1 = -0.2\cdot 5 = -1$, $x_1 = 4$. $v_2 = 0.8\cdot(-1) - 0.2\cdot 4 = -1.6$, $x_2 = 2.4$. $v_3 = 0.8\cdot(-1.6) - 0.2\cdot 2.4 = -1.76$, $x_3 = 0.64$. $v_4 = 0.8\cdot(-1.76) - 0.2\cdot 0.64 = -1.536$, $x_4 = -0.896$. The gradient used in step 4 was $0.64$, still pushing further down, although the ball is about to shoot past $0$.
- Nesterov (gradient at the look-ahead $x + \beta v$). Step 1: look-ahead $5 + 0.8\cdot 0 = 5$, so $v_1 = -1$, $x_1 = 4$. Step 2: look-ahead $4 + 0.8\cdot(-1) = 3.2$, so $v_2 = -0.8 - 0.2\cdot 3.2 = -1.44$, $x_2 = 2.56$. Step 3: look-ahead $2.56 - 1.152 = 1.408$, so $v_3 = -1.152 - 0.2816 = -1.4336$, $x_3 = 1.1264$. Step 4: look-ahead $1.1264 - 1.1469 = -0.0205$, almost exactly the bottom, so the gradient there is about $0$ and $v_4 = -1.1469 + 0.0041 = -1.1428$, $x_4 = -0.0164$.
- Over the first 6 steps, the furthest the ball overshoots is $-2.40$ for momentum and $-1.06$ for Nesterov.
The Nesterov momentum update (the "classical" teaching form) is
$$\tilde{\mathbf{x}}_t = \mathbf{x}_t + \beta\,\mathbf{v}_t, \qquad \mathbf{v}_{t+1} = \beta\,\mathbf{v}_t - \eta\,\nabla f(\tilde{\mathbf{x}}_t), \qquad \mathbf{x}_{t+1} = \mathbf{x}_t + \mathbf{v}_{t+1}.$$The only difference from momentum is where the gradient is measured: at the look-ahead point $\tilde{\mathbf{x}}_t$ instead of $\mathbf{x}_t$.
Equivalent two-line form. Put $\mathbf{y}_t = \mathbf{x}_t + \beta\mathbf{v}_t$. Because $\mathbf{x}_{t+1} - \mathbf{x}_t = \mathbf{v}_{t+1}$, one checks that
$$\mathbf{x}_{t+1} = \mathbf{y}_t - \eta\,\nabla f(\mathbf{y}_t), \qquad \mathbf{y}_{t+1} = \mathbf{x}_{t+1} + \beta\,(\mathbf{x}_{t+1} - \mathbf{x}_t).$$In words: take a plain gradient step from the look-ahead point, then extrapolate along the direction you just moved.
Why it brakes earlier. Use the first-order Taylor expansion (linearization, with Hessian $H$): $\nabla f(\mathbf{x} + \beta\mathbf{v}) \approx \nabla f(\mathbf{x}) + \beta H\mathbf{v}$. Put that into the update:
$$\mathbf{v}_{t+1} \approx \beta\,\mathbf{v}_t - \eta\,\mathbf{g}_t \;-\; \eta\beta\,H\mathbf{v}_t.$$The first two terms are exactly momentum. The extra term $-\eta\beta H\mathbf{v}_t$ opposes the velocity, and more strongly where the curvature $H$ is large. It is friction that is switched on exactly in the steep directions, where overshoot happens. (For $f = \tfrac12 h x^2$ this is exact: $v_{t+1} = \beta(1 - \eta h)\,v_t - \eta h\,x_t$.)
What is proved. For a smooth convex function (gradient Lipschitz with constant $L$), the two-line form with step $\eta = 1/L$ and the special momentum weights $\beta_t = (t-1)/(t+2)$ gives an error that shrinks like $O(1/t^2)$, versus $O(1/t)$ for plain gradient descent. Up to constant factors, no method that only uses gradients can guarantee a better rate on that class of functions (a lower bound by Nemirovski and Yudin). With a fixed $\beta$ such as $0.9$ in deep learning, and non-convex losses, no such guarantee exists: treat it as a good heuristic. (You will meet rates in Chapter 3.5.)
Why do we need it?
Plain momentum reacts to an overshoot one step too late, because it looks at the slope behind it. Looking ahead gives a cheaper, earlier correction and often a smoother path.
Where is it used?
SGD with Nesterov momentum in image models (nesterov=True in the framework), Nadam (Adam with a Nesterov look-ahead), and accelerated methods for convex problems such as FISTA (Chapter 3.13).
How is it used?
Same cost and memory as momentum: one velocity vector. You just evaluate the gradient at $\mathbf{x} + \beta\mathbf{v}$. Libraries re-write the update so that the gradient is taken at the stored point (PyTorch keeps $\mathbf{b} \leftarrow \beta\mathbf{b} + \mathbf{g}$ and $\mathbf{x} \leftarrow \mathbf{x} - \eta(\mathbf{g} + \beta\mathbf{b})$). Their stored parameters are our look-ahead point $\mathbf{x} + \beta\mathbf{v}$, so the numbers differ from ours but the path is the same.
It is a small, not a magic, improvement. In deep learning, Nesterov and ordinary momentum often perform similarly, and the choice is usually settled by trying both. The clearest gains are in smooth, convex problems, where the $O(1/t^2)$ result applies.
Stability needs a smaller $\eta$. On a bowl with curvature $h$, heavy-ball momentum is stable for $\eta h \lt 2(1+\beta)$, but Nesterov only for $\eta h \lt 2(1+\beta)/(1+2\beta)$ (you can check this with the same two-root argument). As $\beta \to 1$ that limit tends to $4/3$, so $\eta \le 1/L$ is always safe. Example: on a bowl with $\kappa = 100$, take the heavy-ball-tuned $\beta = \big(\tfrac{\sqrt\kappa-1}{\sqrt\kappa+1}\big)^2$ and $\eta = 4/(\sqrt L + \sqrt\mu)^2$ from above. Heavy-ball momentum converges, but Nesterov diverges with the same two numbers (here $\eta L \approx 3.3$, above Nesterov's limit of about $1.4$).
Quick check: at the very first step the velocity is $0$. What is the difference between momentum and Nesterov at step 1?
None. The look-ahead point is $\mathbf{x}_0 + \beta\cdot\mathbf{0} = \mathbf{x}_0$, so both measure the gradient at $\mathbf{x}_0$. The two methods only start to differ from step 2 on, once there is a velocity to look ahead with.
Adaptive learning rates: one step size per parameter core
Think of a mixing desk with many sliders. Some sliders are touchy: a tiny nudge changes the sound a lot. Others are sluggish: you must push them far. Turning every slider by the same amount is a bad plan. You want gentle nudges for the touchy ones and big pushes for the sluggish ones.
The parameters of a model are those sliders. The loss may be very sensitive to one parameter and hardly sensitive to another. Adaptive learning rates give every parameter its own step size. The rule they use is simple and cheap:
If a parameter's gradients have recently been big, take smaller steps for it. If they have been small, take bigger steps.
The bowl $f(x, y) = \tfrac12(100x^2 + y^2)$ has gradient $(100x,\ y)$. At the point $(1, 1)$ the gradient is $(100,\ 1)$. Plain GD can only use $\eta \lt 2/100 = 0.02$, so take $\eta = 0.019$.
- Plain GD moves by $-\eta\mathbf{g} = (-1.9,\ -0.019)$. So $x$ jumps from $1$ to $-0.9$ (it overshoots), while $y$ moves from $1$ to $0.981$: almost nothing. To bring $y$ down to $1\%$ takes about $\ln 0.01 / \ln 0.981 \approx 240$ steps.
- Now give each coordinate its own step size: divide the step in $x$ by $100$ and leave $y$ alone. With $\eta = 0.5$ the move is $-0.5\cdot(100/100,\ 1/1) = (-0.5,\ -0.5)$. Both coordinates are halved in every step, and after 7 steps both are about $1/128$ of where they began.
The only difference is dividing each coordinate's step by "how touchy it is".
An adaptive method replaces the single number $\eta$ by one number per coordinate. Let $\mathbf{s}_t$ be a running measure of the squared gradient of each coordinate (a different recipe for $\mathbf{s}_t$ gives a different method). The update is
$$\mathbf{x}_{t+1} = \mathbf{x}_t - \eta\ \frac{\mathbf{g}_t}{\sqrt{\mathbf{s}_t} + \varepsilon},$$where $\varepsilon$ is a tiny number (such as $10^{-8}$) that avoids dividing by zero. The preconditioning view. This is gradient descent with a diagonal rescaling:
$$\mathbf{x}_{t+1} = \mathbf{x}_t - \eta\,P_t^{-1}\,\mathbf{g}_t, \qquad P_t = \mathrm{diag}\big(\sqrt{\mathbf{s}_t} + \varepsilon\big).$$The matrix $P_t$ is called a preconditioner. It reshapes the problem so that the valley looks rounder. Newton's method (Chapter 3.12) is the same idea with the best possible preconditioner, the full Hessian $H$ instead of a diagonal. Here is how the three famous methods choose $\mathbf{s}_t$:
| Method | $\mathbf{s}_t$ (gradient "energy" per coordinate) | Memory of the past |
|---|---|---|
| AdaGrad | $\sum_{i \le t} \mathbf{g}_i^2$ (a sum, only grows) | never forgets |
| RMSProp | $\beta\,\mathbf{s}_{t-1} + (1-\beta)\,\mathbf{g}_t^2$ (an exponentially weighted average) | about $1/(1-\beta)$ steps |
| Adam | like RMSProp (with $\beta_2$), plus momentum on the gradient and a start-up correction | about $1/(1-\beta_2)$ steps |
Two insights.
- The step has a fixed size in parameter units. $g/\sqrt{s}$ has no units: if you rescale the loss by $1000$, the gradients and $\sqrt{s}$ both grow $1000$-fold and the step does not change. That is why $\eta$ becomes "roughly how far a parameter moves per step", whatever the loss scale is. With no averaging at all ($s = g^2$) the step is $\eta\,g/|g| = \eta\,\mathrm{sign}(g)$: sign descent.
- It is a diagonal fix only. Dividing by gradient size is a crude stand-in for dividing by curvature, and a diagonal $P$ can only fix a valley whose axes line up with the coordinate axes. In a rotated valley it helps much less (see the widget below).
Why do we need it?
The parameters of a model live on very different scales (weights of different layers, rare versus common features, embeddings versus biases). One global $\eta$ is always a compromise that is too big for some and too small for others.
Where is it used?
Almost every Transformer, vision-transformer and diffusion-model training run uses Adam or AdamW. Sparse models (text classifiers on word counts, click-prediction, embeddings) were an early home for AdaGrad.
How is it used?
Keep a second number per parameter (the running squared gradient), and divide each gradient by its square root before stepping. You still choose one global $\eta$, but it now means "typical step per parameter". The three methods that follow differ only in how they average the squares.
Adaptive does not mean "always faster". A well-tuned momentum method often matches or beats an adaptive one on a smooth, well-scaled problem. Adaptive methods shine when the parameters have very different gradient scales, or when gradients are sparse or noisy.
It measures gradient size, not curvature. In stochastic training, $\sqrt{s}$ is largely the noise level of a parameter's gradient. The method then moves noisy parameters less. That is often useful, but it is not the same as Newton's method.
Quick check: with no averaging ($\mathbf{s} = \mathbf{g}^2$), what is the step for the gradient $(100,\ 1)$ and learning rate $\eta = 0.1$?
$\eta\,g/\sqrt{g^2} = \eta\,\mathrm{sign}(g)$ in each coordinate, so the step is $-(0.1,\ 0.1)$. The huge difference in gradient size has disappeared: both coordinates move by exactly $\eta$.
AdaGrad: shrink the step by the total gradient seen so far core
Give every parameter a budget. Each time its gradient is large, it spends some budget (the square of the gradient is added to a running total, and the total never goes down). The more a parameter has already been pushed around, the smaller its future steps.
The nice consequence is for rare parameters. Imagine a word that appears in 1 of every 1000 sentences. Its parameter gets a gradient only once in a while, so it has spent almost no budget, and when its gradient finally arrives the step is big. A common word's parameter has spent a lot and takes small steps. Rare things learn fast when they show up; common things settle down.
The price: the budget only grows, so every step size shrinks forever.
Two coordinates, $\eta = 0.1$. Coordinate A gets gradient $1$ at every step (a common feature). Coordinate B gets gradient $1$ only at steps 2 and 4 (a rare one). The rule: $G \leftarrow G + g^2$, then step $= \eta\, g / \sqrt{G}$.
| step $t$ | $g_A$ | $G_A$ | step of A $=0.1/\sqrt{G_A}$ | $g_B$ | $G_B$ | step of B |
|---|---|---|---|---|---|---|
| 1 | 1 | 1 | 0.1000 | 0 | 0 | 0 |
| 2 | 1 | 2 | 0.0707 | 1 | 1 | 0.1000 |
| 3 | 1 | 3 | 0.0577 | 0 | 1 | 0 |
| 4 | 1 | 4 | 0.0500 | 1 | 2 | 0.0707 |
B's first appearance (step 2) gets the full step $0.1$, bigger than A's step in the same round ($0.0707$). A's steps shrink as $0.1,\ 0.0707,\ 0.0577,\ 0.05, \dots$, that is $0.1/\sqrt{t}$.
AdaGrad (Duchi, Hazan and Singer, 2011), per coordinate:
$$\mathbf{G}_t = \mathbf{G}_{t-1} + \mathbf{g}_t^2, \qquad \mathbf{x}_{t+1} = \mathbf{x}_t - \frac{\eta}{\sqrt{\mathbf{G}_t} + \varepsilon}\ \mathbf{g}_t, \qquad \mathbf{G}_0 = \mathbf{0}.$$The squares and the root are taken entry by entry; $\varepsilon \approx 10^{-8}$ protects against $0/0$. This is the adaptive update with $\mathbf{s}_t = \mathbf{G}_t$ (the sum).
- The first step is $\eta$ in every coordinate. At $t = 1$, $G = g^2$, so the step is $\eta\,g/|g| = \eta\,\mathrm{sign}(g)$, whatever the gradient size.
- The learning rate decays. If a coordinate's gradient has about the same size $c$ each step, then $G_t \approx t\,c^2$, so its effective rate is $\eta/(c\sqrt{t})$ and its step is about $\eta/\sqrt{t}$. After $10{,}000$ steps the step is $100$ times smaller than at the start. The steps still add up to infinity (the sum of $1/\sqrt{t}$ diverges), so the iterate can in principle travel as far as needed, but slowly.
- Known theory. For convex problems AdaGrad has a proven guarantee (a regret bound) that is competitive with the best per-coordinate step sizes you could have picked in hindsight, and it is especially good when the gradients are sparse. (This is a statement about convex problems; deep non-convex losses are different.)
Why do we need it?
With one global $\eta$, a feature that appears rarely barely learns, because its gradient is zero most of the time. AdaGrad lets rare features take big steps when they do appear, with no hand-tuning per feature.
Where is it used?
Sparse, high-dimensional models: bag-of-words text classifiers, click-through-rate prediction, learning word embeddings (it was used in early GloVe training). It is also the ancestor of RMSProp, Adam and many optimizers in recommendation systems.
How is it used?
Keep a running sum of squared gradients per parameter; divide each gradient by the square root of its sum. A typical $\eta$ is $0.01$ to $1$ depending on the problem (it is not scale-free for the loss). It is best for convex or sparse problems and short runs.
The learning rate dies. Because $\mathbf{G}$ never shrinks, steps keep getting smaller, even if the problem has changed or you have not yet arrived. In a long deep-learning run AdaGrad's steps can become so small that learning stalls. That is exactly what RMSProp fixes next.
Feature scaling is the other cure. In the widget, the problem came from one feature being ten times bigger than another. Standardising the features (as in classical machine learning) removes much of the need for AdaGrad's per-coordinate scaling.
Quick check: a coordinate has received the gradients $3, 4$ so far. What is $G$, and what is its next step if the next gradient is $0$ and $\eta = 1$?
$G = 3^2 + 4^2 = 25$. With gradient $0$ the step is $\eta\cdot 0/\sqrt{25} = 0$: AdaGrad only moves a coordinate when it receives a gradient. If instead the next gradient were $12$, then $G = 25 + 144 = 169$ and the step would be $12/13 \approx 0.92$.
RMSProp: forget old gradients, so the step size does not die core
AdaGrad's total only grows, so it remembers a huge gradient from step 5 for the whole rest of training. RMSProp keeps the same idea but uses a fading memory (the exponentially weighted average from earlier in the chapter) instead of a total. Old gradients fade out, so if the gradients become small again, the step size recovers.
Read the name as Root Mean Square: we divide the gradient by the square root of the (recent) mean of its squares. In plain words: "divide the gradient by its recent typical size". A gradient that is normal for this parameter gives a step of about $\eta$.
One coordinate, constant gradient $g = 2$, $\eta = 0.01$, $\beta = 0.9$. Rule: $s \leftarrow 0.9\,s + 0.1\,g^2$ and step $= \eta\,g/\sqrt{s}$ (we ignore the tiny $\varepsilon$).
- $s_1 = 0.9\cdot 0 + 0.1\cdot 4 = 0.4$. Step $= 0.01\cdot 2/\sqrt{0.4} = 0.02/0.6325 = 0.03162$. This is $3.16\,\eta$.
- $s_2 = 0.9\cdot 0.4 + 0.1\cdot 4 = 0.76$. Step $= 0.02/\sqrt{0.76} = 0.02/0.8718 = 0.02294$, which is $2.29\,\eta$.
- $s_t = (1 - 0.9^t)\cdot 4$, so after a few dozen steps $s \to 4$ and the step settles at $0.01\cdot 2/2 = 0.01 = \eta$.
If the gradient were $20$ instead of $2$, every number above would be the same: the step does not depend on the size of the gradient, only on how it compares with its own history. Compare AdaGrad on the same stream: after $t$ steps its step is $\eta/\sqrt{t}$, which keeps falling.
RMSProp (proposed by Geoffrey Hinton in a 2012 lecture; there is no original paper), per coordinate:
$$\mathbf{s}_t = \beta\,\mathbf{s}_{t-1} + (1-\beta)\,\mathbf{g}_t^2, \qquad \mathbf{x}_{t+1} = \mathbf{x}_t - \frac{\eta}{\sqrt{\mathbf{s}_t} + \varepsilon}\ \mathbf{g}_t, \qquad \mathbf{s}_0 = \mathbf{0}.$$Typical values: $\beta = 0.9$ (or $0.99$), $\varepsilon = 10^{-8}$, and a small $\eta$ such as $0.001$. $\mathbf{s}_t$ is exactly the exponentially weighted average of $\mathbf{g}^2$ from the earlier section, so $\sqrt{\mathbf{s}_t}$ is the recent root mean square of the gradient, over a window of about $1/(1-\beta)$ steps.
- It fixes AdaGrad's decay. With a constant-size gradient, $\mathbf{s}_t$ settles at $g^2$, so the step stays near $\eta$ forever instead of shrinking like $1/\sqrt{t}$.
- It starts with a big step. Because $\mathbf{s}_0 = \mathbf{0}$, after $t$ steps with a constant gradient $\mathbf{s}_t = (1-\beta^t)\,g^2$, and the step multiplier is $1/\sqrt{1-\beta^t}$. At $t=1$ that is $1/\sqrt{1-\beta}$: $3.16$ for $\beta = 0.9$ and $10$ for $\beta = 0.99$. This start-up bias is exactly what Adam's correction removes.
- It adapts to a change. If the gradients suddenly become 10 times smaller, $\mathbf{s}_t$ follows within about $1/(1-\beta)$ steps and the step size returns to about $\eta$.
Why do we need it?
AdaGrad is great at first, but in a long run its steps shrink to nothing. We want the "per-parameter scaling" without the "learning rate dies" problem, so training can keep going.
Where is it used?
Training recurrent networks (it was popular before Adam), reinforcement-learning agents (the original Atari DQN used it), and as the second-moment half of Adam. It is available in every framework.
How is it used?
Keep a running average of the squared gradient for each parameter ($\beta \approx 0.9$) and divide each gradient by its square root. Use a small $\eta$ (for example $10^{-3}$ to $10^{-4}$), because the step is about $\eta$ per parameter whatever the gradient size.
No magic constant protects you from a bad start. The first RMSProp step is $1/\sqrt{1-\beta}$ times bigger than $\eta$ (ten times bigger at $\beta = 0.99$). On a sharp landscape this alone can throw the parameters far away. That is one of the reasons for learning-rate warmup later in this chapter.
Details differ between libraries. Some put $\varepsilon$ inside the square root, $\sqrt{\mathbf{s}_t + \varepsilon}$, and some add a momentum term or "centre" the average (subtract the mean gradient). The idea is the same; for exact reproduction check the library's formula.
Quick check: with $\beta = 0.99$ and a steady gradient, how big is the very first RMSProp step compared with $\eta$, and how long until the memory has "settled"?
The first step is $1/\sqrt{1-0.99} = 10$ times $\eta$. The memory window is $1/(1-0.99) = 100$ steps, so it takes a few hundred steps before $\mathbf{s}$ has stopped being biased low (for example $1/\sqrt{1-0.99^{100}} = 1.26$ at step 100).
Adam: momentum for the direction, RMSProp for the size core
Adam ("adaptive moment estimation") combines the two ideas you now know:
- Steer with a smoothed gradient: the exponentially weighted average of past gradients (the first moment, $\mathbf{m}$). This is momentum's smoothing.
- Scale each coordinate by its recent gradient size: the square root of the exponentially weighted average of squared gradients (the second moment, $\mathbf{v}$). This is RMSProp's scaling.
There is one extra trick. Both averages start at $0$, so in the first steps they are too small. Adam corrects that start-up bias by dividing by a number that tells it how much of the average has been "filled in" so far.
A useful way to read the result: $\hat m/\sqrt{\hat v}$ is "average gradient divided by typical gradient size". If a parameter's gradient keeps the same sign, this is near $\pm 1$ and the parameter moves by about $\eta$. If the gradient keeps flipping sign (noise), the average is near $0$ and the parameter hardly moves.
Bowl $f = \tfrac12 x^2 + 5y^2$, gradient $(x, 10y)$, start $(4, 1)$. Use $\eta = 0.1$ and the defaults $\beta_1 = 0.9$, $\beta_2 = 0.999$ (we ignore $\varepsilon$).
- Step 1. $\mathbf{g}_1 = (4, 10)$. First moment: $\mathbf{m}_1 = 0.1\cdot(4, 10) = (0.4,\ 1.0)$. Second moment: $\mathbf{v}_1 = 0.001\cdot(16, 100) = (0.016,\ 0.1)$.
- Correct the start-up bias: divide by $1 - 0.9^1 = 0.1$ and $1 - 0.999^1 = 0.001$. Then $\hat{\mathbf{m}}_1 = (4, 10)$ and $\hat{\mathbf{v}}_1 = (16, 100)$. Now $\sqrt{\hat{\mathbf{v}}_1} = (4, 10)$.
- Update: $\mathbf{x}_2 = (4, 1) - 0.1\cdot(4/4,\ 10/10) = (3.9,\ 0.9)$. Both coordinates moved by exactly $\eta = 0.1$, although the gradients were $4$ and $10$.
- Step 2. $\mathbf{g}_2 = (3.9, 9)$. $\mathbf{m}_2 = 0.9\cdot(0.4, 1) + 0.1\cdot(3.9, 9) = (0.75,\ 1.8)$. $\mathbf{v}_2 = 0.999\cdot(0.016, 0.1) + 0.001\cdot(15.21, 81) = (0.031194,\ 0.1809)$. Divide by $1 - 0.9^2 = 0.19$ and $1 - 0.999^2 = 0.001999$: $\hat{\mathbf{m}}_2 = (3.9474,\ 9.4737)$, $\hat{\mathbf{v}}_2 = (15.605,\ 90.495)$. Steps: $0.1\cdot 3.9474/3.9503 = 0.0999$ and $0.1\cdot 9.4737/9.5129 = 0.0996$. So $\mathbf{x}_3 = (3.8001,\ 0.8004)$.
- Without the correction, step 1 would have been $0.1\cdot 0.4/\sqrt{0.016} = 0.316$ in each coordinate, $3.16$ times too big.
Adam (Kingma and Ba, 2015). At step $t = 1, 2, \dots$ with gradient $\mathbf{g}_t$ (all operations entry by entry):
$$\begin{aligned} \mathbf{m}_t &= \beta_1\,\mathbf{m}_{t-1} + (1-\beta_1)\,\mathbf{g}_t &&\text{(first moment: average gradient)}\\ \mathbf{v}_t &= \beta_2\,\mathbf{v}_{t-1} + (1-\beta_2)\,\mathbf{g}_t^2 &&\text{(second moment: average squared gradient)}\\ \hat{\mathbf{m}}_t &= \mathbf{m}_t/(1-\beta_1^t), \qquad \hat{\mathbf{v}}_t = \mathbf{v}_t/(1-\beta_2^t) &&\text{(bias correction)}\\ \mathbf{x}_{t+1} &= \mathbf{x}_t - \eta\,\hat{\mathbf{m}}_t\big/\big(\sqrt{\hat{\mathbf{v}}_t} + \varepsilon\big) \end{aligned}$$with $\mathbf{m}_0 = \mathbf{v}_0 = \mathbf{0}$. Defaults: $\beta_1 = 0.9$ (memory of about $10$ steps), $\beta_2 = 0.999$ (about $1000$ steps), $\varepsilon = 10^{-8}$; the paper suggests $\eta = 0.001$. (Careful: here $\mathbf{v}$ is the second moment, not the velocity of momentum.)
Where does $1-\beta^t$ come from? Unroll the first moment as in the exponentially-weighted-average section: $m_t = (1-\beta_1)\sum_{i=1}^{t}\beta_1^{\,t-i}\,g_i$. Suppose, for the sake of the argument, that every gradient has the same expected value $\mu = \mathbb{E}[g_i]$. Then
$$\mathbb{E}[m_t] = \mu\,(1-\beta_1)\sum_{i=1}^{t}\beta_1^{\,t-i} = \mu\,(1-\beta_1)\,\frac{1-\beta_1^t}{1-\beta_1} = \mu\,\big(1 - \beta_1^t\big).$$So $m_t$ is too small by exactly the factor $1-\beta_1^t$, and dividing by it gives an average with the right expected value: $\mathbb{E}[\hat m_t] = \mu$. The same calculation for $g^2$ gives $\mathbb{E}[v_t] = (1-\beta_2^t)\,\mathbb{E}[g^2]$, hence the second correction. (If the true mean drifts, the correction is no longer exact, but the error is small because old terms carry little weight.) Notice that $1 - \beta^t$ is exactly the sum of the weights we found earlier.
Consequences.
- The first step has size $\approx\eta$ in every coordinate. At $t=1$: $\hat m_1 = g_1$ and $\hat v_1 = g_1^2$, so the step is $\eta\,g_1/(|g_1| + \varepsilon) \approx \eta\,\mathrm{sign}(g_1)$.
- $\eta$ is roughly a per-step distance. Since $|\hat m| \lesssim \sqrt{\hat v}$ in most situations, each parameter moves at most about $\eta$ per step (in the rare extreme of one huge gradient after a long calm period, up to about $\eta(1-\beta_1)/\sqrt{1-\beta_2} \approx 3.2\,\eta$). That is a kind of built-in trust region.
- Why $\beta_2$ is so close to $1$. $\hat v$ needs many samples to be reliable (the warmup section shows how poor it is early on), so it gets a long memory. $\hat m$ should react faster, so $\beta_1$ is smaller.
Why do we need it?
Momentum alone fixes zig-zag but not mixed scales; RMSProp alone fixes scales but is jumpy and starts too big. Adam combines both and corrects the start, giving a method that works reasonably well with little tuning.
Where is it used?
It is the default first choice for training Transformers and language models, GANs, diffusion models, graph networks and most research code, usually in its AdamW form (next section).
How is it used?
Create Adam(params, lr=1e-3) and keep the other defaults ($\beta_1 = 0.9$, $\beta_2 = 0.999$, $\varepsilon = 10^{-8}$). Tune $\eta$ first (try a few powers of 10), then add a schedule and warmup. It stores two extra numbers per parameter.
Adam is not guaranteed to converge. There are simple convex counter-examples where Adam with a constant $\beta_2$ does not converge (Reddi, Kale and Kumar, 2018). Variants such as AMSGrad fix this in theory. In practice Adam trains an enormous range of models well, so treat it as a very good heuristic, not a theorem.
Small variants. The paper's own pseudo-code ends with an equivalent form that folds the two corrections into the learning rate, $\eta_t = \eta\sqrt{1-\beta_2^t}/(1-\beta_1^t)$, and applies $\varepsilon$ to the uncorrected $\sqrt{\mathbf{v}}$. Some implementations do it that way, others follow our form. The two are the same except for a slightly different effect of $\varepsilon$.
$\varepsilon$ matters sometimes. If a parameter's gradients are tiny (so $\sqrt{\hat v}$ is smaller than $\varepsilon$), then $\sqrt{\hat v} + \varepsilon \approx \varepsilon$ and the step is about $\eta\,\hat m/\varepsilon$: not $\eta$ but much smaller. So $\varepsilon$ sets a floor below which Adam stops treating a gradient as "normal size". Some recipes use a larger $\varepsilon$ (such as $10^{-6}$) for stability.
Quick check: Adam, $\eta = 0.01$, $\beta_1 = 0.9$, $\beta_2 = 0.999$. The first gradient of a coordinate is $-250$. What is the first step, with and without bias correction?
With correction: $\hat m_1 = -250$, $\hat v_1 = 62500$, so the step is $0.01\cdot(-250)/250 = -0.01$, and the coordinate increases by $0.01 = \eta$ (we subtract the step). Without correction: $m_1 = -25$, $v_1 = 62.5$, so the step would be $0.01\cdot(-25)/\sqrt{62.5} = -0.0316$, which is $3.16$ times larger. The size of the gradient ($250$) did not matter at all.
AdamW: weight decay done right core
Weight decay means: at every step, shrink each weight a tiny bit toward zero. It stops weights from growing huge and is a standard way to regularize a model (you will study regularization in Chapter 3.11).
There are two ways to build it in:
- L2 penalty. Add $\tfrac\lambda2\|\mathbf{x}\|^2$ to the loss. Its gradient is $\lambda\mathbf{x}$, so the optimizer sees "gradient $+\ \lambda\mathbf{x}$". The shrinking is mixed into the gradient.
- Decoupled decay. Leave the gradient alone, and separately multiply the weights by a number slightly below $1$.
For plain gradient descent these two are exactly the same. For Adam they are not. Adam divides the whole gradient by $\sqrt{\hat v}$, so the decay term gets divided too. A weight whose loss-gradient is big gets almost no shrinking, and a weight whose loss-gradient is tiny gets a lot. AdamW applies the decay outside that division, so every weight is shrunk by the same fraction.
Plain GD first. The L2 gradient is $\mathbf{g} + \lambda\mathbf{x}$. The step is $\mathbf{x} - \eta(\mathbf{g} + \lambda\mathbf{x}) = (1 - \eta\lambda)\,\mathbf{x} - \eta\,\mathbf{g}$. That is "decay by the factor $1-\eta\lambda$, then take the plain step": identical to weight decay.
Now Adam, with numbers. Two weights, both equal to $1$. Take $\lambda = 0.1$, $\eta = 0.01$. Suppose the loss gradient is steadily $g_1 = 5$ for weight 1 and $g_2 = 0.01$ for weight 2.
- Adam with L2: the gradient Adam sees is $g + \lambda w$: $5.1$ for weight 1 and $0.11$ for weight 2. Adam turns each into a step of about $\eta = 0.01$.
- How much of that movement comes from the decay? For weight 1: $\lambda w/(g+\lambda w) = 0.1/5.1 = 1.96\%$. For weight 2: $0.1/0.11 = 90.9\%$. In absolute terms the decay moves weight 1 by $0.01\cdot 0.0196 = 0.0002$ and weight 2 by $0.01\cdot 0.909 = 0.0091$ per step: 46 times more for the weight with the small loss-gradient, although both weights and $\lambda$ are identical.
- AdamW: Adam sees only $g$ ($5$ and $0.01$), takes its step of about $\eta$, and then subtracts $\eta\lambda w = 0.01\cdot 0.1\cdot 1 = 0.001$ from both weights. Same decay for both.
The extreme case: no loss gradient at all. With $g = 0$ and $\eta\lambda = 0.01$ (take $\eta = 0.01$, $\lambda = 1$), starting from $w = 1$: AdamW gives $(1-0.01)^t$: $0.605$ after 50 steps, $0.366$ after 100, $0.049$ after 300. Adam with L2 sees the gradient $\lambda w$, and Adam turns any steady gradient into a step of about $\eta$: the weight falls by about $0.01$ per step at first (a straight line, not an exponential), is at $0.537$ at step 50 and $0.224$ at step 100, and is essentially $0$ ($0.0002$) by step 300. A very different, and much stronger, effect than AdamW's.
AdamW (Loshchilov and Hutter, 2019) computes $\hat{\mathbf{m}}_t$ and $\hat{\mathbf{v}}_t$ from the loss gradient only, exactly as Adam does, and then
$$\mathbf{x}_{t+1} = \mathbf{x}_t - \eta\left(\frac{\hat{\mathbf{m}}_t}{\sqrt{\hat{\mathbf{v}}_t} + \varepsilon} + \lambda\,\mathbf{x}_t\right).$$$\lambda$ is the weight-decay coefficient. The term $-\eta\lambda\mathbf{x}_t$ is the decoupled decay: when the loss gradient is $\mathbf{0}$, every weight is multiplied by exactly $1 - \eta\lambda$ per step.
Compare with Adam + L2 (the older way): replace $\mathbf{g}_t$ by $\mathbf{g}_t + \lambda\mathbf{x}_t$ before computing $\mathbf{m}$ and $\mathbf{v}$. Then the decay enters through $\hat m/\sqrt{\hat v}$, so the effective shrinking of a weight is roughly $\eta\lambda w / \sqrt{\hat v}$: inversely proportional to that weight's gradient size. With a learning-rate schedule, the decay term is scaled by the current $\eta_t$ too (this is what the usual implementations do).
An honest note. AdamW is not "minimising the loss plus an L2 penalty": at a resting point it satisfies $\hat m/\sqrt{\hat v} = -\lambda\mathbf{x}$, which depends on gradient scales. Why it often generalizes better than Adam + L2 is partly empirical. What is solid: the decay strength becomes independent of the gradient scale, which also makes $\lambda$ and $\eta$ easier to tune.
Why do we need it?
We want weight decay to mean "shrink every weight by the same small fraction", so that one number $\lambda$ is meaningful. Inside Adam, an L2 penalty does not do that: its effect depends on how large each weight's gradient happens to be.
Where is it used?
AdamW is the standard optimizer for Transformers: BERT-style encoders, GPT-style language models, Vision Transformers and most fine-tuning recipes. It is torch.optim.AdamW.
How is it used?
Use AdamW(params, lr=..., weight_decay=...). Values like $0.01$ to $0.1$ are common for $\lambda$. Many recipes switch the decay off for biases and normalization-layer gains. Do not also add an L2 term to the loss.
Do not use both. If you use AdamW, set the L2 penalty in the loss to zero; otherwise you have both effects at once.
Names are confusing. In some libraries the parameter called weight_decay of plain Adam is the L2 penalty (coupled). Only the optimizer named AdamW (or a flag like decoupled_weight_decay) decouples it. Check the documentation.
Decay and the learning rate are linked. The decay per step is $\eta\lambda$. If you change $\eta$ with a schedule, the decay changes with it. Tune the two together.
Quick check: AdamW with $\eta = 0.001$ and $\lambda = 0.1$. A weight receives a loss gradient of exactly $0$ for 1000 steps. By what factor has it shrunk?
Each step multiplies it by $1 - \eta\lambda = 1 - 0.0001 = 0.9999$. After 1000 steps: $0.9999^{1000} \approx e^{-0.1} \approx 0.905$. About $10\%$ smaller.
The optimizer race: all of them on the same landscape core
Reading about optimizers is one thing. Watching them run side by side teaches more. Put GD, momentum, Nesterov, AdaGrad, RMSProp and Adam on the same hill, at the same starting point, let each take the same number of steps, and see who gets to the bottom first, who zig-zags, who overshoots and who gets lost.
A fair race needs two rules. Every runner uses one gradient per step, so "steps" really means "cost". And every runner gets its own learning rate, because the best $\eta$ for GD is useless for AdaGrad. Finding each runner's good $\eta$ is half the work, as the next section shows.
The bowl with $\kappa = 100$, valley along the axes ($f = \tfrac12(x^2 + 100y^2)$), start $(-2.5, 2)$, a budget of 200 steps. These are the settings of the first preset in the widget below. The number of steps until $f \lt 0.001$:
| Optimizer | $\eta$ (and $\beta$) | steps until $f \lt 0.001$ |
|---|---|---|
| AdaGrad | $1$ | 12 |
| Nesterov | $0.01$, $\beta = 0.9$ | 25 |
| Adam | $0.3$ | 77 |
| Momentum | $0.01$, $\beta = 0.9$ | 97 |
| GD | $0.0198$ (the best fixed step $2/(L+\mu)$) | not within 200 (still $f \approx 0.06$) |
(RMSProp with $\eta = 0.03$ gets below $0.001$ at step 100 but keeps jittering: it takes steps of a fixed size and never settles.) Rotate the same bowl by $0.6$ radians and the picture changes: AdaGrad and Adam no longer reach $0.001$ in 200 steps, while momentum and Nesterov still get there (after 71 and 24 steps). The widget does exactly this in its second preset.
All the methods in this chapter share one skeleton: at step $t$, compute a gradient, update one or two running averages, move. They differ in what they remember and what they divide by:
| Method | Extra numbers stored per parameter | Knobs | Idea in one line |
|---|---|---|---|
| GD | 0 | $\eta$ | step downhill |
| Momentum, Nesterov | 1 (velocity) | $\eta,\ \beta$ | keep going the way you were going (Nesterov: look ahead first) |
| AdaGrad | 1 (sum of squares) | $\eta$ | divide by total gradient energy seen |
| RMSProp | 1 (average of squares) | $\eta,\ \beta$ | divide by recent gradient energy |
| Adam, AdamW | 2 ($\mathbf{m}$ and $\mathbf{v}$) | $\eta,\ \beta_1,\ \beta_2,\ (\lambda)$ | momentum on top of RMSProp, with start-up correction |
Why do we need it?
Theory only goes so far. Real landscapes are valleys, plateaus, saddles and bumps, and the methods behave differently on each. Seeing the paths and loss curves together builds the instinct for which tool fits which terrain.
Where is it used?
The same comparison is made on a small scale before every large training run: try SGD with momentum and Adam(W) on a small version of the problem, sweep the learning rate, and compare loss curves. Optimizer-benchmark papers do it on a large scale.
How is it used?
Choose a landscape, switch runners on or off, and tune each one's $\eta$. Look at two things: the path on the contour map (zig-zag? overshoot? which minimum?) and the loss curve on a log scale (fast start? slow end? stalls? noise floor?).
A race is not a verdict. The ranking depends on the landscape, the start and how well each $\eta$ was tuned. Tune everyone fairly, or the result says more about your tuning than about the methods. In the next section you will scan the learning rate of every method.
Deterministic is not real training. Here the gradient is exact. In real training it is noisy (a random mini-batch), which changes the story: it is why decaying the step size matters (schedules, below) and why Adam's averaging helps.
Quick check: in the race, why must "steps" be a fair measure of cost? Which optimizer would need two gradients per step if implemented naively?
Computing a gradient is the expensive part (a backward pass through the whole model), and every method in this chapter needs exactly one gradient per step. Nesterov looks like it needs two (one at $\mathbf{x}$ and one at the look-ahead point), but it only needs the one at the look-ahead point. Methods like line search or Newton's method need extra function or Hessian evaluations per step, so counting only "steps" would not be fair to them (Chapter 3.12).
Choosing the learning rate: the sweet spot of each optimizer core
Every optimizer has a Goldilocks zone for its learning rate. Too small: it is safe but crawls. Too large: it bounces around, or blows up. In between: fast and stable. Plot the final loss against $\eta$ (on a log scale) and you almost always see a valley: a flat wall on the right where it blows up, and a slow slope on the left.
What differs between optimizers is where the valley is and how wide it is. For plain GD the best $\eta$ is tied to the steepest curvature ($\eta \lt 2/L$), so it changes with every problem. For Adam the step is "about $\eta$ per parameter per step", so $\eta$ is tied to how far a parameter should move, which is why a value like $10^{-3}$ is a reasonable first guess for very different models.
We scan $\eta$ on 41 values from $10^{-4.5}$ to $10^{1.5}$ (equally spaced on a log scale, $0.15$ apart in the exponent), run each optimizer for 200 steps, and call an $\eta$ usable if the final error $f - f^\star$ is below $1\%$ of the starting error. For the bowl with $\kappa = 100$ (start $(-2.5, 2)$) the usable window is about:
| Optimizer | best $\eta$ in the scan | width of the usable window (in powers of 10) |
|---|---|---|
| GD | 0.016 | 1.2 |
| Momentum ($\beta = 0.9$) | 0.0079 | 2.55 |
| Adam | 1.4 | 3.45 |
On the Rosenbrock valley (start $(-1.2, 1)$) the same scan gives only $0.15$ for GD, $0.75$ for momentum and $1.2$ for Adam. So on this landscape GD has to be tuned to within a factor of about $1.4$, while Adam still works across more than a factor of $10$. This is the sense in which adaptive methods are "easier to tune". It does not say they reach a better final answer.
A learning-rate scan (sweep) is the standard way to tune $\eta$:
- Pick a fixed budget (number of steps or epochs) and a metric (final loss, or loss on held-out data).
- Choose $\eta$ on a logarithmic grid (for example $10^{-4}, 3\cdot 10^{-4}, 10^{-3}, 3\cdot 10^{-3}, \dots$). A linear grid wastes almost all its points on the large values.
- Run each, plot the metric against $\eta$, and take a value near the bottom of the valley, not at the very edge of the blow-up (the edge is fragile).
Rules of thumb (heuristics, not theorems): tune $\eta$ first, then $\beta$ or the schedule. Start Adam at $10^{-3}$ and move by factors of $3$ or $10$. A quick screening trick used in practice, the learning-rate range test, raises $\eta$ exponentially over a short run and watches where the loss starts to explode. And whenever you change the batch size, momentum or schedule, $\eta$ may need re-tuning.
Why do we need it?
The learning rate is the one setting that almost always matters. A wrong $\eta$ wastes a whole training run, so we want a cheap, reliable way to find a good one for each optimizer.
Where is it used?
Every serious training run starts with some version of this scan: hyper-parameter sweeps in research code, tools like Optuna or Weights & Biases sweeps, and learning-rate finders in training libraries.
How is it used?
Run short trials over a log grid of $\eta$ and plot the metric. Pick a value in the lower-middle of the good region. Then confirm with a full-length run, because the best $\eta$ for a short run is often larger than for a long one.
The best $\eta$ depends on the budget and on the other knobs. A larger $\eta$ often wins in a short run, but is too aggressive in a long one. Momentum $\beta$, batch size and the schedule all move the best $\eta$. Treat the scan as a map near your current setting, not a universal number.
"Easier to tune" is not "no tuning". Adam's default $10^{-3}$ is a good start, but the best value for a given model can still differ by a factor of $10$ or more, and large models are often sensitive to it.
Quick check: you scan GD on a problem whose steepest curvature is $L = 50$. Where must the right edge of the valley be, and why does the left side slope down gently?
The right edge is at $\eta = 2/L = 0.04$: above it each step overshoots by more than it corrects and the error grows. Left of the valley the steps are safe but tiny, so after a fixed budget the error has not had time to shrink: the smaller $\eta$, the larger the final error.
Learning-rate schedules: big steps early, small steps late core
Think of parking a car. Far from the spot you drive fast. As you get close you slow down, and for the last metre you creep. A learning-rate schedule does the same for training: a large $\eta$ at the start to make quick progress, and a smaller and smaller $\eta$ as you approach a good solution, so that you can settle into it instead of bouncing around it.
Why does a fixed $\eta$ not settle? In real training each gradient is noisy (it comes from a random mini-batch). With a fixed step, the parameters keep being kicked around the minimum, in a "noise ball" whose size grows with $\eta$. Shrinking $\eta$ shrinks the ball. (Even with exact gradients, RMSProp-style methods take steps of about $\eta$ whatever the gradient size, so they can jitter around the minimum.) You will study this noise in Chapter 3.14.
Base learning rate $\eta_0 = 0.1$ and a horizon of $T = 100$ steps. Each schedule gives a learning rate $\eta_t$ at step $t$:
| Schedule | setting | $\eta_t$ at some steps |
|---|---|---|
| Constant | $0.1$ always | |
| Step decay | multiply by $0.1$ every $30$ steps | $t = 0..29$: $0.1$; $t = 30..59$: $0.01$; $t = 60..89$: $0.001$; $t = 90..$: $0.0001$ |
| Exponential | $\gamma = 0.95$ per step | $t = 10$: $0.0599$; $t = 50$: $0.00769$; $t = 100$: $0.00059$ |
| Inverse-time | $k = 0.1$ | $t = 10$: $0.05$; $t = 100$: $0.00909$ |
| Cosine | $T = 100$, $\eta_{\min} = 0$ | $t = 0$: $0.1$; $t = 25$: $0.0854$; $t = 50$: $0.05$; $t = 75$: $0.0146$; $t = 100$: $0$ |
Check one by hand: the cosine at $t = 25$ is $0.5\cdot 0.1\cdot(1 + \cos(\pi/4)) = 0.05\cdot 1.7071 = 0.0854$.
A schedule is a rule $t \mapsto \eta_t$. The standard ones, with base rate $\eta_0$:
| Name | Formula $\eta_t$ | Shape and use |
|---|---|---|
| Constant | $\eta_0$ | the baseline; fine for short runs and for a quick sweep |
| Step decay | $\eta_0\,\gamma^{\lfloor t/s\rfloor}$ | staircase: drop by a factor $\gamma$ (often $0.1$) every $s$ steps. Classic for image classifiers. |
| Exponential decay | $\eta_0\,\gamma^{t}$ with $\gamma \lt 1$ | smooth geometric decay; the rate falls by the same fraction each step |
| Inverse-time decay | $\eta_0/(1 + k t)$ | slow, $1/t$-style decay. The sum $\sum\eta_t$ diverges while $\sum\eta_t^2$ converges: the classical condition (Robbins–Monro) for stochastic gradient methods to converge (Chapter 3.14). |
| Cosine decay | $\eta_{\min} + \tfrac12(\eta_0 - \eta_{\min})\big(1 + \cos\frac{\pi t}{T}\big)$ | half a cosine wave from $\eta_0$ to $\eta_{\min}$ over $T$ steps: flat at the start, steepest in the middle, flat at the end. Very popular today. |
Where does the cosine come from? $\cos(\pi t/T)$ goes from $1$ (at $t = 0$) to $-1$ (at $t = T$), so $\tfrac12(1+\cos)$ goes smoothly from $1$ to $0$. It is just a convenient smooth "S" curve with zero slope at both ends. There is no deep theory behind the choice. (The linear warmup and the combined warmup + cosine schedule are in the next section.)
Why do we need it?
A step size that is good at the start (far from a solution) is too big at the end (near it), where it only shakes the parameters around the minimum. A schedule gives you both speed and a clean finish with one extra setting.
Where is it used?
Step decay in classic ImageNet training of ResNets; cosine decay in vision models, diffusion models and language models; inverse-square-root style decay in the original Transformer; inverse-scaling decay in scikit-learn's SGDClassifier (learning_rate='invscaling').
How is it used?
Pick a base $\eta_0$ (by a scan), pick the total number of steps $T$ in advance (cosine needs it), and attach a scheduler to the optimizer (torch.optim.lr_scheduler). Call scheduler.step() every step or every epoch, and log the current $\eta_t$.
A schedule is tied to the horizon. Cosine decay needs to know $T$ in advance. If you stop early, you stop at a still-large $\eta$. If you extend training, you must restart the schedule. Some recipes avoid this with "warmup-stable-decay" schedules, which keep $\eta$ constant for most of the run and decay only at the end.
Decay is not magic for deterministic problems. With exact gradients, a constant $\eta$ that is stable still converges (for GD or momentum). The strong benefit of decay appears when the gradient is noisy, as in the widget.
Per-epoch or per-step? Libraries differ. Check whether your scheduler is stepped after every batch or after every epoch, or your decay will be off by a large factor.
Quick check: exponential decay with $\gamma = 0.99$ and $\eta_0 = 0.1$. What is $\eta_t$ after 230 steps, and after how many steps has it dropped to about a tenth of $\eta_0$?
$0.99^{230} = e^{230\ln 0.99} = e^{-2.31} \approx 0.099$, so $\eta_{230} \approx 0.0099$, a tenth of $\eta_0$ (it takes about $\ln 10/(-\ln 0.99) \approx 229$ steps for each factor of $10$).
Warmup: start gently, then speed up core
A schedule that only decays starts at its largest step. That is the worst moment for a big step. At the start the parameters are random, the loss surface can be very steep, and the optimizer's running averages are nearly empty. Warmup does the opposite of decay for a short time: it starts with a tiny learning rate and raises it, step by step, to the full value $\eta_0$. Then the usual decay (for example cosine) takes over.
Think of a cold engine: you do not floor the accelerator at once. Three common explanations of why warmup helps:
- Sharp start. Near the random initial point the curvature can be far higher than it will be later. A step size that is perfectly safe in the flat region later is too large at the start (remember: GD needs $\eta \lt 2/L$, and $L$ is large there). Starting small keeps us inside the stable zone until the parameters have moved to flatter ground.
- Poor second-moment estimate. Adam divides by $\sqrt{\hat v}$, but in the first steps $\hat v$ is an average of only a handful of squared gradients, so it is a very rough estimate of the true gradient scale (see the widget below).
- Big batches and deep networks. Large-batch training uses large learning rates, which are especially risky early; linear warmup is a standard remedy. Many Transformer recipes diverge or train poorly without it.
These are the standard explanations. Warmup is an empirical recipe backed by partial theory, not a proven law.
Linear warmup over $W = 10$ steps to $\eta_0 = 0.1$: $\eta_t = \eta_0\,(t+1)/W$ for $t = 0, 1, \dots, 9$. That gives $0.01,\ 0.02,\ 0.03,\ \dots,\ 0.09,\ 0.10$. From $t = 10$ on, we are at the full rate (or begin the decay).
Warmup + cosine with $W = 10$ and $T = 100$ total steps, $\eta_0 = 0.1$, $\eta_{\min} = 0$: the rate climbs from $0.01$ to $0.1$ during the first 10 steps; then it follows half a cosine over the remaining 90 steps. At $t = 55$ (halfway through the cosine part, $t - W = 45$ of $90$) it equals $0.05$; at $t = 99$ it is about $0.00003$; at $t = 100$ it is $0$.
A toy where warmup is needed. A two-weight "network" $f(a, b) = \tfrac12(ab - 1)^2$ starting at $(3, 3)$. The curvature (largest Hessian eigenvalue) there is $26$, so GD would need $\eta \lt 2/26 \approx 0.077$ locally. Near the solution $(1, 1)$ the largest curvature is only $2$, so any $\eta$ below $1$ would be fine there. With constant $\eta = 0.3$ the very first step throws the weights far away and the run diverges. With a 20-step linear warmup to the same $\eta = 0.3$ it converges to $(1, 1)$. (Constant $\eta = 0.15$ survives, by taking a wild first step that happens to land on the other solution $(-1, -1)$.) You can try all this in the widget below.
Linear warmup over $W$ steps:
$$\eta_t = \eta_0\,\frac{t+1}{W} \quad (t = 0, 1, \dots, W-1), \qquad \eta_t = \eta_0 \quad (t \ge W).$$Linear warmup + cosine decay (the modern default for large models), for a total of $T$ steps:
$$\eta_t = \begin{cases} \eta_0\,\dfrac{t+1}{W} & t \lt W,\\[2mm] \eta_{\min} + \tfrac12(\eta_0 - \eta_{\min})\Big(1 + \cos\dfrac{\pi\,(t - W)}{T - W}\Big) & W \le t \le T.\end{cases}$$Rules of thumb (not laws): the warmup is a small fraction of the whole run (a few percent is common, from a few hundred steps in small runs to thousands in very large ones). Use a longer warmup if training is unstable early, and when you use a bigger batch or a higher $\eta_0$.
How warmup interacts with Adam. Adam's bias correction (previous sections) fixes the size of the early averages; it cannot fix their noise: $\hat v_1 = g_1^2$ is a single squared number. Warmup keeps the early steps small while $\hat v$ collects more samples.
Why do we need it?
The first steps of training are the most fragile: random weights, steep loss, rough running averages. A big step there can produce huge activations, NaNs or a loss that never recovers. A short, gentle start makes the whole run more reliable.
Where is it used?
Almost every Transformer recipe (BERT-style models, GPT-style language models, Vision Transformers), large-batch ImageNet training (Goyal et al., 2017), and fine-tuning recipes for large models.
How is it used?
Use a scheduler that combines a linear ramp over the first $W$ steps with cosine (or another) decay. If the loss spikes or becomes NaN early, lengthen the warmup or lower $\eta_0$. Plot $\eta_t$ to check the shape.
Warmup does not replace a sensible $\eta_0$. If $\eta_0$ is far too large for the flat part of the landscape too, warmup only delays the blow-up. Tune $\eta_0$ first (scan), then add warmup.
It is cheap but not free. A warmup of $W$ steps wastes some compute at small steps. Too long a warmup slows the start; too short fails to protect. The few-percent rule is a starting point, not a result.
Restarting? If you resume training from a checkpoint, restore the optimizer state (the averages $\mathbf{m}$ and $\mathbf{v}$) and the step counter, or the bias correction and the schedule start over.
Quick check: linear warmup over $W = 1000$ steps to $\eta_0 = 3\times 10^{-4}$. What is $\eta$ at step $t = 249$?
$\eta_{249} = \eta_0\,(249+1)/1000 = 3\times 10^{-4}\cdot 0.25 = 7.5\times 10^{-5}$. (We use $t+1$ so that the very first step is not exactly zero.)
Which optimizer should I use? There is no winner core
After the race you might hope for a champion. There is none. Each method fixes some weakness and pays for it somewhere else. Momentum fixes zig-zag but can overshoot. AdaGrad fixes scale and sparsity but its steps die. RMSProp and Adam keep the steps alive but can jitter and are not guaranteed to converge. AdamW fixes the decay but still stores two extra numbers per weight. Which one wins depends on the problem (noise, scale mismatch, curvature), on how well each is tuned, and on what you measure (training loss, or accuracy on new data).
The honest summary of practice is a short list of starting points that work well, not a law.
Typical starting points (common practice, not rules; always check with a learning-rate scan):
| Situation | Common starting recipe |
|---|---|
| Transformers, language models, diffusion models | AdamW, linear warmup then cosine decay, weight decay and gradient clipping |
| Image classifiers (ResNet-style CNNs) | SGD with momentum $0.9$ and a step or cosine schedule. Many report that tuned SGD+momentum generalises as well as or slightly better than Adam here; the evidence is mixed and depends on tuning. |
| Sparse features (text counts, click data) | AdaGrad, or Adam |
| Small smooth problems, full batch | quasi-Newton (L-BFGS, Chapter 3.12) beats all of these |
| Very noisy, non-stationary training (GANs, reinforcement learning) | Adam with small $\eta$ and sometimes a smaller $\beta_1$ |
Memory. The state an optimizer stores per parameter matters for big models. With 32-bit numbers (4 bytes each): GD needs weights and gradients ($2$ numbers per parameter); momentum adds the velocity ($3$); Adam adds $\mathbf{m}$ and $\mathbf{v}$ ($4$). For a model with 7 billion parameters that is $56$ GB, $84$ GB and $112$ GB respectively, before counting activations.
A practical decision checklist (heuristics):
- Start with the community default for your model type (table above). It encodes years of trial and error.
- Scan $\eta$ on a log grid (previous section). Tune $\eta$ before anything else.
- Add a schedule (warmup, then cosine or step decay) and a sensible weight decay (AdamW's $\lambda$).
- If training is unstable: lower $\eta$, lengthen the warmup, clip the gradient norm (Chapter 3.15), or raise $\varepsilon$. If it is slow but stable: raise $\eta$ or $\beta_1$-style momentum, or switch to an adaptive method.
- Judge by validation performance, not only the training loss, and compare optimizers with equally good tuning of each.
What they all share. Every method here is first-order: it uses only gradients. None of them uses curvature beyond a diagonal rescaling. When the problem is small and smooth enough to afford second-order information, Newton-type methods (Chapter 3.12) converge in far fewer steps; they are rarely affordable for deep networks.
Why do we need it?
Picking an optimizer by habit can cost days of compute. A short, honest decision procedure gets you to a good setting with few trial runs, and tells you what to change when training misbehaves.
Where is it used?
At the start of every project: choosing the optimizer, base learning rate, schedule and weight decay for a new model. Published training recipes (for ResNets, BERT, GPT-style models, diffusion models) are exactly such choices written down.
How is it used?
Copy the recipe of a similar published model, run a short scan of $\eta$, watch the training and validation curves, and change one thing at a time (schedule, then decay, then optimizer). Keep a log of what you tried.
Beware simple stories. "Adam always converges faster", "SGD always generalises better", "adaptive methods are better for deep nets": each is true in some experiments and false in others. Results depend on tuning effort, the schedule, the batch size and the architecture. Treat the table above as priors, and let a small experiment overrule them.
Training loss is not the goal. The optimizer can change which minimum you reach, and different minima can generalise differently. Whether optimizers differ systematically in the generalisation of what they find is still an active research question (you will meet it in Chapter 3.15).
Quick check: you must train a model with 3 billion parameters on a machine where memory is tight. Why might you consider SGD with momentum instead of Adam, and what do you give up?
With 4-byte numbers, momentum stores one extra number per parameter ($3\times 10^9\times 4 = 12$ GB for its state), while Adam stores two ($24$ GB), so you save $12$ GB. You give up Adam's per-parameter scaling, which many models (especially Transformers) rely on for stable, quick training, so you may need much more tuning or may not reach the same quality.
Recap, cheat sheet and practice
- Plain GD uses one fixed step size, so it zig-zags in narrow valleys, is slow along flat directions, and cannot treat parameters differently.
- The exponentially weighted average $m_t = \beta m_{t-1} + (1-\beta)g_t$ remembers about $1/(1-\beta)$ steps. Its weights add up to $1-\beta^t$, so it starts too low.
- Momentum $\mathbf{v} \leftarrow \beta\mathbf{v} - \eta\mathbf{g}$, $\mathbf{x} \leftarrow \mathbf{x} + \mathbf{v}$: steady gradients are amplified by $1/(1-\beta)$, alternating ones are damped. On quadratics, tuned momentum needs about $\sqrt\kappa$ steps instead of $\kappa$. Nesterov measures the gradient at the look-ahead point $\mathbf{x} + \beta\mathbf{v}$, which adds curvature-dependent braking.
- Adaptive methods divide each coordinate's gradient by a running size of its past gradients: a diagonal preconditioner. AdaGrad sums the squares (steps decay forever; good for sparse features). RMSProp averages them (steps stay alive).
- Adam = momentum + RMSProp + bias correction ($\hat m = m/(1-\beta_1^t)$, $\hat v = v/(1-\beta_2^t)$). The first step is about $\eta$ in every coordinate. Defaults $\beta_1 = 0.9$, $\beta_2 = 0.999$, $\varepsilon = 10^{-8}$.
- AdamW applies weight decay outside the adaptive division: $-\eta\lambda\mathbf{x}$. Inside Adam, an L2 penalty gets divided by $\sqrt{\hat v}$ and so depends on gradient size.
- Schedules (step, exponential, inverse-time, cosine) reduce $\eta$ over time to settle; warmup starts small, because the start is sharp and the early averages are rough. Linear warmup + cosine decay is the common modern default.
- There is no universal winner: tune $\eta$ on a log grid for each method and compare on validation performance.
Cheat sheet
| Method | Update (per coordinate where it says so) | Defaults / notes |
|---|---|---|
| GD | $\mathbf{x} \leftarrow \mathbf{x} - \eta\mathbf{g}$ | stable only for $\eta \lt 2/L$ |
| Momentum | $\mathbf{v} \leftarrow \beta\mathbf{v} - \eta\mathbf{g};\ \mathbf{x} \leftarrow \mathbf{x} + \mathbf{v}$ | $\beta = 0.9$; steady step $\eta/(1-\beta)$; stable for $\eta h \lt 2(1+\beta)$ |
| Nesterov | $\mathbf{v} \leftarrow \beta\mathbf{v} - \eta\nabla f(\mathbf{x} + \beta\mathbf{v});\ \mathbf{x} \leftarrow \mathbf{x} + \mathbf{v}$ | stable for $\eta h \lt 2(1+\beta)/(1+2\beta)$ |
| AdaGrad | $G \leftarrow G + g^2;\ x \leftarrow x - \eta g/(\sqrt{G} + \varepsilon)$ | first step $=\eta$; step $\approx \eta/\sqrt{t}$ |
| RMSProp | $s \leftarrow \beta s + (1-\beta)g^2;\ x \leftarrow x - \eta g/(\sqrt{s} + \varepsilon)$ | $\beta = 0.9$; first step $\eta/\sqrt{1-\beta}$ |
| Adam | $m, v$ as above; $x \leftarrow x - \eta\,\hat m/(\sqrt{\hat v} + \varepsilon)$ | $\beta_1 = 0.9,\ \beta_2 = 0.999,\ \varepsilon = 10^{-8}$; first step $\approx \eta$ |
| AdamW | Adam step $-\ \eta\lambda x$ | never also add L2 to the loss |
| Cosine | $\eta_{\min} + \tfrac12(\eta_0 - \eta_{\min})(1 + \cos\frac{\pi t}{T})$ | needs the horizon $T$ |
| Warmup + cosine | $\eta_0 (t+1)/W$ for $t \lt W$, then cosine over $T - W$ | $W$ is a small fraction of $T$ |
import numpy as np
# The bowl f(x) = 0.5 * x^T A x with curvatures 1 and 100 (condition number 100)
A = np.diag([1.0, 100.0])
f = lambda x: 0.5 * x @ A @ x
grad = lambda x: A @ x
def run(step, x0, steps=200):
"""Run an optimizer: `step(x, g, t)` returns the new x. Returns the final f."""
x = np.array(x0, dtype=float)
for t in range(steps):
x = step(x, grad(x), t)
return f(x)
x0 = [-2.5, 2.0]
# --- plain gradient descent ------------------------------------------------
gd = lambda x, g, t, lr=0.0198: x - lr * g
# --- momentum (heavy ball): v <- beta v - lr g ; x <- x + v --------------------
def make_momentum(lr=0.01, beta=0.9):
v = np.zeros(2)
def step(x, g, t):
nonlocal v
v = beta * v - lr * g
return x + v
return step
# --- Nesterov: the gradient is measured at the look-ahead point x + beta v ------
def make_nesterov(lr=0.01, beta=0.9):
v = np.zeros(2)
def step(x, g_unused, t):
nonlocal v
v = beta * v - lr * grad(x + beta * v)
return x + v
return step
# --- AdaGrad: G <- G + g^2 ; x <- x - lr g / (sqrt(G) + eps) -------------------
def make_adagrad(lr=1.0, eps=1e-8):
G = np.zeros(2)
def step(x, g, t):
nonlocal G
G = G + g**2
return x - lr * g / (np.sqrt(G) + eps)
return step
# --- RMSProp: s <- beta s + (1-beta) g^2 ---------------------------------------
def make_rmsprop(lr=0.03, beta=0.9, eps=1e-8):
s = np.zeros(2)
def step(x, g, t):
nonlocal s
s = beta * s + (1 - beta) * g**2
return x - lr * g / (np.sqrt(s) + eps)
return step
# --- Adam / AdamW (wd = 0 gives Adam) -------------------------------------------
def make_adam(lr=0.3, b1=0.9, b2=0.999, eps=1e-8, wd=0.0, n=2): # n = number of parameters
m = np.zeros(n); v = np.zeros(n)
def step(x, g, t):
nonlocal m, v
m = b1 * m + (1 - b1) * g
v = b2 * v + (1 - b2) * g**2
m_hat = m / (1 - b1**(t + 1)) # bias correction (t starts at 0)
v_hat = v / (1 - b2**(t + 1))
return x - lr * m_hat / (np.sqrt(v_hat) + eps) - lr * wd * x # last term: AdamW decay
return step
for name, step in [("GD", gd), ("momentum", make_momentum()), ("Nesterov", make_nesterov()),
("AdaGrad", make_adagrad()), ("RMSProp", make_rmsprop()), ("Adam", make_adam())]:
print(f"{name:9s} f after 200 steps = {run(step, x0):.3e}")
# --- Adam's first step is about lr in EVERY coordinate, whatever the gradient size ---
adam = make_adam(lr=0.01, n=3)
g = np.array([0.5, -20.0, 0.001]) # gradients that differ by a factor 20,000
x_new = adam(np.zeros(3), g, 0) # first step (t = 0), starting from x = 0
print("Adam first step:", np.round(-x_new, 6)) # step = x_old - x_new, since x_old = 0
# --- learning-rate schedules ----------------------------------------------------
def warmup_cosine(t, lr0=0.1, W=10, T=100, lr_min=0.0):
if t < W:
return lr0 * (t + 1) / W
return lr_min + 0.5 * (lr0 - lr_min) * (1 + np.cos(np.pi * (t - W) / (T - W)))
print([round(float(warmup_cosine(t)), 5) for t in (0, 9, 10, 55, 99, 100)])
# Output you should see:
# GD f after 200 steps = 6.292e-02 <- still far from 0
# momentum f after 200 steps = 6.001e-08
# Nesterov f after 200 steps = 2.018e-10
# AdaGrad f after 200 steps = 2.862e-65 <- the bowl is aligned with the axes
# RMSProp f after 200 steps = 2.430e-02 <- constant-size steps keep it jittering
# Adam f after 200 steps = 1.994e-09
# Adam first step: [ 0.01 -0.01 0.01]
# [0.01, 0.1, 0.1, 0.05, 3e-05, 0.0]
1. Momentum with $\beta = 0.9$ and a constant gradient $g$. After many steps, how does its step compare with plain GD's step $\eta g$?
2. With bias correction, what is the size of Adam's very first step in a coordinate whose gradient is $g_1 \ne 0$ (ignoring $\varepsilon$)?
3. Why do AdaGrad's step sizes shrink forever while RMSProp's do not?
4. What is the key difference between Adam with an L2 penalty and AdamW?
5. Why does Adam divide $m_t$ by $1 - \beta_1^t$?
6. Cosine decay with $\eta_0 = 0.2$, $\eta_{\min} = 0.02$ and horizon $T = 200$. What is $\eta$ at $t = 100$?
Practice problems
A. Momentum by hand. $f(x) = 2x^2$ (gradient $4x$), $x_0 = 1$, $\eta = 0.1$, $\beta = 0.5$, $v_0 = 0$. Compute $x_1, x_2, x_3$ and compare with plain GD.
Step 1: $g = 4$, $v_1 = 0.5\cdot 0 - 0.1\cdot 4 = -0.4$, $x_1 = 0.6$. Step 2: $g = 2.4$, $v_2 = 0.5\cdot(-0.4) - 0.1\cdot 2.4 = -0.44$, $x_2 = 0.16$. Step 3: $g = 0.64$, $v_3 = 0.5\cdot(-0.44) - 0.1\cdot 0.64 = -0.284$, $x_3 = -0.124$. Plain GD multiplies $x$ by $1 - 0.4 = 0.6$ each time: $0.6,\ 0.36,\ 0.216$. Momentum is much closer to the minimum $0$ at step 2 ($0.16$ against $0.36$), then overshoots slightly to $-0.124$.
B. Adam, $\eta = 0.01$, first gradient $\mathbf{g}_1 = (0.5,\ -20,\ 0.001)$. Give the first update (ignore $\varepsilon$ and note what $\varepsilon = 10^{-8}$ changes).
$\hat m_1 = g_1$ and $\hat v_1 = g_1^2$, so each coordinate's step is $\eta\,\mathrm{sign}(g)$: we subtract $(0.01,\ -0.01,\ 0.01)$, i.e. $\mathbf{x}$ changes by $(-0.01,\ +0.01,\ -0.01)$. Three gradients that differ by a factor $20{,}000$ give equal-sized moves. With $\varepsilon = 10^{-8}$ the third step is $0.01\cdot 0.001/(0.001 + 10^{-8}) = 0.0099999$, still $0.01$ to 5 digits.
C. AdaGrad, $\eta = 1$, one coordinate receives the gradients $3,\ 4,\ 0,\ 12$ in turn. Find each step.
$G$ goes $9,\ 25,\ 25,\ 169$. Steps $g/\sqrt{G}$: $3/3 = 1$, $4/5 = 0.8$, $0/5 = 0$, $12/13 = 0.923$. The first step is exactly $\eta$ whatever $g$ was. A zero gradient gives no step but still keeps $G$. The big gradient $12$ is damped because $G$ already contains $9 + 16$.
D. AdamW with $\eta = 0.001$, $\lambda = 0.1$ and a loss gradient of exactly $0$ for 5000 steps. By what factor does a weight shrink? Would Adam with an L2 penalty give the same?
Each step multiplies by $1 - \eta\lambda = 0.9999$, so after 5000 steps: $0.9999^{5000} \approx e^{-0.5} = 0.607$. Adam with L2 would not give the same: the gradient it sees is $\lambda w$ (not zero), and Adam turns any steady gradient into a step of about $\eta = 0.001$ per step. A weight starting at $1$ then falls about $0.001$ per step at first, and (simulated) is $0.26$ after 1000 steps, $0.02$ after 2000 and essentially $0$ after 3000, far below AdamW's $0.607$ at 5000. (This is the comparison from the AdamW widget.)
E. Cosine schedule with $\eta_0 = 0.2$, $\eta_{\min} = 0.02$, $T = 200$. Compute $\eta$ at $t = 50$, $100$ and $150$.
Use $\eta_t = 0.02 + 0.09\,(1 + \cos(\pi t/200))$. $t = 50$: $\cos(\pi/4) = 0.7071$, so $0.02 + 0.09\cdot 1.7071 = 0.1736$. $t = 100$: $\cos(\pi/2) = 0$, so $0.02 + 0.09 = 0.11$. $t = 150$: $\cos(3\pi/4) = -0.7071$, so $0.02 + 0.09\cdot 0.2929 = 0.0464$. (Symmetric: $0.1736 + 0.0464 = 0.22 = \eta_0 + \eta_{\min}$.)
F. Tuned heavy ball. A quadratic has curvatures $\mu = 1$ and $L = 25$ ($\kappa = 25$). Give Polyak's $\eta$ and $\beta$, the per-step shrink rate, and compare with plain GD at its best step.
$\sqrt\kappa = 5$. $\beta = \big(\frac{5-1}{5+1}\big)^2 = (2/3)^2 = 4/9 \approx 0.444$. $\eta = 4/(\sqrt{25} + \sqrt{1})^2 = 4/36 = 1/9 \approx 0.111$. The rate is $\sqrt\beta = (\sqrt\kappa - 1)/(\sqrt\kappa + 1) = 2/3$ per step, against $(\kappa-1)/(\kappa+1) = 24/26 = 0.923$ for GD with $\eta = 2/(L+\mu) = 1/13$. Steps per factor $10$ in the error: $\ln 10/\ln 1.5 \approx 5.7$ against $\ln 10/(-\ln 0.923) \approx 28.8$. Check: starting at $(1,1)$ on $f = \tfrac12(x^2 + 25y^2)$, after 60 steps the distance to the minimum is $2.8\times 10^{-9}$ for the heavy ball and $0.0116$ for GD.
Convergence
An optimizer does not jump to the answer. It takes many small steps and gets closer and closer. This chapter asks three honest questions: closer to what, how fast, and when do we stop? The answers need a little more mathematics than the chapters before, so we always give the picture first and the proof second.
- Say exactly what "the algorithm has converged" means: the points, the scores or the slopes settle down
- Read a convergence rate from a log-scale plot: sublinear (a bending line), linear (a straight line), quadratic (a dive)
- Understand smooth functions, Lipschitz continuity and Lipschitz gradients ($L$-smoothness), and the "roof" picture behind them
- Derive the descent lemma, the $O(1/k)$ rate of gradient descent on convex functions and the linear rate $(1-\mu/L)^k$ on strongly convex ones
- See Newton's method square its error at every step (quadratic convergence)
- Use the gradient norm, stopping criteria and an optimization tolerance sensibly in real code
How to read this chapter. This is where optimization becomes mathematical. Do not worry if a proof feels long on the first pass. The pictures, the final statements and the numbers are what to remember. Every proof here is only a few lines, and every one is followed by a numerical check you can redo. We use gradients and the Hessian from the Calculus guide, and norms and eigenvalues from the Linear Algebra guide. Throughout, $\eta$ is the learning rate (step size), $\mathbf{x}^\star$ is the optimum and $f^\star = f(\mathbf{x}^\star)$ is the best value.
What does "convergence" mean? core
Picture a hiker walking down a foggy mountain toward the lowest point. When can she say "I have arrived"? There are three different things she could watch:
- Where she stands. Are her footprints settling on one spot?
- How low she is. Is her height above sea level settling at the lowest value?
- How steep the ground feels. Is the slope under her boots going to zero?
In the fog she has no map. She cannot know the lowest spot or the lowest height. But she can feel the slope. That is why the third question matters so much in practice.
Take $f(x) = x^2$. Its minimum is at $x^\star = 0$ with $f^\star = 0$. Gradient descent with $\eta = 0.25$ updates $x \leftarrow x - 0.25\cdot 2x = 0.5\,x$. Start at $x_0 = 4$.
| step $k$ | $x_k$ | distance $\lvert x_k - 0\rvert$ | $f(x_k)$ | slope $\lvert f'(x_k)\rvert = 2\lvert x_k\rvert$ |
|---|---|---|---|---|
| 0 | 4 | 4 | 16 | 8 |
| 1 | 2 | 2 | 4 | 4 |
| 2 | 1 | 1 | 1 | 2 |
| 3 | 0.5 | 0.5 | 0.25 | 1 |
| 4 | 0.25 | 0.25 | 0.0625 | 0.5 |
All three columns go to zero together. The distance and the slope halve at every step; the score falls four times faster, because $f = x^2$ squares the distance.
An algorithm produces a sequence of points $\mathbf{x}_0, \mathbf{x}_1, \mathbf{x}_2, \dots$ We can measure its progress in three ways:
- Convergence of the iterates: $\|\mathbf{x}_k - \mathbf{x}^\star\| \to 0$. The points settle on the optimum.
- Convergence of the function values: $f(\mathbf{x}_k) - f^\star \to 0$. The score reaches its lowest value. (The difference $f(\mathbf{x}_k) - f^\star \ge 0$ is called the optimality gap.)
- Convergence of the gradients: $\|\nabla f(\mathbf{x}_k)\| \to 0$. The ground becomes flat. The points where $\nabla f = \mathbf{0}$ are the stationary points (Chapter 3.2).
Warning: the three are not the same. For a nicely curved bowl they go together. For a function with a very flat bottom (like $x^4$) the gap can be tiny while the point is still far from $\mathbf{x}^\star$. And for a non-convex function a small gradient only tells you that you are near some stationary point: a local minimum, a local maximum or a saddle. Only for convex problems (Chapter 3.6) does "a stationary point" mean "the global optimum".
Why do we need it?
Optimizers never finish exactly. We need a precise way to say "close enough", and a way to compare two algorithms honestly: which one gets closer in fewer steps, and closer in what sense?
Where is it used?
Every training-loop log (loss curve, gradient-norm curve), convergence theorems in papers, solver settings in scikit-learn, SciPy and PyTorch (tolerances, max iterations), and the "has it converged?" warning that logistic regression sometimes prints.
How is it used?
Pick a yardstick you can actually compute. You almost never know $\mathbf{x}^\star$ or $f^\star$ in a real problem, so in code you watch the loss and the gradient norm, and in proofs you bound the gap $f(\mathbf{x}_k)-f^\star$.
"It converged" is a statement about a yardstick. Always ask: converged in distance, in score, or in gradient? A loss curve that looks flat only tells you the score has stopped changing. It does not prove you are near the optimum, and in deep learning (a non-convex problem) there is often no single optimum to be near.
Quick check: for $f(x)=x^2$ with $\eta=0.25$ and $x_0=4$, what is the gradient size $\lvert f'(x_3)\rvert$?
$x_3 = 0.5$, so $\lvert f'(x_3)\rvert = 2\cdot0.5 = 1$. It halves at every step: $8, 4, 2, 1$.
Convergence rate: how fast does the error shrink? core
Two walkers head for the same town. Walker A always covers half of the distance that is left. Walker B takes steps that get shorter and shorter, so after $k$ steps the distance left is about $1/k$. Both arrive "in the end". But A is within one millimetre of the town after about 20 steps, while B needs about a million.
The convergence rate is the answer to "how quickly does the remaining error shrink?" It is the speedometer of an algorithm. Two algorithms can both converge and still differ by a factor of a million in running time.
Let $e_k \ge 0$ be the error after $k$ steps (it could be the distance, or the gap). Three typical sequences, all starting around 1:
- Walker B: $e_k = 1/(k+1)$. After $k = 9, 99, 999$ steps: $0.1,\ 0.01,\ 0.001$. Each extra digit of accuracy costs ten times more steps.
- Walker A: $e_k = 0.5^k$. Every step halves the error: $1, 0.5, 0.25, \dots$ After 20 steps: $9.5\cdot10^{-7}$. Each step gains the same amount of accuracy (about 0.3 of a digit).
- The diver: $e_{k+1} = e_k^2$ starting at $e_0 = 0.5$: $0.5,\ 0.25,\ 0.0625,\ 0.0039,\ 1.5\cdot10^{-5},\ 2.3\cdot10^{-10},\ 5.4\cdot10^{-20}$. The number of correct digits roughly doubles every step: $0.3,\ 0.6,\ 1.2,\ 2.4,\ 4.8,\ 9.6$.
Steps needed to reach an error of $10^{-6}$: B needs $10^6$, A needs 20, the diver needs 5.
Let $e_k \to 0$ be the error. The usual families are:
- Sublinear: $e_k \approx C/k^{p}$ for some $p \gt 0$ (for example $O(1/k)$ or $O(1/\sqrt{k})$). The ratio $e_{k+1}/e_k$ creeps up to $1$. Reaching accuracy $\varepsilon$ takes about $(C/\varepsilon)^{1/p}$ steps.
- Linear (geometric): $e_{k+1} \le \rho\, e_k$ with a fixed $0 \lt \rho \lt 1$, so $e_k \le \rho^k e_0$. The number $\rho$ is the rate. Reaching $\varepsilon$ takes about $\dfrac{\ln(e_0/\varepsilon)}{\ln(1/\rho)}$ steps. Each step gains $-\log_{10}\rho$ digits.
- Superlinear: $e_{k+1}/e_k \to 0$. The ratio itself shrinks (quasi-Newton methods such as BFGS, Chapter 3.12, behave like this).
- Quadratic: $e_{k+1} \le C\,e_k^{2}$. The digits double each step. Reaching $\varepsilon$ takes only about $\log_2\log(1/\varepsilon)$ steps once you are close. (In general "order $q$" means $e_{k+1} \le C e_k^{\,q}$.)
The word linear is confusing. It does not mean slow. It comes from the picture: if you plot $\log e_k$ against $k$, geometric decay $e_k = \rho^k e_0$ gives $\log e_k = k\log\rho + \log e_0$, which is a straight line.
| Family | Typical error | On a log plot | Steps for $10^{-6}$ |
|---|---|---|---|
| Sublinear $1/k$ | $1/k$ | a curve that bends and flattens | $1{,}000{,}000$ |
| Sublinear $1/k^2$ | $1/k^2$ | a bending curve (steeper) | $1{,}000$ |
| Linear, $\rho = 0.9$ | $0.9^k$ | a straight line | 132 |
| Linear, $\rho = 0.5$ | $0.5^k$ | a steeper straight line | 20 |
| Quadratic | $e_{k+1}=e_k^2$ | a curve that dives | 5 |
Why do we need it?
"It converges" is not enough: a method that needs a million steps is useless if another needs twenty. The rate tells you how the cost grows when you ask for one more digit of accuracy.
Where is it used?
Comparing gradient descent ($O(1/k)$ on convex problems), accelerated methods ($O(1/k^2)$), Newton's method (quadratic), and stochastic methods ($O(1/\sqrt{k})$). It also explains why a loss curve plotted on a log axis is the standard diagnostic.
How is it used?
Plot the error (or the gradient norm) against the step number on a log y-axis. A straight line means linear convergence, and its slope gives $\rho$. A line that keeps bending toward flat means sublinear. A dive means quadratic.
Rates are about the long run. A "linear" method can still be slow if $\rho$ is close to $1$. A sequence with $e_{k+1} = 0.99\,e_k$ is linear, but it needs $\ln(10^6)/\ln(1/0.99) \approx 1375$ steps for six digits. The constants matter too: an $O(1/k)$ method with a tiny constant can beat a linear method for modest accuracy. And rates describe worst-case guarantees or asymptotic behaviour, not every single step.
Quick check: errors in a run are $1,\ 0.1,\ 0.01,\ 0.001$. Which family is this, and what is the rate?
Each step multiplies the error by $0.1$, so it is linear with $\rho = 0.1$. It gains exactly one digit per step. On a log plot it is a straight line going down by one unit per step.
Smooth functions core
Ride a bicycle on a road. A smooth road has no sudden bumps or sharp corners: the slope you feel changes gently, so what you feel under your wheels now is a good guide to the next few metres. A road with a sharp ridge is not smooth. At the ridge, the slope jumps from "uphill" to "downhill" in an instant.
Gradient descent reads the slope where you stand and steps that way. This only works well when the slope does not change wildly nearby. That is the whole reason optimization cares about smoothness.
- $f(x) = x^2$ has slope $2x$, which changes gently. Smooth.
- $f(x) = \lvert x\rvert$ has a sharp corner at $0$: the slope jumps from $-1$ to $+1$. Not smooth there.
- The ReLU $\max(0, x)$ and the hinge $\max(0, 1-x)$ have the same kind of corner.
- The Huber function ($\tfrac12x^2$ for $\lvert x\rvert\le1$, and $\lvert x\rvert-\tfrac12$ beyond) glues a parabola to two straight lines. The slope is continuous (it reaches $\pm1$ and stays), but the curvature jumps from $1$ to $0$ at $\lvert x\rvert = 1$.
- The softplus $\ln(1+e^{x})$ is a rounded version of ReLU. Smooth everywhere.
We classify functions by how many times you can differentiate them with continuous results:
- $C^0$: continuous (no jumps). $\lvert x\rvert$ and ReLU are $C^0$ and no better.
- $C^1$: continuously differentiable (the gradient exists and is continuous). Huber is $C^1$ but not $C^2$.
- $C^2$: the Hessian exists and is continuous. $C^\infty$: differentiable as many times as you like (polynomials, $e^x$, softplus, $\sin$).
"Smooth" in everyday ML talk means at least $C^1$. In optimization theory it usually means something a little stronger: the gradient is Lipschitz (two sections ahead). Functions with corners are called non-smooth; they need subgradients or proximal methods (Chapter 3.13).
Why do we need it?
The gradient is only a trustworthy compass when the ground changes gently. Smoothness is the property that lets us say: "a step of this size, in the downhill direction, is guaranteed to help".
Where is it used?
Almost every convergence theorem for gradient descent, momentum and Adam assumes smoothness. In ML: squared error, cross-entropy and softplus are smooth; ReLU networks, hinge loss, L1 penalties and max-pooling are not.
How is it used?
Check the loss: is it $C^1$? If it has corners, either replace it by a smooth cousin (Huber instead of absolute error, softplus instead of ReLU), or use a method built for corners (subgradient descent, proximal methods). In practice, deep-learning libraries just pick a subgradient at the corner and carry on.
Smooth does not mean "flat" or "easy". A smooth function can still have huge curvature somewhere, or millions of local minima. Smoothness only says the slope does not jump. How fast it may change is measured by a number, the Lipschitz constant of the gradient (coming up).
Quick check: is $f(x) = x^{4/3}$ smooth at $0$? What about $f(x) = x^{2}\lvert x\rvert$?
$f'(x) = \tfrac43 x^{1/3}$ is continuous, so $f$ is $C^1$, but its second derivative $\tfrac49 x^{-2/3}$ blows up at $0$, so it is not $C^2$ there. $f(x) = x^2\lvert x\rvert = \lvert x\rvert^3$ has $f'(x) = 3x\lvert x\rvert$ and $f''(x) = 6\lvert x\rvert$, both continuous: it is $C^2$ (but not $C^3$).
Lipschitz continuity: a speed limit for a function core
A road has a speed limit. A function can have one too, but the "speed" is how fast its output changes when you move the input. If moving the input by 1 never changes the output by more than $G$, then the function is $G$-Lipschitz. Its graph is never steeper than slope $G$.
Picture a double cone (an hourglass on its side) with its tip on any point of the graph and walls of slope $\pm G$. A $G$-Lipschitz graph stays inside that cone, for every tip position.
- $f(x) = \lvert x\rvert$: the slope is $\pm1$ everywhere, so $G = 1$. It is Lipschitz even though it is not smooth.
- $f(x) = \sin(2x)$: the slope is $2\cos 2x$, at most $2$ in size, so $G = 2$.
- $f(x) = x^2$ on $[-3, 3]$: the slope $2x$ reaches $6$, so $G = 6$. But on the whole line the slope is unbounded, so $x^2$ is not Lipschitz over all of $\mathbb{R}$.
- $f(x) = \sqrt{\lvert x\rvert}$: near $0$ the slope $\frac{1}{2\sqrt{x}}$ grows without limit, so it is not Lipschitz on any interval around $0$.
So "Lipschitz" and "smooth" are different properties. $\lvert x\rvert$ is Lipschitz and not smooth. $x^2$ is smooth and not (globally) Lipschitz.
A function $f$ is $G$-Lipschitz (on a set $S$) if for all $\mathbf{x}, \mathbf{y}\in S$
$$\lvert f(\mathbf{x}) - f(\mathbf{y})\rvert \le G\,\|\mathbf{x}-\mathbf{y}\|.$$The smallest such $G$ is the Lipschitz constant. For a vector-valued function use $\|F(\mathbf{x})-F(\mathbf{y})\|$ on the left.
For differentiable $f$ on a convex set: $f$ is $G$-Lipschitz exactly when $\|\nabla f(\mathbf{x})\| \le G$ everywhere. (Why one direction holds: by the mean value theorem (somewhere between $\mathbf{x}$ and $\mathbf{y}$ the slope along the segment equals the average slope) $f(\mathbf{x})-f(\mathbf{y}) = \nabla f(\mathbf{z})^\top(\mathbf{x}-\mathbf{y})$ for some point $\mathbf{z}$ between them. By the Cauchy–Schwarz inequality $\lvert\mathbf{a}^\top\mathbf{b}\rvert\le\|\mathbf{a}\|\|\mathbf{b}\|$ this is at most $\|\nabla f(\mathbf{z})\|\,\|\mathbf{x}-\mathbf{y}\|$.)
Why do we need it?
It puts a hard limit on how violently a function can react to a small change. That makes outputs stable, and it lets us prove that simple algorithms (such as subgradient descent) converge even on functions with corners.
Where is it used?
Convergence proofs for non-smooth problems (hinge loss, absolute error, L1 are all Lipschitz), robustness of networks to small input changes, spectral normalisation and the 1-Lipschitz critic in Wasserstein GANs, and gradient clipping (which forces the step to obey a speed limit).
How is it used?
Find a bound on the gradient norm over the region you care about; that is $G$. Convergence guarantees then often read "error $\le G\cdot\text{distance}/\sqrt{k}$". For a network layer $\mathbf{x}\mapsto W\mathbf{x}$ the constant is the largest singular value of $W$.
Quick check: a linear function $f(x) = 3x + 1$. Is it Lipschitz? With what constant?
Yes. $\lvert f(x)-f(y)\rvert = 3\lvert x-y\rvert$, so $G = 3$ exactly. A straight line has the same slope everywhere. (It is also perfectly smooth: its gradient is constant, so it is $L$-smooth with $L = 0$.)
Lipschitz gradients: $L$-smoothness core
Last section put a speed limit on how fast the height can change. Now put a speed limit on how fast the slope can change. If the slope can never change faster than a fixed rate $L$, then the ground never curves more sharply than a parabola of curvature $L$.
That gives a wonderful picture. Stand on the surface at any point $\mathbf{x}$. Hold over it a roof: a parabola that touches the surface at $\mathbf{x}$, has the same slope there, and has curvature $L$. The surface never rises above the roof. A roof is easy to deal with: you can find its lowest point by hand. So wherever you stand you know a guaranteed upper limit for the function. This single picture drives almost every convergence proof for gradient descent.
Take $f(\mathbf{x}) = \tfrac12(x_1^2 + 10x_2^2)$. Its gradient is $\nabla f(\mathbf{x}) = (x_1,\ 10x_2)$. For two points $\mathbf{x},\mathbf{y}$:
$$\nabla f(\mathbf{x}) - \nabla f(\mathbf{y}) = (x_1-y_1,\ 10(x_2-y_2)), \qquad \|\nabla f(\mathbf{x}) - \nabla f(\mathbf{y})\| \le 10\,\|\mathbf{x}-\mathbf{y}\|.$$The factor $10$ is reached when $\mathbf{x}-\mathbf{y}$ points along the $x_2$ axis. So $L = 10$.
The roof at $\mathbf{x} = (2, 1)$. Here $f = \tfrac12(4 + 10) = 7$ and $\nabla f = (2, 10)$. The roof is $R(\mathbf{d}) = 7 + (2,10)\cdot\mathbf{d} + 5\|\mathbf{d}\|^2$ for a move $\mathbf{d}$. Its lowest point is at $\mathbf{d} = -\nabla f/L = (-0.2, -1)$, where $R = 7 + (-0.4 - 10) + 5(0.04 + 1) = 7 - 10.4 + 5.2 = 1.8$. The true function at that spot, $\mathbf{x}+\mathbf{d} = (1.8, 0)$, is $f = \tfrac12(3.24) = 1.62$. And $1.62 \le 1.8$: the surface is below the roof. ✓
A differentiable function $f$ is $L$-smooth (its gradient is $L$-Lipschitz) if for all $\mathbf{x}, \mathbf{y}$
$$\|\nabla f(\mathbf{x}) - \nabla f(\mathbf{y})\| \le L\,\|\mathbf{x}-\mathbf{y}\|.$$(The norm is the usual length.) The consequences you will use:
- The roof (quadratic upper bound). For all $\mathbf{x},\mathbf{y}$: $$f(\mathbf{y}) \le f(\mathbf{x}) + \nabla f(\mathbf{x})^\top(\mathbf{y}-\mathbf{x}) + \frac{L}{2}\|\mathbf{y}-\mathbf{x}\|^2.$$ Proof. Let $\mathbf{d} = \mathbf{y}-\mathbf{x}$. By the fundamental theorem of calculus along the straight line, $f(\mathbf{y}) - f(\mathbf{x}) = \int_0^1 \nabla f(\mathbf{x}+t\mathbf{d})^\top\mathbf{d}\,dt$. Subtract $\nabla f(\mathbf{x})^\top\mathbf{d} = \int_0^1\nabla f(\mathbf{x})^\top\mathbf{d}\,dt$: $$f(\mathbf{y}) - f(\mathbf{x}) - \nabla f(\mathbf{x})^\top\mathbf{d} = \int_0^1 \big(\nabla f(\mathbf{x}+t\mathbf{d}) - \nabla f(\mathbf{x})\big)^\top\mathbf{d}\,dt \le \int_0^1 \underbrace{L\,t\|\mathbf{d}\|}_{\text{Lipschitz}}\,\|\mathbf{d}\|\,dt = \frac{L}{2}\|\mathbf{d}\|^2.$$ (We used Cauchy–Schwarz for the dot product, then the definition of $L$.) $\blacksquare$ The same argument gives a matching floor $f(\mathbf{y}) \ge f(\mathbf{x}) + \nabla f(\mathbf{x})^\top\mathbf{d} - \frac L2\|\mathbf{d}\|^2$.
- In terms of the Hessian. If $f$ is twice differentiable, $f$ is $L$-smooth exactly when every Hessian $\nabla^2 f(\mathbf{x})$ has all its eigenvalues between $-L$ and $L$. For a convex function (all eigenvalues $\ge 0$, see Chapter 3.6) this just says $\lambda_{\max}\big(\nabla^2 f(\mathbf{x})\big) \le L$ everywhere: $L$ is the largest curvature the function ever has.
- Calculation rules. A quadratic $\tfrac12\mathbf{x}^\top A\mathbf{x}$ with $A$ symmetric positive semi-definite has $L = \lambda_{\max}(A)$. Scaling: $cf$ has $|c|L$. Sums: $L_{f+g} \le L_f + L_g$. And a linear map: $\mathbf{w}\mapsto f(X\mathbf{w})$ has $L \le \lambda_{\max}(X^\top X)\,L_f$.
| Function | Smoothness constant $L$ |
|---|---|
| Least squares $\frac{1}{2n}\|X\mathbf{w}-\mathbf{y}\|^2$ | $\lambda_{\max}(X^\top X)/n$ (the Hessian is $X^\top X/n$) |
| Logistic loss $\frac1n\sum \ln(1+e^{-y_i\mathbf{x}_i^\top\mathbf{w}})$ | $\lambda_{\max}(X^\top X)/(4n)$ (each logistic curvature is at most $\tfrac14$) |
| Any of the above plus $\frac\lambda2\|\mathbf{w}\|^2$ | add $\lambda$ |
| $x^4$, $e^x$ | no global $L$ (the curvature is unbounded) |
| ReLU networks, hinge loss, L1 penalty | not smooth: no finite $L$ |
A small $L$ is good news: it means large steps are safe.
Why do we need it?
To choose a safe step size without trial and error, and to prove that gradient descent really does make progress. The roof turns "downhill hopefully helps" into a guarantee: a step of size $1/L$ cannot make things worse.
Where is it used?
The step size rule $\eta = 1/L$ in gradient descent, convergence proofs for GD, SGD, momentum and Nesterov acceleration, and rules of thumb for the learning rate of logistic regression and linear models (compute $\lambda_{\max}(X^\top X)$).
How is it used?
Find (or bound) the largest Hessian eigenvalue over the region you will visit. For a linear model it is $\lambda_{\max}(X^\top X)$ up to a constant. Then set $\eta \le 1/L$. In deep learning $L$ is unknown and changes during training, so people tune $\eta$ instead.
$L$ is a worst case over the whole region. It is set by the sharpest curvature anywhere, even if you only ever visit flat parts. A function can be very flat almost everywhere and still have a huge $L$ because of one sharp spot. Also, $L$-smooth is not "smoother is always better": it is an upper limit on curvature, it says nothing about a lower limit (that is strong convexity, below).
Quick check: least squares on the one-feature data $x = [1, 2, 3]$ with $f(w) = \frac{1}{2\cdot 3}\sum_i (w x_i - y_i)^2$. What is $L$?
$f''(w) = \frac13(1^2 + 2^2 + 3^2) = \frac{14}{3}$ does not depend on $w$ or on $y$. So $L = 14/3 \approx 4.67$, and a safe step size is $\eta = 3/14 \approx 0.214$.
The descent lemma: a step that is guaranteed to help core
Stand at $\mathbf{x}$ with the roof overhead. A gradient step with the right size lands exactly at the bottom of the roof. The real surface is at or below the roof, so the real height after the step is at most the height of the roof's bottom. That is a guaranteed improvement, and we can write down exactly how big it is.
If the step is too long (past twice the right size), we can overshoot the roof's far side where it climbs again, and the guarantee is lost.
Use the same $f(\mathbf{x}) = \tfrac12(x_1^2+10x_2^2)$, $L = 10$, at $\mathbf{x} = (2, 1)$. Then $\nabla f = (2, 10)$ and $\|\nabla f\|^2 = 4 + 100 = 104$. Take $\eta = 1/L = 0.1$.
- The step: $\mathbf{x}^+ = \mathbf{x} - 0.1\cdot(2,10) = (1.8,\ 0)$.
- Before: $f(\mathbf{x}) = 7$. After: $f(\mathbf{x}^+) = \tfrac12\cdot 1.8^2 = 1.62$.
- The lemma promises $f(\mathbf{x}^+) \le f(\mathbf{x}) - \frac{\eta}{2}\|\nabla f\|^2 = 7 - 0.05\cdot 104 = 7 - 5.2 = 1.8$.
- Indeed $1.62 \le 1.8$. ✓ The actual drop (5.38) is larger than the promised drop (5.2).
Descent lemma. Let $f$ be $L$-smooth and take one gradient step $\mathbf{x}^+ = \mathbf{x} - \eta\nabla f(\mathbf{x})$. Then
$$f(\mathbf{x}^+) \le f(\mathbf{x}) - \eta\Big(1 - \frac{L\eta}{2}\Big)\|\nabla f(\mathbf{x})\|^2.$$In particular, for every step size $0 \lt \eta \le 1/L$:
$$\boxed{\,f(\mathbf{x}^+) \le f(\mathbf{x}) - \frac{\eta}{2}\,\|\nabla f(\mathbf{x})\|^2\,}$$Proof. Put $\mathbf{y} = \mathbf{x}^+ = \mathbf{x} - \eta\nabla f(\mathbf{x})$ into the roof inequality, so $\mathbf{y}-\mathbf{x} = -\eta\nabla f(\mathbf{x})$:
$$f(\mathbf{x}^+) \le f(\mathbf{x}) + \nabla f(\mathbf{x})^\top\big(-\eta\nabla f(\mathbf{x})\big) + \frac{L}{2}\eta^2\|\nabla f(\mathbf{x})\|^2 = f(\mathbf{x}) - \eta\|\nabla f(\mathbf{x})\|^2 + \frac{L\eta^2}{2}\|\nabla f(\mathbf{x})\|^2.$$Factor out $\eta\|\nabla f\|^2$ to get the first formula. If $\eta \le 1/L$ then $1 - L\eta/2 \ge \tfrac12$, which gives the boxed one. $\blacksquare$
- Progress is guaranteed whenever $0 \lt \eta \lt 2/L$, because then $1-L\eta/2 \gt 0$. This is the "$\eta \lt 2/L$" rule you saw for a quadratic in the Calculus guide (curvature and the learning rate), now for any smooth function.
- The guaranteed drop $\eta(1 - L\eta/2)\|\nabla f\|^2$ is largest at $\eta = 1/L$, where it equals $\dfrac{1}{2L}\|\nabla f\|^2$.
Corollary (this holds even for non-convex $f$). Run $K$ steps with $\eta = 1/L$ and add up the drops: $\frac{1}{2L}\sum_{k=0}^{K-1}\|\nabla f(\mathbf{x}_k)\|^2 \le f(\mathbf{x}_0) - f(\mathbf{x}_K) \le f(\mathbf{x}_0) - f^\star$ (the right side telescopes). So the smallest of the $K$ squared gradient norms is at most their average:
$$\min_{0\le k \lt K}\|\nabla f(\mathbf{x}_k)\| \le \sqrt{\frac{2L\,\big(f(\mathbf{x}_0)-f^\star\big)}{K}}.$$So gradient descent always finds a point with a small gradient, at rate $O(1/\sqrt K)$, as long as $f$ is $L$-smooth and bounded below. Needing $\|\nabla f\|\le\varepsilon$ is guaranteed after at most $2L(f(\mathbf{x}_0)-f^\star)/\varepsilon^2$ steps (for our example, $f(\mathbf{x}_0) = 23.125$ at $(-2.5, 2)$, so $\varepsilon = 0.1$ is guaranteed within $2\cdot10\cdot23.125/0.01 = 46{,}250$ steps; the real run needs far fewer, because this is a worst-case promise).
Why do we need it?
It is the one inequality that turns "move opposite to the gradient" into a theorem: each step provably lowers the loss by a computable amount. Every convergence proof in the next sections starts from it.
Where is it used?
Proofs for gradient descent, proximal gradient and SGD; guarantees that training reaches a point with small gradient even for non-convex deep networks; and the logic behind backtracking line search, which shrinks the step until the promised decrease is achieved.
How is it used?
Pick $\eta \le 1/L$, then the loss must drop by at least $\frac\eta2\|\nabla f\|^2$. In code, a useful debugging check: with a small enough $\eta$ the loss must go down every step; if it ever goes up, either $\eta$ is above $2/L$ or the gradient is wrong.
The lemma is about one step and one $L$. It says "at least this much decrease", not "exactly this much". Real runs usually do better (the example dropped 5.38 versus 5.2 promised). And it needs the right $L$ for the region you actually move through: if your estimate of $L$ is too small, the "guaranteed" decrease may fail.
Quick check: with $L = 4$ and $\eta = 0.25$, a gradient of length $\|\nabla f\| = 2$. How much must $f$ drop at least?
$\eta = 0.25 = 1/L$, so the drop is at least $\frac{\eta}{2}\|\nabla f\|^2 = 0.125\cdot 4 = 0.5$.
First-order convergence: gradient descent on a convex function gives $O(1/k)$ core
Each step lowers the loss by at least $\frac\eta2\|\nabla f\|^2$ (the descent lemma). That is a big drop when the slope is steep. But as you approach the bottom of a bowl the slope also gets small, so each step helps less than the one before. How much the slope shrinks compared with how far you still have to go is controlled by convexity: for a convex function, the tangent line lies below the graph, so being far above the minimum forces a large enough slope.
Put these together: the progress per step is proportional to (gap)$^2$. A gap that shrinks by an amount proportional to its own square behaves like $1/k$. This is the price of having no lower curvature bound: near the bottom the function may be very flat, so the pull toward the minimum is weak.
Suppose a gap obeys $a_{k+1} \le a_k - c\,a_k^2$ with $c = 0.1$ and $a_0 = 1$. Then $a_1 \le 0.9$, $a_2 \le 0.9 - 0.081 = 0.819$, $a_3 \le 0.819 - 0.0671 = 0.7519$. It is falling, but slowly: the bound $a_k \le \frac{1}{a_0^{-1} + ck}$ gives $a_{10} \le 1/(1+1) = 0.5$, $a_{100} \le 1/11 \approx 0.09$, $a_{1000} \le 1/101 \approx 0.0099$. Each factor of $10$ in accuracy costs a factor of $10$ in steps. That is the signature of $O(1/k)$.
Theorem (gradient descent on a convex, $L$-smooth function). Let $f$ be convex and $L$-smooth with a minimiser $\mathbf{x}^\star$, and run $\mathbf{x}_{k+1} = \mathbf{x}_k - \eta\nabla f(\mathbf{x}_k)$ with $\eta \le 1/L$. Then for every $K\ge1$
$$\boxed{\,f(\mathbf{x}_K) - f^\star \le \frac{\|\mathbf{x}_0 - \mathbf{x}^\star\|^2}{2\eta K}\,}\qquad\Big(\text{with }\eta = \tfrac1L:\ \ \frac{L\,\|\mathbf{x}_0-\mathbf{x}^\star\|^2}{2K}\Big).$$So the gap is $O(1/k)$: reaching accuracy $\varepsilon$ needs about $\dfrac{L\|\mathbf{x}_0-\mathbf{x}^\star\|^2}{2\varepsilon}$ steps. This is a sublinear rate.
Proof. Write $\mathbf{g}_k = \nabla f(\mathbf{x}_k)$ and $\mathbf{r}_k = \mathbf{x}_k - \mathbf{x}^\star$. Three facts:
- Descent lemma: $f(\mathbf{x}_{k+1}) \le f(\mathbf{x}_k) - \frac\eta2\|\mathbf{g}_k\|^2$.
- Convexity (the tangent lies below the graph, Chapter 3.6): $f^\star \ge f(\mathbf{x}_k) + \mathbf{g}_k^\top(\mathbf{x}^\star - \mathbf{x}_k)$, that is, $f(\mathbf{x}_k) \le f^\star + \mathbf{g}_k^\top\mathbf{r}_k$.
- Expand the distance: $\|\mathbf{r}_{k+1}\|^2 = \|\mathbf{r}_k - \eta\mathbf{g}_k\|^2 = \|\mathbf{r}_k\|^2 - 2\eta\,\mathbf{g}_k^\top\mathbf{r}_k + \eta^2\|\mathbf{g}_k\|^2$, so $\mathbf{g}_k^\top\mathbf{r}_k = \dfrac{\|\mathbf{r}_k\|^2 - \|\mathbf{r}_{k+1}\|^2}{2\eta} + \dfrac\eta2\|\mathbf{g}_k\|^2$.
Chain them: $f(\mathbf{x}_{k+1}) \le f(\mathbf{x}_k) - \frac\eta2\|\mathbf{g}_k\|^2 \le f^\star + \mathbf{g}_k^\top\mathbf{r}_k - \frac\eta2\|\mathbf{g}_k\|^2 = f^\star + \dfrac{\|\mathbf{r}_k\|^2 - \|\mathbf{r}_{k+1}\|^2}{2\eta}$. The gradient terms cancel! Now add this up for $k = 0, 1, \dots, K-1$; the right side telescopes (each $\|\mathbf{r}_{k+1}\|^2$ cancels against the next term, so only the first and last survive):
$$\sum_{k=0}^{K-1}\big(f(\mathbf{x}_{k+1}) - f^\star\big) \le \frac{\|\mathbf{r}_0\|^2 - \|\mathbf{r}_K\|^2}{2\eta} \le \frac{\|\mathbf{r}_0\|^2}{2\eta}.$$The descent lemma says $f(\mathbf{x}_k)$ never goes up, so the last term is the smallest of the $K$ terms: $K\,(f(\mathbf{x}_K) - f^\star) \le \sum_{k=0}^{K-1}(f(\mathbf{x}_{k+1}) - f^\star)$. Divide by $K$. $\blacksquare$
Why $1/k$ (a second look). If $a_{k+1} \le a_k - c\,a_k^2$ then $\dfrac{1}{a_{k+1}} \ge \dfrac{1}{a_k(1 - c a_k)} \ge \dfrac{1 + c a_k}{a_k} = \dfrac1{a_k} + c$, so $\dfrac1{a_k} \ge \dfrac1{a_0} + ck$, that is $a_k \le \dfrac{1}{1/a_0 + ck} \le \dfrac{1}{ck}$. For gradient descent the gap $a_k = f(\mathbf{x}_k)-f^\star$ obeys this with $c = \eta/(2R^2)$, where $R = \|\mathbf{x}_0-\mathbf{x}^\star\|$ (the descent lemma gives the drop $\frac\eta2\|\mathbf{g}_k\|^2$, and convexity gives $a_k \le \|\mathbf{g}_k\|\,\|\mathbf{r}_k\|$, where $\|\mathbf{r}_k\| \le R$ because the chain above shows the distance never grows). This quick route gives the weaker constant $2R^2/(\eta k)$; the telescoping proof above is sharper (by a factor 4) but needs a little more work.
Why do we need it?
It is the basic speed guarantee for gradient descent: it tells you roughly how many steps to budget, and it shows the cost of a bad $L$ (a larger $L$ forces a smaller step and a bigger bound) and of a bad starting point (the bound grows with the square of the distance).
Where is it used?
Convex ML models trained with full-batch gradient descent: linear regression, logistic regression, SVMs with smoothed losses. It is also the baseline that momentum methods improve: Nesterov acceleration gets $O(1/k^2)$ on the same class.
How is it used?
Plug in numbers: $L$ from the Hessian, the distance from your initialisation to the solution. The step count is a budget, not a prediction. Real runs on friendly data are usually much faster than the bound. Treat it as the worst case.
A bound is not a prediction. The $O(1/k)$ guarantee is the best statement that holds for all convex smooth functions, not the speed on yours. Strong convexity (next) improves it dramatically, and so does simply having a friendly function. Also, nothing here is about the iterates: the gap goes to zero, but a flat valley can leave you far from the particular $\mathbf{x}^\star$ you wanted.
Quick check: a convex $f$ with $L = 4$, start distance $\|\mathbf{x}_0 - \mathbf{x}^\star\| = 5$, $\eta = 1/L$. After how many steps is the gap guaranteed to be below $0.01$?
The bound is $\frac{L R^2}{2K} = \frac{4\cdot25}{2K} = \frac{50}{K}$. We need $50/K \le 0.01$, so $K \ge 5000$ steps. The guarantee is generous: real runs are usually much quicker.
Strong convexity and the linear rate $(1 - \mu/L)^k$ core
The roof bounds the function from above by a parabola of curvature $L$. Now add a floor: a parabola of curvature $\mu \gt 0$ that the function never dips below. Between the floor and the roof the function is squeezed like a sandwich. Because of the floor, the bowl is definitely curved up everywhere: it has one clear bottom, and near the bottom the slope is large enough to pull you in by a fixed fraction every step. That gives geometric (linear) convergence.
How good the sandwich is depends on the ratio $L/\mu$, the condition number $\kappa$. A round bowl ($\kappa\approx1$) is easy. A long thin valley ($\kappa$ large) is slow.
Take $f(\mathbf{x}) = \tfrac12(x_1^2 + 10x_2^2)$. The Hessian is $\text{diag}(1, 10)$, so $\mu = 1$ (flattest curvature), $L = 10$ (sharpest), $\kappa = 10$. Gradient descent with $\eta = 1/L = 0.1$ updates each coordinate by $x_1 \leftarrow (1 - 0.1\cdot1)x_1 = 0.9\,x_1$ and $x_2 \leftarrow (1 - 0.1\cdot10)x_2 = 0\cdot x_2$.
From $(-2.5, 2)$: $x_2$ is wiped out in the first step and $x_1$ shrinks by $0.9$ each step: $-2.5,\ -2.25,\ -2.025,\ \dots$ After $100$ steps the distance is $2.5\cdot0.9^{100} \approx 6.6\cdot10^{-5}$. Exactly the factor $1 - \mu/L = 1 - 1/\kappa = 0.9$.
With the best fixed step $\eta = \dfrac{2}{L+\mu} = \dfrac{2}{11}$: $x_1\leftarrow(1-\tfrac2{11})x_1 = \tfrac9{11}x_1$ and $x_2\leftarrow(1-\tfrac{20}{11})x_2 = -\tfrac9{11}x_2$. Both shrink by $\tfrac{9}{11}\approx 0.818 = \tfrac{\kappa-1}{\kappa+1}$: faster than $0.9$.
A differentiable $f$ is $\mu$-strongly convex ($\mu\gt0$) if for all $\mathbf{x},\mathbf{y}$
$$f(\mathbf{y}) \ge f(\mathbf{x}) + \nabla f(\mathbf{x})^\top(\mathbf{y}-\mathbf{x}) + \frac{\mu}{2}\|\mathbf{y}-\mathbf{x}\|^2.$$(A floor: the mirror image of the roof.) For twice-differentiable $f$ this says every Hessian eigenvalue is at least $\mu$: $\nabla^2 f \succeq \mu I$. Together with $L$-smoothness: $\mu I \preceq \nabla^2 f \preceq L I$, and $\kappa = L/\mu \ge 1$ is the condition number (conditioning in the Linear Algebra guide). This is only a first look; Chapter 3.6 develops strong convexity fully, with pictures.
A consequence (the Polyak–Łojasiewicz inequality). Minimise the floor over $\mathbf{y}$. Its lowest point is at $\mathbf{y} = \mathbf{x} - \nabla f(\mathbf{x})/\mu$ with height $f(\mathbf{x}) - \frac{1}{2\mu}\|\nabla f(\mathbf{x})\|^2$. Since $f^\star$ is below the floor's lowest point:
$$f^\star \ge f(\mathbf{x}) - \frac{1}{2\mu}\|\nabla f(\mathbf{x})\|^2 \qquad\Longleftrightarrow\qquad \|\nabla f(\mathbf{x})\|^2 \ge 2\mu\,\big(f(\mathbf{x}) - f^\star\big).$$Theorem (linear rate). Let $f$ be $\mu$-strongly convex and $L$-smooth and take $\eta = 1/L$. Then
$$\boxed{\,f(\mathbf{x}_{k+1}) - f^\star \le \Big(1 - \frac{\mu}{L}\Big)\big(f(\mathbf{x}_k) - f^\star\big)\,}\qquad\text{so}\qquad f(\mathbf{x}_k) - f^\star \le \Big(1-\frac1\kappa\Big)^k\big(f(\mathbf{x}_0) - f^\star\big).$$Proof. The descent lemma gives $f(\mathbf{x}_{k+1}) \le f(\mathbf{x}_k) - \frac{1}{2L}\|\nabla f(\mathbf{x}_k)\|^2$. Subtract $f^\star$ and use the consequence above: $f(\mathbf{x}_{k+1}) - f^\star \le (f(\mathbf{x}_k) - f^\star) - \frac{1}{2L}\cdot2\mu\,(f(\mathbf{x}_k)-f^\star) = (1 - \mu/L)(f(\mathbf{x}_k) - f^\star)$. $\blacksquare$
- Steps needed. Since $1 - 1/\kappa \le e^{-1/\kappa}$, we get $f(\mathbf{x}_k) - f^\star \le \varepsilon$ once $k \ge \kappa\ln\dfrac{f(\mathbf{x}_0)-f^\star}{\varepsilon}$. The cost grows linearly in $\kappa$ and only logarithmically in $1/\varepsilon$. (For $\kappa = 10$ and a reduction of $10^{-6}$: at most $10\ln10^6 \approx 138$ steps.)
- The distance converges too. At $\mathbf{x}^\star$ the gradient is $\mathbf{0}$, so the floor gives $f(\mathbf{x}) - f^\star \ge \frac\mu2\|\mathbf{x}-\mathbf{x}^\star\|^2$, hence $\|\mathbf{x}_k - \mathbf{x}^\star\|^2 \le \frac2\mu(1-\frac1\kappa)^k(f(\mathbf{x}_0) - f^\star)$.
- Best fixed step for a quadratic: $\eta = \dfrac{2}{L+\mu}$ gives the factor $\dfrac{\kappa-1}{\kappa+1}$ per step on the distance.
| $\kappa$ | factor, $\eta = 1/L$ ($1-1/\kappa$) | steps for $10^{-6}$ | factor, $\eta=\frac{2}{L+\mu}$ ($\frac{\kappa-1}{\kappa+1}$) | steps for $10^{-6}$ |
|---|---|---|---|---|
| 2 | 0.5 | 20 | 0.333 | 13 |
| 10 | 0.9 | 132 | 0.818 | 69 |
| 100 | 0.99 | 1,375 | 0.980 | 691 |
| 1000 | 0.999 | 13,809 | 0.998 | 6,908 |
The step counts are for a quadratic, shrinking the distance by $10^{-6}$ (computed from $\ln 10^6/\ln(1/\text{factor})$).
Why do we need it?
It explains why some problems train in a few dozen steps and others take forever: the condition number. It also tells you what to fix when training is slow: make the problem rounder.
Where is it used?
Ridge regression and L2-regularised logistic regression (the penalty makes them strongly convex, with $\mu \ge \lambda$), feature scaling and normalisation, preconditioning, and the analysis of momentum and Adam on well-behaved problems.
How is it used?
Estimate $\kappa = \lambda_{\max}/\lambda_{\min}$ of the Hessian. If it is large, standardise your features (this rounds the bowl), add L2 regularisation (this raises $\mu$), or use a method that fights ill-conditioning (momentum, Adam, Newton-type methods).
Linear rate needs both bounds. Without a floor ($\mu = 0$) you only have $O(1/k)$. And a huge $\kappa$ can make a "linear" rate practically as bad as no rate. $\kappa$ is the single number that tells you how hard a quadratic-like problem is for gradient descent. Also remember the guarantee is an upper bound: for the example above the gap bound says $0.9$ per step, but the gap really shrinks by $0.9^2 = 0.81$ per step (the gap is the square of the distance, which shrinks by $0.9$).
Quick check: $\mu = 0.5$, $L = 5$ and $\eta = 1/L$. By what factor does the gap $f - f^\star$ shrink per step at least, and how many steps to cut it by $10^{-6}$ (use the bound $k \ge \kappa\ln 10^6$)?
$\kappa = 10$, so the factor is $1 - 1/10 = 0.9$ per step. Steps: $10\cdot\ln 10^6 = 10\cdot13.8 \approx 138$.
Second-order convergence: Newton's method squares its error
Gradient descent only knows the slope. It feels that the ground tilts, but not how quickly the tilt changes, so it has to take cautious steps. Newton's method also measures the curvature. It replaces the landscape near you by the best-fitting parabola (a Taylor expansion; see also Newton's method in the Calculus guide) and jumps straight to the bottom of that parabola.
Near the solution the real function looks almost exactly like its parabola, so the jump lands extremely close. The error does not just shrink: it gets squared. An error of $0.1$ becomes about $0.01$, then $0.0001$, then $0.00000001$. The number of correct digits doubles every step.
Minimise $f(x) = e^x - 2x$. Then $f'(x) = e^x - 2$ and $f''(x) = e^x$, so $x^\star = \ln 2 = 0.693147\ldots$ Newton's step is $x_{k+1} = x_k - f'(x_k)/f''(x_k) = x_k - (e^{x_k}-2)/e^{x_k} = x_k - 1 + 2e^{-x_k}$. Starting at $x_0 = 0$:
| $k$ | $x_k$ | error $\lvert x_k - x^\star\rvert$ | correct digits | $e_{k+1}/e_k^2$ |
|---|---|---|---|---|
| 0 | 0 | 0.6931 | 0.16 | 0.64 |
| 1 | 1 | 0.3069 | 0.51 | 0.45 |
| 2 | 0.735759 | 0.04261 | 1.4 | 0.49 |
| 3 | 0.694042 | $8.95\cdot10^{-4}$ | 3.0 | 0.4999 |
| 4 | 0.6931476 | $4.00\cdot10^{-7}$ | 6.4 | 0.500 |
| 5 | 0.69314718056 | $8.0\cdot10^{-14}$ | 13.1 |
The correct digits go $0.5,\ 1.4,\ 3.0,\ 6.4,\ 13.1$: roughly doubling each time. The last column shows $e_{k+1}/e_k^2$ settling at $\tfrac12$: that is the constant $C$ in "$e_{k+1} \le C e_k^2$". (Gradient descent on the same function gains only a fixed fraction of a digit per step.)
Newton's method for minimising $f$ takes the step
$$\mathbf{x}_{k+1} = \mathbf{x}_k - \big[\nabla^2 f(\mathbf{x}_k)\big]^{-1}\nabla f(\mathbf{x}_k),$$in practice by solving the linear system $\nabla^2 f(\mathbf{x}_k)\,\mathbf{d} = -\nabla f(\mathbf{x}_k)$ for the step $\mathbf{d}$. (It comes from minimising the second-order Taylor model $f(\mathbf{x}+\mathbf{d}) \approx f + \nabla f^\top\mathbf{d} + \tfrac12\mathbf{d}^\top\nabla^2 f\,\mathbf{d}$: set its gradient $\nabla f + \nabla^2 f\,\mathbf{d}$ to zero.)
Local quadratic convergence (one variable). Let $f$ be three times continuously differentiable, with $f'(x^\star) = 0$ and $f''(x^\star)\ne0$. Take $x$ close enough to $x^\star$ that $f''(x)\ne0$. Taylor-expand $f'$ around $x$ and evaluate at $x^\star$, with some point $\xi$ between them:
$$0 = f'(x^\star) = f'(x) + f''(x)\,(x^\star - x) + \tfrac12 f'''(\xi)\,(x^\star - x)^2.$$Divide by $f''(x)$ and move terms: $\;x - \dfrac{f'(x)}{f''(x)} = x^\star + \dfrac{f'''(\xi)}{2f''(x)}(x - x^\star)^2$. The left side is the Newton step $x^+$, so
$$\boxed{\,x^+ - x^\star = \frac{f'''(\xi)}{2f''(x)}\,(x - x^\star)^2\,}\qquad\Longrightarrow\qquad e_{k+1} \le C\,e_k^2.$$(In the example, $f''' = f'' = e^x$, so $C \approx \tfrac12$, matching the table.)
Many variables (stated without the full proof; it is the same Taylor idea, with the Lipschitz constant of the Hessian playing the role of $f'''$). Suppose that in a ball around $\mathbf{x}^\star$ the Hessian is $M$-Lipschitz (it changes by at most $M\|\mathbf{x}-\mathbf{y}\|$) and $\nabla^2 f \succeq \mu I$. Then for every $\mathbf{x}$ in that ball, $\|\mathbf{x}^+ - \mathbf{x}^\star\| \le \dfrac{M}{2\mu}\|\mathbf{x}-\mathbf{x}^\star\|^2$. This needs the start to be close enough (within about $2\mu/M$) for the squared error to be smaller than the error. On a quadratic function $M = 0$: the Hessian is constant, and Newton lands exactly in one step.
The price. Each step needs the $n\times n$ Hessian ($n^2$ numbers) and a linear solve (about $n^3$ operations). For a network with millions of weights that is impossible, which is why quasi-Newton methods (BFGS, L-BFGS, Chapter 3.12) approximate it. Far from the solution, or where the Hessian is not positive definite, pure Newton can also go badly wrong.
Why do we need it?
To reach high accuracy fast. A quadratic rate means 3 to 6 steps take you from a rough answer to 10 or 12 correct digits, where gradient descent may need thousands. It also ignores the stretching of the bowl: on a quadratic it needs one step however large the condition number is, and near a solution it is not slowed down by a thin valley the way gradient descent is.
Where is it used?
Small and medium convex models solved to high accuracy: logistic regression in classic solvers (the Newton-CG and Newton-Cholesky options of scikit-learn, and IRLS in statistics packages), interior-point methods, and the last polishing steps of many solvers. (The popular lbfgs solver is a quasi-Newton cousin: Chapter 3.12.)
How is it used?
Compute the gradient and Hessian, solve $H\mathbf{d} = -\mathbf{g}$ (never form the inverse), step, repeat. Add a line search or damping so that far-away steps are safe. Stop after a handful of iterations when the gradient norm is tiny.
"Quadratic" is local. The fast rate only starts once you are close to a solution where the Hessian is positive definite. From far away Newton may overshoot, move toward a maximum or a saddle, or take a huge step. That is why practical Newton codes use line search or damping (see Chapter 3.12). Also, "doubling the digits" stops at about 16 digits because of floating-point rounding.
Quick check: a Newton run has errors $10^{-2}$, then $10^{-4}$. Roughly what is the next error, and what would gradient descent with rate $0.9$ have after the same two steps from $10^{-2}$?
Newton squares the error (with a constant near 1): $\approx 10^{-8}$. Gradient descent with $\rho = 0.9$ would only have $10^{-2}\cdot0.9^2 = 8.1\cdot10^{-3}$ after two steps.
The gradient norm: the practical measure of progress core
Back to the hiker in the fog. She cannot see the bottom, but she can feel how steep the ground is under her boots. At the very bottom the ground is flat. So "how steep is it here?" is a progress measure that needs no map.
In code it is just as cheap: the gradient $\nabla f(\mathbf{x}_k)$ is already computed for the next step, so its length $\|\nabla f(\mathbf{x}_k)\|$ is free. You do not need to know where the optimum is, or what its value is.
Take $f(\mathbf{x}) = \tfrac12(x_1^2 + 10x_2^2)$ with $\mu = 1$, $L = 10$, optimum $\mathbf{0}$, $f^\star = 0$. At $\mathbf{x} = (0.1, 0.1)$:
- $\nabla f = (0.1,\ 1)$, so $\|\nabla f\|^2 = 0.01 + 1 = 1.01$ and $\|\nabla f\| = 1.005$.
- True distance: $\|\mathbf{x} - \mathbf{x}^\star\| = 0.1414$. Check: $0.1414 \le \|\nabla f\|/\mu = 1.005$. ✓
- True gap: $f = \tfrac12(0.01 + 0.1) = 0.055$. Check: $\dfrac{1.01}{2\cdot10} = 0.0505 \le 0.055 \le \dfrac{1.01}{2\cdot1} = 0.505$. ✓
A trap. Let $f(x) = 0.005\,x^2$ (so $\mu = L = 0.01$, a very flat bowl). At $x = 1$ the gradient is only $0.01$, which looks tiny, yet the point is a full distance $1$ from the optimum. A tolerance of $10^{-2}$ on the gradient would stop right there.
At a minimiser of a differentiable function, $\nabla f(\mathbf{x}^\star) = \mathbf{0}$. Away from it, the gradient norm is linked to the two other yardsticks:
- Lower bound on the gap (needs only $L$-smoothness): $f(\mathbf{x}) - f^\star \ge \dfrac{1}{2L}\|\nabla f(\mathbf{x})\|^2$. Proof: $f^\star \le f(\mathbf{x} - \tfrac1L\nabla f) \le f(\mathbf{x}) - \tfrac1{2L}\|\nabla f\|^2$ by the descent lemma.
- Upper bound on the gap (needs $\mu$-strong convexity): $f(\mathbf{x}) - f^\star \le \dfrac{1}{2\mu}\|\nabla f(\mathbf{x})\|^2$ (shown in the previous section).
- Bound on the distance ($\mu$-strong convexity): $\|\mathbf{x}-\mathbf{x}^\star\| \le \dfrac{1}{\mu}\|\nabla f(\mathbf{x})\|$. Proof: write the floor inequality at $\mathbf{x}$ looking at $\mathbf{x}^\star$, and at $\mathbf{x}^\star$ looking at $\mathbf{x}$, and add them: $0 \ge \nabla f(\mathbf{x})^\top(\mathbf{x}^\star-\mathbf{x}) + \mu\|\mathbf{x}-\mathbf{x}^\star\|^2$ (the gradient at $\mathbf{x}^\star$ is zero). So $\mu\|\mathbf{x}-\mathbf{x}^\star\|^2 \le \nabla f(\mathbf{x})^\top(\mathbf{x}-\mathbf{x}^\star) \le \|\nabla f(\mathbf{x})\|\,\|\mathbf{x}-\mathbf{x}^\star\|$. Divide.
So a small gradient means a small gap only when $\mu$ is not tiny. And for a non-convex function a small gradient only says "this is near some stationary point": a local minimum, a maximum or a saddle.
Scale. Multiply $f$ by $1000$ and every gradient is $1000$ times larger. So the absolute tolerance "$\|\nabla f\| \le 10^{-6}$" means different things for different problems. Many solvers use the relative test $\|\nabla f(\mathbf{x}_k)\| \le \varepsilon\,\|\nabla f(\mathbf{x}_0)\|$.
Why do we need it?
It is the only progress measure that needs no knowledge of the answer and costs nothing extra. It also has the cleanest meaning: zero gradient is the first-order condition for an optimum (Chapter 3.2).
Where is it used?
A main stopping test in SciPy optimizers (the option gtol of BFGS, CG and L-BFGS-B, default 1e-5), in Newton-type solvers, in scikit-learn's solver tolerances, and as the training-health curve (and clipping trigger) in deep learning.
How is it used?
Log $\|\nabla f\|$ every few steps. Stop when it falls below a tolerance, preferably relative to its starting value. Watch for it exploding (divergence, a learning rate that is too big) or staying flat but non-zero (noise in SGD).
A small gradient is necessary for an optimum but not sufficient for being close. It fails on very flat bowls (small $\mu$), on plateaus, and at saddle points. In stochastic training the gradient of a mini-batch never reaches zero (it is noisy), so people watch the full-data loss or a smoothed average instead.
Quick check: $\mu = 0.1$ and you measure $\|\nabla f(\mathbf{x})\| = 0.02$. How far can $\mathbf{x}$ be from the optimum, at most?
$\|\mathbf{x}-\mathbf{x}^\star\| \le \|\nabla f\|/\mu = 0.02/0.1 = 0.2$. And the gap is at most $\|\nabla f\|^2/(2\mu) = 0.0004/0.2 = 0.002$.
Stopping criteria: when do we stop? core
A loop that never ends is a bug. An optimizer needs a rule that says stop now. There are only a few sensible reasons to stop:
- "The ground is flat enough." (the gradient is small)
- "I have stopped moving." (the steps are tiny)
- "My score is no longer improving." (the loss barely changes)
- "I am out of time." (the maximum number of iterations)
Each reason is a cheap test on numbers the algorithm already has. Each one can also be fooled in its own way: a very slow walker also "stops moving", and a loss on a plateau also "stops improving". So real solvers check several rules and stop when any of them fires.
Gradient descent on $f(x) = x^2$ with $\eta = 0.25$ halves $x$ each step, $x_0 = 4$ (the table in the first section). With tolerance $\varepsilon = 0.1$:
- Gradient rule $\lvert f'(x_k)\rvert = 2\lvert x_k\rvert \le 0.1$. At $x_6 = 0.0625$ the slope is $0.125$ (still too big). At $x_7 = 0.03125$ it is $0.0625$. It stops at step $7$.
- Step rule (the plain version $\lvert x_{k+1}-x_k\rvert \le 0.1$). The step is $\lvert x_k\rvert/2$, so it holds once $\lvert x_k\rvert \le 0.2$: at $x_4 = 0.25$ the step is $0.125$ (too big), at $x_5 = 0.125$ it is $0.0625$. It stops at step 5.
- Maximum iterations $=3$: stops at step 3 with $x_3 = 0.5$, whatever the tolerance.
All stop "near" $0$, but at different places. The rules measure different things, so they stop at different times.
With a tolerance $\varepsilon$ (see the next section), the common rules are:
| Rule | Stop when | Good because | Fooled when |
|---|---|---|---|
| Gradient norm | $\|\nabla f(\mathbf{x}_k)\| \le \varepsilon$ (or $\le \varepsilon\|\nabla f(\mathbf{x}_0)\|$) | Measures flatness directly; free to compute | Very flat bowls (stops far from the optimum); saddles and plateaus in non-convex problems; depends on the scale of $f$ |
| Step size | $\|\mathbf{x}_{k+1}-\mathbf{x}_k\| \le \varepsilon(1+\|\mathbf{x}_k\|)$ | Easy; works without gradients | A tiny learning rate or a crawl along a flat valley: tiny steps, but still far away |
| Relative function change | $\lvert f_k - f_{k+1}\rvert \le \varepsilon(1+\lvert f_k\rvert)$ | Works in the units of the objective; a standard in L-BFGS solvers | Plateaus (the loss is flat, then drops again); noisy losses (SGD); very flat optima |
| Maximum iterations | $k = k_{\max}$ | Always terminates, protects your time | It says nothing about accuracy: if it fires, you have not converged. Treat the solver warning seriously |
Trade-offs. A tight tolerance gives accuracy but costs steps (and may be impossible because of rounding). A loose one is cheap but may stop too early. Combining a gradient test with a budget is the usual safe choice. In machine learning there is one more rule, early stopping: stop when the validation loss stops improving. That is about generalisation, not about optimisation, and it can fire long before any of the rules above.
Why do we need it?
Every iterative algorithm needs an exit. A good rule saves hours of useless computing, a bad one returns a wrong answer that looks finished.
Where is it used?
Solver options in SciPy (gtol, ftol, xtol, maxiter), scikit-learn (tol, max_iter and the ConvergenceWarning), and early stopping in deep-learning training loops.
How is it used?
Choose a tolerance for the main rule, set a generous maximum, and always read the exit status: did it stop because it converged or because it hit the budget? If it hit the budget, scale your features, raise the limit or improve the method.
A solver that stopped is not the same as a solver that succeeded. Always check why it stopped (a convergence flag or message). Also check scaling: tolerances on $\|\nabla f\|$ depend on the units of the loss, and tolerances on steps depend on the units of the parameters. When in doubt, use the relative forms.
Quick check: a training run's loss changes by less than $10^{-6}$ per step for the last 50 steps, but the gradient norm is still $5$. Should you trust "converged"?
No. A gradient of norm $5$ is not small, so the ground is still steep. A tiny loss change with a big gradient usually means the learning rate is far too small (or the steps are being clipped), so the walk is crawling. Check the learning rate before accepting the result.
Optimization tolerance: how close is close enough? core
When you ask a GPS for "a restaurant within 100 metres", you accept an answer that is not exactly at your feet. The tolerance $\varepsilon$ is the same idea for an optimizer: the error you agree to accept. It is a choice that trades accuracy against effort.
Each extra digit of accuracy has a cost, and the cost depends on the rate. For a sublinear method every new digit costs ten times more steps. For a linear method every new digit costs the same number of extra steps. For Newton every new digit costs almost nothing. And there is a hard wall: a computer stores only about 16 digits.
Steps needed to reach accuracy $\varepsilon$ (a "digit" is a factor of 10 in $\varepsilon$):
| $\varepsilon$ | Sublinear $5/\varepsilon$ | Linear, $\kappa=10$: $\kappa\ln(1/\varepsilon)$ | Newton $\approx\log_2(2\cdot\text{digits})$ |
|---|---|---|---|
| $10^{-2}$ | 500 | 47 | 2 |
| $10^{-4}$ | 50,000 | 93 | 3 |
| $10^{-6}$ | 5,000,000 | 139 | 4 |
| $10^{-8}$ | 500,000,000 | 185 | 4 |
(Linear column: $10\ln 100 = 46.05 \to 47$ steps, $10\ln10^4 = 92.1 \to 93$, $10\ln10^6 = 138.2\to139$, $10\ln10^8 = 184.2 \to 185$. These are step counts guaranteed by the bound $k \ge \kappa\ln(1/\varepsilon)$ for a starting gap of $1$. The Newton column is a rule of thumb: start with half a correct digit, double it each step, and count the steps until the digits reach the target. The sublinear column assumes $L\|\mathbf{x}_0-\mathbf{x}^\star\|^2/2 = 5$.)
A point $\mathbf{x}$ is an $\varepsilon$-optimal solution if $f(\mathbf{x}) - f^\star \le \varepsilon$. It is an $\varepsilon$-stationary point if $\|\nabla f(\mathbf{x})\| \le \varepsilon$ (the usual goal for non-convex problems, where a global optimum is out of reach).
- Absolute vs relative tolerance. An absolute test "$\text{quantity} \le \varepsilon$" depends on units. A relative test "$\text{quantity} \le \varepsilon\,(\text{scale})$" does not. Libraries often combine them: $\le \varepsilon_{\text{abs}} + \varepsilon_{\text{rel}}\cdot|\text{value}|$.
- The floating-point floor. Computers keep about 16 significant digits (machine epsilon $\approx 2.2\cdot10^{-16}$, see Numerical Linear Algebra). Near a minimum $f(\mathbf{x}^\star + \boldsymbol\delta) - f^\star \approx \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$, which is smaller than $10^{-16}|f|$ once $\|\boldsymbol\delta\| \lesssim \sqrt{10^{-16}}= 10^{-8}$. So by comparing function values you cannot locate a minimiser better than about 8 digits. Gradients can be used a little further, but a tolerance below about $10^{-12}$ is rarely meaningful.
- The statistical floor. In machine learning the loss itself is measured on a finite, noisy data set. Solving the training problem more accurately than the noise in the data justifies buys nothing (and an over-trained model can generalise worse). This is one reason why tolerances around $10^{-4}$ are common defaults.
Why do we need it?
"Optimal" can never be reached exactly, so "good enough" has to be defined. The tolerance turns an infinite process into a finite, budgeted one and makes results reproducible.
Where is it used?
Defaults such as scikit-learn's LogisticRegression(tol=1e-4, max_iter=100), SciPy's gtol=1e-5 for BFGS, and the "$\varepsilon$-approximate solution" statements in optimization papers ("we reach an $\varepsilon$-optimal point in $O(1/\varepsilon)$ steps").
How is it used?
Start with the library default. If the answer must be more precise (scientific computing), tighten it, but never below what floating point allows. If training is slow and accuracy beyond a few digits does not change your validation score, loosen it and save time.
Tighter is not always better. Demanding $10^{-12}$ from a model trained on noisy data wastes time and may over-fit. And remember what the tolerance is about: the test $\|\nabla f\|\le\varepsilon$ does not promise that the distance or the gap is $\le\varepsilon$ (it depends on $\mu$, see the gradient-norm section).
Quick check: an $O(1/\varepsilon)$ method takes 1000 steps to reach $\varepsilon = 0.01$. About how many for $\varepsilon = 0.001$? And for a linear method that takes 100 steps for $0.01$?
The $O(1/\varepsilon)$ method needs about ten times more: 10,000 steps. The linear method needs only a fixed number of extra steps per digit. If $100$ steps gave two digits it costs about $50$ per digit, so one more digit adds about 50 steps (about $150$ in total).
Recap, cheat sheet and practice
- Convergence can mean the points settle ($\|\mathbf{x}_k-\mathbf{x}^\star\|\to0$), the score settles ($f(\mathbf{x}_k)-f^\star\to0$) or the slope vanishes ($\|\nabla f(\mathbf{x}_k)\|\to0$). They are different yardsticks; on non-convex problems only the gradient one is realistic.
- The rate says how fast: sublinear ($1/k$: each new digit costs 10 times more steps), linear ($\rho^k$: a straight line on a log plot, each digit costs a fixed number of steps), superlinear, quadratic (the digits double every step).
- Smooth means no corners. $G$-Lipschitz bounds the slope ($\|\nabla f\|\le G$). $L$-smooth bounds the curvature: $\|\nabla f(\mathbf{x})-\nabla f(\mathbf{y})\|\le L\|\mathbf{x}-\mathbf{y}\|$, which gives the roof $f(\mathbf{y}) \le f(\mathbf{x}) + \nabla f^\top(\mathbf{y}-\mathbf{x}) + \frac L2\|\mathbf{y}-\mathbf{x}\|^2$ (for convex $f$ the two statements are equivalent), and for convex twice-differentiable $f$ means $\lambda_{\max}(\nabla^2 f)\le L$.
- Descent lemma: $f(\mathbf{x}^+) \le f(\mathbf{x}) - \eta(1-\frac{L\eta}2)\|\nabla f\|^2$; for $\eta \le 1/L$ the drop is at least $\frac\eta2\|\nabla f\|^2$. Progress is guaranteed for $0\lt\eta\lt 2/L$.
- Gradient descent: convex and $L$-smooth gives $f(\mathbf{x}_K)-f^\star \le \frac{\|\mathbf{x}_0-\mathbf{x}^\star\|^2}{2\eta K}$, i.e. $O(1/k)$. $\mu$-strongly convex too gives a linear rate $(1-\mu/L)^k = (1-1/\kappa)^k$, costing about $\kappa\ln(1/\varepsilon)$ steps.
- Newton converges quadratically near a solution ($e_{k+1}\le C e_k^2$), but each step costs a Hessian solve and it can misbehave far from the solution.
- The gradient norm is the practical progress measure: $\frac{\|\nabla f\|^2}{2L}\le f-f^\star\le\frac{\|\nabla f\|^2}{2\mu}$ and $\|\mathbf{x}-\mathbf{x}^\star\|\le\|\nabla f\|/\mu$. A small gradient does not imply closeness if $\mu$ is tiny.
- Stopping: gradient norm, step size, relative change in $f$, maximum iterations (and, in ML, early stopping). Combine them, and check why the solver stopped. A tolerance is a trade between accuracy and cost, bounded below by rounding error (about 16 digits) and by noise in the data.
Cheat sheet
| Idea | Statement | Remember as |
|---|---|---|
| Linear rate | $e_{k+1}\le\rho e_k$ | straight line on a log plot |
| Quadratic rate | $e_{k+1}\le Ce_k^2$ | digits double |
| $G$-Lipschitz | $\lvert f(\mathbf{x})-f(\mathbf{y})\rvert\le G\|\mathbf{x}-\mathbf{y}\|$ | speed limit on the output |
| $L$-smooth | $\|\nabla f(\mathbf{x})-\nabla f(\mathbf{y})\|\le L\|\mathbf{x}-\mathbf{y}\|$ | a roof of curvature $L$ |
| $\mu$-strongly convex | $f(\mathbf{y})\ge f(\mathbf{x})+\nabla f^\top(\mathbf{y}-\mathbf{x})+\frac\mu2\|\mathbf{y}-\mathbf{x}\|^2$ | a floor of curvature $\mu$ |
| Condition number | $\kappa=L/\mu$ | how stretched the bowl is |
| Descent lemma | $f(\mathbf{x}^+)\le f-\frac\eta2\|\nabla f\|^2$ for $\eta\le\frac1L$ | each step must pay for itself |
| GD, convex | $f(\mathbf{x}_K)-f^\star\le\frac{\|\mathbf{x}_0-\mathbf{x}^\star\|^2}{2\eta K}$ | sublinear, $O(1/k)$ |
| GD, strongly convex | $f(\mathbf{x}_k)-f^\star\le(1-\frac1\kappa)^k(f_0-f^\star)$ | linear, about $\kappa\ln\frac1\varepsilon$ steps |
| Best fixed step (quadratic) | $\eta=\frac{2}{L+\mu}$, factor $\frac{\kappa-1}{\kappa+1}$ | balance fastest and slowest directions |
| Newton | $\mathbf{x}^+=\mathbf{x}-H^{-1}\nabla f$ | jump to the bottom of the local parabola |
| Gradient sandwich | $\frac{\|\nabla f\|^2}{2L}\le f-f^\star\le\frac{\|\nabla f\|^2}{2\mu}$ | the slope bounds the gap |
import numpy as np
# ---- 1. The constants mu and L of a least-squares loss f(w) = (1/2n) ||Xw - y||^2
rng = np.random.default_rng(0)
n = 200
X = rng.normal(size=(n, 3)) * np.array([1.0, 2.0, 0.5]) # features with different scales
y = X @ np.array([1.0, -2.0, 0.5]) + 0.1 * rng.normal(size=n)
H = X.T @ X / n # the Hessian (constant for least squares)
mu, L = np.linalg.eigvalsh(H)[[0, -1]] # smallest and largest eigenvalue
print("mu =", round(mu, 3), " L =", round(L, 3), " kappa =", round(L / mu, 2))
# mu = 0.235 L = 3.995 kappa = 16.98
f = lambda w: 0.5 * np.mean((X @ w - y) ** 2)
grad = lambda w: X.T @ (X @ w - y) / n
w_star = np.linalg.solve(H, X.T @ y / n) # exact minimiser (normal equations)
# ---- 2. Gradient descent with eta = 1/L: check the descent lemma and the linear rate
eta = 1 / L
w = np.zeros(3)
gaps, lemma_ok = [f(w) - f(w_star)], True
for k in range(200):
g = grad(w)
w_new = w - eta * g
lemma_ok &= f(w_new) <= f(w) - 0.5 * eta * g @ g + 1e-12 # f(x+) <= f(x) - (eta/2)||grad||^2
w = w_new
gaps.append(f(w) - f(w_star))
gaps = np.array(gaps)
print("descent lemma held at every step:", lemma_ok)
print("worst one-step factor of the gap :", round((gaps[1:50] / gaps[:49]).max(), 4),
" <= 1 - mu/L =", round(1 - mu / L, 4))
# descent lemma held at every step: True
# worst one-step factor of the gap : 0.8857 <= 1 - mu/L = 0.9411
# ---- 3. Gradient-norm sandwich: ||g||^2/(2L) <= f - f* <= ||g||^2/(2 mu)
w_test = np.array([0.3, -0.2, 0.9]); g = grad(w_test); gap = f(w_test) - f(w_star)
print("sandwich:", g @ g / (2 * L) <= gap <= g @ g / (2 * mu)) # True
# ---- 4. Stopping rules: a relative gradient test with a budget
w = np.zeros(3); g0 = np.linalg.norm(grad(w))
for k in range(100000):
g = grad(w)
if np.linalg.norm(g) <= 1e-6 * g0:
print("gradient rule fired at step", k, "| distance to w* =", float(np.linalg.norm(w - w_star)))
break
w = w - eta * g
else:
print("hit the iteration budget: NOT converged")
# gradient rule fired at step 160 | distance to w* = 3.33e-05 (not 1e-6: the test is on the gradient)
# ---- 5. Newton squares the error: minimise f(x) = e^x - 2x (x* = ln 2)
x, xs = 0.0, np.log(2)
errs = []
for k in range(6):
errs.append(abs(x - xs))
x = x - (np.exp(x) - 2) / np.exp(x) # x - f'(x)/f''(x)
print(["%.1e" % e for e in errs])
print("e_{k+1} / e_k^2 :", [round(float(errs[i + 1] / errs[i] ** 2), 3) for i in range(4)])
# ['6.9e-01', '3.1e-01', '4.3e-02', '9.0e-04', '4.0e-07', '8.0e-14']
# e_{k+1} / e_k^2 : [0.639, 0.453, 0.493, 0.5]
1. A run has errors $1,\ 0.5,\ 0.25,\ 0.125,\ 0.0625,\dots$ What is the convergence rate?
2. Which statement about $L$-smoothness is correct?
3. $f$ is $L$-smooth with $L = 5$. For which step sizes does the descent lemma guarantee that one gradient step lowers $f$ (when $\nabla f\ne 0$)?
4. Gradient descent with $\eta = 1/L$ on a $\mu$-strongly convex, $L$-smooth function. About how many steps cut the gap by a factor $10^{-6}$ when $\kappa = L/\mu = 50$?
5. Newton's method has an error of $10^{-3}$ and its constant is $C\approx1$. What is the error after the next step?
6. A solver stops because $\|\nabla f(\mathbf{x}_k)\| \le 10^{-3}$. What can you safely conclude?
Practice problems
A. The error of a method is $e_k = 3\cdot0.8^k$. Which family? How many steps until $e_k\le10^{-6}$?
$e_{k+1}/e_k = 0.8$ is a constant, so it is linear with $\rho = 0.8$. We need $3\cdot0.8^k\le10^{-6}$, i.e. $0.8^k\le 3.33\cdot10^{-7}$, so $k \ge \dfrac{\ln(3\cdot10^6)}{\ln(1/0.8)} = \dfrac{14.91}{0.2231} = 66.8$. After 67 steps.
B. $f(x,y) = x^2 + 3xy + 5y^2$. Find the Hessian, $L$, $\mu$, $\kappa$, a safe step size, and a guaranteed number of steps to cut the gap by $10^{-6}$.
$f_{xx}=2$, $f_{xy}=3$, $f_{yy}=10$, so $H = \begin{bmatrix}2&3\\3&10\end{bmatrix}$ (constant). Its trace is $12$ and its determinant is $20-9=11$, so the eigenvalues solve $\lambda^2 - 12\lambda + 11=0$: $\lambda = 1$ and $\lambda = 11$. Both positive, so $f$ is strongly convex with $\mu = 1$, $L = 11$, $\kappa = 11$. A safe step is $\eta = 1/L = 1/11 \approx 0.091$. The gap shrinks by at least $1-1/11 = 10/11$ per step, so $k \ge 11\ln10^6 = 11\cdot13.82 = 151.97$: 152 steps suffice (a guarantee; the real run may be faster).
C. For the function in B, take $\mathbf{x} = (1, 0)$ and $\eta = 1/11$. Verify the descent lemma.
$\nabla f = (2x+3y,\ 3x+10y) = (2, 3)$, so $\|\nabla f\|^2 = 13$. The step: $\mathbf{x}^+ = (1 - 2/11,\ -3/11) = (9/11, -3/11)$. Values: $f(\mathbf{x}) = 1$, and $f(\mathbf{x}^+) = \frac{81}{121} + 3\cdot\frac{9}{11}\cdot\left(-\frac{3}{11}\right) + 5\cdot\frac{9}{121} = \frac{81 - 81 + 45}{121} = \frac{45}{121}\approx0.372$. The lemma promises $f(\mathbf{x}^+)\le 1 - \frac{\eta}{2}\cdot13 = 1 - \frac{13}{22} = \frac{9}{22}\approx 0.409$. And $0.372\le0.409$. ✓
D. Newton for $f(x) = x - \ln x$ ($x \gt 0$, minimiser $x^\star = 1$). Derive the Newton step, show the error squares exactly, and find a start where it fails.
$f'(x) = 1 - 1/x$, $f''(x) = 1/x^2$. The step: $x^+ = x - \dfrac{1 - 1/x}{1/x^2} = x - (x^2 - x) = 2x - x^2$. Then $1 - x^+ = 1 - 2x + x^2 = (1-x)^2$, so the error squares exactly: $e_{k+1} = e_k^2$. From $x_0 = 0.5$: $0.5,\ 0.75,\ 0.9375,\ 0.99609,\ 0.99998$, with errors $0.5,\ 0.25,\ 0.0625,\ 0.0039,\ 1.5\cdot10^{-5}$. But from $x_0 = 3$: $x_1 = 6 - 9 = -3$, outside the domain where $\ln$ exists. The error $e_0 = 2 \gt 1$, so squaring makes it bigger: Newton's fast rate needs a start close enough ($\lvert e_0\rvert \lt 1$ here).
E. You minimise $f(x) = 0.01x^2$ (so $\mu = L = 0.02$) and stop when $\lvert f'(x)\rvert \le 10^{-3}$. How far from the optimum can you be? What gradient tolerance guarantees a distance of at most $10^{-3}$?
$f'(x) = 0.02x$. The rule fires when $0.02\lvert x\rvert \le 10^{-3}$, i.e. $\lvert x\rvert \le 0.05$: you may be $0.05$, fifty times the tolerance. This agrees with the bound $\lvert x - x^\star\rvert \le \|\nabla f\|/\mu = 10^{-3}/0.02 = 0.05$. To guarantee distance $\le10^{-3}$ we need $\|\nabla f\| \le \mu\cdot10^{-3} = 2\cdot10^{-5}$.
F. Find the Lipschitz constant $G$ and the smoothness constant $L$ of $f(x) = \sin(2x)$. Is $\lvert x\rvert$ Lipschitz? Is it $L$-smooth?
$f'(x) = 2\cos 2x$, so $\lvert f'\rvert\le 2$: $G = 2$. $f''(x) = -4\sin 2x$, so $\lvert f''\rvert\le4$: $L = 4$ (a non-convex function: the eigenvalue condition is $\lvert f''\rvert\le L$). For $\lvert x\rvert$: the slope is $\pm1$, so it is 1-Lipschitz ($\lvert\,\lvert x\rvert-\lvert y\rvert\,\rvert\le\lvert x-y\rvert$). But its derivative jumps from $-1$ to $1$ at $0$, so no finite $L$ satisfies $\lvert f'(x)-f'(y)\rvert\le L\lvert x-y\rvert$ (take $x = \delta$, $y=-\delta$: left side $2$, right side $2L\delta\to0$). It is not smooth.
Convex Optimization
Convexity is the most valuable property a problem can have. In a convex problem there are no false valleys: every valley you can walk into is the valley. Where you start does not matter, and when the slope is zero you are done. This chapter shows what convexity is, how to check it, and why one short argument makes "local minimum = global minimum" true.
- Test whether a set is convex (the segment test) and whether a function is convex (the chord test)
- Know strict convexity, strong convexity (and how it links to the rate $(1-\mu/L)^k$ of Chapter 3.5) and concave functions
- Use Jensen's inequality and see why it makes log-likelihood, KL divergence and cross-entropy arguments work
- Use the first-order test (the tangent lies below the graph) and the second-order test (the Hessian is positive semi-definite)
- Understand deeply that in a convex problem every local minimum is a global minimum: prove it and see it
- Recognise convex problems, and say honestly which ML losses are convex and which are not
Where we are. You know what a minimum is (3.1), how to find stationary points (3.2), how gradient descent works (3.3) and how fast it converges (3.5). Now we ask which problems are friendly. This chapter uses the Hessian from the Calculus guide and positive (semi-)definite matrices and eigenvalues from the Linear Algebra guide.
Convex sets core
Pick any two points inside a shape and pull a rubber band straight between them. If the band always stays inside the shape, for every pair of points, the shape is convex. A disk, a square and a half-plane pass the test. A ring fails (pick two points on opposite sides: the band crosses the hole). A crescent, a star and two separate blobs fail too: the band leaves the shape.
"Convex" means: no dents, no holes, no separate pieces. It is the right shape for a feasible region (the set of allowed solutions, Chapter 3.1), because you can walk from any allowed point to any other in a straight line and never break a rule.
Segment test with numbers. The disk $x^2 + y^2 \le 4$. Take $P = (-2, 0)$ and $Q = (2, 0)$, both on the edge. The midpoint is $(0,0)$, which is inside. In general a point on the segment is $\theta P + (1-\theta)Q = (-2\theta + 2(1-\theta),\ 0) = (2 - 4\theta,\ 0)$ and $|2-4\theta|\le 2$ for $0\le\theta\le1$: always inside. ✓
The ring $1 \le x^2+y^2 \le 4$ with the same $P$ and $Q$: the midpoint $(0,0)$ has $x^2+y^2 = 0 \lt 1$, so it is in the hole. The band leaves the set: not convex.
Examples of convex sets: a line or a plane; a half-plane (everything on one side of a line); a ball $\{\mathbf{x} : \|\mathbf{x}-\mathbf{c}\|\le r\}$; a box (every coordinate between a lower and an upper limit); the set of non-negative vectors; the probability simplex $\{\mathbf{x}\ge0,\ \sum x_i = 1\}$; and any intersection of convex sets (for example a disk cut by a half-plane). Non-examples: a ring, a crescent, a star, two separate blobs, a set with a hole, and the union of two convex sets (in general).
A set $C\subseteq\mathbb{R}^n$ is convex if for all $\mathbf{x},\mathbf{y}\in C$ and every $\theta\in[0,1]$,
$$\theta\,\mathbf{x} + (1-\theta)\,\mathbf{y} \in C.$$The points $\theta\mathbf{x}+(1-\theta)\mathbf{y}$, for $\theta$ from $1$ to $0$, are exactly the straight segment from $\mathbf{y}$ to $\mathbf{x}$ ($\theta=1$ gives $\mathbf{x}$, $\theta = 0$ gives $\mathbf{y}$, $\theta=\tfrac12$ gives the midpoint). So convex means "the segment between any two points of the set stays in the set".
Why the main examples are convex.
- Half-space $\{\mathbf{x}: \mathbf{a}^\top\mathbf{x}\le b\}$: if $\mathbf{a}^\top\mathbf{x}\le b$ and $\mathbf{a}^\top\mathbf{y}\le b$ then $\mathbf{a}^\top(\theta\mathbf{x}+(1-\theta)\mathbf{y}) = \theta\,\mathbf{a}^\top\mathbf{x}+(1-\theta)\,\mathbf{a}^\top\mathbf{y} \le \theta b+(1-\theta)b = b$.
- Ball $\|\mathbf{x}-\mathbf{c}\|\le r$: by the triangle inequality, $\|\theta\mathbf{x}+(1-\theta)\mathbf{y}-\mathbf{c}\| = \|\theta(\mathbf{x}-\mathbf{c})+(1-\theta)(\mathbf{y}-\mathbf{c})\| \le \theta\|\mathbf{x}-\mathbf{c}\| + (1-\theta)\|\mathbf{y}-\mathbf{c}\|\le r$. (This works for every norm: the L1 diamond and the L-infinity square are convex too.)
- Intersection $C_1\cap C_2$: if $\mathbf{x},\mathbf{y}$ lie in both sets, the segment between them lies in $C_1$ (it is convex) and in $C_2$, so in the intersection. The same holds for any number of sets.
Why do we need it?
Convex feasible regions are the "no traps" regions: you can slide between feasible points without leaving, and a minimum found inside is a true minimum (the next sections). Non-convex regions can trap an algorithm in a corner of a dent.
Where is it used?
SVM margin constraints (half-spaces), L1 and L2 constrained regression (a diamond and a disk), probabilities that sum to 1 (the simplex, behind softmax and mixture weights), box constraints on parameters, non-negative matrix factorisation, and projected gradient descent (Chapter 3.17), which needs a well-defined "closest point" in the set.
How is it used?
Write each constraint as a half-space, ball or similar, and intersect them: the intersection is convex automatically. To test a strange set, try the segment test on a few pairs. One violating pair proves the set is not convex; passing a few pairs proves nothing.
One bad pair is enough; one good pair is not. To prove a set is convex you must show the segment test for all pairs (usually by an argument, as above). To prove it is not convex, a single pair whose segment leaves the set is a complete proof. Also note: the union of two convex sets is usually not convex, but the intersection always is.
Quick check: is the set $\{(x, y): y \ge x^2\}$ (everything on or above a parabola) convex? What about $\{(x, y): y \le x^2\}$?
The set above the parabola is convex (it is the epigraph of the convex function $x^2$: take $(-1,1)$ and $(1,1)$; every point $(t,1)$ between them has $1 \ge t^2$). The set below is not: $(-1,1)$ and $(1,1)$ are in it, but their midpoint $(0,1)$ has $1\le 0$ false, so the segment leaves the set.
Convex functions core
A convex function has a graph shaped like a bowl (or a smile). Stretch a string between any two points on the graph. For a convex function the string never dips below the graph: it lies above it, or touches it. For a function with a hump (like a sine wave) you can find two points whose string passes under the hump.
This is the same rubber-band test as for sets, applied to the region above the graph. In fact, a function is convex exactly when the region above its graph (called the epigraph) is a convex set.
Passing. $f(x) = x^2$ with $a = -1$ and $b = 3$, halfway ($\theta = \tfrac12$). The point between them is $\tfrac12(-1) + \tfrac12(3) = 1$ and $f(1) = 1$. The string at that point is at height $\tfrac12 f(-1) + \tfrac12 f(3) = \tfrac12(1) + \tfrac12(9) = 5$. Since $1 \le 5$ the graph is below the string. ✓
Failing. $f(x) = \sin x$ with $a = 0$ and $b = \pi$. The point between them is $\pi/2$ and $f(\pi/2) = 1$. The string is at height $\tfrac12\sin0 + \tfrac12\sin\pi = 0$. Now $1 \le 0$ is false: the graph is above the string, so $\sin$ is not convex.
Let $C$ be a convex set. A function $f: C \to \mathbb{R}$ is convex if for all $\mathbf{x},\mathbf{y}\in C$ and every $\theta\in[0,1]$,
$$f\big(\theta\mathbf{x} + (1-\theta)\mathbf{y}\big) \;\le\; \theta f(\mathbf{x}) + (1-\theta) f(\mathbf{y}).$$Left side: the height of the graph above the point between $\mathbf{x}$ and $\mathbf{y}$. Right side: the height of the string (the chord) above the same point.
Epigraph (awareness). The epigraph of $f$ is the set $\text{epi}\,f = \{(\mathbf{x}, t) : t \ge f(\mathbf{x})\}$, everything on or above the graph. $f$ is convex exactly when $\text{epi}\,f$ is a convex set. Reason: the string between two graph points $(\mathbf{x},f(\mathbf{x}))$ and $(\mathbf{y},f(\mathbf{y}))$ is the set of points $\big(\theta\mathbf{x}+(1-\theta)\mathbf{y},\ \theta f(\mathbf{x})+(1-\theta)f(\mathbf{y})\big)$, and "this point is in the epigraph" says exactly "$\theta f(\mathbf{x})+(1-\theta)f(\mathbf{y}) \ge f(\theta\mathbf{x}+(1-\theta)\mathbf{y})$". This is why convex functions and convex sets are one idea. A useful consequence: every sublevel set $\{\mathbf{x}: f(\mathbf{x})\le c\}$ of a convex function is a convex set (a bowl cut at any height is a convex region).
Convex functions need not be smooth: $\lvert x\rvert$ and $\max(0,x)$ are convex with corners.
Why do we need it?
Convexity is a property you can check that guarantees a friendly landscape: one valley, no traps. Without it, you cannot tell from the formula whether gradient descent will find the best answer or a poor one.
Where is it used?
The loss of linear regression, logistic regression and SVMs; the L1 and L2 penalties; and every "convex relaxation" in classical machine learning, where a hard non-convex problem is replaced by a convex one that can be solved reliably.
How is it used?
Do not test the definition directly on all pairs. Instead check one of the shortcuts later in this chapter (the Hessian is positive semi-definite, or the function is built from convex pieces with the allowed operations). Use the chord test as the picture and as a numerical sanity check.
The direction of the inequality matters. "Convex" means the graph is below the chord ($\le$). If the graph is above the chord you have a concave function (a hill). Also, the domain must itself be convex: $f(x) = 1/x$ is convex on $x \gt 0$ but not on all nonzero $x$ (the domain has a hole).
Quick check: $f(x) = x^2$, $a = 0$, $b = 4$, $\theta = 0.25$. Compute both sides of the convexity inequality.
The point is $0.25\cdot0 + 0.75\cdot4 = 3$ and $f(3) = 9$. The chord height is $0.25\cdot f(0) + 0.75\cdot f(4) = 0 + 0.75\cdot16 = 12$. Since $9 \le 12$ the inequality holds.
Strict convexity core
A convex function may still have flat or straight pieces: the graph of $\lvert x\rvert$ is two straight lines, and a chord drawn along one of them lies exactly on the graph. A strictly convex function has no straight pieces anywhere: every chord (except the two end points) lies strictly above the graph. It is curved everywhere, like $x^2$.
Why care? A flat bottom means a whole stretch of equally good answers. A strictly convex function has at most one best answer.
$x^2$ is strictly convex. Take two different points $a \ne b$ and $\theta = \tfrac12$. The gap between chord and graph is
$$\tfrac12 a^2 + \tfrac12 b^2 - \Big(\tfrac{a+b}{2}\Big)^2 = \frac{2a^2 + 2b^2 - a^2 - 2ab - b^2}{4} = \frac{(a-b)^2}{4} \gt 0.$$$\lvert x\rvert$ is not. Take $a = 1$, $b = 3$, $\theta = \tfrac12$. The point between is $2$, with $f(2) = 2$. The chord height is $\tfrac12(1) + \tfrac12(3) = 2$. The gap is $0$: the chord lies on the graph.
Strictly convex, yet flat at the bottom: $x^4$ is strictly convex (no straight pieces), even though its second derivative is $0$ at $x = 0$.
A convex function $f$ is strictly convex if for all $\mathbf{x}\ne\mathbf{y}$ and every $\theta\in(0,1)$ (strictly between $0$ and $1$)
$$f\big(\theta\mathbf{x} + (1-\theta)\mathbf{y}\big) \;\lt\; \theta f(\mathbf{x}) + (1-\theta) f(\mathbf{y}).$$Consequence: at most one minimiser. Suppose $\mathbf{x}\ne\mathbf{y}$ both minimise $f$, with the same value $m$. Then at the midpoint $f\big(\tfrac{\mathbf{x}+\mathbf{y}}{2}\big) \lt \tfrac12 m + \tfrac12 m = m$, which is lower than the minimum: a contradiction. So the minimiser, if it exists, is unique.
Strict convexity does not guarantee that a minimiser exists: $e^x$ is strictly convex, decreases toward $0$ as $x\to-\infty$, and never reaches it. (Also, "strictly convex" is stronger than "convex" but weaker than "strongly convex", next.)
Why do we need it?
It decides whether the answer is unique. With only convexity there can be a whole flat set of equally good solutions, and different runs (or different solvers) may return different ones.
Where is it used?
Least squares with collinear (duplicated) features has a flat valley of solutions, so it is convex but not strictly convex; adding an L2 penalty (ridge) makes it strictly convex with a unique answer. Logistic regression on separable data has no minimiser at all (the loss only approaches its lowest value).
How is it used?
Check whether the Hessian is positive definite everywhere (a sufficient test). If your problem is only convex, expect many solutions and use a tie-breaker such as regularisation, the minimum-norm solution, or early stopping.
Three levels, easy to mix up. Convex: chords above or on the graph. Strictly convex: chords strictly above (no straight pieces). Strongly convex: curved at least as much as a fixed parabola ($\mu \gt 0$, next). Every strongly convex function is strictly convex, and every strictly convex function is convex, but not the other way round ($x^4$ is strictly but not strongly convex; $\lvert x\rvert$ is just convex).
Quick check: is the function $f(x, y) = (x + y)^2$ strictly convex?
No. It is convex (a square of a linear function), but along the direction $(1,-1)$ it is constant: $f(t, -t) = 0$. The chord between $(1,-1)$ and $(2,-2)$ lies exactly on the graph (gap $0$), and every point of the line $x + y = 0$ is a minimiser. Convex, not strictly convex.
Strong convexity core
A convex function is a bowl. A strongly convex function is a bowl that is at least as curved as a fixed parabola, everywhere. Slide a parabola of curvature $\mu \gt 0$ under the graph so that it touches at any point you like: a strongly convex function never dips below it. This is the floor.
Chapter 3.5 gave the roof: a parabola of curvature $L$ above the graph (smoothness). Put them together and the function is squeezed in a sandwich between a floor with curvature $\mu$ and a roof with curvature $L$. The sandwich is why gradient descent converges geometrically: you can never be on a nearly flat stretch far from the bottom.
A check with numbers. Let $f(x) = x^2 + \sin x$. Then $f''(x) = 2 - \sin x$, which always lies between $1$ and $3$. So $\mu = 1$ and $L = 3$ (condition number $\kappa = 3$). At $x_0 = 0$: $f(0) = 0$ and $f'(0) = 2\cdot0 + \cos 0 = 1$. The floor is $0 + 1\cdot y + \tfrac{1}{2}y^2$ and the roof is $0 + 1\cdot y + \tfrac32 y^2$. At $y = 1$: floor $= 1.5$, roof $= 2.5$, and the function is $f(1) = 1 + \sin 1 = 1.841$. And $1.5\le1.841\le2.5$. ✓
A function that is convex but not strongly convex: $x^4/4$. At $x_0 = 0$ the tangent is flat and the floor would be $\tfrac{\mu}{2}y^2$. But $y^4/4 \lt \tfrac\mu2 y^2$ whenever $y^2 \lt 2\mu$, so the graph dips below the floor for every $\mu\gt0$: no strong convexity (its curvature $3x^2$ is $0$ at the origin).
Making any convex loss strongly convex: add an L2 penalty. If the loss $\ell(\mathbf{w})$ is convex, then $\ell(\mathbf{w}) + \frac\lambda2\|\mathbf{w}\|^2$ is $\lambda$-strongly convex (the penalty has Hessian $\lambda I$).
A differentiable $f$ is $\mu$-strongly convex ($\mu \gt 0$) if for all $\mathbf{x},\mathbf{y}$
$$f(\mathbf{y}) \;\ge\; f(\mathbf{x}) + \nabla f(\mathbf{x})^\top(\mathbf{y}-\mathbf{x}) + \frac{\mu}{2}\|\mathbf{y}-\mathbf{x}\|^2.$$Equivalent forms: (i) the chord version $f(\theta\mathbf{x}+(1-\theta)\mathbf{y}) \le \theta f(\mathbf{x})+(1-\theta)f(\mathbf{y}) - \frac\mu2\theta(1-\theta)\|\mathbf{x}-\mathbf{y}\|^2$; (ii) $f(\mathbf{x}) - \frac\mu2\|\mathbf{x}\|^2$ is still convex; (iii) if $f$ is twice differentiable, $\nabla^2 f(\mathbf{x}) \succeq \mu I$ for all $\mathbf{x}$, that is, every Hessian eigenvalue is at least $\mu$ (positive definiteness with a margin).
What it gives you.
- A unique minimiser exists. The floor parabola grows to $+\infty$ in every direction, so $f$ does too. A continuous function that tends to infinity at infinity has a minimum. And strong convexity implies strict convexity, so the minimiser is unique.
- Close in value means close in position. At $\mathbf{x}^\star$ the gradient is zero, so $f(\mathbf{x}) - f^\star \ge \tfrac\mu2\|\mathbf{x}-\mathbf{x}^\star\|^2$.
- The linear rate of Chapter 3.5. With $L$-smoothness too, gradient descent with $\eta=1/L$ satisfies $f(\mathbf{x}_k)-f^\star \le (1-\mu/L)^k\,(f(\mathbf{x}_0)-f^\star)$ (proved there). The ratio $\kappa = L/\mu \ge 1$ is the condition number: $\kappa=1$ is a perfectly round bowl, a large $\kappa$ is a long thin valley.
Why do we need it?
Plain convexity can leave you with a flat bottom: many answers, slow progress. Strong convexity guarantees one answer and fast, geometric convergence, with the speed controlled by one number, $\kappa$.
Where is it used?
Ridge regression and L2-regularised logistic regression and SVMs (the penalty $\lambda$ gives $\mu\ge\lambda$), the proofs for gradient descent, SGD and Newton's method on "nice" problems, and the theory that explains why feature scaling and weight decay speed up training.
How is it used?
Find the smallest Hessian eigenvalue (for a linear model, the smallest eigenvalue of $X^\top X/n$ plus $\lambda$). If it is zero or tiny, add L2 regularisation or remove redundant features. Then $\kappa = L/\mu$ predicts how many gradient steps you need: about $\kappa\ln(1/\varepsilon)$.
Strong convexity is a statement about the whole region. A function can be strongly convex near its minimum but lose it far away (like $\ln\cosh$). Logistic loss is never strongly convex on its own (its curvature fades in the tails): the L2 penalty is what supplies $\mu$. And remember the order: $\mu \le \lambda_{\min}(H)\le\lambda_{\max}(H)\le L$.
Quick check: a convex loss has Hessian eigenvalues in $[0, 8]$ everywhere. You add the penalty $\frac{\lambda}{2}\|\mathbf{w}\|^2$ with $\lambda = 0.5$. What are $\mu$, $L$ and $\kappa$ now?
The penalty adds $\lambda$ to every eigenvalue: they now lie in $[0.5, 8.5]$. So $\mu = 0.5$, $L = 8.5$ and $\kappa = 17$. Before, $\mu$ was $0$ and the condition number was unbounded.
Concave functions core
Turn a bowl upside down and you get a hill. A concave function has a graph shaped like a hill (a frown): the chord between two graph points lies below the graph. Concave is simply convex flipped upside down.
That one fact connects the two halves of optimization. Maximising a concave function is exactly the same problem as minimising a convex one (flip the sign). So everything in this chapter about minimising convex functions also covers maximising concave ones.
$f(x) = \ln x$ is concave. Take $a = 1$, $b = 4$ and the midpoint $2.5$. Graph: $\ln2.5 = 0.916$. Chord: $\tfrac12\ln1 + \tfrac12\ln4 = 0 + 0.693 = 0.693$. The graph is above the chord: $0.916 \ge 0.693$. ✓
And $-\ln x$ is convex. Maximising $\ln x$ over some range is the same as minimising $-\ln x$ there, and both give the same $x$.
Common concave functions: $-x^2$, $\ln x$, $\sqrt{x}$, $\min(x, 0)$, the entropy $-\sum p_i\ln p_i$, and the minimum of several linear functions. A linear function is both concave and convex.
A function $f$ is concave if $-f$ is convex, that is, for all $\mathbf{x},\mathbf{y}$ and $\theta\in[0,1]$
$$f\big(\theta\mathbf{x} + (1-\theta)\mathbf{y}\big) \;\ge\; \theta f(\mathbf{x}) + (1-\theta) f(\mathbf{y}).$$Everything flips: the chord is below the graph, the tangent line lies above the graph, a stationary point is a global maximum, and the Hessian is negative semi-definite ($\preceq 0$).
The standard translation. $\max_{\mathbf{x}} f(\mathbf{x})$ with $f$ concave $\;\Longleftrightarrow\;$ $\min_{\mathbf{x}}\,(-f(\mathbf{x}))$ with $-f$ convex. Same optimiser, and the optimal values differ only by a sign. A convex optimization problem is therefore "minimise a convex function" or equally "maximise a concave function", over a convex set.
Why do we need it?
Many natural goals are written as "maximise": likelihood, entropy, utility, profit. Knowing that concave means "convex upside down" lets you reuse every convex result for them without proving anything twice.
Where is it used?
Maximum likelihood: for logistic regression and linear regression with Gaussian noise the log-likelihood is concave in the weights, so "maximise log-likelihood" equals "minimise the convex negative log-likelihood", the usual loss. Entropy maximisation, the dual functions in Chapter 3.10 (always concave), and utility maximisation in economics.
How is it used?
Put the problem in the standard form "minimise a convex function": if you are asked to maximise $f$, check that $f$ is concave, then minimise $-f$ (this is why ML code minimises a negative log-likelihood). Not every log-likelihood is concave: for mixture models it is not.
Quick check: is $f(x) = -e^{x}$ concave or convex? What does "maximise $-e^x$ over $x \in [0,2]$" mean as a minimisation?
$e^x$ is convex, so $-e^x$ is concave. Maximising $-e^x$ is the same as minimising $e^x$ (convex). On $[0,2]$ the minimum of $e^x$ is at $x = 0$ (value 1), so the maximum of $-e^x$ is at $x = 0$ with value $-1$.
Jensen's inequality core
You have several numbers. You can average first, then apply $f$: $f(\text{average of } x)$. Or apply $f$ first, then average: $\text{average of } f(x)$. For a convex function (a bowl) these are not equal, and the order is fixed: average first gives the smaller answer.
Picture: put points on a bowl-shaped graph. Their centre of mass lies inside the shape they form, which is above the bowl. The centre of mass has height "average of $f(x)$" and horizontal position "average of $x$". The graph at that position is lower: $f(\text{average of } x)$.
Two numbers $x = 1$ and $x = 5$, each with probability $\tfrac12$. The average is $E[x] = 3$.
- $f(x) = x^2$: $f(E[x]) = 9$ and $E[f(x)] = \tfrac12(1) + \tfrac12(25) = 13$. So $9 \le 13$. ✓ (The gap, $4$, is exactly the variance of $x$.)
- $f(x) = e^x$: $f(3) = 20.09$ and $E[f(x)] = \tfrac12(2.718 + 148.41) = 75.57$. So $20.09 \le 75.57$. ✓
- $f(x) = \ln x$ (concave, so the inequality reverses): $f(3) = 1.099$ and $E[f(x)] = \tfrac12(0 + 1.609) = 0.805$. So $1.099 \ge 0.805$. ✓
Jensen's inequality. Let $f$ be convex and $X$ a random variable (or a list of values $x_i$ with probabilities $p_i \ge 0$, $\sum p_i = 1$). Then
$$f\big(\mathbb{E}[X]\big) \le \mathbb{E}\big[f(X)\big], \qquad\text{that is}\qquad f\Big(\sum_i p_i x_i\Big) \le \sum_i p_i f(x_i).$$For a concave $f$ the inequality reverses: $f(\mathbb{E}[X]) \ge \mathbb{E}[f(X)]$. For a strictly convex $f$, equality holds only when $X$ is constant.
Proof (using the tangent line; it is proved in the section on the first-order test). Let $m = \mathbb{E}[X]$ and let $g$ be the slope of $f$ at $m$ (at a corner, take the slope of any supporting line). A convex function lies above its tangent line, so $f(x) \ge f(m) + g\,(x - m)$ for every $x$. Take the expectation of both sides: $\mathbb{E}[f(X)] \ge f(m) + g\,(\mathbb{E}[X] - m) = f(m) + 0 = f(m)$. $\blacksquare$ (With two points and $p_1 = \theta$, Jensen is exactly the definition of convexity; for more points it follows by induction.)
Four standard uses.
- Variance is non-negative. With $f(x) = x^2$: $\mathbb{E}[X^2] \ge (\mathbb{E}[X])^2$. The gap is $\text{Var}(X)$.
- AM–GM. $\ln$ is concave, so $\ln\!\big(\tfrac{x_1+\dots+x_n}{n}\big) \ge \tfrac1n\sum\ln x_i = \ln\big((x_1\cdots x_n)^{1/n}\big)$: the arithmetic mean is at least the geometric mean.
- KL divergence is non-negative. For distributions $p, q$: $\text{KL}(p\|q) = \sum_i p_i\ln\frac{p_i}{q_i} = -\sum_i p_i\ln\frac{q_i}{p_i} \ge -\ln\sum_i p_i\frac{q_i}{p_i} = -\ln\sum_i q_i = -\ln 1 = 0$ (we used Jensen for the concave $\ln$ with weights $p_i$). Equality only when $q = p$. Consequently the cross-entropy $H(p,q) = -\sum p_i\ln q_i = H(p) + \text{KL}(p\|q)$ is smallest exactly when $q = p$: this is why training with cross-entropy pushes predicted probabilities toward the true ones.
- The log-likelihood bound (ELBO). If a model has hidden variables $z$, then $\ln p(x) = \ln\sum_z q(z)\frac{p(x,z)}{q(z)} \ge \sum_z q(z)\ln\frac{p(x,z)}{q(z)}$ (Jensen again). The right side is the lower bound that the EM algorithm and variational autoencoders maximise instead of the hard log of a sum.
Why do we need it?
It lets you move a function across an average, with a known direction of error. Most proofs in probabilistic ML need exactly this: "the log of an average is at least the average of the log".
Where is it used?
The EM algorithm and variational inference (the ELBO), the proof that KL divergence and cross-entropy are minimised at the true distribution, the bias-variance decomposition, bounds on losses of averaged models, and the AM–GM inequality.
How is it used?
Spot a convex or concave function applied to an average or a sum with weights that add to 1. Then pick the right direction: convex, $f(\text{average})\le\text{average}(f)$; concave, the opposite. The size of the gap grows with the spread of the values.
Direction and conditions. Jensen's inequality needs a convex $f$ and weights that are non-negative and sum to 1. For a concave $f$ it flips. It does not say the two sides are close: the gap grows with the spread of $X$. And for non-convex $f$ (like $\sin$) neither direction is guaranteed.
Quick check: three values $x = 0, 3, 6$ with equal probability. Verify Jensen for $f(x) = x^2$.
$\mathbb{E}[x] = 3$, so $f(\mathbb{E}[x]) = 9$. $\mathbb{E}[f(x)] = \tfrac13(0 + 9 + 36) = 15$. Indeed $9 \le 15$, and the gap $6$ equals the variance $\tfrac13(9 + 0 + 9) = 6$.
First-order characterisation: the tangent lies below the graph core
Stand at any point on a bowl and draw the tangent line there (the straight line that just touches the curve). For a bowl-shaped function the whole graph lies above that line, not only near the point of contact but everywhere. So the tangent is a guaranteed lower bound for the function, for every input.
Now put the contact point at the bottom of the bowl, where the slope is zero. The tangent is a horizontal line at the height of that point, and the whole graph lies above it. So the bottom point is the lowest point of the entire function. This single picture proves the big statement of this chapter: for a convex function, a flat spot is the global minimum.
Passing. $f(x) = x^2$ at $x_0 = 1$: $f(1) = 1$, slope $2$, tangent $T(y) = 1 + 2(y-1) = 2y - 1$. Then $f(y) - T(y) = y^2 - 2y + 1 = (y-1)^2 \ge 0$ for every $y$. For instance at $y = 3$: $f = 9 \ge T = 5$. ✓
Failing. $f(x) = \sin x$ at $x_0 = 0$: tangent $T(y) = y$. At $y = 1$: $f = 0.841 \lt T = 1$. The graph dips below its tangent, so $\sin$ is not convex.
A differentiable function $f$ (on a convex set) is convex if and only if for all $\mathbf{x},\mathbf{y}$
$$\boxed{\,f(\mathbf{y}) \;\ge\; f(\mathbf{x}) + \nabla f(\mathbf{x})^\top(\mathbf{y}-\mathbf{x})\,}$$The right side is the first-order Taylor approximation (the tangent line or tangent plane) at $\mathbf{x}$. Convexity says it never over-estimates $f$. Strict inequality for $\mathbf{y}\ne\mathbf{x}$ means strictly convex; adding $+\frac\mu2\|\mathbf{y}-\mathbf{x}\|^2$ on the right means $\mu$-strongly convex.
Proof (convex $\Rightarrow$ tangent below). Put $\mathbf{d} = \mathbf{y}-\mathbf{x}$ and use convexity with a small weight $\theta$ on $\mathbf{y}$: $f(\mathbf{x}+\theta\mathbf{d}) \le (1-\theta)f(\mathbf{x}) + \theta f(\mathbf{y})$. Move $f(\mathbf{x})$ to the left and divide by $\theta$: $\dfrac{f(\mathbf{x}+\theta\mathbf{d}) - f(\mathbf{x})}{\theta} \le f(\mathbf{y}) - f(\mathbf{x})$. As $\theta\to0$ the left side tends to the directional derivative $\nabla f(\mathbf{x})^\top\mathbf{d}$. So $\nabla f(\mathbf{x})^\top\mathbf{d}\le f(\mathbf{y}) - f(\mathbf{x})$. $\blacksquare$
Proof (tangent below $\Rightarrow$ convex). Let $\mathbf{z} = \theta\mathbf{x}+(1-\theta)\mathbf{y}$. Apply the tangent inequality at $\mathbf{z}$ twice: $f(\mathbf{x}) \ge f(\mathbf{z}) + \nabla f(\mathbf{z})^\top(\mathbf{x}-\mathbf{z})$ and $f(\mathbf{y}) \ge f(\mathbf{z}) + \nabla f(\mathbf{z})^\top(\mathbf{y}-\mathbf{z})$. Multiply the first by $\theta$ and the second by $1-\theta$ and add. Since $\theta(\mathbf{x}-\mathbf{z}) + (1-\theta)(\mathbf{y}-\mathbf{z}) = \mathbf{0}$, the gradient terms cancel: $\theta f(\mathbf{x}) + (1-\theta)f(\mathbf{y}) \ge f(\mathbf{z})$. $\blacksquare$
The big consequence. If $\nabla f(\mathbf{x}^\star) = \mathbf{0}$ then the inequality gives $f(\mathbf{y}) \ge f(\mathbf{x}^\star) + \mathbf{0}^\top(\mathbf{y} - \mathbf{x}^\star) = f(\mathbf{x}^\star)$ for every $\mathbf{y}$. So for a differentiable convex function, a point is a global minimiser if and only if its gradient is zero. No second-derivative test is needed, and no distinction between local and global.
Why do we need it?
It turns the local information at one point (value and slope) into a global fact (a lower bound everywhere). That is what makes "the gradient is zero, so we are done" a theorem for convex functions and only a hope for others.
Where is it used?
The convergence proof of gradient descent in Chapter 3.5, Jensen's inequality (above), duality and the cutting-plane / subgradient methods (every tangent is a valid lower bound on the optimum), and the "stationary point means optimum" argument for linear and logistic regression.
How is it used?
To prove a point optimal, show the gradient is zero there and the function is convex. To bound how far you are from the optimum: $f^\star \ge f(\mathbf{x}) + \nabla f(\mathbf{x})^\top(\mathbf{x}^\star - \mathbf{x}) \ge f(\mathbf{x}) - \|\nabla f(\mathbf{x})\|\,D$ whenever the optimum is known to lie within a distance $D$ of $\mathbf{x}$: a certificate you can compute at a single point.
It must hold for every pair. One tangent that stays below the graph is not a proof; the inequality must hold at every point $\mathbf{x}$ for every $\mathbf{y}$. And a zero gradient only means "global minimum" when the function is convex: for a non-convex function the tangent at a flat spot can cross the graph (a saddle, a local minimum).
For non-differentiable convex functions (like $\lvert x\rvert$) the tangent line is replaced by any supporting line, a subgradient; the same inequality holds with that slope.
Quick check: $f(x) = e^x$. Write the tangent at $x_0 = 0$ and verify the inequality at $y = 1$ and $y = -1$.
$f(0) = 1$, $f'(0) = 1$, so $T(y) = 1 + y$. At $y = 1$: $e = 2.718 \ge 2$. At $y = -1$: $e^{-1} = 0.368 \ge 0$. Both hold. (In fact $e^y \ge 1+y$ for every $y$, a famous inequality.)
Second-order characterisation: the Hessian is positive semi-definite core
The first derivative is the slope. A function is convex when its slope never decreases as you walk to the right: the graph keeps bending upward, like a smile. "The slope never decreases" is the same as "the second derivative is never negative": $f''(x) \ge 0$ everywhere.
With several inputs, walk along any straight line through the landscape and look at how the ground bends along that line. The function is convex if it bends upward or stays straight in every direction and at every point. The Hessian matrix collects the bending in all directions, and "bends up or straight in every direction" is exactly "the Hessian is positive semi-definite". One direction of downward bending (a negative eigenvalue) breaks convexity: that is a saddle.
- $f(x) = e^x$: $f'' = e^x \gt 0$ everywhere. Convex.
- $f(x) = x^4/4 - x^2$: $f'' = 3x^2 - 2$, which is negative for $\lvert x\rvert \lt \sqrt{2/3} = 0.816$. Not convex (it bends down in the middle).
- $f(x,y) = x^2 + xy + y^2$: Hessian $\begin{bmatrix}2&1\\1&2\end{bmatrix}$, eigenvalues $3$ and $1$ (trace $4$, determinant $3$). Both positive: convex, and even strongly convex with $\mu = 1$.
- $f(x,y) = x^2 + 3xy + y^2$: Hessian $\begin{bmatrix}2&3\\3&2\end{bmatrix}$, eigenvalues $5$ and $-1$ (trace $4$, determinant $-5$). One negative eigenvalue: not convex. Check along the direction $(1,-1)$: $f(t,-t) = t^2 - 3t^2 + t^2 = -t^2$, which bends down.
Let $f$ be twice differentiable on an open convex set. Then
$$f \text{ is convex} \iff \nabla^2 f(\mathbf{x}) \succeq 0 \ \text{ for every } \mathbf{x}.$$"$\succeq 0$" means positive semi-definite: $\mathbf{v}^\top\nabla^2 f(\mathbf{x})\,\mathbf{v}\ge0$ for every direction $\mathbf{v}$, equivalently all eigenvalues $\ge 0$. In one variable this is just $f''(x)\ge0$.
- $\nabla^2 f(\mathbf{x}) \succ 0$ (all eigenvalues $\gt0$) everywhere $\Rightarrow$ strictly convex. (Not "if and only if": $x^4$ is strictly convex but $f''(0) = 0$.)
- $\nabla^2 f \succeq \mu I$ $\iff$ $\mu$-strongly convex. And $0 \preceq \nabla^2 f \preceq LI$ means convex and $L$-smooth (Chapter 3.5).
Why it works. In one variable: $f''\ge0$ means $f'$ is non-decreasing, so $f(y) - f(x) - f'(x)(y-x) = \int_x^y \big(f'(t) - f'(x)\big)\,dt \ge 0$, which is the tangent test. In many variables, restrict $f$ to a line: $g(t) = f(\mathbf{x}+t\mathbf{v})$. Then $g''(t) = \mathbf{v}^\top\nabla^2 f(\mathbf{x}+t\mathbf{v})\,\mathbf{v}$. A function is convex exactly when it is convex along every line, so we need $g''\ge0$ for every point and every direction $\mathbf{v}$: that is the definition of positive semi-definite.
Three important ML cases.
- Quadratic $\frac12\mathbf{x}^\top A\mathbf{x} + \mathbf{b}^\top\mathbf{x} + c$ ($A$ symmetric): the Hessian is the constant $A$, so it is convex iff $A\succeq0$.
- Least squares $\frac{1}{2n}\|X\mathbf{w}-\mathbf{y}\|^2$: Hessian $X^\top X/n$, and $\mathbf{v}^\top X^\top X\mathbf{v} = \|X\mathbf{v}\|^2\ge0$ for every $\mathbf{v}$. Always convex.
- Logistic loss: Hessian $\frac1n X^\top D X$ with $D = \text{diag}\big(\sigma_i(1-\sigma_i)\big)$, where each $\sigma_i(1-\sigma_i)\in(0,\tfrac14]$. Then $\mathbf{v}^\top X^\top DX\mathbf{v} = \sum_i D_{ii}(\mathbf{x}_i^\top\mathbf{v})^2\ge0$. Always convex.
Why do we need it?
It is the practical convexity test. The chord and tangent tests need all pairs of points; this one needs only the eigenvalues of one matrix at each point, which you can compute or bound.
Where is it used?
Proving that linear regression, logistic regression and their regularised versions are convex; checking a custom loss (compute the Hessian symbolically or numerically); and diagnosing why a training problem is hard (negative eigenvalues of the Hessian are directions of negative curvature).
How is it used?
Differentiate twice, or use automatic differentiation. Then show $\mathbf{v}^\top H\mathbf{v}\ge0$ by algebra (a sum of squares), or check the eigenvalues numerically at many points as a sanity check (a numerical check at sample points can find non-convexity, but cannot prove convexity).
Sampled eigenvalues can disprove but not prove. Finding one point with a negative Hessian eigenvalue proves the function is not convex there. Finding none on a grid does not prove convexity everywhere: use an algebraic argument (a sum of squares, composition rules) for that. Also, positive definite everywhere is sufficient for strict convexity but not necessary ($x^4$).
Quick check: is $f(x, y) = x^2 - xy + y^2$ convex? Is $f(x,y) = x^2 + 2xy + y^2$ strictly convex?
First: Hessian $\begin{bmatrix}2&-1\\-1&2\end{bmatrix}$, trace $4$, determinant $3$, eigenvalues $3$ and $1$: positive definite, so strictly (even strongly) convex. Second: $f = (x+y)^2$ with Hessian $\begin{bmatrix}2&2\\2&2\end{bmatrix}$, eigenvalues $4$ and $0$: positive semi-definite, so convex, but a zero eigenvalue (along $(1,-1)$) means it is not strictly convex.
Global vs local minima: in a convex problem, local $=$ global core
A local minimum is the bottom of a valley: nothing nearby is lower. A global minimum is the bottom of the deepest valley. On a landscape with several valleys, a ball can settle in a shallow one and stop, even though a much deeper one exists elsewhere. That is the central worry of optimization.
Now take a convex function. It has only one valley. Here is why. Suppose you are at the bottom of a valley (nothing nearby is lower) and somewhere else there is a point $y$ that is strictly lower. Draw the straight string from your point $x$ to $y$. Convexity says the graph lies below the string. But the string itself starts at your height and goes down toward $y$. So just next to you, on the way to $y$, the graph is below the string, which is below your height: there are lower points arbitrarily close to you. That contradicts "nothing nearby is lower". So no such $y$ exists: you were already at the global minimum.
With numbers. $f(x) = x^2$. Suppose someone claims $x = 2$ is a local minimum, and points out $y = 0$ with $f(0) = 0 \lt f(2) = 4$. Walk from $x$ toward $y$ by a fraction $\theta$: $z = (1-\theta)\cdot2 + \theta\cdot0 = 2 - 2\theta$. Convexity gives $f(z) \le (1-\theta)\cdot4 + \theta\cdot0 = 4 - 4\theta \lt 4$. For $\theta = 0.01$: $z = 1.98$ and $f(z) = 3.92\lt 4$. For $\theta=0.0001$: $z = 1.9998$, $f(z)\approx 3.9992 \lt 4$. Points closer and closer to $x = 2$ are lower, so $x = 2$ is not a local minimum.
A non-convex counter-example. $f(x) = x^4/4 - x^2 + 0.3x$ has a right valley with bottom near $x = 1.33$ ($f\approx-0.59$) and a deeper left valley near $x = -1.48$ ($f\approx-1.43$). The point $x = 1.33$ is a local minimum, and $y = -1.48$ is lower. The string from $x$ to $y$ lies below the graph in the middle (there is a hump between), so the argument above breaks, and nothing forces a lower point near $x$.
A point $\mathbf{x}^\star$ is a local minimiser of $f$ over a feasible set $C$ if there is a radius $r\gt0$ such that $f(\mathbf{x}^\star)\le f(\mathbf{x})$ for every feasible $\mathbf{x}$ with $\|\mathbf{x}-\mathbf{x}^\star\|\le r$. It is a global minimiser if $f(\mathbf{x}^\star)\le f(\mathbf{x})$ for every feasible $\mathbf{x}$.
Theorem. Let $f$ be convex and $C$ a convex set. Then every local minimiser of $f$ over $C$ is a global minimiser.
Proof (by contradiction). Let $\mathbf{x}$ be a local minimiser, with radius $r$, and suppose some feasible $\mathbf{y}$ has $f(\mathbf{y}) \lt f(\mathbf{x})$.
- For $\theta\in(0,1]$ let $\mathbf{z}_\theta = (1-\theta)\mathbf{x} + \theta\mathbf{y}$. It is feasible because $C$ is convex (the segment stays in $C$).
- By convexity of $f$: $f(\mathbf{z}_\theta) \le (1-\theta)f(\mathbf{x}) + \theta f(\mathbf{y}) = f(\mathbf{x}) - \theta\big(f(\mathbf{x}) - f(\mathbf{y})\big) \lt f(\mathbf{x})$, because $f(\mathbf{x})-f(\mathbf{y})\gt 0$ and $\theta\gt0$.
- The point $\mathbf{z}_\theta$ is at distance $\|\mathbf{z}_\theta - \mathbf{x}\| = \theta\|\mathbf{y}-\mathbf{x}\|$ from $\mathbf{x}$. Choose $\theta$ so small that this is at most $r$.
- Then $\mathbf{z}_\theta$ is feasible, within the radius $r$ of $\mathbf{x}$, and strictly lower than $\mathbf{x}$: this contradicts "$\mathbf{x}$ is a local minimiser". So no such $\mathbf{y}$ exists. $\blacksquare$
Two companions. (1) The set of global minimisers is convex: if $\mathbf{x},\mathbf{y}$ both have the minimum value $m$ then every point between them has value $\le m$ by convexity, hence exactly $m$. (2) If $f$ is strictly convex the global minimiser is unique (previous sections).
What it needs. Both ingredients: a convex function and a convex feasible set (step 1 uses the set). With a non-convex constraint set, even a convex objective can have bad local minima: minimising $x^2$ over the two intervals $[-3,-2]\cup[1,2]$ has a local minimum at $x = -2$ (value 4) that is not global (the global one is $x=1$, value 1).
Why do we need it?
It removes the biggest fear in optimization: getting stuck. For a convex problem, any algorithm that stops at a point where nothing nearby is lower has found the best answer in the whole space.
Where is it used?
Everywhere a model is "convex": linear and logistic regression, SVMs, lasso, ridge. It is why these models reach the same minimal loss on every run, whatever the initialisation (and the same weights whenever the minimiser is unique), and why their optimisers need no restarts.
How is it used?
First show the problem is convex (a convex function over a convex set) using the tests in this chapter. Then any solver that converges to a stationary point returns the global optimum. If the problem is not convex, plan for several random starts and compare the final values.
Four traps. (1) The theorem needs a convex function and a convex set. (2) "No bad local minima" does not mean "fast": a flat convex bowl is still slow (Chapter 3.5). (3) "A global minimiser exists" is a separate question: $e^x$ is convex and has no minimiser at all. (4) For a non-convex problem, a run that stops where the gradient is zero may be at a local minimum, a saddle or a plateau; compare several random starts before trusting it.
Quick check: where exactly does the proof use convexity of the function, and where does it use convexity of the feasible set?
Step 1 uses the set: the point $\mathbf{z}_\theta$ between $\mathbf{x}$ and $\mathbf{y}$ must be feasible, which holds because the segment lies in $C$. Step 2 uses the function: the chord inequality $f(\mathbf{z}_\theta)\le(1-\theta)f(\mathbf{x})+\theta f(\mathbf{y})$.
Why convex optimization is special core
Think of finding the lowest point of a country. If the whole country is one big smooth bowl, you can drop a ball anywhere and it will find the bottom. You do not need a map, a lucky start or many attempts. If the country is a mountain range, the ball stops in whichever valley it fell into and you can never be sure it is the lowest.
Convexity is the property that turns a hard search into a reliable calculation. It buys you several guarantees at once, and each one removes a worry that non-convex problems keep alive.
Same problem, four methods. Fit a regularised logistic regression from the start $(-3, 3)$ with plain gradient descent, momentum, Nesterov momentum and Adam. With the settings of the widget below (400 steps) all four end at the same weights $(1.033, 0.974)$ with loss $0.2095$ (to 3 digits). Run the same four methods on the non-convex Himmelblau function from the start $(1, -1)$ with the widget's settings (500 steps): gradient descent, momentum and Nesterov reach the minimum near $(3.58, -1.85)$, but Adam reaches a different minimum, $(3, 2)$. This depends on the start and on the settings (other step sizes can change who ends where); the point is that on a convex problem there is only one place to end.
Compared with a general optimization problem, a convex problem (a convex function over a convex set) gives you:
- Local $=$ global. Any local minimum is a global one (the previous section).
- Optimality is checkable. For a differentiable convex $f$ (unconstrained), $\nabla f(\mathbf{x}^\star)=\mathbf{0}$ if and only if $\mathbf{x}^\star$ is a global minimiser. With constraints the condition becomes $\nabla f(\mathbf{x}^\star)^\top(\mathbf{y}-\mathbf{x}^\star)\ge0$ for every feasible $\mathbf{y}$ (you will see it as the KKT conditions in Chapter 3.9, which are necessary and sufficient for convex problems under mild conditions).
- Any sensible descent method works, from any starting point: gradient descent, momentum, Newton, quasi-Newton, coordinate descent, ... with suitable step sizes they converge to the global optimum (for smooth convex problems; a method must also suit the problem, for example coordinate descent can stall on non-smooth ones). No restarts, no clever initialisation.
- You know how fast. The rates of Chapter 3.5 are guarantees: $O(1/k)$ for convex $L$-smooth, linear for strongly convex.
- Duality works well (Chapter 3.10). Tangent planes lie below the graph, and duality turns such lower bounds into a computable lower bound on the optimal value. For convex problems the best lower bound usually matches the optimum (strong duality). So a solver can certify "I am within $10^{-6}$ of the true optimum".
- The solution set is nice: convex, and a single point when strictly convex.
- Efficient algorithms exist: for the standard classes (linear, quadratic, second-order-cone and semidefinite programs) interior-point methods find an $\varepsilon$-accurate global solution in polynomial time.
| Convex problem | General (non-convex) problem | |
|---|---|---|
| Where you start | Does not matter | Decides which valley you reach |
| Flat gradient means | Global minimum | Minimum, maximum or saddle |
| Different solvers | Agree on the optimal value | May disagree |
| Convergence rate | Proven (3.5) | Usually only "reaches a stationary point" |
| Quality certificate | Yes (duality gap) | No |
Why do we need it?
Reliability. A convex model can be trained, shipped and reproduced without worrying about unlucky initialisation or "did it find the best weights?". That is a large part of why classical ML (regression, SVMs, lasso) is so dependable.
Where is it used?
Linear and logistic regression, SVMs, ridge and lasso, convex relaxations of hard problems (the L1 norm replacing a count of non-zeros, the nuclear norm replacing a matrix rank), portfolio optimisation, optimal control and the sub-problems inside bigger non-convex methods.
How is it used?
Whenever you can, formulate your problem as convex (possibly by relaxing it) and use a standard solver. When a problem is truly non-convex, you can often still use convex pieces: alternating minimisation (convex in each block), or convex sub-problems inside each iteration.
Special does not mean easy or fast. A convex problem can be huge, ill-conditioned, or have no minimiser at all. And in deep learning, the losses are not convex; the lessons of this chapter then serve as intuition (and as tools for the convex pieces inside), not as guarantees. Also, "convex" depends on what you optimise over: a loss that is non-convex in the network weights can be convex in a subset of them (for example in the last layer, given fixed features).
Quick check: a friend trains a convex model twice with different random initialisations and gets two different weight vectors with the same loss. Is that a contradiction?
No. A convex function can have a whole set of equally good minimisers (when it is not strictly convex, for example with duplicated features). The set is convex and all its points have the same minimal loss. Local = global is satisfied. If the loss values differed, that would indicate that training had not converged yet.
How to recognise a convex problem core
Nobody checks the chord test on a real loss function. Instead you build the loss from convex Lego bricks joined by safe operations. If every brick is convex and every join is safe, the whole thing is convex, automatically. This is exactly how the modelling tools for convex optimization (like CVXPY) check your problem.
Bricks: straight lines, squares $x^2$, the absolute value $\lvert x\rvert$, $e^x$, $-\ln x$, norms, the ReLU $\max(0,x)$, the log-sum-exp. Safe joins: add them (with non-negative weights), take the largest of them, or plug an affine map $A\mathbf{x}+\mathbf{b}$ into them.
Lasso is convex. $f(\mathbf{w}) = \tfrac{1}{2n}\|X\mathbf{w}-\mathbf{y}\|^2 + \lambda\|\mathbf{w}\|_1$ with $\lambda\ge0$. The map $\mathbf{w}\mapsto X\mathbf{w}-\mathbf{y}$ is affine, and the square of the norm is convex, so the first term is convex (affine composition). $\|\mathbf{w}\|_1 = \sum\lvert w_j\rvert$ is a sum of convex functions. Adding the two with non-negative weights keeps convexity. Done, in two lines.
Logistic regression is convex. For a label $y\in\{0,1\}$ and score $z = \mathbf{w}^\top\mathbf{x}$ the loss is $-y\ln\sigma(z) - (1-y)\ln(1-\sigma(z))$. Since $1-\sigma(z) = \frac{e^{-z}}{1+e^{-z}}$, this simplifies to $\ln(1+e^{z}) - y\,z$. Check $y=1$: $\ln(1+e^z)-z = \ln(1+e^{-z})$ ✓; $y=0$: $\ln(1+e^z)$ ✓. The softplus $\ln(1+e^z)$ is convex, $-yz$ is linear, and $z=\mathbf{w}^\top\mathbf{x}$ is affine in $\mathbf{w}$. So each example's loss is convex in $\mathbf{w}$, and so is their sum.
Not safe: the difference of convex functions ($x^2-\lvert x\rvert$ has a concave kink at $0$), the minimum of convex functions ($\min(x^2,1)$ is not convex), a negative weight, and products ($x\cdot x\cdot x$).
Operations that preserve convexity (each one with its one-line reason):
- Non-negative combinations. If $f_1,\dots,f_m$ are convex and $a_i\ge0$, then $\sum a_if_i$ is convex. Why: multiply each chord inequality by $a_i\ge0$ (the direction is kept) and add them.
- Affine substitution. If $f$ is convex then $g(\mathbf{x}) = f(A\mathbf{x}+\mathbf{b})$ is convex. Why: $A(\theta\mathbf{x}+(1-\theta)\mathbf{y})+\mathbf{b} = \theta(A\mathbf{x}+\mathbf{b}) + (1-\theta)(A\mathbf{y}+\mathbf{b})$, so $g$'s chord test is $f$'s chord test at the transformed points.
- Pointwise maximum. $\max(f_1,\dots,f_m)$ is convex. Why: $f_i(\theta\mathbf{x}+(1-\theta)\mathbf{y}) \le \theta f_i(\mathbf{x}) + (1-\theta)f_i(\mathbf{y}) \le \theta\max_j f_j(\mathbf{x}) + (1-\theta)\max_jf_j(\mathbf{y})$ for every $i$; take the max over $i$ on the left.
- Composition rule. $g(h(\mathbf{x}))$ is convex if $h$ is convex and $g$ is convex and non-decreasing (for example $e^{h(\mathbf{x})}$, or $h(\mathbf{x})^2$ when $h\ge0$). Why (smooth one-variable case): $(g\circ h)'' = g''(h)\,h'^2 + g'(h)\,h''\ge0$ since $g''\ge0$, $g'\ge0$ and $h''\ge0$.
Standard bricks: affine functions; $x^2$ and $\|\mathbf{x}\|^2$; every norm (L1, L2, L-infinity) and $\|A\mathbf{x}-\mathbf{b}\|$; $\lvert x\rvert$, $\max(0,x)$; $e^x$; $-\ln x$ (for $x\gt0$); $x\ln x$; the softplus; the log-sum-exp $\ln\sum_i e^{x_i}$; and $\mathbf{x}^\top A\mathbf{x}$ with $A\succeq0$.
A convex optimization problem has the form
$$\min_{\mathbf{x}}\ f_0(\mathbf{x}) \quad\text{subject to}\quad f_i(\mathbf{x}) \le 0\ \ (i=1,\dots,m),\qquad A\mathbf{x}=\mathbf{b},$$with $f_0$ and every $f_i$ convex, and only affine equality constraints. Then the feasible set is convex (an intersection of sublevel sets of convex functions and flat sets) and the objective is convex over it. (In Chapter 3.7 we will write these constraints as $g_i(\mathbf{x})\le0$ and $h_j(\mathbf{x}) = 0$.) A constraint such as $\|\mathbf{x}\|_2 = 1$ (a sphere, not a ball) is not affine and makes the problem non-convex.
Why do we need it?
Checking the definition on every pair of points is impossible. Rules let you certify convexity by reading the formula, quickly and with no mistakes, and they tell you how to rewrite a problem so that it becomes convex.
Where is it used?
Modelling tools such as CVXPY (they apply exactly these rules, "disciplined convex programming"), proofs that regularised regression, logistic regression, SVMs and softmax regression are convex, and designing new losses (a safe way to build a new convex loss is to add and take maxima of known ones).
How is it used?
Write the objective as a tree of operations on bricks. Check each node against the four rules. If a node breaks a rule (a difference, a minimum, a negative weight), try to rewrite it, or accept that the problem is non-convex and plan accordingly.
The rules are sufficient, not necessary. If a rule breaks, the function may still be convex ($x^2 - \lvert x\rvert + \lvert x\rvert$ is just $x^2$). But you then need another argument (the Hessian test). Also watch the sign: $-\ln x$ is convex, $\ln x$ is concave; and in a composition $g(h(\mathbf{x}))$ the outer function must be non-decreasing (the square of a convex function that takes negative values, like $(x^2-1)^2$, is not convex).
Quick check: is $f(\mathbf{w}) = \sum_i \max\big(0,\,1 - y_i\,\mathbf{w}^\top\mathbf{x}_i\big) + \frac\lambda2\|\mathbf{w}\|^2$ (the SVM objective) convex?
Yes. $1 - y_i\mathbf{w}^\top\mathbf{x}_i$ is affine in $\mathbf{w}$; $\max(0,\cdot)$ of an affine function is a maximum of two affine functions, hence convex; the sum over $i$ is a non-negative combination; $\frac\lambda2\|\mathbf{w}\|^2$ is convex for $\lambda\ge0$. All the rules apply. (Moreover it is $\lambda$-strongly convex for $\lambda\gt0$.)
Convex and non-convex machine-learning losses core
"Is my loss convex?" decides what you can promise. For linear models the answer is almost always yes: the prediction is a linear function of the weights, and the loss is a convex function of that prediction, so the whole loss is convex in the weights (affine substitution into a convex function). The moment you put a non-linear layer with learned weights between the input and the output (a neural network), the answer becomes no.
Honesty matters here. Non-convex does not mean "hopeless": neural networks are trained successfully every day. It means there are no guarantees, and the lessons of this chapter become intuition instead of promises.
A one-line proof that something is not convex: matrix factorisation in the smallest case. Fit the number $3$ by a product $ab$: $f(a,b) = (ab-3)^2$. The points $A = (\sqrt3,\sqrt3)$ and $B = (-\sqrt3,-\sqrt3)$ both give $ab = 3$, so $f(A) = f(B) = 0$, the smallest possible value. Their midpoint is $(0,0)$, where $f = (0-3)^2 = 9$. Convexity would require $f(\text{midpoint})\le\tfrac12 f(A) + \tfrac12 f(B) = 0$. But $9 \gt 0$. So $f$ is not convex. (It has two separate global minima, and in fact a whole curve of them, $ab = 3$.)
The same trick for k-means. One-dimensional data $0, 1, 9, 10$ and two centres $(c_1,c_2)$. The assignment $(0.5,\ 9.5)$ and its swapped twin $(9.5,\ 0.5)$ give the same cost, $0.25$ per point. Their midpoint is $(5,5)$, with cost $\tfrac14(25+16+16+25) = 20.5$. Not convex. This symmetry (relabelling the clusters, or permuting the hidden units of a neural network) is the typical reason for non-convexity.
| Convex in the parameters | Why (rules) |
|---|---|
| Least squares / MSE for a linear model | Hessian $X^\top X/n\succeq0$ |
| Ridge, elastic net, Lasso | convex loss + convex penalty ($\|\cdot\|^2$, $\|\cdot\|_1$) |
| Logistic regression (cross-entropy) | softplus of an affine map, minus a linear term |
| Softmax (multinomial) regression | log-sum-exp of affine maps, minus a linear term |
| SVM hinge loss (+ L2) | max of affine functions, plus a squared norm |
| Huber and absolute-error regression, Poisson regression | convex functions of an affine map |
| L1 and L2 penalties and constraints | norms and norm balls are convex |
| Not convex | Why |
|---|---|
| Neural-network losses (any hidden layer with a non-linearity) | products and compositions of weights; permutation symmetry; plateaus |
| k-means objective | a minimum over assignments; label-swapping symmetry |
| Matrix factorisation / recommender models | products of unknown factors ($ab$ above) |
| Gaussian-mixture log-likelihood | log of a sum over components; label-swapping symmetry |
| Anything with a count or rank constraint ($\|\mathbf{w}\|_0\le k$, rank $\le r$) | the feasible set is not convex |
Nuances (please keep them in mind).
- Non-convex problems can still be tamed. Matrix factorisation is convex in each factor when the other is fixed (alternating least squares exploits that). The last layer of a network, given fixed features, is just logistic or softmax regression: convex. Some non-convex problems (for example PCA, and certain factorisation problems under assumptions) have been proved to have no bad local minima.
- For deep networks there is no general guarantee. That training works well in practice is an empirical fact with partial explanations (over-parameterisation, benign-looking landscapes in experiments), not a theorem for arbitrary networks. Do not read "SGD finds flat minima" or "all local minima are good" as established laws.
- Convex relaxations trade exactness for tractability: the L1 norm replaces the count of non-zeros, giving the Lasso.
Why do we need it?
It tells you what to expect from training. Convex losses: one answer, the same on every run, and theory for speed. Non-convex losses: the answer depends on initialisation and randomness, and you must run several seeds and judge by validation performance.
Where is it used?
Choosing a model family (a convex baseline such as logistic regression is a reliable reference point), debugging (two seeds that give very different losses signal a non-convex problem or a failed run), and deciding whether a convex relaxation or a convex last-layer fit can replace a hard problem.
How is it used?
For each loss in your pipeline ask "linear in the parameters, convex function of the prediction?". If yes, you can use the strong guarantees of this chapter. If no, test sensitivity: train from several seeds, compare final losses, and use techniques built for non-convexity (momentum, Adam, schedules, restarts).
Convexity is about the parameters, not the data. The same logistic loss is convex in $\mathbf{w}$ whatever data you feed it (it can still have no minimiser if the data are perfectly separable, since the loss only approaches zero). And a model is "linear" or "non-linear" in its weights, not in the inputs: polynomial regression with a convex loss is convex in the coefficients even though the curve is non-linear in $x$.
Quick check: $f(w) = \sum_i (y_i - w^2 x_i)^2$ (predicting $y$ by $w^2x$). Is it convex in $w$? What about in $u = w^2$?
In $w$ it is generally not convex: the prediction $w^2x$ is not linear in $w$, and $f(w) = f(-w)$ is symmetric, so if the best fit has $w^\star\ne0$ then $w^\star$ and $-w^\star$ are both minimisers and their midpoint $0$ is worse (because $w=0$ is not the best fit). In terms of $u=w^2$ (with $u\ge0$) the loss $\sum(y_i - ux_i)^2$ is an ordinary least-squares problem: convex. Re-parametrisation can change convexity.
Recap, cheat sheet and practice
- A convex set contains the whole segment between any two of its points (disks, boxes, half-planes, simplices, intersections). A convex function has chords above its graph: $f(\theta\mathbf{x}+(1-\theta)\mathbf{y})\le\theta f(\mathbf{x})+(1-\theta)f(\mathbf{y})$. Equivalently its epigraph is convex.
- Strictly convex: chords strictly above (at most one minimiser). $\mu$-strongly convex: at least as curved as a parabola of curvature $\mu$ (a floor); with $L$-smoothness this gives the linear rate $(1-\mu/L)^k$. Concave means $-f$ is convex: maximising a concave function is minimising a convex one.
- Jensen: $f(\mathbb{E}[X])\le\mathbb{E}[f(X)]$ for convex $f$ (reversed for concave). It gives variance $\ge0$, AM–GM, $\text{KL}\ge0$, "cross-entropy is minimised at the truth" and the ELBO bound.
- First-order test: $f(\mathbf{y})\ge f(\mathbf{x})+\nabla f(\mathbf{x})^\top(\mathbf{y}-\mathbf{x})$ (the tangent lies below the graph). So $\nabla f(\mathbf{x}^\star)=\mathbf{0}$ means a global minimum. Second-order test: $\nabla^2 f\succeq0$ everywhere (all Hessian eigenvalues $\ge0$).
- The key theorem: for a convex function over a convex set, every local minimiser is a global minimiser. Proof: if a lower point $\mathbf{y}$ existed, the points $(1-\theta)\mathbf{x}+\theta\mathbf{y}$ would be lower than $\mathbf{x}$ for every small $\theta\gt0$, contradicting local minimality.
- Convex problems are special: the start does not matter, stationary $=$ optimal, any sensible descent method reaches the optimum, rates are proven, and duality gives certificates.
- Recognise convexity by building from bricks with safe joins: non-negative sums, affine substitution, pointwise maximum, convex non-decreasing outer functions. A difference, a minimum or a negative weight breaks the rules.
- ML losses: least squares, ridge, lasso, logistic, softmax, hinge/SVM, Huber are convex in the parameters. Neural networks, k-means, matrix factorisation and mixtures are not. Say so honestly, and expect no guarantees (only intuition) for the latter.
Cheat sheet
| Test | Convex if ... | Picture |
|---|---|---|
| Set (segment test) | $\theta\mathbf{x}+(1-\theta)\mathbf{y}\in C$ | a rubber band stays inside |
| Function (chord test) | $f(\theta\mathbf{x}+(1-\theta)\mathbf{y})\le\theta f(\mathbf{x})+(1-\theta)f(\mathbf{y})$ | chord above the graph |
| First-order | $f(\mathbf{y})\ge f(\mathbf{x})+\nabla f(\mathbf{x})^\top(\mathbf{y}-\mathbf{x})$ | tangent below the graph |
| Second-order | $\nabla^2 f(\mathbf{x})\succeq0$ for all $\mathbf{x}$ | bends up in every direction |
| Strictly convex | strict chord inequality; $\nabla^2 f\succ0$ is sufficient | no straight pieces |
| $\mu$-strongly convex | $\nabla^2 f\succeq\mu I$, equivalently the floor $+\frac\mu2\|\mathbf{y}-\mathbf{x}\|^2$ | a parabola floor |
| Jensen | $f(\mathbb{E}X)\le\mathbb{E}f(X)$ | average first, then $f$: smaller |
| Optimality | $\nabla f(\mathbf{x}^\star)=\mathbf{0}\iff$ global minimum | a flat tangent below everything |
| Safe operations | $\sum a_if_i$ ($a_i\ge0$), $f(A\mathbf{x}+\mathbf{b})$, $\max_if_i$, $g(h(\mathbf{x}))$ ($g$ convex non-decreasing, $h$ convex) | Lego bricks |
| Convex problem | $\min f_0$ s.t. $f_i\le0$ (convex), $A\mathbf{x}=\mathbf{b}$ (affine) | a bowl over a convex region |
import numpy as np
rng = np.random.default_rng(0)
# ---- 1. The chord test as a numerical sanity check (it can DISPROVE convexity, never prove it)
def looks_convex(f, lo, hi, trials=20000):
a = rng.uniform(lo, hi, trials); b = rng.uniform(lo, hi, trials); t = rng.uniform(0, 1, trials)
return bool(np.all(f(t * a + (1 - t) * b) <= t * f(a) + (1 - t) * f(b) + 1e-9))
print("x^2 :", looks_convex(lambda x: x**2, -3, 3)) # True
print("|x| :", looks_convex(np.abs, -3, 3)) # True (convex, with a corner)
print("sin x :", looks_convex(np.sin, -4, 4)) # False
print("x^2-|x| :", looks_convex(lambda x: x**2 - np.abs(x), -3, 3)) # False (a difference of convex functions)
# ---- 2. Second-order test: eigenvalues of the Hessian of a loss
X = rng.normal(size=(100, 3))
H_ls = X.T @ X / len(X) # least squares: Hessian is constant
print("least squares eig :", np.round(np.linalg.eigvalsh(H_ls), 3), "-> all >= 0, convex")
# least squares eig : [0.771 0.874 1.209] -> all >= 0, convex
w = rng.normal(size=3); s = 1 / (1 + np.exp(-X @ w))
H_log = X.T @ (X * (s * (1 - s))[:, None]) / len(X) # logistic: X^T D X / n
print("logistic eig :", np.round(np.linalg.eigvalsh(H_log), 4), "-> all >= 0, convex")
# logistic eig : [0.1456 0.1638 0.2458] -> all >= 0, convex
A = np.array([[2.0, 3.0], [3.0, 2.0]]) # Hessian of x^2 + 3xy + y^2
print("x^2+3xy+y^2 eig :", np.linalg.eigvalsh(A), "-> one negative: NOT convex") # [-1. 5.]
# ---- 3. Jensen and KL
x = np.array([1.0, 5.0]); p = np.array([0.5, 0.5])
print("Jensen x^2:", (p @ x) ** 2, "<=", p @ x**2) # 9.0 <= 13.0
pt = np.array([0.5, 0.3, 0.2]); q = np.array([0.4, 0.4, 0.2])
kl = float(np.sum(pt * np.log(pt / q)))
print("KL(p||q) =", round(kl, 4), ">= 0") # 0.0253
print("cross-entropy =", round(float(-np.sum(pt * np.log(q))), 4), "= entropy",
round(float(-np.sum(pt * np.log(pt))), 4), "+ KL") # 1.0549 = 1.0297 + 0.0253
# ---- 4. Local = global for a convex loss; not for a non-convex one
def gd(grad, x0, eta, steps):
x = np.array(x0, float)
for _ in range(steps):
x = x - eta * grad(x)
return x
# convex: regularised logistic regression on 6 points
Xd = np.array([[1, .5], [2, 1], [.5, 2], [-1, -.5], [-2, -1], [-.5, -2]]); yd = np.array([1, 1, 1, -1, -1, -1]); lam = 0.1
def g_log(w):
m = yd * (Xd @ w)
return -(Xd * (yd / (1 + np.exp(m)))[:, None]).mean(axis=0) + lam * w
ends = np.array([gd(g_log, rng.uniform(-4, 4, 2), 1.0, 500) for _ in range(20)])
print("convex logistic: spread of 20 end points =", float(np.ptp(ends, axis=0).max())) # about 1e-15: the same point
# non-convex: f(a, b) = (a*b - 3)^2 has a whole curve of minimisers; where we end depends on the start
def g_mf(v):
a, b = v; r = a * b - 3
return np.array([2 * r * b, 2 * r * a])
ends = np.array([gd(g_mf, rng.uniform(-2, 2, 2), 0.02, 3000) for _ in range(20)])
print("non-convex (ab-3)^2: distinct end points =", len({tuple(np.round(e, 1)) for e in ends})) # 14 of 20
print("midpoint test: f(0,0) =", (0 * 0 - 3) ** 2, "> average of f(sqrt3,sqrt3), f(-sqrt3,-sqrt3) = 0") # 9 > 0
1. Which of these sets is convex?
2. When is a local minimum guaranteed to be a global minimum?
3. $X$ takes the values $0$ and $6$ with probability $\tfrac12$ each, and $f(x)=x^2$. Which is correct?
4. Look at $f(x,y)=x^2+4xy+y^2$. Its Hessian is $\begin{bmatrix}2&4\\4&2\end{bmatrix}$. What can you conclude?
5. Which function is convex?
6. A convex, $L$-smooth loss $\ell(\mathbf{w})$ gets the penalty $\frac\lambda2\|\mathbf{w}\|^2$ with $\lambda\gt0$. What is now guaranteed?
Practice problems
A. Which of these sets are convex? (i) $\{x^2+y^2\le4,\ x+y\ge1\}$ (ii) $\{xy\ge1,\ x\gt0,\ y\gt0\}$ (iii) $\{xy\le1,\ x\gt0,\ y\gt0\}$
(i) Convex: a disk intersected with a half-plane. (ii) Convex: it is $\{y\ge1/x,\ x\gt0\}$, the region above the graph of the convex function $1/x$ on $x\gt0$ (an epigraph). (iii) Not convex: $(0.1,\,9.9)$ and $(9.9,\,0.1)$ are in the set (products $0.99\le1$) but their midpoint $(5,5)$ has $xy=25\gt1$.
B. Decide convexity with the Hessian test: (i) $f(x,y)=x^2+4xy+5y^2$ (ii) $f(x,y)=x^2+4xy+3y^2$. For the convex one give $\mu$, $L$ and $\kappa$.
(i) $H = \begin{bmatrix}2&4\\4&10\end{bmatrix}$: trace $12$, determinant $20-16=4$. Eigenvalues $6\pm\sqrt{36-4} = 6\pm5.657$, i.e. $11.657$ and $0.343$. Both positive: strongly convex with $\mu\approx0.343$, $L\approx11.66$, $\kappa\approx34$. (ii) $H = \begin{bmatrix}2&4\\4&6\end{bmatrix}$ has determinant $12-16=-4\lt0$, so one eigenvalue is negative: not convex (a saddle).
C. $X$ takes the values $1,2,4$ with probabilities $0.5, 0.25, 0.25$. Verify Jensen for the convex function $f(x)=1/x$ on $x\gt0$.
$\mathbb{E}X = 0.5 + 0.5 + 1 = 2$, so $f(\mathbb{E}X) = 0.5$. $\mathbb{E}f(X) = 0.5\cdot1 + 0.25\cdot0.5 + 0.25\cdot0.25 = 0.5 + 0.125 + 0.0625 = 0.6875$. Indeed $0.5\le0.6875$.
D. Show that the set of minimisers of a convex function is convex, and describe it for $f(x)=\lvert x-1\rvert+\lvert x+1\rvert$.
Let $\mathbf{x},\mathbf{y}$ both attain the minimum value $m$. For $\mathbf{z}=\theta\mathbf{x}+(1-\theta)\mathbf{y}$ convexity gives $f(\mathbf{z})\le\theta m+(1-\theta)m = m$, and $f(\mathbf{z})\ge m$ because $m$ is the minimum. So $f(\mathbf{z})=m$ and $\mathbf{z}$ is also a minimiser. For $f(x)=\lvert x-1\rvert+\lvert x+1\rvert$ (a sum of convex functions): for $-1\le x\le1$ we get $(1-x)+(x+1)=2$, and outside that interval $f\gt2$. The set of minimisers is the interval $[-1,1]$: convex, but not a single point (so $f$ is not strictly convex).
E. Show that each of these is or is not convex using a violating pair or a rule: (i) $e^x+\lvert x-2\rvert$ (ii) $\min(x^2,1)$ (iii) $x^2-\lvert x\rvert$.
(i) Convex: a sum of convex functions with weights $1\ge0$. (ii) Not convex: $f(0)=0$, $f(2)=1$, and at the midpoint $x=1$, $f(1)=\min(1,1)=1$, but the chord height is $\tfrac12\cdot0+\tfrac12\cdot1 = 0.5\lt1$. (iii) Not convex: $f(-0.5)=f(0.5)=0.25-0.5=-0.25$, and at the midpoint $f(0)=0\gt-0.25$. (A negative weight on $\lvert x\rvert$ breaks the rule.)
F. (i) Is the logistic loss with penalty, $\sum_i\ln(1+e^{-y_i\mathbf{w}^\top\mathbf{x}_i})+\lambda\|\mathbf{w}\|^2$, convex? Strongly? (ii) Show that $L(a,b)=\sum_i\big(y_i - a\tanh(bx_i)\big)^2$ is not convex unless the best fit is $a=0$.
(i) Each term is softplus of an affine function of $\mathbf{w}$ (convex), the sum is convex, and $\lambda\|\mathbf{w}\|^2$ adds Hessian $2\lambda I$: so it is convex, and $2\lambda$-strongly convex when $\lambda\gt0$. (ii) The loss is unchanged by $(a,b)\to(-a,-b)$ because $(-a)\tanh(-bx) = a\tanh(bx)$. At $(0,0)$ the loss is $\sum y_i^2$. Suppose a point $(a,b)$ fits better, $L(a,b)\lt\sum y_i^2$. Then $(a,b)$ and $(-a,-b)$ both have that loss, and their midpoint $(0,0)$ has the larger loss $\sum y_i^2$. A convex function would have midpoint value at most the average. So $L$ is not convex. (Symmetry like this is the standard reason deep-network losses are non-convex.)
Constrained Optimization
So far you could choose any point. Real problems come with rules: a budget, a safety limit, probabilities that must add to 1. This chapter is the big picture of what rules do to an optimization problem, and a first look at the tools (Lagrange multipliers, KKT, duality) that the next three chapters teach in depth.
- Write any constrained problem in one standard form, with the notation we keep for the rest of the guide
- Tell equality constraints from inequality constraints, and read a feasible region: boundary, interior, active and inactive constraints
- See why a constraint can move the best answer onto the boundary, and why $\nabla f = \mathbf{0}$ stops being the test
- Meet Lagrange multipliers, the Lagrangian, the KKT conditions, the primal and dual problems, weak and strong duality, and constraint qualification, each with one worked example
- Recognise constraints in machine learning (norm balls, the simplex, boxes, SVM margins) and know which method fits which problem
The general constrained problem core
Think about shopping with a budget. You want the best basket (the objective), but you may not spend more than 100 dollars (the rule). Two things matter now: what you want and what you are allowed.
Optimization splits the same way. The objective function says what is good. The constraints say which choices are allowed at all. The answer must be the best choice among the allowed ones.
Every constrained problem, however it is worded, can be rewritten in the same tidy shape. Having one shape lets us build one set of tools.
A fence. You have exactly 20 metres of fence and you want a rectangular garden of the largest area. Let $x$ and $y$ be the two side lengths.
- Goal: make the area $xy$ as large as possible.
- Rule: the perimeter is exactly 20, so $2x + 2y = 20$.
- Our standard shape only minimises, so "maximise $xy$" becomes "minimise $-xy$" (the biggest area is the smallest negative area).
- Our standard shape writes every equation as "something $= 0$", so the rule becomes $2x + 2y - 20 = 0$.
So the problem is: minimise $-xy$ subject to $2x + 2y - 20 = 0$. (You will solve it in Chapter 3.8: the answer is the square, $x = y = 5$.)
The standard form of a constrained optimization problem, for $\mathbf{x}\in\mathbb{R}^n$:
$$\min_{\mathbf{x}}\; f(\mathbf{x}) \quad \text{subject to} \quad g_i(\mathbf{x}) \le 0 \;\; (i = 1,\dots,m), \qquad h_j(\mathbf{x}) = 0 \;\; (j = 1,\dots,p).$$- $f$ is the objective. Each $g_i(\mathbf{x}) \le 0$ is an inequality constraint ($m$ of them). Each $h_j(\mathbf{x}) = 0$ is an equality constraint ($p$ of them).
- How to convert: "maximise $f$" becomes "minimise $-f$". "$g(\mathbf{x}) \ge 0$" becomes "$-g(\mathbf{x}) \le 0$". "$h(\mathbf{x}) = c$" becomes "$h(\mathbf{x}) - c = 0$". "$a \le g(\mathbf{x}) \le b$" becomes two constraints, $a - g \le 0$ and $g - b \le 0$.
Notation we keep for the whole guide (Chapters 3.7 to 3.10 and 3.17).
- Objective $f$; inequality constraints $g_i(\mathbf{x}) \le 0$; equality constraints $h_j(\mathbf{x}) = 0$.
- One multiplier per constraint: $\lambda_j$ for each equality constraint (it can have any sign) and $\mu_i \ge 0$ for each inequality constraint (never negative).
- The Lagrangian is $L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu) = f(\mathbf{x}) + \sum_j \lambda_j h_j(\mathbf{x}) + \sum_i \mu_i g_i(\mathbf{x})$ (a plus sign in front of every constraint term).
- $p^\star$ is the optimal value of the problem, $\mathbf{x}^\star$ an optimal point, $\eta$ a learning rate (as before).
These symbols will make sense one by one in this chapter. You do not have to memorise them now.
Why do we need it?
Methods and theorems are written for one shape. If every problem is first put into the same shape, we only have to learn the tools once, and we can compare problems that look different on the surface.
Where is it used?
Textbook statements of KKT and duality, the input format of solvers such as SciPy's minimize (SLSQP), CVXPY, and quadratic-programming code inside support vector machines.
How is it used?
Read the word problem. Pick the variables. Write the objective and turn "maximise" into "minimise the negative". Write each rule as "$\le 0$" or "$= 0$". Then check each piece at a test point, as the widget below does.
Direction of the sign. The standard form says "$g_i(\mathbf{x}) \le 0$" is the allowed side. If your problem says "profit must be at least 4", the constraint is $4 - \text{profit} \le 0$, not $\text{profit} - 4 \le 0$. Writing the wrong sign flips which side is allowed.
Maximise versus minimise. Flipping the sign of $f$ does not change where the best point is, only the value: if the maximum of $xy$ is 25, the minimum of $-xy$ is $-25$.
Quick check: write "maximise $x + y$ subject to $x \ge 1$ and $x + y = 3$" in standard form.
Minimise $-x - y$ subject to $g_1(\mathbf{x}) = 1 - x \le 0$ and $h_1(\mathbf{x}) = x + y - 3 = 0$.
Equality constraints: a rail to walk along core
An equality constraint such as $x + y = 2$ is a rule with an "equals" sign. In the flat plane, the points that obey it form a line. Another, $x^2 + y^2 = 4$, forms a circle. You may stand only on that curve: it is a rail.
That shrinks the problem. With two variables and one equation, only one free choice is left: how far along the rail you stand. So an equality constraint removes one degree of freedom.
Picture the objective as a hill over the plane. The rail draws a path over that hill. You are looking for the lowest point of that path, not the lowest point of the whole hill.
Minimise $f(x,y) = x^2 + y^2$ subject to $x + y = 2$. Here we can use substitution: use the rule to remove one variable.
- From the rule, $y = 2 - x$.
- Put it into $f$: $f = x^2 + (2 - x)^2 = 2x^2 - 4x + 4$. Now it is a one-variable problem.
- Set the derivative to zero: $4x - 4 = 0$, so $x = 1$ and then $y = 2 - 1 = 1$.
- The value is $f(1,1) = 2$. Check: the second derivative $4 \gt 0$, so it is a minimum.
Notice what happened: without the rule the minimum of $x^2+y^2$ is at $(0,0)$ with value $0$. With the rail, the best we can do is $(1,1)$ with value $2$.
An equality constraint is a condition $h(\mathbf{x}) = 0$. The set of points that satisfy $p$ independent equality constraints in $\mathbb{R}^n$ is typically a surface with $n - p$ free directions (a curve for $n=2,\,p=1$; a surface for $n=3,\,p=1$; a curve for $n=3,\,p=2$).
If an equality constraint can be solved for one variable, substitution turns the problem into an unconstrained one in fewer variables. When it cannot be solved by hand, we need a more general tool: Lagrange multipliers.
Why do we need it?
Some rules are exact: probabilities add to 1, a budget is spent completely, a fence has a fixed length. We need a way to optimise while staying exactly on the set where the rule holds.
Where is it used?
Maximum-entropy and probability models (sum to 1), portfolio problems (weights sum to 1), physics and robotics (fixed lengths and joints), and fitting with exact conditions such as "the curve must pass through this point".
How is it used?
Count the free directions ($n - p$). If one variable can be solved out, substitute and minimise the shorter, unconstrained problem. If not, use a Lagrange multiplier for each equation.
Quick check: minimise $x^2 + 2y^2$ subject to $x + y = 3$ by substitution.
Put $x = 3 - y$: $f = (3-y)^2 + 2y^2 = 3y^2 - 6y + 9$. Then $6y - 6 = 0$ gives $y = 1$, $x = 2$, and $f = 4 + 2 = 6$. (Second derivative $6 \gt 0$, so it is a minimum.)
Inequality constraints: a wall you may touch but not cross core
An inequality constraint such as $x \ge 1$ is a wall. You may stand on the wall, or anywhere on the allowed side, but you may not cross it.
An equality constraint forces you onto a curve. An inequality constraint gives you room: a whole region, not just its edge. So two things can happen, and they behave very differently:
- The best place is away from the wall. The wall was not in the way, so it does nothing.
- The best place is pushed against the wall. You would love to go further, but the wall stops you.
Minimise $f(x) = x^2$ (lowest at $x = 0$), with one constraint each time:
- $x \ge 1$, written $g(x) = 1 - x \le 0$. The free minimum $x = 0$ is not allowed (it breaks the rule: $g(0) = 1 \gt 0$). On the allowed side $x \ge 1$ the function $x^2$ only gets bigger as $x$ grows, so the best allowed point is the wall itself: $x^\star = 1$, $f = 1$.
- $x \ge -1$, written $g(x) = -1 - x \le 0$. The free minimum $x = 0$ is allowed ($g(0) = -1 \le 0$). The wall at $x = -1$ is not in the way, so the answer is $x^\star = 0$, $f = 0$, as if there were no rule.
An inequality constraint is a condition $g(\mathbf{x}) \le 0$. It splits the space into the allowed side ($g \le 0$), the wall ($g = 0$) and the forbidden side ($g \gt 0$).
At a point, the constraint is inactive if $g(\mathbf{x}) \lt 0$ (strictly inside, room to move) and active if $g(\mathbf{x}) = 0$ (touching the wall). You will use these two words constantly in the next chapters.
Why do we need it?
Most real rules are limits, not exact amounts: "at most 8 hours", "weights must not be negative", "the error must stay under a tolerance". An inequality lets a rule matter only when it is actually in the way.
Where is it used?
Linear programming (resources), non-negative least squares, support vector machines (every training point must be on the right side of a margin), box bounds on parameters, and trust-region methods.
How is it used?
For each constraint ask one question at the answer: is it touching the wall (active) or not (inactive)? Inactive constraints can be ignored locally; active ones act like equality constraints.
Quick check: minimise $(x - 4)^2$ subject to $x \le 2$. Which side is the answer on, and is the constraint active?
The free minimum $x = 4$ is forbidden. On the allowed side $x \le 2$, the function decreases as $x$ rises toward 2, so the best point is the wall, $x^\star = 2$ with $f = 4$. The constraint $g = x - 2 \le 0$ is active.
The feasible region core
Put all the rules together and look at the ground that satisfies every rule at once. That ground is the feasible region: the part of the map where you are allowed to stand.
It has a shape. Its edge (the boundary) is where at least one rule is exactly tight. Its inside (the interior) is where every inequality rule has room to spare.
Each point of the region is either inside or on a wall. For each wall we can ask: "Am I touching you right now?" The walls you touch are active. The others are inactive: they exist, but they are far enough away not to matter at this spot.
Take three inequality constraints in the plane: $g_1 = x + y - 4 \le 0$, $g_2 = -x \le 0$ and $g_3 = -y \le 0$. The feasible region is the triangle with corners $(0,0)$, $(4,0)$ and $(0,4)$. Classify five points.
| Point | $g_1$ | $g_2$ | $g_3$ | Verdict |
|---|---|---|---|---|
| $(1,1)$ | $-2$ | $-1$ | $-1$ | feasible, interior: all three inactive |
| $(2,2)$ | $0$ | $-2$ | $-2$ | feasible, boundary: $g_1$ active |
| $(0,2)$ | $-2$ | $0$ | $-2$ | feasible, boundary: $g_2$ active |
| $(0,0)$ | $-4$ | $0$ | $0$ | feasible, a corner: $g_2$ and $g_3$ both active |
| $(4,1)$ | $1$ | $-4$ | $-1$ | not feasible: $g_1 = 1 \gt 0$ is violated |
Each cell is just a substitution, for example at $(4,1)$: $g_1 = 4 + 1 - 4 = 1$.
The feasible set (feasible region) is $$\mathcal{F} = \{\mathbf{x} : g_i(\mathbf{x}) \le 0 \text{ for all } i, \;\; h_j(\mathbf{x}) = 0 \text{ for all } j\}.$$ A feasible point is any $\mathbf{x}\in\mathcal{F}$. An optimal point $\mathbf{x}^\star$ is a feasible point with the smallest $f$ among all feasible points, and $p^\star = f(\mathbf{x}^\star)$.
- At a feasible point, the constraint $g_i$ is active if $g_i(\mathbf{x}) = 0$ and inactive if $g_i(\mathbf{x}) \lt 0$. Equality constraints are always active.
- A point where every inequality is inactive is an interior point (ignoring equalities). A feasible point with at least one active inequality is on the boundary.
- If the $g_i$ are convex functions and the $h_j$ are linear ("affine"), then $\mathcal{F}$ is a convex set (Chapter 3.6). Convex feasible regions make life much easier.
- Two warnings: $\mathcal{F}$ can be empty ($x \ge 2$ and $x \le 1$: the problem is infeasible), and a nonempty $\mathcal{F}$ can still have no best point (for example "minimise $x$ subject to $x \gt 0$" has no smallest value; the objective can get arbitrarily close to $0$ but never reach it).
Why do we need it?
The answer must live inside the feasible region, so you cannot judge an answer without knowing the region. Knowing which constraints are active at the answer tells you which rules actually shape it.
Where is it used?
Linear programming (the region is a polygon and the optimum sits at a corner), SVM training (only the points whose margin constraint is active, the support vectors, matter), and solver output that reports "active set" and "slack".
How is it used?
Evaluate each $g_i$ and $h_j$ at the point. All satisfied: feasible. Some $g_i = 0$: boundary. Count the active ones: more than $n$ active constraints, or dependent ones, means a degenerate, hard corner.
Active is about a point, not a constraint. The same constraint can be active at one point and inactive at another. Also "active" means $g_i(\mathbf{x}) = 0$ exactly, not "close to $0$". And an inactive constraint is not useless: it would matter if the answer moved.
Quick check: is $(1,1)$ feasible for $x^2 + y^2 \le 2$ and $x + y \ge 2$? Which constraints are active?
$g_1 = 1 + 1 - 2 = 0$ (active) and $g_2 = 2 - x - y = 0$ (active). Both hold, so it is feasible, and both are active. In fact $(1,1)$ is the only feasible point (the disc touches the line there).
Why constraints change the answer core
Imagine a valley with a lake at the bottom. The lowest ground is in the middle of the lake. A fence runs across the valley and you must stay on the dry side. The best you can do is to walk down until the fence stops you. You end at the fence, not at the bottom.
This is the single most important picture of the chapter:
- If the free minimum is allowed, constraints change nothing.
- If the free minimum is forbidden, the answer lands on the boundary, at the point where the level curves of $f$ just touch the allowed region.
A surprise follows: at that boundary point the ground is still sloping downhill. The usual test "the gradient is zero" no longer holds there. We will need a new test (that is what Lagrange and KKT are).
Minimise $f(x,y) = (x-3)^2 + (y-2)^2$ (a bowl with its bottom at $(3,2)$, value $0$) subject to $x^2 + y^2 \le 4$ (a disc of radius 2).
- Is the free minimum allowed? $3^2 + 2^2 = 13 \gt 4$. No.
- So the answer is on the circle, in the direction of the bowl's centre. The point of the disc closest to $(3,2)$ is the point on the circle along the ray to $(3,2)$: $\mathbf{x}^\star = 2\cdot\dfrac{(3,2)}{\sqrt{13}} \approx (1.664,\,1.109)$.
- Value: the distance from $(3,2)$ to the circle is $\sqrt{13} - 2 \approx 1.606$, so $p^\star = (\sqrt{13}-2)^2 \approx 2.578$. That is worse than the free value $0$, as it must be.
- The gradient there: $\nabla f = 2(\mathbf{x}^\star - (3,2)) \approx (-2.67, -1.78) \ne \mathbf{0}$. The ground still slopes down through the fence.
Now make the disc bigger, radius 4: $13 \le 16$, so $(3,2)$ is inside. Then $\mathbf{x}^\star = (3,2)$, $p^\star = 0$, $\nabla f = \mathbf{0}$, and the constraint is inactive.
Facts. Let $\mathcal{F}$ be the feasible set and $f_{\text{free}} = \min_{\mathbf{x}} f$ the unconstrained minimum value.
- $p^\star \ge f_{\text{free}}$: constraints can only make the best value worse or equal. Equality holds exactly when a free minimiser lies in $\mathcal{F}$.
- If $f$ is convex and no free minimiser is feasible, a constrained minimiser lies on the boundary of $\mathcal{F}$. (If it were strictly inside, nothing nearby would block it, so it would also be a free local minimum with $\nabla f = \mathbf{0}$, and for a convex $f$ that makes it a free global minimum: forbidden.)
- At a boundary optimum, $\nabla f(\mathbf{x}^\star) \ne \mathbf{0}$ in general. Instead, $-\nabla f$ ("downhill") points out of the region, along the wall's outward normal. The wall holds you back, and the strength of that hold is the multiplier you meet next.
Why do we need it?
If you keep using the old test "gradient equals zero" on a constrained problem, you will either miss the answer (it has a non-zero gradient) or return a forbidden point. The picture tells you what to look for instead.
Where is it used?
Reading solver output ("the bound is active"), SVM geometry (the optimal line touches the margin walls), regularised models (the best weights sit on the edge of the allowed norm ball when the data wants more), and sanity checks of any constrained result.
How is it used?
Step 1: solve ignoring the constraints. Step 2: is that point feasible? If yes, done. If no, expect an answer on the boundary and use Lagrange or KKT to find which walls are active.
Quick check: minimise $(x - 5)^2$ subject to $x \le 3$. What are $x^\star$ and $p^\star$, and is $f'(x^\star) = 0$?
The free minimum $x = 5$ is forbidden, so the answer is on the wall: $x^\star = 3$, $p^\star = (3-5)^2 = 4$. And $f'(3) = 2(3-5) = -4 \ne 0$: the function is still decreasing as you push right, but the wall stops you.
Lagrange multipliers: the wall pushes back core
Put a ball in a bowl and thread it on a straight rail that cuts through the bowl. Gravity pulls the ball down the bowl (that is $-\nabla f$). The rail only lets the ball slide along it. The ball rolls along the rail until it stops at the lowest point of the rail.
When it stops, the part of gravity along the rail is zero. What is left is gravity pushing sideways into the rail, and the rail pushes back with exactly the same strength. The rail pushes in a direction perpendicular to itself: along the constraint's gradient $\nabla h$.
So at the answer: $\nabla f$ is a multiple of $\nabla h$. That multiple (with a minus sign) is the Lagrange multiplier $\lambda$: how hard the rail has to push. In short: "$\nabla f + \lambda\nabla h = \mathbf{0}$".
Minimise $f = x^2 + y^2$ subject to $h = x + y - 2 = 0$ (the earlier example).
- Gradients: $\nabla f = (2x,\,2y)$ and $\nabla h = (1,\,1)$.
- Balance: $\nabla f + \lambda\nabla h = \mathbf{0}$ gives $2x + \lambda = 0$ and $2y + \lambda = 0$. So $x = y = -\lambda/2$.
- The rule $x + y = 2$ gives $-\lambda = 2$, so $\lambda = -2$ and $x = y = 1$.
- Check: $\nabla f(1,1) = (2,2)$ and $\lambda\nabla h = -2\,(1,1) = (-2,-2)$. They cancel. ✓
It matches the substitution answer $(1,1)$, found now without eliminating any variable. We used three equations for three unknowns $(x, y, \lambda)$.
A Lagrange multiplier $\lambda_j$ is a number attached to the equality constraint $h_j(\mathbf{x}) = 0$. At a (regular) constrained optimum $\mathbf{x}^\star$ there are numbers $\lambda_j$ with
$$\nabla f(\mathbf{x}^\star) + \sum_{j=1}^{p}\lambda_j\,\nabla h_j(\mathbf{x}^\star) = \mathbf{0}, \qquad h_j(\mathbf{x}^\star) = 0.$$That is $n + p$ equations for the $n + p$ unknowns $(\mathbf{x},\boldsymbol\lambda)$. A multiplier can be positive, negative or zero. It also measures sensitivity: it tells how fast the optimal value changes if the constraint is loosened (Chapters 3.8 and 3.10 develop this). The equality case is derived properly in Chapter 3.8; inequalities get multipliers $\mu_i \ge 0$ in Chapter 3.9.
Why do we need it?
Substitution only works when you can solve the rule for a variable. Multipliers handle any smooth constraint, treat all variables alike, and give you a free extra: a number that says how much the rule is costing you.
Where is it used?
Maximum-entropy distributions (which give softmax), principal component analysis as "best direction of unit length", physics (forces of constraint), economics (shadow prices), and the derivation of the SVM and of the KKT conditions.
How is it used?
Write $\nabla f + \lambda\nabla h = \mathbf{0}$ together with $h = 0$. Solve this system for $\mathbf{x}$ and $\lambda$. The solutions are the candidates for the optimum; compare their values of $f$.
The multiplier is not a point. $\lambda$ is an extra number, one per equality constraint. The unknowns of the Lagrange system are $\mathbf{x}$ and $\boldsymbol\lambda$ together.
The condition finds candidates, not answers. "$\nabla f + \lambda\nabla h = \mathbf{0}$" also holds at the highest point of a rail on a circle. You must compare the values of $f$ at all candidates (see Chapter 3.8).
Quick check: for the ball-on-the-rail example, verify at $(2.5, 1.5)$ that $\lambda = 1$.
$\nabla f = (2(2.5-3),\,2(1.5-2)) = (-1,-1)$ and $\nabla h = (1,1)$. We need $\nabla f + \lambda\nabla h = (-1+\lambda,\,-1+\lambda) = \mathbf{0}$, so $\lambda = 1$. ✓
The Lagrangian: one function for the whole problem core
Here is a neat trick. Instead of forbidding broken rules, charge a price for breaking them. Add to $f$ a term for each rule: "price $\times$ how much the rule is violated". The prices are the multipliers.
With the right prices, the cheapest place in the new, unconstrained landscape is the constrained answer: breaking the rule is no longer worth it. With prices too low, it still pays to cheat. The multipliers are knobs, and the correct setting makes the free minimiser obey the rules. (This works nicely for well-behaved, convex problems like the ones below; the section on strong duality shows a case where no price works.)
For $f = x^2 + y^2$ and $h = x + y - 2$, the Lagrangian is $L(x,y,\lambda) = x^2 + y^2 + \lambda\,(x + y - 2)$. For a fixed $\lambda$, minimise $L$ over $(x,y)$: setting the $x$ and $y$ derivatives to zero gives $2x + \lambda = 0$ and $2y + \lambda = 0$, so $x = y = -\lambda/2$.
| price $\lambda$ | minimiser of $L$ | $h = x + y - 2$ | obeys the rule? |
|---|---|---|---|
| $0$ | $(0,0)$ | $-2$ | no (cheating: sum is too small) |
| $-1$ | $(0.5,0.5)$ | $-1$ | no |
| $-2$ | $(1,1)$ | $0$ | yes |
| $-3$ | $(1.5,1.5)$ | $+1$ | no (over-corrected) |
Only $\lambda = -2$ makes the minimiser feasible, and it is the Lagrange multiplier we found before.
The Lagrangian of $\min f$ s.t. $g_i \le 0,\; h_j = 0$ is
$$L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu) = f(\mathbf{x}) + \sum_{j=1}^{p}\lambda_j\,h_j(\mathbf{x}) + \sum_{i=1}^{m}\mu_i\,g_i(\mathbf{x}), \qquad \mu_i \ge 0.$$- For a feasible $\mathbf{x}$ and any $\mu_i \ge 0$: $L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu) \le f(\mathbf{x})$, because $h_j = 0$ and $\mu_i g_i \le 0$.
- The key identity: $\displaystyle\max_{\boldsymbol\lambda,\;\boldsymbol\mu\ge 0} L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu) = \begin{cases} f(\mathbf{x}) & \mathbf{x}\text{ feasible},\\ +\infty & \text{otherwise.}\end{cases}$ (If a rule is broken, push its multiplier to infinity and $L$ blows up.) So the original problem equals $\displaystyle\min_{\mathbf{x}}\max_{\boldsymbol\lambda,\,\boldsymbol\mu\ge 0} L$.
- The derivative of $L$ with respect to a multiplier recovers the constraint: $\partial L/\partial\lambda_j = h_j(\mathbf{x})$.
Why do we need it?
It turns a constrained problem into a search for stationary points of one function, and it is the source of the KKT conditions and of the dual problem. All the multiplier theory is a property of $L$.
Where is it used?
Deriving the SVM dual, max-entropy models, constrained least squares, augmented-Lagrangian and ADMM solvers, and the min-max (saddle-point) training of GANs, which has the same "min over one, max over the other" shape.
How is it used?
Write down $L$ by adding $\lambda_j h_j$ and $\mu_i g_i$ to $f$. Take derivatives with respect to $\mathbf{x}$ and the multipliers, set them to zero, and solve. Or minimise $L$ over $\mathbf{x}$ to get the dual function (Chapter 3.10).
$L$ is not "minimised" in every variable. The constrained answer is a minimum in $\mathbf{x}$ but, in the multipliers, $L$ is a straight line (a maximum is being taken, or it is flat). So $(\mathbf{x}^\star,\boldsymbol\lambda^\star)$ is typically a saddle point of $L$, not a minimum. Do not run plain gradient descent on all of $L$. (Chapter 3.8 shows a 3D picture.)
Sign convention. We always write $+\lambda_j h_j$ and $+\mu_i g_i$ with $g_i \le 0$ and $\mu_i \ge 0$. Other books write $-\lambda h$ or flip signs; the answers are the same but the signs of $\lambda$ differ. Stay with our convention.
Quick check: write the Lagrangian of "minimise $x^2$ subject to $x \ge 1$".
Standard form: $g(x) = 1 - x \le 0$. So $L(x,\mu) = x^2 + \mu\,(1 - x)$ with $\mu \ge 0$.
The KKT conditions: a checklist for an optimum core
The KKT conditions (named after Karush, Kuhn and Tucker) are Lagrange multipliers extended to inequality constraints. They are a checklist of four common-sense statements about a candidate answer. If all four hold, the point is a serious candidate for the optimum.
- The forces balance. The pull of the objective is cancelled by the pushes of the walls (this is the Lagrange condition).
- You are allowed to stand here. Every rule is obeyed.
- Walls only push inward. A wall can stop you going through it, but it can never pull you toward it. So inequality multipliers are $\ge 0$.
- A wall you are not touching does not push. Its multiplier must be $0$. Only walls that are active (touching) can have a non-zero push.
One variable. Minimise $x^2$ subject to $x \ge 1$, i.e. $g(x) = 1 - x \le 0$. The Lagrangian is $L = x^2 + \mu(1 - x)$. Test the candidate $x = 1$ with $\mu = 2$:
- Stationarity: $\dfrac{\partial L}{\partial x} = 2x - \mu = 2 - 2 = 0$ ✓.
- Primal feasibility: $g(1) = 0 \le 0$ ✓.
- Dual feasibility: $\mu = 2 \ge 0$ ✓.
- Complementary slackness: $\mu\,g(x) = 2\cdot 0 = 0$ ✓.
All four hold. For the version $x \ge -1$ ($g = -1 - x$), the candidate $x = 0$ with $\mu = 0$ passes: $2\cdot0 - 0 = 0$, $g(0) = -1 \le 0$, $\mu = 0 \ge 0$, and $\mu g = 0$. Here the wall is inactive, so its multiplier is $0$.
Two variables. For the disc problem ($f = (x-3)^2 + (y-2)^2$, $g = x^2 + y^2 - 4 \le 0$) the answer is $\mathbf{x}^\star = 2(3,2)/\sqrt{13} \approx (1.664, 1.109)$ and the multiplier is $\mu = (\sqrt{13}-2)/2 \approx 0.803$. Check stationarity: $\nabla f = (-2.672, -1.781)$ and $\mu\nabla g = 0.803\cdot(3.328, 2.219) = (2.672, 1.781)$, so the sum is $\mathbf{0}$ ✓; $g = 0$ (active) so $\mu g = 0$ ✓; $\mu \ge 0$ ✓.
For $\min f(\mathbf{x})$ s.t. $g_i(\mathbf{x}) \le 0$, $h_j(\mathbf{x}) = 0$, a point $\mathbf{x}^\star$ with multipliers $(\boldsymbol\lambda,\boldsymbol\mu)$ satisfies the KKT conditions if:
| Condition | Formula | One-line meaning |
|---|---|---|
| 1. Stationarity | $\nabla f + \sum_j\lambda_j\nabla h_j + \sum_i\mu_i\nabla g_i = \mathbf{0}$ | the objective's pull is balanced by the walls' pushes ($\nabla_{\mathbf{x}} L = \mathbf{0}$) |
| 2. Primal feasibility | $g_i(\mathbf{x}^\star) \le 0,\; h_j(\mathbf{x}^\star) = 0$ | the point obeys every rule |
| 3. Dual feasibility | $\mu_i \ge 0$ | inequality walls only push inward |
| 4. Complementary slackness | $\mu_i\,g_i(\mathbf{x}^\star) = 0$ | for each wall: either it is touched ($g_i = 0$) or it pushes with zero force ($\mu_i = 0$) |
What they promise. (a) If $\mathbf{x}^\star$ is a local minimum and a constraint qualification holds (below), then multipliers exist and KKT holds. (b) If the problem is convex ($f$ and $g_i$ convex, $h_j$ affine), then any point satisfying KKT is a global minimum. So for convex problems KKT is a complete test. Chapter 3.9 studies each condition in depth.
Why do we need it?
"Gradient equals zero" is no longer the test once there are walls. KKT is the replacement: a short list of equations and sign conditions that a constrained optimum must obey, and that a solver can check.
Where is it used?
The convergence test inside constrained solvers (SLSQP, interior-point methods), the derivation of the SVM and of non-negative least squares, and proofs about regularised problems like the Lasso (the Lasso's optimality condition is a KKT condition).
How is it used?
Guess which constraints are active, set the others' multipliers to $0$, solve the equations for the rest, then check the signs: $\mu_i \ge 0$ and the inactive constraints really are satisfied. If a check fails, change your guess.
Complementary slackness is an "or", not an "and". Either the constraint is active ($g_i = 0$) or its multiplier is zero. Both can be zero at once (a wall that just touches the free minimum). The multiplier of an inactive constraint is always $0$.
KKT points are candidates. Without convexity, a KKT point can be a local minimum, a maximum or a saddle. And at an optimum where a constraint qualification fails (see below) the conditions may fail even though the point is optimal.
Quick check: which KKT condition would fail for the candidate $x = 2$, $\mu = 0$ in "minimise $x^2$ s.t. $x \ge 1$"?
Stationarity: $2x - \mu = 4 - 0 = 4 \ne 0$. (Feasibility holds, $\mu \ge 0$ holds, and $\mu g = 0\cdot(-1) = 0$ holds. But the objective still pulls the point toward $x = 1$, and the wall is not touching, so nothing balances that pull.)
The primal and the dual problem core
Every constrained problem has a twin. The original problem is the primal: you choose the variables $\mathbf{x}$ and obey the rules. The twin is the dual: instead of choosing $\mathbf{x}$, you choose the prices of the rules (the multipliers).
Here is the story. A landlord charges a price $\mu$ per unit of rule violation. For each price, you (the tenant) then pick the cheapest $\mathbf{x}$ in the priced landscape $L$. That cheapest value is the dual function $d(\mu)$. The landlord, who wants you to pay as much as possible, chooses the price that makes $d(\mu)$ largest. That is the dual problem.
The surprising fact (next two sections) is that this price game gives a lower bound on the original answer, and often the exact answer.
Primal: minimise $x^2$ subject to $x \ge 1$, i.e. $g(x) = 1 - x \le 0$. The best value is $p^\star = 1$ at $x = 1$.
Dual: first the dual function. For a fixed $\mu \ge 0$, minimise $L(x,\mu) = x^2 + \mu(1 - x)$ over $x$:
- Derivative: $2x - \mu = 0$, so the cheapest $x$ is $x = \mu/2$.
- Plug back in: $d(\mu) = \dfrac{\mu^2}{4} + \mu\Bigl(1 - \dfrac{\mu}{2}\Bigr) = \mu - \dfrac{\mu^2}{4}$.
- The dual problem: maximise $d(\mu) = \mu - \mu^2/4$ over $\mu \ge 0$. The derivative $1 - \mu/2 = 0$ gives $\mu = 2$, and $d(2) = 2 - 1 = 1$.
So $d^\star = 1 = p^\star$. The best price, $\mu = 2$, is the KKT multiplier found earlier.
Primal problem: $p^\star = \min_{\mathbf{x}} f(\mathbf{x})$ subject to $g_i(\mathbf{x}) \le 0$, $h_j(\mathbf{x}) = 0$.
Dual function (for $\boldsymbol\mu \ge 0$ and any $\boldsymbol\lambda$): $\;d(\boldsymbol\lambda,\boldsymbol\mu) = \displaystyle\min_{\mathbf{x}}\; L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu)$ (more precisely the infimum, which may be $-\infty$).
Dual problem: $d^\star = \displaystyle\max_{\boldsymbol\lambda,\;\boldsymbol\mu \ge 0}\; d(\boldsymbol\lambda,\boldsymbol\mu)$.
Fact: $d$ is always concave (an upside-down convex function, Chapter 3.6; it is a minimum of functions that are linear in $(\boldsymbol\lambda,\boldsymbol\mu)$), so the dual problem is always a convex problem, even when the primal is not. The primal is "min over $\mathbf{x}$ of max over prices"; the dual swaps the order: "max over prices of min over $\mathbf{x}$". Details in Chapter 3.10.
Why do we need it?
The dual gives a bound on the best value that holds even if you cannot solve the primal. It is sometimes much easier to solve, it has one variable per constraint (which can be far fewer than $\mathbf{x}$), and its solution is the set of multipliers.
Where is it used?
The SVM is usually trained through its dual (which brings in kernels), linear programming duality, certificates of optimality in solvers, and distributed methods such as dual ascent and ADMM.
How is it used?
Form $L$. Minimise it over $\mathbf{x}$ for fixed prices (a calculus step, often closed form). That gives $d$, a function of the prices only. Then maximise $d$ over the prices.
Quick check: for "minimise $x^2$ subject to $x \ge 2$", find the dual function and its maximum.
$L = x^2 + \mu(2 - x)$. Minimising over $x$: $x = \mu/2$, so $d(\mu) = \mu^2/4 + 2\mu - \mu^2/2 = 2\mu - \mu^2/4$. Maximise: $2 - \mu/2 = 0$ gives $\mu = 4$ and $d(4) = 8 - 4 = 4$. The primal answer is $x = 2$, $p^\star = 4$: they agree.
Weak duality: every price gives a lower bound core
Whatever price the landlord picks, the tenant's cheapest priced cost can never be above the true constrained answer. Why? Because the true answer is a point that obeys the rules, and at such a point the extra "price" terms can only lower or leave unchanged the total cost (a rule you obey costs you nothing or gives you a credit). The tenant, who can pick any $x$ (even rule-breakers), can do at least that well.
So each price gives a certified lower bound: "the optimum is at least this much". The dual problem looks for the best (highest) such bound.
In "minimise $x^2$ subject to $x \ge 1$" we have $p^\star = 1$ and $d(\mu) = \mu - \mu^2/4$. Every legal price gives a bound at most $1$:
| price $\mu$ | $d(\mu)$ | $d(\mu) \le p^\star = 1$ ? |
|---|---|---|
| $0$ | $0$ | yes |
| $1$ | $1 - 0.25 = 0.75$ | yes |
| $2$ | $2 - 1 = 1$ | yes (equal: the best price) |
| $3$ | $3 - 2.25 = 0.75$ | yes |
| $4$ | $4 - 4 = 0$ | yes |
Weak duality. For every $\boldsymbol\lambda$ and every $\boldsymbol\mu \ge 0$: $\;d(\boldsymbol\lambda,\boldsymbol\mu) \le p^\star$. Hence $d^\star \le p^\star$. The difference $p^\star - d^\star \ge 0$ is the duality gap. Weak duality holds for every problem, convex or not.
Proof. Let $\tilde{\mathbf{x}}$ be any feasible point. Then $h_j(\tilde{\mathbf{x}}) = 0$ and $\mu_i g_i(\tilde{\mathbf{x}}) \le 0$ (since $\mu_i \ge 0$, $g_i \le 0$), so $$d(\boldsymbol\lambda,\boldsymbol\mu) = \min_{\mathbf{x}} L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu) \;\le\; L(\tilde{\mathbf{x}},\boldsymbol\lambda,\boldsymbol\mu) = f(\tilde{\mathbf{x}}) + \underbrace{\textstyle\sum_j\lambda_j h_j(\tilde{\mathbf{x}})}_{=\,0} + \underbrace{\textstyle\sum_i\mu_i g_i(\tilde{\mathbf{x}})}_{\le\,0} \;\le\; f(\tilde{\mathbf{x}}).$$ This holds for every feasible $\tilde{\mathbf{x}}$, so it holds for the best one: $d \le p^\star$. $\blacksquare$
Why do we need it?
A lower bound tells you how good a candidate is without knowing the answer. If you have a feasible point with $f = 1.0004$ and a dual point with $d = 1.0000$, the true optimum is trapped between them.
Where is it used?
The "duality gap" is the stopping rule inside interior-point and SVM solvers, branch-and-bound for hard integer problems (bounds from relaxations), and certificates that a solution is optimal.
How is it used?
Run the solver on both sides. Keep the best feasible primal value (upper bound) and the best dual value (lower bound). Stop when the gap is smaller than your tolerance.
Quick check: could a dual value ever be larger than $p^\star$?
No. Weak duality says $d(\boldsymbol\lambda,\boldsymbol\mu) \le p^\star$ for every legal choice ($\boldsymbol\mu \ge 0$), for every problem. If you ever compute a dual value above a primal value, one of them is wrong (or $\mu$ was negative).
Strong duality: the bound is exact core
Weak duality says the best price gives a bound below or equal to the true answer. Strong duality says the two are equal: with the right prices, the cheapest priced cost is exactly the constrained optimum. The landlord can charge exactly enough that the rules are never worth breaking.
This is wonderful when it holds, because then solving the dual solves the primal. It does not always hold. It is guaranteed for well-behaved (convex) problems. For non-convex problems a gap can remain: the price game cannot "see" the true shape of the problem, only its convex outline.
Convex case. "Minimise $x^2$ subject to $x \ge 1$": $p^\star = 1$ and $d^\star = 1$ (shown in the last two sections). The gap is $0$.
Non-convex case. Take $f(x) = (x^2 - 1)^2$ (a double well with two minima at $x = \pm 1$, and a bump between them), with the rule $x = 0$. The only feasible point is $x = 0$, so $p^\star = f(0) = 1$. The dual function is $d(\lambda) = \min_x\,[(x^2-1)^2 + \lambda x]$. At $\lambda = 0$ this is $\min_x (x^2-1)^2 = 0$ (at $x = \pm 1$, which break the rule but are cheap), and for $\lambda \ne 0$ it is negative (the best $x$ gets a bonus of about $-|\lambda|$ from the linear term). So $d^\star = 0$ and the gap is $1 - 0 = 1 \gt 0$.
Strong duality holds when $d^\star = p^\star$ (zero gap). Sufficient condition (Slater's condition): the problem is convex ($f$ and the $g_i$ convex, the $h_j$ affine) and there is a strictly feasible point, i.e. some $\mathbf{x}$ with $g_i(\mathbf{x}) \lt 0$ for every $i$ and $h_j(\mathbf{x}) = 0$. (If every $g_i$ is affine, it is enough that the problem has at least one feasible point.)
When strong duality holds and the optima $\mathbf{x}^\star$ and $(\boldsymbol\lambda^\star,\boldsymbol\mu^\star)$ exist, they satisfy the KKT conditions together. Conversely, for a convex problem, KKT points give zero duality gap. Chapter 3.10 proves and uses this.
Why do we need it?
If the gap is zero you can swap a hard problem for its twin and recover the exact answer, and the multipliers come out as a by-product. If the gap may be positive, the dual only gives a bound.
Where is it used?
Convex problems in ML: SVMs, ridge and Lasso, logistic regression with constraints, linear and quadratic programs. Non-convex cases (neural nets, integer programs) have gaps, which is why bounds there are only bounds.
How is it used?
For a convex problem with a strictly feasible point, solve whichever of the primal or the dual is easier, and trust that the optimal values are equal. For non-convex problems treat the dual value as a lower bound only.
Zero gap needs convexity (or luck). A convex problem can still have a gap in rare pathological cases when Slater's condition fails. A non-convex problem sometimes has zero gap anyway (a few special ones do), but you cannot count on it.
Quick check: for "minimise $x^2$ subject to $x \ge 1$", is Slater's condition true?
Yes. The problem is convex ($x^2$ is convex and $1 - x$ is affine) and $x = 2$ is strictly feasible ($g(2) = -1 \lt 0$). So strong duality holds, as we saw: $d^\star = p^\star = 1$.
Constraint qualification: when the theory needs a safety check awareness
Lagrange and KKT work by replacing the constraint set by its straight-line (linear) picture: near the answer, a smooth wall looks like a flat wall with a normal direction $\nabla g$. That is a fine approximation as long as the wall is well behaved.
At a sharp spike, a cusp, or a place where the wall's gradient vanishes, the straight-line picture is wrong or empty. Then the true optimum can fail to satisfy KKT: no multiplier can balance the forces. A constraint qualification (CQ) is a simple condition that rules these bad places out, so the theory can be trusted.
The standard pathological example. Minimise $f(x) = x$ subject to $g(x) = x^2 \le 0$.
- The only feasible point is $x = 0$ (since $x^2 \le 0$ forces $x = 0$). So $x^\star = 0$ and $p^\star = 0$.
- The gradients at $x^\star$: $f'(0) = 1$ and $g'(0) = 2\cdot 0 = 0$. The wall has no direction at all.
- KKT stationarity needs $f'(0) + \mu g'(0) = 1 + \mu\cdot 0 = 0$. But the left side is always $1$, whatever $\mu$ is. No multiplier exists, although $x = 0$ really is optimal.
A second example with an equality constraint: minimise $x$ subject to $y^2 = x^3$ (a curve with a sharp point at the origin). The optimum is the cusp $(0,0)$, but $\nabla h = (-3x^2, 2y) = (0,0)$ there, so $\nabla f = (1,0)$ cannot equal $-\lambda\nabla h$.
A constraint qualification is a regularity condition on the constraints at $\mathbf{x}^\star$ that guarantees KKT multipliers exist at a local minimum. Common ones:
- LICQ (linear independence): the gradients of all active inequality constraints and of all equality constraints at $\mathbf{x}^\star$ are linearly independent. (In the example above, the single gradient $g'(0) = 0$ is not independent.) LICQ also makes the multipliers unique.
- Slater's condition (for convex problems): some point satisfies every inequality strictly ($g_i \lt 0$) and the equalities. (In the example, $x^2 \lt 0$ is impossible, so Slater fails.)
- Linear constraints: if every constraint is affine (a straight line, plane or half-space), no extra condition is needed.
Theorem (awareness level): if $\mathbf{x}^\star$ is a local minimum and a CQ holds there, then KKT multipliers exist. Without a CQ, they may not.
Why do we need it?
It explains why a correct optimum sometimes fails the KKT test, and stops you from concluding "no answer" just because you could not find a multiplier. It is the fine print of every KKT theorem.
Where is it used?
Proofs of the KKT conditions, convergence theory of constrained solvers (sequential quadratic programming assumes LICQ), and strong duality (Slater). In practice, most ML problems have linear or simple convex constraints and satisfy a CQ automatically.
How is it used?
Check it at the candidate: are the active constraint gradients independent? Is there a strictly feasible point? For linear constraints, just note that it holds. If it fails, expect multipliers to be missing, huge or non-unique, and numerical solvers to struggle.
Do not panic about CQs. If your constraints are linear, or a nice convex set with room inside, a CQ holds and you can ignore this section. It matters when you are proving something, or when a solver reports huge or missing multipliers.
Quick check: does LICQ hold at $(0,0)$ for the single constraint $h(x,y) = x^2 + y^2 = 0$ (whose only solution is $(0,0)$)?
No. $\nabla h = (2x,\,2y) = (0,0)$ at the origin, and the zero vector is not linearly independent (a single zero vector is dependent). The constraint set is one point, and KKT stationarity cannot be relied on there.
Where constraints appear in machine learning core
Most ML constraints are not exotic. They are a handful of simple shapes that keep coming back:
- A ball ("keep the weights small"): $\|\mathbf{w}\|_2 \le r$ (round) or $\|\mathbf{w}\|_1 \le r$ (a diamond).
- The simplex ("be a probability"): non-negative numbers that add up to 1.
- A box ("stay within bounds"): $l \le w \le u$, such as pixel values in $[0,1]$.
- Margin walls ("classify with room to spare"): every training point must be on the correct side of a street of width $2/\|\mathbf{w}\|$.
For the first three, a handy question is: "if I have a point that breaks the rule, what is the closest allowed point?" That is the projection onto the set, and it is cheap for these shapes.
Probabilities. Suppose a model produced the scores $\mathbf{v} = [0.8,\,0.6,\,-0.2]$ and we need a probability vector (non-negative, sum 1). The closest one (in ordinary distance) is found by shifting every entry down by the same amount $\theta$ and clipping at $0$:
- Try dropping the third entry (it is negative anyway): the remaining two sum to $1.4$, which is $0.4$ too much.
- Shift both by $\theta = 0.4/2 = 0.2$: $[0.6,\,0.4]$, and these are positive. The third stays at $0$.
- Result: $[0.6,\,0.4,\,0]$. It sums to 1 and has no negative entries ✓.
Weight clipping. Projecting $w = 2.7$ onto the box $[-1, 2]$ gives $2$ (clip to the nearest bound). Norm ball. Projecting $(2.5, 1.5)$ onto the L2 ball of radius $1.5$ scales it to length $1.5$: $(1.286, 0.772)$.
| Constraint | Set | Standard form | Where in ML |
|---|---|---|---|
| L2 ball | $\|\mathbf{w}\|_2 \le r$ | $g = \|\mathbf{w}\|_2^2 - r^2 \le 0$ | ridge regression in constraint form; weight-norm bounds; trust regions |
| L1 ball | $\|\mathbf{w}\|_1 \le r$ | $g = \sum|w_k| - r \le 0$ | Lasso in constraint form; the corners give exact zeros (sparsity) |
| Simplex | $p_k \ge 0,\;\sum p_k = 1$ | $g_k = -p_k \le 0$, $h = \sum p_k - 1 = 0$ | class probabilities, mixture weights, attention weights, portfolio weights |
| Box | $l \le w_k \le u$ | $g = l - w_k \le 0,\; w_k - u \le 0$ | bounds on parameters, pixel range $[0,1]$, clipping, non-negativity ($l = 0$) |
| SVM margin | $y_i(\mathbf{w}^\top\mathbf{x}_i + b) \ge 1$ | $g_i = 1 - y_i(\mathbf{w}^\top\mathbf{x}_i + b) \le 0$ | hard-margin support vector machine: minimise $\tfrac12\|\mathbf{w}\|^2$ under these (Chapter 3.10) |
Penalty form and constraint form are two views of one idea. "Minimise loss subject to $\|\mathbf{w}\|^2 \le r^2$" and "minimise loss $+\ \lambda\|\mathbf{w}\|^2$" give the same solutions (for convex losses) for matching $r$ and $\lambda$: the multiplier of the constraint is $\lambda$ (Chapter 3.11). A softmax output satisfies the simplex rule automatically (it is built to), so no constraint solver is needed there.
Why do we need it?
Without constraints a model can use absurdly large weights, output "probabilities" that are negative, or separate the data with no safety margin. Constraints build good behaviour into the problem itself.
Where is it used?
Ridge and Lasso (norm balls), softmax and mixture models (simplex), clipped parameters and adversarial-example bounds (boxes), SVMs (margins), PSD constraints on covariance matrices, and projected-gradient training.
How is it used?
Either (1) turn the constraint into a penalty on the loss (Chapter 3.11), (2) take a gradient step and then project back onto the set (Chapter 3.17), or (3) hand it to a constrained solver. Pick by how cheap the projection is.
Projection is not "clip each number" for every set. For a box, clipping each coordinate is the projection. For the L2 ball you rescale; for the simplex you shift and clip; for the L1 ball you shrink toward the axes (soft thresholding, Chapter 3.13). Each set has its own rule.
Quick check: project $[3, -4]$ onto the L2 ball of radius $1$.
Its length is $\sqrt{9 + 16} = 5 \gt 1$, so scale it to length 1: $[3,-4]/5 = [0.6, -0.8]$.
A map of methods core
There is no single tool for constrained problems. There are six main roads, and which you take depends on the kind of constraint:
- Eliminate it (substitution): if the rule can be solved for a variable, remove the variable.
- Balance the forces (Lagrange, KKT): write the optimality conditions and solve them.
- Solve the twin (duality): maximise the dual function over the prices; for convex problems this gives the same answer.
- Walk and snap back (projected gradient): take a normal gradient step, then move to the closest allowed point.
- Charge a fine (penalty): add a growing cost for breaking the rule and minimise without constraints.
- Build an invisible fence (barrier): add a cost that explodes near the wall so iterates stay strictly inside.
All of these reach the same answer on the same problem. They differ in when they are cheap.
One problem, several roads: minimise $f = (x-3)^2 + (y-2)^2$ subject to $x + y = 4$.
- Substitution. $y = 4 - x$, so $f = (x-3)^2 + (2-x)^2$ and $f' = 2(x-3) - 2(2-x) = 4x - 10 = 0$, giving $x = 2.5$, $y = 1.5$.
- Lagrange. $2(x-3) + \lambda = 0$, $2(y-2) + \lambda = 0$, $x + y = 4$. The first two give $x - 3 = y - 2$, i.e. $x = y + 1$, so $2y + 1 = 4$: $y = 1.5$, $x = 2.5$, $\lambda = -2(x-3) = 1$.
- Projection. The bowl's centre $(3,2)$ has $x + y = 5$, one unit too much. Move it back along the normal $(1,1)$ by $\tfrac{5-4}{2} = 0.5$ per unit: $(3,2) - 0.5\cdot(1,1) = (2.5, 1.5)$.
- Penalty. Minimise $f + \tfrac{\rho}{2}(x + y - 4)^2$. By symmetry $x = y + 1$; the sum $s = x + y$ satisfies $(s - 5) + \rho(s - 4) = 0$, so $s = \dfrac{5 + 4\rho}{1 + \rho}$. For $\rho = 10$: $s = 4.09$, slightly off the line; as $\rho \to \infty$, $s \to 4$ and the answer tends to $(2.5, 1.5)$.
| Method | Handles | Core idea | Watch out for | Taught in |
|---|---|---|---|---|
| Substitution | equalities you can solve for a variable | eliminate variables, then minimise freely | needs the constraint solved by hand | here (3.7) |
| Lagrange multipliers | equalities | $\nabla f + \sum\lambda_j\nabla h_j = \mathbf{0}$, $h_j = 0$ | gives candidates, needs regularity | 3.8 |
| KKT conditions | inequalities and equalities | four conditions: stationarity, feasibility, $\mu \ge 0$, complementary slackness | needs a constraint qualification; sufficient only if convex | 3.9 |
| Duality | convex problems (and bounds for others) | maximise over prices the lowest priced cost | a gap if not convex | 3.10 |
| Projected gradient | sets with a cheap projection (box, ball, simplex) | gradient step, then project back | useless if projecting is as hard as the problem | 3.17 |
| Penalty methods | any constraints | add $\tfrac{\rho}{2}(\text{violation})^2$, increase $\rho$ | ill-conditioned as $\rho$ grows; slightly infeasible answers | 3.17 |
| Barrier (interior-point) | inequalities | add $-\tfrac1t\log(-g)$; stay strictly inside; increase $t$ | needs a strictly feasible start | 3.17 |
(Real solvers such as SciPy's SLSQP combine these ideas, repeatedly solving small KKT systems. You rarely implement them yourself.)
Why do we need it?
Picking the wrong method wastes days: a penalty method on a problem with an easy projection is slow, and projected gradient on a problem with a hard set is impossible. A map tells you where to look first.
Where is it used?
Choosing between SciPy's minimize methods, deciding how to enforce non-negativity or sum-to-one in a model, and understanding what an optimiser library does under its hood (SVM libraries, CVXPY, interior-point solvers).
How is it used?
Ask: (1) is it just equalities I can solve out? Substitute. (2) Is it a simple set? Projected gradient. (3) Small and smooth with a few constraints? A KKT-based solver. (4) Huge and soft? Penalty or regularisation. Then verify the answer with KKT.
"Penalty" and "Lagrange" are different. A penalty needs a large $\rho$ to be accurate and is never exactly feasible. A Lagrange multiplier is an exact price. The difference is why the augmented Lagrangian (a mix of both) is popular. Chapter 3.17 teaches penalty and barrier methods in full.
Quick check: which method would you try first for "minimise a loss over probability vectors"?
The feasible set is the simplex, which has a cheap, exact projection (shift and clip). So projected gradient is the natural first choice. If the probabilities come from a model, a softmax parametrisation removes the constraint altogether.
Recap, cheat sheet and practice
- Standard form: $\min f(\mathbf{x})$ subject to $g_i(\mathbf{x}) \le 0$ and $h_j(\mathbf{x}) = 0$. Multipliers: $\lambda_j$ (any sign) for equalities, $\mu_i \ge 0$ for inequalities. Lagrangian $L = f + \sum\lambda_jh_j + \sum\mu_ig_i$.
- The feasible region is where all rules hold. A constraint is active at a point if it is tight ($g_i = 0$), otherwise inactive.
- If the free minimum is forbidden, the answer lands on the boundary, where $\nabla f \ne \mathbf{0}$ and $-\nabla f$ points out through the wall. Constraints can only make $p^\star$ worse or equal.
- Lagrange: at an equality-constrained optimum $\nabla f + \sum\lambda_j\nabla h_j = \mathbf{0}$. KKT adds feasibility, $\mu_i \ge 0$ and complementary slackness $\mu_ig_i = 0$.
- Dual: $d(\boldsymbol\lambda,\boldsymbol\mu) = \min_{\mathbf{x}} L$; weak duality $d \le p^\star$ always; strong duality $d^\star = p^\star$ for convex problems with a strictly feasible point.
- A constraint qualification (LICQ, Slater, or linear constraints) is the fine print that makes KKT necessary. ML constraints: norm balls, simplex, boxes, margins. Methods: substitution, Lagrange, KKT, duality, then projected gradient, penalty and barrier (Chapter 3.17).
Cheat sheet
| Idea | Formula or rule | Picture |
|---|---|---|
| Standard form | $\min f$ s.t. $g_i \le 0,\; h_j = 0$ | objective + allowed region |
| Active / inactive | $g_i(\mathbf{x}) = 0$ / $g_i(\mathbf{x}) \lt 0$ | touching the wall / room to spare |
| Lagrange condition | $\nabla f + \sum\lambda_j\nabla h_j = \mathbf{0}$ | level curve tangent to the constraint |
| Lagrangian | $f + \sum\lambda_jh_j + \sum\mu_ig_i$ | the problem with prices on the rules |
| KKT | stationarity; $g\le0,h=0$; $\mu\ge0$; $\mu g=0$ | forces balance, walls push inward only |
| Dual function | $d = \min_{\mathbf{x}}L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu)$, $\boldsymbol\mu\ge0$ | cheapest priced cost |
| Weak / strong duality | $d \le p^\star$ / $d^\star = p^\star$ | lower bound / bound is exact |
| Constraint qualification | LICQ, Slater, linear constraints | wall has a well-defined direction |
| Projection onto L2 ball | $\mathbf{x}\cdot\min(1, r/\|\mathbf{x}\|)$ | scale back to the sphere |
import numpy as np
from scipy.optimize import minimize
# Problem: minimise (x-3)^2 + (y-2)^2 subject to x^2 + y^2 <= 4
f = lambda z: (z[0] - 3) ** 2 + (z[1] - 2) ** 2
grad_f = lambda z: np.array([2 * (z[0] - 3), 2 * (z[1] - 2)])
g = lambda z: z[0] ** 2 + z[1] ** 2 - 4 # allowed when g <= 0
grad_g = lambda z: np.array([2 * z[0], 2 * z[1]])
# SciPy writes inequality constraints as fun(z) >= 0, so we pass -g.
res = minimize(f, x0=[0.0, 0.0], method="SLSQP",
constraints=[{"type": "ineq", "fun": lambda z: -g(z)}])
x = res.x
print(np.round(x, 4), round(f(x), 4)) # [1.6641 1.1094] 2.5778
print(round(g(x), 6)) # 0.0 -> the constraint is active
# Recover the multiplier from stationarity grad f + mu * grad g = 0 (least squares)
mu = -grad_f(x) @ grad_g(x) / (grad_g(x) @ grad_g(x))
print(round(mu, 4), round((np.sqrt(13) - 2) / 2, 4)) # 0.8028 0.8028
print(np.allclose(grad_f(x) + mu * grad_g(x), 0, atol=1e-5)) # True stationarity holds
print(mu >= 0, abs(mu * g(x)) < 1e-6) # True True dual feasibility, complementary slackness
# Bigger disc: the free minimum (3, 2) is allowed, the constraint becomes inactive
res2 = minimize(f, x0=[0.0, 0.0], method="SLSQP",
constraints=[{"type": "ineq", "fun": lambda z: 16 - (z[0] ** 2 + z[1] ** 2)}])
print(np.round(res2.x, 3), round(f(res2.x), 6)) # [3. 2.] 0.0
# Dual of min x^2 s.t. x >= 1: d(mu) = mu - mu^2 / 4
d = lambda m: m - m ** 2 / 4
mus = np.linspace(0, 5, 501)
print(mus[np.argmax(d(mus))], d(mus).max()) # 2.0 1.0 (= p*)
print(np.all(d(mus) <= 1 + 1e-12)) # True weak duality: every d(mu) <= p*
# Projection onto the simplex (probabilities that sum to 1) via sorting
def project_simplex(v):
u = np.sort(v)[::-1]
css = np.cumsum(u) - 1
k = np.nonzero(u - css / (np.arange(len(v)) + 1) > 0)[0][-1]
return np.maximum(v - css[k] / (k + 1), 0)
print(project_simplex(np.array([0.8, 0.6, -0.2]))) # [0.6 0.4 0. ]
1. "Maximise $3x + 2y$ subject to $x + y \ge 4$" in standard form is…
2. You minimise $x^2$ subject to $x \ge -1$. At the answer, the constraint is…
3. At the optimum of a constrained problem where the answer is pushed onto a wall, the gradient of the objective is…
4. Complementary slackness says that for each inequality constraint…
5. For a problem with $p^\star = 5$, which dual value is impossible (with $\boldsymbol\mu \ge 0$)?
6. Why can minimising $x$ subject to $x^2 \le 0$ have an optimum but no KKT multiplier?
Practice problems
A. Write "minimise $x^2 + y^2$ subject to $x + 2y \ge 3$, $x \ge 0$ and $x - y = 1$" in standard form.
$f = x^2 + y^2$. Inequalities: $g_1 = 3 - x - 2y \le 0$ and $g_2 = -x \le 0$. Equality: $h_1 = x - y - 1 = 0$. Constraints: $m = 2$ inequalities, $p = 1$ equality. Multipliers: $\mu_1, \mu_2 \ge 0$ and $\lambda_1$ of any sign. $L = x^2 + y^2 + \lambda_1(x - y - 1) + \mu_1(3 - x - 2y) + \mu_2(-x)$.
B. For $g_1 = x + y - 4 \le 0$, $g_2 = -x \le 0$, $g_3 = -y \le 0$, classify $(3,1)$, $(0,3)$ and $(2,3)$.
$(3,1)$: $g_1 = 0$, $g_2 = -3$, $g_3 = -1$: feasible, boundary ($g_1$ active). $(0,3)$: $g_1 = -1$, $g_2 = 0$, $g_3 = -3$: feasible, boundary ($g_2$ active). $(2,3)$: $g_1 = 1 \gt 0$: not feasible.
C. Solve "minimise $x^2 + y^2$ subject to $x + y \ge 2$" with KKT.
Standard form: $g = 2 - x - y \le 0$. $L = x^2 + y^2 + \mu(2 - x - y)$. Stationarity: $2x - \mu = 0$, $2y - \mu = 0$, so $x = y = \mu/2$. Complementary slackness: either $\mu = 0$ or $g = 0$. If $\mu = 0$: $x = y = 0$, but then $g = 2 \gt 0$ (infeasible). So $g = 0$: $x + y = 2$, $x = y = 1$, $\mu = 2 \ge 0$ ✓. Answer: $(1,1)$ with $p^\star = 2$; the constraint is active. (It agrees with the equality problem, because the free minimum $(0,0)$ is forbidden, so the answer is on the wall.)
D. Find the dual function and dual optimum of "minimise $x^2 + y^2$ subject to $x + y = 2$".
$L = x^2 + y^2 + \lambda(x + y - 2)$. Minimise over $(x,y)$: $x = y = -\lambda/2$. Then $d(\lambda) = \tfrac{\lambda^2}{4} + \tfrac{\lambda^2}{4} + \lambda(-\lambda - 2) = -\tfrac{\lambda^2}{2} - 2\lambda$. Maximise: $-\lambda - 2 = 0$ gives $\lambda = -2$ and $d(-2) = -2 + 4 = 2 = p^\star$. No gap, as the problem is convex with linear constraints.
E. Solve "minimise $x + y$ subject to $x^2 + y^2 \le 2$" and give $\mu$.
The free problem has no minimum (unbounded), so the wall must be active. $L = x + y + \mu(x^2 + y^2 - 2)$: $1 + 2\mu x = 0$ and $1 + 2\mu y = 0$, so $x = y = -1/(2\mu)$ (this needs $\mu \gt 0$). Active: $x^2 + y^2 = 2$ gives $x = y = \pm1$; with $\mu \gt 0$ we need $x = y = -1$ and $\mu = 1/2$. Value $p^\star = -2$. LICQ holds ($\nabla g = (-2,-2) \ne \mathbf{0}$).
F. Minimise $(x-1)^2 + (y-1)^2$ subject to $x + y \le 1$. Where is the answer, what is $\mu$, and what is $p^\star$?
The free minimum $(1,1)$ has $x + y = 2 \gt 1$: forbidden, so the wall $g = x + y - 1 = 0$ is active. By symmetry $x = y = 0.5$. Then $\nabla f = (2(0.5-1),\,2(0.5-1)) = (-1,-1)$ and $\nabla g = (1,1)$, so $\nabla f + \mu\nabla g = \mathbf{0}$ gives $\mu = 1 \ge 0$ ✓. $p^\star = 0.25 + 0.25 = 0.5$, which is worse than the free value $0$.
Lagrange Multipliers
The classic tool for problems with equality constraints. One beautiful picture explains it: at the best point, the contour of the objective just touches the constraint curve, so their gradients point the same way. This chapter turns that picture into a recipe you can run by hand, then uses it for real problems, including one that produces the softmax.
- See why $\nabla f$ must be parallel to $\nabla h$ at a constrained optimum (tangent level curves), and derive it
- Handle one equality constraint, then several, and know the regularity condition that makes it work
- Build the Lagrangian, understand stationarity, and see why $L$ is a saddle, not a bowl, in $(\mathbf{x},\boldsymbol\lambda)$
- Solve constrained problems with a clear recipe, and compare all candidates (not every stationary point is the minimum)
- Work through the classics: biggest rectangle, closest point on a line or circle, minimum norm, the best unit direction (PCA), maximum entropy (which gives the uniform distribution and softmax)
- Read the multiplier $\lambda$ as a sensitivity: how fast the best value changes when the constraint moves
The geometry: contours that just touch core
Picture a hiking map with contour lines (lines of equal height; see level sets in the Calculus guide). You must walk along one fixed path, such as a straight road or a circular track. Where on the path is the ground lowest?
Look at how the path meets the contour lines:
- If the path crosses a contour line, then walking one way along the path takes you to a higher contour, and walking the other way takes you to a lower one. You can still go down. This is not the lowest point.
- You can no longer improve only when the path just touches a contour line, running along it for a moment without crossing it. There the path and the contour line are tangent (they have the same direction).
Think of the contour as a balloon that you inflate from the lowest level upward. The first moment the balloon touches the path is the lowest point of the path.
The nearest point on a line. Minimise $f(x,y) = x^2 + y^2$ (the squared distance to the origin) on the line $x + y = 2$. The contours of $f$ are circles around the origin.
- A tiny circle (radius $0.5$) is too small to reach the line. A circle of radius $2$ cuts the line at two points. Between these two, the line goes inside the circle, so there are points on the line with a smaller $f$.
- Shrink the circle until it just touches the line. That happens at radius $\sqrt{2} \approx 1.414$, at the point $(1,1)$. There $f = 2$.
- At the touching point, the circle's tangent direction is $(1,-1)$, the same as the line's direction. The line's normal $(1,1)$ and the circle's normal ($\nabla f = (2x,2y) = (2,2)$) point along the same line.
"Tangent curves" and "parallel gradients" are the same statement, because a gradient is perpendicular to its level curve (contour maps).
Geometric condition. Let $\mathbf{x}^\star$ be a local optimum of $f$ on the curve $h(\mathbf{x}) = 0$, and suppose $\nabla h(\mathbf{x}^\star) \ne \mathbf{0}$. Then:
- the level curve of $f$ through $\mathbf{x}^\star$ is tangent to the constraint curve, and
- equivalently, $\nabla f(\mathbf{x}^\star)$ is parallel to $\nabla h(\mathbf{x}^\star)$: one is a multiple of the other.
Moving along the constraint cannot improve $f$ to first order, because along the constraint you move perpendicular to $\nabla h$, hence (by parallelism) perpendicular to $\nabla f$, and the objective does not change at first order. The condition identifies candidates: it also holds at maxima and saddle-like points of $f$ on the curve.
Why do we need it?
It replaces a hard question ("which point of a curve has the lowest $f$?") by a clear test you can check at any point: are two gradients parallel? It also explains why the test is correct, not just how to apply it.
Where is it used?
Every Lagrange-multiplier derivation: principal component analysis (best unit direction), maximum-entropy models, constrained least squares, and the geometric explanation of why the Lasso (L1) ball produces sparse solutions: the loss contours touch the ball at a corner.
How is it used?
Draw (or imagine) the contours of $f$ and the constraint curve. Look for places where they touch without crossing. At such a point compute both gradients and confirm they are parallel.
Parallel does not mean minimum. At a maximum of $f$ on the curve the gradients are parallel too (the balloon touches from the inside). The geometric condition finds all tangency points, so you must still compare them (Section "Solving constrained problems").
Gradient zero is allowed. If $\nabla f = \mathbf{0}$ at a point of the curve, then it is trivially "parallel" to everything ($\lambda = 0$). That happens when the free minimum happens to lie on the constraint.
Quick check: on the circle $x^2 + y^2 = 1$, for $f = x$, which points have parallel gradients?
$\nabla f = (1, 0)$ and $\nabla h = (2x, 2y)$. They are parallel when $y = 0$, so $(1,0)$ and $(-1,0)$. At $(1,0)$ the function $f = x$ is largest, at $(-1,0)$ smallest.
One equality constraint: the rule $\nabla f = -\lambda\nabla h$ core
The picture says: at the best point, $\nabla f$ is a multiple of $\nabla h$. Writing the multiple as $-\lambda$ (the minus sign is only a convention, chosen so that the Lagrangian reads $f + \lambda h$) gives one equation per unknown:
- $n$ equations from $\nabla f + \lambda\nabla h = \mathbf{0}$ (one for each coordinate),
- plus the constraint $h = 0$ itself,
- for $n + 1$ unknowns: the $n$ coordinates of $\mathbf{x}$ and the number $\lambda$.
That is a square system you can try to solve. The multiplier $\lambda$ is an extra unknown we are willing to pay for, because it makes every equation simple.
Biggest rectangle with a fixed fence. A rectangle has sides $x$ and $y$ and a perimeter of 20, so $x + y = 10$. Maximise the area $xy$. Equivalent: minimise $f = -xy$ subject to $h = x + y - 10 = 0$.
- Gradients: $\nabla f = (-y,\,-x)$ and $\nabla h = (1,\,1)$.
- The rule $\nabla f + \lambda\nabla h = \mathbf{0}$ gives $-y + \lambda = 0$ and $-x + \lambda = 0$. So $x = y = \lambda$.
- The constraint: $x + y = 10$ gives $2\lambda = 10$, so $\lambda = 5$ and $x = y = 5$.
- The best rectangle is the square $5 \times 5$ with area $25$.
- Check by substitution: $y = 10 - x$ gives area $x(10 - x) = 10x - x^2$, with derivative $10 - 2x = 0$, so $x = 5$ and area $25$ ✓. Check the neighbours: $4 \times 6 = 24$, $3 \times 7 = 21$, both smaller.
Theorem (Lagrange, one constraint). Let $f$ and $h$ be continuously differentiable. If $\mathbf{x}^\star$ is a local minimiser of $f$ subject to $h(\mathbf{x}) = 0$ and $\nabla h(\mathbf{x}^\star) \ne \mathbf{0}$, then there is a number $\lambda$ with
$$\nabla f(\mathbf{x}^\star) + \lambda\,\nabla h(\mathbf{x}^\star) = \mathbf{0}, \qquad h(\mathbf{x}^\star) = 0.$$Derivation.
- Take any smooth path $\mathbf{x}(t)$ that stays on the constraint curve and passes through $\mathbf{x}^\star$ at $t = 0$, with velocity $\mathbf{v} = \mathbf{x}'(0)$.
- Since $h(\mathbf{x}(t)) = 0$ for all $t$, differentiate with the chain rule (Chapter 2.8): $\nabla h\cdot\mathbf{v} = 0$. So every such velocity is perpendicular to $\nabla h$ (it is a tangent direction).
- Along this path, the function $\varphi(t) = f(\mathbf{x}(t))$ has a local minimum at $t = 0$ (because $\mathbf{x}^\star$ is the best feasible point). So $\varphi'(0) = \nabla f\cdot\mathbf{v} = 0$.
- Every vector perpendicular to $\nabla h$ is the velocity of some path on the constraint (this is what $\nabla h \ne \mathbf{0}$ buys us; it is the implicit function theorem). So $\nabla f$ is perpendicular to every vector that is perpendicular to $\nabla h$.
- The vectors perpendicular to all vectors perpendicular to $\nabla h$ form exactly the line spanned by $\nabla h$ (orthogonal complements of orthogonal complements). Hence $\nabla f = -\lambda\nabla h$ for some number $\lambda$. $\blacksquare$
For a maximum the same equations hold (just apply the theorem to $-f$, which changes the sign of $\lambda$).
Why do we need it?
Substitution fails when you cannot solve the constraint for a variable (or when doing so spoils the symmetry). The multiplier rule works for any smooth constraint, treats all variables the same way, and produces equations that are usually simple.
Where is it used?
Maximising area or volume with a fixed material budget, portfolio weights that must sum to one, "best unit vector" problems (PCA, the largest eigenvalue as a constrained maximum), and maximum-likelihood problems with a normalisation condition.
How is it used?
Write $h = 0$ in the form "something $= 0$". Write $\nabla f + \lambda\nabla h = \mathbf{0}$ coordinate by coordinate. Solve these equations together with $h = 0$. Then check the solutions.
Regularity matters. The proof needs $\nabla h(\mathbf{x}^\star) \ne \mathbf{0}$. At a point where $\nabla h = \mathbf{0}$ (a cusp, a crossing, a pinch) the tangent directions are not well defined and the rule may fail (see the pathological example in Chapter 3.7).
The sign of $\lambda$ depends on the convention. Here the Lagrangian is $f + \lambda h$. With $f - \lambda h$ you would get $-\lambda$. Always check which convention a book uses before comparing numbers.
Quick check: maximise $x + y$ on the circle $x^2 + y^2 = 2$ with the multiplier rule.
Minimise $f = -(x + y)$ with $h = x^2 + y^2 - 2$. $\nabla f = (-1,-1)$, $\nabla h = (2x, 2y)$. So $-1 + 2\lambda x = 0$ and $-1 + 2\lambda y = 0$, giving $x = y = 1/(2\lambda)$. Then $x^2 + y^2 = 2$ gives $2/(4\lambda^2) = 2$, so $\lambda = \pm\tfrac12$. For $\lambda = \tfrac12$: $(1,1)$ with $x + y = 2$ (the maximum). For $\lambda = -\tfrac12$: $(-1,-1)$ with $x + y = -2$ (the minimum).
Worked example: the closest point on a line, and minimum norm core
"What is the closest point on a straight road to my house?" is the most common constrained problem in the world. You know the answer in your gut: drop a perpendicular from the house to the road. The multiplier method gives this same answer by algebra, and, for a general line $ax + by = c$, hands you a clean formula.
Closest point on $2x + y = 10$ to the point $(1, 2)$. Minimise $f = (x-1)^2 + (y-2)^2$ subject to $h = 2x + y - 10 = 0$.
- $\nabla f = (2(x-1),\,2(y-2))$ and $\nabla h = (2, 1)$.
- $\nabla f + \lambda\nabla h = \mathbf{0}$: $2(x-1) + 2\lambda = 0$ gives $x = 1 - \lambda$; and $2(y-2) + \lambda = 0$ gives $y = 2 - \lambda/2$.
- Put these into the constraint: $2(1 - \lambda) + (2 - \lambda/2) = 10$, so $4 - 2.5\lambda = 10$ and $\lambda = -2.4$.
- Then $x = 1 + 2.4 = 3.4$ and $y = 2 + 1.2 = 3.2$. Check: $2(3.4) + 3.2 = 10$ ✓.
- Squared distance: $(3.4 - 1)^2 + (3.2 - 2)^2 = 2.4^2 + 1.2^2 = 5.76 + 1.44 = 7.2$, so the distance is $\sqrt{7.2} \approx 2.683$. The vector from the house to the answer is $(2.4, 1.2) = 1.2\,(2,1)$: parallel to the line's normal $(2,1)$, i.e. perpendicular to the road ✓.
Minimum norm on a line. Minimise $x^2 + y^2$ subject to $ax + by = c$ (with $(a,b) \ne (0,0)$). Then $\nabla f = (2x,2y)$, $\nabla h = (a,b)$, so $2x + \lambda a = 0$ and $2y + \lambda b = 0$, i.e. $(x,y) = -\tfrac{\lambda}{2}(a,b)$. The constraint gives $-\tfrac{\lambda}{2}(a^2 + b^2) = c$, so
$$\lambda = \frac{-2c}{a^2 + b^2}, \qquad (x^\star, y^\star) = \frac{c}{a^2 + b^2}\,(a, b), \qquad \min f = \frac{c^2}{a^2 + b^2}, \qquad \text{distance} = \frac{|c|}{\sqrt{a^2 + b^2}}.$$Vector form. Minimise $\|\mathbf{x}\|^2$ subject to $A\mathbf{x} = \mathbf{b}$ ($A$ has independent rows). With $L = \|\mathbf{x}\|^2 + \boldsymbol\lambda^\top(A\mathbf{x} - \mathbf{b})$: $2\mathbf{x} + A^\top\boldsymbol\lambda = \mathbf{0}$, so $\mathbf{x} = -\tfrac12 A^\top\boldsymbol\lambda$; plugging into $A\mathbf{x} = \mathbf{b}$ gives $\boldsymbol\lambda = -2(AA^\top)^{-1}\mathbf{b}$ and $$\mathbf{x}^\star = A^\top (AA^\top)^{-1}\mathbf{b}.$$ This is the minimum-norm solution of an underdetermined linear system, the same formula as the pseudoinverse (Chapter 1.10).
Why do we need it?
"Closest point satisfying a linear rule" and "the smallest solution that fits the data exactly" appear again and again. The formulas above solve them in one line, with no iteration.
Where is it used?
Minimum-norm solutions of underdetermined systems (more unknowns than equations), the geometry of projecting a point onto a hyperplane, the step in projected-gradient methods, and why ridge-like "smallest weights" solutions arise when many weight vectors fit the data.
How is it used?
For one line, plug into $\mathbf{x}^\star = \tfrac{c}{a^2+b^2}(a,b)$. For several linear constraints, solve $(AA^\top)\mathbf{u} = \mathbf{b}$ and take $\mathbf{x}^\star = A^\top\mathbf{u}$. To project a general point $\mathbf{z}$, apply the same idea to $\mathbf{x} - \mathbf{z}$.
Quick check: use the formula for the nearest point to the origin on $x + 2y = 5$.
$a = 1$, $b = 2$, $c = 5$, $a^2 + b^2 = 5$. So $(x^\star, y^\star) = \tfrac{5}{5}(1,2) = (1,2)$, the minimum of $x^2+y^2$ is $c^2/(a^2+b^2) = 25/5 = 5$ (check: $1 + 4 = 5$ ✓), and $\lambda = -2\cdot5/5 = -2$.
Several constraints at once core
With one constraint, you walk on a surface and the "wall" has one normal direction. With two constraints in 3D (say two planes), you are confined to the line where the planes meet. The walls now have two normal directions, $\nabla h_1$ and $\nabla h_2$, and together they span a whole plane of "blocked" directions. Only the direction along the line is free.
At the best point on that line, the objective's gradient has no component along the line. So $\nabla f$ lies entirely in the plane spanned by the wall normals: it is a combination of them: $\nabla f = -\lambda_1\nabla h_1 - \lambda_2\nabla h_2$. One multiplier per constraint.
Closest point to the origin on a line in 3D. Minimise $f = x^2 + y^2 + z^2$ subject to $h_1 = x + y + z - 6 = 0$ and $h_2 = x + 2y + 3z - 10 = 0$.
- Gradients: $\nabla f = (2x,2y,2z)$, $\nabla h_1 = (1,1,1)$, $\nabla h_2 = (1,2,3)$.
- Stationarity $\nabla f + \lambda_1\nabla h_1 + \lambda_2\nabla h_2 = \mathbf{0}$ gives $2x + \lambda_1 + \lambda_2 = 0$, $\;2y + \lambda_1 + 2\lambda_2 = 0$, $\;2z + \lambda_1 + 3\lambda_2 = 0$.
- So $x = -\tfrac12(\lambda_1 + \lambda_2)$, $y = -\tfrac12(\lambda_1 + 2\lambda_2)$, $z = -\tfrac12(\lambda_1 + 3\lambda_2)$.
- Constraint 1: $x + y + z = -\tfrac12(3\lambda_1 + 6\lambda_2) = 6$, i.e. $\lambda_1 + 2\lambda_2 = -4$.
- Constraint 2: $x + 2y + 3z = -\tfrac12(6\lambda_1 + 14\lambda_2) = 10$, i.e. $3\lambda_1 + 7\lambda_2 = -10$.
- Solve the $2\times2$ system: from the first, $\lambda_1 = -4 - 2\lambda_2$; then $3(-4 - 2\lambda_2) + 7\lambda_2 = -10$ gives $\lambda_2 = 2$ and $\lambda_1 = -8$.
- Then $x = -\tfrac12(-8 + 2) = 3$, $y = -\tfrac12(-8 + 4) = 2$, $z = -\tfrac12(-8 + 6) = 1$.
- Check: $3 + 2 + 1 = 6$ ✓, $3 + 4 + 3 = 10$ ✓, and $\nabla f = (6,4,2) = 8\,(1,1,1) - 2\,(1,2,3)$ ✓. The minimum value is $9 + 4 + 1 = 14$. (Cross-check with $\mathbf{x}^\star = A^\top(AA^\top)^{-1}\mathbf{b}$ from the last section: it gives $(3,2,1)$ as well.)
For constraints $h_1,\dots,h_p$ and a local minimiser $\mathbf{x}^\star$ of $f$ subject to $h_j(\mathbf{x}) = 0$ for all $j$:
$$\nabla f(\mathbf{x}^\star) + \sum_{j=1}^{p}\lambda_j\,\nabla h_j(\mathbf{x}^\star) = \mathbf{0}, \qquad h_j(\mathbf{x}^\star) = 0\;\;(j = 1,\dots,p).$$That is $n + p$ equations for the $n + p$ unknowns $(\mathbf{x},\boldsymbol\lambda)$. In matrix form, with the Jacobian $J$ whose rows are the $\nabla h_j^\top$ (the Jacobian): $\nabla f + J^\top\boldsymbol\lambda = \mathbf{0}$.
Regularity condition (also called LICQ, linear independence constraint qualification): the gradients $\nabla h_1(\mathbf{x}^\star),\dots,\nabla h_p(\mathbf{x}^\star)$ must be linearly independent (the matrix $J$ has full row rank $p$). Then the tangent space is exactly the set of $\mathbf{d}$ with $J\mathbf{d} = \mathbf{0}$, the theorem holds, and the multipliers are unique. If the gradients are dependent, the rule can fail, or multipliers can exist but not be unique.
Why do we need it?
Real problems have many rules at once (sum to one and match the mean). We need to know that one multiplier per rule is enough, and what could go wrong: dependent rules make the multipliers ambiguous and the theory shaky.
Where is it used?
Maximum-entropy models with several moment constraints (the exponential family), constrained least squares, structure from motion and robotics (several kinematic constraints), and every solver for equality-constrained quadratic programs, which solves exactly this linear system.
How is it used?
Introduce one $\lambda_j$ per constraint. Write $\nabla f + \sum\lambda_j\nabla h_j = \mathbf{0}$. Check that the constraint gradients are independent (rank of $J$ equals $p$). Solve the whole system. For quadratic $f$ and linear constraints it is one linear solve.
More constraints than freedom. You need $p \le n$ independent constraints. If $p = n$ independent equations are imposed, usually only a single point (or a few points) is feasible and there is nothing left to optimise.
Redundant constraints. Writing the same rule twice (or a multiple of it) makes $J$ lose rank. The solution is unchanged, but the multipliers are no longer unique and a numerical solver may complain about a singular matrix. Remove the duplicate.
Quick check: with $h_1 = x + y - 2$ and $h_2 = 2x + 2y - 4$, are the gradients independent?
$\nabla h_1 = (1,1)$ and $\nabla h_2 = (2,2) = 2\nabla h_1$. They are dependent (the second constraint is the first one written again, doubled), so the regularity condition fails and the multipliers are not unique: only $\lambda_1 + 2\lambda_2$ is determined.
The Lagrangian, and why it is a saddle core
Remember the trick from Chapter 3.7: bundle the problem and the rules into one function with prices $\lambda_j$. The point of doing so is a bookkeeping miracle: all $n + p$ conditions ($n$ stationarity equations and $p$ constraints) are just "every partial derivative of $L$ is zero". The Lagrangian is a machine that turns "optimise with rules" into "find where a single function is flat".
But be careful. Flat does not mean "bottom of a bowl". In the $\mathbf{x}$ direction the Lagrangian has a valley, but in the $\lambda$ direction it is a straight line, so the point we want is on a saddle: a pass between two hills, flat in both directions, uphill one way and downhill the other.
The smallest example: minimise $f = x^2$ subject to $h = x - 1 = 0$ (the answer is obviously $x = 1$). Then $L(x,\lambda) = x^2 + \lambda(x - 1)$.
- $\dfrac{\partial L}{\partial x} = 2x + \lambda = 0$ and $\dfrac{\partial L}{\partial \lambda} = x - 1 = 0$.
- The second equation is the constraint, $x = 1$. The first gives $\lambda = -2$. So the stationary point is $(x,\lambda) = (1,-2)$ with $L = 1 + (-2)(0) = 1 = f(1)$.
- Slice along $x$ (fix $\lambda = -2$): $L = x^2 - 2x + 2$, a parabola with its minimum at $x = 1$. A valley.
- Slice along $\lambda$ (fix $x = 1$): $L = 1 + \lambda\cdot0 = 1$ for every $\lambda$. Flat. For $x \ne 1$ the slice would be a tilted line.
- The matrix of second derivatives in $(x,\lambda)$ is $\begin{bmatrix} 2 & 1 \\ 1 & 0\end{bmatrix}$, with eigenvalues $1 \pm \sqrt2 \approx 2.414$ and $-0.414$: one positive, one negative, so $(1,-2)$ is a saddle point of $L$.
The Lagrangian of $\min f(\mathbf{x})$ subject to $h_j(\mathbf{x}) = 0$ is $L(\mathbf{x},\boldsymbol\lambda) = f(\mathbf{x}) + \sum_j\lambda_jh_j(\mathbf{x})$. Its partial derivatives are
$$\nabla_{\mathbf{x}}L = \nabla f + \sum_j\lambda_j\nabla h_j, \qquad \frac{\partial L}{\partial\lambda_j} = h_j(\mathbf{x}).$$Theorem (restated). The candidates of the previous sections are exactly the stationary points of $L$ in $(\mathbf{x},\boldsymbol\lambda)$: $\nabla_{\mathbf{x}}L = \mathbf{0}$ and $\nabla_{\boldsymbol\lambda}L = \mathbf{0}$.
Saddle structure. At a constrained local minimiser with its multiplier, $L(\mathbf{x}^\star,\boldsymbol\lambda) = f(\mathbf{x}^\star)$ for every $\boldsymbol\lambda$ (the constraint terms vanish), and in $\mathbf{x}$ (restricted to the tangent directions) $L$ curves upward. So $L$ is a minimum in the $\mathbf{x}$ directions along the constraint but is not minimised in $\boldsymbol\lambda$: in $\boldsymbol\lambda$ it is flat there and linear elsewhere. The Hessian of $L$ in $(\mathbf{x},\boldsymbol\lambda)$ is generally indefinite.
Why do we need it?
One function with all the information lets us use ordinary calculus ("set all partial derivatives to zero") and ordinary linear solvers. It also leads straight to duality (Chapter 3.10) and to saddle-point algorithms.
Where is it used?
Newton's method on the KKT system in constrained solvers, augmented-Lagrangian and ADMM methods (distributed training, large-scale convex problems), and min-max problems such as GAN training, which are also saddle problems.
How is it used?
Form $L$. Set $\nabla_{\mathbf{x}}L = \mathbf{0}$ and $\nabla_{\boldsymbol\lambda}L = \mathbf{0}$. Solve. Do not run plain gradient descent on $L$ in all variables: it would try to minimise in $\boldsymbol\lambda$ too, which has no minimum. Use descent in $\mathbf{x}$ and ascent in $\boldsymbol\lambda$, or solve the equations directly.
Minimising $L$ over all variables does not work. For fixed $\mathbf{x}$ the Lagrangian is a straight line in $\lambda$ (slope $h(\mathbf{x})$), which is unbounded below unless $h = 0$. The correct picture is a min over $\mathbf{x}$ and a max over $\lambda$ (the saddle), which is exactly the primal and dual structure of Chapter 3.10.
Stationary is only necessary. Not every stationary point of $L$ is a constrained minimum (it can be a maximum or something else). That is the next two sections.
Quick check: for $f = x^2 + y^2$ and $h = x + y - 2$, what are the partial derivatives of $L$?
$L = x^2 + y^2 + \lambda(x + y - 2)$. So $\partial L/\partial x = 2x + \lambda$, $\partial L/\partial y = 2y + \lambda$ and $\partial L/\partial\lambda = x + y - 2$. Setting all three to zero gives $x = y = 1$, $\lambda = -2$.
Stationarity: the constrained version of "gradient = 0" core
Without constraints, a point is stationary when $\nabla f = \mathbf{0}$: no direction is downhill. With constraints you may only move along the constraint, so the question becomes: is any allowed direction downhill?
Split the gradient into two pieces: the part pointing along the walls' normals (blocked, nothing you can do about it) and the part along the tangent (the sliding part, the only part that matters). A point is stationary when the sliding part is zero. The multipliers $\lambda_j$ are the numbers that measure the blocked part.
Minimise $f = x^2 + y^2$ on $h = x + y - 2 = 0$. Test two feasible points.
- $(2, 0)$: $\nabla f = (4, 0)$, $\nabla h = (1,1)$. The best $\lambda$ minimises $\|\nabla f + \lambda\nabla h\|$: $\lambda = -\dfrac{\nabla f\cdot\nabla h}{\|\nabla h\|^2} = -\dfrac{4}{2} = -2$. The leftover is $\nabla f + \lambda\nabla h = (4,0) - 2(1,1) = (2,-2)$, which is not zero: it points along the constraint line. Sliding in the direction $-(2,-2) = (-2, 2)$ lowers $f$. Not stationary.
- $(1, 1)$: $\nabla f = (2,2)$, $\lambda = -\tfrac{4}{2} = -2$, leftover $(2,2) - 2(1,1) = (0,0)$. Stationary ✓.
The leftover $\mathbf{r} = \nabla f + \lambda\nabla h$ (with the best $\lambda$) is the tangential part of the gradient. It is called the reduced gradient. Stationarity means $\mathbf{r} = \mathbf{0}$.
A feasible point $\mathbf{x}$ is stationary for $\min f$ subject to $h_j = 0$ if there are multipliers with $\nabla_{\mathbf{x}}L(\mathbf{x},\boldsymbol\lambda) = \nabla f + \sum_j\lambda_j\nabla h_j = \mathbf{0}$.
Equivalent forms:
- $\nabla f(\mathbf{x}) \in \operatorname{span}\{\nabla h_1,\dots,\nabla h_p\}$.
- $\nabla f(\mathbf{x})\cdot\mathbf{d} = 0$ for every tangent direction $\mathbf{d}$ (every $\mathbf{d}$ with $\nabla h_j\cdot\mathbf{d} = 0$ for all $j$).
- The projection of $\nabla f$ onto the tangent space is $\mathbf{0}$ (the "reduced gradient" is zero).
At a non-stationary feasible point, moving along $-\mathbf{r}$ (the negative reduced gradient), then correcting back onto the constraint, decreases $f$. This is the idea behind projected gradient and gradient methods for constrained problems (Chapter 3.17).
Why do we need it?
It gives a numeric test ("is the reduced gradient zero?") that works at any point, and a stopping rule for algorithms: stop when the leftover $\|\mathbf{r}\|$ is tiny.
Where is it used?
Convergence checks in constrained solvers (the "first-order optimality" number printed by solvers such as SciPy's trust-constr and MATLAB's fmincon is closely related to the size of this leftover), reduced-gradient and projected-gradient algorithms, and gradient checks for constrained models.
How is it used?
At a candidate: compute $\nabla f$ and the constraint gradients, find the best $\boldsymbol\lambda$ by least squares ($J J^\top\boldsymbol\lambda = -J\nabla f$), and look at the size of $\nabla f + J^\top\boldsymbol\lambda$.
Zero reduced gradient is not zero gradient. At a stationary point of a constrained problem, $\nabla f$ is typically not zero. It is the multiplier part $-\sum\lambda_j\nabla h_j$ that is non-zero, and it is balanced by the wall.
Slide-downhill finds a local, not a global, answer. Which stationary point you reach depends on where you start (just like gradient descent in Chapter 3.3).
Quick check: is $(0.5, 1.5)$ stationary for $\min x^2 + y^2$ on $x + y = 2$?
$\nabla f = (1, 3)$ and $\nabla h = (1,1)$. Best $\lambda = -\tfrac{1\cdot1 + 3\cdot1}{2} = -2$. Leftover $= (1,3) - 2(1,1) = (-1, 1) \ne \mathbf{0}$. Not stationary. (Moving in direction $(1,-1)$ toward $(1,1)$ lowers $f$.)
Solving constrained optimization problems: a recipe core
All the ideas so far fit into one routine, like a cooking recipe. The recipe does not hand you the answer directly. It hands you a short list of candidates, and then you pick the winner by comparing them.
Why a list? Because "gradients are parallel" is true at the lowest point, at the highest point, and at other special points. Think of a bumpy ring-shaped road: it has several places where the road runs level, but only one is the lowest.
Closest point on a circle. Find the points of the circle $x^2 + y^2 = 4$ that are nearest to, and farthest from, $(3, 2)$. Minimise $f = (x-3)^2 + (y-2)^2$ subject to $h = x^2 + y^2 - 4 = 0$.
- Lagrangian: $L = (x-3)^2 + (y-2)^2 + \lambda(x^2 + y^2 - 4)$.
- $\partial L/\partial x = 2(x-3) + 2\lambda x = 0$ gives $x(1 + \lambda) = 3$. $\;\partial L/\partial y = 2(y-2) + 2\lambda y = 0$ gives $y(1+\lambda) = 2$.
- So $(x,y) = \dfrac{(3,2)}{1 + \lambda}$ (note $1 + \lambda \ne 0$, otherwise $0 = 3$). The constraint gives $\dfrac{9 + 4}{(1 + \lambda)^2} = 4$, so $(1+\lambda)^2 = \dfrac{13}{4}$ and $1 + \lambda = \pm\dfrac{\sqrt{13}}{2}$.
- Candidate 1 ($1 + \lambda = +\tfrac{\sqrt{13}}{2}$, so $\lambda \approx 0.803$): $(x,y) = \dfrac{2}{\sqrt{13}}(3,2) \approx (1.664,\,1.109)$ and $f = (\sqrt{13} - 2)^2 \approx 2.578$.
- Candidate 2 ($1 + \lambda = -\tfrac{\sqrt{13}}{2}$, so $\lambda \approx -2.803$): $(x,y) \approx (-1.664,\,-1.109)$ and $f = (\sqrt{13} + 2)^2 \approx 31.42$.
- The circle is closed and bounded, so a minimum exists and it is the candidate with the smaller $f$: candidate 1 is the nearest point (as expected, it lies on the ray toward $(3,2)$). Candidate 2 is the farthest point.
The recipe for $\min f(\mathbf{x})$ subject to $h_j(\mathbf{x}) = 0$, $j = 1,\dots,p$:
- Standard form. Write every constraint as "$h_j(\mathbf{x}) = 0$". Turn "maximise" into "minimise the negative".
- Lagrangian. $L(\mathbf{x},\boldsymbol\lambda) = f + \sum_j\lambda_jh_j$.
- Equations. Set all $n + p$ partial derivatives of $L$ to zero: $\nabla_{\mathbf{x}}L = \mathbf{0}$ and $h_j = 0$.
- Solve the system. Tips: solve each stationarity equation for $\mathbf{x}$ in terms of $\boldsymbol\lambda$, put that into the constraints, and solve for $\boldsymbol\lambda$; split into cases when you see a product equal to zero; for a quadratic $f$ with linear constraints the system is linear (one matrix solve); otherwise use Newton's method on the system.
- List all candidates, plus any feasible points where the constraint gradients are dependent (the regularity condition fails: the recipe cannot see those, so check them separately).
- Compare. Evaluate $f$ at every candidate. If the feasible set is closed and bounded (and $f$ is continuous), the minimum exists and is the candidate with the smallest $f$. If it is not bounded, add an argument: convexity, or $f \to \infty$ at infinity.
- Second-order check (awareness). Let $Z$ span the tangent directions ($J\mathbf{d} = \mathbf{0}$) and $H_L = \nabla^2 f + \sum_j\lambda_j\nabla^2h_j$. If $\mathbf{d}^\top H_L\,\mathbf{d} \gt 0$ for every nonzero tangent $\mathbf{d}$, the candidate is a strict local minimum; if $\lt 0$ for all of them, a local maximum; if mixed, a saddle. (This compares curvature along the constraint, including the curvature of the constraint itself, which is why the term $\lambda_j\nabla^2h_j$ appears.)
Not every candidate is a minimum. Closest points of the parabola $y = x^2 - 1$ to the origin: $f = x^2 + y^2$, $h = y - x^2 + 1$. Then $\nabla f = (2x, 2y)$, $\nabla h = (-2x, 1)$.
- $2x - 2\lambda x = 0$ gives $x = 0$ or $\lambda = 1$. Also $2y + \lambda = 0$ and $y = x^2 - 1$.
- Case $x = 0$: $y = -1$, $\lambda = 2$, $f = 1$.
- Case $\lambda = 1$: $y = -\tfrac12$, then $x^2 = y + 1 = \tfrac12$, so $x = \pm\tfrac{1}{\sqrt2} \approx \pm0.707$ and $f = \tfrac12 + \tfrac14 = \tfrac34$.
- Three candidates: $(0,-1)$ with $f = 1$ and $(\pm0.707,\,-0.5)$ with $f = 0.75$. The winners are the two with $0.75$.
- Second-order check at $(0,-1)$ ($\lambda = 2$): $H_L = \nabla^2 f + \lambda\nabla^2h = \begin{bmatrix}2 & 0\\0 & 2\end{bmatrix} + 2\begin{bmatrix}-2 & 0\\0 & 0\end{bmatrix} = \begin{bmatrix}-2 & 0\\0 & 2\end{bmatrix}$. The tangent at $(0,-1)$ is $\mathbf{d} = (1,0)$ (since $\nabla h = (0,1)$), and $\mathbf{d}^\top H_L\mathbf{d} = -2 \lt 0$: a local maximum along the curve. Indeed, along the parabola $f = x^4 - x^2 + 1$, which has a local maximum at $x = 0$.
- At $(0.707,-0.5)$ ($\lambda = 1$): $H_L = \begin{bmatrix}0 & 0\\0 & 2\end{bmatrix}$, tangent $\mathbf{d} \propto (1, \sqrt2)$ (perpendicular to $\nabla h = (-\sqrt2, 1)$), and $\mathbf{d}^\top H_L\mathbf{d} = 2\cdot2 = 4 \gt 0$: a local minimum ✓.
Why do we need it?
Knowing the theorem is not the same as solving a problem. A fixed routine tells you what to write, what to solve, and, importantly, what to check so that you do not report a maximum as a minimum.
Where is it used?
Hand derivations in statistics and ML (maximum likelihood with a constraint, PCA, max-entropy), equality-constrained quadratic programs inside SQP and interior-point solvers (they solve the linear KKT system of step 4 at every iteration), and sanity checks of numerical results.
How is it used?
By hand for small problems; with a numerical solver (Newton on the Lagrange system, or SciPy's SLSQP) for larger ones. In both cases finish with steps 6 and 7: compare values and check the second-order condition.
Never skip the comparison. The recipe returns stationary points. Two of the three candidates above are minima; the third is a maximum along the curve.
Lost solutions. When you divide an equation by a variable (for example by $x$ in $2x - 2\lambda x = 0$), you may lose the case $x = 0$. Always split into cases instead of dividing.
Points the recipe cannot see. If the constraint gradients are dependent at some feasible point (a cusp, a crossing), that point can be a minimum without satisfying the multiplier equations. Check such points separately, as in the pathological example of Chapter 3.7.
Quick check: minimise $x^2 + y^2$ on the hyperbola $xy = 1$. Find the candidates.
$h = xy - 1$, $\nabla h = (y, x)$. $2x + \lambda y = 0$ and $2y + \lambda x = 0$. Multiply the first by $x$ and the second by $y$: $2x^2 + \lambda xy = 0$ and $2y^2 + \lambda xy = 0$, so $x^2 = y^2$, i.e. $y = \pm x$. With $xy = 1$ we need $y = x$, so $x = y = \pm1$. Candidates $(1,1)$ and $(-1,-1)$, both with $f = 2$ and $\lambda = -2$. Both are minima (the hyperbola goes off to infinity, where $f \to \infty$).
Worked example: maximum entropy on the simplex (and why softmax appears) core
Suppose you must assign probabilities to four possible outcomes, and you know nothing except that they add up to 1. The least "opinionated" (most honest) choice is to treat them equally: 25% each. Entropy is the number that measures "how spread out" a distribution is, and the uniform distribution has the largest entropy.
Now suppose you also know the average outcome. You are no longer free to use the uniform distribution, but you still want the least opinionated distribution that matches the average. The Lagrange method gives a famous answer: probabilities proportional to $e^{-\lambda x_i}$. That is the softmax.
Part 1: only $\sum p_i = 1$. Maximise the entropy $H(\mathbf{p}) = -\sum_{i=1}^{n} p_i\ln p_i$ subject to $\sum p_i = 1$. (The condition $p_i \gt 0$ is automatically satisfied by the answer.) We minimise $f = \sum p_i\ln p_i$ ($= -H$).
- $L = \sum_i p_i\ln p_i + \lambda\bigl(\sum_i p_i - 1\bigr)$.
- $\dfrac{\partial L}{\partial p_i} = \ln p_i + 1 + \lambda = 0$ for every $i$. (Derivative of $p\ln p$ is $\ln p + 1$.)
- So $\ln p_i = -1 - \lambda$ is the same number for all $i$: all $p_i$ are equal.
- $\sum p_i = 1$ forces $p_i = \tfrac1n$. The multiplier is $\lambda = \ln n - 1$ (from $e^{-1-\lambda} = \tfrac1n$).
- Value: $H = -n\cdot\tfrac1n\ln\tfrac1n = \ln n$. For $n = 4$: $H = \ln 4 \approx 1.386$. Compare $\mathbf{p} = (0.4, 0.3, 0.2, 0.1)$, which has $H \approx 1.280$: smaller ✓.
- Since $p\ln p$ is convex (its second derivative is $1/p \gt 0$) and the constraint is linear, this stationary point is the global minimum of $f$, i.e. the global maximum of entropy.
Part 2: also fix the mean. Outcomes have values $x_1,\dots,x_n$ and we require $\sum_ix_ip_i = m$. Use two multipliers: $\lambda_1$ for $\sum p_i = 1$ and $\lambda_2$ for the mean.
- $L = \sum p_i\ln p_i + \lambda_1\bigl(\sum p_i - 1\bigr) + \lambda_2\bigl(\sum x_ip_i - m\bigr)$.
- $\ln p_i + 1 + \lambda_1 + \lambda_2x_i = 0$, so $p_i = e^{-1-\lambda_1}\,e^{-\lambda_2x_i}$.
- The first factor is the same for all $i$, so it is just a normaliser: $\displaystyle p_i = \frac{e^{-\lambda_2x_i}}{\sum_je^{-\lambda_2x_j}}$. This is exactly the softmax of the scores $z_i = -\lambda_2x_i$.
- The remaining multiplier $\lambda_2$ is set so that the mean equals $m$. For $x = (1,2,3,4)$ and $m = 2$: $\lambda_2 \approx 0.4196$, giving $\mathbf{p} \approx (0.4214,\,0.2770,\,0.1820,\,0.1197)$ and $H \approx 1.284 \lt 1.386$.
Maximum-entropy principle. Among all distributions on $x_1,\dots,x_n$ with $\sum_ip_i = 1$ and prescribed averages $\sum_ip_i\,\phi_k(x_i) = m_k$, the one with the largest entropy has the form
$$p_i = \frac{1}{Z}\exp\Bigl(-\sum_k\lambda_{k}\,\phi_k(x_i)\Bigr), \qquad Z = \sum_i\exp\Bigl(-\sum_k\lambda_k\phi_k(x_i)\Bigr),$$where the $\lambda_k$ are the multipliers of the average-constraints (the Lagrange multiplier of the normalisation is absorbed in $Z$). With no average-constraint, it is the uniform distribution. With one, it is a softmax: in physics $\lambda_2$ is the inverse temperature ($1/T$), and the family is called the Gibbs or exponential-family distribution.
Why do we need it?
It gives a principled answer to "what distribution should I assume, given only what I know?" and explains why the exponential form (softmax, Boltzmann, logistic) appears everywhere, instead of it being an arbitrary choice.
Where is it used?
The softmax output layer and temperature in language models, logistic regression and multinomial models (they are max-entropy models), the Boltzmann distribution in physics, entropy regularisation in reinforcement learning, and the Gibbs distribution in energy-based models.
How is it used?
Write the entropy (negated) plus one multiplier per known average. The stationarity equations give an exponential form. Then adjust the multipliers (by bisection or Newton) until the averages match the data.
The mean must be achievable. With values $1,\dots,4$, a mean outside $(1, 4)$ is impossible (and exactly $1$ or $4$ would need a degenerate distribution with $\lambda_2 = \pm\infty$).
Maximum entropy is a modelling choice. It is the least committed distribution consistent with the constraints, not necessarily the true one.
Quick check: what is the maximum-entropy distribution on 5 outcomes with only $\sum p_i = 1$, and its entropy?
Uniform: $p_i = 0.2$ each, with $H = \ln 5 \approx 1.609$. The multiplier is $\lambda = \ln 5 - 1 \approx 0.609$.
Worked example: the best unit direction is an eigenvector (PCA) core
Principal component analysis (PCA) looks for the direction along which data varies the most. If the data's covariance matrix is $A$, the variance along a direction $\mathbf{x}$ is $\mathbf{x}^\top A\mathbf{x}$. But there is a catch: if $\mathbf{x}$ may be any length, you can make this as large as you like by just stretching $\mathbf{x}$. A direction has no length, so we fix the length to 1: the constraint $\mathbf{x}^\top\mathbf{x} = 1$.
This is an equality-constrained problem on a circle (or sphere). Apply the multiplier rule and a famous fact falls out: the best direction is an eigenvector of $A$, and the multiplier is its eigenvalue.
Let $A = \begin{bmatrix}3 & 1\\1 & 2\end{bmatrix}$. Maximise $\mathbf{x}^\top A\mathbf{x} = 3x^2 + 2xy + 2y^2$ subject to $h = x^2 + y^2 - 1 = 0$. As a minimisation: $f = -\mathbf{x}^\top A\mathbf{x}$.
- Gradients ($\nabla(\mathbf{x}^\top A\mathbf{x}) = 2A\mathbf{x}$ for symmetric $A$): $\nabla f = -2A\mathbf{x}$ and $\nabla h = 2\mathbf{x}$.
- Rule $\nabla f + \lambda\nabla h = \mathbf{0}$: $-2A\mathbf{x} + 2\lambda\mathbf{x} = \mathbf{0}$, so $A\mathbf{x} = \lambda\mathbf{x}$. That is the eigenvalue equation.
- The eigenvalues solve $\det(A - \lambda I) = (3-\lambda)(2-\lambda) - 1 = \lambda^2 - 5\lambda + 5 = 0$, so $\lambda = \dfrac{5 \pm \sqrt5}{2} \approx 3.618,\; 1.382$.
- The value at a candidate: $\mathbf{x}^\top A\mathbf{x} = \mathbf{x}^\top(\lambda\mathbf{x}) = \lambda\|\mathbf{x}\|^2 = \lambda$. So the largest eigenvalue $3.618$ is the maximum and the smallest $1.382$ is the minimum.
- The eigenvector for $\lambda = 3.618$: $(3 - 3.618)x + y = 0$ gives $y = 0.618x$; normalised, $\mathbf{x}^\star = (0.8507,\,0.5257)$ (angle $31.7^\circ$). Check: $A\mathbf{x}^\star = (3.0777,\,1.9021) = 3.618\,\mathbf{x}^\star$ ✓.
Theorem (Rayleigh). For a symmetric matrix $A$, the stationary points of $\mathbf{x}^\top A\mathbf{x}$ on the unit sphere $\|\mathbf{x}\| = 1$ are exactly the unit eigenvectors, with Lagrange multiplier (in the form "minimise $-\mathbf{x}^\top A\mathbf{x}$") equal to the eigenvalue. Hence
$$\max_{\|\mathbf{x}\|=1}\mathbf{x}^\top A\mathbf{x} = \lambda_{\max}(A), \qquad \min_{\|\mathbf{x}\|=1}\mathbf{x}^\top A\mathbf{x} = \lambda_{\min}(A).$$The geometric reading is the same picture as before: the contours of $\mathbf{x}^\top A\mathbf{x}$ are ellipses; the circle $\|\mathbf{x}\| = 1$ touches an ellipse where their normals ($A\mathbf{x}$ and $\mathbf{x}$) are parallel. These are the tips of the ellipse axes (the eigenvectors; PCA in the Linear Algebra guide).
Why do we need it?
It shows that eigenvectors are not a separate trick: they are the answer to a constrained optimization problem. It also shows what the constraint "length 1" does: without it the problem has no maximum.
Where is it used?
Principal component analysis (first principal component = top eigenvector of the covariance), spectral clustering, the power method, normalised weight vectors in "weight-norm" layers, and the largest singular value (the spectral norm) of a weight matrix, which spectral normalisation constrains.
How is it used?
Recognise "maximise a quadratic form over unit vectors" and answer immediately: the top eigenvector, with value $\lambda_{\max}$. For several components, add orthogonality constraints (more multipliers) and you get the next eigenvectors.
Symmetric matrices only. The formula $\nabla(\mathbf{x}^\top A\mathbf{x}) = 2A\mathbf{x}$ and the theorem need $A = A^\top$ (covariance matrices are). For a general $A$ the gradient is $(A + A^\top)\mathbf{x}$.
Two signs, two roles. The multiplier equals the eigenvalue only because of how we wrote the Lagrangian ($-\mathbf{x}^\top A\mathbf{x} + \lambda(\mathbf{x}^\top\mathbf{x} - 1)$). With $+\mathbf{x}^\top A\mathbf{x} + \lambda(\cdots)$ (a minimisation) you would get $\lambda = -$eigenvalue. The eigenvectors themselves do not depend on this.
Quick check: what is $\max \mathbf{x}^\top A\mathbf{x}$ over unit vectors for $A = \begin{bmatrix}5 & 0\\0 & 2\end{bmatrix}$, and where is it attained?
The eigenvalues are $5$ and $2$, so the maximum is $5$, attained at $\mathbf{x} = (\pm1, 0)$ (check: $5x^2 + 2y^2$ on $x^2 + y^2 = 1$ is $2 + 3x^2$, largest at $x^2 = 1$). The minimum is $2$ at $(0,\pm1)$.
The multiplier as a sensitivity: what is a rule worth? core
Here is the most useful meaning of $\lambda$. Suppose a rule says "spend exactly $c$ dollars". If the rule is relaxed a little (to $c + 1$ cent), the best achievable result improves or worsens by a small amount. The multiplier is the exchange rate: how much the best value changes per unit of change of the rule. It is the price of the constraint (a "shadow price" in economics).
A big multiplier says the rule is hurting a lot ("loosening it would help greatly"). A multiplier near zero says the rule is almost irrelevant.
Minimise $f = x^2 + y^2$ subject to $x + y = c$. Rewrite the constraint as $h(\mathbf{x}) - c = 0$ with $h = x + y$.
- Lagrange: $2x + \lambda = 0$, $2y + \lambda = 0$, so $x = y = -\lambda/2$; the constraint $x + y = c$ gives $\lambda = -c$ and $x = y = c/2$.
- The best value as a function of the rule level: $p^\star(c) = 2\cdot(c/2)^2 = c^2/2$.
- Its rate of change: $\dfrac{dp^\star}{dc} = c$. At $c = 2$: $\lambda = -2$ and $\dfrac{dp^\star}{dc} = 2 = -\lambda$ ✓.
- Numerical check by perturbing $c$: $p^\star(2) = 2$. Move $c$ to $2.01$: $p^\star(2.01) = 2.01^2/2 = 2.02005$. The change is $0.02005$, and the prediction $-\lambda\,\Delta c = 2\times0.01 = 0.02$ ✓ (the small difference is second-order).
A maximisation example. The fence problem: minimise $-xy$ s.t. $x + y = c$: $x = y = c/2 = \lambda$, so $\lambda = c/2$ and $p^\star = -c^2/4$, $dp^\star/dc = -c/2 = -\lambda$ ✓. In "area" terms: the best area is $c^2/4$ and it grows at rate $c/2 = \lambda$ per extra unit of $c$. At $c = 10$, the best area $25$ becomes $25.5025$ at $c = 10.1$: change $0.5025 \approx 5\times0.1$ ✓.
Sensitivity theorem (equality constraint). Consider the family of problems $p^\star(c) = \min f(\mathbf{x})$ subject to $h(\mathbf{x}) = c$, with Lagrangian $L = f + \lambda(h - c)$. If $\mathbf{x}^\star(c)$ is a differentiable family of regular optimal solutions with multiplier $\lambda^\star(c)$, then
$$\boxed{\;\frac{dp^\star}{dc} = -\lambda^\star\;}\qquad\text{(and for several constraints, } \tfrac{\partial p^\star}{\partial c_j} = -\lambda_j^\star\text{).}$$Derivation. $p^\star(c) = f(\mathbf{x}^\star(c))$, so by the chain rule $p^{\star\prime}(c) = \nabla f\cdot\dfrac{d\mathbf{x}^\star}{dc}$. The constraint $h(\mathbf{x}^\star(c)) = c$ differentiates to $\nabla h\cdot\dfrac{d\mathbf{x}^\star}{dc} = 1$. Stationarity says $\nabla f = -\lambda\nabla h$. Substituting, $p^{\star\prime}(c) = -\lambda\,\nabla h\cdot\dfrac{d\mathbf{x}^\star}{dc} = -\lambda$. $\blacksquare$
Sign convention warning. The sign depends on writing the Lagrangian as $f + \lambda(h - c)$, as we always do. With the opposite sign convention, the formula reads $+\lambda$. Always check. Units: $\lambda$ is measured in (units of $f$) per (unit of $h$).
Why do we need it?
It turns a by-product of the calculation into an answer to a practical question: "which rule is costing me the most, and how much would I gain by relaxing it?" It also gives a free check of your multiplier.
Where is it used?
Economics (shadow prices of resources, linear programming sensitivity analysis), regularisation (the multiplier of a norm-ball constraint is the penalty strength, so it tells how much error a tighter ball costs), and SVMs (multipliers show which training points matter).
How is it used?
Solve once, read off $\lambda$. Predict the effect of a small change $\Delta c$ on the best value as $-\lambda\,\Delta c$. If you need to be sure, re-solve at $c + \Delta c$ and compare, as in the widget.
It is a local, first-order statement. $\lambda$ predicts the effect of small changes. For big changes, $p^\star(c)$ curves (in the example above, $p^\star = c^2/2$ and the slope itself changes with $c$).
Mind the signs. With the Lagrangian $f + \lambda(h - c)$ the rate is $-\lambda$. If you write $f - \lambda h$ you get $+\lambda$. And whether "better" means "smaller" depends on whether you are minimising $f$ or maximising $-f$.
Quick check: for $\min x^2 + y^2$ s.t. $x + y = c$ with $c = 4$, what is $\lambda$, and how much does $p^\star$ change if $c$ goes to $4.1$?
$\lambda = -c = -4$, so $dp^\star/dc = 4$. The first-order change is $4\times0.1 = 0.4$. (Exactly: $p^\star(4.1) - p^\star(4) = 8.405 - 8 = 0.405$.)
Recap, cheat sheet and practice
- Geometry: at a constrained optimum the contour of $f$ is tangent to the constraint, so $\nabla f \parallel \nabla h$. Moving along the constraint cannot improve $f$ to first order.
- Rule: $\nabla f + \sum_j\lambda_j\nabla h_j = \mathbf{0}$ and $h_j = 0$: $n + p$ equations, $n + p$ unknowns. It needs regularity: the $\nabla h_j$ must be linearly independent.
- Lagrangian $L = f + \sum\lambda_jh_j$: its stationary points in $(\mathbf{x},\boldsymbol\lambda)$ are exactly the candidates. $L$ is a saddle there (min in $\mathbf{x}$ along the constraint, linear in $\boldsymbol\lambda$), never minimised in $\boldsymbol\lambda$.
- Stationarity means the reduced gradient (the part of $\nabla f$ along the tangent) is zero. $\nabla f$ itself is not zero.
- Recipe: write $L$, set all partials to 0, solve, list every candidate, compare $f$-values (closed and bounded set: smallest wins), check second-order (curvature of $L$ on the tangent space).
- Classics: nearest point on a line, $\mathbf{x}^\star = \tfrac{c}{a^2+b^2}(a,b)$; minimum norm $A^\top(AA^\top)^{-1}\mathbf{b}$; max entropy gives uniform (sum only) or softmax (with a mean constraint).
- Sensitivity: $dp^\star/dc = -\lambda^\star$ for the constraint $h = c$ and $L = f + \lambda(h - c)$. The multiplier is the price of the rule.
Cheat sheet
| Idea | Formula | Remember |
|---|---|---|
| Lagrange condition | $\nabla f + \sum\lambda_j\nabla h_j = \mathbf{0}$, $h_j = 0$ | gradients parallel; independent $\nabla h_j$ |
| Lagrangian | $L = f + \sum\lambda_jh_j$ | $\partial L/\partial\lambda_j = h_j$; saddle, not a bowl |
| Reduced gradient | $\mathbf{r} = \nabla f + J^\top\boldsymbol\lambda$, $\boldsymbol\lambda = -(JJ^\top)^{-1}J\nabla f$ | zero at stationary points |
| Min norm on $A\mathbf{x} = \mathbf{b}$ | $A^\top(AA^\top)^{-1}\mathbf{b}$, $\boldsymbol\lambda = -2(AA^\top)^{-1}\mathbf{b}$ | pseudoinverse |
| Equality-constrained quadratic | $\begin{bmatrix}Q & A^\top\\A & 0\end{bmatrix}\begin{bmatrix}\mathbf{x}\\\boldsymbol\lambda\end{bmatrix} = \begin{bmatrix}-\mathbf{c}\\\mathbf{b}\end{bmatrix}$ | one linear solve (for $\tfrac12\mathbf{x}^\top Q\mathbf{x} + \mathbf{c}^\top\mathbf{x}$) |
| Second-order test | $\mathbf{d}^\top(\nabla^2f + \sum\lambda_j\nabla^2h_j)\mathbf{d}\gt0$ for tangent $\mathbf{d}\ne0$ | strict local minimum |
| Max entropy | $p_i\propto e^{-\lambda_2x_i}$ (softmax) | uniform without a mean constraint |
| Sensitivity | $dp^\star/dc = -\lambda^\star$ | price of the rule |
import numpy as np
from scipy.optimize import minimize
# 1. Biggest rectangle: maximise x*y on x + y = 10 (we minimise -x*y)
res = minimize(lambda v: -v[0] * v[1], x0=[1.0, 1.0], method="SLSQP",
constraints=[{"type": "eq", "fun": lambda v: v[0] + v[1] - 10}])
print(np.round(res.x, 4), round(-res.fun, 4)) # [5. 5.] 25.0
grad_f = np.array([-res.x[1], -res.x[0]]) # gradient of f = -xy
grad_h = np.array([1.0, 1.0])
lam = -(grad_f @ grad_h) / (grad_h @ grad_h) # best lambda in grad f + lambda grad h = 0
print(round(lam, 4)) # 5.0
# 2. Equality-constrained quadratic: solve the KKT linear system in one go
# minimise x^2 + y^2 + z^2 s.t. x + y + z = 6 and x + 2y + 3z = 10
A = np.array([[1.0, 1, 1], [1, 2, 3]])
b = np.array([6.0, 10.0])
n, p = 3, 2
K = np.block([[2 * np.eye(n), A.T], [A, np.zeros((p, p))]]) # [[2I, A^T], [A, 0]]
sol = np.linalg.solve(K, np.concatenate([np.zeros(n), b]))
print(np.round(sol, 4)) # [ 3. 2. 1. -8. 2.] = x, y, z, lambda1, lambda2
print(np.round(A.T @ np.linalg.solve(A @ A.T, b), 4)) # [3. 2. 1.] the minimum-norm formula agrees
# 3. Closest and farthest point on the circle x^2 + y^2 = 4 to (3, 2):
# (x, y) = (3, 2) / (1 + lambda) and 13 / (1 + lambda)^2 = 4 -> 1 + lambda = +-sqrt(13)/2
for sign in (+1, -1):
lam = sign * np.sqrt(13) / 2 - 1
pt = np.array([3.0, 2.0]) / (1 + lam)
print(round(lam, 4), np.round(pt, 4), round(((pt - [3, 2]) ** 2).sum(), 4))
# 0.8028 [1.6641 1.1094] 2.5778 <- closest (the minimum)
# -2.8028 [-1.6641 -1.1094] 31.4222 <- farthest (the maximum)
# 4. Sensitivity: for min x^2 + y^2 s.t. x + y = c the best value is c^2/2 and lambda = -c.
pstar = lambda c: c ** 2 / 2
c, d = 2.0, 1e-4
print(round((pstar(c + d) - pstar(c - d)) / (2 * d), 4), "vs -lambda =", 2.0) # 2.0 vs -lambda = 2.0
print(pstar(2.01) - pstar(2.0)) # about 0.02005 (first-order prediction: -lambda * 0.01 = 0.02)
# 5. Maximum entropy on {1,2,3,4} with mean 2: p_i proportional to exp(-lambda2 * x_i)
x = np.array([1.0, 2, 3, 4])
mean = lambda l2: (x * np.exp(-l2 * x)).sum() / np.exp(-l2 * x).sum()
lo, hi = -20.0, 20.0
for _ in range(80): # bisection: the mean decreases as lambda2 grows
mid = (lo + hi) / 2
lo, hi = (mid, hi) if mean(mid) > 2.0 else (lo, mid)
l2 = (lo + hi) / 2
pr = np.exp(-l2 * x) / np.exp(-l2 * x).sum()
print(round(l2, 4), np.round(pr, 4), round(-(pr * np.log(pr)).sum(), 4)) # 0.4196 [0.4214 0.277 0.182 0.1197] 1.2839
print(round(np.log(4), 4)) # 1.3863 = entropy of the uniform distribution (no mean constraint)
1. At a constrained local minimum on the curve $h(\mathbf{x}) = 0$ (with $\nabla h \ne \mathbf{0}$), which statement is true?
2. Minimise $x^2 + y^2$ subject to $x + y = 4$. With $L = x^2 + y^2 + \lambda(x + y - 4)$, the multiplier is…
3. Why is the Lagrangian $L(\mathbf{x},\lambda) = f + \lambda h$ typically a saddle at the solution, not a minimum?
4. On the circle $x^2 + y^2 = 2$, the Lagrange equations for $f = x + y$ give the two candidates $(1,1)$ and $(-1,-1)$. The minimum of $f$ is…
5. For $\min x^2 + y^2$ subject to $x + y = c$ we found $\lambda = -c$. At $c = 2$, if the rule is changed to $c = 2.1$, the best value changes by about…
6. Maximising the entropy $-\sum p_i\ln p_i$ subject only to $\sum p_i = 1$ over $n$ outcomes gives…
Practice problems
A. Maximise $x + 2y$ on the circle $x^2 + y^2 = 5$. Give the maximiser, the maximum and $\lambda$.
Minimise $f = -(x + 2y)$ with $h = x^2 + y^2 - 5$. $\nabla f = (-1,-2)$, $\nabla h = (2x,2y)$. So $-1 + 2\lambda x = 0$ and $-2 + 2\lambda y = 0$, giving $x = \tfrac{1}{2\lambda}$, $y = \tfrac1\lambda$. Constraint: $\tfrac{1}{4\lambda^2} + \tfrac{1}{\lambda^2} = \tfrac{5}{4\lambda^2} = 5$, so $\lambda = \pm\tfrac12$. For $\lambda = \tfrac12$: $(x,y) = (1,2)$ and $x + 2y = 5$ (maximum). For $\lambda = -\tfrac12$: $(-1,-2)$ and $x + 2y = -5$ (minimum). So the maximum is $5$ at $(1,2)$ with $\lambda = \tfrac12$. (Cauchy–Schwarz agrees: $|x + 2y| \le \sqrt{5}\sqrt{x^2+y^2} = 5$.)
B. Find the point of the plane $x + 2y + 2z = 9$ nearest to the origin, and its distance.
Use the formula with $\mathbf{a} = (1,2,2)$, $c = 9$, $\|\mathbf{a}\|^2 = 9$: $\mathbf{x}^\star = \tfrac{9}{9}(1,2,2) = (1,2,2)$, $\lambda = -2c/\|\mathbf{a}\|^2 = -2$, minimum $\|\mathbf{x}\|^2 = c^2/\|\mathbf{a}\|^2 = 81/9 = 9$, distance $3$. Check: $1 + 4 + 4 = 9$ ✓ and $1 + 4 + 4 = 9$ for the constraint ✓.
C. Find the points of the hyperbola $xy = 1$ closest to the origin.
$h = xy - 1$, $\nabla h = (y, x)$. $2x + \lambda y = 0$ and $2y + \lambda x = 0$. Multiplying by $x$ and $y$ respectively: $2x^2 = -\lambda xy = 2y^2$, so $y = \pm x$. With $xy = 1$, $y = x$ and $x = \pm1$. Candidates $(1,1)$ and $(-1,-1)$ with $f = 2$ and $\lambda = -2$ (from $2 + \lambda\cdot1 = 0$). The hyperbola is unbounded and $f \to \infty$ along it, so both are minima; the distance is $\sqrt2$.
D. For the fence problem, use $\lambda$ to estimate how much the best area grows if the half-perimeter goes from $x + y = 10$ to $10.5$. Compare with the exact value.
For "minimise $-xy$" at $c = 10$, $\lambda = c/2 = 5$ and $dp^\star/dc = -5$, i.e. the best area grows at $5$ per unit of $c$. First-order estimate: $5\times0.5 = 2.5$. Exact: best area is $c^2/4$, so $10.5^2/4 - 10^2/4 = 27.5625 - 25 = 2.5625$. The estimate is close; the small gap is the second-order term.
E. Maximum entropy on the values $x = (0, 1, 2)$ with mean $1$. What is $\lambda_2$ and what is the distribution?
The solution has the form $p_i \propto e^{-\lambda_2x_i}$. The values are symmetric around the required mean $1$, and the uniform distribution $(\tfrac13,\tfrac13,\tfrac13)$ already has mean $\tfrac{0 + 1 + 2}{3} = 1$. So the maximum-entropy distribution is uniform and $\lambda_2 = 0$, with entropy $\ln 3 \approx 1.099$.
F. Use the second-order test on the two candidates of "minimise $x + y$ on $x^2 + y^2 = 2$".
$f = x + y$ has $\nabla^2f = 0$, and $h = x^2 + y^2 - 2$ has $\nabla^2h = 2I$. So $H_L = \lambda\cdot2I$. At $(1,1)$: $\nabla f = (1,1) = -\lambda(2,2)$ so $\lambda = -\tfrac12$ and $H_L = -I$. The tangent is $\mathbf{d} = (1,-1)$ and $\mathbf{d}^\top H_L\mathbf{d} = -2 \lt 0$: a local maximum. At $(-1,-1)$: $\lambda = +\tfrac12$, $H_L = +I$, $\mathbf{d}^\top H_L\mathbf{d} = +2 \gt 0$: a local minimum ✓ (it matches the values $f = \pm2$).
KKT Conditions
Fences you must stay inside change what "best" means. The Karush–Kuhn–Tucker (KKT) conditions are a four-line checklist that a best point has to pass. You will learn what each line means, why the signs are what they are, how to solve small problems with it, and how it explains the support vectors of an SVM.
- Tell an active constraint (touching you) from an inactive one, and say why only active ones can push
- State the four KKT conditions: stationarity, primal feasibility, dual feasibility, complementary slackness
- Derive the sign rule $\mu_i \ge 0$ from a picture
- Check a candidate point with ✓ / ✗, and solve small problems by case analysis over active sets
- Know exactly when KKT is necessary (a constraint qualification) and when it is sufficient (convex problems)
- See the geometry: $-\nabla f$ lies in the cone spanned by the active constraint gradients
- Meet the SVM: support vectors are exactly the constraints that complementary slackness leaves active
Notation, fixed in Chapter 3.7. We minimise $f(\mathbf{x})$ subject to inequality constraints $g_i(\mathbf{x}) \le 0$ ($i = 1,\dots,m$) and equality constraints $h_j(\mathbf{x}) = 0$ ($j = 1,\dots,p$). Each equality gets a multiplier $\lambda_j$ (any sign); each inequality gets a multiplier $\mu_i \ge 0$. The Lagrangian is $$L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu) = f(\mathbf{x}) + \sum_j \lambda_j h_j(\mathbf{x}) + \sum_i \mu_i g_i(\mathbf{x}).$$ Chapter 3.8 explained the equality-only case in depth. Here we add the inequalities, which is where the interesting new rules appear. The gradient $\nabla f$ is a column vector, as in the Calculus guide.
Active and inactive constraints core
Picture a ball rolling on a field that is surrounded by a fence. The ground has a lowest spot, where the ball would like to rest. Two things can happen.
- The lowest spot is inside the field. The ball rolls there and stops. It never even touches the fence. You could take the fence away and nothing would change. We call that fence inactive.
- The lowest spot is outside the field. The ball rolls towards it until it hits the fence, and stays there, pressed against it. The fence is stopping it. We call that fence active.
An active fence has to push back. The more strongly the ground pulls the ball outwards, the harder the fence must push. That push is what the multiplier $\mu$ measures.
Minimise $(x - c)^2$ with the two walls $x \le 2$ and $x \ge 1$, so the ball must stay between 1 and 2. The number $c$ is the "wish": the spot where $(x-c)^2$ is smallest.
- $c = 1.5$. The wish is inside. The best point is $x^\star = 1.5$, which is $0.5$ away from each wall. Both walls are inactive.
- $c = 3$. The wish is beyond the right wall. The best allowed point is the wall itself, $x^\star = 2$. The right wall is active, the left one ($1$ away) is inactive.
- $c = 0$. Now the left wall is the active one, with $x^\star = 1$.
Which wall is active depends on the problem, and nobody tells you in advance. Finding out is part of the work.
Write each inequality as $g_i(\mathbf{x}) \le 0$. At a point $\mathbf{x}$ the constraint is
- active if $g_i(\mathbf{x}) = 0$ (you are touching the wall),
- inactive if $g_i(\mathbf{x}) \lt 0$ (strictly inside, with room to spare),
- violated if $g_i(\mathbf{x}) \gt 0$ (outside: not allowed).
The active set is $\mathcal{A}(\mathbf{x}) = \{\, i : g_i(\mathbf{x}) = 0 \,\}$. Equality constraints $h_j = 0$ are active at every feasible point, because they must hold exactly. The number $s_i = -g_i(\mathbf{x}) \ge 0$ is called the slack of constraint $i$: how much room is left before the wall.
Why do we need it?
Real problems have many constraints (an SVM has one per training example), but at any moment only a few of them touch you. Knowing which ones matter turns a huge problem into a small one.
Where is it used?
Support vectors in an SVM (the active margin constraints), the corners of linear programs, box-constrained solvers such as L-BFGS-B that clamp weights at their bounds, and active-set methods for quadratic programs.
How is it used?
Evaluate every $g_i$ at your point. The ones equal to zero are active: treat them like equalities (the tools of Chapter 3.8) and forget the others, which are not holding you back.
Active does not always mean it matters. Minimise $x^2$ subject to $x \ge 0$. The answer $x^\star = 0$ sits exactly on the wall, so the wall is active. But the ground is flat there ($f'(0)=0$), so the wall pushes with strength $\mu = 0$. We say it is weakly active. Such touching-but-idle walls are legal, just a little awkward.
Active is a local word. It describes one point. A wall can be active at one candidate and inactive at the next.
Quick check: minimise $(x-5)^2$ with $1 \le x \le 2$. Which wall is active, and which has multiplier zero?
The wish is at 5, far to the right, so the ball stops at the right wall: $x^\star = 2$. The wall $x \le 2$ is active. The wall $x \ge 1$ is $1$ away, so it is inactive and its multiplier is $0$.
The four KKT conditions, as a set core
Imagine someone hands you a point and says, "this is the best allowed point". You cannot try every other point. The KKT conditions give you four quick questions to ask about that one point, together with the "push strengths" $\mu$ of the walls:
- Are you even allowed here? (primal feasibility)
- Is the pull of the objective cancelled by the push of the walls? (stationarity)
- Do the walls only push, never pull? (dual feasibility)
- Do only the walls that touch you push? (complementary slackness)
If every answer is yes, the point is a KKT point: a serious candidate for the best one.
Take $\min\ (x-2)^2 + (y-1)^2$ subject to $x + y \le 2$, $x \ge 0$, $y \ge 0$. In standard form the walls are $g_1 = x + y - 2 \le 0$, $g_2 = -x \le 0$, $g_3 = -y \le 0$. Test the candidate $\mathbf{x}^\star = (1.5,\ 0.5)$ with multipliers $\boldsymbol\mu = (1,\ 0,\ 0)$.
- Feasible? $g_1 = 1.5 + 0.5 - 2 = 0 \le 0$ ✓, $g_2 = -1.5 \le 0$ ✓, $g_3 = -0.5 \le 0$ ✓.
- Stationary? $\nabla f = (2(1.5-2),\ 2(0.5-1)) = (-1,\ -1)$ and $\nabla g_1 = (1, 1)$, so $\nabla f + 1\cdot\nabla g_1 = (-1,-1) + (1,1) = (0, 0)$ ✓. (The other two walls have $\mu = 0$, so they add nothing.)
- Walls push, never pull? $\mu = (1, 0, 0)$, all $\ge 0$ ✓.
- Only touching walls push? $\mu_1 g_1 = 1\cdot 0 = 0$, $\mu_2 g_2 = 0\cdot(-1.5) = 0$, $\mu_3 g_3 = 0\cdot(-0.5) = 0$ ✓.
All four pass, so $(1.5, 0.5)$ is a KKT point. (Soon you will see that for this kind of problem a KKT point is the global answer.)
Now try the candidate $(1,1)$. It is feasible ($g_1 = 0$), but $\nabla f = (-2, 0)$, and no $\mu_1$ makes $(-2, 0) + \mu_1 (1,1)$ vanish: the first entry needs $\mu_1 = 2$ and the second needs $\mu_1 = 0$. Stationarity fails ✗, so $(1,1)$ is not the answer.
Consider $\min f(\mathbf{x})$ subject to $g_i(\mathbf{x}) \le 0$ and $h_j(\mathbf{x}) = 0$. A point $\mathbf{x}^\star$ with multipliers $(\boldsymbol\lambda, \boldsymbol\mu)$ satisfies the Karush–Kuhn–Tucker (KKT) conditions if all of these hold:
| Name | Condition | In plain words |
|---|---|---|
| 1. Stationarity | $\nabla f(\mathbf{x}^\star) + \sum_j \lambda_j \nabla h_j(\mathbf{x}^\star) + \sum_i \mu_i \nabla g_i(\mathbf{x}^\star) = \mathbf{0}$ | the objective's pull is balanced by the constraints' push |
| 2. Primal feasibility | $g_i(\mathbf{x}^\star) \le 0,\quad h_j(\mathbf{x}^\star) = 0$ | the point obeys every rule |
| 3. Dual feasibility | $\mu_i \ge 0$ for every inequality | a fence can push you in, never pull you out |
| 4. Complementary slackness | $\mu_i\, g_i(\mathbf{x}^\star) = 0$ for every $i$ | a fence pushes only if it is touching you |
Condition 1 says exactly that $\nabla_{\mathbf{x}} L(\mathbf{x}^\star,\boldsymbol\lambda,\boldsymbol\mu) = \mathbf{0}$: the point is a stationary point of the Lagrangian. The names come from W. Karush (1939) and from H. Kuhn and A. Tucker (1951), who found the conditions independently. With only equality constraints, conditions 3 and 4 disappear and you are back to the Lagrange conditions of Chapter 3.8.
Why do we need it?
It replaces the impossible job "compare with every other allowed point" by a few equations and inequalities about one point. It is the inequality version of "set the derivative to zero".
Where is it used?
The derivation of the SVM and its support vectors, the optimality test inside solvers for quadratic and linear programs, convergence checks in constrained solvers (IPOPT, cvxpy backends), and the Lasso's optimality conditions.
How is it used?
Either check a candidate (four questions, ✓/✗), or solve the conditions as a system: guess which walls are active, solve, and keep the answer that passes all four tests.
KKT is a test for a point and its multipliers together. A point can fail with some multipliers and pass with others. When we say "$\mathbf{x}^\star$ is a KKT point" we mean "there exist multipliers that make all four hold".
Standard form first. The signs above assume you minimise and write constraints as $g_i \le 0$. A constraint such as $x \ge 1$ must first become $1 - x \le 0$. If you skip this step, the sign of $\mu$ comes out backwards.
Quick check: why does condition 3 (dual feasibility) not apply to the multipliers $\lambda_j$ of equality constraints?
An equality wall $h_j = 0$ is like a rail that you cannot leave in either direction, so it may need to push you either way. Its multiplier $\lambda_j$ can be positive, negative or zero. An inequality fence only blocks one side, so it can only push inwards: $\mu_i \ge 0$.
1. Stationarity: the pull is balanced by the push core
Think of a tug of war. The objective pulls you downhill, along $-\nabla f$. The walls that touch you can only push you inwards. You stop moving when the pull and the push cancel exactly.
- If no wall touches you, there is nothing to push back, so the pull must already be zero: $\nabla f = \mathbf{0}$. This is the familiar "flat ground" rule from Chapter 3.2.
- If a wall touches you, the ground may still slope downhill towards it. That is fine, as long as the wall pushes back just as hard.
Minimise $f(x,y) = x + y$ inside the disc $x^2 + y^2 \le 2$ (that is, $g = x^2 + y^2 - 2 \le 0$). The ground slopes down towards the bottom left, so we expect the best point at the lower-left edge of the disc. Test $\mathbf{x}^\star = (-1, -1)$ (it lies on the edge: $1 + 1 = 2$).
- The gradient of the objective is $\nabla f = (1, 1)$. It points up and to the right (uphill).
- The gradient of the wall is $\nabla g = (2x, 2y) = (-2, -2)$ at $\mathbf{x}^\star$. It points out of the disc, to the lower left.
- Stationarity asks for a number $\mu$ with $\nabla f + \mu\nabla g = \mathbf{0}$: $(1,1) + \mu(-2,-2) = (1 - 2\mu,\ 1 - 2\mu)$.
- Both entries vanish when $\mu = \tfrac12$. So stationarity holds, with multiplier $\mu = \tfrac12 \ge 0$ ✓.
In tug-of-war language: the pull is $-\nabla f = (-1,-1)$, towards the wall. The push is $-\mu\nabla g = (1, 1)$, back into the disc. They cancel.
Stationarity. There are multipliers such that $$\nabla f(\mathbf{x}^\star) + \sum_{j=1}^{p}\lambda_j \nabla h_j(\mathbf{x}^\star) + \sum_{i=1}^{m}\mu_i \nabla g_i(\mathbf{x}^\star) = \mathbf{0}, \qquad\text{that is,}\qquad \nabla_{\mathbf{x}} L(\mathbf{x}^\star,\boldsymbol\lambda,\boldsymbol\mu) = \mathbf{0}.$$
Three ways to read it. Forces: the pull $-\nabla f$ equals the combined push of the walls $\sum \mu_i\nabla g_i + \sum\lambda_j\nabla h_j$ (up to sign). Calculus: it is the vector equation "derivative = 0" for the Lagrangian, which has the constraints already built in. Counting: it gives $n$ equations (one per coordinate of $\mathbf{x}$). Together with the $p$ equalities and the $m$ slackness equations you get $n + p + m$ equations for the $n + p + m$ unknowns $\mathbf{x}, \boldsymbol\lambda, \boldsymbol\mu$. A square system: that is why solving the conditions can work.
Why do we need it?
On a constrained problem the slope at the best point is not zero (the wall is holding you up), so "set $\nabla f = 0$" fails. Stationarity is the corrected rule: the slope must be exactly cancelled by the walls.
Where is it used?
Every Lagrange-multiplier derivation: the SVM's $\mathbf{w} = \sum \alpha_i y_i \mathbf{x}_i$, the ridge-regression and Lasso conditions, maximum-entropy models, and the optimality residual that constrained solvers print as "dual infeasibility".
How is it used?
Differentiate $L$ with respect to $\mathbf{x}$, set the result to zero, and solve for $\mathbf{x}$ in terms of the multipliers. Or, at a candidate point, solve for the $\mu$ that cancel $\nabla f$ and see whether the leftover is zero.
Stationarity alone does not mean "minimum". A stationary point of the Lagrangian can be a maximum or a saddle too. That is why we need the other three conditions, and (for non-convex problems) more checking.
Flipping the sign of $\mu$. Some books write $L = f - \sum\mu_i g_i$ with constraints $g_i \ge 0$. The physics is identical, only the bookkeeping differs. Here we always use $g_i \le 0$ and $+\mu_i g_i$.
Quick check: minimise $f(x) = x^2$ subject to $1 - x \le 0$. At $x^\star = 1$, what multiplier balances the pull?
$f'(1) = 2$ and $g'(x) = -1$. Stationarity: $2 + \mu\cdot(-1) = 0$, so $\mu = 2$. The pull (towards smaller $x$, strength 2) is cancelled by a push of strength 2 to the right.
2. Primal feasibility: obey every rule
This is the easiest of the four conditions, and the one people forget. Before asking whether a point is good, ask whether it is legal. A route that is the shortest but crosses a river with no bridge is not a valid route.
The word primal only means "about the original variables $\mathbf{x}$". (The multipliers $\lambda,\mu$ are called dual variables. You will see why in Chapter 3.10.)
Use the triangle problem: $g_1 = x + y - 2 \le 0$, $g_2 = -x \le 0$, $g_3 = -y \le 0$. Which of these points are allowed?
- $(1, 0.5)$: $g_1 = -0.5$, $g_2 = -1$, $g_3 = -0.5$. All $\le 0$. Allowed ✓ (inside, nothing touching).
- $(1.5, 0.5)$: $g_1 = 0$, $g_2 = -1.5$, $g_3 = -0.5$. Allowed ✓ (touching wall 1).
- $(2, 1)$: $g_1 = 1 \gt 0$. Not allowed ✗ (outside wall 1). Yet this is exactly where $f$ is smallest ($f = 0$)!
- $(-0.5, 1)$: $g_2 = 0.5 \gt 0$. Not allowed ✗.
Primal feasibility: $\mathbf{x}^\star$ lies in the feasible set $$g_i(\mathbf{x}^\star) \le 0 \ \ (i = 1,\dots,m), \qquad h_j(\mathbf{x}^\star) = 0 \ \ (j = 1,\dots,p).$$ A point that breaks any one of these is infeasible, however small $f$ is there.
Why do we need it?
The other three conditions only describe forces and signs. Without feasibility they would happily accept a point that is outside the allowed region, such as the unconstrained minimum.
Where is it used?
As the "primal residual" every constrained solver reports (how badly the rules are broken), in the feasibility phase of linear-programming solvers, and in checking that a trained SVM separates the training data.
How is it used?
Plug the point into every constraint. Numerically you accept a tiny violation: a solver stops when $\max_i \max(g_i, 0)$ and $|h_j|$ are below a tolerance such as $10^{-8}$.
Equalities are two-sided. $h(\mathbf{x}) = 0$ fails whether $h$ is a little positive or a little negative. An inequality $g \le 0$ only fails on the positive side.
Floating point. A computer almost never gets $h = 0$ exactly. In code, test $|h| \le \text{tol}$, never $h == 0$.
Quick check: is $(1, 1.2)$ feasible for the triangle problem?
$g_1 = 1 + 1.2 - 2 = 0.2 \gt 0$. It breaks the wall $x + y \le 2$, so it is infeasible, no matter what $f$ says.
3. Dual feasibility: walls push, they never pull core
A fence can stop you from walking out of a field. It cannot drag you out. So the force of an inequality wall always points into the allowed region. Mathematically that is the rule $\mu_i \ge 0$.
What if stationarity holds with a negative $\mu$? Then the wall would have to be pulling you outwards to keep you still. That means the ground is sloping downhill inwards, and you are not at a minimum at all: you only paused at a spot where the ground happens to be level along the wall. Step inside and you roll down.
The sign rule from a tiny picture. The allowed region is $x \le 1$, so $g(x) = x - 1 \le 0$, and $g'(x) = +1$ points out of the region (to the right). We stand on the wall at $x = 1$.
- Minimise $f(x) = -x$. The ground slopes down to the right, towards the wall. Stationarity: $f' + \mu g' = -1 + \mu = 0$, so $\mu = 1 \ge 0$ ✓. You want to go right, the wall stops you: $x = 1$ is the true minimum.
- Minimise $f(x) = x$. The ground slopes down to the left, away from the wall. Stationarity: $1 + \mu = 0$, so $\mu = -1 \lt 0$ ✗. At $x = 1$ you could walk left and lower $f$. Indeed $x = 1$ is where $f = x$ is largest on the region.
Both cases satisfy stationarity (with some $\mu$) and feasibility. Only the sign of $\mu$ tells them apart.
Dual feasibility: every multiplier of an inequality constraint is non-negative, $$\mu_i \ge 0 \quad (i = 1,\dots,m).$$ The multipliers $\lambda_j$ of equality constraints may have any sign.
Derivation from the picture. At a point on an active wall, $\nabla g_i$ points out of the region (the direction in which $g_i$ grows). Stationarity with a single active wall reads $-\nabla f = \mu\,\nabla g$. The vector $-\nabla f$ is the downhill direction.
- If $\mu \gt 0$, downhill points outwards, where you are not allowed to go. You are stuck at a local best.
- If $\mu \lt 0$, downhill points inwards, which is allowed. You can still improve, so this is not a minimum.
Why do we need it?
Stationarity cannot tell a constrained minimum from a constrained maximum: both balance forces. The sign of $\mu$ is the extra piece of information that separates them.
Where is it used?
The constraint $\alpha_i \ge 0$ in the SVM dual, the non-negativity of the multiplier in every active-set solver (a negative multiplier is the signal to drop a constraint from the active set), and the sign check in interior-point methods.
How is it used?
After solving the stationarity equations, look at each $\mu_i$. All $\ge 0$: fine. A negative one means "this wall is not really holding me; release it and try again".
It all depends on the form. The rule "$\mu \ge 0$" is for minimising with constraints written as $g \le 0$. If you maximise, or write $g \ge 0$, the sign flips. When a derivation gives you a negative multiplier, first check whether you started from the standard form.
A negative multiplier is a message, not an error. It says "this wall is not holding you in; the ground really slopes inwards. Drop it from the active set".
Quick check: minimise $f(x) = (x-3)^2$ subject to $x \le 5$. If a solver reported the point $x = 5$ with some $\mu$, what sign would $\mu$ have, and is $x = 5$ optimal?
$f'(5) = 2(5-3) = 4$ and $g' = +1$: $4 + \mu = 0$ gives $\mu = -4 \lt 0$. Negative, so dual feasibility fails: the ground slopes down to the left (towards $x = 3$), into the allowed region. The true optimum is $x = 3$ with $\mu = 0$.
4. Complementary slackness: only touching walls push core
Think of a wall with a gap in front of you. Nothing is pressing on it, so it pushes with zero force. Only when you lean against the wall (the gap is zero) can it push back.
So for every wall, one of two things is zero: the gap (slack) in front of it, or its push $\mu$. Maybe both, never neither. "Complementary" means exactly this: the two quantities take turns being zero.
Back to the triangle problem at $\mathbf{x}^\star = (1.5, 0.5)$, with $\mu = (1, 0, 0)$:
| wall | $g_i(\mathbf{x}^\star)$ | slack $-g_i$ | $\mu_i$ | $\mu_i\,g_i$ |
|---|---|---|---|---|
| $x + y \le 2$ | $0$ | $0$ (touching) | $1$ (pushes) | $0$ ✓ |
| $x \ge 0$ | $-1.5$ | $1.5$ (far) | $0$ (idle) | $0$ ✓ |
| $y \ge 0$ | $-0.5$ | $0.5$ (far) | $0$ (idle) | $0$ ✓ |
Only the touching wall has a non-zero push. Every product $\mu_i g_i$ is zero.
Complementary slackness: for every inequality constraint, $$\mu_i\, g_i(\mathbf{x}^\star) = 0.$$ Because $\mu_i \ge 0$ and $g_i \le 0$, this is the same as the pair of statements $$g_i(\mathbf{x}^\star) \lt 0 \;\Rightarrow\; \mu_i = 0 \qquad\text{and}\qquad \mu_i \gt 0 \;\Rightarrow\; g_i(\mathbf{x}^\star) = 0.$$ Inactive walls have no multiplier; a wall with a positive multiplier must be active. (If both $\mu_i = 0$ and $g_i = 0$ the wall is weakly active.)
Why do we need it?
It tells you which multipliers are zero without solving anything: all the far-away walls drop out. That is what turns an inequality problem into a finite list of "which walls are touching?" guesses.
Where is it used?
SVM support vectors ($\alpha_i \gt 0$ only for points on the margin), sparsity of the dual solution, active-set and interior-point solvers (the "complementarity gap" $\sum\mu_i s_i$ is what they drive to zero), and the Lasso (a weight is exactly $0$ unless its gradient hits the boundary $\pm\lambda$).
How is it used?
For each wall decide: touching, so solve with $g_i = 0$ and check $\mu_i \ge 0$; or not touching, so set $\mu_i = 0$ and check $g_i \le 0$. Those are the two branches of the case analysis in the next sections.
"Zero" can happen twice. A wall can have $g_i = 0$ and $\mu_i = 0$: touching but idle (like $\min x^2$ with $x \ge 0$). That is allowed. When exactly one of the two is zero for every wall we say strict complementarity holds, which is the nicer case for algorithms.
It is one equation per wall, not per point. $\mu_i g_i = 0$ is a product. Do not confuse it with "$\mu_i = 0$ and $g_i = 0$ at the same time".
Quick check: at the optimum, a constraint has $g_i(\mathbf{x}^\star) = -3$. What is $\mu_i$?
$\mu_i \cdot (-3) = 0$ forces $\mu_i = 0$. The wall is 3 units away, so it exerts no push.
All four at once: a KKT checker
You have now met each condition on its own. In practice you meet them together: someone gives you a point, and you want a verdict. Build a mental routine and always run it in the same order:
- Which walls touch me? (find the active set)
- Am I allowed here? (feasibility)
- Choose $\mu$ for the touching walls so that the pull is cancelled. Is anything left over? (stationarity)
- Are all those $\mu \ge 0$? (dual feasibility)
- Walls that do not touch me get $\mu = 0$. (complementary slackness)
Test $(1, 0)$ on the box problem "minimise $-(x^2+y^2)$ inside $|x|\le 1,\ |y|\le 1$" (a non-convex objective: it is a hill with its top at the centre). Only the wall $x \le 1$ touches, so only $g_1 = x - 1$ is active.
- Feasible: $g_1 = 0$, $g_2 = -2$, $g_3 = -1$, $g_4 = -1$. All $\le 0$ ✓.
- $\nabla f = (-2x, -2y) = (-2, 0)$. With $\nabla g_1 = (1, 0)$: $(-2,0) + \mu_1(1,0) = 0$ gives $\mu_1 = 2$ ✓ stationary.
- $\mu_1 = 2 \ge 0$ ✓, and the other three have $\mu = 0$ ✓ (slackness).
So $(1,0)$ is a KKT point. But it is not a minimum: sliding up or down the wall lowers $f$ to $-1 - y^2$. The checker below lets you find this trap yourself.
The checking recipe. Given a candidate $\mathbf{x}$:
- Active set $\mathcal{A} = \{i : g_i(\mathbf{x}) = 0\} \cup \{\text{all equalities}\}$.
- Check primal feasibility for all constraints.
- Solve $\nabla f + \sum_{k\in\mathcal{A}} m_k\nabla c_k = \mathbf{0}$ for the active multipliers $m_k$ (for several walls, this is a small linear system or a least-squares fit; the leftover is the residual). Stationarity holds iff the residual is $\mathbf{0}$.
- Dual feasibility: the multipliers of active inequalities are $\ge 0$.
- Complementary slackness: give every inactive inequality $\mu = 0$ (automatic in this recipe).
Why do we need it?
Optimisation code returns a point. You need an independent test that the point deserves to be called a solution, not just "the loop stopped".
Where is it used?
Convergence tests in constrained solvers (SciPy's minimize with SLSQP or trust-constr, IPOPT), unit tests for hand-written solvers, and sanity checks of an SVM fit.
How is it used?
Compute the residuals of the four conditions at the returned point. If all are below a tolerance, accept it. Large residual in one condition tells you which part is wrong.
"Active" is decided with a tolerance. In the widget a wall counts as touching when $|g_i| \le 0.011$. Real solvers use a similar small tolerance, and borderline cases are where they sometimes struggle.
The checker never says "optimal" for a non-convex problem. That is deliberate. The next section explains exactly what a KKT point does and does not guarantee.
Quick check: on the disc problem ($\min x+y$ with $x^2 + y^2 \le 2$), is $(1,1)$ a KKT point? Is it a minimum?
It is on the wall ($1 + 1 = 2$), $\nabla f = (1,1)$ and $\nabla g = (2,2)$, so stationarity gives $1 + 2\mu = 0$, $\mu = -\tfrac12 \lt 0$. Dual feasibility fails, so it is not a KKT point. It is the constrained maximum of $x+y$. The minimum is $(-1,-1)$ with $\mu = +\tfrac12$.
KKT and constrained optimization: necessary or sufficient? core
The KKT conditions are a checklist. Two natural questions follow, and they are different questions:
- Necessary: if a point really is the best, must it pass the checklist? If yes, then we can hunt for the best point among the points that pass.
- Sufficient: if a point passes the checklist, must it be the best? If yes, then passing is a proof of optimality.
The answers are: usually yes to the first (if the walls are "well-behaved"), and yes only for convex problems to the second. A hilly, non-convex landscape has many places where forces balance: valleys, hilltops and passes.
One-dimensional and non-convex: minimise $f(x) = \tfrac14x^4 - x^2 + 0.3x$ with the walls $-2 \le x \le 1.8$. The function has two valleys and a small hill between them.
- In the interior no wall is active, so KKT says $f'(x) = x^3 - 2x + 0.3 = 0$. This cubic has three roots: $x \approx -1.484$, $0.152$, $1.332$.
- Values: $f(-1.484) \approx -1.435$ (deepest valley), $f(0.152) \approx 0.023$ (top of the hill), $f(1.332) \approx -0.588$ (shallower valley).
- At the wall $x = 1.8$: $f'(1.8) = 2.53 \gt 0$, so $\mu = -f'(1.8) = -2.53 \lt 0$ ✗. At the wall $x = -2$: $\mu = f'(-2) = -3.7 \lt 0$ ✗. Neither wall is a KKT point.
So there are three KKT points: a global minimum, a local minimum and a local maximum. Every minimum is among them (necessary), but not every KKT point is a minimum (not sufficient).
Necessity. Suppose $f, g_i, h_j$ are continuously differentiable, $\mathbf{x}^\star$ is a local minimum, and a constraint qualification (next section) holds at $\mathbf{x}^\star$. Then there exist multipliers $\boldsymbol\lambda, \boldsymbol\mu$ such that $(\mathbf{x}^\star,\boldsymbol\lambda,\boldsymbol\mu)$ satisfy all four KKT conditions.
Sufficiency. Suppose the problem is convex: $f$ and all $g_i$ are differentiable convex functions and all $h_j$ are affine (Chapter 3.6). If $(\mathbf{x}^\star,\boldsymbol\lambda,\boldsymbol\mu)$ satisfy the KKT conditions, then $\mathbf{x}^\star$ is a global minimum.
| Problem | Is the minimum a KKT point? | Is every KKT point a minimum? |
|---|---|---|
| Convex, with a constraint qualification | Yes | Yes (global) |
| Non-convex, with a constraint qualification | Yes (local minima too) | No: also maxima, saddles, other local minima |
| Convex, no constraint qualification | Not guaranteed | Yes (sufficiency never needed a qualification) |
Proof of sufficiency (it shows why all four conditions are needed). Let $\mathbf{x}$ be any feasible point.
- For fixed $\boldsymbol\lambda$ and $\boldsymbol\mu \ge 0$, the function $\mathbf{x} \mapsto L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu)$ is convex (a convex $f$, plus non-negative multiples of convex $g_i$, plus affine terms). By stationarity its gradient is zero at $\mathbf{x}^\star$, and a convex function with zero gradient is at its global minimum. So $L(\mathbf{x},\cdot) \ge L(\mathbf{x}^\star,\cdot)$ for every $\mathbf{x}$.
- At the feasible point $\mathbf{x}$ we have $h_j = 0$ and $\mu_i g_i \le 0$, so the extra terms in $L$ are $\le 0$: $L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu) \le f(\mathbf{x})$. (This uses dual feasibility.)
- At $\mathbf{x}^\star$ the extra terms are exactly $0$ ($h_j = 0$, and $\mu_i g_i = 0$ by slackness), so $L(\mathbf{x}^\star,\boldsymbol\lambda,\boldsymbol\mu) = f(\mathbf{x}^\star)$.
- Chain them: $f(\mathbf{x}) \ge L(\mathbf{x},\cdot) \ge L(\mathbf{x}^\star,\cdot) = f(\mathbf{x}^\star)$. Since $\mathbf{x}$ was any feasible point, $\mathbf{x}^\star$ is the global minimum. ∎
Why do we need it?
It tells you how far you can trust the checklist. On a convex problem, passing it is a proof of optimality. On a non-convex problem, it only narrows the search to candidates.
Where is it used?
SVMs, ridge, Lasso and logistic regression with constraints are all convex, so a KKT point is the answer. Non-convex training (neural networks) has no such guarantee: KKT-like conditions are used as stopping tests, and the point reached is only a local candidate.
How is it used?
First decide: is the problem convex (convex $f$ and $g_i$, affine $h_j$)? If yes, find any KKT point and you are done. If not, collect the KKT points and compare their objective values (and check second-order conditions, Chapter 3.2).
"KKT point" does not mean "solution". People often write "we solved the KKT conditions, so we found the minimum". That sentence is only true for convex problems (or after you have shown the point beats all other KKT points).
Local vs global. Necessity is about local minima; a local minimum that is not the global one is also a KKT point. If you need the global answer on a non-convex problem, KKT alone cannot give it.
Quick check: a problem has a convex objective, convex constraints, and you find a point where all four KKT conditions hold. What can you conclude, and which condition is used to show that the extra terms in $L$ disappear at $\mathbf{x}^\star$?
You can conclude $\mathbf{x}^\star$ is a global minimum (sufficiency for convex problems). Complementary slackness ($\mu_i g_i(\mathbf{x}^\star) = 0$) together with $h_j(\mathbf{x}^\star) = 0$ makes $L(\mathbf{x}^\star) = f(\mathbf{x}^\star)$ in step 3 of the proof.
Constraint qualifications: when multipliers exist
The KKT conditions describe the allowed region near $\mathbf{x}^\star$ by its straight-line approximation: each active wall is replaced by its tangent line, whose direction is given by $\nabla g_i$. That works when the walls look like proper walls. It breaks when a wall has no usable normal direction, for example when $\nabla g_i = \mathbf{0}$ at the point, or when two walls squeeze the region to a sharp cusp.
A constraint qualification (CQ) is a promise that the walls are well-behaved, so the straight-line picture is honest.
Minimise $f(x) = x$ subject to $x^2 \le 0$, i.e. $g(x) = x^2 \le 0$. The only allowed point is $x = 0$, so $x^\star = 0$ is the minimum. Is it a KKT point?
- $f'(x) = 1$ and $g'(x) = 2x$, so at $x^\star = 0$: $\nabla g = 0$.
- Stationarity: $1 + \mu\cdot 0 = 1 \ne 0$ for every $\mu$. No multiplier exists.
The minimum fails KKT because the wall has no direction to push along. Now relax the wall a little: $x^2 \le \varepsilon$ with $\varepsilon \gt 0$. The allowed interval is $[-\sqrt\varepsilon, \sqrt\varepsilon]$, the minimum is $x^\star = -\sqrt\varepsilon$, and stationarity $1 + \mu\cdot 2x^\star = 0$ gives $\mu = \dfrac{1}{2\sqrt\varepsilon}$. As $\varepsilon \to 0$ the multiplier grows without bound, and at $\varepsilon = 0$ it no longer exists.
The two qualifications you will meet most often:
- LICQ (linear independence CQ). At $\mathbf{x}^\star$, the gradients of all equality constraints and of the active inequality constraints, $\{\nabla h_j(\mathbf{x}^\star)\} \cup \{\nabla g_i(\mathbf{x}^\star) : i \in \mathcal{A}\}$, are linearly independent (linear independence: none is a combination of the others). It guarantees that the multipliers exist and are unique.
- Slater's condition (for convex problems with affine equalities). There is a strictly feasible point $\hat{\mathbf{x}}$: $g_i(\hat{\mathbf{x}}) \lt 0$ for all $i$ (for affine $g_i$, $\le 0$ is enough) and $h_j(\hat{\mathbf{x}}) = 0$. In words: the allowed region has an inside, it is not paper-thin.
Also: if all constraints are affine (straight walls), no extra qualification is needed: the straight-line picture is exact. In the example above, LICQ fails ($\nabla g = 0$ is not a linearly independent set) and Slater fails (no $x$ has $x^2 \lt 0$).
Why do we need it?
Without it, a perfectly good minimum can fail the KKT test, and a solver that looks for KKT points would never find it. It is the fine print on the necessity theorem.
Where is it used?
In theory it justifies every Lagrange/KKT derivation (SVM, SVR, constrained least squares, maximum entropy), and Slater's condition is what makes strong duality hold in Chapter 3.10. In practice, huge multipliers in a solver's output are a symptom of a nearly failing CQ.
How is it used?
Check the cheap cases first: all constraints affine? Convex with a strictly feasible point? Then you are safe. Otherwise at the candidate, test that the active gradients are linearly independent (their matrix has full row rank).
Failing a CQ does not make a point bad. It only means the KKT test may reject a true minimum. Such cases are unusual in ML, because the constraints we use (non-negativity, linear equalities, margins) are affine.
LICQ is about gradients at the point. It can hold at some points of a problem and fail at others. Checking it everywhere is not required; only at the point you care about.
Quick check: minimise $x + y$ subject to $x \ge 0$ and $y \ge 0$ at the corner $(0,0)$. Do the active gradients satisfy LICQ?
In standard form $g_1 = -x$, $g_2 = -y$, so $\nabla g_1 = (-1, 0)$ and $\nabla g_2 = (0, -1)$. These two vectors are linearly independent, so LICQ holds. (The constraints are affine anyway.) The multipliers are unique: $\nabla f = (1,1)$, so $\mu_1 = \mu_2 = 1$.
Solving problems by case analysis over active sets core
Complementary slackness splits every wall into two possibilities: touching (then solve with $g_i = 0$) or not touching (then $\mu_i = 0$). With $m$ walls there are $2^m$ combinations. For each combination the equations become ordinary ones that you can solve. Then you keep only the combinations whose answer is consistent with the assumption.
It is like a detective's list of suspects: try every guess, and eliminate the guesses that contradict themselves.
The recipe. For each subset $S$ of the inequalities:
- Assume $g_i = 0$ for $i \in S$ and $\mu_i = 0$ for $i \notin S$.
- Solve stationarity together with $g_i = 0$ ($i \in S$) for $\mathbf{x}$ and $\mu_S$.
- Test the assumption: $g_i \le 0$ for $i \notin S$ (feasible) and $\mu_i \ge 0$ for $i \in S$ (right sign).
- Keep the survivors. For a convex problem, a survivor is the global solution.
Example A. $\min (x-2)^2 + (y-1)^2$ s.t. $g_1 = x + y - 2 \le 0$, $g_2 = -x \le 0$, $g_3 = -y \le 0$. There are $2^3 = 8$ cases. Take $S = \{1\}$: stationarity gives $2(x-2) + \mu_1 = 0$ and $2(y-1) + \mu_1 = 0$, so $x = 2 - \mu_1/2$ and $y = 1 - \mu_1/2$. The wall $x + y = 2$ then gives $3 - \mu_1 = 2$, i.e. $\mu_1 = 1$, $\mathbf{x} = (1.5, 0.5)$. Test: $g_2 = -1.5 \le 0$ ✓, $g_3 = -0.5 \le 0$ ✓, $\mu_1 = 1 \ge 0$ ✓. A survivor. The full table:
| active set $S$ | solution $\mathbf{x}$ | multipliers | verdict |
|---|---|---|---|
| none | $(2, 1)$ | ✗ violates $g_1$ ($=1 \gt 0$) | |
| $\{g_1\}$ | $(1.5, 0.5)$ | $\mu_1 = 1$ | ✓ KKT point: the answer |
| $\{g_2\}$ | $(0, 1)$ | $\mu_2 = -4$ | ✗ negative multiplier |
| $\{g_3\}$ | $(2, 0)$ | $\mu_3 = -2$ | ✗ negative multiplier |
| $\{g_1, g_2\}$ | $(0, 2)$ | $\mu_1 = -2,\ \mu_2 = -6$ | ✗ negative |
| $\{g_1, g_3\}$ | $(2, 0)$ | $\mu_1 = 0,\ \mu_3 = -2$ | ✗ negative |
| $\{g_2, g_3\}$ | $(0, 0)$ | $\mu_2 = -4,\ \mu_3 = -2$ | ✗ negative |
| all three | the three walls have no common point | ✗ impossible | |
Exactly one survivor, $\mathbf{x}^\star = (1.5, 0.5)$ with $f = 0.25 + 0.25 = 0.5$. The problem is convex, so it is the global minimum.
Example B (a corner). Move the wish to $(-1, 3)$. Now the case $S = \{1, 2\}$ ($x + y = 2$ and $x = 0$) gives $\mathbf{x} = (0, 2)$; stationarity $2(x + 1) + \mu_1 - \mu_2 = 0$ and $2(y - 3) + \mu_1 = 0$ yield $\mu_1 = 2$ and $\mu_2 = 4$. Both are positive and $g_3 = -2 \le 0$, so this corner is the answer: $f = 1 + 1 = 2$.
Example C (a one-dimensional problem with two walls). $\min (x-3)^2$ s.t. $g_1 = x - 2 \le 0$, $g_2 = 1 - x \le 0$ (so $1 \le x \le 2$). Four cases: none gives $x = 3$, violates $g_1$ ✗. $g_1$ only gives $x = 2$ and $f'(2) + \mu_1 = -2 + \mu_1 = 0$, $\mu_1 = 2 \ge 0$, $g_2 = -1 \le 0$ ✓ answer. $g_2$ only gives $x = 1$ and $f'(1) - \mu_2 = -4 - \mu_2 = 0$, $\mu_2 = -4$ ✗. Both needs $x = 2$ and $x = 1$ at once ✗. Answer: $x^\star = 2$, $\boldsymbol\mu = (2, 0)$, $f = 1$.
Active-set case analysis. For $m$ inequality constraints, enumerate the $2^m$ guesses $S \subseteq \{1,\dots,m\}$ of the active set. Each guess turns the KKT system into equations only: $$\nabla f(\mathbf{x}) + \sum_{i \in S}\mu_i\nabla g_i(\mathbf{x}) + \sum_j\lambda_j\nabla h_j(\mathbf{x}) = \mathbf{0},\qquad g_i(\mathbf{x}) = 0\ (i \in S),\qquad h_j(\mathbf{x}) = 0,$$ with $\mu_i = 0$ for $i \notin S$. A guess is consistent if its solution satisfies $g_i(\mathbf{x}) \le 0$ for $i \notin S$ and $\mu_i \ge 0$ for $i \in S$. Every consistent guess is a KKT point.
For a quadratic $f$ with linear walls, each case is a small linear system. The cost doubles with every extra constraint, so this is a method for tiny problems and for understanding; real solvers walk through the active sets cleverly (active-set methods) or avoid guessing altogether (interior-point methods, Chapter 3.17).
Why do we need it?
It gives you a mechanical, always-works procedure for small constrained problems, and it is the idea behind how the best-known solvers for linear and quadratic programs think.
Where is it used?
Exam-style and hand-derived solutions, active-set QP solvers, the simplex method (each vertex is an active-set guess), and reasoning about which features a Lasso keeps (a feature is "in" or "out").
How is it used?
Write the problem in standard form, list the walls, try the cases from "nothing active" upwards, and use the two tests (feasible? multiplier sign?) to throw out wrong guesses quickly.
Several survivors. In a non-convex problem more than one guess can survive. Then compare their $f$ values; the smallest is the global minimum (provided a global minimum exists and a constraint qualification holds at it, so that it is among the KKT points).
Weakly active walls. If the wish lands exactly on a wall or a corner, a multiplier can be exactly $0$ in a case with that wall assumed active. That guess and its neighbour then give the same point; both are fine.
Quick check: $\min x^2$ subject to $1 - x \le 0$. Solve by cases.
Not active ($\mu = 0$): $2x = 0$ gives $x = 0$, but $g = 1 - 0 = 1 \gt 0$ ✗. Active ($x = 1$): $2x - \mu = 0$ gives $\mu = 2 \ge 0$ ✓ and the only wall is active. So $x^\star = 1$, $\mu = 2$, $f = 1$.
The geometric picture: $-\nabla f$ lies in the cone of the active gradients
Stand at a corner where two walls meet. Each wall has an outward-pointing arrow $\nabla g_i$. Two arrows span a wedge (a "cone"): all the directions you get by taking some non-negative amount of each arrow.
The ground pulls you downhill along $-\nabla f$. If the pull points inside the wedge, both walls can team up to hold you, and you are stuck: a KKT point. If the pull points outside the wedge, some feasible direction still goes downhill, and you can keep moving.
Corner at $(0, 2)$ of the triangle, where the walls $x + y \le 2$ (normal $\mathbf{n}_1 = (1,1)$) and $x \ge 0$ (normal $\mathbf{n}_2 = (-1, 0)$) meet. Let the pull be $\mathbf{d} = -\nabla f = (-2, 2)$. Write $\mathbf{d} = \mu_1\mathbf{n}_1 + \mu_2\mathbf{n}_2 = (\mu_1 - \mu_2,\ \mu_1)$. Then $\mu_1 = 2$, and $\mu_1 - \mu_2 = -2$ gives $\mu_2 = 4$. Both are $\ge 0$, so $\mathbf{d}$ is inside the cone ✓ (and this is exactly Example B from before).
Now take $\mathbf{d} = (1, -1)$: $\mu_1 = -1 \lt 0$ ✗. It is outside the cone. Indeed the feasible direction $\mathbf{v} = (1,-1)$ (sliding down the hypotenuse) has $\mathbf{d}\cdot\mathbf{v} = 2 \gt 0$: it still goes downhill.
The cone generated by vectors $\mathbf{a}_1, \dots, \mathbf{a}_k$ is $$\operatorname{cone}\{\mathbf{a}_1,\dots,\mathbf{a}_k\} = \Big\{\textstyle\sum_{i}\mu_i\mathbf{a}_i \;:\; \mu_i \ge 0\Big\}.$$ One vector gives a ray, two give a wedge, none gives only $\{\mathbf{0}\}$. For $\mathcal{A}$ the active set of inequalities, the KKT conditions (stationarity, dual feasibility, slackness, no equalities) say $$-\nabla f(\mathbf{x}^\star) \in \operatorname{cone}\{\nabla g_i(\mathbf{x}^\star) : i \in \mathcal{A}\}.$$ With equality constraints, their gradients may enter with any sign (a whole subspace, not just a cone).
Why this is the same as "no feasible direction goes downhill". A direction $\mathbf{v}$ is feasible (to first order) if $\nabla g_i\cdot\mathbf{v} \le 0$ for every active wall (it does not point out). It is improving if $-\nabla f\cdot\mathbf{v} \gt 0$. A classical result, Farkas' lemma, says that exactly one of these is true: either $-\nabla f$ is in the cone of the active gradients, or some feasible direction is improving. KKT is the first alternative.
Why do we need it?
It turns the algebra of $\mu \ge 0$ into a picture you can sketch on a napkin, and it explains why the signs are what they are: a cone only contains non-negative combinations.
Where is it used?
Proofs of KKT necessity, the geometry of linear programming (the optimum vertex is where the negative cost vector lies in the cone of the active constraint normals), and projected-gradient methods, which stop when the negative gradient lies in the normal cone of the feasible set.
How is it used?
At a candidate, draw the outward normals of the touching walls and the downhill arrow. If the downhill arrow is between the normals, stop. If it is outside, the side it leans towards tells you which wall to release.
Normals point out, forces point in. The cone is made of the outward normals $\nabla g_i$. The pull $-\nabla f$ lies in it. The wall's actual push, $-\mu_i\nabla g_i$, points the opposite way (inwards). Keep the two arrows apart in your head.
First-order only. The cone argument uses gradients, i.e. first-order information. On a non-convex problem it cannot tell a minimum from a saddle, which is why KKT points need extra checks there.
Quick check: at a point where one wall is active with normal $\mathbf{n} = (0, 1)$, which pulls $\mathbf{d} = -\nabla f$ are KKT?
The cone is the ray $\{\mu\,(0,1) : \mu \ge 0\}$: the pull must point straight up (along the outward normal, into the wall) or be zero. A pull such as $(1, 1)$ has a sideways part, so you could slide along the wall; a pull such as $(0,-1)$ points inwards, giving $\mu = -1 \lt 0$.
Preview: support vectors come from complementary slackness
A support vector machine separates two groups of points with the widest possible "street". Training points are the "walls": each one must stay on its own side of the street. Most points are far from the street, so those walls are inactive and push with force $0$.
Only the points sitting exactly on the edge of the street are touching their wall. They are the only ones that hold the street in place. These are the support vectors. Complementary slackness is the reason that all other points can be deleted without changing the answer.
Four points in the plane: $\mathbf{x}_1 = (1,1)$ and $\mathbf{x}_2 = (3,2)$ with label $y = +1$; $\mathbf{x}_3 = (-1,-1)$ and $\mathbf{x}_4 = (-2,-3)$ with label $y = -1$. The hard-margin SVM minimises $\tfrac12\|\mathbf{w}\|^2$ subject to $y_i(\mathbf{w}^\top\mathbf{x}_i + b) \ge 1$. The solution is $\mathbf{w} = (0.5, 0.5)$, $b = 0$. Look at the margins $y_i(\mathbf{w}^\top\mathbf{x}_i + b)$:
| point | $y_i$ | $\mathbf{w}^\top\mathbf{x}_i + b$ | margin $y_i(\cdot)$ | wall | multiplier $\alpha_i$ |
|---|---|---|---|---|---|
| $\mathbf{x}_1 = (1,1)$ | $+1$ | $0.5 + 0.5 = 1$ | $1$ | active | $0.25$ |
| $\mathbf{x}_2 = (3,2)$ | $+1$ | $1.5 + 1 = 2.5$ | $2.5$ | inactive | $0$ |
| $\mathbf{x}_3 = (-1,-1)$ | $-1$ | $-1$ | $1$ | active | $0.25$ |
| $\mathbf{x}_4 = (-2,-3)$ | $-1$ | $-2.5$ | $2.5$ | inactive | $0$ |
The margin-1 points are the support vectors. Check $\mathbf{w} = \sum_i\alpha_iy_i\mathbf{x}_i = 0.25(1,1) - 0.25(-1,-1) = (0.5, 0.5)$ ✓ and $\sum_i\alpha_iy_i = 0.25 - 0.25 = 0$ ✓. (The multipliers are found in Chapter 3.10.)
Hard-margin SVM in standard form. Variables $\mathbf{w}, b$. Minimise $\tfrac12\|\mathbf{w}\|^2$ subject to $g_i = 1 - y_i(\mathbf{w}^\top\mathbf{x}_i + b) \le 0$ for each training example. One multiplier $\alpha_i \ge 0$ per example (the usual SVM name for $\mu_i$): $$L(\mathbf{w}, b, \boldsymbol\alpha) = \tfrac12\|\mathbf{w}\|^2 + \sum_i\alpha_i\big(1 - y_i(\mathbf{w}^\top\mathbf{x}_i + b)\big).$$ The KKT conditions read:
- Stationarity in $\mathbf{w}$: $\mathbf{w} - \sum_i\alpha_iy_i\mathbf{x}_i = \mathbf{0}$, so $\mathbf{w} = \sum_i\alpha_iy_i\mathbf{x}_i$.
- Stationarity in $b$: $-\sum_i\alpha_iy_i = 0$.
- Primal feasibility: $y_i(\mathbf{w}^\top\mathbf{x}_i + b) \ge 1$ for all $i$.
- Dual feasibility: $\alpha_i \ge 0$.
- Complementary slackness: $\alpha_i\big(1 - y_i(\mathbf{w}^\top\mathbf{x}_i + b)\big) = 0$, so $\alpha_i \gt 0 \Rightarrow y_i(\mathbf{w}^\top\mathbf{x}_i + b) = 1$.
The points with $\alpha_i \gt 0$ are the support vectors. The weight vector $\mathbf{w}$ is a mix of only the support vectors. The problem is convex with affine constraints, so a KKT point is the global solution.
Why do we need it?
It explains the sparsity of an SVM: out of thousands of training points only a few carry weight, so the model is cheap to store and its predictions depend only on those few.
Where is it used?
Every SVM library (LIBSVM, scikit-learn's SVC): its support_vectors_ attribute lists the points with $\alpha_i \gt 0$. The same pattern appears in support vector regression and in the "active constraints" of many ML formulations.
How is it used?
After training, keep only the support vectors and their $\alpha_i y_i$. A new point is classified by the sign of $\sum_i\alpha_iy_i\,\mathbf{x}_i^\top\mathbf{x} + b$. Deeper treatment, with the dual problem and kernels, is in Chapter 3.10.
Support vectors are the points that matter, not the points that are "hard". In the hard-margin case they are exactly the points on the margin. (With a soft margin, in Chapter 3.10, points inside the street or on the wrong side are support vectors too.)
Several points can sit on the margin. Some of them may still have $\alpha_i = 0$ (touching but idle, the weakly active case from before). Then $\alpha_i = 0$ even though they are on the margin.
Quick check: two training points, $\mathbf{x}_1 = (1,1)$ with $y = +1$ and $\mathbf{x}_2 = (-1,-1)$ with $y = -1$, give $\alpha_1 = \alpha_2 = \tfrac14$. Find $\mathbf{w}$ and the street width.
$\mathbf{w} = \tfrac14(+1)(1,1) + \tfrac14(-1)(-1,-1) = (0.25, 0.25) + (0.25, 0.25) = (0.5, 0.5)$. Its length is $\|\mathbf{w}\| = \sqrt{0.5} \approx 0.707$, so the street width is $2/\|\mathbf{w}\| = 2\sqrt2 \approx 2.83$: exactly the distance between the two points ($\|(2,2)\| = 2\sqrt2$).
Recap, cheat sheet and practice
- A constraint $g_i \le 0$ is active if $g_i = 0$ (touching) and inactive if $g_i \lt 0$. Only active ones can push; inactive ones have $\mu_i = 0$.
- The KKT conditions are four tests: stationarity ($\nabla f + \sum\lambda_j\nabla h_j + \sum\mu_i\nabla g_i = \mathbf{0}$, pull = push), primal feasibility (all rules obeyed), dual feasibility ($\mu_i \ge 0$: fences push, never pull) and complementary slackness ($\mu_ig_i = 0$: only touching walls push).
- The sign rule comes from a picture: at an active wall, $-\nabla f = \mu\nabla g$. With $\mu \gt 0$ downhill points into the wall (stuck: a minimum); with $\mu \lt 0$ it points inside (you can still improve).
- Necessary: a local minimum is a KKT point if a constraint qualification holds (LICQ, Slater, or all constraints affine). Sufficient: for a convex problem any KKT point is a global minimum. Otherwise KKT points include maxima and saddles.
- To solve: guess the active set ($2^m$ cases), solve the equations, keep the guesses with $g_i \le 0$ off the set and $\mu_i \ge 0$ on it.
- Geometry: $-\nabla f$ lies in the cone of the active gradients $\nabla g_i$ (and in their span for equalities): no feasible direction goes downhill.
- SVM: $\mathbf{w} = \sum\alpha_iy_i\mathbf{x}_i$, $\sum\alpha_iy_i = 0$, and slackness makes $\alpha_i \gt 0$ only for points on the margin: the support vectors.
Cheat sheet
| Idea | Formula | Picture |
|---|---|---|
| Standard form | $\min f$ s.t. $g_i \le 0$, $h_j = 0$ | fences you stay inside |
| Lagrangian | $L = f + \sum\lambda_jh_j + \sum\mu_ig_i$ | objective + price × violation |
| Stationarity | $\nabla_{\mathbf{x}}L = \mathbf{0}$ | pull cancels push |
| Primal feasibility | $g_i \le 0$, $h_j = 0$ | you are inside the field |
| Dual feasibility | $\mu_i \ge 0$ | walls only push inwards |
| Complementary slackness | $\mu_ig_i = 0$ | far walls are idle |
| Convex problem + KKT | KKT point $\Rightarrow$ global minimum | one valley |
| Constraint qualification | LICQ: active gradients independent; Slater: strictly feasible point | walls have usable normals |
| Gradient cone | $-\nabla f \in \operatorname{cone}\{\nabla g_i, i\in\mathcal{A}\}$ | downhill points into the walls |
| Case analysis | $2^m$ active-set guesses | try, test, discard |
| SVM | $\alpha_i\big(1 - y_i(\mathbf{w}^\top\mathbf{x}_i + b)\big) = 0$ | support vectors on the street edge |
import numpy as np
from itertools import combinations
from scipy.optimize import minimize
# Problem: min (x-2)^2 + (y-1)^2 s.t. x + y <= 2, x >= 0, y >= 0
c = np.array([2.0, 1.0])
f = lambda z: np.sum((z - c) ** 2)
grad_f = lambda z: 2 * (z - c)
# walls written as a_i . z + b_i <= 0 (standard form g_i(z) <= 0)
A = np.array([[1.0, 1.0], [-1.0, 0.0], [0.0, -1.0]])
b = np.array([-2.0, 0.0, 0.0])
# 1) solve with SciPy (SLSQP wants fun(z) >= 0, so we pass -g)
res = minimize(f, [0.5, 0.5], method="SLSQP", tol=1e-12,
constraints=[{"type": "ineq", "fun": lambda z: -(A @ z + b)}])
x = res.x
print("solution :", x.round(4), " f =", round(res.fun, 4)) # [1.5 0.5] f = 0.5
# 2) check the four KKT conditions at x
g = A @ x + b
active = np.abs(g) < 1e-6 # which walls touch us?
mu = np.zeros(3)
mu[active] = np.linalg.lstsq(A[active].T, -grad_f(x), rcond=None)[0]
print("g =", g.round(4)) # [ 0. -1.5 -0.5]
print("mu =", mu.round(4)) # [1. 0. 0.]
print("primal feasible :", bool(np.all(g <= 1e-8))) # True
print("stationarity :", bool(np.linalg.norm(grad_f(x) + A.T @ mu) < 1e-6)) # True
print("dual feasible :", bool(np.all(mu >= -1e-8))) # True
print("compl. slackness :", bool(np.allclose(mu * g, 0, atol=1e-8))) # True
# 3) case analysis: try every guess S of the active set
for k in range(4):
for S in combinations(range(3), k):
S = list(S)
if S:
AS, bS = A[S], b[S]
G = AS @ AS.T
if abs(np.linalg.det(G)) < 1e-9:
print(S, "walls do not meet"); continue
m = np.linalg.solve(G, 2 * (AS @ c + bS)) # multipliers
z = c - 0.5 * AS.T @ m # point
else:
m, z = np.array([]), c.copy()
feasible = np.all(np.delete(A @ z + b, S) <= 1e-9)
signs_ok = np.all(m >= -1e-9)
print(S, z.round(3), m.round(3), "KKT point!" if feasible and signs_ok else "rejected")
# [] [2. 1.] [] rejected
# [0] [1.5 0.5] [1.] KKT point!
# [1] [0. 1.] [-4.] rejected
# [2] [2. 0.] [-2.] rejected
# [0, 1] [0. 2.] [-2. -6.] rejected
# [0, 2] [2. 0.] [ 0. -2.] rejected
# [1, 2] [0. 0.] [-4. -2.] rejected
# [0, 1, 2] walls do not meet
1. Which KKT condition expresses "a fence can push you inwards but never pull you out"?
2. At the optimum, a constraint has $g_i(\mathbf{x}^\star) = -2$. What do the KKT conditions say about $\mu_i$?
3. You minimise a convex function with convex inequality constraints and find a point where all four KKT conditions hold. You can conclude that the point is…
4. Minimise $(x-3)^2$ subject to $x \le 1$. Which pair $(x^\star, \mu)$ satisfies the KKT conditions?
5. Minimise $x$ subject to $x^2 \le 0$. The only feasible point is $x = 0$, so it is the minimum. Why does it fail the KKT stationarity test?
6. In a hard-margin SVM, a training point has $\alpha_i \gt 0$. What follows from complementary slackness?
Practice problems
A. Solve $\min x^2$ subject to $x \ge 1$ with the KKT conditions.
Standard form: $g(x) = 1 - x \le 0$. Stationarity: $2x + \mu\cdot(-1) = 0$, so $\mu = 2x$. Slackness: $\mu(1 - x) = 0$.
Case $\mu = 0$: $x = 0$, but $g(0) = 1 \gt 0$ ✗. Case $x = 1$: $\mu = 2 \ge 0$ ✓. So $x^\star = 1$, $\mu = 2$, $f = 1$. The problem is convex, so this is the global minimum.
B. Mixed constraints: $\min x^2 + y^2$ subject to $x + y = 2$ and $y \le 0.5$.
With only the equality, the answer would be $(1,1)$. That has $y = 1 \gt 0.5$, which breaks the wall, so the wall must be active: $y = 0.5$, $x = 1.5$.
Stationarity: $\nabla f = (2x, 2y) = (3, 1)$, $\nabla h = (1,1)$, $\nabla g = (0,1)$. First entry: $3 + \lambda = 0$, so $\lambda = -3$ (a negative equality multiplier is allowed). Second entry: $1 + \lambda + \mu = 0$, so $\mu = 2 \ge 0$ ✓. All four conditions hold and the problem is convex: $\mathbf{x}^\star = (1.5, 0.5)$, $f = 2.5$.
C. Show that $(-1,-1)$ solves $\min x + y$ subject to $x^2 + y^2 \le 2$, and find $\mu$.
Feasible: $1 + 1 - 2 = 0 \le 0$ ✓ (active). $\nabla f = (1,1)$, $\nabla g = (2x, 2y) = (-2,-2)$. Stationarity: $(1,1) + \mu(-2,-2) = \mathbf{0}$ gives $\mu = \tfrac12 \ge 0$ ✓. Slackness holds because $g = 0$. The problem is convex (linear $f$, convex $g$), so $(-1,-1)$ is the global minimum with $f = -2$.
D. For $\min -(x^2 + y^2)$ inside the box $|x| \le 1$, $|y| \le 1$, show that $(1, 0)$ is a KKT point but not a local minimum.
At $(1,0)$ only $g_1 = x - 1$ is active. $\nabla f = (-2x, -2y) = (-2, 0)$ and $\nabla g_1 = (1,0)$, so $-2 + \mu_1 = 0$, $\mu_1 = 2 \ge 0$. All other multipliers are $0$. All four conditions hold.
But moving along the wall to $(1, t)$ gives $f = -1 - t^2 \lt -1 = f(1,0)$: the value goes down on both sides, so $(1,0)$ is not a local minimum (it is a saddle of the constrained problem). The objective is not convex, so KKT is not sufficient. The true minima are the four corners, with $f = -2$.
E. Explain in one paragraph why $\min x$ subject to $x^2 \le 0$ has a minimiser but no KKT multiplier.
The feasible set is the single point $0$, so $x^\star = 0$ is trivially the minimiser. The wall function $g(x) = x^2$ has $g'(0) = 0$, so the wall has "no direction": the stationarity equation $f'(0) + \mu g'(0) = 1 + 0 = 1$ cannot be solved for any $\mu$. A constraint qualification (LICQ needs $\nabla g(0) \ne 0$; Slater needs a point with $x^2 \lt 0$) fails, which is exactly the case where the necessity theorem is silent.
F. Two training points: $\mathbf{x}_1 = (2, 0)$ with $y_1 = +1$ and $\mathbf{x}_2 = (0, 0)$ with $y_2 = -1$. Given that both are support vectors with $\alpha_1 = \alpha_2 = \tfrac12$, find $\mathbf{w}$ and $b$ and check the margins.
$\mathbf{w} = \sum\alpha_iy_i\mathbf{x}_i = \tfrac12(+1)(2,0) + \tfrac12(-1)(0,0) = (1, 0)$. Since $\mathbf{x}_1$ is a support vector, $y_1(\mathbf{w}^\top\mathbf{x}_1 + b) = 1$: $2 + b = 1$, so $b = -1$.
Check $\mathbf{x}_2$: $y_2(\mathbf{w}^\top\mathbf{x}_2 + b) = -(0 - 1) = 1$ ✓. Also $\sum\alpha_iy_i = \tfrac12 - \tfrac12 = 0$ ✓. The street is $|x - 1| \le 1$ between $x = 0$ and $x = 2$, with width $2/\|\mathbf{w}\| = 2$.
Duality
Every constrained problem has a hidden twin: the dual problem. The twin always gives a lower bound on the best possible answer, often matches it exactly, and its variables are prices that say how much each constraint costs you. This is the idea behind the dual form of the SVM, the kernel trick, and the link between "penalty" and "constraint" in regularization.
- Describe the primal problem and its optimal value $p^\star$, and see why every feasible point gives an upper bound
- Build the dual function $d(\boldsymbol\lambda,\boldsymbol\mu) = \min_{\mathbf{x}} L$ and see why it is always concave and a lower bound
- Prove weak duality ($d^\star \le p^\star$) in two lines and measure the duality gap
- Know when strong duality (zero gap) holds (convex problem plus Slater's condition) and see an example where it fails
- Read dual variables as prices (sensitivities) and verify it by perturbing a constraint
- Work three examples: a 1D dual in closed form, an LP and a QP dual, and least squares with a norm constraint
- Derive the SVM dual step by step, see support vectors and why the dual allows kernels
- Explain why people solve the dual, and how regularization $\lambda$ and a norm budget $r$ are linked by a multiplier
Notation (from Chapter 3.7). $\min f(\mathbf{x})$ subject to $g_i(\mathbf{x}) \le 0$ and $h_j(\mathbf{x}) = 0$; multipliers $\lambda_j$ (any sign) and $\mu_i \ge 0$; $L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu) = f(\mathbf{x}) + \sum_j\lambda_jh_j(\mathbf{x}) + \sum_i\mu_ig_i(\mathbf{x})$. You met multipliers in Chapter 3.8 and the KKT conditions in Chapter 3.9. Here we look at the multipliers from a new angle: as the variables of a problem of their own. We write $\min$ for "the smallest value" even when, strictly speaking, it is an infimum (a greatest lower bound that may not be reached).
The primal problem and its optimal value core
The primal problem is just your original problem, with the original unknowns $\mathbf{x}$ (the primal variables). "Primal" is a label, used because a second problem, the dual, is about to appear.
Here is the key picture. Every allowed point $\tilde{\mathbf{x}}$ has some objective value $f(\tilde{\mathbf{x}})$. The best value over all allowed points is $p^\star$. So any allowed point proves "the best value is at most $f(\tilde{\mathbf{x}})$": an upper bound on $p^\star$.
But how would you ever prove that nothing is better? You would need a lower bound, a guarantee that no allowed point goes below some number. Duality is the machine that produces such guarantees.
Three small problems we will use again and again:
- 1D: $\min x^2$ subject to $1 - x \le 0$ (that is, $x \ge 1$). The answer is $x^\star = 1$, so $p^\star = 1$.
- QP (quadratic program: quadratic objective, linear constraints): $\min \tfrac12(x^2 + y^2)$ subject to $2 - x - y \le 0$. The closest point to the origin on or beyond the line $x + y = 2$ is $(1,1)$, so $p^\star = \tfrac12(1+1) = 1$.
- LP (linear program): $\min x + 2y$ subject to $x + y \ge 1$, $x \ge 0$, $y \ge 0$. Try the corners: $(1,0)$ gives $1$, $(0,1)$ gives $2$. So $p^\star = 1$ at $(1,0)$.
In every case a feasible point such as $(2,2)$ in the QP (value $4$) gives an upper bound ($p^\star \le 4$), but nothing yet tells us that $1$ cannot be beaten.
The primal problem and its optimal value: $$p^\star = \min_{\mathbf{x}} f(\mathbf{x}) \quad\text{subject to}\quad g_i(\mathbf{x}) \le 0\ (i=1,\dots,m),\quad h_j(\mathbf{x}) = 0\ (j=1,\dots,p).$$ A point obeying all constraints is feasible; a minimiser is written $\mathbf{x}^\star$. If no point is feasible, $p^\star = +\infty$; if $f$ can be made as small as you like, $p^\star = -\infty$. For every feasible $\tilde{\mathbf{x}}$: $\;p^\star \le f(\tilde{\mathbf{x}})$.
Why do we need it?
A name for "the real problem" and its best value, so that we can compare it with the twin problem we are about to build.
Where is it used?
SVM training in primal form (weights $\mathbf{w}$, bias $b$), linear programming, constrained least squares, and all solvers that move $\mathbf{x}$ directly, such as projected gradient and interior-point methods in the primal.
How is it used?
Any feasible point you find (a heuristic, a rounded solution, a solver iterate) is an upper bound $f(\tilde{\mathbf{x}})$ on $p^\star$. The dual will supply lower bounds; when the two meet, you have proof of optimality.
"Min" may be an infimum. For $\min e^{-x}$ over all $x$ there is no smallest value, only a greatest lower bound ($0$). Mathematicians write $\inf$ for that. In this guide we write $\min$ and keep the subtlety in mind.
The primal is not "the hard one". The primal and dual are two views of the same truth. Sometimes the primal is easier to solve, sometimes the dual is.
Quick check: in the 1D problem $\min x^2$ s.t. $x \ge 1$, $x = 3$ is feasible. What does it tell you about $p^\star$?
$f(3) = 9$, so $p^\star \le 9$. It is a valid but weak upper bound. The point $x = 1.2$ gives $p^\star \le 1.44$, which is better. Only a lower bound can tell us how far from the truth these upper bounds are.
Lagrangian duality: the dual function $d(\boldsymbol\lambda,\boldsymbol\mu)$ core
Imagine the rule "$x \ge 1$" is lifted, and a boss says: you may break the rule, but you pay a fee of $\mu$ for each unit you fall short. Now you can minimise freely, with no constraint at all: just $f$ plus the fee.
- If the fee is tiny, cheating is a bargain. You break the rule, enjoy a low $f$, and the total can go below the true answer $p^\star$.
- If the fee is large, you stay well inside the rules, and the fee term even turns into a reward (you are "paid" for keeping away from the wall). That lowers the total below $p^\star$ again, in a different way.
- At the right fee, the cheat and the honest answer cost the same.
The best total you can get at fee $\mu$ is the dual function $d(\mu)$. It never exceeds $p^\star$ (the next section proves this), so every fee gives a lower bound.
1D problem: $\min x^2$ subject to $1 - x \le 0$. The Lagrangian is $L(x,\mu) = x^2 + \mu(1 - x)$. Fix a fee $\mu \ge 0$ and minimise over all real $x$:
- $\dfrac{\partial L}{\partial x} = 2x - \mu = 0$, so the minimiser is $x(\mu) = \mu/2$.
- Plug it in: $d(\mu) = \dfrac{\mu^2}{4} + \mu\Big(1 - \dfrac{\mu}{2}\Big) = \dfrac{\mu^2}{4} + \mu - \dfrac{\mu^2}{2} = \mu - \dfrac{\mu^2}{4}$.
- A few values: $d(0) = 0$, $d(1) = 0.75$, $d(2) = 1$, $d(3) = 0.75$, $d(4) = 0$.
All of these are $\le p^\star = 1$, and the best one, at $\mu = 2$, equals $1$ exactly. Notice that $d(\mu) = \mu - \mu^2/4$ is an upside-down parabola: concave.
The Lagrangian dual function is the smallest value of the Lagrangian over all $\mathbf{x}$ (no constraints on $\mathbf{x}$ at all): $$d(\boldsymbol\lambda,\boldsymbol\mu) = \min_{\mathbf{x}}\, L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu) = \min_{\mathbf{x}}\Big[f(\mathbf{x}) + \sum_j\lambda_jh_j(\mathbf{x}) + \sum_i\mu_ig_i(\mathbf{x})\Big],\qquad \boldsymbol\mu \ge \mathbf{0}.$$ It may equal $-\infty$ for some multipliers (then that choice is useless). The multipliers are also called the dual variables, and the unconstrained $\mathbf{x}$ in the minimisation the primal variables.
$d$ is always concave, even if $f$, $g_i$, $h_j$ are not convex. Reason: for a fixed $\mathbf{x}$, the value $L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu)$ is a straight (affine) function of $(\boldsymbol\lambda,\boldsymbol\mu)$. The function $d$ is the pointwise minimum of all these straight lines/planes, one per $\mathbf{x}$. A minimum of affine functions is concave. In symbols, for $0\le\theta\le1$: $d(\theta a + (1-\theta)b) = \min_{\mathbf{x}}[\theta L(\mathbf{x},a) + (1-\theta)L(\mathbf{x},b)] \ge \theta\,d(a) + (1-\theta)\,d(b)$.
Why do we need it?
It turns a constrained problem into a family of unconstrained ones, indexed by the fees. And it produces the lower bounds we were missing, all of which are cheap to compute by an unconstrained minimisation.
Where is it used?
The SVM dual objective is $d(\boldsymbol\alpha)$ for the margin constraints. Dual decomposition (splitting a big problem into pieces coordinated by prices), Lagrangian relaxation bounds in integer programming, and the "dual residual" of ADMM all build on it.
How is it used?
Write the Lagrangian, minimise it over $\mathbf{x}$ (set the gradient to zero when it is smooth), and substitute the minimiser back. What remains is a function of the multipliers only.
Multipliers of inequalities must be $\ge 0$. The dual function is defined for $\boldsymbol\mu \ge \mathbf{0}$. With a negative $\mu$ the "fee" would become a reward for breaking the rule, and the lower-bound guarantee would be lost.
You minimise over all $\mathbf{x}$. In $d(\boldsymbol\mu)$ the minimisation ignores the constraints, so the minimiser can be an infeasible point. That is the whole trick, and also why $d$ is only a bound.
Quick check: for $\min x^2$ s.t. $1 - x \le 0$, compute $d(3)$ directly from the definition.
$L(x,3) = x^2 + 3(1 - x) = x^2 - 3x + 3$. It is smallest at $x = 1.5$, where $L = 2.25 - 4.5 + 3 = 0.75$. So $d(3) = 0.75$, matching $\mu - \mu^2/4 = 3 - 2.25 = 0.75$ ✓.
Weak duality: the dual is always a lower bound core
Take any feasible point $\tilde{\mathbf{x}}$ and any non-negative fees. For a feasible point the fee term cannot hurt: you are inside the rules, so every $g_i \le 0$ and the fee $\mu_ig_i$ is zero or negative (you are even "paid" for staying away from the wall). So the Lagrangian at $\tilde{\mathbf{x}}$ is at most $f(\tilde{\mathbf{x}})$.
The dual function takes the smallest value of the Lagrangian over all $\mathbf{x}$, so it is below the Lagrangian at $\tilde{\mathbf{x}}$, which is below $f(\tilde{\mathbf{x}})$. A chain of three numbers, in this order:
$d(\boldsymbol\lambda,\boldsymbol\mu) \;\le\; L(\tilde{\mathbf{x}},\boldsymbol\lambda,\boldsymbol\mu) \;\le\; f(\tilde{\mathbf{x}})$.
1D problem again, $f = x^2$, $g = 1 - x$. Take the feasible point $\tilde{x} = 1.5$ and the fee $\mu = 1$.
- $f(\tilde x) = 1.5^2 = 2.25$.
- $L(\tilde x, 1) = 2.25 + 1\cdot(1 - 1.5) = 2.25 - 0.5 = 1.75$. Below $f(\tilde x)$, because $g(\tilde x) = -0.5 \lt 0$.
- $d(1) = 1 - \tfrac14 = 0.75$. Below $L(\tilde x,1)$, because $d$ is the minimum of $L$ over all $x$.
The chain reads $0.75 \le 1.75 \le 2.25$ ✓, and $p^\star = 1$ sits between $0.75$ and $2.25$.
Weak duality. For every $\boldsymbol\lambda$ and every $\boldsymbol\mu \ge \mathbf{0}$, $$d(\boldsymbol\lambda,\boldsymbol\mu) \le p^\star.$$ Proof in two lines. Let $\tilde{\mathbf{x}}$ be feasible. Then $h_j(\tilde{\mathbf{x}}) = 0$, $\mu_i \ge 0$ and $g_i(\tilde{\mathbf{x}}) \le 0$, so $$d(\boldsymbol\lambda,\boldsymbol\mu) \;=\; \min_{\mathbf{x}}L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu) \;\le\; L(\tilde{\mathbf{x}},\boldsymbol\lambda,\boldsymbol\mu) \;=\; f(\tilde{\mathbf{x}}) + \underbrace{\textstyle\sum_j\lambda_jh_j(\tilde{\mathbf{x}})}_{=\,0} + \underbrace{\textstyle\sum_i\mu_ig_i(\tilde{\mathbf{x}})}_{\le\,0} \;\le\; f(\tilde{\mathbf{x}}).$$ This holds for every feasible $\tilde{\mathbf{x}}$, in particular for the best one, so $d \le p^\star$. ∎
Taking the best multipliers gives the dual optimal value $d^\star = \max_{\boldsymbol\lambda,\ \boldsymbol\mu\ge 0} d(\boldsymbol\lambda,\boldsymbol\mu)$, and still $d^\star \le p^\star$. The difference $$p^\star - d^\star \;\ge\; 0$$ is the duality gap. Given any feasible $\tilde{\mathbf{x}}$ and any $\boldsymbol\mu \ge 0$, the gap $f(\tilde{\mathbf{x}}) - d(\boldsymbol\lambda,\boldsymbol\mu) \ge 0$ is a number you can compute, and it is at least as large as how far $f(\tilde{\mathbf{x}})$ is from optimal.
Why do we need it?
It is the guarantee that makes the dual useful: whatever fees you try, you get a certified lower bound on the best possible answer, without ever solving the hard constrained problem.
Where is it used?
Branch-and-bound for integer programs (bounds prune the search), stopping rules of interior-point and SVM solvers (they stop when the primal-dual gap is below a tolerance), and certificates of near-optimality in convex optimisation libraries such as CVXPY.
How is it used?
Keep a feasible point (upper bound) and a dual point (lower bound). The true optimum is trapped between them, and the gap tells you how much you could still improve.
A small gap is not proof of a small error in $\mathbf{x}$. It says the value is within the gap of optimal. Two points can have the same value and sit far apart.
Weak duality needs no assumptions. It holds for every problem: non-convex, non-smooth, even with integer variables. Only the "strong" version, next, needs conditions.
Quick check: for the LP $\min x + 2y$ (s.t. $x + y \ge 1$, $x, y \ge 0$) someone shows you a fee vector with $d = 1.2$. What can you say about $p^\star$, and is that possible?
Weak duality says $d \le p^\star$, so $p^\star \ge 1.2$. But we computed $p^\star = 1$ for this LP, so a dual value of $1.2$ is impossible: the claim must contain an error (for instance a negative $\mu$, or an arithmetic mistake). Any valid dual value is $\le 1$.
The dual problem: find the best lower bound core
Every choice of fees gives a lower bound $d(\boldsymbol\lambda,\boldsymbol\mu)$ on $p^\star$. The dual problem is the obvious next question: which fees give the best (highest) lower bound?
There is a lovely bonus. The dual function is always concave, so "maximise a concave function over $\boldsymbol\mu \ge 0$" is an easy, convex problem, even when the original primal is hard and non-convex. So the dual always hands you something solvable, and a guaranteed bound.
QP. $\min \tfrac12(x^2 + y^2)$ s.t. $2 - x - y \le 0$, with Lagrangian $L = \tfrac12(x^2+y^2) + \mu(2 - x - y)$.
- Minimise over $x, y$: $\partial_x L = x - \mu = 0$ and $\partial_y L = y - \mu = 0$, so $x = y = \mu$.
- Substitute: $d(\mu) = \tfrac12(\mu^2 + \mu^2) + \mu(2 - 2\mu) = \mu^2 + 2\mu - 2\mu^2 = 2\mu - \mu^2$.
- Maximise over $\mu \ge 0$: $d'(\mu) = 2 - 2\mu = 0$ gives $\mu^\star = 1$ and $d^\star = 2 - 1 = 1$.
- Compare with the primal: $p^\star = 1$. The bounds meet. And the point that minimises $L$ at the best fee, $x = y = \mu^\star = 1$, is the primal solution $(1,1)$.
LP. $\min x + 2y$ s.t. $1 - x - y \le 0$, $-x \le 0$, $-y \le 0$, with multipliers $\mu_1, \mu_2, \mu_3 \ge 0$:
- $L = x + 2y + \mu_1(1 - x - y) - \mu_2x - \mu_3y = \mu_1 + (1 - \mu_1 - \mu_2)\,x + (2 - \mu_1 - \mu_3)\,y$.
- $L$ is linear in $x, y$. Its minimum over all real $x, y$ is $-\infty$, unless both brackets are zero. So $d = \mu_1$ when $\mu_2 = 1 - \mu_1 \ge 0$ and $\mu_3 = 2 - \mu_1 \ge 0$, and $d = -\infty$ otherwise.
- So the dual problem is: maximise $\mu_1$ subject to $0 \le \mu_1 \le 1$. The answer is $\mu_1 = 1$, $d^\star = 1 = p^\star$ ✓.
The dual problem of the primal problem is $$d^\star = \max_{\boldsymbol\lambda,\ \boldsymbol\mu \ge \mathbf{0}}\ d(\boldsymbol\lambda,\boldsymbol\mu).$$ A pair $(\boldsymbol\lambda,\boldsymbol\mu)$ with $\boldsymbol\mu \ge \mathbf{0}$ and $d \gt -\infty$ is dual feasible (the set where $d$ is finite is the "implicit constraints" of the dual, such as $\mu_1 \le 1$ above). The dual is always a convex problem (concave maximisation).
Two classical recipes.
- Linear program: primal $\min \mathbf{c}^\top\mathbf{x}$ s.t. $A\mathbf{x} \ge \mathbf{b}$, $\mathbf{x} \ge \mathbf{0}$ has dual $\max \mathbf{b}^\top\mathbf{y}$ s.t. $A^\top\mathbf{y} \le \mathbf{c}$, $\mathbf{y} \ge \mathbf{0}$. (Our LP: $A = [1\ 1]$, $b = 1$, $\mathbf{c} = (1,2)$.)
- Quadratic program: $\min \tfrac12\|\mathbf{x}\|^2$ s.t. $A\mathbf{x} \le \mathbf{b}$ has $\mathbf{x}(\boldsymbol\mu) = -A^\top\boldsymbol\mu$ and $d(\boldsymbol\mu) = -\tfrac12\|A^\top\boldsymbol\mu\|^2 - \mathbf{b}^\top\boldsymbol\mu$, to be maximised over $\boldsymbol\mu \ge 0$. (Our QP: $A = [-1\ \ -1]$, $b = -2$, giving $2\mu - \mu^2$.)
Why do we need it?
It gives the best possible certified lower bound, and often a second, easier route to the answer: solve the dual, then recover the primal solution from the multipliers.
Where is it used?
LP duality (max-flow equals min-cut, two-player zero-sum games), the SVM dual, kernel methods, and dual coordinate ascent solvers for linear SVM and logistic regression (such as LIBLINEAR).
How is it used?
Write $L$, minimise over $\mathbf{x}$ in closed form, substitute to get $d$, then maximise $d$ over $\boldsymbol\mu \ge 0$. Recover $\mathbf{x}^\star = \arg\min_{\mathbf{x}} L(\mathbf{x},\boldsymbol\lambda^\star,\boldsymbol\mu^\star)$.
The dual of a maximisation, or of a $\ge$ constraint, looks different. The recipes above assume the standard forms. If your problem is "maximise" or uses $\ge$, convert it first (as in Chapter 3.9) before applying the formulas, or you will get signs wrong.
Dual infeasible is information. If the primal is unbounded below ($p^\star = -\infty$), weak duality forbids any finite lower bound, so the dual is infeasible ($d = -\infty$ for every allowed $\boldsymbol\mu$). Conversely, for a linear program, an infeasible dual means the primal is either infeasible or unbounded.
Quick check: for $\min \tfrac12\|\mathbf{x}\|^2$ with the single wall $x \ge 3$ (in one variable), find $d(\mu)$ and $d^\star$.
$L = \tfrac12x^2 + \mu(3 - x)$. Minimising: $x = \mu$, so $d(\mu) = \tfrac12\mu^2 + \mu(3 - \mu) = 3\mu - \tfrac12\mu^2$. Maximise: $3 - \mu = 0$, $\mu^\star = 3$, $d^\star = 9 - 4.5 = 4.5$. Primal: $x^\star = 3$, $p^\star = 4.5$ ✓.
Strong duality: when the gap is zero core
Weak duality gives a sandwich: $d^\star \le p^\star$. Sometimes the sandwich closes completely: the best lower bound equals the best upper bound. That is strong duality. It means you can certify optimality: exhibit a feasible $\mathbf{x}$ and a dual $\boldsymbol\mu$ with the same value.
Why would it close? Think of the dual as pushing a straight line up from below until it touches the "cost curve" of the problem. If the curve is a smooth bowl (convex), a line can reach the bottom exactly. If the curve has a dent (non-convex), the line gets stuck on the rim and cannot reach down into it: a gap remains.
Gap zero: the 1D convex problem $\min x^2$ s.t. $x \ge 1$. We found $d(\mu) = \mu - \mu^2/4$, with $d^\star = 1 = p^\star$.
Gap positive: a non-convex problem. Minimise $-x^2$ over $0 \le x \le 2$ subject to $x - 1 \le 0$ (the allowed interval $[0,1]$; we keep the simple bounds $0 \le x \le 2$ as "where $x$ may roam" and do not give them multipliers).
- Primal: $-x^2$ is smallest at the far end $x = 1$, so $p^\star = -1$.
- Lagrangian: $L(x,\mu) = -x^2 + \mu(x - 1)$. It is concave in $x$ (a hill), so its minimum over $[0,2]$ is at an end: $L(0) = -\mu$ and $L(2) = -4 + \mu$. Hence $d(\mu) = \min(-\mu,\ \mu - 4)$.
- The two pieces cross at $-\mu = \mu - 4$, i.e. $\mu = 2$. So $d^\star = d(2) = -2$.
- Gap: $p^\star - d^\star = -1 - (-2) = 1 \gt 0$. Weak duality holds ($-2 \le -1$) but the bound is not tight.
Strong duality means $d^\star = p^\star$ (zero duality gap).
When it holds. If the primal problem is convex ($f$ and $g_i$ convex, $h_j$ affine) and Slater's condition holds (some point is strictly feasible: $g_i \lt 0$ for all $i$, $h_j = 0$; for affine $g_i$ a non-strict $\le$ is enough), then strong duality holds, and (when $p^\star$ is finite) the dual optimum is attained. In particular it holds for every feasible and bounded LP and for convex QPs with linear constraints. It can fail for non-convex problems (example above), and, more rarely, for convex ones that violate Slater.
Link to KKT (Chapter 3.9). Suppose $d^\star = p^\star$, with $\mathbf{x}^\star$ primal optimal and $(\boldsymbol\lambda^\star,\boldsymbol\mu^\star)$ dual optimal. Then $$f(\mathbf{x}^\star) = d(\boldsymbol\lambda^\star,\boldsymbol\mu^\star) = \min_{\mathbf{x}} L(\mathbf{x},\boldsymbol\lambda^\star,\boldsymbol\mu^\star) \le L(\mathbf{x}^\star,\boldsymbol\lambda^\star,\boldsymbol\mu^\star) = f(\mathbf{x}^\star) + \underbrace{\textstyle\sum\lambda_j^\star h_j(\mathbf{x}^\star)}_{0} + \textstyle\sum\mu_i^\star g_i(\mathbf{x}^\star) \le f(\mathbf{x}^\star).$$ The two ends are equal, so every "$\le$" is an equality. That says: (a) $\mathbf{x}^\star$ minimises $L(\cdot,\boldsymbol\lambda^\star,\boldsymbol\mu^\star)$, which is stationarity; (b) $\sum\mu_i^\star g_i(\mathbf{x}^\star) = 0$ with every term $\le 0$, so each term is $0$: complementary slackness. Feasibility and $\boldsymbol\mu^\star \ge 0$ hold by construction. So strong duality gives KKT, and for convex problems KKT gives strong duality: they are the same fact seen two ways.
Equivalent picture: $(\mathbf{x}^\star,\boldsymbol\mu^\star)$ is a saddle point of $L$: $\;L(\mathbf{x}^\star,\boldsymbol\mu) \le L(\mathbf{x}^\star,\boldsymbol\mu^\star) \le L(\mathbf{x},\boldsymbol\mu^\star)$ for all $\mathbf{x}$ and all $\boldsymbol\mu \ge 0$: lowest along the $\mathbf{x}$ direction, highest along the $\boldsymbol\mu$ direction.
Why do we need it?
Strong duality is what lets you replace a hard primal problem by its dual without losing anything, and gives a proof of optimality by a matching pair of numbers.
Where is it used?
Solving SVMs through the dual (the dual optimum equals the primal optimum), LP solvers that stop when primal and dual values agree, primal-dual interior-point methods, and convergence certificates in CVXPY, ECOS and OSQP.
How is it used?
Check that the problem is convex and that some point has all inequalities strict (Slater). If so, solve whichever of primal or dual is easier, and trust that the optimal values agree. A reported gap that is not near $0$ means the solver has not converged (or the problem is not convex).
Zero gap does not need a unique solution. Many primal or dual optima can exist; strong duality only says their values agree.
Slater can fail even for convex problems. $\min x$ subject to $x^2 \le 0$: $p^\star = 0$ (only $x = 0$ is feasible), and $L = x + \mu x^2$ has $d(\mu) = -1/(4\mu)$. The values creep up to $0 = p^\star$ as $\mu \to \infty$, but no finite $\mu$ attains it: the dual optimum does not exist. This is the same situation as the missing KKT multiplier in Chapter 3.9.
Quick check: a problem has $p^\star = 5$. Someone gives you a dual point with $d = 5$ and a feasible point with $f = 5$. What have you proved?
Both are optimal. By weak duality $d \le d^\star \le p^\star \le f$, so $5 \le d^\star \le p^\star \le 5$: all four are equal. The pair is a certificate of optimality, and the duality gap is $0$.
Dual variables and what they mean: prices (sensitivities) core
A bakery has a limited number of oven hours. That limit is a constraint. If you could rent one more hour, how much extra profit would you make? That number is the shadow price of oven time. It is the most you should be willing to pay for another hour.
The multiplier of a constraint is exactly that price. A resource that you do not fully use (a constraint that is not touching) has price zero, since more of it would change nothing. That is complementary slackness again, now with a money meaning.
1D. $\min x^2$ s.t. $1 - x \le u$, that is $x \ge 1 - u$. Relaxing the wall by $u$ changes the best value to $p(u) = (1 - u)^2$ (for $u \le 1$). Its slope at $u = 0$: $p'(0) = -2(1 - 0) = -2$. And the dual optimum was $\mu^\star = 2$. So $\mu^\star = -p'(0)$ ✓.
Numerical check. Relax by $u = 0.01$: $p(0.01) = 0.99^2 = 0.9801$, so $\dfrac{p(0.01) - p(0)}{0.01} = \dfrac{0.9801 - 1}{0.01} = -1.99 \approx -2$ ✓.
A bakery (maximise profit, so the prices are rates of profit gain). Cakes $x$ and loaves $y$. Profit $2x + 3y$. Oven: $x + y \le 4$. Flour: $x + 3y \le 6$. The best plan is $(x, y) = (3, 1)$ with profit $9$, and both resources are used up. Solving the dual gives prices $1.5$ per oven unit and $0.5$ per flour unit. Check: $4(1.5) + 6(0.5) = 9$ ✓ (strong duality). Now give the bakery one more oven unit ($4 \to 5$): the new plan is $(4.5, 0.5)$ with profit $2(4.5) + 3(0.5) = 10.5$. Gain: $1.5$ ✓, exactly the price.
Sensitivity (shadow-price) theorem. Let $p(\mathbf{u})$ be the optimal value when each inequality is relaxed: $g_i(\mathbf{x}) \le u_i$ (and an equality becomes $h_j(\mathbf{x}) = v_j$). If strong duality holds at $\mathbf{u} = \mathbf{0}$ with dual optimum $(\boldsymbol\lambda^\star,\boldsymbol\mu^\star)$, then for every $\mathbf{u}$, $\mathbf{v}$: $$p(\mathbf{u},\mathbf{v}) \;\ge\; p^\star - \boldsymbol\mu^{\star\top}\mathbf{u} - \boldsymbol\lambda^{\star\top}\mathbf{v}.$$ Proof. Take any $\mathbf{x}$ feasible for the relaxed problem: $g_i(\mathbf{x}) \le u_i$, $h_j(\mathbf{x}) = v_j$. Then $$f(\mathbf{x}) \ge f(\mathbf{x}) + \textstyle\sum\mu_i^\star\big(g_i(\mathbf{x}) - u_i\big) + \sum\lambda_j^\star\big(h_j(\mathbf{x}) - v_j\big) = L(\mathbf{x},\boldsymbol\lambda^\star,\boldsymbol\mu^\star) - \boldsymbol\mu^{\star\top}\mathbf{u} - \boldsymbol\lambda^{\star\top}\mathbf{v} \ge d^\star - \boldsymbol\mu^{\star\top}\mathbf{u} - \boldsymbol\lambda^{\star\top}\mathbf{v}.$$ The first step uses $\mu_i^\star \ge 0$ and $g_i - u_i \le 0$; the last uses the definition of $d$. Since $d^\star = p^\star$, minimising over $\mathbf{x}$ gives the claim. ∎
Meaning. The straight line $p^\star - \mu^\star u$ lies below the true curve $p(u)$ and touches it at $u = 0$. If $p$ is differentiable at $0$, that line is the tangent, so $$\mu_i^\star = -\frac{\partial p}{\partial u_i}(\mathbf{0}),\qquad \lambda_j^\star = -\frac{\partial p}{\partial v_j}(\mathbf{0}).$$ The multiplier is the rate at which the optimal value improves when you relax that constraint. Units: (objective units) per (constraint units). A touching constraint with $\mu_i^\star = 0$ is one you can tighten slightly at no cost (to first order). If two choices of prices both work (a kink in $p$) the price is not unique.
Why do we need it?
It tells you which constraint is worth relaxing and by how much it pays off, without re-solving the problem for every change. It turns a multiplier from an abstract number into a decision tool.
Where is it used?
Sensitivity reports of LP solvers (shadow prices), resource allocation and pricing in economics and operations research, SVMs ($\alpha_i$ is the price of the margin requirement of example $i$), and regularization, where $\lambda$ is the price of the norm budget (next sections).
How is it used?
Solve once, read the multipliers, and predict the effect of a small change $\Delta$ in a constraint as about $-\mu^\star\Delta$ on the optimal value. For big changes, re-solve: the true curve $p(u)$ bends, and the line is only a bound.
Prices are local. $\mu^\star$ is the slope at the current constraint level. Relax by a lot and the set of touching walls changes, so the price changes too. In the bakery the oven price is $1.5$ while both walls touch, it jumps to $3$ once flour is so plentiful that it stops touching (flour limit at least $3\times$ the oven limit), and it falls to $0$ once the oven stops touching (flour limit below the oven limit). Prices follow which walls are touching.
Sign. With the standard form $g \le u$, relaxing ($u \gt 0$) can only help a minimiser, so $p$ decreases and $\mu^\star = -p'(0) \ge 0$. If you use "maximise profit", the price is the (positive) rate of profit increase, which is the same number.
Quick check: at the optimum of a problem, $\mu_1^\star = 4$ for the constraint $g_1 \le 0$. Roughly how much does the optimal value drop if you relax it to $g_1 \le 0.05$?
About $\mu_1^\star \times 0.05 = 4 \times 0.05 = 0.2$. The optimal value falls by roughly $0.2$ (and by at most that much, by the theorem: $p(u) \ge p^\star - 0.2$).
Worked example: least squares with a norm constraint
You want weights $\mathbf{w}$ that fit the data well, but you refuse to let them grow beyond a budget: $\|\mathbf{w}\| \le r$. Picture the loss as a set of nested ovals around the best unconstrained fit, and the budget as a disc around the origin.
- If the best fit already lies inside the disc, the budget is irrelevant (multiplier $0$).
- If not, the answer is the point where the smallest oval that still touches the disc just kisses its edge. The multiplier $\mu$ measures how hard the disc pushes the answer back towards the origin.
Simplest case: $A = I$ and $\mathbf{y} = (3, 4)$ (so $\|\mathbf{y}\| = 5$): $\min \|\mathbf{w} - \mathbf{y}\|^2$ subject to $\|\mathbf{w}\|^2 \le r^2$, with the budget $r = 2$.
- Lagrangian: $L = \|\mathbf{w} - \mathbf{y}\|^2 + \mu(\|\mathbf{w}\|^2 - r^2)$.
- Minimise over $\mathbf{w}$: $2(\mathbf{w} - \mathbf{y}) + 2\mu\mathbf{w} = \mathbf{0}$, so $\mathbf{w}(\mu) = \dfrac{\mathbf{y}}{1 + \mu}$.
- Substitute: $\|\mathbf{w} - \mathbf{y}\|^2 = \Big(\dfrac{\mu}{1+\mu}\Big)^2\cdot 25$ and $\mu\|\mathbf{w}\|^2 = \dfrac{25\mu}{(1+\mu)^2}$. Their sum is $\dfrac{25(\mu^2 + \mu)}{(1+\mu)^2} = \dfrac{25\mu}{1+\mu}$. So $$d(\mu) = \frac{25\mu}{1 + \mu} - \mu r^2 = \frac{25\mu}{1+\mu} - 4\mu.$$
- Maximise: $d'(\mu) = \dfrac{25}{(1+\mu)^2} - 4 = 0$, so $1 + \mu = 2.5$ and $\mu^\star = 1.5$. Then $d^\star = \dfrac{25(1.5)}{2.5} - 4(1.5) = 15 - 6 = 9$.
- Recover the primal point: $\mathbf{w}^\star = \mathbf{y}/(1 + \mu^\star) = (1.2,\ 1.6)$, with $\|\mathbf{w}^\star\| = 2 = r$ (the budget is fully used). Its loss is $(3 - 1.2)^2 + (4 - 1.6)^2 = 3.24 + 5.76 = 9$ ✓. Primal $=$ dual: zero gap.
For $\min \|A\mathbf{w} - \mathbf{y}\|^2$ s.t. $\|\mathbf{w}\|^2 \le r^2$ (one constraint $g = \|\mathbf{w}\|^2 - r^2 \le 0$): $$\mathbf{w}(\mu) = (A^\top A + \mu I)^{-1}A^\top\mathbf{y},\qquad d(\mu) = \|A\mathbf{w}(\mu) - \mathbf{y}\|^2 + \mu\|\mathbf{w}(\mu)\|^2 - \mu r^2.$$ Slope of the dual. $d'(\mu) = \|\mathbf{w}(\mu)\|^2 - r^2$: the derivative of the dual function equals the constraint value at the Lagrangian minimiser (this holds in general: $\nabla d = $ the constraint values at $\mathbf{x}(\boldsymbol\mu)$). So $d' = 0$ exactly when the budget is used up: complementary slackness in disguise.
With the SVD $A = U\Sigma V^\top$ and $\mathbf{z} = U^\top\mathbf{y}$, $\|\mathbf{w}(\mu)\|^2 = \sum_i \dfrac{\sigma_i^2z_i^2}{(\sigma_i^2 + \mu)^2}$, which decreases in $\mu$. So there is a unique $\mu^\star$ with $\|\mathbf{w}(\mu^\star)\| = r$ whenever the unconstrained fit $\mathbf{w}_{\text{LS}}$ is longer than $r$; otherwise $\mu^\star = 0$. Slater holds ($\mathbf{w} = \mathbf{0}$ is strictly feasible for $r \gt 0$) and the problem is convex, so strong duality holds.
Why do we need it?
It shows the whole duality machine on a problem with real data: a constrained problem in many unknowns becomes a one-dimensional concave problem in the single multiplier $\mu$.
Where is it used?
Ridge regression (the same $\mu$ is the penalty weight, see the regularization section below), trust-region subproblems inside some second-order optimizers, and Tikhonov regularisation of inverse problems.
How is it used?
Find $\mu^\star$ by solving $\|\mathbf{w}(\mu)\| = r$ in one dimension (bisection or Newton on the decreasing function $\|\mathbf{w}(\mu)\|$), then compute $\mathbf{w}^\star = \mathbf{w}(\mu^\star)$ with one linear solve.
The fee scale depends on the loss scale. If you write the loss as $\tfrac12\|A\mathbf{w} - \mathbf{y}\|^2$ instead, the multiplier is divided by two. Always say which form of the loss you used when you quote a $\mu$ or a $\lambda$.
An inactive budget gives $\mu = 0$. When $\|\mathbf{w}_{\text{LS}}\| \le r$ the answer is the plain least-squares fit, with a zero price on the budget.
Quick check: in the example with $A = I$, $\mathbf{y} = (3,4)$, what are $\mu^\star$ and $\mathbf{w}^\star$ if the budget is $r = 5$? And if $r = 10$?
$\mu^\star = 5/r - 1 = 0$ for $r = 5$: the budget just touches the unconstrained fit $\mathbf{y} = (3,4)$ (length $5$), with price $0$. For $r = 10$ the formula would give a negative number, so the constraint is inactive: $\mu^\star = 0$ and $\mathbf{w}^\star = \mathbf{y}$.
ML connection: the SVM dual, support vectors and kernels core
In Chapter 3.9 the SVM gave us the optimality conditions $\mathbf{w} = \sum\alpha_iy_i\mathbf{x}_i$ and $\sum\alpha_iy_i = 0$. Now we substitute them back into the Lagrangian. All of $\mathbf{w}$ and $b$ disappear, and we are left with a problem in the multipliers $\alpha_i$ alone, one number per training example.
The remarkable part: in that problem the data enter only through the dot products $\mathbf{x}_i^\top\mathbf{x}_j$. Replace each dot product by another similarity function (a kernel) and the SVM draws curved boundaries, without ever writing down the curved features.
Derivation, step by step (hard margin). Primal: $\min \tfrac12\|\mathbf{w}\|^2$ s.t. $1 - y_i(\mathbf{w}^\top\mathbf{x}_i + b) \le 0$ for $i = 1,\dots,n$.
- Lagrangian with $\alpha_i \ge 0$: $L = \tfrac12\|\mathbf{w}\|^2 + \sum_i\alpha_i - \sum_i\alpha_iy_i\mathbf{w}^\top\mathbf{x}_i - b\sum_i\alpha_iy_i$.
- Minimise over $\mathbf{w}$: $\nabla_{\mathbf{w}}L = \mathbf{w} - \sum_i\alpha_iy_i\mathbf{x}_i = \mathbf{0}$, so $\mathbf{w} = \sum_i\alpha_iy_i\mathbf{x}_i$.
- Minimise over $b$: $\partial L/\partial b = -\sum_i\alpha_iy_i$. This does not depend on $b$ except through a linear term, so $L \to -\infty$ unless $\sum_i\alpha_iy_i = 0$. That is an implicit dual constraint (the same phenomenon as in the LP example).
- Substitute $\mathbf{w}$ (and $\sum\alpha_iy_i = 0$). The two $\mathbf{w}$-terms become $\tfrac12\|\mathbf{w}\|^2 - \mathbf{w}^\top\sum_i\alpha_iy_i\mathbf{x}_i = \tfrac12\|\mathbf{w}\|^2 - \|\mathbf{w}\|^2 = -\tfrac12\|\mathbf{w}\|^2$, and the $b$-term vanishes. So $$d(\boldsymbol\alpha) = \sum_i\alpha_i - \tfrac12\|\mathbf{w}\|^2 = \sum_i\alpha_i - \tfrac12\sum_{i,j}\alpha_i\alpha_jy_iy_j\,\mathbf{x}_i^\top\mathbf{x}_j.$$
Numerical check on a tiny dataset. $\mathbf{x}_1 = (1,1)$, $\mathbf{x}_2 = (3,2)$ with $y = +1$; $\mathbf{x}_3 = (-1,-1)$, $\mathbf{x}_4 = (-2,-3)$ with $y = -1$. Take $\boldsymbol\alpha = (\tfrac14, 0, \tfrac14, 0)$ ($\sum\alpha_iy_i = \tfrac14 - \tfrac14 = 0$ ✓). The relevant entries of $Q_{ij} = y_iy_j\mathbf{x}_i^\top\mathbf{x}_j$ are $Q_{11} = 2$, $Q_{33} = 2$, $Q_{13} = (+1)(-1)(-2) = 2$. Then $\boldsymbol\alpha^\top Q\boldsymbol\alpha = \tfrac1{16}(2 + 2 + 2\cdot2) = 0.5$ and $d = \sum\alpha - \tfrac12(0.5) = 0.5 - 0.25 = \mathbf{0.25}$. The primal: $\mathbf{w} = \tfrac14(1,1) + \tfrac14(1,1) = (0.5, 0.5)$, so $\tfrac12\|\mathbf{w}\|^2 = \tfrac12(0.5) = \mathbf{0.25}$ ✓. Dual value $=$ primal value: strong duality, as promised for a convex QP with linear constraints.
Kernels, with numbers. The function $k(\mathbf{x},\mathbf{z}) = (\mathbf{x}^\top\mathbf{z})^2$ in 2D equals the dot product $\varphi(\mathbf{x})^\top\varphi(\mathbf{z})$ of the feature maps $\varphi(\mathbf{x}) = (x_1^2,\ \sqrt2x_1x_2,\ x_2^2)$. Check with $\mathbf{x} = (1,2)$, $\mathbf{z} = (3,1)$: $(\mathbf{x}^\top\mathbf{z})^2 = 5^2 = 25$, and $\varphi(\mathbf{x}) = (1, 2\sqrt2, 4)$, $\varphi(\mathbf{z}) = (9, 3\sqrt2, 1)$ give $9 + 12 + 4 = 25$ ✓. The kernel computes the 3D dot product directly from the 2D points.
SVM dual (hard margin). With $Q_{ij} = y_iy_j\,\mathbf{x}_i^\top\mathbf{x}_j$: $$\max_{\boldsymbol\alpha}\ \sum_i\alpha_i - \tfrac12\boldsymbol\alpha^\top Q\boldsymbol\alpha \quad\text{s.t.}\quad \alpha_i \ge 0,\quad \sum_i\alpha_iy_i = 0.$$ Soft margin. Allow slack $\xi_i \ge 0$: $\min\ \tfrac12\|\mathbf{w}\|^2 + C\sum_i\xi_i$ s.t. $y_i(\mathbf{w}^\top\mathbf{x}_i + b) \ge 1 - \xi_i$, $\xi_i \ge 0$. The multiplier $\beta_i$ of $\xi_i \ge 0$ gives $\partial L/\partial\xi_i = C - \alpha_i - \beta_i = 0$, and $\beta_i \ge 0$ forces $\alpha_i \le C$. The dual is the same, with the box $0 \le \alpha_i \le C$.
Reading the solution (complementary slackness, Chapter 3.9). With margin $m_i = y_i(\mathbf{w}^\top\mathbf{x}_i + b)$:
- $\alpha_i = 0$: the example is safely outside the street ($m_i \ge 1$). Not a support vector.
- $0 \lt \alpha_i \lt C$: exactly on the edge of the street ($m_i = 1$). A support vector.
- $\alpha_i = C$ (soft margin only): inside the street or on the wrong side ($m_i \le 1$, with $\xi_i = 1 - m_i \ge 0$). Also a support vector.
Prediction and kernels. $\mathbf{w} = \sum_i\alpha_iy_i\mathbf{x}_i$, so the decision value is $f(\mathbf{x}) = \sum_i\alpha_iy_i\,\mathbf{x}_i^\top\mathbf{x} + b$, a sum over support vectors only. Replace $\mathbf{x}_i^\top\mathbf{x}_j$ by a kernel $k(\mathbf{x}_i,\mathbf{x}_j)$ (common choices: polynomial $(\mathbf{x}^\top\mathbf{z} + 1)^2$ and Gaussian/RBF $e^{-\gamma\|\mathbf{x}-\mathbf{z}\|^2}$) and $f(\mathbf{x}) = \sum_i\alpha_iy_i\,k(\mathbf{x}_i,\mathbf{x}) + b$. The primal cannot do this, because it needs $\mathbf{w}$ explicitly, which may live in an enormous (even infinite-dimensional) feature space.
Why do we need it?
The dual removes $\mathbf{w}$, leaves simple constraints ($0 \le \alpha_i \le C$, one equality), exposes the support vectors, and makes the data appear only through dot products, which unlocks kernels and non-linear boundaries.
Where is it used?
LIBSVM and scikit-learn's SVC (SMO solves this dual), kernel SVMs for text, images and bioinformatics, support vector regression, and the dual coordinate-descent solver in LIBLINEAR.
How is it used?
Choose a kernel and $C$, solve the dual for $\boldsymbol\alpha$ (a QP with one equality and box constraints), keep the examples with $\alpha_i \gt 0$, get $b$ from a point with $0 \lt \alpha_i \lt C$, and predict with the sum over support vectors.
$b$ is not free in the dual. The dual contains no $b$; it is recovered afterwards from a support vector with $0 \lt \alpha_i \lt C$ (where $y_if(\mathbf{x}_i) = 1$ exactly), or as a midpoint if there is none.
A big $C$ means "slack is expensive", not "a stronger model". $C \to \infty$ approaches the hard margin and can overfit noisy points. A small $C$ allows many violations and a wider street.
Quick check: why is "the data appear only through dot products $\mathbf{x}_i^\top\mathbf{x}_j$" such a useful fact?
Because you can replace the dot product by a kernel that equals a dot product in a much bigger feature space, such as $(\mathbf{x}^\top\mathbf{z})^2$, without ever building that space. The SVM then fits a straight street in the big space, which looks like a curved boundary in the original plane. The primal cannot do this, since it needs the explicit weight vector.
ML connection: constrained optimization and why people solve the dual
Imagine you are told "I have a plan that costs $1.44$" and "I can prove no plan costs less than $0.96$". You now know the best cost is between $0.96$ and $1.44$, without having found the best plan. The primal gives the upper number, the dual gives the lower number. When they are close, you can stop.
That is one reason people care about the dual. There are three more.
A certificate. QP: $\min\tfrac12(x^2+y^2)$ s.t. $x + y \ge 2$. The point $\tilde{\mathbf{x}} = (1.2, 1.2)$ is feasible with $f = 1.44$. The fee $\mu = 0.8$ gives $d = 2(0.8) - 0.8^2 = 0.96$. So $0.96 \le p^\star \le 1.44$, gap $0.48$. (The true error of $\tilde{\mathbf{x}}$ is $1.44 - 1 = 0.44 \le 0.48$ ✓.)
Solving the dual by climbing. The slope of $d$ is the constraint value at the Lagrangian minimiser: here $d'(\mu) = 2 - 2\mu$. Climb with steps of $0.3$ and keep $\mu \ge 0$ (dual ascent): $\mu_{k+1} = \max\{0,\ \mu_k + 0.3\,(2 - 2\mu_k)\}$. From $\mu_0 = 0$: $0.6,\ 0.84,\ 0.936,\ 0.974,\dots \to 1 = \mu^\star$. At each step the primal guess $\mathbf{x}(\mu) = (\mu,\mu)$ moves towards $(1,1)$.
Why solve the dual?
- Certificates and stopping rules. The gap $f(\tilde{\mathbf{x}}) - d(\boldsymbol\lambda,\boldsymbol\mu) \ge f(\tilde{\mathbf{x}}) - p^\star$ is computable, so solvers stop when it is below a tolerance.
- Simpler constraints. The dual often has only simple constraints ($\boldsymbol\mu \ge 0$, boxes, one equality) even if the primal has complicated ones. Projecting onto those is cheap.
- Fewer or better-structured variables. The SVM primal has $d + 1$ variables and $n$ constraints; the dual has $n$ variables and trivial constraints, and it exposes the kernel structure. Problems that are sums of independent pieces tied by a shared constraint split into small independent problems once priced (dual decomposition).
- Always convex, even if the primal is not. $d$ is concave, so the dual is an easy problem that gives a lower bound on a hard non-convex or integer problem (used inside branch-and-bound).
Dual ascent. If $\mathbf{x}(\boldsymbol\mu) = \arg\min_{\mathbf{x}}L(\mathbf{x},\boldsymbol\mu)$ is unique, then $\nabla d(\boldsymbol\mu) = \mathbf{g}(\mathbf{x}(\boldsymbol\mu))$ (the constraint values). So gradient ascent on $d$ with projection onto $\boldsymbol\mu \ge 0$ is: $\boldsymbol\mu \leftarrow \max\{0, \boldsymbol\mu + t\,\mathbf{g}(\mathbf{x}(\boldsymbol\mu))\}$: raise the price of a constraint that is violated, lower it when there is slack. This is the logic of market-style coordination of many agents.
Why do we need it?
Sometimes the dual is smaller, simpler, or the only thing you can compute. And even when you solve the primal, the dual gap tells you honestly how far from optimal you are.
Where is it used?
SMO and dual coordinate ascent for SVMs and logistic regression, primal-dual interior-point solvers (their stopping criterion is the gap), ADMM and dual decomposition for distributed training, and Lagrangian relaxation for combinatorial problems.
How is it used?
Maintain a feasible primal point and a dual point, report $f - d$, and stop when it is below your tolerance. When solving the dual directly, update the prices by the rule above and recover $\mathbf{x}$ from the Lagrangian minimiser.
The dual point may not give a feasible $\mathbf{x}$. $\mathbf{x}(\boldsymbol\mu)$ can violate constraints until $\boldsymbol\mu$ is close to optimal. To report an upper bound you need a feasible point, which you sometimes have to repair (project, round, or rescale).
A zero gap needs strong duality. On a non-convex problem the best dual value can stay strictly below the best primal value, so a gap never reaches $0$ even at the true optimum.
Quick check: a solver reports primal value $10.02$ and dual value $9.98$. How far from optimal can the primal solution be (in objective value)?
At most the gap: $10.02 - 9.98 = 0.04$. The true optimum lies in $[9.98,\ 10.02]$, so the primal point is within $0.04$ of the best value.
ML connection: regularization, penalty versus constraint core
There are two ways to keep weights small. Penalty form: add a cost for size, $\min f(\mathbf{w}) + \lambda\|\mathbf{w}\|^2$. Budget form: forbid size beyond a limit, $\min f(\mathbf{w})$ subject to $\|\mathbf{w}\|^2 \le s$. They look different, but they are the same problem seen through the multiplier: the penalty weight $\lambda$ is the price of norm budget.
A large $\lambda$ means each unit of weight is expensive, so the weights end up small. A small budget $s$ means the same thing. Pick one number and the other follows.
Take $f(\mathbf{w}) = \|\mathbf{w} - \mathbf{y}\|^2$ with $\mathbf{y} = (3,4)$ (ridge regression with $A = I$).
- Penalty form with $\lambda = 1.5$: the minimiser of $\|\mathbf{w} - \mathbf{y}\|^2 + 1.5\|\mathbf{w}\|^2$ is $\mathbf{w} = \mathbf{y}/(1 + 1.5) = (1.2,\ 1.6)$, with $\|\mathbf{w}\| = 2$.
- Budget form with $r = 2$: we found in the previous worked example that the best $\mathbf{w}$ is $(1.2, 1.6)$ with multiplier $\mu^\star = 1.5$.
Same point. And the multiplier equals the penalty weight: $\mu^\star = \lambda = 1.5$. Another pair: $\lambda = 0.5$ gives $\mathbf{w} = \mathbf{y}/1.5 = (2, 2.67)$, $\|\mathbf{w}\| = 5/1.5 = 3.33$, so it is the same as the budget $r = 3.33$.
The correspondence. Consider the budget problem $\min f(\mathbf{w})$ s.t. $\|\mathbf{w}\|^2 \le s$. Its Lagrangian is $L(\mathbf{w},\mu) = f(\mathbf{w}) + \mu(\|\mathbf{w}\|^2 - s)$. For a fixed $\mu = \lambda$, minimising over $\mathbf{w}$ is the same as minimising $f(\mathbf{w}) + \lambda\|\mathbf{w}\|^2$ (the term $-\lambda s$ is a constant that does not change the minimiser). So:
- Penalty $\Rightarrow$ budget. If $\mathbf{w}_\lambda$ minimises $f + \lambda\|\cdot\|^2$ (with $\lambda \ge 0$), then it also solves the budget problem with $s = \|\mathbf{w}_\lambda\|^2$. Proof: for any $\mathbf{w}$ with $\|\mathbf{w}\|^2 \le s$, optimality of $\mathbf{w}_\lambda$ gives $f(\mathbf{w}) + \lambda\|\mathbf{w}\|^2 \ge f(\mathbf{w}_\lambda) + \lambda s$, hence $f(\mathbf{w}) \ge f(\mathbf{w}_\lambda) + \lambda(s - \|\mathbf{w}\|^2) \ge f(\mathbf{w}_\lambda)$.
- Budget $\Rightarrow$ penalty. If $f$ is convex and $s \gt 0$ (Slater holds), strong duality gives a multiplier $\mu^\star \ge 0$, and the KKT stationarity condition $\nabla f(\mathbf{w}^\star) + 2\mu^\star\mathbf{w}^\star = \mathbf{0}$ is exactly the condition for $\mathbf{w}^\star$ to minimise the convex function $f + \mu^\star\|\cdot\|^2$. So the budget solution is a penalty solution with $\lambda = \mu^\star$.
Price reading. By the shadow-price theorem, with $p(s)$ the best loss for budget $s$: $\;\lambda = -p'(s)$. A bigger $\lambda$ corresponds to a smaller $s$ (a tighter budget, a steeper price). If the budget is slack ($s \ge \|\mathbf{w}_{\text{LS}}\|^2$) then $\lambda = 0$: no regularisation. The same works for other norms: $\|\mathbf{w}\|_1 \le t$ is linked to the Lasso penalty $\lambda\|\mathbf{w}\|_1$ (you will meet this in Chapter 3.11).
Why do we need it?
It unifies two common ways of controlling model size, and explains what the strength $\lambda$ means: how much loss a unit of norm budget buys. It lets you reason about regularization with pictures of balls and ovals.
Where is it used?
Ridge regression, weight decay in neural networks, the Lasso and elastic net, and the "norm ball" pictures that explain why L1 gives sparse weights (Chapter 3.11). Trust-region radii correspond to a Levenberg-style damping $\lambda$ in the same way.
How is it used?
You usually tune $\lambda$ by cross-validation (it is easier than choosing a radius). When you need to enforce a hard budget on the norm, you find the $\lambda$ that makes $\|\mathbf{w}_\lambda\|$ equal to it.
The equivalence needs convexity. For convex losses the two forms match one-to-one. For non-convex losses (neural networks) a penalised solution still solves a budget problem, but not every budget solution is a penalised one, so the two forms can differ.
Weight decay is not always the same as an L2 penalty. For plain gradient descent they coincide; with adaptive methods such as Adam they differ (AdamW, Chapter 3.4). The constraint view above is about the objective, not the optimizer's update.
Quick check: in ridge regression, you increase $\lambda$. Does the equivalent norm budget $r$ grow or shrink, and what happens to the training loss?
The budget $r = \|\mathbf{w}_\lambda\|$ shrinks, because a higher price per unit of norm makes you buy less of it. A smaller budget cannot fit the data better, so the training loss goes up (or stays equal). Regularisation trades training fit for smaller weights.
Recap, cheat sheet and practice
- The primal problem has optimal value $p^\star$. Every feasible point gives an upper bound $f(\tilde{\mathbf{x}}) \ge p^\star$.
- The dual function $d(\boldsymbol\lambda,\boldsymbol\mu) = \min_{\mathbf{x}}L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu)$ (for $\boldsymbol\mu \ge 0$) is always concave (a minimum of affine functions) and always a lower bound.
- Weak duality: $d(\boldsymbol\lambda,\boldsymbol\mu) \le p^\star$, proved in two lines from $h = 0$, $\mu \ge 0$, $g \le 0$ at a feasible point. The duality gap is $p^\star - d^\star \ge 0$.
- The dual problem maximises $d$ over $\boldsymbol\mu \ge 0$: a convex problem even when the primal is not. LP and QP duals are explicit.
- Strong duality ($d^\star = p^\star$) holds for convex problems under Slater's condition (and for feasible, bounded LPs). It fails for non-convex problems (example: gap 1). It goes hand in hand with KKT (stationarity from minimising $L$, slackness from the equality chain) and with a saddle point of $L$.
- Dual variables are prices: $\mu_i^\star = -\partial p/\partial u_i$, the rate at which the optimal value improves when constraint $i$ is relaxed. Untouched constraints have price $0$.
- SVM: the dual is $\max\sum\alpha_i - \tfrac12\boldsymbol\alpha^\top Q\boldsymbol\alpha$ s.t. $0 \le \alpha_i \le C$, $\sum\alpha_iy_i = 0$; only support vectors have $\alpha_i \gt 0$; data enter via $\mathbf{x}_i^\top\mathbf{x}_j$, so kernels replace dot products.
- Regularization: penalty weight $\lambda$ and norm budget $r$ are linked by the multiplier: $\lambda = \mu^\star = -dp/ds$ with $s = r^2$.
Cheat sheet
| Idea | Formula | Picture |
|---|---|---|
| Primal value | $p^\star = \min f$ s.t. $g \le 0$, $h = 0$ | best allowed point (upper bounds from feasible points) |
| Dual function | $d(\boldsymbol\lambda,\boldsymbol\mu) = \min_{\mathbf{x}}L(\mathbf{x},\boldsymbol\lambda,\boldsymbol\mu)$ | lowest point of the Lagrangian; concave |
| Weak duality | $d(\boldsymbol\lambda,\boldsymbol\mu) \le p^\star$ | every fee gives a lower bound |
| Dual problem | $d^\star = \max_{\boldsymbol\mu \ge 0}d$ | climb to the best lower bound |
| Duality gap | $p^\star - d^\star \ge 0$ | room between the bounds |
| Strong duality | $d^\star = p^\star$ (convex + Slater) | bounds meet: certificate |
| Saddle point | $L(\mathbf{x}^\star,\boldsymbol\mu) \le L(\mathbf{x}^\star,\boldsymbol\mu^\star) \le L(\mathbf{x},\boldsymbol\mu^\star)$ | valley in $\mathbf{x}$, ridge in $\boldsymbol\mu$ |
| Shadow price | $\mu_i^\star = -\partial p/\partial u_i$ | profit gained per unit of relaxation |
| Dual slope | $\nabla d(\boldsymbol\mu) = \mathbf{g}(\mathbf{x}(\boldsymbol\mu))$ | raise the price of violated constraints |
| SVM dual | $\max\sum\alpha_i - \tfrac12\sum\alpha_i\alpha_jy_iy_jk(\mathbf{x}_i,\mathbf{x}_j)$, $0 \le \alpha \le C$, $\sum\alpha_iy_i = 0$ | support vectors, kernels |
| Penalty ↔ budget | $\min f + \lambda\|\mathbf{w}\|^2 \;\leftrightarrow\; \min f$ s.t. $\|\mathbf{w}\|^2 \le s$, $\lambda = \mu^\star$ | price of norm budget |
import numpy as np
from scipy.optimize import minimize, minimize_scalar, linprog
# ---- 1) A 1D dual function: min x^2 s.t. 1 - x <= 0 (p* = 1) ----
def d(mu): # d(mu) = min_x x^2 + mu*(1 - x)
r = minimize_scalar(lambda x: x**2 + mu * (1 - x))
return r.fun
for mu in [0, 1, 2, 3]:
print("mu =", mu, " d(mu) =", round(d(mu), 4), " formula mu - mu^2/4 =", mu - mu**2 / 4)
best = minimize_scalar(lambda m: -d(m), bounds=(0, 6), method="bounded")
print("best mu:", round(best.x, 3), " d* =", round(-best.fun, 4), " (p* = 1)") # 2.0 1.0
# ---- 2) Hard-margin SVM on four points: primal vs dual ----
X = np.array([[1, 1], [3, 2], [-1, -1], [-2, -3.0]])
y = np.array([1, 1, -1, -1.0])
Q = (y[:, None] * y[None, :]) * (X @ X.T) # Q_ij = y_i y_j x_i . x_j
dual = minimize(lambda a: -(a.sum() - 0.5 * a @ Q @ a), np.full(4, 0.1), method="SLSQP", tol=1e-12,
bounds=[(0, None)] * 4, constraints=[{"type": "eq", "fun": lambda a: a @ y}])
alpha = dual.x
w = (alpha * y) @ X # w = sum_i alpha_i y_i x_i
sv = alpha > 1e-6
b = np.mean(y[sv] - X[sv] @ w) # from a support vector: y(w.x + b) = 1
print("alpha =", alpha.round(4)) # [0.25 0. 0.25 0. ]
print("w =", w.round(4), " b =", round(b, 4)) # [0.5 0.5] 0.0
print("margins =", (y * (X @ w + b)).round(3)) # [1. 2.5 1. 2.5]
print("dual value :", round(-dual.fun, 4)) # 0.25
cons = [{"type": "ineq", "fun": (lambda z, i=i: y[i] * (X[i] @ z[:2] + z[2]) - 1)} for i in range(4)]
primal = minimize(lambda z: 0.5 * z[:2] @ z[:2], [1, 1, 0], method="SLSQP", tol=1e-12, constraints=cons)
print("primal value:", round(primal.fun, 4)) # 0.25 -> zero duality gap
# ---- 3) Shadow prices: bakery LP (maximise 2x + 3y; oven x+y <= 4; flour x+3y <= 6) ----
def profit(oven, flour):
r = linprog([-2, -3], A_ub=[[1, 1], [1, 3]], b_ub=[oven, flour], bounds=[(0, None)] * 2, method="highs")
return -r.fun, -r.ineqlin.marginals # best profit, dual variables (prices)
p0, price = profit(4, 6)
print("profit", p0, " prices (oven, flour) =", price) # 9.0 [1.5 0.5]
print("oven +1 unit: profit gain =", profit(5, 6)[0] - p0) # 1.5 = the oven price
# ---- 4) Penalty <-> constraint (ridge): A = I, y = (3, 4) ----
yv = np.array([3.0, 4.0])
lam = 1.5
w_pen = yv / (1 + lam) # minimiser of ||w - y||^2 + lam ||w||^2
print("penalised w =", w_pen, " |w| =", np.linalg.norm(w_pen)) # [1.2 1.6] 2.0
res = minimize(lambda w: np.sum((w - yv)**2), [0.1, 0.1], method="SLSQP", tol=1e-12,
constraints=[{"type": "ineq", "fun": lambda w: 2.0**2 - w @ w}])
print("constrained (r = 2) w =", res.x.round(4)) # [1.2 1.6] same point, multiplier = lam
1. For any inequality multipliers $\boldsymbol\mu \ge 0$, the dual function $d(\boldsymbol\lambda,\boldsymbol\mu)$ and the primal optimal value $p^\star$ satisfy…
2. Which condition guarantees strong duality for a typical ML problem?
3. For $\min -x^2$ with $0 \le x \le 2$ and the constraint $x \le 1$ we found $p^\star = -1$ and $d(\mu) = \min(-\mu,\ \mu - 4)$. What is the duality gap?
4. At the optimum of a problem, the multiplier of constraint 1 is $\mu_1^\star = 3$ (standard form $g_1 \le 0$). If you relax it to $g_1 \le 0.1$, the optimal value falls by about…
5. What is special about the SVM dual that makes kernels possible?
6. In ridge regression you increase the penalty weight $\lambda$. In the equivalent budget form $\|\mathbf{w}\|^2 \le s$, what happens?
Practice problems
A. Find the dual function and the dual optimum of $\min x^2$ subject to $x \ge 2$.
Standard form: $g = 2 - x \le 0$. $L = x^2 + \mu(2 - x)$. Minimise: $2x - \mu = 0$, so $x = \mu/2$. Then $d(\mu) = \mu^2/4 + 2\mu - \mu^2/2 = 2\mu - \mu^2/4$. Maximise: $2 - \mu/2 = 0$, so $\mu^\star = 4$ and $d^\star = 8 - 4 = 4$. Primal: $x^\star = 2$, $p^\star = 4$ ✓ (zero gap). Recover $x = \mu^\star/2 = 2$ ✓. The price is $\mu^\star = 4$: relaxing to $x \ge 2 - u$ lowers the value by about $4u$.
B. Write the dual of the LP $\min 2x + 3y$ subject to $x + y \ge 4$, $x \ge 0$, $y \ge 0$, and solve both.
Using the recipe ($A = [1\ 1]$, $b = 4$, $\mathbf{c} = (2,3)$): dual $\max 4\mu$ subject to $\mu \le 2$, $\mu \le 3$, $\mu \ge 0$. The best is $\mu = 2$, $d^\star = 8$. Primal: $x + y \ge 4$ is cheapest with all weight on $x$ (cost $2$ per unit): $(x,y) = (4, 0)$, value $8$ ✓. The price of the constraint is $2$.
C. For $\min\|\mathbf{w} - \mathbf{y}\|^2$ with $\mathbf{y} = (3,4)$ and budget $\|\mathbf{w}\| \le r$, show that the multiplier is $\mu^\star = 5/r - 1$ for $r \lt 5$. What is the optimal value when $r = 1$?
$d(\mu) = \dfrac{25\mu}{1+\mu} - \mu r^2$, $d'(\mu) = \dfrac{25}{(1+\mu)^2} - r^2 = 0$, so $1 + \mu = 5/r$ and $\mu^\star = 5/r - 1$. For $r = 1$: $\mu^\star = 4$, $\mathbf{w}^\star = \mathbf{y}/5 = (0.6, 0.8)$ (length 1 ✓), and the loss is $(3 - 0.6)^2 + (4 - 0.8)^2 = 5.76 + 10.24 = 16$. Check with the dual: $d^\star = 25(4)/5 - 4(1) = 20 - 4 = 16$ ✓.
D. For $\min x^2$ s.t. $x \ge 1$ the price is $\mu^\star = 2$. Predict the optimal value if the wall moves to $x \ge 0.9$, and compare with the exact value.
Here $u = 0.1$ ($g = 1 - x \le 0.1$). Prediction: the value falls by about $\mu^\star u = 0.2$, to about $0.8$. Exact: $x^\star = 0.9$, $p = 0.81$. The theorem says $p(u) \ge p^\star - \mu^\star u = 0.8$ ✓ ($0.81 \ge 0.8$). The line slightly under-estimates because $p(u) = (1-u)^2$ is curved.
E. Two training points $\mathbf{x}_1 = (2,0)$ with $y_1 = +1$ and $\mathbf{x}_2 = (0,0)$ with $y_2 = -1$. Write the hard-margin dual, solve it, and recover $\mathbf{w}$ and $b$.
$Q_{11} = 4$, $Q_{22} = 0$, $Q_{12} = y_1y_2\,\mathbf{x}_1^\top\mathbf{x}_2 = 0$. The constraint $\alpha_1 y_1 + \alpha_2 y_2 = 0$ gives $\alpha_1 = \alpha_2 = \alpha$. Dual: $\max\ 2\alpha - \tfrac12(4\alpha^2) = 2\alpha - 2\alpha^2$, so $\alpha = \tfrac12$, $d^\star = 0.5$. Then $\mathbf{w} = \tfrac12(2,0) = (1, 0)$ and, from the support vector $\mathbf{x}_1$: $2 + b = 1$, $b = -1$. Primal check: $\tfrac12\|\mathbf{w}\|^2 = 0.5 = d^\star$ ✓. Margins: $\mathbf{x}_2$: $-(0 - 1) = 1$ ✓.
F. Ridge with $A = I$ and $\mathbf{y} = (3, 4)$: the penalty is $\lambda = 0.5$. Find the solution and the equivalent norm budget $r$ and say what it costs in training loss compared with no regularization.
$\mathbf{w} = \mathbf{y}/(1 + \lambda) = (3,4)/1.5 = (2,\ 2.667)$. Its length is $5/1.5 = 3.33$, so the equivalent budget is $r = 3.33$ (and the multiplier of that budget is $\mu^\star = \lambda = 0.5$). Training loss: $\|\mathbf{w} - \mathbf{y}\|^2 = (3 - 2)^2 + (4 - 2.667)^2 = 1 + 1.778 = 2.78$, against $0$ without regularization (where $\mathbf{w} = \mathbf{y}$). The budget cost $2.78$ units of loss in exchange for shrinking $\|\mathbf{w}\|$ from $5$ to $3.33$.
Regularization as Optimization
A model that fits its training data perfectly is often a bad model. Regularization is the cure, and it is pure optimization: instead of minimizing the error alone, we minimize error + λ × a price for complexity. One extra term, one knob, and an enormous change in how well models behave on new data.
- Explain the principle Loss + λ × Penalty and the three reasons we add a penalty: overfitting, ill-posed problems and prior beliefs
- Define the L2 and L1 penalties and describe their pull toward zero (fading versus constant)
- Derive and use ridge regression: the closed form $(X^\top X+\lambda I)\mathbf{w}=X^\top\mathbf{y}$, the shrinkage factor $\sigma^2/(\sigma^2+\lambda)$ and the cure for ill-conditioning
- Derive soft-thresholding, understand why the lasso has no closed form yet gives exact zeros, and see the elastic net as the mix of both
- Read the constraint picture (diamond versus circle), prove that penalised and constrained problems match, and say honestly why corners give sparsity
- Connect regularization to generalization (bias-variance, choosing λ by validation, regularization paths) and to sparsity
- Recognise weight decay, early stopping, dropout and data augmentation as relatives of the same idea
The principle: Loss + λ × Penalty core
What we need from earlier chapters: the objective function (Chapter 3.1), gradients and from the Linear Algebra guide the L1 and L2 norms. The calculus guide already showed what a penalty does to a gradient; here we study it as an optimization problem in its own right.
Imagine a student preparing for an exam. One student memorises every practice answer, including the typos in the answer key. Another understands the topic and keeps the explanation simple. On the practice sheet the first student scores 100%. On a new exam the second one does better.
A flexible model is the first student. Given enough freedom it can bend itself through every training point, noise included, and then fail on new data. We want to say to the model: "Fit the data, but stay simple if you can."
The way to say that in optimization is to give the model a budget. We add to the error a price for complexity, and let the optimizer find the best trade between the two. A knob $\lambda$ sets how expensive complexity is: cheap and the model memorises, expensive and it learns almost nothing.
Two weight vectors, same training error. The two features of our data are exact copies: $x_2 = x_1$, and the target is $y = 2x_1$. Compare two models $\hat y = w_1x_1 + w_2x_2$:
- $\mathbf{w}=(1,\,1)$: prediction $x_1 + x_1 = 2x_1$. Training error $0$.
- $\mathbf{w}=(11,\,-9)$: prediction $11x_1 - 9x_1 = 2x_1$. Training error also $0$.
The loss cannot tell them apart. Now add the price $\lambda\|\mathbf{w}\|^2$ with $\lambda = 0.1$:
- $(1,1)$: $0 + 0.1\,(1^2+1^2) = 0.2$.
- $(11,-9)$: $0 + 0.1\,(11^2+(-9)^2) = 0.1\times202 = 20.2$.
- The penalised objective prefers $(1,1)$ by a factor of 100.
Is the choice sensible? On new data where the second feature is slightly different, say $x_2 = x_1 + 0.1$, and the truth is still $y = x_1 + x_2$, the model $(1,1)$ predicts $2x_1 + 0.1$ (exactly right) while $(11,-9)$ predicts $2x_1 - 0.9$ (off by $1.0$). The huge, cancelling weights were a fragile trick that only worked while the two copies stayed identical.
A regularized (or penalised) problem has the form
$$\min_{\mathbf{w}}\;\; \underbrace{L(\mathbf{w})}_{\text{data loss}} \;+\; \lambda\,\underbrace{R(\mathbf{w})}_{\text{penalty}} ,\qquad \lambda \ge 0 .$$- $L(\mathbf{w})$ is the usual training loss (squared error, cross-entropy, ...). It says how well do I fit the data?
- $R(\mathbf{w})$ is the penalty or regularizer. It says how complex is this $\mathbf{w}$? Usually it measures the size of the weights: $\|\mathbf{w}\|_2^2$ (L2) or $\|\mathbf{w}\|_1$ (L1).
- $\lambda$ is the regularization strength, a hyperparameter: you choose it, the optimizer does not learn it. $\lambda = 0$ removes the penalty. As $\lambda\to\infty$ the penalty wins and $\mathbf{w}\to$ the minimiser of $R$ alone (here $\mathbf{0}$).
Three reasons to add a penalty.
- Overfitting. With many parameters and few examples, the best training fit chases noise. The penalty removes wild solutions.
- Ill-posedness. The loss alone may have no single best answer (copies of a feature give a whole trough of equally good $\mathbf{w}$) or a tiny change in the data may move the answer a lot (ill-conditioning, see Chapter 1.7). A penalty makes the answer unique and stable.
- Prior beliefs. Before seeing data we often believe "most weights are small". Bayes' rule turns that belief into a penalty. Suppose the noise in $y$ is Gaussian with variance $\sigma^2$. If we believe each weight is Gaussian with variance $\tau^2$, the most probable weights (the MAP estimate) minimise $\tfrac{1}{2\sigma^2}\|\mathbf{y}-X\mathbf{w}\|^2 + \tfrac{1}{2\tau^2}\|\mathbf{w}\|^2$. Multiply by $\sigma^2$: this is $\tfrac12\|\mathbf{y}-X\mathbf{w}\|^2 + \tfrac{\lambda}{2}\|\mathbf{w}\|^2$ with $\lambda = \sigma^2/\tau^2$. A Laplace prior (density $\propto e^{-|w|/b}$) gives the L1 penalty in the same way, with $\lambda = \sigma^2/b$. Noisy data (big $\sigma$) or a confident prior (small $\tau$) means a big $\lambda$.
Conventions in this chapter. For a linear model with design matrix $X$ ($n$ examples, $p$ features) and targets $\mathbf{y}$ we use $L(\mathbf{w}) = \tfrac12\|\mathbf{y}-X\mathbf{w}\|^2$ (a sum, not an average) and penalties $\tfrac{\lambda}{2}\|\mathbf{w}\|_2^2$ and $\lambda\|\mathbf{w}\|_1$. The factors $\tfrac12$ are chosen so that the formulas come out clean. Libraries scale things differently (an average loss, a different factor): always check the documentation before comparing a $\lambda$ with a library's alpha or C.
Why do we need it?
A model can fit its training data perfectly and still fail on new data, or have no single best answer at all. The penalty is the cheapest way to keep the weights modest, make the answer unique and bring in what we already believe.
Where is it used?
Ridge and lasso regression, logistic regression (scikit-learn's C is $1/\lambda$ up to scaling), SVMs, weight decay in neural-network training, and the regularizers of matrix factorisation in recommender systems.
How is it used?
Add $\lambda R(\mathbf{w})$ to the loss and minimise the sum with any optimizer. Standardise the features first, leave the bias unpenalised, and choose $\lambda$ by checking performance on data the model did not train on.
You cannot choose $\lambda$ by looking at the training loss. The training loss is smallest at $\lambda=0$ (the penalty can only make the training fit worse). Pick $\lambda$ with data the model was not trained on: a validation set or cross-validation (later in this chapter).
Do not penalise the bias (intercept) and standardise the features (mean 0, same spread) before regularizing. The penalty compares weights with each other; if one feature is measured in thousands and another in fractions, their weights live on different scales and the penalty treats them unfairly.
Penalty is not the same as shrinking everything to zero. With a moderate $\lambda$ the weights that the data really needs stay large; only the ones with weak evidence are pushed down.
Quick check: why can regularization never lower the training error?
The unregularised minimiser $\mathbf{w}_0$ already has the lowest possible training loss. Any regularized solution $\mathbf{w}_\lambda$ is a different point, so $L(\mathbf{w}_\lambda)\ge L(\mathbf{w}_0)$. We accept a slightly worse training fit because the regularized model is expected to do better on new data.
L2 regularization: a spring that pulls harder the further you stray core
What we need from earlier chapters: the L2 norm (the usual length of a vector, Chapter 1.2) and derivatives of $w^2$ (Chapter 2.3).
Attach every weight to the origin with a spring. A spring pulls with a force proportional to how far it is stretched: a weight far from zero feels a strong pull, a weight near zero feels almost none. The data pulls the weights toward values that fit it; the springs pull them back toward zero. The final weights are where the two forces balance.
Because the pull fades as a weight shrinks, the spring makes big weights smaller very quickly but never quite finishes the job: the weights get small, but not exactly zero.
There is a second effect. Squaring punishes one big weight much more than several modest ones. So L2 prefers to share the work among correlated features instead of loading everything on one.
Spreading is cheaper. Two features, and the model needs total weight $w_1+w_2 = 2$ to fit the data. Compare $(2,\,0)$ with $(1,\,1)$:
- $(2,0)$: $\tfrac12(2^2+0^2) = 2$.
- $(1,1)$: $\tfrac12(1^2+1^2) = 1$. Half the price.
One gradient step. $w=2$, $\lambda = 0.5$, step size $\eta=0.1$, and the data's own gradient at this point is $0.5$. The penalty $\tfrac\lambda2 w^2$ has slope $\lambda w = 0.5\times2 = 1$.
- Total gradient: $0.5 + 1 = 1.5$.
- Update: $w \leftarrow 2 - 0.1\times1.5 = 1.85$.
- Same step rearranged: $w\leftarrow(1-\eta\lambda)\,w - \eta\cdot0.5 = 0.95\times2 - 0.05 = 1.85$. First shrink by 5%, then take the data step.
The L2 regularizer (also called Tikhonov regularization or, in neural networks, weight decay) penalises the squared length of the weights:
$$R_2(\mathbf{w}) = \tfrac12\|\mathbf{w}\|_2^2 = \tfrac12\sum_{j=1}^{p} w_j^2,\qquad \nabla R_2(\mathbf{w}) = \mathbf{w},\qquad \nabla^2 R_2 = I .$$So the regularized objective $F(\mathbf{w}) = L(\mathbf{w}) + \tfrac{\lambda}{2}\|\mathbf{w}\|^2$ has gradient $\nabla L + \lambda\mathbf{w}$ and Hessian $\nabla^2L + \lambda I$.
- Smooth. Differentiable everywhere, so every optimizer in this guide works unchanged.
- The pull is proportional to the weight: $\lambda w_j$. It fades to $0$ as $w_j\to0$.
- Strongly convex. Adding $\lambda I$ to the Hessian makes every curvature at least $\lambda$. If $L$ is convex, $F$ has exactly one minimum (Chapter 3.6).
- Round. The set $\{\mathbf{w}:\|\mathbf{w}\|_2\le r\}$ is a ball: the penalty does not care about the coordinate axes.
- Prefers sharing. Among all $\mathbf{w}$ with the same total $\sum w_j = s$, the sum of squares is smallest when all $w_j = s/p$.
Why do we need it?
It keeps weights small and finite, makes the objective strongly convex (one answer, better conditioned) and spreads weight over correlated features instead of letting one explode.
Where is it used?
Ridge regression, the default L2 term of logistic regression in scikit-learn, weight decay in nearly every neural-network training recipe, and the factor-matrix penalty in matrix-factorisation recommenders.
How is it used?
Add $\tfrac\lambda2\|\mathbf{w}\|^2$ to the loss, or equivalently add $\lambda\mathbf{w}$ to the gradient, or set weight_decay=λ in the optimizer. Try $\lambda$ on a log grid ($10^{-4},\dots,10^{2}$) and pick by validation.
L2 shrinks, it does not select. A tiny weight is still tiny-but-nonzero. If you need a model that really ignores some features, L2 alone will not give you one (L1 will, next section).
"$\lambda$" means different things in different books. If the penalty is written $\lambda\|\mathbf{w}\|^2$ (no $\tfrac12$) the gradient is $2\lambda\mathbf{w}$, as in the Calculus guide; with $\tfrac\lambda2\|\mathbf{w}\|^2$ it is $\lambda\mathbf{w}$. The two $\lambda$s differ by a factor of 2. We use $\tfrac\lambda2$ in this chapter.
Quick check: $w=-3$, $\lambda=0.5$, $\eta=0.2$, data gradient $=1$. One step with an L2 penalty?
Penalty gradient $\lambda w = 0.5\times(-3) = -1.5$. Total $=1-1.5 = -0.5$. Update $w\leftarrow -3-0.2\times(-0.5) = -2.9$. Check with the shrink form: $(1-\eta\lambda)w-\eta g = 0.9\times(-3)-0.2\times1 = -2.7-0.2 = -2.9$ ✓. Without the penalty the weight would have moved to $-3.2$: the spring held it back.
L1 regularization: a flat fee per unit of weight core
What we need from earlier chapters: the L1 norm (Chapter 1.2) and the derivative of $|w|$ and its kink at 0 (Chapters 2.2 and 2.3).
Now change the rule. Instead of a spring, charge a flat fee per unit of weight: every unit of $|w|$ costs $\lambda$, whether the weight is big or small. A weight of $0.01$ pays at the same rate as a weight of $5$.
A feature is now worth keeping only if the error it removes is more than the fee. A feature whose benefit is smaller than $\lambda$ per unit is not worth its price, so the optimizer sets its weight to exactly zero. The model switches that feature off.
In the language of forces: the L1 pull has a constant size $\lambda$ however small the weight is. A spring fades near zero, so it never finishes the job. A constant pull keeps pushing all the way to zero and, because the pull flips sign right at zero, holds the weight there.
L1 does not care how the weight is shared. Back to the two features with total weight $2$:
- $(2,0)$: $\|\mathbf{w}\|_1 = |2|+|0| = 2$.
- $(1,1)$: $\|\mathbf{w}\|_1 = 1+1 = 2$. The same!
L2 charged $2$ versus $1$ (it favoured sharing). L1 charges the same, so it has no reason to spread the weight, and the data decides. That is why L1 is happy with solutions that put everything on one feature.
The fee against the benefit. Suppose the data pulls the weight $w$ upward with strength $g = 0.3$ (that is how much the loss drops per unit increase of $w$, at $w=0$) and the fee is $\lambda=0.5$. Increasing $w$ saves $0.3$ per unit but costs $0.5$ per unit: a net loss, so the best weight is $0$. If $g$ were $0.8$ it would be worth it, and the weight would become positive.
The L1 regularizer penalises the sum of absolute values:
$$R_1(\mathbf{w}) = \|\mathbf{w}\|_1 = \sum_{j=1}^{p} |w_j| ,\qquad \text{regularized objective } F(\mathbf{w}) = L(\mathbf{w}) + \lambda\|\mathbf{w}\|_1 .$$- The pull has constant size. For $w_j\neq0$ the slope of $\lambda|w_j|$ is $\lambda\,\mathrm{sign}(w_j)$, of size $\lambda$ whatever $|w_j|$ is.
- It has a kink at $0$. At $w_j=0$ the left slope is $-\lambda$ and the right slope is $+\lambda$, so the derivative does not exist. Any number in $[-\lambda,\lambda]$ is a valid subgradient there. The weight sits at $0$ whenever the data's pull on it is inside $[-\lambda,\lambda]$.
- Convex but not smooth. $F$ is convex if $L$ is, so every local minimum is global (Chapter 3.6). But it is not differentiable on the axes, so plain gradient descent behaves badly; special methods (proximal gradient, coordinate descent) are in Chapter 3.13.
- The ball is a diamond. $\{\|\mathbf{w}\|_1\le r\}$ has corners on the coordinate axes. This shape is the geometric reason for sparsity (see "The constraint picture" below).
- Why this particular norm? What we would really like to penalise is the number of non-zero weights, $\|\mathbf{w}\|_0$, but that is a hard combinatorial problem. The L1 norm is the closest convex stand-in (it is the convex envelope of $\|\mathbf{w}\|_0$ on the box $[-1,1]^p$).
Why do we need it?
We often believe that only a few of many features matter. A penalty with constant pull can set the unneeded weights to exactly zero, so the model chooses its own features.
Where is it used?
The lasso for gene selection and finance, sparse logistic regression on text (millions of word counts), sparse coding and compressed sensing, and L1 penalties on activations or on pruned network weights.
How is it used?
Add $\lambda\|\mathbf{w}\|_1$ to the loss and use a solver built for it (Lasso, LogisticRegression(penalty="l1", solver="saga")). Larger $\lambda$ means more zeros. Choose $\lambda$ by cross-validation and read off which weights survived.
Zeros are not magic. L1 gives exact zeros because of the kink (the pull does not fade) and the diamond shape (corners on the axes). It does not guarantee that the zero weights are the "truly useless" features. When features are strongly correlated, L1 may keep one and drop another almost arbitrarily (the elastic net, below, repairs this).
Subgradient steps chatter. Because the slope jumps at $0$, plain gradient descent bounces around zero and never lands exactly on it (orange above). Practical solvers use soft-thresholding or coordinate descent, so you get true zeros.
Quick check: $\lambda=0.5$ and at $w=0$ the loss $L$ has slope $L'(0)=0.2$. Is $w=0$ optimal for $L(w)+\lambda|w|$? What if $L'(0)=0.8$?
With $L'(0)=0.2$: just to the right of $0$ the total slope is $0.2+0.5=0.7>0$ (rising), and just to the left it is $0.2-0.5=-0.3<0$ (the objective falls as $w$ increases toward $0$). It is higher on both sides, so $w=0$ is a minimum. With $L'(0)=0.8$ the total slope just to the left is $0.8-0.5=0.3>0$: the objective is still rising as $w$ goes from left to right, so lower values lie to the left of $0$ and the optimum is a negative weight. The rule: a weight stays at exactly $0$ when $|L'(0)|\le\lambda$.
Ridge regression: the closed form and the shrinkage factor core
What we need from earlier chapters: the normal equations and the original ridge derivation (Chapter 1.10), the SVD (Chapter 1.13) and the condition number (Chapter 1.7). We do not repeat the Linear Algebra derivation; we add the optimization view.
Ridge regression is ordinary least squares with L2 regularization. Its magic is that it still has a clean formula: the old normal equations with one small number added to the diagonal. (That added diagonal is the "ridge" of the name.)
Here is a way to see what it does. Every direction in weight space has a strength: how much the data actually says about that direction (the squared singular value $\sigma^2$). Some directions are strongly supported by the data; others are barely supported (the data hardly varies along them) and any estimate there is mostly noise.
Ridge treats each direction by comparing its strength $\sigma^2$ with $\lambda$. If $\sigma^2\gg\lambda$: trust the data, leave it alone. If $\sigma^2\ll\lambda$: the evidence is thin, shrink it hard toward zero. Exactly the noisy, unstable directions are the ones that get squashed.
Two orthogonal features. Let $X=\begin{bmatrix}2&0\\0&1\end{bmatrix}$ and $\mathbf{y}=[4,\,3]^\top$, and take $\lambda=1$.
- $X^\top X = \begin{bmatrix}4&0\\0&1\end{bmatrix}$ and $X^\top\mathbf{y}=[8,\,3]^\top$.
- Ordinary least squares: $\mathbf{w}_{\text{ols}} = (X^\top X)^{-1}X^\top\mathbf{y} = [8/4,\ 3/1] = [2,\,3]$.
- Ridge: add $\lambda I$: $X^\top X+I=\begin{bmatrix}5&0\\0&2\end{bmatrix}$, so $\hat{\mathbf{w}} = [8/5,\ 3/2] = [1.6,\,1.5]$.
- Shrinkage factors: $\sigma_1^2=4$ gives $\frac{4}{4+1}=0.8$, and $\sigma_2^2=1$ gives $\frac{1}{1+1}=0.5$. Check: $2\times0.8 = 1.6$ ✓ and $3\times0.5=1.5$ ✓.
The stronger direction ($\sigma^2=4$) kept $80\%$ of its OLS value, the weaker one ($\sigma^2=1$) only $50\%$.
An ill-conditioned matrix. Suppose $X$ has singular values $10$ and $0.1$, so $X^\top X$ has eigenvalues $100$ and $0.01$ and condition number $\kappa = 100/0.01 = 10{,}000$. Adding $\lambda I$ moves the eigenvalues to $100+\lambda$ and $0.01+\lambda$:
- $\lambda = 1$: $\kappa = 101/1.01 = 100$.
- $\lambda = 10$: $\kappa = 110/10.01\approx 11$.
Ridge regression solves
$$\hat{\mathbf{w}}_\lambda = \arg\min_{\mathbf{w}}\ \underbrace{\tfrac12\|\mathbf{y}-X\mathbf{w}\|^2}_{\text{fit}} + \underbrace{\tfrac{\lambda}{2}\|\mathbf{w}\|^2}_{\text{penalty}} .$$Derivation. Set the gradient to zero. The gradient of the fit term is $X^\top(X\mathbf{w}-\mathbf{y})$ and of the penalty is $\lambda\mathbf{w}$, so
$$X^\top(X\mathbf{w}-\mathbf{y})+\lambda\mathbf{w}=\mathbf{0}\quad\Longrightarrow\quad \boxed{(X^\top X+\lambda I)\,\mathbf{w}=X^\top\mathbf{y}}$$The Hessian of the objective is $X^\top X+\lambda I$. For any $\mathbf{v}\ne\mathbf{0}$, $\mathbf{v}^\top(X^\top X+\lambda I)\mathbf{v}=\|X\mathbf{v}\|^2+\lambda\|\mathbf{v}\|^2>0$, so the matrix is positive definite for every $\lambda>0$. Therefore the objective is strictly convex, the solution exists, it is unique, and this stationary point is the global minimum. (In code: solve the system, ideally with a Cholesky factorisation; do not form the inverse.)
The SVD view. Write $X=U\Sigma V^\top$ with singular values $\sigma_1\ge\dots\ge\sigma_p$, left vectors $\mathbf{u}_i$ and right vectors $\mathbf{v}_i$. Then $X^\top X+\lambda I = V(\Sigma^2+\lambda I)V^\top$ and $X^\top\mathbf{y}=V\Sigma U^\top\mathbf{y}$, so
$$\hat{\mathbf{w}}_\lambda=\sum_{i}\frac{\sigma_i}{\sigma_i^2+\lambda}(\mathbf{u}_i^\top\mathbf{y})\,\mathbf{v}_i=\sum_i\underbrace{\frac{\sigma_i^2}{\sigma_i^2+\lambda}}_{\text{shrinkage factor } s_i(\lambda)}\;\underbrace{\frac{\mathbf{u}_i^\top\mathbf{y}}{\sigma_i}\mathbf{v}_i}_{\text{the OLS part along }\mathbf{v}_i}.$$So ridge multiplies the OLS component along $\mathbf{v}_i$ by $s_i(\lambda)=\dfrac{\sigma_i^2}{\sigma_i^2+\lambda}\in(0,1]$. It is $\approx1$ when $\sigma_i^2\gg\lambda$, exactly $\tfrac12$ when $\sigma_i^2=\lambda$, and $\approx0$ when $\sigma_i^2\ll\lambda$. As $\lambda\to0$ all factors tend to $1$ (OLS) and as $\lambda\to\infty$ they tend to $0$.
Conditioning. The eigenvalues of $X^\top X+\lambda I$ are $\sigma_i^2+\lambda$, so $$\kappa(X^\top X+\lambda I)=\frac{\sigma_1^2+\lambda}{\sigma_p^2+\lambda},$$ which decreases from $\sigma_1^2/\sigma_p^2$ (at $\lambda=0$) toward $1$ as $\lambda$ grows. It also fixes singular problems: if $\sigma_p=0$ (copied features) plain OLS has no unique solution, but $\sigma_p^2+\lambda=\lambda>0$ makes ridge well-defined. This matters twice: the linear system is solved more accurately, and any iterative optimizer converges faster (a smaller condition number means gradient descent converges faster: you will measure this in the widget below and study it in Chapter 3.5).
Awareness: two equivalent formulas. $\hat{\mathbf{w}}_\lambda=(X^\top X+\lambda I_p)^{-1}X^\top\mathbf{y}=X^\top(XX^\top+\lambda I_n)^{-1}\mathbf{y}$. The first solves a $p\times p$ system, the second an $n\times n$ one: use whichever is smaller (the second form is the starting point of kernel ridge regression). Effective number of parameters: $\mathrm{df}(\lambda)=\sum_i\frac{\sigma_i^2}{\sigma_i^2+\lambda}$ falls smoothly from $p$ (at $\lambda=0$) toward $0$.
Why do we need it?
Plain least squares breaks when features are nearly copies: the answer is wild or not unique and the linear solve is inaccurate. Ridge gives a unique, stable, well-conditioned answer from the same kind of linear algebra.
Where is it used?
Ridge regression in scikit-learn and R's glmnet, regression with thousands of correlated features (genomics, text, finance), the L2 term in logistic regression, and Tikhonov regularization of ill-posed inverse problems (image deblurring, tomography).
How is it used?
Standardise the columns, centre $\mathbf{y}$ (or leave the intercept unpenalised), choose $\lambda$ on a log grid by cross-validation, and solve $(X^\top X+\lambda I)\mathbf{w}=X^\top\mathbf{y}$ by Cholesky. Look at the shrinkage factors or the condition number to see what $\lambda$ did.
Ridge is not scale-invariant. If you multiply a column of $X$ by 1000, its weight shrinks by 1000 and the penalty barely touches it. Standardise first, or ridge treats your units as if they mattered.
Ridge never gives exact zeros. Every factor $s_i(\lambda)$ is strictly positive, so no direction is ever switched off completely (see the lasso for that).
Shrinkage is bias on purpose. Ridge estimates are systematically a bit too small. We accept that bias because it buys a larger drop in variance (see "Bias, variance and choosing $\lambda$").
Library $\lambda$ differs. scikit-learn's Ridge(alpha=a) minimises $\|\mathbf{y}-X\mathbf{w}\|^2+a\|\mathbf{w}\|^2$; multiplying our objective by 2 shows this is the same problem with $\lambda=a$. The Calculus guide uses a mean loss, $\frac1n\|\mathbf{y}-X\mathbf{w}\|^2+\lambda_c\|\mathbf{w}\|^2$; multiplying that by $\frac n2$ gives our form with $\lambda=n\lambda_c$. Always check which convention a library uses.
Quick check: $X=\begin{bmatrix}3&0\\0&1\end{bmatrix}$, $\mathbf{y}=[6,\,2]^\top$, $\lambda=1$. Find the ridge solution and the two shrinkage factors.
$X^\top X=\mathrm{diag}(9,1)$, $X^\top\mathbf{y}=[18,\,2]$, so OLS is $[2,\,2]$. Ridge: $(X^\top X+I)\mathbf{w}=X^\top\mathbf{y}$ gives $\mathrm{diag}(10,2)\mathbf{w}=[18,2]$, so $\hat{\mathbf{w}}=[1.8,\,1]$. The factors are $\frac{9}{9+1}=0.9$ and $\frac{1}{1+1}=0.5$, and indeed $2\times0.9=1.8$, $2\times0.5=1$ ✓.
Lasso: no closed form, but a beautiful special case core
What we need from earlier chapters: the one-weight L1 calculation of Chapter 2.14 (we redo it here more carefully) and, from the Linear Algebra guide, the normal equations.
The lasso (least absolute shrinkage and selection operator) is least squares with an L1 penalty: it shrinks the weights and selects features by setting some exactly to zero.
Why no formula like ridge's? Ridge's penalty is smooth, so "gradient $=0$" is one clean linear equation. The lasso's penalty has a kink at every axis, so there is no gradient there, and the answer depends on which weights end up zero, something we do not know in advance. That is a combinatorial question, and it needs an iterative solver.
But there is one lovely special case: if the features are uncorrelated and equally scaled (orthonormal), the problem falls apart into one tiny problem per weight, and each is solved by the same rule: take the least-squares value, move it $\lambda$ closer to zero, and if that crosses zero, stop at zero. This rule is called soft-thresholding. Values inside the "dead zone" $[-\lambda,\lambda]$ become exactly $0$.
Take $X=I$ (the simplest orthonormal design, so the least-squares weights are just $\mathbf{y}$), $\mathbf{y}=[3,\,-0.5,\,1.2]$ and $\lambda=1$.
- Weight 1: $|3|>1$, so move $1$ toward zero: $3-1 = 2$.
- Weight 2: $|-0.5|\le1$, so it falls in the dead zone: $0$ exactly.
- Weight 3: $1.2-1=0.2$.
- Lasso answer: $[2,\ 0,\ 0.2]$. For comparison, ridge divides by $1+\lambda=2$: $[1.5,\ -0.25,\ 0.6]$, with no zero.
| weight | least squares $a$ | ridge $a/(1+\lambda)$ | lasso $\mathrm{soft}(a,\lambda)$ |
|---|---|---|---|
| 1 | 3 | 1.5 | 2 |
| 2 | $-0.5$ | $-0.25$ | 0 |
| 3 | 1.2 | 0.6 | 0.2 |
Ridge scales everything down by the same percentage; the lasso subtracts the same amount and kills what falls below it.
The lasso solves $$\hat{\mathbf{w}}^{\text{lasso}}_\lambda=\arg\min_{\mathbf{w}}\ \tfrac12\|\mathbf{y}-X\mathbf{w}\|^2+\lambda\|\mathbf{w}\|_1 .$$ It is a convex problem (a convex loss plus a convex penalty), so a local minimum is the global minimum.
Optimality conditions (the subgradient version of "gradient $=0$"). Let $\mathbf{r}=\mathbf{y}-X\hat{\mathbf{w}}$ be the residual and $\mathbf{x}_j$ the $j$-th column of $X$. The gradient of the fit term with respect to $w_j$ is $-\mathbf{x}_j^\top\mathbf{r}$. So $\hat{\mathbf{w}}$ is optimal exactly when for every $j$ $$\mathbf{x}_j^\top\mathbf{r}=\lambda\,\mathrm{sign}(\hat w_j)\ \ \text{if }\hat w_j\ne0,\qquad |\mathbf{x}_j^\top\mathbf{r}|\le\lambda\ \ \text{if }\hat w_j=0 .$$ In words: a feature in the model must be exactly as correlated with the residual as the fee $\lambda$; a feature left out may be at most that correlated. Taking $\hat{\mathbf{w}}=\mathbf{0}$ (so $\mathbf{r}=\mathbf{y}$) shows that all weights are zero when $\lambda\ge\lambda_{\max}=\max_j|\mathbf{x}_j^\top\mathbf{y}|$.
If we knew which weights are non-zero (the active set $\mathcal{A}$) and their signs $\mathbf{s}_{\mathcal{A}}$, the first condition becomes the linear system $X_{\mathcal{A}}^\top(\mathbf{y}-X_{\mathcal{A}}\hat{\mathbf{w}}_{\mathcal{A}})=\lambda\mathbf{s}_{\mathcal{A}}$, i.e. $\hat{\mathbf{w}}_{\mathcal{A}}=(X_{\mathcal{A}}^\top X_{\mathcal{A}})^{-1}(X_{\mathcal{A}}^\top\mathbf{y}-\lambda\mathbf{s}_{\mathcal{A}})$. The hard part is finding $\mathcal{A}$ and $\mathbf{s}$; this is what the solvers of Chapter 3.13 do.
The orthonormal case: derivation. Suppose $X^\top X=I$ and let $\mathbf{a}=X^\top\mathbf{y}$ be the least-squares weights. Expand the fit term: $$\tfrac12\|\mathbf{y}-X\mathbf{w}\|^2=\tfrac12\|\mathbf{y}\|^2-\mathbf{w}^\top X^\top\mathbf{y}+\tfrac12\mathbf{w}^\top X^\top X\mathbf{w}=\tfrac12\|\mathbf{y}\|^2-\mathbf{w}^\top\mathbf{a}+\tfrac12\|\mathbf{w}\|^2=\tfrac12\|\mathbf{w}-\mathbf{a}\|^2+\text{const}.$$ So the lasso objective is $\sum_j\big[\tfrac12(w_j-a_j)^2+\lambda|w_j|\big]+\text{const}$: separable, one independent one-weight problem per $j$. For one weight, $0\in(w-a)+\lambda\,\partial|w|$:
- $w>0$: $w-a+\lambda=0$, so $w=a-\lambda$, valid only if $a>\lambda$.
- $w<0$: $w-a-\lambda=0$, so $w=a+\lambda$, valid only if $a<-\lambda$.
- $w=0$: need $a\in\lambda[-1,1]$, i.e. $|a|\le\lambda$.
Scaling. The threshold equals $\lambda$ here because we used $\tfrac12\|\mathbf{y}-X\mathbf{w}\|^2$ and $X^\top X=I$. If a library uses $\frac{1}{2n}\|\mathbf{y}-X\mathbf{w}\|^2+\alpha\|\mathbf{w}\|_1$ (scikit-learn's Lasso), multiply by $n$ to see $\lambda=n\alpha$. With standardised columns (so $X^\top X\approx nI$ when features are uncorrelated), $\hat w_j=\mathrm{soft}(a_j,\,n\alpha)/n$ where $a_j=\mathbf{x}_j^\top\mathbf{y}$, which in terms of the least-squares weight $a_j/n$ is again a shift by exactly $\alpha$.
General (correlated) features, preview. Fix all weights but $w_j$. The best $w_j$ is again a soft-threshold: $w_j\leftarrow\mathrm{soft}(\rho_j,\lambda)/\|\mathbf{x}_j\|^2$ with $\rho_j=\mathbf{x}_j^\top(\mathbf{y}-\sum_{k\ne j}w_k\mathbf{x}_k)$. Cycling through the coordinates is coordinate descent, the algorithm that powers sklearn.linear_model.Lasso and glmnet (Chapter 3.13). The widgets in this chapter use it behind the scenes.
Why do we need it?
When there are many candidate features and few matter, we want the model to pick them itself and give a sparse, readable result. The lasso is the convex way to do selection and shrinkage in one optimization problem.
Where is it used?
Selecting genes or biomarkers from thousands of measurements, sparse text classifiers, finance factor models, compressed sensing and signal recovery, and as the building block of many sparse-learning methods.
How is it used?
Standardise the features, fit LassoCV or glmnet over a path of $\lambda$ values, choose $\lambda$ by cross-validation, and read the non-zero weights. Soft-thresholding is the single update that every lasso solver repeats.
Soft-thresholding is exact only for orthonormal features. With correlated features the weights interact, so the dead zone depends on the other weights (that is what the second widget shows: the band is on the residual correlation, which changes as the model changes).
The lasso also shrinks the weights it keeps. Every surviving weight is pulled $\lambda$ closer to zero, so large true weights are underestimated. A common fix is to use the lasso only to choose features and then refit those features without the penalty (the "relaxed" or "refit" lasso).
Sign conventions. $\mathrm{soft}(a,\lambda)$ keeps the sign of $a$. Write it as $\mathrm{sign}(a)\max(|a|-\lambda,0)$, not as $\max(a-\lambda,0)$, which would wrongly kill all negative values.
Quick check: orthonormal features, least-squares weights $\mathbf{a}=[-2.5,\ 0.7]$, $\lambda=1$. Lasso and ridge answers?
Lasso: $\mathrm{soft}(-2.5,1)=\mathrm{sign}(-2.5)\max(2.5-1,0)=-1.5$ and $\mathrm{soft}(0.7,1)=0$ (since $0.7\le1$). So $[-1.5,\ 0]$. Ridge: divide by $1+\lambda=2$: $[-1.25,\ 0.35]$.
Elastic Net: both fees at once core
Each of the two penalties has a weakness the other can fix.
- Lasso selects, but among a group of strongly correlated features it picks one, more or less at random, and may pick a different one if the data change a little. Its answer can be unstable.
- Ridge is stable and shares weight among correlated features, but it never switches anything off.
The elastic net charges both fees: the flat per-unit fee of L1 (to create zeros) plus the spring of L2 (to share and to stabilise). Think of a net of springs plus a rope: it stretches and shares like ridge, but it can still cut loose weights that are not needed.
Two identical features. Suppose columns $\mathbf{x}_1=\mathbf{x}_2$. The loss only sees the total $s=w_1+w_2$, so any split of the same total fits equally well. Say the data wants $s=2$ and consider the splits $(2,0)$, $(1,1)$ and $(0,2)$ (all weights $\ge0$).
- Lasso pays $\lambda(|w_1|+|w_2|)=2\lambda$ for every one of them. A three-way tie: the solver returns whichever it happens to reach, and a tiny change in the data can flip it.
- Elastic net also pays $\tfrac{\lambda(1-\alpha)}{2}(w_1^2+w_2^2)$: $(2,0)$ costs $2\lambda(1-\alpha)$, $(1,1)$ costs only $\lambda(1-\alpha)$. The tie is broken, and the unique winner is the even split $(1,1)$.
Numbers. With the orthonormal example from the lasso section ($\mathbf{a}=[3,-0.5,1.2]$), $\lambda=1$ and mixing $\alpha=0.5$: first soft-threshold by $\lambda\alpha=0.5$, then divide by $1+\lambda(1-\alpha)=1.5$.
- $3\to(3-0.5)/1.5=1.667$.
- $-0.5\to\mathrm{soft}(-0.5,0.5)=0$ (it just fits inside the dead zone).
- $1.2\to(1.2-0.5)/1.5=0.467$.
Answer $[1.667,\ 0,\ 0.467]$: a zero like the lasso, and smaller, tamer survivors like ridge.
The elastic net (Zou and Hastie, 2005) minimises
$$\tfrac12\|\mathbf{y}-X\mathbf{w}\|^2+\lambda\Big[\alpha\|\mathbf{w}\|_1+\tfrac{1-\alpha}{2}\|\mathbf{w}\|_2^2\Big],\qquad \lambda\ge0,\ \alpha\in[0,1].$$
$\alpha=1$ is the lasso, $\alpha=0$ is ridge. ($\lambda$ sets the total strength, $\alpha$ the mix. scikit-learn's ElasticNet calls the mix l1_ratio and uses the loss $\frac1{2n}\|\mathbf{y}-X\mathbf{w}\|^2$, so its alpha equals our $\lambda/n$.)
- Orthonormal features: $\hat w_j=\dfrac{\mathrm{soft}(a_j,\ \lambda\alpha)}{1+\lambda(1-\alpha)}$. Derivation: per weight, $0\in(w-a)+\lambda\alpha\,\partial|w|+\lambda(1-\alpha)w$. For $w\ne0$ this gives $w\,(1+\lambda(1-\alpha))=a-\lambda\alpha\,\mathrm{sign}(w)$, and $w=0$ is optimal exactly when $|a|\le\lambda\alpha$.
- Strongly convex when $\alpha<1$. The Hessian of the smooth part is at least $\lambda(1-\alpha)I$ (plus $X^\top X$), so the minimiser is unique, even with copied features or $p>n$.
- Grouping effect. Highly correlated features receive similar weights. For identical columns the objective with a fixed total $w_1+w_2=s$ is smallest when $w_1=w_2=s/2$, because $w_1^2+w_2^2$ is smallest there.
- It can keep more than $n$ features. With $p>n$ the lasso selects at most $n$ features; the elastic net can keep more.
- It is just a lasso in disguise. $\tfrac12\|\mathbf{y}-X\mathbf{w}\|^2+\tfrac{\lambda(1-\alpha)}{2}\|\mathbf{w}\|^2=\tfrac12\|\mathbf{y}^*-X^*\mathbf{w}\|^2$ for the stacked data $X^*=\begin{bmatrix}X\\ \sqrt{\lambda(1-\alpha)}\,I\end{bmatrix}$, $\mathbf{y}^*=\begin{bmatrix}\mathbf{y}\\ \mathbf{0}\end{bmatrix}$. So any lasso solver solves it. Coordinate descent simply uses $w_j\leftarrow\mathrm{soft}(\rho_j,\lambda\alpha)/(\|\mathbf{x}_j\|^2+\lambda(1-\alpha))$.
Why do we need it?
The lasso is unstable when features are correlated and cannot keep more than $n$ of them. The elastic net keeps the selection but shares weight within correlated groups and gives one stable answer.
Where is it used?
Gene-expression and other genomics models (thousands of correlated features), text and click-prediction models with correlated counts, and as the default regularized regression in glmnet-style pipelines (ElasticNetCV in scikit-learn).
How is it used?
Fix a mixing $\alpha$ such as $0.5$ (or try $0.1,0.5,0.9$), then choose $\lambda$ by cross-validation. Use $\alpha$ near $1$ for sparser answers and near $0$ for more sharing and stability.
Two knobs to tune. The elastic net needs both $\lambda$ and $\alpha$. In practice fix a few values of $\alpha$ and cross-validate $\lambda$ for each.
"Naive" double shrinkage. In the original paper the plain elastic net shrinks the surviving weights twice (a soft-threshold and then a divide by $1+\lambda(1-\alpha)$) and the authors suggest rescaling by $1+\lambda(1-\alpha)$ afterwards; many libraries do not. This is a detail to know exists, not something to worry about on a first pass.
It does not recover truth by magic. The elastic net makes the selection stable and grouped; it does not tell you which of a group is "really" responsible.
Quick check: identical columns, total $w_1+w_2=s$ fixed. Which split does the elastic net choose, and why?
The even split $w_1=w_2=s/2$. The L1 part $|w_1|+|w_2|=s$ is the same for every split with both weights $\ge0$ (no preference), but the L2 part $w_1^2+w_2^2$ is smallest when the weights are equal (for $s=2$: $2$ for $(1,1)$ against $4$ for $(2,0)$). So the L2 term breaks the tie in favour of sharing.
Constraint interpretation: a penalty and a budget are the same thing core
What we need from earlier chapters: the idea of a constrained problem and its feasible region (Chapter 3.7). Lagrange multipliers (Chapter 3.8), KKT (3.9) and duality (3.10) make the equivalence below fully precise; here we prove the easy half and see the rest in pictures.
There are two ways to keep a suitcase light. You can pay per kilogram (a fee on every extra kilo) or you can be told "20 kg maximum" (a budget). Both make you pack less. And they are linked: if the airline raises the price per kilo, you end up carrying fewer kilos, so every price corresponds to some weight you actually carry, and every weight limit (that actually bites) corresponds to some price at which you would have chosen exactly that much.
Regularization has the same two faces:
- Penalised: minimise $L(\mathbf{w})+\lambda R(\mathbf{w})$. Complexity has a price $\lambda$.
- Constrained: minimise $L(\mathbf{w})$ subject to $R(\mathbf{w})\le r$. Complexity has a budget $r$.
Turn up the price and the budget actually used shrinks. The multiplier $\lambda$ is the exchange rate: how much loss you would save per extra unit of budget.
One weight, both ways. The data alone would choose $w=a=3$, so $L(w)=\tfrac12(w-3)^2$.
| penalty price $\lambda$ | L1 penalty: $\mathrm{soft}(3,\lambda)$ | budget form $\vert w\vert\le r$ with that $r$ | L2 penalty: $3/(1+\lambda)$ | budget $\vert w\vert\le r$ with that $r$ |
|---|---|---|---|---|
| 0 | 3 | $r\ge3$ (budget not binding) | 3 | $r\ge3$ |
| 0.5 | 2.5 | $r=2.5$: best allowed point is $2.5$ ✓ | 2 | $r=2$: best is $2$ ✓ |
| 1 | 2 | $r=2$: best is $2$ ✓ | 1.5 | $r=1.5$ ✓ |
| 2 | 1 | $r=1$ ✓ | 1 | $r=1$ ✓ |
| 3 | 0 | $r=0$ ✓ | 0.75 | $r=0.75$ ✓ |
With the budget $|w|\le r$ and $r<3$ the best allowed weight is the boundary point $w=r$, since the loss still wants to go higher. So each price $\lambda$ picks one budget $r(\lambda)$ and the two problems return the same $w$. Notice the price of budget: for the L1 row, $r=3-\lambda$, and the best loss with budget $r$ is $L^\star(r)=\tfrac12(3-r)^2$, whose slope is $\tfrac{dL^\star}{dr}=-(3-r)=-\lambda$. The price is exactly the rate at which the best loss falls as the budget grows.
For a penalty $R$ (such as $\|\mathbf{w}\|_1$ or $\|\mathbf{w}\|_2$) consider the two problems $$\textbf{(P}_\lambda\textbf{)}\ \ \min_{\mathbf{w}}\ L(\mathbf{w})+\lambda R(\mathbf{w})\qquad\qquad \textbf{(C}_r\textbf{)}\ \ \min_{\mathbf{w}}\ L(\mathbf{w})\ \ \text{subject to}\ \ R(\mathbf{w})\le r .$$
Claim 1 (penalised $\Rightarrow$ constrained; elementary). If $\hat{\mathbf{w}}$ solves $(\mathrm{P}_\lambda)$ with $\lambda\ge0$, then it solves $(\mathrm{C}_r)$ for $r=R(\hat{\mathbf{w}})$.
Proof. Take any $\mathbf{w}$ with $R(\mathbf{w})\le r=R(\hat{\mathbf{w}})$. Because $\hat{\mathbf{w}}$ is optimal for $(\mathrm{P}_\lambda)$, $$L(\mathbf{w})+\lambda R(\mathbf{w})\ \ge\ L(\hat{\mathbf{w}})+\lambda R(\hat{\mathbf{w}})\ \Longrightarrow\ L(\mathbf{w})\ \ge\ L(\hat{\mathbf{w}})+\lambda\big[R(\hat{\mathbf{w}})-R(\mathbf{w})\big]\ \ge\ L(\hat{\mathbf{w}}),$$ since $\lambda\ge0$ and $R(\hat{\mathbf{w}})-R(\mathbf{w})\ge0$. So nothing in the allowed set beats $\hat{\mathbf{w}}$. $\blacksquare$
Claim 2 (constrained $\Rightarrow$ penalised; needs convexity). If $L$ and $R$ are convex and the budget $r$ is not trivial (the constraint actually bites), there is a $\lambda\ge0$ such that the solution of $(\mathrm{C}_r)$ also solves $(\mathrm{P}_\lambda)$. That $\lambda$ is the Lagrange multiplier of the constraint, and $\lambda=-\,dL^\star/dr$ is the price of the budget. This is the content of Lagrange/KKT theory and duality (Chapters 3.8 to 3.10).
- As $\lambda$ increases from $0$, the budget actually used, $r(\lambda)=R(\hat{\mathbf{w}}_\lambda)$, decreases from $R(\mathbf{w}_{\text{unpenalised}})$ down to $0$ (L1: reached at $\lambda_{\max}$; L2: only in the limit).
- The map $\lambda\leftrightarrow r$ depends on the data: there is no data-independent formula.
- Penalising $\|\mathbf{w}\|_2^2$ and bounding $\|\mathbf{w}\|_2$ are the same up to relabelling $r\to r^2$.
- Without convexity only Claim 1 is guaranteed.
Why both forms are useful: the penalised form is an unconstrained problem, which is what gradient methods like most. The constrained form is easier to picture and to interpret ("the weights must fit in a box of this size"), and it leads to projected methods that repeatedly push the weights back into the allowed region (Chapter 3.17).
Why do we need it?
Two descriptions of one idea are better than one: the price view tells us how to optimise, the budget view tells us what the answer looks like and why it behaves as it does (the diamond and the circle).
Where is it used?
The original lasso paper was stated as a budget on $\|\mathbf{w}\|_1$; weight-norm constraints and gradient-norm clipping in deep learning; projected gradient methods on balls and simplices; and the sensitivity ("shadow price") reading of $\lambda$ in SVMs and portfolio models.
How is it used?
To get a budget $r$ from a fitted price $\lambda$, compute $r=R(\hat{\mathbf{w}}_\lambda)$. To get the price from a budget, read off the multiplier of the solver, or tune $\lambda$ until $R(\hat{\mathbf{w}})$ equals $r$.
The correspondence is one-to-one only in the convex, "biting" case. If the budget is large enough to contain the unpenalised answer, the constraint does nothing and no positive $\lambda$ matches it (that is $\lambda=0$). For non-convex losses the penalised problem can miss some constrained solutions.
Do not set $r$ and $\lambda$ independently. They are two knobs for one trade-off. Tune one of them (usually $\lambda$) by validation.
Quick check: for $L(w)=\tfrac12(w-3)^2$ and an L1 penalty with $\lambda=1.5$, what is the penalised answer, which budget $r$ gives the same answer, and what is $dL^\star/dr$ there?
Penalised: $\mathrm{soft}(3,1.5)=1.5$. Budget $|w|\le r$ with $r=1.5$: the best allowed point is $1.5$ ✓. The best loss with budget $r$ is $L^\star(r)=\tfrac12(3-r)^2$, so $dL^\star/dr=-(3-r)=-1.5=-\lambda$ ✓: the price equals the slope.
The constraint picture: diamond versus circle core
Here is the famous picture. Plot the two weights $w_1,w_2$. The loss has elliptical contour lines around its best unconstrained point, like the rings of a target. The budget is a region around the origin. The answer is the first place where a growing ring touches the region.
- With L2 the region is a disc: smooth everywhere. The ring touches it at whatever point faces the target. It lands exactly on an axis only if the target happens to be exactly aligned with it: practically never.
- With L1 the region is a diamond with four sharp corners, all on the axes. A corner sticks out the furthest, so a growing ring very often meets it first. At a corner one weight is exactly zero.
Think of lowering a round pebble onto a table that has a spike: the spike usually touches first.
Use the loss $L(\mathbf{w})=\tfrac12(\mathbf{w}-\mathbf{c})^\top A(\mathbf{w}-\mathbf{c})$ with $A=\begin{bmatrix}1.6&0.7\\0.7&0.9\end{bmatrix}$ and unconstrained best point $\mathbf{c}=(2,\,0.8)$ (this is the picture in the widgets). Take budget $t=1.2$.
- L1 budget $|w_1|+|w_2|\le1.2$: the best allowed point is $(1.2,\ 0)$, the right-hand corner, with loss $1.248$. One weight is exactly $0$.
- L2 budget $\sqrt{w_1^2+w_2^2}\le1.2$: the best allowed point is $(1.063,\ 0.557)$, with loss $0.889$. Both weights are non-zero.
(The disc of radius $1.2$ contains the diamond with the same $t$, because $\|\mathbf{w}\|_2\le\|\mathbf{w}\|_1$, so its best loss is lower. The two budgets are measured in different norms and are not comparable like for like.) Now move the target to $\mathbf{c}=(1,2)$: the L1 answer becomes $(0.673,\ 0.527)$, no zero: the ring now touches a flat edge of the diamond, not a corner. So L1 does not always give zeros; it gives them for a whole range of targets.
When does a corner win? For a convex loss and a convex region $C$, a boundary point $\mathbf{v}$ is the solution exactly when the downhill direction $-\nabla L(\mathbf{v})$ points out of the region, that is, lies in the normal cone of $C$ at $\mathbf{v}$ (the set of all directions pointing outward there; this is the KKT condition you will meet in Chapter 3.9).
- Smooth boundary point (circle): the normal cone is a single ray (the outward normal). At the axis point $(t,0)$ it is the ray $(1,0)$, so we need $\partial L/\partial w_2=0$ exactly. This is a condition on one number, so it holds only on a line of targets $\mathbf{c}$: zero area.
- Corner of the diamond at $(t,0)$: the two edges meeting there have outward normals $(1,1)$ and $(1,-1)$, and the normal cone is everything between them: $\{(x,y): x\ge|y|\}$. So the corner wins whenever $-\partial L/\partial w_1\ \ge\ |\partial L/\partial w_2|$ at the corner. This is a wedge of targets with positive area (the shaded zone in the widget below).
- Elastic-net region $\alpha\|\mathbf{w}\|_1+(1-\alpha)\|\mathbf{w}\|_2^2\le t$: still has corners on the axes, but the wedge is narrower (its normals at $(\rho,0)$ are $(\alpha+2(1-\alpha)\rho,\ \pm\alpha)$). As $\alpha\to0$ the wedge shrinks to nothing: ridge.
- Smaller budget, bigger wedges. The set of targets $\mathbf{c}$ captured by the corner $(t,0)$ is a cone with its apex at that corner. As $t$ shrinks the apexes slide toward the origin and the four wedges cover more of the plane. In a window of $\pm3.4$ by $\pm2.5$, for the loss used here, the diamond's corner wedges cover about $60\%$ of the window when $t=1$, $26\%$ when $t=2$ and $3\%$ when $t=3$; for the elastic net with $\alpha=\tfrac12$ the figures are $16\%$, $6\%$ and $3\%$. In $p$ dimensions the L1 ball has $2p$ vertices and also lower-dimensional faces where several weights vanish; the lower-dimensional the face, the larger its normal cone, which is why solutions tend to be sparse when $p$ is large.
An honest statement. This is geometry, not magic. Nothing in the lasso "knows" which features are useful. Zeros appear because the L1 ball has corners on the axes, and a whole wedge of datasets lands on them. Which weights become zero is decided by the data (where the target $\mathbf{c}$ is), and the lasso can drop a useful feature or keep a useless one.
Why do we need it?
The picture explains why L1 gives exact zeros and L2 does not, so you can predict what a penalty will do before running it, and invent new penalties by designing the shape of the region.
Where is it used?
Explaining lasso versus ridge in every textbook, designing group-lasso and nuclear-norm regularizers (regions with corners at "structured sparse" points), and understanding the projection step of projected gradient methods on balls and diamonds.
How is it used?
Ask "what shape is the region, and where are its corners?". Corners on the coordinate axes mean sparse weights; a smooth ball means shrinkage without zeros; a rounded diamond (elastic net) sits in between.
The picture shows two weights; real models have thousands. The mechanism is the same (corners and low-dimensional faces on the axes), but you cannot see it, so trust the algebra of the previous sections.
The diamond only helps if the weights are comparable. If one feature's scale is 1000 times another's, "the same budget" means very different things for each; standardise the features first.
Not every norm-ball region gives sparsity. Only regions with corners on the axes do. The L2 ball does not; an L$\infty$ ball (a square) has corners but not on the axes, so it tends to push weights toward the same size, not toward zero.
Quick check: with $A$ as above, $\mathbf{c}=(2,2)$ and L1 budget $t=1.2$, verify that the right-hand corner $(1.2,\,0)$ is the solution, using the wedge test. What does the L2 disc give?
At $\mathbf{v}=(1.2,0)$: $\mathbf{v}-\mathbf{c}=(-0.8,-2)$ and $\nabla L=A(\mathbf{v}-\mathbf{c})=(1.6\cdot(-0.8)+0.7\cdot(-2),\ 0.7\cdot(-0.8)+0.9\cdot(-2))=(-2.68,\,-2.36)$. The wedge condition is $-\partial_1L\ge|\partial_2L|$: $2.68\ge2.36$ ✓. So the corner wins even though both coordinates of $\mathbf{c}$ are large. The L2 disc of radius $1.2$ gives $(0.966,\,0.712)$ instead: no zero.
Regularization and generalization: bias, variance and choosing $\lambda$ core
A model can be wrong in two different ways. Picture an archer.
- Bias: the arrows land close together but away from the bullseye. The archer is consistent but systematically off. A model that is too rigid (very large $\lambda$) is like this: it cannot bend enough to match the truth, whatever data it sees.
- Variance: the arrows are scattered all around the bullseye. On average they are centred, but any single shot is far off. A model that is too flexible (very small $\lambda$) is like this: it follows the noise of whichever training set it happened to see, so a different training set gives a very different model.
Regularization buys a big drop in variance for a small rise in bias. Too little and variance dominates (overfitting); too much and bias dominates (underfitting). The best $\lambda$ is in the valley between.
One direction, exact numbers. Consider one direction of a ridge problem with strength $a=\sigma_i^2=4$, true weight $w=1$ and noise level $\sigma=1$. Ridge multiplies the least-squares estimate by $s=\frac{a}{a+\lambda}$. The least-squares estimate is unbiased with variance $\sigma^2/a=0.25$. So the ridge estimate has
- mean $=s\,w$, so bias $=(s-1)w=-\frac{\lambda}{a+\lambda}w$, and variance $=s^2\,\sigma^2/a$.
The expected squared error is $\mathrm{bias}^2+\mathrm{variance}=\dfrac{\lambda^2w^2+a\sigma^2}{(a+\lambda)^2}$. Evaluate it:
| $\lambda$ | bias$^2$ | variance | total |
|---|---|---|---|
| 0 (OLS) | 0 | $4/16=0.25$ | 0.25 |
| 0.5 | $0.25/20.25=0.0123$ | $(4/4.5)^2/4=0.1975$ | 0.2099 |
| 1 | $1/25=0.04$ | $(4/5)^2/4=0.16$ | 0.2 |
| 2 | $4/36=0.1111$ | $(4/6)^2/4=0.1111$ | 0.2222 |
| 4 | $16/64=0.25$ | $(4/8)^2/4=0.0625$ | 0.3125 |
A little shrinkage lowers the total error below OLS ($0.25\to0.2$); too much raises it again. The minimum is at $\lambda^\star=\sigma^2/w^2=1$, which is the same $\lambda=\sigma^2/\tau^2$ the Bayesian argument gave at the start of the chapter (with $\tau^2=w^2$).
Overfitting and underfitting. The training error is measured on the data used to fit; the test error on new data. The gap between them is the generalization gap. Overfitting: tiny training error, large test error. Underfitting: both large.
The bias-variance decomposition (squared error). At a fixed input $x$, let $y=f(x)+\varepsilon$ with $\mathbb{E}\varepsilon=0$ and $\mathrm{Var}\,\varepsilon=\sigma^2$, and let $\hat y$ be the prediction of a model trained on a random training set (independent of $\varepsilon$). Write $\bar y=\mathbb{E}\hat y$ for the average prediction over training sets. Then $$\mathbb{E}\big[(y-\hat y)^2\big]=\underbrace{\sigma^2}_{\text{noise}}+\underbrace{(f(x)-\bar y)^2}_{\text{bias}^2}+\underbrace{\mathbb{E}\big[(\hat y-\bar y)^2\big]}_{\text{variance}} .$$ Derivation. Write $y-\hat y=(y-f)+(f-\bar y)+(\bar y-\hat y)$. Square it. The cross terms vanish: $\mathbb{E}[y-f]=0$ and $\varepsilon$ is independent of $\hat y$; and $\mathbb{E}[\bar y-\hat y]=0$ while $(f-\bar y)$ is a constant. What is left is the three squares. $\blacksquare$
- The noise term cannot be reduced by any model.
- Raising $\lambda$ lowers the variance and raises the bias. For ridge (per direction, as in the example) the mean is $s_i\cdot$(truth) and the variance is $s_i^2\sigma^2/\sigma_i^2$; both change smoothly with $\lambda$.
- The total (noise + bias$^2$ + variance) is typically U-shaped in $\lambda$. The sweet spot is where the two competing terms balance.
Choosing $\lambda$ in practice. We cannot compute bias and variance (we do not know $f$), so we estimate the test error with held-out data.
- Hold-out validation. Fit on a training set for each $\lambda$ on a log grid; pick the $\lambda$ with the smallest error on a separate validation set.
- $K$-fold cross-validation (when data are scarce). Split the data into $K$ parts; for each $\lambda$ and each part, train on the other $K-1$ parts and measure the error on the held-out part; average the $K$ errors (their spread gives a standard error). Pick the $\lambda$ with the smallest average.
- Optionally the one-standard-error rule (a heuristic, not a theorem): choose the largest $\lambda$ whose error is within one standard error of the minimum. It prefers the simpler model when the data cannot tell the two apart.
- Refit on all training data with the chosen $\lambda$, and touch a separate test set once, at the end, to report the result.
Why do we need it?
Training error always prefers the most flexible model. The bias-variance view explains why, and cross-validation gives a fair estimate of how the model will do on new data, so we can choose $\lambda$ sensibly.
Where is it used?
RidgeCV, LassoCV and GridSearchCV in scikit-learn, hyperparameter tuning of weight decay in neural networks, and every Kaggle-style pipeline with a validation split.
How is it used?
Try $\lambda$ on a log grid ($10^{-4},\dots,10^{2}$), compute the validation or cross-validation error for each, choose the best (or the one-standard-error choice), refit on all training data, and report the test error once.
Never tune $\lambda$ on the test set. If you pick the best $\lambda$ by looking at test error, the test error is no longer an honest estimate (it has been optimised). Use cross-validation or a validation set for tuning, and the test set once.
Do the preprocessing inside each fold. If you standardise the features using all the data before splitting, information from the held-out fold leaks into training. Fit the scaler on the training part of each fold only.
The U-shape is one regime, not a law of nature. The classical bias-variance picture fits regularised linear models and small models well. Very large modern networks can show different behaviour (test error that falls again as models grow, called double descent), and the decomposition above is only for squared error. Treat it as the first, most useful lens rather than the whole story.
Quick check: a model has training error $0.01$ and test error $0.40$. Should you increase or decrease $\lambda$? What if both errors are $0.40$?
A big gap (tiny training error, large test error) means overfitting: high variance. Increase $\lambda$ (or simplify the model, or get more data). If both errors are large and about equal, the model is underfitting: high bias. Decrease $\lambda$ (or use a more flexible model).
Regularization paths: the whole movie, not one frame
So far we looked at one $\lambda$ at a time. A regularization path is the whole film: start with $\lambda$ so large that every weight is zero, then lower it slowly and watch the weights come alive.
- Ridge: all weights grow together, smoothly, from zero. Nobody is ever exactly zero.
- Lasso: the weights switch on one at a time. The feature that matters most walks in first, then the next, and so on. Between two events every weight moves in a straight line: the path is piecewise linear.
The path shows you which features are strong (they enter early), which are weak (late), and how the whole model changes as you relax the penalty.
In the data set of the widget below (40 examples, 8 standardised features of which only features 1, 2 and 5 truly matter, with true weights $3$, $-2$ and $1.5$), the lasso path reads like this (rounded):
| $\lambda/\lambda_{\max}$ | weights of features $1,\dots,8$ | in the model |
|---|---|---|
| 1.0 | all $0$ | none |
| 0.8 | $[0.64,\,0,\,0,\,0,\,0,\,0,\,0,\,0]$ | 1 |
| 0.5 | $[1.46,\,0,\,0,\,0,\,0.25,\,0,\,0,\,0]$ | 1, 5 |
| 0.3 | $[1.95,\,-0.29,\,0,\,0,\,0.69,\,0,\,0,\,0]$ | 1, 2, 5 |
| 0.05 | $[2.81,\,-1.63,\,0.23,\,0.02,\,1.20,\,0,\,0,\,0]$ | 1, 2, 3, 4, 5 |
The three useful features enter first (feature 1 at $\lambda=\lambda_{\max}$, feature 5 at about $0.62\lambda_{\max}$, feature 2 at about $0.36\lambda_{\max}$). The useless feature 3 enters only at about $0.12\lambda_{\max}$ and feature 4 at about $0.06\lambda_{\max}$. Between $\lambda/\lambda_{\max}\approx0.36$ and $0.12$ the model has exactly the right three features.
The regularization path is the function $\lambda\mapsto\hat{\mathbf{w}}(\lambda)$ for all $\lambda\ge0$.
- Ridge path. $\hat{\mathbf{w}}(\lambda)=(X^\top X+\lambda I)^{-1}X^\top\mathbf{y}$ is a smooth (analytic) function of $\lambda$. It tends to $\mathbf{0}$ only as $\lambda\to\infty$, and the length $\|\hat{\mathbf{w}}(\lambda)\|$ decreases steadily as $\lambda$ grows (each shrinkage factor $\sigma_i^2/(\sigma_i^2+\lambda)$ falls). Individual weights can still change sign or rise and then fall (a weight may first grow while a correlated partner shrinks).
- Lasso path. $\hat{\mathbf{w}}(\lambda)=\mathbf{0}$ for $\lambda\ge\lambda_{\max}=\max_j|\mathbf{x}_j^\top\mathbf{y}|$. The path is piecewise linear in $\lambda$: on each piece the active set and signs are fixed, and then (from the previous section) $\hat{\mathbf{w}}_{\mathcal{A}}=(X_{\mathcal{A}}^\top X_{\mathcal{A}})^{-1}(X_{\mathcal{A}}^\top\mathbf{y}-\lambda\mathbf{s}_{\mathcal{A}})$ is a linear function of $\lambda$. A feature enters when its correlation with the residual reaches $\lambda$; a feature can also leave when its weight returns to zero. (The LARS algorithm traces the exact path.)
- Warm starts. Solvers such as
glmnetandLassoCVcompute the path on a decreasing grid of $\lambda$ values, starting each solve from the previous answer. That is much cheaper than solving each $\lambda$ from scratch, so a whole path costs little more than a single fit at the smallest $\lambda$.
Why do we need it?
We rarely know the right $\lambda$ in advance. The path shows every model at once, tells us the order in which features matter, and lets cross-validation pick one point on it.
Where is it used?
The path plots in every lasso paper, lasso_path and LassoCV in scikit-learn, glmnet in R, feature-ranking in genomics, and the "regularization trace" used to check how stable a model is.
How is it used?
Compute the path on a log grid of $\lambda$ from $\lambda_{\max}$ down to about $10^{-3}\lambda_{\max}$, plot weights against $\lambda$, choose $\lambda$ by cross-validation, and read off which features entered first.
Entering order is not a ranking of "importance". When two features are strongly correlated, whichever has a slightly higher correlation with the target enters first and may then hide the other. And a feature with a small effect can enter late simply because the strong ones already explain most of $\mathbf{y}$.
Ridge and lasso $\lambda$ values are not comparable. The same $\lambda$ means different amounts of shrinkage for the two penalties. In the widget they are plotted against the same ratio $\lambda/\lambda_{\max}$ only for convenience.
Quick check: why is the lasso weight vector exactly $\mathbf{0}$ for every $\lambda\ge\lambda_{\max}=\max_j|\mathbf{x}_j^\top\mathbf{y}|$?
At $\mathbf{w}=\mathbf{0}$ the residual is $\mathbf{y}$. The optimality condition for a zero weight says $|\mathbf{x}_j^\top\mathbf{r}|\le\lambda$, which here reads $|\mathbf{x}_j^\top\mathbf{y}|\le\lambda$ for every $j$. That holds exactly when $\lambda\ge\max_j|\mathbf{x}_j^\top\mathbf{y}|$. So at or above $\lambda_{\max}$ the all-zero vector is optimal (and, the problem being convex, it is the answer).
Sparsity: why zeros are useful core
A weight vector is sparse if most of its entries are exactly zero. A zero weight means "this input is ignored". So a sparse model uses only a few of the available inputs.
Why would you want that?
- Feature selection: the model tells you which inputs it needs. You can stop collecting the others (cheaper sensors, fewer lab tests).
- Interpretability: "the prediction depends on these 12 features" can be read and checked; "it depends on all 50,000" cannot.
- Compression and speed: store only the non-zero entries (an index and a value each) and compute a prediction with only that many multiplications.
- A statistical bet: if the truth really involves few features, a sparse model needs far fewer examples to learn it than a dense one.
The weight vector $\mathbf{w}=(0,\,3,\,0,\,0,\,-2,\,0,\,0,\,0.5)$ has $p=8$ entries but only $k=3$ non-zero: $\|\mathbf{w}\|_0=3$. Stored as (index, value) pairs: $(2,3),(5,-2),(8,0.5)$, that is $6$ numbers instead of $8$.
At scale. A text classifier has $p=10^6$ word weights and the lasso keeps $1\%$ of them: $k=10^4$. Dense storage takes $10^6$ numbers; sparse storage takes $2\times10^4$ (about $50$ times less), and one prediction needs $10^4$ multiplications instead of $10^6$ (about $100$ times fewer, when the inputs are stored sparsely too).
- The support of $\mathbf{w}$ is the set of indices $j$ with $w_j\ne0$. The sparsity is $\|\mathbf{w}\|_0=|\text{support}|$ (the number of non-zeros; not a true norm, but a handy count). A vector with $\|\mathbf{w}\|_0\le k$ is $k$-sparse.
- L1 regularization (the lasso, the elastic net with $\alpha>0$) produces sparse solutions; L2 does not. Larger $\lambda$ means a smaller support.
- Support recovery asks whether the support found equals the true one. Count true positives (useful features kept), false positives (useless features kept) and false negatives (useful features dropped).
- Limits. For $p>n$ the lasso can select at most $n$ features. Under suitable conditions (few true features, features not too correlated, a large enough $\lambda$) it can find the true support, but in practice the support found by cross-validation usually also contains some false positives.
- Also sparse in other senses (awareness): pruned neural networks (set small weights to zero), sparse autoencoders (few active units), and low-rank matrices (the nuclear norm is "L1 on the singular values").
Why do we need it?
Many problems have thousands of candidate inputs and only a few matter. Sparsity lets the model discard the rest, which makes it cheaper to run, easier to read and often more accurate.
Where is it used?
Biomarker discovery from gene-expression data, sparse linear models for ad click prediction and text, magnetic-resonance imaging with compressed sensing, and pruning and sparse autoencoders in deep learning.
How is it used?
Fit an L1 or elastic-net model, choose $\lambda$ by cross-validation, then read the support. For a cleaner model, refit the selected features without the penalty. Treat the selected set as a hypothesis to check, not as proof of cause.
Sparse does not mean true. The lasso can drop a useful feature that is correlated with a kept one, or keep a useless one when $\lambda$ is small. A feature is not "proven important" because it survived.
Large weights are shrunk too. The surviving weights are biased toward zero. If you need unbiased estimates of the survivors, refit them without the penalty (after selection), accepting that the selection step itself was noisy.
If the truth is dense, sparsity hurts. When many features each contribute a little, ridge usually beats the lasso. A common rule of thumb is to "bet on sparsity" only when you believe in it, or to let cross-validation decide between lasso, ridge and the elastic net.
Quick check: a lasso model has $p=5{,}000$ features and keeps $50$. How many numbers does sparse (index, value) storage need, and what fraction of the dense storage is that?
$50$ pairs is $2\times50=100$ numbers, versus $5{,}000$ for the dense vector: $100/5000=2\%$, a $50\times$ saving.
Weight decay: L2 regularization inside gradient descent core
What we need from earlier chapters: the gradient-descent update $\mathbf{w}\leftarrow\mathbf{w}-\eta\nabla L$ (Chapter 3.3).
Picture each weight as water in a slightly leaky bucket. At every training step, a little water leaks out (every weight shrinks by a tiny percentage). The data then pours water back in wherever it is needed (the gradient step). Weights that the data keeps asking for stay full; weights nobody needs slowly drain away. The final level is where leak and refill balance.
This leak is called weight decay. It is the same thing as adding an L2 penalty to the loss, seen from the optimizer's side.
One weight, loss $L(w)=\tfrac12(w-3)^2$ (the data wants $w=3$), decay $\lambda=1$, step size $\eta=0.2$, start at $w_0=0$. The update is $w\leftarrow(1-\eta\lambda)\,w-\eta L'(w)=0.8\,w-0.2\,(w-3)$.
- $w_1=0.8\cdot0-0.2(0-3)=0.6$.
- $w_2=0.8\cdot0.6-0.2(0.6-3)=0.48+0.48=0.96$.
- $w_3=0.8\cdot0.96-0.2(0.96-3)=0.768+0.408=1.176$.
Where does it settle? At a fixed point $\lambda w+(w-3)=0$, so $w=3/(1+\lambda)=1.5$. Each step shrinks the distance to $1.5$ by the factor $1-\eta(1+\lambda)=0.6$ ($0.9\to0.54\to0.324$, matching $1.5-w_k$ above). And $1.5$ is exactly the ridge answer $a/(1+\lambda)$ from earlier.
For the objective $F(\mathbf{w})=L(\mathbf{w})+\tfrac\lambda2\|\mathbf{w}\|^2$ the gradient is $\nabla L+\lambda\mathbf{w}$, so a gradient step is $$\mathbf{w}\leftarrow\mathbf{w}-\eta\big(\nabla L(\mathbf{w})+\lambda\mathbf{w}\big)=\underbrace{(1-\eta\lambda)\,\mathbf{w}}_{\text{decay}}-\eta\nabla L(\mathbf{w}).$$ Weight decay means "first multiply by $(1-\eta\lambda)$, then take the usual step". For plain gradient descent and SGD, weight decay and L2 regularization are the same algorithm.
- Where it settles. The fixed points of the update are the stationary points of $F$. For least squares the unique fixed point is the ridge solution $(X^\top X+\lambda I)^{-1}X^\top\mathbf{y}$ (the widget below shows the path ending exactly there).
- It helps optimization too. $F$ has curvatures in $[\mu+\lambda,\ L+\lambda]$ if $L$ has curvatures in $[\mu,L]$, so the condition number $\frac{L+\lambda}{\mu+\lambda}$ is smaller. Gradient descent is stable for $\eta<2/(L+\lambda)$.
- Conventions. PyTorch's
SGD(weight_decay=wd)addswd * wto the gradient, which is our $\lambda$ with the penalty $\tfrac\lambda2\|\mathbf{w}\|^2$. With momentum, the decay term enters the velocity along with the gradient. - Not the same for adaptive optimizers. In Adam, adding $\lambda\mathbf{w}$ to the gradient sends it through the division by $\sqrt{\hat v}+\epsilon$ together with the data gradient, so the actual shrinkage of each weight depends on the size of its gradients, and no longer looks like a uniform decay. AdamW repairs this by decoupling the decay: $\mathbf{w}\leftarrow\mathbf{w}-\eta\big(\hat{\mathbf{m}}/(\sqrt{\hat{\mathbf{v}}}+\epsilon)+\lambda\mathbf{w}\big)$. The decay is applied straight to the weights, as in plain SGD. (You will meet Adam, AdamW and why decoupling matters in Chapter 3.4, and weight decay in deep networks in Chapter 3.15.)
- What gets decayed. In practice biases and normalisation-layer parameters are usually excluded from weight decay, for the same reason the intercept is excluded from ridge.
Why do we need it?
It gives L2 regularization to any model trained by gradient steps, with one extra multiplication per step and no change to the loss code. It also keeps weights from growing without limit.
Where is it used?
Almost every neural-network recipe: SGD with momentum and weight decay for image models, AdamW for Transformers and large language models, plus the weight_decay argument of PyTorch and TensorFlow optimizers.
How is it used?
Set weight_decay (typical values $10^{-5}$ to $10^{-1}$ depending on the optimizer), exclude biases and norm parameters, and tune it with validation like any $\lambda$. With Adam prefer AdamW.
L2-in-the-loss equals weight decay only for plain (S)GD. With Adam and other adaptive methods the two behave differently; that is why AdamW exists. If you use Adam, check whether weight_decay is coupled (Adam) or decoupled (AdamW).
Step size and decay interact. The per-step shrink is $1-\eta\lambda$; changing the learning rate (or using a schedule) changes how much decay you actually apply per step. Large-scale training recipes tune them together.
Quick check: $w=1.0$, $\eta=0.1$, $\lambda=0.01$, $\nabla L=0$ at this point. What does one step with weight decay do?
$w\leftarrow(1-\eta\lambda)w-\eta\cdot0=(1-0.001)\times1.0=0.999$. With zero data gradient the weight just decays by $0.1\%$ per step. After $1000$ such steps it would be $0.999^{1000}\approx0.37$.
Early stopping: regularization with no penalty at all awareness
Here is a surprise: you can regularize without adding anything to the loss. Just stop training early.
Gradient descent started from zero does not learn everything at the same speed. It learns the strong directions of the data first (the big, clear patterns) and the weak directions last (the faint ones, where noise dominates). If you stop before it has finished the weak directions, those weights are still small: exactly what ridge regression does to weak directions, by shrinking them. Time plays the role of $1/\lambda$: more steps, weaker penalty.
This is called implicit regularization, because the penalty is never written down. It lives in the algorithm.
Two directions with strengths $\sigma_1^2=4$ and $\sigma_2^2=0.25$, step size $\eta=0.1$, started from $\mathbf{w}=\mathbf{0}$. After $t=10$ steps gradient descent has learned this fraction of each OLS component:
- strong direction: $1-(1-0.1\cdot4)^{10}=1-0.6^{10}=0.994$ (almost complete);
- weak direction: $1-(1-0.1\cdot0.25)^{10}=1-0.975^{10}=0.224$ (barely started).
Ridge with $\lambda=\frac{1}{\eta t}=1$ gives $\frac{4}{4+1}=0.8$ and $\frac{0.25}{0.25+1}=0.2$. Not identical (the strong direction is not shrunk as much by early stopping), but the same story: the weak direction is crushed, the strong one survives.
Derivation for least squares. Run gradient descent on $\tfrac12\|\mathbf{y}-X\mathbf{w}\|^2$ from $\mathbf{w}_0=\mathbf{0}$: $\mathbf{w}_{k+1}=\mathbf{w}_k-\eta X^\top(X\mathbf{w}_k-\mathbf{y})$. Using the SVD $X=U\Sigma V^\top$ and writing $u_{k,i}=\mathbf{v}_i^\top\mathbf{w}_k$ for the coordinate along $\mathbf{v}_i$, with $\hat u_i=\mathbf{u}_i^\top\mathbf{y}/\sigma_i$ the OLS value, $$u_{k+1,i}-\hat u_i=(1-\eta\sigma_i^2)\,(u_{k,i}-\hat u_i)\quad\Longrightarrow\quad u_{t,i}=\big[1-(1-\eta\sigma_i^2)^t\big]\,\hat u_i .$$ This converges for $0<\eta<2/\sigma_1^2$. Compare with ridge: $u^{\text{ridge}}_i=\frac{\sigma_i^2}{\sigma_i^2+\lambda}\hat u_i$. Both factors rise from $0$ to $1$ as $\sigma_i^2$ grows, with the transition near $\sigma_i^2\approx\frac{1}{\eta t}$ for early stopping and $\sigma_i^2\approx\lambda$ for ridge. So, as a rule of thumb (same order of magnitude, not an identity), $\lambda\approx\dfrac{1}{\eta t}$.
Early stopping in practice. After each epoch, measure the loss on a validation set. Stop when it has not improved for a number of epochs (the patience) and keep the weights from the best epoch. The number of steps becomes the regularization hyperparameter, and one training run gives the whole "path", just as one pass gives all the $\lambda$'s in a regularization path.
Honest scope. The exact connection to ridge holds for gradient descent on least squares from zero. For deep networks early stopping is an effective, widely used heuristic whose regularizing effect is only partly understood.
Why do we need it?
Training longer lowers the training loss but can raise the validation loss. Early stopping gives a free, effective regularizer and also saves compute.
Where is it used?
Neural-network training (Keras EarlyStopping, PyTorch Lightning callbacks), gradient-boosted trees (stop adding trees when validation error stops improving) and iterative solvers for ill-posed problems.
How is it used?
Hold out validation data, log the validation loss each epoch, stop after a few epochs without improvement, and restore the best weights. Combine with a small explicit weight decay if you like.
Early stopping is not "free". You still need validation data, and the best stopping time depends on the learning rate, schedule and batch size. Changing the learning rate changes where "early" is.
It does not replace explicit regularization. Many strong models use both: weight decay, dropout and early stopping together.
Quick check: you halve the learning rate $\eta$. To get a similar amount of regularization from early stopping, should you train for fewer or more steps?
The effective penalty is $\lambda\approx1/(\eta t)$, so keeping $\eta t$ constant keeps $\lambda$ constant. Halving $\eta$ means you need twice as many steps to reach the same effective regularization.
Dropout and data augmentation: noise as a regularizer awareness
Two more ways to regularize, and both work by injecting noise into training.
- Dropout. During each training step, randomly switch off a fraction $p$ of the units. No unit can rely on a particular partner being there, so the units learn to be useful on their own. It is like a team where a random few members stay home each day: everyone must be able to do the job.
- Data augmentation. Show the model slightly modified copies of the training data (a flipped, shifted or noisy image) with the same label. The model learns that those changes do not matter, which is a kind of built-in prior knowledge.
Both can be seen as penalties in disguise. For a linear model one can work out the penalty exactly.
Dropout, a numerical check. Four units with outputs $\mathbf{a}=(2,1,3,4)$, drop probability $p=0.5$ and inverted dropout: surviving outputs are multiplied by $\frac1{1-p}=2$ so the average stays right. Keep two units at random (the six possible masks):
- keep units $\{1,2\}$: sum $=(2+1)\times2=6$ keep $\{1,3\}$: $(2+3)\times2=10$ keep $\{1,4\}$: $12$
- keep $\{2,3\}$: $(1+3)\times2=8$ keep $\{2,4\}$: $10$ keep $\{3,4\}$: $14$
Average over the six masks $=(6+10+12+8+10+14)/6=10$, the full sum $2+1+3+4$. Any single mask is noisy (6 to 14) but the expectation is exact.
Noise on the inputs, a numerical check. One feature, no intercept, three points $x=(1,2,3)$ with $y=(1,2,3)$, so the least-squares slope is $\sum xy/\sum x^2=14/14=1$. Add noise of variance $\sigma^2=1$ to each $x$. The expected denominator grows from $\sum x^2=14$ to $\sum x^2+n\sigma^2=14+3=17$, so the slope becomes $14/17\approx0.82$: the noise shrinks the weight, exactly like ridge with $\lambda=n\sigma^2=3$.
Dropout. Draw independent masks $m_j\sim\mathrm{Bernoulli}(1-p)$ and use $\tilde a_j=m_ja_j/(1-p)$ in training; at test time use $a_j$ unchanged. Then $\mathbb{E}\tilde a_j=a_j$ and $\mathrm{Var}\,\tilde a_j=\frac{p}{1-p}a_j^2$. For a linear model $\hat y=\sum_jw_jx_jm_j/(1-p)$ the average training loss over masks is $$\mathbb{E}_{\mathbf{m}}\big[(y-\hat y)^2\big]=(y-\mathbf{w}^\top\mathbf{x})^2+\frac{p}{1-p}\sum_jw_j^2x_j^2 ,$$ because the mean of $\hat y$ is $\mathbf{w}^\top\mathbf{x}$ and its variance is $\sum_jw_j^2x_j^2\frac{p}{1-p}$. So dropout acts like an L2 penalty on the weights, scaled by how large each input is (an adaptive ridge); for standardised features it is ridge with strength proportional to $\frac{p}{1-p}$. For deep networks dropout remains a useful heuristic with a richer (not fully understood) effect.
Data augmentation by input noise. If the model sees $\mathbf{x}+\boldsymbol\varepsilon$ with $\mathbb{E}\boldsymbol\varepsilon=\mathbf{0}$ and $\mathrm{Cov}\,\boldsymbol\varepsilon=\sigma^2I$, then for a linear model $$\mathbb{E}\big[(y-\mathbf{w}^\top(\mathbf{x}+\boldsymbol\varepsilon))^2\big]=(y-\mathbf{w}^\top\mathbf{x})^2+\sigma^2\|\mathbf{w}\|^2,$$ since the cross term has zero mean and $\mathbb{E}(\mathbf{w}^\top\boldsymbol\varepsilon)^2=\sigma^2\|\mathbf{w}\|^2$. Summed over $n$ examples and halved, this is ridge with $\lambda=n\sigma^2$. Other augmentations (flips, crops, mixup) are not literally L2 penalties but play the same role: they tell the model which changes it should be insensitive to.
Why do we need it?
Deep networks have far more weights than examples and can memorise the training set. Noise during training stops them from relying on fragile coincidences, with no change to the loss function.
Where is it used?
Dropout layers in fully connected nets and Transformers, random crops, flips and colour jitter in image training, SpecAugment for audio, back-translation and word dropout for text.
How is it used?
Add Dropout(p) layers (typical $p$ from $0.1$ to $0.5$), switch them off at evaluation, and define augmentations that keep the label correct. Tune $p$ and the augmentation strength by validation.
Switch dropout off at test time (or use inverted dropout as above and then just run the full network). Forgetting this makes predictions random and noisy.
Augmentations must preserve the label. Flipping a photo of a cat is fine; flipping the digit "6" can turn it into something else, and rotating "6" by 180 degrees turns it into a "9".
Penalty equivalences are for linear models. For deep networks these are useful intuitions, not exact statements.
Quick check: inputs are standardised, noise of variance $\sigma^2=0.25$ is added to each input, and there are $n=100$ examples. Which ridge $\lambda$ is this equivalent to, for a linear model with loss $\tfrac12\sum(\cdot)^2$?
$\lambda=n\sigma^2=100\times0.25=25$. (Equivalently, the penalty is $\tfrac\lambda2\|\mathbf{w}\|^2=12.5\|\mathbf{w}\|^2$ against a loss summed over 100 examples.)
Recap, cheat sheet and practice
- Principle. Minimise $L(\mathbf{w})+\lambda R(\mathbf{w})$. The penalty fights overfitting, makes ill-posed problems unique and stable, and encodes prior beliefs (L2 is a Gaussian prior with $\lambda=\sigma^2/\tau^2$, L1 a Laplace prior). $\lambda$ is a hyperparameter chosen on held-out data, never on the training loss.
- L2 (ridge). Penalty $\tfrac\lambda2\|\mathbf{w}\|^2$, gradient $\lambda\mathbf{w}$ (a spring: the pull fades near zero). Closed form $(X^\top X+\lambda I)\mathbf{w}=X^\top\mathbf{y}$; each SVD direction is shrunk by $\sigma^2/(\sigma^2+\lambda)$; the condition number becomes $(\sigma_1^2+\lambda)/(\sigma_p^2+\lambda)$. Shrinks, never exactly zero.
- L1 (lasso). Penalty $\lambda\|\mathbf{w}\|_1$, constant pull $\lambda\,\mathrm{sign}(w)$, a kink at $0$. No closed form in general. Optimality: $\mathbf{x}_j^\top\mathbf{r}=\lambda\,\mathrm{sign}(\hat w_j)$ for kept features and $|\mathbf{x}_j^\top\mathbf{r}|\le\lambda$ for dropped ones; $\hat{\mathbf{w}}=\mathbf{0}$ for $\lambda\ge\lambda_{\max}=\max_j|\mathbf{x}_j^\top\mathbf{y}|$. For orthonormal $X$: $\hat w_j=\mathrm{soft}(a_j,\lambda)=\mathrm{sign}(a_j)\max(|a_j|-\lambda,0)$ with $a=X^\top\mathbf{y}$.
- Elastic net. $\lambda[\alpha\|\mathbf{w}\|_1+\frac{1-\alpha}2\|\mathbf{w}\|^2]$: zeros like the lasso, plus grouping, uniqueness and stability. Orthonormal case: $\mathrm{soft}(a_j,\lambda\alpha)/(1+\lambda(1-\alpha))$.
- Constraint view. A penalised problem solves the budget problem $\min L$ s.t. $R\le r$ for $r=R(\hat{\mathbf{w}})$ (elementary proof); for convex problems the converse holds with $\lambda$ the Lagrange multiplier and $\lambda=-dL^\star/dr$. The diamond has corners on the axes, so a whole wedge of targets lands on a corner (a zero weight); the disc has none. It is geometry, not magic: the data decides which weights vanish.
- Generalization. Error $=$ noise $+$ bias$^2+$ variance. Larger $\lambda$: more bias, less variance; the test error is U-shaped. Choose $\lambda$ on a log grid by validation or $K$-fold cross-validation (optionally the one-standard-error rule), refit, and use the test set once.
- Paths and sparsity. The ridge path is smooth; the lasso path is piecewise linear with features entering one at a time. Sparsity gives feature selection, interpretability and compression, but selected features are not proven causes, and survivors are shrunk.
- Relatives. Weight decay $=(1-\eta\lambda)\mathbf{w}-\eta\nabla L$ is L2 for plain (S)GD (AdamW decouples it, Chapters 3.4 and 3.15). Early stopping from zero acts like ridge with $\lambda\approx1/(\eta t)$. Dropout and input-noise augmentation are, for linear models, adaptive and ordinary ridge penalties.
Cheat sheet
| Penalty $R(\mathbf{w})$ | Pull on $w_j$ | Solution (orthonormal $X$, $a=X^\top\mathbf{y}$) | Exact zeros? | Region $R\le r$ | |
|---|---|---|---|---|---|
| None | $0$ | $0$ | $a_j$ | no | everything |
| Ridge (L2) | $\tfrac12\|\mathbf{w}\|_2^2$ | $\lambda w_j$ (fades) | $a_j/(1+\lambda)$ | no | disc (smooth) |
| Lasso (L1) | $\|\mathbf{w}\|_1$ | $\lambda\,\mathrm{sign}(w_j)$ (constant) | $\mathrm{soft}(a_j,\lambda)$ | yes | diamond (corners on axes) |
| Elastic net | $\alpha\|\mathbf{w}\|_1+\frac{1-\alpha}2\|\mathbf{w}\|_2^2$ | $\lambda[\alpha\,\mathrm{sign}+(1-\alpha)w_j]$ | $\mathrm{soft}(a_j,\lambda\alpha)/(1+\lambda(1-\alpha))$ | yes | rounded diamond |
| Ridge, general $X$ | $(X^\top X+\lambda I)^{-1}X^\top\mathbf{y}$, factor $\frac{\sigma^2}{\sigma^2+\lambda}$ per direction | ||||
| Lasso, general $X$ | coordinate descent: $w_j\leftarrow\mathrm{soft}(\rho_j,\lambda)/\|\mathbf{x}_j\|^2$ (Chapter 3.13) | ||||
| Weight decay | $\mathbf{w}\leftarrow(1-\eta\lambda)\mathbf{w}-\eta\nabla L$ | ||||
| Early stopping | GD factor $1-(1-\eta\sigma^2)^t$, so $\lambda\approx1/(\eta t)$ | ||||
| Best $\lambda$, one direction | $\lambda^\star=\sigma_{\text{noise}}^2/w_{\text{true}}^2$ |
import numpy as np
soft = lambda a, t: np.sign(a) * np.maximum(np.abs(a) - t, 0) + 0.0 # soft-thresholding
# ---------- 1. ridge: closed form, SVD shrinkage factors, two equivalent formulas ----------
X = np.array([[2., 0.], [0., 1.]]); y = np.array([4., 3.]); lam = 1.0
w_ols = np.linalg.solve(X.T @ X, X.T @ y)
w_ridge = np.linalg.solve(X.T @ X + lam * np.eye(2), X.T @ y)
print(w_ols, w_ridge) # [2. 3.] [1.6 1.5]
s = np.linalg.svd(X, compute_uv=False) # singular values [2. 1.]
print(s**2 / (s**2 + lam)) # [0.8 0.5] shrinkage factors
rng = np.random.default_rng(0)
A = rng.normal(size=(30, 5)); b = rng.normal(size=30)
w1 = np.linalg.solve(A.T @ A + lam * np.eye(5), A.T @ b) # p x p system
w2 = A.T @ np.linalg.solve(A @ A.T + lam * np.eye(30), b) # n x n system
print(np.allclose(w1, w2)) # True: the two forms agree
s = np.array([10., 0.1]) # singular values of an ill-conditioned X
kappa = lambda l: (s[0]**2 + l) / (s[1]**2 + l) # condition number of X^T X + lam I
print(round(kappa(0), 1), round(kappa(1), 1), round(kappa(10), 2)) # 10000.0 100.0 10.99
# ---------- 2. soft-thresholding: lasso, ridge and elastic net for orthonormal X ----------
a = np.array([3., -0.5, 1.2]); lam = 1.0
print(soft(a, lam)) # [2. 0. 0.2] lasso
print(a / (1 + lam)) # [ 1.5 -0.25 0.6 ] ridge
print(np.round(soft(a, lam * 0.5) / (1 + lam * 0.5), 3)) # [1.667 0. 0.467] elastic net, alpha = 0.5
# ---------- 3. lasso by coordinate descent, and a check of the optimality conditions ----------
def lasso_cd(X, y, lam, sweeps=1000):
n, p = X.shape; w = np.zeros(p); r = y.copy(); cn = (X**2).sum(0)
for _ in range(sweeps):
for j in range(p):
rho = X[:, j] @ r + cn[j] * w[j] # correlation with the partial residual
new = soft(rho, lam) / cn[j]
r += X[:, j] * (w[j] - new); w[j] = new
return w
rng = np.random.default_rng(1)
X = rng.normal(size=(60, 8)); X = (X - X.mean(0)) / X.std(0)
beta = np.array([3., -2., 0., 0., 1.5, 0., 0., 0.])
y = X @ beta + rng.normal(size=60); y = y - y.mean()
lam_max = np.abs(X.T @ y).max()
w = lasso_cd(X, y, 0.3 * lam_max)
print(np.round(w, 2)) # [ 2.19 -0.58 0. 0. 0.58 0. 0. 0. ] (shrunk, 3 of 8 non-zero)
c = X.T @ (y - X @ w) # correlation of each feature with the residual
lam = 0.3 * lam_max
print(np.allclose(np.abs(c[w != 0]), lam), bool((np.abs(c[w == 0]) <= lam + 1e-8).all())) # True True
print(np.abs(lasso_cd(X, y, 1.001 * lam_max)).max()) # 0.0 : above lambda_max everything is zero
# ---------- 4. penalised = constrained (Claim 1): nothing inside the budget beats w_hat ----------
f = lambda w: 0.5 * np.sum((y - X[:, :2] @ w)**2)
w_hat = lasso_cd(X[:, :2], y, 100.0) # penalised answer using only the first two features
r_budget = np.abs(w_hat).sum()
print(np.round(w_hat, 3), round(r_budget, 3)) # [1.571 0. ] 1.571 (second weight exactly 0: a corner)
cand = rng.uniform(-4, 4, size=(200000, 2))
cand = cand[np.abs(cand).sum(1) <= r_budget][:20000] # random points inside ||w||_1 <= r
print(bool(min(f(c_) for c_ in cand) >= f(w_hat) - 1e-9)) # True: nothing inside the budget beats w_hat
# ---------- 5. weight decay in gradient descent ends at the ridge solution ----------
X2 = np.array([[2., 0.], [0., 1.]]); y2 = np.array([4., 3.]); lam = 1.0; eta = 0.1
w = np.zeros(2)
for _ in range(300):
w = (1 - eta * lam) * w - eta * X2.T @ (X2 @ w - y2) # decay, then gradient step
print(np.round(w, 4)) # [1.6 1.5] = the ridge answer from part 1
# ---------- 6. early stopping versus ridge: shrinkage factors (eta = 0.1, t = 10 steps) ----------
eta, t = 0.1, 10
for s2 in (4.0, 0.25):
print(s2, round(1 - (1 - eta * s2)**t, 3), round(s2 / (s2 + 1 / (eta * t)), 3)) # 4.0 0.994 0.8 then 0.25 0.224 0.2
# ---------- 7. bias-variance for one direction: a = sigma_i^2, true weight w, noise sigma ----------
a_, w_, sig = 4.0, 1.0, 1.0
mse = lambda l: (l**2 * w_**2 + a_ * sig**2) / (a_ + l)**2
print([round(mse(l), 4) for l in (0, 0.5, 1, 2, 4)]) # [0.25, 0.2099, 0.2, 0.2222, 0.3125] minimum at lambda = sigma^2 / w^2 = 1
1. You fit ridge regression for $\lambda\in\{0,\ 0.01,\ 0.1,\ 1,\ 10\}$ and choose the $\lambda$ with the smallest training error. Which $\lambda$ do you get, and is that a good method?
2. Orthonormal features, least-squares weights $\mathbf{a}=[0.7,\ -2.0,\ 1.2]$, lasso with $\lambda=1$. The lasso weights are…
3. Why does an L1 penalty set weights exactly to zero, while an L2 penalty only makes them small?
4. In ridge regression, a direction of the data has $\sigma^2=0.5$ and $\lambda=1.5$. By what factor is its least-squares component shrunk?
5. As the regularization strength $\lambda$ increases, what typically happens to the bias and the variance of the fitted model?
6. Two features are exact copies of each other, and the loss only depends on $w_1+w_2$. The lasso gets a tie among many splits. What does the elastic net do and why?
Practice problems
A. $X=\begin{bmatrix}1&1\\1&-1\end{bmatrix}$, $\mathbf{y}=[3,\,1]^\top$. Find the least-squares solution and the ridge solution for $\lambda=2$. What is the shrinkage factor?
$X^\top X=\begin{bmatrix}2&0\\0&2\end{bmatrix}=2I$ and $X^\top\mathbf{y}=[1\cdot3+1\cdot1,\ 1\cdot3-1\cdot1]=[4,\,2]$. Least squares: $\mathbf{w}=\frac12[4,2]=[2,\,1]$. Ridge: $(2I+2I)\mathbf{w}=[4,2]$, so $\mathbf{w}=\frac14[4,2]=[1,\,0.5]$. Both directions have $\sigma^2=2$, so the factor is $\frac{2}{2+2}=0.5$ for both, and indeed $[2,1]\times0.5=[1,0.5]$ ✓.
B. Orthonormal features, $\mathbf{a}=[2.4,\,-0.3,\,0.9,\,-1.5]$, $\lambda=1$. Give the lasso, ridge and elastic-net ($\alpha=0.5$) solutions. How many zeros does each have?
Lasso: $\mathrm{soft}(a,1)=[1.4,\ 0,\ 0,\ -0.5]$ ($0.9\le1$ is killed): 2 zeros. Ridge: $a/2=[1.2,\ -0.15,\ 0.45,\ -0.75]$: 0 zeros. Elastic net: soft-threshold by $\lambda\alpha=0.5$ then divide by $1+\lambda(1-\alpha)=1.5$: $\mathrm{soft}(a,0.5)=[1.9,\ 0,\ 0.4,\ -1.0]$, divided by $1.5$: $[1.267,\ 0,\ 0.267,\ -0.667]$: 1 zero. Elastic net sits between the two.
C. $L(w)=\tfrac12(w-5)^2$ with an L1 penalty $\lambda=2$. Find the penalised answer, the equivalent budget $r$, and $dL^\star/dr$ at that budget. What is the ridge answer ($\tfrac\lambda2w^2$) for the same $\lambda$?
Penalised: $\mathrm{soft}(5,2)=3$. The budget $|w|\le r$ with $r=3$ gives the same point (best allowed point is the boundary $3$). The best loss with budget $r$ is $L^\star(r)=\tfrac12(5-r)^2$, so $dL^\star/dr=-(5-r)=-2=-\lambda$ ✓. Ridge: $w=5/(1+\lambda)=5/3\approx1.667$ (not zero, and a different budget, $r\approx1.667$).
D. Features $\mathbf{x}_1=[1,1,1,1]$, $\mathbf{x}_2=[1,-1,0,1]$, $\mathbf{x}_3=[0,1,1,-1]$ and targets $\mathbf{y}=[2,-1,3,0]$. Find $\lambda_{\max}$, the feature that enters first, and the lasso solution at $\lambda=3$.
Correlations: $\mathbf{x}_1^\top\mathbf{y}=2-1+3+0=4$; $\mathbf{x}_2^\top\mathbf{y}=2+1+0+0=3$; $\mathbf{x}_3^\top\mathbf{y}=0-1+3+0=2$. So $\lambda_{\max}=4$ and feature 1 enters first. At $\lambda=3$ try only feature 1 active: $\hat w_1=(\mathbf{x}_1^\top\mathbf{y}-\lambda)/\|\mathbf{x}_1\|^2=(4-3)/4=0.25$. Residual $\mathbf{r}=\mathbf{y}-0.25\mathbf{x}_1=[1.75,-1.25,2.75,-0.25]$. Check the others: $\mathbf{x}_2^\top\mathbf{r}=1.75+1.25+0-0.25=2.75\le3$ ✓ and $\mathbf{x}_3^\top\mathbf{r}=-1.25+2.75+0.25=1.75\le3$ ✓. So the solution is $[0.25,\,0,\,0]$.
E. One direction with strength $a=\sigma_i^2=1$, true weight $w=1$, noise $\sigma=1$. Compute the expected squared error of the least-squares estimate and of ridge with $\lambda=1$ (the best $\lambda$).
Expected error of ridge in one direction: $\frac{\lambda^2w^2+a\sigma^2}{(a+\lambda)^2}$. OLS ($\lambda=0$): $\frac{0+1}{1}=1$. Ridge $\lambda=1$: $\frac{1+1}{4}=0.5$. Check via bias and variance: $s=\frac{1}{2}$, bias $=-0.5$ so bias$^2=0.25$, variance $=s^2\sigma^2/a=0.25$, total $0.5$ ✓. And $\lambda^\star=\sigma^2/w^2=1$. Regularization halved the error.
F. SGD with weight decay $\lambda=0.5$ and $\eta=0.01$. (i) With zero data gradient, how large is a weight of $4$ after 100 steps? (ii) If you rely on early stopping alone instead, about how many steps $t$ mimic a ridge strength of $0.5$ at this $\eta$?
(i) Each step multiplies by $1-\eta\lambda=0.995$. After 100 steps: $4\times0.995^{100}\approx4\times0.6058=2.42$. (ii) $\lambda\approx\frac{1}{\eta t}$ gives $t\approx\frac{1}{\eta\lambda}=\frac{1}{0.01\times0.5}=200$ steps (a rule of thumb, correct to within a factor of order one).
Second-Order Optimization
Gradient descent only asks "which way is downhill?". Newton's method also asks "how sharply does the ground curve?" and uses the answer to choose both direction and step length. It can solve a bowl in one step. The price is the cost of the curvature, and the rest of this chapter is about paying less of it: quasi-Newton methods (BFGS) and the memory-light L-BFGS.
- Derive Newton's method by minimising the local quadratic model, and see Newton–Raphson as root-finding for $f'=0$
- Compute the Newton update $\mathbf{x}\leftarrow\mathbf{x}-H^{-1}\nabla f$ by solving $H\mathbf{d}=-\nabla f$, and see why it solves a quadratic in one step
- State the advantages (scale invariance, quadratic convergence) and the disadvantages (cost, indefinite Hessians, divergence far from the solution) and the standard fixes: damping, line search, Levenberg shifts
- Count the cost of Newton's method for $n=10^3,10^6,10^9$ parameters and understand why deep learning does not use it directly
- Derive the secant condition and the BFGS update, and see why it keeps the curvature estimate positive definite
- Use L-BFGS: store only the last $m$ pairs, run the two-loop recursion, and know when it is the right tool
- Recognise Gauss–Newton, natural gradient and K-FAC as relatives of the same idea
Newton–Raphson: finding where a function is zero core
What we need from earlier chapters: the derivative as the slope of the tangent line (Chapter 2.3) and linearization (Chapter 2.12): near a point, a smooth curve looks like its tangent line.
Suppose you want the number $x$ where a curve $y=f(x)$ crosses zero (a root). You cannot solve for it directly, but you can stand somewhere on the curve and do the following:
- Draw the tangent line at your position. Close to you, the tangent is almost the same as the curve.
- The tangent is a straight line, so it is easy to see exactly where it crosses zero.
- Jump to that crossing. It is not the root, but it is usually much closer.
- Repeat from there.
This is Newton–Raphson. It is the same trick as "zoom in until the curve looks straight", used as an algorithm. And finding a minimum is a root-finding problem too: a smooth minimum is where the slope $f'$ is zero. So Newton's method for optimization (next section) is Newton–Raphson applied to the derivative.
Computing $\sqrt2$. $\sqrt2$ is the positive root of $f(x)=x^2-2$, and $f'(x)=2x$. The update is $x\leftarrow x-\dfrac{x^2-2}{2x}=\dfrac{x+2/x}{2}$ (the average of $x$ and $2/x$). Start at $x_0=1$:
- $x_1=1-\dfrac{1-2}{2}=1.5$.
- $x_2=1.5-\dfrac{2.25-2}{3}=1.416667$.
- $x_3=1.416667-\dfrac{0.006944}{2.833333}=1.4142157$.
- $x_4=1.41421356237\ldots$, agreeing with $\sqrt2=1.41421356237\ldots$ to 12 digits.
| step | $x_k$ | error $|x_k-\sqrt2|$ |
|---|---|---|
| 0 | 1 | $4\times10^{-1}$ |
| 1 | 1.5 | $8.6\times10^{-2}$ |
| 2 | 1.416667 | $2.5\times10^{-3}$ |
| 3 | 1.4142157 | $2.1\times10^{-6}$ |
| 4 | 1.41421356237 | $1.6\times10^{-12}$ |
Look at the exponents: $-1,\ -2.6,\ -5.7,\ -11.8$. The number of correct digits roughly doubles every step. That is called quadratic convergence.
Newton–Raphson for $f(x)=0$. At $x_k$ the tangent line is $y=f(x_k)+f'(x_k)(x-x_k)$. Set $y=0$ and solve for $x$:
$$x_{k+1}=x_k-\frac{f(x_k)}{f'(x_k)}\qquad(f'(x_k)\ne0).$$- Quadratic convergence. If $f$ is twice continuously differentiable, $f(x^\star)=0$, $f'(x^\star)\ne0$, and $x_0$ is close enough to $x^\star$, then the error satisfies $|x_{k+1}-x^\star|\approx\left|\dfrac{f''(x^\star)}{2f'(x^\star)}\right|\,|x_k-x^\star|^2$. For $x^2-2$ the constant is $\frac{2}{2\cdot2\sqrt2}=\frac{1}{2\sqrt2}\approx0.35$: from error $8.6\times10^{-2}$ it predicts $0.35\times0.0074=2.6\times10^{-3}$, and we saw $2.5\times10^{-3}$.
- It can fail. If $f'(x_k)=0$ the tangent is flat and never crosses zero. A bad start can send it far away, or round in circles: for $f(x)=x^3-2x+2$ starting at $x_0=0$ the iterates go $0\to1\to0\to1\to\cdots$ forever, though a root exists near $-1.769$.
- For optimization: to minimise $f$ find a zero of $g=f'$. Newton–Raphson on $g$ gives $x_{k+1}=x_k-\dfrac{f'(x_k)}{f''(x_k)}$.
Why do we need it?
Many equations have no formula for their solution. Newton–Raphson turns "solve $f(x)=0$" into a few cheap steps, each one solving a straight-line problem, and converges extremely fast near the answer.
Where is it used?
How calculators and numerical libraries compute square roots and reciprocals, finding roots in scipy.optimize.newton, interest-rate solvers, and (applied to the gradient) the training of logistic regression and many other models.
How is it used?
Pick a starting guess, compute $f$ and $f'$, update $x\leftarrow x-f/f'$, and stop when $|f(x)|$ or the step is tiny. Always cap the number of iterations, and check that $f'$ is not close to $0$.
Newton–Raphson finds a root; it does not know which one you wanted. It goes to whatever root the tangent steps lead to. With $e^x-3$ there is only one; with $x^2-2$ the start decides between $+\sqrt2$ and $-\sqrt2$.
Never divide blindly. Check $|f'(x)|$ before dividing, and put a cap on the number of steps.
Quick check: do one Newton–Raphson step for $f(x)=x^2-9$ from $x_0=5$.
$f(5)=16$ and $f'(5)=10$, so $x_1=5-\frac{16}{10}=3.4$. The root is $3$; the error went from $2$ to $0.4$. One more step: $f(3.4)=2.56$, $f'(3.4)=6.8$, $x_2=3.4-0.3765=3.0235$ (error $0.0235$).
Newton's method: jump to the bottom of the local parabola core
What we need from earlier chapters: the second derivative as curvature (Chapter 2.10), the second-order Taylor expansion (Chapter 2.11) and gradient descent (Chapter 3.3). The Calculus guide has a short first look at the method; here we derive it carefully and study when it works.
Gradient descent looks only at the slope: "downhill is that way, take a step of size $\eta\times$ slope". It has no idea how far away the bottom is, so you must guess a step size $\eta$ yourself.
Newton's method also reads the curvature. Near where you stand, it replaces the curve by the parabola that matches the curve's height, slope and curvature. A parabola has an obvious bottom, so you jump straight to it. Then you stand there, build a new parabola, and jump again.
- If the ground really is a parabola, you reach the bottom in one jump.
- The step length adjusts itself: where the curve is sharply curved the parabola is narrow and the jump is short; where it is nearly flat the parabola is wide and the jump is long.
Minimise $f(x)=\tfrac14x^4+\tfrac12x^2-x$. Then $f'(x)=x^3+x-1$ and $f''(x)=3x^2+1$. The minimum is where $x^3+x-1=0$, at $x^\star=0.682328\ldots$ Start at $x_0=0$ and use $x\leftarrow x-f'/f''$:
- $x_0=0$: $f'=-1$, $f''=1$, so $x_1=0-\frac{-1}{1}=1$.
- $x_1=1$: $f'=1$, $f''=4$, so $x_2=1-\frac14=0.75$.
- $x_2=0.75$: $f'=0.171875$, $f''=2.6875$, so $x_3=0.75-0.063953=0.686047$.
- $x_3$: $x_4=0.6823396$. Then $x_5=0.6823278039\ldots$
The errors are $0.32,\ 0.068,\ 3.7\times10^{-3},\ 1.2\times10^{-5},\ 1.5\times10^{-10}$: digits doubling. Plain gradient descent with a safe step $\eta=0.2$ needs 12 steps just to get within $0.001$.
Derivation. Near the current point $x$ use the second-order Taylor expansion (a parabola in the step $\delta$):
$$f(x+\delta)\approx m(\delta)=f(x)+f'(x)\,\delta+\tfrac12f''(x)\,\delta^2 .$$A stationary point of $m$ is where $m'(\delta)=f'(x)+f''(x)\,\delta=0$. So
$$\boxed{\delta=-\frac{f'(x)}{f''(x)},\qquad x_{\text{new}}=x-\frac{f'(x)}{f''(x)}}\qquad(f''(x)\ne0).$$- If $f''(x)>0$ the parabola opens up and $\delta$ is the bottom of the model: a sensible downhill jump.
- If $f''(x)<0$ the parabola opens down, and the same formula jumps to the top of the model. Newton is heading for a local maximum.
- If $f''(x)=0$ the model is a straight line and has no stationary point.
- Relation to gradient descent. A gradient step is $\delta=-\eta f'(x)$. Newton is the same recipe with the fixed guess $\eta$ replaced by $1/f''(x)$, the "perfect learning rate" for a parabola (Chapter 2.10).
- It is the same as Newton–Raphson applied to $g=f'$: the root of $f'$ is the stationary point.
- The model is only local. Far from $x$ the parabola may be a poor picture of $f$, so a long jump can overshoot. The jump has no built-in safety.
Why do we need it?
Gradient descent needs many small steps and a well-chosen step size. Using curvature chooses the step length for us and, near a minimum, converges extremely fast.
Where is it used?
Maximum-likelihood fitting in statistics software (generalised linear models are fitted by Newton's method, called iteratively reweighted least squares), logistic regression solvers, and the final polishing stage of many optimizers.
How is it used?
At each iterate compute $f'$ and $f''$, take $\delta=-f'/f''$, and repeat until $|f'|$ is tiny. Check that $f''>0$ and add a safeguard (damping or a line search) when the jump looks too long.
Newton finds a flat spot of its parabola, not necessarily a minimum. At a point of negative curvature it jumps toward a maximum or (in more dimensions) a saddle. You always need to check the sign of the curvature.
It is local. It converges to a nearby stationary point, so it can land in a local minimum that is not the global one (the second function above).
A long jump is not a good jump. Where the curve is almost flat, $f''$ is tiny and the parabola's bottom is far away; the true function may be nothing like the parabola out there. (Section "When Newton goes wrong".)
Quick check: $f(x)=\tfrac14x^4+\tfrac12x^2-x$. Take one Newton step from $x=2$.
$f'(2)=8+2-1=9$ and $f''(2)=12+1=13$. The step is $-9/13=-0.692$, so $x_{\text{new}}=2-0.692=1.308$. The gradient step with $\eta=0.2$ would have been $-0.2\times9=-1.8$, to $x=0.2$. Newton's step is shorter here because the curvature is large ($13$): it knows the ground rises steeply.
The Hessian and the Newton update: solve, never invert core
What we need from earlier chapters: the Hessian matrix (Chapter 2.10), positive definite matrices (Chapter 1.12) and solving a linear system (Chapter 1.7).
With many variables, "curvature" is no longer one number. The ground can curve sharply in one direction and gently in another, and the directions can interact (a tilted valley). All of that is stored in a matrix, the Hessian $H$: row $i$, column $j$ says how the slope in direction $i$ changes when you move in direction $j$.
The local picture is now a bowl (a paraboloid) instead of a parabola. Newton's method does exactly what it did in 1D: build the bowl that matches height, slope and curvature at your point, and jump to the bottom of the bowl. Finding the bottom of a bowl is a linear algebra problem: solve one linear system.
The gradient points straight uphill, which in a long tilted valley is not the direction of the bottom. The Newton direction corrects the gradient with the curvature matrix, and points at the bottom of the bowl.
Let $f(\mathbf{x})=\tfrac12\mathbf{x}^\top A\mathbf{x}-\mathbf{b}^\top\mathbf{x}$ with $A=\begin{bmatrix}4&1\\1&2\end{bmatrix}$, $\mathbf{b}=\begin{bmatrix}1\\1\end{bmatrix}$. Then $\nabla f=A\mathbf{x}-\mathbf{b}$ and the Hessian is the constant matrix $H=A$. Start at $\mathbf{x}_0=(2,2)$.
- Gradient: $\mathbf{g}=A\mathbf{x}_0-\mathbf{b}=[8+2-1,\ 2+4-1]=[9,\,5]$.
- Solve $H\mathbf{d}=-\mathbf{g}$, that is $4d_1+d_2=-9$ and $d_1+2d_2=-5$. From the second, $d_1=-5-2d_2$. Substituting: $4(-5-2d_2)+d_2=-9\Rightarrow-7d_2=11\Rightarrow d_2=-\tfrac{11}{7}$, and $d_1=-5+\tfrac{22}{7}=-\tfrac{13}{7}$.
- New point: $\mathbf{x}_1=(2-\tfrac{13}{7},\ 2-\tfrac{11}{7})=(\tfrac17,\ \tfrac37)$.
- Check: $\nabla f(\mathbf{x}_1)=[\tfrac47+\tfrac37-1,\ \tfrac17+\tfrac67-1]=[0,\,0]$ ✓. One step landed exactly on the minimum.
Gradient descent from the same point, with the safe step $\eta=1/L=1/4.414=0.227$ ($L$ is the largest eigenvalue of $A$), moves along $-\mathbf{g}=[-9,-5]$ and ends up at $(-0.04,\,0.87)$, still $0.47$ away from the minimum. The Newton step $[-\tfrac{13}{7},-\tfrac{11}{7}]=[-1.857,-1.571]$ points in a different direction altogether, straight at the bottom.
For $f:\mathbb{R}^n\to\mathbb{R}$ the gradient $\mathbf{g}=\nabla f(\mathbf{x})$ is the column vector of first derivatives and the Hessian $H=\nabla^2f(\mathbf{x})$ is the symmetric $n\times n$ matrix with entries $H_{ij}=\dfrac{\partial^2f}{\partial x_i\partial x_j}$. The local quadratic model of a step $\mathbf{d}$ is
$$m(\mathbf{d})=f(\mathbf{x})+\mathbf{g}^\top\mathbf{d}+\tfrac12\mathbf{d}^\top H\,\mathbf{d}.$$Its gradient with respect to $\mathbf{d}$ is $\mathbf{g}+H\mathbf{d}$ (for symmetric $H$). Setting it to zero gives the Newton step
$$\boxed{H\,\mathbf{d}=-\mathbf{g}\quad\Longrightarrow\quad \mathbf{d}=-H^{-1}\mathbf{g},\qquad \mathbf{x}\leftarrow\mathbf{x}+\mathbf{d}=\mathbf{x}-H^{-1}\nabla f(\mathbf{x}).}$$- Solve, never invert. In code you compute $\mathbf{d}$ by solving the linear system $H\mathbf{d}=-\mathbf{g}$ (
np.linalg.solve(H, -g), or a Cholesky factorisation when $H$ is positive definite). Forming $H^{-1}$ explicitly costs more and is less accurate. - If $H$ is positive definite the model is a bowl, $\mathbf{d}$ is its unique minimum, and $\mathbf{d}$ is a descent direction: $\mathbf{g}^\top\mathbf{d}=-\mathbf{g}^\top H^{-1}\mathbf{g}<0$ because $H^{-1}$ is positive definite too.
- If $f$ is a quadratic $\tfrac12\mathbf{x}^\top A\mathbf{x}-\mathbf{b}^\top\mathbf{x}$ with $A\succ0$, the model is $f$, so one Newton step lands on the exact minimiser $A^{-1}\mathbf{b}$ from any start.
- Algorithm. Repeat: compute $\mathbf{g}$ and $H$; stop if $\|\mathbf{g}\|$ is below a tolerance; solve $H\mathbf{d}=-\mathbf{g}$; set $\mathbf{x}\leftarrow\mathbf{x}+\mathbf{d}$.
- A preconditioned gradient step. Gradient descent is $\mathbf{x}\leftarrow\mathbf{x}-\eta I\,\mathbf{g}$. Newton is $\mathbf{x}\leftarrow\mathbf{x}-P\mathbf{g}$ with the "preconditioner" $P=H^{-1}$, which undoes the stretching of the valley. Every method in this chapter is a different choice of $P$.
Why do we need it?
On a long, tilted valley the steepest-descent direction zig-zags and needs thousands of steps. The Hessian knows the shape of the valley and points straight along it.
Where is it used?
Newton's method for logistic regression and other generalised linear models (statsmodels, R's glm), interior-point solvers for linear and quadratic programs, and the Newton polish at the end of non-linear least-squares and scientific optimization codes.
How is it used?
Compute the gradient and Hessian, solve $H\mathbf{d}=-\mathbf{g}$ with a linear solver, step, and repeat. Expect 5 to 15 iterations on a well-behaved smooth problem, each one costing a linear solve.
"One step" is for quadratics. For a general smooth $f$ the model is only an approximation, so Newton needs several steps (the next section shows how fast it converges).
Do not write inv(H) @ g. Use a solver. Better still, if $H$ is large and sparse or structured, use a factorisation or an iterative method (Chapter 3.17 introduces conjugate gradients, which are often used to solve exactly this system).
The Hessian has to be recomputed at every new point (unless the function is quadratic), which is the main cost of the method.
Quick check: $f(x,y)=x^2+3y^2$. Take one Newton step from $(4,\,-2)$.
$\nabla f=[2x,\,6y]=[8,\,-12]$ and $H=\mathrm{diag}(2,\,6)$. Solve $H\mathbf{d}=-\mathbf{g}$: $d_1=-8/2=-4$ and $d_2=12/6=2$. So $\mathbf{x}_{\text{new}}=(4-4,\,-2+2)=(0,0)$, the exact minimum, in one step.
Advantages: no units, no step size, digits that double core
What we need from earlier chapters: the condition number and its effect on gradient descent (Chapter 3.3; the rates are studied in Chapter 3.5), and the condition number (Chapter 1.7).
Newton's method has three real strengths.
- One step on a bowl. If the loss is a quadratic bowl, Newton lands on the bottom in one step, however stretched or tilted the bowl is.
- It does not care about units. Suppose you measure one input in metres instead of millimetres. That stretches one axis of the problem a thousand-fold and wrecks gradient descent (long thin valleys). Newton's method sees the same stretch in the curvature and cancels it exactly: its iterates are the same, just expressed in different units.
- Near the answer it speeds up dramatically. Gradient descent shrinks the error by a fixed factor each step. Newton roughly squares the error each step: $10^{-2}\to10^{-4}\to10^{-8}\to10^{-16}$. The number of correct digits doubles.
And there is no learning rate to tune.
Quadratic convergence on a real ML loss. Fit a logistic regression with two weights (the data are 40 seeded points in two blobs, with a small L2 penalty $\lambda=0.05$; it is the loss in the widget below), starting from $\mathbf{w}=\mathbf{0}$. The distance to the minimum $\mathbf{w}^\star=(1.326,\,0.604)$ after each pure Newton step is
| step | 0 | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|---|
| distance | 1.46 | 0.48 | 0.096 | $4.0\times10^{-3}$ | $7.1\times10^{-6}$ | $2.2\times10^{-11}$ | $2.5\times10^{-16}$ |
From step 3 on the exponent doubles: $-2.4\to-5.1\to-10.7$ (then the computer's precision runs out at $10^{-16}$). Gradient descent with the safe step $\eta=1/L$ ($L=0.867$ bounds the largest curvature) is at $8\times10^{-3}$ after 20 steps and needs about 100 steps to reach $10^{-10}$.
Units. Now multiply the second input feature by $10$ (as if it were measured in different units), with no penalty. Newton still needs 6 or 7 steps (to reach a gradient of size $10^{-9}$). Gradient descent with step $1/L$ needs about $200$ steps to reach a gradient of size $10^{-6}$ in the original units, and about $9{,}700$ after the feature is multiplied by $10$ (about $2{,}800$ if it is multiplied by $0.1$ instead).
Affine invariance. Change variables by an invertible matrix $S$: $\mathbf{x}=S\mathbf{z}$, and let $\tilde f(\mathbf{z})=f(S\mathbf{z})$. By the chain rule, $\nabla\tilde f(\mathbf{z})=S^\top\mathbf{g}$ and $\nabla^2\tilde f(\mathbf{z})=S^\top HS$, where $\mathbf{g}$ and $H$ are the gradient and Hessian of $f$ at $\mathbf{x}$. The Newton step in $\mathbf{z}$-coordinates is $$\mathbf{d}_z=-(S^\top HS)^{-1}S^\top\mathbf{g}=-S^{-1}H^{-1}S^{-\top}S^\top\mathbf{g}=-S^{-1}H^{-1}\mathbf{g},\qquad\text{so}\qquad S\,\mathbf{d}_z=-H^{-1}\mathbf{g}=\mathbf{d}_x .$$ The $\mathbf{z}$-step, mapped back by $S$, is exactly the $\mathbf{x}$-step. Newton's iterates do not depend on the choice of coordinates. A gradient step in $\mathbf{z}$ is $-\eta S^\top\mathbf{g}$, which maps to $-\eta SS^\top\mathbf{g}\ne-\eta\mathbf{g}$ unless $S$ is a rotation: gradient descent does depend on the units. (In machine learning this is why we standardise features before using gradient descent, and why second-order methods need less pre-processing.)
Quadratic convergence (local theorem). Suppose $f$ is twice continuously differentiable, $\nabla f(\mathbf{x}^\star)=\mathbf{0}$, $H(\mathbf{x}^\star)$ is positive definite, and $H$ is Lipschitz continuous near $\mathbf{x}^\star$ (its entries cannot change faster than a constant $M$ times the distance). Then if $\mathbf{x}_0$ is close enough to $\mathbf{x}^\star$, Newton's iterates converge to $\mathbf{x}^\star$ and $$\|\mathbf{x}_{k+1}-\mathbf{x}^\star\|\ \le\ C\,\|\mathbf{x}_k-\mathbf{x}^\star\|^2,\qquad C\ \text{of the order of}\ M/\mu,$$ where $\mu$ is the smallest eigenvalue of $H(\mathbf{x}^\star)$. This is second-order (quadratic) convergence. Compare: gradient descent on a strongly convex quadratic with the best fixed step shrinks the error by $\frac{\kappa-1}{\kappa+1}$ per step, a linear rate that gets close to $1$ (very slow) when the condition number $\kappa$ is large.
Honest scope. "Close enough" is essential: the theorem says nothing about the early iterations, and far from the minimum pure Newton can do badly (next section).
Why do we need it?
When you need a very accurate answer or the problem is badly scaled, first-order methods crawl. Newton's method reaches machine precision in a few steps and needs no step-size tuning.
Where is it used?
Fitting logistic regression and other generalised linear models in statsmodels and R, polishing the final iterates of constrained solvers (interior-point methods use Newton steps inside), and scientific computing where precision matters.
How is it used?
Use it when the number of parameters is small enough to form and solve with $H$ (up to a few thousand), when you can compute $H$ cheaply, and when you want fast, accurate convergence. Start with a safeguarded version (line search).
Fast only near the answer. Quadratic convergence is a local property. The early iterations of Newton's method on a hard problem can be erratic, and the practical algorithms add safeguards (damping, line search; see below).
Fewer iterations is not less work. One Newton iteration costs a Hessian and a linear solve, vastly more than a gradient step. Newton wins when iterations are expensive to repeat and the dimension is moderate.
Affine invariance is about linear changes of units. It does not remove the difficulty of a genuinely non-quadratic landscape, only of a badly scaled one.
Quick check: after a Newton step the error is $10^{-3}$. About how large will it be after the next two steps (assume a constant $C\approx1$)?
Squaring each time: $10^{-3}\to(10^{-3})^2=10^{-6}\to(10^{-6})^2=10^{-12}$. In practice two more steps give roughly $10^{-6}$ and then $10^{-12}$ (the constant $C$ shifts the numbers a little, but the doubling of digits is the signature of quadratic convergence).
Disadvantages, and the standard fixes core
What we need from earlier chapters: positive definite matrices and eigenvalue signs (Chapter 1.12), saddle points and the second-derivative test (Chapters 2.10 and 3.2), and the line search idea (Chapter 3.3).
Newton's method trusts its parabola completely. That trust is the source of both its speed and its weaknesses.
- It needs the ground to curve upward. If the local picture is a dome or a saddle instead of a bowl, the "bottom" it jumps to is a hilltop or a mountain pass. Newton happily walks up.
- It trusts a parabola far from home. Where the ground is almost flat the parabola is very wide and its bottom is miles away, so the jump can overshoot wildly, even to a worse place than where you started.
- It is expensive (next section) and needs second derivatives.
The cures all say the same thing: be cautious. Take a shorter step (damping), check that the step really went down (line search), or bend the curvature matrix so that it always opens upward (the Levenberg shift).
Negative curvature. Take $f(x,y)=x^2-y^2+\tfrac14y^4$. It has a saddle at $(0,0)$ and two minima at $(0,\pm\sqrt2)$ with $f=-1$. At $(1,\ 0.05)$:
- Gradient $\mathbf{g}=[2x,\ -2y+y^3]=[2,\ -0.0999]$ and Hessian $H=\mathrm{diag}(2,\ -2+3y^2)=\mathrm{diag}(2,\ -1.9925)$. The second eigenvalue is negative: $H$ is not positive definite.
- Newton step: $\mathbf{d}=-H^{-1}\mathbf{g}=[-1,\ -0.0501]$, landing at $(0,\ -0.0001)$, essentially the saddle ($f\approx0$, higher than the minima at $-1$). The $y$-coordinate moved toward $0$, where the surface is a hilltop in the $y$ direction. Newton stops there: the gradient is zero.
- Levenberg fix with $\tau=3$. Solve $(H+\tau I)\mathbf{d}=-\mathbf{g}$ with $H+3I=\mathrm{diag}(5,\ 1.0075)$ (positive definite): $\mathbf{d}=[-0.4,\ +0.0991]$. Now $y$ moves away from the saddle, to $(0.6,\ 0.149)$ with $f=0.338$ (down from $0.9975$). Repeating the step 12 times ends at $(0.002,\ 1.414)$, the minimum.
Divergence far from the answer. $f(x)=\sqrt{1+x^2}$ has its minimum at $0$. Here $f'=\dfrac{x}{\sqrt{1+x^2}}$ and $f''=(1+x^2)^{-3/2}$, so $-f'/f''=-x(1+x^2)$ and the Newton update is exactly
$$x_{\text{new}}=x-x(1+x^2)=-x^3 .$$Start at $x_0=1.5$: $\ 1.5\to-3.375\to38.4\to-56{,}815\to1.8\times10^{14}$. Each jump lands further away. Starting at $|x_0|<1$ it converges (to $0$, extremely fast), at $|x_0|=1$ it cycles $1,-1,1,\ldots$ and at $|x_0|>1$ it diverges. Take a half step ($\alpha=0.5$) instead: $1.5\to-0.9375\to-0.0568\to-0.0283\to\ldots$ converges (slowly, halving each step). A backtracking line search takes a half step once and then full steps: $1.5\to-0.9375\to0.824\to-0.559\to0.175\to-0.005\to0$.
The disadvantages.
- Cost. The Hessian has $n^2$ entries (memory $O(n^2)$) and solving $H\mathbf{d}=-\mathbf{g}$ takes $O(n^3)$ operations, per iteration (details in the next section).
- Needs a positive definite $H$. If $H$ has a negative or zero eigenvalue the Newton step may not be a descent direction: $\mathbf{g}^\top\mathbf{d}=-\mathbf{g}^\top H^{-1}\mathbf{g}$ can be positive. With a singular $H$ the step is undefined. On non-convex functions Newton converges to any stationary point near it: a minimum, a saddle or a maximum.
- Only locally reliable. The theorem of the previous section needs a start near the minimum. Far away the step can overshoot or diverge.
- Needs second derivatives, which can be costly or tedious to code (automatic differentiation helps: Chapter 3.16).
The fixes.
- Damping / line search. Take $\mathbf{x}\leftarrow\mathbf{x}+\alpha\mathbf{d}$ with $\alpha\in(0,1]$ chosen by backtracking: start at $\alpha=1$ and halve it until $f(\mathbf{x}+\alpha\mathbf{d})\le f(\mathbf{x})+c\,\alpha\,\mathbf{g}^\top\mathbf{d}$ (the Armijo condition, $c\approx10^{-4}$). Near the solution $\alpha=1$ is accepted and quadratic convergence returns.
- Levenberg (modified Newton). Solve $(H+\tau I)\mathbf{d}=-\mathbf{g}$ with $\tau\ge0$ large enough to make $H+\tau I$ positive definite (the eigenvalues become $\lambda_i+\tau$, so any $\tau>-\lambda_{\min}$ works; test with a Cholesky factorisation and increase $\tau$ until it succeeds). Then $\mathbf{d}$ is always a descent direction. $\tau=0$ is Newton; for huge $\tau$ the step is $\approx-\mathbf{g}/\tau$, a plain gradient step with learning rate $1/\tau$. So $\tau$ is a smooth dial between gradient descent and Newton.
- Trust region (awareness). Minimise the model only inside a ball $\|\mathbf{d}\|\le\Delta$, and grow or shrink $\Delta$ depending on how well the model predicted the drop. It is equivalent to a Levenberg shift chosen automatically.
- Hessian-free Newton (awareness). Solve $H\mathbf{d}=-\mathbf{g}$ only approximately by conjugate gradients, which needs just Hessian-vector products $H\mathbf{v}$ (cheap by automatic differentiation) and never forms $H$.
Why do we need it?
Textbook Newton is elegant but fragile. Real solvers need to survive non-convex losses, bad starting points and flat regions, so every practical Newton-type code contains a safeguard.
Where is it used?
Damped Newton in logistic-regression solvers (newton-cg, statsmodels), Levenberg–Marquardt in scipy.optimize.least_squares and curve fitting, trust-region Newton in scipy.optimize.minimize(method="trust-ncg"), and Hessian-free optimization for neural nets.
How is it used?
Check the sign of the curvature (try a Cholesky factorisation of $H$); if it fails, add $\tau I$ and retry. Then use a backtracking line search along the step, and stop when $\|\mathbf{g}\|$ is small.
A Levenberg shift that is only just big enough gives a huge step. If $\tau$ barely exceeds $-\lambda_{\min}$, the matrix $H+\tau I$ has an eigenvalue near $0$ and the step along that direction is enormous (in the example with $\tau=2$ the step in $y$ is $13$, which lands at $f\approx7800$). Practical codes use a comfortably larger $\tau$ or a trust region.
A positive definite $H$ does not guarantee a good full step. The $\sqrt{1+x^2}$ example has $f''>0$ everywhere and Newton still diverges. Convexity of the model says the step goes downhill for the model, not for $f$.
Near a saddle, Newton does not "escape". It is attracted to saddles, because the gradient is zero there and the step shrinks to nothing. Gradient descent, by contrast, drifts away from a saddle if started slightly off it (more in Chapter 3.15).
Quick check: for $f(x)=-x^2$ (a hilltop), what does one Newton step do from $x=3$?
$f'=-2x=-6$ and $f''=-2$, so $\delta=-f'/f''=-(-6)/(-2)=-3$ and $x_{\text{new}}=0$: it lands exactly on the maximum in one step. Newton finds a stationary point, with no regard for whether it is a minimum. The Levenberg fix with $\tau=3$ gives $(f''+\tau)=1>0$ and $\delta=-f'/(f''+\tau)=6$, moving away from the maximum, to $x=9$, downhill.
Computational cost: why Newton is impractical for deep networks core
What we need from earlier chapters: the Cholesky factorisation (Chapter 1.13) and how the cost of a matrix algorithm is counted (Chapter 1.4, Chapter 1.15).
Newton's method gives you the best possible direction, and sends you the bill. The bill has two lines.
- Storage. The Hessian is a square table with $n\times n$ entries. Double the number of parameters and the table becomes four times bigger. A model with a million parameters has a Hessian with a trillion numbers.
- Work. Solving the linear system $H\mathbf{d}=-\mathbf{g}$ costs about $n^3$ operations. Double the parameters and the work grows eight times.
A gradient step, by comparison, stores $n$ numbers and does about $n$ operations of arithmetic on top of computing the gradient. For a modest model, Newton is affordable. For a neural network, it is out of the question.
Count with 8-byte numbers (double precision) and a machine that does $10^{12}$ operations per second (one TFLOP/s, a good GPU). The Hessian has $n^2$ numbers; a Cholesky solve costs about $\tfrac13n^3$ operations.
| parameters $n$ | gradient (n numbers) | Hessian ($n^2$ numbers) | solve $\approx n^3/3$ operations | time at $10^{12}$ ops/s |
|---|---|---|---|---|
| $10^3$ | 8 KB | 8 MB | $3.3\times10^{8}$ | 0.33 ms |
| $10^6$ | 8 MB | 8 TB | $3.3\times10^{17}$ | 3.9 days |
| $10^9$ | 8 GB | 8 EB (exabytes: $8\times10^{18}$ bytes) | $3.3\times10^{26}$ | about 10 million years |
And that is one iteration. For $n=10^3$ Newton is trivial; for $n=10^6$ the memory alone is hopeless; for a billion-parameter network nothing about it is feasible. Modern language models have $10^9$ to $10^{12}$ parameters.
Cost per iteration (beyond computing the gradient itself), for $n$ parameters:
| method | curvature storage (numbers) | extra work per iteration |
|---|---|---|
| Gradient descent | $0$ | $O(n)$ (about $2n$ operations) |
| Newton | $n^2$ (or $\frac{n(n+1)}{2}$ using symmetry) | build $H$ (second derivatives) and solve: $O(n^3)$ (Cholesky $\approx\frac13n^3$) |
| BFGS (next section) | $n^2$ | $O(n^2)$ (a few matrix-vector products and a rank-two update) |
| L-BFGS with $m$ pairs | $2mn$ | $O(mn)$ (about $8mn$ operations) |
- The $O(n^2)$ memory is usually the first wall (before the time).
- The cost is per iteration. Newton compensates by needing few iterations (often 5 to 20 for a smooth problem).
- Hessian-vector products are cheap. The product $H\mathbf{v}$ (without forming $H$) costs only a small multiple (about 2 to 3 times) of one gradient evaluation, using automatic differentiation. Methods that use only these products (Newton-CG, Hessian-free optimization) avoid the $n^2$ memory and are used in practice for large problems.
- Cheaper approximations are the rest of this chapter: BFGS builds an approximate inverse Hessian from gradients ($O(n^2)$), and L-BFGS stores only a few vectors ($O(mn)$).
Why do we need it?
To choose a method before running it. A quick cost estimate tells you whether a second-order method fits in memory and time, or whether you need a first-order or limited-memory method.
Where is it used?
Choosing between newton-cg, lbfgs and sag solvers in scikit-learn, deciding why large neural networks are trained with SGD and Adam, and sizing scientific optimization problems (dense Newton for hundreds of unknowns, L-BFGS for millions).
How is it used?
Write down $n$. If $n$ is below a few thousand, a dense Newton or BFGS is fine. If it is up to millions, use L-BFGS or Newton-CG. Beyond that, use first-order methods such as SGD and Adam.
These are rough, idealised counts. Real speeds depend on caches, parallelism, sparsity (a sparse Hessian is far cheaper than $n^2$) and on how the Hessian is computed. The message is the growth rates: $n^2$ memory and $n^3$ work against $n$ for first-order methods.
Structure rescues Newton in some problems. If $H$ is banded, sparse, or low-rank-plus-diagonal (as in logistic regression with few features, or in interior-point solvers), it can be factorised far faster than $n^3$.
Cost per iteration is not the whole story. If an iteration of Newton replaces a thousand gradient steps, it may still be cheaper. For problems with $n$ up to a few thousand it often is.
Quick check: a model has $n=20{,}000$ parameters. How much memory does a dense Hessian need (8-byte numbers), and how long does a Cholesky solve take at $10^{12}$ operations per second?
Memory: $n^2=4\times10^8$ numbers $\times8$ bytes $=3.2\times10^9$ bytes $=3.2$ GB. Time: $n^3/3=8\times10^{12}/3\approx2.7\times10^{12}$ operations, divided by $10^{12}$ per second gives about $2.7$ seconds per iteration. Feasible on a big machine, but already heavy; with $n=200{,}000$ it would be 320 GB and about 45 minutes per iteration.
Quasi-Newton methods: learn the curvature from the walk itself core
Newton's method needs the Hessian, and the Hessian is expensive. But notice that you get information for free while you walk. When you step from $\mathbf{x}_k$ to $\mathbf{x}_{k+1}$ the gradient changes. How much the slope changed over the step you took tells you the curvature along that step.
In one dimension this is the familiar idea: curvature $\approx\dfrac{\text{change in slope}}{\text{change in position}}$. The method that uses it is the secant method: instead of the tangent (needs $f''$), use the line through the last two points. In many dimensions you only measure the curvature along the one direction you walked, but you can keep a running estimate of the whole Hessian (or its inverse) and correct it a little with every step, like drawing a map of a landscape as you hike across it. Such methods are called quasi-Newton ("almost Newton"): they take Newton-like steps with an approximate Hessian built only from gradients.
1D, by hand. Let $f(x)=\tfrac14x^4+\tfrac12x^2-x$ again, with slope $g(x)=f'(x)=x^3+x-1$. Take two points $x_0=0$ ($g=-1$) and $x_1=1$ ($g=1$).
- Step $s=x_1-x_0=1$, change of slope $y=g(x_1)-g(x_0)=1-(-1)=2$.
- Curvature estimate $y/s=2$. (The true curvature $f''=3x^2+1$ goes from $1$ to $4$ on $[0,1]$ and is $2.4$ at the minimum $0.68$, so $2$ is a fair average.)
- Next point: where the secant line through $(0,-1)$ and $(1,1)$ crosses zero: $x_2=1-\dfrac{g(1)}{2}=0.5$.
Continuing with the two newest points each time gives $x_3=0.636$, $x_4=0.690$, $x_5=0.68202$, $x_6=0.682326$, $x_7=0.68232780$. The errors are $0.32,\ 0.18,\ 0.046,\ 0.0077,\ 3\times10^{-4},\ 2\times10^{-6},\ 6\times10^{-10}$: faster than a fixed-factor decrease (linear) but not as fast as Newton's squaring. This is superlinear convergence, and it used no second derivative at all.
Let $\mathbf{s}_k=\mathbf{x}_{k+1}-\mathbf{x}_k$ be the step and $\mathbf{y}_k=\nabla f(\mathbf{x}_{k+1})-\nabla f(\mathbf{x}_k)$ the change of the gradient. Why they are related: for a quadratic $f=\tfrac12\mathbf{x}^\top A\mathbf{x}-\mathbf{b}^\top\mathbf{x}$ we have $\nabla f=A\mathbf{x}-\mathbf{b}$, so $\mathbf{y}_k=A(\mathbf{x}_{k+1}-\mathbf{x}_k)=A\mathbf{s}_k$ exactly. For a general smooth $f$, $\mathbf{y}_k=\bar H\mathbf{s}_k$ where $\bar H=\int_0^1\nabla^2f(\mathbf{x}_k+t\mathbf{s}_k)\,dt$ is the average Hessian along the step (the fundamental theorem of calculus applied to $\nabla f$).
A quasi-Newton method keeps an approximation $B_k\approx H$ and demands that the new approximation reproduces what we just measured. This is the secant condition:
$$\boxed{B_{k+1}\,\mathbf{s}_k=\mathbf{y}_k}\qquad\text{equivalently, for the inverse }C_{k+1}=B_{k+1}^{-1}:\quad C_{k+1}\,\mathbf{y}_k=\mathbf{s}_k .$$- It is a requirement, not a formula. $B_{k+1}$ has $\tfrac{n(n+1)}{2}$ free numbers (it is symmetric) and the condition gives only $n$ equations, so many matrices qualify. We pick the one that changes the previous estimate the least (next section).
- The curvature condition. Multiply the secant condition by $\mathbf{s}_k^\top$: $\mathbf{s}_k^\top B_{k+1}\mathbf{s}_k=\mathbf{s}_k^\top\mathbf{y}_k$. A positive definite $B_{k+1}$ makes the left side positive, so we need $\mathbf{s}_k^\top\mathbf{y}_k>0$: the slope must have increased along the step, as it does for a convex function.
- What it teaches. One pair $(\mathbf{s},\mathbf{y})$ tells you the curvature in the direction $\mathbf{s}$ only. After $n$ steps in independent directions on a quadratic, you know $A$ completely.
- In 1D the condition is $b=y/s$, and Newton's step with that $b$ is the secant method.
Why do we need it?
Computing the exact Hessian costs $n^2$ memory and $n^3$ time, and may need second derivatives that are hard to code. The secant condition lets us learn the curvature from gradients alone.
Where is it used?
The secant method for root finding in 1D, and BFGS, DFP and L-BFGS for multi-dimensional smooth optimization (SciPy, MATLAB's fminunc, R's optim, Stan's and many ML libraries' default optimizers).
How is it used?
After each step, record $\mathbf{s}=\mathbf{x}_{\text{new}}-\mathbf{x}_{\text{old}}$ and $\mathbf{y}=\mathbf{g}_{\text{new}}-\mathbf{g}_{\text{old}}$, check $\mathbf{s}^\top\mathbf{y}>0$, and update the curvature estimate so that it maps $\mathbf{s}$ to $\mathbf{y}$.
Only one direction is learned per step. The pair $(\mathbf{s},\mathbf{y})$ says nothing about curvature in directions perpendicular to $\mathbf{s}$; those parts of the estimate are carried over from before. So quasi-Newton methods improve gradually.
$\mathbf{s}^\top\mathbf{y}\le0$ means the function is not convex along the step. A positive definite update is then impossible. Good implementations choose step lengths by a line search that guarantees $\mathbf{s}^\top\mathbf{y}>0$ (next section) or skip the update.
Rounding. If the step $\mathbf{s}$ is tiny, $\mathbf{y}=\nabla f(\mathbf{x}_{k+1})-\nabla f(\mathbf{x}_k)$ is a difference of nearly equal numbers and loses precision (Chapter 3.16). The condition is exact only up to this error.
Quick check: for $f(x)=x^2$ (so $g=2x$) the points $x_0=1$ and $x_1=3$ give what $s$, $y$ and curvature estimate? Is it right?
$s=3-1=2$, $y=g(3)-g(1)=6-2=4$, so $y/s=2=f''$ exactly. For a quadratic the slope is a straight line, so the secant slope is the true curvature at any two points (the secant condition $y=As$ is exact).
BFGS: a self-correcting inverse-Hessian estimate core
What we need from earlier chapters: the line search (Chapter 3.3), outer products and rank-one matrices (Chapter 1.4) and positive definiteness (Chapter 1.12).
BFGS (named after Broyden, Fletcher, Goldfarb and Shanno) keeps a running guess $C_k$ of the inverse Hessian. Because it stores the inverse, the search direction needs no linear solve: just a matrix-vector product, $\mathbf{d}=-C_k\mathbf{g}$.
It starts with $C_0=I$, so the very first step is a plain gradient step. After each step it corrects the guess using the newest pair $(\mathbf{s},\mathbf{y})$, so that the corrected guess satisfies the secant condition, while changing the old guess as little as possible. Each correction is a small, two-dimensional ("rank-two") adjustment. The guess gets better with every step, and on a quadratic bowl it becomes exactly right after $n$ steps.
Use the bowl from before: $f=\tfrac12\mathbf{x}^\top A\mathbf{x}-\mathbf{b}^\top\mathbf{x}$, $A=\begin{bmatrix}4&1\\1&2\end{bmatrix}$, $\mathbf{b}=[1,1]$, start $\mathbf{x}_0=(2,2)$, $C_0=I$, exact line search.
- Step 1. $\mathbf{g}_0=[9,\,5]$, $\mathbf{d}_0=-C_0\mathbf{g}_0=[-9,-5]$. Exact line search along $\mathbf{d}_0$ gives $\alpha_0=\frac{\mathbf{g}_0^\top\mathbf{g}_0}{\mathbf{g}_0^\top A\mathbf{g}_0}=\frac{106}{464}=0.2284$, so $\mathbf{x}_1=(-0.056,\ 0.858)$ and $\mathbf{g}_1=[-0.366,\ 0.660]$.
- Measure. $\mathbf{s}_0=\mathbf{x}_1-\mathbf{x}_0=[-2.056,\,-1.142]$ and $\mathbf{y}_0=\mathbf{g}_1-\mathbf{g}_0=[-9.366,\,-4.341]$. Check: $\mathbf{s}_0^\top\mathbf{y}_0=19.26+4.96=24.22>0$ ✓.
- Update. With $\rho=1/\mathbf{y}_0^\top\mathbf{s}_0=0.0413$, the BFGS formula gives $C_1=\begin{bmatrix}0.352&-0.287\\-0.287&0.882\end{bmatrix}$. Check the secant condition: $C_1\mathbf{y}_0=[-2.056,\,-1.142]=\mathbf{s}_0$ ✓.
- Step 2. $\mathbf{d}_1=-C_1\mathbf{g}_1=[0.318,\,-0.686]$, exact line search $\alpha_1=0.625$, giving $\mathbf{x}_2=(0.1429,\ 0.4286)=(\tfrac17,\tfrac37)$: the exact minimum in two steps.
- The estimate is now exact. The next update gives $C_2=\begin{bmatrix}0.2857&-0.1429\\-0.1429&0.5714\end{bmatrix}=A^{-1}$ (the true inverse is $\frac17\begin{bmatrix}2&-1\\-1&4\end{bmatrix}$).
The BFGS algorithm. Choose $\mathbf{x}_0$ and $C_0=I$ (or a scaled identity). For $k=0,1,2,\dots$
- $\mathbf{g}_k=\nabla f(\mathbf{x}_k)$; stop if $\|\mathbf{g}_k\|$ is small.
- Direction $\mathbf{d}_k=-C_k\mathbf{g}_k$.
- Line search for a step length $\alpha_k$ (satisfying the Wolfe conditions, below); $\mathbf{x}_{k+1}=\mathbf{x}_k+\alpha_k\mathbf{d}_k$.
- $\mathbf{s}_k=\mathbf{x}_{k+1}-\mathbf{x}_k$, $\mathbf{y}_k=\mathbf{g}_{k+1}-\mathbf{g}_k$, $\rho_k=1/(\mathbf{y}_k^\top\mathbf{s}_k)$.
- Update: $$\boxed{C_{k+1}=\big(I-\rho_k\mathbf{s}_k\mathbf{y}_k^\top\big)\,C_k\,\big(I-\rho_k\mathbf{y}_k\mathbf{s}_k^\top\big)+\rho_k\mathbf{s}_k\mathbf{s}_k^\top .}$$
(The same update for the Hessian estimate $B=C^{-1}$ reads $B_{k+1}=B_k-\dfrac{B_k\mathbf{s}\mathbf{s}^\top B_k}{\mathbf{s}^\top B_k\mathbf{s}}+\dfrac{\mathbf{y}\mathbf{y}^\top}{\mathbf{y}^\top\mathbf{s}}$; the two forms are inverses of each other.)
It satisfies the secant condition. Using $\rho\,\mathbf{s}^\top\mathbf{y}=1$: $C_{k+1}\mathbf{y}=(I-\rho\mathbf{s}\mathbf{y}^\top)\,C\,(\mathbf{y}-\rho\mathbf{y}\,\mathbf{s}^\top\mathbf{y})+\rho\mathbf{s}\,(\mathbf{s}^\top\mathbf{y})=(I-\rho\mathbf{s}\mathbf{y}^\top)\,C\,(\mathbf{y}-\mathbf{y})+\mathbf{s}=\mathbf{s}.$ ✓
It keeps $C$ positive definite whenever $\mathbf{s}^\top\mathbf{y}>0$ (and $C_k$ is). For any $\mathbf{z}\ne\mathbf{0}$ put $\mathbf{w}=(I-\rho\mathbf{y}\mathbf{s}^\top)\mathbf{z}=\mathbf{z}-\rho(\mathbf{s}^\top\mathbf{z})\mathbf{y}$. Then $\mathbf{z}^\top C_{k+1}\mathbf{z}=\mathbf{w}^\top C_k\mathbf{w}+\rho(\mathbf{s}^\top\mathbf{z})^2\ge0$, with equality only if $\mathbf{s}^\top\mathbf{z}=0$ and $\mathbf{w}=\mathbf{0}$, which together force $\mathbf{z}=\mathbf{0}$. So $\mathbf{d}_k=-C_k\mathbf{g}_k$ is always a descent direction ($\mathbf{g}^\top\mathbf{d}=-\mathbf{g}^\top C\mathbf{g}<0$).
The Wolfe conditions guarantee $\mathbf{s}^\top\mathbf{y}>0$. They require, for constants $0\lt c_1\lt c_2\lt1$ (typically $10^{-4}$ and $0.9$): (a) sufficient decrease, $f(\mathbf{x}+\alpha\mathbf{d})\le f(\mathbf{x})+c_1\alpha\,\mathbf{g}^\top\mathbf{d}$, and (b) the curvature condition $\nabla f(\mathbf{x}+\alpha\mathbf{d})^\top\mathbf{d}\ge c_2\,\mathbf{g}^\top\mathbf{d}$. From (b), $\mathbf{y}^\top\mathbf{s}=\alpha(\mathbf{g}_{\text{new}}-\mathbf{g})^\top\mathbf{d}\ge\alpha(c_2-1)\mathbf{g}^\top\mathbf{d}>0$, because $\mathbf{g}^\top\mathbf{d}<0$ and $c_2<1$.
- Where the formula comes from (awareness). Among all symmetric matrices $C'$ with $C'\mathbf{y}=\mathbf{s}$, the BFGS update is the one closest to $C_k$ in a weighted matrix norm. It is the minimal change that teaches the matrix the new pair.
- Cost. $O(n^2)$ per iteration (a matrix-vector product and a rank-two update) and $n^2$ numbers of memory. No Hessian, no linear solve, no second derivatives.
- Speed. Under standard conditions (smooth, strongly convex near the minimum, Wolfe line search) BFGS converges superlinearly: $\|\mathbf{x}_{k+1}-\mathbf{x}^\star\|/\|\mathbf{x}_k-\mathbf{x}^\star\|\to0$. That is faster than gradient descent (linear), slower than Newton (quadratic). On a quadratic with exact line searches it terminates in at most $n$ steps.
- Self-correcting. A poor $C_0$ is repaired by the later updates, which is not true of some older methods.
Why do we need it?
We want Newton-like speed on a smooth problem without ever computing a Hessian or solving a linear system. BFGS gives superlinear convergence from gradients alone.
Where is it used?
scipy.optimize.minimize(method="BFGS") (SciPy's default for unconstrained smooth problems), MATLAB fminunc, R's optim(method="BFGS"), maximum-likelihood fitting of statistical models, and small-to-medium scientific problems (up to a few thousand variables).
How is it used?
Supply the function and its gradient (or use automatic differentiation), call the optimizer, and monitor the gradient norm. Provide a good starting point. For many thousands of variables switch to L-BFGS.
The line search matters. The update needs $\mathbf{s}^\top\mathbf{y}>0$, which a plain backtracking (Armijo) search does not guarantee on non-convex functions. Good implementations use a Wolfe line search or skip/damp the update when $\mathbf{s}^\top\mathbf{y}$ is not positive.
$n^2$ memory. BFGS is a medium-size method. For $n$ in the hundreds of thousands the matrix alone is too big (see the cost section above), and L-BFGS is used instead.
BFGS minimises smooth functions. It assumes the gradient exists everywhere; for non-smooth functions (L1 penalty, ReLU kinks) it can stall. For the lasso use the methods of Chapter 3.13.
Not the same as Adam. Both build a "preconditioner" from gradients, but BFGS builds a full matrix from gradient differences in a deterministic setting, whereas Adam keeps a diagonal scaling from running averages of squared (stochastic) gradients (Chapter 3.4).
Quick check: after a step, $\mathbf{s}=[1,\,0]$ and $\mathbf{y}=[2,\,0]$, with $C=I$. What does BFGS do to $C$, and does the secant condition hold?
$\mathbf{s}^\top\mathbf{y}=2>0$, $\rho=\frac12$. $I-\rho\mathbf{s}\mathbf{y}^\top=\begin{bmatrix}1-1&0\\0&1\end{bmatrix}=\begin{bmatrix}0&0\\0&1\end{bmatrix}$ and likewise for the right factor. So $C_{\text{new}}=\begin{bmatrix}0&0\\0&1\end{bmatrix}I\begin{bmatrix}0&0\\0&1\end{bmatrix}+\tfrac12\begin{bmatrix}1&0\\0&0\end{bmatrix}=\begin{bmatrix}0.5&0\\0&1\end{bmatrix}$. Check: $C_{\text{new}}\mathbf{y}=[0.5\cdot2,\,0]=[1,0]=\mathbf{s}$ ✓. The first coordinate learned that the curvature along $\mathbf{s}$ is $y/s=2$ (inverse $0.5$); the second coordinate, which we have not probed, is unchanged.
L-BFGS: BFGS with a short memory core
BFGS carries a big $n\times n$ matrix. But that matrix is nothing more than the identity plus a pile of small corrections, one per step, and each correction is made from just two vectors, $\mathbf{s}_i$ and $\mathbf{y}_i$. So why store the matrix at all? Keep only the last $m$ pairs $(\mathbf{s},\mathbf{y})$ (with $m$ around 5 to 20), and rebuild the product $C\mathbf{g}$ from them whenever you need a direction. Old pairs are dropped: they describe curvature of a part of the landscape you left long ago anyway.
That is L-BFGS (limited-memory BFGS). It needs about $2mn$ numbers and $O(mn)$ work per step instead of $n^2$, so it scales to millions of variables while keeping most of BFGS's speed. It is the workhorse for large, smooth, deterministic problems.
Back to the bowl $A=\begin{bmatrix}4&1\\1&2\end{bmatrix}$. After the first step we know one pair, $\mathbf{s}_0=[-2.056,\,-1.142]$, $\mathbf{y}_0=[-9.366,\,-4.341]$ with $\mathbf{y}_0^\top\mathbf{s}_0=24.22$, and the new gradient $\mathbf{g}_1=[-0.366,\,0.660]$. We want $C_1\mathbf{g}_1$ without forming $C_1$. Use the two-loop recursion with initial matrix $C^0=I$ and $m=1$ stored pair ($\rho=1/24.22=0.0413$):
- First loop (newest pair to oldest): $a=\rho\,\mathbf{s}_0^\top\mathbf{g}_1=0.0413\times0.000=0.000$ (the exact line search made $\mathbf{g}_1\perp\mathbf{s}_0$). So $\mathbf{q}=\mathbf{g}_1-a\,\mathbf{y}_0=\mathbf{g}_1$.
- Apply the initial matrix: $\mathbf{r}=C^0\mathbf{q}=[-0.366,\,0.660]$.
- Second loop (oldest to newest): $b=\rho\,\mathbf{y}_0^\top\mathbf{r}=0.0413\times0.569=0.0235$, then $\mathbf{r}\leftarrow\mathbf{r}+\mathbf{s}_0\,(a-b)=\mathbf{r}-0.0235\,\mathbf{s}_0=[-0.318,\,0.686]$.
The direction is $\mathbf{d}_1=-\mathbf{r}=[0.318,\,-0.686]$, exactly the BFGS direction $-C_1\mathbf{g}_1$ of the previous section, computed from two vectors and never forming a matrix. (Real L-BFGS uses $C^0=\gamma I$ with $\gamma=\frac{\mathbf{s}^\top\mathbf{y}}{\mathbf{y}^\top\mathbf{y}}=0.227$ here, a better-scaled start, which gives a slightly different, scale-aware direction $[0.072,\,-0.156]$.)
Two-loop recursion. Store the last $m$ pairs $(\mathbf{s}_i,\mathbf{y}_i)$, $i=k-m,\dots,k-1$, and $\rho_i=1/(\mathbf{y}_i^\top\mathbf{s}_i)$. To compute $\mathbf{r}=C_k\mathbf{g}$:
q = g
for i = k-1, k-2, ..., k-m: # newest to oldest
a_i = rho_i * (s_i . q)
q = q - a_i * y_i
r = gamma_k * q # gamma_k = (s_{k-1} . y_{k-1}) / (y_{k-1} . y_{k-1})
for i = k-m, ..., k-1: # oldest to newest
b = rho_i * (y_i . r)
r = r + s_i * (a_i - b)
return r # r = C_k g ; the search direction is d = -r
- Why it works. Unrolling the BFGS update $C_{k+1}=V_k^\top C_kV_k+\rho_k\mathbf{s}_k\mathbf{s}_k^\top$ with $V_k=I-\rho_k\mathbf{y}_k\mathbf{s}_k^\top$ gives $C_k$ as a sum of terms built from the stored pairs and $C^0$. Read the unrolled product from the right: the first loop multiplies $\mathbf{g}$ by the factors $V_i=I-\rho_i\mathbf{y}_i\mathbf{s}_i^\top$ (newest first; each one is just $\mathbf{q}\leftarrow\mathbf{q}-a_i\mathbf{y}_i$ with $a_i=\rho_i\mathbf{s}_i^\top\mathbf{q}$), the middle line applies $C^0$, and the second loop applies the transposed factors (oldest first) and adds the $\mathbf{s}_i\mathbf{s}_i^\top$ terms. With $C^0=I$ and all pairs kept it reproduces full BFGS exactly.
- Cost. Each loop does one dot product and one vector update per pair: about $4mn$ multiplications ($\approx8mn$ operations) and $2mn$ numbers of memory for the pairs, against $n^2$ for BFGS. For $n=10^6$ and $m=10$: $2\times10^7$ numbers ($160$ MB) rather than $8$ TB.
- The scaling $\gamma_k$ is an estimate of "one over the curvature along the latest step", so the first guess has the right size.
- Line search. Use a Wolfe line search (try $\alpha=1$ first: it is usually accepted near the solution, as in Newton) so that $\mathbf{s}^\top\mathbf{y}>0$ and the pairs are usable.
- Choice of $m$. Values between $5$ and $20$ are typical; more memory helps a little on hard problems and costs linearly more. $m=0$ would be steepest descent.
- Memory and speed trade-off. L-BFGS usually needs somewhat more iterations than full BFGS but each is far cheaper.
Why do we need it?
Most real problems with thousands to millions of parameters are far too big for a dense matrix, but still smooth and well-behaved. L-BFGS gives quasi-Newton speed with the memory of a few gradient vectors.
Where is it used?
Scikit-learn's LogisticRegression (default solver lbfgs), scipy.optimize.minimize(method="L-BFGS-B"), conditional random fields and maximum-entropy models, classical (non-deep) statistical fitting, neural style transfer and small neural nets (torch.optim.LBFGS), and many physics codes.
How is it used?
Provide the loss and its gradient (full batch, not a noisy mini-batch), pick $m$ (10 is a good default), set a gradient tolerance and a maximum number of iterations. There is no learning rate to tune because the line search chooses the step.
L-BFGS needs accurate, deterministic gradients. With noisy mini-batch gradients the differences $\mathbf{y}=\mathbf{g}_{\text{new}}-\mathbf{g}$ mostly measure noise, not curvature, and the line search is confused. That is why neural networks are trained with SGD and Adam, not L-BFGS (see Chapter 3.14); L-BFGS is for full-batch, smooth problems.
Smoothness. Like BFGS it assumes the loss is smooth. The L-BFGS-B variant adds simple box constraints (lower and upper bounds on variables), but L1 penalties and ReLU kinks need other methods.
Iterations versus time. L-BFGS often beats first-order methods in iterations, but each iteration includes a line search with several function and gradient evaluations. Count total gradient evaluations when you compare methods.
Quick check: with $n=5{,}000{,}000$ parameters and $m=10$, how many numbers does L-BFGS store for its pairs, compared with a dense BFGS matrix?
L-BFGS: $2mn=2\times10\times5\times10^6=10^8$ numbers, about $0.8$ GB. A dense BFGS matrix: $n^2=2.5\times10^{13}$ numbers, about $200$ TB. L-BFGS is roughly $250{,}000$ times smaller.
Relatives of Newton: Gauss–Newton, natural gradient and K-FAC awareness
What we need from earlier chapters: the Jacobian (Chapter 2.5), the normal equations of least squares (Chapter 1.10) and, for the last two, the idea of a probabilistic model's log-likelihood.
Every method in this chapter has the same shape: $\mathbf{x}\leftarrow\mathbf{x}-P\,\mathbf{g}$, a gradient step corrected by a matrix $P$ (a preconditioner) that is supposed to resemble the inverse curvature. The methods differ in how they choose $P$:
| method | preconditioner $P$ | comes from |
|---|---|---|
| gradient descent | $\eta I$ | nothing (a guess) |
| Newton | $H^{-1}$ | second derivatives |
| BFGS / L-BFGS | $C_k\approx H^{-1}$ | gradient differences $(\mathbf{s},\mathbf{y})$ |
| Gauss–Newton / Levenberg–Marquardt | $(J^\top J+\tau I)^{-1}$ | first derivatives of a least-squares model |
| natural gradient | $F^{-1}$ | the Fisher information of a probabilistic model |
| K-FAC (and Shampoo) | a Kronecker-structured $F^{-1}$ | cheap factorised Fisher for each layer |
| Adam, RMSProp | a diagonal $\mathrm{diag}\big(1/(\sqrt{v}+\epsilon)\big)$ | running averages of squared gradients (a heuristic) |
This section is a short awareness tour of three of them. The first, Gauss–Newton, you can use today.
Gauss–Newton on a curve fit. Fit $y\approx a\,e^{bx}$ to seven data points $x=0,0.5,1,1.5,2,2.5,3.5$ with $y=1.979,\ 1.327,\ 0.836,\ 0.609,\ 0.459,\ 0.324,\ 0.134$ (a decay with noise). The loss is $F(a,b)=\tfrac12\sum_i\big(a\,e^{bx_i}-y_i\big)^2$ and its minimum is at $(a,b)=(1.960,\,-0.775)$ with $F=0.0042$. From the start $(a,b)=(1,\,-0.2)$ (the widget below):
- Gauss–Newton reaches within $10^{-3}$ of the answer in 3 steps.
- Gradient descent with step size $0.02$ needs 351 steps.
- From the start $(0.5,\,-2)$ plain Gauss–Newton blows up (the first step is huge and the iterates become infinite), while Levenberg–Marquardt with $\tau=1$ converges in 12 steps.
Gauss–Newton (least squares). Let $\mathbf{r}(\mathbf{x})\in\mathbb{R}^m$ be a vector of residuals (prediction minus data) and $F(\mathbf{x})=\tfrac12\|\mathbf{r}(\mathbf{x})\|^2$. With the Jacobian $J$ ($m\times n$, entries $\partial r_i/\partial x_j$), $$\nabla F=J^\top\mathbf{r},\qquad \nabla^2F=\underbrace{J^\top J}_{\text{first derivatives only}}+\underbrace{\sum_ir_i\,\nabla^2r_i}_{\text{second derivatives}} .$$ Gauss–Newton drops the second sum (small when the residuals are small or the model is nearly linear) and uses $H\approx J^\top J$, which is cheap and always positive semi-definite:
$$\big(J^\top J\big)\,\mathbf{d}=-J^\top\mathbf{r},\qquad \mathbf{x}\leftarrow\mathbf{x}+\mathbf{d}.$$Meaning. $\mathbf{d}$ minimises $\tfrac12\|\mathbf{r}+J\mathbf{d}\|^2$ (gradient: $J^\top(\mathbf{r}+J\mathbf{d})=\mathbf{0}$): linearise the model at the current point and solve the resulting linear least-squares problem, with the normal equations of Chapter 1.10. Levenberg–Marquardt adds the Levenberg shift: $(J^\top J+\tau I)\mathbf{d}=-J^\top\mathbf{r}$, so large $\tau$ behaves like gradient descent (safe, slow) and small $\tau$ like Gauss–Newton (fast, risky). It is the standard algorithm for non-linear curve fitting.
Natural gradient (awareness). For a model of probabilities $p_\theta$ (a classifier, a language model), the Euclidean distance between two parameter vectors is a poor measure of how different the two models are. Measuring change by the KL divergence between the model's outputs leads to the Fisher information matrix $F=\mathbb{E}\big[\nabla_\theta\log p\,\nabla_\theta\log p^\top\big]$ and the step $\mathbf{d}=-\eta F^{-1}\mathbf{g}$. It does not depend on how the parameters are labelled (reparametrisation invariant). For logistic regression the Fisher matrix equals the Hessian of the loss, so Newton's method and natural gradient coincide there (statisticians call it Fisher scoring or iteratively reweighted least squares).
K-FAC (awareness). For a layer with weights $W$ of size $d_{\text{out}}\times d_{\text{in}}$ the Fisher block is huge, $(d_{\text{in}}d_{\text{out}})^2$ entries. K-FAC approximates it by a Kronecker product $A\otimes G$ of two small matrices (the covariance of the layer's inputs and of the gradients at its outputs), and $(A\otimes G)^{-1}=A^{-1}\otimes G^{-1}$, so it inverts two small matrices (costing $d_{\text{in}}^3+d_{\text{out}}^3$) instead of one giant one. Shampoo is a related idea. Such methods are used to speed up training of large networks, at the price of extra memory and complexity.
Why do we need it?
Newton-like speed on specific problem families without the full Hessian: least-squares fitting (Gauss–Newton), geometry-aware steps for probabilistic models (natural gradient) and scalable curvature for deep networks (K-FAC).
Where is it used?
scipy.optimize.least_squares and curve_fit (Levenberg–Marquardt), bundle adjustment in 3D reconstruction and SLAM in robotics, Fisher scoring in statistics, reinforcement-learning methods such as natural policy gradients and TRPO, and the K-FAC and Shampoo optimizers.
How is it used?
For a curve fit, give the residual function (and ideally its Jacobian) to least_squares. For deep networks these are specialist optimizers that you reach for when Adam is too slow in steps and you can afford the extra memory.
Gauss–Newton ignores part of the curvature. When the residuals at the solution are large (a poor model), the dropped term matters, and the method converges only linearly. It is at its best when the model can fit the data well.
Natural gradient and K-FAC are not default tools. They add memory, code and tuning (damping, update frequency) and are used in specific settings. This section is for recognising the names and the idea, not for using them on day one.
"Second-order" for deep nets usually means "approximately". Practical large-scale methods use diagonal, block-diagonal or Kronecker approximations of a curvature matrix, never the full Hessian.
Quick check: why is $J^\top J$ a safer curvature estimate than the true Hessian, and what does the Levenberg shift add?
$J^\top J$ is always positive semi-definite ($\mathbf{v}^\top J^\top J\mathbf{v}=\|J\mathbf{v}\|^2\ge0$), so the step is a descent direction when $J^\top J$ is invertible, unlike the true Hessian that can have negative eigenvalues. It can still be singular or tiny in some directions, giving a huge step; adding $\tau I$ makes the matrix positive definite and bounds the step, and interpolates toward a gradient step as $\tau$ grows.
Recap, cheat sheet and practice
- Newton–Raphson finds a root: $x\leftarrow x-f(x)/f'(x)$ (ride the tangent to zero). A smooth minimum is a root of $f'$, so Newton's method is $x\leftarrow x-f'(x)/f''(x)$, the jump to the bottom of the local parabola $f+f'\delta+\frac12f''\delta^2$. If $f''<0$ the same formula jumps to a maximum.
- In $n$ dimensions: $H\mathbf{d}=-\mathbf{g}$, $\mathbf{x}\leftarrow\mathbf{x}+\mathbf{d}$. Solve the system, never invert $H$. One step solves a quadratic exactly. Gradient descent is the same recipe with $H^{-1}$ replaced by $\eta I$.
- Advantages: no step size, affine invariance (the iterates do not depend on the units), quadratic convergence near a minimum with $H(\mathbf{x}^\star)\succ0$ and a Lipschitz Hessian (the number of correct digits doubles).
- Disadvantages: $O(n^2)$ memory and $O(n^3)$ time per step; needs $H\succ0$ (otherwise it can climb or stop at a saddle); only reliable near the solution (it can overshoot and diverge: for $\sqrt{1+x^2}$ it gives $x\to-x^3$). Fixes: damping and backtracking line search; Levenberg shift $(H+\tau I)\mathbf{d}=-\mathbf{g}$ with $\tau>-\lambda_{\min}$; trust regions.
- Cost: $n=10^3$: 8 MB, 0.3 ms; $n=10^6$: 8 TB, days; $n=10^9$: 8 EB, millions of years. Hence first-order methods for deep nets.
- Quasi-Newton: learn curvature from gradients. Secant condition $B_{k+1}\mathbf{s}=\mathbf{y}$ (equivalently $C_{k+1}\mathbf{y}=\mathbf{s}$), which needs $\mathbf{s}^\top\mathbf{y}>0$. BFGS: $C_{k+1}=(I-\rho\mathbf{s}\mathbf{y}^\top)C_k(I-\rho\mathbf{y}\mathbf{s}^\top)+\rho\mathbf{s}\mathbf{s}^\top$, $\rho=1/\mathbf{y}^\top\mathbf{s}$. It satisfies the secant condition, stays positive definite when $\mathbf{s}^\top\mathbf{y}>0$, costs $O(n^2)$, and converges superlinearly. A Wolfe line search guarantees $\mathbf{s}^\top\mathbf{y}>0$.
- L-BFGS: keep only the last $m$ pairs and compute $C\mathbf{g}$ with the two-loop recursion: $O(mn)$ time and $2mn$ memory. The default for large smooth deterministic problems (logistic regression, CRFs). Not for noisy mini-batch gradients.
- Relatives: Gauss–Newton ($H\approx J^\top J$, plus a Levenberg shift for LM) for least squares; natural gradient ($F^{-1}$) and K-FAC (Kronecker-factored $F^{-1}$) for probabilistic models and deep nets; Adam as a diagonal preconditioner.
Cheat sheet
| Method | Update | Memory | Work per iteration | Convergence near the minimum | Needs |
|---|---|---|---|---|---|
| Gradient descent | $\mathbf{x}-\eta\mathbf{g}$ | $n$ | $O(n)$ | linear, rate $\approx\frac{\kappa-1}{\kappa+1}$ | gradient, a step size |
| Newton | $\mathbf{x}-H^{-1}\mathbf{g}$ (solve) | $n^2$ | $O(n^3)$ | quadratic | gradient and Hessian, $H\succ0$ |
| Damped / Levenberg Newton | $\mathbf{x}-\alpha(H+\tau I)^{-1}\mathbf{g}$ | $n^2$ | $O(n^3)$ | quadratic once $\alpha=1,\tau=0$ | as above, line search |
| BFGS | $\mathbf{x}-\alpha C_k\mathbf{g}$ | $n^2$ | $O(n^2)$ | superlinear | gradient, Wolfe line search |
| L-BFGS (memory $m$) | $\mathbf{x}-\alpha\,\text{two-loop}(\mathbf{g})$ | $2mn$ | $O(mn)$ | linear at least; often close to superlinear in practice | gradient, full batch, smooth $f$ |
| Gauss–Newton / LM | $(J^\top J+\tau I)\mathbf{d}=-J^\top\mathbf{r}$ | $n^2$ | $O(n^3)$ | fast when residuals are small | residuals and Jacobian |
import numpy as np
# ---------- 1. Newton-Raphson: square root of 2 (root of x^2 - 2) ----------
x = 1.0
for k in range(4):
x = x - (x * x - 2) / (2 * x) # x <- x - f(x)/f'(x)
print(k + 1, round(x, 7), f"{abs(x - 2 ** 0.5):.1e}") # 1 1.5 8.6e-02 | 2 1.4166667 2.5e-03 | 3 1.4142157 2.1e-06 | 4 1.4142136 1.6e-12 (digits double)
# ---------- 2. Newton's method for a minimum in 1D: f = x^4/4 + x^2/2 - x ----------
g = lambda x: x**3 + x - 1 # f'
h = lambda x: 3 * x**2 + 1 # f''
x = 0.0
for k in range(4):
x = x - g(x) / h(x)
print(round(x, 7)) # 0.6823396 (the minimum is 0.6823278)
# ---------- 3. Newton in n dimensions: solve H d = -g, never invert; one step on a quadratic ----------
A = np.array([[4., 1.], [1., 2.]]); b = np.array([1., 1.])
x0 = np.array([2., 2.])
g0 = A @ x0 - b # gradient [9. 5.]
d = np.linalg.solve(A, -g0) # Newton step, from a linear solve
print(g0, np.round(d, 3), np.round(x0 + d, 3)) # [9. 5.] [-1.857 -1.571] [0.143 0.429] = (1/7, 3/7), the exact minimum
# ---------- 4. negative curvature and the Levenberg shift: f = x^2 - y^2 + y^4/4 at (1, 0.05) ----------
y = 0.05
g = np.array([2.0, -2 * y + y**3]); H = np.diag([2.0, -2 + 3 * y**2])
print(np.linalg.eigvalsh(H)) # [-1.9925 2. ] not positive definite
print(np.round(np.linalg.solve(H, -g), 4)) # [-1. -0.0501] pure Newton: heads for the saddle at (0, 0)
tau = 3.0
print(np.round(np.linalg.solve(H + tau * np.eye(2), -g), 4)) # [-0.4 0.0991] Levenberg: y moves away from the saddle, f drops
# ---------- 5. the BFGS update: secant condition and positive definiteness ----------
def bfgs_update(C, s, y):
rho = 1.0 / (y @ s); I = np.eye(len(s))
return (I - rho * np.outer(s, y)) @ C @ (I - rho * np.outer(y, s)) + rho * np.outer(s, s)
A = np.array([[4., 1.], [1., 2.]]); b = np.array([1., 1.])
x = np.array([2., 2.]); g = A @ x - b; C = np.eye(2)
for k in range(2):
d = -C @ g; alpha = -(g @ d) / (d @ A @ d) # exact line search on a quadratic
xn = x + alpha * d; gn = A @ xn - b; s = xn - x; y = gn - g
C = bfgs_update(C, s, y)
print(k, np.round(xn, 4), round(s @ y, 3), np.allclose(C @ y, s)) # 0 [-0.056 0.8578] 24.216 True then 1 [0.1429 0.4286] 0.356 True
x, g = xn, gn
print(np.round(C, 4)) # [[ 0.2857 -0.1429] [-0.1429 0.5714]] = inverse of A
print(np.linalg.eigvalsh(C) > 0) # [ True True ] stays positive definite
# ---------- 6. L-BFGS two-loop recursion = the explicit BFGS matrix times g (here with C0 = I) ----------
def two_loop(g, S, Y, gamma=1.0):
q = g.copy(); a = []
for s_, y_ in zip(reversed(S), reversed(Y)): # newest to oldest
a.append((s_ @ q) / (y_ @ s_)); q = q - a[-1] * y_
r = gamma * q
for (s_, y_), a_i in zip(zip(S, Y), reversed(a)): # oldest to newest
r = r + s_ * (a_i - (y_ @ r) / (y_ @ s_))
return r
rng = np.random.default_rng(0); n = 6
Q = rng.normal(size=(n, n)); A = Q @ Q.T + n * np.eye(n); b = rng.normal(size=n)
x = rng.normal(size=n); g = A @ x - b; C = np.eye(n); S, Y = [], []
for k in range(4):
d = -C @ g; alpha = -(g @ d) / (d @ A @ d)
xn = x + alpha * d; gn = A @ xn - b; s = xn - x; y = gn - g
C = bfgs_update(C, s, y); S.append(s); Y.append(y); x, g = xn, gn
print(np.allclose(two_loop(g, S, Y), C @ g)) # True: no n x n matrix is needed
# ---------- 7. what Newton costs (8-byte numbers, 10^12 operations per second) ----------
for n in (1e3, 1e6, 1e9):
print(f"n={n:.0e}: Hessian {8 * n**2 / 1e12:.3g} TB, solve {n**3 / 3 / 1e12:.3g} seconds")
# n=1e+03: Hessian 8e-06 TB, solve 0.000333 seconds | n=1e+06: Hessian 8 TB, solve 3.33e+05 seconds (3.9 days) | n=1e+09: Hessian 8e+06 TB (8 EB), solve 3.33e+14 seconds
# ---------- 8. SciPy on the Rosenbrock valley: iterations ----------
from scipy.optimize import minimize, rosen, rosen_der, rosen_hess
x0 = np.array([-1.2, 1.0])
for method in ("BFGS", "L-BFGS-B", "Newton-CG"):
r = minimize(rosen, x0, jac=rosen_der, hess=rosen_hess if method == "Newton-CG" else None, method=method, tol=1e-8)
print(method, r.nit, np.round(r.x, 4)) # BFGS 34 [1. 1.] | L-BFGS-B 36 [1. 1.] | Newton-CG 85 [1. 1.] (counts can differ a little between SciPy versions; its stopping rules and line searches differ from ours)
1. One Newton–Raphson step for $f(x)=x^2-5$ from $x_0=3$ gives…
2. How many Newton steps does it take to minimise a convex quadratic $f(\mathbf{x})=\tfrac12\mathbf{x}^\top A\mathbf{x}-\mathbf{b}^\top\mathbf{x}$ with $A\succ0$, from any starting point?
3. The Hessian at the current point has eigenvalues $-1.9$ and $+2$. What can go wrong with a pure Newton step, and what is a standard fix?
4. The secant condition of a quasi-Newton method says that the new Hessian estimate $B_{k+1}$ must satisfy…
5. A model has $n=10^6$ parameters. About how much memory does L-BFGS with $m=10$ pairs need for its curvature information, in 8-byte numbers?
6. Which problem is the best fit for L-BFGS?
Practice problems
A. Use Newton–Raphson on $f(x)=x^3-2$ from $x_0=1$ to approximate $\sqrt[3]2=1.259921$. Do two steps and give the errors.
$f'(x)=3x^2$. Step 1: $x_1=1-\frac{1-2}{3}=1.3333$ (error $0.0734$). Step 2: $f(1.3333)=2.3704-2=0.3704$ and $f'=3(1.7778)=5.3333$, so $x_2=1.3333-0.0694=1.26389$ (error $0.0040$). Step 3 would give $1.259933$ (error $1.2\times10^{-5}$) and step 4 about $10^{-10}$: the errors $0.073\to0.004\to10^{-5}\to10^{-10}$ show the digits roughly doubling.
B. Minimise $f(x)=x+4/x$ for $x>0$ with Newton's method from $x_0=1$. Do two steps. (The minimum is at $x=2$.)
$f'(x)=1-\frac{4}{x^2}$, $f''(x)=\frac{8}{x^3}$. At $x_0=1$: $f'=-3$, $f''=8$, so $x_1=1+\frac38=1.375$. At $x_1$: $f'=1-\frac{4}{1.8906}=-1.1157$ and $f''=\frac{8}{2.5996}=3.0775$, so $x_2=1.375+0.3625=1.7375$. Next: $x_3=1.9506$, $x_4=1.99818$, $x_5=1.9999975$. Errors $1,\ 0.625,\ 0.2625,\ 0.049,\ 0.0018,\ 2.5\times10^{-6}$: slow at first (far from the minimum, where $f''$ is small) then quadratic.
C. $f(x,y)=x^2+xy+2y^2-3x-y$. Take one Newton step from the origin and verify that it lands on the minimum.
$\nabla f=[2x+y-3,\ x+4y-1]$, so $\mathbf{g}(0,0)=[-3,\,-1]$. The Hessian is $H=\begin{bmatrix}2&1\\1&4\end{bmatrix}$ (constant). Solve $H\mathbf{d}=[3,\,1]$: from the second equation $d_1=1-4d_2$; substituting, $2(1-4d_2)+d_2=3\Rightarrow-7d_2=1\Rightarrow d_2=-\frac17$ and $d_1=1+\frac47=\frac{11}{7}$. New point $(\frac{11}{7},-\frac17)$. Check: $\nabla f=[\frac{22}{7}-\frac17-3,\ \frac{11}{7}-\frac47-1]=[0,\,0]$ ✓.
D. Apply the BFGS inverse update with $C=I$, $\mathbf{s}=[1,\,1]$, $\mathbf{y}=[3,\,1]$. Verify the secant condition and positive definiteness.
$\mathbf{s}^\top\mathbf{y}=4>0$, $\rho=\frac14$. $I-\rho\mathbf{s}\mathbf{y}^\top=\begin{bmatrix}0.25&-0.25\\-0.75&0.75\end{bmatrix}$ and $I-\rho\mathbf{y}\mathbf{s}^\top=\begin{bmatrix}0.25&-0.75\\-0.25&0.75\end{bmatrix}$. Their product is $\begin{bmatrix}0.125&-0.375\\-0.375&1.125\end{bmatrix}$ and $\rho\mathbf{s}\mathbf{s}^\top=\begin{bmatrix}0.25&0.25\\0.25&0.25\end{bmatrix}$. So $C_{\text{new}}=\begin{bmatrix}0.375&-0.125\\-0.125&1.375\end{bmatrix}$. Secant: $C_{\text{new}}\mathbf{y}=[0.375\cdot3-0.125,\ -0.125\cdot3+1.375]=[1,\,1]=\mathbf{s}$ ✓. Positive definite: trace $1.75>0$ and determinant $0.375\cdot1.375-0.125^2=0.5>0$ ✓.
E. A model has $n=5\times10^4$ parameters. How much memory does a dense Hessian need, how long does one Cholesky solve take at $10^{12}$ operations per second, and what do the same two numbers look like for L-BFGS with $m=10$?
Hessian: $n^2=2.5\times10^9$ numbers $\times8$ bytes $=20$ GB. Solve: $n^3/3=4.2\times10^{13}$ operations, about $42$ seconds per iteration. L-BFGS: $2mn=10^6$ numbers $=8$ MB, and about $8mn=4\times10^6$ operations, $4$ microseconds per direction (plus the cost of the gradient and the line search). L-BFGS is about $2500$ times smaller in memory and $10^7$ times cheaper per direction.
F. At a point $H=\mathrm{diag}(3,\,-1)$ and $\mathbf{g}=[1,\,2]$. Compute the pure Newton step and the Levenberg step with $\tau=2$. Which is a descent direction? Which values of $\tau$ make $H+\tau I$ positive definite?
Pure Newton: $\mathbf{d}=-H^{-1}\mathbf{g}=-[\frac13,\,\frac{2}{-1}]=[-0.333,\,+2]$ and $\mathbf{g}^\top\mathbf{d}=-0.333+4=3.67>0$: uphill. Levenberg: $H+2I=\mathrm{diag}(5,\,1)$, $\mathbf{d}=-[\frac15,\,2]=[-0.2,\,-2]$ and $\mathbf{g}^\top\mathbf{d}=-0.2-4=-4.2<0$: downhill ✓. $H+\tau I=\mathrm{diag}(3+\tau,\,-1+\tau)$ is positive definite exactly when $\tau>1$ (so $\tau=2$ works, and $\tau=1.05$ would work but give a very long step in the second coordinate: $-2/0.05=-40$).
Coordinate & Proximal Methods
Gradient descent needs a smooth landscape. But the most useful penalty in machine learning, the L1 penalty, has a sharp corner. This chapter gives you two tools for such problems: move one variable at a time (coordinate descent) and take a gradient step, then clean up with a small "proximal" problem (proximal gradient). Together they solve the Lasso, and they explain why Lasso answers contain exact zeros.
- Minimise a function one coordinate at a time, in cyclic or random order, and know when this is great and when it stalls or zig-zags
- Define the proximal operator and compute it for four common functions (zero, a convex set, $\lambda|x|$, $\tfrac12\lambda x^2$)
- Derive soft thresholding from the optimality condition and see why it creates exact zeros
- Derive the proximal gradient step from the quadratic upper bound, for $\min f(\mathbf{x}) + g(\mathbf{x})$ with $f$ smooth and $g$ simple but non-smooth
- Solve the Lasso with ISTA, with FISTA (acceleration, awareness) and with coordinate descent, step by step with real numbers, and read the active set and the solution path
Coordinate descent: one variable at a time core
Picture a city whose roads run only north-south and east-west. You want to reach the lowest point of the city, but you may only walk along one road at a time. So you do this: walk along an east-west road to the lowest spot on that road. Then turn, and walk along a north-south road to the lowest spot on that road. Repeat.
Or think of a mixing desk with two knobs. Turn the first knob until the sound is best, then the second knob until the sound is best, then the first knob again. Each turn is a tiny, easy problem with one unknown. That is coordinate descent.
Gradient descent changes all the variables at once, using one big arrow. Coordinate descent changes one variable per step and solves that step exactly.
Minimise $f(x_1, x_2) = x_1^2 + x_1x_2 + x_2^2 - 3x_1 - 3x_2$, starting at $(0, 0)$ (the answer is $(1, 1)$, with $f = -3$).
- Move $x_1$, hold $x_2 = 0$. The function of $x_1$ alone is $x_1^2 - 3x_1$. Its slope is $2x_1 - 3$, which is zero at $x_1 = 1.5$. New point $(1.5,\,0)$, $f = -2.25$.
- Move $x_2$, hold $x_1 = 1.5$. The slope in $x_2$ is $x_1 + 2x_2 - 3 = 1.5 + 2x_2 - 3$, zero at $x_2 = 0.75$. New point $(1.5,\,0.75)$, $f = -2.8125$.
- Move $x_1$ again, hold $x_2 = 0.75$. Slope $2x_1 + 0.75 - 3 = 0$ gives $x_1 = 1.125$. Point $(1.125,\,0.75)$, $f = -2.953125$.
- Move $x_2$ again. Slope $1.125 + 2x_2 - 3 = 0$ gives $x_2 = 0.9375$. Point $(1.125,\,0.9375)$, $f = -2.98828\ldots$
Every step lowers $f$ and the point creeps towards $(1, 1)$: from the second sweep on, each full sweep makes the distance to the answer four times smaller ($0.56$, then $0.14$, then $0.035$). Notice that each step used a one-line formula.
Coordinate descent. To minimise $f(x_1, \dots, x_n)$, repeat: choose an index $i$, then replace only the $i$-th variable by the value that minimises $f$ while all the others stay fixed,
$$x_i \leftarrow \operatorname*{argmin}_{t}\; f(x_1, \dots, x_{i-1},\, t,\, x_{i+1}, \dots, x_n).$$- Cyclic order: $i = 1, 2, \dots, n, 1, 2, \dots$. One pass through all $n$ coordinates is called a sweep (or epoch).
- Random order: pick $i$ at random at every step. This is called randomised coordinate descent; it has cleaner convergence theory and avoids unlucky orderings.
- Exact one-dimensional update for a quadratic. For $f(\mathbf{x}) = \tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x}$ (with $A$ symmetric positive definite, see positive definite matrices), setting $\partial f/\partial x_i = \sum_j A_{ij}x_j - b_i = 0$ and solving for $x_i$ gives $$x_i \leftarrow \frac{b_i - \sum_{j \ne i} A_{ij}\,x_j}{A_{ii}}.$$ Check with the example: $A = \begin{bmatrix}2&1\\1&2\end{bmatrix}$, $\mathbf{b} = [3, 3]$, so $x_1 \leftarrow (3 - x_2)/2$. At $x_2 = 0$ that is $1.5$ ✓. (In linear algebra this exact loop is called the Gauss–Seidel method.)
- Each step never increases $f$: the old value of $x_i$ was one of the candidates in the argmin.
Why do we need it?
Some problems have a huge number of variables, or a penalty with sharp corners, so a full gradient step is clumsy. Solving for one variable is often a one-line formula, and it needs only a tiny slice of the data.
Where is it used?
The Lasso and Elastic Net solvers in scikit-learn and glmnet, the dual solvers for linear SVMs (LIBLINEAR), and alternating least squares for recommender systems (which updates whole blocks of variables at a time).
How is it used?
Loop over the variables. For each one, solve the one-variable problem (exactly if you can, otherwise one cheap step), update a running summary such as the residual, and stop when a full sweep barely changes anything.
"Exact" is for one variable only. Coordinate descent solves each 1-D problem exactly, but it never sees the other variables moving. The overall answer is only reached after many sweeps.
Order can matter. In the cyclic order a bad ordering of strongly linked variables slows things down. Random order is the safe default in theory; cyclic is often a little faster in practice and is what most libraries use.
Quick check: $f(x_1,x_2) = x_1^2 + x_2^2$. Start at $(3, 4)$. How many coordinate steps reach the minimum?
Two. Step 1 sets $x_1 = 0$ (the slope $2x_1$ is zero there): $(0, 4)$. Step 2 sets $x_2 = 0$: $(0, 0)$. The variables are not linked (no $x_1x_2$ term), so each one can be fixed on its own.
When coordinate descent is great, and when it struggles
Great when the knobs do not fight each other. If setting one knob does not change what the best setting of the others is, each knob can be set once and for all. A function like $f = h_1(x_1) + h_2(x_2) + \dots + h_n(x_n)$ (a sum of one-variable pieces, called separable) is solved in a single sweep.
Slow when the knobs are strongly linked. If the best $x_1$ depends heavily on $x_2$, and the best $x_2$ depends heavily on $x_1$, then each exact move undoes most of the previous one. You make many tiny zig-zag steps along a diagonal valley.
Stuck when sharp corners and links meet. If $f$ is not smooth and its variables are linked inside the non-smooth part, you can arrive at a point where no single coordinate move helps, although a diagonal move would. Then coordinate descent stops at a wrong answer.
The L1 penalty is separable, which is the good case. The Lasso penalty $\lambda\|\mathbf{x}\|_1 = \lambda|x_1| + \dots + \lambda|x_n|$ has no links between the variables. The data-fit part $\tfrac12\|A\mathbf{x} - \mathbf{b}\|^2$ is smooth. A smooth part plus a separable non-smooth part is exactly the case where coordinate descent is known to work (see the definition below).
Zig-zag in numbers. Take $f = \tfrac12(x_1^2 + 2\rho x_1x_2 + x_2^2)$. The exact update is $x_1 \leftarrow -\rho x_2$, then $x_2 \leftarrow -\rho x_1$. Start at $(1, 1)$ with $\rho = 0.9$:
- $x_1 \leftarrow -0.9\cdot 1 = -0.9$, then $x_2 \leftarrow -0.9\cdot(-0.9) = 0.81$.
- $x_1 \leftarrow -0.9\cdot 0.81 = -0.729$, then $x_2 \leftarrow -0.9\cdot(-0.729) = 0.6561$.
- Each sweep multiplies $x_2$ by $\rho^2 = 0.81$. To shrink by a factor of 1000 you need about $\ln(1000)/\ln(1/0.81) \approx 33$ sweeps.
With $\rho = 0.3$ the factor is $0.09$ and about 3 sweeps are enough.
Facts about coordinate descent (the first is a statement of arithmetic, the others are known results, quoted with their assumptions):
- Separable $f$: if $f(\mathbf{x}) = \sum_i h_i(x_i)$, one sweep gives the exact minimiser, because every coordinate problem is independent.
- Smooth + separable non-smooth: for $F(\mathbf{x}) = f(\mathbf{x}) + \sum_i g_i(x_i)$ with $f$ smooth and convex and each $g_i$ convex, cyclic coordinate descent produces iterates whose limit points are minimisers (a classical result of Tseng, 2001, which needs technical conditions such as bounded level sets). This covers the Lasso, where the standard conditions hold. It is a statement about convergence, not about speed.
- Strongly linked quadratics: for the 2-variable quadratic with unit diagonal and coupling $\rho$ above, every cyclic sweep (after the first) shrinks the distance to the minimum by $\rho^2$. Gradient descent with its best fixed step shrinks it by $\rho$ per step (condition number $\kappa = (1+\rho)/(1-\rho)$ and rate $(\kappa-1)/(\kappa+1) = \rho$, see Chapter 3.5). For a dense quadratic one sweep costs about as much as one gradient, so CD is cheaper per unit of work here, but both slow down badly as $\rho \to 1$.
- Non-smooth and linked: can stall. Example: $F(x, y) = |x + y| + 3|x - y|$. At $(1,1)$ moving $x$ alone gives $|2 + t| + 3|t| \ge 2$ and moving $y$ alone gives the same, so $(1,1)$ is a "coordinate-wise minimum". Yet $F(1,1) = 2$ while $F(0,0) = 0$.
Why do we need it?
You need to know in advance whether coordinate descent will be a fast, cheap solver or a slow or stuck one, before you spend a week on a problem it cannot solve.
Where is it used?
The good case: Lasso, Elastic Net and sparse logistic regression on very wide data (many features). The slow case: highly correlated features. The stuck case: penalties that tie variables together, such as the fused or group-overlap penalties.
How is it used?
Check two things: is the non-smooth part a sum of one-variable pieces? Are the features strongly correlated? If the first is yes, try coordinate descent. If the second is yes, expect more sweeps, and consider standardising the features first.
"No coordinate move helps" is not the same as "this is the minimum". For smooth $f$ it is (the gradient is zero in every coordinate, so it is zero). For non-smooth $f$ with linked variables it is not, as the picture shows.
Do not oversell speed. Coordinate descent is cheap per step, not magically good. It needs many sweeps when features are strongly correlated.
Quick check: in which case does one sweep of coordinate descent solve the problem exactly: (a) $f = (x_1 - 2)^2 + (x_2 + 1)^2$, or (b) $f = (x_1 - x_2)^2 + x_1^2$?
(a). It is separable: each variable appears alone, so fixing $x_1 = 2$ and $x_2 = -1$ finishes the job. In (b) the term $(x_1 - x_2)^2$ links the variables, so the best $x_1$ depends on $x_2$ and one sweep is not enough.
The proximal operator: "stay close, but obey a rule"
You stand at a point $v$. A rule $g$ tells you which places are cheap (low $g$) and which are expensive (high $g$). You are attached to $v$ by a rubber band. Where do you end up?
- If the rule is very strong, you move far from $v$ to a cheap place.
- If the rubber band is strong, you stay close to $v$.
The final resting place is a balance between the two pulls. That resting place is the proximal operator ("prox" for short) of the rule $g$, applied to $v$. The word "proximal" just means nearby: we want a point near $v$ that is also good for $g$.
A step-size $\eta$ sets the strength of the rubber band: a large $\eta$ means a loose band (the rule wins), a small $\eta$ means a tight one (you barely move).
Take the rule $g(x) = |x|$ (it prefers points close to $0$) and $\eta = 1$. Find the resting place from $v$ by minimising $|x| + \tfrac12(x - v)^2$.
- $v = 3$. Try $x \gt 0$: the function is $x + \tfrac12(x-3)^2$ with slope $1 + (x - 3)$. Slope zero gives $x = 2$, which is positive, so it is valid. The rule pulled us from $3$ down to $2$.
- $v = 0.4$. Try $x \gt 0$: slope $1 + (x - 0.4) = 0$ gives $x = -0.6$, not positive, so it is invalid. Try $x \lt 0$: slope $-1 + (x - 0.4) = 0$ gives $x = 1.4$, not negative, invalid. The only candidate left is the corner $x = 0$. So the answer is exactly $0$.
Other rules (η = 1 unless said otherwise): $g = \tfrac12\lambda x^2$ with $\lambda = 2$, $\eta = 0.5$, $v = 3$ gives $x = v/(1 + \eta\lambda) = 3/2 = 1.5$. The rule "stay inside $[-1, 1]$" with $v = 3$ gives $x = 1$ (the nearest allowed point).
For a convex function $g$ and a step-size $\eta \gt 0$, the proximal operator of $\eta g$ is $$\operatorname{prox}_{\eta g}(\mathbf{v}) \;=\; \operatorname*{argmin}_{\mathbf{x}}\;\Big\{\, g(\mathbf{x}) + \frac{1}{2\eta}\,\|\mathbf{x} - \mathbf{v}\|^2 \,\Big\}.$$ The first term is the rule. The second term is the rubber band (a squared distance to $\mathbf{v}$). Because the squared distance is strongly convex, the minimiser exists and is unique for every $\mathbf{v}$.
| Rule $g$ | $\operatorname{prox}_{\eta g}(v)$ | In words |
|---|---|---|
| $g = 0$ | $v$ | no rule, stay where you are (the identity) |
| $g = $ indicator of a convex set $C$ ($0$ inside, $+\infty$ outside) | $\operatorname{proj}_C(v)$ | jump to the nearest point of $C$ (a projection), for any $\eta$ |
| $g(x) = \lambda|x|$ | $S_{\eta\lambda}(v)$ | soft thresholding (next section) |
| $g(x) = \tfrac12\lambda x^2$ | $\dfrac{v}{1 + \eta\lambda}$ | shrinkage: scale towards zero, never exactly to zero |
Derivation of the last row. Minimise $\tfrac12\lambda x^2 + \tfrac{1}{2\eta}(x - v)^2$. Set the slope to zero: $\lambda x + \tfrac1\eta(x - v) = 0$, so $x(\lambda + \tfrac1\eta) = \tfrac v\eta$, so $x = \dfrac{v}{1 + \eta\lambda}$ ✓.
Separable rules. If $g(\mathbf{x}) = \sum_i g_i(x_i)$ (for example $\lambda\|\mathbf{x}\|_1$), the objective splits into independent one-variable problems, so $\operatorname{prox}$ is applied entry by entry. That is why these proxes cost only $O(n)$.
"Simple" in "$g$ simple but non-smooth" means exactly this: its prox has a cheap formula.
Why do we need it?
A non-smooth rule has no gradient at its corners, so gradient descent cannot handle it. The prox replaces the missing gradient by a small, always-solvable problem with one unique answer.
Where is it used?
The L1 penalty (Lasso), projections onto constraints (projected gradient, which you will meet in Chapter 3.17), nuclear-norm penalties for low-rank matrices, total-variation denoising, and the "proximal" optimisers in sparse and constrained learning libraries.
How is it used?
Look up (or derive) the closed form of the prox for your penalty. Then it is one line of code inside every iteration: apply it to the point after the gradient step (next sections).
Prox is not "the gradient of $g$". It is the answer to a small minimisation problem. It exists even where $g$ has a corner, which is the whole point.
The step-size sits inside the prox. We write $\operatorname{prox}_{\eta g}$ because the rubber band's strength is $1/\eta$. Doubling $\eta$ is the same as doubling the strength of the rule $g$.
Other useful proxes (awareness). For $g(\mathbf{x}) = \lambda\|\mathbf{x}\|_2$ (not squared) the prox is $\max\!\big(0,\,1 - \eta\lambda/\|\mathbf{v}\|\big)\,\mathbf{v}$: it shrinks the whole vector towards $\mathbf{0}$ and can set it to exactly $\mathbf{0}$. This is the "group Lasso" building block. For example with $\mathbf{v} = [3, 4]$, $\eta\lambda = 1$ it gives $0.8\cdot[3, 4] = [2.4, 3.2]$.
Quick check: compute $\operatorname{prox}_{\eta g}(5)$ for $g(x) = \tfrac12 x^2$ with $\eta = 4$.
$\lambda = 1$, so $x = v/(1 + \eta\lambda) = 5/(1 + 4) = 1$. Check by slope: $x + \tfrac14(x - 5) = 1 + \tfrac14(-4) = 0$ ✓. A big $\eta$ (loose rubber band) lets the rule pull $v$ a long way.
Soft thresholding: the prox of the L1 penalty core
Imagine a tax. Everybody's number is pulled towards zero by the same amount $t$. A big number $5$ becomes $4$ (if $t = 1$). A negative number $-3$ becomes $-2$. But if your number is smaller than $t$ in size, the tax would push you past zero, so you simply end exactly at zero.
So there is a dead zone $[-t, t]$: every number inside it is wiped out to $0$. Numbers outside it survive, just a bit smaller. That is soft thresholding.
(The "hard" alternative, "keep it unchanged if big, zero it if small", jumps at the threshold. The "soft" version is continuous: no jump.)
Let $t = 1$ and $\mathbf{v} = [3,\ -0.5,\ 1.2,\ -2,\ 0.9]$.
- $3$: size $3 \gt 1$, so subtract $1$ keeping the sign: $2$.
- $-0.5$: size $0.5 \le 1$, so $0$.
- $1.2$: $1.2 - 1 = 0.2$.
- $-2$: size $2 \gt 1$, so $-(2 - 1) = -1$.
- $0.9$: size $0.9 \le 1$, so $0$.
Result: $S_1(\mathbf{v}) = [2,\ 0,\ 0.2,\ -1,\ 0]$. Two entries became exactly zero.
For a threshold $t \ge 0$, soft thresholding is $$S_t(v) = \operatorname{sign}(v)\,\max(|v| - t,\ 0) = \begin{cases} v - t & v \gt t \\ 0 & |v| \le t \\ v + t & v \lt -t. \end{cases}$$ It equals the proximal operator of the L1 penalty: $\operatorname{prox}_{\eta\lambda|\cdot|}(v) = S_{\eta\lambda}(v)$, with threshold $t = \eta\lambda$. For vectors, apply it to every entry.
Derivation from the optimality condition. We want the minimiser of $h(x) = t|x| + \tfrac12(x - v)^2$. (The prox objective $\lambda|x| + \tfrac{1}{2\eta}(x-v)^2$ equals $\tfrac1\eta$ times this with $t = \eta\lambda$, so it has the same minimiser.) The function is convex, so $x$ is a minimiser exactly when zero is among its slopes. The slope of $\tfrac12(x-v)^2$ is $x - v$. The slope of $|x|$ is $+1$ for $x \gt 0$, $-1$ for $x \lt 0$, and at the corner $x = 0$ every number in $[-1, 1]$ is a valid "slope" (a subgradient: the slope of a line that touches the corner from below, see Chapter 3.6). Now check the three cases of $x$:
- $x \gt 0$: need $t + (x - v) = 0$, so $x = v - t$. This is allowed only if $v - t \gt 0$, i.e. $v \gt t$.
- $x \lt 0$: need $-t + (x - v) = 0$, so $x = v + t$. Allowed only if $v \lt -t$.
- $x = 0$: need $0 \in [-t, t] + (0 - v)$, i.e. $v \in [-t, t]$, i.e. $|v| \le t$.
Exactly one case applies for each $v$, which gives the three-line formula. ∎
Two handy facts. $S_t(v) = v - \operatorname{clip}(v, -t, t)$ (subtract the clipped part). And, in contrast, the prox of the smooth $\tfrac12\lambda x^2$ is $v/(1 + \eta\lambda)$: it is never exactly $0$ unless $v = 0$. The flat piece of $S_t$ is the origin of exact zeros.
Why do we need it?
It is the missing one-line formula that lets us handle the sharp corner of $|x|$, and it is the reason a Lasso answer contains true zeros rather than tiny numbers. Without it there is no cheap way to minimise a loss plus an L1 penalty.
Where is it used?
Every step of ISTA and of coordinate descent for the Lasso; sparse coding and compressed sensing; wavelet denoising (shrink small coefficients to zero); sparse neural-network training with proximal updates.
How is it used?
In code: np.sign(v) * np.maximum(np.abs(v) - t, 0). Choose the threshold as $t = \eta\lambda$: step-size times penalty strength. A larger $t$ gives a bigger dead zone and more zeros.
The threshold is $\eta\lambda$, not $\lambda$. In an algorithm with step-size $\eta$ the dead zone has half-width $\eta\lambda$. This trips up many implementations.
It also shrinks the big values. Even entries that survive are pulled towards zero by $t$. This is a bias: Lasso coefficients are systematically a bit too small. (A common fix is to refit an ordinary regression on the surviving features.)
Quick check: what is $S_{0.5}$ applied to $[2,\ -0.3,\ 0.5,\ -1.5]$?
$2 \to 1.5$; $-0.3$ has size $\le 0.5$, so $0$; $0.5$ has size exactly $0.5 \le 0.5$, so $0$; $-1.5 \to -1$. Result $[1.5,\ 0,\ 0,\ -1]$.
Proximal gradient: a gradient step, then a prox step core
We want to minimise a sum $F = f + g$. Here $f$ is smooth (a data-fit loss: you can take its gradient) and $g$ is simple but non-smooth (an L1 penalty, or "stay inside this set").
Treat the two parts differently:
- Gradient step on $f$ only. Walk downhill as if the rule $g$ did not exist.
- Prox step for $g$. A "referee" nudges you to a point that is close to where you landed and that the rule $g$ likes: soft thresholding for L1, projection for a set.
Repeat. Each part does the job it is good at.
Minimise $F(x) = \tfrac12(x - 4)^2 + |x|$ in one variable. Here $f(x) = \tfrac12(x-4)^2$ ($\nabla f = x - 4$, so $L = 1$) and $g = |x|$. The true minimiser has slope $(x - 4) + 1 = 0$, so $x^\star = 3$. Use $\eta = 0.5$ and start at $x = 0$.
- Iteration 1. Gradient step: $z = 0 - 0.5\,(0 - 4) = 2$. Prox step (threshold $\eta\lambda = 0.5$): $S_{0.5}(2) = 1.5$.
- Iteration 2. $z = 1.5 - 0.5\,(1.5 - 4) = 2.75$. $S_{0.5}(2.75) = 2.25$.
- Iteration 3. $z = 2.25 - 0.5\,(2.25 - 4) = 3.125$. $S_{0.5}(3.125) = 2.625$.
The iterates $0,\ 1.5,\ 2.25,\ 2.625,\ \dots$ close half the remaining gap to $3$ each time. With $\eta = 1 = 1/L$ it would land on $3$ in one step: $z = 4$, $S_1(4) = 3$.
Proximal gradient method for $\min_{\mathbf{x}}\ F(\mathbf{x}) = f(\mathbf{x}) + g(\mathbf{x})$, with $f$ smooth and convex, $g$ convex with a cheap prox: $$\boxed{\;\mathbf{x}_{k+1} = \operatorname{prox}_{\eta g}\big(\mathbf{x}_k - \eta\,\nabla f(\mathbf{x}_k)\big)\;}$$ $\mathbf{z}_k = \mathbf{x}_k - \eta\nabla f(\mathbf{x}_k)$ is the gradient step; the prox is the clean-up.
Derivation from the quadratic upper bound. Suppose $f$ is $L$-smooth (its gradient changes at most at speed $L$, see Chapter 3.5). Then for every $\mathbf{y}$, $$f(\mathbf{y}) \le f(\mathbf{x}) + \nabla f(\mathbf{x})^\top(\mathbf{y} - \mathbf{x}) + \tfrac L2\|\mathbf{y} - \mathbf{x}\|^2 \le f(\mathbf{x}) + \nabla f(\mathbf{x})^\top(\mathbf{y} - \mathbf{x}) + \tfrac{1}{2\eta}\|\mathbf{y} - \mathbf{x}\|^2,$$ where the second inequality holds for any $\eta \le 1/L$ (then $\tfrac{1}{2\eta} \ge \tfrac L2$). Add $g(\mathbf{y})$ to both sides: $$F(\mathbf{y}) \le Q_{\mathbf{x}}(\mathbf{y}) := f(\mathbf{x}) + \nabla f(\mathbf{x})^\top(\mathbf{y} - \mathbf{x}) + \tfrac{1}{2\eta}\|\mathbf{y} - \mathbf{x}\|^2 + g(\mathbf{y}).$$ $Q_{\mathbf{x}}$ is a bowl that sits above $F$ and touches it at $\mathbf{y} = \mathbf{x}$. Minimise the bowl instead of $F$. Complete the square: $\nabla f^\top(\mathbf{y} - \mathbf{x}) + \tfrac{1}{2\eta}\|\mathbf{y} - \mathbf{x}\|^2 = \tfrac{1}{2\eta}\|\mathbf{y} - (\mathbf{x} - \eta\nabla f(\mathbf{x}))\|^2 - \tfrac\eta2\|\nabla f(\mathbf{x})\|^2$ (expand the right side to check). So $$Q_{\mathbf{x}}(\mathbf{y}) = g(\mathbf{y}) + \tfrac{1}{2\eta}\|\mathbf{y} - \mathbf{z}\|^2 + \text{(terms that do not depend on } \mathbf{y}),\qquad \mathbf{z} = \mathbf{x} - \eta\nabla f(\mathbf{x}),$$ and the minimiser of that is, by the definition of the prox, $\operatorname{prox}_{\eta g}(\mathbf{z})$. ∎
- Never goes uphill. Let $\mathbf{x}^+$ be the new point. Then $F(\mathbf{x}^+) \le Q_{\mathbf{x}}(\mathbf{x}^+) \le Q_{\mathbf{x}}(\mathbf{x}) = F(\mathbf{x})$ (the first step is the upper bound, the second is because $\mathbf{x}^+$ minimises $Q_{\mathbf{x}}$). So $F$ never increases when $\eta \le 1/L$.
- Three familiar cases. $g = 0$ (prox = identity) gives plain gradient descent. $g$ = indicator of a set gives projected gradient descent (you will meet it in Chapter 3.17). $g = \lambda\|\cdot\|_1$ gives ISTA, the topic of the Lasso sections below.
- The fixed-point test. A point is a minimiser of $F$ exactly when it does not move: $\mathbf{x}^\star = \operatorname{prox}_{\eta g}(\mathbf{x}^\star - \eta\nabla f(\mathbf{x}^\star))$. (Reason: $\mathbf{x} = \operatorname{prox}_{\eta g}(\mathbf{z})$ means $(\mathbf{z} - \mathbf{x})/\eta$ is a subgradient of $g$ at $\mathbf{x}$; with $\mathbf{z} = \mathbf{x}^\star - \eta\nabla f(\mathbf{x}^\star)$ this reads $-\nabla f(\mathbf{x}^\star) \in \partial g(\mathbf{x}^\star)$, i.e. $\mathbf{0} \in \nabla f(\mathbf{x}^\star) + \partial g(\mathbf{x}^\star)$.) So $\|\mathbf{x}_k - \mathbf{x}_{k+1}\|/\eta$ is a natural stopping measure.
- Rates (assumptions stated): if $f$ is convex and $L$-smooth, $g$ is convex, and $\eta = 1/L$, then $F(\mathbf{x}_k) - F^\star \le \dfrac{L\|\mathbf{x}_0 - \mathbf{x}^\star\|^2}{2k}$, an $O(1/k)$ rate, the same as plain gradient descent on smooth convex functions. If in addition $f$ is $\mu$-strongly convex, then $\|\mathbf{x}_k - \mathbf{x}^\star\| \le (1 - \mu/L)^k\|\mathbf{x}_0 - \mathbf{x}^\star\|$: linear convergence. The non-smooth part costs nothing in the rate.
Why do we need it?
Gradient descent fails when part of the objective has corners, and a plain "subgradient" method is slow and never produces exact zeros. Proximal gradient keeps gradient-descent speed and handles the corners exactly.
Where is it used?
Sparse regression and classification (Lasso, sparse logistic regression), constrained training (projected gradient), low-rank matrix completion (singular-value thresholding), image deblurring and denoising with L1 or total-variation penalties.
How is it used?
Split your objective into $f$ (smooth: you can differentiate it) and $g$ (non-smooth with a known prox). Estimate $L$, pick $\eta = 1/L$ (or use backtracking), and loop: gradient step on $f$, prox of $\eta g$. Stop when the point stops moving.
Use $\eta \le 1/L$ for the guarantee. A bigger step may still work on an easy problem, but the "never goes uphill" proof needs $\eta \le 1/L$. If you do not know $L$, use backtracking: shrink $\eta$ until the upper bound holds.
The prox is applied to $\mathbf{z}$, not to $\mathbf{x}_k$. Applying the prox first and the gradient step second is a different (and wrong) algorithm.
You need a cheap prox. If the prox itself is a hard optimisation problem, proximal gradient loses its appeal.
Quick check: $f(x) = \tfrac12(x-4)^2$, $g(x) = |x|$, $\eta = 1$, current point $x_k = 0$. What is $x_{k+1}$?
Gradient step: $z = 0 - 1\cdot(0 - 4) = 4$. Prox: $S_{1}(4) = 3$. And $3$ is the exact minimiser (slope $(3-4) + 1 = 0$). With $\eta = 1/L = 1$ one step is enough on this 1-D quadratic.
The Lasso objective and why zeros appear exactly core
You are predicting house prices from ten features, but you suspect only a few of them really matter. You want the model itself to choose: keep the useful features, drop the rest.
The Lasso does this by adding a price tag to every weight: each unit of weight costs $\lambda$. A feature is only worth keeping if it reduces the prediction error by more than its price. A feature that helps only a little cannot pay its price, so its weight is set to exactly zero, and the feature disappears from the model.
(How the L1 penalty compares to the L2 penalty, and what the diamond-shaped constraint picture means, is the topic of Chapter 3.11. Here we study how to compute the Lasso answer.)
A tiny regression with three examples and two features: $$A = \begin{bmatrix} 1 & 0 \\ 0 & 1 \\ 1 & 1 \end{bmatrix},\qquad \mathbf{b} = \begin{bmatrix} 3 \\ 0 \\ 2 \end{bmatrix},\qquad \lambda = 2.$$ Each row of $A$ is one example; the model predicts $A\mathbf{x}$ for weights $\mathbf{x} = [x_1, x_2]$. The objective is $F(\mathbf{x}) = \tfrac12\|A\mathbf{x} - \mathbf{b}\|^2 + \lambda(|x_1| + |x_2|)$.
- All weights zero: $F(0, 0) = \tfrac12(9 + 0 + 4) + 0 = 6.5$.
- The answer is $\mathbf{x}^\star = (1.5,\ 0)$. Predictions $A\mathbf{x}^\star = (1.5, 0, 1.5)$, residual $A\mathbf{x}^\star - \mathbf{b} = (-1.5, 0, -0.5)$, so $F = \tfrac12(2.25 + 0 + 0.25) + 2\cdot1.5 = 1.25 + 3 = 4.25$.
- Check that it is better than a tempting alternative that keeps both features. For example $(1.3, 0.1)$: residual $(-1.7, 0.1, -0.6)$, $F = \tfrac12(2.89 + 0.01 + 0.36) + 2\cdot1.4 = 1.63 + 2.8 = 4.43 \gt 4.25$.
- Why is $x_2$ exactly zero? Look at the slope of the smooth part with respect to $x_2$ at the answer: $\mathbf{a}_2^\top(A\mathbf{x}^\star - \mathbf{b}) = 0\cdot(-1.5) + 1\cdot 0 + 1\cdot(-0.5) = -0.5$. Raising $x_2$ lowers the error at rate $0.5$, but it costs $\lambda = 2$ per unit. The price is higher than the benefit, so $x_2$ stays at $0$.
Without the penalty ($\lambda = 0$) the best weights would be $\mathbf{x} = (8/3, -1/3)$, which keeps both features.
The Lasso problem (the name stands for least absolute shrinkage and selection operator) is $$\min_{\mathbf{x} \in \mathbb{R}^n}\ F(\mathbf{x}) = \underbrace{\tfrac12\|A\mathbf{x} - \mathbf{b}\|_2^2}_{f(\mathbf{x}):\ \text{smooth data fit}} + \underbrace{\lambda\|\mathbf{x}\|_1}_{g(\mathbf{x}):\ \text{simple, non-smooth}},\qquad \lambda \ge 0.$$ $A$ is the $m \times n$ data matrix (rows are examples), $\mathbf{b}$ the targets, $\mathbf{x}$ the weights. It has the proximal-gradient form $f + g$ of the previous section, with $\nabla f(\mathbf{x}) = A^\top(A\mathbf{x} - \mathbf{b})$ and $g$'s prox equal to soft thresholding.
- Convention note. scikit-learn minimises $\tfrac{1}{2m}\|\mathbf{b} - A\mathbf{x}\|^2 + \alpha\|\mathbf{x}\|_1$. That is the same problem with $\lambda = m\alpha$.
- Optimality conditions. Let $\mathbf{s} = \nabla f(\mathbf{x}^\star) = A^\top(A\mathbf{x}^\star - \mathbf{b})$ be the vector of slopes of the smooth part ($s_j$ is its slope in direction $j$). Then $\mathbf{x}^\star$ is a solution exactly when, for every coordinate $j$, $$\begin{cases} s_j = -\lambda\,\operatorname{sign}(x^\star_j) & \text{if } x^\star_j \ne 0,\\[2pt] |s_j| \le \lambda & \text{if } x^\star_j = 0. \end{cases}$$ (This is $\mathbf{0} \in \nabla f(\mathbf{x}^\star) + \lambda\,\partial\|\mathbf{x}^\star\|_1$, read coordinate by coordinate, exactly as in the soft-threshold derivation.) Check with the example: $\mathbf{s} = (-2, -0.5)$. For $x_1 = 1.5 \gt 0$: $s_1 = -2 = -\lambda$ ✓. For $x_2 = 0$: $|s_2| = 0.5 \le 2$ ✓.
- When is the answer all zeros? At $\mathbf{x} = \mathbf{0}$ the gradient is $-A^\top\mathbf{b}$, so $\mathbf{0}$ is optimal exactly when $\lambda \ge \lambda_{\max} := \|A^\top\mathbf{b}\|_\infty = \max_j|\mathbf{a}_j^\top\mathbf{b}|$. In the example $A^\top\mathbf{b} = (5, 2)$, so $\lambda_{\max} = 5$.
Why exact zeros? At $x_j = 0$ the penalty $\lambda|x_j|$ has a corner: its slope jumps from $-\lambda$ to $+\lambda$. The smooth part has slope $s_j$ there. If $|s_j| \le \lambda$, then moving $x_j$ in either direction makes the total go up (the penalty rises faster than the data fit can fall). The corner is a trap. A smooth penalty such as $\tfrac12\lambda x_j^2$ has slope $0$ at $0$, so it can only trap if $s_j$ is exactly zero.
Why do we need it?
With many features, ordinary regression keeps all of them, fits noise, and is hard to interpret. The Lasso builds feature selection into the training objective, and it is still a convex problem, so it can be solved reliably.
Where is it used?
Genomics (a few genes out of thousands), text models with huge vocabularies, sparse linear models in finance, compressed sensing, and as a first screening step before fitting a bigger model (scikit-learn's Lasso and LassoCV).
How is it used?
Standardise the features (so the penalty treats them fairly), choose $\lambda$ by cross-validation, solve with coordinate descent or ISTA/FISTA, and read the non-zero weights as the selected features.
Scale the features first. The penalty treats every weight equally. If one feature is measured in millions and another in fractions, the first needs a tiny weight and is almost free, the second needs a big weight and is heavily penalised. Standardise the columns of $A$.
"Zero weight" means "not selected at this $\lambda$". It does not prove the feature is useless. When two features are strongly correlated the Lasso tends to keep one of them and drop the other, somewhat arbitrarily.
Quick check: for the example, would $\lambda = 6$ give the all-zeros answer?
Yes. $\lambda_{\max} = \|A^\top\mathbf{b}\|_\infty = \max(5, 2) = 5$, and $6 \ge 5$. At $\mathbf{x} = 0$ the slopes of the smooth part are $(-5, -2)$ and both have size $\le 6$, so the corner traps every coordinate.
ISTA, and FISTA (acceleration)
Proximal gradient applied to the Lasso has its own name: ISTA, the iterative shrinkage-thresholding algorithm. Each iteration is a gradient step on the data-fit part, then a "shrink everything by a small amount and kill the tiny ones" step.
FISTA ("fast" ISTA) adds momentum, the same rolling-ball idea as in Chapter 3.4: the next gradient step is taken not at the latest point, but at a point pushed a little further along the direction you just travelled.
One full ISTA iteration by hand. Use the tiny example from the last section: $A = \begin{bmatrix}1&0\\0&1\\1&1\end{bmatrix}$, $\mathbf{b} = [3, 0, 2]$, $\lambda = 2$. First collect what we need:
- $A^\top A = \begin{bmatrix}2&1\\1&2\end{bmatrix}$, with eigenvalues $3$ and $1$, so $L = 3$. We choose $\eta = 0.25 \le 1/L$ (a clean number). The threshold is $\eta\lambda = 0.5$.
- $A^\top\mathbf{b} = [5, 2]$.
Start at $\mathbf{x}_0 = (0, 0)$.
- Gradient: $\nabla f = A^\top A\mathbf{x}_0 - A^\top\mathbf{b} = (0, 0) - (5, 2) = (-5, -2)$.
- Gradient step: $\mathbf{z} = \mathbf{x}_0 - \eta\nabla f = (0 + 1.25,\ 0 + 0.5) = (1.25,\ 0.5)$.
- Soft threshold with $0.5$: $S_{0.5}(1.25) = 0.75$. For the second entry $|0.5| \le 0.5$ (it sits exactly on the edge of the dead zone), so $S_{0.5}(0.5) = 0$.
- Result: $\mathbf{x}_1 = (0.75,\ 0)$. The residual is $A\mathbf{x}_1 - \mathbf{b} = (0.75, 0, 0.75) - (3, 0, 2) = (-2.25, 0, -1.25)$, so $F = \tfrac12(5.0625 + 1.5625) + 2\cdot0.75 = 3.3125 + 1.5 = 4.8125$, down from $6.5$.
The next iterations (same recipe):
| $k$ | $\nabla f(\mathbf{x}_{k-1})$ | $\mathbf{z}$ | $\mathbf{x}_k = S_{0.5}(\mathbf{z})$ | $F(\mathbf{x}_k)$ |
|---|---|---|---|---|
| 1 | $(-5,\ -2)$ | $(1.25,\ 0.5)$ | $(0.75,\ 0)$ | $4.8125$ |
| 2 | $(-3.5,\ -1.25)$ | $(1.625,\ 0.3125)$ | $(1.125,\ 0)$ | $4.3906$ |
| 3 | $(-2.75,\ -0.875)$ | $(1.8125,\ 0.21875)$ | $(1.3125,\ 0)$ | $4.2852$ |
The answer is $(1.5, 0)$ with $F^\star = 4.25$. The gap to $F^\star$ is $0.5625,\ 0.1406,\ 0.0352$: it falls by a factor of $4$ each time. And $x_2$ became exactly zero at the very first iteration and stayed there. (Cross-check: a general-purpose optimiser run on the same problem, scipy.optimize.minimize on the split form $\mathbf{x} = \mathbf{u} - \mathbf{v}$ with $\mathbf{u}, \mathbf{v} \ge 0$, returns $(1.5, 0)$ and $F^\star = 4.25$.)
ISTA for the Lasso (proximal gradient with $f = \tfrac12\|A\mathbf{x} - \mathbf{b}\|^2$, $g = \lambda\|\cdot\|_1$): $$\mathbf{x}_{k+1} = S_{\eta\lambda}\!\Big(\mathbf{x}_k - \eta\,A^\top(A\mathbf{x}_k - \mathbf{b})\Big),\qquad \eta = \frac1L,\quad L = \lambda_{\max}(A^\top A) = \sigma_{\max}(A)^2.$$ ($L$ is the largest eigenvalue of the Hessian $A^\top A$ of $f$, see eigenvalues; it equals the square of the largest singular value of $A$.) One iteration costs two matrix–vector products, $A\mathbf{x}$ and $A^\top(\cdot)$, so about $2mn$ operations. By the previous section: $F$ never rises, and $F(\mathbf{x}_k) - F^\star \le \dfrac{L\|\mathbf{x}_0 - \mathbf{x}^\star\|^2}{2k}$ (assumptions: $f$ convex and $L$-smooth, which holds here).
FISTA (Beck and Teboulle, 2009; awareness). Keep two sequences, the iterate $\mathbf{x}_k$ and an extrapolated point $\mathbf{y}_k$. Start with $\mathbf{y}_1 = \mathbf{x}_0$, $t_1 = 1$. Then for $k = 1, 2, \dots$: $$\mathbf{x}_k = S_{\eta\lambda}\big(\mathbf{y}_k - \eta\nabla f(\mathbf{y}_k)\big),\qquad t_{k+1} = \frac{1 + \sqrt{1 + 4t_k^2}}{2},\qquad \mathbf{y}_{k+1} = \mathbf{x}_k + \frac{t_k - 1}{t_{k+1}}\,(\mathbf{x}_k - \mathbf{x}_{k-1}).$$ The only changes from ISTA: the gradient is taken at $\mathbf{y}_k$, and $\mathbf{y}_{k+1}$ pushes past $\mathbf{x}_k$ by a fraction (growing towards $1$) of the last move. Under the same assumptions with $\eta = 1/L$, $F(\mathbf{x}_k) - F^\star \le \dfrac{2L\|\mathbf{x}_0 - \mathbf{x}^\star\|^2}{(k+1)^2}$, an $O(1/k^2)$ rate versus ISTA's $O(1/k)$, for essentially the same cost per iteration. Unlike ISTA, FISTA is not a descent method: $F(\mathbf{x}_k)$ can go up for a moment, like a ball overshooting.
Why do we need it?
ISTA is the simplest method that is guaranteed to solve the Lasso, and it works for any smooth loss plus L1, not only squared error. FISTA gets the same accuracy in far fewer iterations for almost no extra work.
Where is it used?
Large Lasso and sparse-logistic problems where the matrix is too big or too structured for coordinate descent, image deblurring and compressed sensing (where $A$ is a fast transform), and as a building block in deep unrolled networks (LISTA).
How is it used?
Estimate $L$ (the largest singular value squared, from a few power-iteration steps), set $\eta = 1/L$, and loop gradient step then soft threshold. For FISTA keep the extra point $\mathbf{y}$ and the number $t$. Watch the size of the update, or the optimality conditions, to stop.
FISTA is not monotone. The objective can rise from one iteration to the next. If you need guaranteed descent, use ISTA, or a "restart" variant that resets the momentum when the objective rises.
The rates are worst-case bounds. On a specific problem ISTA can converge much faster than $O(1/k)$ (for instance linearly when the active columns are well-conditioned), and FISTA is not always faster in practice.
$L$ matters. If your $\eta$ is much smaller than $1/L$ the method is slow. Compute $L = \sigma_{\max}(A)^2$ (for example with np.linalg.norm(A, 2)**2) instead of guessing.
Quick check: in the hand example, what would the threshold be if we used $\eta = 1/L = 1/3$ instead of $0.25$?
$\eta\lambda = (1/3)\cdot 2 = 2/3 \approx 0.667$. A smaller step makes smaller moves but a wider dead zone; the two effects cancel in the fixed point, which does not depend on $\eta$ (the answer is still $(1.5, 0)$).
Coordinate descent for the Lasso core
The Lasso is the ideal customer for coordinate descent. The L1 penalty is a sum of one-variable pieces, so it does not tie the weights together. The data-fit part is a smooth quadratic. So if we freeze nine weights and solve for the tenth, we get a tiny problem with one unknown and a closed-form answer: a soft threshold.
Think of the one weight as a person deciding whether to join the model. They look at how well their feature explains what the other weights have not explained yet (the partial residual). If the explained amount is smaller than the price $\lambda$, they stay out (exactly $0$). Otherwise they join, with a weight reduced by the price.
One coordinate-descent sweep by hand. Same tiny problem: $A = \begin{bmatrix}1&0\\0&1\\1&1\end{bmatrix}$, $\mathbf{b} = [3, 0, 2]$, $\lambda = 2$. The columns are $\mathbf{a}_1 = (1, 0, 1)$ and $\mathbf{a}_2 = (0, 1, 1)$, with squared lengths $c_1 = c_2 = 2$. Start from $\mathbf{x} = (0, 0)$, so the residual is $\mathbf{r} = \mathbf{b} - A\mathbf{x} = (3, 0, 2)$.
- Coordinate 1. $u_1 = \mathbf{a}_1^\top\mathbf{r} + c_1x_1 = (3 + 0 + 2) + 2\cdot0 = 5$. New value: $x_1 = S_\lambda(u_1)/c_1 = S_2(5)/2 = 3/2 = 1.5$. Update the residual: $\mathbf{r} \leftarrow \mathbf{r} - \mathbf{a}_1(1.5 - 0) = (1.5,\ 0,\ 0.5)$.
- Coordinate 2. $u_2 = \mathbf{a}_2^\top\mathbf{r} + c_2x_2 = (0 + 0 + 0.5) + 2\cdot0 = 0.5$. Since $|0.5| \le \lambda = 2$, we get $x_2 = S_2(0.5)/2 = 0$. Nothing changes.
After one sweep, $\mathbf{x} = (1.5, 0)$. That is already the exact answer (we checked its optimality conditions earlier). A second sweep recomputes $u_1 = (1.5 + 0.5) + 2\cdot1.5 = 5$ and $u_2 = 0.5$ and changes nothing. Compare: ISTA gets closer and closer (the gap to the optimum is $0.56$, $0.14$, $0.035$, and so on) but never lands exactly on $1.5$.
A second problem where the sweeps need to repeat is in the practice section.
The coordinate update for the Lasso. Let $\mathbf{a}_j$ be the $j$-th column of $A$ and $c_j = \|\mathbf{a}_j\|^2$. Freeze all weights except $x_j$. The partial residual is $\mathbf{r}_{-j} = \mathbf{b} - \sum_{k \ne j}\mathbf{a}_kx_k$ and the one-variable problem is $$\min_{t}\ \tfrac12\|\mathbf{r}_{-j} - t\,\mathbf{a}_j\|^2 + \lambda|t| = \min_t\ \tfrac12c_jt^2 - u_jt + \lambda|t| + \text{const},\qquad u_j := \mathbf{a}_j^\top\mathbf{r}_{-j}.$$ (Expand the square: $\tfrac12\|\mathbf{r}_{-j}\|^2 - t\,\mathbf{a}_j^\top\mathbf{r}_{-j} + \tfrac12t^2\|\mathbf{a}_j\|^2$.) Factor out $c_j$: $\tfrac12c_jt^2 - u_jt = c_j\cdot\tfrac12\big(t - u_j/c_j\big)^2 + \text{const}$, so the problem is $c_j\Big[\tfrac12(t - u_j/c_j)^2 + \tfrac{\lambda}{c_j}|t|\Big]$, which is exactly the soft-threshold problem with threshold $\lambda/c_j$ and centre $u_j/c_j$. Therefore $$\boxed{\;x_j \leftarrow \frac{S_\lambda(u_j)}{c_j},\qquad u_j = \mathbf{a}_j^\top\mathbf{r} + c_j\,x_j^{\text{old}}\;}$$ where $\mathbf{r} = \mathbf{b} - A\mathbf{x}$ is the full residual (since $\mathbf{r}_{-j} = \mathbf{r} + \mathbf{a}_jx_j^{\text{old}}$). After the update, fix the residual cheaply: $\mathbf{r} \leftarrow \mathbf{r} - \mathbf{a}_j(x_j^{\text{new}} - x_j^{\text{old}})$.
- Cost. One coordinate: two length-$m$ dot products or updates, $O(m)$. One sweep over all $n$ coordinates: $O(mn)$, about the same as one ISTA iteration. And if $x_j$ does not change, no residual update is needed.
- Standardised features. If every column has $c_j = m$, the update is simply $x_j \leftarrow S_\lambda(u_j)/m$. This is the loop inside scikit-learn's
Lassoand in glmnet (they also add many speed tricks). - Why it works. The penalty is separable and the loss is smooth, so the convergence result of the earlier section applies.
Why do we need it?
It is usually the fastest way to solve the Lasso on tabular data: each update is a one-line formula, uses only one column of the data, and exact zeros come out for free.
Where is it used?
scikit-learn's Lasso, ElasticNet and LassoCV, R's glmnet, and sparse logistic regression solvers. It scales to hundreds of thousands of features, especially with sparse data.
How is it used?
Standardise the columns, start from $\mathbf{x} = \mathbf{0}$ (or the previous solution for a nearby $\lambda$), keep the residual up to date, sweep over the coordinates, and stop when a sweep changes no weight by more than a small tolerance.
Use the full residual and fix it after every change. Forgetting to update $\mathbf{r}$ after changing $x_j$ is the most common bug: the next coordinate then sees a stale residual.
The $+\,c_jx_j^{\text{old}}$ term is easy to forget. It adds the old contribution of feature $j$ back into the residual, so that $u_j$ really is the correlation with the partial residual.
Equal-length columns make the penalty fair. If $c_j$ differ a lot, the same $\lambda$ penalises small-scale features far more. Standardise first.
Quick check: in the hand example, if the residual were $\mathbf{r} = (1.5, 0, 0.5)$, $x_1 = 1.5$, what is $u_1$, and what does the update give?
$u_1 = \mathbf{a}_1^\top\mathbf{r} + c_1x_1 = (1.5 + 0.5) + 2\cdot1.5 = 5$. The update is $S_2(5)/2 = 1.5$, the same value: the coordinate is already optimal, so nothing moves.
The active set and the solution path
At the answer most weights are zero. The weights that are not zero are the active set: the features currently "in the model". Once an algorithm has found out which features are inactive, it can ignore them and spend all its effort on the few active ones. That is why the Lasso solvers get fast: the problem shrinks as the iterations go on.
Now turn the penalty knob $\lambda$ from huge to tiny. At a huge $\lambda$ everything is zero. As $\lambda$ falls, features enter the model one by one, the strongest first. The picture of every weight against $\lambda$ is the solution path.
Use the optimality conditions on the tiny example ($A^\top A = \begin{bmatrix}2&1\\1&2\end{bmatrix}$, $A^\top\mathbf{b} = (5, 2)$) and slide $\lambda$ downwards. Solving the active coordinates and testing the inactive ones gives:
| $\lambda$ | answer $(x_1, x_2)$ | active set | test for the inactive coordinate |
|---|---|---|---|
| $\ge 5$ | $(0,\ 0)$ | empty | $|{-5}| \le \lambda$ and $|{-2}| \le \lambda$ |
| $3.5$ | $(0.75,\ 0)$ | $\{1\}$ | $x_1 = (5-3.5)/2$; $|x_1 - 2| = 1.25 \le 3.5$ ✓ |
| $2$ | $(1.5,\ 0)$ | $\{1\}$ | $|1.5 - 2| = 0.5 \le 2$ ✓ |
| $0.5$ | $(2.25,\ 0)$ | $\{1\}$ | $|2.25 - 2| = 0.25 \le 0.5$ ✓ |
| $0.2$ | $(2.467,\ -0.133)$ | $\{1, 2\}$ | – |
| $0$ | $(2.667,\ -0.333)$ | $\{1, 2\}$ | ordinary least squares |
With only $x_1$ active, $x_1 = (5 - \lambda)/2$ and the test for $x_2$ reads $|x_1 - 2| = |1 - \lambda|/2 \le \lambda$. This stops being true when $(1-\lambda)/2 = \lambda$, i.e. at $\lambda = 1/3$. So feature 2 enters the model only when $\lambda$ falls below $1/3$, and with a negative weight. Between the entry points the answer changes in straight lines (here $x_1 = (5-\lambda)/2$): the path is piecewise linear.
The active set of $\mathbf{x}$ is $\mathcal{A}(\mathbf{x}) = \{\,j : x_j \ne 0\,\}$. The solution path is the map $\lambda \mapsto \mathbf{x}^\star(\lambda)$ of Lasso solutions. Known facts:
- The path starts at $\mathbf{0}$ for $\lambda \ge \lambda_{\max} = \|A^\top\mathbf{b}\|_\infty$; the first feature to enter is the one with the largest $|\mathbf{a}_j^\top\mathbf{b}|$ (for columns of equal length, the one most correlated with the target).
- The path is continuous and piecewise linear in $\lambda$ (this is what the LARS algorithm exploits to compute the whole path exactly).
- Warm starts. To get solutions for many $\lambda$ values (as in cross-validation), solve on a decreasing grid, starting each solve from the previous answer. Neighbouring answers are close, so each solve takes only a few sweeps.
- Active-set strategy. Sweep only over the active coordinates until they converge, then check the optimality test $|\mathbf{a}_j^\top\mathbf{r}| \le \lambda$ for the inactive ones. Any that fail join the active set. Because the active set is small, most sweeps are cheap. (Heuristic "screening rules" can also discard features that are very unlikely to be active, then verify at the end.)
Why do we need it?
You almost never know the right $\lambda$ in advance, and most of the work in a Lasso solver is wasted on weights that end up zero. Warm starts and active sets make "solve for 100 values of $\lambda$" cost little more than one solve.
Where is it used?
LassoCV and lasso_path in scikit-learn, cv.glmnet in R, and feature-selection pipelines on genomics and text data where thousands of features compete.
How is it used?
Build a geometric grid of $\lambda$ values from $\lambda_{\max}$ down to about $10^{-3}\lambda_{\max}$, solve along the grid with warm starts, pick the $\lambda$ with the best validation score, and read the active set at that $\lambda$ as the selected features.
A feature leaving the model is possible too. Along the path a weight can come back to zero after having entered (it is rare for well-behaved data, but it can happen when features are correlated).
Noise features can enter early. Lasso selection is not perfect: with strong correlation or a weak signal, noise features enter before weak true ones. Choose $\lambda$ by validation, not by hope.
Quick check: for the tiny example, what is the smallest $\lambda$ for which $\mathbf{x}^\star = (0,0)$?
$\lambda_{\max} = \|A^\top\mathbf{b}\|_\infty = \max(|5|, |2|) = 5$. For $\lambda \ge 5$ all weights are zero; just below $5$ feature 1 (the one most correlated with $\mathbf{b}$) enters.
Recap, cheat sheet and practice
- Coordinate descent minimises along one coordinate at a time, exactly (cyclic or random order). It is great for smooth + separable non-smooth objectives and cheap per step; it zig-zags when variables are strongly linked and can stall when the non-smooth part ties variables together.
- The proximal operator $\operatorname{prox}_{\eta g}(\mathbf{v}) = \operatorname{argmin}_{\mathbf{x}}\, g(\mathbf{x}) + \tfrac{1}{2\eta}\|\mathbf{x} - \mathbf{v}\|^2$ balances a rule against a rubber band to $\mathbf{v}$. For $g = 0$ it is the identity, for a set indicator it is the projection, for $\lambda|x|$ it is soft thresholding $S_{\eta\lambda}$, for $\tfrac12\lambda x^2$ it is $v/(1 + \eta\lambda)$.
- Soft thresholding $S_t(v) = \operatorname{sign}(v)\max(|v| - t, 0)$ comes from the subgradient optimality condition. Its flat dead zone $[-t, t]$ is why L1 solutions contain exact zeros.
- Proximal gradient for $f + g$: $\mathbf{x}^+ = \operatorname{prox}_{\eta g}(\mathbf{x} - \eta\nabla f(\mathbf{x}))$. It is derived by minimising the quadratic upper bound; with $\eta \le 1/L$ it never increases $F$ and has an $O(1/k)$ rate for convex $f$.
- Lasso $\tfrac12\|A\mathbf{x} - \mathbf{b}\|^2 + \lambda\|\mathbf{x}\|_1$: optimal when $s_j = -\lambda\operatorname{sign}(x_j)$ on the active set and $|s_j| \le \lambda$ off it; all zeros for $\lambda \ge \|A^\top\mathbf{b}\|_\infty$.
- ISTA = proximal gradient for the Lasso ($\eta = 1/\sigma_{\max}(A)^2$); FISTA adds momentum for $O(1/k^2)$ (not monotone); coordinate descent uses $x_j \leftarrow S_\lambda(\mathbf{a}_j^\top\mathbf{r} + c_jx_j)/c_j$ with a running residual. The active set shrinks as solvers converge, and warm starts give the whole solution path.
Cheat sheet
| Idea | Formula | Picture |
|---|---|---|
| Coordinate update (quadratic) | $x_i \leftarrow \dfrac{b_i - \sum_{j\ne i}A_{ij}x_j}{A_{ii}}$ | slide to the lowest point of one street |
| Proximal operator | $\operatorname{argmin}_{\mathbf{x}}\, g(\mathbf{x}) + \tfrac{1}{2\eta}\|\mathbf{x}-\mathbf{v}\|^2$ | tug of war: rule vs rubber band |
| Prox of $0$ / set / $\lambda|x|$ / $\tfrac12\lambda x^2$ | $v$ / projection / $S_{\eta\lambda}(v)$ / $\dfrac{v}{1+\eta\lambda}$ | stay / nearest point / shrink and zero / shrink |
| Soft threshold | $\operatorname{sign}(v)\max(|v|-t,0)$, $t = \eta\lambda$ | dead zone then a shifted diagonal |
| Proximal gradient | $\mathbf{x}^+ = \operatorname{prox}_{\eta g}(\mathbf{x} - \eta\nabla f(\mathbf{x}))$, $\eta\le 1/L$ | gradient step, then referee |
| Lasso optimality | $s_j = -\lambda\operatorname{sign}(x_j)$ or $|s_j|\le\lambda$ if $x_j = 0$ | the corner traps weak features |
| $\lambda_{\max}$ | $\|A^\top\mathbf{b}\|_\infty$ | the penalty that kills every weight |
| ISTA / FISTA | $S_{\eta\lambda}(\mathbf{x} - \eta A^\top(A\mathbf{x}-\mathbf{b}))$ / with momentum | $O(1/k)$ / $O(1/k^2)$ |
| CD for the Lasso | $x_j \leftarrow S_\lambda(\mathbf{a}_j^\top\mathbf{r} + c_jx_j)/c_j$ | each weight decides to join or not |
import numpy as np
def soft(v, t): # soft thresholding = prox of t*|.|
return np.sign(v) * np.maximum(np.abs(v) - t, 0.0) + 0.0 # "+ 0.0" turns -0.0 into 0.0
def ista(A, b, lam, iters, fast=False, eta=None):
if eta is None: eta = 1.0 / np.linalg.norm(A, 2) ** 2 # 1/L, L = largest singular value squared
x = np.zeros(A.shape[1]); y, t = x.copy(), 1.0
for _ in range(iters):
z = (y if fast else x) - eta * A.T @ (A @ (y if fast else x) - b) # gradient step
x_new = soft(z, eta * lam) # prox step
if fast:
t_new = (1 + np.sqrt(1 + 4 * t * t)) / 2
y = x_new + (t - 1) / t_new * (x_new - x); t = t_new
x = x_new
return x
def cd(A, b, lam, sweeps):
m, n = A.shape; x = np.zeros(n); r = b - A @ x; c = (A ** 2).sum(axis=0) # c_j = ||a_j||^2
for _ in range(sweeps):
for j in range(n):
u = A[:, j] @ r + c[j] * x[j] # a_j^T (partial residual without feature j)
new = soft(u, lam) / c[j] # exact 1-D minimiser
r -= A[:, j] * (new - x[j]); x[j] = new
return x
def kkt_violation(A, b, lam, x): # 0 means "x is optimal"
g = A.T @ (A @ x - b)
return max(np.max(np.where(x != 0, np.abs(g + lam * np.sign(x)), np.maximum(np.abs(g) - lam, 0))), 0)
# ---- the hand example of this chapter -------------------------------------------------
A = np.array([[1., 0.], [0., 1.], [1., 1.]]); b = np.array([3., 0., 2.]); lam = 2.0
F = lambda x: 0.5 * np.sum((A @ x - b) ** 2) + lam * np.abs(x).sum()
print(soft(np.array([3, -0.5, 1.2, -2, 0.9]), 1.0)) # [ 2. 0. 0.2 -1. 0. ] two entries became exactly 0
print(np.linalg.norm(A, 2) ** 2) # 3.0 (this is L)
for k in (1, 2, 3, 200):
x = ista(A, b, lam, k, eta=0.25); print(k, x, F(x)) # [0.75 0.] 4.8125 | [1.125 0.] 4.3906 | [1.3125 0.] 4.2852 | [1.5 0.] 4.25
print(cd(A, b, lam, 1), kkt_violation(A, b, lam, cd(A, b, lam, 1))) # [1.5 0.] 0.0 (one sweep is enough)
# ---- a bigger random Lasso, three solvers --------------------------------------------
rng = np.random.default_rng(0)
m, n = 100, 20
A = rng.standard_normal((m, n)) + 1.5 * rng.standard_normal((m, 1)) # features share a common part: correlated
A -= A.mean(axis=0); A /= np.sqrt((A ** 2).mean(axis=0)) # centred, columns of squared length m
x_true = np.zeros(n); x_true[[0, 3, 7]] = [2.0, -1.5, 1.0]
b = A @ x_true + 0.5 * rng.standard_normal(m)
lam = 0.2 * np.max(np.abs(A.T @ b)) # 20% of lambda_max
F = lambda x: 0.5 * np.sum((A @ x - b) ** 2) + lam * np.abs(x).sum()
ref = cd(A, b, lam, 2000)
for name, x in [("ISTA 50 iterations", ista(A, b, lam, 50)), ("FISTA 50 iterations", ista(A, b, lam, 50, True)), ("CD 5 sweeps", cd(A, b, lam, 5))]:
print(name, "gap = %.1e" % (F(x) - F(ref)), "nonzeros =", np.count_nonzero(x), "kkt = %.1e" % kkt_violation(A, b, lam, x))
# ISTA 50: gap about 1.0 (still far) FISTA 50: about 4e-3 CD 5 sweeps: about 1e-3 (all three reach 3 non-zeros except ISTA, which has 2)
print("reference nonzeros:", np.flatnonzero(ref), "kkt = %.1e" % kkt_violation(A, b, lam, ref)) # [0 3 7] kkt ~ 1e-14: exactly the true features
# scikit-learn gives the same weights: Lasso(alpha=lam/m, fit_intercept=False, tol=1e-12).fit(A, b).coef_
1. What is $\operatorname{prox}_{\eta g}(v)$ for $g(x) = \lambda|x|$ with $\lambda = 2$, $\eta = 0.5$ and $v = -2.5$?
2. Which objective can coordinate descent get stuck on, at a point that is not a minimum?
3. In proximal gradient, what happens if $g$ is the indicator function of a convex set $C$ (zero inside, infinite outside)?
4. Why do Lasso solutions contain exactly zero weights?
5. A safe step-size for ISTA on $\tfrac12\|A\mathbf{x} - \mathbf{b}\|^2 + \lambda\|\mathbf{x}\|_1$ is
6. Which statement about FISTA versus ISTA is correct?
Practice problems
A. Apply soft thresholding $S_{1.5}$ to $[4,\ -0.2,\ -3,\ 1]$.
$4 \to 2.5$. $-0.2$ has size $\le 1.5$, so $0$. $-3 \to -1.5$. $1$ has size $\le 1.5$, so $0$. Result: $[2.5,\ 0,\ -1.5,\ 0]$.
B. Do two coordinate-descent sweeps for the Lasso with $A = \begin{bmatrix}1&0\\0&1\\1&1\end{bmatrix}$, $\mathbf{b} = [1, 1, 3]$, $\lambda = 1$, starting from $\mathbf{x} = 0$. The exact answer is $(1, 1)$.
Here $c_1 = c_2 = 2$, $\mathbf{r} = \mathbf{b} = (1, 1, 3)$. Sweep 1. $u_1 = 1 + 0 + 3 = 4$, $x_1 = S_1(4)/2 = 1.5$, $\mathbf{r} = (1,1,3) - 1.5\,(1,0,1) = (-0.5, 1, 1.5)$. $u_2 = 1 + 1.5 = 2.5$, $x_2 = S_1(2.5)/2 = 0.75$, $\mathbf{r} = (-0.5, 1, 1.5) - 0.75\,(0,1,1) = (-0.5, 0.25, 0.75)$. Sweep 2. $u_1 = (-0.5 + 0.75) + 2\cdot1.5 = 3.25$, $x_1 = 2.25/2 = 1.125$, $\mathbf{r} = (-0.125, 0.25, 1.125)$. $u_2 = (0.25 + 1.125) + 2\cdot0.75 = 2.875$, $x_2 = 1.875/2 = 0.9375$. So after two sweeps $(1.125,\ 0.9375)$, creeping towards $(1, 1)$ with the distance falling by a factor of 4 per sweep. Optimality check at $(1,1)$: $A^\top A\mathbf{x} - A^\top\mathbf{b} = (3, 3) - (4, 4) = (-1, -1) = -\lambda\,\operatorname{sign}(\mathbf{x})$ ✓.
C. $A = \begin{bmatrix}1&0\\0&2\end{bmatrix}$, $\mathbf{b} = [2, 1]$, $\lambda = 1$. Find $\lambda_{\max}$ and the Lasso solution, and verify the optimality conditions.
$A^\top\mathbf{b} = (2, 2)$, so $\lambda_{\max} = 2$. The columns are orthogonal ($A^\top A = \operatorname{diag}(1, 4)$), so the weights separate: $x_1 = S_1(2)/1 = 1$ and $x_2 = S_1(2)/4 = 0.25$. Gradient of the smooth part: $A^\top A\mathbf{x} - A^\top\mathbf{b} = (1 - 2,\ 4\cdot0.25 - 2) = (-1, -1)$. Both weights are positive, and $s_j = -1 = -\lambda$ ✓.
D. Take one proximal gradient step for $f(x) = \tfrac12(x - 5)^2$, $g(x) = 2|x|$ from $x = 0$, with $\eta = 0.5$. Then say where the iteration is heading.
Gradient step: $z = 0 - 0.5\,(0 - 5) = 2.5$. Prox with threshold $\eta\lambda = 0.5\cdot2 = 1$: $S_1(2.5) = 1.5$. The fixed point solves $(x - 5) + 2 = 0$, i.e. $x^\star = 3$. (Each iteration halves the gap: $0 \to 1.5 \to 2.25 \to \dots \to 3$.)
E. Why can proximal gradient with $\eta \le 1/L$ never increase $F = f + g$?
The upper bound gives $F(\mathbf{y}) \le Q_{\mathbf{x}}(\mathbf{y})$ for all $\mathbf{y}$, with equality at $\mathbf{y} = \mathbf{x}$. The new point $\mathbf{x}^+$ is the minimiser of $Q_{\mathbf{x}}$, so $Q_{\mathbf{x}}(\mathbf{x}^+) \le Q_{\mathbf{x}}(\mathbf{x}) = F(\mathbf{x})$. Chaining: $F(\mathbf{x}^+) \le Q_{\mathbf{x}}(\mathbf{x}^+) \le F(\mathbf{x})$.
F. Cyclic coordinate descent on $f = \tfrac12(x_1^2 + 2\rho x_1x_2 + x_2^2)$ shrinks the distance to the minimum by $\rho^2$ per sweep (after the first sweep). For $\rho = 0.95$, about how many sweeps shrink the distance by a factor of 1000?
$\rho^2 = 0.9025$. We need $0.9025^k \le 10^{-3}$, so $k \ge \ln(1000)/\ln(1/0.9025) = 6.908/0.1026 \approx 67$ sweeps. Compare $\rho = 0.5$ ($\rho^2 = 0.25$): $\ln(1000)/\ln 4 \approx 5$ sweeps. Strong coupling is expensive.
Stochastic Optimization
Real training sets are too big to look at in full for every step. So we steer by a random sample: a gradient that is noisy, but right on average. This chapter explains exactly how noisy, what the noise costs you, what it does to the end of training, and the practical tricks (bigger batches, smaller steps, averaging) that tame it.
- Write a training loss as an average over examples and see why one example (or a mini-batch) gives a noisy but unbiased estimate of the gradient, with a proof
- Understand gradient noise: it has mean zero, but it does not vanish at the optimum when the step-size is constant
- Derive the variance of a mini-batch gradient, $\sigma^2/B$, and see the diminishing returns of bigger batches
- State what SGD converges to (a noise ball for a constant step, the optimum for decaying steps), with the Robbins–Monro conditions and the rates, each labelled with its assumptions
- Reason about batch size and learning rate (the linear scaling heuristic and its limits), and meet learning-rate decay, iterate averaging and variance reduction (SVRG, SAGA) as the practical fixes for noise
Stochastic gradients: steering by a random sample core
A training loss is an average over all examples: the average error on a million photos. To know exactly which way is downhill, you would have to ask every photo. That is a full pass over the data, just to take one step.
Think of an election poll. You do not ask every voter; you ask a random few hundred. Any single poll is a bit off, but it is unbiased: it is not tilted towards either side, so the average of many polls is the true result. A stochastic gradient is the same trick: ask one random example (or a small random group, a mini-batch) which way is downhill. The answer is noisy, but on average it is exactly right, and it costs a tiny fraction of a full pass.
Four training examples, each with its own loss $f_i(w) = \tfrac12(w - a_i)^2$, where $a = (1, 3, 5, 7)$ and there is a single weight $w$. The training loss is the average $f(w) = \tfrac14\sum_i f_i(w)$. Stand at $w = 2$.
- Per-example gradients. $f_i'(w) = w - a_i$, so at $w = 2$ they are $(1,\ -1,\ -3,\ -5)$.
- Full gradient (the average): $\dfrac{1 - 1 - 3 - 5}{4} = -2$. Check: $f(w) = \tfrac12(w - 4)^2 + \text{const}$, so $f'(2) = 2 - 4 = -2$ ✓. The gradient is negative, so the loss falls when we move right, towards the minimum at $w = 4$.
- One random example. We get $1,\ -1,\ -3$ or $-5$ with probability $\tfrac14$ each. The value $+1$ even points the wrong way, $-1$ is much too small, and $-5$ is far too big. But their average is $(1 - 1 - 3 - 5)/4 = -2$: right on average.
- A mini-batch of 2 (there are 6 pairs). The pair averages are $0,\ -1,\ -2,\ -2,\ -3,\ -4$. Their average is $(0 - 1 - 2 - 2 - 3 - 4)/6 = -2$ again, and the values are less spread out than for single examples.
Many ML losses are a finite sum (an average over $n$ examples), where $f_i$ is the loss on example $i$ and $\mathbf{w}$ are the model parameters (in this chapter we write $\mathbf{w}$ for the variables, as is usual for model weights; see the calculus-to-ML chapter): $$f(\mathbf{w}) = \frac1n\sum_{i=1}^n f_i(\mathbf{w}),\qquad \nabla f(\mathbf{w}) = \frac1n\sum_{i=1}^n \nabla f_i(\mathbf{w}).$$ A stochastic gradient is $\mathbf{g}(\mathbf{w}) = \nabla f_I(\mathbf{w})$ for an index $I$ chosen uniformly at random from $\{1, \dots, n\}$. A mini-batch gradient with batch size $B$ averages $B$ of them: $$\mathbf{g}_B(\mathbf{w}) = \frac1B\sum_{i \in S}\nabla f_i(\mathbf{w}),\qquad S = \text{a random set of } B \text{ indices}.$$ Claim (unbiasedness). For any fixed $\mathbf{w}$, $\ \mathbb{E}[\mathbf{g}_B(\mathbf{w})] = \nabla f(\mathbf{w})$.
Proof. One example: $\mathbb{E}[\nabla f_I(\mathbf{w})] = \sum_{i=1}^n \Pr(I = i)\,\nabla f_i(\mathbf{w}) = \frac1n\sum_i\nabla f_i(\mathbf{w}) = \nabla f(\mathbf{w})$. Batch with replacement ($B$ independent draws $I_1, \dots, I_B$): by linearity of expectation, $\mathbb{E}[\mathbf{g}_B] = \frac1B\sum_{b=1}^B\mathbb{E}[\nabla f_{I_b}] = \frac1B\cdot B\,\nabla f = \nabla f$. Batch without replacement (a random subset): each example $i$ lies in the batch with probability $B/n$, so $\mathbb{E}[\mathbf{g}_B] = \frac1B\sum_i\frac Bn\nabla f_i = \nabla f$. ∎
The SGD step then looks just like gradient descent with the estimate in place of the true gradient: $\ \mathbf{w}_{k+1} = \mathbf{w}_k - \eta\,\mathbf{g}_B(\mathbf{w}_k)$, with learning rate $\eta$ (full study in the SGD section below; plain and mini-batch SGD were introduced in Chapter 3.3).
Why do we need it?
A full gradient costs one pass over all $n$ examples, which for millions of examples is far too slow for a single step. A random sample costs $B \ll n$ evaluations and still points the right way on average, so you can take thousands of cheap steps in the time of one expensive one.
Where is it used?
Training essentially every neural network (SGD, Adam and other optimizers all feed on mini-batch gradients), logistic regression on big data, recommender systems, and any model fitted to a data stream.
How is it used?
Each step: draw a random mini-batch, compute the loss gradient on just those examples, and update the weights with it. In code that is one batch from the data loader, loss.backward(), and optimizer.step().
Unbiased does not mean accurate. A single stochastic gradient can be far from the truth, even pointing the wrong way, as in the 4-example calculation. Unbiased only says there is no systematic tilt. How large the random error is comes next.
The statement is for a fixed point $\mathbf{w}$. After the first step the iterate $\mathbf{w}_k$ is itself random, so what happens over a whole run needs more care (see the SGD section).
Quick check: at the current weights three examples have gradients $2$, $4$ and $9$. What is the full gradient, and what is the average of the three possible one-example estimates?
The full gradient is $(2 + 4 + 9)/3 = 5$. The three one-example estimates are $2$, $4$, $9$ with probability $1/3$ each, so their average is also $5$: the estimator is unbiased, although each single estimate is off (by $-3$, $-1$, $+4$).
Gradient noise: it does not go away at the optimum core
Split the stochastic gradient into two parts: the signal (the true gradient, the same for everyone) and the noise (what is left: how this particular sample differs from the average). The noise has zero average, but it has a size, and the size matters.
Far from the optimum the signal is large (it grows with the distance), and for a reasonable batch it outweighs the noise. At the optimum the signal is exactly zero, but the individual examples still disagree: one wants the weight a bit higher, another a bit lower, and their wishes only cancel on average. Pick one of them and you get a push that is not zero. So with a constant step-size SGD keeps being shoved around the optimum and never quite settles. Picture a marble in a bowl on a table that is being gently, randomly shaken.
Same four examples as before ($f_i = \tfrac12(w - a_i)^2$, $a = (1, 3, 5, 7)$). The minimum of the average loss is at $w^\star = 4$. Per-example gradients there: $w^\star - a_i = (3,\ 1,\ -1,\ -3)$.
- Their average is $(3 + 1 - 1 - 3)/4 = 0$: the true gradient vanishes at the optimum ✓.
- But each single example gives a non-zero gradient. Take an SGD step with $\eta = 0.1$: the weight goes to $4 - 0.1\cdot(3, 1, -1, -3) = 3.7,\ 3.9,\ 4.1$ or $4.3$. We left the optimum, whichever example we drew.
- The noise power is $\sigma^2 = \tfrac14(9 + 1 + 1 + 9) = 5$.
The gradient noise at $\mathbf{w}$ is $\boldsymbol{\xi}(\mathbf{w}) = \mathbf{g}(\mathbf{w}) - \nabla f(\mathbf{w})$, so that $\mathbf{g} = \nabla f + \boldsymbol{\xi}$. By unbiasedness $\mathbb{E}[\boldsymbol{\xi}] = \mathbf{0}$. Its size is measured by the noise power (the total variance) $$\sigma^2(\mathbf{w}) = \mathbb{E}\|\boldsymbol{\xi}\|^2 = \frac1n\sum_{i=1}^n\|\nabla f_i(\mathbf{w})\|^2 - \|\nabla f(\mathbf{w})\|^2.$$ (Derivation: $\mathbb{E}\|\mathbf{g} - \nabla f\|^2 = \mathbb{E}\|\mathbf{g}\|^2 - 2\,\nabla f^\top\mathbb{E}[\mathbf{g}] + \|\nabla f\|^2 = \mathbb{E}\|\mathbf{g}\|^2 - \|\nabla f\|^2$, and $\mathbb{E}\|\mathbf{g}\|^2 = \frac1n\sum_i\|\nabla f_i\|^2$.)
- At the optimum $\nabla f(\mathbf{w}^\star) = \mathbf{0}$, so $\sigma^2(\mathbf{w}^\star) = \frac1n\sum_i\|\nabla f_i(\mathbf{w}^\star)\|^2$. This is zero only if every single example has zero gradient at $\mathbf{w}^\star$. In the example it is $5 \gt 0$.
- Consequence for a constant step. The update $\mathbf{w}_{k+1} = \mathbf{w}_k - \eta\nabla f(\mathbf{w}_k) - \eta\boldsymbol{\xi}_k$ contains the kick $-\eta\boldsymbol{\xi}_k$ forever. Near $\mathbf{w}^\star$ the first term fades and the kick does not, so the iterates keep jittering around $\mathbf{w}^\star$.
- The exception: interpolation. If the model can fit every training example exactly (for example noise-free labels, or a very over-parameterised network), then all $\nabla f_i(\mathbf{w}^\star) = \mathbf{0}$ and $\sigma^2(\mathbf{w}^\star) = 0$. The noise then shrinks together with the gradient, and constant-step SGD can converge cleanly. (Whether a real deep-learning run is close to this regime depends on the model, the data and augmentation; treat it as a useful idealisation.)
- The noise power depends on the point $\mathbf{w}$; it is not a fixed number.
Why do we need it?
It explains why SGD loss curves look jagged, why the training loss stops falling at a plateau with a constant learning rate, and why lowering the learning rate late in training makes the loss suddenly drop.
Where is it used?
The signal-to-noise ratio of the gradient guides the choice of batch size (the "gradient noise scale" idea), learning-rate schedules, and the design of variance-reduced methods (SVRG, SAGA) later in this chapter.
How is it used?
Compare the size of the gradient ($\|\nabla f\|$) with the noise ($\sigma$): when the gradient is much bigger, noise is harmless; when it is smaller, you are noise-dominated and should use a bigger batch, a smaller step, or averaging.
"The gradient is zero" is not what you see in training. Even at a perfect optimum of the training loss, the mini-batch gradient norm you log is not zero: it hovers at the noise level $\sigma/\sqrt B$. So "stop when the gradient is tiny" does not work for SGD (stopping rules are discussed in Chapter 3.5).
The noise is not always Gaussian or constant. It depends on $\mathbf{w}$, on the loss, and on the data. We use only its mean (zero) and power $\sigma^2$.
Quick check: three examples have gradients $-2,\ 0,\ 2$ at the optimum. Is the true gradient zero? What is $\sigma^2$?
The average is $(-2 + 0 + 2)/3 = 0$, so yes, the true gradient is zero. The noise power is the average of the squares: $(4 + 0 + 4)/3 = 8/3 \approx 2.67$, not zero. If all three had been $0$ (interpolation), $\sigma^2$ would be $0$ too.
Mini-batches, sampling and epochs core
You can ask one person (noisy, very cheap), everyone (exact, very expensive), or a small group in between. A mini-batch is that small group. It is the compromise that nearly all training uses: bigger than one (so the noise is smaller and the hardware stays busy), much smaller than the full data set (so each step is cheap).
Counting trick: if a full gradient step costs as much as $n$ example-gradients, then in the same cost you can take $n/B$ mini-batch steps of size $B$. One pass over all the data is called an epoch.
A data set of $n = 1{,}000{,}000$ examples and a batch size of $B = 32$.
- One full-gradient step costs $1{,}000{,}000$ example-gradients.
- One mini-batch step costs $32$. So the same cost buys $1{,}000{,}000/32 = 31{,}250$ steps. That is exactly one epoch of mini-batch SGD.
- After that epoch the weights have moved $31{,}250$ times (each step slightly noisy) rather than once (exact).
Sampling without replacement, tiny case. $n = 6$ examples, $B = 3$. Shuffle the order once, say $(4, 1, 6, 2, 5, 3)$, and cut it into batches: $\{4, 1, 6\}$ then $\{2, 5, 3\}$. After $n/B = 2$ steps every example was used exactly once. If we had drawn each batch with replacement, some examples could appear twice and some not at all.
Batch size $B$ and epoch: one epoch is $n/B$ steps, i.e. one pass worth of example-gradients ($n$ in total). The cost per step is proportional to $B$ example-gradients (on one processor; on a GPU it is nearly flat for small $B$, see the batch-size section).
Three ways to pick the batch:
- With replacement: draw $B$ indices independently, each uniform on $\{1, \dots, n\}$. Easiest to analyse. The variance of the batch gradient is exactly $\sigma^2/B$.
- Without replacement: draw a random subset of $B$ different examples. The variance is a bit smaller: $\ \dfrac{n - B}{n - 1}\cdot\dfrac{\sigma^2}{B}$. (Check on the 4-example problem with $B = 2$: factor $\tfrac{4-2}{4-1} = \tfrac23$; the six pair-averages $0, -1, -2, -2, -3, -4$ have variance $10/6 = 1.67 = \tfrac23 \cdot \tfrac52$ ✓.) For $B = n$ it is $0$: the full gradient has no noise.
- Random reshuffling (what data loaders do): shuffle the data at the start of every epoch and walk through it in batches, so each example is used exactly once per epoch. At a fixed $\mathbf{w}$ the $n/B$ batch gradients of one epoch average to exactly $\nabla f$. (In practice this is often slightly better than sampling with replacement; the theory of why is more delicate than for with-replacement sampling.)
The standard analysis (and the formulas in the next sections) is for sampling with replacement; practice uses reshuffling and behaves similarly.
Why do we need it?
One example per step is too noisy and uses modern hardware badly (a GPU processes hundreds of examples at the price of a few). The full data set per step is too slow. A mini-batch balances noise, speed and hardware use.
Where is it used?
The batch_size argument of every deep-learning data loader (PyTorch DataLoader, TensorFlow tf.data), mini-batch k-means, and mini-batch versions of logistic regression and matrix factorisation.
How is it used?
Shuffle the data every epoch, cut it into batches of size $B$ (typical values are 32 to a few thousand), take one optimizer step per batch, and count training progress in epochs or in steps, not in "iterations over all data".
Count progress in gradient evaluations (or epochs), not steps. A step with $B = 1024$ is not "the same" as a step with $B = 8$. Compare methods at equal cost.
Shuffle every epoch. If the data are sorted (for example all cats, then all dogs), unshuffled batches are not a random sample of the data, so the gradient estimates are biased towards whatever the batch contains.
The last batch may be smaller when $B$ does not divide $n$; data loaders either keep it or drop it (drop_last).
Quick check: $n = 50{,}000$ images and $B = 100$. How many steps are in one epoch, and how many example-gradients does an epoch cost?
$n/B = 500$ steps. Each step costs $100$ example-gradients, so an epoch costs $500\cdot100 = 50{,}000 = n$ example-gradients: exactly the cost of one full-gradient step, but you have made 500 updates.
Variance of the mini-batch gradient: $\sigma^2/B$ core
Measure a table with a wobbly ruler. One measurement may be off by a centimetre either way. Average ten measurements and the errors partly cancel (one too long, another too short), so the average is more accurate. But it is not ten times more accurate. The error shrinks like $1/\sqrt{10} \approx 0.32$, because cancelling is only partial.
A mini-batch gradient is exactly such an average of $B$ noisy single-example gradients. Doubling the batch size does not halve the noise; you must quadruple the batch to halve it. Bigger batches always help, but with diminishing returns.
Back to the four-example problem, where the noise power of a single example is $\sigma^2 = 5$ (so the typical error is $\sigma = \sqrt5 \approx 2.24$). With batches drawn with replacement:
| batch size $B$ | variance $\sigma^2/B$ | typical error $\sigma/\sqrt B$ | cost (example-gradients) |
|---|---|---|---|
| 1 | $5$ | $2.24$ | 1 |
| 4 | $1.25$ | $1.12$ (half) | 4 |
| 16 | $0.3125$ | $0.56$ (quarter) | 16 |
| 64 | $0.078$ | $0.28$ (an eighth) | 64 |
Each halving of the error costs four times as much work.
Variance of a mini-batch gradient (derivation). Draw $B$ examples independently (with replacement). Write the $b$-th single-example gradient as $\mathbf{g}_b = \nabla f(\mathbf{w}) + \boldsymbol{\xi}_b$, where the noise terms $\boldsymbol{\xi}_1, \dots, \boldsymbol{\xi}_B$ are independent, have mean $\mathbf{0}$ and $\mathbb{E}\|\boldsymbol{\xi}_b\|^2 = \sigma^2$. The batch gradient is $$\mathbf{g}_B = \frac1B\sum_{b=1}^B\mathbf{g}_b = \nabla f(\mathbf{w}) + \frac1B\sum_{b=1}^B\boldsymbol{\xi}_b .$$ Its squared error is $$\mathbb{E}\|\mathbf{g}_B - \nabla f\|^2 = \frac{1}{B^2}\,\mathbb{E}\Big\|\sum_b\boldsymbol{\xi}_b\Big\|^2 = \frac{1}{B^2}\Big(\sum_b\mathbb{E}\|\boldsymbol{\xi}_b\|^2 + \sum_{b \ne c}\mathbb{E}[\boldsymbol{\xi}_b^\top\boldsymbol{\xi}_c]\Big) = \frac{1}{B^2}\big(B\sigma^2 + 0\big) = \boxed{\frac{\sigma^2}{B}}.$$ The cross terms vanish because $\boldsymbol{\xi}_b$ and $\boldsymbol{\xi}_c$ are independent with mean zero: $\mathbb{E}[\boldsymbol{\xi}_b^\top\boldsymbol{\xi}_c] = \mathbb{E}[\boldsymbol{\xi}_b]^\top\mathbb{E}[\boldsymbol{\xi}_c] = 0$. ∎
- Typical error $= \sigma/\sqrt B$. To halve it, multiply $B$ by $4$. This is the diminishing return: the cost grows like $B$, the benefit like $\sqrt B$.
- Signal and noise. Since the cross term between $\nabla f$ and the noise vanishes, $\ \mathbb{E}\|\mathbf{g}_B\|^2 = \|\nabla f\|^2 + \sigma^2/B$: "signal squared plus noise squared". The two are equal when $B = \sigma^2/\|\nabla f\|^2$. This number is called the gradient noise scale; as a rule of thumb (an empirical heuristic, not a theorem) batches much smaller than it are noise-dominated and batches much larger give little extra.
- Without replacement the variance is $\dfrac{n-B}{n-1}\cdot\dfrac{\sigma^2}{B}$, smaller by the factor $\frac{n-B}{n-1}$ (see the previous section), which is close to $1$ when $B \ll n$.
Why do we need it?
It tells you what a larger batch actually buys: noise falling like $1/\sqrt B$. This is how you decide whether doubling your batch is worth doubling the cost, and why scaling a run to huge batches stops paying off.
Where is it used?
Choosing the batch size for neural-network training, designing "large batch" training on many GPUs, estimating how noisy a training run is, and variance-reduction theory (SVRG, SAGA) later in this chapter.
How is it used?
Estimate $\sigma^2$ by computing a few single-example gradients (or the spread of a few batch gradients) and compare $\sigma^2/B$ with $\|\nabla f\|^2$. If noise is much bigger than signal, a larger batch helps. If it is already much smaller, spend your compute on more steps instead.
The formula assumes independent draws. If the examples in a batch are strongly related (for example consecutive frames of a video, or sorted data), the variance is larger than $\sigma^2/B$. Shuffle.
Variance is not the same as bias. A bigger batch removes variance, not bias. The mini-batch gradient was already unbiased.
Less noise per step is not the whole story. One step with batch $2B$ costs the same as two steps with batch $B$; which is better depends on the learning rate (the linear scaling heuristic below).
Quick check: a batch of 25 examples has typical gradient error $0.4$. About what error would you expect with a batch of 100, and with a batch of 400?
The error scales like $1/\sqrt B$. $B = 100$ is 4 times larger, so the error halves to about $0.2$. $B = 400$ is 16 times larger, so it is $4$ times smaller: about $0.1$. Each halving costs 4 times more examples.
SGD: the update and the noise ball core
Walk down a foggy hillside with a noisy compass. Each step points roughly downhill with a random wobble. High on the slope, the true slope is steep compared with the wobble, so you make steady progress. Near the valley floor the true slope is tiny, the wobble dominates, and you jiggle around the bottom in a small noise ball instead of stopping.
Two remedies: smaller steps (a smaller ball, but you travel more slowly), or a step-size that shrinks over time (the ball itself shrinks to a point). Both are developed in this and the next sections.
One weight, loss $f(w) = \tfrac12w^2$ (minimum at $0$, curvature $1$), and a gradient with random noise of power $\sigma^2 = 1$: SGD step $w_{k+1} = w_k - \eta(w_k + \xi_k)$. Take $\eta = 0.2$.
- Rewrite: $w_{k+1} = 0.8\,w_k - 0.2\,\xi_k$. The first part shrinks the weight by $0.8$ each step. The second part adds a fresh random kick of size $0.2\,\xi_k$.
- Square and average (the kick is independent of $w_k$ and has mean zero, so the cross term vanishes): $\ \mathbb{E}[w_{k+1}^2] = 0.64\,\mathbb{E}[w_k^2] + 0.04$.
- This settles where $V = 0.64V + 0.04$, i.e. $V = 0.04/0.36 = 0.111$. So the weight jitters with typical distance $\sqrt{0.111} = 0.33$ from the optimum, forever, and the loss gap settles at $\tfrac12V = 0.056$.
- Halve the step to $\eta = 0.1$: $V = 0.01/(1-0.81) = 0.0526$, about half as big. The squared distance $V$ is proportional to $\eta$ (for small $\eta$), so the radius $\sqrt V$ shrinks only by a factor $\sqrt2$.
SGD (mini-batch). $\ \mathbf{w}_{k+1} = \mathbf{w}_k - \eta_k\,\mathbf{g}_B(\mathbf{w}_k)$, where $\mathbf{g}_B$ is an unbiased mini-batch gradient with noise power $\sigma^2/B$ and $\eta_k$ is the learning rate at step $k$ (constant or decaying).
The noise ball, derived on a 1-D quadratic. Take $f(w) = \tfrac12aw^2$ and per-step gradient $aw_k + \xi_k$ with $\mathbb{E}\xi_k = 0$, $\mathbb{E}\xi_k^2 = s^2$ (here $s^2 = \sigma^2/B$ for batch size $B$). Then $w_{k+1} = (1 - \eta a)w_k - \eta\xi_k$ and, as in the example, $$\mathbb{E}[w_{k+1}^2] = (1 - \eta a)^2\,\mathbb{E}[w_k^2] + \eta^2s^2 .$$ For $0 \lt \eta a \lt 2$ this settles at the fixed point $V_\infty = \dfrac{\eta^2s^2}{1 - (1 - \eta a)^2} = \dfrac{\eta\,s^2}{a\,(2 - \eta a)} \approx \dfrac{\eta s^2}{2a}$. Hence the expected loss gap settles at $\tfrac a2V_\infty = \dfrac{\eta s^2}{2(2 - \eta a)} \approx \dfrac{\eta s^2}{4}$. (Check with the example: $a = 1$, $s^2 = 1$, $\eta = 0.2$: $0.2/(1.8) = 0.111$ ✓.) So both the squared distance and the loss gap of the ball are proportional to $\eta\,s^2$: the radius of the ball is proportional to $\sqrt{\eta\sigma^2/B}$. The ball shrinks with a smaller step, a smaller noise power, or a bigger batch.
General statement (assumptions: $f$ is $\mu$-strongly convex and $L$-smooth; the gradient estimate is unbiased with noise power $\le\sigma^2$; constant $\eta \le 1/L$; Bottou, Curtis and Nocedal, 2018): $$\mathbb{E}[f(\mathbf{w}_k)] - f^\star \;\le\; (1 - \eta\mu)^k\,\big(f(\mathbf{w}_0) - f^\star\big) \;+\; \frac{\eta L\sigma^2}{2\mu}.$$ The first term is gradient descent's linear convergence and fades. The second is the noise floor, proportional to $\eta\sigma^2$, that remains. (In the 1-D check it reads $\eta s^2/2$ against the exact $\eta s^2/(2(2-\eta))$: the bound holds.)
Why do we need it?
It tells you exactly what to expect from a constant learning rate: fast progress at first, then a plateau at a level set by $\eta\sigma^2/B$. Without this picture, a flat loss curve looks like "training is stuck" when it is just the noise floor.
Where is it used?
Every training run with a constant learning rate, the reason learning-rate schedules and larger batches lower the final loss, and the theory behind choosing $\eta$ and $B$ together in large-scale training.
How is it used?
If the training loss plateaus and is noisy, you are at the noise floor: lower $\eta$, increase $B$, decay the learning rate, or average the weights (next sections). If it falls steadily, the gradient signal still dominates and you can keep the step as it is.
A constant step does not "converge to the optimum". It converges (linearly, at first) to the noise ball, and then wanders inside it. Do not wait for the loss to reach zero.
The ball size is $\propto\eta\sigma^2/B$ in the loss gap and in the squared distance. Halving $\eta$ halves the floor; the radius shrinks only by $\sqrt2$.
This analysis is for convex problems. For neural networks the loss is not convex. The noise-ball picture still describes what happens near a minimum (where the loss is close to a bowl), but the global story is more complicated (you will meet it in Chapter 3.15).
Quick check: in the 1-D example with $a = 1$, $\sigma^2 = 1$ and $\eta = 0.5$, what is the stationary $\mathbb{E}[w^2]$, and what is it for batch size $B = 4$ (so the noise power is $1/4$)?
$V_\infty = \eta s^2/(a(2 - \eta a))$. For $s^2 = 1$: $0.5/(1.5) = 0.333$. For $B = 4$, $s^2 = 0.25$ and $V_\infty = 0.5\cdot0.25/1.5 = 0.0833$: four times smaller (a 4 times bigger batch gives 4 times less variance, hence a 2 times smaller radius).
How fast is SGD? Rates, and a race against gradient descent
Gradient descent is a few careful, expensive steps: each one reads the whole data set. SGD is many sloppy, cheap steps. Far from the answer, being sloppy does not matter: any rough "downhill" is good enough, and cheap steps win by a huge margin. Close to the answer you need precision: gradient descent keeps improving, while SGD is limited by its noise ball (unless you shrink the step).
So the race has a typical shape: SGD sprints ahead in the first epoch or two, then flattens out; gradient descent starts slowly but keeps going.
Cost to reach a given accuracy (order-of-magnitude, constants and logarithmic factors dropped). Suppose $f$ is $\mu$-strongly convex and $L$-smooth with $\kappa = L/\mu$.
- Gradient descent needs about $\kappa\ln(1/\varepsilon)$ steps to reach a loss gap $\varepsilon$, and each step reads all $n$ examples: total cost $\approx n\,\kappa\ln(1/\varepsilon)$ example-gradients.
- SGD with a decaying step needs about $1/\varepsilon$ steps (times a constant depending on $\sigma^2$, $L$, $\mu$), each reading $B$ examples. The cost does not contain $n$.
For $n = 10^6$ and a modest accuracy such as $\varepsilon \approx 10^{-2}$ (relative to the noise level), SGD is far cheaper. For very small $\varepsilon$ the $1/\varepsilon$ catches up, which is why careful methods (or variance reduction, below) win when high accuracy is needed and $n$ is small.
Rates of SGD (the assumptions matter; each line says what it needs). Throughout: $f$ is $L$-smooth and the gradient estimate is unbiased with noise power at most $\sigma^2$.
| Setting | Step-size | Result |
|---|---|---|
| $\mu$-strongly convex | constant, $\eta \le 1/L$ | $\mathbb{E}f(\mathbf{w}_k) - f^\star \le (1-\eta\mu)^k(f(\mathbf{w}_0) - f^\star) + \dfrac{\eta L\sigma^2}{2\mu}$: linear convergence to a noise floor $\propto\eta\sigma^2$ |
| $\mu$-strongly convex | decaying, $\eta_k = \dfrac{\beta}{\gamma + k}$ with $\beta \gt 1/\mu$ and $\gamma$ large enough that $\eta_0 \le 1/L$ | $\mathbb{E}f(\mathbf{w}_k) - f^\star = O(1/k)$ |
| convex (not strongly) | $\eta \propto 1/\sqrt K$ for a run of $K$ steps; use the averaged iterate $\bar{\mathbf{w}}_K$ | $\mathbb{E}f(\bar{\mathbf{w}}_K) - f^\star = O(1/\sqrt K)$ |
| non-convex | $\eta \propto 1/\sqrt K$ | only a small gradient: $\min_{k \le K}\mathbb{E}\|\nabla f(\mathbf{w}_k)\|^2 = O(1/\sqrt K)$ (a stationary point, not a minimum) |
Under these assumptions, known lower bounds say that these rates (constants omitted) are essentially the best any method can do when it only sees noisy gradients. So a stochastic method cannot match the linear rate of gradient descent per step; it wins on the cost per step. Gradient descent for comparison, with $\eta = 1/L$: $f(\mathbf{w}_k) - f^\star \le (1 - \mu/L)^k(f(\mathbf{w}_0) - f^\star)$ if strongly convex, $O(1/k)$ if only convex. The first row of the table is the one we derived by hand in the 1-D case in the previous section.
Epoch accounting. To compare fairly measure cost in epochs (passes worth of example-gradients). Gradient descent makes 1 step per epoch; SGD with batch size $B$ makes $n/B$ steps per epoch.
Why do we need it?
"Is SGD better than gradient descent?" has no one answer. The rates tell you when: SGD for big data and moderate accuracy, gradient descent (or quasi-Newton, Chapter 3.12) for small data and high accuracy.
Where is it used?
Deciding between SGD and L-BFGS for a logistic regression (SGD for millions of rows, L-BFGS for thousands), explaining training curves that fall fast and then flatten, and picking schedules such as $1/k$ or $1/\sqrt k$ decay.
How is it used?
Always plot loss against epochs (or example-gradients), not steps. Expect SGD to lead early. If you need high precision at the end, decay the learning rate, or switch to a method that removes the noise.
The race depends on the problem. This is one seeded, well-conditioned two-weight regression. SGD's early lead is typical for large data, but the cross-over epoch depends on the data size, the conditioning, and the noise level. It is a demonstration, not a theorem.
Rates are worst-case upper bounds under assumptions. A neural network is not convex, so the first three rows do not apply to it as stated. They describe what happens near a good minimum, and they explain the roles of $\eta$, $B$ and $\sigma^2$.
Plot against epochs. The same SGD run looks good or bad depending on whether the horizontal axis counts steps or passes over the data.
Quick check: why can SGD with a constant step never match gradient descent's final accuracy on a strongly convex problem?
Its bound has a floor $\eta L\sigma^2/(2\mu)$ that does not shrink with $k$: the gradient noise is not zero at the optimum, so the iterates keep jittering inside a ball. Gradient descent has no such floor. SGD can reach the same accuracy only by decaying $\eta$ (or by reducing the noise, e.g. a bigger batch).
Batch-size effects, and the linear scaling heuristic core
The batch size is a dial between many cheap noisy steps (small $B$) and few expensive clean steps (large $B$). Two things change as you turn it:
- How many steps per second. Modern hardware (a GPU) can often handle a batch of 256 almost as fast as a batch of 8, because it does many examples in parallel. Until the hardware is full, a bigger batch is nearly free. After that, time grows in proportion to $B$.
- How good each step is. Noise falls like $1/\sqrt B$. Once the noise is already smaller than the signal, a bigger batch cannot make steps much better: it behaves like gradient descent.
When you make the batch bigger you take fewer steps per epoch, so you must make each step longer to make the same progress. That is the idea behind the linear scaling rule.
Currently $B = 32$ and $\eta = 0.1$. You want to use $B' = 256$, which is $k = 8$ times larger. Compare 8 small steps (batch 32, step $0.1$) with 1 big step (batch 256, step $0.8 = 8\cdot0.1$), assuming the true gradient $\nabla f$ stays about the same over those 8 steps.
- Average movement. 8 small steps move $8\cdot0.1\cdot\nabla f = 0.8\,\nabla f$. The big step moves $0.8\,\nabla f$. Same ✓.
- Noise. Each small step adds a random kick with variance $\eta^2\sigma^2/B = 0.01\,\sigma^2/32$. Eight independent kicks add their variances: $8\cdot0.01\,\sigma^2/32 = 0.0025\,\sigma^2$. The big step adds one kick of variance $0.64\,\sigma^2/256 = 0.0025\,\sigma^2$. Same ✓.
So "multiply the batch by $k$ and the learning rate by $k$" gives (to first order) the same average progress and the same noise per epoch.
Linear scaling heuristic. If you multiply the batch size by $k$, multiply the learning rate by $k$ (and keep the number of epochs).
Derivation. After $k$ steps of size $\eta$ and batch $B$, with the gradient treated as constant, $\mathbf{w}$ has moved by $-\eta\sum_{j=1}^k\mathbf{g}_j = -k\eta\nabla f - \eta\sum_j\boldsymbol{\xi}_j$. The mean is $-k\eta\nabla f$ and the noise covariance is $k\cdot\eta^2\cdot(\sigma^2/B)$. One step with batch $kB$ and step $k\eta$ moves by $-k\eta\nabla f - k\eta\boldsymbol{\xi}'$ with the mean $-k\eta\nabla f$ and noise variance $(k\eta)^2\,\sigma^2/(kB) = k\,\eta^2\sigma^2/B$. The two agree. ∎
Why it is only a heuristic, with its limits.
- It assumes the gradient barely changes over $k$ small steps, i.e. $\eta L \ll 1$. It fails when $\eta$ is already near the stability limit $2/L$: you cannot scale $\eta$ beyond that. (Try the scan below: the best learning rate follows the diagonal and then hits a wall.)
- Early in training the weights change quickly, so large scaled steps can be unstable. A warm-up (starting with a small learning rate, Chapter 3.4) is the usual patch.
- It has a limit in $B$: once $B$ is above the gradient noise scale $\sigma^2/\|\nabla f\|^2$, the steps are already almost noise-free and larger batches add cost with little benefit (an empirical "critical batch size" idea).
- It was reported to work for ImageNet image-classification runs up to batch sizes of several thousand (about 8,000 in Goyal et al., 2017, together with warm-up), with accuracy dropping for larger batches. The best rule depends on the optimizer: for adaptive methods such as Adam, a square-root scaling is sometimes suggested. Treat all of these as rules of thumb, not theorems.
Generalisation remarks (empirical, not theorems). People often report that very large batches, trained for the same number of epochs without re-tuning, end with a worse test score than small batches (the "generalisation gap"; one suggested explanation is that noisier small-batch SGD prefers flatter minima). Later work found that much of the gap can be closed by scaling the learning rate and training longer. So "small batches generalise better" is a tendency that depends on tuning, not a law.
Why do we need it?
You often want a bigger batch (to use more GPUs and finish sooner) without redoing a whole learning-rate search. A scaling rule gives a starting point, and the noise formula tells you how far it can be pushed.
Where is it used?
Large-scale distributed training of image and language models, hyper-parameter transfer between small and large runs, and "train ImageNet in an hour" type experiments that use batch sizes in the thousands.
How is it used?
When you multiply $B$ by $k$, start with $\eta \to k\eta$ plus a warm-up of a few epochs, check that the loss curve per epoch matches the small-batch run, and back off if the training diverges or the final score drops.
"Bigger batch" and "faster training" are not the same. Per epoch a bigger batch gives fewer, better steps. Whether it reaches a target loss sooner in wall-clock time depends on the hardware and on how well the learning rate was scaled.
Do not trust a scaling rule blindly. The diagonal in the scan is a property of this smooth quadratic problem. For deep networks the rule works over some range of $B$ and then stops; always verify.
Small batches are noisy, not "bad". The noise does no harm to the final accuracy if you decay the learning rate, and it is a cheap form of exploration. Whether it also helps generalisation is an empirical question that depends on the model and the tuning.
Quick check: you move from $B = 128$, $\eta = 0.05$ to $B = 1024$. What learning rate does the linear scaling rule suggest, and what could go wrong?
$B$ grows by $k = 8$, so $\eta \to 8\cdot0.05 = 0.4$. Possible problems: $0.4$ may exceed the stable range $2/L$ for this loss; early training may be unstable (use warm-up); and above the gradient noise scale the larger batch gives little benefit, so the bigger step may not buy any speed-up per epoch.
Stochastic approximation and the Robbins–Monro step-sizes
Suppose you can only measure something with noise, and you want to find the input that makes it hit a target. You take a measurement, nudge the input in the right direction, measure again, nudge again. This is stochastic approximation (Robbins and Monro, 1951): finding a root of a function you can only see through noise. SGD is exactly this, where the "function" is the gradient and the target is zero.
How big should the nudges be? Think of the nudge size as fuel and as a noise budget:
- You need enough fuel: the nudges must add up to infinity, or you may run out of fuel before reaching the answer from far away.
- You need a finite noise budget: each nudge carries a random kick whose power grows with the square of its size, so the squares of the nudges must add up to something finite, or the noise never dies out.
Those are the two Robbins–Monro conditions.
Estimating an average as a stream. Numbers $z_1, z_2, \dots$ arrive one at a time ($z = 2, 4, 9, \dots$). We keep one running estimate $w$ and use the SGD-like update $w_k = w_{k-1} - \eta_k(w_{k-1} - z_k)$ (it is SGD on $f(w) = \tfrac12\mathbb{E}(w - z)^2$, whose minimiser is the mean of $z$).
- Choose $\eta_k = 1/k$. $\ w_1 = w_0 - 1\cdot(w_0 - 2) = 2$. (Whatever $w_0$ was is forgotten.)
- $w_2 = 2 - \tfrac12(2 - 4) = 3$, which is the average of $2$ and $4$.
- $w_3 = 3 - \tfrac13(3 - 9) = 5$, which is the average of $2, 4, 9$ ✓.
So the step-size $1/k$ makes SGD exactly the running mean. Check the two conditions: $\sum 1/k = \infty$ (harmonic series, plenty of fuel) and $\sum 1/k^2 = \pi^2/6 \lt \infty$ (finite noise budget). In contrast $\eta_k = 1/k^2$ has $\sum \lt \infty$: it can freeze at a wrong answer, and a constant $\eta$ has $\sum\eta^2 = \infty$: it never settles.
Stochastic approximation (Robbins–Monro, 1951). To find $\mathbf{w}^\star$ with $M(\mathbf{w}^\star) = \mathbf{0}$ when you can only observe $\mathbf{N}(\mathbf{w}) = M(\mathbf{w}) + \text{noise}$ (zero mean), iterate $$\mathbf{w}_{k+1} = \mathbf{w}_k - \eta_k\,\mathbf{N}(\mathbf{w}_k).$$ SGD is the case $M = \nabla f$. Robbins–Monro conditions on the step-sizes: $$\sum_{k=1}^\infty\eta_k = \infty,\qquad \sum_{k=1}^\infty\eta_k^2 \lt \infty .$$ Under these (plus technical assumptions, for example a smooth, strongly convex $f$ and noise with bounded power), $\mathbf{w}_k \to \mathbf{w}^\star$. The family $\eta_k = c/k^p$ satisfies both exactly when $\tfrac12 \lt p \le 1$.
Why both conditions (1-D quadratic $f = \tfrac12aw^2$, error $e_k = w_k$). $e_{k+1} = (1 - \eta_ka)e_k - \eta_k\xi_k$.
- Fuel. The average error obeys $\mathbb{E}e_{k+1} = (1 - \eta_ka)\mathbb{E}e_k$, so $\mathbb{E}e_k = \prod_{i \lt k}(1 - \eta_ia)\,e_0 \approx e^{-a\sum\eta_i}e_0$. This tends to $0$ if and only if $\sum\eta_i = \infty$. If $\sum\eta_i \lt \infty$ the product stays positive: the starting error is never fully forgotten.
- Noise budget. The variance obeys $V_{k+1} = (1 - \eta_ka)^2V_k + \eta_k^2s^2$: each step injects $\eta_k^2s^2$ of fresh variance. If $\sum\eta_k^2 \lt \infty$ the total ever injected is finite, and since old noise is forgotten (fuel condition) the variance tends to $0$. With a constant $\eta$ the injected amount $\eta^2s^2$ never shrinks: the noise ball.
Honest remark. These are sufficient conditions that fit all smooth settings. In a nice strongly convex problem like the one above, $\eta_k \to 0$ slowly with $\sum\eta_k = \infty$ (for example $\eta_k = c/\sqrt k$) already converges, only more slowly. The condition $\sum\eta_k^2 \lt \infty$ is the price for a clean general proof.
Online learning. If data arrive as a stream and each example $z_k$ is used once, SGD $\mathbf{w}_{k+1} = \mathbf{w}_k - \eta_k\nabla\ell(\mathbf{w}_k; z_k)$ is stochastic approximation for the population loss $\mathbb{E}_z\ell(\mathbf{w}; z)$: each gradient is an unbiased estimate of the true population gradient (this is the $n \to \infty$ limit of the finite-sum story). Because the next example has not been used yet, the loss on it before the update is an honest estimate of the test loss (called "progressive validation").
Why do we need it?
It gives the rule for "how must the step-size shrink so that noisy steps still lead to the answer?", and it covers data that arrive in a stream and can never be stored, where you cannot take a second pass.
Where is it used?
Online learning systems (click-through-rate models updated as clicks arrive), adaptive filters (the LMS algorithm in signal processing), temporal-difference learning and Q-learning in reinforcement learning (whose convergence proofs use these exact conditions), and the $1/\sqrt k$ and $1/k$ decay schedules.
How is it used?
Pick $\eta_k = c/(k + k_0)^p$ with $p \in (\tfrac12, 1]$. With $p = 1$ choose $c$ large enough relative to the curvature (for the mean example, $c = 1$). Check on a small run that the loss keeps falling and does not freeze early.
"Ση_k = ∞ and Ση_k² < ∞" is about the whole infinite run. For a finite run of $K$ steps what matters is the partial sums: enough total step to travel the distance, and small enough steps near the end to calm the noise.
$p = 1$ is delicate. The rate depends on $c$. With $\eta_k = c/k$ on a loss of curvature $a$, you need $ca \gt \tfrac12$ for the $O(1/k)$ rate, and a too-small $c$ gives a much slower rate. Slightly smaller $p$ (like $0.6$ to $0.8$) is more forgiving, which is why Polyak averaging (next section) is paired with it.
Constant step-sizes violate the conditions on purpose. Most deep-learning training uses a constant or piecewise-constant learning rate and accepts the noise ball, then lowers the rate at the end.
Quick check: for $\eta_k = 1/(k+1)$ and $w_0 = 0$ on the stream $z = 2, 4, 9$, what is $w_3$, and what is it the average of?
$w_1 = 0 - \tfrac12(0 - 2) = 1$. $\ w_2 = 1 - \tfrac13(1 - 4) = 2$. $\ w_3 = 2 - \tfrac14(2 - 9) = 3.75$. That is the average of $0, 2, 4, 9$: the starting value $w_0 = 0$ acts like one extra data point (that is the effect of $k+1$ instead of $k$).
The practical fixes: learning-rate decay and iterate averaging core
You have shaky hands and want a sharp photograph of the optimum. Two fixes:
- Hold the camera steadier as time goes on: decay the learning rate. Smaller steps mean a smaller noise ball, so eventually the iterate settles on the optimum. The price is slower travel, so you decay only after you have got close.
- Take many shaky photos and average them: iterate averaging. Each iterate jitters around the optimum in a different direction, so their average sits much closer to it than any single one. The iterates themselves are not changed.
Either fix works; using both is common.
Averaging, with made-up numbers to show the idea. A constant-step run keeps jittering around the optimum at $0$. Eight consecutive iterates are $0.4,\ -0.3,\ 0.1,\ -0.2,\ 0.3,\ -0.1,\ 0.2,\ -0.4$. Not one of them is at $0$ (typical size $0.25$), but their sum is $0.4 - 0.3 + 0.1 - 0.2 + 0.3 - 0.1 + 0.2 - 0.4 = 0$, so their average is $0$.
A real calculation. In the 1-D example (curvature $a = 1$, noise power $s^2 = 1$, $\eta = 0.2$) a single iterate has variance $V_\infty = 0.111$. The iterates are correlated (each is $0.8\times$ the last plus a kick), so averaging $N$ of them cuts the variance by a little less than $N$: it is about $\dfrac{s^2}{a^2N} = 1/N$. For $N = 100$ that is about $0.01$, eleven times smaller than $0.111$ (a simulation of 20,000 runs gave $0.0095$ against $0.110$). Note it does not depend on $\eta$.
Decay, in numbers. Noise-ball size scales with $\eta$. With $\eta_k = \eta_0/(1 + k/100)$ the step is $\eta_0/2$ at step 100, $\eta_0/6$ at step 500, and $\eta_0/16$ at step 1500: the ball has shrunk about 16-fold.
Learning-rate decay. Replace the constant $\eta$ by a schedule $\eta_k$ that shrinks. Common choices (plotted in Chapter 3.4): step decay (multiply by $0.1$ every so many epochs), inverse-time $\eta_k = \eta_0/(1 + k/\tau)$, exponential, and cosine decay. The Robbins–Monro family $\eta_k = c/(k + k_0)^p$, $\tfrac12 \lt p \le 1$, is the one with a convergence proof. Rule of thumb: keep $\eta$ large while the loss is falling fast, decay when it plateaus.
Iterate averaging (Polyak–Ruppert). Run SGD as usual and report the running average $$\bar{\mathbf{w}}_k = \frac{1}{k+1}\sum_{i=0}^{k}\mathbf{w}_i .$$ Theory (Polyak and Juditsky, 1992; assumptions: $f$ smooth and strongly convex, step-sizes $\eta_k \propto k^{-\alpha}$ with $\tfrac12 \lt \alpha \lt 1$): $\sqrt k\,(\bar{\mathbf{w}}_k - \mathbf{w}^\star)$ converges to a normal distribution whose covariance is the best possible one for any method using the same data. In words: slowly-decaying steps plus averaging is statistically as good as it gets, and it is forgiving about the exact schedule.
Practical variants. Tail averaging: average only the last half of the run so far (this discards the early iterates, which are far from the optimum and bias the plain average). Exponential moving average (EMA): $\mathbf{m}_k = \beta\,\mathbf{m}_{k-1} + (1-\beta)\,\mathbf{w}_k$ with $\beta$ close to $1$ (such as $0.999$), the usual choice in deep learning when you keep a smoothed copy of the weights for evaluation (an awareness-level technique; its theory is less complete). Note that averaging does not change the training iterates; it only changes which weights you evaluate.
Why do we need it?
A constant step leaves you with noisy final weights and a loss that stops at the noise floor. Decay and averaging remove most of the noise without needing a huge batch.
Where is it used?
Step or cosine decay in almost every image- and language-model training recipe; EMA of the weights in diffusion models, GANs and many vision models; Polyak averaging in classical stochastic approximation and in online convex optimisation; tail averaging in stochastic-gradient solvers for linear models.
How is it used?
Plan the schedule from the start (for example warm-up, then cosine to near zero over the whole budget), and evaluate the EMA or tail-averaged weights alongside the raw ones. If the averaged model is clearly better than the raw one, you were noise-limited.
Decay too early and you stall. A learning rate that falls before the iterate has travelled far leaves you far from the optimum, with a tiny noise ball around the wrong place. This is the "fuel" condition at work.
Averaging far from the optimum hurts. The plain average includes the early, far-away iterates and is biased by them. Use tail averaging or an EMA, or start averaging after a burn-in.
Averaging is easiest to justify near a minimum. On a non-convex deep network, averaging weights taken far apart can land in a bad spot; EMA weights are a common practical tool, but the theory above covers the convex case.
Quick check: a constant step gives a typical distance $0.3$ from the optimum. Roughly what distance would you expect from the average of 100 (correlated) iterates in the example above, and why is it not $0.03$?
The variance is about $1/N = 0.01$, so the distance is about $0.1$, not $0.3/\sqrt{100} = 0.03$. Averaging $N$ independent samples would give $0.03$, but consecutive iterates are correlated (each is $0.8\times$ the previous one plus a kick), so about 9 consecutive iterates carry the information of one independent sample.
Variance reduction: SVRG and SAGA (awareness)
The trouble with SGD is noise that does not vanish at the optimum, because each example's gradient is on its own not zero there. Here is a clever trick. Occasionally stop and compute the full gradient at a snapshot point $\tilde{\mathbf{w}}$. For the example you draw now, compare its gradient at the current point with its gradient at the snapshot. The difference is small when the two points are close, and the snapshot's full gradient supplies the average that is missing.
The result is still unbiased, but its noise shrinks to zero as you approach the optimum. This is a control variate: you subtract the part of the noise you can predict. Then a constant step-size is fine and the algorithm converges as quickly as gradient descent, per step.
The four-example problem again ($f_i = \tfrac12(w - a_i)^2$, $a = (1,3,5,7)$). Use the snapshot $\tilde w = 4$ (here it is the optimum, so $\mu = \nabla f(\tilde w) = 0$) and stand at $w = 2$.
- Gradients at $w = 2$: $(1, -1, -3, -5)$. Gradients at the snapshot $4$: $(3, 1, -1, -3)$.
- Differences $\nabla f_i(w) - \nabla f_i(\tilde w)$: $(-2, -2, -2, -2)$.
- Add $\mu = 0$: every example now reports the estimate $-2$, which is exactly the true gradient $f'(2) = -2$. Zero variance, whichever example you pick (plain SGD gave $1, -1, -3, -5$).
(All four losses have the same curvature here, which makes the cancellation perfect. In general the estimate is not exact, but its variance is small when $w$ is close to $\tilde w$.)
SVRG (stochastic variance-reduced gradient; Johnson and Zhang, 2013). Repeat for rounds $s = 1, 2, \dots$:
- Take a snapshot $\tilde{\mathbf{w}} \leftarrow \mathbf{w}$ and compute its full gradient $\tilde{\boldsymbol{\mu}} = \nabla f(\tilde{\mathbf{w}})$ (one full pass: $n$ example-gradients).
- For $m$ inner steps (often $m \approx n$): draw $i$ at random and set $$\mathbf{g} = \nabla f_i(\mathbf{w}) - \nabla f_i(\tilde{\mathbf{w}}) + \tilde{\boldsymbol{\mu}},\qquad \mathbf{w} \leftarrow \mathbf{w} - \eta\,\mathbf{g}.$$
Unbiased: $\mathbb{E}_i[\nabla f_i(\tilde{\mathbf{w}})] = \nabla f(\tilde{\mathbf{w}}) = \tilde{\boldsymbol{\mu}}$, so $\mathbb{E}[\mathbf{g}] = \nabla f(\mathbf{w}) - \tilde{\boldsymbol{\mu}} + \tilde{\boldsymbol{\mu}} = \nabla f(\mathbf{w})$ ✓.
Variance bound: if every $f_i$ is $L$-smooth then $\mathbb{E}\|\mathbf{g} - \nabla f(\mathbf{w})\|^2 \le \mathbb{E}\|\nabla f_i(\mathbf{w}) - \nabla f_i(\tilde{\mathbf{w}})\|^2 \le L^2\|\mathbf{w} - \tilde{\mathbf{w}}\|^2$. It goes to $0$ as $\mathbf{w}$ and $\tilde{\mathbf{w}}$ approach the optimum, so the noise ball disappears.
Rate (assumptions: each $f_i$ is $L$-smooth and $f$ is $\mu$-strongly convex; $\eta$ of order $1/L$ and $m$ of order $L/\mu$): linear convergence, with total cost of order $(n + L/\mu)\ln(1/\varepsilon)$ example-gradients. Compare gradient descent, $n\,(L/\mu)\ln(1/\varepsilon)$, and SGD, about $1/\varepsilon$ (order of magnitude, constants dropped). One round costs about $3n$ example-gradients ($n$ for the snapshot, $2$ per inner step).
SAGA (Defazio, Bach and Lacoste-Julien, 2014) gets the same effect differently: keep a table with the last gradient seen for every example, $\boldsymbol{\alpha}_1, \dots, \boldsymbol{\alpha}_n$. Draw $i$, use $\mathbf{g} = \nabla f_i(\mathbf{w}) - \boldsymbol{\alpha}_i + \tfrac1n\sum_j\boldsymbol{\alpha}_j$, then overwrite $\boldsymbol{\alpha}_i \leftarrow \nabla f_i(\mathbf{w})$. It is also unbiased, costs one gradient per step, but needs memory for $n$ stored gradients.
Why do we need it?
When you want high accuracy on a finite data set, plain SGD stalls at its noise floor and gradient descent is slow per step. Variance reduction keeps SGD's cheap steps and removes the floor.
Where is it used?
Convex problems on moderately large data: logistic regression, ridge and Lasso regression, SVMs and structured prediction (the SAGA method is available in scikit-learn as LogisticRegression(solver='saga'), and its relative SAG as solver='sag'). It is rarely used for deep networks.
How is it used?
Pick a constant step (of order $1/L$ divided by a small number), refresh the snapshot every epoch or two, and watch the loss gap fall straight on a log plot. For SAGA, make sure there is memory for the gradient table.
Not free. SVRG needs a full-gradient pass every round and two gradients per step. SAGA needs memory for $n$ gradients (for a million examples and a million weights that is far too much). The gain is for problems where accurate convergence on a fixed training set matters.
Why deep learning mostly skips it (commonly cited reasons, an active research question rather than a settled fact): the stored or snapshot gradients go stale quickly in a network, data augmentation and dropout make an example's gradient different on every visit, and early training is not noise-limited anyway.
Quick check: with the snapshot equal to the current point ($\tilde{\mathbf{w}} = \mathbf{w}$), what is the SVRG gradient estimate, and what is its variance?
$\mathbf{g} = \nabla f_i(\mathbf{w}) - \nabla f_i(\mathbf{w}) + \nabla f(\mathbf{w}) = \nabla f(\mathbf{w})$: exactly the true gradient, whichever example is drawn, so the variance is $0$. As $\mathbf{w}$ moves away from the snapshot during the round, the variance grows (like $\|\mathbf{w} - \tilde{\mathbf{w}}\|^2$) until the next snapshot resets it.
Recap, cheat sheet and practice
- A training loss is an average $f = \frac1n\sum f_i$. A stochastic (mini-batch) gradient averages $B$ random examples' gradients and is unbiased: $\mathbb{E}[\mathbf{g}_B] = \nabla f$ (proved above, with or without replacement).
- Gradient noise has mean zero and power $\sigma^2 = \frac1n\sum\|\nabla f_i\|^2 - \|\nabla f\|^2$. At the optimum $\nabla f = \mathbf{0}$ but $\sigma^2 \gt 0$ unless every example is fitted exactly (interpolation), so a constant step never settles.
- A batch gradient has variance $\sigma^2/B$ (times $\frac{n-B}{n-1}$ without replacement): error $\propto 1/\sqrt B$, so 4 times the batch halves the noise: diminishing returns.
- SGD $\mathbf{w} \leftarrow \mathbf{w} - \eta\,\mathbf{g}_B$: with a constant step it reaches a noise ball whose loss gap and squared radius scale like $\eta\sigma^2/B$; with Robbins–Monro steps ($\sum\eta_k = \infty$, $\sum\eta_k^2 \lt \infty$) it converges. Rates (with assumptions): $O(1/k)$ strongly convex, $O(1/\sqrt K)$ convex (averaged), linear to a floor for a constant step.
- Batch size trades steps per second against quality per step; the linear scaling heuristic ($B \to kB$, $\eta \to k\eta$) keeps the mean progress and the noise per epoch (valid when $\eta L \ll 1$, with limits). Generalisation remarks about batch size are empirical.
- The practical fixes for noise are learning-rate decay, iterate averaging (Polyak–Ruppert, tail averaging, EMA) and, for finite sums, variance reduction (SVRG, SAGA: unbiased estimates whose variance vanishes at the optimum).
Cheat sheet
| Idea | Formula | Picture |
|---|---|---|
| Finite sum, stochastic gradient | $f = \frac1n\sum f_i$, $\ \mathbb{E}[\nabla f_I] = \nabla f$ | a poll: noisy but unbiased |
| Noise power | $\sigma^2 = \frac1n\sum\|\nabla f_i\|^2 - \|\nabla f\|^2$ | spread of the per-example gradients |
| Batch variance | $\sigma^2/B$ (with repl.), $\ \frac{n-B}{n-1}\frac{\sigma^2}{B}$ (without) | error shrinks as $1/\sqrt B$ |
| SGD step | $\mathbf{w} \leftarrow \mathbf{w} - \eta_k\mathbf{g}_B(\mathbf{w})$ | noisy compass downhill |
| Noise ball (1-D quadratic) | $\mathbb{E}w^2 \to \dfrac{\eta s^2}{a(2-\eta a)}$, loss gap $\dfrac{\eta s^2}{2(2-\eta a)}$ | jitter around the optimum, $\propto\eta$ |
| Strongly convex, constant $\eta \le 1/L$ | $\le (1-\eta\mu)^k(\cdot) + \dfrac{\eta L\sigma^2}{2\mu}$ | linear, then a floor |
| Robbins–Monro | $\sum\eta_k = \infty$, $\ \sum\eta_k^2 \lt \infty$; $\ \eta_k = c/k^p$, $\tfrac12 \lt p \le 1$ | enough fuel, finite noise budget |
| Rates | $O(1/k)$ str. convex (decaying), $O(1/\sqrt K)$ convex (averaged) | slower per step, cheaper per step |
| Linear scaling (heuristic) | $(B, \eta) \to (kB, k\eta)$ | $k$ small steps $\approx$ one big step |
| Iterate averaging | $\bar{\mathbf{w}}_k = \frac{1}{k+1}\sum_{i\le k}\mathbf{w}_i$ (or tail / EMA) | jitter cancels |
| SVRG estimate | $\nabla f_i(\mathbf{w}) - \nabla f_i(\tilde{\mathbf{w}}) + \nabla f(\tilde{\mathbf{w}})$ | unbiased, variance $\to 0$ |
import numpy as np
rng = np.random.default_rng(0)
# ---- data: least squares, f_i(w) = 0.5 * (x_i . w - y_i)^2, f = mean of the f_i -----------
n, d = 1000, 2
X = rng.standard_normal((n, d)); w_true = np.array([1.5, -1.0])
y = X @ w_true + rng.standard_normal(n) # noisy labels
w_opt = np.linalg.lstsq(X, y, rcond=None)[0]
f = lambda w: 0.5 * np.mean((X @ w - y) ** 2)
per_example = lambda w: (X @ w - y)[:, None] * X # row i = gradient of example i, shape (n, d)
# ---- 1. stochastic gradients are unbiased, with variance sigma^2 / B ----------------------
w = np.array([0.0, 0.0]); G = per_example(w); full = G.mean(axis=0)
sigma2 = ((G - full) ** 2).sum(axis=1).mean() # noise power of ONE example at w
print("sigma^2 =", round(sigma2, 3)) # 12.44 for this seed
for B in (1, 4, 16, 64):
est = np.array([G[rng.integers(0, n, B)].mean(axis=0) for _ in range(20000)])
err = ((est - full) ** 2).sum(axis=1).mean()
print(B, round(err, 3), round(sigma2 / B, 3), round(np.abs(est.mean(axis=0) - full).max(), 3))
# B=1: 12.14 vs 12.44 | B=4: 3.12 vs 3.11 | B=16: 0.79 vs 0.78 | B=64: 0.20 vs 0.19 | bias of the mean ~ 0.01 or less
# ---- 2. SGD: noise ball for a constant step, convergence for a decaying step ---------------
def sgd(eta, B=8, steps=4000, average=False):
w = np.array([-1.0, 2.0]); hist = []
for k in range(steps):
i = rng.integers(0, n, B); g = ((X[i] @ w - y[i])[:, None] * X[i]).mean(axis=0)
w = w - eta(k) * g; hist.append(w)
return np.mean(hist[steps // 2:], axis=0) if average else w # tail average of the last half
gap = lambda w: f(w) - f(w_opt)
print(np.mean([gap(sgd(lambda k: 0.1)) for _ in range(20)])) # about 5e-3 (the noise floor, proportional to eta)
print(np.mean([gap(sgd(lambda k: 0.1 / (1 + k / 100))) for _ in range(20)])) # about 2e-4 (decaying step: the ball shrinks)
print(np.mean([gap(sgd(lambda k: 0.1, average=True)) for _ in range(20)])) # about 6e-5 (average of the iterates: jitter cancels)
# ---- 3. stochastic approximation: step 1/k turns SGD into the running mean ------------------
z = rng.normal(3.0, 1.0, 2000); m = 0.0
for k, zk in enumerate(z, start=1):
m -= (1.0 / k) * (m - zk)
print(m, z.mean()) # 3.0049 3.0049: identical numbers
# ---- 4. SVRG: constant step, yet it converges linearly (the noise vanishes) -----------------
def svrg(eta=0.05, outer=8):
w = np.array([-1.0, 2.0])
for s in range(outer):
w_snap = w.copy(); mu = per_example(w_snap).mean(axis=0) # full gradient at the snapshot
for _ in range(n):
i = rng.integers(0, n); gi = (X[i] @ w - y[i]) * X[i]; gs = (X[i] @ w_snap - y[i]) * X[i]
w = w - eta * (gi - gs + mu) # variance-reduced gradient
print(s + 1, gap(w)) # 0.68, 5e-3, 2e-4, 2e-5, ... 7e-11: linear convergence
svrg()
1. At a fixed point $\mathbf{w}$, what is true about a mini-batch gradient $\mathbf{g}_B$ (batch drawn uniformly at random)?
2. A batch of 16 examples gives a typical gradient error of $0.8$. About what error do you expect from a batch of 64?
3. SGD with a constant learning rate is run on a training loss whose labels contain noise. At the exact minimiser of the training loss, what happens?
4. Which step-size schedule satisfies both Robbins–Monro conditions, $\sum\eta_k = \infty$ and $\sum\eta_k^2 \lt \infty$?
5. You change from $B = 64$, $\eta = 0.05$ to $B = 512$. The linear scaling heuristic suggests a new learning rate of about
6. What makes the SVRG gradient estimate $\nabla f_i(\mathbf{w}) - \nabla f_i(\tilde{\mathbf{w}}) + \nabla f(\tilde{\mathbf{w}})$ useful?
Practice problems
A. At the current weights three examples have gradients $-4,\ 2,\ 8$. Find the full gradient, the noise power $\sigma^2$, and the variance of a batch of size 2 drawn with and without replacement.
Full gradient $(-4 + 2 + 8)/3 = 2$. Deviations from it: $-6, 0, 6$, so $\sigma^2 = (36 + 0 + 36)/3 = 24$. With replacement: $\sigma^2/B = 12$. Without replacement: $\frac{n-B}{n-1}\cdot\frac{\sigma^2}{B} = \frac{3-2}{3-1}\cdot12 = 6$. Check by listing the three pairs: averages $(-4+2)/2 = -1$, $(-4+8)/2 = 2$, $(2+8)/2 = 5$; mean $2$ ✓, deviations $-3, 0, 3$, variance $18/3 = 6$ ✓.
B. A 1-D loss $f(w) = \tfrac12\cdot2\,w^2$ (so $a = 2$) is trained with SGD whose per-step gradient noise power is $s^2 = 1$ and step $\eta = 0.1$. Find the stationary $\mathbb{E}[w^2]$ and the expected loss gap. What happens with $\eta = 0.05$?
$V_\infty = \dfrac{\eta s^2}{a(2 - \eta a)} = \dfrac{0.1}{2\cdot1.8} = 0.0278$. Loss gap $= \tfrac a2V_\infty = 0.0278$. With $\eta = 0.05$: $V_\infty = 0.05/(2\cdot1.9) = 0.01316$, about half (the ratio $2.1$ is slightly above $2$ because of the $2 - \eta a$ factor). The ball size is roughly proportional to $\eta$.
C. Which of these step-size sequences satisfy $\sum\eta_k = \infty$ and $\sum\eta_k^2 \lt \infty$: (a) $1/k$, (b) $0.1$, (c) $1/\sqrt k$, (d) $1/k^2$, (e) $3/(k+10)$?
(a) yes. (b) no: $\sum\eta^2 = \infty$. (c) no: $\sum 1/k = \infty$ for the squares. (d) no: $\sum 1/k^2 \lt \infty$, so the steps run out of fuel. (e) yes: $\sum 3/(k+10) = \infty$ and $\sum 9/(k+10)^2 \lt \infty$.
D. You scale from $B = 32$, $\eta = 0.02$ to $B = 512$ ($k = 16$). Use the linear scaling rule, and verify that 16 small steps and one large step have the same noise variance.
New step $\eta' = 16\cdot0.02 = 0.32$. Noise of 16 small steps: $16\cdot\eta^2\sigma^2/B = 16\cdot0.0004\,\sigma^2/32 = 0.0002\,\sigma^2$. Noise of one big step: $\eta'^2\sigma^2/B' = 0.1024\,\sigma^2/512 = 0.0002\,\sigma^2$ ✓. The mean movement is also equal ($16\cdot0.02 = 0.32$).
E. Prove that the SAGA estimate $\mathbf{g} = \nabla f_i(\mathbf{w}) - \boldsymbol{\alpha}_i + \frac1n\sum_j\boldsymbol{\alpha}_j$ is unbiased when $i$ is uniform on $\{1, \dots, n\}$ (and the table $\boldsymbol{\alpha}$ is fixed at the time of the draw).
$\mathbb{E}_i[\nabla f_i(\mathbf{w})] = \nabla f(\mathbf{w})$. $\ \mathbb{E}_i[\boldsymbol{\alpha}_i] = \frac1n\sum_j\boldsymbol{\alpha}_j$. So $\mathbb{E}[\mathbf{g}] = \nabla f(\mathbf{w}) - \frac1n\sum_j\boldsymbol{\alpha}_j + \frac1n\sum_j\boldsymbol{\alpha}_j = \nabla f(\mathbf{w})$ ✓. (The same argument works for SVRG with $\boldsymbol{\alpha}_i = \nabla f_i(\tilde{\mathbf{w}})$.)
F. In the 1-D example ($a = 1$, noise power $s^2 = 1$, $\eta = 0.2$) a single iterate has variance $0.111$. About what variance does the average of $N = 400$ constant-step iterates have? Why is the answer far smaller than the plain iterate's?
About $s^2/(a^2N) = 1/400 = 0.0025$ (a standard deviation near $0.05$, against $0.33$ for one iterate): roughly $44$ times smaller variance. The iterates jitter in different directions and the jitter cancels in the mean. It is less than a factor of $N$ better than the plain iterate because consecutive iterates are correlated (here each is $0.8\times$ the previous one plus a kick).
Optimization in Deep Learning
Everything so far was clean: smooth bowls, exact rates, proofs. A real neural network is none of those. Its loss is a wild, non-convex landscape in millions of dimensions, and yet we train it with plain gradient methods and it usually works. This chapter explains what makes deep-learning optimization hard, what is proved, what is only observed, and how to diagnose a training run that has gone wrong.
- See why neural-network losses are non-convex, and what that does (and does not) break
- Picture loss landscapes, saddle points and local minima honestly: what is known, what is a heuristic
- Understand vanishing and exploding gradients, and the fixes: initialisation, ReLU, residual connections, normalisation, clipping
- Connect ill-conditioning to slow training, and to why adaptive methods and normalisation help
- Separate optimisation (training loss) from generalisation (test loss)
- Know how batch size, learning rate and weight decay interact in practice
- Use a symptom-to-cause checklist when a training run is NaN, flat, spiking or overfitting
How to read this chapter. Deep-learning optimization mixes three kinds of statements, and we label them: proved (a theorem, with its assumptions), empirical (observed in many experiments, but with no general proof) and rule of thumb (useful, not guaranteed). Where we use the symbols from earlier chapters: parameters $\theta\in\mathbb{R}^n$ (all the weights), loss $L(\theta)$, learning rate $\eta$ (see Chapter 3.3), gradient $\nabla L$ (a column vector, see the Calculus guide) and Hessian $\nabla^2L$ (second derivatives).
Non-convex optimization: why networks are not bowls core
In Chapter 3.6 you met the friendly case: a convex loss is one smooth bowl. Wherever you start, walking downhill ends at the one lowest point. Linear regression and logistic regression are like this.
A neural network is different. Its loss is a landscape with many valleys, ridges, flat plains and passes between hills. Three things make it that way:
- Weights multiply each other. A deep network computes products like $w_2\cdot w_1\cdot x$. A product of two unknowns is not a bowl: it is a saddle.
- Bends in the middle. Activations such as $\tanh$ and ReLU bend the signal between layers, so the output is not a straight function of the weights.
- Copies. Swapping two hidden neurons gives exactly the same network, so every good solution comes in many identical-looking copies. Several copies of a valley with a hill between them cannot be one bowl.
Training still uses the same downhill walk (gradient descent). What we lose is the guarantee that it ends at the best point.
The smallest "network" with a product: prediction $\hat y = a\cdot b\cdot x$ with two weights $a, b$. One data point $x = 1$, target $t = 1$. The loss is $L(a,b) = (ab - 1)^2$.
- $L(1,1) = (1-1)^2 = 0$ and $L(-1,-1) = ((-1)(-1)-1)^2 = 0$. Both are perfect solutions.
- The point halfway between them is $(0,0)$, and $L(0,0) = (0-1)^2 = 1$.
- A convex function must satisfy $L(\text{midpoint}) \le \tfrac12\big(L(p)+L(q)\big) = \tfrac12(0+0) = 0$. But $1 \gt 0$. So $L$ is not convex.
- The gradient is $\nabla L = 2(ab-1)\,(b,\ a)$, which is $(0,0)$ at the origin: a flat spot. The Hessian there is $\begin{bmatrix}0&-2\\-2&0\end{bmatrix}$ with eigenvalues $+2$ and $-2$: uphill one way, downhill the other. A saddle point (Chapter 3.2).
Notice also that $ab = 1$ is a whole curve of perfect solutions ($(2,\tfrac12)$, $(\tfrac12,2)$, …): a "valley floor" with many equal minima, and a saddle in the middle.
A function is non-convex when it is not convex: there exist points $p, q$ with $$L\!\left(\tfrac{p+q}{2}\right) \;\gt\; \tfrac12 L(p) + \tfrac12 L(q).$$ A non-convex optimization problem may have local minima that are not global, saddle points, flat plateaus and ridges. Nothing forces gradient descent to end at the global minimum.
Symmetries. Two parameter vectors $\theta,\theta'$ are equivalent if the network computes the same function with both. Two common sources:
- Permutation: swapping the order of the hidden units in a layer (and the matching rows and columns of the neighbouring weight matrices). A layer with $h$ units gives $h!$ equivalent re-orderings.
- Sign / scale: for an odd activation like $\tanh$, flipping the signs of one unit's incoming and outgoing weights changes nothing ($2^h$ more copies). For ReLU, multiplying a unit's incoming weights by $c>0$ and its outgoing weights by $1/c$ changes nothing (a continuous family).
Why do we need it?
It sets honest expectations. Every convergence guarantee from the convex chapters (local = global, exact rates) is gone for a deep network. We need to know which parts of that story survive, and which are replaced by experiments.
Where is it used?
Training any neural network: CNNs, Transformers, GANs, recurrent models. Also tensor factorisation, matrix completion, mixture models and k-means, which are non-convex too.
How is it used?
Accept that the answer depends on the start. Use random initialisation, run a few seeds, and judge a run by validation performance, not by "did it reach the global minimum" (which you can never check).
What is proved, what is observed, what is open. Be careful with sweeping statements such as "neural-network training has no bad local minima". The honest picture:
| Statement | Status | Details |
|---|---|---|
| Training even a tiny network to its global optimum can be NP-hard | Proved | Worst case, for specific tiny architectures and datasets (Blum and Rivest). It says some instances are hard, not that typical ones are. |
| Deep linear networks have no bad local minima: every local minimum is global, and the other critical points are saddles (some of them degenerate) | Proved | Under assumptions on the data (Kawaguchi, 2016). Linear networks are a toy: the non-linearity is exactly what is missing. |
| Gradient descent from a random start almost never converges to a strict saddle | Proved | For small enough steps and functions whose saddles have a negative-curvature direction (Lee et al., 2016). |
| Very wide networks trained by gradient descent reach near-zero training loss | Proved in a regime | Needs enormous width and a particular initialisation scale ("NTK" or "lazy" regime). Practical networks are usually outside the assumptions. |
| Standard networks trained with SGD reach very low training loss on large datasets | Empirical | Even on random labels (Zhang et al., 2017). No general theorem covers practical architectures. |
| Independently trained solutions are connected by low-loss curved paths | Empirical | Mode-connectivity experiments (Garipov et al., Draxler et al., 2018). Straight lines between them still show a barrier (see the next section). |
| Why the solutions found by SGD generalise to new data | Open | Many partial explanations, no complete theory (see "Optimization vs generalization" below). |
- Non-convex does not mean hopeless. PCA, matrix factorisation and many deep-learning tasks are non-convex yet train reliably.
- "Convex" and "easy" are not the same word. Convex gives a guarantee; non-convex gives none either way. The guarantee is what is missing, not the success.
- Symmetry-copies are harmless by themselves: they have equal loss, so landing in any copy is equally good. Trouble comes from the points between copies (like the origin above).
Quick check: for $L(a,b)=(ab-1)^2$, is $(2,\tfrac12)$ a minimum? Is $(0,0)$?
$L(2,\tfrac12) = (1-1)^2 = 0$, the lowest possible value of a square, so yes, it is a (global) minimum. $(0,0)$ has $L=1$ and is a saddle: the Hessian eigenvalues are $+2$ and $-2$, so you can reduce the loss by moving along the direction $(1,1)$ (where $ab = t^2$ grows towards 1).
Loss landscapes: slices, surfaces and what they hide core
The loss $L(\theta)$ assigns a height to every point $\theta$ in weight space. For two weights that is a real landscape you can look at: hills, valleys, plateaus. For a network with ten million weights it lives in ten million dimensions, and nobody can draw it.
So people take slices: walk along one straight line (or across one flat sheet) through weight space and plot the loss on the way. A line gives a curve (1D). A sheet gives a surface or contour map (2D). It is like judging a whole mountain range from a few hiking trails: useful, but each trail shows only a sliver.
Take the two-unit network from the widget below: loss $L(w_1,w_2)$ with two perfect solutions $\theta_A = (2,-1)$ and $\theta_B = (-1,2)$. They are the same network with the two hidden units swapped. Walk the straight line between them: $\theta(\alpha) = (1-\alpha)\,\theta_A + \alpha\,\theta_B$.
- $\alpha = 0$: $\theta = (2,-1)$, loss $0$. $\alpha = 1$: $\theta = (-1,2)$, loss $0$.
- $\alpha = 0.5$: $\theta = (0.5\cdot 2 + 0.5\cdot(-1),\ 0.5\cdot(-1)+0.5\cdot 2) = (0.5, 0.5)$.
- The loss there is about $0.82$: a high ridge between two identical, perfect solutions.
Moral: two good solutions that look the same can be separated by a barrier on the straight line. Averaging their weights ("weight averaging") works only when the solutions lie in the same valley.
1D slice. Pick a centre $\theta_0$ (often the trained weights) and a direction $\mathbf{d}$. Plot $\;\alpha\mapsto L(\theta_0+\alpha\,\mathbf{d})$. With two trained solutions, use the line through both, as above.
2D slice. Pick two directions $\mathbf{d}_1,\mathbf{d}_2$ and plot $\;(\alpha,\beta)\mapsto L(\theta_0+\alpha\mathbf{d}_1+\beta\mathbf{d}_2)$ as a surface or as contour lines (level sets).
How directions are chosen. Random directions are common. Because rescaling weights can leave the function unchanged (the scale symmetry above), a raw random direction can be misleading. Filter normalisation (Li et al., 2018) rescales each filter of the random direction to the size of the matching filter in $\theta_0$, so that pictures of different networks can be compared.
Caveats. A slice shows $L$ on one line or sheet only. A point that looks like a minimum in your slice can still be a saddle in a direction you did not draw. Smoothness, "flatness" and "number of valleys" in a picture depend on the chosen directions and the axis scaling.
Why do we need it?
We cannot see millions of dimensions, but we want intuition about why training behaves as it does. A slice is the only direct look at the loss we get, and it exposes barriers between solutions.
Where is it used?
Papers on sharp versus flat minima, residual connections smoothing the loss surface, mode connectivity, model averaging (SWA, model soups), and TensorBoard-style loss-surface plots made with the loss-landscape tools.
How is it used?
Save weights $\theta_A$ (and maybe $\theta_B$). Evaluate the loss on about 20 to 50 points along $\theta_A+\alpha(\theta_B-\theta_A)$. Plot loss against $\alpha$. A bump between two good ends is a barrier.
- Pictures flatter. Famous "smooth bowl" and "rough mountain" plots are 2D slices with chosen scales. They support arguments (for example that skip connections smooth the surface) but they do not prove them.
- Axis scale decides what you see. Doubling the direction length changes how sharp a minimum looks. Compare networks only with a consistent normalisation.
- A slice hides the dimension. In 10 million dimensions a random line almost never aligns with the few steep directions (see the Hessian-spectrum idea in the ill-conditioning section), so random slices look much flatter than the true local shape.
Quick check: two trained networks both have training loss $0.01$. The loss at the average of their weights is $2.3$. What does this tell you?
The straight line between them crosses a high-loss region: they are in different basins (or different symmetry copies), and plain weight averaging fails for this pair. It does not mean the two networks are bad, and it does not rule out a curved low-loss path between them.
Saddle points in high dimensions core
A saddle point is a place where the ground is perfectly flat (the gradient is zero) but it is not the bottom of a valley. Think of a mountain pass: along the road you climb up to the pass and then down; across the road the mountain rises on both sides. A ball balanced exactly on the pass would stay there. Any tiny push sends it down one side.
With only two weights a saddle is a rare curiosity. With a million weights, every flat spot has a million independent directions to curve in. For a point to be a true minimum, all of them must curve upward. For a saddle, it is enough that one curves downward. That is why, in high dimensions, saddles are far more common than minima, at least when the loss is high.
Saddles are dangerous in practice for a quiet reason: near one the gradient is tiny, so training slows to a crawl. It is a plateau in disguise.
The simplest saddle is $f(x,y) = x^2 - y^2$. Its gradient $(2x,\,-2y)$ is zero at the origin, and its Hessian is $\mathrm{diag}(2,-2)$. Run gradient descent with learning rate $\eta = 0.1$. Each step multiplies the coordinates:
- $x \leftarrow x - 0.1\cdot 2x = 0.8\,x$ (shrinks to $0$).
- $y \leftarrow y - 0.1\cdot(-2y) = 1.2\,y$ (grows).
- Start at $(1,\ 10^{-6})$. After $k$ steps $y = 10^{-6}\cdot 1.2^k$. It reaches size $1$ when $1.2^k = 10^6$, i.e. $k = \ln(10^6)/\ln(1.2) \approx 75.8$ steps.
- Start at $y = 0$ exactly: then $y$ stays $0$ forever, and gradient descent converges to the saddle.
Now make the saddle flat: $f = x^2 - 0.05\,y^2$, with downhill curvature only $-0.1$. Then $y \leftarrow y - 0.1\cdot(-0.1\,y) = 1.01\,y$, and escaping from $10^{-6}$ takes $\ln(10^6)/\ln(1.01) \approx 1388$ steps instead of 76. Flat saddles are slow.
A critical point has $\nabla L(\theta)=\mathbf{0}$. Look at the eigenvalues $\lambda_1\le\dots\le\lambda_n$ of the Hessian there (Calculus guide: saddle points, Chapter 3.2):
- all $\lambda_i>0$: a local minimum;
- all $\lambda_i<0$: a local maximum;
- some positive, some negative: a saddle. If $\lambda_1<0$ it is a strict saddle: there is a direction of strictly negative curvature, so a small push can escape;
- some $\lambda_i = 0$ and none negative: the second-order test is inconclusive (flat directions).
The index of a critical point is the number of negative eigenvalues. A saddle with index $k$ has $k$ independent directions to escape along.
What is known. Proved: gradient descent from a random start converges to a strict saddle with probability zero (Lee et al., 2016); adding small random perturbations lets gradient descent leave strict saddles in a number of steps that grows only polynomially (Ge et al., 2015) or even only polylogarithmically (Jin et al., 2017) with the dimension, under smoothness assumptions on the loss. Heuristic or empirical: in random models of high-dimensional landscapes the fraction of negative eigenvalues increases with the loss value, so critical points with high loss are mostly saddles and the few with low loss are minima (Bray and Dean; Dauphin et al., 2014). How well this describes real networks is only partly supported.
Why do we need it?
It tells you what "the gradient is almost zero" really means. Small gradients can come from a minimum (good), a saddle or a plateau (bad). You need the Hessian's signs, or an experiment, to tell which.
Where is it used?
The theory of why SGD with noise or momentum works on non-convex problems, perturbed gradient descent, second-order methods that use negative curvature, and diagnosing a loss curve that sits flat for a long time before dropping.
How is it used?
If the loss is stuck on a long plateau with tiny gradients, do not assume convergence. Add noise (smaller batches), use momentum or Adam, or restart with another seed. Check the Hessian's smallest eigenvalue on a small model if you need certainty.
- "Saddles are the main obstacle" is a heuristic, not a theorem. In practice, slow training more often comes from ill-conditioning, plateaus from saturated activations, a bad learning rate or noisy gradients. Saddles are one of several reasons.
- Flat directions. A critical point with many near-zero Hessian eigenvalues is neither a clean saddle nor a clean minimum: the loss barely changes along those directions. Such flat valley floors are very common in over-parameterised networks.
- Escape takes time. The escape time grows like $1/|\lambda_{\min}|$. A saddle with a tiny negative eigenvalue can hold a run for thousands of steps even though it is "strict".
- In the widget, noise is added to every gradient. Real mini-batch noise (Chapter 3.14) is not isotropic, but it plays the same role of kicking the iterate off the unstable line.
Quick check: at a critical point the Hessian has eigenvalues $+3,\ +1,\ 0,\ -0.01$. What can you say?
There is a negative eigenvalue, so it is a (strict) saddle, not a minimum. But the negative curvature is tiny ($-0.01$), so escape along that direction is very slow: it grows by only a factor $1 + 0.01\,\eta$ per step (about $1.001$ when $\eta=0.1$). The zero eigenvalue is another flat direction. Gradient descent can look "converged" here for a long time.
Local minima: "good enough" and "many equivalent" core
A local minimum is the bottom of some valley: every nearby point is higher. It need not be the lowest valley of the whole landscape. Walk downhill from a random spot and you stop in whichever valley you happened to be in.
In the classic picture (the ball on the intro page) that is a disaster: the shallow valley is much worse than the deep one. For modern neural networks the picture is usually kinder, for three reasons:
- Many valleys are copies. By symmetry (neuron swaps) a huge number of valleys have exactly the same depth.
- Valleys are not points. With more weights than data points, the set of zero-loss weights is a continuous surface, not isolated holes.
- "Good enough" is enough. We care about performance on new data, not about the last digit of the training loss. Different local minima often give similar test accuracy (empirical, not proved).
The intro landscape $f(x) = 0.1x^4 - 0.8x^2 + 0.4x$. Set $f'(x) = 0.4x^3 - 1.6x + 0.4 = 0$. The roots are $x\approx-2.115,\ 0.254,\ 1.861$.
- $f(-2.115) \approx -2.424$ (the global minimum), $f(1.861)\approx -0.827$ (a worse local minimum), and $x \approx 0.254$ is the local maximum between them ($f\approx 0.050$).
- The gap between the two valleys is $-0.827-(-2.424) = 1.60$: here the local minimum really is worse.
- Start gradient descent at $x_0=1.5$ and you end at $1.861$; start at $x_0=-1$ and you end at $-2.115$. The start decides the answer.
$\theta^\star$ is a local minimum if there is a radius $r>0$ with $L(\theta)\ge L(\theta^\star)$ for all $\|\theta-\theta^\star\|<r$. It is a global minimum if $L(\theta)\ge L(\theta^\star)$ for all $\theta$. Its basin of attraction is the set of starting points from which gradient descent ends there.
Counting heuristic. If a network has $n$ weights and must fit $m$ examples exactly, that is roughly $m$ equations in $n$ unknowns. When $n>m$, solutions typically form a surface of dimension about $n-m$ (a rough dimension count, not a theorem for every network). For $n=10^7$ and $m=10^6$ the zero-loss set has about $9\times10^6$ free directions.
Known and unknown. Proved: in convex problems every local minimum is global (Chapter 3.6); in deep linear networks too (see the table above). Empirical: in large networks, runs from different seeds typically reach similar training loss. Not known in general: that real networks have no bad local minima. Small networks do have them, and special constructions can be built for larger ones.
Why do we need it?
So that you do not waste effort or lose sleep chasing the global minimum of a network loss. You need to know when a local minimum is a real problem (small or badly-set-up models) and when it is not (big, over-parameterised models).
Where is it used?
Choosing between "train many seeds and pick the best" and "train once": k-means and Gaussian-mixture fitting use many restarts, while large networks usually train once. Also ensembles, learning-rate warm restarts, and model soups.
How is it used?
For small non-convex models, restart from several random initialisations and keep the best validation score. For big networks, change the seed only to estimate the spread, and judge by held-out performance rather than the training-loss digits.
- "Good enough" is task-dependent. It is a statement about validation performance. A training-loss gap that looks tiny can still matter, and one that looks large can be irrelevant.
- Do not claim "there are no bad local minima" for deep networks. This is proved only for special settings (linear networks, extremely wide ones). It is a useful working belief, not a theorem.
- Equal-loss minima are not necessarily equal in quality: two networks with identical training loss can generalise differently (see the generalization section).
Quick check: two runs end with training loss $0.002$ and $0.0021$ but their weights look completely different. Is that a contradiction?
No. Neuron-swap symmetry alone produces completely different-looking weight vectors for the same function, and zero-loss solutions form wide surfaces. Equal loss with different weights is the normal case for big networks.
Vanishing gradients: why depth was hard, and the fixes core
Backpropagation sends an error signal from the loss back to the first layer, and every layer on the way multiplies the signal by "its weights" and "the slope of its activation". If those factors are usually a bit below 1, the signal fades like a whisper passed down a line of 50 people. By the time it reaches the early layers it is almost zero, so those layers stop learning. This is the vanishing gradient problem. (The basic multiplication is in the Calculus guide. Here we ask: what do we do about it?)
There are two culprits, and each has a cure:
- Saturating activations. A sigmoid has slope at most $0.25$, and nearly $0$ when its input is large. Cure: use activations whose slope stays near 1 (ReLU and relatives).
- Weights of the wrong size. If each layer shrinks the signal, depth multiplies the shrinking. Cure: choose the starting weights so that each layer keeps the signal size roughly constant (Xavier or He initialisation).
Two more structural cures change the network itself: residual connections give the signal a bypass, and normalisation layers re-centre and re-scale the signal at every layer.
(a) Saturation. Twenty sigmoid layers: even in the best case (slope $0.25$ at every layer) the slopes alone multiply to $0.25^{20} = 2^{-40}\approx 9\times10^{-13}$.
(b) Initialisation for ReLU. One layer computes $z_i=\sum_{j=1}^{n} W_{ij}h_j$ with independent weights of variance $\sigma^2$ and mean $0$. Then
- $\mathrm{Var}(z_i) = \sum_j \sigma^2\,\mathbb{E}[h_j^2] = n\,\sigma^2\,\mathbb{E}[h^2]$ (the $n$ independent terms add their variances).
- After ReLU, $h=\max(0,z)$. If $z$ is symmetric around $0$, half of the squared mass is cut off: $\mathbb{E}[h^2] = \tfrac12\mathrm{Var}(z)$.
- So one layer maps $\mathbb{E}[h^2_{\text{prev}}]\mapsto \tfrac{n\sigma^2}{2}\,\mathbb{E}[h^2_{\text{prev}}]$. The gain per layer is $\tfrac{n\sigma^2}{2}$.
- Gain $=1$ needs $\sigma^2 = 2/n$: this is He initialisation. With $\sigma^2=1/n$ the gain is $\tfrac12$, and after 30 layers the signal's second moment is $2^{-30}\approx 10^{-9}$ (checked by simulation: $6.8\times10^{-10}$ for $n=512$).
For $\tanh$ (almost linear near $0$, no cut-off) the gain is $n\sigma^2$, so $\sigma^2=1/n$ works. Xavier (Glorot) uses $\sigma^2 = 2/(n_{\text{in}}+n_{\text{out}})$, which also balances the backward pass and equals $1/n$ for square layers.
For a chain $\mathbf{h}_{\ell+1}=\varphi(W_\ell\mathbf{h}_\ell)$, backprop gives $\;\mathbf{g}_\ell = W_\ell^\top D_\ell\,\mathbf{g}_{\ell+1}$ with $\mathbf{g}_\ell=\partial L/\partial\mathbf{h}_\ell$ and $D_\ell=\mathrm{diag}(\varphi'(\mathbf{z}_\ell))$. The gradient reaching layer 0 is a product of $L$ such matrices. Its size behaves like (gain)$^L$: exponentially small when the gain is below 1.
| Fix | What it changes | Caveat |
|---|---|---|
| ReLU, GELU, leaky ReLU | Slope is 1 (or close) for positive inputs: no saturation | ReLU can "die": a unit that is negative for all inputs has zero gradient forever |
| Xavier / He initialisation | Sets weight variance so the signal and gradient keep their size at the start | Only guarantees the initial state; training can drift away |
| Residual connections $\mathbf{h}_{\ell+1}=\mathbf{h}_\ell+F(\mathbf{h}_\ell)$ | Backward: $\mathbf{g}_\ell=\mathbf{g}_{\ell+1}+J_F^\top\mathbf{g}_{\ell+1}$. The identity term is a direct highway for the gradient | Forward signals now add up across layers, so the branch $F$ is usually scaled down (or normalised, or started at zero) |
| Normalisation (BatchNorm, LayerNorm) | Re-centres and re-scales the pre-activations at each layer, so the initial weight scale hardly matters | Batch statistics depend on batch size; LayerNorm adds cost; placement (pre or post) matters |
| Gated cells (LSTM, GRU) | For sequences: an additive memory path, like a residual connection through time | Still need clipping for exploding gradients |
Why do we need it?
Without these fixes a 50-layer network simply does not train: the first layers receive a gradient near zero (or near infinity) and never move. These tricks are why very deep networks became possible.
Where is it used?
Defaults in every framework: kaiming_normal_ (He) for ReLU layers, xavier_uniform_ for tanh/linear layers, ResNet skip connections, Transformer pre-LayerNorm blocks, and LSTM gates.
How is it used?
Pick ReLU-family activations and the matching initialisation (He for ReLU). Add skip connections and normalisation in deep stacks. Then log the gradient norm per layer: all layers within a few orders of magnitude is healthy.
- Residual connections do not remove the problem by magic. In the widget, plain $\mathbf{h}+F(\mathbf{h})$ with a full-size ReLU branch makes the activations (and with them the gradient) grow exponentially, because every block adds to the signal. Real networks scale the branch down, normalise, or start it at zero.
- He and Xavier initialisation are derived for the first step, assuming independent weights. They do not protect you if the learning rate then pushes the weights somewhere bad.
- Dead ReLU. A ReLU unit whose input is negative for every example has gradient exactly 0 and can never recover. Leaky ReLU or GELU avoid this.
- Vanishing gradients are not only a depth problem: a sigmoid output layer with a squared-error loss also saturates (the slope is near 0 when the prediction is confidently wrong). That is one reason cross-entropy is paired with sigmoid and softmax.
Quick check: 30 ReLU layers, width $n$, weights with variance $1/n$. By what factor does the signal's second moment shrink?
The gain per layer is $n\sigma^2/2 = \tfrac12$, so after 30 layers the second moment is multiplied by $2^{-30}\approx 9\times10^{-10}$ and the size (square root) by about $3\times10^{-5}$. With $\sigma^2=2/n$ (He) the factor is $1$.
Exploding gradients: RNNs and deep nets
Exploding gradients are the mirror image of vanishing ones. If each layer multiplies the backward signal by a bit more than 1, the product grows exponentially. The gradient becomes enormous, one update throws the weights far across the landscape, and the loss jumps to a huge value or to NaN ("not a number").
Recurrent networks (RNNs) are the classic victim. An RNN reuses the same weight matrix $W$ at every time step. Backpropagating through $k$ steps multiplies by $W^\top$ about $k$ times. If $W$ stretches vectors even slightly, $W^k$ stretches them exponentially. A 100-step sequence is a 100-layer network with shared weights.
Very deep feed-forward networks are hit too if the weights are too large (see the "gain 2" run in the previous widget).
The simplest RNN has one number as its state: $h_t = w\,h_{t-1}$. The gradient of $h_T$ with respect to $h_{T-k}$ is $w^k$.
- $w = 1.2$, $k=50$: $1.2^{50}\approx 9.1\times10^{3}$. A gradient of size 1 became about nine thousand.
- $w = 0.8$, $k=50$: $0.8^{50}\approx 1.4\times10^{-5}$. It vanished.
- $w = 1.5$, $k=100$: $1.5^{100}\approx 4\times10^{17}$. Multiplied by a learning rate of $0.01$ this is a step of size $4\times10^{15}$: the weights are destroyed.
With a matrix instead of a number, the role of $w$ is played by the spectral radius $\rho$ (the largest absolute eigenvalue) for long products, and by the largest singular value $\|W\|$ for worst-case bounds.
For the linear recurrence $\mathbf{h}_t = W\mathbf{h}_{t-1}+\mathbf{u}_t$, backpropagating a gradient $\mathbf{g}$ from step $T$ back $k$ steps gives $(W^\top)^k\mathbf{g}$, with $$\|(W^\top)^k\mathbf{g}\| \le \|W\|^k\,\|\mathbf{g}\|,\qquad \text{and for a typical }\mathbf{g}:\ \|(W^\top)^k\mathbf{g}\|\approx \rho^{\,k}\ \text{(up to a constant factor, for large }k\text{)}.$$ So $\|W\|\lt1$ forces the gradient to vanish, and $\rho\gt1$ makes it explode for typical $\mathbf{g}$. For $\mathbf{h}_t=\varphi(W\mathbf{h}_{t-1}+\dots)$ each step multiplies by $W^\top D_t$ with $D_t=\mathrm{diag}(\varphi')$, so the slope of the activation enters. For $\tanh$ (slope $\le 1$) vanishing is guaranteed when $\|W\|\lt1$ and explosion needs $\|W\|\gt1$ (Pascanu et al., 2013).
Remedies. Gradient clipping (next sections), orthogonal or identity-like initialisation of $W$ (all singular values 1, so the product neither grows nor shrinks), gated cells (LSTM, GRU), normalisation, a smaller learning rate, and truncating backprop through time to a fixed window.
Why do we need it?
An exploding gradient does not just slow training: it wrecks it. One bad step can reset hours of progress or turn every weight into NaN. Knowing the cause tells you which cure to apply.
Where is it used?
Training RNNs, LSTMs and GRUs on long sequences, very deep plain networks, Transformers at high learning rates, and reinforcement learning with unbounded rewards. It is the reason almost every RNN recipe includes gradient clipping.
How is it used?
Monitor the gradient norm. If it spikes by orders of magnitude, clip the norm (a threshold near the typical value), lower the learning rate, check the initialisation and add warmup. Do not just ignore the spikes: find the cause.
- Saturation can hide an explosion in a $\tanh$ RNN but does not cure it: the price is a vanishing gradient. A network that explodes in one training phase can vanish in another.
- Real RNN explosions often appear suddenly, when the weights drift across a sharp boundary in weight space (a "cliff", see the clipping section), not smoothly as in the linear picture.
- Exploding gradients are easy to detect (huge gradient norm, loss spikes, NaN) and cheap to cure with clipping. Vanishing ones are quieter and need structural fixes.
Quick check: a linear RNN has $\rho = 1.05$. Roughly how big is the gradient 100 steps back?
$1.05^{100}\approx 131$. Only 5% per step, yet over 100 steps it is a factor of about 130. With $\rho = 1.5$ it would be $\approx 4\times10^{17}$.
Ill-conditioning: narrow valleys and the Hessian spectrum core
Picture a long, narrow canyon. The walls are steep, the floor slopes very gently towards the exit. You stand on one wall. The steepest direction points across the canyon, nearly at the opposite wall, not along the floor. Step downhill and you hop over to the other wall. Step again and you hop back. You make progress along the floor only slowly: a zig-zag.
A loss with such a shape is called ill-conditioned: it curves very strongly in some directions and very weakly in others. The step size must be small enough to be safe on the steep walls, but then the walk along the gentle floor takes forever. This is the most common reason gradient descent is slow in practice.
Take $f(x,y)=\tfrac12\,(x^2+100\,y^2)$. The Hessian is $\mathrm{diag}(1,100)$: curvature $\lambda_{\min}=1$ along $x$ and $\lambda_{\max}=100$ along $y$. Gradient descent: $x\leftarrow(1-\eta)x$ and $y\leftarrow(1-100\eta)y$.
- Stability needs $|1-100\eta|<1$, so $\eta<2/100=0.02$. The steep direction sets the step limit.
- Best single step: balance the two factors, $1-\eta=-(1-100\eta)$, giving $\eta=2/101\approx0.0198$. Then $1-\eta=\tfrac{99}{101}$ and $1-100\eta=-\tfrac{99}{101}$: both coordinates shrink by a factor $\tfrac{99}{101}\approx 0.9802$ per step (the second one flips sign each time: that is the zig-zag).
- To shrink the error by $10^3$ (the loss $f$ by $10^6$): $0.9802^{k}=10^{-3}$, so $k=\ln(10^{3})/(-\ln 0.9802)\approx 345$ steps.
- With a perfectly round bowl ($\lambda_{\max}=\lambda_{\min}$) one step would do. Newton's method (Chapter 3.12) rescales by $H^{-1}$ and also needs one step.
For $f(\mathbf{x})=\tfrac12\mathbf{x}^\top H\mathbf{x}$ with $H$ symmetric positive definite (positive definite matrices) and eigenvalues $\lambda_{\min}\le\dots\le\lambda_{\max}$, gradient descent multiplies the error along eigenvector $i$ by $(1-\eta\lambda_i)$ per step. It converges iff $\eta<2/\lambda_{\max}$, and the best fixed step $\eta=\dfrac{2}{\lambda_{\max}+\lambda_{\min}}$ gives the rate $$\frac{\kappa-1}{\kappa+1}\ \text{per step},\qquad \kappa=\frac{\lambda_{\max}}{\lambda_{\min}}\ \ \text{(the condition number)}.$$ Steps to reduce the error by a factor $\epsilon$: about $\tfrac{\kappa}{2}\ln\frac1\epsilon$ for large $\kappa$. Heavy-ball momentum with tuned parameters improves $\kappa$ to $\sqrt\kappa$: rate $\frac{\sqrt\kappa-1}{\sqrt\kappa+1}$ (Chapter 3.4). The same idea for general smooth $f$: near a minimum, $H=\nabla^2 f$ plays this role (condition number, convergence depends on κ).
Hessian spectra of trained networks (empirical). Measured spectra typically show a large bulk of eigenvalues very close to zero and a handful of large outliers (their number is often close to the number of classes). So $\lambda_{\max}$ is huge and most directions are almost flat: $\kappa$ is enormous. Also empirical: for full-batch gradient descent with a fixed learning rate, the largest eigenvalue tends to grow until it hovers near $2/\eta$ (the "edge of stability", Cohen et al., 2021). With mini-batches the picture is less clean.
Why adaptive methods and normalisation help. Adam, RMSProp and AdaGrad rescale each coordinate by a running estimate of its gradient size. If the ill-conditioning comes from coordinates living on different scales (the valley is aligned with the axes), this removes it almost completely. If the valley is rotated (coordinates are correlated), a per-coordinate rescaling cannot fully straighten it. Normalising inputs and layer activations (BatchNorm, LayerNorm) attacks the same problem at its root, by making the scales alike. These are heuristics supported by experiments, not guarantees.
Why do we need it?
It explains why one model trains in minutes and a seemingly similar one takes days, and why a "bigger learning rate" is often impossible: the steepest direction forbids it. It also tells you what to fix: scales, correlations, or the optimizer.
Where is it used?
Standardising features before training, BatchNorm and LayerNorm, Adam as the default for Transformers, preconditioned methods (natural gradient, Shampoo, K-FAC), and Hessian-spectrum tools such as PyHessian.
How is it used?
Standardise inputs (zero mean, unit variance), use normalisation layers, try Adam or AdamW if SGD crawls, and look at the loss curve: a long, slow, steady decline with a tiny learning-rate ceiling is the signature of ill-conditioning.
- Condition number is a property of the local Hessian; for non-convex losses $\lambda_{\min}$ can be zero or negative, so "$\kappa$" is only a loose picture. People look at $\lambda_{\max}$ and at how many eigenvalues are far above zero.
- "Adam fixes ill-conditioning" is too strong. It fixes the axis-aligned part (scale differences between parameters); correlations between parameters remain. Second-order methods (Chapter 3.12) address both but cost more.
- A tiny step size is not the same thing as an ill-conditioned problem. Check the curvature: a loss can be slow because the learning rate is simply too small.
Quick check: a quadratic has $\lambda_{\max}=50$ and $\lambda_{\min}=0.5$. Largest stable gradient-descent step, and the rate for the best fixed step?
$\kappa=100$. Stability needs $\eta<2/\lambda_{\max}=0.04$. The best fixed step is $2/(50+0.5)\approx0.0396$ with contraction $(\kappa-1)/(\kappa+1)=99/101\approx0.980$ per step, so cutting the error by $10^3$ takes about $345$ steps.
Optimization vs generalization core
Optimization has a narrow job: make the training loss small. But the goal of machine learning is a different number: the test loss, the error on data the model has never seen. A student who memorises last year's exam gets full marks on it and fails the new one. A model with enough capacity can do the same: reach zero training loss by memorising noise.
So "better optimization" is not always better learning. Two surprising facts:
- When many different weight vectors give the same training loss, the optimizer chooses one of them. That choice (its implicit bias) depends on the start, the step size, the noise and when you stop. It is the hidden reason some training runs generalise better than others.
- Stopping early can make the model better. Early stopping means: watch the loss on held-out validation data, and stop when it stops improving, even though the training loss would keep falling.
Implicit bias, with an exact answer. One equation, two unknowns: $x_1+2x_2=5$. There are infinitely many exact solutions (a whole line). Run gradient descent on $L(\mathbf{x})=\tfrac12(\mathbf{a}\cdot\mathbf{x}-5)^2$ with $\mathbf{a}=(1,2)$, starting at $\mathbf{0}$.
- The gradient is $(\mathbf{a}\cdot\mathbf{x}-5)\,\mathbf{a}$: always a multiple of $\mathbf{a}$. So every iterate stays on the line $\mathbf{x}=c\,\mathbf{a}$ through the origin.
- At convergence $\mathbf{a}\cdot\mathbf{x}=5$ too, so $c\,\|\mathbf{a}\|^2=5$, i.e. $c=5/5=1$ and $\mathbf{x}=(1,2)$.
- $(1,2)$ has length $\sqrt5\approx2.24$. Another exact solution, $(5,0)$, has length $5$. Gradient descent from $0$ found the minimum-norm solution (Linear Algebra guide, min-norm solution) without being told to.
Early stopping in the same spirit: for least squares, gradient descent stopped after $t$ steps behaves roughly like ridge regression with a penalty of order $1/(\eta t)$. Fewer steps means a stronger penalty. This is a rule-of-thumb equivalence, exact only in simple settings.
Split data into training, validation and test sets. The generalisation gap is test loss minus training loss. Optimization minimises the training loss; generalisation is how well that carries over.
Early stopping. Track the validation loss $V_t$ at each step $t$. Keep the weights from the step with the lowest $V_t$ (or stop after $V_t$ has not improved for a "patience" of $p$ steps).
Implicit bias. Among all minimisers of the training loss, the algorithm picks one in a way that depends on the algorithm. Proved: gradient descent on least squares started at $\mathbf{0}$ converges to the minimum-norm solution (above); on separable classification with logistic loss it converges in direction to the maximum-margin solution (Soudry et al., 2018). Empirical, partly explained: the noise of SGD (small batches, large learning rates) tends to act like an extra regulariser, and runs with larger noise often generalise a little better. Open: a complete description of the implicit bias of SGD on deep networks.
Flat vs sharp minima (a hypothesis). A minimum is sharp if the loss rises quickly when the weights move a little (large top Hessian eigenvalue), flat if it rises slowly. The idea (Hochreiter and Schmidhuber, 1997; Keskar et al., 2017): the test loss is like the training loss shifted a little, so flat minima lose less and generalise better. Caveats: (1) sharpness changes under harmless reparameterisations that do not change the function at all (Dinh et al., 2017: see the second widget below); (2) empirical studies of the link between measured sharpness and test error give mixed results: clear in some settings, weak in others; (3) methods that explicitly seek flat regions (sharpness-aware minimisation) often help in practice, so it is a useful lens but not a theorem.
Why do we need it?
To stop us from optimising the wrong thing. A lower training loss is only progress if it carries over to new data. This is why every training run has a validation set and every serious paper reports test error.
Where is it used?
Early stopping in scikit-learn, Keras (EarlyStopping) and PyTorch Lightning; checkpoint selection; weight decay and dropout as regularisers; sharpness-aware minimisation (SAM); and the theory of overparameterised models ("double descent").
How is it used?
Plot training and validation loss together. Save the checkpoint with the best validation score. If validation loss rises while training loss falls, the model is overfitting: stop earlier, regularise more, or add data.
- Zero training loss is not the goal. It can be a symptom of memorisation (large networks can fit even random labels).
- Do not tune on the test set. Early stopping and hyper-parameter choices use the validation set. The test set is for one final measurement.
- "SGD finds flat minima, and flat minima generalise" is a hypothesis with exceptions, not a law. Treat it as a useful intuition, and look for evidence in your own validation curves.
- Double descent (empirical): in very over-parameterised models, test error can fall again after the classic overfitting peak. So "bigger always overfits more" is not a safe rule either.
Quick check: training loss $0.001$, validation loss $0.8$. Is more optimization the cure?
No. The optimizer is doing its job; the model has fit the training data (including its noise) far better than it can generalise. Cures are on the generalisation side: early stopping, more data, weight decay or other regularisation, a smaller model, data augmentation.
Batch size and its partnership with the learning rate
Each update uses the average gradient over a mini-batch of $B$ examples (Chapter 3.14). That average is a noisy estimate of the true gradient, and its noise shrinks as $B$ grows (the standard deviation falls like $1/\sqrt B$). Think of asking $B$ people for directions: ask more people and the advice is steadier.
A bigger batch is not simply "better", for two reasons. First, in a fixed number of epochs a bigger batch means fewer updates: with 8192 examples, $B=32$ gives 256 updates per epoch and $B=512$ only 16. Second, the noise itself is partly useful (it shakes the iterate out of bad places). So batch size and learning rate cannot be chosen separately: if you make the batch larger and keep the same learning rate, you take fewer, equally-sized steps and learn less per epoch. The partner rule: when the batch grows, the learning rate usually should grow too.
Why the learning rate should grow with the batch. Compare two ways to use $kB$ examples from the same weights $\mathbf{w}_0$.
- Small batches: $k$ steps of batch $B$ and learning rate $\eta$, with gradients $\mathbf{g}_1,\dots,\mathbf{g}_k$. If the gradient hardly changes over these steps: $\mathbf{w}_k\approx\mathbf{w}_0-\eta\,(\mathbf{g}_1+\dots+\mathbf{g}_k)$.
- One big batch of $kB$ examples: its gradient is the average $\bar{\mathbf{g}}=\tfrac1k(\mathbf{g}_1+\dots+\mathbf{g}_k)$. One step with learning rate $\eta'$ gives $\mathbf{w}_0-\eta'\,\bar{\mathbf{g}}$.
- These agree when $\eta'=k\eta$: multiplying the batch by $k$ should multiply the learning rate by $k$. This is the linear scaling rule (Goyal et al., 2017).
The argument breaks when the gradient does change over the $k$ steps: early in training, or when the step is large compared with the curvature (see the first widget). That is why large-batch recipes add a warmup: start with a small learning rate and ramp it up over the first few hundred or thousand steps.
Linear scaling rule (rule of thumb): batch $B\to kB$, learning rate $\eta\to k\eta$ (for SGD with momentum), with a warmup of $W$ steps. For Adam-type optimizers a square-root rule $\eta\to\sqrt k\,\eta$ has theoretical support in simplified models and is a common starting point. Both are heuristics.
Gradient noise scale (McCandlish et al., 2018): $B_{\text{noise}}\approx\dfrac{\operatorname{tr}\Sigma}{\|\nabla L\|^2}$, with $\Sigma$ the covariance of per-example gradients. Empirical picture: below a critical batch size near $B_{\text{noise}}$, doubling the batch roughly halves the number of steps needed (perfect scaling); above it, returns diminish and extra examples per step are mostly wasted. The critical size grows during training.
Large batches and generalisation. Keskar et al. (2017) reported worse test accuracy for very large batches and linked it to sharp minima. Later work found that much of this gap disappears when the learning rate and the number of training steps are tuned for each batch size (Hoffer et al., 2017; Shallue et al., 2019), though a gap can remain in some problems. Treat "large batch generalises worse" as an empirical tendency with caveats, not a law.
Why do we need it?
Hardware loves big batches (more arithmetic per memory read, easy to spread over many GPUs). But only if training still converges as fast per epoch. The learning rate is what makes that possible.
Where is it used?
Large-scale image training (ImageNet in about an hour), LLM pre-training (batches of millions of tokens with warmup and cosine decay), distributed data-parallel training, and gradient accumulation when memory limits the batch.
How is it used?
Start from a learning rate known to work at a small batch. When you multiply the batch by $k$, multiply the learning rate by $k$ (or $\sqrt k$ for Adam), add warmup, and check the validation curve. Stop scaling once the curve stops improving per epoch.
- The scaling rules are rules of thumb. The linear rule worked well for ResNets on ImageNet up to batches of about 8000 (Goyal et al., 2017); beyond that, extra tricks (for example layer-wise rates such as LARS) are usually needed. Always confirm with a validation curve.
- Warmup is usually needed at large $\eta$. Early gradients change fast, and a big learning rate on freshly initialised weights can diverge.
- BatchNorm couples the batch size to the model. Its statistics are estimated from the batch, so very small batches make it noisy and changing $B$ changes the function being trained.
- Gradient accumulation (summing gradients over several small batches before one update) gives you the large batch's optimisation behaviour with the small batch's memory use: you will meet this in Chapter 3.16.
Quick check: a recipe works at batch 256 with learning rate 0.1 (SGD with momentum). You move to batch 1024. What do you try first?
The linear scaling rule suggests $0.1\times(1024/256)=0.4$, with a warmup over the first few epochs. If the recipe used Adam with learning rate $0.001$, the square-root rule suggests starting near $0.001\times\sqrt4=0.002$. In both cases watch the validation curve, and be ready to back off.
The learning rate: the most important hyperparameter core
The learning rate $\eta$ is the size of each step downhill. It is the single setting that most often decides whether a training run works.
- Too small: progress is painfully slow. The run looks fine but is still far from good when your time budget ends, or it stalls on a flat region.
- Just right: the loss falls quickly and smoothly.
- Too large: each step overshoots the valley floor. The loss bounces, shows sudden spikes, and finally diverges (gets larger and larger, or becomes NaN).
The surprise is how narrow the good range is on the high side. In practice the best learning rate is often just a little below the value at which training becomes unstable, so you cannot find it by guessing a "safe" tiny number.
The learning-rate range test (Smith, 2017) finds that range in one short run.
- Train for a few hundred steps while the learning rate grows exponentially from a tiny value $\eta_{\min}$ (say $10^{-4}$) to a huge one $\eta_{\max}$ (say $100$): $\;\eta_t=\eta_{\min}\,(\eta_{\max}/\eta_{\min})^{t/(T-1)}$.
- Record the loss at every step and plot it against $\log\eta$.
- At the start the loss does not move (too small). Then it falls (good range). Then it flattens, spikes up and explodes (too large).
- Pick a learning rate a bit before the loss starts to go up, typically about a factor of 3 to 10 below the lowest point of the curve, or where the loss falls fastest.
In the widget below, the loss reaches its lowest value near $\eta\approx1$ to $2$ and blows up soon after. A safe choice from the test is about $0.1$ to $0.2$.
Largest stable step. For gradient descent on a smooth loss, steps larger than $2/\lambda_{\max}$ (with $\lambda_{\max}$ the largest Hessian eigenvalue) make the iterates grow in the steepest direction (Chapter 3.5, the previous section on conditioning). Because $\lambda_{\max}$ keeps changing during training, the usable learning rate changes too. Empirical: with a fixed learning rate, training tends to move to the "edge of stability" where $\lambda_{\max}\approx2/\eta$.
Practical recipe. (1) Run a range test to find the usable window. (2) Pick a peak learning rate inside it. (3) Use warmup (ramp up from a small value) and a decay schedule (cosine, step or linear decay) as in Chapter 3.4. (4) Different optimizers need different scales: plain SGD often $10^{-2}$ to $1$, Adam often $10^{-4}$ to $10^{-3}$. These are conventions, not laws.
Why do we need it?
No other single knob moves the result as much. A bad learning rate wastes hours of compute and can make a good model look bad. A 10-minute range test avoids that.
Where is it used?
fastai's lr_find, PyTorch Lightning's learning-rate finder, the one-cycle policy, LLM pre-training recipes (peak learning rate with warmup and cosine decay), and fine-tuning (small rates such as $10^{-5}$ for pre-trained weights).
How is it used?
Run the range test on a short training run, read the curve, choose a rate a little below where the loss is lowest, and then watch the real run: a smooth falling curve is good, repeated spikes mean lower the rate or lengthen warmup.
- The lowest point of the range-test curve is too aggressive. It sits just before the cliff, so pick lower (a factor of 3 to 10).
- The range test depends on the batch size, the initialisation and the optimizer. Redo it if you change them. It also ignores the later stages of training, where a smaller rate is needed: that is what decay schedules are for.
- A decreasing loss is not proof of a good rate. A rate that is 100 times too small also gives a smoothly decreasing curve. Compare against a larger rate before accepting it.
- Learning rate and weight decay interact (next section), and so do learning rate and batch size (previous section).
Quick check: the range test shows the lowest loss at $\eta=0.8$ and a spike at $\eta=3$. What is a reasonable peak learning rate?
Somewhere around $0.1$ to $0.3$ (a factor 3 to 8 below the lowest point). Then add warmup and a decay schedule. If the real run still spikes, lower it further.
Weight decay: L2 penalty vs decoupled decay (AdamW) core
Weight decay means: at every step, pull each weight a little towards zero, as if the weights were slowly leaking. Weights that the data does not keep "topping up" fade away. The surviving weights are the ones that earn their keep. Small weights make a smoother, simpler function, which tends to generalise better (Chapter 3.11 on regularisation).
There are two ways to build the leak:
- L2 penalty: add $\tfrac\lambda2\|\mathbf{w}\|^2$ to the loss. The gradient of the penalty, $\lambda\mathbf{w}$, flows into the optimizer together with the data gradient.
- Decoupled decay: leave the loss alone and shrink the weights directly, $\mathbf{w}\leftarrow(1-\eta\lambda)\mathbf{w}$, as a separate step.
For plain SGD these are the same thing. For Adam they are not, and that difference is the reason AdamW exists.
(a) SGD: the two ways agree. With the L2 penalty, $\mathbf{w}\leftarrow\mathbf{w}-\eta(\nabla L+\lambda\mathbf{w})=(1-\eta\lambda)\,\mathbf{w}-\eta\nabla L$. That is exactly "decay, then take the gradient step".
(b) Adam: they differ. Adam divides the gradient by $\sqrt{\hat v}$, a running average of its own squared size. Put the penalty inside the gradient and it gets divided too. Test: a single weight $w=1$ with no data gradient at all, $\eta=0.01$, $\lambda=0.1$, 50 steps:
- SGD with L2, or AdamW: $w=(1-\eta\lambda)^{50}=(1-0.001)^{50}\approx0.951$.
- Adam with L2: the only gradient is $\lambda w$, and Adam divides it by its own size $|\lambda w|$, so each step moves $w$ by about $\eta=0.01$ towards zero whatever $\lambda$ is. After 50 steps $w\approx0.54$ (we measured $0.537$). The strength $\lambda$ has almost no effect on the speed.
- With a real data gradient of size $g$, the L2 term is divided by $\approx g$: the effective decay of Adam+L2 is about $\eta\lambda/g$ per step. Weights with large gradients are decayed less than $\eta\lambda$, weights with tiny gradients much more.
With data gradient $g_t=\nabla L(\mathbf{w}_t)$ and decay strength $\lambda$:
$$\text{SGD + L2: }\ \mathbf{w}\leftarrow\mathbf{w}-\eta\,(g+\lambda\mathbf{w})\qquad \text{Adam + L2: }\ \mathbf{w}\leftarrow\mathbf{w}-\eta\,\frac{\hat m(g+\lambda\mathbf{w})}{\sqrt{\hat v(g+\lambda\mathbf{w})}+\epsilon}$$ $$\text{AdamW: }\ \mathbf{w}\leftarrow\mathbf{w}-\eta\left(\frac{\hat m(g)}{\sqrt{\hat v(g)}+\epsilon}+\lambda\,\mathbf{w}\right)$$Here $\hat m,\hat v$ are Adam's bias-corrected first and second moment estimates, computed from the quantity in brackets (see Chapter 3.4). In AdamW (Loshchilov and Hutter, 2019) the decay $\eta\lambda\mathbf{w}$ is applied directly to the weights, outside the adaptive scaling, so every weight decays at the same rate $\eta\lambda$.
What decay does in networks. Established: it keeps weights small, biases the solution towards small-norm functions and is a standard regulariser; AdamW is the default for Transformers. Empirical, partly explained: for layers followed by normalisation, the output does not change if you rescale the weights, so decay changes the effective learning rate (roughly $\eta/\|\mathbf{w}\|^2$ for such layers) rather than the function's capacity directly. Decay and learning rate together set an equilibrium weight norm. Practice (convention, not theory): typical values are $\lambda\approx5\times10^{-4}$ for SGD on CNNs and $0.01$ to $0.1$ for AdamW; biases and normalisation gains are usually left undecayed.
Why do we need it?
Networks have far more weights than data points, so many weights are only weakly constrained. Decay stops them from drifting to large values and makes training more predictable. And decoupling it from Adam's scaling makes its strength mean the same thing for every weight.
Where is it used?
Almost every modern training recipe: torch.optim.AdamW, weight_decay=5e-4 in ResNet recipes, LLM pre-training with $\lambda=0.1$, and fine-tuning (decay towards zero or towards the pre-trained weights).
How is it used?
Use AdamW (not Adam with an L2 term) when you want decay with Adam. Start with $0.01$ to $0.1$, exclude biases and norm gains, and tune it together with the learning rate: they interact through the product $\eta\lambda$.
- "L2 regularisation" and "weight decay" are only the same for plain SGD. Many libraries still call the argument
weight_decayeven when it implements an L2 gradient term (Adam). Check which one you have. - Decay does not mean "the weights end at the minimum of loss $+\ \tfrac\lambda2\|\mathbf{w}\|^2$" for AdamW: AdamW is an update rule, not a minimiser of a penalised loss.
- Too much decay underfits (the weights cannot grow enough to fit the data); too little does nothing. And since decay interacts with the learning rate, retune one when you change the other.
- With a decaying learning-rate schedule, the decay $\eta_t\lambda$ shrinks too, which is the usual behaviour and is why decay seems to act mainly early in training.
Quick check: AdamW with $\eta=10^{-3}$ and $\lambda=0.1$. How much does a weight with a zero data gradient shrink after 1000 steps?
Each step multiplies it by $1-\eta\lambda=1-10^{-4}$. After 1000 steps: $(1-10^{-4})^{1000}\approx e^{-0.1}\approx0.905$, about a 10% shrink. (With a decayed learning rate, less.)
Gradient clipping: a speed limit for updates core
Most of the time gradients have a sensible size and training is fine. Occasionally one batch produces a gradient that is hundreds of times larger than usual: you have walked to the edge of a cliff in the landscape (see the exponential wall in the widget below). A normal step along that gradient would hurl the weights far away and wreck the run.
Gradient clipping is a speed limit: if the gradient is longer than a threshold $c$, shorten it to length $c$. If it is shorter, leave it alone. Normal steps are untouched; only the dangerous ones are tamed. There are two ways to shorten:
- Clip by norm: scale the whole vector down. The direction stays exactly the same; only the length changes.
- Clip by value: cut each coordinate separately so it lies in $[-c,c]$. Simple, but the direction can change.
Threshold $c=5$.
- $\mathbf{g}=(30,40)$ has $\|\mathbf{g}\|=\sqrt{900+1600}=50$. Clip by norm: factor $\min(1,5/50)=0.1$, so $\mathbf{g}\leftarrow(3,4)$, length $5$, same direction $(0.6,0.8)$.
- Clip by value: each coordinate to $[-5,5]$ gives $(5,5)$, length $\approx7.07$, direction $(0.707,0.707)$. The direction has rotated (from about $53^\circ$ to $45^\circ$) and the length still exceeds $c$.
- $\mathbf{g}=(3,4)$ has $\|\mathbf{g}\|=5\le c$: clip by norm leaves it exactly as it is (factor $\min(1,1)=1$).
Clip by norm with threshold $c>0$: $$\mathbf{g}\;\leftarrow\;\mathbf{g}\cdot\min\!\left(1,\ \frac{c}{\|\mathbf{g}\|}\right).$$ After clipping, $\|\mathbf{g}\|\le c$ and the direction is unchanged (for $\mathbf{g}\neq\mathbf{0}$).
Clip by value: $g_i\leftarrow\max(-c,\min(c,g_i))$ for each coordinate. The result lies in the box $[-c,c]^n$, but is generally not parallel to $\mathbf{g}$.
Practice. The norm is usually the global norm: all parameter gradients concatenated into one vector (torch.nn.utils.clip_grad_norm_). Clip before the optimizer step (and, with mixed precision, after unscaling the gradients). The threshold is chosen from data: log the gradient norm for a while and set $c$ near a high percentile of its typical values (common starting values: 1.0 for Transformers; 0.25 to 5 for RNNs). Watch the fraction of clipped steps: if it is almost 100%, you are not "clipping outliers", you are just running with a smaller learning rate.
Why it works (theory, with a caveat). Near a cliff the gradient changes very quickly, so the usual assumption of a fixed smoothness constant fails. A more realistic model lets the local smoothness grow with the gradient size, and clipped gradient descent is then provably faster than unclipped (Zhang et al., 2020). Clipping also tames heavy-tailed gradient noise. It changes the update on the clipped steps, so it is a small bias, accepted for stability.
Why do we need it?
One enormous gradient can erase hours of training or turn all weights into NaN. Clipping caps the damage of a single bad batch and lets you train with a more ambitious learning rate.
Where is it used?
RNN and LSTM language models (Pascanu et al., 2013), Transformer and LLM pre-training (global norm 1.0 is typical), reinforcement learning (policy-gradient methods), and GAN training. PPO even clips the policy-ratio in its objective, in a similar spirit: to limit how far one update can move.
How is it used?
Call clip_grad_norm_(model.parameters(), max_norm=1.0) between loss.backward() and optimizer.step(). Log the pre-clip norm. If it spikes by orders of magnitude, you have found the instability.
- Clipping treats the symptom. If gradients explode all the time, find the cause: learning rate too high, bad initialisation, missing normalisation, a bug in the loss.
- Pick $c$ from data. Too small: every step is clipped and you are really training with a tiny, fixed step length. Too large: the clipped steps (at most $\eta c$ long) are still huge, so it does not protect you (see the widget at $\eta=1$, $c=5$).
- Clipping by value distorts the direction. That is acceptable for small, noisy networks, but norm clipping is the default because it preserves where you are heading.
- Mixed precision: unscale the gradients before clipping, otherwise you compare the norm of loss-scaled gradients to a threshold meant for unscaled ones (you will meet loss scaling in Chapter 3.16).
- Clipping cannot repair a forward-pass overflow (for example
expof a huge number): that is a numerical-stability problem (you will meet this in Chapter 3.16).
Quick check: the gradient norm is $0.8$ and the threshold is $c=1$. What does clip-by-norm do?
Nothing: the factor is $\min(1,\,1/0.8)=\min(1,1.25)=1$, so the gradient is unchanged. Clipping only acts when $\|\mathbf{g}\|>c$.
Training diagnosis checklist: from symptom to likely cause core
A training run is like a patient: you cannot see inside, but the loss curve is a vital sign. Different illnesses have different curves. If you learn the few common shapes, you can usually guess the cause in a minute and know what to test first, instead of changing ten settings at once.
The golden rule: change one thing at a time, and start with the cheap tests (below). This section collects what the whole chapter taught into one table.
A first-aid sequence for a run that "does not work":
- Overfit one small batch. Train on 10 to 20 examples only. The loss must go to almost zero. If it cannot, the problem is in the code, data labels or optimizer set-up, not in tuning.
- Print the first loss. For $K$-class classification with random weights it should be close to $\ln K$ ($\ln 10\approx2.30$). A very different value means a bug in the loss or in the data.
- Check the learning rate with a range test (previous sections) and print the value that is actually used, including warmup and schedule.
- Log the gradient norm (total and per layer) and the weight norm. Per-layer norms spanning many orders of magnitude point to vanishing or exploding gradients.
- Log the update-to-weight ratio $\|\Delta\mathbf{w}\|/\|\mathbf{w}\|$ per layer, where $\Delta\mathbf{w}$ is the update actually applied (the learning rate is already inside it). A common rule of thumb: about $10^{-3}$. Much smaller: the layer barely learns. Much larger: it is unstable.
- Plot training and validation loss together. The gap and the trend between them separate optimisation problems from generalisation problems.
Rows are symptoms, columns are likely optimisation causes in rough order of how often they are the culprit (an experienced rule of thumb, not a theorem), and what to try first.
| Symptom | Likely causes | Try first |
|---|---|---|
| Loss is NaN or inf | Learning rate too high; exploding gradients; overflow in exp/log/division; float16 overflow; bad values in the data | Lower the learning rate 10×; add warmup and gradient clipping; use stable loss functions (log-sum-exp, BCE-with-logits); check inputs for NaN; use bfloat16 or loss scaling (you will meet this in Chapter 3.16) |
| Loss does not decrease at all | Learning rate far too small or zero; vanishing gradients; dead activations; frozen parameters; missing zero_grad or detached loss; labels shuffled | Overfit one batch; read per-layer gradient norms; range test; check the optimizer got the parameters; check the schedule value |
| Loss falls, then plateaus high | Learning rate too large to settle (noise floor) or too small; ill-conditioning; saddle or plateau; model too small | Cut the learning rate and see if the loss drops again (then it was noise-limited); use momentum or Adam; train longer; bigger model |
| Loss spikes now and then | Learning rate near the edge of stability; exploding gradients on bad batches; missing warmup; mixed-precision overflow | Clip gradient norm; lower peak learning rate or lengthen warmup; inspect the data of the spiking batches |
| Loss is very noisy and never settles | Batch too small for the learning rate; no learning-rate decay | Decay the learning rate; larger batch or gradient accumulation; average the weights (EMA) |
| Training loss ≪ validation loss | Mostly generalisation, not optimisation: too little data or too big a model; trained too long | Early stopping; weight decay (AdamW); augmentation or dropout; more data; smaller model |
| Both losses high (underfit) | Too few steps; learning rate too small; model too small; too much regularisation; ill-conditioning | Longer training; range test; reduce decay; standardise inputs, add normalisation; try Adam |
| Extremely slow progress | Ill-conditioning (unscaled features); small learning rate; saturated activations | Standardise inputs; normalisation layers; Adam; larger learning rate with warmup |
| Different seeds give very different results | Learning rate too close to instability; small dataset; sensitive initialisation | Lower the rate or lengthen warmup; report mean and spread over seeds; normalisation |
Why do we need it?
Training failures are silent: the program runs and prints a number. A systematic checklist turns hours of random tweaking into a few targeted experiments and catches bugs before they cost a full run.
Where is it used?
Every deep-learning project: debugging a new architecture, scaling a model up, monitoring long runs in Weights and Biases or TensorBoard, and the "recipe for training neural nets" style checklists used by practitioners.
How is it used?
Match your curve to the nearest symptom, run the cheap test for the first cause, and change one setting. Keep a log. Use the picker below to see a typical curve and the matching causes.
- One symptom can have several causes, and several symptoms can share one cause (a too-high learning rate shows up as spikes, noise, plateaus and NaN). The tables give likely causes, ranked by experience, not a proof.
- "Training loss ≪ validation loss" is mostly a generalisation symptom. Do not respond with a bigger learning rate or longer training.
- Always change one thing at a time, and keep the seed fixed while testing a change, otherwise seed-to-seed variation hides the effect.
- Bugs look like optimisation problems. If a simple model on a tiny dataset cannot be overfit, look at the code before the hyperparameters.
Quick check: a 10-class model's very first loss is $5.8$ instead of about $2.3$. Where do you look first?
$\ln10\approx2.30$ is what a model that outputs equal probabilities would score. A first loss of $5.8$ means the initial outputs are confidently wrong (initial weights or the final layer too large), or the loss or labels are set up incorrectly. Check the output layer's initialisation and the loss function before tuning anything else.
Recap, cheat sheet and practice
- Neural-network losses are non-convex: weight products, bending activations and symmetry copies. Some facts are proved (strict saddles are avoided almost surely, linear networks have no bad local minima), many are empirical, and why SGD solutions generalise is still open.
- Loss landscapes are seen through slices. Slices flatter, depend on directions and scale, and can hide barriers or saddles. Straight lines between equivalent solutions show barriers.
- Saddles dominate high-loss critical points in heuristic models; escape takes about $\ln(1/\text{offset})/(\eta|\lambda_{\min}|)$ steps, so flat saddles are slow; noise and momentum help. Local minima are often "good enough" and come in many equivalent copies.
- Vanishing gradients come from saturation and from gains below 1 in the chain of layer products. Fixes: ReLU-type activations, He ($\sigma^2=2/n$) or Xavier initialisation, residual connections (with a small branch), normalisation. Exploding gradients (spectral radius above 1, RNNs) are cured by clipping, orthogonal initialisation, gated cells.
- Ill-conditioning: GD contracts by $(\kappa-1)/(\kappa+1)$ per step; Adam fixes the axis-aligned part; normalisation and standardised inputs reduce $\kappa$.
- Optimization is not generalisation: early stopping, implicit bias (GD from $0$ gives the minimum-norm solution), sharpness is a hypothesis with caveats.
- Batch size and learning rate go together (linear scaling rule $\eta\propto B$, with warmup, up to a limit); the learning rate is the most important knob (range test); weight decay should be decoupled for Adam (AdamW); clipping $\mathbf{g}\leftarrow\mathbf{g}\min(1,c/\|\mathbf{g}\|)$ tames cliffs.
- Use the diagnosis table: overfit one batch, check the first loss, log gradient norms, change one thing at a time.
Cheat sheet
| Topic | Key formula or fact | Status / rule of thumb |
|---|---|---|
| Non-convexity test | $L(\tfrac{p+q}2)>\tfrac12(L(p)+L(q))$ for some $p,q$ | proved test; one pair is enough |
| Symmetry copies | $h!\cdot2^h$ for one layer of $h$ tanh units | exact count of equivalent solutions |
| Saddle escape | growth factor $1+\eta|\lambda_{\min}|$ per step | flat saddle = long plateau |
| He / Xavier init | $\sigma^2=2/n$ (ReLU); $\sigma^2=2/(n_{\rm in}+n_{\rm out})$ | keeps signal size at the start |
| Gradient through a chain | $\approx(\text{gain})^{L}$ | gain below 1 vanishes, above 1 explodes |
| Residual block | $\mathbf{g}_\ell=\mathbf{g}_{\ell+1}+J_F^\top\mathbf{g}_{\ell+1}$ | scale branch by about $1/\sqrt L$ |
| GD on a quadratic | $\eta<2/\lambda_{\max}$; rate $\frac{\kappa-1}{\kappa+1}$ | about $\frac\kappa2\ln\frac1\epsilon$ steps |
| Linear scaling | $B\to kB\Rightarrow\eta\to k\eta$ | rule of thumb; needs warmup; has a limit |
| LR range test | $\eta_t=\eta_{\min}(\eta_{\max}/\eta_{\min})^{t/(T-1)}$ | pick 3 to 10× below the lowest loss |
| Weight decay | SGD: $(1-\eta\lambda)\mathbf{w}-\eta\mathbf{g}$; AdamW decays by $\eta\lambda$ per step | Adam + L2 is not the same as AdamW |
| Clip by norm | $\mathbf{g}\leftarrow\mathbf{g}\cdot\min(1,c/\|\mathbf{g}\|)$ | global norm; $c$ near a high percentile of typical norms |
| First loss check | $\ln K$ for $K$ classes | much larger: bug or bad output init |
import numpy as np
rng = np.random.default_rng(0)
# 1) Signal size through 30 ReLU layers: He (2/n) against 1/n initialisation
n = 512
def forward(var_scale, layers=30):
h = rng.standard_normal(n)
for _ in range(layers):
W = rng.standard_normal((n, n)) * np.sqrt(var_scale / n)
h = np.maximum(W @ h, 0.0)
return np.mean(h ** 2)
print("He (2/n):", round(forward(2.0), 3)) # about 0.7 (order 1: the signal survives)
print("LeCun (1/n):", f"{forward(1.0):.1e}") # about 1e-09 = 2^-30: the signal vanished
# 2) Gradient clipping by norm
def clip_by_norm(g, c):
return g * min(1.0, c / (np.linalg.norm(g) + 1e-12))
print(clip_by_norm(np.array([30.0, 40.0]), 5.0)) # [3. 4.] (norm 50 -> norm 5, same direction)
print(clip_by_norm(np.array([3.0, 4.0]), 5.0)) # [3. 4.] (norm 5 is not above c: unchanged)
# 3) Weight decay: L2 inside Adam against decoupled decay (AdamW), with no data gradient
def adam_step(w, g, m, v, t, lr, b1=0.9, b2=0.999, eps=1e-8):
m = b1 * m + (1 - b1) * g
v = b2 * v + (1 - b2) * g * g
mh, vh = m / (1 - b1 ** t), v / (1 - b2 ** t)
return w - lr * mh / (np.sqrt(vh) + eps), m, v
lr, lam = 0.01, 0.1
w_sgd = w_l2 = w_dec = 1.0
m1 = v1 = m2 = v2 = 0.0
for t in range(1, 51):
w_sgd = w_sgd - lr * lam * w_sgd # SGD + L2 (data gradient = 0)
w_l2, m1, v1 = adam_step(w_l2, lam * w_l2, m1, v1, t, lr) # Adam + L2: penalty enters the gradient
w_dec, m2, v2 = adam_step(w_dec, 0.0, m2, v2, t, lr) # AdamW: Adam step on the data gradient (0) ...
w_dec = w_dec - lr * lam * w_dec # ... plus decoupled decay
print(round(w_sgd, 4), round(w_l2, 4), round(w_dec, 4)) # 0.9512 0.5368 0.9512
# 4) k small steps against one big step (the idea behind linear scaling)
s, k = 0.4, 8
print(round((1 - s / k) ** k, 4), round(1 - s, 4)) # 0.6634 0.6
# 5) Escaping the saddle f = x^2 - y^2 with gradient descent
x, y, eta, steps = 1.0, 1e-6, 0.1, 0
while abs(y) < 1:
x, y = x - eta * 2 * x, y + eta * 2 * y
steps += 1
print(steps) # 76 (theory: ln(1e6)/ln(1.2) = 75.8)
1. Why does swapping two hidden neurons (and their weights) matter for the loss landscape?
2. For a ReLU layer with $n$ inputs and independent zero-mean weights of variance $\sigma^2$, which choice keeps the second moment of the signal constant from layer to layer?
3. Clip by norm with $c=2$ is applied to $\mathbf{g}=(6,8)$. What is the result?
4. What is the difference between Adam with an L2 penalty and AdamW?
5. A recipe uses batch 256 and SGD learning rate $0.1$. You move to batch 1024. The linear scaling rule suggests a learning rate near…
6. Training loss $0.01$, validation loss $0.9$. Which response fits best?
Practice problems
A. For $L(a,b)=(ab-1)^2$, compute the Hessian at $(1,1)$ and at $(2,\tfrac12)$. What do the eigenvalues say about sharpness along the valley?
$H=\begin{bmatrix}2b^2&4ab-2\\4ab-2&2a^2\end{bmatrix}$. At $(1,1)$: $\begin{bmatrix}2&2\\2&2\end{bmatrix}$, eigenvalues $4$ and $0$. At $(2,\tfrac12)$: $\begin{bmatrix}0.5&2\\2&8\end{bmatrix}$, trace $8.5$, determinant $0.5\cdot8-4=0$, eigenvalues $8.5$ and $0$. Both points compute the same function ($ab=1$) but the top eigenvalue (sharpness) is $2(a^2+b^2)$: $4$ against $8.5$. The zero eigenvalue is the flat direction along the valley.
B. Gradient descent with $\eta=0.1$ on $f=x^2-0.05\,y^2$ starts at $(1,10^{-6})$. How many steps until $|y|\ge1$?
$y\leftarrow y-0.1\cdot(-0.1y)=1.01\,y$. We need $1.01^k\ge10^6$: $k=\ln(10^6)/\ln(1.01)=13.8155/0.00995\approx1388.4$, so $1389$ steps (against $76$ for the steeper saddle $x^2-y^2$). Flat saddles are slow.
C. A 20-layer ReLU network has weights of variance $1/n$. By what factor does the signal's second moment, and its size, change?
Gain $\tfrac12$ per layer: second moment $\times2^{-20}\approx9.5\times10^{-7}$; size (square root) $\times2^{-10}\approx9.8\times10^{-4}$, about a thousandth. With He initialisation ($2/n$) the factor is about $1$.
D. The gradient norm is $40$ and you clip at $c=1$ with $\eta=0.1$. How long is the step with and without clipping?
Without clipping: $\eta\|\mathbf{g}\|=0.1\cdot40=4$. With clipping: the factor is $\min(1,1/40)=0.025$, so the clipped gradient has norm $1$ and the step has length $0.1$. The step is $40$ times shorter and points the same way.
E. A model with 51,200 training examples trains well at batch 128 with $\eta=0.05$. You switch to batch 2048. What changes, and what do you try?
Steps per epoch fall from $51200/128=400$ to $51200/2048=25$. The linear scaling rule suggests $\eta=0.05\times16=0.8$ with a warmup of a few epochs. Check the validation curve; if it is worse than at batch 128, lower the learning rate, train for more epochs, or settle for a smaller batch: beyond the critical batch size the extra examples per step are wasted.
F. AdamW with $\eta=3\times10^{-4}$ and $\lambda=0.1$: how much does a weight with zero data gradient shrink over $10^4$ steps?
Each step multiplies by $1-\eta\lambda=1-3\times10^{-5}$. After $10^4$ steps: $(1-3\times10^{-5})^{10^4}\approx e^{-0.3}\approx0.741$, a 26% shrink. (Under a decaying schedule the shrink is smaller.)
Numerical Optimization
On paper an optimizer is a clean formula. On a computer every number is rounded, every function evaluation costs time, and every vector costs memory. This chapter is the engineering side: how rounding can break a loss or a gradient, how to compute things stably, how to check a gradient, what each algorithm costs per step, and where the gigabytes of a training run go.
- Know the floating-point formats used in training and what they can and cannot represent
- Recognise rounding, cancellation, overflow and underflow in losses, softmax and gradients
- Write stable formulations: log-sum-exp, log1p, expm1, losses computed from logits
- Connect the condition number of the Hessian to convergence speed and to the sensitivity of the answer, and see what preconditioning does
- Compute derivatives numerically (forward and central differences), choose the step size, and gradient-check your code
- Compare automatic differentiation in forward and reverse mode, and the memory-for-compute trade (checkpointing)
- Estimate time and memory per iteration for GD, SGD, Newton, BFGS, L-BFGS and conjugate gradient, and the memory of a training run (parameters, gradients, optimizer states, activations, mixed precision)
Builds on. The Linear Algebra guide, Chapter 1.15 introduced floating point, machine epsilon, cancellation, stable softmax, conditioning and iterative solvers; the Calculus guide, Chapter 2.9 introduced backpropagation, automatic differentiation and a first gradient check. Here we keep those parts short and focus on what matters for optimization: the learning-rate update, the loss and its gradient, and the cost of the algorithms. Throughout, $u=\varepsilon/2$ is the unit roundoff (the largest relative rounding error) where $\varepsilon$ is machine epsilon, the gap between $1$ and the next representable number.
Floating-point arithmetic, seen from the optimizer core
A computer stores a real number with a fixed number of bits: a few for the size (the exponent: how big or small) and a few for the digits (the mantissa: how precisely). More exponent bits mean a bigger range. More mantissa bits mean more correct digits. (The full story, with the bit layout, is in the Linear Algebra guide.)
Optimization stresses this in a special way. A training step does $\mathbf{w}\leftarrow\mathbf{w}-\eta\mathbf{g}$: a small number is subtracted from a large one, over and over. If the update is smaller than the gap between neighbouring numbers near $\mathbf{w}$, the answer rounds straight back to $\mathbf{w}$ and the update vanishes. That is why training in low precision keeps a high-precision copy of the weights.
Take a weight $w=1.0$ and an update of size $u_0=\eta g = 10^{-3}\times0.3=3\times10^{-4}$. Which formats notice it? The gap between $1$ and the next number is machine epsilon $\varepsilon$; an update smaller than about $\varepsilon/2$ rounds away.
- float16: $\varepsilon=2^{-10}\approx9.8\times10^{-4}$, half of it $\approx4.9\times10^{-4}$. Since $3\times10^{-4}\lt4.9\times10^{-4}$, $1+3\times10^{-4}$ rounds back to $1.0$: lost.
- bfloat16: $\varepsilon=2^{-7}\approx7.8\times10^{-3}$: lost, and so is any update below $0.0039$.
- float32: $\varepsilon\approx1.2\times10^{-7}$: kept (rounded to the nearest multiple of $1.2\times10^{-7}$).
- But even float32 loses an update of $10^{-8}$ on a weight of $1.0$ ($1.0+10^{-8}=1.0$), while float64 keeps it.
The same size problem appears the other way round: a gradient of size $10^{-8}$ is far below the smallest normal float16 number ($6.1\times10^{-5}$) and, below $3\times10^{-8}$, becomes exactly $0$.
A format has $E$ exponent bits and $M$ mantissa bits. Machine epsilon is $\varepsilon=2^{-M}$, the gap between $1$ and the next number; the gap near a value $x$ is about $\varepsilon|x|$ (so relative precision is constant, absolute precision is not). Rounding to nearest makes every stored number accurate to a relative error at most $u=\varepsilon/2$.
| Format | Exponent / mantissa bits | $\varepsilon$ | Largest finite | Smallest normal | Smallest subnormal |
|---|---|---|---|---|---|
float64 | 11 / 52 | $2.2\times10^{-16}$ | $1.8\times10^{308}$ | $2.2\times10^{-308}$ | $4.9\times10^{-324}$ |
float32 | 8 / 23 | $1.2\times10^{-7}$ | $3.4\times10^{38}$ | $1.2\times10^{-38}$ | $1.4\times10^{-45}$ |
float16 | 5 / 10 | $9.8\times10^{-4}$ | $65\,504$ | $6.1\times10^{-5}$ | $6.0\times10^{-8}$ |
bfloat16 | 8 / 7 | $7.8\times10^{-3}$ | $3.4\times10^{38}$ | $1.2\times10^{-38}$ | $9.2\times10^{-41}$ |
Absorption rule. $\mathrm{fl}(w+\delta)=w$ whenever $|\delta|<\tfrac{\varepsilon}{2}|w|$ (roughly). Range rule. A result above the largest finite number becomes $\infty$; a result below the smallest subnormal becomes $0$; $\infty-\infty$ and $0\cdot\infty$ give NaN. float16 has a small range (max $65\,504$) but more digits than bfloat16; bfloat16 has float32's range but only 2 to 3 digits.
Why do we need it?
Precision decides whether a training step changes the weights at all, and range decides whether gradients and losses survive as numbers. Both are properties of the format, so they are your first suspects when a low-precision run stalls or turns into NaN.
Where is it used?
Mixed-precision training (torch.autocast, torch.cuda.amp), bfloat16 on TPUs and recent GPUs, fp8 research, optimizer states stored in 8 bits, and choosing a tolerance for a stopping rule (it can never go below $\varepsilon$ times the size of the numbers).
How is it used?
Keep weights and optimizer states in float32 (a "master copy"), do the heavy matrix products in float16 or bfloat16, and scale the loss when using float16. Ask, for any small quantity: is it larger than $\varepsilon/2$ times the number it is added to?
- Absolute versus relative. Floating point is accurate in relative terms. The same absolute update is fine on a small weight and invisible on a large one.
- Weight decay and tiny learning rates $\eta\lambda w$ can fall below $\varepsilon w/2$: in low precision the decay then does nothing.
- bfloat16 has a huge range but poor digits: it rarely overflows but accumulates rounding error quickly. Sums over many terms should be accumulated in float32.
- Tolerances. A stopping rule such as $\|\nabla f\|<10^{-12}$ may never fire in float32 whose noise floor is about $10^{-7}$ times the size of the numbers involved.
Quick check: in float32, $w=0.5$ and the update is $2\times10^{-8}$. Is it lost?
The threshold is $\tfrac\varepsilon2|w|=6\times10^{-8}\times0.5=3\times10^{-8}$. The update $2\times10^{-8}$ is below it, so $\mathrm{fl}(0.5+2\times10^{-8})=0.5$: lost.
Numerical errors: rounding, cancellation, overflow and underflow core
Four things go wrong in numerical code, from mild to severe.
- Rounding. Every operation rounds its answer. One rounding is harmless (error about $10^{-7}$ in float32). A sum of a million numbers lets the errors pile up.
- Cancellation. Subtracting two nearly equal numbers cancels the leading digits that were right and leaves only the digits that were already wrong. It is the most common cause of silly answers (such as a negative variance).
- Overflow. A result larger than the format's maximum becomes $\infty$. $\infty-\infty$ or $\infty/\infty$ then gives NaN, which poisons everything it touches.
- Underflow. A result smaller than the smallest representable number becomes exactly $0$. Then $\log 0=-\infty$ and $1/0=\infty$.
In machine learning the dangerous places are losses (logs of probabilities, products of many probabilities), softmax (exponentials) and variance or normalisation (differences of large squares).
Cancellation in a variance. Data $x=\{10001,10002,10003,10004,10005\}$. The true (population) variance is $2$, because the data are $10003+\{-2,-1,0,1,2\}$. Compute it in float32 with the textbook shortcut $\mathrm{Var}=\mathbb{E}[x^2]-(\mathbb{E}[x])^2$.
- $\mathbb{E}[x]=10003$ and $(\mathbb{E}[x])^2=100\,060\,009$.
- $\mathbb{E}[x^2]=100\,060\,011$ (it is $2$ larger).
- Near $10^8$ the gap between float32 numbers is $8$. Both numbers round to multiples of $8$ and the "$2$" difference is smaller than the rounding step. The float32 computation gives $0$.
- With the offset $10^6$ it returns $65\,536$, and with $10^7$ it returns $-8\,388\,608$: a negative variance.
- The two-pass formula $\frac1n\sum(x_i-\bar x)^2$ subtracts first (small numbers) and returns $2$ at every offset up to $10^7$ (above that, float32 cannot even store the five numbers exactly, so no formula can recover the answer).
Rounding model. $\mathrm{fl}(a\circ b)=(a\circ b)(1+\delta)$ with $|\delta|\le u$ for $\circ\in\{+,-,\times,/\}$ (in the absence of overflow or underflow).
Cancellation. If $a$ and $b$ are stored with relative errors up to $u$, then $$\frac{|\mathrm{error\ in\ }(a-b)|}{|a-b|}\;\le\; u\,\frac{|a|+|b|}{|a-b|}.$$ The amplification factor $\frac{|a|+|b|}{|a-b|}$ is huge when $a\approx b$. The subtraction itself is exact; it just exposes the earlier errors. Remedy: reformulate to subtract before the numbers become big (two-pass variance), or use an algebraically equivalent form with no subtraction of close numbers.
Sums. Adding $n$ numbers one by one has worst-case error about $n\,u\sum|x_i|$ (typically about $\sqrt n\,u\sum|x_i|$). Pairwise summation (add in a tree) gives about $\log_2 n\cdot u$, and Kahan summation (carry the lost low-order bits along) gives about $u$ independent of $n$ (see the Linear Algebra guide). NumPy's sum uses pairwise summation.
Overflow and underflow. The largest argument of $\exp$ before overflow is about $88.7$ (float32), $709.8$ (float64), $11.09$ (float16). exp(-104) underflows to $0$ in float32.
Why do we need it?
Optimization code runs for millions of steps, so a tiny per-step error that always points the same way, or one cancellation in a loss, can quietly ruin a model. Knowing the four failure modes lets you spot them in a few seconds.
Where is it used?
Computing means and variances in BatchNorm and LayerNorm, summing losses and gradients over large batches, computing log-likelihoods, softmax in attention, and any "difference of two large numbers" such as a numerical gradient.
How is it used?
Look for subtractions of nearly equal quantities, exponentials of unbounded inputs, logs of things that may be zero, and long sums. Rewrite them (two-pass, shift by the maximum, log1p) or accumulate in a wider format.
- "It works on my small test" is not evidence. Cancellation depends on the size of the numbers: the same formula can be fine at offset $10^2$ and garbage at $10^6$.
- NaN spreads. One NaN in a batch makes the loss NaN, then every gradient NaN, then every weight NaN after the update. Detect it early (check
torch.isfinite(loss)). - Do not compare floating-point numbers with
==after arithmetic. Use a tolerance relative to the size of the numbers. - Underflow is silent. Products of many small probabilities underflow to $0$ without any warning. Work with log-probabilities instead (next section).
Quick check: why does $\sqrt{x+1}-\sqrt{x}$ lose accuracy for large $x$, and what is a stable equivalent?
For large $x$ the two square roots are almost equal, so the subtraction cancels the leading digits. Multiply by the conjugate: $\sqrt{x+1}-\sqrt{x}=\dfrac{1}{\sqrt{x+1}+\sqrt{x}}$, which adds two positive numbers (no cancellation).
Stability: writing the same formula in a safe way core
Two different recipes can compute exactly the same number on paper, yet on a computer one is safe and the other explodes. A recipe is called numerically stable if it does not make rounding errors much bigger than the problem forces. It is a property of the recipe (the algorithm). Compare with conditioning (next section), a property of the problem: a badly conditioned problem is sensitive whatever recipe you use (see condition number and stability in the Linear Algebra guide).
The good news for machine learning is that there is a short list of standard rewrites. Learn them once, and use the library function (logsumexp, log1p, binary_cross_entropy_with_logits) instead of writing the naive formula.
Cross-entropy from logits, in float32. Two classes, logits $\mathbf{z}=(z_1,0)$, and the loss for class 1 is $-\log\frac{e^{z_1}}{e^{z_1}+1}$.
- $z_1=100$: $e^{100}\approx2.7\times10^{43}$ overflows float32 (maximum $3.4\times10^{38}$) to $\infty$, so the naive softmax is $\infty/\infty=$ NaN.
- $z_1=-200$: the probability is about $e^{-200}\approx1.4\times10^{-87}$, which underflows to exactly $0$, so the loss is $-\log0=\infty$. The true loss is $\approx200$.
- The fix. Subtract the largest logit $m$ before exponentiating. Because $\dfrac{e^{z_i}}{\sum_j e^{z_j}}=\dfrac{e^{z_i-m}e^{m}}{e^{m}\sum_j e^{z_j-m}}$, nothing changes mathematically. Now every exponent is at most $0$ (no overflow) and at least one term equals $e^{0}=1$ (the sum is at least $1$, so no division by zero and no $\log0$).
- The loss is then $\mathrm{LSE}(\mathbf{z})-z_y$ with $\mathrm{LSE}(\mathbf{z})=m+\log\sum_je^{z_j-m}$. For $z_1=-200$ it gives $200$, for $z_1=100$ it gives $\approx0$.
Stable building blocks (all are algebraically identical to the naive version):
| Naive | Problem | Stable version |
|---|---|---|
| $e^{z_i}/\sum_je^{z_j}$ | overflow, then NaN | $e^{z_i-m}/\sum_je^{z_j-m}$ with $m=\max_j z_j$ |
| $\log\sum_je^{z_j}$ | overflow or underflow | $m+\log\sum_je^{z_j-m}$ (log-sum-exp) |
| $\log(\mathrm{softmax}(\mathbf{z}))$ | $\log0$ when a probability underflows | $\mathbf{z}-\mathrm{LSE}(\mathbf{z})$ (log-softmax) |
| $\log(1+x)$, small $x$ | cancellation: $1+x$ rounds to $1$ | log1p(x) |
| $e^x-1$, small $x$ | cancellation | expm1(x) |
| $1-\cos x$, small $x$ | cancellation | $2\sin^2(x/2)$ |
| $\log(1+e^{d})$ (softplus), binary cross-entropy | overflow for large $d$ | $\max(d,0)+\mathrm{log1p}(e^{-|d|})$ |
| $\prod_ip_i$ (a likelihood) | underflow to $0$ | $\sum_i\log p_i$ (log-likelihood) |
| $\mathbb{E}[x^2]-\mathbb{E}[x]^2$ | cancellation | two-pass or Welford |
| $\sigma(z)=1/(1+e^{-z})$ | $e^{-z}$ overflows for very negative $z$ | for $z<0$ use $e^{z}/(1+e^{z})$ |
Gradients. The gradient of softmax cross-entropy with respect to the logits is $\mathrm{softmax}(\mathbf{z})-\mathbf{y}$: with every entry between $-1$ and $1$. If instead you differentiate the naive "$-\log p_y$" you meet $1/p_y$, which is huge or infinite when $p_y$ underflows. Fused loss functions (cross_entropy, BCEWithLogitsLoss) implement the stable forms for both value and gradient.
A note on the "just add a small constant" trick: $\log(p+10^{-8})$ avoids $\log0$ but changes the loss and the gradient and can still fail in float16 (where $10^{-8}$ is itself $0$). Prefer the exact stable rewrite.
Why do we need it?
The textbook formula for a loss or softmax is fine on paper and returns NaN or infinity on real logits. A stable rewrite gives the right number, a bounded gradient and a loss you can trust at any scale of logits.
Where is it used?
torch.logsumexp, F.log_softmax, F.cross_entropy (takes logits, not probabilities), BCEWithLogitsLoss, np.log1p, attention softmax in Transformers (subtract the row max), HMM and mixture-model likelihoods in log space.
How is it used?
Pass logits to the loss, never probabilities. When you write your own formula, ask: can the argument of exp be large, can the argument of log be zero, are two near-equal numbers subtracted? Then swap in the matching stable form from the table.
- Stable does not mean exact. It means the error stays near the unavoidable level. A badly conditioned problem stays sensitive whatever you do.
- Check the whole chain. A stable softmax followed by a naive $\log$ is unstable again. Use the fused log-softmax or cross-entropy.
- Precision changes the thresholds, not the idea. In float16, exp overflows already at $11.09$, so even "moderate" logits need the max-shift trick.
- Do not use
np.log(np.exp(x))ornp.exp(np.log(x))as identity: either can overflow or underflow.
Quick check: why is it legal to subtract the maximum logit before the softmax?
Multiplying numerator and denominator by $e^{-m}$ does not change the ratio: $\frac{e^{z_i}}{\sum_je^{z_j}}=\frac{e^{z_i-m}}{\sum_je^{z_j-m}}$. So the probabilities are identical, but now all exponents are $\le0$ (no overflow) and one equals $e^0=1$ (the denominator is at least $1$).
Conditioning: one number, three consequences core
The condition number $\kappa$ answers one question: if the input wobbles a little, how much can the answer wobble? For an optimization problem the "input" is the data (or the numbers in the Hessian) and the "answer" is the solution $\mathbf{x}^\star$. The same number $\kappa=\lambda_{\max}/\lambda_{\min}$ of the Hessian (Calculus guide) shows up in three places:
- Speed. Gradient descent needs about $\kappa$ steps: a long narrow valley is slow (Chapter 3.15, the Linear Algebra guide).
- Sensitivity. A relative change of size $\delta$ in the data can change the solution by up to $\kappa\,\delta$.
- Lost digits. You lose about $\log_{10}\kappa$ of your correct digits: with $\kappa=10^{6}$ in float32 (7 digits) about one digit survives.
The cure is preconditioning: change the coordinates so that the valley becomes round. Then all three problems shrink together.
Sensitivity. Solve $A\mathbf{x}=\mathbf{b}$ with $A=\mathrm{diag}(1,\ 0.01)$ (so $\kappa=100$) and $\mathbf{b}=(1,\ 0.01)$. The solution is $\mathbf{x}=(1,\ 1)$.
- Change $b_2$ from $0.01$ to $0.02$: $\delta\mathbf{b}=(0,0.01)$. Relative size: $\|\delta\mathbf{b}\|/\|\mathbf{b}\|=0.01/1.00005\approx0.0100$, a $1\%$ change.
- New solution: $\mathbf{x}'=(1,\ 2)$, so $\delta\mathbf{x}=(0,1)$. Relative size: $1/\sqrt2\approx0.707$, a $70.7\%$ change.
- Amplification: $0.707/0.0100\approx70.7$. The bound says at most $\kappa=100$. ✓ The error was amplified about seventy times.
Preconditioning. For $f=\tfrac12(x^2+100y^2)$ the Hessian is $\mathrm{diag}(1,100)$. Substitute $y=0.1\,v$: $f=\tfrac12(x^2+v^2)$, a perfectly round bowl with Hessian $I$, $\kappa=1$. Gradient descent now needs one step. We changed coordinates with $P=\mathrm{diag}(1,0.1)$.
For a matrix $A$, $\kappa(A)=\|A\|\,\|A^{-1}\|=\sigma_{\max}/\sigma_{\min}$ (for a symmetric positive definite Hessian: $\lambda_{\max}/\lambda_{\min}$). For the linear system $A\mathbf{x}=\mathbf{b}$, to first order $$\frac{\|\delta\mathbf{x}\|}{\|\mathbf{x}\|}\;\le\;\kappa(A)\,\frac{\|\delta\mathbf{b}\|}{\|\mathbf{b}\|},\qquad\text{digits lost}\approx\log_{10}\kappa(A).$$ For least squares the Hessian is $A^\top A$ and $\kappa(A^\top A)=\kappa(A)^2$: the normal equations square the condition number (the reason for QR; see the Linear Algebra guide).
Preconditioning. Substitute $\mathbf{x}=P\mathbf{y}$ and minimise $\tilde f(\mathbf{y})=f(P\mathbf{y})$. Its Hessian is $P^\top HP$, so we want $PP^\top\approx H^{-1}$. Gradient descent on $\tilde f$, written back in $\mathbf{x}$, is preconditioned gradient descent: $$\mathbf{x}\leftarrow\mathbf{x}-\eta\,M\,\nabla f(\mathbf{x}),\qquad M=PP^\top\ (\text{symmetric positive definite}).$$ $M=I$ is plain gradient descent; $M=H^{-1}$ is Newton's method (Chapter 3.12); a diagonal $M$ rescales each coordinate separately (feature standardisation, AdaGrad, RMSProp and Adam act like adaptive diagonal preconditioners). Natural gradient, K-FAC and Shampoo use richer structured $M$.
Limit of the diagonal. For $H=\begin{bmatrix}50.5&49.5\\49.5&50.5\end{bmatrix}$ (eigenvalues $100$ and $1$, $\kappa=100$) both diagonal entries are equal, so Jacobi scaling does nothing: the ill-conditioning comes from the correlation between coordinates, which a per-coordinate rescale cannot remove.
Why do we need it?
A single number tells you three things at once: how long optimization will take, how much to trust the solution, and how many digits you will lose. And preconditioning is the most effective generic way to make a slow problem fast.
Where is it used?
Standardising features before regression or neural-network training, Jacobi/ILU preconditioners in conjugate gradient and iterative solvers, Adam and RMSProp as diagonal preconditioners, L-BFGS and Newton-CG, natural-gradient and Shampoo-style optimizers.
How is it used?
First look at the scales: standardise inputs, normalise activations. If a valley is still long, try an adaptive method, then a quasi-Newton or Newton-CG method. Estimate $\kappa$ from the Hessian's extreme eigenvalues (power iteration) rather than the full spectrum.
- Conditioning is not stability. A bad $\kappa$ cannot be fixed by a smarter recipe; only by changing the problem (reparametrise, regularise, precondition). A bad recipe on a good problem can be fixed by rewriting it.
- Regularisation improves conditioning. Adding $\lambda\|\mathbf{w}\|^2$ adds $\lambda I$ to the Hessian, so $\kappa=\frac{\lambda_{\max}+\lambda}{\lambda_{\min}+\lambda}$ shrinks (ridge regression, jitter in Gaussian processes).
- Never form $A^\top A$ if you can avoid it: it squares $\kappa$. Use QR or SVD, or conjugate gradient on the original system.
- A preconditioner has a price: building and applying $M$ costs time and memory. The best $M=H^{-1}$ costs as much as solving the problem.
Quick check: a regression design matrix $X$ has $\kappa(X)=10^{4}$. What is $\kappa$ of the Hessian $X^\top X/m$ of the least-squares loss, and how many steps does gradient descent need?
$\kappa(X^\top X)=\kappa(X)^2=10^{8}$. Gradient descent needs on the order of $\tfrac{\kappa}{2}\ln\frac1\epsilon\approx5\times10^{7}\ln\frac1\epsilon$ steps: hopeless. Standardise the features (a diagonal preconditioner), regularise, or use a different method (QR, conjugate gradient, L-BFGS).
Finite differences: the derivative from two function values core
The derivative is the slope of the function at a point. A finite difference estimates that slope by evaluating the function at two nearby points and dividing the rise by the run. You need nothing but the function itself: no formula for the derivative. This is the definition of the derivative with $h$ small but not zero (Calculus guide, Chapter 2.3).
The surprise is the choice of $h$. Two errors fight each other:
- Truncation error (the formula is only approximate) shrinks as $h$ gets smaller.
- Round-off error (the computer rounds each function value) grows as $h$ gets smaller, because we divide a tiny difference of two nearly equal numbers by a tiny $h$: cancellation.
So there is a sweet spot: not too big, not too small. On a log–log plot the total error is a V shape: a falling arm on the right, a rising arm on the left.
Take $f(x)=e^x$ at $x=1$ (the true slope is $e=2.71828\ldots$) and $h=0.1$.
- Forward difference: $\dfrac{f(1.1)-f(1)}{0.1}=\dfrac{3.004166-2.718282}{0.1}=2.85884$. Error $0.1406$.
- Central difference: $\dfrac{f(1.1)-f(0.9)}{0.2}=\dfrac{3.004166-2.459603}{0.2}=2.72281$. Error $0.0045$: about thirty times better with the same $h$ (and two evaluations instead of one).
- Halve $h$ to $0.05$: the forward error roughly halves (it is $O(h)$); the central error drops to about a quarter (it is $O(h^2)$).
- On the computer, with $h=10^{-5}$ the central difference is accurate to about $6\times10^{-11}$ in float64. With $h=10^{-12}$ it is only accurate to about $2\times10^{-4}$: the round-off arm.
For $f(x)=x^2$ at $x=1$ the central difference is exactly $2$ for every $h$ (its third derivative is $0$), while the forward difference is $2+h$.
Forward, central and a fourth-order formula: $$D_+(h)=\frac{f(x+h)-f(x)}{h},\qquad D_0(h)=\frac{f(x+h)-f(x-h)}{2h},\qquad D_4(h)=\frac{-f(x+2h)+8f(x+h)-8f(x-h)+f(x-2h)}{12h}.$$ Truncation error (from the Taylor series, Chapter 2.11): $D_+-f'=\tfrac h2f''+O(h^2)$, $\ D_0-f'=\tfrac{h^2}6f'''+O(h^4)$, $\ D_4-f'=-\tfrac{h^4}{30}f^{(5)}+\dots$ Derivation for $D_0$: $f(x\pm h)=f\pm hf'+\tfrac{h^2}2f''\pm\tfrac{h^3}6f'''+\dots$; subtract, the even terms cancel, and divide by $2h$.
Round-off error. Each stored value has an error up to $u|f|$ ($u=\varepsilon/2$). The difference of two values has error up to $2u|f|$, divided by the denominator: $\approx 2u|f|/h$ (forward), $u|f|/h$ (central). Total (worst case): $$E_+(h)\approx\tfrac h2|f''|+\frac{2u|f|}{h},\qquad E_0(h)\approx\tfrac{h^2}6|f'''|+\frac{u|f|}{h}.$$ Setting $dE/dh=0$ gives the sweet spots $$h_+^\star=2\sqrt{u|f|/|f''|}\ \sim\ \varepsilon^{1/2},\qquad h_0^\star=\Big(\tfrac{3u|f|}{|f'''|}\Big)^{1/3}\ \sim\ \varepsilon^{1/3},\qquad h_4^\star\sim\varepsilon^{1/5},$$ with smallest errors about $\varepsilon^{1/2}$, $\varepsilon^{2/3}$ and $\varepsilon^{4/5}$ (times function-dependent constants). In float64 (measured for $e^x$ at $x=1$): $h\approx10^{-8}$, $5\times10^{-6}$, $10^{-3}$ and best errors about $10^{-8}$, $10^{-11}$, $10^{-13}$. In float32 values the sweet spots move right (roughly $10^{-4}$ for forward, a few $10^{-3}$ for central, $10^{-2}$ to $10^{-1}$ for fourth-order) and the best errors rise to about $10^{-5}$, $10^{-5}$ and $10^{-7}$. Rule of thumb: for the central difference take $h\approx\varepsilon^{1/3}\max(|x|,1)$ (about $6\times10^{-6}$ in double precision). These are guides: the best $h$ depends on the function and the point.
Why do we need it?
Sometimes you have a function but no derivative: a simulator, a black-box model, a piece of code you do not trust. A finite difference is the universal fallback, and the standard way to test an analytic gradient.
Where is it used?
Gradient checking (next sections), derivative-free tuning of hyperparameters and simulators, scipy.optimize (approx_grad and jac='2-point' or '3-point'), Hessian-vector products by differencing gradients, and sensitivity analysis.
How is it used?
Evaluate $f$ at $x\pm h$ with $h\approx\varepsilon^{1/3}\max(|x|,1)$, divide by $2h$. Use float64 for the check. Never use $h$ so small that $x+h$ equals $x$, and never shrink $h$ below the sweet spot hoping for more accuracy.
- Smaller $h$ is not better. Below the sweet spot the answer gets worse, and at $h<u|x|$ the points $x+h$ and $x$ are the same number, so the "derivative" is exactly $0$ or garbage.
- Scale $h$ with $x$. A fixed $h=10^{-5}$ is huge for a weight of size $10^{-8}$ and invisible for one of size $10^{6}$. Use $h_i\approx\varepsilon^{1/3}\max(|x_i|,1)$.
- Non-smooth functions (ReLU at 0, absolute value, max) have no derivative at the kink and finite differences give nonsense there. Avoid the kinks when checking.
- Low precision. The sweet spot in float32 is near $5\times10^{-3}$ with an accuracy of only about 5 digits, and in float16 the method is nearly useless. Always do numerical differentiation in float64.
- Complex-step trick (awareness): for analytic code that accepts complex numbers, $f'(x)\approx\mathrm{Im}\,f(x+ih)/h$ has no subtraction at all, so $h=10^{-20}$ works and the error is about $\varepsilon$.
Quick check: in float64, for the central difference, roughly what is the best $h$ and the best error?
$h\approx\varepsilon^{1/3}\approx6\times10^{-6}$ (so around $10^{-5}$ to $10^{-6}$), and the best error is about $\varepsilon^{2/3}\approx4\times10^{-11}$ for a well-scaled function. We measured $6\times10^{-11}$ at $h=10^{-5}$ for $e^x$ at $x=1$ (the exact value depends on how the rounding happens to fall).
Numerical gradients and why they cost $2n$ evaluations
A finite difference wiggles one input and watches the output. But a gradient has one entry for every input. To get all $n$ entries you must wiggle the inputs one at a time: $n$ wiggles with the forward formula, $2n$ with the central one. Each wiggle is a full evaluation of the function, which for a neural network means a whole forward pass over the data.
Compare with backpropagation (Calculus guide, Chapter 2.9): it returns all $n$ entries for about the cost of three evaluations, however large $n$ is. So numerical gradients are a test tool and a last resort, not a training method.
$f(x,y,z)=x^2y+\sin z$ at $(1,2,0.5)$. The analytic gradient is $(2xy,\ x^2,\ \cos z)=(4,\ 1,\ 0.877583)$. Central differences with $h=10^{-5}$ take $2n=6$ evaluations:
- $\partial f/\partial x\approx\dfrac{f(1.00001,2,0.5)-f(0.99999,2,0.5)}{2\cdot10^{-5}}=4.000000000004$
- $\partial f/\partial y\approx\dfrac{f(1,2.00001,0.5)-f(1,1.99999,0.5)}{2\cdot10^{-5}}=1.000000000007$
- $\partial f/\partial z\approx0.877582561887$ (true $0.877582561890$).
Errors about $4\times10^{-12}$: excellent. But now scale up: a network with $n=10^6$ weights, where one forward pass takes $10$ ms. A central-difference gradient needs $2\times10^6$ passes: $2\times10^{6}\times0.01\text{ s}=20\,000$ s, about $5.6$ hours. Backprop needs about $3$ passes: $0.03$ s.
Numerical gradient of $f:\mathbb{R}^n\to\mathbb{R}$: $$\frac{\partial f}{\partial x_i}\approx\frac{f(\mathbf{x}+h_i\mathbf{e}_i)-f(\mathbf{x}-h_i\mathbf{e}_i)}{2h_i},\qquad h_i\approx\varepsilon^{1/3}\max(|x_i|,1).$$ Cost: $2n$ evaluations of $f$ (central) or $n+1$ (forward). Memory: just a copy of $\mathbf{x}$ (cheap!).
One direction, two evaluations. For any direction $\mathbf{v}$, $\ \dfrac{f(\mathbf{x}+h\mathbf{v})-f(\mathbf{x}-h\mathbf{v})}{2h}\approx\nabla f(\mathbf{x})\cdot\mathbf{v}$. Two evaluations check a whole gradient against one random direction: compare it with $\mathbf{g}\cdot\mathbf{v}$ from your analytic gradient. This is how large models are gradient-checked.
Hessian-vector products. The same idea applied to gradients, $\ H\mathbf{v}\approx\dfrac{\nabla f(\mathbf{x}+h\mathbf{v})-\nabla f(\mathbf{x}-h\mathbf{v})}{2h}$, costs two gradient evaluations and never forms the $n\times n$ matrix $H$. Methods such as Newton-CG ("Hessian-free" optimization) rely on this (Chapter 3.12; you will meet conjugate gradient in Chapter 3.17).
Why do we need it?
It needs no derivative code, so it works for any function you can evaluate, and it is the independent check that your analytic gradient is right. Its cost explains why backpropagation was such a breakthrough.
Where is it used?
Testing custom layers and losses (torch.autograd.gradcheck), derivative-free optimization with a handful of parameters (Nelder-Mead, finite-difference BFGS in scipy.optimize.minimize), simulators and black-box objectives, Hessian-free training.
How is it used?
For a small model, perturb each parameter by $\pm h_i$ and recompute the loss. For a big model, pick a random direction $\mathbf{v}$ (or a few random coordinates). Use float64 and $h\approx10^{-5}$ to $10^{-6}$.
- Fixed $h$ for all parameters is a trap. Parameters live on different scales. Use $h_i\propto\max(|x_i|,1)$, or check in a random direction (scale-free).
- The function must be deterministic. Dropout, random augmentation or noisy simulators make $f(\mathbf{x}+h\mathbf{e}_i)$ and $f(\mathbf{x}-h\mathbf{e}_i)$ differ for reasons unrelated to $h$. Fix the seed or switch the randomness off.
- $2n$ is the central rule. A one-sided difference needs $n+1$ evaluations (the base point is shared) but is only first-order accurate.
- Backprop's cost is independent of $n$ in evaluations; its price is memory (it stores the forward values), as we see in the automatic-differentiation section.
Quick check: a model has $n=50\,000$ parameters and a forward pass takes $4$ ms. How long does one central-difference gradient take?
$2n=100\,000$ passes $\times\,0.004$ s $=400$ s, about $6.7$ minutes. Backprop needs about $3\times4$ ms $=12$ ms, which is roughly $33\,000$ times faster.
Gradient checking: thresholds, bug signatures and the random-direction trick core
When you write the gradient of a loss yourself (a custom layer, a new loss, a hand-derived formula), the most dangerous bug is the silent one: the code runs, the loss even goes down a little, and the model is quietly worse. A gradient check catches it by asking the function itself: "nudge each weight a little and tell me how the loss changes". If that disagrees with your formula, one of them is wrong, and it is almost always your formula.
The Calculus guide showed the idea on a tiny network (gradient checking). Here we add what you need to do it for real: how to measure disagreement, what threshold to use, how to read the pattern of failures to find the bug, and how to check a huge model with only two function evaluations.
A softmax classifier with 3 classes, 2 features and 4 training examples (9 parameters: a $3\times2$ weight matrix and 3 biases). The correct gradient at our test point is $(-0.1261,\,-0.0748,\,-0.3635,\,0.1020,\,0.4895,\,-0.0272\ \mid\ -0.1715,\,0.0923,\,0.0792)$.
- Correct code: the largest relative error against central differences ($h=10^{-5}$, float64) is about $10^{-10}$.
- Forget to divide by the number of examples $N=4$: every analytic entry is $4$ times too big, so $\dfrac{|4g-g|}{\max(|4g|,|g|)}=\dfrac34=0.75$ for every entry. A constant ratio is the fingerprint of a missing factor.
- Forget the bias gradient: only the three bias entries fail (analytic $0$, relative error $1$); the six weights pass. The pattern points straight at the bias path.
- Flip the sign: relative error $2$ everywhere (analytic $=-$numeric).
The widget below lets you inject these bugs and read the pattern yourself.
For each parameter compare the analytic value $a_i$ with the numerical one $n_i$ using the relative error $$\mathrm{rel}_i=\frac{|a_i-n_i|}{\max\big(|a_i|,\ |n_i|,\ \tau\big)},\qquad \tau\approx10^{-8}\ (\text{avoids }0/0).$$ (Dividing by $|a_i|+|n_i|$, as in the Calculus guide, differs by at most a factor $2$; just be consistent.) An absolute-error test alone is wrong: a gradient entry of size $10^{-6}$ with error $10^{-6}$ is $100\%$ wrong.
| Largest $\mathrm{rel}_i$ (float64, central, $h\approx10^{-5}$) | Reading |
|---|---|
| below $10^{-7}$ | excellent |
| $10^{-7}$ to $10^{-5}$ | fine, especially for models with non-smooth parts (max, ReLU kinks nearby) |
| $10^{-5}$ to $10^{-3}$ | suspicious: check $h$, look for kinks, recheck in float64 |
| above $10^{-3}$ | almost certainly a bug (real bugs usually give $0.1$ to $2$) |
If the loss is computed in float32, the same code shows relative errors from about $10^{-3}$ up to $10^{-1}$ even when it is correct (the smaller $h$, the worse), because the round-off error in the numerical derivative is about $u|L|/h$ (see the finite-difference section). Do gradient checks in float64.
Random-direction check. Draw a random unit vector $\mathbf{v}$ and compare the two numbers $\mathbf{g}_{\text{analytic}}\cdot\mathbf{v}$ and $\dfrac{L(\theta+h\mathbf{v})-L(\theta-h\mathbf{v})}{2h}$. Two loss evaluations check the whole gradient at once (a bug in any component shows up with probability one). Repeat for a few directions. This is the only practical check for models with millions of parameters.
| Bug | Signature in the check |
|---|---|
| Missing or extra constant factor ($1/N$, $2$) | All entries off by the same ratio |
| Sign error | Relative error $\approx2$ everywhere |
| A forgotten path or term (bias, skip connection) | Those parameters fail with analytic value $0$ (relative error $1$); others pass |
| Wrong index, transpose or broadcast | Many entries wrong with no clear pattern |
Overwrite instead of accumulate (= for +=) | Parameters used by several examples or time steps are wrong |
| Accidental detach or stop-gradient | Zero gradient for everything upstream |
| Kink (ReLU at $0$, $|x|$, max) | One isolated entry fails and changes if you move the point |
| Randomness (dropout, BatchNorm in training mode) | Erratic, different on each run |
Why do we need it?
A wrong gradient does not crash anything. It silently trains a worse model, and you may blame the data or the learning rate for weeks. A five-minute gradient check finds it.
Where is it used?
Unit tests for custom autograd functions (torch.autograd.gradcheck, gradgradcheck), JAX's check_grads, new layers and losses in research code, and numerical libraries before release.
How is it used?
Use a tiny model and a few examples, in float64, with randomness off and away from kinks. Compare analytic and numeric gradients entry by entry, then read the failures against the table. For a large model, use the random-direction check.
- Passing is strong evidence, not proof. Check at a few different points (random weights, several mini-batches). A bug that only matters in rare branches can survive one check.
- Check each component separately in a big model (per layer, per loss term). Then a failure tells you where to look.
- Switch randomness off: fix seeds, put dropout and BatchNorm in evaluation mode, and keep the data order fixed. Otherwise $f(\mathbf{x}+h)$ and $f(\mathbf{x}-h)$ differ for the wrong reason.
- Never use gradient checking during real training: it needs $2n$ evaluations (or at least two with a random direction). Do it once, on a small model, while developing, then turn it off.
- The check validates the gradient of the function as coded. It cannot tell you that the function is the loss you intended.
Quick check: a check shows relative error $\approx1$ for the three bias parameters and $10^{-10}$ for all weights. What do you suspect?
The bias gradient is missing or computed as zero (relative error $1$ means analytic $\approx0$ while numeric is not). The weight path is fine. Look at the backward code for the bias term (for example a forgotten sum over the batch).
Automatic differentiation: forward mode, reverse mode, and the memory bill core
Automatic differentiation (AD) is neither a finite difference (no step $h$, no truncation error) nor symbolic algebra (no giant formula). It applies the chain rule to the program, one tiny operation at a time, and gives the exact derivative of what the code computes (up to rounding). The details, with computational graphs and dual numbers, are in the Calculus guide, Chapter 2.9 (see forward mode and dual numbers and cost and memory). Here is the one fact an optimizer user needs.
The chain rule can be swept through the program in two orders:
- Forward mode: carry a "tangent" next to every value, from the inputs to the output. One sweep gives the derivative of everything with respect to one chosen input.
- Reverse mode (backpropagation): first run the program forward and remember the values; then sweep backwards from the output carrying an "adjoint". One sweep gives the derivative of one output with respect to every input.
A training loss has millions of inputs (the weights) and one output. Reverse mode wins by a factor of the number of weights. The price is memory: it must remember the forward values. Recomputing some of them instead (checkpointing) trades memory for time.
$f(x,y)=xy+\sin x$ at $(x,y)=(2,3)$. Break it into operations: $v_1=x=2$, $v_2=y=3$, $v_3=v_1v_2=6$, $v_4=\sin v_1=0.909297$, $v_5=v_3+v_4=6.909297$.
- Forward mode, seed $\dot x=1,\dot y=0$: $\dot v_1=1$, $\dot v_2=0$, $\dot v_3=\dot v_1v_2+v_1\dot v_2=3$, $\dot v_4=\cos(v_1)\dot v_1=-0.416147$, $\dot v_5=\dot v_3+\dot v_4=\mathbf{2.583853}=\partial f/\partial x$.
- To get $\partial f/\partial y$ forward mode needs a second sweep with seed $\dot y=1$: $\dot v_3=2$, $\dot v_4=0$, $\dot v_5=\mathbf{2}$.
- Reverse mode: after the forward sweep, set $\bar v_5=1$ and go back. $\bar v_4=1$, $\bar v_3=1$, $\bar v_2=\bar v_3v_1=\mathbf{2}$, $\bar v_1=\bar v_3v_2+\bar v_4\cos v_1=3-0.416147=\mathbf{2.583853}$. Both partials from one backward sweep.
- Check: $\partial f/\partial x=y+\cos x=3+\cos2=2.583853$ ✓ and $\partial f/\partial y=x=2$ ✓.
With $n=2$ the saving is small. With $n=10^8$ weights it is the difference between $10^8$ sweeps and one.
Let the program compute $F:\mathbb{R}^n\to\mathbb{R}^m$ with Jacobian $J\in\mathbb{R}^{m\times n}$ (Calculus guide, Chapter 2.5).
- Forward mode computes a Jacobian-vector product $J\mathbf{u}$ (one column if $\mathbf{u}=\mathbf{e}_i$) in one sweep, with cost about $2$ to $3$ times the cost of $F$ and almost no extra memory. The full Jacobian needs $n$ sweeps.
- Reverse mode computes a vector-Jacobian product $\mathbf{w}^\top J$ (one row if $\mathbf{w}=\mathbf{e}_j$) in one backward sweep, with cost about $2$ to $4$ times the cost of $F$ (the cheap gradient principle, a rule of thumb that holds for typical programs), plus memory for the forward values. The full Jacobian needs $m$ sweeps. A scalar loss ($m=1$): the whole gradient in one sweep, independent of $n$.
- Forward-over-reverse gives a Hessian-vector product $H\mathbf{v}$ at a small multiple of the gradient's cost, without forming $H$ (see Chapter 3.12).
Checkpointing (memory versus recomputation). A network of $L$ equal layers stores $L$ layers of activations for the backward sweep. Store only every $k$-th layer (a checkpoint) and recompute the layers in between when the backward sweep reaches them: memory $\approx L/k+k$ layers' worth, extra time about one additional forward pass. The best choice $k=\sqrt L$ gives about $2\sqrt L$ layers of memory (for $L=100$: $20$ instead of $100$).
Why do we need it?
It makes training possible: exact gradients of a function with millions of inputs for a few times the cost of evaluating the function. And knowing the memory bill explains why big models need so much GPU memory and what to do about it.
Where is it used?
PyTorch autograd, JAX (grad, jvp, vjp, hessian), TensorFlow, gradient checkpointing (torch.utils.checkpoint, jax.checkpoint), Hessian-free and implicit-differentiation methods, and scientific computing with few inputs and many outputs (forward mode).
How is it used?
Call loss.backward() or jax.grad for gradients (reverse). Use jvp for directional derivatives and few-input problems (forward). If memory runs out, wrap blocks in checkpoint and pay about 33% more time.
- AD gives the derivative of the code, not of the maths you meant. If the code uses a numerically unstable formula, its exact derivative is unstable too (the stability section). If the code contains a non-differentiable step (
round,argmax) the gradient through it is zero or a convention. - "2 to 4 times" is a rule of thumb, not a theorem for every program. Memory traffic, parallelism and the exact operations change the constant.
- Reverse mode memory scales with the depth of the computation, and with batch size. That, not arithmetic, is usually what limits the size of a trainable model on one GPU.
- Checkpointing is not free: about one extra forward pass, and random operations (dropout) must reuse the same random numbers on recomputation (frameworks handle this by saving the random state).
Quick check: a function has $n=5$ inputs and $m=1000$ outputs. Which mode computes its Jacobian more cheaply, and about how many sweeps?
Forward mode: $n=5$ sweeps (each gives one column of the $1000\times5$ Jacobian). Reverse mode would need $m=1000$ sweeps (each gives one row). With few inputs and many outputs, forward mode wins.
Computational complexity: what one iteration costs core
A good optimizer needs few iterations. A cheap optimizer has a small cost per iteration. You rarely get both. Newton's method takes a handful of enormous steps; stochastic gradient descent takes a million tiny cheap ones. To compare methods fairly you need two numbers: the cost of one iteration, and how many iterations it needs. The total is their product.
This section is about the first number, counted in "how does the cost grow with the size of the problem?" (big-$O$ language). Three sizes matter: the number of parameters $n$, the number of training examples $m$, and for stochastic methods the batch size $b$.
A logistic-regression-like model with $n=1000$ parameters on $m=100\,000$ examples. One pass over the data (computing the loss and the gradient) costs about $m\cdot n=10^{8}$ multiply-adds.
- Gradient descent: one full gradient, about $mn=10^8$ operations. Memory: the weights and the gradient, $2n=2000$ numbers.
- Mini-batch SGD with $b=100$: only $bn=10^5$ operations, a thousand times cheaper per step (and noisier).
- Newton: form the Hessian ($mn^2=10^{11}$ operations) and solve a linear system ($n^3=10^9$). About $10^{11}$ per iteration, a thousand times more than GD, and it stores an $n\times n$ matrix: $10^6$ numbers. But it needs only a few iterations.
- L-BFGS with memory $k=10$: gradient plus about $4kn=4\times10^4$ extra operations, and $2kn=2\times10^4$ stored numbers. It is almost as cheap per step as GD and much better at using each gradient.
Now take $n=10^6$ parameters. Newton's Hessian would have $10^{12}$ entries ($4$ TB in float32): impossible. Even BFGS's $n^2$ matrix is out of reach. That is why deep learning uses first-order methods, and L-BFGS only on smaller problems.
Let $n$ = number of parameters (for a linear or logistic model with $d$ features, $n=d$), $m$ = number of examples, $b$ = batch size, $k$ = L-BFGS history length. One full gradient costs about $mn$ (each example touches each weight a constant number of times; backprop adds a factor of about $3$, which we drop from big-$O$). Rough orders of magnitude, not exact operation counts:
| Method | Work per iteration | Memory (numbers stored) | Iterations needed (typical behaviour) |
|---|---|---|---|
| Gradient descent | $O(mn)$ | $O(n)$ | many: about $\kappa\ln\frac1\epsilon$, linear rate |
| SGD, mini-batch | $O(bn)$ | $O(n)$ | very many, noisy, sublinear in theory |
| Newton | $O(mn^2+n^3)$ | $O(n^2)$ | few: quadratic convergence near the minimum |
| BFGS | $O(mn+n^2)$ | $O(n^2)$ | fewer than GD: superlinear |
| L-BFGS | $O(mn+kn)$ | $O(kn)$ | between GD and BFGS |
| Conjugate gradient (nonlinear) | $O(mn+n)$ | $O(n)$ | on a quadratic: at most $n$ steps in exact arithmetic; often far fewer |
Where the terms come from. Newton: the Hessian of a model with $m$ examples costs $O(mn^2)$ to form (for a linear model, $X^\top DX$), and a Cholesky solve costs $O(n^3)$. BFGS keeps a dense $n\times n$ inverse-Hessian estimate: a rank-two update and a matrix-vector product are both $O(n^2)$. L-BFGS keeps the last $k$ pairs $(\mathbf{s}_i,\mathbf{y}_i)$ and its "two-loop recursion" does about $4kn$ operations (Chapter 3.12). Nonlinear CG does a few vector operations ($O(n)$) on top of the gradient and a line search. For linear CG on $A\mathbf{x}=\mathbf{b}$ one iteration is one matrix-vector product, $O(\text{nonzeros of }A)$ (you will meet conjugate gradient in Chapter 3.17).
Total cost = (cost per iteration) × (iterations). Convergence rates are in Chapter 3.5; stochastic methods in Chapter 3.14.
Why do we need it?
To choose a method before running it: will it finish in a day, and will it fit in memory? A back-of-the-envelope estimate rules out Newton for a million parameters in seconds.
Where is it used?
Choosing between SGD/Adam (deep learning), L-BFGS (classical models with up to millions of parameters, scipy.optimize.minimize(method='L-BFGS-B'), scikit-learn's logistic regression), Newton and IRLS (generalised linear models with few features), and CG (large sparse systems).
How is it used?
Write down $n$, $m$ and the memory you have. Compute cost per iteration for each candidate (the calculator below does it), multiply by a realistic iteration count, and pick the method that fits. First order for huge $n$, L-BFGS for medium, Newton for small $n$ and high accuracy.
- Per-iteration cost is only half the story. Newton wins on iterations, SGD wins on cost per step, and the best method for a given $n$ and accuracy is a trade-off.
- The orders hide constants. The constant for backprop (about 3), memory bandwidth, and parallelism change real timings. Treat these as ranking tools, not predictions.
- Sparse problems change everything. If $A$ or the Hessian is sparse, matrix-vector products cost the number of nonzeros, not $n^2$, and CG or L-BFGS can handle millions of variables.
- Full-batch costs scale with $m$. With $m=10^9$ examples a single GD iteration is already enormous; that is why stochastic methods dominate large-scale learning.
- Newton-CG avoids the $n^2$ Hessian by using Hessian-vector products (about the cost of a gradient each), so its memory is $O(n)$ even though it follows Newton directions.
Quick check: $n=10^5$ parameters. Why can you not store the BFGS matrix in float32, and what do you use instead?
It has $n^2=10^{10}$ entries, which is $40$ GB in float32 (more than most GPUs hold, and one multiply with it costs $10^{10}$ operations). L-BFGS stores only $k\approx10$ pairs of vectors: $2kn=2\times10^6$ numbers, $8$ MB.
Memory complexity: where the gigabytes of a training run go core
A trained model is just its weights. A model being trained needs much more room. Think of building a house: besides the finished house (the weights) you need the scaffolding (gradients), the builder's notebook (optimizer states: running averages), and the drawings of every floor kept until the end (activations, saved for backpropagation).
- Weights: one number per parameter.
- Gradients: one more number per parameter.
- Optimizer state: none for plain SGD, one extra number per parameter for momentum, two for Adam (the running means $\mathbf{m}$ and $\mathbf{v}$).
- Activations: every layer's output for every example in the batch, kept for the backward pass. This grows with batch size and depth, not with the number of parameters.
In units of "one copy of the parameters": weights plus optimizer state is 1× for SGD, 2× for momentum, 3× for Adam; add the gradients and a training run holds 2×, 3× and 4× the parameter count, before activations.
A 7-billion-parameter model ($P=7\times10^9$) trained with Adam in float32.
- Weights: $4$ bytes $\times P=28$ GB. Gradients: $28$ GB. Adam's $\mathbf{m}$ and $\mathbf{v}$: $2\times28=56$ GB.
- Total: $16$ bytes per parameter $=112$ GB, before a single activation. An 80 GB GPU cannot even hold the states.
- Mixed precision (16-bit weights and gradients for the forward and backward passes, plus a float32 "master" copy and float32 Adam states): $2+2+4+4+4=16$ bytes per parameter. The same $112$ GB! Mixed precision speeds up the arithmetic and halves the activation memory, but it does not shrink the optimizer memory.
- What does shrink it: sharding states across GPUs (ZeRO, FSDP), 8-bit optimizer states, or optimizers with less state. Inference needs only $2$ bytes per parameter ($14$ GB for 7B).
Bytes per parameter for training (activations not included):
| Scheme | Weights | Gradients | Master copy | Optimizer state | Total |
|---|---|---|---|---|---|
| float32, SGD | 4 | 4 | — | 0 | 8 |
| float32, momentum | 4 | 4 | — | 4 | 12 |
| float32, Adam | 4 | 4 | — | 8 | 16 |
| mixed (16-bit + float32 master), Adam | 2 | 2 | 4 | 8 | 16 |
| pure bfloat16, Adam (states in bf16: risky, absorption!) | 2 | 2 | — | 4 | 8 |
Activations $\approx$ (batch size) $\times$ (number of stored activation values per example) $\times$ (bytes each). Checkpointing cuts this from $L$ to about $2\sqrt L$ layers' worth, at the cost of one extra forward pass (the autodiff section).
Loss scaling (float16). Gradients are often tiny ($10^{-8}$ to $10^{-4}$) while float16 cannot represent anything below $6\times10^{-8}$ and loses precision below $6.1\times10^{-5}$ (subnormals). Multiply the loss by a scale $S$ before the backward pass (by the chain rule every gradient is multiplied by $S$), do the backward pass, then divide the gradients by $S$ in float32 before clipping and updating. Dynamic loss scaling picks $S$ automatically: start high (PyTorch's default is $2^{16}$), halve $S$ and skip the step whenever a gradient is $\infty$ or NaN, and double $S$ after $2000$ good steps. bfloat16 has float32's range, so it usually needs no loss scaling.
Gradient accumulation. To train with batch $B=K\,b$ on a device that fits only $b$ examples, run $K$ forward-backward passes of size $b$, average (or sum with a $1/K$ factor) the gradients, then take one optimizer step. For a mean loss over independent examples the accumulated gradient equals the big-batch gradient exactly. Memory is that of the batch $b$. Caveats: BatchNorm statistics are computed per micro-batch, and each optimizer step takes $K$ times longer.
Why do we need it?
Memory, not arithmetic, is usually the limit on how big a model or batch fits on one device. A quick estimate tells you before you launch a job whether it will crash with "out of memory", and which remedy to pick.
Where is it used?
Planning LLM fine-tuning and pre-training (the "16 bytes per parameter" rule), choosing mixed precision (torch.autocast plus GradScaler), ZeRO and FSDP sharding, 8-bit optimizers (bitsandbytes), gradient checkpointing and gradient accumulation.
How is it used?
Multiply parameters by bytes per parameter for the static part; add activations (measure one micro-batch). If it does not fit: lower the micro-batch and accumulate, checkpoint, shard the states, or use lower-precision states.
- Mixed precision does not reduce optimizer memory. With a float32 master copy and Adam it is still 16 bytes per parameter. Its savings are in activations and in speed.
- Activations scale with the batch, parameters do not. A small model with a huge batch can run out of memory on activations; a huge model with batch 1 runs out on states.
- Gradient accumulation changes BatchNorm (statistics per micro-batch) and the schedule (count optimizer steps, not micro-steps), and it makes each step $K$ times slower.
- Unscale before clipping. Clipping scaled gradients against a threshold meant for true gradients clips almost everything.
- Numbers here are estimates. Real runs add framework overhead, fragmentation, temporary buffers and communication buffers.
Quick check: a 1.3-billion-parameter model with Adam in float32 (16 bytes per parameter). How much memory for weights, gradients and Adam states?
$1.3\times10^9\times16=20.8\times10^9$ bytes $=20.8$ GB (weights $5.2$, gradients $5.2$, Adam states $10.4$ GB), before activations. It fits on a 24 GB GPU only with a small micro-batch.
Recap, cheat sheet and practice
- Floating point: relative precision $\varepsilon=2^{-M}$ (float64 $2\times10^{-16}$, float32 $1.2\times10^{-7}$, float16 $10^{-3}$, bfloat16 $8\times10^{-3}$) and a finite range (float16 max $65\,504$). An update smaller than about $\tfrac\varepsilon2|w|$ is absorbed, so low-precision training keeps a float32 master copy.
- Errors: rounding, cancellation (subtracting nearly equal numbers amplifies earlier errors by $\frac{|a|+|b|}{|a-b|}$), overflow ($\to\infty$, then NaN) and underflow ($\to0$, then $\log0$).
- Stability: use log-sum-exp (subtract the max), log-softmax,
log1p,expm1, losses computed from logits, log-likelihoods instead of products, two-pass variance. - Conditioning $\kappa$ = speed of GD ($\sim\kappa$ steps) = sensitivity ($\|\delta x\|/\|x\|\le\kappa\|\delta b\|/\|b\|$) = digits lost ($\log_{10}\kappa$). Preconditioning changes variables to reduce $\kappa$; diagonal scaling fixes scale differences, not correlations.
- Finite differences: forward error $O(h)$, central $O(h^2)$, truncation falls and round-off $\approx\varepsilon|f|/h$ rises; sweet spots $h\approx\varepsilon^{1/2},\varepsilon^{1/3},\varepsilon^{1/5}$ (forward, central, fourth-order); central in float64: $h\approx10^{-5}$, error $\approx10^{-11}$.
- Numerical gradient costs $2n$ evaluations; gradient checking uses $\mathrm{rel}=\frac{|a-n|}{\max(|a|,|n|,\tau)}$ in float64 (below $10^{-7}$ excellent, above $10^{-3}$ a bug); read bug signatures; use random directions for large models.
- Automatic differentiation: forward mode = one column per sweep ($n$ sweeps), reverse mode = one row per sweep ($m$ sweeps); a scalar loss needs one backward sweep ($\approx2$ to $4\times$ the cost of the loss) plus memory for the forward values; checkpointing: memory $\approx2\sqrt L$ for about $+33\%$ time.
- Complexity per iteration: GD $O(mn)$, SGD $O(bn)$, Newton $O(mn^2+n^3)$, BFGS $O(mn+n^2)$, L-BFGS $O(mn+kn)$, CG $O(mn+n)$; memory $O(n)$, $O(n)$, $O(n^2)$, $O(n^2)$, $O(kn)$, $O(n)$.
- Memory: weights + gradients + optimizer state (SGD 1×, momentum 2×, Adam 3× the parameters, plus gradients) + activations; Adam in float32 or mixed precision: 16 bytes per parameter; loss scaling rescues small float16 gradients; gradient accumulation reproduces a big batch in small memory.
Cheat sheet
| Topic | Formula or fact | Rule of thumb |
|---|---|---|
| Machine epsilon | $\varepsilon=2^{-M}$, $u=\varepsilon/2$ | $\mathrm{fl}(w+\delta)=w$ if $|\delta|<u|w|$ |
| Cancellation | amplification $\frac{|a|+|b|}{|a-b|}$ | reformulate before subtracting |
| Log-sum-exp | $m+\log\sum e^{z_j-m}$, $m=\max z$ | pass logits, not probabilities |
| Central difference | $\frac{f(x+h)-f(x-h)}{2h}$, error $\tfrac{h^2}6f'''+\tfrac{u|f|}{h}$ | $h\approx\varepsilon^{1/3}\max(|x|,1)$ |
| Gradient check | $\frac{|a-n|}{\max(|a|,|n|,\tau)}$ | float64; below $10^{-7}$ good, above $10^{-3}$ bug |
| Condition number | $\lambda_{\max}/\lambda_{\min}$; $\kappa(A^\top A)=\kappa(A)^2$ | about $\log_{10}\kappa$ digits lost |
| Preconditioned GD | $\mathbf{x}\leftarrow\mathbf{x}-\eta M\nabla f$, $M\approx H^{-1}$ | Jacobi: $M=\mathrm{diag}(H)^{-1}$ |
| Numerical gradient cost | $2n$ evaluations | checking only; random direction for big models |
| Reverse vs forward AD | $m$ sweeps vs $n$ sweeps for $J\in\mathbb{R}^{m\times n}$ | loss: reverse; few inputs: forward |
| Checkpointing | memory $L/k+k$, best $k=\sqrt L$ | +1 forward pass |
| Training memory | 16 bytes per parameter (Adam) | 7B parameters: 112 GB + activations |
| Loss scaling | $S$ up, divide gradients by $S$ after backward | dynamic: halve on inf, double after 2000 good steps |
import numpy as np
# 1) Cancellation: variance of offset + [1..5] in float32 (true variance = 2)
x = (np.float32(1e6) + np.arange(1, 6, dtype=np.float32))
naive = np.mean(x * x) - np.mean(x) ** 2 # E[x^2] - E[x]^2
twopass = np.mean((x - np.mean(x)) ** 2) # subtract the mean first
print(naive, twopass) # 65536.0 2.0
# 2) Stable softmax / log-sum-exp (float32 logits that would overflow exp)
z = np.array([1000.0, 1001.0, 999.0], dtype=np.float32)
with np.errstate(over="ignore", invalid="ignore"):
print(np.exp(z) / np.exp(z).sum()) # [nan nan nan] naive
m = z.max()
e = np.exp(z - m)
print(e / e.sum()) # [0.24472848 0.66524094 0.09003057]
print(m + np.log(e.sum())) # 1001.4076 (log-sum-exp)
# 3) log1p against log(1 + x) for tiny x
xs = 1e-12
print(np.log(1 + xs), np.log1p(xs)) # 1.000088900581841e-12 9.999999999995e-13
# 4) Central difference: error against h (float64, f = exp, x = 1)
hs = 10.0 ** -np.arange(1, 13)
err = [abs((np.exp(1 + h) - np.exp(1 - h)) / (2 * h) - np.e) for h in hs]
best = int(np.argmin(err))
print(f"best h = {hs[best]:.0e}, error = {err[best]:.1e}") # best h = 1e-05, error = 5.9e-11
# 5) Gradient check with the max-based relative error (softmax regression)
rng = np.random.default_rng(0)
X = rng.standard_normal((6, 3)); y = rng.integers(0, 3, 6); W = rng.standard_normal((3, 3)) * 0.5
def loss(W, X=X, y=y):
Z = X @ W; Z = Z - Z.max(1, keepdims=True)
return -np.mean(Z[np.arange(len(y)), y] - np.log(np.exp(Z).sum(1)))
def grad(W, X=X, y=y, bug=False):
Z = X @ W; P = np.exp(Z - Z.max(1, keepdims=True)); P /= P.sum(1, keepdims=True)
P[np.arange(len(y)), y] -= 1
return X.T @ P * (1.0 if bug else 1.0 / len(y)) # bug: forgot the 1/N
def numeric(W, h=1e-5):
G = np.zeros_like(W)
for i in range(W.shape[0]):
for j in range(W.shape[1]):
E = np.zeros_like(W); E[i, j] = h
G[i, j] = (loss(W + E) - loss(W - E)) / (2 * h)
return G
def rel(a, b): return np.abs(a - b) / np.maximum(np.maximum(np.abs(a), np.abs(b)), 1e-8)
print(rel(grad(W), numeric(W)).max()) # about 1e-9 -> passes
print(rel(grad(W, bug=True), numeric(W)).max()) # 0.8333 -> fails: every entry is off by the factor N = 6
v = rng.standard_normal(W.shape); v /= np.linalg.norm(v) # random-direction check: two loss evaluations
h = 1e-5
print(abs((loss(W + h * v) - loss(W - h * v)) / (2 * h) - np.sum(grad(W) * v))) # tiny (about 1e-12)
# 6) Memory per parameter, and gradient accumulation
P = 7e9 # a 7-billion-parameter model
print("fp32 + Adam:", 16 * P / 1e9, "GB") # weights 4 + grads 4 + Adam m, v 8 = 16 bytes per parameter -> 112.0 GB
g_big = grad(W) # one batch of 6 examples
g_acc = (grad(W, X[:3], y[:3]) + grad(W, X[3:], y[3:])) / 2 # two micro-batches of 3, gradients averaged
print(np.allclose(g_big, g_acc)) # True: accumulation reproduces the big-batch gradient
1. In float32, the weight is $w=1.0$ and the update is $10^{-8}$. What is $w+10^{-8}$?
2. Which is the numerically stable way to compute $\log\sum_i e^{z_i}$?
3. For a central difference in float64, which step size is the best starting choice?
4. A gradient check shows relative error $0.75$ on every entry (all analytic values are $4\times$ the numeric ones). What is the most likely bug?
5. A loss depends on $10^{6}$ weights and returns one number. Which differentiation method gives the whole gradient most cheaply?
6. How much memory do weights, gradients and Adam states need for a 3-billion-parameter model in float32?
Practice problems
A. Estimate the best step $h$ for a central difference of $\sin x$ at $x=1$ in float64 and the error bound at that $h$.
$f=\sin1=0.8415$, $|f'''|=|\cos1|=0.5403$, $u=1.1\times10^{-16}$. $h^\star=(3u|f|/|f'''|)^{1/3}=(5.19\times10^{-16})^{1/3}\approx8.0\times10^{-6}$. The bound is $\tfrac{h^2}6|f'''|+\tfrac{u|f|}h=5.8\times10^{-12}+1.2\times10^{-11}\approx1.7\times10^{-11}$. Measured error at $h=8\times10^{-6}$: $7\times10^{-12}$ (round-off is random, so the real error is usually a bit below the worst-case bound).
B. Why does $1-\cos x$ evaluate to exactly $0$ for $x=10^{-8}$ in float64, and what is the stable alternative?
$\cos(10^{-8})=1-5\times10^{-17}$, and the gap just below $1$ is $1.1\times10^{-16}$, so the cosine rounds to exactly $1.0$, and $1-1=0$. The true value is $5\times10^{-17}$. Stable form: $2\sin^2(x/2)=2\,(5\times10^{-9})^2=5\times10^{-17}$ ✓ (no subtraction of close numbers).
C. Compute the softmax of the logits $(800,799,790)$ by hand with the max-shift trick.
Shift by $m=800$: exponents $0,-1,-10$ give $1,\ 0.367879,\ 4.54\times10^{-5}$; sum $=1.367925$. Probabilities $\approx(0.73103,\ 0.26893,\ 3.3\times10^{-5})$. The naive $e^{800}$ overflows even in float64 (limit $e^{709.8}$).
D. $n=10^4$ parameters, $m=10^5$ examples. Compare per-iteration work and memory for GD, Newton, BFGS and L-BFGS ($k=10$).
GD: $mn=10^9$ operations, $2n=2\times10^4$ numbers. Newton: $mn^2+n^3=10^{13}+10^{12}\approx1.1\times10^{13}$ operations (about $10^4\times$ GD), $n^2=10^8$ numbers ($400$ MB). BFGS: $mn+2n^2\approx1.2\times10^9$ operations but $n^2=10^8$ numbers ($400$ MB) of memory. L-BFGS: $mn+4kn\approx10^9$ operations, $2kn=2\times10^5$ numbers ($0.8$ MB). Newton needs far fewer iterations, but each costs $10^4$ times more.
E. A 13-billion-parameter model is trained with Adam in mixed precision. What is the static memory, and what is the minimum number of 80 GB GPUs that can hold it if everything is sharded evenly?
$16$ bytes per parameter: $13\times10^9\times16=208$ GB. Even perfect sharding needs $208/80=2.6$, so at least $3$ GPUs, and in practice more to leave room for activations and buffers.
F. A float16 gradient is $3\times10^{-9}$. What happens without loss scaling, with $S=2^{12}$, and with $S=2^{16}$?
Without scaling: $3\times10^{-9}$ is below $3\times10^{-8}$ (half the smallest subnormal $6\times10^{-8}$), so it becomes exactly $0$. With $S=4096$: $1.23\times10^{-5}$, a subnormal float16 number (accurate to about $0.5\%$). With $S=65536$: $1.97\times10^{-4}$, a normal float16 number. But then any gradient above $65504/65536\approx1$ overflows, which is why $S$ is adjusted dynamically.
Optimization Algorithms — The Final Toolbox
You now know many ways to walk downhill. This last chapter puts them all in one toolbox, teaches the four methods that earlier chapters only named (conjugate gradient, projected gradient, penalty methods and barrier methods), and ends with the question that matters in real work: which method should I use for this problem?
- Sort every algorithm in the syllabus into one of four families, and say what information each one uses
- Read each method's update rule, memory, cost, strengths and weaknesses at a glance
- Understand conjugate gradient: directions that do not spoil each other, exact in at most $n$ steps on a quadratic, and $\sqrt{\kappa}$ instead of $\kappa$
- Use projected gradient (step, then snap back into the allowed set), penalty methods (a fine for crossing the fence) and barrier methods (an invisible wall inside the fence)
- Choose a method with a short list of questions, and watch all of them race on the same landscape
The toolbox map: what does each method know?
Picture hikers in thick fog, all trying to reach the bottom of a valley.
- Hiker A only feels the slope under their boots. They step the way the ground falls. That is a first-order method.
- Hiker B feels the slope and how fast the slope is changing (is the ground curving like a bowl?). From that they can guess where the bottom is and walk straight to it. That is a second-order method.
- Hiker C remembers the last few steps and learns the curve from experience, without measuring it. That is quasi-Newton.
- Hiker D has a fence around them and must stay inside. That is a constrained method.
- Hiker E knows the valley has a special shape (say, a perfect bowl) and uses that. That is a specialised method.
Every algorithm in this guide is one of these hikers, with small changes in what is felt, what is remembered, and how the fence is handled.
Two models of the same function. Take $f(x) = e^{x} - 2x$ and stand at $x_0 = 0$.
- Values: $f(0) = 1$. First derivative: $f'(x) = e^x - 2$, so $f'(0) = -1$. Second derivative: $f''(x) = e^x$, so $f''(0) = 1$.
- First-order model (a straight line): $f(x) \approx 1 - 1\cdot x$. A line has no bottom, so the slope only says "go right". How far is your choice: the learning rate $\eta$. With $\eta = 0.5$, gradient descent goes to $x_1 = 0 - 0.5\cdot(-1) = 0.5$.
- Second-order model (a parabola): $f(x) \approx 1 - x + \tfrac12 x^2$. A parabola has a bottom: set its slope $-1 + x$ to zero, so $x = 1$. This is Newton's step: $x_1 = 0 - \dfrac{f'(0)}{f''(0)} = 0 - \dfrac{-1}{1} = 1$.
- The true minimiser is $\ln 2 \approx 0.6931$. Newton's next points are $0.7358$, then $0.6940$, then $0.693148$. The error goes $0.31 \to 0.043 \to 0.0009 \to 0.0000004$: the number of correct digits roughly doubles each step.
The parabola model tells Newton's method how far to go with no learning rate. The price is that you must compute (and, in many dimensions, store and invert) the second derivatives.
A method is of order $k$ if each step uses the derivatives of $f$ up to order $k$ (and nothing higher). The solver is allowed to ask about the point it stands on:
- Zeroth order: only values $f(x)$. (Random search, Nelder–Mead. Awareness only; used when there is no gradient.)
- First order: values and the gradient $\nabla f(x)$. It builds a tangent plane.
- Second order: also the Hessian $\nabla^2 f(x)$ (calculus guide). It builds a bowl.
- Quasi-Newton: only gradients, but it learns an approximate Hessian from the gradients it has already seen.
- Specialised: uses the structure of $f$ (a quadratic, a smooth part plus an L1 part, a sum over coordinates).
- Constrained: also uses the constraint functions $g_i(x) \le 0$ and $h_j(x) = 0$ (Chapter 3.7).
Notation for the whole chapter: $x \in \mathbb{R}^n$ is the vector of $n$ unknowns, $\eta$ is the step size (learning rate), $g = \nabla f(x)$ is the gradient (a column vector), $x^\star$ is the optimum, and $\kappa$ is the condition number (largest curvature divided by smallest curvature).
Why do we need it?
There is no single best optimizer. Each one trades cheap steps against few steps, and each needs different information. A map of the families tells you in seconds which hiker to send.
Where is it used?
Every training script picks one: Adam or SGD with momentum for neural networks, L-BFGS for logistic regression in scikit-learn, coordinate descent for the Lasso, interior-point solvers for linear and quadratic programs.
How is it used?
Ask three things: what can I compute cheaply (the gradient? the Hessian?), is the problem noisy or huge, and are there constraints. The answers pick the family. The rest of this chapter makes each family concrete.
| Family | Information used | Model of $f$ it builds | Cost of one step | Steps needed (rule of thumb) |
|---|---|---|---|---|
| Zeroth order | values $f(x)$ | none | cheap | very many |
| First order | $f$ and $\nabla f$ | a tangent line or plane | about $n$ operations | hundreds to millions |
| Second order | $f$, $\nabla f$, $\nabla^2 f$ | a parabola or bowl | $n^2$ memory, about $n^3$ to solve | a handful to a few dozen |
| Quasi-Newton | $f$, $\nabla f$ (history) | an estimated bowl | $n^2$ (BFGS) or $mn$ (L-BFGS) | tens to hundreds |
| Specialised | structure of $f$ | exact pieces | depends | depends |
| Constrained | $f$ and the constraints | a landscape with a fence | one inner solve per fence method | depends |
"Second order" does not mean "second best". The word order counts the derivatives a method uses. It says nothing about whether the method is good for your problem. A first-order method is the right tool for a neural network exactly because it avoids the Hessian.
Do not mix this up with the order of convergence of Chapter 3.5 (linear, quadratic). Newton's method has quadratic convergence and uses second derivatives, but those are two different uses of the word "order".
Quick check: in the example, what step size $\eta$ makes gradient descent land exactly where Newton's method does, at $x_0 = 0$?
Gradient descent goes to $x_0 - \eta f'(x_0) = \eta$ (because $f'(0) = -1$). Newton goes to $1$. So $\eta = 1 = 1/f''(0)$. In general, Newton's method is gradient descent with a step size $1/f''$ chosen by the curvature at that spot. In many dimensions that single number becomes a matrix, $(\nabla^2 f)^{-1}$.
Family 1: first-order methods (the gradient is enough)
All first-order methods start from the same move: step against the gradient. They differ in how they use the history of gradients.
- Plain (gradient descent): forget the past, use today's gradient.
- Noisy but cheap (SGD, mini-batch): use a quick estimate of the gradient from a few examples.
- With inertia (momentum, Nesterov): a rolling ball keeps going in a direction it has been going.
- With per-coordinate brakes (AdaGrad, RMSProp, Adam, AdamW): every coordinate gets its own step size, small where gradients have been large and large where they have been small.
The same gradient, seven different steps. Let $f(x) = x^2$ and stand at $x = 4$. The gradient is $f'(4) = 8$. This is the first step of each method (so memories start at zero).
- Gradient descent, $\eta = 0.1$: $x \leftarrow 4 - 0.1\cdot 8 = 3.2$.
- Momentum ($\beta = 0.9$, $\eta = 0.1$): the velocity starts at $0$, so $v = 0.9\cdot 0 - 0.1\cdot 8 = -0.8$ and $x = 4 - 0.8 = 3.2$. Same as GD on step 1; the memory only matters from step 2 on.
- Nesterov (same numbers): the look-ahead point is $x + \beta v = 4$ (as $v = 0$), so again $3.2$.
- AdaGrad, $\eta = 0.5$: $G = 8^2 = 64$, so the step is $0.5\cdot 8/\sqrt{64} = 0.5$ and $x = 3.5$. (Whatever the size of the gradient, the first step is $\eta$.)
- RMSProp ($\beta = 0.9$, $\eta = 0.1$): $E = 0.1\cdot 64 = 6.4$, $\sqrt{E} = 2.530$, so the step is $0.1\cdot 8/2.530 = 0.316$ and $x = 3.684$.
- Adam ($\eta = 0.1$, $\beta_1 = 0.9$, $\beta_2 = 0.999$): $m = 0.1\cdot 8 = 0.8$, $v = 0.001\cdot 64 = 0.064$. Bias-correct: $\hat m = 0.8/0.1 = 8$, $\hat v = 0.064/0.001 = 64$. Step $= 0.1\cdot 8/\sqrt{64} = 0.1$, so $x = 3.9$.
- AdamW (also decay $\lambda = 0.01$): the Adam step, plus the weight decay $\eta\lambda x = 0.1\cdot 0.01\cdot 4 = 0.004$, so $x = 3.896$.
Same place, same gradient, and the steps are $0.8,\ 0.8,\ 0.8,\ 0.5,\ 0.316,\ 0.1,\ 0.104$. The first three scale with the gradient; the adaptive methods mostly ignore its size.
A first-order method only calls an "oracle" that returns $f(x)$ and $\nabla f(x)$. Almost all of them fit one template:
$$x_{k+1} = x_k + \eta_k\, d_k, \qquad d_k = \text{a direction built from the gradients seen so far.}$$Here is the whole family on one page ($g = \nabla f(x)$; $\odot$ and $\sqrt{\cdot}$ act on each coordinate; $\epsilon \approx 10^{-8}$ avoids dividing by zero). "Extra memory" counts the vectors of length $n$ stored beyond $x$ itself.
| Method | Update rule (one line) | Extra memory · cost per step | Strengths | Weaknesses | Details |
|---|---|---|---|---|---|
| Gradient descent | $x \leftarrow x - \eta\, \nabla f(x)$ | none · one full gradient | simple, steady, easy to analyse | slow when $\kappa$ is big; must read all the data each step | 3.3 |
| SGD | $x \leftarrow x - \eta\, \nabla \ell_i(x)$, one random example $i$ | none · one example | cheap steps, scales to huge data | noisy; needs a shrinking $\eta$ to settle | 3.14 |
| Mini-batch SGD | $x \leftarrow x - \eta\, \frac{1}{B}\sum_{i \in \text{batch}} \nabla \ell_i(x)$ | none · $B$ examples | GPU friendly; noise falls like $1/B$ | batch size and $\eta$ must be tuned together | 3.14 |
| Momentum | $v \leftarrow \beta v - \eta g$, then $x \leftarrow x + v$ | 1 vector · one gradient | smooths zig-zags, speeds up along valleys | one more setting; can overshoot | 3.4 |
| Nesterov | $v \leftarrow \beta v - \eta\, \nabla f(x + \beta v)$, then $x \leftarrow x + v$ | 1 vector · one gradient | looks ahead, so it brakes earlier; with the right momentum schedule it reaches the best possible rate among first-order methods on smooth convex problems | same extra setting | 3.4 |
| AdaGrad | $G \leftarrow G + g\odot g$; $x \leftarrow x - \eta\, g/(\sqrt{G}+\epsilon)$ | 1 vector · one gradient | a step size per coordinate; good for rare (sparse) features | $G$ only grows, so steps shrink toward zero and learning stalls | 3.4 |
| RMSProp | $E \leftarrow \beta E + (1-\beta)\,g\odot g$; $x \leftarrow x - \eta\, g/(\sqrt{E}+\epsilon)$ | 1 vector · one gradient | keeps AdaGrad's per-coordinate idea but forgets old gradients | no momentum, no bias correction; needs a small $\eta$ | 3.4 |
| Adam | $m \leftarrow \beta_1 m + (1-\beta_1) g$, $v \leftarrow \beta_2 v + (1-\beta_2)\, g\odot g$; $x \leftarrow x - \eta\, \hat m/(\sqrt{\hat v}+\epsilon)$ | 2 vectors · one gradient | momentum plus per-coordinate scaling; a robust default | can generalise worse than SGD on some tasks; L2 penalty interacts with the scaling | 3.4 |
| AdamW | Adam step, then $x \leftarrow x - \eta\lambda\, x$ (decay kept out of the gradient) | 2 vectors · one gradient | weight decay that behaves as intended; the usual choice for Transformers | one more setting, $\lambda$ | 3.4 |
Why do we need it?
A model with millions of weights can only afford to look at one number per weight (the gradient). Anything that stores or inverts a matrix of $n \times n$ numbers is out of reach, so first-order methods do almost all of deep learning.
Where is it used?
SGD with momentum for image models (ResNets), AdamW for Transformers and language models, Adam for most research code, AdaGrad for sparse features such as click prediction and word embeddings.
How is it used?
Pick an optimizer object (torch.optim.AdamW(params, lr=...)), call loss.backward() to get gradients, then optimizer.step(). The only real decisions are the learning rate, its schedule, and the batch size.
Adaptive does not mean better. Adam, RMSProp and AdaGrad make the step size per coordinate nearly independent of the gradient's size. That helps when coordinates have very different scales, but it also changes where the method ends up (it is not the same path as gradient descent). On some image tasks, well-tuned SGD with momentum generalises as well or better. Treat "Adam is the default" as a habit, not a theorem.
"Adam always converges faster" is false. It often looks faster in the first epochs; the final result depends on the task, the learning-rate schedule and the weight decay.
Quick check: with a constant gradient $g$, what does AdaGrad's step become after $t$ steps?
$G$ grows to $t\,g^2$, so the step is $\eta g/\sqrt{t g^2} = \eta/\sqrt{t}$. It shrinks like $1/\sqrt{t}$. That is why AdaGrad is great early and then slows down, and why RMSProp (which forgets) was invented.
Family 2: second-order and quasi-Newton methods
A first-order method only feels the slope, so in a long narrow valley it keeps bouncing off the walls. A second-order method also feels the curvature: "steep across the valley, nearly flat along it". It uses that to take a long step along the valley and a short step across it, and in a perfect bowl it walks straight to the bottom in one step.
The catch is the bill. Curvature in $n$ directions is a whole $n \times n$ table of numbers. Quasi-Newton methods are the clever compromise: do not measure the curvature, learn it from the steps you already took. L-BFGS goes further and only remembers the last few steps.
What would Newton's method cost on a modest model? Take $n = 1\,000\,000$ parameters (tiny for deep learning), stored as 4-byte numbers.
- The Hessian has $n \times n = 10^{12}$ entries. At 4 bytes each that is $4\times 10^{12}$ bytes $= 4$ terabytes.
- Solving $H d = -g$ (a Cholesky factorisation) costs about $n^3/3 \approx 3\times 10^{17}$ operations. At $10^{13}$ operations per second, one step takes about $3\times 10^{4}$ seconds, nine hours.
- L-BFGS with memory $m = 10$ keeps $2m = 20$ vectors of length $n$: $2\times 10^{7}$ numbers, which is 80 megabytes. Adam keeps 2 vectors: 8 megabytes.
On the Rosenbrock valley (a classic narrow, curved valley used to test optimizers), Newton needs about 21 iterations, BFGS about 33 and L-BFGS about 36, while plain gradient descent has still not finished after 3000 (you can check this in the race at the end of the chapter). Few iterations, but each is expensive.
Newton's method. Replace $f$ near $x_k$ by its parabola (bowl) $f(x_k) + g^\top d + \tfrac12 d^\top H d$ with $g = \nabla f(x_k)$, $H = \nabla^2 f(x_k)$, and jump to its bottom: set the gradient $g + Hd$ to zero.
$$H\,d = -g, \qquad x_{k+1} = x_k + d = x_k - H^{-1} g.$$On a quadratic $f = \tfrac12 x^\top A x - b^\top x$ we have $H = A$ and $g = Ax - b$, so $x_1 = x_0 - A^{-1}(Ax_0 - b) = A^{-1}b$, the exact answer in one step.
Quasi-Newton. Keep an estimate $B_k \approx H$ (or its inverse) and update it from the two differences
$$s_k = x_{k+1} - x_k, \qquad y_k = \nabla f(x_{k+1}) - \nabla f(x_k)$$so that it obeys the secant equation $B_{k+1} s_k = y_k$ ("the new curvature estimate must explain what we just saw"). BFGS keeps the inverse estimate $C_k \approx H^{-1}$ and updates it as
$$C_{k+1} = (I - \rho\, s y^\top)\,C_k\,(I - \rho\, y s^\top) + \rho\, s s^\top, \qquad \rho = \frac{1}{y^\top s},$$then steps along $d = -C_k g$ with a line search. L-BFGS never forms $C_k$: it stores only the last $m$ pairs $(s_i, y_i)$ and computes $C_k g$ with a short loop (the two-loop recursion). Chapter 3.12 derives all of this.
| Method | Update rule (one line) | Memory · cost per step | Strengths | Weaknesses | Details |
|---|---|---|---|---|---|
| Newton | $x \leftarrow x - [\nabla^2 f(x)]^{-1}\nabla f(x)$ | $n^2$ numbers · about $n^3$ to solve | quadratic convergence near the answer; ignores bad scaling ($\kappa$) | costly; can go the wrong way where the Hessian is not positive definite; needs damping or a line search | 3.12 |
| Quasi-Newton (the idea) | $x \leftarrow x - \eta\, B_k^{-1}\nabla f(x)$ with $B_k$ learned from $(s_k, y_k)$ | no Hessian ever computed | nearly Newton-quality steps from gradients alone | needs a good line search | 3.12 |
| BFGS | $d = -C_k g$; $C$ updated by the formula above | $n^2$ numbers · about $n^2$ per step | superlinear convergence; the standard small-problem optimizer (SciPy's default) | $n^2$ memory is too much beyond about $10^4$ parameters | 3.12 |
| L-BFGS | same, but $C_k g$ is built from the last $m$ pairs $(s_i, y_i)$ | $2mn$ numbers · about $4mn$ per step | works for millions of parameters; the default solver for logistic regression in scikit-learn | wants smooth, non-noisy gradients (full batch); hurt by noise from mini-batches | 3.12 |
Why do we need it?
When the gradient is cheap but the landscape is badly scaled, first-order methods need thousands of steps. Curvature information cuts that to tens of steps, and quasi-Newton methods give most of that gain without the $n^2$ or $n^3$ bill of true Newton.
Where is it used?
Logistic regression and CRFs (scikit-learn, L-BFGS), scientific fitting (scipy.optimize.minimize uses BFGS by default), Gaussian-process hyper-parameters, and the final polishing phase of some neural-style-transfer and physics-informed models.
How is it used?
Call scipy.optimize.minimize(f, x0, jac=grad, method="L-BFGS-B") or torch.optim.LBFGS. Supply exact gradients (use autodiff), keep the problem full-batch and smooth, and stop when $\|\nabla f\|$ is small.
Few steps is not the same as fast. Newton wins on the number of iterations, loses on the cost of each one, and in deep learning loses badly. Also, a Hessian that is not positive definite (a saddle or hill) makes the raw Newton step walk uphill; and even with positive curvature, a very flat parabola sends the raw step far past the minimum (start at $x = -1$ in the first widget of this chapter). So real implementations add damping, a line search, or a trust region.
And L-BFGS is for exact gradients. Its line search and its curvature pairs $(s_k, y_k)$ are corrupted by mini-batch noise, so it is rarely used for stochastic training.
Quick check: why can L-BFGS handle $n = 10^6$ when BFGS cannot?
BFGS stores the full $n \times n$ matrix $C_k$: $10^{12}$ numbers. L-BFGS stores only the last $m$ pairs $(s_i, y_i)$ (two vectors each), which is $2mn = 2\times 10^7$ numbers for $m = 10$, and computes $C_k g$ with a loop over those pairs instead of a matrix-vector product.
Family 3: specialised methods (use the structure of the problem)
Some problems come with a built-in shortcut. The trick is to notice it.
- The problem splits into one-dimensional pieces (the Lasso, SVM duals): fix all variables but one and solve that tiny problem exactly. Repeat for every variable. That is coordinate descent.
- The objective is a smooth part plus an awkward part (a loss plus an L1 penalty): take an ordinary gradient step on the smooth part, then fix up the awkward part exactly with one cheap operation (its proximal operator). That is proximal gradient.
- The problem is a perfect bowl (a quadratic, or equivalently a linear system $Ax = b$): use directions that never undo each other. That is conjugate gradient, taught in full in the next section.
Proximal step in one dimension. Minimise $\tfrac12 (x - v)^2 + t\,|x|$. The first term pulls $x$ toward $v$; the second pulls it toward $0$. The answer is "$v$ shrunk toward zero by $t$, but never past zero":
- $v = 3,\ t = 2$: the minimiser is $x = 3 - 2 = 1$.
- $v = -4,\ t = 1$: the minimiser is $x = -4 + 1 = -3$.
- $v = -0.5,\ t = 1$: the pull to zero ($t = 1$) is stronger than the distance ($0.5$), so $x = 0$ exactly.
This is called soft thresholding, $S_t(v) = \operatorname{sign}(v)\max(0, |v| - t)$. A whole Lasso solver is built from this one line (Chapter 3.13).
For a function $g$ and a step $\eta$, the proximal operator is $$\operatorname{prox}_{\eta g}(v) = \arg\min_x \Big\{ \tfrac12\|x - v\|^2 + \eta\, g(x) \Big\}.$$ For $g(x) = \lambda |x|$ (summed over coordinates) it is the soft threshold $S_{\eta\lambda}$ above. For a problem $\min_x\, f(x) + g(x)$ with $f$ smooth and $g$ simple but non-smooth, proximal gradient repeats
$$x \leftarrow \operatorname{prox}_{\eta g}\big(x - \eta\,\nabla f(x)\big)$$(one gradient step on $f$, then one prox step on $g$). For the Lasso this is called ISTA. If $g$ is the indicator of a set $C$ ($0$ inside, $+\infty$ outside), its prox is the projection onto $C$, and the rule becomes projected gradient (taught below). Coordinate descent instead cycles through the coordinates $i = 1, \dots, n$ and sets $x_i \leftarrow \arg\min_z f(x_1, \dots, z, \dots, x_n)$ (or takes a gradient step on that one coordinate). Conjugate gradient is the third member; see the next section.
| Method | Update rule (one line) | Memory · cost per step | Strengths | Weaknesses | Details |
|---|---|---|---|---|---|
| Coordinate descent | $x_i \leftarrow \arg\min_z f(x_1,\dots,z,\dots,x_n)$, cycling over $i$ | no extra memory · one coordinate (cheap if the 1-D problem is closed-form) | very fast for the Lasso, elastic net and SVM duals; no learning rate | can stall on non-smooth functions that couple coordinates; hard to parallelise | 3.13 |
| Proximal gradient (ISTA) | $x \leftarrow \operatorname{prox}_{\eta g}\big(x - \eta\nabla f(x)\big)$ | like GD · one gradient plus one prox | handles non-smooth terms (L1, constraints) exactly; same $O(1/k)$ as GD on convex problems | needs a cheap prox; plain version is slow (accelerated version FISTA helps) | 3.13 |
| Conjugate gradient | $x \leftarrow x + \alpha p$, with $p$ conjugate to all earlier directions | 3 vectors · one matrix-vector product $Ap$ | exact in at most $n$ steps on a quadratic; steps scale with $\sqrt{\kappa}$, not $\kappa$; never forms $A^{-1}$ | quadratics (or smooth problems with line search); wants symmetric positive definite $A$ | next section |
Why do we need it?
A general-purpose method ignores structure it could have used. The L1 penalty is not differentiable at zero, so gradient descent cannot create exact zeros. The prox step can. Quadratic problems have a finite-step finish. Exploiting structure turns "slow and approximate" into "fast and exact".
Where is it used?
Lasso and elastic net (scikit-learn's Lasso uses coordinate descent), linear SVMs (LIBLINEAR's dual coordinate descent), sparse coding, compressed sensing, total-variation image denoising (proximal), and large linear systems (conjugate gradient).
How is it used?
Write the objective as "smooth + simple". Use coordinate descent when each one-variable problem has a formula; use proximal gradient when the simple part has a known prox; use conjugate gradient when the problem is a quadratic or a symmetric positive definite system.
Coordinate descent can get stuck on non-smooth functions. For a smooth objective, or a non-smooth part that separates into one term per coordinate (like $\lambda\sum_i |x_i|$), coordinate descent converges. If the non-smooth part ties coordinates together, for example $|x_1 - x_2|$, every single-coordinate move can look uphill at a point that is not a minimum.
A prox is only useful if it is cheap. Soft thresholding costs $n$ operations. The prox of a complicated function may itself be an optimization problem.
Quick check: what is $S_t(v)$ for $v = 0.8$ and $t = 0.5$? And for $v = -0.3$?
$S_{0.5}(0.8) = \max(0, 0.8 - 0.5) = 0.3$. For $v = -0.3$: $|v| = 0.3 \le 0.5$, so the result is $0$. Values inside $[-t, t]$ are snapped to exactly zero.
Family 4: constrained methods, and one problem solved four ways
When there is a fence, there are four common ways to respect it. Imagine a hiker who must stay on one side of a straight fence:
- Do the maths of the fence (Lagrange multipliers, KKT). At the best point, the pull of the slope exactly balances the push of the fence. You solve equations instead of walking.
- Snap back (projected gradient). Take a normal step downhill. If you ended up on the wrong side, jump to the nearest point on the right side.
- Fine for crossing (penalty). Allow crossing, but charge a fine that grows as you go further over. A bigger fine keeps you closer to the fence. You end up just outside it.
- Invisible wall inside (barrier). Build a force that pushes you away from the fence and grows to infinity at the fence. You stay strictly inside, and weakening the wall lets you creep closer.
Our running problem for this whole chapter. Minimise the squared distance to the point $c = (2, 1)$, but stay in the half-plane $x_1 + x_2 \le 2$:
$$\min_{x}\ f(x) = (x_1 - 2)^2 + (x_2 - 1)^2 \quad \text{subject to} \quad g(x) = x_1 + x_2 - 2 \le 0.$$The unconstrained minimiser is $c = (2, 1)$ itself, but $2 + 1 = 3 > 2$: it breaks the rule. So the answer lies on the fence. By symmetry it is the closest point of the line $x_1 + x_2 = 2$ to $c$.
- KKT check (Chapter 3.9). We need $\nabla f(x) + \mu\,\nabla g(x) = 0$ with $\mu \ge 0$ and $g(x) = 0$. Here $\nabla f = 2(x - c)$ and $\nabla g = (1, 1)$.
- So $2(x_1 - 2) + \mu = 0$ and $2(x_2 - 1) + \mu = 0$. That gives $x_1 = 2 - \mu/2$ and $x_2 = 1 - \mu/2$.
- On the fence: $x_1 + x_2 = 3 - \mu = 2$, so $\mu = 1$.
- The answer: $x^\star = (1.5,\ 0.5)$ with multiplier $\mu = 1 \ge 0$, and $f(x^\star) = 0.25 + 0.25 = 0.5$.
Every method below should reach $(1.5, 0.5)$, in its own way, and the widget at the end of this section lets you watch it happen.
The general problem (notation from Chapter 3.7): $\min_x f(x)$ subject to $g_i(x) \le 0$ ($i = 1..m$) and $h_j(x) = 0$ ($j = 1..p$). Lagrange multipliers: $\lambda_j$ for equalities, $\mu_i \ge 0$ for inequalities. Here is the family at a glance:
| Method | Core rule (one line) | Memory · cost | Strengths | Weaknesses | Details |
|---|---|---|---|---|---|
| Lagrange multipliers | solve $\nabla f + \sum_j \lambda_j \nabla h_j = 0$ and $h_j = 0$ | a system of $n + p$ equations | exact; the multipliers are the sensitivity ("price") of each constraint | equality constraints only; the equations can be hard to solve | 3.8 |
| KKT conditions | stationarity, $g_i \le 0$, $h_j = 0$, $\mu_i \ge 0$, $\mu_i g_i = 0$ | a test, not a loop | necessary (under regularity) and, for convex problems, sufficient for optimality | by itself it does not tell you how to find the point | 3.9 |
| Projected gradient | $x \leftarrow \Pi_C\big(x - \eta\nabla f(x)\big)$ | like GD plus one projection | every iterate feasible; no extra variables | only when the projection is cheap (box, ball, simplex, ...) | section below |
| Penalty method | minimise $f + \frac{\rho}{2}\sum_i \max(0, g_i)^2 + \frac{\rho}{2}\sum_j h_j^2$, raise $\rho$ | one unconstrained solve per $\rho$ | simplest idea; any unconstrained solver works; infeasible starts are fine | never exactly feasible; the problem gets ill-conditioned as $\rho \to \infty$ | section below |
| Barrier method | minimise $f - \frac{1}{t}\sum_i \log(-g_i)$, raise $t$ | one Newton-solved problem per $t$ | iterates strictly feasible; duality-gap bound $m/t$; engine of modern LP/QP/SDP solvers | needs a strictly feasible start; needs Newton's method to stay efficient | section below |
Awareness only: augmented Lagrangian (penalty plus a multiplier estimate, so $\rho$ need not grow forever), SQP (sequential quadratic programming: Newton's method applied to the KKT equations, the workhorse for smooth, nonlinear, medium-sized problems) and ADMM (split a problem into easy pieces and let multipliers glue them) are other members of this family.
Why do we need it?
Unconstrained methods happily walk through a fence. Real problems have fences: probabilities must be non-negative and add to one, a budget cannot be exceeded, a weight vector must stay small. These four ideas are the four standard ways to stay legal.
Where is it used?
Portfolio weights on a simplex, adversarial attacks limited to a small box, non-negative matrix factorisation, SVMs (a constrained problem), LP/QP/SDP solvers for resource planning, trajectory optimization in robotics.
How is it used?
If the allowed set is simple (box, ball, simplex): projected gradient. If you only have a plain unconstrained solver at hand: a penalty. If you need a precise, feasible answer for a convex problem: a barrier (interior-point) solver, usually through a modelling tool such as CVXPY.
The exact answer depends on whether the fence is active. If $c$ lies inside the allowed region, the constraint does nothing and the best point is $c$ itself ($\mu = 0$). The four methods then agree trivially. The interesting case is when the unconstrained minimum is outside.
Quick check: for $c = (2,1)$, which one of the four answers in the widget breaks the constraint, and which ones never do?
The penalty method returns a point outside the allowed region (slightly infeasible for every finite $\rho$). The barrier and projected gradient keep every iterate feasible (the barrier strictly inside, projected gradient inside or on the fence; only the very first start point can be outside). The exact KKT point is on the fence.
Conjugate gradient: a perfect bowl in at most $n$ steps taught here
Why does gradient descent zig-zag? In a long, stretched bowl, the steepest direction points across the valley, not along it. Each step crosses the valley, overshoots, and the next step partly undoes what the last one did.
A thought experiment. Suppose the bowl were perfectly round. Then you could fix the problem one coordinate at a time: slide east-west until it stops dropping, then north-south, and you are at the bottom. Each direction finishes its job for good, because sliding north-south does not change where the east-west minimum is. Two directions, two steps.
A stretched bowl is a round bowl seen through a distorting lens. Conjugate directions are the directions that look perpendicular through the lens. If you slide along them one after another, each new slide leaves all earlier work intact. Conjugate gradient (CG) finds such directions as it goes, without ever being told what the lens is.
The name is honest: it uses gradients (more exactly, residuals) and mixes in a bit of the previous direction so that the new direction is conjugate to the old ones.
Two steps by hand. Minimise $f(x) = \tfrac12 x^\top A x - b^\top x$ with $A = \begin{bmatrix} 3 & 1 \\ 1 & 2 \end{bmatrix}$, $b = \begin{bmatrix} 1 \\ 2 \end{bmatrix}$, starting at $x_0 = (0, 0)$. (Equivalently, solve $Ax = b$; the answer is $x^\star = (0, 1)$, since $3\cdot 0 + 1\cdot 1 = 1$ and $1\cdot 0 + 2\cdot 1 = 2$.) Write $r = b - Ax$ for the residual; it is exactly $-\nabla f(x)$.
- Start. $r_0 = b - A\cdot 0 = (1, 2)$. The first direction is the steepest-descent one: $p_0 = r_0 = (1, 2)$.
- Step length. $A p_0 = (3\cdot 1 + 1\cdot 2,\ 1\cdot 1 + 2\cdot 2) = (5, 5)$. Then $r_0^\top r_0 = 1 + 4 = 5$ and $p_0^\top A p_0 = 1\cdot 5 + 2\cdot 5 = 15$, so $\alpha_0 = \dfrac{5}{15} = \dfrac13$. New point: $x_1 = (0,0) + \frac13 (1, 2) = (\tfrac13, \tfrac23)$.
- New residual. $r_1 = r_0 - \alpha_0 A p_0 = (1, 2) - \frac13(5, 5) = (-\tfrac23, \tfrac13)$.
- Mix in the old direction. $r_1^\top r_1 = \frac49 + \frac19 = \frac59$, so $\beta_0 = \dfrac{5/9}{5} = \dfrac19$. New direction: $p_1 = r_1 + \beta_0 p_0 = (-\tfrac23 + \tfrac19,\ \tfrac13 + \tfrac29) = (-\tfrac59,\ \tfrac59)$.
- Second step. $A p_1 = (-\tfrac{15}{9} + \tfrac59,\ -\tfrac59 + \tfrac{10}{9}) = (-\tfrac{10}{9}, \tfrac59)$, and $p_1^\top A p_1 = \tfrac{50}{81} + \tfrac{25}{81} = \tfrac{25}{27}$. So $\alpha_1 = \dfrac{5/9}{25/27} = \dfrac35$ and $x_2 = (\tfrac13, \tfrac23) + \tfrac35(-\tfrac59, \tfrac59) = (0, 1)$. Exact.
- The conjugacy check. $p_0^\top A p_1 = (1, 2)\cdot(-\tfrac{10}{9}, \tfrac59) = -\tfrac{10}{9} + \tfrac{10}{9} = 0$. The two directions are conjugate, though $p_0\cdot p_1 = -\tfrac59 + \tfrac{10}{9} = \tfrac59 \ne 0$: they are not perpendicular in the ordinary picture.
Compare with steepest descent (same start; "exact line search" means each step slides along its direction to the lowest point on that line). Step 1 is identical. Step 2 goes along $r_1 = (-\tfrac23, \tfrac13)$ with $\alpha = \frac{5/9}{10/9} = \frac12$, landing at $(0, \tfrac56) \approx (0, 0.833)$, not yet at the solution; it needs many more steps. The value of $f$: CG reaches $-1$ (the minimum) in 2 steps; steepest descent has $f(0, \tfrac56) = \tfrac{25}{36} - \tfrac{5}{3} = -0.972$ after 2.
Problem. Minimise the quadratic $f(x) = \tfrac12 x^\top A x - b^\top x$ where $A$ is symmetric positive definite (positive definite matrices). Its gradient is $Ax - b$, so the minimiser solves the linear system $Ax = b$. The residual is $r = b - Ax = -\nabla f(x)$, and $\kappa$ is the ratio of the largest to the smallest eigenvalue of $A$.
Conjugate directions. Two vectors are $A$-conjugate (or $A$-orthogonal) if $p_i^\top A\, p_j = 0$ for $i \ne j$. It is the ordinary dot product measured through the lens $A$.
Algorithm. Start with $r_0 = b - Ax_0$, $p_0 = r_0$. For $k = 0, 1, 2, \dots$:
$$\alpha_k = \frac{r_k^\top r_k}{p_k^\top A p_k}, \quad x_{k+1} = x_k + \alpha_k p_k, \quad r_{k+1} = r_k - \alpha_k A p_k, \quad \beta_k = \frac{r_{k+1}^\top r_{k+1}}{r_k^\top r_k}, \quad p_{k+1} = r_{k+1} + \beta_k\, p_k.$$What you get (in exact arithmetic):
- The directions are conjugate, $p_i^\top A p_j = 0$, and the residuals are orthogonal, $r_i^\top r_j = 0$ for $i \ne j$.
- Finite termination: $x_n = x^\star$. It is even done after at most as many steps as $A$ has distinct eigenvalues.
- Rate: the error in the energy norm $\|e\|_A = \sqrt{e^\top A e}$ ($e = x - x^\star$) obeys $$\|e_k\|_A \le 2\left(\frac{\sqrt{\kappa} - 1}{\sqrt{\kappa} + 1}\right)^{k}\|e_0\|_A, \qquad \text{while steepest descent only guarantees} \quad \left(\frac{\kappa - 1}{\kappa + 1}\right)^{k}.$$ So CG needs about $\tfrac12\sqrt{\kappa}\,\ln(2/\varepsilon)$ steps to shrink the error by a factor $\varepsilon$, steepest descent about $\tfrac12\kappa\ln(1/\varepsilon)$. For $\kappa = 100$ and $\varepsilon = 10^{-6}$ that is about 73 steps against about 690.
- Cost per step: one product $Ap$ plus a few vector operations. It stores 3 vectors ($r$, $p$, $Ap$) beyond $x$. It never forms $A^{-1}$, and it never even needs $A$ as a table: only a routine that returns $Av$ for any $v$.
Where do $\alpha_k$ and $\beta_k$ come from? (Derivation, then a numerical check.)
- Best step along $p$. Along $x + \alpha p$ the function is $\varphi(\alpha) = f(x) - \alpha\, p^\top r + \tfrac12\alpha^2 p^\top A p$ (expand and use $\nabla f = -r$). Setting $\varphi'(\alpha) = 0$ gives $\alpha = \dfrac{p^\top r}{p^\top A p}$.
- The residual update is cheap. $r_{k+1} = b - A(x_k + \alpha_k p_k) = r_k - \alpha_k A p_k$. We already computed $Ap_k$, so no second product is needed.
- Exact steps leave an orthogonal residual: $p_k^\top r_{k+1} = p_k^\top r_k - \alpha_k p_k^\top A p_k = 0$ by step 1. The new gradient is perpendicular to the direction we just used.
- A simpler $\alpha$. Since $p_k = r_k + \beta_{k-1} p_{k-1}$ and $r_k \perp p_{k-1}$ (step 3), $p_k^\top r_k = r_k^\top r_k$. So $\alpha_k = \dfrac{r_k^\top r_k}{p_k^\top A p_k}$, the formula in the box.
- Make the next direction conjugate. Try $p_{k+1} = r_{k+1} + \beta p_k$ and demand $p_k^\top A p_{k+1} = 0$. That forces $\beta = -\dfrac{p_k^\top A r_{k+1}}{p_k^\top A p_k}$.
- Simplify. From step 2, $A p_k = (r_k - r_{k+1})/\alpha_k$. Using $r_k^\top r_{k+1} = 0$ (one of the orthogonality facts from the induction in the next step): $p_k^\top A r_{k+1} = \dfrac{(r_k - r_{k+1})^\top r_{k+1}}{\alpha_k} = -\dfrac{r_{k+1}^\top r_{k+1}}{\alpha_k}$. And from step 4, $p_k^\top A p_k = \dfrac{r_k^\top r_k}{\alpha_k}$. Divide: $\beta_k = \dfrac{r_{k+1}^\top r_{k+1}}{r_k^\top r_k}$. ∎
- The surprise is that being conjugate to the last direction makes $p_{k+1}$ conjugate to all earlier ones, because $r_{k+1}$ is orthogonal to all earlier residuals and directions (a short induction; see Nocedal and Wright, Chapter 5). That is why CG needs only a one-line memory of the past.
- Check with the example: $\alpha_0 = 5/15$ and $\beta_0 = (5/9)/5 = 1/9$ are exactly the numbers above, and $p_0^\top A p_1 = 0$.
Why at most $n$ steps? Conjugate directions are linearly independent: if $\sum_i c_i p_i = 0$, multiply by $p_j^\top A$ to get $c_j\, p_j^\top A p_j = 0$, so $c_j = 0$. So $p_0, \dots, p_{n-1}$ is a basis, and the starting error can be written $x^\star - x_0 = \sum_i \delta_i p_i$. Conjugacy shows that step $k$ removes exactly the term $\delta_k p_k$ and leaves the others alone. After $n$ steps all $n$ terms are gone.
Why do we need it?
Direct solvers cost about $n^3$ operations and fill memory; gradient descent is cheap per step but needs about $\kappa$ steps. Conjugate gradient needs only matrix-vector products and about $\sqrt{\kappa}$ steps: the best of both.
Where is it used?
Large sparse symmetric positive definite systems (finite elements, graph Laplacians), Gaussian-process regression (GPyTorch solves $K^{-1}y$ with CG), ridge regression without forming $X^\top X$, "Newton-CG" and Hessian-free training (solve $Hd = -g$ with Hessian-vector products), and SciPy's trust-ncg and Newton-CG.
How is it used?
Write a function that multiplies by $A$; call scipy.sparse.linalg.cg(A, b, rtol=1e-8, M=precond); stop when the residual $\|r\|$ is small (not after $n$ steps); add a preconditioner $M$, a cheap approximate inverse, which lowers the effective $\kappa$. See the linear-algebra view of the same algorithm.
Nonlinear conjugate gradient (for a general smooth $f$, not only a quadratic). Replace the residual by the negative gradient $g_k = \nabla f(x_k)$ and the exact $\alpha_k$ by a line search (Chapter 3.3, Chapter 3.12):
$$p_0 = -g_0, \quad x_{k+1} = x_k + \alpha_k p_k, \quad p_{k+1} = -g_{k+1} + \beta_k\, p_k,$$ $$\beta_k^{\text{FR}} = \frac{g_{k+1}^\top g_{k+1}}{g_k^\top g_k} \quad \text{(Fletcher–Reeves)}, \qquad \beta_k^{\text{PR}} = \frac{g_{k+1}^\top (g_{k+1} - g_k)}{g_k^\top g_k} \quad \text{(Polak–Ribière)}.$$On a quadratic with exact line searches, $g_{k+1}^\top g_k = 0$, so both reduce to the formula $\beta_k = r_{k+1}^\top r_{k+1}/r_k^\top r_k$ above: the same algorithm. For non-quadratic $f$ they differ. In practice use $\max(0, \beta^{\text{PR}})$ ("PR+", which restarts with steepest descent when $\beta$ would be negative) and restart every $n$ steps. Memory is only a few vectors, which makes it a useful step up from gradient descent when $n$ is large and L-BFGS is still too heavy.
"Exactly $n$ steps" has three conditions: the problem is a quadratic with $A$ symmetric positive definite, the arithmetic is exact, and the line search is exact. With rounding errors the directions slowly lose their conjugacy, so in practice you stop on a residual tolerance, not after exactly $n$ steps (on hard, badly scaled systems you may need more than $n$ steps). Nonlinear CG with an inexact line search has no $n$-step guarantee. Run it in the race at the end of this chapter and compare its iteration count with the exact-quadratic case.
CG needs symmetric positive definite $A$. For other matrices use GMRES or BiCGSTAB, or apply CG to the normal equations $A^\top A x = A^\top b$ (which squares $\kappa$).
A good preconditioner is worth more than tuning. Preconditioned CG solves the same system but in a lens where the effective $\kappa$ is small.
Quick check: a system has $\kappa = 10\,000$. About how many CG steps versus steepest-descent steps for an error reduction of $10^{-6}$?
CG: $\tfrac12\sqrt{\kappa}\ln(2/\varepsilon) = \tfrac12\cdot 100\cdot 14.5 \approx 730$ steps. Steepest descent: $\tfrac12\kappa\ln(1/\varepsilon) = \tfrac12\cdot 10^4\cdot 13.8 \approx 69\,000$ steps. A factor of about 100, which is $\sqrt{\kappa}$.
Quick check: why can CG solve $Ax = b$ for a million unknowns without ever storing $A^{-1}$ (or even $A$)?
Every step only needs the product $A p_k$, which a program can compute from the structure of $A$ (for example, a sparse matrix or a convolution) without writing the matrix down. $A^{-1}$ is never needed: the algorithm builds the solution as a sum of the directions $\alpha_k p_k$.
Projected gradient: step downhill, then snap back into the allowed set taught here
You are walking downhill in a field with a fence. You take your normal step. If it carries you over the fence, you are not allowed to be there, so you walk to the nearest point on the allowed side. Then you take the next downhill step, and snap back again if needed.
Step, then snap back. That is the whole algorithm. The snapping has a name: projection. The projection of a point $y$ onto a set $C$ is the point of $C$ that is closest to $y$ (think of the shadow of $y$ on the set, cast by a light that shines straight at it). If $y$ is already inside, it does not move.
By hand, on a box. Minimise $f(x) = (x_1 - 3)^2 + (x_2 - 1)^2$ over the box $C = [0, 2] \times [0, 2]$. The unconstrained minimum $(3, 1)$ is outside the box, so we expect the answer on the right edge: $x^\star = (2, 1)$. Use $\eta = 0.25$ and start at $(0, 0)$. The gradient is $\nabla f = 2(x - (3, 1))$.
- Step 1. $\nabla f(0,0) = (-6, -2)$, so $y = (0,0) - 0.25(-6,-2) = (1.5, 0.5)$. It is inside the box, so $x_1 = (1.5, 0.5)$.
- Step 2. $\nabla f = (-3, -1)$, so $y = (1.5 + 0.75,\ 0.5 + 0.25) = (2.25, 0.75)$. The first coordinate is past the fence at $2$, so clip it: $x_2 = (2, 0.75)$.
- Step 3. $\nabla f(2, 0.75) = (-2, -0.5)$, so $y = (2.5, 0.875)$ and after clipping $x_3 = (2, 0.875)$.
- Step 4. $y = (2.5, 0.9375)$, so $x_4 = (2, 0.9375)$. The second coordinate's distance to $1$ is $0.5, 0.25, 0.125, 0.0625, \dots$: it halves every step. The first coordinate sits on the fence for good.
The iterates approach $(2, 1)$ and never leave the box.
For a closed convex set $C$, the projection of $y$ onto $C$ is $\Pi_C(y) = \arg\min_{x \in C}\|x - y\|^2$. (For convex $C$ this point exists and is unique.) Projected gradient descent:
$$x_{k+1} = \Pi_C\big(x_k - \eta\,\nabla f(x_k)\big).$$This is the proximal-gradient step of Chapter 3.13 with the "rule" $g$ chosen as the indicator of $C$ (zero inside, $+\infty$ outside): the prox of an indicator is the projection. Projections you can write in one line:
| Set $C$ | Projection $\Pi_C(y)$ | Cost | Example |
|---|---|---|---|
| Box $\{l_i \le x_i \le u_i\}$ | clip each coordinate: $\min(u_i, \max(l_i, y_i))$ | $O(n)$ | $\Pi_{[0,2]^2}(3, -1) = (2, 0)$ |
| Non-negative orthant $\{x \ge 0\}$ | $\max(0, y_i)$ for each $i$ | $O(n)$ | $(-1, 2) \mapsto (0, 2)$ |
| Ball $\{\|x\| \le r\}$ | $y$ if $\|y\| \le r$, else $r\,y/\|y\|$ | $O(n)$ | $\Pi_{r=1}(2, 2) = (\tfrac{1}{\sqrt2}, \tfrac{1}{\sqrt2}) \approx (0.707, 0.707)$ |
| Halfspace $\{a^\top x \le b\}$ | $y$ if $a^\top y \le b$, else $y - \dfrac{a^\top y - b}{\|a\|^2}\,a$ | $O(n)$ | $\Pi(2,1)$ for $x_1 + x_2 \le 2$ is $(1.5, 0.5)$ |
| Simplex $\{x \ge 0,\ \sum_i x_i = 1\}$ | $x_i = \max(0, y_i - \tau)$, with $\tau$ chosen so that the entries sum to $1$ (sort $y$, find $\tau$) | $O(n \log n)$ | $(0.8, 0.6, -0.2) \mapsto (0.6, 0.4, 0)$ with $\tau = 0.2$ |
Why it works (the convergence idea).
- Projection never pushes points apart: $\|\Pi_C(a) - \Pi_C(b)\| \le \|a - b\|$ (it is non-expansive). So adding a projection cannot spoil a step that was already shrinking the distance to $x^\star$.
- The answer is a fixed point. For convex $f$ and convex $C$, $x^\star$ solves the constrained problem exactly when $x^\star = \Pi_C(x^\star - \eta\nabla f(x^\star))$: a gradient step from $x^\star$ followed by the snap lands back on $x^\star$. The size of the move $\|x - \Pi_C(x - \eta\nabla f(x))\|/\eta$ (the gradient mapping) therefore plays the role that $\|\nabla f\|$ plays in unconstrained problems, and is the natural stopping test.
- Rates match gradient descent. For convex, $L$-smooth $f$ and $\eta = 1/L$: $f(x_k) - f^\star \le \dfrac{L\,\|x_0 - x^\star\|^2}{2k}$ (the same $O(1/k)$ as unconstrained GD). If $f$ is also $\mu$-strongly convex: $\|x_k - x^\star\|^2 \le (1 - \mu/L)^k\,\|x_0 - x^\star\|^2$, a linear rate. (Both are standard results; Chapter 3.5 explains the words.)
Why do we need it?
Some constraints cannot be ignored, for example "the weights are probabilities" or "the change must stay within a small box". Projection repairs a step in the cheapest way: it keeps the point as close as possible to where the gradient wanted it to go.
Where is it used?
Adversarial attacks (PGD: project the perturbation onto a small box around the image), portfolio weights and mixture probabilities (simplex), non-negative matrix factorisation and non-negative least squares, weight clipping and norm-ball constraints, bound-constrained model fitting, and the dual of the SVM (box constraints on the multipliers).
How is it used?
Take an ordinary gradient step, apply the projection function, repeat. Use $\eta \le 1/L$. Stop when the gradient mapping (how far one full step moves you) is tiny. Only choose this method if the projection is cheap (box, ball, simplex).
Projection means the nearest point. "Clip, then rescale so it sums to 1" is not the projection onto the simplex, and it gives a different answer; the guarantees above only hold for the true Euclidean projection.
The projection must be cheap. A box, ball or simplex costs about $n$ to $n\log n$. For a general set $\{Ax \le b\}$ the projection is itself a quadratic program, and projected gradient stops being a good idea.
Non-convex sets. The unit circle $\|x\| = 1$ has no unique projection of the origin, and the convergence results above are lost. The method is still used heuristically (for example, re-normalising weights), but treat it as a heuristic.
Not the same as gradient clipping. Clipping bounds the gradient; projection restricts the point.
Quick check: one projected step from $x = (1, 1)$ for $f(x) = \tfrac12\|x\|^2$ onto the box $[0.8, 2]^2$ with $\eta = 0.5$?
$\nabla f = x = (1,1)$, so $y = (1,1) - 0.5(1,1) = (0.5, 0.5)$. That is outside the box (below $0.8$), so clip: $x_{\text{new}} = (0.8, 0.8)$. This is the constrained minimum: the corner of the box that is nearest to the unconstrained minimum $(0,0)$.
Quick check: why is $x^\star = (2, 1)$ in the box example a fixed point of the method?
At $(2, 1)$ the gradient is $2(2-3,\ 1-1) = (-2, 0)$. A step gives $y = (2,1) - 0.25(-2, 0) = (2.5, 1)$. Clipping the first coordinate to $2$ returns $(2, 1)$: the point did not move.
Penalty methods: a fine for crossing the fence taught here
Replace the hard fence by a rubber wall. You may cross it, but the further you push, the harder it pushes back (a fine that grows with the distance you went over). The stiffer the rubber (the bigger the penalty weight $\rho$), the less you can push it, so you end up closer to the fence. But you never end up exactly on it, because at the very spot where the wall starts pushing, the pull of the slope still wins a little. You always stand slightly outside.
The big idea is that the problem becomes unconstrained. Any plain optimizer from earlier chapters can solve it. Then make the rubber stiffer and solve again, starting from the last answer.
One variable. Minimise $(x - 2)^2$ subject to $x \le 1$. The unconstrained minimum is $x = 2$, so the constraint is active and the answer is $x^\star = 1$ (with $f^\star = 1$). Penalised objective: $Q_\rho(x) = (x-2)^2 + \tfrac{\rho}{2}\max(0,\, x - 1)^2$.
- For $x > 1$ the slope is $Q_\rho'(x) = 2(x - 2) + \rho(x - 1)$. Set it to zero: $x(2 + \rho) = 4 + \rho$, so $x(\rho) = \dfrac{4 + \rho}{2 + \rho}$.
- Try some weights: $\rho = 2$: $x = 1.5$. $\rho = 10$: $x = 14/12 = 1.1667$. $\rho = 100$: $x = 104/102 = 1.0196$. $\rho = 1000$: $x = 1.002$.
- The solution is always on the wrong side ($x > 1$) and the violation is $x - 1 = \dfrac{2}{2 + \rho}$: it shrinks like $1/\rho$.
- The objective $f(x) = (x-2)^2$ at these points is $0.25,\ 0.694,\ 0.961,\ 0.996$: always below $f^\star = 1$. Cheating pays, until the fine gets large enough.
- The fine measures the multiplier. The pushback force is $\rho\,(x(\rho) - 1) = \dfrac{2\rho}{2 + \rho}$: that is $1,\ 1.67,\ 1.96,\ 1.996$, heading to $2$. And $2$ is exactly the KKT multiplier: at $x = 1$, $f'(1) = -2$ must be balanced by $\mu \cdot g' = \mu \cdot 1$, so $\mu = 2$.
On our running problem ($c = (2,1)$, $x_1 + x_2 \le 2$): the penalised minimiser is $x(\rho) = (2, 1) - \frac{\rho}{2(1+\rho)}(1, 1)$, so $\rho = 1$ gives $(1.75, 0.75)$, $\rho = 10$ gives $(1.545, 0.545)$, $\rho = 100$ gives $(1.505, 0.505)$: all outside, all sliding toward $(1.5, 0.5)$. The violation $x_1 + x_2 - 2$ is exactly $1/(1+\rho)$ and the push-back $\rho/(1+\rho) \to 1 = \mu$.
For $\min f(x)$ s.t. $g_i(x) \le 0$ ($i = 1..m$), $h_j(x) = 0$ ($j = 1..p$), the quadratic penalty function is
$$Q_\rho(x) = f(x) + \frac{\rho}{2}\sum_{i=1}^{m}\max\big(0,\, g_i(x)\big)^2 + \frac{\rho}{2}\sum_{j=1}^{p} h_j(x)^2, \qquad \rho > 0.$$(Many books call the weight $\mu$; we use $\rho$ because $\mu_i$ already means the KKT multiplier of an inequality.) The penalty method: solve $\min_x Q_\rho(x)$ for $\rho_1 < \rho_2 < \dots$ (say $\rho_{k+1} = 10\rho_k$), starting each solve at the previous answer. Facts for a convex problem with an active constraint:
- The solution is infeasible but better in objective: $f(x(\rho)) \le f^\star$. Proof: the true optimum $x^\star$ is feasible, so its penalty is $0$ and $Q_\rho(x(\rho)) \le Q_\rho(x^\star) = f^\star$; and $f(x(\rho)) \le Q_\rho(x(\rho))$ because penalties are $\ge 0$. So if $x(\rho)$ were feasible it would equal the optimum.
- It converges to the constrained optimum as $\rho \to \infty$, with violation about $\mu^\star/\rho$.
- It estimates the multipliers. Stationarity of $Q_\rho$ reads $\nabla f + \sum_i \rho\max(0, g_i)\nabla g_i + \sum_j \rho h_j \nabla h_j = 0$. Compare with the KKT condition $\nabla f + \sum_i \mu_i\nabla g_i + \sum_j \lambda_j \nabla h_j = 0$: so $\mu_i \approx \rho\max(0, g_i(x(\rho)))$ and $\lambda_j \approx \rho\, h_j(x(\rho))$.
- The price is ill-conditioning. The Hessian of the penalty term at a violated constraint is $\rho\,\nabla g\,\nabla g^\top$ (plus a smaller term), which has one eigenvalue of size about $\rho\,\|\nabla g\|^2$ while all the others stay of ordinary size. So $\kappa$ grows in proportion to $\rho$.
Ill-conditioning, worked out on the running problem. Near the solution, $Q_\rho$ has Hessian $2I + \rho\, a a^\top$ with $a = (1,1)$ (so $aa^\top$ is the matrix of all ones). Its eigenvalues are $2$ (direction $(1,-1)$, along the fence) and $2 + 2\rho$ (direction $(1,1)$, across the fence). So $\kappa = 1 + \rho$.
Gradient descent with the best fixed step has error factor $\frac{\kappa-1}{\kappa+1} = \frac{\rho}{\rho + 2}$ per step, so reducing the error by $10^{-6}$ takes about $\ln(10^6)\,(\rho+2)/2 \approx 6.9\,(\rho + 2)$ steps. Counting the steps in the widget below gives about $13,\ 75,\ 680,\ 6700$ steps for $\rho = 1,\ 10,\ 100,\ 1000$. Ten times the stiffness, ten times the steps. Newton's method does not care: it rescales by the Hessian, and on a quadratic piece it needs one step.
Remedies: (1) raise $\rho$ gradually and warm-start (continuation); (2) solve the inner problems with Newton or L-BFGS, not plain gradient descent; (3) use an exact penalty $\rho\sum_i\max(0, g_i) + \rho\sum_j|h_j|$, which gives the exact answer for a finite $\rho$ (any $\rho$ bigger than the largest multiplier) but is non-smooth at the fence; (4) use the augmented Lagrangian (awareness): add a multiplier term, $f + \lambda h + \frac{\rho}{2}h^2$, and update $\lambda \leftarrow \lambda + \rho\,h(x)$. For $\min x_1^2 + x_2^2$ s.t. $x_1 + x_2 = 1$ with fixed $\rho = 1$ this gives $x_1 = 0.25,\ 0.375,\ 0.4375,\ 0.469, \dots \to 0.5$ with multiplier $\lambda = -0.5,\ -0.75,\ -0.875, \dots \to -1$: the error halves each round, and $\rho$ never has to grow.
Why do we need it?
Sometimes all you have is an unconstrained optimizer, or a constraint is too complicated to project onto. A penalty turns "obey this rule" into "pay for breaking it", so the optimizers you already know can be reused unchanged.
Where is it used?
As a soft constraint in ML itself: weight penalties, physics-informed neural networks (the physics residual is a penalty), constrained fine-tuning, fairness penalties; and as a building block inside augmented-Lagrangian solvers (LANCELOT, ALGENCAN) and ADMM.
How is it used?
Add $\frac{\rho}{2}\max(0, g)^2$ (or $\frac{\rho}{2}h^2$) to the loss. Start with a moderate $\rho$, solve, multiply $\rho$ by 10, solve again from the last answer. Stop when the violation is below tolerance. Prefer an inner solver that tolerates bad conditioning (L-BFGS or Newton).
A penalty solution is not feasible. If the rule is a hard one (a probability must be non-negative, a dimension must be positive), a "slightly outside" answer can be unusable. Then project the result, or use a barrier or a projection method instead.
Do not jump straight to $\rho = 10^{9}$. The inner problem becomes so ill-conditioned that the optimizer stalls or loses digits (Chapter 3.16). Raise $\rho$ in stages.
The squared penalty is smooth but not twice differentiable at the fence ($\max(0, g)^2$ has a continuous slope but a jump in curvature). That is harmless for gradient methods and for most Newton line searches, and it is why the exact penalty (kink) is harder.
Quick check: for $\min (x-2)^2$ s.t. $x \le 1$ and $\rho = 18$, what is $x(\rho)$, and what multiplier does the penalty method estimate?
$x(18) = (4 + 18)/(2 + 18) = 22/20 = 1.1$. Violation $= 0.1$. Estimated multiplier $= \rho\cdot 0.1 = 1.8$ (the true value is $2$; the estimate is $2\rho/(2+\rho)$, so it is always a bit low).
Quick check: why is the penalty solution's objective below the true optimum?
It is allowed to break the rule, and breaking the rule helps the objective. The penalty (a charge we added, not part of $f$) is the only thing holding it back. Remove the charge and measure $f$ alone, and you see a value $\le f^\star$.
Barrier methods: an invisible wall inside the fence taught here
Now build the wall on the inside. Imagine a force field in the room that is almost zero in the middle and grows stronger and stronger as you walk toward a wall, becoming infinitely strong at the wall itself. You can never reach the wall. You always stay strictly inside.
A knob $t$ sets how weak the field is. Turn the field down (raise $t$) and you can stand closer to the wall. As $t \to \infty$ you can stand as close as you like. The sequence of resting places as the knob turns is a curve called the central path: it starts in the middle of the allowed region and ends at the constrained optimum.
Compare with the penalty method: the penalty wall stands outside and you stand slightly beyond it; the barrier wall stands inside and you stay in front of it. This idea is called the interior-point method.
One variable. Minimise $(x - 2)^2$ subject to $x \le 1$, i.e. $g(x) = x - 1 \le 0$. The barrier term is $-\frac{1}{t}\ln(-g) = -\frac{1}{t}\ln(1 - x)$, defined only for $x < 1$; it $\to +\infty$ as $x \to 1^{-}$. Let $s = 1 - x > 0$ be the slack (the distance to the wall).
- Slope of $B_t(x) = (x-2)^2 - \frac1t\ln(1-x)$: $B_t'(x) = 2(x - 2) + \dfrac{1}{t(1 - x)}$. With $x = 1 - s$: $2(-1 - s) + \dfrac{1}{ts} = 0$.
- Multiply everything by $ts$: $-2ts - 2ts^2 + 1 = 0$, i.e. $2ts^2 + 2ts - 1 = 0$. The positive root is $s = \dfrac{-t + \sqrt{t^2 + 2t}}{2t}$.
- Numbers: $t = 1$: $s = 0.366$, $x = 0.634$. $t = 10$: $s = 0.0477$, $x = 0.9523$. $t = 100$: $s = 0.00498$, $x = 0.9950$. $t = 1000$: $x = 0.9995$. Always inside ($x < 1$), creeping toward $x^\star = 1$.
- The gap. At $t = 10$, $f(x) = (0.9523 - 2)^2 = 1.0977$, so $f - f^\star = 0.0977$. The theory below promises it is at most $m/t = 1/10 = 0.1$ ($m = 1$ constraint). ✓.
- The multiplier estimate. $\mu(t) = -\dfrac{1}{t\,g(x)} = \dfrac{1}{ts}$: $2.73,\ 2.10,\ 2.01,\ 2.001$ for $t = 1, 10, 100, 1000$. Again the true KKT multiplier is $2$.
On the running problem ($c = (2,1)$, $x_1 + x_2 \le 2$) the slack $s = 2 - x_1 - x_2$ solves $t s^2 + t s - 1 = 0$, so $s = \frac{-t + \sqrt{t^2 + 4t}}{2t}$ and $x(t) = (2, 1) - \frac{1}{2ts}(1, 1)$. For $t = 1$: $s = 0.618$, $x = (1.191, 0.191)$. For $t = 10$: $x = (1.454, 0.454)$. For $t = 100$: $x = (1.495, 0.495)$. All strictly inside, heading to $(1.5, 0.5)$.
For inequality constraints $g_i(x) \le 0$ ($i = 1..m$), the logarithmic barrier is $\varphi(x) = -\sum_{i=1}^{m}\ln\big(-g_i(x)\big)$, defined on the strict interior $\{x : g_i(x) \lt 0 \text{ for all } i\}$, where it is smooth, and $+\infty$ as any $g_i \to 0^{-}$. The barrier problem is
$$\min_x\ B_t(x) = f(x) - \frac{1}{t}\sum_{i=1}^{m}\ln\big(-g_i(x)\big), \qquad t > 0.$$Its minimiser $x^\star(t)$ is a point on the central path ($t \to \infty$ gives the constrained optimum). The interior-point method is: start at a strictly feasible $x$ and a moderate $t$; (1) minimise $B_t$ with Newton's method, starting from the current $x$; (2) stop if $m/t \lt \varepsilon$; (3) raise $t \leftarrow \beta t$ (a typical $\beta$ is 10 to 50) and repeat.
The duality-gap bound (for convex $f$ and $g_i$): $\;f\big(x^\star(t)\big) - p^\star \le \dfrac{m}{t}$, where $p^\star$ is the optimal value. So you know how far from optimal you are, and you can stop as soon as $m/t$ is small enough.
Derivation of the bound. (Needs the Lagrangian from Chapters 3.7 and 3.10.)
- At the minimiser, $\nabla B_t = 0$: $\ \nabla f(x^\star(t)) + \sum_i \dfrac{-1}{t\,g_i(x^\star(t))}\nabla g_i(x^\star(t)) = 0$. Define $\mu_i(t) = -\dfrac{1}{t\,g_i(x^\star(t))}$. Because $g_i \lt 0$, every $\mu_i(t) \gt 0$.
- So $\nabla f + \sum_i \mu_i(t)\nabla g_i = 0$ at $x^\star(t)$: it is a stationary point of the Lagrangian $L(x, \mu) = f(x) + \sum_i \mu_i g_i(x)$. For convex $f$ and $g_i$ and $\mu \ge 0$, $L(\cdot, \mu)$ is convex, so a stationary point is a minimiser of $L$.
- Hence the dual function (Chapter 3.10) at $\mu(t)$ is $d(\mu(t)) = L(x^\star(t), \mu(t)) = f(x^\star(t)) + \sum_i \mu_i g_i = f(x^\star(t)) - \dfrac{m}{t}$, because every product $\mu_i g_i = -1/t$.
- Weak duality says any dual value is a lower bound: $d(\mu) \le p^\star$. So $f(x^\star(t)) - m/t \le p^\star$, i.e. $f(x^\star(t)) - p^\star \le m/t$. ∎
- Check with the 1-D example at $t = 10$: dual value $= 1.0977 - 0.1 = 0.9977 \le p^\star = 1$ ✓.
Why Newton? As $t \to \infty$ the barrier term's Hessian, $\frac{1}{t}\sum_i \frac{\nabla g_i \nabla g_i^\top}{g_i^2}$, grows large near the wall (since $g_i \to 0$), so $B_t$ becomes badly conditioned, just like the penalty. Gradient descent would crawl. Newton's method is unaffected by such scaling, and the log barrier has an extra property (self-concordance) that makes the number of Newton steps small and predictable. A solver typically needs a few dozen Newton steps in total, almost independent of the problem size. That is why interior-point methods are the standard for linear, quadratic, second-order-cone and semidefinite programs.
Why do we need it?
For constrained convex problems you often want a feasible answer with a certificate of how close to optimal it is. The barrier method keeps every iterate strictly inside, and its duality gap $m/t$ is known at every step.
Where is it used?
The solvers behind linear, quadratic and semidefinite programming (CVXOPT, ECOS, Clarabel, MOSEK, Gurobi's barrier method, SciPy's linprog(method="highs-ipm")), and through them the engines under CVXPY. Also in nonlinear solvers such as IPOPT, and in portfolio optimization, optimal control, and some SVM and robust-regression solvers.
How is it used?
You rarely write it yourself: describe the problem in a modelling tool (CVXPY) and call an interior-point solver. Conceptually: start strictly inside, minimise $f - \frac1t\sum\ln(-g_i)$ with Newton, multiply $t$ by 10, repeat until $m/t$ is below your tolerance.
You must start strictly inside. The logarithm of zero or a negative number does not exist. If you have no strictly feasible point, a first "phase I" problem finds one. If the interior is empty (for example, an equality hidden as two inequalities, $g \le 0$ and $-g \le 0$), the plain log barrier fails. Equality constraints $Ax = b$ are handled inside Newton's linear system instead.
The gap bound needs convexity. For non-convex problems the barrier method still produces a sequence of strictly feasible points, but "$f - p^\star \le m/t$" is no longer a theorem.
Inner solves must be Newton-like. Using plain gradient descent on $B_t$ for large $t$ reproduces the ill-conditioning trap from the penalty method.
Quick check: a linear program has $m = 200$ inequality constraints. You stop the barrier method at $t = 10^{6}$. What is guaranteed about the answer?
The objective is within $m/t = 200/10^6 = 2\times 10^{-4}$ of the true optimum, and the point is strictly feasible (every constraint holds).
Quick check: for $\min (x-2)^2$ s.t. $x \le 1$, compute the barrier solution at $t = 2$ to three decimals.
$s = \dfrac{-t + \sqrt{t^2 + 2t}}{2t} = \dfrac{-2 + \sqrt{8}}{4} = \dfrac{0.8284}{4} = 0.2071$, so $x = 0.793$. Check: $B_t'(x) = 2(0.793 - 2) + \dfrac{1}{2\cdot 0.2071} = -2.414 + 2.414 = 0$ ✓.
Two ways to build a wall: penalty against barrier
Both methods turn a hard rule into a soft cost that you can minimise with an ordinary optimizer, and both need you to turn a knob to infinity. They differ in which side of the fence the wall stands on.
- The penalty wall is a rubber sheet outside. It is flat (zero) inside the allowed region and rises after the fence. You can lean on it.
- The barrier wall is a force field inside. It is gentle in the middle and rises to infinity at the fence. You can never touch the fence.
Picture the objective as a landscape. Adding a penalty bends the landscape up beyond the fence; adding a barrier makes a cliff that rises toward the fence from the inside.
A sandwich around the truth. Take the one-variable problem $\min (x-2)^2$ s.t. $x \le 1$, whose true value is $f^\star = 1$.
- Penalty, $\rho = 10$: $x = 1.1667$ (outside) and $f(x) = 0.694$. This is below $f^\star$.
- Barrier, $t = 10$: $x = 0.9523$ (inside) and $f(x) = 1.098$. This is above $f^\star$.
For convex problems this always happens: the penalty solution gives a value $\le f^\star$ (it cheats), the barrier solution gives a value $\ge f^\star$ (it is a feasible point, so it cannot beat the optimum). So $0.694 \le 1 \le 1.098$. The truth is sandwiched, and both squeeze in as the knob grows.
The two methods side by side (objective $f$, constraint $g(x) \le 0$):
| Penalty (exterior) | Barrier (interior) | |
|---|---|---|
| Modified objective | $f + \frac{\rho}{2}\max(0, g)^2$ | $f - \frac{1}{t}\ln(-g)$ |
| Where the solution lives | slightly outside the allowed set | strictly inside |
| Knob toward the answer | $\rho \to \infty$ | $t \to \infty$ |
| Starting point | anywhere (infeasible is fine) | must be strictly feasible |
| Error control | violation $\approx \mu^\star/\rho$; no simple certificate | duality gap $\le m/t$: a certificate |
| Smoothness | slope continuous, curvature jumps at the fence | infinitely smooth inside |
| Conditioning as the knob grows | worse: $\kappa \propto \rho$ | worse near the wall; fixed by Newton |
| Equality constraints | easy: $\frac{\rho}{2}h^2$ | not directly (use Newton's KKT system) |
| Typical use | soft constraints in ML, simple codes, augmented Lagrangian | LP / QP / SDP solvers (interior-point) |
Why do we need it?
You need to know which wall to build. An unsafe answer (slightly outside) is fine for a soft constraint but not for a hard physical limit, and a feasible answer with a guarantee is worth a more careful solver.
Where is it used?
Penalty-style soft constraints in neural-network losses (physics-informed networks, fairness or norm penalties); barrier-style in convex solvers (CVXPY with ECOS, Clarabel, MOSEK) and in nonlinear programming (IPOPT).
How is it used?
Soft rule, any optimizer, no feasible start: use a penalty and check the violation afterwards. Hard rule, convex problem, certificate wanted: use a barrier (interior-point) solver. If the set is a box, ball or simplex: skip both and use projected gradient.
Neither method is exact at any finite knob value (except the exact penalty, which has a kink). Always report the violation (penalty) or the gap bound $m/t$ (barrier), not just the final point.
The knob schedule matters. Increase $\rho$ or $t$ by a factor of 5 to 50 per round, warm-starting each round from the last. Too small a factor wastes rounds; too large a factor leaves the inner solver far from the new minimiser.
Quick check: you run a penalty method and a barrier method on the same convex problem and get objective values $0.94$ and $1.03$. What do you know about $f^\star$?
$0.94 \le f^\star \le 1.03$. The penalty value is a lower bound (it is allowed to cheat) and the barrier value is an upper bound (it is a feasible point). The true optimum is trapped between them.
Which method for which problem? decision guide
Choosing an optimizer is like a doctor's triage. You do not memorise a treatment per patient. You ask a few quick questions in a fixed order, and the answers narrow the field:
- Are there rules the answer must obey? (constraints). If yes, that decides the family.
- Is the objective smooth? If not, is the rough part a simple known piece (like L1)?
- How big is the problem? Huge data means noisy mini-batch gradients; a modest problem means exact gradients.
- What can you afford? Only gradients, or also the Hessian?
Structure first, then size, then cost.
Eight real problems, eight choices.
| Problem | Questions it answers | Good choice |
|---|---|---|
| Logistic regression, 10 000 examples, 50 features, L2 penalty, full batch | smooth, no constraints, small $n$, exact gradient | L-BFGS (scikit-learn's default), or Newton/IRLS; typically tens of iterations |
| Training a ResNet on a million images | smooth (almost), huge data, noisy, $n$ in millions | Mini-batch SGD with momentum (or AdamW) with a learning-rate schedule |
| Fine-tuning a Transformer language model | same, plus very different gradient scales across layers | AdamW with warm-up and cosine decay, gradient clipping |
| Lasso regression with $10^5$ features | smooth loss + non-smooth L1, no constraints | Coordinate descent or proximal gradient (ISTA/FISTA) |
| Portfolio weights: non-negative, sum to 1, minimise risk | constraint set is a simplex (cheap projection) | Projected gradient (or a QP solver) |
| A linear program or quadratic program with thousands of constraints | convex, general linear constraints | Interior-point (barrier) solver |
| Solving $(X^\top X + \lambda I)w = X^\top y$ for a huge sparse $X$ | a symmetric positive definite linear system | Conjugate gradient (preconditioned) |
| Robot trajectory with 100 variables and a few nonlinear constraints | smooth, non-convex, general constraints | SQP or an interior-point NLP solver such as IPOPT |
The rules in order (a guide, not a theorem; the last word is always a test run):
- Constrained? Simple set with a cheap projection (box, ball, simplex): projected gradient (projected stochastic gradient if the data is huge; coordinate descent for a separable box). A convex problem with general constraints: barrier / interior-point. Smooth, non-convex, general constraints: SQP or an interior-point NLP solver, or (for big soft cases) a penalty / augmented Lagrangian.
- Unconstrained but not smooth? Smooth + a simple known non-smooth term (L1, a constraint indicator): proximal gradient or coordinate descent. Otherwise subgradient methods (awareness: slow but general).
- A quadratic / symmetric positive definite linear system? Conjugate gradient (with a preconditioner).
- Huge data or noisy gradients? SGD with momentum or Adam / AdamW, with a learning-rate schedule.
- Exact gradients, Hessian affordable ($n$ up to a few thousand): Newton (with line search or trust region).
- Exact gradients, Hessian too big: L-BFGS (or nonlinear CG if memory is extremely tight).
Why do we need it?
A wrong choice costs days: L-BFGS on a noisy mini-batch loss misbehaves, plain gradient descent on a badly scaled smooth problem takes thousands of steps, Newton on a 100-million-parameter model cannot even start.
Where is it used?
Every project. scipy.optimize.minimize already makes the first choice for you (BFGS with no constraints, L-BFGS-B with bounds, SLSQP with general constraints); deep-learning libraries default to AdamW or SGD with momentum; CVXPY picks an interior-point solver.
How is it used?
Write down the objective and the constraints, answer the four questions, pick the method, run it with a sensible default, and check: plot the loss curve, watch the gradient norm, test a second method. Treat the rules as a good first guess.
Deep-learning starting points (rules of thumb that many teams start from, not laws; always tune on your own task):
- Transformers and language models: AdamW, $\beta_1 = 0.9$, $\beta_2$ between $0.95$ and $0.999$, weight decay around $0.01$ to $0.1$, a linear warm-up over the first hundreds or thousands of steps, then cosine decay, gradient clipping at norm $1$.
- Convolutional networks (ResNet-style): SGD with momentum $0.9$, a learning rate near $0.1$ for a batch of 256 (scaled with the batch size), weight decay $10^{-4}$ to $5\times 10^{-4}$, step or cosine decay.
- Small classical models (logistic regression, linear SVM, GLMs): L-BFGS or a coordinate/dual method, not Adam.
- If training diverges or is erratic: lower the learning rate first, add warm-up, clip gradients, then look at the data.
The questions are guides, not walls. Adam also works on smooth small problems; L-BFGS sometimes works on mildly stochastic losses (with large batches); an interior-point solver can handle a box. The tree gives a good first choice and a reason; a quick experiment (the race below) is the final judge.
"Convex" is a promise about the problem, not about the data size. A tiny non-convex problem still has local minima that Newton and L-BFGS happily settle into. For deep networks no choice of optimizer removes that (Chapter 3.15).
Quick check: a smooth, unconstrained problem with $n = 200$ parameters, exact gradients, full batch. Which method and why?
$n = 200$ means the Hessian has $40\,000$ numbers: affordable. So Newton (with a line search) or its cousin BFGS, which converges in a few dozen iterations. L-BFGS would also work, but there is no need to throw away curvature information when you can store it.
Quick check: why does "huge data" move you from L-BFGS to Adam or SGD?
A full-batch gradient reads all the data, which is too slow, so you use mini-batches. Mini-batch gradients are noisy, and L-BFGS's line search and curvature pairs $(s_k, y_k)$ are corrupted by that noise. SGD-type methods are designed for noisy gradients, and each step is cheap.
The whole toolbox in one table
You have now met every algorithm in the syllabus. This section is the reference page you come back to: one row per method, with the information it needs, its cost per step, its extra memory and its best use. Use the filters to focus on one family, or on the methods that survive at the scale of modern models.
Reading a row. Take L-BFGS: family "second-order (quasi-Newton)"; it needs gradients and an exact line search; each step costs about $4mn$ operations and it stores $2mn$ numbers ($m \approx 5$ to $20$); it scales to millions of parameters if the gradients are exact; it is best for smooth, full-batch problems such as logistic regression. Compare Adam: first-order; needs (noisy) gradients; $O(n)$ per step; 2 extra vectors; scales to billions of parameters; best for deep networks.
Columns: needs = what you must be able to compute; cost per step in operations for $n$ unknowns (and $B$ = batch size, $m$ = L-BFGS memory); extra memory = numbers stored beyond $x$; scale = the size at which the method is routinely used. "Rule" columns are summaries: details are in the chapters linked in the last column.
Why do we need it?
Twenty-one names are easy to confuse. A single table puts cost, memory and requirements side by side, so a choice becomes a lookup.
Where is it used?
Planning a training run, reviewing a colleague's choice of optimizer, deciding whether a method can fit in GPU memory, and revision before an interview.
How is it used?
Filter by family, then compare cost per step with the number of steps you expect, and check the memory column against your hardware. Then follow the link to the chapter that explains the method.
Every number in the cost and scale columns is an order of magnitude, and "scale" is a rule of thumb about common practice, not a hard limit. A sparse interior-point solver, for instance, handles linear programs with millions of variables, because $n^2$ then really means "the number of non-zeros in a sparse matrix".
Quick check: which methods in the table need no gradient of $f$ at all, only matrix-vector products with a fixed matrix $A$?
Conjugate gradient (on a quadratic $f = \tfrac12 x^\top A x - b^\top x$ the gradient $Ax - b$ is a matrix-vector product, which is why CG only ever asks for $Av$).
Run them all: the scoreboard
Words and tables can only take you so far. The last test is a race: ten optimizers start from the same point on the same landscape and run until the gradient is tiny (or give up). Watch the paths; then read the scoreboard, which counts what really costs money: iterations, function evaluations, gradient evaluations, Hessian evaluations and memory.
Fewer iterations is not the same as less work. A Newton iteration needs a Hessian; a line-search method evaluates $f$ several times per iteration. That is why the scoreboard has several columns.
What to expect (all numbers below are from this very widget, tolerance $\|\nabla f\| \lt 10^{-4}$, at most 3000 iterations):
- Quadratic bowl, $\kappa = 10$: gradient descent 63 iterations, momentum 25, Newton 1, BFGS 3, L-BFGS 5, and linear CG with exact steps exactly 2 (the dimension). Nonlinear CG with a rough line search needs 22: the $n$-step guarantee is only for exact steps.
- Rosenbrock valley: Newton 21, BFGS 33, L-BFGS 36 iterations. Plain gradient descent has not converged after 3000. Nonlinear CG with the simple backtracking line search used here needs 661 iterations (and thousands of function evaluations), because that line search is crude. (SciPy, with a proper Wolfe line search: BFGS 32, L-BFGS-B 36, nonlinear CG 35, as a cross-check.)
- Himmelblau (4 minima): with default settings momentum ends in a different minimum than the others.
How the scoreboard counts. Iterations: steps until $\|\nabla f(x)\| \lt 10^{-4}$ (or "not reached" after 3000). $f$ / $\nabla f$ / $\nabla^2 f$ evaluations: how many times the method called each (line searches call $f$ many times; Newton calls the Hessian once per step). Extra memory: numbers stored beyond $x$, for $n$ unknowns ($n = 2$ in this picture, but read the formula: $n$, $2n$, $n^2$, $2mn$). Learning rates of the six first-order methods were tuned once per landscape; the slider multiplies all of them, so you can see how fragile they are.
Why do we need it?
A claim like "BFGS needs fewer iterations than Adam" hides the cost of an iteration. A scoreboard with counts of function, gradient and Hessian calls lets you compare methods on work, not on slogans.
Where is it used?
Every benchmark paper, every solver comparison (for example SciPy's methods on a test function), and the sanity check an engineer runs before committing to an optimizer for a long job.
How is it used?
Run the candidate methods on a small version of your real problem. Count gradient and function evaluations (that is what costs time), note the final value, and re-run from two or three different starting points before trusting any ranking.
One race is an anecdote. Move the start point and re-run. The ranking of first-order methods changes with the learning rates, and the winner on a smooth 2-D toy problem says little about a million-parameter noisy problem. What carries over is the pattern: curvature methods win on iterations and lose on cost per iteration; adaptive methods are forgiving of scaling; and on non-convex landscapes the method and the start decide which minimum you reach.
"Iterations to tolerance" is not available in deep learning. There you watch the validation loss and stop early. The scoreboard is a classical-optimization tool.
Quick check: on the quadratic bowl with $\kappa = 100$, Newton takes 1 iteration but each iteration needs a Hessian. When would gradient descent still be the cheaper choice?
When $n$ is large. A Newton step needs about $n^3/3$ operations and $n^2$ memory; gradient descent needs about $n$ per step even if it takes hundreds of steps. For $n = 10^6$ the Newton step is out of reach and 700 cheap steps are fine.
Recap, cheat sheet and practice
- Four families, one question each. First-order (gradient only): GD, SGD, mini-batch SGD, momentum, Nesterov, AdaGrad, RMSProp, Adam, AdamW. Second-order (curvature): Newton, quasi-Newton, BFGS, L-BFGS. Specialised (use structure): coordinate descent, proximal gradient, conjugate gradient. Constrained: Lagrange multipliers, KKT, projected gradient, penalty, barrier.
- Cost against iterations. First-order: cheap steps ($O(n)$), many of them. Newton: few steps but $n^2$ memory and $n^3$ time. L-BFGS: $2mn$ memory, tens of steps. This trade-off is the main reason deep learning uses Adam and SGD.
- Conjugate gradient on $\tfrac12 x^\top A x - b^\top x$: $\alpha_k = \dfrac{r_k^\top r_k}{p_k^\top A p_k}$, $\beta_k = \dfrac{r_{k+1}^\top r_{k+1}}{r_k^\top r_k}$, $p_{k+1} = r_{k+1} + \beta_k p_k$. Directions are $A$-conjugate ("perpendicular after the stretch"), exact in at most $n$ steps (at most as many as distinct eigenvalues), the error bound shrinks by a factor $\frac{\sqrt\kappa - 1}{\sqrt\kappa + 1}$ per step (steepest descent: $\frac{\kappa-1}{\kappa+1}$). Nonlinear CG: Fletcher–Reeves and Polak–Ribière $\beta$ with a line search.
- Projected gradient: $x \leftarrow \Pi_C(x - \eta\nabla f(x))$. Projection = nearest point (clip for a box, rescale for a ball, shift-and-clip for the simplex). Needs a cheap projection. Same $O(1/k)$ rate as GD for convex $L$-smooth $f$; the fixed-point test is the gradient mapping.
- Penalty: minimise $f + \frac\rho2\sum\max(0,g_i)^2 + \frac\rho2\sum h_j^2$ for growing $\rho$. The solution is slightly outside (violation $\approx \mu^\star/\rho$), its objective is $\le f^\star$, $\rho\max(0,g_i)$ estimates $\mu_i$, and $\kappa$ grows like $\rho$. Raise $\rho$ gradually; use Newton or L-BFGS inside; augmented Lagrangian avoids $\rho \to \infty$.
- Barrier: minimise $f - \frac1t\sum\ln(-g_i)$. Strictly feasible, central path $x^\star(t)$, $\mu_i(t) = -1/(t g_i)$, and the guarantee $f(x^\star(t)) - p^\star \le m/t$ (convex). Newton inside; this is the interior-point method of LP/QP/SDP solvers.
- Choosing: constraints first (simple set: projection; convex general: interior-point; smooth non-convex: SQP/IPOPT), then smoothness (L1: proximal/coordinate), then size (huge data: AdamW or SGD with momentum), then affordability (Hessian: Newton; otherwise L-BFGS). A quadratic system: preconditioned CG. Always verify with a test run.
Cheat sheet
| Method | Core formula | Key facts | Use when |
|---|---|---|---|
| Gradient descent | $x \leftarrow x - \eta\nabla f$ | $\eta \lt 2/L$; best-step rate $\frac{\kappa-1}{\kappa+1}$ | baseline, smooth |
| Momentum / Nesterov | $v \leftarrow \beta v - \eta g$; $x \leftarrow x + v$ | Nesterov takes $g$ at $x + \beta v$ | narrow valleys |
| Adam / AdamW | $x \leftarrow x - \eta\,\hat m/(\sqrt{\hat v} + \epsilon)$ (+ decay $\eta\lambda x$) | first step $\approx \eta$ per coordinate | deep networks |
| Newton | $x \leftarrow x - H^{-1}g$ | exact on a quadratic; $n^2$ memory | small smooth problems |
| BFGS / L-BFGS | $d = -C_k g$; $C$ from $(s, y)$ pairs | $n^2$ / $2mn$ memory | smooth, exact gradients |
| Coordinate descent | $x_i \leftarrow \arg\min_z f$ | no learning rate | Lasso, SVM duals |
| Proximal gradient | $x \leftarrow \operatorname{prox}_{\eta g}(x - \eta\nabla f)$ | soft threshold $S_t(v)$ for L1 | smooth + simple non-smooth |
| Conjugate gradient | $\alpha = \frac{r^\top r}{p^\top A p}$, $\beta = \frac{r_+^\top r_+}{r^\top r}$ | $\le n$ steps; $\sqrt\kappa$ rate; $p_i^\top A p_j = 0$ | SPD systems, quadratics |
| Projected gradient | $x \leftarrow \Pi_C(x - \eta\nabla f)$ | $O(1/k)$; fixed point = optimum | box, ball, simplex |
| Penalty | $f + \frac\rho2\sum\max(0,g_i)^2$ | outside; violation $\sim 1/\rho$; $\kappa \sim \rho$; $\rho\max(0,g) \approx \mu$ | soft constraints, simple code |
| Barrier | $f - \frac1t\sum\ln(-g_i)$ | inside; gap $\le m/t$; $\mu_i = -1/(tg_i)$ | convex programs (LP, QP, SDP) |
| Decision order | constraints $\to$ smoothness $\to$ data size $\to$ Hessian affordable | ||
import numpy as np
from scipy.optimize import minimize, rosen, rosen_der
from scipy.sparse.linalg import cg as scipy_cg
# ---------- 1. Conjugate gradient from scratch (the example of this chapter) ----------
def conjugate_gradient(A, b, x0=None, tol=1e-10):
x = np.zeros_like(b) if x0 is None else x0.copy()
r = b - A @ x # residual = -gradient
p = r.copy() # first direction: steepest descent
rs = r @ r
steps = []
for k in range(len(b)): # at most n steps
Ap = A @ p
alpha = rs / (p @ Ap) # exact line search along p
x = x + alpha * p
r = r - alpha * Ap # cheap residual update
rs_new = r @ r
beta = rs_new / rs
steps.append((alpha, beta, x.copy()))
if np.sqrt(rs_new) < tol:
break
p = r + beta * p # new direction, conjugate to the old ones
rs = rs_new
return x, steps
A = np.array([[3.0, 1.0], [1.0, 2.0]]); b = np.array([1.0, 2.0])
x, steps = conjugate_gradient(A, b)
print(np.round([s[0] for s in steps], 4), round(steps[0][1], 4)) # [0.3333 0.6 ] 0.1111
print(np.round(x, 6), np.linalg.solve(A, b)) # [0. 1.] [0. 1.]
# conjugacy check: p0^T A p1 = 0
p0 = b.copy(); r1 = b - A @ (steps[0][0] * p0); p1 = r1 + steps[0][1] * p0
print(round(float(p0 @ A @ p1), 12)) # 0.0
# ---------- 2. CG against gradient descent when kappa is large ----------
n = 200
lam = np.linspace(1, 1000, n); Q, _ = np.linalg.qr(np.random.default_rng(0).normal(size=(n, n)))
A = Q @ np.diag(lam) @ Q.T; A = (A + A.T) / 2; b = np.ones(n)
count = {'k': 0}
def cb(xk): count['k'] += 1
xs, info = scipy_cg(A, b, rtol=1e-8, callback=cb) # (older SciPy: tol=1e-8)
print('CG iterations:', count['k'], ' kappa =', round(lam[-1] / lam[0])) # CG iterations: 87 kappa = 1000
# steepest descent with exact line search needs far more:
x = np.zeros(n); k = 0
while np.linalg.norm(b - A @ x) > 1e-8 * np.linalg.norm(b) and k < 200000:
r = b - A @ x; x = x + (r @ r) / (r @ A @ r) * r; k += 1
print('steepest descent iterations:', k) # steepest descent iterations: 7842
# ---------- 3. Projected gradient: box and simplex ----------
def project_simplex(y):
u = np.sort(y)[::-1]; css = np.cumsum(u) - 1
rho = np.nonzero(u - css / (np.arange(len(y)) + 1) > 0)[0][-1]
return np.maximum(y - css[rho] / (rho + 1), 0)
print(project_simplex(np.array([0.8, 0.6, -0.2]))) # [0.6 0.4 0. ]
c = np.array([3.0, 1.0]); x = np.zeros(2)
for k in range(6):
x = np.clip(x - 0.25 * 2 * (x - c), 0, 2) # step, then clip to the box [0,2]^2
print(x) # [2. 0.984375] (heading to (2, 1))
# ---------- 4. Penalty and barrier on the running problem ----------
c = np.array([2.0, 1.0]); g = lambda x: x[0] + x[1] - 2 # allowed: g(x) <= 0
f = lambda x: np.sum((x - c) ** 2)
x_pen = c.copy()
for rho in [1, 10, 100, 1000]: # raise the penalty weight, warm start
Q = lambda x, r=rho: f(x) + 0.5 * r * max(0.0, g(x)) ** 2
x_pen = minimize(Q, x_pen, method='BFGS', options={'gtol': 1e-10}).x
print('penalty', rho, np.round(x_pen, 4), 'violation', round(g(x_pen), 4))
# penalty 1 [1.75 0.75] violation 0.5 / penalty 10 [1.5455 0.5455] violation 0.0909
# penalty 100 [1.505 0.505] violation 0.0099 / penalty 1000 [1.5005 0.5005] violation 0.001
x_bar = np.array([0.0, 0.0]) # strictly feasible start (Nelder-Mead keeps this short; real solvers use Newton)
for t in [1, 10, 100, 1000]:
B = lambda x, t=t: f(x) - np.log(-g(x)) / t if g(x) < 0 else np.inf
x_bar = minimize(B, x_bar, method='Nelder-Mead', options={'xatol': 1e-12, 'fatol': 1e-14, 'maxiter': 4000}).x
print('barrier', t, np.round(x_bar, 4), 'gap', round(f(x_bar) - 0.5, 4), '<= m/t =', 1 / t)
# barrier 1 [1.191 0.191] gap 0.809 <= m/t = 1.0 / barrier 10 [1.4542 0.4542] gap 0.0958 <= m/t = 0.1
# barrier 100 [1.495 0.495] gap 0.0100 <= m/t = 0.01 / barrier 1000 [1.4995 0.4995] gap 0.001 <= m/t = 0.001
# the "real" answer from SciPy's constrained solver (SLSQP; scipy's 'ineq' means fun(x) >= 0)
res = minimize(f, np.zeros(2), method='SLSQP', constraints=[{'type': 'ineq', 'fun': lambda x: -g(x)}])
print(np.round(res.x, 4), round(res.fun, 4)) # [1.5 0.5] 0.5
# ---------- 5. Same start, different methods (Rosenbrock) ----------
for m in ['CG', 'BFGS', 'L-BFGS-B']:
r = minimize(rosen, [-1.2, 1.0], jac=rosen_der, method=m, options={'gtol': 1e-4})
print(m, r.nit, 'iterations,', r.nfev, 'function evaluations') # CG 35, 76 / BFGS 32, 39 / L-BFGS-B 36, 44
1. Conjugate gradient is applied to $\tfrac12 x^\top A x - b^\top x$ with $A$ symmetric positive definite and $n = 50$ unknowns, in exact arithmetic. What is guaranteed?
2. In projected gradient descent, a gradient step lands outside the allowed set $C$. What do you do?
3. A quadratic penalty method is run with $\rho = 10, 100, 1000, \dots$ on a problem whose constraint is active. Which statement is right?
4. A convex problem has $m = 5$ inequality constraints. A barrier method stops at $t = 1000$. What is the guaranteed bound on $f(x) - p^\star$?
5. You must train a 500-million-parameter Transformer on streaming text. Which is the best first choice?
6. Which pairing of problem and method is the worst match?
Practice problems
A. Run conjugate gradient by hand on $A = \begin{bmatrix} 5 & 2 \\ 2 & 2 \end{bmatrix}$, $b = [3, 4]$, from $x_0 = 0$. Check that you land on the solution after two steps and that $p_0^\top A p_1 = 0$.
$r_0 = p_0 = (3, 4)$. $A p_0 = (15 + 8,\ 6 + 8) = (23, 14)$. $r_0^\top r_0 = 25$, $p_0^\top A p_0 = 3\cdot 23 + 4\cdot 14 = 125$, so $\alpha_0 = 25/125 = 0.2$ and $x_1 = (0.6, 0.8)$.
$r_1 = (3, 4) - 0.2\,(23, 14) = (-1.6, 1.2)$, $r_1^\top r_1 = 2.56 + 1.44 = 4$, so $\beta_0 = 4/25 = 0.16$ and $p_1 = r_1 + 0.16\,p_0 = (-1.12, 1.84)$.
$A p_1 = (-5.6 + 3.68,\ -2.24 + 3.68) = (-1.92, 1.44)$, $p_1^\top A p_1 = 2.1504 + 2.6496 = 4.8$, so $\alpha_1 = 4/4.8 = 5/6$ and $x_2 = (0.6, 0.8) + \tfrac56(-1.12, 1.84) = (-\tfrac13, \tfrac73)$.
Check: $A x_2 = (-\tfrac53 + \tfrac{14}{3},\ -\tfrac23 + \tfrac{14}{3}) = (3, 4)$ ✓. Conjugacy: $p_0^\top A p_1 = 3(-1.92) + 4(1.44) = -5.76 + 5.76 = 0$ ✓.
B. Project (i) $y = (3, -2)$ onto the box $[0,2]\times[-1,1]$; (ii) $y = (-3, 4)$ onto the disk of radius $2$; (iii) $y = (0.5, 1.0, 0.9)$ onto the simplex.
(i) Clip each coordinate: $x_1 = \min(2, \max(0, 3)) = 2$, $x_2 = \min(1, \max(-1, -2)) = -1$. Answer $(2, -1)$.
(ii) $\|y\| = 5 > 2$, so scale by $2/5$: $(-1.2, 1.6)$ (length $2$ ✓).
(iii) Sort descending: $1.0, 0.9, 0.5$; cumulative sums $1.0, 1.9, 2.4$. Try all three entries positive: $\tau = (2.4 - 1)/3 = 0.4667$; the smallest, $0.5 - 0.4667 = 0.0333 > 0$, so all stay positive. Answer $(0.0333, 0.5333, 0.4333)$, which sums to $1$ ✓.
C. Minimise $f(x) = (x_1 - 3)^2 + (x_2 - 1)^2$ over the unit disk $\|x\| \le 1$ with projected gradient, $\eta = 0.25$, from $x_0 = (0, 0)$. Do two steps and state the exact answer.
$\nabla f = 2(x - (3,1))$. Step 1: $\nabla f(0,0) = (-6,-2)$, $y = (1.5, 0.5)$, $\|y\| = 1.581 > 1$, so $x_1 = y/\|y\| = (0.9487, 0.3162)$.
Step 2: $\nabla f(x_1) = (-4.103, -1.368)$, $y = x_1 - 0.25\nabla f = (1.974, 0.658)$, $\|y\| = 2.081$, so $x_2 = (0.9487, 0.3162)$.
Exact answer: the projection of $c = (3, 1)$ onto the disk, $c/\|c\| = (3, 1)/\sqrt{10} = (0.9487, 0.3162)$. The iterates are already there, because here every gradient step points along the line through the origin and $c$.
D. Penalty method for $\min x^2$ subject to $x \ge 1$ (write it as $g(x) = 1 - x \le 0$). Find $x(\rho)$, the violation and the multiplier estimate; compare with the KKT multiplier.
$Q_\rho(x) = x^2 + \tfrac\rho2\max(0, 1 - x)^2$. For $x \lt 1$: $Q_\rho' = 2x - \rho(1 - x) = 0$, so $x(\rho) = \dfrac{\rho}{2 + \rho}$. Numbers: $\rho = 2$: $0.5$; $\rho = 10$: $0.833$; $\rho = 100$: $0.980$. Violation $1 - x = \dfrac{2}{2 + \rho}$.
Multiplier estimate $\rho(1 - x) = \dfrac{2\rho}{2 + \rho} \to 2$. KKT: $\nabla f + \mu\nabla g = 2x - \mu = 0$ at $x = 1$, so $\mu = 2$ ✓. The estimate is below $2$ for every finite $\rho$.
E. Barrier method for the same problem: find $x(t)$, check the gap bound at $t = 10$ and give the multiplier estimate.
$B_t(x) = x^2 - \frac1t\ln(x - 1)$ on $x \gt 1$. $B_t' = 2x - \dfrac{1}{t(x-1)} = 0$ gives $2t\,x^2 - 2t\,x - 1 = 0$, so $x(t) = \dfrac{t + \sqrt{t^2 + 2t}}{2t}$.
At $t = 10$: $x = (10 + \sqrt{120})/20 = 1.0477$ (inside, since the allowed region is $x \ge 1$). $f(x) = 1.0977$, gap $= 0.0977 \le m/t = 0.1$ ✓. Multiplier estimate $\mu = -1/(t\,g(x)) = 1/(10\cdot 0.0477) = 2.10 \to 2$ ✓.
F. (i) $\kappa = 400$ and a target error reduction of $10^{-4}$: estimate the steps of steepest descent and of CG. (ii) Pick a method: (a) Lasso with $10^5$ features, (b) a 5000-variable smooth unconstrained problem with exact gradients, (c) an SVM-style problem with a box constraint on each of $10^6$ variables.
(i) Steepest descent: $\tfrac12\kappa\ln(1/\varepsilon) = 200\cdot 9.21 \approx 1840$. CG: $\tfrac12\sqrt\kappa\ln(2/\varepsilon) = 10\cdot 9.90 \approx 99$. A factor of about $\sqrt\kappa = 20$.
(ii-a) Smooth loss plus L1: coordinate descent or proximal gradient (FISTA). (ii-b) $n = 5000$: the Hessian has $2.5\times 10^7$ numbers, affordable, so Newton or BFGS (L-BFGS if memory is tight). (ii-c) A separable box: coordinate descent on the dual (as in LIBLINEAR), or projected gradient (clip to the box).
Glossary
Every important word in this guide, in one place, explained in plain English. Type in the box to filter. Each entry links to the chapter that teaches it.
- Active constraint
- An inequality constraint that holds with equality at the current point ($g_i(x)=0$): the point is touching the boundary. Inactive constraints have slack and do not affect the local answer. 3.7 3.9
- AdaGrad
- An optimizer that gives each parameter its own step size, $\eta/\sqrt{\sum g^2}$. Great for rare features, but the step sizes only ever shrink. 3.4
- Adam
- An optimizer that combines momentum (an average of gradients) with RMSProp-style scaling (an average of squared gradients), with a bias correction for the first steps. 3.4
- AdamW
- Adam with weight decay applied directly to the weights, instead of being added to the gradient. The usual choice for training transformers. 3.4 3.15
- Backtracking line search
- Choosing a step size by starting with a large step and halving it until the objective has decreased enough. 3.3 3.12
- Barrier method
- An interior-point method for inequality constraints: add $-\tfrac1t\sum\log(-g_i)$ to the objective, which blows up at the boundary, so iterates stay strictly inside the feasible region. 3.17
- Batch gradient descent
- Gradient descent that uses the whole training set to compute every gradient. Exact but slow per step. 3.3
- BFGS
- A quasi-Newton method that builds up an estimate of the inverse Hessian from the changes in position and gradient. Needs only gradients. 3.12
- Bias correction
- The factor $1/(1-\beta^t)$ in Adam that removes the pull toward zero caused by starting the moving averages at zero. 3.4
- Bias–variance trade-off
- Simple models underfit (high bias), flexible models overfit (high variance). Regularization moves you along this trade-off. 3.11
- Central path
- The curve of minimisers of the barrier problem as $t$ grows. It runs through the inside of the feasible region to the true optimum. 3.17
- Complementary slackness
- For each inequality constraint, $\mu_i\,g_i(x)=0$: either the constraint is active ($g_i(x)=0$) or its multiplier is zero (or both). An inactive constraint always has multiplier zero. 3.9
- Condition number
- The ratio of the largest to the smallest curvature. Large values make gradient descent zig-zag and slow. 3.3 3.15 3.16
- Conjugate gradient
- An iterative method for quadratic problems (and linear systems $Ax=b$) that uses special search directions and finishes in at most $n$ steps. 3.17
- Constrained optimization
- Minimising an objective while the solution must satisfy equations or inequalities. 3.7
- Constraint qualification
- A mild technical condition on the constraints (for example that the active gradients are independent) under which KKT conditions are guaranteed to be necessary. 3.7 3.9
- Convergence rate
- How fast the error shrinks: sublinear (like $1/k$), linear (geometric, a straight line on a log plot) or quadratic (the number of correct digits doubles). 3.5
- Convex function
- A function where the chord between any two points lies above the graph. Every local minimum is a global minimum. 3.6
- Convex set
- A set that contains the whole line segment between any two of its points. 3.6
- Coordinate descent
- Minimising over one variable at a time while the others stay fixed, cycling through the variables. 3.13
- Cosine decay
- A learning-rate schedule that falls smoothly from its maximum to its minimum along half a cosine wave. 3.4
- Critical point
- A point where the gradient is zero. It may be a minimum, a maximum or a saddle. 3.2
- Decision variable
- A quantity you are free to choose in an optimization problem. In ML, the model parameters. 3.1
- Dual feasibility
- The KKT requirement that inequality multipliers are non-negative: $\mu_i\ge0$. 3.9
- Dual function
- $d(\lambda,\mu)=\min_x L(x,\lambda,\mu)$. It is concave and gives a lower bound on the optimal value. 3.10
- Dual problem
- Maximising the dual function over the multipliers. Its optimal value is a lower bound on the primal optimum. 3.10
- Duality gap
- The difference between the primal and dual values. Zero means strong duality. 3.10
- Early stopping
- Stopping training when validation error starts to rise. It acts as a form of regularization. 3.11 3.15
- Elastic Net
- A penalty that mixes L1 and L2: it keeps the sparsity of the Lasso and the stability of ridge. 3.11
- Exploding gradient
- Gradients that grow exponentially as they travel backward through many layers, causing huge, unstable updates. 3.15
- Feasible solution
- Any point inside the feasible region (not necessarily the best one). 3.1
- Finite difference
- Approximating a derivative from function values only, for example the central difference $(f(x+h)-f(x-h))/2h$. Used for gradient checking. 3.16
- First-order condition
- $\nabla f(x^\star)=0$: a necessary condition for an unconstrained optimum. 3.2
- Global optimum
- The best value over the whole feasible region. 3.1
- Gradient clipping
- Rescaling a gradient whose norm exceeds a limit $c$: $g\leftarrow g\cdot\min(1,c/\|g\|)$. 3.15
- Gradient descent
- Repeatedly stepping against the gradient: $x\leftarrow x-\eta\nabla f(x)$. 3.3
- Gradient norm
- $\|\nabla f\|$. It is zero at a critical point, so it is the usual measure of progress and stopping signal. 3.5
- Ill-conditioning
- A landscape with very different curvatures in different directions: long narrow valleys. 3.15
- Jensen's inequality
- For a convex $f$, $f(\mathbb{E}[x])\le\mathbb{E}[f(x)]$. 3.6
- KKT conditions
- The four conditions (stationarity, primal feasibility, dual feasibility, complementary slackness) that describe a constrained optimum. 3.9
- L1 regularization
- Adding $\lambda\|w\|_1$ to the loss. It pushes many weights exactly to zero. 3.11
- L2 regularization
- Adding $\lambda\|w\|_2^2$ to the loss. It shrinks all weights smoothly. 3.11
- Lagrange multiplier
- A number attached to a constraint that measures how strongly the constraint pushes on the solution. It is also the sensitivity of the optimal value to the constraint. 3.8 3.10
- Lagrangian
- $L(x,\lambda,\mu)=f(x)+\sum\lambda_jh_j(x)+\sum\mu_ig_i(x)$: the objective plus the constraints weighted by their multipliers. 3.7 3.8
- L-BFGS
- BFGS that stores only the last few update pairs, so memory grows linearly with the number of parameters. 3.12
- Learning rate
- The step size $\eta$ in gradient-based methods. 3.3
- Learning-rate schedule
- A rule that changes the learning rate during training: step decay, exponential, cosine, warmup… 3.4
- Lipschitz continuity
- A function whose output cannot change faster than $K$ times the change of its input. 3.5
- Local optimum
- The best value in some neighbourhood of a point, but maybe not the best overall. 3.1 3.2
- Loss landscape
- The surface you get by plotting the loss against the parameters. 3.15
- L-smooth (Lipschitz gradient)
- $\|\nabla f(x)-\nabla f(y)\|\le L\|x-y\|$: the curvature is at most $L$. Gradient descent is safe for $\eta<2/L$. 3.5
- Mini-batch
- A small random subset of the training data used to estimate the gradient in one step. 3.3 3.14
- Momentum
- Keeping a running velocity of past gradients so that steps build up speed along consistent directions and damp oscillations. 3.4
- Nesterov momentum
- Momentum that evaluates the gradient at the look-ahead point, which corrects overshoot earlier. 3.4
- Newton's method
- Jumping to the minimum of the local quadratic model: $x\leftarrow x-H^{-1}\nabla f$. Very fast near a solution but expensive. 3.12
- Non-convex
- Not convex: may have many local minima and saddle points. Neural-network losses are non-convex. 3.6 3.15
- Objective function
- The quantity you want to minimise or maximise. 3.1
- Optimal solution
- A feasible point with the best possible objective value, written $x^\star$. 3.1
- Overshooting
- Taking a step so large that you jump past the minimum to the other side. 3.3
- Parameter
- A number the model learns (a weight or a bias). 3.1
- Penalty method
- Turning constraints into extra cost terms: violations are punished with a growing penalty weight $\rho$ (we avoid $\mu$ because $\mu_i$ is the KKT multiplier). 3.17
- Primal feasibility
- The KKT requirement that the point itself satisfies all constraints. 3.9
- Projected gradient
- A gradient step followed by projecting the point back onto the feasible set. 3.17
- Proximal operator
- $\mathrm{prox}_{\eta g}(v)=\arg\min_x\, g(x)+\tfrac{1}{2\eta}\|x-v\|^2$: the cleverly handled step for the non-smooth part of an objective. 3.13
- Quadratic convergence
- Convergence where the error is squared at every step. Newton's method has it near a solution. 3.5 3.12
- Quasi-Newton method
- A method that approximates the Hessian (or its inverse) from gradient information, as BFGS does. 3.12
- Regularization
- Adding a penalty (or a constraint) that discourages complex solutions to improve generalization. 3.11
- Ridge regression
- Least squares with an L2 penalty: $(X^\top X+\lambda I)w=X^\top y$. 3.11
- RMSProp
- An optimizer that divides each step by a moving average of recent squared gradients. 3.4
- Saddle point
- A critical point that is a minimum in some directions and a maximum in others. 3.1 3.2 3.15
- Second-order condition
- Conditions on the Hessian at a critical point: positive definite gives a minimum, negative definite a maximum, indefinite a saddle. 3.2
- Slater's condition
- A simple constraint qualification for convex problems: there is a point where all inequality constraints hold strictly. It guarantees strong duality. 3.10
- Soft thresholding
- $\operatorname{sign}(v)\max(|v|-t,0)$: shrink toward zero and clip small values to exactly zero. The proximal operator of the L1 norm. 3.13
- Sparsity
- Having many parameters exactly equal to zero. 3.11
- Stationarity
- The KKT condition that the gradient of the Lagrangian with respect to $x$ is zero. 3.9
- Step size
- How far one update moves. In gradient descent, the learning rate. 3.3
- Stochastic gradient
- A gradient computed from a random example or mini-batch: noisy, but correct on average. 3.14
- Stochastic gradient descent (SGD)
- Gradient descent that uses stochastic gradients, so each step is cheap and noisy. 3.3 3.14
- Stopping criterion
- The rule that tells an algorithm when to stop: small gradient norm, small step, little improvement, or a step limit. 3.5
- Strict convexity
- Convexity where the chord lies strictly above the graph: the minimiser, if it exists, is unique. 3.6
- Strong convexity
- Convexity with at least a quadratic bend of size $\mu$ everywhere. It gives linear convergence for gradient descent. 3.5 3.6
- Strong duality
- The primal and dual optimal values are equal. 3.10
- Subgradient
- A generalised slope at a kink, such as any value in $[-1,1]$ for $|x|$ at 0. 3.13
- Tolerance
- The size of error you are happy to accept, used in stopping criteria. 3.5
- Vanishing gradient
- Gradients that shrink toward zero as they travel backward through many layers, so early layers stop learning. 3.15
- Warmup
- Starting training with a small learning rate and increasing it for the first steps. 3.4
- Weak duality
- The dual value is always at most the primal optimal value: $d^\star\le p^\star$. 3.10
Formula cheat sheet
The update rules and conditions you will use most, on one page. Notation: $g_i(x)\le0$ inequality constraints with multipliers $\mu_i\ge0$, $h_j(x)=0$ equality constraints with multipliers $\lambda_j$, learning rate $\eta$, penalty weight $\rho$, barrier sharpness $t$.
First-order update rules
| Method | Update | Chapter |
|---|---|---|
| Gradient descent | $x_{k+1}=x_k-\eta\,\nabla f(x_k)$ | 3.3 |
| Momentum | $v_{k+1}=\beta v_k-\eta\nabla f(x_k)$, $x_{k+1}=x_k+v_{k+1}$ | 3.4 |
| Nesterov | $v_{k+1}=\beta v_k-\eta\nabla f(x_k+\beta v_k)$, $x_{k+1}=x_k+v_{k+1}$ | 3.4 |
| AdaGrad | $G_k=G_{k-1}+g_k^2$, $x_{k+1}=x_k-\eta\,g_k/(\sqrt{G_k}+\varepsilon)$ | 3.4 |
| RMSProp | $E_k=\beta E_{k-1}+(1-\beta)g_k^2$, $x_{k+1}=x_k-\eta\,g_k/(\sqrt{E_k}+\varepsilon)$ | 3.4 |
| Adam | $m_k=\beta_1m_{k-1}+(1-\beta_1)g_k$, $v_k=\beta_2v_{k-1}+(1-\beta_2)g_k^2$, $x_{k+1}=x_k-\eta\,\dfrac{m_k/(1-\beta_1^k)}{\sqrt{v_k/(1-\beta_2^k)}+\varepsilon}$ | 3.4 |
| AdamW | Adam step, then $x\leftarrow x-\eta\lambda_{\text{wd}}\,x$ (decoupled) | 3.4 |
| Cosine schedule | $\eta_t=\eta_{\min}+\tfrac12(\eta_{\max}-\eta_{\min})\bigl(1+\cos\tfrac{\pi t}{T}\bigr)$ | 3.4 |
| Gradient clipping | $g\leftarrow g\cdot\min(1,\,c/\|g\|)$ | 3.15 |
Second-order, proximal, constrained
| Method | Update | Chapter |
|---|---|---|
| Newton | $x_{k+1}=x_k-H^{-1}\nabla f(x_k)$ (solve $Hd=-\nabla f$) | 3.12 |
| BFGS ($C_k\approx H^{-1}$) | $C_{k+1}=(I-\rho sy^\top)C_k(I-\rho ys^\top)+\rho ss^\top$, $\rho=1/y^\top s$, $d=-C_kg$ | 3.12 |
| Soft thresholding | $\operatorname{soft}(v,t)=\operatorname{sign}(v)\max(|v|-t,0)$ | 3.13 |
| Proximal gradient (ISTA) | $x_{k+1}=\operatorname{prox}_{\eta g}\bigl(x_k-\eta\nabla f(x_k)\bigr)$; Lasso: $\operatorname{soft}(\cdot,\eta\lambda)$ | 3.13 |
| Projected gradient | $x_{k+1}=\Pi_C\bigl(x_k-\eta\nabla f(x_k)\bigr)$ | 3.17 |
| Penalty method | $\min f+\tfrac\rho2\sum\max(0,g_i)^2+\tfrac\rho2\sum h_j^2$, $\rho\to\infty$ | 3.17 |
| Barrier method | $\min f-\tfrac1t\sum\log(-g_i)$, $t\to\infty$ (gap $\le m/t$) | 3.17 |
| Conjugate gradient ($Ax=b$) | $\alpha_k=\dfrac{r_k^\top r_k}{p_k^\top Ap_k}$, $\beta_k=\dfrac{r_{k+1}^\top r_{k+1}}{r_k^\top r_k}$; exact in $\le n$ steps | 3.17 |
Conditions, convexity, rates, duality
| Idea | Statement | Chapter |
|---|---|---|
| First- and second-order conditions | $\nabla f(x^\star)=0$; $\nabla^2f\succ0$ ⇒ strict local min, $\prec0$ ⇒ local max, indefinite ⇒ saddle | 3.2 |
| Convexity | $f(y)\ge f(x)+\nabla f(x)^\top(y-x)$; $\nabla^2f\succeq0$; local min = global min | 3.6 |
| $L$-smooth, $\mu$-strongly convex | $\mu I\preceq\nabla^2f\preceq LI$; $\kappa=L/\mu$; GD needs $\eta<2/L$ | 3.5 3.6 |
| Rates for gradient descent | convex, $L$-smooth: $f(x_k)-f^\star=O(1/k)$; strongly convex: $(1-\mu/L)^k$ (linear); Newton: quadratic near $x^\star$ | 3.5 |
| Lagrange (equality) | $\nabla f+\sum\lambda_j\nabla h_j=0$, $h_j=0$ | 3.8 |
| KKT | $\nabla f+\sum\lambda_j\nabla h_j+\sum\mu_i\nabla g_i=0$; $g_i\le0$, $h_j=0$; $\mu_i\ge0$; $\mu_ig_i=0$ | 3.9 |
| Duality | $d(\lambda,\mu)=\min_xL(x,\lambda,\mu)$; weak: $d^\star\le p^\star$; strong (convex + Slater): $d^\star=p^\star$ | 3.10 |
| Ridge, Lasso | $\hat w=(X^\top X+\lambda I)^{-1}X^\top y$; Lasso: soft thresholding gives exact zeros | 3.11 |
| Variance of a mini-batch gradient | $\operatorname{Var}=\sigma^2/B$ | 3.14 |
Where to go next
You finished the tour. Here is how to make it stick, and where to dig deeper.
How to make it stick
- Implement the optimizers yourself. Gradient descent, momentum and Adam are each only a few lines of NumPy. Write them, then run them on Rosenbrock's valley.
- Always plot the loss curve. It is the single best diagnostic of an optimization run. Learn what each shape means (flat, falling, spiking, noisy).
- Check KKT by hand on tiny problems. Two variables and two constraints is enough to see every condition in action.
- Compare against a trusted solver.
scipy.optimize.minimizewithBFGS,L-BFGS-BandSLSQPmakes a great reference for your own code.
Free resources
- Boyd & Vandenberghe, Convex Optimization (free PDF) and the Stanford lectures: the standard text for convexity, duality, KKT and interior-point methods.
- Nocedal & Wright, Numerical Optimization: the standard text for line search, Newton, BFGS, L-BFGS and constrained methods.
- Bottou, Curtis & Nocedal, "Optimization Methods for Large-Scale Machine Learning" (free paper): stochastic optimization for ML.
- Sebastian Ruder, "An overview of gradient descent optimization algorithms" (free article): momentum, AdaGrad, RMSProp, Adam in one place.
- Kingma & Ba, "Adam", and Loshchilov & Hutter, "Decoupled Weight Decay Regularization" (AdamW): the original papers; both are readable.
- Goodfellow, Bengio & Courville, Deep Learning, chapter 8 "Optimization for Training Deep Models".
- Beck, First-Order Methods in Optimization and Bubeck, Convex Optimization: Algorithms and Complexity: for the mathematics of convergence.
Companion guides
Optimization stands on the other two guides: the Calculus guide (gradients, Hessians, Taylor series, backpropagation) and the Linear Algebra guide (positive definite matrices, eigenvalues, least squares, the SVD). Revisit them whenever a step feels shaky.