Calculus for Machine Learning
Learn the maths of change from zero. Every idea starts with a plain-English picture, then a worked example, then the formal definition, and then something you can drag, slide, rotate and play with. Everything points toward one goal: understanding how a machine-learning model learns.
What is calculus, in one sentence? It is the maths of change: how fast something is changing right now (the derivative), and how little changes add up to a total (the integral).
Why care? A model learns by repeating one small move: "if I nudge this number a tiny bit, does the error go up or down, and by how much?" That question is a derivative. Training a neural network is that question asked millions of times. If you understand it, gradient descent and backpropagation stop being magic.
Companion guide. This guide goes hand in hand with the Linear Algebra guide. You can follow both side by side: whenever a vector or matrix appears, we remind you what it is, and link to the exact place that explains it.
How every topic is taught
Each concept follows the same six steps, in this order. The order matters: understanding comes before formulas.
1 · Intuition
A picture or everyday story. No symbols yet. If you only read this part, you will already get the idea.
2 · Example
A small problem with real numbers, solved slowly, step by step. You can redo it with pen and paper.
3 · Definition
Now the precise statement and the notation. Because you have seen the idea, the symbols just give it a name.
4 · Play
An interactive visual. Drag points, move sliders and watch the numbers change. In 3D you can also rotate the whole scene. Each one tells you what to try.
5 · Why · Where · How
Every concept answers three questions in plain English: Why do we need it? Where is it used? How is it used? So you always know why the idea is worth your time.
6 · Check
Short questions to test yourself, a recap of the key points, a NumPy code block to type out, and practice problems with worked answers.
Try your first interactive
The boxes marked Interactive react to you. Drag the dot along the curve and watch the straight line that just touches it.
The roadmap
Fourteen chapters, in an order where each one builds on the last. Click any card to jump there.
- 2.1Functions & FoundationsDomain, range, composition, sigmoid, softmax…
- 2.2Limits & ContinuityWhat "getting closer and closer" means
- 2.3DifferentiationSlope, rate of change, the derivative rules
- 2.4Partial Derivatives & GradientsSlopes of surfaces, steepest ascent
- 2.5Gradients of Vector FunctionsThe Jacobian matrix
- 2.6Matrix CalculusDerivatives with vectors and matrices
- 2.7Useful Gradient IdentitiesHow to derive them, not memorise them
- 2.8The Chain RuleComputational graphs, forward and reverse
- 2.9Backprop & AutodiffHow networks compute all their gradients
- 2.10Higher-Order DerivativesCurvature, the Hessian, saddle points
- 2.11Taylor SeriesApproximating functions near a point
- 2.12LinearizationZoom in far enough and everything is flat
- 2.13Multivariable CalculusFields, divergence, curl, integrals (awareness)
- 2.14Calculus → MLRegression, cross-entropy, SGD, backprop
What matters most for ML. Spend your effort on chapters 2.3 to 2.12: derivatives, gradients, the Jacobian, the chain rule, backpropagation, the Hessian and Taylor expansions. Limits (2.2) only need to be understood well enough to understand a derivative, and the vector-calculus tools in 2.13 (divergence, curl, line and surface integrals) are for awareness.
How to read the maths symbols
You will meet these symbols again and again. Do not memorise them now. Come back to this table whenever one looks strange.
| Symbol | Say it as | Meaning |
|---|---|---|
| $f(x)$ | "f of x" | A rule that turns the number $x$ into another number. |
| $\lim_{x \to a} f(x)$ | "the limit of f as x approaches a" | The value $f(x)$ gets closer and closer to as $x$ gets close to $a$. |
| $f'(x)$, $\dfrac{df}{dx}$ | "f prime", "d f d x" | The derivative: how fast $f$ changes as $x$ changes. |
| $\dfrac{\partial f}{\partial x}$ | "partial f partial x" | How fast $f$ changes when only $x$ changes and everything else is held still. |
| $\nabla f$ | "del f", "the gradient of f" | The list of all partial derivatives. It points uphill. |
| $J$ | "the Jacobian" | The matrix of all partial derivatives of a vector-valued function. |
| $H$, $\nabla^2 f$ | "the Hessian" | The matrix of all second derivatives (curvature). |
| $\displaystyle\int_a^b f(x)\,dx$ | "the integral of f from a to b" | The area under the curve of $f$ between $a$ and $b$. |
| $\Delta x$, $dx$ | "delta x", "d x" | A small change in $x$ ($dx$ means "infinitely small"). |
| $e$, $\ln x$ | "e", "natural log of x" | $e \approx 2.718$, and $\ln$ is the inverse of $e^x$. |
| $\mathbf{x}^\top$, $A^\top$ | "transpose" | Flip a vector or matrix (rows ↔ columns). |
| $\approx$ | "approximately equal" | Close, but not exactly equal. |
Tips for studying
- Play first, read second. If a paragraph confuses you, go to the interactive below it and move things around. Then re-read.
- Do the examples by hand once. It feels slow, but nothing builds intuition faster.
- Always ask "what does it mean for the slope?" Almost every derivative idea is a statement about steepness. Keep that picture in your head.
- Don't rush. One chapter a week is a good pace. Come back to earlier chapters whenever you need to.
- Use the tools. Press / to search, use the sidebar to jump around, switch dark mode with the button at the top, and mark chapters complete as you go.
If a section feels too hard, it is almost always because a word from an earlier section is fuzzy. Go back, find that word in the sidebar, and reread it. Nobody gets calculus in a single pass. Seeing it twice is normal.
Functions & Mathematical Foundations
Calculus studies how functions change, so before we study change we must be comfortable with functions themselves. This chapter builds your toolbox: the function families you will differentiate in the next chapters, and the three special ones that power neural networks: sigmoid, softmax and ReLU. We start from the simplest picture: a function is a machine.
- See a function as a machine, a table, a formula and a graph, and know its domain and range
- Chain functions together (composition) and undo them (inverse functions)
- Know the families: linear, polynomial, exponential, logarithmic, trigonometric
- Understand sigmoid and softmax deeply: their shape, saturation and why every classifier uses them
- Understand piecewise functions such as ReLU, leaky ReLU, absolute value, step, clip and Huber
- Get a first look at functions of many inputs, which is what a loss function really is
You only need school algebra here. Where an idea from the Linear Algebra guide helps, we say so and link to it, for example vectors and matrices. Every symbol is explained the first time it appears. If you feel rusty, that is fine: go slowly and play with each picture.
What is a function? core
Picture a vending machine. You press one button (the input) and one specific snack drops out (the output). Press the same button tomorrow and you get the same snack. That is all a function is: a rule that turns each input into exactly one output, the same way every time.
The square-root key on a calculator is a function: type 9, press the key, and 3 comes out. Type 16 and 4 comes out.
We can describe one rule in four ways, and they are all the same function: in words ("double it, then add 3"), as a table of inputs and outputs, as a formula, and as a graph, a picture of all the (input, output) pairs. Calculus will mostly work with the graph, because a graph lets us see how fast the output changes.
A taxi fare. A taxi charges 3 dollars to start and 2 dollars for every kilometre. If you ride $x$ kilometres, the fare is "double $x$, then add 3". Call this rule $f$ and write $f(x) = 2x + 3$.
- Ride 0 km: $f(0) = 2\cdot 0 + 3 = 3$.
- Ride 1 km: $f(1) = 2\cdot 1 + 3 = 5$.
- Ride 5 km: $f(5) = 2\cdot 5 + 3 = 13$.
- The rule also accepts numbers a taxi would never see: $f(-2) = 2\cdot(-2) + 3 = -4 + 3 = -1$.
Data version. A model that predicts a house price from its area is a function: area goes in, one predicted price comes out.
A function $f$ is a rule that assigns to every allowed input $x$ exactly one output, written $f(x)$ and read "f of x". We write $y = f(x)$.
- $x$ is the input (the independent variable). $y$ is the output (the dependent variable, because it depends on $x$).
- "$f : \mathbb{R} \to \mathbb{R}$" says: $f$ takes a real number and gives a real number. ($\mathbb{R}$ is the set of all real numbers, the whole number line.)
- The graph of $f$ is the set of all points $(x,\, f(x))$ drawn in the plane.
- Vertical line test. A curve is the graph of a function exactly when no vertical line touches it more than once. (Two touches would mean one input with two outputs.)
- The letter does not matter: $f(t) = 2t + 3$ is the same function as $f(x) = 2x + 3$.
Why do we need it?
We need a precise way to say "this quantity depends on that one". Calculus asks how fast the output changes when the input changes, and that question only makes sense once we have a function.
Where is it used?
Everywhere in machine learning: a trained model is a function from features to a prediction, a loss is a function from weights to an error number, and an activation (sigmoid, ReLU) is a function applied to every neuron.
How is it used?
Write the rule, plug in an input, read off the output. To understand a function, plot it. To use it for learning, ask the calculus question: "if I nudge the input a tiny bit, how much does the output move?"
$f(x)$ is not "f times x". It means "the output of $f$ when the input is $x$". And $f(x+1)$ is the output at the input $x+1$, which is usually not $f(x) + 1$. For $f(x)=x^2$: $f(3+1) = 16$ but $f(3)+1 = 10$.
One input, one output. Many inputs may share an output ($x^2$ gives 4 for both 2 and $-2$). That is fine. What is forbidden is one input with two outputs.
Quick check: is $y = \pm\sqrt{x}$ (both signs) a function of $x$?
No. At $x = 4$ it gives both $2$ and $-2$, so a vertical line at $x=4$ touches it twice. The expression $\sqrt{x}$ on its own (the positive root only) is a function.
Domain and range core
Every machine has rules about what it accepts. A toaster accepts bread, not soup. The calculator's square-root key refuses $-4$, and the "divide by" key refuses $0$ and shows an error.
So a function comes with two natural questions:
- What may I put in? The set of allowed inputs is the domain.
- What can come out? The set of outputs that really happen is the range.
On a graph, the domain is the part of the x-axis the curve lives above or below. The range is the part of the y-axis the curve reaches.
| Function | Domain (allowed $x$) | Range (outputs $y$) | Why |
|---|---|---|---|
| $x^2$ | all real numbers | $y \ge 0$ | squares are never negative |
| $\sqrt{x}$ | $x \ge 0$ | $y \ge 0$ | no real root of a negative number |
| $1/x$ | $x \ne 0$ | $y \ne 0$ | cannot divide by zero; and $1/x$ is never 0 |
| $\ln x$ | $x > 0$ | all real numbers | logarithm needs a positive input |
Finding a domain, step by step. For $f(x) = \dfrac{1}{x-3}$: the only danger is dividing by zero. The bottom is zero when $x - 3 = 0$, that is $x = 3$. So the domain is "every $x$ except 3". For $f(x)=\sqrt{x-1}$: we need $x - 1 \ge 0$, so $x \ge 1$.
The domain of $f$ is the set of all inputs $x$ for which $f(x)$ is defined. The range of $f$ is the set of all outputs $f(x)$ that actually occur.
Interval notation. $[a, b]$ means $a \le x \le b$ (square bracket: the end is included). $(a, b)$ means $a < x < b$ (round bracket: the end is left out). $(0, \infty)$ means "every number above 0"; infinity is a direction, not a number, so it always gets a round bracket.
Three danger signs when you look for a domain: (1) dividing by zero, (2) an even root (like $\sqrt{\ }$) of a negative number, (3) a logarithm of zero or of a negative number.
Why do we need it?
Before using a function we must know where it works. Feeding it a forbidden input gives an error or garbage, and knowing the range tells us what values to expect back.
Where is it used?
A probability output must lie in the range $(0,1)$, which is why classifiers end with sigmoid or softmax. The loss $-\ln p$ needs $p>0$. A model trained on ages 20 to 60 has no reliable domain at age 90.
How is it used?
Check the three danger signs before you trust a formula. In code, clip inputs of logs and divisions (for example np.log(p + 1e-12)) so a stray zero cannot crash training.
Domain is about inputs, range is about outputs. Do not mix them up: for $\sqrt{x}$ both are "numbers 0 or more", but they are different sets (one on the x-axis, one on the y-axis).
Never take a log or a divide of an exact zero in code. A probability that rounds to $0$ makes $\ln p = -\infty$ and the whole loss becomes "inf" or "nan".
Quick check: what is the domain of $f(x) = \dfrac{\sqrt{x}}{x-4}$?
Two danger signs. The root needs $x \ge 0$. The division needs $x \ne 4$. Together: all $x \ge 0$ except $x = 4$.
Composition: machines in a row core
Put two machines on an assembly line. The first machine's output drops straight into the second machine. First wash the potato, then chop it. The pair of machines acts like one bigger machine that hides the middle step.
That is composition. Order matters: chopping first and washing second is a different (and messier) recipe.
Let $g(x) = x + 2$ (runs first) and $f(x) = x^2$ (runs second). Start with $x = 3$.
- First machine: $g(3) = 3 + 2 = 5$.
- Second machine takes that 5: $f(5) = 5^2 = 25$.
- So "$f$ after $g$" turns 3 into 25.
Now swap the order. $f(3) = 9$, then $g(9) = 9 + 2 = 11$. We got 11, not 25: order matters.
The formula. Replace $x$ inside $f$ by the whole of $g(x)$: $f(g(x)) = (x+2)^2 = x^2 + 4x + 4$. Check at $x=3$: $9 + 12 + 4 = 25$ ✓.
A one-neuron example. Let $g(x) = 2x - 1$ (a linear step) and $\sigma$ be the sigmoid (you meet it later in this chapter). At $x = 1$: $g(1) = 1$ and $\sigma(1) \approx 0.731$. A neuron is exactly "linear step, then squash".
The composition of $f$ and $g$ is the function $$(f\circ g)(x) = f\big(g(x)\big).$$ Read "$f$ after $g$": $g$ is the inner function (acts first), $f$ is the outer function (acts second).
- Domain: $x$ must be allowed by $g$, and $g(x)$ must be allowed by $f$. Example: $f=\sqrt{\ }$, $g(x)=x-1$ gives $\sqrt{x-1}$, whose domain is $x \ge 1$.
- Order matters: in general $f\circ g \ne g\circ f$.
- Grouping does not matter: $(f\circ g)\circ h = f\circ(g\circ h)$, so we may write $f\circ g\circ h$.
- A deep neural network with $L$ layers is a composition $f_L \circ \dots \circ f_2 \circ f_1$.
Why do we need it?
Complicated behaviour is built by chaining simple steps. Composition lets us name and study each step separately, and later it tells us how nudges travel through a chain.
Where is it used?
Every neural network layer-stack, an image pipeline (resize, normalise, model), a loss that wraps a model, $\mathrm{loss}(\mathrm{model}(x))$, and a model that wraps a feature map, such as $\sigma(\mathbf{w}\cdot\mathbf{x}+b)$.
How is it used?
Work from the inside out: compute the inner output first, then feed it to the outer function. To get a formula, replace the letter $x$ in the outer formula by the whole inner formula (in brackets!).
Looking ahead. If a nudge to the input moves the inner output by some factor, and a nudge to that output moves the final result by another factor, the total effect is the product of the two factors. This is the chain rule (Chapter 2.8), and backpropagation (Chapter 2.9) is that rule applied through a whole network.
Quick check: if $f(x)=3x$ and $g(x)=x-2$, what are $f(g(4))$ and $g(f(4))$?
$g(4) = 2$, then $f(2) = 6$. And $f(4) = 12$, then $g(12) = 10$. So $f(g(4)) = 6$ but $g(f(4)) = 10$.
Inverse functions: the undo button core
Some machines have an undo machine. A function turns Celsius into Fahrenheit; its inverse turns Fahrenheit back into Celsius. Put a number through both and you are back where you started.
On a graph, an inverse swaps the roles of input and output. If $(2, 7)$ is on the graph of $f$, then $(7, 2)$ is on the graph of the inverse. Swapping $x$ and $y$ is a mirror image across the diagonal line $y = x$.
An undo button only works if the machine never maps two different inputs to the same output. Otherwise, looking at the output, you could not tell which input it came from.
Find the inverse of $f(x) = 2x + 3$. Write $y = 2x + 3$ and solve for $x$:
- Start: $y = 2x + 3$.
- Subtract 3: $y - 3 = 2x$.
- Divide by 2: $x = \dfrac{y - 3}{2}$.
- Now rename the input of the new function to $x$: $f^{-1}(x) = \dfrac{x-3}{2}$.
Check: $f(5) = 13$ and $f^{-1}(13) = (13-3)/2 = 5$ ✓. Temperature check: $F = 1.8C + 32$ gives $100^\circ\text{C} \to 212^\circ\text{F}$, and the inverse $C = (F - 32)/1.8$ gives $(212-32)/1.8 = 100$ ✓.
A function with no inverse: $f(x)=x^2$ sends both 2 and $-2$ to 4. Given the output 4, which input was it? If we restrict the domain to $x \ge 0$, the problem disappears and the inverse is $\sqrt{x}$.
A function $f$ is one-to-one if different inputs always give different outputs. (Horizontal line test: no horizontal line touches the graph twice.) Then it has an inverse function $f^{-1}$ that undoes it: $$f^{-1}\big(f(x)\big) = x \qquad\text{and}\qquad f\big(f^{-1}(y)\big) = y.$$
- The domain of $f^{-1}$ is the range of $f$, and the range of $f^{-1}$ is the domain of $f$.
- The graph of $f^{-1}$ is the graph of $f$ reflected in the line $y=x$.
- Careful: $f^{-1}$ is not $1/f$. The small $-1$ means "inverse function", not "to the power $-1$".
- How to find it: write $y = f(x)$, solve for $x$, then swap the names.
Why do we need it?
Often we know the result and want the cause: given the output, which input made it? An inverse answers that. It is also how we build new functions: $\ln$ and $\sqrt{\ }$ are inverses of other functions.
Where is it used?
The logit (inverse of the sigmoid) in logistic regression, $\ln$ as the inverse of $\exp$, undoing data normalisation to get real units back, and sampling random numbers by inverting a cumulative distribution.
How is it used?
First check the function is one-to-one (restrict its domain if not). Then solve $y=f(x)$ for $x$. To verify, compute $f^{-1}(f(x))$ and confirm you get $x$ back.
$f^{-1}(x) \ne \dfrac{1}{f(x)}$. For $f(x)=2x+3$: $f^{-1}(x) = (x-3)/2$, but $1/f(x) = 1/(2x+3)$. They are different functions.
Restricting the domain is a legitimate fix. $\sqrt{x}$ is the inverse of "$x^2$ on $x\ge0$". The function $x^2$ on all numbers has no inverse.
Quick check: find the inverse of $f(x) = 5x - 10$.
$y = 5x - 10 \Rightarrow y + 10 = 5x \Rightarrow x = (y+10)/5$. So $f^{-1}(x) = (x+10)/5 = \tfrac{x}{5} + 2$. Check: $f(3) = 5$, and $f^{-1}(5) = 15/5 = 3$ ✓.
Linear functions: a straight line core
Walk up a straight ramp. Every metre you move forward raises you by the same amount, no matter where you are on the ramp. That constant steepness is the whole idea of a linear function.
Two numbers describe a ramp: how steep it is (the slope), and how high it starts at $x=0$ (the intercept). The taxi fare from the start of the chapter is one: 2 dollars per km, plus 3 dollars to start.
Take $f(x) = 2x + 1$.
| $x$ | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| $f(x)$ | 1 | 3 | 5 | 7 |
Each step of 1 to the right adds exactly 2 to the output. That "2" is the slope.
From two points to a line. A line passes through $(1, 3)$ and $(3, 7)$.
- Slope = rise over run $= \dfrac{7 - 3}{3 - 1} = \dfrac{4}{2} = 2$.
- Intercept: use $y = mx + b$ at $(1,3)$: $3 = 2\cdot 1 + b$, so $b = 1$.
- The line is $y = 2x + 1$.
A linear function has the form $$f(x) = m\,x + b.$$
- $m$ is the slope: $m = \dfrac{\text{rise}}{\text{run}} = \dfrac{\Delta y}{\Delta x} = \dfrac{y_2 - y_1}{x_2 - x_1}$. ($\Delta$, "delta", means "change in".) Positive $m$: the line climbs. Negative: it falls. Zero: flat.
- $b$ is the intercept: $f(0) = b$, where the line crosses the y-axis.
- The key property: $f(x+1) - f(x) = \big(m(x+1) + b\big) - (mx + b) = m$. The change per unit step is the same everywhere.
- A vertical line ($x = 3$) is not a function of $x$ (it fails the vertical line test).
A word on names. In linear algebra, "linear" is stricter: it means $f(cx) = c\,f(x)$ and $f(x+y) = f(x)+f(y)$, which forces $b = 0$. A line with $b\ne0$ is then called affine. Machine learning is relaxed about this and calls $wx + b$ "linear". With many inputs it is $\mathbf{w}\cdot\mathbf{x} + b$, a dot product plus a constant.
Why do we need it?
A line is the simplest possible "cause and effect": one more unit of input always gives $m$ more output. It is the first model to try, and it is what every smooth curve looks like when you zoom in far enough (Chapter 2.12).
Where is it used?
Linear regression, the weighted sum $\mathbf{w}\cdot\mathbf{x}+b$ inside every neuron, linearly decaying learning-rate schedules, and tangent lines (the best straight-line copy of a curve at one point).
How is it used?
Get the slope from two points (rise over run), get the intercept from the value at $x=0$, then predict by plugging in. To learn a line from data, choose the $m$ and $b$ that make the errors smallest.
Looking ahead. The slope of a line is its rate of change, and it is the same everywhere. For a curve the steepness changes from point to point. Finding the steepness at one single point is exactly what the derivative does (Chapter 2.3). The derivative of the line $mx+b$ is just $m$.
Quick check: a line goes through $(0, 5)$ and $(2, 1)$. What is $f(10)$?
Intercept $b = 5$ (the value at $x=0$). Slope $= (1-5)/(2-0) = -2$. So $f(x) = -2x + 5$ and $f(10) = -20 + 5 = -15$.
Polynomial functions: sums of powers core
A polynomial is built from building blocks: $1$, $x$, $x^2$, $x^3$, … Each block is multiplied by a number (a coefficient) and the pieces are added up. With only $1$ and $x$ you get a line. Add $x^2$ and the line can bend once, into a parabola (a bowl). Add $x^3$ and it can bend twice, and so on.
So more powers means more flexibility: a higher-degree polynomial can wiggle more. That is both its power and its danger.
Let $p(x) = x^2 - 4x + 3$.
- Factor it: $x^2 - 4x + 3 = (x-1)(x-3)$ (check: $(x-1)(x-3) = x^2 - 3x - x + 3$ ✓).
- It is zero when a factor is zero, so the roots are $x = 1$ and $x = 3$. Check: $p(1) = 1 - 4 + 3 = 0$ ✓ and $p(3) = 9 - 12 + 3 = 0$ ✓.
- At $x = 0$: $p(0) = 3$ (the constant term).
- The bottom of the bowl is halfway between the roots, at $x = 2$: $p(2) = 4 - 8 + 3 = -1$.
A cubic: $x^3 - x = x(x-1)(x+1)$ has three roots, $-1, 0, 1$, and bends twice.
A polynomial of degree $n$ is $$p(x) = a_n x^n + a_{n-1}x^{n-1} + \dots + a_2 x^2 + a_1 x + a_0, \qquad a_n \ne 0.$$ The numbers $a_0,\dots,a_n$ are the coefficients; the degree is the highest power present. A root is an $x$ with $p(x) = 0$.
- Degree 1 is a line, degree 2 a parabola, degree 3 a cubic.
- A degree-$n$ polynomial has at most $n$ real roots and at most $n-1$ turning points (peaks and valleys).
- End behaviour is decided by the leading term $a_n x^n$ alone, when $|x|$ is huge: even degree means both ends go the same way (both up if $a_n>0$); odd degree means the ends go opposite ways.
- If you know the roots $r_1,\dots,r_n$, then $p(x) = a_n (x - r_1)\cdots(x - r_n)$.
- The domain is all real numbers. Polynomials are smooth: no gaps, jumps or corners.
Why do we need it?
Polynomials are the easiest flexible curves: only multiplication and addition, so they are cheap to compute and (you will see) easy to differentiate. They can also imitate other smooth functions near a point.
Where is it used?
Polynomial regression and polynomial features, the quadratic shape of the mean-squared-error loss in the weights, Taylor series (Chapter 2.11), splines in curve fitting, and polynomial kernels in SVMs.
How is it used?
Choose a degree, then let training find the coefficients. Too low a degree underfits (cannot bend enough); too high overfits (it wiggles through the noise). The degree is a "complexity knob".
"Higher degree" is not "better". More flexibility always lowers the error on the points you trained on, but it can raise the error on new points. We measure fit on data the model has not seen.
Roots can be complex. $x^2+1$ has no real root (its bowl never touches the axis). The count "at most $n$" refers to real roots.
Quick check: how many turning points can a degree-4 polynomial have at most, and how do its ends behave if $a_4 > 0$?
At most $4 - 1 = 3$ turning points. Even degree with positive leading coefficient: both ends go up.
Exponential functions and the number $e$ core
In a linear function you add the same amount every step. In an exponential function you multiply by the same factor every step. A bacterium that splits in two every hour: 1, 2, 4, 8, 16, … It grows slowly at first, then explosively, because the bigger it is, the faster it grows.
Run the film backwards for decay: a cup of coffee's caffeine halves every few hours: 1, 1/2, 1/4, 1/8, … It shrinks toward zero but never reaches it.
One base is special: the number $e \approx 2.71828$. We meet it properly in Chapter 2.2 (where $e$ appears as a limit) and Chapter 2.3 (where its steepness is worked out), but the short story is: $e^x$ is the exponential whose steepness at any point equals its height at that point. That makes it the most natural exponential, and calculus loves it.
$f(x) = 2^x$:
| $x$ | $-2$ | $-1$ | 0 | 1 | 2 | 3 | 10 |
|---|---|---|---|---|---|---|---|
| $2^x$ | $\tfrac14$ | $\tfrac12$ | 1 | 2 | 4 | 8 | 1024 |
Decay: $(1/2)^x$ gives $1, \tfrac12, \tfrac14, \dots$ for $x = 0, 1, 2, \dots$. And $e^0 = 1$, $e^1 \approx 2.718$, $e^2 \approx 7.389$, $e^{-1} \approx 0.368$.
Exponent rules (always true, for $b>0$):
- $b^{x+y} = b^x\, b^y$ (example: $2^{3+2} = 2^5 = 32 = 8\cdot 4$ ✓)
- $b^{0} = 1$ and $b^{-x} = 1/b^x$ (example: $2^{-3} = 1/8$)
- $(b^x)^y = b^{xy}$ (example: $(2^3)^2 = 8^2 = 64 = 2^6$ ✓)
An exponential function with base $b>0$, $b \ne 1$, is $f(x) = b^x$.
- Domain: all real numbers. Range: $y > 0$ (always positive, never zero).
- It always passes through $(0, 1)$, because $b^0 = 1$.
- If $b>1$ it grows (increasing); if $0\lt b<1$ it decays (decreasing).
- The natural exponential uses $b = e \approx 2.71828\ldots$, written $e^x$ or $\exp(x)$.
- Every exponential is secretly an $e$-exponential: since $b = e^{\ln b}$ (the number $\ln b$ is defined in the next section), $b^x = (e^{\ln b})^x = e^{x\ln b}$.
Why do we need it?
Many things grow or shrink by a fixed percentage per step rather than a fixed amount. Exponentials describe these, and $e^x$ is always positive, which is exactly what we need to turn any score into a valid weight or probability.
Where is it used?
Inside sigmoid and softmax, learning-rate decay schedules, exponential moving averages (momentum, Adam), weight decay, the Gaussian bell curve $e^{-x^2/2}$, compound interest and epidemic growth.
How is it used?
Pick a base (or a rate $k$ in $e^{kx}$). For a growth base, the doubling time is $\ln 2/\ln b$; for a decay base, the half-life is $\ln 2/(-\ln b)$. Remember $e^{a+b}=e^a e^b$ to simplify.
$x^2$ is not $2^x$. In $x^2$ the variable is the base (a polynomial); in $2^x$ the variable is the exponent (an exponential). The second one eventually beats every polynomial: $2^{10}=1024$ but $10^2 = 100$, and the gap only grows.
An exponential never reaches 0 or goes negative: $b^x > 0$ for every $x$.
Quick check: simplify $e^{3}\cdot e^{-3}$ and $(e^{2})^{3}$.
$e^3 e^{-3} = e^{3-3} = e^0 = 1$. And $(e^2)^3 = e^{2\cdot3} = e^6$.
Logarithms: the inverse of exponentials core
A logarithm answers one question: "what power?" To what power must 2 be raised to get 8? Three, because $2\cdot2\cdot2=8$. We write $\log_2 8 = 3$.
So the logarithm is the undo button of the exponential (the inverse function from the last section). Exponential: power in, number out. Logarithm: number in, power out.
It also turns multiplication into addition. $1000 \times 100 = 100\,000$: the zeros add up, $3 + 2 = 5$. A logarithm just counts those zeros. That is why earthquakes (Richter), loudness (decibels) and acidity (pH) are measured on log scales: huge ranges become small, friendly numbers.
- $\log_{10} 1000 = 3$, because $10^3 = 1000$.
- $\log_2 32 = 5$, because $2^5 = 32$.
- $\log_2 \tfrac18 = -3$, because $2^{-3} = \tfrac18$.
- $\ln e^2 = 2$ (here $\ln$ means $\log_e$, the "natural log").
- $\log_b 1 = 0$ for any base (since $b^0 = 1$) and $\log_b b = 1$.
Why not $\log 0$ or $\log(-5)$? Because $b^y$ is always positive, no power $y$ can give $0$ or a negative number. So logs only accept positive inputs.
Multiplication becomes addition: $\log_2(8 \cdot 4) = \log_2 32 = 5$, and $\log_2 8 + \log_2 4 = 3 + 2 = 5$ ✓.
For a base $b>0$, $b\ne1$, and $x > 0$: $$y = \log_b x \quad\Longleftrightarrow\quad b^{y} = x.$$ $\ln x = \log_e x$ is the natural logarithm; $\log_{10}$ is the common log; $\log_2$ counts bits. Domain: $x>0$. Range: all real numbers. Its graph passes through $(1, 0)$, has the y-axis as a vertical asymptote, and is the mirror image of $b^x$ in the line $y=x$. As inverses: $$\log_b(b^x) = x, \qquad b^{\log_b x} = x.$$
The rules, derived. Let $p = \log_b x$ and $q = \log_b y$, so $x = b^p$ and $y = b^q$.
- Product: $xy = b^p b^q = b^{p+q}$, so $\log_b(xy) = p + q = \log_b x + \log_b y$.
- Quotient: $x/y = b^p/b^q = b^{p-q}$, so $\log_b(x/y) = \log_b x - \log_b y$.
- Power: $x^k = (b^p)^k = b^{pk}$, so $\log_b(x^k) = k\log_b x$.
- Change of base: let $y = \log_b x$, so $b^y = x$. Take $\ln$ of both sides: $y \ln b = \ln x$, hence $\log_b x = \dfrac{\ln x}{\ln b}$.
Why do we need it?
Logs turn products into sums and powers into multiplications, and they squeeze enormous ranges into small ones. Sums are far easier to handle than products, both on paper and in a computer.
Where is it used?
The log-likelihood and the cross-entropy loss ($-\ln p$), entropy in bits, log-softmax, log-scaled plot axes for losses and learning rates, and feature transforms like $\log(\text{price})$.
How is it used?
To avoid tiny numbers, take logs of probabilities and add instead of multiplying. To undo an exponential, apply $\ln$. To change base, divide by $\ln b$.
$\ln$ vs $\log$. In mathematics and most ML papers, "$\log$" with no base means the natural log $\ln$. In school (and on some calculators) it means $\log_{10}$. Check which one a formula uses; NumPy's np.log is the natural log.
Logs do not split sums: $\ln(A+B) \ne \ln A + \ln B$. The rules only work for products, quotients and powers.
Quick check: simplify $\ln(e^5) + \log_2 8 - \log_{10}(0.01)$.
$\ln(e^5) = 5$, $\log_2 8 = 3$, and $\log_{10}0.01 = \log_{10}10^{-2} = -2$. Total: $5 + 3 - (-2) = 10$.
Trigonometric functions: circles and waves
Watch a point travel round a circle of radius 1 at a steady speed. Look at it from the side (its height) and the height goes up and down smoothly, over and over. Plot that height against the angle travelled and you get a wave. That wave is the sine function. The point's horizontal position gives the cosine, the same wave shifted by a quarter turn.
Angles are measured in radians: the length of the arc you walk along the unit circle. A full circle has length $2\pi \approx 6.283$, so $360^\circ = 2\pi$ radians, $180^\circ = \pi$, $90^\circ=\pi/2$. Calculus (and every programming language) uses radians.
The point on the unit circle at angle $\theta$ is $(\cos\theta,\ \sin\theta)$. Here are the angles you should recognise:
| Angle | $0$ | $30^\circ = \pi/6$ | $45^\circ = \pi/4$ | $60^\circ = \pi/3$ | $90^\circ = \pi/2$ | $180^\circ = \pi$ |
|---|---|---|---|---|---|---|
| $\cos\theta$ | 1 | $\sqrt3/2 \approx 0.866$ | $\sqrt2/2 \approx 0.707$ | $1/2$ | 0 | $-1$ |
| $\sin\theta$ | 0 | $1/2$ | $\sqrt2/2 \approx 0.707$ | $\sqrt3/2 \approx 0.866$ | 1 | 0 |
Converting: radians $=$ degrees $\times \pi/180$. For $30^\circ$: $30\times\pi/180 = \pi/6 \approx 0.524$.
For any angle $\theta$ (in radians, counter-clockwise from the positive x-axis), the point on the unit circle is $(\cos\theta, \sin\theta)$. Also $\tan\theta = \dfrac{\sin\theta}{\cos\theta}$.
- Domain of $\sin$, $\cos$: all real numbers. Range: $[-1, 1]$.
- Periodic: $\sin(\theta + 2\pi) = \sin\theta$ and the same for cosine (period $2\pi$). $\tan$ has period $\pi$ and blows up where $\cos\theta = 0$ (at $\pm\pi/2, \pm3\pi/2,\dots$).
- Pythagoras on the circle: the point $(\cos\theta,\sin\theta)$ is at distance 1 from the origin, so $\cos^2\theta + \sin^2\theta = 1$ always.
- $\sin(-\theta) = -\sin\theta$ (odd) and $\cos(-\theta) = \cos\theta$ (even); $\cos\theta = \sin(\theta + \pi/2)$.
- The wave family: $y = A\sin(\omega x + \varphi) + D$ has amplitude $A$ (height of a peak above the middle), angular frequency $\omega$ (period $2\pi/\omega$), phase $\varphi$ (slides the wave left by $\varphi/\omega$) and offset $D$ (moves the middle line).
Why do we need it?
Many things repeat or rotate: seasons, sound, an angle of a robot arm. Trig functions are the exact mathematics of repeating and of rotating, and they connect angles to lengths.
Where is it used?
Positional encodings in Transformers (sines and cosines of many frequencies), Fourier features and audio signals, rotation matrices, seasonality in time series, cosine similarity (the $\cos\theta$ in the dot product), and cosine learning-rate schedules.
How is it used?
Give angles in radians (NumPy's np.sin expects them). Choose amplitude, frequency and phase to match a repeating pattern. Remember the bounded range $[-1,1]$: sin and cos never blow up.
Degrees vs radians. np.sin(90) is not 1: NumPy reads 90 as 90 radians. Use np.sin(np.pi/2) or np.deg2rad(90).
$\sin^2\theta$ means $(\sin\theta)^2$, not $\sin(\theta^2)$.
Quick check: what are $\sin(\pi)$, $\cos(\pi)$ and the period of $\sin(3x)$?
At $\theta=\pi$ (halfway round, the point $(-1,0)$): $\sin\pi = 0$, $\cos\pi = -1$. The period of $\sin(3x)$ is $2\pi/3\approx 2.094$.
The sigmoid function: a smooth on/off switch core
Imagine a dimmer switch instead of a light switch. Turn the knob far to the left: the light is off (0). Far to the right: fully on (1). In the middle: half bright. The sigmoid is that dimmer. It takes any number (a "score", positive or negative, huge or tiny) and squashes it into a value between 0 and 1.
That is why it is perfect for probabilities. A big positive score means "very likely yes" (close to 1). A big negative score means "very likely no" (close to 0). A score of 0 means "50-50".
The graph is an "S" with two flat tails. In the flat tails the output barely changes even when the input changes a lot. This is called saturation, and it matters a great deal for learning (see the callout below).
The formula is $\sigma(x) = \dfrac{1}{1 + e^{-x}}$. Compute $\sigma(1)$ by hand:
- $e^{-1} \approx 0.3679$.
- $1 + 0.3679 = 1.3679$.
- $\sigma(1) = 1 / 1.3679 \approx 0.7311$.
| $x$ | $-5$ | $-2$ | 0 | 1 | 2 | 5 |
|---|---|---|---|---|---|---|
| $\sigma(x)$ | 0.0067 | 0.1192 | 0.5 | 0.7311 | 0.8808 | 0.9933 |
A spam filter. It computes a score of 2 for an email. $\sigma(2)\approx 0.88$: "88% chance it is spam". A score of $-5$ gives $0.0067$: almost certainly not spam. Going from 2 to 5 only moves the probability from 0.88 to 0.993: that flattening is saturation.
The sigmoid (or logistic function) is $$\sigma(x) = \frac{1}{1 + e^{-x}} = \frac{e^{x}}{1 + e^{x}}.$$ (The two forms agree: multiply top and bottom of the first by $e^x$.)
Properties, each derived:
- Range $(0,1)$. Since $e^{-x}>0$, the bottom $1+e^{-x}$ is bigger than 1, so $0<\sigma(x)<1$. It gets as close as you like to 0 and 1 but never touches them.
- $\sigma(0) = \tfrac12$. Because $e^0 = 1$ gives $1/(1+1)$.
- Symmetry: $\sigma(-x) = 1 - \sigma(x)$. Indeed $1 - \sigma(x) = 1 - \dfrac{1}{1+e^{-x}} = \dfrac{e^{-x}}{1+e^{-x}} = \dfrac{1}{e^{x}+1} = \sigma(-x)$ (multiply top and bottom by $e^x$ in the last step). The curve is point-symmetric about $(0, \tfrac12)$.
- Always increasing, since $e^{-x}$ falls as $x$ grows.
- Inverse = the logit. Solve $p = \dfrac{1}{1+e^{-z}}$ for $z$: $1 + e^{-z} = \dfrac1p$, so $e^{-z} = \dfrac{1-p}{p}$, so $-z = \ln\dfrac{1-p}{p}$, so $$z = \ln\frac{p}{1-p} = \operatorname{logit}(p).$$ Here $\frac{p}{1-p}$ is the odds and $z$ is the log-odds.
- Saturation numbers. $\sigma(x)\ge 0.95$ once $x \ge \ln 19 \approx 2.94$, and $\sigma(x) \ge 0.99$ once $x\ge\ln 99\approx 4.60$. (Solve $\sigma(x) = 0.95$: $e^{-x} = 0.05/0.95 = 1/19$.)
- Slope preview. The steepness of the curve is $\sigma'(x) = \sigma(x)\,(1-\sigma(x))$, at most $\tfrac14$ (at $x=0$) and nearly $0$ in the tails. We prove this in Chapter 2.3.
- Shape knobs: $\sigma(k(x - x_0))$ is the same S, centred at $x_0$ and steeper when $k$ is larger. As $k\to\infty$ it becomes an abrupt step.
- Cousin: $\tanh x = 2\sigma(2x) - 1$ has the same S-shape but ranges over $(-1, 1)$.
Why do we need it?
A model's raw score can be any number, but a probability must lie between 0 and 1. Sigmoid converts one into the other smoothly, so we can still use calculus on it (a hard cut-off at 0 or 1 has no useful slope).
Where is it used?
Logistic regression and every binary classifier's output layer, the gates of LSTM and GRU networks, "attention gates", the Swish/SiLU activation $x\,\sigma(x)$, and mapping any score to a "probability of yes".
How is it used?
Compute a score $z = \mathbf{w}\cdot\mathbf{x} + b$, then the probability $p=\sigma(z)$. Predict "yes" when $p>0.5$, which is the same as $z>0$. Train with the loss slope; the slope of $\sigma$ is the part that shrinks when $|z|$ is large.
Saturation causes slow learning. Training works by nudging the weights in the direction that lowers the error, and the size of the nudge depends on the slope of every function in the chain. In the flat tails of the sigmoid the slope is almost 0 (for example $\sigma'(5) \approx 0.0066$, versus $0.25$ at the centre), so the signal that flows back through a saturated sigmoid is tiny. In deep networks these small factors multiply and the signal vanishes: the vanishing gradient problem (Chapter 2.9). That is one reason hidden layers now often use ReLU (below), while sigmoid is kept for outputs and gates.
$\sigma$ outputs are probabilities only if the score is meaningful; a sigmoid of any number is between 0 and 1, even when the model is wrong.
Quick check: if $\sigma(z) = 0.2$, what is $\sigma(-z)$, and what is $z$ roughly?
By symmetry $\sigma(-z) = 1 - 0.2 = 0.8$. And $z = \operatorname{logit}(0.2) = \ln(0.2/0.8) = \ln 0.25 \approx -1.386$.
The softmax function: scores to probabilities core
Sigmoid answers a yes/no question. Softmax answers "which one of many?". A network looking at a photo produces one raw score (called a logit) per class: say cat 2.0, dog 1.0, bird 0.1. We want a list of probabilities: all positive, adding up to 1, with a bigger score getting a bigger share.
The recipe has two steps. (1) Make every score positive with $e^{\text{score}}$. (2) Divide each by the total, so you get a "share of the pie". The name says it all: it is a soft version of "take the max". The winner gets the biggest slice, but the others still get something, instead of everything going to the winner.
A temperature knob $T$ controls how soft: a small $T$ makes it almost winner-takes-all, a large $T$ makes the slices almost equal.
Scores $z = [2.0,\ 1.0,\ 0.1]$.
- Exponentiate: $e^{2.0} \approx 7.389$, $e^{1.0} \approx 2.718$, $e^{0.1} \approx 1.105$.
- Add them: $7.389 + 2.718 + 1.105 = 11.212$.
- Divide each by the total: $\dfrac{7.389}{11.212} \approx 0.659$, $\dfrac{2.718}{11.212}\approx 0.242$, $\dfrac{1.105}{11.212}\approx 0.099$.
- Check: $0.659 + 0.242 + 0.099 = 1.000$ ✓.
Adding the same number to every score changes nothing: $[102, 101, 100.1]$ gives the same three probabilities, because the common factor $e^{100}$ cancels from top and bottom.
| Temperature | probabilities for $[2.0, 1.0, 0.1]$ | behaviour |
|---|---|---|
| $T = 0.5$ | 0.864, 0.117, 0.019 | sharper: the winner dominates |
| $T = 1$ | 0.659, 0.242, 0.099 | ordinary softmax |
| $T = 5$ | 0.400, 0.327, 0.273 | flatter: nearly equal shares |
For scores $\mathbf{z} = [z_1, \dots, z_n]$, the softmax is the vector with entries $$\operatorname{softmax}(\mathbf{z})_i = \frac{e^{z_i}}{\sum_{j=1}^{n} e^{z_j}}.$$ With temperature: $\operatorname{softmax}(\mathbf{z}/T)$.
- Valid probabilities. Each $e^{z_i}>0$ and is smaller than the total, so each output is in $(0,1)$. Adding them: $\sum_i p_i = \dfrac{\sum_i e^{z_i}}{\sum_j e^{z_j}} = 1$.
- Order is kept. $e^x$ is increasing, so a bigger score always gets a bigger probability.
- Shift invariance. $\dfrac{e^{z_i + c}}{\sum_j e^{z_j + c}} = \dfrac{e^c\,e^{z_i}}{e^c\sum_j e^{z_j}} = \dfrac{e^{z_i}}{\sum_j e^{z_j}}$. Only the differences between scores matter. (Multiplying all scores by a number does change the answer: that is exactly what temperature does.)
- Two classes = sigmoid. $p_1 = \dfrac{e^{z_1}}{e^{z_1}+e^{z_2}}$. Divide top and bottom by $e^{z_1}$: $p_1 = \dfrac{1}{1+e^{-(z_1-z_2)}} = \sigma(z_1 - z_2)$. Softmax generalises the sigmoid.
- Temperature limits. $T\to 0$: all the weight goes to the largest score (the "hard max", argmax). $T\to\infty$: every class gets $1/n$.
- Numerically safe version. Subtract the largest score first, $z_i - \max_j z_j$ (allowed by shift invariance). This keeps every exponent $\le 0$ so $e^{(\cdot)}$ can never overflow. In logs: $\ln p_i = z_i - \ln\sum_j e^{z_j}$ (the "log-sum-exp").
- Slope preview. $\dfrac{\partial p_i}{\partial z_j} = p_i(\delta_{ij} - p_j)$, a table of slopes that you derive in Chapter 2.5 and use in Chapter 2.14. ($\delta_{ij}$ is 1 when $i=j$ and 0 otherwise.)
Why do we need it?
A classifier with many classes must output a proper probability for each, positive and adding to 1. Softmax does this while staying smooth, so calculus can tell the network how to improve each score.
Where is it used?
The output layer of image classifiers and language models (a probability for each of tens of thousands of next words), attention weights in Transformers, policies in reinforcement learning, and the cross-entropy loss. Temperature controls the randomness of text sampling.
How is it used?
Compute logits, subtract their maximum, exponentiate, divide by the sum. Pick the biggest probability (or sample from them). Lower $T$ for confident, repeatable outputs; raise $T$ for more varied ones. For training, take $\ln$ of the right class's probability.
Softmax outputs are not guaranteed to be calibrated. They always look like probabilities, even when the model is guessing. A softmax of $[10, 0, 0]$ says 99.99% for class 1, however wrong the model might be.
Softmax only sees differences. If you add the same number to every logit, nothing changes. That is a feature, but it also means the absolute size of a logit alone carries no meaning.
Quick check: what is softmax of $[0, 0, 0, 0]$, and of $[5, 5]$?
Equal scores give equal shares: $[0.25, 0.25, 0.25, 0.25]$ and $[0.5, 0.5]$. Each $e^0=1$ (or each $e^5$), the sum is $4$ (or $2e^5$), and the shares are $1/4$ (or $1/2$).
Piecewise functions: ReLU, leaky ReLU, absolute value, step, clip core
Some rules change depending on where the input is. A phone plan: "10 dollars flat up to 5 GB, then 2 dollars for every extra GB". Income tax works the same way, with different rates for different brackets. A piecewise function is several simple rules glued together, each used on its own stretch of inputs.
The most famous one in machine learning is ReLU ("rectified linear unit"): if the input is negative, output 0; otherwise pass it through unchanged. It is a flat line that bends into a ramp at 0. Simple, cheap, and it powers almost every modern deep network.
The phone plan. Let $x$ be the gigabytes used. $$\text{cost}(x) = \begin{cases} 10 & \text{if } x \le 5 \\ 10 + 2(x-5) & \text{if } x > 5 \end{cases}$$ For $x = 3$: first rule, cost $=10$. For $x = 8$: second rule, $10 + 2\cdot 3 = 16$. At $x=5$ both rules agree ($10$), so the pieces meet with no jump.
Absolute value. $|x| = x$ if $x \ge 0$, and $-x$ if $x<0$. So $|{-3}| = 3$ and $|4| = 4$.
The ML family (all piecewise):
| Name | Rule | Shape |
|---|---|---|
| ReLU | $\max(0, x)$: $0$ for $x<0$, $x$ for $x\ge0$ | flat, then a 45° ramp |
| Leaky ReLU | $x$ for $x\ge0$, $\alpha x$ for $x<0$ (small $\alpha$, e.g. 0.01) | a gentle slope on the left |
| Absolute value | $|x|$ | a V (used in the L1 loss) |
| Step | $0$ for $x<0$, $1$ for $x\ge0$ | a jump |
| Clip | $\min(\max(x, -c), c)$: stays in $[-c, c]$ | flat, ramp, flat |
| Huber loss | $\tfrac12x^2$ if $|x|\le\delta$, else $\delta(|x| - \tfrac\delta2)$ | a bowl with straight sides |
A piecewise function is defined by cases: $$f(x) = \begin{cases} f_1(x) & \text{if } x \text{ is in region 1} \\ f_2(x) & \text{if } x \text{ is in region 2} \\ \ \vdots \end{cases}$$ The regions must not overlap, and together they must cover the domain, so every $x$ has exactly one output.
- The places where one piece hands over to the next are the break points. If the pieces meet there, the graph is connected. If the graph has a corner there it is a kink (ReLU at 0, $|x|$ at 0). If the pieces do not meet, there is a jump (the step function at 0).
- Max and min give compact forms: $\text{ReLU}(x)=\max(0,x)$, $|x| = \max(x,-x)$, $\text{clip}=\min(\max(x,-c),c)$.
- Slope preview. On each straight piece the slope is constant: ReLU has slope $0$ for $x<0$ and $1$ for $x>0$. At the kink there is no single slope. How calculus deals with kinks is in Chapter 2.3; frameworks simply pick a convention (slope 0 or 1 at exactly 0).
Why ReLU is so popular: (1) it is the cheapest possible non-linear function; (2) for $x>0$ it never saturates: its slope is exactly 1, so learning signals pass through unchanged (compare the flat tails of the sigmoid); (3) many ReLU "hinges" added together can trace any smooth curve, as closely as you like, as a chain of straight pieces (see the second widget below); (4) it outputs exact zeros, which makes layers sparse. Its weakness: for $x<0$ the slope is 0, so a unit that gets stuck on the negative side stops learning (a "dead ReLU"). Leaky ReLU keeps a small slope $\alpha$ there, so the signal never fully dies.
Why do we need it?
Real rules often differ by region, and a network built only from straight (linear) layers can only ever draw straight things. Bending the line, even once, is what lets a network model curves.
Where is it used?
ReLU in most CNNs and the hidden layers of many networks; leaky ReLU in GANs; Huber loss in robust regression and in the DQN reinforcement-learning loss; absolute value in the L1 loss; clipping of gradients and of probability ratios (PPO); the step function in the original perceptron.
How is it used?
Apply the function to every number coming out of a linear layer: np.maximum(0, z). Choose by region: write each piece, test the condition, use the matching piece. For losses, pick the one whose shape penalises errors the way you want.
Check the break points. If two pieces do not agree at the joint, the function jumps (like the step function) and cannot be trained with slopes. If they agree but with different slopes, you get a kink (like ReLU), which is usually fine in practice.
"Leaky" does not mean "smooth". Leaky ReLU still has a corner at 0; it only fixes the zero slope on the left.
Quick check: evaluate leaky ReLU with $\alpha=0.1$ at $x=-4$, and ReLU at $x=-4$ and $x=2.5$.
Leaky ReLU at $-4$: negative input, so $\alpha x = 0.1\cdot(-4) = -0.4$. ReLU at $-4$ is $\max(0,-4) = 0$. ReLU at $2.5$ is $2.5$.
Functions of many inputs: a first look
So far the machine had one slot for input. Real models have many. A house price depends on area and number of bedrooms and age. A neural network's loss depends on millions of weights.
With two inputs we can still draw the picture: the two inputs $x$ and $y$ are positions on a floor, and the output is the height of a surface above that spot. A loss function is a landscape of hills and valleys. Training means walking downhill. The rest of this guide teaches you how to find "downhill" using calculus.
A machine can also have several outputs, such as a layer that turns 3 numbers into 4. Both ideas are previewed here and handled properly in Chapter 2.4 and Chapter 2.5.
Let $f(x, y) = \dfrac{x^2 + y^2}{4}$ (a bowl).
- $f(0, 0) = 0$: the bottom of the bowl.
- $f(2, 0) = 4/4 = 1$.
- $f(2, 2) = (4 + 4)/4 = 2$.
- $f(-2, 2) = (4 + 4)/4 = 2$ as well: the bowl is symmetric.
A function with two outputs: $\mathbf{F}(x) = (2x,\ x^2)$ turns the input 3 into the vector $(6, 9)$.
A function of $n$ inputs is written $f(x_1, \dots, x_n)$, or $f(\mathbf{x})$ with the inputs gathered into a vector $\mathbf{x}\in\mathbb{R}^n$. It maps $\mathbb{R}^n \to \mathbb{R}$. Its graph is a surface in $n+1$ dimensions; for $n=2$ it is a surface over the $xy$-plane.
A function with $m$ outputs, $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$, is a vector-valued function (like one layer of a network). A slice (hold $y$ fixed and let only $x$ move) turns the surface back into an ordinary one-input curve. Slopes of slices are the partial derivatives of Chapter 2.4.
Why do we need it?
Models depend on many numbers at once, so we need functions of many inputs. Seeing the surface gives you the right picture for loss landscapes, even though real ones have millions of directions.
Where is it used?
The loss of any model as a function of its weights, the prediction as a function of many features, a neural-network layer ($\mathbb{R}^n \to \mathbb{R}^m$), and optimisation landscapes in gradient descent.
How is it used?
Think of the inputs as a point on a floor and the output as a height. To study one input at a time, take a slice. To improve a model, move the inputs in the direction that lowers the height.
Quick check: for $f(x,y) = x^2 - y^2$, what are $f(2,0)$ and $f(0,2)$?
$f(2,0) = 4 - 0 = 4$ and $f(0,2) = 0 - 4 = -4$. Same distance from the origin, opposite heights: this is the saddle (up along one axis, down along the other).
Recap, cheat sheet and practice
- A function gives exactly one output for each allowed input. Its domain is the allowed inputs; its range is the outputs that occur.
- Composition $f(g(x))$ chains machines (order matters); an inverse $f^{-1}$ undoes $f$ and is its mirror image in $y=x$, and exists only for one-to-one functions.
- Lines $mx+b$ have a constant slope; polynomials add powers (degree = flexibility, and too much flexibility overfits).
- Exponentials $b^x$ multiply by a fixed factor per step; $e^x$ is the natural one. Logarithms are their inverse and turn products into sums (log-likelihood, no underflow).
- Trig functions are the unit circle seen as waves: use radians, period $2\pi/\omega$.
- Sigmoid squashes any score into $(0,1)$, with saturated flat tails. Softmax turns a score list into probabilities that add to 1 (shift-invariant, temperature-controlled). Two-class softmax is a sigmoid.
- Piecewise functions (ReLU, leaky ReLU, $|x|$, step, clip, Huber) use a different rule per region; glued ReLU hinges can draw any curve.
- Looking ahead: the slope of each function is what calculus will compute, and a loss is a function of many inputs.
Cheat sheet
| Function | Formula | Domain → range | Remember |
|---|---|---|---|
| Line | $mx + b$ | $\mathbb{R}\to\mathbb{R}$ | slope = rise / run |
| Polynomial | $a_nx^n+\dots+a_0$ | $\mathbb{R}\to$ depends | at most $n$ roots, $n-1$ bends |
| Exponential | $b^x$, $e^x$ | $\mathbb{R}\to(0,\infty)$ | $b^{x+y}=b^xb^y$; passes $(0,1)$ |
| Logarithm | $\ln x$, $\log_b x=\frac{\ln x}{\ln b}$ | $(0,\infty)\to\mathbb{R}$ | $\ln(xy)=\ln x+\ln y$ |
| Sine / cosine | $A\sin(\omega x+\varphi)+D$ | $\mathbb{R}\to[D-A,\ D+A]$ (that is $[-1,1]$ for plain $\sin$, $\cos$) | radians; period $2\pi/\omega$; $\sin^2+\cos^2=1$ |
| Sigmoid | $\dfrac{1}{1+e^{-x}}$ | $\mathbb{R}\to(0,1)$ | $\sigma(-x)=1-\sigma(x)$; inverse is logit; max slope $\tfrac14$ |
| Softmax | $\dfrac{e^{z_i}}{\sum_j e^{z_j}}$ | $\mathbb{R}^n\to$ probabilities | sums to 1; subtract max; temperature $z/T$ |
| ReLU / leaky | $\max(0,x)$ / $\max(\alpha x, x)$ | $\mathbb{R}\to[0,\infty)$ / $\mathbb{R}$ | kink at 0; slope 0 or 1 (or $\alpha$) |
| Absolute value, clip | $|x|$, $\min(\max(x,-c),c)$ | $\mathbb{R}\to[0,\infty)$ / $[-c,c]$ | corners, not jumps |
import numpy as np
# 1. A function is a rule: write it as code
def f(x):
return 2 * x + 3
print(f(5), f(-2)) # 13 -1
# 2. Composition and the inverse
g = lambda x: x + 2
sq = lambda x: x ** 2
print(sq(g(3)), g(sq(3))) # 25 11 (order matters)
f_inv = lambda y: (y - 3) / 2
print(f_inv(f(5))) # 5.0
# 3. Exponential and logarithm are inverses
x = np.array([0.5, 1.0, 2.0])
print(np.exp(x)) # [1.64872127 2.71828183 7.3890561 ]
print(np.log(np.exp(x))) # [0.5 1. 2. ]
print(np.log2(32), np.log10(1000)) # 5.0 3.0
# 4. Trig functions use radians
print(np.sin(np.pi / 6)) # 0.49999999999999994 (that is 1/2, up to rounding)
print(np.sin(np.deg2rad(90))) # 1.0
# 5. Sigmoid, logit and tanh
def sigmoid(z):
return 1 / (1 + np.exp(-z))
def logit(p):
return np.log(p / (1 - p))
z = np.array([-5.0, 0.0, 1.0, 5.0])
print(sigmoid(z)) # [0.00669285 0.5 0.73105858 0.99330715]
print(logit(sigmoid(1.0))) # 1.0 (up to rounding)
print(np.tanh(1.0), 2 * sigmoid(2.0) - 1) # 0.7615941559557649 0.7615941559557646
# 6. Softmax (stable version: subtract the max first)
def softmax(z, T=1.0):
z = np.asarray(z, dtype=float) / T
e = np.exp(z - z.max())
return e / e.sum()
print(softmax([2.0, 1.0, 0.1])) # [0.65900114 0.24243297 0.09856589]
print(softmax([2.0, 1.0, 0.1], T=0.5)) # [0.86377712 0.11689952 0.01932336]
print(softmax([1002.0, 1001.0, 1000.1])) # [0.65900114 0.24243297 0.09856589] (same: shift invariance)
# 7. Piecewise functions
relu = lambda x: np.maximum(0, x)
leaky = lambda x, a=0.1: np.where(x >= 0, x, a * x)
clip = lambda x, c=1.5: np.clip(x, -c, c)
v = np.array([-4.0, -1.0, 0.0, 2.5])
print(relu(v)) # [0. 0. 0. 2.5]
print(leaky(v)) # [-0.4 -0.1 0. 2.5]
print(clip(v)) # [-1.5 -1. 0. 1.5]
# 8. Why we add logs: 400 probabilities of 0.1
print(0.1 ** 400) # 0.0 (underflow!)
print(400 * np.log(0.1)) # -921.0340371976182
1. What is the domain of $f(x) = \sqrt{5 - x}$?
2. Let $f(x) = 2x$ and $g(x) = x^2$. What is $f(g(3))$?
3. You are told $\sigma(3) \approx 0.953$. What is $\sigma(-3)$?
4. What happens to the softmax output if you add 10 to every logit?
5. For positive $x$ and $y$, which statement is always true?
6. What is a "dead" ReLU unit?
Practice problems
A. Find the inverse of $f(x) = \dfrac{2x+1}{3}$ and check it at $x=2$.
Write $y = \dfrac{2x+1}{3}$. Multiply by 3: $3y = 2x + 1$. Subtract 1: $3y - 1 = 2x$. Divide by 2: $x = \dfrac{3y-1}{2}$. So $f^{-1}(x) = \dfrac{3x-1}{2}$. Check: $f(2) = 5/3$, and $f^{-1}(5/3) = (5 - 1)/2 = 2$ ✓.
B. Factor $p(x) = x^3 - 4x$, list its roots, say how the ends behave, and compute $p(1)$.
$x^3 - 4x = x(x^2 - 4) = x(x-2)(x+2)$. The roots are $-2$, $0$ and $2$ (three roots, so two bends). Odd degree with positive leading coefficient: the left end goes down and the right end goes up. $p(1) = 1 - 4 = -3$ (negative, as expected between the roots $0$ and $2$).
C. Solve $3e^{2x} = 21$ for $x$.
Divide by 3: $e^{2x} = 7$. Take $\ln$ of both sides: $2x = \ln 7 \approx 1.9459$. So $x = \tfrac12\ln 7 \approx 0.973$. Check: $3e^{1.946} = 3 \times 7.00 = 21$ ✓.
D. A model's score for an email is $z = \ln 4$. What probability does the sigmoid give? What is the logit of that probability?
$\sigma(\ln 4) = \dfrac{1}{1 + e^{-\ln 4}} = \dfrac{1}{1 + \tfrac14} = \dfrac{1}{5/4} = 0.8$. The logit goes back: $\ln\dfrac{0.8}{0.2} = \ln 4 \approx 1.386$ ✓ (the odds are 4 to 1).
E. Compute the softmax of the logits $[0,\ \ln 2,\ \ln 3]$.
$e^0 = 1$, $e^{\ln 2} = 2$, $e^{\ln 3} = 3$. The sum is $6$. So the probabilities are $[\tfrac16, \tfrac26, \tfrac36] = [0.167, 0.333, 0.5]$, and they add to 1 ✓.
F. The Huber loss with $\delta = 1$ is $\tfrac12 x^2$ if $|x|\le1$, else $|x| - \tfrac12$. Compute it at $x = 0.5$ and $x = 3$, and compare with $\tfrac12x^2$ at $x=3$. Why is Huber popular with outliers?
At $x = 0.5$: $|x|\le 1$, so $\tfrac12(0.25) = 0.125$. At $x=3$: $|x|>1$, so $3 - 0.5 = 2.5$. The plain squared loss at 3 would be $\tfrac12\cdot9 = 4.5$. For a big error (an outlier) Huber grows only in a straight line, so one wild point cannot dominate the loss, while near zero it keeps the smooth bowl shape.
Limits & Continuity
A limit asks a gentle question: as the input creeps closer and closer to a spot, what value is the output heading towards? That is all we need. In the next chapter, the derivative will turn out to be one single limit. So this chapter is deliberately short on formalism and long on pictures: we want you to feel what "getting closer and closer" means.
- Understand a limit as "where the output is heading", and read it from a table and a zoomed graph
- Use one-sided and two-sided limits, and know when a limit does not exist
- Understand infinite limits (vertical asymptotes) and limits at infinity (horizontal asymptotes)
- Define continuity and recognise the three kinds of discontinuity: removable, jump, infinite
- Know and derive the key limit identities: $\frac{\sin x}{x}\to1$, $(1+\frac1n)^n\to e$, $\frac{e^x-1}{x}\to 1$, $\frac{\ln(1+x)}{x}\to 1$
- See why the derivative is a limit, so you are ready for Chapter 2.3
This chapter uses the function toolbox from Chapter 2.1 (domain, exponentials, logs, sine, sigmoid). We only need limits well enough to understand derivatives. So the proofs are light, and the one formal definition ($\varepsilon$-$\delta$) is a clearly marked optional glimpse.
The intuition of a limit: sneaking up on a value core
Imagine walking towards a door. You take a step, then half the remaining distance, then half again. You never quite touch the door in this game, but it is perfectly clear where you are heading: the door. A limit is that "where you are heading".
Now the important twist. A limit does not care what happens at the door. It only cares about the values on the way there. The function may be missing a value at that spot, or have a wrong one, and the limit does not change.
Why would anyone need this? Think of a car's speedometer. "Speed at exactly 3:00:00" would be distance divided by zero time, which is $0/0$ and meaningless. But the average speed over the last second, then the last tenth of a second, then the last thousandth, settles down to a clear number. That number is the limit, and it is the speed. This is exactly how the derivative is defined.
Let $f(x) = \dfrac{x^2 - 1}{x - 1}$. At $x = 1$ it is $\frac{0}{0}$, so $f(1)$ does not exist. But what is $f$ heading to as $x$ gets close to 1?
| $x$ | 0.9 | 0.99 | 0.999 | 1.001 | 1.01 | 1.1 | |
|---|---|---|---|---|---|---|---|
| $f(x)$ | 1.9 | 1.99 | 1.999 | ? | 2.001 | 2.01 | 2.1 |
Check one entry: $f(0.9) = \dfrac{0.81 - 1}{0.9 - 1} = \dfrac{-0.19}{-0.1} = 1.9$ ✓. From both sides, the output heads to 2. We write $\displaystyle\lim_{x\to1}\frac{x^2-1}{x-1} = 2$, even though $f(1)$ itself does not exist.
Another one: $\dfrac{\sin x}{x}$ is also $\frac00$ at $x=0$. At $x = 0.1$ it equals $0.99833$, at $x=0.01$ it equals $0.99998$. It is heading to 1.
We write $$\lim_{x \to a} f(x) = L$$ and say "the limit of $f(x)$ as $x$ approaches $a$ is $L$" when $f(x)$ gets as close to $L$ as we like by taking $x$ close enough to $a$ (but $x\ne a$).
- "$x\to a$" means $x$ gets closer and closer to $a$, without ever being equal to it. That is why $0/0$ spots are fine.
- The value $f(a)$ plays no role. It may be missing, or different from $L$.
- If $f(x)$ does not head to a single number, we say the limit does not exist (DNE).
- Two ways to find a limit: numerically (a table of values nearer and nearer) and graphically (zoom in on the point). Later we also use algebra.
Why do we need it?
Instant quantities, such as the speed at one instant or the slope at one point, come out as $0/0$ when computed directly. Limits give a clean meaning to "what the ratio becomes as the gap shrinks to nothing".
Where is it used?
The derivative is a limit (Chapter 2.3), and so is an integral (Chapter 2.13). Limits also describe training that "converges" (loss approaching its minimum), infinite series (Taylor, Chapter 2.11) and a function's behaviour at extreme inputs, as in a saturating sigmoid.
How is it used?
Build a table of $f(x)$ for $x$ closer and closer to $a$ from both sides, or zoom the graph in. If the values settle, that number is the limit. If an algebraic trick (cancelling) is available, use it to be sure.
A limit is not a value of the function. $\lim_{x\to a}f(x)$ and $f(a)$ can be different, or one can exist without the other. We care about the neighbourhood, not the point itself.
"Closer and closer" is not "reaches". $x\to a$ never equals $a$. The sequence $0.9, 0.99, 0.999, \dots$ approaches 1 but is never 1, and 1 is still the limit.
A table can fool you. If a function wiggles very fast, a few sample points may look settled while the true limit does not exist. Pair tables with graphs and algebra (we meet such a case, $\sin(1/x)$, in the next section).
Quick check: if $g(x) = x + 1$ for $x \ne 3$ and $g(3) = 100$, what is $\lim_{x\to3} g(x)$?
The values near 3 (like $2.99 \to 3.99$ and $3.01\to4.01$) head to $4$. The odd value $g(3)=100$ is ignored. The limit is $4$.
One-sided limits: left and right core
Picture a river with a broken bridge. You can walk up to the gap from the left bank or from the right bank. Where you are standing as you reach the edge may be different on each side: the two banks may be at different heights.
A function can do the same. Approaching $a$ from the left (only using $x\lt a$) may head to one value, and approaching from the right (only $x>a$) may head to another. We give each side its own limit.
Let $f(x) = x$ for $x < 1$, and $f(x) = x^2 + 1$ for $x \ge 1$.
- From the left: $f(0.9) = 0.9$, $f(0.99) = 0.99$, $f(0.999)=0.999$: heading to $\mathbf{1}$.
- From the right: $f(1.1) = 1.21+1 = 2.21$, $f(1.01) = 1.0201 + 1 = 2.0201$, $f(1.001) \approx 2.002$: heading to $\mathbf{2}$.
Different banks, different heights: the function jumps at $x=1$. (And $f(1) = 2$ is just the value the function actually takes there.)
More cases. $\sqrt{x}$ near 0 only exists on the right ($x \ge 0$): $\lim_{x\to0^+}\sqrt x = 0$, and the left side makes no sense. The step function (0 for $x<0$, 1 for $x\ge0$) has left limit 0 and right limit 1. ReLU near 0 has left limit 0 and right limit 0 (no jump, only a corner).
The left-hand limit $\displaystyle\lim_{x\to a^-} f(x) = L^-$ uses only $x\lt a$ ("$a^-$" means "from below $a$"). The right-hand limit $\displaystyle\lim_{x\to a^+} f(x) = L^+$ uses only $x>a$.
- Each one can be a number, $+\infty$, $-\infty$, or fail to exist.
- They can differ (a jump), or agree.
- A one-sided limit may not even make sense if the function does not exist on that side of $a$ (like $\sqrt x$ at $0^-$).
- Some functions have neither: $\sin(1/x)$ near 0 oscillates between $-1$ and $1$ faster and faster, and never settles.
Why do we need it?
Functions can behave differently on the two sides of a point: jumps, switches, corners. Looking at each side separately tells us exactly what kind of change happens there.
Where is it used?
The step activation and threshold decisions, piecewise losses like Huber, and ReLU's corner. In Chapter 2.3, ReLU's "slope from the left" is 0 and "slope from the right" is 1: one-sided limits of slopes.
How is it used?
Evaluate the function at points just left and just right of $a$ (for example $a\pm0.001$) and compare. Different results mean a jump; both heading to $\pm\infty$ means a blow-up.
Quick check: for the step function ($0$ for $x<0$, $1$ for $x\ge0$), what are $\lim_{x\to0^-}$ and $\lim_{x\to0^+}$?
From the left the values are all $0$, so the left-hand limit is $0$. From the right the values are all $1$, so the right-hand limit is $1$. They differ: a jump.
Two-sided limits, and how to compute them core
The ordinary limit $\lim_{x\to a}f(x)$ allows approach from both sides at once. It is only meaningful if the function agrees with itself no matter which road you take. Two walkers, one from each bank, must arrive at the same spot. If they arrive at different heights, we say the limit does not exist.
To compute a limit you have three tools: a table (numbers), a graph (picture), and algebra. The algebra trick for the common $0/0$ trouble is: simplify first, so the troublemaker disappears, then substitute.
Example 1: factor and cancel. $\displaystyle\lim_{x\to1}\frac{x^2-1}{x-1}$.
- Substitute $x=1$: $\frac{0}{0}$. Stuck, so simplify first.
- Factor the top: $x^2-1 = (x-1)(x+1)$.
- Cancel $(x-1)$. This is allowed because $x\to1$ means $x\ne1$, so $x-1\ne0$. The expression equals $x+1$ for every $x\ne1$.
- Now substitute: $1+1 = \mathbf{2}$ ✓ (matches the table in the first section).
Example 2: the conjugate trick. $\displaystyle\lim_{x\to0}\frac{\sqrt{x+4}-2}{x}$.
- Substitute $x = 0$: $\frac{2-2}{0} = \frac00$. Stuck.
- Multiply top and bottom by the conjugate $\sqrt{x+4}+2$ (this multiplies by 1): $\dfrac{(\sqrt{x+4}-2)(\sqrt{x+4}+2)}{x(\sqrt{x+4}+2)}$.
- The top is $(x+4) - 4 = x$, using $(A-B)(A+B) = A^2-B^2$.
- Cancel $x$: we get $\dfrac{1}{\sqrt{x+4}+2}$.
- Substitute $x=0$: $\dfrac{1}{2+2} = \mathbf{\tfrac14}$.
Example 3: a limit that does not exist. $\dfrac{|x|}{x}$ at 0: from the right it is $+1$, from the left $-1$. The two sides disagree, so there is no limit.
Two-sided limit rule. $\displaystyle\lim_{x\to a}f(x) = L$ exactly when $$\lim_{x\to a^-}f(x) = L \quad\text{and}\quad \lim_{x\to a^+}f(x) = L.$$ For a limit that is a number, both one-sided limits must exist, be finite, and be equal. (If both sides blow up the same way we write $\pm\infty$, see the next sections.)
Limit laws (if $\lim f = L$ and $\lim g = M$ both exist):
- Sum / difference: $\lim (f \pm g) = L \pm M$. Constant multiple: $\lim (c f) = cL$.
- Product: $\lim (f g) = LM$. Quotient: $\lim (f/g) = L/M$, provided $M\ne0$.
- Power: $\lim (f^n) = L^n$. Compositions of nice functions can be done "from the inside out".
- Direct substitution: for polynomials, exponentials, sine, cosine and (where defined) logs and roots, just plug in $a$: $\lim_{x\to a} f(x) = f(a)$. Trouble ($0/0$) means: simplify first.
Why do we need it?
We need rules that let us compute limits exactly, rather than guess from a table. The $0/0$ case is the important one: it is exactly the form that a derivative takes.
Where is it used?
Deriving every derivative rule in Chapter 2.3 (each one is a limit simplified by algebra), proving that loss functions are well behaved, and checking numerical code where "$0/0$" would crash (for example sin(x)/x at $x=0$).
How is it used?
1. Try direct substitution. 2. If you get $0/0$, factor, cancel or multiply by a conjugate. 3. Substitute again. 4. Cross-check with a table or graph, and test both sides if the function might jump.
"$0/0$" is not an answer; it is a signal. It means "I need to simplify". The limit of a $0/0$ expression can be any number (we got 2, then $\tfrac14$, then 6) or may not exist.
Cancel only common factors, never common terms. In $\frac{x^2-1}{x-1}$ you may cancel the factor $(x-1)$, but you may not "cancel the $x$'s" in $\frac{x+3}{x}$.
Quick check: find $\displaystyle\lim_{x\to2}\frac{x^2-4}{x-2}$.
Substituting gives $0/0$. Factor: $x^2-4 = (x-2)(x+2)$. Cancel $(x-2)$: left with $x+2$. At $x=2$: $4$. The limit is $4$.
Infinite limits: when the output blows up
Divide 1 by a smaller and smaller positive number: $1/0.1 = 10$, $1/0.01 = 100$, $1/0.001 = 1000$. The answer grows without any ceiling. Pick any huge number you like; if $x$ is close enough to 0, $1/x$ is bigger than it.
We describe this by saying the limit is infinity. Infinity is not a number you can reach; it is a short way to say "grows without bound". On a graph, the curve shoots up (or down) along a vertical line it never touches: a vertical asymptote.
- $\dfrac1{x^2}$ near 0: $x = \pm0.1\to100$, $x=\pm0.01\to10\,000$. Both sides go to $+\infty$: $\displaystyle\lim_{x\to0}\frac{1}{x^2} = +\infty$.
- $\dfrac1{x}$ near 0: from the right $0.01\to 100$ (up); from the left $-0.01\to-100$ (down). So $\lim_{x\to0^+}\frac1x = +\infty$ and $\lim_{x\to0^-}\frac1x = -\infty$. The two-sided limit does not exist, since the sides disagree.
- In ML: the cross-entropy loss for a correct class with predicted probability $p$ is $-\ln p$. As $p\to0^+$ it blows up: $-\ln 0.01 = 4.6$, $-\ln 10^{-6} = 13.8$, and it keeps growing without bound. A model that is confidently wrong is punished without limit.
$\displaystyle\lim_{x\to a} f(x) = +\infty$ means: $f(x)$ becomes larger than any chosen number, as soon as $x$ is close enough to $a$. Similarly for $-\infty$ (more negative than any number). One-sided versions: $x\to a^-$, $x\to a^+$.
- Strictly speaking the limit "does not exist" (it is not a real number), but writing $\pm\infty$ tells us how it fails.
- If $\lim_{x\to a^\pm}f(x)=\pm\infty$ (on at least one side), the line $x=a$ is a vertical asymptote of the graph.
- Typical sources: a fraction whose bottom $\to0$ while the top does not, e.g. $\dfrac{1}{x-a}$, $\dfrac{1}{(x-a)^2}$, and $\ln x \to -\infty$ as $x\to0^+$.
- Sign rule: $\dfrac{\text{positive}}{\text{tiny positive}} \to +\infty$, $\dfrac{\text{positive}}{\text{tiny negative}} \to -\infty$.
Why do we need it?
We need to describe functions that explode at some point, and to recognise danger spots where a formula must not be used (division by a vanishing number).
Where is it used?
The cross-entropy loss $-\ln p$ (infinite at $p=0$), numerical instability when dividing by tiny values (for example normalising by a standard deviation that is almost 0), and the logit $\ln\frac{p}{1-p}$, which blows up at $p=0$ and $p=1$.
How is it used?
Find where the denominator (or the log's input) becomes 0, then check each side's sign. In code, add a tiny constant ($10^{-8}$) or clip probabilities away from 0 and 1, so the blow-up never actually happens.
$\infty$ is not a number. You cannot do $\infty - \infty$ or $\frac{\infty}{\infty}$ as ordinary arithmetic. We only use $\infty$ to describe a trend.
Mind the sides. $\frac{1}{x}$ and $\frac{1}{x^2}$ both have a vertical asymptote at 0, but they behave differently on the left side.
Quick check: what are $\lim_{x\to2^+}\frac{1}{x-2}$ and $\lim_{x\to2^-}\frac{1}{x-2}$?
From the right, $x-2$ is a tiny positive number, so $1/(x-2)\to+\infty$. From the left, $x-2$ is a tiny negative number, so $1/(x-2)\to-\infty$.
Limits at infinity: where does the curve level off? core
Now change the question. Instead of asking what happens near a point, ask: what happens far away, as $x$ gets enormous? Does the output settle down to a steady value, grow forever, or keep wiggling?
Think of a cup of coffee cooling in a room. After a long time it settles at room temperature, and that steady value is the limit as time goes to infinity. On a graph this is a horizontal asymptote: a flat line the curve gets closer and closer to.
The sigmoid is the perfect example: as the score gets very large, the output creeps up to 1 and never passes it. That is exactly the "saturation" from Chapter 2.1.
- $\displaystyle\lim_{x\to\infty}\frac{1}{x} = 0$ (since $1/10 = 0.1$, $1/1000 = 0.001$, …).
- $\displaystyle\lim_{x\to\infty}\sigma(x) = 1$ and $\displaystyle\lim_{x\to-\infty}\sigma(x) = 0$. Also $e^{-x}\to0$ as $x\to\infty$.
- A rational function. $\displaystyle\lim_{x\to\infty}\frac{3x^2+1}{x^2+2}$. Divide the top and bottom by the highest power, $x^2$:
- $\dfrac{3x^2+1}{x^2+2} = \dfrac{3 + 1/x^2}{1 + 2/x^2}$.
- As $x\to\infty$, $1/x^2\to0$ and $2/x^2\to0$.
- So the limit is $\dfrac{3+0}{1+0} = \mathbf{3}$.
$\displaystyle\lim_{x\to\infty} f(x) = L$ means $f(x)$ gets as close to $L$ as we like once $x$ is large enough. Then the line $y = L$ is a horizontal asymptote. Likewise for $x\to-\infty$. (The limit can also be $\pm\infty$ if the function grows forever.)
- Rational functions $\dfrac{\text{degree } m}{\text{degree } n}$: if $m\lt n$, the limit is $0$; if $m=n$, it is the ratio of the leading coefficients; if $m>n$, it is $\pm\infty$. (Divide by the highest power of $x$ in the bottom to see why.)
- Exponential beats polynomial: $\displaystyle\lim_{x\to\infty}\frac{x^k}{e^x} = 0$ for every $k$. (Numerically: $20^2/e^{20}\approx 8\times10^{-7}$.) And logarithms grow slower than any power: $\ln x/x\to0$.
- Sine and cosine have no limit at infinity (they keep oscillating), but $\dfrac{\sin x}{x}\to 0$ because the top stays between $-1$ and $1$ while the bottom grows.
Why do we need it?
We want to know what a function does for extreme inputs and in the long run: does it saturate, settle, or explode? This predicts how a model behaves for very large scores or after very many training steps.
Where is it used?
Saturation of sigmoid, tanh and softmax for large logits; learning-rate decay to 0 and weight decay $e^{-\lambda t}$; a loss that levels off at its minimum; and why $e^{-x}$ weights in kernels vanish for far-away points.
How is it used?
Divide top and bottom by the highest power of $x$ to settle rational functions. Compare growth rates (log, polynomial, exponential) for the rest. Then read the horizontal asymptote off the result.
A curve can cross its horizontal asymptote. $\frac{\sin x}{x}$ crosses the line $y=0$ again and again, yet its limit is 0. The asymptote describes the long-run trend, not a barrier.
Computers are not limits. Give a computer $x = 10^{400}$ and it overflows. Limits describe the exact mathematics; code must still avoid overflow (this is why softmax subtracts the maximum).
Quick check: find $\displaystyle\lim_{x\to\infty}\frac{5x+2}{2x-7}$.
Same degree on top and bottom. Divide by $x$: $\dfrac{5 + 2/x}{2 - 7/x} \to \dfrac{5}{2}$. The limit is $2.5$ (the ratio of the leading coefficients).
Continuity: draw it without lifting the pencil core
A function is continuous if you can draw its graph without lifting your pencil: no holes, no jumps, no sudden vertical escapes. In everyday terms: a small change in the input only causes a small change in the output. Nothing "snaps".
A bike's speed is continuous: it cannot jump from 10 to 30 km/h in zero time. A light switch, on the other hand, is not.
Check three functions at $x=1$:
- $f(x)=x^2$: $f(1)=1$, and the limit as $x\to1$ is $1$. They match: continuous.
- $f(x)=\dfrac{x^2-1}{x-1}$: $f(1)$ does not exist (hole). Not continuous at 1. (If we define $f(1)=2$, the hole is filled and it becomes continuous.)
- Step function at $0$: the left limit is 0, the right limit is 1. No single limit: not continuous.
The intermediate value idea. A continuous function cannot skip values. Let $p(x)=x^3-x-1$. Then $p(1)=-1$ and $p(2)=5$: it goes from negative to positive, so it must cross 0 somewhere in between. Halving the interval: $p(1.5) = 0.875>0$, so the root is in $(1, 1.5)$; $p(1.25)\approx-0.297<0$, so it is in $(1.25, 1.5)$; $p(1.375)\approx0.225>0$, so it is in $(1.25, 1.375)$. This "bisection" method homes in on the root $\approx1.3247$.
A function $f$ is continuous at $x=a$ if all three hold:
- $f(a)$ is defined,
- $\displaystyle\lim_{x\to a}f(x)$ exists,
- $\displaystyle\lim_{x\to a}f(x) = f(a)$.
In one line: $\displaystyle\lim_{x\to a}f(x)=f(a)$ ("the limit equals the value"). $f$ is continuous on an interval if it is continuous at every point of it.
- Continuous everywhere: polynomials, $e^x$, $\sin x$, $\cos x$, $|x|$, ReLU, sigmoid, tanh.
- Continuous where defined: $\ln x$ on $x>0$, $\sqrt x$ on $x\ge0$, and rational functions $\frac{p}{q}$ wherever $q\ne0$.
- Building blocks: sums, products, quotients (bottom $\ne0$) and compositions of continuous functions are continuous. So a network made only of continuous layers and activations is continuous.
- Intermediate value theorem: if $f$ is continuous on $[a,b]$ and $f(a)$, $f(b)$ have opposite signs, then $f(x) = 0$ for some $x$ between.
- Preview: to have a derivative at a point, a function must be continuous there (Chapter 2.3). The reverse is false: ReLU and $|x|$ are continuous but have a corner at 0.
Why do we need it?
Continuity is the promise that "nearby inputs give nearby outputs". Without it, a tiny nudge could cause a huge change, and we could neither predict nor train by small steps.
Where is it used?
Gradient-based training needs a loss that varies smoothly with the weights; stable predictions (small input noise should not flip the output wildly); root-finding by bisection; and the guarantee that composing layers keeps a network continuous.
How is it used?
Test the three conditions at suspicious points (where a denominator is 0, where pieces join). Between those points, trust the building-block rules. Prefer continuous activations and losses when you plan to differentiate.
Continuous does not mean smooth. $|x|$ and ReLU are continuous (no pencil lift) but have a sharp corner at 0. Continuity is the first level of "well-behaved"; having a derivative is the next.
Check continuity at a point or on an interval. $\frac1x$ is continuous on $x>0$ and on $x<0$, but it is not continuous on any interval containing 0, because 0 is not in its domain.
Quick check: is $f(x) = \dfrac{x}{x-2}$ continuous at $x=2$? at $x=5$?
At $x=2$ the function is not defined (division by 0), so condition 1 fails: not continuous there (it blows up: an infinite discontinuity). At $x=5$ it is a quotient of continuous functions with a non-zero bottom, so it is continuous: $f(5) = 5/3$.
Discontinuities: removable, jump and infinite core
When the pencil must lift, it can lift in three different ways:
- Removable (a hole). One point is missing or misplaced, but the curve on both sides lines up. Just put the dot where it belongs and the problem is fixed.
- Jump. The two banks of the river sit at different heights. No dot can repair it.
- Infinite. The curve shoots off to $\pm\infty$ along a vertical asymptote.
| Type | Example at the problem spot | One-sided limits | Fixable? |
|---|---|---|---|
| Removable | $\dfrac{x^2-1}{x-1}$ at $x=1$ | both $= 2$ (equal) | yes: define $f(1)=2$ |
| Jump | step function at $0$; a tiered price that leaps from 10 to 20 dollars at 5 GB | $0$ and $1$ (different) | no |
| Infinite | $\dfrac1x$ at $x=0$ | $-\infty$ and $+\infty$ | no |
A fourth, rarer kind is oscillating: $\sin(1/x)$ at 0 swings between $-1$ and $1$ infinitely often, so no one-sided limit exists at all.
In ML. Accuracy is a step-like function of the weights: it only changes when a prediction flips, so it is made of flat pieces with jumps. Its slope is 0 almost everywhere, which tells gradient descent nothing. That is why we train on a smooth, continuous surrogate (cross-entropy or MSE) instead of accuracy itself.
Let $f$ be not continuous at $a$. Look at the one-sided limits $L^-$ and $L^+$:
- Removable: $L^- = L^+ = L$ is a finite number, but $f(a)$ is undefined or $\ne L$. Fix: redefine $f(a)=L$.
- Jump: $L^-$ and $L^+$ are both finite but $L^-\ne L^+$. The jump size is $L^+-L^-$.
- Infinite: at least one of $L^-$, $L^+$ is $\pm\infty$; the line $x=a$ is a vertical asymptote.
Jump and infinite discontinuities are called non-removable. A function is continuous at $a$ exactly when none of these (nor the oscillating kind) happens.
Why do we need it?
Knowing which kind of break a function has tells you what to do: patch a removable hole, treat a jump with care, or avoid an infinite spot completely.
Where is it used?
Spotting the problem of threshold and step activations, analysing piecewise losses, finding where a formula such as $\frac{\sin x}{x}$ or $\frac{p}{1-p}$ needs a special case in code, and deciding which loss can be differentiated.
How is it used?
Compute the two one-sided limits and $f(a)$. Equal limits with a wrong or missing value: removable, and you can patch it in code with an if for that single point. Different finite limits: jump. Infinite: asymptote.
"Removable" means the limit exists. That is the only requirement. A hole does not stop the function from having a clear limit, so it does no harm to calculus; you only need to patch the single value.
A jump is not an asymptote. In a jump, both one-sided limits are ordinary finite numbers; the function stays bounded.
Quick check: classify the discontinuity of $f(x) = \dfrac{x^2-9}{x-3}$ at $x=3$, and of $g(x) = \dfrac{1}{x-3}$ at $x=3$.
$f$: $(x-3)(x+3)/(x-3) = x+3$, so both sides head to $6$ while $f(3)$ is undefined: removable (define $f(3) = 6$). $g$: one side $\to+\infty$, the other $\to-\infty$: infinite.
Important limit identities (and how to derive them) core
A handful of limits turn up again and again, because they are exactly what the derivative of $\sin x$, $e^x$ and $\ln x$ is made of. Each is a $0/0$ puzzle with a clean answer:
- $\dfrac{\sin x}{x}\to 1$: for a tiny angle, the sine is almost the angle itself (a tiny arc is almost straight).
- $\left(1+\dfrac1n\right)^n \to e$: compound interest. Splitting a year into more and more steps gives a bigger and bigger payout, but it levels off at $e\approx2.71828$.
- $\dfrac{e^x-1}{x}\to1$ and $\dfrac{\ln(1+x)}{x}\to1$: near 0, $e^x$ is almost $1+x$ and $\ln(1+x)$ is almost $x$.
You do not need to memorise them. We derive them, step by step, from a picture or from $e$ itself.
Compound interest. Put 1 dollar in a bank that pays 100% per year. If interest is paid once, you have $(1+1)^1 = 2$. If it is paid twice (50% each half-year), you have $(1+\tfrac12)^2 = 2.25$. Four times: $(1+\tfrac14)^4 \approx 2.4414$. Monthly: $(1+\tfrac1{12})^{12}\approx 2.6130$. Daily: $\approx2.7146$. A million times a year: $\approx 2.718280$. It never passes $e = 2.718281828\ldots$
A tiny angle. $\sin(0.1) = 0.0998334$, so $\sin(0.1)/0.1 = 0.99833$. $\sin(0.01)/0.01 = 0.99998$.
Others at $x=0.01$: $\frac{e^{0.01}-1}{0.01} = 1.00502$, $\frac{\ln 1.01}{0.01} = 0.99503$, $\frac{1-\cos0.01}{0.01^2} = 0.499996$.
The identities (all as $x\to0$ unless stated): $$\lim\frac{\sin x}{x}=1,\quad \lim_{n\to\infty}\Big(1+\frac1n\Big)^n=e,\quad \lim\,(1+x)^{1/x}=e,$$ $$\lim\frac{e^x-1}{x}=1,\quad \lim\frac{\ln(1+x)}{x}=1,\quad \lim\frac{a^x-1}{x}=\ln a,\quad \lim\frac{1-\cos x}{x^2}=\frac12.$$ And at infinity: $x^k/e^x\to0$ and $\ln x/x\to0$ ("exponential beats polynomial beats logarithm").
Derivations.
- $e$ as a limit. This is the definition of $e$: $e=\lim_{n\to\infty}(1+\frac1n)^n$. Put $x = 1/n$ (so $x\to0^+$ when $n\to\infty$) to get $(1+x)^{1/x}\to e$. The table above shows the values creeping up to 2.71828.
- $\ln(1+x)/x\to1$. Rewrite with the power rule of logs: $\dfrac{\ln(1+x)}{x} = \tfrac1x\ln(1+x) = \ln\big((1+x)^{1/x}\big)$. As $x\to0$ the inside heads to $e$ (above we showed this from the right, $x=1/n>0$; from the left it also holds, and we use that without proof), and $\ln$ is continuous, so the whole thing heads to $\ln e = 1$.
- $(e^x-1)/x\to1$. Let $u = e^x - 1$. As $x\to0$, $u\to0$, and $x = \ln(1+u)$. Then $\dfrac{e^x-1}{x} = \dfrac{u}{\ln(1+u)} = \dfrac{1}{\ln(1+u)/u}\to\dfrac11 = 1$ (using the previous identity).
- $(a^x-1)/x\to\ln a$. Write $a^x = e^{x\ln a}$ and let $t = x\ln a\to0$: $\dfrac{e^{t}-1}{x} = \ln a\cdot\dfrac{e^t-1}{t}\to\ln a\cdot1$.
- $\sin x/x\to1$ (squeeze). On the unit circle, for $0\lt x<\frac\pi2$, compare three areas: the small triangle ($\tfrac12\sin x$) is inside the circular sector ($\tfrac12 x$), which is inside the big triangle ($\tfrac12\tan x$). So $\sin x\lt x<\tan x$. Divide by $\sin x$: $1<\dfrac{x}{\sin x}<\dfrac1{\cos x}$, and flipping: $\cos x<\dfrac{\sin x}{x}<1$. As $x\to0^+$, $\cos x\to1$, so $\frac{\sin x}{x}$ is squeezed between two things heading to 1. Since $\frac{\sin x}{x}$ has the same value at $-x$, the left side agrees.
- $(1-\cos x)/x^2\to\frac12$. Multiply top and bottom by $1+\cos x$: $\dfrac{1-\cos^2x}{x^2(1+\cos x)} = \dfrac{\sin^2 x}{x^2(1+\cos x)} = \Big(\dfrac{\sin x}{x}\Big)^2\dfrac{1}{1+\cos x}\to 1\cdot\dfrac12$.
The squeeze theorem used above: if $g(x)\le f(x)\le h(x)$ near $a$ and both $g$ and $h$ head to $L$, then $f$ heads to $L$ too.
Why do we need it?
These limits are the raw ingredients of derivatives: $(\sin x)'=\cos x$ comes from $\frac{\sin x}{x}\to1$, $(e^x)'=e^x$ from $\frac{e^x-1}{x}\to1$, and $(\ln x)'=\frac1x$ from $\frac{\ln(1+x)}{x}\to1$ (all in Chapter 2.3).
Where is it used?
The number $e$ in softmax and sigmoid, compound growth and continuous decay ($e^{-\lambda t}$), small-angle approximations in physics and robotics, and numerically stable code: np.expm1 and np.log1p exist precisely because $e^x-1$ and $\ln(1+x)$ lose accuracy near 0.
How is it used?
Spot the pattern (a $0/0$ with $\sin$, $e^x-1$ or $\ln(1+x)$), rewrite so it matches a known identity (sometimes by a substitution like $u=3x$), and read off the answer. Example: $\frac{\sin 3x}{x} = 3\cdot\frac{\sin 3x}{3x}\to3$.
Computers lose digits near 0. $\frac{1-\cos x}{x^2}$ with $x=10^{-8}$ prints 0.0 in floating point, because $\cos x$ rounds to exactly 1 and the subtraction wipes out the answer. The identities say the true value is $\frac12$. Use the stable forms (np.expm1, np.log1p, $2\sin^2(x/2)$).
$(1+\frac1n)^n$ is not "1 to a power". The base tends to 1 and the exponent tends to infinity at the same time, a tug of war. Neither side wins: the result is $e$. And at absurd sizes ($n\approx10^{16}$) the computer rounds $1+\frac1n$ to exactly 1 and returns 1.
Quick check: find $\displaystyle\lim_{x\to0}\frac{\sin 3x}{x}$ and $\displaystyle\lim_{x\to0}\frac{e^{2x}-1}{x}$.
First: $\dfrac{\sin3x}{x} = 3\cdot\dfrac{\sin 3x}{3x}$. With $u=3x\to0$ the fraction $\to1$, so the limit is $3$. Second: $\dfrac{e^{2x}-1}{x} = 2\cdot\dfrac{e^{2x}-1}{2x}\to2\cdot1 = 2$.
Optional: the $\varepsilon$–$\delta$ glimpse
This section is optional. You can skip it and still follow the rest of the guide. It shows how mathematicians make "gets as close as we like" precise.
Think of a game between a challenger and you. The challenger names a tolerance on the output, $\varepsilon$ (the Greek letter epsilon): "I want $f(x)$ to be within $\varepsilon$ of $L$". You must answer with a tolerance on the input, $\delta$ (delta): "Then keep $x$ within $\delta$ of $a$, and I guarantee it."
If you can answer every challenge, however tiny the $\varepsilon$, then $L$ really is the limit.
Let $f(x)=2x+1$, $a=1$, $L=3$. The challenger says $\varepsilon = 0.1$.
- We need $|f(x) - 3| < 0.1$, that is $|2x+1-3| = |2x-2| = 2|x-1| < 0.1$.
- Divide by 2: $|x - 1| < 0.05$.
- So $\delta = 0.05$ works. In general, for any $\varepsilon$ choose $\delta = \varepsilon/2$: then $|x-1|<\delta$ gives $|f(x)-3| = 2|x-1| < 2\delta = \varepsilon$ ✓.
$\displaystyle\lim_{x\to a}f(x)=L$ means: for every $\varepsilon>0$ there is a $\delta>0$ such that $$0<|x-a|<\delta \ \Longrightarrow\ |f(x)-L|<\varepsilon.$$ In words: all inputs within $\delta$ of $a$ (except $a$ itself) give outputs within $\varepsilon$ of $L$. On a graph: the curve over the vertical strip $a\pm\delta$ stays inside the horizontal band $L\pm\varepsilon$.
Why do we need it?
It replaces the fuzzy "gets close" with a checkable promise. Every limit law and every theorem about limits is proved from this definition.
Where is it used?
Mostly in proofs. Its spirit appears in practice as tolerances: stop training when the loss is within $\varepsilon$ of its minimum, or compare floats with np.isclose (tolerance), and in convergence guarantees for optimisers.
How is it used?
To prove a limit, take an arbitrary $\varepsilon$, solve $|f(x)-L|<\varepsilon$ for $|x-a|$, and read off a $\delta$ (often a simple multiple of $\varepsilon$). You will rarely need to do this yourself.
Quick check: for $f(x) = 5x$ at $a=0$ with $L=0$, which $\delta$ works for a given $\varepsilon$?
$|f(x)-0| = 5|x| < \varepsilon$ exactly when $|x| < \varepsilon/5$. So $\delta = \varepsilon/5$ works.
Where this leads: the derivative is a limit core
How steep is a curve at one single point? Steepness needs two points (rise over run), but we have only one. The trick is to take a second point very close by and draw the straight line through both, called a secant. Slide the second point closer and closer to the first. The secant line turns into the line that just touches the curve: the tangent. Its steepness is the slope of the curve at that point.
This is the speedometer story again: the average speed over a shorter and shorter time is the limit that gives the speed at one instant.
Take $f(x)=x^2$ at $a=1$. The secant through $(1, 1)$ and $(1+h,\ (1+h)^2)$ has slope $$\frac{f(1+h)-f(1)}{h} = \frac{(1+h)^2-1}{h}.$$
- At $h=0$ this is $\frac00$: stuck, as in every limit problem of this chapter.
- Expand the top: $(1+h)^2 - 1 = 1 + 2h + h^2 - 1 = 2h + h^2$.
- Cancel an $h$ (allowed: $h\ne0$ while $h\to0$): $\dfrac{2h+h^2}{h} = 2 + h$.
- Let $h\to0$: the slope heads to $\mathbf{2}$.
Numbers: $h=1$ gives $3$, $h=0.1$ gives $2.1$, $h=0.01$ gives $2.01$, $h=0.001$ gives $2.001$. The tangent slope at $x=1$ is 2.
The derivative of $f$ at $a$ is the limit of the secant slopes: $$f'(a)=\lim_{h\to0}\frac{f(a+h)-f(a)}{h}.$$ Every tool of this chapter will be used: the $\frac00$ trouble is a removable-hole type of limit, the algebra tricks (expand, factor, cancel) get it out, and identities such as $\frac{\sin h}{h}\to1$ give the derivatives of $\sin$, $e^x$ and $\ln x$. If the limit exists, $f$ is differentiable at $a$ (needing continuity there). We develop all of this in Chapter 2.3.
Why do we need it?
Machine learning asks "if I nudge this weight, how does the loss change?". That is a slope at a point, and a slope at a single point is only meaningful as a limit.
Where is it used?
Gradient descent, backpropagation, sensitivity analysis and every optimiser in the rest of this guide, since the gradient is a list of such limits. Even the finite-difference check of a gradient in code is this formula with a small, fixed $h$.
How is it used?
Write the secant slope $\frac{f(a+h)-f(a)}{h}$, simplify the algebra, and let $h\to0$. In code, use a small $h$ like $10^{-5}$ with the centred version $\frac{f(a+h)-f(a-h)}{2h}$ to check derivatives numerically.
Quick check: for $f(x)=x^2$ at $a=3$, simplify the secant slope $\frac{(3+h)^2-9}{h}$ and take $h\to0$.
$(3+h)^2 - 9 = 9 + 6h + h^2 - 9 = 6h + h^2$. Divide by $h$: $6 + h$. As $h\to0$ this heads to $6$. So the slope of $x^2$ at $x=3$ is $6$ (compare: at $x=1$ it was $2$; the slope is $2a$).
Recap, cheat sheet and practice
- A limit is where $f(x)$ is heading as $x\to a$ (without ever using $x=a$). The value $f(a)$ plays no role.
- One-sided limits use only $x\lt a$ or only $x>a$. The two-sided limit is a number exactly when both one-sided limits exist, are finite and agree.
- For a $\frac00$ puzzle: factor and cancel, or multiply by the conjugate, then substitute.
- Infinite limits mean a vertical asymptote; limits at infinity mean a horizontal asymptote (divide by the highest power; exponential beats polynomial beats log).
- Continuous at $a$: $f(a)$ defined, limit exists, limit $=f(a)$. A discontinuity is removable (hole), a jump, or infinite.
- Key identities (all derived): $\frac{\sin x}{x}\to1$, $(1+\frac1n)^n\to e$, $\frac{e^x-1}{x}\to1$, $\frac{\ln(1+x)}{x}\to1$, $\frac{1-\cos x}{x^2}\to\frac12$.
- The derivative is the limit of secant slopes, $f'(a)=\lim_{h\to0}\frac{f(a+h)-f(a)}{h}$: the road into Chapter 2.3.
Cheat sheet
| Idea | Notation / test | Picture or trick |
|---|---|---|
| Limit | $\lim_{x\to a}f(x)=L$ | output heads to $L$; $f(a)$ ignored |
| One-sided | $\lim_{x\to a^-}$, $\lim_{x\to a^+}$ | approach from the left / from the right |
| Two-sided exists | $L^-=L^+$ (finite) | both walkers meet at the same height |
| $0/0$ trouble | factor, cancel, conjugate | simplify, then substitute |
| Vertical asymptote | $\lim f=\pm\infty$ at $a$ | curve shoots up or down along $x=a$ |
| Horizontal asymptote | $\lim_{x\to\pm\infty}f=L$ | curve levels off at $y=L$ |
| Continuity | $\lim_{x\to a}f(x)=f(a)$ | no pencil lift |
| Removable / jump / infinite | equal limits, wrong value / unequal finite / $\pm\infty$ | hole / step / asymptote |
| Identities | $\frac{\sin x}{x},\ \frac{e^x-1}{x},\ \frac{\ln(1+x)}{x}\to1$; $(1+\frac1n)^n\to e$ | tiny-angle, tiny-growth approximations |
| Derivative preview | $f'(a)=\lim_{h\to0}\frac{f(a+h)-f(a)}{h}$ | secant swings into tangent |
import numpy as np
# 1. Numeric limits: look at f(a + h) for smaller and smaller h
f = lambda x: (x**2 - 1) / (x - 1)
for h in [0.1, 0.01, 0.001, 1e-6]:
print(h, round(f(1 + h), 5), round(f(1 - h), 5))
# 0.1 2.1 1.9
# 0.01 2.01 1.99
# 0.001 2.001 1.999
# 1e-06 2.0 2.0 (both sides head to 2)
# 2. The sin(x)/x limit (sin(0)/0 is 0/0, so we stay close to 0 but not at 0)
x = np.array([0.5, 0.1, 0.01, 0.001])
print(np.sin(x) / x) # [0.95885108 0.99833417 0.99998333 0.99999983] -> 1
# 3. (1 + 1/n)^n heads to e
for n in [1, 10, 100, 10_000, 1_000_000]:
print(n, round((1 + 1 / n) ** n, 6))
# 1 2.0
# 10 2.593742
# 100 2.704814
# 10000 2.718146
# 1000000 2.71828
print(round(np.e, 6)) # 2.718282
# 4. Stable versions of the identities near 0 (avoid rounding loss)
h = 1e-8
print(round(np.expm1(h) / h, 9)) # 1.000000005 (e^h - 1)/h -> 1
print(round(np.log1p(h) / h, 9)) # 0.999999995 ln(1+h)/h -> 1
print((1 - np.cos(h)) / h**2) # 0.0 naive formula: all digits lost!
print(2 * np.sin(h / 2) ** 2 / h**2) # 0.5 stable form of (1 - cos h)/h^2 -> 1/2
# 5. Infinite limits and limits at infinity
print(1 / 0.001, 1 / 0.000001) # 1000.0 1000000.0 (1/x blows up from the right)
big = np.array([10.0, 100.0, 1000.0])
print((3 * big**2 + 1) / (big**2 + 2)) # [2.95098039 2.9995001 2.999995 ] -> 3
# 6. Secant slopes of x^2 at a = 1 approach the slope 2
for h in [1, 0.1, 0.01, 0.001]:
print(h, round(((1 + h) ** 2 - 1) / h, 5))
# 1 3.0
# 0.1 2.1
# 0.01 2.01
# 0.001 2.001
1. What is $\displaystyle\lim_{x\to3}\frac{x^2-9}{x-3}$?
2. At a point $a$, a function has left-hand limit 2 and right-hand limit 5. Which statement is true?
3. What is $\displaystyle\lim_{x\to\infty}\frac{4x^2+x}{2x^2-3}$?
4. Which of these has a removable discontinuity at $x=2$?
5. What is $\displaystyle\lim_{x\to0}\frac{\sin x}{x}$?
6. $f$ is continuous at $a$ when…
Practice problems
A. Find $\displaystyle\lim_{x\to2}\frac{x^2-5x+6}{x-2}$.
Substitution gives $\frac00$. Factor the top: $x^2-5x+6=(x-2)(x-3)$. Cancel $(x-2)$: left with $x-3$. At $x=2$: $2-3=-1$. The limit is $-1$.
B. Find $\displaystyle\lim_{x\to0}\frac{\sqrt{1+x}-1}{x}$.
It is $\frac00$. Multiply top and bottom by the conjugate $\sqrt{1+x}+1$: the top becomes $(1+x)-1 = x$. So the expression is $\dfrac{x}{x(\sqrt{1+x}+1)} = \dfrac{1}{\sqrt{1+x}+1}$. At $x=0$: $\dfrac{1}{1+1}=\dfrac12$. (Numeric check at $x=0.01$: $0.49876$.)
C. Find $\displaystyle\lim_{x\to\infty}\frac{2x^3-x}{5x^3+4}$ and $\displaystyle\lim_{x\to\infty}\frac{x+1}{x^2+1}$.
First: same degree. Divide by $x^3$: $\dfrac{2-1/x^2}{5+4/x^3}\to\dfrac25$. Second: top degree 1 is lower than bottom degree 2. Divide by $x^2$: $\dfrac{1/x+1/x^2}{1+1/x^2}\to\dfrac01=0$.
D. For $f(x)=\dfrac{|x-3|}{x-3}$, find both one-sided limits at $x=3$ and name the discontinuity.
For $x>3$, $|x-3|=x-3$, so $f=1$: right-hand limit $1$. For $x<3$, $|x-3|=-(x-3)$, so $f=-1$: left-hand limit $-1$. They are different finite numbers: a jump of size $1-(-1)=2$. (Also $f(3)$ is not defined, and no choice of $f(3)$ can repair a jump.)
E. Show that $\displaystyle\lim_{x\to0}\frac{\sin 3x}{x}=3$.
Write $\dfrac{\sin3x}{x} = 3\cdot\dfrac{\sin 3x}{3x}$. Let $u=3x$; as $x\to0$, $u\to0$, and $\frac{\sin u}{u}\to1$. So the limit is $3\cdot1=3$. (Numeric check: at $x=0.01$, $\sin(0.03)/0.01 = 2.99955$.)
F. Compute the slope of the secant for $f(x)=x^2$ between $x=3$ and $x=3+h$, and take $h\to0$. Check it with $h = 0.01$.
Slope $=\dfrac{(3+h)^2-9}{h}=\dfrac{6h+h^2}{h}=6+h\to6$. With $h=0.01$ the secant slope is $6.01$, already close to the limit $6$. This limit is the derivative of $x^2$ at $x=3$, the topic of the next chapter.
Differentiation of Univariate Functions
This is the heart of calculus. A derivative answers one simple question: if I nudge the input a tiny bit, how much does the output move? Learn to see it as a slope, learn to compute it from scratch, and learn the handful of rules that let you do it quickly. Then watch it steer a model downhill.
- See the derivative as the slope of the tangent line and as a rate of change (speed, sensitivity)
- Compute a derivative from first principles: secant line to tangent line, by hand, for $x^2$, $x^3$, $1/x$ and $\sqrt{x}$
- Use and derive the power, product, quotient and chain rules, and the derivatives of $e^x$, $a^x$, $\ln x$, $\sin x$, $\cos x$, $\tan x$
- Know the derivatives of the sigmoid and ReLU, and why a flat sigmoid starves learning
- Meet the second and third derivatives ($f''$, $f'''$)
- Connect it all to ML: loss functions, gradient descent, optimisation, sensitivity, learning curves
What is a derivative? core
Think of a road over hills. On a flat road the steepness is zero. Going uphill the steepness is positive, and going downhill it is negative. A sign at each spot could tell you "the road is climbing at 8% here".
A derivative is that sign, for a graph. At every point of a curve it tells you how steep the curve is right there. A straight line has the same steepness everywhere. A curve changes its steepness from point to point, so the derivative is itself a function: feed in $x$, get back the steepness at $x$.
Another way to say it: nudge $x$ a tiny bit to the right. How much does the output move up or down, for each unit of nudge?
Take $f(x) = x^2$. Remember "slope = rise ÷ run". We will nudge $x$ by $0.001$ and measure.
- At $x = 1$: $f(1) = 1$ and $f(1.001) = 1.002001$. The rise is $0.002001$. Divide by the run $0.001$: slope $\approx 2.001$.
- At $x = 2$: $f(2) = 4$ and $f(2.001) = 4.004001$. The rise is $0.004001$, so slope $\approx 4.001$.
- At $x = 0$: $f(0.001) = 0.000001$, so slope $\approx 0.001$, almost flat (the bottom of the bowl).
- At $x = -1$ the curve is going down: slope $\approx -2$.
The pattern is: slope $= 2x$. That rule, "$x$ goes in, $2x$ comes out", is the derivative of $x^2$.
The derivative of $f$ at $x$ is the slope of the curve at $x$:
$$f'(x) = \lim_{h \to 0} \frac{f(x+h) - f(x)}{h}$$Here $h$ is a tiny nudge in $x$ and $f(x+h)-f(x)$ is how much $f$ moves. ("$\lim_{h\to 0}$" means: see what the ratio gets closer and closer to as $h$ shrinks; you met this in Chapter 2.2.) We prove it in the section on first principles below.
Names for the same thing: $f'(x)$ ("f prime of x"), $\dfrac{df}{dx}$ ("d f d x"), $\dfrac{d}{dx}f(x)$, or $y'$ if $y = f(x)$. A function with a derivative at a point is called differentiable there.
- $f'(x) > 0$: the curve is rising at $x$.
- $f'(x) < 0$: the curve is falling.
- $f'(x) = 0$: the curve is flat (a hilltop, a valley bottom, or a shelf).
Why do we need it?
We need a number that says "which way is up, and how steeply" at one exact point. Without it we could only compare two far-apart points, never describe the curve right where we stand.
Where is it used?
Training every model: the derivative of the loss tells gradient descent which way to move each weight. Also physics (speed), economics (marginal cost) and sensitivity analysis.
How is it used?
Compute $f'(x)$ at your current point. Its sign says "which way is uphill" and its size says "how steep". To go downhill, step against it.
Not every curve has a derivative everywhere. A sharp corner (like the tip of $|x|$ or the kink of ReLU) has no single tangent. A jump has none either. We will meet ReLU's kink again below.
The derivative is a function, and also a number. $f'(x)$ is a whole new function of $x$. $f'(2)$ is one number: the slope at $x=2$.
Quick check: for $f(x)=x^2$, what is the slope at $x=3$, and is the curve rising or falling there?
Using the pattern $f'(x)=2x$: slope $=6$. It is positive, so the curve is rising.
Geometric interpretation: zoom in until it is straight core
Stand in a field. The Earth is round, but the ground looks flat. Zoom in far enough on any smooth curve and the same thing happens: the curve looks like a straight line.
That straight line is the tangent line. It is the best straight-line copy of the curve near one point. Its slope is the derivative. This one picture is the idea behind almost everything in this chapter, and behind linearization later.
Take $f(x)=x^2$ at the point $a=1$, where $f(1)=1$ and the slope is $f'(1)=2$.
- The line through $(1, 1)$ with slope $2$ is $y = 1 + 2(x-1) = 2x - 1$.
- At $x = 1.1$: the curve gives $1.1^2 = 1.21$. The line gives $2(1.1)-1 = 1.2$. Gap: $0.01$.
- At $x = 1.01$: the curve gives $1.0201$, the line gives $1.02$. Gap: $0.0001$.
Come ten times closer and the gap becomes a hundred times smaller. That is "looks straight when you zoom in".
The tangent line to $y=f(x)$ at $x=a$ is the line through the point $(a, f(a))$ with slope $f'(a)$:
$$y = f(a) + f'(a)\,(x - a).$$Near $a$ the curve and the tangent line almost agree: $f(x) \approx f(a) + f'(a)(x-a)$. This is the linear approximation.
Why do we need it?
Curves are hard to work with and straight lines are easy. If a curve looks like a line close up, we can answer "what happens if I move a little?" with simple arithmetic.
Where is it used?
Gradient descent (it assumes the loss is locally a slope), Newton's method, error estimates, and the Taylor series and linearization chapters ahead.
How is it used?
Compute $f(a)$ and $f'(a)$ once. Then predict nearby values with $f(a)+f'(a)(x-a)$ instead of evaluating $f$ again.
Quick check: the tangent line to $f(x)=x^2$ at $a=3$?
$f(3)=9$ and $f'(3)=6$, so $y = 9 + 6(x-3) = 6x - 9$.
Rate of change: speed versus distance core
A car has two dials. The odometer says how far you have gone. The speedometer says how fast you are going right now. Speed is the rate of change of distance: how quickly distance is growing.
You can also compute a speed over a whole trip: "120 km in 2 hours is 60 km/h". That is the average rate. The speedometer shows the instantaneous rate: the average over a trip so short that it is just one moment. That instantaneous rate is the derivative.
The same idea works for any pair of quantities: dollars per item, loss per training step, temperature per hour.
A car's distance is $s(t) = t^2$ metres after $t$ seconds.
- Average over $t=1$ to $t=3$: $s(3)-s(1) = 9-1 = 8$ metres in $2$ seconds: $8/2 = 4$ m/s.
- Squeeze the trip around $t = 2$: from $2$ to $2.1$: $(4.41-4)/0.1 = 4.1$ m/s. From $2$ to $2.01$: $(4.0401 - 4)/0.01 = 4.01$ m/s.
- The averages approach $4$. So the speed at the moment $t=2$ is $4$ m/s, which matches $s'(t)=2t = 4$.
Units: the derivative has units "(output units) per (input unit)": metres per second here.
The average rate of change of $f$ from $x$ to $x+h$ is $\dfrac{f(x+h)-f(x)}{h}$ (rise over run: the slope of a secant line, which is a straight line through two points of the curve).
The instantaneous rate of change at $x$ is its limit as $h\to 0$, which is $f'(x) = \dfrac{df}{dx}$. Read $\dfrac{df}{dx}$ as "a tiny change in $f$ divided by the tiny change in $x$ that caused it".
Why do we need it?
Averages hide what happens at one moment. A trip can average 60 km/h while sometimes standing still. To describe "now", we need the rate over an infinitely short time.
Where is it used?
Physics (velocity, acceleration), finance (marginal cost), epidemic growth rates, and ML: the rate at which the loss falls per training step or per unit change of a weight.
How is it used?
Ask "output change per unit input change". Write it as $df/dx$ and keep the units: they are a quick check that you built the right ratio.
Quick check: a model's loss falls from 5.0 to 3.8 over 6 training steps. What is the average rate of change, with units?
$(3.8-5.0)/6 = -0.2$ loss units per step. Negative means the loss is falling.
Derivative from first principles: secant to tangent core
We want the slope at one point. But slope needs two points (rise over run). The trick: pick a second point a small distance $h$ away and draw the straight line through both. That line is a secant. Its slope is easy: rise ÷ run.
Now slide the second point closer and closer to the first. The secant line swings and settles into the tangent. The secant's slope settles down to one number. That number is the derivative.
Why not just set $h=0$? Then rise and run are both $0$, and $0/0$ means nothing. So we do algebra first, to cancel the $h$, and then let $h$ go to $0$. That is exactly what a limit is for.
Worked example: $f(x)=x^2$. Follow every step.
- Write the secant slope: $\dfrac{f(x+h)-f(x)}{h} = \dfrac{(x+h)^2 - x^2}{h}$.
- Expand the square: $(x+h)^2 = x^2 + 2xh + h^2$.
- Subtract $x^2$: the top becomes $2xh + h^2$.
- Both terms on top contain $h$, so factor it out and cancel: $\dfrac{2xh + h^2}{h} = \dfrac{h(2x + h)}{h} = 2x + h$.
- Now let $h \to 0$. The leftover $h$ vanishes: the slope is $2x$.
So $\dfrac{d}{dx}x^2 = 2x$, which matches the numbers in the first section. See the trend with real numbers at $x=1$: with $h = 1$ the secant slope is $3$; with $h = 0.1$ it is $2.1$; with $h=0.01$ it is $2.01$. Each one equals $2x+h = 2+h$.
The derivative is the limit of the secant slope as the gap shrinks to zero:
$$f'(x) = \lim_{h\to 0} \frac{f(x+h)-f(x)}{h}.$$The recipe, every time: (1) write the ratio, (2) expand and simplify until the $h$ on the bottom cancels, (3) let $h \to 0$. The more careful name for this is the difference quotient.
Why do we need it?
It is the definition: everything else (every rule) is a shortcut that was proven from it. If you can do it from first principles, you never have to trust a formula blindly.
Where is it used?
To prove the derivative rules, and in practice as the "nudge test" (finite differences) that checks gradients in a neural network: nudge a weight by $h$, see how the loss moves.
How is it used?
Write $\frac{f(x+h)-f(x)}{h}$, simplify algebraically until $h$ cancels, then set $h=0$. In code: use a small $h$ such as $10^{-5}$ and compare with your formula.
Worked example: $f(x)=x^3$.
- Ratio: $\dfrac{(x+h)^3 - x^3}{h}$.
- Expand the cube: $(x+h)^3 = x^3 + 3x^2h + 3xh^2 + h^3$.
- Subtract $x^3$: the top is $3x^2h + 3xh^2 + h^3$.
- Divide every term by $h$: $3x^2 + 3xh + h^2$.
- Let $h\to 0$: the terms with $h$ vanish, leaving $3x^2$.
So $\dfrac{d}{dx}x^3 = 3x^2$. Check at $x=1$: the secant slopes are $3+3h+h^2$, which is $3.31$ for $h=0.1$ and gets closer to $3$.
Worked example: $f(x)=1/x$.
- Ratio: $\dfrac{\frac{1}{x+h} - \frac{1}{x}}{h}$.
- Put the top over a common denominator: $\dfrac{1}{x+h} - \dfrac1x = \dfrac{x - (x+h)}{x(x+h)} = \dfrac{-h}{x(x+h)}$.
- Divide by $h$: $\dfrac{-h}{x(x+h)}\cdot\dfrac1h = \dfrac{-1}{x(x+h)}$.
- Let $h\to0$: $\dfrac{-1}{x\cdot x} = -\dfrac{1}{x^2}$.
So $\dfrac{d}{dx}\dfrac1x = -\dfrac1{x^2}$. At $x=2$ the slope is $-\tfrac14$. Check with $h=0.1$: $(1/2.1 - 0.5)/0.1 \approx -0.238$, close to $-0.25$. It is always negative: $1/x$ always falls as $x$ grows (for $x>0$).
Worked example: $f(x)=\sqrt{x}$. A square root on top calls for a trick: multiply top and bottom by the conjugate $\sqrt{x+h}+\sqrt{x}$. (This uses $(A-B)(A+B) = A^2-B^2$.)
- Ratio: $\dfrac{\sqrt{x+h}-\sqrt{x}}{h}$.
- Multiply top and bottom by $\sqrt{x+h}+\sqrt{x}$. The top becomes $(x+h) - x = h$.
- So the ratio is $\dfrac{h}{h\,(\sqrt{x+h}+\sqrt{x})}$. Cancel the $h$: $\dfrac{1}{\sqrt{x+h}+\sqrt{x}}$.
- Let $h\to0$: $\dfrac{1}{\sqrt{x}+\sqrt{x}} = \dfrac{1}{2\sqrt{x}}$.
So $\dfrac{d}{dx}\sqrt{x} = \dfrac{1}{2\sqrt x}$. At $x=4$ the slope is $\tfrac14$. Check with $h=0.1$: $(\sqrt{4.1}-2)/0.1 \approx 0.2485$. The slope is huge near $x=0$ (the curve starts out vertical) and flattens as $x$ grows.
You cannot skip the algebra. Plugging $h=0$ into $\frac{f(x+h)-f(x)}{h}$ gives $\frac00$. Simplify first, then let $h \to 0$.
Computers use a small, not zero, $h$. In code, $h\approx10^{-5}$ works well. Much smaller (like $10^{-15}$) gets ruined by rounding errors. Using the symmetric version $\frac{f(x+h)-f(x-h)}{2h}$ (a "central difference") is more accurate, and is what this guide's "nudge checks" use.
Quick check: use first principles to find the derivative of $f(x)=3x+5$.
$\dfrac{(3(x+h)+5) - (3x+5)}{h} = \dfrac{3h}{h} = 3$. It is $3$ for every $h$, and so $f'(x)=3$: a straight line's slope never changes.
Warm-up rules: constants, multiples and sums
Three tiny facts make big formulas easy.
- A constant never changes, so its slope is $0$. Lifting a whole graph up by 5 does not change how steep it is anywhere.
- Stretching a graph taller by a factor $c$ stretches every slope by $c$.
- Adding two graphs adds their heights, and so adds their slopes.
Together: you can differentiate a long expression one piece at a time.
Differentiate $f(x) = 3x^2 + 5x - 7$. We know $\frac{d}{dx}x^2 = 2x$, $\frac{d}{dx}x = 1$ and $\frac{d}{dx}7 = 0$.
- The $3x^2$ piece: $3 \cdot 2x = 6x$.
- The $5x$ piece: $5\cdot 1 = 5$.
- The $-7$ piece: $0$.
- Add: $f'(x) = 6x + 5$.
Check at $x=2$: $f'(2)=17$. Nudge: $f(2)=15$, $f(2.001)=15.017003$, slope $\approx 17.003$ ✓.
For a constant $c$ and differentiable $f$, $g$:
$$\frac{d}{dx}c = 0, \qquad \frac{d}{dx}\big[c\,f(x)\big] = c\,f'(x), \qquad \frac{d}{dx}\big[f(x)\pm g(x)\big] = f'(x) \pm g'(x).$$Why the sum rule is true (from first principles): the secant slope of $f+g$ is
$$\frac{[f(x+h)+g(x+h)] - [f(x)+g(x)]}{h} = \frac{f(x+h)-f(x)}{h} + \frac{g(x+h)-g(x)}{h},$$and as $h\to0$ each piece goes to its own derivative. The constant-multiple rule works the same way: $c$ comes straight out of the ratio. The word for "these two rules together" is linearity of the derivative.
Why do we need it?
Real functions are sums of simple pieces. Linearity lets us differentiate a messy sum by handling each term alone, and drop any constant.
Where is it used?
Every loss that is a sum over data points (the derivative of a sum of losses is the sum of derivatives), regularisation terms like $\lambda w^2$ added to a loss, and polynomial models.
How is it used?
Split the expression at the plus and minus signs. Differentiate each term. Pull out constant factors. Terms with no $x$ in them vanish.
Quick check: differentiate $f(x) = 4x^2 - 3x + 9$.
$4\cdot2x - 3\cdot 1 + 0 = 8x - 3$.
The power rule core
Look at what first principles gave us: $x^2 \to 2x$ and $x^3 \to 3x^2$. The recipe is a pattern:
Bring the power down in front, then lower the power by one.
So $x^4 \to 4x^3$, $x^{10}\to 10x^9$. The pattern even works for the strange powers we did by hand: $1/x = x^{-1}$ gave $-x^{-2}$, and $\sqrt x = x^{1/2}$ gave $\tfrac12 x^{-1/2}$.
- $\dfrac{d}{dx}x^5 = 5x^4$.
- $\dfrac{d}{dx}x = \dfrac{d}{dx}x^1 = 1\cdot x^0 = 1$.
- $\dfrac{d}{dx}\dfrac{1}{x^2} = \dfrac{d}{dx}x^{-2} = -2x^{-3} = -\dfrac{2}{x^3}$.
- $\dfrac{d}{dx}\sqrt[3]{x} = \dfrac{d}{dx}x^{1/3} = \tfrac13 x^{-2/3}$.
- $\dfrac{d}{dx}\big(4x^3 - 2x^2 + 5x - 9\big) = 12x^2 - 4x + 5$. At $x=2$ this is $48 - 8 + 5 = 45$.
For any real number $n$:
$$\frac{d}{dx}x^n = n\,x^{n-1}.$$Where it comes from. Expand $(x+h)^n$ with the binomial pattern. For small powers:
$$(x+h)^2 = x^2 + 2xh + h^2,\quad (x+h)^3 = x^3 + 3x^2h + 3xh^2 + h^3,\quad (x+h)^4 = x^4 + 4x^3h + 6x^2h^2 + 4xh^3 + h^4.$$In every case the expansion is $x^n + n\,x^{n-1}h + (\text{terms with } h^2 \text{ or higher})$. Subtract $x^n$ and divide by $h$: you get $n\,x^{n-1} + (\text{terms that still contain } h)$. Let $h\to0$ and only $n\,x^{n-1}$ is left. For every whole number $n$ the binomial pattern always gives this, so the rule is proved for whole numbers. For $n=-1$ and $n=\tfrac12$ we proved it by hand above. For other fractions and negative powers it also holds, and we accept that without a full proof here.
Why do we need it?
Polynomials and power laws are everywhere, and differentiating $x^n$ from first principles every time would be slow. One line replaces a page of algebra.
Where is it used?
Squared-error loss $(y-\hat y)^2$, weight decay $w^2$, polynomial regression, scaling laws (loss $\propto$ size$^{-0.05}$), and the Taylor series later on.
How is it used?
Rewrite roots and fractions as powers ($\sqrt x = x^{1/2}$, $1/x^3 = x^{-3}$), multiply by the power, subtract one from the power.
The power rule is for "variable to a fixed power". $x^3$ is a power function. But $3^x$ (fixed base, variable exponent) is an exponential and follows a different rule (below). Do not write $\frac{d}{dx}3^x = x\,3^{x-1}$.
Quick check: differentiate $f(x)=\dfrac{5}{x^3} + 2\sqrt{x}$.
Rewrite: $5x^{-3} + 2x^{1/2}$. Then $5(-3)x^{-4} + 2\cdot\tfrac12 x^{-1/2} = -\dfrac{15}{x^4} + \dfrac{1}{\sqrt x}$.
The product rule core
Picture a rectangle whose width is $u(x)$ and whose height is $v(x)$. Its area is $u \cdot v$. Now nudge $x$ a little. Both sides grow a little. How does the area change?
- A thin strip is added along the right edge: its length is the height $v$, its thickness is how much the width grew.
- A thin strip is added along the top edge: its length is the width $u$, its thickness is how much the height grew.
- A tiny square appears in the corner. It is (tiny) × (tiny), so it is far smaller than the strips and disappears when the nudge shrinks.
So the area's growth is (how fast the width grows) × (height) plus (width) × (how fast the height grows).
Let $f(x) = x^2(3x+1)$, so $u = x^2$ and $v = 3x+1$.
- Derivatives of the parts: $u' = 2x$, $v' = 3$.
- Product rule: $f' = u'v + uv' = 2x(3x+1) + x^2\cdot 3 = 6x^2 + 2x + 3x^2 = 9x^2 + 2x$.
- Check by multiplying out first: $f = 3x^3 + x^2$, so $f' = 9x^2 + 2x$ ✓ (same answer).
- At $x = 2$: $f'(2) = 36 + 4 = 40$. Nudge check: $f(2) = 28$, $f(2.001) = 28.040019$, slope $\approx 40.02$ ✓.
Not every product can be multiplied out (try $x\,e^x$ or $x^2\sin x$). That is where this rule earns its keep.
If $f(x) = u(x)\,v(x)$, then
$$f'(x) = u'(x)\,v(x) + u(x)\,v'(x).$$Derivation from the limit. Write the secant slope and do one clever step: add and subtract $u(x+h)\,v(x)$ on top.
$$\begin{aligned} \frac{u(x+h)v(x+h) - u(x)v(x)}{h} &= \frac{u(x+h)v(x+h) - u(x+h)v(x) + u(x+h)v(x) - u(x)v(x)}{h} \\[4pt] &= u(x+h)\,\frac{v(x+h)-v(x)}{h} + v(x)\,\frac{u(x+h)-u(x)}{h}. \end{aligned}$$Now let $h\to0$. The fraction $\frac{v(x+h)-v(x)}{h}\to v'(x)$ and $\frac{u(x+h)-u(x)}{h}\to u'(x)$. Also $u(x+h)\to u(x)$ (a smooth curve does not jump). This leaves $u(x)v'(x) + v(x)u'(x)$. ∎
In words: "derivative of the first times the second, plus the first times the derivative of the second."
Why do we need it?
Many functions multiply two changing things, and the answer is not "multiply the two derivatives". We need the correct recipe for how a product moves when both factors move.
Where is it used?
Gating in LSTMs and attention (one signal times another), the loss term $x\ln x$ in entropy, "weight times activation" terms in backprop, and physics problems where mass and speed both change (momentum $m\,v$).
How is it used?
Name the two factors $u$ and $v$. Differentiate each separately. Combine as $u'v + uv'$. Simplify at the end.
The derivative of a product is NOT the product of the derivatives. Test with $x\cdot x = x^2$: the true derivative is $2x$, but $u'v' = 1\cdot1 = 1$. The extra pieces ($u'v$ and $uv'$) are the two strips in the picture.
Quick check: differentiate $f(x)=x\,e^x$ at $x=1$.
$u=x$, $v=e^x$. $f' = 1\cdot e^x + x\,e^x = (1+x)e^x$. At $x=1$: $2e \approx 5.437$.
The quotient rule core
A fraction $\dfrac{u}{v}$ changes in two ways. If the top $u$ grows, the fraction grows. If the bottom $v$ grows, the fraction shrinks (a bigger bucket divides the same water into smaller shares). So the rule has a plus and a minus, and a squared bottom to account for dividing.
We do not need new limit work. A fraction is just a product in disguise, so the product rule will give us the answer.
Let $f(x) = \dfrac{x}{x^2+1}$, so $u = x$ and $v = x^2+1$.
- $u' = 1$, $v' = 2x$.
- Top of the rule: $u'v - uv' = 1\cdot(x^2+1) - x\cdot 2x = x^2 + 1 - 2x^2 = 1 - x^2$.
- Bottom of the rule: $v^2 = (x^2+1)^2$.
- So $f'(x) = \dfrac{1 - x^2}{(x^2+1)^2}$.
At $x=1$: $f'(1)=0$, so the graph is flat there (its peak: $f(1)=\tfrac12$). At $x=2$: $f'(2) = \dfrac{1-4}{25} = -0.12$. Nudge check: $f(2) = 0.4$, $f(2.001)\approx 0.39988$, slope $\approx-0.12$ ✓.
Derivation from the product rule. Call the fraction $q = u/v$. Then $u = q\,v$. Differentiate both sides with the product rule:
$$u' = q'\,v + q\,v'.$$Solve for $q'$: $\;q' = \dfrac{u' - q\,v'}{v}$. Now put back $q = u/v$:
$$q' = \frac{u' - \frac{u}{v}v'}{v} = \frac{\frac{u'v - uv'}{v}}{v} = \frac{u'v - uv'}{v^2}. \;∎$$Memory line: "bottom × derivative of top, minus top × derivative of bottom, all over bottom squared". A special case worth knowing: $\dfrac{d}{dx}\dfrac1v = -\dfrac{v'}{v^2}$ (take $u=1$, so $u'=0$), which matches our earlier $\frac{d}{dx}\frac1x = -\frac1{x^2}$.
Why do we need it?
Ratios appear whenever one quantity is divided by another, and dividing by something that also changes needs its own correction (the minus term).
Where is it used?
Sigmoid $1/(1+e^{-x})$, softmax (an exponential divided by a sum), normalising a vector by its length, $\tan x = \sin x/\cos x$, and rates such as "loss per sample".
How is it used?
Name top $u$ and bottom $v$. Compute $u'$ and $v'$. Form $u'v - uv'$ (order matters!) and divide by $v^2$. Often it is simplest to leave the bottom squared.
The order in the top is not optional. It is $u'v - uv'$, not $uv' - u'v$. A swap flips the sign of your answer. A quick sanity check: for $1/x$, the slope should be negative, and the rule with $u=1$ gives $-v'/v^2 = -1/x^2$ ✓.
Often you can skip the rule. $\dfrac{x^3+x}{x} = x^2+1$ is easier by simplifying first.
Quick check: differentiate $f(x) = \dfrac{x+1}{x-1}$.
$u=x+1$, $v=x-1$, $u'=v'=1$. $f' = \dfrac{1\cdot(x-1) - (x+1)\cdot1}{(x-1)^2} = \dfrac{-2}{(x-1)^2}$. It is negative everywhere (the function falls on each side of $x=1$).
The chain rule core
Imagine two machines in a row. Machine $g$ takes $x$ and makes $u$. Machine $f$ takes $u$ and makes $y$. So $x \to u \to y$. This is a composition, written $y = f(g(x))$ (you met it in Chapter 2.1).
Think of three gears in a line. Turn the first gear (the input $x$). The second gear turns, say, 3 times as fast. The third gear turns 2 times as fast as the second. So the third gear turns $3 \times 2 = 6$ times as fast as the first. Rates multiply along the chain.
That is the whole chain rule: "how fast does $y$ change when $x$ changes?" equals "how fast $u$ changes with $x$" times "how fast $y$ changes with $u$".
Let $y = (3x+1)^2$. Break it into two machines: inner $u = 3x+1$, outer $y = u^2$.
- Inner rate: $\dfrac{du}{dx} = 3$.
- Outer rate: $\dfrac{dy}{du} = 2u$.
- Multiply: $\dfrac{dy}{dx} = 2u \cdot 3 = 6u$.
- Put $u$ back: $\dfrac{dy}{dx} = 6(3x+1)$.
At $x=1$: $u = 4$, so the slope is $6\cdot4 = 24$. Nudge check: $y(1) = 16$, $y(1.001) = 16.024009$, so the slope $\approx 24.009$ ✓.
A harder one: $y = \sin(x^2)$. Inner $u = x^2$ (rate $2x$), outer $\sin u$ (rate $\cos u$). So $\dfrac{dy}{dx} = \cos(x^2)\cdot 2x$. At $x = 1$ this is $2\cos 1 \approx 1.0806$.
Three links: $y = e^{\sin(x^2)}$. Let $u = x^2$, $v = \sin u$, $y = e^v$. Then $\dfrac{dy}{dx} = \dfrac{dy}{dv}\cdot\dfrac{dv}{du}\cdot\dfrac{du}{dx} = e^v\cdot\cos u\cdot 2x$. At $x=1$: $e^{\sin 1}\cos 1 \cdot 2 \approx 2.507$. Just keep multiplying rates, one per link.
If $y = f(u)$ and $u = g(x)$, then
$$\frac{dy}{dx} = \frac{dy}{du}\cdot\frac{du}{dx}, \qquad\text{or}\qquad \big(f(g(x))\big)' = f'(g(x))\cdot g'(x).$$Why it is true. Nudge $x$ by $\Delta x$. That nudges $u$ by $\Delta u$, which nudges $y$ by $\Delta y$. Always
$$\frac{\Delta y}{\Delta x} = \frac{\Delta y}{\Delta u}\cdot\frac{\Delta u}{\Delta x}$$(the $\Delta u$ on top and bottom cancel, like ordinary fractions). Now let $\Delta x \to 0$. Then $\Delta u\to0$ too, and the three ratios become $\frac{dy}{dx}$, $\frac{dy}{du}$ and $\frac{du}{dx}$. ∎ (When $g$ happens to leave $u$ unchanged, $\Delta u = 0$, a more careful proof is needed, but the result is the same.)
Recipe: (1) name the inner function $u$; (2) differentiate the outer function with respect to $u$; (3) differentiate the inner function; (4) multiply; (5) replace $u$. Chains can be longer: just multiply one rate per link.
This is the single-variable chain rule. When there are many inputs, it becomes a sum of products over paths (Chapter 2.8), and applying it layer after layer is backpropagation (Chapter 2.9).
Why do we need it?
Almost every function in ML is built by feeding one function into another. We need a way to get the slope of the whole machine from the slopes of its parts.
Where is it used?
Backpropagation (a neural network is a long chain of layers), the derivative of $(y-\hat y)^2$ with respect to a weight, the sigmoid's derivative, $\ln$ of a probability, and every automatic-differentiation library.
How is it used?
Split the formula into layers from the inside out. Write the slope of each layer. Multiply all the slopes together. Each layer only needs to know its own local slope.
Do not forget the inner derivative. The most common mistake: $\frac{d}{dx}\sin(x^2) = \cos(x^2)$. It is missing the factor $2x$. The outer rule alone only handles the outer machine.
Product rule or chain rule? If two pieces are multiplied, use the product rule. If one piece is inside the other, use the chain rule. Many problems need both.
Quick check: differentiate $y = (x^2+1)^3$.
Inner $u = x^2+1$, $u' = 2x$. Outer $u^3$, derivative $3u^2$. So $y' = 3(x^2+1)^2\cdot 2x = 6x(x^2+1)^2$. At $x=1$: $6\cdot4 = 24$.
Exponential derivatives: why $e^x$ is special core
Money in a bank with interest, bacteria in a dish, a rumour on social media: in all of them the more you have, the faster it grows. The growth rate is proportional to the current size. That is the nature of an exponential function.
So the slope of $a^x$ should be a fixed multiple of its height. For some special base the multiple is exactly 1: the slope equals the height at every point. That base is the number $e \approx 2.71828$. The curve $e^x$ is its own derivative.
How steep is $a^x$ at $x=0$? Use the nudge $\dfrac{a^h - 1}{h}$ (because $a^0=1$):
| base $a$ | $h=0.1$ | $h=0.01$ | $h=0.001$ | settles near |
|---|---|---|---|---|
| $2$ | 0.7177 | 0.6956 | 0.6934 | 0.6931 |
| $e$ | 1.0517 | 1.0050 | 1.0005 | 1 |
| $3$ | 1.1612 | 1.1047 | 1.0992 | 1.0986 |
Base 2 is a little too gentle (slope 0.693), base 3 a little too steep (1.099). Somewhere between them, at $e$, the slope at $0$ is exactly $1$.
Derivation. Start from first principles for $a^x$:
$$\frac{a^{x+h} - a^x}{h} = \frac{a^x\,a^h - a^x}{h} = a^x\cdot\frac{a^h - 1}{h}.$$The fraction $\frac{a^h-1}{h}$ does not depend on $x$. As $h\to0$ it goes to a constant $k(a)$ (the slope of $a^x$ at $x=0$, the numbers in the table). So $\dfrac{d}{dx}a^x = k(a)\,a^x$: the slope is always a fixed multiple of the height. The number $e$ is defined as the base for which $k(e)=1$. Hence
$$\frac{d}{dx}e^x = e^x.$$Other bases. Write $a = e^{\ln a}$, so $a^x = e^{x\ln a}$. The chain rule (inner $u = x\ln a$, with rate $\ln a$) gives
$$\frac{d}{dx}a^x = (\ln a)\,a^x, \qquad\text{so } k(a) = \ln a \;(\text{matches: } \ln 2 = 0.693,\ \ln 3 = 1.099).$$With a chain: $\dfrac{d}{dx}e^{g(x)} = g'(x)\,e^{g(x)}$. For example $\dfrac{d}{dx}e^{2x} = 2e^{2x}$, $\dfrac{d}{dx}e^{-x} = -e^{-x}$, $\dfrac{d}{dx}e^{-x^2} = -2x\,e^{-x^2}$.
Why do we need it?
Growth and decay, probabilities and softmax all use exponentials. The special fact "$e^x$ is its own slope" makes their derivatives clean instead of messy.
Where is it used?
Softmax and sigmoid (both are built from $e^x$), exponential learning-rate decay, the Gaussian bell curve $e^{-x^2/2}$, radioactive decay and compound interest.
How is it used?
For $e^{g(x)}$: copy the function and multiply by $g'(x)$. For $a^x$: copy it and multiply by $\ln a$.
$e^x$ is not a power function. The variable is in the exponent. Never write $\frac{d}{dx}e^x = x\,e^{x-1}$. And $\frac{d}{dx}a^x$ is not just $a^x$ unless $a=e$: it has the extra factor $\ln a$.
Quick check: differentiate $f(x) = 5e^{3x}$, and find the slope of $2^x$ at $x=3$.
$f' = 5\cdot3e^{3x} = 15e^{3x}$. For $2^x$: $(\ln 2)\,2^3 = 8\ln2 \approx 5.545$.
Logarithmic derivatives: $\ln x$ and friends core
$\ln x$ answers: "to what power must I raise $e$ to get $x$?" It undoes $e^x$: it is the inverse function (Chapter 2.1). The graph of an inverse is the original graph flipped across the diagonal line $y=x$.
Flip a hill and its steepness flips too: "rise over run" becomes "run over rise". So the slope of the inverse is one over the slope of the original. Where $e^x$ is very steep, $\ln x$ is very flat, and the other way around.
That already hints at the answer: $\ln x$ is steep for small $x$ and gets flatter and flatter. It grows, but slowly.
- At $x = 2$: the slope of $\ln x$ is $\tfrac12$. Nudge check: $(\ln 2.001 - \ln 2)/0.001 \approx 0.49988$ ✓.
- At $x = 0.1$: slope $= 10$ (very steep, just after the curve starts).
- At $x = 10$: slope $= 0.1$ (nearly flat).
- Picture check: the point $(1, e)$ on $e^x$ has slope $e$. Its mirror image $(e, 1)$ on $\ln x$ has slope $1/e$, which is $1/x$ at $x=e$ ✓.
Derivation. Let $y = \ln x$. By definition $e^y = x$. Differentiate both sides with respect to $x$. On the left use the chain rule ($y$ depends on $x$); on the right, $\frac{d}{dx}x = 1$:
$$e^y\cdot\frac{dy}{dx} = 1 \;\Longrightarrow\; \frac{dy}{dx} = \frac{1}{e^y} = \frac1x.\;∎$$Other bases. $\log_a x = \dfrac{\ln x}{\ln a}$ (change of base), and $\ln a$ is just a constant, so $\dfrac{d}{dx}\log_a x = \dfrac{1}{x\ln a}$. For example the slope of $\log_2 x$ at $x=4$ is $\dfrac{1}{4\ln2}\approx0.3607$.
With a chain: $\dfrac{d}{dx}\ln g(x) = \dfrac{g'(x)}{g(x)}$. For example $\dfrac{d}{dx}\ln(x^2+1) = \dfrac{2x}{x^2+1}$, which is $1$ at $x=1$. The quantity $g'/g$ is the relative rate of change (change as a fraction of the current size). It is one reason the log of a likelihood, a product of many probabilities, is so convenient to differentiate (you will meet this in Chapter 2.14).
Why do we need it?
Logs turn products into sums and tame huge or tiny numbers, so ML uses them constantly. We need their slopes to train anything that contains a log.
Where is it used?
Cross-entropy loss $-\ln p$ and log-likelihood (the loss $-\ln p$ has slope $-1/p$: huge when $p$ is tiny, which punishes confident mistakes), softmax, entropy, and log-scaled features.
How is it used?
For $\ln g(x)$: put $g'(x)$ on top of $g(x)$. For $\log_a$: divide by $\ln a$ too. Remember $g(x)$ must be positive.
Only positive numbers have a log. $\ln x$ and its slope $1/x$ exist only for $x>0$. And $\frac{d}{dx}\ln x = \frac1x$ is not $\ln$ of anything: it is a plain fraction.
Quick check: differentiate $f(x) = \ln(5x)$. What do you notice?
Chain rule: $\dfrac{5}{5x} = \dfrac1x$. The same as $\ln x$. (Indeed $\ln 5x = \ln 5 + \ln x$, and the constant $\ln 5$ has slope $0$.)
Trigonometric derivatives: sin, cos, tan core
Walk anticlockwise round a circle of radius 1 at speed 1. Your height is $\sin x$ and your sideways position is $\cos x$, where $x$ is how far round you have gone (an angle in radians).
At each moment you move along the tangent, at speed 1. How fast is your height changing? Only the upward part of your motion counts. At the right-hand side of the circle you are moving straight up (height changes fastest), and at the very top you are moving sideways (height momentarily not changing).
That upward part of the motion is exactly $\cos x$. So the slope of $\sin x$ is $\cos x$. By the same reasoning your sideways position changes at rate $-\sin x$: the slope of $\cos x$ is $-\sin x$ (negative while you are in the top half, moving left).
- At $x = 0$: $\sin$ rises steeply, slope $\cos 0 = 1$.
- At $x = \pi/2$ (the top of the wave): slope $\cos(\pi/2) = 0$. Flat.
- At $x = \pi$: slope $\cos\pi = -1$, falling steeply through zero.
- Small-angle check: $\sin(0.001)/0.001 = 0.9999998 \approx 1$ ✓.
The slope graph of $\sin$ is the cosine wave, shifted a quarter-turn. And the slope graph of cosine is the sine wave flipped upside down.
Derivation of $\sin'$. Use the angle-addition formula $\sin(x+h) = \sin x\cos h + \cos x\sin h$:
$$\frac{\sin(x+h)-\sin x}{h} = \sin x\cdot\frac{\cos h - 1}{h} + \cos x\cdot\frac{\sin h}{h}.$$Two limits do all the work. From Chapter 2.2 we know $\frac{\sin h}{h}\to1$ and $\frac{1-\cos h}{h^2}\to\frac12$. The second one gives $\frac{\cos h-1}{h} = -h\cdot\frac{1-\cos h}{h^2}\to 0\cdot\frac12 = 0$. Numerically, the two ratios in the formula above: at $h=0.1$ they are $0.99833$ and $-0.04996$; at $h=0.01$ they are $0.99998$ and $-0.00500$; at $h=0.001$ they are $0.9999998$ and $-0.0005$. So the sum goes to $\sin x\cdot 0 + \cos x\cdot 1 = \cos x$. ∎
Derivation of $\cos'$. Note $\cos x = \sin(\tfrac\pi2 - x)$. Chain rule: outer $\sin$ gives $\cos(\tfrac\pi2 - x)$, inner $\tfrac\pi2 - x$ has slope $-1$. So $(\cos x)' = -\cos(\tfrac\pi2 - x) = -\sin x$. ∎
Derivation of $\tan'$. $\tan x = \dfrac{\sin x}{\cos x}$. Quotient rule with $u=\sin x$, $v=\cos x$:
$$\frac{\cos x\cdot\cos x - \sin x\cdot(-\sin x)}{\cos^2 x} = \frac{\cos^2x + \sin^2x}{\cos^2x} = \frac{1}{\cos^2 x}.\;∎$$(using $\cos^2x+\sin^2x = 1$). At $x=\pi/4$ this is $1/(\tfrac{1}{\sqrt2})^2 = 2$.
Why do we need it?
Anything that repeats (waves, rotations, seasons) is described by sin and cos, and we want to know how fast those quantities change.
Where is it used?
Positional encodings in Transformers (sines and cosines of different frequencies), signal processing and Fourier features, cosine learning-rate schedules, and rotation matrices.
How is it used?
Use $(\sin)'=\cos$ and $(\cos)'=-\sin$ together with the chain rule: $\frac{d}{dx}\sin(3x) = 3\cos(3x)$.
Radians only! These formulas assume $x$ is in radians. In degrees there is a stray factor: $\frac{d}{dx}\sin(x^\circ) = \frac{\pi}{180}\cos(x^\circ)$. Always do calculus in radians.
Signs. Only $\cos$ has a minus sign in its derivative: $(\sin)'=\cos$ but $(\cos)'=-\sin$. After four derivatives you are back where you started.
Quick check: differentiate $f(x) = x\cos x$.
Product rule: $1\cdot\cos x + x\cdot(-\sin x) = \cos x - x\sin x$.
The rules at a glance: a derivative table and picker
You now own a small toolbox. There are only four combining rules (sum, product, quotient, chain) and a short list of basic derivatives. Any formula built from these pieces can be differentiated by splitting it up and applying them one at a time.
You do not need to memorise the table. Every line in it was derived above, and you can re-derive it. The table is just a place to look things up.
Differentiate $f(x) = x^2 e^{3x}\sin x$? Break it into the pieces it is made of.
- It is a product of three things. Group as $u = x^2e^{3x}$ and $v=\sin x$. Then $f' = u'v + uv'$.
- For $u = x^2\cdot e^{3x}$, the product rule again: $u' = 2x\,e^{3x} + x^2\cdot3e^{3x}$ (chain rule for $e^{3x}$).
- And $v' = \cos x$.
- So $f'(x) = \big(2x + 3x^2\big)e^{3x}\sin x + x^2e^{3x}\cos x$.
Each step used only a rule from this chapter.
Combining rules (with $c$ a constant):
| Rule | Formula |
|---|---|
| Constant multiple | $(c\,f)' = c\,f'$ |
| Sum / difference | $(f\pm g)' = f'\pm g'$ |
| Product | $(uv)' = u'v + uv'$ |
| Quotient | $(u/v)' = \dfrac{u'v - uv'}{v^2}$ |
| Chain | $\big(f(g(x))\big)' = f'(g(x))\,g'(x)$ |
Basic derivatives:
| $f(x)$ | $f'(x)$ | $f(x)$ | $f'(x)$ |
|---|---|---|---|
| $c$ | $0$ | $\ln x$ | $1/x$ |
| $x^n$ | $n\,x^{n-1}$ | $\log_a x$ | $\dfrac{1}{x\ln a}$ |
| $e^x$ | $e^x$ | $\sin x$ | $\cos x$ |
| $a^x$ | $(\ln a)\,a^x$ | $\cos x$ | $-\sin x$ |
| $1/x$ | $-1/x^2$ | $\tan x$ | $1/\cos^2 x$ |
| $\sqrt x$ | $\dfrac{1}{2\sqrt x}$ | $\sigma(x)$ | $\sigma(x)(1-\sigma(x))$ |
| $\tanh x$ | $1-\tanh^2x$ | $\max(0,x)$ | $0$ or $1$ (kink at $0$) |
Why do we need it?
A short, trusted list means you can differentiate nearly any formula you meet without going back to limits each time.
Where is it used?
Deriving gradients by hand for a new loss function, checking what automatic differentiation (autograd) in PyTorch and JAX should return, and exam-style or interview-style derivations.
How is it used?
Split the formula into sums, products, fractions and nested layers. Look up each basic piece, then combine. Always confirm with a nudge test on a number.
Quick check: differentiate $f(x) = \dfrac{\ln x}{x}$.
Quotient rule, $u=\ln x$, $v=x$: $f' = \dfrac{\frac1x\cdot x - \ln x\cdot 1}{x^2} = \dfrac{1-\ln x}{x^2}$. It is zero at $x=e$ (the peak of $\ln x/x$).
Derivatives of sigmoid and ReLU core
Neural networks put a little bend into every neuron using an activation function. Two of the most common:
- Sigmoid $\sigma(x)$ squashes any number into the range $0$ to $1$ with a smooth S shape. It is steepest in the middle and almost flat in both tails.
- ReLU is $\max(0,x)$: flat at $0$ for negative inputs, then a straight line of slope $1$. It has a sharp kink at $0$.
The slope of the activation decides how much of the learning signal gets through. A flat part blocks it: no slope, no learning. A sigmoid in its flat tail is "saturated", and a ReLU on its negative side is "off".
Sigmoid values and slopes, using $\sigma'(x)=\sigma(x)\big(1-\sigma(x)\big)$ (derived below):
| $x$ | $\sigma(x)$ | $\sigma'(x)$ | meaning |
|---|---|---|---|
| $0$ | 0.5000 | 0.2500 | the steepest point |
| $2$ | 0.8808 | 0.1050 | still learning |
| $5$ | 0.9933 | 0.0066 | almost flat |
| $-5$ | 0.0067 | 0.0066 | almost flat (same, by symmetry) |
| $10$ | 0.99995 | 0.000045 | saturated: signal nearly dead |
Stacking makes it worse. Even in the best case the slope is $0.25$. Through $5$ sigmoid layers the signal is multiplied by at most $0.25^5 \approx 0.001$, and through $10$ layers by $0.25^{10}\approx 9.5\times10^{-7}$ (counting only the sigmoid slopes; the weights multiply the signal too). This is the vanishing gradient problem.
ReLU: slope $0$ for $x\lt 0$, slope $1$ for $x>0$. Passing through 10 active ReLUs multiplies the signal by $1^{10}=1$ (again counting only the activation slopes): nothing is lost to the activations.
Sigmoid. $\sigma(x) = \dfrac{1}{1+e^{-x}} = (1+e^{-x})^{-1}$. Derivation with the chain rule (inner $u = 1+e^{-x}$, outer $u^{-1}$):
$$\sigma'(x) = -(1+e^{-x})^{-2}\cdot(-e^{-x}) = \frac{e^{-x}}{(1+e^{-x})^2}.$$Split the fraction into two pieces:
$$\sigma'(x) = \frac{1}{1+e^{-x}}\cdot\frac{e^{-x}}{1+e^{-x}} = \sigma(x)\cdot\big(1-\sigma(x)\big),$$because $1-\sigma = \dfrac{(1+e^{-x}) - 1}{1+e^{-x}} = \dfrac{e^{-x}}{1+e^{-x}}$. ∎ The slope never exceeds $\tfrac14$ (at $x=0$, where $\sigma=\tfrac12$ and $\tfrac12\cdot\tfrac12 = \tfrac14$). A close cousin: $\tanh'(x) = 1-\tanh^2(x)$, with maximum slope $1$.
ReLU. $\text{ReLU}(x) = \max(0,x)$ has $\text{ReLU}'(x) = 0$ for $x\lt 0$ and $1$ for $x>0$. At $x=0$ the two one-sided slopes ($0$ from the left, $1$ from the right) disagree, so the derivative does not exist at exactly $0$. Libraries just pick a value (usually $0$). Leaky ReLU uses a small slope $\alpha$ (like $0.01$) for $x\lt 0$, so the signal never fully dies.
Why do we need it?
The activation's slope is a factor in every gradient flowing back through the network. We must know where it is large (learning passes) and where it is near 0 (learning stalls).
Where is it used?
Backpropagation through every neural network layer, logistic regression (sigmoid), LSTM and GRU gates, and choosing activations such as ReLU, Leaky ReLU, GELU or tanh to avoid vanishing gradients.
How is it used?
During training, the framework multiplies the incoming gradient by $\sigma'(z)=\sigma(z)(1-\sigma(z))$ (re-using the already computed $\sigma(z)$), or by 0/1 for ReLU. You watch for saturated or "dead" units.
Large inputs, not small ones, kill a sigmoid. The slope is largest near $x=0$. If a layer's inputs are huge (say $\pm 10$), the sigmoid is saturated whatever the weights do. This is one reason inputs are normalised.
"The derivative does not exist" is not a disaster. ReLU's kink is a single point. In practice you almost never land exactly on it, and picking either one-sided slope works.
Quick check: what is $\sigma'(3)$, given $\sigma(3)\approx0.9526$?
$\sigma'(3) = 0.9526\times(1-0.9526) = 0.9526\times0.0474 \approx 0.0452$. About 18% of the best-case $0.25$.
Higher-order derivatives: $f''$ and $f'''$
A derivative is itself a function, so we can take its derivative. Back to the car: distance $\to$ (derivative) speed $\to$ (derivative) acceleration. Acceleration is "how fast the speed is changing", the second derivative of distance.
On a graph the second derivative tells you how the curve bends:
- $f''>0$: the slope is increasing, so the curve bends upward like a bowl or smile.
- $f''\lt 0$: the slope is decreasing, so the curve bends downward like a frown or hilltop.
- $f''=0$: momentarily straight (often where the bend switches).
This is a short preview. Curvature and its multi-variable version, the Hessian, get their full treatment in Chapter 2.10.
Take $f(x) = x^4 - 2x^3$.
- $f'(x) = 4x^3 - 6x^2$.
- $f''(x) = 12x^2 - 12x$ (differentiate $f'$).
- $f'''(x) = 24x - 12$.
- $f^{(4)}(x) = 24$, and every derivative after that is $0$.
At $x=1$: $f=-1$, $f'=-2$ (falling), $f''=0$ (the bend switches here), $f'''=12$. For $\sin x$ the derivatives cycle: $\sin,\ \cos,\ -\sin,\ -\cos,\ \sin,\dots$ Every derivative of $e^x$ is $e^x$. For the car $s=3t^2-t^3/3$: speed $v = 6t - t^2$, acceleration $a = 6-2t$.
The second derivative is the derivative of the derivative:
$$f''(x) = \frac{d}{dx}\big[f'(x)\big] = \frac{d^2f}{dx^2}.$$The third is $f'''(x) = \dfrac{d^3f}{dx^3}$, and the $n$-th is written $f^{(n)}(x)$ or $\dfrac{d^nf}{dx^n}$. Read $\dfrac{d^2f}{dx^2}$ as "d two f, d x squared". Meaning: $f'$ is the slope, $f''$ is the rate at which the slope changes (the bend), $f'''$ is the rate at which the bend changes.
Why do we need it?
The slope says which way is downhill but not how quickly the ground bends. Knowing the bend tells us whether we are in a valley or on a hilltop, and how big a step is safe.
Where is it used?
The second-derivative test for minima, Newton's method, curvature of loss surfaces (the Hessian), choosing learning rates, Taylor approximations, and acceleration in physics.
How is it used?
Differentiate once, then differentiate the result. Look at the sign of $f''$: positive means a valley-shaped bend, negative means a hill-shaped bend.
Quick check: find $f''(x)$ for $f(x)=x^3-3x$ and say where the curve bends upward.
$f' = 3x^2 - 3$, $f'' = 6x$. Positive for $x>0$: bends upward there. Negative for $x\lt 0$: bends downward.
Machine learning link 1: a loss is a function of a parameter
A model has knobs called parameters (or weights). For each setting of the knobs, the model makes predictions, and a loss gives one number saying how wrong they are: big = bad, small = good.
Fix the data and turn one knob $w$. The loss changes with $w$, so the loss is just a function $L(w)$, drawn as a curve. Learning means finding the $w$ at the bottom of that curve. And the derivative $L'(w)$ at your current spot tells you which way the bottom is, and how steep the ground is.
Three data points $(x,y) = (1,2),\ (2,4),\ (3,5)$ and a model with one knob: $\hat y = w\,x$. The loss is the mean squared error $L(w) = \tfrac13\sum (w x_i - y_i)^2$.
- Expand: $\tfrac13\big[(w-2)^2 + (2w-4)^2 + (3w-5)^2\big] = \tfrac13\big[14w^2 - 50w + 45\big]$.
- Differentiate with the power rule: $L'(w) = \tfrac13(28w - 50)$.
- Check the sign: $L(0) = 15$, $L'(0) = -16.67$. Negative slope: turning $w$ up lowers the loss. Also $L(1)=3$, $L'(1)=-7.33$; $L(2) = 0.33$, $L'(2) = 2$ (positive: now too big).
- The slope is zero at $28w = 50$, so $w^* = 25/14 \approx 1.786$: the bottom of the curve.
Same answer by the chain rule, for any data: each term is $(w x_i - y_i)^2$. The outer power gives $2(w x_i - y_i)$ and the inner part has slope $x_i$, so $L'(w) = \tfrac{2}{n}\sum x_i(w x_i - y_i)$. Check: $\tfrac23[(w-2) + 2(2w-4) + 3(3w-5)] = \tfrac23(14w - 25) = \tfrac13(28w-50)$ ✓.
A loss function $L(w)$ measures how badly a model with parameter $w$ fits the data. Training means finding $w$ that makes $L$ small. The derivative
$$L'(w) = \frac{dL}{dw}$$is the sensitivity of the loss to the parameter: how much $L$ changes per unit change in $w$. For mean squared error with a one-weight line $\hat y = wx$: $L(w)=\frac1n\sum_i (w x_i - y_i)^2$ and $L'(w)=\frac2n\sum_i x_i\,(w x_i - y_i)$. (Many weights at once: Chapter 2.4.)
Why do we need it?
"Learn from data" has to become a precise question: which parameter value is best? A loss turns that into "find the lowest point of a curve", and calculus knows how to look for it.
Where is it used?
Every supervised model: linear and logistic regression, neural networks, and language models (cross-entropy loss). Mean squared error is used for regression, cross-entropy for classification.
How is it used?
Write the loss as a function of the parameters, compute its derivative, and use the derivative to move the parameters downhill (next sections).
The loss is a function of the parameters, not of the data. The data is fixed while we train. We ask how the loss changes as $w$ changes. (The data points in the widget are movable only so you can see how the curve depends on them.)
Quick check: with the loss $L(w) = (w-4)^2$, what is $L'(1)$ and which way should $w$ move?
$L'(w) = 2(w-4)$, so $L'(1) = -6$. Negative slope means increasing $w$ lowers the loss: move $w$ up, towards 4.
Machine learning link 2: gradient descent in one dimension core
You are on a foggy hillside and want the valley floor. You cannot see it, but you can feel the slope under your feet. If the ground tilts up to your right, step left. If it tilts up to your left, step right. Always step against the slope.
How big a step? Steep ground: take a bigger step (you are far from the bottom). Almost flat: take a tiny step (you are close). So step size = learning rate × slope. The slope shrinks as you near the bottom, so the steps shrink automatically.
The learning rate is how bold you are. Too timid and you crawl. Too bold and you leap clear over the valley and land higher than before.
Minimise $L(w) = (w-3)^2$, with derivative $L'(w) = 2(w-3)$. Start at $w_0=0$ with learning rate $\eta = 0.1$ ($\eta$ is the Greek letter "eta").
- Slope at $0$: $L'(0) = 2(0-3) = -6$. Step: $-\eta L' = -0.1\cdot(-6) = +0.6$. New $w_1 = 0.6$.
- Slope at $0.6$: $2(0.6-3) = -4.8$. Step $= +0.48$. $w_2 = 1.08$.
- Slope at $1.08$: $2(1.08-3) = -3.84$. Step $= +0.384$. $w_3 = 1.464$.
- The steps shrink ($0.6, 0.48, 0.384, \dots$) and $w$ creeps towards $3$.
The distance to the bottom shrinks by a factor $(1 - 2\eta) = 0.8$ every step: $3 \to 2.4 \to 1.92 \to 1.536$. In general: for $0\lt \eta\lt 0.5$ the approach is smooth; at $\eta=0.5$ you land on $3$ in one step; for $0.5\lt \eta\lt 1$ you zig-zag across the bottom but still shrink in; at $\eta=1$ you bounce between two points forever; for $\eta>1$ the distance grows by $|1-2\eta|>1$ each step and you diverge.
Gradient descent (1D) repeats the update
$$w_{\text{new}} = w_{\text{old}} - \eta\,L'(w_{\text{old}}).$$$\eta>0$ is the learning rate (a number you choose). The minus sign sends you downhill. It stops (in practice) when $L'$ is almost $0$, or after a set number of steps. This is the core move of training; with many weights, $L'$ becomes the gradient vector (Chapter 2.4), but the idea is identical.
Why do we need it?
Most losses cannot be solved by algebra. But we can always compute the slope at the current point. Following the slope downhill finds a low point without solving anything.
Where is it used?
Training linear and logistic regression, neural networks and deep learning (as SGD, Adam and friends), matrix factorisation, and almost any model fitted by minimising a loss.
How is it used?
Pick a start and a learning rate. Repeat: compute the slope, subtract $\eta\times$slope from the parameter. Watch the loss curve: too slow means raise $\eta$, exploding means lower it.
The learning rate $\eta$ must be tuned. The "safe" range depends on how curved the loss is. For $L = (w-3)^2$ it is $\eta\lt 1$. For a sharper bowl like $L=10(w-3)^2$ the limit is $\eta\lt 0.1$. The general rule for a parabola-shaped loss: if its second derivative (the bend, $L''$) is a constant, every step multiplies the distance to the bottom by $1-\eta L''$, so descent converges only when $\eta\lt 2/L''$. Here $L''=2$ and $L''=20$. Loss curves that shoot up are the classic sign of a learning rate that is too high.
Gradient descent finds a flat spot, not necessarily the best one. On a curve with several valleys it can end in a shallow one. See the next section.
Quick check: one step of gradient descent on $L(w) = w^2$ from $w=4$ with $\eta=0.25$.
$L'(w) = 2w = 8$. New $w = 4 - 0.25\cdot 8 = 2$. The loss fell from $16$ to $4$.
Machine learning link 3: optimisation, where the derivative is zero core
Walk along a hilly path. At the very top of a hill you stop going up and start going down. At the bottom of a valley you stop going down and start going up. At both turning points the path is, for an instant, flat. So the slope is zero at every hilltop and every valley bottom.
That gives a strategy for finding the best value of anything: find where the derivative is zero. Then tell hill from valley by how the slope changes as you walk through:
- slope goes from $+$ to $-$ (up, then down): a hilltop (maximum);
- slope goes from $-$ to $+$ (down, then up): a valley (minimum);
- slope touches 0 but keeps its sign: a flat shelf, neither.
Find the turning points of $f(x) = x^3 - 3x$.
- $f'(x) = 3x^2 - 3 = 3(x-1)(x+1)$. It is zero at $x=-1$ and $x=1$.
- Test a point in each region. $x=-2$: $f' = 9>0$ (rising). $x=0$: $f' = -3\lt 0$ (falling). $x=2$: $f' = 9>0$ (rising).
- At $x=-1$ the slope goes $+\to-$: a local maximum, $f(-1) = 2$.
- At $x=1$ the slope goes $-\to+$: a local minimum, $f(1) = -2$.
- Second-derivative check: $f''(x)=6x$. $f''(-1)=-6\lt 0$ (bends down: hill ✓). $f''(1)=6>0$ (bends up: valley ✓).
A shelf: $f(x)=x^3$ has $f'(0)=0$, but $f'=3x^2$ is positive on both sides, so $x=0$ is neither a max nor a min.
A critical point (or stationary point) is an $x$ where $f'(x)=0$. At a critical point:
- First-derivative test: look at the sign of $f'$ just left and right. $+\to-$ is a local max; $-\to+$ is a local min; no sign change means neither.
- Second-derivative test: if $f''>0$ it is a local min, if $f''\lt 0$ a local max, and if $f''=0$ the test is inconclusive.
A local minimum is the lowest point in its own neighbourhood. The global minimum is the lowest point of the whole function. Local minima need not be global.
Why do we need it?
Training a model is an optimisation problem: find the parameters with the lowest loss. "Slope equals zero" is the signpost that says "you may have arrived".
Where is it used?
Closed-form solutions (the normal equation of linear regression comes from setting the loss slope to 0), stopping rules for gradient descent ("stop when the gradient is tiny"), and understanding local minima and saddle points in neural-network loss landscapes.
How is it used?
Compute $f'$, solve $f'=0$ if you can, and classify each solution by the sign change or by $f''$. If you cannot solve it, run gradient descent until $f'\approx0$.
$f'=0$ is necessary, not sufficient. It could be a max, a min, or a shelf. Always classify.
Local is not global. Gradient descent on a bumpy loss may settle in a shallow valley. Neural-network losses have many flat spots; the surprise of modern deep learning is that simple gradient descent works well anyway.
Quick check: find and classify the critical point of $f(x) = x^2 - 6x + 5$.
$f' = 2x - 6 = 0$ gives $x=3$. $f''=2>0$, so it is a minimum, with $f(3) = 9 - 18 + 5 = -4$.
Machine learning link 4: sensitivity, $\Delta y \approx f'(x)\,\Delta x$ core
A lamp is wired to a dial. Nudge the dial a tiny bit and watch the brightness. If a small nudge changes the brightness a lot, the lamp is very sensitive to the dial. If hardly anything happens, it is insensitive. The ratio (brightness change) ÷ (dial change) is the derivative.
So the derivative lets us predict the effect of a small nudge without recomputing anything: new output $\approx$ old output + slope × nudge. It works because the curve looks straight up close.
$f(x) = x^2$ at $x = 3$, where $f'(3) = 6$.
- Nudge by $\Delta x = 0.1$. Prediction: $\Delta y \approx 6\times0.1 = 0.6$.
- Truth: $f(3.1) - f(3) = 9.61 - 9 = 0.61$. Error: $0.01$.
- Nudge by $\Delta x = 0.01$. Prediction $0.06$. Truth: $9.0601 - 9 = 0.0601$. Error: $0.0001$.
A nudge ten times smaller gives an error a hundred times smaller (the error is about $(\Delta x)^2$). ML version: loss $L(w) = (w-3)^2$ at $w = 5$ has $L'(5) = 4$. Raising $w$ by $0.1$ should raise the loss by about $0.4$; the truth is $0.41$.
For a small change $\Delta x$ in the input:
$$\Delta y = f(x+\Delta x) - f(x) \;\approx\; f'(x)\,\Delta x.$$$|f'(x)|$ is the sensitivity of the output to the input at $x$. Large $|f'|$: sensitive. $f'=0$: insensitive to first order. The error of this estimate shrinks like $(\Delta x)^2$ (more on this with Taylor series and linearization). The same formula covers error propagation: an uncertainty $\pm\Delta x$ in the input becomes about $\pm|f'(x)|\Delta x$ in the output.
Why do we need it?
We constantly ask "what if this number were a bit different?" The derivative answers it instantly, with one multiplication instead of re-running the model.
Where is it used?
Finding which weights or features matter most (large $|\partial L/\partial w|$), gradient descent (it moves a weight in proportion to its sensitivity), saliency maps, adversarial examples (tiny nudges that flip a prediction), and uncertainty estimates.
How is it used?
Compute $f'(x)$ once. Then $\Delta y\approx f'(x)\Delta x$ for any small $\Delta x$. Compare $|f'|$ across inputs to rank them by sensitivity.
Only for small nudges. The prediction is the tangent line, and the tangent leaves the curve as you move away. For a big nudge the error is big too. Zoom out in the widget (use $\Delta x=1$) to see this.
Zero slope does not mean "no effect". At a minimum $f'=0$, but a big enough nudge still changes the output, as the curvature (the second derivative) kicks in.
Quick check: $f'(2)=5$. Estimate $f(2.02)-f(2)$.
$\Delta y\approx f'(2)\cdot\Delta x = 5\times0.02 = 0.1$.
Machine learning link 5: learning curves and their slope
While a model trains, we write down its loss after every step. Plot loss against step number and you get a learning curve. It is a function of time, so it has a slope too: the slope of the learning curve is how fast the model is improving, right now.
- Steeply falling: fast progress.
- Gently falling: slow progress, getting close to the bottom.
- Slope near zero: a plateau. The model has converged, or is stuck.
- Slope positive: the loss is rising. Something is wrong (learning rate too high), or it is the validation loss turning up (overfitting).
Take the gradient-descent run from before ($L=(w-3)^2$, $\eta=0.1$, start at $0$). The loss after each step is $9,\ 5.76,\ 3.6864,\ 2.3593,\ 1.5099,\dots$ (each is $0.64\times$ the last).
- Slope from step 0 to 1: $5.76 - 9 = -3.24$ per step.
- Slope from step 1 to 2: $3.6864 - 5.76 = -2.0736$.
- Slope from step 2 to 3: $2.3593 - 3.6864 = -1.3271$.
The slope is negative (improving) but its size shrinks: fast progress early, slow progress later. A learning curve that flattens out is the visual signal for convergence.
A learning curve plots the loss $L(t)$ against the training step $t$. Its slope is the rate of improvement:
$$\frac{dL}{dt}\approx\frac{L(t+k)-L(t-k)}{2k}\quad(\text{average over nearby steps}).$$The training loss is measured on data the model learns from; the validation loss on held-out data. Early stopping stops training at the step where the validation curve stops falling, which is where its slope reaches $0$ (the bottom of a valley in $t$!).
Why do we need it?
You cannot see inside a training run. The learning curve and its slope are your instruments: they say whether training is working, finished, stuck or broken.
Where is it used?
Every training run (TensorBoard and Weights & Biases plot them), early stopping, learning-rate schedules and warm-up, and debugging exploding or stalled training.
How is it used?
Watch the curve. Falling quickly: carry on. Flat: stop, or change the learning rate. Rising or jagged: lower the learning rate. Validation turning up while training keeps falling: overfitting, stop early.
Noise hides the slope. Real curves are jagged, so one step's difference is unreliable. Look at the trend over many steps (smooth the curve, or average over a window as the widget does).
A flat curve has two very different meanings: "finished" or "stuck". Only the loss value tells which: low and flat is converged; high and flat is stuck, so change something.
Quick check: the loss goes 4.0, 3.0, 2.4, 2.1, 2.0 over four steps. Is the slope getting steeper or flatter?
Slopes per step: $-1.0,\ -0.6,\ -0.3,\ -0.1$. They are all negative but shrinking: flatter, so training is slowing down towards a plateau.
Recap, cheat sheet and practice
- The derivative $f'(x)=\lim_{h\to0}\frac{f(x+h)-f(x)}{h}$ is the slope of the tangent line, the instantaneous rate of change, and the sensitivity of the output to a nudge: $\Delta y\approx f'(x)\Delta x$.
- First principles: secant slope, simplify until $h$ cancels, let $h\to0$. This gave $x^2\to2x$, $x^3\to3x^2$, $1/x\to-1/x^2$, $\sqrt x\to1/(2\sqrt x)$.
- Rules: power $nx^{n-1}$; product $u'v+uv'$; quotient $(u'v-uv')/v^2$; chain $f'(g(x))\,g'(x)$ ("rates multiply along the chain").
- Special functions: $(e^x)'=e^x$, $(a^x)'=(\ln a)a^x$, $(\ln x)'=1/x$, $(\sin)'=\cos$, $(\cos)'=-\sin$, $(\tan)'=1/\cos^2$.
- ML: $\sigma'=\sigma(1-\sigma)\le0.25$ (vanishing gradients); ReLU has slope 0 or 1 with a kink at 0; a loss is a function $L(w)$ of a parameter; gradient descent $w\leftarrow w-\eta L'(w)$; critical points have $f'=0$ (classify by sign change or $f''$); a learning curve's slope says how fast you are improving.
- $f''$ is the slope of the slope (the bend): $f''>0$ bends up (valley), $f''\lt 0$ bends down (hill).
Cheat sheet
| Idea | Formula | Picture / meaning |
|---|---|---|
| Derivative | $\lim_{h\to0}\frac{f(x+h)-f(x)}{h}$ | slope of the tangent |
| Tangent line at $a$ | $y=f(a)+f'(a)(x-a)$ | best straight-line copy |
| Nudge estimate | $\Delta y\approx f'(x)\Delta x$ | sensitivity |
| Sum / multiple | $(cf+g)'=cf'+g'$ | differentiate piece by piece |
| Power | $(x^n)'=nx^{n-1}$ | bring down, lower by one |
| Product | $(uv)'=u'v+uv'$ | two strips of area |
| Quotient | $(u/v)'=\frac{u'v-uv'}{v^2}$ | product rule in disguise |
| Chain | $\frac{dy}{dx}=\frac{dy}{du}\frac{du}{dx}$ | gears: rates multiply |
| Exponential | $(e^{g})'=g'e^{g}$, $(a^x)'=\ln a\cdot a^x$ | slope proportional to height |
| Logarithm | $(\ln g)'=g'/g$, $(\log_a x)'=\frac1{x\ln a}$ | mirror of $e^x$: reciprocal slope |
| Trig | $(\sin)'=\cos,\ (\cos)'=-\sin,\ (\tan)'=1/\cos^2$ | radians only |
| Sigmoid | $\sigma'=\sigma(1-\sigma)$, max $0.25$ | flat tails: vanishing gradient |
| ReLU | $0$ for $x\lt 0$, $1$ for $x>0$ | kink at $0$ |
| Higher order | $f''=(f')'$ | bend: $>0$ smile, $\lt 0$ frown |
| Gradient descent | $w\leftarrow w-\eta L'(w)$ | step against the slope |
| Critical point | $f'(x)=0$ | hill, valley or shelf |
import numpy as np
def num_deriv(f, x, h=1e-5):
"""Central-difference 'nudge test' for the derivative of f at x."""
return (f(x + h) - f(x - h)) / (2 * h)
# 1. Check hand-derived formulas against the nudge test
print(round(num_deriv(lambda x: x**2, 3.0), 4)) # 6.0 (2x at x=3)
f = lambda x: x**2 * (3*x + 1) # product rule example
print(round(num_deriv(f, 2.0), 4), 9*2.0**2 + 2*2.0) # 40.0 40.0
g = lambda x: np.sin(x**2) # chain rule example
print(round(num_deriv(g, 1.0), 4), round(2*1.0*np.cos(1.0**2), 4)) # 1.0806 1.0806
# 2. Sigmoid and its derivative sigma(1 - sigma)
sigmoid = lambda z: 1 / (1 + np.exp(-z))
z = np.array([-5.0, 0.0, 5.0])
s = sigmoid(z)
print(np.round(s * (1 - s), 5)) # [0.00665 0.25 0.00665]
print(0.25 ** 10) # 9.5367431640625e-07 (10 sigmoid layers, best case)
# 3. A loss as a function of one weight (three data points)
X = np.array([1.0, 2.0, 3.0]); Y = np.array([2.0, 4.0, 5.0])
L = lambda w: np.mean((w * X - Y) ** 2)
dL = lambda w: 2 * np.mean(X * (w * X - Y)) # chain rule by hand
print(round(L(0.0), 4), round(dL(0.0), 4), round(num_deriv(L, 0.0), 4)) # 15.0 -16.6667 -16.6667
print(np.sum(X * Y) / np.sum(X * X), 25 / 14) # 1.7857142857142858 1.7857142857142858 (where L' = 0)
# 4. Gradient descent on L(w) = (w - 3)^2 with three learning rates
for eta in (0.02, 0.1, 1.1):
w = 0.0
for k in range(25):
w = w - eta * 2 * (w - 3) # w_new = w_old - eta * L'(w_old)
print(eta, round(w, 4)) # 0.02 -> 1.9188 (too slow), 0.1 -> 2.9887 (close to 3), 1.1 -> 289.1886 (it diverges)
1. What is the derivative of $f(x)=x^3$ at $x=2$?
2. Which is $\dfrac{d}{dx}\big[x^2e^x\big]$?
3. What is $\dfrac{d}{dx}\cos(3x)$?
4. What is $\dfrac{d}{dx}\ln(x^2)$ for $x>0$?
5. One step of gradient descent on $L(w)=w^2$ from $w=5$ with learning rate $\eta=0.3$ gives a new $w$ of…
6. Why can a deep stack of sigmoid layers have vanishing gradients?
Practice problems
A. Use first principles to differentiate $f(x)=x^2+3x$.
$\dfrac{(x+h)^2+3(x+h) - x^2 - 3x}{h} = \dfrac{2xh + h^2 + 3h}{h} = 2x + h + 3$. Let $h\to0$: $f'(x)=2x+3$. (At $x=1$ the slope is $5$.)
B. Differentiate $f(x)=(2x^3-x)(x^2+5)$ and evaluate at $x=1$.
$u=2x^3-x$, $v=x^2+5$, $u'=6x^2-1$, $v'=2x$. $f'=(6x^2-1)(x^2+5)+(2x^3-x)(2x) = 6x^4+29x^2-5 + 4x^4-2x^2 = 10x^4+27x^2-5$. At $x=1$: $10+27-5=32$. (Check: $f=2x^5+9x^3-5x$, so $f'=10x^4+27x^2-5$ ✓.)
C. Use the quotient rule on $f(x)=\dfrac{e^x}{1+e^x}$ and compare with the sigmoid.
$u=e^x$, $v=1+e^x$, $u'=v'=e^x$. $f' = \dfrac{e^x(1+e^x)-e^x\cdot e^x}{(1+e^x)^2} = \dfrac{e^x}{(1+e^x)^2}$. This is the sigmoid ($\frac{e^x}{1+e^x}=\frac1{1+e^{-x}}$) and its derivative $\sigma(1-\sigma)$ again: $\sigma\cdot(1-\sigma)=\frac{e^x}{1+e^x}\cdot\frac{1}{1+e^x}$ ✓.
D. Differentiate the "softplus" function $f(x)=\ln(1+e^x)$. What famous function do you get?
Chain rule: inner $u=1+e^x$ (derivative $e^x$), outer $\ln u$ (derivative $1/u$). $f'(x)=\dfrac{e^x}{1+e^x}=\sigma(x)$, the sigmoid. So softplus is a smooth version of ReLU whose slope is the sigmoid.
E. Find and classify the critical points of $f(x)=x^3-6x^2+9x+1$.
$f'=3x^2-12x+9=3(x-1)(x-3)$, zero at $x=1$ and $x=3$. $f''=6x-12$: $f''(1)=-6\lt 0$ so $x=1$ is a local maximum, $f(1)=5$. $f''(3)=6>0$ so $x=3$ is a local minimum, $f(3)=1$.
F. (i) Do two gradient-descent steps on $L(w)=(w-2)^2$ from $w_0=6$ with $\eta=0.25$. (ii) For $L(w)=w^3-w$ at $w=2$, estimate the change in loss if $w$ rises by $0.05$.
(i) $L'=2(w-2)$. $L'(6)=8$, so $w_1=6-0.25\cdot8=4$. $L'(4)=4$, so $w_2=4-0.25\cdot4=3$. The distance to $2$ halves each step ($4\to2\to1$), because $1-2\eta=0.5$.
(ii) $L'(w)=3w^2-1$, so $L'(2)=11$. Estimate: $\Delta L\approx11\times0.05=0.55$. Truth: $2.05^3-2.05-6=0.5651$. The error is about $0.015$, from the curvature.
Partial Differentiation & Gradients
Real models have many inputs, not one. This chapter teaches you to find the slope of a landscape: how steep it is east, how steep it is north, and which way is straight uphill. That one arrow, the gradient, is the compass of machine learning.
- See a function of two (or more) inputs as a surface or a coloured height map
- Find a partial derivative: the slope when only one input moves and the others stand still
- Pack all the partial derivatives into the gradient, and say what it means
- Compute the slope in any direction (the directional derivative) with a dot product
- Know why the gradient points to steepest ascent and its negative to steepest descent
- Read level sets and contour maps, and see why the gradient is perpendicular to them
Functions of several inputs: surfaces and height maps core
So far a function took one number in and gave one number out. Real problems have many inputs. A house price depends on its area and its age and its distance from the city. A model's error depends on all of its weights.
The easiest way to picture two inputs is a landscape. Stand on a map. Your position is a pair of numbers: how far east ($x$) and how far north ($y$). At every position the ground has a height. The rule "position in, height out" is a function of two inputs, $f(x, y)$.
Draw the height above every spot and you get a surface. Or look straight down from a helicopter and colour each spot by its height: that is a height map. Both pictures show the same function.
In machine learning the "position" is the set of weights and the "height" is the loss (the error). Training means walking downhill on that landscape.
Let $f(x, y) = x^2 + y^2$. To evaluate it, put in a position and compute the height.
- At $(0, 0)$: $f = 0^2 + 0^2 = 0$. (the lowest point)
- At $(1, 2)$: $f = 1^2 + 2^2 = 1 + 4 = 5$.
- At $(-2, 0)$: $f = (-2)^2 + 0^2 = 4$.
- At $(3, -1)$: $f = 9 + 1 = 10$.
Every point at the same distance from the centre has the same height, so this surface is a round bowl. The function $x^2 - y^2$ gives a different shape, a saddle (up in one direction, down in the other), and $xy$ gives another saddle turned by 45°.
A function of several variables takes a vector of inputs and returns one number:
$$f:\ \mathbb{R}^n \to \mathbb{R}, \qquad \mathbf{x} = \begin{bmatrix} x_1 \\ \vdots \\ x_n \end{bmatrix} \ \longmapsto\ f(\mathbf{x}).$$For two inputs we write $f(x, y)$. Its graph is the set of points $(x, y, f(x, y))$ in 3D: a surface whose height above the floor point $(x, y)$ is the output. (A vector is just a list of numbers, so "many inputs" is "one vector input".)
With $n$ inputs the graph needs $n+1$ dimensions. We cannot draw that, but every idea in this chapter works for any $n$. We draw $n = 2$ and trust it for $n = 1\,000\,000$.
Why do we need it?
One input is rarely enough. To describe anything that depends on several things at once, such as a price, a temperature or an error, we need a rule that accepts several numbers and returns one.
Where is it used?
Every loss function: it takes all the weights of a model (thousands to billions of numbers) and returns one number, the error. Also cost surfaces in optimisation, potential energy in physics, and heat maps of any score.
How is it used?
Think "position in, height out". To see a function, plot its surface or its height map. To train a model, you move the position (the weights) so that the height (the loss) goes down.
The height is the output, not a third input. A function of two inputs $f(x, y)$ has the input pair $(x, y)$ on the floor, and the output $z = f(x, y)$ is the height. Do not confuse the vertical axis with an input.
Pictures need care. A surface hides what is behind it. That is why people also use height maps and contour lines (you will meet them below).
Quick check: for $f(x, y) = x^2 - y^2$, what are $f(2, 1)$ and $f(1, 2)$?
$f(2, 1) = 4 - 1 = 3$ and $f(1, 2) = 1 - 4 = -3$. The order of the inputs matters: $x$ goes first, $y$ second. That is exactly why this surface is a saddle: it is up along the $x$-direction and down along the $y$-direction.
Partial derivatives core
You stand on a hillside. In Chapter 2.3 you learned that the derivative is a slope. On a surface there is no single slope: it depends on which way you face.
- Walk due east (only $x$ changes, $y$ stays fixed). You feel one slope. That is $\partial f/\partial x$.
- Walk due north (only $y$ changes, $x$ stays fixed). You feel a different slope. That is $\partial f/\partial y$.
Picture it as slicing. Cut the surface with a vertical wall along the line $y = $ constant. The edge of the cut is an ordinary curve (a function of $x$ alone). Its slope is the partial derivative $\partial f/\partial x$.
So the recipe is: freeze every input except one, then take an ordinary derivative. The frozen ones are treated like plain constants.
Let $f(x, y) = x^2 y + 2y$. Find both partial derivatives, then evaluate them at $(3, 2)$.
$\partial f/\partial x$: freeze $y$ (treat it as a number, like 5).
- $x^2 y$ is "$x^2$ times a constant $y$". Its derivative in $x$ is $2x\cdot y = 2xy$.
- $2y$ has no $x$ in it, so it is a constant: its derivative is $0$.
- So $\dfrac{\partial f}{\partial x} = 2xy$. At $(3, 2)$: $2\cdot3\cdot2 = 12$.
$\partial f/\partial y$: freeze $x$.
- $x^2 y$ is "a constant $x^2$ times $y$". Its derivative in $y$ is $x^2$.
- $2y$ has derivative $2$.
- So $\dfrac{\partial f}{\partial y} = x^2 + 2$. At $(3, 2)$: $9 + 2 = 11$.
Check by nudging. $f(3, 2) = 9\cdot2 + 4 = 22$. Nudge only $x$ by $0.001$: $f(3.001, 2) = 9.006001\cdot2 + 4 = 22.012002$. The change is $0.012002$, and $0.012002/0.001 = 12.002 \approx 12$ ✓. Nudge only $y$ by $0.001$: $f(3, 2.001) = 18.009 + 4.002 = 22.011$, a change of $0.011$, so the slope is $11$ ✓.
What this means. At $(3, 2)$, a small step in $x$ changes $f$ about 12 times as much as the step; a small step in $y$ changes it about 11 times as much. If we nudge $x$ by $0.01$ and $y$ by $0.02$, the two effects simply add: $12\cdot0.01 + 11\cdot0.02 = 0.12 + 0.22 = 0.34$. The true change is $f(3.01, 2.02) - f(3, 2) = 22.341402 - 22 = 0.341402$ (very close to $0.34$) ✓.
The partial derivative of $f$ with respect to $x$ at the point $(x, y)$ is
$$\frac{\partial f}{\partial x}(x, y) = \lim_{h \to 0} \frac{f(x + h,\, y) - f(x,\, y)}{h},$$and with respect to $y$:
$$\frac{\partial f}{\partial y}(x, y) = \lim_{h \to 0} \frac{f(x,\, y + h) - f(x,\, y)}{h}.$$It is the ordinary derivative with the other inputs held fixed. Notation: $\dfrac{\partial f}{\partial x}$, $f_x$ and $\partial_x f$ all mean the same. The curly $\partial$ ("partial d") warns you that other inputs exist; the straight $d$ is for one-input functions.
For $n$ inputs there are $n$ partial derivatives $\dfrac{\partial f}{\partial x_1}, \dots, \dfrac{\partial f}{\partial x_n}$: one slope for each input, each found by freezing all the others. All the usual rules (power, product, chain, $e^x$, $\sin$ …) still work for the one variable that moves.
Why do we need it?
To fix a model we must know what each single knob does. "If I turn only this weight up a little, does the error rise or fall, and how fast?" Partial derivatives answer that for each knob separately.
Where is it used?
Every training step of linear regression, logistic regression and neural networks needs the partial derivative of the loss with respect to each weight. Also sensitivity analysis (which input matters most) and thermodynamics.
How is it used?
Pick one input. Treat all the others as constants. Differentiate with the normal rules. Repeat for each input. Check one of them by nudging the input by 0.001 and dividing the change in the output by 0.001.
"Frozen" does not mean "zero". When you differentiate in $x$, $y$ is not deleted. It stays in the answer as a constant factor (as in $2xy$).
A partial derivative is a slope in one special direction only. Knowing $f_x$ and $f_y$ does not yet tell you the slope when you walk north-east. (The directional derivative below fixes that.)
You cannot cancel the $\partial$'s like fractions in a casual way. $\dfrac{\partial f}{\partial x}$ is one symbol meaning "slope in the $x$ direction".
Quick check: for $f(x, y) = x^2 y + 2y$, what is $\partial f/\partial y$ at $(1, 5)$?
$\partial f/\partial y = x^2 + 2$. It does not depend on $y$ at all. At $x = 1$ it is $1 + 2 = 3$ (the $y = 5$ is not used).
The gradient core
A hillside gives you two useful numbers at your feet: how steep it is going east, and how steep it is going north. Keeping two numbers separate is clumsy. Better: stack them into one arrow and draw it on the map at your position.
That arrow is the gradient. Think of a weather map with a little wind arrow at every town. The gradient is a "slope arrow" at every position: it points up the hill, and a longer arrow means a steeper hill.
One arrow is also easier to use. To go downhill you simply walk the opposite way.
Example 1. $f(x, y) = x^2 + y^2$. The two partial derivatives are $f_x = 2x$ and $f_y = 2y$, so
$$\nabla f(x, y) = \begin{bmatrix} 2x \\ 2y \end{bmatrix}.$$- At $(1, 2)$: $\nabla f = [2, 4]^\top$. It points up and to the right, away from the centre of the bowl.
- At $(-3, 0)$: $\nabla f = [-6, 0]^\top$. It points left, again away from the centre, and it is longer (the bowl is steeper far out).
- At $(0, 0)$: $\nabla f = [0, 0]^\top$. The bottom of the bowl is flat.
Example 2. From the last section, $f = x^2y + 2y$ has $f_x = 2xy$ and $f_y = x^2 + 2$. At $(3, 2)$ the gradient is $[12, 11]^\top$.
Example 3 (three inputs). $f(x, y, z) = xyz$. Then $f_x = yz$, $f_y = xz$, $f_z = xy$. At $(1, 2, 3)$: $\nabla f = [6, 3, 2]^\top$. Nothing changes except that the list is longer.
The gradient of $f:\mathbb{R}^n\to\mathbb{R}$ is the column vector of all its partial derivatives:
$$\nabla f(\mathbf{x}) = \begin{bmatrix} \dfrac{\partial f}{\partial x_1} \\ \vdots \\ \dfrac{\partial f}{\partial x_n} \end{bmatrix}.$$The symbol $\nabla$ is read "nabla" or "del". Two things to remember:
- $\nabla f(\mathbf{x})$ has the same shape as the input $\mathbf{x}$ ($n$ numbers). That is why an update like $\mathbf{x} - \eta\,\nabla f(\mathbf{x})$ makes sense: you can subtract two vectors of the same size.
- The gradient is a function: it gives one vector for every point. (A vector for every point is called a vector field; see Chapter 2.13.) In this guide the gradient is always a column. Its transpose, a row, is the Jacobian of a scalar function (Chapter 2.5).
Why do we need it?
A model can have a million knobs. We need one object that says, for all of them at once, which way to turn each one to raise the error fastest. The gradient is that object.
Where is it used?
Gradient descent, stochastic gradient descent, Adam and every other optimiser; backpropagation computes the gradient of the loss with respect to all the weights. Also edge detection in images (the gradient of brightness).
How is it used?
Compute every partial derivative, stack them into a column, and evaluate at your current position. Then step against it to lower the loss: new position = old position minus (learning rate) times gradient.
The gradient lives on the floor, not on the surface. It is a vector in the input space (it has one entry per input). It tells you which way to walk, not which way the surface tilts in 3D. (The 3D "tilt" arrow is a different object, drawn in the next section.)
Gradient versus derivative. For one input, $f'(x)$ is a single number. For $n$ inputs, $\nabla f$ is a list of $n$ numbers. When $n = 1$ the two agree.
Quick check: find $\nabla f$ for $f(x, y) = 3x + 5y$ and for $f(x, y) = xy$.
For $3x + 5y$: $f_x = 3$, $f_y = 5$, so $\nabla f = [3, 5]^\top$ at every point (a flat slope that never changes: a tilted plane). For $xy$: $f_x = y$ and $f_y = x$, so $\nabla f = [y, x]^\top$. Notice the swap.
What the gradient tells you (gradient interpretation) core
The gradient at a point tells you three things. Read them like a compass:
- Direction: the arrow points straight uphill, the way the ground rises fastest.
- Length: a long arrow means a steep hill, a short arrow a gentle one.
- Zero: if the arrow has length zero, the ground is flat here. You are at the top of a hill, the bottom of a valley, or the middle of a saddle.
Zoom in far enough on any smooth surface and it looks like a tilted flat sheet (the tangent plane). The gradient is simply the direction in which that sheet tilts upward.
The bowl $f = x^2 + y^2$ has $\nabla f = [2x, 2y]^\top$: two times the position vector. So the gradient always points away from the centre (uphill), and gets longer the farther out you go (steeper). At the centre it is zero: the bottom.
The saddle $f = x^2 - y^2$ has $\nabla f = [2x, -2y]^\top$. At $(1, 1)$: $[2, -2]^\top$, which points right and down: uphill means "more $x$, less $y$". At $(0, 0)$ it is zero, yet the origin is not a minimum or maximum.
Using the gradient to predict. For the bowl at $(1, 2)$: $f = 5$ and $\nabla f = [2, 4]^\top$. Take a small step $\mathbf{h} = [0.1, 0.1]^\top$.
- Prediction: $f \approx 5 + \nabla f\cdot\mathbf{h} = 5 + 2\cdot0.1 + 4\cdot0.1 = 5.6$.
- Truth: $f(1.1, 2.1) = 1.21 + 4.41 = 5.62$.
- The error is $0.02$, tiny compared with the step. The smaller the step, the better the prediction.
If $f$ is smooth near $\mathbf{a}$, then for a small step $\mathbf{h}$
$$f(\mathbf{a} + \mathbf{h}) \approx f(\mathbf{a}) + \nabla f(\mathbf{a})\cdot\mathbf{h}.$$The right side is the tangent plane (or linear approximation) at $\mathbf{a}$: $z = f(\mathbf{a}) + \nabla f(\mathbf{a})\cdot(\mathbf{x} - \mathbf{a})$. The gradient is what tilts it. This gives the three facts:
- $\nabla f(\mathbf{a})$ points in the direction of fastest increase.
- $\|\nabla f(\mathbf{a})\|$ is the steepness (the largest slope in any direction).
- $\nabla f(\mathbf{a}) = \mathbf{0}$ marks a critical point (also stationary point): a minimum, a maximum or a saddle. Telling them apart needs second derivatives (Chapter 2.10).
(The proof that it points to fastest increase comes two sections later, with the directional derivative.)
Why do we need it?
To decide what to do next on a landscape you cannot see all at once. Local information (the tilt under your feet) is enough to choose a good direction.
Where is it used?
Optimisers use the direction to move and the length to judge progress; training stops when the gradient is nearly zero. Linearisation (Chapter 2.12) and Newton's method build on the tangent plane.
How is it used?
Compute $\nabla f$ at your point. Walk along it to go up, against it to go down. If its length is near zero you have reached a flat spot, so check whether it is a minimum, a maximum or a saddle.
Local only. The gradient describes the ground right under your feet. Far away the surface may bend, so the straight-uphill direction can change (the prediction above only works for small steps).
Zero gradient does not mean minimum. Peaks, valleys and saddles all have a zero gradient. In a loss landscape we hope for a valley, but we can get stuck at a saddle.
Quick check: $f(x, y) = x^2 + 3y^2$ at $(2, -1)$. Which way is uphill, and how steep is it?
$\nabla f = [2x, 6y]^\top = [4, -6]^\top$. Uphill is right and down (more $x$, less $y$). Steepness $= \sqrt{16 + 36} = \sqrt{52} \approx 7.21$.
Directional derivatives core
The east-slope and the north-slope are just two special directions. What if you walk north-east, or in any direction you like? The slope you feel in a chosen direction is the directional derivative.
Here is the key idea: the gradient already contains the answer for every direction. To get the slope along a direction, take the dot product of the gradient with that direction. (A dot product multiplies matching entries and adds them up. It is big when two arrows point the same way.)
Why a dot product? If the direction mostly agrees with the uphill arrow, you climb fast. If it is at a right angle to it, you walk along a level path and climb nothing. If it points the opposite way, you go downhill.
Take $f = x^2y + 2y$ at $(3, 2)$, where $\nabla f = [12, 11]^\top$. Walk in the direction $\mathbf{u} = [0.6, 0.8]^\top$ (a unit vector: $0.36 + 0.64 = 1$).
- Dot product: $D_{\mathbf{u}} f = 12\cdot0.6 + 11\cdot0.8 = 7.2 + 8.8 = 16$.
- So walking 1 unit that way raises $f$ by about 16 (for a tiny step, 16 times the step).
Where does the dot product come from? Walk along the line $(3 + 0.6t,\ 2 + 0.8t)$. The height is $g(t) = (3 + 0.6t)^2(2 + 0.8t) + 2(2 + 0.8t)$, an ordinary function of $t$. Its slope at $t = 0$ is, by the product rule,
$$g'(0) = \underbrace{2\cdot3\cdot0.6}_{3.6}\cdot 2 + 9\cdot0.8 + 2\cdot0.8 = 7.2 + 7.2 + 1.6 = 16 \ ✓.$$Compare: $12\cdot0.6 = 7.2$ is the $x$-part and $11\cdot0.8 = 8.8$ the $y$-part. The same $16$.
Special cases. $\mathbf{u} = [1, 0]^\top$ gives $12\cdot1 + 11\cdot0 = 12 = f_x$. $\mathbf{u} = [0, 1]^\top$ gives $f_y = 11$. The partial derivatives are directional derivatives along the axes.
A direction that is not a unit vector. To walk toward $[1, 1]$, first make it length 1: $\mathbf{u} = [1, 1]/\sqrt2 \approx [0.707, 0.707]$. Then $D_{\mathbf{u}}f = (12 + 11)/\sqrt2 = 23/1.4142 \approx 16.26$.
The directional derivative of $f$ at $\mathbf{a}$ in the direction of a unit vector $\mathbf{u}$ is the slope along that line:
$$D_{\mathbf{u}} f(\mathbf{a}) = \lim_{h\to0}\frac{f(\mathbf{a} + h\mathbf{u}) - f(\mathbf{a})}{h} = \nabla f(\mathbf{a})\cdot\mathbf{u}.$$Why the formula holds. Put $g(t) = f(\mathbf{a} + t\mathbf{u}) = f(a_1 + tu_1,\ a_2 + tu_2)$. Each input changes at its own rate ($u_1$ and $u_2$ per unit of $t$), and each change moves $f$ by its partial derivative times that rate. So $g'(0) = f_x u_1 + f_y u_2 = \nabla f\cdot\mathbf{u}$. (This is the chain rule for several variables; you will study it properly in Chapter 2.8.)
Because $\mathbf{u}$ must have length 1, the answer is "slope per unit of distance walked". Without that rule, a longer vector would give a bigger number just from taking a longer step.
Why do we need it?
Real moves are rarely along an axis. We need the rate of change for any direction of travel, and a way to find the best one.
Where is it used?
Line search in optimisation (how the loss changes along the update direction), sensitivity of a model to a change in a chosen mix of features, and gradient checking (nudge the weights in a random direction and compare).
How is it used?
Normalise your direction to length 1, compute the gradient, and take their dot product. The sign tells you uphill (positive), downhill (negative) or level (zero).
Use a unit vector. $\mathbf{u}$ must have length exactly 1. If you are given the direction $[3, 4]^\top$, divide by its length $5$ first: $\mathbf{u} = [0.6, 0.8]^\top$.
A directional derivative is a number, not a vector. It is one slope. The gradient is the vector that holds the slopes for all directions at once.
Quick check: $\nabla f = [3, 4]^\top$ at a point. Find the slope in the direction $[0, 1]^\top$ and in the direction $[4, -3]^\top$.
For $[0, 1]^\top$ (already unit): $3\cdot0 + 4\cdot1 = 4$. For $[4, -3]^\top$, its length is $5$, so $\mathbf{u} = [0.8, -0.6]^\top$ and the slope is $3\cdot0.8 - 4\cdot0.6 = 2.4 - 2.4 = 0$. That direction is perpendicular to the gradient: a level path.
Steepest ascent and steepest descent core
You are on a foggy hill and can only feel the ground at your feet. You want to climb as fast as possible. Which way do you step? Straight along the gradient. Every other direction wastes some of your step going sideways.
Want to go down as fast as possible, like a ball rolling to the valley? Step the exact opposite way: against the gradient.
And if you step at a right angle to the gradient, you walk along a level path: you neither climb nor descend.
This one fact is the whole idea behind gradient descent, the method that trains almost every model: look at the gradient, take a small step the other way, repeat.
Take $f = x^2y + 2y$ at $(3, 2)$ with $\nabla f = [12, 11]^\top$.
- Steepness: $\|\nabla f\| = \sqrt{12^2 + 11^2} = \sqrt{265} \approx 16.28$.
- Steepest ascent direction: divide by the length, $\mathbf{u}_\text{up} = [12, 11]/16.28 \approx [0.737, 0.676]$. Slope there: $+16.28$.
- Steepest descent direction: $\mathbf{u}_\text{down} \approx [-0.737, -0.676]$. Slope there: $-16.28$.
- A level direction, perpendicular to the gradient: $[-11, 12]/16.28 \approx [-0.676, 0.737]$. Slope: $\bigl(12\cdot(-11) + 11\cdot12\bigr)/16.28 = 0$.
- Our earlier direction $[0.6, 0.8]$ gave slope 16, which is $16/16.28 \approx 98\%$ of the best. It is only about $10.6^\circ$ away from the gradient.
One gradient-descent step on the bowl $f = x^2 + y^2$ from $(3, 4)$: $\nabla f = [6, 8]^\top$ and $f = 25$. With step size $\eta = 0.1$, the new point is $[3, 4]^\top - 0.1\,[6, 8]^\top = [2.4, 3.2]^\top$, where $f = 5.76 + 10.24 = 16$. The loss dropped from $25$ to $16$.
Why the gradient is the steepest direction. For a unit vector $\mathbf{u}$ at angle $\theta$ to $\nabla f$, the dot product rule $\mathbf{a}\cdot\mathbf{b} = \|\mathbf{a}\|\|\mathbf{b}\|\cos\theta$ (see the dot product) gives
$$D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u} = \|\nabla f\|\,\underbrace{\|\mathbf{u}\|}_{=1}\cos\theta = \|\nabla f\|\cos\theta.$$The cosine is at most $1$ (at $\theta = 0$, same direction), at least $-1$ (at $\theta = 180^\circ$) and $0$ at $90^\circ$. Therefore:
- Steepest ascent: $\mathbf{u} = \dfrac{\nabla f}{\|\nabla f\|}$, with the largest possible slope $+\|\nabla f\|$.
- Steepest descent: $\mathbf{u} = -\dfrac{\nabla f}{\|\nabla f\|}$, with slope $-\|\nabla f\|$.
- No change: any $\mathbf{u}\perp\nabla f$ has slope $0$.
This gives the gradient descent update, with a small positive number $\eta$ called the learning rate (step size):
$$\mathbf{x}_{\text{new}} = \mathbf{x} - \eta\,\nabla f(\mathbf{x}).$$If $\nabla f = \mathbf{0}$ there is no direction to prefer, and the method stops.
Why do we need it?
Models have too many weights to try combinations at random. The gradient gives, for free, the single best direction to lower the error, so training becomes "repeat one small downhill step".
Where is it used?
Training linear and logistic regression, neural networks, matrix factorisation and embeddings. Gradient ascent is the same idea upside down: maximising a reward or a likelihood.
How is it used?
Start at some weights. Compute the gradient of the loss. Subtract learning rate times gradient. Repeat until the gradient is almost zero or the loss stops improving. If the loss gets worse or explodes, shrink the learning rate.
Steepest locally, not the shortest way to the bottom. The gradient is the best direction for a tiny step. On a long, narrow valley it points mostly toward the valley wall, so the path zig-zags. And a step that is too big overshoots.
You can reach different minima. On a loss with several valleys, where you end up depends on where you start.
Descent needs the minus sign. $+\eta\nabla f$ climbs; $-\eta\nabla f$ descends.
Quick check: at a point $\nabla f = [3, 4]^\top$. What is the steepest descent direction and its slope?
$\|\nabla f\| = 5$. Direction $= -[3, 4]/5 = [-0.6, -0.8]$, slope $= -5$.
Level sets core
On a hiking map, a thin line labelled "100 m" joins every place that is exactly 100 metres above the sea. Walk along that line and you never go up or down. This is a contour line, and mathematicians call it a level set: all the points where the function has one chosen value.
Picture slicing the surface with a perfectly horizontal sheet at height $c$. Where the sheet cuts the surface you get a curve. Drop that curve straight down to the floor and you get the level set.
Do this for several heights and you get a set of nested curves: a map of the whole surface drawn on flat paper.
The bowl $f = x^2 + y^2$. The level set at height $c$ is $x^2 + y^2 = c$:
- $c = 0$: only the single point $(0, 0)$.
- $c = 1$: a circle of radius $1$. $c = 4$: radius $2$. $c = 9$: radius $3$. (In general radius $\sqrt{c}$.)
- $c = -2$: nothing at all, since $x^2 + y^2$ is never negative. The level set is empty.
The saddle $f = x^2 - y^2$: $c = 1$ gives $x^2 - y^2 = 1$, two curved branches opening left and right (a hyperbola). $c = -1$ gives two branches opening up and down. $c = 0$ gives $x^2 = y^2$, i.e. the two crossing lines $y = x$ and $y = -x$.
The product $f = xy$: $c = 2$ gives $y = 2/x$, a hyperbola in the first and third quadrants.
The level set (or level curve when $n=2$) of $f$ at level $c$ is
$$L_c = \{\,\mathbf{x}\in\mathbb{R}^n \;:\; f(\mathbf{x}) = c\,\}.$$For two inputs, $L_c$ is usually a curve in the floor plane. For three inputs it is a surface (a "level surface", for example a sphere for $f = x^2 + y^2 + z^2$). For $n$ inputs it usually has $n-1$ dimensions. Different levels never cross each other, because one point has only one height. (A level set can be a single point, empty, or have a crossing such as at the saddle level $c = 0$, but two different levels never meet.)
Why do we need it?
A surface is hard to draw and hides what is behind it. Level sets turn a 3D landscape into a flat map you can read, and they are the right tool for a function of three or more inputs, which cannot be drawn as a surface at all.
Where is it used?
Contour plots of loss functions in every optimisation tutorial, weather maps (lines of equal pressure), decision boundaries of classifiers (the level set where the score is 0.5), and constrained optimisation (Lagrange multipliers, regularisation balls).
How is it used?
Choose a few levels, find the curve where $f$ equals each, and draw them. Closely spaced curves mean steep ground; closed loops mean a peak or a valley. Equal-loss curves tell you which weights are "equally good".
A level set lives on the floor (in the input space), not up on the surface. The curve on the surface is its "lifted copy"; the level set itself is the shadow on the floor.
Level sets are not the graph. The graph is the surface (inputs plus output). A level set is only the inputs that give one particular output.
Quick check: describe the level set $x^2 + 4y^2 = 4$ of $f = x^2 + 4y^2$ at $c = 4$.
Divide by 4: $\dfrac{x^2}{4} + y^2 = 1$. It is an ellipse that crosses the $x$-axis at $\pm2$ and the $y$-axis at $\pm1$ (wider than tall, because $y$ has the bigger coefficient, so $y$ gets "expensive" faster).
Contour interpretation: reading the map core
A contour map is a flat picture of a landscape. Four rules let you read it, just like a hiker:
- Lines close together = steep. The height changes a lot over a short distance.
- Lines far apart = gentle (nearly flat).
- Closed loops circle a peak or a valley. Look at the labels to tell which.
- The uphill direction is always at a right angle to the contour line. The fastest way up (or down) crosses the lines head-on. Walking along a line is level.
The last rule is the key link to this chapter: the gradient is perpendicular to the level sets.
Perpendicular, with numbers. $f = x^2 + y^2$ has level set $x^2 + y^2 = 5$ through the point $(1, 2)$ (a circle of radius $\sqrt5$).
- The gradient at $(1, 2)$ is $[2, 4]^\top$.
- The circle's tangent direction at $(1, 2)$ is perpendicular to the radius $[1, 2]$, for example $\mathbf{t} = [-2, 1]$.
- Dot product: $\nabla f\cdot\mathbf{t} = 2\cdot(-2) + 4\cdot1 = -4 + 4 = 0$ ✓.
Close lines mean steep. Draw contours of $f = x^2 + y^2$ at heights $c = 1, 2, 3, 4$. They are circles of radius $\sqrt c$: $1,\ 1.414,\ 1.732,\ 2$. The gaps between them are $0.414,\ 0.318,\ 0.268$: they get smaller as you go out. The gradient length is $2r$, so it grows outward: the bowl gets steeper. In fact, the gap between contours that differ in height by $\Delta c$ is about $\Delta c/\|\nabla f\|$. At $r \approx 1.2$: $1/(2\cdot1.2) \approx 0.42$ ✓.
Theorem. At any point where $\nabla f \ne \mathbf{0}$, the gradient is perpendicular to the level set through that point.
Why. Walk along the level curve with a path $\mathbf{r}(t)$. The height never changes: $f(\mathbf{r}(t)) = c$ for all $t$. Differentiate both sides with respect to $t$. The right side is a constant, so its derivative is $0$. The left side, by the chain rule for several variables (Chapter 2.8), is $\nabla f\cdot\mathbf{r}'(t)$. So
$$\nabla f(\mathbf{r}(t))\cdot\mathbf{r}'(t) = 0.$$The velocity $\mathbf{r}'(t)$ is the tangent to the level curve, and its dot product with the gradient is zero: they are perpendicular. $\blacksquare$
Summary of the contour dictionary. (1) $\nabla f\perp$ contours. (2) $\nabla f$ points toward the higher-labelled contour. (3) The spacing of contours is about $\Delta c/\|\nabla f\|$: dense means steep. (4) A closed loop surrounds an extreme point (a peak or a valley); contour lines that cross at a point mark a saddle.
Why do we need it?
The map picture lets you judge a loss landscape at a glance: where it is steep, where it is flat, where the valleys are, and which way an optimiser will move, without computing anything.
Where is it used?
Every picture of gradient descent in a textbook, the explanation of why plain gradient descent zig-zags in narrow valleys (which motivates momentum and Adam), and the geometry of regularised regression and constrained optimisation.
How is it used?
To read a map: find the closest and farthest line spacing (steep versus flat), read the labels to see which side is uphill, then draw the gradient perpendicular to the lines at the point you care about. To plan a step, move across the lines, not along them.
Equal height steps. "Close lines mean steep" is only true when the lines are drawn at equal steps in height (every 100 m, say). Always check the labels.
Perpendicular needs equal axes. If a plot stretches one axis, the gradient will no longer look perpendicular to the contours (it still is, in the true coordinates).
Steep along the arrow is not steep everywhere. A long, thin valley is steep across and nearly flat along. The gradient mostly points across, which is why gradient descent zig-zags there.
Quick check: on a map the 10 m line and the 20 m line are 5 m apart at one place and 50 m apart at another. Where is it steeper, and roughly how steep is it?
Steeper at the first place. The slope is about (height change) / (horizontal gap) $= 10/5 = 2$ there, and $10/50 = 0.2$ at the second place. The gradient is 10 times larger at the first place.
Recap, cheat sheet and practice
- A function of several inputs $f(x, y)$ is a landscape: position in, height out. Its picture is a surface, or a height map seen from above.
- A partial derivative $\partial f/\partial x$ is the slope when only $x$ moves: freeze the other inputs, then differentiate as usual. For small nudges, the effects of $x$ and $y$ simply add.
- The gradient $\nabla f$ is the column of all partial derivatives. It points uphill, its length is the steepness, and it is zero at flat spots (peaks, valleys, saddles).
- The directional derivative is the slope along a unit direction $\mathbf{u}$: $D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u} = \|\nabla f\|\cos\theta$.
- So the steepest ascent is along $\nabla f$ (slope $\|\nabla f\|$), the steepest descent is along $-\nabla f$, and directions perpendicular to $\nabla f$ stay level. Gradient descent: $\mathbf{x}\leftarrow\mathbf{x} - \eta\nabla f(\mathbf{x})$.
- A level set $\{f = c\}$ is a contour line. The gradient is perpendicular to the level sets. Close lines mean steep ground; closed loops circle peaks or valleys.
Cheat sheet
| Idea | Formula | Picture |
|---|---|---|
| Partial derivative | $\dfrac{\partial f}{\partial x} = \lim_{h\to0}\dfrac{f(x+h, y) - f(x, y)}{h}$ | slope of the cut $y = $ const |
| Gradient (column) | $\nabla f = \bigl[\tfrac{\partial f}{\partial x_1}, \dots, \tfrac{\partial f}{\partial x_n}\bigr]^\top$ | arrow pointing uphill |
| Linear approximation | $f(\mathbf{a}+\mathbf{h}) \approx f(\mathbf{a}) + \nabla f\cdot\mathbf{h}$ | tangent plane |
| Directional derivative | $D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u}$ ($\|\mathbf{u}\| = 1$) | slope along $\mathbf{u}$ |
| Steepest ascent / descent | $\pm\nabla f/\|\nabla f\|$, slope $\pm\|\nabla f\|$ | straight up / down the hill |
| No change | $\mathbf{u}\perp\nabla f$ | walk along the contour |
| Gradient descent step | $\mathbf{x}\leftarrow\mathbf{x} - \eta\nabla f(\mathbf{x})$ | small step downhill |
| Level set | $L_c = \{\mathbf{x} : f(\mathbf{x}) = c\}$ | contour line |
| Critical point | $\nabla f = \mathbf{0}$ | flat: min, max or saddle |
import numpy as np
def f(p): # f(x, y) = x^2 * y + 2*y
x, y = p
return x**2 * y + 2*y
def grad(p): # analytic gradient: [2xy, x^2 + 2]
x, y = p
return np.array([2*x*y, x**2 + 2])
def num_grad(f, p, h=1e-6): # nudge ONE input at a time (central difference)
g = np.zeros_like(p, dtype=float)
for i in range(len(p)):
e = np.zeros_like(g); e[i] = h
g[i] = (f(p + e) - f(p - e)) / (2*h)
return g
p = np.array([3.0, 2.0])
print(f(p)) # 22.0
print(grad(p)) # [12. 11.]
print(num_grad(f, p)) # [12. 11.] (agrees with the formula to about 6 decimals)
u = np.array([0.6, 0.8]) # a unit direction (0.36 + 0.64 = 1)
print(grad(p) @ u) # 16.0 directional derivative = gradient . u
h = 1e-6
print((f(p + h*u) - f(p)) / h) # about 16.0000036 the same thing, found by nudging
g = grad(p)
print(np.linalg.norm(g)) # 16.2788... the steepest possible slope
print(g / np.linalg.norm(g)) # [0.7372 0.6757] steepest-ascent direction
# gradient descent on the bowl x^2 + y^2 (gradient = 2x, 2y), step size 0.1
q = np.array([3.0, 4.0]); eta = 0.1
for k in range(3):
q = q - eta * 2*q
print(k + 1, q, q @ q) # 1 [2.4 3.2] 16.0 then 2 [1.92 2.56] 10.24 then 3 [1.536 2.048] 6.5536
# the gradient is perpendicular to the contour: circle x^2 + y^2 = 5 at (1, 2)
gc = np.array([2*1, 2*2]); t = np.array([-2, 1])
print(gc @ t) # 0
1. For $f(x, y) = x^3 y^2$, what is $\partial f/\partial x$ at $(1, 2)$?
2. What is $\nabla f$ at $(2, 5)$ for $f(x, y) = x^2 + 3y$?
3. $\nabla f = [2, -1]^\top$ at a point. What is the slope in the direction $\mathbf{u} = [0.6, 0.8]^\top$?
4. If $\nabla f = [0, -2]^\top$, which direction is steepest descent?
5. On a contour map drawn at equal height steps, where the lines are packed very close together, the ground is…
6. For $f = x^2 + y^2$, one gradient-descent step from $(3, 4)$ with $\eta = 0.5$ lands at…
Practice problems
A. Find $\nabla f$ for $f(x, y) = x^2y^3$ and evaluate it at $(2, 1)$.
$f_x = 2xy^3$ (freeze $y$) and $f_y = 3x^2y^2$ (freeze $x$). At $(2, 1)$: $f_x = 2\cdot2\cdot1 = 4$ and $f_y = 3\cdot4\cdot1 = 12$. So $\nabla f(2, 1) = [4, 12]^\top$.
B. Find $\nabla f$ for $f(x, y) = e^{xy}$ at $(0, 2)$.
Freeze $y$ and use the chain rule: $f_x = y\,e^{xy}$. Freeze $x$: $f_y = x\,e^{xy}$. At $(0, 2)$: $e^0 = 1$, so $f_x = 2\cdot1 = 2$ and $f_y = 0\cdot1 = 0$. The gradient is $[2, 0]^\top$.
C. Find the directional derivative of $f = x^2 + xy$ at $(1, 2)$ toward the vector $[3, 4]$.
$\nabla f = [2x + y,\ x]^\top = [4, 1]^\top$ at $(1, 2)$. Make the direction a unit vector: $\|[3, 4]\| = 5$, so $\mathbf{u} = [0.6, 0.8]$. Then $D_{\mathbf{u}}f = 4\cdot0.6 + 1\cdot0.8 = 2.4 + 0.8 = 3.2$.
D. For $f = xy$ at $(2, 3)$, find the direction of fastest increase and the largest slope. What is the direction of fastest decrease?
$\nabla f = [y, x]^\top = [3, 2]^\top$. Its length is $\sqrt{9 + 4} = \sqrt{13} \approx 3.606$. Fastest increase: $[3, 2]/\sqrt{13} \approx [0.832, 0.555]$, with slope $\approx 3.606$. Fastest decrease: $[-0.832, -0.555]$, with slope $-3.606$.
E. Show that the gradient of $f(x, y) = x - 2y$ is perpendicular to its level sets.
$\nabla f = [1, -2]^\top$ everywhere. The level set $x - 2y = c$ is a straight line $y = (x - c)/2$ with direction vector $\mathbf{t} = [2, 1]$ (go 2 across, 1 up). Then $\nabla f\cdot\mathbf{t} = 1\cdot2 + (-2)\cdot1 = 0$. So the gradient is perpendicular to every level line. (Contours of a tilted plane are parallel straight lines, evenly spaced: the slope is the same everywhere.)
F. Do two steps of gradient descent on $f = x^2 + 4y^2$ from $(2, 1)$ with $\eta = 0.1$. Give the loss after each.
$\nabla f = [2x, 8y]^\top$. At $(2, 1)$: $f = 4 + 4 = 8$ and $\nabla f = [4, 8]$. Step 1: $[2, 1] - 0.1[4, 8] = [1.6, 0.2]$, with $f = 2.56 + 0.16 = 2.72$. Then $\nabla f = [3.2, 1.6]$. Step 2: $[1.6, 0.2] - 0.1[3.2, 1.6] = [1.28, 0.04]$, with $f = 1.6384 + 0.0064 = 1.6448$. The loss went $8 \to 2.72 \to 1.6448$. Notice that $y$ shrinks much faster than $x$: the gradient is bigger in the steep $y$-direction. That is the zig-zag behaviour of narrow valleys.
Gradients of Vector-Valued Functions
A neural-network layer takes a vector in and gives a vector out. How does the output react when you nudge the input? The answer is a table of slopes called the Jacobian matrix. It is the single most important object you will carry into backpropagation.
- Tell scalar-valued and vector-valued functions apart, and meet curves, maps and neural layers
- Build the Jacobian matrix: shape $m\times n$, entry $(i, j) = \partial F_i/\partial x_j$
- Know the shape of every derivative: vector by vector, scalar by vector (the gradient), vector by scalar (velocity)
- See the Jacobian as the local linear map, and its determinant as the local area scale
- Compute Jacobians of a linear map, polar coordinates, a dense layer and softmax
- Apply the chain rule with Jacobians (a product of matrices) and check the shapes
Why this chapter matters. Everything in a neural network is a chain of vector functions (layer after layer). The Jacobian tells you how each layer passes a nudge along, and multiplying Jacobians is exactly what backpropagation does (Chapter 2.9). Take your time here.
Scalar-valued functions: many numbers in, one number out core
A scalar is just a single plain number. A scalar-valued function is a function whose answer is one number, no matter how many numbers go in.
Examples: the temperature at a spot on a map (position in, one temperature out); the price of a house from its size, age and distance to the city; and, most important for us, the loss of a model, which turns all of the weights into one number: "how wrong am I?".
You already know the derivative of such a function: it is the gradient from Chapter 2.4, one slope for each input. This chapter extends the idea to functions whose answer is a list.
$f(x_1, x_2) = x_1^2 + x_2^2$ takes a vector in $\mathbb{R}^2$ and returns a number.
- Input $\mathbf{x} = [3, 4]^\top$ (two numbers).
- Output $f(\mathbf{x}) = 9 + 16 = 25$ (one number).
- Derivative: $\nabla f = [2x_1, 2x_2]^\top = [6, 8]^\top$ (two numbers, one per input).
Shapes: in $n = 2$ numbers, out $m = 1$ number, derivative $n = 2$ numbers.
A scalar-valued function of $n$ variables is a function
$$f:\ \mathbb{R}^n\to\mathbb{R}, \qquad \mathbf{x} = [x_1, \dots, x_n]^\top \mapsto f(\mathbf{x}) \in \mathbb{R}.$$Its derivative is the gradient $\nabla f(\mathbf{x})\in\mathbb{R}^n$: a column vector of the $n$ partial derivatives $\partial f/\partial x_j$. In this chapter we will see it is the special case $m = 1$ of the Jacobian.
Why do we need it?
To measure "how good or bad" with one number, we need a function that squeezes a lot of information down to a single score. Only a single score can be minimised.
Where is it used?
Every loss function: mean squared error, cross-entropy, a regularised loss. Also a model's single output (a house price), a likelihood, and a reward in reinforcement learning.
How is it used?
Compute the score for the current weights, then compute its gradient (one slope per weight) and nudge the weights against it. The output being a single number is what makes "lower is better" well-defined.
"Scalar" refers to the output. A function of a million variables is still scalar-valued if it returns one number. The input can be a vector; the output is what we call scalar or vector.
Quick check: is $f(x, y, z) = xy + z$ scalar-valued? What are the sizes of its input, output and gradient?
Yes. The input is 3 numbers, the output is 1 number, and the gradient $[y, x, 1]^\top$ has 3 numbers.
Vector-valued functions: a list in, a list out core
Now let the answer be a list. A GPS maps the time $t$ to your position: two numbers (latitude and longitude). A paint mixer maps three dials to the three colours (red, green, blue) it produces. A neural-network layer maps its input vector to an output vector.
The easy way to think about it: a vector-valued function is several scalar-valued functions stacked in a column, all sharing the same inputs. Each output number has its own "recipe" that uses all the inputs.
Let $\mathbf{F}:\mathbb{R}^2\to\mathbb{R}^3$ be
$$\mathbf{F}(x_1, x_2) = \begin{bmatrix} F_1 \\ F_2 \\ F_3 \end{bmatrix} = \begin{bmatrix} x_1x_2 \\ x_1 + x_2 \\ x_1^2 \end{bmatrix}.$$- Input $[2, 3]^\top$ (2 numbers).
- $F_1 = 2\cdot3 = 6$, $F_2 = 2 + 3 = 5$, $F_3 = 2^2 = 4$.
- Output $\mathbf{F}(2, 3) = [6, 5, 4]^\top$ (3 numbers).
Notice that $F_3$ does not use $x_2$ at all. We will see that fact show up as a zero in the Jacobian.
A vector-valued function is a function $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$,
$$\mathbf{F}(\mathbf{x}) = \begin{bmatrix} F_1(\mathbf{x}) \\ F_2(\mathbf{x}) \\ \vdots \\ F_m(\mathbf{x}) \end{bmatrix},$$where each component function $F_i:\mathbb{R}^n\to\mathbb{R}$ is scalar-valued. The case $m = 1$ is the scalar-valued case from the last section. Capital bold $\mathbf{F}$ reminds us that the output is a vector. (A linear map $\mathbf{x}\mapsto A\mathbf{x}$ is the simplest example.)
Why do we need it?
Most things we model have several outputs, and most models are built from stages that each turn a vector into another vector. We need to be able to talk about, and differentiate, such functions.
Where is it used?
Every neural-network layer, a softmax that outputs one probability per class, word embeddings, the motion of a robot arm (joint angles in, hand position out), and coordinate changes (polar to Cartesian).
How is it used?
Evaluate it component by component. To differentiate it, differentiate each component with respect to each input. That table of derivatives is the Jacobian (next sections).
Count the numbers. In $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$, $n$ is the number of inputs and $m$ the number of outputs. The two numbers do not have to match. Getting these two straight now saves a lot of confusion with matrix shapes later.
Quick check: $\mathbf{F}(x, y) = [x + y,\ x - y]^\top$. What is $\mathbf{F}(5, 2)$, and which $\mathbb{R}^n\to\mathbb{R}^m$ is it?
$\mathbf{F}(5, 2) = [7, 3]^\top$. It takes 2 numbers and returns 2 numbers, so $\mathbb{R}^2\to\mathbb{R}^2$.
Vector functions you will meet: curves, maps and layers
Vector-valued functions come in three everyday flavours. Each answers a different question:
- A curve $\mathbf{r}:\mathbb{R}\to\mathbb{R}^m$. One number in (time $t$), a position out. Question: where am I at time $t$? The picture is a path.
- A map of the plane $\mathbf{F}:\mathbb{R}^2\to\mathbb{R}^2$. A point in, a point out. Question: where does this point land? The picture is the plane being stretched, rotated or bent, like a rubber sheet.
- A neural layer $\mathbf{x}\mapsto\sigma(W\mathbf{x} + \mathbf{b})$. A vector in, a vector out. Question: what features does the layer compute?
- Curve. $\mathbf{r}(t) = [\cos t,\ \sin t]^\top$. At $t = 0$: $[1, 0]$. At $t = \pi/2$: $[0, 1]$. As $t$ grows it goes round the unit circle.
- Map. Polar to Cartesian, $\mathbf{F}(r, \theta) = [r\cos\theta,\ r\sin\theta]^\top$. The input $(2, \pi/6)$ ("2 steps at 30°") lands at $[2\cdot0.866,\ 2\cdot0.5] = [1.732,\ 1]$.
- Layer. $W = \begin{bmatrix} 1 & -1 \\ 2 & 1 \end{bmatrix}$, $\mathbf{b} = [0.5, -1]^\top$, ReLU activation. For $\mathbf{x} = [1, 2]^\top$: $W\mathbf{x} = [1 - 2,\ 2 + 2] = [-1, 4]$, add $\mathbf{b}$ to get $\mathbf{z} = [-0.5, 3]$, then ReLU (replace negatives by 0): $\mathbf{a} = [0, 3]$.
All three are vector-valued functions $\mathbb{R}^n\to\mathbb{R}^m$, with different sizes:
| Kind | $n$ (inputs) | $m$ (outputs) | Typical name |
|---|---|---|---|
| Curve | 1 | 2 or 3 | $\mathbf{r}(t)$ |
| Map of the plane | 2 | 2 | $\mathbf{F}(x, y)$ |
| Neural layer | any $n$ | any $m$ | $\mathbf{a} = \sigma(W\mathbf{x} + \mathbf{b})$ |
A vector function is simply any function whose output is a vector. We can compose them (feed the output of one into the next), and a neural network is exactly a long composition of layers.
Why do we need it?
These three shapes cover almost everything: following something through time, changing coordinates, and stacking layers of a network. Seeing the common pattern lets one tool (the Jacobian) handle all of them.
Where is it used?
Curves: trajectories in physics and robotics, training paths of the weights. Maps: polar and spherical coordinates, normalising flows and image warping. Layers: every neural network.
How is it used?
Name the sizes $n$ and $m$ first. Then evaluate component by component. Later, differentiate the same way, and the Jacobian will have shape $m\times n$.
A function is the whole machine, not one output. A map of the plane has two output numbers at each point; to draw it we need two pictures (input plane and output plane), or a deformed grid as above.
Quick check: what are $n$ and $m$ for a layer that takes a 784-pixel image and returns 10 class scores?
$n = 784$ inputs and $m = 10$ outputs, so the layer is a function $\mathbb{R}^{784}\to\mathbb{R}^{10}$.
The Jacobian matrix core
A vector function with $n$ inputs and $m$ outputs has $m\times n$ slopes. For each output and each input you can ask: "if I nudge this input, how fast does this output move?" Write all those slopes in a table. That table is the Jacobian matrix.
There is a second, geometric way to see it. Zoom in on a point of a bent map (like the polar grid) until it looks straight. Zoomed in, the map looks like a linear map, a matrix that stretches and turns the tiny neighbourhood. That matrix is the Jacobian. It is the "best straight-line copy" of the function, just as the derivative was the slope of the best straight line in one variable.
How to read the table. Each row belongs to one output (it is the gradient of that output, written sideways). Each column belongs to one input (it says how the whole output vector reacts when that one input is nudged).
Example 1. $\mathbf{F}(x, y) = [x^2y,\ 5x + y^2]^\top$, with $n = 2$ inputs and $m = 2$ outputs, so the Jacobian is $2\times2$.
- Row 1 (output $F_1 = x^2y$): $\partial F_1/\partial x = 2xy$, $\partial F_1/\partial y = x^2$.
- Row 2 (output $F_2 = 5x + y^2$): $\partial F_2/\partial x = 5$, $\partial F_2/\partial y = 2y$.
- So $J = \begin{bmatrix} 2xy & x^2 \\ 5 & 2y \end{bmatrix}$. At $(1, 2)$: $J = \begin{bmatrix} 4 & 1 \\ 5 & 4 \end{bmatrix}$.
Check by nudging. $\mathbf{F}(1, 2) = [2, 9]^\top$. Nudge the input by $\mathbf{h} = [0.01, -0.02]^\top$. The Jacobian predicts the change $J\mathbf{h} = [4\cdot0.01 + 1\cdot(-0.02),\ 5\cdot0.01 + 4\cdot(-0.02)] = [0.02, -0.03]$, so $\mathbf{F}\approx[2.02, 8.97]$. The true value is $\mathbf{F}(1.01, 1.98) = [1.0201\cdot1.98,\ 5.05 + 3.9204] = [2.019798,\ 8.9704]$ ✓.
Example 2: polar to Cartesian. $\mathbf{F}(r, \theta) = [r\cos\theta,\ r\sin\theta]^\top$. Differentiate each component:
$$J = \begin{bmatrix} \partial x/\partial r & \partial x/\partial\theta \\ \partial y/\partial r & \partial y/\partial\theta \end{bmatrix} = \begin{bmatrix} \cos\theta & -r\sin\theta \\ \sin\theta & r\cos\theta \end{bmatrix}, \qquad \det J = r\cos^2\theta + r\sin^2\theta = r.$$At $(r, \theta) = (2, \pi/6)$: $J = \begin{bmatrix} 0.866 & -1 \\ 0.5 & 1.732 \end{bmatrix}$ and $\det J = 0.866\cdot1.732 + 1\cdot0.5 = 1.5 + 0.5 = 2 = r$ ✓.
What the determinant says. A tiny rectangle of size $\Delta r\times\Delta\theta$ in the $(r, \theta)$ plane lands on a patch of area about $r\,\Delta r\,\Delta\theta$ in the $(x, y)$ plane. Far from the origin (large $r$), the same angle step sweeps out a longer arc, so the area is stretched by $r$. This is the famous "extra $r$" that appears when you integrate in polar coordinates: the area scale of any change of coordinates is $|\det J|$.
For $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$, the Jacobian matrix at $\mathbf{x}$ is the $m\times n$ matrix whose $(i, j)$ entry is the partial derivative of output $i$ with respect to input $j$:
$$J_{\mathbf{F}}(\mathbf{x}) = \frac{\partial\mathbf{F}}{\partial\mathbf{x}} = \begin{bmatrix} \dfrac{\partial F_1}{\partial x_1} & \cdots & \dfrac{\partial F_1}{\partial x_n} \\ \vdots & \ddots & \vdots \\ \dfrac{\partial F_m}{\partial x_1} & \cdots & \dfrac{\partial F_m}{\partial x_n} \end{bmatrix} \in\mathbb{R}^{m\times n}.$$- Layout convention (used everywhere in this guide). One row per output, one column per input: shape $m\times n$, entry $(i, j) = \partial F_i/\partial x_j$. This is called numerator layout. Some books use the transpose ($n\times m$, "denominator layout"); always check which one a book uses. We use the same convention as the matrix-calculus chapter of the Linear Algebra guide.
- Row $i$ is the gradient of $F_i$, written as a row: $(\nabla F_i)^\top$. Column $j$ is $\partial\mathbf{F}/\partial x_j$, the response of the whole output to nudging input $j$.
- Best linear copy: $\mathbf{F}(\mathbf{x} + \mathbf{h}) \approx \mathbf{F}(\mathbf{x}) + J_{\mathbf{F}}(\mathbf{x})\,\mathbf{h}$ for small $\mathbf{h}$. (This is multiplication of an $m\times n$ matrix by an $n$-vector, giving an $m$-vector, so the shapes fit.)
- When $m = n$ the Jacobian is square, and $\det J$ is the local area (or volume) scale factor. A negative determinant means the map flips orientation (like a mirror). $\det J = 0$ means the map squashes a patch onto a line or a point.
Why do we need it?
With many outputs and many inputs, one slope is not enough. We need the whole table of "which input moves which output, and how fast" in order to pass a small change through a function.
Where is it used?
Backpropagation multiplies the Jacobians of the layers. Also change of variables in integrals (the determinant), Newton's method for systems of equations, the extended Kalman filter, robot arms (joint speeds to hand speed), and normalising flows.
How is it used?
Differentiate every output component with respect to every input, and arrange the numbers with outputs as rows and inputs as columns. Evaluate at your point. Multiply it by a small input change to predict the output change.
Rows are outputs, columns are inputs. A common slip is to transpose it. A quick test: the Jacobian multiplies an input-sized vector to give an output-sized vector, so it must have $n$ columns and $m$ rows.
The Jacobian depends on the point. Unless the function is linear, the matrix changes as you move (see how the arrows change as you drag).
It is a local picture. The parallelogram only matches the patch for small $h$. Far away the true map bends.
Quick check: find the Jacobian of $\mathbf{F}(x, y) = [x + y,\ xy]^\top$ at $(2, 3)$.
$J = \begin{bmatrix} 1 & 1 \\ y & x \end{bmatrix}$. At $(2, 3)$: $J = \begin{bmatrix} 1 & 1 \\ 3 & 2 \end{bmatrix}$, and $\det J = 2 - 3 = -1$ (this map flips orientation near that point).
The derivative of a vector with respect to a vector core
"Derivative of a vector $\mathbf{F}$ with respect to a vector $\mathbf{x}$" is just another name for the Jacobian. The question it answers is: when the whole input vector is nudged a little, how does the whole output vector change?
Because both sides are vectors, the answer is a matrix. A good habit: always write down the shape first. Output size $m$, input size $n$, so the derivative is $m\times n$. If you can state the shape, you can often spot an error before you compute anything.
The simplest vector function is a linear one, $\mathbf{F}(\mathbf{x}) = A\mathbf{x}$. Its derivative is as nice as it gets: it is the matrix $A$ itself, everywhere. (The scalar version: the derivative of $a x$ is $a$.)
Let $A = \begin{bmatrix} 2 & 1 & 0 \\ -1 & 3 & 4 \end{bmatrix}$ ($2\times3$) and $\mathbf{F}(\mathbf{x}) = A\mathbf{x}$ with $\mathbf{x}\in\mathbb{R}^3$. So $n = 3$, $m = 2$, and the Jacobian will be $2\times3$.
- Write the outputs: $F_1 = 2x_1 + 1x_2 + 0x_3$ and $F_2 = -1x_1 + 3x_2 + 4x_3$.
- Differentiate: $\partial F_1/\partial x_1 = 2$, $\partial F_1/\partial x_2 = 1$, $\partial F_1/\partial x_3 = 0$. And $\partial F_2/\partial x_1 = -1$, $\partial F_2/\partial x_2 = 3$, $\partial F_2/\partial x_3 = 4$.
- Collect: $J = \begin{bmatrix} 2 & 1 & 0 \\ -1 & 3 & 4 \end{bmatrix} = A$.
Why in general. $F_i = \sum_j A_{ij}x_j$. Differentiating with respect to $x_j$, only the term with that $j$ survives: $\partial F_i/\partial x_j = A_{ij}$. That is exactly the $(i, j)$ entry of $A$. Adding a constant vector, $\mathbf{F}(\mathbf{x}) = A\mathbf{x} + \mathbf{b}$, changes nothing, because constants have zero derivative.
(This is why the matrix of a linear map from the Linear Algebra guide is its own derivative: a linear map already is its best linear copy.)
For $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$ the derivative of the vector $\mathbf{F}$ with respect to the vector $\mathbf{x}$ is the Jacobian,
$$\frac{\partial\mathbf{F}}{\partial\mathbf{x}} = J_{\mathbf{F}}\in\mathbb{R}^{m\times n}, \qquad \Bigl(\frac{\partial\mathbf{F}}{\partial\mathbf{x}}\Bigr)_{ij} = \frac{\partial F_i}{\partial x_j}.$$The shapes of every derivative in this chapter (numerator layout):
| Function | Name | Derivative | Shape |
|---|---|---|---|
| $f:\mathbb{R}\to\mathbb{R}$ | ordinary function | $f'(x)$ | $1\times1$ |
| $f:\mathbb{R}^n\to\mathbb{R}$ | scalar-valued | $(\nabla f)^\top$ (gradient, as a row) | $1\times n$ |
| $\mathbf{r}:\mathbb{R}\to\mathbb{R}^m$ | curve | $\mathbf{r}'(t)$ (velocity) | $m\times1$ |
| $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$ | vector-valued | $J_{\mathbf{F}}$ | $m\times n$ |
For a linear map, $\mathbf{F}(\mathbf{x}) = A\mathbf{x} + \mathbf{b}$, $J_{\mathbf{F}} = A$ at every point.
Why do we need it?
A layer of a network is a vector function. To train it we must know how its output vector responds to a change in its input vector (and weights). The Jacobian is that response, in matrix form.
Where is it used?
Linear layers (Jacobian equals the weight matrix), activation layers (diagonal Jacobian), softmax, normalisation layers, and the sensitivity of any model output to its input features.
How is it used?
State $n$ and $m$ and write the shape $m\times n$. Fill each entry $(i, j)$ with $\partial F_i/\partial x_j$. Check that the Jacobian times an input-sized vector gives an output-sized vector.
Do not confuse the Jacobian of $A\mathbf{x}$ with $A\mathbf{x}$'s derivative with respect to $A$. Here $A$ is fixed and $\mathbf{x}$ moves. Taking derivatives with respect to the matrix $A$ is a different (and bigger) object, treated in Chapter 2.6.
Shape mistakes are the number-one bug. If the Jacobian has the wrong shape, the chain rule's matrix products will not even be defined. Always write $m\times n$ first.
Quick check: $\mathbf{F}:\mathbb{R}^4\to\mathbb{R}^2$. What shape is its Jacobian, and how many slopes does it hold?
$2\times4$ (2 outputs = rows, 4 inputs = columns). It holds $2\cdot4 = 8$ slopes.
The derivative of a scalar with respect to a vector (the gradient) core
When the output is a single number, the table of slopes has only one row: one slope per input. That one row is the gradient, laid on its side.
Here is the one subtle point. There are two natural ways to write the same $n$ numbers:
- As a column ($n\times1$): the gradient $\nabla f$. It lives in the same space as the input $\mathbf{x}$, so you can add it to or subtract it from $\mathbf{x}$. This is the form for gradient descent.
- As a row ($1\times n$): the Jacobian $J_f = (\nabla f)^\top$. It is a matrix that multiplies on the left of a vector, which is what the chain rule needs.
Same numbers, different layout. In this guide: the gradient is a column, and the Jacobian of a scalar function is its transpose.
A tiny model predicts $\hat y = \mathbf{w}\cdot\mathbf{x}$ and its loss is $L(\mathbf{w}) = (\hat y - y)^2$. Take the data point $\mathbf{x} = [1, 2, -1]^\top$ with target $y = 3$, and the weights $\mathbf{w} = [1, 0, 2]^\top$.
- Prediction: $\hat y = 1\cdot1 + 0\cdot2 + 2\cdot(-1) = -1$.
- Error: $e = \hat y - y = -1 - 3 = -4$. Loss: $L = e^2 = 16$.
- Since $e = \sum_j w_jx_j - y$, we have $\partial e/\partial w_j = x_j$.
- Chain rule on $L = e^2$: $\partial L/\partial w_j = 2e\cdot x_j$. With $2e = -8$: $\partial L/\partial\mathbf{w} = -8\cdot[1, 2, -1] = [-8, -16, 8]$.
As a column: $\nabla L = [-8, -16, 8]^\top$ ($3\times1$). As a Jacobian: $J_L = [-8,\ -16,\ 8]$ ($1\times3$).
Check by nudging. Raise $w_1$ by $0.001$: $\hat y = -0.999$, $e = -3.999$, $L = 15.992001$. The change is $-0.007999$, and $-0.007999/0.001\approx-8$ ✓. (Gradient descent would now move $\mathbf{w}$ against $\nabla L$: $w_1$ goes up, $w_2$ goes up, $w_3$ goes down.)
For $f:\mathbb{R}^n\to\mathbb{R}$:
$$\nabla f(\mathbf{x}) = \begin{bmatrix} \partial f/\partial x_1 \\ \vdots \\ \partial f/\partial x_n \end{bmatrix}\ (n\times1), \qquad J_f(\mathbf{x}) = \frac{\partial f}{\partial\mathbf{x}} = \bigl[\partial f/\partial x_1\ \cdots\ \partial f/\partial x_n\bigr] = (\nabla f)^\top\ (1\times n).$$So $\nabla f = J_f^\top$. The linear approximation can be written either way: $f(\mathbf{x}+\mathbf{h}) \approx f(\mathbf{x}) + J_f\,\mathbf{h} = f(\mathbf{x}) + \nabla f\cdot\mathbf{h}$ (a $1\times n$ times $n\times1$ is a number).
A useful special case: for $f(\mathbf{x}) = \mathbf{a}^\top\mathbf{x}$ ($=a_1x_1 + \dots + a_nx_n$) we get $\partial f/\partial x_j = a_j$, so $\nabla f = \mathbf{a}$ and $J_f = \mathbf{a}^\top$.
Why do we need it?
Training needs the slope of the loss with respect to every weight, collected in one object, so that the weights can be moved all at once. And we need it in a form that fits the chain rule.
Where is it used?
Every loss in machine learning: for linear and logistic regression, neural networks and SVMs, the gradient of the loss with respect to the weights is what gradient descent follows. In backpropagation the first thing computed is the row $\partial L/\partial\mathbf{a}$ at the output.
How is it used?
Compute the partial derivative for each weight and stack them as a column to get $\nabla L$; update $\mathbf{w}\leftarrow\mathbf{w} - \eta\nabla L$. When chaining, use the row version $(\nabla L)^\top$ on the left of Jacobian matrices.
"Gradient" is a column, "Jacobian of a scalar function" is a row. They hold the same numbers. If a formula multiplies a gradient on the left of a matrix, transpose it first. Books that use the other layout will show the gradient as a row; the maths is the same, only the bookkeeping differs.
Quick check: $f(\mathbf{x}) = 3x_1 - x_2 + 4x_3$. Write $\nabla f$ and $J_f$, with shapes.
$\nabla f = [3, -1, 4]^\top$ ($3\times1$) and $J_f = [3, -1, 4]$ ($1\times3$). It is $\mathbf{a}^\top\mathbf{x}$ with $\mathbf{a} = [3, -1, 4]^\top$.
The derivative of a vector with respect to a scalar (velocity) core
Flip the last case around: one input, several outputs. Time $t$ goes in, a position $\mathbf{r}(t)$ comes out. The derivative says how fast each coordinate changes, and stacked together it is the velocity.
The velocity is an arrow attached to the moving point. It points along the path (the direction of travel), and its length is the speed. The Jacobian has one column: $m\times1$.
- $\mathbf{r}(t) = [t,\ t^2]^\top$ (a parabola). Differentiate each entry: $\mathbf{r}'(t) = [1,\ 2t]^\top$. At $t = 1.5$: $\mathbf{r}' = [1, 3]$, so the point is moving right 1 and up 3 per second; the speed is $\sqrt{1 + 9}\approx3.16$.
- $\mathbf{r}(t) = [\cos t,\ \sin t]^\top$ (a circle). $\mathbf{r}'(t) = [-\sin t,\ \cos t]^\top$. At $t = \pi/2$: the point is at $[0, 1]$ (top) and $\mathbf{r}' = [-1, 0]$ (moving left). The speed is $\sqrt{\sin^2t + \cos^2t} = 1$ always. The velocity is perpendicular to the position (check: $[0, 1]\cdot[-1, 0] = 0$).
- A 3D helix $\mathbf{r}(t) = [2\cos t,\ 2\sin t,\ 0.4(t - 2\pi)]^\top$ has $\mathbf{r}'(t) = [-2\sin t,\ 2\cos t,\ 0.4]^\top$ and constant speed $\sqrt{4 + 0.16}\approx2.04$.
For a curve $\mathbf{r}:\mathbb{R}\to\mathbb{R}^m$ the derivative is taken entry by entry:
$$\mathbf{r}'(t) = \frac{d\mathbf{r}}{dt} = \lim_{h\to0}\frac{\mathbf{r}(t + h) - \mathbf{r}(t)}{h} = \begin{bmatrix} r_1'(t) \\ \vdots \\ r_m'(t) \end{bmatrix}\ (m\times1).$$It is a column vector, tangent to the path. Its length $\|\mathbf{r}'(t)\|$ is the speed. It is the Jacobian when $n = 1$.
Why do we need it?
To describe how something moves or changes over time, we need the rate of change of every coordinate together, as one arrow.
Where is it used?
Physics and robotics (velocity of a moving object), the path of the weights during training (each step is a small piece of a curve in weight space), recurrent networks that evolve a hidden state over time, and neural ODEs.
How is it used?
Differentiate each component with respect to $t$ and stack the results. Then evaluate at a chosen $t$ to get the velocity arrow, and use its length for the speed. Over a short time $\Delta t$, the position changes by about $\mathbf{r}'(t)\,\Delta t$.
Velocity is a vector, speed is a number. Velocity has a direction (along the path), speed is just its length.
Same function, different questions. $\mathbf{r}'(t)$ (derivative of a vector by a scalar) is a column, while the derivative of a scalar by a vector is a row in the numerator layout. Both are just Jacobians with the right shapes.
Quick check: $\mathbf{r}(t) = [3t,\ t^2,\ 5]^\top$. Find $\mathbf{r}'(2)$.
$\mathbf{r}'(t) = [3,\ 2t,\ 0]^\top$ (the constant 5 has derivative 0). At $t = 2$: $[3, 4, 0]^\top$, with speed $\sqrt{9 + 16} = 5$.
Jacobians inside a neural network: dense layer and softmax core
Crucial for neural networks. The two Jacobians in this section appear in every network you will ever train. In Chapter 2.9 (backpropagation) you will multiply them together, layer by layer. Spend time on these.
A dense layer does two things: first a linear step $\mathbf{z} = W\mathbf{x} + \mathbf{b}$ (mix the inputs), then an activation $\mathbf{a} = \sigma(\mathbf{z})$ applied to every entry separately (bend the result, for example ReLU or sigmoid).
Nudge an input $x_j$. It first reaches output $i$ through the weight $W_{ij}$ (the linear step), and then the activation squashes or passes that nudge by its own slope $\sigma'(z_i)$. So the Jacobian is just the weight matrix with each row scaled by the activation's slope.
Softmax turns a vector of scores (logits) into probabilities that add up to 1. Raise one score and its probability rises, but the others must fall to keep the total at 1. So the Jacobian has positive numbers on the diagonal and negative numbers elsewhere.
Dense layer with ReLU. $W = \begin{bmatrix} 1 & -1 \\ 2 & 1 \end{bmatrix}$, $\mathbf{b} = [0.5, -1]^\top$, $\mathbf{x} = [1, 2]^\top$ (the layer from the earlier example). We found $\mathbf{z} = [-0.5, 3]$ and $\mathbf{a} = [0, 3]$.
- ReLU's slope is $1$ where $z>0$ and $0$ where $z<0$. So $\sigma'(\mathbf{z}) = [0, 1]$.
- Scale row 1 of $W$ by $0$ and row 2 by $1$: $J = \operatorname{diag}(0, 1)\,W = \begin{bmatrix} 0 & 0 \\ 2 & 1 \end{bmatrix}$.
- Meaning: output 1 is "switched off" (its $z$ is negative), so nudging the input does nothing to it (row of zeros). Output 2 passes the nudge on, with weights $[2, 1]$.
Softmax. For $\mathbf{z} = [0, 0, 0]$ all three probabilities are $s_i = 1/3$. Then $J_{ii} = s_i(1 - s_i) = \tfrac13\cdot\tfrac23 = \tfrac29$ and $J_{ij} = -s_is_j = -\tfrac19$, so
$$J = \begin{bmatrix} 2/9 & -1/9 & -1/9 \\ -1/9 & 2/9 & -1/9 \\ -1/9 & -1/9 & 2/9 \end{bmatrix}, \quad\text{each row sums to } \tfrac29 - \tfrac19 - \tfrac19 = 0.$$For $\mathbf{z} = [2, 1, 0]$ we get $\mathbf{s}\approx[0.665,\ 0.245,\ 0.090]$ and a Jacobian with first row $[0.223,\ -0.163,\ -0.060]$.
Dense layer. For $\mathbf{a} = \sigma(\mathbf{z})$ with $\mathbf{z} = W\mathbf{x} + \mathbf{b}$ ($W$ is $m\times n$):
$$a_i = \sigma(z_i),\quad z_i = \sum_j W_{ij}x_j + b_i \quad\Longrightarrow\quad \frac{\partial a_i}{\partial x_j} = \sigma'(z_i)\,W_{ij},$$ $$J_{\mathbf{a}}(\mathbf{x}) = \operatorname{diag}\bigl(\sigma'(\mathbf{z})\bigr)\,W \qquad (m\times m)(m\times n) = m\times n.$$(The step $\partial a_i/\partial x_j = \sigma'(z_i)\,\partial z_i/\partial x_j$ is the chain rule from one variable; $\partial z_i/\partial x_j = W_{ij}$ is the linear-map Jacobian from before; the bias drops out.) Common slopes: sigmoid $\sigma' = \sigma(1 - \sigma)$ (at most $0.25$); $\tanh' = 1 - \tanh^2$; ReLU' $= 1$ for $z>0$, else $0$.
Softmax. $s_i = e^{z_i}/S$ with $S = \sum_k e^{z_k}$. Differentiate with the quotient rule:
- $j = i$: $\dfrac{\partial s_i}{\partial z_i} = \dfrac{e^{z_i}S - e^{z_i}e^{z_i}}{S^2} = s_i - s_i^2 = s_i(1 - s_i)$.
- $j \neq i$: $\dfrac{\partial s_i}{\partial z_j} = \dfrac{0\cdot S - e^{z_i}e^{z_j}}{S^2} = -s_is_j$.
It is symmetric, and every row and every column sums to zero. Each column sums to zero because $\sum_i s_i = 1$ never changes, so the rises and falls cancel. Each row sums to zero because adding the same number to all the logits does not change any $s_i$ at all. That second fact is why softmax can safely be computed after subtracting the largest logit (it avoids huge numbers like $e^{1000}$).
Why do we need it?
To train a network we must know how a change in a layer's input (or weights) changes its output, and how that change moves on to the next layer. These two Jacobians do that for the most common building blocks.
Where is it used?
Every fully connected layer in an MLP, the output layer of every classifier (softmax with cross-entropy), attention weights in Transformers (a softmax over scores), and the analysis of vanishing gradients (slopes of $\sigma'$ smaller than 1 shrink the signal).
How is it used?
Dense layer: compute $\mathbf{z}$, evaluate $\sigma'$ at each $z_i$, scale the rows of $W$. Softmax: compute $\mathbf{s}$ and form $\operatorname{diag}(\mathbf{s}) - \mathbf{s}\mathbf{s}^\top$. In practice libraries never build these matrices for big layers; they just apply the same rules to vectors.
Activations act entry by entry, so their Jacobian is diagonal. $a_i$ depends only on $z_i$. Softmax is different: every output depends on all inputs, so its Jacobian is a full matrix.
A zero slope kills the signal. ReLU with $z<0$ and sigmoid with large $|z|$ give (nearly) zero rows. Gradients passing through them become (nearly) zero, a major reason deep networks can be hard to train.
Softmax Jacobian is singular. Its rows sum to zero, so $J\mathbf{1} = \mathbf{0}$: it has no inverse.
Quick check: the sigmoid has $\sigma(0) = 0.5$. What is $\sigma'(0)$, and what does that say about its Jacobian?
$\sigma'(0) = 0.5\cdot0.5 = 0.25$, the largest slope the sigmoid ever has. So the sigmoid step alone shrinks every nudge by a factor of at least 4 (each row of $W$ is multiplied by at most $0.25$).
The chain rule using Jacobians core
Think of a pipeline: $\mathbf{x}\ \xrightarrow{\ \mathbf{F}\ }\ \mathbf{h}\ \xrightarrow{\ \mathbf{G}\ }\ \mathbf{y}$. Nudge the input a little. Function $\mathbf{F}$ turns that nudge into a nudge of $\mathbf{h}$, by multiplying with its Jacobian $J_{\mathbf{F}}$. Then $\mathbf{G}$ turns the nudge of $\mathbf{h}$ into a nudge of $\mathbf{y}$, by multiplying with $J_{\mathbf{G}}$.
So the total effect is two matrices applied one after the other: a matrix product. Like unit conversion (metres to feet, then feet to inches): the conversion factors multiply. In one variable the factors are numbers, $(g\circ f)' = g'(f(x))\,f'(x)$. With vectors the factors are matrices.
The order matters: $\mathbf{F}$ acts first, so $J_{\mathbf{F}}$ sits on the right (next to $\mathbf{x}$), and $J_{\mathbf{G}}$ on the left.
Let $\mathbf{F}:\mathbb{R}^2\to\mathbb{R}^2$, $\mathbf{F}(x, y) = [x^2y,\ 5x + y^2]^\top$ (from before) and $\mathbf{G}:\mathbb{R}^2\to\mathbb{R}^3$, $\mathbf{G}(u, v) = [u + v,\ uv,\ u^2]^\top$. Find the Jacobian of $\mathbf{G}\circ\mathbf{F}$ at $(1, 2)$.
- Shapes first: $\mathbf{G}\circ\mathbf{F}:\mathbb{R}^2\to\mathbb{R}^3$, so the answer is $3\times2$. And $J_{\mathbf{G}}$ is $3\times2$, $J_{\mathbf{F}}$ is $2\times2$: $(3\times2)(2\times2) = 3\times2$ ✓.
- Inner: $\mathbf{F}(1, 2) = [2, 9]$ and $J_{\mathbf{F}}(1, 2) = \begin{bmatrix} 4 & 1 \\ 5 & 4 \end{bmatrix}$.
- Outer, evaluated at the inner output $(u, v) = (2, 9)$: $J_{\mathbf{G}} = \begin{bmatrix} 1 & 1 \\ v & u \\ 2u & 0 \end{bmatrix} = \begin{bmatrix} 1 & 1 \\ 9 & 2 \\ 4 & 0 \end{bmatrix}$.
- Multiply: $\begin{bmatrix} 1 & 1 \\ 9 & 2 \\ 4 & 0 \end{bmatrix}\begin{bmatrix} 4 & 1 \\ 5 & 4 \end{bmatrix} = \begin{bmatrix} 4 + 5 & 1 + 4 \\ 36 + 10 & 9 + 8 \\ 16 + 0 & 4 + 0 \end{bmatrix} = \begin{bmatrix} 9 & 5 \\ 46 & 17 \\ 16 & 4 \end{bmatrix}$.
Check by nudging. With $\mathbf{h} = [0.01, -0.02]$ the prediction is $[9\cdot0.01 - 5\cdot0.02,\ 46\cdot0.01 - 17\cdot0.02,\ 16\cdot0.01 - 4\cdot0.02] = [-0.01,\ 0.12,\ 0.08]$. Directly: $\mathbf{G}(\mathbf{F}(1, 2)) = [11, 18, 4]$ and $\mathbf{G}(\mathbf{F}(1.01, 1.98))$ changes it by $[-0.0098,\ 0.1184,\ 0.0796]$ ✓.
If $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$ and $\mathbf{G}:\mathbb{R}^m\to\mathbb{R}^p$ are differentiable, then
$$J_{\mathbf{G}\circ\mathbf{F}}(\mathbf{x}) = J_{\mathbf{G}}\bigl(\mathbf{F}(\mathbf{x})\bigr)\; J_{\mathbf{F}}(\mathbf{x}) \qquad (p\times m)(m\times n) = p\times n.$$- Check shapes: the inner sizes ($m$ and $m$) must match; the outer sizes ($p$ and $n$) are the answer. If they don't match, you have the factors in the wrong order.
- Why: near $\mathbf{x}$, $\mathbf{F}$ behaves like the linear map $\mathbf{h}\mapsto J_{\mathbf{F}}\mathbf{h}$, and near $\mathbf{F}(\mathbf{x})$, $\mathbf{G}$ behaves like $\mathbf{k}\mapsto J_{\mathbf{G}}\mathbf{k}$. Doing one then the other is the product of the two matrices (matrix product = composing linear maps).
- One variable ($n = m = p = 1$): $g'(f(x))\cdot f'(x)$, the rule you know. Longer chains simply multiply more Jacobians: $J_3J_2J_1$.
- Scalar output (a loss). If the last function is a scalar loss $L$ ($p = 1$), then $J_L$ is a row and $(\nabla_{\mathbf{x}}(L\circ\mathbf{F}))^\top = J_L\,J_{\mathbf{F}}$, or, transposing both sides, $\nabla_{\mathbf{x}} = J_{\mathbf{F}}^\top\,\nabla_{\mathbf{h}}L$. So going backwards, the gradient is multiplied by the transposed Jacobians. That is backpropagation (Chapter 2.9), and the full chain rule gets its own chapter (Chapter 2.8).
Why do we need it?
A network is a long chain of functions. We can only differentiate it by differentiating one piece at a time and then combining the pieces. The Jacobian chain rule is the rule for combining them.
Where is it used?
Backpropagation through every layer of every deep network, automatic differentiation libraries (PyTorch, JAX, TensorFlow), sensitivity analysis of multi-stage pipelines, and the change of variables in several steps.
How is it used?
Write each stage's Jacobian at the right point (the output of the stage before it). Check the shapes. Multiply, with the last stage on the left. In practice you multiply a gradient vector by one Jacobian at a time, never forming the giant product.
Evaluate each Jacobian at the right point. $J_{\mathbf{G}}$ is evaluated at $\mathbf{F}(\mathbf{x})$, not at $\mathbf{x}$. This is the most common slip.
Order matters. Matrix products do not commute. The function applied last has its Jacobian on the left.
Huge Jacobians are never formed in practice. A layer with 4096 inputs and 4096 outputs has a $4096\times4096$ Jacobian (16 million numbers). Backprop multiplies a vector by it instead, which is much cheaper (more in Chapter 2.9).
Quick check: $\mathbf{F}:\mathbb{R}^5\to\mathbb{R}^3$ and $\mathbf{G}:\mathbb{R}^3\to\mathbb{R}^7$. What is the shape of $J_{\mathbf{G}\circ\mathbf{F}}$, and of each factor?
$J_{\mathbf{F}}$ is $3\times5$, $J_{\mathbf{G}}$ is $7\times3$, and the product $J_{\mathbf{G}}J_{\mathbf{F}}$ is $(7\times3)(3\times5) = 7\times5$.
Recap, cheat sheet and practice
- A scalar-valued function $\mathbb{R}^n\to\mathbb{R}$ returns one number; a vector-valued function $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$ returns $m$ numbers, which are $m$ scalar functions stacked. Curves, maps of the plane and neural layers are all vector functions.
- The Jacobian $J_{\mathbf{F}}$ is the $m\times n$ matrix of all partial derivatives, with entry $(i, j) = \partial F_i/\partial x_j$: rows are outputs, columns are inputs. It is the best linear copy: $\mathbf{F}(\mathbf{x} + \mathbf{h})\approx\mathbf{F}(\mathbf{x}) + J\mathbf{h}$. For $m = n$, $\det J$ is the local area scale.
- Special cases: scalar by vector gives the gradient (a column, $n\times1$) and its transpose is the Jacobian (a row, $1\times n$); vector by scalar gives the velocity $\mathbf{r}'(t)$ ($m\times1$); a linear map $A\mathbf{x} + \mathbf{b}$ has $J = A$.
- Dense layer $\sigma(W\mathbf{x} + \mathbf{b})$: $J = \operatorname{diag}(\sigma'(\mathbf{z}))\,W$. Softmax: $J = \operatorname{diag}(\mathbf{s}) - \mathbf{s}\mathbf{s}^\top$ (symmetric, rows sum to 0). Polar: $\det J = r$.
- Chain rule: $J_{\mathbf{G}\circ\mathbf{F}}(\mathbf{x}) = J_{\mathbf{G}}(\mathbf{F}(\mathbf{x}))\,J_{\mathbf{F}}(\mathbf{x})$, shapes $(p\times m)(m\times n) = p\times n$. This product of matrices is what backpropagation does (Chapter 2.9).
Cheat sheet
| Function | Derivative | Shape | Example |
|---|---|---|---|
| $f:\mathbb{R}\to\mathbb{R}$ | $f'(x)$ | $1\times1$ | $x^2\mapsto2x$ |
| $f:\mathbb{R}^n\to\mathbb{R}$ | $\nabla f$ (column), $J_f = (\nabla f)^\top$ (row) | $n\times1$, $1\times n$ | loss $\to$ gradient |
| $\mathbf{r}:\mathbb{R}\to\mathbb{R}^m$ | $\mathbf{r}'(t)$ (velocity) | $m\times1$ | $[\cos t, \sin t]\mapsto[-\sin t, \cos t]$ |
| $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$ | $J_{\mathbf{F}}$, $(J)_{ij} = \partial F_i/\partial x_j$ | $m\times n$ | layer |
| $A\mathbf{x} + \mathbf{b}$ | $A$ | $m\times n$ | linear layer |
| $\sigma(W\mathbf{x} + \mathbf{b})$ | $\operatorname{diag}(\sigma'(\mathbf{z}))W$ | $m\times n$ | dense layer |
| softmax$(\mathbf{z})$ | $\operatorname{diag}(\mathbf{s}) - \mathbf{s}\mathbf{s}^\top$ | $k\times k$ | classifier output |
| polar $\to$ Cartesian | $\begin{bmatrix}\cos\theta & -r\sin\theta \\ \sin\theta & r\cos\theta\end{bmatrix}$, $\det = r$ | $2\times2$ | area scale |
| $\mathbf{G}\circ\mathbf{F}$ | $J_{\mathbf{G}}(\mathbf{F}(\mathbf{x}))\,J_{\mathbf{F}}(\mathbf{x})$ | $p\times n$ | chain rule |
import numpy as np
np.set_printoptions(precision=4, suppress=True)
def num_jac(F, x, h=1e-6):
"""Numerical Jacobian: nudge one input at a time. Shape (m, n)."""
x = np.asarray(x, dtype=float)
cols = []
for j in range(len(x)):
e = np.zeros_like(x); e[j] = h
cols.append((F(x + e) - F(x - e)) / (2*h)) # column j = dF/dx_j
return np.stack(cols, axis=1)
# 1) a 2 -> 2 function and its exact Jacobian
F = lambda x: np.array([x[0]**2 * x[1], 5*x[0] + x[1]**2])
JF = lambda x: np.array([[2*x[0]*x[1], x[0]**2],
[5.0, 2*x[1]]])
x = np.array([1.0, 2.0])
print(F(x)) # [2. 9.]
print(JF(x)) # [[4. 1.] [5. 4.]]
print(num_jac(F, x)) # [[4. 1.] [5. 4.]] (matches)
# 2) polar -> Cartesian: the determinant is r
P = lambda p: np.array([p[0]*np.cos(p[1]), p[0]*np.sin(p[1])])
p0 = np.array([2.0, np.pi/6])
Jp = num_jac(P, p0)
print(Jp) # [[ 0.866 -1. ] [ 0.5 1.7321]]
print(np.linalg.det(Jp)) # about 2.0 (= r)
# 3) a linear map: the Jacobian is the matrix itself
A = np.array([[2., 1., 0.], [-1., 3., 4.]])
print(num_jac(lambda v: A @ v, np.array([1., 2., -1.]))) # [[ 2. 1. 0.] [-1. 3. 4.]] = A
# 4) dense layer a = relu(W x + b): J = diag(relu'(z)) @ W
W = np.array([[1., -1.], [2., 1.]]); b = np.array([0.5, -1.]); xx = np.array([1., 2.])
z = W @ xx + b
print(z) # [-0.5 3. ]
J_layer = np.diag((z > 0).astype(float)) @ W
print(J_layer) # [[0. 0.] [2. 1.]] (neuron 1 is off, so its row is zero)
print(num_jac(lambda v: np.maximum(0, W @ v + b), xx)) # [[0. 0.] [2. 1.]]
# 5) softmax: J = diag(s) - s s^T (symmetric, every row sums to 0)
def softmax(z):
e = np.exp(z - z.max()); return e / e.sum()
zs = np.array([2., 1., 0.]); s = softmax(zs)
Js = np.diag(s) - np.outer(s, s)
print(s) # [0.6652 0.2447 0.09 ]
print(Js) # [[ 0.2227 -0.1628 -0.0599] [-0.1628 0.1848 -0.022 ] [-0.0599 -0.022 0.0819]]
print(Js.sum(axis=1)) # [0. 0. 0.] rows sum to zero
print(np.abs(Js - num_jac(softmax, zs)).max() < 1e-8) # True
# 6) chain rule: J_(G o F) = J_G(F(x)) @ J_F(x) shapes (3x2)(2x2) = 3x2
G = lambda h: np.array([h[0] + h[1], h[0]*h[1], h[0]**2])
JG = lambda h: np.array([[1., 1.], [h[1], h[0]], [2*h[0], 0.]])
print(JG(F(x)) @ JF(x)) # [[ 9. 5.] [46. 17.] [16. 4.]]
print(num_jac(lambda v: G(F(v)), x)) # the same numbers
1. $\mathbf{F}:\mathbb{R}^3\to\mathbb{R}^2$. What is the shape of its Jacobian?
2. The Jacobian of $\mathbf{F}(x, y) = [x + y,\ xy]^\top$ at $(2, 3)$ is…
3. $f:\mathbb{R}^4\to\mathbb{R}$. In this guide's convention, which statement is true?
4. $\mathbf{F}(\mathbf{x}) = A\mathbf{x} + \mathbf{b}$ with a fixed matrix $A$ and vector $\mathbf{b}$. Its Jacobian is…
5. $\mathbf{F}:\mathbb{R}^2\to\mathbb{R}^3$ and $\mathbf{G}:\mathbb{R}^3\to\mathbb{R}^4$. The Jacobian of $\mathbf{G}\circ\mathbf{F}$ is…
6. For polar-to-Cartesian coordinates, $\det J = r$. What does it tell you?
Practice problems
A. Find the Jacobian of $\mathbf{F}(x, y, z) = [xy,\ yz]^\top$ and evaluate it at $(1, 2, 3)$.
$n = 3$, $m = 2$, so $J$ is $2\times3$. Row 1 (for $xy$): $[y, x, 0]$. Row 2 (for $yz$): $[0, z, y]$. So $J = \begin{bmatrix} y & x & 0 \\ 0 & z & y \end{bmatrix}$, and at $(1, 2, 3)$: $J = \begin{bmatrix} 2 & 1 & 0 \\ 0 & 3 & 2 \end{bmatrix}$. The zeros appear where an output does not depend on an input ($xy$ has no $z$).
B. Find the Jacobian and determinant of $\mathbf{F}(x, y) = [e^x\cos y,\ e^x\sin y]^\top$ at $(0, \pi/2)$.
$J = \begin{bmatrix} e^x\cos y & -e^x\sin y \\ e^x\sin y & e^x\cos y \end{bmatrix}$ and $\det J = e^{2x}(\cos^2y + \sin^2y) = e^{2x}$. At $(0, \pi/2)$: $\cos = 0$, $\sin = 1$, $e^0 = 1$, so $J = \begin{bmatrix} 0 & -1 \\ 1 & 0 \end{bmatrix}$ (a rotation by $90^\circ$) and $\det J = 1$: areas near that point are unchanged.
C. For $f(\mathbf{x}) = \|\mathbf{x}\|^2 = x_1^2 + x_2^2 + x_3^2$, write $\nabla f$ and $J_f$ at $\mathbf{x} = [1, -2, 3]^\top$.
$\partial f/\partial x_j = 2x_j$, so $\nabla f = 2\mathbf{x} = [2, -4, 6]^\top$ ($3\times1$) and $J_f = [2, -4, 6]$ ($1\times3$).
D. A ReLU layer has $W = \begin{bmatrix} 2 & 0 \\ 1 & -1 \end{bmatrix}$ and pre-activations $\mathbf{z} = [0.3, -0.7]$. Find its Jacobian.
ReLU' is $1$ for $z>0$ and $0$ for $z<0$, so $\sigma'(\mathbf{z}) = [1, 0]$. $J = \operatorname{diag}(1, 0)\,W = \begin{bmatrix} 2 & 0 \\ 0 & 0 \end{bmatrix}$. The second output is switched off, so its row is zero.
E. $\mathbf{F}(\mathbf{x}) = [2x_1,\ x_1 + x_2]^\top$ and $L(u, v) = u^2 + v^2$. Use the chain rule to find $\nabla(L\circ\mathbf{F})$ at $(1, 1)$, and check it directly.
$\mathbf{F}(1, 1) = [2, 2]$. $J_L = [2u, 2v] = [4, 4]$ ($1\times2$). $J_{\mathbf{F}} = \begin{bmatrix} 2 & 0 \\ 1 & 1 \end{bmatrix}$. Product: $[4, 4]\begin{bmatrix} 2 & 0 \\ 1 & 1 \end{bmatrix} = [8 + 4,\ 0 + 4] = [12, 4]$, so $\nabla(L\circ\mathbf{F}) = [12, 4]^\top$. Directly, $L\circ\mathbf{F} = 4x_1^2 + (x_1 + x_2)^2$, so $\partial/\partial x_1 = 8x_1 + 2(x_1 + x_2) = 12$ and $\partial/\partial x_2 = 2(x_1 + x_2) = 4$ ✓.
F. For two classes with logits $\mathbf{z} = [0, 0]$, find the softmax Jacobian and compare with the sigmoid's slope.
$\mathbf{s} = [0.5, 0.5]$. $J = \operatorname{diag}(\mathbf{s}) - \mathbf{s}\mathbf{s}^\top = \begin{bmatrix} 0.5 - 0.25 & -0.25 \\ -0.25 & 0.5 - 0.25 \end{bmatrix} = \begin{bmatrix} 0.25 & -0.25 \\ -0.25 & 0.25 \end{bmatrix}$. Rows sum to $0$ ✓. The entry $0.25$ is exactly $\sigma'(0)$: with two classes, $s_1 = \sigma(z_1 - z_2)$, so softmax is the sigmoid in disguise.
Matrix Calculus
Calculus when the input is a list of numbers, or a whole table of numbers. You will learn one skill: take any expression with vectors and matrices, find its derivative, and know exactly what shape the answer must have. We derive every result step by step, so nothing has to be memorised.
- Name the kinds of derivative (scalar, vector, matrix) and read the shape of each answer
- Use the notation: gradient $\nabla f$, Jacobian $J$, Hessian $H$, and know the two layout conventions
- Use differentials ($d\mathbf{x}$, $dX$) and the trace trick to derive gradients without index gymnastics
- Derive $\nabla(\mathbf{x}^\top\mathbf{x})$, $A\mathbf{x}$, $\mathbf{a}^\top\mathbf{x}$, $\mathbf{x}^\top A\mathbf{x}$, $\|A\mathbf{x}-\mathbf{b}\|^2$, $\log\mathbf{x}$ and $e^{\mathbf{x}}$, each two ways
- Check any derivative numerically
This chapter builds on partial derivatives and gradients (2.4) and the Jacobian (2.5). It also uses matrices and the dot product from the Linear Algebra guide. That guide has a shorter overview: Matrix calculus in the Linear Algebra guide. Here we go deeper: we derive each identity from scratch, and we add differentials and the trace trick. The conventions are the same: the gradient is a column vector, and the Jacobian of a map from $\mathbb{R}^n$ to $\mathbb{R}^m$ is an $m\times n$ matrix.
The map of derivatives: what depends on what core
Picture a control panel. It has knobs (the inputs) and meters (the outputs). A derivative answers one question: "if I turn this knob a tiny bit, how much does that meter move?"
The knobs can be one number, a list of numbers (a vector), or a whole grid (a matrix). The meters can be the same. So there are nine combinations. You do not need to memorise nine different things. There is one rule behind all of them:
- There is one slope for every (meter, knob) pair. So the number of entries in a derivative is (number of outputs) × (number of inputs).
- The derivative is just those slopes, arranged in a sensible grid.
- One knob, one meter. $f(x)=x^2$. One slope: $f'(x)=2x$.
- Three knobs, one meter. A loss that depends on three weights. Three slopes, one per weight. We stack them into a list (a vector): the gradient.
- Three knobs, two meters. $2\times3=6$ slopes. We lay them in a table with one row per meter and one column per knob: the Jacobian.
- A matrix of knobs, one meter. A weight matrix $W$ with 6 entries and one loss. Six slopes, laid out in the same grid as $W$. This is the gradient with respect to a matrix.
Here is the whole map. A scalar is one number. $\mathbf{x}\in\mathbb{R}^n$ is a vector. $X$ is a $p\times q$ matrix.
| Output ↓ Input → | scalar $x$ | vector $\mathbf{x}$ ($n$ entries) | matrix $X$ ($p\times q$) |
|---|---|---|---|
| scalar $f$ | $f'(x)$, a number | gradient $\nabla f$: $n\times1$ column | gradient matrix $\nabla_X f$: $p\times q$ |
| vector $\mathbf{y}$ ($m$ entries) | $\mathbf{y}'(x)$: $m\times1$ column (a velocity) | Jacobian $J$: $m\times n$ | a 3-index table $m\times p\times q$ (avoid) |
| matrix $Y$ ($r\times s$) | $r\times s$ (derivative of each entry) | 3-index table (avoid) | 4-index table $r\times s\times p\times q$ (avoid) |
Scalar derivatives (top left), vector derivatives (the middle cells) and matrix derivatives (the right column) are the three families this chapter walks through. The "avoid" cells hold a lot of numbers. We never write them out. Instead we use differentials (later in this chapter) to get the answer we need.
Why do we need it?
Models have many inputs (weights) and sometimes many outputs. We need a clear way to say "the sensitivity of everything to everything" without getting lost in indices.
Where is it used?
The gradient in gradient descent, the Jacobian of each layer in backpropagation, the gradient of a loss with respect to a weight matrix $W$ in a linear layer, and the Hessian in second-order optimisers.
How is it used?
Ask two questions: what are my inputs and outputs (number, vector, matrix)? Then count: entries = outputs × inputs. That tells you the shape of the answer before you compute a single slope.
Quick check: a layer maps 4 inputs to 3 outputs. How many slopes are in its derivative, and what is its shape?
$3\times4=12$ slopes. It is a vector-to-vector map, so the derivative is the Jacobian, a $3\times4$ matrix (one row per output, one column per input).
Scalar derivatives: the "nudge" picture
You already know that $f'(x)$ is the slope of the curve. Here is the same idea in the form we will use all chapter: nudge and respond.
Move the input by a tiny amount, called $dx$. The output moves by a tiny amount, called $df$. The slope $f'(x)$ is the conversion rate between them: $df \approx f'(x)\,dx$. For matrices, we will keep exactly this form, only $dx$ will become a vector or a matrix.
Let $f(x)=x^2$ at $x=3$. Then $f(3)=9$ and $f'(3)=6$.
- Nudge by $dx=0.1$: the new value is $3.1^2=9.61$. The true change is $0.61$. The prediction $f'(3)\,dx = 6\times0.1=0.6$.
- Nudge by $dx=0.01$: true change $3.01^2-9=0.0601$. Prediction $0.06$.
- Nudge by $dx=0.001$: true change $0.006001$. Prediction $0.006$.
The prediction gets better as the nudge gets smaller. The leftover error here is exactly $dx^2$ (it is $0.01$, $0.0001$, $0.000001$). It shrinks much faster than the nudge itself.
The derivative is $f'(x)=\dfrac{df}{dx}=\displaystyle\lim_{dx\to0}\frac{f(x+dx)-f(x)}{dx}$. The differential form says the same thing:
$$df = f'(x)\,dx.$$Think of $dx$ as "a nudge so small that squares of it ($dx^2$) can be ignored". The rules you know become rules about differentials:
- Sum: $d(u+v)=du+dv$. Constant factor: $d(cu)=c\,du$.
- Product: $d(uv)=u\,dv+v\,du$.
- Chain: if $y=g(f)$ then $dy=g'(f)\,df$. For example $y=\sin(x^2)$: $dy=\cos(x^2)\cdot 2x\,dx$.
Why do we need it?
The "nudge and respond" form $df = f'(x)\,dx$ extends to vectors and matrices without changing shape. Slopes alone do not, but nudges do.
Where is it used?
Every first-order method: a gradient descent step is "choose $dx$ so that $df$ is negative". It is also the idea behind sensitivity analysis and error propagation.
How is it used?
To differentiate, write the differential of each piece with the rules above, then collect everything that multiplies $dx$. That coefficient is the derivative.
$df$ and $dx$ are small nudges, not the final derivative. The derivative is the ratio $df/dx$. A tangent-line prediction is only trustworthy when $dx$ is small.
Quick check: $f(x)=x^3$ at $x=2$. Predict $df$ for $dx=0.01$, then compare with the truth.
$f'(2)=3\cdot2^2=12$, so the prediction is $12\times0.01=0.12$. The truth is $2.01^3-8=8.120601-8=0.120601$. The error ($0.000601$) is tiny.
Vector derivatives: gradient and velocity core
Now let the input be a vector: several knobs at once. Nudge knob 1 by $dx_1$ and knob 2 by $dx_2$. Each knob adds its own share to the change: (slope of knob 1) × $dx_1$ plus (slope of knob 2) × $dx_2$. That is a dot product:
$df \approx (\text{list of slopes})\cdot(\text{list of nudges})$.
The list of slopes is the gradient. The other direction also exists: one knob (say time $t$) and a vector of meters, like the position $(x(t),y(t))$ of a moving dot. Its derivative is the list of speeds: the velocity.
$f(x_1,x_2)=x_1^2+3x_1x_2$ at $(1,2)$. The value is $f=1+6=7$.
- $\partial f/\partial x_1 = 2x_1+3x_2 = 2+6=8$ and $\partial f/\partial x_2 = 3x_1=3$. So $\nabla f=[8,\,3]^\top$.
- Nudge by $d\mathbf{x}=[0.01,\,-0.02]^\top$. Prediction: $\nabla f^\top d\mathbf{x}=8(0.01)+3(-0.02)=0.08-0.06=0.02$.
- Truth: $f(1.01,\,1.98)=1.0201+3(1.01)(1.98)=1.0201+5.9994=7.0195$. So the true change is $0.0195$. ✓ Close to $0.02$.
Velocity. A point moves on a circle: $\mathbf{r}(t)=[\cos t,\,\sin t]^\top$. Differentiate each entry: $\mathbf{r}'(t)=[-\sin t,\,\cos t]^\top$. At $t=0$ this is $[0,1]^\top$: the dot starts at $(1,0)$ and moves straight up.
Gradient (scalar output, vector input $\mathbf{x}\in\mathbb{R}^n$). It is a column vector with the shape of $\mathbf{x}$:
$$\nabla f(\mathbf{x})=\begin{bmatrix}\partial f/\partial x_1\\ \vdots\\ \partial f/\partial x_n\end{bmatrix},\qquad df=\nabla f^\top d\mathbf{x}=\sum_{i=1}^n\frac{\partial f}{\partial x_i}\,dx_i.$$Derivative of a vector with respect to a number (vector output $\mathbf{y}\in\mathbb{R}^m$, scalar input $t$): the column of ordinary derivatives,
$$\frac{d\mathbf{y}}{dt}=\begin{bmatrix}dy_1/dt\\ \vdots\\ dy_m/dt\end{bmatrix},\qquad d\mathbf{y}=\frac{d\mathbf{y}}{dt}\,dt.$$The gradient lives in the same space as the input. That is why "move opposite to the gradient" makes sense: $\mathbf{x}-\eta\nabla f$ is a vector of the same shape as $\mathbf{x}$.
Why do we need it?
A model has thousands or millions of weights. One list of slopes, one per weight, tells us how to change all of them at once.
Where is it used?
Gradient descent, SGD and Adam for every trained model, gradient clipping, saliency maps (which pixels matter), and velocities in physics simulations and ODE solvers.
How is it used?
Compute the slope for each input, stack them in a column, and update $\mathbf{x}\leftarrow\mathbf{x}-\eta\nabla f$. To predict a small change use $df=\nabla f^\top d\mathbf{x}$.
Some books write the gradient as a row. In this guide it is always a column, the same shape as $\mathbf{x}$. If a formula you find online looks "transposed", this is usually why (see the section on layout below).
Quick check: $f(\mathbf{x})=x_1x_2$ at $(3,5)$. What is $\nabla f$, and what does $d\mathbf{x}=[0.1,0]^\top$ do to $f$?
$\nabla f=[x_2,\,x_1]^\top=[5,3]^\top$. The change is about $\nabla f^\top d\mathbf{x}=5(0.1)+3(0)=0.5$. (True: $3.1\cdot5-15=0.5$.)
Derivatives of vector functions: the Jacobian core
A layer of a neural network takes a vector in and gives a vector out. Every output has its own gradient (its own list of slopes). Stack those lists as rows and you get a table: the Jacobian. Row $i$ answers "how does output $i$ respond to each input?". Column $j$ answers "what does turning input $j$ do to all the outputs?".
The key fact: for a small nudge, the output nudge is the Jacobian times the input nudge. The Jacobian is the best straight-line (linear) description of the function near the point. (Chapter 2.5 develops this picture. Here we focus on notation and shapes.)
$F(x_1,x_2,x_3)=(x_1x_2,\ \ x_2+x_3^2)$. Three inputs, two outputs, so $J$ is $2\times3$.
- Gradient of output 1: $[x_2,\ x_1,\ 0]$. Gradient of output 2: $[0,\ 1,\ 2x_3]$.
- Stack as rows: $J=\begin{bmatrix}x_2&x_1&0\\0&1&2x_3\end{bmatrix}$.
- At $\mathbf{x}=(1,2,3)$: $F=(2,\,11)$ and $J=\begin{bmatrix}2&1&0\\0&1&6\end{bmatrix}$.
- Nudge $d\mathbf{x}=[0.1,\,0,\,-0.1]^\top$. Then $J\,d\mathbf{x}=[0.2,\ -0.6]^\top$.
- Truth: $F(1.1,\,2,\,2.9)=(2.2,\ 10.41)$, so the true change is $[0.2,\ -0.59]^\top$. ✓
For $F:\mathbb{R}^n\to\mathbb{R}^m$ with outputs $F_1,\dots,F_m$, the Jacobian is the $m\times n$ matrix
$$J_F(\mathbf{x})=\begin{bmatrix}\partial F_1/\partial x_1&\cdots&\partial F_1/\partial x_n\\ \vdots&&\vdots\\ \partial F_m/\partial x_1&\cdots&\partial F_m/\partial x_n\end{bmatrix},\quad (J_F)_{ij}=\frac{\partial F_i}{\partial x_j},\quad d\mathbf{y}=J_F\,d\mathbf{x}.$$- Row $i$ is $(\nabla F_i)^\top$. When $m=1$, $J_f=(\nabla f)^\top$ (a single row).
- A linear map $F(\mathbf{x})=A\mathbf{x}$ has $J=A$ everywhere.
- Derivative of a vector function = a derivative of a vector with respect to a vector. This is why the Jacobian is also called "the derivative" of a vector function.
Why do we need it?
Layers output vectors. To pass "how sensitive is the loss?" backwards through a layer we need one object that says how every output reacts to every input.
Where is it used?
Backpropagation through every layer (Chapter 2.8 and 2.9), the Jacobian determinant in normalising flows, Gauss–Newton and Levenberg–Marquardt curve fitting, and robot-arm kinematics.
How is it used?
Compute (or let autograd compute) $J$, check that its shape is outputs × inputs, and use $d\mathbf{y}\approx J\,d\mathbf{x}$ to predict small changes.
Quick check: $F(\mathbf{x})=A\mathbf{x}+\mathbf{b}$ with $A$ a $3\times2$ matrix. What is $J$ and what is its shape?
$J=A$, shape $3\times2$. The constant $\mathbf{b}$ shifts the output but does not change how it reacts to a nudge. (This is the "affine" case: $d\mathbf{y}=A\,d\mathbf{x}$ exactly.)
Matrix derivatives: matrices as inputs and outputs
A layer of a neural network stores its weights in a matrix $W$. The loss $L$ is one number. Every single entry $W_{ij}$ is a knob, so there is one slope $\partial L/\partial W_{ij}$ per entry. We arrange those slopes in the same grid as $W$. The result is the gradient matrix: a "map" showing which weights the loss cares about most.
What if the output is also a matrix, like $Y=X^2$ or $Y=X^{-1}$? Then each of the output entries has its own gradient matrix. That is a four-index table. It is huge and awkward, so in practice we never write it. We use a differential instead (next sections).
- Sum of entries. $f(X)=\sum_{i,j}X_{ij}$. Every entry has slope 1, so $\nabla_Xf$ is a matrix of ones.
- A custom function. $f(X)=X_{11}^2+2X_{12}+3X_{21}X_{22}$. Slopes: $\partial f/\partial X_{11}=2X_{11}$, $\partial f/\partial X_{12}=2$, $\partial f/\partial X_{21}=3X_{22}$, $\partial f/\partial X_{22}=3X_{21}$. At $X=\begin{bmatrix}1&0\\2&1\end{bmatrix}$: $\nabla_Xf=\begin{bmatrix}2&2\\3&6\end{bmatrix}$.
- Sum of squares. $f(X)=\sum X_{ij}^2$. Each slope is $2X_{ij}$, so $\nabla_Xf=2X$.
- A matrix-valued function. $Y=X^2$ for a $2\times2$ matrix $X=\begin{bmatrix}a&b\\c&d\end{bmatrix}$ gives $Y=\begin{bmatrix}a^2+bc&ab+bd\\ac+cd&bc+d^2\end{bmatrix}$. Four output entries, each with a $2\times2$ gradient matrix: for example $\nabla_X Y_{11}=\begin{bmatrix}2a&c\\b&0\end{bmatrix}$. That is $4\times4=16$ numbers.
For a scalar function $f$ of a $p\times q$ matrix $X$, the gradient matrix $\nabla_Xf$ has the same shape as $X$:
$$(\nabla_Xf)_{ij}=\frac{\partial f}{\partial X_{ij}},\qquad df=\sum_{i,j}\frac{\partial f}{\partial X_{ij}}\,dX_{ij}=\operatorname{tr}\!\big((\nabla_Xf)^\top dX\big).$$The last form uses the trace (explained below). It is the matrix version of $df=\nabla f^\top d\mathbf{x}$.
Derivative of a matrix function $Y(X)$: one gradient matrix per entry of $Y$. Two ways to tame it: (1) work with the differential $dY$, which has the same shape as $Y$ and is easy to compute; (2) vectorise: stack the entries of $X$ into one long vector, and then it is an ordinary Jacobian (awareness: this is where the Kronecker product appears).
Why do we need it?
Weights live in matrices. We need the slope of the loss for each weight, laid out so that the update $W\leftarrow W-\eta\,\nabla_WL$ makes sense entry by entry.
Where is it used?
Every linear layer, convolution kernel and attention projection matrix in a neural network; covariance estimation, matrix factorisation and PCA objectives.
How is it used?
Get the gradient matrix with the same shape as the weight matrix, then subtract a small multiple of it. In code, W.grad in PyTorch always has the shape of W.
Quick check: $X$ is $4\times5$ and $f(X)=\sum X_{ij}^2$. What is the shape of $\nabla_Xf$, and what is it?
It is $4\times5$ (the shape of $X$), and it equals $2X$.
Notation: gradient, Jacobian, Hessian, and the layout question core
There are three names to keep straight, and they are three sizes of the same idea:
- Gradient $\nabla f$: the slopes of one output with respect to all inputs.
- Jacobian $J$: the slopes of all outputs with respect to all inputs (a stack of gradients).
- Hessian $H$: the slope of the slope: how each slope changes as each input changes. It measures curvature.
There is also an annoying fact of life: different books arrange these numbers differently (rows or columns). This is the main reason why formulas from two sources can look "transposed" from each other. We will be honest about it and then pick one convention.
$f(x_1,x_2)=x_1^2+3x_1x_2$ again.
- Gradient: $\nabla f=\begin{bmatrix}2x_1+3x_2\\3x_1\end{bmatrix}$.
- Hessian: differentiate each entry of the gradient with respect to $x_1$ and $x_2$. The first entry gives $[2,\ 3]$. The second gives $[3,\ 0]$. So $H=\begin{bmatrix}2&3\\3&0\end{bmatrix}$.
- $H$ is symmetric: $\partial^2f/\partial x_1\partial x_2=\partial^2f/\partial x_2\partial x_1=3$. This always holds when the second derivatives are continuous.
Notation used in this guide.
- Gradient of $f:\mathbb{R}^n\to\mathbb{R}$: $\nabla f$ (or $\nabla_{\mathbf{x}}f$), an $n\times1$ column.
- Jacobian of $F:\mathbb{R}^n\to\mathbb{R}^m$: $J_F$ (also written $\partial F/\partial\mathbf{x}^\top$ or $DF$), an $m\times n$ matrix with $(J_F)_{ij}=\partial F_i/\partial x_j$.
- Hessian of $f$: $H=\nabla^2f$, the $n\times n$ matrix with $H_{ij}=\dfrac{\partial^2f}{\partial x_i\,\partial x_j}$. It is the Jacobian of the gradient: $H=J_{\nabla f}$.
- Differentials: $df=\nabla f^\top d\mathbf{x}$, $\ d\mathbf{y}=J\,d\mathbf{x}$, $\ df=\operatorname{tr}(G^\top dX)\Rightarrow G=\nabla_Xf$.
The layout question. There are two conventions for arranging "$\partial\mathbf{y}/\partial\mathbf{x}$":
| Numerator layout | Denominator layout | |
|---|---|---|
| Rule | shape = (size of $\mathbf{y}$) × (size of $\mathbf{x}$) | shape = (size of $\mathbf{x}$) × (size of $\mathbf{y}$) |
| $\partial f/\partial\mathbf{x}$ for scalar $f$ | a row ($1\times n$) | a column ($n\times1$) |
| Jacobian of $F:\mathbb{R}^n\to\mathbb{R}^m$ | $m\times n$ | $n\times m$ (the transpose) |
| Chain rule order | $J_{f\circ g}=J_f\,J_g$ (natural order) | order reversed |
This guide uses a mixed convention, the most common one in machine learning: the Jacobian is $m\times n$ (numerator layout), but the gradient of a scalar is written as a column of the same shape as the input. They fit together through $\nabla f=J_f^\top$ when $m=1$. The chain rule for gradients then reads $\nabla(f\circ g)=J_g^\top\,\nabla f$. The gradient with respect to a matrix $X$ always has the shape of $X$. We avoid writing $\partial f/\partial\mathbf{x}$ for vectors, to avoid any doubt. The Linear Algebra guide uses the same rules.
Why do we need it?
Without agreed notation, a formula like "$A^\top(A\mathbf{x}-\mathbf{b})$ or $(A\mathbf{x}-\mathbf{b})^\top A$?" is a coin toss. Shapes decide it, but only if you know the convention.
Where is it used?
Optimisation papers and textbooks (each picks a layout), deep-learning libraries (autograd returns gradients with the shape of the parameter), and second-order methods that use the Hessian (Newton's method, L-BFGS, natural gradient).
How is it used?
State the convention once, then check the shape of every formula. If a source looks transposed from yours, check whether it uses the other layout, and transpose the answer.
When you read another book, look at its very first example. If its gradient of a scalar is a row, it uses numerator layout throughout. Do not mix formulas from two layouts without transposing. The Hessian of a smooth function is symmetric, so it never changes.
Quick check: in this guide, $F:\mathbb{R}^5\to\mathbb{R}^3$. What is the shape of $J_F$, and what is the shape of $J_F^\top$?
$J_F$ is $3\times5$ (outputs × inputs). $J_F^\top$ is $5\times3$: it is what multiplies a $3\times1$ output-side gradient to give a $5\times1$ input-side gradient in backpropagation.
Reading the shape of a derivative core
Shapes are your best friend in matrix calculus. Before computing anything, you can usually tell what the answer must look like. And after computing it, a wrong shape instantly tells you that you made a mistake. This is the cheapest bug-check there is.
Two simple rules cover almost everything:
- Scalar output: the gradient has exactly the shape of the input. (Number in, number out. Vector in, vector. Matrix in, matrix.)
- Everything else: output shape first, then input shape. A vector-to-vector map gives (output length) × (input length).
- $f(\mathbf{x})=\|A\mathbf{x}-\mathbf{b}\|^2$ with $\mathbf{x}\in\mathbb{R}^4$: scalar output, so the gradient has shape $4\times1$.
- Check the formula $2A^\top(A\mathbf{x}-\mathbf{b})$ with $A$ of size $3\times4$: $A^\top$ is $4\times3$ and $(A\mathbf{x}-\mathbf{b})$ is $3\times1$. The product is $(4\times3)(3\times1)=4\times1$. ✓ The shape fits.
- Try the wrong formula $2A(A\mathbf{x}-\mathbf{b})$: $(3\times4)(3\times1)$ cannot be multiplied. ✗ We caught the mistake without computing a thing.
- A layer $F:\mathbb{R}^4\to\mathbb{R}^3$: Jacobian is $3\times4$. A weight matrix $W$ ($3\times4$) and a scalar loss: $\nabla_WL$ is $3\times4$.
For input shape $S_{in}$ and output shape $S_{out}$, the derivative has shape:
| Case | Shape of the derivative |
|---|---|
| scalar output, any input | $S_{in}$ (the gradient has the input's shape) |
| vector output ($m$), scalar input | $m\times1$ |
| vector output ($m$), vector input ($n$) | $m\times n$ (Jacobian) |
| any other | $S_{out}$ followed by $S_{in}$ (a higher-index table) |
Counting check: the number of entries is always (size of the output) × (size of the input). A shape check of a formula means: write the shape of each factor and make sure neighbouring inner sizes match, as in matrix multiplication.
Why do we need it?
Matrix formulas are long and easy to garble. A shape check is a free proof-reader: wrong formulas often refuse to multiply.
Where is it used?
Every time you write a backward pass by hand, implement a loss in NumPy or PyTorch, or debug a "size mismatch" error. Gradients with the wrong shape are one of the most common bugs in deep-learning code.
How is it used?
Write the shape under each symbol, from right to left, and check that each product is allowed. The final shape must equal the shape the table above predicts.
A shape check is necessary, not sufficient. $A^\top(A\mathbf{x}-\mathbf{b})$ and $A^\top(\mathbf{b}-A\mathbf{x})$ both fit, but only one has the right sign. Use shapes to catch the wrong arrangement, then use a numerical check (later) for the rest.
Quick check: $W$ is $5\times8$, $\mathbf{x}$ has 8 entries, $L=\mathbf{c}^\top(W\mathbf{x})$ with $\mathbf{c}\in\mathbb{R}^5$. What shape is $\nabla_WL$?
The same as $W$: $5\times8$. (We will derive $\nabla_WL=\mathbf{c}\,\mathbf{x}^\top$, which is $(5\times1)(1\times8)=5\times8$. ✓)
Matrix differential notation: nudge the whole matrix core
Taking a derivative "with respect to a matrix" sounds scary. The trick is to stop asking for the derivative and ask for the nudge response instead. Write $dX$ for "a tiny nudge of the whole matrix": a matrix of the same shape as $X$ whose entries are tiny nudges $dX_{ij}$.
Then work out how everything else responds to that nudge, using the familiar product rule. At the very end, read the gradient off the response. No indices, no big tables.
One warning: matrices do not commute ($AB\neq BA$ in general). So the order of factors in a differential rule matters, even though in one variable it never did.
Take $X=\begin{bmatrix}1&2\\3&4\end{bmatrix}$ and $Y=\begin{bmatrix}0&1\\1&0\end{bmatrix}$. Nudge $X$ by $dX=\begin{bmatrix}0.1&0\\0&0\end{bmatrix}$ and $Y$ by $dY=\begin{bmatrix}0.1&0\\0&0\end{bmatrix}$.
- Original product: $XY=\begin{bmatrix}2&1\\4&3\end{bmatrix}$.
- New product: $(X+dX)(Y+dY)=\begin{bmatrix}1.1&2\\3&4\end{bmatrix}\begin{bmatrix}0.1&1\\1&0\end{bmatrix}=\begin{bmatrix}2.11&1.1\\4.3&3\end{bmatrix}$. True change: $\begin{bmatrix}0.11&0.1\\0.3&0\end{bmatrix}$.
- Rule: $dX\,Y=\begin{bmatrix}0&0.1\\0&0\end{bmatrix}$ and $X\,dY=\begin{bmatrix}0.1&0\\0.3&0\end{bmatrix}$. Sum: $\begin{bmatrix}0.1&0.1\\0.3&0\end{bmatrix}$.
- Compare: true change $0.11$ against predicted $0.1$ in the top-left. The leftover $0.01$ is $dX\,dY$, a product of two tiny nudges (second order), which the rule ignores.
Rules for differentials ($A$ is constant; $X,Y$ vary; $c$ is a number):
| Rule | Why it holds |
|---|---|
| $d(A)=0$, $\ d(cX)=c\,dX$, $\ d(X+Y)=dX+dY$ | constants do not move; nudging is linear |
| $d(XY)=dX\,Y+X\,dY$ | entry $(i,j)$ is $\sum_kX_{ik}Y_{kj}$. Apply the ordinary product rule to each term: $\sum_k(dX_{ik}Y_{kj}+X_{ik}dY_{kj})$, which is entry $(i,j)$ of $dX\,Y+X\,dY$ |
| $d(X^\top)=(dX)^\top$ | transposing just swaps positions of entries; nudges swap with them |
| $d\operatorname{tr}(X)=\operatorname{tr}(dX)$ | $\operatorname{tr}X=\sum_iX_{ii}$ is a sum, and nudging a sum nudges each term |
| $d(X^{-1})=-X^{-1}\,dX\,X^{-1}$ | from $XX^{-1}=I$: nudge both sides, $dX\,X^{-1}+X\,d(X^{-1})=0$. Solve for $d(X^{-1})$ by multiplying on the left by $X^{-1}$ |
| $d(X\circ Y)=dX\circ Y+X\circ dY$, $\ d\,f(X)=f'(X)\circ dX$ | the same product and chain rule, applied entry by entry ($\circ$ = entrywise product) |
Consequence: $d(X^2)=dX\,X+X\,dX$, which is not $2X\,dX$ unless $X$ and $dX$ commute. In one variable, $(1/x)'=-1/x^2$ is the scalar version of $d(X^{-1})$.
Why do we need it?
It lets us differentiate expressions made of matrix products, inverses and traces without ever writing out an entry. Long index sums disappear.
Where is it used?
Deriving the gradients of least squares, ridge regression, Gaussian log-likelihoods (with $\Sigma^{-1}$), PCA objectives, and the backward pass of a layer written in matrix form.
How is it used?
Write $d(\text{your expression})$ using the rules, working from the outside in. Keep every factor in its original order. Then use the trace trick (next section) to read the gradient off the result.
- Keep the order: $d(XY)=dX\,Y+X\,dY$, never $Y\,dX+X\,dY$.
- $d(X^{-1})=-X^{-1}\,dX\,X^{-1}$ has the inverse on both sides.
- $dX$ has the same shape as $X$. If a differential has a different shape from the thing it nudges, something is wrong.
Quick check: what is $d(AXB)$ for constant matrices $A$ and $B$?
$A$ and $B$ do not move, so the product rule gives $d(AXB)=A\,dX\,B$. (Check sizes: if $X$ is $p\times q$ then $A$ is $r\times p$, $B$ is $q\times s$, and the result is $r\times s$.)
The trace trick core
The trace of a square matrix is the sum of its diagonal entries. Why would calculus care? Because of one magic property: you can rotate matrices inside a trace: $\operatorname{tr}(AB)=\operatorname{tr}(BA)$. Even when $AB$ and $BA$ are completely different matrices (even different sizes), their traces are equal.
This is useful because a differential response like $df$ is a single number, and we want to bring the nudge $dX$ to the very end of the expression, where we can read off whatever stands in front. Rotating inside a trace does exactly that.
$A=\begin{bmatrix}1&2&0\\0&1&3\end{bmatrix}$ ($2\times3$) and $B=\begin{bmatrix}1&0\\2&1\\0&1\end{bmatrix}$ ($3\times2$).
- $AB=\begin{bmatrix}5&2\\2&4\end{bmatrix}$ ($2\times2$). Its trace is $5+4=9$.
- $BA=\begin{bmatrix}1&2&0\\2&5&3\\0&1&3\end{bmatrix}$ ($3\times3$). Its trace is $1+5+3=9$.
- Different matrices, different sizes, the same trace.
$\operatorname{tr}(A)=\sum_iA_{ii}$ for a square matrix $A$. Properties:
- Linear: $\operatorname{tr}(A+B)=\operatorname{tr}A+\operatorname{tr}B$, $\ \operatorname{tr}(cA)=c\operatorname{tr}A$. Also $\operatorname{tr}(A^\top)=\operatorname{tr}(A)$.
- Cyclic: $\operatorname{tr}(AB)=\operatorname{tr}(BA)$ whenever both products exist. For three: $\operatorname{tr}(ABC)=\operatorname{tr}(BCA)=\operatorname{tr}(CAB)$. But in general $\operatorname{tr}(ACB)\ne\operatorname{tr}(ABC)$: you may rotate, not swap.
- A scalar is its own trace: a $1\times1$ matrix $a$ has $\operatorname{tr}(a)=a$. So any scalar expression can be wrapped in $\operatorname{tr}$ and then rotated.
- Dot product of matrices: $\operatorname{tr}(A^\top B)=\sum_{i,j}A_{ij}B_{ij}$.
Proof of $\operatorname{tr}(AB)=\operatorname{tr}(BA)$: $\operatorname{tr}(AB)=\sum_i\sum_kA_{ik}B_{ki}=\sum_k\sum_iB_{ki}A_{ik}=\operatorname{tr}(BA)$. We only swapped the order of two sums and two numbers.
Why it gives a gradient. By the last property, $\operatorname{tr}(G^\top dX)=\sum_{i,j}G_{ij}\,dX_{ij}$. Compare with $df=\sum_{ij}\frac{\partial f}{\partial X_{ij}}dX_{ij}$. If $df=\operatorname{tr}(G^\top dX)$ for every nudge, then $G_{ij}=\partial f/\partial X_{ij}$, i.e. $G=\nabla_Xf$.
Why do we need it?
It turns a matrix expression into a form where $dX$ stands alone at the end, so the gradient can be read off. It replaces pages of index algebra by two moves: rotate, then read.
Where is it used?
Deriving gradients for linear and ridge regression, matrix factorisation, Gaussian log-likelihoods with covariance matrices, and the backward pass of linear layers. Also in statistics: for a random vector $\mathbf{x}$ with mean zero and covariance $\Sigma$, the cyclic trick gives $\mathbb{E}[\mathbf{x}^\top A\mathbf{x}]=\operatorname{tr}(A\Sigma)$.
How is it used?
Wrap the scalar in $\operatorname{tr}$. Use the cyclic property to move $dX$ to the right end: $\operatorname{tr}(M\,dX)$. Then $G^\top=M$, so the gradient is $G=M^\top$.
You may rotate the factors of a trace, not shuffle them. $\operatorname{tr}(ABC)=\operatorname{tr}(CAB)$ but usually $\ne\operatorname{tr}(ACB)$. Also, "scalar equals its trace" applies only to $1\times1$ things: $\mathbf{x}^\top A\mathbf{x}$ is a scalar, but $A\mathbf{x}$ is not.
Quick check: rewrite the scalar $\mathbf{a}^\top X\mathbf{b}$ so that $X$ is at the end, inside a trace.
$\mathbf{a}^\top X\mathbf{b}=\operatorname{tr}(\mathbf{a}^\top X\mathbf{b})=\operatorname{tr}(\mathbf{b}\,\mathbf{a}^\top X)$ (rotate $\mathbf{b}$ to the front). Here $X$ is at the right end.
The recipe: differential, rearrange, read off core
Now put the two tools together. Ask the function: "if I nudge the input by $dX$, how much does $f$ change?" The answer is always a number built from $dX$. Rewrite it in the form "(something) multiplying $dX$", and that something is the gradient. It is like reading the price list off a bill: the bill says "3 apples at $a$ each plus 2 pears at $p$ each", and the multipliers tell you the quantities.
$f(X)=\operatorname{tr}(AX)$ for a $2\times2$ matrix $A$. We want $\nabla_Xf$.
- Nudge: $df=\operatorname{tr}(A\,dX)$ ($A$ is constant, trace is linear).
- We need the form $\operatorname{tr}(G^\top dX)$. We already have $\operatorname{tr}(\underbrace{A}_{G^\top}\,dX)$.
- So $G^\top=A$, and therefore $\nabla_Xf=A^\top$.
Check with entries: $\operatorname{tr}(AX)=\sum_{i,j}A_{ij}X_{ji}$. The coefficient of $X_{ji}$ is $A_{ij}$, so the gradient at position $(j,i)$ is $A_{ij}$: that is $A^\top$. ✓
The recipe.
- Differential. Write $df$ using the rules (constants do not move; product rule in order).
- Rearrange. If $df$ is a scalar, wrap it in $\operatorname{tr}$ and use cyclic rotations (and transposes of scalars) so that $dX$ is at the right end, with nothing to its right.
- Read off. The three standard forms:
Why the answer is unique. Suppose $\operatorname{tr}(M^\top dX)=\operatorname{tr}(G^\top dX)$ for every $dX$. Choose $dX$ to have a 1 in position $(i,j)$ and zeros elsewhere. The left side is $M_{ij}$ and the right is $G_{ij}$. So $M=G$.
Why do we need it?
It is one procedure that works for every function built from products, transposes, traces and inverses. You derive instead of looking up.
Where is it used?
Deriving custom loss gradients, checking autograd results by hand, writing backward passes for new layers, and the derivations in papers on Gaussian processes, PCA and matrix completion.
How is it used?
Three lines on paper: write $df$, rotate to put $dX$ last, read off the multiplier (transposing for a matrix gradient). Then run a numerical check and compare shapes.
- The recipe needs $df$ to be a scalar before you wrap it in a trace. If your function is vector-valued, use $d\mathbf{y}=M\,d\mathbf{x}\Rightarrow J=M$ instead.
- Do not forget the final transpose for matrix gradients: $\operatorname{tr}(M\,dX)$ gives $M^\top$, not $M$.
Quick check: use the recipe to find $\nabla_X\operatorname{tr}(X^\top B)$.
$df=\operatorname{tr}(dX^\top B)$. A trace equals the trace of its transpose, so $\operatorname{tr}(dX^\top B)=\operatorname{tr}(B^\top dX)$. Then $G^\top=B^\top$, so $\nabla_Xf=B$. (This is the "dot product of matrices" $\sum X_{ij}B_{ij}$, whose slope for $X_{ij}$ is clearly $B_{ij}$.)
Important derivative 1: $\nabla(\mathbf{x}^\top\mathbf{x})=\nabla\|\mathbf{x}\|^2=2\mathbf{x}$ core
In one variable, $x^2$ has slope $2x$. The vector version of "$x$ squared" is $\mathbf{x}^\top\mathbf{x}=x_1^2+x_2^2+\cdots+x_n^2$: the sum of the squares of the entries. It is also the squared length $\|\mathbf{x}\|^2$ (Pythagoras, from the L2 norm).
Each entry $x_k$ appears in exactly one term, $x_k^2$. So nudging $x_k$ changes the total by $2x_k\,dx_k$ and nothing else. The gradient is just "$2\times$ the vector itself". It points straight away from the origin, and it is longer the further you are from the origin.
$\mathbf{x}=[3,4]^\top$, so $f=\mathbf{x}^\top\mathbf{x}=9+16=25$.
- Prediction: $\nabla f=2\mathbf{x}=[6,8]^\top$.
- Nudge $x_1$ from $3$ to $3.01$: $f=3.01^2+4^2=9.0601+16=25.0601$. The change is $0.0601$. The prediction is $6\times0.01=0.06$ ✓.
- Nudge $x_2$ from $4$ to $4.01$: $f=9+16.0801=25.0801$. Change $0.0801$, prediction $8\times0.01=0.08$ ✓.
- Another case, in 3 dimensions: $\mathbf{x}=[1,-2,2]^\top$ has $f=1+4+4=9$ and $\nabla f=[2,-4,4]^\top$.
The two expressions are the same function, since $\|\mathbf{x}\|^2=\mathbf{x}^\top\mathbf{x}$. The gradient is radial: it is perpendicular to the circular level curves $\|\mathbf{x}\|=\text{const}$. Its length is $2\|\mathbf{x}\|$. A gradient-descent step with learning rate $\eta$ gives $\mathbf{x}-\eta\,2\mathbf{x}=(1-2\eta)\mathbf{x}$: it simply shrinks the vector. (The gradient of the length $\|\mathbf{x}\|$ itself, not squared, is $\mathbf{x}/\|\mathbf{x}\|$. We derive that in Chapter 2.7.)
Why do we need it?
The squared length is the simplest smooth measure of "how big" or "how far". Its gradient is the building block of every squared-error loss.
Where is it used?
L2 regularisation and weight decay (the penalty $\lambda\|\mathbf{w}\|^2$ adds $2\lambda\mathbf{w}$ to the gradient), squared distance in k-means, and mean squared error.
How is it used?
When a loss contains $\|\mathbf{w}\|^2$, add $2\mathbf{w}$ (or $\mathbf{w}$ for $\frac12\|\mathbf{w}\|^2$) to the gradient. Each update then shrinks the weights slightly: "weight decay".
$\nabla\|\mathbf{x}\|^2=2\mathbf{x}$, but $\nabla\|\mathbf{x}\|=\mathbf{x}/\|\mathbf{x}\|$ (a unit vector). Do not confuse the squared length with the length. Also, the gradient of $\|\mathbf{x}\|^2$ at $\mathbf{x}=\mathbf{0}$ is $\mathbf{0}$: the bottom of the bowl.
Quick check: what is $\nabla_{\mathbf{w}}\big(\tfrac\lambda2\|\mathbf{w}\|^2\big)$?
The constant $\tfrac\lambda2$ just comes along: $\tfrac\lambda2\cdot2\mathbf{w}=\lambda\mathbf{w}$. This is exactly the weight-decay term.
Important derivatives 2 and 3: the Jacobian of $A\mathbf{x}$, and $\nabla(\mathbf{a}^\top\mathbf{x})$ core
A linear function has the same slope everywhere. $f(x)=5x$ has slope $5$ at every point. The vector versions are:
- $A\mathbf{x}$: a whole list of weighted sums, one per row of $A$. Its table of slopes is $A$.
- $\mathbf{a}^\top\mathbf{x}$: one weighted sum. Its list of slopes is $\mathbf{a}$. (It is the single-row case of $A\mathbf{x}$.)
Because the slope never changes, a nudge $d\mathbf{x}$ gives an output change of exactly $A\,d\mathbf{x}$ at any point, with no leftover error.
$A=\begin{bmatrix}2&1\\0&3\\1&-1\end{bmatrix}$ ($3\times2$) and $\mathbf{x}=[1,2]^\top$.
- $F(\mathbf{x})=A\mathbf{x}=[2+2,\ 0+6,\ 1-2]^\top=[4,6,-1]^\top$.
- Outputs: $F_1=2x_1+x_2$, $F_2=3x_2$, $F_3=x_1-x_2$. Partial derivatives: $\partial F_1/\partial x_1=2$, $\partial F_1/\partial x_2=1$, and so on. They are exactly the entries of $A$. So $J=A$ ($3\times2$).
- Nudge $d\mathbf{x}=[0.5,-1]^\top$: $A\,d\mathbf{x}=[1-1,\ -3,\ 0.5+1]^\top=[0,-3,1.5]^\top$. And $F(\mathbf{x}+d\mathbf{x})-F(\mathbf{x})$: $A[1.5,1]^\top=[4,3,0.5]^\top$, minus $[4,6,-1]^\top$ gives $[0,-3,1.5]^\top$. Exactly equal ✓.
- Weighted sum. $\mathbf{a}=[3,-1,2]^\top$, $f=\mathbf{a}^\top\mathbf{x}=3x_1-x_2+2x_3$. The slopes are $3$, $-1$, $2$, which is $\mathbf{a}$. At $\mathbf{x}=[1,2,0]^\top$, $f=1$.
For $A$ of size $m\times n$: $J=A$ is $m\times n$ (outputs × inputs ✓). For the affine map $A\mathbf{x}+\mathbf{b}$ the Jacobian is still $A$. The Jacobian of the single row $\mathbf{a}^\top$ is the row $\mathbf{a}^\top$, and the gradient (a column) is its transpose $\mathbf{a}$. The one-variable cousin: $(ax)'=a$.
Why do we need it?
Every linear layer, every linear model and every weighted sum is of this form. Its derivative is the simplest of all, so it is the "base case" of backpropagation.
Where is it used?
Fully connected layers ($\mathbf{z}=W\mathbf{x}+\mathbf{b}$: the Jacobian with respect to $\mathbf{x}$ is $W$), linear and logistic regression, and the final score $\mathbf{w}^\top\mathbf{x}$ of a classifier.
How is it used?
To pass a gradient backwards through $\mathbf{z}=W\mathbf{x}+\mathbf{b}$, multiply by $W^\top$: $\nabla_{\mathbf{x}}L=W^\top\nabla_{\mathbf{z}}L$ (this is the Jacobian rule $J^\top\nabla$ from Chapter 2.8).
Shapes: the Jacobian of $A\mathbf{x}$ is $A$, but the gradient flowing back through $A\mathbf{x}$ uses $A^\top$. Do not mix them up: $J$ goes with a forward nudge ($d\mathbf{y}=J\,d\mathbf{x}$), $J^\top$ goes with a backward gradient.
Quick check: what is $\nabla_{\mathbf{x}}(\mathbf{c}^\top A\mathbf{x})$ for a constant vector $\mathbf{c}$ and matrix $A$?
Write $\mathbf{c}^\top A\mathbf{x}=(A^\top\mathbf{c})^\top\mathbf{x}$, a weighted sum with weights $\mathbf{a}=A^\top\mathbf{c}$. So the gradient is $A^\top\mathbf{c}$.
Important derivative 4: $\nabla(\mathbf{x}^\top A\mathbf{x})=(A+A^\top)\mathbf{x}$ core
In one variable, $ax^2$ has slope $2ax$. The matrix version is $\mathbf{x}^\top A\mathbf{x}=\sum_{i,j}A_{ij}x_ix_j$: a sum of terms of the form (number) × $x_i$ × $x_j$. This is called a quadratic form. The 3D picture is a bowl, a dome or a saddle (Chapter 1.12 of the Linear Algebra guide).
Why does the slope have two parts? Because each input $x_k$ appears twice: once as the left factor $x_i$ (when $i=k$) and once as the right factor $x_j$ (when $j=k$). Each appearance contributes a piece. If $A$ is symmetric the two pieces are equal, and we get $2A\mathbf{x}$, the exact analogue of $2ax$.
$A=\begin{bmatrix}2&1\\0&3\end{bmatrix}$ (not symmetric).
- Multiply out: $A\mathbf{x}=[2x_1+x_2,\ 3x_2]^\top$, so $f=x_1(2x_1+x_2)+x_2(3x_2)=2x_1^2+x_1x_2+3x_2^2$.
- Partial derivatives: $\partial f/\partial x_1=4x_1+x_2$ and $\partial f/\partial x_2=x_1+6x_2$.
- At $\mathbf{x}=[1,2]^\top$: $\nabla f=[6,\,13]^\top$ and $f=2+2+12=16$.
- Formula: $A+A^\top=\begin{bmatrix}4&1\\1&6\end{bmatrix}$, and $(A+A^\top)\mathbf{x}=[4+2,\ 1+12]^\top=[6,13]^\top$ ✓.
- The naive guess $2A\mathbf{x}=2[4,6]^\top=[8,12]^\top$ is wrong. The naive rule holds only when $A$ is symmetric.
A quadratic form only "sees" the symmetric part $\tfrac12(A+A^\top)$ of $A$: the antisymmetric part contributes nothing, because $\mathbf{x}^\top K\mathbf{x}=0$ whenever $K^\top=-K$. The Hessian (the matrix of second derivatives) is the constant $A+A^\top$. One-variable cousin: $(ax^2)'=2ax$, $(ax^2)''=2a$.
Why do we need it?
Quadratic forms measure "energy", "curvature" and "variance". Almost every smooth loss looks like a quadratic bowl near its minimum, so this is the gradient you meet most.
Where is it used?
Least squares and ridge regression (the loss expands into $\mathbf{x}^\top A^\top A\mathbf{x}$), Gaussian distributions (the exponent $-\tfrac12(\mathbf{x}-\boldsymbol\mu)^\top\Sigma^{-1}(\mathbf{x}-\boldsymbol\mu)$), PCA (maximise $\mathbf{w}^\top\Sigma\mathbf{w}$), and Newton's method.
How is it used?
Check whether $A$ is symmetric. If yes, use $2A\mathbf{x}$; if not, use $(A+A^\top)\mathbf{x}$. Setting the gradient to zero finds the bottom of the bowl.
- $2A\mathbf{x}$ is only correct when $A=A^\top$. In general it is $(A+A^\top)\mathbf{x}$. Many bugs come from forgetting this when $A$ is not symmetric (for example a weight matrix).
- $\mathbf{x}^\top A\mathbf{x}$ is a number. Its gradient is a vector. Do not drop the $\mathbf{x}$ from the answer.
Quick check: $A=\begin{bmatrix}3&1\\1&2\end{bmatrix}$ and $\mathbf{x}=[1,-1]^\top$. What is $\nabla(\mathbf{x}^\top A\mathbf{x})$?
$A$ is symmetric, so the gradient is $2A\mathbf{x}=2[3-1,\ 1-2]^\top=2[2,-1]^\top=[4,-2]^\top$. Check with $f=3x_1^2+2x_1x_2+2x_2^2$: $\partial f/\partial x_1=6x_1+2x_2=4$, $\partial f/\partial x_2=2x_1+4x_2=-2$ ✓.
Important derivative 5: $\nabla\|A\mathbf{x}-\mathbf{b}\|^2=2A^\top(A\mathbf{x}-\mathbf{b})$ core
This is the loss of least squares: $A\mathbf{x}$ is the model's prediction, $\mathbf{b}$ the targets, and $\mathbf{r}=A\mathbf{x}-\mathbf{b}$ the vector of errors (the residual). The loss is the squared length of the error vector.
Think of it as "(inside)² → 2 × inside × (slope of the inside)", the chain rule in vector form. The inside is $\mathbf{r}$. Its derivative with respect to $\mathbf{x}$ is $A$ (the Jacobian of $A\mathbf{x}$). The "2 × inside" part is $2\mathbf{r}$. The transpose $A^\top$ appears because $2\mathbf{r}$ lives in output space, and $A^\top$ carries it back to input space: it asks "which inputs caused these errors?".
$A=\begin{bmatrix}1&0\\0&2\\1&1\end{bmatrix}$, $\mathbf{b}=[2,1,4]^\top$, at $\mathbf{x}=[1,1]^\top$.
- Prediction: $A\mathbf{x}=[1,2,2]^\top$. Residual: $\mathbf{r}=A\mathbf{x}-\mathbf{b}=[-1,\ 1,\ -2]^\top$.
- Loss: $\|\mathbf{r}\|^2=1+1+4=6$.
- $A^\top\mathbf{r}=[1(-1)+0(1)+1(-2),\ \ 0(-1)+2(1)+1(-2)]^\top=[-3,\ 0]^\top$.
- Gradient: $\nabla f=2A^\top\mathbf{r}=[-6,\ 0]^\top$.
- Direct check: $f=(x_1-2)^2+(2x_2-1)^2+(x_1+x_2-4)^2$. $\partial f/\partial x_1=2(x_1-2)+2(x_1+x_2-4)=2(-1)+2(-2)=-6$. $\partial f/\partial x_2=4(2x_2-1)+2(x_1+x_2-4)=4+(-4)=0$ ✓.
- Going downhill (the negative gradient) increases $x_1$, because the first and third predictions are both too small.
Setting $\nabla f=\mathbf{0}$ gives the normal equations $A^\top A\,\mathbf{x}=A^\top\mathbf{b}$ (Chapter 1.10 of the Linear Algebra guide). In the example, $A^\top A=\begin{bmatrix}2&1\\1&5\end{bmatrix}$, $A^\top\mathbf{b}=[6,6]^\top$, and the solution is $\mathbf{x}^\star=[8/3,\ 2/3]^\top$ with loss $1$. One-variable cousin: $\big((ax-b)^2\big)'=2a(ax-b)$.
Why do we need it?
It is the gradient of the most-used loss in all of data science: squared error of a linear model. It tells us which way to move the weights to make the predictions better.
Where is it used?
Linear regression, ridge regression (add $2\lambda\mathbf{x}$), the last layer of many networks, signal reconstruction, and the Gauss–Newton method for nonlinear least squares.
How is it used?
Compute the residual $\mathbf{r}=A\mathbf{x}-\mathbf{b}$, then the gradient $2A^\top\mathbf{r}$, then step $\mathbf{x}\leftarrow\mathbf{x}-\eta\,2A^\top\mathbf{r}$. Or solve $A^\top A\mathbf{x}=A^\top\mathbf{b}$ directly.
- The factor $2$ is real. Many books define the loss as $\tfrac12\|A\mathbf{x}-\mathbf{b}\|^2$ so that the gradient is $A^\top(A\mathbf{x}-\mathbf{b})$ with no $2$. Check which one you have.
- $A^\top$ (not $A$) multiplies the residual. The shape check catches this: $A(A\mathbf{x}-\mathbf{b})$ does not fit unless $A$ is square.
Quick check: what is the gradient of the ridge loss $\|A\mathbf{x}-\mathbf{b}\|^2+\lambda\|\mathbf{x}\|^2$?
Add the two gradients: $2A^\top(A\mathbf{x}-\mathbf{b})+2\lambda\mathbf{x}$. Setting it to zero gives $(A^\top A+\lambda I)\mathbf{x}=A^\top\mathbf{b}$.
Important derivatives 6 and 7: $\log\mathbf{x}$ and $e^{\mathbf{x}}$ (elementwise) core
In one variable: $(\ln x)'=1/x$ and $(e^x)'=e^x$. In ML we constantly apply such a function to every entry of a vector: $\log\mathbf{x}=[\ln x_1,\ \ln x_2,\dots]$. Each output depends on only the input in its own position. Turning knob $x_2$ moves meter $2$ and leaves all other meters alone.
So most slopes are zero. The only non-zero slopes sit on the diagonal: the Jacobian is a diagonal matrix. This is very good news: no need to build a big matrix, just multiply entry by entry.
- $\mathbf{x}=[1,2,4]^\top$. Then $\ln\mathbf{x}=[0,\ 0.693,\ 1.386]^\top$ and $J=\operatorname{diag}(1/x_i)=\operatorname{diag}(1,\ 0.5,\ 0.25)$.
- $e^{\mathbf{x}}$ for $\mathbf{x}=[0,1,2]^\top$: $[1,\ 2.718,\ 7.389]^\top$ and $J=\operatorname{diag}(1,\ 2.718,\ 7.389)$.
- Scalar loss. $f(\mathbf{x})=\sum_i\ln x_i$ (like a log-likelihood of independent pieces). Then $\partial f/\partial x_k=1/x_k$, so $\nabla f=[1/x_1,\dots,1/x_n]^\top$. At $[1,2,4]$: $[1,\ 0.5,\ 0.25]^\top$.
- Log-sum-exp $f(\mathbf{x})=\ln\sum_je^{x_j}$. Then $\partial f/\partial x_k=\dfrac{e^{x_k}}{\sum_je^{x_j}}$, which is entry $k$ of softmax. At $\mathbf{x}=[0,1,2]$: $\nabla f\approx[0.090,\ 0.245,\ 0.665]^\top$ (these add to 1).
Let $g$ be a scalar function applied to each entry: $\mathbf{y}=g(\mathbf{x})$, meaning $y_i=g(x_i)$. Write $\circ$ (or $\odot$) for the entrywise product. Then
$$J=\operatorname{diag}\big(g'(x_1),\dots,g'(x_n)\big),\qquad d\mathbf{y}=g'(\mathbf{x})\circ d\mathbf{x}.$$- $\log\mathbf{x}$: $J=\operatorname{diag}(1/\mathbf{x})$ (needs every $x_i>0$). $\quad e^{\mathbf{x}}$: $J=\operatorname{diag}(e^{\mathbf{x}})$.
- Backward pass: a gradient $\mathbf{v}$ arriving at the output becomes $J^\top\mathbf{v}=g'(\mathbf{x})\circ\mathbf{v}$ at the input. No matrix needed: $n$ multiplications instead of $n^2$.
- $\nabla\sum_i\ln x_i=1/\mathbf{x}$, $\quad\nabla\sum_ie^{x_i}=e^{\mathbf{x}}$, $\quad\nabla\ln\sum_je^{x_j}=\operatorname{softmax}(\mathbf{x})$.
Why do we need it?
Activation functions, log-likelihoods and probabilities all apply $\log$, $\exp$ or another scalar function entry by entry. Their derivatives must be cheap, and the diagonal structure makes them so.
Where is it used?
Every activation function (ReLU, sigmoid, tanh) in a neural network, the log in cross-entropy and log-likelihood, and the softmax (log-sum-exp) at the end of every classifier.
How is it used?
In backpropagation, multiply the incoming gradient entry by entry with $g'(\mathbf{x})$. Frameworks do exactly this and never build the diagonal matrix.
- $\log\mathbf{x}$ (a vector) is not a scalar. Its derivative is a matrix (diagonal), not "$1/\mathbf{x}$". Only for a scalar sum like $\sum\ln x_i$ is the gradient the vector $1/\mathbf{x}$.
- $\log$ needs positive inputs. Real code works with $\log(\text{softmax})$ in one combined, stable function, never $\ln$ of a tiny number.
- Do not confuse $e^{\mathbf{x}}$ (entrywise) with the matrix exponential $e^{A}$.
Quick check: what is $\nabla_{\mathbf{x}}\sum_i x_i\ln x_i$ (the negative entropy)?
For each entry, $\frac{d}{dx}(x\ln x)=\ln x+1$ (product rule). So $\nabla f=\ln\mathbf{x}+\mathbf{1}$.
Check any derivative numerically core
How do you know a derivative you derived by hand is right? Test it with brute force. Nudge each input up a tiny bit and down a tiny bit, and measure the slope from the two function values. If that matches your formula, you are almost surely right. It costs two function evaluations per input, so it is far too slow for training. But it is perfect for catching mistakes.
$f(\mathbf{x})=\mathbf{x}^\top\mathbf{x}$ at $\mathbf{x}=[3,4]^\top$, with $h=0.001$.
- Nudge $x_1$: $f(3.001,4)=9.006001+16=25.006001$ and $f(2.999,4)=8.994001+16=24.994001$.
- Slope: $(25.006001-24.994001)/(2\times0.001)=0.012/0.002=6$. The formula $2x_1=6$. ✓
- Same for $x_2$: $8$ ✓. For a quadratic like this, the two-sided difference is exact (apart from round-off).
The central finite difference for each input $x_k$ (with $\mathbf{e}_k$ the vector with a 1 in position $k$):
$$\frac{\partial f}{\partial x_k}\approx\frac{f(\mathbf{x}+h\,\mathbf{e}_k)-f(\mathbf{x}-h\,\mathbf{e}_k)}{2h}.$$The error shrinks like $h^2$ as $h$ gets smaller, until round-off in the computer's arithmetic takes over. A good all-round value is $h\approx10^{-5}$. Compare with the formula using the relative error $\dfrac{\|\text{formula}-\text{numerical}\|}{\max(1,\|\text{formula}\|)}$. Values around $10^{-6}$ or smaller mean "correct". Values around $10^{-2}$ or bigger mean "there is a bug".
Why do we need it?
Hand-derived gradients are easy to get wrong by a transpose, a missing 2 or a sign. A numerical check catches all of these in seconds.
Where is it used?
"Gradient checking" when writing a custom loss or layer, unit tests in ML libraries (for example PyTorch's gradcheck), and verifying hand-written backpropagation.
How is it used?
Pick a random point, compute the analytic gradient and the finite-difference gradient, and compare. Always test at a random point, never at $\mathbf{x}=\mathbf{0}$, where many wrong formulas happen to give 0.
- A numerical check is a test, not a training method: it needs $2n$ function evaluations for $n$ inputs.
- Check at a random point and with non-square, non-symmetric matrices. Special values hide bugs.
- At a kink (like $|x|$ at $0$ or ReLU at $0$) the numerical slope depends on $h$: that is not a bug in your formula (see Chapter 2.7 on subgradients).
Quick check: your formula and the numerical gradient differ by exactly a factor of 2 in every entry. What is the likely mistake?
A missing or extra factor $2$: for instance using $A^\top(A\mathbf{x}-\mathbf{b})$ for $\|A\mathbf{x}-\mathbf{b}\|^2$ (correct for $\tfrac12\|\cdot\|^2$), or the other way round.
Recap, cheat sheet and practice
- A derivative has one slope per (output, input) pair. Entries = outputs × inputs. Scalar out: the gradient has the shape of the input. Vector to vector: the $m\times n$ Jacobian.
- Conventions here: $\nabla f$ is a column; $J_F$ is $m\times n$; $H$ is symmetric $n\times n$; $\nabla_Xf$ has the shape of $X$. Other books may use the transposed layout.
- Differentials replace derivatives: $d(XY)=dX\,Y+X\,dY$, $d(X^\top)=(dX)^\top$, $d(X^{-1})=-X^{-1}dX\,X^{-1}$, $d\operatorname{tr}X=\operatorname{tr}dX$. Keep the order of factors.
- Trace trick: $\operatorname{tr}(AB)=\operatorname{tr}(BA)$, a scalar is its own trace, and $df=\operatorname{tr}(G^\top dX)\Rightarrow\nabla_Xf=G$.
- Recipe: write $df$, rearrange so that $dX$ (or $d\mathbf{x}$) is at the right end, read off the coefficient, transpose if needed, then check the shape and a numerical gradient.
- Results: $\nabla\mathbf{x}^\top\mathbf{x}=2\mathbf{x}$; $J_{A\mathbf{x}}=A$; $\nabla\mathbf{a}^\top\mathbf{x}=\mathbf{a}$; $\nabla\mathbf{x}^\top A\mathbf{x}=(A+A^\top)\mathbf{x}$; $\nabla\|A\mathbf{x}-\mathbf{b}\|^2=2A^\top(A\mathbf{x}-\mathbf{b})$; elementwise $\log$ and $\exp$ have diagonal Jacobians.
Cheat sheet
| Function | Derivative | How we got it |
|---|---|---|
| $\mathbf{x}^\top\mathbf{x}=\|\mathbf{x}\|^2$ | $\nabla=2\mathbf{x}$ | $df=2\mathbf{x}^\top d\mathbf{x}$ |
| $A\mathbf{x}$ | $J=A$ | $d\mathbf{y}=A\,d\mathbf{x}$ |
| $\mathbf{a}^\top\mathbf{x}$ | $\nabla=\mathbf{a}$ | $df=\mathbf{a}^\top d\mathbf{x}$ |
| $\mathbf{x}^\top A\mathbf{x}$ | $\nabla=(A+A^\top)\mathbf{x}$; $H=A+A^\top$ | product rule, then transpose the scalar term |
| $\|A\mathbf{x}-\mathbf{b}\|^2$ | $\nabla=2A^\top(A\mathbf{x}-\mathbf{b})$ | $df=2\mathbf{r}^\top A\,d\mathbf{x}$, $\mathbf{r}=A\mathbf{x}-\mathbf{b}$ |
| $\log\mathbf{x}$, $e^{\mathbf{x}}$ (entrywise) | $J=\operatorname{diag}(1/\mathbf{x})$, $\operatorname{diag}(e^{\mathbf{x}})$ | each output uses only its own input |
| $\ln\sum_je^{x_j}$ | $\nabla=\operatorname{softmax}(\mathbf{x})$ | chain rule on $\ln S$ |
| $\operatorname{tr}(AX)$ | $\nabla_X=A^\top$ | $df=\operatorname{tr}(A\,dX)$ |
| $\mathbf{a}^\top X\mathbf{b}$ | $\nabla_X=\mathbf{a}\mathbf{b}^\top$ | rotate $\mathbf{b}$ to the front of the trace |
import numpy as np
def num_grad(f, x, h=1e-5):
"""Central-difference gradient of a scalar function f at the vector x."""
g = np.zeros_like(x)
for k in range(x.size):
e = np.zeros_like(x); e[k] = h
g[k] = (f(x + e) - f(x - e)) / (2 * h)
return g
# 1) x^T x -> 2x
x = np.array([3.0, 4.0])
print(num_grad(lambda v: v @ v, x)) # [6. 8.] same as 2*x
# 2) x^T A x -> (A + A^T) x (A is NOT symmetric)
A = np.array([[2.0, 1.0], [0.0, 3.0]])
x = np.array([1.0, 2.0])
print((A + A.T) @ x) # [ 6. 13.] the right formula
print(2 * A @ x) # [ 8. 12.] the naive (wrong) formula
print(num_grad(lambda v: v @ A @ v, x)) # [ 6. 13.] numbers agree with (A + A^T) x
# 3) ||Ax - b||^2 -> 2 A^T (Ax - b)
A = np.array([[1.0, 0.0], [0.0, 2.0], [1.0, 1.0]])
b = np.array([2.0, 1.0, 4.0])
x = np.array([1.0, 1.0])
r = A @ x - b
print(r, r @ r) # [-1. 1. -2.] 6.0
print(2 * A.T @ r) # [-6. 0.]
print(num_grad(lambda v: np.sum((A @ v - b) ** 2), x)) # [-6. 0.]
print(np.linalg.solve(A.T @ A, A.T @ b)) # [2.66666667 0.66666667] = [8/3, 2/3], the normal equations
# 4) Jacobian of an elementwise function is diagonal
x = np.array([1.0, 2.0, 4.0])
print(np.diag(1 / x)) # J of log(x): diag(1, 0.5, 0.25)
v = np.array([1.0, 1.0, 1.0])
print(np.diag(1 / x).T @ v, v / x) # [1. 0.5 0.25] twice: the entrywise product is enough
# 5) The trace trick: tr(AB) = tr(BA) even when the sizes differ
A = np.array([[1, 2, 0], [0, 1, 3]]); B = np.array([[1, 0], [2, 1], [0, 1]])
print(np.trace(A @ B), np.trace(B @ A)) # 9 9
# 6) Gradient with respect to a matrix: f(X) = tr(AX) has gradient A^T (the shape of X)
A = np.array([[1.0, 2.0], [3.0, 4.0]]); X = np.array([[0.5, -1.0], [2.0, 1.0]])
G = np.zeros_like(X)
for i in range(2):
for j in range(2):
E = np.zeros_like(X); E[i, j] = 1e-5
G[i, j] = (np.trace(A @ (X + E)) - np.trace(A @ (X - E))) / 2e-5
print(np.round(G, 6)) # [[1. 3.] [2. 4.]]
print(A.T) # [[1. 3.] [2. 4.]] same
1. $F:\mathbb{R}^5\to\mathbb{R}^3$. In this guide, what is the shape of its Jacobian?
2. What is $\nabla_{\mathbf{x}}(\mathbf{x}^\top A\mathbf{x})$ when $A$ is not symmetric?
3. $f(\mathbf{x})=\|A\mathbf{x}-\mathbf{b}\|^2$ with $A$ of size $3\times4$. Which gradient has a valid shape?
4. Which of these is true for matrices $A$ ($2\times3$) and $B$ ($3\times2$)?
5. What is the Jacobian of $\mathbf{y}=e^{\mathbf{x}}$ (entrywise) for $\mathbf{x}\in\mathbb{R}^3$?
6. $d(X^{-1})$ equals…
Practice problems
A. Find $\nabla f$ for $f(\mathbf{x})=\mathbf{c}^\top\mathbf{x}+\mathbf{x}^\top\mathbf{x}$ at $\mathbf{x}=[1,2]^\top$ with $\mathbf{c}=[3,-1]^\top$, using differentials.
$df=\mathbf{c}^\top d\mathbf{x}+2\mathbf{x}^\top d\mathbf{x}=(\mathbf{c}+2\mathbf{x})^\top d\mathbf{x}$, so $\nabla f=\mathbf{c}+2\mathbf{x}=[3,-1]+[2,4]=[5,\,3]^\top$. Check: $f=3x_1-x_2+x_1^2+x_2^2$ gives $\partial_1=3+2x_1=5$, $\partial_2=-1+2x_2=3$ ✓.
B. For $A=\begin{bmatrix}1&2\\3&4\end{bmatrix}$ and $\mathbf{x}=[1,1]^\top$, compute $\nabla(\mathbf{x}^\top A\mathbf{x})$ and compare with $2A\mathbf{x}$.
$A+A^\top=\begin{bmatrix}2&5\\5&8\end{bmatrix}$, so $\nabla=[7,13]^\top$. And $2A\mathbf{x}=2[3,7]^\top=[6,14]^\top$, which is different (since $A$ is not symmetric). Check: $f=x_1^2+5x_1x_2+4x_2^2$; $\partial_1=2x_1+5x_2=7$, $\partial_2=5x_1+8x_2=13$ ✓.
C. Derive $\nabla_{\mathbf{w}}\|X\mathbf{w}-\mathbf{y}\|^2+\lambda\|\mathbf{w}\|^2$ and the solution of the ridge equations, for a data matrix $X$.
With $\mathbf{r}=X\mathbf{w}-\mathbf{y}$, $df=2\mathbf{r}^\top X\,d\mathbf{w}+2\lambda\mathbf{w}^\top d\mathbf{w}$. So $\nabla=2X^\top(X\mathbf{w}-\mathbf{y})+2\lambda\mathbf{w}$. Setting it to $\mathbf{0}$: $(X^\top X+\lambda I)\mathbf{w}=X^\top\mathbf{y}$. Shape check: both sides are $n\times1$.
D. Use the recipe to find $\nabla_X\operatorname{tr}(X^\top AX)$ for a square $A$, in terms of $X$ and $A$.
$df=\operatorname{tr}(dX^\top AX)+\operatorname{tr}(X^\top A\,dX)$. The first term equals its transpose $\operatorname{tr}(X^\top A^\top dX)$. So $df=\operatorname{tr}\big(X^\top(A^\top+A)\,dX\big)$, and $G^\top=X^\top(A+A^\top)$, giving $\nabla_X=(A+A^\top)X$. (Chapter 2.7 revisits this.)
E. $\mathbf{y}=\ln\mathbf{x}$ and $L=\mathbf{c}^\top\mathbf{y}$ for $\mathbf{x}\in\mathbb{R}^3$. Find $\nabla_{\mathbf{x}}L$.
$L=\sum_ic_i\ln x_i$, so $\partial L/\partial x_k=c_k/x_k$ and $\nabla L=\mathbf{c}/\mathbf{x}$ (entrywise). Via Jacobians: $J^\top\mathbf{c}=\operatorname{diag}(1/\mathbf{x})\,\mathbf{c}$, the same thing.
F. Shape check: $W$ is $4\times6$, $\mathbf{x}\in\mathbb{R}^6$, $\mathbf{g}=\nabla_{\mathbf{z}}L\in\mathbb{R}^4$ where $\mathbf{z}=W\mathbf{x}$. Show that $\nabla_WL=\mathbf{g}\mathbf{x}^\top$ and $\nabla_{\mathbf{x}}L=W^\top\mathbf{g}$.
$d\mathbf{z}=dW\,\mathbf{x}+W\,d\mathbf{x}$ and $dL=\mathbf{g}^\top d\mathbf{z}$. With $\mathbf{x}$ fixed: $dL=\mathbf{g}^\top dW\,\mathbf{x}=\operatorname{tr}(\mathbf{x}\mathbf{g}^\top dW)$, so $G^\top=\mathbf{x}\mathbf{g}^\top$, $G=\mathbf{g}\mathbf{x}^\top$: $(4\times1)(1\times6)=4\times6$ ✓. With $W$ fixed: $dL=\mathbf{g}^\top W\,d\mathbf{x}$, so $\nabla_{\mathbf{x}}L=W^\top\mathbf{g}$ ($6\times4$ times $4\times1$) ✓.
Useful Gradient Identities
There are about a dozen gradient formulas that show up again and again in machine learning. This chapter gives you all of them, but with a twist: you are not asked to memorise a single one. Each comes with a derivation you can follow, and all of them come from one reusable recipe.
- Learn the derivation recipe: write the differential, rearrange into $\operatorname{tr}(G^\top dX)$, read off $G$
- Derive the gradients of linear, quadratic and transpose expressions, and of trace expressions $\operatorname{tr}(AX)$, $\operatorname{tr}(X^\top AX)$, $\operatorname{tr}(AXB)$
- See what symmetry buys you, and the gradients of vector norms ($L_2$, $L_1$, $L_\infty$) and matrix norms (Frobenius; nuclear and spectral for awareness)
- Meet $\log\det X$ and $d(X^{-1})$, and the chain-rule identities $\nabla_{\mathbf{x}}f(A\mathbf{x})=A^\top\nabla f$
- Finish with a one-page identity table and a routine to check any gradient numerically
This chapter uses the tools of Chapter 2.6: differentials, the trace trick, the gradient/Jacobian conventions (gradient = column; Jacobian $m\times n$; $\nabla_Xf$ has the shape of $X$), and shape checks. The Linear Algebra guide has a shorter table: Matrix calculus in the Linear Algebra guide. Here, every row is derived, and we add the trace, norm, log-determinant and chain-rule families.
Derive, don't memorise: the recipe core
Imagine you had to remember the multiplication table up to $1000\times1000$. Impossible. Instead you learn how to multiply, and then any product is a few steps away. Gradient identities are the same. There are dozens of them, but they all come out of one short procedure.
The procedure asks the function a question: "if I nudge my input a tiny bit, how much do you change?" Whatever the function says, we tidy the answer until the nudge stands alone at the end. The coefficient in front of the nudge is the gradient.
A first run of the recipe on a function we have not seen yet: $f(W)=\tfrac12\|W\mathbf{x}-\mathbf{y}\|^2$, where $W$ is $m\times n$, $\mathbf{x}\in\mathbb{R}^n$ and $\mathbf{y}\in\mathbb{R}^m$ are fixed.
- Differential. Let $\mathbf{r}=W\mathbf{x}-\mathbf{y}$. Then $d\mathbf{r}=dW\,\mathbf{x}$, and $df=\mathbf{r}^\top d\mathbf{r}=\mathbf{r}^\top dW\,\mathbf{x}$.
- Rearrange. This is a scalar, so wrap it in a trace and rotate $\mathbf{x}$ to the front: $df=\operatorname{tr}(\mathbf{x}\,\mathbf{r}^\top dW)$.
- Read off. $G^\top=\mathbf{x}\mathbf{r}^\top$, so $\nabla_Wf=\mathbf{r}\,\mathbf{x}^\top=(W\mathbf{x}-\mathbf{y})\mathbf{x}^\top$, which is $m\times n$, the shape of $W$ ✓.
This is exactly the gradient for the weights of a linear layer with squared error. We never looked it up.
The recipe for a scalar function $f$ of a vector $\mathbf{x}$ or matrix $X$:
- Differential. Write $df$. Constants do not move. Use $d(XY)=dX\,Y+X\,dY$, $d(X^\top)=(dX)^\top$, $d(X^{-1})=-X^{-1}dX\,X^{-1}$, $d\operatorname{tr}=\operatorname{tr}\,d$. Keep factor order.
- Rearrange. Get every term into the form $\operatorname{tr}(M\,dX)$ (or $\mathbf{g}^\top d\mathbf{x}$), using the toolkit below.
- Read off. $df=\mathbf{g}^\top d\mathbf{x}\Rightarrow\nabla f=\mathbf{g}$. $\ df=\operatorname{tr}(M\,dX)\Rightarrow\nabla_Xf=M^\top$.
- Check. Shape of the answer = shape of the variable. Then a numerical check.
The rearrangement toolkit (only four moves):
| Move | Rule | Use it to… |
|---|---|---|
| wrap a scalar | $s=\operatorname{tr}(s)$ | turn a number into something rotatable |
| rotate | $\operatorname{tr}(PQ)=\operatorname{tr}(QP)$ | move $dX$ to the right end |
| transpose a scalar or trace | $s=s^\top$, $\operatorname{tr}(P)=\operatorname{tr}(P^\top)$, $(PQ)^\top=Q^\top P^\top$ | flip a $dX^\top$ into a $dX$ |
| collect | $\operatorname{tr}(P\,dX)+\operatorname{tr}(Q\,dX)=\operatorname{tr}((P+Q)\,dX)$ | merge terms |
Why do we need it?
Memorised tables fail the first time your loss is slightly different. A procedure that you can run on paper works on any expression made of products, transposes, traces and inverses.
Where is it used?
Writing the backward pass of a new layer, deriving the update of an EM or matrix-factorisation algorithm, checking the output of autograd, and reading the derivations in papers on Gaussian processes, PCA and variational inference.
How is it used?
Do the four steps on paper, then confirm with a finite-difference check (the last section of this chapter). When the two agree, you can trust the formula.
- The recipe only works on scalar functions. For a vector-valued function, use $d\mathbf{y}=M\,d\mathbf{x}\Rightarrow J=M$.
- Do not skip the final transpose ($\operatorname{tr}(M\,dX)\Rightarrow M^\top$). A shape check usually catches a forgotten transpose when $X$ is not square.
- You may rotate a trace, but you may not swap two factors inside it.
Quick check: run the recipe on $f(\mathbf{x})=\mathbf{x}^\top B\mathbf{x}$ and say after which step the answer appears.
$df=(d\mathbf{x})^\top B\mathbf{x}+\mathbf{x}^\top B\,d\mathbf{x}$. Transpose the first term (a scalar): $\mathbf{x}^\top B^\top d\mathbf{x}$. Collect: $df=\mathbf{x}^\top(B+B^\top)\,d\mathbf{x}$. Read off: $\nabla f=(B+B^\top)\mathbf{x}$, the quadratic-form identity of Chapter 2.6.
Derivative of linear functions core
A linear function is the easiest thing to differentiate: its slope is the same everywhere, like a ramp with a constant incline. Whatever the input (a number, a vector, a matrix), the answer is just "the weights".
A linear function of a matrix $X$ is a weighted sum of its entries: $f(X)=\sum_{i,j}C_{ij}X_{ij}$. The slope for entry $X_{ij}$ is its weight $C_{ij}$. So the gradient is the weight matrix $C$ itself. Everything else in this family is a disguise of that one fact.
- Vector. $f(\mathbf{x})=3x_1-x_2+2x_3$: gradient $[3,-1,2]^\top=\mathbf{a}$ for $f=\mathbf{a}^\top\mathbf{x}$.
- Weighted sum of a matrix. $f(X)=1\cdot X_{11}+3\cdot X_{12}+2\cdot X_{21}+4\cdot X_{22}$. The gradient is $\begin{bmatrix}1&3\\2&4\end{bmatrix}=C$, and $f=\operatorname{tr}(C^\top X)$.
- A trace in disguise. $\operatorname{tr}(AX)$ with $A=\begin{bmatrix}1&2\\3&4\end{bmatrix}$ equals $A_{11}X_{11}+A_{12}X_{21}+A_{21}X_{12}+A_{22}X_{22}=X_{11}+2X_{21}+3X_{12}+4X_{22}$. The weights, placed in the grid of $X$, form $\begin{bmatrix}1&3\\2&4\end{bmatrix}=A^\top$ ✓.
- Bilinear. $f(X)=\mathbf{a}^\top X\mathbf{b}=\sum_{ij}a_iX_{ij}b_j$: the weight of $X_{ij}$ is $a_ib_j$, so the gradient is $\mathbf{a}\mathbf{b}^\top$.
With constant $\mathbf{a},\mathbf{b},\mathbf{c}$ and matrices $A,B,C$ of fitting shapes:
| Function | Derivative | Reason |
|---|---|---|
| $\mathbf{a}^\top\mathbf{x}$, also $\mathbf{x}^\top\mathbf{a}$ | $\nabla=\mathbf{a}$ | $df=\mathbf{a}^\top d\mathbf{x}$ |
| $A\mathbf{x}+\mathbf{c}$ | $J=A$ | $d\mathbf{y}=A\,d\mathbf{x}$ (the constant drops out) |
| $\operatorname{tr}(C^\top X)=\sum C_{ij}X_{ij}$ | $\nabla_X=C$ | weights of the entries |
| $\operatorname{tr}(AX)$ | $\nabla_X=A^\top$ | $df=\operatorname{tr}(A\,dX)$ |
| $\mathbf{a}^\top X\mathbf{b}$ | $\nabla_X=\mathbf{a}\mathbf{b}^\top$ | $df=\operatorname{tr}(\mathbf{b}\mathbf{a}^\top dX)$ |
| $\operatorname{tr}(X)$ | $\nabla_X=I$ | $df=\operatorname{tr}(I\,dX)$ |
For a linear function the gradient does not depend on where you stand: the Hessian is zero.
Why do we need it?
Linear pieces are in every model: a prediction $\mathbf{w}^\top\mathbf{x}+b$, a layer $W\mathbf{x}+\mathbf{b}$. They are the building blocks that all bigger gradients are assembled from.
Where is it used?
Linear and logistic regression scores, every dense layer of a neural network, the output projection in a Transformer, and the "linear term" of every quadratic model.
How is it used?
Spot the weights: whatever multiplies the variable (arranged in the shape of the variable) is the gradient. Use the table as a quick lookup, but be able to rebuild each row with the recipe.
- $\nabla_X\operatorname{tr}(AX)=A^\top$ but $\nabla_X\operatorname{tr}(A^\top X)=A$. The transpose depends on how the trace is written. When in doubt, count: the weight of $X_{ij}$.
- Linear in $X$ does not mean the gradient is the same matrix for any shape: use the shape check ($\mathbf{a}\mathbf{b}^\top$ must have the shape of $X$).
Quick check: $X$ is $2\times3$, $\mathbf{a}\in\mathbb{R}^2$, $\mathbf{b}\in\mathbb{R}^3$. What is $\nabla_X(\mathbf{a}^\top X\mathbf{b})$ and what is its shape?
$\mathbf{a}\mathbf{b}^\top$, shape $(2\times1)(1\times3)=2\times3$, matching $X$. Entry $(i,j)$ is $a_ib_j$.
Derivative of quadratic functions core
In one variable a quadratic is $ax^2+bx+c$: a parabola. Its slope is $2ax+b$ and it has one flat spot, at $x=-b/(2a)$.
The vector version has the same three ingredients: a bowl (or dome, or saddle) $\tfrac12\mathbf{x}^\top A\mathbf{x}$, a ramp $\mathbf{b}^\top\mathbf{x}$ that tilts it, and a constant lift $c$. The ramp slides the flat spot away from the origin. Finding where the gradient is zero tells you where the bottom of the bowl is.
$A=\begin{bmatrix}2&0\\0&4\end{bmatrix}$, $\mathbf{b}=[-2,-4]^\top$, $c=0$:
- $f(\mathbf{x})=\tfrac12\mathbf{x}^\top A\mathbf{x}+\mathbf{b}^\top\mathbf{x}=x_1^2+2x_2^2-2x_1-4x_2$.
- Gradient (school way): $[2x_1-2,\ \ 4x_2-4]$. Formula: $A\mathbf{x}+\mathbf{b}=[2x_1-2,\ 4x_2-4]$ ✓.
- Set it to zero: $A\mathbf{x}=-\mathbf{b}$, so $\mathbf{x}^\star=[1,1]^\top$ and $f(\mathbf{x}^\star)=1+2-2-4=-3$.
- Completing the square: $f=\tfrac12(\mathbf{x}-\mathbf{x}^\star)^\top A(\mathbf{x}-\mathbf{x}^\star)-3=(x_1-1)^2+2(x_2-1)^2-3$. Expanding gives back $x_1^2+2x_2^2-2x_1-4x_2$ ✓.
For $f(\mathbf{x})=\tfrac12\mathbf{x}^\top A\mathbf{x}+\mathbf{b}^\top\mathbf{x}+c$:
$$\nabla f=\tfrac12(A+A^\top)\mathbf{x}+\mathbf{b}\ \ \Big(=A\mathbf{x}+\mathbf{b}\ \text{ if }A=A^\top\Big),\qquad \nabla^2f=\tfrac12(A+A^\top).$$Without the $\tfrac12$: $\nabla(\mathbf{x}^\top A\mathbf{x}+\mathbf{b}^\top\mathbf{x})=(A+A^\top)\mathbf{x}+\mathbf{b}$. For a symmetric invertible $A$, the flat spot is $\mathbf{x}^\star=-A^{-1}\mathbf{b}$: a minimum if $A$ is positive definite (a bowl), a maximum if negative definite, a saddle if indefinite.
Shifted form (a squared distance measured with a weight matrix, as in the Gaussian exponent): $\nabla\big[(\mathbf{x}-\mathbf{c})^\top A(\mathbf{x}-\mathbf{c})\big]=(A+A^\top)(\mathbf{x}-\mathbf{c})$, because $d(\mathbf{x}-\mathbf{c})=d\mathbf{x}$.
Why do we need it?
Near its minimum almost every smooth loss looks like a quadratic. Knowing the quadratic's gradient and flat spot lets us understand, and speed up, optimisation.
Where is it used?
Least squares and ridge regression, the Gaussian log-density (Mahalanobis distance), Newton's method (it solves a quadratic model exactly in one step), trust-region methods and the analysis of why gradient descent zig-zags.
How is it used?
Read off $A$ and $\mathbf{b}$, compute $\nabla f=A\mathbf{x}+\mathbf{b}$, and solve $A\mathbf{x}=-\mathbf{b}$ for the flat spot. Never form $A^{-1}$ explicitly in code: use a linear solver.
- Watch the $\tfrac12$. With $\tfrac12\mathbf{x}^\top A\mathbf{x}$ the gradient is $A\mathbf{x}$ (for symmetric $A$); without it, $2A\mathbf{x}$.
- If $A$ is not symmetric, replace $A$ by $\tfrac12(A+A^\top)$ everywhere (the symmetric-matrix section below explains why this loses nothing).
- $A\mathbf{x}=-\mathbf{b}$ has a unique solution only if $A$ is invertible. If $A$ is singular the surface is flat in some direction.
Quick check: $f(\mathbf{x})=\mathbf{x}^\top\mathbf{x}-6x_1$ in two variables. Where is its minimum?
Write $f=\tfrac12\mathbf{x}^\top(2I)\mathbf{x}+\mathbf{b}^\top\mathbf{x}$ with $\mathbf{b}=[-6,0]$. The gradient is $2\mathbf{x}+\mathbf{b}=[2x_1-6,\ 2x_2]$. It is zero at $\mathbf{x}^\star=[3,0]^\top$ and $f=9-18=-9$.
Derivatives involving a transpose
Transposing a table flips rows and columns. A nudge of every entry flips in the same way: nudging a matrix and then transposing it gives the same result as transposing first and then nudging. That is the rule $d(X^\top)=(dX)^\top$.
So a transpose inside the function usually just means a transposed pattern in the gradient. Two small tricks do all the work: a number equals its own transpose (so you can flip a whole term), and $(PQ)^\top=Q^\top P^\top$ (the order reverses).
$X$ is $2\times2$, $\mathbf{a}=[1,2]^\top$, $\mathbf{b}=[3,4]^\top$. Compare the two functions $\mathbf{a}^\top X\mathbf{b}$ and $\mathbf{a}^\top X^\top\mathbf{b}$.
- $\mathbf{a}^\top X\mathbf{b}=\sum_{ij}a_iX_{ij}b_j$: the weight of $X_{ij}$ is $a_ib_j$, so the gradient is $\mathbf{a}\mathbf{b}^\top=\begin{bmatrix}3&4\\6&8\end{bmatrix}$.
- $\mathbf{a}^\top X^\top\mathbf{b}=\sum_{ij}a_jX_{ij}b_i$ (since $(X^\top)_{ji}=X_{ij}$): the weight of $X_{ij}$ is $b_ia_j$, so the gradient is $\mathbf{b}\mathbf{a}^\top=\begin{bmatrix}3&6\\4&8\end{bmatrix}$.
- The two gradients are transposes of each other. Check with the number trick: $\mathbf{a}^\top X^\top\mathbf{b}$ is a number, so it equals its transpose $\mathbf{b}^\top X\mathbf{a}$, whose gradient is $\mathbf{b}\mathbf{a}^\top$ ✓.
- $d(X^\top)=(dX)^\top$, and $\operatorname{tr}(P^\top)=\operatorname{tr}(P)$.
- $\nabla_X(\mathbf{a}^\top X^\top\mathbf{b})=\mathbf{b}\mathbf{a}^\top$ (compare $\nabla_X(\mathbf{a}^\top X\mathbf{b})=\mathbf{a}\mathbf{b}^\top$).
- $\nabla_X\operatorname{tr}(X^\top A)=A$, but $\nabla_X\operatorname{tr}(XA)=A^\top$. Also $\nabla_X\operatorname{tr}(AX^\top)=A$.
- For vectors: $\nabla_{\mathbf{x}}(\mathbf{b}^\top A^\top\mathbf{x})=A\mathbf{b}$ (compare $\nabla_{\mathbf{x}}(\mathbf{b}^\top A\mathbf{x})=A^\top\mathbf{b}$). The Jacobian of $A^\top\mathbf{x}$ is $A^\top$.
- Flipping the argument. If $g(X)=f(X^\top)$, then $\nabla g(X)=\big(\nabla f(X^\top)\big)^\top$: the gradient of a function of the transpose is the transpose of the gradient.
Why do we need it?
Weight matrices appear both as $W$ and as $W^\top$ (forward and backward passes, tied embeddings). We must know how a transpose changes the gradient, or the shapes will not match.
Where is it used?
Backpropagation ($W^\top$ carries gradients backwards), weight tying in language models (the same matrix used as $E$ and $E^\top$), covariance expressions $X^\top X$ and $XX^\top$, and PCA.
How is it used?
Bring any $dX^\top$ to $dX$ by transposing the whole (scalar) term, then read off as usual. Then check the shape. A wrongly transposed gradient usually has the wrong shape for non-square $X$.
- Transposing a number changes nothing. Transposing a matrix changes its shape. Only use "a scalar equals its transpose" on genuinely $1\times1$ expressions.
- $(PQ)^\top=Q^\top P^\top$: the order reverses. Forgetting this is the most common error in this section.
Quick check: $\nabla_X\,\mathbf{u}^\top X\mathbf{v}$ and $\nabla_X\,\mathbf{v}^\top X^\top\mathbf{u}$. Are the answers equal?
The first is $\mathbf{u}\mathbf{v}^\top$. The second is a number equal to its transpose $\mathbf{u}^\top X\mathbf{v}$, so it is the same function, and its gradient is also $\mathbf{u}\mathbf{v}^\top$. Yes, equal.
Trace identities: $\operatorname{tr}(AX)$, $\operatorname{tr}(X^\top AX)$, $\operatorname{tr}(AXB)$ core
A trace of a matrix expression is just a long weighted sum of the entries of $X$, so its gradient is "the weights, in the shape of $X$".
- $\operatorname{tr}(AX)$ and $\operatorname{tr}(AXB)$ are linear in $X$ (a weighted sum).
- $\operatorname{tr}(X^\top AX)$ is quadratic. A nice way to see it: the columns of $X$ are vectors $\mathbf{x}_1,\mathbf{x}_2,\dots$, and $\operatorname{tr}(X^\top AX)=\sum_k\mathbf{x}_k^\top A\,\mathbf{x}_k$: a quadratic form for each column. By the vector rule, each column contributes $(A+A^\top)\mathbf{x}_k$, so the whole gradient is $(A+A^\top)X$.
$A=\begin{bmatrix}1&2\\0&3\end{bmatrix}$ and $X=\begin{bmatrix}1&0\\1&1\end{bmatrix}$, with columns $\mathbf{x}_1=[1,1]^\top$ and $\mathbf{x}_2=[0,1]^\top$.
- $\mathbf{x}_1^\top A\mathbf{x}_1=1+2+0+3=6$ and $\mathbf{x}_2^\top A\mathbf{x}_2=A_{22}=3$. So $\operatorname{tr}(X^\top AX)=9$.
- $A+A^\top=\begin{bmatrix}2&2\\2&6\end{bmatrix}$. Column 1: $(A+A^\top)\mathbf{x}_1=[4,8]^\top$. Column 2: $(A+A^\top)\mathbf{x}_2=[2,6]^\top$.
- So $\nabla_X\operatorname{tr}(X^\top AX)=\begin{bmatrix}4&2\\8&6\end{bmatrix}=(A+A^\top)X$ ✓.
- $\operatorname{tr}(AXB)$. With $A=\begin{bmatrix}1&0\\1&1\end{bmatrix}$ and $B=\begin{bmatrix}2&1\\0&1\end{bmatrix}$ the gradient is $A^\top B^\top=\begin{bmatrix}1&1\\0&1\end{bmatrix}\begin{bmatrix}2&0\\1&1\end{bmatrix}=\begin{bmatrix}3&1\\1&1\end{bmatrix}$.
| Function | $\nabla_X$ | Derivation (the recipe) |
|---|---|---|
| $\operatorname{tr}(AX)$ | $A^\top$ | $df=\operatorname{tr}(A\,dX)$ |
| $\operatorname{tr}(AXB)$ | $A^\top B^\top$ | $df=\operatorname{tr}(A\,dX\,B)=\operatorname{tr}(BA\,dX)$, so $G^\top=BA$ |
| $\operatorname{tr}(AX^\top B)$ | $BA$ | rotate: $\operatorname{tr}(AX^\top B)=\operatorname{tr}(BA\,X^\top)=\sum_{ij}(BA)_{ij}X_{ij}$, so the weights are $BA$ |
| $\operatorname{tr}(X^\top AX)$ | $(A+A^\top)X$ (or $2AX$ if $A=A^\top$) | $df=\operatorname{tr}(dX^\top AX)+\operatorname{tr}(X^\top A\,dX)$, transpose the first |
| $\operatorname{tr}(XAX^\top)$ | $X(A+A^\top)$ | see the recipe widget |
Shape check for all: the result has the shape of $X$. Rule of thumb: write the function as $\sum(\text{weight})\cdot X_{ij}$ and put the weights into the grid of $X$.
Why do we need it?
Many matrix objectives are naturally written as traces: a sum of quadratic forms over all columns, a sum of inner products, a covariance term. These identities give their gradients at once.
Where is it used?
PCA (maximise $\operatorname{tr}(W^\top\Sigma W)$), matrix factorisation and recommender systems, canonical correlation analysis, the Gaussian log-likelihood (the $\operatorname{tr}(\Sigma^{-1}S)$ term), and metric learning.
How is it used?
Write the objective as a trace, apply the matching row of the table (or run the recipe), and update $X\leftarrow X-\eta\,G$. Always check the shape of $G$ and compare with finite differences.
- $\nabla\operatorname{tr}(AXB)=A^\top B^\top$, not $AB$ and not $BA$. The order and both transposes matter. For non-square $A$, $B$ the shape check catches most slips.
- You may rotate $\operatorname{tr}(AXB)=\operatorname{tr}(BAX)$, but not $\operatorname{tr}(ABX)$.
- The "$(A+A^\top)X$" identity needs a square $A$. If $A$ is symmetric it reduces to $2AX$.
Quick check: $X$ is $3\times2$, $A$ is $3\times3$ and symmetric. What is $\nabla_X\operatorname{tr}(X^\top AX)$ and its shape?
$2AX$, shape $(3\times3)(3\times2)=3\times2$, the shape of $X$.
Symmetric-matrix identities core
A symmetric matrix is a table that looks the same when you flip it over its diagonal: $A=A^\top$. Many matrices in ML are symmetric: covariance matrices, Gram matrices $X^\top X$, Hessians, graph Laplacians.
Why does symmetry help? In $\mathbf{x}^\top A\mathbf{x}=\sum A_{ij}x_ix_j$ the product $x_ix_j$ is the same as $x_jx_i$. Only the sum $A_{ij}+A_{ji}$ matters. So any matrix $A$ can be replaced by its symmetric part $S=\tfrac12(A+A^\top)$ without changing the function at all. The leftover (the antisymmetric part) is invisible. After that replacement the gradient is the simple $2S\mathbf{x}$.
$A=\begin{bmatrix}1&3\\1&2\end{bmatrix}$ and $\mathbf{x}=[1,2]^\top$.
- Split: $S=\tfrac12(A+A^\top)=\begin{bmatrix}1&2\\2&2\end{bmatrix}$ and $K=\tfrac12(A-A^\top)=\begin{bmatrix}0&1\\-1&0\end{bmatrix}$. Check: $S+K=A$ ✓.
- $\mathbf{x}^\top A\mathbf{x}=1(1+6)+2(1+4)=17$. $\ \mathbf{x}^\top S\mathbf{x}=1(1+4)+2(2+4)=17$. $\ \mathbf{x}^\top K\mathbf{x}=1(2)+2(-1)=0$.
- So the antisymmetric part contributes nothing: $17=17+0$.
- Gradient: $(A+A^\top)\mathbf{x}=[10,12]^\top$ and $2S\mathbf{x}=2[5,6]^\top=[10,12]^\top$ ✓. But $2A\mathbf{x}=[14,10]^\top$ ✗.
Every square matrix splits uniquely as $A=S+K$ with $S=\tfrac12(A+A^\top)$ (symmetric) and $K=\tfrac12(A-A^\top)$ (antisymmetric: $K^\top=-K$).
Why $\mathbf{x}^\top K\mathbf{x}=0$: it is a number, so it equals its transpose: $\mathbf{x}^\top K\mathbf{x}=\mathbf{x}^\top K^\top\mathbf{x}=-\mathbf{x}^\top K\mathbf{x}$. A number equal to its own negative is $0$.
Identities for symmetric $A=A^\top$:
$$\nabla(\mathbf{x}^\top A\mathbf{x})=2A\mathbf{x},\quad \nabla^2(\mathbf{x}^\top A\mathbf{x})=2A,\quad \nabla_X\operatorname{tr}(X^\top AX)=2AX,\quad \nabla\|B\mathbf{x}\|^2=2B^\top B\mathbf{x}.$$The last uses that $B^\top B$ is symmetric for any matrix $B$. For a general (non-symmetric) $A$: use $S$ instead of $A$, i.e. $\nabla(\mathbf{x}^\top A\mathbf{x})=2S\mathbf{x}$.
Why do we need it?
It removes the awkward "$A+A^\top$" and gives the same simple form as one-variable calculus. It also tells us that when a quadratic form is learned from data, only the symmetric part is identifiable.
Where is it used?
Covariance and precision matrices in Gaussians, PCA ($\mathbf{w}^\top\Sigma\mathbf{w}$), the Gram matrix $X^\top X$ in regression, kernel methods, and Hessians (always symmetric).
How is it used?
Check whether the middle matrix is symmetric. If it is, use $2A\mathbf{x}$. If it is not, symmetrise it first: $S=\tfrac12(A+A^\top)$, or use $(A+A^\top)\mathbf{x}$.
- Symmetry of $A$ is an assumption you must check, not something you can assume. $W$ matrices in neural networks, for example, are usually not symmetric.
- Replacing $A$ by $S$ is only valid for the quadratic form $\mathbf{x}^\top A\mathbf{x}$ (same vector on both sides). It is not valid for $\mathbf{a}^\top A\mathbf{b}$ with two different vectors.
Quick check: $A=\begin{bmatrix}0&5\\-5&0\end{bmatrix}$ (antisymmetric). What is $\nabla(\mathbf{x}^\top A\mathbf{x})$?
$A+A^\top=\mathbf{0}$, so the gradient is $\mathbf{0}$ everywhere. Indeed $\mathbf{x}^\top A\mathbf{x}=5x_1x_2-5x_2x_1=0$ for every $\mathbf{x}$: the function is identically zero.
Vector norm derivatives: $L_2$, squared $L_2$, $L_1$, $L_\infty$ core
A norm measures size (Linear Algebra guide, Chapter 1.2). Its gradient answers: "which way should I move to make the size grow fastest?"
- $L_2$ (length). Walking straight away from the origin lengthens the vector at rate $1$. Walking around a circle does not change the length at all. So the gradient is the unit vector pointing away from the origin: $\mathbf{x}/\|\mathbf{x}\|$.
- Squared $L_2$. The squared length grows faster the further out you are: slope $2\times$ distance. The gradient is $2\mathbf{x}$.
- $L_1$ (taxi size). Each entry adds $|x_i|$. Each one has slope $+1$ on the right of $0$ and $-1$ on the left. At exactly $0$ there is a sharp corner (a kink): no single slope fits.
- $L_\infty$ (largest entry). Only the biggest entry counts, so the gradient is a single $\pm1$ in that position and $0$ elsewhere (with a kink when two entries tie).
- $L_2$: $\mathbf{x}=[3,4]^\top$. $\|\mathbf{x}\|=5$ and $\nabla\|\mathbf{x}\|=[0.6,\ 0.8]^\top$ (a unit vector). Check: nudging $x_1$ by $0.01$ gives $\|[3.01,4]\|=5.0060$, a change of $0.0060\approx0.6\times0.01$ ✓.
- Squared $L_2$: $\|\mathbf{x}\|^2=25$ and $\nabla=[6,8]^\top$ (length 10 = twice the distance).
- $L_1$: $\mathbf{x}=[3,-2,0]^\top$. $\|\mathbf{x}\|_1=5$. A subgradient is $[1,\ -1,\ s]$ for any $s\in[-1,1]$ (the entry at $0$ is a kink).
- $L_\infty$: $\mathbf{x}=[3,-4,1]^\top$. The largest absolute entry is $|-4|$, so $\|\mathbf{x}\|_\infty=4$ and the gradient is $[0,\ -1,\ 0]^\top$ (moving $x_2$ more negative increases the max).
At $\mathbf{x}=\mathbf{0}$ the $L_2$ norm has a cone-shaped tip, so its gradient is not defined there.
Subgradients (awareness). At a kink the ordinary derivative does not exist. For a convex function, a subgradient at $\mathbf{x}$ is any vector $\mathbf{g}$ such that the straight line (plane) $f(\mathbf{x})+\mathbf{g}^\top(\mathbf{y}-\mathbf{x})$ stays below the graph everywhere. The set of all of them is the subdifferential. For $|x|$ at $0$ it is the whole interval $[-1,1]$. In code one usually picks $0$: that is what np.sign(0) returns. For $L_\infty$ at a tie, any weighted average of the tied $\pm\mathbf{e}_i$ is a subgradient.
Other $L_p$ norms ($p\gt1$) are smooth away from $0$: $\nabla\|\mathbf{x}\|_p^p=p\,|\mathbf{x}|^{p-1}\circ\operatorname{sign}(\mathbf{x})$ (entrywise).
Why do we need it?
Penalties and constraints are written with norms. To train with them, or to normalise by them, we need their gradients, including the awkward non-smooth cases.
Where is it used?
L1 regularisation (Lasso) for sparse weights, L2 weight decay, gradient clipping by norm, normalisation layers (dividing by $\|\mathbf{x}\|$), cosine similarity, and adversarial attacks (FGSM steps along $\operatorname{sign}(\nabla)$, the best move inside an $L_\infty$ budget).
How is it used?
Add the gradient of the penalty to the loss gradient: $\lambda\operatorname{sign}(\mathbf{w})$ for L1, $2\lambda\mathbf{w}$ for squared L2. Frameworks use $0$ at the kink. Divide by $\|\mathbf{x}\|$ only after checking it is not (nearly) zero.
- $\nabla\|\mathbf{x}\|_2=\mathbf{x}/\|\mathbf{x}\|$ is not defined at $\mathbf{0}$. In code add a tiny constant, $\mathbf{x}/(\|\mathbf{x}\|+\varepsilon)$, or use $\|\mathbf{x}\|^2$ if you can.
- $\operatorname{sign}(\mathbf{x})$ is a subgradient, not "the" derivative: do not use it in a finite-difference check at points where an entry is exactly $0$.
- $L_1$ and $L_\infty$ gradients keep a constant size however close you are. Optimisers need shrinking steps (or a "proximal" step) to settle at the minimum.
Quick check: what is $\nabla\,\lambda\|\mathbf{w}\|_1$ at $\mathbf{w}=[2,-3,0,5]^\top$, using $\operatorname{sign}(0)=0$?
$\lambda\operatorname{sign}(\mathbf{w})=\lambda[1,-1,0,1]^\top$. Every non-zero weight is pushed towards $0$ by the same amount $\lambda$ per step, regardless of its size. (Squared L2 would push big weights harder.)
Matrix norm derivatives: Frobenius, and a glimpse of nuclear and spectral core
The Frobenius norm treats a matrix as one long vector: square every entry, add, take the square root. So everything we learned about the $L_2$ norm carries over, with "entries of $X$" in place of "entries of $\mathbf{x}$".
The two other matrix norms you will hear about are based on singular values (Chapter 1.13 of the Linear Algebra guide): the spectral norm is the largest singular value (the biggest stretch the matrix makes), and the nuclear norm is the sum of all singular values (a convex stand-in for the rank, just as $L_1$ is the stand-in for counting non-zero entries). They are the matrix versions of $\max$ and of $L_1$, and their gradients are built from the SVD.
- $X=\begin{bmatrix}1&2\\3&4\end{bmatrix}$. $\|X\|_F^2=1+4+9+16=30$. $\nabla\|X\|_F^2=2X=\begin{bmatrix}2&4\\6&8\end{bmatrix}$. And $\|X\|_F=\sqrt{30}\approx5.477$, with gradient $X/\|X\|_F$.
- Matrix least squares. $A=\begin{bmatrix}1&0\\0&2\end{bmatrix}$, $X=\begin{bmatrix}1&1\\1&1\end{bmatrix}$, $B=\begin{bmatrix}0&1\\1&0\end{bmatrix}$. Then $AX=\begin{bmatrix}1&1\\2&2\end{bmatrix}$ and the residual $R=AX-B=\begin{bmatrix}1&0\\1&2\end{bmatrix}$. Loss: $1+0+1+4=6$. Gradient: $2A^\top R=2\begin{bmatrix}1&0\\0&2\end{bmatrix}\begin{bmatrix}1&0\\1&2\end{bmatrix}=\begin{bmatrix}2&0\\4&8\end{bmatrix}$.
- Awareness. $X=\operatorname{diag}(3,1)$ has singular values $3,1$ and $U=V=I$. The spectral norm is $3$, with gradient $\mathbf{u}_1\mathbf{v}_1^\top=\begin{bmatrix}1&0\\0&0\end{bmatrix}$ (only the entry that sets the largest stretch matters). The nuclear norm is $3+1=4$ with gradient $UV^\top=I$. Check: nudging $X_{11}$ by $0.01$ changes $\|X\|_*=3.01+1$ by $0.01$ ✓ (slope $1$).
Awareness (SVD $X=U\Sigma V^\top$ with singular values $\sigma_1\ge\sigma_2\ge\cdots$). The standard first-order fact is $d\sigma_i=\mathbf{u}_i^\top dX\,\mathbf{v}_i$. From it:
$$\nabla\|X\|_2=\mathbf{u}_1\mathbf{v}_1^\top\ \ (\text{spectral norm }\sigma_1,\ \text{if }\sigma_1\gt\sigma_2),\qquad \nabla\|X\|_*=UV^\top=\sum_i\mathbf{u}_i\mathbf{v}_i^\top\ \ (\text{nuclear norm }\textstyle\sum\sigma_i,\ \text{if all }\sigma_i\gt0).$$$UV^\top$ is the matrix version of $\operatorname{sign}(x)$: it keeps the directions and sets every stretch to 1. Where singular values tie or hit $0$, these norms have kinks, and the formulas give just one subgradient.
Why do we need it?
Weights of layers are matrices, and we often want to penalise or control their size. Each matrix norm controls something different, and training needs its gradient.
Where is it used?
Frobenius: weight decay on matrices, multi-output regression $\|XW-Y\|_F^2$, matrix factorisation. Nuclear: low-rank matrix completion and recommender systems. Spectral: spectral normalisation of GAN discriminators and Lipschitz control of networks.
How is it used?
Frobenius: add $2\lambda W$ to the gradient. For the other two, use a library SVD: spectral normalisation estimates $\mathbf{u}_1,\mathbf{v}_1$ with a few power-iteration steps and divides $W$ by $\sigma_1$.
- $\nabla\|X\|_F^2=2X$, but $\nabla\|X\|_F=X/\|X\|_F$. As with vectors, do not mix the squared and un-squared versions.
- $\|AX-B\|_F^2$ gives $2A^\top(AX-B)$; for $\|XA-B\|_F^2$ the $A^\top$ moves to the right: $2(XA-B)A^\top$. The side where $A$ multiplies $X$ decides where $A^\top$ goes.
- The SVD formulas are only valid when the relevant singular values are distinct (spectral) or positive (nuclear). Otherwise they are subgradients.
Quick check: $\nabla_W\big(\|W\|_F^2\cdot\tfrac\lambda2\big)$ for a weight matrix $W$?
$\tfrac\lambda2\cdot2W=\lambda W$. This is weight decay applied to a whole matrix.
The log-determinant and the inverse (awareness) core
The determinant $\det X$ says by what factor the matrix $X$ scales volume (area in 2D; Chapter 1.7 of the Linear Algebra guide). Its logarithm turns products into sums, which is why $\log\det$ appears so often.
How does $\log\det X$ react to a nudge $dX$? In one variable: $d\ln x=dx/x=x^{-1}dx$, the nudge measured in units of $x$. The matrix version is the same: $d\log\det X=\operatorname{tr}(X^{-1}dX)$, the nudge measured in units of $X$. The inverse is the matrix version of "divide by".
$X=\begin{bmatrix}2&1\\1&3\end{bmatrix}$. $\det X=2\cdot3-1\cdot1=5$ and $X^{-1}=\tfrac15\begin{bmatrix}3&-1\\-1&2\end{bmatrix}=\begin{bmatrix}0.6&-0.2\\-0.2&0.4\end{bmatrix}$.
- $\log\det X=\ln(x_{11}x_{22}-x_{12}x_{21})$. By the chain rule, $\partial/\partial x_{11}=x_{22}/\det=3/5=0.6$; $\ \partial/\partial x_{12}=-x_{21}/\det=-0.2$; $\ \partial/\partial x_{21}=-x_{12}/\det=-0.2$; $\ \partial/\partial x_{22}=x_{11}/\det=0.4$.
- So the gradient is $\begin{bmatrix}0.6&-0.2\\-0.2&0.4\end{bmatrix}=X^{-\top}$ ✓ (here equal to $X^{-1}$ because $X$ is symmetric).
- The inverse. Nudge $x_{11}$ by $\varepsilon$: the exact inverse of $\begin{bmatrix}2+\varepsilon&1\\1&3\end{bmatrix}$ has top-left entry $\frac{3}{5+3\varepsilon}\approx0.6-0.36\,\varepsilon$. The formula $-X^{-1}\,dX\,X^{-1}$ with $dX=\varepsilon E_{11}$ gives $-\varepsilon\,(0.6)(0.6)=-0.36\,\varepsilon$ in that entry ✓.
Here $X^{-\top}$ means $(X^{-1})^\top$. For symmetric $X$ (a covariance matrix) all the transposes disappear.
Where the first formula comes from (awareness). For a tiny nudge $\varepsilon E$: $\det(X+\varepsilon E)=\det X\cdot\det(I+\varepsilon M)$ with $M=X^{-1}E$. The determinant of $I+\varepsilon M$ is the product of $(1+\varepsilon\lambda_i)$ over the eigenvalues $\lambda_i$ of $M$. Multiplying out and dropping $\varepsilon^2$ gives $1+\varepsilon\sum\lambda_i=1+\varepsilon\operatorname{tr}M$. So $d\det X=\det X\operatorname{tr}(X^{-1}E)\,\varepsilon$.
Why do we need it?
Gaussian probability densities contain $\log\det\Sigma$ and $\Sigma^{-1}$. To fit them, or to train models that use them, we need their gradients.
Where is it used?
Maximum-likelihood estimation of a covariance, Gaussian processes (the $\log\det K$ term), graphical lasso, normalising flows ($\log|\det J|$ of the Jacobian, Chapter 2.5), and Bayesian model evidence.
How is it used?
Use $\nabla\log\det X=X^{-\top}$ and $d(X^{-1})=-X^{-1}dX\,X^{-1}$ as building blocks in the recipe. In code, never invert: use a Cholesky factorisation and solve. Compute $\log\det$ as twice the sum of logs of the Cholesky diagonal.
- $\nabla\log\det X=X^{-\top}$ needs $X$ invertible ($\det X\ne0$). For the plain $\log$ one also needs $\det X\gt0$; with $\log|\det X|$ the formula holds for any invertible $X$.
- $d(X^{-1})=-X^{-1}dX\,X^{-1}$ has two inverses, one on each side. It is not $-X^{-2}dX$ (that would be right only if $X$ and $dX$ commuted).
- The transposes matter for non-symmetric $X$. For a symmetric covariance matrix, they all vanish.
Quick check: $X=\operatorname{diag}(a,b)$. What is $\nabla_X\log\det X$ at the diagonal entries?
$\log\det X=\ln a+\ln b$. The slopes with respect to the diagonal entries are $1/a$ and $1/b$, which are the diagonal entries of $X^{-\top}=\operatorname{diag}(1/a,1/b)$ ✓.
Chain-rule identities: $\nabla f(A\mathbf{x})=A^\top\nabla f$ core
Most functions in machine learning are compositions: do something to $\mathbf{x}$ to get $\mathbf{u}$, then feed $\mathbf{u}$ into a scalar function $f$. A nudge $d\mathbf{x}$ makes $\mathbf{u}$ move by $d\mathbf{u}=J\,d\mathbf{x}$, and that makes $f$ move by $\nabla f^\top d\mathbf{u}$. Chain the two: $df=\nabla f^\top J\,d\mathbf{x}$.
Reading off, the gradient is $J^\top\nabla f$. The transpose is the heart of backpropagation: a gradient measured at the output, $\nabla f$, travels backwards through the inner function by multiplying with $J^\top$. If the inner function is a matrix, $\mathbf{u}=A\mathbf{x}$, then $J=A$ and the backward step is multiplication by $A^\top$.
- Linear inner function. $f(\mathbf{u})=\tfrac12\|\mathbf{u}\|^2$ and $A=\begin{bmatrix}1&2\\3&0\\0&1\end{bmatrix}$, $\mathbf{x}=[1,1]^\top$. Then $\mathbf{u}=A\mathbf{x}=[3,3,1]^\top$, $\nabla f(\mathbf{u})=\mathbf{u}$, and $\nabla_{\mathbf{x}}f=A^\top\mathbf{u}=[1\cdot3+3\cdot3,\ \ 2\cdot3+1\cdot1]^\top=[12,\ 7]^\top$. Direct check: $f=\tfrac12\big[(x_1+2x_2)^2+9x_1^2+x_2^2\big]$, so $\partial_1f=(x_1+2x_2)+9x_1=3+9=12$ and $\partial_2f=2(x_1+2x_2)+x_2=6+1=7$ ✓.
- Nonlinear inner function. $\mathbf{g}(\mathbf{x})=(x_1x_2,\ x_1+x_2^2)$ and $f(\mathbf{u})=u_1^2+u_2$, at $\mathbf{x}=(2,1)$. Then $\mathbf{u}=(2,3)$, $\nabla f=[2u_1,\,1]=[4,1]^\top$, and $J_{\mathbf{g}}=\begin{bmatrix}x_2&x_1\\1&2x_2\end{bmatrix}=\begin{bmatrix}1&2\\1&2\end{bmatrix}$. So $J_{\mathbf{g}}^\top\nabla f=\begin{bmatrix}1&1\\2&2\end{bmatrix}\begin{bmatrix}4\\1\end{bmatrix}=[5,\ 10]^\top$. Direct check: $f\circ\mathbf{g}=(x_1x_2)^2+x_1+x_2^2$ gives $\partial_1=2x_1x_2^2+1=5$ and $\partial_2=2x_1^2x_2+2x_2=10$ ✓. (The wrong order $J\nabla f=[6,6]$ would fail.)
- Scalar outer function. $f(\mathbf{x})=e^{-\frac12\mathbf{x}^\top\mathbf{x}}$ (a Gaussian bump). Outer $h(q)=e^{-q}$ with $q=\tfrac12\|\mathbf{x}\|^2$, $\nabla q=\mathbf{x}$: $\nabla f=-e^{-q}\mathbf{x}$.
Let $f$ be a scalar function, $A$ a matrix, $\mathbf{g}$ a vector function with Jacobian $J_{\mathbf{g}}$:
$$\nabla_{\mathbf{x}}\,f(A\mathbf{x}+\mathbf{b})=A^\top\,\nabla f(A\mathbf{x}+\mathbf{b}),\qquad \nabla_{\mathbf{x}}\,f(\mathbf{g}(\mathbf{x}))=J_{\mathbf{g}}(\mathbf{x})^\top\,\nabla f(\mathbf{g}(\mathbf{x})),$$ $$\nabla_W\,f(W\mathbf{x}+\mathbf{b})=\nabla f\;\mathbf{x}^\top,\qquad \nabla_{\mathbf{b}}\,f(W\mathbf{x}+\mathbf{b})=\nabla f,\qquad \nabla_{\mathbf{x}}\,h(q(\mathbf{x}))=h'(q)\,\nabla q.$$Shapes: $A^\top$ is $n\times m$ and $\nabla f$ is $m\times1$, so the result is $n\times1$ (the shape of $\mathbf{x}$). $\nabla f\,\mathbf{x}^\top$ is $(m\times1)(1\times n)=m\times n$ (the shape of $W$). Awareness: the Hessian chain rule (Chapter 2.10) gives $\nabla^2_{\mathbf{x}}f(A\mathbf{x})=A^\top(\nabla^2f)\,A$.
Why do we need it?
Nearly every loss is a function of a function of the weights. The chain-rule identities turn "differentiate a big composition" into "multiply a few simple pieces".
Where is it used?
Logistic regression ($f$ applied to $X\mathbf{w}$), every layer of a neural network, backpropagation, Gaussian densities, and any loss written as $\ell(X\mathbf{w})$ on a data matrix $X$.
How is it used?
Differentiate the outer function with respect to its input, evaluated at the inner output. Then multiply by $J^\top$ of the inner function ($A^\top$ if it is linear). For a weight matrix use the outer product $\nabla f\,\mathbf{x}^\top$.
- The transpose is essential: it is $J^\top\nabla f$, not $J\nabla f$. When $J$ is square the shapes still fit, so a shape check cannot save you. Use the numerical check.
- $\nabla f$ must be evaluated at the inner output $\mathbf{u}=\mathbf{g}(\mathbf{x})$, not at $\mathbf{x}$.
- For $\nabla_W f(W\mathbf{x})$ the answer is an outer product (a matrix with the shape of $W$), not a dot product.
Quick check: $\mathbf{x}\in\mathbb{R}^5$, $A$ is $3\times5$, $f$ maps $\mathbb{R}^3\to\mathbb{R}$. What is the shape of $\nabla_{\mathbf{x}}f(A\mathbf{x})$ and what is its formula?
$A^\top\nabla f(A\mathbf{x})$: $(5\times3)(3\times1)=5\times1$, the shape of $\mathbf{x}$.
The one-page identity table, and how to check any gradient numerically core
Here is everything in one place. Use it like a dictionary: look up a pattern, but also know that each line is a few steps from the recipe. And whether you looked it up or derived it, always run the check at the bottom of this section. A gradient formula is a claim, and a numerical test is a quick way to prove the claim wrong, or to believe it.
The test is very simple: nudge each input up and down by a tiny amount, and see if the output changes at the rate your formula says.
Test the claim $\nabla(\mathbf{x}^\top A\mathbf{x})=(A+A^\top)\mathbf{x}$ with $A=\begin{bmatrix}2&1\\0&3\end{bmatrix}$ at $\mathbf{x}=[1,2]^\top$, using $h=10^{-5}$.
- Formula: $(A+A^\top)\mathbf{x}=[6,13]^\top$.
- Nudge $x_1$: $f=2x_1^2+x_1x_2+3x_2^2$, so $f(1.00001,2)=16.0000600002$ and $f(0.99999,2)=15.9999400002$. The slope is $(16.0000600002-15.9999400002)/(2\times10^{-5})=6.0000$ ✓.
- Nudge $x_2$ the same way: $13.0000$ ✓. The relative error is around $10^{-10}$.
- The wrong guess $2A\mathbf{x}=[8,12]^\top$ differs from $[6,13]$ by $2$: the check catches it at once.
The identity table. $\mathbf{x}\in\mathbb{R}^n$; $A,B$ constant matrices; $X$ a matrix variable; $\nabla_X$ has the shape of $X$. Each "key step" is the recipe's central move.
| Function | Gradient | Key step |
|---|---|---|
| $\mathbf{a}^\top\mathbf{x}$ | $\mathbf{a}$ | $df=\mathbf{a}^\top d\mathbf{x}$ |
| $A\mathbf{x}$ (Jacobian) | $A$ | $d\mathbf{y}=A\,d\mathbf{x}$ |
| $\|\mathbf{x}\|^2$ | $2\mathbf{x}$ | number equals its transpose |
| $\mathbf{x}^\top A\mathbf{x}$ | $(A+A^\top)\mathbf{x}$; $2A\mathbf{x}$ if symmetric | product rule, flip one term |
| $\tfrac12\mathbf{x}^\top A\mathbf{x}+\mathbf{b}^\top\mathbf{x}$ | $\tfrac12(A+A^\top)\mathbf{x}+\mathbf{b}$ | sum of the two rows above |
| $\|A\mathbf{x}-\mathbf{b}\|^2$ | $2A^\top(A\mathbf{x}-\mathbf{b})$ | $d\mathbf{r}=A\,d\mathbf{x}$, $df=2\mathbf{r}^\top d\mathbf{r}$ |
| $\|\mathbf{x}\|_2$ / $\|\mathbf{x}\|_1$ / $\|\mathbf{x}\|_\infty$ | $\mathbf{x}/\|\mathbf{x}\|$ / $\operatorname{sign}(\mathbf{x})$ / $\operatorname{sign}(x_k)\mathbf{e}_k$ | chain rule / kink: subgradient / only the max counts |
| $\sum\ln x_i$, $\ \sum e^{x_i}$, $\ \ln\sum e^{x_i}$ | $1/\mathbf{x}$, $\ e^{\mathbf{x}}$, $\ \operatorname{softmax}(\mathbf{x})$ | entrywise; chain rule on $\ln S$ |
| $f(A\mathbf{x}+\mathbf{b})$ / $f(\mathbf{g}(\mathbf{x}))$ | $A^\top\nabla f$ / $J_{\mathbf{g}}^\top\nabla f$ | $df=\nabla f^\top J\,d\mathbf{x}$ |
| $\operatorname{tr}(AX)$ / $\mathbf{a}^\top X\mathbf{b}$ | $A^\top$ / $\mathbf{a}\mathbf{b}^\top$ | $\operatorname{tr}(A\,dX)$ / rotate $\mathbf{b}$ to the front |
| $\operatorname{tr}(X^\top AX)$ | $(A+A^\top)X$ | transpose the first term |
| $\operatorname{tr}(AXB)$ | $A^\top B^\top$ | rotate: $\operatorname{tr}(BA\,dX)$ |
| $\|X\|_F^2$ / $\|X\|_F$ | $2X$ / $X/\|X\|_F$ | $\operatorname{tr}(X^\top X)$ |
| $\|AX-B\|_F^2$ / $\|XA-B\|_F^2$ | $2A^\top(AX-B)$ / $2(XA-B)A^\top$ | $dR=A\,dX$ / $dX\,A$, then rotate |
| $\|X\|_2$ / $\|X\|_*$ (awareness) | $\mathbf{u}_1\mathbf{v}_1^\top$ / $UV^\top$ | $d\sigma_i=\mathbf{u}_i^\top dX\,\mathbf{v}_i$ |
| $\log\lvert\det X\rvert$ / $\det X$ | $X^{-\top}$ / $\det X\,X^{-\top}$ | $d\det X=\det X\operatorname{tr}(X^{-1}dX)$ |
| $\operatorname{tr}(AX^{-1})$ | $-X^{-\top}A^\top X^{-\top}$ | $d(X^{-1})=-X^{-1}dX\,X^{-1}$ |
| $f(W\mathbf{x}+\mathbf{b})$ w.r.t. $W$ | $\nabla f\,\mathbf{x}^\top$ | $df=\operatorname{tr}(\mathbf{x}\nabla f^\top dW)$ |
How to check any gradient numerically.
- Shape first. The gradient must have the shape of the variable. (Free, instant.)
- Pick a random point with non-square, non-symmetric matrices. Avoid special points ($\mathbf{0}$, $I$, kinks like $x_i=0$ for $L_1$).
- Central differences: for each entry, $\dfrac{f(\ldots+h\,\mathbf{e}\ldots)-f(\ldots-h\,\mathbf{e}\ldots)}{2h}$ with $h\approx10^{-5}$ (and double-precision numbers). Forward differences are worse: error $\propto h$ instead of $h^2$.
- Compare with a relative error: $\dfrac{\max|\text{formula}-\text{numeric}|}{\max(1,\max|\text{formula}|)}$. Below $10^{-6}$: correct. Between $10^{-6}$ and $10^{-3}$: suspicious (try another $h$ or point). Above $10^{-3}$: a bug.
- Cheap version for huge models: pick a random direction $\mathbf{v}$ and compare $\nabla f^\top\mathbf{v}$ with the single slope $\frac{f(\mathbf{x}+h\mathbf{v})-f(\mathbf{x}-h\mathbf{v})}{2h}$. Two function evaluations instead of $2n$.
Library versions: scipy.optimize.check_grad, torch.autograd.gradcheck, jax.test_util.check_grads.
Why do we need it?
A wrong gradient does not crash your program: it just trains badly, silently. A ten-second numerical check turns "I think it is right" into "it is right".
Where is it used?
Writing custom layers and losses in PyTorch or JAX, unit tests in ML libraries, reproducing a paper's backward pass, and debugging models whose loss will not go down.
How is it used?
Run the five steps above on a small version of your problem (for example 3 inputs and 2 outputs). When the relative error is below $10^{-6}$ at a few random points, trust the formula and scale up.
- Do not use $h$ smaller than about $10^{-8}$: round-off errors grow. Do not use a huge $h$ either: the slope estimate becomes a poor average.
- In 32-bit floating point (the default in many deep-learning libraries) use bigger steps (about $10^{-3}$) or, better, run the check in 64-bit.
- A mismatch at a point where an entry is exactly $0$ (for $L_1$), or where two values tie (for $\max$), is not a bug: the function has a kink there.
Quick check: formula and numerical gradient differ by a factor of exactly $2$ in every entry, and the shape is right. Name two likely causes.
(1) A missing or extra factor $2$, for example using $A^\top(A\mathbf{x}-\mathbf{b})$ for $\|A\mathbf{x}-\mathbf{b}\|^2$. (2) Using $A\mathbf{x}$ instead of $2A\mathbf{x}$ for $\mathbf{x}^\top A\mathbf{x}$ with a symmetric $A$, or $2A\mathbf{x}$ for $\tfrac12\mathbf{x}^\top A\mathbf{x}$. If the formula is off by exactly the same factor $2$ in every entry, look for the missing or extra $\tfrac12$ first.
Recap, cheat sheet and practice
- The recipe replaces the table: (1) write $df$ with the differential rules, (2) rearrange with the four moves (wrap a scalar in $\operatorname{tr}$, rotate, transpose, collect) until $dX$ is at the right end, (3) read off ($\operatorname{tr}(M\,dX)\Rightarrow\nabla_X=M^\top$, $\ \mathbf{g}^\top d\mathbf{x}\Rightarrow\mathbf{g}$), (4) check shape and numbers.
- Linear functions give "the weights": $\mathbf{a}$, $A$, $A^\top$ (for $\operatorname{tr}(AX)$), $\mathbf{a}\mathbf{b}^\top$. Quadratics give $(A+A^\top)\mathbf{x}$, and $A\mathbf{x}+\mathbf{b}$ for $\tfrac12\mathbf{x}^\top A\mathbf{x}+\mathbf{b}^\top\mathbf{x}$ with symmetric $A$.
- Symmetry: $\mathbf{x}^\top A\mathbf{x}$ only sees $\tfrac12(A+A^\top)$, so $\nabla=2S\mathbf{x}$. $\mathbf{x}^\top K\mathbf{x}=0$ for antisymmetric $K$. $B^\top B$ is always symmetric.
- Trace identities: $\nabla\operatorname{tr}(AXB)=A^\top B^\top$, $\ \nabla\operatorname{tr}(X^\top AX)=(A+A^\top)X$. Transposes flip the pattern of the gradient.
- Norms: $\nabla\|\mathbf{x}\|_2=\mathbf{x}/\|\mathbf{x}\|$, $\nabla\|\mathbf{x}\|^2=2\mathbf{x}$, $\nabla\|\mathbf{x}\|_1=\operatorname{sign}(\mathbf{x})$ (subgradient at kinks), $\nabla\|X\|_F^2=2X$, $\nabla\|AX-B\|_F^2=2A^\top(AX-B)$. Spectral and nuclear: $\mathbf{u}_1\mathbf{v}_1^\top$ and $UV^\top$.
- Log-det and inverse: $d\log\det X=\operatorname{tr}(X^{-1}dX)$, $\nabla=X^{-\top}$, $\ d(X^{-1})=-X^{-1}dX\,X^{-1}$.
- Chain rule: $\nabla_{\mathbf{x}}f(A\mathbf{x})=A^\top\nabla f$, $\ \nabla f(\mathbf{g}(\mathbf{x}))=J_{\mathbf{g}}^\top\nabla f$, $\ \nabla_Wf(W\mathbf{x})=\nabla f\,\mathbf{x}^\top$.
- Always test: shape first, then central differences at a random point ($h\approx10^{-5}$), relative error below $10^{-6}$.
Cheat sheet: the four moves of the recipe
| Situation | Move | Example |
|---|---|---|
| $dX$ is not at the right end | rotate inside a trace | $\operatorname{tr}(P\,dX\,Q)=\operatorname{tr}(QP\,dX)$ |
| You have a number, not a trace | wrap: $s=\operatorname{tr}(s)$ | $\mathbf{a}^\top dX\,\mathbf{b}=\operatorname{tr}(\mathbf{b}\mathbf{a}^\top dX)$ |
| You see $dX^\top$ | transpose the whole term | $\operatorname{tr}(dX^\top M)=\operatorname{tr}(M^\top dX)$ |
| Two terms with $dX$ at the end | collect | $\operatorname{tr}(P\,dX)+\operatorname{tr}(Q\,dX)=\operatorname{tr}((P+Q)\,dX)$ |
| Final answer | transpose the coefficient | $df=\operatorname{tr}(M\,dX)\Rightarrow\nabla_Xf=M^\top$ |
import numpy as np
rng = np.random.default_rng(0)
def num_grad(f, X, h=1e-5):
"""Central-difference gradient of a scalar function f at X (any shape)."""
G = np.zeros_like(X)
for idx in np.ndindex(*X.shape):
E = np.zeros_like(X); E[idx] = h
G[idx] = (f(X + E) - f(X - E)) / (2 * h)
return G
def check(name, f, G, X):
"""Compare a formula G(X) with the numerical gradient. Prints True if they agree."""
a, n = G(X), num_grad(f, X)
rel = np.abs(a - n).max() / max(1.0, np.abs(a).max())
print(f"{name:28s}", rel < 1e-6)
A = rng.normal(size=(3, 3)); B = rng.normal(size=(3, 3)) # NOT symmetric
b = rng.normal(size=3); x = rng.normal(size=3); X = rng.normal(size=(3, 3))
Ai = np.linalg.inv
check("x^T A x", lambda v: v @ A @ v, lambda v: (A + A.T) @ v, x) # True
check("||Ax - b||^2", lambda v: np.sum((A @ v - b) ** 2), lambda v: 2 * A.T @ (A @ v - b), x) # True
check("||x||_2", lambda v: np.linalg.norm(v), lambda v: v / np.linalg.norm(v), x) # True
check("||x||_1", lambda v: np.abs(v).sum(), lambda v: np.sign(v), x) # True (no entry is 0)
check("tr(AX)", lambda M: np.trace(A @ M), lambda M: A.T, X) # True
check("tr(X^T A X)", lambda M: np.trace(M.T @ A @ M), lambda M: (A + A.T) @ M, X) # True
check("tr(A X B)", lambda M: np.trace(A @ M @ B), lambda M: A.T @ B.T, X) # True
check("||AX - B||_F^2", lambda M: np.sum((A @ M - B) ** 2), lambda M: 2 * A.T @ (A @ M - B), X) # True
check("log|det X|", lambda M: np.log(abs(np.linalg.det(M))), lambda M: Ai(M).T, X) # True
check("tr(A X^-1)", lambda M: np.trace(A @ Ai(M)), lambda M: -Ai(M).T @ A.T @ Ai(M).T, X) # True
# a WRONG formula is caught: 2Ax is not the gradient of x^T A x when A is not symmetric
check("x^T A x (wrong: 2Ax)", lambda v: v @ A @ v, lambda v: 2 * A @ v, x) # False
# chain rule: f(Wx) with f(u) = sum(log(cosh(u))) -> grad_x = W^T tanh(Wx), grad_W = tanh(Wx) x^T
W = rng.normal(size=(4, 3))
f = lambda u: np.sum(np.log(np.cosh(u)))
check("f(Wx) w.r.t. x", lambda v: f(W @ v), lambda v: W.T @ np.tanh(W @ v), x) # True
check("f(Wx) w.r.t. W", lambda M: f(M @ x), lambda M: np.outer(np.tanh(M @ x), x), W) # True
# a cheap check for big models: one random direction v instead of every entry
v = rng.normal(size=3)
g = (A + A.T) @ x # formula for the gradient of x^T A x
h = 1e-5
slope = ((x + h * v) @ A @ (x + h * v) - (x - h * v) @ A @ (x - h * v)) / (2 * h)
print(abs(g @ v - slope) < 1e-6) # True
1. $f(X)=\operatorname{tr}(AXB)$ with $A$, $X$, $B$ all $3\times3$. Which gradient is correct?
2. A subgradient of $\|\mathbf{x}\|_1$ at $\mathbf{x}=[2,-1,0]^\top$ is…
3. For a non-symmetric invertible $X$, $\nabla_X\log|\det X|$ equals…
4. $\mathbf{x}\in\mathbb{R}^5$, $A$ is $3\times5$, $f:\mathbb{R}^3\to\mathbb{R}$. The gradient $\nabla_{\mathbf{x}}f(A\mathbf{x})$ is…
5. $K$ is antisymmetric ($K^\top=-K$). Then $\mathbf{x}^\top K\mathbf{x}$ equals…
6. Your formula and the central-difference gradient (with $h=10^{-5}$, at a random point) have a relative error of $0.3$. What do you conclude?
Practice problems
A. Find $\nabla_{\mathbf{x}}(\mathbf{a}^\top\mathbf{x})^2$ using the chain rule, and check it at $\mathbf{a}=[1,2]^\top$, $\mathbf{x}=[3,1]^\top$.
Outer $h(q)=q^2$ with $q=\mathbf{a}^\top\mathbf{x}$, $\nabla q=\mathbf{a}$. So $\nabla=2(\mathbf{a}^\top\mathbf{x})\,\mathbf{a}$. At the point: $\mathbf{a}^\top\mathbf{x}=5$, gradient $=10[1,2]^\top=[10,20]^\top$. Direct: $f=(x_1+2x_2)^2$, $\partial_1=2(x_1+2x_2)=10$, $\partial_2=4(x_1+2x_2)=20$ ✓.
B. Find $\nabla_X\|X-C\|_F^2$, and say what the gradient descent update with step $\eta=\tfrac12$ does.
With $R=X-C$, $dR=dX$, $df=2\operatorname{tr}(R^\top dX)$, so $\nabla=2(X-C)$. The update $X\leftarrow X-\tfrac12\cdot2(X-C)=C$: it lands on $C$ in one step. (Same as $A=I$ in $\|AX-B\|_F^2$.)
C. Compute $\nabla_\Sigma\log\det\Sigma$ at $\Sigma=\operatorname{diag}(2,5)$.
$\Sigma^{-\top}=\Sigma^{-1}=\operatorname{diag}(1/2,\,1/5)$. Check: $\log\det\Sigma=\ln(\Sigma_{11}\Sigma_{22}-\Sigma_{12}\Sigma_{21})$, and $\partial/\partial\Sigma_{11}=\Sigma_{22}/\det=5/10=0.5$, $\partial/\partial\Sigma_{22}=2/10=0.2$, and the off-diagonal slopes are $0$ at a diagonal matrix. ✓
D. Find $\nabla_{\mathbf{x}}\|A\mathbf{x}\|_2$ (assume $A\mathbf{x}\ne\mathbf{0}$).
Chain rule with $\mathbf{u}=A\mathbf{x}$ and outer $f(\mathbf{u})=\|\mathbf{u}\|$, $\nabla f=\mathbf{u}/\|\mathbf{u}\|$, $J=A$. So $\nabla=A^\top\dfrac{A\mathbf{x}}{\|A\mathbf{x}\|}=\dfrac{A^\top A\mathbf{x}}{\|A\mathbf{x}\|}$. Shape $n\times1$ ✓.
E. Use the recipe to find $\nabla_W\big[\operatorname{tr}(W^\top W)+\operatorname{tr}(WA)\big]$.
$df=2\operatorname{tr}(W^\top dW)+\operatorname{tr}(A\,dW)$ (using $\nabla\|W\|_F^2=2W$ and $\operatorname{tr}(A\,dW)$ for the second term). So $G^\top=2W^\top+A$ and $G=2W+A^\top$.
F. Derive $\nabla_X\operatorname{tr}(X^{-1})$.
$df=\operatorname{tr}\big(d(X^{-1})\big)=-\operatorname{tr}(X^{-1}dX\,X^{-1})$. Rotate the last $X^{-1}$ to the front: $-\operatorname{tr}(X^{-2}\,dX)$. So $G^\top=-X^{-2}$ and $G=-(X^{-2})^\top=-X^{-\top}X^{-\top}$.
The Chain Rule
Almost every function in machine learning is a chain: a small step, then another, then another. The chain rule tells you how to get the slope of the whole chain from the slope of each link. It is the one idea behind backpropagation, so we will go slowly and build it up from gears to graphs.
- Use the scalar chain rule and see why it works (slopes multiply)
- Use the multivariate chain rule: add up the contribution of every path
- Write the vector chain rule and the Jacobian chain rule, and check the shapes
- Draw a computation as a computational graph with local derivatives
- Run forward differentiation and reverse differentiation by hand, and compare their cost
- Explain why a scalar loss with millions of parameters needs reverse mode (this leads straight to Chapter 2.9, backpropagation)
The scalar chain rule: slopes multiply core
Think of three gears in a row. Turn the handle once and gear A turns 2 times. Each turn of A makes gear B turn 3 times. So one turn of the handle makes B turn $2\times 3 = 6$ times. When steps are chained, the rates multiply.
A car works the same way. Petrol used per kilometre, times kilometres driven per hour, gives petrol used per hour. Each link only knows its own rate. The whole chain's rate is the product.
A derivative is just a rate: how fast the output moves per unit move of the input (see Chapter 2.3). A function that feeds into another function is a chain of two gears.
Let $y = (2x+1)^3$. Break it into two steps: first $u = 2x + 1$ (the inner step), then $y = u^3$ (the outer step). We want $dy/dx$ at $x = 1$.
- Inner rate: $\dfrac{du}{dx} = 2$.
- Outer rate: $\dfrac{dy}{du} = 3u^2$. At $x=1$ we have $u = 2\cdot1+1 = 3$, so $\dfrac{dy}{du} = 3\cdot 3^2 = 27$. (We use the inner value $u=3$, not $x=1$.)
- Multiply: $\dfrac{dy}{dx} = 27 \cdot 2 = 54$.
Numeric check (nudge $x$ by $0.01$). At $x=1.01$ we get $u = 3.02$ and $y = 3.02^3 = 27.543608$. The rise is $27.543608 - 27 = 0.543608$, and $0.543608 / 0.01 = 54.36 \approx 54$. ✓ (A smaller nudge gets even closer to 54.)
A chain of three links. The sigmoid is $\sigma(z) = (1 + e^{-z})^{-1}$. Write it as $a = -z$, then $b = 1 + e^{a}$, then $\sigma = b^{-1}$:
- $\dfrac{da}{dz} = -1$, $\dfrac{db}{da} = e^{a} = e^{-z}$, $\dfrac{d\sigma}{db} = -b^{-2}$.
- Multiply the three rates: $\dfrac{d\sigma}{dz} = (-b^{-2})\cdot e^{-z}\cdot(-1) = \dfrac{e^{-z}}{(1+e^{-z})^2}$.
- Now split this as $\dfrac{1}{1+e^{-z}}\cdot\dfrac{e^{-z}}{1+e^{-z}}$. The first factor is $\sigma$. The second is $1-\sigma$ (because $1 - \frac{1}{1+e^{-z}} = \frac{e^{-z}}{1+e^{-z}}$). So $\sigma'(z) = \sigma(z)\,(1-\sigma(z))$.
- Check at $z=0$: $\sigma = 0.5$, so $\sigma' = 0.5\cdot0.5 = 0.25$. A nudge: $\sigma(0.01) = 0.502500$, and $(0.502500 - 0.5)/0.01 = 0.25$ ✓.
If $y = f(u)$ and $u = g(x)$, so that $y = f(g(x))$, then
$$\frac{dy}{dx} \;=\; f'\big(g(x)\big)\cdot g'(x) \;=\; \frac{dy}{du}\cdot\frac{du}{dx}.$$Read it as "(outer slope, measured at the inner value) times (inner slope)". For longer chains keep multiplying: if $y = f(v)$, $v = h(u)$, $u = g(x)$ then $\dfrac{dy}{dx} = \dfrac{dy}{dv}\dfrac{dv}{du}\dfrac{du}{dx}$.
The fraction look of $\frac{dy}{du}\cdot\frac{du}{dx}$ is a memory aid: the "$du$" seems to cancel. It is not real cancelling (these are not ordinary fractions), but the pattern is a good way to remember which pieces go together.
Why do we need it?
Real models are functions of functions. Without the chain rule we would have to expand the whole formula before differentiating. With it, we only differentiate one simple link at a time and multiply.
Where is it used?
Every activation inside a neuron (sigmoid, tanh or ReLU of a weighted sum), the derivative of log-loss, exp and log in softmax, and the whole of backpropagation in PyTorch, JAX and TensorFlow.
How is it used?
Name the inner part $u$. Differentiate the outer function with respect to $u$, and the inner function with respect to $x$. Put the inner value back into the outer slope, then multiply.
- Evaluate the outer slope at the inner value. In the example it is $3u^2$ with $u=3$, not $3x^2$ with $x=1$.
- Do not forget the inner slope. The derivative of $\sin(x^2)$ is $\cos(x^2)\cdot 2x$, not just $\cos(x^2)$.
- The chain rule multiplies. The sum rule adds. They are different situations: "one after another" multiplies, "side by side" adds.
Quick check: differentiate $y = e^{3x}$ and find the slope at $x = 0$.
Inner $u = 3x$, so $du/dx = 3$. Outer $y = e^u$, so $dy/du = e^u$. Then $dy/dx = 3e^{3x}$. At $x=0$ this is $3e^0 = 3$.
Why the chain rule works: a derivation from nudges
Nudge the input $x$ by a tiny amount. The inner function reacts and $u$ moves a little. The outer function sees that moved $u$ and reacts too, so $y$ moves. Two small reactions, one after the other.
Each reaction is "slope times the nudge" (that is what a slope means, when the nudge is small). So the second reaction is (outer slope) times (the inner reaction), which is (outer slope) times (inner slope) times (the first nudge).
Take $y = (2x+1)^3$ at $x=1$ and nudge by $h = 0.1$.
- $\Delta x = 0.1$. Then $u$ goes from $3$ to $2(1.1)+1 = 3.2$, so $\Delta u = 0.2$ and $\Delta u / \Delta x = 2$.
- $y$ goes from $27$ to $3.2^3 = 32.768$, so $\Delta y = 5.768$ and $\Delta y/\Delta u = 28.84$.
- Multiply: $\dfrac{\Delta y}{\Delta u}\cdot\dfrac{\Delta u}{\Delta x} = 28.84 \cdot 2 = 57.68$, and indeed $\dfrac{\Delta y}{\Delta x} = \dfrac{5.768}{0.1} = 57.68$. (The two ratios multiply to the third exactly, because $\Delta u$ cancels.)
- As $h$ shrinks the ratios settle at $2$, $27$ and $54$.
Derivation. For a nudge $\Delta x \neq 0$, let $\Delta u = g(x+\Delta x) - g(x)$ and $\Delta y = f(u+\Delta u) - f(u)$. When $\Delta u \neq 0$ we can write
$$\frac{\Delta y}{\Delta x} = \frac{\Delta y}{\Delta u}\cdot\frac{\Delta u}{\Delta x}.$$Now let $\Delta x \to 0$. Then $\Delta u \to 0$ as well (a differentiable function does not jump), so $\Delta y/\Delta u \to f'(u)$ and $\Delta u/\Delta x \to g'(x)$. The product of the limits is $f'(u)\,g'(x)$. ∎
(If $\Delta u$ happens to be exactly $0$ the division is not allowed. Mathematicians handle that case with a small extra argument. The picture above is the heart of the proof.)
Why do we need it?
You asked to learn how to derive identities, not memorise them. This shows the chain rule is just "slope times nudge", applied twice, so you can rebuild it at any time.
Where is it used?
The same nudge argument explains gradient checking (nudge a weight, watch the loss) and the "local derivative times upstream gradient" step inside backpropagation.
How is it used?
Whenever you doubt a derivative, nudge the input by a small $h$, compute $(f(x+h)-f(x))/h$ and compare. Chain-rule results must match.
Quick check: why can the two nudge ratios be multiplied so easily?
Because $\Delta u$ appears once on top and once on the bottom: $\frac{\Delta y}{\Delta u}\cdot\frac{\Delta u}{\Delta x} = \frac{\Delta y}{\Delta x}$. It cancels exactly, as with ordinary fractions. The chain rule is what is left when the nudges become infinitely small.
The multivariate chain rule: add up every path core
A shop's profit depends on the price and on the number sold. Both of those depend on one dial: how much you spend on advertising. Turn the dial a little. The price reacts, and that changes the profit. The number sold reacts, and that also changes the profit.
There are two routes from the dial to the profit. Along each route, the rates multiply (that is the gears idea). To get the total effect of the dial, add the routes.
Draw the routes as arrows: dial $\to$ price $\to$ profit, and dial $\to$ number sold $\to$ profit. The rule is "multiply along a path, add over paths".
Let $f(u,v) = uv + v^2$, with $u = t^2$ and $v = 3t$. Find $df/dt$ at $t = 1$.
- Values at $t=1$: $u = 1$, $v = 3$.
- Rates of the inner steps: $\dfrac{du}{dt} = 2t = 2$ and $\dfrac{dv}{dt} = 3$.
- Partial derivatives of the outer function (Chapter 2.4): $\dfrac{\partial f}{\partial u} = v = 3$ and $\dfrac{\partial f}{\partial v} = u + 2v = 1 + 6 = 7$.
- Path through $u$: $\dfrac{\partial f}{\partial u}\cdot\dfrac{du}{dt} = 3\cdot2 = 6$. Path through $v$: $\dfrac{\partial f}{\partial v}\cdot\dfrac{dv}{dt} = 7\cdot3 = 21$.
- Add the paths: $\dfrac{df}{dt} = 6 + 21 = 27$.
Check by expanding. $f = t^2\cdot3t + 9t^2 = 3t^3 + 9t^2$, so $f' = 9t^2 + 18t = 27$ at $t=1$ ✓. Check by nudging. $f(1.01) = 3(1.030301) + 9(1.0201) = 12.271803$, so $(12.271803 - 12)/0.01 = 27.18 \approx 27$ ✓.
Let $f(u_1,\dots,u_n)$ and let every $u_i$ depend on $t$. Then
$$\frac{df}{dt} \;=\; \sum_{i=1}^{n}\frac{\partial f}{\partial u_i}\,\frac{du_i}{dt}.$$If each $u_i$ depends on several inputs $x_1,\dots,x_k$, apply the rule once for each input (the other inputs stay frozen):
$$\frac{\partial f}{\partial x_j} \;=\; \sum_{i=1}^{n}\frac{\partial f}{\partial u_i}\,\frac{\partial u_i}{\partial x_j}.$$Why it works. A tiny nudge $\Delta t$ moves each $u_i$ by about $\frac{du_i}{dt}\Delta t$. By the linear-approximation rule from Chapter 2.4 (a small step $\Delta\mathbf{u}$ changes $f$ by about $\nabla f\cdot\Delta\mathbf{u}$), $\Delta f \approx \sum_i \frac{\partial f}{\partial u_i}\Delta u_i = \Big(\sum_i \frac{\partial f}{\partial u_i}\frac{du_i}{dt}\Big)\Delta t$. Divide by $\Delta t$.
Same variable used twice. For $f = x\cdot x$ there are two paths from $x$ to $f$ (through the left factor and through the right factor). Each contributes $x$, and the sum is $2x$. This "fan-out" case is why gradients are added in backpropagation.
Why do we need it?
A value is often used in several places (a weight shared by several outputs, an input feeding several neurons). We must count its effect through every place, or the derivative will be too small.
Where is it used?
Hidden neurons that feed many neurons in the next layer, shared weights in convolution and recurrent networks, residual connections ($\mathbf{x} + f(\mathbf{x})$ uses $\mathbf{x}$ twice), and every node with fan-out in a computational graph.
How is it used?
List all paths from the input to the output. Multiply the local derivatives along each path. Add the path products. If it feels hard, draw the graph first.
- Count every path. A forgotten path means a missing term. Fan-out (one value used twice) is the usual place this happens.
- Multiply along a path, add between paths. Never add along a path and never multiply between paths.
- $\partial f/\partial u$ inside the rule is a partial derivative (other inputs of $f$ held still), while $du/dt$ is an ordinary derivative.
Quick check: $f = \sin(x)\cdot x^2$. Use "two paths from $x$" to find $f'(x)$.
Write $f = u\cdot v$ with $u = \sin x$ and $v = x^2$. Then $\partial f/\partial u = v = x^2$, $\partial f/\partial v = u = \sin x$, $du/dx = \cos x$, $dv/dx = 2x$. Sum of paths: $f' = x^2\cos x + \sin x\cdot 2x$. This is exactly the product rule: the product rule is a chain rule over two paths.
The vector chain rule: gradient dotted with velocity core
Picture a hiker on a hilly map. Her position $\mathbf{u}(t)$ changes with time: she walks along a path. The height of the ground under her is $f(\mathbf{u})$. How fast is her height changing?
Two things decide it. The gradient $\nabla f$ points uphill and says how steep the hill is. Her velocity $\mathbf{u}'(t)$ says which way and how fast she walks. If she walks straight uphill she climbs fast. If she walks along a contour line she does not climb at all. The "how much do two arrows agree" measure is the dot product. So the climb rate is gradient · velocity.
Take the function $f(u_1,u_2) = u_1^2 + u_1u_2$ and the path $\mathbf{u}(t) = (t^2,\; 3t)$. At $t = 1$:
- Position: $\mathbf{u} = (1, 3)$.
- Gradient (as a column): $\nabla f = \begin{bmatrix} 2u_1 + u_2 \\ u_1 \end{bmatrix} = \begin{bmatrix} 5 \\ 1 \end{bmatrix}$.
- Velocity: $\mathbf{u}'(t) = (2t,\,3) = (2, 3)$.
- Dot product: $\dfrac{df}{dt} = 5\cdot2 + 1\cdot3 = 13$.
Check: $f(t) = t^4 + 3t^3$, so $f'(t) = 4t^3 + 9t^2 = 13$ at $t=1$ ✓. (This is the sum-over-paths rule with $n = 2$, written as a dot product.)
Many inputs. Now let $\mathbf{u} = g(x,y) = (xy,\; x+y)$ and $f(\mathbf{u}) = u_1u_2$, so $f = xy(x+y)$. At $(x,y) = (2,3)$: $\mathbf{u} = (6,5)$, $\nabla f = (u_2, u_1) = (5, 6)$. The matrix of inner slopes is $\begin{bmatrix} \partial u_1/\partial x & \partial u_1/\partial y \\ \partial u_2/\partial x & \partial u_2/\partial y\end{bmatrix} = \begin{bmatrix} y & x \\ 1 & 1\end{bmatrix} = \begin{bmatrix} 3 & 2 \\ 1 & 1\end{bmatrix}$. Then
$$\begin{bmatrix}\partial f/\partial x\\ \partial f/\partial y\end{bmatrix} = \begin{bmatrix} 3 & 1 \\ 2 & 1\end{bmatrix}\begin{bmatrix} 5 \\ 6\end{bmatrix} = \begin{bmatrix} 15 + 6 \\ 10 + 6\end{bmatrix} = \begin{bmatrix} 21 \\ 16\end{bmatrix}.$$Check by expanding: $f = x^2y + xy^2$, so $f_x = 2xy + y^2 = 12 + 9 = 21$ ✓ and $f_y = x^2 + 2xy = 4 + 12 = 16$ ✓.
Along a path. If $\mathbf{u}:\mathbb{R}\to\mathbb{R}^n$ is a curve and $f:\mathbb{R}^n\to\mathbb{R}$, then
$$\frac{d}{dt}f\big(\mathbf{u}(t)\big) = \nabla f\big(\mathbf{u}(t)\big)^{\top}\,\mathbf{u}'(t) = \nabla f\cdot\mathbf{u}'.$$From many inputs. If $\mathbf{u} = g(\mathbf{x})$ with $\mathbf{x}\in\mathbb{R}^k$, $\mathbf{u}\in\mathbb{R}^n$, let $J_g$ be the $n\times k$ matrix with entry $(i,j) = \partial u_i/\partial x_j$ (the Jacobian, Chapter 2.5). Then the gradient (a column) is
$$\nabla_{\mathbf{x}} f\big(g(\mathbf{x})\big) \;=\; J_g(\mathbf{x})^{\top}\,\nabla_{\mathbf{u}} f\big(g(\mathbf{x})\big).$$Derivation. By the multivariate rule, $\dfrac{\partial f}{\partial x_j} = \sum_i \dfrac{\partial f}{\partial u_i}\dfrac{\partial u_i}{\partial x_j} = \sum_i (J_g)_{ij}\,(\nabla f)_i = \big(J_g^{\top}\nabla f\big)_j$.
Shape check. $J_g^\top$ is $k\times n$ and $\nabla f$ is $n\times1$, so the product is $k\times1$, the same shape as $\mathbf{x}$ ✓. The transpose appears because we are going from "gradient with respect to $\mathbf{u}$" back to "gradient with respect to $\mathbf{x}$".
Why do we need it?
It turns the sum over paths into one dot product (or one matrix times a vector), which is what a computer does fast. It also gives the meaning: how fast a loss changes along the direction that the weights move.
Where is it used?
Gradient descent: the weights follow a path $\mathbf{w}(t)$ and the loss changes at rate $\nabla L\cdot\mathbf{w}'$. Also the directional derivative, the backward step through a layer ($J^\top\nabla$), and sensitivity analysis.
How is it used?
Compute the gradient of the outer function at the inner value, compute the Jacobian (or the velocity) of the inner function, and multiply: velocity with a dot product, or the transposed Jacobian with the gradient.
- The gradient is a column vector. The dot product is $\nabla f^\top\mathbf{u}'$. If you carry the gradient as a row, the transpose sits somewhere else, so keep one convention.
- When you go backwards from $\nabla_{\mathbf{u}}f$ to $\nabla_{\mathbf{x}}f$ you multiply by $J^\top$, not $J$. The shapes tell you: only $J^\top\nabla f$ fits.
Quick check: $f(u_1,u_2) = u_1u_2$ and the path $\mathbf{u}(t) = (t, t^2)$. Find $df/dt$ at $t=2$ with the dot-product rule, and check with $f(t)=t^3$.
$\nabla f = (u_2, u_1) = (4, 2)$ at $\mathbf{u}=(2,4)$. Velocity $= (1, 2t) = (1, 4)$. Dot product $= 4\cdot1 + 2\cdot4 = 12$. Check: $f(t) = t\cdot t^2 = t^3$ and $f'(2) = 3\cdot4 = 12$ ✓.
The Jacobian chain rule: matrices multiply core
Zoom in on a smooth function far enough and it looks like a matrix: a straight, linear map. That matrix is the Jacobian (Chapter 2.5). Now chain two such functions. Zoomed in, the first one is a matrix and the second one is a matrix. Doing one then the other is multiplying the matrices.
That is the scalar gears rule in bigger clothes. A "rate" has become a table of rates, and "multiply the rates" has become "multiply the tables".
Let $G:\mathbb{R}^2\to\mathbb{R}^2$ be $G(x_1,x_2) = (x_1x_2,\; x_1+x_2)$, and $F:\mathbb{R}^2\to\mathbb{R}^3$ be $F(u_1,u_2) = (u_1^2,\; u_1+u_2,\; u_1u_2)$. We want the Jacobian of $F\circ G$ at $\mathbf{x}=(2,3)$. (Convention: a Jacobian of a map $\mathbb{R}^n\to\mathbb{R}^m$ has $m$ rows and $n$ columns, entry $(i,j) = \partial F_i/\partial x_j$.)
- Inner value: $\mathbf{u} = G(2,3) = (6, 5)$.
- $J_G = \begin{bmatrix} x_2 & x_1 \\ 1 & 1\end{bmatrix} = \begin{bmatrix} 3 & 2 \\ 1 & 1\end{bmatrix}$ (shape $2\times2$).
- $J_F$ at $\mathbf{u}=(6,5)$: $\begin{bmatrix} 2u_1 & 0 \\ 1 & 1 \\ u_2 & u_1\end{bmatrix} = \begin{bmatrix} 12 & 0 \\ 1 & 1 \\ 5 & 6\end{bmatrix}$ (shape $3\times2$).
- Multiply, $(3\times2)(2\times2) = 3\times2$: $$J_{F\circ G} = \begin{bmatrix} 12 & 0 \\ 1 & 1 \\ 5 & 6\end{bmatrix}\begin{bmatrix} 3 & 2 \\ 1 & 1\end{bmatrix} = \begin{bmatrix} 36 & 24 \\ 4 & 3 \\ 21 & 16\end{bmatrix}.$$ For the first row: $12\cdot3 + 0\cdot1 = 36$ and $12\cdot2 + 0\cdot1 = 24$. Second row: $1\cdot3+1\cdot1 = 4$ and $1\cdot2+1\cdot1 = 3$. Third row: $5\cdot3+6\cdot1 = 21$ and $5\cdot2+6\cdot1 = 16$.
Check. The third output is $x_1x_2(x_1+x_2)$, whose gradient we found above is $(21, 16)$ ✓. The first output is $(x_1x_2)^2$, so $\partial/\partial x_1 = 2x_1x_2\cdot x_2 = 2\cdot6\cdot3 = 36$ ✓ and $\partial/\partial x_2 = 2x_1x_2\cdot x_1 = 24$ ✓.
If $G:\mathbb{R}^n\to\mathbb{R}^m$ and $F:\mathbb{R}^m\to\mathbb{R}^p$, then
$$\boxed{\;J_{F\circ G}(\mathbf{x}) \;=\; J_F\big(G(\mathbf{x})\big)\;J_G(\mathbf{x})\;}\qquad (p\times n)=(p\times m)(m\times n).$$- Order: the outer function's Jacobian is on the left. Matrix multiplication is not commutative.
- Shape rule: the inner size $m$ (the number of values passed from $G$ to $F$) must match. If it does not, you have mixed up the order.
- Long chains: $F_k\circ\cdots\circ F_1$ has Jacobian $J_k\cdots J_2J_1$, each $J_i$ evaluated at the value that reaches step $i$.
- Special cases: $n=m=p=1$ is the scalar rule. $p=1$ (a scalar output) makes $J_F$ a single row, $\nabla f^\top$, which gives $\nabla_{\mathbf{x}}f = J_G^\top\nabla_{\mathbf{u}} f$ from the previous section.
Derivation. Near $\mathbf{x}$, $G(\mathbf{x}+\boldsymbol{\delta}) \approx G(\mathbf{x}) + J_G\boldsymbol{\delta}$. Near $\mathbf{u}=G(\mathbf{x})$, $F(\mathbf{u}+\boldsymbol{\epsilon}) \approx F(\mathbf{u}) + J_F\boldsymbol{\epsilon}$. Put $\boldsymbol{\epsilon} = J_G\boldsymbol{\delta}$: $F(G(\mathbf{x}+\boldsymbol{\delta})) \approx F(G(\mathbf{x})) + J_FJ_G\boldsymbol{\delta}$. The matrix in front of $\boldsymbol{\delta}$ is the Jacobian of the composite.
Why do we need it?
It gives the derivative of any composition of vector functions from the Jacobians of the small parts, and the shape rule catches mistakes before you compute anything.
Where is it used?
Layer-by-layer gradients in neural networks (each layer has a Jacobian), normalising flows, robot kinematics, Gauss–Newton and Levenberg–Marquardt curve fitting, and sensitivity analysis of a whole pipeline.
How is it used?
Evaluate each step's Jacobian at the value that reaches it, multiply them with the last step on the left, and check that neighbouring sizes match. In code you rarely build the full matrices; you multiply a vector through them.
- The outer Jacobian goes on the left: $J_F J_G$, not $J_GJ_F$. Check the shapes.
- Evaluate each Jacobian at the right point: $J_F$ at $G(\mathbf{x})$, not at $\mathbf{x}$.
- The product of Jacobians is for the value direction (forward). Gradients of a scalar loss move the other way and use the transposes: $J_G^\top J_F^\top\nabla$. We use exactly this in reverse mode below.
Quick check: $G:\mathbb{R}^5\to\mathbb{R}^7$ and $F:\mathbb{R}^7\to\mathbb{R}^2$. What are the shapes of $J_G$, $J_F$ and $J_{F\circ G}$?
$J_G$ is $7\times5$, $J_F$ is $2\times7$, and $J_{F\circ G} = J_FJ_G$ is $(2\times7)(7\times5) = 2\times5$.
Computational graphs: nodes, edges and local derivatives core
Any formula can be cut into tiny steps where each step does one simple thing: add, multiply, take a sine. Write each result on a small box and draw an arrow from every box that was used to the box that it helped make. That picture is a computational graph.
- A node holds one value (an input, an in-between result, or the final output).
- An edge (arrow) says "this value was used to make that one".
- On every edge sits a local derivative: how fast the box at the tip of the arrow changes per unit change of the box at its tail. It depends on that one tiny operation only.
Think of an assembly line. Each worker knows only their own job and how sensitive their output is to each part they receive. Nobody needs to know the whole factory.
Take $f(x,y) = xy\,(x+y)$ at $x=2$, $y=3$. Cut it into three steps:
$$v_1 = x\cdot y,\qquad v_2 = x + y,\qquad f = v_1\cdot v_2.$$- Forward values: $v_1 = 6$, $v_2 = 5$, $f = 30$.
- Local derivatives of each step (one rule per operation): $\dfrac{\partial v_1}{\partial x} = y = 3$, $\dfrac{\partial v_1}{\partial y} = x = 2$, $\dfrac{\partial v_2}{\partial x} = 1$, $\dfrac{\partial v_2}{\partial y} = 1$, $\dfrac{\partial f}{\partial v_1} = v_2 = 5$, $\dfrac{\partial f}{\partial v_2} = v_1 = 6$.
- Paths from $x$ to $f$: $x\to v_1\to f$ with product $3\cdot5 = 15$, and $x\to v_2\to f$ with product $1\cdot6 = 6$. So $\dfrac{\partial f}{\partial x} = 15 + 6 = 21$.
- Paths from $y$ to $f$: $y\to v_1\to f$ gives $2\cdot5 = 10$, and $y\to v_2\to f$ gives $1\cdot6 = 6$. So $\dfrac{\partial f}{\partial y} = 16$.
Check: $f = x^2y + xy^2$, so $f_x = 2xy + y^2 = 12 + 9 = 21$ ✓ and $f_y = x^2 + 2xy = 4+12 = 16$ ✓.
A computational graph is a directed graph with no loops. Each node $v_i$ is computed by one elementary operation from the nodes $v_j$ that point to it. The local derivative on the edge $j\to i$ is $\dfrac{\partial v_i}{\partial v_j}$, computed from that single operation (treating its other inputs as constants).
Chain rule on a graph. The derivative of the output $f$ with respect to an input $x$ is the sum, over all paths from $x$ to $f$, of the product of the local derivatives along the path:
$$\frac{\partial f}{\partial x} \;=\; \sum_{\text{paths }x\to f}\ \prod_{\text{edges }j\to i\text{ on the path}}\frac{\partial v_i}{\partial v_j}.$$This is exactly the multivariate rule ("multiply along, add between"), applied to a graph.
| Operation | Local derivatives | In words |
|---|---|---|
| $v = a + b$ | $\partial v/\partial a = 1,\ \partial v/\partial b = 1$ | an add node passes the gradient to both inputs unchanged |
| $v = a\cdot b$ | $\partial v/\partial a = b,\ \partial v/\partial b = a$ | a multiply node passes back the other input |
| $v = e^{a}$ | $\partial v/\partial a = e^{a} = v$ | reuses its own output |
| $v = \ln a$ | $\partial v/\partial a = 1/a$ | |
| $v = \sin a$ | $\partial v/\partial a = \cos a$ | |
| $v = \sigma(a)$ | $\partial v/\partial a = v(1-v)$ | reuses its own output |
| $v = \max(0,a)$ | $1$ if $a>0$, else $0$ | ReLU lets the gradient through or blocks it |
Why do we need it?
A big formula is hard to differentiate in one go. A graph cuts it into steps so simple that each local derivative is a one-line rule, and a computer can do the bookkeeping.
Where is it used?
PyTorch builds a graph as your code runs ("dynamic graph"), TensorFlow and JAX trace one, and compilers such as XLA optimise it. Every neural network, loss and optimiser step is a node in such a graph.
How is it used?
Run the computation and record each operation as a node with its inputs. Keep the values. Then the chain rule on the graph gives every derivative using only local rules.
- A local derivative uses the values from the forward pass (for a multiply node it is the other input's value). So the graph must remember them.
- One operation = one node. If you hide two operations in one node, its local derivative is no longer a one-line rule.
- "Local" means local to one node. The final derivative needs all the local derivatives along the paths.
Quick check: in the graph of $f = x\cdot x$ (the node uses $x$ twice), what is $\partial f/\partial x$ by paths?
There are two edges from $x$ to the multiply node, one for each use. Each local derivative is the other factor, which is $x$. The two paths add: $x + x = 2x$ ✓.
Forward differentiation: carry the derivative along with the value core
When you compute $f(x,y)$ step by step you carry a number at each node: its value. Forward differentiation carries a second number next to it: how fast that node changes when you nudge one chosen input. Every node holds a pair (value, derivative), and both are updated by the same operation.
You choose one input to nudge (the seed) and press "go". The derivative information travels forward, in the same direction as the computation. It is like pushing a small ripple through the circuit and watching how big it is when it reaches the output.
Again $f = xy(x+y)$ at $(2,3)$. This time we nudge $x$: the seed is $\dot x = 1$ and $\dot y = 0$ (dot means "rate of change with respect to the nudged input"). Each node carries (value, rate).
- $x$: (2, 1) $y$: (3, 0).
- $v_1 = xy$: value $6$. Rate: $\dot v_1 = y\,\dot x + x\,\dot y = 3\cdot1 + 2\cdot0 = 3$.
- $v_2 = x+y$: value $5$. Rate: $\dot v_2 = \dot x + \dot y = 1 + 0 = 1$.
- $f = v_1v_2$: value $30$. Rate: $\dot f = v_2\,\dot v_1 + v_1\,\dot v_2 = 5\cdot3 + 6\cdot1 = 21$.
So $\partial f/\partial x = 21$ after one pass. For $\partial f/\partial y$ we must start again with the seed $\dot x=0,\ \dot y=1$: then $\dot v_1 = 3\cdot0+2\cdot1 = 2$, $\dot v_2 = 1$, $\dot f = 5\cdot2 + 6\cdot1 = 16$. That is a second pass.
In forward mode each node $v_i$ carries its value and its rate $\dot v_i = \dfrac{\partial v_i}{\partial(\text{seed})}$. The rate is computed from the rates of the nodes that feed it, with the local derivatives:
$$\dot v_i \;=\; \sum_{j\to i}\frac{\partial v_i}{\partial v_j}\,\dot v_j.$$One pass has about the cost of evaluating $f$ (a small constant factor more). More generally the seed can be any vector $\mathbf{s}$ (a nudge direction). One pass then returns the Jacobian–vector product $J\mathbf{s}$: the full derivative along that direction. The seeds $\mathbf{e}_1,\dots,\mathbf{e}_n$ give the columns of $J$, so the whole Jacobian costs $n$ passes, one per input.
Why do we need it?
It is the easiest form of automatic differentiation: no storage of the graph is needed, because derivatives are produced in the same order as the values.
Where is it used?
Functions with few inputs and many outputs, Jacobian-vector products (jvp in JAX), directional derivatives, sensitivity analysis in simulations, and Hessian-vector products (forward over reverse).
How is it used?
Pick the input to vary. Set its rate to 1 and all other inputs' rates to 0. Run the program once, updating (value, rate) pairs. Read the rate at the output.
- One forward pass gives the derivative with respect to one input direction only. A function with a million inputs needs a million passes for its full gradient.
- The rate $\dot v$ is not a derivative with respect to $v$. It is how fast $v$ changes per unit change of the seed input.
Quick check: $f = x\cdot y$ at $(4, 5)$. Run forward mode with the seed $\dot x=1,\ \dot y=0$. What do you get?
$\dot f = y\,\dot x + x\,\dot y = 5\cdot1 + 4\cdot0 = 5$. That is $\partial f/\partial x = y = 5$ ✓.
Reverse differentiation: send the sensitivity backwards core
Now flip the question. Instead of asking "if I nudge this input, how does the output change?" ask "how much does the output care about each node?" Start at the output, where the answer is easy: the output changes by exactly 1 per unit change of itself. Then walk backwards. A node's importance is passed back to the nodes that fed it, multiplied by how sensitive the node is to each of them.
Think of blame in a team project that went badly. Start from the final result, ask each member how much of the blame passes to the people who supplied their work, and keep passing it back. At the end, every starting point knows how much it is responsible for the final error, all from one sweep.
$f = xy(x+y)$ at $(2,3)$. First the forward pass (values): $v_1 = 6$, $v_2 = 5$, $f = 30$. Write $\bar v$ for $\partial f/\partial v$ (the sensitivity of the output to that node).
- Start: $\bar f = 1$.
- $f = v_1v_2$: $\bar v_1 = \bar f\cdot v_2 = 5$ and $\bar v_2 = \bar f\cdot v_1 = 6$.
- $v_2 = x + y$: it sends $\bar v_2\cdot 1 = 6$ to $x$ and $6$ to $y$.
- $v_1 = xy$: it sends $\bar v_1\cdot y = 5\cdot3 = 15$ to $x$ and $\bar v_1\cdot x = 5\cdot2 = 10$ to $y$.
- $x$ receives from two places, so add: $\bar x = 6 + 15 = 21$. Likewise $\bar y = 6 + 10 = 16$.
One backward sweep gave both derivatives, $(21, 16)$. The same numbers as forward mode, but forward mode needed two passes.
In reverse mode, after a forward pass that stores all values, we compute for each node its adjoint $\bar v_j = \dfrac{\partial f}{\partial v_j}$ going backwards from $\bar f = 1$:
$$\bar v_j \;=\; \sum_{j\to i}\ \bar v_i\;\frac{\partial v_i}{\partial v_j}.$$In words: (gradient arriving from each consumer) times (the local derivative of that consumer), summed over all consumers. More generally, starting from a vector $\mathbf{u}$ instead of $1$ gives the vector–Jacobian product $\mathbf{u}^\top J$. The seed $\mathbf{u}=\mathbf{e}_i$ gives row $i$ of $J$, so the full Jacobian costs $m$ passes, one per output. For a scalar output ($m=1$) one pass gives the whole gradient, whatever the number of inputs.
In matrix form, for the chain $J = J_3J_2J_1$, reverse mode computes $\bar{\mathbf{x}} = J_1^\top\big(J_2^\top(J_3^\top\,\mathbf{u})\big)$ from the output end (the transposes appear because the gradient is a column).
Why do we need it?
A model has one loss and millions of weights. Reverse mode gives the derivative of that one number with respect to every weight in a single sweep, which forward mode cannot do cheaply.
Where is it used?
Backpropagation in every deep-learning framework (loss.backward() in PyTorch, jax.grad, tf.GradientTape), and gradient-based fitting of any scalar objective: regression, logistic loss, variational inference.
How is it used?
Run the forward pass and keep every value. Set the output's adjoint to 1. Visit the nodes in reverse order: multiply the node's adjoint by each local derivative, and add it into the adjoint of each input.
- A node that is used in several places must add the gradients that come back from each use. Overwriting instead of adding is a classic bug.
- The backward pass needs the forward values (for example the other factor of a multiply). Reverse mode therefore stores them. This costs memory.
- One backward pass gives the gradient of one scalar output. Two outputs need two backward sweeps.
Quick check: for $f = x\cdot y$ at $(4,5)$, what are $\bar x$ and $\bar y$ after the backward pass?
$\bar f = 1$. The multiply node sends back the other factor: $\bar x = 1\cdot y = 5$ and $\bar y = 1\cdot x = 4$.
Forward versus reverse: counting the cost core
Picture the Jacobian $J$ of the whole function as a table with $m$ rows (outputs) and $n$ columns (inputs). The two modes fill the table in different shapes:
- Forward mode fills a whole column per pass (one input nudged, the effect on every output).
- Reverse mode fills a whole row per pass (one output, the sensitivity to every input).
A table with a million columns and one row is filled by one reverse pass, or by a million forward passes. A table with one column and a million rows is the opposite. A training loss has one row and one column per weight.
A tiny network has $n = 1{,}000{,}000$ parameters and one loss ($m=1$). Say one evaluation of the loss takes cost $C$. A forward or reverse pass costs a small multiple of $C$, about $3C$ (a rule of thumb: we use 3 here, and honest values are between about 2 and 4).
- Forward mode: $n$ passes, so about $1{,}000{,}000\times3C = 3{,}000{,}000\,C$.
- Reverse mode: $m = 1$ pass, so about $1\times3C = 3C$.
- Ratio: reverse is about one million times cheaper for the gradient.
- For comparison, finite differences (nudge each weight and re-evaluate) need $2n = 2{,}000{,}000$ evaluations, so about $2{,}000{,}000\,C$.
Reverse mode's price is memory: all forward values must be kept until the backward pass uses them. Forward mode needs almost no extra memory.
For $F:\mathbb{R}^n\to\mathbb{R}^m$ whose evaluation costs $C$:
| Forward mode | Reverse mode | |
|---|---|---|
| One pass computes | $J\mathbf{s}$ (a column combination) | $\mathbf{u}^\top J$ (a row combination) |
| Cost of one pass (rule of thumb) | about $2$–$3\,C$ | about $2$–$4\,C$ (the "cheap gradient" principle) |
| Passes for the full Jacobian | $n$ (one per input) | $m$ (one per output) |
| Extra memory | small | all intermediate values |
| Best when | $n \ll m$ (few inputs) | $m \ll n$ (few outputs, e.g. a loss) |
Because the Jacobian of a chain is a product $J_3J_2J_1$ and matrix multiplication can be grouped either way, the two modes are just two ways of bracketing the same product: forward mode computes $J_3(J_2(J_1\mathbf{s}))$, reverse mode computes $((\mathbf{u}^\top J_3)J_2)J_1$. A thin vector at one end keeps every intermediate product thin.
Why do we need it?
The cost of getting all the derivatives can differ by a factor of a million between the two orders. Choosing the right mode is what makes training large models practical at all.
Where is it used?
Reverse mode for training any model with a scalar loss. Forward mode for a few parameters with many outputs (physics simulations, sensitivity of a curve), and in "forward-over-reverse" Hessian-vector products for second-order methods.
How is it used?
Count inputs $n$ and outputs $m$. If $m$ is much smaller than $n$, use reverse mode (backward(), grad). If $n$ is much smaller, use forward mode (jvp). Budget memory for reverse mode.
- "Reverse is better" is not a law. It is better when there are more inputs than outputs. With few inputs and many outputs forward mode wins.
- The "about 3×" is a typical constant, not an exact number. It depends on the operations. The important thing is that it does not grow with the number of inputs.
- Reverse mode keeps all intermediate values. For a very deep network this is the main memory cost of training. Tricks such as gradient checkpointing recompute some values to save memory.
Quick check: a function has 3 inputs and 500 outputs. Which mode needs fewer passes for the full Jacobian, and how many?
Forward mode needs one pass per input: 3. Reverse mode needs one per output: 500. So forward mode is cheaper here, with 3 passes.
Recap, cheat sheet and practice
- Scalar chain rule: for $y=f(g(x))$, $\dfrac{dy}{dx}=f'(g(x))\,g'(x)$. The slopes of chained steps multiply (gears).
- Why: a nudge $\Delta x$ causes $\Delta u\approx g'\Delta x$, which causes $\Delta y\approx f'\Delta u$. The $\Delta u$ cancels.
- Multivariate: multiply along each path and add the paths: $\dfrac{df}{dt}=\sum_i\dfrac{\partial f}{\partial u_i}\dfrac{du_i}{dt}$.
- Vector form: $\dfrac{d}{dt}f(\mathbf{u}(t))=\nabla f\cdot\mathbf{u}'$, and $\nabla_{\mathbf{x}}f=J_g^\top\nabla_{\mathbf{u}}f$.
- Jacobian chain rule: $J_{F\circ G}=J_F\,J_G$ with shapes $(p\times m)(m\times n)$; the outer map goes on the left.
- Computational graph: nodes are values, edges carry local derivatives, and $\partial f/\partial x$ is the sum over paths of the product of local derivatives.
- Forward mode carries (value, rate) and costs one pass per input. Reverse mode sends adjoints backwards, adds at fan-out, and costs one pass per output (plus memory). A scalar loss with many parameters therefore needs reverse mode. That is backpropagation, the next chapter.
Cheat sheet
| Idea | Formula | Remember |
|---|---|---|
| Scalar chain rule | $\dfrac{dy}{dx}=\dfrac{dy}{du}\dfrac{du}{dx}$ | outer slope at the inner value |
| Paths | $\dfrac{\partial f}{\partial x}=\sum_{\text{paths}}\prod(\text{local derivatives})$ | multiply along, add between |
| Along a curve | $\dfrac{d}{dt}f(\mathbf{u})=\nabla f\cdot\mathbf{u}'$ | gradient · velocity |
| Gradient back | $\nabla_{\mathbf{x}}f=J_g^\top\nabla_{\mathbf{u}}f$ | transpose when going backwards |
| Jacobian chain | $J_{F\circ G}=J_FJ_G$ | $p\times n=(p\times m)(m\times n)$ |
| Forward mode | $\dot v_i=\sum_j\frac{\partial v_i}{\partial v_j}\dot v_j$ | $n$ passes, small memory |
| Reverse mode | $\bar v_j=\sum_i\bar v_i\frac{\partial v_i}{\partial v_j}$ | $m$ passes, stores values |
| Local rules | add: 1, 1; multiply: other input; ReLU: 1 or 0 | each node knows only itself |
import numpy as np
# 1) scalar chain rule: y = (2x + 1)^3 at x = 1
x = 1.0
u = 2 * x + 1 # inner value
dy_dx = (3 * u**2) * 2 # (outer slope at u) * (inner slope)
f = lambda x: (2 * x + 1) ** 3
h = 1e-6
print(dy_dx, (f(x + h) - f(x - h)) / (2 * h)) # 54.0 54.00000000044258
# 2) multivariate chain rule: f(u, v) = u*v + v^2, u = t^2, v = 3t, at t = 1
t = 1.0
u, v = t**2, 3 * t
df_dt = (v) * (2 * t) + (u + 2 * v) * 3 # sum over the two paths
g = lambda t: (t**2) * (3 * t) + (3 * t) ** 2
print(df_dt, (g(t + h) - g(t - h)) / (2 * h)) # 27.0 26.999999998444935
# 3) Jacobian chain rule: J_{F o G} = J_F(G(x)) @ J_G(x)
G = lambda x: np.array([x[0] * x[1], x[0] + x[1]])
F = lambda u: np.array([u[0] ** 2, u[0] + u[1], u[0] * u[1]])
def num_jac(fn, x, h=1e-6):
cols = []
for j in range(len(x)):
e = np.zeros(len(x)); e[j] = h
cols.append((fn(x + e) - fn(x - e)) / (2 * h))
return np.stack(cols, axis=1)
x = np.array([2.0, 3.0])
u = G(x)
JG = np.array([[x[1], x[0]], [1, 1]])
JF = np.array([[2 * u[0], 0], [1, 1], [u[1], u[0]]])
J = JF @ JG
print(J) # [[36. 24.] [ 4. 3.] [21. 16.]]
print(np.allclose(J, num_jac(lambda z: F(G(z)), x), atol=1e-5)) # True
# 4) the same function as a tiny graph: f = (x*y) * (x+y)
def forward_mode(x, y, dx, dy):
v1, d1 = x * y, y * dx + x * dy # carry (value, derivative)
v2, d2 = x + y, dx + dy
return v1 * v2, v2 * d1 + v1 * d2
print(forward_mode(2.0, 3.0, 1.0, 0.0)[1]) # 21.0 (one pass per input)
print(forward_mode(2.0, 3.0, 0.0, 1.0)[1]) # 16.0
def reverse_mode(x, y):
v1, v2 = x * y, x + y # forward pass: store values
f_bar = 1.0 # backward pass
v1_bar, v2_bar = f_bar * v2, f_bar * v1
x_bar = v1_bar * y + v2_bar * 1.0 # x is used twice: add
y_bar = v1_bar * x + v2_bar * 1.0
return x_bar, y_bar
print(reverse_mode(2.0, 3.0)) # (21.0, 16.0) both in ONE pass
# 5) forward vs reverse as matrix products: J = J3 @ J2 @ J1 (n = 5 inputs, m = 2 outputs)
rng = np.random.default_rng(0)
J1, J2, J3 = rng.normal(size=(4, 5)), rng.normal(size=(4, 4)), rng.normal(size=(2, 4))
Jfull = J3 @ J2 @ J1
v = rng.normal(size=5) # forward mode: J @ v, right to left
jvp = J3 @ (J2 @ (J1 @ v))
u_ = rng.normal(size=2) # reverse mode: u @ J, left to right
vjp = ((u_ @ J3) @ J2) @ J1
print(np.allclose(jvp, Jfull @ v), np.allclose(vjp, u_ @ Jfull)) # True True
1. Let $y = (5x-2)^2$. What is $dy/dx$ at $x = 1$?
2. $f(u,v)$ has $\partial f/\partial u=3$ and $\partial f/\partial v=4$ at the point of interest. If $u=2x$ and $v=x^2$, what is $df/dx$ at $x=1$?
3. $G:\mathbb{R}^4\to\mathbb{R}^3$ and $F:\mathbb{R}^3\to\mathbb{R}^5$. What is the shape of the Jacobian of $F\circ G$?
4. A model has 1,000,000 parameters and a single scalar loss. How many backward (reverse-mode) passes give the full gradient?
5. You need the full gradient of $f:\mathbb{R}^{100}\to\mathbb{R}$ using forward mode. How many passes?
6. In the backward pass, a node's value was used in two places. What do you do with the two gradients that come back?
Practice problems
A. Differentiate $y = \sin(3x^2)$.
Inner $u = 3x^2$, $du/dx = 6x$. Outer $y=\sin u$, $dy/du = \cos u$. So $dy/dx = \cos(3x^2)\cdot 6x$.
B. Differentiate the softplus function $y = \ln(1 + e^{x})$ and recognise the answer.
Inner $u = 1 + e^x$, $du/dx = e^x$. Outer $y=\ln u$, $dy/du = 1/u$. So $dy/dx = \dfrac{e^x}{1+e^x}$. Dividing top and bottom by $e^x$ gives $\dfrac{1}{1+e^{-x}} = \sigma(x)$: the derivative of softplus is the sigmoid. Check at $x=0$: slope $= 1/2$, and $(\ln(1+e^{0.01}) - \ln 2)/0.01 \approx 0.5012$ ✓.
C. $f(u,v)=u^2v$ with $u = x+y$ and $v = xy$. Find $\partial f/\partial x$ at $(x,y)=(1,2)$ by summing paths, then check by expanding.
Values: $u = 3$, $v = 2$. Outer partials: $f_u = 2uv = 12$, $f_v = u^2 = 9$. Inner partials with respect to $x$: $\partial u/\partial x = 1$, $\partial v/\partial x = y = 2$. Paths: $12\cdot1 + 9\cdot2 = 12 + 18 = 30$. Check: $f = (x+y)^2xy$, so $f_x = 2(x+y)xy + (x+y)^2y = 2\cdot3\cdot2 + 9\cdot2 = 12 + 18 = 30$ ✓.
D. $f(\mathbf{u}) = u_1^2 + u_2$ and $G(x_1,x_2) = (x_1+x_2,\ x_1x_2)$. Find the gradient of $f\circ G$ at $(1,2)$ with the Jacobian chain rule.
$\mathbf{u} = G(1,2) = (3, 2)$. $J_F = [\,2u_1,\ 1\,] = [6,\ 1]$ (a $1\times2$ row). $J_G = \begin{bmatrix}1 & 1\\ x_2 & x_1\end{bmatrix} = \begin{bmatrix}1&1\\2&1\end{bmatrix}$. Product: $[6\cdot1 + 1\cdot2,\ 6\cdot1+1\cdot1] = [8,\ 7]$, so the gradient is $(8,7)$. Check: $f\circ G = (x_1+x_2)^2 + x_1x_2$, so $\partial_1 = 2(x_1+x_2)+x_2 = 6+2 = 8$ ✓ and $\partial_2 = 6 + 1 = 7$ ✓.
E. A sensor model has 3 tunable inputs and produces 1000 readings. Which mode would you use to get its full Jacobian, and why?
Forward mode: it needs 3 passes (one per input), while reverse mode would need 1000 (one per output). Each forward pass also uses little memory.
F. Draw the graph of $f = x\,(x+y)$ and find $\partial f/\partial x$ at $(1,2)$ by listing the paths.
Nodes: $v = x+y = 3$, $f = x\cdot v = 3$. The input $x$ feeds $f$ directly (local derivative $= v = 3$) and also through $v$ (local derivatives $1$ and then $\partial f/\partial v = x = 1$). Paths: $3 + 1\cdot1 = 4$. Check: $f = x^2 + xy$, $f_x = 2x + y = 4$ ✓.
Backpropagation & Automatic Differentiation
Backpropagation sounds mysterious. It is not. It is the chain rule from Chapter 2.8, applied to a computational graph, in the clever order (from the output backwards). By the end of this chapter you will be able to say why, and to show it with every number written out.
- Compare the three ways to get a derivative: symbolic, numerical and automatic
- Build a computational graph, run the forward pass and read off local derivatives
- Run the backward pass: local derivative times upstream gradient, added at fan-out
- Work a tiny neural network completely by hand (every number), and verify it numerically
- Write backprop for dense layers, activations and losses, with shapes
- Understand forward mode (dual numbers), reverse mode, and the cost in time and memory
- See vanishing and exploding gradients, and check gradients like a professional
- Explain why backpropagation is essentially repeated application of the chain rule
Three ways to get a derivative core
Suppose you want the steepness of a hill at the spot where you stand. There are three ways.
- Symbolic: you have the formula of the hill, and you do algebra on it, like a student with a pen. The answer is another formula.
- Numerical: you ignore the formula. You take one small step, measure how much you rose, and divide by the step. The answer is an approximation.
- Automatic (AD): you watch the computer compute the height, one tiny operation at a time. For each tiny operation you know its exact slope (a one-line rule). The chain rule glues these slopes together as the program runs. The answer is exact, and no big formula is ever written.
Deep-learning libraries use the third way. It is as exact as algebra and almost as cheap as evaluating the function itself.
Differentiate $f(x) = x\sin x$ at $x = 1$ in all three ways.
- Symbolic. Product rule: $f'(x) = \sin x + x\cos x$. At $x=1$: $0.841471 + 0.540302 = 1.381773$.
- Numerical. Take $h = 0.001$. $f(1.001) = 0.842853$ and $f(1) = 0.841471$, so the slope is about $(0.842853 - 0.841471)/0.001 = 1.38189$. Close, but already wrong in the fourth decimal place: the error is about $1.2\times10^{-4}$.
- Automatic (forward). Carry (value, rate) pairs, with $x = (1,\,1)$. The step $\sin x$ gives $(0.841471,\ \cos 1\cdot1 = 0.540302)$. The product step gives value $1\cdot0.841471 = 0.841471$ and rate $1\cdot0.841471 + 1\cdot0.540302 = 1.381773$. Exactly the symbolic answer, with no formula written.
If we make the numerical step smaller, it first gets better and then worse: with $h=10^{-6}$ the error is $10^{-7}$, with $h = 10^{-12}$ it is back to $10^{-4}$, and with $h=10^{-15}$ it is $0.06$. Computers store numbers with about 16 digits, and subtracting two almost equal numbers throws digits away. The widget below shows this.
| Symbolic | Numerical (finite differences) | Automatic (AD) | |
|---|---|---|---|
| What you get | a formula for $f'$ | an approximate number | an exact number (to rounding) |
| Accuracy | exact | truncation error plus rounding error; even with the best $h$ the error is only about $10^{-8}$ (forward) to $10^{-10}$ (central) | exact, about $10^{-16}$ |
| Cost for a gradient of $n$ inputs | can be huge ("expression swell"), then still evaluate it | $n+1$ or $2n$ evaluations of $f$ | reverse mode: a small multiple (rule of thumb: 2–4) of one evaluation, for any $n$ |
| Handles loops, if-statements, program code? | no: needs a closed formula | yes (treats $f$ as a black box) | yes: it follows the actual run |
| Main use | pen-and-paper maths, computer algebra | checking gradients | training models |
Finite differences. Forward: $f'(x)\approx\dfrac{f(x+h)-f(x)}{h}$ (error about $h$). Central: $f'(x)\approx\dfrac{f(x+h)-f(x-h)}{2h}$ (error about $h^2$, usually much better). Both are derived from the definition of the derivative in Chapter 2.3.
Automatic differentiation is neither of the other two. It applies the chain rule to the elementary operations of the program, so its results are exact. It does not manipulate formulas, and it does not take steps.
Why do we need it?
Training needs the gradient of a loss with millions of weights, at every step. We need a method that is exact, fast, and works on real program code. Only automatic differentiation checks all three boxes.
Where is it used?
PyTorch autograd, JAX grad, TensorFlow GradientTape, Stan, and Julia's Zygote use AD. Symbolic differentiation lives in SymPy and Mathematica. Finite differences are used in gradient checks and for black-box functions such as simulators.
How is it used?
You write the forward computation normally. The library records the operations and gives you the gradient (loss.backward()). You may confirm it once with a finite-difference check on a few weights.
- AD is not "numerical" and it is not "symbolic". It has no step size and it never writes the derivative formula. People often confuse these.
- Finite differences are still the right tool for checking a gradient (see the gradient-checking section) and for functions you cannot see inside.
- "Exact" means exact up to the usual 16-digit rounding of the computer.
Quick check: why does the numerical error first fall and then rise as $h$ shrinks?
Two errors compete. The truncation error (from using a straight line instead of the curve) shrinks as $h$ shrinks. The rounding error (from subtracting two nearly equal numbers, each stored with about 16 digits, then dividing by a tiny $h$) grows as $h$ shrinks. The best $h$ balances them.
Computational graphs and the forward pass core
A program that computes a loss is a recipe: a list of tiny steps, each using results of earlier steps. Draw every result as a box and every "was used by" as an arrow. This is the computational graph (you met it in Chapter 2.8).
The forward pass is simply running the recipe from the inputs to the output, filling in the number in each box. Nothing about derivatives yet. But there is one important extra rule: keep every number. The backward pass will need them.
A mini "neuron with a loss": one input $x$, one weight $w$ (we call it $y$ here to fit the widget), and a target $1$.
$$a = x\cdot y,\qquad s = \sigma(a),\qquad e = s - 1,\qquad L = e^2.$$At $x = 2$, $y = 0.5$:
- $a = 2\cdot0.5 = 1$.
- $s = \sigma(1) = \dfrac{1}{1 + e^{-1}} = 0.7311$.
- $e = 0.7311 - 1 = -0.2689$.
- $L = (-0.2689)^2 = 0.0723$.
All four numbers are stored. A real network does exactly this, with millions of boxes: the layer outputs, called activations, are the stored numbers.
The forward pass evaluates the nodes of the graph in an order where every node comes after the nodes it depends on (a topological order). Each node applies its elementary operation to its inputs' values:
$$v_i = \phi_i\big(v_{j_1}, v_{j_2},\dots\big).$$The result is the output (the loss) and a table of all intermediate values $v_i$. That table is the memory that backpropagation reads. In a network, the stored values are the pre-activations $\mathbf{z}$, activations $\mathbf{h}$, and the inputs of every layer.
Why do we need it?
Every local derivative is evaluated at the values the program actually had. Without the forward pass we would not know which numbers to plug in.
Where is it used?
Every prediction a model makes is a forward pass. In training, the forward pass computes the loss and records the activations; PyTorch's model(x) does this and builds the graph at the same time.
How is it used?
Evaluate the steps in order, store each result, and finish at the loss. Then start the backward pass. When you only want a prediction (inference), you can skip storing values and save memory.
- The forward pass is not "just" prediction: in training it also stores the intermediate values. That stored memory is large for deep networks.
- The order must respect dependencies. In a graph with no loops such an order always exists.
Quick check: with $x=2$, $y=0.5$ in the example, which stored value will the sigmoid's local derivative use?
The sigmoid's local derivative is $s(1-s)$, which uses its own stored output $s = 0.7311$: $0.7311\times0.2689 = 0.1966$. This is why the forward values must be kept.
Local derivatives: one rule per operation core
A single node does only one small thing: add, multiply, take $e^a$. For that one thing, the slope is a one-line rule you already know. The node does not care how the rest of the network looks. It only answers: "if my input changes a tiny bit, how much does my output change?"
Backpropagation is built from a short list of local rules, one per kind of node, like a recipe book. A whole deep network is nothing but these few node types used over and over.
The node $s = \sigma(a)$ at $a = 1$ (stored output $s = 0.7311$):
- Rule: $\partial s/\partial a = s(1-s)$ (derived in Chapter 2.8).
- Number: $0.7311\times(1 - 0.7311) = 0.7311\times0.2689 = 0.1966$.
- Check by nudging: $\sigma(1.001) = 0.731255$ and $\sigma(1) = 0.731059$, so the slope is about $(0.731255 - 0.731059)/0.001 = 0.196 \approx 0.1966$ ✓.
The multiply node $v = a\cdot b$ at $a = 3$, $b = 4$: $\partial v/\partial a = b = 4$ and $\partial v/\partial b = a = 3$. Check: nudge $a$ to $3.01$: $v = 12.04$, rise $0.04$ over $0.01$ gives $4$ ✓.
The local derivative of a node $v = \phi(a_1,\dots,a_k)$ is the vector of its partial derivatives $\partial\phi/\partial a_j$ evaluated at the stored input values. The standard list:
| Node | Local derivative(s) |
|---|---|
| $a + b$ | $1$ and $1$ |
| $a - b$ | $1$ and $-1$ |
| $a\cdot b$ | $b$ (for $a$) and $a$ (for $b$) |
| $a / b$ | $1/b$ and $-a/b^2$ |
| $a^k$ (constant $k$) | $k\,a^{k-1}$ |
| $e^a$ | $e^a$ (the stored output) |
| $\ln a$ | $1/a$ |
| $\sin a$, $\cos a$ | $\cos a$, $-\sin a$ |
| $\tanh a$ | $1 - \tanh^2 a$ (uses the stored output) |
| $\sigma(a)$ | $\sigma(a)(1-\sigma(a))$ (uses the stored output) |
| $\mathrm{ReLU}(a)=\max(0,a)$ | $1$ if $a > 0$, otherwise $0$ |
Notice that several rules reuse the node's own output: another reason to store the forward values.
Why do we need it?
If every node type has a tiny, tested derivative rule, a framework can differentiate any program built from them, even one invented tomorrow. Nobody has to differentiate the whole model by hand.
Where is it used?
Every op in PyTorch and JAX (add, matmul, relu, softmax, conv2d…) ships with its own backward rule. Writing a custom layer means writing its local derivative.
How is it used?
During the forward pass the node keeps what its rule needs. During the backward pass it multiplies the incoming gradient by its local derivative and hands the result to each input.
- ReLU has a corner at $0$ where the derivative is not defined. Frameworks simply pick $0$ (or $1$) there. It rarely matters in practice.
- A local derivative is evaluated at the stored input values, not at "x".
- For a node with two inputs there are two local derivatives, one per incoming edge.
Quick check: what are the local derivatives of $v = a\cdot b$ at $a = -2$, $b = 5$?
$\partial v/\partial a = b = 5$ and $\partial v/\partial b = a = -2$.
The backward pass: upstream gradient × local derivative core
After the forward pass we know the loss. Now we ask the question training cares about: "if I nudge this number, how much does the loss move?" We answer it for every number in one sweep, starting at the loss and walking backwards.
At the loss, the answer is trivial: the loss changes by exactly 1 per unit change of itself. Now step back one node. The node before it affects the loss only through the node in front. So its answer is: (the answer in front of it) × (how much I affect the node in front). The first part is called the upstream gradient. The second is the local derivative. That multiplication is the gear rule from the chain rule. Repeat all the way back.
Continue the mini-neuron from before: $a = xy$, $s = \sigma(a)$, $e = s - 1$, $L = e^2$ at $x=2$, $y = 0.5$. Forward values: $a = 1$, $s = 0.7311$, $e = -0.2689$, $L = 0.0723$. Now backwards:
- $\dfrac{\partial L}{\partial L} = 1$.
- $L = e^2$, local derivative $2e = -0.5378$. So $\dfrac{\partial L}{\partial e} = 1\times(-0.5378) = -0.5378$.
- $e = s - 1$, local derivative $1$. So $\dfrac{\partial L}{\partial s} = -0.5378\times1 = -0.5378$.
- $s = \sigma(a)$, local derivative $s(1-s) = 0.1966$. So $\dfrac{\partial L}{\partial a} = -0.5378\times0.1966 = -0.1057$ (more digits: $-0.10575$).
- $a = x\cdot y$: local derivatives $\partial a/\partial x = y = 0.5$ and $\partial a/\partial y = x = 2$. So $\dfrac{\partial L}{\partial x} = -0.10575\times0.5 = -0.0529$ and $\dfrac{\partial L}{\partial y} = -0.10575\times2 = -0.2115$.
Check. Nudge $y$ from $0.5$ to $0.51$: $L$ changes from $0.0723$ to $0.0702$ (rise $-0.0021$), and $-0.0021/0.01 = -0.209 \approx -0.2115$ ✓. A central difference with tiny $h$ gives $-0.21151$, matching to five digits.
The backward pass (reverse-mode differentiation). Given the stored forward values:
- Set the gradient at the output (the loss) to $1$, and all other gradients to $0$.
- Visit the nodes in reverse topological order. For a node $v_i$ with gradient $\bar v_i = \partial L/\partial v_i$ and inputs $v_j$: $$\bar v_j \;\mathrel{+}=\; \bar v_i\cdot\frac{\partial v_i}{\partial v_j}\qquad\text{(for every input }j\text{ of node }i\text{).}$$
- When every node has been visited, $\bar v_j = \partial L/\partial v_j$ for every node, including every weight.
The sign "$+=$" matters: a node used in several places receives a contribution from each, and they add. This is the multivariate chain rule of Chapter 2.8. So the backward pass is just: local derivative × upstream gradient, node after node, and sum at fan-out. Each node needs only its own local rule and the number arriving from the front.
Why do we need it?
Gradient descent needs $\partial L/\partial w$ for every weight. The backward pass delivers all of them at once, at a cost close to one extra forward pass.
Where is it used?
This is what loss.backward() does in PyTorch, jax.grad in JAX and tape.gradient in TensorFlow, for networks from tiny MLPs to large language models.
How is it used?
Run the forward pass, call backward, and read the stored .grad of each parameter. The optimiser then uses it (for example $w \leftarrow w - \eta\,\partial L/\partial w$).
- Gradients add where a value is used more than once. Overwriting is a bug that silently gives wrong gradients.
- The gradient at the output is 1 only because we differentiate the loss with respect to itself. With several outputs you start from a chosen weighting vector.
- Gradients must be reset between training steps (
optimizer.zero_grad()), because PyTorch accumulates them with "+=" just as above.
Quick check: $L = (a\cdot b)^2$ with $a=1$, $b=3$. Do the backward pass for $\partial L/\partial a$.
Forward: $p = ab = 3$, $L = p^2 = 9$. Backward: $\bar L = 1$; $\bar p = 2p = 6$; $\bar a = \bar p\cdot b = 6\cdot3 = 18$ and $\bar b = \bar p\cdot a = 6$. Check: $L = a^2b^2$, so $\partial L/\partial a = 2ab^2 = 18$ ✓.
Backpropagation: a tiny network, every number by hand core
Time to do it on a real (very small) neural network: 2 inputs, 2 hidden neurons, 1 output. The hidden neurons use the sigmoid. The output is a plain weighted sum. The loss is half the squared error. The network has $4 + 2 + 2 + 1 = 9$ numbers to learn: two weights into each hidden neuron (4), a bias for each (2), two weights into the output (2) and an output bias (1).
We will compute the forward pass, then the backward pass, then verify every one of the 9 gradients with a finite difference. After that, you can truthfully say you have done backpropagation.
Input $\mathbf{x} = (1, 2)$, target $t = 1$. Weights: $W_1 = \begin{bmatrix} 0.1 & 0.2 \\ -0.3 & 0.4\end{bmatrix}$, $\mathbf{b}_1 = (0.1, -0.1)$, $W_2 = [\,0.5,\ -0.5\,]$, $b_2 = 0.2$. (Values are shown rounded to 4 decimals, but each step is computed from the unrounded numbers. So a product of two printed numbers can differ from the printed result in the last digit.)
Forward pass.
- Pre-activations $\mathbf{z} = W_1\mathbf{x}+\mathbf{b}_1$: $z_1 = 0.1\cdot1 + 0.2\cdot2 + 0.1 = 0.6$, $\;z_2 = -0.3\cdot1 + 0.4\cdot2 - 0.1 = 0.4$.
- Activations $\mathbf{h} = \sigma(\mathbf{z})$: $h_1 = \sigma(0.6) = 0.6457$, $\;h_2 = \sigma(0.4) = 0.5987$.
- Output $\hat y = W_2\mathbf{h} + b_2 = 0.5\cdot0.6457 - 0.5\cdot0.5987 + 0.2 = 0.3228 - 0.2993 + 0.2 = 0.2235$.
- Error $r = \hat y - t = 0.2235 - 1 = -0.7765$, and loss $L = \tfrac12 r^2 = \tfrac12\cdot0.6030 = 0.3015$.
Backward pass (upstream × local at every step).
- $\dfrac{\partial L}{\partial\hat y} = r = -0.7765$ (since $L=\tfrac12r^2$ and $r = \hat y - t$).
- Output layer: $\dfrac{\partial L}{\partial W_2} = r\cdot\mathbf{h}^\top = (-0.7765\cdot0.6457,\ -0.7765\cdot0.5987) = (-0.5014,\ -0.4649)$ and $\dfrac{\partial L}{\partial b_2} = r = -0.7765$.
- Back to the hidden activations: $\dfrac{\partial L}{\partial h_j} = r\cdot W_2[j]$, so $\dfrac{\partial L}{\partial h_1} = -0.7765\cdot0.5 = -0.3883$ and $\dfrac{\partial L}{\partial h_2} = -0.7765\cdot(-0.5) = 0.3883$.
- Through the sigmoids: $\sigma'(z_j) = h_j(1-h_j)$, so $\sigma'(z_1) = 0.6457\cdot0.3543 = 0.2288$ and $\sigma'(z_2) = 0.5987\cdot0.4013 = 0.2403$. Then $\delta_1 = \dfrac{\partial L}{\partial z_1} = -0.3883\cdot0.2288 = -0.0888$ and $\delta_2 = \dfrac{\partial L}{\partial z_2} = 0.3883\cdot0.2403 = 0.0933$.
- Hidden layer weights: $\dfrac{\partial L}{\partial W_1[j,i]} = \delta_j\,x_i$. So $\dfrac{\partial L}{\partial W_1} = \begin{bmatrix} -0.0888\cdot1 & -0.0888\cdot2 \\ 0.0933\cdot1 & 0.0933\cdot2\end{bmatrix} = \begin{bmatrix} -0.0888 & -0.1777 \\ 0.0933 & 0.1866\end{bmatrix}$, and $\dfrac{\partial L}{\partial\mathbf{b}_1} = (\delta_1,\delta_2) = (-0.0888,\ 0.0933)$.
Verify numerically. Take $W_1[1,1]$. Add and subtract $h = 10^{-6}$, recompute $L$ each time, and divide the difference by $2h$: the result is $-0.088827$, matching $-0.0888$ ✓. All nine gradients agree with their finite differences to about $10^{-10}$ or better (the widget does this check for you).
One gradient step with learning rate $\eta = 0.1$ ($w \leftarrow w - \eta\,\partial L/\partial w$) lowers the loss from $0.3015$ to $0.1959$. Repeating forward, backward, update is called training.
For the network $\mathbf{z} = W_1\mathbf{x}+\mathbf{b}_1,\ \mathbf{h}=\sigma(\mathbf{z}),\ \hat y = W_2\mathbf{h}+b_2,\ L = \tfrac12(\hat y - t)^2$, backpropagation computes
$$\begin{aligned} r &= \hat y - t, & \dfrac{\partial L}{\partial W_2} &= r\,\mathbf{h}^\top, & \dfrac{\partial L}{\partial b_2} &= r,\\ \dfrac{\partial L}{\partial \mathbf{h}} &= W_2^\top r, & \boldsymbol{\delta} = \dfrac{\partial L}{\partial \mathbf{z}} &= \dfrac{\partial L}{\partial \mathbf{h}}\odot\mathbf{h}\odot(1-\mathbf{h}), & &\\ \dfrac{\partial L}{\partial W_1} &= \boldsymbol{\delta}\,\mathbf{x}^\top, & \dfrac{\partial L}{\partial \mathbf{b}_1} &= \boldsymbol{\delta}. & & \end{aligned}$$The symbol $\odot$ means entry-by-entry multiplication. Every line is "(upstream gradient) × (local derivative)". Nothing else is going on.
Why do we need it?
This is the whole of training in miniature: the forward pass gives the loss, the backward pass gives how every weight should change, and a small step downhill improves the model. Seeing every number removes the magic.
Where is it used?
The same pattern, scaled up, trains every MLP, CNN, RNN and Transformer. Real networks have more layers and other activations, but each layer repeats these same steps.
How is it used?
Do forward, then backward, then update the weights by $-\eta$ times each gradient. Verify new code with a finite-difference check on a few weights. The next widgets let you do exactly that.
- The factor $\sigma'(z) = h(1-h)$ is at most $0.25$. If it is tiny (a saturated sigmoid) the gradient behind it is tiny too. We return to this in the vanishing-gradients section.
- $\partial L/\partial W_1$ needs the input $x$ and the upstream $\delta$. $\partial L/\partial W_2$ needs the stored activation $h$. That is why the forward pass stores them.
- Rounding: the numbers here are shown with 4 decimals, so a hand multiplication may differ from the printed one in the last digit.
Quick check: in the example, which of the nine weights has the gradient with the biggest size, and why?
$b_2$ has $-0.7765$ (the error itself) and $W_2[1]$, $W_2[2]$ have $-0.5014$ and $-0.4649$. They sit right next to the output, so no small factor (like $\sigma' \le 0.25$) is multiplied in. The hidden-layer weights have gradients of at most about $0.19$, because they pass through $W_2$ and then $\sigma'$.
Computational graphs in neural networks: dense layer, activation, loss core
A real network is a graph made of only three kinds of building blocks, repeated:
- Dense (linear) layer: $\mathbf{y} = W\mathbf{x} + \mathbf{b}$, a bundle of weighted sums.
- Activation: an entry-by-entry function such as ReLU or sigmoid, $\mathbf{h} = \varphi(\mathbf{z})$.
- Loss: one number that compares the output with the target.
Each block has a tiny backward rule. Backpropagation through the whole network is just these rules, applied in reverse order, layer by layer, instead of node by node. A "layer" is a group of nodes that we treat together to use fast matrix code.
A dense layer with $\mathbf{x} = (1, 2, -1)$, $W = \begin{bmatrix} 1 & 0 & 2 \\ -1 & 3 & 1\end{bmatrix}$ and $\mathbf{b} = (0, 1)$. The upstream gradient (the gradient arriving from the layers in front) is $\mathbf{g} = \partial L/\partial\mathbf{y} = (2, -1)$.
- Forward: $\mathbf{y} = W\mathbf{x}+\mathbf{b} = (1 + 0 - 2 + 0,\ -1 + 6 - 1 + 1) = (-1,\ 5)$.
- $\dfrac{\partial L}{\partial W} = \mathbf{g}\,\mathbf{x}^\top = \begin{bmatrix} 2\cdot1 & 2\cdot2 & 2\cdot(-1)\\ -1\cdot1 & -1\cdot2 & -1\cdot(-1)\end{bmatrix} = \begin{bmatrix} 2 & 4 & -2 \\ -1 & -2 & 1\end{bmatrix}$ (same shape as $W$).
- $\dfrac{\partial L}{\partial\mathbf{b}} = \mathbf{g} = (2, -1)$.
- $\dfrac{\partial L}{\partial\mathbf{x}} = W^\top\mathbf{g} = \begin{bmatrix}1 & -1\\0 & 3\\2 & 1\end{bmatrix}\begin{bmatrix}2\\-1\end{bmatrix} = (2+1,\ 0-3,\ 4-1) = (3,\ -3,\ 3)$.
Check with $L = 2y_1 - y_2$ (so that $\partial L/\partial\mathbf{y} = \mathbf{g}$): $\partial L/\partial x_1 = 2\cdot W_{11} - 1\cdot W_{21} = 2 + 1 = 3$ ✓.
Backward rules for the three blocks (all with a column-vector gradient convention):
| Block (forward) | Given upstream gradient | Backward rule |
|---|---|---|
| Dense: $\mathbf{y} = W\mathbf{x} + \mathbf{b}$ $W$: $m\times n$ | $\mathbf{g} = \partial L/\partial\mathbf{y}$ ($m\times1$) | $\dfrac{\partial L}{\partial W} = \mathbf{g}\mathbf{x}^\top$ ($m\times n$), $\dfrac{\partial L}{\partial\mathbf{b}} = \mathbf{g}$, $\dfrac{\partial L}{\partial\mathbf{x}} = W^\top\mathbf{g}$ ($n\times1$) |
| Activation: $\mathbf{h} = \varphi(\mathbf{z})$ | $\mathbf{g} = \partial L/\partial\mathbf{h}$ | $\dfrac{\partial L}{\partial\mathbf{z}} = \mathbf{g}\odot\varphi'(\mathbf{z})$ |
| Loss: $L = \tfrac12\|\hat{\mathbf{y}}-\mathbf{t}\|^2$ | (start of the backward pass) | $\dfrac{\partial L}{\partial\hat{\mathbf{y}}} = \hat{\mathbf{y}} - \mathbf{t}$ |
| Loss: softmax then cross-entropy | (start) | $\dfrac{\partial L}{\partial\mathbf{z}} = \mathbf{p} - \mathbf{y}_{\text{one-hot}}$ (you will meet this derivation in Chapter 2.14; the log-sum-exp part of it is in Chapter 2.7) |
Derivation for the dense layer. The loss depends on $W_{ij}$ only through $y_i = \sum_j W_{ij}x_j + b_i$, and $\partial y_i/\partial W_{ij} = x_j$. So $\partial L/\partial W_{ij} = g_i\,x_j$, which is the outer product $\mathbf{g}\mathbf{x}^\top$. Also $x_j$ feeds all outputs, with $\partial y_i/\partial x_j = W_{ij}$, so by the multivariate rule $\partial L/\partial x_j = \sum_i g_iW_{ij} = (W^\top\mathbf{g})_j$.
Shape rule. Every gradient has the same shape as the thing it is the gradient of. $\partial L/\partial W$ is $m\times n$ like $W$; $\partial L/\partial\mathbf{x}$ is $n\times1$ like $\mathbf{x}$. With a batch of $B$ examples the weight gradient is the sum over the batch, $\sum_b \mathbf{g}_b\mathbf{x}_b^\top$.
Why do we need it?
Treating a whole layer as one block lets us use fast matrix multiplication on GPUs, for both the forward and the backward pass, with only a few rules to implement per layer type.
Where is it used?
The linear layers, feed-forward blocks and attention projections of Transformers, fully connected heads of CNNs, and every library's Linear module use these rules. Convolutions, normalisation and softmax have their own similar rules.
How is it used?
Walk the layers from the last to the first. At each one, use the upstream gradient to get the weight gradient (using the stored input of that layer), then the gradient for the layer before it (using the transposed weights).
- $\partial L/\partial W = \mathbf{g}\mathbf{x}^\top$ and $\partial L/\partial\mathbf{x} = W^\top\mathbf{g}$. Mixing up the transpose is the commonest hand-written-backprop bug. Let the shapes tell you.
- The activation backward rule is an entry-wise product. It is not a matrix product, because each activation acts on one entry only.
- With a mean loss over $B$ examples, remember the $1/B$ factor.
Quick check: a dense layer maps 784 inputs to 128 outputs. What is the shape of $\partial L/\partial W$, and of $\partial L/\partial\mathbf{x}$?
$W$ is $128\times784$, so $\partial L/\partial W$ is $128\times784$ (the same shape as $W$). $\partial L/\partial\mathbf{x} = W^\top\mathbf{g}$ is $(784\times128)(128\times1) = 784\times1$.
Why backpropagation is just the chain rule, used cleverly core
Here is the one sentence to remember:
Backpropagation is the chain rule applied again and again from the output backwards, saving the partial products so that every weight can reuse them.
The gradient of one weight is always a product of local derivatives along a path from that weight to the loss. Many weights share the front part of their paths. Backprop multiplies that shared part once and passes it back, instead of recomputing it for each weight. That reuse is the "propagation".
In the tiny network, take the weight $W_1[1,1]$. Follow its path to the loss: $W_1[1,1]\to z_1\to h_1\to\hat y\to L$. The chain rule multiplies the four local derivatives:
$$\frac{\partial L}{\partial W_1[1,1]} = \underbrace{\frac{\partial L}{\partial\hat y}}_{r}\;\underbrace{\frac{\partial\hat y}{\partial h_1}}_{W_2[1]}\;\underbrace{\frac{\partial h_1}{\partial z_1}}_{h_1(1-h_1)}\;\underbrace{\frac{\partial z_1}{\partial W_1[1,1]}}_{x_1} = (-0.7765)(0.5)(0.2288)(1) = -0.0888.$$Now see what is shared. The first factor $r$ is used by all nine gradients. The product $r\cdot W_2[1]\cdot\sigma'(z_1) = \delta_1$ is used by $W_1[1,1]$, $W_1[1,2]$ and $b_1[1]$. Backprop computes $r$, then $\delta_1$, then multiplies by $x_1$ or $x_2$ at the very end. Nothing is multiplied twice.
The dictionary.
| Chain-rule idea (Chapter 2.8) | What backprop does |
|---|---|
| Slopes of chained steps multiply | Each backward step multiplies the incoming gradient by a local derivative |
| Sum over all paths | Where a value is used twice, the returning gradients are added |
| Jacobian chain rule $J=J_k\cdots J_1$ | For a scalar loss, multiply from the left with transposes: $J_1^\top(\cdots(J_k^\top\mathbf{1}))$ |
| Reverse mode: one pass per output | One backward pass gives the gradient of the loss for every weight |
| Local derivatives need values | The forward pass stores activations |
Algorithm. (1) Forward pass: compute and store all values. (2) Set $\partial L/\partial L = 1$. (3) For each node in reverse order: multiply its gradient by each local derivative and add to the inputs' gradients. (4) Read the gradients of the weights. This is exactly the chain rule over the graph, and nothing more.
Why do we need it?
Computing every weight's gradient by a separate full chain would repeat huge amounts of work. Reusing the shared factors makes the cost grow in step with the network size, not with its square.
Where is it used?
The reuse idea is why training networks with hundreds of layers is possible at all. It is also why frameworks store activations, and why tricks like gradient checkpointing and mixed precision are about memory (and speed) rather than about maths.
How is it used?
When you debug a gradient, pick one weight and write its chain of local derivatives. Compare each factor with what the framework stored. It must be the same product.
- "Just the chain rule" does not mean "trivial". The clever part is the order (from the loss backwards) and reuse of partial products. Reverse order is what makes one pass enough for all weights.
- Backprop is not a learning algorithm by itself. It only computes gradients. The optimiser (such as gradient descent) uses them to change the weights.
Quick check: say in one sentence why backpropagation is repeated application of the chain rule.
Each weight's gradient is a product of local derivatives along its path to the loss (chain rule), and backprop evaluates these products from the loss backwards, one node at a time, adding where paths merge and reusing the shared partial products for all weights behind them.
Forward-mode differentiation and dual numbers core
In Chapter 2.8 we carried a pair (value, rate) through the graph. There is a neat way to see that pair as one number. Imagine a number $x + \varepsilon$: $x$ plus an infinitely tiny nudge $\varepsilon$, so tiny that $\varepsilon^2$ is zero for all practical purposes.
Run your program on this nudged number. Every operation keeps two parts: the ordinary part and the nudge part. At the end, the nudge part is exactly the derivative. You never wrote a derivative rule for the whole program. The arithmetic of nudged numbers produced it.
Evaluate $f(x) = x^2 + 3x$ at $x = 2 + \varepsilon$, using $\varepsilon^2 = 0$:
$$(2+\varepsilon)^2 + 3(2+\varepsilon) = (4 + 4\varepsilon + \varepsilon^2) + (6 + 3\varepsilon) = 10 + 7\varepsilon.$$The ordinary part $10$ is $f(2)$. The nudge part $7$ is $f'(2) = 2\cdot2+3$ ✓.
A product: $(1+\varepsilon)\cdot\sin(1+\varepsilon)$. The rule $\sin(a + b\varepsilon) = \sin a + b\cos a\,\varepsilon$ gives $\sin(1+\varepsilon) = 0.8415 + 0.5403\varepsilon$. Then $(1+\varepsilon)(0.8415 + 0.5403\varepsilon) = 0.8415 + (0.5403 + 0.8415)\varepsilon + 0.5403\varepsilon^2 = 0.8415 + 1.3818\,\varepsilon$. So $f(1) = 0.8415$ and $f'(1) = 1.3818$ for $f = x\sin x$, the same as in the first section.
A dual number is $a + b\varepsilon$ with real $a, b$ and the rule $\varepsilon^2 = 0$ (but $\varepsilon\neq0$). Its arithmetic:
$$\begin{aligned} (a+b\varepsilon) \pm (c+d\varepsilon) &= (a\pm c) + (b\pm d)\varepsilon\\ (a+b\varepsilon)(c+d\varepsilon) &= ac + (ad + bc)\varepsilon \qquad\text{(the product rule appears!)}\\ \frac{a+b\varepsilon}{c+d\varepsilon} &= \frac{a}{c} + \frac{bc - ad}{c^2}\varepsilon\qquad (c\neq0)\\ \varphi(a+b\varepsilon) &= \varphi(a) + \varphi'(a)\,b\,\varepsilon\qquad\text{(one-input functions)} \end{aligned}$$Why it works. The last line is the first-order Taylor expansion $\varphi(a+\Delta) \approx \varphi(a) + \varphi'(a)\Delta$ with $\Delta = b\varepsilon$ and all higher terms dropped because $\varepsilon^2=0$. Because each operation is computed correctly this way, the chain rule is applied automatically, step by step: $f(x+\varepsilon) = f(x) + f'(x)\,\varepsilon$ for any program built from these operations.
Forward-mode AD is exactly "run the program on dual numbers". To differentiate with respect to input $x_i$, give that input the nudge part $1$ and every other input the nudge part $0$. One run gives one derivative (one column of the Jacobian); a function with $n$ inputs needs $n$ runs. The cost of one run is a small constant times a normal run, since every number is carried as two numbers.
Why do we need it?
It shows that differentiation can be done by ordinary arithmetic on a slightly richer kind of number. Operator overloading gives a working forward-mode AD in a few dozen lines, and it is exact.
Where is it used?
JAX jvp, Julia's ForwardDiff package, C++ libraries such as Eigen's autodiff, and quick sensitivity analysis of physics and finance models with few inputs. Forward mode also gives cheap Hessian-vector products when combined with reverse mode.
How is it used?
Replace the number type with a dual number type (a pair), seed the input of interest with nudge part 1, run the code, and read the nudge part of the result as the derivative.
- $\varepsilon$ is not a small number you pick. It is a formal symbol with $\varepsilon^2 = 0$. That is why there is no truncation error and no step size to tune.
- One run still gives only one directional derivative. For a million parameters you would need a million runs. That is why training uses reverse mode instead.
- Dual numbers carry two numbers instead of one, so each operation costs about two to three times more than the plain number.
Quick check: compute $(3+\varepsilon)^3$ with $\varepsilon^2 = 0$ (and so $\varepsilon^3 = 0$). What are $f(3)$ and $f'(3)$ for $f=x^3$?
$(3+\varepsilon)^3 = 27 + 3\cdot9\,\varepsilon + 3\cdot3\,\varepsilon^2 + \varepsilon^3 = 27 + 27\varepsilon$. So $f(3) = 27$ and $f'(3) = 27 = 3\cdot3^2$ ✓.
Computational complexity: time and memory core
What does backpropagation cost? Two things matter: time and memory.
- Time. The backward pass visits every node once, with about the same amount of arithmetic as the forward pass. For a dense layer it is exactly two matrix products ($\partial L/\partial W$ and $\partial L/\partial\mathbf{x}$), each as big as the forward one. So the backward pass costs about 2× the forward pass, and a whole training step about 3× one prediction. This is a rule of thumb. The key point is that you do not need one extra pass per weight.
- Memory. The backward pass needs the forward values, so the forward pass must keep them: all the layer inputs and activations, for every example in the batch. For big networks this is the number that fills your GPU.
A plain network with $L = 100$ layers, each $1000$ wide, batch size $B = 128$.
- Parameters: $100\times1000\times1000 = 10^8$ numbers.
- Forward multiply-adds: $L\,B\,n^2 = 100\cdot128\cdot10^6 = 1.28\times10^{10}$. Backward: about twice that, $2.56\times10^{10}$. One training step: about $3.84\times10^{10}$.
- Stored activations: $L\,B\,n = 100\cdot128\cdot1000 = 1.28\times10^7$ numbers, which is $51$ MB in 32-bit floats.
- Gradient checkpointing: keep only every $k$-th layer's activations and recompute the others during the backward pass. With $k=\sqrt{L}=10$ you store about $2\sqrt{L} = 20$ layers' worth instead of 100 (memory about 5× smaller) and pay about one extra forward pass in time.
For convolutional networks and Transformers the stored activations are usually much larger than in this plain example, which is why memory, not arithmetic, is often the limit.
Reverse mode (backpropagation). Time: about $2$–$4$ times the cost of evaluating the function (a rule of thumb, the "cheap gradient principle"), for a scalar output, independent of the number of inputs. Memory: proportional to the number of intermediate values stored, that is, the size of the forward computation.
Forward mode. Time: about $2$–$3$ times the cost of the function per input. Memory: little more than the function itself.
Whole Jacobian of $F:\mathbb{R}^n\to\mathbb{R}^m$. Forward: $n$ passes. Reverse: $m$ passes. A scalar loss has $m=1$.
Finite differences. $2n$ function evaluations for a gradient (central differences), which for millions of weights is hopeless.
Why do we need it?
Cost decides what can be trained. Knowing that the backward pass costs about 2× the forward pass, and that memory scales with stored activations, tells you how big a model and batch fit on your hardware.
Where is it used?
Planning GPU memory for training large models, choosing the batch size, gradient checkpointing (torch.utils.checkpoint), mixed precision (16-bit activations halve the memory), and comparing AD modes in scientific computing.
How is it used?
Estimate parameters as the sum of weight sizes, activations as batch × layer widths, and time as about 3× the forward cost. If memory is too high, lower the batch size, use checkpointing, or use smaller precision.
- "About 2–3×" is a rule of thumb. Exact costs depend on the layer types, memory traffic and hardware.
- Forward mode is cheaper in memory, but reverse mode wins whenever there are more inputs than outputs.
- Inference (prediction only) does not need the stored activations, which is why running a model needs much less memory than training it.
Quick check: one forward pass of a network costs $F$ multiply-adds. Roughly what does one full training step (forward, backward) cost?
The backward pass costs about $2F$ (one product for $\partial L/\partial W$ and one for $\partial L/\partial\mathbf{x}$ per layer), so the whole step costs about $3F$.
Automatic differentiation in action: a mini autodiff playground
Everything so far fits in a few dozen lines of code: parse a formula, turn it into a graph, run the forward pass, run the backward pass. This is what PyTorch and JAX do, at a much larger scale and speed. Here is the small version, running in your browser. Type a formula in $x$ and $y$ and watch the graph.
Type x*y + sin(x) with $x = 1$, $y = 2$.
- Graph: $v_1 = x\cdot y$, $v_2 = \sin x$, $f = v_1 + v_2$. Values: $v_1 = 2$, $v_2 = 0.8415$, $f = 2.8415$.
- Backward: $\bar f = 1$. The add node sends $1$ to both. $\bar v_1 = 1$, $\bar v_2 = 1$.
- $x$ gets $\bar v_1\cdot y = 2$ from the product and $\bar v_2\cdot\cos x = 0.5403$ from the sine, and they add: $\partial f/\partial x = 2.5403$.
- $y$ gets $\bar v_1\cdot x = 1$.
- Check with the formula: $\partial f/\partial x = y + \cos x = 2 + 0.5403$ ✓.
What the playground does (a miniature automatic-differentiation system):
- Parse the text into a tree (a safe parser written for this page; nothing is executed as code).
- Build the graph: every operation becomes a node; repeated variables become one node with fan-out.
- Forward pass: compute and store every value.
- Backward pass: local derivative × upstream gradient, added at fan-out (reverse mode).
- Forward mode as a second opinion: two runs with seeds $\dot x=1$ and $\dot y=1$.
- Compare with finite differences and with the symbolic derivative (tidied).
Supported: + − * / ^, exp log sin cos tanh sigmoid relu sqrt, numbers, pi, e, and the names x and y. Write 2*x, not 2x.
Why do we need it?
Seeing a whole AD system in miniature removes the last bit of mystery: there is no magic in loss.backward(), only a graph, stored values and local rules.
Where is it used?
The same design underlies PyTorch autograd, JAX, TensorFlow, Autograd, Tinygrad and Karpathy-style "micrograd" teaching code. Writing your own tiny version is a common way to learn backpropagation.
How is it used?
Type formulas, change $x$ and $y$, and compare the three derivative columns. Test tricky cases: a variable used many times, relu at a negative input, log at a negative input (no value), deep nesting.
- Real frameworks do the same thing with tensors (whole arrays) as the values, so each node is a big matrix operation.
- At corners (like
reluat exactly 0) the derivative is not defined; the engine picks 0 and a finite difference may disagree there. - The symbolic column is for comparison only: it can be long, while AD never builds it.
Quick check: for x*x the graph has the node $x$ used twice. What does the backward pass do at $x$?
The multiply node sends back $x$ to each of its two inputs (the other factor each time). They are the same node, so the gradients add: $x + x = 2x$.
Vanishing and exploding gradients core
Backpropagation multiplies one local derivative per layer. Multiply many numbers that are all a bit smaller than 1 and the result goes to nearly zero: the gradient vanishes, and the early layers barely learn. Multiply many numbers a bit bigger than 1 and it blows up: the gradient explodes, and training becomes unstable.
Think of whispering a message down a line of 30 people, each repeating it a little more quietly (vanishing), or a little more loudly (exploding). The first person's message does not reach the end in a useful form.
The sigmoid's slope $\sigma'(z) = \sigma(1-\sigma)$ is at most $0.25$ (at $z=0$). A chain of 10 sigmoid layers with weights near 1 multiplies the gradient by at most $0.25^{10}$:
- $0.25^2 = 0.0625$, $\;0.25^5 = 0.000977$, $\;0.25^{10} = 0.00000095 \approx 10^{-6}$.
- So the gradient reaching layer 1 is about a millionth of the gradient at the last layer.
- ReLU's slope is exactly $1$ for positive inputs, so $1^{10} = 1$: the signal passes (but blocks to $0$ where the input is negative).
- If the per-layer factor is $1.5$ instead, $1.5^{30} \approx 190{,}000$: the gradient explodes.
Along a chain, the gradient reaching layer $\ell$ from the loss is a product of one factor per layer in between:
$$\frac{\partial L}{\partial\mathbf{h}_{\ell}} = W_{\ell+1}^\top D_{\ell+1}\,W_{\ell+2}^\top D_{\ell+2}\cdots W_L^\top D_L\,\frac{\partial L}{\partial\mathbf{h}_L},\qquad D_k=\mathrm{diag}\big(\varphi'(\mathbf{z}_k)\big).$$Its size behaves roughly like $(\text{typical weight size}\times\text{typical activation slope})^{\text{number of layers}}$. If that base is below 1 the gradient vanishes; above 1 it explodes.
Remedies. ReLU-type activations (slope 1 where active), careful weight initialisation (Xavier or He: scale the weights so the base is near 1), normalisation layers, residual connections (a "+1" path), gradient clipping (cap the size of an exploding gradient), and LSTM/GRU gates in recurrent networks.
Why do we need it?
It explains why very deep networks were hard to train for years, why sigmoid hidden layers fell out of favour, and why tricks like ResNets and normalisation matter.
Where is it used?
Diagnosing a network that "does not learn" (layer-wise gradient norms in TensorBoard or Weights & Biases), choosing initialisation and activations, and gradient clipping when training RNNs and Transformers.
How is it used?
Watch the gradient size per layer. If early layers have gradients many orders of magnitude smaller than late ones, change the activation or initialisation. If the loss suddenly becomes NaN or huge, clip gradients and lower the learning rate.
- The numbers depend on the random weights, but the trend does not: a per-layer factor away from 1 compounds exponentially with depth.
- ReLU can still lose signal: a neuron that is negative for every input has zero gradient forever ("dead ReLU").
- Vanishing and exploding gradients are properties of the product, not of backprop itself; any method that multiplies many Jacobians has them.
Quick check: a 20-layer chain where each layer multiplies the gradient by 0.8. What reaches the first layer?
$0.8^{20} \approx 0.0115$: about 1% of the gradient at the last layer. With a factor of $0.5$ it would be $0.5^{20}\approx10^{-6}$.
Gradient checking: trust, but verify core
Backprop code is easy to get subtly wrong: a missing transpose, a forgotten factor, an overwritten gradient. The loss still goes down a little, so nothing crashes. The test that catches these bugs uses the oldest definition of a derivative: nudge the weight a little in each direction, and see how the loss changes. If it matches what backprop said, the code is right.
It is slow (two extra forward passes per weight), so you do it on a tiny model, a handful of weights, once, while developing. Then you switch it off.
Tiny network from before, weight $W_2[1] = 0.5$. Backprop said $\partial L/\partial W_2[1] = -0.5014$.
- Nudge: $W_2[1] = 0.5 + 10^{-5}$ gives $L(+) = 0.301483\ldots$ and $W_2[1] = 0.5 - 10^{-5}$ gives $L(-) = 0.301493\ldots$ (computed with the full forward pass each time).
- Central difference: $(L(+) - L(-))/(2\cdot10^{-5}) = -0.50136$.
- Relative error: $\dfrac{|a - n|}{|a| + |n|}$ with $a = -0.501362$ and $n = -0.501362$ (they agree to about 12 digits) is below $10^{-12}$. That is tiny, so the gradient passes.
- If we forget the sigmoid slope in the hidden layer, $\partial L/\partial W_1[1,1]$ comes out as $-0.3883$ instead of $-0.0888$, a relative error near $0.6$: the check fails and flags the bug.
Gradient check. For each weight $w_i$ compare the backprop value $a_i$ with the numerical value $n_i = \dfrac{L(w_i+h) - L(w_i-h)}{2h}$, using a small $h$ (about $10^{-5}$ to $10^{-6}$ in 64-bit arithmetic), via the relative error
$$\text{rel}_i = \frac{|a_i - n_i|}{|a_i| + |n_i|}.$$Typical reading: below $10^{-7}$ excellent; around $10^{-5}$ fine for a non-smooth model; above $10^{-3}$ probably a bug. Practical rules: use 64-bit floats, check a few random weights, avoid points where the model has a corner (like ReLU at 0), and turn off dropout and other randomness while checking.
Why do we need it?
A wrong gradient does not crash the program; it silently trains a worse model. A gradient check is the cheapest way to find out whether your backward code matches the true derivative of your forward code.
Where is it used?
Whenever you write a custom layer or loss: torch.autograd.gradcheck, JAX's check_grads, the classic Stanford CS231n and Andrew Ng exercises, and unit tests of numerical libraries.
How is it used?
Pick a few weights. Compute the backprop gradient and the central difference for each. Print the relative error and require it to be tiny. Fix the bug and test again.
- A passing check on a few weights is strong evidence, not proof. Check different weights and a few random inputs.
- Do not use gradient checking during real training: it needs $2n$ forward passes.
- If the check fails only for some weights, look at the layers they belong to (and at corners such as ReLU at $0$).
Quick check: why use the central difference $\frac{L(w+h)-L(w-h)}{2h}$ rather than $\frac{L(w+h)-L(w)}{h}$?
Its error shrinks like $h^2$ instead of $h$, so at the same $h$ it is much more accurate, and the relative error of a correct gradient is small enough to separate it from a bug.
What the gradient does: a loss surface in 3D
All of this work has one purpose: to know which way to move the weights so the loss goes down. Imagine the loss as a landscape: the two weights are the east and north positions, and the loss is the height. Backpropagation tells you the slope of the ground under your feet. Gradient descent takes a step downhill, and repeats.
One surprise: far out on a flat plateau the slope is almost zero, so the arrow is tiny and descent crawls. That is the vanishing gradient again, seen from above.
One sigmoid neuron: prediction $\hat y = \sigma(w x + b)$ with input $x = 2$, target $t = 1$ and loss $L = (\hat y - 1)^2$. At $(w, b) = (0, 0)$:
- Forward: $a = 2\cdot0 + 0 = 0$, $\hat y = \sigma(0) = 0.5$, $L = 0.25$.
- Backward: $\partial L/\partial\hat y = 2(\hat y - 1) = -1$; $\sigma'(0) = 0.25$; so $\partial L/\partial a = -0.25$.
- $\partial a/\partial w = x = 2$ and $\partial a/\partial b = 1$, so $\partial L/\partial w = -0.5$ and $\partial L/\partial b = -0.25$.
- Gradient descent with $\eta = 1$ moves to $(w,b) = (0.5, 0.25)$, where $a = 1.25$, $\hat y = 0.7773$ and $L = 0.0496$. The loss fell from $0.25$ to $0.0496$.
The gradient $\nabla L = (\partial L/\partial w,\ \partial L/\partial b)$ points in the direction of steepest increase. Gradient descent moves the opposite way:
$$(w, b) \leftarrow (w, b) - \eta\,\nabla L(w, b).$$Backpropagation is how $\nabla L$ is computed for a network with millions of weights; the widget below uses our autodiff engine to get it for this neuron.
Why do we need it?
The gradient is the compass of training. Seeing it on a surface connects the abstract numbers from the backward pass to the actual downhill move the weights make.
Where is it used?
Every optimiser (SGD, momentum, Adam) takes the backprop gradient and decides how to step. Plateaus, valleys and saddles of real loss surfaces explain slow training.
How is it used?
Compute the gradient with backprop, step against it with a learning rate $\eta$, and check that the loss falls. Too large a step overshoots; too small crawls.
Quick check: at $(w,b)=(0,0)$ the gradient is $(-0.5,-0.25)$. Which way does gradient descent move $w$ and $b$ with $\eta = 1$?
$w \leftarrow 0 - 1\cdot(-0.5) = 0.5$ and $b \leftarrow 0 - 1\cdot(-0.25) = 0.25$. Both increase, because the loss falls when $w$ and $b$ increase (the neuron needs a larger output to reach the target $1$).
Recap, cheat sheet and practice
- Three ways to differentiate: symbolic (a formula, may swell), numerical (approximate, step-size trouble, good for checking) and automatic (exact, follows the program, cheap).
- Computational graph: one elementary operation per node. The forward pass computes and stores every value. Each node has a local derivative that is a one-line rule.
- Backward pass: set $\partial L/\partial L = 1$, then in reverse order send (upstream gradient) × (local derivative) to each input, adding where a value is used twice. This is reverse-mode differentiation.
- Backpropagation = the chain rule applied from the output backwards, with the shared partial products saved and reused. Layer rules: dense $\partial W = \mathbf{g}\mathbf{x}^\top$, $\partial\mathbf{x} = W^\top\mathbf{g}$; activation $\mathbf{g}\odot\varphi'(\mathbf{z})$.
- Forward mode runs the program on dual numbers $a+b\varepsilon$ ($\varepsilon^2=0$) and costs one pass per input; reverse mode costs one pass per output, so a scalar loss needs only one backward pass.
- Cost (rule of thumb): backward $\approx 2$–$3\times$ forward in time; memory = the stored activations (checkpointing trades time for memory).
- Trouble and checks: a product of many slopes below (above) 1 makes gradients vanish (explode); ReLU, good initialisation, residual links, normalisation and clipping help. Gradient checking with central differences catches backward-pass bugs.
Cheat sheet
| Idea | Formula / rule | Remember |
|---|---|---|
| Backward step | $\bar v_j \mathrel{+}= \bar v_i\,\dfrac{\partial v_i}{\partial v_j}$ | upstream × local, add at fan-out |
| Add node | passes the gradient to both inputs | local derivatives 1, 1 |
| Multiply node | sends back the other input | $\partial(ab)/\partial a = b$ |
| Sigmoid / tanh / ReLU | $\sigma(1-\sigma)$ / $1-\tanh^2$ / $1$ or $0$ | reuse the stored output |
| Dense layer | $\partial W=\mathbf{g}\mathbf{x}^\top,\ \partial\mathbf{b}=\mathbf{g},\ \partial\mathbf{x}=W^\top\mathbf{g}$ | gradient has the shape of the thing |
| Activation layer | $\partial\mathbf{z}=\mathbf{g}\odot\varphi'(\mathbf{z})$ | entry-wise product |
| Dual numbers | $f(a+\varepsilon)=f(a)+f'(a)\varepsilon$ | nudge part = derivative |
| Forward vs reverse | $n$ passes vs $m$ passes | loss: $m=1$, use reverse |
| Training step cost | $\approx 3\times$ one forward pass (rule of thumb) | memory = stored activations |
| Gradient check | $\dfrac{|a-n|}{|a|+|n|}$ with $n=\dfrac{L(w+h)-L(w-h)}{2h}$ | tiny model, 64-bit, $h\approx10^{-5}$ |
import math
class Value:
"""A number that remembers how it was made (a node of the computational graph)."""
def __init__(self, data, parents=(), local=()):
self.data = data
self.grad = 0.0
self.parents = parents # nodes this one was computed from
self.local = local # local derivatives d(self)/d(parent)
def __add__(self, o):
o = o if isinstance(o, Value) else Value(o)
return Value(self.data + o.data, (self, o), (1.0, 1.0))
def __mul__(self, o):
o = o if isinstance(o, Value) else Value(o)
return Value(self.data * o.data, (self, o), (o.data, self.data))
def __sub__(self, o): return self + (o * -1.0)
def __pow__(self, k): return Value(self.data ** k, (self,), (k * self.data ** (k - 1),))
def sigmoid(self):
s = 1 / (1 + math.exp(-self.data))
return Value(s, (self,), (s * (1 - s),))
def relu(self):
return Value(max(0.0, self.data), (self,), (1.0 if self.data > 0 else 0.0,))
def backward(self):
order, seen = [], set()
def visit(v): # topological order: parents before children
if id(v) not in seen:
seen.add(id(v))
for p in v.parents: visit(p)
order.append(v)
visit(self)
self.grad = 1.0 # dL/dL = 1
for v in reversed(order): # walk backwards
for p, loc in zip(v.parents, v.local):
p.grad += v.grad * loc # upstream * local, ADDED at fan-out
# --- the tiny 2-2-1 network of this chapter, built from Value objects ---
def loss(params):
W1a, W1b, W1c, W1d, b1a, b1b, W2a, W2b, b2 = params
x1, x2, t = 1.0, 2.0, 1.0
z1 = W1a * x1 + W1b * x2 + b1a
z2 = W1c * x1 + W1d * x2 + b1b
y = W2a * z1.sigmoid() + W2b * z2.sigmoid() + b2
return (y - t) ** 2 * 0.5
vals = [0.1, 0.2, -0.3, 0.4, 0.1, -0.1, 0.5, -0.5, 0.2]
params = [Value(v) for v in vals]
L = loss(params)
L.backward()
print(round(L.data, 4)) # 0.3015
print([round(p.grad, 4) for p in params])
# [-0.0888, -0.1777, 0.0933, 0.1866, -0.0888, 0.0933, -0.5014, -0.4649, -0.7765]
# --- gradient check: central finite differences ---
def numeric_grad(vals, i, h=1e-6):
up = [Value(v + (h if j == i else 0)) for j, v in enumerate(vals)]
dn = [Value(v - (h if j == i else 0)) for j, v in enumerate(vals)]
return (loss(up).data - loss(dn).data) / (2 * h)
err = max(abs(p.grad - numeric_grad(vals, i)) for i, p in enumerate(params))
print(err < 1e-8) # True: backprop matches finite differences
# --- fan-out: f = x*x uses x twice, so the two gradients ADD ---
x = Value(3.0)
f = x * x
f.backward()
print(f.data, x.grad) # 9.0 6.0
1. Which method gives an exact derivative (up to rounding), works on ordinary program code with loops and if-statements, and costs only a small multiple of one function evaluation for a scalar loss?
2. In the backward pass, a multiply node $v = a\cdot b$ receives upstream gradient $g$. What does it send to input $a$?
3. A value is used in two places. The gradients coming back along the two edges are $3$ and $4$. What is the gradient of that value?
4. Roughly how does the cost of the backward pass compare with the forward pass for a typical network?
5. With dual numbers ($\varepsilon^2 = 0$), what is $(2+\varepsilon)(5+3\varepsilon)$?
6. The slope of the sigmoid is at most $0.25$. Through 12 sigmoid layers (weights of size about 1), about how small can the gradient become at the first layer relative to the last?
Practice problems
A. A neuron has $p = w\cdot x$, $e = p - t$, $L = e^2$ with $w = 2$, $x = 3$, $t = 5$. Run the forward and backward passes and give $\partial L/\partial w$ and $\partial L/\partial x$.
Forward: $p = 6$, $e = 1$, $L = 1$. Backward: $\bar L = 1$; $\bar e = 2e = 2$; $\bar p = \bar e\cdot1 = 2$; $\bar w = \bar p\cdot x = 6$ and $\bar x = \bar p\cdot w = 4$. Check: $L = (wx-t)^2$ gives $\partial L/\partial w = 2(wx-t)x = 2\cdot1\cdot3 = 6$ ✓ and $\partial L/\partial x = 2\cdot1\cdot2 = 4$ ✓.
B. Use dual numbers to find $f(2)$ and $f'(2)$ for $f(x) = x^3 - 2x$.
$(2+\varepsilon)^3 = 8 + 12\varepsilon$ (higher powers vanish) and $2(2+\varepsilon) = 4 + 2\varepsilon$. So $f(2+\varepsilon) = (8-4) + (12-2)\varepsilon = 4 + 10\varepsilon$. Thus $f(2) = 4$ and $f'(2) = 10 = 3\cdot4 - 2$ ✓.
C. A deep chain of 8 tanh layers has an average factor of $0.7$ per layer. What fraction of the gradient reaches the first layer? What if the factor were $1.3$?
$0.7^8 \approx 0.058$: under 6% of the gradient survives (vanishing, mildly). With $1.3$: $1.3^8 \approx 8.2$: the gradient is amplified 8-fold (exploding, mildly). Over 40 layers the same factors give $0.7^{40}\approx6\times10^{-7}$ and $1.3^{40}\approx36{,}000$.
D. A dense layer has $W = \begin{bmatrix}2 & 0\\1 & -1\end{bmatrix}$, input $\mathbf{x} = (1, 3)$ and upstream gradient $\mathbf{g} = (1, 2)$. Find $\partial L/\partial W$ and $\partial L/\partial\mathbf{x}$.
$\partial L/\partial W = \mathbf{g}\mathbf{x}^\top = \begin{bmatrix}1\cdot1 & 1\cdot3\\ 2\cdot1 & 2\cdot3\end{bmatrix} = \begin{bmatrix}1 & 3\\2 & 6\end{bmatrix}$. $\partial L/\partial\mathbf{x} = W^\top\mathbf{g} = \begin{bmatrix}2 & 1\\0 & -1\end{bmatrix}\begin{bmatrix}1\\2\end{bmatrix} = (2+2,\ 0-2) = (4, -2)$. Shapes: $2\times2$ and $2\times1$ ✓.
E. A model has $10^7$ weights and one loss. Compare the number of forward-pass-equivalents needed for the full gradient by central finite differences and by backpropagation (take a backward pass as about 3 forward passes).
Finite differences: $2\times10^7$ evaluations. Backprop: about 3. The ratio is about $7\times10^6$: backprop is millions of times cheaper, and exact.
F. A gradient check for one weight gives backprop $= 0.2500$ and central difference $= 0.5000$. Compute the relative error and say what it suggests.
$\text{rel} = \dfrac{|0.25 - 0.5|}{0.25 + 0.5} = \dfrac{0.25}{0.75} = 0.33$. That is far above $10^{-3}$, so the backward code is wrong. The numerical value is exactly twice the backprop value, which suggests a missing factor of 2 (for example the derivative of $r^2$ was taken as $r$).
Higher-Order Derivatives
The derivative tells you how steep the ground is. The second derivative tells you how the steepness itself is changing: is the ground curving up like a bowl, curving down like a hill, or twisting like a mountain pass? That one extra piece of information decides whether a flat spot is a valley, and how big a step training can safely take.
- Read the second derivative as concavity (smile or frown) and as acceleration; meet the third and higher derivatives
- Build the Hessian matrix of second partial derivatives, and know when mixed partials are equal
- Measure curvature in any direction with $\mathbf{u}^\top H\mathbf{u}$ and read the Hessian through its eigenvalues
- Tell a bowl (positive definite), a hill (negative definite) and a saddle (indefinite) apart, and apply the second-derivative test
- Use curvature to take a smarter step (Newton's method) and to choose a safe learning rate ($\eta \lt 2/\lambda_{\max}$)
The second derivative: concavity and acceleration core
Think about driving. Your position tells where you are. Your speed tells how fast the position is changing (that is the derivative, from Chapter 2.3). Your acceleration tells how fast the speed is changing: press the gas and it is positive, brake and it is negative.
Acceleration is "the derivative of the derivative". That is exactly what a second derivative is. Nothing new is needed: you just differentiate twice.
Now picture a curve as a road. The first derivative is how steep the road is right now. The second derivative says whether the road is getting steeper (curving upward like a smile) or flattening out (curving downward like a frown).
A curve. Let $f(x) = x^3$.
- First derivative (power rule): $f'(x) = 3x^2$.
- Differentiate again: $f''(x) = 6x$.
- At $x = 2$: $f''(2) = 12 \gt 0$, so the curve is a smile there (the slope is growing).
- At $x = -2$: $f''(-2) = -12 \lt 0$, a frown (the slope is shrinking).
- At $x = 0$: $f''(0) = 0$. The curve switches from frown to smile right here.
Numeric check by nudging $x$ by $0.01$: the slope at $2.01$ is $3(2.01)^2 = 12.1203$ and at $1.99$ it is $3(1.99)^2 = 11.8803$. The slope changes by $0.24$ over a step of $0.02$, so the rate of change of the slope is $0.24/0.02 = 12$ ✓.
A falling ball. A dropped ball has fallen $s(t) = 4.9\,t^2$ metres after $t$ seconds. Speed: $s'(t) = 9.8\,t$ metres per second. Acceleration: $s''(t) = 9.8$ metres per second per second, a constant (gravity).
The second derivative of $f$ is the derivative of its derivative:
$$f''(x) \;=\; \frac{d}{dx}\bigl(f'(x)\bigr) \;=\; \lim_{h\to 0}\frac{f'(x+h) - f'(x)}{h}, \qquad\text{also written}\quad \frac{d^2 f}{dx^2}.$$The notation $\dfrac{d^2 f}{dx^2}$ reads "d two f over d x squared": apply $\frac{d}{dx}$ twice.
- $f''(x) \gt 0$: the graph is concave up (a smile; it would "hold water"). The slope is increasing.
- $f''(x) \lt 0$: the graph is concave down (a frown). The slope is decreasing.
- A point where $f''$ changes sign is an inflection point (the bend switches direction).
How to estimate it from values only (derivation). The slope halfway between $x$ and $x+h$ is about $\frac{f(x+h)-f(x)}{h}$, and the slope halfway between $x-h$ and $x$ is about $\frac{f(x)-f(x-h)}{h}$. These two slopes are $h$ apart, so the rate of change of the slope is their difference divided by $h$:
$$f''(x) \approx \frac{1}{h}\left(\frac{f(x+h)-f(x)}{h} - \frac{f(x)-f(x-h)}{h}\right) = \frac{f(x+h) - 2f(x) + f(x-h)}{h^2}.$$Check with $f = x^3$, $x = 2$, $h = 0.1$: $\frac{2.1^3 - 2\cdot 8 + 1.9^3}{0.01} = \frac{9.261 - 16 + 6.859}{0.01} = \frac{0.12}{0.01} = 12$ ✓.
Why do we need it?
The slope alone cannot tell a valley from a hilltop: both have slope zero. Whether the slope is growing or shrinking around a flat spot is what tells them apart. It also tells us whether we are speeding up or slowing down.
Where is it used?
Physics (acceleration), checking whether a loss function is convex ($f'' \ge 0$ everywhere), the second-derivative test for minima, Newton's method, and limits on the learning rate in gradient descent.
How is it used?
Differentiate twice and look at the sign: positive means smile (the bottom of a bowl is nearby), negative means frown. The size tells how quickly the slope changes. In code, use the three-point formula above or an autodiff library.
$f''=0$ does not always mean an inflection point. For $f(x)=x^4$ we get $f''(x)=12x^2$, which is $0$ at $x=0$, yet the curve is a smile on both sides. An inflection needs $f''$ to change sign.
"Concave up" is about the bend, not about being positive. A curve can be below the axis and still smile. The sign of $f''$ only describes how the slope changes.
Quick check: $f(x) = x^2 - 4x$. What is $f''$, and what does it say about the shape?
$f'(x) = 2x - 4$ and $f''(x) = 2$. It is positive everywhere, so $f$ is a smile (a parabola opening upward) at every point, with its lowest point where $f' = 0$, at $x = 2$.
The third derivative and higher-order derivatives
Nothing stops us from differentiating again. The derivative of the acceleration is called the jerk: it is the sudden jolt you feel when a driver stamps on the brake. A smooth ride has a small jerk. That is the third derivative.
And we can keep going: fourth, fifth, and so on. Each one tells how fast the previous one is changing. It is like a ladder: each rung is the slope of the rung below.
Some functions run out of rungs quickly (a polynomial eventually reaches 0). Some never change ($e^x$). Some repeat in a loop ($\sin x$).
Differentiate three functions again and again.
| order | $x^4$ | $e^x$ | $\sin x$ |
|---|---|---|---|
| $f$ | $x^4$ | $e^x$ | $\sin x$ |
| $f'$ | $4x^3$ | $e^x$ | $\cos x$ |
| $f''$ | $12x^2$ | $e^x$ | $-\sin x$ |
| $f'''$ | $24x$ | $e^x$ | $-\cos x$ |
| $f^{(4)}$ | $24$ | $e^x$ | $\sin x$ (back to the start) |
| $f^{(5)}$ | $0$ | $e^x$ | $\cos x$ |
Notice: $x^4$ reaches a constant, then zero. $e^x$ is its own derivative forever. $\sin x$ repeats every 4 steps. The fourth derivative of $x^4$ is $24 = 4\cdot3\cdot2\cdot1 = 4!$ (read "four factorial": the product of all whole numbers from 4 down to 1). Factorials will matter in Chapter 2.11.
The $n$-th derivative is the derivative taken $n$ times:
$$f^{(n)}(x) = \frac{d^n f}{dx^n}, \qquad f^{(0)} = f,\quad f^{(1)} = f',\quad f^{(2)} = f'',\quad f^{(3)} = f''' ,\ \dots$$The small number in brackets is the order, so that it is not mistaken for a power. Useful patterns:
- A polynomial of degree $n$ has $f^{(n)} = n!\times(\text{leading coefficient})$, and every derivative after that is $0$.
- $\dfrac{d^n}{dx^n}e^x = e^x$, and $\sin$/$\cos$ cycle with period 4.
- For $\ln(1+x)$: $f' = \frac{1}{1+x}$, $f'' = -\frac{1}{(1+x)^2}$, $f''' = \frac{2}{(1+x)^3}$, and in general $f^{(n)} = \dfrac{(-1)^{n-1}(n-1)!}{(1+x)^n}$.
A function whose derivatives all exist is called smooth. $\mathrm{ReLU}(x) = \max(0,x)$ is not smooth: it has a kink at $0$. Smooth replacements such as softplus and GELU exist partly for this reason.
Why do we need it?
Each higher derivative adds one more detail about the shape of a function near a point: height, slope, bend, how the bend changes... Together they let us rebuild the function nearby. That is the idea behind Taylor series (next chapter).
Where is it used?
Taylor series (the coefficient of $x^n$ uses $f^{(n)}(0)/n!$), smoothness requirements of optimisers, jerk limits in robot and vehicle motion planning, and comparing ReLU (kinked) with smooth activations such as GELU or softplus.
How is it used?
Differentiate repeatedly and look for a pattern (it often repeats or ends). On a computer, symbolic tools (SymPy diff(f, x, n)) or autodiff do it exactly; finite differences get noisy very quickly as the order grows.
Do not use many finite differences in a row. Each nudge divides a tiny number by another tiny number, so rounding noise grows with the order. Fourth and fifth derivatives from raw values are usually garbage. Use symbolic maths or autodiff.
Quick check: what is the 10th derivative of $x^9$?
$0$. A polynomial of degree 9 has $f^{(9)} = 9!$ (a constant), so $f^{(10)} = 0$.
The Hessian and the Hessian matrix core
With one input there is one slope and one bend. With two inputs (a landscape with east-west $x$ and north-south $y$), the ground has more to say. Walk east: does the slope eastward grow or shrink? Walk north: does the slope northward change? And there is a cross effect: as you walk north, does the east-west slope change? (A twisted roof tile does this.)
Each of those questions is a second partial derivative. For two inputs there are $2\times2 = 4$ of them. We write them in a small table, and that table is the Hessian. It is the "second derivative" for functions of many variables.
In the gradient chapter you collected the first partial derivatives into a column. Now we collect the second ones into a square.
Let $f(x,y) = x^2 y + y^2$. Find the Hessian at the point $(1, 2)$.
- First partials: $f_x = \dfrac{\partial f}{\partial x} = 2xy$ and $f_y = \dfrac{\partial f}{\partial y} = x^2 + 2y$.
- Differentiate $f_x = 2xy$ with respect to $x$: $f_{xx} = 2y$. With respect to $y$: $f_{xy} = 2x$.
- Differentiate $f_y = x^2 + 2y$ with respect to $x$: $f_{yx} = 2x$. With respect to $y$: $f_{yy} = 2$.
- At $(1,2)$: $f_{xx} = 2\cdot2 = 4$, $f_{xy} = f_{yx} = 2\cdot1 = 2$, $f_{yy} = 2$.
Notice that the two off-diagonal entries are equal. That is not an accident (next section).
For $f:\mathbb{R}^n \to \mathbb{R}$ the Hessian matrix is the $n\times n$ table of all second partial derivatives:
$$H(\mathbf{x}) = \nabla^2 f(\mathbf{x}) = \begin{bmatrix} \dfrac{\partial^2 f}{\partial x_1^2} & \dfrac{\partial^2 f}{\partial x_1\partial x_2} & \cdots \\[2mm] \dfrac{\partial^2 f}{\partial x_2\partial x_1} & \dfrac{\partial^2 f}{\partial x_2^2} & \cdots \\ \vdots & \vdots & \ddots \end{bmatrix}, \qquad H_{ij} = \frac{\partial^2 f}{\partial x_i\,\partial x_j}.$$- It is the Jacobian of the gradient. The gradient $\nabla f$ (a column vector, as always in this guide) is a function from $\mathbb{R}^n$ to $\mathbb{R}^n$. Its Jacobian (an $n\times n$ matrix; Chapter 2.5) has entry $(i,j) = \partial(\partial f/\partial x_i)/\partial x_j$, which is exactly $H_{ij}$.
- Shape: $n\times n$. For $n = 1$ it is the $1\times1$ matrix $[f''(x)]$, so the Hessian really is the second derivative again.
- It is symmetric for the smooth functions of ML ($H_{ij} = H_{ji}$, see next section), so it has real eigenvalues and perpendicular eigenvectors (Linear Algebra: eigenvalues).
- The symbol $\nabla^2 f$ is "del squared f". (Some books use it for the Laplacian, the trace of $H$. Here it always means the Hessian.)
Why do we need it?
The gradient only knows the slope. To know whether the ground curves up or down, and in which directions, we need all the second derivatives. One number is not enough once there are many inputs, so we need a matrix.
Where is it used?
Newton's method and quasi-Newton optimisers (L-BFGS), the second-derivative test, the Laplace approximation in Bayesian models, Fisher information, second-order Taylor expansions, and analyses of why training is slow or unstable.
How is it used?
Differentiate the gradient again, entry by entry, and arrange the results in an $n\times n$ symmetric matrix. Then look at its eigenvalues. In code: jax.hessian(f)(x), torch.autograd.functional.hessian, or a finite-difference grid.
The Hessian has $n^2$ entries. A network with a million weights would need a table of $10^{12}$ numbers (about 4 000 gigabytes, i.e. 4 terabytes, at 4 bytes each, and much more work to invert). That is why deep learning mostly avoids the full Hessian and uses first-order methods, or tricks such as Hessian-vector products that never build the whole matrix.
The Hessian is not the Jacobian of a vector function. It belongs to a scalar function $f:\mathbb{R}^n\to\mathbb{R}$, such as a loss. (A vector function has a Jacobian, an $m\times n$ matrix, and no single Hessian.)
Quick check: for $f(x,y)=x^2+3xy+2y^2$, what is $H$?
$f_x = 2x+3y$, $f_y = 3x+4y$. So $f_{xx} = 2$, $f_{xy} = 3$, $f_{yx} = 3$, $f_{yy} = 4$, giving $H = \begin{bmatrix}2&3\\3&4\end{bmatrix}$ at every point (the function is a quadratic, so its second derivatives are constants).
Mixed partial derivatives: does the order matter?
A mixed partial derivative takes one derivative with respect to $x$ and one with respect to $y$. There are two ways to do it: differentiate in $x$ first and then $y$, or in $y$ first and then $x$. Does it matter?
Think of the ground as a twisted sheet. The "twist" is one single fact about the sheet. You can measure it by asking "how does the east-west slope change as I go north?" or "how does the north-south slope change as I go east?". Both questions measure the same twist, so the answers agree.
There is a neat picture. Take four corners of a small square: $(x,y)$, $(x+h,y)$, $(x,y+k)$, $(x+h,y+k)$. The quantity $f(x+h,y+k) - f(x+h,y) - f(x,y+k) + f(x,y)$ is "(top edge difference) minus (bottom edge difference)" and also "(right edge difference) minus (left edge difference)". It is one number, however you group it. Divided by $hk$ and shrunk to a point, it becomes the mixed partial.
Let $f(x,y) = x^2y^3$.
- $x$ first: $f_x = 2xy^3$. Then $y$: $f_{xy} = \dfrac{\partial}{\partial y}(2xy^3) = 6xy^2$.
- $y$ first: $f_y = 3x^2y^2$. Then $x$: $f_{yx} = \dfrac{\partial}{\partial x}(3x^2y^2) = 6xy^2$.
- Same answer: $f_{xy} = f_{yx} = 6xy^2$ ✓.
(Notation warning: $f_{xy}$ usually means "$x$ first, then $y$", and $\frac{\partial^2 f}{\partial y\,\partial x}$ is written right to left. Because the two agree for smooth functions, you rarely need to worry.)
Equality of mixed partials (Schwarz / Clairaut). If the second partial derivatives $f_{xy}$ and $f_{yx}$ exist and are continuous around a point, then at that point
$$\frac{\partial^2 f}{\partial x\,\partial y} = \frac{\partial^2 f}{\partial y\,\partial x}.$$For a function of $n$ variables this says $H_{ij} = H_{ji}$: the Hessian is a symmetric matrix.
The continuity condition cannot be dropped. The function $f(x,y) = \dfrac{xy(x^2-y^2)}{x^2+y^2}$ (with $f(0,0)=0$) is a famous counterexample: at the origin $f_{xy}(0,0) = -1$ but $f_{yx}(0,0) = +1$. Why: along the $y$-axis ($x=0$) its $x$-slope is $f_x(0,y) = -y$, so differentiating that in $y$ gives $-1$; along the $x$-axis its $y$-slope is $f_y(x,0) = x$, so differentiating in $x$ gives $+1$. Its second derivatives are not continuous at the origin.
Why do we need it?
It cuts the work nearly in half and guarantees a nice matrix. A symmetric Hessian has real eigenvalues and perpendicular eigenvectors, so "curvature along each axis" is a meaningful picture.
Where is it used?
Every optimiser that uses the Hessian (Newton, L-BFGS, natural gradient), covariance and Fisher matrices (also symmetric), and checking hand-derived gradients in code by confirming that the Hessian comes out symmetric.
How is it used?
Compute only the upper triangle and copy it. If a computed Hessian is not symmetric, there is a bug (or the function is not smooth, as with ReLU networks at a kink).
Equal mixed partials are a statement about smooth functions. Polynomials, $e^x$, $\log$, sigmoid, tanh, softmax and squared error are all smooth, so their Hessians are symmetric. Networks that use ReLU have kinks, so their second derivatives are only defined away from the kinks (and the ReLU itself adds no curvature there, because it is a straight line on each side).
Quick check: $f(x,y) = e^{xy}$. Find $f_{xy}$ and $f_{yx}$.
$f_x = y\,e^{xy}$. Then $f_{xy} = e^{xy} + y\cdot x\,e^{xy} = (1+xy)e^{xy}$ (product rule). The other order: $f_y = x\,e^{xy}$ and $f_{yx} = e^{xy} + xy\,e^{xy} = (1+xy)e^{xy}$. Equal ✓.
Curvature: how sharply does it bend?
Drive around a bend. A gentle bend is part of a very big circle; a sharp hairpin is part of a very small circle. Curvature is "how tight the circle is". We pick the one circle that hugs the curve best at your spot, the curvature circle (also called the osculating circle, from the Latin for "kissing"). Its radius $R$ says how gently the road bends: big $R$ means gentle, small $R$ means sharp.
A straight road has an infinitely big circle (radius $\infty$): curvature zero.
In machine learning "curvature" almost always means the second derivative. At a flat spot ($f' = 0$) the two ideas are the same number.
The parabola $f(x) = \tfrac14 x^2$ at its bottom, $x = 0$.
- $f'(x) = \tfrac12 x$, so $f'(0) = 0$ (flat).
- $f''(x) = \tfrac12$.
- The curvature is $\kappa = \dfrac{f''}{(1 + f'^2)^{3/2}} = \dfrac{0.5}{(1+0)^{3/2}} = 0.5$.
- The curvature circle has radius $R = 1/\kappa = 2$. Its centre is $2$ above the bottom, at $(0, 2)$.
Away from the bottom, the slope is not zero. At $x = 2$: $f' = 1$, $f'' = 0.5$, so $\kappa = \dfrac{0.5}{(1+1)^{3/2}} = \dfrac{0.5}{2.828} = 0.177$ and $R \approx 5.66$. The parabola flattens out, so the circle gets bigger even though $f''$ stayed the same: the formula divides out the tilt.
For a curve $y = f(x)$ the signed curvature and the radius of curvature are
$$\kappa(x) = \frac{f''(x)}{\bigl(1 + f'(x)^2\bigr)^{3/2}}, \qquad R(x) = \frac{1}{|\kappa(x)|}.$$$\kappa \gt 0$ means the curve bends upward (a smile), $\kappa \lt 0$ downward. The centre of the circle lies on the concave side, at distance $R$ from the curve.
At a critical point ($f'=0$) the denominator is $1$, so $\kappa = f''$. This is why, in optimisation, we say "the curvature is $f''$". Away from flat spots the two differ by the tilt factor $(1+f'^2)^{3/2}$, but $f''$ still tells the sign and how fast the slope changes, which is what a step of gradient descent feels.
Why do we need it?
We want one number for "how sharply does this bend?". A sharp bend means a small safe step; a gentle bend means we can stride. Curvature gives a picture (a circle) for the second derivative.
Where is it used?
Road and rail design, computer graphics (smooth curves), robot path planning, and in ML the idea of sharp versus flat minima: a very sharp bowl is hard for gradient descent and may generalise worse.
How is it used?
Compute $f'$ and $f''$ at the point and plug into $\kappa$. For optimisation at a flat spot just read $f''$: large positive means a narrow, steep bowl; small positive means a wide, gentle bowl.
Quick check: a sharp bowl has $f'' = 10$ at its bottom; a wide one has $f''=0.1$. Which circle is bigger?
At the bottom $f'=0$, so $\kappa = f''$. Sharp bowl: $R = 1/10 = 0.1$. Wide bowl: $R = 1/0.1 = 10$. The wide bowl has the much bigger circle.
What the Hessian tells you: curvature in every direction core
Stand on a hilly landscape and pick a compass direction. Walk a short way along it and watch your height: it traces a curve, a smile or a frown. The curvature of that curve depends on the direction you chose. In a mountain pass it smiles along the ridge (towards the peaks) and frowns along the path that crosses the pass.
The Hessian packs the curvature of all directions into one matrix. Hand it a direction $\mathbf{u}$, and it gives back the curvature $\mathbf{u}^\top H\mathbf{u}$. Among all directions there is a most curved one and a least curved one: they are the eigenvectors of $H$, and the amounts of curvature are the eigenvalues.
Let $H = \begin{bmatrix}1&2\\2&1\end{bmatrix}$. Walk in three directions (unit vectors $\mathbf{u}$):
- East, $\mathbf{u} = [1, 0]$: $\mathbf{u}^\top H\mathbf{u} = 1\cdot1\cdot1 = 1$. (Just the top-left entry.)
- Northeast, $\mathbf{u} = [1,1]/\sqrt2$: $\mathbf{u}^\top H\mathbf{u} = \tfrac12(1 + 2 + 2 + 1) = 3$. A strong smile.
- Southeast, $\mathbf{u} = [1,-1]/\sqrt2$: $\tfrac12(1 - 2 - 2 + 1) = -1$. A frown.
The eigenvalues of $H$ are $3$ and $-1$, with eigenvectors along those last two diagonals. So this ground smiles most (3) along one diagonal, frowns most ($-1$) along the other, and everything else lies in between.
Curvature along a direction. Walk from $\mathbf{x}$ in the unit direction $\mathbf{u}$ and let $g(t) = f(\mathbf{x} + t\mathbf{u})$ be the height after $t$ steps. Derive its second derivative:
- The chain rule gives the slope: $g'(t) = \nabla f(\mathbf{x}+t\mathbf{u})^\top\mathbf{u} = \sum_i f_i(\mathbf{x}+t\mathbf{u})\,u_i$, where $f_i = \partial f/\partial x_i$.
- Differentiate again. Each $f_i(\mathbf{x}+t\mathbf{u})$ changes at rate $\sum_j f_{ij}\,u_j$, so $g''(t) = \sum_i\sum_j u_i\,f_{ij}\,u_j$.
- That double sum is exactly $\mathbf{u}^\top H\mathbf{u}$:
Extremes. Write $H = \sum_k \lambda_k \mathbf{v}_k\mathbf{v}_k^\top$ with orthonormal eigenvectors $\mathbf{v}_k$ (Quadratic Forms & Definiteness). Any unit $\mathbf{u} = \sum_k c_k\mathbf{v}_k$ has $\sum c_k^2 = 1$, so
$$\mathbf{u}^\top H\mathbf{u} = \sum_k \lambda_k c_k^2 = \text{a weighted average of the eigenvalues}.$$A weighted average lies between the smallest and the largest value: $\lambda_{\min} \le \mathbf{u}^\top H\mathbf{u} \le \lambda_{\max}$, with the ends reached exactly along the eigenvectors. Negative curvature in some direction exists exactly when $\lambda_{\min} \lt 0$.
Why do we need it?
In many dimensions, "is the ground curved up or down?" has no single answer: it depends on the direction. We need a tool that answers it for every direction at once, and finds the worst one.
Where is it used?
Finding the stiffest and flattest directions of a loss surface (these control the safe learning rate and the slow directions), detecting negative curvature to escape saddles, and the Laplace approximation, whose width along each eigen-direction is $1/\sqrt{\lambda}$.
How is it used?
Compute eigenvalues with np.linalg.eigvalsh(H) (it is symmetric). The largest tells the sharpest bend, the smallest the gentlest. For a specific direction use u @ H @ u.
The Hessian depends on the point. For a quadratic like the surface above it is the same everywhere. For a general function $H(\mathbf{x})$ changes from place to place, so "bowl" or "saddle" is a local statement.
Use unit vectors. $\mathbf{u}^\top H\mathbf{u}$ scales with the square of the length of $\mathbf{u}$. Always normalise $\mathbf{u}$ first if you want to compare directions fairly.
Quick check: $H = \begin{bmatrix}5&0\\0&1\end{bmatrix}$. Which direction curves most, and what is the curvature along $[0.6, 0.8]$?
The eigenvalues are $5$ (along $x$) and $1$ (along $y$), so the $x$-direction curves most. For $\mathbf{u} = [0.6, 0.8]$: $\mathbf{u}^\top H\mathbf{u} = 5(0.36) + 1(0.64) = 1.8 + 0.64 = 2.44$, between $1$ and $5$ as promised.
Positive definite Hessian: the bowl and the local minimum core
Put a marble on a flat spot. If the ground curves up in every direction around it (like the inside of a bowl), the marble stays: pushed any way, it rolls back. That flat spot is the bottom of a bowl, a local minimum.
"Curves up in every direction" means: the curvature $\mathbf{u}^\top H\mathbf{u}$ is positive for every direction $\mathbf{u}$. A Hessian with that property is called positive definite. It is the matrix version of "$f'' \gt 0$".
Let $f(x,y) = x^2 + xy + 2y^2$. Its gradient is $\nabla f = [2x + y,\; x + 4y]$, which is $[0,0]$ only at the origin. Its Hessian is $H = \begin{bmatrix}2&1\\1&4\end{bmatrix}$ everywhere.
- Trace $= 2 + 4 = 6$ and determinant $= 2\cdot4 - 1\cdot1 = 7$.
- The eigenvalues solve $\lambda^2 - 6\lambda + 7 = 0$, so $\lambda = 3 \pm \sqrt2 \approx 4.41$ and $1.59$. Both are positive: positive definite.
- Check by completing the square: $f = \left(x + \tfrac y2\right)^2 + \tfrac74 y^2$. A sum of squares, so $f \ge 0$, and $f = 0$ only at the origin. The origin is the lowest point ✓.
A symmetric matrix $H$ is positive definite (written $H \succ 0$) when $\mathbf{u}^\top H\mathbf{u} \gt 0$ for every $\mathbf{u}\neq\mathbf{0}$. Equivalently: all its eigenvalues are positive (Quadratic Forms & Definiteness).
Local minimum rule. If $\nabla f(\mathbf{c}) = \mathbf{0}$ and $H(\mathbf{c})$ is positive definite, then $\mathbf{c}$ is a (strict) local minimum. Reason, using the second-order Taylor expansion you will meet in Chapter 2.11: near $\mathbf{c}$,
$$f(\mathbf{c}+\boldsymbol\delta) \approx f(\mathbf{c}) + \underbrace{\nabla f(\mathbf{c})^\top\boldsymbol\delta}_{=\,0} + \tfrac12\,\boldsymbol\delta^\top H\,\boldsymbol\delta \;\gt\; f(\mathbf{c}).$$If $H(\mathbf{x})$ is positive semidefinite (eigenvalues $\ge 0$) at every point, $f$ is convex: a single bowl whose local minimum is the global minimum.
Why do we need it?
It is the guarantee that a flat spot is a true bottom, not a hilltop or a pass. It also tells us optimisation is "easy": a bowl has no traps.
Where is it used?
Least-squares and ridge regression (Hessian $2X^\top X/n$ is positive semidefinite, positive definite with independent features), logistic regression (convex), convex optimisation, and the requirement behind Newton's method that the step goes downhill.
How is it used?
At a point with zero gradient, test the eigenvalues of $H$ (np.linalg.eigvalsh), or try a Cholesky factorisation: it works exactly when the matrix is positive definite. All positive means a local minimum.
"Local" means local. A positive definite Hessian at one flat spot says nothing about other bowls elsewhere. A function can have many local minima (the wells in the classifier below).
A zero eigenvalue breaks the guarantee. If an eigenvalue is $0$ the bowl is a flat-bottomed trough in that direction and the test cannot decide. (We return to this in the second-derivative test.)
Quick check: is $H=\begin{bmatrix}2&3\\3&1\end{bmatrix}$ positive definite?
No. The determinant is $2\cdot1 - 3\cdot3 = -7 \lt 0$, so the eigenvalues have opposite signs (their product is the determinant). Along some direction the ground curves down. Positive diagonal entries are not enough.
Negative definite Hessian: the hill and the local maximum
Turn the bowl upside down and you get a hill. A marble balanced on the very top slides off if you nudge it in any direction. The top is a local maximum, and there the ground curves down in every direction: $\mathbf{u}^\top H\mathbf{u} \lt 0$ for all $\mathbf{u}$. We call that Hessian negative definite.
A hill and a bowl are the same shape flipped. So maximising $f$ is the same job as minimising $-f$.
Let $f(x,y) = -x^2 - xy - y^2$. Then $\nabla f = [-2x - y,\; -x - 2y]$, zero only at the origin, and $H = \begin{bmatrix}-2&-1\\-1&-2\end{bmatrix}$.
- Trace $= -4$, determinant $= 4 - 1 = 3$.
- $\lambda^2 + 4\lambda + 3 = 0$, so $\lambda = -1$ and $\lambda = -3$. Both negative: negative definite.
- Flip the sign: $-f = x^2 + xy + y^2$ has $H = \begin{bmatrix}2&1\\1&2\end{bmatrix}$ with eigenvalues $3$ and $1$, a bowl. Flipping $f$ flipped every eigenvalue ✓.
A symmetric matrix $H$ is negative definite ($H \prec 0$) when $\mathbf{u}^\top H\mathbf{u} \lt 0$ for every $\mathbf{u}\ne\mathbf{0}$, equivalently all eigenvalues are negative.
Local maximum rule. If $\nabla f(\mathbf{c}) = \mathbf{0}$ and $H(\mathbf{c}) \prec 0$, then $\mathbf{c}$ is a strict local maximum.
Link to minimising: $H_{-f} = -H_f$, so the eigenvalues of $-f$ are the negatives of those of $f$. Negative definite for $f$ means positive definite for $-f$.
Why do we need it?
Some problems are about going up: make a probability as large as possible. We need to recognise a true peak just as we recognise a true valley.
Where is it used?
Maximum-likelihood estimation (the log-likelihood should be a hill at the answer), the Laplace approximation (a Gaussian bump matched to the top of the log-posterior), expectation-maximisation, and reinforcement-learning objectives that are maximised.
How is it used?
Check the eigenvalues at the zero-gradient point: all negative means a local maximum. In practice we usually flip the sign and minimise the negative log-likelihood, which then has a positive definite Hessian.
Gradient descent never finds a maximum on purpose. It walks downhill, so a hilltop is an unstable resting place. Only an exact start at the top (or exact arithmetic with zero gradient) leaves it there.
Quick check: $f$ has $H = \begin{bmatrix}-4&0\\0&-1\end{bmatrix}$ at a critical point. What are the eigenvalues of the Hessian of $-f$?
Flip the signs: $4$ and $1$. Both positive, so $-f$ has a bowl there and $f$ has a hill.
Saddle points: flat, but neither a valley nor a hill core
Picture a horse saddle, or a mountain pass between two peaks. Walk along the ridge line of the pass (towards one peak) and you go up on both sides: here the ground smiles. Cross the pass sideways (down into the valleys either side) and it frowns: the pass itself is the highest point of that path. At the centre the ground is level (the gradient is zero), yet it is not a valley bottom and not a hilltop. That spot is a saddle point.
A marble placed exactly there balances for a moment. Nudge it along the smile direction and it rolls back; nudge it along the frown direction and it rolls away.
$f(x,y) = x^2 - y^2$.
- $\nabla f = [2x, -2y] = [0,0]$ only at the origin.
- $H = \begin{bmatrix}2&0\\0&-2\end{bmatrix}$, eigenvalues $2$ and $-2$: mixed signs.
- Along the $x$-axis: $f(x,0) = x^2$, a smile (the origin looks like a minimum). Along the $y$-axis: $f(0,y) = -y^2$, a frown (the origin looks like a maximum).
A tilted example: $H = \begin{bmatrix}2&4\\4&2\end{bmatrix}$ has trace $4$ and determinant $4-16 = -12 \lt 0$, so eigenvalues $6$ and $-2$: also a saddle.
A critical point ($\nabla f = \mathbf{0}$) where the Hessian is indefinite (it has at least one positive and at least one negative eigenvalue) is a saddle point. Along an eigenvector with $\lambda \gt 0$ the function smiles; along one with $\lambda \lt 0$ it frowns.
Why saddles dominate in high dimensions. A critical point of a function of $n$ variables has $n$ curvatures (eigenvalues). It is a minimum only if all $n$ are positive. If each sign were an independent fair coin flip, the chance would be $2^{-n}$: for $n = 20$ that is about one in a million. Real Hessians are not coin flips, but the same effect holds: the more directions there are, the likelier it is that at least one bends down. Neural-network losses have millions of directions, so we expect saddles to far outnumber true minima. Random-matrix arguments and experiments on small networks suggest that critical points at high loss are mostly saddles, while local minima tend to sit at low loss. This is a useful picture, not a theorem about every network.
What gradient descent does near a saddle. The gradient is tiny near the saddle, so progress nearly stops (a plateau), then the small component along the frown direction grows by a factor $(1 + \eta|\lambda|)$ per step and the iterate slides off. Noise (as in SGD) and momentum help it escape.
Why do we need it?
If we only checked "the gradient is zero", we might stop at a saddle and call it a solution. Knowing about saddles explains why training sometimes stalls on a plateau and then suddenly moves on.
Where is it used?
Understanding training curves of deep networks, saddle-escaping optimisers (noise injection, momentum, Adam), the geometry of loss landscapes, and game theory and GANs (a Nash equilibrium is a saddle of the joint objective).
How is it used?
At a flat spot compute the Hessian's eigenvalues. Mixed signs mean saddle, and the eigenvector with a negative eigenvalue is a direction that goes downhill: a clean way to escape.
This is a picture, not a proof about your network. Real Hessians are not random matrices, and trained networks usually reach points with a few slightly negative or near-zero eigenvalues. The lesson is the counting: in many dimensions "all curvatures positive" is a demanding condition.
Zero gradient does not mean "done". Check the Hessian's eigenvalues, or at least watch whether the loss is still falling slowly.
Quick check: a critical point has $H$ with eigenvalues $5, 2, 0.1, -0.01$. What is it?
A saddle point: one eigenvalue is negative (tiny, but negative), so there is a (very gently) downhill direction. Along that direction progress will be slow, which looks like a plateau in training.
The second-derivative test (1D and many dimensions) core
Finding a flat spot (gradient zero) is step one. Step two is to ask: what kind of flat spot? Look at the curvature there. Smile: valley bottom. Frown: hilltop. Mixed: saddle. Zero: the test cannot tell, so look harder.
In one dimension there is one curvature, $f''$. In many dimensions there are $n$ curvatures, the eigenvalues of $H$, and we look at all of them.
One variable. $f(x) = \tfrac14x^4 - x^2$.
- $f'(x) = x^3 - 2x = x(x^2 - 2)$, which is $0$ at $x = 0$ and $x = \pm\sqrt2$.
- $f''(x) = 3x^2 - 2$.
- $f''(0) = -2 \lt 0$: a local maximum. $f''(\pm\sqrt2) = 6 - 2 = 4 \gt 0$: two local minima.
Many variables. $f(x,y) = (x^2-1)^2 + y^2$ has $\nabla f = [4x(x^2-1),\, 2y]$ and $H = \begin{bmatrix}12x^2-4&0\\0&2\end{bmatrix}$. Critical points: $(\pm1, 0)$ with $H = \mathrm{diag}(8, 2)$ (both positive: two minima) and $(0,0)$ with $H = \mathrm{diag}(-4, 2)$ (mixed: a saddle between the wells).
1D test. Suppose $f'(c) = 0$.
- $f''(c) \gt 0$ $\Rightarrow$ local minimum. $f''(c) \lt 0$ $\Rightarrow$ local maximum.
- $f''(c) = 0$ $\Rightarrow$ inconclusive: $x^4$ (a minimum), $-x^4$ (a maximum) and $x^3$ (neither) all have $f'(0) = f''(0) = 0$. Look at higher derivatives: the first non-zero one decides (even order: min or max by its sign; odd order: not an extremum).
$n$-dimensional test. Suppose $\nabla f(\mathbf{c}) = \mathbf{0}$ and look at the eigenvalues of the symmetric $H(\mathbf{c})$:
- All $\gt 0$ (positive definite) $\Rightarrow$ local minimum.
- All $\lt 0$ (negative definite) $\Rightarrow$ local maximum.
- Some $\gt 0$ and some $\lt 0$ (indefinite) $\Rightarrow$ saddle point.
- Some $= 0$ and none of both signs $\Rightarrow$ inconclusive.
Shortcut for two variables. With $D = f_{xx}f_{yy} - f_{xy}^2 = \det H$ (the product of the eigenvalues): $D \gt 0$ and $f_{xx} \gt 0$ means minimum; $D \gt 0$ and $f_{xx} \lt 0$ means maximum; $D \lt 0$ means saddle; $D = 0$ is inconclusive.
Why do we need it?
"The gradient is zero" cannot tell a valley from a hilltop or a pass. The Hessian can. It is the standard way to certify that a point is really a minimum.
Where is it used?
Optimality conditions in optimisation theory, checking an analytic solution (for example that the least-squares solution is a minimum), analysing loss landscapes, and deciding whether an optimiser has stalled at a saddle.
How is it used?
1) Solve $\nabla f = 0$ for candidate points. 2) Compute $H$ there. 3) Get the eigenvalues (eigvalsh) and read off the signs. Treat tiny eigenvalues with care: numerically they may be zero.
"Inconclusive" does not mean "saddle". It means the second derivatives cannot decide. Example: $f = x^2y + y^2$ has $H(0,0) = \mathrm{diag}(0, 2)$, but along the curve $y = -x^2/2$ it equals $-x^4/4 \lt 0$ while along the $y$-axis it is positive, so the origin is actually a (degenerate) saddle.
It finds local, not global, minima. The two wells are both minima. Here they are equally deep, so both are global minima; in general only the lowest well is the global one.
Numerical zero. On a computer an eigenvalue like $10^{-12}$ is really zero. Use a tolerance.
Quick check: at a critical point $f_{xx}=3$, $f_{yy}=2$, $f_{xy}=2$. What is it?
$D = 3\cdot2 - 2^2 = 2 \gt 0$ and $f_{xx} = 3 \gt 0$: a local minimum. (Eigenvalues $(5\pm\sqrt{17})/2 \approx 4.56$ and $0.44$, both positive ✓.)
Newton's method: use the curvature to take a smarter step core
Gradient descent looks only at the slope: "downhill is that way, so take a small step". It has no idea how far the bottom is. Newton's method also looks at the curvature.
Here is the trick. Near where you stand, replace the curve by the parabola that matches its height, slope and curvature. A parabola has an obvious bottom, and you can jump straight to it. Then stand there, build a new parabola, and jump again.
If the ground really is a bowl-shaped parabola, you land on the bottom in one jump. Where the slope is gentle and the curvature is small, the parabola's bottom is far away, so Newton takes a big step; where the curvature is large, it takes a cautious one. The step size adjusts itself.
Minimise $f(x) = e^x - 2x$. Its minimum is where $f' = e^x - 2 = 0$, i.e. $x = \ln 2 \approx 0.6931$. Newton's update is $x \leftarrow x - f'(x)/f''(x)$ with $f'' = e^x$. Start at $x_0 = 0$.
- $x_0 = 0$: $f' = 1 - 2 = -1$, $f'' = 1$. Step $= -(-1)/1 = +1$, so $x_1 = 1$.
- $x_1 = 1$: $f' = e - 2 = 0.7183$, $f'' = e = 2.7183$. Step $= -0.2642$, so $x_2 = 0.7358$.
- $x_2 = 0.7358$: $f' = 0.0877$, $f'' = 2.0877$. Step $= -0.0420$, so $x_3 = 0.6940$.
- $x_3 = 0.6940$: next $x_4 = 0.69315$. The true answer is $0.693147\ldots$ Four steps give five correct digits.
Plain gradient descent with a safe step $\eta = 0.1$ needs about 32 steps to get within $0.001$ of the answer.
Derivation (1D). Near $x$, approximate $f$ by its second-order expansion (Chapter 2.11), a parabola in the step $\delta$:
$$m(\delta) = f(x) + f'(x)\,\delta + \tfrac12 f''(x)\,\delta^2.$$Its bottom is where its slope is zero: $m'(\delta) = f'(x) + f''(x)\,\delta = 0$, so $\delta = -\dfrac{f'(x)}{f''(x)}$. Therefore
$$x_{\text{new}} = x - \frac{f'(x)}{f''(x)}.$$In many dimensions the model is $m(\boldsymbol\delta) = f + \nabla f^\top\boldsymbol\delta + \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$. Setting its gradient (which is $\nabla f + H\boldsymbol\delta$ for symmetric $H$) to zero gives $H\boldsymbol\delta = -\nabla f$, so
$$\mathbf{x}_{\text{new}} = \mathbf{x} - H(\mathbf{x})^{-1}\,\nabla f(\mathbf{x}).$$(In code we solve $H\boldsymbol\delta = -\nabla f$ rather than invert $H$.) Gradient descent is the same recipe with the curvature replaced by a fixed guess: $H \to \frac{1}{\eta} I$ gives $\boldsymbol\delta = -\eta\nabla f$.
Another view: Newton's method is the tangent-line root-finder applied to $f'$: it finds where $f' = 0$.
Why do we need it?
Gradient descent needs many small steps, and a bad step size makes it crawl or blow up. Using curvature chooses the step length (and direction) for us, and near a minimum converges extremely fast.
Where is it used?
Logistic regression solvers (iteratively reweighted least squares is Newton's method), L-BFGS and other quasi-Newton optimisers that approximate the Hessian, natural-gradient and K-FAC ideas in deep learning, and interior-point solvers.
How is it used?
At each iterate compute the gradient and Hessian, solve $H\boldsymbol\delta = -\nabla f$, step to $\mathbf{x}+\boldsymbol\delta$ (often with a line search or damping), and repeat until the gradient is tiny. Cost: about $n^3$ operations and $n^2$ memory per step.
Newton heads for a flat spot of its parabola model, not necessarily a minimum. Where the curvature is negative it happily jumps to a maximum or a saddle (see the widgets). Practical versions fix the Hessian to be positive definite (add $\lambda I$), use a line search or trust region, or only take Newton steps near a minimum.
It is costly. Forming $H$ takes $n^2$ memory and solving takes $\sim n^3$ time. For a million weights that is impossible, which is why deep learning uses first-order methods (SGD, Adam) or cheap Hessian approximations (L-BFGS).
It can overshoot far from the answer. If the curvature is much smaller than it is near the bottom, the parabola's bottom is far away (see $\ln\cosh x$).
Quick check: $f(x) = 3x^2 - 12x + 5$. Take one Newton step from $x = 10$.
$f'(x) = 6x - 12 = 48$ and $f''(x) = 6$. Step $= -48/6 = -8$, so $x_{\text{new}} = 2$. That is the exact minimum ($f'(2) = 0$): on a quadratic Newton needs one step.
Why curvature controls the best learning rate core
Roll a marble down a steep, narrow gully by moving it in fixed hops. If each hop is too big, you jump clean across the gully to the other wall, which is even higher, then back, higher still: it blows up. A steep wall (high curvature) means you must hop small. In a flat, wide valley (low curvature) you can hop large.
So the biggest safe learning rate $\eta$ is set by the sharpest curvature. And the speed is set by the flattest direction: you must hop small to be safe in the steep direction, so you crawl along the flat one. That tension is what makes ill-conditioned problems hard.
One variable, $f(x) = \tfrac12\lambda x^2$ with curvature $\lambda = 5$. Gradient descent: $x \leftarrow x - \eta\,f'(x) = x - \eta\cdot5x = (1 - 5\eta)\,x$.
- $\eta = 0.1$: factor $1 - 0.5 = 0.5$. Start at $x=2$: $2 \to 1 \to 0.5 \to 0.25$ (smooth shrinking).
- $\eta = 0.2$: factor $1 - 1 = 0$. One step lands exactly on $0$. (And $\eta = 1/\lambda = 0.2$ is Newton's step.)
- $\eta = 0.3$: factor $-0.5$. $2 \to -1 \to 0.5 \to -0.25$ (zig-zag but shrinking).
- $\eta = 0.5$: factor $-1.5$. $2 \to -3 \to 4.5 \to -6.75$ (blows up).
The boundary is at factor $-1$, i.e. $\eta = 2/\lambda = 0.4$.
1D quadratic. For $f(x) = \tfrac12\lambda x^2$ ($\lambda \gt 0$), the update gives $x_{k+1} = (1 - \eta\lambda)\,x_k$, so $x_k = (1-\eta\lambda)^k x_0$. This shrinks to $0$ exactly when $|1 - \eta\lambda| \lt 1$, that is
$$0 \lt \eta \lt \frac{2}{\lambda}.$$The best step is $\eta = 1/\lambda$ (one-step convergence). For $\eta\lambda \gt 1$ the iterates alternate sides; for $\eta\lambda \gt 2$ they grow.
Many dimensions. For $f(\mathbf{x}) = \tfrac12\mathbf{x}^\top H\mathbf{x}$, write $H = V\Lambda V^\top$ and use the eigen-directions as new axes (eigendecomposition). The problem splits into $n$ independent 1D problems with curvatures $\lambda_1,\dots,\lambda_n$. Every one must be stable, so
$$\boxed{\;0 \lt \eta \lt \frac{2}{\lambda_{\max}}\;}$$The slowest direction shrinks by $|1 - \eta\lambda_{\min}|$ each step. Choosing $\eta$ so that the fastest and slowest directions have equal-size factors, $1 - \eta\lambda_{\min} = -(1 - \eta\lambda_{\max})$, gives $\eta^* = \dfrac{2}{\lambda_{\max} + \lambda_{\min}}$ and a best possible factor
$$\frac{\lambda_{\max} - \lambda_{\min}}{\lambda_{\max} + \lambda_{\min}} = \frac{\kappa - 1}{\kappa + 1}, \qquad \kappa = \frac{\lambda_{\max}}{\lambda_{\min}} \ (\text{the condition number}).$$With $\lambda = (1, 10)$: $\eta \lt 0.2$, $\eta^* = 2/11 \approx 0.182$, factor $9/11 \approx 0.82$ per step. A bigger $\kappa$ pushes the factor towards 1: slow.
For a general smooth function the Hessian changes from place to place, and the same rule applies locally: the safe step shrinks where the loss is sharp. (Standard theory uses $\eta \le 1/L$, where $L$ bounds $\lambda_{\max}$.) Experiments with full-batch gradient descent on neural networks (reported as the "edge of stability") show the largest Hessian eigenvalue often rising during training until it hovers near $2/\eta$, while the loss keeps falling in a bumpy way. Treat this as an observed phenomenon, not a rule that always holds.
Why do we need it?
"Which learning rate should I use?" is the most common question in training. Curvature gives the principled answer: too big blows up, too small crawls, and the limit is $2/\lambda_{\max}$.
Where is it used?
Choosing and scheduling the learning rate, learning-rate warm-up (curvature can be large at the start), feature scaling (it makes the bowl rounder, reducing $\kappa$) and normalisation layers (often credited with similar effects), preconditioners, Adam and RMSprop (which rescale each direction), and convergence proofs.
How is it used?
If you can estimate $\lambda_{\max}$ (for example by a few power-iteration steps with Hessian-vector products), set $\eta$ below $2/\lambda_{\max}$. Otherwise try a few learning rates on a log scale and watch whether the loss goes down smoothly, zig-zags, or blows up.
The limit comes from the largest curvature, the speed from the smallest. Making the sharpest direction stable forces small steps, which slows the flat direction. That is why rescaling features (making the bowl round, $\kappa \approx 1$) speeds up training dramatically.
Real losses are not quadratic. Curvature changes along the path, so a rate that is safe in a flat region may blow up when you enter a sharp one. Warm-up and gradient clipping help.
Quick check: $f = \frac12(2x^2 + 5y^2)$. What is the largest safe $\eta$, and the best $\eta$?
$\lambda_{\max} = 5$, so $\eta \lt 2/5 = 0.4$. The best step is $2/(5+2) = 2/7 \approx 0.286$, with factor $(5-2)/(5+2) = 3/7 \approx 0.43$ per step.
Recap, cheat sheet and practice
- The second derivative $f''$ is the slope of the slope. Positive: smile (concave up); negative: frown. Higher derivatives $f^{(n)}$ keep going; they are the raw material of Taylor series.
- The Hessian $H_{ij} = \partial^2 f/\partial x_i\partial x_j$ is the $n\times n$ matrix of second partials. It is the Jacobian of the gradient, and symmetric for smooth functions (equal mixed partials).
- The curvature along a unit direction $\mathbf{u}$ is $\mathbf{u}^\top H\mathbf{u}$, always between $\lambda_{\min}$ and $\lambda_{\max}$, reached along the eigenvectors.
- At a critical point: all eigenvalues positive is a bowl (minimum), all negative a hill (maximum), mixed signs a saddle, a zero eigenvalue is inconclusive. In high dimensions saddles are expected to far outnumber minima.
- Newton's method uses curvature: $\mathbf{x}\leftarrow\mathbf{x} - H^{-1}\nabla f$ (exact in one step on a quadratic, but costly and drawn to saddles).
- On a quadratic, gradient descent is stable only if $\eta \lt 2/\lambda_{\max}$; speed is set by the condition number $\kappa = \lambda_{\max}/\lambda_{\min}$.
Cheat sheet
| Idea | Formula | Picture |
|---|---|---|
| Second derivative | $f'' = \dfrac{d}{dx}f' \approx \dfrac{f(x+h)-2f(x)+f(x-h)}{h^2}$ | smile ($+$) or frown ($-$) |
| Hessian | $H_{ij} = \dfrac{\partial^2 f}{\partial x_i\partial x_j}$, $H = H^\top$ | table of curvatures |
| Curvature along $\mathbf{u}$ | $\mathbf{u}^\top H\mathbf{u}$, $\lambda_{\min}\le\cdot\le\lambda_{\max}$ | slice of the surface |
| Curvature circle | $\kappa = \dfrac{f''}{(1+f'^2)^{3/2}}$, $R=1/|\kappa|$ | circle hugging the curve |
| Bowl / hill / saddle | all $\lambda\gt0$ / all $\lambda\lt0$ / mixed signs | min / max / pass |
| 2D test | $D=f_{xx}f_{yy}-f_{xy}^2$; $D\gt0,f_{xx}\gt0$ min; $D\lt0$ saddle | sign of $\det H$ |
| Newton step | $\boldsymbol\delta = -H^{-1}\nabla f$ | jump to the parabola's bottom |
| Safe learning rate | $\eta \lt 2/\lambda_{\max}$; best $2/(\lambda_{\max}+\lambda_{\min})$ | hop smaller than the gully |
| Convergence factor | $(\kappa-1)/(\kappa+1)$ | round bowl is fast |
import numpy as np
# f(x, y) = x^2 * y + y^2 (the worked example of this chapter)
def f(p):
x, y = p
return x**2 * y + y**2
def hess(p): # exact Hessian (from the formulas we derived)
x, y = p
return np.array([[2*y, 2*x],
[2*x, 2.0]])
def hessian_fd(f, p, h=1e-4): # Hessian from function values only (central differences)
p = np.asarray(p, float); n = p.size; H = np.zeros((n, n))
for i in range(n):
for j in range(n):
ei = np.zeros(n); ej = np.zeros(n); ei[i] = h; ej[j] = h
H[i, j] = (f(p+ei+ej) - f(p+ei-ej) - f(p-ei+ej) + f(p-ei-ej)) / (4*h*h)
return H
p = np.array([1.0, 2.0])
print(hess(p)) # [[4. 2.] [2. 2.]]
print(np.round(hessian_fd(f, p), 3)) # same numbers, found by nudging
print(np.linalg.eigvalsh(hess(p))) # [0.7639 5.2361] = 3 -/+ sqrt(5): both positive
def classify(H, tol=1e-8): # second-derivative test for a symmetric Hessian
lam = np.linalg.eigvalsh(H)
if np.all(lam > tol): return "local minimum"
if np.all(lam < -tol): return "local maximum"
if lam.min() < -tol and lam.max() > tol: return "saddle point"
return "inconclusive"
print(classify(np.diag([8., 2.]))) # local minimum (two-wells function at (1, 0))
print(classify(np.diag([-4., 2.]))) # saddle point (two-wells function at (0, 0))
H = np.array([[1., 2.], [2., 1.]])
u = np.array([1., 1.]) / np.sqrt(2)
print(u @ H @ u) # 3.0 curvature along u (about 3, up to rounding)
print(np.linalg.eigvalsh(H)) # [-1. 3.] smallest and largest possible curvature
# Newton's method on g(x) = e^x - 2x (minimum at ln 2 = 0.6931...)
x = 0.0
for k in range(4):
x = x - (np.exp(x) - 2) / np.exp(x) # x - g'(x) / g''(x)
print(k + 1, x) # 1.0, 0.7358, 0.6940, 0.69315
# Learning rate vs curvature on f = 0.5 * (1*x^2 + 10*y^2): stable only if eta < 2/10 = 0.2
lam = np.array([1.0, 10.0])
for eta in (0.15, 0.19, 0.21):
z = np.array([1.0, 1.0])
for _ in range(100):
z = z - eta * lam * z # gradient descent step (gradient = lam * z)
print(eta, np.abs(z).max()) # tiny, tiny, huge
1. At a point where $f'(c) = 0$ and $f''(c) = -3$, what is $c$?
2. Why is the Hessian of a smooth function symmetric?
3. At a critical point the Hessian has eigenvalues $4$ and $-1$. This is a…
4. On $f(x) = \tfrac12\cdot5\,x^2$ gradient descent is run. Which learning rate makes it diverge?
5. Newton's method is applied to $f(x) = 3x^2 - 12x + 5$ starting at $x = 10$. After one step, where is it?
6. Which statement about $\mathbf{u}^\top H\mathbf{u}$ for a unit vector $\mathbf{u}$ is true?
Practice problems
A. Find and classify the critical points of $f(x,y) = x^3 + y^3 - 3xy$.
$f_x = 3x^2 - 3y = 0$ and $f_y = 3y^2 - 3x = 0$. From the first, $y = x^2$; substituting, $x^4 = x$, so $x = 0$ or $x = 1$. Critical points: $(0,0)$ and $(1,1)$. The Hessian is $H = \begin{bmatrix}6x & -3\\ -3 & 6y\end{bmatrix}$.
At $(0,0)$: $H = \begin{bmatrix}0&-3\\-3&0\end{bmatrix}$, eigenvalues $\pm3$: saddle. At $(1,1)$: $H = \begin{bmatrix}6&-3\\-3&6\end{bmatrix}$, eigenvalues $3$ and $9$: local minimum with $f(1,1) = -1$.
B. For $H = \begin{bmatrix}2&1\\1&4\end{bmatrix}$ find the curvature along $\mathbf{u} = [0.6, 0.8]$ and check it lies in the allowed range.
$\mathbf{u}^\top H\mathbf{u} = 2(0.36) + 2\cdot1\cdot(0.6)(0.8) + 4(0.64) = 0.72 + 0.96 + 2.56 = 4.24$. The eigenvalues are $3\pm\sqrt2 \approx 1.59$ and $4.41$, and $1.59 \le 4.24 \le 4.41$ ✓.
C. Take two Newton steps for $f(x) = x^4$ from $x = 2$. What is the pattern?
$f' = 4x^3$, $f'' = 12x^2$, so $x - f'/f'' = x - x/3 = \tfrac23 x$. From $2$: $x_1 = 4/3 \approx 1.333$, $x_2 = 8/9 \approx 0.889$. The error shrinks by a factor $2/3$ each step (slow, because $f''(0) = 0$: the parabola model is poor at a flat-bottomed minimum).
D. $f = \frac12(2x^2 + 5y^2)$. Give the safe range for $\eta$, the best $\eta$ and the best per-step factor.
$\lambda = 2, 5$, so $\kappa = 2.5$. Safe: $\eta \lt 2/5 = 0.4$. Best: $\eta^* = 2/(2+5) = 2/7 \approx 0.286$. Factor: $(\kappa-1)/(\kappa+1) = 1.5/3.5 = 3/7 \approx 0.43$.
E. Show the second-derivative test is inconclusive for $f(x,y) = x^4 + y^2$ at the origin, then decide what the origin is.
$\nabla f = [4x^3, 2y] = \mathbf{0}$ at the origin and $H = \begin{bmatrix}12x^2&0\\0&2\end{bmatrix} = \mathrm{diag}(0, 2)$ there: one eigenvalue is $0$, so the test is inconclusive. But $f = x^4 + y^2 \ge 0$ and equals $0$ only at the origin, so it is a strict (global) minimum.
F. If each of the $n=10$ curvature signs at a critical point were an independent fair coin flip, what is the chance of a minimum? Why does this matter?
All 10 must be positive: $2^{-10} = 1/1024 \approx 0.1\%$. So with many directions, almost every critical point has some downhill direction: most are saddles.
Taylor Series
Almost every function in machine learning is complicated. But close to a single point, any smooth function looks like a simple polynomial: a straight line, then a parabola, then something even closer. Taylor series is the recipe for building that polynomial. It is also the reason gradient descent and Newton's method work at all.
- Build a polynomial that matches a function's value, slope, curvature and more at a point, and see why the coefficients are $f^{(n)}(a)/n!$
- Use the first-order (tangent) and second-order (adds curvature) approximations, in one variable and in many
- Expand $e^x$, $\sin x$, $\cos x$, $\ln(1+x)$, $\frac{1}{1-x}$ and the sigmoid, and use them for quick estimates
- Know how big the error is (about $\delta^2$ for first order, $\delta^3$ for second) and where the approximation stops being good
- Derive gradient descent from the first-order expansion and Newton's method from the second-order one
The Taylor expansion: a polynomial that copies a function core
Polynomials ($1$, $x$, $x^2$, $x^3$, ...) are the easiest functions there are: a computer only needs to add and multiply. Functions like $e^x$ or $\sin x$ are much harder. The big idea: near one point, we can copy a hard function with a polynomial.
How do we copy it? One feature at a time, from the most basic up:
- Match the height at the point (a flat line at the right level).
- Also match the slope (the line now tilts like the curve).
- Also match the curvature (the line bends into a parabola).
- Also match how the curvature changes, and so on. Each extra match makes the copy hug the curve for longer.
Each "feature" is one derivative from Chapter 2.10. The finished polynomial is the Taylor polynomial.
Copy $f(x) = e^x$ near $a = 0$ using $p(x) = c_0 + c_1x + c_2x^2 + c_3x^3 + \cdots$. Every derivative of $e^x$ is $e^x$, and $e^0 = 1$, so at $x = 0$ the height, slope, curvature, ... are all $1$.
- Match the height: $p(0) = c_0$ must be $1$, so $c_0 = 1$.
- Match the slope: $p'(x) = c_1 + 2c_2x + 3c_3x^2 + \cdots$, so $p'(0) = c_1 = 1$.
- Match the curvature: $p''(x) = 2c_2 + 6c_3x + \cdots$, so $p''(0) = 2c_2 = 1$ and $c_2 = \tfrac12$.
- Next: $p'''(0) = 6c_3 = 1$, so $c_3 = \tfrac16$. Next: $p^{(4)}(0) = 24c_4 = 1$, so $c_4 = \tfrac1{24}$.
Numeric check at $x = 1$: $1 + 1 + 0.5 + 0.1667 + 0.0417 + 0.0083 = 2.7167$ after six terms, and $e = 2.7183$. Very close already.
The Taylor polynomial of degree $n$ of $f$ around the point $a$ is
$$P_n(x) = \sum_{k=0}^{n}\frac{f^{(k)}(a)}{k!}(x-a)^k = f(a) + f'(a)(x-a) + \frac{f''(a)}{2!}(x-a)^2 + \cdots + \frac{f^{(n)}(a)}{n!}(x-a)^n.$$Where does $1/k!$ come from? Suppose $p(x) = c_0 + c_1(x-a) + c_2(x-a)^2 + \cdots$ and we want $p^{(k)}(a) = f^{(k)}(a)$ for every $k$. Differentiate the term $c_k(x-a)^k$ exactly $k$ times: the power rule gives $k\cdot(k-1)\cdots2\cdot1 = k!$, a constant. Lower powers have already vanished (their $k$-th derivative is $0$), and higher powers still contain a factor $(x-a)$ that is $0$ at $x=a$. So $p^{(k)}(a) = k!\,c_k$. Setting this equal to $f^{(k)}(a)$ gives
$$c_k = \frac{f^{(k)}(a)}{k!}.$$Taking $n\to\infty$ gives the Taylor series. Around $a=0$ it is also called the Maclaurin series. Degree $0$ is a flat line at height $f(a)$; degree $1$ is the tangent line; degree $2$ is a parabola.
Why do we need it?
Hard functions are hard to compute and hard to reason about. A polynomial copy is easy to evaluate, differentiate and minimise, and it tells us how the function behaves nearby.
Where is it used?
How calculators and math libraries compute $\sin$, $\exp$ and $\log$; justification of gradient descent and Newton's method; Laplace approximations; second-order loss approximations in gradient boosting (XGBoost); and error and sensitivity estimates.
How is it used?
Pick the point $a$ you care about, compute $f$ and its first few derivatives there, divide the $k$-th derivative by $k!$, and add up the terms. Keep as many terms as the accuracy you need.
The order of the polynomial is not the number of non-zero terms. $\sin x$ has only odd powers, so its degree-4 polynomial is the same as its degree-3 one.
A Taylor polynomial is built at one point. Change $a$ and every coefficient changes. It copies the function near $a$, not everywhere (see "Local approximation" below).
Quick check: what is the coefficient of $x^3$ in the Taylor series of $\sin x$ at $0$?
$f'''(x) = -\cos x$, so $f'''(0) = -1$. Divide by $3! = 6$: the coefficient is $-\tfrac16$. So $\sin x \approx x - \tfrac{x^3}{6} + \cdots$.
First-order approximation: the tangent line core
Zoom in on any smooth curve. Zoom in more. Eventually the curve looks like a straight line, and that line is the tangent. So close to a point, the function behaves like its tangent line: "start at the current height, and add slope times how far you moved".
This is the simplest useful approximation, and it is what a gradient is for: the gradient is the slope, and the slope predicts what a small nudge does.
Estimate $e^{0.1}$ without a calculator, using $f(x) = e^x$ near $a = 0$. Here $f(0) = 1$ and $f'(0) = 1$.
- First-order formula: $f(a + h) \approx f(a) + f'(a)\,h$.
- With $a = 0$ and $h = 0.1$: $e^{0.1} \approx 1 + 1\cdot0.1 = 1.1$.
- True value: $e^{0.1} = 1.10517\ldots$ The error is $0.0052$.
Now double the step, $h = 0.2$: estimate $1.2$, true $1.2214$, error $0.0214$. Twice the step gave about four times the error ($0.0214/0.0052 \approx 4.1$). The error grows like $h^2$.
The first-order Taylor approximation (also called the linear approximation or tangent-line approximation) is
$$f(a + h) \approx f(a) + f'(a)\,h, \qquad\text{or}\qquad f(x) \approx f(a) + f'(a)(x - a).$$Its graph is the tangent line at $a$. The error is the part we dropped, which is of size about $\tfrac12 f''(a)\,h^2$: proportional to the square of the step. In words: halve the step, quarter the error.
For $f:\mathbb{R}^n\to\mathbb{R}$ this becomes $f(\mathbf{x}+\boldsymbol\delta) \approx f(\mathbf{x}) + \nabla f(\mathbf{x})^\top\boldsymbol\delta$ (see the multivariable section below). Chapter 2.12 continues this idea as linearisation.
Why do we need it?
It turns a curved, hard question ("what will the loss be after this change?") into a simple one ("height plus slope times step"). That is enough to decide which way to move.
Where is it used?
Gradient descent (the step is chosen using it), sensitivity and error propagation, finite-difference gradient checks, linearising a network around a point (neural tangent kernel ideas), and the Gauss-Newton method.
How is it used?
Compute $f(a)$ and the slope $f'(a)$ (or the gradient). Predict the new value as $f(a)$ plus slope times step. Trust it only for small steps.
It is only good for small steps. The error is $\approx\tfrac12 f''h^2$. If the curvature $f''$ is large, "small" has to be very small. A straight line cannot follow a sharp bend.
The sign of the error tells the curvature. If the curve is a smile ($f''\gt0$) the tangent line lies below the curve; for a frown it lies above.
Quick check: estimate $\sqrt{4.1}$ using the tangent line of $\sqrt x$ at $a = 4$.
$f(4) = 2$ and $f'(x) = \dfrac{1}{2\sqrt x}$, so $f'(4) = \dfrac14$. Then $\sqrt{4.1} \approx 2 + \dfrac14(0.1) = 2.025$. The true value is $2.02485\ldots$ (error about $0.00015$).
Second-order approximation: adding the curvature core
The tangent line is straight, but the curve is not. So the line peels away from the curve, and it peels off to one side: above a frown, below a smile. We can fix this by letting the copy bend. Add a term that matches the second derivative, and the straight line becomes a parabola hugging the curve much longer.
That one extra term, $\tfrac12 f''(a)\,h^2$, is the curvature term from Chapter 2.10.
Estimate $e^{0.1}$ again, now adding the curvature term. At $a=0$: $f = f' = f'' = 1$.
- Formula: $f(a+h) \approx f(a) + f'(a)h + \tfrac12 f''(a)h^2$.
- $e^{0.1} \approx 1 + 0.1 + \tfrac12(0.01) = 1.105$.
- True value $1.105171$. Error $0.000171$, about 30 times smaller than the tangent line's $0.0052$.
Double the step to $h = 0.2$: estimate $1.22$, true $1.221403$, error $0.001403$. Twice the step gives about eight times the error ($0.001403/0.000171 \approx 8.2$): the error now grows like $h^3$.
The second-order Taylor approximation is
$$f(a + h) \approx f(a) + f'(a)\,h + \tfrac12 f''(a)\,h^2.$$Its graph is a parabola that matches $f$ in height, slope and curvature at $a$. The leftover error is of size about $\tfrac16 f'''(a)h^3$: proportional to the cube of the step. Halve the step, divide the error by about eight.
If $f$ is itself a parabola (a quadratic), this approximation is exact, because all higher derivatives are zero.
Why do we need it?
A straight line cannot tell whether a flat spot is a valley or a hilltop. A parabola can: it has a bottom (or a top). Adding curvature also lets us take bigger, smarter steps.
Where is it used?
Newton's method and its relatives (L-BFGS), the Laplace approximation (a Gaussian fitted at a loss minimum), second-order terms in XGBoost, and analysing how big a learning rate can be.
How is it used?
Compute value, slope and curvature at the point and use the parabola as a stand-in for the function. To minimise, jump to the parabola's bottom; to analyse stability, read the curvature.
More terms are not always better far away. The extra terms improve the copy near $a$. Far from $a$ a high-degree polynomial can shoot off wildly. The next sections show where it works.
Quick check: use the second-order formula to estimate $\ln(1.1)$ at $a=0$ for $f=\ln(1+x)$.
$f(0) = 0$, $f'(0) = 1$, $f''(0) = -1$. So $\ln(1+h) \approx h - \tfrac12h^2$. With $h = 0.1$: $0.1 - 0.005 = 0.095$. The true value is $0.09531$.
The famous expansions, derived core
A handful of functions show up everywhere, so it pays to know their Taylor series around $0$ and, even better, to know how to rebuild them. Each one comes from the same two steps: (1) find the derivatives at $0$, (2) divide the $k$-th one by $k!$.
The patterns are simple: $e^x$ has all-equal derivatives; $\sin$ and $\cos$ cycle through $0, 1, 0, -1$; $\frac1{1-x}$ and $\ln(1+x)$ are geometric-looking.
$e^x$: every derivative at $0$ is $1$, so $c_k = \dfrac1{k!}$.
$\sin x$: $f, f', f'', f''', f^{(4)}, \dots$ at $0$ are $0, 1, 0, -1, 0, 1, \dots$ (from $\sin, \cos, -\sin, -\cos, \dots$). Only odd $k$ survive, with alternating signs, so $c_1 = 1$, $c_3 = -\tfrac1{3!}$, $c_5 = \tfrac1{5!}$.
$\cos x$: the derivatives at $0$ are $1, 0, -1, 0, 1, \dots$. Only even $k$: $c_0 = 1$, $c_2 = -\tfrac1{2!}$, $c_4 = \tfrac1{4!}$.
$\dfrac1{1-x}$: $f^{(k)}(x) = \dfrac{k!}{(1-x)^{k+1}}$ (the chain rule brings out one more factor each time). At $0$ it is $k!$, so $c_k = k!/k! = 1$.
$\ln(1+x)$: $f' = \dfrac1{1+x}$, $f'' = -\dfrac1{(1+x)^2}$, $f''' = \dfrac2{(1+x)^3}$, so $f^{(k)}(0) = (-1)^{k-1}(k-1)!$ for $k\ge1$ and $f(0)=0$. Then $c_k = \dfrac{(-1)^{k-1}(k-1)!}{k!} = \dfrac{(-1)^{k-1}}{k}$.
Sigmoid $\sigma(x) = \dfrac1{1+e^{-x}}$: use $\sigma' = \sigma(1-\sigma)$.
- $\sigma(0) = \tfrac12$, so $\sigma'(0) = \tfrac12\cdot\tfrac12 = \tfrac14$.
- Product rule on $\sigma' = \sigma - \sigma^2$: $\sigma'' = \sigma'(1 - 2\sigma)$. At $0$: $\tfrac14\cdot(1 - 1) = 0$.
- Again: $\sigma''' = \sigma''(1-2\sigma) - 2(\sigma')^2$. At $0$: $0 - 2\cdot\tfrac1{16} = -\tfrac18$.
- Coefficients: $c_0 = \tfrac12$, $c_1 = \tfrac14$, $c_2 = 0$, $c_3 = -\tfrac18/6 = -\tfrac1{48}$.
The "valid for" notes are the radius of convergence (see the section on local approximation). Two more used all the time: $\sqrt{1+x} \approx 1 + \tfrac x2 - \tfrac{x^2}{8}$ and $\dfrac{1}{1+x}\approx 1 - x + x^2$. And a link between them (awareness, needs complex numbers): $\cos x$ and $i\sin x$ are the even and odd parts of $e^{ix}$.
Why do we need it?
These functions are the building blocks of activations, losses and probabilities. Their series give quick mental estimates and show how each behaves near zero, where inputs to a network often live.
Where is it used?
Sigmoid $\approx \frac12 + \frac x4$ for small inputs (so a unit fed small inputs behaves almost linearly), $\ln(1+x)\approx x$ (log1p), $e^x - 1\approx x$ (expm1), small-angle $\sin\theta\approx\theta$, softplus $\approx \ln2 + \frac x2 + \frac{x^2}{8}$, and the geometric series in discounted returns in reinforcement learning.
How is it used?
Recognise the pattern, keep the first one to three terms, and check the error is acceptable for your range of $x$. When you need a new series, do not memorise: derive the derivatives at $0$ and divide by $k!$.
Know the "valid for" range. $\ln(1+x)$'s series only works for $-1 \lt x \le 1$ and $\frac{1}{1-x}$'s only for $|x|\lt1$, no matter how many terms you take. At $x=1.5$, adding terms makes the answer worse.
Sigmoid: $\sigma(0)=\tfrac12$, so the series starts at $\tfrac12$, not $0$. It is an odd function plus $\tfrac12$, so only odd powers appear after the constant.
Quick check: use the series to estimate $\sigma(0.4)$ with three terms, and compare with $0.59869$.
$\sigma(0.4) \approx \tfrac12 + \tfrac{0.4}{4} - \tfrac{0.4^3}{48} = 0.5 + 0.1 - \tfrac{0.064}{48} = 0.6 - 0.001333 = 0.598667$. The true value is $0.598688$, so the error is only about $0.00002$.
Multivariate Taylor expansion core
For a function of many inputs the picture is a surface. Zoom in on one spot and the surface looks like a flat tilted sheet (the tangent plane): that is the first-order approximation, built from the gradient. Look a little less closely and you see the sheet bend: it becomes a bowl, hill or saddle (a paraboloid): that is the second-order approximation, built from the Hessian.
The recipe is the same as in one variable: start from the current height, add the slope part, add the curvature part. Only now "slope" is a gradient (a vector) and "curvature" is a Hessian (a matrix), and the step $\boldsymbol\delta$ is a vector.
$f(x,y) = x^2y + y^2$ at $\mathbf{x} = (1,2)$, with a step $\boldsymbol\delta = (0.1,\,-0.1)$. From Chapter 2.10: $f_x = 2xy$, $f_y = x^2 + 2y$, $H = \begin{bmatrix}2y&2x\\2x&2\end{bmatrix}$.
- Height: $f(1,2) = 1\cdot2 + 4 = 6$.
- Gradient (a column): $\nabla f(1,2) = [2\cdot1\cdot2,\; 1 + 4] = [4, 5]$.
- Slope part: $\nabla f^\top\boldsymbol\delta = 4(0.1) + 5(-0.1) = -0.1$. So the first-order estimate is $6 - 0.1 = 5.9$.
- Hessian: $H(1,2) = \begin{bmatrix}4&2\\2&2\end{bmatrix}$. Then $\boldsymbol\delta^\top H\boldsymbol\delta = 4(0.01) + 2\cdot2\cdot(0.1)(-0.1) + 2(0.01) = 0.04 - 0.04 + 0.02 = 0.02$.
- Curvature part: $\tfrac12(0.02) = 0.01$. The second-order estimate is $5.9 + 0.01 = 5.91$.
- True value: $f(1.1,\,1.9) = 1.21\cdot1.9 + 3.61 = 5.909$. Errors: first order $0.009$, second order $0.001$.
For $f:\mathbb{R}^n\to\mathbb{R}$ (smooth), a point $\mathbf{x}$ and a small step $\boldsymbol\delta$:
$$\boxed{\,f(\mathbf{x}+\boldsymbol\delta) \;\approx\; f(\mathbf{x}) \;+\; \nabla f(\mathbf{x})^\top\boldsymbol\delta \;+\; \tfrac12\,\boldsymbol\delta^\top H(\mathbf{x})\,\boldsymbol\delta\,}$$Each term is a single number: $\nabla f^\top\boldsymbol\delta$ is a row $(1\times n)$ times a column $(n\times1)$; $\boldsymbol\delta^\top H\boldsymbol\delta$ is $(1\times n)(n\times n)(n\times 1)$. Keeping only the first two terms gives the first-order (tangent plane) approximation.
Derivation from the 1D formula. Walk along the straight line from $\mathbf{x}$ to $\mathbf{x}+\boldsymbol\delta$ and let $g(t) = f(\mathbf{x} + t\boldsymbol\delta)$, so $g(0) = f(\mathbf{x})$ and $g(1) = f(\mathbf{x}+\boldsymbol\delta)$. Chapter 2.10 showed that $g'(0) = \nabla f(\mathbf{x})^\top\boldsymbol\delta$ and $g''(0) = \boldsymbol\delta^\top H(\mathbf{x})\boldsymbol\delta$ (curvature along a direction, without normalising). The one-variable second-order formula at $t = 1$ is $g(1)\approx g(0) + g'(0)\cdot 1 + \tfrac12 g''(0)\cdot1^2$, which is exactly the boxed formula ∎.
If $f$ is a quadratic $\tfrac12\mathbf{x}^\top A\mathbf{x} - \mathbf{b}^\top\mathbf{x} + c$ ($A$ symmetric) the formula is exact, with $H = A$. Higher orders involve third derivatives (a 3-index array); in ML we almost never go past second order.
Why do we need it?
A loss depends on millions of parameters at once. To predict what a small change of all of them does to the loss, we need a formula that works with vectors: gradient for the slope, Hessian for the bend.
Where is it used?
The derivation of gradient descent and Newton's method (below), trust-region optimisers, the Laplace approximation, natural gradients and K-FAC, Hessian-based pruning (Optimal Brain Surgeon), and sharp-versus-flat-minima analyses.
How is it used?
At the current parameters compute the loss $f$, the gradient $g$ and (if affordable) the Hessian $H$. The model $f + g^\top\delta + \tfrac12\delta^\top H\delta$ predicts the loss after any step $\delta$, and can be minimised over $\delta$.
Order matters in the quadratic term: $\boldsymbol\delta^\top H\boldsymbol\delta$ puts the step on both sides of the matrix; it is not $H\boldsymbol\delta$ (a vector) and not $H\boldsymbol\delta^2$.
The Hessian term is a curvature times the step squared. It is negligible for tiny steps, which is why first order suffices for small learning rates, but it dominates for large ones.
Quick check: for $f(x,y) = x^2 + y^2$ at $(1,1)$, what do first- and second-order approximations predict for $\boldsymbol\delta = (0.5, 0)$, and what is the truth?
$f = 2$, $\nabla f = [2, 2]$, $H = 2I$. First order: $2 + 2(0.5) = 3$. Second order: $+\tfrac12\cdot2\cdot0.25 = 0.25$, giving $3.25$. Truth: $f(1.5, 1) = 2.25 + 1 = 3.25$. Exactly equal, because $f$ is a quadratic.
Taylor approximation in practice
In real work you rarely sum a long series. You keep one, two or three terms and ask: is this good enough for the numbers I care about? Typical uses are quick estimates by hand ("$\sqrt{101}$ is about 10.05"), safer numerics for tiny inputs, estimating how an error in the input spreads to the output, and building simplified models of a loss.
The skill is to match the order to the size of the step: for a very small change one term is enough; for a bigger one add the curvature.
- Estimate by hand. $\sqrt{101} = 10\sqrt{1.01} \approx 10\left(1 + \tfrac{0.01}{2}\right) = 10.05$. True: $10.04988$.
- Error propagation. A circle's radius is measured as $r = 10 \pm 0.1$. Area $A = \pi r^2$, so $dA \approx A'(r)\,dr = 2\pi r\,dr = 2\pi\cdot10\cdot0.1 = 6.28$. (Exact change: $\pi(10.1^2 - 10^2) = 6.31$.)
- Sigmoid at small input. $\sigma(0.5) \approx \tfrac12 + \tfrac{0.5}{4} = 0.625$. True: $0.6225$.
- Why
log1pexists. In floating point $1 + 10^{-17}$ rounds to exactly $1$, so $\ln(1+10^{-17})$ comes out $0$. But $\ln(1+x)\approx x$ gives the right answer $10^{-17}$. Libraries use such expansions internally (np.log1p,np.expm1). - Softplus. $\ln(1+e^x) \approx \ln 2 + \tfrac x2 + \tfrac{x^2}{8}$ near $0$ (derivatives: $\sigma(0) = \tfrac12$ and $\sigma'(0) = \tfrac14$).
A practical recipe.
- Choose the base point $a$ where $f$ and its derivatives are easy (often $a=0$, or a nearby "nice" number).
- Write $f(a+h) \approx f(a) + f'(a)h + \tfrac12 f''(a)h^2$ (add more terms if needed).
- Estimate the size of the next term, $\tfrac16|f'''|\,|h|^3$, to know the error.
- Stop when that error is smaller than what you need.
Propagation of error. If an input is off by $\Delta x$, the output is off by about $|f'(x)|\,\Delta x$. For several inputs, $\Delta f \approx \nabla f^\top\Delta\mathbf{x}$.
Second-order loss models. Gradient-boosted trees (XGBoost) approximate each loss as $\ell(\hat y+\delta) \approx \ell + g\,\delta + \tfrac12 h\,\delta^2$ with first derivative $g$ and second derivative $h$ of the loss (here $h$ is not the step), and then set $\delta = -g/h$: Newton's step.
Why do we need it?
Exact calculations are often too expensive, too unstable, or impossible by hand. A short Taylor expansion gives a fast answer together with an estimate of how wrong it can be.
Where is it used?
Numerical libraries (log1p, expm1, stable softplus), measurement error and uncertainty estimates in science, delta-method statistics, XGBoost and LightGBM, and linearised models of neural networks.
How is it used?
Keep one term for very small changes, two terms for moderate ones. Compare the size of the last kept term with the first dropped one. For tiny inputs use log1p/expm1 instead of log(1+x) or exp(x)-1.
Always say how big the step is compared with the curvature. "$\sin x\approx x$" is excellent for $x = 0.1$ (error $0.00017$) and poor for $x = 1.5$ (it says $1.5$ but the truth is $0.997$).
Do not expand around a bad point. The expansion point must be close to where you evaluate, and the function must be smooth there. Around $x=0$ the series of $\sqrt x$ does not exist (infinite slope), which is why we wrote $\sqrt{101}$ as $10\sqrt{1.01}$.
Quick check: use $e^x \approx 1 + x + \tfrac{x^2}{2}$ to estimate $e^{-0.2}$.
$1 - 0.2 + 0.02 = 0.82$. True $e^{-0.2} = 0.81873$. The error is about $0.0013$.
The remainder: how big is the error? core
A Taylor polynomial of degree $n$ copies the first $n$ features of the function. The error is whatever feature number $n+1$ (and later) contributes, and for a small step that is dominated by the first term we dropped. That term has $h^{n+1}$ in it, and powers of a small number shrink fast:
- Step $h = 0.1$: $h^2 = 0.01$, $h^3 = 0.001$, $h^4 = 0.0001$.
- Halve the step, and an error of size $h^2$ falls by $4$, of size $h^3$ by $8$, of size $h^4$ by $16$.
On a graph that plots the logarithm of the error against the logarithm of the step, these are straight lines with slopes $2$, $3$, $4$: the slope tells the order.
$f(x) = e^x$ at $a=0$, step $h$. First order: $P_1 = 1+h$. Second order: $P_2 = 1 + h + h^2/2$.
| $h$ | error of $P_1$ | $\text{error}/h^2$ | error of $P_2$ | $\text{error}/h^3$ |
|---|---|---|---|---|
| $0.2$ | $0.021403$ | $0.535$ | $0.001403$ | $0.175$ |
| $0.1$ | $0.005171$ | $0.517$ | $0.000171$ | $0.171$ |
| $0.05$ | $0.001271$ | $0.508$ | $0.0000211$ | $0.169$ |
The ratios settle at $\tfrac12 f''(0) = 0.5$ and $\tfrac16 f'''(0) = 0.1667$: the next Taylor coefficients.
Remainder (Lagrange form). If $f$ has $n+1$ derivatives, then
$$f(a+h) = P_n(a+h) + R_n, \qquad R_n = \frac{f^{(n+1)}(\xi)}{(n+1)!}\,h^{n+1}$$for some unknown number $\xi$ between $a$ and $a+h$. We do not know $\xi$, but if we can bound $|f^{(n+1)}|$ by a number $M$ on that interval, then
$$|R_n| \le \frac{M}{(n+1)!}\,|h|^{n+1}.$$- $n=1$ (tangent line): error $\approx \tfrac12 f''\,h^2$, so it is $O(h^2)$.
- $n=2$ (parabola): error $\approx \tfrac16 f'''\,h^3$, so it is $O(h^3)$.
- In general the error after degree $n$ is $O(h^{n+1})$.
Example of a bound. For $\sin 1$ with $P_5(1) = 1 - \tfrac16 + \tfrac1{120} = 0.841667$ (which is also the degree-6 polynomial, since the $x^6$ term is zero), $M = 1$ and $n = 6$: $|R| \le \tfrac1{7!} = 0.000198$. The true error is $0.000196$ ✓. (The Lagrange form is awareness: you rarely compute $\xi$, you just use the bound.)
Check of the Lagrange form for $e^{0.1}$ with $n=1$: the true error is $0.005171 = \tfrac12 e^{\xi}(0.1)^2$, which gives $e^{\xi} = 1.034$, so $\xi = 0.034$, inside $(0, 0.1)$ ✓.
Why do we need it?
An approximation without an error estimate is just a guess. The remainder tells us how far to trust it, and how much better a higher order or a smaller step would be.
Where is it used?
Proofs of convergence for gradient descent and Newton's method (they use exactly these remainder bounds), choosing finite-difference step sizes, designing numerical libraries, and deciding how small a learning rate must be for first-order reasoning to hold.
How is it used?
Look at the first dropped term to estimate the error, or bound the next derivative. In experiments, plot error against step on log-log axes: the slope reveals the order, and a wrong slope reveals a bug.
"$O(h^2)$" hides a constant. It says how the error scales, not how big it is. A large $f''$ makes the constant large, so "small" must be smaller.
Rounding noise. In floating-point arithmetic the error cannot go below about $10^{-16}$ relative to the numbers involved. This is why very tiny finite-difference steps stop helping.
Quick check: a first-order estimate has error $0.04$ for $h = 0.2$. Roughly what error do you expect for $h = 0.1$?
First order means error $\propto h^2$. Half the step gives a quarter of the error: about $0.01$.
Local approximation: only good near the point
A Taylor polynomial is a local copy: it is perfect at the base point and good close to it, but far away it can be badly wrong. A map of your street is a fine approximation of the city near your house and useless on another continent.
For some functions ($e^x$, $\sin x$) adding terms keeps widening the good region without limit. For others the good region has a fixed width however many terms you add: the series simply stops working beyond a certain distance from the base point, the radius of convergence. Beyond it, adding more terms makes things worse.
The geometric series $\dfrac1{1-x} = 1 + x + x^2 + \cdots$. Add up the terms at two points:
- $x = 0.5$: the partial sums are $1,\ 1.5,\ 1.75,\ 1.875,\ 1.9375,\ \dots \to 2$ and indeed $\frac1{1-0.5} = 2$ ✓. Converges.
- $x = 1.5$: the partial sums are $1,\ 2.5,\ 4.75,\ 8.125,\ 13.19,\ \dots$ growing without bound, while the function value is $\frac1{1-1.5} = -2$. Diverges.
The danger point is $x=1$ where the function itself blows up. The series works only for $|x| \lt 1$: radius of convergence $R = 1$.
A Taylor series around $a$ has a radius of convergence $R$: it converges to $f$ for $|x - a| \lt R$ and diverges for $|x-a| \gt R$.
- $e^x$, $\sin x$, $\cos x$: $R = \infty$ (work everywhere).
- $\dfrac{1}{1-x}$ and $\ln(1+x)$ around $0$: $R = 1$ (they break at $x = 1$ and $x=-1$ respectively).
- Sigmoid around $0$: $R = \pi$. This is surprising: $\sigma(4)$ is perfectly well-behaved, yet the series diverges at $x = 4$. The reason is hidden in the complex numbers: $\sigma$ blows up at $x = \pm i\pi$, at distance $\pi$ from $0$ (awareness).
Even inside $R$ you only get an accurate answer if you take enough terms, and the closer to the edge the more you need. Rule of thumb: trust a short Taylor polynomial only for steps much smaller than the distance to the nearest trouble spot.
Why do we need it?
It stops us from trusting an approximation outside the region where it was built. In optimisation this is the idea of a trust region: the model is only reliable for small steps.
Where is it used?
Trust-region methods and the reason learning rates must be small, proximal policy optimisation in reinforcement learning (it limits how far the policy moves because its approximation is only local), and understanding why a quadratic model of a loss fails far from the current weights.
How is it used?
Keep steps small relative to how quickly the curvature changes. If a step makes the actual loss disagree with the model's prediction, shrink the step (line search, smaller learning rate, trust region).
A Taylor polynomial is not a good global model. Even a high-degree polynomial matches the function only near the base point and shoots off elsewhere. Never extrapolate far.
Convergence is not the same as accuracy at a given N. Inside the radius the series converges, but at $x$ near the edge a degree-3 polynomial may still be off by a lot.
Quick check: the series of $\ln(1+x)$ has radius $1$. Is it safe to use at $x = 2$?
No. $|2| \gt 1$, so the series diverges there: the partial sums $2,\ 0,\ 2.67,\ -1.33, \dots$ swing wider and wider, while the true value $\ln3 = 1.0986$ is perfectly fine. Use a base point closer to $x=2$ instead: around $a=1$ the radius is $2$, so $x=2$ is safely inside.
Why gradient descent works: derived from the first-order expansion core
You stand on a foggy hillside. You cannot see the valley. All you can feel is the slope under your feet, so you treat the ground near you as a flat, tilted sheet (the first-order Taylor model). On a tilted sheet, which small step lowers your height the most? Straight against the slope, the way water would run.
That is gradient descent. It is not a rule somebody made up: it is what the first-order Taylor approximation says is the best small move.
$f(x,y) = \tfrac12(x^2 + 4y^2)$ at $P = (2, 1)$, so $f(P) = \tfrac12(4+4) = 4$. The gradient is $\mathbf{g} = [x, 4y] = [2, 4]$, with $\|\mathbf{g}\|^2 = 4 + 16 = 20$. Take the step $\boldsymbol\delta = -\eta\,\mathbf{g}$ with $\eta = 0.1$, i.e. $\boldsymbol\delta = [-0.2, -0.4]$.
- First-order prediction of the change: $\mathbf{g}^\top\boldsymbol\delta = 2(-0.2) + 4(-0.4) = -2.0$, which equals $-\eta\|\mathbf{g}\|^2 = -0.1\cdot20$ ✓.
- Actual change: the new point is $(1.8, 0.6)$ and $f = \tfrac12(3.24 + 4\cdot0.36) = \tfrac12(4.68) = 2.34$. So the change is $2.34 - 4 = -1.66$.
- The difference, $+0.34$, is the curvature term: $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta = \tfrac12\left(1\cdot0.04 + 4\cdot0.16\right) = \tfrac12(0.68) = 0.34$ ✓. The loss did go down, but by a bit less than the flat-sheet model promised.
Derivation.
- Model. For a small step $\boldsymbol\delta$: $f(\mathbf{x}+\boldsymbol\delta) \approx f(\mathbf{x}) + \mathbf{g}^\top\boldsymbol\delta$, where $\mathbf{g} = \nabla f(\mathbf{x})$ (a column).
- Question. Among all steps of a fixed small length $\|\boldsymbol\delta\| = \varepsilon$, which makes $\mathbf{g}^\top\boldsymbol\delta$ as negative (as downhill) as possible?
- Answer by Cauchy–Schwarz (Linear Algebra): $\mathbf{g}^\top\boldsymbol\delta \ge -\|\mathbf{g}\|\,\|\boldsymbol\delta\|$, with equality exactly when $\boldsymbol\delta$ points opposite to $\mathbf{g}$. So the steepest descent direction is $\boldsymbol\delta = -\varepsilon\,\mathbf{g}/\|\mathbf{g}\|$.
- Take a step proportional to $-\mathbf{g}$: $\boldsymbol\delta = -\eta\,\mathbf{g}$ with learning rate $\eta \gt 0$. The model predicts a change of $\mathbf{g}^\top(-\eta\mathbf{g}) = -\eta\|\mathbf{g}\|^2 \le 0$: the loss goes down (unless the gradient is zero).
- Update rule: $\mathbf{x}_{\text{new}} = \mathbf{x} - \eta\,\nabla f(\mathbf{x})$.
Why $\eta$ must be small. The model dropped the curvature term. Adding it back (second-order, exact for quadratics) the true change is
$$\Delta f \approx -\eta\|\mathbf{g}\|^2 + \tfrac12\eta^2\,\mathbf{g}^\top H\mathbf{g}.$$The first term is linear in $\eta$, the second is quadratic. For small $\eta$ the first wins and $f$ falls; for large $\eta$ the second wins and $f$ can go up. The break-even is $\eta = \dfrac{2\|\mathbf{g}\|^2}{\mathbf{g}^\top H\mathbf{g}}$, and since $\mathbf{g}^\top H\mathbf{g} \le \lambda_{\max}\|\mathbf{g}\|^2$, any $\eta \lt 2/\lambda_{\max}$ is safe: exactly the learning-rate limit of Chapter 2.10.
Why do we need it?
It tells us why "step against the gradient" is the right move, and exactly how it can fail (when the step is too big for the curvature). Knowing the reason lets you fix training problems instead of guessing.
Where is it used?
Every training loop: SGD, momentum, Adam all build on this step. It also underlies convergence proofs, learning-rate schedules and the step-size rules in line search.
How is it used?
Compute the gradient, subtract $\eta$ times it from the parameters. If the loss goes up, the first-order model failed: the step was too big compared with the curvature, so reduce $\eta$.
The gradient is the steepest direction only for tiny steps, and only in the usual (Euclidean) sense of distance. With big steps the curvature bends the path; with a different distance (or a different metric) the best direction changes. That is the idea behind natural gradient and, loosely, behind preconditioned methods such as Adam (for the ill-conditioned bowls they try to fix, see Chapter 2.10).
"Steepest" is not "fastest". In a long narrow valley the steepest direction points across the valley, not along it, so gradient descent zig-zags. Curvature (the next section) can fix that.
Quick check: $f(x) = x^2$ at $x = 3$. What does the first-order model predict for one gradient step with $\eta = 0.1$, and what happens?
$f' = 6$, so the step is $-0.6$ and the predicted change is $-\eta f'^2 = -0.1\cdot36 = -3.6$. Actual: $x = 2.4$, $f = 5.76$, change $5.76 - 9 = -3.24$. The loss fell, a little less than predicted (the missing curvature term is $+0.36$).
Why Newton's method works: derived from the second-order expansion core
Same foggy hillside, but now you can feel both the slope and the bend of the ground. Your model is no longer a flat sheet but a bowl (a parabola). A bowl has a lowest point you can compute directly. So instead of a cautious small step, jump straight to the bottom of the bowl, then look again.
Gradient descent and Newton's method are the same idea at two accuracy levels: flat-sheet model (first order) versus bowl model (second order).
Minimise $f(x) = e^x - 2x$ from $x = 0$. There $f = 1$, $f' = e^0 - 2 = -1$, $f'' = e^0 = 1$.
- Second-order model: $m(\delta) = 1 - \delta + \tfrac12\delta^2$.
- Its slope is $m'(\delta) = -1 + \delta$, which is zero at $\delta = 1$ (the bottom of the parabola).
- So the new point is $x = 0 + 1 = 1$. (The true minimum is $\ln 2 = 0.693$, so we overshot a little because the true curve bends up faster than the parabola. One more step from $x=1$ gives $0.736$.)
Derivation.
- Model. $m(\boldsymbol\delta) = f + \mathbf{g}^\top\boldsymbol\delta + \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$, the second-order Taylor expansion at the current point.
- Find its lowest point (requires $H$ positive definite, a bowl). Set its gradient with respect to $\boldsymbol\delta$ to zero. The gradient of $\mathbf{g}^\top\boldsymbol\delta$ is $\mathbf{g}$ and the gradient of $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$ is $H\boldsymbol\delta$ (symmetric $H$; Chapter 2.6). So $\mathbf{g} + H\boldsymbol\delta = \mathbf{0}$.
- Solve: $\boldsymbol\delta = -H^{-1}\mathbf{g}$.
- Update: $\mathbf{x}_{\text{new}} = \mathbf{x} - H^{-1}\nabla f$. In 1D: $x_{\text{new}} = x - f'(x)/f''(x)$.
Gradient descent is Newton with a guessed curvature. Replace the true Hessian in the model by $\frac1\eta I$: $m_{\text{GD}}(\boldsymbol\delta) = f + \mathbf{g}^\top\boldsymbol\delta + \dfrac{1}{2\eta}\|\boldsymbol\delta\|^2$. Setting its gradient to zero, $\mathbf{g} + \boldsymbol\delta/\eta = \mathbf{0}$, gives $\boldsymbol\delta = -\eta\,\mathbf{g}$: the gradient step. So the learning rate $\eta$ is "one over an assumed curvature". If you guess a curvature smaller than the truth, the step is too long and it overshoots.
When it works well: if $f$ is a quadratic the model is exact, so one step lands on the minimum. Near a smooth minimum it converges very quickly (the number of correct digits roughly doubles each step). When it misbehaves: if $H$ is not positive definite the model's flat spot is not a minimum, and if the step is large the model is no longer trustworthy (see "Local approximation").
Why do we need it?
It explains where the Newton step comes from, so we can see its strengths (curvature-aware, scale-free) and weaknesses (needs a bowl, costs a matrix solve) rather than just trusting the formula.
Where is it used?
Logistic regression and generalised linear models (iteratively reweighted least squares), L-BFGS and trust-region optimisers, Gauss–Newton and Levenberg–Marquardt for least squares, and XGBoost's leaf values $-g/h$.
How is it used?
Build the quadratic model at the current point, solve $H\boldsymbol\delta = -\mathbf{g}$, and move. Safeguards: add $\lambda I$ to make $H$ positive definite, limit the step length (trust region) or do a line search.
The model is local. Jumping to the bottom of a parabola that only matches the function near your point can overshoot badly. Newton's method is safest close to the answer.
The model's bottom must be a bottom. If $H$ has a negative eigenvalue the parabola curves down in that direction and the flat spot is a saddle or a peak; the step $-H^{-1}\mathbf{g}$ then walks to it (Chapter 2.10).
Quick check: $f(x) = x^2 - 6x$. Starting from $x=0$, where does one Newton step land, and why?
$f' = 2x - 6 = -6$ and $f'' = 2$. The step is $-(-6)/2 = 3$, so $x = 3$, which is the exact minimum ($f'(3) = 0$). The function is itself a parabola, so its second-order Taylor model is exact.
Recap, cheat sheet and practice
- Near a point $a$, a smooth function is copied by the Taylor polynomial $\sum_k \frac{f^{(k)}(a)}{k!}(x-a)^k$. The $1/k!$ appears because differentiating $(x-a)^k$ exactly $k$ times gives $k!$.
- First order: $f(a+h)\approx f(a)+f'(a)h$ (tangent line). Second order adds $\tfrac12f''(a)h^2$ (curvature). In many variables: $f(\mathbf{x}+\boldsymbol\delta)\approx f+\nabla f^\top\boldsymbol\delta+\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$.
- Know the series of $e^x$, $\sin x$, $\cos x$, $\ln(1+x)$, $\frac1{1-x}$ and the sigmoid $\frac12+\frac x4-\frac{x^3}{48}$, and how to derive any of them.
- The error after degree $n$ is about the first dropped term, $O(h^{n+1})$: $h^2$ for the tangent line, $h^3$ for the parabola. On a log–log plot the slope is the order.
- The approximation is local: it is only trusted for steps small compared with the curvature, and series have a radius of convergence.
- Gradient descent is the best step for the first-order model; Newton's method is the bottom of the second-order model. Gradient descent is Newton with the Hessian replaced by $\frac1\eta I$.
Cheat sheet
| Idea | Formula | Picture |
|---|---|---|
| Taylor polynomial | $P_n(x)=\sum_{k=0}^n \dfrac{f^{(k)}(a)}{k!}(x-a)^k$ | polynomial that copies $f$ near $a$ |
| First order | $f(a)+f'(a)h$; error $\sim\tfrac12f''h^2$ | tangent line |
| Second order | $f(a)+f'(a)h+\tfrac12f''(a)h^2$; error $\sim\tfrac16f'''h^3$ | parabola |
| Multivariate | $f+\nabla f^\top\boldsymbol\delta+\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$ | tangent plane, paraboloid |
| $e^x$, $\sin x$, $\cos x$ | $\sum\frac{x^k}{k!}$; $x-\frac{x^3}{3!}+\cdots$; $1-\frac{x^2}{2!}+\cdots$ | valid for all $x$ |
| $\ln(1+x)$, $\frac1{1-x}$ | $x-\frac{x^2}2+\frac{x^3}3-\cdots$; $1+x+x^2+\cdots$ | radius $1$ |
| Sigmoid | $\frac12+\frac x4-\frac{x^3}{48}+\cdots$ | radius $\pi$ |
| Remainder | $R_n=\dfrac{f^{(n+1)}(\xi)}{(n+1)!}h^{n+1}$ | log–log slope $n+1$ |
| Gradient descent | $\mathbf{x}-\eta\nabla f$ (from the first-order model) | step against the slope |
| Newton | $\mathbf{x}-H^{-1}\nabla f$ (from the second-order model) | jump to the bowl's bottom |
import numpy as np
from math import factorial
# 1) Taylor polynomial of e^x around 0
def exp_taylor(x, n):
return sum(x**k / factorial(k) for k in range(n + 1))
for n in range(5):
print(n, exp_taylor(1.0, n)) # 1.0, 2.0, 2.5, 2.6667, 2.7083 -> e = 2.71828...
# 2) the error shrinks like h^2 (first order) and h^3 (second order)
for h in (0.2, 0.1, 0.05):
e1 = abs(np.exp(h) - (1 + h))
e2 = abs(np.exp(h) - (1 + h + h*h/2))
print(h, e1 / h**2, e2 / h**3) # ratios settle near 0.5 and 0.1667
hs = np.logspace(-2, -0.5, 20)
err1 = np.abs(np.exp(hs) - (1 + hs))
err2 = np.abs(np.exp(hs) - (1 + hs + hs**2/2))
print(np.polyfit(np.log10(hs), np.log10(err1), 1)[0]) # about 2.0 (slope of the log-log line)
print(np.polyfit(np.log10(hs), np.log10(err2), 1)[0]) # about 3.0
# 3) multivariate: f(x, y) = x^2 y + y^2 around (1, 2), step delta = (0.1, -0.1)
f = lambda p: p[0]**2 * p[1] + p[1]**2
grad = lambda p: np.array([2*p[0]*p[1], p[0]**2 + 2*p[1]])
hess = lambda p: np.array([[2*p[1], 2*p[0]], [2*p[0], 2.0]])
p0 = np.array([1.0, 2.0]); d = np.array([0.1, -0.1])
first = f(p0) + grad(p0) @ d
second = first + 0.5 * d @ hess(p0) @ d
print(first, second, f(p0 + d)) # 5.9 5.91 5.909
# 4) one gradient step vs the first-order prediction, f = 0.5*(x^2 + 4 y^2) at (2, 1)
A = np.diag([1.0, 4.0]); g = A @ np.array([2.0, 1.0]); eta = 0.1
q = lambda z: 0.5 * z @ A @ z
z0 = np.array([2.0, 1.0])
print(-eta * g @ g, q(z0 - eta*g) - q(z0)) # predicted -2.0, actual -1.66 (curvature term is +0.34)
# 5) why log1p exists
print(np.log(1 + 1e-17), np.log1p(1e-17)) # 0.0 1e-17
# 6) sigmoid near 0: 1/2 + x/4 - x^3/48
sig = lambda x: 1 / (1 + np.exp(-x))
x = 0.4
print(sig(x), 0.5 + x/4 - x**3/48) # 0.598688 0.598667
1. The Taylor series of $\sin x$ around $0$ starts…
2. A first-order Taylor estimate has error $0.08$ for a step $h = 0.4$. About what error do you expect for $h = 0.2$?
3. Which expression is the second-order multivariate Taylor approximation?
4. Why can the Taylor series of $\ln(1+x)$ not be used at $x = 2$?
5. Newton's method is derived by…
6. A model's first-order Taylor expansion predicts that a gradient step lowers the loss by $2.0$, but the loss goes up. The most likely reason is…
Practice problems
A. Find the degree-3 Taylor polynomial of $f(x)=\ln(1+x)$ at $0$ and use it to estimate $\ln 1.2$.
$f(0)=0$, $f'=\frac1{1+x}\to1$, $f''=-\frac1{(1+x)^2}\to-1$, $f'''=\frac2{(1+x)^3}\to2$. So $P_3 = x - \frac{x^2}{2} + \frac{2x^3}{6} = x - \frac{x^2}2 + \frac{x^3}3$. At $x=0.2$: $0.2 - 0.02 + 0.002667 = 0.182667$. True: $\ln1.2 = 0.182322$; error $0.00034$.
B. Use the tangent line of $f(x)=\sqrt{x}$ at $a=25$ to estimate $\sqrt{26}$, and say whether the estimate is above or below the truth.
$f(25)=5$, $f'(x)=\frac1{2\sqrt x}$, $f'(25)=\frac1{10}$. So $\sqrt{26}\approx5+0.1=5.1$. True: $5.09902$. The estimate is slightly above: $f''=-\frac14x^{-3/2}\lt0$ (a frown), so the tangent line lies above the curve.
C. For $f(x,y)=e^{x}\cos y$ find the second-order Taylor approximation at $(0,0)$ and estimate $f(0.1,\,0.2)$.
$f=e^x\cos y$: at $(0,0)$ $f=1$; $f_x=e^x\cos y=1$; $f_y=-e^x\sin y=0$; $f_{xx}=1$; $f_{xy}=-e^x\sin y=0$; $f_{yy}=-e^x\cos y=-1$. So $f\approx1+x+\tfrac12x^2-\tfrac12y^2$. At $(0.1,0.2)$: $1+0.1+0.005-0.02=1.085$. True: $e^{0.1}\cos0.2=1.10517\times0.98007=1.08314$. Error $0.0019$.
D. Derive the Taylor series of $\cos x$ at $0$ from that of $\sin x$ by differentiating.
$\sin x = x - \frac{x^3}{3!} + \frac{x^5}{5!} - \cdots$. Differentiate term by term: $\cos x = 1 - \frac{3x^2}{3!} + \frac{5x^4}{5!} - \cdots = 1 - \frac{x^2}{2!} + \frac{x^4}{4!} - \cdots$, because $\frac{3}{3!}=\frac1{2!}$, $\frac5{5!}=\frac1{4!}$, and so on.
E. Take one gradient step and one Newton step on $f(x)=x^4$ from $x=1$ with $\eta=0.05$. Compare.
$f'=4x^3=4$, $f''=12x^2=12$. Gradient step: $1-0.05\cdot4=0.8$. Newton step: $1-4/12=0.667$. Newton moves further because it knows the curvature: it uses an effective $\eta=1/12=0.083$ here. (The true minimum is $0$; Newton shrinks the distance by a factor $2/3$ each step for this function.)
F. By what factor does the error of a degree-3 Taylor polynomial drop when the step is halved? And for the tangent line?
Degree 3 means error $O(h^4)$: halving $h$ divides the error by $2^4=16$. The tangent line is $O(h^2)$: by $2^2=4$. (When the leading coefficient happens to vanish, e.g. $f''(a)=0$, the observed factor is larger.)
Linearization
Zoom in far enough on any smooth curve and it looks like a straight line. That one fact lets us swap a hard, bendy function for an easy, flat one, as long as we stay close to where we stand. Almost every optimisation method in machine learning is built on it.
- See a smooth function as "a straight line (or flat plane) up close"
- Write the linear approximation $f(x+\delta) \approx f(x) + f'(x)\,\delta$ and its many-variable forms with the gradient and the Jacobian
- Understand the Jacobian as a local linear transformation: near a point, a curvy map behaves like a matrix
- Measure the error: it is about $\tfrac12\delta^\top H\delta$, so it shrinks like $\delta^2$ (halve the step, quarter the error)
- Know when linearization fails: kinks and very high curvature
- Use it: error propagation, Newton's method, linearizing a sigmoid neuron, gradient descent with small learning rates, and why ReLU networks are piecewise linear
This chapter pulls together the derivative (Chapter 2.3), the gradient (2.4), the Jacobian (2.5), the Hessian (2.10) and Taylor series (2.11). If a symbol looks unfamiliar, jump back to that chapter. Conventions, as everywhere in this guide: the gradient $\nabla f$ is a column vector, and the Jacobian of $F:\mathbb{R}^n\to\mathbb{R}^m$ is an $m\times n$ matrix whose entry in row $i$, column $j$ is $\partial F_i/\partial x_j$. When we talk about matrices as "things that move arrows around", see Linear transformations in the Linear Algebra guide.
Linear approximation: zoom in until it is straight core
The Earth is round, yet a map of your town is flat, and nobody notices the mistake. Zoom in on a curved surface far enough and it looks flat.
A curve on a graph behaves the same way. Pick a point and zoom in. The curve bends less and less, until it is hard to tell apart from a straight line. That line is the tangent line from Chapter 2.3.
So near a point we can pretend the curve is the line. Lines are easy: you only need a starting value and a slope. This pretend-it-is-a-line trick is called linear approximation or linearization.
Take $f(x) = x^2$ near $x = 1$. At $x=1$ the value is $f(1)=1$ and the slope is $f'(1) = 2$. So the tangent line is "start at 1, rise by 2 for every 1 step to the right".
- Go $0.1$ to the right, to $x = 1.1$. The line says: $1 + 2\cdot 0.1 = 1.2$. The true value is $1.1^2 = 1.21$. The gap is $0.01$.
- Go only $0.01$ to the right, to $x = 1.01$. The line says: $1 + 2\cdot 0.01 = 1.02$. The true value is $1.01^2 = 1.0201$. The gap is $0.0001$.
We moved 10 times closer and the gap became 100 times smaller. That is the "zoom in and it gets straighter" effect in numbers.
The linear approximation (or linearization) of $f$ at the point $a$ is the function
$$L(x) = f(a) + f'(a)\,(x - a).$$Writing the step as $h = x - a$, this says
$$f(a + h) \;\approx\; f(a) + f'(a)\,h.$$- $f(a)$ is where we start (the height of the curve at $a$).
- $f'(a)$ is the slope (how fast the height changes per step).
- $h$ is the step. "Height change $\approx$ slope $\times$ step."
Where does it come from? The derivative is defined by $f'(a) = \lim_{h\to 0}\dfrac{f(a+h)-f(a)}{h}$. For a small (but not zero) $h$ the fraction is close to $f'(a)$:
$$\frac{f(a+h)-f(a)}{h} \approx f'(a) \;\;\Longrightarrow\;\; f(a+h) - f(a) \approx f'(a)\,h \;\;\Longrightarrow\;\; f(a+h) \approx f(a) + f'(a)\,h.$$We only multiplied both sides by $h$ and moved $f(a)$ across. That is the whole idea.
Why do we need it?
Most functions in machine learning are curved and messy, but straight lines are easy to reason about and to compute with. Linearization lets us do easy maths and still be nearly right, as long as we stay near the point we know.
Where is it used?
Gradient descent (each step trusts a straight-line model of the loss), Newton's method, error bars on measurements, sensitivity analysis, the Gauss–Newton method in least squares, and the extended Kalman filter in robotics and GPS.
How is it used?
Compute the value $f(a)$ and the slope $f'(a)$ once. Then estimate $f$ at any nearby point with one multiplication and one addition, instead of evaluating the hard function again.
"Approximately equal" is not "equal". The line is only an approximation. It is exactly right at the point $a$ and slowly gets worse as you walk away. The next section studies how fast.
Zoom is relative. A curve that looks straight in a window of width 0.01 may look very curved in a window of width 10. "Straight" always means "straight at this scale".
Quick check: linearize $f(x)=x^2$ at $a=3$, then estimate $3.1^2$.
$f(3)=9$ and $f'(3)=6$, so $L(x) = 9 + 6(x-3)$. At $x = 3.1$: $L(3.1) = 9 + 6\cdot0.1 = 9.6$. The true value is $9.61$, so we are off by only $0.01$.
Local approximation: how near is "near"?
A street map of your town is perfect for walking to the shop and useless for flying to another continent. A linear approximation is a street map: very good close to the spot where you drew it, worse and worse as you travel away.
"Local" simply means "close to the point $a$". The word that matters is how close. That depends on two things: how much accuracy you need (your tolerance), and how sharply the curve bends. A gentle curve keeps the line useful for a long way. A sharp curve spoils it quickly.
Linearize $f(x)=e^x$ at $a=0$. Here $f(0)=1$ and $f'(0)=e^0=1$, so $L(x) = 1 + x$. Compare the line with the truth at three distances:
| distance $h$ | true $e^h$ | line $1+h$ | gap |
|---|---|---|---|
| $0.1$ | $1.10517$ | $1.1$ | $0.0052$ |
| $0.5$ | $1.64872$ | $1.5$ | $0.1487$ |
| $2$ | $7.38906$ | $3$ | $4.389$ |
Close by, the line is excellent. Two steps away it is far off. Moving 5 times further (from $0.1$ to $0.5$) made the gap about 29 times bigger. The gap does not grow in step with the distance. It grows faster.
A function $f$ is differentiable at $a$ exactly when its linear approximation is good locally, in this precise sense:
$$f(a+h) = f(a) + f'(a)\,h + \underbrace{(\text{error})}_{\text{shrinks faster than } h}, \qquad \frac{\text{error}}{h}\to 0 \text{ as } h\to 0.$$"The error shrinks faster than $h$" is the careful way to say "the curve really does become a line". Given a tolerance $\varepsilon$ (the biggest mistake you will accept), the trust zone is the set of $x$ near $a$ where $|f(x) - L(x)| \le \varepsilon$. Inside it you may use the line. Outside it you may not.
Why do we need it?
A tangent line is only a promise about the neighbourhood of one point. To use it safely we must know how big that neighbourhood is, otherwise we may trust it far from home and get nonsense.
Where is it used?
The learning rate in gradient descent (how far a step may go before the straight-line model of the loss stops being true), trust-region optimisers, step-size control in ODE solvers, and numerical derivative checks.
How is it used?
Decide the error you can accept, then keep every step small enough that you stay in the trust zone. If your steps seem to "overshoot", the zone was smaller than you thought: shrink the step.
The trust zone is not symmetric in general. For $e^x$ the curve bends upward, so on one side the line stays good for longer than on the other. Do not assume "plus or minus the same amount".
Quick check: the tangent line to $\sin x$ at $0$ is $L(x)=x$. How wrong is it at $x=0.1$?
$\sin 0.1 = 0.099833\ldots$, so the line's answer $0.1$ is off by about $0.000167$. That is less than two parts in a thousand, so the line is excellent at this distance.
First-order Taylor approximation (many inputs) core
Now let the input be several numbers, such as the weights of a model. Nudge each input a little. How does the output change?
Because things look flat up close, each nudge acts on its own and the effects simply add up. A nudge of $0.1$ in the first input changes the output by "slope in that direction times $0.1$". A nudge of $-0.2$ in the second input does the same with its own slope. The total change is the sum.
Picture a surface (a landscape). Near your feet it looks like a tilted flat sheet, the tangent plane. The linear approximation just reads heights off that sheet instead of off the real, curved ground.
Let $f(x,y) = x^2y$. We stand at $(x,y)=(1,2)$ and step by $\delta = (0.1,\,-0.2)$.
- Value at the start: $f(1,2) = 1^2\cdot 2 = 2$.
- Partial derivatives: $\partial f/\partial x = 2xy = 4$ and $\partial f/\partial y = x^2 = 1$ at $(1,2)$. So $\nabla f = [4,\,1]^\top$.
- First-order change: $\nabla f^\top\delta = 4\cdot 0.1 + 1\cdot(-0.2) = 0.4 - 0.2 = 0.2$.
- Linear estimate: $f(1.1,\,1.8) \approx 2 + 0.2 = 2.2$.
- True value: $1.1^2\cdot 1.8 = 1.21\cdot 1.8 = 2.178$. The estimate is off by $0.022$.
The estimate is within about 1 percent after a step of size roughly $0.22$. Not bad for two multiplications and an addition.
Let $\mathbf{x}$ be a point and $\boldsymbol{\delta}$ a small step. The first-order Taylor approximation (the linearization) is:
| function | approximation |
|---|---|
| $f:\mathbb{R}\to\mathbb{R}$ | $f(x+\delta) \approx f(x) + f'(x)\,\delta$ |
| $f:\mathbb{R}^n\to\mathbb{R}$ (a loss) | $f(\mathbf{x}+\boldsymbol\delta) \approx f(\mathbf{x}) + \nabla f(\mathbf{x})^\top\boldsymbol\delta = f(\mathbf{x}) + \sum_i \dfrac{\partial f}{\partial x_i}\,\delta_i$ |
| $F:\mathbb{R}^n\to\mathbb{R}^m$ (a layer) | $F(\mathbf{x}+\boldsymbol\delta) \approx F(\mathbf{x}) + J(\mathbf{x})\,\boldsymbol\delta$ |
Here $\nabla f^\top\boldsymbol\delta$ is the dot product of the gradient (a column) with the step. The Taylor series of Chapter 2.11 keeps more terms. Keeping only the first gives the linearization.
Derivation for the middle row. Walk along the straight path $g(t) = f(\mathbf{x} + t\boldsymbol\delta)$, so $g(0) = f(\mathbf{x})$ and $g(1) = f(\mathbf{x}+\boldsymbol\delta)$. This is a one-variable function, so the line rule gives $g(1)\approx g(0) + g'(0)$. The chain rule (Chapter 2.8) gives $g'(0) = \nabla f(\mathbf{x})^\top\boldsymbol\delta$. Putting these together:
$$f(\mathbf{x}+\boldsymbol\delta) \approx f(\mathbf{x}) + \nabla f(\mathbf{x})^\top\boldsymbol\delta. \qquad\blacksquare$$Why do we need it?
A model can have millions of weights. We cannot afford to re-run it for every possible change. The gradient is one cheap number-list that predicts the effect of any small change in all the weights at once.
Where is it used?
The gradient-descent update (we choose $\boldsymbol\delta=-\eta\nabla f$ because the formula says it lowers the loss), saliency maps (which pixels matter), adversarial examples, influence functions, and Gauss–Newton.
How is it used?
Evaluate $f$ and $\nabla f$ at the current point once. Then the predicted change for any small step $\boldsymbol\delta$ is the dot product $\nabla f^\top\boldsymbol\delta$. Trust it only while $\boldsymbol\delta$ is small.
The step must be small, in every direction. The formula does not say "linear is good". It says "linear is good for small $\boldsymbol\delta$". A step that is tiny in one input and huge in another is not small.
Transpose matters. $\nabla f$ is a column, so we write $\nabla f^\top\boldsymbol\delta$ (row times column, a single number). Writing $\nabla f\,\boldsymbol\delta$ would multiply a column by a column, which does not make sense.
Quick check: for $f(x,y)=x^2+y^2$ at $(3,4)$, estimate $f(3.1,\,4)$ with the gradient.
$f(3,4)=25$ and $\nabla f = [2x,\,2y]^\top = [6,\,8]^\top$. The step is $\boldsymbol\delta=(0.1,\,0)$, so $\nabla f^\top\boldsymbol\delta = 6\cdot0.1 + 8\cdot 0 = 0.6$. Estimate: $25.6$. The truth is $3.1^2 + 16 = 25.61$.
The Jacobian as a local linear transformation core
A function $F:\mathbb{R}^2\to\mathbb{R}^2$ takes a point of the plane and sends it to another point. Imagine it as a rubber sheet with a grid drawn on it: $F$ pushes, stretches, twists and bends the sheet. Far away the grid may look wildly curved.
Now zoom in on one point. The same trick as before happens: the bent grid lines straighten, and the little grid squares become little parallelograms of equal size. That is exactly what a matrix does to a grid. (See Chapter 1.5 of the Linear Algebra guide: a matrix sends the grid to a tilted, stretched grid, and its columns say where the two basic arrows land.)
So near a point, every smooth map behaves like a matrix. That matrix is the Jacobian. Its first column says "where does a nudge in $x_1$ go?" and its second column says the same for $x_2$.
Let $F(x,y) = (x^2 - y^2,\; 2xy)$. We stand at $(1,1)$.
- Where it lands: $F(1,1) = (1-1,\; 2) = (0,\,2)$.
- The partial derivatives: $\partial F_1/\partial x = 2x$, $\partial F_1/\partial y = -2y$, $\partial F_2/\partial x = 2y$, $\partial F_2/\partial y = 2x$. At $(1,1)$ the Jacobian is $$J = \begin{bmatrix} 2 & -2 \\ 2 & 2 \end{bmatrix}.$$
- Take a small step $\boldsymbol\delta = (0.1,\,0.05)$. Then $J\boldsymbol\delta = (2\cdot0.1 - 2\cdot0.05,\; 2\cdot0.1 + 2\cdot0.05) = (0.1,\,0.3)$.
- Linear estimate: $F(1.1,\,1.05) \approx (0,2) + (0.1,0.3) = (0.1,\,2.3)$.
- True value: $(1.1^2 - 1.05^2,\; 2\cdot1.1\cdot1.05) = (1.21 - 1.1025,\; 2.31) = (0.1075,\,2.31)$.
The estimate is off by only $(0.0075,\,0.01)$. The determinant is $\det J = 2\cdot2 - (-2)\cdot 2 = 8$: tiny patches of area get stretched by a factor of 8 near this point.
For $F:\mathbb{R}^n\to\mathbb{R}^m$, the Jacobian at $\mathbf{x}$ is the $m\times n$ matrix of all first partial derivatives, with entry $\partial F_i/\partial x_j$ in row $i$, column $j$. Linearization says:
$$F(\mathbf{x}+\boldsymbol\delta) \;\approx\; F(\mathbf{x}) + J(\mathbf{x})\,\boldsymbol\delta.$$So near $\mathbf{x}$, $F$ acts like the affine map "shift to $F(\mathbf{x})$, then apply the matrix $J(\mathbf{x})$ to the step". Reading the matrix:
- Column $j$ of $J$ is where the little step $\mathbf{e}_j$ (a nudge of input $j$) goes, per unit of nudge.
- $|\det J|$ (for $n=m$) is the factor by which tiny areas (or volumes) are scaled. A negative determinant means the map flips orientation. See the determinant.
- If $m=1$ (a loss), $J$ is a single row, which is just $\nabla f^\top$. So the gradient is the Jacobian of a scalar function.
Why do we need it?
A neural-network layer is a curved map from vectors to vectors, which is hard to analyse. Its Jacobian replaces it, locally, by one matrix, and we know a lot about matrices (multiplying, inverting, eigenvalues, determinants).
Where is it used?
Backpropagation (each layer's local Jacobian is multiplied along the chain), normalising flows (the determinant of the Jacobian tracks how density is stretched), coordinate changes such as polar to Cartesian, and Gauss–Newton and Levenberg–Marquardt solvers.
How is it used?
Compute (or let autodiff compute) the Jacobian at the current point. Multiply it by a small input change to predict the output change, or look at its determinant or singular values to see how the map stretches space.
The Jacobian depends on the point. Move the dot and the matrix changes. It is a local picture: a different matrix at every location.
Rows or columns? Row $i$ is the gradient of output $i$. Column $j$ is the effect of input $j$ on all the outputs. With the convention used throughout this guide the shape is outputs $\times$ inputs.
Quick check: what is the Jacobian of the linear map $F(\mathbf{x}) = A\mathbf{x}$, and how good is its linearization?
$J = A$ at every point, because the partial derivative of $\sum_j A_{ij}x_j$ with respect to $x_j$ is $A_{ij}$. The linearization $F(\mathbf{x}) + A\boldsymbol\delta$ equals $A(\mathbf{x}+\boldsymbol\delta)$ exactly. A linear map has no curvature, so there is no error at all, for any step size.
How big is the error? It shrinks like $\delta^2$ core
The line ignores one thing: the curve bends. The error is the amount of bending you collected while walking the step. Two facts follow from that picture.
- The bendier the curve (more curvature, the second derivative), the bigger the error.
- If you walk half as far, you collect less than half of the bending. In fact you collect a quarter: the bend builds up with the square of the distance.
So a small step does not just give a small error. It gives a very small error. Halve the step and the error drops 4 times.
One variable. Linearize $e^x$ at $0$ again ($L(h) = 1 + h$). The curvature is $f''(0) = e^0 = 1$, so we guess the error is $\tfrac12 f''\,h^2 = \tfrac12h^2$.
| step $h$ | true error $e^h - 1 - h$ | guess $\tfrac12h^2$ |
|---|---|---|
| $0.1$ | $0.005171$ | $0.005$ |
| $0.05$ | $0.001271$ | $0.00125$ |
Halving the step shrinks the error by $0.005171/0.001271 \approx 4.07$. Almost exactly 4.
Two variables. Back to $f(x,y)=x^2y$ at $(1,2)$ with $\boldsymbol\delta=(0.1,-0.2)$. The Hessian (the matrix of second derivatives) is $H = \begin{bmatrix} 2y & 2x\\ 2x & 0\end{bmatrix} = \begin{bmatrix} 4 & 2\\ 2 & 0\end{bmatrix}$.
- $H\boldsymbol\delta = [\,4\cdot0.1 + 2\cdot(-0.2),\;\; 2\cdot0.1 + 0\,]^\top = [\,0,\; 0.2\,]^\top$.
- $\boldsymbol\delta^\top H\boldsymbol\delta = 0.1\cdot 0 + (-0.2)\cdot0.2 = -0.04$.
- Half of it: $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta = -0.02$.
The true error (earlier) was $2.178 - 2.2 = -0.022$. The guess $-0.02$ is very close.
Keep one more term of the Taylor series (Chapter 2.11):
$$f(\mathbf{x}+\boldsymbol\delta) = f(\mathbf{x}) + \nabla f^\top\boldsymbol\delta + \underbrace{\tfrac12\,\boldsymbol\delta^\top H\,\boldsymbol\delta}_{\text{the error of the linearization}} + (\text{terms of size }\|\boldsymbol\delta\|^3).$$So error $\approx \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$, where $H$ is the Hessian (see Chapter 2.10, and quadratic forms for the shape $\boldsymbol\delta^\top H\boldsymbol\delta$). In one variable it reads $\tfrac12f''(x)\,\delta^2$.
Where does it come from? Walk the path $g(t) = f(\mathbf{x}+t\boldsymbol\delta)$ again. A one-variable Taylor series gives $g(1) = g(0) + g'(0) + \tfrac12 g''(0) + \dots$. We saw $g'(0) = \nabla f^\top\boldsymbol\delta$. Applying the chain rule once more gives $g''(0) = \boldsymbol\delta^\top H\boldsymbol\delta$. Substituting gives the formula above.
A safe bound. If the curvature never exceeds $M$ on the way (in one variable $|f''|\le M$; in many, the Hessian's eigenvalues have size at most $M$), then $\;|\text{error}| \le \tfrac12 M\,\|\boldsymbol\delta\|^2$.
Two consequences: the error is second order (it scales like $\|\boldsymbol\delta\|^2$), and the sign of $\boldsymbol\delta^\top H\boldsymbol\delta$ tells you whether the true function sits above the tangent plane (bowl-like, positive) or below it (hill-like, negative).
Why do we need it?
An approximation without an error estimate is a guess. This formula tells us how wrong the line is, how small a step must be, and whether the true value lies above or below the line.
Where is it used?
Choosing learning rates and trust-region sizes in optimisation, proving that gradient descent converges, deciding step sizes in numerical solvers, and checking a gradient with a finite difference (that check relies on the error shrinking like $\delta^2$ or $\delta$).
How is it used?
To keep the error below a tolerance $\varepsilon$, pick a step with $\tfrac12 M\|\boldsymbol\delta\|^2 \le \varepsilon$. Or test your gradient code: compute the gap $f(\mathbf{x}+\boldsymbol\delta) - f(\mathbf{x}) - \nabla f^\top\boldsymbol\delta$ and shrink the step by 10. A correct gradient makes the gap drop by about 100 (second order). If it only drops by 10, the gradient is wrong.
"Second order" is about the step, not about the function. A function with big curvature still has error $\propto\delta^2$. The constant out front is just bigger. That is why the blue lines in the plot all run parallel at slope 2 but sit at different heights.
Quick check: a linearization has error $0.04$ for a step of $0.2$. Roughly what error do you expect for a step of $0.05$?
The step shrank by a factor of $4$, so the error shrinks by $4^2 = 16$: $0.04/16 = 0.0025$.
When linearization fails: kinks and high curvature
"Zoom in and it gets straight" is true for smooth curves. Two kinds of curve break the promise.
- A kink (corner). Zoom in on the tip of the letter V. The tip is still a tip, however far you zoom. There is no single straight line it turns into, because the left side and the right side have different slopes.
- Very high curvature. A tight bend does flatten if you zoom, but only after you have zoomed in enormously. So the trust zone is tiny and the "small step" you need may be unrealistically small.
A kink. $f(x)=|x|$ at $a=0$. To the right the slope is $+1$, to the left it is $-1$. Try the best flat line, $L(x)=0$. The error at step $h$ is $|h|$, the same size as the step. The ratio error$/h$ stays at $1$ and never tends to $0$. So $|x|$ is not differentiable at $0$, and linearization has nothing good to offer there.
High curvature. From the error formula, the error is about $\tfrac12|f''|h^2$. To keep it below a tolerance $\varepsilon$ we need $h \le \sqrt{2\varepsilon/|f''|}$.
- For $e^x$ at $0$ we have $|f''|=1$. With $\varepsilon = 0.01$: $h\le\sqrt{0.02}\approx 0.14$. (The true error at $h=0.141$ is $0.0105$, just slightly over: we ignored the third-order terms.)
- For $\sin(10x)$ at its peak, $|f''| = 100$. Now $h\le\sqrt{0.02/100}\approx 0.014$.
A curvature 100 times bigger gives a trust zone 10 times smaller.
Linearization at $a$ is trustworthy when:
- $f$ is differentiable at $a$ (no kink, no jump, no vertical tangent), and
- the step satisfies $\tfrac12\,|f''|\,h^2 \le \varepsilon$, i.e. $\;|h| \le \sqrt{2\varepsilon/|f''|}$, where $|f''|$ is the curvature near $a$ (in many variables, the largest size of an eigenvalue of the Hessian).
Condition 1 is a yes/no question. Condition 2 is a how small is small enough question, answered by the curvature.
Why do we need it?
Every method that relies on a line or a plane silently assumes smoothness and a small step. Knowing the two ways it breaks tells you why training sometimes blows up, and what to change.
Where is it used?
Explaining why gradient descent diverges when the learning rate is too large (sharp curvature), why ReLU networks have undefined gradients exactly at a kink (frameworks pick 0), why gradient clipping and warm-up help, and why optimisers use trust regions.
How is it used?
If a step makes things worse than the line predicted, shrink the step (the learning rate). If a function has kinks, use a rule at the kink (a subgradient, such as "ReLU′(0)=0"), or a smooth stand-in.
Kinks are everywhere in modern networks (ReLU, max-pooling, absolute-value losses). In practice you almost never land exactly on a kink, and away from it the function is perfectly linear. The danger is only at (or very near) the corner.
Quick check: a loss has curvature $|f''|\approx 400$. For tolerance $\varepsilon = 0.02$, how big a step can you trust?
$h \le \sqrt{2\cdot0.02/400} = \sqrt{0.0001} = 0.01$.
Sensitivity and error propagation core
Every measurement is a little wrong. You measure a table as 2.00 m but it might really be 2.01 m. If you then compute something from your measurements (an area, a speed, a model's prediction), how wrong is the answer?
Linearization answers instantly. The output wobbles by the slope times the input wobble. A steep slope means the output is very sensitive to that input: a tiny error in it makes a big error in the result. A flat slope means the input hardly matters.
With several inputs, each one adds its own wobble. Whichever input has the biggest "slope times wobble" is the one worth measuring more carefully.
One input. The area of a circle is $A=\pi r^2$. We measure $r = 10$ with an error of $\Delta r = 0.1$.
- Slope: $A'(r) = 2\pi r = 20\pi \approx 62.83$.
- Error in the area: $\Delta A \approx A'(r)\,\Delta r = 62.83\cdot 0.1 = 6.283$.
- Check: $\pi(10.1^2 - 10^2) = \pi\cdot 2.01 = 6.315$. The linear estimate is within 0.5 percent.
- Relative error: $\Delta A/A = 6.283/314.16 = 2\%$, while $\Delta r/r = 1\%$. The radius is squared, so its relative error is doubled.
Two inputs. A rectangle has width $w=5\pm0.1$ and height $h=3\pm0.2$. Its area is $A = wh = 15$.
- Slopes: $\partial A/\partial w = h = 3$ and $\partial A/\partial h = w = 5$.
- Contributions: $3\cdot0.1 = 0.3$ from the width, and $5\cdot 0.2 = 1.0$ from the height.
- Worst case (both errors push the same way): $0.3 + 1.0 = 1.3$. So $A = 15\pm1.3$.
- Independent random errors: they sometimes cancel, so combine them like the sides of a right triangle: $\sqrt{0.3^2 + 1.0^2} = \sqrt{1.09} \approx 1.04$.
The height error dominates, so measuring the height more carefully is what would help.
From $f(\mathbf{x}+\boldsymbol\delta)\approx f(\mathbf{x}) + \nabla f^\top\boldsymbol\delta$ we get, with input errors $\Delta x_i$:
$$\Delta f \;\approx\; \sum_i \frac{\partial f}{\partial x_i}\,\Delta x_i \qquad\text{(one input: } \Delta y \approx f'(x)\,\Delta x\text{)}.$$- Worst case (each error is at most $|\Delta x_i|$): $\;|\Delta f| \le \sum_i \left|\dfrac{\partial f}{\partial x_i}\right|\,|\Delta x_i|$. This is the "sensitivity" of $f$ to each input, times how wrong that input can be.
- Independent random errors with typical sizes $\sigma_i$: $\;\sigma_f \approx \sqrt{\sum_i \left(\dfrac{\partial f}{\partial x_i}\sigma_i\right)^2}$. (Independent random wobbles add as squares, because they partly cancel.)
Both formulas only hold while the errors are small, so that the function is nearly linear over them. That is the linearization assumption, again.
Why do we need it?
To put honest error bars on anything we compute from imperfect numbers, and to find which input matters most. Without it we would have to recompute the answer for many possible input values.
Where is it used?
Lab measurements and engineering tolerances, saliency maps (which input pixels move the output most), adversarial examples such as FGSM (the loss changes by about $\nabla L^\top\boldsymbol\delta$, which is largest for $\boldsymbol\delta=\varepsilon\,\mathrm{sign}(\nabla L)$, giving $\varepsilon\|\nabla L\|_1$), and the condition number of a computation.
How is it used?
Compute the partial derivative for each input, multiply by that input's possible error, and add the pieces (straight sum for the worst case, square-root-of-squares for random errors). The biggest piece tells you what to fix first.
Errors only add up like this when they are small. For a huge error the function bends, the linear estimate is off, and (as the previous sections showed) the error grows like the square of the input error.
Quick check: $v = d/t$ with $d=100$ m (error 1 m) and $t=20$ s (error 0.2 s). Which error hurts more?
$\partial v/\partial d = 1/t = 0.05$, so the distance contributes $0.05\cdot1 = 0.05$. $\partial v/\partial t = -d/t^2 = -0.25$, so the time contributes $0.25\cdot0.2 = 0.05$. They contribute equally: the worst case is $0.05+0.05 = 0.1$ m/s on a speed of $5$ m/s.
Newton's method: repeated linearization
Suppose you want a number $x$ where $f(x)=0$ (a root), but $f$ is curvy and you cannot solve it by hand. Here is a neat trick.
- Stand at a guess $x_0$. Replace the curve by its tangent line. (Linearize.)
- A line is easy: find where the line crosses zero. Jump there. That is your next guess.
- Repeat from the new spot.
Each jump lands much closer to the true root, because the line fits the curve better and better as you get nearer. After a few jumps the number of correct digits doubles each time.
Find $\sqrt2$ by solving $f(x) = x^2 - 2 = 0$. The slope is $f'(x) = 2x$. Start at $x_0=1$.
- $f(1) = -1$, $f'(1)=2$. The line crosses zero at $x_1 = 1 - \dfrac{-1}{2} = 1.5$.
- $f(1.5) = 0.25$, $f'(1.5)=3$. So $x_2 = 1.5 - \dfrac{0.25}{3} = 1.41667$.
- $f(1.41667) = 0.006944$, $f' = 2.83333$. So $x_3 = 1.41667 - 0.002451 = 1.414216$.
- One more step gives $x_4 = 1.414213562375$. The true $\sqrt2 = 1.414213562373\ldots$
The errors are $0.414,\; 0.086,\; 0.0025,\; 0.0000021,\; 0.0000000000016$. The number of correct digits roughly doubles each step.
Newton's method. Linearize $f$ at the current guess $x_n$: $L(x) = f(x_n) + f'(x_n)(x - x_n)$. The new guess is where the line is zero:
$$0 = f(x_n) + f'(x_n)(x_{n+1} - x_n) \;\;\Longrightarrow\;\; x_{n+1} = x_n - \frac{f(x_n)}{f'(x_n)}.$$(We just solved a line equation for $x_{n+1}$: move $f(x_n)$ across, then divide by $f'(x_n)$.)
- Near a root where $f'\neq0$, the error is roughly squared each step: $\varepsilon_{n+1}\approx \dfrac{f''}{2f'}\,\varepsilon_n^2$. This is "quadratic convergence". It comes straight from the $\tfrac12 f''h^2$ error of linearization.
- It can fail: if $f'(x_n)=0$ the line is flat and never crosses zero, and a bad start can bounce between points or fly away.
- For optimisation, minimise $f$ by solving $f'(x)=0$. Newton then reads $x_{n+1} = x_n - f'(x_n)/f''(x_n)$, and in many variables $\mathbf{x}\leftarrow\mathbf{x} - H^{-1}\nabla f$ with the Hessian $H$ (Chapter 2.10).
Why do we need it?
Many equations cannot be solved with algebra. Newton's method turns a hard equation into a short list of easy line problems, and it is very fast once you are close.
Where is it used?
Computing square roots inside your calculator or library, solving equations in physics and finance, logistic regression (the IRLS algorithm is Newton's method), second-order optimisers, and L-BFGS (a cheaper cousin).
How is it used?
Pick a starting guess, compute $f$ and $f'$, update $x\leftarrow x - f/f'$, and stop when $|f(x)|$ is tiny. Always check the starting point: if $f'$ is near 0, choose a different start.
Newton's method is a local method. It trusts the tangent line, so it needs a start that is reasonably near a root, and a function that is not too flat there. Far from the root, the line can send you somewhere worse.
Quick check: do one Newton step for $f(x)=x^2-9$ starting at $x_0=5$.
$f(5)=16$ and $f'(5)=10$, so $x_1 = 5 - 16/10 = 3.4$. The root is $3$, so we already moved from 2 away to $0.4$ away.
Linearizing a sigmoid / neuron
A single neuron computes a weighted sum $z = w_1x_1 + w_2x_2 + b$ and squashes it with the sigmoid $\sigma(z) = 1/(1+e^{-z})$. The sum is already linear. All the bending comes from the S-shaped squash.
Look at the S-curve at the middle: it is almost a straight tilted line. So for small changes around a point, the neuron is a tilted flat sheet: change the inputs a little, and the output changes by (steepness of the S) times (the change in $z$). Out in the flat tails, the S has no steepness, so the neuron barely reacts to anything. That is the saturated neuron.
The sigmoid has the neat slope $\sigma'(z) = \sigma(z)\,(1-\sigma(z))$. At $z=0$: $\sigma(0)=0.5$ and $\sigma'(0) = 0.5\cdot0.5 = 0.25$. So near zero,
$$\sigma(z) \approx 0.5 + 0.25\,z.$$- At $z=0.4$: line gives $0.5 + 0.25\cdot0.4 = 0.6$. True: $\sigma(0.4) = 0.59869$. Off by $0.0013$.
- At $z = 2$: line gives $0.5+0.5 = 1.0$. True: $\sigma(2)=0.8808$. Off by $0.12$. The line is already poor, because the S bends over.
At $z=2$ the slope is $0.8808\cdot0.1192 = 0.105$, less than half of $0.25$. At $z=5$ it is only $0.0066$.
For the neuron $y = \sigma(\mathbf{w}^\top\mathbf{x} + b)$, let $z_0=\mathbf{w}^\top\mathbf{x}_0 + b$ and $y_0=\sigma(z_0)$. The linearization around the input $\mathbf{x}_0$ is
$$y(\mathbf{x}_0+\boldsymbol\delta) \approx y_0 + \sigma'(z_0)\;\mathbf{w}^\top\boldsymbol\delta, \qquad \nabla_{\mathbf{x}}\,y = \sigma'(z_0)\,\mathbf{w} = y_0(1-y_0)\,\mathbf{w}.$$Derivation. Chain rule (Chapter 2.8): the inner function $z=\mathbf{w}^\top\mathbf{x}+b$ has gradient $\mathbf{w}$, the outer function $\sigma$ has slope $\sigma'(z_0)$. Multiply them: $\sigma'(z_0)\,\mathbf{w}$. Then the first-order formula gives the line above.
The level lines (where $y$ is constant) of the linearized neuron are straight, parallel lines perpendicular to $\mathbf{w}$. The real neuron's level lines are also straight lines (since $y$ depends only on $z$), but they are unevenly spaced: crowded in the middle, far apart in the tails.
Why do we need it?
It shows exactly how sensitive a neuron is to its inputs and weights: the slope $\sigma'$. It explains saturation: when $|z|$ is large the slope is nearly zero, so the gradient that flows backwards almost vanishes.
Where is it used?
The vanishing-gradient problem in deep sigmoid and tanh networks, weight-initialisation rules (keep $z$ near the steep middle), logistic regression, and analysing how a small input change moves a prediction.
How is it used?
Compute $z_0$ and $y_0$ in the forward pass. The local slope is $y_0(1-y_0)$ (free, no extra work). The backward pass multiplies by it. If it is close to 0, expect almost no learning signal through that neuron.
The sigmoid's biggest slope is only $0.25$ (at $z=0$). So even a perfectly placed sigmoid multiplies the backward signal by at most $0.25$ per layer (this is the sigmoid's own factor; the weights multiply it too). Stack ten such layers and the sigmoid factors alone shrink the gradient to at most $0.25^{10}\approx 10^{-6}$ of its starting size. This is why deep sigmoid networks were hard to train, and why ReLU took over.
Quick check: at $z=1$, $\sigma(1)=0.7311$. What is the slope $\sigma'(1)$, and what does the line predict for $\sigma(1.2)$?
$\sigma'(1) = 0.7311\cdot(1-0.7311) = 0.7311\cdot0.2689 = 0.1966$. Line: $0.7311 + 0.1966\cdot0.2 = 0.7704$. The true $\sigma(1.2)=0.7685$.
Why small learning rates: gradient descent is a chain of linear approximations core
Stand on a hilly landscape in thick fog. You can only feel the slope under your feet. So you picture the ground as a flat tilted sheet (that is the linearization!) and walk a little way downhill on that sheet. Then you feel the slope again, draw a new sheet, and repeat.
There is a catch. A tilted flat sheet has no bottom: if you believed it forever you would walk downhill for ever. The real ground curves back up. So the sheet is only trustworthy for a short step. The learning rate $\eta$ is exactly "how far do I trust the sheet?". Too big, and you walk off the end of the sheet and may even end up higher than you started.
Let the loss be $L(x,y) = x^2 + 4y^2$ and stand at $\boldsymbol\theta=(2,1)$, where $L = 4+4=8$.
- Gradient: $\mathbf{g} = [2x,\,8y]^\top = [4,\,8]^\top$, so $\|\mathbf{g}\|^2 = 16+64 = 80$.
- Hessian: $H=\mathrm{diag}(2, 8)$, so $\mathbf{g}^\top H\mathbf{g} = 2\cdot16 + 8\cdot64 = 32+512=544$.
- Step with $\eta = 0.1$: the line predicts a drop of $\eta\|\mathbf{g}\|^2 = 8$ (all the way to 0!).
- Reality: the new point is $(2-0.4,\,1-0.8) = (1.6,\,0.2)$, with $L = 2.56 + 0.16 = 2.72$. The true drop is $8 - 2.72 = 5.28$. The curvature "ate" $\tfrac12\eta^2\,\mathbf{g}^\top H\mathbf{g} = 0.5\cdot0.01\cdot544 = 2.72$ of the predicted drop: $8 - 2.72 = 5.28$ ✓.
- Step with $\eta=0.3$: the new point is $(0.8,\,-1.4)$ and $L = 0.64 + 4\cdot1.96 = 8.48$. The loss went up, even though we stepped in the downhill direction.
One gradient-descent step is $\boldsymbol\theta_{\text{new}} = \boldsymbol\theta - \eta\,\mathbf{g}$ with $\mathbf{g}=\nabla L(\boldsymbol\theta)$. Put $\boldsymbol\delta = -\eta\mathbf{g}$ into the linearization and the second-order error term:
$$L(\boldsymbol\theta - \eta\mathbf{g}) \;\approx\; \underbrace{L(\boldsymbol\theta)}_{\text{now}} \;-\; \underbrace{\eta\,\|\mathbf{g}\|^2}_{\text{what the line promises}} \;+\; \underbrace{\tfrac12\,\eta^2\,\mathbf{g}^\top H\mathbf{g}}_{\text{curvature spoils it}}.$$(The middle term: $\nabla L^\top\boldsymbol\delta = \mathbf{g}^\top(-\eta\mathbf{g}) = -\eta\|\mathbf{g}\|^2$. The last term: $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$ with $\boldsymbol\delta=-\eta\mathbf{g}$.) The loss goes down exactly when the promise beats the spoiler:
$$\eta\,\|\mathbf{g}\|^2 > \tfrac12\eta^2\,\mathbf{g}^\top H\mathbf{g} \;\;\Longleftrightarrow\;\; \eta < \frac{2\,\|\mathbf{g}\|^2}{\mathbf{g}^\top H\mathbf{g}}.$$In the example the limit is $2\cdot80/544 \approx 0.294$, which is why $\eta=0.3$ failed. Since $\mathbf{g}^\top H\mathbf{g}\le\lambda_{\max}\|\mathbf{g}\|^2$ (with $\lambda_{\max}$ the largest Hessian eigenvalue), any $\eta<2/\lambda_{\max}$ is safe. Sharper curvature forces a smaller learning rate. See the Hessian and convexity.
Why do we need it?
It explains what the learning rate really is, and why it must be small: it is the distance we trust a straight-line model of the loss. It also explains why training sometimes diverges.
Where is it used?
Choosing learning rates for SGD and Adam, learning-rate warm-up and decay, line search, trust-region methods, and the analysis showing gradient descent converges for $\eta<2/L$ (where $L$ bounds the curvature).
How is it used?
If the loss goes up or oscillates, cut the learning rate. A common rule of thumb: stay below $1/\lambda_{\max}$ of the Hessian. If you cannot compute the Hessian, lower $\eta$ by factors of 3 or 10 until the loss falls smoothly.
The safe-step formula uses the curvature at the current point. Curvature can change as you move (see the flat-bottom bowl), so a learning rate that is safe now can be too big or too small later. That is why schedules and adaptive methods exist.
Quick check: for $L=\tfrac12\lambda\theta^2$ (one variable), what is the largest safe learning rate?
Here $g=\lambda\theta$ and $H=\lambda$, so $\|g\|^2=\lambda^2\theta^2$ and $g^\top Hg=\lambda^3\theta^2$. The limit is $2\lambda^2\theta^2/(\lambda^3\theta^2) = 2/\lambda$. You can also see it directly: each step multiplies $\theta$ by $(1-\eta\lambda)$, which shrinks $|\theta|$ exactly when $0<\eta<2/\lambda$.
Why neural networks look piecewise linear
A ReLU neuron outputs $\max(0, z)$: it is either off (outputs 0) or on (passes $z$ straight through). Inside either state it is perfectly linear. So a whole network built from ReLUs is a patchwork: it is a different straight-line (or flat-sheet) function in each region, with bends where some neuron switches on or off.
Now the link to this chapter. Stand at a point and look at the activation pattern (which neurons are on). Inside that patch the network is a linear map, so its linearization is exact there, and the Jacobian is just a product of the weight matrices with the "off" neurons switched out. Even for smooth networks (sigmoid, tanh) the picture is the same, only with the corners rounded.
A tiny one-input network with three ReLU neurons: $y = 1\cdot\mathrm{ReLU}(x+1.5) - 2\cdot\mathrm{ReLU}(x-0.5) + 3\cdot\mathrm{ReLU}(x-2)$. The neurons switch on at $x=-1.5,\;0.5,\;2$. The slope is the sum of $v_i$ over the neurons that are on:
| region | neurons on | slope |
|---|---|---|
| $x < -1.5$ | none | $0$ |
| $-1.5 < x < 0.5$ | first | $1$ |
| $0.5 < x < 2$ | first, second | $1 - 2 = -1$ |
| $x > 2$ | all three | $1 - 2 + 3 = 2$ |
Check some values: $y(0.5) = 2$, $y(2) = 3.5 - 2\cdot1.5 = 0.5$, $y(3) = 4.5 - 5 + 3 = 2.5$. The graph is four straight pieces joined at the three kinks.
For a ReLU network $F(\mathbf{x}) = W_2\,\mathrm{ReLU}(W_1\mathbf{x} + \mathbf{b}_1) + \mathbf{b}_2$, let $D(\mathbf{x})$ be the diagonal matrix with a $1$ for each neuron that is on at $\mathbf{x}$ and a $0$ for each that is off. In the region where $D$ does not change,
$$F(\mathbf{x}) = W_2\,D\,W_1\,\mathbf{x} + \text{(constant)}, \qquad J = W_2\,D\,W_1.$$This is exactly linear (affine) there, so linearization has zero error inside the region, and the error appears only when you cross a kink. For smooth activations, $D$ is replaced by the diagonal matrix of slopes $\mathrm{diag}(\sigma'(z_i))$, and $J = W_2\,\mathrm{diag}(\sigma'(z))\,W_1$ is still the exact Jacobian at that point (it is the chain rule of Chapter 2.8, a product of Jacobians), but the straight-line model it gives is only good close to the point. (ReLU's slope at exactly 0 is not defined; software uses 0 by convention.)
Why do we need it?
It tells us what a ReLU network is at any point: one linear map. That makes its gradients and Jacobians easy to understand, and explains why the gradient is piecewise constant.
Where is it used?
Counting the linear regions of a network (a measure of its expressiveness), explaining backprop in ReLU networks (gates pass or block the gradient), adversarial robustness analysis, and verifying networks by checking each linear region.
How is it used?
To understand a prediction at an input, look at which neurons are on, multiply the on-paths of weights, and read off the local linear model. To verify or test a gradient, remember it only changes at the kinks.
"Piecewise linear" does not mean "simple". A modest network can have an enormous number of pieces, so the whole function can be extremely wiggly. But up close, at any one input, it is just a matrix.
Quick check: in the example network, what is the slope at $x=1$, and what is $y(1)$?
At $x=1$ the first two neurons are on (since $1>-1.5$ and $1>0.5$) and the third is off ($1<2$). Slope $=1-2=-1$. Value: $y(1) = 1\cdot(2.5) - 2\cdot(0.5) + 0 = 1.5$.
Recap, cheat sheet and practice
- Zoom in on a smooth function and it looks linear: $f(x+\delta)\approx f(x) + f'(x)\,\delta$. For many inputs it is $f(\mathbf{x}+\boldsymbol\delta)\approx f(\mathbf{x}) + \nabla f^\top\boldsymbol\delta$, and for vector outputs $F(\mathbf{x}+\boldsymbol\delta)\approx F(\mathbf{x}) + J\boldsymbol\delta$.
- This is a local statement. The trust zone depends on the tolerance and on the curvature.
- The Jacobian is the matrix a curvy map turns into near a point: its columns say where nudges of each input go, and $|\det J|$ is the local area stretch.
- The error is about $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$: it shrinks like $\delta^2$ (halve the step, quarter the error). Linearization fails at kinks and needs tiny steps where the curvature is high: $|h|\lesssim\sqrt{2\varepsilon/|f''|}$.
- Uses: error propagation $\Delta f\approx\sum \frac{\partial f}{\partial x_i}\Delta x_i$; Newton's method $x\leftarrow x-f/f'$ (repeated linearization, digits double each step); the sigmoid slope $\sigma(1-\sigma)\le0.25$; gradient descent trusts a plane for a step of length $\eta\|\nabla L\|$ and is safe when $\eta<2/\lambda_{\max}$; ReLU networks are piecewise linear, with $J=W_2DW_1$ inside each region.
Cheat sheet
| Idea | Formula | Picture |
|---|---|---|
| Linearization (1D) | $L(x) = f(a) + f'(a)(x-a)$ | tangent line |
| First-order, scalar loss | $f(\mathbf{x}+\boldsymbol\delta)\approx f + \nabla f^\top\boldsymbol\delta$ | tangent plane |
| First-order, vector map | $F(\mathbf{x}+\boldsymbol\delta)\approx F + J\boldsymbol\delta$ | grid becomes parallelograms |
| Error | $\approx \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$, bounded by $\tfrac12 M\|\boldsymbol\delta\|^2$ | gap between curve and line |
| Trust radius | $|h|\le\sqrt{2\varepsilon/|f''|}$ | width of the green band |
| Error propagation | worst case $\sum|\partial_i f||\Delta x_i|$; random $\sqrt{\sum(\partial_i f\,\sigma_i)^2}$ | box over level lines |
| Newton step | $x_{n+1} = x_n - f(x_n)/f'(x_n)$ | tangent hits zero |
| Sigmoid | $\sigma'=\sigma(1-\sigma)$, $\sigma\approx0.5+0.25z$ near 0 | S-curve slope |
| Gradient-descent drop | $\eta\|\mathbf{g}\|^2 - \tfrac12\eta^2\mathbf{g}^\top H\mathbf{g}$; safe if $\eta<2\|\mathbf{g}\|^2/\mathbf{g}^\top H\mathbf{g}$ | line vs bowl |
| ReLU network | $J = W_2\,D\,W_1$ in each region | straight pieces |
import numpy as np
# 1. Linearization of exp at a = 0, and how the error shrinks with the step
f = np.exp
a = 0.0
for h in [0.1, 0.05, 0.025]:
err = f(a + h) - (f(a) + f(a) * h) # f'(a) = e^a = 1
print(h, err, 0.5 * h**2) # error vs the guess (1/2) f'' h^2
# 0.1 0.005170918... 0.005
# 0.05 0.001271096... 0.00125 (half the step: about 4 times smaller error)
# 0.025 0.000315120... 0.0003125
# 2. Jacobian as a local linear map: F(x, y) = (x^2 - y^2, 2xy) at (1, 1)
def F(v):
x, y = v
return np.array([x**2 - y**2, 2 * x * y])
def jacobian(F, v, eps=1e-6): # finite-difference Jacobian (m x n)
v = np.asarray(v, float)
cols = [(F(v + eps * e) - F(v - eps * e)) / (2 * eps) for e in np.eye(len(v))]
return np.stack(cols, axis=1)
p = np.array([1.0, 1.0])
J = jacobian(F, p)
d = np.array([0.1, 0.05])
print(J.round(4)) # [[ 2. -2.] [ 2. 2.]]
print(F(p) + J @ d) # linear estimate [0.1 2.3]
print(F(p + d)) # true value [0.1075 2.31]
print(np.linalg.det(J)) # about 8 (the local area stretch)
# 3. Newton's method for sqrt(2): the error is roughly squared each step
x = 1.0
for n in range(4):
x = x - (x**2 - 2) / (2 * x)
print(n + 1, x, abs(x - np.sqrt(2)))
# 1 1.5 0.0858
# 2 1.41666... 0.00245
# 3 1.4142156862... 2.1e-06
# 4 1.4142135623746... 1.6e-12
# 4. Error propagation for A = w * h (w = 5 +- 0.1, h = 3 +- 0.2)
w, h, dw, dh = 5.0, 3.0, 0.1, 0.2
worst = h * dw + w * dh # sum of |slope| * |error|
rand = np.hypot(h * dw, w * dh) # independent random errors
print(worst, rand) # 1.3 1.0440...
# 5. Gradient descent: line promises vs the truth, for L = x^2 + 4 y^2 at (2, 1)
L = lambda t: t[0]**2 + 4 * t[1]**2
t = np.array([2.0, 1.0]); g = np.array([2 * t[0], 8 * t[1]]); H = np.diag([2.0, 8.0])
for eta in [0.1, 0.3]:
promise = eta * g @ g
spoil = 0.5 * eta**2 * g @ H @ g
print(eta, promise, promise - spoil, L(t) - L(t - eta * g))
# 0.1 8.0 5.28 5.28 (the loss falls by 5.28)
# 0.3 24.0 -0.48 -0.48 (negative drop: the loss went UP)
print(2 * (g @ g) / (g @ H @ g)) # largest safe eta: 0.2941...
1. The linearization of $f$ at $a$ is…
2. A linearization has error $0.08$ for a step of $0.4$. What error do you expect for a step of $0.2$?
3. At which point does linearization fail no matter how far you zoom in?
4. One Newton step for $f(x)=x^2-25$ starting at $x_0=10$ gives $x_1=$…
5. A loss near a point is $L=5\theta^2$ (so the curvature is $L''=10$). Gradient descent lowers the loss for every learning rate below a limit. What is that limit?
6. What is the Jacobian of $F(x,y) = (x+y,\; xy)$ at the point $(2,3)$?
Practice problems
A. Linearize $f(x)=\sqrt{x}$ at $a=4$ and use it to estimate $\sqrt{4.2}$. How big is the error?
$f(4)=2$ and $f'(x)=\dfrac{1}{2\sqrt x}$, so $f'(4)=\dfrac14$. Then $L(x) = 2 + \tfrac14(x-4)$. At $x=4.2$: $L = 2 + 0.25\cdot0.2 = 2.05$. The true value is $2.04939\ldots$, so the error is about $0.0006$. (Check with $\tfrac12|f''|h^2$: $f''(x) = -\tfrac14x^{-3/2}$, so $f''(4) = -\tfrac14\cdot\tfrac18 = -\tfrac1{32}$, giving $\tfrac12\cdot\tfrac1{32}\cdot0.04 = 0.000625$ ✓.)
B. For $f(x,y)=x^2+3xy$ at $(1,2)$ with $\boldsymbol\delta=(0.05,\,-0.1)$: find the linear estimate, the true value, and $\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$.
$f(1,2) = 1 + 6 = 7$. $\nabla f = [2x+3y,\;3x]^\top = [8,\,3]^\top$. So $\nabla f^\top\boldsymbol\delta = 8\cdot0.05 + 3\cdot(-0.1) = 0.4 - 0.3 = 0.1$, and the linear estimate is $7.1$. The truth: $1.05^2 + 3\cdot1.05\cdot1.9 = 1.1025 + 5.985 = 7.0875$, so the error is $-0.0125$.
The Hessian is $H = \begin{bmatrix}2&3\\3&0\end{bmatrix}$. $H\boldsymbol\delta = [0.1 - 0.3,\;0.15]^\top = [-0.2,\,0.15]^\top$, and $\boldsymbol\delta^\top H\boldsymbol\delta = 0.05\cdot(-0.2) + (-0.1)(0.15) = -0.01 - 0.015 = -0.025$. Half is $-0.0125$: exactly the error. (For a quadratic function the third-order terms are zero, so the formula is exact.)
C. A density is $\rho = m/V$ with $m = 50\pm1$ g and $V = 20\pm0.5$ cm³. Give $\rho$, the worst-case error, and the random-error estimate.
$\rho = 50/20 = 2.5$. Slopes: $\partial\rho/\partial m = 1/V = 0.05$ and $\partial\rho/\partial V = -m/V^2 = -50/400 = -0.125$. Contributions: $0.05\cdot1 = 0.05$ and $0.125\cdot0.5 = 0.0625$. Worst case: $0.05 + 0.0625 = 0.1125$ (so $\rho = 2.5\pm0.11$, a 4.5 percent error: the relative errors $1/50 = 2\%$ and $0.5/20 = 2.5\%$ add). Random: $\sqrt{0.05^2 + 0.0625^2} = \sqrt{0.00640625}\approx 0.080$.
D. Do two Newton steps for $f(x) = x^3 - 2$ from $x_0=1$ (to approximate $\sqrt[3]{2}$).
$f'(x) = 3x^2$. Step 1: $f(1) = -1$, $f'(1) = 3$, so $x_1 = 1 + 1/3 = 1.3333$. Step 2: $f(1.3333) = 2.3704 - 2 = 0.3704$, $f'(1.3333) = 3\cdot1.7778 = 5.3333$, so $x_2 = 1.3333 - 0.0694 = 1.2639$. The true value is $1.2599$. Two steps already give 2 correct digits.
E. The polar-to-Cartesian map is $F(r,\theta) = (r\cos\theta,\; r\sin\theta)$. Find its Jacobian and determinant at $(r,\theta)=(2,\pi/2)$ and say what they mean.
$J = \begin{bmatrix}\cos\theta & -r\sin\theta\\ \sin\theta & r\cos\theta\end{bmatrix}$. At $(2,\pi/2)$: $\cos=0$, $\sin=1$, so $J = \begin{bmatrix}0 & -2\\ 1 & 0\end{bmatrix}$ and $\det J = 0\cdot 0 - (-2)(1) = 2$ (in general $\det J = r$). The first column says that a nudge of $r$ moves the point straight up (here $(0,1)$ per unit). The second column says a nudge of $\theta$ moves it left by $r=2$ per radian. The determinant $r$ says small patches in $(r,\theta)$ space are stretched by a factor of $r$: bigger circles have bigger patches.
F. For $L(\theta) = 3\theta^2$ at $\theta = 2$, find the safe range of learning rates, and compare $\eta=0.1$ with $\eta=0.5$.
$L'' = 6$, so the safe range is $0<\eta<2/6\approx 0.333$. The gradient is $g = 6\theta = 12$ and $L = 12$. With $\eta=0.1$: $\theta_{\text{new}} = 2 - 1.2 = 0.8$ and $L = 3\cdot0.64 = 1.92$: a big drop. With $\eta = 0.5$: $\theta_{\text{new}} = 2 - 6 = -4$ and $L = 3\cdot16 = 48$: the loss quadrupled. The step overshot past the minimum and landed higher up the other side of the bowl.
Multivariable Calculus
Fields, flows, spin and flux. This chapter names the tools that describe "a quantity at every point of space": a temperature map, a wind map, the way a flow expands or swirls. It is a tour for awareness: you should know what each tool means and where it appears.
Lower priority for ML (awareness). For machine learning, the tools that matter most are the ones you have already met: gradients (2.4), Jacobians (2.5), Hessians (2.10) and Taylor expansions (2.11). Divergence, curl, line integrals and surface integrals come from physics. They appear in a few corners of ML (diffusion models, normalising flows, physics-informed networks) but you can train most models without ever computing one. So this chapter is deliberately short: one clear picture per idea, and the honest place each one shows up. Sections marked awareness are for recognition, not mastery.
- Tell a scalar field (a number at every point) from a vector field (an arrow at every point)
- Recognise a gradient field and test whether a field is conservative
- Recap the directional derivative, the Jacobian and the Hessian as questions about fields
- Read divergence as "net outflow" and curl as "spin"
- Know what a line integral (work along a path) and a surface integral (flux) measure
- Recognise the big theorems (fundamental theorem for line integrals, Green, Stokes, Divergence) and where the ideas show up in ML
Scalar fields: a number at every point awareness
Look at a weather map showing the temperature. At every spot on the map there is one number: 18 degrees here, 25 degrees there. That whole map is a scalar field. ("Scalar" just means "a single number".)
You can draw it with colours (hot = orange, cold = blue) or with contour lines joining points of equal value (isotherms, or "level curves"). You met both pictures in Chapter 2.4. A mountain map is the same idea: the number is the height.
Let $T(x,y) = 30 - x^2 - y^2$: a hot spot in the middle that cools as you move away.
- $T(0,0) = 30$ (the hottest point).
- $T(1,2) = 30 - 1 - 4 = 25$.
- $T(3,0) = 30 - 9 = 21$. And also $T(0,-3) = 21$, $T(2.12, 2.12)\approx 21$: points the same distance from the middle share a value, so the level curves are circles.
A scalar field is a function $f:\mathbb{R}^n\to\mathbb{R}$ that gives one number $f(\mathbf{x})$ to every point $\mathbf{x}$ of a region. Its level set for the value $c$ is $\{\mathbf{x} : f(\mathbf{x}) = c\}$ (a curve in 2D, a surface in 3D).
You already know examples: every loss function $L(\boldsymbol\theta)$ is a scalar field over "weight space". Its gradient $\nabla f$ turns it into a vector field (next sections).
Why do we need it?
We need a name and a picture for "a quantity that depends on position". It is the setting where gradients, contours and optimisation live.
Where is it used?
Loss surfaces over model weights, probability density functions $p(\mathbf{x})$, image brightness as a function of pixel position, temperature, pressure and potential energy in physics, and energy-based models.
How is it used?
Sample it on a grid and draw a heat map or contours to see its shape (hills, valleys, ridges). Then take its gradient to find which way it climbs.
Quick check: with $T(x,y)=30-x^2-y^2$, is $T(2,1)$ bigger or smaller than $T(0,2)$?
$T(2,1) = 30 - 4 - 1 = 25$ and $T(0,2) = 30 - 0 - 4 = 26$. So $T(0,2)$ is bigger: that point is slightly closer to the hot centre ($\sqrt{4}=2$ versus $\sqrt{5}\approx2.24$).
Vector fields: an arrow at every point awareness
Now picture a wind map. At every spot there is an arrow: which way the air is moving, and how fast (the arrow's length or darkness). That is a vector field: a vector attached to every point.
Drop a leaf into the wind and it travels along a path that is always tangent to the arrows. That path is a flow line (or streamline). Vector fields are how we describe flowing water, moving air, and force fields such as gravity.
The field $\mathbf{F}(x,y) = (-y,\;x)$ attaches to the point $(x,y)$ the arrow $(-y, x)$.
- At $(1,0)$ the arrow is $(0,1)$: straight up.
- At $(0,2)$ the arrow is $(-2,0)$: to the left, and twice as long.
- At $(-1,0)$ the arrow is $(0,-1)$: straight down.
Going around the origin, the arrows turn counter-clockwise. A leaf released here would circle the origin. This is a pure rotation.
A vector field is a function $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^n$ that gives a vector $\mathbf{F}(\mathbf{x})$ to each point $\mathbf{x}$. In 2D: $\mathbf{F}(x,y) = (F_x(x,y),\,F_y(x,y))$, i.e. two scalar fields, one per component.
A flow line is a curve $\mathbf{x}(t)$ whose velocity is the field at every moment: $\dfrac{d\mathbf{x}}{dt} = \mathbf{F}(\mathbf{x})$. The Jacobian of a vector field is a square matrix that says how the arrows change from place to place.
Why do we need it?
Many things have a direction and a size at every location (wind, current, force, "which way is uphill"). A vector field describes them all at once, and its flow lines show where things end up.
Where is it used?
Fluid and weather simulation, electric and gravitational forces, the "score" $\nabla\log p(\mathbf{x})$ that diffusion models learn (a vector field pointing towards likely data), neural ODEs and continuous normalising flows (a learned field moves the data), and gradient descent (the field $-\nabla L$).
How is it used?
Draw arrows on a grid to see the pattern. To follow a particle, take small steps along the arrow at its current position and repeat. That is exactly how an ODE solver, or gradient descent, works.
Quick check: for $\mathbf{F}(x,y)=(x,\,y)$, what is the arrow at $(2,-1)$? What shape is the field?
$\mathbf{F}(2,-1) = (2,-1)$: the arrow equals the position, so it points straight away from the origin, and longer the further you are. This is a source: everything flows outward from the origin.
Gradient fields and conservative fields awareness
Take any scalar field, such as the height of a landscape. At every point, draw the arrow pointing straight uphill, with a length equal to the steepness. The result is a vector field: the gradient field $\nabla f$. It is the bridge between the two kinds of field. (See Chapter 2.4: the gradient is perpendicular to the level curves and points to steepest ascent.)
Not every vector field is the gradient of something. A rotation field circles around: if it were "uphill" arrows, you could walk with the arrows and climb for ever while coming back to where you started, which is impossible for a real height. A field that is a gradient is called conservative, and the scalar field it comes from is its potential.
Let $f(x,y) = x^2y$. Then $\nabla f = (2xy,\;x^2)$, so $\mathbf{F} = (2xy,\;x^2)$ is a gradient field.
A test. If $\mathbf{F} = (F_x,F_y)=\nabla f$, then $\partial F_x/\partial y = f_{xy}$ and $\partial F_y/\partial x = f_{yx}$, and mixed partials are equal. So a gradient field must pass the test $\;\dfrac{\partial F_x}{\partial y} = \dfrac{\partial F_y}{\partial x}$.
- $\mathbf{F} = (2xy,\;x^2)$: $\partial F_x/\partial y = 2x$ and $\partial F_y/\partial x = 2x$. Equal, so it passes.
- $\mathbf{F} = (-y,\;x)$: $\partial F_x/\partial y = -1$ but $\partial F_y/\partial x = +1$. Not equal, so this rotation field is not a gradient. No potential exists.
A vector field $\mathbf{F}$ is a gradient field (or conservative) if $\mathbf{F} = \nabla f$ for some scalar field $f$, the potential. Then:
- $\mathbf{F}$ points perpendicular to the level curves of $f$ and has size $\|\nabla f\|$.
- In 2D, $\partial F_x/\partial y = \partial F_y/\partial x$ (in any dimension: the Jacobian of $\mathbf{F}$ is symmetric; it is the Hessian of $f$, see Chapter 2.10). On a region with no holes this test is also sufficient.
- Its line integrals do not depend on the path, and it has zero curl.
Why do we need it?
A conservative field is the best-behaved kind: everything about it is packed into one scalar potential. And optimisation is exactly "follow a gradient field downhill".
Where is it used?
Gradient descent (the loss and its field $-\nabla L$), energy-based models (the force is the gradient of an energy), potential energy and gravity in physics, and the score function in diffusion models, which is by definition the gradient of $\log p$.
How is it used?
To check whether a field is a gradient, test $\partial F_x/\partial y = \partial F_y/\partial x$. If it passes, integrate to find the potential. To optimise, follow $-\mathbf{F}$ downhill.
"Conservative" does not mean "constant" or "small". It only means "is the gradient of something". The test $\partial F_x/\partial y = \partial F_y/\partial x$ can fail to be sufficient on a region with a hole (a famous example is the swirl around a point), but for ordinary fields on the whole plane it is enough.
Quick check: is $\mathbf{F}=(3x^2y,\;x^3 + 2y)$ conservative? If so, find a potential.
$\partial F_x/\partial y = 3x^2$ and $\partial F_y/\partial x = 3x^2$: equal, so yes. A potential is $f = x^3y + y^2$: check $\partial f/\partial x = 3x^2y$ ✓ and $\partial f/\partial y = x^3 + 2y$ ✓.
Three old friends, seen as questions about fields
Three tools you already know turn out to be three natural questions about fields. No new maths here, only new names for old ideas.
- Directional derivative. Standing in a scalar field and walking in a chosen direction: how fast does the number change? (Chapter 2.4)
- Jacobian. In a vector field, if I nudge my position, how do the arrows change? (Chapter 2.5)
- Hessian. In a scalar field, how does the gradient field itself change? It is the Jacobian of $\nabla f$. (Chapter 2.10)
Directional derivative. $f=x^2y$ at $(1,2)$ has $\nabla f=(4,1)$. Walk in the direction $\mathbf{u} = (3,4)/5$ (a unit vector). Then $D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u} = (4\cdot3 + 1\cdot4)/5 = 16/5 = 3.2$.
Jacobian of a vector field. $\mathbf{F}=(-y,\,x)$ has $J = \begin{bmatrix}0&-1\\1&0\end{bmatrix}$. Its trace is $0+0=0$ and $J_{21}-J_{12} = 1-(-1)=2$. (These two numbers will be the divergence and curl in a moment.)
Hessian as a Jacobian. $\nabla f = (2xy,\,x^2)$ has Jacobian $\begin{bmatrix}2y&2x\\2x&0\end{bmatrix}$, which at $(1,2)$ is $\begin{bmatrix}4&2\\2&0\end{bmatrix}$: the Hessian of $f$.
| tool | of what | definition |
|---|---|---|
| Directional derivative | scalar field $f$, unit direction $\mathbf{u}$ | $D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u} = \|\nabla f\|\cos\theta$ |
| Jacobian | vector field $\mathbf{F}$ | $J_{ij} = \partial F_i/\partial x_j$ |
| Hessian | scalar field $f$ | $H_{ij} = \partial^2 f/\partial x_i\partial x_j$, i.e. $H = J(\nabla f)$ |
The directional derivative is largest ($=\|\nabla f\|$) when $\mathbf{u}$ points along $\nabla f$, and zero along a level curve.
Why do we need it?
These three are the core of everything else in this chapter: divergence and curl are just two numbers read off the Jacobian of a vector field. Seeing the connections means you only have one idea to remember.
Where is it used?
Directional derivatives: checking a gradient numerically, sharpness measures and line searches. Jacobians: backpropagation and normalising flows. Hessians: curvature, Newton's method and second-order optimisers.
How is it used?
For a step in direction $\mathbf{u}$, compute $\nabla f\cdot\mathbf{u}$. For a vector field, compute its Jacobian once and read divergence (trace) and curl (the antisymmetric part) from it.
Quick check: $\nabla f = (3,4)$ at a point. What is the largest directional derivative, and in which direction?
The largest rate is $\|\nabla f\| = \sqrt{9+16} = 5$, in the direction of the gradient, $\mathbf{u} = (0.6,\,0.8)$. Check: $(3\cdot0.6 + 4\cdot0.8) = 1.8+3.2 = 5$ ✓.
Divergence: net outflow awareness
Imagine a tiny imaginary box sitting in a flowing fluid. Count how much fluid leaves through its walls and how much enters. If more leaves than enters, the box is fed from inside: there is a source (a tap), and a blob of fluid placed there would expand. If more enters than leaves, there is a sink (a drain) and the blob would shrink. If the two balance, the fluid just flows through (we call that incompressible).
The divergence is that net outflow, per unit of area (or volume). It is a scalar field made from a vector field.
Take $\mathbf{F}=(x,\,y)$ and a tiny square of half-size $s$ centred at the origin (side $2s$, area $4s^2$).
- Right wall ($x=s$): outward direction $(1,0)$, $F_x = s$. Outflow $= s\cdot 2s = 2s^2$.
- Left wall ($x=-s$): outward direction $(-1,0)$, $F_x=-s$, so $\mathbf{F}\cdot\mathbf{n} = s$. Outflow $=2s^2$.
- Top and bottom walls: the same, $2s^2$ each.
- Total outflow $=8s^2$. Per unit area: $8s^2/4s^2 = 2$.
And $\partial F_x/\partial x + \partial F_y/\partial y = 1 + 1 = 2$ ✓. The divergence of a source field $(x,y)$ is $+2$ everywhere. For the rotation field $(-y,x)$ it is $0+0=0$: fluid circles around and nothing expands.
The divergence of $\mathbf{F}=(F_1,\dots,F_n)$ is the sum of the matching partial derivatives:
$$\nabla\!\cdot\mathbf{F} \;=\; \sum_{i=1}^n \frac{\partial F_i}{\partial x_i} \;=\; \frac{\partial F_x}{\partial x} + \frac{\partial F_y}{\partial y}\;(+\tfrac{\partial F_z}{\partial z}\text{ in 3D}).$$Where it comes from. For a small box of side $2s$ at $(x,y)$: outflow through the right wall $\approx F_x(x+s,y)\cdot2s$, through the left wall $\approx -F_x(x-s,y)\cdot 2s$. Their sum is $[F_x(x+s,y)-F_x(x-s,y)]\cdot2s\approx \dfrac{\partial F_x}{\partial x}\cdot 2s\cdot 2s$, which is $\dfrac{\partial F_x}{\partial x}\times$ area. The top and bottom walls give the $y$-term. Divide by the area to get the formula.
Note: the divergence is the trace of the Jacobian of $\mathbf{F}$ (the sum of its diagonal entries, see the trace), a scalar, not a vector.
Why do we need it?
It measures, at every point, whether a flow is creating or removing "stuff". That is how we express conservation laws such as "mass is neither made nor destroyed here".
Where is it used?
Fluid flow and electromagnetism (Maxwell's equations). In ML: continuous normalising flows (the log-density of a point changes at the rate minus the divergence of the flow), diffusion-model probability-flow equations, and physics-informed networks that penalise a nonzero divergence.
How is it used?
Take the derivative of each component with respect to its own coordinate and add them. In code, it is the trace of the Jacobian (autodiff, or a randomised trace estimator when the dimension is large).
Quick check: what is the divergence of $\mathbf{F}=(x^2,\;3y)$ at $(2,1)$?
$\partial F_x/\partial x = 2x = 4$ and $\partial F_y/\partial y = 3$, so $\nabla\cdot\mathbf{F} = 4+3 = 7$: a strong source at that point.
Curl: rotation awareness
Dip a tiny paddle wheel into a flowing river. If the water pushes one side of the wheel harder than the other, the wheel spins. How fast and which way it spins tells you the curl of the flow at that spot.
Surprise: a flow can have perfectly straight streamlines and still spin the wheel. In a river that is fast in the middle and slow near the bank, the water on one side of the wheel moves faster than on the other, so it turns. And a flow that goes round in circles can have zero curl, if the speed drops off in just the right way. Curl measures local spin, not "does it go around".
In 2D, $\operatorname{curl}\mathbf{F} = \dfrac{\partial F_y}{\partial x} - \dfrac{\partial F_x}{\partial y}$ (positive = counter-clockwise).
- Rotation $\mathbf{F}=(-y,x)$: $\partial F_y/\partial x = 1$ and $\partial F_x/\partial y = -1$. Curl $=1-(-1) = 2$. The wheel spins counter-clockwise.
- Shear $\mathbf{F}=(y,0)$ (wind blows right, faster higher up): $\partial F_y/\partial x = 0$ and $\partial F_x/\partial y = 1$. Curl $= 0 - 1 = -1$. The top of the wheel is pushed right harder than the bottom, so it spins clockwise, although every arrow points right.
- Source $\mathbf{F}=(x,y)$: $0 - 0 = 0$. No spin.
2D. $\operatorname{curl}\mathbf{F} = \dfrac{\partial F_y}{\partial x} - \dfrac{\partial F_x}{\partial y}$, a number: the local spin. It equals the circulation around a tiny loop divided by the loop's area. A paddle wheel turns at angular speed $\tfrac12\operatorname{curl}\mathbf{F}$.
3D. The curl is a vector:
$$\nabla\times\mathbf{F} = \left(\frac{\partial F_z}{\partial y} - \frac{\partial F_y}{\partial z},\;\; \frac{\partial F_x}{\partial z} - \frac{\partial F_z}{\partial x},\;\; \frac{\partial F_y}{\partial x} - \frac{\partial F_x}{\partial y}\right).$$It points along the axis the paddle wheel turns about (right-hand rule: curl your right fingers the way it spins, and your thumb points along the vector), and its length is the spin rate. The last entry is the 2D curl. A gradient field has zero curl, because mixed partial derivatives are equal ($f_{xy}=f_{yx}$), so every difference above is $0$. This is the same test as in the conservative-field section. In terms of the Jacobian, the curl is read from its antisymmetric part ($J_{21}-J_{12}$ in 2D).
Why do we need it?
It tells us, at each point, how much a flow rotates, and whether a field could be a gradient (no spin) or must contain a swirl.
Where is it used?
Fluid vortices and weather, electromagnetism (Maxwell's equations), computer graphics (curl noise for smoke), and as a diagnostic in ML: learned "force" fields in energy-based or score models should have zero curl if they are real gradients.
How is it used?
Compute $\partial F_y/\partial x - \partial F_x/\partial y$ (2D). If it is zero everywhere, the field is conservative (on a region with no holes). If not, the field has swirl and no scalar potential.
Quick check: what is the 2D curl of $\mathbf{F}=(x^2y,\;xy^2)$ at $(1,2)$?
$\partial F_y/\partial x = y^2 = 4$ and $\partial F_x/\partial y = x^2 = 1$. Curl $=4-1 = 3$ (counter-clockwise spin).
Line integrals: work along a path awareness
Push a cart along a winding path while a wind (a vector field) pushes on it. At each tiny step, only the part of the wind pointing along your step helps or hinders you. Add up (wind along the step) $\times$ (step length) over the whole path. That total is the work, and it is called the line integral of the field along the path.
A natural question: does the total depend on which path you took between the same two points? For some fields it does not. Those are exactly the gradient (conservative) fields.
Go from $A=(0,0)$ to $B=(2,1)$. Compare two paths: the straight line, and "right to $(2,0)$, then up".
Field $\mathbf{F}=(y,\,x)$ (this is $\nabla(xy)$):
- Straight: $\mathbf{r}(t)=(2t,\,t)$, $d\mathbf{r}=(2,1)\,dt$, and $\mathbf{F}=(t,\,2t)$. So $\mathbf{F}\cdot d\mathbf{r} = (2t+2t)\,dt = 4t\,dt$, and $\int_0^1 4t\,dt = 2$.
- Corner path, first leg $(2t,0)$: $\mathbf{F}=(0,2t)$, $d\mathbf{r}=(2,0)\,dt$, $\mathbf{F}\cdot d\mathbf{r}=0$. Second leg $(2,t)$: $\mathbf{F}=(t,2)$, $d\mathbf{r}=(0,1)\,dt$, $\mathbf{F}\cdot d\mathbf{r}=2\,dt$, giving $2$. Total $0+2=2$.
- Same answer, $2$, and it equals $f(B)-f(A) = 2\cdot1 - 0 = 2$ for the potential $f=xy$.
Field $\mathbf{F}=(-y,\,x)$ (rotation): straight path: $\mathbf{F}=(-t,2t)$, $\mathbf{F}\cdot d\mathbf{r}=(-2t+2t)\,dt=0$, so the work is $0$. Corner path: first leg $0$, second leg $\mathbf{F}=(-t,2)$, $d\mathbf{r}=(0,1)\,dt$: work $2$. Different answers: the work depends on the path, so this field is not conservative.
The line integral of $\mathbf{F}$ along a curve $C$ given by $\mathbf{r}(t)$, $a\le t\le b$, is
$$\int_C \mathbf{F}\cdot d\mathbf{r} \;=\; \int_a^b \mathbf{F}(\mathbf{r}(t))\cdot\mathbf{r}'(t)\,dt \;\approx\; \sum_i \mathbf{F}(\mathbf{r}_i)\cdot\Delta\mathbf{r}_i.$$(Add up "field dotted with each small step".) For a gradient field $\mathbf{F}=\nabla f$ it collapses to the fundamental theorem for line integrals:
$$\int_C \nabla f\cdot d\mathbf{r} = f(B) - f(A).$$Why. By the chain rule, $\dfrac{d}{dt}f(\mathbf{r}(t)) = \nabla f\cdot\mathbf{r}'(t)$. So the integral above is $\int_a^b \dfrac{d}{dt}f(\mathbf{r}(t))\,dt = f(\mathbf{r}(b)) - f(\mathbf{r}(a))$, by the ordinary fundamental theorem of calculus. Only the endpoints matter, and the work around any closed loop is $0$.
Why do we need it?
It adds up an effect that varies along a path: work done by a force, flow along a pipe, or the change of a quantity along a trajectory. And it gives a clean test for "is this a gradient?": path independence.
Where is it used?
Physics (work and energy, circulation), and in ML the idea behind path-based attribution methods such as integrated gradients (which integrate the gradient of a model along a straight path from a baseline input to the real input), and thinking of training as a path through weight space.
How is it used?
Parameterise the path, take many small steps, and add $\mathbf{F}\cdot\Delta\mathbf{r}$. If the field is a gradient, skip all that and subtract the potential at the two ends.
Quick check: for $\mathbf{F}=\nabla f$ with $f(x,y)=x^2+y$, find the work from $(0,0)$ to $(1,3)$ along any path.
It depends only on the ends: $f(1,3)-f(0,0) = (1+3) - 0 = 4$. Check the gradient: $\nabla f=(2x,\,1)$; along the straight path $(t,3t)$, $\mathbf{F}\cdot d\mathbf{r} = (2t\cdot1 + 1\cdot3)\,dt$ and $\int_0^1(2t+3)\,dt = 1+3 = 4$ ✓.
Surface integrals and flux awareness
Hold a net in a steady wind. How much air passes through it? That amount is the flux of the wind through the net. It depends on how big the net is, how strong the wind is, and how the net is tilted: face-on to the wind catches the most, edge-on catches none.
Cut the surface into tiny patches. For each, take the part of the wind that goes straight through it (the wind dotted with the patch's unit normal arrow) times the patch's area. Add them all. That sum is a surface integral of the field.
Uniform wind $\mathbf{F} = (0,0,2)$ (straight up, strength 2) through a flat $3\times3$ square lying horizontally. Its normal is $\mathbf{n}=(0,0,1)$, so $\mathbf{F}\cdot\mathbf{n}=2$ and the flux is $2\times 9 = 18$.
Tilt the square so that its normal is $\mathbf{n}=(0,\,0.6,\,0.8)$ (a unit vector, since $0.36+0.64=1$). Now $\mathbf{F}\cdot\mathbf{n} = 1.6$ and the flux drops to $1.6\times9 = 14.4$. Turn it edge-on ($\mathbf{n}\perp\mathbf{F}$) and the flux is $0$.
The flux of $\mathbf{F}$ through a surface $S$ (with a chosen unit normal $\mathbf{n}$) is the surface integral
$$\iint_S \mathbf{F}\cdot\mathbf{n}\;dS \;\approx\; \sum_{\text{patches}} (\mathbf{F}\cdot\mathbf{n})\,\Delta S.$$For a surface that is a graph $z=g(x,y)$ with the normal pointing up, $\mathbf{n}\,dS = (-g_x,\,-g_y,\,1)\,dx\,dy$, so
$$\text{flux} = \iint \big(F_z - F_x\,g_x - F_y\,g_y\big)\,dx\,dy.$$For a closed surface with the normal pointing outward, the flux is the net amount flowing out, and the divergence theorem says this equals the total divergence inside (see the table below).
Why do we need it?
It counts how much of a flow crosses a boundary: how much heat leaves a room, how much fluid goes through a membrane. It is the "per-surface" cousin of divergence.
Where is it used?
Physics (Gauss's law for electric flux, heat and fluid flow). In ML it is rare in everyday work: it shows up in the maths of continuity equations behind normalising flows and diffusion models, and in graphics (rendering integrates light over surfaces).
How is it used?
Describe the surface, find the normal on each patch, dot it with the field and add up area-weighted values (by formula, or numerically on a grid). Or use the divergence theorem to turn a hard surface sum into a volume integral.
Quick check: wind $\mathbf{F}=(1,0,0)$ blows through a flat vertical $2\times 5$ rectangle whose normal is $(1,0,0)$. What is the flux?
$\mathbf{F}\cdot\mathbf{n} = 1$ and the area is $2\cdot5 = 10$, so the flux is $10$. (If the rectangle were turned to face another way, e.g. $\mathbf{n}=(0,1,0)$, the flux would be $0$.)
The big theorems, in plain words awareness
All of these theorems say the same thing in different clothes: to add up a "change" over a region, you only need to look at the boundary. You already know the simplest one. To find the total change of a quantity from $a$ to $b$ you do not need to add up every tiny change in between: just subtract the value at $a$ from the value at $b$. The vector-calculus theorems extend this from a line segment to curves, flat regions, surfaces and solids.
Check Green's theorem on the unit circle with the rotation field $\mathbf{F}=(-y,x)$. Going round the circle $\mathbf{r}=(\cos\phi,\sin\phi)$, the velocity is $(-\sin\phi,\cos\phi)$ and $\mathbf{F}=(-\sin\phi,\cos\phi)$ too. So $\mathbf{F}\cdot d\mathbf{r} = (\sin^2\phi+\cos^2\phi)\,d\phi = d\phi$ and the circulation is $\int_0^{2\pi}d\phi = 2\pi$. On the other side: the curl is $2$ everywhere and the disc has area $\pi$, so $\iint \text{curl}\,dA = 2\pi$. They match.
| Theorem | Formula | In plain words |
|---|---|---|
| Fundamental theorem of calculus (1D, the ancestor) | $\displaystyle\int_a^b f'(x)\,dx = f(b)-f(a)$ | The total change is the sum of the small changes. |
| Fundamental theorem for line integrals | $\displaystyle\int_C \nabla f\cdot d\mathbf{r} = f(B)-f(A)$ | The work of a gradient field depends only on where you start and end. |
| Green's theorem (2D) | $\displaystyle\oint_C \mathbf{F}\cdot d\mathbf{r} = \iint_D \Big(\frac{\partial F_y}{\partial x}-\frac{\partial F_x}{\partial y}\Big)\,dA$ | The circulation around a loop equals the total spin inside it. |
| Stokes' theorem (3D) | $\displaystyle\oint_{\partial S} \mathbf{F}\cdot d\mathbf{r} = \iint_S (\nabla\times\mathbf{F})\cdot\mathbf{n}\,dS$ | The circulation around the rim of a surface equals the total spin passing through it. |
| Divergence theorem (Gauss) | $\displaystyle\oiint_{S}\mathbf{F}\cdot\mathbf{n}\,dS = \iiint_V \nabla\!\cdot\mathbf{F}\,dV$ | The total outflow through a closed surface equals the total source strength inside. |
Conditions (each theorem is only true when they hold):
- Line integrals: $f$ is smooth (continuously differentiable) and $C$ is a smooth curve from $A$ to $B$.
- Green: $C$ is a simple closed curve (it does not cross itself) bounding the region $D$, travelled counter-clockwise (so $D$ is on your left), and $\mathbf{F}$ is smooth on all of $D$, with no holes or blow-ups inside.
- Stokes: $S$ is an oriented surface with edge $\partial S$. The direction round the edge and the choice of normal $\mathbf{n}$ go together by the right-hand rule (curl the fingers of your right hand along the edge, and your thumb points along $\mathbf{n}$). $\mathbf{F}$ is smooth on $S$.
- Divergence: $S$ is the closed surface that bounds the solid $V$, with $\mathbf{n}$ pointing outward, and $\mathbf{F}$ is smooth everywhere inside $V$.
Green's theorem also has a "flux form" that says the same thing for the divergence in 2D: $\oint_C \mathbf{F}\cdot\mathbf{n}\,ds = \iint_D \nabla\cdot\mathbf{F}\,dA$ (the flow out through a loop equals the total divergence inside). The second comparison in the widget below is exactly this.
Why do we need it?
They let us trade a hard sum for an easy one: a sum over a whole region can be replaced by a sum over just its edge (or the reverse). They also connect the local ideas (divergence, curl) to the global ones (flux, circulation).
Where is it used?
Deriving conservation laws in physics, electromagnetism, and fluid mechanics. In ML they sit in the background: the divergence theorem is used to derive the continuity equation (how a density is carried along by a flow), and that equation is the starting point for the theory of continuous normalising flows and diffusion models.
How is it used?
To compute a flux or circulation, check whether replacing it by a divergence or curl integral is easier. Mostly, for ML, it is enough to recognise the names and what they say.
Where these ideas appear in ML, honestly.
- Gradient flow. Gradient descent with very small steps follows the flow lines of the field $-\nabla L$. This is the most direct use of vector fields in ML, and it is fully covered by earlier chapters.
- Normalising flows and neural ODEs. A learned vector field moves points around; for continuous flows the log-density changes at the rate $-\nabla\!\cdot\mathbf{F}$ (minus the trace of the Jacobian). Divergence is therefore a real working quantity there.
- Diffusion and score models. The learned "score" $\nabla_{\mathbf{x}}\log p(\mathbf{x})$ is a gradient field, and sampling moves points along it (plus random noise in the stochastic version). The theory uses the divergence theorem, but practitioners rarely compute it by hand.
- Physics-informed models. Physics-informed neural networks add terms such as "divergence $=0$" to the loss to force a flow to conserve mass.
- Not needed for ordinary supervised learning. Training a classifier or a regressor uses gradients, Jacobians (backprop) and sometimes Hessians. Line and surface integrals almost never appear. That is why this chapter is for awareness.
Quick check: a vector field has divergence $3$ everywhere. How much flows out of a closed surface that encloses a region of volume $2$?
By the divergence theorem, outflow $=\iiint\nabla\cdot\mathbf{F}\,dV = 3\times 2 = 6$.
Recap, cheat sheet and practice
- A scalar field gives a number at every point; a vector field gives an arrow. A flow line follows the arrows: $d\mathbf{x}/dt=\mathbf{F}(\mathbf{x})$.
- $\nabla f$ is a gradient field, perpendicular to the level curves. A field is conservative if it is a gradient; in 2D test $\partial F_x/\partial y=\partial F_y/\partial x$.
- Directional derivative $\nabla f\cdot\mathbf{u}$; Jacobian $\partial F_i/\partial x_j$; Hessian $=$ Jacobian of $\nabla f$.
- Divergence $\nabla\cdot\mathbf{F}=\sum\partial F_i/\partial x_i$ = net outflow = trace of the Jacobian. Curl (2D) $\partial F_y/\partial x-\partial F_x/\partial y$ = spin; in 3D it is a vector along the spin axis.
- Line integral $\int_C\mathbf{F}\cdot d\mathbf{r}$ = work along a path; for $\mathbf{F}=\nabla f$ it equals $f(B)-f(A)$. Flux $\iint\mathbf{F}\cdot\mathbf{n}\,dS$ = how much flows through a surface.
- The big theorems all say "total of a derivative inside = value on the boundary". For ML, this chapter is awareness: the daily tools are gradients, Jacobians, Hessians and Taylor expansions.
Cheat sheet
| Tool | Formula | Picture |
|---|---|---|
| Gradient | $\nabla f = (\partial f/\partial x_i)_i$ | arrow uphill, perpendicular to contours |
| Conservative test (2D) | $\partial F_x/\partial y = \partial F_y/\partial x$ | no swirl, has a potential |
| Divergence | $\partial F_x/\partial x + \partial F_y/\partial y$ | little box expands or shrinks |
| Curl (2D) | $\partial F_y/\partial x - \partial F_x/\partial y$ | paddle wheel spins |
| Line integral | $\int_C \mathbf{F}\cdot d\mathbf{r}$ | work along a path |
| Flux | $\iint_S \mathbf{F}\cdot\mathbf{n}\,dS$ | wind through a net |
| Green / Stokes | circulation $=\iint$ curl | loop = spin inside |
| Divergence theorem | flux out $=\iiint$ div | outflow = sources inside |
import numpy as np
# Numerical divergence and 2D curl of a field F(x, y) -> [Fx, Fy]
def div_curl(F, x, y, h=1e-5):
dFx_dx = (F(x + h, y)[0] - F(x - h, y)[0]) / (2 * h)
dFx_dy = (F(x, y + h)[0] - F(x, y - h)[0]) / (2 * h)
dFy_dx = (F(x + h, y)[1] - F(x - h, y)[1]) / (2 * h)
dFy_dy = (F(x, y + h)[1] - F(x, y - h)[1]) / (2 * h)
return dFx_dx + dFy_dy, dFy_dx - dFx_dy # (divergence, curl)
rot = lambda x, y: np.array([-y, x])
source = lambda x, y: np.array([x, y])
shear = lambda x, y: np.array([y, 0.0])
print(np.round(div_curl(rot, 1.0, 2.0), 4)) # [0. 2.] spins, does not expand
print(np.round(div_curl(source, 1.0, 2.0), 4)) # [2. 0.] expands, does not spin
print(np.round(div_curl(shear, 1.0, 2.0), 4)) # [ 0. -1.] straight flow that still spins
# Line integral of F along a straight segment P -> Q (midpoint rule)
def line_integral(F, P, Q, n=1000):
P, Q = np.array(P, float), np.array(Q, float)
t = (np.arange(n) + 0.5) / n
dr = (Q - P) / n
return sum(F(*(P + ti * (Q - P))) @ dr for ti in t)
def along(F, pts): # a path made of several segments
return sum(line_integral(F, pts[i], pts[i + 1]) for i in range(len(pts) - 1))
grad = lambda x, y: np.array([y, x]) # = gradient of f(x, y) = x*y
A, M, B = (0, 0), (2, 0), (2, 1)
print(along(grad, [A, B]), along(grad, [A, M, B])) # 2.0 2.0 same: path independent (= f(B) - f(A))
print(along(rot, [A, B]), along(rot, [A, M, B])) # 0.0 2.0 different: rot is not conservative
# Green's theorem on the unit circle for the rotation field: circulation = integral of curl
phi = np.linspace(0, 2 * np.pi, 2000, endpoint=False)
dphi = 2 * np.pi / 2000
pts = np.stack([np.cos(phi), np.sin(phi)], axis=1)
tangent = np.stack([-np.sin(phi), np.cos(phi)], axis=1)
circulation = sum(rot(*p) @ t for p, t in zip(pts, tangent)) * dphi
print(circulation, 2 * (np.pi * 1.0**2)) # 6.2831... 6.2831... (curl 2 times area pi)
1. Which of these is a scalar field?
2. Is $\mathbf{F}=(2x+y,\;x+3)$ a gradient field?
3. The divergence of $\mathbf{F}=(3x,\;-y)$ is…
4. A paddle wheel sits in the field $\mathbf{F}=(y,\,0)$ (wind to the right, stronger higher up). It…
5. For $\mathbf{F}=\nabla f$ with $f(x,y)=x^2y$, the line integral along any path from $(0,0)$ to $(2,1)$ is…
6. Where do divergence, curl and surface integrals matter most in everyday ML practice?
Practice problems
A. For $f(x,y) = x^2 + 3y^2$, sketch (describe) the gradient field and find $D_{\mathbf{u}}f$ at $(1,1)$ for $\mathbf{u}=(1,0)$.
$\nabla f = (2x,\;6y)$. At $(1,1)$: $(2,6)$. The arrows point away from the origin (uphill), steeper in $y$, and are perpendicular to the elliptical contours. $D_{\mathbf{u}}f = (2,6)\cdot(1,0) = 2$.
B. Is $\mathbf{F}=(y\cos x,\;\sin x + 2y)$ conservative? Find a potential.
$\partial F_x/\partial y = \cos x$ and $\partial F_y/\partial x = \cos x$: equal, so yes. Integrate $F_x$ with respect to $x$: $f = y\sin x + g(y)$. Then $\partial f/\partial y = \sin x + g'(y)$ must equal $\sin x + 2y$, so $g'(y)=2y$ and $g=y^2$. Potential: $f = y\sin x + y^2$.
C. Compute the divergence and 2D curl of $\mathbf{F}=(x^2-y^2,\;2xy)$ at $(1,1)$.
$\partial F_x/\partial x = 2x = 2$, $\partial F_y/\partial y = 2x = 2$: divergence $4$. $\partial F_y/\partial x = 2y = 2$ and $\partial F_x/\partial y = -2y = -2$: curl $2-(-2)=4$. The field both expands and spins counter-clockwise at that point.
D. Compute $\int_C \mathbf{F}\cdot d\mathbf{r}$ for $\mathbf{F}=(-y,\,x)$ along the straight segment from $(1,0)$ to $(0,1)$.
$\mathbf{r}(t) = (1-t,\;t)$, $d\mathbf{r}=(-1,\,1)\,dt$, $\mathbf{F} = (-t,\;1-t)$. So $\mathbf{F}\cdot d\mathbf{r} = (t + 1 - t)\,dt = dt$, and the integral is $\int_0^1 dt = 1$. (Compare: along the quarter circle between the same points the work is $\pi/2\approx1.57$, since $\mathbf{F}\cdot d\mathbf{r}=d\phi$ there. Different paths, different work: not conservative.)
E. Wind $\mathbf{F}=(0,\,3,\,4)$ crosses a flat square of area $5$ whose unit normal is $\mathbf{n}=(0,\,0.6,\,0.8)$. Find the flux.
$\mathbf{F}\cdot\mathbf{n} = 3\cdot0.6 + 4\cdot0.8 = 1.8 + 3.2 = 5$. Flux $=5\cdot 5 = 25$. (The wind is exactly face-on to the square, since $\mathbf{F}=5\mathbf{n}$.)
F. Use Green's theorem to find the circulation of $\mathbf{F}=(-y,\,x)$ around a circle of radius $3$.
The curl is $2$ everywhere and the disc has area $\pi\cdot3^2 = 9\pi$. So the circulation is $2\cdot9\pi = 18\pi\approx 56.5$. Direct check: on the circle $\mathbf{F}\cdot d\mathbf{r} = 3\cdot3\,d\phi = 9\,d\phi$ (speed 3, radius 3), and $\int_0^{2\pi}9\,d\phi = 18\pi$ ✓.
Calculus → ML Connection
This is where everything comes together. A model learns by one repeated move: measure how wrong it is (the loss), ask the derivative which way is downhill (the gradient), and take a small step that way. In this chapter we derive the gradient of every classic model by hand, link each one to a picture, and finish with the training loop that PyTorch runs.
- See training as "walk downhill on the loss surface", and know exactly what the update rule does
- Derive the gradients of mean squared error, linear regression, the sigmoid, logistic regression, cross-entropy and softmax, step by step
- Understand why cross-entropy trains classifiers well and squared error struggles (saturation)
- Choose learning rates, use momentum, and add L1 / L2 regularisation, knowing what each does to the gradient
- Follow the chain rule through a small neural network and watch one train live
- Explain full-batch, mini-batch and stochastic gradient descent, and write the "forward, loss, backward, update" loop
The big picture: learning is going downhill core
What we need from earlier chapters: the derivative as a slope (Chapter 2.3), the gradient as the direction of steepest ascent (Chapter 2.4) and linearization (Chapter 2.12: close to a point, a curve looks like its tangent line).
A model is a machine with knobs. The knobs are called weights. For every setting of the knobs the model makes some predictions, and we can add up how wrong they are. That total wrongness is the loss.
So the loss is a landscape: one knob gives a curve, two knobs give a surface with hills and valleys, a million knobs give a landscape too big to draw (but the maths is the same). Training means: find a low point of this landscape.
We cannot see the whole landscape. But we can feel the ground under our feet: the slope tells us which way is down. So we take a small step downhill, feel the slope again, and repeat. That is all of machine-learning training. The rest of this chapter is about computing the slope for each kind of model.
Take one knob $w$ and three data points $(x, y) = (1, 2),\ (2, 3),\ (3, 7)$. Our model is $\hat y = w\,x$ ("multiply the input by $w$"). Start with $w = 1$.
- Forward: predictions $\hat y = [1, 2, 3]$.
- Errors (prediction minus truth): $[1-2,\ 2-3,\ 3-7] = [-1, -1, -4]$.
- Loss (average of squared errors): $(1 + 1 + 16)/3 = 6$.
- Slope of the loss at $w = 1$: $\dfrac{dL}{dw} = \dfrac{2}{3}\big(1\cdot(-1) + 2\cdot(-1) + 3\cdot(-4)\big) = \dfrac23(-15) = -10$ (we derive this formula in the next section). Negative slope: the loss goes down when $w$ goes up.
- Update with step size $\eta = 0.1$: $w \leftarrow 1 - 0.1\cdot(-10) = 2$.
The new loss is $(0+1+1)/3$: with $w=2$ the predictions are $[2,4,6]$ and the errors are $[0, 1, -1]$, so $L = 2/3 \approx 0.67$. One step took the loss from 6 down to 0.67.
Four ingredients. A model $f(\mathbf{x};\mathbf{w})$ with weights $\mathbf{w}$. A loss per example $\ell(\hat y, y)$ that is small when the prediction is good. The training loss, the average over the $n$ examples:
$$L(\mathbf{w}) = \frac1n\sum_{i=1}^{n}\ell\big(f(\mathbf{x}_i;\mathbf{w}),\,y_i\big).$$And the gradient descent update, where $\eta > 0$ is the learning rate (step size):
$$\mathbf{w} \leftarrow \mathbf{w} - \eta\,\nabla L(\mathbf{w}).$$The gradient $\nabla L$ is a column vector with one slope $\partial L/\partial w_j$ for each weight. Why does this step go downhill? Linearization (Chapter 2.12) says that for a small step $\boldsymbol{\Delta}$, $L(\mathbf{w}+\boldsymbol{\Delta}) \approx L(\mathbf{w}) + \nabla L^\top\boldsymbol{\Delta}$. Choose $\boldsymbol{\Delta} = -\eta\nabla L$. Then the change is $-\eta\,\nabla L^\top\nabla L = -\eta\,\|\nabla L\|^2 \le 0$. A squared length is never negative, so the loss can only go down (for small enough $\eta$).
Why do we need it?
A model has far too many possible settings to try them all. The slope of the loss tells us, for free, which tiny change to every knob lowers the error. That turns "search everywhere" into "keep walking downhill".
Where is it used?
Training linear and logistic regression, every neural network (CNNs, Transformers, diffusion models), matrix factorisation for recommenders, and almost any model with a differentiable loss.
How is it used?
Repeat four steps: run the model forward to get predictions, compute the loss, compute the gradient of the loss with respect to every weight (backward), and move the weights a small step against the gradient.
"Gradient descent" never promises the best possible answer. It walks to a low point. For the bowl-shaped losses in this chapter's first sections (linear and logistic regression) there is only one low point, so we reach the best answer. For neural networks there are many, and we usually accept a good one (see Chapter 2.10 on curvature and saddle points).
Quick check: at some $w$ the slope is $+6$ and $\eta = 0.1$. Which way does $w$ move, and by how much?
$w \leftarrow w - 0.1\cdot 6 = w - 0.6$. A positive slope means the loss rises as $w$ rises, so we move $w$ left (down), by $0.6$.
Mean squared error: why squared, and what its gradient looks like core
What we need from earlier chapters: the power rule and the chain rule for one variable (Chapters 2.3 and 2.8), and the idea of a length (norm) from the Linear Algebra guide: MSE is a squared length.
For one example, the error is how far the prediction is from the truth: $e = \hat y - y$. To score many examples with one number we cannot just add the errors: a $+3$ and a $-3$ would cancel to zero and look perfect. So we square each error first. Squares are never negative, and they punish big mistakes much more than small ones (an error of 4 costs 16, an error of 1 costs only 1).
There is a second reason, and it is the one calculus cares about: the slope of a squared error is proportional to the error itself. Big mistake, big push. Small mistake, small push. When the model is nearly right the steps shrink by themselves, so we glide into the bottom instead of jumping over it.
Same data as before: $(1,2),\ (2,3),\ (3,7)$ and $\hat y = w\,x$. Let us derive the slope of the loss step by step.
- Loss for one example: $\ell_i = (w x_i - y_i)^2$. Write $u = w x_i - y_i$, so $\ell_i = u^2$.
- Chain rule (Chapter 2.8): $\dfrac{d\ell_i}{dw} = \dfrac{d\ell_i}{du}\cdot\dfrac{du}{dw} = 2u\cdot x_i = 2\,(w x_i - y_i)\,x_i$.
- Average over the three examples: $\dfrac{dL}{dw} = \dfrac{2}{3}\sum_i (w x_i - y_i)\,x_i$. Each term is error × input.
- At $w = 1$ the errors are $-1, -1, -4$, so the sum is $(-1)(1) + (-1)(2) + (-4)(3) = -15$, and $\dfrac{dL}{dw} = \dfrac23(-15) = -10$.
- Numeric check: nudge $w$ by $0.001$. $L(1.001) - L(0.999)$ divided by $0.002$ gives $-10.0000$ (to 4 decimals), which matches.
- Set the slope to zero to find the best $w$: $\sum_i x_i(w x_i - y_i) = 0 \Rightarrow w = \dfrac{\sum x_i y_i}{\sum x_i^2} = \dfrac{29}{14} \approx 2.07$.
The second formula is the key fact: the gradient with respect to each prediction is proportional to that prediction's error. To get the gradient with respect to a weight, we then pass it back through the model with the chain rule: $\dfrac{\partial L}{\partial w_j} = \sum_i \dfrac{\partial L}{\partial \hat y_i}\,\dfrac{\partial \hat y_i}{\partial w_j}$.
- Compare with absolute error $|e|$: its slope is $\pm 1$ whatever the size of the error, with a sharp kink at $e=0$. It treats all mistakes alike and the step does not shrink near the bottom.
- RMSE $=\sqrt{\text{MSE}}$ is in the same units as $y$. It has the same minimum but a different gradient scale.
- Many books write $\tfrac12\text{MSE}$ so the gradient has no "2". PyTorch's
mse_losshas no $\tfrac12$. We keep the 2 here, and say so whenever it matters.
Why do we need it?
We need one number that says how wrong a whole set of predictions is, that never lets errors cancel, and whose slope is smooth so that gradient descent can use it. The squared error does all three.
Where is it used?
The loss for regression: house prices, temperature forecasts, the reconstruction loss of autoencoders, the "value" loss in reinforcement learning, and as a metric for almost every numeric prediction.
How is it used?
Compute the errors, square and average them. For training, take the gradient $\tfrac{2}{n}(\hat y - y)$ at the output and send it backwards. Watch out for outliers: one huge error dominates the sum.
Outliers. One wild data point with error $20$ contributes $400$ to the sum, more than a hundred points with error $2$. MSE bends the fit toward outliers. If that is a problem, use absolute error or the Huber loss (square for small errors, straight line for big ones).
Where the "2" goes. The 2 from the power rule only rescales the gradient (it is the same as doubling $\eta$). It never changes where the minimum is.
Quick check: for one example with $\hat y = 5$ and $y = 8$, what is $\partial(\hat y - y)^2/\partial\hat y$, and which way should $\hat y$ move?
$2(\hat y - y) = 2(5 - 8) = -6$. The slope is negative, so increasing $\hat y$ lowers the loss: move $\hat y$ up toward 8. Note the size, $6$, is twice the error $3$.
Linear regression: derive the gradient, then solve it core
What we need from earlier chapters: the gradient (Chapter 2.4), matrix calculus and the identity $\nabla\|X\mathbf{w}-\mathbf{y}\|^2 = 2X^\top(X\mathbf{w}-\mathbf{y})$ (Chapters 2.6 and 2.7), the Hessian (Chapter 2.10) and, from the Linear Algebra guide, matrix–vector products, the normal equations and linear regression.
Linear regression draws the best straight line (or flat plane) through a cloud of points. The line has two knobs: a slope $w$ and an intercept $b$. Every pair $(b, w)$ is a different line, and each line has a loss (its mean squared error). Plot the loss above the $(b, w)$ plane and you get a smooth bowl. The bottom of the bowl is the best line.
Two ways down: (1) walk with gradient descent, or (2) notice that at the very bottom the ground is flat (the gradient is zero) and solve "gradient $= 0$" directly. That equation is the famous normal equations, the same ones the Linear Algebra guide found from a projection picture. Calculus gives the same answer by a different road.
Data: $(1,2),\ (2,3),\ (3,7)$. To include the intercept, give every input a leading 1. Rows of $X$ are $[1, x_i]$ and the weight vector is $\mathbf{w} = [b, w]^\top$:
$$X = \begin{bmatrix}1&1\\1&2\\1&3\end{bmatrix},\quad \mathbf{y} = \begin{bmatrix}2\\3\\7\end{bmatrix},\quad \mathbf{w}=\begin{bmatrix}0\\1\end{bmatrix}\ (\text{intercept } 0,\ \text{slope } 1).$$- Predictions: $X\mathbf{w} = [1, 2, 3]^\top$. Residual $\mathbf{r} = X\mathbf{w}-\mathbf{y} = [-1, -1, -4]^\top$.
- Loss: $L = \frac13\mathbf{r}^\top\mathbf{r} = \frac13(1+1+16) = 6$.
- $X^\top\mathbf{r} = [\,1(-1)+1(-1)+1(-4),\ \ 1(-1)+2(-1)+3(-4)\,]^\top = [-6, -15]^\top$.
- Gradient $\nabla L = \frac23 X^\top\mathbf{r} = [-4, -10]^\top$. The second entry, $-10$, matches the slope we found in the last section.
- One step with $\eta = 0.1$: $\mathbf{w} \leftarrow [0,1] - 0.1\,[-4,-10] = [0.4,\ 2.0]$. New residual $[0.4, 1.4, -0.6]$, new loss $(0.16+1.96+0.36)/3 \approx 0.83$.
- Zero gradient instead. $X^\top X = \begin{bmatrix}3&6\\6&14\end{bmatrix}$ and $X^\top\mathbf{y} = [12, 29]^\top$. Solving $X^\top X\mathbf{w} = X^\top\mathbf{y}$ gives $\mathbf{w}^\star = [-1,\ 2.5]^\top$: the line $\hat y = -1 + 2.5x$. Its residual is $[-0.5, 1, -0.5]$, and $X^\top\mathbf{r} = [0, 0]^\top$ (the residual adds to zero and is perpendicular to the $x$ column). The loss is $0.5$.
Model $\hat{\mathbf{y}} = X\mathbf{w}$ ($X$ is $n\times d$, one row per example). Loss (the mean squared error, a squared length divided by $n$):
$$L(\mathbf{w}) = \frac1n\|X\mathbf{w}-\mathbf{y}\|^2 = \frac1n\,\mathbf{r}^\top\mathbf{r},\qquad \mathbf{r} = X\mathbf{w}-\mathbf{y}.$$Derivation, one weight at a time. $L = \frac1n\sum_i r_i^2$ and $r_i = \sum_j X_{ij}w_j - y_i$, so $\partial r_i/\partial w_j = X_{ij}$. The chain rule gives
$$\frac{\partial L}{\partial w_j} = \frac1n\sum_i 2\,r_i\,\frac{\partial r_i}{\partial w_j} = \frac2n\sum_i X_{ij}\,r_i = \frac2n\,(X^\top\mathbf{r})_j.$$Stack the $d$ partial derivatives into a column (the gradient is a column vector):
$$\boxed{\ \nabla L(\mathbf{w}) = \frac2n X^\top(X\mathbf{w}-\mathbf{y})\ }\qquad\text{(shape check: } d\times n \text{ times } n\times 1 = d\times 1\text{)}$$Set it to zero: $X^\top X\mathbf{w} = X^\top\mathbf{y}$ (the normal equations), so $\mathbf{w}^\star = (X^\top X)^{-1}X^\top\mathbf{y}$ when $X^\top X$ is invertible. Read $X^\top\mathbf{r} = \mathbf{0}$ as a picture: the residual is perpendicular to every column of $X$, which is exactly the projection picture of Chapter 1.10.
Why the bottom is the bottom. The Hessian is $\frac2n X^\top X$, which is positive semi-definite (it never curves downward), so $L$ is a bowl: every flat point is a lowest point (and there is exactly one when $X^\top X$ is invertible). For gradient descent the update is $\mathbf{w} \leftarrow \mathbf{w} - \eta\,\frac2n X^\top(X\mathbf{w}-\mathbf{y})$: "move the weights against (the inputs weighted by the errors)".
Why do we need it?
We want the best line, plane or hyperplane through data. Calculus gives a recipe that works for any number of features: take the gradient of the loss, then either walk down it or set it to zero.
Where is it used?
Predicting prices and demand, calibrating sensors, the last layer of many networks (a linear layer fitted by MSE), the "baseline" model for almost any regression problem, and the building block of ridge regression.
How is it used?
For small $d$, solve the normal equations (np.linalg.solve or lstsq). For huge data, run gradient descent with the gradient $\frac2n X^\top(X\mathbf{w}-\mathbf{y})$. Always check: after fitting, $X^\top\mathbf{r}$ should be almost zero.
Gradient zero ≠ always one answer. If two columns of $X$ are copies of each other, $X^\top X$ is singular and the bowl has a flat-bottomed trough: infinitely many best lines. Gradient descent still works (it lands somewhere on the trough). Ridge regularisation (below) fixes it.
Inverting is not how you solve it. Libraries solve $X^\top X\mathbf{w} = X^\top\mathbf{y}$ (or use QR / SVD) instead of forming $(X^\top X)^{-1}$: it is faster and more accurate. See Chapter 1.10.
Quick check: with one feature and no intercept, show that $\nabla L = 0$ gives $w = \sum x_iy_i/\sum x_i^2$.
Here $X$ is a single column, so $X^\top X = \sum x_i^2$ and $X^\top\mathbf{y} = \sum x_iy_i$. The normal equation $\big(\sum x_i^2\big)w = \sum x_iy_i$ gives the formula. With our data: $29/14 \approx 2.07$.
The sigmoid: turning a score into a probability core
What we need from earlier chapters: the sigmoid function and exponentials (Chapter 2.1), and the chain rule and derivative of $e^x$ (Chapters 2.3 and 2.8).
A model that predicts a yes/no outcome first computes a score $z$ that can be any number, positive or negative. But a probability must sit between 0 and 1. The sigmoid is a smooth S-shaped curve that squeezes every score into that range: a very negative score becomes almost 0, a very positive score becomes almost 1, and a score of exactly 0 becomes $\tfrac12$ ("no idea").
The slope of the S tells us how much the probability reacts to a small change in the score. In the middle the curve is steep: the model is unsure, so evidence matters. At the far ends the curve is almost flat: the model is already confident, and extra evidence barely moves it. That flatness is called saturation, and it will matter a lot in this chapter.
Let $z = 2$. Then $\sigma(2) = \dfrac{1}{1+e^{-2}} = \dfrac{1}{1+0.1353} \approx 0.8808$.
Derive the slope. Write $\sigma(z) = (1+e^{-z})^{-1}$.
- Chain rule with outer power $u^{-1}$ and inner $u = 1+e^{-z}$: $\sigma'(z) = -(1+e^{-z})^{-2}\cdot\dfrac{d}{dz}(1+e^{-z}) = -(1+e^{-z})^{-2}\cdot(-e^{-z}) = \dfrac{e^{-z}}{(1+e^{-z})^2}$.
- Split the fraction into two factors: $\dfrac{e^{-z}}{(1+e^{-z})^2} = \dfrac{1}{1+e^{-z}}\cdot\dfrac{e^{-z}}{1+e^{-z}}$.
- The first factor is $\sigma(z)$. The second is $1 - \sigma(z)$, because $1 - \dfrac{1}{1+e^{-z}} = \dfrac{e^{-z}}{1+e^{-z}}$.
- So $\boxed{\sigma'(z) = \sigma(z)\,\big(1-\sigma(z)\big)}$. The slope can be computed from the output alone.
- At $z=2$: $\sigma' = 0.8808\times 0.1192 \approx 0.1050$. Numeric check: $\big(\sigma(2.001)-\sigma(1.999)\big)/0.002 = 0.1050$ ✓.
- Output is in $(0, 1)$; $\sigma(0) = \tfrac12$.
- The slope is biggest at $z=0$, where $\sigma' = \tfrac12\cdot\tfrac12 = \tfrac14$, and it is never larger than $\tfrac14$. It tends to 0 as $z\to\pm\infty$ (saturation).
- The inverse is the logit (log-odds): $z = \ln\dfrac{p}{1-p}$.
- For a vector of scores, apply $\sigma$ to each entry; the Jacobian is diagonal, $\mathrm{diag}(\sigma(1-\sigma))$ (Chapter 2.5).
Why do we need it?
Models compute unbounded scores, but probabilities live between 0 and 1. The sigmoid is the smooth bridge, and its simple derivative $\sigma(1-\sigma)$ keeps backpropagation cheap.
Where is it used?
The output of logistic regression and of binary classifiers, gates inside LSTMs and GRUs, "is this token a yes?" heads, and (historically) hidden layers of early neural networks.
How is it used?
Compute $p=\sigma(z)$ and predict "yes" when $p\ge\frac12$ (that is, $z\ge0$). In backprop, multiply the gradient arriving at $p$ by $p(1-p)$ to get the gradient at $z$. Beware: when $p$ is near 0 or 1 this factor is tiny.
Vanishing gradients. In a deep network with sigmoids, the backward pass multiplies by $\sigma'\le0.25$ at every layer. The product shrinks fast, so early layers barely learn. This is one reason ReLU replaced the sigmoid in hidden layers (see vanishing gradients in Chapter 2.9). The sigmoid is still the right choice for a final yes/no probability.
Cousin. $\tanh z = 2\sigma(2z) - 1$ has range $(-1,1)$ and derivative $1-\tanh^2 z$ (maximum 1, so less shrinking).
Quick check: what is $\sigma'(0)$, and what is $\sigma'(5)$ roughly?
$\sigma'(0) = 0.5\times0.5 = 0.25$. At $z=5$, $\sigma\approx0.9933$, so $\sigma' \approx 0.9933\times0.0067 \approx 0.0066$: about 38 times smaller.
Logistic regression: the gradient and its beautiful cancellation core
What we need from earlier chapters: the chain rule (Chapter 2.8), the sigmoid and $\sigma'=\sigma(1-\sigma)$ (the previous section), the gradient of a dot product (Chapter 2.6) and, from the Linear Algebra guide, logistic regression and its decision boundary.
Now the answer is yes or no: is this email spam? We compute a score $z = \mathbf{w}^\top\mathbf{x}$ (a weighted sum of the features, with the bias folded in as a weight on a constant 1), squeeze it into a probability $p = \sigma(z)$, and predict spam when $p \ge \frac12$. The place where $z = 0$ is the decision boundary: a line in 2D.
How do we score a guess? By the probability the model gave to the true answer. If the email really is spam, a good model gave $p$ close to 1. We turn that into a loss by taking $-\ln$ of it: $-\ln(\text{probability of the truth})$ is near 0 when the model was right and confident, and huge when it was confident and wrong. (The next section shows where this loss comes from.)
The gradient has a lovely surprise: after the chain rule, almost everything cancels, and what is left is "prediction minus truth, times the input", the same shape as in linear regression.
Two examples with one feature plus a bias: $x = -1$ has label $y=0$ and $x = 1$ has label $y=1$. So
$$X=\begin{bmatrix}1&-1\\1&1\end{bmatrix},\quad \mathbf{y}=\begin{bmatrix}0\\1\end{bmatrix},\quad \mathbf{w} = \begin{bmatrix}0\\0\end{bmatrix}\ (\text{bias},\ \text{weight}).$$- Scores $X\mathbf{w} = [0, 0]$, probabilities $\sigma(0) = [0.5, 0.5]$.
- Loss: $-\ln(1-0.5)$ for the first, $-\ln 0.5$ for the second; the mean is $\ln 2 \approx 0.693$.
- Prediction minus truth: $\mathbf{p}-\mathbf{y} = [0.5,\ -0.5]$.
- $X^\top(\mathbf{p}-\mathbf{y}) = [\,0.5 - 0.5,\ \ (-1)(0.5) + (1)(-0.5)\,] = [0, -1]$. Divide by $n=2$: $\nabla L = [0, -0.5]$.
- Step with $\eta = 1$: $\mathbf{w}\leftarrow [0, 0.5]$. New probabilities $[\sigma(-0.5), \sigma(0.5)] = [0.378, 0.622]$, new loss $\approx 0.474$. The model now leans the right way for both points.
Model and loss for one example (label $y\in\{0,1\}$):
$$z = \mathbf{w}^\top\mathbf{x},\qquad p = \sigma(z),\qquad \ell = -\,y\ln p-(1-y)\ln(1-p).$$Derivation, four small derivatives multiplied (the chain rule):
- Loss with respect to $p$: $\dfrac{\partial\ell}{\partial p} = -\dfrac{y}{p}+\dfrac{1-y}{1-p} = \dfrac{-y(1-p) + (1-y)p}{p(1-p)} = \dfrac{p-y}{p(1-p)}$.
- Probability with respect to score: $\dfrac{\partial p}{\partial z} = \sigma'(z) = p(1-p)$.
- Score with respect to weights: $\dfrac{\partial z}{\partial\mathbf{w}} = \mathbf{x}$.
- Multiply: $\dfrac{\partial\ell}{\partial\mathbf{w}} = \dfrac{p-y}{\cancel{p(1-p)}}\cdot\cancel{p(1-p)}\cdot\mathbf{x} = (p-y)\,\mathbf{x}$. The awkward $p(1-p)$ cancels exactly.
Average over $n$ examples and stack into matrix form (the gradient is a column):
$$\boxed{\ \nabla L(\mathbf{w}) = \frac1n X^\top\big(\sigma(X\mathbf{w}) - \mathbf{y}\big)\ }$$- Compare linear regression: $\frac2n X^\top(X\mathbf{w}-\mathbf{y})$. Same pattern: inputs times (prediction − target). Only the prediction changed from $X\mathbf{w}$ to $\sigma(X\mathbf{w})$.
- The Hessian is $\frac1n X^\top\,\mathrm{diag}\big(p_i(1-p_i)\big)\,X$, positive semi-definite, so the loss is a convex bowl: one valley, no false minima.
- Setting the gradient to zero gives equations with $\sigma$ inside them. There is no closed form, so we use gradient descent (or Newton's method, Chapter 2.10).
Why do we need it?
Many problems are yes/no, and a straight line is a bad fit for 0/1 answers. Logistic regression gives a probability and a clean gradient, so a classifier can be trained with plain gradient descent.
Where is it used?
Spam filters, credit and medical risk scores, click-through prediction in online ads, and the final layer of every binary neural-network classifier (where the same $\hat p - y$ appears in backprop).
How is it used?
Compute $\mathbf{p} = \sigma(X\mathbf{w})$, then the gradient $\frac1nX^\top(\mathbf{p}-\mathbf{y})$, then step. Watch the loss go down and the boundary rotate. Scikit-learn and PyTorch do exactly this internally.
Perfectly separable data. If a line can split the classes with no mistakes, the loss keeps shrinking as the weights grow (the model gets ever more confident), so the "best" weights are infinite. In practice we stop early or add a regulariser (see the regularisation section).
Forgetting the bias. Without the constant-1 column the boundary must pass through the origin.
Quick check: one example $\mathbf{x} = [1, 3]$ (bias, feature) with $y=1$ and current $p = 0.8$. What is its gradient $\partial\ell/\partial\mathbf{w}$?
$(p - y)\,\mathbf{x} = (0.8-1)[1, 3] = [-0.2, -0.6]$. Both entries are negative, so gradient descent will increase both weights, raising the score and so the probability of class 1.
Cross-entropy: where it comes from and why it beats MSE for classifying core
What we need from earlier chapters: the logarithm and its properties (Chapter 2.1: $\ln(ab)=\ln a+\ln b$), the derivative of $\ln x$ (Chapter 2.3, $1/x$), and the logistic-regression gradient from the previous section.
Think of the loss as surprise. If the model says "99% sure it is a cat" and it is a cat, you are not surprised: tiny loss. If it says "1% cat" and it is a cat, you are very surprised: huge loss. The right way to measure surprise is $-\ln(\text{probability given to what really happened})$.
Why this and not squared error? Squared error says a wrong answer costs at most $1$ (since $p$ is between 0 and 1). So a confidently wrong model gets only a mild loss and a tiny gradient, and it stays stuck. Cross-entropy says a confidently wrong answer costs a lot, and gives a big gradient to fix it.
From likelihood to loss. Three predictions $p = [0.9, 0.2, 0.6]$ (probability of class 1) with true labels $y = [1, 0, 1]$. The probability the model gave to what actually happened: $0.9,\ 1-0.2 = 0.8,\ 0.6$.
- Likelihood (probability of all three outcomes, if examples are independent): $0.9\times0.8\times0.6 = 0.432$. Training should maximise this.
- Take the log, so the product becomes a sum (sums are easier to differentiate, and do not underflow): $\ln 0.432 = \ln0.9+\ln0.8+\ln0.6 = -0.105-0.223-0.511 = -0.839$.
- Flip the sign (we minimise losses) and average: $\dfrac{0.105+0.223+0.511}{3} \approx 0.280$. That is the binary cross-entropy.
Punishing confident mistakes (true label is 1; the model says $p$ for class 1):
| $p$ given to the truth | 0.9 | 0.5 | 0.1 | 0.01 | 0.001 |
|---|---|---|---|---|---|
| cross-entropy $-\ln p$ | 0.105 | 0.693 | 2.303 | 4.605 | 6.908 |
| squared error $(1-p)^2$ | 0.01 | 0.25 | 0.81 | 0.980 | 0.998 |
Cross-entropy keeps growing without limit. Squared error flattens out near 1.
Derivation. A model that outputs $p_i = P(y_i=1\mid\mathbf{x}_i)$ assigns probability $p_i^{\,y_i}(1-p_i)^{1-y_i}$ to the label that happened (for $y=1$ this is $p$, for $y=0$ it is $1-p$). The likelihood of all the data is the product of these. Taking $\ln$ (which keeps the same maximiser, because $\ln$ only goes up) and flipping the sign gives
$$L = -\frac1n\sum_i\Big[y_i\ln p_i + (1-y_i)\ln(1-p_i)\Big]\quad(\text{binary cross-entropy}),\qquad \ell = -\sum_{k}y_k\ln p_k\quad(K\text{ classes, one-hot } \mathbf{y}).$$("One-hot" means $\mathbf{y}$ has a 1 for the true class and 0 for every other class, for example $[0,1,0]$.)
For one-hot labels the $K$-class loss is just $-\ln p_{\text{true class}}$. (Name: $-\sum_k q_k\ln p_k$ is the cross-entropy of the model's distribution $\mathbf{p}$ against the true distribution $\mathbf{q}$. Minimising it makes $\mathbf{p}$ match $\mathbf{q}$.)
A neat form in terms of the score $z$ (this is what BCEWithLogitsLoss computes, and it is numerically safe): since $-\ln\sigma(z) = \ln(1+e^{-z})$ and $-\ln(1-\sigma(z)) = \ln(1+e^{z})$,
Gradient sizes for $y = 1$. Squared error through a sigmoid: $\ell=(p-y)^2$, so $\dfrac{\partial\ell}{\partial z} = 2(p-y)\cdot\sigma'(z) = 2(p-y)\,p(1-p)$. The extra factor $p(1-p)$ is the culprit: it is almost 0 whenever the model is very sure, including when it is very wrong. Cross-entropy's gradient is just $p-y$, with no such factor.
| score $z$ (true label 1) | $-6$ (very wrong) | $-3$ | $0$ | $3$ | $6$ (very right) |
|---|---|---|---|---|---|
| cross-entropy gradient $p-1$ | $-0.9975$ | $-0.9526$ | $-0.5$ | $-0.0474$ | $-0.0025$ |
| MSE gradient $2(p-1)p(1-p)$ | $-0.0049$ | $-0.0861$ | $-0.25$ | $-0.0043$ | $-0.00001$ |
At $z=-6$ the cross-entropy push is about 200 times bigger. MSE saturates: it gives almost no signal exactly where a fix is most needed.
Why do we need it?
A classifier outputs probabilities, so we need a loss that measures how well probabilities match reality, and that pushes hard on confident mistakes. The negative log-likelihood does both and has the clean gradient $p-y$.
Where is it used?
Almost every classifier: logistic regression, image classifiers, and language models (next-token prediction is cross-entropy over the vocabulary). "Perplexity" is $e^{\text{cross-entropy}}$.
How is it used?
Use the built-in cross_entropy / BCEWithLogitsLoss on raw scores (logits), not on probabilities you computed yourself: it is faster and avoids $\ln 0$. Report the average loss per example.
Never feed probabilities you computed with sigmoid/softmax into a log yourself. If $p$ rounds to exactly 0 you get $\ln 0 = -\infty$. Use the combined "with logits" functions, which use the stable form $\ln(1+e^z)-yz$.
MSE is not "wrong" for regression, where the output is not squeezed by a sigmoid. The saturation problem is specific to bounded outputs.
Quick check: the true class gets probability $0.25$. What is the cross-entropy, and what is it if the model had given $0.5$?
$-\ln0.25 = \ln4 \approx 1.386$. With $0.5$: $-\ln 0.5 = \ln 2 \approx 0.693$. Halving the probability doubled the loss here.
Gradient descent: convergence, the learning rate, and momentum core
What we need from earlier chapters: the gradient (Chapter 2.4), the second derivative and the Hessian as curvature (Chapter 2.10), linearization (Chapter 2.12) and, from the Linear Algebra guide, eigenvalues and gradient descent on the loss bowl.
Imagine walking down a foggy hill. You cannot see the valley, but you can feel the slope under your boots. The rule is: step in the direction that feels steepest downhill, then feel again. The only choice is how long each step is. Tiny steps are safe but you will be there all day. Giant steps are fast, but you may jump clean across the valley and land higher up the other side, or even fly off the map.
On a bowl-shaped valley the "just right" step lands you near the bottom quickly. The steepness of the bowl (its curvature, from the second derivative) decides what "too big" means: a steep, narrow bowl needs small steps. That is why the learning rate is the most important setting in training.
The simplest bowl: $L(w) = w^2$, with slope $L'(w) = 2w$ and minimum at $w=0$. Start at $w_0 = 4$. The update is $w \leftarrow w - \eta\cdot 2w = (1-2\eta)\,w$.
| $\eta$ | multiplier $1-2\eta$ | $w_0,\ w_1,\ w_2,\ w_3$ | what happens |
|---|---|---|---|
| 0.1 | 0.8 | 4, 3.2, 2.56, 2.05 | smooth but slow |
| 0.5 | 0 | 4, 0, 0, 0 | perfect: lands on the minimum in one step |
| 0.9 | $-0.8$ | 4, $-3.2$, 2.56, $-2.05$ | jumps over the bottom each time, but shrinks |
| 1.0 | $-1$ | 4, $-4$, 4, $-4$ | bounces forever, no progress |
| 1.1 | $-1.2$ | 4, $-4.8$, 5.76, $-6.91$ | diverges: each step is bigger than the last |
Everything is decided by that one multiplier.
Update rule: $\ \mathbf{w}_{t+1} = \mathbf{w}_t - \eta\,\nabla L(\mathbf{w}_t)$.
Convergence on a quadratic (derived). Take $L(w) = \tfrac12 a\,(w-w^\star)^2$ with curvature $a>0$ (the second derivative). Then $L'(w) = a(w-w^\star)$ and
$$w_{t+1}-w^\star = w_t - w^\star - \eta\,a\,(w_t - w^\star) = (1-\eta a)\,(w_t-w^\star)\ \Longrightarrow\ w_t - w^\star = (1-\eta a)^t\,(w_0-w^\star).$$So the error shrinks to 0 exactly when $|1-\eta a| < 1$, that is, $0 < \eta < 2/a$. The best step is $\eta = 1/a$ (multiplier 0). A multiplier in $(0,1)$ gives a smooth approach, in $(-1,0)$ a zig-zag that still converges, exactly $-1$ bounces forever, and below $-1$ a blow-up.
Many weights. For $L=\tfrac12(\mathbf{w}-\mathbf{w}^\star)^\top H(\mathbf{w}-\mathbf{w}^\star)$ with Hessian $H$, the same argument works along each eigenvector of $H$ with $a=\lambda_i$. Stability needs $\eta < 2/\lambda_{\max}$, but the slowest direction (smallest $\lambda_{\min}$) shrinks by only $1-\eta\lambda_{\min}$ per step. The ratio $\kappa=\lambda_{\max}/\lambda_{\min}$ (the condition number) says how long and thin the valley is, and so how slow descent will be.
Momentum (awareness: a standard extra). Keep a "velocity" $\mathbf{v}$ that remembers past gradients:
$$\mathbf{v}\leftarrow\beta\,\mathbf{v}+\nabla L(\mathbf{w}),\qquad \mathbf{w}\leftarrow\mathbf{w}-\eta\,\mathbf{v},\qquad \beta\in[0,1)\ (\text{often }0.9).$$Unrolled, $\mathbf{v}_t = \mathbf{g}_t+\beta\mathbf{g}_{t-1}+\beta^2\mathbf{g}_{t-2}+\dots$, a running average of recent gradients. Where gradients agree (down the valley floor) they add up and speed us along: a steady gradient $\mathbf{g}$ gives $\mathbf{v}\to\mathbf{g}/(1-\beta)$, ten times larger for $\beta=0.9$. Where they flip sign (across the valley) they cancel and the zig-zag fades. Related ideas: Nesterov momentum (look ahead before taking the gradient) and Adam (also rescales each weight by a running size of its gradient).
Why do we need it?
For almost every model there is no formula for the best weights. Gradient descent only needs the slope, so it works for any smooth loss. But it needs a good step size, or it crawls or explodes.
Where is it used?
Training every neural network. Plain gradient descent for convex models (logistic regression), SGD with momentum for image models, Adam and AdamW for Transformers and language models.
How is it used?
Try learning rates on a log scale (0.1, 0.01, 0.001) and watch the loss curve. If the loss rises or jumps wildly, $\eta$ is too big. If it creeps, it is too small. Add momentum 0.9, and decay $\eta$ later in training.
The best learning rate changes during training. Early on a large step makes quick progress. Near the bottom the same step makes you bounce around the minimum, so many training runs decay the learning rate (see the batch-optimisation section).
Too large a learning rate shows up as a loss that goes UP or turns into NaN. That is the "diverges" row of the table, not a bug in your gradients.
Quick check: $L(w)=3(w-1)^2$. Which learning rates converge, and which is perfect?
Here $L'(w)=6(w-1)$, so the curvature is $a=6$ and the multiplier is $1-6\eta$. It converges when $|1-6\eta|<1$, i.e. $0<\eta<2/6 \approx 0.333$. The perfect step is $\eta=1/6\approx0.167$.
Regularisation: how a penalty changes the gradient core
What we need from earlier chapters: the derivative of $x^2$ and of $|x|$ (and its kink) (Chapters 2.2 and 2.3), the gradient and level sets (Chapter 2.4) and, from the Linear Algebra guide, the L1 and L2 norms and ridge and lasso.
A model with a lot of freedom can fit the training data perfectly by using huge weights that cancel each other out, and then fail on new data. Regularisation adds a price for large weights to the loss: "fit the data, but keep the weights small if you can".
In the gradient this price is a constant pull toward zero, in addition to the usual pull toward fitting the data. There are two flavours, and the difference is in the shape of the pull:
- L2 (ridge): the pull is proportional to the weight. A big weight is pulled hard, a small weight only gently. So weights shrink toward zero but almost never reach it.
- L1 (lasso): the pull has a fixed size, whatever the weight. Even a tiny weight keeps getting pushed, so it can be driven all the way to exactly zero. Those zero weights switch features off: a sparse model.
L2 as "weight decay". One weight $w = 2$, $\lambda = 0.1$, $\eta = 0.1$, and the data gradient is $\nabla L = 0.5$. The penalty $\lambda w^2$ has slope $2\lambda w = 2(0.1)(2) = 0.4$.
- Total gradient: $0.5 + 0.4 = 0.9$.
- Update: $w \leftarrow 2 - 0.1\times0.9 = 1.91$.
- Same thing, rearranged: $w\leftarrow(1-2\eta\lambda)\,w - \eta\nabla L = 0.98\times 2 - 0.05 = 1.91$ ✓. Each step first shrinks the weight by 2% ("decay"), then applies the data gradient.
L1 in one dimension. Minimise $\tfrac12(w-a)^2+\lambda|w|$, where $a$ is the weight the data alone would choose.
- If $w>0$, the slope is $(w-a)+\lambda$. Setting it to 0 gives $w = a-\lambda$, which is valid only if $a>\lambda$.
- If $w<0$, the slope is $(w-a)-\lambda$, giving $w=a+\lambda$, valid only if $a<-\lambda$.
- Otherwise the minimum sits at the kink $w=0$: just left of 0 the slope is $-a-\lambda\le0$ and just right it is $-a+\lambda\ge0$ when $|a|\le\lambda$.
So $w^\star = \operatorname{sign}(a)\max(|a|-\lambda,\,0)$ (the soft threshold). With $\lambda=0.5$: $a=0.8\Rightarrow w^\star=0.3$, and $a=0.4\Rightarrow w^\star=0$ exactly. Ridge with the same $a=0.8$, minimising $\tfrac12(w-a)^2+\lambda w^2$: $w^\star = a/(1+2\lambda) = 0.4$, never exactly 0.
- Ridge gradient step: $\mathbf{w}\leftarrow(1-2\eta\lambda)\mathbf{w}-\eta\nabla L$, which is why L2 regularisation is called weight decay.
- Ridge for linear regression, closed form: from $\frac2nX^\top(X\mathbf{w}-\mathbf{y})+2\lambda\mathbf{w}=\mathbf{0}$ we get $(X^\top X+n\lambda I)\mathbf{w}=X^\top\mathbf{y}$. The matrix $X^\top X+n\lambda I$ is always invertible (every eigenvalue is at least $n\lambda>0$), which cures the flat trough of copied features. See Chapter 1.10.
- Lasso: $|w|$ has a kink at 0, where the derivative does not exist (Chapter 2.2: the left and right slopes, $-1$ and $+1$, disagree). Any slope in $[-1,1]$ is allowed there (a subgradient). A weight sits at exactly 0 when the data's pull on it is smaller than $\lambda$.
- The picture. Minimising $L$ subject to $\|\mathbf{w}\|\le t$ is (for a matching $t$) the same as the penalised problem. The loss contours (ellipses) grow until they first touch the constraint region. A round L2 ball is touched at a generic point; an L1 diamond has corners on the axes, and the ellipses usually hit a corner first, where a weight is exactly 0.
Usually the bias is not penalised. (Awareness: L2 equals a Gaussian prior on the weights, L1 a Laplace prior; AdamW applies the decay separately from the gradient.)
Why do we need it?
With few examples, many features, or copied features, the plain loss has wild or non-unique solutions that fit noise. A penalty keeps the weights small and the answer stable, which usually generalises better.
Where is it used?
Ridge regression, lasso for choosing features (genomics, finance), weight decay in nearly every deep-learning optimiser (SGD, AdamW), and the L2 term in logistic regression (C in scikit-learn).
How is it used?
Add $\lambda\|\mathbf{w}\|^2$ or $\lambda\|\mathbf{w}\|_1$ to the loss (or set weight_decay in the optimiser). Pick $\lambda$ by checking performance on held-out data. Larger $\lambda$ means smaller weights and a simpler model.
Standardise features before regularising. The penalty treats all weights alike. If one feature is measured in thousands, its weight is tiny and escapes the penalty, so features on different scales get unfair treatment.
L1 and gradient descent. Plain gradient descent on an L1 loss chatters around 0 because the slope flips sign at the kink. Real lasso solvers use soft-thresholding (the formula derived above) after every step, an idea called the proximal step.
Quick check: $w=-3$, $\lambda=0.05$, $\eta=0.2$, data gradient $\nabla L=1$. One L2 step?
Penalty slope $2\lambda w = 2(0.05)(-3) = -0.3$, so the total gradient is $1-0.3 = 0.7$ and $w\leftarrow -3 - 0.2(0.7) = -3.14$. Without the penalty the step would have gone to $-3.2$. The penalty held the weight back, so it ended closer to zero. Check with the decay form: $1-2\eta\lambda = 0.98$, and $0.98\cdot(-3) - 0.2\cdot 1 = -3.14$ ✓.
Softmax and the gradient $\mathbf{p}-\mathbf{y}$ core
What we need from earlier chapters: the softmax function (Chapter 2.1), the Jacobian matrix (Chapter 2.5), the chain rule with vectors (Chapter 2.8) and, from the Linear Algebra guide, softmax regression and the dot product.
With more than two classes (cat, dog, bird) the model produces one score per class, called logits. Softmax turns the list of scores into a list of probabilities that are all positive and add to 1. It does this in two moves: exponentiate every score (this makes them positive and exaggerates the winner), then divide by the total (so they sum to 1).
Because the probabilities all share one total, they compete: raising one class's score must lower the others' probabilities. That competition shows up as the minus signs in the derivative. And once again, when we combine softmax with the cross-entropy loss, the messy parts cancel and the final gradient is as simple as it can be: predicted probabilities minus the true answer.
Logits $\mathbf{z} = [2, 1, 0]$, and the true class is class 2 (so $\mathbf{y} = [0, 1, 0]$).
- Exponentiate: $[e^2, e^1, e^0] = [7.389,\ 2.718,\ 1]$. Their total is $11.107$.
- Divide: $\mathbf{p} = [0.665,\ 0.245,\ 0.090]$ (these add to 1).
- Loss: $-\ln p_2 = -\ln0.245 \approx 1.408$.
- Gradient with respect to the logits: $\mathbf{p}-\mathbf{y} = [0.665,\ -0.755,\ 0.090]$. (Derived below.) Gradient descent will raise the true class's logit (negative entry) and lower the others.
- Numeric check: nudge $z_2$ by $0.001$ and recompute the loss: the change per unit is $-0.7553$ ✓.
Step 1: the Jacobian. Take logs: $\ln p_i = z_i - \ln\sum_k e^{z_k}$. Differentiate with respect to $z_j$ (the derivative of the second term is $e^{z_j}/\sum_k e^{z_k} = p_j$):
$$\frac{\partial\ln p_i}{\partial z_j} = \delta_{ij} - p_j\ \Longrightarrow\ \frac{\partial p_i}{\partial z_j} = p_i\,(\delta_{ij}-p_j),\qquad J = \operatorname{diag}(\mathbf{p}) - \mathbf{p}\mathbf{p}^\top,$$where $\delta_{ij}$ is 1 if $i=j$ and 0 otherwise. For our example $J = \begin{bmatrix}0.223&-0.163&-0.060\\-0.163&0.185&-0.022\\-0.060&-0.022&0.082\end{bmatrix}$. It is symmetric, and every row sums to zero: adding the same number to all logits changes nothing, so $J\mathbf{1}=\mathbf{0}$ (and $J$ is singular). That is also why we subtract $\max_k z_k$ before exponentiating (it avoids overflow and changes nothing).
Step 2: add the cross-entropy loss. With one-hot $\mathbf{y}$, $\ell = -\sum_k y_k\ln p_k$, so $\partial\ell/\partial p_k = -y_k/p_k$. Chain rule over all $k$:
$$\frac{\partial\ell}{\partial z_j} = \sum_k\Big(-\frac{y_k}{p_k}\Big)p_k(\delta_{kj}-p_j) = -\sum_k y_k(\delta_{kj}-p_j) = -y_j + p_j\sum_k y_k = p_j - y_j,$$because $\sum_k y_k = 1$. The factor $p_k$ from the Jacobian cancels the $1/p_k$ from the log, exactly like $p(1-p)$ cancelled for the sigmoid.
$$\boxed{\ \nabla_{\mathbf{z}}\,\ell = \mathbf{p}-\mathbf{y}\ }\qquad\text{Next layer back: if } \mathbf{z}=W\mathbf{h}+\mathbf{b}:\ \ \frac{\partial\ell}{\partial W} = (\mathbf{p}-\mathbf{y})\,\mathbf{h}^\top,\ \ \frac{\partial\ell}{\partial\mathbf{h}} = W^\top(\mathbf{p}-\mathbf{y}).$$- Why so clean: $\ell = -z_c + \ln\sum_k e^{z_k}$ for true class $c$. The first part has slope $-1$ on $z_c$ only; the second is the "soft maximum" whose slope on $z_j$ is its own probability $p_j$.
- The sigmoid is the 2-class case: $\sigma(z) = \operatorname{softmax}([z, 0])_1$, and $\sigma(z)-y$ is the same formula.
- Temperature (awareness): $\operatorname{softmax}(\mathbf{z}/T)$. Large $T$ gives nearly equal probabilities, small $T$ gives nearly one-hot.
Why do we need it?
A classifier with many classes needs probabilities that are positive and sum to 1. Softmax provides them, and paired with cross-entropy it has the simplest possible gradient, $\mathbf{p}-\mathbf{y}$, so training is stable and fast.
Where is it used?
The last layer of image classifiers, the next-word prediction layer of every language model (a softmax over tens of thousands of words), and attention weights in Transformers (a softmax over positions).
How is it used?
Feed raw logits to cross_entropy (it applies a stable log-softmax inside). In the backward pass, the gradient at the logits is just softmax(logits) - one_hot(label), then it flows back through the layer by $W^\top$.
Softmax is shift-invariant but not scale-invariant. Adding 100 to every logit changes nothing; doubling every logit makes the output more confident (it is the temperature).
Do not apply softmax and then $\ln$ yourself. A probability can underflow to 0 and give $-\infty$. Use log_softmax (it computes $z_i-\ln\sum e^{z_k}$ stably) or cross_entropy on logits.
Quick check: logits $[0, 0]$ (two classes), true class 1. What are $\mathbf{p}$, the loss and the gradient?
$\mathbf{p} = [0.5, 0.5]$. Loss $= -\ln0.5 = \ln 2 \approx 0.693$. Gradient $\mathbf{p}-\mathbf{y} = [0.5-1,\ 0.5-0] = [-0.5, 0.5]$: raise logit 1, lower logit 2.
The neural-network loss: the chain rule through the layers core
What we need from earlier chapters: the chain rule and computational graphs (Chapter 2.8), backpropagation and automatic differentiation (Chapter 2.9: this section is a worked recap), the Jacobian (Chapter 2.5) and, from the Linear Algebra guide, backprop as Jacobian products.
A neural network is a chain of simple steps: multiply by a weight matrix, add a bias, bend with a simple function (like ReLU), and repeat. The loss at the end is a single number that depends on every weight in every layer at once. A big network has millions or billions of weights, so its loss is a landscape in that many dimensions.
To train it we need the slope of the loss with respect to each of those weights. The chain rule gives it: change a weight early in the network, and the change travels forward through the later layers to the loss. So the slope of the loss with respect to an early weight is a product of the local slopes along the path. Backpropagation is simply the chain rule, organised so each local slope is computed once and reused. We start from the loss and walk backwards, carrying "how much does the loss care?" with us.
Tiny network: 2 inputs, 2 hidden ReLU units, 1 sigmoid output, cross-entropy loss. Input $\mathbf{x}=[1,2]$, label $y=1$. Weights: $W_1=\begin{bmatrix}1&0.5\\-0.5&1\end{bmatrix}$, $\mathbf{b}_1=\mathbf{0}$, $\mathbf{w}_2=[0.5,\,-1]$, $b_2=0.25$.
Forward (compute and remember):
- $\mathbf{z}_1 = W_1\mathbf{x}+\mathbf{b}_1 = [1\cdot1+0.5\cdot2,\ -0.5\cdot1+1\cdot2] = [2,\ 1.5]$.
- $\mathbf{h} = \mathrm{ReLU}(\mathbf{z}_1) = [2,\ 1.5]$ (both positive, so unchanged).
- $z_{\text{out}} = \mathbf{w}_2\cdot\mathbf{h}+b_2 = 0.5\cdot2 - 1\cdot1.5 + 0.25 = -0.25$, so $p=\sigma(-0.25) = 0.4378$.
- Loss: $-\ln p = 0.826$ (the model gave only 44% to the truth).
Backward (the arrows reversed, each step is one local derivative):
- At the output: $\delta_{\text{out}} = \partial L/\partial z_{\text{out}} = p - y = -0.5622$ (the sigmoid + cross-entropy cancellation again).
- Output weights: $\partial L/\partial\mathbf{w}_2 = \delta_{\text{out}}\,\mathbf{h} = [-1.1244,\ -0.8433]$ and $\partial L/\partial b_2 = -0.5622$.
- Into the hidden layer: $\partial L/\partial\mathbf{h} = \delta_{\text{out}}\,\mathbf{w}_2 = [-0.2811,\ 0.5622]$.
- Through ReLU (slope 1 where $z>0$, else 0): $\boldsymbol{\delta}_1 = [-0.2811,\ 0.5622]\odot[1,1] = [-0.2811,\ 0.5622]$.
- First-layer weights: $\partial L/\partial W_1 = \boldsymbol{\delta}_1\mathbf{x}^\top = \begin{bmatrix}-0.2811&-0.5622\\0.5622&1.1244\end{bmatrix}$, and $\partial L/\partial\mathbf{b}_1=\boldsymbol{\delta}_1$.
- Numeric check of one entry: nudge $W_1[1,2]$ by $\pm10^{-6}$ and recompute the loss: the slope is $-0.5622$ ✓.
Network with layers $l=1,\dots,\mathcal{L}$: $\ \mathbf{a}_0=\mathbf{x}$, $\ \mathbf{z}_l = W_l\mathbf{a}_{l-1}+\mathbf{b}_l$, $\ \mathbf{a}_l = \varphi(\mathbf{z}_l)$. Training loss $L(\theta)=\frac1n\sum_i\ell(\mathbf{a}_{\mathcal{L}}(\mathbf{x}_i),\mathbf{y}_i)$, where $\theta$ is the list of all $W_l,\mathbf{b}_l$. Define the error signal $\boldsymbol{\delta}_l=\partial L/\partial\mathbf{z}_l$ (a column vector). By the chain rule with Jacobians:
$$\frac{\partial\mathbf{z}_{l+1}}{\partial\mathbf{z}_l} = W_{l+1}\,\operatorname{diag}\big(\varphi'(\mathbf{z}_l)\big)\ \Longrightarrow\ \boxed{\ \boldsymbol{\delta}_l = \varphi'(\mathbf{z}_l)\odot\big(W_{l+1}^\top\boldsymbol{\delta}_{l+1}\big)\ },\qquad \frac{\partial L}{\partial W_l}=\boldsymbol{\delta}_l\,\mathbf{a}_{l-1}^\top,\quad \frac{\partial L}{\partial\mathbf{b}_l}=\boldsymbol{\delta}_l.$$$\odot$ means entry-by-entry multiplication. The recipe: forward pass stores every $\mathbf{z}_l,\mathbf{a}_l$; backward pass starts with $\boldsymbol{\delta}_{\mathcal{L}}$ (for sigmoid or softmax + cross-entropy, simply $\mathbf{p}-\mathbf{y}$) and applies the boxed rule layer by layer. The backward pass costs about as much as the forward pass, no matter how many weights there are: that is the power of reverse mode (Chapter 2.9).
- The loss is generally not convex in $\theta$: it has valleys, ridges and saddle points (Chapter 2.10), and training finds a good valley, not a provably best one.
- Each backward step multiplies by $\varphi'$ and by $W^\top$. With $\varphi'\le0.25$ (sigmoid) the signal shrinks layer after layer (vanishing gradients); with large weights it can grow (exploding). ReLU, good initialisation and normalisation exist for this reason.
- Gradient checking: compare backprop with a finite difference, as in the numeric check above. It is the standard way to test hand-written backward code.
Why do we need it?
A network has millions of weights and we need the slope for every one, every step. Trying a nudge per weight would take millions of forward passes. The chain rule, run backwards, gets all of them in about one extra pass.
Where is it used?
Training every deep network: CNNs, Transformers, diffusion models. It is what loss.backward() does in PyTorch and what tf.GradientTape does in TensorFlow.
How is it used?
You write only the forward model and the loss. The framework records the computational graph, then walks it backwards multiplying local derivatives and fills a .grad for each weight. You use these gradients in the update step.
A weight's gradient needs two things: the error signal $\boldsymbol{\delta}$ arriving at its layer's output and the activation $\mathbf{a}_{l-1}$ that went in (the outer product $\boldsymbol{\delta}\mathbf{a}^\top$). That is why the forward pass must store its values: backward needs them.
Dead ReLU units. If $z\le0$ for every example, $\varphi'=0$, so no gradient reaches that unit's weights and it never recovers.
Quick check: in the worked example, what would $\partial L/\partial W_1[1,1]$ be if the label were $y=0$ instead?
Then $\delta_{\text{out}} = p-y = 0.4378-0 = 0.4378$, $\partial L/\partial\mathbf{h} = 0.4378\cdot[0.5,-1] = [0.2189,-0.4378]$, and $\boldsymbol\delta_1 = [0.2189,-0.4378]$ (both ReLUs still on). The entry is $\delta_1[1]\cdot x_1 = 0.2189\cdot1 = 0.2189$. The sign flipped, because now we want to make $p$ smaller.
A tiny network, trained live core
What we need from earlier chapters: the previous section's backward recipe (Chapters 2.8 and 2.9), the logistic-regression gradient, and the derivative of $\tanh$ (Chapter 2.3).
Logistic regression can only draw one straight line. Some problems need more. Picture four blobs of points where opposite corners share a class (the "XOR" pattern): no single line separates the classes.
A hidden layer fixes this. Each hidden unit $\tanh(\mathbf{w}\cdot\mathbf{x}+b)$ is a soft line: it is about $-1$ on one side and $+1$ on the other. The output unit then mixes several soft lines into a more complicated shape. Training slides all the lines around, and bends the shape, until the classes are separated. You are about to watch that happen, one gradient step at a time.
The model has $H$ hidden units. It has $2H$ first-layer weights, $H$ first-layer biases, $H$ output weights and 1 output bias: $4H+1$ parameters (13 for $H=3$). The data is 48 points. Every step does the following for the whole data set:
- Forward: for each point compute $h_j = \tanh(W_{j1}x+W_{j2}y+b_j)$, then $z=\sum_j w_{2j}h_j+b_2$ and $p=\sigma(z)$.
- Loss: the average cross-entropy over the 48 points. At the start the model is nearly clueless, so the loss is about $\ln2\approx0.69$.
- Backward: for each point $\delta=p-y$, $\ \delta_j=\delta\,w_{2j}(1-h_j^2)$. Average the gradients over the 48 points.
- Update every parameter by $-\eta\times$ its average gradient.
With $H=3$ and $\eta=0.5$, about 600 steps bring the loss from about $0.7$ down to about $0.02$.
Network: $h_j=\tanh(\mathbf{W}_{j}\cdot\mathbf{x}+b_j)$, $\ z=\sum_j w_{2j}h_j+b_2$, $\ p=\sigma(z)$, loss $\ell=\ln(1+e^z)-yz$. The gradients (the chain rule, as in the last section; $\tanh'(u)=1-\tanh^2(u)$ by the quotient rule on $\frac{e^u-e^{-u}}{e^u+e^{-u}}$):
$$\delta = p-y,\quad \frac{\partial\ell}{\partial w_{2j}}=\delta\,h_j,\quad \frac{\partial\ell}{\partial b_2}=\delta,\quad \delta_j=\delta\,w_{2j}\big(1-h_j^2\big),\quad \frac{\partial\ell}{\partial\mathbf{W}_j}=\delta_j\,\mathbf{x},\quad \frac{\partial\ell}{\partial b_j}=\delta_j.$$Training is the loop "average these over the data, subtract $\eta$ times them", repeated. The loss curve typically drops quickly, then slowly: the network first finds the coarse pattern and then fine-tunes the boundary.
Why do we need it?
Real data is rarely separated by one straight line. Hidden layers let the model build curved boundaries by combining many simple ones, and the gradient recipe still tells every weight how to move.
Where is it used?
This exact loop, at a much bigger scale, trains image classifiers, speech recognisers and language models. A Transformer has the same ingredients: linear layers, a simple non-linearity, softmax, cross-entropy and backprop.
How is it used?
Pick the number of hidden units and a learning rate, then run the loop while watching the loss curve and the decision regions. A different random start can end in a different valley, so try several.
Different start, different valley. Because the loss is not convex, two random starts can end in different places. Small networks (2 units) can get stuck; larger ones usually find a good solution more reliably.
A falling training loss is not the whole story. We have only shown the training loss. Real projects also measure loss on held-out data to check the model generalises.
Quick check: why can a network with only linear hidden units (no tanh) not solve the XOR pattern?
A linear unit followed by a linear output is still one linear function of $\mathbf{x}$, so the decision boundary is still a single straight line. The bend in $\tanh$ is what lets the units combine into curved boundaries.
Batch optimisation: full-batch, mini-batch and stochastic gradient descent core
What we need from earlier chapters: the training loss as an average over examples (the first section of this chapter), the learning rate, and the linearity of derivatives (Chapter 2.3: the gradient of an average is the average of the gradients).
The training loss is an average over all $n$ examples, so its gradient is an average too. With a million examples, one honest gradient means looking at a million examples, and then we take just one small step. That is wasteful.
Think of a poll. To learn the country's average opinion you do not ask everyone: a random sample of a few hundred gives a good estimate. Likewise we can estimate the gradient from a small random mini-batch. The estimate is a bit noisy, but it points roughly the right way, and it is hundreds of times cheaper, so we can take many more steps in the same time. Noisy steps in the right direction beat rare perfect steps.
The estimate is unbiased: a tiny check. Four examples have one-weight gradients $g=[1, 3, 5, 7]$, so the full gradient is the mean, $\frac{1+3+5+7}{4} = 4$. Draw a mini-batch of 2. There are 6 equally likely batches:
| batch | {1,3} | {1,5} | {1,7} | {3,5} | {3,7} | {5,7} |
|---|---|---|---|---|---|---|
| mini-batch gradient | 2 | 3 | 4 | 4 | 5 | 6 |
One batch can be off (2 or 6), but the average over all batches is $(2+3+4+4+5+6)/6 = 4$, exactly the full gradient. Counting steps: $n=1000$ examples, batch size $m=50$: one pass over the data (an epoch) makes $1000/50 = 20$ updates. Full-batch makes just 1 update per epoch, so ten epochs give 10 updates versus 200.
Let $\mathbf{g}_i=\nabla\ell_i(\mathbf{w})$ be the gradient from example $i$. Three flavours, all with the same update $\mathbf{w}\leftarrow\mathbf{w}-\eta\,\mathbf{g}$:
- Full-batch GD: $\mathbf{g}=\frac1n\sum_{i=1}^n\mathbf{g}_i=\nabla L$.
- Mini-batch SGD: pick a random batch $B$ of $m$ examples, $\mathbf{g}_B=\frac1m\sum_{i\in B}\mathbf{g}_i$.
- Stochastic GD (SGD) proper: $m=1$. (In practice, people say "SGD" for any mini-batch version.)
Unbiased: a randomly chosen example has expected gradient $E[\mathbf{g}_i] = \frac1n\sum_i\mathbf{g}_i=\nabla L$, and an average of $m$ such terms has the same expectation: $E[\mathbf{g}_B]=\nabla L$. Noisy: if one example's gradient has variance $\sigma^2$, the average of $m$ roughly independent ones has variance $\sigma^2/m$, so the noise (standard deviation) shrinks like $1/\sqrt m$. A batch 4 times bigger halves the noise but costs 4 times more per step.
Practicalities. Shuffle the data each epoch and walk through it in batches. Smaller batches: cheaper and more frequent steps, noisier path (the noise can help escape sharp valleys). Larger batches: smoother and better for GPUs, but fewer updates per epoch, and often they need a larger learning rate (rule of thumb: scale $\eta$ with the batch size, with a short warm-up). Near a minimum the noise keeps us bouncing around instead of settling, so we decay the learning rate late in training.
Awareness: smoothed gradients and schedules. Momentum (previous section) averages recent gradients. Adam keeps two running averages, $\mathbf{m}\leftarrow\beta_1\mathbf{m}+(1-\beta_1)\mathbf{g}$ (a smoothed gradient) and $\mathbf{v}\leftarrow\beta_2\mathbf{v}+(1-\beta_2)\mathbf{g}^2$ (a smoothed squared gradient), and steps $\eta\,\hat{\mathbf{m}}/(\sqrt{\hat{\mathbf{v}}}+\epsilon)$ (here $\hat{\mathbf{m}},\hat{\mathbf{v}}$ are $\mathbf{m},\mathbf{v}$ corrected for starting at zero, and $\epsilon$ is a tiny number that avoids dividing by 0), so each weight gets its own step size. Common learning-rate schedules: step decay, cosine decay, and linear warm-up followed by decay.
Why do we need it?
Modern datasets have millions of examples, and models need thousands of updates. Using a small random batch per step makes each update cheap while staying pointed in the right direction on average.
Where is it used?
Every deep-learning training run: batch sizes from 32 to several thousand, with SGD+momentum for vision models and Adam/AdamW with warm-up and cosine decay for Transformers and language models.
How is it used?
Use a DataLoader with batch_size and shuffle=True. Choose the batch size by memory and speed, tune the learning rate, and add a schedule. If you double the batch size, try raising the learning rate too.
"Batch" is overloaded. In "full-batch" it means all the data; in "batch size 32" it means the mini-batch. Papers and code usually mean the latter.
A noisy loss curve is normal with mini-batches. Each step's loss is measured on a different small batch. Look at a moving average, or at the loss over a whole epoch.
Quick check: $n=1200$ examples and batch size 32. How many updates in one epoch?
$1200/32 = 37.5$, so 38 updates (37 full batches plus a last smaller batch of 16). Libraries either keep the small last batch or drop it (drop_last), giving 37.
The training loop on one page: forward, loss, backward, update core
What we need from earlier chapters: every section above: the gradient of a loss, backpropagation (Chapter 2.9), the update rule and mini-batches. From the Linear Algebra guide: machine-learning applications.
All of this chapter fits in a four-beat rhythm, repeated thousands of times:
- Forward: push a batch of data through the model to get predictions.
- Loss: compare predictions with the truth and get one number.
- Backward: use the chain rule to get the slope of that number for every weight.
- Update: move every weight a little downhill.
That is exactly what a PyTorch training loop does. There is no more to it: bigger models change what happens inside each beat, never the rhythm.
Our running example: $(1,2),(2,3),(3,7)$, a line $\hat y=b+wx$, mean squared error, $\eta=0.1$, starting at $b=0,\ w=1$. Iteration 1:
- Forward: $\hat{\mathbf{y}} = X\mathbf{w} = [1, 2, 3]$.
- Loss: residual $\mathbf{r}=\hat{\mathbf{y}}-\mathbf{y} = [-1,-1,-4]$, $L=\frac13(1+1+16)=6$.
- Backward: $\nabla L=\frac23X^\top\mathbf{r}=\frac23[-6,-15]=[-4,-10]$.
- Update: $\mathbf{w}\leftarrow[0,1]-0.1[-4,-10]=[0.4,\ 2.0]$.
Iteration 2 starts from the new weights: $\hat{\mathbf{y}}=[2.4,4.4,6.4]$, $L\approx0.827$, gradient $[0.8, 0.933]$, and so on. The loss falls fast at first, then creeps along the floor of the long thin valley (the condition number is about 46, as in the 3D bowl earlier). It takes about 200 iterations at this $\eta$ to get within 0.01 of the closed-form $[-1, 2.5]$ from the earlier section, and about 1000 to match it to ten decimals.
Pseudo-code, with the PyTorch name for each beat:
for epoch in range(num_epochs): # one pass over the data
for xb, yb in loader: # one mini-batch at a time
pred = model(xb) # 1. FORWARD (build the computational graph)
loss = loss_fn(pred, yb) # 2. LOSS (one number)
optimizer.zero_grad() # clear old gradients (they ADD UP otherwise)
loss.backward() # 3. BACKWARD (chain rule fills p.grad for every weight)
optimizer.step() # 4. UPDATE (SGD: p -= lr * p.grad)
zero_grad()is needed becausebackward()adds to the stored gradients. That is the right behaviour whenever one weight is used in several places (its gradient is a sum over paths, Chapter 2.8), but it means you must clear them each step.- Swapping
optimizer(SGD, SGD with momentum, Adam) changes only beat 4. Swappingloss_fnchanges only beat 2. Adding regularisation changes beat 2 (a penalty term) or beat 4 (weight_decay). - Typical extras: a learning-rate scheduler (
scheduler.step()), evaluation on held-out data each epoch, gradient clipping beforestep().
Why do we need it?
Knowing the loop turns "the model trains itself" into four understandable calls. When training misbehaves (loss not falling, NaN, slow), you know which beat to inspect.
Where is it used?
Every PyTorch, TensorFlow and JAX training script, from a 10-line regression to a large language model. Libraries such as Hugging Face Trainer and PyTorch Lightning wrap this exact loop.
How is it used?
Write the model and the loss; keep the four beats in order. Debug by printing the loss each step, checking one gradient against a finite difference, and trying to overfit a single tiny batch first.
The same loop written by hand in plain NumPy for our linear model (no framework). Run it and compare with the widget:
import numpy as np
X = np.array([[1., 1.], [1., 2.], [1., 3.]]) # first column of ones = intercept
y = np.array([2., 3., 7.])
w = np.array([0., 1.]) # [intercept b, slope w]
eta = 0.1
for it in range(1, 4):
pred = X @ w # 1. forward
r = pred - y
loss = np.mean(r ** 2) # 2. loss
grad = 2 / len(y) * X.T @ r # 3. backward (the gradient we derived)
w = w - eta * grad # 4. update
print(it, round(loss, 4), grad.round(4), w.round(4))
# 1 6.0 [ -4. -10.] [0.4 2. ]
# 2 0.8267 [0.8 0.9333] [0.32 1.9067]
# 3 0.7525 [ 0.2667 -0.2578] [0.2933 1.9324] (then a slow creep toward [-1, 2.5])
Forgetting zero_grad() makes gradients pile up across steps: effectively a huge, wrong learning rate. Updating inside no_grad: the update itself must not be recorded in the graph (the optimiser handles that).
Overfit one batch first. A good test of the whole loop: train on a single tiny batch until the model has memorised it. The loss should go to almost zero. If it does not, the bug is in the loop, not the data.
Quick check: what goes wrong if you call loss.backward() twice without zero_grad() in between?
The second call adds its gradients to the stored ones, so p.grad becomes the sum of two gradients (twice as large if nothing changed), and the next update takes a wrongly large step.
Recap, cheat sheet and practice
- Training = going downhill on the loss. $\mathbf{w}\leftarrow\mathbf{w}-\eta\nabla L(\mathbf{w})$. The step lowers the loss because $L(\mathbf{w}-\eta\nabla L)\approx L-\eta\|\nabla L\|^2$ (linearization).
- MSE: the gradient with respect to a prediction is $\frac2n(\hat y-y)$, proportional to the error. Linear regression: $\nabla L=\frac2nX^\top(X\mathbf{w}-\mathbf{y})$; setting it to zero gives the normal equations $X^\top X\mathbf{w}=X^\top\mathbf{y}$ (same as Chapter 1.10).
- Sigmoid: $\sigma'=\sigma(1-\sigma)\le\frac14$. Logistic regression: $\nabla L=\frac1nX^\top(\sigma(X\mathbf{w})-\mathbf{y})$; the $p(1-p)$ factors cancel.
- Cross-entropy is the negative log-likelihood. It punishes confident mistakes ($-\ln p\to\infty$) and its gradient with respect to the score is $p-y$; MSE through a sigmoid has an extra factor $p(1-p)$ and saturates.
- Softmax: Jacobian $\operatorname{diag}(\mathbf{p})-\mathbf{p}\mathbf{p}^\top$; with cross-entropy the gradient at the logits is $\mathbf{p}-\mathbf{y}$.
- Learning rate: on a quadratic the error is multiplied by $1-\eta\lambda$ each step: converge if $\eta<2/\lambda_{\max}$, perfect at $1/\lambda$. Momentum averages past gradients.
- Regularisation: L2 adds $2\lambda\mathbf{w}$ to the gradient (weight decay: shrink by $1-2\eta\lambda$ each step); L1 adds $\lambda\,\mathrm{sign}(\mathbf{w})$, a constant pull with a kink, which gives exact zeros (sparsity).
- Neural networks: the loss depends on all weights; backprop is the chain rule: $\boldsymbol\delta_l=\varphi'(\mathbf{z}_l)\odot W_{l+1}^\top\boldsymbol\delta_{l+1}$, $\ \partial L/\partial W_l=\boldsymbol\delta_l\mathbf{a}_{l-1}^\top$.
- Mini-batches: the batch gradient is an unbiased, noisy estimate of the full gradient (noise $\propto1/\sqrt m$). The loop is always: forward, loss, backward (after
zero_grad), update.
Cheat sheet
| Model | Loss | Gradient at the output | Gradient for the weights |
|---|---|---|---|
| Linear regression | $\frac1n\|X\mathbf{w}-\mathbf{y}\|^2$ | $\frac2n(\hat y_i-y_i)$ | $\frac2nX^\top(X\mathbf{w}-\mathbf{y})$ |
| Ridge (L2) | $L+\lambda\|\mathbf{w}\|^2$ | unchanged | $\nabla L+2\lambda\mathbf{w}$ |
| Lasso (L1) | $L+\lambda\|\mathbf{w}\|_1$ | unchanged | $\nabla L+\lambda\,\mathrm{sign}(\mathbf{w})$ |
| Logistic regression | $-\frac1n\sum[y\ln p+(1-y)\ln(1-p)]$ | $\partial\ell/\partial z=p-y$ | $\frac1nX^\top(\sigma(X\mathbf{w})-\mathbf{y})$ |
| Softmax regression | $-\sum_ky_k\ln p_k$ | $\partial\ell/\partial\mathbf{z}=\mathbf{p}-\mathbf{y}$ | $(\mathbf{p}-\mathbf{y})\mathbf{h}^\top$ per example |
| Layer $\mathbf{a}=\varphi(W\mathbf{a}'+\mathbf{b})$ | any | $\boldsymbol\delta=\varphi'(\mathbf{z})\odot(\text{incoming gradient})$ | $\boldsymbol\delta\,\mathbf{a}'^\top$, and $W^\top\boldsymbol\delta$ goes back |
| Sigmoid | $\sigma'=\sigma(1-\sigma)$ | ||
| Softmax | $J=\operatorname{diag}(\mathbf{p})-\mathbf{p}\mathbf{p}^\top$ | ||
| Gradient descent | $\mathbf{w}\leftarrow\mathbf{w}-\eta\,\mathbf{g}$; stable if $\eta<2/\lambda_{\max}$ | ||
| Momentum | $\mathbf{v}\leftarrow\beta\mathbf{v}+\mathbf{g}$, $\ \mathbf{w}\leftarrow\mathbf{w}-\eta\mathbf{v}$ | ||
| Mini-batch | $\mathbf{g}_B=\frac1m\sum_{i\in B}\nabla\ell_i$, $E[\mathbf{g}_B]=\nabla L$ |
import numpy as np
def num_grad(f, w, h=1e-6): # finite-difference gradient, for checking
g = np.zeros_like(w)
for j in range(len(w)):
e = np.zeros_like(w); e[j] = h
g[j] = (f(w + e) - f(w - e)) / (2 * h)
return g
# ---------- 1. linear regression: gradient, closed form, gradient descent ----------
X = np.array([[1., 1.], [1., 2.], [1., 3.]])
y = np.array([2., 3., 7.])
mse = lambda w: np.mean((X @ w - y) ** 2)
grad_mse = lambda w: 2 / len(y) * X.T @ (X @ w - y)
w0 = np.array([0., 1.])
print(grad_mse(w0)) # [ -4. -10.]
print(num_grad(mse, w0).round(4)) # [ -4. -10.] (finite differences agree)
print(np.linalg.solve(X.T @ X, X.T @ y)) # [-1. 2.5] normal equations: gradient = 0
w = w0.copy()
for _ in range(3000):
w -= 0.1 * grad_mse(w)
print(w.round(3)) # [-1. 2.5] gradient descent reaches the same point
# ---------- 2. logistic regression: gradient = X^T (sigmoid(Xw) - y) / n ----------
sigmoid = lambda z: 1 / (1 + np.exp(-z))
Xc = np.array([[1., -1.], [1., 1.]])
yc = np.array([0., 1.])
def ce(w): # stable form: log(1 + e^z) - y z
z = Xc @ w
return np.mean(np.logaddexp(0, z) - yc * z)
grad_ce = lambda w: Xc.T @ (sigmoid(Xc @ w) - yc) / len(yc)
wc = np.zeros(2)
print(round(ce(wc), 4), grad_ce(wc)) # 0.6931 [ 0. -0.5] (ln 2, and the gradient we derived)
print(num_grad(ce, wc).round(4)) # [ 0. -0.5]
wc = wc - 1.0 * grad_ce(wc)
print(wc, round(ce(wc), 4)) # [0. 0.5] 0.4741
# ---------- 3. softmax + cross-entropy: gradient = p - y ----------
def softmax(z):
e = np.exp(z - z.max()) # subtract the max for stability
return e / e.sum()
z = np.array([2., 1., 0.]); yo = np.array([0., 1., 0.])
loss = lambda z: -np.log(softmax(z)[1])
p = softmax(z)
print(p.round(3), round(loss(z), 3)) # [0.665 0.245 0.09 ] 1.408
print((p - yo).round(3)) # [ 0.665 -0.755 0.09 ]
print(num_grad(loss, z).round(3)) # [ 0.665 -0.755 0.09 ]
print((np.diag(p) - np.outer(p, p)).sum(axis=1).round(10)) # [0. 0. 0.] Jacobian rows sum to zero
# ---------- 4. ridge: the gradient adds 2*lambda*w, and the closed form changes ----------
lam, n = 0.1, len(y)
ridge_w = np.linalg.solve(X.T @ X + n * lam * np.eye(2), X.T @ y)
ridge_grad = lambda w: grad_mse(w) + 2 * lam * w
print(ridge_w.round(4), ridge_grad(ridge_w).round(8)) # [-0.2145 2.118 ] [0. 0.] gradient is 0 there
print(np.linalg.norm(ridge_w[1:]) < np.linalg.norm(np.linalg.solve(X.T @ X, X.T @ y)[1:])) # True: the slope shrank
# ---------- 5. mini-batches: unbiased, but noisy ----------
g = np.array([1., 3., 5., 7.]) # per-example gradients
from itertools import combinations
means = [float(np.mean(c)) for c in combinations(g, 2)]
print(means, float(np.mean(means))) # [2.0, 3.0, 4.0, 4.0, 5.0, 6.0] 4.0
1. For $L(\mathbf{w})=\frac1n\|X\mathbf{w}-\mathbf{y}\|^2$, the gradient is…
2. In the logistic-regression gradient, the factors $\dfrac{1}{p(1-p)}$ (from the loss) and $p(1-p)$ (from the sigmoid) multiply to give…
3. Softmax gives $\mathbf{p}=[0.7, 0.2, 0.1]$ and the true class is class 1. The gradient of the cross-entropy loss with respect to the logits is…
4. A sigmoid classifier is confidently wrong ($z=-6$, true label 1). Compared with cross-entropy, the gradient of squared error with respect to $z$ is…
5. Which statement about L1 and L2 regularisation is correct?
6. Minimising $L(w)=w^2$ by gradient descent from $w_0=1$ with $\eta=1.2$. What happens?
Practice problems
A. Data $(1,3),(2,5)$, model $\hat y=b+wx$, MSE, start $b=w=0$, $\eta=0.1$. Do one gradient-descent step.
$X=\begin{bmatrix}1&1\\1&2\end{bmatrix}$, $\mathbf{y}=[3,5]$. Residual $\mathbf{r}=\hat{\mathbf{y}}-\mathbf{y}=[-3,-5]$. $X^\top\mathbf{r}=[-3-5,\ -3-10]=[-8,-13]$. Gradient $=\frac2nX^\top\mathbf{r}=\frac22[-8,-13]=[-8,-13]$. Update: $[0,0]-0.1[-8,-13]=[0.8,\ 1.3]$. The loss went from $(9+25)/2=17$ to $\hat{\mathbf{y}}=[2.1,3.4]$, $\mathbf{r}=[-0.9,-1.6]$, $L=(0.81+2.56)/2=1.685$.
B. Logistic regression, one example $\mathbf{x}=[1,2]$ (bias, feature) with $y=1$, weights $\mathbf{w}=[0,0]$, $\eta=1$. Find the gradient, the new weights and the new loss.
$z=0$, $p=0.5$. Gradient $=(p-y)\mathbf{x}=-0.5\,[1,2]=[-0.5,-1]$. New weights $[0,0]-1\cdot[-0.5,-1]=[0.5,1]$. New score $z=0.5+1\cdot2=2.5$, $p=\sigma(2.5)\approx0.924$, loss $=-\ln0.924\approx0.079$ (down from $\ln2\approx0.693$).
C. Softmax with logits $[1,1,1]$ and true class 3. Give $\mathbf{p}$, the loss and the gradient.
All logits are equal, so $\mathbf{p}=[\tfrac13,\tfrac13,\tfrac13]$. Loss $=-\ln\tfrac13=\ln3\approx1.099$. Gradient $\mathbf{p}-\mathbf{y}=[\tfrac13,\tfrac13,-\tfrac23]$ (sums to 0).
D. One weight with data-best value $a=3$. Find the minimiser of $\tfrac12(w-a)^2$ with an L2 penalty $\lambda w^2$ and with an L1 penalty $\lambda|w|$, for $\lambda=0.25$. What if $a=0.2$?
Ridge: $(w-a)+2\lambda w=0\Rightarrow w=a/(1+2\lambda)=3/1.5=2$. Lasso: $w=\mathrm{sign}(a)\max(|a|-\lambda,0)=3-0.25=2.75$. For $a=0.2$: ridge gives $0.2/1.5\approx0.133$ (small but not 0), lasso gives $\max(0.2-0.25,0)=0$ exactly, since $|a|\le\lambda$.
E. For $L(w)=4(w-2)^2$, find the range of learning rates that converge and the learning rate that lands on the minimum in one step.
$L'(w)=8(w-2)$, curvature $a=8$. The error multiplier is $1-8\eta$. It converges when $|1-8\eta|<1$, i.e. $0<\eta<2/8=0.25$. The perfect step is $\eta=1/a=0.125$: $w\leftarrow w-0.125\cdot8(w-2)=2$.
F. A data set has 50,000 examples and the batch size is 128. How many updates does one epoch make, and how many for 5 epochs? How does the noise of the batch gradient change if the batch size grows to 512?
$50000/128=390.6$, so 391 updates per epoch (the last batch is smaller), and $5\times391=1955$ updates. The gradient noise shrinks like $1/\sqrt m$; going from 128 to 512 multiplies $m$ by 4, so the noise is $\sqrt4=2$ times smaller, but each step costs 4 times as much and there are 4 times fewer updates per epoch.
Glossary
Every important word in this guide, in one place, explained in plain English. Type in the box to filter. Each entry links to the chapter that teaches it.
- Antiderivative
- A function whose derivative is the one you started with. Finding it is the reverse of differentiating, and the integral uses it. 2.13
- Asymptote
- A line that a curve gets closer and closer to but never quite reaches (for example the x-axis for $e^{-x}$). 2.2
- Automatic differentiation
- Computing exact derivatives of a program by applying the chain rule to every small step it performs. Backpropagation is its reverse mode. 2.9
- Backpropagation
- The way neural networks compute all their gradients: one forward pass, then one backward pass that applies the chain rule layer by layer. 2.9
- Batch / mini-batch
- The group of training examples used to estimate the gradient in one step. Full batch uses all of them, a mini-batch uses a small random subset. 2.14
- Chain rule
- The derivative of a function inside another function is the product of the two derivatives: $(f\circ g)'(x) = f'(g(x))\,g'(x)$. 2.3 2.8
- Computational graph
- A diagram of a calculation: each node is a simple operation, each arrow carries a value forward and a derivative backward. 2.8 2.9
- Concave / convex
- A curve that bends downward (second derivative negative) is concave. One that bends upward (second derivative positive) is convex, like a bowl. 2.10
- Continuity
- A function is continuous at a point if you can draw it through that point without lifting your pencil: the limit equals the value. 2.2
- Contour line
- A line joining all points where a function has the same value, like the height lines on a map. Also called a level set. 2.4
- Critical point
- A point where the gradient is zero. It can be a minimum, a maximum or a saddle. 2.3 2.10
- Cross-entropy
- The standard loss for classification: $-\sum y_k \log p_k$. It punishes confident wrong answers very strongly. 2.14
- Curl
- How much a vector field rotates around a point, like a tiny paddle wheel placed in flowing water. 2.13
- Curvature
- How quickly the slope itself is changing. For a function it is measured by the second derivative, for a surface by the Hessian. 2.10
- Derivative
- The instantaneous rate of change of a function: the slope of its tangent line, $f'(x) = \lim_{h\to0}\frac{f(x+h)-f(x)}{h}$. 2.3
- Differentiable
- Having a derivative at every point. A smooth curve is differentiable, a sharp corner (like $|x|$ at 0) is not. 2.3
- Directional derivative
- The slope of a surface in a chosen direction $\mathbf{u}$: $D_{\mathbf{u}}f = \nabla f\cdot\mathbf{u}$. 2.4
- Discontinuity
- A point where a function breaks: a hole (removable), a jump, or a blow-up to infinity. 2.2
- Divergence
- How much a vector field spreads out from a point (a source) or flows in (a sink): $\nabla\cdot\mathbf{F} = \sum \partial F_i/\partial x_i$. 2.13
- Domain
- The set of inputs a function accepts. 2.1
- Dual number
- A number $a + b\varepsilon$ with $\varepsilon^2 = 0$. Calculating with dual numbers gives value and derivative together, which is forward-mode differentiation. 2.9
- e (Euler's number)
- About 2.71828. The base for which $e^x$ has slope equal to its own height. 2.1 2.2
- Exploding gradient
- Gradients that grow larger and larger as they are passed back through many layers, making training unstable. 2.9
- Finite difference
- Estimating a derivative by nudging the input a tiny bit: $f'(x) \approx [f(x+h) - f(x-h)]/2h$. Used for checking gradients. 2.3 2.9
- First-order approximation
- Replacing a function by its tangent line (or tangent plane): $f(x+\delta)\approx f(x)+f'(x)\delta$. 2.11 2.12
- Forward mode
- Differentiation that carries derivatives forward with the values. One pass per input variable. 2.8 2.9
- Function
- A rule that gives exactly one output for each input. 2.1
- Gradient
- The vector of all partial derivatives, $\nabla f$. It points in the direction of steepest increase, and its length is the steepness. 2.4
- Gradient checking
- Comparing a hand-derived or backprop gradient with a finite-difference estimate to catch mistakes. 2.9
- Gradient clipping
- Shrinking a gradient whose length is too big, to keep a training step under control. 2.14
- Gradient descent
- Repeatedly stepping against the gradient: $\mathbf{w}\leftarrow\mathbf{w}-\eta\nabla L(\mathbf{w})$. The basic way models learn. 2.3 2.14
- Gradient field
- The vector field made of the gradient of a scalar field. Always perpendicular to the level lines. 2.4 2.13
- Hessian
- The matrix of all second partial derivatives. It describes the curvature of a surface. 2.10
- Integral
- The total accumulated amount, shown as the area under a curve between two points. 2.13
- Inverse function
- A function that undoes another one: $f^{-1}(f(x)) = x$. Its graph is the mirror image across the line $y=x$. 2.1
- Jacobian
- The matrix of all first partial derivatives of a vector-valued function. It is the best local linear approximation of the function. 2.5
- L1 / L2 regularization
- Adding $\lambda\|\mathbf{w}\|_1$ (L1) or $\lambda\|\mathbf{w}\|_2^2$ (L2) to the loss to keep weights small. L1 produces zeros, L2 shrinks smoothly. 2.14
- Learning curve
- A plot of the loss against training steps (or training-set size). Its slope tells you if learning has stalled. 2.3 2.14
- Learning rate
- The step size $\eta$ in gradient descent. Too small is slow, too large overshoots. 2.3 2.14
- Level set
- All the inputs where a function has the same value. In 2D it is a contour line. 2.4
- Limit
- The value a function gets closer and closer to as the input approaches a point. 2.2
- Line integral
- Adding up a field along a path, for example the work done by a force along a route. 2.13
- Linear function
- A function with a constant slope: $f(x)=mx+b$ (or a matrix times a vector, plus a constant). 2.1
- Linearization
- Approximating a smooth function near a point by its tangent line, plane or Jacobian map. 2.12
- Local derivative
- The derivative of a single small operation in a computational graph. Backpropagation multiplies local derivatives. 2.8 2.9
- Logarithm
- The inverse of an exponential: $\ln x$ answers 'what power of $e$ gives $x$?'. It turns products into sums. 2.1
- Mean squared error (MSE)
- The average of the squared differences between predictions and targets. 2.14
- Mixed partial derivative
- A second derivative taken with respect to two different variables, like $\partial^2 f/\partial x\,\partial y$. For smooth functions the order does not matter. 2.10
- Newton's method
- Using the first and second derivative to jump to where the local quadratic model is lowest (or to where the tangent crosses zero). 2.10 2.12
- Numerator layout
- The convention where the Jacobian of $\mathbf{F}:\mathbb{R}^n\to\mathbb{R}^m$ is an $m\times n$ matrix and the gradient of a scalar function is a column vector. 2.5 2.6
- Partial derivative
- The derivative with respect to one variable while all the others are held fixed: $\partial f/\partial x$. 2.4
- Piecewise function
- A function defined by different formulas on different pieces of its domain, like ReLU or $|x|$. 2.1
- Polynomial
- A sum of powers of $x$ with constant coefficients, such as $3x^2-2x+1$. 2.1
- Power rule
- $\dfrac{d}{dx}x^n = n x^{n-1}$. 2.3
- Product rule
- $(fg)' = f'g + fg'$. 2.3
- Quotient rule
- $(f/g)' = (f'g - fg')/g^2$. 2.3
- Range
- The set of outputs a function can produce. 2.1
- ReLU
- $\max(0,x)$: zero for negative inputs, the identity for positive ones. The most common activation function. 2.1 2.3
- Remainder (error term)
- What is left over when a Taylor polynomial is used instead of the true function. 2.11
- Reverse mode
- Differentiation that runs backward from the output, giving the gradient with respect to all inputs in one pass. This is backpropagation. 2.8 2.9
- Riemann sum
- Approximating the area under a curve by adding the areas of thin rectangles. 2.13
- Saddle point
- A critical point that is a minimum in one direction and a maximum in another, like a mountain pass. 2.10
- Scalar field
- A number attached to every point in space, like temperature. 2.13
- Secant line
- A line through two points of a curve. As the points merge it becomes the tangent line. 2.3
- Second derivative
- The derivative of the derivative. Positive means curving up, negative means curving down. 2.3 2.10
- Sensitivity
- How much the output changes when an input is nudged: $\Delta y\approx f'(x)\,\Delta x$. 2.3 2.12
- Sigmoid
- $\sigma(x)=1/(1+e^{-x})$: squashes any number into the range 0 to 1. Used for probabilities. 2.1 2.3
- Softmax
- Turns a list of scores into probabilities that add to 1: $p_i = e^{z_i}/\sum_j e^{z_j}$. 2.1 2.14
- Stochastic gradient descent (SGD)
- Gradient descent using a small random mini-batch per step, so each step is cheap but noisy. 2.14
- Subgradient
- A generalised slope at a corner, such as any value from −1 to 1 for $|x|$ at 0. 2.7
- Surface integral
- Adding up a field over a surface, for example the flow through it (flux). 2.13
- Tangent line
- The straight line that just touches a curve at a point and has the same slope there. 2.3
- Taylor series
- Writing a function as an infinite sum of powers built from its derivatives at one point. 2.11
- Trace trick
- Using $\operatorname{tr}(AB)=\operatorname{tr}(BA)$ and that a number equals its own trace to rearrange matrix derivatives. 2.6 2.7
- Vanishing gradient
- Gradients that shrink toward zero as they are passed back through many layers, so early layers stop learning. 2.9
- Vector field
- A vector (an arrow) attached to every point in space, like wind. 2.13
- Vector-valued function
- A function whose output is a vector, such as a neural-network layer. 2.5
- Weight decay
- Shrinking weights a little at every step. It is the gradient step of L2 regularization. 2.14
Formula cheat sheet
The formulas you will use most, on one page. If a formula does not make sense, the chapter link tells you where to relearn the idea. Remember: learn to derive them, do not just memorise them.
Derivative rules (one variable)
| Rule | Formula | Chapter |
|---|---|---|
| Definition | $f'(x)=\lim_{h\to0}\dfrac{f(x+h)-f(x)}{h}$ | 2.3 |
| Power | $(x^n)'=nx^{n-1}$ | 2.3 |
| Sum, constant multiple | $(f+g)'=f'+g'$, $(cf)'=cf'$ | 2.3 |
| Product | $(fg)'=f'g+fg'$ | 2.3 |
| Quotient | $(f/g)'=\dfrac{f'g-fg'}{g^2}$ | 2.3 |
| Chain | $(f\circ g)'(x)=f'(g(x))\,g'(x)$ | 2.3 2.8 |
| Exponential, log | $(e^x)'=e^x$, $(a^x)'=a^x\ln a$, $(\ln x)'=1/x$ | 2.3 |
| Trigonometric | $(\sin x)'=\cos x$, $(\cos x)'=-\sin x$, $(\tan x)'=1/\cos^2 x$ | 2.3 |
| Sigmoid, tanh, ReLU | $\sigma'=\sigma(1-\sigma)$, $\tanh'=1-\tanh^2$, $\mathrm{ReLU}'=\mathbb{1}[x>0]$ | 2.3 |
Gradients, Jacobians and Hessians
| Idea | Formula | Chapter |
|---|---|---|
| Gradient (column vector) | $\nabla f=\left[\dfrac{\partial f}{\partial x_1},\dots,\dfrac{\partial f}{\partial x_n}\right]^\top$ | 2.4 |
| Directional derivative | $D_{\mathbf u}f=\nabla f\cdot\mathbf u$, maximal along $\nabla f$ | 2.4 |
| Jacobian ($m\times n$) | $J_{ij}=\partial F_i/\partial x_j$ | 2.5 |
| Hessian ($n\times n$) | $H_{ij}=\partial^2 f/\partial x_i\partial x_j$ (symmetric) | 2.10 |
| Chain rule with Jacobians | $J_{f\circ g}=J_f\,J_g$; $\nabla_{\mathbf x}f(g(\mathbf x))=J_g^\top\nabla f$ | 2.5 2.8 |
| Taylor (1st, 2nd order) | $f(\mathbf x+\boldsymbol\delta)\approx f+\nabla f^\top\boldsymbol\delta+\tfrac12\boldsymbol\delta^\top H\boldsymbol\delta$ | 2.11 |
| Gradient descent, Newton | $\mathbf w\leftarrow\mathbf w-\eta\nabla L$; $\mathbf w\leftarrow\mathbf w-H^{-1}\nabla L$ | 2.10 2.14 |
Matrix-calculus identities
| Function | Gradient / derivative | Chapter |
|---|---|---|
| $\mathbf a^\top\mathbf x$ | $\mathbf a$ | 2.6 |
| $A\mathbf x$ | Jacobian $A$ | 2.6 |
| $\mathbf x^\top\mathbf x=\|\mathbf x\|^2$ | $2\mathbf x$ | 2.6 |
| $\mathbf x^\top A\mathbf x$ | $(A+A^\top)\mathbf x$ ($=2A\mathbf x$ if $A$ symmetric) | 2.6 2.7 |
| $\|A\mathbf x-\mathbf b\|^2$ | $2A^\top(A\mathbf x-\mathbf b)$ | 2.6 |
| $\|\mathbf x\|_2$, $\|\mathbf x\|_1$ | $\mathbf x/\|\mathbf x\|$, $\operatorname{sign}(\mathbf x)$ (subgradient) | 2.7 |
| $\|X\|_F^2$, $\operatorname{tr}(AX)$ | $2X$, $A^\top$ | 2.7 |
| $\log\det X$, $d(X^{-1})$ | $X^{-\top}$, $-X^{-1}\,dX\,X^{-1}$ | 2.7 |
The gradients behind machine learning
| Model / loss | Gradient | Chapter |
|---|---|---|
| Linear regression, MSE | $\dfrac{2}{n}X^\top(X\mathbf w-\mathbf y)$ | 2.14 |
| Logistic regression, cross-entropy | $\dfrac{1}{n}X^\top(\sigma(X\mathbf w)-\mathbf y)$ | 2.14 |
| Softmax + cross-entropy (logits $\mathbf z$) | $\mathbf p-\mathbf y$ | 2.14 |
| Softmax Jacobian | $\operatorname{diag}(\mathbf p)-\mathbf p\mathbf p^\top$ | 2.14 |
| L2 / L1 regularization | $2\lambda\mathbf w$, $\lambda\operatorname{sign}(\mathbf w)$ | 2.14 |
| Backprop through a dense layer ($\mathbf y=W\mathbf x+\mathbf b$) | $\partial L/\partial W=\boldsymbol\delta\,\mathbf x^\top$, $\partial L/\partial\mathbf x=W^\top\boldsymbol\delta$, $\partial L/\partial\mathbf b=\boldsymbol\delta$ | 2.9 |
Where to go next
You finished the tour. Here is how to make it stick, and where to dig deeper.
How to make it stick
- Check every gradient numerically. After you derive a gradient, nudge each input by a tiny amount and compare. This one habit catches almost every mistake.
- Derive, don't memorise. If you can derive $\nabla\|A\mathbf x-\mathbf b\|^2$ in two minutes, you will never forget it.
- Write a tiny autograd. A few dozen lines that build a graph, run forward and run backward will make backpropagation feel obvious.
- Train something. Fit a logistic regression with gradient descent using only NumPy, and watch the loss curve.
Free resources
- 3Blue1Brown, "Essence of Calculus" (YouTube): the best visual intuition for derivatives and the chain rule.
- Mathematics for Machine Learning (Deisenroth, Faisal, Ong), chapter 5 "Vector Calculus": free PDF.
- The Matrix Calculus You Need For Deep Learning (Parr & Howard): free article.
- Andrej Karpathy, "The spelled-out intro to neural networks and backpropagation" (micrograd): builds backprop from scratch.
- The Matrix Cookbook (Petersen & Pedersen): free reference of matrix-derivative identities.
- Deep Learning (Goodfellow, Bengio, Courville), chapters 4 and 6: numerical computation and backpropagation.
Next in the series
Calculus tells you which way is downhill. Optimization is the art of actually getting to the bottom, quickly and safely: gradient descent, momentum, Adam, constraints and convexity. Continue with the Optimization for Machine Learning guide.
Companion guide
Linear algebra and calculus work together in every model. Revisit the Linear Algebra guide whenever a vector or matrix feels shaky, in particular least squares, eigenvalues and the SVD.